iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
FCC Bans Foreign-Made Robot Vacuums Over National Security Risks Apple Warns of 'Significant' Supply Constraints for Mac, iPhone, and iPad India seeks to cut reliance on imported strawberry varieties with indigenous breeding Burnham Confirms Pragmatic North Sea Oil Stance in Trump Call, Fueling Drilling Debate Leaked Memo Links Iranian Hackers to Minnesota Water Utility Cyberattacks Everyone Is Freaking Out About OpenAI and Anthropic’s Race for Dominance Govt Debunks AI-Generated Fake Video of Finance Minister Nirmala Sitharaman Promoting Investment Scheme UPS Unveils Digital Tools to Attract Small Businesses Amid Strategic Shift from Low-Margin E-Commerce CPKC sets second-quarter revenue record as operating income rises 10% Your Freight Funnel Is Leaking Margin: What Your Reports Won't Show FCC Bans Foreign-Made Robot Vacuums Over National Security Risks Apple Warns of 'Significant' Supply Constraints for Mac, iPhone, and iPad India seeks to cut reliance on imported strawberry varieties with indigenous breeding Burnham Confirms Pragmatic North Sea Oil Stance in Trump Call, Fueling Drilling Debate Leaked Memo Links Iranian Hackers to Minnesota Water Utility Cyberattacks Everyone Is Freaking Out About OpenAI and Anthropic’s Race for Dominance Govt Debunks AI-Generated Fake Video of Finance Minister Nirmala Sitharaman Promoting Investment Scheme UPS Unveils Digital Tools to Attract Small Businesses Amid Strategic Shift from Low-Margin E-Commerce CPKC sets second-quarter revenue record as operating income rises 10% Your Freight Funnel Is Leaking Margin: What Your Reports Won't Show
Home ›› Technology ›› Ai ›› Llms ›› Transformer Feed-Forward Block Linearity: Learned, Not Architectural, According to New Research

Transformer Feed-Forward Block Linearity: Learned, Not Architectural, According to New Research

A new study introduces R^2_lin, a measure of linearity for transformer feed-forward blocks. Across models like GPT-2 and Pythia-160m, R^2_lin varies widely and is not determined by activation function. The findings offer targeted compression signals and reveal pitfalls in training linear baselines.

iG
iGEN Editorial
June 20, 2026
Transformer Feed-Forward Block Linearity: Learned, Not Architectural, According to New Research

A research paper from arXiv presents a novel method for measuring exactly how linear each block in a transformer feed-forward network (FFN) actually is. The study, titled "How Linear Is a Transformer Feed-Forward Block? Per-Block Linear Recoverability Is Learned, Not Architectural," challenges the common assumption that FFN blocks serve uniformly as nonlinear stores of computation. Instead, the authors find that linearity is a learned property of individual trained blocks and varies dramatically across layers.

Measuring Per-Block Linearity

The authors treat each FFN block as a position-wise input-to-output map and decompose it into an exact least-squares linear approximation plus a residual. The key metric introduced is R^2_lin, which measures the held-out variance explained by the closed-form linear map. This provides an optimiser-free, per-block assessment of linearity. The study applies this method to three transformer models: GPT-2, Pythia-160m, and llama-160m, spanning all twelve blocks of each.

Heterogeneous and Learned Behavior

The results show that R^2_lin is highly heterogeneous and non-monotone with depth, ranging from near-linear (>0.99) to strongly nonlinear (<0.3) between adjacent blocks. Critically, this variation is not set by the activation function. For example, both GPT-2 and Pythia-160m use GELU activations and have the same width, yet their R^2_lin profiles are sharply different. The authors conclude that recoverability is a learned property of individual trained blocks, not an architectural one.

Model R^2_lin Range (across 12 blocks) Activation Function
GPT-2 <0.3 to >0.99 (non-monotone) GELU
Pythia-160m <0.3 to >0.99 (different profile) GELU
llama-160m <0.3 to >0.99 (heterogeneous) (assumed similar)

The paper also explores the nature of the residual—the part not captured by the linear approximation. A low-rank bilinear probe of the residual recovers only a few additional points of R^2, with gain uncorrelated with residual nonlinearity. This indicates that the unrecovered computation is not a single position-wise product but likely higher-order or distributed structure.

Implications for Model Compression

The R^2_lin measurement serves as a targeted compression signal. Recoverable blocks (those with high R^2_lin) admit large single-layer replacements. For instance, in GPT-2's early FFN, a replacement with 8x fewer parameters results in only +0.77 perplexity increase. Conversely, low-recoverability blocks flag where such compression would be unsafe. This insight could enable more efficient deployment of transformer models in resource-constrained environments.

The study also exposes a methodological pitfall: trained linear baselines can badly under-converge on ill-conditioned transformer activations. The authors address this by reporting the exact closed-form least-squares ceiling throughout, ensuring reliable comparisons.

For enterprise technology decision-makers, this research provides a principled way to identify which layers of a transformer model can be replaced with simpler linear operations without significant performance loss. Such compression could reduce computational costs and memory footprint, making AI more feasible for edge and real-time applications in supply chain and logistics, such as document processing or demand forecasting.


Sources:

Keep Reading

Recommended Stories

Yann LeCun's new AI startup AMI Labs raises $1bn to build flexible intelligence beyond LLMs Technology

Yann LeCun's new AI startup AMI Labs raises $1bn to build flexible intelligence beyond LLMs

Yann LeCun, former Meta chief AI scientist, has founded AMI Labs to develop a new AI architecture called JEPA, which aims to overcome the limitations of large language models (LLMs) in understanding the physical world. The startup raised over $1bn in seed funding from Nvidia and Jeff Bezos' private investment fund, marking one of Europe's largest seed rounds.

July 2, 2026
New Unified Definition of AI Hallucination Pins It on Inaccurate World Modeling Technology

New Unified Definition of AI Hallucination Pins It on Inaccurate World Modeling

A new arXiv paper by Liu et al. proposes a unified definition of hallucination in large language models, defining it as inaccurate internal world modeling observable to the user. The framework subsumes prior definitions and distinguishes true hallucinations from planning or reward errors, and introduces the HalluWorld benchmark for stress-testing models.

June 16, 2026
Z-Plane Neural Networks Replace ReLU and LayerNorm with Bounded Geometric Activation Technology

Z-Plane Neural Networks Replace ReLU and LayerNorm with Bounded Geometric Activation

Researchers propose Z-Plane Neural Networks, which replace traditional ReLU activations and LayerNorm with a bounded geometric activation called Radial Bounding. This new approach maintains 1-Lipschitz continuity, prevents gradient vanishing, and preserves directional information. A 100-layer Z-Plane MLP achieved 98.34% accuracy on MNIST without any ReLU or LayerNorm, demonstrating numerical stability.

June 16, 2026
New Drift-RAE Method Distills Transformers Efficiently Using Representation Autoencoders Technology

New Drift-RAE Method Distills Transformers Efficiently Using Representation Autoencoders

A new research paper proposes Drift-RAE, a method for distilling pretrained flow models in representation autoencoder latent spaces. It overcomes anisotropy and large curvature challenges, achieving 1.77 FID on ImageNet 256 with only 10,000 distillation steps, outperforming existing RAE distillation methods.

June 16, 2026