Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

Why "Classic" Transformers Are Shallow and A Depth-Enabling Technique

Дата публикации: 17-08-2026 20:26:00


Since its introduction in 2017, the Transformer has emerged as the leading neural network architecture, catalyzing revolutionary advancements in many AI disciplines. The key innovation in Transformer is a Self-Attention (SA) mechanism designed to capture contextual information. However, stacking up more layers of the same design has failed to produce trainable deeper Transformers. Thus far, various architectural modifications to the original design have been proposed to enable deeper depths for Transformer models, but a thorough understanding of this depth issue remains lacking. In this paper, we conduct a comprehensive investigation to substantiate the claim that the depth problem is caused by a phenomenon called token similarity escalation; that is, tokens grow increasingly alike after repeated applications of the SA mechanism. Our analysis reveals that, driven by the invariant leading eigenspace and large spectral gaps of attention matrices, token similarity provably escalates at a linear rate as the depth increases. This insight suggests a simple technique that surgically removes excessive token similarity without reducing the overall role of the SA mechanism, as is done by existing approaches. We perform a set of proof- of-concept, small-scale experiments to show the viability of the proposed depth-enabling technique.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1 Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation Perspective 04.8817-08-2026
2 Limiting Over-Smoothing and Over-Squashing of Graph Message Passing by Deep Scattering Transforms 010.8717-08-2026
3 Statistical Test for Attention in Transformers for Images and Time Series 09.1817-08-2026
4 Beyond Unconstrained Features: Neural Collapse for Shallow Neural Networks with General Data 04.3617-08-2026
5 Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks 0517-08-2026
6 Enhancing Accuracy in Generative Models via Knowledge Transfer 06.6617-08-2026
7 High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks 08.717-08-2026
8 Nested Subspace Learning with Flags 0717-08-2026
9 Gradient Span Algorithms Make Predictable Progress in High Dimension 06.3817-08-2026
10DeepSeek-V3 Model: Theory, Config, and Rotary Positional Embeddings022.512-03-2026

Классификация: Пресс-релизы. Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 8.42. Источник: jmlr.org.