Abstract
The self-attention mechanism has revolutionized sequence modeling but suffers from quadratic computational complexity with respect to sequence length, limiting its applicability to long sequences. We propose Sparse Projection Attention (SPA), a novel attention variant that leverages learnable sparse projections to reduce the effective dimensionality of queries and keys while maintaining expressive power. Our method is grounded in the Johnson–Lindenstrauss lemma and provides theoretical guarantees on distance preservation for fixed random projection variants. We introduce a comprehensive mathematical framework including error bounds, convergence analysis, and gradient dynamics. Experimental results on language modeling, machine translation, and long-range sequence classification demonstrate that SPA achieves up to speedup in attention score computation, and approximately end-to-end speedup, while maintaining competitive performance compared to standard attention and other efficient variants. The proposed approach offers an effective trade-off between computational efficiency and model expressivity for long-sequence tasks, making transformers more accessible for resource-constrained environments and real-time applications.
Keywords:
attention mechanism; sparse projection; computational efficiency; transformers; long sequence modeling; Johnson–Lindenstrauss lemma; randomized algorithms MSC:
68T07; 65F50; 68W40
1. Introduction
The Transformer architecture [1] has emerged as the dominant paradigm in modern deep learning, achieving state-of-the-art performance across diverse domains including natural language processing [2,3], computer vision [4,5], and speech processing [6]. At the heart of this architecture lies the self-attention mechanism, which enables the model to capture global dependencies by computing pairwise interactions between all positions in a sequence.
Despite its remarkable success, the standard self-attention mechanism exhibits quadratic computational complexity and memory requirements with respect to sequence length N, where represents the dimensionality of queries and keys. This computational bottleneck becomes particularly problematic for long sequences encountered in domains such as document understanding [7], genomic sequence analysis [8], high-resolution image processing [5], and long-term temporal forecasting [9].
The limitations of standard attention have motivated extensive research into efficient alternatives. Current approaches can be broadly categorized into:
- Pattern-based Sparse Attention [10,11];
- Low-rank Approximations [12,13];
- Locality-Sensitive Hashing [14];
- Kernel-Based Methods [15,16].
1.1. Precise Positioning Against Related Work
SPA is related to but conceptually distinct from existing projection-based efficient attention methods in three important ways.
(i) vs. Linformer [12]. Linformer projects the key and value matrices along the sequence dimension from N to a fixed rank k, achieving complexity. SPA instead projects the feature dimension of both queries and keys symmetrically, preserving the full attention structure. These are fundamentally different trade-offs: Linformer sacrifices per-token interaction fidelity; SPA reduces the inner-product space dimensionality while retaining all pairwise interactions.
(ii) vs. Performer [15]. Performer approximates via random feature maps (FAVOR+), effectively approximating the attention matrix itself at cost. SPA reduces the inner-product space in which scores are computed, provides JL-style guarantees for fixed random projections, and supports learnable sparse patterns, which Performer does not.
(iii) vs. Nyströmformer [17]. Nyströmformer approximates the attention matrix using a data-dependent Nyström decomposition with m landmark tokens. SPA operates in query- and key-feature space before score computation, is independent of token selection, and provides input-agnostic theoretical guarantees for its random projection variants.
1.2. Additional Related Work
Random Feature Attention [18] approximates the full attention map at cost but does not support learnable sparse patterns. Scatterbrain [19] combines sparse and low-rank approximations of the attention matrix; SPA is complementary as its projection operates in feature space. cosFormer [20] achieves complexity via cosine-based re-weighting. We refer the reader to the comprehensive survey by Tay et al. [21] for broader context. SPA is the only method that (a) projects both Q and K symmetrically in feature space, (b) provides JL-style guarantees for fixed projections, and (c) supports end-to-end learnable sparse patterns. A conceptual comparison is provided in Table 1.
Table 1.
Conceptual comparison of SPA with related projection-based efficient attention methods. † JL guarantees apply to fixed random projection variants (SPA-Random, SPA-Block, SPA-Banded) only.
In this paper, we introduce Sparse Projection Attention (SPA). The key contributions are:
- A mathematically grounded sparse projection mechanism reducing score computation to , with JL guarantees for fixed random projection variants;
- Comprehensive theoretical analysis including error bounds, convergence properties, and generalization bounds for the projection component;
- An adaptive sparsity learning algorithm;
- Extensive empirical evaluation demonstrating competitive performance with significant computational savings;
- Detailed ablation studies.
2. Theoretical Foundations
2.1. Johnson–Lindenstrauss Lemma and Random Projections
The Johnson–Lindenstrauss (JL) lemma [22] forms the cornerstone of our theoretical framework.
Lemma 1
(Johnson–Lindenstrauss). For any and set X of m points in , there exists a linear map with such that for all :
Theorem 1
(JL for Attention Preservation). Let . For any , there exists with such that with probability , for all :
Proof.
Step 1: Distribution of . Let have i.i.d. -sub-Gaussian entries with (e.g., ). By the quantitative JL lemma [23], for any fixed vector u and failure probability :
Step 2: From distances to inner products (polarization identity). For all , the polarization identity gives
and analogously . Set and , and apply Step 1 to each of these vectors: with probability at least ,
Combining via the polarization identity and using yields
This is the natural symmetric bound delivered by the polarization identity, and it is the form we adopt throughout.
Step 3: Union bound. Set in Step 1 applied to each of the vectors . By union bound, all inequalities hold simultaneously with probability , requiring .
Step 4: Normalization. From Step 2 (Equation (1)),
By the triangle inequality:
The first term is bounded by . For the second, the Cauchy–Schwarz inequality and Step 1 give (using AM-GM in its correct direction ), while since . Setting and absorbing constants into yields the bound . □
Remark 1
(Product form under norm normalization). Under the standard normalization enforced by LayerNorm in Transformer architectures, AM-GM gives , so the symmetric bound of Theorem 1 implies the product form
We use the symmetric form throughout the analysis since it is distribution-free; the product form is invoked only when explicitly indicated.
Remark 2
(Scope of JL-Based Guarantees). The guarantees of Theorem 1 rely on S being drawn from a sub-Gaussian distribution and remaining fixed. They apply rigorously to SPA-Random, SPA-Block, and SPA-Banded. For SPA-Learn and SPA-Adaptive, S is initialized randomly but subsequently optimized by gradient descent; after training it is deterministic and data-dependent, and no longer satisfies the JL distributional assumptions. The theoretical treatment of learnable variants is provided in Section 4.4.
2.2. Sparse Random Projections
Definition 1
(Sparse Random Matrix). A sparse random matrix has entries:
where controls the sparsity level.
Proposition 1
(Sparse JL Guarantee). For the sparse random matrix of Definition 1 with and , the distance preservation guarantee of Lemma 1 holds with high probability.
Proof.
The distance preservation bound is established in Achlioptas [24], Theorem 1.1: for m points and , choosing ensures -isometry with probability . Setting (union bound over all query-key pairs and their sums and differences as in Theorem 1) gives . The computational gain follows from sparsity: each column of S has on average nonzero entries, reducing the projection cost to . □
Corollary 1
(Sparse JL for Attention Scores). Under the conditions of Proposition 1, the attention score preservation guarantee of Theorem 1 holds for the sparse matrix S of Definition 1.
Proof.
The distance preservation of Proposition 1 implies inner-product preservation via the polarization identity (Step 2 of Theorem 1). Theorem 1 then follows by Lemma 1 and the same argument as for dense projections. □
2.3. Attention Score Preservation
Lemma 2
(Softmax Lipschitz in Frobenius Norm). Let be applied row-wise. For any :
Proof.
For each row i, the Jacobian of at is with . All eigenvalues of lie in , so . Since maps to the probability simplex , which is compact and convex, the mean-value theorem applies: for any segment there exists such that . Summing over rows: . □
Theorem 2
(Attention Score Preservation). Let and be a sparse random projection. For any and , with probability :
where is a universal constant.
Proof.
Let and . By Lemma 2: . By the triangle inequality:
Bounding the first term. From Theorem 1 with , with probability , for all : . Squaring and summing:
Hence the first term is .
Bounding the second term. By sub-multiplicativity and Proposition 1 (norm preservation): for . Since : .
Combining and identifying C. Adding both terms:
where . □
3. Sparse Projection Attention Model
3.1. Mathematical Formulation
3.1.1. Standard Attention Recap
, with bottleneck at cost .
3.1.2. Sparse Projection Layer
.
3.1.3. Sparse Projection Attention (SPA)
3.2. Sparsity Pattern Definitions
We introduce two families of sparsity patterns for the projection matrix S: fixed patterns, which preserve JL guarantees, and learnable patterns, which adapt to the task at the cost of formal theoretical guarantees.
- Fixed Sparsity Patterns: Random sparse, block sparse, banded sparse—all satisfying the JL guarantees of Theorem 1.
- Learnable Sparsity Patterns: with binary mask M and learnable weights W. The mask is optimized via Gumbel-softmax reparameterization [25] with temperature annealed from 1.0 to 0.1, regularization, or magnitude pruning. Unlike fixed projections, learnable patterns are data-driven and the JL guarantees of Theorem 1 do not apply post-training (see Section 4.4).
- Structured Sparsity Patterns: Hierarchical, conv, and head-specific patterns.
3.3. Visual Architecture Overview
Figure 1 contrasts the standard self-attention pipeline with the proposed SPA. The key architectural difference is the insertion of a learnable (or random) sparse projection S that maps queries and keys from to before the attention-score computation. The projection S is shared across all positions in a given layer, so the additional memory cost is independent of the sequence length N.
Figure 1.
Architecture comparison between standard self-attention (left) and Sparse Projection Attention (right). SPA reduces the feature dimension from to , lowering score computation from to .
3.4. Gradient Analysis and Training Dynamics
3.4.1. Gradient Flow
3.4.2. Convergence Analysis
Assumption 1
(L-smoothness). for all .
Assumption 2
(Bounded gradients). for all .
Assumption 3
(Bounded projection error). , where and .
Theorem 3
(SPA Convergence). Under Assumptions 1–3, with :
Proof. Step 1: Descent lemma.
By L-smoothness: .
Step 2: Cross-term. By Cauchy-Schwarz and Assumptions 2 and 3: .
Step 3: Quadratic term. (for ), so .
Step 4: Substitution and summation. Substituting and summing over :
Step 5: Computing the sums. With : and . Dividing by yields the stated bound. □
3.4.3. Interpretation of the Bound
The right-hand side of the bound decomposes into three contributions:
- an optimization term , which decays as with the number of iterations;
- a step-size bias , controlled by the learning rate ;
- a projection bias , controlled by the projection-error term from Assumption 3.
Convergence to a stationary point therefore requires both the step-size and the projection bias to vanish. This holds, for instance, under a decreasing learning-rate schedule combined with , since by Theorem 1 the projection error satisfies . With these two conditions, the algorithm converges to a stationary point of the true (unprojected) objective; with the constant schedule used above, the bound only guarantees convergence to a stationary region whose radius is controlled by and . This is consistent with classical results for perturbed SGD [26].
Scope note. This result is a corollary of general perturbed-SGD theory [26]. It establishes that the projection approximation error does not prevent convergence; it does not claim a convergence advantage over standard attention.
3.5. Algorithms
This subsection presents the two procedures underlying the SPA framework. Algorithm 1 formalizes the forward pass of SPA, making explicit the per-step computational cost of each operation. Algorithm 2 extends the procedure to the learnable-sparsity setting, alternating between standard gradient-based updates of the projection matrix S and structured pruning/regrowth steps that enforce the target sparsity level.
| Algorithm 1 Sparse Projection Attention (SPA) | |
| |
1: Init , sparsity | |
2: | ▹ |
3: | ▹ |
4: | ▹ |
5: | ▹ |
6: | ▹ |
7: return
O | |
| Algorithm 2 Learnable Sparse Projection Training |
|
4. Theoretical Analysis
4.1. Complexity Analysis
4.1.1. End-to-End Cost Analysis
The speedup of SPA must be understood carefully. The total forward pass cost is summarized below; a side-by-side comparison with competing methods is provided in Table 2.
compared to . The speedup factor applies only to score computation. The asymptotic end-to-end speedup is:
For , : end-to-end, vs. for score computation alone.
Table 2.
Theoretical complexity comparison. Column “JL Guarantee” indicates whether JL distance-preservation guarantees apply after training.
4.1.2. Hardware Considerations
Sparse GEMM kernels yield practical speedups on A100 GPUs only above 70–90% sparsity [27]. Reported wall-clock times in Section 6 reflect actual GPU execution and capture these hardware-level effects.
4.2. Error Bound Analysis
Theorem 4
(SPA Approximation Error). Let and . With probability :
where is a universal constant.
Proof.
By Lemma 2 and the triangle inequality:
where from Theorem 1, and the first numerator is bounded by via the entry-wise argument of Theorem 2. Combining yields the stated bound. □
4.3. Generalization Bounds
Assumption 4
(Bounded projection class). Let be the class of loss functions of transformers with SPA layers where all parameters except S are fixed. . The loss is B-bounded. The model is -Lipschitz in S:
for all inputs x.
Theorem 5
(Projection-Component Generalization Bound). For a transformer with SPA layers, treating all parameters except S as fixed, with probability :
where bounds and is the number of training samples.
Proof.
We apply Rademacher complexity theory [28]. Under Assumption 4, the covering number of satisfies . By Dudley’s entropy integral:
The effective parameter count is independent of sequence length N. Substituting into the standard generalization inequality yields the bound. □
Remark 3
(Scope of the Generalization Bound). This theorem covers only the projection component’s contribution to model complexity. The full model contains many additional parameters (, feed-forward layers) whose complexity is not reduced by SPA. The benefit of reducing is best understood as implicit regularization of the projection, not as a reduction of the full model’s VC dimension. The original bound contained a factor in the parameter count, which is not warranted for a matrix of fixed size ; this factor has been removed.
4.4. Theoretical Validity of Learnable Projections
4.4.1. Scope of JL-Based Guarantees
Theorems 1, 2 and 4 apply to fixed random projections. The convergence (Theorem 3) and generalization (Theorem 5) results are valid for all variants. For SPA-Learn and SPA-Adaptive, JL guarantees hold at initialization only.
4.4.2. Alternative Framework: Restricted Isometry Property
The RIP [29] provides a weaker condition sufficient for distance preservation by non-random matrices.
Proposition 2
(Learned Projection Quality under RIP). Let be the matrix after training. If satisfies RIP of order 2 with constant :
Learned projections perform task-specific dimensionality reduction, preserving distances most relevant to the downstream task. Extending rigorous JL-style guarantees to learned matrices remains an open theoretical problem.
5. Experimental Setup
5.1. Datasets and Tasks
- Language Modeling: Wikitext-103 [30], PG-19 [31], ArXiv [14].
- Machine Translation: WMT14 EN-DE, WMT16 EN-RO.
- LRA [32]: ListOps, Text, Retrieval, Image, Pathfinder.
5.2. Baseline Models
We compare against Standard Transformer [1], Sparse Transformer [10], Longformer [7], Linformer [12], Performer [15], and BigBird [11].
5.3. Implementation Details
All models implemented in PyTorch 2.1, trained on NVIDIA A100 (80 GB) GPUs with mixed-precision (fp16) and gradient clipping at norm 1.0.
5.3.1. Optimizer
AdamW [33], , , weight decay .
5.3.2. Learning Rate Schedule
Linear warmup for 4000 steps followed by inverse square-root decay: , with (Small), (Base), (Large).
5.3.3. Stopping Criterion
Early stopping on validation loss, patience 5. Maximum: 100 epochs (LRA), 50 (LM), 30 (MT).
5.3.4. Batch Sizes
512 tokens: 128 seq; 1024 tokens: 64 seq; 2048 tokens: 32 seq.
5.3.5. Hyperparameter Selection
by grid search over on validation sets: (LM), (MT), (LRA). Sparsity: 80% for fixed patterns; coefficient for learnable patterns.
5.3.6. Learnable Variants
SPA-Learn: Gumbel-softmax [25] with temperature annealed from 1.0 to 0.1. SPA-Adaptive: additionally uses a per-head scalar gate.
6. Results
6.1. Language Modeling
SPA-Adaptive achieves perplexity within 0.6 points of the Standard Transformer at sequence length 512 while providing speedup in score computation (≈ end-to-end). The difference at sequence length 2048 is not statistically significant (). Against other efficient baselines, SPA-Adaptive consistently outperforms by 1–3 perplexity points. Detailed quantitative results are reported in Table 3.
Table 3.
Language modeling on Wikitext-103 (Perplexity, mean ± std, 3 seeds). Bold values indicate the best score per column. †: not significant vs. Standard Transformer (, paired t-test).
A graphical comparison of performance versus computational cost across sequence lengths is shown in Figure 2.
Figure 2.
Language modeling performance and efficiency across sequence lengths. Panel (b) compares memory scaling for Standard Attention, SPA (), Longformer, and Linformer.
6.2. Machine Translation
SPA-Adaptive achieves BLEU within 0.2 of the Standard Transformer (28.2 vs. 28.4 EN→DE; ) while reducing inference time by . Against other efficient baselines, SPA-Adaptive improves BLEU by 0.4–1.3 points while also being faster. Full results are reported in Table 4.
Table 4.
Machine Translation on WMT14 EN-DE (BLEU Score, mean ± std, 3 seeds). Bold values indicate the best score per column. †: not significant vs. Standard Transformer (, paired t-test, two-tailed).
The BLEU–latency trade-off is visualized in Figure 3.
Figure 3.
Machine translation results: BLEU score and inference time comparison.
6.3. Long-Range Arena
SPA-Adaptive achieves accuracy within 0.1–0.3 percentage points of the Standard Transformer on all five LRA tasks while being faster at inference. SPA does not surpass the Standard Transformer on any individual task; the contribution is near-parity performance at significantly reduced computational cost. Compared to other efficient baselines, SPA consistently improves accuracy by 2–3 points. Per-task results are reported in Table 5 and visualized in Figure 4.
Table 5.
Long Range Arena (Accuracy %, mean ± std, 3 seeds). SPA-Adaptive is within 0.1–0.3 points of the Standard Transformer on all tasks (); the contribution is competitive performance at reduced cost, not superior performance.
Figure 4.
Performance on Long Range Arena benchmark.
6.4. Ablation Studies
6.4.1. Projection Dimensionality
Diminishing returns beyond suggest most task-relevant information is captured in a reduced subspace, as illustrated in Figure 5.
Figure 5.
Performance vs. projection dimension .
6.4.2. Effect of the Sparsity Pattern
We now study the sensitivity of SPA to the structure of the sparsity pattern at fixed projection dimension. Table 6 reports perplexity, training time, and memory for the six fixed patterns introduced in Section 4.4 as well as for the two learnable variants. The two learnable variants (Learnable and Adaptive) achieve the lowest perplexity at the cost of moderately higher training time and memory.
Table 6.
Ablation: Sparsity patterns (Wikitext-103 Perplexity). Bold value indicates the best perplexity achieved across patterns.
7. Discussion
7.1. Theoretical Implications
The JL lemma provides a strong theoretical foundation for dimensionality reduction in attention for fixed random projection variants. Sparse projections act as implicit regularizers. For learnable variants, the theoretical justification is provided by the RIP framework (Section 4.4).
7.2. Practical Implications
SPA enables transformer training on consumer hardware, faster inference, and reduced energy consumption. The practical end-to-end speedup is for typical configurations (, ), meaningful for resource-constrained settings even if below the theoretical maximum for score computation.
7.3. Limitations
(i) Theoretical gaps. JL-based guarantees apply only to fixed random projection variants. For SPA-Learn and SPA-Adaptive, formal distance-preservation guarantees do not hold post-training. The convergence result is a corollary of generic perturbed-SGD theory. The generalization bound covers only the projection component.
(ii) Memory complexity. SPA does not reduce the space complexity. For sequences above tokens, memory remains the primary bottleneck.
(iii) Practical efficiency gap. The speedup applies to score computation only; end-to-end wall-clock speedup is ≈ on A100.
(iv) Limited reproducibility data. Results are over 3 seeds. Some differences between SPA variants and baselines may not be statistically significant.
(v) Learnable variant sensitivity. SPA-Learn and SPA-Adaptive require careful tuning of and the temperature schedule. Training collapsed to degenerate patterns for in preliminary experiments.
7.4. Future Work
Future directions: extending formal guarantees to learnable projections via the RIP framework; combining SPA with methods to address memory; hardware-aware sparse projection kernels; cross-modal applications.
8. Conclusions
We presented Sparse Projection Attention (SPA), a novel attention mechanism that leverages mathematically principled sparse projections to reduce attention score computation complexity while maintaining competitive performance. SPA achieves up to speedup in score computation and ≈ end-to-end, remaining within 1 perplexity point of the Standard Transformer on language modeling, within 0.3 BLEU on machine translation, and within 0.5 accuracy points on LRA benchmarks. Key innovations include a rigorous mathematical framework with explicit scope conditions, multiple sparsity patterns with JL guarantees for fixed random variants, and learnable variants. SPA represents a practical efficiency-accuracy trade-off for long-sequence tasks. We do not claim state-of-the-art performance on any benchmark; the contribution is the efficiency-accuracy trade-off profile demonstrated across multiple tasks.
Author Contributions
Conceptualization, M.C.A., N.-E.J. and M.E.; methodology, M.C.A.; software, M.C.A.; validation, M.C.A., N.-E.J. and M.E.; formal analysis, M.C.A.; investigation, M.C.A.; resources, N.-E.J. and M.E.; data curation, M.C.A.; writing—original draft preparation, M.C.A.; writing—review and editing, N.-E.J. and M.E.; visualization, M.C.A.; supervision, M.E.; project administration, M.E.; and funding acquisition, M.E. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.
Acknowledgments
The authors thank the anonymous reviewers for their valuable comments.
Conflicts of Interest
The authors declare no conflicts of interest.
Abbreviations
| SPA | Sparse Projection Attention |
| JL | Johnson–Lindenstrauss |
| LRA | Long Range Arena |
| STE | Straight-Through Estimator |
| RIP | Restricted Isometry Property |
References
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
- Devlin, J.; Chang, M.-W.; Lee, K.; Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT; Association for Computational Linguistics: Minneapolis, MN, USA, 2019; pp. 4171–4186. [Google Scholar]
- Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I. Language models are unsupervised multitask learners. OpenAI Blog 2019, 1, 9. [Google Scholar]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16 × 16 words: Transformers for image recognition at scale. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2021. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF ICCV; IEEE: Los Alamitos, CA, USA, 2021; pp. 10012–10022. [Google Scholar]
- Gulati, A.; Qin, J.; Chiu, C.-C.; Parmar, N.; Zhang, Y.; Yu, J.; Han, W.; Wang, S.; Zhang, Z.; Wu, Y.; et al. Conformer: Convolution-augmented transformer for speech recognition. In Proceedings of Interspeech 2020; ISCA: Graz, Austria, 2020; pp. 5036–5040. [Google Scholar]
- Beltagy, I.; Peters, M.E.; Cohan, A. Longformer: The long-document transformer. arXiv 2020, arXiv:2004.05150. [Google Scholar] [CrossRef] [Scilit]
- Avsec, Ž.; Agarwal, V.; Visentin, D.; Ledsam, J.R.; Grabska-Barwinska, A.; Taylor, K.R.; Assael, Y.; Jumper, J.; Kohli, P.; Kelley, D.R. Effective gene expression prediction from sequence by integrating long-range interactions. Nat. Methods 2021, 18, 1196–1203. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhou, H.; Zhang, S.; Peng, J.; Zhang, S.; Li, J.; Xiong, H.; Zhang, W. Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of AAAI; AAAI Press: Palo Alto, CA, USA, 2021; Volume 35, pp. 11106–11115. [Google Scholar]
- Child, R.; Gray, S.; Radford, A.; Sutskever, I. Generating long sequences with sparse transformers. arXiv 2019, arXiv:1904.10509. [Google Scholar] [CrossRef] [Scilit]
- Zaheer, M.; Guruganesh, G.; Dubey, K.A.; Ainslie, J.; Alberti, C.; Ontanon, S.; Pham, P.; Ravula, A.; Wang, Q.; Yang, L.; et al. Big bird: Transformers for longer sequences. Adv. Neural Inf. Process. Syst. 2020, 33, 17283–17297. [Google Scholar]
- Wang, S.; Li, B.Z.; Khabsa, M.; Fang, H.; Ma, H. Linformer: Self-attention with linear complexity. arXiv 2020, arXiv:2006.04768. [Google Scholar] [CrossRef] [Scilit]
- Katharopoulos, A.; Vyas, A.; Pappas, N.; Fleuret, F. Transformers are RNNs: Fast autoregressive transformers with linear attention. In Proceedings of ICML; PMLR: Vienna, Austria, 2020; pp. 5156–5165. [Google Scholar]
- Kitaev, N.; Kaiser, Ł.; Levskaya, A. Reformer: The efficient transformer. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2020. [Google Scholar]
- Choromanski, K.; Likhosherstov, V.; Dohan, D.; Song, X.; Gane, A.; Sarlos, T.; Hawkins, P.; Davis, J.; Mohiuddin, A.; Kaiser, L.; et al. Rethinking attention with performers. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2021. [Google Scholar]
- Lee-Thorp, J.; Ainslie, J.; Eckstein, I.; Ontanon, S. FNet: Mixing tokens with Fourier transforms. In Proceedings of NAACL-HLT; Association for Computational Linguistics: Seattle, WA, USA, 2022; pp. 4296–4313. [Google Scholar]
- Xiong, Y.; Zeng, Z.; Chakraborty, R.; Tan, M.; Fung, G.; Li, Y.; Singh, V. Nyströmformer: A Nyström-based algorithm for approximating self-attention. In Proceedings of AAAI; AAAI Press: Palo Alto, CA, USA, 2021; Volume 35, pp. 14138–14148. [Google Scholar]
- Peng, H.; Pappas, N.; Yogatama, D.; Schwartz, R.; Smith, N.A.; Kong, L. Random feature attention. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2021. [Google Scholar]
- Chen, B.; Dao, T.; Winsor, E.; Song, Z.; Rudra, A.; Ré, C. Scatterbrain: Unifying sparse and low-rank attention approximation. Adv. Neural Inf. Process. Syst. 2021, 34, 17413–17426. [Google Scholar]
- Qin, Z.; Sun, W.; Deng, H.; Li, D.; Wei, Y.; Lv, B.; Yan, J.; Kong, L.; Zhong, Y. cosFormer: Rethinking softmax in attention. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2022. [Google Scholar]
- Tay, Y.; Dehghani, M.; Bahri, D.; Metzler, D. Efficient transformers: A survey. ACM Comput. Surv. 2022, 55, 109. [Google Scholar] [CrossRef] [Scilit]
- Johnson, W.B.; Lindenstrauss, J. Extensions of Lipschitz mappings into a Hilbert space. Contemp. Math. 1984, 26, 189–206. [Google Scholar]
- Dasgupta, S.; Gupta, A. An elementary proof of a theorem of Johnson and Lindenstrauss. Random Struct. Algorithms 2003, 22, 60–65. [Google Scholar] [CrossRef] [Scilit]
- Achlioptas, D. Database-friendly random projections: Johnson–Lindenstrauss with binary coins. J. Comput. Syst. Sci. 2003, 66, 671–687. [Google Scholar] [CrossRef] [Scilit]
- Jang, E.; Gu, S.; Poole, B. Categorical reparameterization with Gumbel-softmax. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2017. [Google Scholar]
- Bottou, L.; Curtis, F.E.; Nocedal, J. Optimization methods for large-scale machine learning. SIAM Rev. 2018, 60, 223–311. [Google Scholar] [CrossRef] [Scilit]
- Gale, T.; Elsen, E.; Hooker, S. The state of sparsity in deep neural networks. arXiv 2020, arXiv:1902.09574. [Google Scholar]
- Bartlett, P.L.; Mendelson, S. Rademacher and Gaussian complexities: Risk bounds and structural results. J. Mach. Learn. Res. 2002, 3, 463–482. [Google Scholar]
- Candès, E.J.; Romberg, J.; Tao, T. Stable signal recovery from incomplete and inaccurate measurements. Commun. Pure Appl. Math. 2006, 59, 1207–1223. [Google Scholar] [CrossRef] [Scilit]
- Merity, S.; Keskar, N.S.; Socher, R. Pointer sentinel mixture models. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2017. [Google Scholar]
- Rae, J.W.; Potapenko, A.; Jayakumar, S.M.; Hillier, C.; Lillicrap, T.P. Compressive transformers for long-range sequence modelling. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2020. [Google Scholar]
- Tay, Y.; Dehghani, M.; Abnar, S.; Shen, Y.; Bahri, D.; Pham, P.; Rao, J.; Yang, L.; Ruder, S.; Metzler, D. Long range arena: A benchmark for efficient transformers. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2021. [Google Scholar]
- Loshchilov, I.; Hutter, F. Decoupled weight decay regularization. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2019; Available online: https://openreview.net/forum?id=Bkg6RiCqY7 (accessed on 10 May 2026).
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.




