Skip to Content
MathematicsMathematics
  • Article
  • Open Access

23 May 2026

16 Pages

Sparse Projection Attention: A Computationally Efficient Framework for Long Sequence Modeling

,
and
Faculty of Sciences and Techniques, University Sidi Mohammed Ben Abdellah, Fez 30050, Morocco
*
Author to whom correspondence should be addressed.

Abstract

The self-attention mechanism has revolutionized sequence modeling but suffers from quadratic computational complexity with respect to sequence length, limiting its applicability to long sequences. We propose Sparse Projection Attention (SPA), a novel attention variant that leverages learnable sparse projections to reduce the effective dimensionality of queries and keys while maintaining expressive power. Our method is grounded in the Johnson–Lindenstrauss lemma and provides theoretical guarantees on distance preservation for fixed random projection variants. We introduce a comprehensive mathematical framework including error bounds, convergence analysis, and gradient dynamics. Experimental results on language modeling, machine translation, and long-range sequence classification demonstrate that SPA achieves up to 8 × speedup in attention score computation, and approximately 2 × end-to-end speedup, while maintaining competitive performance compared to standard attention and other efficient variants. The proposed approach offers an effective trade-off between computational efficiency and model expressivity for long-sequence tasks, making transformers more accessible for resource-constrained environments and real-time applications.

1. Introduction

The Transformer architecture [1] has emerged as the dominant paradigm in modern deep learning, achieving state-of-the-art performance across diverse domains including natural language processing [2,3], computer vision [4,5], and speech processing [6]. At the heart of this architecture lies the self-attention mechanism, which enables the model to capture global dependencies by computing pairwise interactions between all positions in a sequence.
Despite its remarkable success, the standard self-attention mechanism exhibits quadratic computational complexity O ( N 2 d k ) and memory requirements O ( N 2 ) with respect to sequence length N, where d k represents the dimensionality of queries and keys. This computational bottleneck becomes particularly problematic for long sequences encountered in domains such as document understanding [7], genomic sequence analysis [8], high-resolution image processing [5], and long-term temporal forecasting [9].
The limitations of standard attention have motivated extensive research into efficient alternatives. Current approaches can be broadly categorized into:
  • Pattern-based Sparse Attention [10,11];
  • Low-rank Approximations [12,13];
  • Locality-Sensitive Hashing [14];
  • Kernel-Based Methods [15,16].

1.1. Precise Positioning Against Related Work

SPA is related to but conceptually distinct from existing projection-based efficient attention methods in three important ways.
(i) vs. Linformer [12]. Linformer projects the key and value matrices along the sequence dimension from N to a fixed rank k, achieving O ( N k ) complexity. SPA instead projects the feature dimension  d k of both queries and keys symmetrically, preserving the full N × N attention structure. These are fundamentally different trade-offs: Linformer sacrifices per-token interaction fidelity; SPA reduces the inner-product space dimensionality while retaining all pairwise interactions.
(ii) vs. Performer [15]. Performer approximates softmax ( Q K T ) via random feature maps (FAVOR+), effectively approximating the attention matrix itself at O ( N ) cost. SPA reduces the inner-product space in which scores are computed, provides JL-style guarantees for fixed random projections, and supports learnable sparse patterns, which Performer does not.
(iii) vs. Nyströmformer [17]. Nyströmformer approximates the attention matrix using a data-dependent Nyström decomposition with m landmark tokens. SPA operates in query- and key-feature space before score computation, is independent of token selection, and provides input-agnostic theoretical guarantees for its random projection variants.

1.2. Additional Related Work

Random Feature Attention [18] approximates the full attention map at O ( N ) cost but does not support learnable sparse patterns. Scatterbrain [19] combines sparse and low-rank approximations of the attention matrix; SPA is complementary as its projection operates in feature space. cosFormer [20] achieves O ( N ) complexity via cosine-based re-weighting. We refer the reader to the comprehensive survey by Tay et al. [21] for broader context. SPA is the only method that (a) projects both Q and K symmetrically in feature space, (b) provides JL-style guarantees for fixed projections, and (c) supports end-to-end learnable sparse patterns. A conceptual comparison is provided in Table 1.
Table 1. Conceptual comparison of SPA with related projection-based efficient attention methods. † JL guarantees apply to fixed random projection variants (SPA-Random, SPA-Block, SPA-Banded) only.
In this paper, we introduce Sparse Projection Attention (SPA). The key contributions are:
  • A mathematically grounded sparse projection mechanism reducing score computation to O ( N 2 d k ′ ) , with JL guarantees for fixed random projection variants;
  • Comprehensive theoretical analysis including error bounds, convergence properties, and generalization bounds for the projection component;
  • An adaptive sparsity learning algorithm;
  • Extensive empirical evaluation demonstrating competitive performance with significant computational savings;
  • Detailed ablation studies.

2. Theoretical Foundations

2.1. Johnson–Lindenstrauss Lemma and Random Projections

The Johnson–Lindenstrauss (JL) lemma [22] forms the cornerstone of our theoretical framework.
Lemma 1
(Johnson–Lindenstrauss). For any 0 < ϵ < 1 and set X of m points in R d , there exists a linear map f : R d → R k with k = O ( ϵ − 2 log m ) such that for all u , v ∈ X :
( 1 − ϵ ) ∥ u − v ∥ 2 ≤ ∥ f ( u ) − f ( v ) ∥ 2 ≤ ( 1 + ϵ ) ∥ u − v ∥ 2 .
Theorem 1
(JL for Attention Preservation). Let Q , K ∈ R N × d k . For any ϵ , δ > 0 , there exists S ∈ R d k × d k ′ with d k ′ = O ( ϵ − 2 log ( N / δ ) ) such that with probability ≥ 1 − δ , for all i , j ∈ [ N ] :
⟨ q i , k j ⟩ d k − ⟨ q i S , k j S ⟩ d k ′ ≤ ϵ 2 ∥ q i ∥ 2 + ∥ k j ∥ 2 .
Proof. 
Step 1: Distribution of S . Let S ∈ R d k × d k ′ have i.i.d. σ -sub-Gaussian entries with σ 2 = 1 / d k ′ (e.g., S i j ∼ N ( 0 , 1 / d k ′ ) ). By the quantitative JL lemma [23], for any fixed vector u and failure probability η ∈ ( 0 , 1 ) :
P ( 1 − ϵ ′ ) ∥ u ∥ 2 ≤ ∥ u S ∥ 2 ≤ ( 1 + ϵ ′ ) ∥ u ∥ 2 ≥ 1 − η .
Step 2: From distances to inner products (polarization identity). For all i , j ∈ [ N ] , the polarization identity gives
⟨ q i , k j ⟩ = 1 4 ∥ q i + k j ∥ 2 − ∥ q i − k j ∥ 2 ,
and analogously ⟨ q i S , k j S ⟩ = 1 4 ( ∥ ( q i + k j ) S ∥ 2 − ∥ ( q i − k j ) S ∥ 2 ) . Set u + = q i + k j and u − = q i − k j , and apply Step 1 to each of these vectors: with probability at least 1 − 2 η ,
∥ u ± S ∥ 2 − ∥ u ± ∥ 2 ≤ ϵ ′ ∥ u ± ∥ 2 .
Combining via the polarization identity and using ∥ u + ∥ 2 + ∥ u − ∥ 2 = 2 ( ∥ q i ∥ 2 + ∥ k j ∥ 2 ) yields
⟨ q i S , k j S ⟩ − ⟨ q i , k j ⟩ ≤ ϵ ′ 4 ∥ u + ∥ 2 + ∥ u − ∥ 2 = ϵ ′ 2 ∥ q i ∥ 2 + ∥ k j ∥ 2 .
This is the natural symmetric bound delivered by the polarization identity, and it is the form we adopt throughout.
Step 3: Union bound. Set η i j = δ / ( 2 N 2 ) in Step 1 applied to each of the 2 N 2 vectors { u i j + , u i j − } i , j ∈ [ N ] . By union bound, all inequalities hold simultaneously with probability ≥ 1 − δ , requiring d k ′ = O ( ϵ − 2 log ( N / δ ) ) .
Step 4: Normalization. From Step 2 (Equation (1)),
| ⟨ q i S , k j S ⟩ − ⟨ q i , k j ⟩ | ≤ ϵ ′ 2 ∥ q i ∥ 2 + ∥ k j ∥ 2 .
By the triangle inequality:
⟨ q i , k j ⟩ d k − ⟨ q i S , k j S ⟩ d k ′ ≤ | ⟨ q i , k j ⟩ − ⟨ q i S , k j S ⟩ | d k + | ⟨ q i S , k j S ⟩ | · 1 d k − 1 d k ′ .
The first term is bounded by ϵ ′ 2 d k ( ∥ q i ∥ 2 +   ∥ k j ∥ 2 )   ≤ ϵ ′ 2 ( ∥ q i ∥ 2 +   ∥ k j ∥ 2 ) . For the second, the Cauchy–Schwarz inequality and Step 1 give | ⟨ q i S , k j S ⟩ |   ≤   ∥ q i S ∥ ∥ k j S ∥   ≤   ( 1 + ϵ ′ ) ∥ q i ∥ ∥ k j ∥   ≤   1 + ϵ ′ 2 ( ∥ q i ∥ 2 +   ∥ k j ∥ 2 ) (using AM-GM in its correct direction ∥ q i ∥ ∥ k j ∥   ≤   1 2 ( ∥ q i ∥ 2 +   ∥ k j ∥ 2 ) ), while 1 / d k − 1 / d k ′ ≤ 1 / d k ′ since d k ′ ≤ d k . Setting ϵ ′ = ϵ / ( 2 + 2 ϵ ) ≈ ϵ / 2 and absorbing constants into O ( d k ′ ) yields the bound ϵ 2 ( ∥ q i ∥ 2 +   ∥ k j ∥ 2 ) .    □
Remark 1
(Product form under norm normalization). Under the standard normalization ∥ q i ∥ , ∥ k j ∥ ≤ 1 enforced by LayerNorm in Transformer architectures, AM-GM gives ∥ q i ∥ ∥ k j ∥   ≤ 1 2 ( ∥ q i ∥ 2 +   ∥ k j ∥ 2 ) ≤ 1 , so the symmetric bound of Theorem 1 implies the product form
⟨ q i , k j ⟩ d k − ⟨ q i S , k j S ⟩ d k ′ ≤ ϵ .
We use the symmetric form throughout the analysis since it is distribution-free; the product form is invoked only when explicitly indicated.
Remark 2
(Scope of JL-Based Guarantees). The guarantees of Theorem 1 rely on S being drawn from a sub-Gaussian distribution and remaining fixed. They apply rigorously to SPA-Random, SPA-Block, and SPA-Banded. For SPA-Learn and SPA-Adaptive, S is initialized randomly but subsequently optimized by gradient descent; after training it is deterministic and data-dependent, and no longer satisfies the JL distributional assumptions. The theoretical treatment of learnable variants is provided in Section 4.4.

2.2. Sparse Random Projections

Definition 1
(Sparse Random Matrix). A sparse random matrix S ∈ R d k × d k ′ has entries:
S i j = s d k ′ × + 1 with probability 1 2 s 0 with probability 1 − 1 s − 1 with probability 1 2 s ,
where s ≥ 1 controls the sparsity level.
Proposition 1
(Sparse JL Guarantee). For the sparse random matrix of Definition 1 with s = O ( d k ) and d k ′ = O ( ϵ − 2 log N ) , the distance preservation guarantee of Lemma 1 holds with high probability.
Proof. 
The distance preservation bound is established in Achlioptas [24], Theorem 1.1: for m points and ϵ ∈ ( 0 , 1 ) , choosing d k ′ ≥ C ϵ − 2 log m ensures ( 1 ± ϵ ) -isometry with probability ≥ 1 − m − β . Setting m = 2 N 2 (union bound over all query-key pairs and their sums and differences as in Theorem 1) gives d k ′ = O ( ϵ − 2 log N ) . The computational gain follows from sparsity: each column of S has on average d k / s nonzero entries, reducing the projection cost to O ( N d k d k ′ / s ) .    □
Corollary 1
(Sparse JL for Attention Scores). Under the conditions of Proposition 1, the attention score preservation guarantee of Theorem 1 holds for the sparse matrix S of Definition 1.
Proof. 
The distance preservation of Proposition 1 implies inner-product preservation via the polarization identity (Step 2 of Theorem 1). Theorem 1 then follows by Lemma 1 and the same argument as for dense projections.    □

2.3. Attention Score Preservation

Lemma 2
(Softmax Lipschitz in Frobenius Norm). Let softmax : R N × N → R N × N be applied row-wise. For any A , B ∈ R N × N :
∥ softmax ( A ) − softmax ( B ) ∥ F ≤ ∥ A − B ∥ F .
Proof. 
For each row i, the Jacobian of softmax at a ∈ R N is J a = diag ( p ) − p p T with p = softmax ( a ) . All eigenvalues of J a lie in [ 0 , 1 ] , so ∥ J a ∥ 2 ≤ 1 . Since softmax maps R N to the probability simplex Δ N − 1 , which is compact and convex, the mean-value theorem applies: for any segment [ a , b ] there exists c ∈ [ a , b ] such that ∥ softmax ( a ) − softmax ( b ) ∥ 2 =   ∥ J c ( a − b ) ∥ 2 ≤   ∥ J c ∥ 2 ∥ a − b ∥ 2 ≤ ∥ a − b ∥ 2 . Summing over rows: ∥ softmax ( A ) − softmax ( B ) ∥ F 2 ≤ ∑ i ∥ A i − B i ∥ 2 2   = ∥ A − B ∥ F 2 .    □
Theorem 2
(Attention Score Preservation). Let Q , K ∈ R N × d k and S ∈ R d k × d k ′ be a sparse random projection. For any ϵ > 0 and δ ∈ ( 0 , 1 ) , with probability ≥ 1 − δ :
softmax Q K T d k − softmax Q S S T K T d k ′ F ≤ C ∥ Q ∥ F ∥ K ∥ F log ( N / δ ) d k ′ 1 d k + 1 d k ′ ,
where C > 0 is a universal constant.
Proof. 
Let A raw = Q K T / d k and A ˜ raw = Q S S T K T / d k ′ . By Lemma 2: ∥ softmax ( A raw ) − softmax ( A ˜ raw ) ∥ F   ≤   ∥ A raw − A ˜ raw ∥ F . By the triangle inequality:
∥ A raw − A ˜ raw ∥ F   ≤   ∥ Q K T − Q S S T K T ∥ F d k + ∥ Q S S T K T ∥ F · 1 d k − 1 d k ′ .
Bounding the first term. From Theorem 1 with ϵ 0 = C ′ log ( N / δ ) / d k ′ , with probability ≥ 1 − δ , for all i , j : | ( Q K T ) i j − ( Q S S T K T ) i j |   ≤   ϵ 0 ∥ q i ∥ ∥ k j ∥ . Squaring and summing:
∥ Q K T − Q S S T K T ∥ F 2   =   ∑ i , j | ( Q K T ) i j − ( Q S S T K T ) i j | 2 ≤ ϵ 0 2 ∑ i , j ∥ q i ∥ 2 ∥ k j ∥ 2   =   ϵ 0 2 ∥ Q ∥ F 2 ∥ K ∥ F 2 .
Hence the first term is ≤ ϵ 0 ∥ Q ∥ F ∥ K ∥ F / d k .
Bounding the second term. By sub-multiplicativity and Proposition 1 (norm preservation): ∥ Q S S T K T ∥ F   ≤ ∥ Q S ∥ F ∥ K S ∥ F ≤ ( 1 + ϵ 0 ) 2 ∥ Q ∥ F ∥ K ∥ F ≤ 4 ∥ Q ∥ F ∥ K ∥ F for ϵ 0 ≤ 1 . Since d k ′ ≤ d k : | 1 / d k − 1 / d k ′ |   ≤ 1 / d k ′ .
Combining and identifying C. Adding both terms:
∥ softmax ( A raw ) − softmax ( A ˜ raw ) ∥ F ≤   ϵ 0 ∥ Q ∥ F ∥ K ∥ F d k + 4 ∥ Q ∥ F ∥ K ∥ F d k ′ ≤   C ∥ Q ∥ F ∥ K ∥ F log ( N / δ ) d k ′ 1 d k + 1 d k ′ ,
where C = max ( C ′ , 4 / log ( N / δ ) ) .    □

3. Sparse Projection Attention Model

3.1. Mathematical Formulation

3.1.1. Standard Attention Recap

Attention ( Q , K , V ) = softmax ( Q K T / d k ) V , with bottleneck Q K T at cost O ( N 2 d k ) .

3.1.2. Sparse Projection Layer

Q ′ = Q S , K ′ = K S , Q ′ , K ′ ∈ R N × d k ′ .

3.1.3. Sparse Projection Attention (SPA)

SPA ( Q , K , V ) = softmax Q ′ ( K ′ ) T d k ′ V .
Score complexity: O ( N 2 d k ′ ) , speedup d k / d k ′ over standard attention. The full end-to-end cost is analyzed in Section 4.1.

3.2. Sparsity Pattern Definitions

We introduce two families of sparsity patterns for the projection matrix S: fixed patterns, which preserve JL guarantees, and learnable patterns, which adapt to the task at the cost of formal theoretical guarantees.
  • Fixed Sparsity Patterns: Random sparse, block sparse, banded sparse—all satisfying the JL guarantees of Theorem 1.
  • Learnable Sparsity Patterns: S i j = M i j · W i j with binary mask M and learnable weights W. The mask is optimized via Gumbel-softmax reparameterization [25] with temperature annealed from 1.0 to 0.1, L 0 regularization, or magnitude pruning. Unlike fixed projections, learnable patterns are data-driven and the JL guarantees of Theorem 1 do not apply post-training (see Section 4.4).
  • Structured Sparsity Patterns: Hierarchical, conv, and head-specific patterns.

3.3. Visual Architecture Overview

Figure 1 contrasts the standard self-attention pipeline with the proposed SPA. The key architectural difference is the insertion of a learnable (or random) sparse projection S that maps queries and keys from R d k to R d k ′ before the attention-score computation. The projection S is shared across all positions in a given layer, so the additional memory cost is independent of the sequence length N.
Figure 1. Architecture comparison between standard self-attention (left) and Sparse Projection Attention (right). SPA reduces the feature dimension from d k to d k ′ ≪ d k , lowering score computation from O ( N 2 d k ) to O ( N 2 d k ′ ) .

3.4. Gradient Analysis and Training Dynamics

3.4.1. Gradient Flow

∇ Q L = ( ∇ Q ′ L ) S T , ∇ S L = Q T ( ∇ Q ′ L ) .

3.4.2. Convergence Analysis

Assumption 1
(L-smoothness). ∥ ∇ L ( θ ) − ∇ L ( ϕ ) ∥   ≤ L ∥ θ − ϕ ∥ for all θ , ϕ ∈ Θ .
Assumption 2
(Bounded gradients). E [ ∥ ∇ L ( θ ) ∥ 2 ] ≤ G 2 for all θ ∈ Θ .
Assumption 3
(Bounded projection error). ∥ ∇ L ( θ ) − ∇ ˜ L ( θ ) ∥   ≤ δ ϵ , where δ ϵ = C 1 ϵ ∥ Q ∥ F ∥ K ∥ F and C 1 > 0 .
Theorem 3
(SPA Convergence). Under Assumptions 1–3, with η t = η / T :
1 T ∑ t = 1 T E ∥ ∇ L ( θ t ) ∥ 2 ≤ 2 ( L ( θ 0 ) − L * ) η T + η L G 2 + 2 G δ ϵ .
Proof. Step 1: Descent lemma. 
By L-smoothness: L ( θ t + 1 ) ≤ L ( θ t ) − η t ⟨ ∇ t , ∇ ˜ t ⟩ + L η t 2 2 ∥ ∇ ˜ t ∥ 2 .
Step 2: Cross-term. By Cauchy-Schwarz and Assumptions 2 and 3: ⟨ ∇ t , ∇ ˜ t ⟩ =   ∥ ∇ t ∥ 2 + ⟨ ∇ t , ∇ ˜ t − ∇ t ⟩ ≥   ∥ ∇ t ∥ 2 − ∥ ∇ t ∥ ∥ ∇ ˜ t − ∇ t ∥   ≥   ∥ ∇ t ∥ 2 − G δ ϵ .
Step 3: Quadratic term. ∥ ∇ ˜ t ∥   ≤   ∥ ∇ t ∥   +   δ ϵ ≤ G + δ ϵ ≤ 2 G (for δ ϵ   ≤   G ), so ∥ ∇ ˜ t ∥ 2   ≤ 4 G 2 .
Step 4: Substitution and summation. Substituting and summing over t = 1 , … , T :
∑ t = 1 T η t ∥ ∇ t ∥ 2 ≤ L ( θ 0 ) − L * + G δ ϵ ∑ t = 1 T η t + 2 L G 2 ∑ t = 1 T η t 2 .
Step 5: Computing the sums. With η t = η / T : ∑ t = 1 T η t = η T and ∑ t = 1 T η t 2 = η 2 . Dividing by η T yields the stated bound.    □

3.4.3. Interpretation of the Bound

The right-hand side of the bound decomposes into three contributions:
  • an optimization term  2 ( L ( θ 0 ) − L * ) η T , which decays as O ( 1 / T ) with the number of iterations;
  • a step-size bias  η L G 2 , controlled by the learning rate η ;
  • a projection bias  2 G δ ϵ , controlled by the projection-error term δ ϵ = C 1 ϵ ∥ Q ∥ F ∥ K ∥ F from Assumption 3.
Convergence to a stationary point therefore requires both the step-size and the projection bias to vanish. This holds, for instance, under a decreasing learning-rate schedule η t = η 0 / t combined with d k ′ → ∞ , since by Theorem 1 the projection error satisfies δ ϵ = O log ( N / δ ) / d k ′ → 0 . With these two conditions, the algorithm converges to a stationary point of the true (unprojected) objective; with the constant schedule η t = η / T used above, the bound only guarantees convergence to a stationary region whose radius is controlled by η and d k ′ . This is consistent with classical results for perturbed SGD [26].
Scope note. This result is a corollary of general perturbed-SGD theory [26]. It establishes that the projection approximation error does not prevent convergence; it does not claim a convergence advantage over standard attention.

3.5. Algorithms

This subsection presents the two procedures underlying the SPA framework. Algorithm 1 formalizes the forward pass of SPA, making explicit the per-step computational cost of each operation. Algorithm 2 extends the procedure to the learnable-sparsity setting, alternating between standard gradient-based updates of the projection matrix S and structured pruning/regrowth steps that enforce the target sparsity level.
Algorithm 1 Sparse Projection Attention (SPA)
Require: 
Q , K ∈ R N × d k , V ∈ R N × d v , d k ′ , s
Ensure: 
O ∈ R N × d v
 1: Init S ∈ R d k × d k ′ , sparsity 1 / s
 2:  Q ′ ← Q S
▹ O ( N d k d k ′ / s )
 3:  K ′ ← K S
▹ O ( N d k d k ′ / s )
 4:  A ← Q ′ ( K ′ ) T / d k ′
▹ O ( N 2 d k ′ )
 5:  W ← softmax ( A )
▹ O ( N 2 )
 6:  O ← W V
▹ O ( N 2 d v )
 7: return  O
Algorithm 2 Learnable Sparse Projection Training
Require: 
D , S, sparsity target ρ
Ensure: 
Trained S
  1:
for epoch = 1 to E do
  2:
    for batch ( Q , K , V ) in D  do
  3:
        Forward: SPA output via Algorithm 1
  4:
        Compute L
  5:
        Backward: ∇ S L
  6:
        Update S (Adam)
  7:
        Prune to sparsity ρ
  8:
        [Optional] Regrow via gradient magnitude
  9:
    end for
10:
end for
11:
return  S

4. Theoretical Analysis

4.1. Complexity Analysis

4.1.1. End-to-End Cost Analysis

The speedup of SPA must be understood carefully. The total forward pass cost is summarized below; a side-by-side comparison with competing methods is provided in Table 2.
T SPA = O ( N d k d k ′ / s ) ︸ project Q , K + O ( N 2 d k ′ ) ︸ scores + O ( N 2 d v ) ︸ output W V ,
compared to T std = O ( N 2 d k ) + O ( N 2 d v ) . The speedup factor d k / d k ′ applies only to score computation. The asymptotic end-to-end speedup is:
T std T SPA → N → ∞ d k + d v d k ′ + d v .
For d k = d v = 512 , d k ′ = 64 : ≈ 1.9 × end-to-end, vs. 8 × for score computation alone.
Table 2. Theoretical complexity comparison. Column “JL Guarantee” indicates whether JL distance-preservation guarantees apply after training.

4.1.2. Hardware Considerations

Sparse GEMM kernels yield practical speedups on A100 GPUs only above 70–90% sparsity [27]. Reported wall-clock times in Section 6 reflect actual GPU execution and capture these hardware-level effects.

4.2. Error Bound Analysis

Theorem 4
(SPA Approximation Error). Let A = softmax ( Q K T / d k ) and A ˜ = softmax ( Q S S T K T / d k ′ ) . With probability ≥ 1 − δ :
∥ A − A ˜ ∥ F   ≤ C ∥ Q ∥ F ∥ K ∥ F log ( N / δ ) d k ′ 1 d k + 1 d k ′ ,
where C > 0 is a universal constant.
Proof. 
By Lemma 2 and the triangle inequality:
∥ A − A ˜ ∥ F ≤   ∥ Q K T − Q S S T K T ∥ F d k + ( 1 + ϵ 0 ) ∥ Q ∥ F ∥ K ∥ F / d k ′ ,
where ϵ 0 = C ′ log ( N / δ ) / d k ′ from Theorem 1, and the first numerator is bounded by ϵ 0 ∥ Q ∥ F ∥ K ∥ F via the entry-wise argument of Theorem 2. Combining yields the stated bound. □

4.3. Generalization Bounds

Assumption 4
(Bounded projection class). Let F be the class of loss functions of transformers with SPA layers where all parameters except S are fixed. S ∈ S ≜ { S :   ∥ S ∥ F   ≤ W max } . The loss ℓ : Y × Y → [ 0 , B ] is B-bounded. The model is L f -Lipschitz in S:
| f S ( x ) − f S ′ ( x ) | ≤ L f ∥ S − S ′ ∥ F for all inputs x.
Theorem 5
(Projection-Component Generalization Bound). For a transformer with SPA layers, treating all parameters except S as fixed, with probability ≥ 1 − δ :
L ( θ ) ≤ L ^ ( θ ) + O d k d k ′ W max 2 + log ( 1 / δ ) N train ,
where W max bounds ∥ S ∥ F and N train is the number of training samples.
Proof. 
We apply Rademacher complexity theory [28]. Under Assumption 4, the covering number of S satisfies log N ( S , ε , ∥ · ∥ F ) ≤ d k d k ′ log ( 1 + 2 W max / ε ) . By Dudley’s entropy integral:
R ^ N train ( F ) ≤ 4 L f W max d k d k ′ N train · log ( 1 + 2 N train ) = O d k d k ′ W max 2 N train .
The effective parameter count p eff = d k d k ′ is independent of sequence length N. Substituting into the standard generalization inequality yields the bound. □
Remark 3
(Scope of the Generalization Bound). This theorem covers only the projection component’s contribution to model complexity. The full model contains many additional parameters ( W Q , W K , W V , feed-forward layers) whose complexity is not reduced by SPA. The benefit of reducing d k ′ is best understood as implicit regularization of the projection, not as a reduction of the full model’s VC dimension. The original bound contained a log N factor in the parameter count, which is not warranted for a matrix of fixed size d k × d k ′ ; this factor has been removed.

4.4. Theoretical Validity of Learnable Projections

4.4.1. Scope of JL-Based Guarantees

Theorems 1, 2 and 4 apply to fixed random projections. The convergence (Theorem 3) and generalization (Theorem 5) results are valid for all variants. For SPA-Learn and SPA-Adaptive, JL guarantees hold at initialization only.

4.4.2. Alternative Framework: Restricted Isometry Property

The RIP [29] provides a weaker condition sufficient for distance preservation by non-random matrices.
Proposition 2
(Learned Projection Quality under RIP). Let S * be the matrix after training. If S * satisfies RIP of order 2 with constant δ 2 < 1 :
| ⟨ q i , k j ⟩ − ⟨ q i S * , k j S * ⟩ | ≤ δ 2 ∥ q i ∥ 2 ∥ k j ∥ 2 , ∀ i , j ∈ [ N ] .
Learned projections perform task-specific dimensionality reduction, preserving distances most relevant to the downstream task. Extending rigorous JL-style guarantees to learned matrices remains an open theoretical problem.

5. Experimental Setup

5.1. Datasets and Tasks

  • Language Modeling: Wikitext-103 [30], PG-19 [31], ArXiv [14].
  • Machine Translation: WMT14 EN-DE, WMT16 EN-RO.
  • LRA [32]: ListOps, Text, Retrieval, Image, Pathfinder.

5.2. Baseline Models

We compare against Standard Transformer [1], Sparse Transformer [10], Longformer [7], Linformer [12], Performer [15], and BigBird [11].

5.3. Implementation Details

All models implemented in PyTorch 2.1, trained on 4 × NVIDIA A100 (80 GB) GPUs with mixed-precision (fp16) and gradient clipping at norm 1.0.

5.3.1. Optimizer

AdamW [33], β 1 = 0.9 , β 2 = 0.98 , weight decay 10 − 2 .

5.3.2. Learning Rate Schedule

Linear warmup for 4000 steps followed by inverse square-root decay: η t = η max · min ( t − 0.5 , t · t warm − 1.5 ) , with η max = 5 × 10 − 4 (Small), 3 × 10 − 4 (Base), 10 − 4 (Large).

5.3.3. Stopping Criterion

Early stopping on validation loss, patience 5. Maximum: 100 epochs (LRA), 50 (LM), 30 (MT).

5.3.4. Batch Sizes

512 tokens: 128 seq; 1024 tokens: 64 seq; 2048 tokens: 32 seq.

5.3.5. Hyperparameter Selection

d k ′ by grid search over { 16 , 32 , 48 , 64 , 96 , 128 } on validation sets: d k ′ = 32 (LM), d k ′ = 48 (MT), d k ′ = 32 (LRA). Sparsity: 80% for fixed patterns; L 0 coefficient λ = 10 − 5 for learnable patterns.

5.3.6. Learnable Variants

SPA-Learn: Gumbel-softmax [25] with temperature annealed from 1.0 to 0.1. SPA-Adaptive: additionally uses a per-head scalar gate.

6. Results

6.1. Language Modeling

SPA-Adaptive achieves perplexity within 0.6 points of the Standard Transformer at sequence length 512 while providing 2.4 × speedup in score computation (≈ 1.9 × end-to-end). The difference at sequence length 2048 is not statistically significant ( p = 0.08 ). Against other efficient baselines, SPA-Adaptive consistently outperforms by 1–3 perplexity points. Detailed quantitative results are reported in Table 3.
Table 3. Language modeling on Wikitext-103 (Perplexity, mean ± std, 3 seeds). Bold values indicate the best score per column. †: not significant vs. Standard Transformer ( p > 0.05 , paired t-test).
A graphical comparison of performance versus computational cost across sequence lengths is shown in Figure 2.
Figure 2. Language modeling performance and efficiency across sequence lengths. Panel (b) compares memory scaling for Standard Attention, SPA ( d k ′ = 64 ), Longformer, and Linformer.

6.2. Machine Translation

SPA-Adaptive achieves BLEU within 0.2 of the Standard Transformer (28.2 vs. 28.4 EN→DE; p > 0.05 ) while reducing inference time by 2.1 × . Against other efficient baselines, SPA-Adaptive improves BLEU by 0.4–1.3 points while also being faster. Full results are reported in Table 4.
Table 4. Machine Translation on WMT14 EN-DE (BLEU Score, mean ± std, 3 seeds). Bold values indicate the best score per column. †: not significant vs. Standard Transformer ( p > 0.05 , paired t-test, two-tailed).
The BLEU–latency trade-off is visualized in Figure 3.
Figure 3. Machine translation results: BLEU score and inference time comparison.

6.3. Long-Range Arena

SPA-Adaptive achieves accuracy within 0.1–0.3 percentage points of the Standard Transformer on all five LRA tasks while being 3.1 × faster at inference. SPA does not surpass the Standard Transformer on any individual task; the contribution is near-parity performance at significantly reduced computational cost. Compared to other efficient baselines, SPA consistently improves accuracy by 2–3 points. Per-task results are reported in Table 5 and visualized in Figure 4.
Table 5. Long Range Arena (Accuracy %, mean ± std, 3 seeds). SPA-Adaptive is within 0.1–0.3 points of the Standard Transformer on all tasks ( p > 0.05 ); the contribution is competitive performance at reduced cost, not superior performance.
Figure 4. Performance on Long Range Arena benchmark.

6.4. Ablation Studies

6.4.1. Projection Dimensionality

Diminishing returns beyond d k ′ = 64 suggest most task-relevant information is captured in a reduced subspace, as illustrated in Figure 5.
Figure 5. Performance vs. projection dimension d k ′ .

6.4.2. Effect of the Sparsity Pattern

We now study the sensitivity of SPA to the structure of the sparsity pattern at fixed projection dimension. Table 6 reports perplexity, training time, and memory for the six fixed patterns introduced in Section 4.4 as well as for the two learnable variants. The two learnable variants (Learnable and Adaptive) achieve the lowest perplexity at the cost of moderately higher training time and memory.
Table 6. Ablation: Sparsity patterns (Wikitext-103 Perplexity). Bold value indicates the best perplexity achieved across patterns.

7. Discussion

7.1. Theoretical Implications

The JL lemma provides a strong theoretical foundation for dimensionality reduction in attention for fixed random projection variants. Sparse projections act as implicit regularizers. For learnable variants, the theoretical justification is provided by the RIP framework (Section 4.4).

7.2. Practical Implications

SPA enables transformer training on consumer hardware, faster inference, and reduced energy consumption. The practical end-to-end speedup is ≈ 1.9 × for typical configurations ( d k = d v = 512 , d k ′ = 64 ), meaningful for resource-constrained settings even if below the 8 × theoretical maximum for score computation.

7.3. Limitations

(i) Theoretical gaps. JL-based guarantees apply only to fixed random projection variants. For SPA-Learn and SPA-Adaptive, formal distance-preservation guarantees do not hold post-training. The convergence result is a corollary of generic perturbed-SGD theory. The generalization bound covers only the projection component.
(ii) Memory complexity. SPA does not reduce the O ( N 2 ) space complexity. For sequences above ≈ 4000 tokens, memory remains the primary bottleneck.
(iii) Practical efficiency gap. The 8 × speedup applies to score computation only; end-to-end wall-clock speedup is ≈ 2 × on A100.
(iv) Limited reproducibility data. Results are over 3 seeds. Some differences between SPA variants and baselines may not be statistically significant.
(v) Learnable variant sensitivity. SPA-Learn and SPA-Adaptive require careful tuning of λ and the temperature schedule. Training collapsed to degenerate patterns for λ > 10 − 4 in preliminary experiments.

7.4. Future Work

Future directions: extending formal guarantees to learnable projections via the RIP framework; combining SPA with O ( N ) methods to address memory; hardware-aware sparse projection kernels; cross-modal applications.

8. Conclusions

We presented Sparse Projection Attention (SPA), a novel attention mechanism that leverages mathematically principled sparse projections to reduce attention score computation complexity while maintaining competitive performance. SPA achieves up to 8 × speedup in score computation and ≈ 2 × end-to-end, remaining within 1 perplexity point of the Standard Transformer on language modeling, within 0.3 BLEU on machine translation, and within 0.5 accuracy points on LRA benchmarks. Key innovations include a rigorous mathematical framework with explicit scope conditions, multiple sparsity patterns with JL guarantees for fixed random variants, and learnable variants. SPA represents a practical efficiency-accuracy trade-off for long-sequence tasks. We do not claim state-of-the-art performance on any benchmark; the contribution is the efficiency-accuracy trade-off profile demonstrated across multiple tasks.

Author Contributions

Conceptualization, M.C.A., N.-E.J. and M.E.; methodology, M.C.A.; software, M.C.A.; validation, M.C.A., N.-E.J. and M.E.; formal analysis, M.C.A.; investigation, M.C.A.; resources, N.-E.J. and M.E.; data curation, M.C.A.; writing—original draft preparation, M.C.A.; writing—review and editing, N.-E.J. and M.E.; visualization, M.C.A.; supervision, M.E.; project administration, M.E.; and funding acquisition, M.E. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Acknowledgments

The authors thank the anonymous reviewers for their valuable comments.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

SPASparse Projection Attention
JLJohnson–Lindenstrauss
LRALong Range Arena
STEStraight-Through Estimator
RIPRestricted Isometry Property

References

  1. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
  2. Devlin, J.; Chang, M.-W.; Lee, K.; Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT; Association for Computational Linguistics: Minneapolis, MN, USA, 2019; pp. 4171–4186. [Google Scholar]
  3. Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I. Language models are unsupervised multitask learners. OpenAI Blog 2019, 1, 9. [Google Scholar]
  4. Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16 × 16 words: Transformers for image recognition at scale. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2021. [Google Scholar]
  5. Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF ICCV; IEEE: Los Alamitos, CA, USA, 2021; pp. 10012–10022. [Google Scholar]
  6. Gulati, A.; Qin, J.; Chiu, C.-C.; Parmar, N.; Zhang, Y.; Yu, J.; Han, W.; Wang, S.; Zhang, Z.; Wu, Y.; et al. Conformer: Convolution-augmented transformer for speech recognition. In Proceedings of Interspeech 2020; ISCA: Graz, Austria, 2020; pp. 5036–5040. [Google Scholar]
  7. Beltagy, I.; Peters, M.E.; Cohan, A. Longformer: The long-document transformer. arXiv 2020, arXiv:2004.05150. [Google Scholar] [CrossRef] [Scilit]
  8. Avsec, Ž.; Agarwal, V.; Visentin, D.; Ledsam, J.R.; Grabska-Barwinska, A.; Taylor, K.R.; Assael, Y.; Jumper, J.; Kohli, P.; Kelley, D.R. Effective gene expression prediction from sequence by integrating long-range interactions. Nat. Methods 2021, 18, 1196–1203. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Zhou, H.; Zhang, S.; Peng, J.; Zhang, S.; Li, J.; Xiong, H.; Zhang, W. Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of AAAI; AAAI Press: Palo Alto, CA, USA, 2021; Volume 35, pp. 11106–11115. [Google Scholar]
  10. Child, R.; Gray, S.; Radford, A.; Sutskever, I. Generating long sequences with sparse transformers. arXiv 2019, arXiv:1904.10509. [Google Scholar] [CrossRef] [Scilit]
  11. Zaheer, M.; Guruganesh, G.; Dubey, K.A.; Ainslie, J.; Alberti, C.; Ontanon, S.; Pham, P.; Ravula, A.; Wang, Q.; Yang, L.; et al. Big bird: Transformers for longer sequences. Adv. Neural Inf. Process. Syst. 2020, 33, 17283–17297. [Google Scholar]
  12. Wang, S.; Li, B.Z.; Khabsa, M.; Fang, H.; Ma, H. Linformer: Self-attention with linear complexity. arXiv 2020, arXiv:2006.04768. [Google Scholar] [CrossRef] [Scilit]
  13. Katharopoulos, A.; Vyas, A.; Pappas, N.; Fleuret, F. Transformers are RNNs: Fast autoregressive transformers with linear attention. In Proceedings of ICML; PMLR: Vienna, Austria, 2020; pp. 5156–5165. [Google Scholar]
  14. Kitaev, N.; Kaiser, Ł.; Levskaya, A. Reformer: The efficient transformer. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2020. [Google Scholar]
  15. Choromanski, K.; Likhosherstov, V.; Dohan, D.; Song, X.; Gane, A.; Sarlos, T.; Hawkins, P.; Davis, J.; Mohiuddin, A.; Kaiser, L.; et al. Rethinking attention with performers. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2021. [Google Scholar]
  16. Lee-Thorp, J.; Ainslie, J.; Eckstein, I.; Ontanon, S. FNet: Mixing tokens with Fourier transforms. In Proceedings of NAACL-HLT; Association for Computational Linguistics: Seattle, WA, USA, 2022; pp. 4296–4313. [Google Scholar]
  17. Xiong, Y.; Zeng, Z.; Chakraborty, R.; Tan, M.; Fung, G.; Li, Y.; Singh, V. Nyströmformer: A Nyström-based algorithm for approximating self-attention. In Proceedings of AAAI; AAAI Press: Palo Alto, CA, USA, 2021; Volume 35, pp. 14138–14148. [Google Scholar]
  18. Peng, H.; Pappas, N.; Yogatama, D.; Schwartz, R.; Smith, N.A.; Kong, L. Random feature attention. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2021. [Google Scholar]
  19. Chen, B.; Dao, T.; Winsor, E.; Song, Z.; Rudra, A.; Ré, C. Scatterbrain: Unifying sparse and low-rank attention approximation. Adv. Neural Inf. Process. Syst. 2021, 34, 17413–17426. [Google Scholar]
  20. Qin, Z.; Sun, W.; Deng, H.; Li, D.; Wei, Y.; Lv, B.; Yan, J.; Kong, L.; Zhong, Y. cosFormer: Rethinking softmax in attention. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2022. [Google Scholar]
  21. Tay, Y.; Dehghani, M.; Bahri, D.; Metzler, D. Efficient transformers: A survey. ACM Comput. Surv. 2022, 55, 109. [Google Scholar] [CrossRef] [Scilit]
  22. Johnson, W.B.; Lindenstrauss, J. Extensions of Lipschitz mappings into a Hilbert space. Contemp. Math. 1984, 26, 189–206. [Google Scholar]
  23. Dasgupta, S.; Gupta, A. An elementary proof of a theorem of Johnson and Lindenstrauss. Random Struct. Algorithms 2003, 22, 60–65. [Google Scholar] [CrossRef] [Scilit]
  24. Achlioptas, D. Database-friendly random projections: Johnson–Lindenstrauss with binary coins. J. Comput. Syst. Sci. 2003, 66, 671–687. [Google Scholar] [CrossRef] [Scilit]
  25. Jang, E.; Gu, S.; Poole, B. Categorical reparameterization with Gumbel-softmax. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2017. [Google Scholar]
  26. Bottou, L.; Curtis, F.E.; Nocedal, J. Optimization methods for large-scale machine learning. SIAM Rev. 2018, 60, 223–311. [Google Scholar] [CrossRef] [Scilit]
  27. Gale, T.; Elsen, E.; Hooker, S. The state of sparsity in deep neural networks. arXiv 2020, arXiv:1902.09574. [Google Scholar]
  28. Bartlett, P.L.; Mendelson, S. Rademacher and Gaussian complexities: Risk bounds and structural results. J. Mach. Learn. Res. 2002, 3, 463–482. [Google Scholar]
  29. Candès, E.J.; Romberg, J.; Tao, T. Stable signal recovery from incomplete and inaccurate measurements. Commun. Pure Appl. Math. 2006, 59, 1207–1223. [Google Scholar] [CrossRef] [Scilit]
  30. Merity, S.; Keskar, N.S.; Socher, R. Pointer sentinel mixture models. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2017. [Google Scholar]
  31. Rae, J.W.; Potapenko, A.; Jayakumar, S.M.; Hillier, C.; Lillicrap, T.P. Compressive transformers for long-range sequence modelling. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2020. [Google Scholar]
  32. Tay, Y.; Dehghani, M.; Abnar, S.; Shen, Y.; Bahri, D.; Pham, P.; Rao, J.; Yang, L.; Ruder, S.; Metzler, D. Long range arena: A benchmark for efficient transformers. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2021. [Google Scholar]
  33. Loshchilov, I.; Hutter, F. Decoupled weight decay regularization. In Proceedings of ICLR; OpenReview.net: Alameda, CA, USA, 2019; Available online: https://openreview.net/forum?id=Bkg6RiCqY7 (accessed on 10 May 2026).
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.