Next Article in Journal
Eigenfunction Expansion and Parseval Identity of β-Dirac Operators on the Whole Line
Next Article in Special Issue
Mathematical Sociology: Models, Structures, and Open Problems
Previous Article in Journal
Weighted Stein–Weiss and John–Nirenberg Inequalities for Fractional Singular Integral Operators on Fofana Spaces
Previous Article in Special Issue
A Dynamical Systems Model of Port–Industry–City Co-Evolution Under Data Constraints
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Value Matrix Dynamics and Attractor Selection in a Reduced Transformer Model

1
Institut Camille Jordan, UMR 5208 CNRS, University Lyon 1, 69622 Villeurbanne, France
2
S.M. Nikolskii Mathematical Institute, Peoples’ Friendship University of Russia, 6 Miklukho-Maklaya St, 117198 Moscow, Russia
Mathematics 2026, 14(18), 3350; https://doi.org/10.3390/math14183350 (registering DOI)
Submission received: 31 July 2026 / Revised: 8 September 2026 / Accepted: 11 September 2026 / Published: 15 September 2026
(This article belongs to the Special Issue Dynamical Systems and Complex Systems)

Abstract

We study an interacting particle system on the unit sphere S d 1 motivated by the forward-pass dynamics of a transformer model that combines a value matrix V and a u-aligned multi-layer perceptron (MLP) Φ, a particular form of MLP specified below. For a symmetric value matrix and this MLP, the one-cluster dynamic is the Riemannian gradient flow of an explicit Lyapunov function L = f V + g , which combines the Rayleigh quotient of V with the MLP potential. Every trajectory converges to a fixed point, limit cycles are excluded, and the fixed points are the critical points of L: local maxima are stable attractors and saddles are unstable. Isolating the value matrix ( Φ 0 , with V possibly non-symmetric and a real strictly dominant eigenvalue λ 1 ), we show that the dominant eigenvector v 1 is the selected attractor and that its basin is almost the whole sphere, the two basins of ± v 1 being separated by an invariant subsphere carried by the left eigenvector of V. The subdominant eigenvectors form a hierarchical structure of saddles, and a complex dominant pair replaces the point attractor by a great-circle limit cycle. When the tokens move independently, softmax attention contributes no first-order clustering force: the tokens collapse to ± v 1 at a consensus rate equal to λ 1 and independent of the attention matrix, and whether two clusters in opposite basins coexist or merge is governed by the single order parameter s = v 1 A v 1 . Thus, the value matrix, rather than the attention, controls the location and the selection of the limiting cluster. All results are confirmed by numerical integration of the continuous-time ordinary differential equation (ODE).

1. Introduction

1.1. Token Dynamics and Interacting Particle Systems

Transformer models [1] process a sequence of tokens through a succession of attention and feed-forward layers. When the layers are regarded as discrete time steps, the forward pass becomes a dynamical system acting on the token representations, and its continuous-time (deep) limit is a system of ordinary differential equations. If the representations are normalised, the tokens evolve as interacting particles on the unit sphere S d 1 (in d-dimensional space R d ), coupled through the softmax attention weights. This point of view has been developed recently [2,3,4], and reveals that the tokens tend to cluster as the depth increases: the particles aggregate toward a small number of limiting configurations, determined by the initial data and weight matrices. Clustering is the dynamical counterpart of an empirical observation about deep transformers, namely that the token representations tend to become similar to each other, a phenomenon known in the machine learning literature as rank collapse, oversmoothing, or token uniformity [5,6].
The forward pass combines three ingredients: the attention weights, which couple the tokens; the value matrix V, which acts linearly on each token before the attention-weighted sum; and the feed-forward (MLP) block, which acts on each token separately. The question addressed in the present work is how these ingredients interact: which of them selects the limiting cluster, and at what rate the tokens collapse toward it.

1.2. Related Work

The interpretation of self-attention as an interacting particle system on the sphere was introduced by Geshkovski et al. [2,3]. They consider the pure-attention dynamics, in which the value matrix is the identity and no feed-forward block is present, and prove that the tokens cluster in infinite time; the type of limiting configuration depends on the spectrum of the value matrix, and the attention matrix converges to a low-rank Boolean matrix. The mean-field description of these dynamics, in which the empirical measure of the tokens satisfies a nonlinear transport equation, is studied in [4]. These works connect transformer dynamics to classical models of collective behaviour and consensus, such as synchronisation of coupled oscillators [7], self-propelled particle and flocking models [8,9], and opinion dynamics with bounded confidence [10], and to the view of deep residual networks as discretisations of ordinary differential equations [11].
A parallel line of work, motivated by the design of transformer architectures, studies the loss of expressivity caused by clustering. Dong et al. [5] show that pure attention, without skip connections or feed-forward blocks, loses rank doubly exponentially with depth, and that the skip connections and the MLP are precisely what prevents the collapse to a rank-one representation. The role of attention normalisation and layer normalisation in this collapse has been subsequently examined [6], and doubly stochastic variants of attention have been proposed [12]. In all of these works, the value matrix is either the identity or is not the object of study, and the feed-forward block is either absent or treated only qualitatively. Related recent developments include the analysis of clustering under causal attention masking [13], the mean-field study of the perceptron (feed-forward) block, which shows that its critical points are generically atomic and localised on the sphere [14], and the white-box, rate-reduction view of transformer layers [15].
In contrast to these works, which treat the three ingredients of the forward pass (attention, the value matrix, and the feed-forward block) in isolation, the present paper studies their joint action and disentangles their distinct roles: attention drives the collapse, the value matrix selects its location, and the feed-forward block shapes the attractor landscape. The value matrix is added to the model precisely to determine its role in the selection of the limiting cluster.

1.3. Contributions and Outline

We study the token dynamics (1) under two successive simplifications: first, the value matrix acting alone ( Φ 0 ), then, the value matrix together with a u-aligned MLP (a particular form of MLP, defined in Section 2). The main results are as follows.
For the value matrix alone, the one-cluster flow x ˙ = P x ( V x ) admits a complete and explicit description that requires no symmetry of V. When the dominant eigenvalue λ 1 is real and simple, the corresponding eigenvector v 1 is the selected attractor (Theorem 1), and its basin is almost the whole sphere: every trajectory converges to + v 1 or v 1 except on a measure-zero invariant subsphere, which is carried by the left eigenvector w 1 and is tilted off the geometric equator when V is non-symmetric (Theorem 2, Lemma 4). The subdominant eigenvectors organise into a hierarchical structure of saddles, and a complex dominant pair replaces the point attractor by an invariant great-circle limit cycle (Proposition 1, Remark 6). Whether nearby tokens actually cluster is then governed by the sign of the dominant eigenvalue: λ 1 > 0 gives collapse to a point, λ 1 < 0 gives one-cluster stability but token spreading, and a complex dominant pair gives no point cluster at all. Releasing the tokens to move independently under the value matrix, we show that softmax attention contributes no first-order clustering force, so the collapse is governed by V alone. When λ 1 > 0 , the tokens collapse to ± v 1 according to sign w 1 , x , with a consensus rate equal to λ 1 and independent of the attention matrix (Theorem 3), and the positive-eigenvalue hemispheres and cones are trapping (Propositions 2 and 3). Attention cannot override the value matrix: collapse from generic data is robust across all attention classes, and the only genuinely attention-dependent question, whether two clusters in opposite basins coexist or merge, is decided by a single order parameter s = v 1 A v 1 , the alignment of the attention along the cluster axis, with coexistence for s > 0 and merging for s < 0 (Section 4.3).
Adding a u-aligned MLP Φ to a symmetric value matrix V, we show that the combined one-cluster ODE is again a Riemannian gradient flow, now of the explicit Lyapunov function L = f V + g that combines the Rayleigh quotient of V with the MLP potential g (Theorem 4). Every trajectory converges to a fixed point and limit cycles are excluded; the fixed points are the critical points of L, with local maxima stable and saddles unstable (Corollary 1). Near a stable maximum, the tokens collapse to a single cluster, while near a saddle, the cluster escapes toward the global maximum of L yet remains clustered, under an explicit fluctuation-stability condition (Theorem 5). The dominant eigenvalue of V controls which direction is selected as the global attractor (Remark 11).
These two settings are unified by a general observation, valid for arbitrary V and Φ: a fixed point that is stable for the one-cluster ODE need not be stable for the multi-token system, because the value matrix enters the mean-mode operator but not the fluctuation operator (Proposition 4). The u-aligned MLP with symmetric V is the exactly solvable case in which this instability is structurally excluded. All the analytical results are confirmed by numerical integration of the continuous-time ODE.
The paper is organised as follows. Section 3 treats the value matrix alone in the one-cluster regime: fixed points, the dominant-eigenvector attractor, almost-global bistability, and the hierarchical flag. Section 4 releases the tokens and studies multi-token collapse, trapping, and the role of the attention matrix. Section 5 adds the u-aligned MLP and establishes the Lyapunov structure together with the uncollapsed token dynamics. Section 6 presents the general one-cluster versus multi-token mechanism, and Section 7 is devoted to a discussion of the results.
  • Highlights:
  • For the value matrix alone, the dominant eigenvector v 1 is the selected attractor and its basin is almost the whole sphere; the two basins of ± v 1 are separated by an invariant subsphere carried by the left eigenvector w 1 , tilted off the geometric equator when V is non-symmetric (Theorems 1 and 2).
  • Softmax attention exerts no first-order clustering force (Proposition 5): the collapse rate equals the dominant eigenvalue λ 1 and is independent of the attention matrix, and whether two antipodal clusters coexist or merge is governed by the single order parameter s = v 1 A v 1 .
  • For symmetric V with a u-aligned MLP, the one-cluster flow is the Riemannian gradient flow of L = f V + g ; every trajectory converges to a critical point and limit cycles are excluded (Theorem 4).
  • One-cluster stability need not imply multi-token stability, because V enters the mean operator L col but not the fluctuation operator L fluct (Proposition 4).

2. Model

  • Setting
Let n tokens x 1 , , x n S d 1 R d evolve in continuous time. Fix r d orthonormal directions u 1 , , u r S d 1 and write s i k = u k , x i . The continuous-time ODE (the Euler limit of the discrete normalised residual update; cf. the neural-ODE view of residual networks [11]) is:
x ˙ i = P x i j w i j V x j + Φ ( x i ) , P x y = y x , y x ,
where w i j = softmax j ( β x i , A x j / d ) are attention weights, A is a coupling matrix, and β > 0 . Throughout, · denotes the Euclidean norm on R d and · , · the associated inner product, so S d 1 is the unit sphere in this norm.
  • Symmetric value matrix
We take V = V symmetric with eigenvectors u 1 , , u r and eigenvalues ν 1 ν r 0 . Its Rayleigh quotient is f V ( x ) = 1 2 x , V x . In a full transformer, the attention output is W O j w i j W V x j , so the value/output map V = W O W V acts on the values inside the attention-weighted sum; in the reduced model (1), the same V acts linearly on each token before the attention-weighted average.
  • u -aligned MLP
We consider a particular form of feed-forward block, in which the input and output directions are tied to the same orthonormal feature basis { u k } . This u-aligned MLP is
Φ ( x ) = k = 1 r γ k σ ( α k s k + b k ) u k , s k = u k , x ,
where γ k , α k , b k R and σ are continuously differentiable (e.g., tanh). The name refers to this alignment with the basis { u k } , not to coordinatewise action in the standard basis; we avoid the term “diagonal MLP”, which would misleadingly suggest the latter. The same orthonormal directions u k are used to read each feature s k = u k , x and to write Φ back (tied input/output weights). Each component Φ i ( x ) therefore depends on all coordinates of x through the s k , so Φ is not coordinatewise in the standard basis; what matters is that its Jacobian,
D Φ ( x ) = k = 1 r γ k α k σ ( α k s k + b k ) u k u k ,
is symmetric (diagonal in the u-basis, with eigenvectors u k and eigenvalues γ k α k σ ( α k s k + b k ) ). It is this symmetry, not standard-basis diagonality, that the Lyapunov structure of Section 5 relies on. The associated scalar potential is
g ( x ) = k = 1 r γ k 0 s k σ ( α k t + b k ) d t .
  • Numerical methods
The results below are illustrated by direct numerical integration of the continuous-time system (1) with an explicit Runge–Kutta scheme, and, for the one-cluster map, by the explicit Euler–retraction iteration x ( I + h V ) x / ( I + h V ) x ; the dimension d, token number n, inverse temperature β, spectrum of V, and step size are stated in each figure caption. Ensembles of initial conditions are drawn uniformly on S d 1 , and the number of initial conditions behind each reported statistic is given in the corresponding caption (for example, 600 one-cluster initial conditions in Figure 1). Two groups seeded in opposite basins are declared to coexist when their inter-cluster geodesic distance stays bounded away from zero over the integration horizon, and to merge otherwise; this coincides with the sign of the order parameter s = v 1 A v 1 of Section 4.3.

3. One Cluster Without MLP

Setting Φ 0 isolates the value matrix. The resulting one-cluster flow x ˙ = P x ( V x ) admits a sharp, fully explicit analysis that requires no symmetry of V: we locate its fixed points, identify the dominant eigenvector as the selected attractor, show that the selection is almost-global, and organise the remaining dynamics into a hierarchical flag culminating in great-circle limit cycles.

3.1. Fixed Points for the One-Cluster Flow

With Φ 0 , the one-cluster ODE is
x ˙ = F ( x ) : = P x ( V x ) = V x x , V x x .
Lemma 1
(Fixed points are real eigenvectors).  x S d 1 is a fixed point of (5) iff V x = λ x for some real λ, necessarily λ = x , V x . In particular, complex eigenvalues of V produce no fixed point on S d 1 .
Proof. 
F ( x ) = 0 means V x = x , V x x , i.e., x is an eigenvector with eigenvalue λ = x , V x R . Conversely, an eigenvector with real eigenvalue gives P x ( V x ) = P x ( λ x ) = 0 . A complex eigenvalue has no real unit eigenvector in R d . □
Lemma 2
(Linearisation). At a fixed point x = v k with V v k = λ k v k , the linearisation of (5) on T v k S d 1 is
A k : = P v k V | T v k S d 1 λ k I ,
and its spectrum is { λ j λ k : j k } , where λ 1 , , λ d are the eigenvalues of V. This holds for arbitrary (possibly non-symmetric) V.
Proof. 
Differentiating F ( x ) = V x ( x V x ) x at x = v k in a direction ξ T v k S d 1 (so v k , ξ = 0 ) gives:
D F ( v k ) ξ = V ξ ξ V v k + v k V ξ v k ( v k V v k ) ξ .
Now, ξ V v k = λ k ξ , v k = 0 , and v k V v k = λ k . The remaining inner term v k V ξ v k = V v k , ξ v k is parallel to v k , hence normal to T v k S d 1 ; the tangential generator is its projection,
A k ξ = P v k D F ( v k ) ξ = P v k V ξ λ k ξ ,
which is (6); the asymmetric part of V drops out through P v k . For the spectrum, complete v k to an orthonormal basis { v k , q 2 , , q d } of R d , where q 2 , , q d is an arbitrary orthonormal basis of v k (not eigenvectors of V which, V being non-symmetric, need be neither orthogonal nor real). Block-triangularity below uses only that v k is an eigenvector. Write J = [ q 2 | | q d ] for the tangent block (so J J = I , J v k = 0 and J J = P v k ) and Q = [ v k | J ] (orthogonal); using V v k = λ k v k , the similarity V ˜ = Q V Q has first column ( λ k , 0 , , 0 ) , hence
V ˜ = λ k b 0 M , det ( V ˜ s I ) = ( λ k s ) det ( M s I ) ,
with b R d 1 generally nonzero (the asymmetry, lying in the normal direction) and M = J V J the lower-right block. Thus, spec ( M ) = spec ( V ) { λ k } = { λ j } j k , read off the characteristic polynomial without diagonalising M. Finally, M = J V J is exactly the matrix of the compression P v k V | T in the basis { q m } : for q l , q m T ,
q l , P v k V q m = P v k q l , V q m = q l , V q m = ( J V J ) l m ,
using P v k q l = q l for q l T (the off-diagonal b = v k V J is the discarded normal component). Hence, A k = M λ k I and spec ( A k ) = spec ( M ) λ k = { λ j λ k } j k . □
A single observation underlies the global picture of the next subsections: the flow is the radial projection of a linear flow.
Lemma 3
(Projectivised flow). For any y 0 0 , let y ( t ) = e t V y 0 and x ( t ) = y ( t ) / y ( t ) . Then, x ( t ) solves x ˙ = P x ( V x ) , and conversely, every solution of the latter is the radial projection of a ray of y ˙ = V y .
Proof. 
Differentiating x = y / y and using y ˙ = V y and d d t y = y , V y / y ,
x ˙ = V y y y y , V y y 3 = V x x , V x x = P x ( V x ) .
Conversely, given a solution x ( t ) on S d 1 , the scalar ODE d d t ρ = x , V x ρ reconstructs y = ρ , and y = ρ x solves y ˙ = V y . □
Lemmas 1–3 are elementary and classical. Lemma 3 exhibits the one-cluster flow as the radial projection (projectivisation) of the linear flow e t V , and the convergence of that projection to the dominant eigendirection is the continuous-time counterpart of the power method [16]. We include self-contained proofs for completeness; the contribution of this section is not these lemmas individually but their assembly into the almost-global attractor-selection picture, the left-eigenvector separator, and the hierarchical flag developed in the following subsections.

3.2. The Dominant-Eigenvector Attractor

Theorem 1
(One-cluster attractor). Suppose λ 1 R is a simple eigenvalue of V with a strictly maximal real part, λ 1 > Re λ j for every other eigenvalue λ j of V, and let v 1 be the corresponding unit eigenvector. Then, v 1 is locally asymptotically stable for (5). Every other real eigenvector v k ( k 1 ) is unstable, with unstable dimensions equal to the number of { j : Re λ j > λ k } 1 .
Proof. 
By Lemma 2, spec ( A 1 ) = { λ j λ 1 } j 1 . Strict dominance gives Re ( λ j λ 1 ) = Re λ j λ 1 < 0 for every j 1 , so A 1 is a possibly non-normal matrix with eigenvalues with negative real parts, and v 1 is locally asymptotically stable; non-normality permits transient growth but not loss of asymptotic stability. At any other real eigenvector v k , spec ( A k ) = { λ j λ k } j k contains λ 1 λ k > 0 (and more generally one positive-real eigenvalue for each λ j with Re λ j > λ k ), so v k is unstable with the stated index. □
Remark 1
(Complex dominant eigenvalue ⇒ no point attractor). If the eigenvalue of the maximal real part is a complex pair, no real eigenvector carries it, so by Theorem 1, every real-eigenvector fixed point has at least one unstable direction and none are stable. On S 2 this forces, by the Poincaré–Bendixson theorem, a periodic orbit; this is the limit-cycle mechanism of the asymmetric-V analysis. The clean statement that the top eigenvalue is the unique attractor is thus exactly the real-dominant-eigenvalue case.

3.3. Almost-Global Bistability and the Tilted Separator

Informally, almost-global bistability states that, outside a measure-zero set of initial data, every trajectory ends at one of the two poles ± v 1 . The reason is transparent from the projectivised-flow viewpoint of Lemma 3: writing x 0 = k c k v k in the eigenbasis, the solution is x ( t ) = e t V x 0 / e t V x 0 , and the dominant mode c 1 e λ 1 t v 1 outgrows all others whenever c 1 0 , so the limit is determined by the sign of c 1 = w 1 , x 0 alone. The exceptional set is exactly { c 1 = 0 } = w 1 S d 1 , a great subsphere of one lower dimension. Two features are specific to a non-symmetric V: the selector is the left eigenvector w 1 , so the separating subsphere is w 1 rather than the geometric equator v 1 , and is tilted off it; further, the sign of w 1 , x is exactly conserved along the flow, so the two basins never mix. The following lemma and theorem make this precise.
Lemma 4
(Invariant separating sphere). Let w 1 be the left eigenvector of V for λ 1 ( V w 1 = λ 1 w 1 ). Along any solution of (5),
d d t w 1 , x = λ 1 x , V x w 1 , x .
Hence, w 1 , x ( t ) = w 1 , x 0 exp 0 t λ 1 x ( s ) , V x ( s ) d s ; the sign of w 1 , x is conserved, and the great subsphere H = { x S d 1 : w 1 , x = 0 } = w 1 S d 1 is invariant.
Proof. 
Using P x ( V x ) = V x x , V x x and w 1 , V x = V w 1 , x = λ 1 w 1 , x ,
d d t w 1 , x = w 1 , P x ( V x ) = w 1 , V x x , V x w 1 , x = λ 1 x , V x w 1 , x ,
which is (7). The integrating factor is positive, so w 1 , x never changes sign, and w 1 , x 0 = 0 forces w 1 , x 0 . □
The use of the left eigenvector is essential: it is w 1 , not v 1 , that satisfies w 1 V = λ 1 w 1 , so only the coordinate w 1 , x decouples. By biorthogonality w 1 , v k = δ 1 k , the invariant subspace w 1 equals span { v 2 , , v d } (the complement of v 1 in the eigenbasis), not the geometric complement v 1 ; the two agree if V = V .
Theorem 2
(Almost-global bistability). Under the hypotheses of Theorem 1, assume V is diagonalisable (Remark 3 removes this) and normalise w 1 , v 1 = 1 . Then ± v 1 are asymptotically stable, and for every x 0 with w 1 , x 0 0 , the ω-limit set is the single point ω ( x 0 ) = { sign w 1 , x 0 v 1 } . The basins are the open half-spheres W ± = { ± w 1 , x > 0 } , which exhaust S d 1 H ; their common boundary is the invariant subsphere H, which contains every fixed point other than ± v 1 together with all of their stable manifolds, and has an empty interior. No globally attracting point exists.
Proof. 
Write y 0 = k c k v k in the (right) eigenbasis, with c 1 = w 1 , y 0 by biorthogonality. By Lemma 3, x ( t ) = e t V y 0 / e t V y 0 . If c 1 0 ,
e λ 1 t e t V y 0 = c 1 v 1 + k 1 e ( λ k λ 1 ) t ( ) t c 1 v 1 ,
since every other mode carries a factor e ( Re λ k λ 1 ) t 0 (the imaginary parts only rotate the shrinking components; the approach may spiral but dist ( x ( t ) , ± v 1 ) 0 ). Since e λ 1 t > 0 is a scalar and normalisation is scale-invariant, we may insert this factor in both numerator and denominator without changing x ( t ) :
x ( t ) = e t V y 0 e t V y 0 = e λ 1 t e t V y 0 e λ 1 t e t V y 0 t c 1 v 1 c 1 v 1 = sign ( c 1 ) v 1 ,
the magnitude | c 1 | cancelling so that only the sign of the (real) dominant coefficient survives. Hence ω ( x 0 ) = { sign ( c 1 ) v 1 } , with sign ( c 1 ) = sign w 1 , x 0 constant by Lemma 4. In particular, no x 0 off H can lie on the stable manifold of any other fixed point: such a fixed point is a real eigenvector v k ( k 2 ) with w 1 , v k = 0 , and convergence to it would require c 1 0 , impossible when c 1 0 is conserved in sign. All other fixed points and their stable manifolds therefore lie in H = { w 1 , x = 0 } ; concretely, they form the heteroclinic skeleton of the deflated dynamics on H (Remark below). Local asymptotic stability of ± v 1 is Theorem 1; oddness of P x ( V x ) gives two disjoint basins, neither the whole sphere, ruling out a global attractor. □
Remark 2
(Mode of approach). Convergence to the point v 1 ( dist ( x ( t ) , v 1 ) 0 ) need not be monotone: when the subdominant eigenvalues are complex, the trajectory spirals into v 1 , winding infinitely while the distance decays like e ( λ 1 Re λ 2 ) t . This is still ω ( x 0 ) = { v 1 } .
Remark 3
(Non-diagonalisable V). The diagonalisability hypothesis is only a convenience. If some subdominant eigenvalue is defective, e t V carries Jordan terms t m e λ k t , but since Re λ k < λ 1 strictly we still have t m e ( λ k λ 1 ) t 0 , so e λ 1 t e t V y 0 ( projection of y 0 onto the λ 1 - eigenline ) and Theorem 2 hold verbatim, with w 1 , x 0 the coordinate read off by the left eigenvector. Only λ 1 itself must be simple.
Remark 4
(Phase portrait and the separating sphere). On H = w 1 S d 1 the flow is, by Lemma 3, the projectivisation of y ˙ = V y restricted to the invariant subspace w 1 , on which V acts as the deflated block M with spectrum { λ j } j 1 (Lemma 2). Thus, the dynamics on H S d 2 are the same problem one dimension down, and the structure of the basin boundary follows by recursion on the subdominant spectrum:
  • If the eigenvalue of M of the maximal real part is real and strictly dominant, H carries its own antipodal pair of attractors, and the remaining real eigenvectors are lower equatorial saddles (their indices given by deflation), forming a heteroclinic skeleton;
  • If it is a complex pair, H has no equatorial point at its top; for d = 3 (so H S 1 ), Poincaré–Bendixson forces H to be a single periodic orbit, a limit cycle bounding the two basins.
For non-symmetric V, the subdominant right eigenvectors v k ( k 2 ) need not lie on the geometric equator; relative to w 1 they sit at w 1 , v k = 0 , hence on H, so they are exactly the equatorial unstable fixed points.
Remark 5
(Discrete-time threshold). The distinction is between the real part and the modulus: the ODE selects the eigenvalue of the largest real part, whereas the discrete map, being power iteration, selects the eigenvalue of the largest modulus  | 1 + h λ | . These two orderings can disagree at finite step h, and the threshold below is exactly where they do. The explicit Euler–retraction scheme x i + 1 = ( I + h V ) x i / ( I + h V ) x i is power iteration for I + h V , and converges to the eigenvector of maximal modulus | 1 + h λ k | . Since | 1 + h λ | 2 = 1 + 2 h Re λ + h 2 | λ | 2 ,
log | 1 + h λ | = h Re λ + h 2 2 ( Im λ ) 2 ( Re λ ) 2 + O ( h 3 ) ,
so the leading O ( h ) term reproduces the ODE Lyapunov exponent Re λ , while the O ( h 2 ) correction adds a spurious + h 2 2 ( Im λ ) 2 favouring rotation. A subdominant complex pair can therefore win the modulus competition at finite step. For spec ( V ) = { 1 , ± 3 i } , | 1 + h | > | 1 ± 3 h i | ( 1 + h ) 2 > 1 + 9 h 2 h < 1 4 , so the discrete map converges to v 1 only for h < 1 4 and otherwise rotates (Figure 2). As h 0 , the real-part ordering is restored, recovering the ODE result of Theorem 2; in numerics, d t must stay below this threshold.

3.4. Complex Dominant Pair: Great-Circle Limit Cycle and the Hierarchical Flag

In decreasing dimensions, the attractor and its separating sphere is the same problem one dimension down. Order the eigenvalues of V by strictly descending the real part and group them into blocks B 1 , , B p , each a single real eigenvalue (real eigenline E j , dim 1 ) or a single complex-conjugate pair (real invariant plane E j = span { Re v j , Im v j } , dim 2 ), with Re ( B 1 ) > > Re ( B p ) . Let W m = E 1 E m collect the m fastest blocks and F m = E m + 1 E p its V-invariant complement; by biorthogonality, F m is the joint kernel of the left eigen-functionals of B 1 , , B m . Set H m = F m S d 1 , H 0 = S d 1 .
Proposition 1
(Flag of invariant subspheres). For V diagonalisable with strictly separated real block parts, the H m form an invariant flag
S d 1 = H 0 H 1 H p 1 H p = , codim H m 1 H m = dim E m { 1 , 2 } ,
and the flow restricted to H m 1 has a unique attractor A m :
  • If B m is real, A m = { ± v m } , an antipodal pair of fixed points;
  • If B m is a complex pair α m ± i β m , A m = E m S d 1 is a great circle: a single periodic orbit of period 2 π / | β m | (uniform rotation when E m is orthonormal, a reparameterisation otherwise).
The basin of A m within H m 1 is the open dense stratum H m 1 H m ; A m is normally hyperbolic with transverse contraction governed by the real-part gap α m Re ( B m + 1 ) > 0 , and its unstable dimension equals the depth dim W m 1 = j < m dim E j . The strata { H m 1 H m } m = 1 p partition S d 1 into nested basins (an “onion”), each flowing to its A m .
Proof. 
F m is V-invariant, so H m is invariant (Lemma 3), and the left-functional description of F m together with Lemma 4 gives the sign-preservation nesting the strata. On H m 1 , the dominant surviving block is B m , strictly dominant in the real part, so e Re ( B m ) t e t V y 0 converges to the E m -component of y 0 (the argument of Theorem 2 applied to V | F m 1 ): a point ± v m if B m is real, the rotating E m -component if complex, whose projectivisation is the great circle. The index is the deflation count of faster blocks; normal hyperbolicity follows from the spectral gaps. □
Remark 6
(Limit cycles are exact great circles, not Hopf cycles). The periodic orbits A m = E m S d 1 are exact invariant great circles, present for every V carrying the pair, at full amplitude and with rotation number fixed by β m = | Im λ m | (Figure 3). They are not small-amplitude cycles born from a Hopf bifurcation: there is no amplitude parameter to grow, since the circle is the projectivisation of the V-invariant eigenplane and is one-dimensional by construction at all parameter values. What does vary is the circle’s stability, set entirely by where Re λ m sits in the real-part order— A m is the level-m attractor when B m is the dominant survivor, and a saddle or in-plane repeller otherwise. A “bifurcation” is thus an exchange of dominance (the global attractor switching between a point-pair and a circle as two real parts cross), not the nucleation of a cycle from a fixed point. Two complex pairs with equal real parts are the degenerate exception: their joint four-plane projectivises to an invariant S 3 with a two-frequency linear flow (invariant tori); strict separation excludes this.
Remark 7
(The symmetric case). When V = V all blocks are real, every A m is a point-pair, the eigenvectors are orthogonal ( w m = v m ), and the flow is the gradient ascent of the Rayleigh quotient f ( x ) = 1 2 x , V x : a Morse function on S d 1 with critical pairs ± v k , values 1 2 λ k , and Morse index = depth equals the number of { j : λ j > λ k } (Figure 4). The onion is the orthogonal nested-equator decomposition. For non-symmetric V, the gradient/Morse structure is lost exactly when a complex block appears: A m becomes a Morse–Bott circle carrying rotation and f ceases to be monotone, but the inclusions and the index-by-depth bookkeeping survive unchanged.

4. Multi-Token, No MLP

Section 3 followed a single cluster. We now release the n tokens to move independently under the value matrix ( Φ 0 ). Softmax attention couples them through configuration-dependent convex means, but cannot override V: when λ 1 > 0 , the tokens collapse to ± v 1 with an attention-independent consensus rate, positive hemispheres and cones are trapping, and the one genuinely attention-dependent question, whether two antipodal clusters coexist, is decided by the single scalar s = v 1 A v 1 . We use the terminology consistently: clustering (equivalently, collapse) means convergence of the tokens to a common point of S d 1 , whereas spreading means growth of the inter-token fluctuations; the two are governed, respectively, by the operators L col and L fluct introduced in Section 6.

4.1. Collapse to the Dominant Eigenvector

Theorem 3
(Multi-token attractor, conditional). Under the hypotheses of Theorem 1, consider the n-token system x ˙ i = P x i j w i j V x j near the collapsed state x i = v 1 . The linearised dynamics split into:
  • A mean mode ξ ¯ ˙ = ( P v 1 V P v 1 λ 1 I ) ξ ¯ = L col ξ ¯ , with spectrum { λ j λ 1 } j 1 (stable, by Lemma 2);
  • Fluctuation modes ζ ˙ i = L fluct ζ i , with L fluct = λ 1 I .
Consequently, if λ 1 > 0 the fluctuations decay at rate λ 1 and the tokens collapse to v 1 , which is an attractor of the full system; if λ 1 < 0 , the fluctuations grow and the tokens spread while the mean remains at v 1 .
Proof. 
Write x i = v 1 + ξ i , ξ i T v 1 S d 1 , and split ξ i = ξ ¯ + ζ i into mean and fluctuation ( ξ ¯ = 1 n i ξ i , i ζ i = 0 ). Since Φ 0 , the force is S i = j w i j V x j . At first order, the attention weights satisfy j w i j V x j = V v 1 + V ξ ¯ + O ( ξ 2 ) : writing w i j = 1 n + O ( ξ ) , the O ( ξ ) correction to the weights multiplies V v 1 (an i-independent vector) and sums to zero after using j (correction) = 0 , while the 1 n part of the weights gives 1 n j V ( v 1 + ξ j ) = V v 1 + V ξ ¯ ; the fluctuations ζ j = ξ j ξ ¯ do not survive this averaging since j ζ j = 0 . Hence, S i = λ 1 v 1 + V ξ ¯ + O ( ξ 2 ) , independent of i to this order.
Expanding x ˙ i = P x i S i exactly as in the proof of Lemma 2 (replacing V x there by the common force S i , using P v 1 ( λ 1 v 1 ) = 0 and v 1 , ξ i = 0 ) yields:
ξ ˙ i = P v 1 ( V ξ ¯ ) λ 1 ξ i + O ( ξ 2 ) .
Averaging over i gives the mean equation ξ ¯ ˙ = ( P v 1 V P v 1 λ 1 I ) ξ ¯ = L col ξ ¯ , with spectrum { λ j λ 1 } j 1 by Lemma 2, hence stable. Subtracting the mean equation from the ξ i equation, the P v 1 ( V ξ ¯ ) term cancels (it is i-independent) and only the curvature term survives:
ζ ˙ i = λ 1 ζ i + O ( ξ 2 ) = : L fluct ζ i , L fluct = λ 1 I .
The sign of λ 1 decides collapse versus spreading. □
Remark 8
(Relation to the general mechanism). Theorem 3 is the MLP-free instance of Proposition 4. Collapse needs λ 1 > 0 , the special case of η > 0 with η = λ 1 > 0 (for V with λ 1 > 0 and V 0 on the tangent space). The spreading regime λ 1 < 0 (a negative dominant eigenvalue) is the MLP-free realisation of Proposition 4(ii): one-cluster stable, multi-token unstable.

4.2. Trapping and Drain

Throughout this subsection, Φ 0 and we write the n-token system as
x ˙ i = P x i S i , S i : = j w i j V x j , ρ i : = x i , S i ,
with W = ( w i j ) , the row-stochastic attention matrix ( w i j > 0 , j w i j = 1 ), and D ρ = diag ( ρ 1 , , ρ n ) . Both W and D ρ depend on x i ; no claim below assumes them constant.
  • The left-eigenvector reduction
Lemma 5
(Scalar reduction). Let w k be a left eigenvector of V, V w k = λ k w k , and set p i ( k ) : = w k , x i , p ( k ) = ( p 1 ( k ) , , p n ( k ) ) . Then, along (8),
p ˙ i ( k ) = λ k j w i j p j ( k ) ρ i p i ( k ) , i . e . , p ˙ ( k ) = λ k W D ρ p ( k ) .
Proof. 
By differentiating p i ( k ) = w k , x i and using P x i S i = S i x i , S i x i ,
p ˙ i ( k ) = w k , P x i S i = w k , S i x i , S i w k , x i = w k , S i ρ i p i ( k ) .
Since w k is a left eigenvector, w k , V x j = V w k , x j = λ k w k , x j = λ k p j ( k ) , so w k , S i = j w i j w k , V x j = λ k j w i j p j ( k ) , which gives (9). □
The reduction is exact: the value-matrix vector dynamics become one scalar consensus-with-growth equation per token, per left eigenvector. Two consequences follow: every positive-eigenvalue coordinate is trapping (Proposition 2), but only the dominant one is basin-separating (made precise below), while every subdominant trap merely confines a coordinate that itself converges to zero (Proposition 3).
  • Trapping
Proposition 2
(Hemisphere and cone trapping). If λ k > 0 is real, then λ k W D ρ has positive off-diagonal entries λ k w i j > 0 , and the closed orthants { p ( k ) 0 } and { p ( k ) 0 } are forward-invariant: if all tokens start on one side of the separator w k , they remain there. Writing K + = { k : λ k > 0 real } , for any sign vector ( ϵ k ) k K + { ± 1 } K + , the polyhedral cone
K ϵ = k K + { x : ϵ k p i ( k ) 0 i }
(all tokens jointly on prescribed sides of every positive-eigenvalue separator) is forward-invariant. This holds for every real positive eigenvalue, not only the dominant one.
Proof. 
By Lemma 5, p ˙ ( k ) = M k ( t ) p ( k ) with M k ( t ) = λ k W ( t ) D ρ ( t ) , a closed linear (time-varying) equation; its off-diagonal entries λ k w i j are positive when λ k > 0 . On the boundary face { p i ( k ) = 0 , p j ( k ) 0 ( j i ) } , the inward normal is e i and p ˙ i ( k ) = j i λ k w i j p j ( k ) 0 , so the field satisfies Nagumo’s subtangentiality condition and the closed orthant is forward-invariant; { p ( k ) 0 } follows by oddness. The cone K ϵ is a finite intersection of such invariant sets, hence invariant. (Individual token signs are not conserved: a lone token may be pulled across w k by its neighbours, but the joint orthant, all tokens on one side, is.) □
  • Drain: only the dominant trap is a basin separator
Since w k , v 1 = δ k 1 , the two attractors ± v 1 lie in opposite hemispheres of w 1 ( w 1 , ± v 1 = ± 1 ) but on every subdominant separator ( w k , ± v 1 = 0 , k 2 ). The dominant coordinate is therefore marginal; it defines the basins, while every subdominant one decays, as the next proposition makes quantitative (see also Figure 5).
Proposition 3
(Subdominant drain). Let λ k ( k 2 ) be real with λ k < λ 1 , and p i ( k ) = w k , x i = O ( ξ ) near x i = v 1 (since w k , v 1 = 0 ). The linearisation of (9) at the collapsed configuration is
p ˙ ( k ) = λ k n 1 1 λ 1 I p ( k ) + O ( ξ 2 ) .
Here, 1 = ( 1 , , 1 ) R n , so that 1 1 is the n × n all-ones matrix and 1 n 1 1 is the averaging (barycentre) projector onto the consensus direction 1 . The operator in (10) has eigenvalues λ k λ 1 on 1 (the consensus direction) and λ 1 on 1 . For λ 1 > 0 , both are negative, so p ( k ) 0 : the consensus part drains at the spectral gap rate λ 1 λ k and the fluctuation part at λ 1 —matching the isotropic rate of Theorem 3, since p i ( k ) = w k , ζ i here and ζ ˙ i = λ 1 ζ i already gives the second rate directly. Hence, the collective set { x i w k } is attracting (the multi-token shadow of the stable manifold of the saddle v k ), not a basin separator: which attractor is reached is fixed by sign w 1 , x , transverse to the w k -constraint, so a subdominant positive eigenvalue confines a dynamically irrelevant coordinate, not a competing cluster. Only k = 1 gives a marginal coordinate, whose consensus direction has an eigenvalue of 0 and is therefore basin-defining.
Proof. 
At collapse W = 1 n 1 1 and ρ i = λ 1 + O ( ξ ) , the O ( ξ ) correction is i-independent to this order, by the same computation as in the proof of Theorem 3: S i = λ 1 v 1 + V ξ ¯ + O ( ξ 2 ) gives ρ i = v 1 + ξ i , S i = λ 1 + v 1 , V ξ ¯ + O ( ξ 2 ) , using ξ i , v 1 = 0 ). Hence, λ k W D ρ = λ k n 1 1 λ 1 I + O ( ξ ) , giving (10); since p ( k ) = O ( ξ ) , the neglected terms are O ( ξ 2 ) . The rank-one matrix 1 n 1 1 has an eigenvalue of 1 on 1 and 0 on 1 , so the eigenvalues of (10) are λ k λ 1 (on 1 ) and λ 1 (on 1 ); both are negative when λ 1 > 0 > λ k λ 1 . Setting k = 1 replaces the first eigenvalue by λ 1 λ 1 = 0 . For k = 1 , however, p ( 1 ) is O ( 1 ) rather than O ( ξ ) , so the linearisation above is not applied to it: the vanishing rate λ 1 λ 1 = 0 only records that this direction has no exponential decay (it is marginal), while the sign of the dominant coordinate is preserved exactly and independently by the orthant invariance of Proposition 2. The asymptotic stability of ± v 1 is provided by the fluctuation and mean modes of Theorem 3, not by this zero. For a complex subdominant pair, the two real functionals Re w k and Im w k give a coupled pair of coordinates with the same gap-rate decay of λ 1 Re λ k , together with rotation at | Im λ k | . □
Remark 9
(Two boundary cases). λ k < 0 makes λ k w i j < 0 , breaking the positivity of the off-diagonal entries, so no trapping (the spreading regime); a complex pair with a positive real part does not trap either, since the real reduction of w k , x i carries a rotation (frequency | Im λ k | ) that destroys orthant invariance. By contrast, in the one-cluster flow d d t w k , x = ( λ k x , V x ) , w k , x is sign-preserving for every real eigenvalue regardless of sign—the pure invariant-subspace structure of the invariant-subspace lattice (Supplementary Material), with no off-diagonal positivity hypothesis, because there is no inter-token coupling.

4.3. The Role of Attention: Numerical Results

The preceding results show that, analytically, the attention matrix A cannot override V: the field is x ˙ i = P x i ( V m i ) , with m i = j w i j x j , a convex mean, so A only reweights the value force (Remark 10). We now report what this means quantitatively, established numerically.
Remark 10
(Clustering dichotomy: when a cluster forms). Three mutually exclusive cases, governed entirely by the dominant spectrum of V:
  • λ 1 > 0 real and dominant. The collapsed configuration is fluctuation-stable ( L fluct = λ 1 I 0 ) and tokens cluster, settling on + v 1 or v 1 by sign w 1 , x (Theorem 3).
  • λ 1 < 0 . Then, L fluct = | λ 1 | I 0 : every transverse direction grows, so no cluster forms—tokens spread regardless of initial data, even within a single hemisphere { w 1 , x > 0 } , since that constraint is transverse to the fluctuation mode that decides coherence.
  • Complex dominant pair. No point cluster exists: the configuration is carried onto the great-circle limit cycle of the dominant eigenplane (Proposition 1, Remark 6).
This is the pure value-matrix statement; with a u-aligned MLP, the fluctuation operator becomes P v 1 D Φ ( v 1 ) P v 1 λ 1 I , so a sufficiently expansive Φ can hold a cluster together even when λ 1 < 0 —the L col versus L fluct tension of Proposition 4.
  • Collapse is robust to A .
Writing S i = V m i as above, attention enters only through the convex weights m i : for symmetric V with λ 1 > 0 and generic one-hemisphere initial data, tokens collapse to the single cluster v 1 for every attention matrix tested: A = I , symmetric positive-definite, symmetric indefinite, strongly non-symmetric, and negative-definite, as well as across temperatures β (Figure 6a). The only stable clusters are ± v 1 . A tight cluster placed at a subdominant eigenvector v k ( k 2 , λ k > 0 ) stays internally coherent—its fluctuations decay at rate λ k —but its centre sits at a saddle of the one-cluster flow (the unstable direction λ 1 λ k > 0 , Proposition 1), so the cluster migrates to ± v 1 rather than persisting at v k . The genuine multi-cluster state is therefore the antipodal pair { + v 1 , v 1 } (Figure 6b,c).
  • Coexistence versus merger: the order parameter s = v 1 Av 1
Whether two groups occupying the opposite basins coexist or merge is governed by a single scalar,
s : = v 1 A v 1 ,
with the alignment of attention along the cluster axis. The leading cross-cluster logit between a + v 1 token and a v 1 token is β v 1 A ( v 1 ) = β s , so the clusters coexist when within-cluster attention dominates ( s > 0 ) and merge when cross-cluster attention dominates ( s < 0 ), the transition lying at s = 0 (Figure 6d). Since the antisymmetric part of A drops out of the quadratic form, s = v 1 A sym v 1 and a non-symmetric A obey the same law—the relevant case, as a trained attention matrix A = W Q W K is generically non-symmetric and indefinite; negative-definiteness is merely one route to s < 0 , not a distinct mechanism. At large β, the s < 0 side becomes metastable: a + v 1 token attending to the v 1 cluster feels P + v 1 ( λ 1 ( v 1 ) ) = 0 , a vanishing radial force, so the merger stalls; the clean transition is the diffuse-attention limit.

5. Diagonal MLP

We now switch on the u-aligned MLP Φ together with a symmetric value matrix V. This is the regime in which the one-cluster flow acquires a Lyapunov function and the uncollapsed dynamics inherit multi-token stability automatically; both rest on D Φ being symmetric. The gradient, Hessian, and linearisation identities used throughout are collected for reference in the Supplementary Material.

5.1. Lyapunov Structure and One-Cluster Convergence

  • One-cluster ODE
When all tokens coincide, x i ( t ) = x ( t ) for all i, the weights are uniform ( w i j = 1 / n ), and (1) reduces to
x ˙ = P x ( V x + Φ ( x ) ) .
Theorem 4
(Lyapunov function). For symmetric V and u-aligned Φ as above, the ODE (11) is the Riemannian gradient flow on S d 1 of
L ( x ) = 1 2 x , V x + k = 1 r γ k 0 u k , x σ ( α k t + b k ) d t .
Along every trajectory, L ˙ = x ˙ 2 0 , with equality if x is a fixed point. By LaSalle’s invariance principle, every trajectory converges to a fixed point. Periodic orbits and chaotic dynamics are impossible.
Proof. 
The construction proceeds in six steps.
  • Step 1 ( D Φ is symmetric). The u-aligned MLP is Φ ( x ) = k = 1 r γ k σ ( α k u k , x + b k ) u k . Its Jacobian is:
                D Φ ( x ) i j = Φ i x j ( x ) = k = 1 r γ k σ ( α k s k + b k ) · α k · ( u k ) j · ( u k ) i = k = 1 r γ k α k σ ( α k s k + b k ) · ( u k ) i ( u k ) j ,
where s k = u k , x = j ( u k ) j x j . In matrix form:
D Φ ( x ) = k = 1 r γ k α k σ ( α k s k + b k ) u k u k .
Since each term u k u k is symmetric and the coefficients γ k α k σ ( α k s k + b k ) R are scalars, D Φ ( x ) is symmetric for all x.
  • Step 2 ( Φ = g ). Define
g ( x ) = k = 1 r γ k 0 u k , x σ ( α k t + b k ) d t .
We verify g ( x ) = Φ ( x ) by differentiating with respect to x j :
g x j ( x ) = k = 1 r γ k · x j 0 u k , x σ ( α k t + b k ) d t = k = 1 r γ k · σ ( α k u k , x + b k ) · ( u k ) j = Φ j ( x ) ,
using j u k , x = ( u k ) j ). Hence, Φ = g .
  • Step 3 ( V x + Φ ( x ) = L ( x ) ).
With f V ( x ) = 1 2 x , V x and V = V :
f V ( x ) = V x .
Indeed,
x j ( 1 2 i , k V i k x i x k ) = 1 2 k V j k x k + 1 2 i V i j x i = k V j k x k = ( V x ) j
using V = V .
Therefore,
L = f V + g = V x + Φ ( x ) .
  • Step 4 (The ODE is a Riemannian gradient flow). We use the following criterion. A vector field of the form F ( x ) = P x h ( x ) on S d 1 (with h : R d R d ) is a Riemannian gradient flow of some potential if its covariant Jacobian P x D h ( x ) P x is self-adjoint on T x S d 1 for every x; this is the integrability (symmetric-Hessian) condition for a gradient on a Riemannian manifold, and for h = L it is automatic, but we verify it directly from the model. Here, h ( x ) = V x + Φ ( x ) , so D h ( x ) = V + D Φ ( x ) and the covariant Jacobian is P x ( V + D Φ ( x ) ) P x . Since V = V and D Φ ( x ) = D Φ ( x ) (Step 1), the sum V + D Φ ( x ) is symmetric, so the compression P x ( V + D Φ ( x ) ) P x is self-adjoint on T x S d 1 . The criterion is therefore met at every x, confirming x ˙ = P x ( V x + Φ ( x ) ) = grad S L ( x ) .
  • Step 5 ( L ˙ = x ˙ 2 ). Along a trajectory of the ODE, set w = L ( x ) = V x + Φ ( x ) and split it into its tangential and normal parts at x,
w = P x w + x , w x , P x w = grad S L ( x ) T x S d 1 ,
with the tangential part being, by definition, the Riemannian gradient. Using x ˙ = P x w together with (19) and bilinearity,
d d t L ( x ( t ) ) = L ( x ) , x ˙ = w , P x w = P x w + x , w x , P x w = grad S L , P x w + x , w x , P x w .
The second term vanishes because P x w = grad S L T x S d 1 is orthogonal to x, i.e., x , P x w = 0 (equivalently, the ODE preserves x   =   1 , so x , x ˙ = 0 ). Therefore:
L ˙   =   grad S L ( x ) , x ˙ = grad S L ( x ) , grad S L ( x ) = grad S L ( x ) 2 = x ˙ 2 0 .
  • Step 6 (Convergence). L is continuous on the compact manifold S d 1 , hence bounded. L ˙ = x ˙ 2 0 , so L is non-decreasing along trajectories. By LaSalle’s invariance principle, every trajectory converges to the largest invariant set contained in { L ˙ = 0 } = { x ˙   =   0 } = { x : grad S L ( x ) = 0 } , which is the set of critical points of L on S d 1 .
Since S d 1 is compact and L is bounded and non-decreasing, the ω-limit set of every trajectory is non-empty and invariant, hence contained in the critical point set. Convergence to a single critical point follows when the critical points are isolated, which holds for generic ( V , Φ ) . When they are not isolated (for instance, when V has a repeated eigenvalue, so that the critical set contains submanifolds rather than isolated points), L is a Morse–Bott function along the gradient flow, and every trajectory still converges to a single connected component of the critical set (the ω-limit set of a gradient flow of an analytic, or Morse–Bott, potential is a single critical component). The convergence conclusion is therefore unchanged; only the limit may be a critical manifold rather than a point.
Periodic orbits are excluded because L is strictly increasing along any non-constant trajectory. □
Corollary 1
(Fixed points and stability). The fixed points of (11) are the critical points of L on S d 1 : solutions of V x + Φ ( x ) = μ x . When the u k are eigenvectors of V with eigenvalues ν k , this decouples as
ν k s k + γ k σ ( α k s k + b k ) = μ s k , k = 1 , , r , s 1 2 + . . . + s r 2 = 1 .
A critical point x is a stable attractor if it has a strict local maximum of L (Riemannian Hessian negative definite; derived in the Supplementary Material); it is an unstable saddle if the Hessian has at least one positive eigenvalue. The global attractor (from generic initial conditions) is the global maximum of L.
Remark 11
(Role of V in attractor selection). The global maximum of L = f V + g balances the Rayleigh quotient of V (favouring eigenvectors of V with large eigenvalues) against the MLP potential g (favouring directions where g is large). Increasing ν k shifts the global maximum toward u k : a large ν k makes S d 1 tilt in the u k direction. The value matrix thus participates directly in attractor selection, in contrast to the attention matrix A, which drives the collapse to a single cluster but does not select among critical points.

5.2. Uncollapsed Token Dynamics

Proposition 4 below shows that for general V and Φ, a stable one-cluster fixed point need not be stable for the multi-token system. We now show that the u-aligned MLP with symmetric V is a special case where this instability is excluded: the gradient flow structure automatically guarantees multi-token stability at every attractor.
We consider n tokens near a critical point x of L: x i = x + ξ i , decomposed as ξ i = ξ ¯ + ζ i (mean plus fluctuation, i ζ i = 0 ).
  • Mean mode
At first order in ξ : j w i j V x j V ( x + ξ ¯ ) and Φ ( x i ) Φ ( x ) + D Φ ( x ) ( ξ ¯ + ζ i ) . Averaging over i and using i ζ i = 0 :
ξ ¯ ˙ = P x ( 2 L ( x ) ξ ¯ ) + O ( ξ 2 )   =   ( Hess S L | x ) ξ ¯ + O ( ξ 2 ) .
At a local maximum of L, Hess S L | x is negative definite, so ξ ¯ 0 : the cluster mean converges to x . At a saddle, the mean grows in the unstable Hessian direction, driving the cluster toward the global maximum of L.
  • Fluctuation modes
Since Φ acts independently on each token and i ζ i = 0 , the MLP contributes P x ( D Φ ( x ) ζ i ) to ζ ˙ i and the curvature correction contributes μ ζ i (the same normal-force term as in L col ); the softmax attention contributes nothing at first order (Supplementary Material). Combined,
ζ ˙ i = P x D Φ ( x ) ζ i MLP     μ ζ i Curvature + O ( ξ 2 ) = L fluct ζ i + O ( ξ 2 ) .
Writing Hess S g | x = P x D Φ ( x ) P x μ Φ I for the Riemannian Hessian of the MLP potential and using μ = μ V + μ Φ , the operator is L fluct = Hess S g | x μ V I . At a local maximum of L, one has Hess S g | x ( μ V η ) I (because Hess S L η I and P x V P x 0 for V 0 ), hence L fluct η I 0 : all fluctuation eigenvalues are negative near a local maximum of L.
Theorem 5
(Uncollapsed token dynamics). Let x be a critical point of L on S d 1 .
(i) 
(Collapse.) If x is a local maximum of L, both mean and fluctuation modes decay. Every trajectory of (1) near ( x , , x ) converges to ( x , , x ) : tokens collapse to the stable attractor.
(ii) 
(Saddle escape.) Suppose x is a saddle of L and the fluctuation-stability condition
L fluct = P x D Φ ( x ) P x μ I 0 on T x S d 1
holds (equivalently, Hess S g | x μ V I ; a sufficient explicit form is max 1 k r γ k α k σ ( α k s k + b k ) + < μ , which, in particular, forces μ > 0 ). Then, the mean mode grows in the unstable Hessian direction, driving the cluster toward the global maximum of L, while every fluctuation mode decays: the tokens stay clustered while the cluster escapes the saddle. Condition (24) holds automatically at a local maximum (part (i)); at a saddle, it is a genuine restriction, and if it fails, the saddle drives the tokens apart instead of keeping them clustered.
Proof. 
Part (i): At a local maximum, Hess S L is negative definite, so both (22) and the fluctuation operator in (23) have all negative eigenvalues. Part (ii): The mean operator is L col = Hess S L | x and the fluctuation operator is L fluct = Hess S g | x μ V I . At a saddle, L col has an eigenvalue λ + > 0 with eigenvector v + T x S d 1 , so, by (22), the mean grows, ξ ¯ ( t ) e λ + t v + , and the cluster leaves the neighbourhood of x along v + toward the global maximum of L. Condition (24) makes L fluct negative definite, so, by (23), every fluctuation mode decays ζ i ( t ) 0 at an exponential rate bounded away from zero. Hence, the inter-token spread max i ζ i contracts while the common mean escapes: the tokens remain clustered throughout. If, instead, (24) fails, then v , L fluct v > 0 for some v T x S d 1 and that fluctuation direction grows—the saddle drives the tokens apart, which is the multi-token instability of Proposition 4(ii). □
Remark  12
(Why the u-aligned case resolves the generic instability). Proposition 4(ii) below identifies a potential conflict: V may stabilise the one-cluster ODE without stabilising the fluctuations. The u-aligned MLP with symmetric V avoids this by construction. Since V and D Φ ( x ) share eigenvectors u k , the operators differ only by L fluct L col = P x V P x , so for symmetric positive semi-definite V the eigenvalue correction is q V ( v ) 0 , so λ fluct ( v ) λ col ( v ) < 0 . The gradient flow structure ensures that λ col ( v ) < 0 everywhere, and since the correction is negative, λ fluct ( v ) < λ col ( v ) < 0 : fluctuation modes are more stable than collapsed modes.
  • Numerical verification of the u -aligned MLP results
We use d = 64 , r = 2 , n = 10 , β = 2 , A = 2 u 1 u 1 , γ k = 1.5 , α k = 2 , b k = 0 , tanh activation, and step size d t = 0.05 . Figure 7 shows six experiments with two parameter regimes: (i) stable regime with ν c = ( 5 , 0.1 ) , global max of L near u 1 ; (ii) saddle regime using ν s = ( 0.5 , 5 ) , for which x c u 1 is a saddle of L s and u 2 is the global maximum.
  • Panels (a) and (b): stable attractor
Starting near the global max x c u 1 under L c : L increases monotonically and ζ F 0 (panel (a)). Starting near the same point under L s (where it is a saddle): L s still increases monotonically and ζ F 0 , but | u 2 , x ¯ | 1 (panel (b)). The cluster escapes the saddle while staying clustered, confirming Theorem 5(ii).
  • Panel (c): Token positions
Projected onto ( u 1 , u 2 ) : stable-regime tokens (green) all converge to u 1 ; saddle-escape tokens (red) converge to u 2 . In both cases, tokens form a single cluster at the final position.
  • Panel (d): Attractor shift
Varying ν 1 with ν 2 = 5 fixed: the global maximum transitions from u 2 -dominated (small ν 1 ) to u 1 -dominated (large ν 1 ), with a smooth crossover near ν 1 ν 2 = 5 . This confirms Remark 11.
  • Panel (e): L monotone for all  ν 1
From a common random initialisation, L increases monotonically to different attractor values for five values of ν 1 , confirming Theorem 4 uniformly.
  • Panel (f): Robustness
Six perturbation sizes near the stable x all collapse at machine precision, confirming Theorem 5(i).

6. One-Cluster vs. Multi-Token: The General Mechanism

We establish a general relationship between one-cluster stability and multi-token stability for arbitrary V and Φ.
  • Linearisation at a fixed point
Let x be a fixed point: P x ( V x + Φ ( x ) ) = 0 , i.e., V x + Φ ( x ) = μ x with μ = x , V x + Φ ( x ) . Write μ V = x , V x for the normal component of the value matrix. For standard softmax attention, the linearised attention force is independent of the attention matrix and common to all tokens at first order, so it acts only on the mean and contributes no first-order term to the fluctuations (this is established self-containedly in Proposition 5 below). The clustering of fluctuations is therefore carried entirely by the curvature term μ V , and the relevant linearisation operators on T x S d 1 are as follows (derived below in (39)):
L col = P x ( V + D Φ ( x ) ) P x μ I ,
L fluct = P x D Φ ( x ) P x μ I ,
where both operators carry the same curvature term μ I , μ = μ V + μ Φ , arising from the normal force μ x at the collapsed fixed point. Note that the tangential action of V (the form q V ) is absent from L fluct : the value matrix term j w i j V x j V x ¯ contributes only to the mean mode, not to fluctuations at first order. (V’s normal component μ V is present in both operators through μ and cancels when they are compared.)
  • Relationship
Throughout this comparison, we write λ col ( v ) = v , L col v , λ fluct ( v ) = v , L fluct v , and q V ( v ) = v , P x V v for the quadratic forms (Rayleigh quotients) of the linearisation operators on the tangent space. Since L col L fluct = P x V P x , these satisfy
λ fluct ( v ) = λ col ( v ) q V ( v )
as an identity of quadratic forms, valid for arbitrary (possibly non-symmetric) V.
For a symmetric operator, the Rayleigh quotient attains the eigenvalues at its stationary values, so max v λ col ( v ) is the top eigenvalue and the sign of the form decides stability. For a non-symmetric L col (non-symmetric V), this fails: max v λ col ( v ) is the numerical abscissa (the top eigenvalue of the symmetric part 1 2 ( L col + L col ) ), an upper bound for the spectral abscissa max j Re spec ( L col ) that can be strictly larger. Consequently, the collapse criterion of Proposition 4(i) is a sufficient condition: it forces the numerical abscissa of L fluct below zero, hence giving monotone (not merely asymptotic) decay. The exact one-cluster spectrum for non-symmetric V is obtained spectrally by the deflation Lemma 2 ( spec ( A k ) = { λ j λ k } ), not from these quadratic forms; we use the form-level relation (27) only for the sufficient comparison below.
Proposition 4
(One-cluster vs. multi-token stability). Let η = max v λ col ( v ) > 0 be the one-cluster stability margin (in the numerical-range sense above).
(i) 
(Sufficient condition for collapse.) If V is positive semi-definite on T x S d 1 ( q V ( v ) 0 for all v x ), then λ fluct ( v ) η < 0 for every unit v x : all fluctuation eigenvalues are negative and the tokens collapse to x . No further condition is required.
(ii) 
(Possible instability.) For general V and Φ, there exist configurations where λ col ( v ) < 0 (one-cluster stable) but λ fluct ( v ) > 0 (tokens spread). This occurs when V is contractive in direction v (stabilising the mean but not the fluctuations), while D Φ ( x ) is expansive in v.
Proof. 
The eigenvalue relation. For any v T x S d 1 with v   =   1 , write q V ( v ) = v , P x V v and q D Φ ( v ) = v , P x D Φ ( x ) v . From operators (25) and (26):
λ col ( v ) = q V ( v ) + q D Φ ( v ) μ ,
λ fluct ( v ) = q D Φ ( v ) μ .
Subtracting:
λ fluct ( v ) λ col ( v ) = q V ( v ) .
This is the key relation. It shows that the tangential action of V, encoded in the quadratic form q V ( v ) , enters λ col but not λ fluct . (The normal component μ V enters both operators equally through μ = μ V + μ Φ and therefore cancels in the difference (30).) For positive semi-definite V ( q V 0 ), the fluctuation mode is more stable than the collapsed mode: λ fluct ( v ) λ col ( v ) . For indefinite V, q V ( v ) < 0 is possible; either effect can make λ fluct ( v ) > λ col ( v ) , so that fluctuations are less stable than the mean and potentially unstable even when λ col < 0 .
Part (i) Sufficient condition. We want to show that λ fluct ( v ) < 0 for all v x , given that x is one-cluster stable ( η = max v λ col ( v ) > 0 ) and V is positive semi-definite on T x S d 1 .
From (30):
λ fluct ( v ) = λ col ( v ) q V ( v ) .
Since V 0 on T x S d 1 , we have q V ( v ) 0 . From the one-cluster stability margin, λ col ( v ) η for all units v x . Substituting into (31) and using q V ( v ) 0 yields:
λ fluct ( v ) = λ col ( v ) q V ( v ) η = η < 0 ,
the last inequality by the hypothesis η > 0 . This holds for any V 0 and any η , C with η > 0 . The value matrix enters the fluctuation eigenvalue only through the stabilising q V ( v ) 0 ; the only role of the attention is through the bound, involving neither λ max ( V ) nor μ V .
Part (ii) (An explicit instability). We construct explicit parameters. Let d = 2 , x = e 1 = ( 1 , 0 ) , T x S 1 = span ( e 2 ) . Set V = κ e 2 e 2 (contractive on T x S 1 , q V ( e 2 ) = κ < 0 , so V is indefinite), D Φ ( x ) e 2 = ϕ e 2 with ϕ > C (MLP expansive), and μ V = e 1 , V e 1 = 0 .
Then, writing μ = μ V + μ Φ = μ Φ (since μ V = 0 ), with μ Φ = x , Φ ( x ) the normal component of the MLP map from the Supplementary Material (not the Jacobian normal component x , D Φ ( x ) x ), yields:
λ col ( e 2 ) = q V ( e 2 ) + q D Φ ( e 2 ) μ = κ + ϕ μ Φ ,
λ fluct ( e 2 ) = q D Φ ( e 2 ) μ = ϕ μ Φ ,
at the fixed point V e 1 = 0 , so Φ ( e 1 ) = μ x = μ Φ e 1 ; choosing the MLP so that Φ ( e 1 ) = 0 gives μ Φ = 0 . This is consistent with D Φ ( e 1 ) e 2 = ϕ e 2 0 : the map vanishes at e 1 while its Jacobian there is expansive in the tangent direction. Then,
λ col ( e 2 ) = κ + ϕ ,
λ fluct ( e 2 ) = ϕ .
Choose C < ϕ < κ (possible since C and ϕ can be chosen freely with ϕ > C , then κ > ϕ ):
λ col ( e 2 ) = ϕ κ < 0 ( stable ) , λ fluct ( e 2 ) = ϕ > 0 ( unstable ) .
This confirms that a stable one-cluster fixed point can be multi-token unstable. The mechanism: V is contractive on the tangent space ( κ < 0 ) providing one-cluster stability, but does not appear in λ fluct , which is governed by ϕ > 0 . □
Remark 13
(The curvature term is the full μ ). Both L col and L fluct contain the same curvature term μ I with the full μ = μ V + μ Φ , originating from the normal force μ x at the collapsed fixed point (see (25) and (26) and the linearisation in the Supplementary Material). In particular, the curvature term is not μ Φ I (which would omit the normal component of the value matrix) and not μ D Φ I (the Jacobian normal component, a different object, cf. the Supplementary Material). Because μ V appears identically in both eigenvalues, it cancels in the difference (30), leaving λ fluct ( v ) = λ col ( v ) q V ( v ) , with no residual μ V .
Remark 14
(Interpretation). The key asymmetry is that the tangential action of V (the form q V ) appears in L col but not in L fluct . A value matrix that stabilises the one-cluster ODE by being contractive in some tangential direction v provides no corresponding stability to the fluctuation modes in that direction. The fluctuation stability depends entirely on D Φ ( x ) (and, in the MLP-free case, on the sign of μ V ). This asymmetry is the source of the potential discrepancy between one-cluster and multi-token stability for general V and Φ.
  • The attention force and the linearised operators (main-text derivation)
We close the section by recording the derivation of (25) and (26) and the proof that softmax attention contributes no first-order clustering force. Place n tokens at x i = x + ξ i with ξ i T x S d 1 , and split ξ i = ξ ¯ + ζ i , ξ ¯ = 1 n i ξ i , i ζ i = 0 .
Proposition 5
(No first-order attention clustering force). For standard softmax attention with any coupling matrix A, the attention-weighted value force satisfies
j w i j V x j = V x + V ξ ¯ + O ( ξ 2 )
independent of the token index i to first order. Hence, the attention force drives only the mean mode ξ ¯ and contributes no first-order term to the fluctuations ζ i .
Proof. 
Linearising the score β x i , A x j / d about x i = x j = x gives, after softmax normalisation,
w i j = 1 n + β n d ( A x ) ( ξ j ξ ¯ ) + O ( ξ 2 ) ,
in which the query deviation ξ i cancels in the softmax denominator, so the first-order correction depends on the key index j only. Hence,
j w i j V x j = 1 n j V ( x + ξ j ) + j w i j 1 n V x + O ( ξ 2 ) .
The first sum is V x + V ξ ¯ ; in the second, j ( w i j 1 / n ) = 0 (the rows of W sum to one) multiplies the i-independent vector V x to zero, while the correction acting on V ξ j is O ( ξ 2 ) . This gives (37), common to all i; subtracting the mean of the linearised Equation (39) removes this common term, so it contributes nothing to ζ ˙ i at first order. □
The linearised operators follow. Expanding x ˙ i = P x i j w i j V x j + Φ ( x i ) to first order, using (37), Φ ( x i ) = Φ ( x ) + D Φ ( x ) ξ i + O ( ξ 2 ) , V x + Φ ( x ) = μ x , and P x + ξ i = P x ξ i x x ξ i + O ( ξ 2 ) , gives
ξ ˙ i = P x V ξ ¯ + D Φ ( x ) ξ i μ ξ i + O ( ξ 2 ) .
The value matrix acts on the mean ξ ¯ only, whereas the MLP acts on the full deviation ξ i = ξ ¯ + ζ i . Averaging (39) over i gives ξ ¯ ˙ = ( P x ( V + D Φ ( x ) ) P x μ I ) ξ ¯ = L col ξ ¯ , which is (25); subtracting the mean equation, the common term P x V ξ ¯ cancels and ζ ˙ i = ( P x D Φ ( x ) P x μ I ) ζ i = L fluct ζ i , which is (26). The value matrix is thus absent from L fluct : it enters only through the common mean.

7. Discussion

The paper establishes two results at different levels of generality.
  • Summary
The first (Proposition 4) is a general structural observation: for arbitrary V and Φ, a stable fixed point of the one-cluster ODE need not be stable for the multi-token system. The asymmetry arises because V enters the mean operator L col but not the fluctuation operator L fluct (Section 6): the value matrix can stabilise the mean without stabilising the fluctuations, and a spreading instability occurs when V is contractive in some tangential direction while D Φ ( x ) is expansive there. The second result (Theorems 4 and 5) shows that the u-aligned MLP with symmetric V is the exactly solvable case in which this instability is structurally excluded: the one-cluster flow is the gradient ascent of L = f V + g , every trajectory converges to a critical point, and the tokens collapse to a single cluster at every stable attractor, escaping saddles while remaining clustered.
  • The value matrix, not the attention, selects the cluster
A recurring conclusion of the analysis is that the value matrix governs where the tokens collapse, while the attention governs only that they collapse and how fast. For softmax attention, the first-order clustering force vanishes (Proposition 5), so the consensus rate is set by the dominant eigenvalue λ 1 of V and is independent of the attention matrix A (Theorem 3), and the selected direction is the dominant eigenvector v 1 , tilted off the geometric equator through the left eigenvector w 1 when V is non-symmetric (Theorem 2). The attention matrix re-enters only in the multi-cluster question: whether two clusters occupying opposite basins coexist or merge is decided by the single scalar s = v 1 A v 1 (Section 4.3). This gives a sharp, finite-n complement to the mean-field picture of [2,3,4], in which the limiting configuration is likewise governed by the spectrum of the value matrix; here, the basin structure is described explicitly and holds without a mean-field limit. In the same direction, the mean-field analysis of the feed-forward block [14] shows that its critical points are generically atomic and localised on the sphere, in agreement with the value-matrix and MLP selection described here. The two viewpoints are complementary: the mean-field results characterise where the critical points lie as n , whereas our finite-n analysis identifies which of them is reached from a given initial condition, at what rate, and by what mechanism. A rigorous mean-field limit of the present finite-n statements is left as an open problem.
  • Collapse, spreading, and representation degeneracy
The sign of the dominant eigenvalue separates two qualitatively different behaviours. For λ 1 > 0 , the tokens collapse to a point, the dynamical counterpart of the rank collapses, and oversmoothing or token uniformity is observed in deep transformers [5,6]. For λ 1 < 0 , the one-cluster state is stable but nearby tokens spread rather than collapse, and a complex dominant pair produces no point cluster at all. In this reduced model, the value matrix thus carries a lever over representation degeneracy: a value matrix whose dominant eigenvalue is small, negative, or complex slows or prevents the collapse to a rank-one representation, in the same spirit as the architectural mechanisms (skip connections, feed-forward blocks, normalisation) known to mitigate it [5,6].
  • Beyond the gradient case: limit cycles
When V is symmetric, the flow is the gradient ascent of the Rayleigh quotient and the dynamics are of Morse type, with an antipodal pair of attractors and a hierarchy of equatorial saddles indexed by spectral depth (Remark 7). This gradient structure is lost precisely when a complex block becomes dominant: the point attractor is then replaced by an exact invariant great circle carried by the dominant eigenplane (Proposition 1, Remark 6). These cycles are not born from a Hopf bifurcation: they are present at full amplitude for every V carrying the complex pair, and a change of the global attractor between a point-pair and a circle is an exchange of dominance as two real parts cross. Non-normality of V leaves the asymptotic selection unchanged but permits transient growth and a spiralling approach (Remark 2), and the discrete-time iteration adds a genuine rotation threshold absent from the ODE (Remark 5). Even the value-matrix-only dynamics can therefore fail to converge to a point, a mechanism for persistent oscillation of the token representations.
  • Scope and limitations
The model is deliberately reduced. It uses a single value matrix and a single coupling matrix rather than multi-head, multi-layer weights, tied input/output directions in the u-aligned MLP, and sphere normalisation in place of LayerNorm; the Lyapunov structure requires V = V and a u-aligned Φ, and the multi-token collapse and trapping statements are local, obtained by linearisation about the collapsed state. The vanishing of the first-order attention force is specific to standard softmax normalisation: a different attention mechanism could contribute a genuine first-order term, which would sharpen collapse when contractive but could not by itself produce the spreading instabilities, which arise through V. Within these restrictions, the results are exact and hold for a finite particle number, without mean-field approximation.
  • Modeling scope
The reductions are chosen to retain the mechanism under study while keeping the dynamics analytically tractable, and each preserves a specific feature of the transformer forward pass. Sphere normalisation retains the norm constraint that LayerNorm imposes and is the natural geometry for the projected ODE. The single coupling matrix A carries the query–key score geometry A = W Q W K (and the effective A eff = h W Q ( h ) W K ( h ) for multi-head attention), and the value/output map, absent from pure-attention models, is exactly the object V = W O W V added here. The u-aligned MLP is the tractable case in which the Lyapunov structure is exact; the general dense MLP, whose Jacobian need not be symmetric, is left as an open problem (below). These reductions omit causal masking, positional encodings, and layer-dependent weights, so the model is not a quantitative surrogate for a trained transformer; rather, the phenomena it exhibits (collapse driven by attention, location selected by the value matrix, and the collapse/spreading dichotomy set by sign λ 1 ) are the reduced-model counterparts of the rank collapse, oversmoothing, and token uniformity reported for deep transformers [5,6], and are offered as mechanisms rather than as claims about production models.
  • Open problems
Several directions remain. (i) The general (dense) MLP, where D Φ ( x ) need not be symmetric, so that the spreading instability of Proposition 4(ii) can genuinely occur and a general feed-forward block may produce limit cycles on S 2 . (ii) The multi-cluster regime, where the cluster locations depend jointly on V and A through a self-consistency condition, and the order parameter s = v 1 A v 1 governing coexistence should generalise. (iii) Deeper spectra: a higher-dimensional hierarchical flag and the associated 2 p lattice of invariant subspheres, defective or equal-real-part spectra (invariant tori), and the recursive basin structure on the separating subsphere. (iv) The relation to training, where V and A are learned and the attractor structure evolves, and a rigorous mean-field limit of the finite-n results obtained here.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/math14183350/s1. Full derivations and complete proofs of all statements are provided in the accompanying supplement.

Funding

The research was carried out within the state assignment (project no. FSSF-2026-0013).

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the author.

Acknowledgments

The author has been supported by the RUDN University Strategic Academic Leadership Program. During the preparation of this manuscript, the author used AI tools to assist with language editing, improve the presentation of the manuscript, and discuss possible mathematical formulations and interpretations. All scientific results, mathematical derivations, analyses, and conclusions were developed and verified by the author. The author critically reviewed all AI-generated suggestions and assumes full responsibility for the content of the manuscript.

Conflicts of Interest

The author declares no conflict of interest.

References

  1. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Long Beach, CA, USA, 4–9 December 2017; Volume 30, pp. 5998–6008. [Google Scholar]
  2. Geshkovski, B.; Letrouit, C.; Polyanskiy, Y.; Rigollet, P. The emergence of clusters in self-attention dynamics. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), New Orleans, LA, USA, 10–16 December 2023; Volume 36, pp. 57026–57037. [Google Scholar] [CrossRef] [Scilit]
  3. Geshkovski, B.; Letrouit, C.; Polyanskiy, Y.; Rigollet, P. A mathematical perspective on transformers. Bull. Am. Math. Soc. 2025, 62, 427–479. [Google Scholar] [CrossRef] [Scilit]
  4. Rigollet, P. The mean-field dynamics of transformers. arXiv 2025, arXiv:2512.01868. [Google Scholar]
  5. Dong, Y.; Cordonnier, J.-B.; Loukas, A. Attention is not all you need: Pure attention loses rank doubly exponentially with depth. In Proceedings of the 38th International Conference on Machine Learning (ICML), Virtual, 18–24 July 2021; Volume 139, pp. 2793–2803. [Google Scholar]
  6. Wu, X.; Ajorlou, A.; Wang, Y.; Jegelka, S.; Jadbabaie, A. On the role of attention masks and LayerNorm in transformers. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 10–15 December 2024; Volume 37, pp. 14774–14809. [Google Scholar] [CrossRef] [Scilit]
  7. Kuramoto, Y. Chemical Oscillations, Waves, and Turbulence; Springer: Berlin/Heidelberg, Germany, 1984. [Google Scholar]
  8. Vicsek, T.; Czirók, A.; Ben-Jacob, E.; Cohen, I.; Shochet, O. Novel type of phase transition in a system of self-driven particles. Phys. Rev. Lett. 1995, 75, 1226–1229. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Cucker, F.; Smale, S. Emergent behavior in flocks. IEEE Trans. Autom. Control 2007, 52, 852–862. [Google Scholar] [CrossRef] [Scilit]
  10. Hegselmann, R.; Krause, U. Opinion dynamics and bounded confidence: Models, analysis and simulation. J. Artif. Soc. Soc. Simul. 2002, 5, 1–33. [Google Scholar]
  11. Chen, R.T.Q.; Rubanova, Y.; Bettencourt, J.; Duvenaud, D. Neural ordinary differential equations. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Montréal, QC, Canada, 3–8 December 2018; Volume 31, pp. 6572–6583. [Google Scholar]
  12. Sander, M.E.; Ablin, P.; Blondel, M.; Peyré, G. Sinkformers: Transformers with doubly stochastic attention. In Proceedings of the 25th International Conference on Artificial Intelligence and Statistics (AISTATS), Virtual, 28–30 March 2022; Volume 151, pp. 3515–3530. [Google Scholar]
  13. Karagodin, N.; Polyanskiy, Y.; Rigollet, P. Clustering in causal attention masking. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 10–15 December 2024; Volume 37, pp. 115652–115681. [Google Scholar] [CrossRef] [Scilit]
  14. Álvarez-López, A.; Geshkovski, B.; Ruiz-Balet, D. Perceptrons and localization of attention’s mean-field landscape. arXiv 2026, arXiv:2601.21366. [Google Scholar]
  15. Yu, Y.; Buchanan, S.; Pai, D.; Chu, T.; Wu, Z.; Tong, S.; Bai, X.; Zhai, Y.; Haeffele, B.D.; Ma, Y. White-box transformers via sparse rate reduction: Compression is all there is? J. Mach. Learn. Res. 2024, 25, 1–128. [Google Scholar]
  16. Golub, G.H.; Van Loan, C.F. Matrix Computations, 4th ed.; Johns Hopkins University Press: Baltimore, MD, USA, 2013. [Google Scholar]
Figure 1. Almost-global bistability for a non-symmetric V with real dominant eigenvalue ( d = 3 , spec V = { 2 , 1 , 1 } , tilt 28 . 8 ). Six hundred one-cluster initial conditions are coloured by their limit: every trajectory converges to v 1 (blue initial conditions) or v 1 (red initial conditions) according to sign w 1 , x 0 (left eigenvector w 1 ), and sign w 1 , x is conserved along each trajectory (both 100 % ). The basin boundary is the invariant great subsphere w 1 (solid black), tilted off the geometric equator v 1 (dashed green). Verifies Theorem 2 and Lemma 4.
Figure 1. Almost-global bistability for a non-symmetric V with real dominant eigenvalue ( d = 3 , spec V = { 2 , 1 , 1 } , tilt 28 . 8 ). Six hundred one-cluster initial conditions are coloured by their limit: every trajectory converges to v 1 (blue initial conditions) or v 1 (red initial conditions) according to sign w 1 , x 0 (left eigenvector w 1 ), and sign w 1 , x is conserved along each trajectory (both 100 % ). The basin boundary is the invariant great subsphere w 1 (solid black), tilted off the geometric equator v 1 (dashed green). Verifies Theorem 2 and Lemma 4.
Mathematics 14 03350 g001
Figure 2. Discrete-time rotation threshold for spec V = { 1 , ± 3 i } . The explicit map x ( I + h V ) x / ( I + h V ) x converges to the dominant real eigenvector v 1 (alignment | x , v 1 | 1 ) only for h < 1 4 ; above the threshold, the subdominant complex pair wins the modulus competition and the iterate rotates. Measured threshold h 0.245 (grid spacing 0.005 ), matching the predicted 1 4 . The ODE is the h 0 limit and has no such threshold. Verifies Remark 5.
Figure 2. Discrete-time rotation threshold for spec V = { 1 , ± 3 i } . The explicit map x ( I + h V ) x / ( I + h V ) x converges to the dominant real eigenvector v 1 (alignment | x , v 1 | 1 ) only for h < 1 4 ; above the threshold, the subdominant complex pair wins the modulus competition and the iterate rotates. Measured threshold h 0.245 (grid spacing 0.005 ), matching the predicted 1 4 . The ODE is the h 0 limit and has no such threshold. Verifies Remark 5.
Mathematics 14 03350 g002
Figure 3. Great-circle limit cycle from a complex dominant pair ( d = 3 , spec V = { 0.5 ± 2 i , 0.6 } ). (Left): A trajectory spirals onto the invariant great circle E S 2 , the projectivisation of the dominant eigenplane. (Right): The distance to that plane decays exponentially, and the measured period π matches the prediction 2 π / | Im λ | to integrator accuracy ( 0 % error). The circle is an exact invariant object, not a Hopf-born cycle. Verifies Proposition 1 and Remark 6.
Figure 3. Great-circle limit cycle from a complex dominant pair ( d = 3 , spec V = { 0.5 ± 2 i , 0.6 } ). (Left): A trajectory spirals onto the invariant great circle E S 2 , the projectivisation of the dominant eigenplane. (Right): The distance to that plane decays exponentially, and the measured period π matches the prediction 2 π / | Im λ | to integrator accuracy ( 0 % error). The circle is an exact invariant object, not a Hopf-born cycle. Verifies Proposition 1 and Remark 6.
Mathematics 14 03350 g003
Figure 4. Hierarchical flag for symmetric V ( d = 4 , spec V = { 3 , 2 , 1 , 1 } ). (Left): Along a generic trajectory, the Rayleigh quotient x , V x increases monotonically, confirming that it is a Lyapunov/Morse function. (Right): The measured unstable index of each eigenvector v k equals the depth # { j : λ j > λ k } (here, 0 , 1 , 2 , 3 ); the deflation growth rates λ j λ k are recovered to within 0.05 . Verifies Proposition 1 and Remark 7.
Figure 4. Hierarchical flag for symmetric V ( d = 4 , spec V = { 3 , 2 , 1 , 1 } ). (Left): Along a generic trajectory, the Rayleigh quotient x , V x increases monotonically, confirming that it is a Lyapunov/Morse function. (Right): The measured unstable index of each eigenvector v k equals the depth # { j : λ j > λ k } (here, 0 , 1 , 2 , 3 ); the deflation growth rates λ j λ k are recovered to within 0.05 . Verifies Proposition 1 and Remark 7.
Mathematics 14 03350 g004
Figure 5. Hemisphere/cone trapping in the multi-token system ( n = 10 , softmax A = I , β = 4 , spec V = { 1.5 , 0.5 , 0.5 } , two positive eigenvalues). (Left): All tokens initialised in the cone { p i ( 1 ) > 0 ,   p i ( 2 ) > 0 i } keep both coordinates non-negative for all t (joint-orthant invariance, Proposition 2). (Right): The subdominant coordinate max i | p i ( 2 ) | converges to zero at the gap rate λ 1 λ 2 = 1 , confirming that w 2 is attracting, not basin-separating (Proposition 3).
Figure 5. Hemisphere/cone trapping in the multi-token system ( n = 10 , softmax A = I , β = 4 , spec V = { 1.5 , 0.5 , 0.5 } , two positive eigenvalues). (Left): All tokens initialised in the cone { p i ( 1 ) > 0 ,   p i ( 2 ) > 0 i } keep both coordinates non-negative for all t (joint-orthant invariance, Proposition 2). (Right): The subdominant coordinate max i | p i ( 2 ) | converges to zero at the gap rate λ 1 λ 2 = 1 , confirming that w 2 is attracting, not basin-separating (Proposition 3).
Mathematics 14 03350 g005
Figure 6. Clustering and the attention matrix ( n = 14 , symmetric V, spec V = { 1 , 0.3 , 0.2 } , d = 3 ; coordinates rotated so v 1 is the north pole, grey dots initial, red dots final, stars ± v 1 ). (a) From generic one-hemisphere data, the tokens collapse to a single cluster at v 1 for every attention class and temperature—collapse is governed by V, not A. (b,c) Two groups seeded in the opposite basins ± v 1 : they either coexist as an antipodal two-cluster state (b) or merge into one (c). (d) The outcome is set by the single order parameter s = v 1 A v 1 : coexistence for s > 0 , merge for s < 0 , transition at s = 0 . A non-symmetric A falls on the same curve (orange), since its antisymmetric part drops out of s; negative-definite A is merely one route to s < 0 . A cluster placed at a subdominant v k is not shown: it is a saddle and migrates to ± v 1 .
Figure 6. Clustering and the attention matrix ( n = 14 , symmetric V, spec V = { 1 , 0.3 , 0.2 } , d = 3 ; coordinates rotated so v 1 is the north pole, grey dots initial, red dots final, stars ± v 1 ). (a) From generic one-hemisphere data, the tokens collapse to a single cluster at v 1 for every attention class and temperature—collapse is governed by V, not A. (b,c) Two groups seeded in the opposite basins ± v 1 : they either coexist as an antipodal two-cluster state (b) or merge into one (c). (d) The outcome is set by the single order parameter s = v 1 A v 1 : coexistence for s > 0 , merge for s < 0 , transition at s = 0 . A non-symmetric A falls on the same curve (orange), since its antisymmetric part drops out of s; negative-definite A is merely one route to s < 0 . A cluster placed at a subdominant v k is not shown: it is a saddle and migrates to ± v 1 .
Mathematics 14 03350 g006
Figure 7. Numerical verification of the u-aligned MLP results ( d = 64 , n = 10 , r = 2 , β = 2 , A = 2 u 1 u 1 , γ k = 1.5 , α k = 2 , b k = 0 , tanh activation, d t = 0.05 ). Green marks the stable regime ν c = ( 5 , 0.1 ) , whose global maximum of L lies near u 1 ; red marks the saddle-escape regime ν s = ( 0.5 , 5 ) , for which u 1 is a saddle of L and u 2 the global maximum. (a) Stable regime: starting near x c u 1 , the Lyapunov function L increases monotonically to its maximum (green) while the inter-token spread ζ F decays to machine precision (grey); the tokens collapse to u 1 . (b) Saddle-escape regime: starting near the saddle u 1 , L s still increases monotonically and the spread ζ F decays to machine precision (grey), while the cluster mean rotates to the global maximum, | u 2 , x ¯ | 1 (violet): the tokens escape the saddle yet stay clustered, confirming Theorem 5(ii). (c) Final token positions projected on ( u 1 , u 2 ) : the stable-regime tokens (green) form a single cluster at u 1 and the saddle-escape tokens (red) a single cluster at u 2 (stars mark u 1 , u 2 ; the red arc traces the escaping mean). (d) Attractor shift: varying ν 1 with ν 2 = 5 fixed, the u 1 -weight of the global maximiser of L moves from 0 ( u 2 -dominated) to 1 ( u 1 -dominated), the crossover lying at ν 1 ν 2 (dashed), confirming Remark 11. (e) From a common random initialisation, L increases monotonically to distinct attractor values for five values of ν 1 , confirming Theorem 4 uniformly. (f) Robustness: for six perturbation sizes about the stable attractor, the spread collapses to machine precision, confirming Theorem 5(i).
Figure 7. Numerical verification of the u-aligned MLP results ( d = 64 , n = 10 , r = 2 , β = 2 , A = 2 u 1 u 1 , γ k = 1.5 , α k = 2 , b k = 0 , tanh activation, d t = 0.05 ). Green marks the stable regime ν c = ( 5 , 0.1 ) , whose global maximum of L lies near u 1 ; red marks the saddle-escape regime ν s = ( 0.5 , 5 ) , for which u 1 is a saddle of L and u 2 the global maximum. (a) Stable regime: starting near x c u 1 , the Lyapunov function L increases monotonically to its maximum (green) while the inter-token spread ζ F decays to machine precision (grey); the tokens collapse to u 1 . (b) Saddle-escape regime: starting near the saddle u 1 , L s still increases monotonically and the spread ζ F decays to machine precision (grey), while the cluster mean rotates to the global maximum, | u 2 , x ¯ | 1 (violet): the tokens escape the saddle yet stay clustered, confirming Theorem 5(ii). (c) Final token positions projected on ( u 1 , u 2 ) : the stable-regime tokens (green) form a single cluster at u 1 and the saddle-escape tokens (red) a single cluster at u 2 (stars mark u 1 , u 2 ; the red arc traces the escaping mean). (d) Attractor shift: varying ν 1 with ν 2 = 5 fixed, the u 1 -weight of the global maximiser of L moves from 0 ( u 2 -dominated) to 1 ( u 1 -dominated), the crossover lying at ν 1 ν 2 (dashed), confirming Remark 11. (e) From a common random initialisation, L increases monotonically to distinct attractor values for five values of ν 1 , confirming Theorem 4 uniformly. (f) Robustness: for six perturbation sizes about the stable attractor, the spread collapses to machine precision, confirming Theorem 5(i).
Mathematics 14 03350 g007
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Volpert, V. Value Matrix Dynamics and Attractor Selection in a Reduced Transformer Model. Mathematics 2026, 14, 3350. https://doi.org/10.3390/math14183350

AMA Style

Volpert V. Value Matrix Dynamics and Attractor Selection in a Reduced Transformer Model. Mathematics. 2026; 14(18):3350. https://doi.org/10.3390/math14183350

Chicago/Turabian Style

Volpert, Vitaly. 2026. "Value Matrix Dynamics and Attractor Selection in a Reduced Transformer Model" Mathematics 14, no. 18: 3350. https://doi.org/10.3390/math14183350

APA Style

Volpert, V. (2026). Value Matrix Dynamics and Attractor Selection in a Reduced Transformer Model. Mathematics, 14(18), 3350. https://doi.org/10.3390/math14183350

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop