2. Model
Let
n tokens
evolve in continuous time. Fix
orthonormal directions
and write
. The continuous-time ODE (the Euler limit of the discrete normalised residual update; cf. the neural-ODE view of residual networks [
11]) is:
where
are attention weights,
A is a coupling matrix, and
. Throughout,
denotes the Euclidean norm on
and
the associated inner product, so
is the unit sphere in this norm.
We take
symmetric with eigenvectors
and eigenvalues
. Its Rayleigh quotient is
. In a full transformer, the attention output is
, so the value/output map
acts on the values inside the attention-weighted sum; in the reduced model (
1), the same
V acts linearly on each token before the attention-weighted average.
We consider a particular form of feed-forward block, in which the input and output directions are tied to the same orthonormal feature basis
. This
u-aligned MLP is
where
and
σ are continuously differentiable (e.g., tanh). The name refers to this alignment with the basis
, not to coordinatewise action in the standard basis; we avoid the term “diagonal MLP”, which would misleadingly suggest the latter. The same orthonormal directions
are used to read each feature
and to write Φ back (tied input/output weights). Each component
therefore depends on all coordinates of x through the
, so Φ is not coordinatewise in the standard basis; what matters is that its Jacobian,
is symmetric (diagonal in the
u-basis, with eigenvectors
and eigenvalues
). It is this symmetry, not standard-basis diagonality, that the Lyapunov structure of
Section 5 relies on. The associated scalar potential is
The results below are illustrated by direct numerical integration of the continuous-time system (
1) with an explicit Runge–Kutta scheme, and, for the one-cluster map, by the explicit Euler–retraction iteration
; the dimension
d, token number
n, inverse temperature
β, spectrum of
V, and step size are stated in each figure caption. Ensembles of initial conditions are drawn uniformly on
, and the number of initial conditions behind each reported statistic is given in the corresponding caption (for example, 600 one-cluster initial conditions in
Figure 1). Two groups seeded in opposite basins are declared to
coexist when their inter-cluster geodesic distance stays bounded away from zero over the integration horizon, and to
merge otherwise; this coincides with the sign of the order parameter
of
Section 4.3.
3. One Cluster Without MLP
Setting isolates the value matrix. The resulting one-cluster flow admits a sharp, fully explicit analysis that requires no symmetry of V: we locate its fixed points, identify the dominant eigenvector as the selected attractor, show that the selection is almost-global, and organise the remaining dynamics into a hierarchical flag culminating in great-circle limit cycles.
3.1. Fixed Points for the One-Cluster Flow
With
, the one-cluster ODE is
Lemma 1 (Fixed points are real eigenvectors)
. is a fixed point of (
5)
iff for some real λ, necessarily . In particular, complex eigenvalues of V produce no fixed point on . Proof. means , i.e., is an eigenvector with eigenvalue . Conversely, an eigenvector with real eigenvalue gives . A complex eigenvalue has no real unit eigenvector in . □
Lemma 2 (Linearisation)
. At a fixed point with , the linearisation of (
5)
on is and its spectrum is , where are the eigenvalues of V. This holds for arbitrary (possibly non-symmetric) V.
Proof. Differentiating
at
in a direction
(so
) gives:
Now,
, and
. The remaining inner term
is parallel to
, hence normal to
; the tangential generator is its projection,
which is (
6); the asymmetric part of
V drops out through
. For the spectrum, complete
to an orthonormal basis
of
, where
is an arbitrary orthonormal basis of
(not eigenvectors of
V which,
V being non-symmetric, need be neither orthogonal nor real). Block-triangularity below uses only that
is an eigenvector. Write
for the tangent block (so
,
and
) and
(orthogonal); using
, the similarity
has first column
, hence
with
generally nonzero (the asymmetry, lying in the normal direction) and
the lower-right block. Thus,
, read off the characteristic polynomial without diagonalising
M. Finally,
is exactly the matrix of the compression
in the basis
: for
,
using
for
(the off-diagonal
is the discarded normal component). Hence,
and
. □
A single observation underlies the global picture of the next subsections: the flow is the radial projection of a linear flow.
Lemma 3 (Projectivised flow). For any , let and . Then, solves , and conversely, every solution of the latter is the radial projection of a ray of .
Proof. Differentiating
and using
and
,
Conversely, given a solution
on
, the scalar ODE
reconstructs
, and
solves
. □
Lemmas 1–3 are elementary and classical. Lemma 3 exhibits the one-cluster flow as the radial projection (projectivisation) of the linear flow
, and the convergence of that projection to the dominant eigendirection is the continuous-time counterpart of the power method [
16]. We include self-contained proofs for completeness; the contribution of this section is not these lemmas individually but their assembly into the almost-global attractor-selection picture, the left-eigenvector separator, and the hierarchical flag developed in the following subsections.
3.2. The Dominant-Eigenvector Attractor
Theorem 1 (One-cluster attractor)
. Suppose is a simple eigenvalue of V with a strictly maximal real part, for every other eigenvalue of V, and let be the corresponding unit eigenvector. Then, is locally asymptotically stable for (
5)
. Every other real eigenvector () is unstable, with unstable dimensions equal to the number of . Proof. By Lemma 2, . Strict dominance gives for every , so is a possibly non-normal matrix with eigenvalues with negative real parts, and is locally asymptotically stable; non-normality permits transient growth but not loss of asymptotic stability. At any other real eigenvector , contains (and more generally one positive-real eigenvalue for each with ), so is unstable with the stated index. □
Remark 1 (Complex dominant eigenvalue ⇒ no point attractor). If the eigenvalue of the maximal real part is a complex pair, no real eigenvector carries it, so by Theorem 1, every real-eigenvector fixed point has at least one unstable direction and none are stable. On this forces, by the Poincaré–Bendixson theorem, a periodic orbit; this is the limit-cycle mechanism of the asymmetric-V analysis. The clean statement that the top eigenvalue is the unique attractor is thus exactly the real-dominant-eigenvalue case.
3.3. Almost-Global Bistability and the Tilted Separator
Informally, almost-global bistability states that, outside a measure-zero set of initial data, every trajectory ends at one of the two poles . The reason is transparent from the projectivised-flow viewpoint of Lemma 3: writing in the eigenbasis, the solution is , and the dominant mode outgrows all others whenever , so the limit is determined by the sign of alone. The exceptional set is exactly , a great subsphere of one lower dimension. Two features are specific to a non-symmetric V: the selector is the left eigenvector , so the separating subsphere is rather than the geometric equator , and is tilted off it; further, the sign of is exactly conserved along the flow, so the two basins never mix. The following lemma and theorem make this precise.
Lemma 4 (Invariant separating sphere)
. Let be the left eigenvector of V for (). Along any solution of (
5)
,Hence, ; the sign of is conserved, and the great subsphere is invariant. Proof. Using
and
,
which is (
7). The integrating factor is positive, so
never changes sign, and
forces
. □
The use of the left eigenvector is essential: it is , not , that satisfies , so only the coordinate decouples. By biorthogonality , the invariant subspace equals (the complement of in the eigenbasis), not the geometric complement ; the two agree if .
Theorem 2 (Almost-global bistability). Under the hypotheses of Theorem 1, assume V is diagonalisable (Remark 3 removes this) and normalise . Then are asymptotically stable, and for every with , the ω-limit set is the single point . The basins are the open half-spheres , which exhaust ; their common boundary is the invariant subsphere H, which contains every fixed point other than together with all of their stable manifolds, and has an empty interior. No globally attracting point exists.
Proof. Write
in the (right) eigenbasis, with
by biorthogonality. By Lemma 3,
. If
,
since every other mode carries a factor
(the imaginary parts only rotate the shrinking components; the approach may spiral but
). Since
is a scalar and normalisation is scale-invariant, we may insert this factor in both numerator and denominator without changing
:
the magnitude
cancelling so that only the sign of the (real) dominant coefficient survives. Hence
, with
constant by Lemma 4. In particular, no
off
H can lie on the stable manifold of any other fixed point: such a fixed point is a real eigenvector
(
) with
, and convergence to it would require
, impossible when
is conserved in sign. All other fixed points and their stable manifolds therefore lie in
; concretely, they form the heteroclinic skeleton of the deflated dynamics on
H (Remark below). Local asymptotic stability of
is Theorem 1; oddness of
gives two disjoint basins, neither the whole sphere, ruling out a global attractor. □
Remark 2 (Mode of approach). Convergence to the point () need not be monotone: when the subdominant eigenvalues are complex, the trajectory spirals into , winding infinitely while the distance decays like . This is still .
Remark 3 (Non-diagonalisable V). The diagonalisability hypothesis is only a convenience. If some subdominant eigenvalue is defective, carries Jordan terms , but since strictly we still have , so and Theorem 2 hold verbatim, with the coordinate read off by the left eigenvector. Only itself must be simple.
Remark 4 (Phase portrait and the separating sphere). On the flow is, by Lemma 3, the projectivisation of restricted to the invariant subspace , on which V acts as the deflated block M with spectrum (Lemma 2). Thus, the dynamics on are the same problem one dimension down, and the structure of the basin boundary follows by recursion on the subdominant spectrum:
If the eigenvalue of M of the maximal real part is real and strictly dominant, H carries its own antipodal pair of attractors, and the remaining real eigenvectors are lower equatorial saddles (their indices given by deflation), forming a heteroclinic skeleton;
If it is a complex pair, H has no equatorial point at its top; for (so ), Poincaré–Bendixson forces H to be a single periodic orbit, a limit cycle bounding the two basins.
For non-symmetric V, the subdominant right eigenvectors () need not lie on the geometric equator; relative to they sit at , hence on H, so they are exactly the equatorial unstable fixed points.
Remark 5 (Discrete-time threshold)
. The distinction is between the real part and the modulus: the ODE selects the eigenvalue of the largest real part
, whereas the discrete map, being power iteration, selects the eigenvalue of the largest modulus
. These two orderings can disagree at finite step h, and the threshold below is exactly where they do. The explicit Euler–retraction scheme is power iteration for ,
and converges to the eigenvector of maximal modulus . Since ,so the leading term reproduces the ODE Lyapunov exponent , while the correction adds a spurious favouring rotation. A subdominant complex pair can therefore win the modulus competition at finite step. For , , so the discrete map converges to only for and otherwise rotates (Figure 2). As , the real-part ordering is restored, recovering the ODE
result of Theorem 2; in numerics, must stay below this threshold. 3.4. Complex Dominant Pair: Great-Circle Limit Cycle and the Hierarchical Flag
In decreasing dimensions, the attractor and its separating sphere is the same problem one dimension down. Order the eigenvalues of V by strictly descending the real part and group them into blocks , each a single real eigenvalue (real eigenline , ) or a single complex-conjugate pair (real invariant plane , ), with . Let collect the m fastest blocks and its V-invariant complement; by biorthogonality, is the joint kernel of the left eigen-functionals of . Set , .
Proposition 1 (Flag of invariant subspheres)
. For V diagonalisable with strictly separated real block parts, the form an invariant flagand the flow restricted to has a unique attractor :If is real, , an antipodal pair of fixed points;
If is a complex pair , is a great circle: a single periodic orbit of period (uniform rotation when is orthonormal, a reparameterisation otherwise).
The basin of within is the open dense stratum ; is normally hyperbolic with transverse contraction governed by the real-part gap , and its unstable dimension equals the depth . The strata partition into nested basins (an “onion”), each flowing to its .
Proof. is V-invariant, so is invariant (Lemma 3), and the left-functional description of together with Lemma 4 gives the sign-preservation nesting the strata. On , the dominant surviving block is , strictly dominant in the real part, so converges to the -component of (the argument of Theorem 2 applied to ): a point if is real, the rotating -component if complex, whose projectivisation is the great circle. The index is the deflation count of faster blocks; normal hyperbolicity follows from the spectral gaps. □
Remark 6 (Limit cycles are exact great circles, not Hopf cycles)
. The periodic orbits are exact invariant great circles,
present for every
V carrying the pair, at full amplitude and with rotation number fixed by (Figure 3). They are not small-amplitude cycles born from a Hopf bifurcation: there is no amplitude parameter to grow, since the circle is the projectivisation of the V-invariant eigenplane and is one-dimensional by construction at all parameter values. What does
vary is the circle’s stability,
set entirely by where sits in the real-part order— is the level-m attractor when is the dominant survivor, and a saddle or in-plane repeller otherwise. A “bifurcation” is thus an exchange of dominance
(the global attractor switching between a point-pair and a circle as two real parts cross), not the nucleation of a cycle from a fixed point. Two complex pairs with equal
real parts are the degenerate exception: their joint four-plane projectivises to an invariant with a two-frequency linear flow (invariant tori); strict separation excludes this. Remark 7 (The symmetric case)
. When all blocks are real, every is a point-pair, the eigenvectors are orthogonal (), and the flow is the gradient ascent of the Rayleigh quotient : a Morse function on with critical pairs , values , and Morse index = depth equals the number of (Figure 4). The onion is the orthogonal nested-equator decomposition. For non-symmetric V, the gradient/Morse structure is lost exactly when a complex block appears: becomes a Morse–Bott circle carrying rotation and f ceases to be monotone, but the inclusions and the index-by-depth bookkeeping survive unchanged. 4. Multi-Token, No MLP
Section 3 followed a single cluster. We now release the n tokens to move independently under the value matrix (
). Softmax attention couples them through configuration-dependent convex means, but cannot override
V: when
, the tokens collapse to
with an attention-independent consensus rate, positive hemispheres and cones are trapping, and the one genuinely attention-dependent question, whether two antipodal clusters coexist, is decided by the single scalar
. We use the terminology consistently: clustering (equivalently, collapse) means convergence of the tokens to a common point of
, whereas spreading means growth of the inter-token fluctuations; the two are governed, respectively, by the operators
and
introduced in
Section 6.
4.1. Collapse to the Dominant Eigenvector
Theorem 3 (Multi-token attractor, conditional). Under the hypotheses of Theorem 1, consider the n-token system near the collapsed state . The linearised dynamics split into:
A mean mode , with spectrum (stable, by Lemma 2);
Fluctuation modes , with .
Consequently, if the fluctuations decay at rate and the tokens collapse to , which is an attractor of the full system; if , the fluctuations grow and the tokens spread while the mean remains at .
Proof. Write , , and split into mean and fluctuation (, ). Since , the force is . At first order, the attention weights satisfy : writing , the correction to the weights multiplies (an i-independent vector) and sums to zero after using (correction), while the part of the weights gives ; the fluctuations do not survive this averaging since . Hence, , independent of i to this order.
Expanding
exactly as in the proof of Lemma 2 (replacing
there by the common force
, using
and
) yields:
Averaging over
i gives the mean equation
, with spectrum
by Lemma 2, hence stable. Subtracting the mean equation from the
equation, the
term cancels (it is
i-independent) and only the curvature term survives:
The sign of
decides collapse versus spreading. □
Remark 8 (Relation to the general mechanism). Theorem 3 is the MLP-free instance of Proposition 4. Collapse needs , the special case of with (for V with and on the tangent space). The spreading regime (a negative dominant eigenvalue) is the MLP-free realisation of Proposition 4(ii): one-cluster stable, multi-token unstable.
4.2. Trapping and Drain
Throughout this subsection,
and we write the n-token system as
with
, the row-stochastic attention matrix (
,
), and
. Both
W and
depend on
; no claim below assumes them constant.
Lemma 5 (Scalar reduction)
. Let be a left eigenvector of V, , and set , . Then, along (
8)
, Proof. By differentiating
and using
,
Since
is a
left eigenvector,
, so
, which gives (
9). □
The reduction is exact: the value-matrix vector dynamics become one scalar consensus-with-growth equation per token, per left eigenvector. Two consequences follow: every positive-eigenvalue coordinate is trapping (Proposition 2), but only the dominant one is basin-separating (made precise below), while every subdominant trap merely confines a coordinate that itself converges to zero (Proposition 3).
Proposition 2 (Hemisphere and cone trapping)
. If is real, then has positive off-diagonal entries , and the closed orthants and are forward-invariant: if all tokens start on one side of the separator , they remain there. Writing , for any sign vector , the polyhedral cone(all tokens jointly on prescribed sides of every positive-eigenvalue separator) is forward-invariant. This holds for every
real positive eigenvalue, not only the dominant one. Proof. By Lemma 5, with , a closed linear (time-varying) equation; its off-diagonal entries are positive when . On the boundary face , the inward normal is and , so the field satisfies Nagumo’s subtangentiality condition and the closed orthant is forward-invariant; follows by oddness. The cone is a finite intersection of such invariant sets, hence invariant. (Individual token signs are not conserved: a lone token may be pulled across by its neighbours, but the joint orthant, all tokens on one side, is.) □
Since
, the two attractors
lie in
opposite hemispheres of
(
) but
on every subdominant separator (
,
). The dominant coordinate is therefore marginal; it defines the basins, while every subdominant one decays, as the next proposition makes quantitative (see also
Figure 5).
Proposition 3 (Subdominant drain)
. Let () be real with , and near (since ). The linearisation of (
9)
at the collapsed configuration isHere, , so that is the all-ones matrix and is the averaging (barycentre) projector onto the consensus direction . The operator in (
10)
has eigenvalues – on (the consensus direction) and on . For , both are negative, so : the consensus part drains at the spectral gap rate – and the fluctuation part at —matching the isotropic rate of Theorem 3, since here and already gives the second rate directly. Hence, the collective set is attracting
(the multi-token shadow of the stable manifold of the saddle ), not
a basin separator: which attractor is reached is fixed by , transverse to the -constraint, so a subdominant positive eigenvalue confines a dynamically irrelevant coordinate, not a competing cluster. Only gives a marginal coordinate, whose consensus direction has an eigenvalue of 0
and is therefore basin-defining. Proof. At collapse
and
, the
correction is
i-independent to this order, by the same computation as in the proof of Theorem 3:
gives
, using
). Hence,
, giving (
10); since
, the neglected terms are
. The rank-one matrix
has an eigenvalue of 1 on
and 0 on
, so the eigenvalues of (
10) are
–
(on
) and
(on
); both are negative when
. Setting
replaces the first eigenvalue by
. For
, however,
is
rather than
, so the linearisation above is not applied to it: the vanishing rate
only records that this direction has no exponential decay (it is marginal), while the sign of the dominant coordinate is preserved exactly and independently by the orthant invariance of Proposition 2. The asymptotic stability of
is provided by the fluctuation and mean modes of Theorem 3, not by this zero. For a complex subdominant pair, the two real functionals
and
give a coupled pair of coordinates with the same gap-rate decay of
–
, together with rotation at
. □
Remark 9 (Two boundary cases)
. makes , breaking the positivity of the off-diagonal entries, so no trapping (the spreading regime); a complex pair with a positive real part does not trap either, since the real reduction of carries a rotation (frequency ) that destroys orthant invariance. By contrast, in the one-cluster flow is sign-preserving for every
real eigenvalue regardless of sign—the pure invariant-subspace structure of the invariant-subspace lattice (Supplementary Material), with no off-diagonal positivity hypothesis, because there is no inter-token coupling. 4.3. The Role of Attention: Numerical Results
The preceding results show that, analytically, the attention matrix A cannot override V: the field is , with , a convex mean, so A only reweights the value force (Remark 10). We now report what this means quantitatively, established numerically.
Remark 10 (Clustering dichotomy: when a cluster forms). Three mutually exclusive cases, governed entirely by the dominant spectrum of V:
real and dominant. The collapsed configuration is fluctuation-stable () and tokens cluster, settling on or by (Theorem 3).
. Then, : every transverse direction grows, so no cluster forms—tokens spread regardless of initial data, even within a single hemisphere , since that constraint is transverse to the fluctuation mode that decides coherence.
Complex dominant pair. No point cluster exists: the configuration is carried onto the great-circle limit cycle of the dominant eigenplane (Proposition 1, Remark 6).
This is the pure value-matrix statement; with a u-aligned MLP, the fluctuation operator becomes , so a sufficiently expansive Φ can hold a cluster together even when —the versus tension of Proposition 4.
Writing
as above, attention enters only through the convex weights
: for symmetric
V with
and generic one-hemisphere initial data, tokens collapse to the single cluster
for every attention matrix tested:
, symmetric positive-definite, symmetric indefinite, strongly non-symmetric, and negative-definite, as well as across temperatures
β (
Figure 6a). The only stable clusters are
. A tight cluster placed at a subdominant eigenvector
(
,
) stays internally coherent—its fluctuations decay at rate
—but its centre sits at a saddle of the one-cluster flow (the unstable direction
–
, Proposition 1), so the cluster migrates to
rather than persisting at
. The genuine multi-cluster state is therefore the antipodal pair
(
Figure 6b,c).
Whether two groups occupying the opposite basins coexist or merge is governed by a single scalar,
with the alignment of attention along the cluster axis. The leading cross-cluster logit between a
token and a
token is
, so the clusters coexist when within-cluster attention dominates (
) and merge when cross-cluster attention dominates (
), the transition lying at
(
Figure 6d). Since the antisymmetric part of A drops out of the quadratic form,
and a non-symmetric
A obey the same law—the relevant case, as a trained attention matrix
is generically non-symmetric and indefinite; negative-definiteness is merely one route to
, not a distinct mechanism. At large
β, the
side becomes metastable: a
token attending to the
cluster feels
, a vanishing radial force, so the merger stalls; the clean transition is the diffuse-attention limit.
5. Diagonal MLP
We now switch on the
u-aligned MLP Φ together with a symmetric value matrix
V. This is the regime in which the one-cluster flow acquires a Lyapunov function and the uncollapsed dynamics inherit multi-token stability automatically; both rest on
being symmetric. The gradient, Hessian, and linearisation identities used throughout are collected for reference in the
Supplementary Material.
5.1. Lyapunov Structure and One-Cluster Convergence
When all tokens coincide,
for all
i, the weights are uniform (
), and (
1) reduces to
Theorem 4 (Lyapunov function)
. For symmetric V and u-aligned Φ as above, the ODE (11) is the Riemannian gradient flow on ofAlong every trajectory, , with equality if x is a fixed point. By LaSalle’s invariance principle, every trajectory converges to a fixed point. Periodic orbits and chaotic dynamics are impossible. Proof. The construction proceeds in six steps.
where
. In matrix form:
Since each term
is symmetric and the coefficients
are scalars,
is symmetric for all
x.
We verify
by differentiating with respect to
:
using
). Hence,
.
With
and
:
Indeed,
using
.
Step 4 (The ODE is a Riemannian gradient flow). We use the following criterion. A vector field of the form on (with ) is a Riemannian gradient flow of some potential if its covariant Jacobian is self-adjoint on for every x; this is the integrability (symmetric-Hessian) condition for a gradient on a Riemannian manifold, and for it is automatic, but we verify it directly from the model. Here, , so and the covariant Jacobian is . Since and (Step 1), the sum is symmetric, so the compression is self-adjoint on . The criterion is therefore met at every x, confirming .
Step 5 (). Along a trajectory of the ODE, set and split it into its tangential and normal parts at x,
with the tangential part being, by definition, the Riemannian gradient. Using
together with (
19) and bilinearity,
The second term vanishes because
is orthogonal to
x, i.e.,
(equivalently, the ODE preserves
, so
). Therefore:
Step 6 (Convergence). L is continuous on the compact manifold , hence bounded. , so L is non-decreasing along trajectories. By LaSalle’s invariance principle, every trajectory converges to the largest invariant set contained in , which is the set of critical points of L on .
Since is compact and L is bounded and non-decreasing, the ω-limit set of every trajectory is non-empty and invariant, hence contained in the critical point set. Convergence to a single critical point follows when the critical points are isolated, which holds for generic . When they are not isolated (for instance, when V has a repeated eigenvalue, so that the critical set contains submanifolds rather than isolated points), L is a Morse–Bott function along the gradient flow, and every trajectory still converges to a single connected component of the critical set (the ω-limit set of a gradient flow of an analytic, or Morse–Bott, potential is a single critical component). The convergence conclusion is therefore unchanged; only the limit may be a critical manifold rather than a point.
Periodic orbits are excluded because L is strictly increasing along any non-constant trajectory. □
Corollary 1 (Fixed points and stability)
. The fixed points of (
11)
are the critical points of L on : solutions of . When the are eigenvectors of V with eigenvalues , this decouples asA critical point is a stable attractor if it has a strict local maximum of L (Riemannian Hessian negative definite; derived in the Supplementary Material); it is an unstable saddle if the Hessian has at least one positive eigenvalue. The global attractor (from generic initial conditions) is the global maximum of L. Remark 11 (Role of V in attractor selection). The global maximum of balances the Rayleigh quotient of V (favouring eigenvectors of V with large eigenvalues) against the MLP potential g (favouring directions where g is large). Increasing shifts the global maximum toward : a large makes tilt in the direction. The value matrix thus participates directly in attractor selection, in contrast to the attention matrix A, which drives the collapse to a single cluster but does not select among critical points.
5.2. Uncollapsed Token Dynamics
Proposition 4 below shows that for general V and Φ, a stable one-cluster fixed point need not be stable for the multi-token system. We now show that the u-aligned MLP with symmetric V is a special case where this instability is excluded: the gradient flow structure automatically guarantees multi-token stability at every attractor.
We consider n tokens near a critical point of L: , decomposed as (mean plus fluctuation, ).
At first order in
:
and
. Averaging over
i and using
:
At a local maximum of
L,
is negative definite, so
: the cluster mean converges to
. At a saddle, the mean grows in the unstable Hessian direction, driving the cluster toward the global maximum of
L.
Since Φ acts independently on each token and
, the MLP contributes
to
and the curvature correction contributes
(the same normal-force term as in
); the softmax attention contributes nothing at first order (
Supplementary Material). Combined,
Writing
for the Riemannian Hessian of the MLP potential and using
, the operator is
. At a local maximum of
L, one has
(because
and
for
), hence
: all fluctuation eigenvalues are negative near a local maximum of
L.
Theorem 5 (Uncollapsed token dynamics). Let be a critical point of L on .
- (i)
(Collapse.) If is a local maximum of L, both mean and fluctuation modes decay. Every trajectory of (
1)
near converges to : tokens collapse to the stable attractor. - (ii)
(Saddle escape.) Suppose is a saddle of L and the fluctuation-stability condition
holds (equivalently, ; a sufficient explicit form is , which, in particular, forces ). Then, the mean mode grows in the unstable Hessian direction, driving the cluster toward the global maximum of L, while every fluctuation mode decays: the tokens stay clustered while the cluster escapes the saddle. Condition (
24)
holds automatically at a local maximum (part (i)); at a saddle, it is a genuine restriction, and if it fails, the saddle drives the tokens apart instead of keeping them clustered.
Proof. Part (i): At a local maximum,
is negative definite, so both (
22) and the fluctuation operator in (
23) have all negative eigenvalues. Part (ii): The mean operator is
and the fluctuation operator is
. At a saddle,
has an eigenvalue
with eigenvector
, so, by (
22), the mean grows,
, and the cluster leaves the neighbourhood of
along
toward the global maximum of
L. Condition (
24) makes
negative definite, so, by (
23), every fluctuation mode decays
at an exponential rate bounded away from zero. Hence, the inter-token spread
contracts while the common mean escapes: the tokens remain clustered throughout. If, instead, (
24) fails, then
for some
and that fluctuation direction grows—the saddle drives the tokens apart, which is the multi-token instability of Proposition 4(ii). □
Remark 12 (Why the u-aligned case resolves the generic instability). Proposition 4(ii) below identifies a potential conflict: V may stabilise the one-cluster ODE without stabilising the fluctuations. The u-aligned MLP with symmetric V avoids this by construction. Since V and share eigenvectors , the operators differ only by , so for symmetric positive semi-definite V the eigenvalue correction is , so . The gradient flow structure ensures that everywhere, and since the correction is negative, : fluctuation modes are more stable than collapsed modes.
We use
,
,
,
,
,
,
,
, tanh activation, and step size
.
Figure 7 shows six experiments with two parameter regimes: (i) stable regime with
, global max of
L near
; (ii) saddle regime using
, for which
is a saddle of
and
is the global maximum.
Starting near the global max under : L increases monotonically and (panel (a)). Starting near the same point under (where it is a saddle): still increases monotonically and , but (panel (b)). The cluster escapes the saddle while staying clustered, confirming Theorem 5(ii).
Projected onto : stable-regime tokens (green) all converge to ; saddle-escape tokens (red) converge to . In both cases, tokens form a single cluster at the final position.
Varying with fixed: the global maximum transitions from -dominated (small ) to -dominated (large ), with a smooth crossover near . This confirms Remark 11.
From a common random initialisation, L increases monotonically to different attractor values for five values of , confirming Theorem 4 uniformly.
Six perturbation sizes near the stable all collapse at machine precision, confirming Theorem 5(i).
6. One-Cluster vs. Multi-Token: The General Mechanism
We establish a general relationship between one-cluster stability and multi-token stability for arbitrary V and Φ.
Let
be a fixed point:
, i.e.,
with
. Write
for the normal component of the value matrix. For standard softmax attention, the linearised attention force is independent of the attention matrix and common to all tokens at first order, so it acts only on the mean and contributes no first-order term to the fluctuations (this is established self-containedly in Proposition 5 below). The clustering of fluctuations is therefore carried entirely by the curvature term
, and the relevant linearisation operators on
are as follows (derived below in (
39)):
where both operators carry the same curvature term
,
, arising from the normal force
at the collapsed fixed point. Note that the tangential action of
V (the form
) is absent from
: the value matrix term
contributes only to the mean mode, not to fluctuations at first order. (
V’s normal component
is present in both operators through
μ and cancels when they are compared.)
Throughout this comparison, we write
,
, and
for the quadratic forms (Rayleigh quotients) of the linearisation operators on the tangent space. Since
, these satisfy
as an identity of quadratic forms, valid for arbitrary (possibly non-symmetric)
V.
For a symmetric operator, the Rayleigh quotient attains the eigenvalues at its stationary values, so
is the top eigenvalue and the sign of the form decides stability. For a non-symmetric
(non-symmetric
V), this fails:
is the numerical abscissa (the top eigenvalue of the symmetric part
), an upper bound for the spectral abscissa
that can be strictly larger. Consequently, the collapse criterion of Proposition 4(i) is a sufficient condition: it forces the numerical abscissa of
below zero, hence giving monotone (not merely asymptotic) decay. The exact one-cluster spectrum for non-symmetric
V is obtained spectrally by the deflation Lemma 2 (
), not from these quadratic forms; we use the form-level relation (
27) only for the sufficient comparison below.
Proposition 4 (One-cluster vs. multi-token stability). Let be the one-cluster stability margin (in the numerical-range sense above).
- (i)
(Sufficient condition for collapse.) If V is positive semi-definite on ( for all ), then for every unit : all fluctuation eigenvalues are negative and the tokens collapse to . No further condition is required.
- (ii)
(Possible instability.) For general V and Φ, there exist configurations where (one-cluster stable) but (tokens spread). This occurs when V is contractive in direction v (stabilising the mean but not the fluctuations), while is expansive in v.
Proof. The eigenvalue relation. For any
with
, write
and
. From operators (
25) and (26):
Subtracting:
This is the key relation. It shows that the tangential action of
V, encoded in the quadratic form
, enters
but not
. (The normal component
enters both operators equally through
and therefore cancels in the difference (
30).) For positive semi-definite
V (
), the fluctuation mode is more stable than the collapsed mode:
. For indefinite
V,
is possible; either effect can make
, so that fluctuations are less stable than the mean and potentially unstable even when
.
Part (i) Sufficient condition. We want to show that for all , given that is one-cluster stable () and V is positive semi-definite on .
From (
30):
Since
on
, we have
. From the one-cluster stability margin,
for all units
. Substituting into (
31) and using
yields:
the last inequality by the hypothesis
. This holds for any
and any
with
. The value matrix enters the fluctuation eigenvalue only through the stabilising
; the only role of the attention is through the bound, involving neither
nor
.
Part (ii) (An explicit instability). We construct explicit parameters. Let , , . Set (contractive on , , so V is indefinite), with (MLP expansive), and .
Then, writing
(since
), with
the normal component of the MLP map from the
Supplementary Material (not the Jacobian normal component
), yields:
at the fixed point
, so
; choosing the MLP so that
gives
. This is consistent with
: the map vanishes at
while its Jacobian there is expansive in the tangent direction. Then,
Choose
(possible since
C and
ϕ can be chosen freely with
, then
):
This confirms that a stable one-cluster fixed point can be multi-token unstable. The mechanism:
V is contractive on the tangent space (
) providing one-cluster stability, but does not appear in
, which is governed by
. □
Remark 13 (The curvature term is the full
)
. Both and contain the same
curvature term with the full , originating from the normal force at the collapsed fixed point (see (25) and (26) and the linearisation in the Supplementary Material). In particular, the curvature term is not
(which would omit the normal component of the value matrix) and not
(the Jacobian normal component, a different object, cf. the Supplementary Material). Because appears identically in both eigenvalues, it cancels in the difference (
30)
, leaving , with no residual . Remark 14 (Interpretation). The key asymmetry is that the tangential action of V (the form ) appears in but not in . A value matrix that stabilises the one-cluster ODE by being contractive in some tangential direction v provides no corresponding stability to the fluctuation modes in that direction. The fluctuation stability depends entirely on (and, in the MLP-free case, on the sign of ). This asymmetry is the source of the potential discrepancy between one-cluster and multi-token stability for general V and Φ.
We close the section by recording the derivation of (
25) and (26) and the proof that softmax attention contributes no first-order clustering force. Place
n tokens at
with
, and split
,
,
.
Proposition 5 (No first-order attention clustering force)
. For standard softmax attention with any coupling matrix A, the attention-weighted value force satisfiesindependent of the token index i to first order. Hence, the attention force drives only the mean mode and contributes no first-order term to the fluctuations . Proof. Linearising the score
about
gives, after softmax normalisation,
in which the query deviation
cancels in the softmax denominator, so the first-order correction depends on the key index
j only. Hence,
The first sum is
; in the second,
(the rows of
W sum to one) multiplies the
i-independent vector
to zero, while the correction acting on
is
. This gives (
37), common to all
i; subtracting the mean of the linearised Equation (
39) removes this common term, so it contributes nothing to
at first order. □
The linearised operators follow. Expanding
to first order, using (
37),
,
, and
, gives
The value matrix acts on the mean
only, whereas the MLP acts on the full deviation
. Averaging (
39) over
i gives
, which is (
25); subtracting the mean equation, the common term
cancels and
, which is (26). The value matrix is thus absent from
: it enters only through the common mean.
7. Discussion
The paper establishes two results at different levels of generality.
The first (Proposition 4) is a general structural observation: for arbitrary
V and Φ, a stable fixed point of the one-cluster ODE need not be stable for the multi-token system. The asymmetry arises because
V enters the mean operator
but not the fluctuation operator
(
Section 6): the value matrix can stabilise the mean without stabilising the fluctuations, and a spreading instability occurs when
V is contractive in some tangential direction while
is expansive there. The second result (Theorems 4 and 5) shows that the
u-aligned MLP with symmetric
V is the exactly solvable case in which this instability is structurally excluded: the one-cluster flow is the gradient ascent of
, every trajectory converges to a critical point, and the tokens collapse to a single cluster at every stable attractor, escaping saddles while remaining clustered.
A recurring conclusion of the analysis is that the value matrix governs where the tokens collapse, while the attention governs only that they collapse and how fast. For softmax attention, the first-order clustering force vanishes (Proposition 5), so the consensus rate is set by the dominant eigenvalue
of
V and is independent of the attention matrix
A (Theorem 3), and the selected direction is the dominant eigenvector
, tilted off the geometric equator through the left eigenvector
when
V is non-symmetric (Theorem 2). The attention matrix re-enters only in the multi-cluster question: whether two clusters occupying opposite basins coexist or merge is decided by the single scalar
(
Section 4.3). This gives a sharp, finite-
n complement to the mean-field picture of [
2,
3,
4], in which the limiting configuration is likewise governed by the spectrum of the value matrix; here, the basin structure is described explicitly and holds without a mean-field limit. In the same direction, the mean-field analysis of the feed-forward block [
14] shows that its critical points are generically atomic and localised on the sphere, in agreement with the value-matrix and MLP selection described here. The two viewpoints are complementary: the mean-field results characterise where the critical points lie as
, whereas our finite-
n analysis identifies which of them is reached from a given initial condition, at what rate, and by what mechanism. A rigorous mean-field limit of the present finite-
n statements is left as an open problem.
The sign of the dominant eigenvalue separates two qualitatively different behaviours. For
, the tokens collapse to a point, the dynamical counterpart of the rank collapses, and oversmoothing or token uniformity is observed in deep transformers [
5,
6]. For
, the one-cluster state is stable but nearby tokens spread rather than collapse, and a complex dominant pair produces no point cluster at all. In this reduced model, the value matrix thus carries a lever over representation degeneracy: a value matrix whose dominant eigenvalue is small, negative, or complex slows or prevents the collapse to a rank-one representation, in the same spirit as the architectural mechanisms (skip connections, feed-forward blocks, normalisation) known to mitigate it [
5,
6].
When V is symmetric, the flow is the gradient ascent of the Rayleigh quotient and the dynamics are of Morse type, with an antipodal pair of attractors and a hierarchy of equatorial saddles indexed by spectral depth (Remark 7). This gradient structure is lost precisely when a complex block becomes dominant: the point attractor is then replaced by an exact invariant great circle carried by the dominant eigenplane (Proposition 1, Remark 6). These cycles are not born from a Hopf bifurcation: they are present at full amplitude for every V carrying the complex pair, and a change of the global attractor between a point-pair and a circle is an exchange of dominance as two real parts cross. Non-normality of V leaves the asymptotic selection unchanged but permits transient growth and a spiralling approach (Remark 2), and the discrete-time iteration adds a genuine rotation threshold absent from the ODE (Remark 5). Even the value-matrix-only dynamics can therefore fail to converge to a point, a mechanism for persistent oscillation of the token representations.
The model is deliberately reduced. It uses a single value matrix and a single coupling matrix rather than multi-head, multi-layer weights, tied input/output directions in the u-aligned MLP, and sphere normalisation in place of LayerNorm; the Lyapunov structure requires and a u-aligned Φ, and the multi-token collapse and trapping statements are local, obtained by linearisation about the collapsed state. The vanishing of the first-order attention force is specific to standard softmax normalisation: a different attention mechanism could contribute a genuine first-order term, which would sharpen collapse when contractive but could not by itself produce the spreading instabilities, which arise through V. Within these restrictions, the results are exact and hold for a finite particle number, without mean-field approximation.
The reductions are chosen to retain the mechanism under study while keeping the dynamics analytically tractable, and each preserves a specific feature of the transformer forward pass. Sphere normalisation retains the norm constraint that LayerNorm imposes and is the natural geometry for the projected ODE. The single coupling matrix
A carries the query–key score geometry
(and the effective
for multi-head attention), and the value/output map, absent from pure-attention models, is exactly the object
added here. The
u-aligned MLP is the tractable case in which the Lyapunov structure is exact; the general dense MLP, whose Jacobian need not be symmetric, is left as an open problem (below). These reductions omit causal masking, positional encodings, and layer-dependent weights, so the model is not a quantitative surrogate for a trained transformer; rather, the phenomena it exhibits (collapse driven by attention, location selected by the value matrix, and the collapse/spreading dichotomy set by
) are the reduced-model counterparts of the rank collapse, oversmoothing, and token uniformity reported for deep transformers [
5,
6], and are offered as mechanisms rather than as claims about production models.
Several directions remain. (i) The general (dense) MLP, where need not be symmetric, so that the spreading instability of Proposition 4(ii) can genuinely occur and a general feed-forward block may produce limit cycles on . (ii) The multi-cluster regime, where the cluster locations depend jointly on V and A through a self-consistency condition, and the order parameter governing coexistence should generalise. (iii) Deeper spectra: a higher-dimensional hierarchical flag and the associated lattice of invariant subspheres, defective or equal-real-part spectra (invariant tori), and the recursive basin structure on the separating subsphere. (iv) The relation to training, where V and A are learned and the attractor structure evolves, and a rigorous mean-field limit of the finite-n results obtained here.