Foreign-exchange decisions rest on hierarchically organized evidence whose latent structure is inadequately captured by Euclidean representations. Reinforcement-learning agents trained on flat embeddings inherit stability guarantees that do not transfer to the manifold supporting the latent state. We address both limitations through a hybrid architecture in which a schema-constrained structured chain-of-thought is embedded into a Poincaré ball, transported to a qubit register via angle encoding, and processed by an
L-layer hardware-efficient variational ansatz on a state-vector backend. The circuit exposes two read-outs to the policy, namely, a scalar Pauli-
Z observable and a projected quantum kernel inducing a fidelity-based similarity over magnet-price attractors, the latter identified via kernel-weighted recurrence density and finite-time Lyapunov statistics. The Lipschitz constraint on the action-value function is lifted from the hyperbolic geodesic distance to a joint metric on
. A stability theorem yields an explicit bound depending on the read-out operator norm, on the depth–width product of the ansatz, and on the curvature–Hilbert balance. The pipeline is evaluated on nine major FX crosses over a 2015–2025 out-of-sample window, with rolling-origin walk-forward retraining and broker-published transaction costs. The system attains
pair-averaged non-compounded monthly P&L and
maximum drawdown, with Sharpe
, Calmar
, and Probabilistic Sharpe Ratio exceeding
on every cross. The gain remains significant under a deflated-Sharpe-ratio test with
correction. Block-wise ablations exhibit strictly monotone degradation: removing the projected kernel costs
p.p. on annualized P&L, the joint Lipschitz penalty
p.p., the attractor module
p.p., and the hyperbolic embedding
p.p. The quantum block thereby instantiates a structurally non-classical, geometry-aware regularizer identifiable through ablation rather than asymptotically advantageous.
Full article