Next Article in Journal
Numerical Error Propagation in Contemporary Molecular Dynamics Simulations of Lithium Battery Components
Previous Article in Journal
From 2D Vision–Language Models to Volumetric Medical AI: Large Language Models and Foundation Models for 3D Medical Imaging
Previous Article in Special Issue
Computational Modelling of Institutional Perceptions in a Civilisational Intergovernmental Organisation
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Stochastic Semantic Fields for Sentiment-Driven Models

Department of Computer Science, College of Mathematics, University of Verona, Strada le Grazie 15, 37134 Verona, Italy
Computation 2026, 14(9), 216; https://doi.org/10.3390/computation14090216
Submission received: 7 August 2026 / Revised: 27 August 2026 / Accepted: 1 September 2026 / Published: 14 September 2026

Abstract

This article studies a time-dependent latent sentiment representation governed by a semilinear stochastic evolution equation on a separable Hilbert space. The objective is to derive uncertainty, interpretability, and robustness from one stochastic law. The mathematical analysis establishes covariance propagation and spectral uncertainty attribution, Fréchet differentiability of the forcing-to-state and forcing-to-readout maps with quadratic remainders, trace-norm differentiation of the covariance, and an adjoint influence kernel whose norm equals the exact worst-case displacement over an energy-bounded forcing ball in the linear regime. The dependence of the certificates on dissipativity, forecast horizon, noise geometry, and spectral truncation is made explicit. A Galerkin state-space reduction yields exact linear transitions, a likelihood-based calibration scheme, and structural identifiability conditions. A reproducible three-mode study verifies the covariance and duality identities, evaluates the sensitivity constants, and compares Gaussian with split-conformal predictive intervals under controlled synthetic conditions. The results provide a rigorous operator calculus and a tractable finite approximation. They establish internal mathematical and numerical validity without asserting a generalised superiority.

1. Introduction

Contextual encoders and large language models have increased the expressive power of sentiment systems while also making their temporal behaviour harder to analyse. A terminal score may appear accurate even when its dispersion is poorly calibrated, its response to earlier semantic evidence is opaque, or a modest perturbation of the information stream creates a large displacement. The mathematical problem studied here is the construction of uncertainty, interpretability, and robustness from one time-dependent stochastic law. This problem matters whenever sentiment enters a sequential decision process, since the reliability of the terminal output depends on the evolution that produced it.
Classical work established the linguistic and computational foundations of opinion mining [1,2], and transformer encoders supplied a high-dimensional geometry for contextual text representations [3]. Current large language model studies show that strong aggregate performance does not remove weaknesses on structured sentiment phenomena [4]. Progressive task adaptation improves generic affective capabilities [5], while causal analysis reveals that apparently simple sentiment decisions may involve distinct internal reasoning routes [6]. The faithfulness of model-generated explanations remains dependent on the architecture and the task [7]. These findings motivate a dynamical representation in which semantic evidence evolves before the readout is formed and in which the resulting diagnostic quantities have operator-level definitions.
Bayesian approximations and post-training calibration remain central finite-dimensional tools [8,9,10]. Conformal prediction complements them by providing distribution-free marginal coverage under exchangeability [11,12]. Recent evidence also shows that in-context language model probabilities can remain substantially miscalibrated [13]. The present framework uses covariance propagation to describe the internal transport of stochastic uncertainty. The synthetic study then applies split conformal calibration to terminal projected residuals. The covariance operator supplies a structural decomposition, whereas the conformal layer supplies a finite-sample coverage statement for the projected predictor.
Interpretability and robustness are derived from the same solution map. Local surrogate methods and Shapley-type attributions remain influential [14,15], while adversarial analyses have documented the fragility of text systems under constructed and naturally occurring perturbations [16,17,18]. In the present setting, the Fréchet derivative of the terminal readout with respect to the forcing path defines the explanation. Its Riesz representative is an adjoint influence kernel. The norm of this kernel equals the exact worst-case displacement of a linear readout over an energy-bounded forcing ball; consequently, the Hilbert-space duality links interpretability and robustness exactly.
The latent state is modelled in a separable Hilbert space through
d X t = A X t + F ( X t ) + B u t d t + G ( X t ) d W t , X 0 = x .
A time-indexed text encoder produces the forcing path u. The bounded map B lifts that path into the state space. The generator A propagates information across semantic directions, while F describes state-dependent response, and G transmits unresolved variability. A scalar prediction is obtained from Φ ( X T ) . When the semantic representation carries a neighbourhood structure, a graph Laplacian or a kernel operator gives a concrete approximation of A. Large eigenvalues then describe coordinates with high Dirichlet energy across neighbouring semantic prototypes. The phrase high semantic frequency refers to this graph-spectral variation and not to token frequency.
The analytic setting follows the semigroup theory of stochastic evolution equations [19,20,21]. Approximation theory for semilinear stochastic evolution equations provides the connection between the infinite-dimensional field and its finite projections [22]. Stochastic reaction-diffusion equations on networks provide a direct precedent for random diffusion over graph-like structures [23]. These two articles by Luca Di Persio and coauthors are pertinent to the Galerkin and network mechanisms used below.
Computational work on stochastic partial differential equations has also moved towards learned solution operators. Neural stochastic PDEs aim at resolution-invariant learning of continuous random dynamics [24]. Regularity-informed architectures and Bayesian deep solvers address stochastic forward and inverse problems while retaining information about numerical uncertainty [25,26]. Such methods can approximate the projected state equation. The operator identities proved here provide independent checks for covariance propagation, adjoint sensitivity, and spectral truncation.
A broader numerical PDE literature clarifies the methodological position of the paper. Dissipative PDE evolution can serve as a denoising mechanism [27]. Deep learning can recover forward solutions and unknown coefficients for coupled nonlinear wave equations [28]. Lie-group methods show how a high-order discretisation can preserve a prescribed invariant [29], while the Riemann–Hilbert method provides exact spectral information for integrable coupled systems [30]. Artificial-intelligence-enhanced mathematical derivation combines learned search with symbolic verification of exact formulae [31]. The stochastic semantic field has a dissipative probabilistic structure. Its numerical approximations must therefore preserve contractivity, covariance positivity, and consistency between the forward and adjoint equations. The cited methods supply complementary inverse, geometric, spectral, and symbolic techniques. Their mathematical targets differ from the sentiment model.
Appendix A contains the stochastic-convolution estimates, the fixed-point existence proof, finite-horizon input stability, fourth-moment estimates, Galerkin consistency, covariance identities, and the Yosida passage used in the main proofs. Theorem 1 and Theorem 13 adapt established synchronous-coupling and Gronwall arguments to the semantic-field setting. The model-specific contribution begins with the joint covariance, sensitivity, and perturbation calculus. It includes a genuine Fréchet differentiability theorem with a quadratic remainder, trace-norm differentiation of covariance, an adjoint influence representation, exact influence-robustness duality, and explicit truncation and ablation formulae.
The computational part uses a reproducible three-mode Galerkin experiment. In particular, exact transitions are used to estimate a stable state-space model, to compare it with finite-dimensional baselines, evaluate Gaussian and split-conformal intervals, and verify the covariance and duality identities. The experiment also measures the dependence of the diagnostics on dissipation and noise geometry. It establishes internal numerical validity under controlled synthetic conditions. Sentiment140, SemEval, TweetEval, IMDb, and the NRC Emotion Lexicon are discussed as candidate resources for a subsequent study [32,33,34,35,36].
Section 2 develops the analytical results and their proofs. Section 3 gives the finite-dimensional calibration map. Section 4 presents the synthetic study. Section 5 states the conclusions and the scope of the claims. The appendices contain the standard analytical material and the conformal coverage argument.

2. Original Theoretical Contributions and Proofs

The analytical architecture separates classical stochastic-evolution theory from the results that are specific to sentiment-driven diagnostics. Appendix A supplies the existence theory and the approximation devices required to manipulate mild solutions. The contraction theorem below and the later finite-horizon perturbation theorem adapt standard Hilbert-space stability arguments to the semantic field notation and make the dependence on the dissipativity and Lipschitz constants explicit. The original contribution lies in the joint operator calculus built on top of that state equation. The covariance of the latent field is pushed to an output uncertainty and decomposed spectrally. The solution map with respect to semantic forcing is differentiated in the Fréchet sense with a uniform quadratic remainder. This derivative yields a trace-norm covariance derivative and an adjoint influence kernel. The kernel norm is proved to equal the exact worst-case displacement over an input energy ball. Spectral Galerkin truncation is then controlled directly at the level of uncertainty and influence, so the finite diagnostics inherit a quantified relation to the infinite dimensional objects.
The standard lineage of Theorem 1 is stated explicitly. Its proof uses synchronous coupling, the Hilbert-space Itô formula, and Gronwall’s lemma. The sentiment interpretation and the later exact duality are model-specific, while the contraction mechanism itself belongs to the established stability theory of stochastic evolution equations. The same distinction applies to Theorem 13.
The construction starts from an abstract semantic space. Let ( E , E , μ ) be a sigma-finite measure space. The set E represents an abstract semantic domain. Its elements may be interpreted as latent semantic locations, topics, contextual directions, source strata, or any continuum-level representation of evaluative content. Fix an integer d 1 and define
H = L 2 ( E , E , μ ; R d ) .
The inner product and norm on H are denoted by · , · H and · H . When no ambiguity is possible, the subscript is omitted.
The dimension d allows one to represent several affective components, such as polarity, intensity, stance, or uncertainty-aware latent channels. However, the present article does not assign empirical meaning to these components and treats them instead as coordinates of a theoretical semantic field.
Let U and K be separable real Hilbert spaces, where U carries semantic forcing signals and K carries the cylindrical noise. The notation L ( U , H ) denotes bounded linear operators from U to H, while L 2 ( K , H ) denotes Hilbert–Schmidt operators from K to H, endowed with the norm · L 2 ( K , H ) .
The semantic diffusion is described by a densely defined operator. Let A : Dom ( A ) H H be a densely defined closed linear operator generating a strongly continuous semigroup S ( t ) , t 0 . The operator A is the semantic diffusion and relaxation operator, determining how sentiment intensity spreads, smooths, and dissipates across the latent semantic domain.
The following assumption is used throughout the main part of the article.
Assumption 1
(Semigroup structure). There exist constants M 1 and ω R such that
S ( t ) L ( H ) M e ω t , t 0 .
When spectral uncertainty attribution is discussed, the stronger condition is imposed that A is self-adjoint and satisfies
A e k = λ k e k , k N
for an orthonormal basis ( e k ) k 1 of H, where 0 < λ 1 λ 2 and λ k .
The spectral condition covers the canonical case of an elliptic diffusion operator on a bounded semantic domain with dissipative boundary conditions, and the abstract statement is broader, without requiring a concrete geometric representation. The probabilistic structure is fixed, i.e., let
( Ω , F , ( F t ) t [ 0 , T ] , P )
be a complete filtered probability space satisfying the usual conditions, and consider W = ( W t ) t [ 0 , T ] to be a cylindrical Wiener process on K adapted to ( F t ) . If ( f j ) j 1 is an orthonormal basis of K, then
W t = j 1 β j ( t ) f j
formally, where ( β j ) j 1 are independent real Brownian motions. The stochastic integral
0 t Ψ s d W s
is well defined as an H-valued square-integrable random variable whenever Ψ is predictable and
E 0 T Ψ s L 2 ( K , H ) 2 d s < .
With these spaces fixed, let B L ( U , H ) , and let
F : H H , G : H L 2 ( K , H )
be measurable maps. The stochastic semantic field driven by a predictable forcing process u is defined by the SPDE
d X t = A X t + F ( X t ) + B u t d t + G ( X t ) d W t , X 0 = x .
The corresponding mild formulation is
X t = S ( t ) x + 0 t S ( t s ) F ( X s ) + B u s d s + 0 t S ( t s ) G ( X s ) d W s .
The state X t is the sentiment field, the forcing u t is the abstract representation of incoming sentiment-bearing information, and the operator B lifts this forcing into the latent semantic field. The drift F captures nonlinear sentiment reaction, whereas the diffusion coefficient G determines how uncertainty enters the semantic dynamics.
Assumption 2
(Lipschitz and growth conditions). There exist constants L F , L G , C F , C G 0 such that for all x , y H ,
F ( x ) F ( y ) L F x y , G ( x ) G ( y ) L 2 ( K , H ) L G x y ,
and
F ( x ) C F ( 1 + x ) , G ( x ) L 2 ( K , H ) C G ( 1 + x ) .
Definition 1
(Mild sentiment field). Let x L 2 ( Ω , F 0 , P ; H ) , and let u L 2 ( Ω × [ 0 , T ] , P , P d t ; U ) , where P is the predictable sigma field. A predictable process X : [ 0 , T ] × Ω H is a mild sentiment field on [ 0 , T ] if X is mean-square continuous, which satisfies
sup t [ 0 , T ] E X t 2 < ,
and satisfies (3) for every t [ 0 , T ] , as an equality in L 2 ( Ω ; H ) .
The stochastic field becomes a sentiment-driven prediction through a readout Φ : H R , so that the real random variable Φ ( X T ) represents the sentiment-driven output at horizon T, and the three central quantities are defined as follows:
Definition 2
(Uncertainty, influence, and robustness modulus). Let X u be the mild field generated by a deterministic input u L 2 ( 0 , T ; U ) . For a readout Φ : H R with finite second moment, define
Unc Φ , T ( u ) = Var [ Φ ( X T u ) ] 1 / 2 .
When the map J T ( u ) = E Φ ( X T u ) is Fréchet differentiable, its influence in the direction h L 2 ( 0 , T ; U ) is
Inf Φ , T ( u ; h ) = D J T ( u ) h .
Let Θ denote the collection of model ingredients, and let d T be a specified discrepancy on admissible models. If D T ( Θ ) is a diagnostic with values in a metric space with distance d D , its robustness modulus is
Rob D , T ( ρ ; Θ ) = sup d T ( Θ , Θ ˜ ) ρ d D D T ( Θ ) , D T ( Θ ˜ ) .
Theorems below specify d T and prove finite moduli for the terminal law, expected readout, output standard deviation, covariance, and influence kernel.
The analysis below proves that these quantities are well defined and controlled by the same stochastic evolution equation. The classical result ensuring that X u exists uniquely is Theorem A1 in Appendix A, and the classical finite-horizon continuity of the solution map is Proposition A1 in the same appendix.
For later use, we define T > 0 , M T as the space of predictable mean-square continuous processes X : [ 0 , T ] × Ω H such that
X M T 2 = sup t [ 0 , T ] E X t 2 < ,
which is a Banach space.
The following synchronous-coupling estimate exposes the dissipativity margin used later in the certification bounds.
Assumption 3
(Mean-square dissipativity). The semigroup generator satisfies
A z , z α z 2 , z Dom ( A ) ,
for some α > 0 . Moreover, there exists η R such that
2 F ( x ) F ( y ) , x y + G ( x ) G ( y ) L 2 ( K , H ) 2 η x y 2
for all x , y H .
Theorem 1
(Dissipative mean-square contraction). Let Assumptions 1–3 hold. Let X and Y be solutions of (2) with the same forcing u, the same Wiener process, and initial states x and y. Then, for every t [ 0 , T ] ,
E X t Y t 2 e ( 2 α η ) t E x y 2 .
If 2 α > η , the stochastic semantic dynamics are exponentially contractive in mean square.
Proof. 
The proof is first given for strong solutions with values in Dom ( A ) , while the general mild case follows by Yosida approximation, as detailed in Appendix A. Let
Z t = X t Y t .
Since the forcing is the same, Z satisfies
d Z t = [ A Z t + F ( X t ) F ( Y t ) ] d t + [ G ( X t ) G ( Y t ) ] d W t .
Apply the Hilbert-space Itô–Döblin formula to Z t 2 . This gives
Z t 2 = Z 0 2 + 2 0 t A Z s , Z s d s + 2 0 t F ( X s ) F ( Y s ) , Z s d s + 0 t G ( X s ) G ( Y s ) L 2 ( K , H ) 2 d s + 2 0 t Z s , [ G ( X s ) G ( Y s ) ] d W s .
The stochastic integral has expectation zero because the integrand is square-integrable. Taking expectations and using Assumption 3,
E Z t 2 E Z 0 2 2 α 0 t E Z s 2 d s + η 0 t E Z s 2 d s = E Z 0 2 ( 2 α η ) 0 t E Z s 2 d s .
Let r = 2 α η . The last inequality implies
f ( t ) f ( 0 ) r 0 t f ( s ) d s , f ( t ) = E Z t 2 .
If r is any real number, the standard differential form of Gronwall’s lemma yields
f ( t ) e r t f ( 0 ) .
This is (4). The Yosida passage in Appendix A shows that the same inequality holds for mild solutions because the approximating strong solutions converge in M T . □
The contraction estimate also permits an analytic ablation of the constants that determine whether a robustness certificate is informative. Set
γ = 2 α η .
The parameter γ is the mean-square dissipativity margin. Its role can be separated from the forcing amplitude and from the spectral placement of the noise.
Theorem 2
(Ablation modulus for dissipation, forcing, and modal noise). Assume that γ > 0 . Let X and Y solve (2) with the same coefficients and the same Wiener process but with initial states x , y and forcing paths u , v . For every ϑ ( 0 , γ ) and t [ 0 , T ] ,
E X t Y t 2 e ( γ ϑ ) t E x y 2 + B 2 ϑ 0 t e ( γ ϑ ) ( t s ) E u s v s U 2 d s .
If E u s v s U 2 d 2 for almost every s, then
sup 0 t T E X t Y t 2 E x y 2 + B 2 d 2 ϑ 1 e ( γ ϑ ) T γ ϑ .
The horizon-uniform upper envelope obtained by replacing the last fraction by ( γ ϑ ) 1 is minimised at ϑ = γ / 2 , where the forcing amplification factor is 4 B 2 / γ 2 .
For the remaining assertions, assume that the initial state x H and the forcing path u L 2 ( 0 , T ; U ) are deterministic. Assume further that the equation is linear with F = 0 and constant diffusion G ( x ) = Σ , where Σ L 2 ( K , H ) , and set Q = Σ Σ * . Suppose that A e k = λ k e k , that Q e k = q k e k , and that Φ a ( x ) = a , x , where a k = a , e k . Define
h T ( λ ) = 1 e 2 λ T 2 λ .
The contribution of mode k to the terminal output variance is
v k ( T ) = a k 2 q k h T ( λ k ) .
For every T > 0 and λ > 0 ,
T h T ( λ ) = e 2 λ T > 0 , λ h T ( λ ) = e 2 λ T ( 2 λ T + 1 ) 1 2 λ 2 < 0 ;
therefore, a fixed value of a k 2 q k produces less terminal variance when it is assigned to a mode with a larger damping eigenvalue.
Finally, define
R T ( γ ) = 0 T e γ r d r 1 / 2 = 1 e γ T γ 1 / 2 .
It is the L 2 -norm of the exponential memory kernel that appears in the perturbation certificate. It satisfies
T R T ( γ ) 2 = e γ T > 0 , γ R T ( γ ) 2 = e γ T ( γ T + 1 ) 1 γ 2 < 0 ,
so that the certificate deteriorates monotonically with the forecasting horizon and improves monotonically with the dissipativity margin.
Proof. 
For strong solutions, let Z t = X t Y t . The Itô formula used in the proof of Theorem 1 contains the additional term
2 Z t , B ( u t v t ) .
Young’s inequality gives
2 Z t , B ( u t v t ) ϑ Z t 2 + B 2 ϑ u t v t U 2 .
After taking expectations, Assumption 3 yields
d d t E Z t 2 ( γ ϑ ) E Z t 2 + B 2 ϑ E u t v t U 2
for almost every t. Multiplication by e ( γ ϑ ) t and integration prove (5). The bounded-input estimate follows by evaluating the exponential integral. The horizon-uniform envelope uses
0 t e ( γ ϑ ) ( t s ) d s 1 γ ϑ .
The function ϑ [ ϑ ( γ ϑ ) ] 1 is minimised at γ / 2 . The Yosida argument in Proposition A3 passes the estimate to mild solutions without changing its constants.
Formula (7) follows from Theorem 4. Direct differentiation gives (8). Its second numerator is negative because
e 2 λ T > 1 + 2 λ T .
The same calculation with γ in place of 2 λ proves (10). The strict inequalities follow from T > 0 , λ > 0 , and γ > 0 . □
Theorem 2 identifies the constants responsible for amplification before a dataset is selected. Section 4 evaluates the resulting functions in a parameter regime where the exact law is available.
We now formalise uncertainty through the covariance operator of the sentiment field and the variance of the sentiment readout. The nonlinear theory gives general bounds, the linear Gaussian theory gives exact formulas and a weak Lyapunov equation, and the spectral theory gives an interpretable decomposition by semantic modes.
Definition 3
(Covariance operator). Let Y L 2 ( Ω ; H ) . Its mean is m = E Y H . Its covariance operator is the operator C Y : H H , defined by
C Y h = E Y m , h ( Y m ) , h H .
The standard trace identity for Hilbert-space covariance operators is recalled in Appendix A as Lemma A3; it will be used below to pass from field-level covariance to scalar readout uncertainty.
Proposition 1
(Output uncertainty bound). Let X T L 2 ( Ω ; H ) , and let Φ : H R be Lipschitz with constant L Φ . Then,
Unc Φ , T 2 = Var [ Φ ( X T ) ] L Φ 2 Tr C X T .
Proof. 
For any real random variable Z L 2 ( Ω ) ,
Var ( Z ) = inf a R E | Z a | 2 .
Choosing a = Φ ( E X T ) , we obtain
Var [ Φ ( X T ) ] E | Φ ( X T ) Φ ( E X T ) | 2 L Φ 2 E X T E X T 2 .
The trace identity in Lemma A3 gives (11). □
Corollary 1
(Finite uncertainty for Lipschitz readouts). Under the assumptions of Theorem A1, every Lipschitz readout Φ : H R satisfies
Unc Φ , T 2 L Φ 2 C T 1 + E x 2 + E 0 T u s U 2 d s .
Proof. 
By Lemma A3,
Tr C X T = E X T E X T 2     E X T 2 .
The moment estimate (A1) and Proposition 1 complete the proof. □
It is worth mentioning that the nonlinear model is appropriate for general theory, while exact uncertainty formulae are most transparent in the linear Gaussian regime, which is also important because it gives the local covariance calculus around nonlinear sentiment trajectories. The Gaussian-measure facts used in this passage are standard; see [37].
Assume throughout the linear Gaussian analysis that F = 0 , that the diffusion is the constant operator Σ L 2 ( K , H ) , and that u L 2 ( 0 , T ; U ) is deterministic. The notation
Q = Σ Σ * L 1 ( H )
is reserved for the positive trace-class covariance injection operator. This distinction avoids treating a general Hilbert–Schmidt map from K to H as an operator square root on H. The equation is
d X t = ( A X t + B u t ) d t + Σ d W t , X 0 = x H .
Theorem 3
(Gaussian law, covariance formula, and weak Lyapunov equation). Let Assumption 1 hold. For x H , u L 2 ( 0 , T ; U ) , and Σ L 2 ( K , H ) , the mild solution of (12) is Gaussian in H at every time. Its mean and covariance are
m t = S ( t ) x + 0 t S ( t s ) B u s d s
and
C t = 0 t S ( r ) Q S ( r ) * d r , Q = Σ Σ * ,
where the integral converges in the trace norm. For every φ , ψ Dom ( A * ) , the map
c φ , ψ ( t ) = C t φ , ψ
is continuously differentiable and satisfies
d d t c φ , ψ ( t ) = C t A * φ , ψ + C t φ , A * ψ + Q φ , ψ .
Moreover,
Tr C t M 2 e 2 ω + T t Σ L 2 ( K , H ) 2 , t [ 0 , T ] .
If S ( t ) e α t for some α > 0 , then
Tr C t 1 e 2 α t 2 α Σ L 2 ( K , H ) 2 .
Proof. 
The mild solution is
X t = m t + Z t , Z t = 0 t S ( t s ) Σ d W s .
For any h 1 , , h m H , the vector with components Z t , h j is a finite family of stochastic integrals with deterministic integrands. It is therefore a centred Gaussian vector, which proves that Z t , and hence, X t , are Gaussian in the Hilbert-space sense. The stochastic integral has zero mean, and (13) follows.
For h , g H , the Itô isometry gives
E Z t , h Z t , g = 0 t Σ * S ( t s ) * h , Σ * S ( t s ) * g K d s = 0 t S ( t s ) Q S ( t s ) * h , g d s .
After the substitution r = t s , this is the weak form of (14). The ideal property of trace-class operators yields
S ( r ) Q S ( r ) * 1 S ( r ) 2 Q 1 = S ( r ) 2 Σ L 2 ( K , H ) 2 ,
implying that the covariance integral exists as a Bochner integral in L 1 ( H ) .
We next justify the weak Lyapunov equation without applying A to the range of Q. Since t S ( t ) * φ is continuously differentiable for φ Dom ( A * ) , the fundamental theorem of calculus gives
C t A * φ , ψ + C t φ , A * ψ = 0 t d d r Q S ( r ) * φ , S ( r ) * ψ d r = Q S ( t ) * φ , S ( t ) * ψ Q φ , ψ .
It remains to justify the differentiation of the covariance integral in the trace norm. The map r S ( r ) Σ is continuous with values in L 2 ( K , H ) . To see this, approximate Σ in Hilbert–Schmidt norm by finite-rank operators, use strong continuity of S on each finite-dimensional range, and use the uniform semigroup bound on compact time intervals for the approximation error. The inequality
R R * T T * 1 ( R L 2 + T L 2 ) R T L 2
then shows that r S ( r ) Q S ( r ) * is trace-norm continuous. The fundamental theorem of calculus for Bochner integrals gives
d d t C t φ , ψ = Q S ( t ) * φ , S ( t ) * ψ .
Combining the last two identities proves (15). This argument uses only the domain invariance of the adjoint semigroup on Dom ( A * ) and does not require Q H Dom ( A ) .
Finally, for any orthonormal basis ( e k ) of H, Tonelli’s theorem and the Hilbert–Schmidt ideal property imply
Tr C t = 0 t k 1 Σ * S ( r ) * e k K 2 d r = 0 t S ( r ) Σ L 2 ( K , H ) 2 d r M 2 e 2 ω + T t Σ L 2 ( K , H ) 2 ,
which is (16). Under exponential stability, the same calculation gives
Tr C t Σ L 2 ( K , H ) 2 0 t e 2 α r d r ,
and (17) follows. □
Corollary 2
(Exact readout uncertainty and Gaussian confidence bound). Under the assumptions of Theorem 3, let a H and Φ a ( x ) = a , x . Then
v t ( a ) = Var [ Φ a ( X t ) ] = C t a , a .
If v t ( a ) > 0 , then, for every r > 0 ,
P | Φ a ( X t ) E Φ a ( X t ) | r 2 exp r 2 2 v t ( a ) .
When v t ( a ) = 0 , the centred readout vanishes almost surely, and the left-hand side is zero.
Proof. 
The covariance identity gives (18). The centred readout is a real Gaussian random variable with variance v t ( a ) . If Z is standard Gaussian, by the exponential Markov inequality, it holds
P ( Z z ) inf θ > 0 e θ z E e θ Z = inf θ > 0 e θ z + θ 2 / 2 = e z 2 / 2 .
Applying this estimate to both tails with z = r / v t ( a ) 1 / 2 proves (19). □
The spectral structure is informative even when the covariance injection is not diagonal in the eigenbasis of the generator. The following formula distinguishes direct modal uncertainty from cross-semantic covariance.
Theorem 4
(Cross-semantic spectral covariance decomposition). Assume that A is self-adjoint and that A e k = λ k e k for an orthonormal basis ( e k ) k 1 , with λ k > 0 . Set
q j k = Q e k , e j .
Then
C t e k , e j = q j k 1 e ( λ j + λ k ) t λ j + λ k .
For a = k 1 a k e k , the output variance satisfies
Var [ Φ a ( X t ) ] = lim N j , k = 1 N a j a k q j k 1 e ( λ j + λ k ) t λ j + λ k .
The limit is the quadratic-form limit and does not require absolute convergence of the double series.
Proof. 
Since S ( r ) e k = e λ k r e k , and S ( r ) is self-adjoint,
C t e k , e j = 0 t Q S ( r ) e k , S ( r ) e j d r = q j k 0 t e ( λ j + λ k ) r d r ,
which proves (20). Let P N be the orthogonal projection onto the first N eigenvectors. Since P N a a in H and C t L ( H ) ,
C t P N a , P N a C t a , a .
Expanding the finite-dimensional quadratic form and using Corollary 2 gives (21). □
Corollary 3
(Diagonal semantic uncertainty attribution). If Q e k = q k e k , where q k 0 and k q k < , then
C t e k = q k ( 1 e 2 λ k t ) 2 λ k e k
and
Var [ Φ a ( X t ) ] = k 1 a k 2 q k ( 1 e 2 λ k t ) 2 λ k .
Proof. 
The diagonal assumption gives q j k = q k when j = k and zero otherwise. Equations (22) and (23) therefore follow from Theorem 4. The series converges because it equals the finite quadratic form C t a , a and also because
0 q k ( 1 e 2 λ k t ) 2 λ k t q k .
Theorem 5
(Stationary covariance and long-time uncertainty). Assume that S ( t ) e α t for some α > 0 . Then
C = 0 S ( r ) Q S ( r ) * d r
exists in L 1 ( H ) , is positive and self-adjoint, and satisfies
Tr C 1 2 α Σ L 2 ( K , H ) 2 .
For every φ , ψ Dom ( A * ) solves the weak algebraic Lyapunov equation
C A * φ , ψ + C φ , A * ψ + Q φ , ψ = 0 .
Moreover,
C C t 1 e 2 α t 2 α Σ L 2 ( K , H ) 2 ,
and consequently,
C a , a Var [ Φ a ( X t ) ] e 2 α t 2 α a 2 Σ L 2 ( K , H ) 2 .
Proof. 
The trace-class estimate
S ( r ) Q S ( r ) * 1 e 2 α r Q 1 = e 2 α r Σ L 2 ( K , H ) 2
is integrable on the positive half-line, which proves existence, positivity, self-adjointness, and (25). The same derivative computation used in Theorem 3, now integrated over [ 0 , R ] , gives
C R A * φ , ψ + C R φ , A * ψ + Q φ , ψ = Q S ( R ) * φ , S ( R ) * ψ .
The right-hand side tends to zero and C R C in trace norm, proving (26). Finally,
C C t = t S ( r ) Q S ( r ) * d r ,
and integration of the trace-norm bound proves (27). The scalar estimate follows from | T a , a | T 1 a 2 . □
The diagonal summand in (23) separates noise injection, readout sensitivity, and semantic damping. Formula (21) adds a fourth mechanism because off-diagonal entries of Q quantify the simultaneous stochastic excitation of distinct semantic directions. This distinction is material when ambiguity couples topics or affective coordinates. Independent modal perturbations form a special case.
Remark 1
(Semantic meaning of the uncertainty spectrum). The modal contribution
a k 2 q k 1 e 2 λ k t 2 λ k
contains a readout factor, a noise factor, and a propagation factor. The term high-frequency semantic mode has a precise meaning after a graph or kernel basis is chosen. For a graph Laplacian, λ k equals the Dirichlet energy in (71); a large eigenvalue therefore indicates rapid variation across semantically adjacent nodes. It does not refer to the corpus frequency of a word or token. The derivative calculation in Theorem 2 proves that the propagation factor decreases strictly with λ k . The interpretive statement is therefore conditional on the geometry encoded by the chosen basis and can be checked directly from the fitted graph.
Interpretability is formalised through the response of the state and the readout to a perturbation of the forcing path. A directional limit is insufficient for a robustness statement because it need not be uniform over small directions. The assumptions below yield a Fréchet derivative and a quadratic remainder.
Assumption 4
(Lipschitz differentiable coefficients). The maps F : H H and G : H L 2 ( K , H ) are continuously Fréchet differentiable. There are finite constants L D and L D , 1 such that
D F ( x ) L ( H ) + D G ( x ) L ( H , L 2 ( K , H ) ) L D , D F ( x ) D F ( y ) L ( H ) + D G ( x ) D G ( y ) L ( H , L 2 ( K , H ) ) L D , 1 x y
for all x , y H .
Let L 2 ( 0 , T ; U ) = L 2 ( 0 , T ; U ) . Let M T 4 be the space of predictable processes that are continuous from [ 0 , T ] to L 4 ( Ω ; H ) and satisfy
Y M T 4 = sup t [ 0 , T ] E Y t 4 1 / 4 < .
The fourth-moment estimates used below are proved in Proposition A2 of Appendix A.
Theorem 6
(Fréchet differentiability of the forcing-to-state map). Let Assumptions 1, 2, and 4 hold. Let x L 4 ( Ω , F 0 , P ; H ) , and let X u denote the mild solution driven by a deterministic path u L 2 ( 0 , T ; U ) . The map
S : L 2 ( 0 , T ; U ) M T , S ( u ) = X u
is Fréchet differentiable, and its derivative in the direction h L 2 ( 0 , T ; U ) is the unique predictable solution Z u , h of
Z t u , h = 0 t S ( t s ) D F ( X s u ) Z s u , h + B h s d s + 0 t S ( t s ) D G ( X s u ) Z s u , h d W s .
There is a constant C T , depending only on the structural constants and on T, such that
Z u , h M T 4 C T h L 2 ( 0 , T ; U )
and
X u + h X u Z u , h M T C T h L 2 ( 0 , T ; U ) 2 .
Proof. 
Fix u , h L 2 ( 0 , T ; U ) . The random operator coefficients D F ( X s u ) and D G ( X s u ) are predictable and bounded by L D . The fourth-moment deterministic and stochastic convolution estimates in Proposition A2 show that the map
( Γ Y ) t = 0 t S ( t s ) D F ( X s u ) Y s + B h s d s + 0 t S ( t s ) D G ( X s u ) Y s d W s
is a contraction on M τ 4 when τ > 0 is sufficiently small. Indeed, if c S = M e ω + T , then,
Γ Y Γ Y ˜ M τ 4 4 C c S 4 L D 4 τ 4 + τ 2 Y Y ˜ M τ 4 4 .
The forcing convolution satisfies
sup t T 0 t S ( t s ) B h s d s 4 c S 4 B 4 T 2 h L 2 ( 0 , T ; U ) 4 .
Local fixed points can therefore be concatenated to produce a unique process Z u , h M T 4 . Applying the same convolution inequalities directly to (28) gives
f ( t ) C T h L 2 ( 0 , T ; U ) 4 + C T 0 t f ( s ) d s , f ( t ) = sup r t E Z r u , h 4 .
Gronwall’s lemma proves (29). Linearity in h follows from uniqueness.
Set Δ h = X u + h X u . Proposition A2 yields
Δ h M T 4 C T h L 2 ( 0 , T ; U ) .
Let R h = Δ h Z u , h . Define
r F h ( s ) = F ( X s u + Δ s h ) F ( X s u ) D F ( X s u ) Δ s h , r G h ( s ) = G ( X s u + Δ s h ) G ( X s u ) D G ( X s u ) Δ s h .
The fundamental theorem of calculus in Banach spaces and the Lipschitz property of the derivatives give
r F h ( s ) L D , 1 2 Δ s h 2 , r G h ( s ) L 2 ( K , H ) L D , 1 2 Δ s h 2 .
For example,
r F h ( s ) = 0 1 D F ( X s u + θ Δ s h ) D F ( X s u ) Δ s h d θ ,
and integration of θ L D , 1 Δ s h 2 proves the first estimate. The proof for G is identical.
Subtracting (28) from the difference equation for Δ h gives
R t h = 0 t S ( t s ) D F ( X s u ) R s h + r F h ( s ) d s + 0 t S ( t s ) D G ( X s u ) R s h + r G h ( s ) d W s .
The deterministic convolution estimate, the Itô isometry, boundedness of the derivatives, and (32) imply
sup r t E R r h 2 C T 0 t sup q s E R q h 2 d s + C T 0 t E Δ s h 4 d s .
Gronwall’s lemma and (31) give
sup t T E R t h 2 C T h L 2 ( 0 , T ; U ) 4 .
Taking square roots proves (30). The bound is uniform over directions with small L 2 ( 0 , T ; U ) -norm, which is the Fréchet remainder condition. □
Theorem 7
(Fréchet derivative of the expected readout). Assume the hypotheses of Theorem 6. Let Φ : H R be continuously Fréchet differentiable, with bounded derivative and global Lipschitz derivative. Then,
J : L 2 ( 0 , T ; U ) R , J ( u ) = E Φ ( X T u )
is Fréchet differentiable, and
D J ( u ) h = E D Φ ( X T u ) , Z T u , h .
For each fixed u,
| J ( u + h ) J ( u ) D J ( u ) h | C T , Φ h L 2 ( 0 , T ; U ) 2 .
Proof. 
Write X = X u , Δ = X u + h X u , Z = Z u , h , and R = Δ Z . If L Φ , 1 is the Lipschitz constant of D Φ , Taylor’s formula gives
| Φ ( X T + Δ T ) Φ ( X T ) D Φ ( X T ) , Δ T | L Φ , 1 2 Δ T 2 .
Consequently,
| J ( u + h ) J ( u ) E D Φ ( X T ) , Z T | L Φ , 1 2 E Δ T 2 + D Φ E R T .
Proposition A1 and (30) bound the right-hand side by C T , Φ h L 2 ( 0 , T ; U ) 2 . This proves both claims. □
For v , w H , define the rank-one operator v w by
( v w ) z = w , z v .
It is trace class and v w 1 = v w .
Theorem 8
(Trace-norm derivative of the sentiment covariance). Under the assumptions of Theorem 6, define
m T ( u ) = E X T u , C T ( u ) = E ( X T u m T ( u ) ) ( X T u m T ( u ) ) .
Then C T : L 2 ( 0 , T ; U ) L 1 ( H ) is Fréchet differentiable. If
X ¯ T u = X T u m T ( u ) , Z ¯ T u , h = Z T u , h E Z T u , h ,
then,
D C T ( u ) h = E Z ¯ T u , h X ¯ T u + X ¯ T u Z ¯ T u , h .
Moreover,
D C T ( u ) h 1 2 X ¯ T u L 2 ( Ω ; H ) Z T u , h L 2 ( Ω ; H ) .
Proof. 
Let Δ = X T u + h X T u , Z = Z T u , h , and R = Δ Z . Centring is an orthogonal projection in L 2 ( Ω ; H ) , so it is contractive. Expanding the covariance gives
C T ( u + h ) C T ( u ) = E [ Δ ¯ X ¯ T u ] + E [ X ¯ T u Δ ¯ ] + E [ Δ ¯ Δ ¯ ] .
After subtraction of (35), the trace-norm remainder is bounded by
2 R ¯ L 2 X ¯ T u L 2 + Δ ¯ L 2 2 .
Theorem 6 and Proposition A1 make this quantity O ( h L 2 ( 0 , T ; U ) 2 ) . Formula (36) follows from the rank-one trace identity and Cauchy–Schwarz. □
Theorem 9
(Derivatives of output variance and uncertainty). Under the hypotheses of Theorem 7, set
Y ( u ) = Φ ( X T u ) , Y ˙ ( u ; h ) = D Φ ( X T u ) , Z T u , h .
The map u Y ( u ) from L 2 ( 0 , T ; U ) to L 2 ( Ω ) is Fréchet differentiable with derivative Y ˙ ( u ; h ) . The variance
V Φ ( u ) = Var [ Φ ( X T u ) ]
is Fréchet differentiable and
D V Φ ( u ) h = 2 E Y ( u ) E Y ( u ) Y ˙ ( u ; h ) .
At every u satisfying V Φ ( u ) > 0 , the uncertainty map is Fréchet differentiable and
D Unc Φ , T ( u ) h = E [ ( Y ( u ) E Y ( u ) ) Y ˙ ( u ; h ) ] Unc Φ , T ( u ) .
For Φ a ( x ) = a , x ,
D V Φ a ( u ) h = D C T ( u ) h a , a .
Proof. 
Taylor’s formula for Φ , estimate (31), and the state remainder (30) give an O ( h L 2 ( 0 , T ; U ) 2 ) remainder in L 2 ( Ω ) . Thus, Y is Fréchet differentiable. Let P 0 be the orthogonal projection of L 2 ( Ω ) onto the zero-mean subspace. Since
V Φ ( u ) = P 0 Y ( u ) L 2 2 ,
the Hilbert-space chain rule yields
D V Φ ( u ) h = 2 P 0 Y ( u ) , P 0 Y ˙ ( u ; h ) L 2 .
The first factor has zero mean, so the mean of Y ˙ gives no contribution. This proves (37). Formula (38) follows from the scalar chain rule on the positive half-line. The final identity follows from Theorem 8. □
The derivative formula becomes explicit when the dynamics are linear and the noise is additive, and this explicit representation is the core interpretability statement of the framework.
Assume that Σ L 2 ( K , H ) and
d X t = ( A X t + B u t ) d t + Σ d W t , X 0 = x ,
with deterministic u L 2 ( 0 , T ; U ) . Let Φ : H R be Fréchet differentiable with bounded derivative.
Theorem 10
(Semantic influence kernel). In the linear additive setting, the derivative of
J ( u ) = E Φ ( X T u )
in the direction h L 2 ( 0 , T ; U ) is
D J ( u ) h = 0 T I s , h s U d s ,
where the influence kernel is
I s = B * S ( T s ) * g T , g T = E [ D Φ ( X T u ) ] H .
If Φ a ( x ) = a , x , then
I s = B * S ( T s ) * a .
Proof. 
For the linear additive equation, the sensitivity process has no stochastic part and satisfies
Z T h = 0 T S ( T s ) B h s d s .
By Theorem 7,
D J ( u ) h = E D Φ ( X T u ) , 0 T S ( T s ) B h s d s H = 0 T E S ( T s ) * D Φ ( X T u ) , B h s H d s = 0 T B * S ( T s ) * E [ D Φ ( X T u ) ] , h s U d s .
The interchange of expectation and time integration is justified by the boundedness of D Φ , the semigroup bound, and h L 2 ( 0 , T ; U ) . This proves (39) and (40). For Φ a ( x ) = a , x , D Φ a ( x ) = a for all x, so g T = a , giving (41). □
Corollary 4
(Spectral influence attribution). Assume that A is self-adjoint, and A e k = λ k e k . Let Φ a ( x ) = a , x , where a = k 1 a k e k . Then,
I s = k 1 a k e λ k ( T s ) B * e k
in U for every s [ 0 , T ] . The series defines an element of L 2 ( 0 , T ; U ) and satisfies
I L 2 ( 0 , T ; U ) 2 B 2 k 1 a k 2 1 e 2 λ k T 2 λ k T B 2 a 2 .
Proof. 
The spectral theorem gives
S ( T s ) a = k 1 a k e λ k ( T s ) e k
in H. Since B * : H U is bounded, its application to the partial sums preserves convergence and proves (42). Moreover,
I s U 2 B 2 S ( T s ) a 2 = B 2 k 1 a k 2 e 2 λ k ( T s ) .
Tonelli’s theorem and direct integration prove (43). □
Remark 2
(Semantic interpretation of spectral influence). When A is obtained from a semantic graph Laplacian, λ k measures the variation of e k across adjacent semantic prototypes. The factor e λ k ( T s ) suppresses graph-oscillatory directions when the forcing time s lies far from the terminal horizon. A high-frequency semantic component may still be influential near T, or when the readout coefficient and semantic lift strongly activate that component. This statement concerns graph geometry and has no relation to the empirical frequency of a token.
The influence kernel has an exact robustness meaning. Put L 2 ( 0 , T ; U ) = L 2 ( 0 , T ; U ) , endowed with its Hilbert norm, and define the controllability operator
K T h = 0 T S ( T s ) B h s d s , h L 2 ( 0 , T ; U ) .
The semigroup estimate and the Cauchy–Schwarz inequality show that K T L ( L 2 ( 0 , T ; U ) , H ) . Its adjoint satisfies
K T * a s = B * S ( T s ) * a for almost every s [ 0 , T ] .
Theorem 11
(Exact influence-robustness duality). Consider the linear additive equation and the linear readout Φ a ( x ) = a , x . Let
J T ( u ) = E a , X T u , u L 2 ( 0 , T ; U ) ,
and let I = K T * a be the influence kernel in (41). For every radius ρ 0 ,
sup h L 2 ( 0 , T ; U ) ρ | J T ( u + h ) J T ( u ) | = ρ I L 2 ( 0 , T ; U ) .
If I 0 , the supremum is attained by
h ρ ± = ± ρ I I L 2 ( 0 , T ; U ) .
If I = 0 , the output is invariant under every perturbation in L 2 ( 0 , T ; U ) .
Proof. 
The mild formula gives
X T u + h X T u = K T h .
The difference is deterministic because the two fields are driven by the same additive noise. Hence,
J T ( u + h ) J T ( u ) = a , K T h H = K T * a , h L 2 ( 0 , T ; U ) = 0 T I s , h s U d s .
The Cauchy–Schwarz inequality yields
| J T ( u + h ) J T ( u ) | I L 2 ( 0 , T ; U ) h L 2 ( 0 , T ; U ) ρ I L 2 ( 0 , T ; U ) .
When I 0 , substitution of either direction in (47) gives equality. When I = 0 , the adjoint identity implies a , K T h = 0 for every h, which proves invariance. □
The exact identity extends locally to nonlinear readouts and nonlinear dynamics whenever the expected output is Fréchet differentiable. For u L 2 ( 0 , T ; U ) , define
r u ( ρ ) = sup h L 2 ( 0 , T ; U ) ρ | J ( u + h ) J ( u ) | .
Proposition 2
(Local nonlinear duality). Assume that J : L 2 ( 0 , T ; U ) R is Fréchet differentiable at u. Let g u L 2 ( 0 , T ; U ) be the Riesz representative of D J ( u ) . Then,
lim ρ 0 r u ( ρ ) ρ = D J ( u ) L 2 ( 0 , T ; U ) * = g u L 2 ( 0 , T ; U ) .
If the dynamics are linear and additive, then g u = I . Under the differentiability assumptions of Theorem 7, the same conclusion holds with g u determined by the linearised SPDE in (28).
Proof. 
Fréchet differentiability gives a function ε u : [ 0 , ) [ 0 , ) , with ε u ( r ) 0 as r 0 , such that
| J ( u + h ) J ( u ) D J ( u ) h | ε u ( h L 2 ( 0 , T ; U ) ) h L 2 ( 0 , T ; U ) .
For h L 2 ( 0 , T ; U ) ρ ,
| J ( u + h ) J ( u ) | ρ D J ( u ) L 2 ( 0 , T ; U ) * + ρ ε u ( ρ ) .
Taking the supremum gives
lim   sup ρ 0 r u ( ρ ) ρ D J ( u ) L 2 ( 0 , T ; U ) * .
If g u 0 , choose h ρ = ρ g u / g u L 2 ( 0 , T ; U ) . Then,
| J ( u + h ρ ) J ( u ) | | D J ( u ) h ρ | ρ ε u ( ρ ) = ρ g u L 2 ( 0 , T ; U ) ρ ε u ( ρ ) .
This proves the matching lower limit. The case g u = 0 follows from the upper estimate. The Riesz theorem gives D J ( u ) L 2 ( 0 , T ; U ) * = g u L 2 ( 0 , T ; U ) . In the linear additive case, Theorem 10 identifies the representative with I . □
The finite-dimensional approximation can be quantified directly at the level of the diagnostics. Let P N denote the orthogonal projection onto span { e 1 , , e N } .
Theorem 12
(Exact spectral truncation error for uncertainty and influence). Assume that A is self-adjoint with A e k = λ k e k , that Q e k = q k e k , and that Φ a ( x ) = a , x . Let V T be the terminal variance in (23), and let V T ( N ) be the variance obtained after retaining the first N modes. Then,
V T V T ( N ) = k > N a k 2 q k 1 e 2 λ k T 2 λ k .
Assume in addition that U = H and B e k = b k e k . Let I ( N ) = P N I . Then,
I I ( N ) L 2 ( 0 , T ; H ) 2 = k > N a k 2 b k 2 1 e 2 λ k T 2 λ k .
If a Dom ( ( A ) r ) for some r 0 , q k q ¯ , and | b k |     b ¯ , then,
0 V T V T ( N ) q ¯ 2 λ N + 1 2 r + 1 ( A ) r a 2
and
I I ( N ) L 2 ( 0 , T ; H ) 2 b ¯ 2 2 λ N + 1 2 r + 1 ( A ) r a 2 .
Proof. 
Formula (50) is obtained by subtracting the first N summands from (23). Under the diagonal lifting assumption,
I s = k 1 a k b k e λ k ( T s ) e k .
Orthogonality, Tonelli’s theorem, and the change in variable r = T s give
I I ( N ) L 2 ( 0 , T ; H ) 2 = 0 T k > N a k 2 b k 2 e 2 λ k ( T s ) d s = k > N a k 2 b k 2 1 e 2 λ k T 2 λ k ,
which proves (51). Since 1 e 2 λ k T 1 ,
V T V T ( N ) q ¯ 2 k > N λ k 2 r 1 λ k 2 r a k 2 q ¯ 2 λ N + 1 2 r + 1 k > N λ k 2 r a k 2 .
The final sum does not exceed ( A ) r a 2 . The same argument with b ¯ 2 in place of q ¯ proves (53). □
The estimates identify a concrete approximation criterion. A Galerkin dimension is sufficient for a chosen readout once the omitted variance and influence energy fall below the prescribed tolerances. The spectral regularity of the readout determines the rate. The ambient dimension alone does not determine it.
Robustness is expressed as quantitative stability under perturbations, and the estimates below are deterministic inequalities at the level of laws and derived functionals. Quadratic Wasserstein distance is used in its standard form [38], so that the estimates provide theoretical certification independent of numerical experiments.
Consider two stochastic semantic models driven by the same cylindrical Wiener process. Their generators may be different and generate semigroups S and S ˜ . In mild form, the models are
X t = S ( t ) x + 0 t S ( t s ) [ F ( X s ) + B u s ] d s + 0 t S ( t s ) G ( X s ) d W s
and
X ˜ t = S ˜ ( t ) x ˜ + 0 t S ˜ ( t s ) [ F ˜ ( X ˜ s ) + B ˜ u ˜ s ] d s + 0 t S ˜ ( t s ) G ˜ ( X ˜ s ) d W s .
Assume that both semigroups satisfy
max { S ( t ) , S ˜ ( t ) } M e ω t , t [ 0 , T ]
and that both pairs of coefficients satisfy Assumption 2 with common constants. Define
d S ( T ) = sup 0 r T S ( r ) S ˜ ( r ) L ( H )
and the along-trajectory coefficient discrepancies
d F 2 = E 0 T F ( X ˜ s ) F ˜ ( X ˜ s ) 2 d s , d G 2 = E 0 T G ( X ˜ s ) G ˜ ( X ˜ s ) L 2 ( K , H ) 2 d s .
The reference-model moment budget is
M T ( Θ ˜ ) = E x ˜ 2 + E 0 T 1 + X ˜ s 2 + u ˜ s U 2 d s ,
where Θ ˜ denotes the collection of ingredients in (55). The aggregate model discrepancy is
d T ( Θ , Θ ˜ ) 2 = E x x ˜ 2 + E 0 T u s u ˜ s U 2 d s + B B ˜ 2 E 0 T u ˜ s U 2 d s + d F 2 + d G 2 + d S ( T ) 2 M T ( Θ ˜ ) .
The use of along-trajectory discrepancies avoids the unnecessarily restrictive assumption that two nonlinear coefficients remain uniformly close on the whole infinite-dimensional state space.
Theorem 13
(Robustness under structural model perturbations). Under the preceding assumptions, there is a constant C T depending only on T, the common semigroup bound, the common Lipschitz and growth constants, B , and B ˜ , such that
sup t [ 0 , T ] E X t X ˜ t 2 C T d T ( Θ , Θ ˜ ) 2 .
Thus, the solution is stable under simultaneous perturbations of the initial field, forcing path, generator, semantic lift, drift, and diffusion.
Proof. 
Set D t = X t X ˜ t . Add and subtract terms so that the Lipschitz differences are always evaluated under S. The two mild equations give
D t = S ( t ) ( x x ˜ ) + [ S ( t ) S ˜ ( t ) ] x ˜ + 0 t S ( t s ) [ F ( X s ) F ( X ˜ s ) ] d s + 0 t S ( t s ) [ F ( X ˜ s ) F ˜ ( X ˜ s ) ] d s + 0 t S ( t s ) B ( u s u ˜ s ) d s + 0 t S ( t s ) ( B B ˜ ) u ˜ s d s + 0 t [ S ( t s ) S ˜ ( t s ) ] [ F ˜ ( X ˜ s ) + B ˜ u ˜ s ] d s + 0 t S ( t s ) [ G ( X s ) G ( X ˜ s ) ] d W s + 0 t S ( t s ) [ G ( X ˜ s ) G ˜ ( X ˜ s ) ] d W s + 0 t [ S ( t s ) S ˜ ( t s ) ] G ˜ ( X ˜ s ) d W s ,
for ten vectors, j = 1 10 z j 2 10 j = 1 10 z j 2 . Then, exploiting the semigroup bound, Cauchy–Schwarz in time, and the Itô isometry, it holds
E D t 2 C T E x x ˜ 2 + C T d S ( T ) 2 E x ˜ 2 + C T 0 t E D s 2 d s + C T d F 2 + C T d G 2 + C T E 0 T u s u ˜ s U 2 d s + C T B B ˜ 2 E 0 T u ˜ s U 2 d s + C T d S ( T ) 2 E 0 t F ˜ ( X ˜ s ) + B ˜ u ˜ s 2 + G ˜ ( X ˜ s ) L 2 2 d s .
The common linear-growth condition gives
F ˜ ( X ˜ s ) + B ˜ u ˜ s 2 + G ˜ ( X ˜ s ) L 2 2 C 1 + X ˜ s 2 + u ˜ s U 2 .
Consequently,
E D t 2 C T d T ( Θ , Θ ˜ ) 2 + C T 0 t E D s 2 d s .
Taking the supremum over [ 0 , t ] and applying Gronwall’s lemma proves (60). □
Corollary 5
(Robustness of laws, predictions, and output uncertainty). Under the assumptions of Theorem 13, let Φ : H R be Lipschitz with constant L Φ . Then,
W 2 Law ( X T ) , Law ( X ˜ T ) C T 1 / 2 d T ( Θ , Θ ˜ ) ,
| E Φ ( X T ) E Φ ( X ˜ T ) | L Φ C T 1 / 2 d T ( Θ , Θ ˜ ) ,
and
Var [ Φ ( X T ) ] 1 / 2 Var [ Φ ( X ˜ T ) ] 1 / 2 L Φ C T 1 / 2 d T ( Θ , Θ ˜ ) .
Proof. 
The pair ( X T , X ˜ T ) is an admissible coupling of the two terminal laws. The definition of quadratic Wasserstein distance and Theorem 13 give (61). The Lipschitz property and Cauchy–Schwarz give
| E Φ ( X T ) E Φ ( X ˜ T ) | L Φ X T X ˜ T L 2 ( Ω ; H ) ,
which proves (62). For any Y L 2 ( Ω ) , its standard deviation is the distance from Y to the closed subspace of constants. The reverse triangle inequality for distances to a closed set provides
| Var Y Var Y ˜ |     Y Y ˜ L 2 ( Ω ) .
Apply this inequality to Y = Φ ( X T ) and Y ˜ = Φ ( X ˜ T ) , then use the Lipschitz property and Theorem 13. □
The next estimate is independent of any Gaussian or linear structure and transfers solution robustness to the full latent covariance operator.
Theorem 14
(Nonlinear trace-norm covariance robustness). Let Y , Z L 2 ( Ω ; H ) , with covariance operators C Y and C Z . Then,
C Y C Z 1 Y L 2 ( Ω ; H ) + Z L 2 ( Ω ; H ) Y Z L 2 ( Ω ; H ) .
Consequently, the terminal covariances of (54) and (55) satisfy
C X T C X ˜ T 1 C T 1 / 2 X T L 2 + X ˜ T L 2 d T ( Θ , Θ ˜ ) .
Proof. 
Let Y ¯ = Y E Y and Z ¯ = Z E Z . Then,
C Y C Z = E [ ( Y ¯ Z ¯ ) Y ¯ ] + E [ Z ¯ ( Y ¯ Z ¯ ) ] .
The rank-one trace identity and Cauchy–Schwarz imply
C Y C Z 1 Y ¯ Z ¯ L 2 Y ¯ L 2 + Z ¯ L 2 .
Centring is an orthogonal projection in L 2 ( Ω ; H ) and therefore a contraction. This proves (64); Equation (65) follows from Theorem 13. □
Consider now two linear Gaussian fields with semigroups S 1 and S 2 and covariance injections Q 1 = Σ 1 Σ 1 * and Q 2 = Σ 2 Σ 2 * . Let
C t ( i ) = 0 t S i ( r ) Q i S i ( r ) * d r , i = 1 , 2 .
Theorem 15
(Robustness of linear covariance-based uncertainty). Assume S i ( t )   e α t for the same α > 0 , and set
d S , 12 ( t ) = sup 0 r t S 1 ( r ) S 2 ( r ) .
Then,
C t ( 1 ) C t ( 2 ) 1 1 e 2 α t 2 α Q 1 Q 2 1 + 2 ( 1 e α t ) α d S , 12 ( t ) Q 2 1 .
For every a H ,
| Var a , X t ( 1 ) Var a , X t ( 2 ) | a 2 C t ( 1 ) C t ( 2 ) 1 .
Proof. 
For each r, add and subtract S 1 ( r ) Q 2 S 1 ( r ) * and S 2 ( r ) Q 2 S 1 ( r ) * . This gives
S 1 ( r ) Q 1 S 1 ( r ) * S 2 ( r ) Q 2 S 2 ( r ) * = S 1 ( r ) ( Q 1 Q 2 ) S 1 ( r ) * + [ S 1 ( r ) S 2 ( r ) ] Q 2 S 1 ( r ) * + S 2 ( r ) Q 2 [ S 1 ( r ) * S 2 ( r ) * ] .
The trace ideal property bounds the first term by e 2 α r Q 1 Q 2 1 and the sum of the final two terms by 2 d S , 12 ( t ) e α r Q 2 1 . Integration over [ 0 , t ] proves (66). The scalar variance estimate follows from Corollary 2 and | T a , a | T 1 a 2 . □
Influence kernels should remain stable when the semantic lift, generator, or readout is perturbed. The next theorem includes all three mechanisms.
Theorem 16
(Robustness of linear influence kernels). Assume S ( t ) e α t and S ˜ ( t ) e α t . Let
I s = B * S ( T s ) * a , I ˜ s = B ˜ * S ˜ ( T s ) * a ˜ .
With d S ( T ) defined by (56),
I s I ˜ s U e α ( T s ) B B ˜ a + B ˜ d S ( T ) a + e α ( T s ) B ˜ a a ˜ .
Consequently,
I I ˜ L 2 ( 0 , T ; U ) 1 e 2 α T 2 α 1 / 2 B B ˜ a + B ˜ a a ˜ + T 1 / 2 B ˜ d S ( T ) a .
Proof. 
Add and subtract B ˜ * S ( T s ) * a and B ˜ * S ˜ ( T s ) * a to obtain
I s I ˜ s = ( B * B ˜ * ) S ( T s ) * a + B ˜ * [ S ( T s ) * S ˜ ( T s ) * ] a + B ˜ * S ˜ ( T s ) * ( a a ˜ ) .
Taking norms proves (68). Integrating the first and third terms against e 2 α ( T s ) and the middle term against one proves (69) by Minkowski’s inequality. □
Proposition 3
(Certified robustness moduli). Fix a reference model Θ and an admissible class whose semigroup, Lipschitz, growth, and operator bounds are uniform. Let
P ρ ( Θ ) = { Θ ˜ : d T ( Θ , Θ ˜ ) ρ } .
For a Lipschitz readout Φ, every model in P ρ ( Θ ) satisfies
W 2 ( Law ( X T ) , Law ( X ˜ T ) ) C T 1 / 2 ρ ,
| E Φ ( X T ) E Φ ( X ˜ T ) | L Φ C T 1 / 2 ρ ,
and
Var [ Φ ( X T ) ] Var [ Φ ( X ˜ T ) ] L Φ C T 1 / 2   ρ .
Hence, the robustness moduli of Definition 2 are finite and at most linear in the perturbation radius.
Proof. 
The three inequalities are the uniform versions of Corollary 5 over the defining ball P ρ ( Θ ) . Taking the corresponding suprema proves the statement. □
The preceding results give a single operator calculus. The terminal covariance controls the dispersion of the readout. Its derivative measures the first-order change of the uncertainty operator under a semantic intervention. The adjoint influence kernel represents the derivative of the expected output, and its Hilbert norm gives the exact local robustness modulus. Structural perturbation estimates then control the same quantities when the generator, semantic lift, nonlinear coefficients, forcing path, or covariance injection are changed. The truncation theorem determines how accurately a finite semantic basis preserves these diagnostics.

3. Implementation and Calibration

Figure 1 shows the passage from text to the three diagnostics. A corpus provides time-indexed documents. An encoder maps each document into a finite semantic representation. Aggregation over a prescribed time grid produces the forcing path. A Galerkin model propagates the latent state, while an observation equation links the state to ratings, labels, lexicon scores, or external responses. Filtering and likelihood estimation recover finite operators whose covariance and adjoint sensitivity approximate the infinite-dimensional quantities.
Let { d j , t j , m j } j = 1 N doc be a time-indexed corpus. The symbol d j denotes a document, and m j records its source together with any available annotation. A fixed encoder E maps d j to z j R p . The encoder may be a transformer representation, a task-adapted language model, or an interpretable lexicon map. For a grid 0 = τ 0 < τ 1 < < τ L , the forcing vector on the interval [ τ , τ + 1 ) is
u = j : t j [ τ , τ + 1 ) a j z j .
The deterministic weights a j may correct source imbalance or annotation reliability. Their normalisation is fixed before estimation in order to remove the scale ambiguity between u and the lifting operator.
A semantic graph can be formed from encoded documents or from a set of semantic prototypes. Let w i j 0 be a symmetric similarity weight, and let L n be the corresponding graph Laplacian. If e k is a unit eigenvector with eigenvalue λ k , then
e k T L n e k = 1 2 i , j w i j e k ( i ) e k ( j ) 2 = λ k .
The eigenvalue is therefore the Dirichlet energy of the coordinate. Large values indicate rapid variation across semantically adjacent prototypes. Selecting the first n eigenvectors produces a low-frequency Galerkin space H n . The truncation error for uncertainty and influence is controlled by Theorem 12.
Let P n : H H n be the orthogonal projection. A consistent projected equation has the form
d X t ( n ) = A n X t ( n ) + F n ( X t ( n ) ) + B n u t d t + G n ( X t ( n ) ) d W t ( n ) .
When H n Dom ( A ) and the basis is domain compatible, one may take A n = P n A | H n . A data-adapted basis requires a separately constructed generator whose semigroup converges consistently to S. The approximation theorem in Appendix A states a sufficient consistency condition.
In the linear additive regime, write Q n = G n G n T . If the forcing is constant on an interval of length Δ , the exact discrete transition is
x + 1 = F x + D u + ξ , ξ N ( 0 , Ω ) ,
where
F = e A n Δ , D = 0 Δ e A n s B n d s , Ω = 0 Δ e A n s Q n e A n T s d s .
The use of the exact transition removes time-discretisation bias from the linear calibration problem. The observation equation is
y = R n x + ε , ε N ( 0 , V n ) .
The matrix R n links the latent state to the observed sentiment proxy. The covariance V n represents annotation or measurement variation, while Q n represents unresolved variation in the latent evolution.
For fixed parameters, the Kalman recursion supplies the likelihood and the filtered state covariance. With the standard notation for conditional means and covariances,
x ^ + 1 | = F x ^ | + D u , P + 1 | = F P | F T + Ω .
The innovation covariance and update are
S + 1 = R n P + 1 | R n T + V n , K + 1 = P + 1 | R n T S + 1 1 , x ^ + 1 | + 1 = x ^ + 1 | + K + 1 y + 1 R n x ^ + 1 | .
Up to an additive constant, the Gaussian negative log-likelihood is
L ( θ ) = 1 2 log det S + ν T S 1 ν , ν = y R n x ^ | 1 .
The parameter vector contains the finite representations of A, B, the process covariance, the observation map, and the observation covariance. A dissipative parameterisation is
A n = L A L A T α 0 I n + K A , K A T = K A , α 0 > 0 .
Its symmetric part is bounded above by α 0 I n . Cholesky factors enforce positive semidefiniteness of Q n and V n . A dense Kalman likelihood costs O ( L n 3 ) . Diagonal spectral dynamics reduce state propagation to O ( L n ) , while low-rank observation maps permit intermediate complexity through matrix identities.
The operator B n is estimated from the conditional response to the aggregated text forcing. The generator A n is estimated under the dissipative constraint. A nonlinear drift can be parameterised by a differentiable map whose Jacobian has a controlled one-sided Lipschitz constant, thereby preserving a positive dissipativity margin. The process covariance is estimated from the residual state innovation after deterministic forcing has been accounted for. The observation covariance is identified from replicated annotation, known label reliability, or an explicit restriction separating measurement noise from process noise.
Proposition 4
(Identifiability of a linear projected model). Assume that R n = I n , that V n is known, and that the sampling interval is a fixed number Δ > 0 . Suppose that the collection of conditional laws
Law y + 1 x = x , u = u , x R n , u R p
is known. Assume that F = e A n Δ has no eigenvalue on the closed negative real axis and that
σ ( Δ A n ) { z C : | Im z | < π } .
Assume further that 0 Δ e A n s d s is invertible. Then, the conditional laws determine A n , B n , and Ω Δ . If the map
K Δ : Q 0 Δ e A n s Q e A n T s d s
is injective on the chosen covariance family, then the same laws determine Q n . For finite-sample estimation of F and D , the corresponding regressor matrix with rows ( x T , u T ) must have full column rank.
Proof. 
The conditional law is Gaussian with mean
F x + D u
and covariance Ω Δ + V n . Equality of the conditional means for every x and u identifies F and D . The spectral assumptions place Δ A n in the uniqueness strip of the principal matrix logarithm, so
A n = Δ 1 log F .
The identity
D = 0 Δ e A n s d s B n
and invertibility of the bracketed matrix identify B n . The conditional covariance identifies Ω Δ because V n is known. Injectivity of K Δ then identifies Q n . The final rank condition is the standard uniqueness condition for estimating the coefficients of the conditional mean from finitely many observed pairs. □
The candidate corpora play distinct roles. Time-stamped social media data can construct u and provide noisy labels. Review corpora provide longer evaluative units and document-level observations. Lexica provide externally anchored coordinates for H n or regularisation for R n . These resources do not determine the operators automatically. They provide the observations, entering the likelihood, and the geometry, entering the basis construction.
The comparison with intelligent and structure-preserving solvers is made at the level of preserved objects. AI-enhanced symbolic derivation can propose closed-form identities and algebraic factorisations [31]. Physics-informed inverse methods can estimate projected coefficients from observations [28]. PDE denoising motivates a dissipative filtering stage [27]. Lie-group methods show how a numerical scheme can be designed around a geometric invariant [29]. Riemann–Hilbert analysis shows the value of exact spectral structure when integrability is present [30]. For the semantic-field equation, dissipativity is the discrete target. Hamiltonian energy preservation addresses a different geometry. A reliable high-order method should preserve contractivity up to a controlled defect, maintain positive semidefinite covariance, and converge jointly for the forward state and the adjoint influence equation. Neural SPDE solvers provide an alternative when the Galerkin dimension makes dense filtering impractical [25,26]. The covariance, duality, and truncation identities in Section 2 supply verification criteria for such solvers.

4. Synthetic Numerical Study

The numerical study uses a three-mode linear Galerkin model for which every transition and diagnostic is available in closed form. Its purpose is to test internal consistency, calibration feasibility, and the magnitude of the theoretical constants. The study does not represent a benchmark on natural language data.
Let X t R 3 solve
d X t = Λ X t + B u t d t + Q 1 / 2 d W t , Y t = c T X t + ε t ,
where
Λ = diag ( 0.45 , 1.10 , 2.40 ) , B = 0.90 0.20 0.25 0.70 0.15 0.45 , Q = diag ( 0.040 , 0.035 , 0.025 ) , c = ( 0.65 , 0.55 , 0.35 ) T .
The observation noise has standard deviation 0.07 . The time step is Δ = 0.15 , the number of transitions is 80, and the horizon is T = 12 . Each forcing coordinate combines a smooth oscillatory component with a localised pulse. Random amplitudes, phases, and pulse centres create trajectory-level heterogeneity. Initial states are Gaussian. The initial state, Brownian increments, observation errors, and random forcing parameters are mutually independent across trajectories. The random seed is 18082026.
The exact matrices in (74) are used for simulation. Twenty trajectories form the training sample, eighty independent trajectories form the conformal calibration sample, and four hundred independent trajectories form the test sample. The structured estimator uses a diagonal transition matrix with entries constrained to ( 0 , 1 ) , together with an unrestricted forcing matrix. A ridge regression estimates each modal transition and the two forcing coefficients. The readout is estimated by ridge regression on the latent states. The residual covariance estimates Ω Δ . The unrestricted vector autoregression with exogenous forcing and the persistence predictor provide finite-dimensional baselines.
The estimated damping rates are
0.470754 , 1.023585 , 2.370493 .
The corresponding generating values are 0.45 , 1.10 , and 2.40 . The estimated readout is
c ^ = ( 0.654022 , 0.551861 , 0.338787 ) T .
Table 1 reports one-step performance on the test set.
For each test trajectory, the readout RMSE is computed over the eighty one-step forecasts. A paired bootstrap with five thousand resamples gives a mean structured-minus-VARX difference of 9.6960 × 10 6 , with percentile interval
[ 1.3130 × 10 5 , 3.3299 × 10 5 ] .
The two-sided Wilcoxon signed-rank p-value is 0.4494 . The structured restriction therefore causes no detectable loss relative to the unrestricted VARX model in this experiment. The structured-minus-persistence difference is 9.4908 × 10 3 , with interval
[ 1.0052 × 10 2 , 8.9277 × 10 3 ] ,
and the two-sided Wilcoxon signed-rank p-value is 1.30 × 10 66 . These tests concern synthetic trajectories generated from (80). They do not support a claim about performance on a natural language benchmark.
The estimated one-step predictive standard deviation is 0.094752 . Coverage is evaluated at the final one-step forecast of each independent test trajectory, so the calibration scores and the test scores are exchangeable conditional on the fitted predictor. A nominal Gaussian interval at level 0.90 has empirical coverage 0.8800 , with Wilson interval [ 0.8445 , 0.9083 ] and mean width 0.31170 . Split conformal calibration uses the eighty absolute standardised terminal residuals. The order statistic multiplier is 1.81583 . The resulting interval has empirical coverage 0.9175 , with Wilson interval [ 0.8864 , 0.9407 ] and mean width 0.34411 . Proposition A4 in Appendix B gives the finite-sample marginal coverage statement. Figure 2 displays one representative trajectory and the interval obtained from the same conformal multiplier.
For the continuous model, the influence kernel of the terminal linear readout is
I s = B T e Λ ( T s ) c .
Numerical quadrature gives
I L 2 ( 0 , T ; R 2 ) = 0.752295 .
For the input-energy radius ρ = 0.25 , Theorem 11 gives the exact worst-case terminal displacement
ρ I L 2 = 0.1880738 .
Substitution of the normalised extremal direction in (47) gives 0.1880738 by numerical quadrature. The agreement verifies the dual identity to the displayed precision. Figure 3 shows how the two forcing coordinates contribute over time.
The exact variance contributions at horizon T are
0.0187774 , 0.0048125 , 0.0006380 .
Their shares are 77.50 percent, 19.86 percent, and 2.63 percent. For the identity lift diagnostic associated with Theorem 12, the corresponding influence-energy shares are 74.22 percent, 21.74 percent, and 4.04 percent. The first two modes retain 97.37 percent of the output variance and 95.96 percent of the identity-lift influence energy. Figure 4 displays these quantities. The third coordinate represents local graph variation, so its small long-horizon contribution is the concrete counterpart of the high-frequency damping statement.
The analytic ablation in Theorem 2 is evaluated by multiplying every damping rate by 0.5 , 1, and 2. Table 2, see also Figure 5, shows the predictive standard deviation, the exact influence norm, and the semigroup factor
κ T ( λ 1 ) = 1 e 2 λ 1 T 2 λ 1 1 / 2 .
Stronger damping lowers both uncertainty and influence. The values demonstrate that the dependence identified by the theorem is numerically informative in this regime.
A second ablation keeps Tr Q fixed and changes the noise geometry. The predictive standard deviation is 0.170669 for diagonal noise, 0.151759 for the stated positively correlated covariance, and 0.161189 for isotropic noise with equal trace. The differences show that Tr Q alone cannot determine output uncertainty. Alignment with the readout and the damping spectrum is decisive.
The experiment gives an operational interpretation of the mathematical quantities. The covariance determines a predictive dispersion and a Gaussian reference interval. The influence kernel identifies the times and forcing directions that move the terminal mean. Its norm gives the exact energy-ball robustness radius in the linear regime. The truncation formula states how much uncertainty and influence are lost when modes are removed. The dissipativity margin controls the scale of all three diagnostics.

Limitations

The numerical evidence is synthetic and low dimensional. It verifies identities and tractability under controlled conditions, while leaving real-language validity unresolved. The training trajectories expose the latent state, whereas a corpus application would infer that state through the observation model. The global Lipschitz assumptions exclude discontinuous threshold responses and polynomial drifts without truncation or variational reformulation. The exact covariance study uses additive Gaussian noise. Heavy-tailed jumps, state-dependent non-Gaussian perturbations, and annotation processes with temporal dependence require different uncertainty tools. Split conformal coverage relies on exchangeability of the independent trajectory-level calibration and test scores. Distribution shift invalidates that guarantee unless a shift-aware conformal method is used. A large Galerkin basis can produce ill-conditioned likelihoods and weak identifiability when the forcing lacks persistent excitation. The synthetic baseline comparison does not establish superiority over transformer classifiers, large language models, conformal language model methods, or post hoc explanation algorithms. Such a claim requires a common real-data task, a fixed encoder, nested hyperparameter selection, and external benchmark evaluation.

5. Conclusions and Future Developments

A stochastic semantic field provides one evolution law for predictive uncertainty, forcing sensitivity, and perturbation stability. The covariance operator propagates latent variability to a terminal readout. The forcing-to-state map is Fréchet differentiable with a quadratic remainder, so the adjoint influence kernel represents a genuine bounded derivative. In the linear regime, the norm of that kernel equals the exact worst-case displacement over an input-energy ball. The spectral truncation formula then quantifies how finite semantic coordinates approximate both uncertainty and influence.
The projected model admits an exact continuous-to-discrete transition in the linear Gaussian case. This transition supports likelihood estimation and filtering without introducing a time-stepping error into calibration. The three-mode study confirms the covariance and duality identities, exhibits non-vacuous robustness constants, and shows how dissipativity changes predictive dispersion and influence. Split conformal calibration corrects the finite-sample coverage of the Gaussian interval in the synthetic experiment. These findings establish mathematical coherence and computational feasibility at the scale studied here.
A real-data investigation should begin with a fixed text encoder and a time-indexed corpus. The semantic graph, forcing aggregation, observation map, and projected dimension should be selected inside a nested validation design. Comparisons with transformer and large language model baselines should use a common predictive target and should report calibration, coverage, explanation stability, and perturbation response together with ordinary predictive loss. Distribution shift requires weighted or adaptive conformal methods. Partially observed latent states call for smoothing and data assimilation beyond the fully observed synthetic calibration.
The global Lipschitz setting can be extended through locally monotone or maximal monotone methods. Lévy noise can represent abrupt semantic shocks, while delay equations can model persistent influence from earlier information. Statistical consistency under a growing time horizon and a growing Galerkin dimension remains an important inverse problem. High-order numerical methods should preserve the dissipative structure and covariance positivity that enter the present estimates. Neural stochastic PDE solvers may reduce computational cost when their approximation errors are controlled in the state and adjoint equations.
Artificial-intelligence-enhanced mathematical derivation offers a further direction. A symbolic system guided by learned search can propose Lyapunov functionals, semigroup factorisations, covariance identities, and candidate adjoint representations. Every proposed identity must still be verified in the operator topology required by the theorem. Domain invariance for an unbounded generator, trace-class convergence, adaptedness of stochastic integrands, and passage from Yosida approximations to mild solutions remain mathematical obligations. The appropriate role of an automated derivation system is therefore the generation and checking of algebraic candidates within a proof workflow whose analytic hypotheses remain explicit.
The framework consequently defines a research programme in which infinite dimensional analysis determines the quantities that a finite estimator must preserve, numerical calibration makes those quantities observable, and external benchmarks determine whether the resulting representation is useful for a particular language domain.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/computation14090216/s1.

Funding

This research received no external funding.

Data Availability Statement

The synthetic data used are solely the ones generated by the exact state transition stated in Section 4. The simulation script, the numerical summaries, and the figure-generation code accompany this article as Supplementary Material.

Conflicts of Interest

The author declares no conflicts of interest.

Appendix A. Standard Hilbert-Space, Stochastic Evolution, and Approximation Facts

This appendix collects the background results that are standard in Hilbert-space stochastic calculus and semigroup theory, but are stated and proved in the notation of the paper in order to keep the manuscript self-contained. The underlying theory is developed in compatible formulations in [19,20,21,39,40]. The main text cites this appendix whenever one of these standard facts enters a proof.
Lemma A1
(Deterministic convolution estimate). Let a L 2 ( Ω × [ 0 , T ] , P ; H ) . Define
Y t = 0 t S ( t s ) a s d s .
Then, Y is mean-square continuous and
sup t [ 0 , T ] E Y t 2 T M 2 e 2 ω + T E 0 T a s 2 d s .
Proof. 
Fix t [ 0 , T ] . By the Cauchy–Schwarz inequality in time and by Assumption 1,
Y t 2 = 0 t S ( t s ) a s d s 2 t 0 t S ( t s ) a s 2 d s T M 2 e 2 ω + T 0 T a s 2 d s .
Taking expectations and the supremum over t gives the estimate.
It remains to prove mean-square continuity. Let t n t . Assume first t n t . Then,
Y t n Y t = 0 t S ( t n s ) S ( t s ) a s d s + t t n S ( t n s ) a s d s .
The first term converges to zero in L 2 ( Ω ; H ) because S ( r ) z is continuous in r for every z H and because
[ S ( t n s ) S ( t s ) ] a s 2 4 M 2 e 2 ω + T a s 2 .
Dominated convergence in Ω × [ 0 , t ] applies. The second term is bounded in mean square by
( t n t ) M 2 e 2 ω + T E t t n a s 2 d s ,
which converges to zero. The case t n t is identical. Hence, Y is mean-square continuous. □
Lemma A2
(Stochastic convolution estimate). Let Ψ L 2 ( Ω × [ 0 , T ] , P ; L 2 ( K , H ) ) . Define
Z t = 0 t S ( t s ) Ψ s d W s .
Then, Z has a mean-square continuous version, and
sup t [ 0 , T ] E Z t 2 M 2 e 2 ω + T E 0 T Ψ s L 2 ( K , H ) 2 d s .
Proof. 
For fixed t, the Hilbert-space Itô isometry gives
E Z t 2 = E 0 t S ( t s ) Ψ s L 2 ( K , H ) 2 d s M 2 e 2 ω + T E 0 T Ψ s L 2 ( K , H ) 2 d s .
Taking the supremum over t gives the estimate.
We prove mean-square continuity. Let t n t . Consider first t n t . Then,
Z t n Z t = 0 t [ S ( t n s ) S ( t s ) ] Ψ s d W s + t t n S ( t n s ) Ψ s d W s .
By the Itô isometry, the mean square of the first term is
E 0 t [ S ( t n s ) S ( t s ) ] Ψ s L 2 ( K , H ) 2 d s .
For each s < t , strong continuity of S and the Hilbert–Schmidt expansion of Ψ s imply pointwise convergence to zero. The integrand is bounded by
4 M 2 e 2 ω + T Ψ s L 2 ( K , H ) 2 .
Dominated convergence gives convergence to zero. The mean square of the second term equals
E t t n S ( t n s ) Ψ s L 2 ( K , H ) 2 d s M 2 e 2 ω + T E t t n Ψ s L 2 ( K , H ) 2 d s ,
which converges to zero by absolute continuity of the integral. The case t n t follows in the same way. Thus, Z has a mean-square continuous version. □
Theorem A1
(Existence and uniqueness of the sentiment SPDE). Let Assumptions 1 and 2 hold. Let
x L 2 ( Ω , F 0 ; H ) , u L 2 ( Ω × [ 0 , T ] , P ; U ) .
Then, there exists a unique mild sentiment field X M T satisfying (3). Moreover, there exists a constant C T , depending only on T , M , ω , L F , L G , C F , C G , B L ( U , H ) , such that
sup t [ 0 , T ] E X t 2 C T 1 + E x 2 + E 0 T u s U 2 d s .
Proof. 
Define, for Y M T ,
( Γ Y ) t = S ( t ) x + 0 t S ( t s ) [ F ( Y s ) + B u s ] d s + 0 t S ( t s ) G ( Y s ) d W s .
Lemmas A1 and A2 show that Γ Y is mean-square continuous and predictable. The growth conditions in Assumption 2 imply
E 0 T F ( Y s ) + B u s 2 d s 2 C F 2 E 0 T ( 1 + Y s ) 2 d s + 2 B 2 E 0 T u s U 2 d s < ,
and
E 0 T G ( Y s ) L 2 ( K , H ) 2 d s C G 2 E 0 T ( 1 + Y s ) 2 d s < .
Thus, Γ maps M T into itself.
We first prove local contraction. Let Y , Z M τ , where 0 < τ T . For t τ , the elementary inequality a + b 2 2 a 2 + 2 b 2 , Lemma A1, the Itô isometry, and the Lipschitz assumptions yield
E ( Γ Y ) t ( Γ Z ) t 2 2 E 0 t S ( t s ) [ F ( Y s ) F ( Z s ) ] d s 2 + 2 E 0 t S ( t s ) [ G ( Y s ) G ( Z s ) ] d W s 2 2 t M 2 e 2 ω + T L F 2 0 t E Y s Z s 2 d s + 2 M 2 e 2 ω + T L G 2 0 t E Y s Z s 2 d s q τ Y Z M τ 2 ,
where
q τ = 2 M 2 e 2 ω + T ( τ 2 L F 2 + τ L G 2 ) .
Choose τ > 0 such that q τ < 1 . Then, Γ is a contraction on M τ . By the Banach fixed point theorem, there is a unique mild solution on [ 0 , τ ] .
The same argument can be restarted at time τ , because the random variable X τ is square-integrable and F τ -measurable whenever X M τ . On [ τ , 2 τ ] , the mild equation is written with initial value X τ and shifted semigroup, and the constants are unchanged. Repeating this finite continuation procedure gives a unique solution on [ 0 , T ] . If two solutions on [ 0 , T ] existed, their restrictions to the first subinterval would coincide, then their restrictions to the second subinterval would coincide, and the same induction would cover the whole interval. Uniqueness follows.
It remains to prove the moment estimate. For t [ 0 , T ] , use a + b + c 2 3 a 2 + 3 b 2 + 3 c 2 . The semigroup bound gives
E S ( t ) x 2 M 2 e 2 ω + T E x 2 .
The deterministic convolution estimate gives
E 0 t S ( t s ) [ F ( X s ) + B u s ] d s 2 t M 2 e 2 ω + T E 0 t F ( X s ) + B u s 2 d s 2 T M 2 e 2 ω + T C F 2 E 0 t ( 1 + X s ) 2 d s + 2 T M 2 e 2 ω + T B 2 E 0 t u s U 2 d s .
Since ( 1 + r ) 2 2 ( 1 + r 2 ) , this is bounded by
C 1 T + E 0 t X s 2 d s + E 0 T u s U 2 d s
for a constant C 1 depending only on the indicated parameters. The stochastic convolution estimate and the growth of G give
E 0 t S ( t s ) G ( X s ) d W s 2 M 2 e 2 ω + T C G 2 E 0 t ( 1 + X s ) 2 d s C 2 T + E 0 t X s 2 d s .
Combining the estimates yields
E X t 2 C 3 1 + E x 2 + E 0 T u s U 2 d s + C 3 0 t E X s 2 d s .
Gronwall’s lemma gives
E X t 2 C 3 e C 3 T 1 + E x 2 + E 0 T u s U 2 d s .
Taking the supremum over t [ 0 , T ] proves (A1). □
Proposition A1
(Stability with respect to initial state and semantic forcing). Let Assumptions 1 and 2 hold. Let X and X ˜ be the mild solutions associated with data ( x , u ) and ( x ˜ , u ˜ ) , with the same coefficients A , F , G , B and the same cylindrical Wiener process. Then,
sup t [ 0 , T ] E X t X ˜ t 2 C T E x x ˜ 2 + E 0 T u s u ˜ s U 2 d s ,
where C T depends only on T , M , ω , L F , L G , B .
Proof. 
Subtracting the two mild equations gives
X t X ˜ t = S ( t ) ( x x ˜ ) + 0 t S ( t s ) [ F ( X s ) F ( X ˜ s ) ] d s + 0 t S ( t s ) B ( u s u ˜ s ) d s + 0 t S ( t s ) [ G ( X s ) G ( X ˜ s ) ] d W s .
Using a + b + c + d 2 4 a i 2 , the semigroup estimate, Cauchy–Schwarz, and the Itô isometry, we obtain
E X t X ˜ t 2 4 M 2 e 2 ω + T E x x ˜ 2 + 4 t M 2 e 2 ω + T L F 2 0 t E X s X ˜ s 2 d s + 4 t M 2 e 2 ω + T B 2 E 0 t u s u ˜ s U 2 d s + 4 M 2 e 2 ω + T L G 2 0 t E X s X ˜ s 2 d s .
Define
f ( t ) = sup r [ 0 , t ] E X r X ˜ r 2 .
Then,
f ( t ) C 1 E x x ˜ 2 + E 0 T u s u ˜ s U 2 d s + C 1 0 t f ( s ) d s .
Gronwall’s lemma gives (A2). □
Proposition A2
(Fourth-moment estimates and fourth-moment input stability). Let Assumptions 1 and 2 hold. If x L 4 ( Ω , F 0 ; H ) , and u L 2 ( 0 , T ; U ) is deterministic, then the mild solution belongs to M T 4 and
sup t [ 0 , T ] E X t u 4 C T 1 + E x 4 + u L 2 ( 0 , T ; U ) 4 .
If u , v L 2 ( 0 , T ; U ) are deterministic and the corresponding solutions have the same initial state and are driven by the same Wiener process, then,
sup t [ 0 , T ] E X t u X t v 4 C T u v L 2 ( 0 , T ; U ) 4 .
The constant depends only on the structural constants and T.
Proof. 
Write c S = M e ω + T . We first record the fourth-moment estimates and their continuity consequences. If a is predictable and belongs to L 4 ( Ω × [ 0 , T ] ; H ) , Hölder’s inequality gives
E 0 t S ( t s ) a s d s 4 c S 4 t 3 0 t E a s 4 d s .
Let t j t , and suppose first that t j t . Splitting the difference at t and applying the same inequality gives
E 0 t [ S ( t j s ) S ( t s ) ] a s d s 4 t 3 0 t E [ S ( t j s ) S ( t s ) ] a s 4 d s 0
by strong continuity and dominated convergence. The integral over [ t , t j ] is bounded by
c S 4 ( t j t ) 3 t t j E a s 4 d s 0 .
The case t j t is analogous. Thus, deterministic convolution maps such integrands into C ( [ 0 , T ] ; L 4 ( Ω ; H ) ) .
For a predictable Hilbert–Schmidt integrand Ψ L 4 ( Ω × [ 0 , T ] ; L 2 ( K , H ) ) , the Hilbert-space Burkholder–Davis–Gundy inequality and Cauchy–Schwarz in time yield
E 0 t S ( t s ) Ψ s d W s 4 C BDG E 0 t S ( t s ) Ψ s L 2 2 d s 2 C BDG c S 4 t 0 t E Ψ s L 2 4 d s .
For the stochastic convolution difference over [ 0 , t ] , the same two inequalities give
E 0 t [ S ( t j s ) S ( t s ) ] Ψ s d W s 4 C BDG t 0 t E [ S ( t j s ) S ( t s ) ] Ψ s L 2 4 d s 0 .
Here, strong continuity is lifted from H to Hilbert–Schmidt operators by finite-rank approximation, and the integrand is dominated by 16 c S 4 E Ψ s L 2 4 . The contribution over [ t , t j ] is at most
C BDG c S 4 ( t j t ) t t j E Ψ s L 2 4 d s ,
which tends to zero. Hence, the stochastic convolution also belongs to C ( [ 0 , T ] ; L 4 ( Ω ; H ) ) .
We now construct the solution in the fourth moment. On [ 0 , τ ] , let Γ be the Picard map from the proof of Theorem A1, now acting on M τ 4 . The initial orbit t S ( t ) x belongs to C ( [ 0 , T ] ; L 4 ( Ω ; H ) ) by strong continuity, the uniform semigroup bound on [ 0 , T ] , and dominated convergence. The deterministic forcing convolution is a continuous deterministic H-valued function because u L 2 ( 0 , T ; U ) , and
sup t T 0 t S ( t s ) B u s d s 4 c S 4 B 4 T 2 u L 2 ( 0 , T ; U ) 4 .
The linear-growth assumptions and the two continuity results above show that Γ maps M τ 4 into itself. For Y , Y ˜ M τ 4 , the forcing terms cancel and
Γ Y Γ Y ˜ M τ 4 4 C c S 4 τ 4 L F 4 + τ 2 L G 4 Y Y ˜ M τ 4 4 .
Choose τ so that the coefficient on the right is strictly less than one. Banach’s fixed point theorem gives a unique fixed point in M τ 4 . Repeating the construction on finitely many adjacent intervals yields a process in M T 4 . Since M T 4 embeds continuously into M T 2 , uniqueness in Theorem A1 identifies this fixed point with the mild solution already constructed there.
The linear-growth assumptions imply
F ( z ) 4 + G ( z ) L 2 4 C ( 1 + z 4 ) .
Applying z 1 + z 2 + z 3 + z 4 4 4 3 j = 1 4 z j 4 to the mild equation, using (A5), (A6), and the deterministic forcing estimate, gives
E X t u 4 C T 1 + E x 4 + u L 2 ( 0 , T ; U ) 4 + C T 0 t E X s u 4 d s .
Gronwall’s lemma proves (A3).
For the difference D = X u X v , the initial term vanishes and
D t = 0 t S ( t s ) [ F ( X s u ) F ( X s v ) ] d s + 0 t S ( t s ) B ( u s v s ) d s + 0 t S ( t s ) [ G ( X s u ) G ( X s v ) ] d W s .
The Lipschitz assumptions, the fourth-moment convolution inequalities, and
sup t T 0 t S ( t s ) B ( u s v s ) d s 4 c S 4 B 4 T 2 u v L 2 ( 0 , T ; U ) 4
give
E D t 4 C T u v L 2 ( 0 , T ; U ) 4 + C T 0 t E D s 4 d s .
A second application of Gronwall’s lemma proves (A4). □
Theorem A2
(Consistency of Galerkin approximations). Let x L 2 ( Ω , F 0 ; H ) , and let u L 2 ( Ω × [ 0 , T ] , P ; U ) . Let H n H be finite-dimensional subspaces with orthogonal projections P n converging strongly to the identity. Let S n be strongly continuous semigroups on H n , extended to H through S n ( t ) P n , and assume
sup n 1 sup 0 t T S n ( t ) P n M T , sup 0 t T S n ( t ) P n z S ( t ) z 0 for every z H .
Let F n : H n H n and G n : H n L 2 ( K , H n ) satisfy Lipschitz and growth bounds that are uniform in n, and let B n L ( U , H n ) satisfy sup n B n < . Let X solve (2) with data ( x , u ) , and assume
E 0 T F n ( P n X s ) P n F ( X s ) 2 d s 0 , E 0 T G n ( P n X s ) P n G ( X s ) L 2 ( K , H ) 2 d s 0 , E 0 T B n u s P n B u s 2 d s 0 .
If X ( n ) is the mild solution of the projected equation with initial state P n x , then,
sup t [ 0 , T ] E X t ( n ) X t 2 0 .
Proof. 
The strong convergence in (A7) is uniform in time on compact intervals for each fixed vector. Because t X t is continuous in L 2 ( Ω ; H ) and [ 0 , T ] is compact, the set { X t : t [ 0 , T ] } is compact in L 2 ( Ω ; H ) . Strong convergence of the uniformly bounded projections therefore implies
sup t [ 0 , T ] E P n X t X t 2 0 .
Write the projected mild equation in H, and subtract the mild equation for X. Add and subtract F n ( P n X s ) and G n ( P n X s ) . The difference is the sum of the initial semigroup error, two Lipschitz terms involving X ( n ) P n X , the coefficient-consistency terms in (A8), the input-consistency term, and the deterministic and stochastic semigroup-consistency convolutions applied to F ( X ) + B u and G ( X ) . The deterministic convolution estimate and the Itô isometry give
sup r t E X r ( n ) X r 2 ε n + C T 0 t sup q s E X q ( n ) X q 2 d s ,
where ε n is the sum of all consistency errors.
We verify that ε n 0 . The initial term converges by (A7) and approximation of x by simple H-valued random variables. The three coefficient terms converge by (A8). For the remaining deterministic convolution, approximate the square-integrable random integrand F ( X ) + B u in L 2 ( Ω × [ 0 , T ] ; H ) by a finite-valued simple function. Uniform boundedness of the semigroups and their strong convergence make the error for the simple function tend to zero uniformly in terminal time. In contrast, the approximation error is controlled uniformly in n. The stochastic convolution is treated identically after applying the Itô isometry and approximating G ( X ) in L 2 ( Ω × [ 0 , T ] ; L 2 ( K , H ) ) by finite-rank simple integrands. Finally, (A10) absorbs the occurrences of X P n X in the Lipschitz terms. Thus ε n 0 , and Gronwall’s lemma proves (A9). □
Lemma A3
(Trace identity for covariance). Let Y L 2 ( Ω ; H ) . Then C Y is self-adjoint, positive, and trace class. Moreover,
Tr C Y = E Y E Y 2 .
Proof. 
Let m = E Y . The defining Bochner expectation is well defined because, for every h H ,
C Y h E | Y m , h | Y m E Y m 2 h .
Thus, C Y L ( H ) . For h , g H ,
C Y h , g = E [ Y m , h Y m , g ] .
This identity shows self-adjointness because the right-hand side is symmetric in h and g. It also shows positivity because
C Y h , h = E | Y m , h | 2 0 .
Let ( e k ) k 1 be an orthonormal basis of H. Then,
k = 1 n C Y e k , e k = E k = 1 n | Y m , e k | 2 E Y m 2 < .
By monotone convergence and Parseval’s identity,
k = 1 C Y e k , e k = E k = 1 | Y m , e k | 2 = E Y m 2 .
For a positive self-adjoint operator, the finiteness of this sum for one orthonormal basis is equivalent to the trace-class property, and the sum equals the trace. Thus, (A11) holds. □
Lemma A4
(Distance to constants and standard deviation). Let Y L 2 ( Ω ; R ) . Then,
( Var Y ) 1 / 2 = inf a R Y a L 2 ( Ω ) .
Consequently, for Y , Z L 2 ( Ω ; R ) ,
| ( Var Y ) 1 / 2 ( Var Z ) 1 / 2 |     Y Z L 2 ( Ω ) .
Proof. 
For each a R ,
E | Y a | 2 = E | Y E Y + E Y a | 2 = E | Y E Y | 2 + | E Y a | 2 + 2 ( E Y a ) E ( Y E Y ) = Var Y + | E Y a | 2 .
The infimum over a is attained at a = E Y and equals Var Y . Taking square roots gives the first claim. For the second claim, let C be the closed subspace of constants in L 2 ( Ω ) . The map Y dist ( Y , C ) is one-Lipschitz in any Hilbert space. Hence,
| dist ( Y , C ) dist ( Z , C ) | Y Z L 2 ( Ω ) .
Using the first claim completes the proof. □
Lemma A5
(Trace norm ideal inequality). Let R , L L ( H ) , and let T be trace class on H. Then, R T L is trace class and
R T L 1 R T 1 L .
Proof. 
Let ( s n ( T ) ) n 1 be the singular values of T. The trace norm is T 1 = n s n ( T ) . The singular value inequality
s n ( R T L ) R L s n ( T ) ,
and bounded left and right multiplication follows from the min-max characterisation of singular values. Summing over n gives
R T L 1 = n s n ( R T L ) R L n s n ( T ) = R L T 1 .
Thus, R T L is trace class, and the inequality holds. □
Lemma A6
(Yosida preservation of strict dissipativity). Assume that A generates a strongly continuous semigroup and satisfies
A z , z α z 2 , z Dom ( A ) ,
for some α > 0 . For n > 0 , define
A n = n A ( n I A ) 1 .
Then, A n L ( H ) , and
A n z , z α n z 2 , α n = n α n + α .
In particular, α n α .
Proof. 
Set C = A + α I . The assumed inequality says that C is dissipative, and C is the generator of the shifted semigroup e α t S ( t ) . Because C is a generator, Range ( λ I C ) = H for all sufficiently large λ > 0 . Together with dissipativity, this range condition and the Lumer–Phillips theorem show that C is maximally dissipative and that its given semigroup is contractive. In particular,
λ ( λ I C ) 1 1 , λ > 0 .
The shifted resolvent identity gives
( n I A ) 1 = ( ( n + α ) I C ) 1 .
Consequently, the operator
J n = ( n + α ) ( ( n + α ) I C ) 1
is a contraction. Since
A n = n n ( n I A ) 1 I = n n n + α J n I ,
we obtain
A n z , z = n n n + α J n z , z z 2 n n n + α 1 z 2 = n α n + α z 2 .
This is (A12). □
Proposition A3
(Yosida justification of the contraction estimate). Under the hypotheses of Theorem 1, the Itô calculation carried out there for strong solutions is valid for general mild solutions through Yosida approximation.
Proof. 
Let S n ( t ) = e t A n . The standard Yosida theory gives
S n ( t ) z S ( t ) z
for every z H , uniformly for t in compact intervals, and the semigroups have a common finite-horizon bound. Consider
d X t n = [ A n X t n + F ( X t n ) + B u t ] d t + G ( X t n ) d W t , X 0 n = x ,
and the corresponding equation for Y n with initial state y. Since A n is bounded, both solutions are Hilbert-space semimartingales and the Itô formula applies directly. Apply it to e ( 2 α n η ) t X t n Y t n 2 and stop the local martingale by the same quadratic-variation stopping times used in the proof of Theorem 1. Lemma A6, Assumption 3, and Fatou’s lemma give
E X t n Y t n 2 e ( 2 α n η ) t E x y 2 .
No fourth-moment assumption on the initial states is required.
It remains to justify convergence. Subtracting the mild equations for X n and X, and adding and subtracting terms with the same semigroup, gives Lipschitz terms controlled by
C T 0 t sup q s E X q n X q 2 d s
and consistency terms involving
S n ( t ) x S ( t ) x , 0 t [ S n ( t s ) S ( t s ) ] [ F ( X s ) + B u s ] d s , 0 t [ S n ( t s ) S ( t s ) ] G ( X s ) d W s .
The first converges in L 2 ( Ω ; H ) by approximation of x by simple random variables. The second and third converge uniformly in terminal time by the deterministic convolution estimate, the Itô isometry, strong semigroup convergence, uniform boundedness, and approximation of the integrands by finite-valued simple processes. Gronwall’s lemma yields
X n X M T 0 .
The same proof applies to Y n . Therefore, for every fixed t,
E X t n Y t n 2 E X t Y t 2 .
Since α n α , passage to the limit in the approximating inequality proves
E X t Y t 2 e ( 2 α η ) t E x y 2 ,
which is the estimate used in Theorem 1. □

Appendix B. Split-Conformal Coverage for the Projected Readout

The coverage statement used in Section 4 is recorded here to distinguish a finite-sample calibration guarantee from the model-based Gaussian covariance calculation. Let a predictor be fitted on data independent of the calibration and test observations. For a covariate-response pair Z = ( V , Y ) , let μ ^ ( V ) be a point prediction, and let σ ^ ( V ) > 0 be a scale estimate. Define the nonconformity score
R ( Z ) = | Y μ ^ ( V ) | σ ^ ( V ) .
Proposition A4
(Finite-sample split-conformal coverage). Let Z 1 , , Z m , Z m + 1 be exchangeable conditional on the fitted predictor. Fix α ( 0 , 1 ) , set
k = ( m + 1 ) ( 1 α ) ,
and assume k m . Let R ( k ) be the k-th order statistic of the calibration scores R ( Z 1 ) , , R ( Z m ) . Then, the interval
C α ( V m + 1 ) = μ ^ ( V m + 1 ) R ( k ) σ ^ ( V m + 1 ) , μ ^ ( V m + 1 ) + R ( k ) σ ^ ( V m + 1 )
satisfies
P Y m + 1 C α ( V m + 1 ) 1 α .
No distributional assumption on the score is required.
Proof. 
Condition the training sample used to construct μ ^ and σ ^ . The scores R 1 , , R m + 1 are then exchangeable. Introduce an independent continuous random tie breaker, and rank the pairs formed by each score and its tie breaker in lexicographic order. The rank of the test pair is uniform on { 1 , , m + 1 } . Whenever that rank does not exceed k, the test score is no larger than the k-th calibration score after the deterministic tie convention is restored. Ties can only increase the probability of this event. Hence,
P { R m + 1 R ( k ) } k m + 1 1 α .
The event on the left is exactly the coverage event in (A13). Removing the conditioning proves (A14). □
In the synthetic study, one terminal one-step score is taken from each independently generated trajectory. Conditional on the training sample, the calibration scores and each test score are exchangeable, implying that the proposition applies to the reported terminal coverage. It does not imply validity under distribution shift or after adaptively reusing the test trajectories for model selection.

References

  1. Pang, B.; Lee, L. Opinion mining and sentiment analysis. Found. Trends Inf. Retr. 2008, 2, 1–135. [Google Scholar] [CrossRef] [Scilit]
  2. Liu, B. Sentiment Analysis and Opinion Mining; Morgan and Claypool Publishers: San Rafael, CA, USA, 2012. [Google Scholar] [CrossRef] [Scilit]
  3. Devlin, J.; Chang, M.-W.; Lee, K.; Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT 2019; Association for Computational Linguistics: Minneapolis, MN, USA, 2019; pp. 4171–4186. [Google Scholar] [CrossRef] [Scilit]
  4. Zhang, W.; Deng, Y.; Liu, B.; Pan, S.; Bing, L. Sentiment analysis in the era of large language models: A reality check. In Findings of NAACL 2024; Association for Computational Linguistics: Mexico City, Mexico, 2024; pp. 3881–3906. [Google Scholar] [CrossRef] [Scilit]
  5. Hou, G.; Shen, Y.; Lu, W. Progressive tuning: Towards generic sentiment abilities for large language models. In Findings of ACL 2024; Association for Computational Linguistics: Bangkok, Thailand, 2024; pp. 14392–14402. [Google Scholar] [CrossRef] [Scilit]
  6. Lyu, Z.; Jin, Z.; Gonzalez Adauto, F.; Mihalcea, R.; Schölkopf, B.; Sachan, M. Do LLMs think fast and slow? A causal study on sentiment analysis. In Findings of EMNLP 2024; Association for Computational Linguistics: Miami, FL, USA, 2024; pp. 9353–9372. [Google Scholar] [CrossRef] [Scilit]
  7. Madsen, A.; Chandar, S.; Reddy, S. Are self-explanations from large language models faithful? In Findings of ACL 2024; Association for Computational Linguistics: Bangkok, Thailand, 2024; pp. 295–337. [Google Scholar] [CrossRef] [Scilit]
  8. Gal, Y.; Ghahramani, Z. Dropout as a Bayesian approximation: Representing model uncertainty in deep learning. In Proceedings of the 33rd International Conference on Machine Learning; PMLR: New York, NY, USA, 2016; Volume 48, pp. 1050–1059. [Google Scholar]
  9. Kendall, A.; Gal, Y. What uncertainties do we need in Bayesian deep learning for computer vision? In Advances in Neural Information Processing Systems; Curran Associates: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
  10. Guo, C.; Pleiss, G.; Sun, Y.; Weinberger, K.Q. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning; PMLR: Sydney, Australia, 2017; Volume 70, pp. 1321–1330. [Google Scholar]
  11. Angelopoulos, A.N.; Bates, S. Conformal prediction: A gentle introduction. Found. Trends Mach. Learn. 2023, 16, 494–591. [Google Scholar] [CrossRef] [Scilit]
  12. Campos, M.; Farinhas, A.; Zerva, C.; Figueiredo, M.A.T.; Martins, A.F.T. Conformal prediction for natural language processing: A survey. Trans. Assoc. Comput. Linguist. 2024, 12, 1497–1516. [Google Scholar] [CrossRef] [Scilit]
  13. Li, C.; Zhou, H.; Glavaš, G.; Korhonen, A.; Vulić, I. Large language models are miscalibrated in-context learners. In Findings of ACL 2025; Association for Computational Linguistics: Vienna, Austria, 2025; pp. 11575–11596. [Google Scholar] [CrossRef] [Scilit]
  14. Ribeiro, M.T.; Singh, S.; Guestrin, C. Why should I trust you? Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; ACM: San Francisco, CA, USA, 2016; pp. 1135–1144. [Google Scholar] [CrossRef] [Scilit]
  15. Lundberg, S.M.; Lee, S.-I. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems; Curran Associates: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
  16. Miyato, T.; Dai, A.M.; Goodfellow, I. Adversarial training methods for semi-supervised text classification. In Proceedings of the 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, 24–26 April 2017. [Google Scholar]
  17. Belinkov, Y.; Bisk, Y. Synthetic and natural noise both break neural machine translation. In Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada, 30 April–3 May 2018. [Google Scholar]
  18. Jia, R.; Liang, P. Adversarial examples for evaluating reading comprehension systems. In Proceedings of EMNLP 2017; Association for Computational Linguistics: Copenhagen, Denmark, 2017; pp. 2021–2031. [Google Scholar] [CrossRef] [Scilit]
  19. Da Prato, G.; Zabczyk, J. Stochastic Equations in Infinite Dimensions, 2nd ed.; Cambridge University Press: Cambridge, UK, 2014. [Google Scholar] [CrossRef] [Scilit]
  20. Prévôt, C.; Röckner, M. A Concise Course on Stochastic Partial Differential Equations; Lecture Notes in Mathematics; Springer: Berlin, Germany, 2007; Volume 1905. [Google Scholar] [CrossRef] [Scilit]
  21. Liu, W.; Röckner, M. Stochastic Partial Differential Equations: An Introduction; Springer: Cham, Switzerland, 2015. [Google Scholar] [CrossRef] [Scilit]
  22. Marinelli, C.; Di Persio, L.; Ziglio, G. Approximation and convergence of solutions to semilinear stochastic evolution equations with jumps. J. Funct. Anal. 2013, 264, 2784–2816. [Google Scholar] [CrossRef] [Scilit]
  23. Cordoni, F.; Di Persio, L. Stochastic reaction-diffusion equations on networks with dynamic time-delayed boundary conditions. J. Math. Anal. Appl. 2017, 451, 583–603. [Google Scholar] [CrossRef] [Scilit]
  24. Salvi, C.; Lemercier, M.; Gerasimovics, A. Neural stochastic PDEs: Resolution-invariant learning of continuous spatiotemporal dynamics. In Advances in Neural Information Processing Systems; Curran Associates: Red Hook, NY, USA, 2022; Volume 35, pp. 1333–1344. [Google Scholar] [CrossRef] [Scilit]
  25. Gong, S.; Hu, P.; Meng, Q.; Wang, Y.; Zhu, R.; Chen, B.; Ma, Z.; Ni, H.; Liu, T.-Y. Deep latent regularity network for modeling stochastic partial differential equations. Proc. AAAI Conf. Artif. Intell. 2023, 37, 7740–7747. [Google Scholar] [CrossRef] [Scilit]
  26. Jung, J.; Shin, H.; Choi, M. Bayesian deep learning framework for uncertainty quantification in stochastic partial differential equations. SIAM J. Sci. Comput. 2024, 46, C57–C76. [Google Scholar] [CrossRef] [Scilit]
  27. Tian, C.; Chen, Y. Image segmentation and denoising algorithm based on partial differential equations. IEEE Sens. J. 2020, 20, 11935–11942. [Google Scholar] [CrossRef] [Scilit]
  28. Qiu, W.-X.; Geng, K.-L.; Zhu, B.-W.; Liu, W.; Li, J.-T.; Dai, C.-Q. Data-driven forward-inverse problems of the 2-coupled mixed derivative nonlinear Schrödinger equation using deep learning. Nonlinear Dyn. 2024, 112, 10215–10228. [Google Scholar] [CrossRef] [Scilit]
  29. Yin, F.; Xu, Z.; Fu, Y. Novel high-order explicit energy-preserving schemes for NLS-type equations based on the Lie-group method. Math. Comput. Simul. 2024, 225, 570–585. [Google Scholar] [CrossRef] [Scilit]
  30. Wei, H.-Y.; Fan, E.-G.; Guo, H.-D. Riemann-Hilbert approach and nonlinear dynamics of the coupled higher-order nonlinear Schrödinger equation in the birefringent or two-mode fiber. Nonlinear Dyn. 2021, 104, 649–660. [Google Scholar] [CrossRef] [Scilit]
  31. Zhao, Z.-L.; Zhang, R.-F. Artificial intelligence-enhanced mathematical derivation method: Exact solutions of the Benjamin-Bona-Mahony equation. Eng. Appl. Artif. Intell. 2026, 173, 114464. [Google Scholar] [CrossRef] [Scilit]
  32. Go, A.; Bhayani, R.; Huang, L. Twitter Sentiment Classification Using Distant Supervision; CS224N Project Report; Stanford University: Stanford, CA, USA, 2009. [Google Scholar]
  33. Rosenthal, S.; Farra, N.; Nakov, P. SemEval-2017 Task 4: Sentiment analysis in Twitter. In Proceedings of the 11th International Workshop on Semantic Evaluation; Association for Computational Linguistics: Vancouver, BC, Canada, 2017; pp. 502–518. [Google Scholar] [CrossRef] [Scilit]
  34. Barbieri, F.; Camacho-Collados, J.; Neves, L.; Espinosa-Anke, L. TweetEval: Unified benchmark and comparative evaluation for tweet classification. In Findings of EMNLP 2020; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 1644–1650. [Google Scholar] [CrossRef] [Scilit]
  35. Maas, A.L.; Daly, R.E.; Pham, P.T.; Huang, D.; Ng, A.Y.; Potts, C. Learning word vectors for sentiment analysis. In Proceedings of ACL-HLT 2011; Association for Computational Linguistics: Portland, OR, USA, 2011; pp. 142–150. Available online: https://aclanthology.org/P11-1015 (accessed on 16 August 2026).
  36. Mohammad, S.M.; Turney, P.D. Crowdsourcing a word-emotion association lexicon. Comput. Intell. 2013, 29, 436–465. [Google Scholar] [CrossRef] [Scilit]
  37. Bogachev, V.I. Gaussian Measures; American Mathematical Society: Providence, RI, USA, 1998. [Google Scholar] [CrossRef] [Scilit]
  38. Villani, C. Optimal Transport: Old and New; Springer: Berlin, Germany, 2009. [Google Scholar] [CrossRef] [Scilit]
  39. Da Prato, G. An Introduction to Infinite-Dimensional Analysis; Springer: Berlin, Germany, 2006. [Google Scholar] [CrossRef] [Scilit]
  40. Pazy, A. Semigroups of Linear Operators and Applications to Partial Differential Equations; Springer: New York, NY, USA, 1983. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Conceptual pipeline from time-indexed text to a projected stochastic semantic field and its uncertainty, influence, and robustness diagnostics.
Figure 1. Conceptual pipeline from time-indexed text to a projected stochastic semantic field and its uncertainty, influence, and robustness diagnostics.
Computation 14 00216 g001
Figure 2. Observed synthetic readout, one-step conditional mean, and split-conformal interval at nominal level 0.90 .
Figure 2. Observed synthetic readout, one-step conditional mean, and split-conformal interval at nominal level 0.90 .
Computation 14 00216 g002
Figure 3. Adjoint influence kernel for the two forcing coordinates in the three-mode model.
Figure 3. Adjoint influence kernel for the two forcing coordinates in the three-mode model.
Computation 14 00216 g003
Figure 4. Modal uncertainty shares and identity lift influence-energy shares.
Figure 4. Modal uncertainty shares and identity lift influence-energy shares.
Computation 14 00216 g004
Figure 5. Sensitivity of predictive dispersion and influence to a common damping multiplier.
Figure 5. Sensitivity of predictive dispersion and influence to a common damping multiplier.
Computation 14 00216 g005
Table 1. One-step prediction performance on four hundred synthetic test trajectories.
Table 1. One-step prediction performance on four hundred synthetic test trajectories.
MethodReadout RMSEReadout MAELatent RMSE
Structured dissipative model0.0951820.0760100.065492
Unrestricted VARX0.0951720.0760040.065534
Persistence0.1048030.0837640.075702
Table 2. Dissipativity ablation at horizon T = 12 .
Table 2. Dissipativity ablation at horizon T = 12 .
Damping MultiplierPredictive Standard DeviationInfluence NormSemigroup Factor
0.50.2306231.0621811.487342
1.00.1706690.7522951.054082
2.00.1304380.5320250.745356
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Persio, L.D. Stochastic Semantic Fields for Sentiment-Driven Models. Computation 2026, 14, 216. https://doi.org/10.3390/computation14090216

AMA Style

Persio LD. Stochastic Semantic Fields for Sentiment-Driven Models. Computation. 2026; 14(9):216. https://doi.org/10.3390/computation14090216

Chicago/Turabian Style

Persio, Luca Di. 2026. "Stochastic Semantic Fields for Sentiment-Driven Models" Computation 14, no. 9: 216. https://doi.org/10.3390/computation14090216

APA Style

Persio, L. D. (2026). Stochastic Semantic Fields for Sentiment-Driven Models. Computation, 14(9), 216. https://doi.org/10.3390/computation14090216

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop