Next Article in Journal
A Novel Hybrid Adaptive Multi-Resolution Feature Extraction Method for Power Quality Disturbance Detection
Next Article in Special Issue
The Geometry of Quantum Walks on Graphs—Theory and Applications
Previous Article in Journal
Bio-Inspired Constant-Time Arithmetic Kernels in Hybrid Membrane–Neural Spiking P Systems
Previous Article in Special Issue
Information-Geometric Models in Data Analysis and Physics
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Information-Geometric Models in Data Analysis and Physics II

1
Department of Artificial Intelligence, National University of Distance Education (UNED), 28040 Madrid, Spain
2
Department of Genetics, Microbiology and Statistics, Faculty of Biology, Universitat de Barcelona, 08028 Barcelona, Spain
*
Author to whom correspondence should be addressed.
Mathematics 2026, 14(5), 785; https://doi.org/10.3390/math14050785
Submission received: 27 January 2026 / Revised: 19 February 2026 / Accepted: 21 February 2026 / Published: 26 February 2026

Abstract

This paper continues the development of information-geometric models for data analysis and physics by focusing on their formulation and interpretation through variational principles. Building on the geometric framework introduced previously, we investigate how fundamental variational structures—such as information-theoretic functionals—naturally encode the laws of nature. In the first manuscript, we showed that a wide class of physical problems can be expressed as constrained variational problems on spaces of probability distributions, leading to geodesic flows, gradient dynamics, and generalized Hamiltonian formulations on statistical manifolds. In this second part, we extend the variational formalism by utilizing an extended metric, clarifying the geometric origin of the dynamical equations commonly used in modern physics and providing a coherent interpretation of physical laws in terms of information optimization. By emphasizing variational foundations, this paper strengthens the conceptual and mathematical links between information geometry, data analysis, and physics, and it provides a flexible framework for extending geometric methods to complex, high-dimensional, and dynamical systems.

1. Introduction

In recent years, the interaction between data and geometry has surged as a transformative frontier in both information science and physics. The concept of information-induced geometries suggests that the structure and behavior of physical systems can be understood and modeled at a fundamental level through geometric frameworks informed by data-driven insights. This innovative approach not only opens new pathways for analyzing complex datasets but also has significant implications for our understanding of the universe’s very fabric. By integrating principles from information theory, differential geometry, and quantum physics, we aim to develop methods that revolutionize data analysis and deepen our understanding of physical phenomena, paving the way for breakthroughs across theoretical insights and practical applications.
Building upon this conceptual foundation, we introduced a robust mathematical framework that leverages advanced geometric and statistical techniques to analyze data residing on Riemannian manifolds [1]. Central to our approach are methods that respect the intrinsic curvature and nonlinear structure of manifold data, such as flow lines of gradient vector fields, exponential and logarithmic maps, and kernel-based principal component analysis (PCA). These tools facilitate faithful, low-dimensional representations and the insightful visualization of complex data structures, capturing both the local and global geometric relationships that are critical for understanding the underlying phenomena. By integrating these geometric insights with an information-theoretic perspective, our framework aims to uncover the mechanisms through which information manifests in physical systems, advancing the notion that informational principles fundamentally shape the universe’s fabric. This interdisciplinary approach offers promising avenues for both theoretical explorations and practical applications in data science, physics, and beyond.
The idea that information is a fundamental building block of the universe is gaining significant traction in theoretical physics and mathematics [2,3,4]. This enlightening perspective suggests that beyond the realms of matter and energy, information intricately weaves the very fabric of reality, shaping the structure and dynamics of everything from subatomic particles to vast cosmic phenomena. Recognizing information as a core component unlocks deep insights into the nature of physical laws, the origins of the universe, and the intricate interconnectedness of all existence. This paradigm shift not only enhances our understanding but also paves the way for groundbreaking explorations into the most profound enigmas of the universe (see Figure 1).
This paper is poised to advance a robust mathematical framework and pioneering conceptual tools that delve into the profound relationship between information and the physical universe. By formalizing the principles governing information processing and its manifestation in physical systems, we aim to uncover the intricate mechanisms by which information may emerge from or be embedded in the fundamental laws of physics [5,6]. This approach not only establishes strong connections between abstract informational concepts and concrete physical models but also provides a solid foundation for investigating how information-theoretic principles could transform our understanding and inspire the development of new theories of reality [7,8].
By representing parametric models through complex-valued functions that encapsulate both magnitude and phase, we construct a Riemannian manifold structure on the parameter space informed by divergence functions and complex model functions. This approach generalizes classical information geometry by incorporating phase information relevant to physical phenomena, such as quantum states and wave propagation. We explore how these geometric structures facilitate the analysis of model sensitivities, estimation procedures, and information encoding with particular emphasis on invariance properties and their implications for physical and statistical inference [9].
In particular, we further develop variational principles rooted in the geometric formalism with novel constraints, yielding differential equations reminiscent of quantum wave equations that govern optimal plausibility functions representing observer beliefs or physical states. Using an illustrative example with Gaussian-modulated wave functions, we demonstrate the applicability of this framework to models with deep physical significance.
Our results establish a foundational bridge between information theory, differential geometry, and physics, offering new tools for modeling, analyzing, and interpreting complex systems. This interdisciplinary integration opens avenues for exploring the informational fabric of the universe, providing insights into the emergence of physical laws from informational principles and fostering the development of novel theoretical and experimental paradigms.

2. Materials and Methods

As in [1], we assume that all relevant aspects of the observation process—whether determined by the object, the subject, the environment, or any combination thereof—can be adequately described by elements of the parameter space Θ through a suitable family of complex-valued maps, which will be introduced below. To motivate this framework, we first consider a measure space ( χ , A , ν ) , where χ is an arbitrary set (the sample space), A is a σ –algebra of subsets of χ , and ν is a positive σ –finite measure on the measurable space ( χ , A ) . Let Θ be a manifold and let p : χ × Θ R be a measurable function such that for each θ Θ , P θ ( d x ) = p ( x , θ ) ν ( d x ) defines a probability measure on ( χ , A ) . This constitutes the underlying statistical model. Each point of the manifold Θ represents, potentially, a distinct information source from which sample data are obtained under independence assumptions. Let x χ n , say x = ( x 1 , , x n ) , denote a sample with values in χ . For fixed θ , the joint density with respect to the product measure ν n is given by p ˜ ( x , θ ) = i = 1 n p ( x i , θ ) . The sample is assumed to provide a statistically consistent estimator T ( x ) = T ( x 1 , , x n ) of the true parameter θ Θ characterizing the information source under study.
Fix θ Θ and an estimator T. The mapping T induces a probability measure on ( Θ , B ( Θ ) ) defined by P θ T ( B ) = P θ T 1 ( B ) , for any Borel set B B ( Θ ) , where B ( Θ ) is the Borel σ –algebra generated by the open sets of the manifold Θ . Under appropriate regularity conditions, this induced measure admits a density with respect to a conveniently defined measure on the manifold Θ , d μ , namely f ( ω , θ ) = d P θ T d μ ( ω ) , which is the Radon–Nikodym derivative of P θ T with respect to d μ . In what follows, we focus on this probability density on Θ , or more precisely, on a not necessarily real square root of it, in contrast with the standard real-valued framework of information geometry. To investigate the behavior and structural properties of information sources, we introduce the following notion.
A regular parametric family of functions  F Θ is defined as a family of measurable maps h : Θ × Θ C , which admit a complex exponential representation h ( ω , θ ) = r ( ω , θ ) , e i ζ ( ω , θ ) , where r ( ω , θ ) 0 and ζ ( ω , θ ) are real-valued functions representing, respectively, the modulus and the argument of h ( ω , θ ) . This construction generalizes the classical square-root embedding of statistical models by allowing h ( ω , θ ) = r ( ω , θ ) e i ζ ( ω , θ ) , where f ( ω , θ ) is a probability density. The statistical model is completely determined by | h | 2 , whereas the phase function ζ provides an additional geometric degree of freedom. By multiplying the density by a scalar field on Θ prior to taking the square root, one may assume without loss of generality that the phase ζ is independent of the choice of reference measure. For each fixed θ Θ , the squared modulus r 2 ( ω , θ ) = h ( ω , θ ) h ( ω , θ ) ¯ is required to be a probability density with respect to d μ ; that is, P θ T ( d ω ) = r 2 ( ω , θ ) μ ( d ω ) defines a probability measure on ( Θ , B ( Θ ) ) for all θ Θ .
Although the probability measure P θ T is independent of the particular choice of reference measure, the function r ( · , θ ) depends on d μ . Indeed, if μ ^ is another reference measure such that d μ / d μ ^ = q for some positive measurable function q , then by the Radon–Nikodym theorem, P θ T ( d ω ) = r 2 ( ω , θ ) q ( ω ) μ ( d ω ) = r ^ 2 ( ω , θ ) μ ^ ( d ω ) , where r ^ ( ω , θ ) = r ( ω , θ ) , q ( ω ) . This expression makes explicit the dependence of the modulus r on the chosen reference measure.
Throughout, we implicitly assume the regularity conditions required for subsequent developments, including sufficient smoothness with respect to θ and the assumption that the closure of the set on which r ( · , θ ) is strictly positive does not depend on θ . We shall refer to μ as the reference measure, to h as the model function, and to f ( ω , θ ) = r 2 ( ω , θ ) as the model density.

2.1. The Riemannian Structure on the Parameter Space

The exploration of Riemannian structures on parameter spaces provides a profound geometric perspective on models and their underlying function spaces. By endowing these spaces with a Riemannian metric, we can investigate the intrinsic geometric properties that influence model behavior, estimation, and inference. Such a perspective enables a more nuanced understanding of the parameter space’s complexity and curvature, offering insights that extend beyond traditional Euclidean approaches.
Letting L 2 ( Θ , μ ) be the Hilbert space of these complex-valued measurable functions q in Θ such that the Lebesgue integral Θ | q ( ω ) | 2 μ ( d ω ) < , where | z | = z z ¯ , z C , if we denote by D 2 the square of the Hilbert distance
D 2 ( p , q ) = Θ | p ( ω ) q ( ω ) | 2 μ ( d ω ) ,
once given the model function h ( · , θ ) of a parametric family, F Θ , since these functions are naturally identifiable with elements of L 2 ( Θ , μ ) , we are able to build the function
δ κ 2 ( γ , θ ) = κ 2 D 2 ( h ( · , γ ) , h ( · , θ ) ) = κ 2 Θ | h ( ω , γ ) h ( ω , θ ) | 2 μ ( d ω )   = κ 2 2 Θ h ( ω , γ ) h ( ω , θ ) ¯ + h ( ω , θ ) h ( ω , γ ) ¯ μ ( d ω ) ,
where κ > 0 is a positive constant which determines the units of the dissimilarity measure δ κ 2 on Θ , and when γ = θ , we have taken into account Θ h ( ω , θ ) h ( ω , θ ) ¯ μ ( d ω ) = Θ r 2 ( ω , θ ) μ ( d ω ) = Θ f ( ω , θ ) μ ( d ω ) = 1 θ Θ and therefore δ κ 2 ( θ , θ ) = 0 . Notice that since | h | = | h ¯ | = r , if we change the reference measure μ to ν accordingly, r changes to r ^ , and when we integrate (2) now with respect to ν , we find that δ κ 2 ( γ , θ ) remains invariant.
We are going to consider the Riemannian metric induced in Θ by δ κ 2 , following the same basic ideas of [10]. If we define d κ , θ 2 ( γ ) = δ κ 2 ( γ , θ ) , we will have d κ , θ 2 ( θ ) = 0 , and taking into account that the Jacobian of d κ , θ 2 ( γ ) in γ = θ is also null, since the function has an absolute minimum at θ , the second-order expansion of d κ , θ 2 ( γ ) at θ will be given by its Hessian at θ , ( d κ , θ 2 ( θ ) ) = 2 d κ , θ 2 γ i γ j γ = θ ( m × m ) , which will be positive definite or semidefinite. Moreover, the Hessian components at γ = θ are the components of a second-order covariant symmetric tensor at each tangent space T θ Θ , see [1] for details, in which tensor may be used to define a Riemannian metric in Θ with fundamental tensor
g i j ( θ ) = 1 2 2 d κ , θ 2 γ i γ j | γ = θ i , j = 1 , , m ,
i.e., the components of the metric tensor will be equal to one half times the second partial derivatives of the function d κ , θ 2 , and the line element corresponding to the above-mentioned induced Riemannian metric, under θ = ( θ 1 , , θ m ) , is given through its square by
d s 2 = i = 1 m j = 1 m g i j ( θ ) d θ i d θ j .
This formulation is typical in classical information geometry, where the metric is derived from second derivatives of divergence functions or similar measures.
If we express the complex-valued function h ( · , θ ) in their exponential form, that is, h ( · , θ ) = r ( · , , θ ) e i ζ ( · , θ ) for a convenient real-valued function ζ of ( · , θ ) , from (2), taking into account that if z = r e i ζ then z ¯ = r e i ζ , we obtain after some laborious but straightforward computations at the point γ = θ , and we have
1 2 2 d κ , θ 2 γ β γ α γ = θ = κ 2 ( Θ ( 2 r γ β γ α ( ω , γ ) γ = θ r ( ω , θ ) r 2 ( ω , θ ) ζ γ α ( ω , γ ) γ = θ ζ γ β ( ω , γ ) γ = θ ) μ ( d ω ) ) ,
see [1], but if we take into account that r ( ω , θ ) = f ( ω , θ ) is the real square root of a probability density with respect to the reference measure μ , we can express (5) as
1 2 2 d κ , θ 2 γ β γ α γ = θ = κ 2 4 Θ ( ln f γ β ( ω , γ ) γ = θ ln f γ α ( ω , γ ) γ = θ f ( ω , θ ) 2 2 f γ β γ α ( ω , γ ) γ = θ + 4 f ( ω , θ ) ζ γ α ( ω , γ ) γ = θ ζ γ β ( ω , γ ) γ = θ ) μ ( d ω ) = κ 2 4 ( Θ ln f γ β ( ω , γ ) γ = θ ln f γ α ( ω , θ ) γ = θ f ( ω , θ ) μ ( d ω ) + 4 Θ ζ γ α ( ω , γ ) γ = θ ζ γ β ( ω , γ ) γ = θ f ( ω , θ ) μ ( d ω ) ) .
Moreover, let I = ( i α β ) be the ( m × m ) Fisher information matrix; for more information refer to the original work by Fisher (1922) [11], corresponding to the density of the model, that is, i α β = Θ ln f θ α ln f θ β f d μ and additionally define
ϰ α β ( θ ) = 4 Θ ζ γ α ( ω , γ ) γ = θ ζ γ β ( ω , γ ) γ = θ f ( ω , θ ) μ ( d ω ) ,
we obtain that the fundamental tensor of Θ , (3), converting it as a Riemannian manifold, is equal to
g α β ( θ ) = κ 2 4 i α β ( θ ) + ϰ α β ( θ ) ,
and the Riemannian metric, using repeated index summation convention, can also be expressed as
d s 2 = g α β d θ α d θ β = κ 2 4 i α β d θ α d θ β + ϰ α β d θ α d θ β = κ 2 4 d s I 2 + d s A 2 ,
where d s I 2 is the information metric and d s A 2 is a metric specifically related to the imaginary part of the function that defines the model. Observe that locally, the Riemannian metric induced in Θ is the distance induced by the Hilbert metric structure in L 2 ( Θ , d μ ) on a radius 2 sphere and coincides, when κ = 2 , with the information metric when ζ is constant. More results on the information metric can be found in [10,12,13], among others.
Please note that while we assume that all properties of an information source are determined by θ or h ( · , θ ) , the exact value of θ has to be estimated from the data generated by the source. Although we assume that these data can remain partially hidden, it should still allow a reasonably good estimate of both θ and h ( · , θ ) .

2.2. Physical Applications

In the first manuscript [1], we introduced the concept of extended information, codified by the estimation ω , obtained from an information source, of the true parameter θ , which is information relative to this true value of the parameter, and referred to an arbitrary point θ 0 , as
I ω ( θ ) = ι log h ( ω , θ ) log h ( ω , θ 0 )   = ι 1 2 ln r 2 ( ω , θ ) + i ζ ( ω , θ ) 1 2 ln r 2 ( ω , θ 0 ) i ζ ( ω , θ 0 ) ,
where log denotes a version of the complex logarithm defined to satisfy log h ( ω , θ ) = ln r ( ω , θ ) + i ζ ( ω , θ ) , ln being the standard natural logarithm, and ι is a constant that determines the units of information with which we will work. The implicit dependence of (10) on θ 0 is omitted from the notation, since its choice will not play any further role when we calculate its gradient in the parametric space. Additionally, (10) will remain invariant under appropriate data changes and, fixed ω is also invariant under coordinate changes in the parametric manifold, since it is a scalar field on Θ .
The information provided by a source external to the observer is represented within him, allowing him to increase his understanding of the objects. Although we can consider different levels and types of the said representation, we will focus on two critical aspects: the parameter space with its natural extended information geometry and the observer’s ability to construct a plausibility regarding the true value of the parameter, Ψ , in the parameter space Θ , once given a particular estimation of the true parameter, essentially a complex square root of a kind of subjective conditional probability density with respect to the Riemannian volume, which is induced by the information metric over the parameter space, Θ , up to a normalization constant. Specifically, we shall write at the beginning
Θ Ψ ( θ ) Ψ ( θ ) ¯ V ( d θ ) = a > 0 ,
although we shall be particularly interested in the case a = 1 . Observe that in (11), we are integrating with respect to the Riemannian measure defined in (3), and therefore this expression is invariant under coordinate changes. Furthermore, if we intend to define on the parametric manifold a probability, interpretable as a plausibility about the true value of the parameter, we can take the function Ψ Ψ ¯ as the Radon–Nikodym derivative of said probability with respect to the Riemannian volume, which is also a measure in the same parameter space. Both measures are independent of the coordinate system used and, therefore, Ψ Ψ ¯ will be an invariant scalar field on the parametric manifold Θ . Then, if we let ψ = | Ψ | , we can simply define the information encoded by the subjective plausibility Ψ ( θ ) = ψ ( θ ) e i ϱ ( θ ) on the parameter space Θ relative to the true parameter θ as
Λ ( θ ) = ι log ( Ψ ( θ ) ) = ι 1 2 ln ψ 2 ( θ ) + i ϱ ( θ ) .
This quantity (12) remains also invariant under coordinate changes in the parametric manifold, being another scalar field in Θ ; see also [14].

The Variational Principle

In this context, trusting that many of the abilities of the observer have been efficiently shaped by natural selection in the process of biological evolution, we propose that the subjective information mentioned above adjusts in some way to the information provided by the source and in particular satisfying the following variational principle; see also [14]. However, in this paper, we are going to solve a slightly more flexible variational principle considering that given ω , Ψ vanishes outside the submanifold Θ ω Θ , which is a submanifold that is possibly dependent on ω but eventually could be the whole Θ , and, therefore, the functional to be minimized or at least make stationary will become the following
Ω ( Ψ ) = Θ ω | grad I ω ( θ ) grad Λ ( θ ) | 2 Ψ ( θ ) Ψ ( θ ) ¯ V ( d θ ) ,
subject to the constraint (11), which may be written as, subject to the constraint
Θ ω ψ 2 ( θ ) d V = a > 0 ,
and, potentially, to another constraint
Θ ω H ( θ ) ψ 2 ( θ ) V ( d θ ) = b .
This additional constraint is intended as a basic first attempt to model the subject’s limitations in apprehending the object of knowledge. The subject seeks to describe the object by exploiting as fully as possible the information generated by it, as quantified in (10). However, this exploitation—formulated in terms of the variational principle (13)—cannot be achieved without a bound; rather, it can only be partially attained. This limitation arises fundamentally from the constraints inherent in the subject, which are precisely represented by the additional constraint aforementioned (15).
Moreover, it will also be assumed that Ψ and its gradient grad Ψ vanish at the boundary Θ ω or at infinity, which is the way to model that we have strong reasons to believe that the true parameter θ should belong far from the boundary and clearly inside of Θ ω . Notice that the functional Ω is equal to the expected value corresponding to the probability in Θ given by the density Ψ Ψ ¯ with respect to the Riemannian volume induced by the metric (9), V ( d θ ) , of the square of the norm, corresponding to the vector field grad I ω ( θ ) minus grad Λ ( θ ) .
Observe that the difference between the gradients grad I ω ( θ ) and grad Λ ( θ ) is a complexified vector field in a Riemannian manifold. If we denote by r ω ( θ ) = r ( ω , θ ) and ζ ω = ζ ( ω , θ ) for a specific ω , and omitting for simplicity the dependence with variable θ , we have grad I ω grad Λ = ι ( 1 2 grad ( ln ψ 2 r ω 2 ) ) + i ( grad ( ϱ ζ ω ) ) and therefore | grad I ω grad Λ | 2 = ι 2 ( 1 4 | grad ( ln ψ 2 r ω 2 ) | 2 + | grad ( ϱ ζ ω ) | 2 ) ; then, the variational principle (13) may be expressed in terms of ψ and ϱ as follows
Ω ( ψ , ϱ ) = ι 2 Θ ω 1 4 | grad ln ψ 2 r ω 2 | 2 + | grad ( ϱ ζ ω ) | 2 ψ 2 d V ,
where d V indicates the integration with respect to the Riemannian measure on Θ , subject to the constraints (14) and (15) and assuming that ψ and its gradient grad ψ vanish at the boundary Θ ω or at infinity.
Observe that (16) is invariant under coordinate changes in Θ , since the square of the norm inside the integral is invariant and ψ 2 d V = Ψ Ψ ¯ V ( d θ ) = Ψ Ψ ¯ g d θ is also, where g is the determinant of the metric tensor (3).
Notice also that the source is considered as something objective or at least, strictly speaking, intersubjective, while the parameter space, with its geometric properties, in some sense is built by the observer, and is therefore subjective, although it is strongly conditioned by the source. In future work, some additional restrictions will be added to the variational principle (13) with the aim of more accurately modeling the observation process as a whole.
Any change in the information encoded by ω , caused by considering a change at the source in the parameter space, should correspond to a change in the subjective information proposed by the observer. For this reason, we propose that the squared difference of both gradients grad I ω and grad Λ would be on average Ω as small as possible, at least locally.
This variational principle is an extension of a previous one presented by the authors in [14], where only regular parametric statistical models with the standard information metric were considered, and, based on the former variational principle applied to these models and playing only with basic statistical tools, in particular the information carried by the data in a given statistical model, see [15,16,17]; among many others, a probability density was obtained in the parameter space Θ by solving a system of partial differential equations. This probability can be viewed as a Bayesian posterior probability over all possible probabilistic mechanisms that have generated the data identifiable with the parameters. Furthermore, if we apply this procedure of analyzing data to the simple statistical model corresponding to a multivariate normal distribution with a constant covariance matrix, which is a model that can be considered at least as an approximation of many regular statistical models, for large samples via extensions of the Central Limit Theorem, we obtain, from the partial differential equations in the parameter space mentioned above, a specific differential equation already studied in Physics and known as the stationary (time-independent) Schrödinger equation applied to a quantum harmonic oscillator; see [14] for further details.

2.3. Solving the Variational Problem

Since (16) is an optimization problem with, at least, constraint (14), we may introduce the augmented Lagrangian
L λ 1 , λ 2 , a , b ( ψ , ϱ ) = ι 2 Θ ω 1 4 | grad ( ln ( ψ 2 r ω 2 ) ) | 2 + | grad ( ϱ ζ ω ) | 2 ψ 2 d V   λ 1 Θ ω ψ 2 d V a λ 2 Θ ω H ψ 2 d V b ,
where λ 1 and λ 2 are constant Lagrange multipliers. Following the previously introduced notation, if we define L ω ( θ ) = f ( ω , θ ) = r ω 2 ( θ ) , in fact the likelihood corresponding to ω , since grad ( ln ψ 2 ) = 1 ψ 2 ( 2 ψ grad ψ ) = 2 ψ grad ψ , we have the following
| grad I ω grad Λ | 2 = ι 2 4 | 2 ψ grad ψ grad ln L ω | 2   + ι 2 grad ( ϱ ζ ω ) 2 .
This expression is invariant under coordinate changes in Θ . Let η 1 and η 2 be an arbitrary real-valued function of θ but assume that ( ψ + ϵ η 1 ) e i ( ϱ + ϵ η 2 ) satisfies (14). Then, if we let L L λ 1 , λ 2 , a , b ( ψ + ϵ η 1 , ϱ + ϵ η 2 ) , we have
L = ι 2 Θ ω { | 2 ψ + ϵ η 1 grad ( ψ + ϵ η 1 ) grad ln L ω | 2 ( ψ + ϵ η 1 ) 2 + | grad ( ϱ + ϵ η 2 ζ ω ) | 2 ( ψ + ϵ η 1 ) 2 } d V λ 1 Θ ω ( ψ + ϵ η 1 ) 2 d V a λ 2 Θ ω H ( ψ + ϵ η 1 ) 2 d V b = ι 2 Θ ω { 4 | grad ( ψ + ϵ η 1 ) | 2 + | grad ln L ω | 2 ( ψ + ϵ η 1 ) 2 4 grad ln L ω , grad ( ψ + ϵ η 1 ) ( ψ + ϵ η 1 ) + | grad ( ϱ + ϵ η 2 ζ ω ) | 2 ( ψ + ϵ η 1 ) 2 } d V λ 1 Θ ω ( ψ + ϵ η 1 ) 2 d V a λ 2 Θ ω H ( ψ + ϵ η 1 ) 2 d V a ,
and therefore
L = ι 2 Θ ω { ϵ 2 4 | grad η 1 | 2 + | grad ln L ω | 2 η 1 2 4 grad ln L ω , grad η 1 η 1 + | grad ( ϱ ζ ω ) | 2 η 1 2 + | grad η 2 | 2 ψ 2 + 4 grad ( ϱ ζ ω ) , grad η 2 ψ η 1 λ 1 η 1 2 λ 2 H η 1 2 + ϵ ( 8 grad ψ , grad η 1 + 2 | grad ln L ω | 2 ψ η 1   4 grad ln L ω , grad ( ψ η 1 ) + 2 grad ( ϱ ζ ω ) , grad η 2 ψ 2   + 2 | grad ( ϱ ζ ω ) | 2 ψ η 1 2 λ 1 ψ η 1 2 λ 2 H ψ η 1 ) + 4 | grad ψ | 2 + | grad ln L ω | 2 ψ 2 4 grad ln L ω , grad ψ ψ   + | grad ( ϱ ζ ω ) | 2 ψ 2 λ 1 ψ 2 λ 2 H ψ 2 } d V + λ 1 a + λ 2 b ,
where we have taken into account grad ( η 1 ψ ) = η 1 grad ψ + ψ grad η 1 . Then, the derivative of (20) with respect to ϵ will be
d L λ 1 , λ 2 , a , b d ϵ = ι 2 Θ ω 2 ϵ 4 grad η 1 2 + grad ln L ω 2 η 1 2 4 grad ln L ω , grad η 1 η 1 + grad ϱ ζ ω 2 η 1 2 + grad η 2 2 ψ 2 + 4 grad ϱ ζ ω , η 2 ψ η 1 ( λ + μ H ) η 1 2 + 8 grad ψ , grad η 1 + 2 grad ln L ω 2 ψ η 1 4 grad ln L ω , grad ψ η 1 + 2 grad ϱ ζ ω , grad η 2 ψ 2 + 2 grad ϱ ζ ω 2 ψ η 1 2 λ 1 + λ 2 H ψ η 1 d V ,
and the first variation of the augmented Lagrangian L λ 1 , λ 2 , a , b
δ L λ 1 , λ 2 , a , b ( ψ , ϱ , η 1 , η 2 ) d L λ 1 , λ 2 , a , b d ϵ | ϵ = 0 = ι 2 Θ ω { 8 grad ψ , grad η 1 + 2 | grad ln L ω | 2 ψ η 1   4 grad ln L ω , grad ( ψ η 1 ) + 2 grad ( ϱ ζ ω ) , grad η 2 ψ 2   + 2 | grad ( ϱ ζ ω ) | 2 ψ η 1 2 ( λ 1 + λ 2 H ) ψ η 1 } d V .
Observe now that we can use a well-known differential operator expression in a Riemannian manifold, div ( η 1 grad ψ ) = η 1 Δ ψ + grad η 1 , grad ψ , where Δ stands for the Laplacian operator. For further details, see, for example, [18]. Thus, we shall have
Θ ω grad ψ , grad η 1 d V = Θ ω div ( η 1 grad ψ ) d V Θ ω η 1 Δ ψ d V ,
but by the Gauss divergence theorem, we have
Θ ω div ( η 1 grad ψ ) d V = Θ ω η 1 grad ψ , ν d A = Θ ω η 1 grad ψ , ν d A = 0 ,
where ν is a unitary vector field on Θ ω pointing out Θ , and d A is the surface element in Θ ω induced by the information Riemannian metric on Θ , and taking into account that by the boundary conditions, grad ψ vanishes at Θ ω or at infinity, we obtain (24). Thus, substituting this result into (23), we obtain
Θ ω grad ψ , grad η 1 d V = Θ ω η 1 Δ ψ d V .
Additionally, since div η 1 ψ grad ln L ω ) = η 1 ψ Δ ln L ω + grad ( η 1 ψ ) , grad ln L ω , we shall have
Θ ω grad ( η 1 ψ ) , grad ln L ω d V = Θ ω div ( η 1 ψ grad ln L ω ) d V Θ ω η 1 ψ Δ ln L ω d V ,
but again, by the Gauss divergence theorem and the bi-linearity of the scalar product, we have
Θ ω div ( η 1 ψ grad ln L ω ) d V = Θ ω η 1 ψ grad ln L ω , ν d A = 0 ,
since, by the boundary conditions, ψ vanishes at Θ ω or at infinity, and therefore we have (27). Then, substituting this equation into (26), we obtain
Θ ω grad ( η 1 ψ ) , grad ln L ω d V = Θ ω η 1 ψ Δ ln L ω d V .
And since div η 2 ψ 2 grad ( ϱ ζ ω ) = η 2 ψ 2 Δ ( ϱ ζ ω ) + grad ( ψ 2 η 2 ) , grad ( ϱ ζ ω ) , we shall have
Θ ω grad ( η 2 ψ 2 ) , grad ( ϱ ζ ω ) d V = Θ ω d i v ( η 2 ψ 2 grad ( ϱ ζ ω ) ) d V   Θ ω η 2 ψ 2 Δ ( ϱ ζ ω ) d V ,
but once again by the Gauss divergence theorem,
Θ ω d i v ( η 2 ψ 2 grad ( ϱ ζ ω ) d V = Θ ω η 2 ψ 2 grad ( ϱ ζ ω ) ν d A = 0 ,
where ν is a unitary vector field on Θ ω pointing out Θ , and d A is the surface element in Θ ω , since, by the boundary conditions, ψ 2 vanishes at Θ ω or at infinity, and thus we have (31).
Θ ω grad ( η 2 ψ 2 ) , grad ( ϱ ζ ω ) d V = Θ ω η 2 ψ 2 Δ ( ϱ ζ ω ) d V ,
and taking into account that grad ( ϱ ζ ω ) , grad η 2 ψ 2 = grad ( ϱ ζ ω ) , grad ( ψ 2 η 2 ) 2 grad ( ϱ ζ ω ) , grad ψ ψ η 2 , we have
Θ ω grad η 2 , grad ( ϱ ζ ω ) ψ 2 d V = Θ ω η 2 ψ 2 Δ ( ϱ ζ ω ) d V   2 Θ ω η 2 ψ grad ψ , grad ( ϱ ζ ω ) d V ,
and combining (22), (25), (28) and (32), the first variation δ L δ L λ 1 , λ 2 , a , b ( ψ , ϱ , η 1 , η 2 ) can be expressed as
δ L = ι 2 Θ ω 2 η 1 { 4 Δ ψ + | grad ln L ω | 2 ψ + 2 ψ Δ ln L ω   + | grad ( ϱ ζ ω ) | 2 ψ ( λ 1 + λ 2 H ) ψ } 2 η 2 { ψ 2 Δ ( ϱ ζ ω ) + 2 ψ grad ψ , grad ( ϱ ζ ω ) } d V .
this expression (33) must vanish for all η 1 and η 2 , which leads to the following fundamental equations
4 Δ ψ + | grad ln L ω | 2 + | grad ( ϱ ζ ω ) | 2 + 2 Δ ln L ω λ 1 λ 2 H ψ = 0 ,
ψ 2 Δ ( ϱ ζ ω ) + 2 ψ grad ψ , grad ( ϱ ζ ω ) = 0 .
In summary, using standard techniques from the calculus of variations, we derive the necessary conditions for the variational problem (13) subject to the constraints (14) and (15). These conditions can be expressed as two coupled systems of partial differential equations: the first system determines the modulus of the function Ψ , while the second governs its argument. The resulting fundamental Equation (34) may be written as
4 Δ ψ + ( | grad ln L ω | 2 + | grad ( ϱ ζ ω ) | 2 ) ψ = ( 2 Δ ln L ω + λ 1 + λ 2 H ) ψ
and Equation (35) becomes
div ( ψ 2 grad ( ϱ ζ ω ) ) = 0 ,
which suggests a steady-state or conservation law.
We defer to future work the investigation of situations in which the above boundary conditions must be replaced by alternative ones. Such modifications would introduce additional terms into the fundamental equations, arising from the corresponding changes in Equations (27) and (32), which were obtained through an application of Gauss’s divergence theorem. Consequently, Equations (36) and (37) would also require appropriate modification.

2.4. An Example

To illustrate the significant implications of the two equations, we attempt to model a situation in which we will have two approximately independent sources of information and several replicas of size k 1 > 0 and k 2 > 0 , respectively, that are not necessarily equal from each of them, although k 1 and k 2 may be linked to each other, as we will discuss later. Based on this basic model, we will consider that we will have random samples of size n > 0 . The first source of information may be modeled by a univariate Gaussian distribution with mean μ and constant variance σ 1 2 . The second source, at the beginning, will be modeled by an absolutely continuous translation invariant family of probability functions defined through a convenient function q; specifically, if we let x = ( x 1 , , x k 1 ) and y = ( y 1 , , y k 2 ) , we shall have the statistical model defined by the joint density
p ( x , y , μ , τ ) = α = 1 k 1 1 2 π σ 1 e 1 2 σ 1 2 ( x α μ ) 2 β = 1 k 2 q 2 ( y β τ ) ,
where q is the real-valued smooth function of a real variable such that
q 2 ( u ) d u = 1 and 0 < 4 ( q ) 2 ( u ) d u = 4 a 2 < , with a > 0 .
but for simplicity, hereafter we explore the case q ( u ) = 1 2 π σ 2 2 4 e u 2 4 σ 2 2 , which is the real non-negative square root of a normalized Gaussian distribution. Then, we shall have a = n k 2 2 σ 2 , and the parameter space of this model is Θ = R 2 ; thus, we shall identify that the points θ Θ with their coordinates ( μ , τ ) and ( μ , τ ) will be the corresponding basis vector field in T Θ . The likelihood, based on a random sample of size n of the joint model (38), will be
L ( x , y ) ( μ , τ ) = j = 1 n { α = 1 k 1 1 2 π σ 1 e 1 2 σ 1 2 ( x j α μ ) 2 β = 1 k 2 1 2 π σ 2 e 1 2 σ 2 2 ( y j β τ ) 2 } = σ 1 k 1 σ 2 k 2 2 π n e n 2 k 1 σ 1 2 s x 2 + k 2 σ 2 2 s y 2 e n 2 k 1 σ 1 2 ( x ¯ n k 1 μ ) 2 + k 2 σ 2 2 ( y ¯ n k y τ ) 2 ,
where x j α and y j β are the α and β replicas of the underlying random variables X and Y corresponding to the sample j, and
x ¯ n k 1 = 1 n k 1 j = 1 n α = 1 k 1 x j α , s x 2 = 1 n k 1 j = 1 n α = 1 k 1 ( x j α x ¯ n k 1 ) 2 , y ¯ n k 2 = 1 n k 2 j = 1 n β = 1 k 2 y j β , s y 2 = 1 n k 2 j = 1 n β = 1 k 2 ( y j α y ¯ n k y ) 2
From (40), it is straightforward to obtain the maximum-likelihood estimation of ( μ , τ ) as ω = ( x ¯ n k 1 , y ¯ n k 2 ) . If θ = ( μ , τ ) and ω = ( u , v ) , the joint distribution of the maximum-likelihood estimator is
f ( ω , θ ) = f ( u , v , μ , τ ) = n k 1 2 π σ 1 2 e n k 1 2 σ 1 2 ( u μ ) 2 n k 2 2 π σ 2 2 e n k 2 2 σ 2 2 ( v τ ) 2 ,
This expression (41) leads us to consider the regular parametric family of functions F Θ defined as
h ( u , v , μ , τ ) = n 2 π σ 1 σ 2 k 1 k 2 4 e n k 1 4 σ 1 2 ( u μ ) 2 e n k 2 4 σ 2 2 ( v τ ) 2 e i ζ ( v τ ) ,
where ζ is a convenient real-valued smooth function of a real variable such that
( ζ ) 2 ( v ) d v = k 2 < , where k > 0 .
Observe that f = h h ¯ and the corresponding density of the joint model is just (41).
If we let κ = 2 , we have ϰ 11 = ϰ 12 = ϰ 21 = 0 and ϰ 22 = 4 k 2 , and it is straightforward to obtain the induced Riemannian geometry (8) in the parameter space Θ = R 2 , which is given by
g 11 = n k 1 σ 1 2 , g 12 = g 21 = 0 , g 22 = n k 2 σ 2 2 + 4 k 2 ,
the metric is Euclidean and, taking into account (40), we have
ln L ( x , y ) μ = n k 1 σ 1 2 ( x ¯ n k 1 μ ) , ln L ( x , y ) τ = n k 2 σ 2 2 ( y ¯ n k 2 τ ) ,
Then, we have
| grad ln L ( x , y ) | 2 = n k 1 σ 1 2 ( x ¯ n k 1 μ ) 2 + n k 2 σ 2 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) ( y ¯ n k 2 τ ) 2 ,
Δ ψ = σ 1 2 n k 1 2 ψ μ 2 + σ 2 2 n k 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) 2 ψ τ 2 ,
Δ ln L ( x , y ) = σ 1 2 n k 1 2 ln L ( x , y ) μ 2 + σ 2 2 n k 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) 2 ln L ( x , y ) τ 2   = 1 1 1 + 4 σ 2 2 n k 2 k 2 .
Then, Equation (34) will be written as
4 { σ 1 2 n k 1 2 ψ μ 2 + σ 2 2 n k 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) 2 ψ τ 2 } + { n k 1 σ 1 2 ( x ¯ n k 1 μ ) 2 + n k 2 σ 2 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) ( y ¯ n k 2 τ ) 2 + | grad ( ϱ ζ ) | 2 } ψ = 2 + 2 1 + 4 σ 2 2 n k 2 k 2 + λ 1 + λ 2 H ψ ,
Additionally, assume that H and | grad ( ϱ ζ ) | 2 are only functions of τ and do not depend on μ . Then, Equation (49) admits a solution of separate variables of the form ψ ( μ , τ ) = ϕ ( μ ) υ ( τ ) , and it can be written as
4 σ 1 2 n k 1 υ 2 ϕ μ 2 + n k 1 σ 1 2 ( x ¯ n k 1 μ ) 2 2 υ ϕ = 4 σ 2 2 n k 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) ϕ 2 υ τ 2 + ( 2 1 + 4 σ 2 2 n k k 2 + λ 1 + λ 2 H n k 2 σ 2 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) ( y ¯ n k 2 τ ) 2 | grad ( ϱ ζ ) | 2 ) υ ϕ ,
and therefore, dividing by ϕ υ , it follows that
4 σ 1 2 n k 1 1 ϕ 2 ϕ μ 2 + n k 1 σ 1 2 ( x ¯ n k 1 μ ) 2 2 = 4 σ 2 2 n k 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) 1 υ 2 υ τ 2 + 2 1 + 4 σ 2 2 n k 2 k 2 + λ 1 + λ 2 H n k 2 σ 2 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) ( y ¯ n k 2 τ ) 2 | grad ( ϱ ζ ) | 2 = ξ ,
where ξ is a constant that is independent of μ and τ . Then, with minor changes, we have the following.
2 σ 1 2 n k 1 2 ϕ μ 2 + n k 1 2 σ 1 2 ( x ¯ n k 1 μ ) 2 ϕ = 1 + ξ 2 ϕ
2 υ τ 2 = { n k 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) 4 σ 2 2 ξ + | grad ( ϱ ζ ) | 2 λ 1 λ 2 H + n k 2 4 σ 2 2 n k 2 σ 2 2 ( y ¯ n k 2 τ ) 2 2 } υ ,
With the present model (42) and assumptions, in components, since grad ( ϱ ζ ) = ( 0 , σ 2 2 n k 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) ( ϱ τ ζ τ ) ) and if we let Z = ψ 2 grad ( ϱ ζ ) in components and taking into account, the repeated index summation convention, and recalling that since the geometry is Euclidean, the Christoffel symbols are null, then we have div ( Z ) = Z α ,   α = Z 2 τ and Equation (37), and we obtain
div ( ψ 2 grad ( ϱ ζ ) ) = τ ϕ 2 υ 2 σ 2 2 n k 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) ϱ τ ζ τ = ϕ 2 σ 2 2 n k 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) τ υ 2 ϱ τ ζ τ = 0 ,
which implies
υ 2 ϱ τ ζ τ = C t e ,
where C t e is a real constant. Therefore,
| grad ( ϱ ζ ) | 2 = σ 2 2 n k 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) ϱ τ ζ τ 2 = σ 2 2 n k 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) C t e 2 υ 4 ,
| grad ( ϱ ζ ) | = σ 2 n k 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) | C t e | υ 2 ,
and
ϱ τ = C υ 2 + ζ τ .
where C = ± σ 2 | C t e | / n k 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) . Equation (52) is the time-independent Schrödinger equation for the quantum harmonic oscillator. Together with the prescribed boundary conditions, it admits a discrete family of solutions for the functions ϕ and ξ (see [14]). In particular, the parameter ξ takes the form ξ = 4 ν , where ν N . Furthermore, the function ϱ depends on a constant C, which can be related to the energy spectrum of the system. These energy levels depend on the parameter ν , as will be discussed later.
Assume that the subject-related constraint is obtained requiring that the part of the gradient that depends on τ is constant, that is, something like
H = n k 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) σ 2 2 ( y ¯ n k 2 τ ) 2 + | grad ( ϱ ζ ) | 2 = b .
Then, Equation (53) becomes
2 υ τ 2 = n k 2 4 σ 2 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) ξ λ 1 + ( 1 λ 2 ) b 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) υ .
which is a second-order linear and homogeneous differential equation with constant coefficients. In order to obtain a function υ whose square is a density on the real line, we probably need to think in a large interval [ y ¯ n k 2 L , y ¯ n k 2 + L ] , i.e., with L > 0 , making the length of the interval large and then taking the limit in case it would be necessary. We are going to study different cases, depending on the sign of the constant C E defined as
C E = n k 2 4 σ 2 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) ξ λ 1 + ( 1 λ 2 ) b 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) ,
or, since ξ = 4 ν where ν N , we have
C E = n k 2 4 σ 2 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) 4 ν λ 1 + ( 1 λ 2 ) b 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) .
We may also impose υ ( y ¯ n k 2 L ) = υ ( y ¯ n k 2 + L ) = 0 and y ¯ n k 2 L y ¯ n k 2 + L υ 2 d τ = 1 .
In order to solve (60) with the boundary conditions mentioned previously, it is convenient to observe that if we define u = τ y ¯ n k 2 and the real function υ ^ ( u ) = υ ^ ( τ y ¯ n k 2 ) = υ ( τ ) , i.e., υ ^ ( u ) = υ ( u + y ¯ n k 2 ) , then the function υ ^ also satisfies (60) but with the boundary conditions expressed as υ ^ ( L ) = υ ^ ( L ) = 0 and L L υ ^ 2 d u = 1 .
  • CASE (1) CE = 0.
In this case,
2 υ ^ u 2 = 0 υ ^ ( u ) = A 1 u + A 2 ,
but the boundary conditions in L and L, since L > 0 , imply A 1 = A 2 = 0 , but L L υ ^ 2 d u = 0 1 and therefore this case does not provide an admissible solution for υ ^ nor, therefore, for υ .
  • CASE (2) CE > 0.
In this case,
2 υ ^ u 2 = C E υ ^ υ ^ ( u ) = A 1 e C E u + A 2 e C E u ,
but in the boundary, we have υ ^ ( L ) = A 1 e C E L + A 2 e C E L = 0 and L L ( A 1 e C E u + A 2 e C E u ) 2 d u = 1 . Therefore, if we let α = e C E L > 0 , β = e C E L > 0 , we have
A 1 β + A 2 α = 0 ,
A 1 α + A 2 β = 0 ,
L L A 1 2 e 2 C E u + A 2 2 e 2 C E u + 2 A 1 A 2 d u = 1 , A 1 2 2 C E e 2 C E u A 2 2 2 C E e 2 C E u + 2 A 1 A 2 u L L = 1 , 1 2 C E A 1 2 + A 2 2 β 2 α 2 + 2 A 1 A 2 L = 1 .
From (65) and (66) we obtain that if α β , then A 1 = A 2 = 0 and L L υ ^ 2 d u = 0 1 , which are not admissible solutions, but in the case α = β , then A 1 + A 2 = 0 , and L L υ ^ 2 d u = 2 A 1 2 L < 0 1 , which are also not admissible solutions. Then, if C E > 0 , there are no admissible solutions for υ ^ nor, therefore, for υ .
  • CASE (3) CE < 0.
In this case,
υ ^ u 2 = C E υ ^ υ ^ ( u ) = A 1 cos C E u + A 2 sin C E u ,
but, as in the previous case, on the boundary, we have υ ^ ( L ) = A 1 cos C E L A 2 sin C E L = 0 and υ ^ ( L ) = A 1 cos C E L + A 2 sin C E L = 0 and also L L A 1 cos C E u + A 2 sin C E u 2 d u = 1 . Therefore, if we let α = cos ( C E L ) , β = sin ( C E L ) , we have
A 1 α A 2 β = 0 ,
A 1 α + A 2 β = 0 ,
L L ( A 1 2 cos 2 C E u + A 2 2 sin 2 C E u + 2 A 1 A 2 cos C E u sin C E u ) d u = 1 , L L ( A 1 2 1 + cos 2 C E u 2 + A 2 2 1 cos 2 C E u 2 + A 1 A 2 sin 2 C E u ) d u = 1 , ( A 1 2 + A 2 2 ) L + A 1 2 A 2 2 2 C E sin 2 C E L = 1 .
From Equations (69) and (70), we have several possible solutions.
Case A 1 , A 2 0 .
In this case, the linear system with α and β as unknowns has no null determinant, and then α and β must be equal to zero, but it is not possible to vanish cos ( x ) and sin ( x ) at the same time: there are no admissible solutions for υ ^ in this case.
Case A 1 = 0 .
Then, A 2 0 , otherwise (71) cannot be satisfied, but in that case sin ( C E L ) = 0 , and since L > 0 , then C E L = m π with m Z equivalent to C E L 2 = 4 m 2 π 2 for m N . Moreover, since sin ( 2 C E L ) = sin ( 2 m π ) = 0 and (71), we have A 2 = ± 1 L and
υ ^ ( u ) = ± 1 L sin C E u = ± 1 L sin m π L u ,
and therefore
υ ( τ ) = ± 1 L sin m π L ( τ y ¯ n k 2 ) ,
but in order to be ϱ well defined, since Equation (58), we need υ ( τ ) 0 when τ [ y ¯ n k 2 L , y ¯ n k 2 + L ] and (73) is not an admissible solution, since υ ( y ¯ n k 2 ) = 0 .
Case A 2 = 0 .
Then, A 1 0 ; otherwise, (71) cannot be satisfied, but in that case cos ( C E L ) = 0 , but since L > 0 , then ± C E L = π 1 2 + m with m Z , which implies that C E L 2 = 1 2 + m 2 π 2 for m N . Moreover, since sin ( 2 C E L ) = sin ( 2 π 1 2 + m ) = 0 and (71), we have A 1 = ± 1 L and
υ ^ ( u ) = ± 1 L cos C E u = ± 1 L cos π L 1 2 + m u ,
and therefore
υ ( τ ) = ± 1 L cos π L 1 2 + m ( τ y ¯ n k 2 ) ,
and, in order for ϱ to be well defined, since Equation (58), we need υ ( τ ) 0 when τ [ y ¯ n k 2 L , y ¯ n k 2 + L ] . From (58), we have the following
ϱ ( τ ) = ϱ ( y ¯ n k 2 ) + y ¯ n k 2 τ C L cos 2 ± π L 1 2 + m ( τ y ¯ n k 2 ) d τ + ζ ( τ ) ζ ( y ¯ n k 2 )   = ϱ ( y ¯ n k 2 ) + C L ± π L 1 2 + m tan ± π L 1 2 + m ( τ y ¯ n k 2 ) + ζ ( τ ) ζ ( y ¯ n k 2 ) .
When τ y ¯ n k 2 and ϱ ( y ¯ n k 2 ) = 0 , we can say that approximately
ϱ ( τ ) C L ( τ y ¯ n k 2 ) + ζ ( τ ) ζ ( y ¯ n k 2 ) .
If C C E = π L 1 2 + m , then we approximately have
ϱ ( τ ) π 1 2 + m ( τ y ¯ n k 2 ) + ζ ( τ ) ζ ( y ¯ n k 2 ) .
Using the expression for the energy levels of the quantum harmonic oscillator, we can conclude that
ϱ ( τ ) π n 2 E m τ y ¯ n k 2 + ζ ( τ ) ζ ( y ¯ n k 2 ) ,
where E m = 2 n 1 2 + m represents the energy levels of the quantum harmonic oscillator as a function of m N , which follows the notation in [14] with respect to ν .
Observe that with a Taylor expansion at x = 0 and for a given constant a, we shall have tan ( a x ) + ( a x ) 2 cot ( a x ) = a x + ( a x ) 5 9 + and therefore
b a tan ( a x ) ( a x ) 2 cot ( a x ) 2 b x + O ( x 5 ) .
If now we consider that a = ± π L 1 2 + m , b = C L , ϱ ( y ¯ n k 2 ) = 0 , ζ ± π L 1 2 + m ( y ¯ n k 2 τ ) = ζ π L 1 2 + m ( τ y ¯ n k 2 ) , and x = ( τ y ¯ n k 2 ) , we have
ϱ ( τ ) = C L ± π L 1 2 + m ( tan ( ± π L 1 2 + m ( τ y ¯ n k 2 ) )   ± π L 1 2 + m ( τ y ¯ n k 2 ) 2 cot π L 1 2 + m ( τ y ¯ n k 2 ) )   = 2 C L τ y ¯ n k 2 + O ( ( τ y ¯ n k 2 ) 5 ) .
If C 1 2 C E = π 2 L 1 2 + m , then we approximately have
ϱ ( τ ) π 1 2 + m τ y ¯ n k 2 .
Now, using the expression for the energy levels of the quantum harmonic oscillator E m , we have
ϱ ( τ ) π n 2 E m τ y ¯ n k 2 .
  • Final comments on CE
Since C E < 0 and n k 2 4 σ 2 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) > 0 , we have
4 ν λ 1 + ( 1 λ 2 ) b 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) < 0 ,
and
C E = π 2 L 2 1 2 + m 2 .
Then,
π 2 L 2 1 2 + m 2 = n k 2 4 σ 2 2 + k 2 4 ν λ 1 + ( 1 λ 2 ) b 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) .
therefore
π L 1 2 + m = n k 2 4 σ 2 2 + k 2 λ 1 4 ν ( 1 λ 2 ) b + 2 ( 1 + 4 σ 2 2 n k 2 k 2 ) .
For large σ 2 , the Gaussian distribution becomes nearly flat across the real line, serving as the opposite of Dirac’s δ . In this limit, we have a highly diffuse probability distribution, and the previous equation becomes
π L 1 2 + m = k λ 1 4 ν ( 1 λ 2 ) b ,
and if we choose λ 2 1 , we have
π L 1 2 + m = k λ 1 4 ν .
Now, by leveraging the expression for the energy levels of the quantum harmonic oscillator E m , we can write that
π L 1 2 + m = π L n 2 E m k λ 1 4 ν .

3. Discussion

This paper presents a rigorous framework that integrates concepts from functional analysis, differential geometry, and information theory to study parametric models, notably in the context of physical applications. It proposes an innovative geometric perspective on statistical models by endowing the parameter space with a Riemannian structure derived from divergence functions and complex-valued model functions. This approach offers deep insights into the intrinsic properties of models, the estimation procedures, and the nature of information encoding, which are central to both theoretical and applied sciences.
At the core of this framework is the notion of a family of complex-valued functions, h ( ω , θ ) , that encode the observation process, with θ representing the parameters of the model and ω denoting the source or data space. The complex exponential form of h, written as h ( ω , θ ) = r ( ω , θ ) e i ζ ( ω , θ ) , encapsulates both magnitude and phase information, which can be particularly relevant in physical contexts such as quantum mechanics.
The introduction of a reference measure μ for the model space ensures that the squared modulus r 2 ( ω , θ ) serves as a density with respect to μ , thereby facilitating the representation of the model as a family of probability measures P θ . This structure aligns with the classical statistical framework, where densities with respect to a base measure underpin likelihood functions and subsequent inference.
The assumption that the phase function ζ ( ω , θ ) is independent of μ , but that the magnitude r ( ω , θ ) depends on μ , highlights an essential subtlety: the model’s probabilistic content is primarily governed by the magnitude, while the phase adds a geometric or physical layer of complexity. The invariance of certain quantities under the change in the reference measure underscores the robustness of the model representation.
In this paper, we use a significant contribution introduced in [1] where a Riemannian metric on the parameter space Θ is constructed via the concept of a divergence-induced metric applying [10] by embedding the family of model functions h ( · , θ ) into the Hilbert space L 2 ( Θ , μ ) , and we leverage the Hilbert distance D 2 ( p , q ) to quantify the dissimilarity between models corresponding to different parameters.
The divergence δ κ 2 ( γ , θ ) acts as a squared distance measure scaled by a positive constant κ , which induces a Riemannian structure through its second derivatives at the point θ . The fundamental tensor g i j ( θ ) , obtained as the Hessian of the divergence, encapsulates the local geometric curvature of the parameter space and provides a natural metric for statistical inference, estimation, and hypothesis testing.
This geometric perspective aligns with the principles of information geometry, where the Fisher information metric is a canonical example. Here, the divergence function generalizes the notion of a distance and can be tailored to include phase information via the complex exponential form of h ( ω , θ ) . The invariance of the divergence δ κ 2 under the change in the reference measure further emphasizes its fundamental nature.
The line element d s 2 , expressed via the metric tensor g i j ( θ ) , enables the measurement of infinitesimal distances within the parameter space, providing a geometric interpretation of model sensitivity and the local structure of the statistical manifold. This geometric structure can be exploited to analyze the efficiency of estimators, the complexity of models, and the behavior of inference algorithms.
The detailed derivation of the metric tensor g α β ( θ ) from the divergence function involves second-order derivatives evaluated at the point θ . When the model function h ( ω , θ ) is expressed in exponential form, the Hessian can be decomposed into contributions from the magnitude r ( ω , θ ) and phase ζ ( ω , θ ) . The resulting metric tensor combines the Fisher information matrix i α β , derived from the log-likelihood derivatives, and an additional term ϰ α β , which accounts for the phase ζ ( ω , θ ) .
This decomposition highlights that the geometric structure on the parameter space is richer than the classical Fisher metric; it incorporates phase information that can be crucial in physical models like quantum mechanics or wave phenomena. When ζ ( ω , θ ) is constant, the metric reduces to a scaled version of the classical Fisher information, reaffirming the connection to traditional information geometry.
The explicit expression of the metric tensor as g α β ( θ ) = κ 2 4 ( i α β + ϰ α β ) provides a versatile mathematical tool for analyzing models. It facilitates the computation of geodesics, curvature, and volume elements, which are instrumental in understanding the model’s complexity, the efficiency of estimators, and the behavior of Bayesian posterior distributions.
The framework extends beyond purely statistical models into physical applications, where the concept of extended information plays a central role. The estimation ω of the true parameter θ from an information source encapsulates the notion of how data or signals encode information about underlying physical quantities.
The complex scalar I ω ( θ ) , which represents information relative to a reference point θ 0 , incorporates both magnitude and phase. Its invariance properties under data transformations and coordinate changes make it a robust measure for physical and informational interpretation. Observe that this quantity is essentially an extension of the information-theoretic entropy introduced by Shannon in [19], which is deeply related to thermodynamic entropy. In classical statistical mechanics, the Gibbs entropy coincides up to a multiplicative constant. This conceptual identification was clarified and systematized by Jaynes in [20,21], who showed that equilibrium ensembles arise from the constrained maximization of Shannon entropy, thereby casting statistical mechanics as an inference theory on probability measures.
The introduction of a subjective plausibility function Ψ ( θ ) , normalized with respect to the Riemannian volume element V ( d θ ) , enables the construction of a probability-like measure over the parameter space that reflects the observer’s belief about the true parameter. The scalar quantity Λ ( θ ) , representing the encoded information of this plausibility, is invariant under coordinate changes, aligning with the geometric viewpoint that the model’s structure is intrinsic.
This approach resonates with principles in physics, particularly in quantum mechanics, where wave functions encode subjective or epistemic information about physical states and the geometric structure of the state space influences the evolution and measurement outcomes. The invariance under the change of coordinates ensures that physical predictions are independent of the observer’s particular representation.
A key methodological innovation is the formulation of a variational principle that seeks to minimize or stationarize a functional Ω ( Ψ ) , which measures the discrepancy between the gradients of the information I ω and the plausibility Λ . This principle embodies a form of optimality or consistency condition: the observer’s subjective model should, in some sense, align with the information provided by the source.
The functional Ω ( Ψ ) is expressed as an expectation over the parameter space with the measure given by the plausibility Ψ Ψ ¯ . The integrand involves the norm of the difference of gradients, which, when expanded into real and imaginary parts, yields a natural interpretation akin to a least-squares fitting in a geometric setting.
The derivation of the equations associated with this variational problem involves sophisticated calculus on Riemannian manifolds, incorporating divergence, Laplacian, and boundary conditions. The resulting Equations (34) and (35) resemble Schrödinger-like equations, with potential terms that reflect the likelihood and phase structure, thereby bridging statistical inference with physical models.
The method for solving the variational problem uses the classical technique of Lagrange multipliers, yielding a system of coupled differential equations. The invariance properties under coordinate transformations, along with the boundary conditions, ensure that the solutions are both physically and statistically meaningful.
The detailed derivation of the first variation, involving integration by parts, a divergence theorem, and boundary conditions, exemplifies the rigorous mathematical treatment necessary for such advanced models. The resulting Euler–Lagrange equations provide conditions for the optimal plausibility functions Ψ , which can be interpreted as wave functions or probability amplitudes within the geometric framework.
The example involving a parametric family defined in Equation (42) demonstrates how the abstract framework can be instantiated in concrete models such as Gaussian-modulated wave functions with phase factors. The explicit calculations of the metric tensor, likelihood derivatives, and Laplacians serve to connect the geometric formalism with familiar quantum harmonic oscillator equations.
The separation of variables and reduction to differential equations analogous to Schrödinger equations exemplify the deep connection between information geometry and quantum physics. The eigenvalue problems arising from these equations suggest that the models’ structure naturally encodes physical phenomena such as quantization, energy levels, and wave propagation.
The example underscores the potential of the formalism to analyze complex physical systems, where the geometry of the parameter space influences the system’s behavior, and variational principles yield well-understood differential equations governing the model’s dynamics.
The framework presented opens the door to integrating geometric, statistical, and physical theories into a unified approach to modeling complex systems. By emphasizing invariance, curvature, and metric structures, it provides tools for analyzing model complexity, estimation efficiency, and the physical interpretation of information.
Future research could extend these ideas to non-parametric models, incorporate dynamics (e.g., in time-dependent systems), and explore the role of phase information in quantum information theory. Additionally, the connection to Schrödinger’s equations suggests that quantum-like models can be derived from informational principles, fostering interdisciplinary insights among physics, information theory, and geometry.

4. Conclusions

This paper continues the development of information-geometric models for data analysis and physics by focusing on their formulation and interpretation through variational principles. The work exploits the idea that the information is determined by a complex-valued function that characterizes the model. Starting from this function, it is possible to construct the action principle and probability-like measure over the parameter space that reflects the observer’s belief about the true parameter. The final equation for the subjective plausibility function resembles the Schrodinger equation from quantum mechanics.

Author Contributions

Conceptualization, D.B.-C. and J.M.O.; writing—original draft preparation, D.B.-C. and J.M.O.; writing—review and editing, D.B.-C. and J.M.O. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

No new data were created or analyzed in this study.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Bernal-Casas, D.; Oller, J.M. Information-Geometric Models in Data Analysis and Physics. Mathematics 2025, 13, 3114. [Google Scholar] [CrossRef] [Scilit]
  2. Wheeler, J.A.; Zurek, W.H. (Eds.) Quantum Theory and Measurement; Princeton University Press: Princeton, NJ, USA, 1983. [Google Scholar] [CrossRef] [Scilit]
  3. Lloyd, S. Programming the Universe: A Quantum Computer Scientist Takes on the Cosmos; Alfred A. Knopf: New York, NY, USA, 2006. [Google Scholar]
  4. Bekenstein, J.D. Information in the holographic universe. Sci. Am. 2003, 289, 56–63. [Google Scholar] [CrossRef] [Scilit]
  5. Albert, D.Z. Quantum Mechanics and Experience; Harvard Up: Cambridge, MA, USA, 1992. [Google Scholar]
  6. Hardy, L. Quantum Theory From Five Reasonable Axioms. arXiv 2001, arXiv:quant-ph/0101012. [Google Scholar] [CrossRef] [Scilit]
  7. Moyal, J.E. Quantum mechanics as a statistical theory. Math. Proc. Camb. Philos. Soc. 1949, 45, 99–124. [Google Scholar] [CrossRef] [Scilit]
  8. Hall, B.C. Quantum Theory for Mathematicians; Number 267 in Graduate Texts in Mathematics; Springer: New York, NY, USA, 2013. [Google Scholar]
  9. Chentsov, N.N. Statistical Decision Rules and Optimal Inference; Translations of Mathematical Monographs—Translated from the Russian by the Israel Program for Scientific Translations; American Mathematical Society: Providence, RI, USA, 1982; Volume 53. [Google Scholar] [CrossRef] [Scilit]
  10. Burbea, J.; Rao, C.R. Entropy differential metric, distance and divergence measures in probability spaces: A unified approach. J. Multivar. Anal. 1982, 12, 575–596. [Google Scholar] [CrossRef] [Scilit]
  11. Fisher, R. On the mathematical foundations of theoretical statistics. Philos. Trans. R. Soc. Lond. Ser. A Contain. Pap. Math. Phys. Character 1922, 222, 309–368. [Google Scholar] [CrossRef] [Scilit]
  12. Rao, C. Information and Accuracy Attainable in Estimation of Statistical Parameters. Bull. Calcutta Math. Soc. 1945, 37, 81–91. [Google Scholar]
  13. Atkinson, C.; Mitchell, A.F.S. Rao’s Distance Measure. Sankhyā Indian J. Stat. Ser. A (1961–2002) 1981, 43, 345–365. [Google Scholar]
  14. Bernal-Casas, D.; Oller, J.M. Variational Information Principles to Unveil Physical Laws. Mathematics 2024, 12, 3941. [Google Scholar] [CrossRef] [Scilit]
  15. Heyer, H. Theory of Statistical Experiments; Springer: New York, NY, USA, 1983. [Google Scholar]
  16. Berger, J.O. Statistical Decision Theory and Bayesian Analysis; Springer: New York, NY, USA, 1985. [Google Scholar]
  17. Strasser, H. Mathematical Theory of Statistics: Statistical Experiments and Asymptotic Decision Theory; Walter de Gruyter: Berlin, Germany, 1985. [Google Scholar]
  18. Chavel, I. Eigenvalues in Riemannian Geometry; Elsevier: Philadelphia, PA, USA, 1984. [Google Scholar] [CrossRef] [Scilit]
  19. Shannon, C.E. A Mathematical Theory of Communication. Bell Syst. Tech. J. 1948, 27, 379–423+623–656. [Google Scholar] [CrossRef] [Scilit]
  20. Jaynes, E.T. Information Theory and Statistical Mechanics. Phys. Rev. 1957, 106, 620–630. [Google Scholar] [CrossRef] [Scilit]
  21. Jaynes, E.T. Information Theory and Statistical Mechanics II. Phys. Rev. 1957, 108, 171–190. [Google Scholar] [CrossRef] [Scilit]
Figure 1. A diagram illustrating the relationships among the key elements: the environment, the object, and the knowing subject.
Figure 1. A diagram illustrating the relationships among the key elements: the environment, the object, and the knowing subject.
Mathematics 14 00785 g001
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Bernal-Casas, D.; Oller, J.M. Information-Geometric Models in Data Analysis and Physics II. Mathematics 2026, 14, 785. https://doi.org/10.3390/math14050785

AMA Style

Bernal-Casas D, Oller JM. Information-Geometric Models in Data Analysis and Physics II. Mathematics. 2026; 14(5):785. https://doi.org/10.3390/math14050785

Chicago/Turabian Style

Bernal-Casas, D., and José M. Oller. 2026. "Information-Geometric Models in Data Analysis and Physics II" Mathematics 14, no. 5: 785. https://doi.org/10.3390/math14050785

APA Style

Bernal-Casas, D., & Oller, J. M. (2026). Information-Geometric Models in Data Analysis and Physics II. Mathematics, 14(5), 785. https://doi.org/10.3390/math14050785

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop