Next Article in Journal
Data-Driven Sliding-Mode Predictive Tracking Control for Networked Nonlinear Systems Under Random Deception Attacks: A Symmetry Perspective
Next Article in Special Issue
Scaling Symmetry in Symplectic Thermodynamics
Previous Article in Journal
Improving the Robustness of Scene-Aware Neuro-Symbolic Solving for Arithmetic Word Problems Under Input Perturbations
Previous Article in Special Issue
Lie Symmetries and Invariants of General Time-Dependent Quadratic Hamiltonian System
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

What Is Special About the Kirkwood–Dirac Distributions? Only They Produce Natural Conditional Expectations

by
Matéo Spriet
1,*,
Christopher Langrenez
1,
Raymond Brummelhuis
2 and
Stephan De Bièvre
1
1
Laboratoire Paul Painlevé, Univ. Lille, CNRS, Inria, UMR 8524, F-59000 Lille, France
2
Laboratoire Mathématique de Reims, Univ. Reims Champagne-Ardenne, CNRS, UMR 9008, F-51687 Reims, France
*
Author to whom correspondence should be addressed.
Symmetry 2026, 18(6), 1008; https://doi.org/10.3390/sym18061008
Submission received: 17 March 2026 / Revised: 28 May 2026 / Accepted: 30 May 2026 / Published: 11 June 2026

Abstract

Among the many quasiprobability representations of quantum mechanics, the family of Kirkwood–Dirac (KD) representations has come to the foreground in recent years. Each such KD representation is determined by the choice of two complementary complete sets of commuting observables A ^ and B ^ with respect to which it is Born-compatible, meaning that it correctly reproduces their Born probabilities for every state. In this paper, we identify what property uniquely characterizes the KD representations among all such A ^ and B ^ Born-compatible quasiprobability representations. For that purpose, we first define a natural notion of a quantum conditional expectation of an observable X ^ , given an observable Y ^ , in a state ρ ^ , as a best estimator, and we show that it has the basic properties generally expected of a conditional expectation. We then show that only the KD representations provide a notion of conditional, expectation given B ^ (or given A ^ ) that coincides with the above quantum conditional expectation. As a byproduct of our analysis, we show a state-dependent no-go theorem. We prove that, if the quantum conditional expectation of an observable X ^ , given an observable Y ^ in a state ρ ^ admits an anomalous value (meaning a value lying outside the interval [ x min , x max ] ), then there cannot exist a Born-compatible joint probability distribution μ ( x , y ) for X ^ and Y ^ in the state ρ ^ for which the associated conditional probability μ ( x | y ) yields a conditional expectation that coincides with the quantum conditional expectation. We further apply our findings to revisit a standard model for phase estimation in quantum metrology. We show in particular that, within the real sector of a given KD representation, the classical Fisher information of this phase estimation problem vanishes identically.

1. Introduction

Within quantum mechanics, it is not possible to properly define the joint probability distribution of two incompatible (i.e., non-commuting) observables in a coherent manner for all states [1,2,3,4]. This impossibility is one of the characteristic features of quantum mechanics, distinguishing it from classical mechanics. For the prototypical incompatible observables of position and momentum, Wigner famously found a way around this impossibility by associating to each quantum state a quasiprobability distribution for position and momentum, having the correct marginals, but that can take negative values [5]. A large variety of other quasiprobability distributions for position and momentum have subsequently been developed and have been proven useful in various fields of physics [6,7,8,9,10,11,12,13], even though not all these constructions yield joint quasiprobability distributions: their marginals do not always coincide with the Born-rule probabilities for position and momentum. Extensions of these constructions to abelian groups other than the space translations and their associated Heisenberg groups have also been developed [14,15,16,17,18,19,20,21,22,23,24,25]. Other approaches include coherent state constructions [26,27,28,29]. All these approaches fit in a general unifying setup using frames and dual frames on the quantum Hilbert space H . This setup allows one to define a quasiprobability distribution Q ( ρ ^ ) for all quantum states ρ ^ and symbols Q ˜ ( X ^ ) for all operators X ^ L ( H ) [3,30], both defined on a space Λ , thus forming a quasiprobability representation of quantum mechanics (see Section 4 for details). Here, L ( H ) is the set of linear operators on H ; we suppose throughout that dim H = d < + .
Among the countless quasiprobability representations of quantum mechanics so obtained feature the ones introduced by Dirac [8]. Suppose one is given two complete sets of commuting operators (CSCO) A ^ and B ^ . Dirac then proposed, for each quantum state ρ ^ , two joint quasiprobability distributions Q KD , ( ρ ^ ) and Q KD , r ( ρ ^ ) (both with the correct marginals, therefore) on the space σ ( A ^ ) × σ ( B ^ ) where σ ( A ^ ) and σ ( B ^ ) denote the spectra of A ^ and B ^ :
Q a , b KD , ( ρ ^ ) = Tr ( Π ^ a A ^ ρ ^ Π ^ b B ^ ) , Q a , b KD , r ( ρ ^ ) = Tr ( Π ^ b B ^ ρ ^ Π ^ a A ^ ) ,
where Π ^ a A ^ and Π ^ b B ^ are the one-dimensional spectral projectors of A ^ onto the eigenspace of a σ ( A ^ ) and of B ^ onto the eigenspace of b σ ( B ^ ) , respectively. Throughout, we will assume that A ^ and B ^ are complementary CSCO, meaning that Π ^ a A ^ Π ^ b B ^ 0 for all a , b . The reason for the choice of superscript , for “left” and r for “right” will become clear below. Q KD , and Q KD , r are today referred to as the Kirkwood–Dirac (KD) distributions [8]: they do indeed generalize the construction of Kirkwood [7] for the special case of position and momentum. The naturally associated KD symbols Q ˜ KD , ( X ^ ) and Q ˜ KD , r ( X ^ ) for any operator X ^ ,
Q ˜ a , b KD , ( X ^ ) = Tr ( Π ^ a A ^ X ^ Π ^ b B ^ ) Tr ( Π ^ a A ^ Π ^ b B ^ ) , Q ˜ a , b KD , r ( X ^ ) = Tr ( Π ^ b B ^ X ^ Π ^ a A ^ ) Tr ( Π ^ b B ^ Π ^ a A ^ ) ,
correspond to a choice of ordering between the operators A ^ and B ^ : A ^ before B ^ for and B ^ before A ^ for r . Indeed, the KD symbol Q ˜ KD , ( X ^ ) of the operator X ^ = f ( A ^ ) g ( B ^ ) is readily verified to be the function ( a , b ) f ( a ) g ( b ) , whereas the KD symbol Q ˜ KD , r ( X ^ ) of X ^ = g ( B ^ ) f ( A ^ ) is also ( a , b ) f ( a ) g ( b ) . (See Section 4 for details.) The flexibility of this construction, applying as it does to general non-commuting observables A ^ and B ^ , has proven to be an asset in an increasing number of applications in the last few years; we refer to the recent reviews [4,31] for details.
The question then arises: what singles out the Kirkwood–Dirac quasiprobability representations of quantum mechanics among all possible such representations? The central result of this paper is that they are, in a precise sense to be explained below (see Theorem 1), the only ones that are well-behaved with respect to a natural notion of conditional expectation in quantum mechanics that we introduce. Equivalently, the KD conditional expectation is the only one that can be naturally interpreted as a best estimator and that, as such, coincides with the conditional expectation proposed in weak value physics. We present several applications of these results. We establish a simple state-dependent no-go result for the existence of joint probabilities in quantum mechanics (Section 3.3). In Section 6.3, we show that in a standard model for phase estimation in quantum metrology, the Fisher information vanishes within the real sector of a given KD representation. This result provides a novel interpretation of the real sector of KD representations and allows us to show that the optimal measurement yielding the quantum Fisher information cannot be KD real (Proposition 7). We further explain to what extent KD-positive states (for which Q a , b KD ( ρ ^ ) 0 ), can be considered “classical” by exhibiting some of their nonclassical features.
In the rest of this introduction, we outline the paper. To precisely state and then establish our results, we first need to revisit the notions of conditional expectation used in both probability theory (Section 2) and quantum mechanics (Section 3 and Section 4). We shall use two distinct approaches to the definition of quantum conditional expectation. One approach consists in adapting to quantum mechanics the standard conditional expectation E P ( X | Y ) of a random variable X given a random variable Y, defined as the complex-valued function f ( Y ) of Y that minimizes the mean squared error
E P ( | X f ( Y ) | 2 ) : = Ω | X ( ω ) f ( Y ( ω ) ) | 2 d P ( ω ) ,
P being the underlying probability measure on the space Ω on which the random variables X and Y are defined. (See Section 2). In statistics, E P ( X | Y ) is referred to as the best predictor or best estimator of X among all complex-valued functions of Y. The idea that underlies this definition is that the law of E P ( X | Y ) provides some information about the law of X. For example, it is easily verified that their first moments agree:
E P ( E P ( X | Y ) ) = E P ( X ) .
This relation, known as the iterated expectation formula, expresses the fact that the best estimator E P ( X | Y ) of X is “unbiased”. In addition, one has the following well-known relation between their variances, referred to as the law of total variance or the variance decomposition formula:
Δ 2 X = Δ 2 E P ( X | Y ) + E P ( | X E P ( X | Y ) | 2 ) = Δ 2 E P ( X | Y ) + E P Δ P 2 ( X | Y ) .
This relation decomposes the total variability of X in a term due to the variability of the conditional expectation E P ( X | Y ) , plus an error term evaluating the fluctuations of X around this conditional expectation. This error term equals the expected value of the conditional variance Δ P 2 ( X | Y ) of X, given Y. In particular, the smaller is this mean squared error, the closer are their variances. In that case, the variability of X is considered to be well explained by that of E P ( X | Y ) .
In Section 3, we define a conditional expectation in the quantum mechanical context, using a similar minimization argument [32,33,34,35]. We will then show that identities (see Equations (68) and (69)) analogous to, but in important ways strikingly different from Equations (4) and (5) hold for this quantum conditional expectation. Let X ^ be an operator on a Hilbert space H and let Y ^ be a CSCO. Then, for any given mixed state ρ ^ (positive trace-class operator of trace 1), we consider the following expression,
Tr ( ρ ^ ( X ^ f ( Y ^ ) ) ( X ^ f ( Y ^ ) ) ) = Tr ( ( X ^ f ( Y ^ ) ) ρ ^ ( X ^ f ( Y ^ ) ) ) ,
where f is complex valued, and which is a natural analog of Equation (3). We then define, following [34,35], the conditional expectation E ρ ^ ( X ^ | Y ^ ) as that operator function f ( Y ^ ) of Y ^ which minimizes the expression in Equation (6) (see Definition 2). As in classical probability theory, where the conditional expectation depends on the underlying probability P , the quantum conditional probability E ρ ^ ( X ^ | Y ^ ) depends on the state ρ ^ . It can be computed explicitly:
E ρ ^ ( X ^ | Y ^ ) = y σ ( Y ^ ) φ y Y ^ , X ^ ρ ^ φ y Y ^ φ y Y ^ , ρ ^ φ y Y ^ Π ^ y Y ^ ,
where, for technical reasons and to ensure uniqueness of the solution, ρ ^ is supposed to belong to
D Y ^ = { ρ ^ D ( H ) : y σ ( Y ^ ) , Tr ( ρ ^ Π ^ y Y ^ ) 0 } .
Note that E ρ ^ ( X ^ | Y ^ ) is not necessarily self-adjoint, even if X ^ is. This is the first marked difference from what happens in probability theory: if X is a real random variable, then the conditional expectation E P ( X | Y ) is also real-valued. The physical meaning of the real and imaginary parts of E ρ ^ ( X ^ | Y ^ ) will be further discussed in Section 3 and Section 6.
The coefficients in Equation (7) are so-called “weak values”. They were introduced in [36], where they were given an operational meaning through an experimental procedure referred to as a weak measurement, that involves the weak coupling of the system to a meter whose position and/or momentum can be measured. For completeness and the reader’s convenience, we describe this setup and the relevant analysis in some detail in Appendix A. The nature of this experimental procedure has subsequently suggested the interpretation of the weak values as conditional expectations [37,38]. This course of events may leave one with the impression that the notion of conditional expectation in quantum mechanics depends on the experimental setup associated with weak values. Our results in Equation (7) show that this is not the case. On the contrary, weak values and their interpretation in terms of quantum conditional expectations arise naturally, in an experiment-independent manner, from the definition of conditional expectations via a quadratic minimization problem analogous to the one used to define conditional expectations in probability theory. In particular, the conditional expectation of X ^ given Y ^ can be viewed as the “best estimator” of X ^ by a function of Y ^ . Following this line of thought, the possibility of experimental determination of the coefficients of E ρ ^ ( X ^ | Y ^ ) in Equation (7) in a weak value experiments is then an a posteriori observation that provides an operational interpretation to the quantum conditional expectation.
Even if one does agree that defining conditional expectations via a minimization procedure is appealing and elegant, one may still wonder why one should choose to minimize a quadratic error. In order to give additional justification for this choice beyond the fact that it provides a natural analogue of the classical definition and the operational interpretation in terms of weak values just given, we provide, in Section 3.2, a characterization of the quantum conditional expectation in Equation (7): we show in Theorem 5 that, for ρ ^ D Y ^ , the map X ^ E ρ ^ ( X ^ | Y ^ ) is uniquely characterized by the following two properties, naturally associated with a conditional expectation and which also characterize conditional expectations in probability theory:
(i)
Pull-out formula: E ρ ^ ( f ( Y ^ ) X ^ | Y ^ ) = f ( Y ^ ) E ρ ^ ( X ^ | Y ^ ) , for all functions f ( Y ^ ) of Y ^ and for all X ^ L ( H ) ;
(ii)
Iterated expectations: E ρ ^ ( E ρ ^ ( X ^ | Y ^ ) ) = E ρ ^ ( X ^ ) .
In (ii), we wrote E ρ ^ ( X ^ ) : = Tr ( X ^ ρ ^ ) to stress the similarity with the analogous properties that hold for the probabilistic notion of conditional expectation, with ρ ^ playing the rôle of the probability measure P . Property (ii) expresses the fact that the best estimator E ρ ^ ( X ^ | Y ^ ) of X ^ is “unbiased” in the sense that it has the same expected value (in the state ρ ^ ) as X ^ itself. It is the analog of Equation (4) in probability theory.
Finally, in Proposition 2 we show a quantum version of the law of total variance
Δ ρ ^ / r , 2 X ^ = Δ ρ ^ / r , 2 E ρ ^ / r ( X ^ | Y ^ ) + E ρ ^ ( Δ ρ ^ / r , 2 ( X ^ | Y ^ ) ) .
Here, Δ ρ ^ / r , 2 ( X ^ | Y ^ ) is the conditional variance of X ^ , given Y ^ , defined in Equation (66). This equation is clearly analogous to Equation (5). It shows that the variance of X ^ equals the sum the variance of E ρ ^ ( X ^ | Y ^ ) and of the expected value of the conditional variance of X ^ , given Y ^ . We use these results to derive a sharpened version of an additive uncertainty principle first proposed in [37], see Section 6.2.
Having stressed the analogies between the quantum and classical (by which we mean here probabilistic) conditional expectations, we then identify a number of marked differences between them in Section 3.3 and explain their physical interpretation. We show in particular a simple but effective no-go theorem for the existence of joint probabilities. Consider ρ ^ , X ^ and Y ^ such that the quantum conditional expectation has a value E ρ ^ ( X ^ | Y ^ = y ) that falls outside the interval [ x min , x max ] (where x min , max are the extremal eigenvalues of X ^ ): such values are said to be anomalous. Then, there does not exist a joint probability distribution for X ^ and Y ^ with the right Born marginals in ρ ^ , and such that the corresponding conditional expectation for X ^ , given Y ^ , constructed with Bayes’ rule, equals the quantum conditional expectation E ρ ^ ( X ^ | Y ^ = y ) . We refer to Lemma 3 for a precise statement. The strength of this result lies in the fact that it puts a requirement only on a fixed triplet ( ρ ^ , X ^ , Y ^ ) to preclude the existence of a joint probability. It also gives a precise meaning to the statement that the presence of anomalous values of the quantum conditional expectation points to “nonclassicality”.
The superscript appearing in E ρ ^ ( X ^ | Y ^ ) stands for “left” because in property (i) above f ( Y ^ ) appears to the left of X ^ . In fact, although the above definition of a quantum conditional expectation as a best estimator is very natural; the choice of Equation (6) as the quantum equivalent to Equation (3) of the quantity to be minimized surreptitiously hides a choice of operator ordering. This should be expected, since the passage from the classical observable E P ( X | Y ) to a quantum equivalent involves a choice of quantization of observables, and as such can be expected to imply a choice of ordering. Indeed, an equally natural quantum analog of Equation (3) is
Tr ( ρ ^ ( X ^ f ( Y ^ ) ) ( X ^ f ( Y ^ ) ) ) = Tr ( ( X ^ f ( Y ^ ) ) ρ ^ ( X ^ f ( Y ^ ) ) ) .
Note that this quantity differs from Equation (6) only by the order in which ( X ^ f ( Y ^ ) ) and ( X ^ f ( Y ^ ) ) appear. The same minimization procedure as above then leads to a different conditional expectation, which we denote using E ρ ^ r ( X ^ | Y ^ ) . Again, the map X ^ E ρ ^ r ( X ^ | Y ^ ) is uniquely characterized by property (ii) above as well as by the following version of (i):
E ρ ^ r ( X ^ f ( Y ^ ) | Y ^ ) = f ( Y ^ ) E ρ ^ r ( X ^ | Y ^ ) , f ( Y ^ ) , X ^ L ( H ) ,
in which f ( Y ^ ) appears to the right of X ^ . The “left” and “right” conditional expectation are simply related to each other:
E ρ ^ r ( X ^ | Y ^ ) = E ρ ^ ( X ^ | Y ^ ) .
Contrary to the classical conditional expectation of a real random variable, which is real, the quantum conditional expectations E ρ ^ / r ( X ^ | Y ^ ) are not necessarily self-adjoint, even if X ^ is. When, on the other hand, either E ρ ^ ( X ^ | Y ^ ) or E ρ ^ r ( X ^ | Y ^ ) is self-adjoint, for some self-adjoint X ^ , we have
E ρ ^ r ( X ^ | Y ^ ) = E ρ ^ r ( X ^ | Y ^ ) = E ρ ^ ( X ^ | Y ^ ) = E ρ ^ ( X ^ | Y ^ ) .
In other words, in that case the choice of ordering plays no role and the “left” and “right” conditional expectations are identical.
There exists a second definition of conditional expectation that can be used in quantum mechanics, and that we revisit in Section 4. It starts from the observation that, in probability theory, the notion of conditional expectation E P ( X | Y ) of a random variable X, given a random variable Y, both defined on an underlying probability space Ω with probability P (both assumed to be discrete), can be defined in terms of their conditional probability distribution P ( X = x | Y = y ) , which in turn is defined in terms of their joint probability distribution P ( X = x , Y = y ) . Since, as pointed out above, such joint probabilities do not exist in quantum mechanics for non-commuting observables, this definition cannot straightforwardly be adapted to the quantum context. One approach to resolve this issue consists in replacing probabilities by quasiprobabilities and then proceeding in analogy with the classical treatment [38,39]. As we explain in Section 4, this approach leads to a notion of conditional expectation E ρ ^ Q ( X ^ | Y ^ ) as an operator that depends not only on ρ ^ , but also on the quasiprobability representation used in its definition. In addition, as we will show, the conditional expectation E ρ ^ Q ( X ^ | Y ^ ) so defined does not generally have all the natural properties one would expect from a conditional expectation. In particular, whereas it does always satisfy the law of iterated expectations (property (ii) above), it does not in general satisfy the pull-out property (property (i) above). This is not surprising since, as we saw, the pull-out property together with the law of iterated expectations does uniquely fix the definition of conditional expectation (Theorem 5).
After these preparatory developments, we turn in Section 5.2 and Section 5.3 to our principal result, which is the unique characterization of the KD quasiprobability representations, introduced in Section 5.1. We consider the set of all quasiprobability representations ( Q , Q ˜ ) , defined on some finite set Λ with d 2 elements, that are Born-compatible with two given CSCO A ^ and B ^ . This means that the marginals of the quasiprobability distribution Q ( ρ ^ ) of any state ρ ^ with respect to the symbols Q ˜ ( A ^ ) and Q ˜ ( B ^ ) yield the correct quantum mechanical Born probabilities for these two observables (see Definition 3 for the precise definition). Depending on the goal pursued, the set Λ is sometimes referred to as the “ontic space” or simply as the “phase space” of the quasiprobability representation. Our central result concerning the KD representations of quantum mechanics is then summed up in the following theorem:
Theorem 1.
Let A ^ and B ^ be complementary CSCO . Let ( Q , Q ˜ ) be an A ^ and B ^ -compatible quasiprobability representation of quantum mechanics defined on a set Λ ( Λ = d 2 ) . Then the following are equivalent:
(i) 
ρ ^ D B ^ , X ^ L ( H ) , E ρ ^ Q ( X ^ | B ^ ) = E ρ ^ ( X ^ | B ^ )
(ii) 
ρ ^ D A ^ , X ^ L ( H ) , E ρ ^ Q ( X ^ | A ^ ) = E ρ ^ r ( X ^ | A ^ )
(iii) 
There exists a bijective map Φ : Λ σ ( A ^ ) × σ ( B ^ ) such that for all ( a , b ) σ ( A ^ ) × σ ( B ^ )
ρ ^ D ( H ) , Q Φ 1 ( a , b ) ( ρ ^ ) = Q a , b KD , ( ρ ^ ) and X ^ L ( H ) , Q ˜ Φ 1 ( a , b ) ( X ^ ) = Q ˜ a , b KD , ( X ^ ) .
The theorem can be paraphrased by saying that of all A ^ and B ^ Born-compatible quasiprobability representations of quantum mechanics, only the left KD representation has conditional expectations that coincide with the best estimators E ρ ^ ( X ^ | B ^ ) and E ρ ^ r ( X ^ | A ^ ) . Theorem 1 is a generalization of a result presented by the authors in [40]. In that work, the class of quasiprobability representations among which the KD representation is proven to be unique was considerably restricted by the a priori assumption that Λ = σ ( A ^ ) × σ ( B ^ ) . When using quasiprobability representations in the study of foundational issues such as contextuality and hidden variables, this is not a natural assumption, and thus it was lifted here.
We will provide two independent proofs of this result. One uses the characterization of E ( X ^ | Y ^ ) via the pull-out property (Theorem 7). The other shows it follows from a slightly stronger statement (Theorem 9) that is itself based on Theorem 8 which is of interest in its own right: it shows that, quite generally, A ^ and B ^ Born-compatible quasiprobability representations of quantum mechanics are always uniquely determined by their conditional expectations given A ^ or given B ^ . The proof of Theorem 8 uses the techniques developed here with an argument found in [41], where a partial result similar to Theorem 9, but limited to the case where Λ = σ ^ ( A ^ ) × σ ^ ( B ^ ) , was recently announced and partially proven.
As we saw, one of the striking differences between classical and quantum conditional expectations is that the latter can be non-self-adjoint. In addition, even if they are self-adjoint, they can take classically forbidden values (Lemma 3) and, in this sense, still signal typical quantum behavior such as, for example, weak value amplification or quantum advantage in quantum metrology [31,42,43,44]. In Section 6, we further analyze the meaning of the imaginary part of the quantum conditional expectation values and more specifically of its vanishing. For that purpose, we first recall the role played by the imaginary part of the quantum conditional expectation in a well-known problem of phase estimation in quantum mechanics. We consider a one-parameter family of states
ρ ^ X ^ ( θ ) = exp ( i θ X ^ ) ρ ^ exp ( i θ X ^ ) ,
where θ R is referred to as the phase and X ^ is self-adjoint. The outcome probabilities of repeated measurements of a CSCO Y ^ in ρ ^ X ^ ( θ ) are given by p ( y ; θ ) = Tr ( Π ^ y Y ^ ρ ^ X ^ ( θ ) ) ( y σ ( Y ^ ) ), and we write I F ( Y ^ ; θ ) for the Fisher information of p ( y ; θ ) . We first then recall the well-known relation [37,45] between the above Fisher information I F ( y ; θ ) and the variance of the imaginary part of E ρ ^ X ^ ( θ ) / r ( X ^ | Y ^ ) :
I F ( Y ^ ; θ ) = 4 Tr Im E ρ ^ X ^ ( θ ) / r ( X ^ | Y ^ ) 2 ρ ^ X ^ ( θ ) .
Since it is well known that, in order to be able to obtain a good estimate on θ , one needs a large value of the Fisher information, the imaginary part of the conditional expectation E ρ ^ X ^ ( θ ) / r ( X ^ | Y ^ ) contains information on the phase θ . For that reason, we will say that a state is phase insensitive (for a given choice of X ^ and of Y ^ ) if its Fisher information I F ( Y ^ ; 0 ) vanishes. To further justify this terminology, we give an operational interpretation of the vanishing of Fisher information in Appendix D. As pointed out above, in the context of the phase estimation problem considered, vanishing of the associated Fisher information is equivalent to saying that the conditional expectation E ρ ^ / r ( X ^ | Y ^ ) is self-adjoint. When this is the case, the probability distribution p ( y ; θ ) cannot be used to efficiently estimate θ . Using the quantum law of total variance (Equation (9)), we sharpen an additive uncertainty relation first established in [37] between the Fisher information I F ( Y ^ , 0 ) and the variance of the real part of the conditional expectation, denoted by E ρ ^ sa ( X ^ | Y ^ ) . We show
min Y ^ Δ ρ ^ 2 E ρ ^ sa ( X ^ | Y ^ ) Δ ρ ^ 2 E ρ ^ sa ( X ^ | L ^ ) Δ ρ ^ 2 X ^ 1 4 I QF ( 0 ) Δ ρ ^ 2 X ^ = max Y ^ Δ ρ ^ 2 E ρ ^ sa ( X ^ | Y ^ ) .
Here, I QF ( 0 ) is the quantum Fisher information of ρ ^ and X ^ and L ^ is the symmetric logarithmic derivative of ρ ^ X ^ ( θ ) . (See Appendix E for an introduction to the quantum Fisher information.) When ρ ^ is pure, one more precisely has
min Y ^ Δ ρ ^ 2 E ρ ^ sa ( X ^ | Y ^ ) = Δ ρ ^ 2 E ρ ^ sa ( X ^ | L ^ ) = Δ ρ ^ 2 X ^ 1 4 I QF ( 0 ) = 0 , Δ ρ ^ 2 X ^ = max Y ^ Δ ρ ^ 2 E ρ ^ sa ( X ^ | Y ^ ) ,
where the maximum on the right is reached when ρ ^ is phase insensitive for the given choice of X ^ and Y ^ . In other words, when Y ^ is chosen close to L ^ , then the Fisher information I F ( Y ^ ; 0 ) is close to its maximal value I QF and the variance of E ρ ^ sa ( X ^ | Y ^ ) is close to its minimal value and vice versa.
In light of this analysis, it is of interest to know for which triplets ρ ^ , X ^ , Y ^ , the Fisher information I F ( Y ^ ; 0 ) vanishes; we call such triplets phase insensitive. For that purpose, we place ourselves in the framework of KD representations and exploit Theorem 1. We show that, if ρ ^ and X ^ have a real KD symbol with respect to some CSCO A ^ and B ^ = Y ^ , then I F ( Y ^ ; 0 ) vanishes (Section 6.3); hence, ( ρ ^ , X ^ , B ^ ) form a phase-insensitive triplet. This observation provides an interpretation of the “real sector” of a given KD representation and shows it is inefficient for the above phase estimation problem. Since recent work has shown that the set of all self-adjoint operators with a real KD symbol can in many cases be explicitly identified [23,24,46], these results also provide many examples of phase insensitive states. Finally, supposing that ρ ^ and X ^ are KD real, we investigate their associated quantum Fisher information. We show that the latter is non-vanishing if and only if ρ ^ and X ^ do not commute. It is well known that it is obtained by optimizing the choice of observable Y ^ , taking it to be equal to the symmetric logarithmic derivative L ^ of ρ ^ X ^ ( 0 ) . We then show that L ^ cannot be KD real when ρ ^ is pure (Proposition 7). In other words, if ρ ^ and X ^ are KD-real, then the optimal observable L ^ is not KD-real.
It should be noted that a triplet ( ρ ^ , X ^ , B ^ ) can be phase-insensitive, with ρ ^ KD positive, while the conditional expectation E ρ ^ ( X ^ | B ^ ) still allows for anomalous values. This shows once again that the word “classical” needs to be given a precise meaning when used.
We have completed this paper with a number of appendices with the goal to make it reasonably self-contained and to provide, for the reader’s convenience, proofs of important properties not necessarily easily accessible in the literature. In Appendix A, we provide a short but rather complete introduction to weak value theory and in particular to weak value measurement. We have included a discussion of the conditional expectations and variances of meter positions in the weak limit which connects with recent work [47]. In Appendix B, we briefly explain the link between our definition of conditional expectations and a definition used in the context of von Neumann algebras, restricted to our finite dimensional setup. Appendix C contains illustrative examples and counterexamples. Appendix D gives operational interpretations of the vanishing of the Fisher information for parameter estimation. Appendix E provides an introduction to the quantum Fisher information. Using the ideas developed in the paper, we give a simple proof of the result of Braunstein and Caves showing that the Helström–Holevo quantum Fisher information is obtained as the maximum over all possible measurements of the classical Fisher information associated to those measurements. In Appendix F, we provide a variational characterization of the KD distribution of a state.
It is a pleasure to dedicate this work to Jean-Pierre Gazeau on the occasion of his 80th birthday. Jean-Pierre has been for decades an efficient and enthusiastic advocate of phase space representations of quantum mechanics, coherent states and frames, and their applications in quantum theory and signal analysis. We hope he will find our contribution of interest. SDB is particularly grateful to Jean-Pierre for having welcomed him in the French scientific world over 35 years ago and for his continued support and friendship throughout all these years.

2. Classical Conditional Expectation and Conditional Variance

In this section, we collect some basic definitions and properties of the classical conditional expectation as used in probability theory. This allows us to fix notation and prepare for the quantum formulation.
Let Ω be a finite set. Let X : Ω C be a complex random variable, and let Y : Ω R m be a real vector-valued random variable. We write Ran ( X ) for the range of X, meaning
Ran ( X ) = { x C ω Ω , X ( ω ) = x } = X ( Ω ) ,
and similarly for Y. We use D Y to denote the set of probability measures P on Ω such that for each y Ran ( Y ) , P ( Y = y ) > 0 . The space of complex random variables on Ω is denoted by F ( Ω ; C ) . We first recall the elementary definition of “conditional expectation” of a random variable.
Definition 1.
Let P D Y . For y Ran ( Y ) and x C , the conditional probability that X takes the value x knowing Y = y is given by
P ( X = x | Y = y ) = P ( X = x , Y = y ) P ( Y = y ) .
The conditional expectation of X knowing that Y = y is
E P ( X | Y = y ) = x Ran ( X ) x P ( X = x | Y = y ) = x Ran ( X ) x P ( X = x , Y = y ) P ( Y = y ) .
We will, as usual, define the conditional expectation of X, knowing Y, denoted by E P ( X | Y ) , as the function
E P ( X | Y ) ( y ) = E P ( X | Y = y ) .
Note that E P ( X | Y ) is a random variable that is a function of Y. Also,
E P ( f ( Y ) | Y ) = f ( Y ) and E P E P ( X | Y ) = E P ( X ) ,
as is readily verified. Finally, if X is a real random variable, then E P ( X | Y ) is also real-valued. As we will see, this is a marked difference from what happens in quantum mechanics, where the conditional expectation of an observable can be non self-adjoint.
We note that this definition, and all the properties that follow in this section, are strongly dependent on the initially chosen probability P on Ω . We indicate this dependence in the notation to anticipate the quantum conditional probability defined in the next section, which also depends on the ρ ^ state considered.
We now provide two distinct characterizations of E P ( X | Y ) that will be essential in the following sections in order to define a quantum conditional expectation. For P D Y , we equip F ( Ω ; C ) with the following sesquilinear form:
X , X P = E P ( X ¯ X ) = Ω X ¯ ( ω ) X ( ω ) d P = ω Ω X ¯ ( ω ) X ( ω ) P ( { ω } ) .
Theorem 2.
Let P D Y . The conditional expectation of X knowing Y is the unique minimizer of
f E P ( | X f ( Y ) | 2 ) = X f ( Y ) , X f ( Y ) P ,
over the set of complex valued functions f defined on Ran ( Y ) .
We use F C , Y to denote the set
F C , Y = f ( Y ) f : Ran ( Y ) C .
F C , Y is thus the set of all complex random variables on Ω that are functions of Y. In other words, Z : Ω C belongs to F C , Y provided that Z is constant on all level sets of Y. This means that there exists a function f : Ran ( Y ) C such that Z = f ( Y ) . Note that F C , Y is a complex vector space ( F C , Y F ( Ω ; C ) ); it is in addition closed under multiplication of functions. In the usual language of Hilbert space theory, the minimizer of Equation (24)—which is the conditional expectation of X given Y—is the orthogonal projection of X onto F C , Y . In more advanced treatments of probability, where the space Ω is not finite and where the probability measures are general, this property provides a simple way to define the conditional expectation, since Definition 1 then does not necessarily make sense. We will see in the next section that this definition adapts nicely to quantum theory as well.
Proof. 
For f : Ran ( Y ) C , we have
f ( y ) = y Ran ( Y ) f ( y ) 1 y ( y )
where 1 y ( y ) is 1 if y = y and 0 otherwise. One then computes
E P ( | X f ( Y ) | 2 ) = E P ( | X | 2 ) 2 y Ran ( Y ) Re f ( y ) ¯ E P ( X 1 y ( Y ) ) + y Ran ( Y ) | f ( y ) | 2 P ( Y = y ) = E P ( | X | 2 ) + y Ran ( Y ) P ( Y = y ) | f ( y ) | 2 2 Re f ( y ) ¯ E P ( X 1 y ( Y ) ) P ( Y = y ) = E P ( | X | 2 ) y Ran ( Y ) E P ( X 1 y ( Y ) ) 2 P ( Y = y ) + y Ran ( Y ) P ( Y = y ) f ( y ) E P ( X 1 y ( Y ) ) P ( Y = y ) 2 .
This quantity is minimal if and only if for all y Ran ( Y ) ,
f ( y ) = E P ( X 1 y ( Y ) ) P ( Y = y ) .
Moreover, we have that
E P ( X 1 y ( Y ) ) = ( x , y ) Ran ( X ) × Ran ( Y ) P ( X = x , Y = y ) x 1 y ( y ) = x Ran ( X ) x P ( X = x , Y = y ) = P ( Y = y ) E P ( X | Y = y ) .
This concludes the proof. □
From Equation (22), we recall the law of iterated expectations,
E P E P ( X | Y ) = E P X
and we note, for later reference, that Equation (27) can be rewritten in the form
Δ 2 X = Δ 2 E P ( X | Y ) + E P ( | X E P ( X | Y ) | 2 ) ,
which is the law of total variance, where
Δ 2 Z : = E P ( | Z E P ( Z ) | 2 )
is the variance of the random variable Z : Ω C . In other words, X and its best estimator E ( X | Y ) have the same expected value and their variances differ by the minimal “cost”. Introducing the conditional variance
Δ P 2 ( X | Y ) = E P ( | X E P ( X | Y ) | 2 | Y ) ,
the law of total variance reads alternatively as
Δ 2 X = Δ 2 E P ( X | Y ) + E P Δ P 2 ( X | Y ) .
We note for later reference that
Δ P 2 ( X | Y ) = E P ( | X | 2 ) | E P ( X | Y ) | 2 .
As we shall now prove, the conditional expectation is also characterized by two simple properties:
Theorem 3.
Let Y be a real vector-valued random variable on Ω and P D Y . Then, there exists a unique linear map E cond , Y : F ( Ω ; C ) F C , Y F ( Ω ; C ) that satisfies the following properties:
1. 
Pull-out formula. For any complex random variable X on Ω and any f : Ran ( Y ) C ,
E cond , Y ( f ( Y ) X ) = f ( Y ) E cond , Y ( X ) ;
2. 
The law of iterated expectations. For any complex random variable X on Ω,
E P ( E cond , Y ( X ) ) = E P ( X ) .
One has E cond , Y ( X ) = E P ( X | Y ) .
Proof. 
Its existence is clear, since it is easily verified that X E P ( X | Y ) satisfies the two desired properties. Uniqueness remains to be shown. Let E cond , Y : X E cond , Y ( X ) F C , Y satisfy properties 1 and 2. Let X be a random variable. For any f ( Y ) F C , Y , we have
f ( Y ) , X E cond , Y ( X ) P = E P ( f ( Y ) ¯ ( X E cond , Y ( X ) ) ) = E P ( f ( Y ) ¯ X ) E P ( E cond , Y ( f ( Y ) ¯ X ) ) by 1 = E P ( f ( Y ) ¯ X ) E P ( f ( Y ) ¯ X ) by 2 = 0 .
We therefore also have f ( Y ) , X E P ( X | Y ) P = 0 . It follows that
E cond , Y ( X ) E P ( X | Y ) , E cond , Y ( X ) E P ( X | Y ) P = E cond , Y ( X ) X + X E P ( X | Y ) , E cond , Y ( X ) E P ( X | Y ) P = X E cond , Y ( X ) , E cond , Y ( X ) E P ( X | Y ) P + X E P ( X | Y ) , E cond , Y ( X ) E P ( X | Y ) P = 0
since E cond , Y ( X ) E P ( X | Y ) F C , Y . Therefore, E P ( X | Y ) = E cond , Y ( X ) almost everywhere with respect to P . Let ω Ω , then Y ( ω ) Ran ( Y ) and thus, since P D Y , there exists ω Ω such that P ( { ω } ) > 0 and Y ( ω ) = Y ( ω ) . The fact that E P ( X | Y ) = E cond , Y ( X ) almost everywhere with respect to P then implies E P ( X | Y ) ( ω ) = E cond , Y ( X ) ( ω ) . Moreover, since they are both functions of Y we have E P ( X | Y ) ( ω ) = E P ( X | Y ) ( ω ) and E cond , Y ( X ) ( ω ) = E cond , Y ( X ) ( ω ) . In the end, we have indeed E P ( X | Y ) = E cond , Y ( X ) on Ω . This concludes the proof. □
Let us point out that one can see quite directly that Equations (35) and (36) imply that I E cond , Y ( I ) is orthogonal to the algebra F C , Y , as well as belonging to F C , Y . This implies
E cond , Y ( I ) = I ,
where I is the constant function equal to 1 on Ω . Using Equation (35) again, we obtain that
E cond , Y ( f ( Y ) ) = f ( Y ) .
These identities can also be retrieved from E cond , Y ( X ) = E P ( X | Y ) .
We finish this section by showing that one can recover the joint probability distribution of the random variables X and Y using the expectation and conditional expectation. We will discuss a similar relation between joint quasiprobabilities and the conditional expectation in the quantum realm in Proposition 3.
Proposition 1.
For any ( x , y ) Ran ( X ) × Ran ( Y ) , we have
P ( X = x , Y = y ) = E P [ 1 y ( Y ) E P ( 1 x ( X ) | Y ) ] .
Proof. 
We compute
E P [ 1 y ( Y ) E P ( 1 x ( X ) | Y ) ] = y Ran ( Y ) P ( Y = y ) 1 y ( y ) E P ( 1 x ( X ) | Y ) ( y ) = P ( Y = y ) E P ( 1 x ( X ) | Y ) ( y ) = P ( Y = y ) x Ran ( X ) 1 x ( x ) P ( X = x , y = y ) P ( Y = y ) = P ( X = x , Y = y ) .
We finally point out that Y can also be taken complex vector-valued in what precedes. But, to stress the analogy with the quantum mechanical treatment, we have taken Y real vector-valued here.

3. Quantum Conditional Expectation and Variance via Minimization

In Section 3.1, we show how to adapt the classical definition of a conditional expectation based on Theorem 2 to the quantum context in order to define a quantum conditional expectation: see Theorem 4. This procedure defines the quantum conditional expectation as a “best estimator” of the operator X ^ by functions of the observable Y ^ . Such definition based on a minimization procedure in analogy with the classical definition clearly depends on the quantity (the cost function) that is chosen to be minimized. We will, as in the classical context, use a quadratic minimizer and see that its choice is closely linked to the familiar “ordering problem” in quantization. Indeed, whereas for random variables X and Y, one has f ( Y ) X = X f ( Y ) , this is no longer generally true when X ^ and Y ^ are operators on a Hilbert space. This observation naturally leads to two conditional expectations, a “left” one and a “right” one, among many other possibilities. In Section 3.2, we provide in Theorem 5 a characterization of the “left” and “right” quantum conditional expectations introduced that is analogous to its classical counterpart, Theorem 3 in that it relies on a “pull-out property”. In Section 3.3, we introduce the conditional variance of X ^ given Y ^ , and we prove relations analogous to Equations (29) and (30), relating the mean and variance of the operator X ^ to those of its quantum conditional expectation and variance. We also discuss some of the marked differences between the classical and quantum notions, in particular with regard to pure states. Finally, in Section 3.4, we show that the self-adjoint part of the quantum conditional expectation can itself be interpreted as a best estimator of X ^ by self-adjoint functions of Y ^ .

3.1. The Quantum Conditional Expectation: Definition

Let H be a Hilbert space of dimension d < + . Let L ( H ) be the space of operators on H , and D ( H ) the set of density matrices (positive semi-definite operators of trace 1). We will first adapt the classical construction of the conditional expectation of Theorem 2 to the quantum context. For that purpose, we need to identify a suitable quantum analog of the sesquilinear form between classical observables in Equation (23). One may expect such analogs not to be uniquely determined since the sesquilinear form involves the product of two classical observables, and in general, quantizations of observables do not respect products: the quantization of the product is usually not the product of the quantizations due to ordering problems. We will see this phenomenon manifests itself here as well. One choice that presents itself naturally is to define, for any two X ^ , X ^ L ( H ) , the sesquilinear form
X ^ , X ^ : = X ^ , X ^ ρ ^ , : = Tr ( ρ ^ X ^ X ^ ) = Tr ( X ^ ρ ^ X ^ ) ,
where ρ ^ D ( H ) is a fixed quantum state, which we will mostly drop from the notations (except when we will consider varying states in Section 6 below). Now, let Y ^ = ( Y ^ 1 , Y ^ m ) be a set of commuting observables on H . We define the joint spectrum of Y ^ as
σ ( Y ^ ) = σ ( Y ^ 1 ) × × σ ( Y ^ m ) ,
where σ ( Y ^ j ) is the spectrum of Y ^ j . We will say Y ^ is complete when, for each y = ( y 1 , , y m ) σ ( Y ^ ) , there exists a unique (up to an irrelevant global phase) normalized common eigenvector φ y Y ^ , for all Y ^ j :
Y ^ j φ y Y ^ = y j φ y Y ^ .
In other words, Y ^ is complete if and only if its joint spectrum is non-degenerate. Such a family is called a complete set of commuting observables (CSCO). While demanding that Y ^ is a family of commuting observables rather than an observable itself may seem superfluous, it is in fact of importance in many cases. A relevant example is found in spin systems, where H = C 2 m , and where the Pauli operators are given by Y ^ j σ ^ = I I σ ^ I I . Here, σ ^ is a 2 × 2 Pauli matrix, acting on the j-th sector.
In what follows, we shall write Π ^ y Y ^ for the orthogonal projector onto φ y Y ^ . We consider the algebra of all functions of Y ^ :
F C , Y ^ = { f ( Y ^ ) | f : σ ( Y ^ ) C } .
Note that F C , Y ^ is a commutative *-algebra. We use F C , Y ^ to denote the commutant of F C , Y ^ . F C , Y ^ is maximal in the following sense:
Lemma 1.
F C , Y ^ = F C , Y ^ .
Proof. 
It is clear that F C , Y ^ F C , Y ^ . Suppose therefore C ^ F C , Y ^ . Then, [ C ^ , Π ^ y Y ^ ] = 0 , for all y σ ( Y ^ ) . Hence, for all y σ ( Y ^ ) ,
C ^ φ y Y ^ = φ y Y ^ Tr ( C ^ Π ^ y Y ^ ) .
Defining the function g ( y ) = Tr ( C ^ Π ^ y Y ^ ) , one concludes that C ^ = g ( Y ^ ) and hence that C ^ F C , Y ^ . □
For a density matrix ρ ^ , analogous to Theorem 2, one can then define (following [32,33,34,35] and generalized in [48,49]) a quantum conditional expectation of X ^ , knowing Y ^ , that we shall denote using E ρ ^ ( X ^ | Y ^ ) as the minimizer of the following quantity:
f Tr ( ρ ^ | X ^ f ( Y ^ ) | 2 ) = X ^ f ( Y ^ ) , X ^ f ( Y ^ ) ,
over all f : σ ( Y ^ ) C . Here, for Z L ( H ) , we write | Z | = Z Z . As in the classical case, we think of the right-hand side of this equation as a mean squared error cost. We shall discuss the existence and uniqueness of this minimizer below. We shall refer to E ρ ^ ( X ^ | Y ^ ) as the left quantum conditional expectation of X ^ knowing Y ^ in the state ρ ^ , for reasons that will become clear shortly. There is a second, equally natural choice for the sesquilinear product, analogous to the inner product in Equation (23), namely
X ^ , X ^ r = Tr ( ρ ^ X ^ X ^ ) = Tr ( X ^ ρ ^ X ^ ) .
Note that, as anticipated, the difference between Equations (40) and (45) lies indeed in the change of order of appearance of X ^ and X ^ . We then define the right quantum conditional expectation of X ^ , knowing Y ^ , that we shall denote using E ρ ^ r ( X ^ | Y ^ ) , as the minimizer over complex-valued functions of
f X ^ f ( Y ^ ) , X ^ f ( Y ^ ) r .
The existence and uniqueness of these conditional expectations is guaranteed by the following theorem, under a suitable hypothesis on ρ ^ . We define
D Y ^ = { ρ ^ D ( H ) : y σ ( Y ^ ) , Tr ( ρ ^ Π ^ y Y ^ ) 0 } .
Theorem 4.
Assume that ρ ^ D Y ^ . Then, the functions given in Equations (44) and (46) admit unique minimizers given by
f 0 ( y ) = φ y Y ^ , X ^ ρ ^ φ y Y ^ φ y Y ^ , ρ ^ φ y Y ^ , f 0 r ( y ) = φ y Y ^ , ρ ^ X ^ φ y Y ^ φ y Y ^ , ρ ^ φ y Y ^ ,
respectively, for all y σ ( Y ^ ) .
Proof. 
We only prove the “left” case, the “right” case being similar. We compute
X ^ f ( Y ^ ) , X ^ f ( Y ^ ) = Tr ( ρ ^ | X ^ | 2 ) Tr ( ρ ^ X ^ f ( Y ^ ) ) Tr ( ρ ^ f ( Y ^ ) X ^ ) + Tr ( ρ ^ | f ( Y ^ ) | 2 ) = Tr ( ρ ^ | X ^ | 2 ) + y σ ( Y ^ ) φ y Y ^ , ρ ^ φ y Y ^ | f ( y ) | 2 2 Re f ( y ) ¯ φ y Y ^ , X ^ ρ ^ φ y Y ^ φ y Y ^ , ρ ^ φ y Y ^ = Tr ( ρ ^ | X ^ | 2 ) y σ ( Y ^ ) | φ y Y ^ , X ^ ρ ^ φ y Y ^ | 2 φ y Y ^ , ρ ^ φ y Y ^ + y σ ( Y ^ ) φ y Y ^ , ρ ^ φ y Y ^ f ( y ) φ y Y ^ , X ^ ρ ^ φ y Y ^ φ y Y ^ , ρ ^ φ y Y ^ 2 .
Hence, this quantity is minimal if and only if
f ( y ) = φ y Y ^ , X ^ ρ ^ φ y Y ^ φ y Y ^ , ρ ^ φ y Y ^
for all y σ ( Y ^ ) . □
Definition 2.
We define the left quantum conditional expectation of the operator X ^ knowing Y ^ in the state ρ ^ D Y ^ as
E ρ ^ ( X ^ | Y ^ ) = f 0 ( Y ^ ) = y σ ( Y ^ ) φ y Y ^ , X ^ ρ ^ φ y Y ^ φ y Y ^ , ρ ^ φ y Y ^ Π ^ y Y ^ ,
where Π ^ y is the orthogonal projector to the eigenvector φ y Y ^ for y σ ( Y ^ ) . Similarly, we define the right quantum conditional expectation of X ^ knowing Y ^ in the state ρ ^ as
E ρ ^ r ( X ^ | Y ^ ) = f 0 r ( Y ^ ) = y σ ( Y ^ ) φ y Y ^ , ρ ^ X ^ φ y Y ^ φ y Y ^ , ρ ^ φ y Y ^ Π ^ y Y ^ .
We will write E ρ ^ / r ( X ^ | Y ^ = y ) = Tr E ρ ^ / r ( X ^ | Y ^ ) Π ^ y Y ^ .
It is readily verified that for any f : σ ( Y ^ ) C ,
E ρ ^ / r ( f ( Y ^ ) | Y ^ ) = f ( Y ^ ) .
Remark 1. 
( i ) If the state ρ ^ is not an element of D Y ^ , i.e., if there exists y σ ( Y ^ ) such that φ y Y ^ , ρ ^ φ y Y ^ = 0 , then the functional (44) still admits minimizers, but they are no longer unique. Since φ y Y ^ , ρ ^ φ y Y ^ = 0 implies that ρ ^ 1 / 2 φ y Y ^ = 0 , and therefore ρ ^ φ y Y ^ = 0 , the sums in (49) only extend over y’s with φ y Y ^ , ρ ^ φ y Y ^ 0 and the minimizers are in fact given by
f 0 ( Y ^ ) = y σ ( Y ^ ) φ y Y ^ , ρ ^ φ y Y ^ 0 φ y Y ^ , X ^ ρ ^ φ y Y ^ φ y Y ^ , ρ ^ φ y Y ^ Π ^ y Y ^ + y σ ( Y ^ ) φ y Y ^ , ρ ^ φ y Y ^ = 0 g y Π ^ y Y ^ ,
where ( g y ) φ y Y ^ , ρ ^ φ y Y ^ = 0 C can be any complex numbers [33]. In the usual interpretation of quantum mechanics, the quantity φ y Y ^ , ρ ^ φ y Y ^ represents the probability of the event “ Y ^ takes the value y”. It follows in accordance with classical probability theory that the minimizers of the functional (44) are unique up to functions of Y ^ which are “ ρ ^ -almost everywhere 0”, such functions corresponding to the term y σ ( Y ^ ) φ y Y ^ , ρ ^ φ y Y ^ = 0 g y Π ^ y Y ^ . In this article, we only consider non-degenerate states, except for in Appendix E, where we discuss the connection with quantum Fisher information. Note that here by non-degenerate we mean states ρ ^ D Y ^ , which does not necessarily imply that the state ρ ^ is positive definite. For example the set D Y ^ contains pure states. This non-degeneracy condition allows nice characterizations of the conditional expectations. Moreover, it is not too restrictive, as shown by Lemma 2.
( i i ) The expressions · , · / r are not inner products because X ^ , X ^ / r may vanish even if X ^ 0 . This will happen if ρ ^ has a non-vanishing kernel, foremost when the state ρ ^ is pure. One may however think of them as degenerate inner products or as quadratic cost functions. Note that this cost depends on the state ρ ^ . From this point of view, one may think of the operator E ρ ^ / r ( X ^ | Y ^ ) as the operator function of the observables Y ^ that approximates X ^ at minimal cost. It is also clear that such minimizers necessarily depend on the choice of cost function, that in turn necessarily depends on the physics of the problem considered. In the next subsection, we show the “left” and “right” choices made here are natural in the sense that they can be uniquely characterized by a pull-out property analogous to the classical conditional expectation.
( i i i ) In view of what precedes, the quantum conditional expectations E , r are orthogonal projections with respect to the two “inner products” induced by the state ρ ^ of the operator X ^ onto the algebra of functions of the operators Y ^ .
( i v ) The conditional expectations E ρ ^ / r ( X ^ | Y ^ ) are not self-adjoint in general, even if we take X ^ to be self-adjoint. This is markedly different from the classical situation, where the conditional expectation of a real random variable X, given a real random variable Y, is always real. We will extensively elaborate on this point below. Note that
E ρ ^ r ( X ^ | Y ^ ) = E ρ ^ ( X ^ | Y ^ ) .
( v ) When X ^ is self-adjoint, the coefficients
φ y , X ^ ρ ^ φ y φ y , ρ ^ φ y y σ ( Y ^ )
appearing in E ρ ^ ( X ^ | Y ^ ) are weak values with as the initial state ρ ^ , a weak measurement of the observable X ^ and a post-selection Π ^ y Y ^ for each y σ ( Y ^ ) . This observation links the conditional expectation of X ^ to the physics of weak values, which has been studied extensively, see [31,45,50,51,52,53]. For the readers’ convenience, we provide, in Appendix A, a concise, complete, self-contained and mathematically clean introduction to weak value measurement, which is the procedure by which the above weak values can be measured. This procedure relies in particular on a postselection, which has historically provided the intuitive link between weak value measurement and conditioning. We refer to the appendix for details. The definition of quantum conditional expectation through minimization that we propose here allows to sharpen this link as we shall further discuss is Section 3.3 and Section 6.

3.2. Characterization of the Left/Right Quantum Conditional Expectations

Let us point out that there are many other quantum analogs of the inner product in Equation (3), among which one could for example consider
X ^ , X ^ α = α X ^ , X ^ + ( 1 α ) X ^ , X ^ r .
for 0 α 1 , interpolating between X ^ , X ^ and X ^ , X ^ r . Minimizing X ^ f ( Y ^ ) , X ^ f ( Y ^ ) α leads to a candidate notion of conditional expectation, whose explicit expression can be easily computed to be
E ρ ^ α ( X ^ | Y ^ ) = y σ ( Y ^ ) φ y Y ^ , ( α X ^ ρ ^ + ( 1 α ) ρ ^ X ^ ) φ y Y ^ φ y Y ^ , ρ ^ φ y Y ^ Π ^ y Y ^ .
Other, more elaborate choices have been considered in [49].
The following theorem shows that to obtain a conditional expectation that satisfies the properties of the classical conditional expectation given in Theorem 3, α has to be either 0 or 1.
Theorem 5.
Let ρ ^ D Y ^ . For any # { , r } , there exist a unique map E cond , Y # : L ( H ) F C , Y ^ satisfying the following properties:
1. 
Pull-out formula. For any X ^ L ( H ) and any f : σ ( Y ^ ) C ,
for # = , E cond , Y ( f ( Y ^ ) X ^ ) = f ( Y ^ ) E cond , Y ( X ^ ) = E cond , Y ( X ^ ) f ( Y ^ ) ,
for # = r , E cond , Y r ( X ^ f ( Y ^ ) ) = f ( Y ^ ) E cond , Y r ( X ^ ) = E cond , Y r ( X ^ ) f ( Y ^ ) .
2. 
Quantum law of iterated expectations. For any X ^ L ( H ) ,
E ρ ^ ( E cond , Y # ( X ^ ) ) = E ρ ^ ( X ^ ) .
This map is given by E cond , Y # ( X ^ ) = E ρ ^ # ( X ^ | Y ^ ) .
This theorem singles out two natural definitions of quantum conditional expectation through the conditions Equations (59) and (61) or Equations (60) and (61).
As in the classical case, the uniqueness in Theorem 5 proves that Equation (53) can be deduced from Equations (59)–(61). Again, one could have shown this fact directly.
Proof. 
Again, we write the proof for # = , the other case being similar. We begin by showing that the left quantum conditional expectation satisfies these two properties. For point 1, we compute
E ρ ^ ( f ( Y ^ ) X ^ | Y ^ ) = y σ ( Y ^ ) φ y Y ^ , f ( Y ^ ) X ^ ρ ^ φ y Y ^ φ y Y ^ , ρ ^ φ y Y ^ Π ^ y = y σ ( Y ^ ) f ( y ) φ y Y ^ , X ^ ρ ^ φ y Y ^ φ y Y ^ , ρ ^ φ y Y ^ Π ^ y = f ( Y ^ ) E ρ ^ ( X ^ | Y ^ ) = E ρ ^ ( X ^ | Y ^ ) f ( Y ^ ) .
For point 2, we compute
E ρ ^ ( E ρ ^ ( X ^ | Y ^ ) ) = Tr ρ ^ y σ ( Y ^ ) φ y Y ^ , X ^ ρ ^ φ y Y ^ φ y Y ^ , ρ ^ φ y Y ^ Π ^ y Y ^ = y σ ( Y ^ ) φ y Y ^ , X ^ ρ ^ φ y Y ^ = Tr ( X ^ ρ ^ ) = E ρ ^ ( X ^ ) .
Now, let E cond , Y : L ( H ) F C , Y ^ be a map that satisfies Equations (59) and (61). We will show that E cond , Y ( X ^ ) = E ρ ^ ( X ^ | Y ^ ) for any X ^ L ( H ) . Let X ^ L ( H ) . For any f ( Y ^ ) F C , Y ^ , we have
f ( Y ^ ) , X ^ E cond , Y ( X ^ ) = Tr ( ρ ^ f ( Y ^ ) ( X ^ E cond , Y ( X ^ ) ) ) = Tr ( ρ ^ f ( Y ^ ) X ^ ) Tr ( ρ ^ E cond , Y ( f ( Y ^ ) X ^ ) ) by Equation ( 59 ) = Tr ( ρ ^ f ( Y ^ ) X ^ ) Tr ( ρ ^ f ( Y ^ ) X ^ ) by Equation ( 61 ) = 0 ,
and of course f ( Y ^ ) , X ^ E ρ ^ ( X ^ | Y ^ ) = 0 as well. It follows
E cond , Y ( X ^ ) E ρ ^ ( X ^ | Y ^ ) , E cond , Y ( X ^ ) E ρ ^ ( X ^ | Y ^ ) = E cond , Y ( X ^ ) X ^ + X ^ E ρ ^ ( X ^ | Y ^ ) , E cond , Y ( X ^ ) E ρ ^ ( X ^ | Y ^ ) = X ^ E cond , Y ( X ^ ) , E cond , Y ( X ^ ) E ρ ^ ( X ^ | Y ^ ) + X ^ E ρ ^ ( X ^ | Y ^ ) , E cond , Y ( X ^ ) E ρ ^ ( X ^ | Y ^ ) = 0
since E cond , Y ( X ^ ) E ρ ^ ( X ^ | Y ^ ) F C , Y ^ . Therefore, Tr ( ρ ^ | E cond , Y ( X ^ ) E ρ ^ ( X ^ | Y ^ ) | 2 ) = 0 . Since E cond , Y ( X ^ ) E ρ ^ ( X ^ | Y ^ ) F C , Y ^ , there exists C ρ ^ : σ ( Y ^ ) C such that E cond , Y ( X ^ ) E ρ ^ ( X ^ | Y ^ ) = C ρ ^ ( Y ^ ) . One gets
0 = Tr ( ρ ^ | C ρ ^ ( Y ^ ) | 2 ) = y σ ( Y ^ ) | C ρ ^ ( y ) | 2 φ y Y ^ , ρ ^ φ y Y ^ ,
which implies, since ρ ^ D Y ^ , that C ρ ^ = 0 , i.e., E cond , Y ( X ^ ) = E ρ ^ ( X ^ | Y ^ ) . □
It is legitimate to ask if one can define a conditional expectation which satisfies both the left- and right-pullout properties, see Equations (59) and (60). This has been performed in the context of von Neumann algebras in [54]. In Appendix B, we briefly explain the link between that approach and ours.
The conditional expectation of X ^ knowing Y ^ is defined only for states ρ ^ D Y ^ , to avoid dividing by 0. As we now prove in the following Lemma, this condition on the state ρ ^ is not too restrictive, as the set D Y ^ is, in a topological sense, a large set (it is dense in the set of all density matrices).
Lemma 2.
Let Y ^ be a self-adjoint operator on H . Then D Y ^ , defined in Equation (47) is dense in D ( H ) (for any norm topology).
Proof. 
Since we are working in finite dimensions, all norms are equivalent. We chose to work with the trace-class norm. Let ρ ^ 0 D ( H ) . If ρ ^ 0 D Y ^ there is nothing to do. Otherwise, let
I 0 = { y σ ( Y ^ ) : Tr ( ρ ^ 0 Π ^ y Y ^ ) = 0 } .
We use I to denote the cardinal of the finite set I. For ε > 0 , let
ρ ^ ε = 1 1 + ε | I 0 | ρ ^ 0 + ε y I 0 Π ^ y Y ^
Then, it is easily verified that ρ ^ ε D Y ^ for any ε > 0 and moreover
Tr ( | ρ ^ ε ρ ^ 0 | ) = Tr ρ ^ 0 1 1 + ε | I 0 | 1 + ε 1 + ε | I 0 | y I 0 Π ^ y Y ^ 1 1 1 + ε | I 0 | + ε | I 0 | 1 + ε | I 0 | ε 0 0 .
This proves the lemma. □
Remark 2. 
( i ) If Y ^ is an incomplete set of commuting observables, the associated projectors ( Π ^ y Y ^ ) y σ ( Y ^ ) are not all of rank 1. However, the results in this section still adapt to this case. For ρ ^ D Y ^ and X ^ L ( H ) , it is easily shown that the minimizers of Equation (48) become
y σ ( Y ^ ) , f 0 ( y ) = Tr ( X ^ ρ ^ Π ^ y Y ^ ) Tr ( ρ ^ Π ^ y Y ^ ) , f 0 r ( y ) = Tr ( ρ ^ X ^ Π ^ y Y ^ ) Tr ( ρ Π ^ y Y ^ ) ,
and that one can define the left and right conditional expectations by
E ρ ^ ( X ^ | Y ^ ) = y σ ( Y ^ ) Tr ( X ^ ρ ^ Π ^ y ) Tr ( ρ ^ Π ^ y ) Π ^ y , E ρ ^ r ( X ^ | Y ^ ) = y σ ( Y ^ ) Tr ( ρ ^ X ^ Π ^ y ) Tr ( ρ ^ Π ^ y ) Π ^ y .
The characterization given in Theorem 5 remains true in this setting, as do the equivalences of Proposition A6 but not Equation (A44). In the following sections, we require the completeness of Y ^ .
( i i ) We finally note that in the case of commuting operators X ^ and Y ^ , the quantum conditional expectation reduces to the classical conditional expectation (1), when considering the two operators as classical random variables with joint probability P ( X ^ = x , Y ^ = y ) = Tr ( Π ^ x Π ^ y ρ ^ ) , the Born measure of the commuting system ( X ^ , Y ^ ) .

3.3. The Quantum Laws of Iterated Expectations and of Total Variance

The first two moments of the quantum conditional expectations E ρ ^ / r ( X ^ | Y ^ ) have properties that are reminiscent of those of the classical conditional expectation: the law of iterated expectations and the law of total variance. There are however some important differences between the classical and quantum versions that we shall point out. The results are summed up in Proposition 2 and in the discussion following it.
We introduce, for any operator C ^ L ( H ) , the left and right variance of C ^ in the state ρ ^ by
Δ ρ ^ , 2 ( C ^ ) = Tr ρ ^ ( C ^ E ρ ^ ( C ^ ) ) ( C ^ E ρ ^ ( C ^ ) ) , Δ ρ ^ r , 2 ( C ^ ) = Tr ρ ^ ( C ^ E ρ ^ ( C ^ ) ) ( C ^ E ρ ^ ( C ^ ) ) .
When C ^ = C ^ or C ^ = C ^ , we will write
Δ ρ ^ 2 C ^ = Δ ρ ^ , 2 ( C ^ ) = Δ ρ ^ r , 2 ( C ^ ) .
We further introduce the left and right conditional variances of C ^ , analogously to the classical definition:
Δ ρ ^ , 2 ( C ^ | Y ^ ) = E ρ ^ ( | C ^ E ρ ^ ( C ^ | Y ^ ) | 2 | Y ^ ) , Δ ρ ^ r , 2 ( C ^ | Y ^ ) = E ρ ^ r ( C ^ E ρ ^ r ( C ^ | Y ^ ) ) ( C ^ E ρ ^ r ( C ^ | Y ^ ) ) | Y ^ .
Recall that, for any operator Z, one has | Z ^ | 2 = Z ^ Z ^ . In addition, a simple computation allows establishing that the expected value of the conditional variance equals the error
E ρ ^ ( Δ ρ ^ / r , 2 ( C ^ | Y ^ ) ) = C ^ E ρ ^ / r ( C ^ | Y ^ ) , C ^ E ρ ^ / r ( C ^ | Y ^ ) / r .
The notations are cumbersome, but the treatment is similar for the “left” and “right” cases. Recall that, even if X ^ is self-adjoint, its conditional expectation may not be. The left/right conditional variances will generally then differ. The following proposition sums up the basic properties of the conditional expectation and the conditional variance.
Proposition 2.
If ρ ^ D Y ^ , then for any X ^ L ( H ) the quantum law of iterated expectations holds:
E ρ ^ ( E ρ ^ / r ( X ^ | Y ^ ) ) = E ρ ^ ( X ^ ) .
The quantum law of total variance reads as
Δ ρ ^ / r , 2 X ^ = Δ ρ ^ / r , 2 E ρ ^ / r ( X ^ | Y ^ ) + X ^ E ρ ^ / r ( X ^ | Y ^ ) , X ^ E ρ ^ / r ( X ^ | Y ^ ) / r
= Δ ρ ^ / r , 2 E ρ ^ / r ( X ^ | Y ^ ) + E ρ ^ ( Δ ρ ^ / r , 2 ( X ^ | Y ^ ) ) .
If ρ ^ D Y ^ is pure,
Δ ρ ^ , 2 ( X ^ | Y ^ ) = 0
and thus,
Δ ρ ^ / r , 2 X ^ = Δ ρ ^ / r , 2 E ρ ^ / r ( X ^ | Y ^ ) .
The quantum law of iterated expectations, Equation (68), says that the best estimator of X ^ by a function of Y ^ has the same expected value as X ^ itself; it is, in this sense, unbiased, as the classical conditional expectation. Equation (69) is in turn a quantum analog of the classical law of total variance Equation (30). It provides a relation between the variance of X ^ , the variance of the conditional expectation E ρ ^ / r ( X ^ | Y ^ ) , and the expected value of the conditional variance Δ / r , 2 ( X | Y ) , in complete analogy with the classical situation. The error term in this relation (the second term on the right-hand side of Equation (69)) can, as in the classical case, be expressed as the expected value of the conditional variance: this is the content of Equation (67).
Proof. 
We give the proof for the left conditional expectation E ρ ^ ( X ^ | Y ^ ) . Note that Equation (68) is exactly Equation (61), and it has already been shown in the proof of Theorem 5.
To prove Equation (69), we note that, from Equation (49), we obtain
min f F C , Y ^ ( X ^ f ( Y ^ ) ) , ( X ^ f ( Y ^ ) ) = Tr ( X ^ X ^ ρ ^ ) y | φ y Y ^ , X ^ ρ ^ φ y Y ^ | 2 φ y Y ^ , ρ ^ φ y Y ^ = X ^ , X ^ E ρ ^ ( X ^ | Y ^ ) , E ρ ^ ( X ^ | Y ^ ) .
Thus,
X ^ E ρ ^ ( X ^ | Y ^ ) , X ^ E ρ ^ ( X ^ | Y ^ ) = X ^ , X ^ E ρ ^ ( X ^ | Y ^ ) , E ρ ^ ( X ^ | Y ^ ) .
We have that
Δ ρ ^ , 2 ( X ^ ) = Tr ( ρ ^ X ^ X ^ ) E ρ ^ ( X ^ ) ¯ Tr ( ρ ^ X ^ ) E ρ ^ ( X ^ ) Tr ( ρ ^ X ^ ) + E ρ ^ ( X ^ ) E ρ ^ ( X ^ ) ¯ = X ^ , X ^ E ρ ^ ( X ^ ) 2 = X ^ E ρ ^ ( X ^ | Y ^ ) , X ^ E ρ ^ ( X ^ | Y ^ ) + E ρ ^ ( X ^ | Y ^ ) , E ρ ^ ( X ^ | Y ^ ) | E ρ ^ ( E ρ ^ ( X ^ | Y ^ ) ) | 2 = X ^ E ρ ^ ( X ^ | Y ^ ) , X ^ E ρ ^ ( X ^ | Y ^ ) + Δ ρ ^ , 2 ( E ρ ^ ( X ^ | Y ^ ) )
where we used Equations (68) and (74) on the third line.
Equation (67) now follows:
E ρ ^ ( Δ , 2 ( X ^ | Y ^ ) ) = E ρ ^ E ρ ^ ( | X ^ E ρ ^ ( X ^ | Y ^ ) | 2 | Y ^ ) = E ρ ^ | X ^ E ρ ^ ( X ^ | Y ^ ) | 2 = Tr ρ | X ^ E ρ ^ ( X ^ | Y ^ ) | 2 = X ^ E ρ ^ ( X ^ | Y ^ ) , X ^ E ρ ^ ( X ^ | Y ^ )
where we used the quantum law of iterated expectations.
We finally prove Equation (71). We have that
Δ ρ ^ , 2 ( X ^ | Y ^ ) = E ρ ^ ( | X ^ E ρ ^ ( X ^ | Y ^ ) | 2 | Y ^ ) = E ρ ^ ( X ^ X ^ | Y ^ ) E ρ ^ ( X ^ E ρ ^ ( X ^ | Y ^ ) | Y ^ ) E ρ ^ ( E ρ ^ ( X ^ | Y ^ ) X ^ | Y ^ ) + E ρ ^ ( | E ρ ^ ( X ^ | Y ^ ) | 2 | Y ^ ) = E ρ ^ ( X ^ X ^ | Y ^ ) E ρ ^ ( X ^ E ρ ^ ( X ^ | Y ^ ) | Y ^ ) ,
We now compute
E ρ ^ ( X ^ E ρ ^ ( X ^ | Y ^ ) | Y ^ ) = y σ ( Y ^ ) Tr ( Π ^ y Y ^ X ^ E ρ ^ ( X ^ | Y ^ ) ρ ^ ) Tr ( Π ^ y Y ^ ρ ^ ) Π ^ y Y ^ = y σ ( Y ^ ) 1 Tr ( Π ^ y Y ^ ρ ^ ) y σ ( Y ^ ) Tr ( Π ^ y Y ^ X ^ ρ ^ ) Tr ( Π ^ y Y ^ X ^ Π ^ y Y ^ ρ ^ ) Tr ( Π ^ y Y ^ ρ ^ ) Π ^ y Y ^ .
As ρ ^ is a pure state, we note ρ ^ = ψ ψ , and we obtain that
y σ ( Y ^ ) Tr ( Π ^ y Y ^ X ^ ρ ^ ) Tr ( Π ^ y Y ^ X ^ Π ^ y Y ^ ρ ^ ) Tr ( Π ^ y Y ^ ρ ^ ) = y σ ( Y ^ ) ψ , φ y Y ^ φ y Y ^ , X ^ ψ ψ , φ y Y ^ φ y Y ^ , X ^ φ y Y ^ φ y Y ^ , ψ | φ y Y ^ , ψ | 2 = y σ ( Y ^ ) ψ , φ y Y ^ φ y Y ^ , X ^ φ y Y ^ φ y Y ^ , X ^ ψ = ψ , φ y Y ^ φ y Y ^ , X ^ X ^ ψ = Tr ( Π ^ y Y ^ X ^ X ^ ρ ^ ) .
Consequently, we conclude that
E ρ ^ ( X ^ E ρ ^ ( X ^ | Y ^ ) | Y ^ ) = y σ ( Y ^ ) Tr ( Π ^ y Y ^ X ^ X ^ ρ ^ ) Tr ( Π ^ y Y ^ ρ ^ ) Π ^ y Y ^ = E ρ ^ ( X ^ X ^ | Y ^ ) .
This proves Equation (71). Moreover, Equation (72) is a direct consequence of Equation (69), Equations (67) and (71). □
Having stressed the analogies between the classical and quantum conditional expectations and their basic properties, we now discuss four important differences between them.
First, classically, one has, if X is bounded ( X min X ( ω ) X max ), that
X min E P ( X | Y = y ) X max .
We have already pointed out that quantum mechanically, even if X ^ is self-adjoint, E P ( X ^ | Y ^ = y ) may be complex non-real. In addition, when it is indeed real, it may not be confined to lie between the minimal and maximal eigenvalue of X ^ . One says a value E ρ ^ ( X ^ | Y ^ = y ) of the conditional expectation E ρ ^ ( X ^ | Y ^ ) is anomalous if E ρ ^ ( X ^ | Y ^ = y ) [ x min , x max ] , where x min , max are the extremal eigenvalues of X ^ . The historic example of such anomalous values corresponds to a spin 1/2 [36]. Let H = C 2 , X ^ = σ ^ z , Y ^ = σ ^ x and ρ ^ = | α α | , | α = cos α | 1 + sin α | 1 . Then,
E ρ ^ ( X ^ | Y ^ = 1 ) = sin α cos α sin α + cos α , E ρ ^ ( X ^ | Y ^ = 1 ) = sin α + cos α sin α cos α .
Choosing α close to π / 4 , for example, the second term can be made arbitrary large, while the first one tends to 0. This phenomenon lies at the basis of weak value amplification; see [31] for a review on this topic. Note that, when α = π / 4 , then | α does not belong to D Y ^ , since it is then an eigenvector of Y ^ = σ ^ x .
The existence of anomalous values is a strong indicator of typically quantum mechanical behavior as the following immediate consequence of Equation (77) shows.
Lemma 3.
Let X ^ and Y ^ be observables and ρ ^ D Y ^ . Suppose there exists a probability μ : ( x , y ) σ ( X ^ ) × σ ( Y ^ ) μ ( x , y ) [ 0 , 1 ] satisfying
(i) 
Born compatibility: for all x σ ( X ^ ) , for all y σ ( Y ^ ) ,
x σ ( X ^ ) μ ( x , y ) = Tr ( Π ^ y Y ^ ρ ^ ) , y σ ( Y ^ ) μ ( x , y ) = Tr ( Π ^ x X ^ ρ ^ ) ;
(ii) 
For all y σ ( Y ^ ) , E ρ ^ ( X ^ | Y ^ = y ) = x σ ( X ^ ) x μ ( x , y ) ( ρ ^ Π ^ y Y ^ ) .
Then,
x min E ρ ^ ( X ^ | Y ^ = y ) x max .
The same statement holds with replaced by r . Condition (i) asserts that, for the given triple ( X ^ , Y ^ , ρ ^ ) , μ is a joint probability distribution on the joint spectrum of X ^ and Y ^ that has the correct quantum mechanical Born marginals in the state ρ ^ and X ^ and Y ^ . Condition (ii) adds a condition on the natural associated conditional probability obtained from Bayes’ rule, μ ( x | y ) = μ ( x , y ) Tr ( ρ ^ Y ^ ) : it must lead to a conditional expectation for X ^ , given Y ^ , that coincides with the quantum mechanical conditional expectation E ρ ^ ( X ^ | Y ^ = y ) . This then implies that the quantum mechanical conditional expectation cannot take on anomalous values. This lemma is of interest because its contrapositive provides a no-go theorem for the existence of joint probabilities: if, for a given triple ( X ^ , Y ^ , ρ ^ ) , the conditional expectation E ρ ^ ( X ^ | Y ^ = y ) admits at least one anomalous value, then there does not exist a joint probability distribution satisfying (i) and (ii) of the lemma. The usual “no-go” theorems of this type, ruling out the existence of Born-compatible joint probability distributions in quantum mechanics, rely on assumptions that need to hold for all  ρ ^ , given a pair of observables X ^ and Y ^ or, alternatively, for all ρ ^ and all X ^ (see for example [1,4,30]). Here, we consider a single-state ρ ^ and add to the Born-compatibility of this state a second compatibility condition of the same type, still only for this unique state; indeed, (ii) requires equality between the “classical” conditional expectation of X ^ and the quantum one defined in the previous section. If the quantum conditional expectation admits at least one anomalous value (i.e., if Equation (80) is violated) then such a joint probability distribution cannot exist. It should be noted that the Born-compatibility condition in and by itself is very weak and can always be satisfied by taking, for example:
μ ( x , y ) = Tr ( ρ ^ Π ^ x X ^ ) Tr ( ρ ^ Π ^ y Y ^ ) .
However, this treats x and y as independent random variables, which is generally not coherent with the second condition. The current standard way to avoid this no-go theorem is to allow μ ( x , y ) to be complex-valued, i.e., a quasiprobability distribution rather than a probability distribution, as we will see in the next section.
To explain the second difference between the classical and quantum conditional expectations, we express Equation (69) in terms of the self-adjoint and the anti-self-adjoint part of E ρ ^ / r ( X ^ | Y ^ ) . For the self-adjoint part of E ρ ^ / r ( X ^ | Y ^ ) , we write
E ρ ^ , sa ( X ^ | Y ^ ) : = 1 2 E ρ ^ ( X ^ | Y ^ ) + E ρ ^ ( X ^ | Y ^ ) = y σ ( Y ^ ) Re φ y , X ^ ρ ^ φ y φ y , ρ ^ φ y Π ^ y
E ρ ^ r , sa ( X ^ | Y ^ ) : = 1 2 E ρ ^ r ( X ^ | Y ^ ) + E ρ ^ r ( X ^ | Y ^ ) = y σ ( Y ^ ) Re φ y , ρ ^ X ^ φ y φ y , ρ ^ φ y Π ^ y ,
for any ρ ^ D Y ^ and any X ^ L ( H ) . Similarly, we introduce
Im E ρ ^ / r ( X ^ | Y ^ ) = 1 2 i E ρ ^ / r ( X ^ | Y ^ ) E ρ ^ / r ( X ^ | Y ^ )
As the self-adjoint and anti-self-adjoint parts of E ρ ^ / r ( X ^ | Y ^ ) are both functions of Y ^ , they commute. Then,
Δ ρ ^ / r , 2 E ρ ^ / r ( X ^ | Y ^ ) = Δ ρ ^ 2 E ρ ^ / r , sa ( X ^ | Y ^ ) + Δ ρ ^ 2 Im E ρ ^ / r ( X ^ | Y ^ ) .
In conclusion, the law of total variance, Equation (69), shows that the variance of X ^ differs from the sum of the variances of the self-adjoint and anti-self-adjoint part of its conditional expectation given Y ^ by the mean quadratic error between the two. Note that, for self-adjoint X ^ , the underlying chosen ordering is irrelevant:
E ρ ^ , sa ( X ^ | Y ^ ) = E ρ ^ r , sa ( X ^ | Y ^ ) = 1 2 y σ ( Y ^ ) φ y , { ρ ^ , X ^ } φ y φ y , ρ ^ φ y ,
Im E ρ ^ ( X ^ | Y ^ ) = Im E ρ ^ r ( X ^ | Y ^ ) = 1 2 i y σ ( Y ^ ) φ y , [ ρ ^ , X ^ ] φ y φ y , ρ ^ φ y ,
where { · , · } denotes the anti-commutator and [ · , · ] the commutator. Thus, for self-adjoint X ^ , we will write
E ρ sa ( X ^ | Y ^ ) : = E ρ ^ , sa ( X ^ | Y ^ ) = E ρ ^ r , sa ( X ^ | Y ^ ) .
It should be noted that this equality need not be true if X ^ is not a self-adjoint operator. In particular, we have that, generally,
E ρ ^ , sa ( f ( Y ^ ) X ^ | Y ^ ) E ρ ^ r , sa ( f ( Y ^ ) X ^ | Y ^ )
for f : σ ( Y ^ ) R , even if X ^ is self-adjoint. The law of total variance, Equation (69), then becomes, for self-adjoint X ^ ,
Δ ρ ^ 2 X ^ = Δ ρ ^ 2 E ρ ^ sa ( X ^ | Y ^ ) + Δ ρ ^ 2 Im E ρ ^ / r ( X ^ | Y ^ ) + ( X ^ E ρ ^ / r ( X ^ | Y ^ ) ) , ( X ^ E ρ ^ / r ( X ^ | Y ^ ) ) / r .
Comparing this to the classical equality Equation (30), one notices the extra term given by the variance of Im E ρ ^ / r ( X ^ | Y ^ ) which comes from the non-self-adjointness of the conditional expectation. This term depends on the commutator [ ρ ^ , X ^ ] and as such, its absence classically is expected. We will see in Section 6 that it can vanish even if ρ ^ and X ^ do not commute, and we will use results on the KD representations of quantum mechanics to show many examples where Im E ρ ^ / r ( X ^ | Y ^ ) vanishes. We also further explore there the link between Im E ρ ^ / r ( X ^ | Y ^ ) and an appropriately defined Fisher information associated to a parameter estimation protocol naturally associated with a triple ρ ^ , X ^ , and Y ^ .
The third difference with the classical situation concerns what happens for pure states. In classical probability theory, the pure states are Dirac delta measures δ ω 0 , on Ω , with ω 0 Ω . If X , Y are random variables, E δ ω o ( X | Y = y ) = X ( ω 0 ) δ Y ( ω 0 ) ( y ) , and is therefore also a Dirac delta measure, on R . As a result, the classical analogues to Equations (71) and (72) also hold, but with one major difference: since the Dirac delta measures are dispersion-free, both sides of this classical equivalent of Equation (72) then vanish. In the quantum case, pure states ρ ^ = | ψ ψ | are never dispersion-free, and neither side vanishes as soon as | ψ is not an eigenstate of X ^ . So, in quantum theory, the error term in the law of total variation vanishes for pure states, and the variance of X ^ equals that of its conditional expectation E | ψ ψ | ( X ^ | Y ^ ) .
Fourth, and finally, we note that the well-known relation that holds in classical probability
Δ 2 ( X | Y ) = E P ( X 2 | Y ) E P ( X | Y ) 2 ,
does not hold in the quantum case. Indeed, it follows from Equation (76) that, for the left quantum conditional expectation, for example:
Δ , 2 ( X ^ | Y ^ ) = E ρ ^ ( X ^ X ^ | Y ^ ) E ρ ^ ( X ^ E ρ ^ ( X ^ | Y ^ ) | Y ) E ρ ^ ( | X ^ | 2 | Y ) E ρ ^ ( X ^ | Y ^ ) 2 .
One can also define an alternative conditional variance by the final expression above, specifically,
Δ ˜ ρ ^ 2 , ( X ^ | Y ^ ) : = E ρ ^ ( | X ^ | 2 | Y ) E ρ ^ ( X ^ | Y ^ ) 2 ,
with an obvious right variant obtained by replacing | X ^ | 2 = X ^ X ^ by X ^ X ^ (the two coincide if X ^ is self-adjoint). These will then still satisfy the law of total variance, in the form
Δ 2 , ( X ^ ) = Δ 2 , E ρ ^ ( X ^ | Y ^ ) + E ρ ^ Δ ˜ 2 , ( X ^ | Y ^ ) .
Another notion of conditional variance that has been proposed in the literature is the weak variance introduced in [47], which with our notations is defined as
σ w , ρ ^ ( X ^ | Y ^ ) = E ρ ^ ( X ^ 2 | Y ^ ) E ρ ^ ( X ^ | Y ^ ) 2 .
This weak variance can be given an operational meaning in the context of weak measurements (cf. Appendix A, in particular Formula (A33)) but does not satisfy the law of total variance, unlike (89) which can be given a similar operational significance: see Corollary A1. There exists in fact a whole range of plausible definitions of conditional variances, including variants which can be introduced in the context of quasiprobabilistic representations of quantum mechanics, whose relative merits still have to be investigated.

3.4. The Real Part of E ρ ^ / r ( X ^ | Y ^ ) as a Best Estimator

As one might expect, the left and right self-adjoint quantum conditional expectations E ρ ^ / r , sa ( X ^ | Y ^ ) satisfy properties similar to the ones given in Theorem 5. We specify them for the left self-adjoint quantum conditional expectation. They also hold for the right self-adjoint one, with adapted ordering. We use F R , Y ^ to denote the set of self-adjoint functions of Y ^ :
F R , Y ^ = f ( Y ^ ) f : σ ( Y ^ ) R .
Slightly adapting the proof of Theorem 5, one easily shows that for ρ ^ D Y ^ , E ρ ^ , sa ( · | Y ^ ) is the unique map with values in the set F R , Y ^ with the two following properties:
1.
For any f : σ ( Y ^ ) R and any X ^ L ( H ) ,
E ρ ^ , sa ( f ( Y ^ ) X ^ | Y ^ ) = f ( Y ^ ) E ρ ^ , sa ( X ^ | Y ^ ) .
2.
For any X ^ L ( H ) ,
E ρ ^ ( E ρ ^ , sa ( X ^ | Y ^ ) ) = Re ( E ρ ^ ( X ^ ) ) .
Finally, it is also easily seen that this left/right conditional expectations can be obtained as the minimizer of the mean square error functional
f X ^ f ( Y ^ ) , X ^ f ( Y ^ ) / r ,
where the minimization is considered over the set of real valued functions f : σ ( Y ^ ) R . The self-adjoint left conditional expectation of X ^ is therefore the observable closest to X ^ (with respect to the above degenerate left inner product) that is a self-adjoint function of Y ^ .
For self-adjoint X ^ , this real minimization problem was previously addressed in [32,33,34,35,49], where the link between the minimizer, viewed as a conditional expectation, and the real part of weak values was also established. As we pointed out above, our analysis in Section 3.1 shows that, in fact, the conditional expectation E ρ ^ / r ( X ^ | Y ^ ) itself is a best estimator of X ^ (self-adjoint or not), when the minimization is performed over all complex valued operator functions of Y ^ , and not only over the self-adjoint ones.
As in the proof of Proposition 2, one finds that for any X ^ L ( H )
X ^ E ρ / r , sa ( X ^ | Y ^ ) , X ^ E ρ / r , sa ( X ^ | Y ^ ) / r = E ρ ^ ( | X ^ | 2 ) E ρ ^ E ρ ^ / r , sa ( X ^ | Y ^ ) 2 .
Furthermore, if X ^ is self-adjoint, we have E ρ ^ ( E ρ ^ sa ( X ^ | Y ^ ) ) = Re E ρ ^ ( X ^ ) = E ρ ^ ( X ^ ) . We deduce
Δ ρ ^ 2 X ^ = Δ ρ ^ 2 E ρ ^ sa ( X ^ | Y ^ ) + X ^ E ρ ^ sa ( X ^ | Y ^ ) , X ^ E ρ ^ sa ( X ^ | Y ^ ) / r .
Comparing this to Equation (69), we have that the error made when approximating the variance of X ^ by the variance of E ρ ^ sa ( X ^ | Y ^ ) is greater than the error made when approximating the variance of X ^ by the variance of E ρ ^ / r ( X ^ | Y ^ ) . This is a consequence of the fact that the mean squared error functional is minimized over a larger class of functions in the latter case than in the former. More precisely, from Equation (88), we have
X ^ E ρ ^ sa ( X ^ | Y ^ ) , X ^ E ρ ^ sa ( X ^ | Y ^ ) / r = Δ ρ ^ 2 Im E ρ ^ / r ( X ^ | Y ^ ) + ( X ^ E ρ ^ / r ( X ^ | Y ^ ) ) , ( X ^ E ρ ^ / r ( X ^ | Y ^ ) ) / r .

4. Quantum Conditional Expectation via Quasiprobability Representation of Quantum Mechanics

As shown in the previous section, defining the quantum conditional expectation via a minimization problem as in Equation (44) or Equation (46) is a way of mimicking what happens classically, see Theorem 2. It also allowed us to compare the classical and quantum notions and to emphasize their differences. One can alternatively try to imitate the classical conditional expectation by using Definition 1. This definition uses the joint probability distribution of the two random variables under consideration, which is naturally defined in classical probability theory. In quantum mechanics, as recalled in the introduction, a notion of joint probability distribution does not exist for non-commuting observables, but one can try to use a quasiprobability distribution instead.
To that end, we first describe the class of quasiprobability representations of quantum mechanics that we consider. We will use the formalism of frames, as detailed in [3,30], to do so. Let H be a Hilbert space of dimension d, let Λ be a finite set, with Λ = d 2 , where A is the cardinal of the finite set A. Let ( S λ ) λ Λ be a basis of L ( H ) , satisfying
λ S λ = I d .
Let ( T λ ) λ Λ denote its unique dual basis, which satisfies
Tr ( T λ S λ ) = δ λ , λ .
Given an operator C ^ on H , we define
Q λ ( C ^ ) = Tr ( C ^ S ^ λ ) , Q ˜ λ ( C ^ ) = Tr ( C ^ T λ ) .
It follows that the maps Q , Q ˜ : L ( H ) C Λ are bijective and that
C ^ = λ Q ˜ λ ( C ^ ) S λ , C ^ = λ Q λ ( C ^ ) T λ .
Given C ^ , D ^ L ( H ) , one then has
Tr C ^ = λ Q λ ( C ^ ) , Tr ( C ^ D ^ ) = λ Q ˜ λ ( C ^ ) ¯ Q λ ( D ^ ) .
The pair ( Q , Q ˜ ) , which is completely determined by the choice of Λ and of the basis ( S λ ) λ , is referred to as a quasiprobability representation of quantum mechanics. It associates to each density matrix ρ ^ a quasiprobability Q λ ( ρ ^ ) on Λ , which is a complex-valued function that satisfies
1 = λ Q λ ( ρ ^ ) ,
and to each observable C ^ = C ^ its symbol Q ˜ λ ( C ^ ) , with
Tr ( ρ ^ C ^ ) = λ Q λ ( ρ ^ ) Q ˜ λ ( C ^ ) . ¯
Depending on context, one thinks of Λ as an ontic space or as a classical phase space, on which the quantum states and observables are represented by quasiprobabilities and functions respectively. Many quasiprobability representations of quantum mechanics fitting in the above framework have been introduced and studied. A number of those [14,15,16,17,18,20,21] aim at defining quasiprobability representations for quantum systems on finite dimensional Hilbert space reproducing many of the known properties of the Wigner–Weyl–Moyal representation [5,9,12,13] associated with conjugate variables and the Fourier transform on L 2 ( R 2 ) . Extensions of such constructions to locally compact Abelian groups have been considered more recently in [22,25]. A more versatile family of quasiprobability representation is provided by the Kirkwood–Dirac representations: they are defined using two arbitrary observables A ^ and B ^ and we shall describe them in detail in the next section. Further examples of quasiprobability representations can be found in [30].
Definition 3.
Let Y ^ be a CSCO and let ( Q , Q ˜ ) be a quasiprobability representation of quantum mechanics. Writing Q ˜ λ ( Y ^ ) : = ( Q ˜ λ ( Y ^ 1 ) , , Q ˜ λ ( Y ^ m ) ) , we define, for all y C m :
Λ y Y ^ = { λ Λ Q ˜ λ ( Y ^ ) = y } .
We say the quasiprobability representation ( Q , Q ˜ ) of quantum mechanics is Y ^ -compatible provided for all density matrices ρ ^ , one has
λ Λ y Y ^ Q λ ( ρ ^ ) = Tr ( ρ ^ Π ^ y Y ^ ) , when y σ ( Y ^ ) ,
= 0 , when y σ ( Y ^ ) .
If Q λ ( ρ ^ ) is a probability on Λ (meaning it is non-negative for all λ ), then what this condition means is that the joint probability law of the vector-valued random variable Q ˜ ( Y ^ ) = ( Q ˜ ( Y ^ 1 ) , , Q ˜ ( Y ^ m ) ) ( C m ) Λ induced by the probability Q λ ( ρ ^ ) on Λ is identical to the quantum mechanical Born-probability law of Y ^ when the system is in the state ρ ^ . More generally, to be Y ^ -compatible, a quasiprobability representation must reproduce the correct Born-probabilities for the observable Y ^ and for all states ρ ^ , even those for which the quasiprobability Q λ ( ρ ^ ) takes on non-positive values and is therefore only a quasiprobability distribution.
The following lemma expresses the fact that, if ( Q , Q ˜ ) is Y ^ -compatible, then it agrees naturally with the functional calculus of Y ^ in the sense that the symbol of f ( Y ^ ) equals λ f ( Q ˜ λ ( Y ^ ) ) .
Lemma 4.
If ( Q , Q ˜ ) is a Y ^ -compatible quasiprobability representation, then the family ( Λ y Y ^ ) y σ ( Y ^ ) is a partition of Λ and
f ( Y ^ ) F C , Y ^ , λ Λ , Q ˜ λ ( f ( Y ^ ) ) = f ( Q ˜ λ ( Y ^ ) ) .
The lemma states that the symbol of a function f of Y ^ equals the same function f of the symbol of Y ^ . Note that, generally, if X ^ and Z ^ are operators on H , then it is not true that Q ˜ ( X ^ Z ^ ) = Q ˜ ( X ^ ) Q ˜ ( Z ^ ) . When both X ^ and Z ^ are functions of Y ^ , this is true, however, by the above lemma, for Y ^ -compatible quasiprobability representations.
Proof. 
From Equation (104), one finds that for all ρ ^ and for all y σ ( Y ^ )
Tr ρ ^ λ Λ y Y ^ S λ = Tr ρ ^ Π ^ y Y ^ .
Therefore,
Π ^ y Y ^ = λ Λ y Y ^ S λ = λ Λ y Y ^ S λ .
Equation (100) then implies that
Q ˜ ( Π ^ y Y ^ ) = 1 Λ y Y ^ ,
where 1 E denotes the indicator function of a set E, and therefore
1 = Q ˜ ( I d ) = y σ ( Y ^ ) Q ˜ ( Π ^ y Y ^ ) = y σ ( Y ^ ) 1 Λ y Y ^ .
Moreover, it is clear that Λ y Y ^ Λ y Y ^ = if y y . Since none of the Λ y Y ^ are empty, by Equation (107), this implies that ( Λ y Y ^ ) y σ ( Y ^ ) is a partition of Λ . Consequently, one has that
Q ˜ ( Y ^ ) = y σ ( Y ^ ) y 1 Λ y Y ^ .
Finally,
Q ˜ ( f ( Y ^ ) ) = y σ ( Y ^ ) f ( y ) Q ˜ ( Π ^ y Y ^ ) = y σ ( Y ^ ) f ( y ) 1 Λ y Y ^ = f ( Q ˜ ( Y ^ ) ) .
Suppose now that we have a ( Q , Q ˜ ) that is Y ^ -compatible. Given ρ ^ D Y ^ , we then define the quasiprobability of λ , given y σ ( Y ^ ) , as follows
Q λ | y ( ρ ^ ) = Q λ ( ρ ^ ) Q ( Λ y Y ^ ) = Q λ ( ρ ^ ) φ y | ρ ^ | φ y λ Λ y Y ^ ,
= 0 λ Λ y Y ^ .
This allows us to define a notion of conditional expectation naturally associated to the quasiprobability representation ( Q , Q ˜ ) , as follows.
Definition 4.
For Y ^ -compatible ( Q , Q ˜ ) we define the Q-conditional expectation of X ^ L ( H ) knowing Y ^ in the state ρ ^ D Y ^ by
E ρ ^ Q ( X ^ | Y ^ ) = y σ ( Y ^ ) λ Λ y Y ^ Q ˜ λ ( X ^ ) ¯ Q λ | y ( ρ ^ ) Π ^ y Y ^ .
Note that, whenever ρ ^ is Q-positive, by which we mean that Q λ ( ρ ^ ) 0 for all λ , Q λ | y ( ρ ^ ) is equal to the conditional probability of λ , given that the random variable Q ˜ ( Y ^ ) takes the value y: Q ˜ ( Y ^ ) = y . Nevertheless, even in that case, and even if X ^ is self-adjoint, E ρ ^ Q ( X ^ | Y ^ ) is not necessarily self-adjoint. This will be the case, provided Q ˜ λ ( X ^ ) is real.
The following result is now immediate:
Proposition 3.
The Q-conditional expectation has the following property: for any ρ ^ D Y ^ and any X ^ L ( H ) ,
E ρ ^ ( E ρ ^ Q ( X ^ | Y ^ ) ) = E ρ ^ ( X ^ ) .
In particular, for any ρ D Y ^ and any λ Λ
Q λ ( ρ ^ ) = Tr ( ρ ^ S λ ) = E ρ ^ ( E ρ ^ Q ( S λ | Y ^ ) ) .
We can now formulate our first main result. It provides a characterization of all quasiprobability distributions that are compatible with the projective measurement of a given CSCO Y ^ and the associated conditional expectation of which coincides with the one defined in terms of minimization, as in the previous section. Recall that λ Q ˜ λ ( Y ^ ) sends Λ onto the spectrum of Y ^ , by Equation (111).
Theorem 6.
Let ( Q , Q ˜ ) be a Y ^ -compatible quasiprobability representation of quantum mechanics on H . Then, the following statements are equivalent:
(i) 
X ^ L ( H ) , ρ ^ D Y ^ , E ρ ^ Q ( X ^ | Y ^ ) = E ρ ^ ( X ^ | Y ^ ) ;
(ii) 
X ^ L ( H ) , ρ ^ D Y ^ , f : σ ( Y ^ ) C , E ρ ^ Q ( f ( Y ^ ) X ^ | Y ^ ) = f ( Y ^ ) E ρ ^ Q ( X ^ | Y ^ ) .
(iii) 
λ Λ , ρ ^ D Y ^ , S λ = S λ Π ^ y Y ^ , T λ = T λ Π ^ y Y ^ , with y = Q ˜ λ ( Y ^ ) σ ( Y ^ ) .
Condition (iii) implies that both the frame operators S λ and their duals T λ are rank one operators, which is quite a restrictive condition. It is this property that will allow us to single out the KD distributions in Section 5.2 below.
Proof. 
The statement that (i) and (ii) are equivalent follows from Theorem 5 and Equation (116). We now prove that (i) implies (iii). From (i), one finds that, for all X ^ L ( H ) , for all y σ ( Y ^ ) and ρ ^ D Y ^ ,
λ Λ y Y ^ Q ˜ λ ( X ^ ) ¯ Q λ | y ( ρ ^ ) = Tr ( ρ ^ Π ^ y Y ^ X ^ ) Tr ( ρ ^ Π ^ y Y ^ ) .
Inserting X ^ = S λ and using that Q ˜ λ ( S λ ) = δ λ , λ , one finds that
λ Λ y Y ^ δ λ λ Q λ ( ρ ^ ) = Tr ( ρ ^ Π ^ y Y ^ S λ ) .
Since both sides are continuous in ρ ^ , this equality holds for all ρ ^ by Lemma 2 and hence, for all y σ ( Y ^ ) , λ Λ ,
λ Λ y Y ^ δ λ λ S λ = Π ^ y Y ^ S λ .
We know from Lemma 4 that Q ˜ λ ( Y ^ ) σ ( Y ^ ) . Consequently, Equation (120) implies that, if λ Λ y Y ^ , then S λ Π ^ y Y ^ = 0 and if λ Λ y Y ^ , then S λ = S λ Π ^ y Y ^ . Since y σ ( Y ^ ) Π ^ y Y ^ = I d , it follows that, for all λ Λ
S λ = y σ ( Y ^ ) S λ Π ^ y Y ^ = S λ Π ^ Q ˜ λ ( Y ^ ) Y ^ + y σ ( Y ^ ) { Q ˜ λ ( Y ^ ) } S λ Π ^ y Y ^ = S λ Π ^ Q ˜ λ ( Y ^ ) Y ^ .
Introducing, for y σ ( Y ^ ) ,
L y = span { S λ λ Λ y Y ^ } ,
it follows that these spaces are orthogonal with respect to the Hilbert–Schmidt inner product,
L y L y if y y .
Since the Λ y Y ^ form a partition of Λ , one therefore concludes that
L ( H ) = y σ ( Y ^ ) L y ,
where the sum is an orthogonal direct sum for the Hilbert–Schmidt inner product. It follows from this that
T λ = T λ Π ^ Q ˜ λ ( Y ^ ) Y ^ .
Indeed, T λ is by definition orthogonal to all S λ with λ λ . It is therefore orthogonal to all S λ with λ such that Q ˜ λ ( Y ^ ) Q ˜ λ ( Y ^ ) . Hence T λ is orthogonal to each of the subspaces L y with y Q ˜ λ ( Y ^ ) and therefore
T λ = λ Λ Q ˜ λ ( Y ^ ) c λ S λ = λ Λ Q ˜ λ ( Y ^ ) c λ S λ Π ^ Q ˜ λ ( Y ^ ) Y ^ = T λ Π ^ Q ˜ λ ( Y ^ ) Y ^ .
This implies (iii).
We finally show that (iii) implies (ii). For that purpose, we compute
Q ˜ λ ( X ^ f ( Y ^ ) ) = Tr ( X ^ f ( Y ^ ) T λ ) = Tr ( X ^ f ( Y ^ ) Π ^ Q ˜ λ ( Y ^ ) Y ^ T λ ) = f ( Q ˜ λ ( Y ^ ) ) ¯ ( X ^ T λ ) = f ( Q ˜ λ ( Y ^ ) ) ¯ Q ˜ λ ( X ^ ) ,
where we used for the third line that f ( Y ^ ) Π ^ Q ˜ λ ( Y ^ ) Y ^ = f ( Q ˜ λ ( Y ^ ) ) ¯ Π ^ Q ˜ λ ( Y ^ ) Y ^ . One then computes
E ρ ^ Q ( f ( Y ^ ) X ^ | Y ^ ) = y σ ( Y ^ ) λ Λ y Y ^ f ( Q ˜ λ ( Y ^ ) ) Q ˜ λ ( X ^ ) ¯ Q λ | y ( ρ ^ ) Π ^ y Y ^ = y σ ( Y ^ ) f ( y ) λ Λ y Y ^ Q ˜ λ ( X ^ ) ¯ Q λ | y ( ρ ^ ) Π ^ y Y ^ = f ( Y ^ ) E ρ ^ Q ( X ^ | Y ^ ) ,
which establishes (ii). □

5. Characterizing Quasiprobability Representations

In Section 4, we have associated a notion of conditional expectation to every quasiprobability representation ( Q , Q ˜ ) that is Born-compatible with a CSCO Y ^ . In this section, we consider all quasiprobability representations that are compatible with two given complementary CSCO A ^ and B ^ (see Definition 5 for the notion of complementarity of CSCO).
In Section 5.1, we recall the definition of the Kirkwood–Dirac representations Born-compatible with A ^ and B ^ . In Section 5.2, we show that those KD representations are the only quasiprobability representations for which the associated conditional expectation (given A ^ or given B ^ ) satisfies the “pull-out formula” (see Theorem 7). This structural result provides a unique characterization of the KD quasiprobability representations that allows us to show Theorem 1, which states that the (left/right) Kirkwood–Dirac representation is the only quasiprobability representation for which the associated conditional expectations (given A ^ or given B ^ ) agree with the (left/right) conditional expectations defined as best estimators in Section 3.
In Section 5.3, we show that all A ^ and B ^ Born-compatible quasiprobability representations with ontic space Λ = σ ( A ^ ) × σ ( B ^ ) are completely determined by their associated conditional expectations (Theorem 9). This result provides a strengthening of Theorem 1 and allows for some further applications that we discuss.
Section 5.2 and Section 5.3 provide independent proofs of Theorem 1 and can be read in any order.

5.1. The Kirkwood–Dirac Distribution

The definition of a KD quasiprobability representation of quantum mechanics depends on the choice of two bases in the Hilbert space [4,7,8,23,24,31]. Let A ^ = ( A ^ 1 , , A ^ n ) and B ^ = ( B ^ 1 , , B ^ m ) be two CSCO. We write ( φ a A ^ ) a σ ( A ^ ) , ( φ b B ^ ) b σ ( B ^ ) for the corresponding eigenbases:
A ^ i φ a A ^ = a i φ a A ^ , B ^ j φ b B ^ = b j φ b B ^ .
We use d i A ^ and d j B ^ to denote the number of eigenvalues of A ^ i and B ^ j for each i 1 , n , respectively, j 1 , m . For any p = ( p 1 , , p n ) N n , we write
A ^ p : = A ^ 1 p 1 A ^ n p n ,
and for any a = ( a 1 , , a n ) σ ( A ^ ) , we write
a p : = a 1 p 1 a n p n .
We define Γ A ^ as
Γ A ^ = 0 , d 1 A ^ 1 × × 0 , d n A ^ 1 ,
and Γ B ^ is obtained analogously.
We consider, as in Equation (42), the maximal commutative algebras
F C , A ^ = { f ( A ^ ) | f : σ ( A ^ ) C } , F C , B ^ = { f ( B ^ ) | f : σ ( B ^ ) C } .
Several times, we will use the fact that, as we are working on a finite dimensional Hilbert space, any f ( A ^ ) F C , A ^ can be written as a polynomial f ( A ^ ) = p Γ A ^ c p A ^ p , where ( c p ) p Γ A ^ C . This is a direct consequence of the Lagrange interpolation. Of course, a similar statement holds for the elements of F C , B ^ .
We have the following lemma:
Lemma 5.
Let A ^ , B ^ be two CSCO , as above. Consider the statements
(i) 
( a , b ) σ ( A ^ ) × σ ( B ^ ) , φ a A ^ , φ b B ^ 0 ;
(ii) 
span C F C , A ^ F C , B ^ = L ( H ) = span C F C , B ^ F C , A ^ ;
(iii) 
F C , A ^ F C , B ^ = C I d .
Then, ( i ) ( i i ) ( i i i ) .
Proof. 
That ( i ) implies ( i i ) follows from the fact that Π ^ a A ^ Π ^ b B ^ , with ( a , b ) σ ( A ^ ) × σ ( B ^ ) form a basis of L ( H ) and from the observation that Π ^ a A ^ F C , A ^ , Π ^ b B ^ F C , B ^ .
To see that ( i i ) implies ( i ) , let us note that
f ( A ^ ) g ( B ^ ) = ( a , b ) σ ( A ^ ) × σ ( B ^ ) f ( a ) g ( b ) Π ^ a A ^ Π ^ b B ^ .
Hence,
span C F C , A ^ F C , B ^ = span C { Π ^ a A ^ Π ^ b B ^ φ a A ^ , φ b B ^ 0 } ,
which implies the result.
We now show that ( i ) implies ( i i i ) . Let f ( A ^ ) F C , A ^ . We need to show that, if f ( A ^ ) F C , B ^ , then it is a multiple of the identity. But, since F C , B ^ = F C , B ^ , this means we need to show that if f ( A ^ ) F C , B ^ , then it is a multiple of the identity. But, f ( A ^ ) F C , B ^ is equivalent to [ f ( A ^ ) , Π ^ b B ^ ] = 0 , for all b σ ( B ^ ) , which is equivalent to
a , a σ ( A ^ ) , b σ ( B ^ ) , 0 = φ a A ^ , [ f ( A ^ ) , Π ^ b B ^ ] φ a A ^ = ( f ( a ) f ( a ) ) φ a A ^ , φ b B ^ φ b B ^ , φ a A ^ .
This shows that f ( A ^ ) = c I d for some c C . □
We point out that it is easy to construct two bases in dimension 3 for which condition ( i i i ) of Lemma 5 is fulfilled but condition ( i ) fails. This shows that those conditions are not equivalent.
Definition 5.
We say two complete sets of commuting observables A ^ ,   B ^ are complementary if they satisfy any one of conditions (i) and (ii) in Lemma 5.
The term “complementary” is used for a variety of different notions in the literature [39,55,56,57]. In the context of finite dimensional Hilbert space, it is sometimes used as synonymous to “mutually unbiased”, in other words to | a | b | = d 1 / 2 , for all a σ ( A ^ ) , b σ ( B ^ ) . Our definition here is therefore a relaxation of the latter definition. We refer to [57] for the link of this definition to various notions of incompatibility.
We now introduce the left and right Kirkwood–Dirac distributions associated with two CSCO A ^ and B ^ , assumed to be complementary. Two natural choices of Born-compatible frames for A ^ and B ^ are
S a , b = Π ^ a A ^ Π ^ b B ^ and S a , b r = Π ^ b B ^ Π ^ a A ^ .
It is then easily shown that the dual frames are
T a , b = 1 | φ a A ^ , φ b B ^ | 2 Π ^ a A ^ Π ^ b B ^ and T a , b r = 1 | φ a A ^ , φ b B ^ | 2 Π ^ b B ^ Π ^ a A ^ .
The associated quasiprobability representations are the left and right KD representations
Q a , b KD , ( ρ ^ ) = φ b B ^ , φ a A ^ φ a A ^ , ρ ^ φ b B ^ , Q ˜ a , b KD , ( X ^ ) = 1 φ a A ^ , φ b B ^ φ a A ^ , X ^ φ b B ^ Q a , b KD , r ( ρ ^ ) = φ a A ^ , φ b B ^ φ b B ^ , ρ ^ φ a A ^ , Q ˜ a , b KD , r ( X ^ ) = 1 φ b B ^ , φ a A ^ φ b B ^ , X ^ φ a A ^ .
For later reference, we note that, for all X ^ L ( H ) , one has that
Q ˜ a , b KD , ( X ^ ) ¯ = Q ˜ a , b KD , r ( X ^ )
They are called the left and right Kirkwood–Dirac representations because of the following fact. For f : σ ( A ^ ) C and g : σ ( B ^ ) C , one gets
Q ˜ a , b KD , ( f ( A ^ ) g ( B ^ ) ) = f ( a ) g ( b )
so the symbol is well-behaved when f ( A ^ ) is “to the left”, and
Q ˜ a , b KD , r ( g ( B ^ ) f ( A ^ ) ) = g ( b ) f ( a ) ,
the symbol is well-behaved when f ( A ^ ) is “to the right”.
We now analyze the right and left KD-conditional expectations of arbitrary observables, given A ^ or B ^ .
Lemma 6.
Given two complementary CSCO A ^ and B ^ , we have that, for every X ^ L ( H ) ,
E ρ ^ Q KD , ( X ^ | B ^ ) = E ρ ^ ( X ^ | B ^ ) , E ρ ^ Q KD , r ( X ^ | B ^ ) = E ρ ^ r ( X ^ | B ^ ) .
and
E ρ ^ Q KD , ( X ^ | A ^ ) = E ρ ^ r ( X ^ | A ^ ) , E ρ ^ Q KD , r ( X ^ | A ^ ) = E ρ ^ ( X ^ | A ^ ) .
Proof. 
Following Definition 4, one computes, for all ρ ^ D B ^ and any X ^ L ( H ) ,
E ρ ^ Q KD , ( X ^ | B ^ ) = ( a , b ) σ ( A ^ ) × σ ( B ^ ) 1 φ a A ^ , φ b B ^ φ a A ^ , X ^ φ b B ^ ¯ φ b B ^ , φ a A ^ φ a A ^ , ρ ^ φ b B ^ φ b B ^ , ρ ^ φ b B ^ Π ^ b B ^ = ( a , b ) σ ( A ^ ) × σ ( B ^ ) φ b B ^ , X ^ φ a A ^ φ a A ^ , ρ ^ φ b B ^ φ b B ^ , ρ ^ φ b B ^ Π ^ b B ^ = b σ ( B ^ ) φ b B ^ , X ^ ρ ^ φ b B ^ φ b B ^ , ρ ^ φ b B ^ Π ^ b B ^ .
One recognizes the left conditional expectation given in Definition 2, proving the first equality in Equation (139). The other three equalities follow from a similar computation. □
In conclusion, both the left and right KD representations are not only A ^ and B ^ compatible, but in addition, the associated conditional expectations given either A ^ or B ^ both can be interpreted as best estimators with respect to either the right or left sesquilinear forms introduced in Section 3. In particular, their associated conditional expectation satisfy the pull-out formula, see the first point in Theorem 5.

5.2. Characterization of the Kirkwood–Dirac Quasiprobability Representations via the Pull-Out Property

We shall prove in this section that the KD representations are the only A ^ and B ^ Born-compatible quasiprobability representations that give rise to Q-conditional expectations satisfying the pull-out property: this is the content of Theorem 7.
The main technical result we need is contained in the following proposition, which is formulated for the left KD representation. An analogous statement holds for the right KD representation.
Proposition 4.
Let H be a d dimensional Hilbert space and let ( Q , Q ˜ ) be a quasiprobability representation of quantum mechanics on H over a space Λ with Λ = d 2 . Suppose that there exists a CSCO   B ^ such that
(i) 
( Q , Q ˜ ) is B ^ -compatible;
(ii) 
X ^ L ( H ) , f ( B ^ ) F C , B ^ , ρ ^ D B ^ , E ρ ^ Q ( f ( B ^ ) X ^ | B ^ ) = f ( B ^ ) E ρ ^ Q ( X ^ | B ^ ) ;
Then, the following holds. If there exists a second CSCO   A ^ , complementary to B ^ , and for which ( Q , Q ˜ ) is A ^ -compatible, then there exists a bijective map Φ : Λ σ ( A ^ ) × σ ( B ^ ) such that for all ( a , b ) σ ( A ^ ) × σ ( B ^ )
ρ ^ D ( H ) , Q Φ 1 ( a , b ) ( ρ ^ ) = Q a , b KD , ( ρ ^ ) and X ^ L ( H ) , Q ˜ Φ 1 ( a , b ) ( X ^ ) = Q ˜ a , b KD , ( X ^ ) .
We illustrate this proposition in Appendix C where we compare the conditional expectations arising from the Gross-Wigner and the KD quasiprobability representations for a qutrit system and show that they are different. We show in particular that the conditional expectation associated to the Gross–Wigner representation does not satisfy the pull-out formula.
The pull-out formula ( ( i i ) above) is an intrinsic condition on the Q-conditional expectation that can be weakened in the presence of two complementary CSCO, as shown in the following lemma:
Lemma 7.
Suppose that A ^ and B ^ are complementary CSCO on a Hilbert space H and that ( Q , Q ˜ ) is a B ^ -compatible quasiprobability representation of quantum mechanics on H . Then, the following statements are equivalent for all ρ ^ D B ^ :
(i) 
p Γ A ^ , q Γ B ^ , E ρ ^ Q ( B ^ q A ^ p | B ^ ) = B ^ q E ρ ^ Q ( A ^ p | B ^ ) ;
(ii) 
X ^ L ( H ) , f ( B ^ ) F C , B ^ , E ρ ^ Q ( f ( B ^ ) X ^ | B ^ ) = f ( B ^ ) E ρ ^ Q ( X ^ | B ^ ) ;
Proof. 
We only need to prove that (i) implies (ii). By Lemma 5, one can assume that X ^ = g ( B ^ ) h ( A ^ ) . By Lagrange interpolation and linearity, one can therefore assume X ^ = g ( B ^ ) A ^ p , for p Γ A ^ . By Lagrange interpolation again, one has that f ( B ^ ) g ( B ^ ) = q Γ B ^ c q B ^ q . One has
E ρ ^ Q ( f ( B ^ ) X ^ | B ^ ) = E ρ ^ Q q Γ B ^ c q B ^ q A p | B ^ = q Γ B ^ c q B ^ q E ρ ^ Q ( A ^ p | B ^ ) = f ( B ^ ) g ( B ^ ) E ρ ( A ^ p | B ^ ) = f ( B ^ ) E ρ Q ( g ( B ^ ) A ^ p | B ^ ) = f ( B ^ ) E ρ ^ Q ( X ^ | B ^ ) ,
where we used (i) twice and the fact that g ( B ^ ) F C , B ^ can also be expressed as a polynomial. This proves the result. □
As a direct consequence of Proposition 4 and Lemma 7, we get the following characterization of the KD distribution, in terms of the pull-out formula with respect to powers of the CSCO A ^ and B ^ :
Theorem 7.
Let H be a d dimensional Hilbert space and let ( Q , Q ˜ ) be a quasi-probability representation of quantum mechanics on H over a space Λ with Λ = d 2 . Suppose that there exist A ^ and B ^ two complementary CSCO such that
(i) 
( Q , Q ˜ ) is A ^ and B ^ -compatible;
(ii) 
p Γ A ^ , q Γ B ^ , ρ ^ D B ^ , E ρ ^ Q ( B ^ q A ^ p | B ^ ) = B ^ q E ρ ^ Q ( A ^ p | B ^ ) ;
Then, there exists a bijective map Φ : Λ σ ( A ^ ) × σ ( B ^ ) such that for all ( a , b ) σ ( A ^ ) × σ ( B ^ )
ρ ^ D ( H ) , Q Φ 1 ( a , b ) ( ρ ^ ) = Q a , b KD , ( ρ ^ ) and X ^ L ( H ) , Q ˜ Φ 1 ( a , b ) ( X ^ ) = Q ˜ a , b KD , ( X ^ ) .
We shall now prove Proposition 4.
Proof of Proposition 4. 
As in the proof of Theorem 6, conditions (i) and (ii) of the proposition imply that, for all λ Λ and for all b σ ( B ^ ) , S λ = S λ Π ^ B ( λ ) B ^ , where we wrote B ( λ ) = Q ˜ λ ( B ^ ) . Consequently, S λ Π ^ b B ^ = S λ Π ^ B ( λ ) B ^ Π ^ b B ^ = 0 , when b B ( λ ) . Since ( Q , Q ˜ ) is both B ^ and A ^ compatible, and since A ^ and B ^ are complementary and complete, one concludes, via Theorem 6, for all ( a , b ) σ ( A ^ ) × σ ( B ^ ) , that
Π ^ b B ^ Π ^ a A ^ = λ Λ b B ^ λ Λ a A ^ S λ S λ = λ Λ b B ^ λ Λ a A ^ S λ Π ^ b B ^ S λ = λ Λ b B ^ λ Λ a A ^ Λ b B ^ S λ Π ^ b B ^ S λ = λ Λ b B ^ λ Λ a A ^ Λ b B ^ S λ S λ .
Since we assume that A ^ and B ^ are complementary, the left-hand side of this equality does not vanish for any ( a , b ) σ ( A ^ ) × σ ( B ^ ) . Hence, there exists λ Λ a A ^ Λ b B ^ such that Λ a A ^ Λ b B ^ . Consequently, writing A ( λ ) = Q ˜ λ ( A ^ ) as above, the map
Φ : = ( A , B ) : λ Λ ( A ( λ ) , B ( λ ) ) σ ( A ^ ) × σ ( B ^ )
is surjective and since Λ = d 2 = σ ( A ^ ) × σ ( B ^ ) , it is a bijection. It follows that
Λ a A ^ = b 1 , d Λ a A ^ Λ b B ^ , Λ a A ^ Λ b B ^ = 1 , Λ a A ^ = d .
We can therefore identify Λ with σ ( A ^ ) × σ ( B ^ ) , by considering ( A ( λ ) , B ( λ ) ) as coordinates on Λ . We can now conclude as follows. From Equation (141), we find
Π ^ b B ^ Π ^ a A ^ = Π ^ b B ^ Λ a A ^ Λ b B ^ S λ = Π ^ b B ^ S λ ( a , b ) ,
for the unique λ ( a , b ) Λ a A ^ Λ b B ^ . Since λ ( a , b ) Λ b B ^ , Theorem 6 (iii) implies that S λ ( a , b ) = S λ ( a , b ) Π ^ b B ^ . This in turn implies that S λ = Π ^ A ( λ ) A ^ Π ^ B ( λ ) B ^ .
This concludes the proof with Φ ( λ ) = ( A ( λ ) , B ( λ ) ) . □
This first characterization of the left KD distribution relies on an intrisic property of its quantum conditional expectation: among the A ^ and B ^ compatible quasiprobability representations of quantum mechanics, the only one for which the associated quantum conditional expectation satisfies a left pull-out property is the left Kirkwood–Dirac distribution.
We now prove Theorem 1: Lemma 6 proves that ( i i i ) implies ( i ) and ( i i i ) implies ( i i ) , after direct computations.
We prove that ( i ) implies ( i i i ) : from Theorem 6, we know that ( i ) of Theorem 1 implies that
X ^ L ( H ) , f ( B ^ ) F C , B ^ , E ρ ^ Q ( f ( B ^ ) X ^ | B ^ ) = f ( B ^ ) E ρ ^ Q ( X ^ | B ^ ) .
Thus, as ( Q , Q ˜ ) satisfies a left pull out property (given just above) and is A ^ and B ^ compatible, ( Q , Q ˜ ) satisfies the hypothesis of Proposition 4. This proves ( i i i ) . The fact that ( i i ) implies ( i i i ) also follows from the same reasoning.

5.3. Characterization of Quasiprobability Representations via Their Quantum Conditional Expectations

In this subsection, we show that A ^ and B ^ Born-compatible quasiprobability representations are uniquely determined by their associated conditional expectations. To do so, we combine the framework we developed in Section 4 and Section 5.1 and Section 5.2 with an idea first developed in [41].
Theorem 8.
Let A ^ and B ^ be two complementary CSCO. Let ( R , R ˜ ) be a A ^ and B ^ -compatible quasiprobability representation of quantum mechanics on σ ( A ^ ) × σ ( B ^ ) . Let ( Q , Q ˜ ) be a A ^ and B ^ -compatible quasiprobability representation of quantum mechanics on Λ, with Λ = d 2 . Then, the following statements are equivalent:
(i) 
for all ρ ^ D B ^ , for all p Γ A ^ ,
E ρ ^ Q ( A ^ p | B ^ ) = E ρ ^ R ( A ^ p | B ^ ) ;
(ii) 
for all ρ ^ D B ^ , for all f ( A ^ ) F C , A ^ , E ρ ^ Q ( f ( A ^ ) | B ^ ) = E ρ ^ R ( f ( A ^ ) | B ^ ) ;
(iii) 
for all ρ ^ D B ^ , for all ( p , q ) Γ A ^ × Γ B ^ ,
E ρ ^ Q ( A ^ p B ^ q | B ^ ) = E ρ ^ R ( A ^ p B ^ q | B ^ ) .
(iv) 
for all ρ ^ D B ^ , for all X ^ L ( H ) ,
E ρ ^ Q ( X ^ | B ^ ) = E ρ ^ R ( X ^ | B ^ ) .
(v) 
there exists a bijective map Φ : Λ σ ( A ^ ) × σ ( B ^ ) such that
( a , b ) σ ( A ^ ) × σ ( B ^ ) , R a , b = Q Φ 1 ( a , b ) ;
Proof. 
Let us prove first that ( v ) ( i v ) . Let S R and S Q be the respective frames associated to ( R , R ˜ ) and ( Q , Q ˜ ) . The fact that R a , b = Q Φ 1 ( a , b ) for all ( a , b ) σ ( A ^ ) × σ ( B ^ ) implies that S a , b R = S Φ 1 ( a , b ) Q for all ( a , b ) σ ( A ^ ) × σ ( B ^ ) . Moreover, from Equation (98), the symbols also satisfy R ˜ a , b = Q ˜ Φ 1 ( a , b ) for all ( a , b ) σ ( A ^ ) × σ ( B ^ ) . Recall from Lemma 4 that the fact that ( Q , Q ˜ ) is B ^ -Born compatible implies that ( Λ b B ^ ) b σ ( B ^ ) is a partition of Λ . In particular, we can define the projection π 0 : Λ σ ( B ^ ) which verifies that for all λ Λ , λ Λ π 0 ( λ ) B ^ . For b σ ( B ^ ) , we have, from B ^ -Born compatibility of ( R , R ˜ ) ,
Π ^ b B ^ = a σ ( A ^ ) S a , b R = a σ ( A ^ ) S Φ 1 ( a , b ) Q = λ Φ 1 ( σ ( A ^ ) × { b } ) S λ Q .
Moreover, from B ^ -compatibility of ( Q , Q ˜ ) ,
Π ^ b B ^ = λ Λ b B ^ S λ Q .
Taking the Hilbert–Schmidt inner product on these two expressions with the dual frame T Q yields that Φ is, for any b σ ( B ^ ) , a bijection between Λ b B ^ and σ ( A ^ ) × { b } . In particular, we have that π 0 ( Φ 1 ( a , b ) ) = b for all ( a , b ) σ ( A ^ ) × σ ( B ^ ) . Let X ^ L ( H ) and ρ ^ D B ^ . One has
E ρ ^ R ( X ^ | B ^ ) = ( a , b ) σ ( A ^ ) × σ ( B ^ ) R ˜ a , b ( X ^ ) ¯ R a , b ( ρ ^ ) φ b B ^ , ρ ^ φ b B ^ Π ^ b B ^ = ( a , b ) σ ( A ^ ) × σ ( B ^ ) Q ˜ Φ 1 ( a , b ) ( X ^ ) ¯ Q Φ 1 ( a , b ) ( ρ ^ ) φ π 0 ( Φ 1 ( a , b ) ) B ^ , ρ ^ φ π 0 ( Φ 1 ( a , b ) ) B ^ Π ^ π 0 ( Φ 1 ( a , b ) ) B ^ = λ Λ Q ˜ λ ( X ^ ) ¯ Q λ ( ρ ^ ) φ π 0 ( λ ) B ^ , ρ ^ φ π 0 ( λ ) B ^ Π ^ π 0 ( λ ) B ^ = b σ ( B ^ ) λ Λ b B ^ Q ˜ λ ( X ^ ) ¯ Q λ ( ρ ^ ) φ b B ^ , ρ ^ φ b B ^ Π ^ b B ^ = E ρ ^ Q ( X ^ | B ^ ) ,
proving ( i v ) .
It is clear that ( i v ) ( i i i ) ( i ) . Moreover, as already mentioned, any f ( A ^ ) F C , A ^ can be written as a polynomial P f ( A ^ ) and thus, ( i ) ( i i ) by linearity. We now prove that ( i ) ( v ) to conclude the proof. We compute, for all ρ ^ D B ^ , for all p Γ A ^ :
E ρ ^ Q ( A ^ p | B ^ ) = b σ ( B ^ ) λ Λ b B ^ Q ˜ λ ( ( A ^ p ) ) ¯ Q λ | b ( ρ ^ ) Π ^ b B ^ = b σ ( B ^ ) λ Λ b B ^ Q ˜ λ ( A ^ p ) ¯ Q λ | b ( ρ ^ ) Π ^ b B ^ = b σ ( B ^ ) λ Λ b B ^ Q ˜ λ ( A ^ ) p ¯ Q λ | b ( ρ ^ ) Π ^ b B ^ = b σ ( B ^ ) a σ ( A ^ ) λ Λ a A ^ Λ b B ^ Q ˜ λ ( A ^ ) p ¯ Q λ | b ( ρ ^ ) Π ^ b B ^ = b σ ( B ^ ) a σ ( A ^ ) λ Λ a A ^ Λ b B ^ a p ¯ Q λ | b ( ρ ^ ) Π ^ b B ^ = b σ ( B ^ ) a σ ( A ^ ) a p λ Λ a A ^ Λ b B ^ Q λ | b ( ρ ^ ) Π ^ b B ^ .
where Q ˜ λ ( A ^ p ) = Q ˜ λ ( A ^ ) p follows directly from Lemma 4. We also have that
E ρ ^ R ( A ^ p | B ^ ) = b σ ( B ^ ) a σ ( A ^ ) a p R ( a , b ) | b ( ρ ^ ) Π ^ b B ^ .
Thus, for all polynomial P with for all i 1 , n deg i ( P ) d i A ^ 1 and all ρ ^ D B ^ , we have that
b σ ( B ^ ) , a σ ( A ^ ) P ( a ) λ Λ a A ^ Λ b B ^ Q λ | b ( ρ ^ ) = a σ ( A ^ ) P ( a ) R ( a , b ) | b ( ρ ^ ) .
We fix a σ ( A ^ ) . By Lagrange interpolation, there exists P a such that P a ( a ) = δ a , a for all a σ ( A ^ ) and deg i ( P ) d i A ^ 1 for all i 1 , n . Thus, we have that, for all ρ ^ D B ^
b σ ( B ^ ) , λ Λ a A ^ Λ b B ^ Q λ | b ( ρ ^ ) = R ( a , b ) | b ( ρ ^ ) .
Using Equations (99) and (113) multiplied by Tr ( ρ ^ Π ^ b B ^ ) , we obtain that, for all ( a , b ) σ ( A ^ ) × σ ( B ^ ) and for all ρ ^ D B ^
Tr ρ ^ λ Λ a A ^ Λ b B ^ S λ Q S a , b R = 0 .
Consequently, we conclude that, for all ( a , b ) σ ( A ^ ) × σ ( B ^ )
λ Λ a A ^ Λ b B ^ S λ Q = S a , b R
As S a , b R 0 , we have that for all ( a , b ) σ ( A ^ ) × σ ( B ^ ) , Λ a A ^ Λ b B ^ . Consequently, the map
Φ : Λ σ ( A ^ ) × σ ( B ^ ) λ ( Q λ ( A ^ ) , Q λ ( B ^ ) )
is surjective and as Λ = σ ( A ^ ) × σ ( B ^ ) = d 2 , Φ is a bijection. In particular, Λ a A ^ Λ b B ^ = 1 and this element is Φ 1 ( a , b ) for all ( a , b ) σ ( A ^ ) × σ ( B ^ ) . We finally obtain, by Equation (144), that for all ( a , b ) σ ( A ^ ) × σ ( B ^ )
S Φ 1 ( a , b ) Q = S a , b R
proving ( v ) . □
This theorem shows that a quasiprobability representation of quantum mechanics on a Hilbert space H that is Born-compatible for two complementary CSCO A ^ and B ^ , is completely determined by its associated quantum conditional expectations of A ^ p given B ^ with p Γ A ^ . Combining Theorem 8 with Lemma 6, and setting ( R , R ˜ ) = ( Q KD , , Q ˜ KD , ) , we immediately obtain the following theorem:
Theorem 9.
Let A ^ and B ^ be complementary CSCO . Let ( Q , Q ˜ ) be an A ^ and B ^ -compatible quasiprobability representation of quantum mechanics defined on a set Λ ( Λ = d 2 ) . Then, the following are equivalent:
(i) 
ρ ^ D B ^ , p Γ A ^ , E ρ ^ Q ( A ^ p | B ^ ) = E ρ ^ ( A ^ p | B ^ )
(ii) 
ρ ^ D A ^ , p Γ B ^ , E ρ ^ Q ( B ^ p | A ^ ) = E ρ ^ r ( B ^ p | A ^ )
(iii) 
There exists a bijective map Φ : Λ σ ( A ^ ) × σ ( B ^ ) such that for all ( a , b ) σ ( A ^ ) × σ ( B ^ )
ρ ^ D ( H ) , Q Φ 1 ( a , b ) ( ρ ^ ) = Q a , b KD , ( ρ ^ ) and X ^ L ( H ) , Q ˜ Φ 1 ( a , b ) ( X ^ ) = Q ˜ a , b KD , ( X ^ ) .
Theorem 9 provides a direct proof of Theorem 1 as it is a strenghtened version thereof: indeed, Theorem 9 relies on the hypothesis that the conditional expectations coincides on powers of A ^ or on powers of B ^ whereas Theorem 1 relies on the equality of the conditional expectation on all operators in L ( H ) .
The left Kirkwood–Dirac distribution is therefore the only quasiprobability representation of quantum mechanics that is A ^ and B ^ compatible and for which the quantum conditional expectation of A ^ p given B ^ is given by the “left weak-value” formalism, meaning by the values of Tr ( Π ^ b B ^ A ^ p ρ ^ ) Tr ( Π ^ b B ^ ρ ) b σ ( B ^ ) , p Γ A ^ . For the right Kirkwood–Dirac representation, it is characterized by the “right weak-value formalism”, meaning by the values of Tr ( ρ ^ A ^ p Π ^ b B ^ ) Tr ( Π ^ b B ^ ρ ) b σ ( B ^ ) , p Γ A ^ .
Theorem 8 can further be used to characterize any fixed A ^ and B ^ Born-compatible quasiprobability representation of quantum mechanics. If one briefly goes to the infinite dimensional setting for illustrative purposes, the Wigner representation is then uniquely determined by the associated conditional expectations of P ^ n given Q ^ that will not coincide with the ones arising from the weak value formalism. We illustrate this fact in Appendix C where we compare the KD and Wigner conditional expectations of P ^ n knowing Q ^ , for n = 1 , 2 , showing they are different when n = 2 . This also means that the conditional expectation given by the Wigner quasiprobability representation cannot be the one defined by minimization in Section 3.

6. Quantum Conditional Expectation and Parameter Estimation

In this section, we explore properties of the quantum conditional expectations E ρ ^ / r ( X ^ | Y ^ ) introduced in Section 3 in relation to the following well-known parameter estimation problem. Let ρ ^ be any quantum state and consider the state
ρ ^ X ^ ( θ ) : = e i θ X ^ ρ ^ e i θ X ^
evolved under the unitary flow generated by some observable X ^ . The problem is then to estimate the “phase” θ from the measurement of an appropriately chosen observable Y ^ in ρ ^ X ^ ( θ ) [37,45,58,59,60].
We first show that the variance of the imaginary part of the above quantum conditional expectation is equal to the classical Fisher information associated with this phase estimation problem (Proposition 5). This relation was previously established and elaborated upon in [37,45] within the context of weak value physics. We then apply the results of Section 3.3, and in particular the law of iterated expectations and the law of total variance to relate the mean and the variance of X ^ in the state ρ ^ X ^ ( θ ) to those of its best estimator E ρ ^ X ^ ( θ ) / r ( X ^ | Y ^ ) and of its conditional variance Δ ρ ^ X ( θ ) / r , 2 ( X ^ | Y ^ ) in the same state. This sharpens a result of [37] and corroborates the idea that E ρ ^ / r ( X ^ | Y ^ ) is a “best estimator” of X ^ .
These developments allow us to introduce and then study, in Section 6.3, the notion of phase insensitive state, which is a state for which the above Fisher information vanishes. We then show that, modulo some technical assumptions, KD-real states are phase insensitive (Proposition 6). More precisely, we will observe that if the metrological protocol described above involves an observable X ^ and and a state ρ ^ that are KD-real, with Y ^ = B ^ or Y ^ = A ^ , then their Fisher information I F ( Y ^ ; 0 ) vanishes. Such states and observables are therefore not useful for this particular phase estimation protocol. But, as noted in Section 6.2, they can provide simultaneous information on the Born probability distributions of non-commuting observables X ^ , Y ^ through a weak measurement protocol. This result further suggests that, if ρ ^ and X ^ are KD real, and if the associated quantum Fisher information is non-zero, then the logarithmic derivative L ^ of ρ ^ cannot be KD real. We show this to be the case when ρ ^ is pure and provide two families of examples where ρ ^ is mixed and L ^ is indeed not KD real.

6.1. The Imaginary Part of E ρ ^ ( X ^ | Y ^ ) : Fisher Information

If the observable Y ^ is measured when the system is in the state ρ ^ X ^ ( θ ) , the outcome y σ ( Y ^ ) is observed with probability
p ( y ; θ ) : = Tr ( Π ^ y Y ^ ρ ^ X ^ ( θ ) ) .
The Fisher information [61,62,63,64] of this θ -dependent probability is defined as the expected value of ( θ ln ( p ( y ; θ ) ) ) 2 :
I F ( Y ^ ; θ ) : = y ( θ ln ( p ( y ; θ ) ) ) 2 p ( y ; θ ) .
Note that this Fisher information depends not only on Y ^ and θ , but also on ρ ^ and X ^ , although we do not indicate this dependence in the notation. The Fisher information derives its importance principally from the famous Cramèr–Rao inequality which states that the variance V ( θ ˜ ) of any unbiased estimator θ ˜ ( Y ^ ) of θ is bounded from below by the inverse of the Fisher information:
V ( θ ˜ ) ( I F ( Y ^ ; θ ) ) 1 .
The Fisher information I F ( Y ^ ; θ ) is related to the (imaginary part of) the conditional expectation E ρ ^ / r ( X ^ | Y ^ ) , as follows [37,45]:
Proposition 5.
Let Y ^ be a CSCO , X ^ = X ^ L ( H ) and suppose that ρ ^ X ^ ( θ ) D Y ^ . Then
I F ( Y ^ ; θ ) = 4 Tr Im E ρ ^ X ^ ( θ ) / r ( X ^ | Y ^ ) 2 ρ ^ X ^ ( θ ) .
Proof. 
A straightforward computation yields [37,45]
θ ln p ( y ; θ ) = i Tr ( Π ^ y Y ^ [ X ^ , ρ ^ X ^ ( θ ) ] ) Tr ( Π ^ y Y ^ ρ ^ X ^ ( θ ) ) = 2 Im Tr ( Π ^ y Y ^ X ^ ρ ^ X ^ ( θ ) ) Tr ( Π ^ y Y ^ ρ ^ X ^ ( θ ) ) ,
where we used the self-adjointness of X ^ . This is the imaginary part of the coefficient of Π ^ y in E ρ ^ X ^ ( θ ) / r ( X ^ | Y ^ ) , see Definition 2. This proves Equation (150). □
Note that E ρ ^ ( Im ( E ρ ^ / r ( X ^ | Y ^ ) ) ) = 0 . Hence Equation (150) states that the classical Fisher information of the quantum mechanical Born probabilities p ( y ; θ ) equals (up to a factor 4) the variance of the imaginary part of the (left/right) quantum conditional expectation. This observation provides an operational meaning to the imaginary part of the conditional expectation E ρ ^ ( X ^ | Y ^ ) . Since in classical probability theory the conditional expectation of a real random variable is always real, the presence of an imaginary part to E ρ ^ ( X ^ | Y ^ ) can be considered as a typical quantum feature. This raises the question precisely when Im E ρ ^ ( X ^ | Y ^ ) vanishes, depending on ρ ^ , X ^ , and Y ^ . From Equation (51) one obtains that
Im E ρ ^ ( X ^ | Y ^ ) = 1 2 y σ ( Y ^ ) φ y Y ^ , [ X ^ , ρ ^ ] φ y Y ^ φ y Y ^ , ρ ^ φ y Y ^ Π ^ y Y ^ = Im E ρ ^ r ( X ^ | Y ^ ) .
Hence, if
[ ρ ^ , X ^ ] = 0 , or [ ρ ^ , Π ^ y Y ^ ] = 0 , or [ Π ^ y Y ^ , X ^ ] = 0 ,
then Im E / r ( X ^ | Y ^ ) = 0 . This, in turn, implies that the Fisher information I F ( Y ^ ; 0 ) vanishes. One may therefore conclude that the absence of incompatibility between ρ ^ , X ^ and Y ^ —which is a form of “classicality”—is a sufficient condition for the absence of an imaginary part to the quantum conditional expectation and hence of the vanishing of the Fisher information I F ( Y ^ ; 0 ) . We will see below that it is not a necessary condition, as the Fisher information can vanish in cases where the commutators of Equation (153) are all non-zero. An example is given in Equation (78). Generally, self-adjointness of the conditional expectation E ρ ^ ( X ^ | Y ^ ) does not preclude the manifestation of other nonclassical features associated to the triplet ( ρ ^ , X ^ , Y ^ ) : as we will see in the next section, anomalous values of E ρ ^ ( X ^ | Y ^ = y ) point to the possibility of weak value amplification, which is a typical quantum phenomenon as well.
In order to address the question under what circumstances the Fisher information vanishes, we introduce the notion of phase insensitive states:
Definition 6.
Given an observable X ^ and a CSCO Y ^ , we say that a state ρ ^ D Y ^ is phase insensitive for a measurement of Y ^ if I F ( Y ^ ; 0 ) = 0 .
The phase insensitivity of a state ρ ^ depends both on the generator X ^ associated to the phase and on the observable Y ^ . We will simply say “ ρ ^ is phase insensitive” when the choice of X ^ and Y ^ is clear from the context; otherwise, we will say the triplet ( ρ ^ , X ^ , Y ^ ) is phase-insensitive. The operational meaning of phase insensitivity is discussed in some detail in Appendix D. We recall there that the Fisher information can be viewed as the squared relative rate of change of the probabilities p ( y ; θ ) , when the phase θ varies by a small amount δ θ . Comparing this quantity to the mean squared relative error δ p F 2 of the experimentally determined probabilities, one obtains an estimate on the minimal variation δ θ m i n that is experimentally detectable:
δ θ 2 δ θ m i n 2 : = δ p F 2 I F ( Y ^ ; θ ) .
For small, and a fortiori, for vanishing Fisher information, this minimal phase variation becomes very large, thereby justifying the definition of phase insensitivity. We refer to Appendix D for more details. We will show how to identify phase insensitive triplets ( ρ ^ , X ^ , Y ^ ) using the notion of KD-reality in Section 6.3, where we will also further explore the link with the quantum Fisher information.

6.2. An Additive Uncertainty Principle

As a direct application of the results of Section 3, and in particular of the quantum law of total variance derived there, we can now relate the variance of X ^ to the variances of the real and imaginary parts of its best estimator E ρ ^ X ^ ( θ ) / r ( X ^ | Y ^ ) and to the expected value of its conditional variance Δ ρ ^ X ^ ( θ ) / r , 2 ( X ^ | Y ^ ) . In what follows, for Z L ( H ) , we write | Z | = Z Z . Since Δ ρ ^ X ^ ( θ ) 2 X ^ = Δ ρ ^ 2 X ^ for any value of θ , it follows from the law of total variance in Proposition 2 combined with Equation (88) and Proposition 5 that, for any X ^ = X ^ L ( H ) and CSCO   Y ^ ,
Δ ρ ^ 2 X ^ = Δ ρ ^ X ( θ ) 2 E ρ ^ X ( θ ) sa ( X ^ | Y ^ ) + 1 4 I F ( Y ^ ; θ ) + Tr ( ρ ^ X ^ ( θ ) X ^ E ρ ^ X ( θ ) ( X ^ | Y ^ ) 2 ) ,
and
Δ ρ ^ 2 X ^ = Δ ρ ^ X ( θ ) 2 E ρ ^ X ( θ ) sa ( X ^ | Y ^ ) + 1 4 I F ( Y ^ ; θ ) + Tr ( ρ ^ X ^ ( θ ) X ^ E ρ ^ X ( θ ) r ( X ^ | Y ^ ) 2 ) ,
where we recall the notation E ρ ^ X ^ ( θ ) sa ( X ^ | Y ^ ) : = E ρ ^ X ^ ( θ ) / r , sa ( X ^ | Y ^ ) . As we saw in Section 3.3, the error term (last term) can be understood as the expected value of the conditional variance of X ^ . When ρ ^ is pure, this error term vanishes and one finds
Δ ρ ^ 2 X ^ = Δ ρ ^ X ( θ ) 2 E ρ ^ X ( θ ) sa ( X ^ | Y ^ ) + 1 4 I F ( Y ^ ; θ ) .
We therefore have, for all ρ ^
Δ ρ ^ 2 X ^ Δ ρ ^ X ( θ ) 2 E ρ ^ X ( θ ) sa ( X ^ | Y ^ ) + 1 4 I F ( Y ^ ; θ ) .
Equations (156) and (157) were first derived in [37] (at θ = 0 ), using a formulation in terms of weak values instead of in terms of conditional expectations. The identification of the error term in Equations (154) and (155), the concurrent link with the quantum law of total variance, and the interpretation of the error term as the expected value of a quantum conditional variance are, to the best of our knowledge, new here.
It is pointed out in [37] that one can think of Equation (156) (at θ = 0 ) as an (additive) uncertainty relation. Indeed, both terms in the right-hand side depend on Y ^ , whereas the left-hand side does not. So if one term is small, the other must be large, and vice versa; this result has the following interpretation. If, for a given ρ ^ and X ^ , Y ^ can be chosen such that the Fisher information is large, meaning close to (four times) the variance of X ^ , then, according to the Cramer–Rao bound, the variance on the estimation of θ can be made small, meaning one can obtain a good estimate on the phase θ from measurements of Y ^ in the state ρ ^ X ^ ( θ ) . In that case, however, E ρ ^ X ^ ( θ ) sa ( X ^ | Y ^ ) will have a small variance that will be far from the variance of X ^ . So then the fluctuations in the conditional expectation E ρ ^ X ^ ( θ ) sa ( X ^ | Y ^ ) do not provide much information on the probability distribution of X ^ for the state ρ ^ . For pure states ρ ^ , the extreme case occurs when Y ^ is taken to be equal to the symmetric logarithmic derivative L ^ of ρ ^ X ^ ( θ ) at θ = 0 , because then the Fisher information is maximal. This maximum Fisher information, which is the quantum Fisher information I QF ( 0 ) , then equals 4 Δ ρ ^ 2 X ^ (See Appendix E for details). On the other hand, if Y ^ is such that the Fisher information vanishes or is small, one can obtain only little information on the phase θ from measurements of Y ^ in the state ρ ^ X ^ ( θ ) . However, in that case, the variance of E ρ ^ X ( θ ) sa ( X ^ | Y ^ ) will be close or equal to the one of X ^ . Specifically, when the Fisher information vanishes and ρ ^ is pure, the former is equal to the latter (see Equation (156)). In conclusion, if ρ ^ is pure, we find that
min Y ^ Δ ρ ^ 2 E ρ ^ sa ( X ^ | Y ^ ) = Δ ρ ^ 2 E ρ ^ sa ( X ^ | L ^ ) = Δ ρ ^ 2 X ^ 1 4 I QF ( 0 ) = 0 , Δ ρ ^ 2 X ^ = max Y ^ Δ ρ ^ 2 E ρ ^ sa ( X ^ | Y ^ ) ,
where the maximum on the right is reached on phase insensitive triplets ( ρ ^ , X ^ , Y ^ ) . When ρ ^ is mixed, one has
min Y ^ Δ ρ ^ 2 E ρ ^ sa ( X ^ | Y ^ ) Δ ρ ^ 2 E ρ ^ sa ( X ^ | L ^ ) Δ ρ ^ 2 X ^ 1 4 I QF ( 0 ) Δ ρ ^ 2 X ^ = max Y ^ Δ ρ ^ 2 E ρ ^ sa ( X ^ | Y ^ ) ,
where the maximum is reached when Y ^ = X ^ , for example.
We note that the real and imaginary parts of the conditional expectations E ρ ^ , r ( X ^ | Y ^ ) can be experimentally accessed via a weak measurement procedure that we describe in Appendix A. We conclude that there is therefore a necessary trade-off between obtaining information on θ from p ( y ; θ ) , for which a large Fisher information is needed, or on the probability distribution of X ^ in ρ ^ through the weak measurement procedure of X ^ , conditioned on Y ^ , for which the Fisher information should optimally vanish, such that the triplet ( ρ ^ , X ^ , Y ^ ) is phase insensitive. Phase insensitivity will be related to KD reality in Section 6.3.
We finally point out that, for all ρ ^ and θ
0 I F ( Y ^ ; θ ) I QF ( θ ) 4 Δ ρ ^ , 2 X ^ ( x max x min ) 2 ,
where x max and x min are the largest and smallest, eigenvalues of X ^ , respectively, and we used Popoviciu’s inequality, which states that the variance of a random variable taking values in some bounded interval [ a , b ] is bounded from above by 1 4 ( b a ) 2 . For example, see [65], where this inequality is derived as a consequence of a stronger inequality which states that
1 4 I F ( Y ^ ; θ ) Δ ρ ^ , 2 X ^ ( x max E ρ ^ ( X ^ ) ) ( E ρ ^ ( X ^ ) x min ) ,
but for which one has to know the expected value of X ^ .

6.3. An Application: KD Reality Implies Phase Insensitivity

In this subsection, we exploit the link between the KD representations of quantum mechanics and the conditional expectations E ρ ^ / r ( X ^ | Y ^ ) (Theorem 1) with Y ^ = A ^ or Y ^ = B ^ in order to construct phase-insensitive triplets using the KD-real sector of a KD representation of quantum mechanics.
We have seen that, if any two of ρ ^ , X ^ , and Y ^ are compatible (commute), then the associated Fisher information vanishes and hence ρ ^ is phase insensitive. However, these conditions are very strong, and in fact rather trivial. For example, if [ ρ ^ , X ^ ] = 0 , then ρ ^ X ^ ( θ ) = ρ ^ is independent of θ and if [ X ^ , Π ^ y Y ^ ] = 0 , the two observables involved are compatible. We will now show that non-trivial examples of phase insensitivity arise naturally within the KD-real sector of KD representations.
As a first set of examples, consider the following situation. Let A ^ and B ^ be two complementary CSCO and suppose
X ^ = f 1 ( A ^ ) + g 1 ( B ^ ) , ρ ^ = f 2 ( A ^ ) + g 2 ( B ^ ) ,
where f 1 , g 1 are real-valued and f 2 , g 2 are positive. Then
[ X ^ , ρ ^ ] = [ f 1 ( A ^ ) , g 2 ( B ^ ) ] + [ g 1 ( B ^ ) , f 2 ( A ^ ) ] .
Note that this commutator will not vanish if the functions f i , g i are non-trivial. It nevertheless follows straightforwardly from Equation (152) that Im E ρ ^ ( X ^ | B ^ ) = 0 = Im E ρ ^ r ( X ^ | B ^ ) and hence that I F ( B ^ ; 0 ) = 0 and, similarly, I F ( A ^ ; 0 ) = 0 . In the following, we shall extend this family of examples by using the KD representation of quantum mechanics associated to A ^ and B ^ .
For that purpose, we define the space of left/right KD-real operators as
V KD , real / r = C ^ L ( H ) ( a , b ) σ ( A ^ ) × σ ( B ^ ) , Q ˜ a , b KD , / r ( C ^ ) R .
So, C ^ V KD , real / r if and only if C ^ has a real left/right KD-symbol. In view of Equation (136), one has that C ^ V KD , real if and only if C ^ V KD , real r . Consequently,
V KD , real sa = V KD , real L sa ( H ) = V KD , real r L sa ( H ) .
In short, V KD , real sa is the set of all observables with a real KD symbol. We will refer to V KD , real sa as the KD-real sector of quantum mechanics on H , and we will say that a state ρ ^ or an observable X ^ is KD real when it belongs to V KD , real sa .
Before turning to the study of KD-real pairs ( ρ ^ , X ^ ) , let us take a brief look at the special case where ρ ^ is KD positive and X ^ is KD real. Analogous to the discussion in Section 3.3, KD-positive states can be thought of as “classical”, at least to the extent that they allow for a joint probability distribution for A ^ and B ^ . Moreover, we know from Theorem 1 that the associated conditional expectations E ρ ^ Q KD ( X ^ | B ^ ) equal E ρ ^ ( X ^ | B ^ ) , for all X ^ and can be thought as a classical conditional expectation (see Definition 4). When in addition X ^ = A ^ and ρ ^ is KD positive, we are in the situation of Lemma 3: the KD distribution of ρ ^ is then a joint probability for A ^ and B ^ and satisfies (ii) of Lemma 3, such that in particular, E ρ ^ sa ( A ^ | B ^ ) does not have anomalous values. In addition, such positive states, when injected into a KD-positivity preserving quantum circuit, can be efficiently simulated classically [66,67,68]. Note however the following. Even if ρ ^ is KD-positive, and X ^ is KD-real, such that E ρ ^ Q KD ( X ^ | B ^ ) = E ρ ^ sa ( X ^ | B ^ ) is self-adjoint, it only satisfies the operator bound
min Q ˜ KD ( X ^ ) E ρ ^ sa ( X ^ | B ^ ) max Q ˜ KD ( X ^ ) ;
it is therefore still possible for E ρ ^ sa ( X ^ | B ^ ) to take on anomalous values, since nothing guarantees that x min Q ˜ ( X ^ ) x max . This is then still a form of nonclassicality, despite the positivity of the KD distribution of ρ ^ . In addition, there is, for general X ^ no straightforward notion of “marginals”. In conclusion, even if KD positive states have some classical features, they may still display quantum characteristics, depending on the problem considered.
We now turn to KD-real pairs ( ρ ^ , X ^ ) . Our first result, which is an immediate consequence of Lemma 6, shows that the KD-real states (and not only the KD-positive ones) are phase-insensitive provided X ^ is KD real and Y ^ = B ^ :
Proposition 6.
Let A ^ , B ^ be two complementary CSCO and let Q KD , / r be the corresponding left/right KD quasiprobability representations. Let X ^ = X ^ belong to V KD , real sa and ρ ^ belong to V KD , real sa D B ^ . Then, the left/right conditional expectations of X ^ , given B ^ , are self-adjoint such that
E ρ ^ sa ( X ^ | B ^ ) = E ρ ^ ( X ^ | B ^ ) = E ρ ^ r ( X ^ | B ^ ) .
Similarly, if ρ ^ belongs to V KD , real sa D A ^ ,
E ρ ^ sa ( X ^ | A ^ ) = E ρ ^ ( X ^ | A ^ ) = E ρ ^ r ( X ^ | A ^ ) .
Consequently, ρ ^ is phase insensitive for the measurement of B ^ and for the measurement of A ^ .
Proof. 
ρ ^ V KD , real sa implies that Q a , b KD , ( ρ ^ ) and therefore also Q a | b KD , ( ρ ^ ) : = Q a , b KD , ( ρ ^ ) / φ b B ^ , ρ ^ φ b B ^ are real. If X ^ = X ^ it follows from Lemma 6 that
E ρ ^ ( X ^ | B ^ ) = b a Q ˜ a , b KD , ( X ^ ) ¯ Q a | b KD , ( ρ ^ ) Π ^ b B ^ .
This operator is self-adjoint if X ^ has a real KD symbol. Moreover, E ρ ^ r ( X ^ | B ^ ) = E ρ ^ ( X ^ | B ^ ) = E ρ ^ ( X ^ | B ^ ) , proving Equation (166). Finally, Equation (150) implies that I F ( A ^ ; 0 ) = I F ( B ^ ; 0 ) = 0 , meaning that ρ ^ X ^ ( θ ) is phase insensitive for the measurement of B ^ and of A ^ at θ = 0 . □
The proposition provides an interpretation of the KD-real sector, as follows. If ρ ^ and X ^ belong to the KD-real sector, then the conditional expectations E ρ ^ / r ( X ^ | A ^ ) and E ρ ^ / r ( X ^ | B ^ ) are self-adjoint, which implies that the Fisher information I F ( A ^ ; 0 ) and I F ( B ^ ; 0 ) vanishes such that the state ρ ^ is phase insensitive for measurements of A ^ or of B ^ .
Consequently, in order to extract information about the phase θ from information about the family of states ρ ^ X ^ ( θ ) , one needs to perform a different measurement than the one of A ^ or of B ^ . This result is of interest because the space V KD , real sa can in a number of cases be described quite explicitly. As already pointed out above, the operators f ( A ^ ) + g ( B ^ ) , with f : σ ( A ^ ) R and g : σ ( B ^ ) R , always belong to V KD , real sa . It was proven in [69] that all operators in V KD , real sa are of this form with probability one when the eigenbases of A ^ and B ^ are chosen randomly from the uniform (Haar) distribution. In specific cases, the space V KD , real sa can be considerably larger, such as for the KD representation associated to the natural position-momentum Heisenberg representation of qdits (with d not a prime number) and of finite and certain second countable LCA groups [23,24]. Let us finally note that the space V KD , real sa is particularly simple when the transition matrix between A ^ and B ^ only has real entries: V KD , real sa is then the space of all real symmetric matrices [46]. This, occurs, for example, for n-qubit systems [66]. In fact, we note in passing that, quite generally, if X ^ and ρ ^ both have real symmetric matrices in the Y ^ basis, then I F ( Y ^ ; 0 ) = 0 , always, as can be seen directly from Equation (152).
Suppose X ^ and ρ ^ are fixed and belong to the KD-real sector. To estimate θ , a measurement different from A ^ or B ^ then has to be chosen. As is known, the measurement L ^ yielding the largest Fisher information, referred to as the quantum Fisher information,
I QF = I F ( L ^ , 0 ) = 4 Tr Im E ρ ^ X ^ ( 0 ) / r ( X ^ | L ^ ) 2 ρ ^ X ^ ( 0 ) ,
is provided by the symmetric logarithmic derivative L ^ of ρ ^ X ^ ( θ ) at θ = 0 . We refer to Appendix E for a self-contained discussion on the quantum Fisher information and its properties; see in particular Corollary A2. In view of what precedes, it is tempting to conjecture that, if ρ ^ and X ^ are both KD real, the optimal observable L ^ must not be KD-real itself, provided I QF does not vanish. For pure states, this conjecture is true, as proved in the proposition below. To that effect, we first prove the following lemma.
Lemma 8.
The following statements are equivalent:
1. 
I QF ( 0 ) = 0 ,
2. 
ρ ^ commutes with X ^ .
In other words, if ρ ^ and X ^ do not commute, then the triplet ( ρ ^ , X ^ , L ^ ) is not phase insensitive.
Proof. 
We use L ^ to denote the symmetric logarithmic derivative of ρ ^ ( θ ) = ρ ^ X ^ ( θ ) at θ = 0 . If I QF ( 0 ) = 0 then by definition, we have that
I QF ( 0 ) = Tr ( ρ ^ L ^ 2 ) = 0 .
We thus obtain that 0 = Tr ( ρ ^ L ^ 2 ) = Tr ( ( ρ ^ 1 / 2 L ^ ) ρ ^ 1 / 2 L ^ ) . Thus, ρ ^ 1 / 2 L ^ = 0 , which implies ρ ^ L ^ = 0 . The same argument proves that L ^ ρ ^ = 0 . Consequently, we obtain that
θ ρ ^ ( θ ) | θ = 0 = 1 2 { L ^ , ρ ^ ( 0 ) } = 1 2 ( L ^ ρ ^ + ρ ^ L ^ ) = 0
and by definition of ρ ^ ( θ ) , we obtain that i [ X ^ , ρ ^ ] = θ ρ ^ ( θ ) | θ = 0 = 0 and thus, X ^ and ρ ^ commute. Conversely, if [ X ^ , ρ ^ ] = 0 then we still have that
0 = i [ X ^ , ρ ^ ] = θ ρ ^ ( θ ) | θ = 0 = 1 2 ( L ^ ρ ^ + ρ ^ L ^ )
and consequently,
I QF ( 0 ) = Tr ( ρ ^ L ^ 2 ) = Tr 1 2 ρ ^ L ^ + L ^ ρ ^ L ^ = 0 ,
ending the proof. □
Proposition 7.
Let A ^ and B ^ be two complementary CSCO. Suppose that ρ ^ = ψ ψ D B ^ is a left KD-real pure state and that X ^ = X ^ is a left KD-real observable such that I QF ( 0 ) > 0 (or equivalently, that ψ is not an eigenvector of X ^ ) . Let L ^ be the symmetric logarithmic derivative of ρ ^ X ( θ ) at θ = 0 .
Then, the conditional expectation of L ^ given B ^ , E ρ ^ ( L ^ | B ^ ) , is non-zero and anti-Hermitian. In particular, L ^ is not left KD real.
Proof. 
Suppose, by contradiction, that E ρ ^ ( L ^ | B ^ ) = 0 . Then, as ρ ^ D B ^ , we have that, for all b σ ( B ^ ) :
φ b B ^ , L ^ ψ = 0 ,
meaning that L ^ ψ = 0 as B ^ is a CSCO. Moreover, by definition of L ^ , we have that
1 2 { L ^ , ρ ^ } = 1 2 { L ^ , ρ ^ X ^ ( 0 ) } = θ ρ X ^ ( θ ) | θ = 0 = i [ X ^ , ρ ^ ] .
Applying this equality with ρ ^ = | ψ ψ | we obtain
i X ^ ψ ψ , X ^ ψ ψ = 1 2 L ^ ψ + ψ , L ^ ψ ψ = 0 ,
as L ^ ψ = 0 . This implies that
X ^ ψ = ψ , X ^ ψ ψ ,
proving that ψ is an eigenvector of X ^ , which is a contradiction. Thus, E ρ ^ ( L ^ | B ^ ) 0 .
According to Proposition 6, as ρ ^ and X ^ are left KD real, we have that E ρ ^ ( X ^ | B ^ ) is self-adjoint, meaning that, for all b σ ( B ^ ) :
Im ( Tr ( Π ^ b B ^ X ^ ρ ^ ) ) = 0 .
Moreover, from Equation (173), we compute that, for all b σ ( B ^ ) :
Re Tr ( Π ^ b B ^ L ^ ρ ^ ) = Im Tr ( Π ^ b B ^ X ^ ρ ^ ) .
This shows that E ρ ^ sa , ( L ^ | B ^ ) = 0 , proving that E ρ ^ ( L ^ | B ^ ) is anti-hermitian.
Finally, as E ρ ^ ( L ^ | B ^ ) is both non-zero and anti-hermitian, and as ρ ^ is left KD real, Proposition 6 implies that L ^ cannot be left KD real. This concludes the proof. □
We now show that this conjecture also holds for mixed states in two examples.
For any state ρ ^ , there exists a CSCO A ^ such that ρ ^ = ρ ( A ^ ) for some ρ : σ ( A ^ ) R . If ρ ^ has simple eigenvalues, one can take ρ ^ = A ^ . Let B ^ be a CSCO complementary to A ^ . ρ ^ is then automatically KD real. For a σ ( A ^ ) , we use ρ a to denote the eigenvalue of ρ ^ associated to the vector φ a A ^ . From Equation (A65), one easily computes that a logarithmic derivative of ρ ^ X ^ ( θ ) at θ = 0 , L ^ , is given by
L ^ a , a = i ( ρ a ρ a ) ρ a + ρ a φ a A ^ , X ^ φ a A ^ , if ρ a + ρ a 0 0 otherwise ,
where L ^ a , a are the matrix coefficients of L ^ in the eigenbasis of A ^ . The KD symbol of L ^ is then given by
Q ˜ a , b KD , ( L ^ ) = i a , b ρ a + ρ a 0 ρ a ρ a ρ a + ρ a Q ˜ a , b KD , ( X ^ ) φ a A ^ , φ b B ^ φ a A ^ , φ b B ^ φ b B ^ , φ a A ^ φ a A ^ , φ b B ^ ,
( a , b ) σ ( A ^ ) × σ ( B ^ ) . From this expression we can give at least two situations in which the KD symbol of L ^ is purely imaginary:
1.
If the transition matrix between A ^ and B ^ has real coefficients and X ^ is KD real, then, the right-hand side of Equation (174) is purely imaginary. The condition that the transition matrix between A ^ and B ^ has real coefficients is for example fulfilled in systems of n qbits, where the transition matrix between A ^ and B ^ is given by the matrix of the Fourier transform on the abelian group Z / 2 Z n [23,66].
2.
If X ^ is of the form X ^ = f ( A ^ ) + Π ^ b 0 B ^ , where f : σ ( A ^ ) R and b 0 σ ( B ^ ) . It is clear in this case that X ^ belongs to the KD real sector. Moreover, it is easily computed in this case that for any a σ ( A ^ ) ,
Q ˜ a , b 0 KD , ( L ^ ) = i a ρ a + ρ a 0 ρ a ρ a ρ a + ρ a | φ b 0 B ^ , φ a A ^ | 2 ,
which is also purely imaginary.
One may conclude from this analysis that, in order to test the phase sensitivity of a KD-real state ρ ^ under the unitary flow generated by a KD-real observable X ^ , one cannot limit oneself to measuring the observable B ^ : indeed, under these hypotheses, the triplet ( ρ ^ , X ^ , B ^ ) is phase-insensitive. This analysis does not rule out the possibility that some KD-real observable C ^ exists for which I F ( C ^ , 0 ) > 0 . Nevertheless, we proved that the optimal choice L ^ of measurement observable is not KD real, provided ρ ^ is pure. We leave the question if whether this remains true for all mixed states open. Another way to state the above result is to say that, if a pure state ρ ^ and X ^ do not commute, then there does not exist a KD representation defined by two CSCO A ^ and B ^ for which L ^ = g ( B ^ ) and for which both ρ ^ and X ^ are KD real.
We end this section with a short remark on post-selected phase estimation, as studied in [42,43]. It is proven in those works that a quantum advantage can be obtained in that case, even though the state is KD real, but not KD positive. This does not contradict our findings for two reasons. First of all, the phase estimation problem considered there differs from the one considered here because of the post-selection process. As a result, the Fisher information is different as well. In addition, the relevant KD distribution used to obtain the results of [42,43] is different from the one considered here.

7. Discussion and Conclusions

There is elegance in defining concepts through optimization: the principles of least action and of maximal entropy are eminent examples. There is also convenience in defining concepts in quantum mechanics analogously to concepts in classical mechanics and probability theory. We have combined both approaches in order to define the notion of conditional expectation of an observable X ^ given an observable Y ^ in quantum mechanics: it is the best predictor of X ^ by a function of Y ^ in the sense that it is the function of Y ^ that minimizes a quadratic error in a given state ρ ^ . In the first part of this paper, we have shown that, because of the non-commutativity between observables inherent in quantum mechanics, several different choices of quadratic error mimicking the classical choice exist, and lead to different definitions. We have explained how two of these choices are particularly natural. First, they are intimately linked to weak value physics, thereby sharpening the known links between weak values and quantum conditional expectations. Second, our approach has allowed us to establish that the quantum conditional expectations so defined are characterized fully by a “pull out” property, and obey a law of iterated expectations as well as a law of total variance, similar to the original notion in probability theory. In this manner, the definition of the quantum conditional expectation through minimization of a quadratic error provides a unified picture within which the essential properties of the quantum conditional expectation naturally emerge. In addition, by working in analogy with classical probability theory, this approach has allowed us to highlight precisely under what circumstances the quantum object differs from its classical counterpart: in other words, it sheds light on aspects of the classical–quantum boundary. Notably, the conditional expectation of an observable is not necessarily an observable, but can fail to be self-adjoint, a feature that has no equivalent in the classical theory. Its imaginary part is thus a typical quantum feature. Its variance provides an extra term in the law of the total variance, absent in the classical theory. This imaginary part has an interpretation as the Fisher information of a phase estimation problem associated to the triplet ρ ^ , X ^ , Y ^ . We have also shown how the quantum conditional expectation can be used to establish a simple but efficient state-dependent no-go theorem for the existence of joint probabilities in quantum mechanics.
We conclude from this analysis that the proposed approach to quantum conditional expectations via minimization of a quadratic error provides a satisfying unifying picture that allows to analyze efficiently its properties and to pinpoint its similarities and—most importantly—its differences with the classical notion.
In the second part of the paper, we established a relation between these results and the theory of quasiprobability representations of quantum mechanics. The origin of the difficulties with the notion of conditional expectations in quantum mechanics is well known to be that joint probabilities of incompatible observables cannot be naturally defined. Quasiprobability representations of quantum mechanics circumvent this difficulty by allowing for non-positive and even non-real distributions to represent quantum states. This has led to a definition of joint quasiprobabilities and, through an analog of Bayes’ rule, to an alternative notion of conditional expectation from the one through minimization, that however depends very strongly on the choice of quasiprobability representation. Our main result in this paper is then that, among all possible quasiprobability representations that produce the correct Born probabilities associated to two chosen observables A ^ and B ^ , only the Kirkwood–Dirac quasiprobability representations reproduce the above conditional expectations defined through miminization of a quadratic error. This provides therefore a defining feature for the Kirkwood–Dirac representations, singling them out among all Born-compatible quasiprobability representations.
In the last part of the paper, we have applied our insights to provide an interpretation of the KD-real sector of quantum mechanics in terms of phase-insensitivity.
Our work leaves a number of questions unanswered. First, whereas it is certainly elegant and natural to require a quantum conditional expectation to obey a “pull-out” formula, as we did, we did not find a direct operational meaning to such a pull-out formula. Such an interpretation would certainly strengthen the case for the approach promoted here. As we pointed out, other approaches exist; see, for example, [48,49,70], where no pull-out property is required and [54], where in a more algebraic context, on the contrary, both the left and right pull-out properties are required, thereby effectively restricting the states for which the notion makes sense. Note, however, that giving up the “pull-out” formula implies giving up the link with weak-value physics, as we proved. Second, although a quantum conditional variance appears naturally within our approach, we did not identify an operational meaning for this quantity, contrary to what is known for the conditional expectations, which are linked to weak value physics. Indeed, although the expected value of this conditional variance plays an analogous role to the so-called “unexplained variation” familiar from statistics, no such interpretation seems to arise in an evident manner in quantum mechanics. It would be clearly of interest to develop such an interpretation. In addition, the definition of the conditional variance does not seem to be unique, several possibilities presenting themselves, that appear natural from different points of view.
Finally, to avoid technical complications and to allow for mathematical precision, we have restricted ourselves in this paper to finite probability spaces Ω and to finite dimensional Hilbert spaces H . While many of our results extend, at least on a formal level, to the infinite dimensional setting, technical difficulties arise that we have not tried to address; see [34] for quantum conditional expectations given a self-adjoint operator when the underlying state is pure.

Author Contributions

Conceptualization, M.S., C.L., R.B. and S.D.B.; Formal analysis, M.S., C.L., R.B. and S.D.B.; Investigation, M.S., C.L., R.B. and S.D.B.; Writing—original draft, M.S., C.L., R.B. and S.D.B.; Writing—review and editing, M.S., C.L., R.B. and S.D.B. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported in part by the CNRS through the MITI interdisciplinary programs. S. De Bièvre, C. Langrenez and M. Spriet acknowledge the support of the CDP C2EMPI, as well as the French State under the France-2030 programme, the University of Lille, the Initiative of Excellence of the University of Lille, the European Metropolis of Lille for their funding and support of the R-CDP-24-004-C2EMPI project.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Acknowledgments

The authors thank David Arvidsson-Shukur, Justin Dressel, and Andrew Jordan for stimulating and insightful discussions on the subject of this paper. They also thank Arthur Parzygnat as well as two anonymous referees for their constructive comments on (a previous version of) this paper.

Conflicts of Interest

The author declares no conflicts of interest.

Appendix A. Conditioned Measurements in Von Neumann’s Measurement Scheme

Appendix A.1. Quantum Measurements According to Von Neumann

In von Neumann’s model for the measurement of a quantum mechanical observable X ^ = X ^ acting on some Hilbert space H [71], X ^ is coupled to a measuring device or meter with pointer position q R . Both the system and the meter are treated quantum mechanically, and we let Q ^ and P ^ = i 1 q be the pointer’s position and momentum operator, acting on L 2 ( R ) , where Planck’s constant is set equal to 1. If the system is in an eigenstate of X ^ with eigenvalue x σ ( X ^ ) , the measuring device is assumed to cause the pointer to undergo a displacement γ x , where the constant γ is proportional to the interaction strength and to the duration of the system–meter interaction. The effect of the measurement is therefore modeled by the unitary transformation
U ^ γ : = e i γ X ^ P ^ ,
acting on the tensor product Hilbert space H L 2 ( R ) . Let ρ ^ be the initial state of the system and assume that the initial state of the pointer is a pure state defined by a normalized wave function φ = φ ( q ) L 2 ( R ) . In the weak value literature, φ is often taken to be a Gaussian, but we will allow it to be an arbitrary Schwartz function here, sometimes restricted to be real-valued to simplify formulas. This condition can be weakened to requiring that φ belong to the domain of suitable powers of P ^ and Q ^ , depending on the order of approximation in γ .
The initial state of the compound system is the tensor product state ρ ^ | φ φ | and its state after the system–meter interaction is given by
σ ^ γ : = σ ^ γ ( ρ ^ , φ ) : = U ^ γ ( ρ ^ | φ φ | ) U ^ γ .
The quantum mechanical expectation of the pointer position Q ^ , post-interaction, is then found to be
E σ ^ γ ( Q ^ ) = Tr ( ( I H Q ^ ) σ ^ γ ) = φ , Q ^ φ + γ Tr ( X ^ ρ ^ ) ,
and its variance post-interaction is given by
var γ ( Q ^ ) : = Δ σ ^ γ ( Q ^ ) = var φ ( Q ^ ) + γ 2 var ρ ^ ( X ^ )
where var φ is the variance associated to the state φ (denoted Δ φ 2 in the main text), as a direct computation starting from Equation (A12) shows. The first term on the right of (A1) is the initial mean position of the pointer, which we can always assume to be 0 by selecting the pointer’s zero reading, and Tr ( X ^ ρ ^ ) can in principle be inferred from measurements of Q ^ , provided γ is known. In other words, a first point to make here is that some information about the observable X ^ of the system, while it is in the state ρ ^ , is obtained from a measurement of Q ^ on the meter. Note however that, if the variance Δ Q ^ of Q ^ in the state ρ ^ is large, and γ Tr ( X ^ ρ ^ ) is small, then the signal-to-noise ratio will be unfavourable and, supposing γ is known, an accurate determination of Tr ( X ^ ρ ^ ) will require a very large number of individual measurements.
Similarly, if γ is not known, but Tr ( X ^ ρ ^ ) is, this procedure allows for the determination of γ , but with the same caveat when γ Tr ( X ^ ρ ^ ) is small. In that case, a conditional measurement of Q ^ can improve the signal-to-noise ratio, and lead to what is referred to as weak-value amplification, as we explain in the next subsection. The conditioning is based on the following procedure.
If Y ^ is a second observable on H , then Y ^ I L 2 ( R ) and Q ^ = I H Q ^ commute as operators on H L 2 ( R ) , unlike (in general) Y ^ and X ^ , and can be simultaneously measured using standard projective measurements. This then leads to a well-defined joint probability distribution of those two observables. In particular, simply writing Q ^ for I H Q ^ and Z ^ for Z ^ I L 2 ( R ) for operators Z ^ on H , the classical conditional expectation, post interaction, of Q ^ = I H Q ^ given that Y ^ = y with y σ ( Y ^ ) is, as we saw (see Remark 1(vi)), equal to the quantum conditional expectation E σ ^ γ ( Q ^ | Y ^ ) of Q ^ , given Y ^ , and is given by
E σ ^ γ ( Q ^ | Y ^ = y ) = Tr ( Π ^ y Y ^ Q ^ σ ^ γ ) Tr ( Π ^ y Y ^ σ ^ γ ) .
This conditional expectation can be experimentally determined by doing successive measurements of Q ^ and Y ^ on the state σ ^ γ and by post-selecting on Y ^ = y or, what is equivalent since we are dealing with commuting operators, by only measuring Q ^ when Y ^ = y , which can be implemented by passing the joint system through a filter which filters out the appropriate eigenstate of Y ^ before reading the meter position. We show in the next subsection that in the limit of small γ the real and imaginary parts of E ρ ^ ( X ^ | Y ^ = y ) determine the first-order correction to the value of E σ ^ γ ( Q ^ | Y ^ ) , for γ = 0 . When E ρ ^ ( X ^ | Y ^ ) is anomalous, meaning that it lies outside the range of the spectrum of X ^ , this can enhance the signal-to-noise ratio referred to above, a phenomenon referred to as weak-value amplification [31,36,45,50,51,52,53,72,73,74,75,76]. This phenomenon is considered of typical quantum nature since such anomalous values cannot occur classically.

Appendix A.2. Conditioned Measurements in the Weak Interaction Limit

In the analysis presented in this section only the left quantum conditional expectation intervenes. In order to lighten the notation, we therefore drop the “” upper-script. Let us define the quantum mechanical covariance of two observables C ^ 1 and C ^ 2 in a state σ ^ , all defined on some unspecified Hilbert space, by
cov σ ^ ( C ^ 1 , C ^ 2 ) : = Re Tr ( C ^ 1 C ^ 2 σ ^ ) Tr ( C ^ 1 σ ^ ) Tr ( C ^ 2 σ ^ ) = 1 2 Tr ( { C ^ 1 , C ^ 2 } σ ^ ) Tr ( C ^ 1 σ ^ ) Tr ( C ^ 2 σ ^ )
where we recall that { C ^ 1 , C ^ 2 } = C ^ 1 C ^ 2 + C ^ 2 C ^ 1 is the anti-commutator. If C ^ 1 = C ^ 2 = C ^ , we write cov σ ^ ( C ^ , C ^ ) = var σ ^ ( C ^ ) = Δ σ ^ 2 ( C ^ ) , the variance of C ^ . If the state is a pure state | φ φ | we will simply write cov φ instead of cov | φ φ | , and similarly for the variance.

Appendix A.2.1. Small γ Expansions of Conditional Expectations of Pointer Variables

We first compute the first-order expansion in γ of the conditional quantum expectations of Q ^ and of P ^ . Recall our standing assumption that φ is Schwartz class.
Proposition A1.
As γ 0 , we have that
E σ ^ γ ( Q ^ | Y ^ = y ) = φ , Q ^ φ + γ · Re E ρ ^ ( X ^ | Y ^ = y ) + 2 Im E ρ ^ ( X ^ | Y ^ = y ) · cov φ ( Q ^ , P ^ ) + O ( γ 2 ) ,
where cov φ ( Q ^ , P ^ ) : = Re Q ^ φ , P ^ φ φ , Q ^ φ φ , P ^ φ is the quantum mechanical covariance of the pointer position and momentum operators Q ^ and P ^ in the initial pointer state. Moreover, we also find that
E σ ^ γ ( P ^ | Y ^ = y ) = φ , P ^ φ + 2 Im E ρ ^ ( X ^ | Y ^ = y ) var φ ( P ^ ) · γ + O ( γ 2 ) ,
where var φ ( P ^ ) = cov φ ( P ^ , P ^ ) = P ^ φ , P ^ φ φ , P ^ φ 2 is the quantum mechanical variance of the pointer momentum operator in the initial pointer state φ.
The proof will show that more generally, for functions of the position operator Q ^ , if f is a real-valued differentiable function of at most polynomial growth, then
d d γ E σ ^ γ ( f ( Q ) | Y = y ) | γ = 0 = Re E ρ ^ ( X ^ | Y ^ = y ) · φ , f ( Q ^ ) φ + 2 Im E ρ ^ ( X ^ | Y ^ = y ) · cov φ ( f ( Q ^ ) , P ^ ) .
For the momentum operator P ^ , we find more generally that
d d γ E σ ^ γ ( f ( P ^ ) | Y ^ = y ) | γ = 0 = 2 Im E ρ ^ ( X ^ | Y ^ = y ) · cov φ ( f ( P ^ ) , P ^ ) .
Proof. 
For the proof of Equation (A5), we will simply compute the derivative at γ = 0 and show that
d d γ E σ ^ γ ( Q ^ | Y ^ = y ) | γ = 0 = Re E ρ ^ ( X ^ | Y ^ = y ) + 2 Im E ρ ^ ( X ^ | Y ^ = y ) · cov φ ( Q ^ , P ^ ) .
If p γ ( q , y ) is the joint probability (density) of having pointer position Q ^ = q while Y ^ = y , then the conditional pointer position density is given by
p γ ( q | y ) : = p γ ( q , y ) p γ ( y ) , p γ ( y ) = R p γ ( q , y ) d q ,
and
E σ ^ γ ( Q ^ | Y ^ = y ) = R q p γ ( q | y ) d q .
We compute p γ ( q , y ) and its derivative with respect to γ . Since e i γ X ^ P ^ = x σ ( X ^ ) e i γ x P ^ Π ^ x X ^ , we have
σ ^ γ = x , x σ ( X ^ ) Π ^ x X ^ ρ ^ Π ^ x X ^ | φ ( · γ x ) φ ( · γ x ) | .
It follows that
p γ ( q , y ) = x , x Tr ( Π ^ y Y ^ Π ^ x X ^ ρ ^ Π ^ x X ^ ) φ ( q γ x ) φ ( q γ x ) ¯ .
The derivative with respect to γ in γ = 0 is
γ p γ ( q , y ) | γ = 0 = x , x σ ( X ^ ) x Tr ( Π ^ y Y ^ Π ^ x X ^ ρ ^ Π ^ x X ^ ) q φ ( q ) φ ( q ) ¯ + x Tr ( Π ^ y Y ^ Π ^ x X ^ ρ ^ Π ^ x X ^ ) φ ( q ) q φ ( q ) ¯ = Tr ( Π ^ y Y ^ X ^ ρ ^ ) q φ ( q ) φ ( q ) ¯ Tr ( Π ^ y Y ^ ρ ^ X ^ ) φ ( q ) q φ ( q ) ¯ = 2 Re Tr ( Π ^ y Y ^ X ^ ρ ^ ) φ ( q ) ¯ q φ ( q ) ,
and since φ ( q ) ¯ q φ ( q ) = 1 2 q | φ ( q ) | 2 + i Re φ ( q ) ¯ P ^ φ ( q ) , this equals
γ p γ ( q , y ) | γ = 0 = Re ( Tr ( Π ^ y Y ^ X ^ ρ ^ ) ) q | φ ( q ) | 2 + 2 Im ( Tr ( Π ^ y Y ^ X ^ ρ ^ ) ) Re ( φ ( q ) ¯ P ^ φ ( q ) ) .
Integrating with respect to q,
γ p γ ( y ) = γ p ( q , y ) d q = 2 Im ( Tr ( Π ^ y Y ^ X ^ ρ ^ ) ) φ , P ^ φ ,
Finally, differentiating (A10) while using that p 0 ( q , y ) = | φ ( q ) | 2 Tr ( Π ^ y Y ^ ρ ^ ) and p 0 ( y ) = Tr ( Π ^ y Y ^ ρ ^ ) , we find
γ p ( q | y ) | γ = 0 = Re E ρ ^ ( X ^ | Y ^ = y ) q | φ | 2 + 2 Im E ρ ^ ( X ^ | Y ^ = y ) Re ( φ ¯ P ^ φ ) 2 Im E ρ ^ ( X ^ | Y ^ = y ) φ , P ^ φ | φ | 2 .
and (A9) follows from (A11) after an integration by parts. Equation (A8) is proven similarly by differentiating E σ ^ γ ( f ( Q ^ ) | Y ^ = y ) = f ( q ) p γ ( q | y ) d q .
We next turn to the proof of Equation (A6). This equation is shown as before by computing the derivative with respect to γ , a computation that we will carry in momentum representation. If we (somewhat idiosyncratically, but the letter p is already used to designate probability densities) denote the conjugate (or dual) variable of q by k and let φ ^ ( k ) be the Fourier transform of φ (normalized so as to be unitary on L 2 ( R ) ), then the probability density that P ^ = k while Y ^ = y is given by
p γ ( k , y ) = x , x Tr ( Π ^ y Y ^ Π ^ x X ^ ρ ^ Π ^ x X ^ ) e i γ ( x x ) k | φ ^ ( k ) | 2 ,
and
γ p γ ( k , y ) | γ = 0 = 2 Im Tr ( Π ^ y Y ^ X ^ ρ ^ ) k | φ ^ ( k ) | 2 .
If p γ ( k | y ) = p γ ( k , y ) / p γ ( y ) , with p γ ( y ) = p γ ( k , y ) d k , then we find that
γ p γ ( k | y ) = Im E ρ ^ ( X ^ | Y ^ = y ) ( k | φ ^ ( k ) | 2 φ , P ^ φ | φ ^ ( k ) | 2 )
(note that γ p γ ( y ) | γ = 0 = 2 Im Tr ( Π ^ y Y ^ X ^ ρ ^ ) k | φ ^ ( k ) | 2 d k ). The theorem follows by differentiating E σ ^ γ ( P ^ | Y ^ = y ) = k p γ ( k | y ) d k under the integral sign. Similarly, for (A8) by differentiating f ( k ) p γ ( k | y ) d k .
As we saw, therefore, the imaginary part of the weak values/quantum conditional expectations manifest themselves in the conditional Q ^ and P ^ measurements of the meter, provided the latter is prepared in appropriate states [45].
Equation (A5), with E ρ ^ ( X ^ | Y ^ = y ) replaced by its explicit expression Tr ( Π ^ y Y ^ X ^ ρ ^ ) / Tr ( Π ^ y Y ^ ρ ^ ) goes back to the paper by Aharonov et al. [36], where the term weak value is introduced. Equation (A6) has also been noted before [45]. The measurements of Q ^ or of P ^ are said to be “weak” when the post-interaction state σ ^ γ is not much perturbed from the initial state. This will clearly happen in the extreme case when | φ = | p = 0 is the momentum eigenstate with zero momentum, independently of how large γ is. Since the momentum eigenstates are not physical, in practice, one mostly considers a centered Gaussian state with a large Q ^ -variance Δ , which is therefore sharply localized in momentum about p = 0 . This leads then to a small perturbation of the initial state ρ ^ | φ φ | of the system plus probe under U ^ γ , provided γ is not too large, depending on Δ . For a detailed analysis of the conditions needed, we refer to [74]: a rule of thumb is that one needs γ ( x max x min ) Δ . For our purposes here, we only consider Taylor expansions of the relevant observers, to leading order, without explicitly controlling the error terms.
The definition of weak value as introduced in the literature depends on the measurement protocol described above. It has, a posteriori , been interpreted as a conditional expectation because of its appearance in the low order expansions in γ of the classical conditional expectations associated to the joint probability distribution of the commuting observables Q ^ and Y ^ . Our perspective is different: we have defined the quantum conditional expectation E ρ ^ ( X ^ | Y ^ ) as the solution to an appropriate minimization problem, in direct analogy with the conditional expectation in classical probability. We then showed that the conditional expectation so defined has the natural properties expected from a conditional expectation. Finally, Equations (A5) and (A8) give an operational meaning to this abstract definition.
The “weak measurement” terminology is somewhat confusing since the two measurements that are performed, of the meter position Q ^ and of the system variable Y ^ are both projective, hence “strong” measurements, but it is the interaction between system and meter that is weak. Also, as a result, the information about the probability distribution p ( x ) = Tr ( Π ^ x X ^ ρ ^ ) obtained in this manner is partial, and the effect of the measurement on the system limited, or “weak”.
We can go one step further and compute the term of order γ 2 for the conditional Q ^ measurements. We define
E 2 ( X ^ ; ρ ^ , y ) : = Tr Π ^ y Y ^ X ^ ρ ^ X ^ Tr ( Π ^ y Y ^ ρ ^ ) 0 .
Note that E 2 ( X ^ ; ρ ^ , y ) already appears in [47]. Its physical/probabilistic interpretation is not clear, except in the two following examples: first, when [ X ^ , ρ ^ ] = 0 or when [ X ^ , Y ^ ] = 0 , then
E 2 ( X ^ ; ρ ^ , y ) = E ρ ^ ( X ^ 2 | Y = y )
Second, when ρ ^ = | ψ ψ | is a pure state, and Y ^ is a CSCO. In that case, Π ^ y Y ^ is the projection onto the joint eigenvector φ y , and
E 2 ( X ^ ; ρ ^ , y ) = Tr | φ y φ y | X ^ | ψ ψ | X ^ | φ y , ψ | 2 = | φ y , X ^ ψ | 2 | φ y , ψ | 2 = E | ψ ψ | ( X ^ | Y ^ = y ) 2 ,
as already observed in [47]. Note that this is no longer true if Π ^ y Y ^ is of rank 2 or larger.
Proposition A2.
If φ is real-valued and as γ 0 , we have that
E σ ^ γ ( Q ^ | Y ^ = y ) = φ , Q ^ φ + Re ( E ρ ^ ( X ^ | Y ^ = y ) ) · γ γ 2 · Re E ρ ^ ( X ^ 2 | Y ^ = y ) E 2 ( X ^ ; ρ ^ , y ) cov φ ( P ^ 2 , Q ^ ) + O ( γ 3 )
Note that, when [ X ^ , ρ ^ ] = 0 or when [ X ^ , Y ^ ] = 0 , the second-order coefficient in Equation (A19) is zero, meaning that Equation (A5) is in fact a second-order approximation in this case. We also note that, when ρ ^ = | ψ ψ | is a pure state and Y ^ is a CSCO, the second-order coefficient in Equation (A19) can be interpreted as the real part of a conditional variance, see Equation (89).
Proof. 
We have already seen that for real-valued φ ,
γ p γ ( q , y ) | γ = 0 = Re Tr ( Π y Y ^ X ^ ρ ^ ) q φ 2 ( q ) = p 0 ( y ) Re E 1 q φ 2 ( q ) ,
since p 0 ( y ) = Tr ( Π ^ y Y ^ ρ ^ ) . Differentiating (A13) twice gives
γ 2 p γ ( q , y ) | γ = 0 = 2 Re Tr ( Π y Y ^ X ^ 2 ρ ^ ) φ ( q ) q 2 φ ( q ) + Tr ( Π y Y ^ X ^ ρ ^ X ^ ) q φ ( q ) 2 = 2 p 0 ( y ) Re E 2 φ ( q ) q 2 φ ( q ) + E q φ ( q ) 2
where E ν = E ρ ^ ( X ^ ν | Y ^ = y ) and E = E 2 ( X ^ ; ρ ^ , y ) . The second-order Taylor expansion in γ of p γ ( q , y ) at γ = 0 is thus given by
p γ ( q , y ) = p 0 ( y ) φ 2 ( q ) γ · Re E 1 q φ 2 ( q ) + γ 2 Re E 2 φ ( q ) q 2 φ ( q ) + E q φ ( q ) 2 + O ( γ 3 ) .
Integrating this expression with respect to q, we find
p γ ( y ) = R p γ ( q , y ) d q = p 0 ( y ) 1 + γ 2 E Re E 2 φ , P 2 φ + O ( γ 3 ) ,
since
R q φ 2 ( q ) d q = 0 and R φ ( q ) q 2 φ ( q ) d q = R ( q φ ( q ) ) 2 d q = P φ , P φ = φ , P 2 φ .
We next compute the Taylor expansion of R q p γ ( q , y ) d q :
R q p γ ( q , y ) d q = p 0 ( y ) φ , Q φ + γ · Re ( E 1 ) + γ 2 · ( E Re ( E 2 ) φ , P Q P φ + O ( γ 3 )
where we used, as before, that
R q q φ 2 ( q ) d q = R φ 2 ( q ) d q = 1 ,
while
R q φ ( q ) q 2 φ ( q ) d q = R q ( q φ ( q ) ) 2 d q = P φ , Q P φ = φ , P Q P φ .
Putting everything together, we finally conclude that
E σ ^ γ ( Q ^ | Y ^ = y ) = φ , Q φ + γ · Re ( E 1 ) + γ 2 · ( E Re ( E 2 ) φ , P Q P φ + O ( γ 3 ) 1 + γ 2 E Re E 2 φ , P 2 φ + O ( γ 3 ) = φ , Q φ + γ · Re ( E 1 ) + γ 2 E Re E 2 φ , P Q P φ φ , P 2 φ φ , Q φ + O ( γ 3 ) = φ , Q φ + γ · Re ( E 1 ) + γ 2 E Re E 2 cov φ ( P 2 , Q ) + O ( γ 3 )
where we used that
φ , P Q P φ = 1 2 φ , ( P 2 Q + Q P 2 ) φ φ , [ P , [ P , Q ] ] φ = Re φ , P 2 Q φ .
This ends the proof. □
We note that if φ is even, and in particular if it is a centered Gaussian, cov φ ( P ^ 2 , Q ^ ) = 0 and the approximation of the mean pointer position by γ Re E ρ ^ ( X ^ | Y ^ = y ) is of third-order in γ .

Appendix A.2.2. Small γ Behavior of the Conditional Variance of the Pointer Position

We examine the small- γ behavior of the conditional variance of Q ^ , which is given by
var σ ^ γ ( Q ^ | Y ^ = y ) : = E σ ^ γ ( Q ^ 2 | Y ^ = y ) E σ ^ γ ( Q ^ | Y ^ = y ) 2 .
Equation (A7) allows us to quickly compute the first-order expansion in γ of this conditional variance. Its derivative
d d γ var σ ^ γ ( Q ^ | Y ^ = y ) = d d γ E σ ^ γ ( Q ^ 2 | Y ^ = y ) 2 E σ ^ γ ( Q ^ | Y ^ = y ) d d γ E σ ^ γ ( Q ^ | Y ^ = y )
in γ = 0 becomes, on using (A7),
Re ( E 1 ) · 2 φ , Q ^ φ + 2 Im ( E 1 ) · cov φ ( Q ^ 2 , P ^ ) 2 φ , Q ^ φ E σ ^ 0 ( Q ^ | Y ^ = y ) Re ( E 1 ) + 2 Im ( E 1 ) cov φ ( Q ^ , P ^ ) = 2 Im ( E 1 ) · cov φ ( Q ^ 2 , P ^ ) 2 φ , Q ^ φ cov φ ( Q ^ , P ^ ) ,
where we recall that E 1 stands for the quantum conditional expectation E ( X ^ | Y ^ = y ) . We therefore have the following corollary, providing yet another interpretation for the imaginary part of the weak value:
Proposition A3.
Assuming for simplicity that the initial pointer position φ , Q ^ φ = 0 , the pointer’s conditional variance has the weak interaction expansion
var σ ^ γ ( Q ^ | Y ^ = y ) = φ , Q ^ 2 φ + 2 γ Im E ρ ^ ( X ^ | Y ^ = y ) · cov φ ( Q ^ 2 , P ^ ) + O ( γ 2 )
The corresponding expression when φ , Q ^ φ 0 follows on replacing Q ^ by Q ^ φ , Q ^ φ . This corollary is interesting in that the (empirical) conditional variance of Q ^ can be directly computed from experimental pointer position data alongside with the conditional expectation of Q ^ itself. It provides an estimator for the imaginary part of the weak value, provided that cov φ ( Q ^ 2 , P ^ ) 0 . This covariance is 0 when φ is real but, perhaps somewhat surprisingly, also for all Gaussian wave functions φ , even if these are complex, as a direct computation shows. In fact, if φ = R e i S then
cov φ ( Q ^ 2 , P ^ ) = q 2 Re ( φ ¯ P φ ) d q q 2 | φ | 2 d q Re ( φ ¯ P φ ) = q 2 q S R 2 d q q 2 R 2 d s q S R 2 d q ,
since Re ( φ ¯ P φ ) = Im ( φ ¯ q φ ) = ( q S ) R 2 . This is clearly 0 if φ is real-valued or if S is linear and q S is a constant. A complex Gaussian wave function can be written in the form
φ = C e 1 2 α ( q q 0 ) 2 + i p 0 q ,
with Re α > 0 , where C = C Re α is a normalization constant and q 0 = φ , Q ^ φ and p 0 = φ , P ^ φ are the mean position respectively momentum. If q 0 = 0 then q S = Im α q + p 0 , R 2 = e Re α q 2 and cov φ ( Q ^ 2 , P ^ ) is easily found to be 0. We note in passing, with reference to (A5), that cov φ ( Q ^ , P ^ ) = Im α / 2 Re α for such a Gaussian.
If cov φ ( Q ^ 2 , P ^ ) = 0 , then var σ ^ γ ( Q ^ | Y ^ = y ) = var φ ( Q ^ ) + O ( γ 2 ) , showing that to first order the system-meter interaction does not alter the variance of the position measurements Q ^ ; in particular, it does not increase it and the O ( γ ) shift of the conditional mean of Q ^ is not accompanied by a decrease in precision.
If we can prepare the meter in an initial state for which cov φ ( Q ^ 2 , P ^ ) < 0 , then the conditional variance of Q ^ would, to first order, decrease with γ , making the measurement of Q ^ (marginally) more accurate after the interaction than before.
Using the same method and Equation (A7), one can derive the small- γ asymptotics of order 1 for the conditional variance var σ ^ γ ( P ^ | Y ^ = y ) : we leave the details to the reader.
We will now compute the second-order correction to the conditional variance of the position operator Q ^ . We allow general φ but will assume these to be real, to simplify the calculations somewhat. As we have seen, this implies in particular that cov φ ( Q ^ 2 , P ^ ) = 0 .
Proposition A4.
Suppose φ is real-valued. Then,
var σ ^ γ ( Q ^ | Y ^ = y ) = var φ ( Q ^ ) + c 2 γ 2 + O ( γ 3 )
with
c 2 = E 2 ( X ^ ; ρ ^ , y ) Re ( E ρ ^ ( X ^ | Y ^ = y ) ) 2 + E 2 ( X ^ ; ρ ^ , y ) Re ( E ρ ^ ( X ^ 2 | Y ^ = y ) ) · cov φ ( P ^ 2 , Q ^ 2 ) 2 cov φ ( P 2 , Q ) φ , Q φ
Proof. 
The second-order Taylor expansion (A22) of p γ ( q , y ) implies that
q 2 p γ ( q , y ) d q = p 0 ( y ) ( φ , Q ^ 2 φ + 2 φ , Q ^ φ Re ( E 1 ) · γ + ( Re ( E 2 ) + ( E Re ( E 2 ) ) | | Q ^ P ^ φ | | 2 ) · γ 2 + O ( γ 3 ) ) ,
where we used that
q 2 q φ 2 d q = 2 q φ 2 d q = 2 φ , Q ^ φ ,
and
q 2 φ q 2 φ d q = 2 q φ q φ d q q 2 ( q φ ) 2 d q = φ 2 d q | | q q φ | | 2 = 1 | | Q ^ P ^ φ | | 2 .
Dividing by p γ ( y ) = p 0 ( y ) 1 + ( E Re ( E 2 ) ) | | P ^ φ | | 2 · γ 2 + O ( γ 3 ) (cf. (A23) we see that
E σ ^ γ ( Q ^ 2 | Y ^ = y ) = φ , Q ^ 2 φ + + 2 φ , Q ^ φ Re ( E 1 ) · γ + Re ( E 2 ) + ( E Re ( E 2 ) ) ( | | Q ^ P ^ φ | | 2 φ , Q ^ 2 φ φ , P ^ 2 φ ) · γ 2 + O ( γ 3 ) .
Since | | Q ^ P ^ φ | | 2 = φ , P ^ Q ^ 2 P ^ φ and
1 2 ( P 2 Q 2 + Q 2 P 2 ) = 1 2 P ( Q 2 P + [ P , Q 2 ] ) + ( P Q 2 P + [ Q 2 , P ] ) P = P Q 2 P + 1 2 [ P , [ P , Q 2 ] ] = P Q 2 P 1
we find that the coefficient of γ 2 can be re-written as E + ( E Re ( E 2 ) ) cov φ ( P ^ 2 , Q ^ 2 ) . The proposition now follows from the definition (A25) of conditional variance and the second-order expansion of E σ ^ γ ( X ^ | Y ^ = y ) of Proposition A2. □
Ogawa et al. in [47] computed the second-order correction to var σ ^ γ ( cos α Q ^ + sin α P ^ ) for arbitrary α , but for Gaussian meter states φ only. To make the connection with [47], let us again write E ν for E ( X ^ ν | Y ^ = y ) and E for E 2 ( X ^ ; ρ ^ , y ) , and put
Γ φ : = cov φ ( P ^ 2 , Q ^ 2 ) 2 cov φ ( P 2 , Q ) φ , Q φ
Ogawa et al. define the weak variance of X ^ in the post-selected state φ y and the preselected state ρ ^ , by
σ w 2 : = σ w , ρ ^ 2 ( X ^ | Y ^ = y ) : = E ρ ^ ( X ^ 2 | Y ^ = y ) E ρ ^ ( X ^ | Y ^ = y ) 2 = E 2 E 1 2 .
The coefficient c 2 of the second-order term can then be re-written in terms of this weak variance as
c 2 = E ( Re E 1 ) 2 ( Γ φ + 1 ) + ( Im E 1 ) 2 Γ φ Re ( σ w 2 ) Γ φ
If ρ ^ is pure, E = | E 1 | 2 by (A18) and
c 2 = Im ( E 1 ) 2 ( 2 Γ φ + 1 ) Re ( σ w 2 ) Γ φ .
If we now specialize to φ ( q ) = π 1 / 4 e q 2 / 2 , a centered Gaussian, then Γ φ = cov φ ( P ^ 2 , Q ^ 2 ) = 1 2 by direct computation, and we recover [47]’s result that
var σ ^ γ ( Q ^ | Y ^ = y ) = var φ ( Q ^ ) + 1 2 Re σ w 2 ( X ^ ; ρ ^ , y ) γ 2 + 0 ( γ 3 ) .
In fact, since Γ φ is invariant under dilations φ ( q ) α 1 / 4 φ ( α q ) , α > 0 and can be written as Γ φ = cov φ P ^ 2 , ( Q ^ φ , Q ^ φ ) 2 , the result holds for arbitrary Gaussians φ ( q ) = π 1 / 4 α 1 / 4 e α ( q β ) 2 / 2 , where we can even allow complex α in the right half plane, Re α > 0 , by analytic continuation. The weak variance (eq:Ogawawv) does not satisfy the law of total variance of Section 3: if we introduce a “weak variance operator” by
σ ^ w , ρ ^ 2 ( X ^ | Y ^ ) = y σ ( Y ^ ) σ w , ρ ^ 2 ( X ^ | Y ^ = y ) Π ^ y Y ^ ,
then its expectation will in general not be equal to var ( X ) var ρ ^ ( E ρ ^ ( X ^ | Y ^ ) (except when E ρ ^ ( X ^ | Y ^ ) happens to be Hermitian), unlike the conditional variances of Section 3. It is natural to try and express the coefficient c 2 in terms of the latter. We cannot use Δ ρ ^ 2 , ( X ^ | Y ^ ) , since this is 0 for pure states, but this is easily done in terms of Δ ˜ ρ ^ 2 , ( X ^ | Y ^ ) . Dropping the from the notation, and writing
Δ ˜ 2 : = Δ ˜ ρ ^ 2 ( X ^ | Y ^ = y ) = E ρ ^ ( X ^ 2 | Y ^ = y ) | E ρ ^ ( X ^ | Y ^ = y ) | 2 = E 2 | E 1 | 2 .
we find on replacing Re ( E 2 ) by Re ( Δ ˜ ) + | E 1 | 2 in (A27) that
c 2 = ( E Re ( E 1 ) 2 ) ( Γ φ + 1 ) Im ( E 1 ) 2 Γ φ Re ( Δ ˜ 2 ) .
In particular, for pure states we arrive at the following corollary of proposition A4:
Corollary A1.
If ρ ^ is pure, then
var σ ^ γ ( Q ^ | Y ^ = y ) = var φ ( Q ^ ) + c 2 γ 2 + O ( γ 3 ) ,
with
c 2 = c 2 ( X ^ , ρ ^ , y ) = ( Im E ρ ^ ( X ^ | Y ^ = y ) ) 2 Re Δ ˜ ρ ^ 2 ( X ^ | Y ^ = y ) · Γ φ
An interesting class of examples generalizing the Gaussians is that of the Hermite functions h n = h n ( q ) , which are the eigenfunctions of the Harmonic oscillator P ^ 2 + Q ^ 2 . Using the well-known recurrence relations for the derivative of h n and for the multiplication of h n by q,
h n = n 2 h n 1 n + 1 2 h n + 1 , q h n = n 2 h n 1 + n + 1 2 h n + 1 ,
together with their orthonormality, one easily shows that h n , Q ^ h n = 0 , h n , Q ^ 2 h n = φ , P ^ 2 φ = n + 1 2 and
Γ h n = cov h n ( P ^ 2 , Q ^ 2 ) = 1 2 ( n 2 + n + 1 ) .
This implies that, for pure states again,
var σ ^ γ ( Q ^ | Y ^ = y ) = n + 1 2 + Im ( E 1 ) 2 + 1 2 ( n 2 + n + 1 ) Re Δ ˜ 2 · γ 2 + O ( γ 3 ) .

Appendix A.3. Conditioned Measurements in the Case of Strong System-Meter Interactions

We turn to the large γ - or strong interaction limit of the conditional expectations with respect to σ γ . It turns out that these are still given by quantum conditional expectations, but with respect to a new state, which is the post-interaction system state: the system not only acts on the meter, but there is also a back-action of the meter on the system. We assume for simplicity that the initial reader-wave function is compactly supported, though this is by no means necessary.
Proposition A5.
Suppose that φ is compactly supported. Then, for sufficiently large γ,
E σ ^ γ ( Q ^ | Y ^ = y ) = φ , Q ^ φ + γ E ρ ^ X ^ ( X ^ | Y ^ = y ) ,
where
ρ ^ X ^ : = x σ ( X ^ ) Π ^ x X ^ ρ ^ Π ^ x X ^ = Tr L 2 ( R ) ( σ ^ γ ) ,
is the reduced post-interaction state of the system.
Note that E ρ ^ X ^ ( X ^ | Y ^ = y ) is real since X ^ and ρ ^ X ^ commute. This theorem shows that the weak value as usually defined can also naturally appear in scenarios where the interaction is not weak and corresponds to a quantum conditional expectation, as in the weak regime.
Proof. 
Since φ is compactly supported, the supports of φ ( q γ x ) and φ ( q γ x ) for x x will be disjoint if γ is sufficiently large, and therefore
p γ ( q , y ) = x σ ( X ^ ) Tr ( Π ^ y Y ^ Π ^ x X ^ ρ ^ Π ^ x X ^ ) | φ ( q γ x ) | 2
Integrating over q,
p γ ( y ) = x σ ( X ^ ) Tr ( Π ^ y Y ^ Π ^ x X ^ ρ ^ Π ^ x X ^ ) = Tr ( Π ^ y Y ^ ρ ^ X ^ )
while
q p γ ( q , y ) d q = x σ ( X ^ ) Tr ( Π ^ y Y ^ Π ^ x X ^ ρ ^ Π ^ x X ^ ) ( q + γ x ) | φ ( q ) | 2 d q = Tr ( Π ^ y Y ^ ρ ^ X ^ ) φ , Q ^ φ + γ x x Tr ( Π ^ y Y ^ Π ^ x X ^ ρ ^ Π ^ x X ^ ) = Tr ( Π ^ y Y ^ ρ ^ X ^ ) φ , Q ^ φ + γ Tr ( Π ^ y Y ^ X ^ ρ ^ X ^ )
Taking the quotient proves (A38). Concerning the interpretation of ρ ^ X ^ , Formula (A12) for σ ^ γ implies that for sufficiently large γ ,
Tr L 2 ( R ) ( σ ^ γ ) = x , x σ ( X ^ ) Π ^ x X ^ ρ ^ Π ^ x X ^ R φ ( q γ x ) φ ( q γ x ) ¯ d q = x σ ( X ^ ) Π ^ x X ^ ρ ^ Π ^ x X ^ R | φ ( q γ x ) | 2 d q = x σ ( X ^ ) Π ^ x X ^ ρ ^ Π ^ x X ^
If φ is not compactly supported, but if φ ( q ) tends to 0 at ± while φ ( q ) and q φ ( q ) are bounded, the theorem remains true in the form
lim γ E σ ^ γ ( Q ^ | Y ^ = y ) φ , Q ^ φ γ = E ρ ^ X ^ ( X ^ | Y ^ = y ) .
We also mention two other interpretations of ρ ^ X ^ : it is Umegaki conditional expectation of ρ ^ relative to F C , X ^ and also the conditional expectation of ρ ^ given X ^ with respect to the maximally mixed state.
The large γ -behavior of the pointer momentum P ^ is very different: if γ is sufficiently large, then
E σ ^ γ ( P ^ | Y ^ = y ) = φ , P ^ φ ,
independently of γ , X ^ and Y ^ , as follows by taking the trace of P ^ Π ^ Y ^ σ ^ γ with σ ^ γ given by (A12) and using the disjointedness of the supports of the different translates of φ . This result is the same as for unconditional P ^ measurements, for which E σ ^ γ ( P ^ ) = φ , P ^ φ = E σ 0 ( P ^ ) , indicating that there is no net momentum transfer from the system to the meter and vice versa.

Appendix B. Quantum Conditional Expectation Versus Von Neumann Algebraic Conditional Expectation

According to the theory of von Neumann algebras, following Umegaki [54], a conditional expectation of L ( H ) given a subalgebra A is a trace-preserving projection map E A : L ( H ) A which satisfies E A ( A 1 X A 2 ) = A 1 E A ( X ) A 2 for all A 1 , A 2 A (this notion is in fact introduced for type II von Neumann algebras in [54]). The map E A is unique and, in the finite dimensional case, given by the orthogonal projection with respect to the Hilbert–Schmidt inner product, as can be easily verified. If A = F C , Y ^ for a CSCO Y ^ , it is given by
E Y ^ ( X ^ ) : = E F C , Y ^ ( X ^ ) = y σ ( Y ^ ) Π ^ y Y ^ X ^ Π ^ y Y ^ .
The left quantum conditional expectation E ρ ^ ( X ^ | Y ^ ) can be expressed in terms of the von Neumann algebraic conditional expectation E Y ^ : L ( H ) F C , Y ^ as
E ρ ^ ( X ^ | Y ^ ) = E Y ^ ( X ^ ρ ^ ) E Y ^ ( ρ ^ ) 1
where E Y ^ ( ρ ^ ) is invertible since ρ ^ D Y ^ (cf. [34] for an analogous formula for E ρ ^ sa ( X ^ | Y ^ ) when X ^ is self-adjoint). This expression can be verified by direct computation or by using Theorem 5: The right-hand side of Equation (A43) clearly satisfies Equation (59), and Equation (61) is verified as follows:
Tr ( E Y ^ ( X ^ ρ ^ ) E Y ^ ( ρ ^ ) 1 ρ ^ ) = Tr ( E Y ^ E Y ^ ( X ^ ρ ^ ) E Y ^ ( ρ ^ ) 1 ρ ^ ) = Tr E Y ^ ( X ^ ρ ^ ) E Y ^ ( ρ ^ ) 1 E Y ^ ( ρ ^ ) = Tr E Y ^ ( X ^ ρ ^ ) = Tr ( X ^ ρ ^ ) .
One similarly checks that E ρ ^ r ( X ^ | Y ^ ) = E Y ^ ( ρ ^ X ^ ) E Y ^ ( ρ ^ ) 1 .
It is legitimate to wonder if, given a CSCO Y ^ and a state ρ ^ , it is possible that the left and right conditional expectations coincide, for all X ^ . The answer is provided by the following proposition.
Proposition A6.
Let Y ^ be a CSCO and let ρ ^ D Y ^ . Then, the following are equivalent:
(i) 
E ρ ^ ( X ^ | Y ^ ) = E ρ ^ r ( X ^ | Y ^ ) , for all X ^ ;
(ii) 
The left quantum conditional expectation E ρ ^ ( · | Y ^ ) has the “right pull-out property” given by Equation ();
(iii) 
Y ^ and ρ ^ commute.
In that case, one has
E ρ ^ ( X ^ | Y ^ ) = E Y ^ ( X ^ ) = y σ ( Y ^ ) Π ^ y Y ^ X ^ Π ^ y Y ^ = y σ ( Y ^ ) φ y , X ^ φ y Π ^ y Y ^ .
We point out that Equation (A44) is independent of the state ρ ^ D Y ^ ; it is in particular the conditional expectation of the maximally mixed state ρ ^ = 1 d I d , where d is the dimension of H .
Proof. 
That (i) implies (ii) is immediate. We prove that (ii) implies (iii). If E ρ ^ ( X ^ | Y ^ ) satisfies (60) for all X ^ and f, then it is equal to the right quantum conditional expectation, by the characterization of Theorem 5. This means that for all y and all X ^ ,
Tr ( X ^ ρ ^ Π ^ y Y ^ ) = Tr ( X ^ Π ^ y Y ^ ρ ^ )
which implies that ρ ^ Π ^ y Y ^ = Π ^ y Y ^ ρ ^ , i.e., ρ ^ commutes with all spectral projectors of Y ^ and therefore with Y ^ . To show that (iii) implies (i) it suffices to note that if ρ ^ and Y ^ commute, then the left and right quantum conditional expectations coincide.
Finally, to show Equation (A44), we note that, as Y ^ and ρ ^ commute, ( φ y Y ^ ) y σ ( Y ^ ) are also eigenfunctions of ρ ^ and their associated eigenvalues are not 0 as ρ ^ D Y ^ . Consequently, for all y σ ( Y ^ ) :
φ y Y ^ , X ^ ρ ^ φ y Y ^ φ y Y ^ , ρ ^ φ y Y ^ = φ y Y ^ , X ^ φ y Y ^ φ y Y ^ , ρ ^ φ y Y ^ φ y Y ^ , ρ ^ φ y Y ^ = φ y Y ^ , X ^ φ y Y ^ .
Thus, we obtain that
E ρ ^ ( X ^ | Y ^ ) = y σ ( Y ^ ) φ y Y ^ , X ^ ρ ^ φ y Y ^ φ y Y ^ , ρ ^ φ y Y ^ Π ^ y Y ^ = y σ ( Y ^ ) φ y Y ^ , X ^ φ y Y ^ Π ^ y Y ^ = y σ ( Y ^ ) Π ^ y Y ^ X ^ Π ^ y Y ^ .
We finish by mentioning that, as a consequence of this result, the only state ρ ^ that satisfies
E ρ ^ ( X ^ | Y ^ ) = E ρ ^ r ( X ^ | Y ^ ) , X ^ L ( H ) , Y ^ CSCO ,
is the maximally mixed state ρ ^ = 1 d I d . Indeed from Proposition A6, such a state must commute with all CSCOs; hence, it must commute with all rank one projectors and thus with all of L ( H ) . This implies that it is a multiple of the identity.

Appendix C. Comparing the Wigner and KD Distributions

To illustrate our findings, and in particular the distinguishing features of the KD representations, we compare here the KD representation and the Gross–Wigner representation of a single qutrit with Hilbert space C 3 .
For that purpose we introduce the “position operator” Q ^ = 0 | 0 0 | + 1 | 1 1 | + 2 | 2 2 | with eigenbasis
| 0 = 1 0 0 , | 1 = 0 1 0 , | 2 = 0 0 1 ,
as well as the conjugate “momentum” operator P ^ = 0 | 0 ^ 0 ^ | + 1 | 1 ^ 1 ^ | + 2 | 2 ^ 2 ^ | , with eigenbasis
| 0 ^ = 1 3 1 1 1 , | 1 ^ = 1 3 1 j j 2 , | 2 ^ = 1 3 1 j 2 j ,
where j = e i 2 π 3 . The transition matrix between these two bases is provided by the discrete Fourier transform in 3 dimensions:
U = 1 3 1 1 1 1 j j 2 1 j 2 j .
We will consider the KD representation of quantum mechanics on H = C 3 with A ^ = P ^ and B ^ = Q ^ . It was proven in [46] that in this case X ^ V KD , real sa if and only if X ^ = f ( P ^ ) + g ( Q ^ ) . In other words, any such X ^ can be written
X ^ = i = 1 3 f i | i ^ i ^ | + j = 1 3 g j | j j | , f i , g j R .
The same holds true for qudits, when d is a prime number [23,46].
On the other hand, we recall the definition of the Gross–Wigner representation, defined in [21]. Let Λ = σ ( P ^ ) × σ ( Q ^ ) = { 0 , 1 , 2 } 2 . Consider the operators
S p , q W = 1 3 0 p , q 2 e i 2 π 3 ( p q q p ) e i π 3 p q z ^ ( p ) x ^ ( q ) , 0 p , q 2 ,
where z ^ ( p ) , x ^ ( q ) are generalized Pauli operators for ( p , q ) Λ , satisfying
x ^ ( q ) | q = | q + q , z ^ ( p ) | q = e i 2 π 3 p q | q ,
for all q { 0 , 1 , 2 } . It is readily verified that the family ( S p , q W ) 0 p , q 2 is a frame, whose dual frame is given by
T p , q W = 3 S p , q W , 0 p , q 2 .
The Gross–Wigner function of a state ρ ^ is the associated quasiprobability, given by
W p , q ( ρ ^ ) = Tr ρ ^ S p , q W = 1 3 0 ξ 2 e i 2 π 3 ξ p ρ ^ q + ξ 2 , q ξ 2 , 0 p , q 2 .
The associated symbol of an operator X ^ L ( H ) is given by
W ˜ p , q ( X ^ ) = Tr X ^ T p , q W = ξ e i 2 π 3 ξ p X ^ q + ξ 2 , q ξ 2 = 3 W p , q ( X ^ ) , 0 p , q 2 .
Theorem 6 implies that the conditional expectation E ρ ^ W ( X ^ | Q ^ ) does not coincide with E ρ ^ / r ( X ^ | Q ^ ) for all X ^ and ρ ^ . Indeed, the frame of Equation (A47) is not composed of rank one operators. For example, S 0 , 0 W is the parity operator satisfying S 0 , 0 W | q = | q , and is therefore of full rank. We now illustrate this fact with a specific example. We have
E ρ ^ W ( X ^ | Q ^ ) = p , q = 0 2 W ˜ p , q ( X ^ ) ¯ W p , q ( ρ ^ ) q | ρ ^ | q Π ^ q Q ^ .
It is easily verified that the Gross–Wigner symbol satisfies W ˜ p , q ( X ) = W ˜ p , q ( X ) ¯ for any X L ( H ) and any p , q Λ . It follows that
E ρ ^ W ( X ^ | Q ^ ) = p , q = 0 2 W ˜ p , q ( X ^ ) W p , q ( ρ ^ ) q | ρ ^ | q Π ^ q Q ^ .
Consider the state ρ ^ given in the position basis by
ρ ^ = 1 4 1 1 0 1 2 0 0 0 1 .
One readily checks that ρ ^ D Q ^ . The Gross–Wigner function of ρ ^ is then given by
( W p , q ( ρ ^ ) ) 0 p , q 2 = 1 12 1 2 3 1 2 0 1 2 0 .
The KD distribution of ρ ^ is
( Q a , b KD , ( ρ ^ ) ) 0 a , b 2 = 1 12 2 3 1 1 + j 2 2 + j 1 1 + j 2 + j 2 1 .
Note that ρ ^ is Wigner-positive; in fact, it can be verified that ρ ^ is a convex combination of stabilizer states but not KD-positive. We recall that the pure Wigner-positive states are the so-called stabilizer states and that the convex hull of the stabilizer states does not exhaust the set of all Gross–Wigner-positive states [21]. We note that, in this particular case, the set of KD-positive states is the convex hull of the basis states [23] and thus all KD-positive states are convex combinations of stabilizer states. Note that this is no longer true for any system of n qubits [66], for n 2 . Thus, we consider the state ρ ^ to be in the convex hull of the stabilizer states but outside of the convex hull of pure KD-positive states.
Now let X ^ = f ( P ^ ) for f : σ ( P ^ ) = { 0 , 1 , 2 } C . Then, its Gross–Wigner function is
( W ˜ p , q ( f ( P ^ ) ) ) 0 p , q 2 = f ( 0 ) f ( 0 ) f ( 0 ) f ( 1 ) f ( 1 ) f ( 1 ) f ( 2 ) f ( 2 ) f ( 2 ) .
It follows from Equation (A51) that
E ρ ^ W ( f ( P ^ ) | Q ^ ) = 1 3 ( f ( 0 ) + f ( 1 ) + f ( 2 ) ) ( | 0 0 | + | 1 1 | ) + f ( 0 ) | 2 2 | .
Note that, when f is real-valued, this is a self-adjoint operator. In fact, as the Gross–Wigner function of self-adjoint operators is real, the Gross–Wigner conditional expectation E ρ ^ W ( X ^ | Q ^ ) is self-adjoint when X ^ is self-adjoint. On the other hand, it is easily computed that
E ρ ^ ( f ( P ^ ) | Q ^ ) = 1 3 ( 2 f ( 0 ) + ( 1 + j 2 ) f ( 1 ) + ( 1 + j ) f ( 2 ) ) | 0 0 | + 1 6 ( 3 f ( 0 ) + ( 2 + j ) f ( 1 ) + ( 2 + j 2 ) f ( 2 ) ) | 1 1 | + 1 3 ( f ( 0 ) + f ( 1 ) + f ( 2 ) ) | 2 2 | .
Hence, the Gross–Wigner-conditional expectation is indeed different from the left conditional expectation and therefore also from the left KD-conditional expectation. In particular, the latter may have a non-zero imaginary part even if f is a real valued function as is readily seen from Equation (A53).
To illustrate that the Gross–Wigner conditional expectation does not satisfy the pull-through property, we consider the function f ( p ) = sin 2 π 3 p and the operator X ^ = f ( P ^ ) . From Equation (A52), one has
E ρ ^ W ( X ^ | Q ^ ) = 0 .
On the other hand, one computes
( W ˜ p , q ( Q ^ X ^ ) ) 0 p , q 2 = i 3 4 2 1 1 1 j 2 j 1 j j 2 ,
from which follows
E ρ ^ W ( Q ^ X ^ | Q ^ ) = i 3 4 | 2 2 | .
Hence,
E ρ ^ W ( Q ^ X ^ | Q ^ ) Q ^ E ρ ^ W ( X ^ | Q ^ ) ,
which shows that the left pull-through property is not satisfied by the Gross–Wigner conditional expectation. One could similarly verify that the right pull-through property isn’t satisfied either. While the conditional expectations associated to the Wigner and Kirkwood–Dirac representations are different, they may share similarities for specific observables or states. To illustrate this fact, let us briefly go to the infinite dimensional setting, with H = L 2 ( R ) . Let us use W to denote the standard Wigner distribution and by Q KD the Kirkwood–Dirac distribution associated to the position and momentum operators Q ^ and P ^ . It is easily computed that for any pure state ψ L 2 ( R ) and any q R ,
E ψ W ( P ^ | Q ^ = q ) = Re i ψ ( q ) ψ ( q ) ¯ | ψ ( q ) | 2 , E ψ Q KD ( P ^ | Q ^ = q ) = i ψ ( q ) ψ ( q ) ¯ | ψ ( q ) | 2 ,
such that the Wigner conditional expectation of P ^ knowing Q ^ is exactly is the self-adjoint part of the Kirkwood–Dirac conditional expectation. This property does not hold in general, as
E ψ W ( P ^ 2 | Q ^ = q ) = | ψ ( q ) | 2 Re ψ ( q ) ψ ( q ) ¯ 2 | ψ ( q ) | 2 , E ψ Q KD ( P ^ 2 | Q ^ = q ) = ψ ( q ) ψ ( q ) ¯ | ψ ( q ) | 2 ,
hence E ψ W ( P ^ 2 | Q ^ ) Re E ψ Q KD ( P ^ 2 | Q ^ ) in general.

Appendix D. Interpretation of Phase Insensitivity

In this Appendix, we briefly explore the physical and operational meaning of vanishing Fisher information. We will in particular give an alternative interpretation of the Fisher information that is independent of parameter estimation and the Cramer–Rao bound and as such useful in the following discussions.
First, remark that the Cramer–Rao bound seems to indicate that vanishing Fisher information implies that the variance of any estimator of θ is infinite. This however, is not possible since the set σ ( Y ^ ) is finite, such that any random variable on σ ( Y ^ ) necessarily has a finite variance. In fact, I F ( Y ^ ; 0 ) = 0 is equivalent to θ p ( y , 0 ) = 0 , for all y σ ( Y ^ ) . But, for an unbiased estimator θ ˜ ( Y ^ ) , one has
y θ ˜ ( y ) p ( y ; θ ) = θ ,
for all θ in a neighborhood of θ = 0 . Taking a derivative with respect to θ on both sides of this equation, one immediately concludes that such an unbiased estimator does not exist if θ p ( y , 0 ) = 0 , for all y σ ( Y ^ ) . In other words, one concludes that repeated measurements of Y ^ cannot be used to produce an unbiased estimate on the phase θ of the states ρ ^ X ^ ( θ ) .
A different interpretation of the Fisher information, not directly related to parameter estimation, sheds light on this situation, and further justifies the above definition of phase-insensitive state. We consider the set of all probabilities ( p y ) y σ ( Y ^ ) on σ ( Y ^ ) , which form a convex polytope. We will suppose p y 0 for all y, meaning that p y does not lie on any of the facets of the probability polytope. When such a probability is experimentally determined, it is natural to think of the experimental errors ( δ p y ) y σ ( Y ^ ) as forming a tangent vector to the set of probability distributions on the spectrum of Y ^ , at the probability distribution ( p y ) y σ ( Y ^ ) . Given two such tangent vectors ( δ p y ) y σ ( Y ^ ) , ( δ p y ) y σ ( Y ^ ) , one may consider the Riemannian metric (i.e., the symmetric positive definite form)
( δ p ) , ( δ p ) F = y σ ( Y ^ ) δ p y 1 p ( y ) δ p y ,
known as the Fisher metric. This transforms the probability polytope into a Riemannian manifold. The square of the Fisher norm of δ p y is naturally interpreted as the mean squared relative error of the experimentally determined probabilities:
δ p F 2 = δ p , δ p F = y σ ( Y ^ ) δ p y p ( y ) 2 p ( y ) .
Next, we consider the curve θ ( p ( y ; θ ) ) y , where p ( y ; θ ) = Tr ( ρ ^ X ^ ( θ ) Π ^ y Y ^ ) . One then sees that
I F ( Y ^ ; θ ) = d p d θ F 2 .
This interprets the Fisher information as the squared relative rate of change of the probabilities p ( y ; θ ) , as a result of a variation in the phase θ . As such, I F ( Y ^ ; θ ) measures the sensitivity of p ( y ; θ ) to variations in θ . In fact, for small δ θ , one has that
p ( θ + δ θ ) p ( θ ) F I F ( Y ^ ; θ ) 1 / 2 δ θ .
One concludes that, for a variation δ θ to induce an experimentally noticeable change in the relative changes of the probabilities p ( y ; θ ) , this variation must exceed the mean square relative experimental error, meaning that
δ θ 2 I F ( Y ^ ; θ ) δ p F 2 , or δ θ 2 δ θ min 2 : = δ p F 2 I F ( Y ^ ; θ ) .
In other words, only when δ θ exceeds δ θ min is one sure that the probability distribution p ( y , θ + δ θ ) leaves the ball of Fisher radius δ p F around p ( y ; θ ) , thereby making the change in θ experimentally detectable. From this point of view, the probability distributions p ( y ; θ ) are indeed phase-insensitive when the Fisher information is small: δ θ min is then large. This interpretation does not refer to the estimation of θ .
Conversely, one may also conclude that, given the experimental precision δ p y , it is not possible to infer θ with a precision smaller than δ θ min . In other words, when the Fisher information is large, a small change in the phase θ will be experimentally noticeable in variations of the Born probabilities p ( y ; θ ) . Under these same circumstances, one can also expect to determine θ with high accuracy: this is the idea behind the Cramer–Rao bound.
If I F ( Y ^ ; θ ) = 0 , on the other hand, θ p ( y , 0 ) vanishes for all y: variations of θ have then little effect on the measurement outcome probabilities p ( y , θ ) . A second-order Taylor expansion of p ( y , θ ) then yields analogously
δ θ 2 δ θ min 2 : = 2 δ p F J ( Y ^ ; θ ) ,
where
J ( Y ^ ; θ ) = y σ ( Y ^ ) θ 2 p ( y , θ ) 1 p ( y , θ ) θ 2 p ( y , θ ) ,
assuming not all θ 2 p ( y , θ ) vanish. Note that in this case δ θ min 2 behaves like δ p F and is therefore larger than when the Fisher information does not vanish. The state is therefore less sensitive to variations in θ when the Fisher information vanishes than when it does not.
It should be noted that, in what precedes, we have only invoked the classical Fisher information of the θ -dependent probabilities p ( θ , y ) . We briefly elaborate on the link with the quantum Fisher information of the states ρ ^ X ^ ( θ ) in Appendix E where we will use the techniques developed here concerning the quantum conditional expectation to give a simple proof of the result of Braunstein and Caves [62] linking the definition of quantum Fisher information of Helstrom and Holevo [63,64] to the optimization of classical Fisher information over all possible quantum measurements.

Appendix E. Quantum Fisher Information

Quantum conditional expectations can be used to give a quick proof of the Braunstein and Caves theorem expressing the quantum Fisher information as a supremum of classical Fisher information. To be able to do so in full generality, we will have to extend the definition of Fisher information I F ( Y ^ ; θ ) to cases when ρ ^ D Y ^ and also revisit the quantum conditional expectations for such ρ ^ .
First of all, if ρ ^ ( θ ) is an arbitrary C 1 -family of density operators on H , not necessarily of the form considered in Section 6, and if Y ^ is an arbitrary observable, we define the Fisher information by
I F ( Y ^ ; θ ) : = y : p ( y ; θ ) 0 ( θ p ( y ; θ ) ) 2 p ( y ; θ )
where p ( y ; θ ) : = Tr ( Π ^ y Y ^ ρ ^ ( θ ) ) and the sum extends over elements y of the spectrum σ ( Y ^ ) , here and below.
Remark A1.
That this is the correct definition of the Fisher information when ρ ^ D Y ^ follows from the continued validity of the Cramèr–Rao inequality for the variance of an unbiased estimator θ ( Y ^ ) for θ, which itself is a consequence of the identity y : p ( y ; θ ) 0 θ ( y ) θ θ p ( y ; θ ) = 1 . The latter follows as usual from differentiating the unbiasedness condition y θ ( y ) p ( y ; θ ) = θ and using that y θ p ( y ; θ ) = 0 . Since p ( y , θ ) = 0 implies that θ p ( y ; θ ) = 0 , the sum extends only over y’s for which p ( y ; θ ) 0 .
Turning to quantum conditional expectations, as noted in Remark 1, minimizers f 0 ( y ) of (44)) (or minimizers of (46)) are no longer unique (all the minimizers are given in (54)). However, E ρ ^ ( | f 0 ( Y ^ ) | 2 ) is independent of the choice of minimizer and given by
E ρ ^ ( | E ρ ^ , 0 ( X ^ | Y ^ ) | 2 ) = Tr ( Π ^ y Y ^ ρ ^ ) 0 | Tr ( Π ^ y Y ^ X ^ ρ ^ ) | 2 Tr ( Π ^ y Y ^ ρ ^ ) ,
where it is convenient to introduce a distinguished minimizer
E ρ ^ , 0 ( X ^ | Y ^ ) : = Tr ( Π ^ y Y ^ ρ ^ ) 0 Tr ( Π ^ y Y ^ X ^ ρ ^ ) Tr ( Π ^ y Y ^ ρ ^ ) Π ^ y Y ^ ,
which is the minimizer with minimal Hilbert–Schmidt norm. We have that E ρ ^ ( | E ρ ^ , 0 ( X ^ | Y ^ ) | 2 ) E ρ ^ ( X ^ X ^ ) with equality if Y ^ = X ^ . Similarly, if we define E ρ ^ sa , 0 ( X ^ | Y ^ ) = Re E ρ ^ , 0 ( X ^ | Y ^ ) , then for self-adjoint X ^ , E ρ ^ ( | E ρ ^ sa , 0 ( X ^ | Y ^ ) | 2 ) E ρ ^ ( X ^ 2 ) , with equality if Y ^ = X ^ . This will be used below for the proof of Braunstein–Caves.
To finally make the connection with the quantum Fisher information, recall that a symmetric logarithmic derivate of a C 1 -family ρ ^ ( θ ) is defined to be any self-adjoint operator L ^ ( θ ) satisfying
θ ρ ^ ( θ ) = 1 2 { L ^ ( θ ) , ρ ^ ( θ ) } ,
where { C ^ , D ^ } = C ^ D ^ + D ^ C ^ is the anti-commutator [62,63,64,77]. In general, if R ^ 0 and S ^ are self-adjoint operators on H , the operator (or matrix) equation L ^ R ^ + R ^ L ^ = S ^ has a solution if and only if Π ^ 0 R ^ S ^ Π ^ 0 R ^ = 0 , where Π ^ 0 R ^ is the orthogonal projection onto the kernel of R ^ . If | ψ j is an orthogonal basis of eigenvectors of R ^ with eigenvalues r j (repeated according to their multiplicity), then the general solution is given by
L ^ = r j + r k > 0 ψ j , S ^ ψ k r j + r k | ψ j ψ k | + Π ^ 0 R ^ C ^ Π ^ 0 R ^ ,
where C ^ L ( H ) is arbitrary. Since S ^ is self-adjoint, L ^ is self-adjoint if and only if the final term on the right is, and E ( L ^ 2 R ^ ) is independent of the choice of self-adjoint solution L ^ .
We can apply this with R ^ = ρ ^ ( θ ) and S ^ = θ ρ ^ ( θ ) . The necessary and sufficient condition for solvability of the defining equation for L ^ ( θ ) is automatically satisfied, and a consequence of the positivity of ρ ^ ( θ ) : suppose that Ker ( ρ ^ ( θ 0 ) ) 0 , and let Π ^ 0 ( θ 0 ) : = Π ^ 0 ρ ^ ( θ 0 ) be the corresponding orthogonal projector. Then the positive (= non-negative) function v , Π ^ 0 ( θ 0 ) ρ ^ ( θ ) Π ^ 0 ( θ 0 ) v vanishes in θ = θ 0 , for any vector v H , which implies that its derivative in θ 0 vanishes:
v , Π ^ 0 ( θ 0 ) θ ρ ^ ( θ 0 ) Π ^ 0 ( θ 0 ) v = 0 ,
Since this holds for arbitrary v, Π ^ 0 ( θ 0 ) θ ρ ^ ( θ 0 ) Π ^ 0 ( θ 0 ) = 0 , by the usual polarization argument.
We can now, following Helstrom [63], define the quantum Fisher information or QFI by
I QF ( θ ) = E ρ ^ ( θ ) ( L ^ ( θ ) 2 ) ,
where the expectation on the right is, as we have seen, independent of the choice of symmetric logarithmic derivative. Helmstrom introduced the notion of symmetric logarithmic derivative, and used it to prove the quantum Cramèr–Rao bound that the variance of any unbiased estimator θ ( Y ^ ) of θ constructed from (observations of) some observable Y ^ is bounded from below by I QF ( θ ) 1 by adapting the proof of the classical Cramèr–Rao bound: see also [64].
Symmetric logarithmic derivatives can now be used to express the classical Fisher information I F ( Y ^ , θ ) in terms of the quantum conditional expectations.
Proposition A7.
If L ^ ( θ ) is a symmetric logarithmic derivative of the C 1 -family of states ρ ^ ( θ ) , then
I F ( Y ^ , θ ) = E ρ ^ ( θ ) E ρ ^ ( θ ) sa , 0 ( L ^ ( θ ) | Y ^ ) 2
Proof. 
If p ( y , θ ) 0 , then
θ log p ( y ; θ ) = Tr ( Π ^ y Y ^ θ ρ ^ ( θ ) ) Tr ( Π ^ y Y ^ ρ ^ ( θ ) ) = Re Tr ( Π ^ y Y ^ L ^ ( θ ) ρ ^ ( θ ) ) Tr ( Π ^ y Y ^ ρ ^ ( θ ) ) = E ρ ^ ( θ ) sa , 0 ( L ^ ( θ ) | Y ^ = y ) ,
from which the formula follows after squaring, multiplying by p ( y ; θ ) and summing over the y’s for which p ( y ; θ ) 0 .
Since the expectation on the right-hand side of (A67) is maximized if Y ^ = L ^ ( θ ) , we obtain the Braunstein and Caves theorem as an immediate corollary:
Corollary A2
([62]). If ρ ^ ( θ ) is a C 1 -family of non-singular states, then
max Y ^ observ . I F ( Y ^ ; θ ) = E ρ ^ ( θ ) L ^ ( θ ) 2 = I QF ( θ ) .
The reason we cannot just work with quantum conditional expectations for which ρ ^ D Y ^ is that it is not necessarily true that ρ ^ ( θ ) D L ^ ( θ ) , preventing us to then take Y ^ = L ^ ( θ ) . For example, if ρ ^ ( θ ) = | ψ ( θ ) ψ ( θ ) | is a family of pure states with ψ ( θ ) a C 1 -family of unit vectors, then ρ ^ ( θ ) 2 = ρ ^ ( θ ) implies that L ^ ( θ ) = 2 θ ρ ^ ( θ ) = 2 ( | θ ψ ( θ ) ψ ( θ ) | + | ψ ( θ ) θ ψ ( θ ) | ) is a symmetric logarithmic derivative, and if χ ψ ( θ ) , θ ψ ( θ ) , then χ is an eigenvector of L ^ ( θ ) with eigenvalue 0, for which χ , ρ ^ ( θ ) χ = 0 , so ρ ^ ( θ ) D L ^ ( θ ) .
We note for completeness that a short computation shows that the quantum Fisher information of such a family of pure states is given by the well-known formula
I QF ( θ ) = 4 ( | | θ ψ | | 2 | ψ , θ ψ | 2 ) ,
with the right-hand side evaluated in θ , and where we used that ψ , θ ψ = θ ψ , ψ .
If we take ρ ^ ( θ ) to be the Born state associated to an observable Y ^ ,
ρ ^ ( θ ) = y p ( y ; θ ) Π ^ y Y ^ ,
then the quantum Fisher information associated to this state is precisely the classical Fisher information, I F ( Y ^ ; θ ) . We examine what happens near an isolated point θ 0 where this state is singular: p ( y ; θ 0 ) = 0 for certain values of y. Assuming that p ( y ; θ ) is analytic (or at least not flat at θ 0 ), positivity implies that the first non-zero derivative in θ 0 is of even order, say of order 2 k y , k y N * . Letting π 2 k ( y ) : = θ 2 k y p ( y ; θ ) , Taylor expansion of the numerator and denominator shows that
( θ p ( y ; θ ) ) 2 p ( y ; θ ) = ( 2 k ) ! ( 2 k 1 ) ! 2 ( θ θ 0 ) 2 k 2 + ,
whose limit as θ θ 0 is 0 unless k = 1 , in which case it is 2 π 2 . We therefore have the curious result that
lim θ θ 0 I F ( Y ^ ; θ ) = I F ( Y ^ ; θ 0 ) + y : p ( y ; θ 0 ) = 0 2 θ 2 p ( y ; θ 0 ) ,
showing that the Fisher information is in general not continuous as function of the parameter, unless all maxima and minima of the probabilities p ( y ; θ ) are degenerate (non-Morse). Doing a similar analysis of the multi-parameter case would be more complicated, the limit of the quotient of two quadratic forms does not exist.

Appendix F. Variational Characterizations of KD

The identification of the conditional expectations defined by the left or right KD distribution with the corresponding quantum conditional expectations leads to a variational characterization of the KD quasiprobabilities. In the proof of Theorem 4 the minimization is done for each y σ ( Y ^ ) separately (see Equation (49)) and the proof shows in fact that
E ρ ^ ( X ^ | Y ^ = y ) = Tr ( Π ^ y Y ^ X ^ ρ ^ ) Tr ( Π ^ y Y ^ ρ ^ ) = Argmin λ C E ρ ^ | X ^ λ Π ^ y Y ^ | 2 .
Applying this with X ^ = Π ^ a A ^ and Y ^ = B ^ , where A ^ and B ^ are two complementary CSCOs, this translates into
Q a , b KD , φ b B ^ , ρ ^ φ b B ^ = Argmin λ C E ρ ^ | Π ^ a A ^ λ Π ^ b B ^ | 2
Now, if a function f ( λ ) of a complex variable λ attains its minimum in λ 0 and if β C , then the function f β ( λ ) : = f ( λ / β ) has its minimum in β λ 0 since f ( λ / β ) f ( λ 0 ) = f β ( β λ 0 ) ; this remains of course true for the function β f β ( λ ) if β > 0 . Applying this to the function which is minimized in the right-hand side of (A70) with β = φ b B ^ , ρ ^ φ b B ^ we find
Proposition A8.
Q a , b KD , ( ρ ^ ) = Argmin λ C E ρ ^ φ b B ^ , ρ ^ φ b B ^ Π ^ a A ^ λ Π ^ b B ^ 2 .
In words, Q a , b KD , ( ρ ^ ) Π ^ b B ^ is the best approximation of the spectral projector Π ^ a A ^ of a σ ( A ^ ) , weighted by the Born probability of b σ ( B ^ ) by a multiple of the spectral projector Π ^ b B ^ .
There is a second variational characterization which follows from the observation that if ρ ^ 0 : = d 1 I d is the maximally mixed state (d being the dimension of H ), then
Q a , b KD , ( ρ ^ ) = Tr ( Π ^ b B ^ Π ^ a A ^ ρ ^ ) = E ρ ^ 0 ( Π ^ a A ^ ρ ^ | B ^ = b ) ,
Tr ( Π ^ b B ^ ρ ^ 0 ) being d 1 since B ^ is a CSCO. Formula (A69) then immediately implies
Proposition A9.
Q a , b KD , ( ρ ^ ) = Argmin λ C Tr | Π ^ a A ^ ρ ^ λ Π ^ b B ^ | 2 .
Another way to derive (A73) is by observing that
b σ ( B ^ ) Q a , b KD , ( ρ ^ ) Π ^ b B ^ = E B ^ Π ^ a A ^ ρ ^ ,
where E B ^ is defined in (A42). Since E B ^ is the orthogonal projection with respect to the Hilbert–Schmidt norm of L ( H ) onto F B ^ , C , this implies (A73).
We finally note that Formulas (A71) and (A73) can also easily be verified directly.

References

  1. Ludwig, G. Foundations of Quantum Mechanics I; Springer: Berlin/Heidelberg, Germany, 1983. [Google Scholar] [CrossRef] [Scilit]
  2. Busch, P. Quantum States and Generalized Observables: A Simple Proof of Gleason’s Theorem. Phys. Rev. Lett. 2003, 91, 120403. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Ferrie, C.; Morris, R.; Emerson, J. Necessity of negativity in quantum theory. Phys. Rev. A 2010, 82, 044103. [Google Scholar] [CrossRef] [Scilit]
  4. Lostaglio, M.; Belenchia, A.; Levy, A.; Hernández-Gómez, S.; Fabbri, N.; Gherardini, S. Kirkwood-Dirac quasiprobability approach to the statistics of incompatible observables. Quantum 2023, 7, 1128. [Google Scholar] [CrossRef] [Scilit]
  5. Wigner, E. On the Quantum Correction For Thermodynamic Equilibrium. Phys. Rev. 1932, 40, 749–759. [Google Scholar] [CrossRef] [Scilit]
  6. Weyl, H. Quantenmechanik und Gruppentheorie. Z. Phys. 1927, 46, 1–46. [Google Scholar] [CrossRef] [Scilit]
  7. Kirkwood, J.G. Quantum statistics of almost classical assemblies. Phys. Rev. 1933, 44, 31. [Google Scholar] [CrossRef] [Scilit]
  8. Dirac, P.A.M. On the Analogy Between Classical and Quantum Mechanics. Rev. Mod. Phys. 1945, 17, 195–199. [Google Scholar] [CrossRef] [Scilit]
  9. Moyal, J.E. Quantum mechanics as a statistical theory. Math. Proc. Camb. Philos. Soc. 1949, 45, 99–124. [Google Scholar] [CrossRef] [Scilit]
  10. Sudarshan, E.C.G. Equivalence of Semiclassical and Quantum Mechanical Descriptions of Statistical Light Beams. Phys. Rev. Lett. 1963, 10, 277–279. [Google Scholar] [CrossRef] [Scilit]
  11. Glauber, R.J. Coherent and Incoherent States of the Radiation Field. Phys. Rev. 1963, 131, 2766–2788. [Google Scholar] [CrossRef] [Scilit]
  12. Cahill, K.; Glauber, R.J. Ordered Expansions in Boson Amplitude Operators. Phys. Rev. 1969, 177, 1857. [Google Scholar] [CrossRef] [Scilit]
  13. Cahill, K.; Glauber, R.J. Density Operators and Quasi-Probability Distributions. Phys. Rev. 1969, 177, 1882. [Google Scholar] [CrossRef] [Scilit]
  14. Wootters, W.K. A Wigner-function formulation of finite-state quantum mechanics. Ann. Phys. 1987, 176, 1–21. [Google Scholar] [CrossRef] [Scilit]
  15. Cohendet, O.; Combe, P.; Sirugue, M.; Sirugue-Collin, M. A stochastic treatment of the dynamics of an integer spin. J. Phys. A Math. Gen. 1988, 21, 2875. [Google Scholar] [CrossRef] [Scilit]
  16. Leonhardt, U. Quantum-State Tomography and Discrete Wigner Function. Phys. Rev. Lett. 1995, 74, 4101–4105. [Google Scholar] [CrossRef] [Scilit]
  17. Leonhardt, U. Discrete Wigner function and quantum-state tomography. Phys. Rev. A 1996, 53, 2998–3013. [Google Scholar] [CrossRef] [Scilit]
  18. Bouzouina, A.; De Bièvre, S. Equipartition of the eigenfunctions of quantized ergodic maps on the torus. Commun. Math. Phys. 1996, 178, 83–105. [Google Scholar] [CrossRef] [Scilit]
  19. Feichtinger, H.G.; Kozek, W. Quantization of TF Lattice-Invariant Operators on Elementary LCA Groups. In Gabor Analysis and Algorithms: Theory and Applications; Feichtinger, H.G., Strohmer, T., Eds.; Birkhäuser: Basel, Switzerland, 1998; pp. 233–266. [Google Scholar] [CrossRef] [Scilit]
  20. Wootters, W.K. Picturing qubits in phase space. IBM J. Res. Dev. 2004, 48, 99–110. [Google Scholar] [CrossRef] [Scilit]
  21. Gross, D. Hudson’s Theorem for Finite-Dimensional Quantum Systems. J. Math. Phys. 2006, 47, 122107. [Google Scholar] [CrossRef] [Scilit]
  22. Bény, C.; Crann, J.; Lee, H.H.; Park, S.J.; Youn, S.G. Gaussian quantum information over general quantum kinematical systems I: Gaussian states. Lett. Math. Phys. 2025, 115, 32. [Google Scholar] [CrossRef] [Scilit]
  23. De Bièvre, S.; Langrenez, C.; Radchenko, D. The Kirkwood-Dirac Representation Associated to the Fourier Transform for Finite Abelian Groups: Positivity. Ann. Henri Poincaré 2025. [Google Scholar] [CrossRef] [Scilit]
  24. Spriet, M. Characterizing the Kirkwood-Dirac positivity on second countable LCA groups. arXiv 2025, arXiv:2507.23628. [Google Scholar] [CrossRef] [Scilit]
  25. Nicola, F.; Riccardi, F. The Hudson Theorem in LCA Groups and Infinite Quantum Spin Systems. arXiv 2025, arXiv:2507.13154. [Google Scholar] [CrossRef] [Scilit]
  26. Klauder, J.R.; Skagerstam, B.S. Coherent States: Applications in Physics and Mathematical Physics; World Scientific: Singapore, 1985. [Google Scholar]
  27. Robert, D.; Combescure, M. Coherent States and Applications in Mathematical Physics; Theoretical and Mathematical Physics; Springer International Publishing: Berlin/Heidelberg, Germany, 2021. [Google Scholar] [CrossRef] [Scilit]
  28. Gazeau, J.P. Coherent States in Quantum Physics; Wiley-VCH Verlag GmbH & Co., KGaA: Weinheim, Germany, 2009. [Google Scholar] [CrossRef] [Scilit]
  29. Perelomov, A. Generalized Coherent States and Their Applications; Springer: Berlin/Heidelberg, Germany, 1986. [Google Scholar] [CrossRef] [Scilit]
  30. Ferrie, C. Quasi-Probability Representations of Quantum Theory with Applications to Quantum Information Science. Rep. Prog. Phys. 2011, 74, 116001. [Google Scholar] [CrossRef] [Scilit]
  31. Arvidsson-Shukur, D.R.M.; Braasch, W.F., Jr.; De Bièvre, S.; Dressel, J.; Jordan, A.N.; Langrenez, C.; Lostaglio, M.; Lundeen, J.S.; Yunger Halpern, N. Properties and applications of the Kirkwood–Dirac distribution. New J. Phys. 2024, 26, 121201. [Google Scholar] [CrossRef] [Scilit]
  32. Dressel, J. Weak values as interference phenomena. Phys. Rev. A 2015, 91, 032116. [Google Scholar] [CrossRef] [Scilit]
  33. Brummelhuis, R. Conditional expectations in Quantum Mechanics and causal interpretations: The Bohm momentum as a best predictor. arXiv 2024, arXiv:2411.08532. [Google Scholar] [CrossRef] [Scilit]
  34. Brummelhuis, R. Conditional expectations of quantum observables: I. Bohm momentum as best predictor of momentum given position. J. Phys. A Math. Theor. 2025, 58, 455304. [Google Scholar] [CrossRef] [Scilit]
  35. Brummelhuis, R. Conditional expectations of quantum observables: II. A causal model for the Pauli equation. J. Phys. A Math. Theor. 2025, 58, 455305. [Google Scholar] [CrossRef] [Scilit]
  36. Aharonov, Y.; Albert, D.Z.; Vaidman, L. How the result of a measurement of a component of the spin of a spin-1/2 particle can turn out to be 100. Phys. Rev. Lett. 1988, 60, 1351–1354. [Google Scholar] [CrossRef] [Scilit]
  37. Hofmann, H.F. Uncertainty limits for quantum metrology obtained from the statistics of weak measurements. Phys. Rev. A 2011, 83, 022106. [Google Scholar] [CrossRef] [Scilit]
  38. Steinberg, A.M. Conditional probabilities in quantum theory and the tunneling-time controversy. Phys. Rev. A 1995, 52, 32–42. [Google Scholar] [CrossRef] [Scilit]
  39. Johansen, L.M. Quantum theory of successive projective measurements. Phys. Rev. A 2007, 76, 012119. [Google Scholar] [CrossRef] [Scilit]
  40. Spriet, M.; Langrenez, C.; Brummelhuis, R.; De Bièvre, S. What is special about the Kirkwood-Dirac distribution. arXiv 2025, arXiv:2511.01996. [Google Scholar] [CrossRef] [Scilit]
  41. Jordan, A.N.; Arvidsson-Shukur, D.R.M.; Steinberg, A.M. Theory of direct measurement of the quantum pseudo-distribution via its characteristic function. arXiv 2026, arXiv:2602.06145. [Google Scholar] [CrossRef] [Scilit]
  42. Arvidsson-Shukur, D.R.M.; Yunger Halpern, N.; Lepage, H.V.; Lasek, A.A.; Barnes, C.H.W.; Lloyd, S. Quantum advantage in postselected metrology. Nat. Commun. 2020, 11, 3775. [Google Scholar] [CrossRef] [Scilit]
  43. Lupu-Gladstein, N.; Yilmaz, Y.B.; Arvidsson-Shukur, D.R.M.; Brodutch, A.; Pang, A.O.; Steinberg, A.M.; Yunger Halpern, N. Negative quasiprobabilities enhance phase estimation in quantum-optics experiment. Phys. Rev. Lett. 2022, 128, 220504. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  44. Arvidsson-Shukur, D.R.M.; McConnell, A.G.; Yunger Halpern, N. Nonclassical Advantage in Metrology Established via Quantum Simulations of Hypothetical Closed Timelike Curves. Phys. Rev. Lett. 2023, 131, 150202. [Google Scholar] [CrossRef] [Scilit]
  45. Dressel, J.; Jordan, A.N. Significance of the Imaginary Part of the Weak Value. Phys. Rev. A 2012, 85, 012107. [Google Scholar] [CrossRef] [Scilit]
  46. Langrenez, C.; Arvidsson-Shukur, D.R.M.; De Bièvre, S. Characterizing the geometry of the Kirkwood–Dirac-positive states. J. Math. Phys. 2024, 65, 072201. [Google Scholar] [CrossRef] [Scilit]
  47. Ogawa, K.; Abe, N.; Kobayashi, H.; Tomita, A. Complex counterpart of variance in quantum measurements for pre- and postselected systems. Phys. Rev. Res. 2021, 3, 033077. [Google Scholar] [CrossRef] [Scilit]
  48. Tsang, M. Generalized conditional expectations for quantum retrodiction and smoothing. Phys. Rev. A 2022, 105, 042213. [Google Scholar] [CrossRef] [Scilit]
  49. Tsang, M. Operational meanings of a generalized conditional expectation in quantum metrology. Quantum 2023, 7, 1162. [Google Scholar] [CrossRef] [Scilit]
  50. Dressel, J.; Jordan, A.N. Contextual-Value Approach to the Generalized Measurement of Observables. Phys. Rev. A 2012, 85, 022123. [Google Scholar] [CrossRef] [Scilit]
  51. Williams, N.S.; Jordan, A.N. Weak Values and the Leggett-Garg Inequality in Solid-State Qubits. Phys. Rev. Lett. 2008, 100, 026804. [Google Scholar] [CrossRef] [Scilit]
  52. Jozsa, R. Complex weak values in quantum measurement. Phys. Rev. A 2007, 76, 044103. [Google Scholar] [CrossRef] [Scilit]
  53. Dressel, J.; Jordan, A.N. Weak Values Are Universal in Von Neumann Measurements. Phys. Rev. Lett. 2012, 109, 230402. [Google Scholar] [CrossRef] [Scilit]
  54. Umegaki, H. Conditional expectation in an operator algebra, I. Tohoku Math. J. 1954, 6, 177–181. [Google Scholar] [CrossRef] [Scilit]
  55. Kraus, K. Complementary observables and uncertainty relations. Phys. Rev. D 1987, 35, 3070–3075. [Google Scholar] [CrossRef] [Scilit]
  56. De Bièvre, S. Complete Incompatibility, Support Uncertainty, and Kirkwood-Dirac Nonclassicality. Phys. Rev. Lett. 2021, 127, 190404. [Google Scholar] [CrossRef] [Scilit]
  57. De Bièvre, S. Relating incompatibility, noncommutativity, uncertainty, and Kirkwood–Dirac nonclassicality. J. Math. Phys. 2023, 64, 022202. [Google Scholar] [CrossRef] [Scilit]
  58. Giovannetti, V.; Lloyd, S.; Maccone, L. Quantum-Enhanced Measurements: Beating the Standard Quantum Limit. Science 2004, 306, 1330–1336. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  59. Giovanetti, V.; Lloyd, S.; Maccone, L. Advances in quantum metrology. Nat. Photonics 2011, 5, 222–229. [Google Scholar] [CrossRef] [Scilit]
  60. Smith, J.G.; Barnes, C.H.W.; Arvidsson-Shukur, D.R.M. Adaptive Bayesian quantum algorithm for phase estimation. Phys. Rev. A 2024, 109, 042412. [Google Scholar] [CrossRef] [Scilit]
  61. Fisher, R.A. Theory of Statistical Estimation. Math. Proc. Camb. Philos. Soc. 1925, 22, 700–725. [Google Scholar] [CrossRef] [Scilit]
  62. Braunstein, S.L.; Caves, C.M. Statistical distance and the geometry of quantum states. Phys. Rev. Lett. 1994, 72, 3439–3443. [Google Scholar] [CrossRef] [Scilit]
  63. Helstrom. Quantum Detection and Estimation Theory; Academic Press: Cambridge, MA, USA, 1976. [Google Scholar]
  64. Holevo, A.S. Probabilistic and Statistical Aspects of Quantum Theory; North-Holland Publishing Company: Amsterdam, The Netherlands, 1982. [Google Scholar]
  65. Bhatia, R.; Davis, C. A Better Bound on the Variance. Am. Math. Mon. 2000, 107, 353–357. [Google Scholar] [CrossRef]
  66. Thio, J.J.; Yang, S.; De Bièvre, S.; Barnes, C.H.; Arvidsson-Shukur, D.R. Kirkwood-Dirac Nonpositivity is a Necessary Resource for Quantum Computing. arXiv 2025, arXiv:2506.08092. [Google Scholar] [CrossRef] [Scilit]
  67. Pashayan, H.; Wallman, J.J.; Bartlett, S.D. Estimating Outcome Probabilities of Quantum Circuits Using Quasiprobabilities. Phys. Rev. Lett. 2015, 115, 070501. [Google Scholar] [CrossRef] [Scilit]
  68. Delfosse, N.; Allard Guerin, P.; Bian, J.; Raussendorf, R. Wigner Function Negativity and Contextuality in Quantum Computation on Rebits. Phys. Rev. X 2015, 5, 021003. [Google Scholar] [CrossRef] [Scilit]
  69. Langrenez, C.; Salmon, W.; De Bièvre, S.; Thio, J.J.; Long, C.K.; Arvidsson-Shukur, D.R.M. The set of Kirkwood-Dirac positive states is almost always minimal. arXiv 2024, arXiv:2405.17557. [Google Scholar] [CrossRef] [Scilit]
  70. Parzygnat, A.J.; Fullwood, J. From Time-Reversal Symmetry to Quantum Bayes’ Rules. PRX Quantum 2023, 4, 020334. [Google Scholar] [CrossRef] [Scilit]
  71. Von Neumann, J. Mathematical Foundations of Quantum Mechanics: New Edition; Princeton University Press: Princeton, NJ, USA, 2018; Volume 53. [Google Scholar]
  72. Leggett, A.J. Comment on “How the result of a measurement of a component of the spin of a spin-1/2 particle can turn out to be 100”. Phys. Rev. Lett. 1989, 62, 2325. [Google Scholar] [CrossRef] [Scilit]
  73. Peres, A. Quantum measurements with postselection. Phys. Rev. Lett. 1989, 62, 2326. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  74. Duck, I.M.; Stevenson, P.M.; Sudarshan, E.C.G. The sense in which a “weak measurement” of a spin-½ particle’s spin component yields a value 100. Phys. Rev. D 1989, 40, 2112–2117. [Google Scholar] [CrossRef] [Scilit]
  75. Tamir, B.; Cohen, E. Introduction to Weak Measurements and Weak Values. Quanta 2013, 2, 7–17. [Google Scholar] [CrossRef] [Scilit]
  76. Dressel, J.; Malik, M.; Miatto, F.M.; Jordan, A.N.; Boyd, R.W. Colloquium: Understanding quantum weak values: Basics and applications. Rev. Mod. Phys. 2014, 86, 307. [Google Scholar] [CrossRef] [Scilit]
  77. Paris, M. Quantum estimation for quantum technology. Int. J. Quantum Inf. 2009, 7, 125–137. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Spriet, M.; Langrenez, C.; Brummelhuis, R.; De Bièvre, S. What Is Special About the Kirkwood–Dirac Distributions? Only They Produce Natural Conditional Expectations. Symmetry 2026, 18, 1008. https://doi.org/10.3390/sym18061008

AMA Style

Spriet M, Langrenez C, Brummelhuis R, De Bièvre S. What Is Special About the Kirkwood–Dirac Distributions? Only They Produce Natural Conditional Expectations. Symmetry. 2026; 18(6):1008. https://doi.org/10.3390/sym18061008

Chicago/Turabian Style

Spriet, Matéo, Christopher Langrenez, Raymond Brummelhuis, and Stephan De Bièvre. 2026. "What Is Special About the Kirkwood–Dirac Distributions? Only They Produce Natural Conditional Expectations" Symmetry 18, no. 6: 1008. https://doi.org/10.3390/sym18061008

APA Style

Spriet, M., Langrenez, C., Brummelhuis, R., & De Bièvre, S. (2026). What Is Special About the Kirkwood–Dirac Distributions? Only They Produce Natural Conditional Expectations. Symmetry, 18(6), 1008. https://doi.org/10.3390/sym18061008

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop