Next Article in Journal
Boltzmann–Loschmidt Dispute Reloaded: Quantum 150 Years Later
Previous Article in Journal
S2-HGNN: Scale-Aware Hypergraph Node Classification with Spectral Inductive Bias
Previous Article in Special Issue
Coded Caching Scheme for Multiaccess Cache-Assisted Partially Connected Linear Network via Multi-Antenna Placement Delivery Array
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Rate–Distortion Limits for Task-Oriented Compression with Side Information

1
School of Cyber Science and Engineering, Southeast University, Nanjing 210096, China
2
Ningbo Institute of Digital Twin, Eastern Institute of Technology, Ningbo 315200, China
3
Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Ningbo 315200, China
4
School of Economics, University of Nottingham Ningbo China, Ningbo 315100, China
*
Author to whom correspondence should be addressed.
Entropy 2026, 28(6), 593; https://doi.org/10.3390/e28060593
Submission received: 20 April 2026 / Revised: 21 May 2026 / Accepted: 22 May 2026 / Published: 26 May 2026
(This article belongs to the Special Issue Network Information Theory and Its Applications)

Abstract

This paper analyzes the semantic rate–distortion problem motivated by task-oriented data compression with side information. The semantic information related to a task is not directly accessible to the encoder but implicitly impacts the observations through a joint probability distribution. The decoder aims to simultaneously recover the observation and infer the semantic information under certain distortion constraints. Notably, this paper advances the related research by involving side information and the observation of two semantic segments at both the encoder and decoder, which significantly complicates the theoretic analysis. We establish the information-theoretic limits for the tradeoff between compression rates and distortions by fully characterizing the rate–distortion function. Additionally, we explicitly derive the corresponding rate–distortion functions under specific Markov conditions for two scenarios: (i) the task is a binary classification of an integer observation as even and odd; and (ii) Gaussian-correlated task and observation. Furthermore, we validate the information-theoretic analysis by conducting a classification-oriented lossy image compression based on deep learning. The results are consistent with theoretical expectations, demonstrating the effectiveness of side information on both distortion and classification accuracy and the rationality of semantic segmentation.

1. Introduction

It is known that lossy source coding under given fidelity criterion was introduced in Shannon’s landmark paper [1], and the rate–distortion function was further proposed in [2], characterizing the optimal tradeoff between compression rates and distortion measurements from the perspective of mutual information. However, the classical lossy compression aims to reconstruct only the original source under certain distortion constraints, which is not aligned with the efficiency requirements of intelligent systems in the current AI era. Taking the traffic violation detection as an example, one would expect that a single-step compression algorithm optimized for the detection accuracy would be significantly more efficient than the straightforward recover-and-detect approach, where the video/image is first decompressed before detection. Indeed, this has been validated in related studies, e.g., [3,4,5].
Therefore, task-oriented or semantic compression [6,7,8,9] tailored for a specific task (e.g., video inference, decision making, classification, etc.) is gaining increasing attention. As shown in [3,4,5], the compression efficiency is improved by recovering only the most relevant semantic information corresponding to a given task rather than the entire source as in classical Shannon rate–distortion setups. Accordingly, in order to build its information-theoretic foundation, the classical rate–distortion theory has to be revisited considering the emerging semantic constraints.
Recently, the classical indirect source coding problem [10] was revisited from the semantic point of view, establishing the theoretical limits for point-to-point semantic compression in [11]. Specifically, they considered the lossy compression of sources ( S , X ) , where S is interpreted as intrinsic semantic information and X is the corresponding extrinsic observation. Leveraging the classical indirect rate–distortion theory, they characterized the semantic rate–distortion function, when the decoder reconstructs both S ^ and X ^ with two separate distortion constraints. Subsequently, the authors of [12] investigated the impacts of different distortion measures, especially the context-dependent distortion in semantic compression. Moreover, a variant of the Blahut–Arimoto algorithm was developed to compute the point-to-point semantic rate–distortion functions for discrete sources [12]. Most recently, the semantic compression with rate–distortion, perception and classification constraints was studied [13] from the perspective of information bottleneck. In addition, the semantic lossy compression has been extended to the joint source-channel coding scenario [14].
Nevertheless, the aforementioned information-theoretic analysis solely concentrates on the point-to-point semantic compression. Tracing the footsteps of classical rate–distortion theory research, the introduction of side information could greatly improve compression efficiency. The rate–distortion function was investigated when side information is available at the encoder or/and decoder; see [15,16,17,18] and the references therein.
Specifically, in the case where the side information is only available at the encoder, then no benefit could be achieved. In contrast, when side information is only accessible at the decoder, the corresponding rate–distortion function was considered by Wyner and Ziv in [17] with its extensions being discussed in [18]. Finally, if both the encoder and decoder have access to the same side information, the optimal tradeoff is called conditional rate–distortion function, which was given by [15,16]. In addition, Slepian and Wolf [19] solved a more general problem of distributed compression. In practice, lossy source coding with side information has been applied in video compression standards such as H.264 and HEVC, as well as in learning-based multi-view video compression [20].
Given the pivotal role of side information in classical rate–distortion theory, it is natural to wonder whether it has a positive impact on semantic lossy compression. This paper affirms that it does by introducing side information Y to both the encoder and decoder and fully characterizing the corresponding rate–distortion function. In doing so, this paper also advances the prior point-to-point semantic rate–distortion theory [11,12,13] into the distributed case.
Moreover, considering the practical scenario where the observation (e.g., images in the street view house numbers (SVHN) dataset) typically contains different semantic information (e.g., numbers, colors, background elements such as walls) that indicates varying levels of relevance for a given task (e.g., the numbers are more relevant than the background for classification tasks), this paper further generalizes the theoretical analysis by partitioning the observation as a pair of variables ( X 1 , X 2 ) to model different relevance levels. Without loss of generality, X 1 denotes the most relevant semantic in the observation and X 2 represents the remaining. In practical applications such as autonomous driving and intelligent surveillance, although downstream machine vision tasks primarily rely on the highly relevant features in X 1 , recovering the background or appearance segment X 2 remains crucial for human operators to maintain situational awareness. Under resource-constrained channels, a hybrid human–machine vision system can allocate fewer bits to X 2 (by allowing a larger distortion D 2 ) while ensuring high fidelity for X 1 and high accuracy for task inference. We point out that [21] has already investigated the varying relevance of different objects in an image for downstream machine vision tasks by proposing a semantically disentangled image coding framework, which necessitates the information-theoretic observation model in this paper.
In addition to the theoretical semantic rate–distortion limits, we implement a classification-oriented image compression pipeline using autoencoders. The extensive experimental results confirm the positive impact of side information on both classification accuracy and distortion, and they affirm the rationality of the generalized observation ( X 1 , X 2 ) , which aligns with the derived theoretical analysis.
To summarize, this paper primarily advances the information-theoretic analysis of semantic lossy compression with side information, particularly by modeling the semantic source as ( S , X 1 , X 2 , Y ) , which results in a notably non-trivial characterization of rate–distortion functions. The main contributions are listed as follows:
  • We formulate a theoretical semantic rate–distortion framework involving side information Y and the observation of two semantic segments ( X 1 , X 2 ) . The corresponding optimal rate–distortion function is fully characterized, which is followed by an exploration of some key properties.
  • We reveal that the separate compression of X 1 and X 2 is optimal if they are conditionally independent given the side information Y.
  • The rate–distortion functions are explicitly discussed for cases when the following apply: (i) ( S , X 2 , Y ) are binary sources and X 1 is an integer source (i.e., an odd-even classification of integers) and (ii) ( S , X 1 , X 2 , Y ) are all Gaussian sources.
  • An autoencoder-based image compression scheme is implemented for a classification task. The experimental results substantiate the positive effect of side information on both distortion and classification accuracy, which is in line with the theoretical analysis.
The remainder of this paper is organized as follows. We first formulate the problem and present some preliminary results in Section 2. In Section 3, we characterize the rate–distortion function and some useful properties. Numerical evaluations of the rate–distortion functions for binary, integer, and Gaussian sources are presented in Section 4.1 and Section 4.2, respectively. We conduct experiments on a deep learning-based image compression scheme for a classification task in Section 5. Finally, the paper is concluded in Section 6. Some essential proofs can be found in the appendices.

2. Problem Formulation and Preliminaries

2.1. Problem Formulation

The proposed theoretical semantic compression model is illustrated in Figure 1, where the side information is known at both the encoder and decoder sides.
The problem is formally defined as follows. A collection of discrete memoryless sources (DMS) is described by generic random variables ( S , X 1 , X 2 , Y ) taking values in finite alphabets S × X 1 × X 2 × Y according to probability distribution p ( x 1 , x 2 , y ) p ( s | x 1 ) . In particular, this indicates the Markov chain S X 1 ( X 2 , Y ) . We interpret S as a latent variable, i.e., the intrinsic semantic information (e.g., the state of a cognitive system), which is not observable by the encoder. We assume that the extrinsic observation consists of two parts:
  • X 1 dynamically varies according to the semantic information S, which captures the most relevant segment, e.g., the cars and red lights in a frame showing a traffic violation;
  • X 2 is the remaining “appearance” of the observation, e.g., the remaining elements in the frame capturing the violation.
We interpret Y as the side information that improves compression efficiency such as previous frames in a video. For length-n source sequences, ( S n , X 1 n , X 2 n , Y n ) , the encoder has access to only the observed ones ( X 1 n , X 2 n , Y n ) and outputs the encoded message W. Upon observing local information Y n and receiving W, the decoder reconstructs the source sequences as ( X ^ 1 n , X ^ 2 n ) drawn values from X ^ 1 × X ^ 2 , within distortions D 1 and D 2 , respectively. Given the reconstructions, the classifier aims to recover the semantic information as S ^ n from alphabet S ^ with distortion constraint D s . Here, for simplicity, we assume a perfect classifier; i.e., it is equivalent to recovering S ^ n directly at the decoder.
Formally, an n , 2 n R code is defined by the encoding function
E n : X 1 n × X 2 n × Y n { 1 , 2 , , 2 n R }
and the decoding function
D e : { 1 , 2 , , 2 n R } × Y n X ^ 1 n × X ^ 2 n × S ^ n ,
where R is the coding rate. Let R + be the set of non-negative real numbers. We consider bounded per-letter distortion functions d 1 : X 1 × X ^ 1 R + , d 2 : X 2 × X ^ 2 R + , and d s : S × S ^ R + . The distortions between two length-n sequences are defined as average distortion, e.g.,
d 1 ( x 1 n , x ^ 1 n ) 1 n i = 1 n d 1 ( x 1 , i , x ^ 1 , i ) .
A non-negative rate–distortion tuple ( R , D 1 , D 2 , D s ) is said to be achievable if for sufficiently large n, there exists an n , 2 n R code such that
lim n E d 1 ( X 1 n , X ^ 1 n ) D 1 ,   lim n E d 2 ( X 2 n , X ^ 2 n ) D 2 ,   lim n E d s ( S n , S ^ n ) D s .
The semantic rate–distortion function R ( D 1 , D 2 , D s ) is the infimum of coding rate R for distortions ( D 1 , D 2 , D s ) such that the quadruple ( R , D 1 , D 2 , D s ) is achievable. The goal of this paper is to completely characterize R ( D 1 , D 2 , D s ) .

2.2. Preliminaries

2.2.1. Conditional Rate–Distortion Function

The elegant rate–distortion function was investigated and fully characterized in [2]. Assume the length-n source sequence X n is independent and identically distributed (i.i.d.) over X with generic random variable X and let d : X × X ^ R + be a bounded per-letter distortion measure. The rate–distortion function for a given distortion criterion D is given by
R ( D ) = min p ( x ^ | x ) : E d ( X , X ^ ) D I ( X ; X ^ ) .
It was proved in [2] and also introduced in [22,23] that R ( D ) is a non-increasing and convex function of D.
If both the encoder and decoder are allowed to observe side information Y n (with generic variable Y over Y jointly distributed with X), then the tradeoff is called the conditional rate–distortion function [15,16], which is characterized as
R X | Y ( D ) = min p ( x ^ | x , y ) : E d ( X , X ^ ) D I ( X ; X ^ | Y ) .
If ( X , Y ) is a doubly symmetric binary source (DSBS) with parameter p 0 , i.e.,
p ( x , y ) = 1 p 0 2 ,   if   x = y p 0 2 ,   if   x y .
then the conditional rate–distortion function is given in [16] by
R X | Y ( D ) = h b ( p 0 ) h b ( D ) · 1 0 D p 0 ,
where h b ( q ) = q log q ( 1 q ) log ( 1 q ) is the entropy for a Bernoulli(q) distribution and 1 A is the indicator function of whether event A happens.

2.2.2. Rate–Distortion Function with Two Constraints

This scenario was discussed in [23] that one wishes to encode the i.i.d. source sequence X n at rate R and recover two reconstructions X ^ a n and X ^ b n with distortion criteria E d a ( X n , X ^ a n ) D a and E d b ( X n , X ^ b n ) D b , respectively. The rate–distortion function is given by
R 2 d ( D a , D b ) = min p ( x ^ a , x ^ b | x ) :   E d a ( X , X ^ a ) D a E d b ( X , X ^ b ) D b I ( X ; X ^ a , X ^ b ) .
Comparing (1) and (5), one easily obtains that
max { R ( D a ) , R ( D b ) } R 2 d ( D a , D b ) R ( D a ) + R ( D b ) .
For the special case where X ^ a = X ^ b and d a ( x , x ^ ) = d b ( x , x ^ ) for all x X and x ^ X ^ a , it suffices to recover only one sequence X ^ a n = X ^ b n with distortion min { D a , D b } . Then, both distortion constraints are satisfied since E d a ( X n , X ^ a n ) = E d b ( X n , X ^ b n ) = min { D a , D b } , which is upper bounded by both D a and D b . This implies
R 2 d ( D a , D b ) = R ( min { D a , D b } ) = max { R ( D a ) , R ( D b ) } ,
where the second equality follows from the non-increasing property of R ( D ) .
When side information is available at the decoder for only one of the two reconstructions, e.g., X ^ b , it was proved in [18] that successive encoding (first X ^ a , then X ^ b ) is optimal. For the case where the two reconstructions have access to different side information, respectively, the rate–distortion tradeoff was characterized in [22].

2.2.3. Rate–Distortion Function of Two Sources

The problem of compressing two i.i.d. source sequences X a n and X b n at the same encoder is considered in [23] [Problem. 10.14]. The rate–distortion function is given therein, i.e.,
R 2 s ( D a , D b ) = min p ( x ^ a , x ^ b | x a , x b ) :   E d a ( X a , X ^ a ) D a E d b ( X b , X ^ b ) D b I ( X a , X b ; X ^ a , X ^ b ) .
It is also shown that for two independent sources, compressing simultaneously is the same as compressing separately in terms of the rate and distortions, i.e.,
R 2 s ( D a , D b ) = R ( D a ) + R ( D b ) .
If the two sources are dependent, the equality in (8) can be false, and the Slepian–Wolf rate region [19] indicates that joint entropy of the two source variables is sufficient and optimal for lossless reconstructions. Taking into account distortions, Gray showed via an example in [16] that the compression rate can be strictly larger than R ( D a ) + R X b | X a ( D b ) in general. At last, some related results for compressing compound sources can be found in [24].

3. Optimal Rate–Distortion Tradeoff

3.1. The Complete Rate–Distortion Function

Theorem 1. 
The complete rate–distortion function for semantic compression with side information is given as the solution to the following optimization problem
R ( D 1 , D 2 , D s ) = min I ( X 1 , X 2 ; X ^ 1 , X ^ 2 , S ^ | Y )
  s . t . E d 1 ( X 1 , X ^ 1 ) D 1 ,       E d 2 ( X 2 , X ^ 2 ) D 2 ,       E d s ( X 1 , S ^ ) D s ,
where the minimum is taken over all conditional pmf p ( x ^ 1 , x ^ 2 , s ^ | x 1 , x 2 , y ) and the modified distortion measure is defined by
d s ( x 1 , s ^ ) = 1 p ( x 1 ) s S p ( x 1 , s ) d s ( s , s ^ ) .
Proof. 
We provide a rigorous technical proof in Appendix A. □
Remark 1. 
We can interpret the problem as the combination of rate–distortion with two sources ( X 1 and X 2 ), rate–distortion with two constraints ( X 1 is recovered with two constraints D 1 and D s ), and conditional rate–distortion (conditioning on Y). Then, Theorem 1 can be obtained informally by combining the rate–distortion functions in (2), (5), and (7).
Remark 2. 
Theorem 1 considers the observation of two semantic segments, and it can be easily extended to arbitrary T segments by expanding the notations X t and X ^ t , 1 t T .

3.2. Some Properties

Similar to the rate–distortion function (1), we collect some properties in the following lemma. The proof simply follows the same procedure as that for (1) in [2,22,23]. We omit the details here.
Lemma 1. 
The rate–distortion function R ( D 1 , D 2 , D s ) is non-increasing and convex in ( D 1 , D 2 , D s ) .
Recall from (8) that compressing two independent sources is the same as compressing them simultaneously. Then, one may naturally ask whether the optimality of separate compression remains to hold here. We answer the question in the following lemma.
Lemma 2. 
If X 1 Y X 2 forms a Markov chain, then
R ( D 1 , D 2 , D s ) = R 2 d , X 1 | Y ( D 1 , D s ) + R X 2 | Y ( D 2 ) ,
where the conditional rate–distortion function with two constraints is given by
R 2 d , X 1 | Y ( D 1 , D s ) = min p ( x ^ 1 , s ^ | x 1 , y ) : E d 1 ( X 1 , X ^ 1 ) D 1 E d s ( X 1 , S ^ ) D s I ( X 1 ; X ^ 1 , S ^ | Y )
and the conditional rate–distortion function is given in (2) and can be written as
R X 2 | Y ( D ) = min p ( x ^ 2 | x 2 , y ) : E d 2 ( X 2 , X ^ 2 ) D 2 I ( X 2 ; X ^ 2 | Y ) .
Proof. 
The proof is given in Appendix B. □
Remark 3. 
Compared to the unconditional independence assumption required for the equality in (8), the optimality of separate compression in Lemma 2 requires conditional independence I ( X 1 ; X 2 | Y ) = 0 , which is equivalent to the Markov chain X 1 Y X 2 . Moreover, this is the necessary and sufficient condition for separate compression optimality in the presence of side information Y. This is intuitive because both the encoder and decoder have access to the side information Y, meaning the rate–distortion tradeoff is evaluated under the conditional probability space given Y.
Direct (unconditional) independence of X 1 and X 2 is neither necessary nor sufficient for separate compression to be optimal here. On one hand, if X 1 and X 2 are unconditionally independent but conditionally dependent given Y (for instance, when X 1 and X 2 are independent Bernoulli (0.5) variables and Y = X 1 X 2 ), joint compression conditioned on Y can exploit this conditional correlation to achieve a strictly lower rate than separate compression. On the other hand, X 1 and X 2 can be unconditionally dependent, but as long as their mutual correlation is fully mediated by the side information Y (i.e., X 1 Y X 2 holds), separate compression remains optimal.

3.3. Rate–Distortion Function for Semantic Information Only

The indirect rate–distortion problem can be viewed as a special case of Theorem 1 that only recovers the semantic information S; i.e., X 2 and Y are constants and D 1 = . Denote the minimum achievable rate for a given distortion constraint D s by R s ( D s ) .
Consider doubly symmetric binary sources S and X 1 , i.e.,
p ( s , x 1 ) = 1 p 2 ,   if   s = x 1 p 2 ,   if   s x 1 .
Without loss of generality, assume p 0.5 , which means that X 1 has a higher probability to reflect the same value as S. Let d s : S × S ^ { 0 , 1 } be the Hamming distortion measure. Then, the evaluation of R s ( D s ) is given in the following lemma.
Let R d s ( · ) be the Shannon rate–distortion function in (1) under the distortion measure d s (c.f. (10)). For notational simplicity, and for D s p , define
D s 0 D s p 1 2 p .
Lemma 3. 
For binary sources in (11) and Hamming distortion, the rate–distortion function for semantic information is
R s ( D s ) = R d s ( D s ) ,
where R d s ( D s ) = R D s 0 = 1 h b D s p 1 2 p · 1 p D s 0.5 .
Proof. 
The evaluation of R s ( D s ) was given in [25]. A simpler proof is in Appendix C. □
Remark 4. 
By the properties of the rate–distortion function in (1) and the linearity between D s 0 and D s , we see that R s ( D s ) is also non-increasing and convex in D s .
Remark 5. 
It is easy to check that D p 1 2 p < D for D < 0.5 . This implies that R s ( D ) > R ( D ) for D < 0.5 , where R ( D ) is the Shannon rate–distortion function (1). The inequality is intuitive from the data processing inequality that under the same distortion constraint D, recovering S directly (with rate R ( D ) ) is easier than recovering it from the observation X 1 (with rate R s ( D ) ). Moreover, we see from the lemma that D s p , which means that the semantic information can never be losslessly recovered for p > 0 . This can be deduced from the fact that even though we know the complete information of X 1 , the best distortion for reconstructing S is the distortion between S and X 1 , which is equal to p.
The rate–distortion functions R s ( D ) and R ( D ) are illustrated in Figure 2 for p = 0.1 , which verifies the above observations. For general source and distortion measure, we have R s ( D ) R ( D ) where the equality holds only when X 1 determines S. This can be easily proved by the data processing inequality, and we omit the details here.
Remark 6. 
We can imagine that d s measures the distortion between the observation and reconstruction of semantic information. Furthermore, it was shown in [10,11] that d s and d s measure equivalent distortions, i.e.,
E d s ( X 1 , S ^ ) = E d s ( S , S ^ ) ,         E d s ( X 1 n , S ^ n ) = E d s ( S n , S ^ n ) .
Then, we can regard the system of compressing X 1 n and reconstructing S ^ n as the Shannon rate–distortion problem with distortion measure d s . Thus, R s ( D s ) is equivalent to the Shannon rate–distortion function in (1) under distortion measure d s , which rigorously proves (13).

4. Case Studies

4.1. Binary Semantic Sources

Assume S and X 1 are doubly symmetric binary sources with distribution in (11) and X 2 and Y are both Bernoulli ( 1 2 ) sources. The reconstructions are all binary, i.e., X ^ 1 = X ^ 2 = S ^ = { 0 , 1 } . The distortion measures d 1 , d 2 , and d s are all assumed to be Hamming distortion. We further assume that any two of X 1 , X 2 , and Y are doubly symmetric binary distributed with parameters p 1 , p 2 , and p 3 , respectively. Specifically, ( X 1 , X 2 ) DSBS ( p 1 ) , ( X 1 , Y ) DSBS ( p 2 ) , ( X 2 , Y ) DSBS ( p 3 ) .

4.1.1. Conditionally Independent Binary Sources

We further assume the Markov chain X 1 Y X 2 ; i.e., X 1 and X 2 are conditionally independent given Y. This assumption coincides with the intuitive understanding of X 1 and X 2 in Section 2.1, indicating that the most relevant semantic segment can be independent with the remaining parts. (Note that the Markov chain X 1 Y X 2 indicates p 1 = p 2 p 3 p 2 ( 1 p 3 ) + p 3 ( 1 p 2 ) .) Then, from Lemma 2, compressing X 1 n and X 2 n simultaneously is the same as compressing them separately in terms of the optimal compression rate and distortions, which implies the following theorem.
Theorem 2. 
For the conditionally independent sources satisfying the Markov chain X 1 Y X 2 , the rate–distortion function is given by
R ( D 1 , D 2 , D s ) = h b ( p 3 ) h b ( D 2 ) · 1 0 D 2 p 3 + h b ( p 2 ) h b min { D 1 , D s 0 } · 1 0 min { D 1 , D s 0 } p 2 ,
where D s 0 = D s p 1 2 p is defined in (12).
Proof. 
The rate–distortion function in Theorem 1 satisfies
R ( D 1 , D 2 , D s )   = R 2 d , X 1 | Y ( D 1 , D s ) + R X 2 | Y ( D 2 )   = min E d 1 ( X 1 , X ^ 1 ) D 1 E d s ( X 1 , S ^ ) D s I ( X 1 ; X ^ 1 , S ^ | Y ) + min E d 2 ( X 2 , X ^ 2 ) D 2 I ( X 2 ; X ^ 2 | Y )   = h b ( p 2 ) h b min { D 1 , D s 0 } · 1 0 min { D 1 , D s 0 } p 2             + h b ( p 3 ) h b ( D 2 ) · 1 0 D 2 p 3 ,
where the last step follows from the rate–distortion functions in (4), (6) and (13). □

4.1.2. Binary Classification of Integers

We consider the classification of integers as even or odd. Let X 1 be uniformly distributed over X 1 = [ 1 : N ] with N 4 being an even integer. The semantic information S is a binary random variable that probabilistically indicates whether X 1 is even or odd. The transition probability can be defined by a parameter p, which is similar to that in (11) by replacing the value of X 1 with “even” and “odd”. The binary side information Y is correlated with X 1 and also indicates its oddity (even/odd) with parameter p 2 .
Assume the Markov chain X 1 Y X 2 holds, and the Bernoulli( 1 2 ) source X 2 is independent with Y. We can verify that X 2 is independent with ( X 1 , Y ) . By Lemma 2, compressing X 1 n and X 2 n simultaneously is the same as compressing them separately. For simplicity, we consider only small distortions such that ( D 1 , D 2 , D s ) D 1 where
D 1 = ( D 1 , D 2 , D s ) : 0 min { D 1 , D s 0 } 2 ( N 1 ) p 2 N   and   0 D 2 0.5 .
Theorem 3. 
For ( D 1 , D 2 , D s ) D 1 , the rate–distortion function for integer classification is
R ( D 1 , D 2 , D s ) = h b ( p 2 ) + log ( N / 2 ) h b ( min { D 1 , D s 0 } ) D 1 log ( N 1 ) + 1 h b ( D 2 ) ,
where D s 0 = D s p 1 2 p is defined in (12).
Proof. 
The proof is given in Appendix D. □

4.1.3. Numerical Results for Binary Classification

The rate–distortion function for integer classification in Theorem 3 with p = p 1 = p 2 = 0.25 , D 2 = 0.5 , and N = 8 is illustrated in Figure 3. Note that in Theorem 3, D 2 = 0.5 indicates that X 2 can be recovered by random guessing, which further implies that X 2 can also be regarded as side information at both sides.
By comparing the rates along the D 1 and D s axes in Figure 3, it is evident that recovering only the semantic information can reduce the coding rate compared to recovering the original source message.

4.2. Gaussian Sources

Assume S and X 1 are jointly Gaussian sources with zero mean and covariance matrix
σ S σ S X 1 σ S X 1 σ X 1 .
Similarly, we assume the Markov chain X 1 Y X 2 where X 2 and Y are jointly Gaussian sources with zero mean and covariance matrix
σ X 2 σ X 2 Y σ X 2 Y σ Y .
Thus, X 1 is conditionally independent of X 2 given Y. Let the covariance of X 1 and Y be σ X 1 Y . The reconstructions are real scalars, i.e., X ^ 1 = X ^ 2 = S ^ = R , and the distortion metrics are squared error.

4.2.1. Theoretical Bounds and Dominance Analysis

Given the Markov chain, we see from Lemma 2 that compressing X 1 n and X 2 n simultaneously is the same as compressing them separately in terms of the optimal compression rate and distortions. Then, we have the following theorem.
Theorem 4. 
For the Gaussian sources, if the Markov chain X 1 Y X 2 holds, the rate–distortion function is
R ( D 1 , D 2 , D s ) = 1 2 log σ X 2 σ X 2 Y 2 σ Y D 2 + + 1 2 log max σ X 1 σ X 1 Y 2 σ Y D 1 , σ S X 1 2 σ X 1 σ X 1 Y 2 σ Y σ X 1 2 D s D MMSE + ,
where ( x ) + = max ( x , 0 ) and D MMSE is the minimum mean squared error for estimating S from X 1 , which is given by
D MMSE = σ S σ S X 1 2 σ X 1 .
Proof. 
The proof is given in Appendix E. □
To theoretically characterize the dominant roles of D 1 and D s in Theorem 4, we compare the two terms inside the max ( · , · ) operator in (17). The boundary between the reconstruction-dominant and semantic-dominant regions is given by the linear threshold:
D s = D M M S E + D 1 σ S X 1 2 σ X 1 2 .
Specifically, when D s D M M S E + D 1 σ S X 1 2 σ X 1 2 , the reconstruction constraint D 1 dominates the rate; otherwise, the semantic constraint D s dominates the rate.

4.2.2. Numerical Results for Gaussian Sources

We then present the numerical analysis of the Gaussian rate–distortion function in Theorem 4, where we take all of the variances as 2, all of the covariances as 1, and D 2 = 1 . The resultant optimal tradeoff between the coding rate R ( D 1 , D 2 , D s ) and distortions ( D 1 , D s ) is illustrated in Figure 4. It can be observed that the rate–distortion function is decreasing and convex in ( D 1 , D s ) , and the minimum rate is given as R X 2 | Y ( D 2 ) = 1 2 log σ X 2 σ X 2 Y 2 σ Y D 2 + = 0.20 .
The corresponding contour plot is also shown in Figure 4, where the slanted line represents the theoretical boundary in (19). Under our specific numerical setup, this boundary simplifies to the linear relation D s = 1.5 + 0.25 D 1 , which perfectly matches the slanted line starting at D s = 1.5 when D 1 = 0 and passing through the corner points of the L-shaped contour lines. As predicted by our theoretical analysis, when the distortion pair ( D 1 , D s ) lies in the reconstruction-dominant region above this slanted line (where D s 1.5 + 0.25 D 1 ), the rate is determined solely by the reconstruction distortion constraint of X 1 . Conversely, in the semantic-dominant region below the slanted line (where D s < 1.5 + 0.25 D 1 ), the rate only needs to meet the distortion constraint D s for reconstructing the semantic information S.

5. Experimental Results and Discussion

This section validates the theoretical analysis by implementing a deep learning-based image compression scheme for a classification task.

5.1. Deep Learning-Based Classification-Oriented Image Compression Scheme

By employing an autoencoder (AE), the lossy compression procedure can be realized through a parameterized encoder and decoder, which is denoted as p X ^ | X Y X ^ | X , Y . Let ϕ = { ϕ e , ϕ d } represent the parameters of the entire encoder–decoder network. Similar to [13,26], we adopt a neural network with a stochastic encoder and decoder together with universal quantization [27]. As shown in Figure 5, the system consists of an encoder E, a decoder D, and a classifier C.
In the experiments, we take the central region of the image as the most relevant semantic segment, i.e., X 1 , and the surrounding region of the image is regarded as X 2 . The size of X 1 is one half of each image; e.g., for an image of size 32 × 32 , the central region of 16 × 16 is designated as X 1 . Both X 1 and X 2 are fed into two identical encoder–decoder pairs without parameter sharing. In addition, the top 1 / 3 of each image is taken as side information Y, which is accessible by both the encoder and the decoder. The classifier takes only the estimation X ^ 1 of the most relevant semantic region as its input and predicts a vector S ^ , where each entry represents the probability that the image belongs to each class. The true label S and the predicted label S ^ are used to compute the classification accuracy. Moreover, we take the conventional mean squared error (MSE) as the distortion measure, i.e., E X X ^ 2 . As for the classification evaluation, we employ the commonly used cross-entropy (CE) loss, i.e.,
CE S , S ^ = S X ^ 1 p ϕ S , X ^ log 1 p ψ S | X ^ 1 = 1 N X ^ 1 log p ψ i X 1 | X ^ 1 ,
where S ^ denotes the output vector of the classifier given X ^ 1 , and its i-th entry corresponds to the conditional probability of the i-th class, i.e., p ψ S | X ^ 1 , with ψ being the parameters of the classifier. Given encoder–decoder parameter pair ( ϕ e , ϕ d ) , the reconstruction X ^ 1 = D ϕ d E ϕ e ( X 1 ) is a function of the corresponding input X 1 . Let the label of X 1 be i X 1 , and let the batch size be N. If S = i X 1 , then p ϕ S , X ^ 1 = 1 N ; otherwise, p ϕ S , X ^ 1 = 0 .
To evaluate the compression rate R, we follow the settings in [13,26] and define R as the upper bound k × log 2 ( L ) , where k denotes the dimensionality of the encoder output and L is the number of quantization levels for each entry. As discussed in [28], setting R to its upper bound greatly simplifies the framework and is found to be only slightly suboptimal. Finally, given a rate R = k × log 2 ( L ) , the overall loss function L of the deep learning-based image compression framework is defined as a weighted sum of MSE and CE, i.e.,
L λ d   E X X ^ + λ c   CE S , S ^ ,
where λ d and λ c are hyper-parameters balancing the weights of distortion and classification losses, respectively.

5.2. Results and Discussion

Figure 6 illustrates the tradeoff between MSE and classification accuracy on the MNIST and SVHN datasets, where λ d is fixed to 1 and the variations are introduced by changing λ c and R. Note that each color represents a fixed compression rate R, which is achieved by applying the same quantizer throughout the experiments. For the MNIST and SVHN datasets, we take ( k , L ) R { ( 2 , 4 ) 4 , ( 3 , 4 ) 6 , ( 4 , 4 ) 8 , ( 5 , 4 ) 10 , ( 3 , 8 ) 9 , ( 4 , 8 ) 12 , ( 5 , 8 ) 15 } , respectively, ( k , L ) R { ( 10 , 8 ) 30 , ( 15 , 8 ) 45 , ( 20 , 8 ) 60 , ( 25 , 8 ) 75 , ( 10 , 16 ) 40 , ( 20 , 16 ) 80 , ( 25 , 16 ) 100 } . Moreover, the weights λ c { 0.001 , 0.0025 , 0.004 , 0.005 , 0.006 , 0.008 , 0.009 , 0.01 , 0.011 , 0.013 } and λ c { 0.0001 , 0.0003 , 0.0005 , 0.0008 , 0.001 , 0.0012 , 0.0017 } are utilized for the MNIST and SVHN datasets, respectively. As shown in Figure 6, each point corresponds to an encoder–decoder pair and a classifier jointly trained with given R and λ c . For fair comparison, consider the following two cases:
  • Case A: the classifier takes only X ^ 1 as input;
  • Case B: the classifier takes the pair ( X ^ 1 , X ^ 2 ) as input.
The results of Case A are notably marked with a black outline. Note that the classification accuracy corresponds to D s , while the MSE distortions for Cases A and B correspond to D 1 and D 2 in Theorem 1, respectively. Next, we discuss the experimental results in detail.

5.2.1. Positive Impact of Side Information

Figure 7 compares the distortion MSE and classification accuracy of the semantic compression scheme when side information is incorporated versus when it is excluded. As expected, incorporating side information leads to improvements in both distortion MSE and classification accuracy, and these gains are more substantial in the low-rate regime. The performance gap between the two settings gradually diminishes with growing rate R. This indicates that for sufficiently large rates, both classification and distortion have already reached satisfactory levels, thereby reducing the relative benefit provided by side information.

5.2.2. Semantic Relevance Preserves Classification Accuracy

Recall that Case A employs only the most semantically relevant region, X ^ 1 , as input to the classifier. As shown in Table 1, the distortion MSE of Case A varies within a narrow range compared to Case B. Although a decrease in classification accuracy is observed, it is generally limited to less than 7.5%. Specifically, Case A yields a modest average increase of 0.29% in distortion MSE compared to Case B, which is accompanied by a slight decrease in classification accuracy of 1.28% for the MNIST dataset. Moreover, for the SVHN dataset, Case A leads to an average reduction of 2.53% in MSE and an accuracy decline of 6.18%.
These results align with our theoretical analysis, confirming that retaining the most critical semantic features does not substantially compromise classification performance. In addition, the visualization results presented in Figure 8 further affirm our findings: across different values of λ c , the reconstruction quality remains nearly identical between Case A and Case B.

5.2.3. Trade-Off Between Classification Accuracy and Distortion

It can be observed from Figure 6 that for both Case A and Case B, higher classification accuracy is generally achieved at the expense of increased distortion MSE. As shown in Table 1, when λ c increases from 0.001 to 0.01 for the MNIST dataset, the MSE increases by 6.47% (Case A) and 4.09% (Case B), while the classification accuracy improves by 11.93% (Case A) and 4.74% (Case B). Moreover, it is noted that with rising R, the fitting curves shift upward and to the left corner, indicating that both distortion and classification accuracy improve with larger rates. Taking the MNIST dataset with λ c = 0.001 as an example, when rate R increases from 4 to 15, the MSE improves by 52.63% (Case A) and 51.46% (Case B), while the classification accuracy grows by 3.80% (Case A) and 4.67% (Case B). Similar trends can also be found for the SVHN dataset.

5.2.4. The Role of Rate in Balancing Distortion and Classification

In addition to the MSE–accuracy pairs, Figure 6 also plots the corresponding MSE–accuracy fitting curves for each rate R with solid and dashed lines representing Case B and Case A, respectively.
We first discuss the results from the SVHN dataset (see Figure 6). Since the available transmission rate is limited in the low-rate regime (blue region), the compression system has to trade off between reconstruction quality (a smaller MSE corresponds to better reconstruction quality) and classification accuracy. In other words, relaxing the constraint on distortion (i.e., allowing a higher MSE) can contribute to improved classification accuracy. However, this tradeoff gradually weakens with increasing compression rate R. Notably, when R is sufficiently large (yellow region), the system would have adequate transmission resources, making it unnecessary to sacrifice reconstruction quality for accuracy improvement. Under such conditions, distortion MSE and classification accuracy exhibit a positive correlation; i.e., better reconstruction quality (lower MSE) results in higher classification accuracy (see the fitting curves with R > 45 ).
In contrast, due to the considerably lower image resolution and dimensionality of the MNIST dataset (approximately one fourth that of SVHN dataset), even when the compression rate is relatively low, the system still has sufficient resources to optimize simultaneously the reconstruction quality and classification performance.
Moreover, the simple and regular structure of MNIST images renders the classifier relatively less sensitive to minor distortions, whereas the more diverse content in SVHN images leads to the classification accuracy having greater susceptibility to compression-induced loss. These inherent differences account for the distinct MSE–accuracy relationships observed in Figure 6.

5.2.5. Practical Design Implications

The theoretical limits and case studies developed in this paper offer several critical guidelines for designing practical task-oriented semantic compression systems.
(i)
Distributed Network Architecture: Lemma 2 suggests that if different semantic segments of an observation (e.g., foreground X 1 and background X 2 in video frames) are conditionally independent given temporal side information Y (e.g., key frames), we can safely design separate, parallel network branches to compress them without losing rate–distortion optimality.
(ii)
Dynamic Bit Allocation: The dominance threshold boundary discussed in Theorem 4 indicates that when downstream task requirements are loose ( D s is large), the compression network can safely optimize for reconstruction quality ( D 1 ) using MSE loss alone. Conversely, when strict task accuracy is required, semantic classification loss must be heavily weighted to prevent task failure.
(iii)
Side Information Utilization: The significant rate savings shown in Figure 7 demonstrate that predicting and sharing correlated side information Y at both ends is highly effective for both saving bandwidth and preserving task-classification accuracy.

6. Conclusions

This paper investigated the semantic rate–distortion theory with side information and the observation of two semantic variables. The corresponding optimal rate–distortion function was fully characterized, which particularly advanced the information-theoretic analysis of existing point-to-point semantic compression. In addition, we explicitly provided rate–distortion functions for binary and Gaussian sources under common distortion measures, and we revealed that the recovery of the semantic information is more efficient than reconstructing the original source. Moreover, experimental results on a classification-oriented image compression model validate the positive impact of side information, as anticipated by the theoretical analysis. For future work, our approach will be extended to more complex multi-user systems like the distributed compression system, the multiple description system, where the side information is available only at the decoder, etc.

Author Contributions

Conceptualization, T.G. and H.W.; methodology, Z.S.; software, Z.S.; validation, T.G. and H.W.; formal analysis, T.G.; investigation, Z.S. and Y.L.; resources, T.G.; data curation, Z.S.; writing—original draft preparation, T.G. and Z.S.; writing—review and editing, H.W. and Y.L.; visualization, Y.L.; supervision, T.G.; project administration, T.G.; funding acquisition, T.G., H.W. and Y.L. All authors have read and agreed to the published version of the manuscript.

Funding

This research was supported in part by the National Natural Science Foundation of China (Grants 62301144 and 72404155), and in part by the Zhishan Young Scholar Fund (No. 2242025RCB0032) and Ningbo Yongjiang Talent Program (No. 2025A-418-G).

Data Availability Statement

No new data were created or analyzed in this study.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A. Proof of Theorem 1

Proof. 
The achievability part is a straightforward extension of the joint typicality coding scheme for lossy source coding. We simply present the coding ideas and analysis as follows. Fix the conditional pmf p ( x ^ 1 , x ^ 2 , s ^ | x 1 , x 2 , y ) such that the distortion constraints are satisfied, i.e., E d 1 ( X 1 , X ^ 1 ) D 1 , E d 2 ( X 2 , X ^ 2 ) D 2 , and E d s ( X 1 , S ^ ) D s .
Let p ( x ^ 1 , x ^ 2 , s ^ | y ) = x 1 , x 2 p ( x 1 , x 2 | y ) p ( x ^ 1 , x ^ 2 , s ^ | x 1 , x 2 , y ) . Randomly and independently generate 2 n R sequence triples ( x ^ 1 n , x ^ 2 n , s ^ n ) indexed by m [ 1 : 2 n R ] , each according to p ( x ^ 1 , x ^ 2 , s ^ | y ) . The whole codebook C , consisting of these sequence triples, is revealed to both the encoder and decoder. When observing the source messages ( x 1 n , x 2 n , y n ) , find an index m such that its indexing sequence ( x ^ 1 n , x ^ 2 n , s ^ n ) satisfies ( x 1 n , x 2 n , y n , x ^ 1 n , x ^ 2 n , s ^ n ) T ϵ n . If there is more than one such index, randomly choose one of them; if there is no such an index, set m = 1 . Upon receiving the index m, the decoder reconstructs the messages and inference by choosing the codeword ( x ^ 1 n , x ^ 2 n , s ^ n ) indexed by m. By the law of large numbers, the source sequences are joint typical with probability 1 as n . Then, we define the “encoding error” event as
E = X 1 n , X 2 n , Y n , X ^ 1 n , X ^ 2 n , S ^ n T ϵ n ,   m 1 : 2 n R .
Then, we can bound the error probability as follows
P ( E )   = P x 1 n , x 2 n , y n , X ^ 1 n , X ^ 2 n , S ^ n T ϵ n ,   m 1 : 2 n R   = m = 1 2 n R P x 1 n , x 2 n , y n , X ^ 1 n , X ^ 2 n , S ^ n T ϵ n   = 1 P x 1 n , x 2 n , y n , X ^ 1 n , X ^ 2 n , S ^ n T ϵ n 2 n R   ( x 1 n , x 2 n , y n ) T ϵ n p ( x 1 n , x 2 n , y n ) · 1 2 n [ I ( X 1 , X 2 ; X ^ 1 , X ^ 2 , S ^ | Y ) + δ ( ϵ ) ] 2 n R   exp 2 n [ R I ( X 1 , X 2 ; X ^ 1 , X ^ 2 , S ^ | Y ) δ ( ϵ ) ] ,
where δ ( ϵ ) 0 as n , the first inequality follows from the joint typicality lemma in [22], and the last inequality follows from the fact that ( 1 z ) t exp ( t z ) for z [ 0 , 1 ] and t 0 . We see that P ( E ) 0 as n if R > I ( X 1 , X 2 ; X ^ 1 , X ^ 2 , S ^ | Y ) + δ ( ϵ ) . If the error event does not happen, i.e., the reconstruction is joint typical with the source sequences, then from the distortion constraints assumed for the conditional pmf, the expected distortions can achieve D 1 , D 2 , and D s , respectively. This proves the achievability.
Define R I ( D 1 , D 2 , D s ) as the rate–distortion function characterized by Theorem 1. For the converse part, we show that
n R H ( W ) H ( W | Y n ) I ( X 1 n , X 2 n ; W | Y n )   I ( X 1 n , X 2 n ; X ^ 1 n , X ^ 2 n , S ^ n | Y n )   = I ( X 1 n , X 2 n , Y n ; X ^ 1 n , X ^ 2 n , S ^ n ) I ( Y n ; X ^ 1 n , X ^ 2 n , S ^ n )   = i = 1 n I ( X 1 , i , X 2 , i , Y i ; X ^ 1 n , X ^ 2 n , S ^ n | X 1 , 1 i 1 , X 2 , 1 i 1 , Y 1 i 1 ) I ( Y i ; X ^ 1 n , X ^ 2 n , S ^ n | Y 1 i 1 )   = i = 1 n I ( X 1 , i , X 2 , i , Y i ; X ^ 1 n , X ^ 2 n , S ^ n , X 1 , 1 i 1 , X 2 , 1 i 1 , Y 1 i 1 ) I ( Y i ; X ^ 1 n , X ^ 2 n , S ^ n , Y 1 i 1 )   = i = 1 n I ( X 1 , i , X 2 , i ; X ^ 1 , i , X ^ 2 , i , S ^ i | Y i )               + I ( X 1 , i , X 2 , i , Y i ; X ^ 1 , 1 i 1 , X ^ 1 , i + 1 n , X ^ 2 , 1 i 1 , X ^ 2 , i + 1 n , S ^ 1 i 1 , S ^ i + 1 n , X 1 , 1 i 1 , X 2 , 1 i 1 , Y 1 i 1 | X ^ 1 , i , X ^ 2 , i , S ^ i )               I ( Y i ; X ^ 1 , 1 i 1 , X ^ 1 , i + 1 n , X ^ 2 , 1 i 1 , X ^ 2 , i + 1 n , S ^ 1 i 1 , S ^ i + 1 n , Y 1 i 1 | X ^ 1 , i , X ^ 2 , i , S ^ i )   = i = 1 n I ( X 1 , i , X 2 , i ; X ^ 1 , i , X ^ 2 , i , S ^ i | Y i ) + I ( X 1 , i , X 2 , i , Y i ; X 1 , 1 i 1 , X 2 , 1 i 1 | X ^ 1 n , X ^ 2 n , S ^ n , Y 1 i )               + I ( X 1 , i , X 2 , i ; X ^ 1 , 1 i 1 , X ^ 1 , i + 1 n , X ^ 2 , 1 i 1 , X ^ 2 , i + 1 n , S ^ 1 i 1 , S ^ i + 1 n , Y 1 i 1 | X ^ 1 , i , X ^ 2 , i , S ^ i , Y i )
  i = 1 n I ( X 1 , i , X 2 , i ; X ^ 1 , i , X ^ 2 , i , S ^ i | Y i )
  i = 1 n R I E d 1 ( X 1 , i , X ^ 1 , i ) , E d 2 ( X 2 , i , X ^ 2 , i ) , E d s ( X 1 , i , S ^ i )
  n R I E d 1 ( X 1 n , X ^ 1 n ) , E d 2 ( X 2 n , X ^ 2 n ) , E d s ( X 1 n , S ^ n )
  n R I E d 1 ( X 1 n , X ^ 1 n ) , E d 2 ( X 2 n , X ^ 2 n ) , E d s ( S n , S ^ n )
  n R I ( D 1 , D 2 , D s ) ,
where (A2) follows from the non-negativity of mutual information, (A3) follows from the definition of R I ( D 1 , D 2 , D s ) , (A4) follows from the convexity of R I ( D 1 , D 2 , D s ) , (A5) follows from E d s ( X 1 n , S ^ n ) = E d s ( S n , S ^ n ) which is proved in [11], and the last inequality follows from the non-increasing property of R I ( D 1 , D 2 , D s ) . This completes the converse proof. □

Appendix B. Proof of Lemma 2

Proof. 
The Markov chain X 1 Y X 2 indicates that
H ( X 2 | X 1 , Y ) = H ( X 2 | Y ) .
Then, the mutual information in (9) can be bounded by
I ( X 1 , X 2 ; X ^ 1 , X ^ 2 , S ^ | Y )   = H ( X 1 , X 2 | Y ) H ( X 1 , X 2 | X ^ 1 , X ^ 2 , S ^ , Y )   = H ( X 1 | Y ) + H ( X 2 | Y ) H ( X 1 | X ^ 1 , X ^ 2 , S ^ , Y ) H ( X 2 | X 1 , X ^ 1 , X ^ 2 , S ^ , Y )   H ( X 1 | Y ) + H ( X 2 | Y ) H ( X 1 | X ^ 1 , S ^ , Y ) H ( X 2 | X ^ 2 , Y )   = I ( X 1 ; X ^ 1 , S ^ | Y ) + I ( X 2 ; X ^ 2 | Y ) ,
where the inequality follows from the fact that conditioning does not increase entropy. Now, we have
R ( D 1 , D 2 , D s )   = min p ( x ^ 1 , x ^ 2 , s ^ | x 1 , x 2 , y ) E d 1 ( X 1 , X ^ 1 ) D 1 E d 2 ( X 2 , X ^ 2 ) D 2 E d s ( X 1 , S ^ ) D s I ( X 1 , X 2 ; X ^ 1 , X ^ 2 , S ^ | Y )   min p ( x ^ 1 , x ^ 2 , s ^ | x 1 , x 2 , y ) E d 1 ( X 1 , X ^ 1 ) D 1 E d 2 ( X 2 , X ^ 2 ) D 2 E d s ( X 1 , S ^ ) D s I ( X 1 ; X ^ 1 , S ^ | Y ) + I ( X 2 ; X ^ 2 | Y )   = min p ( x ^ 1 , s ^ | x 1 , y ) E d 1 ( X 1 , X ^ 1 ) D 1 E d s ( X 1 , S ^ ) D s I ( X 1 ; X ^ 1 , S ^ | Y ) + min p ( x ^ 2 | x 2 , y ) E d 2 ( X 2 , X ^ 2 ) D 2 I ( X 2 ; X ^ 2 | Y )   = R 2 d , X 1 | Y ( D 1 , D s ) + R X 2 | Y ( D 2 ) .
For the other direction, we show that the rate–distortion quadruple ( R 2 d , X 1 | Y ( D 1 , D s ) + R X 2 | Y ( D 2 ) , D 1 , D 2 , D s ) is achievable. To see this, let p * ( x ^ 1 , s ^ | x 1 , y ) and p * ( x ^ 2 | x 2 , y ) be the optimal distributions that achieve the rate–distortion tuples R 2 d , X 1 | Y ( D 1 , D s ) , D 1 , D s and R X 2 | Y ( D 2 ) , D 2 , respectively. Now, we consider the distribution p * ( x 1 , x 2 , x ^ 1 , x ^ 2 , s ^ | y ) p * ( x 1 , x ^ 1 , s ^ | y ) p * ( x 2 , x ^ 2 | y ) which requires the Markov chain ( X 1 , X ^ 1 , S ^ ) Y ( X 2 , X ^ 2 ) and is consistent with the condition X 1 Y X 2 . Then, the corresponding random variables satisfy
I ( X 1 , X 2 ; X ^ 1 , X ^ 2 , S ^ | Y )   = H ( X 1 | Y ) + H ( X 2 | Y ) H ( X 1 | X ^ 1 , X ^ 2 , S ^ , Y ) H ( X 2 | X 1 , X ^ 1 , X ^ 2 , S ^ , Y )   = H ( X 1 | Y ) + H ( X 2 | Y ) H ( X 1 | X ^ 1 , S ^ , Y ) H ( X 2 | X ^ 2 , Y )   = I ( X 1 ; X ^ 1 , S ^ | Y ) + I ( X 2 ; X ^ 2 | Y )   = R 2 d , X 1 | Y ( D 1 , D s ) + R X 2 | Y ( D 2 ) ,
where the first equality follows from (A7), the second equality follows from the Markov chain ( X 1 , X ^ 1 , S ^ ) Y ( X 2 , X ^ 2 ) , and the last equality follows from the optimality of p * ( x ^ 1 , s ^ | x 1 , y ) and p * ( x ^ 2 | x 2 , y ) . Lastly, by the minimization in the expression of the rate–distortion function in (9), we conclude that R ( D 1 , D 2 , D s ) R 2 d , X 1 | Y ( D 1 , D s ) + R X 2 | Y ( D 2 ) , which completes the proof of the lemma. □

Appendix C. Proof of Lemma 3

Proof. 
As we are in the binary Hamming setting, we first calculate d s ( x 1 , s ^ ) (c.f. (10)) by
d s ( 0 , 0 ) = 1 p ( x 1 = 0 ) s = 0 , 1 p ( x 1 = 0 , s ) d s ( s , 0 )   = 1 p ( x 1 = 0 ) p ( x 1 = 0 , s = 0 ) d s ( 0 , 0 ) + p ( x 1 = 0 , s = 1 ) d s ( 1 , 0 )   = p ( x 1 = 0 , s = 1 ) p ( x 1 = 0 )   = p ( s = 1 | x 1 = 0 ) = p .
The other values follow similarly, and we obtain the distortion function
d s ( x 1 , s ^ ) = p , if   s ^ = x 1 1 p , if   s ^ x 1 .
Then,
E d s ( X 1 , S ^ ) = x 1 , s ^ p ( x 1 , s ^ ) d s ( x 1 , s ^ )   = P ( X 1 S ^ ) · ( 1 p ) + P ( X 1 = S ^ ) · p   = P ( X 1 S ^ ) 1 2 p + p .
The distortion constraint E d s ( X 1 , S ^ ) D s for D s p implies P ( X 1 S ^ ) D s p 1 2 p . Now, we can follow the rate–distortion evaluation for Bernoulli source and Hamming distortion in [23] while only changing the probability P ( X 1 S ^ ) and obtain
R s ( D s ) = 1 h b D s p 1 2 p · 1 p D s 0.5 .
This proves the lemma. □

Appendix D. Proof of Theorem 3

Proof. 
Since Lemma 2 holds here, we first calculate the rate–distortion function for X 2 , which is
R X 2 | Y ( D 2 ) = 1 h b ( D 2 ) · 1 0 D 2 0.5 .
Now, it remains to calculate R 2 d , X 1 | Y ( D 1 , D s ) . We first consider the case that D 1 D s 0 and provide a lower bound of the mutual information as follows
I ( X 1 ; X ^ 1 , S ^ | Y )   I ( X 1 ; X ^ 1 | Y ) = H ( X 1 | Y ) H ( X 1 | X ^ 1 , Y )   H ( X 1 | Y ) H ( X 1 | X ^ 1 )   = h b ( p 2 ) + log ( N / 2 ) H ( X 1 | X ^ 1 )   h b ( p 2 ) + log ( N / 2 ) h b ( D 1 ) D 1 log ( N 1 ) ,
where the last inequality follows from P ( X ^ 1 X 1 ) D 1 , h b ( x ) is an increasing function for x [ 0 , 0.5 ] , and uniform distribution maximizes entropy. (Note that we can also obtain the above lower bound by directly applying the log-sum inequality.)
Next, we show the lower bound is tight by finding a joint distribution that achieves the above rate and distortions D 1 and D s . For 0 D 1 2 ( N 1 ) p 2 N , we choose S ^ = X ^ 1 and ( X 1 , Y , X ^ 1 ) by the test channel p ( x 1 | x ^ 1 ) in Figure A1 and the Markov chain Y X ^ 1 X 1 . The conditional probability of the test channel in Figure A1 is given as
p ( x ^ 1 | x ) = 1 D 1 ,   if   x ^ 1 = x D 1 N 1 ,   if   x ^ 1 x .
Solving the equations
q i ( 1 D 1 ) + j [ 1 : N ] , j i q j D 1 N 1 = 1 N ,   i [ 1 : N ] ,
we obtain that q i = 1 N for i [ 1 : N ] , i.e., X ^ 1 is also uniformly distributed over [ 1 : N ] . For the joint distribution of Y and X ^ 1 , we define
p ( y | x ^ 1 ) = p 2 N D 1 2 ( N 1 ) 1 N D 1 N 1 , if   y = 0 , x ^ 1   is   odd       or   y = 1 , x ^ 1   is   even     1 p 2 N D 1 2 ( N 1 ) 1 N D 1 N 1 , if   y = 0 , x ^ 1   is   even       or   y = 1 , x ^ 1   is   odd .
Figure A1. Test channel from X ^ 1 to X 1 : X 1 Uniform( 1 N ) and the transition probability p ( x 1 | x ^ 1 ) is given by (A11).
Figure A1. Test channel from X ^ 1 to X 1 : X 1 Uniform( 1 N ) and the transition probability p ( x 1 | x ^ 1 ) is given by (A11).
Entropy 28 00593 g0a1
We see that p ( y | x ^ 1 ) 0 for any D 1 2 ( N 1 ) p 2 N , i.e., ( D 1 , D 2 , D s ) D 1 . Then, we can verify using p ( y | x 1 ) = x ^ 1 p ( y | x 1 , x ^ 1 ) p ( x ^ | x ) = x ^ 1 p ( y | x ^ 1 ) p ( x ^ | x ) that the above distribution can induce the same conditional probability p ( y | x 1 ) (with transition probability p 2 ) as defined at the beginning of this section. Thus, we have constructed a feasible p ( x 1 , y , x ^ 1 ) that can achieve expected distortion E d 1 ( X , X ^ 1 ) = D 1 , D s 0 D 1 , and mutual information
I ( X 1 ; X ^ 1 , S ^ | Y )   = I ( X 1 ; X ^ 1 | Y ) = H ( X 1 | Y ) H ( X 1 | X ^ 1 , Y )   = H ( X 1 | Y ) H ( X 1 | X ^ 1 )   = h b ( p 2 ) + log ( N / 2 ) h b ( D 1 ) D 1 log ( N 1 ) ,
where the first equality follows from S ^ = X ^ 1 , the third equality follows from the Markov chain Y X ^ 1 X 1 , and the last equality follows from the joint distribution of ( X 1 , Y ) and the distribution in (A11).
For the other case that D 1 D s 0 , we only need to switch the role of ( X ^ 1 , D 1 ) and ( S ^ , D s 0 ) . Then, the rate and distortions can be obtained accordingly, which can prove the theorem. □

Appendix E. Proof of Theorem 4

Proof. 
The rate–distortion function in Theorem 1 satisfies
R ( D 1 , D 2 , D s ) = R 2 d , X 1 | Y ( D 1 , D s ) + R X 2 | Y ( D 2 ) .
The second term is the solution to the quadratic Gaussian source coding problem with side information [22] [Chapter 11], which is given as
R X 2 | Y ( D 2 ) = 1 2 log σ X 2 | Y D 2 + ,
where σ X 2 | Y is the conditional variance of X 2 given Y. Note that X 2 is Gaussian conditioning on Y, i.e., X 2 | Y N ( σ X 2 Y σ Y Y , σ X 2 σ X 2 Y 2 σ Y ) , which implies σ X 2 | Y = σ X 2 σ X 2 Y 2 σ Y .
For the first term, note that R 2 d , X 1 | Y ( D 1 , D s ) is lower bounded by both
R X 1 | Y ( D 1 ) = min E d 1 X 1 , X ^ 1 D 1 I ( X 1 ; X ^ 1 | Y ) ,
and
R S | Y ( D s ) = min E d S X 1 , S ^ D s I ( X 1 ; S ^ | Y ) .
Obviously, (A14) is the solution of the quadratic Gaussian source coding with side information similarly to (A13).
We proceed to calculate (A15), which is actually the semantic rate–distortion function of the indirect source coding with side information. Observing X 1 , Y , S is conditionally Gaussian as S | X 1 , Y N σ S X 1 σ X 1 X 1 , σ S σ S X 1 2 σ X 1 . It is shown in [29] that we can rewrite the semantic distortion as
E d S X 1 , S ^ = E S S ˜ MMSE 2 + E S ˜ MMSE S ^ 2 ,
where S ˜ MMSE = σ S X 1 σ X 1 X 1 is the MMSE estimator upon observing X 1 and Y, and the first term on the right-hand side is the corresponding minimum mean squared error ( D MMSE , c.f. (18)), i.e.,
E S S ˜ MMSE 2 = D MMSE = σ S σ S X 1 2 σ X 1 .
Then, we consider a specific encoder which first estimates the semantic information using the MMSE estimator and then compresses the estimation under mean squared error distortion constraint D s D MMSE with side information. The resulting achievable rate provides an upper bound on R S | Y ( D s ) , i.e.,
R S | Y ( D s ) 1 2 log σ S ˜ MMSE | Y D s D MMSE +   = 1 2 log σ S X 1 2 σ X 1 | Y σ X 1 2 D s D MMSE +
  = 1 2 log σ S X 1 2 σ X 1 σ X 1 Y 2 σ Y σ X 1 2 D s D MMSE + .
Furthermore, for D s D MMSE + σ S X 1 2 σ X 1 | Y σ X 1 2 , we have R S | Y ( D s ) = 0 , which can be obtained by setting S ^ = E S ˜ MMSE | Y = σ S X 1 σ X 1 E X 1 | Y . For D s < D MMSE + σ S X 1 2 σ X 1 | Y σ X 1 2 , we derive a lower bound for R S | Y ( D s ) as follows
R S | Y ( D s ) I ( X 1 ; S ^ | Y ) = H ( X 1 | Y ) H ( X 1 | S ^ , Y )   = 1 2 log 2 π e σ X 1 | Y H ( X 1 σ X 1 σ S X 1 S ^ | S ^ , Y )   1 2 log 2 π e σ X 1 | Y H ( X 1 σ X 1 σ S X 1 S ^ )
  1 2 log 2 π e σ X 1 | Y 1 2 log 2 π e E X 1 σ X 1 σ S X 1 S ^ 2
  1 2 log 2 π e σ X 1 | Y 1 2 log 2 π e σ X 1 2 D s D MMSE σ S X 1 2
  = 1 2 log σ S X 1 2 σ X 1 | Y σ X 1 2 D s D MMSE = 1 2 log σ S X 1 2 σ X 1 σ X 1 Y 2 σ Y σ X 1 2 D s D MMSE ,
where (A19) is due to the fact that the Gaussian distribution maximizes the entropy for a given variance, and (A20) follows from (A16) and the semantic distortion constraint. Combining the upper and lower bounds in (A18) and (A21), we obtain
R S | Y ( D s ) = 1 2 log σ S X 1 2 σ X 1 σ X 1 Y 2 σ Y σ X 1 2 D s D MMSE + .
Thus, we have
R ( D 1 , D 2 , D s ) max R X 1 | Y ( D 1 ) , R S | Y ( D s ) + R X 2 | Y ( D 2 )   = 1 2 log σ X 2 σ X 2 Y 2 σ Y D 2 + + 1 2 log max σ X 1 σ X 1 Y 2 σ Y D 1 , σ S X 1 2 σ X 1 σ X 1 Y 2 σ Y σ X 1 2 D s D MMSE + .
To show the achievability, consider the following two cases.
  • For D s D MMSE σ S X 1 2 D 1 σ X 1 2 , we first reconstruct X ^ 1 and X 2 subject to distortion constraints D 1 and D 2 and hence achieve R X 1 | Y ( D 1 ) + R X 2 | Y ( D 2 ) . Then, we recover the semantic information by S ^ = σ S X 1 σ X 1 X ^ 1 , and the semantic distortion satisfies
    E d S X 1 , S ^ = D MMSE + E σ S X 1 σ X 1 X 1 σ S X 1 σ X 1 X ^ 1 2 D MMSE + σ S X 1 2 σ X 1 2 D 1 D s .
  • For D s D MMSE σ S X 1 2 < D 1 σ X 1 2 , we first reconstruct S ^ and X 2 subject to distortion constraints D s and D 2 and hence achieve R S | Y ( D s ) + R X 2 | Y ( D 2 ) . Then, we recover X ^ 1 = σ X 1 σ S X 1 S ^ , and the distortion satisfies
    E d 1 X 1 , X ^ 1 = E X 1 σ X 1 σ S X 1 S ^ 2 = σ X 1 2 σ S X 1 2 E d S X 1 , S ^ D MMSE < D 1 .
This establishes the achievability and thus completes the proof. □

References

  1. Shannon, C.E. A mathematical theory of communication. Bell Syst. Tech. J. 1948, 27, 379–423. [Google Scholar] [CrossRef] [Scilit]
  2. Shannon, C.E. Coding theorems for a discrete source with a fidelity criterion. IRE Nat. Conv. Rec. 1959, 7, 142–163. [Google Scholar]
  3. Shao, J.; Zhang, X.; Zhang, J. Task-Oriented Communication for Edge Video Analytics. IEEE Trans. Wirel. Commun. 2024, 23, 4141–4154. [Google Scholar] [CrossRef] [Scilit]
  4. Wang, Y.; Chan, P.H.; Donzella, V. Semantic-Aware Video Compression for Automotive Cameras. IEEE Trans. Intell. Veh. 2023, 8, 3712–3722. [Google Scholar] [CrossRef] [Scilit]
  5. Al-Shakarji, N.; Bunyak, F.; Aliakbarpour, H.; Seetharaman, G.; Palaniappan, K. Multi-cue vehicle detection for semantic video compression in georegistered aerial videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, Vancouver, BC, Canada, 18–22 June 2023; pp. 56–65. [Google Scholar]
  6. Chen, J.; Fang, Y.; Khisti, A.; Özgür, A.; Shlezinger, N. Information Compression in the AI Era: Recent Advances and Future Challenges. IEEE J. Sel. Areas Commun. 2025, 43, 2333–2348. [Google Scholar] [CrossRef] [Scilit]
  7. Liang, G.; Zhang, X.; Zhang, J.; Sun, Y.; Cui, Q.; Tao, X. Semantic codebook-based HARQ for wireless image transmission. IEEE Trans. Commun. 2025, 73, 14332–14346. [Google Scholar] [CrossRef] [Scilit]
  8. Gündüz, D.; Qin, Z.; Aguerri, I.E.; Dhillon, H.S.; Yang, Z.; Yener, A.; Wong, K.K.; Chae, C.B. Beyond Transmitting Bits: Context, Semantics, and Task-Oriented Communications. IEEE J. Sel. Areas Commun. 2023, 41, 5–41. [Google Scholar] [CrossRef] [Scilit]
  9. Xu, W.; Yang, Z.; Ng, D.W.K.; Levorato, M.; Eldar, Y.C.; Debbah, M. Edge Learning for B5G Networks With Distributed Signal Processing: Semantic Communication, Edge Computing, and Wireless Sensing. IEEE J. Sel. Top. Signal Process. 2023, 17, 9–39. [Google Scholar] [CrossRef] [Scilit]
  10. Witsenhausen, H. Indirect rate distortion problems. IEEE Trans. Inf. Theory 1980, 26, 518–521. [Google Scholar] [CrossRef] [Scilit]
  11. Liu, J.; Shao, S.; Zhang, W.; Poor, H.V. An indirect rate-distortion characterization for semantic sources: General model and the case of Gaussian observation. IEEE Trans. Commun. 2022, 70, 5946–5959. [Google Scholar] [CrossRef] [Scilit]
  12. Stavrou, P.A.; Kountouris, M. The Role of Fidelity in Goal-Oriented Semantic Communication: A Rate Distortion Approach. IEEE Trans. Commun. 2023, 71, 3918–3931. [Google Scholar] [CrossRef] [Scilit]
  13. Wang, Y.; Wu, Y.; Ma, S.; Angela Zhang, Y.J. Task-Oriented Lossy Compression With Data, Perception, and Classification Constraints. IEEE J. Sel. Areas Commun. 2025, 43, 2635–2650. [Google Scholar] [CrossRef] [Scilit]
  14. Stavrou, P.A.; Kountouris, M. Goal-Oriented Single-Letter Codes for Lossy Joint Source-Channel Coding. In Proceedings of the IEEE Information Theory Workshop (ITW 2023), Saint-Malo, France, 23–28 April 2023; pp. 64–69. [Google Scholar]
  15. Gray, R.M. Conditional Rate-Distortion Theory; Technical Report, 6502-2; Stanford Electronics Lab: Stanford, CA, USA, 1972; Volume 7. [Google Scholar]
  16. Gray, R.M. A new class of lower bounds to information rates of stationary sources via conditional rate-distortion functions. IEEE Trans. Inf. Theory 1973, 19, 480–489. [Google Scholar] [CrossRef]
  17. Wyner, A.; Ziv, J. The rate-distortion function for source coding with side information at the decoder. IEEE Trans. Inf. Theory 1976, 22, 1–10. [Google Scholar] [CrossRef] [Scilit]
  18. Kaspi, A. Rate-distortion function when side-information may be present at the decoder. IEEE Trans. Inf. Theory 1994, 40, 2031–2034. [Google Scholar] [CrossRef] [Scilit]
  19. Slepian, D.; Wolf, J. Noiseless coding of correlated information sources. IEEE Trans. Inf. Theory 1973, 19, 471–480. [Google Scholar] [CrossRef] [Scilit]
  20. Zhang, X.; Shao, J.; Zhang, J. LDMIC: Learning-based Distributed Multi-view Image Coding. In Proceedings of the 11th International Conference on Learning Representations (ICLR 2023), Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
  21. Liu, J.; Wei, Y.; Lin, J.; Zhao, S.; Sun, H.; Chen, Z.; Zeng, W.; Jin, X. Tell Codec What Worth Compressing: Semantically Disentangled Image Coding for Machine with LMMs. In Proceedings of the IEEE International Conference on Visual Communications and Image Processing (VCIP 2024), Tokyo, Japan, 8–11 December 2024; pp. 1–5. [Google Scholar]
  22. El Gamal, A.; Kim, Y.H. Network Information Theory; Cambridge University Press: Cambridge, UK, 2011. [Google Scholar] [CrossRef] [Scilit]
  23. Cover, T.M.; Thomas, J.A. Elements of Information Theory, 2nd ed.; Wiley Series in Telecommunications and Signal Processing; Wiley-Interscience: New York, NY, USA, 2006. [Google Scholar]
  24. Carter, M. Source Coding of Composite Sources. Ph.D. Thesis, Department of Computer Information and Control Engineering, University of Michigan, Ann Arbor, MI, USA, 1984. [Google Scholar]
  25. Kipnis, A.; Rini, S.; Goldsmith, A.J. The indirect rate-distortion function of a binary i.i.d source. In Proceedings of the 2015 IEEE Information Theory Workshop (ITW 2015), Jerusalem, Israel, 26 April–1 May 2015; pp. 352–356. [Google Scholar]
  26. Zhang, G.; Qian, J.; Chen, J.; Khisti, A. Universal rate-distortion-perception representations for lossy compression. Adv. Neural Inf. Process. Syst. 2021, 34, 11517–11529. [Google Scholar] [CrossRef] [Scilit]
  27. Agustsson, E.; Theis, L. Universally quantized neural compression. Adv. Neural Inf. Process. Syst. 2020, 33, 12367–12376. [Google Scholar]
  28. Agustsson, E.; Tschannen, M.; Mentzer, F.; Timofte, R.; Gool, L.V. Generative adversarial networks for extreme learned image compression. In Proceedings of the IEEE/CVF International Conference on Computer Vsion, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 221–231. [Google Scholar]
  29. Wolf, J.; Ziv, J. Transmission of noisy information to a noisy receiver with minimum distortion. IEEE Trans. Inf. Theory 1970, 16, 406–411. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Illustration of system model with side information.
Figure 1. Illustration of system model with side information.
Entropy 28 00593 g001
Figure 2. Comparison of rate–distortion functions R s ( D ) and R ( D ) for p = 0.1 .
Figure 2. Comparison of rate–distortion functions R s ( D ) and R ( D ) for p = 0.1 .
Entropy 28 00593 g002
Figure 3. The rate–distortion function for integer classification with p = p 1 = p 2 = 0.25 , D 2 = 0.5 , and N = 8 .
Figure 3. The rate–distortion function for integer classification with p = p 1 = p 2 = 0.25 , D 2 = 0.5 , and N = 8 .
Entropy 28 00593 g003
Figure 4. The rate–distortion function for Gaussian sources and its contour plot, for σ S = σ X 1 = σ X 2 = σ Y = 1 , σ S X 1 = σ X 1 Y = σ X 2 Y = 2 , and D 2 = 1 .
Figure 4. The rate–distortion function for Gaussian sources and its contour plot, for σ S = σ X 1 = σ X 2 = σ Y = 1 , σ S X 1 = σ X 1 Y = σ X 2 Y = 2 , and D 2 = 1 .
Entropy 28 00593 g004
Figure 5. Illustration of rate–distortion classification setup.
Figure 5. Illustration of rate–distortion classification setup.
Entropy 28 00593 g005
Figure 6. The classification accuracy versus distortion with varying rates R and weights λ c . The results of Case A and Case B are denoted by points with and without black outline, respectively. The fitting curve is also plotted for each R, where the solid and dashed lines represent, respectively Case B and Case A.
Figure 6. The classification accuracy versus distortion with varying rates R and weights λ c . The results of Case A and Case B are denoted by points with and without black outline, respectively. The fitting curve is also plotted for each R, where the solid and dashed lines represent, respectively Case B and Case A.
Entropy 28 00593 g006
Figure 7. Illustration of performance improvement with side information on the MNIST dataset, when R = 4 , λ d = 1 and λ c = 0.001 .
Figure 7. Illustration of performance improvement with side information on the MNIST dataset, when R = 4 , λ d = 1 and λ c = 0.001 .
Entropy 28 00593 g007
Figure 8. Visual reconstructions for the SVHN ( R = 30 ) dataset with different λ c .
Figure 8. Visual reconstructions for the SVHN ( R = 30 ) dataset with different λ c .
Entropy 28 00593 g008
Table 1. Performance comparison of Case A and Case B for MNIST ( R = 4 ) and SVHN ( R = 30 ) datasets. Here, λ d is fixed to 1 and λ c varies. Moreover, ‘diff.’ denotes the relative difference between two cases, i.e., diff . = result ( Case   A ) result ( Case   B ) result ( Case   B ) × 100 % .
Table 1. Performance comparison of Case A and Case B for MNIST ( R = 4 ) and SVHN ( R = 30 ) datasets. Here, λ d is fixed to 1 and λ c varies. Moreover, ‘diff.’ denotes the relative difference between two cases, i.e., diff . = result ( Case   A ) result ( Case   B ) result ( Case   B ) × 100 % .
λ c CaseMSEAccuracyMSE diff.Accuracy diff.
(a) MNIST0.0010Case B0.01710.91620.00%−7.40%
Case A0.01710.8484
0.0025Case B0.01730.94100.58%−1.21%
Case A0.01740.9296
0.0060Case B0.01790.9517−1.12%−1.08%
Case A0.01770.9414
0.0100Case B0.01780.95971.69%−1.35%
Case A0.01810.9467
(b) SVHN0.0010Case B0.00280.8397−3.57%−7.40%
Case A0.00270.7776
0.0030Case B0.00280.84250.00%−6.75%
Case A0.00280.7856
0.0080Case B0.00300.8446−3.33%−5.52%
Case A0.00290.7980
0.0120Case B0.00310.8489−3.23%−5.04%
Case A0.00300.8061
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Guo, T.; Song, Z.; Wu, H.; Li, Y. Rate–Distortion Limits for Task-Oriented Compression with Side Information. Entropy 2026, 28, 593. https://doi.org/10.3390/e28060593

AMA Style

Guo T, Song Z, Wu H, Li Y. Rate–Distortion Limits for Task-Oriented Compression with Side Information. Entropy. 2026; 28(6):593. https://doi.org/10.3390/e28060593

Chicago/Turabian Style

Guo, Tao, Zhangyao Song, Huihui Wu, and Yang Li. 2026. "Rate–Distortion Limits for Task-Oriented Compression with Side Information" Entropy 28, no. 6: 593. https://doi.org/10.3390/e28060593

APA Style

Guo, T., Song, Z., Wu, H., & Li, Y. (2026). Rate–Distortion Limits for Task-Oriented Compression with Side Information. Entropy, 28(6), 593. https://doi.org/10.3390/e28060593

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop