Skip to Content
EntropyEntropy
  • Article
  • Open Access

3 August 2026

36 Pages

Undetected Error Bounds for Hybrid Integrity Protection Using Reed–Muller Codes, Algebraic Manipulation Detection, and Universal Hashing

,
,
,
,
,
,
,
,
and
1
Department of Artificial Intelligence, Gachon University, Seongnam 13120, Republic of Korea
2
Department of Artificial Intelligence, Samarkand State University Named After Sharof Rashidov, Samarkand 140104, Uzbekistan
3
Department of Construction Engineering, Samarkand State Technical University, Samarkand 140143, Uzbekistan
4
Department of IT, Samarkand Institute of Economics and Service, Samarkand 140100, Uzbekistan

Abstract

Ensuring information integrity requires not only reducing decoding errors but also reducing the probability that corrupted data are accepted as valid. This research presents a hybrid integrity protection system that incorporates seeded universal hash verification, algebraic manipulation detection (AMD), and a binary Reed–Muller outer code. Transmission over the binary symmetric channel B S C ( p ) , outer encoding using R M ( r ,   m ) , bounded-distance decoding, an ε A M D -secure AMD layer, and a seeded 2-universal hash family with l -bit output define the model used in the analysis. Under explicitly stated freshness and conditional-independence assumptions, the system-level undetected error probability is upper-bounded by the residual decoder-miscorrection probability multiplied by the AMD acceptance bound and the seeded universal hash collision bound. A conservative alternative is also provided for settings in which the required conditional independence cannot be guaranteed. In this context, an explicit upper bound for the undetected error probability is derived. The outcome makes clear the different functions of outer coding and post-decoding verification and results in a direct dependency on the parameters r, m, p, and l. Finite-length Monte Carlo validation for a concrete instantiation based on RM(2, 5) complements the theoretical study and verifies that the hybrid construction offers a lower empirical undetected error probability compared to the comparable outer-only, AMD-only, and hash-only variations. The study does not propose new coding or verification primitives. Its contribution is a finite-length layered acceptance model and a Reed–Muller-specific undetected error analysis that incorporates the code weight distribution and bounded-distance decoding regions. The resulting spectrum-based bound distinguishes decoder miscorrection from the broader event of exceeding the guaranteed correction radius and is evaluated together with post-decoding verification and redundancy overhead. The model’s formal manipulation detection and collision guarantees are provided by AMD and universal hash layers, while Reed–Muller code parameters and their standard distance formulas are conventional.

1. Introduction

In communication systems, coding theory, and reliability-critical digital infrastructures, ensuring the integrity of transmitted and processed data continues to be a major challenge. In many real-world situations, the goal is to lower both the likelihood of decoding error and the likelihood of corrupted data being regarded as legitimate. This distinction is crucial because, in high-assurance systems, an undetected error can be even more damaging than a typical detected failure because it permits an inaccurate output to flow through the system without causing recovery or retransmission. Because of this, while designing integrity protection measures, the undetected error probability is therefore a critical performance measure. The fundamental mathematical basis for managing such failures is provided by classical error-correcting codes, but the issue of undetected acceptance cannot be entirely solved by distance-based protection alone [1].
Because of its clear algebraic structure, effective implementation, and well-understood distance qualities, linear block codes continue to be among the most crucial technologies for dependable communication. Reed–Muller codes are particularly well-known among these. Because of their explicit parameters, rich algebraic interpretation, and applicability to both theoretical and algorithmic issues, they are among the oldest and most researched code families in coding theory and continue to garner significant attention. A recent survey and related studies show research on decoding, weight distribution, threshold behavior, and performance under random errors continues to rely heavily on Reed–Muller codes [2,3,4]. Despite these advantages, the possibility of undetected acceptance cannot be eliminated by outer error-control coding alone. A decoding error may result in the acceptance of an incorrect decoded candidate even in cases where a code has strong minimum-distance properties. This encourages the introduction of extra verification methods that function after decoding and check for consistency in the decoded output. According to this more comprehensive perspective, the integrity problem naturally splits into two components: first, lowering the likelihood that the channel and decoder generate an incorrect candidate, and second, lowering the likelihood that such a candidate is recognized as legitimate. Therefore, extra post-decoding verification should be combined with distance-based protection in a mathematically adequate integrity architecture.
In this situation, two verification methods are very pertinent. The first is algebraic manipulation detection (AMD), which was developed as a primitive for identifying additive or algebraic modification of encoded data by Cramer, Dodis, Fehr, Pade, and Wichs. AMD codes have become a typical technique in integrity protection and associated cryptographic contexts because they offer formal guarantees that a modified encoding would be rejected except with a tiny probability [1,2,3,4,5]. The second is universal hashing, which was developed by Carter and Wegman and offers strict collision boundaries for seeded hash families. As a result, it offers a mathematically sound method for controlling the likelihood that a message that has been incorrectly decoded will pass a fingerprint-based consistency test [6,7]. AMD codes target algebraic manipulation, while universal hashing provides compact probabilistic verification with explicit collision guarantees.
This work examines a layered integrity protection system under an explicitly defined two-stage threat model. In the first stage, stochastic bit errors are introduced during transmission and are modeled by a binary symmetric channel, BSC(p). A binary Reed–Muller outer code and bounded-distance decoder are used to reduce the probability that these errors produce an incorrect decoded candidate. In the second stage, the decoded candidate is subjected to AMD and seeded universal hash verification. These mechanisms do not replace channel decoding; rather, they reduce the probability that a residual incorrect or algebraically modified candidate is accepted as valid. Thus, the proposed architecture distinguishes transmission reliability from acceptance integrity and combines the corresponding mechanisms within one end-to-end probabilistic model. The model is relevant to telemetry, industrial control, low-rate command transmission, safety-critical messaging, and untrusted storage or relay systems, where silent acceptance of corrupted information may be more harmful than an explicit decoding or verification failure. All coding and verification algorithms are assumed to be public. The security arguments rely on the stated AMD guarantee, the 2-universal collision property, and explicitly defined randomness assumptions rather than on secrecy of the coding architecture.
An explicit undetected error bound for the suggested hybrid approach is the paper’s primary contribution. The resultant bound under the fully stated model factors into a verification term determined by the AMD and hash layers and a residual decoding-failure term determined by the Reed–Muller code and the channel. This factorization produces a direct parameterized dependence on r, m, p, and l and makes clear the distinct responsibilities of coding and verification. Furthermore, a concrete instantiation based on RM(2, 5), a short Reed–Muller code that is also well-known in implementation-oriented decoding studies [1,2,3,4,5,6], is validated for finite length in this study.

3. Mathematical Model and Notation

This section introduces the concrete mathematical model used throughout the paper. In contrast to generic layered integrity architectures, the present work focuses on a fully specified setting consisting of a binary symmetric channel, a binary Reed–Muller outer code, bounded-distance decoding, an explicit algebraic manipulation detection layer, and a seeded universal hash verification layer. This choice allows the undetected error probability to be expressed directly in terms of the code parameters and the channel crossover probability (Table 3). Reed–Muller codes are particularly suitable for this purpose because their length, dimension, and minimum distance are available in closed form.
Table 3. Principal notation used throughout the manuscript.

3.1. Channel Model and Notation

Let F 2 denote the binary field, and let the source message be a vector
M ∈ F 2 k .
The B S C ( p ) represents stochastic, non-adversarial corruption of the transmitted Reed–Muller codeword. It is not intended to represent a fully adaptive manipulation adversary. The AMD threat model is defined separately at the tagged-representation level. Consequently, the random error vector E and an admissible algebraic manipulation offset Δ are different mathematical objects: E acts on the transmitted codeword through the communication channel, whereas Δ represents a nonzero algebraic modification considered under the security definition of the AMD construction. The baseline theorem evaluates residual incorrect candidates caused by E ; the AMD and hash layers then bound the probability that such a candidate passes post-decoding verification. The transmitted codeword is sent through a binary symmetric channel B S C ( p ) , where each coordinate is flipped independently with probability p , 0 < p < 1 / 2 . Thus, if
X ∈ F 2 n
is the transmitted codeword and
E = ( E 1 , … , E n ) ∈ F 2 n
is the random error vector with i.i.d. components E i ∼ B e r n o u l l i ( p ) , then the received word is
Y = X ⊕ E .
The Hamming weight of a vector x ∈ F 2 n is denoted by w H ( x ) , and the Hamming distance between x , z ∈ F 2 n is denoted by d H ( x , z ) . Since the channel is binary and memoryless, the distribution of w H ( E ) is binomial:
P r { w H ( E ) = i } = n i p i ( 1 − p ) n − i ,     i = 0 , 1 , … , n .
This binomial tail will play a central role in the undetected error bounds developed later.

Probability Space and Order of Random-Variable Generation

The system analysis is defined over a joint probability space containing the source message, AMD randomness, universal hash seed, channel-error vector, and decoder output. The order in which these random variables are generated is important because the post-decoding verification bound depends on their conditional relationships.
The experiment proceeds as follows:
1.
A source message M is selected according to a specified source distribution.
2.
Fresh AMD randomness R is sampled according to the AMD construction.
3.
The AMD-protected representation is generated from M   and R .
4.
A hash seed S is sampled independently of M and R from the prescribed seed space.
5.
The hash tag T H = h S M is computed.
6.
The complete tagged representation is embedded into the information coordinates of the Reed–Muller encoder.
7.
A channel-error vector E is generated according to the memoryless B S C ( p ) , independently of R and S .
8.
The receiver applies bounded-distance decoding and obtains either a decoder failure or a candidate tagged representation T ^ .
9.
If a candidate is produced, the receiver applies the AMD and universal hash verification rules.
Let D denote the event that the decoder outputs an incorrect but syntactically valid tagged candidate. Let A A M D denote the event that this candidate passes the AMD verification rule, and let A H denote the event that it passes the universal hash verification rule. The system-level undetected error event is
U = D ∩ A A M D ∩ A H .
The baseline analysis assumes that R , S , and E are generated independently according to their prescribed distributions. It further assumes that the incorrect candidate to which the verification guarantees are applied is fixed before the fresh hash seed is evaluated, or equivalently, that conditioning on the decoder event and the AMD acceptance event does not bias the distribution of S . These assumptions are stated explicitly because the product-form verification bound is not valid under arbitrary dependence.

3.2. Reed–Muller Outer Coding

As the outer error-control code, we use the binary Reed–Muller code R M ( r , m ) , where 0 ≤ r ≤ m . This code has length
n = 2 m ,
dimension
k R M = ∑ i = 0 r m i ,
and minimum distance
d = 2 m − r .
These parameters are classical and provide the main reason for selecting Reed–Muller codes in the present work: they allow the correction radius and residual decoding-failure region to be described explicitly. In addition, Reed–Muller codes admit efficient decoding algorithms and remain one of the most studied structured code families in modern coding theory.
The receiver is assumed to use a bounded-distance decoder for R M ( r , m ) . Its guaranteed correction radius is
t = d − 1 2 = 2 m − r − 1 2 = 2 m − r − 1 − 1 .
Accordingly, every received vector Y satisfying
w H ( E ) ≤ t
is decoded correctly. Any decoding failure or incorrect decoding can therefore occur only when
w H ( E ) > t .
This observation provides the first structural ingredient in system-level reliability analysis.

3.3. AMD Verification Layer

To detect structured additive manipulations that may survive outer decoding, we incorporate an algebraic manipulation detection (AMD) layer. AMD codes were introduced by Cramer, Dodis, Fehr, Padro, and Wichs as keyless mechanisms for detecting additive tampering, with security defined through the probability that a nonzero additive offset transforms a valid encoding into another valid encoding.
Formally, let C A M D = A M D E n c M R denote the AMD-protected representation generated using fresh randomness R . An admissible algebraic adversary selects a nonzero additive offset Δ from the alphabet of the AMD construction and causes the verifier to receive C A M D + Δ . The adversary may know the AMD construction and all public system parameters but does not control the fresh encoder randomness R before selecting an offset in the security experiment assumed here. The AMD construction is ε A M D -secure if, for every admissible nonzero Δ ,
P r R   AMDDec A M D E n c M R + Δ ≠ ⊥ ≤ ε A M D
Here, ⊥ denotes rejection. This guarantee applies to algebraically structured manipulation of the AMD-protected representation and should not be interpreted as a complete model of an arbitrary adaptive communication-channel adversary.

3.4. Seeded Universal Hash Verification

As a second verification layer, we use a seeded universal hash family in the sense of Carter and Wegman. Let
H = { h s : M → { 0 , 1 } l } s ∈ S
be a 2-universal family of hash functions with l -bit output. The hash family and its evaluation algorithm are public. A seed S is sampled independently of M and the AMD randomness R , and the corresponding hash value is T h = h S M . In the present construction, S is included in the tagged representation to allow receiver-side verification; therefore, it is not treated as a secret authentication key. The collision guarantee applies when S is sampled according to the prescribed seed distribution and is not maliciously selected as a function of a competing message pair. Adaptive manipulation after observing or controlling the seed, repeated seed reuse, and compromised seed generation are outside the present theorem and are identified as limitations. This means that for any two distinct messages M ≠ M ′ ,
P r S { h S ( M ) = h S ( M ′ ) } ≤ 2 − l ,
when the seed s is chosen uniformly at random. Universal hashing is especially appropriate here because it gives a direct and rigorous upper bound on the probability that an incorrect decoded message collides with the stored integrity tag.
For a given source message M , the hash tag is
τ H = h S ( M ) ,
and the hash-augmented representation is
T H ( M ; s ) = ( M , s , τ H ) .
At the receiver, a candidate message M ^ passes the hash test only if
τ ^ H = h s ( M ^ ) .
The seed S is sampled independently of the source message, AMD randomness, channel-error vector, and admissible manipulation event. For any fixed pair of distinct messages M ≠ M ^ , the 2-universal property gives
P r S   h S M = h S M ^ ≤ 2 − l .
The sequential theorem below additionally requires that conditioning on the decoder event D and on the AMD acceptance event A A M D does not alter the prescribed distribution of S . Under this condition,
P r   ( A H ′ ∣ D A A M D ) ≤ 2 − l .
This condition is satisfied when the hash seed is fresh, sampled independently, and not chosen adaptively as a function of the incorrect candidate or AMD verification outcome. It may fail under malicious seed selection, correlated seed generation, adaptive seed-dependent manipulation, or repeated seed reuse combined with information leakage.

3.5. Combined Tagged Representation and Transmission

The proposed hybrid construction combines both verification mechanisms before outer encoding. Let the tagged message be
T ( M ; R , s ) = ( M ,   R ,   f ( M , R ) ,   s ,   h s ( M ) ) .
Assume that this tagged object is embedded into the information part of the Reed–Muller encoder, yielding the transmitted codeword
X = E n c R M ( T ( M ; R , s ) ) ∈ F 2 n .
After transmission through B S C ( p ) , the receiver obtains
Y = X ⊕ E
and applies bounded-distance decoding to produce a candidate tagged vector
T ^ = ( M ^ , R ^ , τ ^ A M D , s ^ , τ ^ H ) .
The receiver accepts the decoded output only if both verification conditions hold:
τ ^ A M D = f M ^ , R ^     and   τ ^ H = h s ^ M ^ .
In the implementation considered here, the seed κ is transmitted as part of the tagged object, so the hash check remains fully specified and reproducible to the receiver. The resulting undetected error event is therefore the event that an incorrect decoded candidate is produced and simultaneously passes both the AMD and hash verification stages.

3.6. Performance Metric

The main performance metric is the probability of undetected error,
P u e = P r { M ^ ≠ M   and   the   receiver   accepts } .
This is a stricter event than decoder failure alone. A decoder error does not yet imply system failure, because the incorrect candidate may still be rejected by the AMD or hash verification stages. To avoid ambiguity, four receiver outcomes are distinguished throughout the manuscript. A decoder failure occurs when the bounded-distance decoder does not return a candidate codeword and outputs ⊥ . A decoder miscorrection occurs when the decoder returns a codeword or tagged representation different from the transmitted one. A detected error occurs when channel corruption or decoder miscorrection is identified either through decoder failure or through failure of at least one post-decoding verification check. A rejected candidate is any decoder output that is not accepted because the AMD check, the hash check, or a syntactic validity check fails. Finally, an undetected error occurs only when an incorrect source-message candidate is delivered as valid after all enabled verification checks have accepted it. Accordingly, the purpose of the hybrid architecture is not only to reduce decoding error through outer coding, but also to reduce the probability that a residual decoding error is accepted as valid. The analysis in Section 5 shows that these two effects can be separated explicitly in the fully specified model introduced above.

4. Proposed Hybrid Integrity Protection Scheme

This section specifies the hybrid integrity protection scheme under the mathematical model and notation introduced in Section 3.
The construction combines three components: a binary Reed–Muller outer code, an AMD tagging layer, and a seeded universal hash verification layer as shown in Figure 1. The purpose of the architecture is to ensure that a decoding error at the outer-code level does not automatically imply acceptance of an incorrect source message. Instead, a decoded candidate is accepted only if it also satisfies the AMD and hash consistency conditions.
Figure 1. End-to-end structure of the proposed layered integrity protection system.

4.1. Encoder Structure

Let   M ∈ F 2 k be the source message. Before outer encoding, the message is augmented by two verification layers. First, an AMD randomness variable R is generated, and the AMD tag is computed as τ A M D = f ( M , R ) , where f is the AMD encoding function. Second, a hash seed s is selected and the universal hash tag is computed as τ H = h s ( M ) , where h s is a member of the seeded 2-universal hash family introduced in Section 2.
The resulting tagged object is
T ( M ; R , s ) = ( M ,   R ,   f ( M , R ) ,   s ,   h s ( M ) ) .
Let L denote the binary length of this tagged representation. The tagged vector is then embedded into the information part of the outer Reed–Muller encoder. Denoting the binary Reed–Muller encoding map by E n c R M : F 2 L → F 2 n , the transmitted codeword is
X = E n c R M ( T ( M ; R , s ) ) ,
where n = 2 m for the chosen code R M ( r , m ) .
Thus, the encoder consists of two logically distinct stages:
  • Message augmentation, which adds integrity-check information.
  • Outer coding, which provides distance-based protection against random channel perturbations.
This separation is central to the hybrid design, because the verification layers and the outer code serve different mathematical purposes.
Encoder-side sequence:
  • Generate source message M .
  • Generate fresh AMD randomness R .
  • Compute the AMD-protected representation.
  • Generate a fresh hash seed S , independently of M and R .
  • Compute T h = h S M .
  • Form the tagged representation.
  • Embed the tagged representation into the information positions of R M ( r , m ) .
  • Transmit the resulting codeword through B S C ( p ) .

4.2. Transmission Model

The codeword X is transmitted through the B S C ( p ) defined in Section 3.1, and the receiver observes Y according to Equation (4).
The outer Reed–Muller code is intended to correct sufficiently small perturbations. However, if the noise pattern lies outside the guaranteed correction radius of the bounded-distance decoder, then the receiver may output an incorrect tagged candidate. The role of the verification layers is precisely to reduce the probability that such an incorrect candidate is nevertheless accepted.

4.3. Decoder and Verification Procedure

At the receiver, the first stage is bounded-distance decoding with respect to the outer Reed–Muller code R M ( r , m ) . Let
D e c R M ( Y )
denote the decoder output. If the received vector lies within Hamming distance
t = 2 m − r − 1 − 1
of the transmitted codeword, then the decoder returns the correct tagged vector. Otherwise, the decoder may fail or return an incorrect candidate. Suppose the decoder outputs
T ^ = ( M ^ , R ^ , τ ^ A M D , s ^ , τ ^ H ) .
The receiver accepts this candidate only if both of the following conditions hold:
τ ^ A M D = f ( M ^ , R ^ ) ,
and
τ ^ H = h s ^ ( M ^ ) .
Therefore, the acceptance rule is conjunctive:
Accept   M ^   ⟺   ( τ ^ A M D = f ( M ^ , R ^ ) )   and   ( τ ^ H = h s ^ ( M ^ ) ) .
If either verification condition fails, the decoded output is rejected. In that case, the system may declare an integrity violation, request retransmission, or invoke a higher-layer recovery procedure, depending on the application context. These operational responses are external to the coding model itself and are therefore not included in the mathematical analysis.
Receiver-side sequence:
  • Apply bounded-distance decoding.
  • Reject immediately if the decoder declares failure or if the output cannot be parsed as a valid tagged representation.
  • Extract the candidate message, AMD data, seed, and hash tag.
  • Apply the AMD verification condition.
  • Recompute the universal hash using the recovered seed and candidate message.
  • Accept only when both checks succeed.
  • Otherwise, return ⊥ or declare an integrity violation.

4.4. Protection Modes Considered in the Paper

For analytical comparison, the proposed framework includes four protection modes derived from the same general architecture.
Outer-code-only mode. In the baseline mode, only the source message is encoded by the outer Reed–Muller code. No AMD or hash verification is applied after decoding. Thus, any incorrectly decoded source message is accepted automatically.
AMD-only mode. In this variant, the encoded payload is
T ( A ) = ( M , R , T A M D ) .
After outer decoding, the receiver accepts the candidate only if the AMD consistency condition is satisfied.
Hash-only mode. In this variant, the encoded payload is
T ( H ) = ( M , S , T h ) .
After outer decoding, the receiver accepts the candidate only if the hash consistency condition holds.
Hybrid mode. In the full hybrid construction, both verification layers are included:
T ( A H ) = ( M , R , T A M D , S , T h ) .
A decoded candidate is accepted only when it passes both checks.
These four modes make it possible to compare the separate and joint effects of AMD and universal hash verification under the same outer-code and channel model.
Under the construction above, the system-level undetected error event is defined as
E u e = { M ^ ≠ M   and   the   receiver   accepts   M ^ } .
This event is stronger than decoder error alone. A decoding error produces a system-level failure only if the incorrect candidate also survives the verification stage associated with the chosen protection mode. Consequently, the hybrid architecture separates two logically distinct failure mechanisms:
  • The failure of the outer code to recover the correct tagged object.
  • The failure of the verification stage to reject an incorrect decoded candidate.
This separation is exactly what enables the factorized undetected error to bound, derived in Section 5.

4.5. Redundancy Structure of the Construction

The hybrid scheme introduces redundancy at two different levels. If L is the binary length of the tagged object T ( M ; R , κ ) , then the verification redundancy is
ρ v e r = L − k ,
while the outer-code redundancy is
ρ o u t = n − L .
Hence, the total redundancy is
ρ t o t = n − k = ρ v e r + ρ o u t .
This decomposition is important because the two redundancy terms play different roles. The outer-code redundancy determines the distance-based correction capability of the Reed–Muller layer, whereas the verification redundancy determines the strength of the AMD and hash checks. The proposed architecture therefore allows reliability improvement to be distributed between coding protection and verification protection in a mathematically transparent way.

5. Theoretical Analysis

This section derives upper bounds on the probability of undetected error defined in Equation (27). Section 5.1 establishes the sequential verification bound under explicit randomness and conditional-independence assumptions. Section 5.2 derives a general system-level bound using the bounded-distance correction radius in Equation (10) and the BSC tail in Equation (5). Section 5.3 then refines the decoder term using Reed–Muller-specific weight distribution and decoding region information.

5.1. Explicit Bound for the Hybrid Construction

Lemma 1. 
Let  D  be the event that the decoder outputs an incorrect tagged candidate. Let  A A M D  and  A H  denote acceptance by the AMD and universal hash checks, respectively. Assume that:
  • For every admissible incorrect decoded candidate,
    P r A A M D ∣ D ≤ ε A M D ;
    the hash seed is sampled independently of the variables determining  D  and  A A M D ;
  • Conditioned on  D  and  A A M D , the candidate message pair remains fixed and distinct before evaluation over the hash seed;
  • The hash family is seeded 2-universal with   l -bit output. Then,
    P r A A M D ∩ A H ∣ D ≤ ε A M D 2 − l .
Proof. 
By the conditional chain rule,
P r A A M D ∩ A H ∣ D = P r A A M D ∣ D P r   ( A H ′ ∣ D A A M D ) .
The AMD security assumption gives
P r A A M D ∣ D ≤ ε A M D .
By the independence and freshness assumptions on the hash seed, conditioning on D and A A M D does not alter the prescribed distribution of S . The incorrect decoded message remains distinct from the transmitted message, so the 2-universal collision property gives
P r   ( A H ′ ∣ D A A M D ) ≤ 2 − l .
Substituting Equations (32) and (33) into Equation (31) yields
P r A A M D ∩ A H ∣ D ≤ ε A M D 2 − l .
This proves the claim. □
Theorem 1. 
General hybrid undetected error bound under sequential verification.
Let C = R M ( r ,   m ) be a binary Reed–Muller code of length with minimum distance and bounded-distance decoding radius. Assume transmission over B S C p ( p ) . Let D denote the event that the bounded-distance decoder outputs an incorrect tagged candidate. Suppose the AMD and universal hash verification layers satisfy the assumptions of Lemma 1. Then,
P U E = P r D ∩ A A M D ∩ A H ≤ P r D ε A M D 2 − l .
Since a bounded-distance decoder is guaranteed to return the transmitted codeword whenever w t E ≤ t ,
P r D ≤ P r w t E > t .
Therefore,
P U E ≤ ε A M D 2 − l ∑ j = t + 1 n n j p j 1 − p n − j .
Proof. 
By definition, the system-level undetected error event is
U = D ∩ A A M D ∩ A H .
Applying conditional probability,
P r U = P r D P r A A M D ∩ A H ∣ D .
Lemma 1 gives
P r A A M D ∩ A H ∣ D ≤ ε A M D 2 − l .
Consequently,
P r U ≤ P r D ε A M D 2 − l .
For bounded-distance decoding, the decoder is guaranteed to recover the transmitted codeword whenever the channel-error weight does not exceed t . Therefore,
D ⊆ w t E > t ,
and hence,
P r D ≤ P r w t E > t .
Since E is generated by B S C ( p ) , its Hamming weight is binomially distributed:
P r w t E > t = ∑ j = t + 1 n n j p j 1 − p n − j .
Combining Equations (40)–(42) yields Equation (37). □
Theorem 2. 
RM-spectrum-based hybrid undetected error bound. Let  C = R M ( r , m )  be used over  B S C ( p )  with a bounded-distance decoder of radius  t . Let  A w  be the weight distribution of  C , and let  P m i s c R M p  be defined by Equation (43). Suppose that, conditional on a fixed incorrect decoded tagged candidate, the joint probability that the candidate passes both verification layers is at most  ε v e r . Then, 
P U E h y b ≤ P m i s c R M p ε v e r .
Under the additional assumptions required to establish independent AMD and hash verification bounds,
ε v e r ≤ ε A M D 2 − l ,
and therefore,
P U E h y b ≤ ε A M D 2 − l ∑ w = d m i n n A w ∑ j = m a x 0 w − t m i n n w + t N t n w j p j 1 − p n − j .
Proof. 
Because C is linear and the BSC is symmetric, the all-zero codeword may be assumed to have been transmitted. A bounded-distance miscorrection occurs when the received vector lies within radius t of a nonzero codeword c ∈ C . For a codeword of weight w , N t ( n , w , j ) counts the weight- j vectors lying within its decoding sphere. Each such vector occurs with BSC probability p j 1 − p n − j . Summing over channel weights, codeword weights, and the corresponding multiplicities A w yields P m i s c R M p .
An undetected system error requires both a bounded-distance miscorrection and acceptance by the enabled verification mechanisms. By conditioning on the incorrect decoded candidate and applying the uniform joint verification bound ε v e r ,
P U E h y b = ∑ T ^ ≠ T P T ^   is   decoded P T ^   is   accepted ∣ T ^   is   decoded ≤ ε v e r ∑ T ^ ≠ T P T ^   is   decoded = P m i s c R M p ε v e r .
The final expression follows when the revised verification theorem establishes
ε v e r ≤ ε A M D 2 − l .
□

5.2. Refined Finite-Length Bound Based on Decoder Miscorrection

The correction-radius tail used in the general theorem is convenient because it depends only on the code length, minimum distance, and channel crossover probability. However, it is generally not equal to the decoder-miscorrection probability. For a bounded-distance decoder, the event
W > t
includes both incorrect decoding and explicit decoder failure. Consequently,
P m i s c o r r ≤ P r W > t .
A tighter system-level representation is obtained by conditioning directly on the decoder-miscorrection event D m i s :
P U E = P r D m i s P r A A M D ∩ A h ∣ D m i s .
Let
P m i s c o r r = P r D m i s
and
P v e r , a c c ∣ m i s c o r r = P r A A M D ∩ A h ∣ D m i s .
Then,
P U E = P m i s c o r r P v e r , a c c ∣ m i s c o r r .
Under the sequential verification assumptions of Lemma 1,
P v e r , a c c ∣ m i s c o r r ≤ ε A M D 2 − l ,
and therefore
P U E ≤ P m i s c o r r ε A M D 2 − l .
Equation (53) is tighter than the generic correction-radius result whenever P m i s c o r r is evaluated or upper-bounded more accurately than P r W > t .
Reed–Muller-specific miscorrection bound. Let C = R M ( r , m ) , and let A w denote the number of codewords of Hamming weight w .
Assume that the all-zero codeword is transmitted, which is valid for a linear code over a symmetric channel. For a nonzero codeword c of weight w , define the bounded-distance decoding sphere
B t c = y ∈ F 2 n : d H y c ≤ t .
A miscorrection to c occurs only when
Y ∈ B t c .
Therefore,
P m i s c o r r = ∑ c ∈ C ∖ 0 P r Y ∈ B t c ,
provided the bounded-distance decoding spheres are disjoint and the decoder outputs a codeword only when the received vector lies in one such sphere. If the implementation uses another decoder rule, use ≤ rather than equality.
For a codeword of weight w , the number of weight- j error vectors lying within radius t of that codeword is
N t n w j = ∑ a = a m i n a m a x w a n − w j − w + a ,
where
a m i n = m a x 0 w − j ,
and
a m a x = min w t + w − j 2 .
The RM-specific miscorrection bound is then
B R M p = ∑ w = d m i n n A w ∑ j = max 0 w − t min n w + t N t ( n , w , j ) p j 1 − p n − j .
Thus,
P m i s c o r r ≤ B R M p ,
and
P U E ≤ B R M p ε A M D 2 − l .
The relationship between the bounds is
P m i s c o r r ≤ B R M p ≤ B t a i l p ,
where
B t a i l p = ∑ j = t + 1 n n j p j 1 − p n − j .
Theorem 3. 
Miscorrection-based hybrid undetected error bound. Let  C = R M ( r , m )  be used over  B S C p  with a bounded-distance decoder of radius  t . Let  D m i s  denote the event that the decoder outputs an incorrect tagged candidate. Suppose the verification assumptions of Lemma 1 hold. Then, 
P U E ≤ P m i s c o r r ε A M D 2 − l .
If  B R M p  is the Reed–Muller decoding region bound defined in Equation (57), then 
P U E ≤ B R M p ε A M D 2 − l .
Moreover, 
B R M p ≤ B t a i l p ,
so the code-specific result is no weaker than the generic correction-radius tail bound.
Proof. 
By definition,
P U E = P r D m i s P r A A M D ∩ A h ∣ D m i s .
Lemma 1 gives
P r A A M D ∩ A h ∣ D m i s ≤ ε A M D 2 − l .
Therefore,
P U E ≤ P m i s c o r r ε A M D 2 − l .
For a bounded-distance decoder, an incorrect output can occur only when the received vector enters the radius- t decoding region of a nontransmitted codeword. Summation over nonzero codewords, grouped according to the Reed–Muller weight distribution, yields the bound B R M p . This proves Equation (63).
Finally, every miscorrection requires more than t channel errors, whereas not every error pattern of weight greater than t causes a miscorrection. Hence,
D m i s ⊆ W > t ,
which yields Equation (64). □
Empirical miscorrection refinement. If the exact RM weight enumerator or decoding region calculation is not available for every code, then it provides an empirical finite-length refinement.
Define
P ^ m i s c o r r = N m i s N M C ,
and
P ^ v e r , a c c ∣ m i s c o r r = N U E N m i s .
Then,
P ^ U E = P ^ m i s c o r r P ^ v e r , a c c ∣ m i s c o r r .
A semi-empirical upper estimate is
B s e m i = P ^ m i s c o r r   U ε A M D 2 − l ,
where P ^ m i s c o r r   U is the upper endpoint of a confidence interval for the empirical miscorrection probability.
This approach is useful because it separates:
  • Looseness due to the decoder term;
  • Looseness due to the verification term.

5.3. Reed–Muller Parameterization

Substituting the Reed–Muller parameters into Theorem 1 gives the explicit form
P u e ( r , m , p , l ) ≤ ∑ i = 2 m − r − 1 2 m 2 m i p i ( 1 − p ) 2 m − i ε A M D   2 − l ,
where the lower summation index is equivalent to t + 1 because t = 2 m − r − 1 − 1 . This expression displays the dependence of the undetected error probability on all principal system parameters. The code order r , the code length exponent m , the BSC crossover probability p , the hash length l , and the AMD security parameter ε A M D .
The interpretation is immediate. Increasing m increases block length and changes the binomial tail governing bounded-distance decoding failure. Decreasing r increases the Reed–Muller minimum distance and therefore enlarges the guaranteed correction radius. Increasing l decreases the seeded hash collision term exponentially. Strengthening the AMD construction decreases ε A M D . Accordingly, outer-code design and verification-layer design remain mathematically separable while contributing jointly to the final integrity guarantee.

5.4. Comparison with Outer-Code-Only Protection

The same framework yields an immediate comparison with the corresponding outer-code-only architecture. If no AMD or hash verification is used, then the system-level undetected error probability is bounded only by the residual decoding-failure probability,
P u e o u t e r ≤ P r { w H ( E ) > t } .
Therefore, by Theorem 1,
P u e h y b r i d ≤ P u e o u t e r   ε A M D   2 − l .
This shows that, within the specified model, the hybrid construction improves the outer-code-only bound by an explicit multiplicative factor
ε A M D   2 − l .
Whenever ε A M D < 1   a and l ≥ 1 , the hybrid scheme provides a strictly stronger upper bound than the corresponding outer-code-only architecture. This formalizes the main design advantage of the proposed layered construction.

5.5. Exponential Decay Regime

The explicit form of Theorem 1 also permits an asymptotic interpretation when the Reed–Muller order r is fixed and m → ∞ . In this regime,
n = 2 m ,     t n = 2 m − r − 1 − 1 2 m = 2 − r − 1 − 2 − m .
Hence, for sufficiently large m , the correction threshold is asymptotically close to 2 − r − 1 n . If the BSC crossover probability satisfies p < 2 − r − 1 , then standard large-deviation estimates for the binomial tail imply exponential decay of   P r { w H ( E ) > t } in the block length n . Therefore, for fixed ε A M D and fixed l , the undetected error probability also decays exponentially in n . This is stated below.
Corollary 1. 
Fix  r ,  ε A M D , and  l . If  p < 2 − r − 1 ,  then there exists a positive constant  c ( r , p )  such that
P u e ≤ e − c ( r , p ) n   ε A M D   2 − l
for all sufficiently large  m , where  n = 2 m .
Proof. 
By Theorem 1,
P u e ≤ P r { w H ( E ) > t } ε A M D 2 − l .
Since w H ( E ) ∼ B i n o m i a l ( n , p ) and t / n → 2 − r − 1 , the assumption p < 2 − r − 1 implies that the tail probability P r { w H ( E ) > t } decays exponentially in n by standard large-deviation bounds for binomial distributions. Multiplying by the fixed factor ε A M D 2 − l yields the claim.
Corollary 1 shows that hybrid architecture preserves the exponential reliability gain associated with the outer code while adding verification-level suppression through the AMD and hash factors. In this sense, the outer code controls the large-deviation decay regime, whereas the inner verification layers shift the final undetected error probability downward by additional multiplicative factors. □

6. Simulation Validation

This section evaluates the proposed hybrid integrity protection framework under the binary symmetric channel B S C ( p ) . The empirical implementation uses the original R M ( r , m ) configuration together with the GF(4)-based AMD construction and a 2-bit universal hash. To examine the influence of code parameters and verification strength more broadly, analytical parameter sweeps are additionally reported for several Reed–Muller codes, hash lengths, and AMD security levels.
The empirical and analytical results are kept separate. Empirical probabilities are reported only for the implemented R M ( 2 ,   5 ) system, whereas the additional Reed–Muller configurations are compared using the generic BSC correction-radius bound.

6.1. Simulation Configuration

For a binary Reed–Muller code R M ( r , m ) the block length, dimension, minimum distance, and bounded-distance decoding radius. Table 4 summarizes the Reed–Muller configurations considered in the analytical comparison.
Table 4. Reed–Muller configurations considered in the validation.
The empirical baseline uses R M ( 2 ,   5 ) a 2-bit source message, the GF(4)-based AMD construction, and a seeded 2-bit universal hash. The remaining information positions are filled with zeros. Four protection modes are evaluated:
T O = M ,
T A = M R T A M D ,
T H = M S T h ,
T A H = M R T A M D S T h .
These correspond to outer-code-only, AMD-assisted, hash-assisted, and hybrid protection, respectively. The principal validation parameters are listed in Table 5.
Table 5. Empirical and analytical validation settings.
For each Monte Carlo trial, a source message is encoded, transmitted through B S C ( p ) , and decoded. A decoder output is classified as correct, explicitly rejected, miscorrected and rejected by verification, or miscorrected and accepted. The last event is the system-level undetected error. The empirical undetected error probability is
P ^ U E = N U E N M C ,
where N U E is the number of undetected error events and N M C is the total number of trials. When no undetected event is observed, the result is interpreted as a finite-sample zero rather than as exact impossibility.

6.2. Empirical Results for R M ( 2 , 5 )

Table 6 and Figure 2 report the empirical results for the implemented R M ( 2 ,   5 ) system. The decoder-error estimate varies slightly among protection modes because each mode uses a different subset of valid payloads inside the same outer code.
Table 6. Empirical undetected error probabilities for R M ( 2 ,   5 ) .
Figure 2. Empirical undetected error probability of the four protection modes for R M 2 5 , together with the analytical hybrid upper bound. Zero-event results are shown as finite-sample upper limits rather than exact zero probabilities.
The outer-code-only mode produces the largest undetected error probability for every tested value of p , because any incorrect decoded source message is accepted automatically. Both single-layer verification modes reduce the probability substantially. The hybrid mode produces the smallest empirical undetected error probability at all nonzero observed points.
At p = 0.05 , for example, the hybrid probability is 8.0 × 10 − 5 , compared with 1.232 × 10 − 3   for the outer-code-only mode. This corresponds to an improvement factor of approximately
1.232 × 10 − 3 8.0 × 10 − 5 ≈ 15.4 .
At p = 0.10 , the corresponding improvement factor is
1.1068 × 10 − 2 7.28 × 10 − 4 ≈ 15.2 .
The analytical bound remains above the measured hybrid probability for all tested channel conditions. The gap is expected because the bound uses the complete binomial tail beyond the guaranteed correction radius and worst-case verification-acceptance factors. It therefore includes channel patterns that may lead to explicit decoder rejection rather than miscorrection.
The zero entries at p = 0.01 indicate that no undetected event was observed in the performed trials. They should not be interpreted as exact zero probabilities.

6.3. Influence of Verification Strength

The joint verification factor under the assumptions of Section 5 is ε v e r = ε A M D 2 − l .  Table 7 shows how this factor changes with the hash output length when ε A M D = 2 − 4 .
Table 7. Joint verification bound versus hash output length.
Each additional hash bit reduces the hash collision bound by a factor of two. Increasing l from two to 16 reduces the joint verification bound from approximately 1.56 × 10 − 2 to 9.54 × 10 − 7 , at the cost of 14 additional hash bits. Table 8 shows the corresponding effect of the AMD security level for a fixed 4-bit hash.
Table 8. Joint verification bound versus AMD security level for l = 4 .
Table 7 and Table 8 are analytical sensitivity results. They show the expected reduction in verification-acceptance probability but do not represent additional empirical AMD implementations.

6.4. Influence of Reed–Muller Parameters

For a bounded-distance decoder, an incorrect output can occur only when the channel-error weight exceeds t . The generic BSC-tail bound is
B t a i l p = ∑ j = t + 1 n n j p j 1 − p n − j .
Using ε A M D = 2 − 4 ,   l = 4 , the generic hybrid bound becomes
B h y b p = 2 − 8 B t a i l p .
Table 9 compares this bound across the Reed–Muller configurations in Table 4.
Table 9. Generic BSC-tail and hybrid bounds for different Reed–Muller codes.
The analytical comparison illustrates the expected rate–distance trade-off. The configurations R M ( 1 ,   5 ) and R M ( 2 ,   6 ) , both with correction radius t = 7 , provide substantially smaller BSC tails than the configurations with t = 3 at low and moderate crossover probabilities. However, this advantage is accompanied by lower code rates. For example, R M ( 1 ,   5 ) has the smallest analytical tail among the tested configurations but also the lowest rate, 6 / 32 = 0.1875 . In contrast, R M ( 3 ,   6 ) has the highest rate, 42 / 64 = 0.6563 , but the largest analytical tail at p = 0.05 and p = 0.10 (see Figure 3).
Figure 3. Generic hybrid upper bound versus BSC crossover probability for the evaluated Reed–Muller configurations. The curves illustrate the trade-off among code rate, minimum distance, and guaranteed correction radius.
The values in Table 10 are upper bounds rather than measured decoder-miscorrection probabilities. They count all error patterns beyond the guaranteed correction radius, including patterns that may cause explicit decoding failure.
Table 10. Decomposition and tightness of the hybrid bound for R M ( 2 ,   5 ) .
The B t a i l values and ratios should be recalculated from the final code rather than copied manually. The purpose of the table is to show that the largest source of looseness is usually the replacement of actual decoder miscorrection with the full correction-radius tail. For the current baseline, the verification factor is
ε A M D 2 − l = 0.25 × 0.25 = 0.0625 .
At p = 0.05 ,
B t a i l ≈ 7.38 × 10 − 2 ,
while the empirical decoder-error estimate is approximately
2.99 × 10 − 3 .
Thus, the generic decoder term is approximately
7.38 × 10 − 2 2.99 × 10 − 3 ≈ 24.7
times larger than the measured decoder-error estimate. This already explains a substantial part of the final gap.

6.5. Redundancy Interpretation

The four protection modes differ in verification overhead. Let q denote the source-message length, ρ A M D the combined AMD randomness and tag length, s the hash-seed length, and l the hash output length. The total redundancy is
ρ t o t a l = ρ v e r + ρ p a d + n − k ,
where ρ p a d denotes unused information positions.
The hybrid mode provides the strongest empirical protection in Table 11, but it also introduces the largest verification overhead. Consequently, its advantage should be interpreted as an integrity–redundancy trade-off rather than as an unconditional improvement.
Table 11. Redundancy structure of the protection modes.
The empirical R M ( 2 ,   5 ) results confirm the principal ordering predicted by the proposed model. Outer-code-only protection produces the highest undetected error probability, both single verification layers provide substantial reductions, and the combined hybrid mode produces the smallest measured probability.
The analytical parameter sweeps show that stronger AMD and hash parameters reduce the verification-acceptance bound multiplicatively under the assumptions of Section 5. They also show that Reed–Muller performance depends jointly on rate and correction radius. Low-rate configurations with larger t provide smaller BSC-tail bounds, whereas higher-rate configurations provide greater payload efficiency at the cost of weaker correction guarantees. The empirical evidence remains limited to the implemented RM 2 ,   5 configuration. The additional Reed–Muller results in Table 9 are analytical comparisons and should not be interpreted as measured scalability results.

6.6. Complexity, Redundancy, and Fair-Budget Comparisons

Comparing only undetected error probability favors schemes with greater verification redundancy. The hybrid mode includes an AMD-protected representation, a hash seed, and a hash tag in addition to the redundancy introduced by the outer Reed–Muller code. Therefore, the four protection modes are also evaluated in terms of transmitted bits, effective source rate, computation, storage, and verification latency.
Let q denote the original source-message length, ρ R denote the AMD randomness length, ρ T denote the AMD tag length, s denote the universal hash seed length, l denote the hash output length, k denote the outer-code information dimension, n denote the transmitted codeword length, and ρ p a d denote the number of unused information positions.
The AMD redundancy is
ρ A M D = ρ R + ρ T ,
and the hash redundancy is
ρ h = s + l .
The total verification redundancy is
ρ v e r = ρ A M D + ρ h ,
while the total transmitted overhead relative to the original source message is
ρ t o t a l = n − q .
Equivalently,
ρ t o t a l = n − k + ρ p a d + ρ v e r .
The effective source rate is
R e f f = q n .
These definitions ensure that fixed padding positions are counted as overhead rather than being omitted.
The exact bit values depend on the current GF(4)-based AMD and hash implementations. For the baseline system described in the manuscript, a 2-bit source message is used and the GF(4)-based AMD construction operates with one GF(4) randomness symbol and one GF(4) tag symbol. Since one GF(4) symbol corresponds to 2 bits, the baseline AMD representation contributes
ρ R = 2 ,   ρ T = 2 ,   ρ A M D = 4 .
The manuscript also states that the hash has a 2-bit output. In the implemented hash family, the seed length is s = 2 bits.
Under the baseline assumption q = 2 ,   ρ A M D = 4 ,   s = 2 ,   l = 2 , and using R M ( 2 ,   5 ) with k = 16 and n = 32 , the payload and overhead are as follows (see Table 12).
Table 12. Redundancy and effective-rate comparison for the baseline R M ( 2,5 ) implementation.
This table reveals an important limitation of the original fixed- R M ( 2 ,   5 ) experiment: all modes transmit the same 32-bit codeword and protect the same 2-bit source message. Therefore, their total transmitted overhead is identical, but the outer-only mode wastes more information positions as fixed zero padding. The hybrid mode uses more of the available k = 16 information positions for verification rather than increasing the transmitted block length. This is a fair comparison under the equal block length, equal source payload, and equal outer code (see Table 13). However, it is not a fair comparison under equal verification redundancy, because the outer-only mode has no post-decoding verification data.
Table 13. Asymptotic computational and storage cost of the evaluated protection modes.
The asymptotic order of the outer coding stage is unchanged by the verification layers. AMD and universal hashing introduce additive rather than multiplicative computational overhead. However, they consume information positions and require additional storage for randomness, seed, and tags.
Comparisons under equal block length and equal source payload show the effect of using available information positions for verification rather than padding (see Table 14). In the baseline RM 2 ,   5 experiment, all modes transmit a 32-bit block and protect a 2-bit source message. Thus, the hybrid mode does not increase channel block length relative to the outer-only mode, but it uses more of the 16 information positions for integrity data. Comparisons with CRC-assisted schemes require a distinction between random error detection and adversarial manipulation detection. CRC is computationally efficient and provides strong practical detection of random error patterns, but it does not offer the same algebraic manipulation security definition as AMD. Universal hashing provides a probabilistic collision guarantee under its seed model, while the AMD layer provides a guarantee for the specified algebraic tampering family. Therefore, no single baseline dominates under every threat model.
Table 14. Structure of the fair-budget comparison.
BCH and Polar baselines are also decoder-dependent. Their performance must be reported together with decoder type, rate, block length, and verification overhead. Raw undetected error probability alone is insufficient for declaring one architecture superior.

7. Discussion

The theoretical results in Section 5 and the numerical results in Section 6 provide a consistent interpretation of layered integrity protection. In the proposed framework, the outer Reed–Muller code and the two verification layers address different parts of the failure process. The outer code acts at the transmission-decoding stage and reduces the probability that channel perturbations produce an incorrect decoded candidate. By contrast, the AMD and universal hash layers act after decoding and reduce the probability that such an incorrect candidate is accepted as valid. This functional separation is the key structural feature of the hybrid architecture and explains why the resulting undetected error probability is lower than that of the corresponding outer-code-only system.
A central outcome of the analysis is that the undetected error probability can be factorized into two conceptually distinct terms: a decoder-miscorrection term determined by the outer code and the channel, and a verification-acceptance term determined by the AMD and hash layers. In the fully specified model adopted in this paper, the first term is expressed through the binomial tail
P r { W H ( E ) > t } ,
where t is the bounded-distance decoding radius of the outer Reed–Muller code, while the second term is controlled by the product
ε A M D 2 − l .
This decomposition is mathematically useful because it isolates the contributions of coding and verification, making it possible to study their influence separately. In particular, the analysis shows that improving the outer code and strengthening the verification layers are not interchangeable operations: they affect different logical stages of the integrity protection process.
The simulation study supports this interpretation. For the instantiated finite-length system based on R M ( 2,5 ) , the outer-only mode consistently produced the largest empirical undetected error probability, while the AMD-only and hash-only modes each yielded a substantial reduction. The combined hybrid mode achieved the smallest undetected error probability across all tested values of the channel crossover probability. This confirms the practical relevance of the factorized theoretical structure derived in Section 4. Although the analytical bound is conservative, the ordering of the empirical results agrees with the theorem and demonstrates that the layered screening principle remains effective in a concrete short-block implementation.
Another important point concerns the role of redundancy. In the present framework, total redundancy is naturally divided into verification redundancy and outer-code redundancy. This distinction is not merely notational. Outer-code redundancy increases the distance-based correction capability and therefore improves the probability that the correct tagged object reaches the verification stage. Verification redundancy, on the other hand, does not improve decoding itself, but reduces the probability that an incorrect decoded object is accepted. From a design perspective, this means that reliability improvement can be distributed across two independent mechanisms. In practical terms, once the outer code has reached a satisfactory correction capability, further reduction in undetected error probability may be achieved more efficiently by increasing the strength of the AMD or hash layer than by increasing the outer-code redundancy alone.
The paper also highlights the importance of using a fully specified model when analyzing hybrid integrity schemes. In the earlier version of the manuscript, the framework remained too abstract, and the numerical section was only illustrative. By fixing the channel to B S C ( p ) , the outer code to R M ( r , m ) , the decoder to bounded-distance decoding, and the verification mechanisms to explicit AMD and universal hash constructions, the analysis becomes both mathematically sharper and experimentally verifiable. This strengthens the contribution of the paper and makes the derived results reproducible.
The gap between the analytical upper bound and the empirical undetected error probability arises from several successive relaxations. First, the general theorem replaces the decoder-miscorrection event with the broader event that the number of channel errors exceeds the guaranteed correction radius. This binomial tail includes both miscorrections and explicit decoder failures. Second, the generic bound treats every error pattern beyond the correction radius as potentially leading to an incorrect accepted codeword, although many received vectors lie outside the bounded-distance region of every competing codeword. Third, the theorem applies the worst-case AMD acceptance probability to every incorrect candidate, even though the actual acceptance probability may be much smaller for the candidate distribution produced by the decoder. Fourth, the universal hash factor uses the worst-case collision bound 2 − l , which need not be attained by the observed candidate pairs. Fifth, the AMD and hash bounds may each be maximized by different incorrect candidates, so their product can remain conservative even when the sequential conditional assumptions are valid. Finally, the generic theorem does not use the Reed–Muller weight distribution, coset structure, or exact decoding regions.
To address the first and final sources of looseness, the revised manuscript introduces a miscorrection-based representation and an RM-specific decoding region bound. The revised result replaces the BSC correction-radius tail with P m i s c o r r , or with the code-specific upper bound B R M p . The remaining gap reflects the use of worst-case verification guarantees rather than the empirical distribution of incorrect decoded candidates.

Case Reed–Muller Codes May Not Be the Optimal Choice

Although Reed–Muller codes provide an algebraically structured setting for the proposed finite-length integrity analysis, they are not universally optimal outer codes. Their available dimensions and rates are discrete, and some configurations may result in low effective payload rates or a substantial number of unused information positions, particularly for very short messages. For high-rate or high-throughput communication, LDPC, Polar, BCH, or other rate-flexible codes may provide a more favorable balance among error-correction performance, decoder complexity, and implementation efficiency. Reed–Muller codes may also be less attractive at long block lengths when near-maximum-likelihood, list, recursive-projection, or ensemble decoding is required, because such decoding methods can introduce considerable computational and latency costs. In soft-decision AWGN or fading channels, modern LDPC and Polar decoders may make more effective use of channel-reliability information than the hard-decision bounded-distance model considered in this study. Similarly, channels dominated by burst errors, insertion/deletion errors, or synchronization failures may be better served by interleaved BCH or Reed–Solomon codes, marker codes, or other channel-specific constructions. Reed–Muller coding may also be unnecessary in systems that already employ standardized CRC-assisted LDPC or Polar architectures. In such cases, introducing an RM decoder could increase implementation complexity without providing a proportional system-level benefit. The proposed AMD and universal hash verification layers are therefore not inherently restricted to Reed–Muller codes and may be combined with alternative outer codes. The present study uses Reed–Muller codes because their explicit parameters, algebraic structure, and finite-length decoding regions facilitate the analytical treatment, rather than because they are claimed to be optimal for every communication scenario.
At the same time, several limitations should be acknowledged. First, the theoretical bound in Theorem 1 is an upper bound and is not intended to be tight for every finite-length configuration. The simulation results indeed show that the measured hybrid failure probability lies well below the theorem bound, especially in low-noise regimes. Second, the finite-length validation in Section 5 is based on a short Reed–Muller code, namely R M ( 2,5 ) , and therefore should be interpreted as a proof-of-concept validation rather than a large-scale performance study. Third, the specific AMD and hash instantiations used in the simulation were deliberately simple to make exhaustive finite-length validation feasible. More sophisticated AMD constructions or longer hash outputs would further reduce the empirical undetected error probability, but they would also increase verification redundancy. A further limitation is that the present work focuses on the binary symmetric channel and bounded-distance decoding. These assumptions were chosen to obtain an explicit and transparent analytical result, but they do not cover all practically relevant corruption models. In particular, the current analysis does not address erasures, burst errors, insertions, deletions, decoder mismatches, or adversarial channels with non-memoryless behavior. Likewise, only one outer-code family was considered. Reed–Muller codes are mathematically attractive because of their explicit parameters and classical structure, but the same layered verification principle could be studied with other code families as well. The multiplicative verification factor also depends on randomness and dependence assumptions. The product ε A M D 2 − l is justified only when the AMD guarantee applies to the incorrect candidate and the fresh universal hash seed remains independent after conditioning on the decoder and AMD acceptance events. The bound may not hold under malicious seed selection, adaptive manipulation after observing the seed, compromised random-number generation, repeated seed reuse, or dependence between the AMD and hash verification data. Under such conditions, a conservative intersection bound based on the smaller marginal acceptance probability may be used if the marginal conditional guarantees remain valid. If even those guarantees fail, the analysis reduces to the decoder-error bound alone. Extending the model to adaptive seed-aware adversaries and reused verification randomness is left for future work.
Despite these limitations, the revised results establish a clear methodological contribution. The paper shows that hybrid integrity protection can be analyzed within a single mathematical framework in which the outer code governs residual decoding failure and the verification layers govern acceptance failure. This perspective clarifies why a layered architecture can outperform a standalone distance-based code even when both operate under the same outer-code parameters. It also provides a principled basis for parameter selection: the outer code should be chosen to control the residual decoding-failure region, whereas the AMD and hash layers should be chosen to control the acceptance probability of incorrect decoded candidates.
From a broader viewpoint, the proposed framework is relevant to systems in which accepted corruption is more critical than ordinary decoding failure. In such systems, the objective is not merely to reduce the number of erroneous outputs, but to reduce the probability that an erroneous output is treated as valid. The hybrid construction studied here is well suited to this objective because it introduces additional acceptance tests after decoding rather than relying entirely on distance-based protection.
Overall, the discussion confirms that the main value of the proposed approach lies in its layered mathematical structure. The outer Reed–Muller code, the AMD layer, and the universal hash layer perform complementary functions, and their joint use yields stronger protection against undetected error than any of the individual components alone. The analytical and numerical results together support the conclusion that hybrid coding with post-decoding verification is a viable and mathematically grounded strategy for high-assurance information integrity.

8. Conclusions

This paper developed a hybrid coding framework for information integrity based on the joint use of a binary Reed–Muller outer code, an algebraic manipulation detection layer, and a seeded universal hash verification layer. In contrast to purely distance-based protection, the proposed architecture separates two logically distinct tasks: the outer code reduces the probability that channel perturbations produce an incorrect decoded candidate, while the AMD and hash layers reduce the probability that such a candidate is accepted as valid. This separation makes the framework especially suitable for the analysis of undetected error events, which constitute the central reliability criterion in high-assurance integrity systems.
The present study does not claim novelty for the individual use of Reed–Muller coding, AMD, universal hashing, or bounded-distance decoding. Its contribution lies in connecting these established mechanisms through a formally specified finite-length acceptance model. In this model, the outer decoder determines whether an incorrect tagged candidate is produced, whereas the verification layers determine whether that candidate is accepted. The principal coding-theoretic refinement is the replacement of the generic correction-radius tail by a Reed–Muller-specific miscorrection expression based on the code’s weight distribution and bounded-distance decoding regions. This refinement separates channel patterns that merely exceed the guaranteed correction radius from those that actually reach the decoding sphere of an incorrect codeword. Consequently, it provides a more informative finite-length integrity bound and clarifies how Reed–Muller structure affects the final probability of undetected acceptance. The numerical results should be interpreted as finite-length evidence for the stated models and parameter settings rather than as proof that Reed–Muller codes are universally optimal for integrity protection. The appropriate outer code depends on rate, block length, channel, decoder, latency, and implementation constraints. Reed–Muller codes are useful here because their algebraic structure and finite-length spectra permit a transparent code-specific analysis; BCH, Polar, LDPC, and CRC-assisted alternatives may be preferable in other operating regimes. The analytical results provide conservative integrity guarantees rather than exact performance predictions. The generic bound is simple and parameterized by r , m , p , l , and ε A M D , while the refined result incorporates decoder miscorrection and Reed–Muller decoding region information. The empirical results confirm the expected ordering of the protection modes, but the numerical gap shows that worst-case system-level bounds should not be interpreted as tight estimates of finite-length performance.
Another important conclusion is that the proposed framework admits a meaningful redundancy decomposition. Outer-code redundancy and verification redundancy play different roles and can therefore be tuned separately. The former determines the distance-based correction capability of the code, while the latter governs the acceptance probability of incorrect decoded candidates. This provides useful design flexibility and suggests that, once the outer code has reached a suitable correction regime, further improvements in undetected error performance may be achieved efficiently by strengthening the verification layers. The revised formulation of the paper also clarifies the mathematical novelty of the approach. Rather than presenting a generic architecture, the study analyzes a concrete hybrid model in which coding and post-decoding verification are integrated into a single probabilistic framework. This makes the contribution more rigorous, more reproducible, and more appropriate for a mathematics-oriented treatment of integrity protection.
Several directions remain open for future work. First, the present analysis could be extended to longer Reed–Muller codes and to additional code families with different distance and decoding properties. Second, stronger AMD constructions and larger universal hash outputs could be investigated to study the trade-off between verification redundancy and empirical performance. Third, the current channel model could be generalized to include burst errors, erasures, insertions, deletions, or adversarial perturbation models. Finally, it would be valuable to derive sharper finite-length bounds and asymptotic trade-off results for optimal allocation between outer-code redundancy and verification redundancy.
In summary, the results of this paper show that combining outer Reed–Muller coding with AMD-based and universal hash-based post-decoding verification provides a mathematically coherent and practically effective strategy for reducing undetected error probability. The proposed hybrid framework therefore offers a useful foundation for the design and analysis of high-assurance information-integrity systems.

Author Contributions

Conceptualization, B.A.S., A.A. (Akmal Abduvaitov) and J.I.; methodology, B.A.S., A.A. (Akmal Abduvaitov), K.H. and S.B.; software, A.A. (Abbos Abduvaytov) and B.A.S.; validation, A.A. (Abbos Abduvaytov), B.A.S. and A.A. (Akmal Abduvaitov); formal analysis, R.R.; investigation, B.A.S., J.I. and K.H.; resources, S.B. and O.M.; data curation, A.A. (Akmal Abduvaitov), A.A. (Abbos Abduvatov) and S.B.; writing—original draft preparation, B.A.S. and H.S.J.; writing—review and editing, H.S.J. and A.A. (Aziza Akhmedova); visualization, J.I. and K.H.; supervision, A.A. (Aziza Akhmedova) and R.R.; project administration, H.S.J., J.I. and A.A. (Abbos Abduvaytov); funding acquisition, H.S.J. and O.M. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (No. RS-2024-00412141). This research was also supported by the Regional Innovation System & Education (RISE) program through the (Chungbuk Regional Innovation System & Education Center), funded by the Ministry of Education (MOE) and the (Chungcheongbuk-do), Republic of Korea (2026-RISE-11-003-03).

Data Availability Statement

Data are contained within the article.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Abbe, E.; Shpilka, A.; Ye, M. Reed-Muller codes: Theory and algorithms. IEEE Trans. Inf. Theory 2021, 67, 3251–3277. [Google Scholar] [CrossRef] [Scilit]
  2. Ye, M.; Abbe, E. Recursive projection-aggregation decoding of Reed-Muller codes. IEEE Trans. Inf. Theory 2020, 66, 4948–4965. [Google Scholar] [CrossRef] [Scilit]
  3. Li, S. On the Weight Distribution of Second-Order Reed–Muller Codes and Their Relatives. Des. Codes Cryptogr. 2020, 88, 2447–2465. [Google Scholar]
  4. Muhiddinov, M.; Ochilov, N.; Umurkulov, B.; Kholnazarov, U.; Sapaev, I.; Bazarova, N.; Kholova, M.; Aripova, G. Privacy-Aware Information Security for E-Learning Platforms in History Using Attribute-Based Encryption Algorithm. J. Internet Serv. Inf. Secur. 2025, 15, 305–315. [Google Scholar] [CrossRef] [Scilit]
  5. Dumer, I. Recursive decoding of Reed-Muller codes. arXiv 2017, arXiv:1703.05303. [Google Scholar]
  6. Gajraj, S.; Umarov, S.; Kamilova, S.; Abdullayev, I.; Nayimov, S.; Sandeep, D.; Enoch, A. Machine Learning-Assisted Automated VLSI Design for Bioinformatics Hardware Accelerators With Embedded Cryptographic Security. J. VLSI Circuits Syst. 2025, 7, 30–36. [Google Scholar] [CrossRef] [Scilit]
  7. Hauck, P.; Huber, M.; Bertram, J.; Brauchle, D.; Ziesche, S. Efficient majority-logic decoding of short-length Reed-Muller codes at information positions. IEEE Trans. Commun. 2013, 61, 930–938. [Google Scholar] [CrossRef] [Scilit]
  8. Kaufman, T.; Lovett, S.; Porat, E. Weight distribution and list-decoding size of Reed-Muller codes. IEEE Trans. Inf. Theory 2012, 58, 2689–2696. [Google Scholar] [CrossRef] [Scilit]
  9. Bhowmick, A.; Lovett, S. List decoding Reed-Muller codes over small fields. In Proceedings of the forty-seventh annual ACM symposium on Theory of Computing (STOC’15); Association for Computing Machinery: New York, NY, USA, 2015; pp. 277–285. [Google Scholar] [CrossRef] [Scilit]
  10. Sharanya; Vani, A.; Thilagam, K.; Ramkumar, S.; Bakhritdinov, F.; Alzubaidi, Y.T.; Al-Hiti, A.S.; Munimathan, A.; Rahma, M.K.; Khishe, M. A hybrid deep learning framework for cooperative resource management in 5G using spike-driven transformers and cycle-consistent adaptation. Results Control Optim. 2026, 24, 100773. [Google Scholar] [CrossRef] [Scilit]
  11. Rasulov, A.; Sattorova, Z.; Yuldoshov, Y.; Pardaev, A.; Musabekova, M.; Abdullayeva, U.; Sapaev, I.B. Integrated Approaches to Enhancing Cross-Layer Communication and Antenna Design for Next-Generation Wireless Networks. Natl. J. Antennas Propag. 2025, 7, 116–122. [Google Scholar] [CrossRef] [Scilit]
  12. Cramer, R.; Fehr, S.; Padró, C. Algebraic manipulation detection codes. Sci. China Math. 2013, 56, 1349–1358. [Google Scholar] [CrossRef] [Scilit]
  13. Cramer, R.; Padró, C.; Xing, C. Optimal algebraic manipulation detection codes in the constant-error model. In Theory of Cryptography; Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2015; Volume 9014, pp. 481–501. [Google Scholar]
  14. Shao, M. Algebraic manipulation detection codes via highly nonlinear functions. arXiv 2020, arXiv:2002.03724. [Google Scholar]
  15. Jongsma, E. Algebraic Manipulation Detection Codes. Bachelor’s Thesis, Leiden University, Leiden, The Netherlands, 2009. [Google Scholar]
  16. Carter, J.L.; Wegman, M.N. Universal classes of hash functions. J. Comput. Syst. Sci. 1979, 18, 143–154. [Google Scholar] [CrossRef] [Scilit]
  17. Wegman, M.N.; Carter, J.L. New hash functions and their use in authentication and set equality. J. Comput. Syst. Sci. 1981, 22, 265–279. [Google Scholar] [CrossRef] [Scilit]
  18. Stinson, D.R. Universal hashing and authentication codes. In Advances in Cryptology-CRYPTO ’91; Springer: Berlin/Heidelberg, Germany, 1992; pp. 74–85. [Google Scholar]
  19. Lin, S.; Costello, D.J. Error Control Coding: Fundamentals and Applications, 2nd ed.; Pearson: Upper Saddle River, NJ, USA, 2004. [Google Scholar]
  20. Huffman, W.C.; Pless, V. Fundamentals of Error-Correcting Codes; Cambridge University Press: Cambridge, UK, 2003. [Google Scholar]
  21. McEliece, R.J. The Theory of Information and Coding, 2nd ed.; Cambridge University Press: Cambridge, UK, 2002. [Google Scholar]
  22. van Lint, J.H. Introduction to Coding Theory, 3rd ed.; Springer: Berlin/Heidelberg, Germany, 1999. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.