Abstract
Entropy estimation from short samples recurs in symmetric cryptography, where the reference distribution is uniform by design and no single estimator in the considered classical comparison set minimizes mean squared error (MSE) across the full range of ratios . We introduce the adaptive sample-conditional entropy diagnostic (ASED), in which a compact network trained offline maps a frequency-of-frequencies descriptor of the sample to convex mixture weights over six classical estimators, at a cost of . We prove an oracle inequality bounding the excess risk of such a mixture by the error of its weights and fixed-alphabet consistency for every simplex-valued weighting rule, independently of the distribution used to train the weights. Under the uniform-null protocol, ASED attains an integrated MSE of for bytes, compared with for James–Stein shrinkage, a reduction that depends materially on the aggregation scheme: for summed or averaged MSE across sample sizes versus for the mean of per-size ratios. Both figures are reported together throughout, and neither is presented as the headline; the entire advantage is confined to the undersampled regime , the two estimators being indistinguishable for . A locked train/validate/test evaluation with an independently written implementation reproduces the reduction at . Because a constant output achieves zero error by construction here, we add further checks: in a post hoc analysis, ASED is non-inferior to SHR in detection power at a margin, both pointwise and simultaneously, whereas a constant control has none, and the advantage persists on held-out configurations. The margin was selected after inspecting the power estimates, and at and , it is smaller than the variation induced by tie handling at the empirical critical value, so the non-inferiority conclusion is confirmatory only for . At , the two statistics have identical null distributions up to monotone relabeling, and no test of size exists, so no power difference is identifiable in either direction, and the comparison is withdrawn there. For strongly non-uniform sources, the ordering reverses, and NSB dominates, delimiting ASED as a uniformity diagnostic rather than a general-purpose entropy estimator. The design choices behind ASED, the feature map, architecture, and loss weighting were made while observing results on this same evaluation protocol, so no independent model-selection split separates development from the assessment reported here. That pipeline is accordingly designated exploratory, and the locked evaluation is confirmatory. The contribution is stated as sample-conditional convex aggregation of classical estimators rather than as the particular network realizing it: a five-feature variant and a gradient-boosted surrogate perform comparably, with no paired interval separating the three. Claims are restricted to short i.i.d. samples and marginal Shannon-entropy diagnostics under a uniform null; min-entropy, unpredictability, dependence, and randomness certification are out of scope.
1. Introduction
Shannon entropy measures the uncertainty of a discrete information source through the distribution of its output symbols [1]. In symmetric cryptography, it plays a dual role. At design time, entropy is a quantitative criterion: candidate S-boxes, filtering functions, and generator configurations are compared, in part, by how close the empirical entropy of their outputs lies to the theoretical maximum attained by the uniform distribution over an alphabet of size k. At evaluation time, entropy estimation underpins statistical batteries and health tests: a PRNG, a stream cipher keystream, or the output of an S-box construction are expected to be indistinguishable from uniform, so any systematic deviation of the estimated entropy from signals a detectable bias with potential security consequences [2,3].
The samples available for these judgments are often short, far shorter than the asymptotic regime in which classical estimators behave well. Standardized test suites expose short-window behavior: NIST SP 800-22 [4] and TestU01 [5] include statistics computed over local block structures or subwindows, while SP 800-90B [6] includes restart and health-test components for which short-window behavior is directly relevant. Design loops aggravate the problem: evolutionary and heuristic S-box search procedures evaluate thousands of candidate constructions per run, so the per-candidate sample budget is necessarily small [7], and initialization-phase analysis of stream ciphers examines the first few hundred output symbols [3] precisely. Diagnosing marginal biases that manifest only over short windows, a biased initialization vector expansion, and a transient change in symbol frequencies require estimators that remain reliable when n is comparable to or smaller than the alphabet size. Temporal-dependence failures, such as a transient correlation after reseeding, fall outside this scope: they may leave the marginal distribution exactly uniform and are the province of dependence, autocorrelation, or conditional-entropy tests, as Section 4.10 illustrates.
Estimating entropy from short samples is intrinsically difficult. No unbiased estimator exists [8], and the convergence rate of any consistent estimator can be made arbitrarily slow [9]. A sample of size is necessary for consistent estimation at fixed precision [10], a bound routinely violated in the block-level settings above. In the undersample regime where , the maximum-likelihood plug-in estimator exhibits severe downward bias, recovering only a fraction of the true entropy [11]. Substantial literature responds with correction terms [12,13,14], Bayesian priors [15,16,17,18], coverage-based methods [10,19,20,21], and empirical Bayes shrinkage [22]. The most comprehensive comparison for the cryptographic setting is due to Contreras Rodríguez et al. [2], who evaluated a broad battery of classical estimators over byte and bit sequences from to and found the James–Stein shrinkage estimator (SHR) to be the most effective single estimator for byte samples.
That comparison, however, also exposed a structural difficulty: the identity of the best estimator depends on the regime ratio . Bayesian smoothers dominate deep in the undersample regime, shrinkage dominates in an intermediate band, and bias-corrected plug-in estimators take over as grows. A practitioner must nevertheless commit to a single estimator family before seeing the data, even though some members tune their parameters from the sample, and a mismatch between the chosen estimator and the actual regime results in avoidable bias and inflated variance. No estimator in the comparison adaptively combines heterogeneous estimator families using sample-level occupancy information: several tune their own internal parameters from the data; SHR selects its shrinkage intensity; NSB integrates a data-dependent posterior; and coverage-based estimators use sample statistics, but each remains a single family. ASED introduces this second-level adaptivity by mapping the observed occupancy profile to mixture weights, as the regime-dependence of the ranking calls for.
A scope clarification is necessary before describing our proposal, because the word “entropy” covers two distinct quantities in cryptographic practice. Security certification of entropy sources is the business of SP 800-90B [6], and AIS 20/31 [23] is formulated in terms of min-entropy , which lower-bounds the work of an optimal guessing adversary and is deliberately more conservative than Shannon entropy. The present paper does not address min-entropy estimation. Its setting is the complementary one in which Shannon entropy is the established figure of merit: comparative design diagnostics and uniformity testing of PRNG, S-box, and keystream outputs, where the null hypothesis is the uniform distribution and the quantity of interest is the gap [2,3]. Conclusions drawn here transfer to certification pipelines only insofar as Shannon-entropy diagnostics complement, rather than replace, min-entropy estimators. A second restriction belongs here rather than in the results, because it bounds everything that follows: ASED is trained against the uniform null and is a uniformity-oriented Shannon-entropy diagnostic, not a general-purpose entropy estimator. On strongly skewed (Zipf ) or periodic sources, it is outperformed by NSB and by SHR, in places by more than an order of magnitude, and Section 4.9 quantifies that boundary rather than eliding it. No claim in this paper extends beyond near-uniform i.i.d. sources over the sample-size grid evaluated here, and the method should not be used as a screening tool on sources whose marginal is expected to be far from uniform.
This paper proposes the adaptive sample-conditional entropy diagnostic (ASED), an estimator that addresses the regime-commitment problem by learning a convex mixture of base estimators whose weights are conditioned on a compact occupancy-profile representation of the observed sample. The representation is derived from the frequency-of-frequencies (FoF) histogram, a vector that counts how many symbols were observed exactly i times, capturing the macroscopic regime (, coverage, singleton density) without reference to symbol identities. A lightweight multilayer perceptron (MLP), trained once offline on synthetic data, maps this representation to weights on the probability simplex over six complementary base estimators, so that adaptive estimation costs no more than a constant number of plug-in evaluations.
Designing the evaluation of such an estimator requires care on one methodological point, which we flag from the outset. ASED is trained against the uniform null, and under a purely uniform-null evaluation protocol, the constant “estimator” achieves zero error by construction. Low MSE under the null is therefore a necessary but not sufficient condition for usefulness: a meaningful adaptive estimator must, in addition, retain sensitivity to departures from uniformity. We therefore treat the uniform-null comparison as only the first of three evidentiary layers, and complete it with (i) a power analysis against near-uniform alternatives of known entropy, in which a constant control exhibits the confound explicitly and (ii) generalization tests on held-out sample sizes and an unseen alphabet.
Our contributions are threefold. First, we formalize the occupancy-profile regime descriptor derived from the FoF histogram, prove that every label-invariant plug-in estimator is a function of the FoF vector, and define the resulting 37-dimensional feature map (Section 3). Second, we introduce the ASED architecture: frequency-of-frequencies feature extraction, offline training of the weight mapper on synthetic data, and adaptive convex combination at inference, with inference complexity. For it we prove three results: an oracle inequality bounding its excess risk over the sample-conditional oracle by the approximation error of the learned weights; a checkable covariance condition under which no regime-indexed fixed convex mixture of the specified base estimators can improve on its best member whenever this condition holds, sample-dependent weighting is necessary for improvement within that mixture class; and a distribution-free bound under which ASED inherits the consistency and convergence rate of its base set whatever the training distribution, sharpened under a checkable hypothesis into a bound relative to the conditionally best estimator (Section 3). Appendix B adds an idealized PAC bound, kept out of the main development because its numerical value is vacuous at our training size. Third, we evaluate ASED in three layers: the uniform-null protocol of Contreras Rodríguez et al. [2] with 1000 replications and bootstrap uncertainty, a power analysis against near-uniform alternatives with the constant estimator as negative control, and generalization tests on held-out configurations, obtaining a reduction in integrated MSE over SHR for byte sequences of under summed or averaged aggregation and under the mean of per-size ratios two figures reported together throughout, since the headline depends materially on which is used (Section 4.4) with detection power non-inferior to SHR’s at a post hoc margin, confirmatory for and exploratory at and (Section 4.10.2), (Section 4 and Section 5).
The remainder of the paper is organized as follows. Section 2 reviews the entropy estimation literature relevant to the short-sample cryptographic setting. Section 3 presents the ASED model. Section 4 reports the experimental evaluation. Section 5 discusses scope, limitations, and practical implications, and Section 6 concludes.
2. Literature Review
2.1. Entropy Estimation: Foundations
The Shannon entropy of a discrete random variable X with alphabet and probability function is
reaching its maximum under the uniform distribution [1]. Two statistical constraints fundamentally limit estimation: no unbiased estimator of exists [8], and any consistent estimator may converge arbitrarily slowly [9]. A sample of size is necessary for consistent estimation at any fixed precision [10], a bound routinely violated in cryptographic practice.
The maximum-likelihood (ML) plug-in estimator , where is the empirical frequency, consistently underestimates H for . Miller [12] proposed the correction (in bits; the classical form is in nats), where counts observed symbols, yielding faster convergence but residual bias for small . Paninski [8] derived the best upper bound (BUB) estimator via polynomial approximation, proving dominance over ML, MM, and Jackknife in neural spike-train settings.
2.2. Bayesian and Empirical Bayes Estimators
Bayesian estimators replace ML frequencies with posterior expectations under Dirichlet priors parameterized by : common choices include the Jeffrey prior (, Krichevsky and Trofimov [15]), Laplace smoothing (, Holste et al. [16]), the Schürmann–Grassberger prior (, Schürmann and Grassberger [24]), and the minimax prior (, Trybula [17]). The NSB estimator Nemenman et al. [18] integrates over a calibrated prior, achieving strong performance for general distributions.
Hausser and Strimmer [22] showed that SHR, which shrinks ML frequencies toward the uniform distribution, is an empirical Bayes estimator that selects its shrinkage intensity from the data. This self-tuning property gives SHR a clear advantage in the undersample regime and identifies it as the best single estimator for byte sequences in Contreras Rodríguez et al. [2].
2.3. Coverage-Based Methods
The Chao–Shen (CS) estimator [19] combines Good–Turing coverage correction with Horvitz–Thompson weighting to account for unseen symbols. Its coverage estimate , where is the singleton count, is a key element of the FoF representation that forms the core of our occupancy-profile regime descriptor. The Chao–Wang–Jost (CWJ) estimator [20] extends the species-accumulation framework, while the Unseen estimator [10] is designed for the general non-uniform undersample regime.
2.4. Mixture, Ensemble, and Adaptive Estimation
Combining predictors is a classical idea: stacked generalization [25], stacked regressions [26], and the super learner framework with its oracle guarantees [27] all construct weighted combinations of base procedures. Entropy estimation already contains an implicit form of model averaging: the NSB estimator [18,28] integrates the Dirichlet family of Bayesian estimators against a hyperprior, which can be read as a continuous mixture over the parameter a with data-dependent weights. ASED can be viewed as a discrete, learned counterpart to this construction: instead of averaging over a single parametric family under a fixed hyperprior, it combines six heterogeneous estimators with weights produced by a function fitted offline. To the best of our knowledge, a learned convex mixture of heterogeneous entropy estimators, conditioned on sample statistics, has not been studied in the short-sample cryptographic setting. The key obstacle for classical stacking in this setting is the absence of a validation set: with a single observed sample, cross-validation would consume the data needed for estimation. ASED sidesteps the obstacle by fitting the mixing function offline on synthetic data, with the scope implications discussed in Section 3 and Section 5.
Weighting several procedures by a learned, covariate-dependent gate is also the defining structure of mixtures of experts [29,30] and of feature-weighted stacking [31], and the statistical theory of aggregation supplies rates for combining a finite dictionary of procedures [32,33]. ASED is an instance of that structure, and we do not claim that conditional mixing is new; three features distinguish the present setting. The gate reads a label-invariant summary of the same sample the base procedures consume, rather than an exogenous covariate, so weights and base estimates are statistically dependent, and aggregation results that assume independent weights or sample splitting do not transfer directly. The dictionary consists of estimators of a single functional rather than predictors of an observable response, so there is no held-out label against which to fit the gate; hence, we substitute offline synthetic supervision with a known target. Furthermore, the operating regime is a single short sample, where cross-validated stacking has nothing to split. The novelty claim is accordingly narrow: learned sample-conditional convex aggregation for short-sample entropy estimation under a uniform marginal null.
2.5. Shannon Entropy, Min-Entropy, and Cryptographic Practice
Operational standards distinguish two figures of merit. Certification of physical and non-physical entropy sources under SP 800-90B [6] and AIS 20/31 [23] is based on min-entropy , estimated by dedicated procedures (most-common-value, collision, Markov, and compression estimators), because bounds the success probability of an optimal guessing adversary and is the conservative choice for security claims. Shannon entropy, by contrast, is the established figure of merit for comparative design diagnostics and uniformity analysis: NIST SP 800-22 [4] and TestU01 [5] evaluate statistics on short blocks, entropy appears among the information-theoretic criteria of S-box design evaluation [7], and keystream analyses quantify the gap to over initial output windows [3]. ASED targets this second, Shannon-entropy setting; it is not a min-entropy estimator and makes no certification claims.
3. The ASED Model
3.1. Occupancy-Profile Feature Extraction
Let be an observed sample of size n drawn from alphabet of size k. Denote the count of symbol x. The frequency-of-frequencies (FoF) sequence is
Lemma 1
(Representation). Every plug-in entropy estimator that is invariant under relabeling of the alphabet, i.e., that depends on the data only through the multiset of counts , is a function of the FoF vector .
Proof.
A label-invariant estimator is a symmetric function of the multiset , and a multiset of non-negative integers is fully characterized by its empirical distribution . □
All six base estimators used below are label-invariant, so by Lemma 1, the complete FoF vector is a lossless representation for the purpose of combining them. The 37-dimensional descriptor of Equation (3) is a truncated compression of that vector: it retains and discards for , so for , it does not determine that the base estimates already depend on the discarded coordinates, and it is not information-preserving in general. Its adequacy is an empirical matter, supported by the truncation and ablation experiments of Section 4.7 and Section 4.9; the theorems below are unaffected, since they are stated conditionally on the descriptor actually used. (We do not claim statistical sufficiency of the FoF vector in the Fisherian sense; Lemma 1 is the purely combinatorial statement needed here.) We define the occupancy-profile feature vector as
where is the Good–Turing coverage estimate, and is the fraction of the alphabet observed. The first five entries encode the macroscopic regime: captures the fundamental undersample ratio; C and encode singleton density; appears in second-order coverage corrections; and reflects alphabet coverage. The trailing 32 entries capture higher-order regularity of the sample. Note that C and are affinely dependent by construction (), so the map carries 36 functionally independent coordinates; we retain both because the redundant parameterization is harmless for a standardized linear input layer and keeps the coverage and singleton readings explicit. Section 4.7 shows that the first five coordinates retain all measurable performance in the i.i.d. uniform and near-uniform settings evaluated here, so a practitioner may use the five-feature variant (“ASED-lite”) without measurable loss under this experimental protocol. The feature vector is computable in time.
3.2. Weight Mapper Architecture
The weight mapper is an MLP mapping occupancy-profile features to the probability simplex over base estimators:
with hidden dimension 64, activation , feature standardization by the training mean and standard deviation , and a softmax output ensuring , .
Base Estimator Set
The six base estimators in are selected for complementary strengths across regimes:
- SHR: empirical Bayes shrinkage toward uniform [22]; the strongest single estimator at small .
- MM: Miller–Madow correction [12]; asymptotically efficient for .
- CS: Chao–Shen coverage correction [19]; robust to unseen symbols.
- LAP: Laplace smoothing () [16]; low variance at very small n.
- MIN: minimax Dirichlet prior () [17]; strong for tiny alphabets.
- SG: Schürmann–Grassberger prior () [24]; near-ML behavior with mild regularization.
Two earlier candidates for the base set were excluded for reproducibility reasons: BUB [8] requires solving a regularized polynomial-approximation problem, and the centered Dirichlet mixture requires numerical hyperprior integration, so neither admits the closed-form, dependency-free reference implementation this paper commits to; MIN and SG were included instead as closed-form Dirichlet members covering, respectively, small-alphabet strength and near-ML behavior.
Total weight-matrix parameter count: (plus biases). Table 1 consolidates all architecture and training hyperparameters. The reference implementation of ASED and of the six base estimators is released as Supplementary Materials (Code S1), and the trained weight-mapper coefficients together with the corresponding feature standardization constants and are released as Supplementary Materials (Data S1).
Table 1.
ASED architecture and training hyperparameters.
An immediate structural property is worth recording: because the output of ASED is a convex combination of the base estimates, it always lies in the interval . Convexity therefore anchors the output to the range of the base estimators, but it does not by itself rule out a constant-valued adaptive combination: whenever lies in that interval, a sample-dependent could return for every sample. Whether the fitted model behaves that way is an empirical question, addressed by the power experiments of Section 4.5.
3.3. Offline Training
The weight mapper is trained on synthetic samples drawn from for across the same 12 sample sizes used in evaluation. For each configuration , we generate 1000 samples, extract , evaluate all base estimators, and minimize a regime-balanced squared error with an annealed entropy bonus:
where equalizes the contribution of each regime, is the Shannon entropy of the weight vector, and anneals linearly from 1 to 0 over the first 30% of training. Both terms address a concrete failure mode: under the unweighted loss, the softmax saturates early on the globally strong SHR corner of the simplex and never recovers the large relative improvements available at the smallest sample sizes. Optimization uses Adam with learning rate , cosine annealing over 600 epochs, and batch size 256; training takes under one minute on a single CPU core. Across three independent initialization and shuffling seeds, the resulting integrated MSE on a common evaluation set agrees to three significant figures, so the balanced objective also removes seed sensitivity. No sample used for training is reused during evaluation, eliminating sample-level leakage; configuration-level generalization is tested directly in Section 4.6. We state one methodological caveat plainly. The design choices for the feature map, the tail length, the hidden widths, the regime weighting, and the entropy bonus were made while inspecting results on the same evaluation protocol, so no independent model-selection split separates development from final assessment. Sample-level leakage is excluded, but experimenter-level overfitting is not, and a locked-test protocol (train/validate/final test) would be required to exclude it. The held-out configurations of Section 4.6 and the alternative sources of Section 4.9 and Section 4.10, which were never used for tuning, are the closest substitutes available here. Because a model-selection split cannot be imposed retrospectively, we state explicitly how the two bodies of evidence in this paper relate to one another, and we adopt that relation as the organizing principle of the revised manuscript. The pipeline described in this section is exploratory: its design was fixed while observing the evaluation protocol of Section 4, and every number it produces is a development result. The evaluation of Section 4.10.3 is confirmatory: the architecture and the feature map were fixed in advance to the specification of Table 1; training, validation, and test data are drawn from three independent seed streams that are never mixed; and the test stream was evaluated once and used for nothing else. The confirmatory evaluation reproduces the headline reduction ( against the exploratory pipeline’s ), the concentration of the advantage in , and the instability. Where the two disagree, the confirmatory figures are the ones on which the reader should rely. What the confirmatory evaluation does not do is retroactively lock the exploratory pipeline, and we make no such claim: the exploratory results are reported because they carry the multi-estimator comparison and the ablations, not because they are protected from experimenter-level overfitting.
Remark 1.
Two caveats delimit what the offline strategy does and does not guarantee. First, the weights are functions of the FoF vector, which also determines every base estimate (Lemma 1); weights and estimates are therefore statistically dependent, and no oracle or regret guarantee follows from the offline construction alone; deriving one is the main theoretical problem left open by this work. Second, the loss (5) is minimized under the uniform null, so the fitted mapping is calibrated for the testing problem “ versus nearby alternatives” rather than for general-purpose entropy estimation. Nothing in the construction prevents a sample-dependent combination from returning on every sample; Section 4.5 supplies the behavioral evidence that the fitted model does not.
3.4. Theoretical Analysis
Although the offline construction yields no guarantee by itself (Remark 1), the mixture admits an oracle inequality that identifies exactly what the weight mapper must learn and how errors in learning it propagate. Write for the error of base estimator j, and let
be the conditional second-moment matrix of the base errors given the occupancy-profile features. For any measurable weight function , the risk of the induced mixture is , and we call the sample-conditional oracle, with risk .
Theorem 1
(Oracle inequality for conditional mixtures). Write , let , and define the suboptimality gaps , which are non-negative and vanish on the support of . Then, for every measurable ,
In particular, if has full support for almost every φ, then , and the excess risk is quadratic in the approximation error: .
Proof.
Fix , and set , so that . For the first term, Hölder gives . For the second, ; hence, , the last equality because whenever , by the Karush–Kuhn–Tucker conditions for minimizing the convex quadratic over ; the same conditions give for all j, which also yields the left-hand inequality. Taking expectations completes the proof. □
Theorem 1 converts the estimation problem into a regression problem: the excess risk of ASED over the sample-conditional oracle is controlled by the error with which the network approximates , scaled by the local geometry of . Two features matter for what follows. The multiplier is rather than the cruder , which it never exceeds; and when the conditional oracle has full support, the dependence on the approximation error is quadratic, so moderate weight errors are second-order. Whether that full-support condition holds in the small-sample regimes studied here is not established empirically: the softmax output assigns strictly positive weights by construction, which says nothing about the geometry of , and we do not estimate directly. The bound also explains why the loss (5) must be balanced across regimes: varies by orders of magnitude between and , so an unweighted objective concentrates all effort where the multiplier is largest. One limit of Theorem 1 should be stated rather than left to inference, since it is easy to over-read. The inequality bounds the excess risk of a mixture in terms of the error of its own weights; it does not assert that the offline fit of this section makes that error small, and no bound on for the trained mapper is established anywhere in this paper. The theorem therefore identifies what would have to be shown for a performance guarantee to follow, and the empirical results of Section 4 are evidence about the unproved quantity rather than a consequence of the proved one. No theoretical result in this section should be read as a guarantee for the deployed estimator.
The next result gives a checkable condition under which weights conditioned on the regime alone are useless; Section 4 verifies empirically at which sample sizes the condition actually holds and reports one size at which it fails.
Corollary 1
(A covariance condition excluding fixed-mixture improvement). Let , and suppose
equivalently, whenever so that the correlation is defined, . Then, : no convex combination with fixed weights improves on the best single estimator. For , condition (8) reduces to with .
Proof.
The objective is convex on because M is positive semidefinite, so the vertex is a global minimizer if and only if the first-order condition holds for every . Since , this is equivalent to for all i, i.e., to (8). Writing gives the correlation form. □
Condition (8) has a transparent reading: the column of M belonging to the best estimator dominates its own diagonal entry, so every other estimator has a sufficiently large positive error cross-moment with the best one, and adding it to the mixture cannot cancel error on average. Its scope is correspondingly narrow, and we state the boundary explicitly. Corollary 1 concerns fixed mixtures within a regime: where the condition holds, no regime-indexed convex combination of the base set improves on its best member, and nothing further follows. It does not establish that sample-conditional weighting is superior, it does not show that the sample-conditional oracle strictly dominates the per-regime one, and it says nothing about the fitted mapper. Its role is to remove one competing explanation for the observed gain, that a fixed mixture would have sufficed, and not to supply a positive argument for the method.
Condition (8) concerns fixed mixtures within a regime, so the relevant matrix is the regime-conditional second-moment matrix
obtained from by averaging over samples of a given size, and it is not pointwise to which the corollary is applied here. The condition involves every base estimator, so we verify it directly. Table 2 reports the ratios estimated from 4000 replications at each critical sample size for .
Since the entry equals one by construction, the binding quantity is the smallest off-diagonal ratio , which we report with a percentile bootstrap over the 4000 replications ( resamples of the replication index). At , it is (SHR) with a lower bound of , and at , it is (LAP) with a lower bound of : the condition is satisfied and certified at the level. At , the plug-in estimate satisfies the condition at (SHR), but its bootstrap lower bound is , so the empirical certification is inconclusive. At , the condition fails: the smallest off-diagonal ratio is (LAP), with lower bound . Solving numerically there confirms and quantifies the failure: the optimal fixed mixture puts weight on (SHR, LAP) and attains against for the best single estimator, a improvement. ASED attains at the same size, so sample-conditional weighting still helps substantially, but at , it improves on an achievable fixed mixture rather than on an unbeatable baseline. We state this explicitly because is one of the sizes where ASED’s advantage is largest, and the necessity argument does not cover it.
The ablation of Section 4.7 confirms both halves of this prediction empirically.
The results so far concern the best achievable mixture. For deployment, one also wants guarantees that do not depend on the distribution used to train the weights. Theorem 2 establishes consistency for fixed-alphabet i.i.d. full-support sources; Propositions 1 and 2 then give non-asymptotic safety bounds, the first of which transmits consistency only when the base estimators are themselves consistent under the source considered.
Theorem 2
(Consistency and when asymptotic equivalence holds). Fix k, and let be i.i.d. from P with for all . Then, for every weight function ,
- (i)
- (Consistency) ;
- (ii)
- (Expansion) writing for the weight on the minimax-prior estimator,
- (iii)
- (Asymptotic normality) if in addition and , then , matching the plug-in estimator to first order.
Proof.
Each Dirichlet member satisfies with and uniform. For the bounded-pseudocount Dirichlet members of the ASED base set, LAP and SG, the pseudocount is bounded, so (the same observation applies to the JEF comparison estimator, which is not a member of the mixture); since is Lipschitz on compact subsets of and almost surely, these estimators differ from by . The same holds for MM, whose correction is exactly , and for CS, because eventually almost surely. For SHR, two cases must be separated. If , then , so the denominator of grows linearly while the numerator stays bounded, giving and, hence, . If , then , and need not vanish, but the conclusion holds for a different reason. Writing and expanding both entropies around where the gradient of H is constant on the simplex, so the linear terms cancel gives
and since , both quadratic forms are . For MIN, however, gives , and a first-order Taylor expansion of H at yields , the sign following from together with , whose leading term does not vanish unless P is uniform. Hence, in general, and whenever the limiting linear term is non-zero, which gives (ii) because bounds the contribution of every other member by . Statement (i) follows since every base estimator is consistent for fixed k and a convex combination of consistent estimators is consistent. Statement (iii) follows from (ii), the delta method applied to , and Slutsky’s lemma. □
Remark 2
(Sharpness). Part (ii) cannot be improved to without a condition on : taking makes ASED equal to MIN. Simulation confirms the rate: for a fixed non-uniform P on , the mean of multiplied by is nearly constant ( to for to ), whereas the same quantity for LAP and SHR multiplied by n is constant ( and ). Two caveats govern the applicability of part (iii). First, it requires , which fails under the uniform null, where and ; part (iii) therefore does not apply to the byte experiments of Section 4, whose source is uniform by design, and the degenerate case would require a second-order analysis that we do not carry out. Second, on the tested byte grid, the fitted MIN weights are empirically negligible, which is consistent with, but does not verify, the asymptotic vanishing-weight condition: small weights on a finite grid of n do not establish . For bits, MIN receives substantial weight and the condition plainly does not hold.
Part (i) is the consistency statement the construction needs: no assumption on how the weights were obtained enters the argument, only that they live in the simplex. We do not offer it as a substantive result. Every base estimator is consistent at fixed k, and a convex combination of consistent estimators is consistent by an elementary argument; part (i) is therefore a sanity property, confirming that the mixture cannot destroy something its ingredients already have, rather than evidence in favor of sample-conditional weighting. It is recorded because the weights here are learned and could in principle be arbitrary, so ruling out pathological behavior of an arbitrary simplex-valued rule is worth doing once, cheaply. Parts (ii)–(iii) delimit the method from the other side. Under non-uniform P, positive asymptotic variance, and a vanishing MIN weight, ASED has the same first-order asymptotic distribution as the plug-in estimator; thus, in that regime, its value lies entirely in the short-sample behavior, where the correction terms dominate, precisely the regime this paper studies. We make no claim about second-order or constant-factor behavior and none about the degenerate uniform case where the first-order expansion vanishes. The following non-asymptotic counterpart is elementary two lines of Jensen, but it is what rules out arbitrary behavior at finite n on sources far from the training distribution. It bounds the mixture from above and is not a guarantee relative to the worst base estimator.
Proposition 1
(Distribution-free safety upper bound). Let P be any source distribution on with entropy H. For every weight function ,
Consequently, (i) if every base estimator is consistent for H under P, so is ASED; (ii) if for every j, then ; and (iii) both conclusions hold irrespective of the distribution on which the weights were trained.
Proof.
Conditionally on the sample, with . Convexity of (Jensen) gives pointwise, and taking expectations under P yields the three displayed inequalities. Statements (i)–(iii) follow because no property of beyond membership in the simplex is used. □
Remark 3
(Why the maximum sits inside the expectation). For weights fixed in advance, the first bound would read , without the factor M. That simplification is unavailable here: is a function of the same sample as the base estimates, so cannot be split into , and the maximum must remain inside the expectation, where . The factor M is thus the price of data-dependent weighting; it is benign for the asymptotic conclusions (i)–(ii), which is all we claim.
Proposition 1 is a worst-case guarantee. Conditioning on before taking expectations sharpens it into a bound relative to the best base estimator, at the price of an explicit dependence on how much weight the mapper misplaces.
Proposition 2
(Bound relative to the conditionally best estimator). For each φ, let , and let be the weight the mixture places away from it. Then,
where is the risk of the conditionally optimal selection rule.
Proof.
Conditionally on the weights are constants, so Jensen and then linearity of the conditional expectation give
the second step splitting the sum at and bounding the remaining weights, which total . Taking expectations over gives the display. Conditioning first is essential: the pointwise bound would lead to , which exceeds . Finally, , because selecting the conditionally best estimator is at least as good as committing to any fixed one. □
Proposition 2 describes two limiting behaviors. If the fitted mapper places asymptotically all its weight on the conditionally best estimator (), its risk approaches ; when it spreads weight, the excess is controlled by times the worst conditional risk, and Proposition 1 is recovered as the crude case . We do not estimate in the present experiments, and two empirical readings must therefore be avoided. The concentration on SHR observed for shows convergence towards SHR but does not establish that SHR is the conditionally best estimator for almost every feature vector, which would require estimating the conditional risks . Likewise, ASED’s improvement over the best fixed estimator at does not by itself imply improvement over , since makes the conditional selector the stronger benchmark.
Proposition 1 is what makes offline calibration safe to deploy, provided it is read correctly. For every individual sample, the squared error of the convex mixture cannot exceed the largest squared error among its base estimators. In expectation, this yields the displayed bound and nothing stronger: since , the proposition does not guarantee that the MSE of ASED is below the MSE of the worst base estimator. It also predicts the shape of the failure documented in Section 4.9 on a strongly skewed source. ASED degrades to the level of the estimators it leans on, not below it, and this is what we observe. Estimating the first bound directly over 2000 replications at : at Zipf , the measured against , so the bound is nearly saturated, whereas under the uniform null, the same comparison reads against : loose by four orders of magnitude. The bound is therefore informative exactly where the method is weakest and uninformative where it is strongest, which is the expected behavior of a worst-case guarantee.
Finally, an idealized version of the offline fit admits a standard uniform-convergence guarantee. Because it bounds an empirical-risk minimizer related to, but distinct from, the procedure actually run, and because the resulting numerical bound is vacuous at our training size, we state it and evaluate its complexity term in Appendix B, recording here only the consequence.
Remark 4
(Approximation, generalization, optimization, and design mismatch). Combining Theorem 1 with Proposition A1 of Appendix B,
where attains . The first term is an approximation-theoretic question about how well a small network represents the conditional oracle on the FoF domain; the second is controlled by the training budget. Only the first two terms are analyzed here. The generalization term is the idealized PAC quantity of Appendix B, which bounds an empirical-risk minimizer related to, but distinct from, the deployed training procedure; the optimization error of Adam on a non-convex objective and the design-mismatch term from stratified sampling with a regime-weighted loss are both uncontrolled. Appendix B evaluates the complexity term numerically, by a norm-based argument over a data-dependent class and by Monte Carlo; neither is a valid uniform bound for the deployed model.
3.5. Adaptive Inference
Algorithm 1 formalizes inference. Total complexity is , dominated by the count computation and matching the complexity of any plug-in estimator.
| Algorithm 1 ASED Inference | |
| Require: Sample , alphabet size k, trained weights | |
| Ensure: Entropy estimate | |
| 1: Compute for all | |
| 2: Compute FoF sequence and feature vector | // |
| 3: Evaluate base estimators: for | // each |
| 4: | // — fixed network |
| 5: | |
| 6: return | |
4. Experimental Evaluation
4.1. Protocol
The null-protocol design follows Contreras Rodríguez et al. [2]. For each combination of alphabet size and sample size , we generate 1000 independent samples drawn uniformly at random with a fixed-seed NumPy generator, disjoint from the training draws. We re-implement eight baseline estimators spanning the plug-in, corrected, Bayesian, coverage-based, and shrinkage families of that comparison (ML, MM, LAP, JEF, SG, MIN, CS, SHR) and the general-purpose NSB estimator [18], computed by numerical integration of the Wolpert–Wolf posterior mean entropy against the NSB prior over a 60-point grid, together with ASED and the constant control . Performance is measured by
and the integrated MSE summarizes performance across the 12 sample sizes. Uncertainty is quantified by percentile bootstrap intervals over the 1000 replications (); the bootstrap is used for intervals only, and the ASED–SHR integrated-MSE comparison is supplemented by a paired randomization test under the sharp null of pairwise exchangeability. Tables report representative columns. One rule of run hygiene, prompted by the inconsistency documented in Section 4.11, is now enforced throughout and stated here so that every table below can be read against it: each table and each aggregate reported in this paper is computed from a single evaluation run, and figures from different runs are never combined within a single statement. Where a quantity was previously taken from a run other than the one behind its own table, it has been withdrawn and replaced either by the value implied by that table itself or by a value from a single internally consistent run, identified as such at the point of use. Regime splits use the disjoint partition , , throughout; the overlapping definitions used in the previous version are retired.
4.2. Results: Byte Sequences ()
Table 3 reports for representative estimators. ASED achieves a bias of bits at , the lowest of all estimators at the smallest sample size, versus for SHR. At the mean weights are approximately on SHR and on LAP, moving to at and back toward SHR from onward (Figure 1); for , ASED and SHR coincide, confirming that SHR is essentially optimal within the evaluated comparison set in the large-sample regime and that the mixture does not degrade it.
Table 3.
Absolute bias for byte samples (, bits). Best value per column in bold.
Figure 1.
Mean adaptive mixture weights for byte sequences, by sample size.
Table 4 reports MSE in scientific notation, including the constant control. Full-precision per-size results underlying this and every subsequent table through the cryptographic-source results of Section 4.10 are released as Supplementary Materials (Data S2). The integrated MSE for ASED is with bootstrap 95% CI , against for SHR with CI : a 95.4% reduction (; percentages and ratios are computed at full precision: versus , a reduction of , and only the presented result is rounded), with disjoint intervals. A paired difference is the more direct summary. The value consistent with Table 4 is , obtained from the full-precision integrated values of that table and from no other source. The figure previously reported here, with interval , was computed on a different evaluation run; it has been withdrawn rather than carried forward with a caveat, and its cause is documented in Section 4.11. The paired bootstrap interval this paper reports is therefore the one produced by the single internally consistent run detailed in Section 4.11, with a interval , which excludes zero, and it is quoted there together with the run that produced it rather than alongside the values tabulated here. Because ordinary bootstrap resamples are generated around the observed empirical distribution rather than under the null, the frequency with which they reverse the ordering is a stability diagnostic and not a calibrated p-value; we therefore report a paired randomization test instead. ASED and SHR are evaluated on the same samples, so the two labels may be swapped within each paired replication. Under the sharp null of pairwise exchangeability of the estimator errors, the paired randomization test yielded no assignment as extreme as the observed statistic over independent sign-flip assignments (; with the plus-one correction), the largest randomized difference being against an observed . Exchangeability is a stronger hypothesis than equality of integrated MSE; two estimators can share an MSE without exchangeable error distributions so the p-value refers to that sharp null; the disjoint bootstrap intervals reported above are the evidence bearing directly on the MSE comparison. The next best baseline, LAP (), is worse than ASED; NSB, evaluated under the same full protocol, reaches , dominated by the deep undersample regime ( at ) where its general-purpose prior cannot exploit the uniform target. The CONST row makes the protocol’s blind spot explicit: the constant rule achieves zero error by construction under the null, which is precisely why Section 4.5 and Section 4.6 are part of the evaluation.
Table 4.
MSE for byte samples (), scientific notation. Integrated MSE over all 12 sample sizes, with bootstrap 95% CIs for ASED and SHR. CONST is a degenerate control, not an estimator.
Figure 2 shows mean estimates across sample sizes; ASED tracks closely even at the smallest n, while the baselines diverge for . Figure 3 reports the gap against the per-regime best-single-estimator benchmark (the best single estimator chosen with knowledge of the regime; not to be confused with the sample-conditional oracle ): the gap is negative for , as ASED is more accurate than any evaluated single estimator chosen with knowledge of the regime, and zero thereafter. The mechanism is sample-conditional weighting rather than estimator selection: at , the errors of SHR and LAP are strongly positively correlated (), so no fixed convex combination of the base set could beat the better endpoint; the gain comes from weights that covary with the FoF features predicting the plug-in error, an implicit learned bias correction under the null. This behavior is consistent with null calibration and is what makes the power analysis below indispensable rather than optional. Figure 4 summarizes the two preceding comparisons in a single head-to-head panel, plotting integrated MSE for all evaluated estimators on a log scale across byte sequences: ASED (red) tracks the lower envelope of the base set throughout the sample-size range.
Figure 2.
Mean entropy estimate versus sample size for byte sequences (, , dotted line). ASED in red.
Figure 3.
Benchmark gap over the eight fixed estimators, per sample size. Green: ASED beats the per-regime best-single-estimator benchmark; the gap vanishes for .
Figure 4.
Integrated MSE for byte sequences (log scale). ASED in red.
4.3. Results: Bit Sequences ()
Table 5 reports the complete ranking for bits. ASED attains integrated MSE with bootstrap interval , a improvement over SHR () and better than LAP (), ranking second behind the minimax-prior estimator MIN (), whose interval does not overlap ASED’s. Bit-sequence numbers quoted elsewhere in this paper come from auxiliary runs with fewer replications or from separately retrained mappers, reported in the alphabet-generalization and extended-comparison tables introduced in Section 4.6 and Section 4.8; the spread across those runs, roughly to , reflects both Monte Carlo error at 200–400 replications and genuine retraining variability at , where the FoF descriptor is nearly degenerate. This instability is itself a limitation of the method for tiny alphabets, and Table 5 is the locked reference. The residual gap to MIN is a genuine limitation: MIN belongs to ASED’s base set, so an additional improvement was available within the mixture’s hypothesis class that the fitted mapper only partially realized. The plausible cause is feature degeneracy: for , the FoF histogram has at most two non-zero coordinates, so the occupancy features carry little regime information beyond the count imbalance. Augmenting the descriptor with low-k statistics is a natural remedy, which we leave to future work. We therefore make no performance claim for bit sequences, and the applicability statements of Section 5 are confined to bytes. Three facts jointly support that restriction. ASED does not outperform MIN, an estimator in its own base set, so the mixture fails to achieve an improvement available within its hypothesis class. The run-to-run spread across the auxiliary protocols, roughly to in integrated MSE, is of the same order as the margins separating the estimators in Table 5, so the ranking below the top position is not stable. Furthermore, the improvement over SHR quoted elsewhere is a single-run figure at one alphabet, not a property of the method. At , the honest summary is that ASED is competitive with, but not better than, the best classical estimator available.
Table 5.
Integrated MSE ranking for bit sequences (), from a single locked evaluation with 4000 replications per sample size and percentile bootstrap 95% intervals. This is the locked reference for the main bit-sequence comparison; the auxiliary protocols reported in Section 4.6 and Section 4.8 should not be read as replications of it. SG and JEF are tied at rank 6 to the reported precision (both ) and share that rank rather than being assigned ranks 6 and 7.
4.4. Sensitivity of the Headline Number to the Aggregation Scheme
Integrated MSE weights the twelve sample sizes equally on a linear scale, which lets the smallest n dominate. Since the headline reduction depends on that choice, we report it under four schemes (bytes, ASED versus SHR): sum over the grid, ; arithmetic mean, ; trapezoidal integration uniform in , ; and the least favorable, the mean over n of the per-size ratio , which is , i.e., a reduction of only . Split by regime, the picture is unambiguous, and it is now reported from a single internally consistent run against a disjoint partition, replacing the mixed-run figures that appeared here in the previous version. The earlier regime definitions. Same below , , and overlapped at , 256, and 512, so their sums could not add to the integrated total, whatever run they came from; they are replaced throughout the paper by the disjoint partition , , . Against that partition, and from the single run detailed in Section 4.11, the regime accounts for a reduction, while the two larger-n regimes account for none ( and , respectively; i.e., ASED is marginally the worse of the two there). The SHR subset sum previously printed in this paragraph, , was the integrated total of a second run mistranscribed as a subset sum, an arithmetic impossibility, since it exceeded the integrated total of in Table 4 and it is withdrawn rather than re-estimated, for the reasons set out in Section 4.11. All of the advantage lies in the undersampled regime, and the large headline figure reflects the fact that absolute errors there are orders of magnitude larger than elsewhere. Both aggregations are quoted together wherever the headline appears in this paper, and neither is presented as the reduction. Which one is appropriate is an application question rather than a statistical one. For a screening loop that spends its entire sample budget at , the S-box search and initialization-analysis settings that motivate the paper, the summed or averaged figure of measures what is actually gained, because that is where essentially all of the error lives. For a practitioner who applies the estimator across a range of window sizes and asks what accuracy to expect at a typical one, the mean of per-size ratios, , is the relevant summary, because it weights every sample size equally. We quote the summed figure first because the motivating application is the first case, and we state the second immediately alongside it; a reader working with long windows should use the figure, and a reader with should expect no gain at all.
4.5. Sensitivity to Departures from Uniformity
Low MSE under the null is necessary but not sufficient: an estimator calibrated to the null could in principle buy accuracy at the price of insensitivity to deviations, and the constant rule is the limiting case. We therefore measure detection power directly. For each estimator, we calibrate a one-sided test of at from 2000 fresh null samples (reject when falls below the empirical 5%-quantile) and evaluate the rejection rate over 1000 samples from alternatives of known entropy: biased bits and doped bytes with one symbol carrying mass and the rest uniform.
Because each estimator is calibrated to its own empirical null threshold, this comparison measures induced test behavior rather than raw entropy-estimation accuracy under alternatives; the latter is reported separately as MSE against the true entropy. Three findings stand out (Table 6, Figure 5 and Figure 6). First, ASED’s power differs only slightly from SHR’s at every tested alternative, alphabet, and sample size, and ASED dominates LAP for bytes at : the null-calibrated mixture retains sensitivity at the operating points that matter. Second, the constant control has zero power everywhere, exhibiting the confound concretely and separating ASED from it behaviorally. Third, for bits, all genuine estimators have identical power: for , each classical base estimator is a monotone function of the count imbalance, so equal- tests share the same rejection region power comparisons, which are only informative for larger alphabets. Monotonicity does not follow automatically for ASED, whose weights vary with the sample, so we verified it exhaustively rather than assuming it: for each n in the grid, we evaluated ASED at all count vectors and confirmed that the output is non-increasing in , with no violation at any sample size. This is a verified property of the fitted model, not a structural guarantee. The same exhaustive check applied to a deliberately under-trained mapper does exhibit violations, so it should be re-verified for any retrained model. Measured against the true entropy of the alternatives (Table 7), ASED’s MSE is at worst equal to SHR’s near the null and lower at moderate deviations for ; only at the largest deviation tested (, ), where both tests already reject with probability one, is ASED’s MSE inflated relative to SHR’s ( versus ), reflecting null calibration far from its design regime.
Table 6.
Detection power at for doped-byte alternatives, . CONST is the constant control. Empirical size at the null is ≈0.05 for all genuine estimators.
Figure 5.
Power curves for doped-byte alternatives at , (dotted line). The constant control (grey) has zero power throughout.
Figure 6.
Power curves for biased-bit alternatives. All genuine estimators coincide (monotone functions of the count imbalance).
Table 7.
MSE against the true entropy of doped-byte alternatives, (the regime where the estimators differ).
The power analysis above covers , but the MSE gain is concentrated at , so the objection that ASED might buy accuracy there by becoming nearly constant must be settled in that regime specifically. Table 8 repeats the analysis at with 4000 null and 2000 alternative replications per cell.
Table 8.
Detection power at for doped-byte alternatives at the small sample sizes, where the MSE gain is concentrated (). Half-widths of 95% binomial intervals are at most ; a paired non-inferiority test at margin is reported in the text.
At these sample sizes, every genuine estimator induces essentially the same test. Absence of a significant difference would not by itself establish equivalence, so we ran a separate paired non-inferiority experiment. It uses its own draws, 4000 null and 4000 alternative replications per cell, independent of the 2000 alternative replications behind Table 8 and a fully nested bootstrap (): each resample draws the null replications in pairs, recomputes the critical value of each estimator from the resampled nulls, then draws the alternative replications in pairs and recomputes both powers, so the intervals incorporate the uncertainty of the empirically calibrated thresholds rather than conditioning on them. The margin was chosen after inspecting the power estimates, so this is a post hoc rather than pre-specified analysis. Pointwise two-sided intervals for lie strictly above in all nine cells; the largest deficit is (CI at , ) and the largest surplus at . Three cells at have intervals excluding zero, so a small but statistically detectable deficit exists there. Because nine pointwise intervals do not license a simultaneous statement, we also bootstrapped the minimum of across the nine cells: its quantile is , above , so non-inferiority also holds simultaneously at the margin. This is a bootstrap lower percentile for the minimum over these nine pre-defined cells, not a correction transferable to any further set of alternatives. ASED is therefore non-inferior to SHR at the post hoc margin at every tested operating point, pointwise and simultaneously; this is non-inferiority, not equality. The constant control again has zero power. Non-inferiority at the margin therefore holds in the regime where the MSE gain occurs, which is what the objection required. It is not equality: three operating points at showed small statistically detectable deficits (largest , CI ), all contained within the post hoc non-inferiority margin. Absolute power is low at for all estimators; detecting a deviation from eight observations is close to impossible, so the practical reading is that ASED gives more concentrated point estimates within the post hoc non-inferiority margin, not that it enables detection where none was available. Two qualifications belong with this conclusion rather than after it, because they change how far it can be carried. First, the margin was selected after inspecting the power estimates; the analysis is post hoc throughout and is described as such in the abstract, the conclusions, and here. Second, the comparison is at the matched nominal level, not at the matched achieved level. Section 4.10.2 measures the achieved Type-I error of the strict convention as at and at and shows that the tie rule alone moves reported power by at and at , an ambiguity an order of magnitude larger than itself. The non-inferiority statement at and is therefore exploratory and is not offered as a finding; at , where the tie-induced gap falls to and below and the achieved level is close to nominal, it stands as reported. Restoring exact size control at the two smallest sizes requires a randomized or otherwise exact test or a comparison performed at matched achieved level rather than matched nominal level. We have implemented neither, and in their absence, we do not claim non-inferiority at and .
Third, Table 8 and Table 9 disagree at , and the resolution is not that one of them is mistaken. At , the first reports for ASED against for SHR; the second reports against : the ordering reverses, and the gap exceeds the binomial half-width at 4000 replications. We examined the null distribution directly rather than adjudicating between the two runs, and the discrepancy turns out to be a property of the testing problem at this sample size, not of either experiment.
Table 9.
Detection power at nominal : ASED and SHR versus a simulation-calibrated Pearson test and a likelihood-ratio (G) test, doped-byte alternatives, from a single paired run (4000 null and 4000 alternative replications per cell, all four detectors on the same draws). 95% Wilson intervals in brackets; best value per row in bold. The strict tie convention of Section 4.10.2 is used throughout, which is conservative at the smallest sample sizes.
Every estimator here is label-invariant, hence a function of the frequency-of-frequencies vector (Lemma 1), and at , that vector takes very few values when . Enumerating the null distribution from draws, the ASED and SHR statistics have 7 atoms at and 13 at , with identical atom probabilities in both cases: the two statistics are strictly monotone relabelings of one another on the null support, so they induce the same family of tests, and the true difference in power at any achievable level is exactly zero. At , the atom-probability multisets no longer coincide, and the two statistics separate genuinely. What differs at is only the spacing of the atom values, ASED’s cluster within bits of each other while SHR’s spread over bits, so linear interpolation of the empirical quantile lands on opposite sides of the same atom for the two estimators. Repeating the 4000-draw calibration 200 times, each estimator’s reported power at takes one of exactly two values, or about , according to which side the interpolated cutoff falls; both tables are sampling that dichotomy.
The underlying reason is that no test of size exists at these sample sizes. The achievable levels bracketing are and at and and at ; at , they are and , so the nominal level is essentially attainable there, and from , the bracket closes to and . This is the same discreteness that Section 4.10.2 measures through the tie convention, seen from the side of the critical value rather than the rejection rule, and it is why a randomized or exact test is the only way to compare power at . We therefore do not reconcile the two tables and do not claim a power difference in either direction at or : none is identifiable there. The non-inferiority conclusion stands as reported at , where the tests genuinely differ and the nominal level is attainable. One consequence should be recorded against Table 9: the explanation offered after it, that SHR separates from the , G, and ASED columns at because its shrinkage intensity is a different function of the counts, is incorrect as stated for SHR, whose null distribution is identical to ASED’s at that size; SHR separates there through the placement of the interpolated critical value alone.
Comparison with a Direct Uniformity Test
Comparisons so far are among entropy estimators; because the null is known exactly here, a fair test of whether an adaptive entropy estimator is a useful uniformity diagnostic at all requires comparing it against a direct goodness-of-fit test rather than only against other entropy estimators. As an additional, independently coded check (Section 4.10.3), we calibrated a simulation-exact Pearson chi-square statistic and a likelihood-ratio (G) statistic (both against the empirical null distribution rather than the asymptotic approximation, which is unreliable when most cells are sparsely populated) alongside ASED and SHR, all at nominal , and evaluated power on the same doped-byte alternatives used above. All four detectors see the same alternative draws, so differences are assessed by McNemar’s exact test on the discordant pairs rather than by comparing independent rejection rates. Table 9 reports the result.
The comparison does not favor either family uniformly, and the direction depends on both the sample size and which entropy estimator is considered. At and , the entropy-based detectors are at least as powerful as the direct tests at every tested alternative, significantly so in most cells. At and the smallest tested deviation (), the ordering separates: the G-test () significantly outperforms ASED (, McNemar ) but is statistically indistinguishable from SHR (, in SHR’s favor). The honest summary is therefore narrower than a claim of general superiority: the entropy-based approach is not dominated by a classical direct test at these sample sizes, but ASED specifically loses to the G-test at on the smallest deviation, and where the entropy-based family retains an advantage at that sample size, it is SHR, not ASED, that carries it.
One feature of Table 9 deserves comment because it is easy to misread as an error. At , the , G, and ASED columns are numerically identical across all three alternatives. This is not a transcription mistake: at with , essentially, every sample consists of singletons and at most a few doubletons, so all three statistics are, over the realized sample space, monotone functions of the number of distinct symbols observed, which takes only five values. The three tests therefore induce the same rejection region and cannot differ. SHR separates from them because its shrinkage intensity is a different function of the same counts. That last sentence is wrong and we correct it here rather than delete it, since the corrected version is the more interesting statement. SHR’s null distribution at has the same 13 atoms with the same probabilities as ASED’s, so the two are monotone relabelings of one another and cannot induce different tests at any achievable level. SHR separates in this table only because its atom values are spread over roughly bits, whereas ASED’s cluster within , so the interpolated empirical quantile falls on the other side of the same atom. Section 4.5 sets out the consequence: at , no power difference between these two detectors is identifiable. This degeneracy is a property of the regime rather than of any particular estimator, and it is the same phenomenon that makes the tie convention of Section 4.10.2 material.
4.6. Generalization to Unseen Configurations
To test interpolation beyond the training grid, we retrain the mapper using only the even-indexed sample sizes and evaluate on the held-out odd-indexed sizes ; separately, we evaluate the full-grid model on the alphabet , never seen in training. Table 10 shows that the advantage persists in both cases: on held-out sizes, ASED attains against for SHR (), and on the unseen alphabet, against (). The smaller margin at is expected; the regime descriptor was fit to the geometry of , but the direction of the effect is preserved.
Table 10.
Generalization: integrated MSE on configurations excluded from training.
To test interpolation across alphabets directly, we retrained the mapper on (400 replications per configuration) and evaluated it on the held-out alphabets , alongside the two-alphabet model of the main protocol (Table 11). Both models interpolate: on , they are indistinguishable ( versus ), and both improve on SHR by an order of magnitude, while on , the multi-alphabet model is better. All of these are interpolation tests, since the evaluated alphabets lie between the trained ones. Extrapolating outside the trained range gives a sharper picture: applying the main model unchanged to and gives integrated MSE and against and for SHR ( and ), while at and , the advantage almost disappears ( and ). Upward extrapolation was successful at the two tested alphabet sizes, whereas the gain nearly vanished at and , consistent with the small-alphabet degeneracy discussed for ; two sizes do not establish that upward extrapolation is safe in general. Training on four alphabets buys little for interpolation at intermediate k, which indicates that the descriptor already transports across alphabet sizes rather than memorizing the two seen in training; it does cost slightly at , where diluting the training set weakens the degenerate small-alphabet case discussed above.
Table 11.
Interpolation across alphabets: integrated MSE on held-out alphabets and on the training alphabets, for a mapper trained on versus the main model trained on .
4.7. Ablation: What Does the Descriptor Contribute?
A natural objection is that the network might merely have learned the sample size. We test this by retraining the mapper under the same protocol with progressively poorer inputs: a constant input (a single fixed global mixture), the ratio alone, the five-dimensional regime summary of Equation (3), and the full 37-dimensional descriptor; Table 12 reports byte-sequence integrated MSE on a common evaluation stream.
Table 12.
Ablation on the feature map (bytes, integrated MSE, common evaluation stream, 500 replications per size). SHR on the same stream: .
The result is sharper than the objection anticipates. Neither a fixed global mixture nor weights conditioned on alone improve materially on SHR in integrated MSE. The reason is regime-specific and must be stated carefully. At , where condition (8) is satisfied and certified by the bootstrap, no per-regime convex combination of the six estimators can improve on the best member, so knowing the regime is not enough. At , the plug-in estimate of satisfies the condition, but the bootstrap certification is inconclusive, so the same conclusion is supported by the point estimate only. At , the condition fails, and a fixed mixture does improve on SHR by , as computed above, yet it remains a factor of three worse than ASED ( versus ). The integrated-MSE figures for the -only mapper therefore do not by themselves establish that every per-regime mixture is useless; the per-regime computation does, at the sizes where the condition holds, and quantifies the residual gap where it does not. Essentially the entire gain appears once the weights can respond to the per-sample coverage statistics , which predict the plug-in error of the individual sample; the FoF tail adds no measurable improvement for the i.i.d. alternatives studied here. This localizes the mechanism with precisely sample-conditional bias correction driven by coverage features and implies that a five-feature “ASED-lite” suffices in the i.i.d. setting, while the full descriptor is retained because it may provide headroom for structured deviations (Markov, periodic), whose signature would appear in the FoF tail; whether it does is an empirical question left open here.
4.8. Comparison with General-Purpose Estimators
The base set is deliberately restricted to closed-form estimators, but the comparison should not be. Table 13 adds two further general-purpose estimators evaluated on a common stream: the Chao–Wang–Jost estimator CWJ [20] and the Grassberger correction GR [13] alongside NSB, which is also reported under the full 1000-replication protocol in Table 4 and Table 5.
Table 13.
Integrated MSE including general-purpose estimators, common evaluation stream (200 replications per sample size). Absolute values differ slightly from Table 4 because of the smaller replication count; the ranking is the object of interest.
Our NSB implementation was validated before use: on three distributions of known entropy at (uniform, Zipf , and a half-mass two-point mixture), its bias at is , , and bits against , , and for the plug-in, and the estimate is stable to four decimals as the integration grid is refined from 20 to 200 points (, , ). It therefore behaves as documented; its poor showing below is a property of the uniform-null criterion, not of the implementation. NSB, CWJ, and GR are not competitive under the uniform null, and the reason is structural rather than incidental: they are built to be accurate for arbitrary distributions and therefore cannot exploit the uniform target that SHR shrinks toward and that ASED is calibrated on. The comparison is thus informative in one direction only; it confirms that the gain is specific to the uniformity-diagnostic setting, and Section 4.9 examines the converse regime, where the ordering reverses. We also implemented a simplified linear-programming reconstruction of the Unseen estimator [10]; it was not competitive in extreme undersampling (integrated MSE two orders of magnitude above the others, with estimates exceeding at ). We attribute this to our reconstruction rather than to the method and therefore exclude it from Table 13. The BUB estimator of Paninski [8] is omitted for the same reason: its published behavior depends on a tuning parameter and on a linear-programming construction whose released implementations we could not reproduce to the accuracy required for a fair comparison at these sample sizes, and reporting a self-coded approximation under our protocol would misrepresent it. A faithful comparison against the authors’ implementation remains open.
4.9. Beyond the Null: Strongly Non-Uniform and Dependent Sources
The power analysis probes alternatives close to the null. We now go deliberately far from it to delimit where ASED must not be used. Table 14 reports MSE against the true entropy for three families, specified here in full. Zipf: for with , sampled i.i.d. Markov: a first-order chain on with and for , and ; the initial state is uniform, which is also the stationary distribution of this doubly stochastic kernel, so the marginal entropy is exactly 8 bits at every position while the sequence is strongly dependent. Periodic: a phase u is drawn uniformly from , and the window is for , the 64 attainable values being regarded as a subset of the 256-symbol alphabet. Each position then has the same time-marginal symbol distribution under a uniformly random phase uniform on those 64 symbols, so bits exactly, and for n, a multiple of 64, each symbol occurs exactly times in the window. Drawing the phase from the full alphabet instead would make every positional marginal uniform on 256 symbols and the target 8 bits while leaving the within-window count pattern unchanged; we use the former so that the tabulated target matches the quantity the estimators are asked to recover. No block permutation is applied. Note that this is the time-marginal symbol entropy under a random phase, not the entropy rate: a deterministic periodic sequence has entropy rate zero, and neither nor is the target here.
Table 14.
MSE against the true (marginal) entropy for structured sources, , , 400 replications.
The picture is clear and partly unfavorable to ASED, which is why we report it. Near the null mild Zipf skew, where the true entropy is within a few hundredths of , ASED matches the best estimator. Far from the null, it fails: at , its MSE is that of NSB and that of SHR, and on a periodic source, it is worse than both. The failure is exactly the one predicted by the scope statement: a mixture whose weights are calibrated where inherits the behavior of its Bayesian-smoother components when the sample is strongly non-uniform, and those components are the wrong ones there. ASED is a uniformity diagnostic, not a general-purpose entropy estimator, and Table 14 is the quantitative statement of that boundary. One result runs the other way: under strong dependence with a uniform marginal, ASED and LAP estimate the marginal entropy an order of magnitude more accurately than SHR or NSB, because the FoF profile of a dependent sample resembles that of a smaller effective sample. We note, however, that no marginal-entropy estimator detects the dependence itself: the marginal is genuinely uniform, so this regime falls outside what any estimator in this comparison is designed to diagnose.
4.9.1. Is Uniform-Only Training the Cause?
To separate the effect of the training distribution from the effect of the base set, we retrained the mapper on a 50/50 mixture of uniform samples and near-uniform perturbations (one symbol carrying mass ), with the true entropy of each source as the target. The retrained mapper is indistinguishable from the original under the null (integrated MSE versus ) and improves only marginally far from it (Zipf : versus ; Zipf : versus ). Broadening the training distribution over near-uniform sources is therefore not what limits performance on strongly skewed sources; the results indicate that the convex hull of the current base set is the principal observed limitation, its members being themselves inaccurate in exactly the regime in which Proposition 1 becomes nearly saturated. Extending ASED beyond uniformity diagnostics would likely require adding general-purpose members such as NSB to the base set, not merely retraining on a wider family. The general question of what happens under an arbitrary family of source distributions has a partial answer in three parts and one genuine gap. Proposition 1 gives the distribution-free safety upper bound: for any source, the finite-sample inequality holds, and consistency and convergence rates are inherited whenever every member of the base set has the corresponding property under that source. Table 14 gives the empirical ceiling for six structured families, and the retraining experiment above indicates that, under the training distributions and optimization variants evaluated, the current base set is the primary observed constraint rather than the training measure, without establishing it as the unique limiting factor. What is missing is a characterization of the transfer regime: for which perturbation families around the uniform does a mapper fitted under retain the oracle behavior of Theorem 1 at ? Since the theorem is stated conditionally on , the natural route is a bound in terms of how much the conditional error geometry moves between P and Q, which we leave open.
4.9.2. Why a Neural Mapper?
The weight function need not be a network, and the objective admits two distinct strategies: fitting the weights directly against the mixture loss of Equation (5) or fitting a surrogate that predicts each base estimator’s squared error from and converting those predictions into weights. Table 15 evaluates six families under both strategies on the byte protocol.
Table 15.
Weight-function families under the identical protocol (bytes, integrated MSE). Direct: fitted against Equation (5). Surrogate: per-estimator squared-error regression, converted to weights by softmax or by selecting the predicted minimum.
The honest reading is that the architecture is not what carries the result. Every sufficiently flexible family evaluated, except the linear-softmax and Nadaraya–Watson specifications, lands within a factor of two of the MLP and an order of magnitude below the best fixed estimator; gradient boosting with hard selection is in fact marginally better (). Only the two genuinely restrictive families fall behind: a linear softmax, which cannot express the non-linear dependence, and kernel regression, which degrades because the FoF features are highly non-uniformly distributed across regimes. We retain the MLP because it optimizes the deployment objective directly, produces smooth weights, and costs nothing at inference, not because it dominates. We can now be precise about the comparison rather than merely cautious about it. Using the independent implementation of Section 4.10.3, we fitted all three candidates, the full 37-feature mapper, the five-feature ASED-lite variant, and the gradient-boosted surrogate with hard selection, on a common training stream and evaluated them on a single locked evaluation stream never used in fitting, then bootstrapped the differences by resampling the replication index so that all three see the same resamples. Table 16 reports the result. None of the three pairwise differences exclude zero. The gradient-boosted selector has the lowest point estimate, as seen in Table 15, but the paired interval for its advantage over the full mapper spans zero, as does the interval for ASED-lite against either alternative. No statistically supported ordering among the three families is established by these data, and none should be inferred from the point estimates alone; a practitioner may choose among them on grounds of interpretability, inference cost, or implementation convenience without accepting an accuracy penalty that these experiments can detect.
Table 16.
Paired comparison of three weight-function families on one locked evaluation stream (independent implementation, 1000 replications per sample size, 4000 paired bootstrap resamples of the replication index). Absolute values are higher than in Table 15 because these variants were fitted under a shorter common training budget for comparability; the differences, not the levels, are the object of interest.
4.9.3. Why 37 Features?
The tail length was fixed a priori to cover FoF coordinates up to . Retraining with truncated descriptors of dimension 5, 13, 21, and 37 gives integrated MSE , , , and , respectively: the choice is not load-bearing for i.i.d. sources, consistent with the ablation of Section 4.7. We report the full descriptor because it is the natural occupancy-profile object and leaves room for structured alternatives, while noting that the five-feature variant matches the integrated-MSE performance of the full descriptor under the evaluated i.i.d. protocols, with no measurable gain from the FoF tail. For the domain validated here, we therefore recommend ASED-lite as the parsimonious implementation, treating the 37-dimensional descriptor as an extensible architecture whose extra coordinates are motivated by, but not yet demonstrated on, structured alternatives. Read together with Table 16, this settles what the contribution of the paper is and is not. Of the three weight-function families compared on a common locked stream, the full 37-feature mapper, the five-feature variant, and a gradient-boosted surrogate with hard selection, no pairwise difference has a bootstrap interval excluding zero, and the surrogate has the lowest point estimate. Neither the MLP nor the 37-dimensional descriptor is therefore load-bearing, and we claim neither as a contribution. What the experiments support is the more general proposition: convex mixtures of classical entropy estimators whose weights are conditioned on per-sample occupancy statistics outperform any fixed member of the mixture in the undersampled regime under a uniform null, essentially independently of the function class used to produce the weights. The MLP over 37 features is the realization we fitted and report in full; ASED-lite is the implementation we recommend, and a practitioner preferring gradient boosting may use it without an accuracy penalty these experiments can detect.
4.10. Cryptographic and Defective Sources
All experiments so far use parametric families. Because the paper motivates the problem by cryptographic practice, we close the loop on actual generator output. We evaluate five sources at with 400 independent windows each, every window using an independent seed. (i) A reference stream generated by SHA-256 in counter mode: with a fresh 32-byte key per window, successive output blocks are for , concatenated and truncated to n bytes, where is the 64-bit big-endian encoding, and ‖ is byte concatenation. (ii) An RC4 keystream from a fresh 16-byte key, discarding the first 768 output bytes. (iii) RC4 initial output from a fresh 16-byte key with no discard. (iv) A truncated linear congruential generator with a fresh uniform seed, emitting the byte . (v) A source with controlled marginal bias on one symbol and the remaining mass uniform.
Three observations follow. Of the three cryptographic streams in Table 17, we treat SHA-256 in counter mode and RC4 after discarding 768 bytes as the references without known defects under this protocol, with the initial output of RC4 analyzed separately below. For these two, ASED is closer to than SHR at and identical from , matching the synthetic uniform results and providing evidence that the uniform-null behavior transfers to the two reference streams under the tested protocol. The truncated LCG deserves a precise statement, because its marginal entropy is available in closed form. The multiplier 1103515245 is odd, so and the affine update is a bijection on ; under a uniform seed, every state is therefore exactly uniform. The emitted byte is , and since , the shift leaves only attainable values, each with exactly preimages. The one-byte marginal is thus exactly uniform on a support of half the nominal alphabet, giving bits, exactly a genuine marginal deficit of one bit relative to , caused by the truncation discarding the alphabet rather than by serial dependence. A direct computation over uniform seeds returns , matching 7 minus the Miller–Madow bias for . Read against this target, both estimators are above the true value at every window length and descend towards it as n grows ( at ), which is the expected behavior of null-calibrated estimators on a source whose support is half the assumed alphabet. Both estimators therefore register a growing deficit relative to , with ASED slightly higher at short windows. Because this source lies under the alternative , failing to register the deficit is not the avoidance of a false alarm but a loss of detection: its upward null-calibration bias makes the diagnostic conservative at short windows and reduces sensitivity to this genuine one-bit marginal-entropy deficit. The overestimation is substantial, with mean estimates of about at and at against a true 7, so a support-restricted alternative of this kind is detected only slowly as the window grows. Third, pooling successive output positions within a window does not resolve position-specific structure, and RC4 is the standard example.
Table 17.
Mean entropy estimate on generator output (, 400 windows). “Target H” is the marginal entropy the source is designed to have: 8 bits for the two reference streams and for RC4 initial output when positions are pooled (its position-specific deficit is analyzed separately below) and the exact value for the biased source. The truncated LCG has an exactly uniform marginal on a support of 128 values; hence, bits (see text).
4.10.1. Position-Conditional Analysis of Early RC4 Output
Early RC4 keystream carries well-documented position-specific marginal biases, notably that of the second output byte toward zero [34,35]. A pooled-window entropy cannot resolve them, because averaging over positions dilutes each one; the comparison above is therefore silent about these biases rather than evidence against their detectability. The correct protocol conditions on the output position. We generated independent 16-byte keys, recorded the first sixteen keystream bytes of each, and analyzed the marginal distribution of across keys separately for each position t. The reference computation over all keys reproduces the known effect: against the uniform , a factor of two as predicted, while positions give –. The entropic consequence is nevertheless small: position-wise marginal entropies are bits at against – elsewhere, a deficit near bits. Estimating the entropy of from n independent keys, neither ASED nor SHR detect the anomaly at the sizes of interest: rejection rates at are , , and for (1000 replications, so binomial half-widths near ), close to the nominal size and identical for the two estimators; the value at is nominally above but corresponds to very low power in any case.
The conclusion is sharper than a pooled design would license. The position-conditional design does resolve the distributional bias, and it is now possible to say precisely how much sharper. As an additional check using the independent implementation of Section 4.10.3 (not the primary pipeline behind the results above), we compared three detectors under identical null calibration and the same resampled draws from the reference pool: ASED, SHR, and a direct test on the count of occurrences among n keys, calibrated by simulation against the exact null binomial. Table 18 reports rejection rates with 95% Wilson intervals. The direct test dominates throughout: at , it already has power against for ASED and SHR combined, and it reaches near-certain detection by , roughly an order of magnitude earlier in sample size than the entropy-based estimators. This confirms directly what was previously only a plausible reading: a test targeting is the right instrument for this specific bias, and the entropy-based estimators, while correctly calibrated in general, are the wrong tool for it. Other purpose-built statistics beyond this one direct test were not evaluated, so the comparison is between three specific detectors rather than an exhaustive search over possible tests; it nonetheless illustrates the scope limit stated throughout with a concrete, quantified alternative rather than a speculative one.
Table 18.
RC4 detection: rejection rate at , ASED and SHR (entropy-based, calibrated by simulation) versus a direct test on the count of occurrences (calibrated against the exact null binomial). 95% Wilson intervals in brackets.
Other entropy estimators and purpose-built statistics beyond the one direct test above were not evaluated, so this is an observation about the three detectors tested rather than an exhaustive impossibility result for entropy-based detection more generally; it nonetheless illustrates the scope limit stated throughout, now with a quantified alternative rather than a speculative one.
4.10.2. Tie-Breaking at Empirical Critical Values and Its Consequences
Every test in this paper calibrates its critical value as an empirical quantile of a simulated null distribution. When that null distribution is highly discrete, as it is throughout the regime that motivates the paper, a non-negligible probability mass sits exactly at the critical value, and the rule for handling it stops being a rounding detail. We state the convention explicitly and quantify its cost because the effect is larger than we anticipated.
The convention used throughout this paper is strict: the null is rejected only when the statistic falls strictly below (for entropy estimators) or strictly above (for and G) the empirical critical value, with ties resolved in favor of the null. Table 19 shows why the choice matters. At , the ASED null distribution takes only four distinct values over 8000 draws, and roughly of the null mass sits exactly at the quantile; switching from the strict to the inclusive convention moves reported power at a fixed alternative from to , a swing of 28 percentage points attributable entirely to the tie rule.
Table 19.
Discreteness of the ASED null distribution and the resulting sensitivity of reported power to the tie convention (, nominal , 8000 null and 4000 alternative replications). “Mass at crit” is the null probability exactly at the empirical critical value. “Gap” is the change in reported power from the tie rule alone.
The consequence is more serious than a loss of precision, and we report it rather than leave it implicit. Because the mass at the critical value is substantial, neither convention attains the nominal level at the smallest sample sizes. Measured on fresh null draws at , the strict convention gives an achieved Type-I error of at and at , while the inclusive convention gives and , respectively; nominal is attained by neither at , the strict rule being conservative and the inclusive rule anti-conservative by more than a factor of two. Exact size control with a discrete statistic would require a randomized test, which we have not implemented.
Two consequences follow for how the power comparisons in this paper should be read. First, the small-sample power figures in Table 8 and Table 9 are comparisons at the matched nominal level, not at the matched achieved level, and at , the achieved level is well below the nominal level for all detectors under the strict convention. Second, and more importantly for the non-inferiority analysis of Section 4.5, the margin is of the same order as the tie-induced ambiguity at and ; the non-inferiority conclusion at those two sample sizes should therefore be regarded as resting on a convention rather than as a sharp finding, which reinforces rather than softens the existing description of that analysis as post hoc and exploratory. The conclusion at and above, where the gap falls to and below, is not affected in the same way.
4.10.3. Independent Verification of the Exploratory Pipeline: The Locked Confirmatory Evaluation
Section 3 discloses that the design choices behind the primary pipeline, the feature map, architecture, and loss weighting were made while observing results on the same evaluation protocol used to report them, so no independent model-selection split separates development from the assessment reported in Section 4.1, Section 4.2, Section 4.3, Section 4.4, Section 4.5, Section 4.6, Section 4.7, Section 4.8, Section 4.9 and Section 4.10. That limitation is not resolved by what follows. As a partial, external check on whether the headline result is nonetheless reproducible under a genuinely locked protocol, we built a separate implementation, its own code, its own reconstruction of the 37-dimensional feature descriptor from the description in Section 3 rather than a copy of the descriptor actually used, and its own weight-mapper training run, and evaluated it under a design in which the training, validation, and test data are drawn from three independent random-seed streams that are never mixed. The architecture (two hidden layers of width 64 over the 37-dimensional input, six-way softmax output) was fixed in advance to match Table 1; this yields exactly 6848 weights and 134 biases, matching the primary model’s parameter count exactly and giving some confidence that the reconstruction is structurally faithful. The test stream was evaluated once, at the end, and used for nothing else. This independent implementation, together with the auxiliary scripts for the RC4 analysis, the Zipf sweep, the support-restricted generator, and the fresh-null power calibration, is released as Supplementary Materials (Code S2).
Under this protocol, the independent implementation reaches an integrated MSE of for ASED against for SHR at , a reduction, closely matching the primary pipeline’s reported against and reduction. Table 20 reports the full per-size breakdown, along with the integrated totals and the internal-consistency checks that follow. Four further qualitative findings from the primary pipeline were checked in the same independent run without being specifically targeted for reproduction. The instability described in Section 3 reproduced as elevated validation-loss variance during training. The support-restricted overestimation of Section 4.10 reproduced closely ( at and at against the primary pipeline’s and ). The Zipf degradation of Section 4.9 was reproduced in direction and rough magnitude but not exactly ( against NSB at , versus the primary pipeline’s , using an independently implemented and separately validated NSB estimator that runs at roughly 55– of the primary implementation’s reference bias scale). Finally, the covariance condition of Corollary 1 was re-verified by directly solving with a constrained quadratic program (confirmed to global optimality: 150 independent random restarts across 10 batches converged to the same solution to within , as guaranteed by M’s positive semi-definiteness) on an independently estimated error covariance matrix at . The qualitative pattern reproduced exactly: the condition holds at (the best single estimator is the unique optimal fixed mixture, improvement available) and fails at , where the optimal fixed mixture places weight on (SHR, LAP) and improves on the best single estimator by , the same qualitative finding as the primary pipeline’s weighting and improvement, on an independently estimated covariance matrix.
Table 20.
Independent verification, byte sequences (): per-size MSE from a separate implementation with a locked train/validate/test split (three independent random-seed streams, test evaluated once). The per-size column sums to the integrated total exactly, and the subset ( for ASED, for SHR) is strictly below it, as it must be.
Table 20 is internally consistent by construction, and we note the checks explicitly because their absence in the primary run is the subject of Section 4.11: the twelve per-size values sum to the reported integrated total; the subset sum ( for ASED against for SHR) is strictly smaller than that total, accounting for and of it, respectively; and the paired difference of the integrated values, , follows arithmetically from the two columns. The per-size breakdown also shows where the aggregate advantage comes from and where it does not: the reduction is at and falls monotonically to approximately zero by , beyond which ASED and SHR are indistinguishable and ASED is marginally worse in several cells. This is the same regime-concentration reported for the primary pipeline in Section 4.4 and is the reason the headline figure is so sensitive to the aggregation scheme.
This exercise has clear limits, stated plainly. The feature descriptor is a reconstruction from the manuscript’s own prose description, not a byte-for-byte copy of the code behind Section 4; a single training run is reported, without restart variance; and the NSB comparison used to check the Zipf result is calibrated to a different scale than the primary implementation’s, so the figure should be read as directionally, not numerically, confirmatory. Most importantly, this is external corroboration by a second, independently written codebase, not a resolution of the primary pipeline’s own model-selection contamination: the results reported in Section 4.1, Section 4.2, Section 4.3, Section 4.4, Section 4.5, Section 4.6, Section 4.7, Section 4.8, Section 4.9 and Section 4.10 still lack an independent split of their own, and closing that gap, rather than corroborating it externally, remains the priority stated in Section 3.
Finally, we record the treatment of the feature standardization constants and of Equation (4), since their status was previously left implicit. They are estimated once, from the first training batch of a given run, and held fixed thereafter: they are not recomputed at validation or test time, and each independently trained model, including each retraining in Section 4.6 and Section 4.9, and the verification model of this section, carries its own set. Reusing one model’s constants with another model’s weights is therefore not valid. For the verification model, all 37 pairs for each of and accompany the released coefficients. We note one respect in which the verification model is not a substitute for the primary one on this point: its descriptor is a reconstruction from the prose of Section 3 and does not coincide coordinate-for-coordinate with Equation (3), so its constants document the verification pipeline rather than the primary one, whose corresponding values must come from the primary implementation.
4.11. Diagnosis and Correction of a Numerical Inconsistency in the Exploratory Run
Two related numerical inconsistencies were identified in review, and we can now state their common cause precisely rather than merely acknowledging them. Section 4.4 reports the regime-sum for SHR as , which exceeds the integrated total of given in Table 4, impossible for a subset of a sum of non-negative terms. Separately, the paired difference reported in Section 3 is , whereas the full-precision integrated values in Table 4 imply .
Both arise from the same defect: two distinct evaluation runs were used within a single section. The main tables derive from one run, giving integrated values (ASED) and (SHR); the paired randomization analysis derives from a second run, giving and . The second run is internally consistent; its stored paired difference agrees with the difference of its own integrated values to twelve decimal places, and it is the source of the figure, which is correct for that run but not for the tabulated one. The figure in Section 4.4 is the same second run’s integrated SHR value, reported there as though it were the subset sum; the ratio is exactly , confirming it is the total rather than a subset. The ASED figure printed alongside it, , is by contrast a genuine subset of that run’s total, at of it, closely matching the subset share measured independently in Section 4.10.3. The defect is therefore confined to a single transcribed quantity, not to the underlying computation.
A third defect, independent of the first two, concerns the regime definitions themselves. The three regimes named in Section 4.4, , , and are not a partition of the twelve-point grid: falls in the first two, and and fall in the last two. Sums reported against these definitions therefore cannot add to the integrated total, whatever run they come from, since three of the twelve sizes are counted twice. A disjoint definition is required; we use , , and below.
To establish what an internally consistent regeneration looks like, and to give a template against which the primary run can be checked, we regenerated the complete byte-sequence evaluation from a single run of the independent implementation of Section 4.10.3 (1000 replications per size, one fixed seed, disjoint regimes), with the consistency conditions enforced as assertions rather than checked afterward. Table 21 reports it. Every aggregate follows from the per-size column: the three regime sums add to the integrated total exactly; each subset is strictly below its total; the paired difference equals the difference of the integrated values; and all four aggregation schemes are computed from the same per-size values. The four schemes give reductions of (summed), (arithmetic mean), (log-trapezoidal), and (mean of per-size ratios), reproducing the qualitative spread reported in Section 4.4 from a source in which the spread can be verified. The paired difference is with a bootstrap interval of , and a sign-flip randomization test over assignments yields no assignment as extreme as the observed statistic.
Table 21.
Internally consistent single-run regeneration of the byte-sequence evaluation, from the independent implementation of Section 4.10.3 (seed fixed, 1000 replications per size, disjoint regimes). All aggregates below the rule are computed from the per-size column above it.
We are explicit about what this table is and about the status of the affected figures in the present version. It is an internally consistent evaluation produced by the independent implementation, not a re-run of the exploratory pipeline: the feature descriptor is a reconstruction, so its values differ from the exploratory run’s in the third significant figure, and it does not license substituting these numbers for those of Table 4. What it establishes is that the structure of the result, a reduction above under summed aggregation, falling to roughly under per-size-ratio aggregation, with essentially all of the advantage confined to , survives regeneration under a protocol in which every aggregate is verifiable from its own constituents.
Three corrective actions have been taken in this revision rather than deferred, so that no figure in the manuscript now carries a provisional status. First, the overlapping regime definitions are retired and replaced throughout by the disjoint partition , , ; every regime split reported in the paper uses it. Second, the two quantities transcribed from a second run, the SHR subset sum of and the paired difference of with its interval, have been removed from Section 4 and Section 4.4, not merely annotated. The paired difference is replaced by the value implied by Table 4 itself () and the interval by the one produced by the single run of this table, attributed to it; the subset sum is replaced by the corresponding figure from this table. Third, the single-run rule stated in Section 4 is applied to every aggregate in the paper, and the aggregation code enforces the consistency conditions as run-time assertions (regime sums against the integrated total, subset sums strictly below their totals, paired differences against the difference of integrated values) rather than checking them afterwards.
The result is that no claim in this manuscript rests on a quantity drawn from a run other than the one behind its own table, and the word “provisional” no longer applies to any reported figure. One item remains genuinely outstanding, and we state it rather than conceal it: the exploratory pipeline of Section 4 has not itself been re-executed under the assertion-checked procedure, so the third significant figure of Table 4 may move when it is. Nothing in the paper’s conclusions depends on that digit. The locked confirmatory evaluation of Section 4.10.3, which was produced under exactly that procedure by an independently written implementation, reaches the same conclusion on its own evidence ( against , with the same regime concentration), and it is the evaluation on which the reader is asked to rely.
5. Discussion
The central finding is that a learned, sample-conditional mixture of classical estimators improves substantially on the best fixed estimator for short cryptographic byte sequences, in the undersampled regime and under a uniform marginal null: a reduction in integrated MSE over SHR of under summed or averaged aggregation and under the mean of per-size ratios, confined entirely to and reproduced at by the locked confirmatory evaluation of Section 4.10.3, with statistically separated bootstrap intervals, detection power non-inferior at a post hoc margin against the near-uniform alternatives tested (confirmatory for , exploratory at and ; Section 4.10.2), an advantage that survives held-out sample sizes and an unseen alphabet, and an ablation localizing the gain in the per-sample coverage features rather than in knowledge of the regime. Theorem 1 and Corollary 1 supply the matching explanation: the excess risk over the sample-conditional oracle is governed by how well the weights are approximated, and where condition (8) holds, no fixed or regime-indexed mixture of the base set can improve on its best member. Proposition 1 bounds the other side, and its near-saturation on skewed sources ( measured against a bound of ) shows that the failure in Section 4.9 is the degradation the theory predicts rather than an unmodeled pathology.
5.1. Why the Three-Layer Evaluation Is the Result
Because the degenerate rule attains zero null error by construction (Section 1), low null MSE is not by itself evidence of estimation ability, and convexity does not rule out a constant-valued adaptive combination. The decisive evidence is therefore behavioral: the constant control has zero power everywhere, while ASED’s power is non-inferior to SHR’s at a post hoc margin (Section 4.5). The benchmark-gap analysis sharpens the interpretation of what was learned. At , ASED outperforms every evaluated fixed estimator and the per-regime best-single-estimator benchmark, even though the errors of its dominant components are positively correlated; the gain therefore comes from weights that covary with the FoF features predicting the plug-in error, a learned bias correction calibrated to the null. The power analysis bounds the cost of that calibration: it is contained within the post hoc non-inferiority margin, although three operating points at exhibit a small statistically detectable deficit, with a measurable MSE penalty only at gross deviations where every test already rejects with probability one. The ablation also clarifies what the contribution is and is not. Although the full FoF tail does not improve the i.i.d. uniform-null results over the five-feature coverage summary, it is retained because it represents the occupancy-tail extension that may be needed for future structured alternatives; in the present experiments, the effective signal is concentrated in the low-order coverage coordinates, and the method is most accurately described as an interpretable convex mixture of classical estimators whose weights respond to per-sample coverage statistics.
5.2. Practical Implications
For frameworks that apply marginal-entropy statistics to short output blocks, ASED offers SHR’s detection behavior with substantially more concentrated point estimates under the null and detection power non-inferior to SHR’s at a margin. We say “more concentrated point estimates” rather than “narrower confidence intervals”: lower MSE does not by itself deliver shorter intervals with correct coverage, and we construct no interval procedure here. In S-box design loops that screen thousands of candidates on small output samples [7], the accuracy at beyond the per-regime best-single-estimator benchmark may permit earlier or cheaper comparisons between candidates. In stream cipher initialization analysis [3], the held-out results support applying the trained mapper at window sizes between the grid points without retraining. We should not overstate what the generator experiments add to this case, because they are the weakest evidence in the paper for a practical advantage over SHR rather than the strongest. On SHA-256 counter mode and post-discard RC4, the two estimators differ only at , by about bits, and coincide from onward. On the truncated LCG, both overestimate the true 7 bits substantially (≈7.96 at and at ), with ASED slightly the worse of the two at short windows. On the RC4 bias, a purpose-built direct test detects the defect roughly an order of magnitude earlier in sample size than either estimator, and the two entropy-based detectors are indistinguishable from each other throughout. The case for preferring ASED to SHR therefore rests on the null-regime MSE at and on non-inferior power at and not on the cryptographic-source experiments, which show the two behaving very similarly wherever they were directly compared. These statements are confined to marginal-symbol Shannon-entropy diagnostics under a uniform marginal null. ASED does not assess cryptographic randomness, unpredictability, independence, or min-entropy, and it is not a certification procedure: Section 4.10 exhibits a source (early RC4 output) whose known position-specific defect went undetected by both estimators tested at realistic sample sizes, and certification remains a min-entropy problem per SP 800-90B [6] and AIS 20/31 [23].
5.3. Limitations
First, the theory is open: weights and base estimates are dependent functions of the same FoF vector (Remark 1), so no learning-theoretic oracle guarantee for the fitted mapper follows from the offline construction alone, in contrast to the guarantees available for the super learner [27] or BUB [8]. Second, the neural mapper is a convenience rather than a necessity: Table 15 shows several simpler families within a factor of two and a gradient-boosted surrogate marginally ahead. Third, ASED is not a general-purpose entropy estimator, and Section 4.9 quantifies the boundary: on strongly skewed (Zipf ) or periodic sources, its error exceeds that of NSB by more than an order of magnitude, and retraining on a broader near-uniform family does not repair this: under the tested training distributions and optimization variants, the convex hull of the current base set appears to be the principal observed limitation, without being established as the unique limiting factor. The bit-sequence mapper likewise realizes only part of the improvement available in its own base set, leaving MIN ahead; the FoF descriptor is nearly degenerate at . Fourth, the alternatives tested are single-parameter families (biased bits, doped bytes), and ASED is calibrated under i.i.d. sampling: temporal dependence, Markovian artifacts, and block correlations may alter the FoF distribution in ways the current descriptor does not fully capture. Fifth, the headline reduction depends on the aggregation scheme: it is for summed or averaged MSE but for the mean of per-size ratios (Section 4.4), and no independent model-selection split separates development from final evaluation. Sixth, the comparison now includes NSB, CWJ, and Grassberger (Section 4.8) but still omits BUB, whose faithful implementation requires the regularized minimax program, and reports no fair figure for Unseen, since our simplified reconstruction was not competitive, and we do not attribute that to the original method. Seventh, the small-sample power comparison is not identifiable: at , the ASED and SHR statistics take the same few null values up to relabeling, no test of size exists, and the power differences reported in Table 8 and Table 9 are artifacts of where an interpolated empirical quantile falls between atoms. The tests reported here are calibrated at nominal rather than achieved level. The null statistic is highly discrete for , where neither tie convention attains , and the choice of convention alone moves reported power by up to ; the resulting ambiguity exceeds the non-inferiority margin, so the power conclusions at the two smallest sample sizes rest on a convention and are exploratory (Section 4.10.2). Correcting this requires a randomized or exact test, which we have not implemented.
5.4. What the Theory Does and Does Not Settle
Remark 4 separates four contributions. The approximation and idealized generalization terms are partially analyzed; optimization error and design mismatch remain uncontrolled. The idealized generalization term is standard and decreases with the training budget; the approximation term, how well a small network represents the conditional oracle on the FoF domain, has no rate here, and supplying one, together with the Rademacher complexity of this specific architecture, is the main problem left open. A second open question concerns what Corollary 1 does not provide: it characterizes when a vertex of the simplex minimizes the fixed-weight risk, but it does not establish strict dominance of the sample-conditional oracle over the per-regime one. Conditions on under which that dominance is strict, and a quantification of the gap, remain open. Super-learner arguments [27] do not transfer directly, since weights and base estimates are functions of the same statistic.
6. Conclusions
We introduced ASED, a sample-conditional mixture entropy estimator for short sequences in the marginal uniformity-testing setting of symmetric cryptography. A compact MLP (6848 weight parameters), trained once offline with a regime-balanced objective, maps a 37-dimensional frequency-of-frequencies descriptor to convex weights over six classical estimators, at inference cost.
On the theoretical side, we prove an oracle inequality bounding the excess risk of any conditional mixture by the geometry of the conditional error matrix and the weight approximation error; a checkable covariance condition under which sample-dependent weighting is necessary for any convex-mixture improvement, which explains the regime-specific ablation findings where it holds and quantifies why a fixed mixture remains insufficient at ; and fixed-alphabet consistency for every simplex-valued weighting rule, independently of the training distribution, together with non-asymptotic safety bounds. The precise hypotheses of full support of the conditional oracle for the quadratic rate, and a non-degenerate asymptotic variance together with a vanishing minimax-prior weight for asymptotic normality, the former failing under the uniform null, are stated with the results in Section 3.
The evaluation proceeds in layers, each closing a specific evidentiary gap. Under the uniform-null protocol of Contreras Rodríguez et al. [2], ASED attains an integrated MSE of for bytes against for the best single estimator, SHR, a reduction of under summed or averaged aggregation and under the mean of per-size ratios, confined to and reproduced at under the locked confirmatory protocol of Section 4.10.3 with a paired bootstrap interval for the integrated-MSE difference excluding zero, and outperforms the per-regime best-single-estimator benchmark for ; for bits, it improves on SHR by , ranking second behind the minimax-prior estimator, a single-alphabet figure we do not advance as a claim for the method. MIN belongs to ASED’s own base set, and the run-to-run spread at is comparable to the margins separating the estimators. Against near-uniform alternatives of known entropy, ASED’s detection power was non-inferior to SHR’s within a post hoc margin of at every tested operating point, pointwise and simultaneously, with one exclusion: at , the ASED and SHR statistics are monotone relabelings of one another on the null support, and no test of size exists, so no power difference is identifiable and the comparison is withdrawn there (Section 4.5), while the constant control, which achieves zero error by construction under the null, has zero power separating genuine adaptive estimation from calibration artifacts. On held-out sample sizes and an unseen alphabet, the advantage persists.
On strongly non-uniform or periodic sources, the ordering reverses, and general-purpose estimators such as NSB dominate by an order of magnitude, a boundary we quantify rather than elide. The contribution is therefore best stated without reference to the network or to the 37-dimensional descriptor, neither of which is load-bearing: it is the sample-conditional calibration of convex mixtures of classical entropy estimators using occupancy statistics. The neural mapper and the extended descriptor are one convenient realization: a five-feature MLP and, separately, a gradient-boosted surrogate each achieved comparable performance, the two ablations having been run independently. For uniformity diagnostics on short cryptographic byte sequences, these results justify ASED as a candidate for prospective validation against SHR in short-sample marginal-uniformity diagnostics: a nested bootstrap supported pointwise and simultaneous non-inferiority over the nine pre-defined small-sample cells at the post hoc margin, estimation error under the null is substantially lower, and the computational overhead is negligible. That prospective validation matters concretely: the integrated-MSE reduction falls to under a per-size-ratio aggregation (Section 4.4), and the feature map, architecture, and loss weighting were all selected while observing this same evaluation protocol, so a locked train/validate/test split, rather than a restatement of the present results, is what “prospective” is intended to require here. The agenda this work opens is the generalization half of the theory, a bound on for the offline fit, together with base sets extended by general-purpose members such as NSB, descriptors informative at small alphabets, and a faithful comparison against BUB and Unseen.
Supplementary Materials
The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/e28101073/s1. Code S1: reference implementation of ASED and of the six base estimators, the weight-mapper trainer, the locked evaluation harness with its three disjoint random-seed streams, and the aggregation module in which the internal-consistency conditions of Section 4.11 are enforced as run-time assertions; a regression test suite covering each of those conditions is included, and run_all.py regenerates the byte-sequence and bit-sequence evaluations end to end in under a minute per configuration on one CPU core. Data S1: trained weight-mapper coefficients (the 6848 weights and 134 biases of Equation (4)) together with the feature standardization constants and , one set per trained model, since one model’s constants are not valid with another model’s weights (Section 4.10.3). Data S2: full-precision per-size results for every estimator, sample size, and configuration reported in Table 3, Table 4, Table 5, Table 6, Table 7, Table 8, Table 9, Table 10, Table 11, Table 12, Table 13, Table 14, Table 15, Table 16 and Table 17, in the disjoint regime partition used throughout. Code S2: the independent implementation used for the locked confirmatory evaluation of Section 4.10.3, together with the auxiliary scripts for the RC4 analysis, the Zipf sweep, the support-restricted generator, and the fresh-null power calibration; it is not the pipeline behind the exploratory results and is not guaranteed to share every implementation detail with it.
Funding
This research received no external funding.
Data Availability Statement
The complete implementation is released in support of the results reported here: the reference implementation of ASED and of all base estimators; the trained weight-mapper coefficients together with the feature-standardization constants and belonging to each trained model, including every retrained variant of Section 4.6, Section 4.9 and Section 4.10.3; the independently written implementation used for the locked confirmatory evaluation; the random seeds and seed-stream definitions of every experiment; and the evaluation and aggregation scripts that regenerate each table and figure in this paper from those seeds. In view of the inconsistency documented in Section 4.11, the aggregation scripts carry the internal-consistency conditions as run-time assertions, so that regime sums, subset sums, and paired differences are checked against their own per-size constituents whenever a table is produced. The archive is deposited in a public repository under a permanent identifier, to be inserted at proof stage; until then, it is available from the corresponding author on request.
Conflicts of Interest
The author declares no conflicts of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| ASED | Adaptive sample-conditional entropy diagnostic |
| FoF | Frequency-of-frequencies |
| MSE | Mean squared error |
| SHR | James–Stein shrinkage estimator |
| NSB | Nemenman–Shafee–Bialek estimator |
| MM | Miller–Madow estimator |
| CS | Chao–Shen estimator |
Appendix A. Base Estimator Definitions
All estimators operate on the counts , with , , and entropies in bits.
- Dirichlet family.
For a prior pseudocount , the smoothed frequencies and plug-in entropy are
with the four members used here written explicitly (note that the MIN pseudocount is , not , so that and the denominator is ; the displayed frequencies sum to one exactly):
Plug-in, corrected, and coverage-based estimators.
with the guard when .
- Shrinkage estimator.
SHR [22] plugs in the convex combination with the uniform target,
i.e., the data-driven shrinkage intensity toward the uniform target with explicit clipping to . If the denominator vanishes (the empirical distribution equals the uniform target), we set ; for , the uniform plug-in is returned directly.
Appendix B. Uniform-Convergence Bound for the Offline Fit
Proposition A1
(PAC bound for an idealized empirical-risk minimizer). Assume (H1) the training pairs , , are i.i.d. draws from the training distribution; (H2) for the idealized procedure considered only in this proposition, base estimates are clipped to with , so that both the clipped estimates and the true entropy lie in and the loss takes values in ; and (H3) minimizes the empirical risk over a class Λ of weight functions with values in , specified before seeing the data (for instance, by norm constraints fixed in advance). Then, for any , with probability at least over the training sample,
where , and is the Rademacher complexity with respect to the training distribution.
Remark A1
(What this does and does not describe). Proposition A1 bounds an idealized procedure, not the training actually run in Section 3, which differs in four respects: sampling is stratified with exactly 1000 replications per configuration rather than i.i.d.; the loss is regime-weighted and carries an epoch-dependent entropy term; the objective is non-convex and optimized by Adam, so no global minimizer is certified; and the class is not norm-constrained a priori. The honest form of the guarantee for the deployed model is therefore
and adapting the concentration step to a stratified design with a weighted loss, together with bounding the optimization error, is left open. For the same reason the norm-based evaluation below, which uses the norms of the trained weights, describes a data-dependent class and would require a localization, peeling, or sample-splitting argument to be a valid uniform bound; we report it as a numerical illustration only.
Proof.
Write . By (H2), the loss lies in , so changing one training pair changes by at most ; McDiarmid’s inequality gives with probability , and by (H1), standard symmetrization gives . The map is -Lipschitz on the range of because there, so Talagrand’s contraction principle yields . Hence, . Let attain ; by (H3), , and a second application of the same concentration argument to the single function (or to the supremum, as here) bounds . Combining the two displays with gives (A7). □
Remark A2
(Evaluating the complexity term). For the architecture of Equation (4), admits an explicit norm-based bound: composing the depth-3 ReLU bound of Golowich et al. [36] with the vector-contraction inequality of Maurer [37] for the softmax layer gives , with . Evaluated at the trained weights (, , , , ), this yields a value of order : the bound is numerically vacuous at our training size, as norm-based bounds for networks typically are. Moreover, these are the norms of the trained weights, so the class is data-dependent, and the evaluation is not a valid uniform bound; we report it only to make the constants and the scaling explicit. A direct Monte Carlo estimate of the empirical Rademacher complexity of maximizing by gradient ascent over five independent sign draws on a 5000-sample subset and gives , three to four orders of magnitude below the norm-based bound. Since local optimization attains the supremum only approximately, this is an estimate rather than a certificate; closing the gap between the two numbers is precisely the open problem noted in Remark 4.
References
- Cover, T.M.; Thomas, J.A. Elements of Information Theory, 2nd ed.; Wiley-Interscience: Hoboken, NJ, USA, 2006. [Google Scholar] [CrossRef] [Scilit]
- Contreras Rodríguez, L.; Madarro-Capó, E.J.; Legón-Pérez, C.M.; Rojas, O.; Sosa-Gómez, G. Selecting an Effective Entropy Estimator for Short Sequences of Bits and Bytes with Maximum Entropy. Entropy 2021, 23, 561. [Google Scholar] [CrossRef] [Scilit]
- Madarro-Capó, E.J.; Legón-Pérez, C.M.; Rojas, O.; Sosa-Gómez, G. Information Theory Based Evaluation of the RC4 Stream Cipher Outputs. Entropy 2021, 23, 896. [Google Scholar] [CrossRef] [Scilit]
- Rukhin, A.; Soto, J.; Nechvatal, J.; Smid, M.; Barker, E.; Leigh, S.; Levenson, M.; Vangel, M.; Banks, D.; Heckert, A.; et al. A Statistical Test Suite for Random and Pseudorandom Number Generators for Cryptographic Applications; Technical Report NIST Special Publication 800-22 Rev. 1a; National Institute of Standards and Technology: Gaithersburg, MD, USA, 2010. [Google Scholar] [CrossRef] [Scilit]
- L’Ecuyer, P.; Simard, R. TestU01: A C Library for Empirical Testing of Random Number Generators. ACM Trans. Math. Softw. 2007, 33, 1–40. [Google Scholar] [CrossRef] [Scilit]
- Turan, M.S.; Barker, E.; Kelsey, J.; McKay, K.A.; Baish, M.L.; Boyle, M. Recommendation for the Entropy Sources Used for Random Bit Generation; Technical Report NIST Special Publication 800-90B; National Institute of Standards and Technology: Gaithersburg, MD, USA, 2018. [Google Scholar] [CrossRef] [Scilit]
- Zahid, A.H.; Arshad, M.J.; Ahmad, M. A Novel Construction of Efficient Substitution-Boxes Using Cubic Fractional Transformation. Entropy 2019, 21, 245. [Google Scholar] [CrossRef] [Scilit]
- Paninski, L. Estimation of Entropy and Mutual Information. Neural Comput. 2003, 15, 1191–1253. [Google Scholar] [CrossRef] [Scilit]
- Antos, A.; Kontoyiannis, I. Convergence Properties of Functional Estimates for Discrete Distributions. Random Struct. Algorithms 2001, 19, 163–193. [Google Scholar] [CrossRef] [Scilit]
- Valiant, G.; Valiant, P. Estimating the Unseen: Improved Estimators for Entropy and Other Properties. J. ACM 2017, 64, 1–41. [Google Scholar] [CrossRef] [Scilit]
- Verdú, S. Empirical Estimation of Information Measures: A Literature Guide. Entropy 2019, 21, 720. [Google Scholar] [CrossRef] [Scilit]
- Miller, G.A. Note on the Bias of Information Estimates. In Information Theory in Psychology: Problems and Methods; Quastler, H., Ed.; Free Press: Glencoe, IL, USA, 1955; pp. 95–100. [Google Scholar]
- Grassberger, P. Entropy Estimates from Insufficient Samplings. arXiv 2003, arXiv:physics/0307138. [Google Scholar]
- Schürmann, T. Bias Analysis in Entropy Estimation. J. Phys. A Math. Gen. 2004, 37, L295–L301. [Google Scholar] [CrossRef] [Scilit]
- Krichevsky, R.E.; Trofimov, V.K. The Performance of Universal Encoding. IEEE Trans. Inf. Theory 1981, 27, 199–207. [Google Scholar] [CrossRef] [Scilit]
- Holste, D.; Große, I.; Herzel, H. Bayes’ Estimators of Generalized Entropies. J. Phys. A Math. Gen. 1998, 31, 2551–2566. [Google Scholar] [CrossRef] [Scilit]
- Trybula, S. Some Problems of Simultaneous Minimax Estimation. Ann. Math. Stat. 1958, 29, 245–253. [Google Scholar] [CrossRef] [Scilit]
- Nemenman, I.; Shafee, F.; Bialek, W. Entropy and Inference, Revisited. In Proceedings of the Advances in Neural Information Processing Systems; MIT Press: Cambridge, MA, USA, 2002; Volume 14, pp. 471–478. [Google Scholar]
- Chao, A.; Shen, T.J. Nonparametric Estimation of Shannon’s Index of Diversity When There Are Unseen Species in Sample. Environ. Ecol. Stat. 2003, 10, 429–443. [Google Scholar] [CrossRef] [Scilit]
- Chao, A.; Wang, Y.T.; Jost, L. Entropy and the Species Accumulation Curve: A Novel Entropy Estimator via Discovery Rates of New Species. Methods Ecol. Evol. 2013, 4, 1091–1100. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Z. Entropy Estimation in Turing’s Perspective. Neural Comput. 2012, 24, 1368–1389. [Google Scholar] [CrossRef] [Scilit]
- Hausser, J.; Strimmer, K. Entropy Inference and the James–Stein Estimator, with Application to Nonlinear Gene Association Networks. J. Mach. Learn. Res. 2009, 10, 1469–1484. [Google Scholar]
- Killmann, W.; Schindler, W. A Proposal for: Functionality Classes for Random Number Generators (AIS 20/31); BSI Technical report; Bundesamt für Sicherheit in der Informationstechnik (BSI): Bonn, Deutschland, 2011. [Google Scholar]
- Schürmann, T.; Grassberger, P. Entropy Estimation of Symbol Sequences. Chaos An. Interdiscip. J. Nonlinear Sci. 1996, 6, 414–427. [Google Scholar] [CrossRef] [Scilit]
- Wolpert, D.H. Stacked Generalization. Neural Netw. 1992, 5, 241–259. [Google Scholar] [CrossRef] [Scilit]
- Breiman, L. Stacked Regressions. Mach. Learn. 1996, 24, 49–64. [Google Scholar] [CrossRef] [Scilit]
- van der Laan, M.J.; Polley, E.C.; Hubbard, A.E. Super Learner. Stat. Appl. Genet. Mol. Biol. 2007, 6, 25. [Google Scholar] [CrossRef] [Scilit]
- Nemenman, I. Coincidences and Estimation of Entropies of Random Variables with Large Cardinalities. Entropy 2011, 13, 2013–2023. [Google Scholar] [CrossRef] [Scilit]
- Jacobs, R.A.; Jordan, M.I.; Nowlan, S.J.; Hinton, G.E. Adaptive Mixtures of Local Experts. Neural Comput. 1991, 3, 79–87. [Google Scholar] [CrossRef] [Scilit]
- Jordan, M.I.; Jacobs, R.A. Hierarchical Mixtures of Experts and the EM Algorithm. Neural Comput. 1994, 6, 181–214. [Google Scholar] [CrossRef] [Scilit]
- Sill, J.; Takács, G.; Mackey, L.; Lin, D. Feature-Weighted Linear Stacking. arXiv 2009, arXiv:0911.0460. [Google Scholar]
- Yang, Y. Adaptive Regression by Mixing. J. Am. Stat. Assoc. 2001, 96, 574–588. [Google Scholar] [CrossRef] [Scilit]
- Tsybakov, A.B. Optimal Rates of Aggregation. In Proceedings of the Learning Theory and Kernel Machines (COLT 2003); Learning Theory and Kernel Machines; Springer: Berlin/Heidelberg, Germany, 2003; Volume 2777, pp. 303–313. [Google Scholar] [CrossRef] [Scilit]
- Mantin, I.; Shamir, A. A Practical Attack on Broadcast RC4. In Proceedings of the Fast Software Encryption (FSE 2001); Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2002; Volume 2355, pp. 152–164. [Google Scholar] [CrossRef] [Scilit]
- Sen Gupta, S.; Maitra, S.; Paul, G.; Sarkar, S. (Non-)Random Sequences from (Non-)Random Permutations—Analysis of RC4 Stream Cipher. J. Cryptol. 2014, 27, 67–108. [Google Scholar] [CrossRef] [Scilit]
- Golowich, N.; Rakhlin, A.; Shamir, O. Size-Independent Sample Complexity of Neural Networks. In Proceedings of the 31st Conference on Learning Theory (COLT), Stockholm, Sweden, 6–9 July 2018; PMLR: Cambridge, MA, USA, 2018; Volume 75, pp. 297–299. [Google Scholar]
- Maurer, A. Vector-Contraction Inequality for Rademacher Complexities. In Proceedings of the Algorithmic Learning Theory (ALT 2016); Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2016; Volume 9925, pp. 3–17. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.





