Skip to Content
StatsStats
  • Article
  • Open Access

23 September 2026

25 Pages

Information-Preserving and Model-Aware Voronoi Dequantization of Weighted Discrete Laws

,
,
,
and
1
Department of Mathematics and Statistics, Utah State University, Logan, UT 84322, USA
2
Department of Statistics, Mathematics and Insurance, Faculty of Commerce, Benha University, Benha 13511, Egypt
3
Department of Statistics, Nahda University in Beni Suef (NUB), Beni Suef 62511, Egypt
4
Department of Management Information Systems, College of Business and Economics, Qassim University, Buraidah 52571, Saudi Arabia

Abstract

Continuous dequantization embeds discrete data into a continuous space, but relaxation can alter the statistical information carried by the original categories. We study a complementary regime in which a specified weighted discrete law is the target and the dequantizer is required to be lossless under a prescribed quantizer. Any partition-respecting, parameter-independent kernel produces a continuous experiment Blackwell equivalent to the original discrete one, so likelihood ratios, maximum-likelihood estimators, score and Fisher information, Bayes posteriors, and optimal risks are unchanged before downstream approximation. We then characterize what happens after a finite-capacity continuous model is fitted. A Kullback–Leibler (KL) chain rule separates continuous approximation error into the categorical error of the requantized model plus a within-cell shape term; the information-preservation guarantee therefore does not extend automatically to an arbitrary fitted downstream model. We also provide a plug-in analysis for empirically estimated weights, showing that dequantization preserves rather than removes the discrete estimation error. Within each cell, maximizing entropy minus a displacement penalty yields a truncated Gibbs kernel. The apparent concentration and geometric-scale parameters reduce exactly to one effective concentration, α = λ / s c 2 . We further derive exact transport costs, a sharp W 1 = Θ ( h 2 ) reconstruction rate for the special reconstruction-from-exact-bin-masses setting, leakage-specific and geometry-perturbation bounds, and a bounded-domain multidimensional formulation. A finite-candidate Monte Carlo selector satisfies a nonasymptotic O ( m − 1 / 2 ) oracle inequality. Experiments verify the identities and show that exact dequantization can materially reduce requantized error when a smooth continuous downstream model is imposed; this is a conditional representational advantage, not a claim that dequantization dominates an unrestricted categorical model.

1. Introduction

Let P X = ∑ i = 1 n p i δ x i be a known discrete probability law. A common machine-learning workflow replaces the atoms with a continuous random variable before fitting a continuous-density model. For integer-valued data, uniform dequantization adds a fixed-width uniform perturbation to each category; this device is standard in neural autoregressive models and normalizing flows [1,2,3,4]. More general dequantization objectives and learned conditional dequantizers have been developed specifically to bridge discrete and continuous likelihood models [5]. Fixed width is natural when the discrete values are equally spaced, but not when the support carries a meaningful, irregular numeric geometry.
More broadly, the representation problem is not unique to dequantization: recent statistical work has emphasized that continuous model construction should respect support, identifiability, entropy, and the information contained in the observed representation. Examples include bounded and lifetime distribution models in which reparameterization or shape constraints determine the effective model class [6,7,8], and uncertainty-aware methods that retain interval or soft-data information rather than collapsing it to a single point [9,10]. These studies address different inferential problems, but they reinforce the general principle motivating the present work: a convenient continuous representation should not silently discard structure that is relevant to the downstream statistical task.
The geometric idea of spreading point mass over Voronoi cells is not new. Classical Voronoi density estimators assign density inversely proportional to cell volume, and recent work has developed compactified and continuous variants for density estimation from samples [11,12]. Chen, Amos, and Nickel [13] introduced learned Voronoi dequantization inside semi-discrete normalizing flows. These papers make it inappropriate to claim novelty for a piecewise-constant Voronoi formula itself.
This paper asks a narrower but, we believe, useful question: what exact probabilistic structure is obtained when a known weighted discrete law is dequantized on a prescribed geometric partition? That change in viewpoint produces several properties that are directly relevant to statistical computation and that are not the usual focus of sample-based Voronoi density estimation.
Our contributions are as follows:
1.
We formulate partition-respecting dequantization as a Markov kernel and prove exact requantization. For parameterized weight vectors, the discrete and continuous experiments are Blackwell equivalent [14,15]; hence, exact dequantization is lossless for statistical decision making.
2.
We prove exact likelihood and f-divergence preservation for a common parameter-independent kernel. We then derive a downstream KL chain rule: after fitting an arbitrary continuous density, its KL error decomposes exactly into requantized categorical KL plus a weighted within-cell approximation term.
3.
We use that decomposition to separate representation fidelity from model capacity. A continuous model can waste approximation capacity on within-cell shape, even when all exact dequantizers contain the same discrete information. This yields a model-aware kernel-selection criterion based on requantized loss, together with a finite-sample oracle inequality when categorical risk is estimated by Monte Carlo requantization.
4.
We solve an entropy–transport variational problem over all exact cell-confined kernels. The unique solution is a truncated Gibbs family k i , λ ∝ exp { − λ c i } ; λ = 0 is the maximum-entropy uniform kernel, while increasing λ decreases expected displacement monotonically without sacrificing exact requantization.
5.
We quantify two specific perturbation effects: category leakage bounds requantization error for approximate kernels, while a geometry-perturbation bound controls the effect of using perturbed support sites or quantization boundaries.
6.
For Voronoi cells, we obtain the exact W r transport cost to the atomic law, closed-form one-dimensional density/CDF/quantile formulas, and a sharp W 1 = Θ ( h 2 ) result for the distinct special case of reconstructing a smooth law from its exact bin masses. A bounded-domain formulation handles multidimensional cells correctly.
7.
Reproducible experiments verify the exact inference/divergence/KL identities, compare ten-seed Gaussian-mixture capacity curves, validate model-aware selection with bounded rational-quadratic spline normalizing flows, audit finite-sample selector regret and geometry perturbations, and include a real irregular-spacing example.
The proposed framework should therefore be viewed as an information-preserving and model-aware dequantization framework, not as a new nonparametric density estimator. The geometry fixes which continuous draws decode to each discrete state; the within-cell kernel can then be chosen to balance entropy, transport, and downstream approximation while preserving the target PMF exactly.
The model-aware component also makes validation design part of the method. Recent work on adaptive thresholds, deployment auditing under distribution shift, and risk-gated recalibration has similarly separated data-driven model or threshold selection from held-out or prospective evaluation [16,17,18]. The present problem is different–the candidate kernels are exact right inverses before downstream approximation–but the same validation principle applies once kernel selection and continuous-model fitting are learned from finite data.

2. Relation to Existing Voronoi and Dequantization Methods

2.1. Uniform and Learned Dequantization

Uniform dequantization adds independent uniform noise to integer-like observations. The continuous likelihood can then be related to the discrete likelihood, while variational and importance-weighted dequantization learn or optimize more flexible conditional noise distributions [2,4,5]. Nielsen and Winther [19] identified the usual variational dequantization gap as a conditional KL term, and that truncated-flow dequantizers can enforce non-overlapping bounded supports while remaining trainable [20]. These studies optimize a generative likelihood for unknown data distributions. The present regime instead takes the weighted discrete law as the target itself: support points may be irregularly spaced, exact requantization is required, and we ask which exact within-cell kernel is most convenient for a restricted downstream continuous model.
Chen et al. [13] used Voronoi cells inside a learned semi-discrete flow, including multidimensional learned quantization boundaries. Their contribution is substantially more expressive computationally. Our goal is complementary: a closed-form deterministic operator with exact requantization, exact divergence preservation, and no learned tessellation. To test whether the model-aware effect persists beyond Gaussian mixtures, Section 6.5 uses monotone rational-quadratic spline flows, a flexible analytically invertible normalizing-flow transformation introduced by Durkan et al. [21].

2.2. Voronoi Density Estimation

The VDE literature starts from samples of an unknown continuous law. A Voronoi tessellation is constructed from the sample, and density is related to the inverse cell volume. Polianskii et al. [11] emphasized both the computational difficulty of high-dimensional Voronoi cells and the need for compactification; Marchetti et al. [12] developed a continuous radial VDE. In contrast, here, p i is known, and the cell integral is required to be exactly p i . On a bounded domain, the resulting density formula coincides with a weighted piecewise-constant VDE. We explicitly acknowledge this overlap; the contribution lies in the dequantization/right-inverse and information-preservation characterization.

3. Partition-Respecting Dequantization

3.1. General Construction

Let Ω ⊂ R d be bounded and measurable. Let C 1 , … , C n form a measurable partition of Ω up to null sets, with  0 < Vol ( C i ) < ∞ . Associate one discrete state x i with each cell and define the quantizer
Q ( y ) = x i , y ∈ C i .
For a weight vector p = ( p 1 , … , p n ) in the probability simplex, define
( T C p ) ( y ) = ∑ i = 1 n p i Vol ( C i ) 1 C i ( y ) .
Equivalently, sample X = x i with probability p i , then sample Y ∣ X = x i ∼ Unif ( C i ) .
Theorem 1 
(Exact stochastic right inverse). Under the coupled construction above, Q ( Y ) = X almost surely. Consequently,
Q # ( T C p ) = P X , P { Q ( Y ) = x i } = p i for every i .
Proof. 
Conditional on X = x i , the draw Y belongs to C i with probability one, and  Q ( y ) = x i on that cell. Integrating over X provides the stated marginal identity.    □
This property is stronger than saying that the continuous density integrates to one: it says the continuous representation retains the discrete law exactly under its associated quantizer.

3.2. Blackwell Equivalence and Likelihood Preservation

The right-inverse property has a stronger statistical interpretation. Let { p θ : θ ∈ Θ } be any parameterized family of weights on the fixed support, and let K ( d y ∣ x i ) = k i ( y ) d y be a kernel that does not depend on θ , with  k i supported on C i . Let E X = { P θ X } denote the discrete experiment and E Y = { P θ Y } its dequantized version.
Theorem 2 
(Lossless statistical experiment). The experiments E X and E Y are Blackwell equivalent. The kernel K maps P θ X to P θ Y , while the deterministic kernel induced by Q maps P θ Y back to P θ X for every θ. Therefore, they have the same attainable risk set for every statistical decision problem.
If additionally p θ , i > 0 on the parameter region of interest, then for y ∈ C i ,
log f θ ( y ) = log p θ , i + log k i ( y ) .
For an i.i.d. sample, the continuous and discrete log-likelihoods differ only by a θ-free additive term. Consequently, their likelihood ratios, MLEs, score functions, observed information, and Fisher information coincide; with the same prior, their Bayes posteriors also coincide after applying Q.
Proof. 
The first statement follows because P θ Y = P θ X K , and Theorem 1 provides P θ X = P θ Y Q for every θ , which is precisely mutual garbling in the sense of Blackwell. On  C i , disjoint support yields f θ ( y ) = p θ , i k i ( y ) , proving the likelihood identity. Every stated likelihood-based consequence follows because the added kernel term is independent of θ .    □
The general comparison-of-experiments theory is classical [14,15]; the point here is that exact partition-respecting dequantization furnishes the two Markov kernels explicitly. Thus, continuous relaxation can be statistically lossless for inference on the weight vector even though a finite-capacity downstream density model may later introduce approximation error.

3.3. Exact Preservation of Information Divergences

Let p and q be two weight vectors on the same support and use the same partition. For a convex function ϕ with ϕ ( 1 ) = 0 , the Csiszár f-divergence is
D ϕ ( P ∥ Q ) = ∫ q ( y ) ϕ p ( y ) q ( y ) d y ,
with the usual conventions at zero mass.
Theorem 3 
(f-divergence isometry). For every f-divergence for which the two sides are defined,
D ϕ ( T C p ∥ T C q ) = D ϕ ( p ∥ q ) = ∑ i = 1 n q i ϕ p i q i .
In particular,
TV ( T C p , T C q ) = TV ( p , q ) , KL ( T C p ∥ T C q ) = KL ( p ∥ q ) ,
and the squared Hellinger distance is also preserved exactly.
Proof. 
On C i , the ratio ( T C p ) / ( T C q ) equals p i / q i . Therefore,
∫ C i q i Vol ( C i ) ϕ p i q i d y = q i ϕ p i q i .
Summing over cells proves the result.    □
Remark 1. 
Theorem 3 does not require uniformity inside a cell. It remains true whenever both p and q are dequantized with the same conditional density k i ( y ) supported on C i , because the factor k i cancels in the likelihood ratio. Uniformity is selected by entropy, not by divergence preservation.
  • Estimated weights and plug-in dequantization.
The exact right-inverse and Blackwell statements above concern the discrete law supplied to the dequantizer. When the population weights p are unknown, and an empirical vector p ^ is inserted instead, the exact guarantees hold conditionally for p ^ ; they do not erase the sampling error between p and p ^ . This distinction is especially simple for a common fixed cell-confined kernel K. Because the cells are disjoint,
TV ( T K p , T K p ^ ) = TV ( p , p ^ ) ,
and, under the same absolute-continuity conditions as Theorem 3, D ϕ ( T K p ∥ T K p ^ ) = D ϕ ( p ∥ p ^ ) . Thus, plug-in dequantization neither amplifies nor contracts the discrepancy caused by estimating the weights; it transports that discrepancy exactly into the continuous representation.
If p ^ is the empirical PMF from N independent draws from an n-category law p, then
E TV ( p , p ^ ) ≤ 1 2 n − 1 N , TV ( p , p ^ ) ≤ 1 2 n − 1 N + log ( 1 / δ ) 2 N
with probability at least 1 − δ . The expectation bound follows from Cauchy–Schwarz and E ∥ p ^ − p ∥ 2 2 = ( 1 − ∥ p ∥ 2 2 ) / N ; the deviation term follows from bounded differences because replacing one observation changes TV by at most 1 / N . For KL, an unsmoothed empirical PMF can assign zero mass to a population-positive category and, therefore, produce infinite KL; smoothing or a model that enforces positive category mass is required when a finite KL comparison is desired.
When the same finite dataset is used to estimate p, choose the dequantizer and fit the downstream model. The selection procedure itself can overfit even though the conditional dequantization map is exact. A practical implementation should therefore separate fitting and selection—for example, by a training/validation split, nested validation, or cross-fitting. The real-data experiments below use outer held-out folds and an inner validation subset for choosing the Gibbs concentration.

3.4. Downstream Approximation and the Categorical KL Decomposition

Blackwell equivalence describes the relaxation before an approximating continuous model is imposed. The next identity isolates the error created by such an approximation.  Let
f p ( y ) = ∑ i p i k i ( y ) , supp ( k i ) ⊆ C i ,
and let q be any density on the same domain. Define its induced category probabilities
p ^ i = ∫ C i q ( y ) d y ,
and, whenever p ^ i > 0 , its within-cell conditional density r i ( y ) = q ( y ) / p ^ i for y ∈ C i .
Theorem 4 
(Downstream KL decomposition). Assume p i > 0 implies p ^ i > 0 and k i ≪ r i . Then,
KL ( f p ∥ q ) = KL ( p ∥ p ^ ) + ∑ i = 1 n p i KL ( k i ∥ r i ) .
Consequently, KL ( p ∥ p ^ ) ≤ KL ( f p ∥ q ) . More generally, for every f-divergence,
D ϕ { p ∥ Q # q } ≤ D ϕ ( f p ∥ q ) ,
by data processing through Q.
Proof. 
On C i , f p = p i k i and q = p ^ i r i . Therefore,
∫ C i p i k i log p i k i p ^ i r i = p i log p i p ^ i + p i KL ( k i ∥ r i ) .
Summation proves (4). The f-divergence inequality is the data-processing inequality for the deterministic Markov kernel Q.    □
The coverage condition in Theorem 4 is substantive. For every target-positive cell, the fitted density must place positive mass on the cell, and its within-cell conditional density must dominate k i on the support of k i . If this fails, the relevant KL terms are + ∞ and the identity remains only an extended-valued equality, not a finite categorical-versus-shape decomposition. The Gaussian-mixture fits used below have positive density on all of R , and the bounded spline flows used in the audits have positive density throughout the modeled domain, so cell coverage is 100% in the experiments where the finite decomposition is interpreted. A truncated or compact-support model that leaves part of a target cell uncovered must instead be reported as a coverage failure.
The first term in (4) is the error in the discrete law recovered after fitting q; the second is a within-cell shape penalty that can consume finite model capacity without changing the target categories. Thus, two exact dequantizers can be statistically equivalent before approximation but differ substantially in how easy they are for a restricted class F to fit. For a family of exact kernels k λ , a natural model-aware criterion is
λ F * ∈ arg min λ D p , Q # q ^ λ , F , q ^ λ , F ≈ arg min q ∈ F KL ( f p , λ ∥ q ) ,
where D may be TV or categorical KL. Because p is the specified target law, this criterion can be evaluated directly. For parametric inference across unknown p θ , however, the dequantization kernel must remain θ -independent for Theorem 2 to apply.
Theorem 5 
(Finite-sample oracle inequality for Monte Carlo kernel selection). Let Λ be a finite set of M exact dequantizers and condition on fitted downstream densities { q λ : λ ∈ Λ } . Write r λ = Q # q λ and R ( λ ) = TV ( p , r λ ) . For each λ, generate m draws from q λ , requantize them, let r ^ λ , m be the resulting empirical categorical law, and define
λ ^ ∈ arg min λ ∈ Λ TV ( p , r ^ λ , m ) .
Then, for every 0 < δ < 1 , with probability at least 1 − δ over the Monte Carlo draws,
R ( λ ^ ) ≤ min λ ∈ Λ R ( λ ) + n − 1 m + 2 log ( M / δ ) 2 m .
If the categorical masses r λ ( C i ) are available analytically, the Monte Carlo term is unnecessary and the best candidate in Λ can be selected exactly conditional on the fitted models.
Proof. 
For fixed λ , the reverse triangle inequality provides | TV ( p , r ^ λ , m ) − TV ( p , r λ ) | ≤ TV ( r ^ λ , m , r λ ) . Moreover,
E TV ( r ^ λ , m , r λ ) ≤ 1 2 n − 1 m ,
because E ∥ r ^ − r ∥ 1 ≤ n E ∥ r ^ − r ∥ 2 2 and E ∥ r ^ − r ∥ 2 2 = ( 1 − ∥ r ∥ 2 2 ) / m ≤ ( n − 1 ) / ( n m ) . Changing one generated category changes TV by at most 1 / m , so McDiarmid’s inequality and a union bound over M candidates imply, with probability at least 1 − δ ,
sup λ ∈ Λ | TV ( p , r ^ λ , m ) − R ( λ ) | ≤ 1 2 n − 1 m + log ( M / δ ) 2 m .
The empirical minimizer is therefore within twice this uniform error of the oracle risk, proving (5).    □

3.5. Approximate Right Inverses and a Leakage Bound

Exact cell confinement is a sharp idealization for learned or globally smoothed dequantizers. Let an arbitrary kernel K i generate Y from X = x i , retain the same quantizer Q, and define the weighted category-mismatch probability
ϵ p = ∑ i = 1 n p i K i { y : Q ( y ) ≠ x i } .
Proposition 1 
(Requantization bound under category leakage). If P ˜ X = Q # P Y is the law obtained after dequantization and requantization, then
TV ( P X , P ˜ X ) ≤ ϵ p .
In particular, every exact partition-respecting dequantizer has ϵ p = 0 and zero requantization error.
Proof. 
Use the natural coupling that first draws X and then draws Y ∣ X . Under this coupling, P { Q ( Y ) ≠ X } = ϵ p . The coupling characterization of total variation yields the bound.    □
This inequality supplies a model-independent diagnostic for approximate dequantizers. It also separates two effects that are sometimes conflated: probability mass may leave a chosen bounded plotting domain without changing category after nearest-neighbor requantization, whereas crossing an internal quantization boundary contributes directly to  ϵ p .
A second perturbation question arises when the dequantizer is exact for one geometry but decoding uses perturbed sites or boundaries. Let Q ^ have cells C ^ i carrying the same labels as C i .
Proposition 2 
(Bound under quantizer-geometry perturbation). If Y ∣ X = x i ∼ K i with K i ( C i ) = 1 and p ^ is the law of Q ^ ( Y ) , then
TV ( p , p ^ ) ≤ ∑ i p i K i ( C i ∖ C ^ i ) ≤ ∑ i p i K i ( C i ▵ C ^ i ) .
For the one-dimensional uniform Voronoi kernel, if the ordered sites are perturbed to x ^ i with max i | x ^ i − x i | ≤ δ and their order is preserved, then
TV ( p , p ^ ) ≤ ∑ i p i min 1 , 2 δ Δ i .
Proof. 
Couple X , Y as before and decode Y using Q ^ . Conditional error in category i is K i ( C i ∖ C ^ i ) , so the first inequality follows from the coupling bound for TV. The second is immediate. In one dimension, every internal midpoint moves by at most δ ; hence, at most 2 δ of an original interval of width Δ i can be relabeled. Uniformity within the cell provides (6).    □

3.6. Maximum-Entropy and Entropy–Transport Exact Kernels

Consider all partition-respecting kernels k i supported on C i and integrating to one. Their mixture density is f ( y ) = ∑ i p i k i ( y ) .
Theorem 6 
(Maximum entropy among exact right inverses). Assume p i > 0 for all i. Among all partition-respecting dequantizers,
h ( Y ) ≤ H ( p ) + ∑ i = 1 n p i log Vol ( C i ) ,
with equality if and only if k i is uniform on C i almost everywhere for every i.
Proof. 
Because the cells are disjoint,
h ( Y ) = H ( p ) + ∑ i p i h ( k i ) .
For a density supported on a set of finite volume, h ( k i ) ≤ log Vol ( C i ) , with equality uniquely for the uniform density. Summing provides the result.    □
Maximum entropy is only one possible design principle. To control geometric displacement while retaining exact decoding, let c i : C i → [ 0 , ∞ ) be measurable and define, for  λ ≥ 0 ,   
k i , λ ( y ) = exp { − λ c i ( y ) } Z i ( λ ) 1 C i ( y ) , Z i ( λ ) = ∫ C i exp { − λ c i ( u ) } d u .
Theorem 7 
(Entropy–transport Gibbs frontier). Assume 0 < Z i ( λ ) < ∞ . For each fixed λ ≥ 0 , k i , λ is the unique maximizer over densities k supported on C i of
h ( k ) − λ E k { c i ( Y ) } .
Every mixture f p , λ = ∑ i p i k i , λ remains an exact right inverse and satisfies Theorems 2 and 3 whenever λ and c i do not depend on the unknown parameter. Writing D i ( λ ) = E k i , λ c i ( Y ) and H i ( λ ) = h ( k i , λ ) , wherever differentiation is justified,
D i ′ ( λ ) = − Var k i , λ { c i ( Y ) } ≤ 0 , H i ′ ( λ ) = − λ Var k i , λ { c i ( Y ) } ≤ 0 .
Thus, λ = 0 recovers the maximum-entropy uniform kernel, while increasing λ monotonically reduces expected cost at the price of entropy.
Proof. 
For any admissible k,
KL ( k ∥ k i , λ ) = − h ( k ) + λ E k c i + log Z i ( λ ) ≥ 0 ,
with equality only at k = k i , λ . The derivative identities follow from D i ( λ ) = − ∂ λ log Z i ( λ ) and H i ( λ ) = log Z i ( λ ) + λ D i ( λ ) .    □
For the one-dimensional experiments, we write c i ( y ) = { ( y − x i ) / s c } 2 . Then, k i , λ is a truncated Gaussian inside the same Voronoi cell; sampling is available by truncated-normal inversion. Importantly, λ and s c are not two separately identifiable tuning parameters, because
exp { − λ [ ( y − x i ) / s c ] 2 } = exp { − α ( y − x i ) 2 } , α = λ / s c 2 .
Hence, every pair ( λ , s c ) with the same α defines exactly the same conditional kernel. We retain s c = 0.5 in the multiscale experiments only as a reporting convention tied to the median local spacing; changing s c simply rescales the numerical value of λ quadratically. The Gibbs form itself is a standard consequence of entropy-penalized optimization, and the statistically relevant family is one-dimensional in the effective concentration α .
The uniform kernel is the least committal exact right inverse under a fixed partition. This characterization also clarifies a limitation: maximum entropy does not imply that the density has a useful score ∇ y log f ( y ) . In fact, the uniform density is flat inside each cell and discontinuous at cell boundaries. Its role is dequantized data for a separate differentiable model, not a smooth score model by itself.

4. One-Dimensional Voronoi Dequantization

Assume x 1 < ⋯ < x n . Internal nearest-neighbor boundaries are
b i = x i + x i + 1 2 , i = 1 , … , n − 1 .
The full Euclidean Voronoi cells of x 1 and x n are unbounded. A probability density with constant positive height on them would therefore be impossible. We use the finite mirrored endpoint convention
b 0 = x 1 − x 2 − x 1 2 , b n = x n + x n − x n − 1 2 .
This is equivalent to restricting the nearest-neighbor quantizer to the bounded interval Ω = [ b 0 , b n ] . It is a boundary convention, and should not be confused with the unrestricted Voronoi tessellation of R .
Let Δ i = b i − b i − 1 . Equation (1) becomes
f p ( y ) = ∑ i = 1 n p i Δ i 1 ( b i − 1 , b i ] ( y ) .
For y ∈ ( b k − 1 , b k ] ,
F p ( y ) = ∑ i < k p i + p k y − b k − 1 Δ k .
If all p i > 0 , the quantile is also explicit. Let P k = ∑ i ≤ k p i and P 0 = 0 . For  u ∈ ( P k − 1 , P k ] ,
F p − 1 ( u ) = b k − 1 + Δ k u − P k − 1 p k .
Zero-probability cells simply have zero density and are omitted when inverse-CDF sampling is implemented.

4.1. Exact Dequantization Displacement

Let ℓ i = x i − b i − 1 and r i = b i − x i . The continuous relaxation necessarily moves probability away from the atoms; for nearest-neighbor Voronoi cells, this displacement can be quantified exactly.
Proposition 3 
(Exact Wasserstein cost to the discrete law). For every r ≥ 1 , the nearest-neighbor coupling Y ↦ Q ( Y ) is an optimal transport from the Voronoi-uniform law T C p to P X , and
W r r ( T C p , P X ) = ∑ i = 1 n p i Δ i ℓ i r + 1 + r i r + 1 r + 1 .
More generally, on a bounded Voronoi partition in R d , the right side is
∑ i p i Vol ( C i ) ∫ C i ∥ y − x i ∥ r d y .
Proof. 
For every y, Q ( y ) is a nearest support point, so any coupling from y to the discrete support has conditional cost at least min j ∥ y − x j ∥ r = ∥ y − Q ( y ) ∥ r . The deterministic coupling induced by Q attains this pointwise lower bound and has target mass p i because P Y ( C i ) = p i . The one-dimensional formula follows by integrating | y − x i | r on the two sides of x i .    □
For r = 1 , the uniform kernel therefore has conditional mean displacement ( ℓ i 2 + r i 2 ) / ( 2 Δ i ) . This is a different question from the O ( h 2 ) reconstruction result below, which compares the dequantized law with an underlying smooth distribution whose exact bin masses generated p.

4.2. Special Case: Reconstruction from Exact Bin Masses of a Smooth Law

This subsection addresses a distinct histogram-reconstruction problem and is not a general statement about dequantizing an arbitrary specified discrete law. Suppose now that the probabilities are not arbitrary but arise by quantizing a continuous law. Let G have density g on [ a , b ] and let a = b 0 < ⋯ < b n = b be a one-dimensional partition with cell masses p i = G ( b i ) − G ( b i − 1 ) . The dequantized CDF in each cell is exactly the linear interpolant of G at the endpoints.
Theorem 8 
(Sharp W 1 rate for reconstruction from exact bin masses). If g ∈ C 1 [ a , b ] and sup | g ′ | ≤ M , with mesh h = max i Δ i , then
W 1 ( G , T C p ) = ∫ a b | G ( y ) − F p ( y ) | d y ≤ M 8 ( b − a ) h 2 .
The order h 2 cannot be improved uniformly over this class.
Proof. 
On [ b i − 1 , b i ] , F p is the linear interpolant of G. Since G ″ = g ′ and | G ″ | ≤ M , the standard interpolation remainder yields
sup [ b i − 1 , b i ] | G − F p | ≤ M Δ i 2 / 8 .
Integrating and summing provides the upper bound.
For sharpness, take [ a , b ] = [ 0 , 1 ] , an equal grid of width h = 1 / n , and 
g ϵ ( x ) = 1 + ϵ ( 2 x − 1 ) , 0 < ϵ < 1 .
Then, G ϵ ″ = 2 ϵ and, on each cell [ u , u + h ] , the chord error is exactly F p ( x ) − G ϵ ( x ) = ϵ ( x − u ) ( u + h − x ) . Hence,
W 1 = n ϵ ∫ 0 h t ( h − t ) d t = ϵ 6 h 2 .
Thus, the h 2 order is attained.    □
The assumption that p i are the exact bin masses of a smooth G is essential. The rate, therefore, does not apply automatically to an arbitrary empirical PMF such as those used in the Wine, Abalone, or RAND examples; those applications use the general exact cell-confined construction, not this reconstruction theorem.

4.3. A Continuous Exact-Mass Alternative

A continuous density may be desirable for visualization or for a downstream method that benefits from a non-flat input density. A simple universal construction is a triangular tent in every cell:  
t i ( y ) = 2 ( y − b i − 1 ) Δ i ( x i − b i − 1 ) , b i − 1 < y ≤ x i , 2 ( b i − y ) Δ i ( b i − x i ) , x i < y < b i , 0 , otherwise .
Then, f tent ( y ) = ∑ i p i t i ( y ) is continuous because every tent is zero at shared cell boundaries, and  ∫ C i f tent = p i exactly. It is not maximum entropy and is not claimed to share the sharp h 2 reconstruction rate. The same pointwise optimality argument as in Proposition 3 provides its mean absolute displacement explicitly:
W 1 ( f tent , P X ) = ∑ i p i ℓ i 2 + r i 2 3 Δ i ,
which is exactly two-thirds of the Voronoi-uniform displacement in Proposition 3. The two canonical kernels thus expose a transparent entropy–displacement trade-off: uniformity maximizes entropy, while the tent concentrates more mass near each atom and remains continuous.

5. Higher-Dimensional Formulation

For a finite set of sites in all of R d , ordinary Voronoi cells belonging to convex-hull sites are unbounded. A finite-volume probability construction must therefore specify an ambient domain. Let Ω ⊂ R d be bounded and let
C i = Ω ∩ { y : ∥ y − x i ∥ ≤ ∥ y − x j ∥ for all j } .
If every C i has positive d-dimensional volume, Equation (1) is a valid density and Theorems 1–7, Proposition 1, and the multidimensional transport identity in Proposition 3 apply without change. Taking Ω = conv { x i } is a parameter-free geometric choice when the sites span R d , but it encodes a boundary decision and is not universally appropriate. This explicit domain dependence corrects the misleading idea that a finite Euclidean Voronoi tessellation is automatically a collection of finite cells.

6. Experiments

The experiments are organized to test distinct claims: deterministic approximation, category leakage, downstream approximation capacity, model-aware selection under both Gaussian mixtures and spline normalizing flows, real irregular geometry, and the exact likelihood/information identities.

6.1. Sharp Approximation Rate

We bin two compactly supported continuous distributions on increasingly fine equal grids and provide the dequantizers with the exact cell masses. This isolates reconstruction from sampling noise. The first truth is Beta ( 2 , 5 ) on [ 0 , 1 ] ; the second is a two-component Gaussian mixture truncated to [ − 5 , 5 ] . We compare maximum-entropy Voronoi uniform dequantization, the continuous tent variant, and a weighted Gaussian KDE with Silverman bandwidth. The uniform method yields slopes of 1.999 and 2.017 , essentially the theoretical value 2. The tent method is about first order in these experiments, while the KDE bandwidth decreases too slowly relative to the deterministic mesh. Table 1 reports the fitted convergence slopes, while Figure 1 shows the full error curves.
Table 1. Empirical slope of log W 1 against log h . The Voronoi-uniform slopes agree with the sharp h 2  theorem.
Figure 1. 1-Wasserstein distance ( W 1 ) reconstruction error under deterministic mesh refinement. The left panel shows the Beta ( 2 , 5 ) distribution, and the right panel shows the truncated Gaussian mixture. Blue circles represent Voronoi-uniform dequantization, orange squares represent the continuous tent kernel, and green triangles represent Gaussian kernel density estimation (KDE). The red dashed line indicates the theoretical h 2 reference rate. Both axes use logarithmic scales.

6.2. Multi-Scale Non-Uniform Support

We next use 12 support points with local spacings 0.2 , 0.5 , and 5. We measure (i) total variation between the original PMF and the PMF induced by nearest-neighbour requantization on the full line, (ii) the analytic mismatch probability ϵ p in Proposition 1, (iii) probability mass placed outside the finite mirrored dequantization domain, and (iv) the fraction of that domain on which the density is positive. In addition to global-width baselines, we include a stronger inscribed local uniform baseline: each atom receives the largest symmetric interval centered at x i that stays inside its own Voronoi cell. It therefore preserves the PMF exactly, but leaves gaps. Table 2 summarizes the resulting requantization error, mismatch probability, outside-domain mass, and domain coverage across the competing relaxations.
Table 2. Multi-scale support. The analytic mismatch probability ϵ upper-bounds the observed requantized TV as predicted by Proposition 1. Exact right-inverse methods have both quantities equal to zero.
The experiment makes three points. First, exact recovery is not unique: any kernel confined to the assigned cell has it. Second, every approximate baseline satisfies the predicted inequality TV ≤ ϵ p . Third, maximum-entropy Voronoi uniformity is distinguished by filling each cell completely with no width choice, whereas the inscribed local uniform baseline leaves substantial gaps.

6.3. Downstream Model Capacity: No Favorable K Selected Post Hoc

The original version of this project used one 24-component Gaussian mixture and, therefore, risked conflating relaxation quality with model capacity. We instead fit mixtures with K ∈ { 4 , 8 , 12 , 24 } , using ten random seeds per method/capacity combination. The training set contains 1200 dequantized draws and each fitted model is evaluated by requantizing 10,000 generated draws. Table 3 summarizes the resulting TV errors across model capacities, and Figure 2 visualizes the corresponding capacity sensitivity.
Table 3. Mean (SD) requantized TV error over ten seeds. Low capacity can favor smoother but biased relaxation; as capacity increases, the exact right-inverse relaxations become best.
Figure 2. Capacity sensitivity of the downstream Gaussian-mixture model. The horizontal axis shows the number of Gaussian-mixture components (K), and the vertical axis shows total variation (TV) error after requantization. Blue, orange, green, and red lines represent Voronoi-uniform dequantization, continuous tent dequantization, fixed-width uniform dequantization ( w = 1 ), and Gaussian kernel density estimation (KDE), respectively. Circular markers indicate mean TV errors, and error bars represent standard deviations across ten random seeds. The ordering changes at low K, demonstrating the importance of reporting downstream model capacity.
At K = 4 , fixed w = 1 noise is easiest for the low-capacity smooth mixture to fit even though it is not probability preserving. At  K = 24 , Voronoi uniform has a mean TV of 0.038 compared with 0.051 for fixed-width dequantization and roughly 0.312 for KDE. The result is more informative than a single favorable capacity: exact representation and downstream approximation are separate design criteria.

6.4. Model-Aware Exact Kernels and the KL Decomposition

The capacity reversal above suggests that maximum entropy need not be the easiest exact representation for a restricted smooth model. We therefore apply Theorem 7 on the same 12-point support with c i ( y ) = { ( y − x i ) / 0.5 } 2 and λ ∈ { 0 , 0.25 , 1 , 4 } . Every value is an exact right inverse; only the within-cell concentration changes. For each K ∈ { 4 , 8 , 12 , 24 } , we again fit ten independent Gaussian mixtures and evaluate requantized TV on generated draws. Table 4 reports the capacity-specific performance of the exact Gibbs family, and Figure 3 displays the corresponding model-aware concentration patterns.
Table 4. Mean (SD) requantized TV for the exact Gibbs family. The optimal concentration depends on downstream capacity, even though every column is exactly information preserving before approximation.
Figure 3. Model-aware exact dequantization. All curves use the same cells and preserve the target PMF exactly before the Gaussian-mixture approximation; only the within-cell Gibbs concentration changes.
At K = 4 , moving from the maximum-entropy endpoint λ = 0 to λ = 1 reduces mean TV from about 0.132 to 0.045 ; excessive concentration at λ = 4 is slightly worse. At  K = 24 , a milder λ = 0.25 is best among the tested exact kernels, with mean TV about 0.036 . The preferred kernel therefore changes with approximation capacity rather than being universally uniform or maximally concentrated.
The same experiment also audits Theorem 4. For an eight-component fit to the λ = 0 target, continuous KL equals categorical KL plus the weighted within-cell KL to numerical quadrature precision; Table 5 reports this numerical decomposition.
Table 5. Numerical audit of the exact downstream KL decomposition.
The decomposition exposes why continuous NLL alone can be a misleading proxy for discrete fidelity: part of the fitting budget is spent matching within-cell shape, which disappears after requantization.
  • Geometric-scale reparameterization.
We examined s c ∈ { 0.25 , 0.5 , 1 , 2 } . Because the kernel depends on the parameters only through the effective concentration
α = λ / s c 2 ,
parameter pairs yielding the same value of α are mathematically equivalent. For example, α = 4 is represented by
( λ , s c ) = ( 0.25 , 0.25 ) , ( 1 , 0.5 ) , ( 4 , 1 ) , ( 16 , 2 ) .
Table 6 confirms this equivalence numerically on the multiscale support, with a maximum density difference of zero to machine precision. Thus, λ and s c are not separately identifiable from the kernel, and there is no unique optimal ( λ , s c ) pair. If the optimal effective concentration is α * , then the corresponding parameterization satisfies   
λ * ( s c ) = α * s c 2 .
Table 6. Scale reparameterization audit. Pairs with the same α = λ / s c 2 define the same exact Gibbs kernel, so s c changes the numerical labeling of concentration rather than adding an independent tuning dimension.
For example, at  s c = 0.5 , the Gaussian-mixture optima are λ = 1 for K = 4 , 8 and λ = 0.25 for K = 12 , 24 , corresponding to effective concentrations α = 4 and α = 1 , respectively.

6.5. Rational-Quadratic Spline-Flow Validation and Finite-Sample Selection

The Gaussian-mixture experiment deliberately varies a classical finite-capacity density family. We next ask whether the same model-aware effect persists for a modern invertible density model. We fit one-dimensional monotone rational-quadratic spline normalizing flows [21] on the same bounded dequantization domain. The implementation uses a uniform base on [ 0 , 1 ] , followed by an analytically invertible rational-quadratic spline and an affine map to the dequantization interval; hence, likelihoods, CDF values, sampling, and induced cell probabilities are all available exactly. Flow capacities use 8, 16, or 32 spline bins. For every capacity and seed, the same underlying discrete training sample is used across λ ∈ { 0 , 0.1 , 0.25 , 0.5 , 1 , 2 , 4 } ; only the exact within-cell dequantization changes. Five seeds are used per capacity–concentration pair. Table 7 summarizes the exact requantized TV errors across flow capacities and Gibbs concentrations, while Figure 4 visualizes the corresponding capacity-dependent pattern.
Table 7. Mean exact requantized TV for rational-quadratic spline flows. The best exact Gibbs concentration changes with flow capacity; at 32 spline bins, λ = 1 reduces the mean TV from 0.125 at the uniform endpoint to 0.073 .
Figure 4. Model-aware exact dequantization with a rational-quadratic spline flow. Each target kernel is an exact right inverse before downstream fitting; the plotted error is introduced only by the fitted flow.
The flow experiment reinforces two conclusions. First, uniform dequantization is not uniformly easiest to approximate: the mean TV optimum is λ = 0.25 with 8 spline bins, λ = 0.1 with 16 bins, and  λ = 1 with 32 bins. Second, selecting the representation by continuous fit error need not select the representation with the best categorical fidelity. Across the 15 capacity–seed fits, the  λ minimizing Monte Carlo continuous KL and the λ minimizing exact requantized TV disagree in 14 cases. This is the behavior predicted by Theorem 4: continuous KL includes the within-cell approximation term, whereas the downstream categorical target does not.
For black-box models whose cell probabilities cannot be integrated exactly, Algorithm 1 uses generated samples to estimate categorical risk. Table 8 audits Theorem 5 by repeatedly selecting among the seven fitted Gibbs candidates using m requantized generated draws per candidate. The corrected theorem has the required m − 1 / 2 Monte Carlo order. Its bound is intentionally distribution-free and conservative, but empirical regret decreases rapidly with the Monte Carlo budget even when the exact oracle candidate is not selected.
Algorithm 1 Model-aware exact dequantization
Given a fixed partition, target weights p, a finite tuning set Λ , and a downstream model class F : (i) for each λ ∈ Λ , generate training data from the exact kernel k λ and fit q ^ λ , F ; (ii) compute the induced categorical law Q # q ^ λ , F exactly when cell probabilities are available, or estimate it by generating from q ^ λ , F and requantizing; (iii) select the λ minimizing a categorical loss such as TV or KL; and (iv) optionally refit the chosen continuous model using all available simulation/training budget. The selection criterion is deliberately categorical: Theorem 4 shows that continuous fit error contains a within-cell term that disappears after requantization. When p is estimated from observations rather than specified externally, the candidate fitting and tuning steps should use data splitting or cross-fitting so that the same observations do not determine both the plug-in target and its reported out-of-sample performance.
Table 8. Monte Carlo model-aware selection for the fitted spline flows. “Oracle hit” is the frequency of selecting the exact-TV best candidate; regret is measured in true requantized TV. The final column is the finite-sample bound in (5) at δ = 0.05 .

6.6. Geometry Perturbation Audit

To test Proposition 2, we perturb each site of the multiscale support independently by at most δ , preserve the ordering, generate from the original Voronoi-uniform cells, and decode with the perturbed nearest-neighbor quantizer. Each row averages 200 perturbations. Table 9 compares the observed requantized TV, category-mismatch probability, and deterministic perturbation bound.
Table 9. Sensitivity to perturbed support geometry. Requantized TV is below the actual category-mismatch probability, which is itself below the deterministic bound in (6).
The bound is intentionally conservative but monotone and fully deterministic once p, the cell widths, and a site-error radius are specified. This provides a simple audit when support locations or embeddings are estimated rather than known exactly.

6.7. Real Irregular Support: Wine Proline Measurements

The scikit-learn Wine dataset (scikit-learn; https://scikit-learn.org/; accessed on 20 September 2026) contains 178 observations and is distributed with the package. We use the empirical distribution of its “proline” feature: 121 distinct observed values, with adjacent gaps ranging from 1 to 133 and a median gap of 7.5 . Treating the empirical mass function as the specified discrete target produces a real irregular-spacing example without inventing a synthetic geometry. The exact dequantization statements are conditional on these empirical weights; they do not remove the sampling uncertainty involved in estimating the population PMF from 178 observations. Table 10 reports the resulting category-mismatch and requantization diagnostics for this irregular-support example.
Table 10. Wine proline empirical PMF. The mismatch probability ϵ controls the requantized TV error; exact Voronoi/tent kernels have ϵ = 0 .
The exact methods have zero category mismatch by construction. For fixed-width and Gaussian relaxations, the analytic mismatch probability is reported beside the resulting requantized TV, making the coupling bound directly checkable on real irregular support. The outside-domain column is kept separate because leaving the finite mirrored plotting domain need not imply a category error after nearest-neighbor clipping. These metrics are not intended as a density-estimation contest: they measure faithfulness to the specified empirical discrete law.

6.8. Two Real Discrete-to-Continuous Distributional Applications

We finally evaluate the specific use case that motivates dequantization in practice: a genuinely discrete numeric variable is represented either by its original atoms or by an exact continuous relaxation, after which the same smooth downstream density model is fitted. We use two real datasets with meaningful numeric category geometry rather than arbitrary class codes. The UCI Abalone data contain 4177 specimens and an integer “Rings” variable, with ring count plus 1.5 used as an age proxy [22]. The RAND Health Insurance Experiment subset contains 20,190 observations; we use mdvis, the number of annual outpatient visits to a physician [23,24].
For each dataset, we fix the integer lattice over the published observed range and use its half-integer Voronoi cells. Five outer folds are used for evaluation; within each training fold, 20% of the remaining observations are reserved for selecting the exact Gibbs concentration λ ∈ { 0 , 0.25 , 1 , 4 } by categorical validation NLL. Thus, the test observations are not used to estimate the training PMF, select the kernel, or fit the continuous model. Every continuous fit uses the same 32-bin rational-quadratic spline flow. The matched “raw discrete” comparator fits that flow directly to the integer atoms without dequantization. After fitting, flow probabilities are integrated exactly over the Voronoi cells, so all reported test criteria are again discrete: held-out categorical NLL and total variation from the held-out empirical PMF. A Jeffreys-smoothed categorical PMF is included as a direct discrete reference, not as a model we expect dequantization to dominate. Table 11 summarizes held-out NLL and TV across the five outer folds for both applications.
Table 11. Two real discrete-to-continuous applications. Entries are shown as the mean (SD across five held-out folds). The proposed continuous representation is compared with fitting the identical spline flow directly to the original atoms.
The representation effect is consistent in every fold. For Abalone rings, model-aware dequantization lowers mean held-out NLL from 2.558 to 2.506 and TV from 0.140 to 0.065; the dequantized flow wins all five folds on both criteria. For RAND physician visits, NLL falls from 2.237 to 2.189 and TV from 0.091 to 0.044, again with five of five fold-wise wins. Thus, the continuous relaxation cuts the PMF approximation error of the matched smooth flow by roughly one half in both applications. Uniform Voronoi dequantization already captures most of this gain, while validation-selected Gibbs concentration yields a small additional improvement.
The categorical reference is intentionally informative about what this result does not claim. In the Abalone data, the selected continuous flow essentially matches the direct categorical reference (NLL 2.506 versus 2.505; TV 0.065 versus 0.065), while in the much larger RAND sample the unrestricted categorical reference remains better. This is consistent with the Blackwell result: dequantization cannot create information. Its benefit is computational and representational—when a smooth continuous model is to be used, an exact cell-respecting relaxation is substantially easier to fit faithfully than the original atomic sample.
These results also clarify when the within-cell kernel can matter practically. The strongest evidence in the present experiments is for downstream capacity: the preferred effective concentration changes with Gaussian-mixture or spline-flow capacity, and the raw-atom singularity is especially difficult for a smooth finite-capacity density model. Irregular cell widths and highly uneven masses can make the continuous target more heterogeneous and may therefore magnify representation effects, but the present experiments do not isolate a monotone causal relationship with spacing irregularity or mass imbalance. We therefore avoid claiming such a law without a dedicated factorial study.
Figure 5 and Figure 6 show, respectively, the averaged PMF reconstructions and the fold-wise gains from exact dequantization relative to fitting the same spline flow directly to the atoms.
Figure 5. Real-data PMF reconstruction. Curves average the five outer-fold fits. Exact dequantization allows the same spline flow to track the discrete target substantially better than fitting the raw atoms directly. The RAND panel is shown through 30 visits for readability.
Figure 6. Fold-wise gain from transforming the discrete observations to the proposed exact continuous representation before fitting the spline flow. Positive values mean the dequantized representation has lower test error. All ten dataset–fold comparisons are positive for both NLL and TV.

6.9. Likelihood-Equivalence Audit

To verify the finite-sample likelihood statement in Theorem 2, we use a one-parameter softmax family on five irregular support points, draw one discrete sample of size 400, and dequantize that same sample with either the Voronoi-uniform or continuous-tent kernel. For each kernel we optimize the continuous likelihood and compare its centered likelihood curve with the original discrete likelihood. Table 12 reports the resulting agreement in MLEs and centered log-likelihood values.
Table 12. Numerical audit of likelihood preservation. The continuous likelihood differs from the discrete likelihood only by a parameter-free kernel term.
The MLEs agree to numerical optimization tolerance, and the centered log-likelihood curves agree to floating-point precision. This experiment is intentionally simple: it checks the implementation of an exact algebraic identity rather than supplying evidence for an asymptotic claim.

6.10. Numerical Checks of Exact Identities and Multidimensional Mass

Table 13 numerically integrates three divergences for two random weight vectors on the same irregular support. Errors are below 3 × 10 − 6 on a million-point integration grid, consistent with the exact identity in Theorem 3.
Table 13. Numerical verification of exact divergence preservation.
In two dimensions, we use Ω = conv { x i } for 31 irregular support points. The clipped Voronoi cells cover the convex hull to numerical precision (total area 148.1506935 ), and ∑ i ( p i / | C i | ) | C i | = 1 to machine precision. Figure 7 shows the bounded-domain construction.
Figure 7. Two-dimensional weighted Voronoi dequantization on the explicit bounded domain Ω = conv { x i } . The black outline marks the convex-hull boundary, the white lines delineate the internal Voronoi cells, red dots mark the support sites, and the color scale represents the density.

6.11. Estimated-Weight Plug-In Audit

Equation (3) is distribution-free. To illustrate its scale, Table 14 repeatedly estimates the 12-category synthetic PMF used in the model-aware experiments. The empirical TV decreases at the expected N − 1 / 2 rate, and the stated high-probability bound is conservative. By Equation (2), the same TV values apply exactly after dequantization with any common fixed cell-confined kernel.
Table 14. Finite-sample plug-in audit for an empirically estimated 12-category PMF (5000 repetitions). The final column is the fraction of repetitions satisfying the 1 − δ bound with δ = 0.05 .

6.12. Computational Scope of the Current Implementation

The closed-form one-dimensional construction is inexpensive because only adjacent support points are needed once the support is ordered. The present multidimensional implementation explicitly constructs a two-dimensional Voronoi diagram and clips its cells to the convex hull; that is, a computational-geometry implementation, not a claim that explicit clipped-cell enumeration remains practical in high dimension. Table 15 reports a machine-specific timing audit included for order-of-magnitude guidance rather than as a hardware-independent benchmark. In this environment, one-dimensional boundary construction remains negligible through 10 6 support points, while the bounded two-dimensional implementation remains sub-second through 1000 support points. The reproducibility code does not implement explicit bounded-cell enumeration above two dimensions. In higher dimensions, Voronoi complexity can grow rapidly and practical use would require specialized polytope libraries, approximate/implicit nearest-neighbor representations, or sampling methods that avoid enumerating every cell.
Table 15. Implementation scaling audit (median of five runs). Times are machine-specific and are reported only to indicate the computational scope of the supplied code.

7. Discussion

The central formula in this paper overlaps with weighted Voronoi density constructions; pretending otherwise would weaken the contribution. The useful distinction is the object being represented. For density estimation, the cells arise from samples, and cell volume is used to estimate an unknown continuous density. Here, the finite weight vector is already the probability law of interest, and the continuous density is required to be an invertible relaxation under a specified quantizer.
This perspective yields several practical consequences. Exact requantization is guaranteed rather than assessed approximately. More strongly, Theorem 2 shows that when the parameter enters only through the discrete weights, dequantization preserves the entire statistical experiment before a downstream approximation is imposed. This is a statement about equivalence of statistical experiments, not a guarantee about an arbitrary model subsequently fitted to the continuous outputs. Theorem 4 identifies exactly where that later approximation error enters: categorical error and within-cell shape error add in KL when the fitted model covers the target cells. This resolves the apparent tension in the capacity experiments. Exact kernels carry the same discrete information, yet a restricted continuous model can approximate them very differently. The Gibbs family makes that representation choice explicit instead of treating uniform noise as inevitable. The spline-flow experiment shows that this is not peculiar to Gaussian mixtures: the preferred exact kernel changes with flow capacity, and continuous KL and categorical fidelity select different concentrations in 14 of 15 seed-wise comparisons. The finite-sample oracle inequality then makes the selection step operational even when the induced category probabilities of a fitted continuous model are available only by simulation. The two additional real-data studies in Section 6.8 show the same representation effect on natural count variables: under a matched spline-flow model, exact continuousization improves held-out requantized NLL and TV in every fold for both Abalone ring counts and RAND physician visits.
A related distinction appears in a recent stochastic-modeling study, where mechanistic model construction is evaluated separately from out-of-family numerical and stochastic verification [25]. Here, the analogous separation is between an exact information-preserving representation and the approximation properties of the continuous model fitted afterward.
The limitations are equally important. First, the geometric support itself must be meaningful; categorical labels with arbitrary numeric codes should not be assigned Euclidean Voronoi cells without a defensible embedding. Second, high-dimensional Voronoi construction can be expensive, as emphasized by the VDE literature [11,12]; the supplied implementation is deliberately closed-form in one dimension and explicit only in bounded two-dimensional geometry. Third, a bounded ambient domain is part of the multidimensional model. Fourth, the piecewise-uniform density is not differentiable at cell boundaries and has zero score within cells, so it should not be presented as a ready-made score model. Fifth, the finite KL decomposition requires downstream cell coverage; truncated models that leave target-positive regions uncovered yield infinite within-cell KL rather than a finite categorical/shape split. Finally, model-aware tuning is not free of design choices. If the weights are estimated from data, exactness is conditional on the plug-in vector, population error remains governed by p − p ^ , and data splitting or cross-fitting is appropriate when kernel selection and downstream fitting are data-driven. If the weights themselves are an unknown parameter in an inferential model, allowing the kernel to depend on that parameter would invalidate the parameter-independent likelihood-preservation statement. Geometry-perturbation bounds also require a meaningful labeling correspondence between original and perturbed cells. The selector guarantee in Theorem 5 is conditional on the fitted downstream models and controls only Monte Carlo selection error; it is not a generalization bound for the training procedure itself.

8. Conclusions

We developed an information-preserving and model-aware theory for dequantizing weighted discrete laws on prescribed partitions. Exact cell-confined kernels are stochastic right inverses and produce a Blackwell-equivalent continuous experiment; inference and information are unchanged until a downstream approximation is imposed. At that point, an exact KL chain rule separates the discrete error that matters after requantization from within-cell shape error. This distinction leads to an entropy–transport Gibbs family of exact kernels: uniform dequantization is one endpoint, while concentration can be tuned to the capacity of the downstream model without sacrificing exact category recovery. Leakage and geometry-perturbation bounds quantify departures from the ideal construction, and Voronoi geometry supplies exact transport costs and a sharp Θ ( h 2 ) reconstruction result. A finite-sample selector provides an oracle guarantee for choosing among candidate exact kernels when categorical masses are estimated by simulation, and rational-quadratic spline-flow experiments confirm that the capacity-aware effect survives beyond Gaussian mixtures. Two real count-data applications further show that, when the same smooth flow is mandated, exact continuousization can reduce requantized PMF error by roughly one half relative to fitting that flow directly to the original atoms. The unrestricted categorical reference matches or outperforms the dequantized pipeline, as expected: the contribution is an information-preserving continuous representation for continuous-model workflows, not a replacement for direct categorical modeling. The resulting framework shifts the question from “which noise is convenient?” to “which information-preserving continuous representation is easiest for the model class that will actually be used?”

Author Contributions

Conceptualization, M.M., M.F.A. and A.S.E.; methodology, K.A.E.G. and M.F.A.; software, K.A.E.G.; validation, M.M., K.A.E.G., H.M.H. and M.F.A.; formal analysis, M.F.A. and A.S.E.; investigation, K.A.E.G. and A.S.E.; resources, H.M.H. and A.S.E.; data curation, H.M.H.; writing—original draft preparation, K.A.E.G. and A.S.E.; writing—review and editing, M.M.; visualization, M.M.; supervision, M.M.; project administration, M.M. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding. The article processing charge (APC) was funded by Qassim University (QU-APC-2026).

Data Availability Statement

The Abalone dataset is publicly available from the UCI Machine Learning Repository (https://archive.ics.uci.edu/dataset/1/abalone, accessed on 15 August 2026). The RAND Health Insurance Experiment dataset is available through the statsmodels dataset collection (https://www.statsmodels.org/stable/datasets/generated/randhie.html, accessed on 15 August 2026). The analysis code and additional numerical results supporting this study are available from the corresponding author upon reasonable request.

Acknowledgments

The Researchers would like to thank the Deanship of Graduate Studies and Scientific Research at Qassim University (www.qu.edu.sa) for financial support (QU-APC-2026).

Conflicts of Interest

The authors declare no conflicts of interest. The APC funder had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

References

  1. Uria, B.; Murray, I.; Larochelle, H. RNADE: The real-valued neural autoregressive density-estimator. In Advances in Neural Information Processing Systems 26; Curran Associates, Inc.: Red Hook, NY, USA, 2013; Available online: https://proceedings.neurips.cc/paper/2013/hash/53adaf494dc89ef7196d73636eb2451b-Abstract.html (accessed on 20 September 2026).
  2. Theis, L.; van den Oord, A.; Bethge, M. A Note on the Evaluation of Generative Models. In Proceedings of the International Conference on Learning Representations, San Juan, Puerto Rico, 2–4 May 2016; Available online: https://arxiv.org/abs/1511.01844 (accessed on 20 September 2026).
  3. Dinh, L.; Sohl-Dickstein, J.; Bengio, S. Density Estimation Using Real NVP. In Proceedings of the International Conference on Learning Representations, Toulon, France, 24–26 April 2017; Available online: https://openreview.net/forum?id=HkpbnH9lx (accessed on 20 September 2026).
  4. Ho, J.; Chen, X.; Srinivas, A.; Duan, Y.; Abbeel, P. Flow++: Improving Flow-Based Generative Models with Variational Dequantization and Architecture Design. In Proceedings of the International Conference on Machine Learning, Long Beach, CA, USA, 9–15 June 2019; Volume 97. Available online: https://proceedings.mlr.press/v97/ho19a.html (accessed on 20 September 2026).
  5. Hoogeboom, E.; Cohen, T.S.; Tomczak, J.M. Learning Discrete Distributions by Dequantization. arXiv 2020, arXiv:2001.11235. [Google Scholar] [CrossRef] [Scilit]
  6. Gad, K.A.E.; Atia, F.A.; Abouelenein, M.F.; Hamouda, H.M.; Moussa, M.; Elshafae, S. Conformable Calculus Derivation of the Power Function Distribution: Reparameterization and Applications. Next Mater. 2026, 13, 102790. [Google Scholar] [CrossRef] [Scilit]
  7. Gad, K.A.E.; Atia, F.A.; Moussa, M.; Hamouda, H.M.; Abouelenein, M.F.; Abdul-Moniem, I.B.; Eldeeb, A.S. The Fractional Exponential Distribution: A Gamma Subfamily from Conformable Calculus. Results Eng. 2026, 31, 111224. [Google Scholar] [CrossRef] [Scilit]
  8. Gad, K.A.E.; Atia, F.A.; Moussa, M.; Abouelenein, M.F.; Hamouda, H.M.; Abdul-Moniem, I.B.; Eldeeb, A.S. The Exponentiated New Failure Distribution: Theory and Applications. Results Eng. 2026, 32, 111907. [Google Scholar] [CrossRef] [Scilit]
  9. Duah, K.; Sun, Y.; Moussa, M. A Random Quantile Approach for Prior Covariance Estimation in Bayesian Maximum Entropy. Spat. Stat. 2026, 73, 100974. [Google Scholar] [CrossRef] [Scilit]
  10. Grzegorzewski, P.; Moussa, M.; Sun, Y.; Elshafae, S. Uncertainty-Aware Repeated-Measures ANOVA for Interval-Valued Data with an Application to Longitudinal CT Tumor Imaging. In Information Processing and Management of Uncertainty in Knowledge-Based Systems; Communications in Computer and Information Science; Springer: Cham, Switzerland, 2026; Volume 3019, pp. 477–491. [Google Scholar] [CrossRef] [Scilit]
  11. Polianskii, V.; Marchetti, G.L.; Kravberg, A.; Varava, A.; Pokorny, F.T.; Kragic, D. Voronoi Density Estimator for High-Dimensional Data: Computation, Compactification and Convergence. In Proceedings of the Conference on Uncertainty in Artificial Intelligence, Eindhoven, The Netherlands, 2–4 August 2022; Volume 180, pp. 1644–1653. Available online: https://proceedings.mlr.press/v180/polianskii22a.html (accessed on 20 September 2026).
  12. Marchetti, G.L.; Polianskii, V.; Varava, A.; Pokorny, F.T.; Kragic, D. An Efficient and Continuous Voronoi Density Estimator. In Proceedings of the International Conference on Artificial Intelligence and Statistics, Valencia, Spain, 25–27 April 2023; Volume 206, pp. 4732–4744. Available online: https://proceedings.mlr.press/v206/marchetti23a.html (accessed on 20 September 2026).
  13. Chen, R.T.Q.; Amos, B.; Nickel, M. Semi-Discrete Normalizing Flows through Differentiable Tessellation. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2022; Volume 35. [Google Scholar] [CrossRef] [Scilit]
  14. Blackwell, D. Comparison of Experiments. In Proceedings of the Second Berkeley Symposium on Mathematical Statistics and Probability, Berkeley, CA, USA, 31 July–12 August 1950; pp. 93–102. Available online: https://projecteuclid.org/euclid.bsmsp/120050022 (accessed on 20 September 2026).
  15. Blackwell, D. Equivalent Comparisons of Experiments. Ann. Math. Stat. 1953, 24, 265–272. [Google Scholar] [CrossRef] [Scilit]
  16. Moussa, M.; Ratterman, C.; Zhang, W.; Zhe, S.; Sun, Y. DeepCut: Adaptive Neural Network Thresholds for Precipitation Phase Partitioning. Mach. Learn. Earth 2026, 2, 015009. [Google Scholar] [CrossRef] [Scilit]
  17. Moussa, M. Auditing Conformal Prediction under Distribution Shift: A Detectability Boundary, Exact Label-Budget Design, and Repair. Res. Sq. 2026. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Moussa, M.A.; Gad, K.A.E.; Sharabati, W.K.; Hamouda, H.M. Risk-Guided Maintenance of Sports AI Under Seasonal Shift: Empirical-Bayes Recalibration with Anytime-Valid Certification. Preprints 2026. [Google Scholar] [CrossRef] [Scilit]
  19. Nielsen, D.; Winther, O. Closing the Dequantization Gap: PixelCNN as a Single-Layer Flow. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, Available online: https://proceedings.neurips.cc/paper/2020/hash/4c5bcfec8584af0d967f1ab10179ca4b-Abstract.html (accessed on 20 September 2026).
  20. Tan, S.; Huang, C.W.; Sordoni, A.; Courville, A. Learning to Dequantise with Truncated Flows. In Proceedings of the International Conference on Learning Representations, Virtual Event, 25–29 April 2022; Available online: https://openreview.net/forum?id=fExcSKdDo_ (accessed on 20 September 2026).
  21. Durkan, C.; Bekasov, A.; Murray, I.; Papamakarios, G. Neural Spline Flows. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2019; Volume 32, Available online: https://proceedings.neurips.cc/paper/2019/hash/7ac71d433f282034e088473244df8c02-Abstract.html (accessed on 20 September 2026).
  22. Nash, W.; Sellers, T.; Talbot, S.; Cawthorn, A.; Ford, W. Abalone. Dataset, UCI Machine Learning Repository. 1994. Available online: https://archive.ics.uci.edu/dataset/1/abalone (accessed on 20 September 2026).
  23. Cameron, A.C.; Trivedi, P.K. Microeconometrics: Methods and Applications; Cambridge University Press: Cambridge, UK, 2005. [Google Scholar] [CrossRef] [Scilit]
  24. Seabold, S.; Perktold, J. Statsmodels: Econometric and Statistical Modeling with Python. In Proceedings of the 9th Python in Science Conference, Austin, TX, USA, 28 June–3 July 2010; pp. 92–96. Available online: https://conference.scipy.org/proceedings/scipy2010/seabold.html (accessed on 20 September 2026).
  25. Moussa, M.; Youssef, S.; Elrayany, K.; Nafea, A. Mechanism-Resolved Fractional–Volterra Modeling for Thermal-Hazard Screening: Lithium-Ion Battery Intervention Evidence, PDE Transfer, and Stochastic Verification. SSRN Preprint, 2026. SSRN 7322883. Available online: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7322883 (accessed on 20 September 2026).
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.