1. Introduction
In many applied fields, including engineering, environmental sciences, economics, and biomedicine, the variable of interest is a rate, fraction, or proportion bounded in the unit interval
. Typical examples include conversion rates in online marketing, defect fractions in quality control, success proportions in clinical trials, and coverage or usage ratios in environmental and ecological studies. For this type of data, the beta distribution has become a standard reference model because of its flexibility in terms of skewness and kurtosis, as well as the availability of well-established inferential and regression methodologies, such as beta regression for modeling rates and proportions [
1,
2].
The Kumaraswamy distribution is an important alternative to the beta distribution for modeling bounded random variables. It was originally introduced in the context of double-bounded random processes [
3]. Its cumulative distribution and quantile functions are available in closed form, which facilitates random variate generation as well as estimation by maximum likelihood and moments. In the last decade, both beta and Kumaraswamy distributions have often been used as baseline models to construct more general families via different generator schemes, leading to a broad class of unit distributions with one or more additional shape parameters.
In parallel, the literature on general methods for generating new families of continuous distributions has grown considerably. Influential contributions include the Marshall–Olkin approach, which adds a parameter at the level of the survival function [
4], the T-X method based on transforming a baseline variable through a generator function [
5], and the so-called “odds” or “odds-
x” families, which rely on transformations of the form
[
6] and related odds-based constructions [
7]. Recent reviews systematize these approaches in terms of transformation-based generators and compounding mechanisms, highlighting the central role of beta-generated and T-X-type constructions in building flexible univariate families [
8].
Despite these advances, in the specific context of unit distributions, many of the available models are based on baseline distributions with relatively light tails, such as the normal, gamma, or Weibull distributions, or on direct extensions of the beta and Kumaraswamy laws. This may limit their ability to capture heavy-tailed behavior or complex risk structures when the data are naturally linked, through a monotonic transformation, to heavy-tailed positive variables. In many applications, a proportion can be interpreted as a monotonic transform of a positive variable Y, for example, when Y represents failure times, traffic intensities, or claim sizes, and Z is a ratio-type quantity built from these latent variables.
The Lomax distribution (also known as Pareto type II) is a classical model for non-negative data with heavy tails, widely used in reliability, insurance, traffic analysis, and economics [
9]. Its cumulative distribution function on
exhibits a power-law tail and a decreasing hazard rate, which makes it suitable for describing survival phenomena with “wear-out” or cleansing effects. Building on this baseline, an extensive family of Lomax-based extensions has been proposed in the literature, such as the extended Lomax distribution [
10] and the Weibull–Lomax distribution [
11], among others. These models display a rich variety of density and hazard shapes, including increasing, decreasing, unimodal, and bathtub-shaped hazards.
The explicit incorporation of this kind of heavy-tail behavior into models with support on has received comparatively less attention. From a conceptual perspective, it is natural to consider unit distributions obtained by applying the inverse logit or other monotonic mappings from to for Lomax-distributed variables. This strategy allows the model to inherit part of the heavy-tail structure of the baseline Lomax distribution on the odds scale while producing a distribution that is directly interpretable in terms of rates or proportions.
Related unit-interval models based on Lomax-type baselines have recently appeared in the literature, including exponentiated and regression-oriented extensions on
[
12] and the unit ratio-extended Weibull family of [
13], of which the present model is a particular case. In addition, several recent contributions have proposed transformed or unit-interval models for proportional data, including unit Burr-XII and quantile-regression formulations [
14], alternative unit-Lindley-type models [
15], unit exponential and related transformed models [
16], new bounded and power-unit constructions [
17,
18], and recent Chen- and Weibull-type unit models [
19,
20,
21]. This active literature makes it important to distinguish between algebraic reformulations, genuinely different generators, and equivalent constructions obtained through monotone maps.
Motivated by this point, we study a bounded model with a cumulative distribution function
which arises from the odds transformation
with
. This representation is not merely a reparametrization device; it provides a transparent structural link between heavy-tail behavior on the positive half-line and interpretable shape behavior on the unit interval while preserving analytic tractability.
The contribution of this paper is primarily structural, inferential, and computational, and it should be interpreted as a complement to previously reported unit-Lomax-type constructions rather than as the introduction of a genuinely new distributional family. More precisely, the contribution is threefold. First, we establish explicitly the equivalence between the present odds-scale formulation and a previously studied unit-Lomax-type construction, up to complementation and reparametrization, thereby clarifying issues of nomenclature and avoiding an overstatement of novelty. Second, we emphasize the representation , , which gives the model a direct odds-scale interpretation in terms of a latent positive heavy-tailed variable. Third, we study the inferential consequences of this representation. In particular, we show that the likelihood for the unit-interval model is equivalent to the classical Lomax likelihood applied to the odds-transformed observations . This observation explains why maximum likelihood estimation may become unstable in small samples or under heavy-tailed parameter configurations. To address this issue, the revised inferential framework includes maximum product of spacings estimation as a more stable alternative in challenging finite-sample settings.
Accordingly, the paper should be viewed as a detailed structural, inferential, and computational reassessment of an odds-scale representation of a Lomax-based unit model. This perspective is especially relevant when bounded observations are naturally interpreted through latent positive mechanisms with heavy-tailed behavior on the odds scale.
The remainder of the paper is organized as follows.
Section 2 presents the model, its equivalence with previous unit-Lomax-type constructions, and its main structural properties.
Section 3 develops likelihood-based inference, diagnoses finite-sample maximum likelihood estimator instability, and introduces maximum product of spacings estimation.
Section 4 describes random variate generation.
Section 5 reports the Monte Carlo study comparing the maximum likelihood estimator and MPS.
Section 6 presents the data applications and model comparisons.
Section 7 concludes the paper.
2. Structural Properties and Tractability
In this section, we record the structural facts that are used later in the inferential and computational analysis. Since the model is generated by the monotone transformation
, the cdf, pdf, survival, hazard, and quantile functions follow directly from the Lomax baseline and are therefore stated in compact form. Greater attention is given to endpoint behavior, hazard shapes, and modality since these properties are more informative for applications and for the finite-sample inference discussed in
Section 3.
2.1. Definition, Density and Hazard Functions
Definition 1. Let Y be a Lomax (Pareto type II) random variable with shape parameter and scale parameter , that is,with the corresponding probability density function Consider the monotonic transformation
whose inverse is
with Jacobian
A random variable
Z is said to follow an odds-Lomax distribution with parameters
, denoted by
, if it admits the stochastic representation (
3) with
as in (
1). By construction,
Z takes values in the open unit interval
.
Proposition 1. Let U be a random variable on with a distribution functionfor some and . Define . Then Z has the cdfwith and . Hence, the odds-Lomax formulation is equivalent, up to simple complementary transformation and reparametrization, to the unit-Lomax-type construction considered in [22]. Proof. For
,
Substituting into
T gives
Setting
and
yields the stated form. □
Remark 1. Proposition 1 shows that the present formulation is best understood as an odds-scale representation of a Lomax-based unit-interval model. The terminology odds-Lomax is adopted here to emphasize the transformation and to avoid confusion with other non-equivalent families that have also been described in the literature under the label “unit-Lomax”.
Remark 2. In addition, the present model is a particular case of the unit ratio-extended Weibull family introduced by [13], with the present formulation corresponding to the Lomax-based odds-scale case within that family. The present paper should therefore be understood not as introducing a new distributional family, but as providing a detailed structural, inferential, and computational study of this specific odds-scale Lomax formulation, with particular emphasis on the MLE instability mechanism and the role of MPS as a robust alternative. Proposition 2. If , then for , These expressions follow directly from the monotone transformation , its inverse , and the change-of-variables formula.
The odds-Lomax model has parameter-independent support
, and for
and
, its density (
6) is continuously differentiable in
z on
. These properties support the standard maximum likelihood regularity conditions used in
Section 3.
Figure 1 displays representative shapes of the odds-Lomax density for selected combinations of
, highlighting the variety of forms that the model can exhibit on the unit interval, including decreasing, increasing, unimodal, and U-shaped patterns.
2.2. Limiting Behavior and Shape Properties
The behavior of the density and hazard functions near the boundaries and , as well as the main shapes that can arise for the hazard function, is examined next.
Proposition 3. Let . ThenIn particular,whereas the survival function decays to zero at the polynomial rate as . These limits follow directly from the explicit forms of the density and survival functions in (
6) and (
7) by evaluating the leading terms as
and
.
Proposition 3 shows that the odds-Lomax density is always finite and strictly positive at , while its behavior near depends on the shape parameter . For , the density vanishes at 1, whereas for , it diverges as . This allows the model to accommodate both relatively light right tails on the probability scale (for large ) and highly concentrated mass near (for small ), which is particularly useful when modeling proportions close to one.
The tail representation (
11) also highlights the role of
as a tail index on the odds scale: the survival function of
Z inherits the polynomial decay of the Lomax baseline under the transformation
, with
controlling the order of decay and
acting as a scale parameter.
Proposition 4. For , the hazard function (8) satisfiesand its derivative is given byConsequently: - (i)
If , then for all , so the hazard is strictly increasing.
- (ii)
If , then there exists a uniquesuch that for and for , so the hazard is strictly decreasing on and strictly increasing on , displaying a bathtub-shaped hazard.
The result follows by differentiating the closed-form hazard function in (
8) and analyzing the sign of the resulting affine numerator.
Remark 3. Proposition 4 shows that the odds-Lomax distribution can produce two qualitatively distinct hazard shapes on . When , the hazard function is strictly increasing, which may be appropriate for processes where the instantaneous risk grows as the proportion approaches one. In contrast, for , the hazard is bathtub-shaped, reflecting an initial decrease in risk followed by an increasing phase, a pattern often encountered in reliability and survival applications. Interestingly, these shapes are controlled entirely by the scale parameter β, whereas the shape parameter α acts as a vertical scaling factor for the hazard function.
2.3. Quantiles and Modes
In this subsection, the quantile function of the odds-Lomax distribution is derived, and the location of its mode (or modes) is characterized as a function of the parameters .
Proposition 5. Let with cdf given by (5). Then the quantile function , defined by , isIn particular, the median m of Z is given by . Equation (
15) follows from the direct inversion of the cdf and is retained because it provides the inversion sampler used in
Section 4.
Attention now turns to the mode. Let
denote the density (
6) and consider the log-density
Differentiating with respect to
z,
The stationary points of the density are obtained by solving . The next result characterizes the modality of the odds-Lomax density, including the existence of an interior stationary point and the boundary-mode behavior. Since the support is the open interval , statements such as “mode at ” are understood in the limiting sense as (analogously for ).
Proposition 6. Let with and .
- (i)
If , then with density In this case, the density is
- (a)
Strictly decreasing (L-shaped) for ;
- (b)
Constant (uniform) for ; and
- (c)
Strictly increasing (J-shaped) for .
- (ii)
If , then the log-density (16) admits at most one stationary point in , given by Moreover, the following cases hold:
- (a)
If and , then , and is unimodal, with a unique interior mode at .
- (b)
If and , then is strictly increasing on , with (unbounded if and finite if ).
- (c)
If and , then is strictly decreasing on , with .
- (d)
If and , then and has a unique interior minimum at , decreasing on and increasing on .
The modality classification follows from differentiating the log-density in (
16), solving
, and combining the sign of
with the endpoint behavior in Proposition 3. The odds-Lomax family can therefore accommodate a wide variety of density shapes on the unit interval, including L- and J-shaped, unimodal, and U-shaped patterns, with the parameters
and
exerting distinct control over tail behavior and modality.
2.4. Moments and Entropies
This section derives expressions for the raw moments of the distribution using its Lomax-based stochastic representation and discusses the existence conditions and some consequences for moment-based measures, such as the mean and variance. Since entropy is not central to the inferential contribution of the paper, a compact Shannon entropy expression is recorded only as an ancillary descriptive quantity.
2.4.1. Raw Moments via the Lomax Representation
Let
and recall that
with Lomax density
For
, the
rth raw moment of
Z can be written as
Using the odds-scale Lomax representation and the standard Euler-type integral identity, the raw moments admit the following compact expression.
Proposition 7. Let with and . Then, for every , the power moment exists and is given byIn particular, , and all integer raw moments of order exist. For integer
, the beta function simplifies to
Thus, the first four raw moments are obtained from (
19) by setting
.
Let
denote the central moments. Then
Consequently, the standardized skewness and kurtosis coefficients are
Both coefficients are well defined for all
and
, since all positive integer moments of
Z exist. They can be computed directly from (
19)–(
21).
2.4.2. Shannon Entropy
For completeness, the Shannon entropy of
can be written as
. Using the density in (
6), this gives
Both expectations are finite for every
and
. Since entropy is not used in the subsequent estimation or model-comparison procedures, we retain this expression only as a descriptive property. When required, it can be evaluated by the one-dimensional integral
using standard adaptive quadrature on
.
2.5. Odds-Scale Identities Used in Inference
The following identities are used repeatedly in the inferential development:
Thus, likelihood-based inference for the unit-interval model can be interpreted through the classical Lomax likelihood applied to the odds-transformed sample. The corresponding inversion representation for random variate generation is given separately in
Section 4.
3. Statistical Inference
Let
be a random sample from the
distribution with pdf (
6). Statistical inference for
is developed through likelihood-based methods and maximum product of spacings estimation. Particular attention is paid to the finite-sample instability of the maximum likelihood estimator (MLE) under heavy-tailed regimes, since the odds transformation maps observations close to one into extreme values on the positive half-line.
3.1. Maximum Likelihood Estimation
Let
be an observed sample from
Z that follows
, the log-likelihood function is
The corresponding score vector is denoted by
. Define the odds-transformed sample
for
. Since the transformation
does not depend on
, the log-likelihood based on the sample coincides with the Lomax log-likelihood based on
up to an additive term that does not depend on
. In particular,
where
does not depend on
. Consequently, the MLEs for the odds-Lomax model based on
coincide with the MLEs for the Lomax model based on
.
This observation is important for interpretation. On the one hand, it provides a substantial computational simplification, since likelihood-based inference for the odds-Lomax model can be carried out through the corresponding Lomax likelihood on the transformed sample. On the other hand, it shows that the main inferential contribution of the present paper is not the introduction of a fundamentally new likelihood structure, but rather the explicit unit-scale interpretation and organization of that structure under the odds transformation.
Proposition 8. Let be an observed sample from Z that follows and define . Then the score functions are The score functions are obtained by direct differentiation of the odds-scale log-likelihood in (
23). The MLEs
are defined as any solution to the system
and
provided the maximizer lies in the interior of the parameter space
. The first equation admits an explicit solution for
as a function of
.
Proposition 9. For each fixed , the equation has a unique solutionMoreover, for fixed , the log-likelihood is strictly concave as a function of α, so (24) is the unique maximizer in α. The expression follows by solving for ; strict concavity in follows from . In practice, maximization of the profile log-likelihood over may be used, followed by . The following implementation strategy is convenient:
- (i)
Optimize over to enforce and automatically;
- (ii)
Use and as stable starting values, exploiting the proportional relationship between the scale parameter and the median of the baseline Lomax distribution;
- (iii)
Compute standard errors from the inverse observed information .
3.2. Finite-Sample Instability of the MLE
Although the odds-scale representation provides a simple likelihood structure, it also explains why the MLE may be unstable in finite samples. The transformation maps observations close to the upper endpoint of the unit interval into very large positive values. Under heavy-tailed parameter configurations, especially for small values of , a few large transformed observations may produce a nearly flat or ridge-like likelihood surface. As a consequence, the observed information matrix may become ill-conditioned, leading to extreme estimates, inflated Wald standard errors, and very long confidence intervals.
This phenomenon is not merely a numerical convergence issue. In fact, optimization algorithms may report successful convergence even when the resulting estimates are statistically unreliable. Therefore, convergence rates must be interpreted jointly with robust measures of estimation error, interval lengths, and coverage probabilities.
The profile representation in (
24) is useful for diagnosing this behavior. For each fixed
, the estimator
is explicit, so numerical instability is mainly associated with the profile likelihood in
. In the simulation study, this structure is used to distinguish parameter settings where the MLE is reliable from those where it is unstable.
3.3. Maximum Product of Spacings Estimation
To provide a more stable alternative to maximum likelihood estimation in challenging finite-sample settings, we also consider the maximum product of spacings (MPS) estimator. The maximum product of spacings estimator is a well-known alternative to maximum likelihood estimation for continuous distributions [
23,
24]. Here, it is used as a robustness-oriented inferential procedure in regimes where the likelihood may be affected by extreme odds-transformed observations. Let
denote the ordered sample and define
and
. For a given
, define the spacings
where
is given in (
5). The MPS estimator is defined as
In implementation, the optimization is carried out on the log-parameter scale, and , which automatically enforces the positivity constraints. If a spacing is numerically zero, a small machine-level constant is used only to avoid evaluating the logarithm at zero.
For interval estimation based on MPS, we use approximate Wald-type intervals on the log-parameter scale. Let
be the negative log-spacing objective function. After obtaining
, the inverse of the observed Hessian matrix
is used as a numerical approximation to the covariance matrix of the log-parameter estimators. This approximation is motivated by the standard large-sample theory of maximum product of spacings estimators for continuous distributions [
23,
24]. However, because the MPS criterion is based on spacings rather than on a likelihood function,
should not be interpreted as an inverse Fisher information matrix. Instead, it is used here only to construct approximate Wald-type intervals and to provide a diagnostic measure of local curvature of the spacing objective.
Approximate
confidence intervals are first constructed on the log scale as
and then transformed back to the original parameter scale by exponentiation. For example,
This construction is used in the Monte Carlo study to compute empirical coverage probabilities and average interval lengths for MPS-based intervals.
3.4. Parametric Bootstrap for the MPS Estimator
To complement the Hessian-based intervals for the MPS estimator, we also implemented a parametric bootstrap procedure. This is particularly useful in small samples and in regimes where the spacing objective may be locally flat or affected by observations close to the boundary of the unit interval.
The procedure is as follows. First, the model is fitted by MPS to obtain
. Second, for
, a bootstrap sample
is generated from the
using the inversion method described in
Section 4. Third, the MPS estimator is recomputed from each bootstrap sample, producing
Percentile bootstrap confidence intervals are then obtained from the empirical quantiles of the bootstrap estimates. For example,
with an analogous expression for
.
When bias correction is reported, it is computed on the log-parameter scale in order to preserve the positivity constraints. Let
and let
denote the bootstrap mean of the log-parameter estimates. The log-scale bias-corrected estimator is
and the corresponding estimator on the original scale is
This log-scale correction avoids producing inadmissible negative estimates for positive parameters. However, it should be interpreted with caution in small samples or heavy-tailed regimes. In such cases, the bootstrap distribution of the MPS estimates may be highly skewed or may contain extreme values, especially when some bootstrap samples generate observations close to the upper endpoint of the unit interval. Consequently, the simple linear bias correction on the log scale may overcorrect and move the resulting estimates toward numerically unreasonable regions of the parameter space. For this reason, in the empirical applications, we report percentile bootstrap intervals primarily as finite-sample uncertainty diagnostics, while the log-scale bias-corrected estimator is treated only as an optional descriptive correction rather than as a default replacement for the original MPS estimate.
3.5. Fisher Information and Asymptotic Theory
The Hessian matrix of the log-likelihood (
22) for a sample
is
From Proposition 8, direct differentiation yields
The observed information matrix is defined as and is evaluated at the MLE in applications.
Closed-form expressions for the Fisher information matrix
are now obtained. Since the observations are independent and identically distributed,
, where
denotes the Fisher information matrix for a single observation. For a generic random variable
, define
and
. The entries of
follow from (
25)–(
27) by setting
.
The diagonal element
is immediate:
so that
.
For the off-diagonal and
elements, the Lomax-based stochastic representation is useful. Recall that if
, then
, and conversely
. A short algebraic manipulation shows that
and therefore
where
.
Lemma 1. If , then for any , The entries are obtained by taking the expectations of the negative second derivatives of the log-likelihood, using the odds-scale Lomax representation and standard regularity arguments.
By (
28) and Lemma 1 with
,
The remaining entries of
can now be computed. From the single-observation versions of (
26) and (
27),
and therefore
Collecting the entries, the Fisher information matrix for a single observation is
and for a sample of size
n,
The determinant of
can be written in factored form as
so
is positive definite for all
and
.
The inverse of
is available in closed form. Writing
, one has
, and straightforward matrix inversion gives
Under the usual regularity conditions for maximum likelihood estimation (open parameter space, differentiability of the log-likelihood, existence and continuity of the Fisher information, and dominated convergence), the MLE is consistent and asymptotically normal.
Proposition 10. Let be an observed sample from Z that follows with and , and let denote the MLE. Then, as ,where is given by (29) evaluated at . In particular, the asymptotic variances and covariance of the MLE are Approximate Wald-type confidence intervals for and can be constructed in the usual way using the diagonal elements of , and likelihood ratio tests for nested hypotheses on follow standard asymptotic theory. In particular, the submodel (which reduces to ) can be assessed by a one-degree-of-freedom likelihood ratio test under the usual regularity conditions.
5. Monte Carlo Simulation Study
This section investigates the finite-sample behavior of two estimation procedures for the odds-Lomax model: MLE and MPS. The purpose is twofold. First, we diagnose the instability of the MLE under heavy-tailed and small-sample configurations. Second, we assess whether MPS provides a more stable alternative in such settings.
Monte Carlo experiments were carried out using
independent replications for each configuration. Random samples were generated from the odds-Lomax distribution for
,
, and
. For each simulated sample, the parameters were estimated by both MLE and MPS. The MLE was computed by numerical maximization of the log-likelihood using the odds-transformed representation described in
Section 3. The MPS estimator was computed by maximizing the logarithm of the spacings induced by the fitted cdf, as described in
Section 3.3.
5.1. Performance Measures
For each parameter
and each estimator
the following quantities were computed.
- (i)
- (ii)
- (iii)
- (iv)
- (v)
Empirical coverage probability of nominal
intervals:
- (vi)
- (vii)
The median absolute error is reported because RMSE may be dominated by a small number of extreme estimates, particularly in heavy-tailed small-sample configurations. Thus, RMSE and robust summaries are interpreted jointly.
5.2. Reliability Regimes Considered
The simulation design was chosen to distinguish between parameter regimes in which likelihood-based inference is expected to be reliable and regimes in which instability may occur. Small values of correspond to heavier tails on the Lomax odds scale and stronger concentration of probability mass near the upper endpoint of the unit interval. In these cases, observations close to one generate large transformed values , which may produce flat or ill-conditioned likelihood surfaces.
Accordingly, the setting is treated as a heavy-tailed regime, as an intermediate regime, and as moderate or lighter-tailed regimes on the odds scale. However, the simulation results also show that this classification alone does not guarantee finite-sample stability. The reliability of likelihood-based inference depends on the joint configuration , the occurrence of extreme odds-transformed observations, and the conditioning of the observed information matrix. This design allows us to characterize regimes in which the MLE behaves satisfactorily and regimes in which an alternative procedure, such as MPS, should be reported as a robustness check.
5.3. Simulation Results
The full Monte Carlo results are provided in the
Supplementary Material and in the reproducibility repository, as summarized in
Appendix A. The
Supplementary Material also includes the complete results for the case
, allowing direct comparison between MLE and MPS in the moderate-sample setting emphasized by the Academic Editor. To keep the main text concise, we focus here on the main patterns.
Table 1 summarizes representative heavy-tailed cases using the same
Monte Carlo replications reported in
Appendix A. These cases illustrate that MLE instability may appear not only through large RMSE values but also through extremely inflated average interval lengths. The MPS estimator substantially reduces this effect in the most problematic configurations.
5.4. Discussion of Simulation Results
The simulation study reveals three main findings. First, the MLE has a mixed finite-sample behavior. For several moderate configurations and sufficiently large samples, the MLE provides stable estimates, moderate RMSE values, and confidence intervals with reasonable empirical coverage. However, the results also show that large values of alone do not guarantee numerical or inferential stability. Certain combinations of , especially when accompanied by ill-conditioned observed information matrices, may still produce extreme estimates or inflated interval lengths. Therefore, reliability must be assessed through RMSE, MedAE, AIL, coverage, and convergence diagnostics jointly.
Second, in heavy-tailed regimes and small samples, the MLE may be severely unstable. This instability is reflected in very large RMSE values and inflated average interval lengths. Importantly, the instability may occur even when the numerical optimizer reports successful convergence. Therefore, the convergence rate alone is not a sufficient indicator of inferential reliability.
Third, the MPS estimator substantially mitigates the instability observed for the MLE. In the most problematic configurations, MPS reduces the effect of extreme estimates and produces more stable empirical medians as well as smaller RMSE values. However, the improvement should not be interpreted as a complete elimination of the difficulty: very small samples and strongly heavy-tailed regimes remain challenging.
These results lead to the following practical recommendation. Standard MLE is appropriate for moderate tail behavior and sufficiently large samples, provided that the observed information matrix is well conditioned and the interval lengths are reasonable. When the fitted model suggests heavy-tailed behavior on the odds scale, or when several observations are close to one, MPS should be reported as a robustness check or used as the primary estimator.
5.5. Practical Diagnostic Recommendation
For applied work, we recommend the following diagnostic protocol. After fitting the model by MLE, the analyst should inspect:
- (i)
The maximum value of the odds-transformed observations, ;
- (ii)
The condition number of the observed information matrix;
- (iii)
The length of the Wald confidence intervals;
- (iv)
The agreement between MLE and MPS estimates.
Large transformed observations, ill-conditioned information matrices, extremely long confidence intervals, or substantial disagreement between MLE and MPS indicate that MLE-based inference may be unreliable. In such cases, MPS estimates and spacing-based diagnostics should be preferred, or at a minimum, reported alongside MLE estimates.
6. Real-World Data Applications
In this section, we illustrate the empirical usefulness of the proposed odds-Lomax model through three real-world datasets on the unit interval. The first two datasets were originally reported in [
25] and correspond to the SC16 and P3 algorithms. Although these datasets are commonly described as computation-time data, the observations used here are bounded values within
. Therefore, in the present analysis, they are interpreted as normalized relative computation-time measures rather than as raw positive-valued running times. This distinction is important because the purpose of using a unit distribution is to model the bounded relative performance scale based on which the observations are available. Modeling the original unnormalized computation times, if available, would constitute a different positive-support modeling problem and is outside the scope of the present empirical illustration. These samples were also studied in [
26] in the context of estimating the unit capacity factor and were later analyzed in [
27] using the Topp-Leone distribution for comparative purposes. The third dataset is taken from the OECD Better Life Index (BLI) database (
https://stats.oecd.org/) and consists of 38 observations of the long-term unemployment rate, defined as the proportion of individuals who have been unemployed for one year or more relative to the total labor force. For completeness and reproducibility, the raw observations for the three datasets are reported in
Appendix B. In the main text, we summarize their main empirical features through the descriptive statistics displayed in
Table 2.
To align the empirical analysis with the diagnostic recommendations in
Section 5, we report both MLE and MPS estimates for the odds-Lomax model in all datasets. In addition to parameter estimates and standard errors, we include goodness-of-fit summaries and numerical stability diagnostics, namely the maximum odds-transformed observation
, the condition number of the observed Hessian or information matrix, and the minimum Hessian eigenvalue. These quantities help identify situations in which numerical instability may affect standard errors and Wald-type confidence intervals.
For Datasets 1 and 2, the term “computation-time data” should therefore be understood in a normalized sense. The values do not represent raw unbounded times but rather bounded relative measures on a unit scale. Consequently, these examples are used as illustrative bounded performance-type datasets, whereas Dataset 3 is a direct proportion-type variable.
The descriptive summaries reveal clear differences among the samples. Datasets 1 and 2 have very similar centers and dispersions, with means close to , medians near , and relatively large standard deviations compared with their central values. Their positive skewness coefficients also indicate noticeable right-skewness. By contrast, Dataset 3 is much more concentrated near zero, with a mean of , a median of , and a substantially larger skewness coefficient, reflecting a stronger right tail generated by a few relatively large observations. Hence, while all three datasets are bounded in , Dataset 3 presents a more pronounced lower-end concentration and a more challenging asymmetric structure than Datasets 1 and 2.
These features make the three samples suitable benchmarks for assessing the flexibility of bounded distributions. In particular, the Topp-Leone distribution has often been used as a reference model for such data because of its simplicity and tractability, but its one-parameter form may be too restrictive to accommodate the heterogeneity seen in
Table 2. In contrast, the proposed odds-Lomax distribution offers additional flexibility through its two-parameter specification, allowing for a broader range of density shapes and tail behaviors. This makes it a natural candidate for improving the fit to unit data exhibiting marked asymmetry, substantial spread, or strong concentration near one of the boundaries.
In the following subsections, the odds-Lomax model is fitted to each dataset and compared with the beta, Kumaraswamy, Topp-Leone, UL2, UL3, and UEL distributions using likelihood-based criteria, KS, CvM, and AD goodness-of-fit diagnostics, and graphical assessments based on fitted densities, empirical and fitted distribution functions, and P-P plots.
6.1. Competing Models
For comparative purposes, we also fitted several established bounded-data models, including the alternative unit-Lindley distribution [
15] and the unit Burr-XII distribution [
14], along with beta, Kumaraswamy, and Lomax-based competitors, with all model parameters estimated via maximum likelihood. In particular, the following distributions were considered: beta distribution, Kumaraswamy distribution, Topp-Leone distribution, UL2 model, UL3 model, the unit-exponentiated Lomax (UEL) distribution [
12], and the proposed odds-Lomax distribution. The UL2 and UL3 models are defined, respectively, by
and
6.2. Model Comparison Criteria
Model performance was evaluated using likelihood-based criteria and empirical goodness-of-fit diagnostics. The likelihood-based criteria include the maximized log-likelihood, Akaike information criterion (AIC), and Bayesian information criterion (BIC). The empirical goodness-of-fit diagnostics include the Kolmogorov–Smirnov (KS), Cramér–von Mises (CvM), and Anderson–Darling (AD) statistics.
The KS, CvM, and AD statistics were computed from the fitted cdf evaluated at the ordered observations using their standard empirical-distribution-function forms. The KS statistic measures the maximum pointwise discrepancy, the CvM statistic measures integrated squared discrepancy, and the AD statistic gives greater weight to tail discrepancies. To account for parameter estimation, KS p-values were obtained through a parametric bootstrap with parameter re-estimation. Specifically, for each fitted model, bootstrap samples were generated under the fitted parameter values, the model was re-fitted to each bootstrap sample, and the bootstrap distribution of the KS statistic was used to approximate the corresponding p-value. The estimated parameters are reported in the format estimated parameters (SE), thereby preserving the natural parametrization of each competing model. For models whose maximum likelihood estimates approach the boundary of the parameter space, standard errors should be interpreted with caution since Wald-type approximations may become unstable in such cases.
We also considered the reviewer’s suggestion regarding Kullback–Leibler-type discrepancies. For two densities
g and
on
, the Kullback–Leibler divergence is
In the present setting, g is unknown and only a small sample is available. Therefore, a formal KL-based goodness-of-fit test would require an additional nonparametric estimate of g and a bootstrap calibration of the resulting statistic. This would introduce bandwidth and smoothing choices that are not central to the main objective of the paper. For this reason, we do not develop a separate KL-based test here. Instead, we use likelihood-based criteria, which are directly related to empirical cross-entropy and KL risk, together with KS, CvM, AD, and graphical diagnostics. A full KL-based goodness-of-fit procedure for odds-scale unit models is left as a topic for future work.
The results in
Table 3 and
Table 4 connect the empirical applications with the diagnostic recommendations from the simulation study. For each dataset, we report both MLE and MPS estimates, together with goodness-of-fit statistics and numerical stability diagnostics. The maximum odds-transformed observation
, the Hessian condition number (Cond. no.), and the minimum Hessian eigenvalue (Min. eig.) provide information about the local stability of the estimation procedure. Moderate condition numbers and positive minimum eigenvalues indicate locally stable estimation, whereas large condition numbers, very small eigenvalues, inflated standard errors, or substantial disagreement between MLE and MPS estimates should be interpreted as evidence of finite-sample instability.
6.3. Dataset 1
Results for Dataset 1 are reported in
Table 5 and illustrated in
Figure 2. The odds-Lomax model provides the most favorable likelihood-based fit among the models considered, according to the log-likelihood, AIC, and BIC. The KS statistic and the bootstrap
p-value are also consistent with an adequate fit. The Kumaraswamy and beta distributions remain competitive, whereas UL2 and UL3 show clearer signs of misspecification. The inclusion of the UEL model is especially relevant here because it represents the closest Lomax-based competitor. However, its fitted values suggest numerical instability, with one parameter effectively collapsing to the boundary of the parameter space. Accordingly, the UEL fit should be interpreted with caution in this dataset, and the present comparison supports the odds-Lomax model primarily in terms of relative fit and numerical stability under the current sample.
Graphical diagnostics provide more detailed support for these conclusions. The histogram with fitted densities shows a pronounced concentration of observations near zero together with a sparse right tail. In this panel, the odds-Lomax curve reproduces the overall decreasing pattern more closely than the competing models, whereas UL2 concentrates too much density near the lower boundary, and UL3 concentrates too much density over the middle and upper parts of the support. The empirical-versus-fitted CDF panel also favors the proposed model, whose curve tracks the empirical step function more closely over the lower and central portions of the sample. Finally, in the P-P plot, the odds-Lomax model remains closest to the 45-degree reference line, with only mild deviations in the middle-to-upper probability range, indicating a satisfactory global fit over the unit interval.
6.4. Dataset 2
Results for Dataset 2 are presented in
Table 6 and
Figure 3. Again, the odds-Lomax distribution provides the most favorable likelihood-based fit among the fitted competitors and is also supported by the KS-based evidence. The comparison with the UEL model is informative, but it also reveals a numerical issue: the UEL estimates approach the boundary of the parameter space, and some reported standard errors are unavailable. This behavior suggests weak local identifiability or instability of the Wald approximation for that model under the present dataset. Hence, although UEL is retained as an important Lomax-based benchmark, its fitted results should be interpreted with these limitations in mind.
The graphical diagnostics reinforce this ranking. The histogram with fitted densities again reveals a marked concentration near zero, and the odds-Lomax density follows this shape more convincingly than the main competitors, while UL2 and UL3 display a clear mismatch in body and tail behavior. In the empirical-versus-fitted CDF panel, the proposed model remains closer to the empirical distribution over most of the support, especially in the lower and intermediate regions where an accurate fit is most important for this dataset. The P-P plot confirms this visual advantage: the odds-Lomax curve stays closest to the diagonal benchmark, whereas the beta and Kumaraswamy models show more systematic departures. Although the UEL curve is visually competitive in parts of the support, its numerical instability prevents a reliable inferential interpretation.
6.5. Dataset 3
The results for Dataset 3 are shown in
Table 7 and
Figure 4. This dataset presents a more challenging structure due to the strong concentration of observations near zero. The odds-Lomax model remains highly competitive according to the likelihood-based criteria, while alternative models such as UL3 and Kumaraswamy exhibit competitive behavior in some diagnostic aspects. The UEL model again yields estimates close to the boundary, indicating that its numerical fit is comparatively unstable for this dataset. Therefore, although it remains a natural Lomax-based comparator, the practical interpretation of its parameter estimates and standard errors is limited in this case.
The graphical diagnostics show why this dataset is more difficult. The histogram with fitted densities indicates a very strong concentration near zero, together with a thin but non-negligible spread over the remaining support. In this panel, no single competitor dominates uniformly: some alternative models adapt reasonably well in restricted regions, but the odds-Lomax density provides a more balanced representation of the sharp lower-tail behavior and the subsequent decay. The empirical-versus-fitted CDF panel shows a similar behaviour. While several curves remain close to the empirical step function, the proposed model maintains good agreement across most of the support without the pronounced distortions displayed by the clearly misspecified models. The P-P plot further indicates that the comparison is tighter than in Datasets 1 and 2, since Kumaraswamy and UL3 also track the empirical probabilities reasonably well. Even so, the odds-Lomax model preserves the best overall compromise between lower-tail adaptation, global fit, and consistency with the likelihood-based criteria.
The bootstrap intervals in
Table 8 should be interpreted as complementary finite-sample diagnostics rather than as replacements for the Hessian-based intervals. For the SC16 and P3 datasets, the bootstrap intervals are relatively moderate and support the use of MPS as a stable, robustness-oriented alternative. For the BLI dataset, however, the bootstrap interval for
is very wide, indicating substantial finite-sample uncertainty. This reinforces the need to interpret parameter estimates cautiously when the sample is small, strongly asymmetric, or highly concentrated near a boundary.
6.6. Overall Assessment
The empirical applications should be interpreted primarily as illustrative examples rather than as definitive evidence of estimator superiority. This caution is particularly relevant for Datasets 1 and 2, whose sample sizes are small, with
and
, respectively. Although these datasets are useful for illustrating the behavior of the odds-Lomax model on bounded performance-type data, small samples may lead to unstable estimates, inflated standard errors, and wide confidence intervals. Therefore, the empirical results are interpreted together with the simulation study and the stability diagnostics reported in
Table 4.
Some fitted models produce relatively large standard errors and wide confidence intervals, especially in small-sample settings or when the data are highly concentrated near the boundary of the unit interval. These cases should not be interpreted as strong empirical evidence in favor of a particular estimator or model. Rather, they provide empirical manifestations of the same instability mechanisms identified in the simulation study: observations close to one are mapped into large odds-scale values, which may flatten or ill-condition the objective function. For this reason, large standard errors, wide intervals, high Hessian condition numbers, and small minimum Hessian eigenvalues are interpreted as warning signals. In such cases, MPS estimates, bootstrap intervals, and the stability diagnostics in
Table 3 and
Table 4 are reported as complementary tools for assessing the numerical stability and inferential validity of the fitted results.
Across the three datasets, the odds-Lomax distribution provides a competitive fit relative to the bounded-data models considered here. The model is particularly useful when the data exhibit marked asymmetry or concentration near one of the endpoints. The comparison with the unit exponentiated Lomax model is especially relevant because it represents a close Lomax-based competitor. In the present applications, the odds-Lomax fits were numerically interpretable and did not exhibit the boundary-estimation problems observed for some competing models. However, in view of the Monte Carlo results, this empirical stability should be interpreted as dataset-specific rather than as a general guarantee of uniform MLE stability.
6.7. Discussion of Empirical Results
The empirical analysis across the three datasets highlights both the flexibility and the limitations of the proposed formulation. In Datasets 1 and 2, the histogram panels show that the odds-Lomax model captures the strong concentration of observations near the lower part of the unit interval, while the empirical-versus-fitted CDF plots indicate close agreement with the observed distribution over most of the support. The corresponding P–P plots show only minor deviations from the diagonal reference line, which is consistent with the favorable goodness-of-fit statistics.
For Dataset 3, the graphical comparison is more nuanced. The histogram and fitted-density panel reveal a more difficult tail structure, and some competitors remain visually close to the proposed model in restricted regions of the support. Nevertheless, when numerical fit criteria, graphical diagnostics, and interpretability are considered jointly, the odds-Lomax model remains a competitive option.
These applications should be read together with the simulation study. While the fitted examples indicate that the odds-Lomax model can be useful in practice, the Monte Carlo results show that standard MLE may be unstable in small-sample or heavy-tailed regimes. Therefore, in empirical work, likelihood estimates should be accompanied by convergence diagnostics, interval-length checks, and, when necessary, alternative estimation methods such as MPS.
7. Conclusions
This paper has provided a structural, inferential, and computational study of an odds-Lomax model on induced by the transformation with . A central contribution is the explicit clarification of the relationship between this odds-scale formulation and previously used unit-Lomax-type constructions. In particular, the model is not presented as a genuinely new distributional family, but rather as an odds-scale representation that is equivalent, up to complementation and reparametrization, to an existing Lomax-based unit construction. This clarification helps resolve issues of parametrization and nomenclature while preserving a direct probabilistic interpretation in terms of a heavy-tailed latent variable.
The model retains substantial analytical tractability. Closed-form expressions were obtained for the main distributional functions, quantile function, hazard rate, endpoint behavior, and stochastic representations. Moment and entropy-related quantities can also be evaluated through tractable formulas involving standard special functions or stable numerical procedures.
From an inferential perspective, the odds transformation shows that likelihood inference for the unit-interval model reduces to classical Lomax inference on the transformed sample . This representation is useful computationally, but it also explains the finite-sample instability of the MLE in heavy-tailed regimes. Observations close to the upper boundary of the unit interval generate extreme values on the odds scale, which may lead to flat likelihood surfaces, ill-conditioned observed information matrices, inflated Wald intervals, and very large RMSE values.
The Monte Carlo study therefore provides a more nuanced conclusion than a standard likelihood analysis alone. The MLE performs satisfactorily in several moderate-tail configurations and sufficiently large samples, but its reliability is not uniform over the entire parameter space. The simulation results show that isolated extreme estimates may still occur for some parameter combinations, making interval-length checks, MedAE, and comparison with MPS essential diagnostic tools.
The empirical applications show that the odds-Lomax formulation can be competitive for bounded data exhibiting asymmetry or endpoint concentration. However, the practical use of the model should be accompanied by diagnostic checks, including convergence assessment, interval-length inspection, and comparison between MLE and MPS estimates when the data suggest heavy-tailed behavior on the odds scale.
Future work may consider regression extensions, multivariate constructions on the simplex, copula-based dependence structures, profile-likelihood or bootstrap interval estimation, and systematic comparisons between likelihood-based, spacing-based, and fiducial procedures for bounded data.