1. Introduction
Count data modeling plays a central role in applied statistics and arises in diverse fields such as reliability engineering, epidemiology, environmental science, and sports analytics. The Poisson distribution remains the most classical and parsimonious model for count data; however, its fundamental assumption of equidispersion—equality of the mean and variance—is often violated in practice. Empirical datasets frequently exhibit overdispersion and latent heterogeneity, which motivates the development of more flexible count models.
One of the most successful strategies for accommodating overdispersion is the Poisson mixture (or compounding) framework. This approach has generated several influential distributional families, including the Poisson–Gamma (negative binomial), Poisson–Lindley [
1], and Poisson–Lindley–Quasi XGamma [
2] distributions. Further extensions, such as the new three-parameter Poisson–Lindley model [
3], have expanded this framework by allowing finer control over dispersion and capturing more intricate variability patterns encountered in reliability and risk analysis.
Building on this line of research, Yousfi and Zeghdoudi [
4] propose the Poisson–X–Exponential (PXED) distribution by incorporating the X–Exponential kernel into the Poisson mixing structure, thereby addressing overdispersion while preserving analytical tractability. In the multivariate context, Arrar et al. [
5] propose the Bivariate Poisson–XLindley (BPXL) distribution with applications to soccer match outcomes, whereas Haddari et al. [
6] develop the Modified Bivariate Poisson–Lindley model (BPNXLD), offering enhanced flexibility in modeling dependence structures. These developments are closely related to reliability modeling under exponential-type distributions, where Bayesian estimation techniques have been explored under various loss functions [
7]. Moreover, the application of bivariate Poisson models to sports data analysis has been well established in the literature, particularly in modeling match scores and dependence between competing teams [
8].
Beyond mixture-based constructions, alternative mechanisms have been proposed to introduce dependence in bivariate count models. For example, Genest et al. [
9] developed a bivariate Poisson common-shock model capable of generating a wide range of positive dependence through a shared latent shock component. Using a conditional specification approach, Ghosh et al. [
10] introduced a tractable bivariate Poisson model with explicit joint and conditional structures. From a compound distribution perspective, Abdelghani et al. [
11] proposed a bivariate model based on Poisson maxima of Gamma variates, providing additional flexibility in tail behavior. More recently, Maya et al. [
12] developed bivariate Poisson extended exponential distributions together with associated BINAR(1) processes, enabling joint modeling of contemporaneous and temporal dependence in multivariate count time series. Collectively, these contributions highlight the diversity of strategies used to relax the restrictive assumptions of the classical bivariate Poisson model while maintaining mathematical and computational tractability.
Despite these advances, many existing models are motivated primarily by specific application domains or rely predominantly on Lindley-type mixing structures. Moreover, comprehensive evaluations of a single bivariate count model across substantially different scientific fields remain scarce.
In this context, the present work introduces a unified bivariate Poisson-type model designed for broad applicability across multiple domains. To the best of our knowledge, few studies have systematically assessed the same bivariate count model in diverse applied settings such as:
Sports analytics (e.g., soccer goal modeling and outcome prediction);
Reliability engineering (e.g., dependent system failures); and
Astronomy and physics (e.g., correlated photon-count data).
We propose a new bivariate discrete distribution—the Bivariate Poisson–X–Exponential Distribution (BPXED)—constructed by compounding two Poisson variables with a shared latent X–Exponential mixing variable. This formulation induces positive dependence through a common latent factor while permitting flexible marginal dispersion. The BPXED model encompasses the classical Bivariate Poisson (BP) and the Bivariate Poisson–Lindley (BPLD) as special or limiting cases, thereby offering enhanced interpretability and dispersion control within a coherent and analytically tractable framework.
The remainder of the paper is organized as follows.
Section 2 introduces the proposed BPXED model, presents its construction and principal distributional properties, and develops the associated estimation procedures, including maximum likelihood, regression-based, and Bayesian approaches.
Section 3 reports the Monte Carlo simulation study and empirical applications to sports scores, reliability failure counts, and correlated photon-count data, together with comparisons to competing bivariate Poisson-type models.
Section 4 interprets the main findings, highlights practical implications, and examines the strengths and limitations of the proposed model. Finally,
Section 5 summarizes the key contributions and outlines directions for future research.
4. Discussion
This study introduced the Bivariate Poisson–X–Exponential Distribution (BPXED), a mixture-based bivariate count model obtained by compounding two Poisson variables with a shared X–Exponential latent mixing distribution. The proposed construction yields a closed-form joint pmf and a simple moment structure, while accommodating marginal overdispersion and inducing positive dependence through shared heterogeneity.
4.1. Interpretation of the Dependence Mechanism
A key feature of BPXED is that dependence arises solely from the common latent factor
. Conditionally on
, the counts are independent, but marginally they are positively correlated with
which highlights a clear interpretation: larger heterogeneity in the underlying intensity (i.e., larger
) produces stronger association between the two counts. In practical terms, BPXED is well suited to settings where both outcomes are driven by an unobserved common environment, such as team strength and match tempo in soccer, operating conditions in reliability, or source variability in photon emission.
The role of is particularly informative. Because governs the mixing distribution, it controls (i) the overall level of overdispersion in each margin and (ii) the magnitude of cross-dependence. Larger corresponds to weaker latent heterogeneity and, consequently, smaller dispersion and weaker dependence. The parameters and primarily determine the marginal intensities, providing a clean separation between marginal level () and shared variability ().
4.2. Empirical Performance Across Domains
The applications demonstrate that BPXED can provide a competitive and often improved fit relative to classical and recent bivariate Poisson-type models. In the soccer dataset, BPXED achieved the smallest information criteria and Pearson , indicating that the model can capture both overdispersion and positive association beyond the classical BP benchmark. Similar improvements were observed in the reliability data, where failure counts typically exhibit heterogeneity due to changing operating conditions, maintenance, or load fluctuations. In the photon-count application, BPXED reproduced the empirical correlation closely and produced the lowest AIC/BIC among the competitors considered, suggesting that the shared-mixing interpretation is plausible for correlated count processes observed under the same underlying emission state.
Across all datasets, the consistent ranking of BPXED suggests that the X–Exponential mixing distribution provides an effective compromise between flexibility and analytic tractability. While more complex constructions can increase flexibility, they often sacrifice closed-form expressions or impose heavier numerical burdens. BPXED remains computationally manageable through standard likelihood maximization while retaining an interpretable hierarchical representation.
4.3. Estimation and Practical Considerations
The simulation study supports the feasibility of parameter estimation under common sample sizes used in practice. As expected, estimation accuracy improves with
n, and Bayesian estimation can stabilize inference in smaller samples by incorporating mild regularization through prior information. For applied work, method-of-moments starting values provide a practical route to stable likelihood optimization. In addition, the hierarchical form
facilitates Monte Carlo generation and posterior computation via MCMC, making BPXED convenient both for frequentist and Bayesian workflows.
4.4. Model Limitations
Despite its advantages, BPXED has limitations that should be acknowledged. First, because dependence is induced via common mixing, the model is restricted to nonnegative correlation. This is appropriate for many real-world applications with shared risk or shared environmental effects, but it may be unsuitable when negative dependence is present (e.g., competitive substitution effects). Second, as with many parametric count models, goodness-of-fit can deteriorate if data exhibit strong zero inflation, structural zeros, or multimodality not well represented by the chosen mixing distribution. Third, the current formulation assumes a single common latent factor for both margins; in some applications, partial sharing or multiple latent components may be needed to capture more complex dependence patterns.
4.5. Implications and Future Directions
The results suggest several practical directions. First, extending BPXED to include covariates through log-linear links for can enhance interpretability and predictive performance in regression settings. Second, incorporating zero-inflated or hurdle mechanisms may broaden applicability to sparse datasets. Third, higher-dimensional generalizations with shared or partially shared mixing variables could extend the model to multivariate count vectors while preserving interpretability. Finally, implementation in open-source software would facilitate broader adoption and reproducibility.
Overall, BPXED provides a tractable and interpretable framework for modeling positively dependent and overdispersed bivariate count data. Theoretical derivations, simulation evidence, and cross-domain applications collectively indicate that the model can serve as a useful alternative to existing bivariate Poisson-type distributions, particularly when shared latent heterogeneity is a natural scientific explanation for dependence.
5. Conclusions
This paper introduced the Bivariate Poisson–X–Exponential Distribution (BPXED), a new bivariate count model obtained by compounding Poisson variables with a shared X–Exponential mixing distribution. The proposed specification extends classical bivariate Poisson and Lindley-type constructions while preserving analytical tractability. Closed-form expressions were derived for the joint probability mass function, probability generating function, and main moment characteristics. The model accommodates marginal overdispersion and induces positive dependence through a common latent factor.
Empirical analyses from sports, reliability, and astrophysics demonstrate that BPXED provides competitive and, in several cases, improved fit relative to existing bivariate Poisson-type models. Across applications, the model yielded favorable likelihood-based criteria and reproduced observed dependence patterns. The parameters admit a clear interpretation: and determine marginal intensities, while controls the degree of latent heterogeneity and the strength of induced dependence.
The shared mixing structure implies that BPXED is restricted to a nonnegative correlation. This feature aligns with many applications involving common environmental or systemic influences, although it limits applicability when negative dependence is present.