Next Article in Journal
Financial Flexibility, Corporate Governance, and Firm Performance: Evidence from Chinese A-Share Listed Firms
Previous Article in Journal
Digital Financial Inclusion and Household Financial Fragility: Evidence of a U-Shaped Relationship in China
Previous Article in Special Issue
The Kerper–Bowron Method: A Foundational Change for Service Contract Claim Estimation and Accounting
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Modeling Data with Nonlinear LogNormal–Pareto Regression via the Approximate Bayesian Computation

by
Mostafa S. Aminzadeh
Department of Mathematics, Towson University, Towson, MD 21252, USA
Risks 2026, 14(7), 165; https://doi.org/10.3390/risks14070165
Submission received: 1 June 2026 / Revised: 10 July 2026 / Accepted: 14 July 2026 / Published: 16 July 2026
(This article belongs to the Special Issue Advances in Risk Models and Actuarial Science)

Abstract

The development of regression models for composite distributions has received insufficient attention in the literature. The purpose of this research is to provide maximum likelihood (ML) and approximate Bayesian computation (ABC) estimators for the parameters of a regression model with a response variable following the LogNormal–Pareto composite distribution. Composite models such as Exponential–Pareto, LogNormal–Pareto, and Inverse-Gamma–Pareto, which separate small-to-moderate and significant losses using a threshold parameter, have been developed using classical and Bayesian methods and applied to insurance data. We derive closed-form formulas for MLEs in regression models with two and three covariates. For models with more than three covariates, we provide MLEs using Mathematica code. Simulation studies show that the ABC method is more accurate than the ML method. Mathematica code written specifically for the computations of the proposed methods is provided.

1. Introduction

The main characteristic of loss data in the insurance industry is that small losses occur frequently and large losses occur infrequently. Composite distributions such as Exponential–Pareto, LogNormal–Pareto, and Inverse-Gamma–Pareto, which separate small-to-moderate and significant losses using a threshold parameter, have been developed using classical and Bayesian methods and applied to insurance data. Risk modeling associated with insurance losses is a crucial part of an actuary’s work. It is important to utilize the correct probability distribution that captures the characteristics of claim data in order to provide an accurate estimate of the desired risk. Klugman et al. (2012) provide a comprehensive review of modeling data sets with composite models in actuarial science.
Gamma-Type II–Pareto was developed by Teodorescu and Vernic (2013). The authors derived the Gamma–Pareto model as a special case with three free parameters. Cooray and Ananda (2005) developed a 2-parameter LogNormal–Pareto model and applied it to analyze Danish Fire Insurance data. Deng et al. (2021) used LogNormal–Pareto as one of the possible models for natural disaster losses. Grün and Miljkovic (2019) presented a comprehensive analysis of composite models for the Danish Fire data set, using 16 parametric distributions most often employed in actuarial science. Mutali and Vernic (2022) considered estimating the threshold parameter for the LogNormal–Pareto using two approaches: when the parameter is a random variable and when it is a fuzzy number. Liu and Ananda (2023) introduced the exponentiated Inverse Gamma–Pareto model and applied it to several insurance data sets, concluding that the new model fitted some insurance data better than the Inverse-Gamma–Pareto model. Kumi et al. (2026) considered composite distributions to model automobile insurance from Ghana using 11,879 claims. The article considered 240 composite models for automobile insurance claims and selected the top 10 using goodness-of-fit measures; the LogNormal–Pareto model was the best. Marambakuyana and Shongwe (2024) provided composite and mixture models for insurance claims data. The article analyzed two real insurance data sets from different industries and provided risk metrics for the top 20 models in each category (composite/mixture). Sarabia and Calderín-Ojeda (2018) provided expressions for several actuarial and statistical quantities, such as Value-at-Risk and tail moments, for a general class of composite models. Li and Liu (2023) provided new composite models for use with individual claims in the insurance industry. In contrast to common composite models with two components, the article used a Weibull distribution for the smallest claims, a LogNormal distribution for medium-sized claims, and a long-tailed distribution for very large claims. The models were then applied to two real-world insurance data sets using both ML and Bayesian methods.
Regression models for composite distributions have not received sufficient attention in the literature. In a regression context, covariates enable actuaries to build composite models that are sensitive to individual risk profiles. Instead of a static threshold separating standard claims from tail-risk catastrophes, a covariate-dependent threshold parameter allows the threshold to vary by policyholder. This captures how specific driver or property traits affect both frequency and extreme severity. By incorporating covariates, models such as the LogNormal–Pareto can handle both everyday severity and extreme tail events within a unified framework. To our knowledge, Konşuk Ünlü (2022) is the only article to date that has employed the LogNormal–Pareto Type II composite distribution to formulate a regression model for analyzing the National Household Budget Study data conducted by the Turkish Statistical Institute. The model uses two sets of regression parameters, one for each component of the composite distribution, and estimation is performed using Particle Swarm Optimization on household budget data. The Particle Swarm Optimization method requires extensive computations and coding. In contrast, the regression model in the current article is based on a single set of regression parameters, and ML estimation is performed by directly optimizing the likelihood function and identifying insignificant covariates using the estimated Fisher information matrix. Additionally, for the two- and three-covariate cases, the exact MLEs are derived, and corresponding Mathematica programs are provided. Furthermore, in contrast to Konşuk Ünlü (2022), we provide an algorithm based on approximate Bayesian computation (ABC) for estimating regression parameters, and we demonstrate that ABC outperforms MLEs in terms of accuracy. Another novel contribution of the current article is the use of a data-driven approach, using MLEs and the corresponding Fisher information matrix, to identify hyperparameters associated with the regression parameters. In a recent article, Aminzadeh and Deng (2025) provided ML estimates for the parameters of a new two-parameter Gamma–Pareto composite distribution and used ML and ABC methods to estimate regression parameters. However, the estimation algorithm requires that the shape parameter ( α ) of the Gamma distribution be approximated using Mathematica code before optimizing the likelihood function to compute the MLEs of the regression parameters. In contrast, in this research, a specific parameterization of the LogNormal–Pareto model is used so that none of the parameters need to be approximated before obtaining ML and ABC estimates of the regression parameters.
In Section 2, the two-parameter composite LogNormal–Pareto distribution, also considered by other authors, is developed by requiring that the composite PDF be continuous and differentiable at the threshold parameter. Section 3 presents an algorithm to compute MLEs and to determine the correct value of m, where y m is the loss value at the threshold parameter. Section 4 presents the likelihood function in a regression context, modeling the relationship between the response variable and the covariates via a link function. In this section, the general approach to computing regression parameters for p covariates is presented, and closed-form formulas for MLEs are provided for models with p = 2 and p = 3 .
Section 5 uses the ABC algorithm with a multivariate normal prior to estimate regression parameters. A data-driven approach based on the MLEs of the regression parameters is used to determine the hyperparameters. This section also outlines the steps in the Mathematica programs written specifically for this article to compute regression parameter estimates using the ML and ABC methods. Our simulation results are presented in Section 6. In Section 7, a detailed analysis of two data sets is presented.

2. Derivation of LogNormal–Pareto Composite Distribution

In the following, we derive the PDF of the proposed LogNormal–Pareto model. Let Y be a random variable with the probability density function
f Y ( y ) = c f 1 ( y ) 0 < y θ c f 2 ( y ) θ y <
where
f 1 ( y ) = 1 2 π σ y e ( ln ( y ) μ ) 2 2 σ 2 , σ > 0 , μ > 0 , 0 < y θ
and
f 2 ( y ) = α θ α y α + 1 y θ , θ > 0 , α > 0 .
The normalizing constant c, can be found via
c 1 = 0 θ f 1 ( y ) d y + θ f 2 ( y ) d y .
The closed formula is given below, once the number of parameters is reduced to two; f 1 ( y ) and f 2 ( y ) are, respectively, the PDFs of the LogNormal and Pareto distributions. For the composite density function to be smooth, the two PDFs must be continuous and differentiable at the threshold parameter θ . That is,
f 1 ( θ ) = f 2 ( θ ) , f 1 ( θ ) = f 2 ( θ ) .
From f 1 ( θ ) = f 2 ( θ ) , we get
α = 1 σ 2 π e w 2 2 ,
and from the equation f 1 ( θ ) = f 2 ( θ ) , we get
1 2 π σ 2 ( w + σ ) e w 2 2 = α ( 1 + α ) ,
where w = ln ( θ ) μ σ . Using (1) and (2) yields
α = w σ .
Equating Equations (1) and (3) gives the following non-linear equation:
e w 2 2 = 2 π w w = 0.372239 .
It can be verified that the normalizing constant c for the LogNormal–Pareto composite PDF is
c = 2 3 erf w 2 = 0.60785 ,
where
e r f ( z ) = 2 π 0 z e x 2 d x .
Using ( ln ( y ) μ ) 2 = ( ln ( θ ) μ ) + ln ( y / θ ) ) 2 and (3), the composite LogNormal–Pareto density function can be written as
f Y ( y ) = c α θ α y ( α + 1 ) e α 2 ln 2 ( y / θ ) 2 w 2 0 < y θ c α θ α y ( α + 1 ) θ y <
and the CDF F Y ( y ) can be derived as
F Y ( y ) = F 1 ( y ) = 0.303925 1 + erf 0.707107 ln y θ σ + 0.263213 0 < y θ F 2 ( y ) = 1 0.60785 y θ α θ y <
We observe that the original four parameters are reduced to two: α and θ . Figure 1a–c show that for a fixed value of α the PDF is a smooth curve. For smaller values of α (corresponding to larger values of σ ), the PDF is more heavily tailed.
  • Value- at-Risk (VaR)
In actuarial science applications, an important risk measure is Value-at-Risk (VaR), which is used to assess risk exposure. VaR quantifies the potential loss an investment may incur at a given confidence level 1 γ and is defined as
V a R ( 1 γ ) = i n f y | F Y ( y ) 1 γ .
The common values of 1 γ are 0.99 and 0.95. For composite distributions as a loss model, VaR lies in the second distribution F 2 ( y ) . Based on the LogNormal–Pareto model,
V a R ( 1 γ ) = F 2 1 ( 1 γ ) = θ c γ ( 1 α ) .
For example, with θ = 50 and α = 2.0 , and with ( 1 γ ) = 0.95 , we get V a R ( 0.95 ) = 174.334 .

3. Maximum Likelihood Method for Parameters and Finding m

This section discusses computing the ML estimates of the parameters α and, consequently, θ for the PDF f Y ( y ) given in Section 2. Assume y 1 , y 2 , , y n LogNormal Pareto ( θ , α ) , with y 1 < y 2 < y m < y m + 1 < < y n for an integer m such that 1 < m < n . Assume that y 1 , y 2 , , y m and y m + 1 , , y n are, respectively, from the first and second components of the composite distribution. The likelihood function L is written as
L = c n α n θ n α i = 1 n y i ( α + 1 ) e i = 1 m α 2 2 w 2 ln 2 ( y i / θ )
Following the algorithm in Deng et al. (2021), the MLEs are obtained as follows:
If m = 1 ,
θ ^ = x 1 i = 1 m x i / x 1 w 2 , α ^ = m i = 1 m ln ( x i / x 1 ) 1
otherwise,
θ ^ = e n w 2 m α ^ i = 1 m x i 1 m , α ^ = w 2 B + w 4 B 2 + 4 m n w 2 A 2 A ,
where A = m i = 1 m ( l n ( x i ) ) 2 ( i = 1 m l n ( x i ) ) 2 , B = n i = 1 m l n ( x i ) m i = 1 n l n ( x i ) .
In the simulation studies, a code similar to Code A (see Supplementary Materials) uses selected “true” values α t and θ t to generate N = 100 samples of size n from the composite distribution. For each generated sample, the correct value of m is determined and then used to compute the MLEs. The mean of the MLE and the average squared error (ASE), denoted by ξ , are listed in Table 1.
Table 1 reveals that as the sample size n increases, the ASE for the estimators, defined as
ξ ( α ^ ) = j = 1 N ( α ^ j α t ) 2 N ξ ( θ ^ ) = j = 1 N ( θ ^ j θ t ) 2 N ,
decreases for all values of α t and θ t listed in Table 1. Also, for specific combinations of ( α t , θ t ) with the same θ t , a larger α t yields a smaller ξ ( θ ^ ) . Similarly, when α t is held fixed a smaller θ t yields a smaller ξ ( α ^ ) . A visual summary of the results in Table 1 is shown in Figure 2.

4. Regression Model

Suppose there are p potential covariates, x 1 , , x p associated with the loss variable Y. For each row x i = ( x i 1 , , x i p ) , i = 1 , 2 , , k in the design matrix X ( k × p ) let y i be a realization of the response variable Y from the LogNormal–Pareto composite PDF. Without loss of generality, assume the ordered sample y 1 < y 2 < < y m < y m + 1 , < y k , where y 1 , , y m come from the LogNormal and y m + 1 , , y k from the Pareto.
Let us rewrite PDF (4) utilizing α = w σ ; we get
f Y ( y ) = c ( w σ ) θ ( w σ ) y ( w σ + 1 ) e ln 2 ( y / θ ) 2 σ 2 0 < y θ c ( w σ ) θ ( w σ ) y ( w σ + 1 ) θ y <
The advantage of the above parameterization is that it directly involves the scale parameter σ of the LogNormal distribution, with larger σ yielding a more heavily tailed PDF. The regression parameters are denoted by β = ( β 1 , β 2 , , β p ) . To define a link function in regression modeling, the threshold parameter θ i must be associated with the p covariates. The constraint is that θ i , for given covariate values x i 1 , , x i p , i = 1 , 2 , , k , must be positive. Choices like θ i = x i β or θ i = ln ( x i β ) are not mathematically feasible because, due to dependence on the covariate data, θ i may not be positive. In that article, we chose θ i = e x i β , i = 1 , 2 , , k , as the link function. If covariate values are large, rescaling them is recommended to avoid overflow when applying the link function.
With the link function and covariates, the loglikelihood function l for the LogNormal–Pareto composite distribution can be expressed as
l k ln ( σ ) ( w σ ) i = 1 k x i β ( 1 + w σ ) i = 1 k ln ( y i ) 0.5 σ 2 i = 1 m ( ln ( y i ) x i β ) 2 .
There are p + 1 parameters, including σ , that should be estimated by optimizing the likelihood function.

ML Estimation for Regression Parameters

This section presents the computation of MLEs of the regression parameters. The variance–covariance matrix of the estimators is obtained from the Fisher information matrix, which is then used to identify insignificant covariates. Using matrix notation, the log-likelihood function, based on a random sample y ̲ on Y, and a k × p design matrix X,
y ̲ = y 1 y 2 y m y m + 1 y k X k × p = x 1 x 2 x m x m + 1 x k
can be expressed as,
l k ln ( σ ) + ( w σ ) 1 X β ( 1 + w σ ) 1 ln ( y * ) 0.5 σ 2 ln ( y * ) X * β ln ( y * ) X * β ,
where
y ̲ * = y 1 y 2 y m X m × p * = x 1 x 2 x m 1 k × 1 = 1 1 1 1 m × 1 * = 1 1 1 .
Differentiating l with respect to σ and β 1 , , β p yields the following equations:
l σ = k σ 2 + σ w 1 X β σ w 1 ln ( y ̲ ) l n ( y ̲ * ) l n ( y ̲ * ) + 2 ln ( y ̲ * ) X * β β X * X * β = 0 ,
l β g = w σ 1 x g l n ( y ̲ * ) x g * + β g x g * x g * h g m β h x g * x h * = 0 , g , h = 1 , 2 , , p .
Equations (5) and (6) can be solved simultaneously for the p + 1 parameters to obtain exact solutions; there is no need to use an optimization program, such as Mathematica’s NMaximize, to maximize the likelihood function. In the following, we present explicit formulas for the MLEs of the regression parameters for the p = 2 and p = 3 cases.
  • The p = 2 Case
From (5) and (6), we get
β 1 = β 2 ( A 1 A 2 A 2 A 5 ) + A 2 A 6 A 1 A 7 A 2 A 3 A 1 A 5 , σ = S 1 + S 1 2 + 4 k S 2 2 k ,
where
S 1 = w i = 1 k β 1 x i , 1 + β 2 x i , 2 ln ( y i ) , S 2 = i = 1 k ln ( y i ) β 1 x i , 1 β 2 x i , 2 2 ,
A 1 = i = 1 k x i , 1 , A 2 = i = 1 k x i , 2 , A 3 = i = 1 m x i , 1 2 , A 4 = i = 1 m x i , 2 2 ,
A 5 = i = 1 m x i , 1 x i , 2 , A 6 = i = 1 m x i , 1 ln ( y i ) , A 7 = i = 1 m x i , 2 ln ( y i ) ,
Note that σ , throughout S 1 and S 2 , depends on β 2 . Therefore, using (7), β ^ 2 is the unique solution to the equation
l β 1 = w σ A 1 + A 6 β 1 A 3 + β 2 A 5 = 0 .
Consequently, β ^ 1 and σ ^ can be computed using (7). Code B (see Supplementary Materials) computes the MLEs of the regression parameters for the p = 2 case.
  • The p = 3 Case
From (6), we get three equations:
B 10 + β 1 B 4 + β 3 B 7 + β 2 B 8 B 1 = B 11 + β 2 B 5 + β 1 B 8 + β 3 B 9 B 2
B 10 + β 1 B 4 + β 3 B 7 + β 2 B 8 B 1 = B 12 + β 3 B 6 + β 2 B 9 + β 1 B 7 B 3
B 11 + β 2 B 5 + β 1 B 8 + β 3 B 9 B 2 = B 12 + β 3 B 6 + β 2 B 9 + β 1 B 7 B 3 ,
where
B 1 = i = 1 k x i , 1 , B 2 = i = 1 k x i , 2 , B 3 = i = 1 k x i , 3 , B 4 = i = 1 m x i , 1 2 , B 5 = i = 1 m x i , 2 2 , B 6 = i = 1 m x i , 3 2
B 7 = i = 1 m x i , 1 x i , 3 , B 8 = i = 1 m x i , 1 x i , 2 , B 9 = i = 1 m x i , 2 x i , 3 , B 10 = i = 1 m x i , 1 ln ( y i ) ,
B 11 = i = 1 m x i , 2 ln ( y i ) , B 12 = i = 1 m x i , 3 ln ( y i ) , B 13 = i = 1 m ln ( y i ) , B 14 = i = 1 m ln 2 ( y i ) .
Utilizing Equation (8) simultaneously, β 2 and β 3 are expressed in terms of β 1 . From (5), we get
w σ B 1 = B 10 + β 1 B 4 + β 3 B 7 + β 2 B 8 ,
which expresses σ in terms of β 1 . Consequently, the positive solution to (10) is σ ^ . Then, the MLEs of β g , g = 1 , 2 , 3 are computed:
k σ 2 + w σ ( β 1 B 1 + β 2 B 2 + β 3 B 3 B 13 ) B 14 + 2 y * X * β ( X * β ) ( X * β ) = 0 .
Code C (see the Supplementary Materials) computes the MLEs of the regression parameters for the p = 3 case.
Some predictors may be insignificant. We propose using the Fisher information matrix and Wald’s test to identify them. Using the log-likelihood function l in Section “ML Estimation for Regression Parameters”, let
I β = l β I β β = I β β .
I β is a p × 1 vector. The objective is to compute IFM = ( E [ I β β ] ) 1 | β = β ^ , which is the inverse of the estimated Fisher information matrix. IFM is the variance–covariance matrix of β ^ . In this article, the Fisher scoring algorithm for the parameter estimates,
β ^ v + 1 = β ^ v ( E [ I β β ] ) 1 · I β | β = β ^ v
where v is the iteration index, is not employed because closed-form formulas are presented for the p = 2 and p = 3 cases, and for the general case p > 3 a numerical optimization in Mathematica can be used to maximize the log-likelihood and determine the MLEs. Computation of the Fisher information matrix is necessary to identify statistically significant predictors in the regression model.
It can be verified, based on the likelihood function, l, that I β β does not involve Y 1 , , Y m , Y m + 1 , , Y n ; consequently, the computation of IFM does require estimating the expected values,
E [ Y i | Y i < θ i ] = 0.303925 e σ 2 ( 0.0692809 0.138562 w ) σ w 2 θ i ( 1 + E r f ( 0.263213 0.263213 σ w ) ,
E [ Y i | Y i > θ i ] = c w θ i ( w σ ) ,
by replacing the parameters in (13) with their MLEs.
We propose using Wald’s test statistic
W g = β ^ g I F M ( g g ) , g = 1 , 2 , , p ,
where the diagonal element of IFM ( g g ) estimates the variance of β g ^ . Due to the asymptotic properties of MLEs, β ^ 1 , , β ^ p are asymptotically normally distributed when the sample size k is large. | W g | can be compared with a critical value from the standard normal distribution to identify insignificant predictors. This process continues until all remaining predictors in the model are significant.

5. Bayesian Estimation of Regression Parameters

Selecting a conjugate prior distribution can be challenging, especially when the goal is for the posterior to belong to the same class as the prior. Additionally, the posterior PDF may not be in a recognizable form, complicating the derivation of the Bayes estimator as the posterior’s expected value. In these situations, a computational approach is to use a Markov Chain Monte Carlo (MCMC) algorithm. It requires selecting a proposal distribution for the posterior, which may not belong to any well-known class of distributions. The algorithm then runs simulations until convergence is confirmed. Aminzadeh and Deng (2022) used an MCMC algorithm to obtain a Bayes estimate of the renewal function when the interarrival times follows a Pareto distribution.
Since the regression parameters β 1 , β 2 , , β p can take any value, and their MLEs based on a large sample are approximately normally distributed, we propose the multivariate normal distribution as a prior. Let
β ̲ N p ( ϕ , Σ ) , β ̲ = ( β 1 , , β p ) ,
where E ( β j ) = ϕ j , j = 1 , 2 , , p . The matrix Σ is the p × p positive-definite variance–covariance matrix of β 1 , , β p .
Using the likelihood function and (14), the posterior can be written as
i = 1 k c w σ y i ( 1 + w σ ) i = 1 m e 1 2 σ 2 ln ( y i ) x i β 2
× ( 2 π ) p / 2 [ det ( Σ ) ] 0.5 e 0.5 ( β ̲ ϕ ̲ ) Σ 1 ( β ̲ ϕ ̲ ) .
The number of hyperparameters is p + p 2 in the prior distribution. They are ϕ j ( j = 1 , 2 , , p ) as well as the covariances Cov ( β i , β j ) (for i j and i , j = 1 , 2 , , p ). Data-driven approaches based on MLEs are used in the literature to assign hyperparameter values, as in Aminzadeh and Deng (2025) and Aminzadeh and Deng (2022), when relevant information about the parameters is unavailable. This involves using MLEs to compute the hyperparameter governing the prior mean while simultaneously minimizing the prior variance. The simulation studies reported in the articles confirm that this method yields more accurate Bayesian estimates than MLEs.

5.1. Approximate Bayesian Computation

The ABC algorithm dates back to the 1980s, when Donald Rubin introduced it. Lintusaari et al. (2017) provided a comprehensive overview of recent developments in the ABC method. Researchers have employed this method, mainly when the posterior probability density function (PDF) is intractable. We turned to the ABC algorithm, which relies on extensive simulations, because the posterior in (15) is challenging to work with. It is worth noting that the proposed ABC method presented in this article is preferred over alternative Bayesian methods, such as classical Bayesian methods, because with many covariates in the regression model selecting hyperparameter values for the regression parameters is extremely difficult, if not impossible. Alternative Bayesian computational methods, such as MCMC, are not useful because the posterior distribution (15) is intractable. In addition, the multivariate normal distribution is an appropriate prior given the asymptotic properties of MLEs.

5.2. Computation Steps for ML and ABC Estimates Using Simulated Data

For selected “true” values of σ t and β t = ( β t , 1 , , β t , p ) and a chosen k × p design matrix X, we generate N S = 300 samples y j 1 , , y j k for j = 1 , 2 , , N S from the LogNormal–Pareto composite distribution. The following steps in Mathematica Code A compute the MLEs and approximate Bayesian estimates of β t 1 , , β t p :
1.
In the j t h iteration j = 1 , 2 , , N S , software such as Mathematica (as shown for p = 2 and p = 3 cases in Section 4) is used to solve
l σ = k σ 2 + σ w 1 X β σ w 1 * ln ( y ̲ * ) l n ( y ̲ * ) l n ( y ̲ * ) + ln ( y ̲ * ) X * β + β X * ln ( y ̲ * ) β X * X * β = 0 ,
l β g = w σ 1 x g l n ( y ̲ * ) x g * + β g x g * x g * h g m β h x g * x h * = 0 , g , h = 1 , 2 , , p
simultaneously to find the MLEs: σ ^ , β ^ 1 , , β ^ p and to use Wald’s test statistic to identify significant predictors to retain in the model. Assuming that all predictors are retained in the model, the MLEs based on N S samples are computed as
σ ^ ML = j = 1 N S σ ^ j N S , β ^ ML , g = j = 1 N S β ^ j g N S , β ^ M L = ( β ^ ML , 1 , , β ^ ML , p )
and with ASEs,
ξ ( σ ^ ML ) = j = 1 N S ( σ ^ j σ t ) 2 N S , ξ ( β ^ ML , g ) = j = 1 N S ( β ^ ML , g β t , g ) 2 N S , g = 1 , 2 , , p .
For the jth generated sample, j = 1 , 2 , , N S , IFM j matrix is computed, and then the overall variance–covariance matrix
IFM = j = 1 N S IFM j N S
is obtained.
2.
Generate N s s = 3000 random samples from a multivariate Normal distribution N p ( β ^ M L , IFM ) . β 1 v , , β p v is the vth generated sample, v = 1 , 2 , , N s s .
3.
Use β v = ( β 1 v , , β p v ) to generate a sample of size k from the composite distribution with the ith row x i of the matrix X and the link function θ i = e x i β v . The simulated samples are denoted by y 1 v , y 2 v , , y k v , v = 1 , 2 , , N s s .
4.
Compare the simulated samples y 1 v , y 2 v , , y k v with y j 1 , . . , y j k j = 1 , 2 , , N S using selected summary statistics. Per the ABC method, it is desired to have “small” differences in the selected summary statistics. Add the v t h simulated β v from Step 3 to the set of “accepted samples” from the posterior distribution if the difference in the selected summary statistics is less than the tolerance ϵ . We use the 30th and 90th percentiles as summary statistics. This choice reflects the characteristics of loss data in the insurance industry, which include frequent small losses and infrequent large losses. Therefore, percentiles such as the 30th and 90th are used to assess agreement between the original and simulated data. If we compare summary statistics such as means or sums, a few extreme values at the high end of the simulated and original samples can heavily distort the metric, whereas percentiles isolate values by their relative position. Therefore, for a more accurate comparison, percentiles are employed; ϵ 1 and ϵ 2 respectively denote the tolerance errors for the 30th and 90th percentiles. There are N S original samples from Step 1 and N s s samples in Step 3. In the simulation studies, all N s s samples were compared with a random sample drawn from the N S samples. This choice reflected the practical case, as in an actual application only one sample is available for analysis.
5.
Let N a denote the number of acceptable simulated samples β 1 ν , , β p ν . Define an ABC-based estimate as
β ^ Bayes , g = h = 1 N a β g h N a ,
with the ASE,
ξ ( β ^ Bayes , g ) = h = 1 N a ( β ^ Bayes , g β t , g ) 2 N a , g = 1 , 2 , , p .
The choice of the tolerance parameters ϵ 1 and ϵ 2 depends mainly on the y 1 , y 2 , , y k and the simulated sample. Small tolerance values would lead to a small N a , which is undesirable, and large tolerance values would force N a to be close to N s s , which is also undesirable, as MLE and ABC estimates would be almost identical. Before making all the comparisons in Step 4, it is recommended to print a few absolute differences between the percentiles to inform the appropriate choices of ϵ 1 and ϵ 2 . As we implement the ABC method in the code and check the acceptance status of each simulated sample, if the majority (e.g., 80%) of the first few samples are accepted or only a very few (e.g., 5%) are accepted, then the tolerance parameter values must be adjusted.

6. Simulation: Regression Model

This section presents our simulation results from the LogNormal–Pareto model using a variety of selected values for the number of covariates (p), “true” values for the regression parameters β 1 , , β p , and “true” values for σ . The objective of these simulations was to demonstrate that the proposed ABC method outperforms ML in estimator accuracy, as measured by ASE.
Using sample sizes k = 50 , k = 100 , k = 300 and β t , σ t , and p, we generated N S = 300 samples from the composite distribution to obtain the MLEs. We then generated N s s = 3000 samples from a multivariate normal distribution and implemented the ABC algorithm to obtain Bayesian estimates of the regression parameters.
In our simulation studies, for convenience, a k × p design matrix X was generated from a Uniform distribution. Notably, the same matrix X was used across all N S simulations, based on input parameters k , σ t , and β t . The choice of distribution for X and the parameter values for the Uniform distribution were irrelevant and did not affect the main results reported in the article regarding the accuracies of the ML and ABC estimates. Table 2, Table 3 and Table 4 reveal that the ABC algorithm outperformed ML, as it had smaller ASE values. Additionally, for larger values of σ , which yielded extreme observations in the generated data, the ASE of the regression parameters increased for both the ABC and the ML estimates. Figure 3, Figure 4 and Figure 5, respectively, provide a visual summary of the numbers in Table 2, Table 3 and Table 4 and support the same conclusion. From an insurance industry perspective, a small ASE in regression parameters ensures highly accurate and stable predictions of insurance loss severity, which combines heavy-tailed extreme events with frequent small claims separated by a threshold. This minimizes volatility in capital allocation and premium pricing models. Most often, insurers rely on composite distributions such as LogNormal–Pareto to model both frequent, low-cost attritional losses and rare, catastrophic tail risks. A small ASE directly limits the variance of parameter estimates, preventing severe over- or underestimation of reserve requirements.

7. Goodness-of-Fit and Sensitivity Analysis

This section discusses the goodness of fit of the generated and real data sets to the LogNormal–Pareto composite distribution.
Table 5 lists a generated sample of size 100 drawn from the LogNormal–Pareto. For this data set, using Code B, the MLEs are
σ ^ = 0.71 , θ ^ = 16.96 , m = 39 .
The Anderson–Darling (AD) test is performed in Mathematica to assess the goodness of fit of a data set to the LogNormal–Pareto model, using the command: AndersonDarlingTest[y, f, “TestDataTable”]; y is the list of observations, and f is the piecewise PDF for the LogNormal–Pareto (4). We obtained AD = 0.3680 with a p-value = 0.879452, confirming that the data set fits the composite distribution well, as expected.
To demonstrate the computations involved in our proposed method for the regression model, we considered a real data set representing the total payment (y) for auto claims, in thousands of Swedish Kronor, available at https://www.kaggle.com/datasets/redwankarimsony/auto-insurance-in-sweden (accessed on 15 May 2026). We modified the data by adding five values ( 601.2 , 820.1 , 1530.5 , 1930.5 , 2631.5 ) to y and x (# of claims) to obtain a better fit to the LogNormal–Pareto model; otherwise, the p-value for the AD test based on the original data set was 0.0390:
y = { 4.4 , 6.6 , 11.8 , 12.6 , 13.2 , 14.6 , 14.8 , 15.7 , 20.9 , 21.3 , 23.5 , 27.9 , 31.9 , 32.1 , 38.1 , 39.6 , 39.9 , 40.3 , 46.2 , 48.7 , 48.8 , 50.9 , 52.1 , 55.6 , 56.9 , 57.2 , 58.1 , 59.6 , 65.3 , 69.2 , 73.4 , 76.1 , 77.5 , 77.5 , 87.4 , 89.9 , 92.6 , 93 . , 95.5 , 98.1 , 103.9 , 113 , 119.4 , 133.3 , 134.9 , 137.9 , 142.1 , 152.8 , 161.5 , 162.8 , 170.9 , 181.3 , 187.5 , 194.5 , 202.4 , 209.8 , 214 . , 217.6 , 244.6 , 248.1 , 392.5 , 422.2 , 601.2 , 820.1 , 1530.5 , 1930.5 , 2631.5 }
x 1 = the vector of 1 s taking into account the intercept in the regression model .
x 2 = { 3 , 2 , 4 , 4 , 3 , 6 , 6 , 13 , 5 , 11 , 11 , 7 , 13 , 15 , 4 , 23 , 13 , 5 , 19 , 9 , 7 , 6 , 9 , 18 , 23 , 11 , 12 , 16 , 10 , 25 , 41 , 8 , 7 , 14 , 19 , 13 , 27 , 13 , 14 , 20 , 29 , 24 , 24 , 17 , 37 , 22 , 55 , 57 , 41 , 26 , 30 , 35 , 31 , 45 , 16 , 33 , 48 , 29 , 44 , 36 , 15 , 23 , 32 , 35 }
Using Code B, we got
θ ^ = 44.8571 , σ ^ = 0.689201 , m = 19 , n = 67
and via Mathematica, AD = 1.8873, with p-value = 0.1062.
Goodness-of-fit (GOF) of the data set was also assessed for several composite distributions proposed in the literature. The MLEs of the corresponding parameters for each distribution were obtained, and BIC values and Anderson–Darling (AD) test p-values were computed. Although the BIC value for LogNormal was higher than Gamma–Pareto and InverseGamma–Pareto, Table 6 shows that the AD p-values for the Inverse-Gamma–Pareto, Exponential–Pareto, and Gamma–Pareto composite models were very small, indicating poor fit to the data. Based on the AD p-value, the LogNormal–Pareto model was a reasonable fit to the data set. It is worth mentioning that the Anderson–Darling (AD) test was more sensitive to the selected distribution. Unlike the Bayesian Information Criterion (BIC), which primarily penalizes model complexity to avoid overfitting, the AD test focuses exclusively on goodness of fit and places heavy weight on discrepancies in the tails of the distribution.
Consequently, at the 0.01 and 0.05 significance levels, the claim amount was concluded to follow the LogNormal–Pareto model.
As for the regression model, the MLEs for the parameters were
β ^ 0 = 1.3374 , β ^ 1 = 0.24915 ,
with
IFM = 0.0273398 0.002169 0.002169 0.000246772 , Wald test = ( 8.0884 , 15.8607 ) .
The Wald test statistics indicate that the covariate (# of claims) is significant and that β 0 should be retained in the model.
In the context of regression analysis, using the significant covariate x 1 (the number of claims) in the model, V a R ( 1 γ ) (see Section 2) can be computed. For example, with ( 1 γ ) = 0.95 , σ ^ = 0.689201 ( α = w σ = 0.5401 ) , θ ^ = e β ^ 0 + β ^ 1 x 1 , and the number of claims x 1 = 12 , we get V a R ( 0.95 ) = 7724.05 .
Following the method proposed in the article, 1000 samples were generated from the multivariate Normal distribution with mean β ^ = ( 1.3374 , 0.24915 ) and the variance–covariance matrix, IFM in (17). The ABC-based estimates were computed as
β ^ 0 , B a y e s = 1.33979 , β ^ 1 , B a y e s = 0.250318 , N a = 426 .
Code D (see the Supplementary Materials) provided computations for the ABC estimates of the regression parameters.
The original data, which excluded the last 5 numbers at the end of the vector y above and the last 5 rows of the design matrix X above, was also analyzed, and the results are given below.
θ ^ = 43.5786 , σ ^ = 0.586609 , m = 18 , n = 62
and via Mathematica, AD = 2.77657 with p-value = 0.035766. We can see that the original data did not fit the LogNormal–Pareto model as well as the modified data, as indicated by the smaller p-value. The parameter estimates differed slightly, and m = 18 . The MLEs for the parameters were
β ^ 0 = 1.35251 , β ^ 1 = 0.235205
with
IFM = 0.025064 0.002127 0.002127 0.000258 , Wald test = ( 8.5431 , 15.7428 )
Note that the MLEs of the regression parameters and the FIM did not change drastically.
The ABC-based estimates were
β ^ 0 , B a y e s = 1.38892 , β ^ 1 , B a y e s = 0.248915 , N a = 520 .
Sensitivity Analysis
In the following, the sensitivity of the proposed method was investigated with respect to the model assumption about the loss data.
1.
Instead of the LogNormal–Pareto model, the Exponential–Pareto model was considered for the original claim data with n = 62 observations. The Exponential–Pareto composite model’s PDF is
f Y ( y ) = 0.775 θ e 1.35 y θ 0 < y θ 0.2 θ 35 y 1.35 θ y <
Using the data and the algorithm given in Aminzadeh and Deng (2018), we obtained the MLE θ ^ = 77.1118 , m = 34 , and AD = 50.9115, with p-value = 0, indicating that the data did not fit the Exponential–Pareto model. Using the algorithm proposed in the current article, the MLEs for the regression parameters were computed as
β ^ 0 = 1.58439 , β ^ 1 = 1.00246 .
It is noted that β ^ 1 , in particular, was markedly different from its value under the LogNormal–Pareto model. In addition, the expected values, similar to (13), of the Exponential–Pareto model did not converge. As a result, the IFM could not be computed under the Exponential–Pareto model assumption. Therefore, the proposed method for computing ABC estimates of regression parameters could not be applied either. It is concluded that the requirements for applying the proposed ML and ABC methods are that the data should fit the composite model and that the expected values of the composite model should exist to apply ABC via IFM.
2.
A data set of size n = 300 generated from the LogNormal–Pareto model was used to compute regression estimates based solely on the LogNormal PDF. The same link function θ i = e x i β was used with p = 3 . The likelihood optimization did not converge, so the MLEs were not computable, implying that the proposed method in the article is sensitive to data characteristics and that the appropriate distributional model must be used.
  • Summary
We introduced a nonlinear regression model using the LogNormal–Pareto composite distribution. From an insurance industry perspective, precise regression parameters yield highly accurate, stable predictions of insurance loss severity that combine heavy-tailed extreme events with frequent small claims, separated by a threshold. This minimizes volatility in capital allocation and premium pricing models. An exponential link function is proposed to compute MLEs and ABC-based estimates for regression parameters. Our approach provides closed-form formulas for the MLEs of the regression parameters in the cases with two and three covariates. Wald’s test statistic, based on the Fisher information matrix, is used to identify insignificant covariates. A multivariate normal distribution is used as the prior for the regression parameters, and its parameters are estimated via MLE. The Anderson–Darling goodness-of-fit test and the Bayesian information criterion (BIC) assess the fit between the response variable and the composite distribution. Our simulations confirmed that the ABC method outperforms the ML method, yielding lower ASE values. Simulated data from the LogNormal–Pareto distribution and a real-world data set were used as illustrative examples for computations. The computation of VaR was also illustrated in the context of regression modeling on the real data set. It was demonstrated that the proposed method was sensitive to the model assumption for the loss data. When incorrect distributions, such as Exponential–Pareto or LogNormal models, were applied to data generated by a LogNormal–Pareto model, they led to computational challenges for the MLEs of the regression parameters, as the likelihood function of the incorrect model failed to converge, rendering the MLEs infeasible. The choice of tolerance parameter for the ABC method is crucial. The selection of the tolerance parameters ϵ 1 and ϵ 2 for the application of the ABC method depends mainly on the available y 1 , y 2 , , y k and the simulated data based on generated regression parameter values from the posterior distribution. To better select the tolerance parameters, it is suggested to print a few values of the absolute differences between the percentiles to inform the appropriate choices for ϵ 1 and ϵ 2 .
Potential future research involves modeling loss data using composite models and modeling the threshold parameter of the composite distribution via time-series models.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/risks14070165/s1, Mathematica Codes A-D for computations of ML and ABC estimates presented in the article.

Funding

This research received no external funding.

Data Availability Statement

The data presented in this study are openly available at https://www.kaggle.com/datasets/redwankarimsony/auto-insurance-in-sweden, accessed on 10 July 2026.

Conflicts of Interest

The author declares no conflicts of interest.

References

  1. Aminzadeh, Mostafa S., and Min Deng. 2018. Bayesian Predictive Modeling for Exponential-Pareto Composite Distribution. Variance Journal 12: 59–68. [Google Scholar] [CrossRef] [Scilit]
  2. Aminzadeh, Mostafa S., and Min Deng. 2022. Bayesian estimation of renewal function based on Pareto-distributed Inter-Arrival times via an MCMC algorithm. Variance 15. [Google Scholar] [CrossRef] [Scilit]
  3. Aminzadeh, Mostafa S., and Min Deng. 2025. Extending Approximate Bayesian Computation to Non-Linear Regression Models: The Case of Composite Distributions. Risks 13: 220. [Google Scholar] [CrossRef] [Scilit]
  4. Cooray, Kahadawala, and Malwane M. A. Ananda. 2005. Modeling actuarial data with a composite Lognormal-Pareto model. Scandinavian Actuarial Journal 5: 321–34. [Google Scholar] [CrossRef] [Scilit]
  5. Deng, Min, Mostafa Aminzadeh, and Min Ji. 2021. Bayesian Predictive Analysis of Natural Disaster Losses. Risks 9: 12. [Google Scholar] [CrossRef] [Scilit]
  6. Grün, Bettina, and Tatjana Miljkovic. 2019. Extending composite loss models using a general framework of advanced computational tools. Scandinavian Actuarial Journal 8: 642–60. [Google Scholar] [CrossRef] [Scilit]
  7. Klugman, Stuart A., Harry H. Panjer, and Gordon E. Willmot. 2012. Loss Models from Data to Decisions, 3rd ed. New York: John Wiley. [Google Scholar]
  8. Konşuk Ünlü, Hande. 2022. A new composite Lognormal-Pareto type II regression model to analyze household budget Data via Particle Swarm Optimization. Soft Computing 26: 2391–408. [Google Scholar] [CrossRef] [Scilit]
  9. Kumi, Williams, Henry Otoo, Charles Kwofie, and Sampson Takyi Appiah. 2026. Composite Distributions and their Associated Risk Measures for Automobile Insurance Claims Data. Journal of Statistics Applications & Probability 15: 203–17. [Google Scholar] [CrossRef] [Scilit]
  10. Li, Jackie, and Jia Liu. 2023. Claims Modeling with Three-Component Composite Models. Risks 11: 196. [Google Scholar] [CrossRef] [Scilit]
  11. Lintusaari, Jarno, Michael U. Gutmann, Ritabrata Dutta, Samuel Kaski, and Jukka Corander. 2017. Fundamentals and recent developments in approximate Bayesian Computation. Systematic Biology 66: 66–82. [Google Scholar]
  12. Liu, Bowen, and Malwane M. A. Ananda. 2023. Analyzing insurance data with an exponentiated composite inverse Gamma-Pareto model. Communications in Statistics–Theory and Methods 52: 7618–31. [Google Scholar]
  13. Marambakuyana, Walena Anesu, and Sandile Charles Shongwe. 2024. Composite and Mixture Distributions for Heavy-Tailed Data—An Application to Insurance Claims. Mathematics 12: 335. [Google Scholar] [CrossRef] [Scilit]
  14. Mutali, Sezen, and Raluca Vernic. 2022. On the composite Lognormal–Pareto distribution with uncertain threshold. Communications in Statistics–Simulation and Computation 51: 4492–508. [Google Scholar]
  15. Sarabia, José María, and Enrique Calderín-Ojeda. 2018. Analytical expressions of risk quantities for composite models. Journal of Risk Model Validation 12: 75–101. [Google Scholar] [CrossRef] [Scilit]
  16. Teodorescu, Sandra, and Raluca Vernic. 2013. On composite Parerio models. Mathematical Reports 15: 11–29. [Google Scholar]
Figure 1. LogNormal–Pareto PDF with selected values of α and θ .
Figure 1. LogNormal–Pareto PDF with selected values of α and θ .
Risks 14 00165 g001
Figure 2. Plot of ASE for selected values of the parameters.
Figure 2. Plot of ASE for selected values of the parameters.
Risks 14 00165 g002
Figure 3. Plot of ξ versus selected “true” values of parameters with sample size k = 50 .
Figure 3. Plot of ξ versus selected “true” values of parameters with sample size k = 50 .
Risks 14 00165 g003
Figure 4. Plot of ξ versus selected “true” values of parameters with sample size k = 100 .
Figure 4. Plot of ξ versus selected “true” values of parameters with sample size k = 100 .
Risks 14 00165 g004
Figure 5. Plot of ξ versus selected “true” values of parameters with sample size k = 300 .
Figure 5. Plot of ξ versus selected “true” values of parameters with sample size k = 300 .
Risks 14 00165 g005
Table 1. Accuracy of ML estimator for θ and α .
Table 1. Accuracy of ML estimator for θ and α .
n α t α ^ ¯ ξ ( α ^ ) θ t θ ^ ¯ ξ ( θ ^ )
1000.50.5047(0.0019)22.0297(0.0835)
1500.50.5014(0.0010)22.0091(0.0457)
1000.50.5110(0.0022)1010.1583(1.6353)
1500.50.5019(0.0016)1010.0368(1.1477)
1000.50.5111(0.0023)1515.2412(4.7998)
1500.50.5058(0.0018)1515.4338(2.7620)
1001.51.5145(0.0209)22.0025(0.0085)
1501.51.5105(0.0130)21.9992(0.0064)
1001.51.5355(0.0232)1010.042(0.2242)
1501.51.5271(0.0145)109.9592(0.1022)
1001.51.5187(0.0193)1514.9662(0.4716)
1501.51.4880(0.0144)1515.0986(0.3921)
Table 2. Comparison of ML and ABC accuracies for k = 50.
Table 2. Comparison of ML and ABC accuracies for k = 50.
β 1 = 0.2 β 2 = 0.5
ML Bayes
σ t σ ^ ¯ M L ξ σ ^ M L E β 1 ^ ¯ M L E ξ β 1 M L E β 2 ^ ¯ M L E ξ β 2 M L E β 1 ^ ¯ B a y e s ξ β 1 B a y e s β 2 ^ ¯ B a y e s ξ β 2 B a y e s
0.60.5979(0.0045)0.2185(0.0531)0.5283(0.0543)0.2247(0.0375)0.5174(0.0451)
1.21.1814(0.0235)02224(0.1745)0.5164(0.1172)0.2344(0.0540)0.4679(0.0675)
β 1 = 1.2 β 2 = 0.5
σ t σ ^ ¯ M L ξ σ ^ M L E β 1 ^ ¯ M L E ξ β 1 M L E β 2 ^ ¯ M L E ξ β 2 M L E β 1 ^ ¯ B a y e s ξ β 1 B a y e s β 2 ^ ¯ B a y e s ξ β 2 B a y e s
0.60.7064(0.0198)1.6020(0.1025)0.4174(0.0384)1.5726(0.0252)0.4666(0.0241)
1.21.2431(0.0320)1.3773(0.1858)0.5034(0.0494)1.3738(0.0806)0.5198(0.0347)
β 1 = 0.6 β 2 = 1.5
σ t σ ^ ¯ M L ξ σ ^ M L E β 1 ^ ¯ M L E ξ β 1 M L E β 2 ^ ¯ M L E ξ β 2 M L E β 1 ^ ¯ B a y e s ξ β 1 B a y e s β 2 ^ ¯ B a y e s ξ β 2 B a y e s
0.60.6832(0.0198)0.6402(0.0484)1.6648(0.0689)0.6494(0.0255)1.6615(0.0370)
1.21.2551(0.0313)0.6465(0.2012)1.6385(0.1358)0.6204(0.0604)1.6433(0.0492)
β 1 = 1.2 β 2 = 1.5
σ t σ ^ ¯ M L ξ σ ^ M L E β 1 ^ ¯ M L E ξ β 1 M L E β 2 ^ ¯ M L E ξ β 2 M L E β 1 ^ ¯ B a y e s ξ β 1 B a y e s β 2 ^ ¯ B a y e s ξ β 2 B a y e s
0.60.6699(0.0123)1.2905(0.0792)1.6831(0.1068)1.2276(0.0329)1.6993(0.0477)
1.21.2564(0.0234)1.4067(0.1471)1.5512(0.1126)1.3774(0.0603)1.5700(0.0837)
Table 3. Comparison of ML and ABC accuracies for k = 100.
Table 3. Comparison of ML and ABC accuracies for k = 100.
β 1 = 0.2 β 2 = 0.5
ML Bayes
σ t σ ^ ¯ M L ξ σ ^ M L E β 1 ^ ¯ M L E ξ β 1 M L E β 2 ^ ¯ M L E ξ β 2 M L E β 1 ^ ¯ B a y e s ξ β 1 B a y e s β 2 ^ ¯ B a y e s ξ β 2 B a y e s
0.60.6260(0.0042)0.2023(0.0201)0.5467(0.0161)0.2185(0.0144)0.5323(0.0067)
1.21.1814(0.0126)0.1842(0.0395)0.5085(0.0277)0.2175(0.0153)0.5209(0.0093)
β 1 = 1.2 β 2 = 0.5
σ t σ ^ ¯ M L ξ σ ^ M L E β 1 ^ ¯ M L E ξ β 1 M L E β 2 ^ ¯ M L E ξ β 2 M L E β 1 ^ ¯ B a y e s ξ β 1 B a y e s β 2 ^ ¯ B a y e s ξ β 2 B a y e s
0.60.6791(0.0101)1.5607(0.2019)0.4313(0.0212)1.2548(0.0107)0.4349(0.0093)
1.21.2397(0.0126)1.3955(0.2239)0.4734(0.0404)1.2828(0.0328)0.4602(0.0205)
β 1 = 0.6 β 2 = 1.5
σ t σ ^ ¯ M L ξ σ ^ M L E β 1 ^ ¯ M L E ξ β 1 M L E β 2 ^ ¯ M L E ξ β 2 M L E β 1 ^ ¯ B a y e s ξ β 1 B a y e s β 2 ^ ¯ B a y e s ξ β 2 B a y e s
0.60.6621(0.0069)0.5855(0.0164)1.6616(0.0448)0.5717(0.0095)1.6699(0.0331)
1.21.2809(0.0164)0.6030(0.0545)1.6405(0.0638)0.5750(0.0359)1.6534(0.0512)
β 1 = 1.2 β 2 = 1.5
σ t σ ^ ¯ M L ξ σ ^ M L E β 1 ^ ¯ M L E ξ β 1 M L E β 2 ^ ¯ M L E ξ β 2 M L E β 1 ^ ¯ B a y e s ξ β 1 B a y e s β 2 ^ ¯ B a y e s ξ β 2 B a y e s
0.60.6995(0.0152)1.4424(0.1442)1.6078(0.0489)1.4682(0.0856)1.5766(0.0150)
1.21.2734(0.0130)1.4983(0.2175)1.5071(0.0599)1.3952(0.0957)1.4967(0.0207)
Table 4. Comparison of ML and ABC accuracies for k = 300.
Table 4. Comparison of ML and ABC accuracies for k = 300.
β 1 = 0.8 β 2 = 1.5 β 3 = 0.9
ML Bayes
σ t σ ^ ¯ M L ξ σ ^ M L E β 1 ^ ¯ M L E ξ β 1 M L E β 2 ^ ¯ M L E ξ β 2 M L E β 3 ^ ¯ M L E ξ β 3 M L E β 1 ^ ¯ B a y e s ξ β 1 B a y e s β 2 ^ ¯ B a y e s ξ β 2 B a y e s β 3 ^ ¯ B a y e s ξ β 3 B a y e s
0.50.5195(0.0003)0.8471(0.0281)1.5544(0.0184)0.8825(0.0245)0.8326(0.0126)1.5649(0.0108)0.8811(0.0201)
1.21.2157(0.0191)0.7790(0.2172)1.5481(0.1009)0.9220(0.1909)0.7754(0.1045)1.5912(0.0757)0.8532(0.1120)
β 1 = 1.8 β 2 = 0.6 β 3 = 2.0
ML Bayes
σ t σ ^ ¯ M L ξ σ ^ M L E β 1 ^ ¯ M L E ξ β 1 M L E β 2 ^ ¯ M L E ξ β 2 M L E β 3 ^ ¯ M L E ξ β 3 M L E β 1 ^ ¯ B a y e s ξ β 1 B a y e s β 2 ^ ¯ B a y e s ξ β 2 B a y e s β 3 ^ ¯ B a y e s ξ β 3 B a y e s
0.50.5195(0.0003)2.0918(0.1226)0.4295(0.0499)2.1874(0.0823)2.1039(0.1137)0.4194(0.0472)2.1861(0.0583)
1.21.2139(0.0157)1.9169(0.1570)0.5251(0.0771)2.1108(0.1336)1.9773(0.1350)0.5460(0.0609)2.0184(0.1188)
β 1 = 1.5 β 2 = 1.3 β 3 = 2.0
ML Bayes
σ t σ ^ ¯ M L ξ σ ^ M L E β 1 ^ ¯ M L E ξ β 1 M L E β 2 ^ ¯ M L E ξ β 2 M L E β 3 ^ ¯ M L E ξ β 3 M L E β 1 ^ ¯ B a y e s ξ β 1 B a y e s β 2 ^ ¯ B a y e s ξ β 2 B a y e s β 3 ^ ¯ B a y e s ξ β 3 B a y e s
0.51.2258(0.0122)1.5300(0.1415)1.2879(0.0974)2.1133(0.1248)1.6466(0.1271)1.2524(0.0592)2.0875(0.1196)
1.20.6013(0.0263)1.8275(0.1922)0.9343(0.1802)2.6341(0.4903)1.8447(0.1514)0.9349(0.1560)2.6180(0.4114)
Table 5. Generated sample from LogNormal–Pareto ( θ t = 15 , σ t = 0.6).
Table 5. Generated sample from LogNormal–Pareto ( θ t = 15 , σ t = 0.6).
4.163864.825155.66116.173226.904037.256877.278967.294327.435427.45443
7.533167.689817.742067.759078.236948.562529.133059.3355210.038310.4195
10.469610.519310.872212.654112.683312.726313.168313.847314.309514.3217
14.552514.963514.99115.451216.053316.482416.614416.827716.949217.3335
18.847720.109620.404920.719720.795722.123324.323125.042825.169626.2846
26.728626.757926.850228.900829.237829.344430.372930.394530.430231.3167
31.627531.911132.168134.940235.542838.53839.418840.092543.418447.8926
47.92348.311749.202751.413354.419463.119865.43272.772373.312574.3392
100.284119.793122.711135.812138.105138.803145.339171.228213.533242.595
308.472434.651527.658563.648652.523774.342833.141477.941908.152089.69
Table 6. GOF of composite models to the data set based on BIC and AD.
Table 6. GOF of composite models to the data set based on BIC and AD.
LogNormal–ParetoExponential–ParetoGamma–ParetoInverseGamma–Pareto
BIC835.4931185.13402.434766.853
AD test1.88737.775621.45644.7602
AD-p-value0.10620.00010.00000.0037
m19393214
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Aminzadeh, M.S. Modeling Data with Nonlinear LogNormal–Pareto Regression via the Approximate Bayesian Computation. Risks 2026, 14, 165. https://doi.org/10.3390/risks14070165

AMA Style

Aminzadeh MS. Modeling Data with Nonlinear LogNormal–Pareto Regression via the Approximate Bayesian Computation. Risks. 2026; 14(7):165. https://doi.org/10.3390/risks14070165

Chicago/Turabian Style

Aminzadeh, Mostafa S. 2026. "Modeling Data with Nonlinear LogNormal–Pareto Regression via the Approximate Bayesian Computation" Risks 14, no. 7: 165. https://doi.org/10.3390/risks14070165

APA Style

Aminzadeh, M. S. (2026). Modeling Data with Nonlinear LogNormal–Pareto Regression via the Approximate Bayesian Computation. Risks, 14(7), 165. https://doi.org/10.3390/risks14070165

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop