1. Introduction
The accurate estimation of population parameters is critical in survey sampling, which is used to make valid inferences about a target population using sample data. One of these parameters is the population variance, which is important for interpreting the dispersion of a variable and is required in many statistical problems, such as the construction of a confidence interval, hypothesis testing, sample sizes, and quality control. A good and restricted approximation of population variance ensures more reliability of the inferences and less uncertainty of data-driven decisions. One of the best sampling techniques that is used is simple random sampling (SRS) because it is simple, easy to use, and has positive statistical characteristics. In SRS, each unit of the population has a balanced and independent probability of being selected. Although SRS is used in the design of most inferential methods, the computation of population variance in this arrangement is often performed using conventional unbiased or nearly unbiased estimators, including the sample variance.
Even though these classical estimators are effective in ideal situations, they can be limited in regard to efficiency, particularly in cases with available auxiliary data. In a real survey situation, auxiliary information, which can be quantitative or qualitative data on the variable under study, is likely to be available in prior surveys, censuses, administrative data, or any other reliable sources. These would be demographic information, income, crop production, or any other useful variable that provides a statistically significant correlation with the variable of interest. Auxiliary information, when used properly, can dramatically enhance the efficiency of estimators, resulting in improved accuracy without any increment in the sample size. Although the use of such information has been well-researched in the context of mean estimation, its use in variance estimation has not been explored as much.
Using two auxiliary variables, [
1] recommended an improved estimator for the estimation of population variance. Ref. [
2] developed an improved estimator for the estimation of population variance under stratified random sampling. Using dual auxiliary information, a new approach is discussed for the estimation of population variance in [
3]. The authors in [
4] discussed population variance using real-life data and a simulation study. Dual use of auxiliary information for estimating population variance is performed by [
5]. A generalized estimation for the estimation of population variance is given in detail in [
6]. A generalized class of estimators for the estimation of variance using auxiliary information is given in [
7]. A transformed log-type estimator for the estimation of population variance is given in [
8]. Based on successive sampling, a new approach for the estimation of variance under stratified random sampling is provided in [
9].
Recent developments in sampling theory have placed greater emphasis on the need to establish more general estimation methods that readily take into account the auxiliary detail in adaptable manners. The generalized estimators are usually built by adjusting the traditional estimators with the known population parameters of the auxiliary variable (the mean, variance, skewness, kurtosis, or any other). Such estimators can be optimized to play well in given conditions by reducing mean squared error (MSE) or improving efficiency. Here, the generalized estimation frameworks in the problem of population variance estimation are a prospective avenue of enhancing the statistical inference of SRS. The reason behind performing this study is the necessity of overcoming inefficiencies of traditional variance estimators by incorporating the auxiliary information, which is easily accessible but overlooked. The research will help fill the gap in the literature by providing improved generalized estimators, which would be useful in both theoretical statistics and applications in survey methodology. In particular, the proposed estimators will minimize MSE and be more efficient in relation to the currently available estimators, making them more dependable in providing information about the variance in real-world sampling situations. This paper proposes a new class of generalized statistical estimators of the population variance for simple random sampling that uses auxiliary information to improve the accuracy of the estimations in an organized manner. The overall structure of the proposed estimators is obtained based on the known population parameters of the auxiliary variable, and their theoretical characteristics, like bias and MSE, are discussed in detail. The circumstances under which the proposed estimators outperform the existing ones are derived through a comparative analysis that is performed both analytically and empirically.
In order to evaluate the performance of the proposed estimators, a set of numerical examples is given with the use of real population data. The numerical results show that the proposed methods are superior based on the lower MSE and increasing the relative efficiency in a wide variety of population settings and parameter organizations. Also, sensitivity studies are performed to determine the strength of the proposed estimators as far as variations in population features (sample size, variability, and correlation between the auxiliary and the study variable) are concerned. The practical implications of the study findings are very important to the statisticians, researchers, and policy makers concerned with designing and analyzing sample surveys. Better efficiency in variance estimation, especially in the context of limited resources or high data collection costs, can save significant costs and yield more valid conclusions in most applications. The suggested methodology will offer a feasible guideline for using auxiliary information more effectively, thereby improving the overall quality of statistical inference in finite population sampling.
In addition to being a methodologically important study that creates a new class of estimators, this study also provides suggestions on how to apply it in practice to a survey field. This work gives practitioners the power to make informed decisions when choosing the methods of estimating the variances by providing specific circumstances that will result in better estimators being applied. The combination of theoretical derivations and empirical validations ensures that the methods suggested are statistically sound and can be implemented practically.
1.1. Novelty over Existing Work
Even though population variance estimation has been extensively examined in the context of simple random sampling (SRS), most traditional and modern methods make substantial use of the sample data itself without much or efficient use of additional information. Although the use of auxiliary information has been adopted widely in the methods of estimating the mean, including ratio estimators, regression estimators, and exponential estimators, little has been performed on its use in estimating the variance. In some of the existing literature, better variance estimators have been suggested based on the use of auxiliary information, although these are often restricted to some type, e.g., ratio-type or regression-type estimators, and usually assume known auxiliary means only. The techniques are not normally flexible, and they tend to be optimized under constraining assumptions, such as linear correlations among the research and auxiliary variables or homoscedastic error models. Additionally, many current estimators do not fully exploit the potential of unknown auxiliary parameters, including the coefficient of variation, kurtosis, and skewness.
The key novelty of the current study over the existing literature is as follows:
In this article, a novel, generalized, and flexible class of estimators is introduced using a variety of known values of auxiliary information. These are the correlation, coefficient of variation, skewness, and kurtosis, which enable it to be more extensively applicable and more accurate in various population structures.
Contrary to many other previous works in which the estimators are suggested without any rigorous conditions on performance, our work derives the expression of the bias and the mean squared error (MSE) of the proposed estimators in a rigorous manner. This can help us determine whether the exact circumstances under the proposed methods can be better than the existing ones.
The suggested work provides a single model that includes traditional estimators (the conventional sample variance, ratio, and regression-type estimators) as special cases. This consistency not only furthers theoretical knowledge but also provides a versatile mechanism for practitioners.
The proposed estimators are demonstrated to be strong when used in different conditions through extensive real-life data using different sample sizes, levels of correlation, and skewness of the population. This strength contrasts with the low optimality range of most previous estimators.
Empirical comparisons on real and artificial populations indicate that the estimation methods suggested will always be characterized by a higher relative efficiency and a lower MSE than the more traditional and some of the more recent estimation methods, even in situations where the auxiliary variable is only modestly related to the study variable.
Unlike some previous estimators, which are mathematically beautiful but hard to compute in the real world, the suggested family of estimators is easy to compute, requires readily available auxiliary information, and can be integrated into standard survey processes.
Overall, the study contributes to the existing body of knowledge by suggesting a more effective and comprehensive way of estimating population variance with simple random sampling. Using a wider scope of auxiliary information and providing the theoretical and empirical justification, this work provides an important gap in the literature on the topic of survey sampling and also provides practical measures to apply in real-life situations.
1.2. Objectives of the Study
The main aim of this research is to construct and compare better statistical estimates of the population variance, using the simple random sampling (SRS) design, by effectively using auxiliary information. Due to the shortcomings of the traditional variance estimators and the excessive use of the existing literature in terms of using the available auxiliary information, this research undertaking attempts to address the gap in the methodology and creates a more generalized and robust class of variance estimators.
The following are the specific objectives of the study:
To critically examine existing estimators of population variance when using simple random sampling, using auxiliary information, and to find their limitations in terms of both efficiency and applicability.
To build a generalized class of estimators of the population variance by including the known parameters of a substitute variable, i.e., the population mean, coefficient of variation, skewness, and kurtosis.
To obtain the theoretical properties of the proposed estimators, such as expressions of bias and mean squared error (MSE), and in which cases they are superior to conventional estimators.
To make a comparative analysis of the performance of proposed estimators and existing estimators through theoretical analysis, emphasizing the enhancement in efficiency and robustness.
To confirm the practical usefulness of the proposed estimators with real-life population data in order to show that they can be useful in applied survey research.
To offer suggestions on the choice and use of variance estimators in various survey conditions, and thus make practical mechanisms for the sampling theory and survey methodology.
The rest of this paper will have the following structure:
Section 2 provides notations and methodology related to population variance based on simple random sampling.
Section 3 provides a review of the existing literature in the framework of SRS using auxiliary information.
Section 4 presents the suggested class of generalized estimators, its mathematical formulation, and the theoretical character.
Section 5 contains an extensive empirical analysis in terms of real data application.
Section 6 covers the numerical findings and their practical implications for the study.
Section 7 is the final part of the paper that will give ideas and recommendations regarding future research.
The goal of this work is to develop a novel generalized class of estimators for the estimation of population variance using auxiliary variables.
2. Methodology and Symbols
Consider a finite population = consisting of N distinct and identified units. A sample of size n was selected from with the help of simple random sampling without replacement (SRSWOR). The study and auxiliary information for the population are denoted with the help of Y and X. The population variances for the study and auxiliary information are represented with the help of and , and their sample variances are denoted with the help of and . Mathematically, we represented the sample and population variance as given by the following:
= represents the population variances of the study variable Y.
= represents the population variances of the auxiliary variable X.
and
= represents the sample variances of the study variable y.
= represents the sample variances of the auxiliary variable x.
= , = represents the population standard deviation of the study and auxiliary variable.
= , = represents the population means of the study and auxiliary variables.
= , = represents the sample means of the study and auxiliary variables.
= represents population covariance between the study and auxiliary variables.
= and = represents the coefficients of variation of the study and auxiliary variables.
To obtain the expressions for the bias and mean squared error of the existing and the proposed class of estimators, we used the following error terms, which are given by the following:
where
Here,
and
are the population coefficients of kurtosis:
where
r and
s are the positive numbers, and
is the moment ratio.
3. Some Adopted Existing Estimators for Variance
In this section, we discussed some adopted estimators for the estimation of population variance under simple random sampling using auxiliary information, which are given as follows:
- (i)
The usual variance estimator is given by the following:
- (ii)
The traditional ratio estimator suggested by the following [
1]:
The bias and variance of
is given by the following:
and
- (iii)
The regression estimator, along with variance, is given by the following:
- (iv)
The exponential estimator developed by [
2] for variance, along with bias and mean squared error, is given by the following:
The bias and mean squared error of
are given by the following:
and
- (v)
The enhanced exponential type estimator developed by [
3] is given by the following:
The bias and mean squared error of
are given by the following:
and
where
- (vi)
The enhanced estimators for estimation of variance developed by [
4] are given by the following:
The properties of
,
, and
are given by the following:
and
where
and
where
and
where
4. Suggested Generalized Class of Estimators
It is necessary to estimate population parameters accurately to make informed decisions, articulate policies, and conduct scientific studies in survey sampling. Although much focus has been paid to the estimation of the population mean, the estimation of the population variance, which is one of the essential measures of dispersion and reliability of the data, has not been studied as thoroughly, particularly when it comes to the use of auxiliary information. Conventional variance estimation under a simple random sampling (SRS) is based on conventional estimators (such as sample variance), which, though unbiased, are not generally very efficient in terms of mean squared error (MSE). These estimators fail to take advantage of auxiliary information, which is often available in real-world surveys (as in forms of census, administrative records, or prior studies). With the prevalence and potential of such information, there is a great desire to use it in the estimation process to enhance accuracy and minimize sampling error.
The existing literature contains a number of ratio-type and regression-type estimators, which are aimed at utilizing some auxiliary information; however, these methods are commonly constrained to using only auxiliary means. They can also be sub-optimal, where the relationship between the study and the auxiliary information is not linear and other auxiliary parameters are known, including the coefficient of variation, skewness, or kurtosis, and may be used to do a better estimation. In addition, in most practical cases, it is either expensive or not possible to increase the size of a sample to enhance the efficiency of estimators. This also highlights the necessity of methodologically efficient techniques that can increase the accuracy of estimations without having to incur extra cost in data collection. This can be achieved by a well-designed generalized estimator that effectively incorporates known auxiliary characteristics at lower cost and greater statistical efficiency. Hence, the idea of this research is to accomplish the following:
Overcome the shortcomings of the existing estimators of variance.
Apply a greater variety of auxiliary information.
Develop a new generalized and flexible class of estimators.
Increase the efficiency, strength, and usefulness of population variance estimation under SRS.
The authors in [
1] developed an improved estimator for the estimation of population variance using two auxiliary pieces of information based on simple random sampling. Taking motivation from [
1], we developed an improved generalized class of estimators using auxiliary information for estimating population variance. The aim of this research is to add value to both theory and practice by offering better tools of inference that are not merely mathematically correct but also applicable in general in real-world survey situations. This is motivated by the aim to close the gap between the theoretical possibility and its actual application in statistical surveys. The recommended class of estimator is given by the following:
where G and
represents scalers that assume real values, and
and
are unknown constants.
To obtain the bias and mean squared error of the suggested class of estimators, we used the following error terms:
where
and
Here,
and
are the population coefficient of kurtosis:
where
and
After simplification of (32), we obtain the following expression:
Apply expectation on both sides of (33), and we have the following:
Apply the expected values (34) and simplify, and we obtain a bias of
:
For the MSE expression, we take expectation and Square (33), and we obtain the following:
where
and
To find out the unknown values of
and
, minimizing (37), we have the following:
and
Putting
and
in (36), we obtain the minimum MSE of
, and we have the following:
Simplify (41), and we obtain the minimum MSE of
:
5. Numerical Study
To check the efficiency of estimators, we used two real data sets, and their source, summary statistics, the details of mean squared error, and percentage of relative efficiency of the estimators are given in tables.
Population-I: [Source: [
10]].
Population-II: [Source: [
11]].
6. Simulation Study
In this section, we conduct a simulation study to check the performance of the suggested class of estimators as compared to existing estimators. We generate a population of size 5000 with the help of standard normal distribution having a mean of 0 and a variance of 1, and we take a different sample size from each, having
n = 50, 100, 200, and 500. To select a sample of size
n, we used simple random sampling without replacement:
The number of simulations is repeated 10,000 times and it is considered to be fixed. The estimators are then evaluated using conventional evaluation metrics such as empirical mean squared error (MSE) and percentage relative efficiency (PRE), with PRE being evaluated in comparison with the standard unbiased estimator of variance:
where i =
,
,
,
,
,
,
,
,
, …,
.
7. Discussion
In the current article, we proposed a novel generalized class of estimators for the estimation of population variance when the sampling is performed through a simple random sample (SRS) with the help of auxiliary information. The proposed estimators are much better than the traditional estimators, which involve more known variables of the auxiliary variable, including the population mean, the coefficient of variation, the skewness, and the kurtosis, into the estimation. The theoretical derivations and analysis of the empirical evidence give very strong arguments in support of the superiority of the proposed estimators over the existing estimators.
Table 1 includes some existing estimators which were generated from our generalized class of estimators by putting some suitable constant, or parameters like skewness, kurtosis, coefficient of variation and correlation among the study and auxiliary variables. Similarly,
Table 2 includes some newly developed estimators obtained from our generalized class of estimators.
Table 3 consists of a summary statistic based on real data sets. The mean squared error of all considered estimators in the article is included in
Table 4. Similarly,
Table 5 includes the result of the percentage relative efficiency of all considered estimators in the article. For further understanding, we also visualized the numerical results with the help of a radar chart and a line chart.
Figure 1 shows MSEs of all considered estimators using Population-I and Population-II.
Figure 2 represents MSEs with the help a line graph, using Population-I and Population-II.
Figure 3 shows PREs with the help of a radar chart using Population-I and Population-II.
Figure 4 shows PREs with the help a line graph using Population-I and Population-II. We also check the efficiency of an estimator, with the help of a simulation study, by using different sample sizes. The numerical results of MSEs and PREs of all considered estimators are given in
Table 6. The numerical result of simulation study is also visualized and shown in
Figure 5 and
Figure 6.
The theoretical findings show the mean squared error (MSE) of the proposed estimators is smaller than the mean squared error of the traditional sample variance and the current auxiliary-based estimators, provided that the relationship that exists between the study variable and the auxiliary variable satisfy reasonable assumptions. In particular, the proposed estimates gain significant efficiency when the study and auxiliary variables have a strong positive correlation. This is similar to the conclusions in classical sampling theory, where as the correlation increases, the more valuable the auxiliary information will be. In addition, the robustness and efficiency of the proposed estimators is also confirmed in the simulation studies carried out in different population structures (normal, skewed, and kurtosis distributions). The generalized estimators perform better in comparison with conventional ones, even in those situations when the relationship between variables does not follow the ideal linearity, which proves that the generalized estimators are more adaptable to the practical sampling issues. This strength is especially needed in practical use, where the population statistics are frequently unknown or nonstandard.
The estimators are also flexible because most existing estimators are special cases. This implies that the proposed framework can be applied in the most diverse cases, when practitioners do not have to change the methodology. This homogenous method of implementation makes this easier and provides a common point for variance estimation regardless of survey design variations.
Another strength of the work is its focus on higher-order auxiliary parameters, including skewness and kurtosis, which are mostly overlooked in the previous studies. The fact that they are incorporated brings in another level of accuracy to the estimation process. In cases where these parameters are known or are estimable with reasonable accuracy by an analysis of previous data, they provide great improvement to the performance of the estimators. It must, however, be noted that the benefit realized is related to the accuracy and the availability of the auxiliary information. Misuse or inaccurate estimation of auxiliary parameters can negatively affect the performance of the proposed estimators.
In practice, the suggested estimators are easy to compute and do not involve intricate optimization algorithms and iterative methods. This makes them particularly appropriate for large-scale surveys or official statistics, where efficiency and ease of implementation are of great concern. The empirical demonstration using real-world data also indicates that the estimators suggested are simple to implement in any standard statistical software operating environment, and they also result in more accurate estimates compared to the traditional methods.
However, despite their benefits, the proposed estimators are not without limitations. The quality of the auxiliary variables and their relevance is critical to their efficiency. Where the auxiliary variable is not significantly or loosely positively correlated with the study variable, the improvement on the traditional estimators is reduced. Moreover, the theoretical benefits of the auxiliary parameters cannot always be completely realized in practice, if the auxiliary parameters are not known or estimated correctly.
Also, although this research has restricted itself to simple random sampling, the methodology has been used to give future researchers the chance to generalize such estimators to more elaborate sampling frameworks, such as stratified, systematic, or cluster sampling. The addition of auxiliary information to such settings would also enhance variance estimation and extend the applicability of the suggested methods.
To conclude, the paper provides a solid theoretical rationale and empirical findings that the proposed generalized estimators represent an important breakthrough in the estimation of population variance in the case of SRS. These estimators have greater accuracy, flexibility, and usability than the current techniques, due to their power to utilize a wider variety of supportive data. Their presentation not only contributes to the statistical toolkit that can be used in the survey sampling but also contributes to more informed decision-making within the data-driven world.
8. Conclusions
In this article, we propose a new generalized class of estimators of the population variance under simple random sampling (SRS) that uses auxiliary information in a more universal way than in existing methods. The proposed estimators show a great enhancement of the efficiency and precision of the conventional estimators by adding known parameters of an auxiliary variable, namely the population mean, coefficient of variation, skewness, and kurtosis. Theoretical discussion supports the fact that the suggested estimators have the desirable statistical characteristics, such as lower bias and mean squared error (MSE), under suitable conditions. Such gains are especially noticeable when the research and auxiliary variables are highly correlated. Moreover, the proposed estimators provide a very general framework in which most of the existing estimators are considered as special cases and thus are generalized into a variety of different situations in surveys. The theoretical results are supported by empirical analyses, including real-world applications. The proposed estimators are observed to be much more effective than the traditional estimators in relative efficiency, particularly in skewed or kurtosis populations where the traditional methods are not very accurate. Practically, the suggested method is easy to adopt and does not entail complicated calculations, thereby making it convenient for large-scale surveys and statistical agencies that are in need of improving the accuracy of their estimates without necessarily having to increase the sample sizes.
Overall, the study makes a significant contribution to sampling theory by providing a more effective and more powerful method for estimating variance in the case of SRS. It points to the unutilized opportunities of auxiliary information for statistical inference, and it supplies a basis for future investigation using more complicated sampling plans or multivariate estimation models.
9. Future Recommendations
Based on the results and the input of this research, we suggest some of the possible directions for future research and practice:
The extension of complex sampling designs: Although the present research is based on simple random sampling, in the future, the research should be extended to determine how the generalized estimators can be adapted and also perform in more complex sampling designs like stratified, cluster, systematic, and multistage sampling. Such designs are very common in practice, and the extension of the methodology would increase their usability extensively.
Addition of several auxiliary variables: The current research takes into account mainly one auxiliary variable. Future research may be aimed at examining the effectiveness of incorporating several auxiliary variables that may be related to each other to better estimate variance.
Model mis-specification: robust estimation and strong approximation: Use of real-world data tends to be in contrast with idealized assumptions. This would be important when coming up with strong forms of the suggested estimators that do not lose efficiency, in the event that the auxiliary data, or the relationships they is assumed to hold, are partially mis-specified or are corrupted by outliers.
Adaptive procedures of estimation: The future of estimator choice might be to have a data-driven or adaptive approach to choosing the best form of the generalized estimator, given the observed sample properties, and therefore, automating the process of estimator choice to practitioners.
Big data applications: non-traditional data sources: As the accessibility of large and complicated data sets based on administrative records, sensor data, and other non-survey information increases, the ability to utilize auxiliary information provided by such sources in variance estimation is an interesting future direction.
Empirical testing in a wide range of areas: In addition to the initial case studies carried out in various fields, agricultural, health survey, economical, and environmental studies will be carried out extensively to test and perfect the proposed methods in diverse practical scenarios. With these guidelines in place, future studies can expand on the basis of this study, increasing the accuracy, strength, and the usefulness of population variance estimation in contemporary statistical research.