1. Introduction
Survey sampling is one of the fundamental areas of statistical inference as a component of the areas associated with sampling of data from a subset of a population for the purposes of collection, analysis, and interpretation. In most real-world applications, it is impractical to conduct a complete survey of a population due to the time, cost, and resources required. Sampling techniques are therefore very common for estimating unknown parameters of a population with a reasonable degree of accuracy and reliability. The population proportion is one of the most important parameters of survey sampling, much used in social sciences, medical research, agriculture, economics, quality control, public health, and opinion surveys. The population proportion is an important piece of information about the occurrence of a specific attribute in a population and is crucial for statistical and policy decision-making.
The sample proportion estimator is the traditional estimator of the population proportion in simple random sampling without replacement (SRSWOR). This estimator is unbiased and easy to compute, but may not be precise enough in some cases, especially if there is further information concerning the study attribute that is not being fully utilized. For this reason, researchers have paid close attention to developing better estimation procedures that use auxiliary information to improve estimator efficiency and reduce sampling error.
Auxiliary information is additional information related to the study variable or the study attribute that is known before the study or can be obtained from outside sources like census records, administrative records, previous surveys, or historical databases. The use of auxiliary information has become one of the most powerful methods in survey sampling, greatly improving the accuracy of estimation without requiring a larger sample size. Traditionally, auxiliary variables and auxiliary attributes have been added in the ratio, product, regression, exponential, and generalized types of estimators to enhance parameters of the estimation in a population. The problem has been addressed in recent years by using multiple auxiliary attributes simultaneously. Dual auxiliary attributes can be added, which results in a more comprehensive recording of variation associated with the study attribute and, thus, more accurate and reliable estimates. Two auxiliary attribute estimators have been shown to achieve better performance than a single auxiliary attribute estimator, both in terms of bias, the size of mean squared error, and the percentage relative efficiency. In the current era of applications, especially in statistical data applications, the development of the efficient estimation techniques using two auxiliary attributes has increased significantly due to the growing number of data sources and auxiliary information.
The authors in [
1] developed an improved generalized class of estimators for the estimation of population proportion using two auxiliary attributes. The main focus of the study was improving estimation efficiency through simulation experiments. In [
2], the authors introduced a general class of estimators for the population proportion based on the auxiliary attributes and applied them to radiation data sets. The study showed that the use of auxiliary information greatly enhances the accuracy of the estimation. Ref. [
3] proposed efficient estimators of population mean based on auxiliary attributes. The study investigated the use of auxiliary qualitative information to improve estimation when simple random sampling is used. Ref. [
4] proposed a class of estimators for population proportion based on auxiliary attributes. The study comprised two parts: simulation analysis and practical application. Ref. [
5] proposed new estimators of the population mean, which are obtained by using the proportion of a population on an auxiliary character. The study revealed that the use of auxiliary qualitative information can improve the efficiency of estimators in simple random sampling.
The authors in Ref. [
6] discussed a predictive approach for estimating the population means using auxiliary attributes. The authors determined that the accuracy of these population proportion estimates using regression and exponential methods is better than that of the existing estimators and that the bias is lower than theirs. Ref. [
7] recommended a new approach for estimating population proportion using auxiliary attribute under non-response stratified random sampling. Ref. [
8] suggested better estimates of the population proportion by auxiliary attributes. The study emphasized that auxiliary information can be used effectively to reduce the sampling error and to make the estimators more reliable. To evaluate the population using an auxiliary attribute, a new estimator is recommended by Ref. [
9]. The authors in Ref. [
10] recommended the population mean using two auxiliary attributes. The authors suggested new estimators to leverage two auxiliary information to improve the estimator accuracy. Ref. [
11] estimated population means with the help of auxiliary attributes when some elements are missing. In Ref. [
12], an improved estimator with two auxiliary attributes was proposed. Their research suggests that employing several auxiliary attributes together can greatly improve the accuracy of estimators. The superiority of the proposed estimator over the traditional estimators was verified theoretically and empirically. In sample surveys, reference [
13] focused on estimating finite-population variances using auxiliary attributes. In the last few years, considerable progress has been made in the area of statistical modeling with an emphasis on flexible methods that can work with complex data structures that arise in reliability and survival studies. Ref. [
14] introduced the hierarchical Bayesian multivariate Wiener process model which can simultaneously consider the dependent degradation rates and volatilities. Likewise, Ref. [
15] proposed a single model for survival data analysis allowing for the inclusion of clustering effects, cure proportions and competing risks. The authors demonstrated via simulation experiments and real-data applications that the proposed approach provided more accurate estimates of the parameters and less bias in parameter estimation, as well as inferential efficiency.
Although significant progress has been made in this area, there are still problems developing efficient population proportion estimators based on auxiliary attributes. Many current estimators have drawbacks, including strong assumptions, complex mathematical structures, limited flexibility or poor performance across different correlation scenarios. Furthermore, some estimators offer efficiency gains only under particular circumstances; others may not remain robust across varying population structures. Thus, there remains a need for generalized and flexible classes of estimators that can be more effectively applied to the dual auxiliary attributes under SRS. These concerns prompted the present study to suggest an improved general class of estimators of a population proportion based on two simple random sampling without replacement auxiliary attributes. The proposed generalized class is designed to merge information from two auxiliary attributes simultaneously to achieve better estimation precision and lower sampling variability. The estimator combines the merits of ratio and generalized estimation methods and provides a more flexible structure suitable for specific circumstances.
Empirical and simulation studies are performed on real and artificial populations to confirm the practical usefulness of the proposed methodology. We expect these studies to show that the proposed generalized class will yield smaller mean square error and larger percentage relative efficiency when compared to some of the existing estimators. These research results can be used as a good contribution in the survey sampling field because they can be used for estimating the population proportion in practical applications, which is reliable and efficient.
With its wide-ranging applications in resource allocation, public health assessment, policymaking, quality monitoring, and socioeconomic planning, the importance of accurate estimation of population proportions has grown in modern statistical analysis. To achieve precise information at minimum survey cost and effort, efficient estimation procedures based on auxiliary attributes can be developed, enabling researchers and policymakers to obtain the information they require. So, using dual auxiliary attributes to develop an improved generalized estimator has significant theoretical and practical value.
This study is novel in the sense that it introduces a novel generalized estimation system for population proportion that leverages two auxiliary attributes in a single flexible system. The class of estimators proposed here is designed differently from existing estimators that are generally optimized for a particular type of auxiliary information and work well only when certain conditions hold; instead, the proposed class has adjustable parameters that enable the estimator to change its behavior with respect to different population characteristics and correlation patterns. This flexibility makes it possible to have a number of well-known estimators as special cases within the proposed framework, as well as to produce a much larger family of new estimators that have not been explored in the literature. Another theoretical contribution is the derivation of the optimum values of the parameters using the minimum mean square error (MSE) criterion which results in a more efficient estimator than the others available. The proposed methodology also employs various auxiliary measures such as mean, median, 1st quartile, and 3rd quartile, which improve the robustness of the methodology for various distributions of the population. The study, therefore, extends the current theory of auxiliary-attribute-based estimation, offers a unified dual-auxiliary estimation framework, lays a foundation for the statistical properties of the proposed estimator, outlines analytical efficiency conditions and illustrates, by means of empirical and simulation examples, the effectiveness of the proposed estimator in terms of its lower MSE and higher PRE as compared with other estimators. Thus, the contribution is not just the aggregation of the existing estimators, but rather the emergence of a new theoretical framework, more applicable, more flexible, and more efficient in population proportion estimation.
1.1. Research Contribution
The present work is an addition to the literature on survey sampling, where it is suggested that a new generalized class of estimators for the population proportion be developed based on two auxiliary attributes for simple random sampling without replacement. The following is a summary of the major contributions of the study:
In the study, a new generalized class of estimators is introduced for estimating the population proportion using two auxiliary attributes simultaneously. The proposed methodology builds on the current estimation framework and offers a more flexible approach to efficiently working with auxiliary information.
The proposed estimator can estimate using two auxiliary attributes, thereby improving estimation accuracy unlike most existing estimators which use a single auxiliary attribute. Two auxiliary information allows for greater variation represented by the study attribute and make the estimation more efficient overall.
Theoretically, the proposed generalized class is designed to yield lower mean squared error than the sample proportion estimator and several estimators in the literature. This will result in accurate and reliable estimates of the proportion of the population.
The study derives mathematical expressions for the bias and mean square error of the proposed estimator using first-order approximation techniques. The derivations also give some good theory for studying the statistical properties and efficiency of the estimator.
The research provides theoretical efficiency conditions under which the proposed estimator is more efficient than the traditional and the existing estimators. The analytical comparison helps to validate and transfer the proposed methodology.
Through numerical and simulation studies using both real and artificially generated populations, the proposed generalized class is verified in terms of its performance. The empirical results validate the advantages of the proposed estimator including higher percentage relative efficiency (PRE) and lower sampling variability.
The study contributes to the existing literature about auxiliary attribute estimation techniques and generalized estimators under simple random sampling. It presents an advanced estimation method that could stimulate future research on the methodological development of survey sampling theory.
The proposed estimator is useful for applications in practical areas such as public health, agriculture, economics, social sciences, environmental studies, and quality control, where accurate estimation of population proportions is imperative for decision-making and policy-making.
The general structure of the proposed estimator lends opportunities for further extensions to sampling designs, multiple auxiliary attributes, non-response cases, and a robust estimation framework.
In practical survey applications, the proposed methodology would enable more accurate and efficient estimates of a population’s proportion, thereby enabling improved statistical inference, planning, and evidence-based decision-making.
1.2. Objectives of the Study
The main aim of this study is to find an enhanced generalized class of estimators for population proportion based on two auxiliary attributes in simple random sampling without replacement (SRSWOR). Specific objectives of the study are given below:
To introduce a new general class of estimators to estimate the population proportion based on two simple auxiliary attributes for simple random sampling.
To include two auxiliary attributes at the same time to enhance the accuracy and efficiency of population proportion estimation.
To obtain the mathematical expressions for the bias and mean square error (MSE) of the proposed generalized class by the first-order approximation.
To examine theoretical properties and statistical behavior of the suggested estimator in various correlation configurations.
To compare the efficiency of the proposed new generalized class with the conventional sample proportion estimator and other estimators found in the survey sampling literature.
To formulate theoretical conditions on which the proposed estimator is better than the other estimators in terms of minimum mean square error.
To assess, from empirical and simulation studies of real and artificially generated data sets, the practical applicability of the proposed estimator.
To calculate the percentage relative efficiency (PRE) of the proposed estimator and its superiority over the traditional estimation procedures.
To illustrate the need and utility of using dual auxiliary attributes to reduce the sampling variability and improve the estimation accuracy.
To offer a dependable and adaptable estimation method that may be used in a variety of practical applications, including public health, agriculture, economics, social sciences, and environmental studies.
To advance survey sampling theory with a new generalized estimation framework for the population proportion in simple random sampling.
To establish a basis for future studies with multiple auxiliary attributes, non-response situations with generalized estimation techniques under other sampling schemes.
The rest of this study is organized as follows. The notations and methodology are given in
Section 2. The relevant literature is presented in
Section 3, which reviews estimation techniques based on auxiliary attributes. A general class of estimators is then proposed, and its theoretical properties are developed in
Section 4. Comparisons of efficiency and analytical results are then discussed, and numerical and simulation studies are conducted to assess the performance of the estimator in
Section 5 and
Section 6. The discussion of the study is given in
Section 7. Assumptions for optimal performance are given in
Section 8 and conditions under which the proposed estimator perform better are given in
Section 9. Lastly,
Section 10 and
Section 11 include findings and implications, which are summarized in the conclusion and directions for future research.
2. Materials and Methods
Suppose a population Ψ =
contains
N independent and identical elements. We choose a sample size of
n with the help of simple random sampling without replacement. Suppose
and
,
and
be the characteristics of the study attributes
and auxiliary attributes
and
. Let us consider
ij (
j =
,
x,
z), if
units are selected and zero otherwise. Consider
=
and
=
represents statistic and population parameter values, where
=
and
=
as (
j =
,
x,
z). Let us define the following error terms:
= , = , = , denoted correlation among the study and auxiliary attributes.
= , = , = , denoted covariance among the study and auxiliary attributes.
where
= , = , = , denoted standard deviation among the study and auxiliary attributes.
=
,
j =
y,
x,
z. denoted coefficient of variations for the study and auxiliary attributes.
3. Existing Estimators
This section includes some adopted existing estimators for the estimation of population proportion along with properties, which are given by:
- (i)
The usual estimator for population proportion is given by:
The variance of
is given by:
- (ii)
The adopted ratio estimator for population proportion developed by [
16] is given by:
The bias and MSE of
is given by:
and
- (iii)
The author in [
17] suggested that a product estimator for population proportion is given by:
And
- (iv)
The regression estimator for population proportion is given by:
where
is a constant and its value is given by:
where
- (v)
Ref. [
18] suggested the following estimator, which is given by:
and
where
Using
,
we obtained:
- (vi)
The authors in [
19] recommended the following exponential type estimators, given by:
and
- (vii)
Ref. [
20] suggested the following estimator, which is given by:
The estimator
reduces to
and
when
and
, respectively.
and
where
- (viii)
The author in [
21] suggested the following improved exponential estimator is given by:
and
4. Suggested Generalized Estimator
Population proportion estimation is a key goal of survey sampling, as it is a critical component of statistical analysis, policy formulation, planning, and decision-making. Population proportions are used in many fields of science, such as in public health, economics, agriculture, social sciences, engineering, and quality control, to determine the prevalence or occurrence of an attribute within a population. They can be used to predict the proportion of literate people, the number of people with a specific disease, the number of people who do not work, the number of people who prefer a particular candidate or product, or the number of people living below the poverty line. Hence, the problem of efficiently estimating a population proportion has continued to attract interest in survey sampling theory.
The existing estimation procedures for population proportions are based on a single auxiliary attribute or auxiliary variable. These estimators are improvements over the traditional estimator; but if there is more than one auxiliary attribute related to the study attribute, these estimators cannot fully utilize available information to improve. In many practical situations the correlation between the auxiliary attribute and the characteristic under study is very high, and there may be two or more of them. Without considering other auxiliary information, the precision and efficiency of the estimator can consequently be sacrificed.
In recent years, as the number of data sources in statistical applications has increased, researchers have been interested in studying dual auxiliary-attribute-based estimation. Two auxiliary attributes can be used together to capture more of the variation in the study attribute, thereby increasing the accuracy of the estimates and decreasing the mean square error. Other estimators can be efficiently applied only under certain correlation structures and might not yield satisfactory results in more general practical applications. Therefore, it is always necessary to develop better generalized estimators that can be used efficiently with the dual auxiliary attributes under simple random sampling. A second driving force for undertaking this study is the need for highly accurate statistical estimates in data-driven research settings. Survey results are increasingly becoming a key asset for effective planning, monitoring, and evaluation on which governments, policy-makers, researchers and organizations depend. Any efficiency gains by the estimator can result in significant cost, time, and effort savings on a survey without compromising on quality and reliability.
In view of these problems, the current research is conceived to construct a better class of estimators for the population proportion with two auxiliary attributes in simple random sampling. The proposed methodology aims to improve accuracy, minimize mean squared error, and offer a more flexible and reliable alternative to current estimators. It also aims to enhance the theoretical and practical foundations of survey sampling by proposing a generalized framework for estimation that can be effectively applied in various practical situations. By taking inspiration from
,
,
,
and
, we suggested the following proposed generalized class of the estimator, which is given by:
Equation (31) can further be expressed as:
Table 1 contains some family members obtained from the suggested class of estimators by putting different parameter values.
Subtract
from both side of (32), we obtain Equation (33):
Apply expectation to (33), we obtain bias of
, which is given by:
and
Differentiate (35) w.r.t
,
and
and equate to zeror, we obtain the following expression:
and
where
5. Numerical Analysis
We take a numerical analysis to compare the existing and the suggested classes of estimators. Six actual data sets are used for this purpose.
Table 2,
Table 3,
Table 4,
Table 5,
Table 6 and
Table 7 present aggregate statistics for the provided data. PRE of an estimator
concerning
is given by:
where
i =
,
,
, …,
(
j =
,
, …,
).
The thresholds used in Situations I–IV are based on four commonly used measures of location: mean, median, first quartile (Q1), and third quartile (Q3). They were chosen a priori because they combine various features of the distribution of the population and offer auxiliary information. The median is the middle value of the data and is very informative if the distribution is somewhat symmetric. The median is the 50th percentile and is a good measure of central tendency that is not affected by extreme values. The first quartile (Q1, 25th percentile) indicates the bottom 25% of the data, while the third quartile (Q3, 75th percentile) indicates the top 25%. The combination of Q1 and Q3 gives information about the spread and shape of the distribution and is less influenced by outliers.
The range of these thresholds enables the proposed methodology to test the performance of the estimators under various distributional conditions. If the population is symmetric, thresholds based on the mean might be very effective, but if the population is skewed, has heavy tails, or extreme observations, thresholds based on the median and quartiles may be more robust. The choice of these four thresholds was not based on an optimization process based on the data but instead based on statistical theory and their common use in descriptive and inferential analysis. The thresholds were a priori defined without bias and before the analysis, so they do not cause over fitting and allow unbiased and balanced evaluation of the suggested estimators in different population structures.
Population-I: [Source: [
22]]:
Y = Number of teachers in Turkey in 2007
X = Number of students in primary level in Turkey 2007
Z = Number of students in secondary level in Turkey 2007
Situation-I
= Proportion i1 < 1 for (Y ≤ 436.4) and i1 > 1 for (Y > 436.4)
= Proportion i2 < 1 for (X ≤ 11,440) and i2 > 1 for (X > 11,440)
= Proportion i2 < 1 for (X ≤ 333.2) and i2 > 1 for (X > 333.2)
N = 923; n = 180; = 0.004472132, = 0.2296858, = 0.2307692, = 0.979415, = 0.8930209, = 0.07916363, = 0.07940598, = 1.832324, = 1.826732, = 0.1450534
Situation-II
= Proportion i1 < 1 for (Y ≤ 171) and i1 > 1 for (Y > 171)
= Proportion i2 < 1 for (X ≤ 4123) and i2 > 1 for (X > 4123)
= Proportion i2 < 1 for (X ≤ 180) and i2 > 1 for (X > 180)
N = 923; n = 180, = 0.004472132, = 0.4983749, = 0.4994583, = 0.9956663, = 0.8461553, = 0.06575983, = 0.06590247, = 1.003799, = 1.001627, = 0.06600968,
Situation-III
= Proportion i1 < 1 for (Y ≤ 84) and i1 > 1 for (Y > 84)
= Proportion i2 < 1 for (X ≤ 1734) and i2 > 1 for (X > 1734)
= Proportion i2 < 1 for (X ≤ 93) and i2 > 1 for (X > 93)
N = 923; n = 180, = 0.004472132, = 0.744312, = 0.7497291, = 0.9989166, = 0.8711024, = 0.05618972, = 0.05700088, = 0.5864257, = 0.5780805, = 0.0329511.
Situation IV
= Proportion i1 < 1 for (Y ≤ 394.5) and i1 > 1 for (Y > 394.5)
= Proportion i2 < 1 for (X ≤ 9975) and i2 > 1 for (X > 9975)
= Proportion i2 < 1 for (X ≤ 367) and i2 > 1 for (X > 367)
N = 923; n = 180; = 0.004472132, = 0.2502709, = 0.2502709, = 0.971831, = 0.8902923, = 0.09836563, = 0.09836563, = 1.731739, = 1.731739, = 0.1703436.
6. Simulation Study
Suppose a population N = 1000 and the sample size is n = 100. Assume that Y is the study attribute and X and Z are auxiliary attributes. The following steps are taken in the simulation study.
Step 1: Specify population parameter
The average vector is considered to be:
and the covariance matrix is defined as:
which is a tri-variate normal population with high positive correlation between the variables.
Step 2: Generate Population
The mean vector μ and covariance matrix Σ are used to generate a tri-variate normal population with
N = 1000 observations.
Use a simple random sample to select a subsample from the data set. Also convert variables to attributes, as discussed in the methodology section.
Compute the sample proportions , , ; and the coefficients of variation , , and as well as the correlation coefficients , , and .
Step 3: Review and improve the existing estimators. Using the mathematical expressions of the existing estimators, calculate the Mean Squared Errors (MSEs).
Step 4: Calculate the percentage relative efficiency.
Step 5: Compare the MSEs and PRE of each estimator for comparison to the usual estimator.
7. Discussion
The mean, median, first quartile (Q1), and third quartile (Q3) were parameters of the study and auxiliary variables that summarize different aspects and can be useful in enhancing estimation efficiency in different situations. When the auxiliary variable is approximately symmetric and does not have any extreme observations, the mean is the most widely used measure of central tendency and is very informative. But it can be affected by the presence of outliers, which can make it less useful in case the data is skewed or contains outliers. To overcome this, the median was used as a good measure of the central tendency. The median is not as greatly influenced by extreme observations and therefore can give more stable and reliable auxiliary information for skewed populations. Likewise, Q1 and Q3 were represented since they encompass the bottom and top part of the distribution, respectively. Quartile-based measures are especially suitable if the auxiliary variable is asymmetric, has heavy tails, or does not have homogenous dispersion, because they are not so much affected by outliers and allow more information about the distributional structure of the population than can be gleaned from the mode.
These measures considered together increase the flexibility of the suggested estimator due to the fact that the estimator can use various distributional properties of the auxiliary variable. In the real world, the mean is usually good for almost symmetric populations, while the median and quartiles may have more benefits for skewed or non-normal populations. Thus, instead of using just one location measure, the proposed framework uses several location measures to maximize efficiency of estimation over a larger set of population structures.
Table 1 contains some family members obtained from the suggested class of estimators by putting different parameter values. The MSEs and PREs of the proposed generalized estimators and the existing estimators for real data sets are presented in
Table 2,
Table 3,
Table 4,
Table 5,
Table 6,
Table 7,
Table 8 and
Table 9 for different choices of auxiliary measures. In all cases, the estimator with the minimum MSE and the maximum PRE is said to be the most efficient estimator. The MSE values for the presence of auxiliary information based on the population means are presented in
Table 2. The proposed generalized proportional class is significantly better than the classes of
and
estimators, as well as all other estimators of the traditional classes. Among all estimators
produces the smallest MSE indicating the highest precision. The MSEs of the other estimators are also high; the traditional estimator
is 0.00079211360, and the other estimators, namely the regression estimator and the ratio-difference estimator, are even larger. The
yield moderate improvement over the traditional estimator while
gives good efficiency gains. The proposed class is, however, easily generalized and the resulting class is clearly dominant over all other classes considered in this paper.
The PRE values corresponding to the values in
Table 2 are reported in
Table 3. The results are confirmation of superiority of the proposed estimators. The estimator
has the highest PRE value of 61861.97 which is very high compared to PRE values of the existing estimators. The PREs of the
are approximately 495 while the Singh estimators are approximately 100–147. The traditional estimator is by definition PRE = 100. The results show that the generalized proportional estimators proposed are very efficient compared to classical estimators. The MSEs for the auxiliary measures based upon medians are given in
Table 4. Again, the proposed generalized proportional estimators are better than the other estimators. The minimum MSE is provided by
, which is much smaller than the other MSEs of the traditional, ratio, product, regression, and difference-type estimators. The Grover class also gives results of efficiency near MSE 0.000317 and the
gives moderate reduction in MSE than the traditional estimator. Overall, the proposed class is still superior with regard to median auxiliary information. The PREs for the median-based estimators are given in
Table 5. The same can be concluded for the estimator
which has the highest PRE value of 46008.97 and is the most efficient one of all the existing estimators. The PRE values of Grover class are ~353 while Singh class ranges from ~129 to 139. The efficiencies of the regression and the ratio-difference estimators are found to be reasonably better than traditional estimator; however, still not as efficient as the proposed generalized proportional class.
Table 6 shows the MSEs for the case of the first quartiles being used as auxiliary measures
.
Once again, the proposed generalized class of estimators among all considered estimators has the lowest MSEs. Specifically,
gives the minimum MSE. The
also exhibits some different efficiencies depending on the type of transformation used, as do the
when the MSE is close to 0.000205. Traditional estimators and product estimators show relatively higher MSE value, which is less precise. The PREs for the estimators based on the first quartiles in the table are the following:
Table 7 includes the PRE of the estimator
is the highest value, 52781.82, indicating that the estimator is very efficient. The PREs of the Grover estimators are more than 414 and the Singh estimators are from around 100 to 203 PREs. The proposed generalized proportional class always achieves the highest efficiencies, hence the effectiveness of incorporating quartile-based auxiliary information is validated.
The MSEs are obtained using the third quartiles as auxiliary measures
and are reported in
Table 8. As before, the generalized proportional class is shown to be dominant. The estimator
achieves the minimum MSE, outperforming all existing estimators. It is noted that MSEs of the Grover estimators are comparatively smaller than the MSEs of the Singh estimators and the MSEs of traditional and product estimators are the least efficient. These results suggest that the estimation accuracy is further enhanced with the use of third quartile information. Finally, the PREs for third quartile-based estimators are given in
Table 9. In this case, the optimum PRE value is 60546.34, achieved by
, and this is clear evidence that this is the most efficient estimator in this situation. The Grover class PREs are around 484 and these of Singh estimators range from 100 to 146. The proposed generalized proportional estimators give remarkably high efficiencies as compared to the available estimators like ratio, regression, and difference estimators. In general, the empirical (
Table 2,
Table 3,
Table 4,
Table 5,
Table 6,
Table 7,
Table 8 and
Table 9) results show that the proposed class of generalized proposed estimators are superior to all other existing estimators based on the MSE and PRE for all the measures of auxiliary data considered, such as means, medians, first quartiles and third quartiles. This verifies the high robustness, stability, and efficiency of the proposed estimation approach for estimating the population proportion. The results in
Table 10,
Table 11,
Table 12,
Table 13,
Table 14,
Table 15,
Table 16 and
Table 17 are for the simulated population using various auxiliary measures. Similarly to the real data sets, the numerical results from the simulation study verify that the suggested generalized class of estimators performs better in terms of minimum MSE and higher PREs across all population parameters.
8. Assumption for Optimal Performance
The following are assumptions and conditions that must be met for optimal performance:
Simple random sampling without replacement (SRSWOR) is used for selecting the sample from the population.
The number of individuals N in the population is known or can be reliably estimated, and the proportion of individuals with a particular type of auxiliary attribute can be known.
Study attribute and auxiliary attributes have a positive correlation.
The number of samples is adequate for the first order Taylor series approximation to be accurate.
Sampling errors are small and have finite moments of order required.
There is no measurement error for the auxiliary attributes, and they are available for all population units.
The population proportions are not too near to 0 or 1, which is good for the validity of the linearization approach.
9. Conditions Under Which the Proposed Estimator Perform Optimally
The conditions under which the proposed Estimator performs optimally.
There is a strong positive correlation between the study attribute and the auxiliary attributes.
The auxiliary attributes are large and provide a considerable amount of information about the study attributes.
Sample size is moderate to large, minimizing the approximation errors and sampling variability.
The auxiliary proportions are accurately known and representative of the population.
There are no significant parts of the population that are very high or very low.
The optimum values of the estimator weights are used according to the minimum Mean Squared Error (MSE) criterion.
Various auxiliary summary measures, including the mean, median, the first quartile (Q1), and the third quartile (Q3), provide useful information about the distributional structure of the auxiliary variables.
In that case, the proposed estimator is more efficient than the considered estimators in the study, and has the lowest mean square error, thus it allows obtaining a more accurate estimation of the population proportion.
10. Conclusions
This Study recommends a generalized class of estimators for the population proportion based on two auxiliary attributes, under simple random sampling without replacement. The proposed methodology was designed to improve the accuracy of population proportion estimation by effectively leveraging auxiliary attribute information in the estimation process. The estimator has a generalized form, which provides greater flexibility and enables the use of available auxiliary information to reduce the sampling variability. Expressions for bias and mean-squared error (MSE) were derived and used to evaluate the estimator’s statistical efficiency. Analytical comparisons showed that under appropriate conditions, the proposed estimator is better than the sample proportion estimator and several estimators currently found in the survey sampling literature. The efficiency improvement was mainly achieved due to the simultaneous use of two highly associated auxiliary attributes.
The empirical results showed that the proposed generalized class yielded lower mean square error and higher percentage relative efficiency (PRE) than the traditional estimators. These results confirm that dual auxiliary attribute use significantly enhances the accuracy and reliability of population proportion estimation. The study emphasizes the role of auxiliary attribute information in modern survey sampling and shows that combining two auxiliary attributes can significantly improve estimation efficiency. The proposed generalized class offers a reliable, efficient, and flexible estimation approach which can be used in a wide range of applied applications, including public health, agriculture, economics, environmental sciences, social sciences, and quality control studies where the population proportion is needed.
In general, the proposed methodology advances the theory of survey sampling by providing an enhanced generalized estimation process for population proportion when sampling is simple random. The results of this study could also serve as a basis for further research on the use of multiple auxiliary attributes, generalized estimation methods, and more complex sampling designs.
11. Some Future Direction of the Current Study
The suggested generalized class of estimators could be extended to other sampling methods, such as systematic sampling, cluster sampling, stratified sampling, and probability proportional to size (PPS) sampling.
Future studies could develop an estimation procedure based on multiple auxiliary attributes to further improve the precision of population proportion estimation.
The methodology can be extended to the cases of non-response, missing observations, and missing survey data which are prevalent in real-life sampling applications.
The proposed generalized class can be adjusted to estimate other population parameters, such as the population mean, variance, median and distribution function, based on auxiliary attributes.
Future research could consider combining both variables and attributes as auxiliary variables to build more efficient auxiliary estimation methods.
To address the uncertainty and imprecision of the survey data, the performance of the estimator is evaluated in fuzzy, neutrosophic and interval-valued statistical environments.
Dual auxiliary attributes based on generalized exponential, logarithmic and calibration-based estimators may be studied in the future to improve the efficiency of estimation.
The suggested estimator can be generalized to multistage and adaptive sampling designs where auxiliary information is important for improving survey accuracy.
Higher-order approximations could then be used in comparative studies to yield more accurate theoretical efficiency estimates for the proposed estimator.
The researchers can explore the potential of these proposed approaches in real-world contexts like healthcare surveys, environmental measurements, agricultural research, quality control, and socioeconomic studies.
The proposed general structure in this study could be a starting point to develop resilient and efficient estimation methods for large-scale surveys and high-dimensional data environments.