Next Article in Journal
GO-PILL: A Geometry-Aware OCR Pipeline for Reliable Recognition of Debossed and Curved Pill Imprints
Previous Article in Journal
Qualitative Analysis for Modifying an Unstable Time-Fractional Nonlinear Schrödinger Equation: Bifurcation, Quasi-Periodic, Chaotic Behavior, and Exact Solutions
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

The Efficiency of Missing at Random Planned Missing Designs

by
David G. Steel
1,* and
James Chipperfield
2
1
School of Mathematics and Applied Statistics, University of Wollongong, Northfields Ave, Gwynneville, NSW 2522, Australia
2
Methodology and Data Science Division, Australian Bureau of Statistics, Belconnen, ACT 2616, Australia
*
Author to whom correspondence should be addressed.
Mathematics 2026, 14(2), 355; https://doi.org/10.3390/math14020355
Submission received: 1 December 2025 / Revised: 9 January 2026 / Accepted: 16 January 2026 / Published: 21 January 2026

Abstract

Planned Missing Designs (PMDs) allow for different sets or patterns of variables to be collected from sample units. While the typical motivation for PMDs is to manage respondent burden, they can also reduce data collection costs and provide flexibility in meeting reliability requirements for key survey outputs. Almost all PMD applications involve data being Missing Completely At Random (MCAR). That is, the pattern of variables to be collected from a sample unit is determined prior to collecting any variables from the unit. Here we generalise this approach by considering designing PMDs that allow data to be Missing At Random (MAR). That is, the set of variables collected from a unit is allowed to depend upon the value of some of the variables collected from the unit. At the design stage no data have been observed and so we consider the expected information function associated with maximum likelihood estimation for any specified PMD. We show how the missing information principle can be used to determine the loss of information arising from the use of a PMD. This paper considers the multinomial distribution in detail and conducts an empirical evaluation to illustrate potential efficiency gains associated with MCAR and MAR PMDs.

1. Introduction

1.1. Planned Missing Designs

Consider a survey which collects information on P variables, from a sample, s, of size n. For i s the variables are x i = ( x 1 i , x 2 i , , x P i ) . The sample data collected under the standard or single-phase design (SPD) is the complete sample data x = { x i ; i s } . The data is complete in the sense that there are no missing values. Suppose some variables are missing for some sample units. Let I i p = 1 indicate that the variable x p is observed for unit i and I i p = 0 if x p is missing for unit i. I i = ( I i 1 , , I i p , I i P ) , indicates which variables are observed for sample unit i and the matrix I = ( I 1 , I 2 , , I n ) summarises the missingness for the sample. The observed values are x o b s and the observed data includes I and so it is d o = { x o b s , i ,   I i ;   i s } . The complete sample data is then x s = [ x o b s , x m i s ] where x m i s denote the data values that are missing for sample units.
Here we will call a design that plans to only collect x o b s according to some process specified at the design stage a Planned Missing Design (PMD). A PMD is defined by the mechanism that gives rise to d o given s and can also be called a Split Questionnaire Design (SQD).
There are two types of missing data mechanisms that can be planned. They are missing completely at random (MCAR) and missing at random (MAR). A MCAR mechanism assumes that the joint density of I and x o b s is p ( I , x o b s ) = p ( I ) p ( x o b s ) , which means the pattern of missing data is independent of the observed data. A MCAR PMD is defined by p ( I ) . A MAR PMD allows the pattern of missing data to depend upon the observed data in some way, so p ( I , x o b s ) = p ( I | x o b s ) p ( x o b s ) . A MAR PMD is more complicated than an MCAR PMD as specifying a distribution for p ( I | x o b s ) is more involved than specifying a distribution for p ( I ) . The MCAR design is a special case of the MAR design.
An MAR PMD depends on x o b s . At the design stage when the PMD is being developed, no data has been observed. Thus, we need to consider the properties of estimates with respect to the joint density p ( I , x o b s ) . For maximum likelihood estimates (MLEs) this leads to considering the expected information function. In Section 2 we show how a PMD affects the observed and expected information in general and how the loss of information can be determined. Section 3 considers the common case of the multinomial distribution.
In some applications there may be variables available at the design stage, some of which may be used in the sample design. These variables can be used in a PMD and this is considered in Section 2.4.
Some data items may be missing due to non-response, but unlike PMDs, the missingness is not designed and which data items are missing may depend on the variables in an unknown way, leading to biased estimates. In this case the data are missing not at random (MNAR) and the non-response is informative. In practice it is often assumed that survey non-response is MAR with I depending on some fully observed data items. However, because which data items are missing is not controlled by the survey organisation, it is possible that I and x o b s are not independent. See [1] (Chapter 7) and [2] (Chapter 8) for a discussion of informative non-response. For this paper we will assume there is no data item non-response and focus on missingness due to the PMD. The approach developed here can also cover data item non-response that is MCAR or MAR.
An MCAR PMD is by far the most common form of PMD in the literature. There are a total of T = 2 P 1 patterns or possible values of I i . For example, Table 1 illustrates all possible patterns for P = 3 indexed by j = 1 , , 7 . In an MCAR PMD, sample units would be allocated to a pattern at random. The development of a PMD would involve specifying what proportion or number of units would be allocated to each pattern. An example of an MAR PMD would be one in which x 1 is observed for all units and the probability of which of x 2 and/or x 3 are observed depends on the value of x 1 . This means only patterns j = 1 , 3 , 6 , 7 are used, but the probability of which one is used for unit i depends on its value of x 1 . A multi-phase design (MPD) is a special case of a PMD in which the patterns used follow a monotone pattern. For example, patterns j = 1 , 3 , 7 . A standard or SPD survey would collect all variables from all sample units (i.e., j = 7 in Table 1 in which I i = ( 1 , 1 , 1 ) ).
MAR PMDs are more flexible and can be more efficient than MCAR PMDs as they allow the missingness pattern to depend on the cost and loss of information associated with different categories of key variables. Their potential benefit will be greater when some categories lead to higher costs or are relatively rare. For example, in a health survey the response to a question on general self-assessed health status of poor or fair may lead to more detailed data on health conditions being collected, increasing costs. Thus, having a higher rate of missingness for the additional data items for people of poor or fair health can decrease costs while still allowing the more detailed data items to be collected within the available budget. In a cancer survey those that report never having being diagnosed can have a lower probability of collecting other lifestyle and dietary data. In both situations the data collected can be fully exploited in maximum likelihood estimation.

1.2. Main Issues in the Literature

The typical motivation for PMDs is to manage respondent burden, which allows fewer than P data items to be collected from a sample unit. Next we review the main issues in a PMD and how they are addressed in the literature.
The first issue is to define the target statistic or objective of the survey. PMDs have been used to estimate a wide range of statistics. Examples include population totals ([3,4]), the power of a hypothesis test ([5]), the variance of analytic parameters (e.g., ref. [6] who use A- and D-optimality, i.e., the trace and determinant of the variance matrix of the parameter estimates), or the likelihood itself (see [7]).
Second, the number of possible patterns, T, becomes large even for a small number of variables. There are several reasons for limiting the set of allowable patterns. The target parameter itself may impose a natural constraint on allowable patterns. For example, when regression coefficients are the target, then the dependent variable should always be collected (e.g., [8,9]). To equalize the reporting burden across respondents, another common constraint is that each allowable pattern has the same number of variables. Moreover, if the interaction between x 1 and x 2 is important, then we may ensure that they are always collected together. Restricting the number of allowable patterns may reduce the complexity of administering a survey using a PMD and of analysing the data once it is collected.
A common way of restricting the number of allowable patterns is to create modules, whereby a variable appears in only one module and all variables in a module are collected together. The creation of modules has been widely studied. Ref. [10] recommends that a variable in a module should be well correlated with variables in different modules, assuming only one module is collected per sample unit. Ref. [11] forms modules explicitly for the purpose of minimising imputation error. Ref. [12] considers simple algorithms to allocate variables to a module so as to create diverse topics or a single topic in a module.
Third, once the set of permissible patterns or values of I have been determined, each must be assigned a probability of occurring in the sample. Finding the optimal p ( I ) involves optimisation algorithms that typically aim to maximise or minimise the statistical objective. Refs. [3,4] use a simple grid search algorithm since P is small. Ref. [13] searches for the optimal PMD patterns using Monte Carlo simulation to estimate the standard error of estimates in growth curve models. Ref. [6] minimises the A-optimality and Ref. [14] minimises the Kullback–Leibler Distance when the number of cells in the multinomial distribution is large.
There are limitations of a PMD. A PMD collects less information about interactions than a standard design. This may be acceptable when key interactions are identified at the design stage but may be unacceptable if a survey is to support analysis of a wide range of interactions. Different routings through a questionnaire can sometimes affect response values. This is particularly the case for sensitive questions (for discussion see [15] chapter 5). A PMD tends to complicate data collection and restrict modes of data collection. Analysis of missing data is typically more complex than analysis of complete data [16]. A range of methods are available when there are planned or unplanned missing data. These include full information maximum likelihood (FIML), the E-M algorithm, multiple imputation and Bayesian methods. Refs. [2,17] describe modern methods in the analysis of missing data.
The range of problems to which PMDs have been applied has increased over time, supported by estimation methods and technology that make their implementation easier. We give a few interesting examples. An experience survey ([18]) involves a person responding, via a mobile phone, at random points throughout the day, to a pattern of questions about the amount of time spent on a variety of tasks. Ref. [19] uses a PMD to estimate smoking behaviour, a latent variable measured by a question–answer, a saliva test and a breath test. PMDs are applied in ecology and animal welfare studies ([20]) where the data are measurements from the environment. Ref. [21] considers a PMD in a panel survey under various longitudinal correlations structures. Ref. [22] considers PMDs within the context of an official statistics agency integrating multiple surveys over time. Ref. [23] reviews PMDs for epidemiological research and [24] considers PMDs in educational psychology research.
Here we consider a PMD from the perspective of efficient sample design. In general there will be a vector of population parameters of interest and we focus on the efficiency of a design based on A- and D-optimality, for fixed total survey cost. PMDs have two efficiency-based advantages over standard surveys. Firstly, they allow variables with high reporting and collection cost or low importance to be collected from fewer units than variables with relatively low cost and high importance. Secondly, the correlation between variables can be exploited to minimise the information loss due to missing data.
Relatively little work has considered PMDs where the data are MAR. An MAR design has more control over I than an MCAR design. In an MAR design, I can depend upon x o b s , while in an MCAR design it cannot. This is important since d o determines the variance of parameter estimates in a Maximum Likelihood (ML) framework. Ref. [21] briefly mentions an MAR PMD. MAR PMDs involving screening questions or missingness based on a dependent variable have been considered in the literature. Ref. [25] allows I to depend upon screening questions. Ref. [8] considers allowing the pattern of missing data to depend upon the value of the dependent variable in a logistic regression model. A feature of MAR PMDs is there flexibility and adaptivity as decisions on which variables to collect depend on the observed values some key variables. This also improves robustness as it can be used to ensure particular combinations of variables are collected from an adequate number of units. For some distributions of x , an MAR and MCAR PMD will be equally efficient. This is the case if p ( x m i s | x o b s ) is homoskedastic, as is the case for the multivariate normal, as it implies V a r ( x m i s | x o b s ) = V a r ( x m i s | I ) and the pattern (not the observed variables) affects the uncertainty in the missing data. We consider the multinomial distribution in detail, which is heteroskedastic in p ( x m i s | x o b s ) .
In this paper we develop a general and systematic information-based framework for the design of MCAR and MAR PMDs assuming ML estimation is to be used once the data have been collected. Our approach focusses on the information loss and how this can be determined at the design stage. For many surveys the key outputs are tables of frequencies or proportions, and so the framework is applied to the reasonably general case of a P-way table using the multinomial distribution.
Section 2 considers the general effect of a PMD on information for MLEs. Section 3 derives the information and loss of information of maximum likelihood estimates under the multinomial distribution for an MCAR and MAR PMD. This is done for the observed information, which is relevant when the data have been collected, and expected information given the PMD, which is relevant at the design stage. A sequential and efficient MAR PMD is described. Section 4 considers the development of MCAR and MAR designs based on A- and D- optimality, and provides an empirical evaluation. It also considers ways of identifying the marginal effect of observing a particular data pattern for a single unit. Section 5 makes concluding remarks. Appendix A provides a summary of the notation used in the paper.

2. Effect of PMDs on Information for MLEs

The use of a PMD will affect the variance of parameter estimates. There are several approaches available for estimation once data have been obtained using a PMD. In this section we consider cases where estimation is based on the widely used maximum likelihood approach. The variances of the MLEs are given by the inverse of the information matrix and so we consider the effect of a PMD on the information matrix and the associated information loss. At the design stage, before data have been collected, means different possible PMDs need to be evaluated and an efficient design identified. It is necessary to have clear criteria on which to base these evaluations. The variances of parameter estimates associated with MLEs obtained from the information function, and functions of them, provide relevant criteria at the design stage.

2.1. Complete Sample Data

To help focus on the development of a PMD, we will assume a simple sample design in which no auxiliary variables are used in the selection of the sample of units and it is noninformative, so that p ( s | x U ) = p ( s ) , where x U is the matrix of population values. Many simple designs such as simple random sampling are noninformative. More complex noninformative sample designs that depend on auxiliary variables available for the population are briefly considered in Section 2.4.
We assume that the density of the variables of interest depends on a vector parameter θ . Given the complete sample data, the density is treated as a function of θ , which is the likelihood. Let δ U be the sample membership indicator vector. For a noninformative sample design, the joint density p ( x , δ U ) = p ( x ) p ( δ U ) and the sampling can be ignored and the sample taken as fixed for maximum likelihood estimation of θ . Standard maximum likelihood estimation would be obtained using the observed score function, S c ( θ ; x s ) , which is the derivative of the log-likelihood with respect to θ . Under standard regularity conditions, setting the score function equal to zero gives the MLE, θ ^ . The observed information function, i n f o ( θ ; x s ) , is minus the derivative of the score function. The expected information function is I n f o ( θ ) = E[ i n f o ( θ ; x s ) ]—this expectation takes δ U as fixed. The variance of the asymptotic distribution of the MLE is V ( θ ^ ) = I n f o ( θ ) 1 . Given the complete sample data, the variance can be estimated by i n f o ( θ ^ ; x s ) 1 or I n f o ( θ ^ ) 1 . The former has advantages for conditional analysis but the latter is often simpler and more stable. See [1] for a comprehensive treatment of maximum likelihood estimation using sample survey data.

2.2. Information with Missing Data

For a PMD the observed data are d o = { x o b s , I } . In general the density p ( d o ; ϕ ) will depend on the parameter ϕ that characterises the distribution of d o leading to the score function S c ( ϕ ; d o ) , and observed and expected information functions i n f o ( ϕ ; d o ) and I n f o M ( ϕ ) = E[ i n f o ( ϕ ; d o ) ]. The subscript M is used to distinguish I n f o M ( ϕ ) from the expected information based on the complete data, I n f o ( ϕ ) = E[ i n f o ( ϕ ; d ) ], where d = { x s , I } .
The MLE can be obtained directly from p ( d o ; ϕ ) . In many applications the score and information function based on the complete data are well known and easily handled. The missing information principle (MIP) (see [1,26]) can be used to derive the score function by replacing complete data statistics by their expectation given the observed data:
S c ( ϕ ; d o ) = E [ S c ( ϕ ; d ) | d o ] ,
For the observed information function, the MIP gives a key result that highlights the loss of information:
i n f o ( ϕ ; d o ) = E [ i n f o ( ϕ ; d ) | d o ] V a r [ S c ( ϕ ; d ) | d o ]
The expected information is
I n f o M ( ϕ ) = I n f o ( ϕ ) E V a r [ S c ( ϕ ; d ) | d o ]
The variance term in (2) directly reflects the information loss induced by some variables being missing for some units. It explicitly shows the uncertainty associated with the missing components of the score function being replaced by their expectations given what is observed in Equation (1).
These results are quite general and apply when d o is a subset of d. In general, for missing data items we write p ( I , x ; ϕ ) = p ( I | x ; ψ ) p ( x ; θ ) , where ψ are the parameters of the missing data items process. We assume ψ , θ are distinct. Then p ( I , x o b s ; ϕ ) = p ( I | x ; ψ ) p ( x ; θ ) d x m i s . If the data items are MAR p ( I | x ; ψ ) = p ( I | x o b s ; ψ ) and the likelihood will split into two factors, p ( I , x o b s ; ϕ ) = p ( I | x o b s ; ψ ) p ( x o b s ; θ ) . Estimation of ψ and θ can proceed separately and the process that led to missing data can be ignored for estimation of θ . Estimation and inference for θ can be based on p ( x o b s ; θ ) and the resulting score and information functions. For a specified PMD the parameter ψ is known.

2.3. Designing a PMD Using Expected Information

While we can consider using the observed information once the PMD has been implemented and data items collected, at the design stage, d o is not available. Thus, we need to consider the expected information, I n f o M ( ϕ ) . Since we assume a simple noninformative design, this expectation takes δ U as fixed. When the parameters ψ , θ are distinct, i n f o ( ϕ ; d o ) will be block diagonal and the relevant component for θ will be i n f o ( θ ; d o ) . Taking expectations with respect to x o b s a n d I , the relevant component of the expected information is I n f o M ( θ ) = E[ i n f o ( θ ; d o ) ].
Developing a PMD involves deciding upon p ( I | x o b s ; ψ ) . To evaluate different PMDs we can consider V M ( θ ^ ) = I n f o M ( θ ) 1 or functions of it such as t r V M ( θ ^ ) or d e t V M ( θ ^ ) .

2.4. PMDs for Noninformative Sample Designs

Many sample designs use auxiliary variables available for all population units before the sample is selected, for example, through use of stratified sampling. Denote these population values by z U . These variables may also be used in a PMD. Maximum likelihood estimation for data items that are MCAR for noninformative sample designs is considered in Section 2.4 of [1]. Here we extend this to the MAR case and focus on the expected information that can be used for developing the PMD and the sample design.
To allow for the sample design, we consider the sample membership indicator vector δ U . The complete population data is d U = ( x U , I U , δ U , z U ) . The observed data are d o = ( x o b s , I s , δ U , z U ) . Three assumptions can be made:
A1: Sampling is noninformative given z U , so p ( δ U | x U , I U , z U ) = p ( δ U | z U ) .
A2: Data items missing for sample units does not depend on non-sampled units, so p ( I s | x U , z U ) = p ( I s | x s , z U ) .
A3: In the sample data items are MAR, so p ( I s | x s , z U ) = p ( I s | x o b s , z U ) .
  • With these conditions the likelihood factors
p ( d o ) = p ( δ U | z U ) p ( I s | x o b s , z U ) p ( x o b s , z U )
Assuming the parameters for these three factors are distinct estimation and inference, we can use p ( x o b s , z U ) . We assume that once the auxiliary variables are recognised, the parameter of interest θ is redefined to characterise p ( x U , z U ) . If the parameters of the marginal distribution p ( x U ) are of interest, then we need to consider how they are related to the parameters of the joint distribution, since p ( x U ) = p ( x U , z U ) d z U = p ( x U | z U ) p ( z U ) d z U .
From (4) the observed information will be block diagonal with the component relevant to θ being i n f o ( θ ; d o ) . At the stage of designing a PMD, we can consider the expected information conditional on δ U and z U , I n f o M ( θ ) = E[ i n f o ( θ ; d o ) | δ U , z U ].
We will not directly consider sample designs using auxiliary variables. However, Equation (4) indicates how more complex designs can be handled by including z U in p ( x o b s , z U ) . A practical feature is that auxiliary variables available for sample design can also be used in the PMD.

3. Multinomial Model

To explore PMDs in more detail we consider a P dimensional table and associated multinomial model for the cell frequencies. Such tables are often a key output from a sample survey.

3.1. Complete Data

Consider the P vector x i = ( x 1 i , , x P i ) of categorical data items, where x p i has l p levels. Denote the complete sample data for unit i by d i = { x i } . The complete data d defines a P-way contingency table with C = Π p = 1 P l p cells. Assume that x i and x j are independent for i j and the sample design for s is noninformative and no auxiliary variables are used.
Define W i = ( W i 1 , , W i a , , W i C ) to be a C x 1 vector where W i a = 1 if unit i belongs to the ath cell of the contingency table and W i a = 0 otherwise, where a = 1 , C . The distribution of the cell counts n a = Σ i W i a in the contingency table is assumed to be multinomial with parameter π = ( π 1 , , π a , , π C ) , where π a is the probability for cell a. The elements of π are constrained to sum to 1. For identifiability, the constraint is removed by dropping a parameter, say π C , from π and noting π C = 1 Σ c = 1 C 1 π c . Accordingly, we consider the parameter vector θ = ( π 1 , , π a , , π C 1 ) . Thus, θ consist of the first C 1 elements of π . Once data have been obtained, the observed log-likelihood is
Σ i Σ a = 1 C W i a l n π a = Σ a = 1 C n a l n π a
Differentiation with respect to θ gives the score function:
S c ( θ ; d ) = S c ( π 1 ; d ) , , S c ( π a ; d ) , , S c ( π C 1 ; d )   where S c ( π a ; d ) = Σ i W i a π a 1 Σ i W i C π C 1 = Σ i S c i a = n a π a 1 n C π C 1 ,
The contribution of unit i to the score for π a is S c i a = W i a π a 1 W i C π C 1 . The ( a , b ) th element of the observed information i n f o ( θ ; d ) = θ S c ( θ ; d ) = i i n f o i ( θ ; d ) where i n f o i ( θ ; d ) has the ( a , b ) th element
= W i a π a 2 + W i C π C 2           i f         a = a = W i C π C 2           i f         a   b ,
The ( a , b ) th element of i n f o ( θ ; d )
= n a π a 2 + n C π C 2           i f         a = a = n C π C 2           i f         a   b ,
In the likelihood framework, the solution to S c ( θ ; d ) = 0 leads to the Maximum Likelihood Estimate θ ^ = ( π ^ 1 , . . . , π ^ a , . . . , π ^ C 1 ) where π ^ a = i s W i a / n = n a / n (see p. 17 [27]). Given the observed data the variance of θ ^ could be estimated by i n f o 1 ( θ ^ ; d ) .
The expected information can be obtained by taking the expectation of the observed information with respect to the complete data, noting E ( W i a ) = π a . This is I n f o ( θ ) = E [ i n f o ( θ ; d ) ] = i I n f o i ( θ ) where I n f o i ( θ ) is a ( C 1 ) × ( C 1 ) matrix with ( a , b ) th element
= π a 1 + π C 1           i f         a = b = π C 1           i f         a   b
The variance of the asymptotic distribution of θ ^ is I n f o ( θ ) 1 , which has ( a , b ) th element
= 1 n π a ( 1 π a )           i f         a = b = 1 n π a π b           i f         a   b
At the design stage no data are observed and so the observed information cannot be calculated. Instead, we use the expected information with respect to the complete data.
When the data has been collected variance estimation can be based on plugging the MLE, θ ^ , into the observed or expected information. In this case this results in the same estimated variance matrix.
These results are usually expressed solely in terms of the cell counts, n a . We show that for situations involving missing data items there are advantages in considering score and information in terms of W i a .

3.2. Observed Information in Missing Data

An important constraint on a PMD applied to the multinomial model is that all variables need to be collected from a subsample of units, otherwise the ML estimator of θ cannot be identified. This is because it is the saturated model allowing for all the interactions. The observed data on i is d o , i = { x o b s , i , I i } and for the sample is d o . The MLE of θ using d o can be found using the missing information principle, in which the score function is obtained by replacing a complete data statistic by its expectation conditional on the observed data. Thus, in (6) W i a is replaced by Π i a = E W i a d o , i . Given d o , i , define F i as the set of cells that are feasible, (i.e., to which unit i could belong to given d o , i ) and W i a o = 1 if a F i and 0 otherwise, which is observed. The total probability of the feasible cells for i is ϕ i = c F i π c = c W i c o π c and so Π i a = W i a o π a / ϕ i and the ath element of the score function is
S c ( π a ; d o ) = Σ i W i a o ϕ i 1 Σ i W i C o ϕ i 1
Conditioning on d o , i narrows the set of cells to which unit i could belong to F i and the probability of belonging to cell a F i is proportional to π a .
In general, W i a is not observed, so we do not know which cell i belongs to, but W i a o , which indicates which cells i could belong to, is observed. Comparing (11) with (6), we see that W i a is replaced by W i a o , which is non-zero for all feasible cells given what has been observed. Moreover, π a is replaced by ϕ i , the probability of the feasible cells.
To illustrate the notation, consider the data, d, in Table 2 which is taken from Table 13.8, p. 284 of [16]. The variables are ( x 1 , x 2 , x 3 ) where x p = 1 , 2 and n = 715 . The sample data define a 2 × 2 × 2 table with C = 8 . Let the indexing of cells by a follow the ordering of top down and left to right so that ( π 1 , π 2 , , π 8 ) = (3, 176, 4, 293, 17, 197, 2, 23)/715. If only x 1 i = 1 is collected from unit i, then d o , i = { x 1 i = 1 , I i = ( 1 , 0 , 0 ) } , the set of the cells unit i could belong to is F i = { 1 , 2 , 3 , 4 } , the probability of these cells is ϕ i = (3 + 17 + 4 + 2)/715, and the conditional probability that unit i belongs to cell a = 2 given d o , i is Π i 2 = 17 / ( 3 + 17 + 4 + 2 ) .
When data are missing, there is a loss of information that increases the variance of the MLE, which is reflected in the information function. The ( a , b ) th element of the observed information function can be obtained directly by differentiating S c ( π a ; d o ) with respect to π b . Using ϕ i 1 / π b = ϕ i 2 ( W i b o W i C o ) the element is
Σ i ( W i a o W i C o ) ( W i b o W i C o ) ϕ i 2
For unit i the coefficient of ϕ i 2 is 1 when a and b are both feasible and C is not or when neither a or b are feasible but C is, and 0 otherwise. Define F ( a ) = { i : a F i } as the set of units for which cell a is feasible, and then the ( a , b ) th element is
i F ( a b C ¯ ) ϕ i 2 + i F ( a ¯ b ¯ C ) ϕ i 2
The contribution of unit i to the observed information can be determined from (12) and depends on the feasibility of cells a, b and C given d o , i , resulting in the following cases:
( C a s e   1 ) = 0           i f         a = b F i         a n d         C F i ( C a s e   2 ) = ϕ i 2           i f         a = b F i         a n d         C F i ( C a s e   3 ) = 0           i f           a   b ,         a       a n d / o r     b , C F i         ( C a s e   4 ) = ϕ i 2           i f           a       a n d     b F i , C F i ,     ( i n c     a = b ) ( C a s e   5 ) = ϕ i 2           i f         a b ,         a       a n d     b     F i , C F i ( C a s e   6 ) = 0           i f         a       a n d / o r     b F i , C F i     ( i n c     a = b )
For example, if x 1 i = 1 , then the contribution of that unit to each of cells a = 1 , 2 , 3 , 4 is ϕ i 2 and zero for all the other cells.
The loss of information can be considered from the MIP
i n f o ( θ ; d o ) = E [ i n f o ( θ ; d ) | d o ] V a r [ S c ( θ ; d ) | d o ] .
Here the ( a , b ) th element of E [ i n f o ( θ ; d ) | d o ] is:
= Σ i W i a o π a 1 ϕ i 1 + Σ i W i C o π C 1 ϕ i 1           i f         a = b = Σ i W i C o π C 1 ϕ i 1           i f         a   b ,
The information loss from observing d o instead of d is
V a r [ S c ( θ ; d ) | d o ] = i l i
where l i = C o v [ S c i , S c i | d o b s , i ] . After expanding S c i , l i has ( a , b ) th element
l i a b = π a 1 π b 1 C o v [ W i a , W i b | d o , i ] π C 1 π b 1 C o v [ W i C , W i b | d o , i ]   π a 1 π C 1 C o v [ W i a , W i C | d o , i ] + π C 2 V a r W i C | d o , i ]
Given the observed W i a o , W i a = W i a m W i a o where W i a m = 1 if i a and 0 otherwise. Conditional to W i a o , the vector W i m = ( W i 1 m , , W i a m , , W i C m ) has a categorical distribution (i.e., multinomial with n = 1 ) where W i a m has probability Π i a . Hence
C o v [ W i a , W i b | d o , i ]   = W i a o W i b o C o v [ W i a m , W i b m | d o , i ]   = W i a o Π i a ( 1 Π i a )           i f         a = c   = W i a o W i b o Π i a Π i b           i f         a   b
Since Π i a = W i a o π a / ϕ i
π a 1 π b 1 C o v [ W i a , W i b | d o , i ]   = W i a o ( π i a 1 ϕ i 1 ϕ i 2 )           i f         a = b   = W i a o W i b o ϕ i 2           i f         a   b
Thus, Equation (20) becomes
  = π a 1 ϕ i 1 ϕ i 2           i f         a = b       a n d       a F i   = ϕ i 2           i f         a   b       a n d       a , b F i   = 0           i f         a     o r       b     F i     ( i n c     a = b )
After substituting this into l i , we see it has ( a , b ) th element
( C a s e   1 ) = ϕ i 1 π a 1 + π C 1 ϕ i 1           i f         a = b F i         a n d         C F i ( C a s e   2 ) = ϕ i 1 π a 1 ϕ i 2           i f         a = b F i         a n d         C F i ( C a s e   3 ) = ϕ i 1 π C 1           i f           a   b ,         a       a n d / o r     b , C F i         ( C a s e   4 ) = ϕ i 1 π C 1 ϕ i 2           i f           a       a n d     b F i , C F i ,     ( i n c     a = b ) ( C a s e   5 ) = ϕ i 2           i f         a b ,         a       a n d     b     F i , C F i ( C a s e   6 ) = 0           i f         a       a n d / o r     b F i , C F i     ( i n c     a = b )
For the sample, the ( a , b ) th element of the loss is l a b
  = i F ( a ) ϕ i 1 π a 1 + i F ( C ) ϕ i 1 π C 1 i F ( a C ¯ ) ϕ i 2 i F ( a ¯ C ) ϕ i 2           i f         a = b   = i F ( C ) ϕ i 1 π C 1 i F ( a b C ¯ ) ϕ i 2 i F ( a ¯ b ¯ C ) ϕ i 2           i f         a   b ,
The loss can also be expressed in terms of the observed W i a o
= Σ i ( W i a o π a 1 + W i C o π C 1 ) ϕ i 1 Σ i ( W i a o W i C o ) 2 ϕ i 2             i f         a = b = Σ i W i C o π C 1 ϕ i 1 Σ i ( W i a o W i C o ) ( W i b o W i C o ) ϕ i 2             i f         a   b ,
These results can also be obtained using V a r [ S c ( θ ; d ) | d o ] = E [ i n f o ( θ ; d ) | d o ] (given by (16))— i n f o ( θ ; d o ) (given by (12)).

3.3. Expected Information Given an MCAR PMD

The observed information and information loss for unit i are a function of x o b s , i and I i . For MCAR PMDs whose variables are missing and independent of the values of the variables, p ( I i , x o b s , i ; ϕ ) = p ( I i ; ψ ) p ( x o b s , i ; θ ) . At the design stage of an MCAR we can fix I i or, in other words, the condition on unit i being assigned data pattern j. Consider a fixed missing data pattern j in which P j variables are observed and the number of cells in the P j w a y contingency table produced by the variables is K j , so that x o b s , i can take values k for k = 1 , , K j . Let F j k be the set of feasible cells in the full contingency table given x o b s , i = k . At the design stage of an MCAR we have not yet observed x o b s , i , so we are interested in the expected information and information loss with respect to the distribution p ( x o b s , i | I i = j ) . For given pattern j, x o b s , i = k for any cell a F j k , so the total probability of the feasible cells is p ( x o b s , i = k | I i = j ) = ϕ j k = c F j k π c = c W k c j π c . Here W k a j = 1 if a F i and 0 otherwise.
When x o b s , i = k and I i = j , W i a o = W k a j and ϕ i = ϕ j k , and from (12) the contribution to the observed information for ( a , b ) is ( W k a j W k C j ) ( W k b j W i C j ) ϕ j k 2 . The expected information is obtained by multiplying by p ( x o b s , i = k | I i = j ) = ϕ j k and summing over k. This gives
I n f o M a b j = k = 1 K j ( W k a j W k C j ) ( W k b j W i C j ) ϕ j k 1
Note that (25) is useful and novel in that it explicitly shows the contribution of pattern j to the expected information.
The contribution of the possible value k to the expected information depends on the feasibility of cells a, b and C given k, resulting in the following cases:
( C a s e   1 ) = 0           i f         a = b F j k         a n d         C F j k ( C a s e   2 ) = ϕ j k 1           i f         a = b F j k         a n d         C F j k ( C a s e   3 ) = 0           i f           a   b ,         a       a n d / o r     b , C F i         ( C a s e   4 ) = ϕ j k 1           i f           a       a n d     b F j k , C F j k ,     ( i n c     a = b ) ( C a s e   5 ) = ϕ j k 1           i f         a b ,         a       a n d     b     F j k , C F j k ( C a s e   6 ) = 0           i f         a       a n d / o r     b F j k , C F j k     ( i n c     a = b )
Suppose we consider a PMD in which n j units are allocated to pattern j, then the ( a , b ) th element of the expected information for the sample is j n j I n f o M a b j .
The same approach can be used to obtain the expected information loss, L i j = E ( l i | I i = j ) , which has element ( a , b )
k = 1 K j ( W k a j π a 1 + W k C j π C 1 ) k = 1 K j ( W k a j W i C j ) 2 ϕ j k 1             i f         a = b k = 1 K j W k C j π C 1   k = 1 K j ( W k a j W k C j ) ( W i b j W k C j ) ϕ j k 1             i f         a   b ,
The contribution of the possible value k to the expected information loss depends on the feasibility of cells a, b and C given k, resulting in the following cases:
( C a s e   1 ) = π a 1 + π C 1           i f         a = b F j k         a n d         C F j k ( C a s e   2 ) = π a 1 ϕ j k 1           i f         a = b F j k         a n d         C F j k ( C a s e   3 ) = π C 1           i f           a   b ,         a       a n d / o r     b , C F j k         ( C a s e   4 ) = π C 1 ϕ j k 1           i f           a       a n d     b F j k , C F j k ,     ( i n c     a = b ) ( C a s e   5 ) = ϕ j k 1           i f         a b ,         a       a n d     b     F j k , C F j k ( C a s e   6 ) = 0           i f         a       a n d / o r     b F j k , C F j k     ( i n c     a = b )
Suppose we consider a PMD in which n j units are allocated to pattern j, then the expected information loss for the sample is j n j L j . Subtracting L j from E [ i n f o ( θ ; d ) | I i = j ] also gives the expected information (25).

3.4. Expected Information Given an MAR PMD

In an MAR the missing data pattern depends on at least some elements of x o b s . In a simple example the variable x 1 is observed and which of the remaining variables is obtained depends on the value of x 1 . In a health survey x 1 might be categories of BMI. If the data items are MAR, p ( I | x ; ψ ) = p ( I | x o b s ; ψ ) . For example, people with a BMI > 35 are given a higher probability of receiving detailed questions on diet and exercise. In general the distribution of the observed data for a unit in the MAR PMD can be factorised as
p ( d o ) = p ( I , x o b s ; ϕ ) = p ( I | x o b s ; ψ ) p ( x o b s ; θ )
where p ( x o b s ; θ ) is defined, in this section, by the multinomial distribution, p ( I | x o b s ; ψ ) defines the missing data mechanism and ψ are the parameters of the PMD. More generally, p ( x o b s ; θ ) is given by the relevant density of the model being considered.
Unlike an MCAR PMD, in an MAR PMD we do not fix the missing data patterns at the design stage and there is more flexibility in defining d o . The expected information from an MAR SQD design is obtained by taking the expectation of the observed information with respect to the joint distribution of d o = ( I , x o b s ) . Thus, the expected information is
I n f o M ( θ ) = d o p ( d o ) i n f o ( θ ; d o )
where d o is a summation over all possible values of observed data, i.e., ( I , x o b s ) and i n f o ( θ , d o ) is the observed information function given in Section 3.2.
We use the MIP
I n f o M ( θ ) = I n f o ( θ ) d o p ( d o ) l ( θ ; d o )
where l ( θ ; d o ) is the loss of the observed information function given in Section 3.2. The second terms on the RHS is the expected information loss for a given PMD. Assuming that p ( d o , i ) are independent over i,
I n f o M ( θ ) = i s d o , i p ( d o i ) i n f o ( θ ; d o , i )   = I n f o ( θ ) i s d o , i p ( d o , i ) l ( θ ; d o , i )
This form is useful when considering planned missing designs, as we do in Section 4, as it gives the information loss for specific values of I i . For example, if the expected information loss is high when d o , i = ( x 1 , x 2 ) = ( 1 , 2 ) , then it would seem sensible to make this impossible in an MAR PMD. For example, the probability that d o , i = ( x 1 , x 2 ) = ( 1 , 2 ) could be set to zero if the design always collects x 3 when ( x 1 , x 2 ) = ( 1 , 2 ) .
Equations (25)–(32) explicitly show the contribution of individual missing data patterns to the expected information for MCAR and MAR PMDs. These results enable the expected information for a specific proposed PMD to be evaluated. They also can serve as the objective function in the search for optimal or efficient designs. Next we specify p ( I i | x i ) or p ( d o , i ) for an MAR PMD using a sequential approach.

3.5. Sequential PMD

In general an MAR PMD is any design for which p ( I | x ; ψ ) = p ( I | x o b s ; ψ ) . This could involve a two phase approach. Let I ( k ) denote the variables collected at phase k. At the first phase a subset of variables is collected from unit i using an MCAR approach, leading to I ( 1 ) and x o b s ( 1 ) . At the second phase a subset of the remaining variables is collected using MAR based on x o b s ( 1 ) , with probability p ( I ( 2 ) | x o b s ( 1 ) ; ψ ) .
We will consider a sequential MAR PMD in which there is a fixed sequence in which the elements of x are collected: x 1 is always collected first, x 2 may be collected second, x 3 may then be collected, and so on to x P . Under a sequential MAR PMD, let the conditional probability of collecting x 2 be ψ 2 | x 1 = E ( I 2 | I 1 = 1 , x 1 ) . In general, let the conditional probability of collecting x p be ψ p | d o b s , ( p 1 ) = E ( I p | d o b s , ( p 1 ) ) where d o b s , ( p 1 ) is the observed data on the first p 1 variables in the sequence. The first p 1 elements in the sequence of x have missing elements replaced by ‘.’. The MAR PMD design parameters are ψ ψ = { ψ p | d o b s , ( p 1 ) , p = 2 , P } , the set of all sequential parameters for all d o b s , ( p 1 ) allowed under the sequential design. A sequential design is not Markov in nature since the probability variable x p is collected can depend on more than the value of x p 1 , although this restriction could be imposed.
Under the sequential design, we can consider the factorisation,
p ( I | x o b s ; ψ ) = Π p = 2 P p ( I p | d o b s , ( p 1 ) )
and so
p ( I | x o b s ; ψ ) = Π p = 2 P ψ d o b s , ( p 1 ) I p ( 1 ψ d o b s , ( p 1 ) ) 1 I p
Figure 1 illustrates the notation in the 2 × 2 × 2 case. If x 1 = 1 , then the probability of collecting x 2 is ψ 2 | x 1 = 1 or ψ 2 | ( 1 ) for short. If x 2 is collected and x 2 = 1 , then the probability of collecting x 3 is ψ 3 | x 1 = 1 , x 2 = 1 or ψ 2 | ( 1 , 1 ) . If x 3 is not collected, then the sample unit could either belong to cells indicated by a = 1 (when x 3 = 1 ) or a = 2 (when x 3 = 2 ), so the set of feasible cells is F i = { 1 , 2 } , d o = ( 1 , 1 , . ) where the ‘.’ indicates the missing value of x 3 , and the probability of the observed data from an MAR SQD is p ( d o ) = p ( I , x o b s ) = ( π 1 + π 2 ) ψ 2 | ( 1 ) ( 1 ψ 3 | ( 1 , 1 ) ) . This p ( d o ) is then used in (32).
Different orderings can be considered which may lead to different efficiencies, and two or more options can be evaluated. However, there will often be practical reasons to adopt a particular ordering, for example, easy to collect and important variables can be included early in the sequence.

4. Developing MCAR and MAR PMDs

4.1. Missing Data Effects and A and D Optimality

Here we consider a standard survey of size n involving complete cases (CCs), and MCAR and MAR PMDs. For a PMD the variance matrix of θ ^ is V M ( θ ^ ) = I n f o M 1 ( θ ) , which has elements V M ( a , b ) . For the sample comprising data for complete cases, the variance matrix is V C C ( θ ^ ) = I n f o 1 ( θ ) , which has elements V C C ( a , b ) . For convenience we will suppress θ ^ . A simple way to examine the effect of a PMD on π ^ a is through V M ( a , a ) / V C C ( a , a ) , where both variances are based on n, which we will call the Missing Data effect (MDeff(a)). This is analogous to the design effect (Deff) which measures the variance of an estimate under a complex sample design divided by the variance of the same estimate under a Simple Random Sample of the same sample size [28]. We can also define the effective sample size for the PMD as n a = n / MDeff(a). Thus, an MDeff(a) of 2 corresponds to a reduction in the effective sample size of 50 percent for π ^ a .
The MDeff(a) is parameter specific. An overall measure can be obtained using the concept of A-optimality, which is based on A M = t r V M . Define A C C = t r V C C . This leads to the overall missing data effect, MDeff A = A M / A C C , and associated effective sample size, n A * = n / MDeff A .
Since MDeff A = a = 1 C 1 MDeff(a) V C C ( a , a ) / a = 1 C 1 V C C ( a , a ) , it is a weighted average of the variables’ specific MDeff(a), with weights proportional to V C C ( a , a ) , which is equal to V ( π ^ a ) . We can also consider giving each variable the same weight, leading to MDeff A W = 1 C 1 a = 1 C 1 MDeff(a). This is equivalent to using weights π a 1 ( 1 π a ) 1 in A M , which gives A W .
Corresponding measures based on D-optimality, that is the determinant of the variance matrices, D M = d e t V M and D C C = d e t V C C , can also be used. This leads to the overall missing data effect, MDeff D = D M / D C C , and associated effective sample size, n D * = n / MDeff D .
The parameter θ does not include π C because π C = a = 1 C 1 π a = 1 1 T θ . Then π ^ C = 1 1 T θ and V M ( π ^ C ) = 1 T V M ( θ ^ ) 1 and is included in global A optimality but not included in global D optimality.
These measures can be used to evaluate any specified PMD being considered. In developing a PMD the expected costs also need to be considered. For design M a simple cost function is C M = n M ( c f + c M ) where c f is the fixed cost per sample unit, n M is the number of units selected for design M and c M is the expected cost of the variables collected per unit. In general c M = E M [ p = 1 P c p I p ] , where c p is the cost of collecting variable p. Cost can include collection and processing costs for the survey organisation and also a quantification of the burden imposed on respondents. In the equal cost scenario, we assume c p = 1 so the cost of collecting each variable is the same. This will be a reasonable assumption in many surveys, but there will also be surveys where some variables are more costly to collect. Then for a standard survey in which all the variables are collected, the cost is n C C ( c f + p ) . For a survey using MCAR there are up to 2 P 1 possible patterns and the cost is n M C A R ( c f + c M C A R ) , where c M C A R = j = 1 2 P 1 p j P j , where p j is the probability of pattern j. For a survey using MAR, the cost is n M A R ( c f + c M A R ) , where c M A R = d 0 p ( d 0 ) P d 0 .
In some situations some variables may impose more reporting burden or collection costs than others or costs may vary in a disproportionate manner. For example, the cost of collecting a subset of variables may be more than the sum of the cost of collecting each variable. Table 3 describes other cost scenarios. In the unequal cost scenario, the cost of collecting x 3 is three times the cost of collecting x 1 or x 2 . The disproportionate cost scenario is the same as the equal cost scenario, except that collecting all variables together costs six times the cost of collecting a single variable.
The multinomial distribution allows for all interactions between elements of x . Accordingly, for estimation of π to be supported, all variables must be collected by a core sample of sufficient size to ensure the parameters are identified and the large sample ML approximation applies. Thus, we assume a core sample of size n c o r e is also selected. The potential gains of a PMD are greater when the core sample is small so that its cost is small compared with the total cost. This should be balanced against identifiability considerations which suggest having at least two cases in each cell. The PMD may also involve all variables being collected from some sample units. It would be possible to not use a core sample, and in searching for optimal MCAR and MAR designs, reject those for which I n f o M ( θ ) is non-singular. In practice having a small core sample of size n c o r e will make the search for optimal designs quicker and easier. To ensure that using a core sample is not seriously affecting the efficiency, its size could be varied to assess the sensitivity of the PMD to the core sample size. An indication that the core sample is too large would be the PMD not having any complete cases.
The cost of the core sample will be added to the cost for each design, so the total sample size is n c o r e + n M . To examine the cost effectiveness of a design, M, n M will be set to have the same costs as the complete case sample of size n C C , so n M C A R , n M A R n C C , and the use of a PMD allows a larger overall sample size for given total cost.

4.1.1. Example 1

Consider the design data in Table 2, where P = 3 and each variable has l p = 2 categories. There are eight parameters for the sequential MAR PMD: two values for ψ 2 | x 1 = P ( I 2 | x 1 ) , two values for ψ 3 | x 1 ,   . = P ( I 3 | I 2 = 0 , x 1 ) and four values for ψ 3 | x 1 , x 2 = P ( I 3 | I 2 = 1 , x 1 , x 2 ) . These are illustrated in Figure 1. There are four patterns of missing data under the sequential pattern given by j = 1 , 3 , 6 , 7 in Table 1. To make a fair comparison, an MCAR PMD is restricted to these same four patterns and has the design parameter p j , which is the proportion of the sample assigned to pattern j. In our example we set n c o r e = 715 , the sample size in Table 2, which meets the criterion of having at least two cases in each cell, and n C C = 10,000, so the total cost when c f = 0 is 32,145.
The optimal design parameters, p j , for MCAR and ψ . | . s for MAR, were found by a grid search, which is feasible when the number of design parameters is small. A large number of parameters can arise when there are more dimensions (i.e., P > 3) or the number of levels, l p , is large. In this situation more efficient algorithms, such as stochastic optimization, need to be considered.
Table 4 gives the reduction in A-, A W - and D-optimality for optimal MAR and sequential MCAR designs relative to a CC design for the same total cost. Zero reduction occurs when the optimal PMD corresponds to the standard complete case design. The results show that MCAR is marginally more efficient than a CC design when costs are disproportionate (e.g., 3.5% for A optimality when c f = 0 ) but is not more efficient than CC when costs are equal or unequal. In contrast, the MAR design is significantly more efficient (up to 92% reduction ) than a CC design for D optimality and still more efficient for A- and A W -optimality when costs are unequal or disproportionate. This is because when costs vary considerably, there is more cost reduction possible when not collecting the variables that are more costly to collect. The greater reductions obtained when considering D-optimality suggest that results depend on the criteria used to evaluate the PMD. In practice several criteria can be examined.
Table 5 shows how the optimal PMD designs change according to c f and the optimality criteria when costs are disproportionate. The case of c f = 0 reflects cases when the reporting and collection costs are much greater, or of more concern, than the cost of the initial contact. PMDs allow for an increase in sample size, which is how the reduction in variance compared with the standard survey can arise. D optimality leads to the highest sample sizes (over 30,000 when c f = 0 ). As c f increases the total sample size approaches 10,000, the sample size for CC, since less of the cost is associated with collecting each variable. Let p ˜ 1 , p ˜ 12 , p ˜ 13 and p ˜ 123 be the proportion of units where only x 1 , ( x 1 , x 2 ) , ( x 1 , x 3 ) and ( x 1 , x 2 , x 3 ) are collected, respectively. Then p ˜ 1 + p ˜ 12 + p ˜ 13 + p ˜ 123 = 1 . We see that D-optimality leads to designs where a proportion of the sample collects only x 1 , whereas A-optimality never leads to such designs, and A W optimality only once. The fact that ψ 2 | 1 = ψ 3 | 12 = ψ 3 | 11 = 1 for PMDs using D- and A W -optimality means all variables are collected when x 1 = 1 . This is because P r o b ( x 1 = 1 ) is small and any missing data when x 1 = 1 leads to high information loss. Other common features of the PMDs are: ψ 3 | 1 . = ψ 3 | 2 . = 0 , which means collecting x 3 without x 2 is not worthwhile, regardless of the value of x 1 ; and ψ 3 | 21 = ψ 3 | 22 , which means collecting x 3 does not depend on the value of x 2 if x 1 = 2 .
A-, A W - and D-optimality are overall measures. The effect of a PMD can also be examined for the individual parameters by comparing the ratio of variance under the optimal design with that of a CC design of the same cost, which are the diagonal elements of the variance matrices used for Table 4. Table 6 gives results for the optimal design for a sequential MAR PMD, with c f = 0 and disproportionate costs for these different overall criteria. We see that for D optimality the PMD reduced the variance of π ^ j for j = 1 , 2 , 3 , 4 to a third of that of a CC design at the cost of increasing the variance of the other parameters, particularly for j = 4 , 5 .
Table 7 gives the MDeff(a) for PMDs when c f = 0 and variable costs are disproportionate. MDeff(a) measures the impact of missing data on the variance of parameter estimates at the design stage. The MDeff can be obtained from the preceding tables by multiplying by ( n c o r e + n M ) / ( n c o r e + n C C ) . It shows the PMD for A-optimality loses efficiency in all parameter estimates due to the missing data. For example, the variance of π ^ 8 is 2.3 times larger than a CC design with the same sample size, n. In contrast, for the PMDs for A W - and D-optimality, the variance of π ^ 1 to π ^ 4 is the same as that for CC with the same sample size. As mentioned earlier, this is because all variables are collected when x 1 = 1 and so there is no loss in efficiency when estimating the four corresponding multinomial parameters.
MAR PMDs that are developed based on the overall criteria may have detrimental effects on the parameter estimates for individual cells, for example, low probability cells, if the variables defining those cells are missing more frequently, although the reverse can also occur with the PMD leading to those cells being missing less frequently. This suggests that the effects on individual parameters can be examined, as in Table 6 and Table 7.

4.1.2. Example 2

We repeated the process using A-optimality and D-optimality for the marginal means (i.e., proportions): y 1 = δ ( x 1 = 1 ) , y 2 = δ ( x 2 = 1 ) and y 3 = δ ( x 3 = 1 ) with corresponding mean parameters μ 1 = 1 π 5 π 6 π 7 π 8 , μ 2 = π 3 + π 4 + π 7 + π 8 , and μ 3 = π 2 + π 4 + π 6 + π 8 , respectively. If μ ^ = ( μ ^ 1 , μ ^ 2 , μ ^ 3 ) , V a r M ( μ ^ ) has ( k , k ) element C o v M ( μ ^ k , μ ^ k ) = c k T V ( θ ^ ) c k , where c k depends on the expression for μ k in terms of the π parameters (e.g., c 1 = ( 0 , 0 , 0 , 1 , 1 , 1 , 1 ) . Table 8 presents the reduction in variance for the mean estimates for the MAR and MCAR optimal designs using the A-, A W - and D-optimality criteria compared to a sample of CCs with the same total cost. The results show that MCAR designs are more efficient than the CC design when costs are unequal or disproportionate, with the strongest gains for D-optimal design. The MAR designs are more efficient than the MCAR designs. As before, gains are greater when c f is smaller.
The results show that MAR designs can dramatically reduce the variance for some parameter estimates while increasing it for others. There are trade-offs that may need to be considered, especially if the survey objectives prioritize specific parameters. As part of the design, the importance of different parameters needs to be identified and the parameter specific evaluations considered as in Table 6 and Table 7, as well as the overall criteria. In some cases the importance of a specific parameter can be incorporated into the criteria, for example, importance weight in A W -optimality.
To develop a PMD some assumptions have to be made on the values of the cell probabilities, π . This is a general issue affecting the development of any survey design—the efficiency of a possible design will depend on the values of some unknown population parameters. Practice values are obtained from previous or related surveys, pilot data or administrative data. It is useful to check the sensitivity of the design to misspecification of the underlying multinomial probabilities used at the design stage by varying the parameters. For a contingency table there may be data available on the marginal distribution of the variables. These can be combined with data or assumptions on interactions, including that some may be zero, to produce likely values of the cell probabilities.

4.2. Marginal Effects of Missing Data for a Single Unit

At the sample design stage, there are advantages in having an analytic expression for the marginal effect or influence that a particular value for d o , i has on the optimality of an MAR PMD. (It may be worthwhile incorporating cost into the assessment of the marginal effect, but we do not here for simplicity.) First, this could be incorporated into criteria for splitting a questionnaire, as in the data-driven approach in [7], that avoids having to recalculate the information matrix for each possible split. Second, it can be used to constrain elements of ψ so that undesirable values of d o , i either do not occur or occur with a low probability. This simplifies the design and reduces the computation required to find the optimum. Below we consider how to assess this impact and illustrate the approach using data in Table 2.
Take the situation where a large sample of size n is taken of complete cases. Consider the option of instead collecting d o i from a single unit and incurring a loss of information given by l ( θ ; d o i ) . We can apply (32) where we have set p ( d o , i ) = 1 for the unit, and for all other units complete data are obtained. Then
I n f o M ( θ ) = I n f o ( θ ) l ( θ ; d o i )
let l i = l ( θ ; d o i ) . Taking the inverse of I n f o M ( θ ) gives
V M = V C C [ I V C C l i ] 1   V C C [ I + V C C l i ]   = V C C + Δ
Here I is the identity matrix and we have used the approximation [ I B ] 1 I + B when B is small and set Δ = V C C V C C l i . The i subscript is dropped for convenience.
Consider the effect on t r V M . Define ϕ A = n t r Δ / t r V C C ; then MDeff A = 1 + ϕ A / n and the effective sample size n A = n ( 1 + ϕ A / n ) 1 . It is useful to consider the reduction in effective sample size, which is n n A = ϕ A ( 1 + ϕ A / n ) 1 ϕ A when n is large. When the data are independently and identically distributed, V C C = V 1 / n and ϕ A = t r [ V 1 l i V 1 ] / t r V 1 . In the multinomial case
V 1 ( a , b ) = π a ( 1 π a )           i f         a = b   = π a π b           i f         a   b
For d e t V M we use the approximation d e t [ I + B ] 1 + t r B ; then similar results apply with ϕ A replaced by ϕ D = n t r [ V C C l i ] = t r [ V 1 l i ] .
For the parameter π a , using (33) gives similar results with ϕ A replaced by ϕ A ( a ) = n Δ ( a , a ) / V C C ( a , a ) = n 2 Δ ( a , a ) / V 1 ( a , a ) = Δ 1 ( a , a ) / V 1 ( a , a ) , where Δ 1 = V 1 V 1 l i . Similarly, for D-optimality we could replace ϕ D with ϕ D ( a ) , the ( a , a ) th element of V 1 l i .
Table 9 shows how the reduction in effective sample size depends on the observed values in d o i , for example, 1. It gives ϕ A and ϕ D , as well as ϕ A ( a ) and ϕ D ( a ) for a = 2 , , 8 for all possible values of d o i . Because what is missing depends on the observed variables and the values they take, these results correspond to MAR. The table shows the reduction in the effective sample size depends greatly on the observed data. For both A and D optimality, the reduction in the effective sample size for π 3 when we observe only x 1 = 1 and x 3 = 1 on a unit is equivalent to dropping the CC sample size by 102, and when we observe only x 1 = 1 , the reduction is 28 for π 2 , π 3 and π 4 . When the reduction in the effective sample size is significantly greater than 1, ϕ A ( a ) ϕ D ( a ) , meaning there is little difference whether the A and D optimality criteria are used. However, when the reduction in effective sample is close to 1, differences between ϕ A ( a ) and ϕ D ( a ) emerge. For example, for the observed data x 1 = 2 , ϕ A ( 8 ) = 1.3 and ϕ D ( 8 ) = 1.0 .
By definition, the reduction in the effective sample must be non-negative when data are missing for a unit. However, when x 1 = 1 , some ϕ A ( a ) values are negative. This occurs because l i is not positive definite and, because its elements can be much larger than V 1 1 , V 1 V i l i has negative diagonal elements. The value of x 1 = 1 is an influential data point because it is 26 times less likely than x 1 = 2 , and is the reason the approximation used to derive ϕ A ( a ) leads to negative values. This is an artefact of the approximation affecting the trace criterion for specific parameters associated with relatively rare variables and suggests the need for improved approximations in such cases. The negative values of ϕ A ( a ) for A optimality are replaced by 0s in the corresponding values for D-optimality, ϕ D ( a ) .
Another interesting feature of Table 9 is the number of zeros, indicating no difference between collecting the observed and complete data on the accuracy of a parameter. The pattern of 0s reflects, for instance, if we observe x 1 = 1 , there is no additional benefit to estimating π 5 to π 8 by further collecting x 2 or x 3 .
Table 9 also shows the expected reduction in effective sample size depending on what variables are observed. These results do not depend on the values of the variables and take the expectation of the loss for all possible values of the observed variables and correspond to MCAR. The values for MAR are much more variable than MCAR, which are between 0 and 2.0 for A optimality. This highlights the fact that knowing the values of observed variables, not just the pattern of observed variables, is more useful in measuring the impact of missing data on the accuracy of estimates.

5. Discussion

A PMD enables a reduction in the reporting load for respondents and also a reduction in data collection costs. For a PMD that uses an MCAR or MAR approach, MLEs can be obtained once the data have been collected using the observed likelihood and resulting score function. This ensures all the data collected are effectively used and correspond to full information maximum likelihood. The Missing Information Principle can be used to obtain the observed score function from the complete data score function and also the observed information function. It clearly shows the information lost due to the variables that are missing compared to estimation based on complete cases. A PMD can be more efficient than a standard survey collecting all variables in some situations since the reduction in reporting and collection costs enables a larger total sample to be used for the same total cost.
At the design stage no data have been collected, and to develop a PMD we need to consider the expected information and associated information loss. With few exceptions, previous works on PMDs have only allowed the data to be MCAR. This paper gives a framework for designing a PMD that allows the data to be MAR, so that the pattern of variables for which data are collected depends on the values of some of the variables. We focussed on a simple MAR PMD, called a sequential PMD, but there are many other options to be considered. A feature of MAR PMDs is there flexibility and adaptivity as decisions on which variables to collect depend on the observed values of some key variables. This also improves robustness as it can be used to ensure particular combinations of variables are collected from an adequate number of units.
For many surveys the key outputs consist of contingency tables. We derive the observed and expected score and information for a multinomial model for the cell frequencies. The expected information can be used to evaluate the likely variances for any specified PMD being considered. They are then used to develop and evaluate possible MCAR and MAR designs using A- and D-optimality. The effect on the variance of the estimates of cell probabilities and the marginal probabilities are also considered. We also develop ways of assessing the marginal effect of reducing the variables collected from a single sample unit.
A number of criteria to develop and evaluate potential PMDs have been considered here. A-optimality is very easy to understand and is proportional to the average of the variance of the parameter estimates. It is easily modified by applying weights to the variances of the parameter estimate to reflect their relative importance. D-optimality minimizes the volume of the confidence ellipsoid for the parameter estimates and is a good overall measure of efficiency. In some cases, particular criteria or parameters may be of interest, but in many surveys a range of outputs are to be produced. In practice we would look at several criteria to help guide us on a sensible design.
It would be useful to have mathematically precise sufficient conditions to ensure identifiability of the multinomial parameters. We see this as a topic for further research. Identifiability will occur when the expected information matrix is non-singular. Because we are focusing on the design of a PMD, different choices of the size of the core sample can be tried and the resulting expected information matrix checked to see if it is non-singular.
In an empirical evaluation, we show that an MAR PMD can be more efficient than an optimal MCAR PMD or standard survey which collects all the variables, particularly when the costs for each variable are unequal or vary in a disproportionate way. When costs for each variable are equal, the PMDs may not be any more efficient than a standard survey. The evaluation focuses on estimating parameters of the multinomial model, which allows for all interactions between variables. A PMD can lose considerable information about interactions. Further work into designing an MAR PMD can consider tables for which higher order interactions are not of interest or other target parameters, including parameters in various analytical models (e.g., loglinear, linear and logistic regression).
Developing more effective algorithms to identify efficient PMDs is a major challenge for future research. The fundamental theory on information loss developed in this paper provides the basis for such development. Using the unit level information loss considered in Section 4.2 could be useful in developing more efficient search algorithms. If some interactions are not of interest or are zero, then that will reduce the number of parameters and the size of the search space.
As in all modelling the model and parameterisation need to be considered. This will involve structural considerations and interpretability, as well as robustness to model and parameter misspecification. We focus on a reasonably general survey producing key tables. We have not made any structural assumptions and allowed for all interactions to be present. It may be possible to assume some specific structure in some situations corresponding to some higher order interaction being absent. Considering how the efficiency of PMDs changes as higher order interactions are set to zero can be part of robustness considerations. The robustness of the PMDs to uncertainty associated with parameters values assumed at the design stage can also be examined.
This paper provides the general theory for developing MCAR and MAR PMDs based on the expected information. It considers the reasonably general case of the multinomial distribution for a contingency table, providing explicit formulas for the information, and information loss, which can be used to developed MCAR and MAR PMDs. There are a number of issues such as sensitivity to model assumptions and challenges in specifying realistic MAR mechanisms at the design stage that need further consideration. The general approach can be used to consider extensions beyond the multinomial model, such as generalized linear models or models for continuous outcomes. Ref. [3] considers multivariate normal and Ref. [8] considers logistic regression with binary covariates. An issue for analyses involving models for an outcome variable given a set of covariates is that to use the MIP it is necessary to use a model for the joint distribution of the outcome variable and the covariates, even though only the conditional distribution is of interest.

Author Contributions

Conceptualization, D.G.S. and J.C.; methodology, D.G.S. and J.C.; software, J.C.; writing—original draft, D.G.S. and J.C.; project administration, D.G.S. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A. Summary of Notation Used

Table A1. Nomenclature.
Table A1. Nomenclature.
NotationDescription
ssample
s core sample of units where all variables are collected
x p i variable x p for unit i
x i vector of variables for unit i
I i p variable indicating that variable x p is observed for unit i
d c complete sample data
d o i observed data for unit i
d o observed data
x s complete sample data
x o b s observed values
x m i s data values missing for sample units
I matrix of indicators of observed variables
S c ( ϕ ; d o ) observed score function for ϕ given d o
i n f o ( ϕ ; d o ) observed information for ϕ given d o
I n f o M ( ϕ ) expected information for ϕ for PMD M
I n f o ( ϕ ) expected information for ϕ using complete case design
l p number of levels for x p
Cnumber of cells in contingency table
π a probability for cell a in contingency table
π vector of cell probabilities
θ vector of cell probabilities removing π C
W i a variable indicating whether unit i belongs to cell a
W i vector of cell indicator variables for unit i
n a number of sample units in cell a
Π i a expectation of W i a given d o i
F i set of feasible cells given d o , i
W i a o observed indicator that cell a is feasible for unit i given d o i
ϕ i total probability of feasible cells for unit i
W i a m indicator that unit i in cell a given W i a o
F ( a ) set of units for which cell a is feasible given d o
jMCAR data pattern
F j k feasible cells for unit i for MCAR pattern j given x o b s , i = k
W k a j indicator that a is feasible unit for MCAR pattern j given x o b s , i = k
ψ p | d o b s , ( p 1 ) probability of collecting x p in sequential MAR given observed values
in first p 1 variables
V M ( θ ^ ) variance matrix of θ ^ for PMD M
MDeff(a)ratio of variance of parameter estimate for cell a for PMD to variance for complete case design
A M trace of V M ( π ^ )
D M determinant of V M ( θ ^ )
c p cost of collecting x p
l i observed information loss for unit i
ϕ A large sample reduction in effective sample size for trace criterion due to missing data for single unit i
ϕ D large sample reduction in effective sample size for determinant criterion due to missing data for single unit i
ϕ ( a ) large sample reduction in effective sample size for parameter estimate for cell a missing data for single unit i

References

  1. Chambers, R.L.; Steel, D.G.; Wang, S.; Welsh, A. Maximum Likelihood Estimation for Sample Surveys; CRC Press Taylor and Francis Group: Boca Raton, FL, USA, 2012. [Google Scholar]
  2. Kim, J.; Shao, J. Statistical Methods for Handling Incomplete Data; CRC Press: Boca Raton, FL, USA, 2021. [Google Scholar]
  3. Chipperfield, J.O.; Steel, D.G. Efficiency of Split Questionnaire Surveys. J. Stat. Plan. Inference 2012, 141, 1925–1933. [Google Scholar] [CrossRef] [Scilit]
  4. Chipperfield, J.O.; Steel, D.G. Design and Estimation for Split Questionnaire Surveys. J. Off. Stat. 2009, 25, 227–244. [Google Scholar]
  5. Schoemann, A.M.; Miller, P.; Pornprasertmanit, S.; Wu, W. Using Monte Carlo simulations to determine power and sample size for planned missing designs. Int. J. Behav. Dev. 2014, 38, 471–479. [Google Scholar] [CrossRef] [Scilit]
  6. Jang, D.G.; Zhu, Z.; Yu, C. A-Optimal Split Questionnaire Designs for Multivariate Continuous Variables. arXiv 2022, arXiv:2203.04883. [Google Scholar] [CrossRef] [Scilit]
  7. Adiguzel, F.; Wedel, M. Split Questionnaire Designs for Massive Surveys. J. Mark. Res. 2008, 45, 608–617. [Google Scholar] [CrossRef] [Scilit]
  8. Chipperfield, J.O.; Barr, M.L.; Steel, D.G. Split Questionnaire Designs: Collecting only the data that you need through MCAR and MAR designs. J. Appl. Stat. 2018, 45, 1465–1475. [Google Scholar] [CrossRef] [Scilit]
  9. Lin, C.; Peng, J.; Qin, Y.; Li, Y.; Yang, Y. Optimal Integrating Learning for Split Questionnaire Design Type Data. J. Comput. Graph. Stat. 2022, 32, 1009–1023. [Google Scholar] [CrossRef] [Scilit]
  10. Raghunathan, T.E.; Grizzle, J.E. A split questionnaire survey design. J. Am. Stat. Assoc. 1995, 90, 54–63. [Google Scholar] [CrossRef]
  11. Axenfeld, J.B.; Blom, A.G.; Bruch, C.; Wolf, C. Split Questionnaire Designs for Online Surveys: The Impact of Module Construction on Imputation Quality. J. Surv. Stat. Methodol. 2022, 10, 1236–1262. [Google Scholar] [CrossRef] [Scilit]
  12. Lang, K.M.; Moore, E.W.G.; Grandfield, E.M. A novel item-allocatioran procedure for the three-form planned missing data design. MethodsX 2020, 7, 100941. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. Wu, W.; Jia, F.; Rhemtulla, M.; Little, T.D. Search for efficient complete and planned missing data designs for analysis of change. Behav. Res. Methods 2016, 48, 1047–1061. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Stuart, M.; Yu, C. A computationally efficient method for selecting a split questionnaire design. Commun. Stat.-Simul. Comput. 2022, 51, 2464–2486. [Google Scholar] [CrossRef] [Scilit]
  15. Lyberg, L.; Biemer, P.; Collins, M.; de Leeuw, E.; Dippo, C.; Schwarz, N.; Trewin, D. Survey Measurement and Process Quality; John Wiley and Sons: Hoboken, NJ, USA, 1997. [Google Scholar]
  16. Rubin, D.B.; Little, R.J.A. Statistical Analysis of Missing Data; John Wiley and Sons: Hoboken, NJ, USA, 1987. [Google Scholar]
  17. Enders, C.K. Applied Missing Data Analysis; Guilford Press: New York, NY, USA, 2022. [Google Scholar]
  18. Silvia, P.J.; Kwapil, T.R.; Walsh, M.A.; Myin-Germeys, I. Planned missing-data designs in experience-sampling research: Monte Carlo simulations of efficient designs for assessing within-person constructs. Behav. Res. Methods 2014, 46, 41–54. [Google Scholar] [CrossRef] [Scilit]
  19. Lawes, M.; Schultze, M.; Eid, M. Making the Most of Your Research Budget: Efficiency of a Three-Method Measurement Design with Planned Missing Data. Assessment 2020, 27, 903–920. [Google Scholar] [CrossRef] [Scilit]
  20. Noble, D.W.A.; Nakagawa, S. Planned missing data designs and methods: Options for strengthening inference, increasing research efficiency and improving animal welfare in ecological and evolutionary research. Evol. Appl. 2021, 14, 1958–1968. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Imbriano, P.M.; Raghunathan, T.E. Three-Form Split Questionnaire Design for Panel Surveys. J. Off. Stat. 2020, 36, 827–854. [Google Scholar] [CrossRef] [Scilit]
  22. Ioannidis, E.; Merkouris, T.; Zhang, L.C.; Karlberg, M.; Petrakos, M.; Reis, F.; Stavropoulos, P. On a Modular Approach to the Design of Integrated Social Surveys. J. Off. Stat. 2016, 32, 259–286. [Google Scholar] [CrossRef] [Scilit]
  23. Rioux., C.; Lewin, A.; Odejimi, O.A.; Little, T.D. Reflection on modern methods: Planned missing data designs for epidemiological research. Int. Stat. Rev. 2020, 49, 1702–1711. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Rhemtulla, M.; Hancock, G.R. Planned Missing Data Designs in Educational Psychology Research. Educ. Psychol. 2016, 51, 305–316. [Google Scholar] [CrossRef] [Scilit]
  25. Gonzalez, J.M.; Eltinge, J.L. Multiple Matrix Sampling: A Review; Office of Survey Methods Research: Washington, DC, USA, 2007. [Google Scholar]
  26. Orchard, T.; Woodbury, M.A. A missing information principle: Theory and applications. In Proceedings of the 6th Berkeley Symposium on Mathematical Statistics and Probability; University of California Press: Berkeley, CA, USA, 1972; Volume 1, pp. 697–715. [Google Scholar]
  27. Agresti, J. An Introduction to Categorical Data Analysis; John Wiley and Sons: Maitland, FL, USA, 1996. [Google Scholar]
  28. Kish, L. Survey Sampling; John Wiley & Sons: Hoboken, NJ, USA, 1965. [Google Scholar]
Figure 1. Monotonic sequential MAR SQD with three categorical variables: ψ is the conditional probability of observing the indicated sequence; F i give feasible cells for the observed sequence; I indicates the variables observed; d o is the values of observed variables; p ( d 0 ) is the probability of d 0 .
Figure 1. Monotonic sequential MAR SQD with three categorical variables: ψ is the conditional probability of observing the indicated sequence; F i give feasible cells for the observed sequence; I indicates the variables observed; d o is the values of observed variables; p ( d 0 ) is the probability of d 0 .
Mathematics 14 00355 g001
Table 1. All possible MCAR data patterns for P = 3 , (1 indicates x p is collected).
Table 1. All possible MCAR data patterns for P = 3 , (1 indicates x p is collected).
(j) I 1 I 2 I 3
1100
2010
3110
4001
5011
6101
7111
Table 2. Example: count data with cell counts indexed by a.
Table 2. Example: count data with cell counts indexed by a.
x 3 = 1 x 3 = 2
x 1 = 1 x 2 = 1 3 ( a = 1 ) 17 ( a = 2 )
x 2 = 2 4 ( a = 3 ) 2 ( a = 4 )
x 1 = 2 x 2 = 1 176 ( a = 5 ) 197 ( a = 6 )
x 2 = 2 293 ( a = 7 ) 23 ( a = 8 )
Table 3. Variable costs scenarios.
Table 3. Variable costs scenarios.
ScenarioCost of Variables Collected
x 1 x 1 , x 2 x 1 , x 3 x 1 , x 2 , x 3
Equal1223
Unequal1245
Disproportionate1226
Table 4. Percentage reduction in A-, A W - and D-optimality for MCAR and sequential MAR PMDs relative to CC design for same total cost, data in Table 2, for fixed and variable cost scenarios given in Table 3.
Table 4. Percentage reduction in A-, A W - and D-optimality for MCAR and sequential MAR PMDs relative to CC design for same total cost, data in Table 2, for fixed and variable cost scenarios given in Table 3.
Fixed CostVariable CostMCAR PMDMAR PMD
( c f ) A A W D A AW D
0Equal0001.02.342
1Equal000000
2Equal000000
0Unequal0008.51392
1Unequal0005.06.259
2Unequal0003.73.144
0Disproportionate3.51.3128.51392
1Disproportionate0.9005.06.259
2Disproportionate0003.73.144
Table 5. Optimal sequential MAR PMD for A-, A W - and D-optimality, same total cost: total sample size, non-core sample proportions (e.g., p ˜ 12 ), conditional probabilities (e.g., ψ 2 | 2 ), for data in Table 2, c f = 0 , disproportionate costs given in Table 3.
Table 5. Optimal sequential MAR PMD for A-, A W - and D-optimality, same total cost: total sample size, non-core sample proportions (e.g., p ˜ 12 ), conditional probabilities (e.g., ψ 2 | 2 ), for data in Table 2, c f = 0 , disproportionate costs given in Table 3.
FixedDesignSampleProportionsDesign Parameters
Cost Criteria Size ( n ) p ˜ 1 , p ˜ 12 , p ˜ 13 , p ˜ 123 ψ 2 | 1 ψ 2 | 2 ψ 3 | 1 ψ 3 | 2 ψ 3 | 11 ψ 3 | 12 ψ 3 | 21 ψ 3 | 22
0A13,5220, 0.38, 0.01, 0.610.61100.610.80.4
1A11,2070, 0.18, 0, 0.820.8100.81110.6
2A11,0400, 0.18, 0.01, 0.810.81100.8110.6
0 A W 19,0610.39, 0.23, 0, 0.3810.600110.60.6
1 A W 12,8240, 0.38, 0, 0.621100110.60.6
2 A W 12,8730, 0.39, 0, 0.611100110.60.6
0D31,4490.77, 0.06, 0, 0.1710.200110.60.8
1D18,9950.38, 0.35, 0, 0.2710.600110.40.4
2D14,6270.19, 0.39, 0, 0.4210.800110.40.6
Note: complete case non-core sample size = 10,000.
Table 6. Ratio of variances of estimates of the parameters: optimal sequential MAR PMD relative to CC design with same total cost, for A-, A W - and D-optimality, data in Table 2, disproportionate costs, c f = 0 .
Table 6. Ratio of variances of estimates of the parameters: optimal sequential MAR PMD relative to CC design with same total cost, for A-, A W - and D-optimality, data in Table 2, disproportionate costs, c f = 0 .
ParameterDesign Criteria
π a A A W D
11.40.540.33
20.840.540.33
30.970.540.33
41.20.540.33
50.871.32.10
60.871.22.1
70.880.931.5
81.71.41.9
Table 7. MAR PMD, MDeff for parameters by optimal design criteria, for data in Table 2, disproportionate costs, c f = 0 .
Table 7. MAR PMD, MDeff for parameters by optimal design criteria, for data in Table 2, disproportionate costs, c f = 0 .
ParameterDesign Criteria
π a A A W D
11.911
21.111
31.311
41.611
51.22.36.4
61.22.36.2
71.21.74.6
82.32.65.5
Table 8. Means: percentage reduction in A-, A W - and D-optimality for MCAR and sequential MAR PMDs relative to CC design for same total cost, data in Table 2, for fixed and variable cost scenarios given in Table 3.
Table 8. Means: percentage reduction in A-, A W - and D-optimality for MCAR and sequential MAR PMDs relative to CC design for same total cost, data in Table 2, for fixed and variable cost scenarios given in Table 3.
Fixed CostVariable CostMCARMAR
( c f ) A AW D A AW D
0Equal0002.33.27.8
1Equal0000.81.72.3
2Equal00000.81.0
0Unequal7.41335111541
1Unequal2.86.8156.49.321
2Unequal1.13.65.44.06.413
0Disproportionate111950152255
1Disproportionate6.010289.81436
2Disproportionate2.86.8156.79.722
Table 9. Reduction in the effective sample size for the possible data observed on a single unit, data in Table 2, for each parameter, ϕ ( a ) , and for overall A- and D-optimality, ϕ .
Table 9. Reduction in the effective sample size for the possible data observed on a single unit, data in Table 2, for each parameter, ϕ ( a ) , and for overall A- and D-optimality, ϕ .
x 1 x 2 x 3 ϕ ( 2 ) ϕ ( 3 ) ϕ ( 4 ) ϕ ( 5 ) ϕ ( 6 ) ϕ ( 7 ) ϕ ( 8 ) ϕ
Observed Values MARA-optimality
1..27.5027.5027.50−0.61−0.63−0.77−0.470.62
11.35.75−0.73−0.72−0.96−1.0−1.2−0.750.18
1.1−0.33102−0.33−0.43−0.45−0.55−0.340.35
12.0408000000.64
1.23.9703400000.27
2..0001.11.10.851.30.97
21.0001.41.2000.72
2.10001.11.10.851.30.97
22.000000.173.00.19
2.200000.3503.70.27
Observed Variables MCAR
1.01.01.01.11.00.791.30.73
1.000.310.650.690.600.041.30.47
0.101.00.900.730.810.552.00.69
0000000
Observed Values MARD-optimality
1..282828000083
11.3800000038
1.1010200000102
12.039790000119
1.24.0034000038
2..0000.770.740.601.03.1
21.0001.00.90001.9
2.10000.80.70.61.03.1
22.000000.22.102.3
2.20000002.93.2
Observed Variables MCAR
1.01.01.00.740.710.570.965.0
1.00.310.660.520.470.070.924.0
0.101.00.890.500.590.401.556.0
00000000
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Steel, D.G.; Chipperfield, J. The Efficiency of Missing at Random Planned Missing Designs. Mathematics 2026, 14, 355. https://doi.org/10.3390/math14020355

AMA Style

Steel DG, Chipperfield J. The Efficiency of Missing at Random Planned Missing Designs. Mathematics. 2026; 14(2):355. https://doi.org/10.3390/math14020355

Chicago/Turabian Style

Steel, David G., and James Chipperfield. 2026. "The Efficiency of Missing at Random Planned Missing Designs" Mathematics 14, no. 2: 355. https://doi.org/10.3390/math14020355

APA Style

Steel, D. G., & Chipperfield, J. (2026). The Efficiency of Missing at Random Planned Missing Designs. Mathematics, 14(2), 355. https://doi.org/10.3390/math14020355

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop