1.1. Planned Missing Designs
Consider a survey which collects information on P variables, from a sample, s, of size n. For the variables are . The sample data collected under the standard or single-phase design (SPD) is the complete sample data . The data is complete in the sense that there are no missing values. Suppose some variables are missing for some sample units. Let indicate that the variable is observed for unit i and if is missing for unit i. , indicates which variables are observed for sample unit i and the matrix summarises the missingness for the sample. The observed values are and the observed data includes and so it is . The complete sample data is then where denote the data values that are missing for sample units.
Here we will call a design that plans to only collect according to some process specified at the design stage a Planned Missing Design (PMD). A PMD is defined by the mechanism that gives rise to given s and can also be called a Split Questionnaire Design (SQD).
There are two types of missing data mechanisms that can be planned. They are missing completely at random (MCAR) and missing at random (MAR). A MCAR mechanism assumes that the joint density of and is , which means the pattern of missing data is independent of the observed data. A MCAR PMD is defined by . A MAR PMD allows the pattern of missing data to depend upon the observed data in some way, so . A MAR PMD is more complicated than an MCAR PMD as specifying a distribution for is more involved than specifying a distribution for . The MCAR design is a special case of the MAR design.
An MAR PMD depends on
. At the design stage when the PMD is being developed, no data has been observed. Thus, we need to consider the properties of estimates with respect to the joint density
. For maximum likelihood estimates (MLEs) this leads to considering the expected information function. In
Section 2 we show how a PMD affects the observed and expected information in general and how the loss of information can be determined.
Section 3 considers the common case of the multinomial distribution.
In some applications there may be variables available at the design stage, some of which may be used in the sample design. These variables can be used in a PMD and this is considered in
Section 2.4.
Some data items may be missing due to non-response, but unlike PMDs, the missingness is not designed and which data items are missing may depend on the variables in an unknown way, leading to biased estimates. In this case the data are missing not at random (MNAR) and the non-response is informative. In practice it is often assumed that survey non-response is MAR with
depending on some fully observed data items. However, because which data items are missing is not controlled by the survey organisation, it is possible that
and
are not independent. See [
1] (Chapter 7) and [
2] (Chapter 8) for a discussion of informative non-response. For this paper we will assume there is no data item non-response and focus on missingness due to the PMD. The approach developed here can also cover data item non-response that is MCAR or MAR.
An MCAR PMD is by far the most common form of PMD in the literature. There are a total of
patterns or possible values of
. For example,
Table 1 illustrates all possible patterns for
indexed by
. In an MCAR PMD, sample units would be allocated to a pattern at random. The development of a PMD would involve specifying what proportion or number of units would be allocated to each pattern. An example of an MAR PMD would be one in which
is observed for all units and the probability of which of
and/or
are observed depends on the value of
. This means only patterns
are used, but the probability of which one is used for unit
i depends on its value of
. A multi-phase design (MPD) is a special case of a PMD in which the patterns used follow a monotone pattern. For example, patterns
. A
standard or
SPD survey would collect all variables from all sample units (i.e.,
in
Table 1 in which
).
MAR PMDs are more flexible and can be more efficient than MCAR PMDs as they allow the missingness pattern to depend on the cost and loss of information associated with different categories of key variables. Their potential benefit will be greater when some categories lead to higher costs or are relatively rare. For example, in a health survey the response to a question on general self-assessed health status of poor or fair may lead to more detailed data on health conditions being collected, increasing costs. Thus, having a higher rate of missingness for the additional data items for people of poor or fair health can decrease costs while still allowing the more detailed data items to be collected within the available budget. In a cancer survey those that report never having being diagnosed can have a lower probability of collecting other lifestyle and dietary data. In both situations the data collected can be fully exploited in maximum likelihood estimation.
1.2. Main Issues in the Literature
The typical motivation for PMDs is to manage respondent burden, which allows fewer than P data items to be collected from a sample unit. Next we review the main issues in a PMD and how they are addressed in the literature.
The first issue is to define the target statistic or objective of the survey. PMDs have been used to estimate a wide range of statistics. Examples include population totals ([
3,
4]), the power of a hypothesis test ([
5]), the variance of analytic parameters (e.g., ref. [
6] who use
A- and
D-optimality, i.e., the trace and determinant of the variance matrix of the parameter estimates), or the likelihood itself (see [
7]).
Second, the number of possible patterns,
T, becomes large even for a small number of variables. There are several reasons for limiting the set of allowable patterns. The target parameter itself may impose a natural constraint on allowable patterns. For example, when regression coefficients are the target, then the dependent variable should always be collected (e.g., [
8,
9]). To equalize the reporting burden across respondents, another common constraint is that each allowable pattern has the same number of variables. Moreover, if the interaction between
and
is important, then we may ensure that they are always collected together. Restricting the number of allowable patterns may reduce the complexity of administering a survey using a PMD and of analysing the data once it is collected.
A common way of restricting the number of allowable patterns is to create modules, whereby a variable appears in only one module and all variables in a module are collected together. The creation of modules has been widely studied. Ref. [
10] recommends that a variable in a module should be well correlated with variables in different modules, assuming only one module is collected per sample unit. Ref. [
11] forms modules explicitly for the purpose of minimising imputation error. Ref. [
12] considers simple algorithms to allocate variables to a module so as to create diverse topics or a single topic in a module.
Third, once the set of permissible patterns or values of
have been determined, each must be assigned a probability of occurring in the sample. Finding the optimal
involves optimisation algorithms that typically aim to maximise or minimise the statistical objective. Refs. [
3,
4] use a simple grid search algorithm since
P is small. Ref. [
13] searches for the optimal PMD patterns using Monte Carlo simulation to estimate the standard error of estimates in growth curve models. Ref. [
6] minimises the
A-optimality and Ref. [
14] minimises the Kullback–Leibler Distance when the number of cells in the multinomial distribution is large.
There are limitations of a PMD. A PMD collects less information about interactions than a standard design. This may be acceptable when key interactions are identified at the design stage but may be unacceptable if a survey is to support analysis of a wide range of interactions. Different routings through a questionnaire can sometimes affect response values. This is particularly the case for sensitive questions (for discussion see [
15] chapter 5). A PMD tends to complicate data collection and restrict modes of data collection. Analysis of missing data is typically more complex than analysis of complete data [
16]. A range of methods are available when there are planned or unplanned missing data. These include full information maximum likelihood (FIML), the E-M algorithm, multiple imputation and Bayesian methods. Refs. [
2,
17] describe modern methods in the analysis of missing data.
The range of problems to which PMDs have been applied has increased over time, supported by estimation methods and technology that make their implementation easier. We give a few interesting examples. An experience survey ([
18]) involves a person responding, via a mobile phone, at random points throughout the day, to a pattern of questions about the amount of time spent on a variety of tasks. Ref. [
19] uses a PMD to estimate smoking behaviour, a latent variable measured by a question–answer, a saliva test and a breath test. PMDs are applied in ecology and animal welfare studies ([
20]) where the data are measurements from the environment. Ref. [
21] considers a PMD in a panel survey under various longitudinal correlations structures. Ref. [
22] considers PMDs within the context of an official statistics agency integrating multiple surveys over time. Ref. [
23] reviews PMDs for epidemiological research and [
24] considers PMDs in educational psychology research.
Here we consider a PMD from the perspective of efficient sample design. In general there will be a vector of population parameters of interest and we focus on the efficiency of a design based on A- and D-optimality, for fixed total survey cost. PMDs have two efficiency-based advantages over standard surveys. Firstly, they allow variables with high reporting and collection cost or low importance to be collected from fewer units than variables with relatively low cost and high importance. Secondly, the correlation between variables can be exploited to minimise the information loss due to missing data.
Relatively little work has considered PMDs where the data are MAR. An MAR design has more control over
than an MCAR design. In an MAR design,
can depend upon
, while in an MCAR design it cannot. This is important since
determines the variance of parameter estimates in a Maximum Likelihood (ML) framework. Ref. [
21] briefly mentions an MAR PMD. MAR PMDs involving screening questions or missingness based on a dependent variable have been considered in the literature. Ref. [
25] allows
to depend upon screening questions. Ref. [
8] considers allowing the pattern of missing data to depend upon the value of the dependent variable in a logistic regression model. A feature of MAR PMDs is there flexibility and adaptivity as decisions on which variables to collect depend on the observed values some key variables. This also improves robustness as it can be used to ensure particular combinations of variables are collected from an adequate number of units. For some distributions of
, an MAR and MCAR PMD will be equally efficient. This is the case if
is homoskedastic, as is the case for the multivariate normal, as it implies
and the pattern (not the observed variables) affects the uncertainty in the missing data. We consider the multinomial distribution in detail, which is heteroskedastic in
.
In this paper we develop a general and systematic information-based framework for the design of MCAR and MAR PMDs assuming ML estimation is to be used once the data have been collected. Our approach focusses on the information loss and how this can be determined at the design stage. For many surveys the key outputs are tables of frequencies or proportions, and so the framework is applied to the reasonably general case of a P-way table using the multinomial distribution.
Section 2 considers the general effect of a PMD on information for MLEs.
Section 3 derives the information and loss of information of maximum likelihood estimates under the multinomial distribution for an MCAR and MAR PMD. This is done for the observed information, which is relevant when the data have been collected, and expected information given the PMD, which is relevant at the design stage. A sequential and efficient MAR PMD is described.
Section 4 considers the development of MCAR and MAR designs based on
A- and
D- optimality, and provides an empirical evaluation. It also considers ways of identifying the marginal effect of observing a particular data pattern for a single unit.
Section 5 makes concluding remarks.
Appendix A provides a summary of the notation used in the paper.