Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
67,598 characters · 11 sections · 29 citation commands
Difference-in-differences Design with Outcomes Missing Not at Random
\paragraph{Keywords} Difference-in-differences, Causal Inference, Missingness, Panel Data, Principal Strata
\setstretch{1.5}
Difference-in-differences (DID) design is a quasi-experimental method in social science widely used to estimate causal effects of a treatment on an outcome variable using repeated observations of units over time. The DID design is particularly useful when the treatment is not randomly assigned, and the researcher is concerned with unobserved time-invariant confounder. Yet, one of the most prevalent problems encountered by researchers working with panel data is missingness. For example, this problem is evident in DID studies that utilize survey data to measure outcome variables, where respondent's non-response is a frequent concern. chiu2023, in their extensive review and replication of articles from three leading political science journals using observational panel data with binary treatments, highlighted this issue with unbalanced panels. They especially emphasized that the missingness pattern is appeared to be nonrandom or extremely prevalent in some studies.
Despite the prevalence of missing data in panel studies, most methodological work presumes balanced panels without missing data. When faced with the methodological challenge of addressing potential bias due to missing data, researchers have often resorted to complete case analysis (listwise deletion) or imputation methods. However, these methods are not without their own limitations. Complete case analysis can lead to biased estimates when the missingness is not completely at random (MCAR) or missing at random (MAR). Imputation methods can also be problematic when the missingness is not at random (MNAR) or the covariates are also susceptible to missingness. Alternatively, when the quantity of interest is a causal estimand, a more principled approach to address missing data is to use nonparametric bounds of causal effects horowitz2000, zhang2003, imai2008, lee2009 or inverse probability weighting using baseline covariates. However, these methods are general remedies that under-utilize the assumptions already imposed on panel structure for causal identification.
This implies that the intersection of causal inference, panel data, and missing values presents a unique challenge in social science studies that remains underexplored. Recently, there have been attempts in disciplines adjacent to social science to address attrition bias by using parallel trends assumptions and the changes-in-changes approach ghanem2022, dukes2022. ghanem2022 extends the changes-in-changes condition so that the distribution of unobserved heterogeniety to be stable across time within treatment-response subpopulations. Using this assumption, they identify the average treatment effect of the treated-respondents and also show that the average treatment effect can be identified with an additional assumption that the distribution is homogeneous across treatment-response subpopulations. dukes2022 consider two alternative strategies, one based on the parallel trends assumption and the other based on the `bespoke instrumental variable' approach, yet their approach is limited to randomized experiments.
In this paper, I provide an alternative approach of addressing missing outcome variable in panel data with DID design for observational studies. My discussion aligns with the recent studies but also differs from them by setting the DID design and parallel trends assumption for causal identification at the center and exploring the assumptions and data structure within the DID design to address the missing data problem. Specifically, I answer the following questions using the DID design and the principal stratification framework: Under which extension of parallel trends assumptions can we justify the complete case analysis? What are the main identification challenges with missing data in the DID design? How can we identify the average treatment effect for treated (ATT) using auxiliary variables from the panel data that offer additional information yet have not been explored?
What must not be overlooked is that missing indicator in this setup is a post-treatment intermediate variable. In this vein, I proceed to introduce principal stratification frangakis2002--namely, always-respondents, if-treated-respondents, if-control-respondents, and never-respondents--to induce parallel trends assumptions conditioning on these latent groups that shares the same missingness pattern. I interpret the complete case analysis under this framework and discuss the main identification challenges with missing data: (1) potential dependence between selection into treatment with principal strata and (2) heterogenous effect across principal strata. I then extend the DID design using auxiliary variables (e.g. outcome variables in multiple pretreatment periods), and use its response indicator as an instrumental variable for identifying ATT. The essential intuition behind this approach is that the response indicator of the auxiliary variables can be used as an instrumental variable that affects the time trend of the outcome variable only through the treatment selection and the missingness of the post-treatment outcome variable. Lastly, I propose alternative approaches with the principal strata specific parallel trends assumption to partially identify the pricipal strata specific ATT. To be specific, I impose parallel trends assumption within each principal stratum and leverage missingness rates over time to estimate the proportions of these groups. Building on this, I tailor Lee bounds lee2009, a well-known nonparametric bounds under selection bias, to partially identify the causal effect within the DID design.
In this section, I review the parallel trends assumptions under which the standard DID design with complete case analysis can be justified. Using the principal stratification framework, I describe the limitations of this standard methodology and discuss the main identification challenges with missing data in the DID design.
To illustrate the problem of missing data in the DID design, I revisit two application studies using two-periods DID design with survey data. In the first application, I revisit sexton2023 which studies how small development aid projects affect the public perception and attitudes toward government. Specifically, the authors are interested in the impacts of German development aid during 2017–18 on subsequent political attitudes in northern Afghanistan. The study uses two waves (2016 and 2018) of geocoded public opinion survey data with five different main outcome variables. As shown in Table (ref), the missingness of the survey responses in the second wave is substantial, and the pattern of missingness appears to be different across outcome variables. Particularly, the missingness ratio differ in time trend (before-after), difference between treated and control groups (treated-control), and its interaction (differece-in-differences).
In the second study, I revisit bisgaard2018 which examines the effects of elite partisan cues on economic perception. The study uses five waves of panel surveys collected from 2010 to 2011 in Denmark to track public opinion on economic issues. After the second wave of survey data, the Center-Right government in Denmark dramatically changed its partisan cue on the severity of the public budget deficit, which led to a change in the economic perception of incumbent supporters according the their findings. The missingness of the survey responses in this study is shown in Table (ref). Given a non-ignorable amount of missingness, the authors provided additional regression analysis with the second and third waves, and concluded the missingness does not appear to be systematically different across treatment groups and time periods, conditioning on prior perceptions of the national economy.
Several questions arise from these examples, particularly related to the unique features of panel data. What does the trend of missingness within each treatment group imply for potential bias in the DID estimates? Can we directly compare such trends between different treatment groups, and if they are similar across groups, would the DID estimate from complete case analysis be unbiased? Answering these questions necessitates exploring diverse variants of canonical parallel trends assumptions and their substantive implications. As I will discuss in the following sections, the short answer is no. It requires an additional assumption regarding the parallel trends of the outcome variable between respondents and nonrespondents within each treatment group, which may restrict the heterogeneity of the treatment effect.
Another important aspect of these questions is how we define the groups in terms of missingness and how the composition of these groups relates to causal identification. In particular, it is crucial to recognize that the missingness of the outcome variable is a post-treatment variable, meaning there exists a latent group of units with different missingness patterns under different treatment statuses. This motivates the need for principal stratification, which examines the potential missingness patterns of the outcome variable under different treatment statuses. Given that we are interested in observational studies where the treatment is not randomly assigned, the composition of these latent groups may vary across treatment groups. This may be due to selection bias, where the treatment selection is correlated with missingness, heterogenous treatment effect across these latent groups, or both.
Lastly, it is worth noting that panel data provides additional information that can be crucial for identifying causal effects. At a minimum, researchers always have access to the missingness rate of the outcome variable over time, which can be used in causal identification. In more favorable cases, researchers may also have access to auxiliary variables from the panel data, such as other outcome measures or multiwave data, which can further aid in identifying the causal effect. In this paper, I propose a novel identification strategy that combines the principal strata framework with these auxiliary variables to address the missing data problem in the DID design.
I consider a two-period DID design with binary treatment and missingness in the outcome variable. Let $D_{i}$ be a binary treatment group indicator of unit $i$, where $D_{i} = 1$ for the treated group and $0$ for the control group. Let $Y_{it}$ be the outcome variable of unit $i$ at time $t = 1,2$, where time $1$ is the pre-treatment period and $2$ is the post-treatment period. I assume the consistency and no anticipation assumptions for the potential outcomes:
Let $R_{i}$ denote a binary response indicator at time $t = 2$. $R_{i} = 1$ if $Y_{i2}$ is observed, and $0$ if it is missing. Here, $R_{i}(d)$ is the potential response indicator if $D_{i} = d$. I assume the same consistency and no anticipation assumptions for $R_{i}$.
One example of this setting is a study using a DID design (or two-way fixed effects) where the data comes from a survey conducted in two waves, with attrition observed in the second wave. Another potential example of missing outcomes is the “don't know” response to survey questions. A common practice in this case is to treat “don't know” as missing, resulting in a similar setup to survey data with attrition. Although I illustrate the problem of missing data in the DID design using survey attrition to facilitate understanding, the proposed methodology can be applied to other types of missing data and is not limited to surveys.
Based on the joint distribution of the response indicator and treatment group indicator we can define the following groups:
One critical limitation of such a group is that it does not account for the fact that the response itself is a post-treatment variable. This group is considered crude because it only captures the realized response under a given treatment assignment and does not consider the counterfactual response (i.e., whether these individuals would have responded if they had been selected into the other treatment group). For example, “treated-respondents” in the first motivating example correspond to those people whose districts received aid and responded to the survey question afterward. We do not know whether this group of people would have also responded to the question if their district had not received such aid.
Alternatively, we can introduce the principal strata framework frangakis2002 in this setup based on the joint distribution of the potential outcomes of the response indicator. Let $S_i = (R_{i}(1), R_{i}(0))$ denote a principal strata define as below. For example, $S_i = (1,1)$ implies that unit $i$'s outcome is always observed, no matter the treatment status.
Note that this can be viewed as a latent pre-treatment covariate, distinct from the observed post-treatment response indicator $R_{i}$. With this, we can consider the joint distribution of this principal stratum and treatment group (e.g. never-respondents who are treated), as opposed to the realized response and treatment group (e.g. treated-respondents, which is a mix of never-respondents and if-treated-respondents within the treated group). We use the term “respondents” for ease of interpretation, but it may not necessarily refer to survey respondents. In general, this term should be understood as a group of units that exhibit different missingness patterns under two possible treatment regimes.
The principal strata framework is widely utilized in the causal inference literature, particularly for examining issues of truncation by death and noncompliance. In the context of truncation by death in medical studies, for example, “always-respondents” corresponds to “survivors,” who would have always survived no matter what the medical treatment (e.g. an uptake of a medicine) was. In the context of noncompliance, “always-respondents” correspond to “always-taker,” who would have always taken the medicine no matter what the treatment assignment, an encouragement to take the medicine, was.
Similar to these issues, the potential outcome of the response indicator in our context is significant because it defines latent subgroups of units with substantively different characteristics. For example, in the first motivating application, if-treated-respondents represent a group of people who respond to these sensitive survey questions only if their district received aid. In contrast, if-control-respondents are those who respond to the survey questions only if their district did not receive the aid. One can imagine that there might be a systematic reason for these opposing behaviors, which could be related to the impact of aid. In the second motivating application, always-respondents are individuals who consistently respond to survey questions about economic issues, regardless of partisan cues. In contrast, never-respondents are those who never respond to the survey questions on economic issues. These two groups may possess distinct characteristics, such as varying levels of political engagement or interest in economic matters, which could correlate with partisanship.
Two main issues with different principal strata are that (1) the proportions of these latent groups may vary between the treated and control groups, and (2) the treatment effect may differ across these latent groups. I will formally demonstrate this intuition in this section, where I discuss the identification challenges within this principal strata framework. To be specific, in this paper, I treat the principal strata as one of the conditioning variables when imposing parallel trends assumptions for causal identification, a topic that will be elaborated upon in this section. For simplicity, I assume there is no missing data in the first wave (alternatively, one could assume MCAR for the pre-treatment outcome missingness for now). However, this assumption will be relaxed in the following section with the proposed solution to the missing data problem.
To begin with, we discuss how the general practice of complete case analysis can be justified under the parallel trends assumptions. In the DID design, researchers are interested in identifying the ATT defined as
We employ the canonical parallel trends assumption to identify the counterfactual outcome $Y_{i2}(0)$ for treated units.
With this assumption and the consistency assumption, we have
The final expression is not identified with missing data in the outcome variable. Upon the missingness in panel data with a randomized treatment, dukes2022 consider imposing the parallel trends assumption between the respondents and nonrespondents, for the time trend under each treatment status respectively.
What this assumption implies is as follows. We first consider the expression under the control group, $d = 0$.
It assumes that the time trend of outcome, $Y_{i2}(0) - Y_{i1}$, is on average parallel between control respondents and control nonrespondents. For example, in the motivating application, within the districts that received aid, the average change in perception over time among respondents is equal to that among nonrespondents. This can be considered a variant of the canonical parallel trends assumption, where we are interested in DID of control-respondent and control-nonrespondents groups. Since the assumption is made on the time trends, not the outcome levels, it allows for the missingness to be correlated with the outcome level. For instance, if respondents who are more positive towards the government are more likely to respond to the survey, yet the change in perception over time remains constant, the assumption may still hold. Additionally, if there are multiple pre-treatment periods without missingness, this assumption can be indirectly tested by examining the parallel trends of the outcome between control-respondent and control-nonrespondents groups.
Next, let's consider (ref) under the treated group, $d = 1$.
This assumption is more restrictive than the former, as it requires the parallel trends of the sum of causal effect and time trend of the outcome between treated-respondents and treated-nonrespondents. For example, within the districts that received the aid, the change in perception pertaining to the aid and a common time shock among the respondents is on average equivalent to that of the nonrespondents. Hypothetically, this can be violated if the missingness is correlated with the treatment effect size, holding the time trend constant. For example, if the respondents who received the aid and became more positive towards the government are more likely to respond to the survey, the assumption may be violated.
Under (ref) and (ref), the ATT can be identified by the DID estimand using complete case analysis. The intuition is as follows: As illustrated in the left panel of Figure (ref), the canonical parallel trends assumption ((ref)) allows us to identify the ATT with the DID estimand. Here, the red vertical line represents the difference in outcome for the treated group ($\mathbb{E}[Y_{i2} - Y_{i1} \mid D_{i} = 1]$), and the blue vertical line represents the difference in outcome for the control group ($\mathbb{E}[Y_{i2} - Y_{i1} \mid D_{i} = 0]$). By taking the difference between these two lines, we can identify the ATT as shown in Equation (ref). Yet, due to the missingness in the outcome variable, these quantities are not directly observed. Under (ref), however, we can identify each of these quantities by the difference in outcome of the treated-respondents and control-respondents, respectively, as shown in the middle and right panels of Figure (ref). For example, two red vertical lines in the middle panel represent the difference in outcome for the treated-respondents ($\mathbb{E}[Y_{i2} - Y_{i1} \mid D_{i} = 1, R_{i} = 1]$) and treated-nonrespondents ($\mathbb{E}[Y_{i2} - Y_{i1} \mid D_{i} = 1, R_{i} = 0]$), respectively, and are identical to the red vertical line in the left panel by (ref). Since the red vertical line at the bottom is observed ($\mathbb{E}[Y_{i2} - Y_{i1} \mid D_{i} = 1, R_{i} = 1]$), we can identify the difference in outcome for the treated group ($\mathbb{E}[Y_{i2} - Y_{i1} \mid D_{i} = 1]$). Similarly, we can identify that of the control group as well. I formally state this identification result in the following proposition.
In simpler terms, if we adopt the canonical parallel trends assumption ((ref)), and additionally assume that there is no selection bias in the missingness of the before-after difference in the outcome variable within each treatment group ((ref)), then the ATT can be identified using DID estimator with complete case analysis. As previously discussed, the parallel trends assumption between treated-respondents and treated-nonrespondents is more restrictive than the canonical parallel trends assumption. This is particularly the case when heterogeneous treatment effects related to missingness are present. In the subsequent section, we will delve further into the main identification challenges associated with missing data in the DID design.
In a general setup, when the parallel trends assumption is less plausible, the researcher may alternatively consider conditional parallel trends assumption where we assume parallel trends for each subgroup defined by pretreatment covariates. The rationale behind this is that the parallel trends assumption will become more plausible when the baseline covariates are balanced between the treated and control groups. Given that the missingness is a post-treatment binary variable, we can consider the parallel trends assumptions conditioning on the principal strata, which can be viewed as a conditional parallel trends assumption, conditioning on a latent subgroup: never-respondents, if-treated-respondents, if-control-respondents, and always-respondents.
This can be understood as a weaker assumption than the canonical parallel trends assumption ((ref)). It is worth pointing out that (ref) does not imply (ref) in general, and vice versa.
Intuitively speaking, even if the time trend of outcome is parallel between treated and control groups within a specific principal strata (e.g. always-respondents), unless the distribution of principal strata is equivalent between treated and control groups, the canonical parallel trends assumption may not hold. For example, in the second motivating application, if the incumbent supporters are more likely to respond to the questions on economic issues regardless of the partisan cues, the proportion of always-respondents may be higher in the treated group than the control group, which may potentially violate the canonical parallel trends assumption.
Now, we discuss and clarify the main identification challenges with outcome MNAR using this principal strata parallel trends assumption.
This decomposition is intuitive in the sense that it follows a natural logic starting from the respondents data and then correcting the potential bias due to the missingness. First, suppose that we naively compute $\mathbb{E}[Y_{i2} - Y_{i1} \mid D_{i} = 1, R_{i} = 1] \Pr(R_{i} = 1 \mid D_{i} = 1)$, which is before-after difference of treated-respondents weighted by the proportion of respondents. Then, we need to adjust for the time trend for the respondents: always-respondents and if-treated-respondents. This corresponds to the second and third lines, using the principal strata parallel trends assumption. Also, we need to incorporate the effect for the nonrespondents: never-respondents and if-control-respondents. This corresponds to the last two lines, which cannot be identified with the principal strata parallel trends assumption since the outcome is missing.
In other words, if we assume that the missingness pattern is independent of selection into treatment (i.e., $S_{i} \!\perp\!\!\!\perp D_{i}$) and that the treatment effect is homogeneous across principal strata within the treated group (i.e., $\mathbb{E}[Y_{i2}(1) - Y_{i2}(0) \mid D_{i} = 1, S_{i} = s]$ is constant for all $s$), then the ATT can be identified using the principal strata parallel trends assumption and complete case analysis as before. While these assumptions may be plausible in some applications, they are generally strong and may not hold. In particular, the assumption of a homogeneous treatment effect across principal strata within the treated group is restrictive and may be violated under missing not at random (MNAR).
In the subsequent section, we will propose an alternative approach to identify the ATT using auxiliary variables from the panel data without imposing the homogeneous treatment effect assumption across principal strata. Note that this can be viewed as MNAR in the missing data literature, where the missingness is dependent on the unobserved outcome variable. Specifically, I first adapt an instrumental variable approach from dukes2022, where the response indicator of the auxiliary variables is used as an instrumental variable for the missingness of the outcome variable and thus point identification of ATT is possible. I also propose a partial identification approach based on lee2009 using pre-treatment missingness trend. This allows us to partially identify the ATT for always-respondents, in a more general setup of the dependence between the missingness and the treatment selection.
In this section, I propose two alternative approaches that do not require the homogeneous treatment effect assumption across principal strata. The first approach is based on the instrumental variable (IV) method, where the key idea is to utilize an IV that is associated with baseline missingness probability but not with the magnitude of the bias. I motivate this IV approach using randomized incentives for participation in the survey and also consider the response indicator of the auxiliary variables. The second approach is based on the partial identification method, where I tailor Lee bounds lee2009 to address the missing data problem in the DID design. Based on the parallel trends assumptions of response rate over time, I show that the ATT for always-respondents can be partially identified without requiring the homogeneous treatment effect assumption across principal strata or the independence assumption between missingness and treatment selection.
We first consider the point identification of the ATT using an IV that captures the baseline probability of response under a canonical parallel trends assumption. This approach is motivated by `bespoke instrument variable (IV)' from dukes2022, in which they introduce a special type of IV that can be leveraged to identify the selection bias due to missingness tchetgen2017, richardson2022. Here, I adapt the basic idea of (1) introducing an IV that satisfies a certain exclusion restriction and (2) assuming that the bias is homogeneous across the subgroups defined by the IV. In dukes2022, the focus was on a general framework for identifying selection bias due to missingness, specifically the difference $\mathbb{E}[Y_{i2} - Y_{i1} \mid R_{i} = 1] - \mathbb{E}[Y_{i2} - Y_{i1} \mid R_{i} = 0]$, under the assumption of a randomized experiment. However, this paper considers observational studies, where we employ parallel trends assumptions to identify the ATT. Our primary focus is on identifying selection bias due to missingness within each treatment group, i.e. $\delta_{d} \equiv \mathbb{E}[Y_{i2} - Y_{i1} \mid D_{i} = d, R_{i} = 1] - \mathbb{E}[Y_{i2} - Y_{i1} \mid D_{i} = d, R_{i} = 0]$ for $d = 0,1$.
To illustrate this IV approach, hypothetically assume that we offer a randomized incentive for survey participation. Let $\widetilde{R}_{i}$ denote the binary indicator of whether unit $i$ received the incentive. We consider the following set of assumptions on this indicator variable.
The first assumption, relevance to missingness applies if $\widetilde{R}_{i}$ is correlated with the missingness of the post-treatment outcome. For example, if the incentive encourages the respondents to participate in the survey, then the missingness of the outcome may be negatively correlated with the receipt of the incentive. Intuitively speaking, if we consider the post-treatment missingness to be a combination of MAR (in terms of baseline response probability) and MNAR mechanisms, $\widetilde{R}_{i}$ helps us control for the MAR component of the missingness. Unlike others, this assumption can be empirically tested.
The second assumption, parallel trends of observed outcome, corresponds to (ref) (\@nameuse{title@a:pt-observed}) with regards to $\widetilde{R}_{i}$ instead of $R_{i}$. As discussed in a previous section, this assumption can be violated if $\widetilde{R}_{i}$ is correlated with the treatment effect size while the time trend remains constant. This suggests that a crucial criterion for $\widetilde{R}_{i}$ as an IV is the expectation of homogeneous effects between groups defined by this indicator. Under the incentive scenario, this assumption is plausible since the incentive is randomized.
The last assumption, bias homogeneity, is the core assumption for using $\widetilde{R}_{i}$ to correct the bias. This assumption holds if the bias resulting from post-treatment missingness is consistent across individuals who received the incentive and those who did not, within each treatment group. This assumption might be violated if there is an unmeasured confounder, denoted as $U_{i}$, that interacts with the outcome and $\widetilde{R}_{i}$. For instance, in an additive outcome model, the presence of an interaction term $U_{i} \times \widetilde{R}_{i}$ suggests that the bias in post-treatment outcomes between two groups defined by $\widetilde{R}_{i}$ could be heterogeneous.
Intuitively speaking, when we consider all these assumptions together, we can deduce that any observed differences between the $\{\widetilde{R}_{i} = 0\}$ and $\{\widetilde{R}_{i} = 1\}$ groups are due to their association with post-treatment missingness, which induces bias. The identification result is formally stated in the following theorem.
So far, I have used a randomized incentive for survey participation as a potential choice of an IV. Alternatively, one might consider using the response indicator of auxiliary variables as an IV. These auxiliary variables can be constructed from other pre-treatment variables, such as those survey questions unrelated to the main outcome of interest. For instance, in the first motivating example, we could use other outcome variables (e.g., “Afghanistan right direction?”) from the 2016 wave as auxiliary variables. However, it is crucial to ensure that the auxiliary variable satisfies the assumptions in (ref). For example, in the initial motivating application, if “local government confidence” is the primary outcome of interest and there is concern that the missingness of the auxiliary variable correlates with the treatment effect size, it would be advisable to choose another outcome variable that is not directly associated with local governance.
In Appendix (ref), I generalize the setup to allow for missingness in pre-treatment periods and discuss the identification of the ATT using the IV approach. In Appendix (ref), I also present a variant of the IV method discussed above, where the assumption of parallel trends in the observed outcome for the auxiliary variable is relaxed. Instead, this variant employs multiple auxiliary variables that exhibit consistent bias between respondents and nonrespondents. This approach offers a more flexible framework, accommodating scenarios where the missingness of the auxiliary variables may correlate with the treatment effect size.
Despite its potential, the proposed IV approach has several limitations. First and foremost, the bias homogeneity assumption is crucial for the identification of the ATT, yet its validity is difficult to verify. tchetgen2017 provides an example of semiparametric shared parameter model that satisfies the bias homogeneity assumption. A future research direction could be to show which types of models in our setup satisfy this assumption. For example, a separable model of the post-treatment missingness in terms of $\widetilde{R}_{i}$ and an unmeasured confounder $U_{i}$ could be a potential candidate \footnote{For instance, consider the following semiparametric model:
}. Alternatively, as mentioned in tchetgen2017, one may consider a hypothesis test of $\mathbb{E}[Y_{i2} - Y_{i1} \mid D_i = d, \widetilde{R}_{i} = 1, R_{i} = 1] - \mathbb{E}[Y_{i2} - Y_{i1} \mid D_i = d, \widetilde{R}_{i} = 1, R_{i} = 0] = 0$ without making such bias homogeneity assumption.
More importantly, we apply the canonical parallel trends assumption ((ref)) in this approach instead of the principal strata parallel trends assumption ((ref)). It's crucial to highlight that, should we choose to adopt assumption (ref) in lieu of (ref), it may be required to introduce an additional assumption such as the selection into treatment is independent of the principal strata (see Remark (ref) for more details). In the subsequent section, I modify this assumption by allowing the selection into treatment to depend on the principal strata, and propose a partial identification approach to estimate the ATT for the subgroup of always-respondents.
In this section, I introduce a partial identification of the ATT specifically for the subgroup of always-respondents, leveraging the trend observed in pre-treatment missingness. The partial identification follows the `trimming bounds' approach of zhang2003 and lee2009, where the bounds are derived by considering the extreme cases based on the observed distribution. Here, we are interested in the following quantity:
ATT-AR is an average treatment effect among always-respondents who are selected into the treatment. As discussed in the previous section, if researchers believe that the missingness is MNAR, but are reluctant to make additional assumptions about effect heterogeneity across principal strata, then the ATT-AR is the only quantity that can be partially identified with minimal assumptions including (ref) (\@nameuse{title@a:pt-principal}). When the primary interest is in the ATT, this approach serves as a valuable starting point to understand the potential biases introduced by missingness and to develop a corresponding sensitivity analysis.
Since we utilize pre-treatment missingness in this approach, we slightly modify the notation to distinguish the missingness indicator at different time points. Specifically, we use $R_{it}$ instead of $R_{i}$ to denote the response indicator of the outcome variable at time $t$. That is, $R_{it} = 1$ if $Y_{it}$ is observed, and $0$ if it is missing, for $t = 1,2$. We also assume that we are interested in the ATT for the population who responded to the survey at time 1, while omitting $R_{i1} = 1$ in the conditioning set for simplicity of notation. Alternatively, we could consider an exclusion restriction type of assumption for pre-treatment missingness and time trends of the outcome, as described in (ref) in (ref).
\paragraph{Trimming bounds.} We first review the main idea of the trimming bounds approach. Let $\pi_{r_{1}, r_{0}}(d) \equiv \Pr(R_{i2}(1) = r_{1}, R_{i2}(0) = r_{0} \mid D_{i} = d)$. For now, suppose that we identified the proportion of principal strata in the treated group. We will discuss more about this in the later part.
Under (ref) (\@nameuse{title@a:pt-principal}), we have
where the right hand side is the DID among always respondents. Our goal here is to bound this quantity using the observed data:
Suppose $\frac{\pi_{11}(1)}{\pi_{11}(1) + \pi_{10}(1)} = p/100$. Following `trimming bounds' approach of zhang2003 and lee2009, we can get the lower bound of $\mathbb{E}[Y_{i2} - Y_{i1} \mid D_{i} = 1, R_{i2} (1) = 1, R_{i2} (0) = 1]$ by `trimming' lower $p\%$ of treated-respondent group and taking average. That is,
where $q^{\text{low}}_{1}(f_{1})$ is the $\frac{\pi_{10}(1)}{\pi_{11}(1) + \pi_{10}(1)}$ quantile of $f_{1}(\cdot)$. Similarly, we can get the upper bound as follows:
where $q^{\text{high}}_{1}(f_{1})$ is the $\frac{\pi_{11}(1)}{\pi_{11}(1) + \pi_{10}(1)}$ quantile of $f_{1}(\cdot)$.
Using the same strategy, we can bound $\mathbb{E}[Y_{i2} - Y_{i1} \mid D_{i} = 0, R_{i2} (1) = 1, R_{i2} (0) = 1]$:
where $f_{0} \equiv \Pr(Y_{i2} - Y_{i1} \le y \mid D_{i} = 0, R_{i2} = 1)$, $q^{\text{low}}_{0}(f_{0})$ is the $\frac{\pi_{01}(0)}{\pi_{11}(0) + \pi_{01}(0)}$ quantile of $f_{0}(\cdot)$, $q^{\text{high}}_{0}(f_{0})$ is the $\frac{\pi_{11}(0)}{\pi_{11}(0) + \pi_{01}(0)}$ quantile of $f_{0}(\cdot)$. Consequently, our remaining task involves identifying the proportion of principal strata within the treated group. This analysis diverges from the standard principal strata analysis, which typically focuses on randomized experiments. Here, we are delving into observational studies and employing a DID design instead of relying on the assumption of ignorability (refer to wang2017). In this paper, we introduce the parallel trends assumption in the missingness of outcomes as a strategy to address the challenges associated with identifying the proportion of principal strata in the treated group.
\paragraph{Principal strata proportion.} In this paper, we consider two approaches for identifying the principal strata proportion within the treated group: one that incorporates a monotonicity assumption and another that does not. We first start with the method that assumes monotonicity alongside parallel trends in the response indicator.
The monotonicity assumption suggests that the treatment effect on the response indicator for any individual is always non-negative. This assumption is reasonable in certain contexts. For example, consider our first motivating study where we are interested in the effect of aid on the public perception toward the government. If this aid positively affects individuals' perceptions, it is plausible to expect a higher response rate in the post-treatment survey under treatment compared to the control. This is based on the premise that individuals with negative perceptions are more hesitant to participate in the survey. Thus, a positive treatment effect on perception would likely increase the likelihood of response, reflecting the monotonicity stated above. With this assumption, we have
Thus, we only need to identify $\Pr(R_{i2} (0) = 1 \mid D_{i} = 1)$ for always respondents in the treated group. To this end, we introduce the following assumption.
With this assumption, we can identify $\Pr(R_{i2} (0) = 1 \mid D_i = 1) = \Pr(R_{i2} = 1 \mid D_{i} = 0) - \Pr(R_{i1} = 1 \mid D_{i} = 0) + \Pr(R_{i1} = 1 \mid D_{i} = 1)$.
In certain scenarios, the monotonicity assumption might not be applicable. Take, for instance, our second motivating example where the treatment effect—altering cues may not uniformly result in an increased response rate in the post-treatment survey. In such instances, one may consider the following assumption.
\paragraph{Partial identification of ATT-AR.} Combining these identification results, we can bound ATT-AR.
As noted earlier, the results in (ref) provide bounds on the ATT-AR, accommodating both the dependence between treatment selection and principal strata, as well as heterogeneous treatment effects across principal strata. This approach is particularly useful when the missingness rate is moderate, raising concerns about potential bias in the ATT while the proportion of always-respondents remains substantial.
In this section, I revisit the study by sexton2023 to illustrate the proposed method. The paper examines the impact of small development aid projects on public perception and attitudes toward the government. The study employs a DID design, aggregating data at the village cluster level to address inferential challenges (“spatial spillovers and incorrect standard errors”). The aim here is to demonstrate the application of the proposed method rather than to replicate the original study in its entirety. Therefore, for the purpose of this illustration, I simplify the data structure and employ a canonical two-period DID estimator.\footnote{The results presented in this section are provided for illustrative purposes only and should not be interpreted as supporting substantive conclusions.}
Table (ref) shows the partial identification of ATT-AR for the five outcome variables considered in the original study. The table presents the lower and upper bounds of ATT-AR, along with the DID estimates. There are two main takeaways here: across the outcome variables, except for “Local Government Confidence,” the proportion of always-respondents is about $0.8$ to $0.9$ in both treatment and control groups. This result is valuable for researchers as ATT-AR provides a useful reference for ATT with the correction of bias, given a large proportion of this latent subgroup. Particularly for “Confidence in President” and “National Government Good Job,” where the missingness ratio was relatively small (see Table (ref)), the proportion of always-respondents is close to $1$. In contrast, for “Local Government Confidence,” the proportion of always-respondents in the control group is about $0.8$, while it is close to $0$ in the treatment group. This is intuitive, given the fact that the missingness rate is about $70\%$ and $65\%$ for the treated group in pre- and post-treatment surveys, respectively. Secondly, the bounds of ATT-AR are comparatively tight, except for “Local Government Confidence,” where the bounds range from $-1$ to $1$, given the large gap between the treated and control groups.
Table (ref) presents the ATT estimates obtained using the IV method outlined in Theorem (ref). These estimates were derived using the self-reported “Employment opportunity” response indicator from the 2016 (pre-treatment) wave as an instrumental variable. The findings indicate that the difference between the bias-corrected DID estimates and the standard DID estimates is proportional to the missingness ratio observed in the post-treatment survey.
Missingness in panel data is a prevalent issue that has not received sufficient attention in the context of DID studies. In this paper, I address the challenges posed by missing data in DID analyses by exploring different variants of the parallel trends assumption and additional sources of information, such as the trend of missingness rates. Employing the principal strata framework, I identify two primary challenges when outcomes are Missing Not At Random (MNAR): (1) the potential dependence of treatment selection on principal strata and (2) heterogeneous effects across principal strata. To overcome these challenges, I propose two alternative strategies with weaker assumptions: (1) partial identification of the ATT for always-respondents and (2) an IV method that employs baseline missingness indicators as instruments.
The paper also suggests several avenues for extending these methods. For example, the nonparametric identification of ATT-AR could be achieved by adopting the assumptions from wang2017. Additionally, integrating the principal strata framework into the IV approach could provide a way to account for the dependency between treatment selection and principal strata. Exploring partial identification through the bracketing relationship among multiple auxiliary variables, as inspired by ye2023, offers another potential extension. Moreover, generalizing the approach to accommodate staggered adoption and multiple treatment groups represents a promising direction for future research. Lastly, addressing missingness in covariates within DID studies is also worth investigating (see Appendix (ref) for a relevant discussion).