Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
96,578 characters · 17 sections · 70 citation commands
Design-Based Uncertainty for Quasi-Experiments
{\singlespacing
}
In the social sciences, researchers often have data on the full population of interest. For example, we may observe aggregate data on all 50 U.S. states or administrative data on all individuals in Denmark. Traditional approaches to statistical inference that view the sample as being drawn from a super-population may be unnatural in such settings ManskiPepper(18). One possible alternative in such settings is a model-based approach wherein the units are viewed as fixed, but one develops a statistical model for the outcome. In practice, however, researchers may have difficulty specifying the outcome formation process AbadieEtAl(22).
The literature on design-based inference addresses these difficulties by conditioning on both the units in the finite population and their potential outcomes, and instead viewing the stochastic assignment of treatment as the sole source of randomness in the data. This provides an alternative approach to inference in settings where the researcher does not wish to model the statistical process governing the sampling or formation of potential outcomes. However, existing work on design-based inference has primarily focused on settings where treatment probabilities are known, as in a randomized experiment neyman_application_1923,imbens_causal_2015, li_general_2017, or where treatments are determined independently of potential outcomes conditional on covariates abadie_sampling-based_2020, AbadieEtAl(22).
In contrast, social scientists often study non-experimental settings in which the assumption of (conditional) random assignment of treatment may be questionable due to concerns about selection into treatment based on unobservable factors. Researchers therefore typically turn to strategies such as difference-in-differences (DID) or instrumental variables (IVs). Researchers often refer to these strategies as “quasi-experimental” or “natural experiments,” because the treatments are determined in part by factors such as delays in court systems that affect the timing of state-level policy changes jackson_effects_2016, fluctuations in local weather patterns MadestamEtAl(13), DeryuginaEtAl(19), or exposure to natural disasters Hornbeck(12)-DustBowl, HornbeckNaidu(14), Deryugina(17), NakamuraEtAl(21)-Moving that might reasonably be viewed as stochastic.
In this paper, we develop a design-based approach to inference for such quasi-experimental settings. In line with design-based approaches developed for experiments, we condition on the units in the finite population and their potential outcomes, thus avoiding the need to model the sampling or formation of the potential outcomes. The stochastic nature of the data arises solely from the realization of the quasi-experimental factors, such as court delays or weather shocks, that determine treatment assignment. While we view these factors as stochastic, we importantly do not assume that they generate treatment assignments mimicking a completely randomized experiment. Rather, we view the realization of the quasi-experimental factors as mimicking an unequal-probability experiment wherein each unit $i$ is assigned to treatment with marginal probability $\pi_i$. For example, while it may be reasonable to view court delays as the realization of a stochastic legal process, some states may have a higher probability of realizing such delays than others, leading to heterogeneous $\pi_i$. Of course, if the $\pi_i$ were known, or estimable as functions of observable characteristics, it would be straightforward to adjust for the unequal assignment probabilities. In practice, researchers may not know the $\pi_i$, and they may suspect that they depend on unobservable factors. They therefore proceed using estimators that do not fully adjust for the $\pi_i$.
Our main results concern the properties of common estimators for quasi-experimental settings under this data-generating process. We provide identifying conditions under which common estimators and their associated confidence intervals are valid for finite-population causal estimands. We characterize the biases and coverage distortions that arise when these conditions are violated, and we demonstrate how researchers can conduct sensitivity analyses if they are concerned about possible violations. Altogether, we provide a framework for analyzing quasi-experimental estimators in settings where researchers do not wish to statistically model the sampling or formation of potential outcomes.
As a building block toward understanding popular quasi-experimental estimators, we analyze the difference-in-means (DIM) estimator, which compares the average outcome for the treated and untreated units, under this data-generating process. This allows us to connect our results with the existing design-based literature, which has often focused on the DIM owing to its popularity in randomized experiments. We later generalize our results for the DIM to study least squares regression adjustment, the instrumental variables estimator, and the difference-in-differences estimator.
We derive design-based analogs to the familiar omitted variable bias formula for the DIM. Its expectation can be decomposed into two terms: a finite-population analog to the average treatment effect on the treated, which we call the expected average treatment effect (EATT), and a bias term that depends on the finite-population covariance between the (unknown) treatment probabilities and the untreated potential outcomes. The DIM is unbiased for the EATT if the treatment probabilities are uncorrelated with the untreated potential outcomes in the finite population. The DIM is further unbiased for the average treatment effect (ATE) if the treatment probabilities are also uncorrelated with the treated potential outcomes.
We next establish that the DIM is approximately normally distributed with a particular variance that depends on the finite-population variances of the potential outcomes and treatment effects. We provide a finite-population central limit theorem and Berry-Esseen bound, which imply that the DIM is approximately normally distributed when the finite population is large. We further show that the usual heteroskedasticity-robust variance estimator is consistent for an upper bound on the variance of the DIM. These results follow from exploiting connections between our assignment process with unequal probabilities and rejective sampling from a finite population hajek_asymptotic_1964. Taken together, these results imply that when the finite population is large, conventional confidence intervals yield valid but potentially conservative inference for the expectation of the DIM (which corresponds with a causal estimand under the identifying conditions described above).
A novel feature of our setting is that when the individual treatment probabilities $\pi_i$ are heterogeneous across units, conventional standard errors can be strictly conservative even under homogeneous treatment effects. This contrasts with the celebrated result from neyman_application_1923 for completely randomized experiments, which states that conventional standard errors are strictly conservative if and only if treatment effects are heterogeneous. As a result, even when the DIM is biased, conventional confidence intervals for the EATT or ATE need not necessarily undercover if the conservativeness of the variance estimator dominates the bias. In practice, it is difficult to know which effect will dominate, as neither the conservativeness of the variance estimator nor the bias can be consistently estimated.
Our results suggest a natural form of sensitivity analysis based on the DIM estimator. Given researcher-specified bounds on the magnitude of selection bias, we show how researchers can construct bounds on and confidence intervals for the EATT or ATE. Researchers can use these bounds to report the “breakdown” value of selection bias that would be needed to overturn particular causal conclusions. The (potentially strict) conservativeness of conventional standard errors discussed above implies that such sensitivity analyses yield a (potentially strictly) conservative lower-bound on the robustness of the conclusions to violations of the identifying conditions.
Our analysis of the DIM estimator immediately applies to the canonical two-period DID estimator CardKrueger(94), bertrand_how_2004, one of the most influential quasi-experimental estimators in the social sciences, which can be viewed as a DIM for a first-differenced outcome. Our results imply that the DID estimator is unbiased for the EATT under a design-based analog to the parallel trends assumption, which imposes that the treatment probabilities are uncorrelated with the trends in untreated potential outcomes in the finite population. Our results also enable researchers to conduct sensitivity analyses for violations of this assumption. Similar to the approach in rambachan_more_2023 from the super-population perspective, we can benchmark reasonable values for the violations of parallel trends using data from pre-treatment periods.
We illustrate our theoretical results in both a Monte Carlo simulation based on real data and an empirical application. In our Monte Carlo simulations, we conduct two-period DID analyses of simulated state-level treatments using aggregated data from Longitudinal Household-Employer Dynamics (LEHD) data from the U.S. Census. Since the aggregated data cover over 95% of all private sector jobs in the United States, the LEHD program writes that “no sampling error measures are applicable” qwi-census. Our simulations therefore analyze uncertainty as arising from the realization of placebo state-level policy changes. We allow the state-level treatment probabilities $\pi_i$ to depend on a state's voting results in the 2016 presidential election. While the placebo law has no treatment effect for any state, the untreated potential outcomes may vary in a way that is related to state-level voting patterns, leading to violations of the design-based parallel trends assumption. We illustrate how varying the strength of the relationship between the treatment probabilities $\pi_i$ and state-level voting patterns affects bias and the coverage of conventional confidence intervals for the EATT. Strengthening the relationship between the $\pi_i$ and state-level voting patterns increases bias but has ambiguous effects on the coverage of conventional confidence intervals, due to its competing effects on bias and the conservativeness of conventional standard errors. Robust confidence intervals that account for the bias have correct coverage for the EATT, but are conservative when the $\pi_i$ differ across units.
We next revisit empirical work studying the causal effect of Medicaid expansions across U.S. states. Due to the Affordable Care Act, all U.S. states could expand Medicaid eligibility in 2014, but not all state governments decided to do so. Researchers have used this variation to measure the causal effect of Medicaid expansion on health insurance coverage ($Y_i$) by reporting two-period DID estimates that compare states that expanded Medicaid ($D_i = 1$) against those that did not ($D_i = 0$) wherry_early_2016, MillerWherry(17). We view the 50 U.S. states and their potential outcomes as fixed, and model each state as having an unknown probability of expanding Medicaid based on the realization of stochastic political factors. For example, Ohio famously expanded Medicaid in 2014 only due to a narrow 4-3 ruling by its Supreme Court; but one can imagine that the political process could have played out differently such that Ohio did not expand Medicaid. Although all states are subject to the vagaries of the political process, some states would require a much rarer realization of the political process in order to adopt Medicaid expansion, leading to potential violations of the design-based parallel trends assumption. We conduct sensitivity analyses based on the two-period DID estimator in which we calculate how much the design-based parallel trends assumption must be violated in order to overturn conclusions about the causal effect of Medicaid expansions on health insurance coverage.
We conclude with several extensions that are useful for empirical applications. First, we extend our framework to settings with clustered treatments where, for example, we observe individual-level data but treatment is determined in an unknown manner at a more aggregate level (e.g., states or counties). The cluster-robust variance estimator is valid but potentially conservative, justifying the popular heuristic to cluster standard errors at the level at which treatment is assigned in quasi-experimental settings. Second, we provide sufficient conditions under which adjusting for differences in baseline covariates can address the bias of the DIM estimator. Finally, we study two popular quasi-experimental estimators: instrumental variables (IV) estimators and multi-period difference-in-differences (DID) estimators. We provide conditions under which their estimands have a causal interpretation and conventional confidence intervals are valid, and we illustrate how researchers can report sensitivity analyses to violations of these assumptions.
Rather than suggesting a new estimator or method for calculating standard errors, our analysis shows that canonical estimators and standard errors can be coherently interpreted from an alternative, design-based perspective. This perspective aligns with the empirical descriptions provided by researchers, in which statistical uncertainty arises from quasi-experimental factors that partially determine treatments. Our framework clarifies the identifying conditions under which conventional estimators and standard errors are valid for finite-population causal estimands, and it further provides simple methods for sensitivity analyses based on standard estimators and inferential tools.
\paragraph{Related work:} We build on the literature on design-based inference, which dates to neyman_application_1923 and fisher_design_1935 and has received substantial attention recently. See, for example, Freedman(2008)-regadj_to_experimental_data, Lin(13), AronowMiddleton(15), li_general_2017,kang_inference_2018, BojinovShephard(19)-ts_experiments, WuDing(21) in statistics, and abadie_sampling-based_2020, Xu_2021, Bojinov_Rambachan_Shephard_2020, roth_efficient_2021, AbadieEtAl(22) in econometrics, among many others. Much existing work on design-based inference has focused primarily on settings where treatment probabilities are known to the researcher, as in completely randomized experiments or more complex experimental designs. By contrast, we analyze a setting in which treatment probabilities are unknown to the researcher and may be related to the potential outcomes.
Our framework is related to the design-based framework in abadie_sampling-based_2020, who in Section 3 of their paper consider a setting where treatment assignments are i.n.i.d., and thus can differ across units. Xu_2021 extends these results to non-linear estimators. However, the causal interpretation of the parameters in abadie_sampling-based_2020 relies on the assumption that treatment probabilities are linear in observable characteristics, whereas we consider estimation and inference for analogs to the ATE or ATT under arbitrary forms of selection. We provide a novel analysis of the factors determining the conservativeness of the variance when there is selection into treatment, and the bias and undercoverage that can result from violations of the selection-on-observables assumption. In the other direction, abadie_sampling-based_2020 study both binary and continuous treatments, whereas we focus on the binary case only. Finally, a technical difference between our framework and that in abadie_sampling-based_2020 is that, as in neyman_application_1923 and much of the statistics literature that followed, we view the number of treated units $N_1$ as fixed, whereas abadie_sampling-based_2020 view $N_1$ as stochastic.
Consider a finite population of $N$ units. Each unit is associated with potential outcomes $Y_i(\cdot) := (Y_i(0), Y_i(1))$ corresponding to their outcomes under control and treatment. Individuals also have fixed observable covariates $W_i$. The observed outcome is $Y_i = D_i Y_i(1) + (1 - D_i) Y_i(0)$, where $D_i \in \{0, 1\}$ denotes the treatment of unit $i$. The collection of potential outcomes is $Y(\cdot) := \{ Y_i(\cdot) \colon i = 1, \hdots, N \}$ and covariates $W := \{W_i \colon i = 1, \hdots, N\}$ are viewed as fixed (or conditioned on).
Treatment is realized for each unit according to $D_i \sim Bernoulli(p_i)$, where $p_i$ is an unknown, individual-specific treatment probability that may be arbitrarily related to the potential outcomes $Y_i(\cdot)$ and covariates $W_i$. Thus, treatment assignment is determined as if we had an experiment with unequal treatment probabilities $p_i$. We analyze the distribution of the treatment vector $D := \left(D_1, \hdots, D_{N}\right)^\prime$ conditional on the number of treated units and the potential outcomes and covariates (see pashley_conditional_2021 for discussion of why it is desirable to condition on $N_1$).
The special case with $p_i = \bar{p}$ for all $i = 1, \hdots, N$ nests the completely randomized experiment in which any treatment assignment vector with $N_1$ treated units is equally likely. We have in mind that the stochastic treatment assignment $D_i$ corresponds to the realization of some quasi-experimental process, such as court delays or weather. However, some units may be more likely to have a realization of this factor that leads them to adopt treatment than others. This is captured by the individual-specific treatment probability $p_i$. To make this more concrete, we consider the following example.
\paragraph{Example: Effects of Medicaid expansions across U.S. states.} As part of the Affordable Care Act, all U.S. states were eligible to expand Medicaid eligibility in 2014, yet not all state governments chose to do so. Researchers use this variation in Medicaid expansions across U.S. states to study its effects on state-level health insurance coverage, health care usage, and various health outcomes ($Y_i$) by comparing states that expanded Medicaid ($D_i = 1$) and those that did not ($D_i = 0$) wherry_early_2016, HuEtAl(18), MillerJohnsonWherry(21). Justifying these analyses from a sampling or model-based perspective requires viewing the 50 U.S. states as being drawn from some hypothetical super-population of states or modeling these outcomes as a random process. By contrast, our framework views the 50 U.S. states ($i = 1, \hdots, 50$) and their potential outcomes $(Y_i(0), Y_i(1))$ as fixed. The randomness in the data comes from the realization of state-level expansion decisions $D_i \sim Bernoulli(p_i)$, which we view as the stochastic realization of a state-level political process. For example, Ohio expanded Medicaid in 2014 due to a narrow 4-3 ruling by its Supreme Court, but one could imagine a different realization of the political process in which Ohio chose not to expand in 2014. Indeed, similar states such as Wisconsin and Pennsylvania did not expand in 2014. While all states are subject to the whims of their Supreme Court justices and other political processes, we expect the probability of these processes resulting in Medicaid expansion to differ across states. This is reflected in the heterogeneous treatment probabilities $p_i$, which we would expect, for example, to be higher in more liberal states. The $p_i$ are likely to be complicated functions of state characteristics, some of which may be unobserved, and thus we treat them as unknown to the researcher. $\blacktriangle$
The treatment assignment process captured in Assumption (ref) is compatible with rich models of selection bias (i.e., “endogeneity” in econometrics heckman_common_1976, Heckman(78) or “non-ignorability” in statistics Rubin(78)), because it allows for the treatment probabilities $p_i$ to be related to the potential outcomes. For example, it allows for treatment to be determined by the threshold-crossing model $D_i = 1\{ g(W_i, Y_i(1), Y_i(0)) - \epsilon_i \geq 0 \}$, where $g(\cdot)$ is an arbitrary function of the potential outcomes and covariates, and $\epsilon_i \sim U([0,1]$) is a uniform individual-level shock. Finally, we emphasize that the interpretation of the treatment probabilities $p_i$ depends on the particular, stochastic determinants of treatment that the researcher has in mind (e.g., court delays or weather); uncertainty is then interpreted relative to that source, holding other determinants of treatment fixed.
\paragraph{Notation:} Let $N_1 := \sum_{i=1}^{N} D_i$ and $N_0 := \sum_{i=1}^{N} (1-D_i)$ denote the number of treated and untreated units, respectively. We refer to the distribution of $D$ given in Assumption (ref) as the “randomization distribution”, and we denote probabilities over the randomization distribution by $\mathbb{P}_{R}\left(\cdot\right) := \mathbb{P}\left( \cdot \,|\, \sum_{i=1}^{N} D_i =N_1, W, Y(\cdot) \right)$. We define expectations $\mathbb{E}_{R}\left[\cdot\right]$ and variances $\mathbb{V}_R\left[\cdot\right]$ analogously.
While treatment $D_i$ is unconditionally assigned to unit $i$ with probability $p_i$, we conduct our analysis conditional on $N_1 = \sum_i D_i$ (see Assumption (ref)). We denote the marginal probability of treatment for unit $i$ after this conditioning by $\pi_i := \mathbb{P}_{R}\left(D_i =1\right)$. (It turns out that when the finite population is large, the results in hajek_asymptotic_1964 imply that the $\pi_i$ are approximately equal to the $p_i$ up to a re-scaling; for our results, however, it will typically be easier to work with the $\pi_i$ directly.)
For non-stochastic weights $w_i$ and a non-stochastic attribute $X_i$, we define $\mathbb{E}_{w}\left[X_i\right] := \frac{1}{\sum_{i=1}^{N} w_i} \sum_{i=1}^{N} w_i X_i$ and $\mathbb{V}\text{ar}_{w}\left[X_i\right] := \frac{1}{\sum_{i=1}^{N} w_i} \sum_{i=1}^{N} w_i \left( X_i - \mathbb{E}_{w}\left[X_i\right] \right)^2$ to be the finite-population weighted expectation and variance, respectively. The finite-population weighted covariance $\mathbb{C}\text{ov}_{w}\left[\cdot, \cdot\right]$ is defined analogously. So, for example, $\mathbb{E}_{1}\left[Y_i(0)\right] = \frac{1}{N} \sum_{i=1}^{N} Y_i(0)$ is the equal-weighted average of the untreated potential outcome across the $N$ units in the finite population.
If the marginal treatment probabilities $\pi_i$ were known to the researcher, it would be straightforward to obtain an unbiased estimate of the average treatment effect using the Horvitz-Thompson estimator, $\frac{1}{N}\sum_i (\frac{D_i}{\pi_i} - \frac{1-D_i}{1-\pi_i}) Y_i$. In practice, however, the treatment probabilities $\pi_i$ are unknown, and may not be consistently estimable if the $\pi_i$ are functions of unobservables. Thus, in practice, researchers will typically estimate a treatment effect using other approaches such as DID or IV that do not explicitly adjust for the differences in treatment probabilities across units.
As a stepping stone, we study the difference in means (DIM) estimator
that compares the average outcome for treatment and control units. We derive its expectation and distribution under Assumption (ref), and show how one can conduct sensitivity analyses that account for bias from non-random assignment. In Section (ref), we show that these results apply immediately to the DID estimator, which can be viewed as a DIM for a first-differenced outcome. For simplicity, we abstract away from observable covariates in this section; see Section (ref) for an extension to covariate-adjusted estimators. We consider extensions to IV in Section (ref).
We first analyze the expectation of the DIM over the randomization distribution, characterizing its bias for the finite-population average treatment effect and average treatment effect on the treated.
Proposition (ref) decomposes the expectation of the DIM in two ways. First, it equals the finite-population average treatment effect ($\tau_{ATE}$) plus a bias term that depends on the finite-population covariances between the individual treatment probabilities $\pi_i$ and potential outcomes. Second, it can also be written in terms of a finite-population average treatment effect on the treated, $\tau_{EATT}$, which we refer to as the expected ATT (EATT). The EATT is the expected value (over the randomization distribution) of what imbens_nonparametric_2004 and sekhon_inference_2020 refer to as the “sample average treatment effect on the treated” (SATT). Equivalently, it is a convex weighted average of the treatment effects $\tau_i$, with weights proportional to the individual treatment probabilities $\pi_i$.
Proposition (ref) implies that the DIM is unbiased for the EATT if the finite-population covariance between individual treatment probabilities $\pi_i$ and the untreated potential outcomes $Y_i(0)$ is equal to zero, i.e. $\sum_{i=1}^{N} (\pi_i - \frac{N_1}{N}) Y_i(0) = 0$. This is satisfied in a completely randomized experiment with $\pi_i \equiv \frac{N_1}{N}$. It can also be satisfied if the individual treatment probabilities vary across units but in a way that is not systematically related to the untreated potential outcomes on average in the finite population. Proposition (ref) analogously implies the DIM is unbiased for the finite-population ATE if the finite-population covariance between $\pi_i$ and both potential outcomes is zero.
Since our framework views the potential outcomes as fixed (or conditioned on), we note that both $\tau_{ATE}$ and $\tau_{EATT}$ are functions of the fixed potential outcomes for the $N$ units in the population. Such parameters may be easier to interpret than a super-population ATE or ATT in settings where it is difficult to conceptualize sampling from a super-population or the DGP generating the potential outcomes. On the other hand, in many cases researchers may be interested in what the effect of the treatment would be if it were applied in a new, different context, and it may not be entirely obvious how to extrapolate from $\tau_{ATE}$ or $\tau_{EATT}$ to the new setting. As argued in reichardt_justifying_1999, however, it is also not entirely clear that imagining the $N$ units as having been drawn from a hypothetical super-population helps with extrapolation to different contexts. We thus view $\tau_{ATE}$ and $\tau_{EATT}$ as coherent, internally-valid estimands, while cautioning that they may not be externally valid when extrapolated to new settings.
We next analyze the behavior of $\hat{\tau}$ over the randomization distribution. We shown that when the finite population is large, $\hat{\tau}$ is approximately normally distributed with a particular variance and the heteroskedasticity-robust variance estimator is a conservative estimator for this variance.
Existing results on the distribution of the DIM in randomized experiments Freedman(2008)-regadj_to_experimental_data, Lin(13), li_general_2017 exploit the fact that random treatment assignment is closely-connected to simple random sampling from a finite-population Cochran(77). Because in our setting treatment probabilities $\pi_i$ differ across units, the DIM estimator no longer corresponds to a sample mean under simple random sampling. A key observation for deriving our results, however, is that the DIM is analogous to a Horwitz-Thompson estimator under what is referred to as rejective sampling. We can rewrite the DIM as $\hat{\tau} = \sum_{i=1}^{N} \frac{D_i}{\pi_i} (\pi_i \tilde{Y}_i) - \frac{1}{N_0} \sum_{i=1}^{N} Y_i(0)$, where $\tilde{Y}_i := \frac{1}{N_1} Y_i(1) + \frac{1}{N_0} Y_i(0)$. (We can accommodate the case where $\pi_i = 0$ for some $i$, if $\frac{D_i}{\pi_i}$ is defined to be 0 whenever $\pi_i = 0$.) The second term, $\frac{1}{N_0} \sum_{i=1}^{N} Y_i(0)$, is non-stochastic, and therefore does not affect the variance or higher-order moments of the distribution of $\hat\tau$. The first term, $\sum_{i=1}^{N} \frac{D_i}{\pi_i} (\pi_i \tilde{Y}_i)$, is a Horvitz-Thompson estimator for $\sum_{i=1}^{N} (\pi_i\tilde{Y}_i)$ under rejective sampling, which was first studied by hajek_asymptotic_1964. Our results on the distribution of $\hat\tau$ below are then obtained by applying results on rejective sampling from hajek_asymptotic_1964 and others, and then translating these results back into conclusions about the underlying potential outcomes and causal effects (which, in many cases, are non-trivial).
The exact variance of $\hat\tau$ depends on the second-order treatment probabilities, $\mathbb{P}_{R}\left(D_i = 1, D_j =1\right)$, which in general are complicated functions of $(p_1, \hdots, p_N)$. Fortunately, a simple approximation to the variance is available which becomes accurate when $\sum_{i=1}^{N} \mathbb{V}_R\left[D_i\right] =\sum_{i=1}^{N} \pi_i (1-\pi_i)$ is large.
Lemma (ref) shows that the variance of $\hat{\tau}$ depends on the weighted finite-population variances of the potential outcomes and the treatment effects, where unit $i$ is weighted proportionally to the variance of their treatment status, $\mathbb{V}_R\left[D_i\right] = \pi_i (1-\pi_i)$. The leading constant term $C$ is less than or equal to one, with equality when $\pi_i$ is constant across units. In the special case of a completely randomized experiment, the right-hand side of ((ref)) reduces to $\left( \frac{1}{N_1} \mathbb{V}\text{ar}_{1}\left[Y_i(1)\right] + \frac{1}{N_0}\mathbb{V}\text{ar}_{1}\left[Y_i(0)\right] - \frac{1}{N} \mathbb{V}\text{ar}_{1}\left[\tau_i\right] \right) $, matching neyman_application_1923's celebrated formula for completely randomized experiments up to a degrees-of-freedom correction.
We further provide an approximate expression for the expectation of the heteroskedasticity-robust variance estimator $\hat{s}^2$ White(80). Define $\hat{s}^2 = \frac{1}{N_1} \hat{s}_1^2 + \frac{1}{N_0} \hat{s}_0^2$, where $\hat{s}_1^2 := \frac{1}{N_1} \sum_i D_i (Y_i - \bar{Y}_1)^2$ and $\hat{s}_0^2 := \frac{1}{N_0} \sum_i (1-D_i) (Y_i - \bar{Y}_0)^2$ for $\bar{Y}_1 := \frac{1}{N_1}\sum_i D_i Y_i$, $\bar{Y}_0 := \frac{1}{N_0}\sum_i (1-D_i) Y_i.$
By combining the previous two lemmas, our next result shows the heteroskedasticity-robust variance estimator $\hat{s}^2$ is (weakly) conservative for the true variance of $\hat{\tau}$ over the randomization distribution, up to the approximation errors described above.
In a completely randomized experiment, ((ref)) is satisfied if and only if treatment effects are constant, and thus Proposition (ref) nests the well-known result from neyman_application_1923 that in a completely randomized experiment, the usual variance estimator is weakly conservative and is strictly conservative if and only if there are heterogeneous treatment effects (i.e., $\mathbb{V}\text{ar}_{1}\left[\tau_i\right] >0$). Interestingly, Proposition (ref) implies that even when there are constant effects, $\hat{s}^2$ will generally be strictly conservative whenever the marginal treatment probabilities $\pi_i$ differ across units, except in knife-edge cases.
Corollary (ref) establishes that when treatment effects are constant and $\hat\tau$ is unbiased, the heteroskedasticity-robust variance estimator is non-conservative if and only if the treatment probabilities $\pi_i$ are equal (as in an experiment) for all units $i$ with $Y_i(0) \neq \mathbb{E}_{\pi}\left[Y_i(0)\right]$. More generally, (ref) shows that under constant effects (but without unbiasedness of $\hat\tau$), the variance estimator will be strictly conservative unless the odds ratio $\pi_i / (1-\pi_i)$ is exactly proportional to a factor depending on the inverse of $Y_i(0) - \mathbb{E}_{\pi}\left[Y_i(0)\right]$ for all $i$.
To develop one intuition, note that if $\pi_i$ converges to either zero or one, then $\mathbb{V}_R\left[D_i\right] = \pi_i (1-\pi_i)$ converges to zero. Thus, when all individual treatment probabilities are close to either zero or one, the variance of $\hat{\tau}$ over the randomization distribution is small. It is less obvious that when treatment effects are constant and $\hat{\tau}$ is unbiased, the variance of $\hat{\tau}$ is in fact maximized when all treatment probabilities are equal (as in a randomized experiment). Notice, however, that the sum of the variances of the treatments, $\sum_{i} \pi_i (1 - \pi_i)$, is maximized when $\pi_i = N_1/N$ for all $i$, by Jensen's inequality. The proofs of Proposition (ref) and Corollary (ref) establish that this is sufficient for the variance of $\hat{\tau}$ to be maximized under equal treatment probabilities. In Appendix (ref), we discuss how the conservativeness of the usual variance estimator is intuitively related to, but distinct from, the well-known fact that a conditional variance must on average be less than an unconditional one by the law of total variance.
The proof of Proposition (ref) also suggests that the conservativeness of $\hat{s}^2$ will tend to be larger when there is more heterogeneity in $\pi_i$. For example, under the setting in Corollary (ref) when $b=0$, $\mathbb{E}_{R}^{approx}\left[\hat{s}^2\right] - \mathbb{V}_R^{approx}\left[\hat{\tau}\right]$ is bounded below by a term proportional to $\mathbb{V}\text{ar}_{1}[ (\pi_i - \frac{N_1}{N}) \cdot \allowbreak (Y_i(0) - \mathbb{E}_{\pi}\left[Y_i(0)\right])]$. Thus, $\hat{s}^2$ will tend to be quite conservative when the heterogeneity in $\pi_i$ is large, especially if $\pi_i-\frac{N_1}{N}$ is large for units with extreme values of $Y_i(0)$. The fact that conventional variance estimates tend to become more conservative when the $\pi_i$ are more heterogeneous has important implications for the coverage of conventional confidence intervals, as we formalize next and explore in Monte Carlo simulations below.
So far we established that the heteroskedasticity-robust variance estimator is conservative in the sense that its expectation is weakly larger than the true variance of $\hat{\tau}$. This suggests standard confidence intervals based on $\hat{s}$ will be conservative for $\mathbb{E}_{R}\left[\hat\tau\right]$ if (i) $\hat\tau$ is approximately normally distributed, and (ii) $\hat{s}^2$ is close to its expectation with high probability. We formalize this argument by considering sequences of finite populations indexed by $m$ of size $N_m$, with $N_{1,m}$ treated units, potential outcomes $\{ Y_{i,m}(\cdot) \,:\, i = 1,...,N_{m} \}$, and assignment probabilities $\pi_{1,m},...,\pi_{N_m,m}$. For brevity, we leave the subscript $m$ implicit; all limits are implicitly taken as $m\rightarrow\infty$. We provide a central limit theorem (CLT) and variance consistency result under the following mild regularity conditions on the sequence of finite populations.
Recall $\pi_i (1-\pi_i)$ is the variance of the Bernoulli random variable $D_i$, so Assumption (ref)(ref) implies that the sum of the variances of the $D_i$ grows large. It also implies that both $N_1$ and $N_0$ go to infinity, since $\sum_{i=1}^N \pi_i (1-\pi_i) \leq \min\{\sum_i \pi_i, \sum_i (1-\pi_i)\} = \min\{N_1, N_0 \}$. Assumption (ref)(ref) is similar to the condition for the Lindeberg central limit theorem, and imposes that the weighted finite-population variance of $\tilde{Y}_i$ is not dominated by a small number of observations. Assumption (ref)(ref) bounds the influence that any single observation has on the $\pi$- and ($1-\pi$)-weighted variances of the potential outcomes. Under the conditions introduced above, we have the following finite-population central limit theorem and consistency result for the heteroskedasticity-robust variance estimator.
These results allow us to formalize the conditions under which conventional confidence intervals of the form $\hat\tau \pm z_{1-\alpha/2} \cdot \hat{s}$ will be valid for $\tau_{EATT}$ (or $\tau_{ATE}$) when the finite population is large, where $z_{1-\alpha/2}$ is the $1-\alpha/2$ quantile of the standard normal distribution.
Condition (i) of Proposition (ref) imposes that the sequence of finite populations is such that the bias of $\hat\tau$ is of the same order of magnitude as its standard deviation over the randomization distribution (i.e., local to zero). Condition (ii) of the proposition imposes that the conservativeness of the typical variance estimator stabilizes asymptotically (recall $\mathbb{E}_{R}^{approx}\left[\hat{s}^2\right] \geq \mathbb{V}_R^{approx}\left[\hat{\tau}\right]$ by Proposition (ref)).
When $\hat\tau$ is unbiased, so that $b^* = 0$, Proposition (ref) shows that confidence intervals based on the normal approximation will have correct but generally conservative coverage. Interestingly, it also implies that conventional confidence intervals will maintain correct coverage provided the bias of $\hat\tau$ is sufficiently small relative to the conservativeness of the variance estimator. For example, a sufficient condition to ensure at least 95% coverage is that $|b^*| \leq z_{0.975} \cdot \left(\frac{1}{r}-1\right)$. Conventional confidence intervals can therefore accommodate some bias owing to the fact that heterogeneity in treatment probabilities $\pi_i$ or treatment effects $\tau_i$ typically induces conservativeness of the heteroskedasticity-robust variance estimator. In practice which effect dominates will be difficult to gauge, as neither the bias of the estimator nor the conservativeness of the variance are consistently estimable. Nevertheless, this conservativeness has implications for the interpretation of sensitivity analyses that account for the bias, as we discuss in the following section.
In Appendix (ref), we provide Berry-Esseen type bounds on the approximation quality of the CLT in any finite population of fixed size, applying a result by berger_rate_1998 for rejective sampling. This result establishes that the distribution of $\hat\tau$ will be approximately normally distributed in sufficiently large finite populations without appealing to a sequence of finite populations of increasing size.
Our framework lends itself to sensitivity analyses based on the DIM. While unobserved selection cannot be estimated from the data itself, researchers may place assumptions on the magnitude of selection bias (specifically the finite-population covariance between treatment probabilities $\pi_i$ and potential outcomes). Under such assumptions, identified sets for the EATT and the ATE can be obtained and researchers can conduct valid yet conservative inference on the now partially identified, finite-population causal estimands. Concretely, suppose we assume $\mathbb{C}\text{ov}_{1}\left[\pi_i, Y_i(0)\right]$ lies in the interval $[\underline{b}, \overline{b}]$. Proposition (ref) implies that $\tau_{EATT}$ lies in the interval $[\tau_{EATT}^{lb}, \tau_{EATT}^{ub}]$, where
Natural estimators plug-in the DIM $\hat{\tau}$ for $\mathbb{E}_{R}\left[\hat{\tau}\right]$ in ((ref)), yielding unbiased estimates for the bounds $\hat{\tau}_{EATT}^{lb} = \hat{\tau} - \frac{N}{N_0} \frac{N}{N_1} \overline{b}$ and $\hat{\tau}_{EATT}^{ub} = \hat{\tau} -\frac{N}{N_0} \frac{N}{N_1} \underline{b}$. Bounds on the finite-population ATE could be obtained analogously if the researcher also places bounds on $\mathbb{C}\text{ov}_{1}\left[\pi_i, Y_i(1)\right]$.
By combining our analysis in Section (ref) with existing results from the partial identification literature in econometrics, we can obtain valid yet typically conservative confidence intervals for the partially identified EATT. In particular, in a super-population setting, ImbensManski(04) construct valid confidence intervals for partially identified parameters. Letting $\Delta = \frac{N}{N_0} \frac{N}{N_1} (\overline{b}-\underline{b})$ be the length of the identified set, the Imbens-Manski confidence for the EATT takes the form $[\hat\tau^{lb}_{EATT}- C \hat{s}, \hat\tau^{ub}_{EATT} + C \hat{s}]$, where the constant $C$ is chosen to solve $\Phi\left( \frac{\Delta}{\hat{s}} + C \right) - \Phi(-C) = 1-\alpha$, for $\Phi(\cdot)$ the standard normal cumulative distribution function. In Appendix (ref), we show that this confidence interval has correct but potentially conservative coverage in our design-based framework.
Altogether, our results imply that researchers can report design-based sensitivity analyses directly based on the DIM and assumptions on the magnitude of selection bias. As we illustrate below, a natural statistic to report is the “breakdown” value of selection bias needed to overturn their causal conclusions---for example, how large must $\left| \mathbb{C}\text{ov}_{1}\left[\pi_i, Y_i(0)\right] \right|$ be in order for the Imbens-Manski interval to contain a null effect. \Copy{conservativeness-breakdown}{Since conventional standard errors are conservative for the standard deviation of $\hat\tau$ over the randomization distribution (see Proposition (ref)), such a sensitivity analysis will be conservative about the robustness of causal conclusions. In particular, we show in Appendix (ref) that $\liminf_{N \to \infty} P_R(\hat{b}^* \leq b^*) \geq 1-\alpha$, where $\hat{b}^*$ is the estimated breakdown value using the Imbens-Manski interval and $b^*$ is the true breakdown value for which the identified set includes zero.}
\Copy{exploit-conservativeness}{An interesting question is whether the conservativeness of the typical variance estimator could be exploited to produce less conservative sensitivity analyses. In general, the conservativess of the variance estimator is not consistently estimable since it is a function of the unknown $\pi_i$ (see Proposition (ref)). However, with some auxiliary assumptions one could potentially obtain a lower bound on the conservativeness of the variance. This strikes us an interesting avenue for future work.}
Our analysis immediately applies to the classic two-period difference-in-differences estimator CardKrueger(94), bertrand_how_2004, one of the most influential quasi-experimental estimators. (We show in Appendix (ref) that our discussion extends directly to non-staggered, difference-in-differences estimators with multiple time periods.) Suppose we observe aggregate outcomes $(Y_{it})$ of U.S. states over two periods $t \in \{1,2\}$. Some states ($D_i = 1$) are treated beginning in period 2, whereas other states ($D_i=0$) are untreated in both periods. The observed outcome for state $i$ in period $t$ is $Y_{it} = D_i Y_{it}(1) + (1-D_i)Y_{it}(0)$. In this setting, the DIM estimator for the first-differenced outcome $Y_i := Y_{i2} - Y_{i1}$ is equivalent to the DID estimator between treated and control states, $\hat\tau_{DID} = \frac{1}{N_1} \sum_{i:D_i=1} (Y_{i2} - Y_{i1}) - \frac{1}{N_0} \sum_{i:D_i=0} (Y_{i2} - Y_{i1})$. Under the “no-anticipation” assumption that $Y_{i1}(0) = Y_{i1}(1)$, Proposition (ref) implies
where $\tau_{i2} = Y_{i2}(1) - Y_{i2}(0)$ is unit $i$'s treatment effect in period 2. The first term is the EATT in period 2. The second term is proportional to the finite-population covariance between individual treatment probabilities $\pi_i$ and trends in the untreated potential outcomes. Thus, in our framework, the DID estimator is unbiased for $\tau_{EATT,2}$ provided the treatment probabilities $\pi_i$ are uncorrelated in the finite-population with changes in potential outcomes $Y_{i2}(0) - Y_{i1}(0)$. This is a finite-population parallel trends assumption since it is equivalent to the condition $\mathbb{E}_{R}\left[\frac{1}{N_1} \sum_i D_i (Y_{i2}(0) - Y_{i1}(0))\right] = \mathbb{E}_{R}\left[\frac{1}{N_0} \sum_i (1-D_i) (Y_{i2}(0) - Y_{i1}(0))\right]$.
Furthermore, in this setting, the variance estimator $\hat{s}^2$ is equivalent to the cluster-robust (at the unit level) variance estimator for $\hat{\tau}_{DID}$ from the panel OLS regression $Y_{it} = \alpha_i + \lambda_t + D_i \cdot 1[t=2] \tau_{DID} + \epsilon_{it}.$ Therefore, Proposition (ref) implies that the cluster-robust variance estimator for $\hat{\tau}_{DID}$ is weakly conservative for the variance of the DID estimator over the randomization distribution, and will typically be strictly conservative if treatment probabilities differ across units. As a consequence, provided the finite-population parallel trends assumption holds, conventional confidence intervals of the form $\hat{\tau}_{DID} \pm z_{1-\alpha} \cdot \hat{s}$ will be valid (but typically conservative) in our framework. Since empirical researchers are often unsure about the validity of the parallel trends assumption in practice, it will often be useful to conduct sensitivity analyses on conclusions about the EATT under possible violations of the finite-population parallel trends assumption using the approach described in Section (ref) above. We provide an empirical example of this approach in Section (ref) below.
We conduct Monte Carlo simulations based on the Quarterly Workforce Indicators (QWI) from the Longitudinal Household-Employer Dynamics (LEHD) Program at the U.S. Census qwi-census, which provides aggregate statistics from linked employer-employee microdata covering over 95% of all private sector jobs in the United States. The LEHD program writes, “Because the estimates are not derived from a probability-based sample, no sampling error measures are applicable” qwi-census. Our simulations therefore view uncertainty as arising from the stochastic realization of state-level policy changes.
\paragraph{Simulation design:} We use aggregate data on the 50 U.S. states and Washington D.C. from the QWI (indexed by $i=1,...,N$) for the first quarter of 2012 and 2016 (indexed by $t=1,2$). For each state and year, we set the potential outcomes $Y_{it}(1)$ and $Y_{it}(0)$ equal to the state's observed outcome in the QWI ($Y_{it}$). Mimicking a two-period DID analysis, we simulate treatment by randomly generating placebo laws across states. Our simulated treatments have no causal effect for any state, and so $\tau_{EATT,2} = \tau_{ATE,2} = 0$. The potential outcomes are held fixed throughout our simulations; the simulation draws differ in that each corresponds with a different realization of the generated placebo laws $D=(D_1,...,D_N)'$.
Following Assumption (ref), we draw $D_1, \hdots, D_{N}$ as independent Bernoulli random variables with (unconditional) state-level treatment probabilities $p_i$, discarding any draws where $\sum_i D_i \neq N_1$. Based on state-level results from the 2016 presidential election MITElectionData, the state-level unconditional treatment probabilities $p_i$ are chosen such that, for some $p^1 \in [0,1]$, states that voted for Clinton have $p_i = p^1$, and states that voted for Trump have $p_i = 1 - p^1$. When $p^1 =0.5$, all states have the same probability of adopting treatment, as in a completely randomized experiment, whereas when $p^1 > 0.5$, Democratic states are more likely to adopt the treatment. We report results as $p^1$ varies over $p^1 \in \{0.50, 0.75, 0.90\}$ and fix the number of treated and untreated states at $N_1=25$ and $N_0=26$, respectively.
For each draw of the assignment vector, we calculate the two-period DID estimator $\hat{\tau}_{DID}$ and a nominal 95% confidence interval $\hat{\tau}_{DID} \pm z_{0.975} \cdot \hat{s}$, where $\hat{s}$ is the heteroskedasticity-robust standard error for the first-differenced outcome. We also calculate a nominal ImbensManski(04) 95% confidence interval for the partially identified EATT under the assumption that $|\mathbb{C}\text{ov}_{1}\left[\pi_i, Y_{i2}(0)-Y_{i1}(0)\right]| \leq \tilde{b}$, as discussed in Section (ref). We choose the bound $\tilde{b}$ corresponding to the actual bias of the estimator, $\tilde{b} = |\mathbb{C}\text{ov}_{1}\left[\pi_i, Y_{i2}(0)-Y_{i1}(0)\right]|$, to evaluate the properties of a robust confidence interval that properly accounts for the bias. We report results for two choices of the outcome $Y_{it}$: the log employment level and the log of state-level average monthly earnings for state $i$ in period $t$.
\paragraph{Simulation results:} We first report the bias of the two-period DID estimator. While the placebo law has no treatment effect for any state, the change in untreated potential outcomes $Y_{i2}(0) - Y_{i1}(0)$ varies across states in a way that is related to state-level voting patterns in the 2016 presidential election. As a result, the design-based parallel trends assumption, $\mathbb{C}\text{ov}_{1}\left[\pi_i, Y_{i2}(0) - Y_{i1}(0)\right] = 0$, is violated when $p^1 \neq 0.5$, and hence the DID estimator is biased for the EATT over the randomization distribution in these simulations. The first row of Table (ref) reports the normalized bias of the DID estimator (i.e., $\mathbb{E}_{R}\left[\hat{\tau}_{DID}\right]/\sqrt{\mathbb{V}\text{ar}_{R}\left[\hat{\tau}_{DID}\right]}$) as $p^1$ varies for both of these two outcomes. For $p^1 = 0.5$, the bias is zero up to simulation error. The magnitude of the bias increases as we increase $p^1$, since the average value of $Y_{i2}(0) - Y_{i1}(0)$ differs between Democratic and Republican states for both of our outcomes. Appendix Figure (ref) plots the distribution of the DID estimator over the randomization distribution. The distributions are approximately normally distributed, illustrating the finite-population CLT from Section (ref).
The conservativeness of the usual heteroskedasticity-robust variance estimator is summarized in the second row of Table (ref), which shows the ratio of the average estimated variance for $\hat\tau$ to the actual variance of the estimator, $\frac{\mathbb{E}_{R}\left[\hat{s}^2\right]}{\mathbb{V}\text{ar}_{R}\left[\hat{\tau}\right]}$. In line with the results in Proposition (ref) and Corollary (ref), $\hat{s}^2$ becomes conservative when there is variation in the treatment probabilities. For simulations with $p^1 = 0.5$, $\hat{s}^2$ is, on average, approximately equal to the true variance of the DID estimator. As $p^1$ increases, however, it becomes more conservative: in the most extreme case when $p^1 = 0.9$, the average estimated variance is approximately 2.5 times as large as the true variance. Since there is no treatment effect heterogeneity, this conservativeness is the result of heterogeneity in the $\pi_i$.
The third row of Table (ref) reports the coverage of a standard 95% confidence interval. When $p^1 = 0.5$, the standard confidence intervals have approximately 95% coverage for both outcomes. As we increase $p^1$, there is a tradeoff between the fact that the estimator is biased (which leads to lower coverage) and the fact that the variance estimator is conservative (which leads to higher coverage), as formalized in Proposition (ref). For the log earnings outcome, the bias dominates and coverage decreases in $p^1$---coverage of the EATT is only about 88.8% when $p^1 = 0.9$. By contrast, for the state-level log average employment outcome, the bias is smaller, and so the conservativeness of the variance estimator dominates---the coverage rate is 99.1% when $p^1 = 0.9$. For comparison, the last row of Table (ref) reports the coverage of an “oracle” 95% confidence interval that uses the true variance of the DID estimator instead of the estimated variance $\hat{s}^2$. When $p^1 = 0.9$ for log-earnings, for example, coverage would be only 51.6% using the oracle variance, but is 88.8% using the conventional conservative variance estimator.
Finally, Table (ref) highlights the implications of the heteroskedasticity-robust variance estimator's conservativeness for constructing robust confidence intervals for the partially identified EATT, as discussed in Section (ref). The Imbens-Manski CIs that account for the bias have coverage of at least 93.9% in all specifications. As $p^1$ increases, coverage becomes more conservative---for the state-level log-average employment outcome, the coverage rate is 99.5% when $p^1 = 0.9$. For comparison, we again report the coverage of an “oracle” 95% confidence interval for the identified set that uses the true variance of the DID estimator, which remains approximately 95% for both outcomes as $p^1$ varies. These results illustrate that robust confidence intervals that account for the bias provide a conservative estimate of how much bias can be accommodated to reach particular conclusions.
Appendix (ref) presents several extensions. We consider simulation designs that vary the number of treated units and finite population sizes. We also consider designs with treatment effect heterogeneity, which we find leads conventional confidence intervals to be even more conservative.
We return to the example of analyzing the impact of Medicaid expansions introduced in Section (ref). wherry_early_2016 study the impact of state-level Medicaid expansions on statewide health insurance coverage using a two-period difference-in-differences estimator that compares the percentage of uninsured individuals ($Y_{it}$) in states that expanded Medicaid in 2014 ($D_i = 1$) against those that did not ($D_i = 0$). The authors estimate $\hat{\tau}_{DID} = -7.1$ and report a 95% CI of $[-11.1, -3.0]$, which implies a standard error of $\hat{s} \approx 2.09$ (see their Table 2). The authors indicate that the standard error is clustered at the state-level. To interpret this standard error from the traditional sampling perspective, we would have to imagine the 50 U.S. states as sampled from an infinite super-population of states. As discussed in Section (ref), it may be more natural to think of the 50 states as fixed, and the state-level treatment assignments as stochastic---e.g. owing to stochastic realizations of state political processes. Our framework implies that if the finite-population parallel trends assumption is satisfied, the CI of $[-11.1, -3.0]$ can alternatively be interpreted as a valid, but possibly conservative 95% confidence for the EATT on the fraction of uninsured individuals.
We may worry that the finite-population parallel trends assumption is violated---we would expect liberal-leaning states to have higher treatment probabilities than conservative states, and they may have different potential outcomes. To address such concerns, we conduct a sensitivity analysis on the authors' conclusions about the EATT. We calculate 95% Imbens-Manski confidence intervals under the assumption that the covariance between the treatment probabilities $\pi_i$ and trends in potential outcomes $Y_{i2}(0) - Y_{i1}(0)$ is bounded in magnitude by a constant $\tilde{b}$, i.e. assuming $\left| \mathbb{C}\text{ov}_{1}\left[\pi_{i}, Y_{i2}(0) - Y_{i1}(0)\right]\right| \leq \tilde{b}$. Figure (ref) shows the resulting confidence intervals for different values of $\tilde{b}$. The “breakdown” value for concluding there is a significant negative effect is $\hat{b}^* \approx 0.9$, i.e. the robust CI excludes zero for all $\tilde{b} < 0.9$. As discussed in Section (ref), this is a conservative estimate of the true “breakdown” value $b^*$ for which the identified set includes 0. Similar to the analysis in rambachan_more_2023 from the super-population perspective, we can benchmark the magnitudes of $\tilde{b}$ using data from years prior to treatment. The authors' Appendix Table 6 suggests that the largest in magnitude finite-population covariance between treated probabilities and trends in untreated potential outcomes occurred between 2012-2013, with a point estimate of $-0.37$ (SE $0.48$); the magnitude of this estimate is well below the breakdown value of 0.9, although its 95% confidence interval includes values larger in magnitude than the breakdown value.
In this section, we present several extensions that illustrate practical implications of our framework for empirical research. First, we consider the common setting where the researcher has data on individuals but treatment is assigned at a more aggregate level. We show that the cluster-robust variance estimator is valid but potentially conservative, justifying the popular heuristic to cluster at the level at which treatment is determined in quasi-experimental settings. Second, we provide two sufficient conditions under which adjusting for differences in baseline covariates can address the bias of the DIM estimator. Finally, we apply our framework to study instrumental variable (IV) estimators, showing conditions under which they have a causal interpretation and how sensitivity analyses can be conducted for violations of these assumptions. Our analysis also extends directly to non-staggered difference-in-differences estimators with multiple time periods (see Appendix (ref)).
We consider the common setting where treatment is determined at a more aggregate level than the unit of observation. Specifically, each unit $i = 1, \hdots, N$ now belongs to one of $C$ clusters, where $c(i)$ denotes the cluster membership of unit $i$. We assume treatment is determined at the cluster level. For example, units $i$ may be individuals living in states $c(i)$, and policy is determined at the state level.
Assumption (ref) is the cluster-level analog to the assignment mechanism considered throughout the paper (Assumption (ref)). Mirroring our earlier notation, let $C_1 := \sum_{c} D_c$ and $C_0 := \sum_{c} (1 - D_c)$ denote the number of treated and untreated clusters respectively, $\pi_c := \mathbb{P}_{R}\left(D_c = 1\right)$ denote the marginal treatment probability for cluster $c$ under Assumption (ref), and $D_i = D_{c(i)}$ denote unit $i$'s treatment assignment. As before, we analyze the behavior of the DIM estimator $\hat{\tau}$ constructed using the outcomes and treatment at the individual level, except we now consider the randomization distribution generated by the clustered treatment assignments. Since the regularity conditions are natural extensions of those in Section (ref) to the clustered design, we defer them to Appendix (ref) and summarize the key takeaways here. Proposition (ref) in the Appendix provides conditions under which $\hat{\tau}$ converges in probability to $\tau_{EATT}^{cluster} + \delta_{cluster}$, where $\tau_{EATT}^{cluster} = \mathbb{E}_{\pi_{c(i)}}\left[\tau_i\right]$ and $\delta_{cluster} = \frac{N}{N - \sum_{i} \pi_{c(i)} } \frac{N}{\sum_i \pi_{c(i)} } \mathbb{C}\text{ov}_{1}\left[\pi_{c(i)}, Y_i(0)\right]$. The first term, $\tau_{EATT}^{cluster}$, is analogous to the EATT discussed earlier, except it uses the cluster-level treatment probabilities $\pi_{c(i)}$ instead of the individual-level probabilities $\pi_i$. Likewise, the bias $\delta_{cluster}$ is proportional to the finite-population covariance between the cluster-level treatment probabilities $\pi_{c(i)}$ and the potential outcome $Y_i(0)$. Proposition (ref) also shows that $\sqrt{C} (\hat{\tau} - \tau^{cluster}_{EATT} - \delta_{cluster})$ converges to a Gaussian distribution, and the LiangZeger(86) cluster-robust variance estimator is consistent for an upper bound on this variance. By contrast, Proposition (ref) shows that the heteroskedasticity-robust variance estimator that ignores clustering can be either too large or small, and thus CIs based on this standard error may not have correct coverage even if $\delta_{cluster} = 0$. Taken together, these results imply that if the need for clustering in quasi-experimental settings arises from the stochastic assignment of treatment, then the researcher should cluster at the level at which treatment is assigned.
Suppose each unit $i$ is associated with fixed covariates $W_i \in \mathbb{R}^{k}$, and consider the OLS regression of the observed outcome on a constant, the treatment $D_i$, and the covariates $W_i$. This is the “covariate-adjusted” DIM studied by Freedman(2008)-regadj_to_experimental_data and Lin(13), among others, in the context of completely randomized experiments. We provide two characterizations of the estimand associated with the OLS coefficient on $D_i$ in our framework.
Proposition (ref) gives two decompositions of the OLS estimand $\beta_D$, the first involving an adjusted outcome and the second involving an adjusted treatment probability. Specifically, part (ref) decomposes the covariate-adjusted DIM into the EATT plus a bias term that depends on the finite-population covariance between the treatment probabilities $\pi_i$ and the covariate-adjusted untreated potential outcomes, $Y_i(0) - \gamma^\prime W_i$, where the coefficient $\gamma$ is a weighted average of the projections of each of the potential outcomes onto the covariates. Thus $\beta_D$ corresponds to the EATT if the treatment probabilities $\pi_i$ are orthogonal to the adjusted potential outcomes. Similar to Section (ref), one could also conduct sensitivity analyses for the bias based on conjectured values for the covariance between $\pi_i$ and the adjusted potential outcomes. Part (ref) alternatively decomposes the covariate-adjusted DIM into $\tau_{OLS}$, which is a particular weighted average of unit-specific treatment effects, and a bias that depends on the finite-population covariance between $Y_i(0)$ and the residualized treatment probability, $\pi_i - \hat\pi_i$. The covariate adjusted DIM estimand thus recovers a weighted average of treatment effects whenever the finite-population covariance between the untreated potential outcomes and the residualized treatment probabilities is equal to zero (note that some of the weights could be negative if $\hat\pi_i > 1$ for some $i$). If the $\pi_i$ are linear in the covariates, then $\hat\pi_i =\pi_i$ and this bias equals zero. Part (ref) thus nests the known result that when the propensity score is linear, the covariate-adjusted DIM gives a variance-weighted average of treatment effects; see angrist_estimating_1998 and abadie_sampling-based_2020 for similar results in a super-population and design-based setting, respectively. Our more general results, however, provide a causal interpretation to the covariate-adjusted DIM if the propensity is not linear in covariates but satisfies the orthogonality conditions described above. Our results also allow us to understand the biases that will result if the propensity score is mis-specified in a way that is related to the potential outcomes.
In Appendix (ref), we provide regularity conditions under which $\sqrt{N}(\hat\beta_D - \beta_D)$ is asymptotically normally distributed, and show that the typical heteroskedasticity-robust standard errors are consistent for an upper bound on the asymptotic variance. Typical standard errors will thus yield conservative inference on $\beta_D$, and sensitivity analyses for the inference that account for the bias will typically be conservative.
In many settings, the researcher has access to an instrumental variable $Z_i$. In some cases, such as a randomized trial with imperfect compliance, the instrument $Z_i$ is completely randomly assigned. However, in other settings the instrument is not explicitly randomized, but the researcher may argue that it is at least partially determined by quasi-experimental factors. For example, in studying the effects of childbearing, angrist_children_1998, Angrist_Lavy_2010 consider having twins at a woman's second birth as an instrument for whether the woman has a third child. The birth of twins $Z_i = 1$ depends on the realization of random biological processes, such as whether a fertilized eggs splits, yet different individuals may have different probabilities of realizing $Z_i = 1$ due to genetic factors, age, or other health risks. Our results can be used to interpret and assess the sensitivity of IV estimates when the instrument may not be completely randomly assigned.
Let $Z_i \in \{0, 1\}$ be a binary instrument, $D_i(z) \in \{0, 1\}$ be the potential treatment status for $z \in \{0, 1\}$, and $Y_i(d)$ be the potential outcome for $d \in \{0, 1\}$. The notation $Y_i(d)$ encodes the exclusion restriction that $Y$ depends on $Z$ only through $d$. We further impose the monotonicity assumption that $D_i(1) \geq D_i(0)$ for all units $i = 1, \hdots, N$. The observed data is then $(Y_i, D_i, Z_i)$, where $Y_i = Y_i(D_i(Z_i))$ and $D_i = D_i(Z_i)$. We view the instrument as stochastic, holding fixed the potential treatments $D(\cdot) = \{D_i(\cdot) \colon i = 1, \hdots, N\}$ and potential outcomes $Y(\cdot) = \{Y_i(\cdot) \colon i = 1, \hdots, N\}$. We let $N_1^Z$ be the number of units with $Z_i = 1$ and $N_0^Z$ be the number of units with $Z_i = 0$.
We write $\mathbb{P}_{R}\left(\cdot\right)$, $\mathbb{E}_{R}\left[\cdot\right]$, $\mathbb{V}_R\left[\cdot\right]$ as probabilities, expectations, and variances respectively under Assumption (ref) and define $\pi_i^{Z} := \mathbb{P}_{R}\left(Z_i = 1\right)$ to be the marginal probability that $Z_i=1$. Similar to Assumption (ref) for the treatment in earlier sections, Assumption (ref) models the instrument assignment as a random experiment with unequal probabilities. In the example of the twin birth instrument, the stochastic instrument assignment corresponds to the realization of the biological process that determines whether a fertilized egg splits in two. The model, however, allows for different women to have different probabilities of having an egg split in two, owing to different biological risk factors, in ways that may be related to their potential outcomes. By contrast, existing IV frameworks typically assume the instrument to be fully independent of the potential treatments and outcomes (see, e.g., AngristImbens(94), angrist_identification_1996 for a sampling-based setting, and kang_inference_2018, hong_inference_2020 for a design-based setting). Our framework will thus allow us to assess the interpretation of the IV estimand in settings where the instrument may not be completely randomly assigned.
We analyze the popular two-stage least-squares (2SLS) estimator, $\ensuremath{\hat{\beta}}_{2SLS} := \hat{\tau}_{RF}/\hat{\tau}_{FS}$, with
corresponding to the reduced form and first-stage, respectively. Proposition (ref) and the monotonicity assumption imply that
where $\mathcal{C} := \{i \colon D_i(1) > D_i(0)\}$ is the set of complier units. We define the 2SLS estimand as $\beta_{2SLS} := \frac{\mathbb{E}_{R}\left[\hat{\tau}_{RF}\right]}{\mathbb{E}_{R}\left[\hat{\tau}_{FS}\right]}$. In Appendix (ref), we show that under conditions similar to those in Section (ref), $\sqrt{N}(\hat\beta_{2SLS} - \beta_{2SLS})$ converges to a Gaussian distribution, and the usual delta-method standard errors for 2SLS are consistent for an upper bound on this variance. (Note we impose “strong instrument” asymptotics where the first-stage is strong relative to sampling variation.) What is the causal interpretation of the estimand $\beta_{2SLS}$? If $\pi_i^Z \equiv \frac{N_1^Z}{N}$, so that all units receive $Z_i=1$ with equal probability, then $\beta_{2SLS} = \frac{1}{|\mathcal{C}|} \sum_{i \in \mathcal{C}} (Y_i(1)-Y_i(0))$, which is a design-based local average treatment effect (LATE) angrist_identification_1996, kang_inference_2018. Our results imply that $\beta_{2SLS}$ maintains a causal interpretation under the weaker orthogonality restriction $\mathbb{C}\text{ov}_{1}\left[\pi_i^Z, Y_i(D_i(0))\right] = \mathbb{C}\text{ov}_{1}\left[\pi_i^Z, D_i(0)\right] = 0$. In this case, $\beta_{2SLS}$ is a weighted average treatment effect among the compliers,
where the weights are proportional to $\pi_i^Z$, the probability that $Z_i = 1$ under Assumption (ref).
Researchers can conduct simple, design-based sensitivity analyses on the two-stage least-squares estimator by placing restrictions on the finite-population covariance between the instrument probabilities and the potential outcomes and treatments to obtain an identified set for the weighted average treatment effect among the compliers. Specifically, assuming $\mathbb{C}\text{ov}_{1}\left[\pi_{i}^{Z}, Y_i(D_i(0))\right] \in [b_{RF}^{lb}, b_{RF}^{ub}]$ and $\mathbb{C}\text{ov}_{1}\left[\pi_i^Z, D_i(0)\right] \in [b_{FS}^{lb}, b_{FS}^{ub}]$, our decompositions of the expectations of $\hat{\tau}_{RF}$, $\hat{\tau}_{FS}$ imply the bounds
Provided the lower bound on $\frac{1}{N_1^Z} \sum_{i \in \mathcal{C}} \pi_i^Z$ is strictly positive, $LATE_{\pi^z}$ must lie in the interval $$ \left[ \frac{\mathbb{E}_{R}\left[\hat{\tau}_{RF}\right] - \frac{N}{N_1^Z} \frac{N}{N_0^Z} b_{RF}^{ub}}{\mathbb{E}_{R}\left[\hat{\tau}_{FS}\right] - \frac{N}{N_1^Z} \frac{N}{N_0^Z} b_{FS}^{lb}}, \frac{\mathbb{E}_{R}\left[\hat{\tau}_{RF}\right] - \frac{N}{N_1^Z} \frac{N}{N_0^Z} b_{RF}^{lb}}{\mathbb{E}_{R}\left[\hat{\tau}_{FS}\right] - \frac{N}{N_1^Z} \frac{N}{N_0^Z} b_{FS}^{ub}} \right]. $$ It is straightforward to estimate these bounds by plugging in $\hat{\tau}_{RF}, \hat{\tau}_{FS}$ in place of the expectations. This will yield consistent estimates of the bounds (under appropriate regularity conditions) in large populations. Likewise, we can further conduct (typically conservative) inference on the bounds based on conventional delta-method standard errors, and we can construct (typically conservative) confidence intervals for $LATE_{\pi^z}$ as in Section (ref).
This paper develops a design-based framework for analyzing quasi-experimental settings in the social sciences in which uncertainty arises from stochastic realizations of treatment assignment, holding fixed the population and their potential outcomes. This perspective is natural in settings where the researcher does not wish to model the statistical process governing the sampling or formation of potential outcomes and the researcher describes the variation being used as the result of quasi-experimental factors that influence treatment status. We derive conditions under which conventional estimators and CIs are valid for interpretable causal parameters in this framework and characterize the bias and size-distortions that arise when these conditions are violated. This leads to natural forms of sensitivity analysis. Altogether, we show that the design-based perspective can also be coherently applied in quasi-experimental settings where there is concern about selection into treatment. While our framework views only treatment assignment as stochastic, an interesting direction for future research could be to study quasi-experimental settings under a finite-population data-generating process that adopts a statistical model for both the outcome and treatment assignments.
Disclosure statement: The authors report there are no competing interests to declare.
\singlespacing
\singlespacing