EconBase
← Back to paper

A Negative Correlation Strategy for Bracketing in Difference-in-Differences

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

69,629 characters · 11 sections · 68 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Negative Correlation Strategy for Bracketing in Difference-in-Differences

\newtheorem{prop}{Proposition} \newtheorem{assum}{Assumption} \newtheorem{theorem}{Theorem} \newtheorem{lemma}{Lemma}

\crefname{equation} \crefname{theorem}{Theorem}{Theorems} \crefname{lemma}{Lemma}{Lemmas}

\def\spacingset#1{ {#1}} \spacingset{1}

\if11 \fi

\if01 {

center[center omitted — 105 chars of source]

} \fi

abstractThe method of difference-in-differences (DID) is widely used to study the causal effect of policy interventions in observational studies. DID employs a before and after comparison of the treated and control units to remove bias due to time-invariant unmeasured confounders under the parallel trends assumption. Estimates from DID, however, will be biased if the outcomes for the treated and control units evolve differently in the absence of treatment, namely if the parallel trends assumption is violated. We propose a general identification strategy that leverages two groups of control units whose outcomes relative to the treated units exhibit a negative correlation, and achieves partial identification of the average treatment effect for the treated. The identified set is of a union bounds form that involves the minimum and maximum operators, which makes the canonical bootstrap generally inconsistent and naive methods overly conservative. By utilizing the directional inconsistency of the bootstrap distribution, we develop a novel bootstrap method to construct confidence intervals for the identified set and parameter of interest when the identified set is of a union bounds form, and we theoretically establish the uniform asymptotic validity of the proposed method. We develop a simple falsification test and sensitivity analysis. We apply the proposed strategy for bracketing to study whether minimum wage laws affect employment levels.

{\it Keywords:} bootstrap, parallel trends, partial identification, sensitivity analysis, uniform inference

\spacingset{1.5}

Introduction

The method of difference-in-differences (DID) is one of the most widely used strategies for policy evaluation in non-experimental settings. The simplest DID estimate is based on a comparison of the outcome differences for the treated units before and after adopting the treatment and the outcome differences for the control units. {In classic DID settings when the treated groups adopt the treatment at the same time}, the DID estimate can be obtained using fixed effects regression models and one can adjust for observed variables Angrist:2009. The key advantage of DID is that it removes time-invariant bias from unobserved confounders. However, the DID method depends on a key assumption that the outcomes in the treated and control units are, in the absence of treatment, evolving in the same way over time. This key assumption is often referred to as the parallel trends assumption, which may not hold in many applications.

The effects of minimum wage laws on employment is a key area of investigation in labor economics. In the United States, minimum wages are often set at the state or local level, which creates numerous opportunities for investigating the effects of these laws. For example, six cities in the U.S., including Seattle and Chicago, increased the minimum wage to \$15 over the last few years dube2019impacts. Minimum wage studies have long relied on DID methods. Even very early studies of first minimum wage laws employed DID obenauer1915effect,lester1946shortcomings. In this context, the parallel trends assumption says that absent the minimum wage laws, employment levels in the treated industries (or states) and in the control industries (or states) would have evolved in the same way over time. However, different industries (or states) react in different ways to the business cycle fluctuation of the economy berman1997industries, and thus the parallel trends assumption is likely violated. Given these threats to validity, researchers need tools that exploit the strengths of DID, but depend on less stringent assumptions.

A growing literature has developed more robust inference strategies for DID designs. For example, Abadie:2005 and callaway2019difference assume that the parallel trends assumption holds after conditioning on observed covariates. Athey:2006 assume a changes-in-changes model that is general but rules out classic measurement error on the outcome. Daw:2018aa and Ryan:2018aa focus on matching in DID analyses. freyaldenhoven2019pre propose to net out the effect of the unmeasured confounding by utilizing a covariate that is affected by the same confounding factors as the outcome but is unaffected by the treatment. manski2013deterrence, Manski:2017aa and Rambachan:2019aa consider partial identification approaches to DID when the variation in outcomes or violations of parallel trends are restricted to some known set.

In this article, we propose a general strategy for DID that addresses the concerns about heterogeneity in different units' outcome dynamics. We leverage two groups of control units whose untreated potential outcome relative to that for the treated units exhibit a negative correlation across the study period, i.e., when the relative outcome for one control group increases, the relative outcome for the other control group decreases (illustrated in Figure (ref) and discussed in Section (ref)). In this case, DID estimates can be constructed using these two control groups and used to bound (bracket) the average treatment effect for the treated units. This general strategy builds on an idea in hasegawadid2018 but requires a weaker assumption, which can be motivated in a much broader context (see Section 3.1). We derive the identification assumption for this general strategy, which is shown to accommodate many commonly adopted assumptions in the DID literature, most noticeably the interactive fixed effects model with a single interactive fixed effect.

The identified set for our proposed bracketing method takes a “union bounds” form, namely the bounds can be expressed as the {union} of several intervals. {Inference for the identified set or the parameter of interest that belongs to this identified set is challenging because the union bounds form involves the minimum and maximum operators, which makes the canonical bootstrap generally inconsistent Shao:1994aa, romano2008inference, Andrews:2009aa, Bugni:2010aa, Canay:2010aa. For this reason, the percentile bootstrap can be overly conservative. The “intersection-union” approach of berger1996bioequivalence can also be quite conservative. A related but different problem is inference on intersection bounds chernozhukov2013intersection, which falls into a broader class of problems where the identified set is defined by moment inequalities. Much progress has been made in this direction tamer2010partial, canay2017practical; however, these methods cannot be directly used for union bounds because the union bounds cannot generally be represented using moment inequalities. Therefore, it is important to develop valid and informative inference methods for the union bounds. To this end, by utilizing the directional inconsistency of the bootstrap distribution, we develop a novel and easy-to-implement bootstrap method to construct confidence intervals (CIs) for the identified set and parameters of interest, and we theoretically establish the uniform asymptotic validity of the proposed method. This new inference method for union bounds is itself an important contribution to the fast growing literature on inference for partially identified parameters.

Our paper proceeds as follows. In Section (ref), we introduce notation and our causal framework, and review the DID method. In Section (ref), we develop the general bracketing strategy in DID. We introduce the identification assumption and derive the identification formula. Then we develop a bootstrap inference method and study its theoretical properties. We also develop a falsification test and sensitivity analysis. In Section (ref), we examine the finite sample empirical performance of the proposed bootstrap inference method for union bounds in a simulation study. In Section (ref), we apply the proposed methods to study the employment effect of minimum wage laws. Section (ref) concludes with a discussion.

Review: Causal Effects Based on DID

We consider applications where data are observed for the treated and control units before and after the treated units adopt the treatment, while the control units are never treated. Suppose the treated units adopt the treatment between two time periods, which we denote as $t=1$ and $t=2$, and remain treated afterwards. We will refer to time period $t= 1$ as the pre-treatment period and time periods $t=2, \dots, T$ as the post-treatment periods. We write $D_{i}=1$ if individual $i$ belongs to the treated units, $D_{i}=0$ if individual $i$ belongs to the control units. One common data configuration is where the units are states, and the treatment is a change in state policy for one or more states. {In the case of staggered adoption where the treatment is adopted by multiple states over time, we can group treated states according to their treatment adoption time and consider each group separately; see Section S1.2 of the Supplement for details.}

As in Neyman:1923a and Rubin:1974, we define treatment effects in terms of potential outcomes. Let $Y_{it}^{(1)}$ represent the potential outcome for individual $i$ at time $t$ if being treated, let $Y_{it}^{(0)}$ represent the potential outcome for individual $i$ at time $t$ if being untreated. We assume throughout that at each time $t$, the potential outcomes and the treated unit indicator $( Y_{it}^{(1)}, Y_{it}^{(0)}, D_i), i=1,\dots, N_t$, are identically and independently distributed (i.i.d.) realizations of $( Y_t^{(1)}, Y_t^{(0)}, D)$. {Some caveats and reflections on such a random sampling assumption can be found in Manski:2017aa.} Relatedly, there have been studies that evaluate the impact of within-cluster correlation arising from a random unit-time specific component Bertrand:2004, Donald:2007; see imbens2008recent for a review. In this article, we instead view the unit-time specific components as fixed effects and propose to bracket these fixed effects rather than modeling them as random. As such, the unit-time specific components can be accounted for using our bracketing method without creating within-cluster correlation. More specific discussion on this point is given in Section (ref) below. Depending on the treatment status, the observed outcomes can be expressed as $Y_{it}=Y_{it}^{(0)}$ for $t\leq 1$, $Y_{it}=D_{i} Y_{it}^{(1)}+(1-D_{i})Y_{it}^{(0)}$ for $t=2, \dots, T$. The observed data $ \{Y_{i1}, D_{i} \}_{i=1,\dots, N_1}, \dots, \{Y_{iT}, D_{i} \}_{i=1,\dots, N_T}$ can be obtained from a longitudinal study of the same units over time or a repeated random sample. Hereafter, we drop the subscript $i$ to simplify the notation.

We are interested in the average treatment effect for the treated units in the post-treatment periods, \[ ATT_t=E[Y_t^{(1)}-Y_t^{(0)}|D=1], \qquad t=2, \dots, T, \] where the expectation is taken with respect to the distribution of $(Y_t^{(1)}, Y_t^{(0)}, D)$. Note that $E[Y_t^{(1)}|D=1]=E[Y_t|D=1]$ for $t=2, \dots, T$ can be identified from the observed data, but we never observe $Y_{t}^{(0)}$ for the treated units in post-treatment periods. One approach to causal identification is to use the method of DID under the parallel trends assumption

eqnarray[eqnarray omitted — 124 chars of source]

which says that the treated units and the control units would have exhibited parallel trends in the potential outcomes in the absence of treatment. With this assumption, we can use the control units to identify the change in the potential outcomes for the treated units had the units counterfactually not been treated. Thus, $ATT_t$ can be identified through

align[align omitted — 242 chars of source]

where the third equality holds because of the parallel trends assumption in ((ref)). In the simplest scenario, the DID estimator replaces the above conditional expectations with the corresponding sample averages.

A General DID Bracketing Strategy

In extant work, researchers have sought to relax the parallel trends assumption in various ways. The DID bracketing method in hasegawadid2018 is one proposal that connects the outcome levels and outcome dynamics in the absence of treatment, such that the changes in outcome for different groups are ordered according to their outcome levels. Facilitated by the control group construction approach discussed in hasegawadid2018 in which units are designated to the lower (upper) control group if the average outcome is lower (higher) than the average outcome for the treated group in a prior-study period, the two standard DID estimators based on these two control groups can bound the true treatment effect. The models and assumptions required by hasegawadid2018 are reviewed in detail in Section S1.1 of the Supplement. In Section (ref), we will establish that the bracketing relationship holds much more broadly.

Identification

We present our key partial identification assumption based on two control groups which we denote as $ a,b $. Let $\Delta_t (g)=E[Y_t^{(0)}-Y_{t-1}^{(0)}|G=g]$ be the expected change in potential untreated outcome for group $g$ from time $t-1$ to time $t$. Let $ \Gamma_t (g) = E[Y_t^{(0)}|G=g]- E[Y_t^{(0)}|G=trt]$ be the expected potential untreated outcome for control group $ g $ at time $ t $ relative to the treated group. The only identification assumption in our new proposal is

assum(monotone trends) For $t= 2, \dots, T$, \begin{align}\min \left\{ \Delta_t (a), \Delta_t (b) \right\}\leq \Delta_t(trt)\leq \max\left\{ \Delta_t (a), \Delta_t (b) \right\}. \end{align}

One can show that $ \Delta_t (g) - \Delta_t (trt) = \Gamma_t{(g)}- \Gamma_{t-1}{(g)}$. Therefore, Assumption (ref) has an equivalent formulation, which we state as Lemma (ref).

lemmaAssumption (ref) is equivalent to the following: for $t= 2, \dots, T$, \begin{align} \left\{\Gamma_t(a)- \Gamma_{t-1} (a)\right\} \left\{\Gamma_t(b)- \Gamma_{t-1} (b)\right\} \leq 0. \end{align}

What does Assumption (ref) imply about the behavior of the two control groups relative to the treated group? According to Assumption (ref), in every pair of adjacent time periods, the change in outcome for control group $ a $ and that for control group $ b $ provide bounds on the change in outcome for the treated group if it were untreated. The equivalent formulation in Lemma (ref) provides another interesting perspective: the untreated potential outcome for the two control groups $ a,b $ relative to that for the treated units exhibit a negative correlation across the study period. In other words, if the relative outcome for control group $ a $ increases (decreases), the relative outcome for control group $ b $ decreases (increases).

Figure (ref) provides an illustrative example of Assumption (ref) and the equivalent formulation as in Lemma (ref). In the left figure, the treated group has the lowest outcome level across all three time periods. However, for every pair of adjacent time periods, the slope for the treated group (i.e., $\Delta_t (trt)$) is bounded by the slopes for the two control groups (i.e., $\Delta_t (a)$ and $\Delta_t (b)$). Specifically, $\Delta_2 (b)> \Delta_2 (trt)> \Delta_2 (a)$, because from $ t=1 $ to $ t=2 $, the expected outcome for the control group $ b $ has a larger increase compared to the treated group and the expected outcome for the control group $ a $ decreases, and similarly $\Delta_3 (a)> \Delta_3 (trt)>\Delta_3 (b)$, which implies that Assumption (ref) holds. The left figure is then translated into the figure on the right by plotting the untreated potential outcomes for the two control groups relative to the treated group (i.e., $ \Gamma_t (a)$ and $ \Gamma_t (b) $). From the right figure, for every pair of adjacent time periods, the relative outcome for one control group increases and that for the other control group decreases. Specifically, $ \Gamma_2 (a)-\Gamma_1 (a)<0, \Gamma_2 (b)-\Gamma_1 (b)>0 $, so that ((ref)) holds for $ t=2$; $ \Gamma_3 (a)-\Gamma_2 (a)>0, \Gamma_3 (b)-\Gamma_2 (b)<0 $, so that ((ref)) holds for $ t=3$; this implies the equivalent condition in Lemma (ref) is satisfied.

figure[figure omitted — 5,319 chars of source]

Note that the models and conditions assumed in hasegawadid2018 (reviewed in Section S1.1 of the Supplement) imply Assumption (ref), so the bracketing method in hasegawadid2018 is a special case of our general strategy.

Assumption (ref) also accommodates many existing assumptions in the literature. For example, the standard DID method assumes the parallel trends assumption that requires the outcome dynamics for every group are the same, i.e., $\Delta_t (a)=\Delta_t (trt)=\Delta_t (b)$ for every $t$, under which Assumption (ref) holds. Therefore, the bracketing method is valid under the standard DID assumptions. Assumption (ref) also relates to the parallel growth assumption mora2012treatment, which requires that $\Delta_t(a)-\Delta_{t-1}(a)=\Delta_t(trt)-\Delta_{t-1}(trt)=\Delta_t(b)-\Delta_{t-1}(b)$ for every $t$. If we construct the control groups such that $\Delta_0 (a)\leq \Delta_0 (trt)\leq\Delta_0 (b)$, then the parallel growth assumption implies that $\Delta_t (a)\leq \Delta_t (trt)\leq\Delta_t (b)$ for every $t$, and thus also implies Assumption (ref).

More generally, Assumption (ref) can be motivated from the group-level interactive fixed effects model with a single interactive fixed effect:

align[align omitted — 93 chars of source]

for $ g\in \{a, b, trt\} $ and $ t=1,\dots, T $. Here, $ Y_{gt}^{(0)} $ is group $ g $'s untreated potential outcome at time $ t $, $ \alpha_t $ is a time fixed effect, $ (\eta_g, \lambda_g )$ are unobserved, time-invariant group characteristics, and $ F_t $ is an unobserved time-varying factor, $ \epsilon_{gt} $ is a mean zero random error. Model (ref) accommodates the heterogeneity in different groups' outcome dynamics in the absence of treatment due to the heterogeneous impact of a common latent trend, i.e., $ \lambda_g F_t $. It also includes the parallel trends and parallel growth assumptions as special cases, which can be seen by setting $ F_t=0 $ and $ F_t-F_{t-1} = F_{t-1}-F_{t-2} $ for every $ t $, respectively. Based on model (ref), if we find two control groups such that $ \lambda_a\leq \lambda_{trt}\leq \lambda_b $ or $ \lambda_a\geq \lambda_{trt}\geq \lambda_b $, then Assumption (ref) holds. Unlike previous work on treatment effects in interactive fixed effects models Bai:2009aa, xu2017generalized that typically requires the number of time periods to go to infinity, Theorem (ref) achieves partial identification of the treatment effect with as few as two time periods by utilizing two control groups whose $ \lambda_a, \lambda_b $ bound $ \lambda_{trt} $.

Recall that we define the average treatment effect for the treated $ATT_t=E[Y_t^{(1)}-Y_t^{(0)}|G=trt]$ for $t=2,\dots, T$. For $t=2$, we can relate the DID parameter using each of the two control groups $g\in \{a,b\}$ to $ATT_2$,

equation[equation omitted — 143 chars of source]

where $\tau_2(a), \tau_2(b)$ are standard DID parameters. For the case where $t>2$, define

equation[equation omitted — 154 chars of source]

where $\tau_t(a), \tau_t(b)$ are not standard DID parameters because $t-1$ is also a post-treatment period when $t>2$. Under Assumption (ref), it is true that for every $t$, $\min \{\Delta_t (trt) -\Delta_t (a),\Delta_t (trt) -\Delta_t (b) \}\leq 0$ and $\max \{\Delta_t (trt) -\Delta_t (a),\Delta_t (trt) -\Delta_t (b) \}\geq 0$. Therefore, when $ t=2 $, $\tau_2(a)$ and $\tau_2(b)$ bound the $ATT_2$, i.e., $\min\{\tau_2(a), \tau_2(b)\} \leq ATT_2 \leq \max\{\tau_2(a), \tau_2(b)\}$; when $t>2$, we have $\min\{\tau_t(a), \tau_t(b)\} \leq ATT_t-ATT_{t-1} \leq \max\{\tau_t(a), \tau_t(b)\}$.

We state the key partial identification result as follows. {In Section S1.2 of the Supplement, we provide an extension to adjust for covariates.}

theoremUnder Assumption (ref), the average treatment effect for the treated $ATT_t, t= 2, \dots, T$ can be partially identified through \begin{eqnarray} &&\sum_{s=2}^{t}\min\{\tau_s(a), \tau_s(b)\} \leq ATT_t \leq \sum_{s=2}^{t}\max\{\tau_s(a), \tau_s(b)\}, \end{eqnarray} where $\tau_2(a)$ and $\tau_2(b)$ are defined in ((ref)), $\tau_s(a)$ and $\tau_s(b)$ for $s>2$ are defined in ((ref)).

The tightness of the bounds depends on the magnitude of violation of the parallel trends assumption, since the width of the bounds equals \[ \sum_{s=2}^{t}\big[\max\{\tau_s(a), \tau_s(b)\} -\min\{\tau_s(a), \tau_s(b)\}\big]=\sum_{s=2}^t\left|\Delta_s(b)-\Delta_s (a)\right|, \] where $|\cdot|$ denotes the absolute value. If the parallel trends assumption holds over the study period, i.e., $\Delta_s (a)=\Delta_s (b) $ for every $s$, the width of the bounds equals zero for every post-treatment period, and (ref) becomes the identification formula from the standard DID.

In fact, for the control groups $a$ and $b$, Theorem (ref) holds under a weaker assumption than Assumption (ref), that is the monotone trends assumption holds in a cumulative fashion. Specifically, for the bounds in ((ref)) to be valid, the sufficient and necessary condition is

eqnarray[eqnarray omitted — 214 chars of source]

This cumulative monotone trends assumption is useful for scenarios when Assumption (ref) is subject to mild violations for brief time periods, but ((ref)) still holds.

{Lastly, several papers have noted that the usual parallel trends assumption may be sensitive to the functional form chosen for the outcome, for example, it may hold for $Y$ but not $\log(Y)$ or vice versa Athey:2006, kahn2020promise. These concerns are applicable to Assumption (ref) as well. roth2020parallel show that the parallel trends assumption holds for all functional forms under a parallel trends type assumption on the full distribution. It would be interesting to see if similar results can be obtained for the monotone trends assumption. We leave that to future work. }

Estimation and Inference

In Theorem (ref), we assume that two control groups satisfying Assumption (ref) are given. In practice, we typically must identify these two control groups from the candidate control units. In the application in Section (ref), the control groups are selected based on domain knowledge. {If data prior to the study period are available, we can also select two control groups in a data-driven way; see Section S1.3 of the Supplement for details.} Once the two control groups are given, we can proceed with inference of the treatment effect.

First, it is important to discuss the implications of the i.i.d. assumption we invoked at the beginning of Section (ref). Consider the standard model in a DID design \[ Y_{gti} = \alpha_t+ \eta_g+\tau D_{gt} + \lambda_{gt} + \epsilon_{gti}, \] where $ g $ indexes group, $ t $ indexes time, $ i $ indexes individual imbens2008recent. Here, $ Y_{gti} $ is the outcome, $ D_{gt} $ is a indicator variable that equals one if group $ g $ is being treated at time $ t $, $ \alpha_t, \eta_g, \tau $ are unknown parameters, respectively representing time effects, time-invariant group effects, and the treatment effect of interest, $ \lambda_{gt} $ is an unobserved group-time specific effect, and $ \epsilon_{gti} $ is an individual-level random error. Bertrand:2004 and Donald:2007 note that in some applied work, the $ \lambda_{gt} $ (i.e., time specific factors which affect the whole group) were effectively set to zero and ignored, which can severely understate the standard error of the DID estimator. Donald:2007 outline an approach which models the $ \lambda_{gt} $ as mean zero random factors. In our approach, we instead view $ \lambda_{gt} $ as fixed effects and propose to bracket the $ \lambda_{gt} $'s of treated and control groups rather than modeling them as random. As such, the existence of non-zero $ \lambda_{gt}$ can be accounted for using our bracketing method without creating within-group correlation.

In what follows, we introduce a novel bootstrap method to construct confidence intervals (CIs) for the partially identified parameter of interest $ATT_t$ and its identified set in ((ref)) for $ t=2,\dots, T$. Differences between CIs for the identified set and for the parameter of interest within that set have been well-addressed in the prior literature; see, for instance, Imbens:2004aa and Stoye:2009aa. Algebra reveals that the identified set in ((ref)) for $ATT_t$ can be equivalently formulated as

eqnarray[eqnarray omitted — 192 chars of source]

where the proof is in the Supplement. For example, when $ t=2 $, the bounding parameters are $\{ \tau_2(a), \tau_2(b)\} $, and their minimum and maximum form the bounds for $ ATT_2 $; when $ t=3 $, the bounding parameters are $ \{ \tau_2(a)+\tau_3(a), \tau_2(a)+\tau_3(b), \tau_2(b)+\tau_3(a), \tau_2(b)+\tau_3(b)\} $, and their minimum and maximum form the bounds for $ ATT_3 $. In general, there are $ 2^{t-1} $ bounding parameters for $ t\geq 2 $. We call such bounds “union bounds”.

To describe the bootstrap inference procedure, we first rearrange the data as ${\bm O}_i=(Y_{i1}, Y_{i2}, \dots, Y_{iT}, R_{i1}, \dots, R_{iT}, G_i)$, $i=1, \dots, N$, which are assumed to be i.i.d. sequence of random vectors with distribution $P\in{\cal P}$, where $ {\cal P} $ is a family of distributions that satisfy a very weak uniform integrability condition stated in Theorem (ref), $Y_{it} $ is the outcome for individual $ i $ at time $ t $, $ R_{it} $ indicates whether $ Y_{it} $ is observed or not, which equals 1 if we observe $ Y_{it} $, equals 0 if not, and $N $ is the total number of individuals we observe. We assume $ R_{it} $ is independent of $ (Y_{it}, G_i) $ for every $ t $, and $ P(R_{it}=1) $ and $ P(G_i=g) $ are strictly bounded away from zero for every $ t $ and $ g $. This data configuration enables the proposed inferential method to account for arbitrary serial correlation among multiple observed outcomes for an individual. Suppose based on $ \bm{O}_1, \dots, \bm{O}_N $, we compute a vector of sample means denoted by $\bar{{\bm X}}_N$, and ${\bm \mu} (P)=E_P( \bar{{\bm X}}_N)$. Let $\{ \theta_j (P)=\theta_j({\bm \mu (P)}), j=[k]\}$ be the set of bounding parameters and let $\hat{\theta}_{Nj}=\theta_j(\bar{{\bm X}}_N )$ be their estimators, where $ k $ is a finite number and $[k] = \{1,\dots, k\}$. The parameter of interest $ \psi_0 (P)$ belongs to the identified set $ \Psi_0(P)=[\theta_{\min}(P), \theta_{\max} (P)] $, where $\theta_{\min} (P) = \min_{j\in[k]} \theta_j (P) $ and $\theta_{\max} (P) = \max_{j\in[k]} \theta_j (P) $. The goal is to construct uniformly valid CIs for $ \Psi_0(P) $ and $ \psi_0(P)$ in the asymptotic sense.

As an illustration, in the setting when the parameter of interest is $ ATT_2 $,

displaymath\bar{{\bm X}}_N= \left[ \begin{array}{c} \frac{\sum_{G_i=trt} Y_{i2} R_{i2}}{\sum_{G_i=trt} R_{i2} } -\frac{\sum_{G_i=trt} Y_{i1} R_{i1}}{\sum_{G_i=trt} R_{i1}}\\ \frac{\sum_{G_i=a} Y_{i2} R_{i2}}{\sum_{G_i=a} R_{i2}} -\frac{\sum_{G_i=a} Y_{i1} R_{i1}}{\sum_{G_i=a} R_{i1}} \\ \frac{\sum_{G_i=b} Y_{i2} R_{i2}}{\sum_{G_i=b} R_{i2}} -\frac{\sum_{G_i=b} Y_{i1} R_{i1}}{\sum_{G_i=b} R_{i1}} \end{array}\right], \qquad {\bm \mu} (P)= \left[ \begin{array}{c} E_P(Y_2-Y_1|G=trt) \\ E_P(Y_2-Y_1|G=a)\\ E_P(Y_2-Y_1|G=b) \end{array}\right],

and $ \theta_1(P)= E_P(Y_2-Y_1|G=trt)- E_P(Y_2-Y_1|G=a)$, $ \theta_2(P)= E_P(Y_2-Y_1|G=trt)- E_P(Y_2-Y_1|G=b)$. The goal is to construct uniformly valid CIs for the identified set $ [\min(\theta_1(P), \theta_2(P)), \max(\theta_1(P), \theta_2(P))] $ and parameter of interest $ ATT_2 $.

Suppose we obtain a nonparametric bootstrap sample ${\bm O}_1^*, \dots, {\bm O}_N^*$ that is drawn from the empirical distribution based on $ {\bm O}_1, \dots, {\bm O}_N $. Let $ \bar{{\bm X}}_N^* $ and $\hat{\theta}_{Nj}^*=\theta_j( \bar{{\bm X}}_N^* ), j\in [k] $ be the bootstrap analogues calculated based on the bootstrap sample. The bootstrap inference method can be implemented following Algorithm 1. According to Theorem (ref), the random interval in (ref) is a uniformly valid $ 1-\alpha $ level CI for the identified set $ \Psi_0 (P) $, and thus the parameter of interest $ \psi_0 (P) $; the random interval in (ref) is a more refined uniformly valid $ 1-\alpha $ level CI for $ \psi_0 (P) $ that adapts to the width of the identified set.

Next, we derive the theoretical properties of the CIs ((ref))-(ref). For $m= m_N$ a sequence of positive integers tending to infinity but satisfying $ m/N\rightarrow 0$ or $m= N$, let $\hat\theta_{N,\min}, \hat\theta_{N,\max}, \hat\theta_{N, \min}^{*b}, \hat\theta_{N, \max}^{*b}, \hat\theta_{m,\min}, \hat\theta_{m,\max}$ be as defined in Algorithm 1, $L_m(x)=P\{\sqrt{m}(\hat\theta_{m,\min} - \theta_{\min} (P) )\leq x\}$ be the true distribution of $\hat\theta_{m,\min}$, $\hat{L}_{N,\rm mod}(x)=P_*\{\sqrt{N}(\hat\theta_{N, \min}^{*b}-\hat\theta_{N,\rm min})\leq x\}$ be the proposed bootstrap estimator of $ L_m(x) $, where $P_*$ is the conditional probability with respect to the random generation of bootstrap sample given the original data. Analogously, define $R_m(x)=P\{\sqrt{m}(\hat{\theta}_{m,\max} -\theta_{\max}(P))\leq x\}$ and $\hat{R}_{N,\rm mod}(x)=P_*\{\sqrt{N}(\hat\theta_{N, \max}^{*b} -\hat\theta_{N,\max})\leq x\}$. Theorem (ref) lays the theoretical foundation for the proposed inference procedure.

theoremFor $ t=1,\dots, T $ and $ g=a, b, trt $, suppose that $r_{igt}= \{Y_{it}- E(Y_{it}\mid G_i=g)\} I(G_i= g) R_{it} / P(G_i=g, R_{it}=1)$ is uniformly integrable in the sense that \[ \lim_{\lambda\rightarrow \infty} \sup_{P\in \mathcal{P}} E_P\left\{ \frac{r_{igt}^2}{\mbox{Var} (r_{igt}) } I\left( \frac{|r_{igt}|}{\mbox{Var} (r_{igt})^{1/2}} >\lambda\right) \right\} =0, \] $ \theta_j(\bm\mu) = \bm c_j^T \bm\mu $ with $\bm c_j $ a vector of fixed constants for $ j \in[k]$, and either (S1) or (S2) holds: \\ (S1) $\lim_{N\rightarrow \infty }\inf_{P\in \mathcal{P}} \inf_{j\in [k]: \theta_j (P) \neq \theta_{\max}(P)} \sqrt{N} ( \theta_{\max}(P) - \theta_j(P) )= \infty $ and \\ $\lim_{N\rightarrow \infty }\inf_{P\in \mathcal{P}} \inf_{j\in [k]: \theta_j (P) \neq \theta_{\min}(P)} \sqrt{N} ( \theta_j(P) - \theta_{\min}(P) )= \infty $. Set $m= N$.\\ (S2) $m \rightarrow \infty$ and $m/N\rightarrow 0$, as $N\rightarrow \infty$. Then, (a) $\lim_{N\rightarrow\infty} \sup_{P\in{\cal P}} \sup_{x\in {\cal R}} \{\hat{L}_{N, \rm mod}(x)-L_m(x)\} \leq 0$ and $\lim_{N\rightarrow\infty} \inf_{P\in{\cal P}} \inf_{x\in {\cal R}} \{\hat{R}_{N,\rm mod}(x)-R_m(x)\} \geq 0$. \\ (b) Let $c_L^* (p) =\inf\{ x\in{\cal R}: \hat{L}_{N, \rm mod}(x)\geq p\}$, $c_U^*(p)=\sup\{ x\in{\cal R}: \hat{R}_{N,\rm mod}(x)\leq p\}$, then \begin{align} &\lim_{N\rightarrow\infty} \inf_{P\in{\cal P}} P\left\{ \sqrt{m}(\hat{\theta}_{m, \min} -\theta_{\min} (P))\leq c_L^* (p)\right\}\geq p, \nonumber\\ &\lim_{N\rightarrow\infty} \inf_{P\in{\cal P}} P\left\{ \sqrt{m}(\hat{\theta}_{m, \max}- \theta_{\max}(P))\geq c_U^* (1-p)\right\}\geq p \nonumber. \end{align} (c) Set $ p=1-\alpha/2 $, then \begin{align} CI_{1-\alpha}\equiv \left[ \hat{\theta}_{m, \min} - m^{-1/2}c_L^*(1-\alpha/2), \hat{\theta}_{m, \max} - m^{-1/2}c_U^*(\alpha/2)\right] \end{align} is a uniformly valid $1-\alpha$ level CI for the identified set $\Psi_0 (P)=[ \theta_{\min}(P), \theta_{\max}(P) ]$, i.e., $ \lim_{N\rightarrow\infty} \inf_{P\in{\cal P}} P\big( \Psi_0(P) \subseteq CI_{1-\alpha} \big) \geq 1-\alpha $.\\ (d) Let $ \hat{w}^{+} = \hat{w} I(\hat{w} >0) $, where $ \hat{w}= \{ \hat{\theta}_{m,\max}- m^{-1/2}c_U^*(1/2)\}- \{ \hat{\theta}_{m, \min}- m^{-1/2}c_L^*(1/2)\}$, and $ \hat{p}= 1- \Phi( \rho \hat{w}^+) \alpha$, where $ \Phi(\cdot) $ is the standard normal cumulative distribution function, $ \rho $ is a sequence of constants satisfying $ \rho\rightarrow\infty $, $ m^{-1/2} \rho\rightarrow 0 $ and $ \rho| \hat{w}^+ - (\theta_{\max}(P)-\theta_{\min}(P)) | \xrightarrow{P}0$, where $ \xrightarrow{P} 0$ denotes convergence in probability, then \begin{align} CI^{\psi}_{1-\alpha}\equiv \left[ \hat{\theta}_{m, \min}- m^{-1/2}c_L^*(\hat{p}), \hat{\theta}_{m, \max}- m^{-1/2}c_U^*(1-\hat{p})\right] \end{align} is a uniformly valid $1-\alpha$ level CI for the partially identified parameter $ \psi_0(P)$, i.e., {$ \lim_{N\rightarrow\infty} \inf_{P\in{\cal P}}\inf_{\psi_0(P)\subseteq \Psi_0(P)} P\big(\psi_0(P)\in CI^\psi_{1-\alpha} \big)\geq 1-\alpha $.}

The proof is in the Supplement. Note that the uniform integrability condition is also required for the standard percentile bootstrap Romano:2012aa and is important for the uniform asymptotic validity of our bootstrap procedure. As discussed in Romano:2012aa, CIs satisfying the uniform asymptotic validity are more desirable compared to CIs that only satisfy the pointwise asymptotic validity.

From Theorem (ref)(a), the bootstrap estimators $ \hat{L}_{N,\rm mod}(x), \hat{R}_{N,\rm mod}(x)$ are not consistent for the true distributions $ L_m(x) , R_m(x)$, and this is caused by the possibility that there can be more than one bounding parameters equal to $ \theta_{\min} (P) $ and $\theta_{\max} (P) $. Also, this inconsistency is directional, particularly, $ \hat{L}_{N,\rm mod}(x) $ tends to be smaller than $ L_m(x) $, and $ \hat{R}_{N,\rm mod}(x)$ tends to be larger than $ R_m(x) $. Given that the goal is to construct a CI with asymptotic coverage probability at least $ 1-\alpha $, using critical values $ c_L^*(p), c_U^*(p) $ satisfying $ \hat{L}_{N, \rm mod} (c_L^*(p))\geq p, \hat{R}_{N, \rm mod}(c_U^*(p)) \leq p$, we have asymptotically $ L_m (c_L^*(p))\geq \hat{L}_{N, \rm mod} (c_L^*(p)) \geq p $, $ R_m(c_U^*(p))\leq \hat{R}_{N, \rm mod}(c_U^*(p)) \leq p$. This property lays the ground for the proposed bootstrap inference method.

Theorem (ref) contains two choices of $m$. When it is safe to assume the conditions in Scenario (S1) hold, i.e., uniformly over $P\in \mathcal{P}$, the bounding parameters that achieve the minimum and maximum are well-separated from the other parameters, then we can set $m= N$, which makes $d_{j, \min } = d_{j,\max} = 0$ for all $j$. Setting $m=N$ makes more efficient use of the data and tends to have shorter CIs because the widths of the CIs are proportional to $m^{-1/2}$. In this case, the proposed bootstrap procedure is still not the same as the standard percentile bootstrap CI

align[align omitted — 201 chars of source]

that uses the $ \alpha/2 $ quantile as the lower end and $ 1-\alpha/2 $ quantile as the upper end. The proposed bootstrap inference procedure in Algorithm 1 subtracts the $ 1-\alpha/2 $ quantile from the lower end and the $ \alpha/2 $ quantile from the upper end. As shown by the simulation in Table (ref), due to the skewness of the distributions of $ \min_j \hat{\theta}_{Nj}^{*b} $ and $ \max_j \hat{\theta}_{Nj}^{*b} $, the proposed bootstrap inference procedure in Algorithm 1 not only guarantees adequate coverage probability but also produces shorter CIs compared to the standard percentile bootstrap. When the conditions in Scenario 1 are in doubt, Theorem (ref) indicates that a more conservative bootstrap procedure with $m$ satisfying the conditions in (S2) can still ensure uniform validity. In practice, we can set $m= N/\log(\log(N))$. This bootstrap procedure is motivated from guo2021inference, but with modifications to ensure uniform validity in our setting. Note that the idea of using a subset of data to have strong validity guarantee under irregular settings has also appeared in wasserman2020universal. {Based on extensive simulations included in Section S1.4 of the Supplement, we find that for DID bracketing, the bootstrap with $m=N$ has adequate coverage probability even with shrinking difference between the bounding parameters that achieve the minimum (and maximum) and the other parameters. Therefore, we still recommend using the bootstrap with $m=N$ because it is more efficient.}

The interval ((ref)) is in fact a Monte Carlo approximation of ((ref)) in Theorem (ref)(c). Notice that the $ CI_{1-\alpha} $ in Theorem (ref)(c) is also a uniformly valid CI for the parameter of interest $ \psi_0 (P)$, because the event that $ CI_{1-\alpha} $ covers the identified set $\Psi_0(P)$ implies the event that $ CI_{1-\alpha} $ covers the parameter of interest $ \psi_0(P) $. This also means that $ CI_{1-\alpha} $ as a CI for $ \psi_0 (P)$ can be further improved by taking into consideration the width of the identified set, which motivates $ CI^\psi_{1-\alpha} $ in Theorem (ref)(d).

The idea in Theorem (ref)(d) is to set $ \hat{p}=1-\alpha $ when the bounds are wide enough in the sense that $ \rho( \theta_{\max} (P) - \theta_{\min}(P)) \rightarrow \lambda\in(0,\infty]$, and set $ \hat{p}=1-\alpha/2 $ if $ \rho( \theta_{\max} (P) - \theta_{\min}(P)) \rightarrow 0$, where $ \rho $ satisfies $ \rho\rightarrow \infty, N^{-1/2} \rho\rightarrow 0 $. The use of $ \Phi(\cdot) $ is simply to smoothly connect these two ends because $ \Phi(\rho\hat{w}^+) \in [0,1/2]$. The intuition behind this construction is that if the bounds are wide relative to the measurement error, the parameter of interest $ \psi_0 (P)$ can only be close to at most one boundary of the identified set, so the asymptotic probability that $ \psi_0(P) $ is more extreme than the other boundary is negligible and the noncoverage risk is one-sided. This reasoning appears in Imbens:2004aa, Stoye:2009aa and CLR2009. In practice, we set $ \rho=\{m^{-1/2} \max[ c_U^* (3/4)-c_U^* (1/4), c_L^* (3/4)-c_L^* (1/4)]\}^{-1}/\log(m) $ as in Algorithm 1, which is proportional to $\sqrt{m}/\log(m)$ and is reminiscent of the parameter used in generalized moment selection for moment inequalities andrews2010inference. A caveat of using these types of sequences of tuning parameters is that although they don't affect the asymptotic performance of the inference procedure, they may affect the finite sample performance and there is limited guidance on how to choose these parameters. There has been much effort in the moment inequality literature to produce tests that are valid in a finite-sample normal model without relying on sequences of tuning parameters, e.g., andrews2012inference and romano2014practical.

We emphasize that $ CI^\psi_{1-\alpha} $ is valid uniformly with respect to the location of $ \psi_0 (P)$ in the identified set $\Psi_0(P) $ and the width of $ \Psi_0 (P)$. This uniformity is important because it ensures that the coverage probability is adequate even when $ \psi_0(P)$ is at the boundary of $ \Psi_0 (P)$, or when the width of $ \Psi_0 (P)$ shrinks towards zero and point identification is established. In particular, the point identification scenario is very salient for the union bounds developed in Section (ref), because when the parallel trends assumption holds over the study period, the width of the identified set in ((ref)) equals zero. Theorem (ref)(d) guarantees that $ CI^\psi_{1-\alpha} $ is valid when the parallel trends assumption holds.

{Theorem (ref) also provides bias-corrected estimators for the identified set $ \Psi_0(P)=[\theta_{\min}(P), \theta_{\max} (P)] $. In particular, $ \hat{\theta}_{m, \min}^{\rm med} = \hat{\theta}_{m, \min} - m^{-1/2}c_L^*(1/2)$ is a half-median-unbiased estimator chernozhukov2013intersection for $\theta_{\min}(P)$ in the sense that the lower bound estimator falls below $\theta_{\min} (P)$ with probability at least 1/2 asymptotically, i.e., $\lim_{N\rightarrow\infty} \inf_{P\in{\cal P}} P\{\hat{\theta}_{m, \min}^{\rm med}\leq \theta_{\min} (P) \}\geq 1/2$. Analogously, $ \hat{\theta}_{m, \max}^{\rm med}= \hat{\theta}_{m, \max} - m^{-1/2}c_U^*(1/2)$ is a half-median-unbiased estimator for $\theta_{\max}(P)$ in the sense that the upper bound exceeds $\theta_{\max}(P)$ with probability at least 1/2 asymptotically, i.e., $\lim_{N\rightarrow\infty} \inf_{P\in{\cal P}} P\{ \hat{\theta}_{m, \max}^{\rm med} \geq \theta_{\max} (P) \}\geq 1/2$. }

Lastly, as discussed in the introduction, the union bounds we focus on in this article are different from the intersection bounds considered in CLR2009, chernozhukov2013intersection, in which the minimum operator appears in the upper bound and the maximum operator appears in the lower bound. Applying methods developed in CLR2009, chernozhukov2013intersection to the union bounds tends to produce a CI with insufficient coverage probability. The proposed inference method generally applies to an identified set for a parameter that can be expressed as union bounds, i.e., the lower bound can be formulated as the minimum of a set of bounding parameters, and the upper bound can be formulated as the maximum of another set of bounding parameters. These two sets of bounding parameters can be different. This union bounds problem can also be addressed using the “intersection-union” approach by berger1996bioequivalence, which uses $\max_j (\hat\theta_{Nj} + z_{1-\alpha/2} \hat\sigma_{Nj})$ as an upper $1-\alpha/2 $ level CI for $\theta_{\max} (P)$, where $z_{1-\alpha/2}$ is the $1-\alpha/2$ quantile of the standard normal distribution and $\hat\sigma_{Nj}$ is the standard error of $\hat\theta_{Nj}$. The lower $1-\alpha/2$ level CI is $\min_j (\hat\theta_{Nj} - z_{1-\alpha/2} \hat\sigma_{Nj})$. By Bonferroni's inequality, $\big[\min_j (\hat\theta_{Nj} - z_{1-\alpha/2} \hat\sigma_{Nj}), \max_j (\hat\theta_{Nj} + z_{1-\alpha/2} \hat\sigma_{Nj})\big]$ is an $1-\alpha$ level CI for the identified set $ \Psi_0 (P)$; the proof is in the Supplement. This approach can be conservative if the $\theta_j$'s are close together, but yields a uniformly valid CI under mild assumptions without any tuning parameters; see the Supplement for a proof. Another alternative approach is the bootstrap of fang2019inference using the directional differentiability of the max/min function.

Assessing the Validity of the Design

To enhance the reliability of an observational study, it is useful to include additional analyses that explore whether the key assumptions appear plausible and whether conclusions are robust to violations of the key assumptions Rosenbaum:2010. In this section, methods for a falsification test and a sensitivity analysis are developed.

Falsification Testing

In the context of DID, investigators often conduct a falsification test by testing for parallel trends in pre-treatment time periods Angrist:2009. Our falsification test is similar in principle, where we test whether the monotone trends relationship holds in a pair of unused adjacent time periods prior to the study period, say $t= t_1^*, t_2^*$.

The null hypothesis that the monotone trends assumption holds when $t=t_1^*, t_2^*$ is $$ H_0: \min \left[ \Delta_{t_2^*} (a), \Delta_{t_2^*} (b) \right]\leq \Delta_{t_2^*} (trt)\leq \max \left[ \Delta_{t_2^*} (a), \Delta_{t_2^*} (b)\right], $$ and the alternative hypothesis includes two possible scenarios when $H_0$ is not true:\\ (i) $\min \left\{ \Delta_{t_2^*} (a), \Delta_{t_2^*} (b) \right\}> \Delta_{t_2^*} (trt)$; (ii) $\max \left\{ \Delta_{t_2^*} (a), \Delta_{t_2^*} (b) \right\}< \Delta_{t_2^*}(trt)$. Next, we define a set of simple hypotheses:

eqnarray[eqnarray omitted — 308 chars of source]

The null hypothesis $H_0$ can be written in a form of a composite null hypothesis, that is \[ H_0: (H_{a}^i\cap H_{b}^i)\cup (H_{a}^d\cap H_{b}^d). \] Let the p-values testing each individual hypothesis $H_{a}^{i}, H_{b}^{i}, H_{a}^{d}, H_{b}^{d}$ be $p_{a}^i, p_{b}^i, p_{a}^d$, $p_{b}^d$. From the definition of one-sided p-values, $ p_a^d=1-p_a^i, p_b^d=1-p_b^i.$ Therefore, following Bonferroni's method and berger1982multiparameter, we reject $H_0$ if \[ \max(\min(p_{a}^i, p_{b}^i), \min(1-p_{a}^i, 1-p_{b}^i))\leq \alpha/2. \]

{It is important to emphasize the limitations of falsification tests as they pertain to the prior-study period whereas Assumptions 1 is for the study period. Moreover, non-rejection of the falsification tests hypotheses does not provide any evidence in favor of the identifying assumptions. At best, a non-rejection provides some assurance that the data does not outright refute the premises of the main analysis. Lastly, a number of studies have noted that DID falsification tests may have low power in finite samples and they may distort estimation and inference roth2019pre,kahn2020promise,bilinski2018seeking,hartman2018equivalence. Similar issues arise in our setting as well. The limitations of falsification tests motivate the sensitivity analysis we outline next. }

Sensitivity Analysis

We develop a sensitivity analysis to evaluate how sensitive the conclusion is to violations of Assumption (ref). A sensitivity analysis is used to quantify the degree to which a key identification assumption must be violated in order for a researcher's original conclusion to be reversed. There is a large and growing literature on sensitivity analysis, e.g., Rosenbaum:1987, Imbens:2003, ding2016sensitivity and fogary2017avg.

There are two scenarios when Assumption (ref) is violated at time $t$: (i) $\min \left\{ \Delta_t (a), \Delta_t (b) \right\}> \Delta_t(trt)$; or (ii) $\max \left\{ \Delta_t (a), \Delta_t (b) \right\}< \Delta_t(trt)$. We use two sensitivity parameters $\gamma_t$ and $\delta_t$ for these two scenarios at time $t$, where $\gamma_t$ is for scenario (i) and $\delta_t$ is for scenario (ii), and we introduce the following sensitivity assumption.

assum[Sensitivity] For the given non-negative sensitivity parameters\\ $\{\gamma_t, \delta_t\}_{t\geq 2}$, the two control groups $a $ and $b$, and $t=2,\dots, T$, $$\min \left\{ \Delta_t (a), \Delta_t (b) \right\}-\gamma_t\leq \Delta_t(trt)\leq \max \left\{ \Delta_t (a), \Delta_t (b) \right\}+\delta_t.$$

{Assumption (ref) states that violation to the monotone trends assumption is bounded, which is similar in principle to the bounded-variation assumption in Manski:2017aa.} When $\gamma_t=\delta_t=0$ for every $t$, Assumption (ref) degenerates to Assumption (ref), under which the bounds in ((ref)) are valid. Next, we derive the bounds and CI for $ATT_t$ under Assumption (ref), which will serve as a basis for the sensitivity analysis.

theoremUnder Assumption (ref), for $ t=2,\dots, T $, \\ (a) The treatment effect for treated $ATT_t$ can be partially identified through \begin{eqnarray} \sum_{s=2}^{t}\min\{\tau_s(a), \tau_s(b)\} - \sum_{s=2}^t \delta_s\leq ATT_t \leq \sum_{s=2}^{t}\max\{\tau_s(a), \tau_s(b)\}+\sum_{s=2}^t \gamma_s , \end{eqnarray} where $\tau_s(a), \tau_s(b)$ are in ((ref))-((ref)), $\delta_s , \gamma_s$ are sensitivity parameters defined in Assumption (ref).\\ (b) Further assume the conditions in Theorem (ref), let $[\hat{l}_{t}, \hat{r}_t]$ be the uniformly valid $1-\alpha$ CI for $ATT_t$ (or the identified set) developed using Theorem (ref) (c)-(d) under Assumption (ref), then $[\hat{l}_t- \sum_{s=2}^t \delta_s, \hat{r}_t+\sum_{s=2}^t \gamma_s]$ is a uniformly valid $1-\alpha$ CI for $ATT_t$ (or the identified set) under Assumption (ref).

The proof is in the Supplement. The CI $ [\hat{l}_t, \hat{r}_t] $ can be constructed using the method developed in Theorem (ref). The form of the CI under Assumption (ref) illustrates how large $\gamma_s, \delta_s$ have to be in order for the study conclusion to be materially altered. If $\hat{l}_{t}, \hat{r}_t$ are both positive, $\sum_{s=2}^t\delta_s=\hat{l}_t$ would suffice to explain away the treatment effect, which requires that $\sum_{s=2}^t \Delta_s (trt) \geq \sum_{s=2}^t\max \{\Delta_2 (a), \Delta_2 (b)\}+\hat{l}_t$. If this hypothesized scenario is unlikely to happen in practice, the observed positive treatment effect is robust. Similarly, if $\hat{l}_t, \hat{r}_t$ are both negative, $\sum_{s=2}^t\gamma_s= \hat{r}_t$ would suffice to explain away the treatment effect, which requires that $\sum_{s=2}^t \Delta_s (trt) \leq \sum_{s=2}^t\min \{\Delta_2 (a), \Delta_2 (b)\}-\hat{r}_t$. If this hypothesized scenario is unlikely to happen in practice, the observed negative treatment effect is robust.

Simulations

We empirically evaluate the performance of the proposed bootstrap inference methods for union bounds. {Here we only present results using the recommended bootstrap method (i.e., with $m=N$) under the following two scenarios.} Additional simulations including the bootstrap with $m/N\rightarrow 0$ and under more scenarios are in Section S1.4 of the Supplement. \\ {\bf Case I: parallel trends:} $ E[Y^{(0)}_1 |G=trt] = 3, E[Y^{(0)}_1|G=a] = 10, E[Y^{(0)}_1|G=b] = 4$, $ \Delta_t (trt)= \Delta_t (a)= \Delta_t (b)\equiv \Delta_t $ for every $ t $, where $ \Delta_2=1, \Delta_3=-2, \Delta_4=-1 $.\\ {\bf Case II: partially parallel trends:} $ E[Y^{(0)}_1 |G=trt] = 3, E[Y^{(0)}_1|G=a] = 10, E[Y^{(0)}_1|G=b] = 4$, $ \Delta_2 (trt)=1, \Delta_3 (trt)=-4, \Delta_4 (trt)=1, \Delta_2 (a)=1, \Delta_3 (a)=-1, \Delta_4 (a)=1, \Delta_2 (b)=2, \Delta_3 (b)=-4, \Delta_4 (b)=1 $.

In both scenarios, the simulated data resembles a longitudinal study where $ N=1000 $ individuals are followed for $ T=4 $ time points. The group indicators $ G_i $ are given, where $ P(G=trt)=0.3 $, $ P(G=a) =0.2$, $ P(G=b)=0.5 $. The average treatment effects for the treated group are $ATT_1=0, ATT_2=2, ATT_3=3, ATT_4=1 $. The observed outcomes are generated from $ Y_{it}=E(Y_t^{(0)}|G) + ATT_t +\varepsilon_{it}$, for $ t=1,\dots, T $, where $ \varepsilon_{it}$'s are independent from the standard normal distribution. We set the number of bootstrap iterations $ B=300 $ and significance level $ \alpha=0.05 $.

Table (ref) reports the simulation average of half-median unbiased estimators of the bounds. It also shows the average length and the empirical probability that the CI covers the parameter of interest $ ATT_t$ (i.e., coverage probability) for the proposed bootstrap CIs in (ref)-(ref), respectively for the identified set and $ ATT_t $. We compare with the intersection-union method of berger1996bioequivalence formed by taking the union of every $ 1-\alpha $ level CI for $ \sum_{s=2}^t \tau_s (g_s),$ $ g_s\in\{a,b\}$, and the percentile bootstrap CI in (ref).

We summarize the key findings as follows. First, all the methods produce CIs with adequate coverage probability, indicating the validity of all four types of CIs. Second, the intersection-union and the standard percentile bootstrap methods can both be unnecessarily conservative, especially when the width of the bounds is small (Case I) and the number of the bounding parameters is large ($ t=4 $). Overall, the proposed bootstrap methods improve significantly over the two comparison methods, as they produce shorter CIs with correct coverage probability. When one is interested in a CI for the partially identified parameter itself rather than the identified set, the bootstrap CI tailored for the parameter can be tighter than that for the identified set, and the improvement is more evident when the width of the bounds is large (Case II). Lastly, the proposed bootstrap methods also provide an easy recipe to generate estimates of the bounds. In Case I where the treatment effect is point identified, the true lower and upper bounds are equal and are respectively equal to 2, 3, 1 for $t=2,3,4$, respectively. The half-median unbiased estimators $\hat\theta_{N, \min}^{\rm med}, \hat\theta_{N, \max}^{\rm med}$ are closer to the true lower and upper bounds compared to $\hat\theta_{N, \min}, \hat\theta_{N, \max}$. The result is similar but maybe to a smaller extent for Case II where the treatment effect is partially identified and the lower (upper) bounds are equal to 1, -1, -3 (2,3,1) for $t=2,3,4$, respectively.

Minimum Wage Data Application

The Fair Labor Standards Act (FLSA) of 1938 introduced the federal minimum wage in the United States, which covered economic sectors such as manufacturing, transportation and communication, wholesale trade, finance and real estate, and affected about 54% of the U.S. workforce. More economic sectors were included in 1961, 1966, 1974, and 1986 Amendments to the FLSA. A classification of industry by date of FLSA coverage can be found in derenoncourt2018minimum. derenoncourt2018minimum used Current Population Survey data and applied a very nice cross-industry DID design to study the effects of the 1966 FLSA. Next, we apply the proposed DID bracketing strategy to investigate the employment effect of the 1974 FLSA that extended the federal minimum wage coverage to employees in the federal government and to private household domestic service workers.

The treated group is comprised of 797 employees that work in either private households or for the federal government -- the two industries for which the minimum wage was added by the 1974 FLSA. To select two control groups, we use work in economics on the correlations between industries and gross domestic product (GDP). According to berman1997industries, employment in private households and federal government (the two industries in the treated group) tends to have weak positive correlation with GDP. Employment in construction and retail trade (two industries covered by the 1961 FLSA), however, tends to have strong positive correlation with GDP, while employment in agriculture, entertainment and recreation services, nursing homes and other professional services, and hospitals (four industries covered by the 1966 FLSA) tends to be negatively correlated with GDP. We designate the 798 employees in the first set of industries as control group $ a $, and the 1568 employees in the second set of industries as control group $ b $. Under the interactive fixed effects model in (ref), Assumption (ref) is plausible with GDP conceptualized as the time-varying latent factor $F_t$.

Following derenoncourt2018minimum, we restrict our sample to all prime-age workers aged 25 to 55, run the cross-industry design at the industry $ \times $ state $ \times $ year level, and define the outcome variable as the log employment rate. We focus on the average effect of the 1974 FLSA on employment in the affected industries in 1975. We report both adjusted and unadjusted results. As in the original study, we control for the average age, share of males, share of white workers, share of married persons, share of low education workers, and industry fixed effects using a linear model.

First, we plot trends in the log employment rate. The upper panel of Figure (ref) plots the average outcomes for the treated and two control groups, and the lower panel of Figure (ref) plots the average outcomes for the two control groups after subtracting the average outcomes for the treated group. A visual inspection of the lower panel in Figure (ref) shows that the relative average outcomes for the two control groups exhibit a negative correlation during 1972-1974, providing visual evidence that Assumption (ref) is plausible. We also apply the proposed falsification test using the 1972-1973 data, which is non-overlapping with the 1974-1975 data that we will soon use for analysis. The p-value for the falsification test (unadjusted) equals 0.26, for the falsification test (adjusted) equals 0.25. Therefore, we find no evidence that data in the prior study period are inconsistent with Assumption (ref).

figure[figure omitted — 334 chars of source]

Table (ref) contains the results from the DID bracketing strategy. We also include point estimates and 95% CIs from the standard DID method for comparison, which are directly from linear regression. We report estimates for $ATT_t$ and CIs obtained from the proposed bootstrap method in (ref)-(ref) for the identified set and $ ATT_t $ respectively, using 300 repetitions and $m=N$. In summary, DID bracketing finds no detectable effect of the 1974 FLSA on employment and this result is robust to adjusting for observed confounders. In this example, the CIs for the identified set and the parameter of interest are nearly identical, which is because the widths of the bounds are not large compared to the measurement error and thus $ \hat p $ used in (ref) is very close to $ 1-\alpha/2 $. Based on the CIs for $ ATT_t $, we are able to rule out a negative effect on the log employment rate more extreme than 0.029. The results based on standard DID are similar. Our results are also similar to the results presented in derenoncourt2018minimum.

Concluding Remarks

The method of difference-in-differences (DID) is widely used for policy evaluation in non-experimental settings. The DID method has the advantage of allowing flexible data structures (e.g., it works for repeated cross sectional studies or longitudinal data) and being able to remove time-invariant systematic differences between the treated and control groups. However, it is also well known that the DID method relies on a strong parallel trends assumption, which requires that in the absence of the treatment, the treated and control groups would experience the same outcome dynamics. To relax the stringent parallel trends assumption, recent work by hasegawadid2018 outlined a bracketing method, that uses two control groups with different outcome levels in the prior-study period and addresses the bias arising from a historical event interacting with the groups.

In this work, we propose a general strategy for bracketing in DID that addresses the concerns about the heterogeneity in different units' outcome dynamics in the absence of treatment, and includes the original bracketing method in hasegawadid2018 as a special case. Critically, we leverage two control groups whose untreated potential outcomes relative to the treated group exhibit a negative correlation, and such negative correlation naturally exists under the interactive fixed effects model with a single interactive fixed effect. This negative correlation appears to be reasonable in our minimum wage application, as industries fluctuate with the business cycle in different ways, some are cyclical while some are countercyclical. {More broadly, our bracketing strategy is applicable to studies where the main concern about confounding is differential economic shocks, which includes many topics in labor economics. Another general area that the method is applicable is in public health. For example, in state-level DID analyses of policies to address drug overdose deaths, a major concern is differential shocks in the availability of fentanyl. We can use states such as Ohio that were first hit by fentanyl and had high per capita overdose rates as one control group, and states such as Nebraska that were relatively insulated from fentanyl and had low per capita overdose rates as the other control group zoorob2019fentanyl. Under an interactive fixed effects model with the nationwide fentanyl supply as the latent time-varying factor, these two control groups can be used to bound the potential overdose outcomes of the other states in the absence of policy interventions. }

Another important contribution in this work is that we develop a novel and easy-to-implement bootstrap inference method to construct uniformly valid CIs for union bounds (i.e., the identified set can be formulated as the union of several intervals). This bootstrap inference method for union bounds has the potential of being applied to broader settings.