Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
106,249 characters · 19 sections · 56 citation commands
Design-Robust Two-Way-Fixed-Effects Regression For Panel Data \@thefnmark\@footnotetextGenerous support from the Office of Naval Research through ONR grants N00014-17-1-2131 and N00014-19-1-2468 is gratefully acknowledged.
\thispagestyle{empty}
Keywords: fixed effects, panel data, causal effects, treatment effects, double robustness, staggered adoption.
\baselineskip=20pt \setcounter{page}{1}
Difference-in-difference (DiD) methods ({\it e.g.}, \citet*{ashenfelter1985using, angrist1999empirical}) are commonly used in empirical economics to establish causal relationships (see \citet*{currie2020technology} for some evidence regarding the usage in the empirical literature). In particular, researchers estimate regression functions of the form
using ordinary least squares (OLS), treating $\alpha_i$ and $\lambda_t$ as fixed parameters -- the fixed effects, leading to the two-way fixed effect (TWFE) estimator. Here $Y_{it}$ is the outcome variable of interest, $W_{it}$ is a binary treatment, $X_{it}$ are observed exogenous characteristics, and $\tau$ is the main object of interest. Practitioners routinely justify regression ((ref)) by appealing to “quasi-experimental” variation in treatment paths $\boldsymbol{W}_i = (W_{i1},\dots, W_{iT})$. Formal and informal arguments are invoked to make a case that this variation is not associated with unobserved unit and time-specific components $\epsilon_{it}$. In other words, to motivate ((ref)), researchers reason about the underlying model for $\boldsymbol{W}_i$. This model, however, does not explicitly enter the estimation process. Moreover, econometric assumptions that justify the OLS estimation apply conditionally on $\boldsymbol{W}_i$ and do not appeal to randomness in the treatment paths (e.g., arellano2003panel).
In this paper, we develop new methods for estimating causal effects that explicitly incorporates design assumptions on the assignment process without abandoning the transparency and simplicity of the two-way model. We incorporate assumptions about the assignment mechanism by augmenting the specification ((ref)) with unit-specific weights $\gamma_i$, leading to
We compute the weights $\gamma_i$ using the assignment model for $\boldsymbol{W}_i$.
We start our analysis by assuming that the assignment process for $\boldsymbol{W}_i$ is known. In Section (ref), we show how to use this knowledge to construct oracle weights $\gamma^{\star}$ and conduct design-based inference. Under the correct specification of the assignment model, our inference procedure is valid regardless of the underlying model for potential outcomes, and in particular, we do not rely on the validity of the equation ((ref)). Our results substantially generalize the properties established in athey2018design, allowing for an arbitrary assignment process (subject to overlap restrictions).
To construct $\gamma^{\star}$, we need to solve a nonlinear equation that depends on the support of $\boldsymbol{W}_i$. Practically, this means that the construction and the values of the weights vary across different types of assignment processes. In Appendix (ref) we provide solutions for several prominent examples, including staggered adoption, {\it i.e.}, a situation where units opt into treatment sequentially. Another input we need for $\gamma^{\star}$ is the probability distribution of $\boldsymbol{W}_i$ (generalized propensity score, imbens2000).
After establishing design-based properties of the oracle estimator $\hat \tau(\gamma^{\star})$ based on knowledge of the assignment process, we turn to the robustness -- the behavior of the estimator in settings where the postulated assignment model can be incorrect. At this point, we use the structure of the regression problem ((ref)) to demonstrate that $\hat \tau(\gamma^{\star})$ has a strong double-robustness property (\citet*{robins1994estimation,kang2007demystifying, bang2005doubly, chernozhukov2018double}): it has a small bias whenever either the assignment or the regression model is approximately correct. We view these results as the primary motivation for using our estimator in practice, where we cannot expect either the TWFE model or the assignment model to be fully correct.
In practice, the assignment model is rarely completely known -- unless $\boldsymbol{W}_i$-s are assigned in the controlled experiment (e.g., \citet*{attanasio2012education,broda2014economic, colonnelli2022corruption})-- and has to be estimated. We use the insights from the known assignment setting as a building block in Section (ref), where the assignment process is unknown but can be estimated consistently from the data. In Section (ref) we use an empirical example to show how to estimate this distribution for the staggered adoption design using duration models. This approach is connected to shaikh2019randomization that uses a duration model to test a sharp null hypothesis that specifies no treatment effects. Our general strategy of explicitly using the assignment model for estimation is directly connected to the recent literature on quasi-experimental designs (e.g., borusyak2022nonrandom). Our results on robustness are especially appealing in such contexts because in quasi-experimental settings researchers cannot rule out the misspecification of the assignment model.
Our focus on TWFE regression ((ref)) is motivated by its increased popularity in economics \citep*{currie2020technology}. In applications, this model provides an effective and parsimonious approximation for the baseline outcomes, allowing researchers to capture unobserved confounders and to improve the efficiency of the resulting estimator by reducing noise. At the same time, recent research shows that regression estimators for average treatment effects based on TWFE models might have undesirable properties, particularly negative weights for unit-time specific treatment effects. These concerns are particularly salient in settings with heterogeneity in treatment effects and general assignment patterns (e.g., \citet*{de2020two,goodman2018difference,sun2021estimating,callaway2018difference,borusyak2017revisiting}). Our results show that the concerns raised in this literature regarding negative weights lose some of their force under random assignment, or more generally once we properly reweight the observations.
Our main analysis assumes that the treatment affects only contemporaneous outcomes, ruling out dynamic effects. We make this choice to crystallize the connection between the TWFE regression model ((ref)) and the assignment process. We do not restrict heterogeneity in contemporaneous treatment effects that can vary over units and periods. To test for, or estimate, dynamic treatment effects, one has to compare units that receive treatment at different times. Such comparisons are justified only if we restrict individual heterogeneity in treatment effects or treat the assignment as random. Consequently, and this is of course a key insight from the causal inference literature in cross-section settings since rosenbaum1983central, it is imperative to model both the assignment mechanism and the outcome model. In \citet*{bojinov2020panel} the authors show how to use the assignment process to estimate dynamic treatment effects (see also blackwell2021adjusting for the related analysis in large-$T$ setup). Our results suggest that a fruitful approach may be to construct robust estimators by combining bojinov2020panel approach to estimation with conventional dynamic panel regression models using the weighting methods derived in the current paper for the static case. We discuss a particular realization of this in Section (ref).
Our results are related to recent literature on doubly robust estimators with panel data. Conceptually the closest paper to us is arkhangelsky2019double that also emphasizes the role of the assignment process in the same setting and shows double robustness. In arkhangelsky2019double the focus is on a class of estimators defined as a linear function of realized outcomes, with the coefficients in that linear representation chosen to lead to consistent estimators for average treatment effects under either assumptions on the outcome model or on the assignment mechanism. Here, we start with a different class of estimators, restricted to weighted versions of the TWFE estimator in ((ref)). We also show how to estimate a flexible class of average treatment effects with user-specified weights over units and time. The robustness property in our paper is distinct from the double robustness analyzed recently in the difference-in-difference literature (e.g., sant2020doubly): our estimator is robust to arbitrary violations of parallel trends assumptions, as long as the assignment model is correctly specified. At the same time, our estimator is not necessarily semiparametrically efficient in environments where, as in sant2020doubly, the conditional parallel trends assumption holds.
We also connect to recent work on causal panel model with experimental data (e.g., athey2018design,bojinov2020panel,roth2023efficient). Similar to these papers, we establish properties of regression estimators under design assumptions. Importantly, we consider a general setting without restricting our attention to staggered adoption design. Our contribution to this literature is the characterization of the behavior of $\hat \tau(\gamma)$ for a large class of weighting functions and general designs. By establishing a connection between weighting functions and limiting estimands, we allow users to construct consistent estimators for a pre-specified weighted average treatment effect of interest.
Finally, the form of our estimator ((ref)) connects it to the Synthetic Difference in Differences (SDID) estimator introduced in \citet*{arkhangelsky2019synthetic}. The difference between these two procedures is in the way they construct the weights $\gamma^{\star}$. The SDID estimator uses pretreatment outcomes to build a synthetic control unit that follows the path of the average treated unit as closely as possible (up to an additive shift). This strategy is infeasible if $W_{it}$ varies over time. However, precisely in situations with enough variation in $\boldsymbol{W}_i$, we can estimate the assignment process and use it to construct the weights $\gamma^{\star}$. As a result, the two estimators are complementary and can be used in applications with different assignment patterns.
Throughout the paper, we adopt the standard probability notation $O(\cdot), o(\cdot), O_{\mathbb{P}}(\cdot),$ $o_{\mathbb{P}}(\cdot)$. For any vector $v$, denote by $v^\top$ the transpose of $v$, $\|v\|_{2}$ the $L_2$ norm of $v$, and by $\operatorname*{diag}(v)$ the diagonal matrix with the coordinates of $v$ being the diagonal elements. For a pair of vectors $v_1, v_2$, we write $\left\langle v_1, v_2\right\rangle$ for their inner product $v_1^\top v_2$. Furthermore, let $[m]$ denote the set $\{1, \ldots, m\}$, $I_{m}$ the $m\times m$ identity matrix, and $\textbf{1}_{m}$ the $m$-dimensional vector with all entries $1$. Finally, the support of a discrete distribution $F$ is the set of elements with positive probabilities under $F$.
We consider a setting with $n$ units and each unit is characterized by potential outcomes $\boldsymbol{Y}_i(1) = (Y_{i1}(1), \ldots, Y_{iT}(1)), \boldsymbol{Y}_i(0) = (Y_{i1}(0), \ldots, Y_{iT}(0))$ and a set of covariates $\boldsymbol{X}_i = (X_{i1}, \ldots, X_{iT})$.\footnote{Time-invariant covariates can be handled by letting $X_{i1} = \ldots = X_{iT}$.} By writing the potential outcomes in this form, we assume away any dynamic effects of past treatments on current outcomes, thus focusing on static models. Analysis of such models is useful both theoretically and practically. First, they constitute a building block for more general environments, which we consider in Section (ref). Second, when the treatment is irreversible, as in staggered adoption designs, we are likely interested in its average (over time) effect on the outcome rather than the transitory dynamics. This makes the static model a reasonable approximation for a more complicated dynamic model. Finally, if we observe the data at a lower frequency than the one that is relevant for dynamics (e.g., days vs. months), then the static model is the only available option.
Given the realized treatment assignment $W_{it}$, the observed outcomes are defined in the usual way:
Throughout the paper, we treat covariates as fixed and consider $\{(\boldsymbol{Y}_i(1), \boldsymbol{Y}_i(0), \boldsymbol{W}_i): i\in [n]\}$ as a random vector (jointly) drawn from a distribution (conditional on $\{\boldsymbol{X}_i: i\in [n]\}$). We let $\mathbb{P}$ denote the joint distribution of the entire random vector $\{(\boldsymbol{Y}_i(1), \boldsymbol{Y}_i(0), \boldsymbol{W}_i): i\in [n]\}$ (conditional on $\{\boldsymbol{X}_i: i\in [n]\}$) and $\mathbb{E}$ denote the expectation over this distribution. We consider the asymptotic regime with $n$ going to infinity and fixed $T\ge 2$.
This structure nests the conventional sampling-based framework, which is common in panel data analysis, going back to chamberlain1984panel, and which was used to establish statistical results in the recent DiD literature abadie2005semiparametric, callaway2018difference. It also extends the standard fixed effects framework, where the distribution for each unit is characterized by unit-specific parameters, but units themselves are usually assumed independent neyman1948consistent,lancaster2000incidental. Even in the absence of any covariates we do not assume that unit-level observations $(\boldsymbol{Y}_i(1), \boldsymbol{Y}_i(0), \boldsymbol{W}_i)$ are independent or exchangeable, which brings two practical advantages. First, it allows us to accommodate correlated potential outcomes among units, which is natural in applications involving networks or multilevel structures. Second, it allows the assignments to be correlated across units, which is natural for many commonly used experimental designs. We elaborate on this point in the next section.
In this section we study a special case where the assignment mechanism is known. This assumption is natural for experimental settings brown2006stepped, attanasio2012education, broda2014economic, hemming2015stepped, chandar2019drivers, chandar2019design, colonnelli2022corruption, but it has also been used to analyze the quasi-experimental settings borusyak2023non. It allows us to derive inferential results under mild assumptions. We will consider the case of unknown designs in Section (ref) at the cost of stronger (yet standard) assumptions.
We assume that, for any $i \in \{1, \ldots, n\}$ and $\boldsymbol{w}\in \{0, 1\}^T$,
where $\bm{\pi}_i$ is a distribution known to the analyst. We call it the generalized propensity score (imbens2000, athey2018design, bojinov2020panel, bojinov2020design) -- the marginal probability of the treatment path.
This structure allows for covariate-adaptive designs, where the probability of $W_{it}$ depends on past covariates. However, we rule out sequentially-adaptive designs where the assignment can depend on past outcomes, even if the randomization protocol is known.\footnote{Even if units are i.i.d. and $\mathbb{P}(W_{it}\mid Y_{i1}, \ldots, Y_{i(t-1)})$ is known, $\mathbb{P}(W_{it}\mid Y_{i1}, \ldots, Y_{iT})$ would depend on the unknown conditional distribution of $(Y_{it}, \ldots, Y_{iT})$ given $W_{it}$} Furthermore, our framework places no restriction on the support of $\boldsymbol{W}_i$ and substantially generalizes the previous works that focus on simple random sampling for non-staggered difference-in-differences rambachan2020design and staggered adoption athey2018design,roth2023efficient.
If the treatment paths $\{\boldsymbol{W}_i, i \in [n]\}$ are independent across units, then the marginal distributions $\{\bm{\pi}_i(\boldsymbol{w}), i \in [n]\}$ characterize the joint distribution of $\{\boldsymbol{W}_i, i \in [n]\}$. However, as discussed above, we allow the assignments to be correlated across units. In practice, this correlation can range from being very mild, as in the case of completely randomized experiments with a fixed share of treated units neyman23, to being sizable, as in cases of cluster-level randomization such as cluster randomized design abadie2023should and two-stage randomization. We impose technical restrictions on the dependence across units in Section (ref).
We define the unit and time-specific treatment effect as:
Note that $\tau_{it}$ can vary with both $i$ and $t$ since we assume neither identically distributed units nor time-homogeneous treatment effects. For time period $t$, we define the time-specific ATE as:
and consider a broad class of weighted average of time-specific ATE:
for some user-specified deterministic weights $\xi = (\xi_{1}, \ldots, \xi_{T})^{\top}$ such that
We refer to (ref) as a doubly average treatment effect (DATE). For example, the weights $\xi_{t}=1 / T$ yield the usual ATE over units and time periods. In the difference-in-differences setting with two time periods, $\xi_{t} = \boldsymbol{1}_{t = 2}$. In a particular application, one might also be interested in an effect with time discounting factor that puts more weight on initial periods, i.e. $\xi_{t} \propto \beta^{t}$ for some $\beta < 1$.
We allow $\boldsymbol{W}_i$ to be dependent across units to capture different assignment processes. Such dependence arises in applications, sometimes for technical reasons (e.g., in case of sampling without replacement as in athey2018design), and sometimes by the nature of the assignment process (spatial experiments). To quantify this dependence as well as the dependence among the potential outcomes, we follow renyi1959measures and define the maximal correlation:
where the supremum is taken over all real-valued measurable functions $f, g$.
In the standard design-based framework where potential outcomes are assumed fixed, it reduces to the $\rho$-mixing coefficient between $\boldsymbol{W}_i$ and $\boldsymbol{W}_j$. In the main text, we maintain a simplified restriction on $\{\rho_{ij}\}_{ij}$ leaving a more general one to Appendix (ref). The assumption is stated as follows:
By definition $\frac{1}{n}\le (1/n^2)\sum_{i,j=1}^{n}\rho_{ij}\le 1$ with lower bound being attained if the observations are independent, and the upper bound being attained if they are perfectly dependent. As a result, one can view $q$ as measuring the strength of the correlation. When $(\boldsymbol{Y}_i(1), \boldsymbol{Y}_i(0), \boldsymbol{W}_i)$ are independent across units, (ref) holds with $q = 1$. More generally, when $\{(\boldsymbol{Y}_i(1), \boldsymbol{Y}_i(0), \boldsymbol{W}_i): i\in [n]\}$ have a network dependency with $\rho_{ij} = 0$ if there is no edge between $i$ and $j$, (ref) is satisfied if the number of edges is $O(n^{2(1 - q)})$. Note that it imposes no constraint on the maximum degree of the dependency graph. Even if the network is fully connected, it can still hold if the pairwise dependence is weak, e.g., sampling without replacement; see Appendix (ref). On the other hand, (ref) excludes the case where all units are perfectly correlated or equicorrelated with a positive maximal correlation that is bounded away from $0$.
We also impose minimal overlap restrictions on each $\bm{\pi}_i$:
Our final assumption restricts the second moment of outcomes:
It is presented here only for simplicity. We relax it substantially in Appendix (ref).
We consider a class of weighted TWFE regression estimators without covariates. We refer to them as reshaped inverse propensity weighted (RIPW) estimators, and formally define them as follows:
where $\bm{\Pi}(\boldsymbol{w})$ is a density function on $\{0, 1\}^{T}$, i.e.,
We refer to the distribution $\bm{\Pi}$ as a reshaped distribution, and the weight $\bm{\Pi}(\boldsymbol{W}_{i}) / \bm{\pi}_{i}(\boldsymbol{W}_{i})$ as a RIP weight. To ensure that the RIPW estimator is well-defined, we require $\bm{\Pi}$ to be absolutely continuous with respect to each $\bm{\pi}_{i}$, i.e.
The estimator ((ref)) is feasible for any such $\bm{\Pi}$ because $\bm{\pi}_i$ is assumed to be known.
Adding covariates to the objective function ((ref)) is relatively straightforward. However, it considerably complicates the notation without contributing substantially to the primary narrative. We will explicitly incorporate covariates in the objective function in Section (ref). Note that the covariates still play a role in the RIPW estimator through $\bm{\pi}_i$ for covariate-adaptive designs.
The reshaped distribution $\bm{\Pi}$ can be interpreted as an experimental design. If $\boldsymbol{W}_{i}\sim \bm{\Pi}$, then $\bm{\pi}_{i} = \bm{\Pi}$ and (ref) reduces to the standard unweighted TWFE regression. If this is not the case, then $\bm{\Pi}(\boldsymbol{W}_i) / \bm{\pi}_i(\boldsymbol{W}_i)$ acts like a likelihood ratio that changes the original design to one provided by $\bm{\Pi}$. For cross-sectional data, we would like to shift the distribution to uniform $\{0, 1\}$, making the weights equal to $1 / 2\bm{\pi}_i(\boldsymbol{W}_i)$ if the fixed effects are not included. This would yield the standard IPW estimator. However, as we alluded to in the introduction, the situation is more complicated with panel data, and shifting towards the uniform design might not deliver consistent estimators for the DATE of interest. We explore this formally in the next section, where we characterize the set of $\bm{\Pi}$ that one can use. This interpretation of $\bm{\Pi}$ has one caveat: RIP weights only shift the marginal distribution of $\boldsymbol{W}_i$ to $\bm{\Pi}$, but they do not say anything about the joint distribution of $\{\boldsymbol{W}_i,i\in[n]\}$ which can remain complicated.
We now derive sufficient conditions under which the RIPW estimator is a consistent estimator for a given DATE of interest. The following theorem presents a precise condition for consistency of $\hat{\tau}(\bm{\Pi})$ for $\tau^{*}(\xi)$:
This result has two user-specified parameters: time weights $\xi$, and the reshaped distribution $\bm{\Pi}$. They are naturally connected: to guarantee consistency for $\tau^{*}(\xi)$ we can select $\bm{\Pi}$ such that the following holds:
Alternatively, for a given $\bm{\Pi}$, we can look for $\xi$ such that ((ref)) is satisfied. We call (ref) the DATE equation hereafter. For a fixed $\xi$, it is a quadratic system with $\{\bm{\Pi}(\boldsymbol{w}): \boldsymbol{w}\in \{0, 1\}^{T}\}$ being the variables. Together with the density constraint (ref) and the support constraint in Theorem (ref) that $\bm{\Pi}(\boldsymbol{w}) = 0$ for $\boldsymbol{w}\not\in \mathbb{S}^{*}$, there are $T + 1 + 2^{T} - |\mathbb{S}^{*}|$ equality constraints and $|\mathbb{S}^{*}|$ inequality constraints that impose the positivity of $\bm{\Pi}(\boldsymbol{w})$ for each $\boldsymbol{w}\in \mathbb{S}^{*}$. We will show in Appendix (ref) that the DATE equation have closed-form solutions in various examples and provide a generic solver based on nonlinear programming in Appendix (ref).
Without further restrictions on $\bm{\tau}_i$, we can show that the DATE equation is also a necessary condition for consistency of $\hat{\tau}(\bm{\Pi})$ for $\tau^{*}(\xi)$. To see this assume that
for some vector $z$ that is not proportional to $\xi$. Because we can vary individual treatment effects without changing the average one, we can find a set $\{\bm{\tau}_{i}: i\in [n]\}$ that yields the same DATE but $\left\langle z, (1/n)\sum_{i=1}^{n}\left( \bm{\tau}_{i} - \tau^{*}(\xi)\textbf{1}_{T}\right)\right\rangle \not= 0$, leading to inconsistency. For $z=b \xi$ we get that the inner product of the LHS of ((ref)) and $\textbf{1}_{T}$ is $0$ because $\textbf{1}_{T}^\top (\operatorname*{diag}(\boldsymbol{W}) - \xi \boldsymbol{W}^\top) = W^\top (1 - \textbf{1}_T^\top \xi) = 0$, while that of the right-hand side and $\textbf{1}_{T}$ is equal to $b$. This implies that $z$ has to be equal to zero, thus proving the necessity of DATE equation.
Notably, when the DATE equation has a solution, our estimator is consistent without any restrictions on the potential outcomes, except Assumption (ref). This is in sharp contrast to usual results about TWFE estimators, which typically require the trends to be parallel among units, at least conditionally on observed covariates callaway2018difference, sant2020doubly. Theorem (ref) shows that if the assignment process is known and the DATE equation has a solution, we can correct the potentially misspecified TWFE regression model by simply reweighting the objective function. We want to stress that this result relies on the knowledge of the assignment process, whereas the analysis based on conditional parallel trends does not require such knowledge.
To further parse the DATE equation, we discuss two alternative interpretations. First, fix $\xi$ and let $\bm{\Pi}$ be the solution of the DATE equation. Then consider a class of complete randomized experiments where all propensity scores $\bm{\pi}_i$ are identical and are equal to $\bm{\Pi}$. Then, by definition, the RIPW estimator with reshaped distribution $\bm{\Pi}$ reduces to the standard (unweighted) TWFE estimator. Theorem (ref) guarantees that this estimator converges to $\tau^{*}(\xi)$. Since the DATE equation is a necessary condition, all experimental designs that do not satisfy this restriction cannot lead to a consistent estimator for $\tau^{*}(\xi)$. As a result, DATE equation characterizes all complete randomized experiments under which the unweighted two-way estimator converges to a given estimand. This can be interpreted as a general converse of the results established in athey2018design.
As an alternative interpretation, consider a fixed $\bm{\Pi}$ instead. For any such $\bm{\Pi}$ the equation (ref) can be rewritten as
It is easy to see that
where $\tilde{\boldsymbol{W}} = J\boldsymbol{W}$. Since the support of $\bm{\Pi}$ involves a point $\boldsymbol{w}\not\in \{\textbf{0}_{T}, \textbf{1}_{T}\}$, for which $\boldsymbol{w}' \not= 0$ the quantity in (ref) is strictly positive. Therefore, (ref) implies that
By Theorem (ref), in a randomized experiment with $\bm{\pi}_{i} \triangleq \bm{\Pi}$ athey2018design, roth2023efficient, the effective estimand of the unweighted TWFE regression is the DATE with weight vector $\xi$.
The following result shows that the induced weights are guaranteed to be non-negative for arbitrary design.
This result generalizes the conventional cross-sectional logic that says that in randomized experiments, regression estimators are consistent for average effects lin2013agnostic. However, in the case of the TWFE regression, the situation is more nuanced. While the resulting estimand always corresponds to a weighted average effect with non-negative weights, it still depends on the experimental design. As a result, if two analysts were to split a given population into two random subpopulations and conduct two experiments with different designs on each part, the resulting estimands would have been different.
There are two reasons for this unusual behavior. First, in the cross-sectional case, $\boldsymbol{W}_i$ has two points of support, while in the panel case the support of $\boldsymbol{W}_i$ ranges from $2$ to $2^T$ points (as long as Assumption (ref) is satisfied). For example, if none of the units is treated in the first period, it is impossible to identify any DATE that puts positive weight on the first period. Second, fixed effects lead to a familiar incidental parameter problem neyman1948consistent, albeit in a mild form. To see this, consider $ \bm{\pi}_i= \bm{\Pi} \propto 1$, in which case the RIPW estimator corresponds to the conventional TWFE regression. The effective estimand for this regression is equal to the solution of ((ref)) and is different from the effective estimand for the regression without the unit fixed effects.
To enable statistical inference of DATE, we first present an asymptotic expansion showing the asymptotic linearity of RIPW estimators.
Note that the asymptotic linear expansion holds under a fairly general dependency structure in the treatment assignments. Below, we derive a valid confidence intervals for $\tau^*(\xi)$ when $\{(\boldsymbol{Y}_i(1), \boldsymbol{Y}_i(0),\boldsymbol{W}_i): i\in [n]\}$ are independent. The general case is discussed in Appendix (ref). If $\{\mathcal{V}_i: i\in [n]\}$ are well-behaved Theorem (ref) implies that \[\frac{\mathcal{D}\cdot \sqrt{n}(\hat{\tau}(\bm{\Pi}) - \tau^{*}(\xi))}{\sigma_{n}^{*}} \approx N(0, 1), \quad \text{where }\sigma_{n}^{*2} = (1/n)\sum_{i=1}^{n}\mathrm{Var}(\mathcal{V}_i),\] where $\mathcal{D}$ is known by design.\footnote{By well-behaved $\mathcal{V}_i$ we mean that they are sufficiently regular for the appropriate version of the Central Limit Theorem to hold. In the simplest case, when data is i.i.d., this reduces to standard moment restrictions.} If $\{\mathcal{V}_i: i\in [n]\}$ were known, a natural estimator for $\sigma_{n}^{*2}$ would be the empirical variance: \[\hat{\sigma}_{n}^{*2} = \frac{1}{n - 1}\sum_{i=1}^{n}(\mathcal{V}_i - \bar{\mathcal{V}})^2, \quad \text{where }\bar{\mathcal{V}} = \frac{1}{n}\sum_{i=1}^{n}\mathcal{V}_i.\] We should not expect the difference between $\hat{\sigma}_{n}^{*}$ and $\sigma_{n}^{*}$ to converge to zero since $\mathbb{E}[\mathcal{V}_i]$ in general varies over $i$. Nonetheless, $\hat{\sigma}_{n}^{*}$ is an asymptotically conservative estimate of $\sigma_{n}^{*}$ since
where the second term measures the heterogeneity of $\mathbb{E}[\mathcal{V}_i]$ and is always non-negative, implying that $\hat{\sigma}_{n}^{*2}$ is a conservative estimator for $\sigma_{n}^{*2}$. This is unsurprising because even in the cross-section case, the asymptotic design-based variance is only partially identifiable due to the unknown correlation structure between two potential outcomes; see, e.g., Neyman's variance formula neyman23, rubin74.
In general, $\mathcal{V}_i$ is unknown due to $\tau^{*}(\xi)$ and the expectation terms. Nonetheless, we can estimate $\mathcal{V}_i$ by replacing each expectation with the corresponding plug-in estimate, i.e.
and use them to compute the variance:
This yields a Wald-type confidence interval for $\tau^{*}(\xi)$ as
where $z_{\eta}$ is the $\eta$-th quantile of the standard normal distribution. Properties of this confidence interval are established in the next theorem.
In Appendix (ref), we discuss a generic result for general dependent assignments (Theorem (ref)), which covers completely randomized experiments, blocked and matched pair experiments, two-stage randomized experiments, and so on. We present a detailed result (Theorem (ref)) for completely randomized experiments where potential outcomes are fixed and $\boldsymbol{W}_i$'s are sampled without replacement from a user-specified subset of $\{0, 1\}^T$.\footnote{Specifically, given any support $\mathbb{S}^{*}$ and a pre-specified vector $\{n_{\boldsymbol{w}}: \boldsymbol{w}\in \mathbb{S}^*\}$ with $\sum_{\boldsymbol{w}\in \mathbb{S}^*}n_{\boldsymbol{w}} = n$, the experimenter sample assignments $(\boldsymbol{W}_1, \ldots, \boldsymbol{W}_n)$ with probability $\prod_{\boldsymbol{w}\in \mathbb{S}^{*}}n_{\boldsymbol{w}} ! / n!$.} This substantially generalizes the setting of athey2018design and roth2023efficient, where the assignments are sampled without replacement from the set of $T+1$ staggered assignments. At the same time, compared to roth2023efficient, we cannot provide efficiency guarantees for our estimator.
Theorem (ref) and Proposition (ref) might appear counter-intuitive given well-understood problems of TWFE estimators (e.g., de2020two,goodman2018difference,sun2021estimating). To put our result in context, we emphasize two important features of the setup. First, we restrict attention to static models, and second, we use the randomness that is coming from $\boldsymbol{W}_i$. Both of these restrictions play a key role in Theorem (ref). The absence of dynamic effects implies that we can meaningfully average units with different histories of past treatments. A version of this assumption is inescapable if we want the method to work for general designs where controlling for past history is practically infeasible. As we explain below, the randomness of assignments helps to resolve the issue that TWFE estimators put negative weights on some individual treatment effects.
In de2020two,goodman2018difference,sun2021estimating the authors show that treated units are averaged with potentially negative weights, but these results are conditional on the assignments $\boldsymbol{W} = (\boldsymbol{W}_1, \ldots, \boldsymbol{W}_n)$ being fixed. Let $\xi_{it}(\gamma; \boldsymbol{W})$ be these weights for the general weighted least squares estimator $\hat{\tau}(\gamma)$ defined in (ref) such that \[\mathbb{E}[\hat{\tau}(\gamma)\mid \boldsymbol{W}] = \sum_{i=1}^{n}\sum_{t=1}^{T}\xi_{it}(\gamma; \boldsymbol{W})\tau_{it},\] where we now explicitly allow the weights to depend on $\boldsymbol{W}$. When the assignments are treated as random, the large sample limit of $\hat{\tau}(\gamma)$ is \[\mathbb{E}[\hat{\tau}(\gamma)] = \sum_{i=1}^{n}\sum_{t=1}^{T}\xi_{it}(\gamma)\tau_{it},\] where $\xi_{it}(\gamma) = \mathbb{E}_{\boldsymbol{W}}[\xi_{it}(\gamma; \boldsymbol{W})]$. While $\{(i, t): \xi_{it}(\gamma; \boldsymbol{W}) < 0\}$ is non-empty almost surely for every realization of $\boldsymbol{W}$, it is still possible that all $\xi_{it}(\gamma)$ are positive due to the averaging over $\boldsymbol{W}$. For illustration, we consider a simulation study with $n = 100, T = 4$ and other details specified in Section (ref). We consider the conditional and unconditional weights induced by the unweighted and RIP-weighted TWFE estimator in Figure (ref) and Figure (ref), respectively. We plot the histograms of $\{(nT)\cdot\xi_{it}(\gamma; \boldsymbol{W}): i\in [n], t \in [T]\}$ for three realizations of $\boldsymbol{W}$ and the histogram of $\{(nT)\cdot \xi_{it}(\gamma): i\in [n], t\in [T]\}$, approximately by averaging over a million realizations of $\boldsymbol{W}$, where the multiplicative factor $nT$ is chosen to normalize the weights into a more interpretable scale. Clearly, despite the large fraction of negative weights in each realization, their averages do not have any negatives. Therefore, the criticism on TWFE estimators does not apply in this case. Indeed, it never applies to the RIPW estimator by Proposition (ref). In this study, all weights are designed to be $1/nT > 0$ when $\bm{\Pi}$ is a solution of the DATE equation with $\xi = \textbf{1}_{T}/T$, as shown in Figure (ref)(b), regardless of the data generating process.
The discussion above demonstrates that while for each cell $(i,t)$, a particular realization of weights can be negative, this fact is not systematic. If we use the RIPW estimator designed for the equally weighted DATE, then all cells will receive the same weight on average. An alternative description of the same phenomenon is that once correctly weighted, the realized treatment paths $\boldsymbol{W}_i$ are independent of potential outcomes. This independence implies that there cannot be systematic differences in treatment effects among units with distinct assignment paths, and thus negative weights do not create complications for the interpretation of the estimates. As we illustrate in Section (ref), this interpretation remains valid even when certain dynamic effects are present.
In this section, we move to non-experimental settings where the assignment mechanism is not controlled by the researcher and is unknown. We assume that researchers constructed unit-level estimates $\{\hat{\bm{\pi}}_i, i \in [n]\}$. In addition, we assume that the researchers have access to a set of estimates $\{(\hat{\bm{\mu}}_{i}(0), \hat{\bm{\mu}}_{i}(1)): i\in [n]\}$ of $\{(\mathbb{E}[\boldsymbol{Y}_{i}(0)], \mathbb{E}[\boldsymbol{Y}_{i}(1)]): i\in [n]\}$. Further, let $\hat{m}_{it}$ be the double-centered version of $\hat{\mu}_{it}(0)$ and $\hat{\nu}_{it}$ be a shifted version of $\hat{\mu}_{it}(1) - \hat{\mu}_{it}(0)$:
For notational convenience, we write $\hat{\boldsymbol{m}}_i$ for the vector $(\hat{m}_{i1}, \ldots, \hat{m}_{iT})$ and $\hat{\bm{\nu}}_i$ for the vector $(\hat{\nu}_{i1}, \ldots, \hat{\nu}_{iT})$. Given a set of estimates $\{(\hat{\bm{\pi}}_i, \hat{\boldsymbol{m}}_{i}, \hat{\bm{\nu}}_i): i\in [n]\}$, we define the RIPW estimator as
The above estimator generalizes (ref) by allowing for regression adjustment. Throughout the rest of the paper, we will abuse the notation by denoting it as $\hat{\tau}(\bm{\Pi})$. This two-stage formulation replaces the regression with covariates by regression on the modified outcome $(Y_{it} - \hat{m}_{it} - \hat{\nu}_{it}W_{it})$ without covariates, yielding a simplified structure which allows us to use previously established results. In the rest of this section, we discuss formal properties they need to satisfy to guarantee consistency and asymptotic normality of $\hat{\tau}(\bm{\Pi})$.
In the previous section, we assumed that the researcher controlled the assignment process, which led to the restriction ((ref)). In observational studies, the assignment process is unknown, and we must substitute this restriction with a different assumption. Throughout this section, we impose a high-level restriction on the relationship between unit-specific potential outcomes and assignment paths.
Recall that we do not assume that $(\boldsymbol{Y}_i(1), \boldsymbol{Y}_i(0),\boldsymbol{W}_i)$ are identically distributed across units. As a result, Assumption (ref) imposes $n$ separate restrictions, one for each unit. It follows the tradition of the part of the panel data literature that treats unit-specific unobservables as fixed parameters lancaster2000incidental,hahn2004jackknife, rather than random variables as in chamberlain1984panel. It is trivially satisfied in an extreme case where $(\boldsymbol{Y}_i(1), \boldsymbol{Y}_i(0))$ has a degenerate distribution for each $i \in [n]$, which corresponds to the finite population analysis abadie2020sampling. In applications where $(\boldsymbol{Y}_i(1), \boldsymbol{Y}_i(0))$ is random, this assumption imposes a strict exogeneity restriction. It describes the average behavior of the outcomes conditional on the whole treatment path and does not allow the current treatment to depend on past outcomes. To illustrate this connection, consider the classical linear TWFE model where
Assumption (ref) is equivalent to $\mathbb{E}[\epsilon_{it}\mid \boldsymbol{W}_i] = 0$ for $t = 1, \ldots, T$, which is a strict exogeneity restriction. In contrast, if $\{\epsilon_{it}: i\in [n], t\in [T]\}$ only satisfies contemporenous restrictions $\mathbb{E}[\epsilon_{it}\mid W_{it}] = 0$, Assumption (ref) does not necessarily hold. We want to note that in the DiD literature it is common to impose restrictions only on $\boldsymbol{Y}_i(0)$, while Assumption (ref) restricts both potential outcomes. This is necessary given our focus on the ATE, defined in Section (ref).
Assumption (ref) is also related to the recent cross-sectional literature on quasi-experimental designs borusyak2023non. A typical restriction in that literature is that while the distribution of the treatment of interest varies over units in a complicated way it still can be estimated and then used to construct counterfactuals. For this approach to be valid, one needs to impose a version of Assumption (ref). In the panel data literature, this type of quasi-experimental variation was also exploited. For example, wojtaszek2022sensitivity studied the effect of military bonuses on charitable giving and found that the timing of receiving the bonus is (nearly) as-if random. Depending on the choice of outcome variable, the bonus can be viewed as a staggered or one-off treatment with a uniform generalized propensity score.
To construct estimators $\{(\hat{\bm{\pi}}_i, \hat{\boldsymbol{m}}_{i}, \hat{\bm{\nu}}_i): i\in [n]\}$ we use the observed covariates $\{\boldsymbol{X}_i, i \in [n]\}$. Our assumptions implicitly restrict the set of feasible covariates. In particular to respect Assumption (ref), we do not allow any parts of the observed outcomes $\boldsymbol{Y}_i$ to be used as covariates. The situation is more delicate for $\boldsymbol{W}_i$, and we allow functions of $\boldsymbol{W}_i$ to be part of $\boldsymbol{X}_i$ as long as Assumption (ref) holds. We elaborate on this in the next two sections.
In strictly exogenous panel models, the distribution of $\boldsymbol{W}_i$ is commonly left unspecified and the analysis is based on the outcome model alone. In particular, the distribution of $\boldsymbol{W}_i$ can be degenerate for each $i \in [n]$, which is another extreme case where Assumption (ref) trivially holds. However, researchers often informally appeal to random or quasi-random variation in $\boldsymbol{W}_i$ as a source of identification, even though they continue using outcome-based methods, such as the TWFE regression. We interpret these informal statements as statistical restrictions on $\{\bm{\pi}_i, i \in [n]\}$ that go beyond Assumption (ref).
Precisely because the arguments used in the applied work are often informal, we cannot offer and analyze a general methodology of how to use them to construct $\{\hat{\bm{\pi}}_i, i \in [n]\}$. Instead, we discuss several strategies that are potentially relevant for a large class of applications. Our goal is to demonstrate how to utilize the information used to construct the outcome-based estimators and thus is readily available. In practice, researchers can have other sources of information that we do not incorporate in our analysis. After this discussion, we continue our formal analysis under high-level assumptions on $\{\hat{\bm{\pi}}_i, i \in [n]\}$.
We use $\{(\boldsymbol{W}_i, \boldsymbol{X}_i): i\in [n]\}$ to estimate $\bm{\pi}_i$. At first glance, it might appear to be challenging to estimate the distribution of the whole vector. Nevertheless, treatment paths often have restricted support with a size much smaller than $2^T$, such as staggered adoption and/or special structures that reduce the complexity of the distribution, such as the Markov structure. We present a few examples below for illustration.
In the staggered adoption designs $\boldsymbol{W}_i$ is equivalent to an adoption time $A_i \in \{1, \ldots, T, \infty\}$, where $A_i = \infty$ for never-treated units and $A_i = t$ for units initially treated at time $t$. Then $A_i$ can be viewed as an event or during outcome, and one can apply any survival or duration model, such as the Cox proportional hazard model and accelerated failure time model, to estimate its distribution which yields $\bm{\pi}_i$ by taking the difference between the consecutive points; see Section (ref) for an empirical illustration that uses this strategy and additional discussion. For transient treatments that occur at most once during the study period, $\boldsymbol{W}_i$ can be expressed by the adoption time $A_i \in \{1, \ldots, T, \infty\}$ as above. The propensity score $\bm{\pi}_i$ can then be estimated via a discrete choice model.
For general designs where the treatment can be alternated on and off, $\bm{\pi}_i$ can be reparametrized as a sequence of conditional distributions $\mathbb{P}(W_{it}\mid W_{i(t-1)}, \ldots, W_{i1}, \boldsymbol{X}_{i})$ and estimated by a Markov model. In particular, arkhangelsky2019double show that if $\boldsymbol{X}_i$ incorporates appropriate sufficient statistics, then conditioning on $\boldsymbol{X}_i$ eliminates the ex-ante present unobserved heterogeneity from the distribution of $\boldsymbol{W}_i$. aguirregabiria2021sufficient show that these assumptions are satisfied by a large class of models that are widely used in economic applications, including structural models with forward-looking agents as well as myopic (backward-looking) dynamic logit models. They also provide explicit characterizations for sufficient statistics in such models. These results can be directly applied in our setting.
Given an estimate $\hat{\bm{\pi}}_i$, we say that it estimates the assignment model well if $\hat{\bm{\pi}}_{i}$ is close to $\bm{\pi}_{i}$ in $L^2$ distance. Specifically, for each unit $i$ we define the accuracy of $\hat{\bm{\pi}}_{i}$ as
Here, the expectation is taken over both $\boldsymbol{W}_i$ and $\hat{\bm{\pi}}_i$ (conditional on $\{\boldsymbol{X}_i: i\in [n]\}$). In the setting of Section (ref), $\delta_{\pi i} = 0$ because $\hat{\bm{\pi}}_i = \bm{\pi}_{i}$.
In this section, we discuss the construction of the terms $\{(\hat{\boldsymbol{m}}_{i}, \hat{\bm{\nu}}_i): i\in [n]\}$ which we use to build the estimator ((ref)). We start with unit specific quantities $(\hat{\bm{\mu}}_{i}(0), \hat{\bm{\mu}}_i(1))$, which we view as estimators for $(\mathbb{E}[\boldsymbol{Y}_i(0)], \mathbb{E}[\boldsymbol{Y}_i(1)])$. There are many ways of constructing such estimators, and our results require only high-level restrictions on these objects. For example, one can consider a generalization of the linear TWFE model (ref):
Then $(\hat{\mu}_{it}(0), \hat{\mu}_{it}(1))$ can be chosen as
where the parameters are estimated by regressing $Y_{it}$ on $X_{it}, W_{it}$, the covariate-treatment interaction $X_{it}W_{it}$, and a set of fixed effects. When we estimate $\hat{\mu}_{it}(w)$ for a new unit whose unit fixed effect is not estimated, we can simply set $\hat{\mu}_{it}(w) = \hat{\mu} + \hat{\lambda}_t + X_{it}^\top \hat{\beta} + (\hat{\tau} +X_{it}^\top \hat{\phi})w$.
In the cross-sectional case, an estimate $(\hat{\bm{\mu}}_{i}(0), \hat{\bm{\mu}}_i(1))$ is considered an accurate estimate of $(\mathbb{E}[\boldsymbol{Y}_{i}(0)], \mathbb{E}[\boldsymbol{Y}_i(1)])$ if $\{\|\hat{\bm{\mu}}_{i}(0) - \mathbb{E}[\boldsymbol{Y}_{i}(0)]\|_{2} + \|\hat{\bm{\mu}}_i(1) - \mathbb{E}[\boldsymbol{Y}_i(1)]\|_{2}: i\in [n]\}$ is small on average robins1994estimation, kang2007demystifying. Constructing such estimators for panel models with fixed effects and a finite number of periods is impossible. Thus, the standard approach of measuring accuracy does not apply in our setting, and we need to consider alternative measures.
We start by defining the estimands $(m_{it}, \nu_{it})$ that $(\hat{m}_{it}, \hat{\nu}_{it})$ attempt to estimate:
To measure the degree of mispecification of the outcome model we introduce the following quantity:
The first term captures the estimation accuracy of $\boldsymbol{m}_i$, and the second term captures the estimation accuracy of $\bm{\tau}_i$. By definition, $\delta_{yi}$ is invariant if we replace $\mathbb{E}[Y_{it}(0)]$ by $\mathbb{E}[Y_{it}(0)] + \mu' + \alpha_i' + \lambda_t'$ and $\tau_{it}$ by $\tau_{it} + \tau'$ for any $\mu', \tau', \{\alpha_i': i\in [n]\}$, and $\{\lambda_t': t\in [T]\}$. Thus, requiring $\delta_{yi}$ to be small is strictly less stringent than requiring the standard measure of outcome model accuracy for cross-sectional data to be small.
In the simplest TWFE model (ref) without covariates, $\delta_{yi} = 0$ if we choose $\hat{\mu}_{it}(0) = \hat{\mu}_{it}(1) = 0$. For the more general TWFE model (ref), regardless whether unit $i$ is used for fitting the TWFE regression, \[\delta_{yi} = \sqrt{\mathbb{E}\left[\sum_{t=1}^{T}\{(X_{it} - \bar{X}_{i\cdot} - \bar{X}_{\cdot t} + \bar{X}_{\cdot \cdot})^\top (\hat{\beta} - \beta)\}^2 + \{(X_{it} - \sum_{t=1}^{T}\xi_{t}\bar{X}_{\cdot t})^\top (\hat{\phi} - \phi)\}^2\right]},\] Standard assumptions arellano2003panel, wooldridge2010econometric guarantee that $(\hat{\beta}, \hat{\phi})$ are consistent for $(\beta, \phi)$ even with a finite number of periods. We can further generalize the model by replacing $X_{it}^\top\beta$ and $X_{it}^\top \phi$ with nonlinear functions $g(X_{it})$ and $\tau(X_{it})$ and estimate them by nonparametric TWFE regressions boneva2015semiparametric.
The requirement that $\delta_{yi} \approx 0$, at least on average, puts restrictions on the treatment effects. These requirements, however, can be redundant, depending on the structure of $\boldsymbol{X}_i$. For example, wooldridge2021two shows that the problems with heterogeneous treatment effects can be solved, under conditional parallel trends and linearity, by including a sufficiently rich set of controls, which includes functions of $\boldsymbol{W}_i$. In the staggered adoption case, one needs to include interactions with all the adoption dates. Unfortunately, including such interactions into $\boldsymbol{X}_i$ violates the overlap assumption (ref).
In this and the next subsection, we consider a simplified case where the estimates $\{(\hat{\bm{\pi}}_i, \hat{\boldsymbol{m}}_{i}, \hat{\bm{\nu}}_i): i\in [n]\}$ are independent of the data (e.g., obtained from external data). While this is not always possible in practice, the theory of consistency and asymptotic normality can be stated without much mathematical complication. Moreover, these results are building blocks for the theory of cross-fitting estimator described at length in Appendix (ref). To ease implementation, we provide a self-contained description of the (derandomized) cross-fitting RIPW estimator in Algorithm (ref) at the end of the next subsection.
Theorem (ref) implies that the RIPW estimator with $\bm{\Pi}$ being a solution of the DATE equation, if any, is a consistent estimator of DATE without any outcome model when $\hat{\bm{\pi}}_i = \bm{\pi}_{i}$ is known. On the other hand, when the outcome model is correctly specified, $Y_{it} - \hat{m}_{it} - \hat{\nu}_{it}W_{it}\approx Y_{it} - m_{it} - (\tau_{it} - \tau^{*}(\xi))W_{it}$ is a linear model with two-way fixed effects and a single predictor $W_{it}$ and $\hat{\tau}$ is approximately a weighted least squares estimator which is consistent under mild conditions on the weights wooldridge2010econometric. This shows a weak double robustness property that $\hat{\tau}(\bm{\Pi})$ is consistent if either the outcome model or the assignment model is exactly correct.
For cross-sectional data, the augmented IPW estimator enjoys a strong double robustness property, which states that the asymptotic bias is the product of estimation errors of the outcome and assignment models robins1994estimation, kang2007demystifying, chernozhukov2017double, chernozhukov2018double. Clearly, this implies the weak double robustness. It further implies the estimator has higher asymptotic precision than estimators based on merely the outcome or assignment modeling when both models are estimated well. The next result provides a sufficient condition for strong double robustness of $\hat{\tau}(\bm{\Pi})$ when the estimated treatment and outcome models are independent of the data.
Assumptions (ref) and (ref) guarantee that $\bar{\delta}_y$ is bounded. Thus, the RIPW estimator is consistent whenever $\bm{\pi}_i$ is consistently estimated without any requirement on the rate of convergence. On the other hand, under the TWFE model (ref) or nonparametric TWFE models discussed in the last subsection, $\bar{\delta}_{y} = o(1)$ and the estimator is consistent even if the assignment model is globally misspecified.
Similar to Theorem (ref), we can derive an asymptotic linear expansion for $\mathcal{D}\cdot \sqrt{n}(\hat{\tau}(\bm{\Pi}) - \tau^{*}(\xi))$.
Similar to Section (ref), we can estimate $\hat{\mathcal{V}}_i$ and the asymptotic variance via (ref) and construct the Wald-type confidence interval as (ref) when units are independent. This is a special case of Theorem (ref) in Appendix (ref) for general dependent designs.
Under Assumption (ref), Theorem (ref) and Theorem (ref) strictly generalize Theorem (ref) and Theorem (ref) -- when $\bm{\pi}_i$ is known, $\bar{\delta}_{\pi} = 0$ and hence $\bar{\delta}_{\pi}\bar{\delta}_{y} = 0 = o(1/\sqrt{n})$ regardless of the accuracy of the outcome model estimates. When $\bm{\pi}_i$ is unknown, $\bar{\delta}_{\pi}$ and $\bar{\delta}_{y}$ are typically no less than $O(1/\sqrt{n})$ without external data. As a result, both models should be consistently estimated to achieve $\bar{\delta}_{\pi}\bar{\delta}_{y} = o(1/\sqrt{n})$ though the estimates can have a slower convergence rate than $O(1/\sqrt{n})$. For example, it would be satisfied if $\bar{\delta}_{\pi}, \bar{\delta}_{y} = o(n^{-1/4})$. We emphasize that this rate requirement is standard for inference with cross-sectional data chernozhukov2017double, chernozhukov2018double. Under this rate condition, by virtue of the asymptotic linear expansion in Theorem (ref), the researcher can safely ignore the variability of the model estimates and use them in the variance calculation as if they are the truth.
Even when condition $\bar{\delta}_{\pi}\bar{\delta}_{y} = o(1/\sqrt{n})$ is violated, the asymptotically valid inference is still possible at the cost of more involved variance estimation. Doubly robust inference in this regime is generally hard benkeser2017doubly. We consider the setting where parametric models are used to fit the generalized propensity score and regression adjustment. This setting has been studied in the literature for cross-sectional data cao2009improving. Our formal results are deferred in Appendix (ref) due to the mathematical complication. Roughly speaking, if the estimators $\{(\hat{\bm{\pi}}_i, \hat{\boldsymbol{m}}_i, \hat{\bm{\nu}}_i): i \in [n]\}$ come from a smooth parametric model, then one can use their asymptotic expansion (around their limits, which do not necessarily correspond to the true parameters) to compute the asymptotic variance. Similar to the case discussed Section (ref) and this section, we can obtain an asymptotically conservative variance estimator without knowing which model is misspecified apriori.
In practice, it is uncommon to obtain estimates of $(\hat{\bm{\pi}}_i, \hat{\boldsymbol{m}}_i, \hat{\bm{\nu}}_i)$ that are independent of the data, except in the design-based inference where $\hat{\bm{\pi}}_i = \bm{\pi}_i$ and $\hat{\boldsymbol{m}}_i = \hat{\bm{\nu}}_i = \textbf{0}_{T}$, or when external data is available. Usually, these parameters need to be estimated from the data. The resulting dependence invalidates the assumptions of Theorem (ref) and (ref). However, as we show in Appendix (ref) similar results hold if we use a particular version of cross-fitting. Note that this implies that $\hat{\bm{\pi}}_i, \hat{\boldsymbol{m}}_i, \hat{\bm{\nu}}_i$ cannot contain unit-specific fixed effects. Moreover, we propose a simple approach to mitigate the randomness introduced by sample splitting. We describe the estimator in Algorithm (ref). More details can be found in Appendix (ref). We implemented this method in an R package ripw that is available at \url{https://github.com/lihualei71/ripw}.
A key limitation of our analysis in previous sections is the focus on static models. This is important both theoretically and practically. Theoretically, some policies of interest are transient in nature, e.g., a large infrastructure investment, but policymakers expect them to have a lasting impact, which requires a dynamic model. Practically, a large part of applied work in economics uses regression models that explicitly incorporate lags of treatment variables.
We consider a relatively simple class of linear potential outcome modes to address these concerns. For every $i$ and $t$, we specify the potential outcomes as a function of the current treatment $w$ and its $p$ lags:
As in Section (ref), the expectation is conditional on covariates, and we do not require the units to be independent or identically distributed. This model does not restrict the baseline outcomes but puts structure on the dynamic effects of the treatment. First, the effect of the treatment is present only for $p$ periods after it is implemented. Second, the effect is linear, i.e., the causal effect of being treated one period ago, $w_{-1}$, does not depend on whether the unit was treated two periods ago $w_{-2}$. Finally, the effects are homogenous over time, meaning that $\tau_{i,l}$ do not depend on calendar time $t$. These restrictions are important: the first eliminates the possibility of long-term effects, while the other two eliminate state dependence. Still, we think this model is flexible enough to be useful for a large class of empirical applications.
Interestingly, if the treatment timing is fixed and common across units, and $p$ is large enough, then (ref) is a parametrization of all realizable potential outcomes and thus does not impose any testable restrictions. To see this, let $q+1$ denote the adoption time and set $p = T-q+1$. Then each unit $i$ has $T + (T - q)$ potential outcomes $\{Y_{it}(\textbf{0}_{T}): t\in [T]\}$ and $\{Y_{it}(\textbf{0}_{q}, \textbf{1}_{T-q}): t \in \{q+1, \ldots, T\}\}$. It is easy to see that (ref) holds with $\mu_{it} = \mathbb{E}[Y_{it}(\textbf{0}_{T})]$, $\tau_{i, 0} = \mathbb{E}[Y_{i(q+1)}(\textbf{0}_{q}, \textbf{1}_{T-q})] - \mathbb{E}[Y_{i(q+1)}(\textbf{0}_{T})]$, and \[\tau_{i, -\ell} = \mathbb{E}[Y_{i(q+\ell+1)}(\textbf{0}_{q}, \textbf{1}_{T-q})] - \mathbb{E}[Y_{i(q+\ell+1)}(\textbf{0}_{T})] - (\mathbb{E}[Y_{i(q+\ell)}(\textbf{0}_{q}, \textbf{1}_{T-q})] - \mathbb{E}[Y_{i(q+\ell)}(\textbf{0}_{T})]), \quad \ell \in [p].\] Similar logic extends to staggered adoption designs as long as we treat the assignment as fixed. However, it breaks if we assume that the adoption time is randomly assigned. In this case, we can test the static model from Section (ref) and the dynamic model ((ref)) by comparing outcomes across units that were previously treated at different periods. This emphasizes the importance of the assignment model for the analysis of dynamic effects.
In this case, it is natural to consider the RIPW estimator coupled with an event-study regression model, i.e.,
where $W_{it}$ is defined as $0$ whenever $t \le 0$. Our next result describes the probability limit of $(\hat{\tau}_{0}, \hat{\tau}_{-1}, \ldots, \hat{\tau}_{-p})$. The proof is presented in Appendix (ref).
This result justifies using the RIPW estimator in a large class of applications. If $\bm{\pi}_i$-s are unknown, then one can estimate them using one of the strategies discussed in the previous section. Similarly, one can introduce covariates in this model in the same way as before. Also, applied researchers often consider leads in addition to lags in their regressions, especially in the context of staggered adoption designs. To incorporate this practice into our framework, one simply needs to shift the treatment path $\boldsymbol{W}_{i}$ appropriately. The resulting estimators for the leads can then be used to test for the validity of the underlying model.
We do not establish analogs of Theorems (ref) - (ref) for this estimator, but we expect them to hold under appropriate technical conditions. In particular, under (ref), if the TWFE model holds for the baseline potential outcomes such that $\mu_{it} = \mu + \alpha_{i} + \lambda_{t}$ and $\tau_{i, -\ell} = \tau_{-\ell}$, then (ref) is consistent for $(\tau_0, \tau_{-1}, \ldots, \tau_{-p})$ since it is a weighted least squares estimator for a correctly specified linear model. Compared to our analysis in previous sections, the reshaping distribution $\bm{\Pi}$ does not play a major role in these results. The reason for this behavior is that the model for treatment effects is time-homogeneous. If we relax this assumption and allow for time-varying dynamic effects $\tau_{i,-l,t}$, then the distribution $\bm{\Pi}$ becomes important again. The corresponding DATE equation for this problem is more complicated than the one presented in Section (ref), and its analysis is beyond the scope of this paper.
In this section, we investigate the properties of our estimator in simulations and show how to apply it to real datasets. The R programs to replicate all results in this section is available at \url{https://github.com/xiaomanluo/ripwPaper}.
To highlight the central role of the reshaping function in eliminating the bias, we focus on inference with known assignment mechanisms. Put another way, in such settings, the bias of the unweighted or IPW estimators is purely driven by the wrong reshaping function rather than other sources of variability. We consider the DATE with $\xi = \textbf{1}_{T} / T$ for simplicity. We also design a simulation study with unknown assignment mechanisms and present the results in Appendix (ref), which involves all $2$-by-$2$ settings with correct/incorrect assignment/outcome model and a detailed comparison between the RIPW estimator and several other competing estimators.
We consider a short panel with $T = 4$ and sample size $n = 1000$. We generate a single time-invariant covariate $X_{it} = X_{i}$ with $P(X_i = 1) = 0.7$ and $P(X_i = 2) = 0.3$ and a single time-invariant unobserved confounder $U_{it} = U_{i}$ with $U_i\sim \mathrm{Unif}(\{1, \ldots, 10\})$. Within each experiment, the covariates and unobserved confounders are only generated once and then fixed to ensure a fixed design. For treatment assignments, we consider a staggered adoption design, i.e., $\boldsymbol{W}_i\in \mathcal{W}^{\mathrm{sta}}$. We assume that $\boldsymbol{W}_i$ is less likely to be treated when $X_i = 1$. In particular, \[\big(\bm{\pi}_i(\boldsymbol{w}_{(0)}), \bm{\pi}_i(\boldsymbol{w}_{(1)}), \bm{\pi}_i(\boldsymbol{w}_{(2)}), \bm{\pi}_i(\boldsymbol{w}_{(3)}), \bm{\pi}_i(\boldsymbol{w}_{(4)})\big) = \left\{
\right..\]
The potential outcome $Y_{it}(0)$ and the treatment effect $\tau_{it}$ are generated as follows: \[Y_{it}(0) = \mu + \alpha_i + \lambda_t + m_{it} + \epsilon_{it}, \quad m_{it} = \sigma_{m}X_i\beta_{t}, \quad \tau_{it} = \sigma_{\tau} a_{i}b_{t},\] where $\mu = 0$, $\beta_t = t - 1$, $\alpha_i = 0.5U_i$, $\lambda_t\stackrel{i.i.d.}{\sim} \mathcal{N}(0,1)$, $b_t \stackrel{i.i.d.}{\sim} \mathcal{N}(0,1)$, and $\epsilon_{it} \stackrel{i.i.d.}{\sim} N(0, 1)$. For $a_i$, we consider two settings: we either set $a_i = 1$ thus making $\tau_{it}$ unit-invariant; or $a_{i}\stackrel{i.i.d.}{\sim}\mathrm{Unif}([0, 1])$, in which case $\tau_{it}$ varies over units and periods. As with the covariates $X_i$, the time fixed effects $\lambda_t$ and factors $a_{i}, b_t$ are generated once for each setting and then fixed over runs. In contrast, $\epsilon_{it}$ will be resampled in every run as the stochastic errors. Note that both $m_{it}$ and $\tau_{it}$ are generated from rank-one factor models.
The parameters $\sigma_{m}$ and $\sigma_{\tau}$ measures two types of deviations from the TWFE model: $\sigma_{m}$ measures the violation of parallel trend because we will not adjust for $X_i$ in the design-based inference, and $\sigma_{\tau}$ measures the violation of constant treatment effects. We consider two settings: we either set $\sigma_{m} = 1, \sigma_{\tau} = 0$ --- a model without parallel trends, but constant treatment effects; alternatively, we set $\sigma_{m} = 0, \sigma_{\tau} = 1$ --- a TWFE model with heterogeneous effects, but parallel trends. In the first setting $\tau_{it} = 0$ regardless of the model for $a_i$, thus, we have $3$ different scenarios in total.
We consider three estimators: the unweighted TWFE estimator, the IPW estimator, and the RIPW estimator with $\bm{\Pi}$ given by (ref). For each of the three experiments, we resample $W_{it}$'s and $\epsilon_{it}$'s, while keeping other quantities fixed, for $1000$ times and collect the estimates and the confidence intervals. Figure (ref) presents the boxplots of the bias $\hat{\tau}(\bm{\Pi}) - \tau^{*}(\xi)$. In all settings, the unweighted estimator is clearly biased, demonstrating that both the parallel trend and treatment effect homogeneity are indispensible for classical TWFE regression. In contrast, the IPW estimator is biased when the treatment effects are heterogeneous, but unbiased otherwise even if the parallel trend assumption is violated. This is by no means a coincidence; in this case, $\bm{\tau}_i = \tau^{*}(\xi)\textbf{1}_{T}$ for all $i$ and, by Theorem (ref), the asymptotic bias $\Delta_{\tau}(\xi) = 0$ for RIPW estimators with any reshaped function including the IPW estimator. Finally, as implied by our theory, the RIPW estimator is unbiased in all settings. Moreover, the coverage of confidence intervals for the RIPW estimator is $94.6\%, 95.2\%$, and $94.6\%$ in these three settings, respectively, confirming the inferential validity stated in Theorem (ref).
On February 29th, 2020, Washington declared a state of emergency in response to the COVID-19 pandemic. A state of emergency is a situation in which a government is empowered to perform actions or impose policies that it would normally not be permitted to undertake. It alerts citizens to change their behaviors and urges government agencies to implement emergency plans. As the pandemic has swept across the country, more states declared a state of emergency in response to the COVID-19 outbreak.
The state of emergency restricts various human activities. It would be valuable for governments and policymakers to get a sense of the short-term effect of this urgent action. Since mid-February 2020, OpenTable has been releasing daily data of year-over-year seated diners for a sample of restaurants on the OpenTable network through online reservations, phone reservations, and walk-ins. This provides an opportunity to study how the state of emergency affects the restaurant industry in a short time. The data covers 36 states in the United States, which we will focus our analysis on. Policy evaluation in the pandemic is extremely challenging due to the complex confounding and endogeneity issues chetty2020did, chinazzi2020effect, goodman2020using, holtz2020interdependence, kraemer2020effect, abouk2021immediate. Fortunately, compared to the policies later in the pandemic, the state of emergency suffered from less confounding since it was the first policy that affected the vast majority of the public in the US. On the other hand, the restaurant industry is responding to the policy swiftly because the restaurants are forced to limit and change operations, thereby eliminating some confounders that cannot take effect in a few days.
Despite being more approachable, the problem remains challenging due to the effect heterogeneity and the difficulty of building a reliable model for the dine-in rates in a short time window. In contrast, the declaration time of the state of emergency is arguably less complex to model because it is mainly driven by the progress of the pandemic and the authority's attitude towards the pandemic.
We demonstrate our RIPW estimator on this data. The summary statistics and data sources can be found in Appendix (ref). The outcome variable is the daily state-level year-over-year percentage change in seated diners provided by OpenTable. The treatment variable is the indicator of whether the state of emergency has been declared. We also include the state-level accumulated confirmed cases to measure the progress of the pandemic, the vote share of Democrats based on the 2016 presidential election data to measure the political attitude towards COVID-19, and the number of hospital beds as a proxy for the amount of regular medical resources. For demonstration purposes, we restrict the analysis to February 29th -- March 13th, the first 14 days since the first declaration by Washington. As of March 13th, 34 out of 36 states have declared a state of emergency; thus, the declaration times are right-censored. The treatment paths are plotted in Figure (ref).
For the treatment model, we fit a Cox proportional hazard model on the declaration date to derive an estimate of the generalized propensity scores. Specifically, letting $T_i$ be declaration time of state $i$, a Cox proportional hazard model with time-varying covariates $X_{it}$ assumes that \[h_i(t\mid X_{it}) = h_0(t)\exp\{X_{it}^\top \beta\}\] where $h_i(t\mid \cdot)$ denotes the hazard function for state $i$, and $h_0(t)$ denotes a nonparametric baseline hazard function. The estimates $\hat{h}_0$ and $\hat{\beta}$ yield an estimate $\hat{F}_i(t)$ of the survival function $\mathbb{P}(T_i \ge t)$ for state $i$, differencing which yields an estimate of the generalized propensity score \footnote{For discrete event times, an alternative is the discrete Cox model introduced in Section 6 of cox1972regression. Here, we stick with the standard Cox model for simplicity.} \[\hat{\bm{\pi}}_i(\boldsymbol{W}_i) = \left\{
\right..\] Here, we include as the time-varying covariates the logarithms of the accumulated confirmed cases and as the time-invariant covariates the logarithms of the number of hospital beds and the vote share. Note that fixed effects cannot be added into the Cox model because each state has only one outcome. To address unobserved heterogeneity, we include region fixed effects (Northeast, North Central, South, and West). While we will cross-fit the Cox model for the RIPW estimator, we fit the model on the entire data to illustrate the effect of covariates on the adoption time. Table (ref) summarizes the exponentiated parameter estimates along with their standard errors with and without region fixed effects. It also reports the p-value of the joint significance test for the null hypothesis that all coefficients are zero. While most of the coefficients are not significant individually, they are jointly significant, suggesting that the generalized propensity score is non-constant.
The proportional hazard assumption imposed by the Cox model is often controversial. Here, we apply the standard statistical tests based on Schoenfeld residuals schoenfeld1980chi as a specification test for the Cox model. Figure (ref) presents the p-values yielded by Schoenfeld's test. Clearly, none of them show evidence against the proportional hazard assumption. The p-value of Schoenfeld's test is $0.311$, suggesting no evidence against the specification.
For the outcome model, we fit an interacted TWFE regression in the form of (ref) with the same set of covariates. Since unit fixed effects are included, no time-invariant covariate can be added to the main effects due to perfect collinearity. Thus, we add log confirmed cases, treatment, and the interactions between treatment and all variables, including region fixed effects, into the TWFE regression. Table (ref) summarizes the results. The first row gives the treatment effect estimates by the TWFE regressions, though these estimates are irrelevant in the regression adjustment for our RIPW estimator, which only depends on the other rows. Table (ref) reports the p-value of the joint significance test for the null hypothesis that all coefficients other than the two-way fixed effects are zero. Again, the null hypothesis that $m_{it} = 0$ for all $(i, t)$ is rejected in both settings.
Finally, we compute the RIPW estimator for equally-weighted DATE with the reshaped distribution (ref) in Appendix (ref) for staggered adoption and $10$-fold cross-fitting that is described in Algorithm (ref) and discussed at length in Appendix (ref). Since this problem has a small sample size, the estimate exhibits large variation across different data splits. We thus apply the de-randomization procedure discussed in Appendix (ref) with $10,000$ splits (i.e., $B = 10,000$ in Algorithm (ref)). Our de-randomized cross-fitted RIPW estimate is reported in Table (ref), together with the estimates obtained using the TWFE regressions reported in Table (ref). It is significant at the $10\%$ level and the magnitude is larger than that given by the unweighted TWFE regressions shown in Table (ref). Recall that the joint F-test p-value for the assignment model presents strong evidence of selection and hence the difference between the RIPW estimator and the unweighted TWFE regression are likely due to the bias of the latter.
We demonstrate both theoretically and empirically that the unit-specific reweighting of the OLS objective function improves the robustness of the resulting treatment effects estimator in applications with panel data. The proposed weights are constructed using the assignment process (either known or estimated) and thus appropriate in situations with substantial cross-sectional variation in the treatment paths. Practically, our results allow applied researchers to exploit domain knowledge about outcomes and assignments, thus resulting in a more balanced approach to identification and estimation.