Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
77,957 characters · 8 sections · 21 citation commands
Controlling for Unmeasured Confounding in Panel Data Using Minimal Bridge Functions: From Two-Way Fixed Effects to Factor Models
\if11 \fi
\if01 {
} \fi
{\it Keywords:} Difference-in-Difference, Synthetic Controls, Matrix Completion, Regularization, Generalized Method of Moments, Negative Controls.
\spacingset{1.9}
Panel data, where the researcher has multiple observations on the same units over time, are ubiquitous in applications and offer a unique opportunity to control for unmeasured confounding and draw credible causal inferences. For example, under a two-way fixed effect (TWFE) model, we may remove confounding effects by simply subtracting pre-treatment outcomes. The resulting estimator is known as the difference-in-differences (DID) estimator. Because the assumptions underlying a TWFE model can be controversial in practice, more recent estimators such as synthetic controls and matrix completion allow for generalizations such as linear factor models, which may be more realistic. However, these and related methods require both the number of units and the number of time periods to grow large for consistent estimation of the effects of interest.
In this paper, we tackle causal inference from panel data under a linear factor model with a fixed number of time periods. We focus on the setting where units either start being treated at the same time period and then remain treated afterward, or are never treated at all. We show how causal effect identification is possible when we have observations in both pre-treatment and post-treatment periods. To do this, we interpret the DID differencing transformation in TWFE as a bridge function, a function that transforms pre-treatment variables so that the effect of unmeasured confounders on the transformed variables is the same as on the unobserved counterfactual outcome. In the DID setting the bridge function is very simple: any linear combination of the pretreatment outcomes with the weights summing to one will work. We show that, analogously, bridge functions also exist in the more general linear factor model, if pre-treatment variables are sufficiently informative about unobserved confounding factors. These bridge functions can effectively control for unmeasured confounding, facilitating the identification of average causal effects. Although bridge functions are defined in terms of unmeasured confounders, we can learn them by using moment equations based on post-treatment observations under an additional serial independence assumption and use them to learn average causal effects. Importantly, the number of pre-treatment or post-treatment periods does not need to grow to infinity, but only to be sufficiently large in order to account for all unmeasured confounders.
However, we show that there often exist many different bridge functions even when the average causal effect is uniquely identified. Although each of these bridge functions is valid in identifying the true average causal effect, the multiplicity of bridge functions creates challenges for estimation and inference. We solve this using a new regularized generalized method of moments (GMM) estimator that targets the minimal bridge function, \ie, the bridge function whose unknown parameters have the smallest size among all valid ones. We prove that the final average causal effect estimator based on this approach is consistent and asymptotically normal, and provide asymptotically valid confidence intervals based on a simple plug-in asymptotic variance estimator. We also prove that using the inverse moment covariance matrix as the weighting matrix in a class of regularized GMM estimators is asymptotically optimal in a certain sense. Notably, our estimator and inferential procedure are agnostic to whether bridge functions are nonunique nor not. When the bridge function happens to be unique, our results recover standard GMM estimation theory, but when this is not true, our results are still valid but standard GMM estimators may be ill-performing.
This paper is organized as follows. In (ref), we set up the linear factor model and average causal effect estimand. Then we use the TWFE model as a special example to motivate the concept of bridge functions ((ref)), review existing methods for panel data causal inference, and show their limitations in requiring infinitely many cross-sectional units and time periods to consistently estimate causal effects ((ref)). In (ref), we derive bridge functions for the linear factor model and use them to identify the average causal effect based on a fixed number of pre-treatment and post-treatment outcomes. In (ref), we propose a regularized estimator for the minimal bridge function and present the estimation and inferential theory. In (ref), we extend our methodology to alternative causal effect estimands and linear factor models with time-varying unmeasured factors, and connect this paper to the negative control framework proposed in recent literature cui2020semiparametric,tchetgen2020introduction,miao2018identifying,deaner2021proxy. We finally conclude in (ref).
Consider a balanced panel with $N$ cross-sectional units observed in time periods $t \in \braces{-T_0, \dots, -1, 0, 1, \dots, T_1}$. A treated unit $i$, with treatment indicator $A_i = 1$, is exposed to a treatment of interest for the first time at time $t = 0$ and remains in the treatment condition until the end of time horizon, whereas a control unit $i'$, with $A_{i'} = 0$, is never exposed to the treatment. We use $Y_{i, t}(0)$ to denote the potential outcome that would be observed for unit $i$ at time $t$ in absence of the treatment condition, and use $Y_{i, t}(1)$ for the potential outcome in the treatment condition. Accordingly, we observe $Y_{i, t} = Y_{i, t}\prns{1}$ only when $A_i = 1$ and $t \ge 0$, and observe $Y_{i, t} = Y_{i, t}\prns{0}$ otherwise. See (ref) for an illustration. Additionally, we may observe covariates $X_i \in \R{d}$ for each unit. We assume that $\prns{A_i, X_i, Y_{i, t}\prns{0}, Y_{i, t}\prns{1}: -T_0 \le t \le T_1}$ for $i = 1, \dots, N$ are independent and identically distributed ({i.i.d}) draws from a common (infinite) population denoted as $\prns{A, X, Y_{t}\prns{0}, Y_{t}\prns{1}: -T_0 \le t \le T_1}$.
In this paper, we are interested in the average causal effect for the treated at time $t = 0$:
Because the first term $\Eb{Y_{0} (1)\mid A = 1}$ is trivially identified, we focus on studying the identification and estimation of the second term, \ie, the counterfactual mean for the treated:
In this paper, we assume that the control potential outcomes satisfy a linear factor model.
Linear factor models for causal effects have been the focus of a growing literature doudchenko2016balancing,Abadie2007,bai2009panel,athey2021matrix,xu_2017,ruoxuan2020. One of its key property is that the unobserved confounders $U_i$ has time-varying effects on the potential outcomes, captured by the vectors $V_t$. Although here the covariates $X_i$ appear to be time invariant, actually we can still allow time-varying covariates. For example, we can consider $X_i = \prns{X_{i, -T_0}, \dots, X_{i, T_1}}$, and set coefficients $b_{t}$ to have nonzero entries only for the covariates $X_{i, t}$. The condition $\epsilon_{t} \perp \prns{U, X, A}$ requires the idisyncratic errors to be exogenous. Under this condition, we have
but we allow $U$ to be unmeasured confounders, \ie, $A \not \perp U \mid X$ and
We also impose a standard positivity assumption on the treatment assignment.
In this paper, we aim to study the identification and estimation of the counterfactual mean for the treated, \ie, $\gamma^*$ in (ref), based on the following observed data:
where $Y_{i, \mathtt{pre}} = \prns{Y_{i, t}: t < 0}$ and $Y_{i, \mathtt{post}} = \prns{Y_{i, t}: t > 0}$ denote pre-treatment outcomes and post-treatment outcomes, respectively.
\paragraph{Notation.} We use $\mathcal C = \braces{i: A_i = 0}$ and $\mathcal T = \braces{i: A_i = 1}$ to denote the sets of control units and treated units respectively, and denote their sizes as $N_0$ and $N_1$ respectively. We let $\mathtt{pre} = \braces{t: t < 0}$ and $\mathtt{post} = \braces{t: t > 0}$ denote the pre-treatment and post-treatment periods, respectively. We use them in subscripts to represent vectors or matrices whose corresponding components belong to these index subsets. For example, we define a vector $Y_{\mathcal C, 0} \in \R{N_0}$ as $\prns{Y_{i, 0}: i \in \mathcal C}$, and define $\mb Y_{\mathcal C, \mathtt{pre}}$ as a $N_0 \times T_0$ matrix whose rows correspond to control units in $\mathcal C$ and columns correspond to pre-treatment periods in $\mathtt{pre}$. We also define $\mathbf{{U}}_\mathcal C \in \mathbb{R}^{N_0 \times r}$, $\mathbf{{V}}_\mathtt{pre} \in \mathbb{R}^{T_0 \times r}$, and $\mathbf{{B}}_\mathtt{pre} \in \mathbb{R}^{T_0 \times d}$ as matrices whose rows correspond to $\prns{U_i^\top: i \in \mathcal C}$, $\prns{V_t^\top: t \in \mathtt{pre}}$, and $\prns{\beta_t^\top: t \in \mathtt{pre}}$ respectively. For positive integers $n$ and $n'$, we define $\mathbf{{1}}_{n}$ as an all-one vector of length $n$, $\mathbf{0}_{n \times n'}$ as an all-zero matrix of size $n \times n'$, and $I_{n \times n}$ as an $n \times n$ identity matrix. Other vectors and matrices with similar subscripts can be understood analogously. For a function $f(O)$ of observed variables $O$, we denote its sample average with respect to data $\prns{O_1, \dots, O_N}$ as $\hat\E_N\bracks{{f(O)}} = \frac{1}{N}\sum_{i=1}^N f(O_i)$.
One special example of our model in (ref) is the two-way fixed effect model (TWFE):
where $U_i, b_t \in \mathbb{R}$ denote unit effects and time effects respectively and additional covariates are ignored for simplicity. This corresponds to (ref) with $V_t = X_t = 1$. This model imposes a strong additivity assumption on the fixed effects, which requires the effects of unobserved confounders on the counterfactual outcomes to be time invariant (\ie, $V_t = 1$). This also implies the so-called “parallel trend” assumption, \ie, the average counterfactual outcomes of treated and control units follow parallel paths. However, this assumption is often controversial in practice ({\it e.g.,} callaway2020difference, goodman2018difference, sun2020estimating). The linear factor model in (ref) is substantially more general and in particular does not impose this “parallel trend” assumption.
The TWFE model accommodates a particularly simple estimation procedure, known as the difference-in-difference (DID) estimator for $\Eb{Y_{0}\prns{0} \mid A = 1}$:
To interpret the DID estimator, first we note that it is a sample analogue of the population estimand
Next, we note that this estimand can be written as:
where $h(Y_{\mathtt{pre}}; \theta) = \theta^{\top}_{1} Y_{\mathtt{pre}} + \theta_2$, $\theta^*_{\mathrm{DID}, 1} = \mathbf{{1}}_{T_0}/T_0$, and $\theta^*_{\mathrm{DID}, 2} = \Eb{-\frac{1}{T_0}\sum_{t\in\mathtt{pre}}Y_{t} + Y_{0} \mid A = 0} = b_0 - b^\top_\mathtt{pre} \theta_1^*$. Here $h(Y_{\mathtt{pre}}; \theta^*_{\mathrm{DID}})$ is a transformation of the pre-treatment outcomes, whose expectation for the treated units is exactly equal to the target parameter $\gamma^*$. The DID estimator learns this transformation based on control units' observed outcomes, \ie, $Y_{i, t}$ for $i \in \mathcal C$ and $t \in \mathtt{pre} \cup \braces{0}$.
Actually, the transformation $h(Y_{\mathtt{pre}};\theta^*_{\mathrm{DID}})$ learned by the DID estimator is only one among a many valid transformations: we in fact have
These functions all have a special property that
The condition in (ref) says that the effect of unmeasured confounders $U$ on the transformed pre-treatment outcomes $h(Y_{\mathtt{pre}})$ is exactly the same as the unmeasured confounding effect on $Y_{0}(0)$. Consequently, any such transformation, which we call a bridge function, can recover the parameter $\gamma^*$.
According to (ref) , whenever $T_0 > 1$, there are infinitely many different bridge functions, whose coefficients on the pre-treatment outcomes correspond to different solutions to the equation $\mathbf{{1}}_{T_0}^\top\theta_1^* = 1$. Among them, the DID bridge function $h(Y_{\mathtt{pre}}; \theta^*_{\mathrm{DID}})$ is particularly simple, in that its coefficient $\theta^*_{1, \mathrm{DID}}$ has the smallest $L_2$ norm.
In this paper, we generalize this bridge function approach to the linear factor model in (ref). We will show that for this more general model, there also exist (usually nonunique) linear bridge functions of pre-treatment outcomes that can control for the unmeasured confounding effects on the primary outcome. However, learning the bridge functions becomes more difficult: the coefficients on the pre-treatment outcomes depend on unknown parameters, so we can no longer directly pick a particular one as we do in the DID estimator. Instead, we have to estimate them from data, which we realize by leveraging post-treatment outcomes (see (ref)). The nonuniqueness of the bridge functions poses a unique challenge to estimation and inference, which we tackle using regularization (see (ref)).
In this part, we review some existing methods for estimating causal effects under the linear factor model and show their limitations. For simplicity, we ignore the covariates $X$ in (ref), \ie, setting $b_t = 0$ for all $t$. \paragraph*{Horizontal Regressions.} In (ref), we show that under the TWFE model, some transformations of the pre-treatment outcomes, \ie, the bridge functions, can control for the unmeasured confounding effects. For the linear factor model, we also have similar transformations: in (ref) (ref), we show that when $\mathbf{{V}}_\mathtt{pre}$ has full column rank (\ie, equal to the number of unmeasured confounders $r$), there exist $\theta^*_1 \in \mathbb R^{T_0}$ such that
Obviously, once we can learn a coefficient $\theta^*_1$ satisfying (ref) above, we can immediately estimate $\gamma^* = \Eb{Y_{0}\prns{0}\mid A = 1}$. One straightforward idea is to view (ref) as a regression equation, and run a linear regression of $Y_{i, 0}$ against $Y_{i, \mathtt{pre}}$, based on the data for control units up to time $t = 0$, possibly with additional regularization. This amounts to using $N_0$ data points to estimate the $T_0$-dimensional parameter $\theta^*$ in (ref). We call this as a horizontal regression, because the regressors $Y_{i, \mathtt{pre}}$ are horizontally laid out in the observed data matrix (see (ref)). Variants of horizontal regressions are also considered in hazlett2018trajectory,athey2021matrix.
However, this horizontal regression approach is susceptible to estimation bias. Indeed, the regressors $Y_{i, \mathtt{pre}}$ are dependent with errors $\xi_i$, since both include common components $\epsilon_{i, \mathtt{pre}}$. This is analogous to the well-known problem of error-in-variable regressions, where using proxy variables in place of unobserved true variables as regressors leads to coefficient estimates with nonvanishing bias wooldridge2010econometric. Similarly, here we can view pre-treatment outcomes as proxy variables for the unmeasured confounders, so even when the horizontal regression has infinite sample size, \ie, $N_0 = \infty$, the resulting counterfactual mean estimator will still have persistent bias. In (ref) (ref), we prove that the bias can vanish when the dimension of regressors in the horizontal regression, \ie, $T_0$, also grows to infinity. This means that we need both $N_0 \to \infty$ and $T_0 \to \infty$ to consistently estimate $\gamma^*$ based on horizontal regressions, which is infeasible with only observations in a limited number of time periods.
\paragraph*{Vertical Regressions.} In (ref) (ref), we also show that when $\mathbf{{U}}_{\mathcal C}$ has full column rank (which holds with high probability if components of $U$ are not collinear), there exists $w^* \in \R{N_0}$ such that
We may also view (ref) as a regression function, and run a linear regression of $\frac{1}{N_1}\sum_{i \in \mathcal T} Y_{i, t}$ against $Y_{\mathcal C, t}$ based on data up to time $t = -1$. This amounts to using $T_0$ data points to estimate the $N_0$ dimensional parameter $w^*$ in (ref). We call this as a vertical regression, because the regressors $Y_{\mathcal C, t}$ are vertically laid out in the observed data matrix (see (ref)). The vertical regression recovers the synthetic control method Abadie2007,abadie2003conflict if the regression coefficients are additionally constrained to be nonnegative and sum to one. Other variants of vertical regressions also appear in doudchenko2016balancing,chernozhukov2021ttest,BenMichael2021.
However, the vertical regression approach also has the error-in-variable regression problem, because the regressors $Y_{\mathcal C, t}$ and the errors $\nu_t$ share common components $\epsilon_{\mathcal C, t}$. Therefore, the resulting counterfactual mean estimators also have nonvanishing biases when the sample size of the vertical regressions, \ie, $T_0$, grows to infinity. Instead, we show in (ref) (ref) that we also need the dimension of regressors $N_0 \to \infty$ to consistently estimate the counterfactual mean. Similar observations were also noted by ferman2021synthetic,ferman2020properties,Gobillon2016, using different analyses.
\paragraph{Matrix Estimation.} Alternatively, some recent literature propose to impute missing counterfactuals by directly learning the factor model structure in (ref), \eg, by estimating $U_i$ and $V_t$ factors for all $i$ and $t$ xiong2019large,xu_2017,bai2021matrix, by matrix norm regularization methods athey2021matrix,farias2021learning, or by singular value thresholding amjad2018robust. However, learning the factor model structure is a very difficult high-dimensional estimation problem. To consistently estimate the factor model structure and causal parameters, these existing estimators need both $N_0$ and $T_0$ to grow to infinity (see (ref) (ref) for details), with post-treatment outcomes or not. Note that this is very different from the TWFE model: although learning all fixed effects $U_i, b_t$ for $i, t$ is also a difficult high-dimensional estimation problem, we do not need to estimate them at all. Instead, we can directly estimate the causal parameters consistently when $N \to \infty$ but $T_0$ is fixed.
\paragraph{Summary.} These existing approaches require both $N_0 \to \infty$ and $T_0 \to \infty$ to consistently estimate the causal parameter, either because of the error-in-variable regression problem, or because of the need to learn the factor model directly. Note that the post-treatment data (\ie, gray cells in (ref)) are not important in these approaches. Actually, it is not immediately clear how to use post-treatment data in horizontal/vertical regressions, since in post-treatment periods the counterfactual outcomes for the treated units are all missing. In the matrix estimation methods, although we can indeed incorporate post-treatment outcomes, they also bring in more missing entries and cannot relax the requirement of $N_0 \to \infty$ and $T_0 \to \infty$ for consistent estimation of $\gamma^*$ (see discussions in (ref)).
In this paper, we will show that even when the number of pre-treatment outcomes $T_0$ is fixed, we can still identify the counterfactual mean parameter $\gamma^*$ and estimate it consistently, provided that we have access to some post-treatment observations for a fixed number of periods (\ie, fixed $T_1$). We will build on a generalization of (ref) that relates the missing counterfactual to the pre-treatment outcomes and covariates, and show that post-treatment outcomes are valuable in addressing the problem of “error-in-variable” regressions. Importantly, our results reveal that even though the factor model structure cannot be identified or consistently estimated with data only in a fixed number of time periods, identification and consistent estimation of causal effects is still possible. Our methods thus uniquely enable effective causal inference under the popular linear factor model with big-$N$-small-$T$ panel data.
In (ref), we show that in the TWFE model, the so-called bridge functions can effectively control for unmeasured confounding and lead to the familiar DID estimator. In this section, we derive bridge functions for the linear factor model in (ref), by generalizing the formulation in (ref). Although horizontal regression methods based on this formulation may not consistently the target parameter when $T_0$ is fixed, we show that this can be realized by leveraging post-treatment outcomes.
We first introduce the definition of bridge functions.
According to (ref), bridge functions give some transformations of the pre-treatment outcomes and covariates, such that the unmeasured confounding effects on this transformation exactly reproduce those on the counterfactual outcome. This formalizes the requirement that bridge functions can control for unmeasured confounding.
In the following lemma, we show that under our assumptions, the treatment has no direct causal effects on the pre-treatment outcomes. As a result, their bridge function transformations, despite being defined in terms of the control population in (ref), can be applied to the treated population to recover the counterfactual mean for the treated.
In the following lemma, we prove the existence of bridge functions in the linear factor model, by generalizing (ref) to incorporate additional covariates.
(ref) shows the form of bridge functions in the linear factor model:
Any of these bridge functions can identify the target counterfactual mean parameter $\gamma^*$ by (ref), or equivalently by the following moment equation:
These bridge functions exist when solutions $\theta_1^*$ to the equation $\mathbf{{V}}_\mathtt{pre}^\top \theta_1^* = V_0$ exist, which is ensured for $\mathbf{{V}}_\mathtt{pre}$ with full column rank $r$. Intuitively, this rank condition means that pre-treatment outcomes are informative proxy variables for the unmeasured confounders: their number must be no smaller than the number of unmeasured confounders, \ie, $T_0 \ge r$, and the unmeasured confounding effects on them, \ie, $U^\top \mathbf{{V}}_\mathtt{pre}$, captures the variations of confounders $U$ in any direction. This is why some linear transformations of pre-treatment outcomes can control for the confounding effects on the counterfactual outcome $Y_{0}\prns{0}$. Moreover, it is obvious that the confounding bridge function is unique only when the number of pre-treatment outcomes is equal to the number of unmeasured confounders, \ie, $T_0 = r$ (also see (ref)).
However, it remains unclear how to learn the bridge functions from observed data, since their definition involves unmeasured confounders $U$ (see (ref)), and their coefficients depend on unknown factors $\mathbf{{V}}_\mathtt{pre}, V_0$ and unknown coefficients $\mathbf{{B}}_\mathtt{pre}, b_0$. This is in stark contrast to bridge functions in the TWFE model, where all but one coefficients are specified by fully known equations (see (ref)). We already show that directly regressing $Y_{i, 0}$ against $Y_{i, \mathtt{pre}}$ and $X_i$, just like the horizontal regressions in (ref), cannot learn bridge functions without bias due to the error-in-variable regression problem. In the following lemma, we show that post-treatment outcomes provide new opportunities for learning the bridge functions from observed data.
In (ref), we assume that the idiosyncratic errors in the post-treatment periods are independent with those up to period $0$. This condition trivally holds when $\epsilon_t$ for $t = -T_0, \dots, T_1$ are all serially independent. Under this condition, post-treatment outcomes $Y_{\mathtt{post}}$ are conditionally independent with the pre-treatment outcomes $Y_{\mathtt{pre}}$ and target outcome $Y_0$. Serial independence assumption is often not assumed in the TWFE model ((ref)), or in many previous literature on regression-based methods and matrix estimation ((ref)), with some exceptions like Abadie2007,amjad2018robust. However, given the general factor model with only limited pre-treatment outcomes, additional assumptions like this become important in causal effect identification. Note that this assumption does not rule out dependence among the outcomes in different periods. The post-treatment outcomes $Y_{\mathtt{post}}$ can be still dependent with the pre-treatment outcomes $Y_{\mathtt{pre}}$ and the target outcome $Y_0$, but their dependence is completely mediated by unmeasured confounders $U$ and covariates $X$. In (ref), we further allow confounders to be time-varying, so that dependence structure of the outcomes can be even more complex.
(ref) gives a moment equation characterization of bridge functions in (ref), which only depends on observed data. Although this moment equation looks very similar to moment equations in instrumental variable estimation wooldridge2010econometric, it is based on substantially different assumptions. Actually, post-treatment outcomes $Y_{\mathtt{post}}$ must not be valid instrumental variables, since below they are assumed to be strongly dependent with the unmeasured confounders. Note that this moment equation characterization does not suffice to show the identifiability of counterfactual mean $\gamma^*$: some of its solutions may not be valid bridge function coefficients, and based on only observed data, there is no way to distinguish invalid coefficients from valid ones. In the following theorem, we stregthen (ref) by showing that when post-treatment outcomes are also informative proxy variables for the unmeasured confounders, the moment equation in (ref) sharply characterizes all valid bridge functions, which proves the identifiability of the target parameter $\gamma^*$.
Here condition 1 rules out multicollinearity of the unmeasured confounders $U$ and $X$, which is a common identification condition in linear models. Condition 2 roughly means that the unobserved confounding effects on the post-treatment outcomes, \ie, $U^\top \mathbf{{V}}_\mathtt{post}$, after accounting for the covariates $X$, can still capture variations of confounders $U$ in any direction. This condition implicitly requires the number of post-treatment outcomes to be no smaller than the number of unmeasured confounders, \ie, $T_1 \ge r$. Under these two conditions, we can characterize all valid bridge functions by the moment equation in (ref) that only involves observed data, and plug any of them into (ref) to identify the counterfactual mean parameter $\gamma^*$. This shows that $\gamma^*$ can be completely determined by observed data so it is identifiable.
In (ref), we show that when ${\epsilon_{\mathtt{post}} \perp \prns{\epsilon_{\mathtt{pre}}, \epsilon_0}}$, post-treatment outcomes can be used to learn bridge functions and identify the causal parameter. If we further assume that the idiosyncratic errors are serially independent, then we can achieve this with other alternative observations. Indeed, if this is the case and given that the unmeasured confounders $U$ are time-invariant, then the temporal order of data is not important. We may use additional pre-treatment outcomes not in $Y_{\mathtt{pre}}$ (\eg, outcomes before time $-T_0$) or a mix of these additional pre-treatment outcomes and the post-treatment outcomes $Y_{\mathtt{post}}$ to form the marginal moments in (ref). Nevertheless, in (ref), we show that when unmeasured confounders themselves are time-varing, the temporal order of data is indeed important, and we must only use post-treatment outcomes to learn the bridge functions (see discussions below (ref)). Thus we focus on using post-treatment outcomes to learn bridge functions as it is robust to time-varying unmeasured confounding.
In (ref), we prove the identifiability of the counterfactual mean for the treated parameter $\gamma^*$ by bridge functions, based on both pre-treatment outcomes and post-treatment outcomes. In particular, (ref) show the identification of $\gamma^*$ by the following moment equations of $\prns{\theta, \gamma}$:
In this section, we study how to estimate $\gamma^*$ based on these moment equations. One immediate challenge in this estimation task is that solutions to $\Eb{m\prns{O; \theta}} = 0$, \ie, valid bridge functions, may not be unique.
Actually, nonunique bridge functions are likely to be prevalent. In practice, we never know the number of unmeasured confounders, so we can at best use as many pre-treatment outcomes as possible to safeguard the existence of bridge functions. It is rather unlikely that the number of negative controls is exactly the same as the number of unmeasured confounders, so that bridge functions uniquely exist. Instead, we may tend to use more than enough, \ie, $T_0 > r$, and end up with nonunique bridge functions.
Although any valid bridge function identifies the true counterfactual mean parameter $\gamma^*$, nonunique bridge functions pose serious estimation challenges. Standard methods, such as Generalized Method of Moments (GMM) hansen1982large that solves sample analogues of (ref), may fail to perform well. Their estimates are not guaranteed to converge to any fixed elements in $\Theta^*$. Worse yet, the set of valid bridge function coefficients $\Theta^*$ is an unbounded linear subspace, so standard GMM estimators may give very extreme values, which causes the resulting counterfactual mean estimator to be highly unstable. Technically, (ref) shows that the rank of the Jacobian matrix of the moment equation corresponding to the function $m$ is strictly smaller than the dimension of the bridge function coefficients to be estimated, which violates key conditions in the asymptotic guarantees for standard GMM estimators newey1994large.
To understand how to overcome this challenge, recall the TWFE model in (ref). In (ref), we show that the bridge functions under the TWFE model, characterized by $\Theta^*_{\op{FE}}$ in (ref), are also nonunique whenever the number of pre-treatment outcomes $T_0 > 1$. But we can easily target a specific bridge function and obtain the DID estimator. Motivated by this, we propose to target the minimal bridge function, whose coefficients have the smallest norm among all valid bridge function coefficients:
and $[\cdot]^+$ denotes the Moore–Penrose pseudoinverse. Obviously, when there happen to be a unique bridge function, \ie, $T_0 = r$, the minimal bridge function is trivally the only valid bridge function. Thus our approach of targeting the minimal bridge function is valid regardless of whether bridge functions are unique or not. In (ref), we also show that it is possible to target alternative bridge functions achieving smallest $\theta^\top M\theta$ for a positive semidefinite matrix $M$.
It might be tempting to estimate $\theta_{\min}^*$ by substituting empirical averages for all true expectations in (ref). However, this simple estimator may not even be consistent. Note that in (ref), we need to the pseudoinverse of a $(T_1 + d)\times(T_0 + d)$ matrix with rank at most $r + d$ (which corresponds to $\nabla\Eb{m\prns{Z; \theta^{*}_{\min}}}$ in (ref)). It is well known that Moore–Penrose pseudoinverse is not a continuous operation at singular matrices. Thus the pseudoinverse of the empirical average matrix may not converge to the pseudoinverse of the limiting expectation matrix. Therefore, this simple estimator can be often ill-performing.
Instead, we propose a regularized GMM estimator:
where $\mathcal{W}_{m, N} \in \mathbb{R}^{\prns{T_1 + d} \times \prns{T_1 + d}}$ is a (possibly data-dependent) positive definite weighting matrix that converges (in probability) to a fixed positive definite matrix $\mathcal{W}_{m, \infty}$, and $\lambda_N > 0$ is a regularization parameter that converges to $0$ as $N \to \infty$. The $L_2$ norm regularization in (ref) encourages small norm solutions to the sample moment equations. So it is reasonable to expect the regularized estimator $\hat\theta$ to target the minimal bridge function whose coefficients are given in (ref).
In the following lemma, we prove that this is indeed the case: when the regularization parameter $\lambda_N$ converges at an appropriate rate, the regularized estimator converges to the minimal bridge function coefficient.
According to (ref), $\hat\theta$ converges to the minimal bridge function coefficient $\theta_{\min}^*$ in (ref) when $\lambda_N \to 0$ but $N\lambda_N \to \infty$, which shows the validity of the regularized GMM estimator $\hat\theta$. Then we can use $\hat\theta$ to get the final estimator for $\gamma^*$:
In the following theorem, we further show that this counterfactual mean estimator has an asymptotically linear expansion with a closed-form influence function.
(ref) shows that our estimator $\hat\gamma$ based on the regularized GMM estimator for bridge functions has desirable asymptotic properties. Note that the influence function $\psi\prns{O_i; \theta_{\min}^*, \gamma^*, \mathcal{W}_{m, \infty}}$ of estimator $\hat\gamma$ has mean zero. So when $\lambda_N \sqrt{N} \to 0$ and $\lambda_N N \to \infty$, we can use Law of Large Number and Central Limit Theorem to show that estimator $\hat\gamma$ is $\sqrt{N}$-consistent with an asymptotic normal distribution. The asymptotic variance of $\hat\gamma$ is given by the variance of the influence function, \ie,
This can be estimated by a straightforward plug-in estimator:
In the following theorem, we further prove that the variance estimator is consistent and it can be used to construct asymptotically valid confidence intervals.
In (ref), we derive the asymptotic property of estimator $\hat\gamma$ and the associated confidence intervals, both based on a generic weighting matrix $\mathcal{W}_{m, N}$ with a limit $\mathcal{W}_{m, \infty}$. We may wonder how to choose this weighting matrix. In standard GMM estimation, it is well known that the asymptotically optimal weighting matrix is the inverse moment covariance matrix hansen1982large. In the following theorem, we show that the inverse moment covariance matrix is also optimal for regularized GMM estimation.
(ref) shows that among the class of regularized GMM estimators given in (ref), the ones that achieve the smallest asymptotic variance have $\Sigma_{m}^{-1}$ as the limit of their weighting matrix in (ref). Note that the optimal asymptotic variance $\sigma^2\prns{\Sigma_{m}^{-1}}$ depends on the minimal bridge function coefficient $\theta^*_{\min}$ that we choose to target. If we target a different one, then the optimal asymptotic variance will also change accordingly, and the best target depends on the unknown covariance of pre-treatment idiosyncratic errors (see (ref) (ref)).
We can construct an asymptotically optimal estimator by a two-stage approach:
\paragraph*{Effects on Multiple Time Periods. } In previous sections, we only study the treatment effects on the outcome in period $t = 0$. We can also consider effects on outcomes aggregated from multiple time periods, \eg, the ATT parameter $\frac{1}{L}\sum_{t=0}^L \Eb{Y_{t}\prns{1} - Y_{t}\prns{0} \mid A = 1}$ for an integer $0 < L < T_1$. In this case, we can straightforwardly adapt the results in (ref) to the identification and estimation of the aggregated counterfactual mean $\frac{1}{L}\sum_{t=0}^L \Eb{Y_{t}\prns{0} \mid A = 1}$. In particular, the corresponding bridge functions still have the form in (ref), with the set of coefficients being $$\Theta^* = \braces{\theta^*: \mathbf{{V}}_\mathtt{pre}^\top \theta_1^* = \frac{1}{L}\sum_{t=0}^L V_t, \theta_2^* = {b_0 - \mathbf{{B}}_\mathtt{pre}^\top\theta^*_1}}.$$ For the rest of part, we only need to redefine $\mathtt{post} = \braces{L+1, \dots, T_1}$ and revise the assumptions and estimation procedures accordingly.
\paragraph*{Counterfactual Mean for the Whole Population.} In previous sections, we focus on the counterfactual mean for a subpopulation, \ie, the treated units. We can also study the counterfactual mean for the whole population, \ie,
This means that we need to solve the following moment equations:
Note that here the moment function $\tilde g$ and $m$ are generally correlated, while the previous moment function $g$ and $m$ in (ref) have zero correlations. Because of the latter, when estimating the counterfactual mean for the treated, we can solve the two moment equations in (ref) separately without loss of efficiency. However, when estimating the counterfactual mean for the whole population, explicitly accounting for the correlations between the two moment equations in (ref) can improve the asymptotic estimation efficiency. See (ref) for details.
\paragraph*{Time-Varying Unmeasured Confounders.} In the linear factor model in (ref), although the unmeasured confounders have time-varying effects (characterized by matrices $\braces{V_t: t= - T_0, \dots, T_1}$ for all $t$), the confounders themselves are time-invariant. Now we drop this restriction, and consider unmeasured confounders that follow an autoregression model. Under this model, the unmeasured confounders are also time-varying, so the serial dependence structure of counterfactual outcomes also becomes more complex than before.
In (ref), we assume that $A \perp U_t \mid X, U_0$, which means that the treatment assignment $A$ only depends on confounders $U_0$ at the time of the treatment, but not any past or future confounder. This condition implies that $A \perp Y_{\mathtt{pre}} \mid X, U_0$, an analogue of (ref), ensuring that bridge functions can be applied to the treated units to recover the counterfactual mean parameter (see (ref)).
Note that we should not deal with this model by redefining $U_i = \prns{U_{i, -1}, \dots, U_{i, -T_0}} \in \R{T_0 r}$ and casting it as a special example of (ref). Otherwise the dimensionality of $U_i$ exceeds the number of pre-treatment outcomes $T_0$ unless $r = 1$, so the rank condition in (ref) is violated. Therefore, we have to handle the model in (ref) directly.
In the following theorem, we show that under additional assumptions for the unmeasured confounders and the transition innovations, this model again has linear bridge functions.
In (ref), condition 1 is the key condition for the bridge functions to be linear in the pre-treatment outcomes and the covariates. It assumes linear regression functions of transition innovations $\eta_t$ and the unmeasured confounders $U_{-T_0}$ with respect to unmeasured confounders $U_0$, in the control subpopulation. This condition is satisfied, for example, when the innovations and unmeasured confounders in the control subpopulation have a joint normally distribution. In (ref) (ref), we relax this condition to allow the two conditional expectations to depend on covariates $X$ as well, and prove a similar conclusion, albeit with more complex notations. Condition 2 rules out collinear components in the unmeasured confounders $U_0$. The matrix in condition 3 characterizes the effects of unmeasured confounders on the pre-treatment outcomes that can be attributed to $U_0$. We require this matrix to be invertible, as an anologue of the rank condition in (ref), to ensure that pre-treatment outcomes are sufficiently informative proxies for confounders $U_0$. In (ref), we discuss that condition 1 in (ref) is not necessary for the existence of bridge functions, but without this condition bridge functions may not be linear, which is out of the scope of this paper.
In the following theorem, we further show that under conditions analogous to those in (ref), we can again use post-treatment outcomes to learn the bridge functions.
In (ref), we assume that confounding innovations in post-treatment periods are conditionally independent with those in pre-treatment periods and the initial unmeasured confounders $U_{-T_0}$. The former condition holds trivially when the innovations are serially independent. Under these two condition, the dependence between unmeasured confounders in the pre-treatment periods and those in the post-treatment periods is fully mediated by $U_0$ and covariates $X$, \ie, $\prns{U_t, t \in \mathtt{post}} \perp \prns{U_s, s \in \mathtt{pre}} \mid U_0, X$, which in turn ensures (ref). According to (ref), (ref) shows the identifiability of the target counterfactual mean parameter $\gamma^*$. Then we can estimate it by applying the regularized GMM estimation procedure in (ref) to the moment equation in (ref).
Below (ref), we mention that when unmeasured confounders are time-invariant and idiosyncratic errors are serially independent, it is possible to use some extra pre-treatment outcomes not in $Y_{\mathtt{pre}}$ to learn the bridge functions. However, this is no longer feasible when unmeasured confounders are time-varying. In this time-varying setting, unmeasured confounders in pre-treatment periods follow the autoregressive process in (ref), so their dependence cannot be fully mediated by covariates $X$ and confounders $U_0$ taking place after the pre-treatment periods. As a result, pre-treatment outcomes are all dependent even conditionally on $U_0$ and $X$, so we cannot use extra pre-treatment outcomes in place of $Y_{\mathtt{post}}$ in (ref). Therefore, when unmeasured confounders are time-varying, we must use only post-treatment outcomes to learn the bridge functions.
\paragraph{Connections to Negative Controls.}
Recently, a series of works propose a negative control framework to deal with the challenge of unmeasured confounding cui2020semiparametric,tchetgen2020introduction,miao2018identifying,deaner2021proxy,shi2020multiply. This framework requires two types of proxy variables for the unmeasured confounders: negative control outcomes $W$ and negative control treatments $Z$. These proxy variables are informative in that they are dependent with the unmeasured confounders. Importantly, they have special causal relations with other variables: the negative control outcomes $W$ cannot be directly caused by the primary treatment of interest, and the negative control treatments $Z$ cannot directly cause either the negative control outcomes or the primary outcome of interest. In (ref), with slight abuses of notations, we show a typical causal diagram of negative controls when studying the causal effect of a primary treatment $A$ on a primary outcome $Y$ in presence of unmeasured confounders $U$.
In our panel data setting, we have natural candidates for the negative control variables (see (ref) for illustrations): the pre-treatment outcomes $Y_{\mathtt{pre}}$ can be considered as negative control outcomes $W$ and the post-treatment outcomes $Y_{\mathtt{post}}$ can be considered as negative control treatments $Z$. Indeed, pre-treatment outcomes are realized before the treatment takes place so they may not be directly caused by the treatment $A$, and post-treatment outcomes are realized after the primary outcome $Y_0$ and pre-treatment outcomes $Y_\mathtt{pre}$, so they may not directly cause the latter. These conditions are formalized in (ref) respectively.
In (ref), we identify the causal parameters through bridge functions. This concept is originally proposed in the negative control literature miao2018a,cui2020semiparametric,tchetgen2020introduction,deaner2021proxy,Miao2016, where some also apply it to panel data. The connection between negative controls and (nonlinear) difference-in-differences is also noted by Sofer2016. Our identification results in (ref) can be viewed as an application of the general negative control identification strategy to linear factor models in panel data. By focusing on this important model, we explicitly characterize when bridge functions exist and when post-treatment outcomes can be used to learn them, shedding light on the abstract conditions assumed in previous literature. More importantly, our analyses elucidate the benefits and prices of the negative control identification relative to existing ones: using negative controls only needs finite time periods, but has to additionally assume some serial independence assumptions on the idiosyncratic errors (see (ref)).
However, our estimation method in (ref) is distinct from those in previous negative control literature, in that we use regularization to deal with the prevalent problem of nonunique bridge functions (see discussions below (ref)). Remarkbly, our estimator enjoys $\sqrt{n}$-consistency, asymptotic normality, and simple plug-in confidence intervals, regardless of whether bridge functions are unique. In contrast, previous negative control literature often explicitly or implicitly assume a unique bridge function to facilitate estimation miao2018a,cui2020semiparametric,shi2020multiply,QiZhengling2021PLfI,pmlr-v139-mastouri21a,SinghRahul2020KMfU,GhassamiAmirEmad2021MKML. deaner2021proxy,kallus2021proxy recognize this problem and derive convergence rates of their proposed causal effect estimators even when bridge functions are nonunique, but they either rely on inefficient sample splitting and computationally intensive bootstrap methods, or do not have inferential procedures. Our regularization approach that targets the minimal bridge function provides a new solution to this important problem. Extending it to more general negative control settings is an exciting future direction.
In this paper, we study the identification and estimation of average causal effects in panel data under a linear factor model. Previous regression-based methods (\eg, synthetic controls) and matrix estimation methods (\eg, matrix completion) all require both the number of units and the number of time periods to grow to infinity to consistently estimate the causal effects. So they may not be suitable when only observations in a relatively small number of time periods are available.
Motivated by the differencing transformation in the DID estimator for the simpler TWFE model, we propose to identify the causal effects using bridge functions, namely some tansformations of pre-treatment outcomes to control for unmeasured confounding. Learning bridge functions for causal effect estimation requires sufficiently many informative pre-treatment and post-treatment outcomes to account for all unmeasured confounders, but not infinitely many. Noting that bridge functions are often nonunique in practice, we propose an novel regularized GMM estimator to target the minimal bridge function. We prove that the resulting causal effect estimators and confidence intervals have desirable asymptotic guarantees, regardless of whether bridge functions are unique or not. Our proposal thus features a novel approach to handle unmeasured confounding in panel data with observations in only a limited number of time periods.