Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
38,408 characters · 7 sections · 19 citation commands
Identification of Average Treatment Effects in Nonparametric Panel Models
\newtheorem{result}{Result} \newtheorem{example}{Example} \newtheorem{aside}{Aside} \newtheorem{comments}{Comment} \newtheorem{conjecture}{Conjecture}
\listoftodos
There is a large literature on estimating causal effects and predictive estimands using panel data. Various estimators have been proposed, often motivated by different models of the underlying data generating process. There is not a widely used model with an accompanying set of assumptions that is the focal point of this literature, a situation that contrasts with the cross-section setting, where there are several commonly used sets of assumptions such as unconfounded treatment assignment and overlap in the distribution of propensity scores. The panel data literature most commonly relies on estimators that are justified at least in part by functional form assumptions, but the precise role of these assumptions has not been fully described. In this paper we propose a model of the data generating process together with a set of conditions under which the conditional expectations of outcomes given unobserved unit and time components are identified, where the conditions do not include functional form restrictions. Thus, our approach is analogous in spirit to the approach taken in the literature on identification of causal effects under the unconfoundedness and overlap assumptions (see e.g. imbens2015causal). We the extend these results for predictive estimands to the identification of average causal effects and decomposition estimands.
The outcome model is a generative structural model that takes the form of a Nonparametric Panel Model (NPM):
where the unobserved components $\alpha_i$, $\beta_t$ and $\varepsilon_{it}$ may be vector-valued. Our main results rely on several assumptions about these components. First, the sets $\{\alpha_i\}_{i=1}^{N},$ $\{\beta_t\}_{t=1}^{T},$ and $\{\varepsilon_{it}\}_{i,t},$ are independent of each other. In addition, the $\alpha_i$ are independent across units. However, the $\beta_t$ and $\varepsilon_{it}$ may be autocorrelated over time, but we assume stationarity of the process. These assumptions are satisfied if $g$ has an additive form, as in the popular Two-Way-Fixed-Effect (TWFE) models, or if it takes the form of a linear factor model. However, much richer models are consistent with this setup.
Although we cannot identify the unobserved unit and time components $\alpha_i$ and $\beta_t$ under our assumptions, we can both identify and consistently estimate the expected value $\mu(\alpha_i,\beta_t)=\mathbb{E}[Y_{it}|\alpha_i,\beta_t]$, which is a key building block for both our average treatment effect and decomposition results. Beyond the generative model, this result relies solely on smoothness and restrictions on the asymptotic sequences. We do so by constructing sets of units ${\cal J}(i)\subset\{1,\ldots,N\}$ with increasing cardinality such that for units $j\in{\cal J}(i)$ in this set the function $\mu(\alpha_j,\beta)$ is similar to $\mu(\alpha_i,\beta)$ as a function of $\beta$. We construct these sets by finding units $j$ such that, for all other units $k$, the covariance over time between $Y_{it}$ and $Y_{kt}$ is similar to the covariance between $Y_{jt}$ and $Y_{kt}$. We then estimate $\mu_{it}$ by averaging $Y_{jt}$ for units in such sets ${\cal J}(i).$
For the causal problem, we focus on a setting with a binary treatment $W_{it}$ that affects only unit $i$ in period $t$ (thus ruling effects of the treatment that persist over time). This allows the potential outcomes to be defined in terms of the contemporaneous treatment, $Y_{it}(0)$ and $Y_{it}(1)$. We focus on identification of the average treatment effect for the treated,
although our results can be generalized to other average treatment effects. The primary condition for our result is that conditional on $\alpha_i$ and $\beta_t$, assignment to treatment is independent of $Y_{it}(0)$:
In addition we need an overlap condition that the probability of receiving the treatment conditional on $\alpha_i$ and $\beta_t$ is bounded away from zero. This Latent Factor Unconfoundedness assumption is analogous to unconfoundedness in the cross-section setting. A final condition is that for $\mu(\alpha,\beta)\equiv \mathbb{E}[Y_{it}|\alpha_i=\alpha,\beta]$, $\mathbb{E}_\beta[\mu(\alpha,\beta)-\mu(\alpha',\beta))^2]=0$ implies $\alpha=\alpha'$. Our main result for the treatment effect setting is that given sufficient smoothness on $g(\cdot)$, the average effect for the treated $\tau$ is identified in this setting when both $T$ and $N$ go to infinity.
In addition to causal problems, the methods may also be applied to other purely predictive problems that rely on identification and estimation of $\mu_{it}$. One such category of problems is decomposition of a difference in outcomes (e.g. by demographic group or other categorization) into explained and unexplained components, as is common in the literature on gender wage gaps blau2017gender. We develop this application further in Section (ref).
There are a number of distinct literatures that the current paper builds on. One is the econometric literature on panel data. Much of this literature has focused on Two-Way-Fixed-Effect (TWFE) models and related Difference-In-Differences (DID) estimators, and more recently linear factor models bai2009panel. See chamberlain1984panel, arellano2001panel, arellano2003panel,baltagi2008econometric,hsiao2022analysis, wooldridge2010econometric,arellano2011nonlinear for textbook discussions and arkhangelsky2024causal,roth2022s for recent surveys.
Extensions allowing for nonlinear functions of linear factor models are considered in chen2021nonlinear, yalcin2001nonlinear, feng2020causal, feng2023optimal, freyberger2018non, zeleneev2020identification, abadiecausal.
The current study is also related to the literature on synthetic control methods, abadie2003,abadie2010synthetic, abadie2019using, arkhangelsky2021synthetic, liu2020practical, fry2024method. In that literature the generative models are often not specified explicitly, with the focus on algorithms that lead to effective estimators.
Another literature that is relevant is the research on row and column exchangeable arrays, see aldous1981representations, lynch1984canonical, mccullagh2000resampling. Nonparametric models motivated by the double exchangeability have been considered in network settings, where they are referred to as graphons. The focus in this literature is on estimating the conditional probability of links between nodes in the network, assuming the probability depends in an unrestricted way on unit specific components for both nodes. See for key insights bickel2009nonparametric, bickel2011method, zhang2017estimating, graham2024sparse, chiang2023inference.
Very relevant is also the work by bonhomme2015grouped, bonhomme2022discretizing, cytrynbaum2020blocked on grouped heterogeneity. In this line of work units are grouped through k-means clustering so that units in the same group have the same values for the unobserved unit components.
We are interested in estimating causal effects in a panel data setting. We focus on a setting with a binary treatment, with ${\bf W}$ the $N\times T$ matrix of treaments, with typical element $W_{it}\in\{0,1\}$. There are two matrices of potential outcomes ${\bf Y}(0)$ and ${\bf Y}(1)$. We observe the matrix of realized outcomes ${\bf Y}$ with typical element $Y_{it}=Y_{it}(W_{it})$. This notation already imposes the stable unit treatment values assumption or SUTVA rubin1978bayesian, imbens2015causal, ruling out dynamic effects.
One of the estimands that is a main focal point of the current study is the average effect on the treated,
Our starting point, and a key contribution, is a fully nonparametric generative model that captures the relation between outcomes for different units and time periods:
To be able to be precise about consistency we first need to be precise about the sequences of matrices or populations. Let us index the sequence of matrices by $m=1,2,\ldots.$ The dimensions of the matrix ${\bf Y}_m$ are $N_m$ and $T_m$.
We also need some smoothness conditions. Note that the row and column exchangeability implies the existence of a measurable function $g_m(\cdot)$. In the smoothness assumption we strengthen that to include continuity and Lipschitz conditions.
Before focusing on the average treatment effect or decomposition results, we show a preliminary identification result for the expected value of $Y_{it}(0)$ given $\alpha_i$ and $\beta_t$. This is a key technical result that is a building block for the more substantively interesting identification results discussed later. For this result we focus on a predictive setting where we observe the entire matrix ${\bf Y}(0)$, and do not use the Latent Factor Unconfoundedness assumption. Clearly identification of the entire function $g(\cdot)$, or of the unit and time components $\alpha_i$ and $\beta_t$ is not feasible based on Assumptions (ref), (ref), and (ref) without more structure. However, we do not need to estimate either the function $g(\cdot)$ itself or its arguments. For estimating average treatment effects and for our decomposition results it suffices to estimate the conditional expectation of $g(\alpha_i,\beta_t,\varepsilon_{it})$, conditional on $(\alpha_i,\beta_t)$, for all the treated pairs of indices $(i,t)$. Define the conditional expectations \[ \mu(\alpha,\beta)\equiv \mathbb{E}[g(\alpha,\beta,\varepsilon)|\alpha,\beta],\qquad \mu_{it}\equiv \mu(\alpha_i,\beta_t) \] and the residual functions \[ \eta(\alpha,\beta,\varepsilon)\equiv g(\alpha,\beta,\varepsilon)-\mu(\alpha,\beta).\qquad\eta_{it}\equiv\eta(\alpha_i,\beta_t,\varepsilon_{it}).\]
In this section we focus on consistently estimating $\mu_{it}$ for a single unit/period pair $(i,t)$ as a building block. So we want to have an estimator $\hat Y_{it}$ that is a function of ${\bf Y}$ such that
under Assumptions (ref), (ref), and (ref).
We establish this identification result by constructing a set ${\cal J}_m(i)\subset \{1,2,\ldots,N_m\}$ with two properties under Assumptions (ref), (ref), and (ref) (we do not need Assumption (ref) for this discussion): $(i)$ units $j$ in this set have $\mu(\alpha_j,\beta)$ close to $\mu(\alpha_i,\beta)$ for all $\beta$ with high probability, and $(ii)$ the cardinality of the set increases without bounds with $m$. (And of course a similar strategy would be to look for the corresponding set of time indices.) The key insight is that in order to check that unit $j$ and unit $i$ have similar values of the function $\mu(\alpha_i,\beta)$ and $\mu(\alpha_j,\beta)$ for all $\beta$, it is {\it not} sufficient to compare outcomes $Y_{it}$ and $Y_{jt}$ for units $i$ and $j$ over time. Instead we compare their covariances with other units. Specifically, we compare the covariance between unit $i$ and unit $k$ with the covariance between unit $j$ and unit $k$. If for large $T_m$ and $N_m$ those covariances are similar for {\it all} other comparison units $k$, then it must be the case that $\mu(\alpha_i,\beta)$ and $\mu(\alpha_j,\beta)$ are similar for all $\beta$. Note that it does {\it not} mean that $\alpha_i$ and $\alpha_j$ are close, but it means that we can use unit $j$ for estimating the conditional mean for unit $i$ in the same period. A second insight is that although we do not directly observe $\mu(\alpha_i,\beta)$ for unit $i$, we can estimate the covariance between $\mu(\alpha_i,\beta)$ and $\mu(\alpha_j,\beta)$ using the covariance between $Y_{it}$ and $Y_{kt}$ over time.
First we define an infeasible version of the set ${\cal J}_m(i)$ that helps develop the intuition for the identification result. Define \[ {\cal J}^*(\alpha)\equiv \left\{ \alpha'\in[0,1],{\rm\ s.t.\ } \sup_{\alpha''} \biggl|\mathbb{E}_\beta\left[\left.\left\{\mu(\alpha,\beta)-\mu(\alpha',\beta)\right\}\mu(\alpha'',\beta)\right|\alpha,\alpha',\alpha''\right] \biggr|=0 \right\} .\]
We cannot use this first infeasible identification result directly to create, for a given unit $i$, a set of units with similar functions $\mu(\alpha,\beta)$ because we do not know either the values of the $\alpha_i$ nor the function $\mu(\alpha,\beta)$. This leads to three complications and corresponding modifications of the set ${\cal J}^*(\alpha)$ to turn this into an actual identification result. First, we compare the covariances only at the sample values $\alpha_i$, still assuming we know the entire functions $\mu(\alpha_i,\beta)$ as a function of $\beta$, but now only for all the realized values $\alpha_i$. This implies that we cannot find pairs of units with exactly the same $\mu(\alpha_j,\beta)$, only approximately so, with the apporoximation improving as the number of units $N_m$ increases. Second, we only evaluate the $\mu(\alpha,\beta)$ at the sample values of $\beta_t$, instead of taking the expectation over $\beta_t$. In other words, we focus on the average $\sum_t \mu(\alpha_i,\beta_t)\mu(\alpha_k,\beta_t)/T_m$ being close to the average $\sum_t \mu(\alpha_i,\beta_t)\mu(\alpha_k,\beta_t)/T_m$, for all $k$. The quality of the approximation this induces relies on $T_m$ being large. Third, we do not observe the values of the products $\mu(\alpha_i,\beta_t)\mu(\alpha_k,\beta_t)$ for pairs of units $i$ and $k$ at all $t$, we need to estimate these products using the averages $\sum_{t=1}^{T_m} Y_{it}Y_{kt}/T_m$. For this we rely on $T_m$ being large relative to the number of units $i$ and $k$ where we evaluate these averages to ensure those averages approximate the expectations accurately.
The first of our main identification results presents a feasible version of Lemma (ref). It constructs a set of indices ${\cal J}_{m,\nu}(i)\subset\{1,\ldots,N_m\}$. The set is indexed by $\nu$ which governs how close the functions $\mu(\alpha_j,\beta)$ are to $\mu(\alpha_j,\beta)$ for $j\in{\cal J}_{m,\nu}(i) $, with high probability. Define \[ {\cal J}_{m,\nu}(i)\equiv \left\{j=1,\ldots,N_m,{\rm\ s.t.\ } \max_{k\neq i,j} \left|\frac{1}{T_m}\sum_{t=1}^{T_m} (Y_{it}-Y_{jt}) Y_{kt}\right|\leq \nu \right\}.\]
Let $\overline{\mu}$ be an upper bound on the absolute value of $\mu(\alpha,\beta)$, and let $\overline{\mu'}$ be an upper bound on the absolute value of the derivatives of $\mu(\alpha,\beta)$ with respect to $\alpha$ and $\beta.$ Let $C_Y$ be an upper bound on the absolute value of $g(\alpha,\beta,\varepsilon).$
In this section we use the results from the previous section to estimate average treatment effects under an unconfoundedness type assumption. We start from a slightly different point though, without the generating model $g(\alpha_i,\beta_t,\varepsilon_{it})$.
First, we postulate the existence of a pair of potential outcomes $Y_{it}(0)$ and $Y_{it}(1),$ corresponding to the outcomes without and given the intervention. There is a binary treatment $W_{it}\in\{0,1\}$, with the realized outcome corresponding to the potential outcome given the treatment received, \[ Y_{it}\equiv \left\{
\right.\] We are interested in the average treatment effect for the treated, \[ \tau\equiv \frac{1}{\sum_{i,t} W_{it}}\sum_{i,t} W_{it} \Bigl(Y_{it}(1)-Y_{it}(0)\Bigr)\qquad {\bf (ATT)}\] The insights extend to other estimands such as the overall average treatment effect.
In order to identify $\tau$ we need some restrictions on the assignment mechanism. Our key assumption is closely related to the standard unconfoundedness/ignorability assumptions in the cross-section evaluation literature rosenbaum1983central, imbens2015causal. Such assumptions are typically stated as independence of the assignment and the set of potential outcomes conditional on observed confounders. Here the critical assumption is formulated differently in two aspects. The first, minor, difference is that we only use {\it weak} unconfoundedness, where we assume independence of the assignment and the control potential outcome, as introduced in imbens2000. Second, the conditioning is on {\it unobserved} confounders, rather than observed confounders. In particular, we assume there are latent unobserved confounders such that conditional on those unconfounders assignment would be as good as random. That by itself is without loss of generality, in other words such an assumption does not have any direct content. Our assumption strengthens this by assuming that this set of latent confounders can be separated into some that are unit-specific and some that are time-specific, and {\it none} that are unit-time specific.
Here we want to extend the insights from the previous section to the case where we are interested in estimating the average treatment effect for the treated, $\tau=\sum_{i,t} W_{it}(Y_{it}(1)-Y_){it}(0))/\sum_{i,t} W_{it}.$ This involves estimating $\mu_{it}=\mathbb{E}[Y_{it}(0)|\alpha_i,\beta_t]$ for all treated unit/period pairs. The primary challenge relative to estimating $\mu_{it}$ in the previous section is that we cannot estimate the marginal covariance between outcomes for units $i$ and $k$ because we do not observe $Y_{it}(0)$ and $Y_{kt}(0)$ for all units and periods. Moreover, the time units and time periods where we do observe $Y_{it}(0)$ may be systematically different from those where we do not do so. To deal with that we modify the definition of the sets ${\cal J}^*(\alpha)$ and ${\cal J}_{m,\nu}(i)$. To avoid notational confusion we refer to the new sets as ${\cal C}^*(\alpha)$ and ${\cal C}_{m,\nu}(i)$ (where the ${\cal C}$ stands for Causal).
Define the propensity score and the marginal treatment assignment probability \[ e(\alpha_i,\beta_t)\equiv{\rm pr}(W_{it}=1|\alpha_i,\beta_t). \] and the set \[ {\cal C}^*(\alpha)\equiv \left\{ \alpha'\in[0,1],{\rm\ s.t.\ } \sup_{\alpha''} \biggl|\mathbb{E}_\beta\left[\left.\left\{\mu(\alpha,\beta)-\mu(\alpha',\beta)\right\}\mu(\alpha'',\beta)\right|\alpha,\alpha',\alpha''\right] \biggr|=0 \right\} .\]
First define for a subset ${\cal N}$ of $\{1,\ldots,N_m\}$ the set of periods \[{\cal S}({\cal N})=\{t=1,\ldots,T_m| W_{jt}=0\forall j\in{\cal N}\}\] where all the units in the set ${\cal N}$ are in the control group. By the overlap assumption the cardinality of the set ${\cal S}({\cal N})$ will be increasing as $T_m$ increases for a finite set ${\cal N}.$ For a given unit $i$ we now want to construct sets ${\cal J}_{m,\nu}(i)$ such that if $j\in{\cal J}_{m,\nu}(i) $, then $\mu(\alpha_i,\beta)$ is close to $\mu(\alpha_j,\beta)$ for all $\beta$. Previously we operationalized this by ensuring that the expectation of $(\mu(\alpha_i,\beta_t)-\mu(\alpha_j,\beta_t)^2$ is close to zero. This in turn we ensured by requiring that the expectation of $\mu(\alpha_i,\beta_t)-\mu(\alpha_j,\beta_t))\mu(\alpha_k,\beta_t)$ is close to zero for all other units $k$. In fact we only used the fact that this expetation was close to zero at two choices for $k$: at $k=i$ and at $k=j$. Now to deal with the missing $Y_{it}(0)$ we modify this by ensuring that the {\it conditional} expectation of $\mu(\alpha_i,\beta_t)-\mu(\alpha_j,\beta_t))\mu(\alpha_k,\beta_t)$ is close to zero for all other units $k$. The conditioning is designed to ensure that we can compare the expectations for two choices of $k$. This means that we now require that for all pairs $(k,l)$ the expectations $(\mu(\alpha_i,\beta_t)-\mu(\alpha_j,\beta_t))\mu(\alpha_k,\beta_t)$ and $(\mu(\alpha_i,\beta_t)-\mu(\alpha_j,\beta_t))\mu(\alpha_l,\beta_t)$ are close to zero, {\it conditional} on $W_{it}=W_{jt}=W_{kt}=W_{lt}=0$.
Define \[ {\cal J}_{m,\nu}(i)\equiv \left\{j=1,\ldots,N_m,{\rm\ s.t.\ } \max_{k,l\neq i,j}\left\{ \left|\frac{\sum_{t\in{\cal S}(\{i,j,k,l\})}(Y_{it}-Y_{jt}) Y_{kt}}{ \sum_{t\in{\cal S}(\{i,j,k,l\}} 1} \right|\right. \right.\] \[\left.\left.+ \left|\frac{\sum_{t\in{\cal S}(\{i,j,k,l\})}(Y_{it}-Y_{jt}) Y_{lt}}{ \sum_{t\in{\cal S}(\{i,j,k,l\}} 1} \right|\right\} \leq \nu \right\},\]
Then, under suitable rate conditions on $\nu$ we can estimate $\tau$ as \[ \hat\tau=\frac{\sum_{i,t} W_{it}\Bigl( Y_{it}-\hat Y_{it,\nu}(0)\Bigr)} {\sum_{i,t} W_{it}},\qquad {\rm where}\quad \hat Y_{it,\nu}(0)=\frac{\sum_{j\in{\cal J}_{m,\nu}(i)} Y_{jt}}{\sum_{j\in{\cal J}_{m,\nu}(i)} 1}.\]
In this section, we will change the interpretation of $W_{it}$ to denote group membership, and we do not impose unconfoundedness assumptions. Rather, we make use of identities to break out differences in group means into several components. We further modify notation slightly and assume that there is a parallel NPM for units where $W_{it}=1$, so that for $w\in\{0,1\}$:
where the unobserved components $\alpha_i^w$, $\beta_t^w$ and $\varepsilon^w_{it}$ may be vector valued. We further let
Then, we develop a decomposition by first using iterated expectations and substituting in ((ref)), and then adding and subtracting a term:
If $\alpha_i$ and $\beta_t$ were observed, the analysis above would be an example of standard decompositions (e.g. blau2017gender), where((ref)) and ((ref)) are referred to as the “unexplained” and “explained” gap, respectively. The unexplained gap is the part that remains when the two groups have the same distribution of observables, while the explained gap is the difference that arises when both groups have the same expected outcome function (but different distributions of characteristics). In our case, $\alpha_i$ and $\beta_t$ are unobserved, but the analysis of Section (ref) can be adapted to show that ((ref)) and ((ref)) are identified and can be estimated under the overlap condition.
We propose a nonparametric version of a factor model for estimating average treatment effects in a panel data setting. The model substantially generalizes the commonly used linear factor models and allows for a clear statement of the critical assumptions.
We establish that there is a consistent estimator of the expected non-treated outcome for any given unit and time period, and we use this result to provide conditions under which the average treatment effect is identified.
Our results can also be applied in other settings, such as to problems of decomposing gaps in outcomes by group, as is common in the literature on the gender wage gap.