EconBase
← Back to paper

Identification of Average Treatment Effects in Nonparametric Panel Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

38,408 characters · 7 sections · 19 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification of Average Treatment Effects in Nonparametric Panel Models

\newtheorem{result}{Result} \newtheorem{example}{Example} \newtheorem{aside}{Aside} \newtheorem{comments}{Comment} \newtheorem{conjecture}{Conjecture}

titlepageThis paper studies identification of average treatment effects in a panel data setting. It introduces a novel nonparametric factor model and proves identification of average treatment effects. The identification proof is based on the introduction of a consistent estimator. Underlying the proof is a result that there is a consistent estimator for the expected outcome in the absence of the treatment for each unit and time period; this result can be applied more broadly, for example in problems of decompositions of group-level differences in outcomes, such as the much-studied gender wage gap.

\listoftodos

Introduction

There is a large literature on estimating causal effects and predictive estimands using panel data. Various estimators have been proposed, often motivated by different models of the underlying data generating process. There is not a widely used model with an accompanying set of assumptions that is the focal point of this literature, a situation that contrasts with the cross-section setting, where there are several commonly used sets of assumptions such as unconfounded treatment assignment and overlap in the distribution of propensity scores. The panel data literature most commonly relies on estimators that are justified at least in part by functional form assumptions, but the precise role of these assumptions has not been fully described. In this paper we propose a model of the data generating process together with a set of conditions under which the conditional expectations of outcomes given unobserved unit and time components are identified, where the conditions do not include functional form restrictions. Thus, our approach is analogous in spirit to the approach taken in the literature on identification of causal effects under the unconfoundedness and overlap assumptions (see e.g. imbens2015causal). We the extend these results for predictive estimands to the identification of average causal effects and decomposition estimands.

The outcome model is a generative structural model that takes the form of a Nonparametric Panel Model (NPM):

equation[equation omitted — 105 chars of source]

where the unobserved components $\alpha_i$, $\beta_t$ and $\varepsilon_{it}$ may be vector-valued. Our main results rely on several assumptions about these components. First, the sets $\{\alpha_i\}_{i=1}^{N},$ $\{\beta_t\}_{t=1}^{T},$ and $\{\varepsilon_{it}\}_{i,t},$ are independent of each other. In addition, the $\alpha_i$ are independent across units. However, the $\beta_t$ and $\varepsilon_{it}$ may be autocorrelated over time, but we assume stationarity of the process. These assumptions are satisfied if $g$ has an additive form, as in the popular Two-Way-Fixed-Effect (TWFE) models, or if it takes the form of a linear factor model. However, much richer models are consistent with this setup.

Although we cannot identify the unobserved unit and time components $\alpha_i$ and $\beta_t$ under our assumptions, we can both identify and consistently estimate the expected value $\mu(\alpha_i,\beta_t)=\mathbb{E}[Y_{it}|\alpha_i,\beta_t]$, which is a key building block for both our average treatment effect and decomposition results. Beyond the generative model, this result relies solely on smoothness and restrictions on the asymptotic sequences. We do so by constructing sets of units ${\cal J}(i)\subset\{1,\ldots,N\}$ with increasing cardinality such that for units $j\in{\cal J}(i)$ in this set the function $\mu(\alpha_j,\beta)$ is similar to $\mu(\alpha_i,\beta)$ as a function of $\beta$. We construct these sets by finding units $j$ such that, for all other units $k$, the covariance over time between $Y_{it}$ and $Y_{kt}$ is similar to the covariance between $Y_{jt}$ and $Y_{kt}$. We then estimate $\mu_{it}$ by averaging $Y_{jt}$ for units in such sets ${\cal J}(i).$

For the causal problem, we focus on a setting with a binary treatment $W_{it}$ that affects only unit $i$ in period $t$ (thus ruling effects of the treatment that persist over time). This allows the potential outcomes to be defined in terms of the contemporaneous treatment, $Y_{it}(0)$ and $Y_{it}(1)$. We focus on identification of the average treatment effect for the treated,

equation[equation omitted — 138 chars of source]

although our results can be generalized to other average treatment effects. The primary condition for our result is that conditional on $\alpha_i$ and $\beta_t$, assignment to treatment is independent of $Y_{it}(0)$:

equation[equation omitted — 150 chars of source]

In addition we need an overlap condition that the probability of receiving the treatment conditional on $\alpha_i$ and $\beta_t$ is bounded away from zero. This Latent Factor Unconfoundedness assumption is analogous to unconfoundedness in the cross-section setting. A final condition is that for $\mu(\alpha,\beta)\equiv \mathbb{E}[Y_{it}|\alpha_i=\alpha,\beta]$, $\mathbb{E}_\beta[\mu(\alpha,\beta)-\mu(\alpha',\beta))^2]=0$ implies $\alpha=\alpha'$. Our main result for the treatment effect setting is that given sufficient smoothness on $g(\cdot)$, the average effect for the treated $\tau$ is identified in this setting when both $T$ and $N$ go to infinity.

In addition to causal problems, the methods may also be applied to other purely predictive problems that rely on identification and estimation of $\mu_{it}$. One such category of problems is decomposition of a difference in outcomes (e.g. by demographic group or other categorization) into explained and unexplained components, as is common in the literature on gender wage gaps blau2017gender. We develop this application further in Section (ref).

Related Literature

There are a number of distinct literatures that the current paper builds on. One is the econometric literature on panel data. Much of this literature has focused on Two-Way-Fixed-Effect (TWFE) models and related Difference-In-Differences (DID) estimators, and more recently linear factor models bai2009panel. See chamberlain1984panel, arellano2001panel, arellano2003panel,baltagi2008econometric,hsiao2022analysis, wooldridge2010econometric,arellano2011nonlinear for textbook discussions and arkhangelsky2024causal,roth2022s for recent surveys.

Extensions allowing for nonlinear functions of linear factor models are considered in chen2021nonlinear, yalcin2001nonlinear, feng2020causal, feng2023optimal, freyberger2018non, zeleneev2020identification, abadiecausal.

The current study is also related to the literature on synthetic control methods, abadie2003,abadie2010synthetic, abadie2019using, arkhangelsky2021synthetic, liu2020practical, fry2024method. In that literature the generative models are often not specified explicitly, with the focus on algorithms that lead to effective estimators.

Another literature that is relevant is the research on row and column exchangeable arrays, see aldous1981representations, lynch1984canonical, mccullagh2000resampling. Nonparametric models motivated by the double exchangeability have been considered in network settings, where they are referred to as graphons. The focus in this literature is on estimating the conditional probability of links between nodes in the network, assuming the probability depends in an unrestricted way on unit specific components for both nodes. See for key insights bickel2009nonparametric, bickel2011method, zhang2017estimating, graham2024sparse, chiang2023inference.

Very relevant is also the work by bonhomme2015grouped, bonhomme2022discretizing, cytrynbaum2020blocked on grouped heterogeneity. In this line of work units are grouped through k-means clustering so that units in the same group have the same values for the unobserved unit components.

Set Up

We are interested in estimating causal effects in a panel data setting. We focus on a setting with a binary treatment, with ${\bf W}$ the $N\times T$ matrix of treaments, with typical element $W_{it}\in\{0,1\}$. There are two matrices of potential outcomes ${\bf Y}(0)$ and ${\bf Y}(1)$. We observe the matrix of realized outcomes ${\bf Y}$ with typical element $Y_{it}=Y_{it}(W_{it})$. This notation already imposes the stable unit treatment values assumption or SUTVA rubin1978bayesian, imbens2015causal, ruling out dynamic effects.

One of the estimands that is a main focal point of the current study is the average effect on the treated,

equation[equation omitted — 140 chars of source]
remarkThe focus on the average effect for the treated is for ease of exposition and not essential for the ideas developed in this paper. Similar results hold for the overall average effect.
remarkIn most of the discussion we abstract from the presence of exogenous covariates.

Our starting point, and a key contribution, is a fully nonparametric generative model that captures the relation between outcomes for different units and time periods:

assumptionThe control outcomes satisfy \begin{equation} Y_{it}(0)=g(\alpha_i,\beta_t,\varepsilon_{it}),\hskip5cm {\bf (NFM)}\end{equation} with\\ \\ $(i)$ $\alpha_i$, $\beta_{t}$ and $\varepsilon_{it}$ are finite dimensional, \\ $(ii)$ the $\{\alpha_i\}_{i=1}^N$ are $\{\beta_{t}\}_{t=1}^T$ and $\{\varepsilon_{it}\}_{i,t}$ are jointly independent. In addition the $\alpha_i$ are independent across units, although the $\beta_t$ and $\varepsilon_{it}$ can be autocorrelated, although they need to have stationary distributions.
remarkOne special case and leading example of this structure in empirical work is the Two-Way-Fixed-Effect (TWFE) model, where the function $g(\cdot)$ satisfies \begin{equation} g(\alpha_i,\beta_t,\varepsilon_{it})=\alpha_i+\beta_t+\varepsilon_{it},\end{equation} often with the assumption that $\varepsilon_{it}$ is mean zero, homoskedastic, and independent over time. A second special case is the Linear Factor Model (LFM), where \begin{equation} g(\alpha_i,\beta_t,\varepsilon_{it})=\alpha_i^\top\beta_t+\varepsilon_{it}.\end{equation}
remarkIn both the TWFE and LFM literature the $\alpha_i$ and $\beta_t$ are typically viewed as parameters to be estimated, with assumptions made only about the stochastic properties of the $\varepsilon_{it}$ conditional on any sequence of values for $\alpha_i$ and $\beta_t$. Here we treat the $\alpha_i$ and $\beta_t$ explictly as random variables and make assumptions about their properties. These assumptions do not restrict their distributions, nor do they impose restrictions on the association with covariates in settings when we include those later, in the spirit of Chamberlain's correlated random effects model chamberlain1984panel.
remarkFor any given size matrix ${\bf Y}$, with $N$ rows and $T$ columns, one can write that matrix as a rank $\min(T,M)$ linear factor model $Y_{it}=\alpha_i^\top\beta_t$ without any $\varepsilon_{it}$. One may be able to approximate ${\bf Y}$ using a LFM with substantially fewer factors. This suggest that a linear factor model in itself need not be a restrictive structure. Why then, do we wish to consider nonparametric/nonlinear factor models? One reason is that although the approximation works for a given matrix, the same number of factors and factor loadings need not be sufficient for larger $N$ and $T$, making the representation less attractive as a generative model. To formalize this discussion we later introduce a sequence of populations with restrictions on the sequence of models.
remarkThe generative TFM model can be motivated by row and column exchangeability of ${\bf Y},$ meaning that the distribution of ${\bf Y}$ is not changed by arbitrarily permuting the time or unit indices. See for a general discussion on exchangeability de2017theory, and for definitions in settings with arrays aldous1981representations, lynch1984canonical, mccullagh2000resampling. If ${\bf Y}(0)$ is row and column exchangeable, then there is a representation \[ Y^*_{it}=g(\alpha_i,\beta_t,\varepsilon_{it}),\] with $g(\cdot)$ measurable, such that the distribution of ${\bf Y}^*$ is the same as that of ${\bf Y}$, with $\alpha_i$, $\beta_t$, and $\varepsilon_{it})$ all scalar, jointly independent, and uniformly distributed.
remarkNote that the exchangeability result implies there is a function $g(\cdot)$ in the representation that has scalar arguments. We allow in the TFM model in ((ref)) to have vector-valued arguments. At this stage allowing for vector-valued components does not add generality. However, once we make smoothness assumptions on the function $g(\cdot)$ beyond measurability, this may affect the quality of the approximations. In addition, allowing for vector-valued components facilitates comparisons to the linear factor model in ((ref)).

To be able to be precise about consistency we first need to be precise about the sequences of matrices or populations. Let us index the sequence of matrices by $m=1,2,\ldots.$ The dimensions of the matrix ${\bf Y}_m$ are $N_m$ and $T_m$.

assumption{\sc (Asymptotic Sequences)}\\ $(i)$ $N_m,T_m\rightarrow \infty$,\\ $(ii)$ $g_m(\cdot)$ is identical for all $m$ (so we can omit the subscript $m$).
remarkThere is some important content to this assumption although there is nothing testable. We cannot deal in general with the case where we have $g(\cdot)$ indexed by the population, for example, $g_m(\alpha,\beta,\varepsilon)=h(m\alpha,\beta,\varepsilon)$. In that case, for any given value of $\alpha_1$, we would not be guaranteed to have a lot of units $i$ that have a value $m\alpha_j$ that is close to $m\alpha_1$. So this assumption about the asymptotic sequences captures the idea that as we get more units and more time periods they are filling in the space, not purely expanding it, similar to the way infill asymptotics is utilized in time series analyses and spatial econometrics zhang2005towards.

We also need some smoothness conditions. Note that the row and column exchangeability implies the existence of a measurable function $g_m(\cdot)$. In the smoothness assumption we strengthen that to include continuity and Lipschitz conditions.

assumption{\sc (Smoothness Condition)}\\ $g(\alpha,\beta,\varepsilon)$ is continuously differentiable in all its arguments with first derivatives bounded.
assumption{\sc (Rank Condition)}\\ If $\mathbb{E}_\beta[(\mu(\alpha,\beta)-\mu(\alpha',\beta))^2]=0$, then $\alpha=\alpha'.$

Identification of the Conditional Means

Before focusing on the average treatment effect or decomposition results, we show a preliminary identification result for the expected value of $Y_{it}(0)$ given $\alpha_i$ and $\beta_t$. This is a key technical result that is a building block for the more substantively interesting identification results discussed later. For this result we focus on a predictive setting where we observe the entire matrix ${\bf Y}(0)$, and do not use the Latent Factor Unconfoundedness assumption. Clearly identification of the entire function $g(\cdot)$, or of the unit and time components $\alpha_i$ and $\beta_t$ is not feasible based on Assumptions (ref), (ref), and (ref) without more structure. However, we do not need to estimate either the function $g(\cdot)$ itself or its arguments. For estimating average treatment effects and for our decomposition results it suffices to estimate the conditional expectation of $g(\alpha_i,\beta_t,\varepsilon_{it})$, conditional on $(\alpha_i,\beta_t)$, for all the treated pairs of indices $(i,t)$. Define the conditional expectations \[ \mu(\alpha,\beta)\equiv \mathbb{E}[g(\alpha,\beta,\varepsilon)|\alpha,\beta],\qquad \mu_{it}\equiv \mu(\alpha_i,\beta_t) \] and the residual functions \[ \eta(\alpha,\beta,\varepsilon)\equiv g(\alpha,\beta,\varepsilon)-\mu(\alpha,\beta).\qquad\eta_{it}\equiv\eta(\alpha_i,\beta_t,\varepsilon_{it}).\]

In this section we focus on consistently estimating $\mu_{it}$ for a single unit/period pair $(i,t)$ as a building block. So we want to have an estimator $\hat Y_{it}$ that is a function of ${\bf Y}$ such that

equation[equation omitted — 90 chars of source]

under Assumptions (ref), (ref), and (ref).

We establish this identification result by constructing a set ${\cal J}_m(i)\subset \{1,2,\ldots,N_m\}$ with two properties under Assumptions (ref), (ref), and (ref) (we do not need Assumption (ref) for this discussion): $(i)$ units $j$ in this set have $\mu(\alpha_j,\beta)$ close to $\mu(\alpha_i,\beta)$ for all $\beta$ with high probability, and $(ii)$ the cardinality of the set increases without bounds with $m$. (And of course a similar strategy would be to look for the corresponding set of time indices.) The key insight is that in order to check that unit $j$ and unit $i$ have similar values of the function $\mu(\alpha_i,\beta)$ and $\mu(\alpha_j,\beta)$ for all $\beta$, it is {\it not} sufficient to compare outcomes $Y_{it}$ and $Y_{jt}$ for units $i$ and $j$ over time. Instead we compare their covariances with other units. Specifically, we compare the covariance between unit $i$ and unit $k$ with the covariance between unit $j$ and unit $k$. If for large $T_m$ and $N_m$ those covariances are similar for {\it all} other comparison units $k$, then it must be the case that $\mu(\alpha_i,\beta)$ and $\mu(\alpha_j,\beta)$ are similar for all $\beta$. Note that it does {\it not} mean that $\alpha_i$ and $\alpha_j$ are close, but it means that we can use unit $j$ for estimating the conditional mean for unit $i$ in the same period. A second insight is that although we do not directly observe $\mu(\alpha_i,\beta)$ for unit $i$, we can estimate the covariance between $\mu(\alpha_i,\beta)$ and $\mu(\alpha_j,\beta)$ using the covariance between $Y_{it}$ and $Y_{kt}$ over time.

remarkThree natural approaches for constructing sets ${\cal J}_m(i)$ with the two aforementioned properties do not work. The first is a simple matching strategy where for a given unit $i$ we look for the closest units $j$ in terms of the average value over time, $\overline{Y}_{i\cdot}=\sum_{t=1}^{T_m} Y_{it}/T_m$. This strategy works in a TWFE setting but not in more general settings. To see why this does not work in general suppose $\alpha_i$ and $\beta_j$ are both uniformly distributed on $[0,1]$ and $g(\kappa,\alpha,\beta,\varepsilon)=\alpha(\beta-1/2)$ so that $\mu(\alpha)=0$ for all $\alpha$. The second approach is also a matching approach. Suppose we look for units $j$ that are the closest to unit $i$, that is, units that minimize $\sum_{t=1}^{T_m} (Y_{jt}-Y_{it})^2/T_m$. Although this approach may work in additive settings where $\eta(\alpha_i,\beta_t,\varepsilon_{it})$ does not depend on $\alpha_i,\beta_t)$, it does not work in general. To see that this does not work in general, consider an example where $g(\kappa,\alpha,\beta,\varepsilon)=\alpha(\beta+\varepsilon)$, with $\alpha$, $\beta$ and $\varepsilon$ ${\cal N}(0,1)$ and $\varepsilon$ has a ${\cal N}(0,1)$ distribution. Suppose the treated unit $i$ has $\alpha_i=1$. Then for a unit $j$ with $\alpha_j=\alpha_i=1$, the expected value of $\sum_{t=1}^{T_m} (Y_{jt}-Y_{it})^2/T_m$ is equal to $\mathbb{E}[(\varepsilon_{it}-\varepsilon_{jt})^2]=2.$ In that case unit $i$ will get matched with units $j$ with $\alpha_j\approx 1/2$. A third approach starts with the observation that if $\alpha_j$ is approximately equal to $\alpha_i$, it must be the case that the distribution of $Y_{it}$ over time identical to the distribution of $Y_{jt}$ over time. Let $F_i(y)$ denote the cumulative distribution function of $Y_{it}=g(\alpha_i,\beta_t,\varepsilon_{it})$ conditional on $\alpha_i$. A strategy based on this would look for $j$ such that $\sup_y |F_i(y)-F_j(y)|$ is close to zero. (In practice of course one would need to use estimated versions of these cumulative distribution functions, but with $T_m$ large one could estimate them precisely.) To see that this strategy does not work, consider the case where $g(\alpha,\beta,\varepsilon)=(\alpha-1/2)^2(\beta-1/2)$, with both $\alpha_i$ and $\beta_t$ uniformly distributed on $[0,1]$. Consider a unit $i$ with $\alpha_i=0$. The cumulative distribution function for such a unit is the same as for a unit $j$ with $\alpha_j=1$. Note that if the TWFE model holds, then all three of the aforementioned strategies do work. However, outside of the TWFE model these strategies do not generally work. These three examples of failed strategies show that a challenge is finding a metric that differentiates between different values of the unit (or time) component that lead to different functions $\mu(\alpha_i,\beta_t)$. Our strategy will focus on metrics that involve multiple units rather than simply focus on the marginal distribution of the outcome for a specific unit.

First we define an infeasible version of the set ${\cal J}_m(i)$ that helps develop the intuition for the identification result. Define \[ {\cal J}^*(\alpha)\equiv \left\{ \alpha'\in[0,1],{\rm\ s.t.\ } \sup_{\alpha''} \biggl|\mathbb{E}_\beta\left[\left.\left\{\mu(\alpha,\beta)-\mu(\alpha',\beta)\right\}\mu(\alpha'',\beta)\right|\alpha,\alpha',\alpha''\right] \biggr|=0 \right\} .\]

lemmaIf $\alpha'\in{\cal J}^*(\alpha)$ then $\mathbb{E}_\beta[\{\mu(\alpha,\beta)-\mu(\alpha',\beta)\}^2|\alpha,\alpha']=0.$
remarkIf $\alpha'\in{\cal J}^*(\alpha)$ then $\mu(\alpha,\beta)$ is close to $\mu(\alpha',\beta)$ for all $\beta.$ It is not, however, the case that $\alpha$ and $\alpha'$ are necessarily close.
remarkThe expression $\sup_{\alpha''} |\mathbb{E}_\beta[\{\mu(\alpha,\beta)-\mu(\alpha',\beta)\}\mu(\alpha'',\beta)|\alpha,\alpha',\alpha''] |$ is a $L_\infty$ version of the similarity distance lovasz2010regularity, lovasz2012large.
remarkWe only need the covariances to be the same for $\alpha''\in\{\alpha,\alpha'\}$. However, because we do not observe the values, we end up having to check this for all values of $\alpha''\in[0,1].$
remarkIt is interesting to consider the set ${\cal J}^*(\alpha)$ in the TWFE and linear factor model settings.

We cannot use this first infeasible identification result directly to create, for a given unit $i$, a set of units with similar functions $\mu(\alpha,\beta)$ because we do not know either the values of the $\alpha_i$ nor the function $\mu(\alpha,\beta)$. This leads to three complications and corresponding modifications of the set ${\cal J}^*(\alpha)$ to turn this into an actual identification result. First, we compare the covariances only at the sample values $\alpha_i$, still assuming we know the entire functions $\mu(\alpha_i,\beta)$ as a function of $\beta$, but now only for all the realized values $\alpha_i$. This implies that we cannot find pairs of units with exactly the same $\mu(\alpha_j,\beta)$, only approximately so, with the apporoximation improving as the number of units $N_m$ increases. Second, we only evaluate the $\mu(\alpha,\beta)$ at the sample values of $\beta_t$, instead of taking the expectation over $\beta_t$. In other words, we focus on the average $\sum_t \mu(\alpha_i,\beta_t)\mu(\alpha_k,\beta_t)/T_m$ being close to the average $\sum_t \mu(\alpha_i,\beta_t)\mu(\alpha_k,\beta_t)/T_m$, for all $k$. The quality of the approximation this induces relies on $T_m$ being large. Third, we do not observe the values of the products $\mu(\alpha_i,\beta_t)\mu(\alpha_k,\beta_t)$ for pairs of units $i$ and $k$ at all $t$, we need to estimate these products using the averages $\sum_{t=1}^{T_m} Y_{it}Y_{kt}/T_m$. For this we rely on $T_m$ being large relative to the number of units $i$ and $k$ where we evaluate these averages to ensure those averages approximate the expectations accurately.

The first of our main identification results presents a feasible version of Lemma (ref). It constructs a set of indices ${\cal J}_{m,\nu}(i)\subset\{1,\ldots,N_m\}$. The set is indexed by $\nu$ which governs how close the functions $\mu(\alpha_j,\beta)$ are to $\mu(\alpha_j,\beta)$ for $j\in{\cal J}_{m,\nu}(i) $, with high probability. Define \[ {\cal J}_{m,\nu}(i)\equiv \left\{j=1,\ldots,N_m,{\rm\ s.t.\ } \max_{k\neq i,j} \left|\frac{1}{T_m}\sum_{t=1}^{T_m} (Y_{it}-Y_{jt}) Y_{kt}\right|\leq \nu \right\}.\]

Let $\overline{\mu}$ be an upper bound on the absolute value of $\mu(\alpha,\beta)$, and let $\overline{\mu'}$ be an upper bound on the absolute value of the derivatives of $\mu(\alpha,\beta)$ with respect to $\alpha$ and $\beta.$ Let $C_Y$ be an upper bound on the absolute value of $g(\alpha,\beta,\varepsilon).$

theoremSuppose Assumptions (ref), (ref), and (ref) hold. For any $\xi,\epsilon>0$, if \[N_m\geq\frac{\ln(\epsilon\xi/(16\overline{\mu'}\overline{\mu})}{\ln(1-\xi/16\overline{\mu'}\overline{\mu})},\qquad{\rm and}\quad T_m\geq \frac{256N_m^2 C^2_Y}{\epsilon^2\xi^2},\] then $(i)$ \[ {\rm pr}\left(\mathbb{E}_\beta\left[\left.\left\{\mu(\alpha_i,\beta)-\mu(\alpha_j,\beta)\right\}^2\right|\alpha_i,\alpha_j,j\in {\cal J}_{m,\xi/8}(i) \right]>\xi\right)<\varepsilon,\] and $(ii)$ the cardinality of the set ${\cal J}_{m,\xi/8}(i)$ diverges for all $i$.

Estimating Average Treatment Effects

In this section we use the results from the previous section to estimate average treatment effects under an unconfoundedness type assumption. We start from a slightly different point though, without the generating model $g(\alpha_i,\beta_t,\varepsilon_{it})$.

First, we postulate the existence of a pair of potential outcomes $Y_{it}(0)$ and $Y_{it}(1),$ corresponding to the outcomes without and given the intervention. There is a binary treatment $W_{it}\in\{0,1\}$, with the realized outcome corresponding to the potential outcome given the treatment received, \[ Y_{it}\equiv \left\{

array[array omitted — 87 chars of source]

\right.\] We are interested in the average treatment effect for the treated, \[ \tau\equiv \frac{1}{\sum_{i,t} W_{it}}\sum_{i,t} W_{it} \Bigl(Y_{it}(1)-Y_{it}(0)\Bigr)\qquad {\bf (ATT)}\] The insights extend to other estimands such as the overall average treatment effect.

In order to identify $\tau$ we need some restrictions on the assignment mechanism. Our key assumption is closely related to the standard unconfoundedness/ignorability assumptions in the cross-section evaluation literature rosenbaum1983central, imbens2015causal. Such assumptions are typically stated as independence of the assignment and the set of potential outcomes conditional on observed confounders. Here the critical assumption is formulated differently in two aspects. The first, minor, difference is that we only use {\it weak} unconfoundedness, where we assume independence of the assignment and the control potential outcome, as introduced in imbens2000. Second, the conditioning is on {\it unobserved} confounders, rather than observed confounders. In particular, we assume there are latent unobserved confounders such that conditional on those unconfounders assignment would be as good as random. That by itself is without loss of generality, in other words such an assumption does not have any direct content. Our assumption strengthens this by assuming that this set of latent confounders can be separated into some that are unit-specific and some that are time-specific, and {\it none} that are unit-time specific.

assumption{\sc (Latent Factor Ignorability)}\\ There are unobserved unit components $\alpha_i$ and unobserved time components $\beta_t$ such that the assignment mechanism satisfies $(i)$ (Latent Factor Unconfoundedness): \[ W_{it}\ \perp\!\!\!\perp\ Y_{it}(0)\ \Bigl|\ \alpha_i,\beta_t,\] and $(ii)$ (Latent Overlap): for some $c>0$, \[ c<{\rm pr}(W_{it}=1|\alpha_i,\beta_t)<1-c.\]
remarkIf we strengthened the condition to require only conditioning on the unit component, \[ W_{it}\ \perp\!\!\!\perp\ Y_{it}(0)\ \Bigl|\ \alpha_i,\] the problem would be straightforward. In that case we could immediately use the average outcome for the treated unit during control periods as an estimator for the missing potential outcome. If a treated unit period pair is $(i,t)$ the imputed value would be \[ \hat Y_{it}(0)=\frac{1}{T-1}\sum_{s\neq t}Y_{is}.\] In that case we would discard observations on units other than the treated unit for the purposes of estimating $Y_{it}(0)$. Similarly the problem would be simple to solve if we strengthened the condition to \[ W_{it}\ \perp\!\!\!\perp\ Y_{it}(0)\ \Bigl|\ \beta_t.\] The fundamental challenge we address in this paper is that the Latent Factor Unconfoundedness condition requires conditioning on both $\alpha_i$ and $\beta_t$, and we do not observe either of them. Of course this challenge arises exactly from the concern that units differ in unobserved but relevant aspects, and time periods differ in unobserved but relevant aspects, and which is often the motivation for collecting panel data.
remarkLatent Factor Unconfoundedness has no testable implication. This can be seen directly by considering the case where we observe $\alpha_i$ and $\beta_t$. In that case Latent Factor Unconfoundedness is directly equivalent to a standard unconfoundedness assumption when we view the unit of analysis the unit/time-period pair.
remarkWe do not restrict the dependence of the assignment on the unit and time components.

Here we want to extend the insights from the previous section to the case where we are interested in estimating the average treatment effect for the treated, $\tau=\sum_{i,t} W_{it}(Y_{it}(1)-Y_){it}(0))/\sum_{i,t} W_{it}.$ This involves estimating $\mu_{it}=\mathbb{E}[Y_{it}(0)|\alpha_i,\beta_t]$ for all treated unit/period pairs. The primary challenge relative to estimating $\mu_{it}$ in the previous section is that we cannot estimate the marginal covariance between outcomes for units $i$ and $k$ because we do not observe $Y_{it}(0)$ and $Y_{kt}(0)$ for all units and periods. Moreover, the time units and time periods where we do observe $Y_{it}(0)$ may be systematically different from those where we do not do so. To deal with that we modify the definition of the sets ${\cal J}^*(\alpha)$ and ${\cal J}_{m,\nu}(i)$. To avoid notational confusion we refer to the new sets as ${\cal C}^*(\alpha)$ and ${\cal C}_{m,\nu}(i)$ (where the ${\cal C}$ stands for Causal).

Define the propensity score and the marginal treatment assignment probability \[ e(\alpha_i,\beta_t)\equiv{\rm pr}(W_{it}=1|\alpha_i,\beta_t). \] and the set \[ {\cal C}^*(\alpha)\equiv \left\{ \alpha'\in[0,1],{\rm\ s.t.\ } \sup_{\alpha''} \biggl|\mathbb{E}_\beta\left[\left.\left\{\mu(\alpha,\beta)-\mu(\alpha',\beta)\right\}\mu(\alpha'',\beta)\right|\alpha,\alpha',\alpha''\right] \biggr|=0 \right\} .\]

First define for a subset ${\cal N}$ of $\{1,\ldots,N_m\}$ the set of periods \[{\cal S}({\cal N})=\{t=1,\ldots,T_m| W_{jt}=0\forall j\in{\cal N}\}\] where all the units in the set ${\cal N}$ are in the control group. By the overlap assumption the cardinality of the set ${\cal S}({\cal N})$ will be increasing as $T_m$ increases for a finite set ${\cal N}.$ For a given unit $i$ we now want to construct sets ${\cal J}_{m,\nu}(i)$ such that if $j\in{\cal J}_{m,\nu}(i) $, then $\mu(\alpha_i,\beta)$ is close to $\mu(\alpha_j,\beta)$ for all $\beta$. Previously we operationalized this by ensuring that the expectation of $(\mu(\alpha_i,\beta_t)-\mu(\alpha_j,\beta_t)^2$ is close to zero. This in turn we ensured by requiring that the expectation of $\mu(\alpha_i,\beta_t)-\mu(\alpha_j,\beta_t))\mu(\alpha_k,\beta_t)$ is close to zero for all other units $k$. In fact we only used the fact that this expetation was close to zero at two choices for $k$: at $k=i$ and at $k=j$. Now to deal with the missing $Y_{it}(0)$ we modify this by ensuring that the {\it conditional} expectation of $\mu(\alpha_i,\beta_t)-\mu(\alpha_j,\beta_t))\mu(\alpha_k,\beta_t)$ is close to zero for all other units $k$. The conditioning is designed to ensure that we can compare the expectations for two choices of $k$. This means that we now require that for all pairs $(k,l)$ the expectations $(\mu(\alpha_i,\beta_t)-\mu(\alpha_j,\beta_t))\mu(\alpha_k,\beta_t)$ and $(\mu(\alpha_i,\beta_t)-\mu(\alpha_j,\beta_t))\mu(\alpha_l,\beta_t)$ are close to zero, {\it conditional} on $W_{it}=W_{jt}=W_{kt}=W_{lt}=0$.

Define \[ {\cal J}_{m,\nu}(i)\equiv \left\{j=1,\ldots,N_m,{\rm\ s.t.\ } \max_{k,l\neq i,j}\left\{ \left|\frac{\sum_{t\in{\cal S}(\{i,j,k,l\})}(Y_{it}-Y_{jt}) Y_{kt}}{ \sum_{t\in{\cal S}(\{i,j,k,l\}} 1} \right|\right. \right.\] \[\left.\left.+ \left|\frac{\sum_{t\in{\cal S}(\{i,j,k,l\})}(Y_{it}-Y_{jt}) Y_{lt}}{ \sum_{t\in{\cal S}(\{i,j,k,l\}} 1} \right|\right\} \leq \nu \right\},\]

Then, under suitable rate conditions on $\nu$ we can estimate $\tau$ as \[ \hat\tau=\frac{\sum_{i,t} W_{it}\Bigl( Y_{it}-\hat Y_{it,\nu}(0)\Bigr)} {\sum_{i,t} W_{it}},\qquad {\rm where}\quad \hat Y_{it,\nu}(0)=\frac{\sum_{j\in{\cal J}_{m,\nu}(i)} Y_{jt}}{\sum_{j\in{\cal J}_{m,\nu}(i)} 1}.\]

Nonparametric Decompositions

In this section, we will change the interpretation of $W_{it}$ to denote group membership, and we do not impose unconfoundedness assumptions. Rather, we make use of identities to break out differences in group means into several components. We further modify notation slightly and assume that there is a parallel NPM for units where $W_{it}=1$, so that for $w\in\{0,1\}$:

equation[equation omitted — 88 chars of source]

where the unobserved components $\alpha_i^w$, $\beta_t^w$ and $\varepsilon^w_{it}$ may be vector valued. We further let

equation[equation omitted — 102 chars of source]

Then, we develop a decomposition by first using iterated expectations and substituting in ((ref)), and then adding and subtracting a term:

center[center omitted — 630 chars of source]

If $\alpha_i$ and $\beta_t$ were observed, the analysis above would be an example of standard decompositions (e.g. blau2017gender), where((ref)) and ((ref)) are referred to as the “unexplained” and “explained” gap, respectively. The unexplained gap is the part that remains when the two groups have the same distribution of observables, while the explained gap is the difference that arises when both groups have the same expected outcome function (but different distributions of characteristics). In our case, $\alpha_i$ and $\beta_t$ are unobserved, but the analysis of Section (ref) can be adapted to show that ((ref)) and ((ref)) are identified and can be estimated under the overlap condition.

Conclusion

We propose a nonparametric version of a factor model for estimating average treatment effects in a panel data setting. The model substantially generalizes the commonly used linear factor models and allows for a clear statement of the critical assumptions.

We establish that there is a consistent estimator of the expected non-treated outcome for any given unit and time period, and we use this result to provide conditions under which the average treatment effect is identified.

Our results can also be applied in other settings, such as to problems of decomposing gaps in outcomes by group, as is common in the literature on the gender wage gap.