EconBase
← Back to paper

Identification and estimation of treatment effects in a linear factor model with fixed number of time periods

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

43,562 characters · 8 sections · 33 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification and estimation of treatment effects in a linear factor model with fixed number of time periods

abstractThis paper provides a new approach for identifying and estimating the Average Treatment Effect on the Treated under a linear factor model that allows for multiple time-varying unobservables. Unlike the majority of the literature on treatment effects in linear factor models, our approach does not require the number of pre-treatment periods to go to infinity to obtain a valid estimator. Our identification approach employs a certain nonlinear transformations of the time invariant observed covariates that are sufficiently correlated with the unobserved variables. This relevance condition can be checked with the available data on pre-treatment periods by validating the correlation of the transformed covariates and the pre-treatment outcomes. Based on our identification approach, we provide an asymptotically unbiased estimator of the effect of participating in the treatment when there is only one treated unit and the number of control units is large.

Introduction

The use of panel data has been increasingly popular in empirical socio-economic studies. An essential advantage of panel data is that researchers can obtain consistent estimates of important parameters while controlling for unobserved heterogeneity. The most common version of this approach for identifying the causal effect of a binary treatment on an outcome of interest is the difference in differences (DID) method. The main motivating model for the DID approach is one in which untreated ‘‘potential’’ outcomes are generated by a two-way fixed effects model, a model that allows for unobserved time invariant individual and time fixed effects. This model is consistent with the so-called ‘‘parallel trends’’ assumption underlying the DID approach (see, for example, callaway2022difference for a recent review on the DID approach). This assumption requires that the treated group's counterfactual path of untreated potential outcomes is the same as the path of actual outcomes for the untreated group. However, the parallel trends assumption may only hold in some applications. The main concern with this assumption is that the effect of some time invariant unobserved variable may change over time and therefore cause untreated potential outcomes to follow different paths for the treated group relative to the untreated group.

In this paper, we develop a new approach for identifying and estimating the Average Treatment Effect on the Treated (ATT) when untreated potential outcomes are generated by a linear factor model (also known as an interactive fixed effects model). This model allows for unobserved multiple time-varying individual effects. We establish the identification of the ATT when a researcher observes a vector of time invariant covariates. Based on our identification result, we propose an estimation method that is asymptotically valid when the number of control units goes to infinity. There is a growing literature on treatment effects in linear factor models, and one of the most popular methods is the Synthetic Control (SC) method, proposed in a series of influential papers by abadie2003economic, abadie2010synthetic, and abadie2015comparative. Unlike the majority of the literature, such as synthetic controls and more recent methods based on the original SC estimator (see, for example, abadie2021using for a recent review on the SC method), our approach does not require the number of pre-treatment periods to go to infinity to obtain an asymptotically unbiased estimator of the effect of participating in the treatment. In many applications of SC and related methods to comparative case studies and micro-data, the number of pre-treatment periods may not be large enough for asymptotic justification, and the number of control units may be as large or larger than that of pre-treatment periods (e.g., doudchenko2016balancing).

The key assumption of our identification approach is that certain nonlinear transformations of the observed covariates are sufficiently correlated with the unobserved variables. Although the effects of the covariates and the unobserved variables are not identified, we show that the covariates help identify the effect of participating in the treatment. We can check the relevance condition with the available data by validating the correlation of the transformed covariates and the pre-treatment outcomes instead of the unobserved variables.

The basic idea of our identification approach is to use the transformed covariates as instruments for the unobserved variables. The corresponding instrumental variable estimator is infeasible due to the unobserved variables. However, we need not recover the effects of these variables to identify the effect of participating in the treatment. We show that, when we have access to pre-treatment periods, it suffices to recover the effects of the pre-treatment outcomes in a regression where the individual fixed effects are eliminated. In our approach, we employ an alternative estimator that replaces the unobserved variables with the pre-treatment outcomes for recovering the effect of participating in the treatment.

We develop estimators of the ATT using our identification approach. We propose a two-step estimator where, in the first step, we estimate the effects of the pre-treatment outcomes in a differenced out model. In the second step, we plug these in to obtain estimates of the ATT. For estimation, we focus on the case when only one treated unit exists. This case is the main application of SC methods to comparative case studies. The SC method and our approach are based on opposite ideas. SC method is based on the idea that a weighted average of the outcomes of the control units reconstructs the pre-treatment outcomes of the treated unit and use the pre-treatment periods to estimate these weights. Our approach is based on the idea that a linear combination of the pre-treatment outcomes recovers the outcomes after the treatment period and uses the control units to estimate the corresponding parameters.

This paper is also related to works on linear factor models with a small number of periods. The closely related studies are ahn2013panel, imbens2021controlling, callaway2022treatment, and brown2022unified; similar to the latter three studies, we generalize the DID-type approach and obtain identification by providing some additional conditions. imbens2021controlling employ post-treatment observations under an additional serial independence assumption, and callaway2022treatment employ some covariates that have time invariant effects on untreated potential outcomes to recover average causal effects. Our approach may have broader applicability than these methods. Because our approach does not require knowledge of post-treatment observations, our approach can be applied just after the timing of treatment. Although our estimator is constructed to have good properties for short periods, our estimator can also be applied when the number of periods is large enough for all the covariates to have time varying effects on untreated potential outcomes. Recently, brown2022unified propose a general framework for incorporating factor-model estimators for treatment effect estimation with a small number of periods. Their framework is more general and similar to the identification approach of callaway2022treatment and imbens2021controlling. Our identification approach is also similar to that of brown2022unified, but differs in how we use covariates. For factor-model estimation, they mainly focus on the estimation approach of ahn2013panel, who propose an estimator of linear factor models with a small number of periods. The estimation approach of ahn2013panel allows general instruments as well as time varying covariates and is more general than our approach. Our approach focus on the case where we have time invariant covariates that is included in the model and we employ the linear structure for identifying the treatment effect.

The outline of the paper is as follows. Section (ref) introduces the linear factor model and our basic assumptions. Section (ref) provides our main identification arguments. Section (ref) proposes an estimation method based on the identification results. Section (ref) provides Monte Carlo simulations that compare our approach to DID and SC methods. Section (ref) concludes. We provide the the proofs of the results in the Appendix.

Settings

In this section, we introduce the linear factor model and our basic assumptions. Suppose we have a balanced panel of $N$ units observed on a total of $T$ periods. Let $Y_{it}\in \mathbb{R}$ denote an observed outcome in a particular period $t$. Individuals either belong to a treated group or an untreated group. We set $D_i\in\{0,1\}$ as a treatment indicator so that $D_i = 1$ for individuals in the treated group and $D_i = 0$ for the untreated group. We assume that treatment occurs in period 0, which is the same across all individuals, and we have access to $T_0$ and $T_1+1$ periods of data indexed by $t=-T_0,\ldots,-1$ and $t=0,\ldots,T_1$ before and after the policy change, respectively. Individuals have treated potential outcomes and untreated potential outcomes at each period. We denote these variables by $Y_{it}(1)$ and $Y_{it}(1)$, respectively, for $t = -T_0,\ldots,T_1$. In pre-treatment periods, we have $Y_{it} = Y_{it} (0)$ for all $i$ and $t < 0$; that is, we observe untreated potential outcomes for all individuals in these periods. When $t \geq 0$, we have $Y_{it} = D_i Y_{it}(1) +(1-D_i) Y_{it}(0)$; that is, in period 0 and subsequent periods, we observe treated potential outcomes for individuals in the treated group and untreated potential outcomes for individuals in the untreated group.

We also suppose that we observe a vector of time-invariant covariates $Z_i\in\mathbb{R}^d$. A key difference between our approach and standard panel data models is that we treat the treatment variable asymmetrically from the other covariates. Unlike the traditional estimation approach, the objective in the treatment effects literature is not to estimate the effect of the covariates but to control for them.

The parameter of interest is the Average Treatment Effect on the Treated (ATT); the average treatment effect on the treated conditional on $Z=z$ in period $t$ is

equation*[equation* omitted — 108 chars of source]

Because $E[Y_{it}|D_i=1, Z_i=z]$ is immediately identified from the sampling process, we consider the identification and estimation of $E[Y_{it}(0)|D_i=1, Z_i=z]$ for $t = 0, \ldots, T_1$ when $(T_0,T_1)$ is fixed.

We first impose the following structure on the data generating process in each period.

assumptionThe data $(\{Y_{it}(0), Y_{it}(1)\}_{t=-T_0}^{T_1},D_i,Z_i)_{i=0}^{N-1}$ is independent over $i$ and satisfies the following model: For $i=0,1,\ldots,N-1$, \begin{eqnarray} Y_{it} &=& \begin{cases} D_i Y_{it}(1) +(1-D_i) Y_{it}(0), & t=0,\ldots,T_1, \\ Y_{it}(0), & t=-T_0,\ldots,-1, \end{cases}, \\ Y_{it}(0) &=& b_t' Z_i + F_t' \lambda_i + \epsilon_{it}, \ \ \ t = -T_0. \ldots, T_1, , \end{eqnarray} where $b_t\in\mathbb{R}^d$ is a vector of structural parameters, $\lambda_i\in\mathbb{R}^r$ is a vector of factor loadings, $F_t\in\mathbb{R}^r$ is a vector of common factors, and $\epsilon_{it}\in\mathbb{R}$ is an idiosyncratic error. We treat $F_t$ as a non-random vector.

We assume that in all periods, untreated potential outcomes are generated by a linear factor model with multiple time invariant unobservables whose ‘‘effects’’ can change over time. \footnote{We can include an individual fixed effect into the term $F_t' \lambda_i$ by having $\lambda_i$ include an element where the corresponding element of $F_t$ is equal to one. Similarly, we can include a time fixed effect into the term $b_t' Z_i$ by having $Z_i$ include an intercept where the corresponding element of $b_t$ is equal to the time fixed effect.} The SC methods also employ linear factor models to establish their validity.

We highlight some important comments on the model in Assumption (ref). First, Assumption (ref) only imposes structure on how untreated potential outcomes are generated and allows for essentially unrestricted treatment effect heterogeneity. This structure is standard in the literature on treatment effects with panel data. Second, Assumption (ref) allows the distributions of $\lambda_i$ and $Z_i$ to be different for individuals in the treated and untreated groups. As in fixed effects models, Assumption (ref) also allows for $\lambda_i$ and $Z_i$ to be arbitrarily correlated.

Contrary to the DID literature, the linear factor structure in (ref) allows the effect of the unobserved $\lambda_i$ to vary over time. Suppose one removes the linear factor structure from the model in Assumption (ref). In that case, this model is a two-way fixed effects model consistent with the conditional parallel trends assumptions. Therefore, if we observe both $\lambda_i$ and $Z_i$, we could use the conditional DID approach. Because we do not observe $\lambda_i$, we need to make some modifications.

We also make the following assumptions throughout the paper:

assumptionWe have $E[\epsilon_{it}|D_i,Z_i]=0$, $E[Z_i Z_i'|D_i=0]$ is full rank, and the support of $(D_i,Z_i)$ is $\{0,1\} \times \mathcal{Z}$, where $\mathcal{Z}$ is the support of $Z_i$.

First, Assumption (ref) assumes that the idiosyncratic error $\epsilon_{it}$ is mean independent of $D_i$ and $Z_i$. In the literature on linear factor models for causal effects, many papers assume $E[\epsilon_{it}|\lambda_i, D_i, Z_i]=0$; i.e., $\epsilon_{it}$ is mean independent of $\lambda_i$ as well as $D_i$ and $Z_i$ (e.g., gobillon2016regional, freyberger2018non, and imbens2021controlling.) Our approach allows an arbitrary correlation between $\epsilon_{it}$ and $\lambda_i$. Second, Assumption (ref) assumes that the covariates $Z_i$ have no perfect multicollinearity and that the support of $Z_i$ is sufficiently large for both treated and untreated groups. These requirements are standard in linear regression models.

Identification

Intuition of our identification challenge

In this section, we outline the challenges that we face in identifying ATT when untreated potential outcomes are generated by the interactive fixed effects model in Assumption (ref). To explain the intuition of our identification challenge, we consider a simple case with no covariates where $b_t = 0$ for all $t$. For this case, brown2022unified proposed a general identification scheme. Although the following explanation can be seen as interpreting the framework presented in brown2022unified, we discuss this here for completeness.

Let $\bm{Y}_{i, pre} \equiv \left( Y_{i,-1}, \ldots, Y_{i,-T_0} \right)'$, $\bm{\epsilon}_{i,pre} \equiv (\epsilon_{i,-1}, \ldots , \epsilon_{i,-T_0})'$, and $\bm{F} \equiv (F_{-1}, \ldots , F_{-T_0})$. Assume that we have at least $r$ pre-treatment periods, i.e., $T_0\geq r$, and that $\bm{F}$ is full row rank, which requires there is sufficient variation in $F_t$ in the pre-treatment periods. Because no one is treated before period 0, we observe untreated potential outcomes in the pre-treatment periods for both treated and untreated individuals, and $\bm{Y}_{i, pre}$ can be written as follows:

equation[equation omitted — 94 chars of source]

If $T_0\geq r$ and $\bm{F}$ is full row rank, we can obtain an expression for $\lambda_i$ by re-arranging the terms in (ref):

equation[equation omitted — 93 chars of source]

where for any matrix $\bm{B}$, $\bm{B}^{+}$ is the Moore-Penrose generalized inverse matrix of $\bm{B}$. Plugging in the expression for $\lambda_i$ in (ref) to (ref) and re-arranging terms give, for all $t=0,\ldots,T_1$,

equation[equation omitted — 115 chars of source]

where we define $F^*_t = \bm F^+F_t$. Note that (ref) holds for both the treated and untreated groups. Consequently, we can eliminate the time invariant unobservables from the expression of $Y_{it}(0)$ if we know the value of $F^*_t$.

Then, for identifying ATT, notice that by Assumptions (ref) and (ref),

equation[equation omitted — 88 chars of source]

holds for any time period $t\geq0$ because $F^*_t$ is non-random. Hence, if $F_{t}^*$ is identified, then $E[Y_{it}(0)|D_i=1]$ is also identified because all the other terms in the right-hand side of (ref) are immediately identified from the sampling process.

The above argument suggests that the key identification challenge is to recover the parameter $F^*_t$. It is also worth pointing out that identifying $F^*_t$ is easier than identifying $F_t$ itself in all the periods. In particular, we do not need to impose normalizations on $F_t$ that are common in the interactive fixed effects literature.

Because, in the post-treatment periods $t=0,\ldots,T_1$, $Y_{it}(0)$ is observed for individuals in the untreated group, (ref) suggests that we can recover $F^*_t$ using the untreated group if $\bm{Y}_{i, pre}$ is uncorrelated with $\bm \epsilon_{i,pre}$. However, because they are correlated by construction, $F^*_t$ is not identified by simply regressing $Y_{it}(0)$ on $\bm{Y}_{i, pre}$ for individuals in the untreated group.

remarkAs mentioned above, brown2022unified propose a general identification scheme for incorporating factor-model estimators for treatment effect estimation with a small number of periods. Equation (ref) can be obtained by applying Theorem 1 of brown2022unified. However, our approach differs in how we use covariates to recover the key parameter $F^*_t$.

Main identification results

In this section, we present our main identification strategy. As a first step, in the spirit of Frisch-Waugh-Lovell Theorem, we rewrite (ref) as

equation[equation omitted — 108 chars of source]

where we define $\xi_{it} \equiv Y_{it}(0) - E[Y_{it} Z_i'|D_i=0] E[Z_i Z_i'|D_i=0]^{-1} Z_i$ and $\eta_i \equiv \lambda_i - E[\lambda_i Z_i'|D_i=0] E[Z_i Z_i'|D_i=0]^{-1} Z_i$. From an argument analogous to Section (ref), we obtain an expression for $\xi_{it}$ in terms of their observed counterparts $\bm{\xi}_{i, pre} \equiv (\xi_{i,-1},\ldots,\xi_{i,-T_0})'$:

equation[equation omitted — 120 chars of source]

Similar to (ref), (ref) follows from plugging in the expression of $\eta_i$ obtained from the pre-treatment periods to (ref). Then, by Assumptions (ref) and (ref),

equation[equation omitted — 103 chars of source]

holds for any period $t\geq0$. Therefore, if $F_{t}^*$ is identified, then $E[\xi_{it}|D_i=1,Z_i=z]$ and hence $E[Y_{it}(0)|D_i=1,Z_i=z]$ is also identified because all the other terms in the right-hand side of (ref) are immediately identified from the sampling process.

The central idea of our identification strategy is to employ correlation between $\xi_{it}$ and certain nonlinear transformations of $Z_i$ in the pre-treatment periods. For $j=1,\ldots,R$, let $\omega_j:\mathcal{Z} \to \mathbb{R}$ be some nonlinear weight function. Using the weight functions, we construct nonlinear transformations of $Z_i$. We then construct a $R \times T_0$ matrix $\bm{\Omega}$ as follows: \[ \bm{\Omega} \ \equiv \

pmatrix[pmatrix omitted — 205 chars of source]

. \] This matrix is the correlation matrix between $\omega(Z_i) \equiv \left( \omega_1(Z_i), \ldots , \omega_R(Z_i) \right)'$ and $\bm{\xi}_{i, pre}$ for individuals in the untreated group.

We next introduce our main identifying assumption.

assumptionA matrix $\bm{\Omega}$ has rank $r$.

Assumption (ref) is the rank condition and requires that, for individuals in the untreated group, covariates with certain nonlinear transformations are sufficiently correlated with the outcome after eliminating the linear effects of covariates. This rank condition is the key assumption of the identification. Because $\bm{\Omega}$ is a $R \times T_0$ matrix, $T_0 \geq r$ and $R \geq r$ are the necessary conditions for Assumption (ref) to hold. Hence, Assumption (ref) implies that we need at least $r$ pre-treatment periods and $r$ different weight functions.

Assumption (ref) also requests that $\omega(Z_i)$ contains nonlinear functions of $Z_i$. If $\omega_j(Z_i) = c'Z_i$ for some $c \in \mathbb{R}^d$, then we have

eqnarray*[eqnarray* omitted — 85 chars of source]

because $\xi_{it}$ is orthogonal to $Z_i$ by the definition of $\xi_{it}$. Hence, for satisfying Assumption (ref), we need to use nonlinear functions as the weight functions. This requirement is not strong, and we could use, for example, orthogonal polynomials such as Hermite polynomials when $Z_i$ takes sufficiently many values.

An important implication of Assumption (ref) is that, under Assumptions (ref) and (ref), we could obtain a rank decomposition of $\bm{\Omega}$ with two meaningful matrices: $$ \bm{\Omega} \ = \ \bm{H} \bm{F}, $$ where we define a $R \times r$ matrix $\bm{H}$ as \[ \bm{H} \equiv \left( E[ \omega_1(Z_i) \eta_i | D_i=0], \cdots, E[\omega_R(Z_i) \eta_i | D_i=0] \right)'. \] This matrix is the correlation matrix between $\omega(Z_i)$ and $\eta_i$ for individuals in the untreated group. Because $\bm{H}$ is a $R \times r$ matrix and $\bm{F}$ is a $r \times T_0$ matrix, Assumption (ref) implies not only that $\bm{F}$ is full row rank, but also that $\bm{H}$ is full column rank. Hence, Assumption (ref) requires that $\eta_i$, that contains elements of $\lambda_i$ that are nonlinear in $Z_i$, is sufficiently correlated with $\omega(Z_i)$.

To discuss the applicability of Assumption (ref), suppose that $\lambda_i$ is represented as $\lambda_i = h(Z_i) + u_i$, where $u_i$ is defined as $u_i\equiv\lambda_i - h(Z_i)$ and $h:\mathcal{Z} \to \mathbb{R}^r$ is a function of $Z_i$. Without loss of generality, we assume that $h$ is a nonlinear function of $Z_i$ when $h(Z_i)\neq 0$ because linear terms of $Z_i$ are absorbed into the term $b_t' Z_i$ in the model (ref). If $h(Z_i)\neq 0$ and $h$ is nonlinear, Assumption (ref) then requires sufficient correlation between $\lambda_i$ and $\omega(Z_i)$. When $h(Z_i)=0$ and $u_i$ is exogenous, then there is no endogeneity because the terms $F_t' \lambda_i$ and $\epsilon_{it}$ are exogenous in the model (ref), and we could use existing methods such as pooled ordinary least squares for estimation. When $h(Z_i)=0$ and $u_i$ are correlated with either $D_i$ or $Z_i$, Assumption (ref) may hold when $u_i$ is sufficiently correlated with $Z_i$.

Next, we show that Assumption (ref) provides the identification of $F_{t}^*$. When we do not have access to pre-treatment periods, the available moment conditions relevant for $F_t$ are

eqnarray[eqnarray omitted — 168 chars of source]

Suppose that we observe $\eta_i$ as well as $Z_i$. Then, the moment conditions in (ref) suggest that we could recover $F_t$ by using $\omega(Z_i)$ as the instruments for $\eta_i$. However, this estimation approach is infeasible because $\eta_i$ is unobservable. Although $F_t$ itself is not identified, we only need to identify $F^*_t$, which is $F_t$ multiplied by $\bm F^+$, and we show that this parameter is identified using data on pre-treatment periods.

Notice that (ref) combined with Assumption (ref) provides some moment conditions relevant for identifying $F^*_t$. The available moment conditions are

eqnarray[eqnarray omitted — 221 chars of source]

The moment conditions in (ref) suggest that we could recover $F^*_t$ by using $\omega(Z_i)$ as the instruments for $\bm \xi_{i,pre}$. Under Assumption (ref), we prove that $F^*_t$ is expressed as, for any $R \times R$ positive definite weighting matrix $\bm W$,

equation[equation omitted — 92 chars of source]

where we define a $R\times 1$ vector $\bm{\Omega}_t$ as \[ \bm{\Omega}_t \ \equiv \ \left( E[\omega_1(Z_i) \xi_{it} |D_i=0], \ldots, E[\omega_R(Z_i) \xi_{it} |D_i=0] \right)', \] and $\bm W^{1/2}$ is a $R \times R$ positive definite matrix that satisfies $\bm W=\bm W^{1/2}\bm W^{1/2}$. Specifically, when $T_0=r$, the above expression of $F^*_t$ in (ref) is

equation*[equation* omitted — 92 chars of source]

This expression is the population counterpart of the generalized method of moments (GMM) estimator. Although $\bm \Omega$ may not have full rank, $F^*_t$ has an expression analogous to the GMM estimator for $\bm{\xi}_{i, pre}$ instead of $\eta_i$ using $\omega(Z_i)$ as the instruments. Because the sample counterpart of $\bm \Omega$ is observed from the pre-treatment periods, (ref) shows that $F^*_t$ is identified.

We then state our main identification result.

theoremUnder Assumptions (ref)-(ref), $E[Y_{it}(0)|D_i=1, Z_i = z]$ is identified as follows for all $t \in \{0, \ldots, T_1\}$ and $z \in \mathcal{Z}$ from the distribution of $(\{Y_{it}\}_{t=-T_0}^{0},D_i,Z_i)$. For $t \in \{0, \ldots, T_1\}$, we have \begin{equation} E[Y_{it}(0)|D_i=1,Z_i=z] \ = \ \bm{\Omega}_t'\bm W^{1/2} (\bm{\Omega}' \bm W^{1/2})^{+} (E[\bm{Y}_{i,pre}|D_i=1,Z_i=z] - \bm{\beta}' z) + \beta_t' z, \end{equation} where $\beta_t \equiv E[Z_iZ_i'|D_i=0]^{-1} E[Z_i Y_{it}|D_i=0]$ and $\bm{\beta} \equiv (\beta_{-1}, \ldots, \beta_{-T_0})$.

Theorem (ref) establishes identification of $E[Y_{it}(0)|D_i=1, Z_i = z]$ for all $t \in \{0, \ldots, T_1\}$ under Assumptions (ref)-(ref).

Estimation

In this section, we propose an estimation procedure based on our identification result. Specifically, we consider that the treatment effect of a policy change affects only one unit. This case is the main consideration of SC methods and is common in many applications of comparative case studies. We consider the following setting with $N_0=N-1$:

equation*[equation* omitted — 87 chars of source]

We propose a two-step procedure for estimation. In the first step, we estimate all the parameters $\beta_t$, $\bm{\Omega}_t$, and $\bm{\Omega}$. Then, we plug these into the expression for ATT in Theorem (ref).

The expression for ATT in Theorem (ref) suggests the following estimator of $Y_{0,t}(0)$, \[ \hat{Y}_{0,t}(0) \ \equiv \ (\hat F_t^*)' (\bm{Y}_{0,pre} - \hat{\bm{\beta}}' Z_0) + \hat{\beta}_t' Z_0, \] where, for some $R \times R$ matrix $\hat{\bm W}^{1/2}$ that converges in probability to $\bm W^{1/2}$, $\hat F_t^*$ is defined as \[ \hat F_t^* \equiv (\hat{\bm W}^{1/2} \hat{\bm{\Omega}})^+ \hat{\bm W}^{1/2} \hat{\bm{\Omega}}_t, \] $\hat{\bm{\beta}} \equiv (\hat{\beta}_{-1}, \ldots, \hat{\beta}_{-T_0})$, $\hat{\beta}_t$, $\hat{\bm{\Omega}}_t$, and $\hat{\bm{\Omega}}$ are the sample counterparts of $\bm{\Omega}_t$ and $\bm{\Omega}$, respectively. $\hat{\beta}_t$ is defined as \[ \hat{\beta}_t \ \equiv \ \left( \frac{1}{N_0} \sum_{i=1}^{N_0} Z_i Z_i' \right)^{-1} \frac{1}{N_0} \sum_{i=1}^{N_0} Z_i Y_{it} \ \ \text{for $t=-T_0, \ldots , 0$.} \] The $(j,t)$ element of $\hat{\bm{\Omega}}$ is defined as \[ \hat{\Omega}_{j,-t} \equiv \frac{1}{N_0} \sum_{i=1}^{N_0} \omega_j(Z_i) \left( Y_{i,-t} - \hat{\beta}_{-t}'Z_i \right) \] and $\hat{\bm{\Omega}}_t$ is defined as $\hat{\bm{\Omega}}_t \equiv (\hat{\Omega}_{1,t},\ldots,\hat{\Omega}_{R,t})'$.

We make the following assumption:

assumptionThe data $(\{Y_{it}(0), Y_{it}(1)\}_{t=-T_0}^{T_1},D_i,Z_i)_{i=1}^{N_0}$ is identically distributed over $i$, and we have $E[||Y_{it}||_2^4|D_i=0]<\infty$ $E[||Z_i||_2^4|D_i=0]<\infty$, and $E[||\omega_j(Z_i)||_2^4|D_i=0]<\infty$.

Assumption (ref) is a standard condition for establishing the probability limit by the law of large numbers (LLN). Under Assumption (ref), $\hat{\bm{\Omega}}_t$, $\hat{\bm{\Omega}}$, and $\hat{\beta}_t$ converge in probability to the population counterparts.

Notice that the above estimator may not be valid when $\bm{\Omega}$ does not have full rank because the Moore-Penrose generalized inverse of a matrix is not necessarily a continuous function of the elements of the matrix. Hence, $\hat{\bm{\Omega}}^+$ may not converge in probability to $\bm{\Omega}^+$. stewart1969continuity gives a necessary and sufficient condition for the continuity of the Moore-Penrose generalized inverse, and the continuity holds if and only if $\mathrm{rank}(\hat{\bm{\Omega}})=\mathrm{rank}(\bm{\Omega})$ holds for all sufficiently large $N_0$. However, this condition may not be satisfied in our settings, and we suggest using the above estimator when $\bm{\Omega}$ has full rank, i.e., either $R=r$ or $T_0=r$ holds.

The following result shows that, under $R=r$ or $T_0=r$, $\hat{Y}_{0,t}(0)$ is an asymptotically unbiased estimator of $Y_{0,t}(0)$ when $N_0\to\infty$.

theoremSuppose that either $R=r$ or $T_0=r$ holds. Then, under Assumptions (ref)-(ref), \begin{equation*} \hat{Y}_{0,t}(0) = Y_{0,t}(0) - \epsilon_{0,t} + (F_t^*)' \bm{\epsilon}_{0,pre} + O_p\left(\frac{1}{\sqrt{N_0}}\right). \end{equation*}

In general, we do not know the number of factors $r$, and we are likely to set $R$ and $T_0$ large enough for $R\geq r$ and $T_0\geq r$ to hold. When $R> r$ and $T_0> r$, we propose the following alternative estimator \[ \tilde{Y}_{0,t}(0) \ \equiv \ (\tilde F_t^*)' (\bm{Y}_{0,pre} - \hat{\bm{\beta}}' Z_0) + \hat{\beta}_t' Z_0, \] where $\tilde F_t^*$ is defined as \[ \tilde F_t^* \equiv (\hat{\bm{\Omega}}' \hat{\bm W} \hat{\bm{\Omega}} + \delta_{N_0} \bm I_{T_0} )^{-1} \hat{\bm{\Omega}}' \hat{\bm W} \hat{\bm{\Omega}}_t. \] We set $\delta_{N_0}>0$ as a tuning parameter that converges to 0 when $N_0 \to \infty$. The above estimator is motivated by a well-known result (e.g., see ben2003generalized) that for any matrix $\bm B$, we have $(\bm B' \bm B + \delta \bm I)^{-1} \bm B' \to \bm B^+$ when $\delta \to 0$. Instead of using $(\bm W^{1/2} \hat{\bm{\Omega}}')^+$, we use $(\hat{\bm{\Omega}}' \bm W \hat{\bm{\Omega}} + \delta_{N_0} \bm I_{T_0} )^{-1} \hat{\bm{\Omega}}' \bm W^{1/2}$ for an estimator of $(\bm W^{1/2} \bm{\Omega}')^+$.

The following result shows that, when the tuning parameter $\delta_{N_0}$ converges to 0 at an appropriate rate, $\tilde{Y}_{0,t}(0)$ is an asymptotically unbiased estimator of $Y_{0,t}(0)$ when $N_0\to\infty$.

theoremUnder Assumptions (ref)-(ref), \begin{equation*} \tilde{Y}_{0,t}(0) = Y_{0,t}(0) - \epsilon_{0,t} + (F_t^*)' \bm{\epsilon}_{0,pre} + O_p\left(\frac{1}{\sqrt{N_0}} + \delta_{N_0} + \frac{1}{\delta_{N_0} N_0}\right). \end{equation*}

Theorem (ref) states that $\tilde{Y}_{0,t}(0)$ is a valid estimator when the convergence of the tuning parameter $\delta_{N_0}$ is slower than $1/N_0$.

We summarize some interesting features of our estimation approach in the remarks below.

remarkAlthough our estimator is constructed to have good properties even when $T_0$ is small, the bias of our estimator is expected to be smaller when $T_0$ is larger. Observe that the remained bias of our estimator when $N_0$ is sufficiently large is \[ -\epsilon_{0,t} + F_t' (\bm{F} \bm{F}')^{-1}\bm{F} \bm{\epsilon}_{0,pre} = -\epsilon_{0,t} + F_t' \left( \frac{1}{T_0} \sum_{s=1}^{T_0} F_{-s} F_{-s}' \right)^{-1} \left( \frac{1}{T_0} \sum_{s=1}^{T_0} F_{-s} \epsilon_{0,-s} \right). \] Hence, when $\frac{1}{T_0} \sum_{s=1}^{T_0} F_{-s} F_{-s}'$ converges to a positive definite matrix and $\sum_{s=1}^{T_0} F_{-s} \epsilon_{N,-s}$ converges in probability to 0 when $T_0 \to \infty$, the bias term is expected to be as small as the error term $\epsilon_{0,t}$ when $T_0$ as well as $N_0$ is sufficiently large.
remarkOur alternative estimator $\tilde F_t^*$ of $F_t^*$ is a unique minimizer of the following minimization problem: \begin{equation*} \min_{f\in \mathbb{R}^{T_0}} (\hat{\bm{\Omega}}_t - \hat{\bm{\Omega}} f)' \hat{\bm W} (\hat{\bm{\Omega}}_t - \hat{\bm{\Omega}} f ) + \delta_{N_0} ||f||_2^2. \end{equation*} This problem is a type of well known Tikhonov regularization (e.g., see ben2003generalized), and our estimator of $F_t^*$ is related to ridge regression (hoerl1970ridge). For linear factor models with a small number of periods, imbens2021controlling also propose a similar type of regularized estimator in their context and obtain the same rate of convergence (they use the whole sample size) as in Theorem (ref) for the tuning parameter.

Simulation study

In this section, we provide Monte Carlo simulations to illustrate the finite sample properties of our estimation procedure. In particular, we compare the performance of our approach to DID and SC methods. Suppose that the potential outcomes are generated as follows: \[ Y_{it}(0) = b_{1t} + b_{2t} Z_i + F_{1t} \lambda_{1i} + F_{2t} \lambda_{2i} + \epsilon_{it}, \ \ \ t = -T_0, \ldots,0. \] We consider the case where $T_0=5,10$ and $T_1=0$ for the number of periods and $N=40,100$ for the number of units. We set $Y_{00} = 1 + Y_{00}(0)$, and the effect of participating in the treatment is 1. For the observed covariates, we set $Z_i\sim N(0,1)$ for $i=1,\ldots,N_0$ and $Z_0\sim N(1,1)$ for the treated unit. For the effects of these covariates, we set \[

casesb_{1t} = \{(t - 1)/T_0\} + 1 & \\ b_{2t} = \{(t - 1)/T_0\}^2 + 1. &

\] For the unobserved factors, we set \[

cases\lambda_{1i} = \log(1 + Z_i^4) - E[\log(1 + Z_1^4)] + u_{1i} & \\ \lambda_{2i} = 0.5 \{ \exp(-0.2 Z_i) - E[\exp(-0.2 Z_1)] \} + u_{2i}, &

\] where we set $u_{1i},u_{2i}\sim N(0,0.04)$. We generate $2(T_0+1)$ values from $N(0,1)$ and set them as the (fixed) values of $F_{1t}$ and $F_{2t}$. Finally, we set $\epsilon_{it}\sim N(0,1)$, and $(Z_i,\epsilon_{it},u_{1i},u_{2i})$ are independently generated.

Under this setting, DID is not correctly specified because of multiple factors. Similarly, the methods of imbens2021controlling and callaway2022treatment cannot be applied because we have only one post-treatment period, and the effects of covariates are not time invariant.

For our methods, we compare two cases $R=2$ and $3$. The weighting functions are Hermite polynomials, that are $\omega_1(z)=4z^{2}-2$ and $\omega_2(z)=8z^{3}-12z$, and $\omega_3(z)=16z^4 - 48z^2 + 12$, normalized with sample means and sample variances. In this setting, the number of factors is $r=2$. For the $R=2$ case, we consider two types of estimators. The first one directly uses the Moore-Penrose generalized inverse of the sample matrix for estimation. This can be applied when the number of factors is correctly specified. The second one is the alternative estimator with a tuning parameter $\delta_{N_0}$. We select $\delta_{N_0}$ using (delete-one) cross-validation or generalized cross-validation. For the $R=3$ case, we employ the alternative estimator with a tuning parameter $\delta_{N_0}$. For the SC methods, we compare two types of predictors; predictors I: All pre-treatment outcome lags, and predictors II: Half of the pre-treatment outcome lags and covariates. These specifications are based on the suggestions of ferman2020cherry.

table[table omitted — 3,762 chars of source]

Tables (ref)--(ref) contain the results of this experiment for $T_0 = 5$ and $10$. The number of replications is set at 500 throughout. Tables (ref) and (ref) show the bias, standard deviation, and root mean squared error of the proposed estimators. Tables (ref) and (ref) show the bias, standard deviation, and root mean squared error of the DID and SC methods. As $N$ and $T_0$ increase, the bias and standard deviations tend to decrease for the SC and the proposed estimators. Because of the factor structure, the DID method does not perform well for all settings. The SC method with predictors I has a larger bias than the SC method with predictor II. The bias of the proposed estimator is smaller than other methods for most settings especially when the number of weighted functions is appropriately selected and one directly uses the Moore-Penrose generalized inverse of the sample matrix for estimation. The proposed estimators have a smaller RMSE than the DID and the SC methods.

Conclusion

In this paper, we have developed a new approach for identifying and estimating the ATT when untreated potential outcomes have a linear factor structure. We highlight some attractive features of our approach for use in applied work. First, our approach generalizes the common approaches for policy evaluation, such as DID methods, to the cases where the parallel trends assumption may not hold due to the time-varying unobservables. Second, our approach does not require the number of pre-treatment periods to be sufficiently large for valid estimation. Our estimation approach is expected to have good properties against other methods on linear factor models when the researcher does not have access to enough periods of data to validate their large $T_0$ asymptotics.

The main requirement for using our approach is that the researcher has access to time invariant variables that are sufficiently correlated with the pre-treatment outcomes in that a certain nonlinear transformation of these variables has enough correlation with the pre-treatment outcomes. A good candidate for this variable would be a variable that takes many values. Even when there are only discrete variables with small supports, our approach may be applicable if the researcher has access to a variety of covariates and their whole support is large enough for certain nonlinear transformations.