EconBase
← Back to paper

Confidence Intervals of Treatment Effects in Panel Data Models with Interactive Fixed Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

113,926 characters · 13 sections · 100 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Confidence Intervals of Treatment Effects in Panel Data Models with Interactive Fixed Effects

\setlength\abovedisplayskip{2pt plus 0pt minus 2pt} \setlength\belowdisplayskip{2pt plus 0pt minus 2pt}

\thispagestyle{fancy}

abstractWe consider the construction of confidence intervals for treatment effects estimated using panel models with interactive fixed effects. We first use the factor-based matrix completion technique proposed by bai2021matrix to estimate the treatment effects, and then use bootstrap method to construct confidence intervals of the treatment effects for treated units at each post-treatment period. Our construction of confidence intervals requires neither specific distributional assumptions on the error terms nor large number of post-treatment periods. We also establish the validity of the proposed bootstrap procedure that these confidence intervals have asymptotically correct coverage probabilities. Simulation studies show that these confidence intervals have satisfactory finite sample performances, and empirical applications using classical datasets yield treatment effect estimates of similar magnitudes and reliable confidence intervals. Keywords: Bootstrap, Confidence interval, Treatment effects, Panel data analysis, Interactive effects, Matrix completion JEL classification: C01, C21, C31, I18

Introduction

Treatment effects of certain policy interventions on economic entities are often of major interest in economic and econometric studies. In panel data models, estimation of and inference on treatment effects can be formulated as missing data problems. For instance, consider a sample of $N$ units with a policy intervention on units $N_{0}+1,\ldots ,N$ at time $T_{0}$, so that the pre-treatment periods are $1,\ldots ,T_{0}$, the post-treatment periods are $T_{0}+1,\ldots ,T$, and the control units are $1,\ldots ,N_{0}$. Let $ y_{i,t}$ be the potential outcome in the absence of policy intervention for unit $i$ at time $t$, and $y^{+}_{i,t}=y_{i,t}+\Delta _{i,t}$ be the potential outcome under policy intervention for unit $i$ at time $t$. We are interest in the treatment effect $\Delta _{i,t}$ for $(i,t)\in \mathcal{I}_{1}=\{N_{0}+1,\ldots ,N\}\times \left\{ T_{0}+1,\ldots ,T\right\} $, and it is easy to see that calculating $\Delta _{i,t}= y^{+}_{i,t}-y_{i,t}$ requires knowing $y_{i,t}$, which is actually unobservable (missing) for $(i,t)\in \mathcal{I}_{1}$.

Many existing methods of estimating $\Delta _{i,t}$ make efforts to construct “counterfactuals" for $y_{i,t}$, $(i,t)\in \mathcal{I}_{1}$ (e.g., abadie2010synthetic; hsiao2012panel; bai2021matrix; athey2021matrix; among others). The underlying models in counterfactual frameworks can be simplified as $y_{i,t}=\xi _{i,t}+e_{i,t}$, where $\xi _{i,t}$ is the systematic part and $e_{i,t}$ is the idiosyncratic error. Then the predicted values $\left\{ \widehat{\xi }_{i,t}:(i,t)\in \mathcal{I} _{1}\right\} $ serve the role as the counterfactuals for $\left\{ y_{i,t}:(i,t)\in \mathcal{I}_{1}\right\} $, and the treatment effects are estimated by $\widehat{\Delta }_{i,t}=y^{+}_{i,t}-\widehat{\xi }_{i,t}$ for every $(i,t)\in \mathcal{I}_{1}$. According to model specifications, assumptions, and estimation strategies, we can roughly classify the existing methods into the following 4 categories.

The first category is the synthetic control method proposed by abadie2003economic and abadie2010synthetic. It approximates $y_{N,t}$ by a convex combination of $\left( y_{1,t},\ldots ,y_{N-1,t}\right) $, which can be obtained from a constrained regression. Generalisations of the synthetic control method can be found in, for example, amjad2018robust, abadie2021penalized, arkhangelsky2021synthetic, ben2021augmented, ben2021synthetic, kellogg2021combining, and masini2021counterfactual. One can see abadie2021using for a comprehensive review. The second category is the the panel data approach proposed by hsiao2012panel, which constructs the counterfactual by solving a linear regression of $y_{N,t}$ on $\left( y_{1,t},\ldots ,y_{N-1,t}\right)$. This approach is later extended by ouyang2015treatment, li2017estimation, and hsiao2022multiple. Methods in the third category use the factor model technique (e.g., bai2002determining; bai2003inferential; pesaran2006estimation; bai2009panel; moon2015linear) to construct counterfactuals. Examples include kim2014divorce, xu2017generalized, and bai2021matrix. The last category starts from the statistical learning literature on matrix completion (e.g., candes2009exact; candes2010matrix; mazumder2010spectral; gamarnik2016note), and a recent example is athey2021matrix, which completes the potential outcome matrix via nuclear norm regularisation.

In practice, a mere point estimator of the treatment effect $\Delta _{i,t}$ is not sufficient either theoretically or empirically, and a valid inference is needed, which usually requires the (asymptotic) distribution of $\left( \widehat{ \Delta }_{i,t}-\Delta _{i,t}\right) $. As $\widehat{\Delta }_{i,t}-\Delta _{i,t}=\left( \xi _{i,t}-\widehat{\xi }_{i,t}\right) +e_{i,t}$, the idiosyncratic error term $e_{i,t}$ will dominate the distribution of $\left( \widehat{\Delta }_{i,t}-\Delta _{i,t}\right) $ as $ N_{0},T_{0}\rightarrow \infty $ whenever the systematic part can be consistently approximated. Due to this issue, most studies make inferences on time series average treatment effects as the number of post-treatment periods $\left( T-T_0 \right) \to\infty$ (e.g., hsiao2012panel; li2017estimation; li2020statistical, and chernozhukov2018t), or make inferences on cross-sectional average treatment effects as the number of treated units $\left( N-N_0 \right) \to\infty$ (e.g., xu2017generalized; arkhangelsky2021synthetic; ben2021synthetic; bai2021matrix), or assume the normality of error terms (e.g., fujiki2015disentangling; bai2021matrix). However, these approaches can not simultaneously meet the following needs in empirical studies.

enumerate[label=(\arabic*), labelindent=\parindent, leftmargin=*, nosep] • Inferences do not rely on any specific distributional assumptions ( e.g., normality) on the error terms. • Inference are built on small number of post-treatment period, either because the observed post-treatment time span is short or in order to avoid confounding effects from other interventions in longer time span. • Inferences are performed not only for the (cross-sectional or time series) average treatment effect, but also for the treatment effect on every treated unit at each post-treatment period.

To the best of our knowledge, there are only few inferential studies simultaneously satisfy the three conditions above. abadie2010synthetic use a cross-sectional permutation strategy, where they apply the synthetic control method to every potential control unit to approximate the true distribution of treatment effect estimator under the null hypothesis of $\Delta _{N,t}=0$. Obviously, the validity of this permutation approach relies on the cross-sectional exchangeability of the units. Furthermore, if one wants to construct the confidence intervals of the treatment effects in this fashion, then the equality between distributions of outcomes in the absence of and under intervention up to a location shift is required. chernozhukov2021exact apply time series permutation test to the inference on treatment effects. For a given null hypothesis $H_{0}:\Delta _{N}=\delta _{N}$, they plug $\delta_{N}$ into the observed data matrix to compute the full outcome matrix in the absence of intervention, and then perform time series permutations of the residuals from the treated unit to obtain an approximation of the distribution of the test statistic under $ H_{0}$ and the critical value. Confidence intervals can be obtained by grid searching on a set of different values of $\delta_{N}$, which may lead to intensive computation. As the latest work, cattaneo2021prediction build prediction intervals for synthetic control methods on conditional prediction intervals and non-asymptotic concentration. Since concentration inequalities only guarantee lower bounds of coverage probabilities, the proposed prediction intervals may suffer from conservativeness.

This paper aims to provide a theoretically non-conservative and computationally simple inferential procedure for the estimated treatment effects that simultaneously satisfy the above conditions, without imposing the cross-sectional exchangeability assumption. To be more specific, we apply bootstrap method to the treatment effect estimators proposed by bai2021matrix. Note that bai2021matrix's model keeps a simple factor structure but is quite general in the sense that it allows for multiple treated units and heterogeneous time of intervention. Bootstrapping is one of most commonly used tools to approximate an unknown distribution via resampling, and is shown to be asymptotically valid under mild conditions. We construct bootstrap-based confidence intervals of treatment effects in panel data models with interactive fixed effects. Specifically, we first apply the factor-based matrix completion technique proposed by bai2021matrix to obtain point estimators of treatment effects. Then we follow gonccalves2017bootstrap to approximate the distribution of $\widehat{ \Delta }_{i,t}-\Delta _{i,t}$ by (block) wild bootstrap and bootstrap, and use the quantiles of bootstrapped distribution to construct confidence intervals of $\Delta _{i,t}$ {for every} $(i,t)\in \mathcal{I}_{1}$.

Our paper contributes to the literature in the following folds. First, we establish the validity of confidence intervals by proving that they have asymptotically correct coverage probabilities as $\left( N_{0},T_{0}\right) \rightarrow \infty $, and our proposed confidence intervals do not require any specific distributional assumption on $e_{i,t}$ and are robust to small number of post-treatment periods. Thus, it is theoretically non-conservative and can meet the three needs for inferential purpose of treatment effects estimation in empirical studies. Second, we extend the bootstrap procedure in gonccalves2017bootstrap to a more general case. Since gonccalves2017bootstrap intend to make inference on factor-augmented regression models bai2006confidence where the factors serves as intermediate variables, they only require valid bootstrap approximations of factors. In this paper, we intend to make inference on the factor model (with missing values) itself, and the estimators involve products of estimated factors, their associated loadings, and a rotation matrix. This implies that we need valid bootstrap approximations of factors, loadings, and the rotation matrix, which creates theoretical challenges. Nevertheless, we successfully establish the asymptotic validity of our proposed bootstrap for such a scenario.

The finite sample properties of proposed bootstrap construction of confidence intervals for estimated treatment effects are investigated through Monte Carlo simulation, using different data generating processes (DGPs) for models with or without exogenous covariates, and for models with heteroscedastic errors or serially correlated errors. The simulation results show that the proposed bootstrap procedure works remarkably well for constructing confidence intervals for estimated treatments. Namely, the empirical coverage ratios are quite close to nominal values in all cases we considered regardless of whether the number of unobserved factors are priorly given or estimated from data.

Finally, we revisit the benefits of political and economic integration of Hong Kong with Mainland China in hsiao2012panel, as well as the effectiveness of the California Tobacco Control Program (CTCP) on per capita cigarettes consumption and personal health expenditures in abadie2010synthetic. We note that the treatment effects estimated using our new method are of similar magnitudes to those in the literature, while our bootstrap procedure provides reliable confidence intervals showing that the impacts of CTCP were significant over time for both per capita cigarettes consumption and personal health expenditures, and that the impacts of political and economic integration of Hong Kong with Mainland China vanished over time, although significant in the first few years.

The rest of this article is organised as follows. Section (ref) establishes the model and the assumptions. Section (ref) focuses on the estimation and inference in a model without covariates, and Section (ref) deals with a model with covariates. We conduct Monte Carlo simulations in Section (ref), and apply our method to classical datasets in Section (ref). Section (ref) concludes. Mathematical proofs of main results are left to online appendices.

Notations: we introduce some notations that are frequently used throughout this paper. $\mathbb{Z}_{+}=\{1,2,\ldots \}$ is the set of positive integers. $I_r$ is the $r\times r$ identity matrix. $A^{\mathsf{T}}$ denotes the transpose of matrix $A$. For a vector $z$, let $\left\Vert z\right\Vert =\sqrt{z^{\mathsf{T}}z}$ be the Euclidean norm of $z$. For a matrix $A$, let $\left\Vert A\right\Vert =\sqrt{ \mathrm{tr}\left( A^{\mathsf{T}}A\right) }$ be the Frobenius norm of $A$. $ \mathbb{N}\left( \mu ,\Sigma \right) $ stands for a normal distribution with mean $\mu $ and variance $\Sigma $. And $\xrightarrow{\mathrm{a.s.}}$, $ \xrightarrow{\mathbb{P}}$ and $\xrightarrow{\mathrm{d}}$ denote almost sure convergence, convergence in probability and convergence in distribution, respectively. We let $M$ denote a generic finite positive constant, whose value does not depend on $N$ or $T,$ and may vary case by case.

The Model

Suppose there are observations $\left( \mathscr{Y}_{i,t},x_{i,t}\right) $ for $i=1,\ldots ,N$ and $t=1,\ldots ,T$. Let the dummy variable $d_{i,t}$ indicate the $i$-th unit's treatment status at time $t$ with $d_{i,t}=1$ if under the treatment and $d_{i,t}=0$ if not. The observed data takes the form

equation[equation omitted — 63 chars of source]

where $y_{i,t}$ is the latent outcome of unit $i$ at time $t$ in the absence of treatment, and $\Delta_{i,t}$ is the treatment effect on unit $i$ at time $t$. For the ease of exposition, we assume $d_{i,t}=1$ for all $(i,t)\in \mathcal{I}_{1}$ and $d_{i,t}=0$ for all $(i,t)\in \mathcal{I}\backslash \mathcal{I}_{1}$, where

equation*[equation* omitted — 190 chars of source]

with $1\leq N_{0}<N$ and $1\leq T_{0}<T$. That is, we assume the last $ N-N_{0}$ units are intervened by the treatment at time $T_{0}$. Therefore, the observed matrix of outcomes is

equation*[equation* omitted — 570 chars of source]

Note that the above specification can be generated to allow for heterogeneous intervention time periods by letting $T_{0}$ be the time of the earliest intervention. This is consistent with the definition of $\mathcal{I}_{1}$ in bai2021matrix that $\mathcal{I}_{1}$ corresponds to the \textquotedblleft largest possible" missing block of the matrix.

We are interested in measuring the treatment effects of the policy intervention for the treated units after time $T_{0}$, which are $\Delta _{i,t}$ for $(i,t)\in \mathcal{I}_{1}$. Note that $\Delta _{i,t}$ measures the difference of the outcomes with and without the intervention of the treatment in the post-treatment periods. Unfortunately, the outcomes with and without the intervention of treatment cannot be simultaneously observed in reality for the same unit at a given time. This is because, once the policy intervention is in effect, then the researchers can only observe $\mathscr{Y} _{i,t}=y_{i,t}+\Delta _{i,t}$, not $y_{i,t}$. Thus, in order to estimate $ \Delta _{i,t}$, we need to generate the counterfactual of $y_{i,t}$ for $ (i,t)\in \mathcal{I}_{1}$, denoted as $\widehat{y}_{i,t}$, and thus the treatment effect can be estimated as

equation[equation omitted — 108 chars of source]

Our goal in this paper is to estimate $\Delta _{i,t}$ for $(i,t)\in \mathcal{ I}_{1}$ when neither treated units $N_{1}=N-N_{0}$ nor the post treated periods $T_{1}=T-T_{0}$ is large, where the former can be formulated as average treatment effects across units (see hsiao2022multiple for the aggregation of multiple treatment effects), and the latter can be formulated as the average treatment effects across post-treatment periods (see fujiki2015disentangling and li2017estimation).

Following the literature of treatment effects estimation using panels (e.g., abadie2010synthetic; hsiao2012panel; and bai2021matrix), we assume that $y_{i,t}$ is generated by the following panel data model with interactive fixed effects:

equation[equation omitted — 148 chars of source]

where $x_{i,t}$ is a $p$-dimensional vector of observed covariates of unit $ i $ at time $t$, $e_{i,t}$ is the idiosyncratic error term of unit $i$ at time $t$, $f_{t}$ is an $r$-dimensional time-variant unobserved factor at time $t$, and $\lambda _{i}$ is an $r$-dimensional individual specific factor loading of unit $i$, where $r$ is the number of unobserved factors and is usually unknown to researchers. \footnote{ Even if $r$ is unknown in practice, it can be consistently estimated using the method described in Section 4 and 5 of bai2002determining or the method proposed by alessi2010improved, and thus we can treat $r$ as known in the theoretical analyses below. This can be formally justified as follows. Let $\theta $ be a parameter in the model of interest and $\widehat{ \theta }$ be the estimator of $\theta $ based on estimated number of factors $\widehat{r}$. Note that $r$ takes values in $\mathbb{Z}_{+}$, and then the consistency of its estimator $\widehat{r}$ implies that $\mathbb{P}\left( \widehat{r}=r\right) \rightarrow 1$ as $\left( N,T\right) \rightarrow \infty $. Therefore,

align*[align* omitted — 410 chars of source]

}

Let $y_{i}=\left( y_{i,1},\ldots ,y_{i,T}\right) ^{\mathsf{T}}$, $ X_{i}=\left( x_{i,1},\ldots ,x_{i,T}\right) ^{\mathsf{T}}$, and $ e_{i}=\left( e_{i,1},\ldots ,e_{i,T}\right) ^{\mathsf{T}}$ for every $i\in \{1,\ldots ,N\}$. Define $Y=\left( y_{1},\ldots ,y_{N}\right) $, $F=\left( f_{1},\ldots ,f_{T}\right) ^{\mathsf{T}}$, and $\Lambda =\left( \lambda _{1},\ldots ,\lambda _{N}\right) ^{\mathsf{T}}$. We make the assumptions below.

assumThe factors and loadings satisfy the following conditions. \begin{enumerate}[label=(\arabic*), labelindent=\parindent, leftmargin=*, nosep] • $\mathbb{E} \left( \left\Vert f_t \right\Vert^8 \right)\le M$ for all $ t\in \{ 1,\ldots, T \}$. • $\left\Vert \lambda_{i} \right\Vert \le M$ for all $i\in\{ 1, \ldots, N \}$. • There exists an $r\times r$ positive definite matrix $\Sigma_{F}$, so that \begin{align*} \frac{1}{T}\sum_{t=1}^T f_t f_t^\mathsf{T} \xrightarrow{\mathbb{P}} \Sigma_{F}, \qquad \frac{1}{T_0}\sum_{t=1}^{T_0} f_t f_t^\mathsf{T} \xrightarrow{\mathbb{P}} \Sigma_{F}, \qquad \frac{1}{T-T_0} \sum_{t=T_0+1}^{T} f_t f_t^\mathsf{T} \xrightarrow{\mathbb{P}} \Sigma_{F} \end{align*} as $T$, $T_0\to\infty$. • There exists an $r\times r$ positive definite matrix $\Sigma_{\Lambda}$ , so that \begin{align*} \frac{1}{N}\sum_{i=1}^N \lambda_{i} \lambda_{i}^\mathsf{T} \to \Sigma_{\Lambda}, \qquad \frac{1}{N_0}\sum_{i=1}^{N_0} \lambda_{i} \lambda_{i}^\mathsf{T} \to \Sigma_{\Lambda}, \qquad \frac{1}{N-N_0} \sum_{i=N_0+1}^N \lambda_{i} \lambda_{i}^\mathsf{T} \to \Sigma_{\Lambda} \end{align*} as $N$, $N_0\to\infty$. • The eigenvalues of $\Sigma_{F}\Sigma_{\Lambda}$ are distinct. \end{enumerate}
assumThe distributions of the idiosyncratic errors have the following properties. \begin{enumerate}[label=(\arabic*), labelindent=\parindent, leftmargin=*, nosep] • The idiosyncratic errors have no cross-sectional dependence, i.e., the sequences $\left\{ e_{1,t} \right\}_{t=1}^T$, $\left\{ e_{2,t} \right\}_{t=1}^T$, $\ldots$, $\left\{ e_{N,t} \right\}_{t=1}^T$ are mutually independent. • For all $i\in \{1, \ldots, N \} $, the process $\left\{ e_{i,t} \right\}_{t=1}^\infty$ is strictly stationary and ergodic. • For all $(i,t)\in \mathcal{I}$, the cumulative distribution function of $e_{i,t}$ is everywhere continuous. • The error terms $\left\{ e_{i,t}:(i,t)\in\mathcal{I} \right\}$ are independent of the treatment status $\left\{ d_{i,t}:(i,t)\in \mathcal{I} \right\}$. \end{enumerate}
assumThe moment conditions below hold for the full sample indexed by $\mathcal{I}$, the balanced subsample indexed by $\{1,\ldots, N_0\}\times \{ 1, \ldots, T_0\}$, the control subsample indexed by $\{ 1, \ldots, N_0\}\times \{1, \ldots, T \}$, the pre-treatment subsample indexed by $\{1, \ldots, N \}\times \{ 1, \ldots, T_0\}$, and the missing subsample indexed by $\mathcal{I}_1$. \begin{enumerate}[label=(\arabic*), labelindent=\parindent, leftmargin=*, nosep] • $\mathbb{E}\left( e_{i,t}\right) =0$ and $\mathbb{E}\left( e_{i,t}^{8}\right) \leq M$ for all $(i,t)\in \mathcal{I}$. • For all $t\in\{1,\ldots, T \}$, \begin{align*} \sum_{s=1}^T \left\vert \mathbb{E} \left( \frac{1}{N}\sum_{i=1}^N e_{i,t} e_{i,s} \right) \right\vert \le M. \end{align*} • \begin{align*} \frac{1}{NT}\sum_{i=1}^N \sum_{t=1}^T \sum_{s=1}^T \left\vert \mathbb{E} \left( e_{i,t} e_{i,s} \right) \right\vert \le M. \end{align*} • For all $(s,t)\in \{1,\ldots, T\}^2$, \begin{align*} \mathbb{E} \left( \left\vert \frac{1}{\sqrt{N}}\sum_{i=1}^N \left[ e_{i,s}e_{i,t} -\mathbb{E}\left( e_{i,s}e_{i,t} \right) \right] \right\vert^4 \right) \le M. \end{align*} • \begin{align*} \mathbb{E}\left( \frac{1}{N}\sum_{i=1}^N \left\Vert \frac{1}{\sqrt{T}} \sum_{t=1}^T f_t e_{i,t} \right\Vert^2 \right) \le M. \end{align*} • For all $t\in \{ 1,\ldots, T\}$, \begin{align*} \mathbb{E}\left( \left\Vert \frac{1}{\sqrt{NT}}\sum_{s=1}^T\sum_{i=1}^N f_s \left[ e_{i,s}e_{i,t}-\mathbb{E}\left( e_{i,s}e_{i,t} \right) \right] \right\Vert^2 \right) \le M. \end{align*} • \begin{align*} \mathbb{E}\left( \left\Vert \frac{1}{\sqrt{NT}}\sum_{t=1}^T\sum_{i=1}^N f_t \lambda_i^\mathsf{T} e_{i,t} \right\Vert^2 \right) \le M. \end{align*} \end{enumerate}
assumThe factors, loadings and errors satisfy central limit theorems. \begin{enumerate}[label=(\arabic*), labelindent=\parindent, leftmargin=*, nosep] • For all $t\in\{1,\ldots, T\}$, there exists an $r\times r$ positive definite matrix $\Gamma_{t}$, so that \begin{align*} \frac{1}{\sqrt{N}}\sum_{i=1}^N \lambda_{i} e_{i,t} \xrightarrow{\mathrm{d}} \mathbb{N} \left( 0, \Gamma_{t} \right), \qquad \frac{1}{\sqrt{N_0}} \sum_{i=1}^{N_0} \lambda_{i} e_{i,t} \xrightarrow{\mathrm{d}} \mathbb{N} \left( 0, \Gamma_{t} \right) \end{align*} as $N, N_0 \to\infty$. • For all $i\in \{1,\ldots, N\}$, there exists an $r\times r$ positive definite matrix $\Phi_{i}$, so that \begin{align*} \frac{1}{\sqrt{T}}\sum_{t=1}^T f_t e_{i,t} \xrightarrow{\mathrm{d}} \mathbb{N } \left( 0, \Phi_{i} \right), \qquad \frac{1}{\sqrt{T_0}}\sum_{t=1}^{T_0} f_t e_{i,t} \xrightarrow{\mathrm{d}} \mathbb{N} \left( 0, \Phi_{i} \right) \end{align*} as $T, T_0 \to\infty$. \end{enumerate}
assumThe quantities $N$, $T$, $N_0$, $T_0$ satisfy the conditions below. \begin{enumerate}[label=(\arabic*), labelindent=\parindent, leftmargin=*, nosep] • $TN_0> r\left( T+N_0 \right)$ and $T_0 N > r \left( T_0+N \right)$. • $N$, $T$, $N_0$, $T_0$ are of the same order, i.e., \begin{align*} \lim_{N,N_0\to\infty} \frac{N_0}{N}=c_1\in (0,1], \qquad \lim_{T,T_0\to\infty} \frac{T_0}{T}=c_2\in (0,1], \qquad \lim_{N,T\to\infty} \frac{N}{T}=c_3\in (0,\infty). \end{align*} \end{enumerate}
assumThe conditions below hold for the full sample indexed by $\mathcal{I}$, the control subsample indexed by $\{ 1, \ldots, N_0\}\times \{1, \ldots, T \}$, and the pre-treatment subsample indexed by $ \{ 1, \ldots, N\}\times \{1, \ldots, T_0 \}$. \begin{enumerate}[label=(\arabic*), labelindent=\parindent, leftmargin=*, nosep] • $\mathbb{E}\left( \left\Vert x_{i,t} \right\Vert^8 \right) \le M$ for every $(i,t)\in\mathcal{I}$. • $\inf \left\{ \mathcal{D}(\mathcal{F}) : \mathcal{F}\in \mathscr{F} \right\}>0$, where \begin{align*} \mathcal{D}(\mathcal{F})&=\left[ \frac{1}{NT}\sum_{i=1}^N X_i^\mathsf{T} \left( I_T-\frac{\mathcal{F}\mathcal{F}^\mathsf{T}}{T} \right) X_i \right] - \left[ \frac{1}{TN^2} \sum_{i=1}^N \sum_{k=1}^N X_i^\mathsf{T} \left( I_T- \frac{\mathcal{F}\mathcal{F}^\mathsf{T}}{T} \right) X_k a_{i,k} \right], \\ a_{i,k}&=\lambda_{i}^\mathsf{T} \left( \frac{\Lambda^\mathsf{T} \Lambda}{N} \right)^{-1} \lambda_{k}, \\ \mathscr{F}&=\left\{ \mathcal{F}\in \mathbb{R}^{T\times r}: \left. \mathcal{F }^\mathsf{T} \mathcal{F} \right/ T =I_r \right\}. \end{align*} \end{enumerate}
assumThe conditions below hold for the full sample indexed by $\mathcal{I}$, the control subsample indexed by $\{ 1, \ldots, N_0\}\times \{1, \ldots, T \}$, and the pre-treatment subsample indexed by $\{ 1, \ldots, N\}\times \{1, \ldots, T_0 \}$. \begin{enumerate}[label=(\arabic*), labelindent=\parindent, leftmargin=*, nosep] • \begin{align*} \frac{1}{T}\sum_{t=1}^T \sum_{s=1}^T \max_{1\le i \le N} \left\vert \mathbb{E } \left( e_{i,t} e_{i,s} \right) \right\vert \le M. \end{align*} • $\left\Vert \mathbb{E}\left( e_{i}e_{i}^\mathsf{T} \right) \right\Vert_S \le M$ for all $i\in\{1,\ldots, N\}$, where $\left\Vert \cdot \right\Vert_S$ is the spectral norm of a matrix. • \begin{align*} \frac{1}{T^2 N}\sum_{t=1}^T \sum_{s=1}^T \sum_{u=1}^T \sum_{v=1}^T \sum_{i=1}^N \left\vert \mathrm{Cov} \left( e_{i,t}e_{i,s}, e_{i,u}e_{i,v} \right) \right\vert \le M. \end{align*} • \begin{align*} \frac{1}{TN^2}\sum_{t=1}^T \sum_{s=1}^T \sum_{i=1}^N \sum_{j=1}^N \left\vert \mathrm{Cov} \left( e_{i,t}e_{j,t}, e_{i,s}e_{j,s} \right) \right\vert \le M. \end{align*} • There exists a $p\times p$ positive definite matrix $\Omega$, so that \begin{align*} \frac{1}{\sqrt{NT}}\sum_{i=1}^N \left[ \left( I_T-\frac{FF^\mathsf{T}}{T} \right) X_i -\frac{1}{N}\sum_{k=1}^N a_{i,k} \left( I_T-\frac{FF^\mathsf{T}}{ T} \right) X_k \right]^\mathsf{T} e_i \xrightarrow{\mathrm{d}} \mathbb{N} (0,\Omega). \end{align*} \end{enumerate}
assum$\left\{ e_{i,t}: (i,t)\in \mathcal{I} \right\}$ is independent of $\left\{ f_t \right\}_{t=1}^T$ and $ \left\{ x_{i,t}: (i,t)\in\mathcal{I} \right\}$.

The moment and convergence conditions in Assumptions (ref), (ref), (ref), (ref), and (ref) are used to establish probability bounds and asymptotic distributions. They are quite standard in the literature for panel models with interactive fixed effects, see, for example, bai2003inferential , bai2009panel and bai2021matrix, among others. The distribution and order conditions in Assumptions (ref), (ref), and (ref) mainly work for the validity of our bootstrap procedure.

We make further remarks on some of the assumptions. Assumption (ref)(4) requires the strict exogeneity of treatment status, and one can find similar conditions in Assumption 5 of hsiao2012panel and Assumption 2 of xu2017generalized. Assumption (ref)(3) states that the factors in the pre-treatment subsample and post-treatment subsample have the same asymptotic second sample moment matrices as those in the full sample. Assumption (ref)(4) states that the factor loadings in the control subsample and treated subsample have the same asymptotic second sample moment matrices as those in the full sample. When either $T-T_0$ or $N-N_0$ is finite, Assumption (ref)(3) or (ref)(4) can be replaced with equality as in the identifying restriction PC1 on Page 19 of bai2013principal to fulfil the condition that either the factors or factor loadings possess the same properties in the full sample as well as in the post-treatment periods or for the treated units.

Estimation and Inference in a Model without Covariates

To highlight the essence of our approach for estimation and inference of treatment effects, in this section we focus on a special case of Model (ref), where there is no covariate (i.e., $ \beta=0$) and the model is reduced to a pure approximate factor model:

align[align omitted — 116 chars of source]

We will extend our suggested approach to the full model (ref) in Section (ref).

Estimation of Treatment Effects

Following bai2021matrix, we consider the factor-based approach to estimate $y_{i,t}$ for $(i,t)\in\mathcal{I}_1$. Let $Y$, $F$, and $\Lambda$ follow their definitions in Section (ref). Since $y_{i,t}$ is unobserved for all $(i,t)\in \mathcal{I}_1$, the south-east block of $Y$ is missing in reality. Formally, we define two sub-matrices of $Y$:

align[align omitted — 461 chars of source]

That is, $Y_\mathrm{tall}$ corresponds to the control subsample, and $Y_\mathrm{wide}$ corresponds to the pre-treatment subsample. Let $\left( {F}_\mathrm{tall}, {\Lambda}_\mathrm{tall} \right)$ and $\left( { F}_\mathrm{wide}, {\Lambda}_\mathrm{wide} \right)$ be the factor and loading matrices associated with $Y_\mathrm{tall}$ and $Y_\mathrm{wide}$, respectively. It is easy to see that ${F}_\mathrm{tall}=F$, ${\Lambda}_ \mathrm{wide}=\Lambda$, ${F}_\mathrm{wide}$ is a sub-matrix formed by the first $T_0$ rows of $F$, and ${\Lambda}_\mathrm{tall}$ is a sub-matrix formed by the first $N_0$ rows of $\Lambda$.

In lieu of bai2021matrix, we can use the algorithm below to estimate the treatment effects $\Delta_{i,t}$ for all $(i,t)\in\mathcal{I}_1 $.

algoTreatment effect estimation via factor-based matrix completion (without covariates). \begin{enumerate}[labelindent=\parindent, leftmargin=*, nosep] • Perform a singular value decomposition of $\dfrac{Y_\mathrm{tall}}{ \sqrt{T N_0}}$, and let $D_\mathrm{tall}$ be an $r\times r$ diagonal matrix with the $r$ largest singular values of $\dfrac{Y_\mathrm{tall}}{\sqrt{T N_0} }$ on the diagonal in descending order. Then let $P_\mathrm{tall}$ and $Q_ \mathrm{tall}$ be $T\times r$ and $N_0\times r$ matrices containing the left and right singular vectors of $\dfrac{Y_\mathrm{tall}}{\sqrt{T N_0}}$ respectively, corresponding to $D_\mathrm{tall}$. Compute \begin{align} &\widehat{F}_\mathrm{tall} =\left( \widehat{f}_{\mathrm{tall},1}, \ldots, \widehat{f}_{\mathrm{tall},T} \right)^\mathsf{T} =\sqrt{T}P_\mathrm{tall}, \nonumber \\ & \widehat{\Lambda}_\mathrm{tall} =\left( \widehat{\lambda}_{\mathrm{tall},1}, \ldots, \widehat{\lambda}_{\mathrm{tall},N_0} \right)^\mathsf{T} =\sqrt{N_0} Q_\mathrm{tall} D_\mathrm{tall}. \end{align} • Perform a singular value decomposition of $\dfrac{Y_\mathrm{wide}}{ \sqrt{T_0 N}}$, and let $D_\mathrm{wide}$ be an $r\times r$ diagonal matrix with the $r$ largest singular values of $\dfrac{Y_\mathrm{wide}}{\sqrt{T_0 N} }$ on the diagonal in descending order. Then let $P_\mathrm{wide}$ and $Q_ \mathrm{wide}$ be $T_0\times r$ and $N\times r$ matrices containing the left and right singular vectors of $\dfrac{Y_\mathrm{wide}}{\sqrt{T_0 N}}$ respectively, corresponding to $D_\mathrm{wide}$. Compute \begin{align} &\widehat{F}_\mathrm{wide} =\left( \widehat{f}_{\mathrm{wide},1}, \ldots, \widehat{f}_{\mathrm{wide},T_0} \right)^\mathsf{T} =\sqrt{T_0}P_\mathrm{wide}, \nonumber \\ & \widehat{\Lambda}_\mathrm{wide} =\left( \widehat{\lambda}_{\mathrm{wide},1}, \ldots, \widehat{\lambda}_{\mathrm{wide},N} \right)^\mathsf{T} =\sqrt{N} Q_\mathrm{wide} D_\mathrm{wide}. \end{align} • Let $\widehat{\Lambda}_{\mathrm{wide},0}$ be the first $N_0$ rows of $ \widehat{\Lambda}_\mathrm{wide}$. Compute \begin{align} \widehat{H}_\mathrm{miss}=\widehat{\Lambda}_\mathrm{tall}^\mathsf{T} \widehat{\Lambda}_{\mathrm{wide},0} \left( \widehat{\Lambda}_{\mathrm{wide} ,0}^\mathsf{T} \widehat{\Lambda}_{\mathrm{wide},0} \right)^{-1} , \end{align} and then let $\widehat{C}=\widehat{F}_\mathrm{tall}\widehat{H}_\mathrm{miss} \widehat{\Lambda}_\mathrm{wide}^\mathsf{T}$. • Let $\widehat{c}_{i,t}$ denote the $(t,i)$-th entry of $\widehat{C}$, then compute the residuals $\widehat{e}_{i,t}=y_{i,t}-\widehat{c}_{i,t}$ for $(i,t)\in \mathcal{I}\backslash\mathcal{I}_1$. • For $(i,t)\in\mathcal{I}_1$, the variance of $\widehat{c}_{i,t}$ is estimated by \begin{align} \widehat{\mathbb{V}}_{i,t}&=\frac{1}{T_0} \widehat{f}_{\mathrm{tall},t}^ \mathsf{T} \left( \frac{\widehat{F}_\mathrm{tall}^\mathsf{T} \widehat{F}_ \mathrm{tall}}{T} \right)^{-1} \widehat{\Phi}_i \left( \frac{\widehat{F}_ \mathrm{tall}^\mathsf{T} \widehat{F}_\mathrm{tall}}{T} \right)^{-1} \widehat{ f}_{\mathrm{tall},t} \nonumber \\ &\phantom{=\:\:} +\frac{1}{N_0} \widehat{\lambda}_{\mathrm{wide},i}^\mathsf{T } \left( \frac{\widehat{\Lambda}_\mathrm{wide}^\mathsf{T} \widehat{\Lambda}_ \mathrm{wide}}{N} \right)^{-1} \widehat{\Gamma}_t \left( \frac{\widehat{ \Lambda}_\mathrm{wide}^\mathsf{T} \widehat{\Lambda}_\mathrm{wide}}{N} \right)^{-1} \widehat{\lambda}_{\mathrm{wide},i}, \end{align} where \begin{align} \widehat{\Gamma}_t &=\frac{1}{N_0}\sum_{j=1}^{N_0} \widehat{e}_{j,t}^2 \widehat{\lambda}_{\mathrm{wide},j} \widehat{\lambda}_{\mathrm{wide},j}^ \mathsf{T}, \qquad L_{k,i}=\frac{1}{T_0}\sum_{s=k+1}^{T_0} \widehat{f}_{ \mathrm{tall},s} \widehat{e}_{i,s} \widehat{e}_{i,s-k} \widehat{f}_{\mathrm{ tall},s-k}^\mathsf{T}, \nonumber \\ \widehat{\Phi}_i&= L_{0,i}+\sum_{k=1}^K \left( 1-\frac{k}{K+1} \right) \left( L_{k,i}+L_{k,i}^\mathsf{T} \right), \end{align} with $K\to\infty$ and $\dfrac{K}{T_0^{1/4}}\to 0$ as $T_0\to\infty$. • For every $N_0< i \le N$, compute $\displaystyle \widehat{\sigma}^2_i= \frac{1}{T_0} \sum_{s=1}^{T_0} \widehat{e}_{i,s}^2$. • For $(i,t)\in\mathcal{I}_1$, the estimated treatment effect is $ \widehat{\Delta}_{i,t}=\mathscr{Y}_{i,t}-\widehat{c}_{i,t}$, and the standard error of $\widehat{\Delta}_{i,t}$ is $\sqrt{\widehat{\mathbb{V}} _{i,t}+\widehat{\sigma}^2_i}$. \end{enumerate}
remarkSince $\widehat{\mathbb{V}}_{i,t}=O_{ \mathbb{P}}\left( \dfrac{1}{T_{0}}+\dfrac{1}{N_{0}}\right) $ for every $ (i,t)\in \mathcal{I}_{1}$, it is asymptotically negligible in the standard error of $\widehat{\Delta }_{i,t}$. Here we include $\widehat{\mathbb{V}} _{i,t}$ in the standard error of $\widehat{\Delta }_{i,t}$ to improve finite sample performance.
remarkIn the above algorithm, we discuss the estimation for individual treatment effects for each treated unit. If the interest is the average treatment effects (ATE) across treated units or across post-treatment periods, not individual treatment effects, see hsiao2022multiple for discussion of the aggregation of multiple treatment effects across treated units, and see fujiki2015disentangling and li2017estimation for discussion of the average treatment effects over post-treatment periods.

Construction of Confidence Intervals

When neither $N_1$ nor $T_1$ is large, the average treatment effects either across treated units or post-treatment periods may not be a good measure for inferential purpose since it is hard to establish the asymptotic properties of such effects. In order to provide inferential procedure for individual treatment effects especially when neither $N_1$ nor $T_1$ is large, motivated by gonccalves2017bootstrap, we adapt bootstrap method to construct the confidence intervals of $\Delta_{i,t}$ for $(i,t)\in\mathcal{I} _1$, without a specific distributional assumption on the error term $e_{i,t}$ or particular quantity on $N_1$ or $T_1$ .

algoConfidence intervals of the treatment effects via bootstrap (without covariates). \begin{enumerate}[labelindent=\parindent, leftmargin=*, nosep] • Apply Algorithm (ref) to the sample and obtain $ \left\{ \widehat{c}_{i,t}: (i,t)\in\mathcal{I} \right\}$, $\left\{ \widehat{e }_{i,t}: (i,t)\in\mathcal{I} \right\}$, $\left\{ \widehat{\mathbb{V}}_{i,t}: (i,t)\in\mathcal{I}_1 \right\}$, and $\left\{ \widehat{\sigma}^2_i: i=N_0+1, \ldots, N \right\}$. • For $b=1,2,\ldots, B$: \begin{enumerate}[nosep] • For $(i,t)\in\mathcal{I}\backslash\mathcal{I}_1$, let $ e_{i,t}^*=u_{i,t}\widehat{e}_{i,t}$, where $\left\{ u_{i,t} \right\}$ are i.i.d.\ or block i.i.d.\ from a standard normal distribution $\mathbb{N} (0,1)$ and are independent of the raw sample. • For $(i,t)\in\mathcal{I}_1$, let $e_{i,t}^*$ be independently drawn from a discrete uniform distribution on the set $\left\{ \widehat{e}_{i,1}- \overline{\widehat{e}}_i ,\widehat{e}_{i,2}-\overline{\widehat{e}}_i, \ldots, \widehat{e}_{i,T_0}-\overline{\widehat{e}}_i \right\}$, where $ \displaystyle\overline{\widehat{e}}_i=\frac{1}{T_0}\sum_{s=1}^{T_0}\widehat{e }_{i,s}$. • For $(i,t)\in\mathcal{I}$, let $y_{i,t}^*=\widehat{c}_{i,t}+e^*_{i,t}$ . Use $\left\{ y_{i,t}^* \right\}$ to construct $Y_\mathrm{tall}^*$ and $Y_ \mathrm{wide}^*$ in the same fashion as Equation (ref). • Apply Algorithm (ref) to the bootstrapped sample $Y_ \mathrm{tall}^*$ and $Y_\mathrm{wide}^*$, and obtain $\left\{ \widehat{c} _{i,t}^*: (i,t)\in\mathcal{I}_1 \right\}$, $\left\{ \widehat{\mathbb{V}} _{i,t}^*: (i,t)\in\mathcal{I}_1 \right\}$, and $\left\{ \left(\widehat{\sigma }^*_i\right)^2: i=N_0+1,\ldots, N \right\}$. • For $(i,t)\in\mathcal{I}_1$, compute \begin{align} s_{i,t}^*=\frac{\widehat{c}_{i,t}^*-y_{i,t}^*} {\sqrt{\widehat{\mathbb{V}} _{i,t}^*+\left(\widehat{\sigma}^*_i\right)^2}}. \end{align} \end{enumerate} • In the previous step, we generate $B$ statistics denoted by $ s_{i,t}^*(1), s_{i,t}^*(2),\ldots, s_{i,t}^*(B)$ for every $(i,t)\in\mathcal{ I}_1$. Let $q_{1-\alpha,i,t}$ be the $(1-\alpha)$ empirical quantile of $ \left\{ s_{i,t}^*(1), s_{i,t}^*(2), \ldots, s_{i,t}^*(B)\right\} $, and let $ p_{1-\alpha,i,t}$ be the $(1-\alpha)$ empirical quantile of $\left\{ \left\vert s_{i,t}^*(1)\right\vert, \left\vert s_{i,t}^*(2)\right\vert, \ldots, \left\vert s_{i,t}^*(B) \right\vert \right\} $. • For $(i,t)\in\mathcal{I}_1$, the equal tailed $(1-\alpha)$ confidence interval of $\Delta_{i,t}$ is \begin{align} \mathrm{EQ}_{1-\alpha,i,t}=\left[ \widehat{\Delta}_{i,t}+ q_{\alpha/2, i,t} \sqrt{\widehat{\mathbb{V}}_{i,t}+\widehat{\sigma}^2_i}, \; \widehat{\Delta} _{i,t}+ q_{1-(\alpha/2),i,t} \sqrt{\widehat{\mathbb{V}}_{i,t}+\widehat{\sigma }^2_i} \right], \end{align} and the symmetric $(1-\alpha)$ confidence interval of $\Delta_{i,t}$ is \begin{align} \mathrm{SY}_{1-\alpha,i,t}=\left[ \widehat{\Delta}_{i,t}- p_{1-\alpha,i,t} \sqrt{\widehat{\mathbb{V}}_{i,t}+\widehat{\sigma}^2_i}, \; \widehat{\Delta} _{i,t}+ p_{1-\alpha, i,t} \sqrt{\widehat{\mathbb{V}}_{i,t}+\widehat{\sigma} ^2_i} \right]. \end{align} \end{enumerate}
remarkAs is shown in the proof of Theorem (ref), the distribution of $\left( \widehat{\Delta} _{i,t}-\Delta_{i,t} \right)$ is dominated by $e_{i,t}$ and the effect of $ \left( \widehat{c}_{i,t}-c_{i,t} \right)$ is asymptotically negligible. However, the bootstrap procedure above takes $\left( \widehat{c} _{i,t}-c_{i,t} \right)$ into consideration in order to improve the finite sample performances.
remarkWhen the error terms $\left\{ e_{i,t}\right\} $ are suspected of having serial correlation, a block wild bootstrap can be used to address this issue (e.g., gonccalves2017bootstrap). That is, a for block width $b\in \mathbb{Z}_{+}$, we let $u_{i,(j-1)b+s}=\overline{u}_{i,j}$ for every $j\in \mathbb{Z}_{+}$ and $s\in \left\{ 1,2,\cdots ,b\right\} $ such that $ (j-1)b+s\leq T_{0}$ in Step 2(1).

Under Assumptions (ref)--(ref) in this paper, the asymptotic properties of the bootstrapped confidence intervals (ref) and (ref) are summarized in the following theorem.

thryIf Assumptions (ref)--(ref) hold, then \begin{align} \lim_{N_0,T_0\to\infty} \mathbb{P} \left( \Delta_{i,t}\in \mathrm{EQ} _{1-\alpha,i,t} \right) =1-\alpha\qquad and\qquad \lim_{N_0,T_0\to\infty} \mathbb{P} \left( \Delta_{i,t}\in \mathrm{SY} _{1-\alpha,i,t} \right) =1-\alpha \end{align} for every $(i,t)\in\mathcal{I}_1$.

From Theorem (ref), we can conclude that the bootstrap confidence intervals will provide a correct coverage for the estimated treatment effects in the post-treatment periods, and thus one could conduct statistical inference about the treatment effects based on the bootstrap confidence intervals. It also worths noting that even if our bootstrap procedure follows the idea of gonccalves2017bootstrap, the proof of bootstrap validity becomes quite complicated due to the differences in purposes and model specifications. Since gonccalves2017bootstrap intend to make inference on factor-augmented regression models bai2006confidence where the factors serves as intermediate variables, they only require valid bootstrap approximations of factors. In this paper, we intend to make inference on the factor model (with missing values) itself, and the estimators involve products of estimated factors, their associated loadings, and a rotation matrix. This implies that we need valid bootstrap approximations of factors, loadings, and the rotation matrix, which leads to laborious works.

Estimation and Inference in a Model with Covariates

In the above section, we discussed the estimation of the treatment effects and associated bootstrap confidence intervals for a pure factor model. Besides the unobserved factors, in practice, the outcome variable could also be affected by some exogenous regressors. To accommodate such a scenario, we extend the factor-based approach to a model with covariates, i.e., Model (ref) in Section (ref).

Estimation of Treatment Effects

For a factor model with covariates, i.e., a panel data model with interactive fixed effect, our estimation strategies involve an interactive fixed effect estimation (IFEE), which is described and studied in bai2009panel. If $\left\{ y_{i,t}: (i,t)\in\mathcal{I} \right\}$ were fully observed, we could directly use IFEE to estimate Model (ref). The general framework of IFEE is summarised in the algorithm to follow.

algoInteractive fixed effect estimation. \begin{enumerate}[labelindent=\parindent, leftmargin=*, nosep] • Input arguments. $\mathcal{Y}=\left( \mathcal{Y}_1, \ldots, \mathcal{Y} _\mathcal{N} \right)$, a $\mathcal{T}\times \mathcal{N}$ matrix. $\mathcal{X} =\left( \mathcal{X}_1, \ldots, \mathcal{X}_\mathcal{N} \right)$, a $\mathcal{ T}\times (p\mathcal{N})$ matrix, where $\mathcal{X}_i$ is a $\mathcal{T} \times p$ matrix for every $i$. • Compute the starting value \begin{align} \beta^{(0)}=\left( \sum_{i=1}^\mathcal{N} \mathcal{X}_i^\mathsf{T} \mathcal{X }_i \right)^{-1} \sum_{i=1}^\mathcal{N} \mathcal{X}_i^\mathsf{T} \mathcal{Y} _i. \end{align} • Perform the following iteration until $\left\Vert \beta^{(M)}- \beta^{(M-1)} \right\Vert_2< \varepsilon$ for some $M\in\mathbb{Z}_+$, where $\varepsilon$ is a sufficiently small positive number. \begin{enumerate}[nosep] • Let \begin{align} \mathcal{R}^{(k)}= \dfrac{\mathcal{Y}-\mathcal{X}\left( I_\mathcal{N} \otimes \beta^{(k-1)} \right)}{\sqrt{\mathcal{N}\mathcal{T}}}. \end{align} • Perform a singular value decomposition of the matrix $\mathcal{R} ^{(k)} $, and let $D^{(k)}$ be an $r\times r$ diagonal matrix with the $r$ largest singular values of $\mathcal{R}^{(k)}$ on the diagonal in descending order. Then let $U^{(k)}$ and $V^{(k)}$ be $\mathcal{T}\times r$ and $ \mathcal{N}\times r$ matrices containing the left and right singular vectors of $\mathcal{R}^{(k)}$ corresponding to $D^{(k)}$. • Let $F^{(k)}=\sqrt{\mathcal{T}}U^{(k)}$ and $H^{(k)}= I_\mathcal{T}- \dfrac{F^{(k)} \left( F^{(k)} \right)^\mathsf{T}}{\mathcal{T}} $. • Let \begin{align} \beta^{(k)}=\left[ \sum_{i=1}^{\mathcal{N}} \mathcal{X}_i^\mathsf{T} H^{(k)} \mathcal{X}_i \right]^{-1} \left[ \sum_{i=1}^{\mathcal{N}} \mathcal{X}_i^ \mathsf{T} H^{(k)} \mathcal{Y}_i \right]. \end{align} \end{enumerate} • Output arguments. $\beta^{(M)}$, $F^{(M)}$, and $\Lambda^{(M)}= \sqrt{ \mathcal{N}}V^{(M)}D^{(M)}$. \end{enumerate}

Now we turn to the estimation of treatment effect when $\left\{ y_{i,t}:(i,t)\in\mathcal{I} \right\}$ are not observable. Let $Y$, $Y_ \mathrm{tall}$, $Y_\mathrm{wide}$, $X_i$, $F$, ${F}_\mathrm{tall}$, ${F}_ \mathrm{wide}$, $\Lambda$, ${\Lambda}_\mathrm{tall}$, and ${\Lambda}_\mathrm{ wide}$ follow their definitions in Sections (ref) and (ref). Let $ X=\left( X_1, \ldots, X_N \right)$, $X_\mathrm{tall}= \left( X_1, \ldots, X_{N_0} \right)$, and $X_\mathrm{wide}$ to be a sub-matrix formed by the first $T_0$ rows of $X$.

Then we can use the algorithm below to estimate the treatment effects $ \Delta_{i,t}$ for all $(i,t)\in\mathcal{I}_1$.

algoTreatment effect estimation via factor-based matrix completion (with covariates). \begin{enumerate}[labelindent=\parindent, leftmargin=*, nosep] • Apply Algorithm (ref) to $\left( Y_\mathrm{tall}, X_\mathrm{ tall} \right)$, and obtain $\widehat{\beta}_\mathrm{tall}$, $\widehat{F}_\mathrm{tall}= \left( \widehat{f}_{\mathrm{tall},1}, \ldots, \widehat{f}_{\mathrm{tall},T} \right)^\mathsf{T}$ and $\widehat{\Lambda}_\mathrm{tall}= \left( \widehat{\lambda}_{\mathrm{tall},1}, \ldots, \widehat{\lambda}_{\mathrm{tall},N_0} \right)^\mathsf{T}$. • Apply Algorithm (ref) to $\left( Y_\mathrm{wide}, X_\mathrm{ wide} \right)$, and obtain $\widehat{F}_\mathrm{wide} =\left( \widehat{f}_{\mathrm{wide},1}, \ldots, \widehat{f}_{\mathrm{wide},T_0} \right)^\mathsf{T}$ and $\widehat{\Lambda}_\mathrm{wide} =\left( \widehat{\lambda}_{\mathrm{wide},1}, \ldots, \widehat{\lambda}_{\mathrm{wide},N} \right)^\mathsf{T}$. • Let $\widehat{\Lambda}_{\mathrm{wide},0}$ be the first $N_0$ rows of $ \widehat{\Lambda}_\mathrm{wide}$. Compute \begin{align} \widehat{H}_\mathrm{miss}=\widehat{\Lambda}_\mathrm{tall}^\mathsf{T} \widehat{\Lambda}_{\mathrm{wide},0} \left( \widehat{\Lambda}_{\mathrm{wide} ,0}^\mathsf{T} \widehat{\Lambda}_{\mathrm{wide},0} \right)^{-1}, \end{align} and then let $\widehat{C}=\widehat{F}_\mathrm{tall}\widehat{H}_\mathrm{miss} \widehat{\Lambda}_\mathrm{wide}^\mathsf{T}$. • Let $\widehat{c}_{i,t}$ denote the $(t,i)$-th entry of $\widehat{C}$, then compute the residuals $\widehat{e}_{i,t}=y_{i,t}-x_{i,t}^\mathsf{T} \widehat{\beta}_\mathrm{tall} -\widehat{c}_{i,t}$ for $(i,t)\in \mathcal{I} \backslash\mathcal{I}_1$. • For $(i,t)\in\mathcal{I}_1$, the variance of $\widehat{c}_{i,t}$ is estimated by \begin{align} \widehat{\mathbb{V}}_{i,t}&=\frac{1}{T_0} \widehat{f}_{\mathrm{tall},t}^ \mathsf{T} \left( \frac{\widehat{F}_\mathrm{tall}^\mathsf{T} \widehat{F}_ \mathrm{tall}}{T} \right)^{-1} \widehat{\Phi}_i \left( \frac{\widehat{F}_ \mathrm{tall}^\mathsf{T} \widehat{F}_\mathrm{tall}}{T} \right)^{-1} \widehat{ f}_{\mathrm{tall},t} \nonumber \\ &\phantom{=\:\:} +\frac{1}{N_0} \widehat{\lambda}_{\mathrm{wide},i}^\mathsf{T } \left( \frac{\widehat{\Lambda}_\mathrm{wide}^\mathsf{T} \widehat{\Lambda}_ \mathrm{wide}}{N} \right)^{-1} \widehat{\Gamma}_t \left( \frac{\widehat{ \Lambda}_\mathrm{wide}^\mathsf{T} \widehat{\Lambda}_\mathrm{wide}}{N} \right)^{-1} \widehat{\lambda}_{\mathrm{wide},i}, \end{align} where \begin{align} \widehat{\Gamma}_t &=\frac{1}{N_0}\sum_{j=1}^{N_0} \widehat{e}_{j,t}^2 \widehat{\lambda}_{\mathrm{wide},j} \widehat{\lambda}_{\mathrm{wide},j}^ \mathsf{T}, \qquad L_{k,i}=\frac{1}{T_0}\sum_{s=k+1}^{T_0} \widehat{f}_{ \mathrm{tall},s} \widehat{e}_{i,s} \widehat{e}_{i,s-k} \widehat{f}_{\mathrm{ tall},s-k}^\mathsf{T}, \nonumber \\ \widehat{\Phi}_i&= L_{0,i}+\sum_{k=1}^K \left( 1-\frac{k}{K+1} \right) \left( L_{k,i}+L_{k,i}^\mathsf{T} \right), \end{align} with $K\to\infty$ and $\dfrac{K}{T_0^{1/4}}\to 0$ as $T_0\to\infty$. • For every $N_0< i \le N$, compute $\displaystyle \widehat{\sigma}^2_i= \frac{1}{T_0} \sum_{s=1}^{T_0} \widehat{e}_{i,s}^2$. • For $(i,t)\in\mathcal{I}_1$, the estimated treatment effect is $ \widehat{\Delta}_{i,t}=\mathscr{Y}_{i,t}-x_{i,t}^\mathsf{T} \widehat{\beta}_ \mathrm{tall} -\widehat{c}_{i,t}$, and the standard error of $\widehat{\Delta }_{i,t}$ is $\sqrt{\widehat{\mathbb{V}}_{i,t}+\widehat{\sigma}^2_i}$. \end{enumerate}
remarkBecause \begin{equation} \widehat{\Delta }_{i,t}-\Delta _{i,t}=x_{i,t}^{\mathsf{T}}\left( \beta - \widehat{\beta }_{\mathrm{tall}}\right) +\left( c_{i,t}-\widehat{c} _{i,t}\right) +e_{i,t} \end{equation} for every $(i,t)\in \mathcal{I}_{1}$, the variance of $\widehat{\Delta } _{i,t}$ is comprised of $\mathrm{Var}\left( e_{i,t}\right) $, $\mathrm{Var} \left( \widehat{c}_{i,t}\right) $ and $\mathrm{Var}\left( \widehat{\beta }_{ \mathrm{tall}}\right) $. In spirit of Remark (ref), we should include an estimate of $\mathrm{Var}\left( \widehat{\beta } _{\mathrm{tall}}\right) $ in the standard error of $\widehat{\Delta }_{i,t}$ to improve the finite sample performances. However, by Theorem 3 of bai2009panel, $\mathrm{Var}\left( \widehat{\beta }_{\mathrm{tall} }\right) =O\left( \dfrac{1}{N_{0}T}\right) $, which is of higher order than $ \mathrm{Var}\left( \widehat{c}_{i,t}\right) =O\left( \dfrac{1}{N_{0}}+\dfrac{ 1}{T_{0}}\right) $ for $(i,t)\in \mathcal{I}_{1}$. Furthermore, a consistent estimation of $\mathrm{Var}\left( \widehat{\beta }_{\mathrm{tall}}\right) $ is quite complicated. As a result of the cost-benefit trade-off, we do not include an estimate of $\mathrm{Var}\left( \widehat{\beta }_{\mathrm{tall} }\right) $ in the standard error of $\widehat{\Delta }_{i,t}$.

Construction of Confidence Intervals

For the estimated $\widehat{\Delta}_{i,t}$ for $(i,t)\in\mathcal{I}_1$, we shall again use the bootstrap procedure discussed above to construct their confidence intervals.

algoConfidence intervals of the treatment effects via bootstrap (with covariates). \begin{enumerate}[labelindent=\parindent, leftmargin=*, nosep] • Apply Algorithm (ref) to the sample and obtain $ \left\{ \widehat{c}_{i,t}: (i,t)\in\mathcal{I} \right\}$, $\left\{ \widehat{e }_{i,t}: (i,t)\in\mathcal{I} \right\}$, $\left\{ \widehat{\mathbb{V}}_{i,t}: (i,t)\in\mathcal{I}_1 \right\}$, and $\left\{ \widehat{\sigma}^2_i: i=N_0+1, \ldots, N \right\}$. • For $b=1,2,\ldots, B$: \begin{enumerate}[nosep] • For $(i,t)\in\mathcal{I}\backslash\mathcal{I}_1$, let $ e_{i,t}^*=u_{i,t}\widehat{e}_{i,t}$, where $\left\{ u_{i,t} \right\}$ are i.i.d.\ or block i.i.d.\ from a standard normal distribution $\mathbb{N} (0,1)$ and are independent of the raw sample. • For $(i,t)\in\mathcal{I}_1$, let $e_{i,t}^*$ be independently drawn from a discrete uniform distribution on the set $\left\{ \widehat{e}_{i,1}- \overline{\widehat{e}}_i ,\widehat{e}_{i,2}-\overline{\widehat{e}}_i, \ldots, \widehat{e}_{i,T_0}-\overline{\widehat{e}}_i \right\}$, where $ \displaystyle\overline{\widehat{e}}_i=\frac{1}{T_0}\sum_{s=1}^{T_0}\widehat{e }_{i,s}$. • For $(i,t)\in\mathcal{I}$, let $r_{i,t}^*=\widehat{c}_{i,t}+e^*_{i,t}$ . Use $\left\{ r_{i,t}^* \right\}$ to construct $R_\mathrm{tall}^*$ and $R_ \mathrm{wide}^*$ in the same fashion as Equation (ref). • Apply Algorithm (ref) to the bootstrapped sample $R_ \mathrm{tall}^*$ and $R_\mathrm{wide}^*$, and obtain $\left\{ \widehat{c} _{i,t}^*: (i,t)\in\mathcal{I}_1 \right\}$, $\left\{ \widehat{\mathbb{V}} _{i,t}^*: (i,t)\in\mathcal{I}_1 \right\}$, and $\left\{ \left(\widehat{\sigma }^*_i\right)^2: i=N_0+1,\ldots, N \right\}$. • For $(i,t)\in\mathcal{I}_1$, compute \begin{align} s_{i,t}^*=\frac{\widehat{c}_{i,t}^*-r_{i,t}^*} {\sqrt{\widehat{\mathbb{V}} _{i,t}^*+\left(\widehat{\sigma}^*_i\right)^2}}. \end{align} \end{enumerate} • In the previous step, we generate $B$ statistics denoted by $ s_{i,t}^*(1), s_{i,t}^*(2),\ldots, s_{i,t}^*(B)$ for every $(i,t)\in\mathcal{ I}_1$. Let $q_{1-\alpha,i,t}$ be the $(1-\alpha)$ empirical quantile of $ \left\{ s_{i,t}^*(1), s_{i,t}^*(2), \ldots, s_{i,t}^*(B)\right\} $, and let $ p_{1-\alpha,i,t}$ be the $(1-\alpha)$ empirical quantile of $\left\{ \left\vert s_{i,t}^*(1)\right\vert, \left\vert s_{i,t}^*(2)\right\vert, \ldots, \left\vert s_{i,t}^*(B) \right\vert \right\} $. • For $(i,t)\in\mathcal{I}_1$, the equal tailed $(1-\alpha)$ confidence interval of $\Delta_{i,t}$ is \begin{align} \mathrm{EQ}_{1-\alpha,i,t}=\left[ \widehat{\Delta}_{i,t}+ q_{\alpha/2, i,t} \sqrt{\widehat{\mathbb{V}}_{i,t}+\widehat{\sigma}^2_i}, \; \widehat{\Delta} _{i,t}+ q_{1-(\alpha/2),i,t} \sqrt{\widehat{\mathbb{V}}_{i,t}+\widehat{\sigma }^2_i} \right], \end{align} and the symmetric $(1-\alpha)$ confidence interval of $\Delta_{i,t}$ is \begin{align} \mathrm{SY}_{1-\alpha,i,t}=\left[ \widehat{\Delta}_{i,t}- p_{1-\alpha,i,t} \sqrt{\widehat{\mathbb{V}}_{i,t}+\widehat{\sigma}^2_i}, \; \widehat{\Delta} _{i,t}+ p_{1-\alpha, i,t} \sqrt{\widehat{\mathbb{V}}_{i,t}+\widehat{\sigma} ^2_i} \right]. \end{align} \end{enumerate}
remarkBy Equation (ref) and in spirit of Remark (ref), a bootstrap procedure should take the distributions of all the 3 terms $x_{i,t}^\mathsf{T} \left( \beta-\widehat{ \beta}_\mathrm{tall} \right)$, $\left( c_{i,t}-\widehat{c}_{i,t} \right)$ and $e_{i,t}$ into consideration to ensure the finite sample performances. But in Algorithm (ref), resampled observations are generated by a pure factor model and only the interactive fixed effects $ \left\{ \widehat{c}_{i,t}^* \right\}$ are estimated in every bootstrapped sample. Thus the bootstrap procedure only approximates the distribution of $ \left( c_{i,t}-\widehat{c}_{i,t} \right)+e_{i,t}$, ignoring the effect of $ x_{i,t}^\mathsf{T} \left( \beta-\widehat{\beta}_\mathrm{tall} \right)$. The underlying rationale is also a cost-benefit trade-off mentioned in Remark (ref). On the one hand, by bai2009panel and bai2021matrix, we have $\left( c_{i,t}- \widehat{c}_{i,t} \right)=O_\mathbb{P}\left( \dfrac{1}{\sqrt{N_0\wedge T_0}} \right)$ and $x_{i,t}^\mathsf{T} \left( \beta-\widehat{\beta}_\mathrm{tall} \right)=O_\mathbb{P} \left( \dfrac{1}{\sqrt{N_0 T}} \right)$ as $N_0, T_0\to\infty$ for every $(i,t)\in\mathcal{I}_1$. On the other hand, the computation of interactive fixed effect estimation is far more intensive than that of estimating a pure factor model. These two facts motivate us to ignore the effect of $x_{i,t}^\mathsf{T} \left( \beta-\widehat{\beta}_ \mathrm{tall} \right)$ in the bootstrap procedure. Furthermore, evidence from Section (ref) shows the current bootstrap procedure has already yielded satisfactory finite sample performances.
remarkIf the error terms $\left\{ e_{i,t} \right\}$ are suspected of having serial correlation, we suggest to use block wild bootstrap in Step 2(1). See Remark (ref) for details.

As above, we can establish the validity of the proposed confidence intervals (ref) and (ref) in the sense that they have asymptotically correct coverage probabilities as $N_0,T_0\to \infty$.

thryIf Assumptions (ref)--(ref) hold, then \begin{align} \lim_{N_0,T_0\to\infty} \mathbb{P} \left( \Delta_{i,t}\in \mathrm{EQ} _{1-\alpha,i,t} \right) =1-\alpha\qquad and\qquad \lim_{N_0,T_0\to\infty} \mathbb{P} \left( \Delta_{i,t}\in \mathrm{SY} _{1-\alpha,i,t} \right) =1-\alpha \end{align} for every $(i,t)\in\mathcal{I}_1$.

Theorem (ref) suggests that we can apply the proposed bootstrap procedure to conduct statistical inference for estimated treated effects from a panel with interactive fixed effects model.

Simulation Studies

In this section, we conduct several Monte Carlo experiments to investigate the finite sample properties of the proposed confidence intervals. \footnote{MATLAB codes for simulation are available from the authors upon request.}

In the data generating processes below, we assume that the common factors $ \left\{ f_{t}\right\} $ are i.i.d.\ as $\mathbb{N}\left( 0,I_{3}\right) $, and the factor loadings $\left\{ \lambda _{i}\right\} $ are also i.i.d.\ as $ \mathbb{N}\left( 0,I_{3}\right) $. For the model with covariates, we assume that the covariates $\left\{ x_{i,t}\right\} $ are i.i.d.\ as $\mathbb{N} \left( 0,AA^{\mathsf{T}}\right) $, where each entry of the $2\times 2$ matrix $A$ is drawn from $\mathbb{N}(0,1)$. Let $\beta \sim \mathbb{N}\left( 0,I_{2}\right) $. We consider the following data generating processes. \footnote{We also consider AR(1) factors, exponentially distributed errors, other variance structures, and dependence between factors and covariates. These results are available from the authors upon request.}

itemize[labelindent=\parindent, leftmargin=*, nosep] • DGP1: Model without covariates, $y_{i,t}=f_t^\mathsf{T} \lambda_i+ e_{i,t}$. • DGP2: Model with covariates, $y_{i,t}=x_{i,t}^\mathsf{T} \beta +f_t^ \mathsf{T} \lambda_i+ e_{i,t}$.

The error term is defined as

align[align omitted — 68 chars of source]

where $v_{i,t}=\rho_i v_{i,t-1}+\varepsilon_{i,t}$ and $\left\{ \varepsilon_{i,t} \right\}$ are i.i.d.\ drawn from $\mathbb{N} (0,1)$. We specify two variance structures of $\left\{ e_{i,t} \right\}$.

itemize[labelindent=\parindent, leftmargin=*, nosep] • Case 1: $\rho_i=0$ and $\sigma_{i}^2=1$. • Case 2: $\left\{ \rho_i \right\}$ are i.i.d.\ drawn from $\mathrm{Unif} \left( [-0.8,-0.2]\cup[0.2, 0.8] \right)$, and $\left\{ \sigma_{i}^2 \right\} $ are i.i.d.\ drawn from $\mathrm{logNormal}\left( 0,1 \right)$.

Moreover, we consider the following marginal distributions of $\left\{ v_{i,t}\right\}$.

itemize[labelindent=\parindent, leftmargin=*, nosep] • Margin 1: $v_{i,t}\sim \left[ \chi^2(1)-1 \right]/\sqrt{2}$. • Margin 2: $v_{i,t}\sim \left( \mathrm{Unif}[-0.5,0.5] \right)/ \sqrt{12 }$.

For demonstration, we assume there is only one treated unit $i=N$ and $5$ post-treatment periods. The treatment effects are assumed to be constants equal to 1, i.e., $ \Delta _{N,t}=1$ for $t=T_{0}+1,\ldots ,T_{0}+5$. The number of control units $N_{0}\in \{30,50,100\}$ and the number of pre-treatment periods $ T_{0}\in \{20,40\}$. We construct the $90\%$ and $95\%$, equal-tailed and symmetric confidence intervals for $\Delta _{N,t}$, $t=T_{0}+1,\ldots ,T_{0}+5$. In Step 2(1) of Algorithms (ref) and (ref), we use ordinary wild bootstrap procedure for Case 1, and block wild bootstrap procedure with block width equal to 4 for Case 2. The number of factors is either treated as known or estimated using the method of bai2002determining. For computational simplicity, a warp-speed method giacomini2013warp is applied with 2000 replications for each scenario. We report the coverage rates (in percent) of confidence intervals in Tables (ref)--(ref). In these tables, EQ stands for equal-tailed confidence intervals and SY stands for symmetric confidence intervals.

Several interesting findings can be observed from the simulation results. On the first hand, the results in Table (ref)--(ref) clearly show that the empirical coverage ratio for treatment effects from a pure factor model is quite close to the nominal values (both 90% and 95% level) when the post-treatment period is short, regardless of whether the idiosyncratic errors are heteroscedastic or serially correlated, or whether the number of unobserved factors is known or estimated from the data. On the other hand, when exogenous covariates are included for treatment effects estimation, Tables (ref)--(ref) also show that our proposed bootstrapped confidence intervals are able to provide accurate coverage ratios for the estimated treatment effects in a panel model with exogenous regressors and with heteroscedastic or serially correlated errors. In general, the simulation results confirm the validity of our proposed bootstrap procedure in providing accurate and robust confidence intervals for estimated treatment effects using a panel with interactive fixed effects.

table[table omitted — 5,069 chars of source]
table[table omitted — 5,061 chars of source]
table[table omitted — 5,068 chars of source]
table[table omitted — 5,060 chars of source]
table[table omitted — 5,089 chars of source]
table[table omitted — 5,081 chars of source]
table[table omitted — 5,088 chars of source]
table[table omitted — 5,080 chars of source]

Empirical Applications

In this section, we re-evaluate the impacts of Hong Kong's Political and Economic Integration with Mainland China as well as the effects of California's Tobacco Control Program using our proposed bootstrap procedure. \footnote{MATLAB, Python and R codes for application are available from the authors upon request.}

Hong Kong's Political and Economic Integration with Mainland China Revisited

In this subsection, we revisit the impacts of political and economic integration of Hong Kong with Mainland China, which has been analysed in hsiao2012panel. Since hsiao2012panel specify a pure factor model without covariates, we apply the methods in Section (ref) of this paper to the dataset of hsiao2012panel. For the results about political integration, quarterly real GDP growth rates from 1993Q1 to 1997Q2 of 10 countries and districts are used to form the counter-factual path of Hong Kong from 1997Q3 up to 2003Q4. The 10 countries and districts are Mainland China, Indonesia, Japan, Korea, Malaysia, Philippines, Singapore, Taiwan, Thailand and US. For the analysis of economic integration, quarterly real GDP growth rates of 24 countries and districts from 1993Q1 to 2003Q4 are used to form counter-factual path of Hong Kong from 2004Q1 to 2008Q1.

The results for the impact of political integration and economic integration with Mainland China on Hong Kong's economic growth are provided in (ref) and (ref), respectively.

figure[figure omitted — 223 chars of source]
figure[figure omitted — 222 chars of source]

As is shown in Figure (ref), the estimated treatment effects of Hong Kong political integration are of similar magnitudes and patterns as those in hsiao2012panel. In Figure (ref), the estimated treatment effects of Hong Kong economic integration are positive, as in hsiao2012panel, and confidence intervals cover 0 in most of the periods. This suggests that the impact of political and economic integration of Hong Kong with Mainland China is significant in the first few years after integration, and vanishes afterward. This observation is in general consistent with the insignificant average treatment effects across post-treatment periods in hsiao2012panel.

California's Tobacco Control Program Revisited

Now we revisit the effectiveness of CTCP on per capita cigarettes consumption and personal healthcare expenditures using the methods discussed in Section 4 of this article. In November 1988, California passed the Proposition 99, which increased California's cigarettes tax by 25 cents per pack and earmarked the tax revenue for health and anti-smoking measures. Proposition 99 triggered a wave of local clean-air ordinances in California. abadie2010synthetic used the synthetic control method for the period 1970-2000 to show that the California Tobacco Control Program had a significant impact on per capita cigarette consumption for the period 1989-2000, and that its impact continued to be enhanced over time.

Note that in the model and dataset of abadie2010synthetic, covariates do not change over time, which violates Assumption (ref)(2) of this study. To accommodate time-variant covariates, we use the dataset of hsiao2019panel, who also revisit the impact of CTCP on per capita cigarettes consumption and personal healthcare expenditures but use a set of time-variant covariates: per capita GDP obtained from abadie2010synthetic; poverty rates obtained from the National Census Bureau; educational attainment, defined as the percentage of obtaining college degree of population 25 years and over, obtained from the National Census Bureau.

In our analysis, the number of factors is estimated using the method proposed by alessi2010improved, which shows better performance than bai2002determining in this and the next applications. Both ordinary wild bootstrap procedure and block wild bootstrap procedure with block width equal to 3 are considered, and equal tailed confidence intervals are reported. The results for the impact of CTCP on per capita cigarette consumption and personal healthcare expenditures are provided in (ref) and (ref), respectively.

figure[figure omitted — 190 chars of source]
figure[figure omitted — 193 chars of source]

Figure (ref) shows that the estimated treatment effects are of similar magnitudes as those in abadie2010synthetic; and that the confidence intervals indicate negative and significant impacts over time, consistent with their findings using permutation tests. Figure (ref) shows that the estimated treatment effects of CTCP on health expenditure are of similar magnitudes as reported in hsiao2019panel. Consistent with their findings, confidence intervals indicate that the effects are short-lived. In both figures, we can see that the confidence intervals based on ordinary and block wild bootstrap procedures produce similar results.

Conclusion

In this paper, we consider the construction of confidence intervals for treatment effects estimated in panel models with interactive fixed effects, which serves as an alternative inferential approach. We first use the factor-based matrix completion technique proposed by bai2021matrix for panel models to estimate the treatment effects, and then use bootstrap method to construct confidence intervals of the treatment effects for treated units at each post-treatment period. Our construction of confidence intervals requires neither specific distributional assumptions on the error terms nor large number of post-treatment periods. We also establish the validity of proposed bootstrap procedure that these confidence intervals have asymptotically correct coverage probabilities. Simulation studies show that these confidence intervals have satisfactory finite sample performances, and empirical applications using classical datasets yield treatment effect estimates of similar magnitude and reliable confidence intervals.

center[center omitted — 529 chars of source]