Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
112,054 characters · 13 sections · 46 citation commands
Individual Causal Inference Using Panel Data With Multiple Outcomes
The main focus of the policy evaluation literature has been the average treatment effect and more recently the heterogeneous treatment effects or conditional average treatment effects, which are the average treatment effects for heterogeneous subgroups defined by the observed covariates (for reviews of these methods, see \citealp*{athey2017state,abadie2018econometric}). Ubiquitous in these studies is the unconfoundedness assumption, or the strong ignorability assumption, which requires all the covariates correlated with both the potential outcomes and the treatment assignment to be observed rosenbaum1983central.\footnote{This is also known as selection on observables or the conditional independence assumption.} Under this assumption, the potential outcomes and the treatment status are independent conditional on the observed covariates, and the difference between the mean outcomes of the treated and the untreated groups with the same values of the observed covariates is an unbiased estimator of the average treatment effect for the units in the groups. The unconfoundedness assumption is satisfied in randomised controlled experiments, but may not be plausible otherwise even with a rich set of covariates, since the access to certain essential individual characteristics remains limited for the researchers due to privacy or ethical concerns, despite the explosive growth of data availability in the big data era.
One popular method to circumvent the unconfoundedness assumption is difference-in-differences (DID), which assumes that the effect of the unobserved confounder on the untreated potential outcome is constant over time, so that the average outcomes of the treated and untreated units would follow parallel trends in the absence of the treatment.\footnote{Alternative methods that do not rely on the unconfoundedness assumption include the instrumental variables method and the regression discontinuity design, which estimate the average treatment effect for specific subpopulations (the compliers or those with values of the running variable near the cutoff).} This is also a strong assumption, and in many cases is not supported by data. The interactive fixed effects model relaxes the “parallel trends” assumption and allows the unobserved confounders to have time-varying effects on the outcomes, by modeling them using an interactive fixed effects term, which incorporates the additive unit and time fixed effects model or difference-in-differences as a special case bai2009panel.
Several methods have been developed based on the interactive fixed effects model to estimate the treatment effect on a single or several treated units, where the units are observed over an extended period of time before the treatment abadie2010synthetic,hsiao2012panel,xu2017generalized. These methods exploit the cross-sectional correlations attributed to the unobserved common factors to predict the counterfactual outcomes for the treated units, and are mainly used in macroeconomic settings with a large number of pretreatment periods, which is crucial for the results to be credible. For example, abadie2015comparative point out that “the applicability of the method requires a sizable number of preintervention periods” and that “we do not recommend using this method when the pretreatment fit is poor or the number of pretreatment periods is small”, while xu2017generalized states that users should be cautious when there are fewer than 10 pretreatment periods. As a consequence, despite the potential to estimate individual treatment effects without imposing the unconfoundedness assumption, these methods have not seen much use in empirical microeconomics, since the individuals are rarely tracked for more than a few periods that justify the use of these methods.
The main contribution of this paper is that we propose a method for estimating the individual treatment effects in applied microeconomic settings, characterised by multiple related outcomes being observed for a large number of individuals over a small number of time periods. The method is based on the interactive fixed effects model, which assumes that an outcome of interest can be well approximated by a linear combination of a small number of observed and unobserved individual characteristics. Analogous to hsiao2012panel who predict the posttreatment outcomes using pretreatment outcomes in lieu of the unobserved time factors, we use a subsample of the pretreatment outcomes to replace the unobserved individual characteristics in the models, and use the remaining pretreatment outcomes as instrumental variables. Although our method does not require a large number of pretreatment periods, the number of pretreatment outcomes needs to be at least as large as the number of unobserved individual characteristics, which may still be difficult to satisfy in microeconomic datasets if we use only a single outcome, especially if the treatment assignment took place in the early stages of the survey or if the study subjects are children or youths. Utilising multiple related outcomes allows our method to be applicable in cases where there is only a single period before the treatment. Under the assumption that these outcomes depend on roughly the same set of observed covariates and unobserved individual characteristics with time-varying and outcome-specific coefficients shared by all individuals, our method exploits the correlations across related outcomes and over time, which are induced by the unobserved individual characteristics, to predict the counterfactual outcomes and estimate the treatment effects for each individual in the posttreatment periods.
Our method has several advantages. First, with the assumption on the model specification, it relaxes the arguably much stronger unconfoundedness assumption, and allows the treatment assignment to be correlated with the unobserved individual characteristics. Second, it enables the estimation of treatment effects on the individual level, which may be helpful for designing more individualised policies to maximize social welfare, as well as for other fields such as precision medicine and individualised marketing. It also has the potential to be combined with more flexible machine learning methods to work with big datasets and more general nonlinear function forms. Third, it is intuitive. In real life, we may never know a person through and through, and a viable approach to predicting the outcome of a person is using his or her related outcomes in the past, assuming that the outcomes are affected by the underlying individual characteristics and that these characteristics are stable over time, at least within the study period. For example, past academic performance is an important consideration when recruiting a student into college, as it is believed that a student that excelled in the past is likely to continue to have outstanding performance. To the extent that we may never observe all the confounders, this is perhaps the only way to predict potential outcomes and estimate treatment effects on the individual level in social sciences without going deeper to the levels of neuroscience or biology. Fourth, our method has wide applicability, as it is common to have multiple related outcomes collected in microeconomics data. For example, we may observe several health related outcomes such as health facility usage, health related cost, general health, etc.
The rest of the study is organised as follows. Section (ref) presents the theoretical framework. Section (ref) examines the small sample performance of our method using Monte Carlo simulation, and compares it with related methods. Section (ref) provides an empirical example of estimating the effect of health insurance coverage on individual usage of hospital emergency departments using the Oregon Health Insurance Experiment data. Section (ref) concludes and discusses potential directions for future research. The proofs are collected in the appendix.
Suppose that we observe $K$ outcomes in domain $\mathcal{K}=\{1,2,\dots,K\}$ for $N$ individuals or units over $T\ge 2$ time periods, where a domain refers to a collection of related outcomes that depend on the same set of observed covariates and unobserved characteristics. For example, health-related outcomes may be affected by observed covariates such as age, education, occupation and income, as well as unobserved individual characteristics such as genetic inheritance, health habits and risk preferences. Assume that the $N_1$ individuals in the treated group $\mathcal{T}$ receive the treatment at period $T_0+1\le T$ and remain treated afterwards, while the $N_0=N-N_1$ individuals in the control group $\mathcal{C}$ remain untreated throughout the $T$ periods. Denoting the binary treatment status for individual $i$ at time $t$ as $D_{it}$, we have $D_{it}=1$ for $i\in\mathcal{T}$ and $t>T_0$, and $D_{it}=0$ otherwise.
Following the “Rubin Causal Model” rubin1974estimating, the treatment effect on outcome $k\in\mathcal{K}$ for individual $i$ at time $t$ is given by the difference between the treated and untreated potential outcomes
where $Y_{it,k}^1$ is the treated potential outcome, the outcome that we would observe for individual $i$ at time $t$ if $D_{it}=1$, and $Y_{it,k}^0$ is the untreated potential outcome, the outcome that we would observe if $D_{it}=0$. Instead of assuming the unconfoundedness condition, we characterise the two potential outcomes for individual $i$ at time $t$ and $k\in\mathcal{K}$ using the interactive fixed effects models:
where $\boldsymbol{X}_{it}$ is the $r\times 1$ vector of observed covariates unaffected by the treatment, $\boldsymbol{\mu}_i$ is the $f\times 1$ vector of unobserved individual characteristics, $\boldsymbol{\beta}_{t,k}^1$ and $\boldsymbol{\lambda}_{t,k}^1$ are the $r\times 1$ and $f\times 1$ vectors of coefficients of $\boldsymbol{X}_{it}$ and $\boldsymbol{\mu}_i$ respectively for the treated potential outcome, $\boldsymbol{\beta}_{t,k}^0$ and $\boldsymbol{\lambda}_{t,k}^0$ are the coefficients for the untreated potential outcome, and $\varepsilon_{it,k}^1$ and $\varepsilon_{it,k}^0$ are the idiosyncratic shocks.
The regularity conditions on the observed covariates and the unobserved individual characteristics are stated in Assumption (ref), and the assumptions on the idiosyncratic shocks are given in Assumption (ref).
Given the models for $Y_{it,k}^1$ and $Y_{it,k}^0$ in (ref) and (ref), the individual treatment effect is identified by the observed covariates and the unobserved individual characteristics, i.e., two persons with the same values for these underlying predictors have identical individual treatment effect. Denote the set of observed covariates and unobserved individual characteristics as $\boldsymbol{H}_{it}=\left[\boldsymbol{X}_{it}'\enspace\boldsymbol{\mu}_i'\right]'$, then the individual treatment effect for individual $i$ with $\boldsymbol{H}_{it}=\boldsymbol{h}_{it}$ is given by
which may appear similar to the conditional average treatment effect, but is different by conditioning not only on the observed covariates, but also on the unobserved individual characteristics.\footnote{As we assume the parametric models for the potential outcomes in (ref) and (ref) for all individuals, there is no need to impose additional assumptions on the propensity distribution for the individual treatment effect to be identified on its full support.}
Our goal is to estimate the individual treatment effects $\bar{\tau}_{it,k}$, $i=1,\dots,N$. Once we have estimated the individual treatment effects, the estimates of the average treatment effects for heterogeneous subgroups defined by some observed covariates, also known as the conditional average treatment effects, and the estimate for the average treatment effect for the sample or the population are also readily available using the average of the estimated individual treatment effects in the corresponding groups.
As $\boldsymbol{\mu}_i$ is not observed, a direct application of least squares estimation to estimate the models in (ref) and (ref) would suffer from omitted variables bias. Since we have multiple outcomes that depend on the same set of underlying predictors, and we observe the untreated potential outcomes for all individuals prior to the treatment, we can use these pretreatment outcomes to replace $\boldsymbol{\mu}_i$ in the models.\footnote{This is analogous to the first step of the approach in hsiao2012panel, who predict the posttreatment outcomes using pretreatment outcomes in lieu of the unobserved time factors in a small $N$, big $T$ environment.} Stacking the $K$ outcomes observed in $t\le T_0$, we have
where $\boldsymbol{Y}_{it}^0$ and $\boldsymbol{\varepsilon}_{it}^0$ are $K\times 1$, $\boldsymbol{\beta}_{t}^0$ is $K\times r$, and $\boldsymbol{\lambda}_{t}^0$ is $K\times f$. Let $\mathcal{P}\subseteq\left\{1,\cdots,T_0\right\}$ be a set of $P$ pretreatment periods. We can further stack the outcomes over these periods to get
where $\boldsymbol{\delta}_i^\mathcal{P}=\left[\cdots \enspace \left(\boldsymbol{\beta}_{s}^0\boldsymbol{X}_{is}\right)' \enspace \cdots\right]'$ with $s\in\mathcal{P}$ is $KP\times 1$, $\boldsymbol{\lambda}^\mathcal{P}$ is $KP\times f$, and $\boldsymbol{\varepsilon}_i^\mathcal{P}$ is $KP\times 1$.
To be able to recover $\boldsymbol{\mu}_i$ from the covariates and outcomes observed in $\mathcal{P}$, we need the following full rank condition, which ensures that there is enough variation in the effects of the unobserved individual characteristics over time or across different outcomes.
Under Assumption (ref), we can pre-multiply both sides of equation (ref) by $({\boldsymbol{\lambda}^\mathcal{P}}'\boldsymbol{\lambda}^\mathcal{P})^{-1}{\boldsymbol{\lambda}^\mathcal{P}}'$ to obtain
Substituting (ref) into $Y_{it,k}^0=\boldsymbol{X}_{it}'\boldsymbol{\beta}_{t,k}^0+\boldsymbol{\mu}_i'\boldsymbol{\lambda}_{t,k}^0+\varepsilon_{it,k}^0$, $t>T_0$, and with a little abuse on the notation by omitting the superscript $\mathcal{P}$ on the new coefficients and error term, we have
where
Let $Z=r(P+1)+KP$. If we denote the $Z\times 1$ vector of observables $[\boldsymbol{X}_{it}' \cdots \boldsymbol{X}_{is}' \cdots {\boldsymbol{Y}_i^\mathcal{P}}']'$ as $\boldsymbol{Z}_{it}$, and the $Z\times 1$ vector of coefficients $[{\boldsymbol{\beta}_{t,k}^0}' \cdots {\boldsymbol{\alpha}_{st,k}^0}' \cdots {\boldsymbol{\gamma}_{t,k}^0}']'$ as $\boldsymbol{\theta}_{t,k}^0$, then equation (ref) can be abbreviated as
Similarly, substituting (ref) into $Y_{it,k}^1=\boldsymbol{X}_{it}'\boldsymbol{\beta}_{t,k}^1+\boldsymbol{\mu}_i'\boldsymbol{\lambda}_{t,k}^1+\varepsilon_{it,k}^1$, $t>T_0$, we have
where
Under Assumption (ref), we have $\mathbb{E}(e_{it,k}^1\mid \boldsymbol{H}_{it})=0$ and $\mathbb{E}(e_{it,k}^0\mid \boldsymbol{H}_{it})=0$. This suggests using $\widehat{\tau}_{it,k}=\boldsymbol{Z}_{it}'(\widehat{\boldsymbol{\theta}}_{t,k}^1-\widehat{\boldsymbol{\theta}}_{t,k}^0)$, where $\widehat{\boldsymbol{\theta}}_{t,k}^1$ and $\widehat{\boldsymbol{\theta}}_{t,k}^0$ are some estimators of ${\boldsymbol{\theta}}_{t,k}^1$ and ${\boldsymbol{\theta}}_{t,k}^0$, to estimate $\bar{\tau}_{it,k}$. Note, however, that the error terms $e_{it,k}^1$ and $e_{it,k}^0$ are correlated with the regressors, since $\boldsymbol{Z}_{it}$ contains $\boldsymbol{Y}_i^\mathcal{P}$ which is correlated with $\boldsymbol{\varepsilon}_i^\mathcal{P}$. This renders the OLS estimators biased and inconsistent, which can be seen as a classical measurement errors in variables problem.\footnote{This is noted in ferman2019synthetic as well, who also suggested using pre-treatment outcomes as instrumental variables to deal with the problem. Our method is also related to the quasi-differencing approach in holtz1988estimating and the GMM approach in ahn2013panel. While these studies focus on estimating the coefficients on the observed covariates, our focus is on estimating the individual treatment effects.} We thus use the remaining outcomes as instrumental variables for $\boldsymbol{Y}_i^\mathcal{P}$ to consistently estimate ${\boldsymbol{\theta}}_{t,k}^1$ and ${\boldsymbol{\theta}}_{t,k}^0$ in each period, which would then allow us to obtain asymptotically unbiased estimates for the individual treatment effects.\footnote{We may construct the vectors of regressors and instruments differently under alternative assumptions on the dependence structure of the idiosyncratic shocks. For example, if the idiosyncratic shocks are correlated across time but are independent across outcomes, then we can split different outcomes into regressors and instruments. This would be similar to using the characteristics of similar products berry1995automobile or trading countries (see the Trade-weighted World Income instrument in acemoglu2008income) as instrumental variables. Incorporating more complex structures of the idiosyncratic shocks in the model is left for future research.} Since the outcomes depend on about the same set of observed and unobserved individual characteristics, the remaining outcomes are strongly correlated with the outcomes included in $\boldsymbol{Y}_i^\mathcal{P}$. Additionally, given that the idiosyncratic shocks are independent across time, the remaining outcomes are not correlated with $e_{it,k}^1$ or $e_{it,k}^0$. Thus, both the relevance and exogeneity conditions are satisfied, and the remaining outcomes can serve as valid instrumental variables.
Let $\boldsymbol{R}_{it}=[\boldsymbol{X}_{it}' \cdots \boldsymbol{X}_{is}' \cdots {\boldsymbol{Y}_i^{-\mathcal{P}}}']'$ be the $R\times1$ vector of instruments, where the $(KT-KP-1)\times1$ vector $\boldsymbol{Y}_i^{-\mathcal{P}}$ comprises the remaining pretreatment outcomes as well as the posttreatment outcomes other than $Y_{it,k}$.\footnote{In the special case of $T_1=1$ and $T_0=1$, we can include $K-1$ pretreatment outcomes as regressors, and use the posttreatment outcomes other than $Y_{it,k}$ as instruments so that $R\ge Z$.} Stacking $\boldsymbol{Z}_{it}$, $\boldsymbol{R}_{it}$ and $\boldsymbol{Y}_{it,k}^0$ respectively over the $N_0$ untreated individuals, we obtain the $N_0\times Z$ matrix of regressors $\boldsymbol{Z}_{t}^0$, the $N_0\times R$ matrix of instruments $\boldsymbol{R}_{t}^0$ and the $N_0\times 1$ matrix of outcomes $\boldsymbol{Y}_{t,k}^0$ for the untreated individuals. We can obtain $\boldsymbol{Z}_{t}^1$, $\boldsymbol{R}_{t}^1$ and $\boldsymbol{Y}_{t,k}^1$ similarly for the $N_1$ treated individuals. The GMM estimator for the individual treatment effect $\bar{\tau}_{it,k}$ can then be constructed as
where
with $\boldsymbol{W}^1$ and $\boldsymbol{W}^0$ being some $R\times R$ positive definite matrices.
The following result shows that the bias of the GMM estimator for the individual treatment effect in (ref) goes away as both the number of treated individuals and the number of untreated individuals become larger.
Once we have the estimates for the individual treatment effects, the average treatment effect $\tau_{t,k}=\mathbb{E}\left(\tau_{it,k}\right)$ can be conveniently estimated using the average of the estimated individual treatment effects $\widehat{\tau}_{t,k}=\frac{1}{N}\sum_{i=1}^N\widehat{\tau}_{it,k}$, which can be shown to be consistent.
To satisfy Assumption (ref), we need the number of pretreatment outcomes that we include as regressors in the model to be at least as large as $f$. Including more pretreatment outcomes may increase the variance of the estimator by increasing the variances of $\widehat{\boldsymbol{\theta}}_{t,k}^1$ and $\widehat{\boldsymbol{\theta}}_{t,k}^0$, but may also reduce the variance of the estimator when the sample is large and the variances of $\widehat{\boldsymbol{\theta}}_{t,k}^1$ and $\widehat{\boldsymbol{\theta}}_{t,k}^0$ are small, since
in the prediction error converges in probability to 0 as $KP$ grows.\footnote{Consistency of the individual treatment effect estimator may also be shown by allowing both $N$ and $KP$ to grow, with restrictions on the relative growth rate, e.g., $\frac{KP}{\min\left(\sqrt{N_1},\sqrt{N_0}\right)}\rightarrow0$. We do not pursue this path in this study, as the number of pretreatment outcomes in empirical microeconomics that we focus on is usually not large.}
To select the number of pretreatment outcomes to include in the model, we follow a model selection procedure similar to that in hsiao2012panel, where for each usable number of pretreatment outcomes, we construct many different models by including a random subset of the pretreatment outcomes as regressors and the remaining outcomes as instruments. We then estimate the models using GMM and obtain the leave-one-out prediction errors for all or a subsample of the individuals. The best set of pretreatment outcomes is chosen as the one that minimises the mean squared leave-one-out prediction error.\footnote{An alternative way to select the best set of pretreatment outcomes is to use information criteria such as GMM-BIC and GMM-AIC andrews1999consistent. To avoid the potential problem of post-selection inference, we may also randomly split the sample into two parts, where we select the best model on one part, and conduct inference on the other.}
In addition to the models using only a subset of the pretreatment outcomes, we also consider averaging different models that use the same number of pretreatment outcomes. Since the estimators constructed using only a subset of the pretreatment outcomes are asymptotically unbiased, as long as the number of pretreatment outcomes is larger than $f$, this property is passed on to the averaged estimator. The averaged estimator may also be more efficient as it uses more information in the sample and reduces uncertainty caused by a small number of sample splits.\footnote{We stick with simple averaging in this paper. More flexible averaging scheme, e.g., with larger weights on those with smaller out of sample prediction errors, would be an interesting direction for future research.} The leave-one-out prediction errors are also averaged over the models, and the best number of pretreatment outcomes to be used for the averaged estimator is similarly determined by minimising the mean squared leave-one-out prediction error.
An alternative approach to estimating the treatment effects is to follow hsiao2012panel and assume that
where $\boldsymbol{C}=\mathbb{E}\left(\boldsymbol{Z}_{it}\boldsymbol{Z}_{it}'\right)^{-1}\mathbb{E}\left(\boldsymbol{Z}_{it}{\boldsymbol{\varepsilon}_i^\mathcal{P}}'\right)$ is $Z\times KT_0$.\footnote{This assumption holds in special cases, e.g., when the unobserved predictors and the idiosyncratic shocks all follow the normal distribution li2017estimation. In more general cases, this assumption may be considered to hold approximately.} We can then separate the error term into a part correlated with the regressors and a part that has zero conditional mean, and rewrite the untreated potential outcome $Y_{it,k}^0$ as
where $u_{it,k}^0=\varepsilon_{it,k}^0-{\boldsymbol{\gamma}_{t,k}^0}'\boldsymbol{\varepsilon}_i^\mathcal{P}+{\boldsymbol{\gamma}_{t,k}^0}'\boldsymbol{C}'\boldsymbol{Z}_{it}$. Similarly, the treated potential outcome $Y_{it,k}^1$ can be rewritten as
where $\boldsymbol{\theta}_{t,k}^{*1}={\boldsymbol{\theta}_{t,k}^1}-\boldsymbol{C}\boldsymbol{\gamma}_{t,k}^1$, and $u_{it,k}^1=\varepsilon_{it,k}^1-{\boldsymbol{\gamma}_{t,k}^1}'\boldsymbol{\varepsilon}_i^\mathcal{P}+{\boldsymbol{\gamma}_{t,k}^1}'\boldsymbol{C}'\boldsymbol{Z}_{it}$.
Since $\mathbb{E}(u_{it,k}^1\mid\boldsymbol{Z}_{it})=\mathbb{E}[e_{it,k}^1-\mathbb{E}(e_{it,k}^1\mid\boldsymbol{Z}_{it})\mid \boldsymbol{Z}_{it}]=0$ and $\mathbb{E}(u_{it,k}^0\mid\boldsymbol{Z}_{it})=0$, it is straightforward to show that the least squares estimators $\widehat{\boldsymbol{\theta}}_{t,k}^{*1}=({\boldsymbol{Z}_{t}^1}'\boldsymbol{Z}_{t}^1)^{-1}{\boldsymbol{Z}_{t}^1}'\boldsymbol{Y}_{t,k}^1$ and $\widehat{\boldsymbol{\theta}}_{t,k}^{*0}=({\boldsymbol{Z}_{t}^0}'\boldsymbol{Z}_{t}^0)^{-1}{\boldsymbol{Z}_{t}^0}'\boldsymbol{Y}_{t,k}^0$ are the unbiased estimators of $\boldsymbol{\theta}_{t,k}^{*1}$ and $\boldsymbol{\theta}_{t,k}^{*0}$ respectively.\footnote{The linear conditional mean assumption also implies that the unconfoundedness assumption is satisfied, as $\mathbb{E}(Y_{it,k}^0\mid\boldsymbol{Z}_{it},D_{it}=1)=\mathbb{E}(Y_{it,k}^0\mid\boldsymbol{Z}_{it},D_{it}=0)$ and $\mathbb{E}(Y_{it,k}^1\mid\boldsymbol{Z}_{it},D_{it}=1)=\mathbb{E}(Y_{it,k}^1\mid\boldsymbol{Z}_{it},D_{it}=0)$.}
We can then construct an estimator as
which is an unbiased estimator for the average treatment effect for individuals with the same values of $\boldsymbol{Z}_{it}$, or the conditional average treatment effect. It follows that the average of the conditional average treatment effects estimators $\tilde{\tau}_{t,k}=\frac{1}{N}\sum_{i=1}^N\tilde{\tau}_{it,k}$ is an unbiased estimator for the average treatment effect $\tau_{t,k}$. In addition, it can also be shown that $\tilde{\tau}_{t,k}$ is a consistent estimator without imposing the linear conditional mean assumption li2017estimation.
Instead of replacing the unobserved confounders with the observed pretreatment outcomes, bai2009panel models the unobserved fixed effects directly by iterating between estimating the coefficients on the observed covariates and estimating the unobserved factors and factor loadings using the principal component analysis, given some initial values. This approach allows more general structures in the error terms, but requires both $N$ and $T$ to be large, and is also more restrictive on the model specification: the observed covariates need to be time-varying, while the coefficients are assumed constant over time. xu2017generalized adapts this method to the potential outcomes framework to estimate the average treatment effects on the treated, assuming that the untreated potential outcomes for both the treated and untreated units follow the interactive fixed effects model, and proposes a cross-validation procedure to choose the number of unobserved factors and a parametric bootstrap procedure for inference.
This approach has the desired feature of being less computationally expensive compared with repeated pretreatment set splitting and averaging, and is potentially more efficient compared with using only the best set of pretreatment outcomes and discarding the remaining information when all outcomes are related. However, its potential to be adapted to our settings is limited by the restrictions discussed above. In particular, we may assume the coefficients to be constant over time, but it would be unrealistic to assume that they are the same across different outcomes, if we were to use multiple related outcomes.
Another closely related study is athey2021matrix, which generalises the results from the matrix completion literature in computer science to impute the missing elements of the untreated potential outcome matrix for the treated units in the posttreatment periods, where the matrix is assumed to have a low rank structure, similar to that of the interactive fixed effects model. The bias of the estimator is shown to have an upper bound that goes to 0 as both $N$ and $T$ grow. This method allows staggered adoption of the treatment, i.e., the treated units receive the treatment at different time periods.
abadie2010synthetic estimate the treatment effect on a treated unit by predicting its untreated potential outcome using a synthetic control constructed as a weighted average of the control units. The synthetic control method applies to cases where the pretreatment characteristics of the treated unit can be closely approximated by the synthetic control constructed using a small number of control units over an extended period of time before the treatment, which may not generally hold. In terms of implementation, the objective function for the synthetic control method is similar to that of the linear regression approach in hsiao2012panel. However, the weights on the control units in the synthetic control method are restricted to be nonnegative to avoid extrapolation. This reduces the risk of overfitting, but may also limit its applicability by making it difficult to find a set of weights that satisfy the restrictions.
\sloppy To measure the conditional variance of the individual treatment effect estimator, $\text{Var}\left(\widehat{\tau}_{it,k}\mid \boldsymbol{H},\boldsymbol{D}\right)$, where $\boldsymbol{H}$ is the matrix of observed covariates and unobserved individual characteristics and $\boldsymbol{D}$ is the matrix of the treatment status for all individuals and all time periods in the sample, we follow xu2017generalized and employ a parametric bootstrap procedure.
First, we apply our method to all outcomes in all periods to obtain $\widehat{Y}_{it,k}^1$ and $\widehat{e}_{it,k}^1$ for the treated individuals in the posttreatment periods, and $\widehat{Y}_{it,k}^0$ and $\widehat{e}_{it,k}^0$ for the untreated individuals in the posttreatment periods and for all individuals in the pretreatment periods. Note that the residuals $\widehat{e}_{it,k}^1$ and $\widehat{e}_{it,k}^0$ are estimates for $\varepsilon_{it,k}^1-{\boldsymbol{\gamma}_{t,k}^1}'\boldsymbol{\varepsilon}_i^\mathcal{P}$ and $\varepsilon_{it,k}^0-{\boldsymbol{\gamma}_{t,k}^0}'\boldsymbol{\varepsilon}_i^\mathcal{P}$, respectively, rather than the idiosyncratic shocks in the original model, $\varepsilon_{it,k}^1$ and $\varepsilon_{it,k}^0$. Thus, the variance of the individual treatment effect estimator tends to be overestimated using the parametric bootstrap by resampling these residuals, especially when the number of pretreatment outcomes is small.\footnote{See the discussion on equation (ref).} Correcting for this bias would be a necessary step for future research.
These fitted values of the outcomes can be stacked into a $TK\times 1$ vector $\widehat{\boldsymbol{Y}}_i$ for each individual, where $\widehat{\boldsymbol{Y}}_i$ for $i\in\mathcal{T}$ contains $\widehat{Y}_{it,k}^1$ in the posttreatment periods and $\widehat{Y}_{it,k}^0$ in the pretreatment periods, and $\widehat{\boldsymbol{Y}}_i$ for $i\in\mathcal{C}$ contains $\widehat{Y}_{it,k}^0$ in all periods. The $TK\times 1$ vector of residuals $\widehat{\boldsymbol{e}}_i$ can be obtained similarly.
We then start bootstrapping for $B$ rounds:
The variance for the individual treatment effect estimator is computed using the bootstrap estimates as
and the $100(1-\alpha)\%$ confidence intervals for $\bar{\tau}_{it,k}$, $i=1,\dots,N$ can be constructed as
where the superscript denotes the index of the bootstrap estimates in ascending order. Alternatively, we can use a normal approximation and construct the confidence intervals as
where $\Phi\left(\cdot\right)$ is the cumulative distribution function for the standard normal distribution, and $\widehat{\sigma}_{it,k}=\sqrt{\text{Var}\left(\widehat{\tau}_{it,k}\mid \boldsymbol{H},\boldsymbol{D}\right)}$.
The variance for the average treatment effect estimator $\widehat{\tau}_{t,k}=\frac{1}{N}\sum_{i=1}^N\widehat{\tau}_{it,k}$ and the confidence interval for the average treatment effect $\tau_{t,k}$ can be obtained in similar manners using the bootstrap estimates $\widehat{\tau}_{t,k}^{(b)}=\frac{1}{N}\sum_{i=1}^N\widehat{\tau}_{it,k}^{(b)}$, $b=1,\dots,B$.
In this section, we conduct Monte Carlo simulations to assess the performance of our estimator in small samples, and compare it with related methods in relevant settings. The number of posttreatment period $T_1$ is fixed at 1, and the number of related outcomes $K$ is fixed at 5 in all settings.
The untreated potential outcomes are generated from
where $\boldsymbol{X}_{it}$ contains 2 observed covariates, and $\boldsymbol{\mu}_{i}$ contains 2 unobserved individual characteristics as well as the constant 1. The 2 observed covariates are i.i.d. $N(0,1)$ in period 1, and then follow an AR(1) process, $\boldsymbol{X}_{it}=0.9\boldsymbol{X}_{i,t-1}+\xi_{it}$, where $\xi_{it}$ are i.i.d. $N(0,\sqrt{1-0.9^2})$, so that the observed covariates are correlated across time and the variances stay 1. The 2 unobserved individual characteristics are also i.i.d. $N(0,1)$. The coefficients $\boldsymbol{\beta}_{t,k}^0$ and $\boldsymbol{\lambda}_{t,k}^0$ are i.i.d. $N(\omega_k,1)$ with $\omega_k\sim N(1,1)$, for $k\in\mathcal{K}$, so that the means of the coefficients differ across outcomes, and the idiosyncratic shocks $\varepsilon_{it,k}^0$ are i.i.d. $N(0,1)$.
The individual treatment effect in the posttreatment period $\bar{\tau}_{iT_0+1,k}$ is a deterministic function of $\boldsymbol{X}_{it}$ and $\boldsymbol{\mu}_i$ with the coefficients being i.i.d. $N(0.5,0.5)$, for $k\in\mathcal{K}$. And the observed outcomes $Y_{it,k}$, $k\in\mathcal{K}$ are equal to $Y_{it,k}^0-\varepsilon_{it,k}^0+\bar{\tau}_{iT_0+1,k}+\varepsilon_{it,k}^1$, where $\varepsilon_{it,k}^1$ are i.i.d. $N(0,1)$, for the treated individuals in the posttreatment period, and $Y_{it,k}^0$ otherwise. $\boldsymbol{X}_{it}$ and $\boldsymbol{\mu}_i$ as well as their coefficients for the untreated potential outcomes and the treatment effects are drawn 5 times, and for each set of $\{\boldsymbol{X}_{it},\boldsymbol{\mu}_i\}$ and their coefficients drawn, $\varepsilon_{it,k}^0$ and $\varepsilon_{it,k}^1$ are drawn 1000 times, which allows us to compute the bias and variance of the estimator conditional on the observed covariates and the unobserved individual characteristics.
\sloppy To measure the performances of the estimators, we compute the biases and standard deviations for the estimates of the individual treatment effects and the average treatment effect for outcome $K$ in the posttreatment period. Specifically, the bias of the individual treatment effect estimator $\widehat{\tau}_{iT_0+1,K}$ is measured by $\frac{1}{N}\sum_{i=1}^N\frac{1}{5}\sum_{d=1}^5\left\vert\mathbb{E}\left(\widehat{\tau}_{iT_0+1,K}^{(d,s)}\right)-\bar{\tau}_{iT_0+1,K}^{(d)}\right\vert$, where the superscript $d$ denotes the $d$th draw of $\{\boldsymbol{X}_{it},\boldsymbol{\mu}_i\}$ and $s$ denotes the $s$th draw of $\varepsilon_{it,k}^0$ and $\varepsilon_{it,k}^1$, and the standard deviation is constructed as $\frac{1}{N}\sum_{i=1}^N\frac{1}{5}\sum_{d=1}^5\sqrt{\mathbb{E}\left(\widehat{\tau}_{iT_0+1,K}^{(d,s)}-\mathbb{E}\widehat{\tau}_{iT_0+1,K}^{(d,s)}\right)^2}$. Similarly, the bias of the average treatment effect estimator $\widehat{\tau}_{T_0+1,K}$ is measured by $\frac{1}{5}\sum_{d=1}^5\left\vert\mathbb{E}\left(\widehat{\tau}_{T_0+1,K}^{(d,s)}\right)-\bar{\tau}_{T_0+1,K}^{(d)}\right\vert$, and the standard deviation is constructed as $\frac{1}{5}\sum_{d=1}^5\sqrt{\mathbb{E}\left(\widehat{\tau}_{T_0+1,K}^{(d,s)}-\mathbb{E}\widehat{\tau}_{T_0+1,K}^{(d,s)}\right)^2}$.\footnote{The performance of the estimators can also be measured using RMSE, which is computed as $\frac{1}{5}\sum_{d=1}^5\sqrt{\mathbb{E}\left(\widehat{\tau}_{iT_0+1,K}^{(d,s)}-\mathbb{E}{\bar{\tau}}_{iT_0+1,K}^{(d,s)}\right)^2}$ for $\widehat{\tau}_{iT_0+1,K}$ and $\frac{1}{5}\sum_{d=1}^5\sqrt{\mathbb{E}\left(\widehat{\tau}_{T_0+1,K}^{(d,s)}-\mathbb{E}{\bar{\tau}}_{T_0+1,K}^{(d,s)}\right)^2}$ for $\widehat{\tau}_{T_0+1,K}$. Since the biases of our estimators are small, these measures are quite similar to SD and are thus omitted from reporting.}
Table (ref) compares the GMM estimator constructed using only the best set of pretreatment outcomes with that constructed by averaging estimators from different models with the same number of pretreatment outcomes. We see that the best number of pretreatment outcomes, $P$, is slightly larger than the number of unobserved individual characteristics ($f=2$) for both estimators, and increases when the sample size is larger and when there are more pretreatment outcomes available, which is in line with our discussions in section (ref). The estimators constructed by model averaging also tends to select a slightly larger $P$ than the estimator using only the best set of pretreatment outcomes.
In almost all settings, the estimator using only the best set of pretreatment outcomes tends to have a smaller bias, whereas the estimator constructed from model averaging tends to have a smaller variance, except for estimating the individual treatment effects when the number of pretreatment outcomes is very small. The bias and SD also become smaller for both estimators when the sample size as well as the number of pretreatment outcomes grow.
In the following simulations, we fix $P$ at 2 when $T_0=1$, and 3 when $T_0=2$, and construct the GMM estimator using only the best set of pretreatment outcomes, with the best set of pretreatment outcomes selected at the first simulation and used for the remaining simulations for each setting.\footnote{This is mainly to save computing time and does not fundamentally change the conclusions.}
Table (ref) reports the bias and SD of the GMM estimator, as well as the coverage probability of the 95% confidence interval, for estimating the individual treatment effects and the average treatment effect. Panel A shows that the bias and SD for the estimators are small even with a small sample size and a small number of pretreatment outcomes. However, the 95% confidence intervals tend to have larger coverage probabilities, especially when the number of pretreatment outcomes is small. This distortion is alleviated as more pretreatment outcomes are available.
Since the validity of the GMM estimator relies on the assumption that $\varepsilon_{it,k}$ are uncorrelated across time or outcomes, we examine the performance of the estimator when this assumption is violated in Panel B, where the idiosyncratic shocks follow an AR(1) process over time with the autoregression coefficient being 0.1, and are correlated across outcomes by sharing a common component for different outcomes in the same period. This slightly increases the biases and SD's of the estimators, but the performance of the estimators are still quite good, especially in comparison with related methods as shown in the following tables.
Table (ref) compares our method with the OLS approach in hsiao2012panel. In panel A, both the unobserved individual characteristics and the idiosyncratic shocks are normally distributed so that the linear conditional mean assumption is satisfied. The results show that the GMM estimator outperforms the OLS estimator by having a smaller bias in estimating both the individual treatment effects and the average treatment effect, although the variance of the GMM estimator is also larger.
In panel B, the unobserved individual characteristics are drawn from the uniform distribution, and the linear conditional mean assumption is no longer satisfied li2017estimation. We see that the results are virtually unchanged for the GMM estimator, while the OLS estimator performs slightly worse by having larger biases and SD's, which is more pronounced in estimating the individual treatment effects. The results indicate that the linear conditional mean assumption is not a very strong one. Indeed, the distribution of the sum of several random variables would become more bell-shaped like the normal distribution under fairly general conditions, as a result of the central limit theorem. The simulation results are very similar when the unobserved individual characteristics are drawn from a mix of other distributions.
Table (ref) compares our method with the method of estimating the interactive fixed effects model directly, which was first developed in bai2009panel and then adapted into the potential outcomes framework by xu2017generalized to allow heterogeneous treatment effects. We fix the number of treated individuals at 5, and compare the performance of the two methods in estimating the individual treatment effect on the treated and the average treatment effect on the treated.
We consider two scenarios that are relevant in the context of empirical microeconomics. In panel A, the observed covariates are constant over time. This is plausible for covariates such as gender, race or education level, which are likely to be stable over time. Since the IFE method requires the observed covariates to be time-varying, the covariates that are constant over time are dropped from the estimation and become part of the unobserved individual characteristics, which makes the model equivalent to a pure factor model with 4 unobserved factors. As we have 5 related outcomes, this model should still be estimable by the IFE method. However, we see that IFE method perform poorly when there are only a small number of pretreatment outcomes to recover the unobserved individual characteristics. The bias and SD of the IFE estimator become smaller as more pretreatment outcomes are available, but are still quite large compared with our method.
To accommodate the restrictive model specification for the IFE method, we allow the covariates to be time-varying while keeping the coefficients constant over time in panel B, although the coefficients are allowed to vary across outcomes since it is unlikely that the coefficients for different outcomes would be the same in practice. We see that the IFE estimator has poor performance since the model is still misspecified in their method, whereas the results for our method are virtually unchanged.
Table (ref) compares our method with the synthetic control method abadie2010synthetic. In panel A, the unobserved individual characteristics for both the treated individuals and the untreated individuals are drawn from $N(0,1)$, while in panel B, the unobserved individual characteristics for the treated individuals are drawn from $N(1,1)$. Since the synthetic control method requires the treated units to be in the convex hull of the control units by restricting the weights assigned to the control units to be nonnegative, their method may perform poorly when the support of the unobserved individual characteristics are different for the treated and untreated individuals. While our method should be unaffected by the degree of overlapping in the distributions of the unobserved individual characteristics for the two treatment groups. The simulation results show that indeed the synthetic control estimator performs worse in panel B. Perhaps somewhat surprising is that its performance is also poor compared with our method in panel A. This is because the coefficients are outcome-specific, so that the levels of the outcomes are also likely to vary across outcomes, which makes it more difficult to obtain a good pretreatment fit under the nonnegativity restriction. In comparison, our method has good performance in both panels.
Overall, the simulation results show that our method has good performance in terms of the bias and SD in estimating the individual treatment effects and the average treatment effect under various settings, and has superior performance than related methods. The shortcoming of our method is that the confidence intervals tend to be too wide, especially when the number of pretreatment outcomes is small.
We illustrate our method by estimating the effect of health insurance coverage on the individual usage of hospital emergency departments.
Although the usage of emergency departments applies to only a small proportion of the population, it imposes great financial pressure on the health care system. In addition, it is not clear ex ante what the direction of the effect should be. E.g., taubman2014medicaid argues that health insurance coverage could either increase emergency-department use by reducing its cost for the patients, or decrease emergency-department use by encouraging primary care use or improving health.
The findings on emergency-department use have been mixed. Using survey data collected from the participants of the Oregon Health Insurance Experiment (OHIE) about a year after they were notified of the selection results, finkelstein2012oregon find no discernible impact of health insurance coverage on emergency-department use.\footnote{The Oregon Health Insurance Experiment (OHIE) was initiated in 2008, targeting at low-income adults in Oregon who had been without health insurance for at least 6 months. Among the 89,824 individuals who signed up, 35,169 individuals were randomly selected by the lottery and were eligible to apply for the Oregon Health Plan (OHP) Standard program, which provided relatively comprehensive medical benefits with no consumer cost sharing, and the monthly premiums was only between \$0 and \$20 depending on the income. As a randomised controlled experiment, the OHIE offers an opportunity for researchers to study the effect of health insurance coverage on various health outcomes without confounding factors.} While using the visit-level data for all emergency-department visits to twelve hospitals in the Portland area probabilistically matched to the OHIE study population on the basis of name, date of birth, and gender, taubman2014medicaid find that health insurance coverage significantly increases emergency-department use by 0.41 visits per person, from an average of 1.02 visits per person in the control group in the first 15 months of the experiment. They also examine whether the effect differs across heterogeneous groups, and find statistically significant increases in emergency-department use across most subgroups in terms of the number of pre-experiment emergency-department visits, hospital admission (inpatient or outpatient visits), timing (on-hours or off-hours visits), the type of visits (emergent and not preventable, emergent and preventable, primary care treatable, and non-emergent), as well as gender, age, and health condition.
In this application, we wish to estimate the effect of health insurance coverage on emergency-department use for each individual in the sample. This would potentially help us better understand whether and how health insurance coverage affects emergency-department use, compared with using only the average treatment effect for the whole sample or for some preassigned subgroups (conditional average treatment effects).
Our data combines both the hospital emergency-department visit-level data and the survey data. There are two time periods, one before the randomisation and one after.\footnote{The pre-randomisation period in the hospital visit-level data was from January 2007 to March 2008, and the post-randomisation period was from March 2008 to September 2009. The two surveys were collected shortly after the randomisation and about a year after randomisation, respectively, each covering a 6-month period before the survey.} To estimate the individual treatment effects, we include 3 observed covariates including gender, birth year, and household income as a percentage of the federal poverty line, and 10 related outcomes including different types of emergency-department visits and medical charges. We also consider a rich list of variables on which we make comparisons for individuals with different estimated treatment effects. There are 2154 individuals with complete information on these variables.\footnote{Note that our sample size is significantly smaller than the other studies using the OHIE data, due to the inclusion of the extensive list of variables. For example, the sample size in finkelstein2012oregon is 74,922, and the sample size in taubman2014medicaid is 24,646. So our sample may not be representative of the OHIE sample and the results in different studies may not be directly comparable.}
\resizebox{\columnwidth}{!}{
}
The first 3 columns in Table (ref) present the mean values of the covariates and outcomes in the pretreatment period for individuals selected by the lottery and for individuals not selected by the lottery, as well as the difference between the two groups. Since a considerable number of observations with incomplete information are dropped, being selected by the lottery is negatively correlated with different types of emergency-department visits in the pretreatment period in our sample, which suggests that the lottery assignment is not likely to be a valid instrument for health insurance coverage. Table (ref) also compares the mean pretreatment characteristics for individuals covered by health insurance and those not covered, which shows that individuals who were covered were poorer and used emergency-department more frequently in the pretreatment period than people who were not covered by health insurance.
Figure (ref) shows the distribution of the estimated individual treatment effects using our method.\footnote{As mentioned earlier, since only individuals with complete information on the variables are selected into our sample, the distribution may not be representative of the OHIE participants. The distribution of the estimated individual treatment effects may be more spread out than the distribution of the true effects due to noise, or less spread out since the estimates are based on the parametric models for the potential outcomes, which may be over-simplifying compared with the true models.} The mean of estimated individual treatment effects, or the estimated average treatment effect is 0.33, which is significant at 1% level. 114 individuals have treatment effects that are significant at 10% level, among which 23 are negative and 91 are positive.\footnote{Note that if we were to adjust for multiple testing, e.g., using the Benjamini–Hochberg procedure to control the false discovery rate (FDR) at 10% level, then we would be left with only one individual whose treatment effect is significant. Although the small number of individuals with significant treatment effects may also be attributed to the overestimation of the variance of the individual treatment effect estimator.} We then move on to compare the characteristics of the individuals based on their estimated treatment effects, which are presented in Table (ref)-(ref). Column (1) shows the mean characteristics of individuals whose treatment effects are not significant at 10% level, column (2) shows the mean characteristics of individuals whose treatment effects are significantly negative, column (3) shows the differences between column (2) and column (1), column (4) shows the mean characteristics of individuals whose treatment effects are significantly positive, and column (5) shows the differences between column (4) and column (1).
Compared with individuals who would not be significantly affected by the treatment, individuals who would significantly decrease or increase their emergency-department visits if covered by health insurance both had more emergency-department visits and more medical charges in the pretreatment period. However, these two groups were also distinct in some characteristics.
The individuals who would have fewer emergency-department visits if covered by health insurance were on average 7 years younger than individuals in the control group and 10 years younger than the positive group, more likely to be female with less education, and in particular, were much poorer than individuals in the other groups. They were more likely to be diagnosed with depression but not other conditions. Importantly, they were less likely to have any primary care visits, and more likely to use emergency department as the place for medical care. In terms of emergency-department use in the pretreatment period, they had fewer visits resulting in hospitalisation, more outpatient visits, more preventable and non-emergent visits, more visits to hospitals with a low fraction of uninsured patients, fewer visits for chronic conditions, and more visits for injury. Although their medical charges were not as high as those for individuals in the positive group, they owed more money for medical expenses.
In comparison, individuals who would have more emergency-department visits if covered by health insurance were more likely to be older, male, and with household income right above the federal poverty line, which means that they were not as poor as the individuals in the other groups. They were in worse health conditions, more likely to be diagnosed with diabetes and high blood pressure, and were taking more prescription medications. They also had more emergency-department visits of all types in the pretreatment period, including visits resulting in hospitalisation and visits for more severe conditions such as chronic conditions, chest pain and psychological conditions, and they incurred more medical charges.
Overall, these comparisons suggest that the individuals who would have fewer emergency-department visits if covered by health insurance were younger and not in very bad physical conditions. However, their access to primary care were limited due to being in much more disadvantaged positions financially, which made them resort to using the emergency department as the usual place for medical care. In contrast, the individuals who would have more emergency-department visits if covered by health insurance were more likely to be older and in poor health. So even with access to primary care, they still used emergency departments more often for severe conditions, although sometimes for primary care treatable and non-emergent conditions as well.
All in all, it seems that both mechanisms discussed by taubman2014medicaid are playing a role. For people who used emergency department for medical care because they did not have access to primary care service, health insurance coverage decreases emergency-department use because it increases access to primary care and may also lead to improved health. Whereas for people who had access to primary care and still used emergency department due to worse physical conditions, health insurance coverage increases emergency-department use because it reduces the out-of-pocket cost of the visits.
This application shows the potential value of estimating individual treatment effects for policy evaluation. Our findings would not have been possible by only estimating conditional average treatment effects, as we would not be able to distinguish individuals with positive or negative treatment effects at the first place.
\resizebox{\columnwidth}{!}{
}
\resizebox{\columnwidth}{!}{
}
\resizebox{\columnwidth}{!}{
}
In this paper, we propose a method for estimating the individual treatment effects using panel data, where multiple related outcomes are observed for a large number of individuals over a small number of pretreatment periods. The method is based on the interactive fixed effects model, and allows both the treatment assignment and the potential outcomes to be correlated with the unobserved individual characteristics. Monte Carlo simulations show that our method outperforms related methods. We also provide an example of estimating the effect of health insurance coverage on individual usage of hospital emergency departments using the Oregon Health Insurance Experiment data.
There are several directions for future research. First, our method requires the idiosyncratic shocks in the pretreatment outcomes to be uncorrelated either over time or across outcomes. It would be a valuable addition to allow (or detect and adjust for) more general dependence structure in the idiosyncratic shocks. Second, since the residuals of the rearranged models are not estimates of the idiosyncratic shocks, the variance of our estimator may be over-estimated, especially when the number of pretreatment outcomes is small. A necessary step for future research is to correct for this bias. Third, the repeated pretreatment set splitting and averaging approach in our method is computationally expensive. It would be an interesting direction for future research to find better ways to select related outcomes or use more flexible averaging scheme. Fourth, the linear model specification may be restrictive. There is potential to extend our method, perhaps in combination with more flexible machine learning methods, to work with more general nonlinear outcomes.