Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
70,823 characters · 21 sections · 90 citation commands
Quantile Treatment Effects in High Dimensional Panel Data
Over the past two decades, Synthetic Control Method (SCM) abadie2003economic,abadie2010synthetic has been extensively applied in empirical research in economics and social sciences. It estimates the average treatment effects on the treated (ATT) by constructing a counterfactual of a single or small group of treated units using the control group. Due to its interpretability, transparency, and plausible identification assumptions, SCM has gained popularity among practitioners, and is considered “arguably the most important innovation in the policy evaluation literature in the last 15 years” athey2017state. In particular, xu_2017 proposes the generalized synthetic control method (GSCM) adapting SCM to a high dimensional panel data setting (large $N$ and $T$) with numerous control units. This approach naturally incorporates a factor model to manage the abundance of control units, thereby attracting significant interest for estimating ATT in high dimensional panel data.
Beyond the mean effect, researchers might be concerned with the distributional changes caused by certain policies or interventions. For example, adrian2019vulnerable argue that traditional economic forecasting, which focuses primarily on the conditional mean of GDP growth, tends to obscure the importance of downside risks. They propose assessing the vulnerability of economic growth by examining the full conditional distribution of GDP, with a particular emphasis on its lower tails. Such distinct behaviors across different parts of the distribution of GDP growth raise a natural question: How will a macroeconomic policy affect different quantiles of the GDP growth? In our empirical application, where we study the impact of 2008 Chinese Economic Stimulus Program on GDP growth, we try to answer the following question: How will this policy alter GDP growth during economic downturns (lower quantiles) or booms (upper quantiles)?
In this paper, we consider the quantile treatment effects on the treated (QTT) in the SCM context with high dimensional panel data, where we have a single or a few treated units and a large number of control units. QTT delivers the shifts in outcome distributions caused by a policy at different quantiles, allowing one to fully characterize the impact of the intervention. Our estimators take a similar spirit to xu_2017's GSCM, which posits a factor model bai2003inferential,bai2009panel to capture the unobserved time-variant heterogeneities on the panel data. The GSCM first estimates the unknown factors using the control group, then estimates the factor loadings for the treated units, and finally imputes the counterfactual outcome for the treated units to calculate the ATT. To study the QTT, we consider a quantile factor model chen2021quantile, which allows the number and values of the factors as well as its loadings to vary across different quantiles. This setting is motivated by the increasing evidence of comovements of some economics and finance variables adrian2019vulnerable,de2019dynamic and admits much more flexibility in the model. Our QTT estimation proceeds in two stages. First, we estimate the unobserved quantile-dependent factors using the control units, with a data-driven method to determine the number of factors, following chen2021quantile. Next, we take the estimated factors into quantile regressions on the treated units to deliver QTT estimates. Our proposal complements the flourishing SCM literature by providing methodologies for evaluating distributional causal effects.
We consider two QTT estimators differing at their estimation of common factors in the first stage, where the first one relies on a nonsmooth objective function (the check function) and the second one is based on a smoothed alternative. We refer to them as NQTT (Nonsmooth QTT) and SQTT (Smoothed QTT), respectively. At the cost of more computational burden, SQTT achieves a faster convergence rate of the estimated factors, lending us the inference theory for the final QTT estimator. We prove the consistency for NQTT and the asymptotic normality for SQTT. For SQTT, we provide a blockwise bootstrap method horowitz2019bootstrap to construct confidence intervals, since the asymptotic variance might not be accurately estimated, as is well addressed in quantile regressions koenker2017handbook. We further establish the consistency of our bootstrap procedure, offering a theoretical justification for its validity. Monte-Carlo simulations are conducted to demonstrate a good finite sample performance of our method. To illustrate the implementation of our proposal, we apply it to analyze the macroeconomic effects of the 2008 Chinese Economic Stimulus Program ouyang2015treatment. Our QTT estimation supports that the stimulus package has a significant effect on GDP growth for quantiles lower than 50%. This finding reinforces the widely circulated perception that China's fiscal policies in 2008 served as a safety net, preserving the bottom line of economic growth and ensuring that the Chinese economy could maintain high-speed growth even amidst unfavorable global economic conditions. Finally, we discuss several possible extensions for our method, including time varying QTT, multiple treated units, and the use of covariates.
Despite its importance, the literature on quantile treatment effects within the SCM framework is relatively sparse. chen2020distributional extends the SCM abadie2003economic to evaluate the distributional effects of policy interventions in the possible presence of poor matching, but the “distributional" information comes from subunits (or individuals) within an aggregate-level unit. That is, researchers must observe $\{y_{i\ell t}\}_{\ell=1}^L$ to identify the distribution of unit $i$ at time $t$. gunsilius2023distributional imposes a similar data requirement and constructed counterfactual distribution for the treated unit by taking a weighted average of the untreated $y_{it}$'s distributions, while chen2020distributional focuses on a given quantile. This line of methodologies, while innovative, may not always be applicable in empirical settings where we only observe the aggregate level outcome $y_{it}$, such as in cases involving macroeconomic adrian2019vulnerable, or financial ando2020quantile variables. cai2022estimating also consider quantile treatment effects estimation in panel data building on hsiao2012panel. They propose to construct the counterfactual distribution with the pre-treatment data by first nonparametrically estimating the conditional cumulative distribution function (CDF), and then transforming it into unconditional CDF to obtain the estimated quantile function. However, this method may falter when the total number of covariates and control units is large due to the curse of dimensionality of the nonparametric estimation. When facing a high dimensional panel data, they provide an alternative estimation scheme that relies on penalized quantile regressions, whereas its theoretical properties are underexplored.
Compared to existing works, our method has the following advantages. Firstly, it can be directly extended to the multiple treated units cases. abadie2021using highlighted the challenge that classical matching methods with simplex conditions may provide multiple (even infinitely many) solutions. This would lead to a practical issue of choosing proper weights, when the multiple treated units have quite different characteristics. Secondly, our method effectively handles the high-dimensionality in panel data, as the factor models are inherently designed for dimension reduction. Therefore, we are able to capture the common trends from a large number of control units and use them to identify the underlying causal effects on the treated unit. Finally, compared to the standard factor models as in bai2009panel and xu_2017, QFM captures hidden factors that may shift characteristics (moments or quantiles) of the outcome distribution beyond its mean, permitting the factors, factor loadings, and number of factors to vary across the quantile levels of the outcome distribution.
The remainder of the paper is organized as follows. Section (ref) describes the model, outlines the estimation procedure, and proposes the blockwise bootstrap method to construct confidence intervals. Section (ref) establishes the asymptotic theory for the proposed estimators. Section (ref) shows the results of the Monte Carlo simulation. Section (ref) illustrates the empirical implementation of our method by evaluating the policy effect of the 2008 Chinese Economic Stimulus Program. Section (ref) explores potential extensions of our estimators. Section (ref) concludes. The proofs as well as additional simulation and empirical results are presented in the Appendix.
We posit the potential outcome framework in a panel data environment. Suppose $d_{it}\in\{0,1\}$ is a binary treatment indicating whether unit $i=1,\cdots,N,N+1$ is treated at time $t=1,\cdots,T$. Let $y^1_{it}$ and $y^0_{it}$ denote the potential outcomes under and in the absence of treatment, respectively. Following the convention in the SCM literature, we consider the case that only the first unit $i=1$ receives the treatment which occurs at time $T_0+1$, and continues till the last time period $T$. Denote $T_1=T-T_0$ as the number of treated periods. Resembling the Stable Unit Treatment Value Assumption (SUTVA) rubin1980randomization, we suppose there is no “anticipation effect" so that the treatment does not affect the treated unit in the pretreatment period, and there is no “interference effect" so that the control units are not influenced by the treatment abadie2010synthetic. Therefore, the observed outcome is given by $y_{it} = d_{it}y^1_{it} + (1-d_{it})y^0_{it}$.
Generally, researchers are interested in ATT, i.e., $\mathbb{E}[y_{1t}^1-y_{1t}^0]$ for $t>T_0$, and numerous ATT estimators and their properties have been developed in the SCM literature, e.g., abadie2010synthetic, hsiao2012panel, xu_2017, and carvalho2018arco, to name a few. If, however, the distributions of potential outcomes poorly concentrate around the mean, the distributional change of the outcomes, instead of ATT, would better represent the treatment effects. Next, we will present our model and the QTT estimand.
We consider that the correlations among cross-sectional units are due to some common factors that drive all cross-sectional units, whose impacts on each cross-sectional unit may be different, and allow that hidden factors may shift characteristics (moments or quantiles) of the distribution of outcome variable. Moreover, the common factors, factor loadings, as well as the number of factors could vary across distributional characteristics. Therefore, we introduce the following QFM chen2021quantile at some quantile $\tau \in (0,1)$ :
where $Q_{\tau}(y|x)$ denotes the $\tau$-th quantile of $y$ conditional on $x$. Here $f_t(\tau)$ is an $r(\tau) \times 1$ vectors of common factors, $\lambda_i(\tau)$ is an $r(\tau) \times 1$ vector of factor loadings, with $r(\tau) \ll N$. The notation of $f_t(\tau)$, $\lambda_i(\tau)$ and $r(\tau)$ indicates that they are all allowed to be quantile-dependent. Our causal parameter of interest QTT is $\delta(\tau)$ as defined in ((ref)), as in firpo2007efficient, chernozhukov2005iv, etc. If $y_{it}$ is GDP growth and $d_{it}$ is the economic stimulus program, then QTT $\delta(\tau)$ represents the difference of the $\tau$-th quantile of GDP growth distribution between receiving and not receiving the program. For a small $\tau$, $\delta(\tau)$ reflects the policy impact under economic downturns. Notice that the model would turn to xu_2017's GSCM setting when one replaces the quantile operators with expectation operators. The above QFM can be written as:
where $\epsilon_{it}(\tau)$ has zero $\tau$-th quantile conditional on $f_t(\tau)$ , i.e. $Q_{\tau}(\epsilon_{it}|f_t(\tau))=0$.
Consider a DGP used in our simulation as an illustrative example:
where $\alpha_{t} = u_{1t} + 0.5$ and $u_{it}\stackrel{\text{i.i.d}}{\sim}\mathcal{N}(0,1)$. The mean of $y^0_{it}$ is determined by $f_{1t}$ and $f_{2t}$, while the quantile of $y^0_{it}$ is additionally determined by $f_{3t}$. We may fit this location-scale shift model for $y_{it}^0$ into QFM by writing $\lambda_i(\tau)=(\lambda_{1i},\lambda_{2i})'$, $f_t(\tau) = (f_{1t},f_{2t})'$ if $\tau= 0.5$, and $\lambda_i(\tau)=(\lambda_{1i},\lambda_{2i},\lambda_{3i} \Phi^{-1}(\tau))'$, $f_t(\tau) = (f_{1t},f_{2t},f_{3t})'$ if $\tau\neq 0.5$. The standard factor model would fail to capture $f_{3t}$ at $\tau \neq 0.5$, while QFM is capable of describing such a quantile-dependent feature. In this case, $\epsilon_{it}(\tau) = y^0_{it} - Q(y^0_{it} \mid f_t)=\lambda_{3i} f_{3t} (u_{it}-Q_\tau(u_{it}|f_t))$.
In the above example, we see that the true QTT is $\delta(\tau) = 0.5 + \Phi^{-1}(\tau)$, where $\Phi$ is the CDF function of the standard normal distribution. Although the QTT varies over $\tau$ and reflects the distributional changes due to the treatment, it remains constant for a given quantile. Different from chen2020distributional and gunsilius2023distributional who could utilize subunit level data $\{y_{1\ell t}\}_{\ell=1}^L$ to obtain the quantile functions $Q_\tau(y_{1t})$ and identify QTT $\delta_t$ for each $t$, we only observe the aggregate outcome $y_{1t}$, so it is only possible to identify a finite-dimensional parameter for the distribution of treatment effect.\footnote{In this example, the quantile of treatment effect $\alpha_t$ equals to our defined QTT. See hansen2022econometrics for further discussions on the definitions of quantile treatment effects.} In this sense, we share a similar spirit with cai2022estimating. However, our quantile regression framework in the second stage could be extended to allow QTT to be functions of time $t$ or factors $f_t$, which enables a parametric dependence of QTT on time. We discuss this extension in Section (ref).
We propose our QTT estimator in this subsection. We first estimate the common factors in QFM using the control units following chen2021quantile, with a data-driven method to determine the factor numbers. Next, we estimate QTT by quantile regressions based on the estimated factors using the treated unit.
To simplify the notations, we suppress hereafter the dependence of $\delta(\tau)$, $f_t(\tau)$, $\lambda_i(\tau)$, $r(\tau)$ and $\epsilon_{it}(\tau)$ on $\tau$. Denote the true values of $\{f_t\}$, $\{\lambda_i\}$, and $\delta$ as $\{f_{0t}\}$, $\{\lambda_{0i}\}$, and $\delta_0$, respectively. Let $\Delta\subset \mathbb{R}$ be the parameter space for $\delta$. Following chen2021quantile, we treat $f_{0t}$ and $\lambda_{0i}$ as parameters to be estimated. Let $\theta_0=(\theta_{0\lambda}',\theta_{0f}')'=(\lambda_{01}',\lambda_{02}',...,\lambda_{0N}',f_{01}',...,f_{0T}')'$, and $\theta=(\theta_{\lambda}',\theta_{f}')'=(\lambda_1',\lambda_2',...,\lambda_N',f_1',...,f_T')'$. Following the convention in the factor model literature, we impose normalization conditions for the identification of factors and loadings for any positive integer $T$ and $N$, i.e.,
Let $M=(N+T)r$, $\mathcal{A},\mathcal{F}\subset\mathbb{R}^r$, and define
chen2021quantile propose two estimators for the common factors in the QFM. The first method considers the minimization of the following objective function.
where $\rho_{\tau}(a)=a(\tau-\mathbf{1}(a \leq 0))$ is the check function. We estimate $\theta_0$ by $\widehat{\theta}= \underset{\theta\in\Theta^r}{\text{argmin }} \mathbb{M}_{NT}(\theta)$ where the common factors $f_{0t}$ are a part of $\theta_0$. Since $\mathbb{M}_{NT}(\theta)$ is not convex in $\theta$, chen2021quantile introduce the Iterative Quantile Regression (IQR) algorithm. Let $F=(f_1,...,f_T)'$, $\Lambda=(\lambda_1,\lambda_2,...,\lambda_N)'$ and
Based on the fact that $\mathbb{M}_{i,T}(\lambda,F)$ is convex in $\lambda$ for each $i$, and $\mathbb{M}_{N,t}(\Lambda,f)$ is convex in $f$ for each $t$, an iterative procedure of optimizations for $\mathbb{M}_{i,T}(\lambda,F)$ and $\mathbb{M}_{N,t}(\Lambda,f)$ could lead to the optimization of $\mathbb{M}_{NT}(\theta)$. Algorithm (ref) illustrates the IQR algorithm for a given $\tau$ and $r$.
While the IQR algorithm is computationally efficient and convenient, its convergence rate for the estimated factors $\widehat{f}_t$ is only sufficient for the consistency, but not for the asymptotic normality of the QTT estimator in the second stage. In particular, as shown in the proof of our Theorem (ref), the asymptotic normality result requires that
where $c_t$ is a bounded sequence, and $\widehat{S}=\operatorname{sgn}(\widehat{F}'F/T)$. However, chen2021quantile shows that $T^{-1/2}\sum_{t=1}^T\|\widehat{f}_t-\widehat{S}f_{0t}\|^2=O_p(L_{NT}^{-1})$ where $L_{NT}=\min\{\sqrt{N},\sqrt{T}\}$, so that one can only show
which is at best $O_p(1)$. The underlying reason is the non-smoothness of the check function in the objective function (ref). bai2008extremum also considers M-estimation where the regressors are estimated from the standard factor models bai2003inferential, and establish asymptotic normality for the regression coefficients. However, they manage to achieve the rate in (ref) using the stochastic expansion of $\widehat{f}_t-\widehat{S}f_{0t}$ in standard factor models bai2002determining, and its extension to QFM based on IQR appears not trivial chen2021quantile.
In view of that, we consider a second method for estimating $f_{0t}$ proposed in chen2021quantile which admits asymptotic normality for the QTT estimator. It replaces the non-smooth objective function (ref) by a smoothed alternative:
where $K(z)=1-\int_{-1}^z k(z) dz$, $ k(z)$ is a continuous kernel function with support $[-1,1]$, and $h$ is a bandwidth parameter. Define $ \widetilde{\mathbb{M}}_{i,T}$ and $ \widetilde{\mathbb{M}}_{N,t}$ as the smoothed version of (ref):
and the estimation procedure is just modifying Algorithm (ref) by using the smoothed objective functions.
In Lemma (ref) in the Appendix it can be seen that the estimated factors $\widetilde{f}_t$ possess the desired convergence rate in (ref). In the next section, we show that QTT estimator based on $\widetilde{f}_t$ is asymptotically normal which facilitates valid statistical inference.
While in the above algorithms, $r$ is supposed to be known, the researcher needs to first determine it and usually prefer a data-driven method to do so. Let $r_0$ denote the true factor number, and $P_{NT}$ to be a sequence goes to 0 as $N,T\rightarrow \infty$. chen2021quantile use Algorithm (ref) to choose a proper factor number by rank minimization, and show that it produces a consistent estimator for $r_0$.
The intuition of the rank minimization is, if we select a number of factors $k$ larger than the true $r_0$, as $N$ and $T$ grows, only the largest $r_0$ diagonal elements of $(\widehat{\Lambda}^k)'\widehat{\Lambda}^k /N$ would stay positive, while the rest $k-r_0$ diagonal elements would converge to zero. The efficiency of the rank minimization comes from the fact that one single proper value of $k$ leads to a consistent estimator $\widehat{r}_{RM}(\tau)$, compared to other information criterion techniques. Though we fail to observe the $r_0$ in a real world, we always assume the true factor numbers, under all possible quantiles, should be “small" enough. In practice, chen2021quantile suggest $P_{NT} = \widehat{\sigma}_{N,1}^k \cdot L_{NT}^{-2/3}$, where $ L_{NT} = \min\{\sqrt{N},\sqrt{T}\}$, and $k \geq 8$ is recommended.
Now we are ready to move to the estimation of the QTT. From Algorithm (ref) and (ref) we have the estimated common factors ${\widecheck{f}_t}$, where we use ${\widecheck{f}_t}$ to denote $\widehat{f}_t$ or $\widetilde{f}_t$. To find the QTT, we simply conduct a quantile regression of the outcome on the estimated factors as well as the treatment, using the treated unit data. For a given quantile $\tau$, our QTT estimators is summarized in Algorithm (ref).
Though the asymptotic normality of the QTT estimator $\widetilde{\delta}(\tau)$ based on ISQR is established in the next section, a direct estimation of the asymptotic variance could be unstable. This is because the asymptotic variance involves the conditional density function of the idiosyncratic errors (see Theorem (ref) below), whose nonparametric estimation might be inaccurate in finite samples koenker2017handbook. Following the quantile regression literature, we use bootstrap method for inference. Since the common factors $f_{0t}$ might be time dependent, we apply the blockwise bootstrap kunsch1989jackknife which is developed for inferences with time series data and has been adopted in quantile regressions gregory2018smooth. Compared to the conventional pairwise bootstrap, the blockwise bootstrap respects the sequential order of the data by putting a certain number of sequential observations into one block, and taking random draws among the blocks. Our bootstrap inference procedure is summarized in Algorithm (ref).
In the next section, we show that this blockwise bootstrap is consistent, in the sense that as $T\to\infty$, the bootstrap distribution becomes close to the asymptotic distribution of the QTT estimator, so that the confidence interval constructed from the bootstrap sample standard deviation is valid.
In this section, we provide asymptotic properties for the proposed QTT estimators. All proofs are presented in the Appendix. We first elaborate on the required assumptions. Denote $w_{0t}=(f_{0t}',d_{1t})'$.
Assumption (ref) is the strong factor assumption which is satisfied when each factor has a non-trivial contribution to the outcome, and we can order the factors by the distinct $\sigma_j$. This is necessary for the true number of factors to be $r_0$, which is standard in the factor model literature bai2003inferential. Assumption (ref) regulates the density functions of the idiosyncratic errors conditional on true factors. In particular, Assumption (ref)(i) requires that the conditional densities of errors are all positive for the control units, but does not require their moments to exist. This allows for heavy-tailed idiosyncratic errors such as standard Cauchy or some Pareto distributions. In the simulation study in Appendix, we show that our method is robust to a $t(2)$ distributed error. Assumption (ref)(ii) states that for the treated unit, the conditional densities are uniformly upper bounded, which is satisfied by many commonly used distributions. Assumption (ref) restricts the idiosyncratic errors of the control units to be mutually independent given the common factors. However, cross-sectional and temporal correlations among $\epsilon_{it}$ driven by common factors are still admitted. Moreover, we allow for cross-sectional and serial heteroscedasticity among $\epsilon_{it}$, making it a more general framework. Assumption (ref)-(ref) lead to consistent estimates for the common factors using the control units in the first stage of our estimation chen2021quantile.
Assumption (ref) posits the stationarity and weak dependence structure on the random shocks on the treated unit. Since $\epsilon_{1t}$ represents the modeling errors that are not captured by common factors, it is natural to assume that they follow this standard assumption about idiosyncratic shocks. Assumption (ref) lends us the law of large numbers and the central limit theorem for quantile regression with time series data gregory2018smooth, which is commonly used in the asymptotic analysis of synthetic control methods li2017estimation,cai2022estimating. Assumption (ref)(i) posits the strict exogeneity which is typical to identify quantile treatment effects firpo2007efficient. It requires that there are nothing other than the common factors that jointly affect the outcome and treatment. Assumption (ref)(ii) establishes the uniqueness of the parameters $\beta_0=(\lambda_{01},\delta_0)$ in the quantile regression in Algorithm (ref). Denote $\Upsilon=\lim_{T\to\infty}T^{-1}\sum_{t=1}^T w_{0t}w_{0t}'$. The positive definiteness of $\Gamma$ holds if $\Upsilon$ is positive definite, $h_{1t}(0)$ is uniformly bounded away from zero and $h_{1t}(\epsilon)$ does not depend on $f_{0t}$ at $\epsilon=0$. Assumption (ref)(iii) requires that we have enough observations in both pre-treatment and post-treatment periods to identify the quantile treatment effect.
Our first theoretical result is that the NQTT estimator $\widehat{\delta}$ based on Algorithm (ref) (IQR) is consistent.
As illustrated in the previous section, the estimated factor $\widehat{f}_t$ from Algorithm (ref) may not achieve the desired convergence rate for $\widehat{\delta}$ to be asymptotically normal. However, $\widetilde{f}_t$ from Algorithm (ref) is able to attain such rate due to its smoothed objective function. To establish the asymptotic normality for $\widetilde{\delta}(\tau)$, we need some additional notations and assumptions. Define $\psi_\tau(u)=\tau-\mathbf{1}(u<0)$, and
Assumption (ref) requires $\Sigma$, a component of the asymptotic variance of $\widetilde{\delta}(\tau)$, to be positive definite. $\Sigma$ is the long-run variance-covariance structure of $\epsilon_{1t}$ that contains the temporal dependence in the data davidson1994stochastic. If $\epsilon_{1t}$ is uncorrelated across $t$, then $\Sigma$ equals $\tau(1-\tau)\Upsilon$ as the quantile regression for i.i.d. data. Assumption (ref) resembles Assumption 2 in chen2021quantile, which is standard in smoothed quantile regressions galvao2016smoothed. Assumption (ref)(i) is a similar version of Assumption (ref)(ii) for the control units. Assumption (ref)(ii) is mild once we properly define the parameter space $\mathcal{A}$ and $\mathcal{F}$. Assumption (ref)(iii) is a technical requirement for high-order kernels used in the smoothed objective function and many candidates are available in the literature li2007nonparametric. Assumption (ref)(iv) demands more smoothness for conditional density functions. Assumption (ref)(v) states that $N$ and $T$ should grow at the same rate, rendering a balanced panel, and regulates the rate of bandwidth $h$ of the kernel.
Theorem (ref) enables pointwise inference for the SQTT estimator $\widetilde{\delta}$, e.g., the confidence interval for each $\tau$-th quantile as shown in Figure (ref) and (ref) in Section (ref) for some given $\tau\in(0,1)$. It would be interesting to examine the uniform inference for an interval of quantiles $\mathcal{T}=[\tau_1,\tau_2]\subset (0,1)$, which allows researchers to test global hypotheses about the entire distribution of potential outcomes en bloc, and facilitates comparisons between different quantiles. Establishing uniform inference for $\widetilde{\delta}$ might require the use of empirical process theory vaart2023empirical to deal with the estimated regressors $\widetilde{f}_t$ in the quantile regression process $\{\sqrt{T}(\widetilde{\delta}(\tau)-\delta_0(\tau)),\tau\in\mathcal{T}\}$, which is essentially different from proving the pointwise asymptotic distribution angrist2006quantile. We leave this for the future research avenue.
Remarkably, the asymptotic properties of our QTT estimators are not affected when the number of factors is consistently estimated as in Algorithm (ref). Specifically, chen2021quantile show that $\lim_{N,T\rightarrow\infty} \mathbb{P}[\widehat{r}_{RM}=r_0]=1$. Following the argument in bai2003inferential, for $\widecheck{f}_t$ being $\widehat{f}_t$ or $\widetilde{f}_t$, we have $\mathbb{P}[\widecheck{f}_t\leq z]=\mathbb{P}[\widecheck{f}_t\leq z,\hat{r}_{RM}=r_0]+\mathbb{P}[\widecheck{f}_t\leq z,\hat{r}_{RM}\neq r_0]$. Observe that $\mathbb{P}[\widecheck{f}_t\leq z,\widehat{r}_{RM}\neq r_0]\leq \mathbb{P}[\widehat{r}_{RM}\neq r]=o(1)$. Thus,
since $\mathbb{P}[\widehat{r}_{RM}=r_0]\to 1$. Therefore, asymptotic properties of the estimated factors $\widecheck{f}_t$ using $\widehat{r}_{RM}$ are the same as using $r_0$, and so are our QTT estimators.
The last part of this section is to establish the bootstrap validity for Algorithm (ref). We need an additional assumption for the bootstrap block size.
Assumption (ref) simply requires that the block size increases with $T$ at a slow enough rate, which is a mild restriction. The next theorem shows that the block boostrap procedure in Algorithm (ref) is consistent for the asymptotic distribution in Theorem (ref) horowitz2019bootstrap, so that the bootstrap confidence interval is valid.
In this section, we report the finite sample performance of our method using Monte-Carlo experiments. The sample sizes are set to be $N+1 \in \{51,101,201\}$ and $T \in \{100,200,400\}$. A treatment starts to affect the outcome of interest $y$ at period $T_0 = T/2+1$. We consider a DGP for the potential outcome in the absence of treatment following chen2021quantile:
where $f_{1t}=0.8f_{1,t-1}+\nu_{1t}$, $f_{2t}=0.5f_{2,t-1}+\nu_{2t}$, $f_{3t}=|g_{t}|$, $\lambda_{1i}, \lambda_{2i}, \nu_{1t}, \nu_{2t}, g_{t}, u_{it}\stackrel{\text{i.i.d}}{\sim}\mathcal{N}(0,1)$, and $\lambda_{3i}\stackrel{\text{i.i.d}}{\sim}\mathcal{U}(1,2)$. The observed outcome is the following:
We set $\alpha_{t} = u_{1t} + 0.5$ so that the true quantile treatment effect is $\delta_{0}(\tau) = 0.5 + \Phi^{-1}(\tau)$, where $\Phi$ is the CDF function of the standard normal distribution. One may fit (ref) into QFM by writing $\lambda_i(\tau)=(\lambda_{1i},\lambda_{2i})'$, $f_t(\tau) = (f_{1t},f_{2t})'$ if $\tau= 0.5$, and $\lambda_i(\tau)=(\lambda_{1i},\lambda_{2i},\lambda_{3i} \Phi^{-1}(\tau))'$, $f_t(\tau) = (f_{1t},f_{2t},f_{3t})'$ if $\tau\neq0.5$.
We compare the proposed estimators with two other estimation results, where NQTT (Non-smoothed QTT) and SQTT (Smoothed QTT) denote the estimators $\widehat{\delta}$ and $\widetilde{\delta}$ following Algorithm (ref) and (ref), respectively. In SQTT, we use the kernel function suggested by chen2021quantile\footnote{$k(z)=\mathbf{1}(|z|<1)\frac{3465}{8192} \left(7 - 105 z^{2} + 462 z^{4} - 858 z^{6} + 715 z^{8} - 221 z^{10} \right)$}, and set $h=0.5$ (since $h$ would not change much with $T$ given the rate in Assumption (ref)(v)). Oracle denotes the direct quantile regression had we observed the complete set of factors $f_t$, and GSCM denotes the general synthetic control method of xu_2017. While Oracle simplifies to a standard quantile regression estimation, \emph{GSCM} first estimates the factor number $\widehat{r}_{IC}$ based on the information criteria considered in bai2002determining\footnote{$\widehat{r}_{IC} = \operatorname*{argmin}_{r \in \{2,\dots,5\}} \text{IC}(r) = \operatorname*{argmin}_{r \in \{2,\dots,5\}} \ln \left( \frac{1}{NT} \sum_{i=2}^{N+1}\sum_{t=1}^T (y_{it}-\widehat{\lambda}_i'\widehat{f}^r_{t,PCA})^2 \right)+ r\frac{N+T}{NT} \ln ( \frac{NT}{N+T} ).$}, and then uses principle component analysis (PCA) to estimate the factor $\widehat{f}^r_{t, PCA}$, given $r=\widehat{r}_{IC}$. Finally, for each quantile level $\tau$, \emph{GSCM} runs the standard quantile regression with $\widehat{f}^r_{t, PCA}$ and an indicator. The coefficient of the indicator is considered as the estimated treatment effect. Note that $\widehat{f}^r_{t, PCA}$ remains the same across $\tau$, which is different from our two QTT estimators.\footnote{We note that \emph{GSCM} is simplified to PCA in our simulations, since we do not consider the existence of covariates. When covariates are included, \emph{GSCM} relies on an iterative procedure to estimate the covariate coefficients, factors, and loadings.}
First, we assess the performance of the estimators by computing the bias and rooted mean square error (RMSE) at $\tau\in\{0.1,0.25,0.5,0.75,0.9\}$ across $R=1000$ simulations:
where $m \in \{NQTT, SQTT, Oracle, GSCM\}$. Table (ref) presents the result. We see that NQTT and SQTT both produce consistent estimates across all quantiles compared to Oracle. In contrast, GSCM exhibits severe biases at $\tau\neq0.5$, as GSCM could not capture $f_{3t}$ and thus fails to recover the true quantile structure in this case.
Second, we validate our inference theory by reporting the standard deviation and coverage probability for the 95% confidence intervals using the blockwise bootstrap procedure. The bootstrap sample size is $B=1000$. The standard deviation and coverage probabilities are calculated as follows:
where the estimator $\widecheck{\delta}$ is either NQTT $\widehat{\delta}(\tau)$ or SQTT $\widetilde{\delta}(\tau)$. Note that although we only provide the asymptotic normality and bootstrap validity theory for SQTT (Theorem (ref) and (ref)), we also conduct the same bootstrap inference for NQTT in this simulation to examine its finite sample performance. Table (ref) shows that for both NQTT and SQTT, the SD declines as the sample size increases and the accuracy of the coverage probability is satisfactory. This demonstrates the effectiveness of our asymptotic theory for SQTT, and suggests that valid inference might also hold for NQTT.
In the Appendix, we conduct further simulations to safeguard our estimators under different DGP setups, including heavy-tailed errors (allowed by our theory), serially and cross-sectionally correlated errors (not allowed by our theory), and more complicated factor structures (allowed by our theory). We arrive at similar results as the above baseline simulation in which our two proposed estimators consistently recover the true quantile treatment effects and permit valid bootstrap inference.
In this section, we apply our proposed method to evaluate the effects of the 2008 Chinese Economic Stimulus Program.\footnote{We revisit California’s tobacco control program abadie2010synthetic in Section (ref) in the Appendix.} At the end of 2008, the Chinese government launched an economic stimulus package amounting to four trillion RMB (equivalent to 586 billion USD) as a response to the global financial crisis, with the aim of counteracting its adverse impacts on the domestic economy. Assessing the effectiveness of this fiscal intervention will deepen our understanding on large-scale economic stimulus programs.
The economic outcomes following the policy’s implementation are influenced by various latent factors, such as global financial crises and trade dynamics, alongside the stimulus itself. Recognizing these complexities, ouyang2015treatment improve the SCM approach originally proposed by hsiao2012panel, which is grounded in a factor model structure, to estimate the ATT of the stimulus. Their findings revealed that the fiscal stimulus program boosted China’s annual real GDP growth by approximately 3.2% on average.
Although the ATT provides valuable insights, policymakers might also be interested in how the stimulus program influenced the distribution of macroeconomic outcomes. Specifically, examining the QTT offers a deeper understanding of the policy impact under varying economic conditions. For instance, QTT at the 90% quantile sheds light on the policy’s effectiveness in optimistic scenarios, while QTT at lower quantiles highlights its role during economic downturns. Next, we apply our QTT estimators to explore the impact of the stimulus program at different quantiles of GDP growth and investment.
To evaluate QTT, we collect quarterly real GDP growth data from the OECD statistical databases, following the same data source as ouyang2015treatment, but with an expanded dataset. The control group comprises 40 countries, and the GDP data spans from the first quarter of 1999 to the fourth quarter of 2015, giving a pre-treatment period of $T_0 = 40$ quarters and a post-treatment period of $T_1 = 28$ quarters. The real GDP growth is measured as the annual growth rate—the difference between the GDP level of the current quarter and the same quarter of the previous year—to ensure stationarity and eliminate seasonality.
The pre-treatment trend of GDP growth rate is described in Figure (ref). We plot the actual China GDP growth rate together with the quantile predictions $\widehat{Q}_\tau(y_{1t}|\widehat{f}_t)$, where the estimated factors $\widehat{f}_t$ is from QFM using the control group. We see that the median prediction closely traces the observed outcome, suggesting that the estimated factors from the control group explain the treated unit fairly well. This aligns with the convex hull condition abadie2021using in SCM that a combination of unaffected units can approximate the pre-treated outcome of the affected unit. Our QTT, for example, at $80\%$ quantile, is to uncover how $\widehat{Q}_{0.8}(y_{1t}|\widehat{f}_t)$ would change after the policy. The difference between observed outcome and the quantile predictions, i.e., $y_{1t}-\widehat{Q}_\tau(y_{1t}|\widehat{f}_t)$, is just the (estimated) idiosyncratic error $\widehat{\epsilon}_{it}(\tau)$, which may reflect the random shock unexplained by latent factors.
We apply our QTT estimators to study the impact of the 2008 economic stimulus program on China's real GDP growth. Figure (ref) presents the estimated QTTs across different quantiles, along with 95% confidence intervals based on blockwise bootstrap with $1000$ replications.
The results are consistent between NQTT and SQTT, indicating that the stimulus program had significantly positive effects on GDP growth at quantiles below 50%, suggesting that the policy served as a safety net to sustain economic growth during adverse global economic conditions. However, for the quantiles above 50%, while the estimated effects remain positive, they are not statistically significant. This highlights the program’s primary function as a stabilizer rather than as a mechanism to maximize economic growth in favorable scenarios. The findings underscore that fiscal stimulus policies are more effective in mitigating economic downturns than boosting growth ceilings.
We further examine the influence of the stimulus program on one of the key components of real GDP: Gross Fixed Capital Formation (investment). ouyang2015treatment report that the stimulus program increased China’s real investment growth by 22.15% and note that investment responded with the fastest and largest magnitude among the economic components.
Our analysis in Figure (ref) uncovers an intriguing pattern: the QTTs are significant for quantiles between 30% and 50%, where investment growth significantly outpaces GDP growth. This suggests that the influence of the stimulus program on investment is most pronounced under relatively bad economic conditions, while its effects diminish in other scenarios: either favorable or highly unfavorable. Moreover, the finding also indicates that the fiscal stimulus may skew resource allocation excessively toward investment at the expense of other GDP components. This imbalance highlights the potential unintended consequences of stimulus measures and underscores the importance of carefully considering their broader economic implications in future policy designs.
As a conclusion, the empirical findings provide a comprehensive view of the 2008 Chinese Economic Stimulus Program’s impacts. The analysis of QTTs reveals that the stimulus effectively stabilized economic growth and bolstered investment under economic downturns, potentially at the cost of directing resources toward investment and suppressing the growth of other GDP components. Such findings reinforce the view that fiscal stimulus should be considered as a temporary measure, and policymakers should carefully weigh its unintended consequences.
In this section, we discuss some possible extensions of our method. As before, we omit the subscript $\tau$ for $\delta$ and $f_t$ whenever it does not cause confusion.
Our model (ref) implies a constant QTT $\delta = Q_{\tau}(y^1_{it} | f_t) - Q_{\tau}(y^0_{it} | f_t)$ for a given $\tau$. In many application, the policy might only exert transient impact so that it is more reasonable to consider a treatment effect that gradually fades out over time. Therefore, we discuss two extensions of our model in which the QTT is allowed to vary over time.
We could consider QTT as a function of time, i.e., $\delta_t = g(t)$, where $g(\cdot)$ is a deterministic function. For instance, one may let $g(t)=\zeta_0 + \zeta_1 / (t - T_0)$ for all $t \geq T_0 + 1$, which describes the idea that QTT gradually decays as time increases. Alternatively, if the researcher believes the intervention gradually builds up to a peak within a certain period after $T_0$ and subsequently diminishes, $g(t)$ can be modeled as a quadratic function of $t$.
We could directly extend the estimation algorithms to this case. While estimating common factors remains completely unchanged, the QTT estimator only requires a minor modification of (ref), which becomes:
where $\delta_t=g(t;\zeta)$ is parametrized by $\zeta$. However, including a function of $t$ in the quantile regression may lead to a different asymptotic property of the estimator $(\widecheck{\lambda}_{1},\widecheck{\zeta})$, which requires more involved and case-by-case theoretical analysis. For example, for $g(t)=\zeta_0+\zeta_1t$, the convergence rate for $\widehat{\zeta}_1$ would be $T^{3/2}$ instead of $T^{1/2}$ hamilton2020time.
We could also allow QTT to be a function of common factors. For example, we may extend (ref) to be $Q_{\tau}(y^1_{it}|f_t) =\zeta_0 + (\lambda_i+\zeta_1)' f_t$. In other words, the treatment may influence the potential outcome through the factor loadings as well, and the true QTT becomes $\delta_t=\zeta_0+\zeta_1'f_t$. This is analogous to adding an interactive term in the linear model. For estimation, we can still apply Algorithm (ref), with the following modification in Step 2:
One may consider more complicated interactions between treatment and factors, or even semi/nonparametric functional forms, and our theoretical results can be extended to this case. However, caution should be taken when interpreting the parameters $\zeta$, since $f_t$ are unobserved latent factors and lack real-world meaning.
Our QTT estimation method could be easily extended to the multiple treated units case, which typically raises heavy computational burdens in SCM. Under the traditional weighting scheme SCM, the weights might be non-unique when there are multiple treated units abadie2021using. To address the issue of multiplicity, abadie2021penalized propose the use of penalized regression. Such penalized regressions should be executed for every treated unit $i$ that may bring in substantial computation as in cai2022estimating. In contrast, our method uses the factors $f_t$ capture the common underlying characteristics in all units, which is estimated only once using the control group. In particular, to extend our method to the case with multiple treated units, we just need to apply Algorithm (ref) for every treated unit $i$:
Since the computation of quantile regression is relatively efficient, our QTT estimators could be easily adapted to the case with multiple treated units.
While our method does not require predictors for $y_{it}$, incorporating information from observed covariates might assist the estimation of the quantile factor model. Suppose we have a $K$-dimensional vector of continuous covariates $x_t=(x_{1t},...,x_{Kt})'$ which shares the same quantile factor structure with $y_{it}^0$:
Then our method could be directly extended by adding $x_t$ to the control units $(y_{it})_{i=2}^{N+1}$ to estimate the latent factors $f_t(\tau)$. Good candidates of $x_t$ may include predictors for the outcome of the treated unit $y_{1t}$, or variables that capture the common time trends. As always, substantive knowledge is helpful for finding good covariates. For instance, in the asset pricing literature, some well-known factors, such as those proposed by Fama and French, are found to be significantly correlated with the factors estimated from empirical studies ando2020quantile.
ando2020quantile consider a quantile regression model with interactive fixed effects which is close to the quantile factor model of chen2021quantile and allow for covariates. Since our goal is QTT, recovering the quantile factor structure from the control units using chen2021quantile's quantile factor model is sufficient for us to uncover the latent factor structure for the treated unit. Moreover, chen2021quantile's method has the advantage of efficiently estimating for the number of factors and allowing for heavy tails in error terms. Nonetheless, including covariates using the framework of ando2020quantile is a promising research revenue and is left for future exploration.
This paper proposes quantile treatment effects on treated (QTT) estimators under high dimensional panel data, which extends xu_2017's generalized synthetic control method (GSCM) for estimating average treatment effects on treated (ATT). Our framework estimates QTT in a panel data setting where $N,T\to\infty$ and allows the underlying common factor structure to change with quantiles, where the issue of high-dimensionality is handled via the first-stage estimation of a quantile factor model. We provide asymptotic properties for our proposed estimators, along with a valid inference approach based on blockwise bootstrap. Monte-Carlo simulations demonstrate the effectiveness of our method, which is suitable for studying policy effect in the fields such as macroeconomics and finance where high frequency panel data can be approached. We apply our method to study the impact of the 2008 China Stimulus Program's on China's economy. Our empirical results show that QTT serves as a valuable complement to ATT.
However, our proposal exhibits certain limitations. First, estimations in small sample sizes at the tail quantiles may be less precise, as shown in the simulations. Second, There is potential for improvement in the blockwise bootstrap method, particularly in cases of small samples where the inference at some quantile levels are not precise enough. Moreover, for the QTT estimator $\widehat{\delta}$ based on Algorithm (ref), we only provide a consistency result due to technical reasons illustrated in Section (ref). We conjecture that asymptotic normality may hold under certain conditions and we show its valid inference via simulations, while a formal proof is highly encouraged in the future study.
Beyond the extensions discussed in Section (ref), we believe that the core idea developed in this paper can be fruitfully applied to other emerging panel data causal inference literature, particularly those based on low-rank factor structures, such as matrix completion athey2021matrix. However, such extensions might not be straightforward. While matrix completion determines the rank adaptively through regularization and leverages information from both treated and control units in estimating the factor structure, its implementation requires iterative procedures (see Section 4.2 in athey2021matrix) involving matrix shrinkage operators, which may become computationally burdensome when combined with the IQR/ISQR framework. Furthermore, inference in matrix completion typically relies on subsampling methods and lacks a formal asymptotic theory. These challenges highlight that extending these new-generation panel causal estimators into factor-based QTT frameworks is a nontrivial yet promising direction for future research.
\appendixpage