Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
77,061 characters · 17 sections · 2 citation commands
Distributional Counterfactual Analysis in High-Dimensional Setup
Treatment effect estimation has always been an active research topic in Economics. A causal statement in empirical studies usually concerns the effects of a given treatment (policy, intervention, event) on the population of interest. Since each unit in the population is either treated or untreated at a given time, one of the outcomes is unobservable. Therefore, causal statements often rely on constructing counterfactuals based on the outcomes of the treated and untreated units (controls). Notwithstanding, definitive cause-and-effect statements are usually problematic to formulate, given economists' constraints in finding sources of exogenous variation. For a comprehensive account in the context of program evaluation, see \citet*{Abadie2018EconometricMF} and references therein. Recently, there has been a renewed interest in causal inference using panel (longitudinal) data, particularly when one observes only a single (or a few) treated unit and several potential controls across time. Most of this development has focused on the mean treatment effect. However, much more can be learned by taking a more comprehensive approach. For that reason, we propose a new methodology to evaluate the impact of interventions on the entire distribution of the variable of interest.
The setup consists of a panel structure where a single unit is treated at a given time, and only a few (possibly one) post-intervention periods are available. Although not required, the methodology accommodates situations when the number of units is much larger than the number of time periods (high-dimensional setup). As previously mentioned, we depart from the typical approach of modeling only the conditional mean as in \citet*{cHhsCskW2012}. Instead, we simultaneously model all the conditional (on the untreated units) quantiles of the unit of interest, allowing us to re-construct the conditional distribution without intervention.
The benefits of this approach are threefold. First, we have a comprehensive picture of the intervention effect as we have available the (estimated) counterfactual distribution and not only its mean. For instance, we can detect mean-preserving treatments, such as those that affect only the distribution's tails or variance. Second, although we might take the median as our point estimate, we have readily available confidence intervals for the (possibly random) treatment effect at each post-intervention period. Third, it allows us to construct a new hypothesis test for the null of the no-distributional effect. The test procedure is based on a quasi-pivotal test statistic whose critical values can be easily simulated.
Specifically, we estimate the conditional quantile function using a $\ell_1$-penalized linear quantile regression to account for a potentially large number of units using the pre-intervention period. We provide non-asymptotic probabilistic bounds for the deviation of the estimated quantile function from the true one and the coverage of the estimated conditional quantile in terms of the true one. Both results are valid uniformly over any quantile in a compact set. The bounds are expressed explicitly in terms of the pre-intervention period, the number of peers, the tail condition, the number of relevant peers, and the measure of serial dependency. This makes evident the trade-off between these quantities in the convergence rate. As a corollary, we propose confidence intervals for the treatment effect, valid under simple rate conditions.
Furthermore, the proposed test for the null hypothesis of the no-distribution effect is based on the $\mathcal{L}_p$-norm of the conditional quantile function before and after the intervention without relying on permutation tests. Interestingly, the test statistic has a quasi-pivotal distribution under the null, in the sense that it only depends on a nuisance parameter that can be estimated using the pre-intervention period and, therefore, is valid even for a single post-intervention period. This test's critical values or p-values can be efficiently obtained via a simple Monte Carlo simulation that only draws from independent standard uniform random variables.
The literature on treatment effects is quite vast and diverse. Nonetheless, we see our methodology based on three key features: distributional effects, few treat units, and high-dimensional setup. We borrow core ideas from the quantile treatment effect literature for the first one; see, for instance, \citet*{vCcH2005}. However, we do not exploit heterogeneity among individuals as the source of our distributional effect; this approach was recently taken in the Distributional synthetic controls (gunsilius2021distributional). Instead, we apply this idea in a panel structure assuming a single treated unit. In other words, we only observe a single realization of a treated unit at any given time period, which motivates our second design feature. By focusing only on “few treated units” techniques, our setup is closely related to \citet*{cHhsCskW2012} and \citet*{cCrMmM2016}, the latter also considers a high-dimensional setup, which leads to our third feature. High-dimensional methods and treatment effects are reviewed in \citet*{BCH2014} with the focus on the idea of dealing with treatment effect estimation with a high-dimensional control group for the case when the treatment is considered exogenous conditional on a very large number of unaffected (by the treatment) characteristic of the treated unit. In our case, these observable characteristics will be replaced by the control units' outcome.
In all those papers considering a high-dimensional setup, some sparsity assumption is usually evoked to impose a low-dimensional structure. Exploiting a strong factor structure in a high-dimensional environment, \citet*{fan2021we} proposes a counterfactual of the conditional mean, which can also be seen as a sparsity assumption in terms of the eigenvalue of the design matrix. Finally, the methodology can also be loosely connected to the Synthetic Control Methods (SCM). For a modern reference, see \citet*{abadie2021using}. In particular, \citet*{chen2020distributional} extends the canonical SCM and considers macro-level intervention in a single unit that consists of several individuals (or sub-unit as labeled by the author). In that sense, there are, in effect, several treated (sub-) units at each given period, and distributional effects can be measured as the heterogeneity effect across individuals in the same unit of interest. In a different setup, the same idea of leveraging cross-sectional variation (of a single unit at a given time) is used in \citet*{gunsilius2021distributional}.
The rest of the paper is organized as follows. Section (ref) presents the setup, defines the estimator, and discusses the main assumptions underlying our results. In particular, Subsection (ref) is dedicated to an overview of the main results, which are formally presented in Section (ref), together with additional technical assumptions. We describe the Hypothesis Testing procedure in Subsection (ref), which includes a viable implementation of a simulation procedure to obtain critical values. Section (ref) described the Monte Carlo study and brief analysis of the estimator's finite-sample performance. The proposed methodology is illustrated in an empirical application in Section (ref), and Section (ref) concludes. All proofs are relegated to the Supplemental Material, which also includes some additional figures from the empirical illustration.
Suppose we have $n\geq 2$ units (countries, states, municipalities, firms, etc.) indexed by $i\in\{1,\dots,n\}$. For each unit and for every time period $t\in\{1,\ldots, T\}$ for $T\geq 2$, we observe a realization of the random variable $Z_{it}$. Furthermore, we consider that there is only one unit that suffers the intervention\footnote{The proposed methodology would still be applicable when there are a few treated units by applying it to each treated unit separately. That is the case, for instance, in our empirical application.} (treatment) at time $1< T_0<T$. Without loss of generality, we assume the treated unit to be the unit one ($i=1$). Let $D_{t}$ be a binary variable flagging the periods when the intervention is in place. Using the potential outcome notation, we can express our observable variable as
where, following the literature on treatment effects, $Z_{it}^{(1)}$ denotes the outcome when the unit $i$ is exposed to the intervention at time $t$ and $Z_{it}^{(0)}$ when it is not.
We are interested in cases where once the unit one suffers the intervention after $t=T_0$ it is kept treated throughout the remaining sample, i.e., $D_{s}=1$ implies $D_{t}=1$ for all $t \geq s$. Hence, we write $D_t=\1\{t\geq T_0\}$. It is important to stress that even though the intervention variable $D_{t}$ might seem deterministic, it is, in effect, a random variable to the extent that $T_0$ is a random variable that might depend on both potential outcomes $\{Z_i^{(0)}, Z_i^{(1)}:1\leq i\leq n, 1\leq t\leq T\}$.
Clearly, in our sample, we do not observe $Z_{1t}^{(0)}$ for $ T_0<t\leq T$; for that reason, we refer to it as the counterfactual, i.e., what would the unit of interest has been like had there been no intervention in place. The next assumption restricts the treatment dependency on potential outcomes. Write $Z^{(0)}_{t}:=(Z^{(0)}_{1t},\dots, Z^{(0)}_{nt})'$ and $Z^{(1)}_{t}:=(Z_{1t}^{(1)},\dots, Z_{nt}^{(1)})'$. We make the following identification assumption
Part (a) requires the treatment (or equivalently, the timing of the treatment $T_0$) to be independent of the potential outcomes under no intervention for all units. We do not, however, require the treatment to be independent of the potential outcome under the intervention, namely $Z_t^{(1)}$. Since we are only interested in the treatment effect on the treated, it is analog to the well-known fact in the treatment effect literature that one can consistently estimate the average effect even when $\mathbb{E}(Z_t|D_t)\neq\mathbb{E}(Z_t)$. Part (b) states that the untreated units (peers) are unaffected by the intervention in the unit of interest, or equivalently, its potential outcome is the same with or without the intervention.
Specifically, Assumption (ref)(a) ensures that, conditional on the treatment $D := (D_1,\dots, D_T)'$ (or equivalently on $T_0$), $Z_1,\dots, Z_{T_0}$ and $Z_1^{(0)},\dots ,Z_{T_0}^{(0)}$ share the same distribution. Thus, by postulating an appropriate model for the variables in the absence of intervention
say, one can, in principle, estimate it using the pre-intervention sample $\{ Z_1,\dots Z_{T_0}\}$. Denote the estimated model by $\widehat{\mathcal{M}}$. Assumption (ref)(b) then allows us to extrapolate the (estimated) model to the post-intervention period to construct the desired counterfactual for the unit of interest. In particular, it allow us to claim that $\mathcal{M}(Z^{(0)}_{2t}, \dots, Z^{(0)}_{nt})$ has the same distribution as $\mathcal{M}( Z_{2t}, \dots, Z_{nt})$. We then use the post-intervention sample $\{ Z_{T_0+1},\dots Z_{T}\}$ to compute the counter factual as $\widehat{Z}_{1t}^{(0)}:=\widehat{\mathcal{M}}( Z_{2t}, \dots, Z_{nt})$ for $T_0< t \leq T$.
I consider the following data generating process (DGP) for the units in the absence of the intervention
The GDP described in Assumption (ref) is quite flexible as it leaves the dependency among the $n$ entries of $Z_{t}^{(0)}$ and the serial dependence unspecified. It also does not assume the sequence to be identically distributed or stationary. However, we require the conditional quantiles to be identically distributed as per Assumption (ref) below. The latter is necessary to model the entire (conditional) distribution.
We use absolute regularity ($\beta-$mixing coefficients) as the measure of serial weak dependency of $\{Z_t^{(0)}:1\leq t\leq T\}$. There are several equivalent definitions of $\beta$-mixing coefficients. In our context, the most appropriate appears on page 3 of \citet*{Doukhan1994}, reproduced below for convenience. As opposed to being stated in terms of sigma-algebras as originally defined, the coefficients can be represented as the largest distance (in the total variation norm) between the joint distribution of the past and the future and the product of its marginals. Formally, for $0\leq m <T$, define the $\beta$-mixing coefficient by
where $\P_{\mathcal{Z}_s^t}$ denote the joint distribution of $(Z_{s}^{(0)},\ldots, Z_t^{(0)})$ for $1\leq s\leq t\leq T$. Note that $\beta_m$ might depend on $T$ and $n$, but $\{\beta_m\}$ is non-decreasing sequence in $m$ for each fixed $T$ and $n$. Therefore the process is required to be $\beta$-mixing in Assumption (ref) in in the sense that $\beta_m\rightarrow 0$ as $m\rightarrow\infty$ for each fixed $T$ and $n$.
Numerous studies have been conducted on the mixing properties of strictly stationary linear processes, encompassing but not restricted to strictly stationary ARMA processes, “non-causal” linear processes, and linear random fields. Additionally, research has been developed on related processes like bilinear, ARCH, or GARCH models. For comprehensive insights into the mixing properties of these and other associated processes, refer to Chapter 2 of \citet*{Doukhan1994}.
To exemplify, Assumption (ref) nests the latent (possibly dynamic) factor model structure by setting \[Z_{it}^{(0)}= \mu_i + \lambda_i' F_t+V_{it},\] where $\mu_i$ is a (possibly random) fixed effect for unit $i$, $\lambda_i\in \mathbb{R}^r$; $\lambda_i$ are loading of the $(r\times 1)$ common factor $ F_t$ and $V_{it}$ are the idiosyncratic shocks.
The presence of a common unknown deterministic time-trend $\{\zeta_t:1\leq t\leq T\}$ commonly encountered in the Synthetic Control can also be accommodated. Suppose that \[Z_{it}^{(0)}= \zeta_t + U_{it}\] where $\{U_{t}:=(U_{1t},\dots, U_{nt})':1 \leq t\leq T\}$ fullills Assumption (ref). Then as long as the conditional quantiles are stable in the sense of Assumption (ref), the methodology described below can be directly applied. The case of stochastic trends is subtle due to the fact that an infinite sum of a (underlying) mixing process is not necessarily mixing. Hence, even an integrated linear process would require a careful analysis. For conterfactual analysis of the conditional mean involving stochastic trends refer to masini_jasa.
The presence of observable (covariates) does not present any particular issue as long as we assume that those covariates are unaffected by the treatment $D_t$. If, for instance, we postulate that \[Z_{it}^{(0)}= \pi_i'W_{it} + U_{it},\] where $\{U_{t}:1 \leq t\leq T\}$ fullills Assumption (ref) and we carry on the analsys with the “observable” $\{\widehat{U}_{t}:1 \leq t\leq T\}$ where $\widehat{U}_{t}:= Z_{it} + \widehat{\pi_i}'W_{it}$ for some estimator $\widehat{\beta}_i$ as long as we are able to control for $\sup_{it}|\widehat{U}_{it} - U_{it}|$. Refer to Theorem 1 in \citet*{fan2021bridging} for primitive conditions to bound this quantity when $\widehat{\pi}_i$ is the ordinary least squares estimator.
Recall that unit 1 is the unit of interest (treated). To ease on the notation we set, for $1\leq t\leq T$, \[ Y_t := Z_{1t}^{(0)}\qquad X_{t}:=(1,Z_{2t}^{(0)},\dots, Z_{nt}^{(0)})'. \] In most counterfactual exercises, the model $\mathcal{M}$ appearing in (ref) is the conditional mean or just the linear projection of $Y_t$ onto $X_t$ (see, for example, \citet*{cHhsCskW2012} and \citet*{cCrMmM2016}). Since the focus is to investigate the potential distributional effects of the intervention of interest, we must model the entire counterfactual distribution. Therefore, we take $\mathcal{M}$ as the collection of the conditional quantiles of $Y_t$ given $X_t$. We could also choose to model the conditional distribution as both fully characterize the distributional effect of the intervention. An interesting comparison between those two approaches can be found in \citet*{koenker2013distributional}. Heuristically, we measure the distributional treatment effect by the differences it may cause to the conditional quantiles of $Y_t|{X}_t$, which under Assumption (ref) are attributed to treatment on the unit of interest.
Let $F(\cdot|x)$ denote the conditional distribution function of $Y_t$ given $X_t =x$ i.e., $F(y|x) := \P(Y_t\leq y|X_t=x)$ for $y\in\mathbb{R}$ and $x$ in the support of $X_t$. Note that under Assumption (ref) below, $F(y|x)$ does not depend on $t$. Also, let $Q(\cdot|x)$ denotes the conditional quantile function of $Y_t$ given $X_t$ namely $Q(\tau|x):=\inf\{y\in\mathbb{R}:F(y|x)\geq\tau\}$ for $x$ in the support of $X$ and $\tau\in(0,1)$. We are left to specify the functional form of the conditional quantiles.
Assumption (ref) postulate a correctly specified linear model for the conditional quantile function for all quantiles in $\mathcal{T}$. The consequences of missepecification in the quantile regression model are treated in \citet*{10.2307/3598810}. We could also consider a more flexible specification where we allow the functional form to vary with $\tau$, such that $Q_{Y_t|X_t}(\tau|x)=g_\tau( x,\theta_0(\tau))$ for some know class of function $g_\tau$. Or even a non-parametric specification if, in the empirical application, one expects to have $n$ is much smaller than $T_0$. However, we are interested in covering cases when $n$ might be of the same order or much larger than $T_0$ when a more parsimonious model is desirable.
Also, note that $\theta_0(\tau)$ might not be unique. This lack of identification might come from the fact that the density of $Y_t$ given $X_t$ is not bounded away from zero at the conditional $\tau$-quantile (which is ruled out by Assumption (ref) below) or a rank deficiency of $\mathbb{E} \sum_{t=1}^{T_0} X_tX_t'$ which we allow here. For each $\tau\in\mathcal{T}$, we denote by $\Theta^0_\tau \subseteq \mathbb{R}_n$ the set of all vectors that fulfill Assumption (ref). It is well known that $\Theta^0_\tau$ is the set of minimizers of the population objective function $\theta\mapsto \mathbb{E}\rho_\tau(Y_t - X_t'\theta)$ for each $\tau\in \mathcal{T}$, where $\rho_\tau(z):=z(\tau - 1\{z<0\})$ is known as the check function. Our model can then be expressed for some compact $\mathcal{T}\subseteq (0,1)$ as
Motivated by the points discussed in the last two paragraphs, we define our estimator $\widehat{\theta}(\tau)$ as a solution to the following optimization problem
where $\lambda>0$ is a common (to all quantiles) penalty (regularizer) parameter to be specified later.
Similar to $\theta_0(\tau)$, $\widehat{\theta}(\tau)$ is not necessarily unique for each $\tau\in\mathcal{T}$ so we let $\widehat{\Theta}_\tau\subset\mathbb{R}^n$ denote the set of all solutions to the optimization problem (ref). In principle, we could set $\lambda=0$ so that the second component of (ref) would vanish if we treat $n$ as fixed or growing at a slower rate as the sample size increases. However, it will be necessary to ensure that $\Theta_\tau^0$ and $ \widehat{\Theta}_\tau$ are close in the appropriate sense for cases when $n>T_0$. In high-dimensional setup $(n>T_0)$, the $\ell_1$-penalization component will induce a sparsity structure on $\widehat{\theta}(\tau)$ which will capture the assumed sparsity structure on $\theta_0(\tau)$. We resume this discussion in Section (ref) after Assumption (ref).
It is easy to verify that using the fact that $\rho_\tau(z) + \rho_\tau(-z) = |z|$ for $z\in\mathbb{R}$, we can re-write (ref) as an unconstrained optimization problem in the $\lambda$-augmented variables. Precisely, (ref) is equivalent to
where $Y_\lambda:= (Y_1^\lambda, \dots, Y_{T_0+2n}^\lambda)' = (Y_1, \dots, Y_{T_0}, 0,\dots ,0)'$ and $X_\lambda:=(X_t^\lambda, \dots, X_{T_0+2n}^\lambda) = (X_1,\dots, X_{T_0}, \lambda T I_n ,-\lambda T I_n)$.
For each $x$, the estimated quantile function $\tau\mapsto \widetilde{Q}(\tau|x):= x' \widehat{\theta}(\tau)$ might not be non-decreasing. Under Assumption (ref), this is a finite sample issue but will prevent us from constructing meaningful finite sample confidence intervals whenever the quantiles “cross”. To circumvent the problem, we use the rearranged conditional quantile estimator proposed in \citet*{chernozhukov2010quantile}, which, in our setup, can be expressed as
Therefore, using only the pre-intervention sample $\{(Y_t,X_t):1\leq t\leq T_0\}$, we estimate model (ref) using
Recall we do not observe $Y_{t}^{(0)}$ for $ T_0<t\leq T$ in our sample. However, we can estimate all its conditional quantiles by
Section (ref) state conditions under which non-asymptotic probabilistic bounds hold for some “distances” between $\mathcal{M}$ to $\widehat{\mathcal{M}}$. In this subsection, we explore the consequence of these results (refer to Section (ref) for a formal statement) as the number of pre-intervention periods diverges ($T_0\rightarrow\infty$) while the number of post-intervention periods ($T_1:= T-T_0)$ is kept fixed. We call it a Partial Asymptotics analysis\footnote{As opposed to the Full Asymptotics analysis, when the entire sample size diverges ($T\rightarrow\infty$) and the ratio $T_0/T$ is kept fixed as $T\rightarrow\infty$}. This approach is tailored to accommodate empirical situations where the number of pre-intervention periods $T_0$ is much larger than $T_1$ (it is not uncommon in empirical applications that we have a single or a couple of post-intervention periods), which justifies the sampling error from the estimation of $\theta_0(\tau)$ by $\widehat{\theta}(\tau)$ to be of smaller order when compared to the inference using the post-intervention periods.
Our first result concerns the uniform approximation of the conditional quantile function in probability. For each post-intervention period ($T_0<t\leq T$), \[\widehat{Q}(\tau|X_t)- {Q}(\tau|X_t)=o_\P(1),\] uniformly in $\tau\in\mathcal{T}$, where $\mathcal{T}$ is any compact subset of $(0,1)$. The next two results are for inference procedures. The second result states that, for any $\alpha<0.5$, the interval \[\mathscr{C}_{t}(\alpha):=[Y_t -\widehat{Q}(1-\alpha/2|X_t), Y_t -\widehat{Q}(\alpha/2|X_t)]\] is an asymptotically (as $T_0\rightarrow\infty$) $1-\alpha$ confidence prediction interval for the treatment effect $\delta_t:=Y_t^{(1)}-Y_t^{(0)}$ for each post intervention period.
Finally, the last result proposes a statistical test for the null hypothesis that the intervention in question had no effect, i.e.,
To that aim, we use for test statistic, $\phi_{\mathcal{T},p}:= \tfrac{1}{T_1}\sum_{t=T_0+1}^T\|\1\{Y_t\leq \widehat{Q}(\tau|X_t)\} -\tau\|_{\mathcal{T},p}$, where $\|\cdot\|_{\mathcal{T},p}$ in the $\mathcal{L}_p$ norm with the respect to the uniform law on $\mathcal{T}$ for any $p\in[1,\infty]$, and show that a test that rejects when \[\{\phi_{\mathcal{T},p}>c_p(\alpha)\}\] has the correct asymptotic (as $T_0\rightarrow\infty$) size $\alpha$, where $c_p(\alpha)$ is a critical value from a pivotal distribution that only depends on $T_1$, $\mathcal{T}$ and $p$. Refer to Section (ref) for further discussion on the test procedure.
The critical value can be obtained by simulation using the empirical quantile of $\phi_{\mathcal{T},p}^0$ with $\1\{Y_t\leq \widehat{Q}(\tau_k|X_t)\}$ replaced by $\1\{ U\leq \tau_k\}$, where $U$ are independent uniform (0,1) random variables. Figure (ref) shown the estimated density for selects $\mathcal{L}_p$ norms and interval $\mathcal{T}\subseteq (0,1)$.
To state the results in a unified manner covering a broad range of tail distributions (polynomial and exponential), we control the tail by its Orlicz norm. Recall that the $\psi$-Orlicz norm of a random variable $X$, which we denote by $\|X\|_\psi$, is define as $\inf\{c>0:\mathbb{E}\psi(X/c)\leq 1\}$ for a convex, even, lower semi-continuous $\psi:\mathbb{R}\rightarrow [0,\infty]$, such that $\psi(0)=0$. In particular, we are concerned with polynomial tails, exponential tails, and bounded variables.
Therefore, we define $\Psi$ the class of functions containing:
Cases (i) and (ii) together deal with $\mathcal{L}_{q,\P}$-random variables whereas case $(iii)$ includes sub-Gaussian ($\mathfrak{p}=2$), sub-Exponential ($\mathfrak{p}=1$) and sub-Weibull ($\mathfrak{p}\in(0,1)$) random variables as special instances.
Assumption (ref)(b) bounds the second moment of $Z_{it}^{(0)}$ away from zero. That does not rule out multicollinearity among the units, but it is important to avoid each unit behaving like a constant as the sample size diverges. As opposed to $B$ in Assumption (ref)(a), $\underline{B}$ is assumed to be a universal constant. While it is possible to allow $\underline{B}$ to slowly converge to zero (as a function of $n$ and $T$), that would add yet another rate term in the final oracle with no extra generality.
Recall we assume the conditional quantile function to be linear in Assumption (ref). The next assumption further imposes structure on conditional quantile function by requiring some level of sparsity (for cases when $n>T_0$) and smoothness of the map $\tau\mapsto \theta_0(\tau)$.
Specifically, the assumption above requires that at most $s_0 \leq n$ elements of $\theta_0(\tau)$ differ from zero in the conditional quantile function. Notice that these non-zeros elements might differ from quantile to quantile. The assumption is similar to the one considered in \citet*{HeWangHong2013} in the context of nonlinear quantile regression models. In a factor model structure with a sparse loading matrix, one expects the non-zeros elements to be the same across all quantiles. We might have a dense model in that $s_0=n$ provided that $n$ grows slower than $T$ (low-dimensional setup). Otherwise, $s_0$ must be strictly smaller than $n$. The smoothness requirement is not strict as we allow the Lipschitz constant to grow polynomial with $T\lor n$. For instance, it is implied whenever $\sup_j|\theta_j(\tau)-\theta_j(\tau')|\leq C|\tau-\tau'|$ for all $\tau,\tau'\in\mathcal{T}$.
Assumption (ref) above is a regularity condition on the conditional distribution of $Y_t|X_t$. Parts (a) and (b) are used to ensure the (restricted) strong convexity of the population objective function in the neighborhood of the quantiles of interest. Similar conditions are needed to ensure the parameter identification even in the fixed $n$ case.
We need to establish an additional piece of notation to state our next assumption clearly and concisely. For $x\in\mathbb{R}^n$ and index set $A\subseteq \{1,\dots,n \}$ we write $x_A$ for the vector $x$ with the entries index by $A^c:=\{1,\dots,n \}\setminus A$ set to zero. For each $\tau\in\mathcal{T}$, let $\mathcal{S}_\tau\subseteq \{1,\dots,n \}$ denote the support of $\theta_0(\tau)$, i.e., $\mathcal{S}_\tau = \{j:\theta_{0,j}(\tau)\neq 0\}$, and define the family of cones $\mathcal{C}:=\{\mathcal{C}_\tau:\tau\in\mathcal{T}\}$ in $\mathbb{R}^n$ as follows \[ \mathcal{C}_\tau:=\{x\in\mathbb{R}^n:\| x_{\mathcal{S}^c}\|_1\leq 3\| x_{\mathcal{S}}\|_1\}. \]
The compatibility condition is standard in LASSO literature; refer to Chapter 6 of \citet*{buhlmann2011statistics} for further details and applications. It is also closely related to the well-known restricted eigenvalue condition on the $\Sigma$. As a matter of fact, the latter is implied by the former and is sufficient for a restricted identification of $\theta_0(\tau)$. Furthermore, Assumption (ref) combined with Assumption (ref)(a) ensure that $\theta_0(\tau)$ as a maximizer of $\theta\mapsto J(\tau):=\mathbb{E}(\rho_\tau(Y_t-Z_t'\theta)$ is unique restricted to the cone $\mathcal{C}_\tau$, uniformly in $\tau\in\mathcal{T}$. Clearly, a sufficient condition for Assumption (ref) would be the minimum eigenvalue of $\Sigma$ to be bounded away for zero by an absolute constant.
In this subsection, we present the two main results of the paper, including two useful corollaries. Recall the serial dependence measure in terms of $\beta$-mixing coefficients defined in (ref). With a slight abuse of notation, define its continuous version by the function $\beta:\mathbb{R}^+\rightarrow[1/4,0]$ given by $\beta(t):= \beta_{\lfloor t\rfloor}$ for $t\geq 0$. Since $\beta(t)$ is a non-increasing function, it admits a left-inverse, which we denote by $\beta^{-1}$. We use the function $\beta^{-1}$ to present the rate dependency of the mixing coefficients of the series $\{Z_t^{(0)}\}$. Similarly the tails of $\{Z_t^{(0)}\}$ are controlled via the $\psi$-Orlicz norm for $\psi\in\Psi$ (refer to Assumption (ref)) and $\psi^{-1}$ denotes the (left) inverse of $\psi$ restricted to $\mathbb{R}_+$.
The first result in Theorem (ref) is related to \citet*{belloni2011} and \citet*{zheng2015} with essential contrasts. First, note that (ref) is an $\ell_1$-convergence rate as opposed to $\ell_2$ in the aforementioned papers. This difference is vital to be able to derive result (ref), which does not follow from the second inequality in Theorem 2 in \citet*{belloni2011}. Second, Theorem 1 accommodates non-independent data via the $\beta$-mixing coefficients with an arbitrary decay and a wide range of distribution tails. Notice that for geometric $\beta$-mixing decay (for instance, most ARMA processes), we pay the price of an extra $\log T_0$ term in the numerator as compared to the independent counterpart. Also, we recover the $\ell_1$-sharp rate $s_0\sqrt{\log n/T_0}$ uniformly in $\tau\in\mathcal{T}$ under independence ($\beta_m =0$ for $m\geq 1$) and absolutely bounded random variables (c.f. Corollary 6.2 in \citet*{buhlmann2011statistics}).
Furthermore, note that our estimator differs from \citet*{belloni2011}, who, in the independence setting, consider a self-normalized process by using a penalty specific for each variable equal to its sample standard deviation. From the practical point of view, this can always be achieved by normalizing the variable before running the LASSO estimator. However, this modification considerably complicates the theoretical analysis of the estimator. That is because, with a self-normalized process under dependency, we need to guarantee the Berbee's Coupling construction (refer to Proposition 2 in \citet*{DMR_1995}) conditional on the $X_{t}$. This explains why in \citet*{belloni2011}, the tail $X_t$ influence appears on the condition that the sample normalization converges uniformly to the population normalization. Refer to the definition of $\gamma$ in Condition D.3 in \citet*{belloni2011}. In most cases, there will be a dependence on the number of covariates implicit in $\gamma$. Our estimator also differs from \citet*{zheng2015}, who consider a $\ell_1$-adaptive penalization in ultra high-dimensional setup. Under independence and compact support assumptions, the authors improve the $\ell_2$-rate appearing in \citet*{belloni2011} from $\sqrt{s_0\log n/T_0}$ to $\sqrt{s_0\log s_0/T_0}$ which allow them to accommodate ultra-high dimensional, i.e., when $\log n$ increases polynomial in $T$.
For our purposes, (ref) is an intermediate step for the final results appearing in Theorem (ref), which gives non-asymptotic bounds for the estimation of the conditional quantile function of the counterfactual uniformly over the quantiles in $\mathcal{T}$. The results are presented both in terms of its “closeness” to the true conditional quantile function, inequality (ref); and the probability of deviation from it when we use the estimated quantile function instead of the true one, inequality (ref). The term $n^{1-\vartheta^2}$ captures the variation in the “empirical process” component, which is controlled by $\lambda$, whereas the terms of the form $\psi^{-1}(\cdot)$ represents the dependency on the tail of $X_t$. For instance, if $X_t$ has $q\geq 4$ finite moments then $\psi^{-1}(x) = x^{1/q}$, whereas in the sub-Gaussissian case, we may take $\psi^{-1}(x)$ to be of order $\sqrt{\log (x)}$.
The main consequence of Theorem (ref) is that we can estimate and draw inference on the intervention effect. For instance, from (ref) we conclude under mild conditions that $\widehat{Q}(\tau|X_t)-Q(\tau|X_t) = O_\P\left(s_0\psi^{-1}(n)r\right)$ uniformly in $\tau\in\mathcal{T}$. In that case, $\widehat{Q}(\tau|X_t)$ is a consistent estimator of the $\tau$ conditional quantile of the the counterfactual uniformly in $\tau\in\mathcal{T}$, provided that $s\psi^{-1}(n)r=o(1)$ as $T_0\rightarrow\infty$. We could for instance take the median $\{\widehat{Q}(0.5|X_t)\}_{t> T_0}$ as the point estimate for the counterfactual $\{Y_t^{(0)}\}_{t>T_0}$. Furthermore, since we have a model for each conditional quantile, we can construct confidence bands of asymptotically correct size for the unobservable $Y_t^{(0)}$ after the intervention. As a concrete example, take the sub-Gaussian case ($\psi(x) = \exp(x^2)-1$) with geometric mixing $\beta_m\lesssim \exp(-cm)$, the last result becomes \[ \sup_{\tau\in\mathcal{T}}|\P(Y_t^{(0)}\leq \widehat{Q}(\tau|X_t)) -\tau| \ = O\left(s_0\sqrt{\frac{\log (n\lor T_0)\log T_0)(\log n)^2}{T_0}} \right). \]
We formalize the discussion in the last paragraph in the following Corollary.
For now, suppose we have only a single post-intervention period $(T_1=1)$ and $Q(\cdot|\cdot)$ is known. Consider a test that rejects when $Y_T< Q(\tau_1|X_t) $ or $Y_T> Q(\tau_2|X_t) $ for some $\tau_1\leq \tau_2\in\mathcal{T}$. The test has the correct size $\alpha$ as soon as we choose $\tau_1,\tau_2$ such that $\tau_2-\tau_1\geq 1 - \alpha$. Clearly, there are infinite as many tests indexed by the pairs $(\tau_1,\tau_2)\in\mathcal{T}^2$. If we take the rejection region to be symmetric around zero, we can show the test is powerful against alternatives of the type $Y_T^{(1)} = Y_T^{(0)} + c$ for some constant $c\in\mathbb{R}$. Consider the case when the treatment is mean preserving but reduces the dispersion; for instance, $Y^{(1)}_T = \sigma Y^{(0)}_T$ for some $\sigma>0$, the test now looses power for $0< \sigma<1$ and it can its power approach size as $\sigma$ gets smaller.
In this situation, it would be better to consider the rejection region centered instead of on the tails. However, we loss power against alternatives when $\sigma>1$ or when the treatment is of the type $Y_T^{(1)} = Y_T^{(0)} + \epsilon$ for some non-degenerate random variable $\epsilon$. We could also consider different rejection regions, such as non-symmetric, non-connected, or any other shape, aiming to gain power against more complex alternatives. There are several (parametric) weak null hypotheses implied by (ref), notably $\mathbb{E} Y_T^{(1)}=\mathbb{E} Y_T^{(0)}$, the classical null hypothesis on ATT estimation in cross-sectional setup.
We propose to test the null (ref) by testing the stability of all the quantiles of the conditional distribution of $Y_T$ given $X_T$ before and after the intervention. To expose the idea, consider the stochastic process $\{M_T^0(\tau):\tau\in \mathcal{T}\}$ where $M^0_T(\tau):=\1\{Y_T\leq Q(\tau|X_T)\}-\tau$ and $\mathcal{T}$ is any compact subset of $(0,1)$. By construction, under the null hypothesis, $\{M_T^0(\tau):\tau\in \mathcal{T}\}$ is a zero-mean process and $\1\{Y_T\leq Q(\tau|X_T)\}$ is distributed as a $\text{Bernoulli} (\tau)$ for each $\tau\in\mathcal{T}$. Thus, $\{M_T^0(\tau):\tau\in \mathcal{T}\}$ can be seen as a process of dependent over $\tau\in\mathcal{T}$ centered Bernoullis.
Obviously, any hypothesis test based on $\{M_T^0(\tau):\tau\in \mathcal{T}\}$ is infeasible. However, apart from the unobservable mapping $\tau\mapsto Q(\tau|\cdot)$, which can be estimated (uniformly) by $\widehat{Q}$, the distribution of $\{M_T^0(\tau):\tau\in \mathcal{T}\}$ is pivotal under the null. To see that, notice that $\{Y_T\leq Q(\tau|X_T)\} = \{U\leq \tau\}$ almost surely under the null, where $U$ denotes a uniformly distributed random variable over the interval $(0,1)$. Then, $\{M_T^0(\tau):\tau\in \mathcal{T}\}$ can be (almost surely) equivalently expressed as $\{\1\{U\leq \tau\}-\tau:\tau\in\mathcal{T}\}$. As a consequence, finite-sample (in terms of post-intervention periods) inference is possible as long as we can compute critical values for $\|M^0_{T}\|_{p,\mathcal{T}}$, where $\|\cdot\|_{p,\mathcal{T}}$ denotes the $\mathcal{L}_p$ norm for $p\in[1,\infty]$ under the uniform measure over $\mathcal{T}$.
For a given $\mathcal{T}$ and $p\in[1,\infty]$, the proposed test-statistic takes the form
The following result bounds the difference in the distribution of $\phi_{{\mathcal{T}},p}$ and $\phi_{{\mathcal{T}},p}^0$ where the latter is the former with $\widehat{Q}$ replaced by $Q$. The result is a “convergence in distribution type” except that the target distribution also depends on the sample size $T$ and number of peers $n$, and, for that reason, it does not necessarily weak-converge in the usual sense.
We propose a test that rejects $\mathcal{H}_0$ whenever $\phi_{{\mathcal{T}},p} > c$ for some constant $c$ and given $\mathcal{T}$ and $p\in[1,\infty]$. To control for the test size $\alpha$, say, we set $c=c_{\mathcal{T},p}(\alpha/T_1)$ where $c_{\mathcal{T},p}(\mathfrak{q})$ denotes the $\mathfrak{q}$-quantile of $\|M_t^0\|_{p,\mathcal{T}}$ under the null. Notice that that the distribution of $M_t^0$ is independent of $t$ under Assumption (ref).
The condition that interval $\mathcal{T}$ must be large enough is to ensure that it includes the $1-\alpha$ quantile of $\phi_{\mathcal{T},p}^0$, which then becomes a continuity point and the result above follows directly from Theorem (ref) and the union bound.
To implement the hypothesis test described above, we need to compute the critical value $c_{\mathcal{T},p}(\cdot)$ or p-values. Fortunately, as previously mentioned, the distribution of $\{M^0_t(\tau):\tau\in\mathcal{T}\}$ is pivotal under the null so that a direct Monte Carlo simulation can be used. For every $u\in[0,1]$ define
Therefore, $\|M_t^0\|_{\mathcal{T},p}$ and $g_{\mathcal{T},p}(U)$ share the same distribution under the null for every $t>T_0$ where $U$ is a uniform random variable on the unit interval. If we take $\mathcal{T}$ to be a closed interval with endpoints $0<\underline{\tau}\leq \overline{\tau}< 1$, since $\tau\mapsto \1\{u\leq\tau\}-\tau$ is piece-wise linear for each $u\in[0,1]$, we can solve the integral analytically for $p\in[1,\infty)$ and the suprema appearing in the case $p=\infty$ is straightforward to compute. Specifically, the last display becomes
where $\tau^*(u):=\underline{\tau}\lor(u\land\overline{\tau})$.
Equation (ref) is key to computing critical values by simply drawing a large number of uniform random variables in a Monte Carlo simulation. Figure (ref) above was constructed using this approach. Furthermore, since $g_{\mathcal{T},p}$ is continuously differentiable except at finitely many points, we conclude that $\phi_{\mathcal{T},p}^0$ is an absolutely continuous random variable for each $p\in[1,\infty]$.
Also, in practice, we cannot compute the test statistics $\phi_{\mathcal{T},p}$ defined by (ref) for an infinite set $\mathcal{T}$. So, we take a large (but finite) $\mathcal{T}$ as an approximation. For instance, $\mathcal{T}$ could be the set of the $n_\mathcal{T} \geq 2$ equally-spaced quantiles from the lowest chosen quantile, $0<{\underline{\tau}}$, to the largest chosen quantile ${\overline{\tau}}<1$. For finite $\mathcal{T}$ we have \[ \phi_{\mathcal{T},p} =
\]
Finally, the choice of the norm $p\in[1,\infty]$ ultimately depends on the expected alternative. Similarly to the non-parametric hypothesis test situation, it is not possible to have the optimum $p$ in the sense of maximizing the test power against all potential alternatives. Recall the alternative in our case is defined by $Y_t^{(0)} \neq Y_t^{(1)}$ in distribution for the post-intervention periods. While it is possible to compute the test power function for a given alternative, i.e., for a fixed pair of laws $(Y_t^{(0)}, Y_t^{(1)})$, in most empirical applications, the expected distribution effect is usually unknown by the econometrician.
Heuristically, we have that when all (or most) of the quantiles of $Y_t$ are expected to change due to the intervention, for instance, when a shift to distribution location occurs, smaller norms such as $p\in\{1,2\}$ are more powerful when compared to $p=\infty$. The opposite is true, i.e., when the change is localized around a single quantile, for instance, only the extreme right tail of the distribution changes, then $p=\infty$ is preferable as opposed to $p\in\{1,2\}$. Both behaviors are intuitive and are corroborated in the simulations. As a rule of thumb, we recommend computing the hypothesis test for $p\in\{1,2,\infty\}$ to guide one's conclusion in empirical investigation. We follow this recommendation in the empirical illustration presented in Section (ref).
We conducted a Monte Carlo study to evaluate the performance of the estimator (ref) and the test (ref) in finite-sample. The simulated DGP is a linear location and scale model given by
where $X_t=(1,X_{2t},\dots, X_{nt})'$ is a $n$-dimensional random vector with covariance structure $\Omega:=(\omega_{ij})$ among the last $n-1$ entries given by $\omega_{ij} = \rho_X^{|i-j|}$ for $\rho_X\in[0,1)$ and $U_t$ is zero-mean unit-variance absolutely continuous random variable independent of $X_t$ with cdf $G$ . Also, $\{(X_t,U_t):t\geq 1\}$ is an autoregressive process with coefficient $\rho_Z\in[0,1)$ for each element. Finally, $\beta_0=(\alpha_0, \widetilde{\beta}'_0)'\in\mathbb{R}^{n}$ is a sparse vector and $x\mapsto \sigma(x)> 0$ the scale function both to be specified later.
The conditional quantile function of (ref) becomes \[Q(\tau|x) = x' \beta_0 + \sigma(x)G^{-1}(\tau).\] Clearly, we have a correctly specified linear quantile model (Assumption (ref)) if the scale function is a constant $\sigma(x)=\sigma^*> 0$ or, more generally when the scale function is linear on $x$, i.e., $\sigma(x)=\sigma^* + x'\gamma$ for some $\gamma\in\mathbb{R}^{n}$ provided that $\gamma'x>0$ for all $x$ in the support of $X_t$. In that case, the the conditional quantile becomes $Q(\tau|x) = x' \theta_0(\tau)$ where $\theta_0(\tau) = \big( \alpha_{0} + \sigma^*G^{-1}(\gamma), (\widetilde{\beta}_0 + G^{-1}(\tau)\gamma)'\big)'$ We simulate (ref) with non-linear scale function to evaluate departures from Assumption (ref). The conditional density of $Y_t$ given $X_t=x$ is given by $f(y|x) = g(y)/\sigma(x)$ where $g$ denotes the density of $G$. Assumption (ref) is fulfilled provided that $\sigma(x)$ is bounded away from zero and infinity over the support of $X_t$, and $g$ is bounded away from zero and infinity over the support of the conditional quantiles for a given compact interval $\mathcal{T}\subset (0,1)$.
The baseline for our simulation is $n = 200$ units ( including the treated one) with $s_0=5$ relevant units for all quantiles, $T_0=100$ pre-intervention periods, $T_1 = 3$ post-intervention periods. All marginal distributions are Gaussian, and the correlation coefficient among the regressors is $\rho_X=0.25$ and the serial correlation coefficient $\rho_Z=0.5$. The linear quantile model is correctly specified. For comparison, we show in the first column of Tables (ref) and (ref) the results for the “Oracle” estimation, i.e., the estimation using unpenalized quantile regression if we knew which are the relevant regressors ($\mathcal{S}_\tau$ is known for each $\tau\in\mathcal{T}$).
We also consider different configurations around a baseline scenario. We investigate the influence of decreasing the number of pre-intervention periods to $T_0 = 80$ and $50$. An increase in the number of peers to $n=300$ and $500$. Change the Signal-to-Noise ratio (SNR) to $\text{SNR}=0.1$ and $10$. Alter the distribution of the $G$ to a Chi-squared and t-distribution with 4 and 10 degrees of freedom, respectively. Finally, we consider a missecifed case where $\sigma(x) = 1 + (\gamma'x)^2$ with $\gamma= (1,1/2,\dots, 1/5)$. See the header of Table (ref) and (ref) for a summary. For all cases, $\mathcal{T}$ is set to be an equally-spaced grid over $(0,1)$ containing 100 points.
In every replication, we compute the estimator defined in (ref) and the test statistics described in (ref) over $\mathcal{T}$. Table (ref) organizes the CQF estimation results. Panel A evaluates the estimator performance in terms of the distance between $\tau\mapsto \widehat{Q}(\tau|X_t)$ and $\tau\mapsto Q(\tau|X_t)$ in the $\mathcal{L}_p$ norm for $p\in\{1,2,\infty\}$ over the uniform distribution in the grid. The result is then averaged over the post-intervention periods. Panel B presents the empirical conditional quantiles, i.e., the conditional quantiles implied by the estimated CQF. Finally, Panel (C) shows the empirical coverage of confidence intervals created as described in Corollary (ref). Table (ref) collects the results of the hypothesis test based on test statistic (ref) for the norm $p\in\{1,2,\infty\}$. It shows the rejection rates under the null (empirical size), computed via the simulated null distribution using $10,000$ replications of (ref).
Overall, the test seems to be correctly sized in all scenarios considered. Interestingly, the size distortions for the Oracle are systematically larger than the ones observed in the Baseline, which might be partially explained by the increase in the estimation error shown in Panel A of Table (ref). Also, the test based on the $\mathcal{L}_\infty$ appears to be consistently slightly undersized, whereas the tests based on $\mathcal{L}_1$ and $\mathcal{L}_2$ are somewhat oversized in most cases. Therefore, all distributional tests seem satisfactory for practical purposes even in the Misspecified model considered (last column of Table (ref)).
To describe the methodology from the practitioner perspective, we revisit the empirical study in \citet*{acemoglu2016value}, henceforth AJKKM. Below, we discuss the basic setup, including a brief discussion of the authors' motivation and essential details about the dataset. For a comprehensive account, refer to AJKKM.
On November 24, 2008, Tim Geithner was nominated by President-elect Obama as Treasury Secretary. The task is to understand the effects of the appointment announcement on the stock returns of financial firms supposedly connected to him. In the authors' words, In a time of crisis the position of Treasury Secretary has unusual powers, including over the financial system, a point that was confirmed by the emergency legislation passed in October 2008. When immediate action is necessary, social connections are likely to become more important as sources of both ideas and human capital for the U.S. executive branch. In complex, stressful situations, it is natural to tap private sector friends, associates, and acquaintances with relevant expertise. The authors find compelling evidence that firms linked to Geithner experienced higher cumulative abnormal returns (on average) than non-connected firms.
The sample consists of 583 firms traded on NYSE or Nasdaq categorized as banks or financial services firms. The data consists of daily stock returns based on closing prices from 2007 and 2008. Out of those firms, 22 were considered treated, i.e., connected to Geithner. The firm's connection to Geithner was labeled based on three criteria: (A) Personal connection, the number of shared board memberships between the firm's executives and Geithner ; (B) Schedule Connections, the number of times that Geithner interacted with executives from each firm while he was president of the New York Fed based on his schedule for each day from January 2007 through January 2009; (C) N.Y. Headquartered, a dummy variable equal to one if a firm's headquarters is identified as New York City since Geithner was president of the Federal Reserve Bank of New York. Those labels are not mutually exclusive.
Even though Geithner's nomination was officially announced on Monday, November 24, news of his impending nomination was leaked to the press late in the trading day on Friday, November 21, 2008, at approximately 3:00 p.m., which coincides with the beginning of a stock market rally. Therefore, we treat the first-day post-treatment ($T_0+1$ is our notation) as Friday, November 21, and refer to it as day $1$. Day $2$ refers to the following Monday, November 24, and so on throughout the subsequent trading days.
We apply the methodology described in Section (ref) independently on each of the 22 connected firms using all the remaining $n=583-22$ firms as potential controls. The pre-intervention data are stock returns for 250 days ending 30 days before the announcement, $T_0=250$. We consider the window of up to 10 days after the announcement, $T_1$. We take $\mathcal{T}$ as an equally-spaced grid on $(0,1)$ with $1000$ points. We present the results for one firm on each connection category with a single label.
The distribution counterfactual for selected firms in each connection type is shown in Figure (ref) together with the actual return for ten days after the announcement, while Figure (ref) plots the median and the 95% confidence interval for announcement effect on returns for each of the ten post-announcement periods for the same firms selected in Figure (ref).
Finally, we test the null hypothesis of no effect over 2, 5, and 10 days after the announcement. The null is tested individually for each of the 22 connected units using the test statistic (ref) for the norms $p\in\{1,2,\infty\}$. The p-value for each test is presented in Table (ref)., and are computed via a Monte Carlo simulation using $10,000$ replications of (ref) for each combination of $p$ and $T_1$.
In roughly half of the cases, the null is rejected at $5\%$ across all the post-announcement windows and all $\mathcal{L}_p$ norms considered. Overall, the rejection is more decisive for the shortest window $(T_1=2)$, which includes only the day of the leak and the official announcement. We have solid rejections for all connection types, which indicates no systematic relation between the kind of connection and the effect on daily return.
Suppose we aggregate the results over all the connected units for each period. In that case, we conclude that the median daily return increase is about $3\%$ at day one, around $1\%$ at day two, and smaller than $0.3\%$ for the subsequent periods in a non-cumulative way. Interestingly, the distribution of the medians of the connected units at each given period is right-skewed, which might explain why AJKKM reports a mean daily return increase of about $6\%$ after a full day of trading, which corresponds to day 1 in our analysis. We also have two other results using the same dataset as a comparison. The first is based on a penalized synthetic control method in \citet*{abadie2021}, and the second is obtained by applying a nonlinear factor model proposed in \citet*{feng2021causal}. The former reports a mean return increase of 6.1% while the former reports an increase between $8$-$9\%$, depending on the chosen parameter for the local PCA, for the first full day of trading.
In summary, based on these four independent analyses, there seems to be compelling evidence to support a positive effect on the daily returns for connected firms after Geithner was nominated Treasury Secretary. The degree of the impact depends on how we measure this effect (mean vs. median), the aggregation level (all units vs. individual units), the type of connection, and the time span after the nomination (day one or cumulative over ten days).
We propose a new methodology to carry out counterfactual analysis to evaluate the impact of interventions on the distribution of the variable of interest. Our approach is especially suited to situations when we observe a single (or just a few) treated unit and several potential controls. The setup is partially motivated by what is encountered in previous work. However, we depart from the standard approach of modeling only the conditional mean and model all the conditional quantiles. The methodology also allows for the number of units to be equal to or much larger than the number of time periods, which is usually the case in event studies or when working with aggregated data.
Therefore, we added to the counterfactual analysis toolkit. Moreover, the proposed methodology's estimation and test procedure can be quickly implemented using standard quantile estimation software and straightforward simulation\footnote{The Monte Carlo simulation results and the empirical application were coded in R based on the quantreg package. The code is available from the author upon request.}, which is the reason we believe in its broad applicability to empirical work. As our empirical illustration demonstrates, virtually every empirical application that falls within the “few treated units” panel data setup can be re-considered under the proposed methodology to gain further insight via the treatment distributional effects.
This supplemental material is organized into two appendices. Appendix (ref) contains the proofs of the main results, Theorem (ref) and (ref) stated in Section (ref) and its supporting lemmas. Appendix (ref) collects additional figures from the empirical exercises described in Section (ref).
\setcounter{equation}{0} \setcounter{section}{0} \setcounter{figure}{0} \setcounter{table}{0} \setcounter{page}{1} \makeatletter