Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
62,032 characters · 12 sections · 91 citation commands
Panel Data Quantile Regression for Treatment Effect Models
\doublespacing
In the literature on program evaluation, it is important to learn about the distributional effects beyond the average effects of the treatment. Policymakers are more likely to prefer a policy that tends to increase outcomes in the lower tail of the outcome distribution to one that tends to increase outcomes in the middle or upper tail of the outcome distribution. Such effects can be captured by comparing the quantiles of the treated and control potential outcomes. The parameter of interest here is the quantile treatment effects (QTE) or the quantile treatment effects on the treated (QTT). For example, abadie2002instrumental estimated the distributional impact of the Job Training Partnership Act (JTPA) program on earnings. They showed that, for women, the JTPA program had the largest proportional impact at low quantiles. However, the training impact for men was largest in the upper half of the distribution, with no significant effect on the lower quantiles. This result could not have been achieved using a mean impact analysis. Empirical researchers have estimated distributional effects, such as the QTE or QTT, in many areas of empirical economic researches. For example, chernozhukov2004effects estimated the QTE of participation in a 401(k) plan on several measures of wealth; james2006mean estimated the QTE of welfare reforms on earnings, transfers, and income; martincus2010beyond estimated the QTE of trade promotion activities; and havnes2015universal and kottelenberg2017targeted estimated the QTT of universal child care.
There is also a rich literature on the identification and estimation of the QTE and QTT in various contexts. firpo2007efficient explored the identification and estimation of the QTE under unconfoundedness. abadie2002bootstrap, chernozhukov2005iv, chernozhukov2006instrumental, and frolich2013unconditional showed how instrumental variables can be used to identify the QTE. athey2006identification, melly2015changes, and callaway2019quantile provided the identification and estimation results for the QTT in a difference-in-differences (DID) setting by using repeated cross-sections or panel data. Further, d2013nonlinear studied the identification of nonseparable models with continuous treatments using repeated cross sections.
In this study, we use use panel data to develop a novel estimation method for the QTE under rank invariance and rank stationarity assumptions. We propose a two-step estimator, based on the quantile regression and minimum distance methods. The rank invariance assumption is used in many nonseparable models, such as those in matzkin2003nonparametric, chernozhukov2005iv, d2015identification, torgovitsky2015identification, feng2020estimation, and ishihara2020identification. This assumption implies that a scalar unobserved factor determines the potential outcomes across treatment status. The rank stationarity assumption implies that the conditional distribution of the unobserved factor, given explanatory variables and covariates, does not change over time. In the literature on nonseparable panel data models, similar assumptions were employed by athey2006identification, hoderlein2012nonparametric, graham2012identification, d2013nonlinear, chernozhukov2013average, chernozhukov2015nonparametric, and ishihara2020identification.
ishihara2020identification also explores the identification of the nonseparable panel data model under the rank invariance and rank stationarity assumptions. In this work, the structural function depends on the time period in an arbitrary way and does not require the existence of "stayers" - individuals with the same regressor values in two time periods. It is important to consider nonlinear time trends when modeling the quantile function. In this case, additive time trends may be restrictive. For example, if the quantile function of $Y_t$ is written as $q_t(\tau) = g(\tau) + \mu_t$, the distribution of $Y_t$ is the same across time, up to the location. However, such an assumption is not valid for many empirical applications. In contrast, the nonseparable panel data model proposed by ishihara2020identification captures nonlinear time effects.
Many nonseparable panel data models require the existence of stayers; this is included in evdokimov2010identification, hoderlein2012nonparametric, and chernozhukov2015nonparametric. In particular, evdokimov2010identification requires the existence of stayers for any value of the treatment variable. However, many empirically important models do not satisfy this assumption. For example, in standard DID models, no individuals are treated during both time periods. The identification approach of ishihara2020identification does not require the existence of stayers and allows the support conditions that are employed in standard DID models.
ishihara2020identification also proposes a parametric estimation based on the minimum distance method. However, when the dimensionality of the covariates is large, the minimum distance estimator is computationally demanding. Hence, when we add many covariates into the model, it is difficult to compute the estimator. To overcome this problem, we propose a two-step estimation method based on the quantile regression and minimum distance methods. Using quantile regression, we can obtain an estimator of the QTE by optimizing the objective function over a low-dimensional parameter. This two-step estimation method is similar to the instrumental variable quantile regression method proposed by chernozhukov2006instrumental.
In a DID setting, our model is similar to the changes-in-changes (CIC) model. athey2006identification suggest the CIC model as an alternative to the DID model. The CIC model allows QTT estimation. Their model is less restrictive than our model because their approach does not require the rank invariance assumption. However, their approach does not work when the treatment is a continuous variable or a discrete variable with many different possible values. There exist many empirical applications in which the treatment variable is continuous, such as when many researchers use panel data to estimate the effect of class size on children's test scores. Our estimation method, contrary to the CIC model, works when the treatment variable is continuous.
d2013nonlinear study the identification of nonseparable models with continuous treatments using repeated cross sections. They allow for nonlinear time effects by assuming that the structural function $g_t(x,u)$ can be written as $m_t(h(x,u))$, where $m_t$ is a monotonic transformation. Their study proposes a nonparametric estimation method of the QTT. However, if we add many covariates into the model, their estimation method does not work because of the curse of dimensionality.
melly2015changes, kottelenberg2017targeted, and sawada2019noncompliance also consider the estimation of the CIC model in the presence of covariates. melly2015changes suggest a flexible semiparametric estimator based on a quantile regression analysis. They estimate the conditional distribution of outcomes for both treatment and control groups and both periods by using quantile regression, and then apply the changes-in-changes transformations. Similar to melly2015changes, sawada2019noncompliance proposes a semiparametric estimator based on distribution regressions. kottelenberg2017targeted rely on the Firpo's (2007) extension to quantiles of the inverse propensity scores method. However, none of them allow for continuous treatments.
An alternative approach estimates the distributional effects using panel data. callaway2019quantile provide identification and estimation results for the QTT under a straightforward extension of the most common DID assumption. To identify the QTT, they employ two key assumptions: the distributional difference-in-differences assumption and the copula stability assumption. The first assumption means that the distribution of the change in potential untreated outcomes does not depend on whether the individual belongs to the treatment or control group. The second assumption means that the copula between the change in the untreated potential outcomes for the treated group and the initial untreated outcome for the treated group is stable over time.
The rest of the paper is organized as follows. Section 2 introduces the model and assumptions and demonstrates that our model is nonparametrically identified. In Section 3, we review the minimum distance estimator, as suggested by ishihara2020identification. We then propose a two-step estimator and show the uniform asymptotic properties of our estimator and the validity of the nonparametric bootstrap. Section 4 contains the results of several Monte Carlo simulations and we illustrate our estimation method in two empirical settings in Section 5. The paper concludes in Section 6. The proofs of the theorems and auxiliary lemmas are provided in the Appendix.
We consider the following potential outcome framework. The potential outcomes are indexed against the potential values $x$ of the treatment variable $X_{it} \in \mathbb{R}^{d_X}$ and denoted by $Y_{it}(x)$. We cannot observe $Y_{it}(x)$ directly and the observed outcome is given by $Y_{it}\equiv Y_{it}(X_t)$. Furthermore, we observe a vector of covariates, $Z_{it}$. We define $\mathbf{Y}_i \equiv (Y_{i1}, \cdots , Y_{iT})'$, $\mathbf{X}_i \equiv (X_{i1}', \cdots , X_{iT}')'$, $\mathbf{Z}_i \equiv (Z_{i1}', \cdots, Z_{iT}')'$, $W_{it} \equiv (Y_{it},X_{it}',Z_{it}')'$ and $\mathbf{W}_i \equiv (W_{i1}',\cdots,W_{iT}')'$. Let $\mathcal{X}_t$, $\mathcal{X}_{1:T}$, and $\mathcal{Z}$ denote the supports of $X_{it}$, $\mathbf{X}_i$, and $\mathbf{Z}_i$.
We assume that the potential outcome can be expressed as
where $q_t(x,z_t,\tau)$ is the conditional $\tau$-th quantile of $Y_{it}(x)$ conditional on $\mathbf{Z}_i=(z_1, \cdots, z_T)'$ and $U_{it}(x)$ is uniformly distributed conditional on $\mathbf{Z}_i$. This implies that the conditional distribution of $Y_{it}(x)$ conditional on $\mathbf{Z}_i$ depends only on $Z_{it}$. When all covariates are time-invariant, this condition does not restrict the conditional distribution and expression ((ref)) is known as the Skorohod representation. Following chernozhukov2005iv, we refer to $U_{it}(x)$ as the rank variable. Additionally, we allow $q_t$ to depend on the time period in an arbitrary manner, similar to the work of ishihara2020identification.
First, we impose the rank invariance assumption.
Assumption 1 (i) is referred to as the rank invariance assumption. For example, matzkin2003nonparametric, chernozhukov2005iv, d2015identification, torgovitsky2015identification, feng2020estimation, and ishihara2020identification also employ similar assumptions. This model is restrictive because the potential outcomes $\{Y_{it}(x)\}_{x \in \mathcal{X}_t}$ are not truly multivariate and, are jointly degenerate. As discussed in chernozhukov2005iv, we can relax the rank invariance assumption to the rank similarity assumption. That is, $U_{it}(x)|\mathbf{X}_i,\mathbf{Z}_i \overset{d}{=} U_{it}(\tilde{x})|\mathbf{X}_i,\mathbf{Z}_i$ for all $x$ and $\tilde{x}$.
Under the rank invariance assumption, the observed outcome can be written as
This is the nonseparable model with a scalar unobserved variable and model ((ref)) is the same as the model proposed by ishihara2020identification when there are no covariates. If $U_{it}$ is independent of $X_{it}$ and $Z_{it}$, then this model is identical with the usual quantile regression model. However, our model allows for correlation between $U_{it}$ and the treatment variable. Hence, to achieve point identification, we require additional assumptions.
Next, we impose the rank stationarity assumption.
Assumption 2 implies that the rank variable is stationary across the time period. In the literature on nonseparable panel data models, similar assumptions were employed by athey2006identification, hoderlein2012nonparametric, graham2012identification, d2013nonlinear, chernozhukov2013average, chernozhukov2015nonparametric, and ishihara2020identification. chernozhukov2013average referred to Assumption 2 as “time is randomly assigned” or “time is an instrument.”
Assumption 2 can be viewed as a quantile version of the identification condition of the following conventional linear panel data model: $$ Y_{it} = X_{it}'\alpha + A_i + \epsilon_{it}, \ \ \ \ E[X_{is} \epsilon_{it}]=0 \ \text{for all $t$ and $s$,} $$ where $A_i$ is a fixed effect and $\epsilon_{it}$ is a time-variant unobserved variable. Let $\bar{E}[\cdot|\mathbf{X}_i]$ denote the linear projection on $\mathbf{X}_i$, as in chamberlain1982multivariate. chernozhukov2013average show that the above equation is satisfied if and only if there is $\tilde{\epsilon}_{it}$ with $$ Y_{it} = X_{it}'\alpha + \tilde{\epsilon}_{it}, \ \bar{E}[\tilde{\epsilon}_{it}|\mathbf{X}_i]=\bar{E}[\tilde{\epsilon}_{is}|\mathbf{X}_i] \ \text{for all $t$ and $s$.} $$ In contrast, if the conditional quantile function is linear in $X_{it}$ and there are no covariates, then we can rewrite model ((ref)) as
where $\epsilon_{it}(\tau) = X_{it}'(\alpha(U_{it})-\alpha(\tau))$. Then, under Assumption 2, $\epsilon_{it}(\tau)$ satisfies $F_{\epsilon_t(\tau)| \mathbf{X}}(0|\mathbf{x}) = F_{\epsilon_s(\tau)| \mathbf{X}}(0|\mathbf{x})$ for all $t\neq s$ and $\mathbf{x}$. Hence, the rank stationarity assumption can be viewed as a quantile version of the identification condition of the conventional linear panel data model.
Under Assumptions 1 and 2 and additional assumptions in Appendix 1, we can show that $q_t(x,z_t,\tau)$ is nonparametrically identified. The following proposition is essentially the same as Corollary 1 in ishihara2020identification.
From the proof of Proposition 1, for any $t \neq s$, we have $$ F_{Y_t|\mathbf{X},\mathbf{Z}}\left( q_t(x_t,z_t,\tau) | \mathbf{x},\mathbf{z} \right) = F_{Y_s|\mathbf{X},\mathbf{Z}}\left( q_s(x_s,z_s,\tau) | \mathbf{x},\mathbf{z} \right), $$ where $\mathbf{x} = (x_1, \cdots , x_T)'$ and $\mathbf{z} = (z_1, \cdots , z_T)'$. ishihara2020identification demonstrates that this condition provides point identification when the support of $\mathbf{X}_i$ satisfies Assumption A.2. In Section 3, we propose an estimation method based on this condition.
To illustrate our model, we consider the following two examples:
In this section, we consider the estimation method of the QTE. First, in Section 3.1, we review the minimum distance method proposed by ishihara2020identification and show that the minimum distance estimator does not work when there are many covariates. Second, in Section 3.2, we propose a two-step estimator based on the quantile regression and minimum distance methods and show that our estimator is computationally convenient. Finally, in Sections 3.3 and 3.4, we demonstrate the consistency and uniform asymptotic normality of our estimator.
In this section, for simplicity, we assume that $T=2$ and there are no covariates. ishihara2020identification considers the following parametric model: \[ q_t(x_t,\tau) = g_t(x_t,\tau ; \theta_0). \] The structural functions are parameterized by $\theta \in \Theta \subset \mathbb{R}^{d_{\theta}}$, where $\theta_0 \in \Theta$ is the true parameter. Then, from Assumption 2, we obtain
Thus, ishihara2020identification proposes a minimum distance estimator based on ((ref)).
Let $\| \cdot \|_{\mu}$ denote the $L_2$-norm with respect to a probability measure $\mu$ with support $[0,1] \times \mathcal{V}$. The minimum distance estimator $\hat{\theta}$ is then obtained from the following optimization:
where $\omega(\mathbf{x},v)$ is a weight function. Since $\hat{D}_{\theta}(\tau,v)$ is not continuous in $\theta$, the minimum distance estimator requires minimizing the discontinuous objective function over $\theta \in \Theta$. If the dimension of $\theta$ is large, the optimization ((ref)) is computationally demanding. Therefore, adding many covariates into the model makes it difficult to compute the minimum distance estimator.
For estimation, we focus on the following linear-in-parameter model:
The observed outcome is then written as $$ Y_{it} = X_{it}'\alpha(U_{it})+Z_{it}'\beta_t(U_{it}), \ \ \ U_{it}|\mathbf{Z}_i \sim U(0,1), $$ where $Z_{it}$ contains a constant term. Hereafter, we set $\mathbf{Z}_i \in \mathbb{R}^{d_z}$ as a vector of all the variables of $Z_{i1}, \cdots, Z_{iT}$. For example, if all covariates are time-invariant, we have $Z_{i1} = \cdots = Z_{iT} = \mathbf{Z}_i$. We assume that $\mathcal{X}_{t}$ and $\mathcal{Z}$ are bounded. In this model, we have $\partial q_t(x,z,\tau) / \partial x = \alpha(\tau)$; hence, our target parameter is $\alpha(\tau)$. Because $\beta_t(\tau)$ depends on the time period, this model captures nonlinear time effects. This model is similar to the IV quantile regression model proposed by chernozhukov2006instrumental.
Using Proposition 1, we can identify $\alpha(\tau)$ and $\beta_t(\tau)$ using the following conditions:
where $\mathbf{x} \equiv (x_1, \cdots ,x_T)'$ and $\mathbf{z} \equiv (z_1, \cdots ,z_T)'$, respectively. Similar to ((ref)), we can construct a minimum distance estimator using ((ref)) and ((ref)). However, if the dimensionality of covariates is high, the minimum distance approach cannot be directly applied because the minimum distance estimator is computationally demanding.
We propose the following two-step estimator based on the quantile regression and minimum distance methods. Fix $\tau \in (0,1)$. In the first step, we define $\tilde{\beta}_t(a,\tau)$ as
where $\rho_{\tau}(u) \equiv (\tau - \mathbf{1}\{u<0\})u$, $\mathcal{B}_t$ is the parameter space of $\beta_t(\tau)$, and $R_{\tau}(W_{it};a,b_t) \equiv \rho_{\tau}\left( Y_{it} - X_{it}'a - Z_{it}'b_t \right)$. This is an ordinary quantile regression of $Y_{it}-X_{it}'a$ on $Z_{it}$. Then, from ((ref)), $\tilde{\beta}_t \left( \alpha(\tau),\tau \right)$ becomes a consistent estimator of $\beta_t(\tau)$.
In the second step, we construct an estimator of $\alpha(\tau)$ using the minimum distance approach. We define
where $b = (b_1', \cdots , b_T')'$, $v = (v_{\mathbf{x}}', v_{\mathbf{z}}')'$, and $\tilde{\mathbf{X}}_i$ and $\tilde{\mathbf{Z}}_{i}$ are standardized versions of $\mathbf{X}_i$ and $\mathbf{Z}_i$, where each component has a mean of 0 and a standard deviation of 1. It follows from ((ref)) that we have
where $\beta(\tau) \equiv (\beta_1(\tau)', \cdots, \beta_T(\tau)')'$. As shown in stinchcombe1998consistent, if ((ref)) holds for all $t$ and $v \in \mathcal{V} \equiv [-0.5,0.5]^{d_{X}\cdot T + d_{Z}}$, the conditional moment condition ((ref)) is satisfied. Let $\|\cdot\|_{L_2}$ be the $L_2$-norm over a compact set $\mathcal{V}$; that is, $\|f(v)\|_{L_2}^2 = \int_{\mathcal{V}} f(v)^2 dv$. Using this norm, we obtain the following estimator of $\alpha(\tau)$:
where $\tilde{\beta}(a,\tau) \equiv (\tilde{\beta}_1(a,\tau)', \cdots , \tilde{\beta}_T(a,\tau))'$, $\hat{D}^t_n(v;a,b)\equiv \frac{1}{n} \sum_{i=1}^n g_t(\mathbf{W}_i;a,b,v)$, and $\mathcal{A}$ is the parameter space of $\alpha(\tau)$. Finally, we estimate $\beta_t(\tau)$ by $\hat{\beta}_t(\tau) \equiv \tilde{\beta}_t(\hat{\alpha}_t(\tau),\tau)$.
We briefly explain our two-step estimation method. As discussed above, because the covariates are independent of $U_{it}$, $\tilde{\beta}_t \left( \alpha(\tau),\tau \right)$ becomes a consistent estimator of $\beta_t(\tau)$. Using this result, under regularity conditions, we obtain \[ \hat{D}_n^t \left( v ; \alpha(\tau), \tilde{\beta}(\alpha(\tau),\tau) \right) \rightarrow_p E[g_t(\mathbf{W}_i; \alpha(\tau), \beta(\tau),v)] = 0, \] which implies that the objective function of ((ref)) converges to zero for $a = \alpha(\tau)$. Hence, we expect to obtain a consistent estimator of $\alpha(\tau)$ by minimizing ((ref)).
In practice, we can implement this estimation procedure as follows:
Our estimator is similar to that proposed by chernozhukov2006instrumental. They consider the IV quantile regression for heterogeneous treatment effect models and simultaneous equation models with nonadditive errors. Similarly, our estimator is attractive from a computational point of view. As ordinary quantile regressions are obtained by convex optimization, our first step estimation ((ref)) is computationally convenient. Our second step estimation ((ref)) requires non-convex optimization; hence, it seems to be computationally demanding. However, we can obtain ((ref)) by optimizing the objective function over the $\alpha$ parameter (typically one-dimensional). This fact makes our estimator computationally convenient.
In this section, we show that $\alpha(\tau)$ and $\beta(\tau) = (\beta_1(\tau)', \cdots , \beta_T(\tau)')'$ uniquely solve the limit problems. We define
and
where $\beta(a,\tau) \equiv (\beta_1(a,\tau)', \cdots , \beta_T(a,\tau)')'$ and $D^t(v;a,b) \equiv E[g_t(\mathbf{W}_{i};a,b,v)]$. Hence, to prove consistency, we need to show that $\alpha^*(\tau)$ is unique and $\alpha^*(\tau) = \alpha(\tau)$.
We define $e_t(a,\tau,\mathbf{z}) \equiv P ( Y_{it} \leq X_{it}'a + Z_{it}'\beta_t(a,\tau) | \mathbf{Z}_i=\mathbf{z} )$ and impose the following assumptions.
When the support of $(X_{i1},X_{i2})$ is $\{(0,0),(0,1)\}$, we have $E[X_{i1}^2] = 0$ but $E[X_{i2}^2]$ is positive. Hence, Assumption 3 (i) holds in standard DID settings. Assumption 4 is a technical condition that is satisfied in many situations. Using the proof of Theorem 2 in angrist2006quantile, it follows from the first-order condition of ((ref)) that we have $E\left[ \left( \mathbf{1}\{Y_{it} \leq X_{it}'a + Z_{it}'\beta_t(a,\tau)\} - \tau \right) Z_{it} \right]=0$, which implies that $E\left[ \left( e_t(a,\tau,\mathbf{Z}_i) - \tau \right) Z_{it} \right]=0$. Hence, we have $E[e_t(a,\tau,\mathbf{Z}_{i})]=\tau$ because $Z_{it}$ contains a constant. When $\mathbf{Z}_i$ has continuous covariates and $e_t(a,\tau,\mathbf{z})$ is continuous in $\mathbf{z}$, $e_t(a,\tau,\mathbf{z}) = \tau$ holds for some $\mathbf{z} \in \mathcal{Z}$. Even when all covariates are discrete, if $Z_{it}$ is time invariant and the model is saturated, that is, the cardinality of $\mathcal{Z}$ is equal to the dimension of $\beta_t(a,\tau)$, then we have $e_t(a,\tau,\mathbf{z})= \tau$ for all $\mathbf{z} \in \mathcal{Z}$.
Theorem 1 implies that $\alpha(\tau)$ minimizes $\frac{1}{T} \sum_{t=1}^T \left\|D^t(v;a,\beta(a,\tau)) \right\|_{L_2}^2$. Hence, if the objective function of ((ref)) converges to $\frac{1}{T} \sum_{t=1}^T \left\|D^t(v;a,\beta(a,\tau)) \right\|_{L_2}^2$ uniformly, then we obtain the consistency of $\hat{\alpha}(\tau)$.
In this section, we show the uniform asymptotic normality of our estimator and prove the validity of the nonparametric bootstrap. Our asymptotic result also implies that our estimator is consistent.
Let $\mathcal{T}$ be a closed subset of $[\epsilon, 1- \epsilon]$ for $\epsilon > 0$. In addition, we define $J_t^b(a,\tau) \equiv E\left[ f_{Y_t-X_t'a|Z_t}(Z_{it}'\beta_t(a,\tau)|Z_{it}) Z_{it} Z_{it}' \right]$ and $J_t^b(\tau) \equiv J_t^b(\alpha(\tau),\tau)$. The following assumption is sufficient for the consistency of $\hat{\alpha}(\tau)$ and $\hat{\beta}_t(\tau)$.
Condition (iii) imposes that $X_{it}$ and $Z_{it}$ are bounded. If $X_{it}$ and $Z_{it}$ are unbounded, then for $q_t(x_t,z_t,\tau)$ to be monotonically increasing in $\tau$, $\alpha(\tau) = \alpha$ and $\beta_t(\tau) = \beta_t$ must hold. Hence, we assume the boundedness of $\mathcal{X}_{1:T}$ and $\mathcal{Z}$. Condition (v) means that there exists a continuous density $f_{Y_t-X_t'a|Z_t}(y|z_t)$ for all $a \in \mathcal{A}$. Because we have
we obtain $f_{Y_t-X_t'a|Z_t}(y|z_t) = \int f_{Y_t|X_t,Z_t}(y + x_t'a |x_t, z_t) dF_{X_t}(x_t)$ if $f_{Y_t|X_t,Z_t}(y|x_t,z_t)$ is bounded. Hence, condition (v) holds if $f_{Y_t|X_t,Z_t}(y|x_t,z_t)$ is bounded and continuous in $y$. In addition, this implies that $J_t^b(a,\tau)$ is continuous in $a$ if $\beta_t(a,\tau)$ is continuous in $a$.
We define
$\Gamma_1^t(v;\tau) \equiv \Gamma_1^t(v;\alpha(\tau),\tau)$, and $\Gamma_2^t(v;\tau) \equiv \Gamma_2^t(v;\alpha(\tau),\beta(\tau))$. Then, the following assumption is required to derive the asymptotic distribution of the estimator.
To derive the asymptotic distribution of $\hat{\alpha}(\tau)$, we need to show that $\sqrt{n} (\tilde{\beta}_t(a,\tau) - \beta_t(a, \tau) )$ converges in distribution uniformly in $a \in \mathcal{A}$. We use condition (ii) to show this result. Similarly, we need condition (iv) to show a uniform approximation of $\hat{\alpha}(\cdot)$. Condition (v) indicates that the rank condition holds uniformly in $\tau \in \mathcal{T}$.
The proof of this theorem is based on arguments similar to those in brown2002weighted, chen2003estimation, and torgovitsky2017minimum.
The following corollary follows immediately from Theorem 2.
We consider the case in which $T=2$, $X_{it}$ is scalar, and there are no covariates. In this case, we have
where $Y_{it}(\tau) \equiv \alpha(\tau) X_{it} + \beta_t(\tau)$,
and $\gamma_2^t(v;\tau) \equiv E\left[ f_{Y_t|\mathbf{X}}(Y_{it}(\tau)|\mathbf{X}_i)\omega(\mathbf{X}_i,v) \right]$. Because $J_t^b(\tau)$, $\gamma_2^1(v;\tau)$, and $\gamma_2^2(v;\tau)$ are positive, the variances of $\xi(\mathbf{W}_i;\tau)$ and $\Delta_{12}(\tau) l(\mathbf{W}_i;\tau)$ become small when $U_{i1}$ and $U_{i2}$ are positively correlated. Specifically, if $U_{i1}=U_{i2}$, then $\xi(\mathbf{W}_i;\tau)$ is exactly equal to zero.
\if0
\fi
Let $\{\mathbf{W}_{i}^*\}_{i=1}^n$ denote a bootstrap sample drawn with replacement from $\{\mathbf{W}_i\}_{i=1}^n$. That is, $\{\mathbf{W}_{i}^*\}_{i=1}^n$ are independently and identically distributed from the empirical measure, conditional on the realizations $\{\mathbf{W}_i\}_{i=1}^n$. We define $\hat{\alpha}^*(\tau)$ as the bootstrap counterpart to $\hat{\alpha}(\tau)$. Then, we can obtain the following theorem.
Simulation 1.\, Suppose that the potential outcomes are given by
where $Z_{i} \sim U(0,1)$ and $\Phi$ is the standard normal distribution function. The observed outcomes are generated from $Y_{it}=Y_{it}(X_{it})$. We assume that $X_{it} = \Phi(\tilde{X}_{it})$, $U_{it} = \Phi(A_i + \tilde{U}_{it})$, $(\tilde{X}_{i1},\tilde{X}_{i2},A_i)' \sim N(0,\Sigma_{\mathbf{X} A})$, and $\tilde{U}_{it} \sim N(0,1-\rho^2)$, where $\rho \in [0,1]$ and
Then, $U_{it}$ is uniformly distributed and $\rho$ represents the dependence between $U_{i1}$ and $U_{i2}$. When $\rho = 0$, $U_{i1}$ and $U_{i2}$ are uncorrelated and when $\rho = 1$, $U_{i1}$ and $U_{i2}$ are perfectly correlated. Here, we have $\alpha(0.25) = 0.66$, $\alpha(0.5) = 1$, and $\alpha(0.75) = 1.34$.
Table 1 contains the results of this experiment for two different choices of the sample size, $1000$ and $2000$, and three different choices of $\rho^2$, $0.1$, $0.5$, and $0.9$. The number of replications is set at $1000$ throughout. Table 1 shows the bias, standard deviation, and MSE of the estimates of $\alpha(\tau)$ for $\tau = 0.25, 0.5$, and $0.75$. For all settings, the bias is quite small. Table 1 shows that the standard deviation and MSE decrease in all experiments as the sample size increases. As expected, when the correlation between $U_{i1}$ and $U_{i2}$ is high (i.e. $\rho^2 = 0.9$), the standard deviation decreases.
We also verify that the nonparametric bootstrap procedure works for $(n,\rho^2) = (2000,0.9)$. We calculate 90% and 95% confidence intervals of $\alpha(\tau)$ to obtain the coverage probabilities for $\tau = 0.25, 0.5$, and $0.75$. Table 2 shows the nominal and actual coverage probabilities are close in all settings.
Simulation 2.\, To compare our estimation method with that of athey2006identification, we consider the following model. We assume that $\mathcal{X}_{1:2} = \{(0,0),(0,1)\}$ and the potential outcomes are given by
The observed outcomes are generated from $Y_{i1} = Y_{i1}(0)$ and $Y_{i2}=Y_{i2}(X_{i2})$. It is assumed that $X_{i2}= \mathbf{1}\{\tilde{X}_i+A_i \geq 0\}$, $U_{it} = \Phi(A_i + \tilde{U}_{it})$, $\tilde{X}_i \sim N(0,1)$, $A_i \sim N(0,\rho^2)$, and $\tilde{U}_{it} \sim N(0,1-\rho^2)$, where $\rho \in [0,1]$. Then, $G_i \equiv \mathbf{1}\{X_{i2}=1\}$ denotes an indicator for the treatment group.
Since $U_{it}$ satisfies the rank stationarity assumption, we have \[ E[Y_{i2}(0)-Y_{i1}(0)|G_i = g] \ = \ - 0.5 E\left[ \Phi(U_{i2}) | G_i =g \right]. \] The conditional distribution of $U_{i2}|G_i=0$ is different from that of $U_{i2}|G_i=1$; therefore, this model does not satisfy the parallel trend assumption employed in standard DID models and we cannot estimate the average treatment effect on the treated (ATT) using the standard DID estimation method. By contrast, using our estimation method, we can estimate the quantile functions of the potential outcomes and obtain an estimate of the ATT.
From ((ref)) and ((ref)), we can estimate $F_{Y_2(0)|G=1}(y)$ and $F_{Y_2(1)|G=0}(y)$ by
where $\hat{F}_{Y_t|G=g}(\cdot)$ and $\hat{F}_{Y_t|G=g}^{-1}(\cdot)$ are the empirical distribution and quantile functions, respectively. The marginal distributions of the potential outcomes and QTE can be obtained using these estimators and the empirical distributions of $Y_{i2}|G_i=0$ and $Y_{i2}|G_i=1$. We refer to this estimator as AI estimator.
Table 3 presents the bias, standard deviation, and MSE of our estimator and the AI estimator for $n = 500$ and three different choices of $\rho^2$, $0.1$, $0.5$, and $0.9$. For all settings, the results of our estimator are similar to those of the AI estimator. Hence, when there are no covariates, our estimator is not worse than the AI estimator.
In this section, we use our method to study the impact of an agricultural insurance program on household production. We use the data employed by cai2016impact to estimate the QTE of insurance provision on tobacco production.
This empirical analysis is based on data obtained from 12 tobacco production counties in the Jiangxi province of China. Across these 12 counties, only tobacco farmers in the county of Guangchang were eligible to buy the tobacco insurance policy. In 2003, the People's Insurance Company of China (PICC) designed and offered the first tobacco production insurance program to households in Guangchang. Hence, we use this county as a treatment group.
The sample includes information on approximately 3,400 tobacco households during 2002 and 2003. Table 4 provides summary statistics for 2002 and shows that treatment regions are quite different from control regions in terms of their observed characteristics. For example, control regions include more educated people than treatment regions. The proportion of high school- or college-educated people in the treatment regions is $0.025$, whereas that in the control regions is $0.257$. Hence, controlling the observed characteristics is important for adjusting the differences between the treatment and control regions.
We estimate the following linear-in-parameter model:
where $Y_{it}$ is the tobacco production area (mu), $G_i$ is a treatment indicator equal to one for the treatment regions and zero for the control regions, and $Z_i$ is a vector of covariates including a constant term. We estimate $\alpha(\tau)$ for $\tau = 0.1, ... , 0.9$. Following cai2016impact, we employ the age of the household head, household size, and education level indicators as control variables.
The main results from our method are presented in Figure 1. The DID estimate is $0.239$, and the $95$ % confidence interval is $[0.078,0.388]$. We use the nonparametric bootstrap method to construct this confidence interval. Figure 1 shows that the estimates of $\alpha(\tau)$ differ across $\tau$, and the QTE increases in $\tau$. The impact of the insurance provision is nearly zero at the lower and middle quantiles and positive at the upper quantiles. For $\tau = 0.8$ and $0.9$, the QTE is statistically significant. In addition, we consider the null hypothesis that the QTEs are constant along $\tau$. We conduct the test described in Remark 3 and calculate the test statistic and the critical value at the $0.05$ significance level. These values are $571.9$ and $296.3$, respectively; hence, the null hypothesis is rejected.
cai2016impact analyzes the welfare impact of the insurance program through the calibration. The parameter values of the production function are chosen to match the DID (or triple difference) estimate. From this analysis, she concludes that providing a heavily subsidized compulsory insurance program has a positive welfare impact on rural households. However, our results show that the insurance program does not significantly change households' investment behavior at the lower and middle quantiles, and hence, may not affect household welfare at such quantiles.
Next, we use our method to study the effect of TV on child cognitive development. We use the data employed by huang2010dynamic to estimate the QTE of TV watching on children's cognitive development.
This empirical analysis is based on a childhood longitudinal sample from NLSY79 (National Longitudinal Survey of Youth 1979). Following huang2010dynamic, we use a longitudinal sample of approximately 2,400 children and treat the Peabody Individual Achievement Test (PIAT) reading scores at ages 6--7 and 8--9 as $Y_{i1}$ and $Y_{i2}$, respectively. The PIAT reading score at ages 6--7 has mean 103.0 and SD 11.7, and that at ages 8--9 has mean 104.3 and SD 14.6. The outcome distribution at ages 8--9 is more dispersed than that at ages 6--7. Hence, in these cases, additive time trends may not be plausible. We use daily TV watching hours at ages 6--7 and 8--9 as the treatment variables $X_{i1}$ and $X_{i2}$, respectively and estimate the following linear-in-parameter model:
where $Z_{it}$ is a vector of covariates including a constant term, dummy variables of race and gender, and an indicator of whether a child has 10 or more children's books at home. In addition, we employ the Home Observation Measurement of the Environment variable (HOME), where is often used in child development research as an aggregate quality indicator of the home environment.
The main results from using our method are presented in Figure 2. We find that the estimates of $\alpha(\tau)$ differ slightly across $\tau$ and the QTE is nearly zero at the lower and middle quantiles. However, the impact of TV watching on the PIAT reading score is statistically significant at $\tau = 0.8$. We consider the null hypothesis that the QTEs are zero at all quantiles. We conduct the test described in Remark 3 and calculate the test statistic and the critical value at the $0.05$ significance level. These values are $145.7$ and $290.5$, respectively; hence, the null hypothesis is not rejected. Similar to huang2010dynamic, the magnitude of the effect is quite small compared to the standard deviation of the PIAT reading score. Therefore, the effect of TV on child cognitive development is neither statistically nor economically significant.
In this study, we developed a novel estimation method for the QTE under rank invariance and rank stationarity assumptions. Although ishihara2020identification also explores the identification and estimation of the nonseparable panel data model under these assumptions, the minimum distance estimation using this process is computationally demanding when the dimensionality of covariates is large. To overcome this problem, we proposed a two-step estimation method based on the quantile regression and minimum distance methods. We then showed the uniform asymptotic properties of our estimator and the validity of the nonparametric bootstrap. The Monte Carlo studies indicated that our estimator performs well in finite samples. Finally, we presented two empirical illustrations to estimate the distributional effects of insurance provision on household production, and TV watching on child cognitive development.
\setcounter{equation}{0}