Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
82,377 characters · 12 sections · 56 citation commands
Estimation of the Local Conditional Tail Average Treatment Effect
\onehalfspacing
\doublespacing
The treatment effect of a policy change is often heterogeneous among individuals. Accounting for such heterogeneity is crucial for evaluating the effects of and understanding the mechanisms underlying the policy HSC_1997. To capture the heterogeneity of treatment effects, the quantile treatment effect (QTE) is a frequently used measure, which is computed as the difference between a given quantile of the distribution of the potential outcome subject to the policy being changed and that of the potential outcome under the policy had it not been changed. In this paper, we consider an alternative measure for evaluating heterogeneous treatment effects: the conditional tail average treatment effect (CTATE), defined as a difference between the conditional tail expectations (CTEs) of two potential outcomes at a given quantile level. The CTATE amounts to a (rescaled) integral of the QTE over a specified range of quantiles and is thus useful for delivering aggregated local information on treatment effect.
Recently, FZ_2016 proposed a class of consistent loss functions (henceforth the FZ loss) for jointly estimating the quantile and CTE of a random variable. Through the FZ loss, we develop a semiparametric estimation procedure to jointly estimate the CTATE and QTE for the group of compliers under the two-sided noncompliance framework IR_2015, which is commonly employed in addressing endogeneity with instrumental variables and in developing various estimators for the local treatment effects in the causal inference literature IA_1994, AIR_1996, AAI_2002, Abadie_2003, FM_2013, DHL_2014, BCFH_2017, FH_2017, CHW_2020, FFHL_2020, WPZF_2021, HLL_2022. Under this framework, if some regularity conditions hold AAI_2002, Abadie_2003, endogeneity can be eliminated within samples of compliers. Using this fact, we demonstrate that both the QTE and CTATE locally for compliers can be estimated through a weighted FZ loss minimization-based estimation approach. We then establish the asymptotic theory for the resulting local CTATE (LCTATE) estimator and provide an efficient and stable algorithm for its implementation.
We now further remarks on the usefulness of the CTATE in empirical policy evaluation research. The conditional tail expectation (CTE) is related to the notion of second order stochastic dominance (SOSD). Stochastic dominance is a uniform order relation between the distributions of stochastic outcomes and is often used in ranking individuals' preferences under uncertainty. Conventionally, the formulations of first and second order stochastic dominances (FOSD and SOSD) are based on the properties of the cumulative distributions of competing random outcomes. Such formulations can also be equivalently established using their quantiles and CTEs Levy_2016. For policy evaluation, if the QTE is nonnegative over all and positive over some quantile levels, the policy being changed first order stochastically dominates that that had not been changed. Likewise, if the CTATE is nonnegative over all and positive over some quantile levels, we can then attest that the policy being changed second order stochastically dominates that that had not been changed. Therefore we can use the CTATE to rank counterfactual outcome distributions in the SOSD sense when ranking by the FOSD criterion is empirically inconclusive. In addition, the CTATE has recently attracted growing interest in the study of optimal treatment assignment policies. LSW_2023 show the importance of the CTATE in constructing optimal regret policies in the worst-case welfare when sample selection is biased. Their study demonstrates the usefulness of the CTATE for robust policy learning.
The CTATEs at two quantile levels can be used to calculate an average of the QTEs between the two quantile levels, which we term the inter-quantile average treatment effect (IQATE). The IQATE is also useful for summarizing heterogeneous treatment effects to reveal an overall picture especially when the QTEs exhibit a large fluctuation over a specified range of quantiles of interest. Moreover, related to the distributional comparison of economic outcomes, the CTE is directly connected to the Lorenz curve for comparing income inequalities. Hence our proposed procedure for estimating the CTATE may be used to quantify the Lorenz effect, which captures the shift in the Lorenz curve due to changes in policy regimes CFM_2013.
In this paper, we estimate the CTEs of potential outcomes by minimizing the FZ loss FZ_2016. To the best of our knowledge, the FZ loss is the only known class of consistent loss functions for estimating the CTE of a random variable. Using consistent loss to estimate parameters of statistical functionals results in an M-estimation problem, which facilitates computation and statistical analysis of parameter estimators through empirical risk minimization. The FZ loss is non-smooth in model parameters, hence rendering the estimation problem computationally challenging. To address this issue, we propose an iterative scheme that decomposes the estimation problem into a smooth optimization and a weighted quantile regression estimation problem. This then yields a computation algorithm that can be efficiently and stably implemented in practice. We apply our proposed method to analyze the effects of the Job Training Partnership Act (JTPA) program participation. Our empirical findings indicate that, when endogeneity is considered, the estimated LCTATE on the earnings for adult women is nonnegative over all and statistically significantly positive over some quantile levels, but that for adult men is not. Our results also reveal that for both adult men and women, evidence for the positivity of local QTEs on earnings is much weaker than that of LCTATEs over all quantile levels. For adult women, conditional on the group of compliers, there hence appears to be second order stochastic dominance of the distribution of potential earnings from participating in the JTPA programs over that from not participating. However, the SOSD results do not seem to hold for the case of adult men. As a risk-averse individual prefers the distribution of a random outcome that second order stochastically dominates that of another competing outcome, our empirical results suggest that at least, the JTPA could be beneficial for risk-averse female workers who complied with assignment of the JTPA offer.
Conditional tail expectation (CTE), as a parameter of interest, has already received considerable attention in finance and risk management. In particular, CTE of an asset's return, known as the expected shortfall, is now a benchmark risk assessment measure in the industry. In the previous literature, the estimation of expected shortfalls is often based on estimated truncated means of the assets' returns LX_2013,Hill_2015. Recently, there has been a growing interest in structural models for forecasting the expected shortfall using the FZ loss PFC_2019, Taylor_2019, MT_2020, CYY_2022. These studies address the problem of finding the best model to predict the expected shortfall. In contrast, our paper, which also builds upon the FZ loss minimization approach, focuses on the CTATE estimation in a causal inference framework and thus has a fundamentally different objective from these previous works. Finally, in a study independent of ours, WTH_2024 developed a method for estimating the causal effect related to the CTE of the potential outcome for compliers. However, their estimation relies on a plug-in conditional quantile estimate and is quite different from ours, which uses the FZ loss.
The remainder of this paper is organized as follows. In Section 2, we introduce the notion of CTATE, set forth the causal framework, and present our estimation approach as well as the implementation algorithm. In Section 3, we establish the asymptotic properties of our CTATE estimator. In Section 4, we conduct simulation experiments to examine the finite-sample performance of the proposed method. In Section 5, we illustrate the usefulness of our approach in an empirical application using the JTPA data. Finally, we conclude the paper in Section 6. The proofs of all the technical results are provided in the appendix.
We follow the Rubin causal model with potential outcomes IR_2015. Let $D$ denote the treatment status having a value in $\{0,1\}$ (binary variable). Following convention, here an individual is deemed treated if and only if her treatment status $D$ takes a value of $1$, and $0$ otherwise. We use $Y_{d}$ to denote the potential outcome under the $d$th treatment, where $d\in\{0,1\}$, and $Y:=DY_{1}+(1-D)Y_{1}$ to denote the observed outcome. To formally define the CTATE, let $Q_{Y_{d}|X}(\tau)$ denote the $\tau$-quantile of the distribution of the potential outcome $Y_{d}$ conditional on $X$. Let $CTE_{Y_{d}|X}\left( \tau \right) :=E\left[Y_{d}|X,Y_{d}\leq Q_{Y_{d}|X}(\tau)\right]$ denote the CTE of $Y_{d}$ at the $\tau$-quantile level, given $X$. Note that, for any continuous random variable $Y_{d}$,
We define the CTATE at the $\tau $th quantile level as
In the literature, the QTE conditional on $X$ at the $\tau $th quantile is defined as
Using ((ref)), ((ref)) and ((ref)), we note that the CTATE is related to QTE through the equation:
Therefore, if the function $Q_{Y_{d}|X}(\tau )$ is linear in quantile-specific parameters (e.g., CH_2006, CHJ_2007, CH_2008 and CHJ_2009), then both $CTE_{Y_{d}|X}\left(\tau\right)$ and $CTATE(\tau )$ are also linear in quantile-specific parameters. Accordingly, for such a linear specification setting, if the treatment status $D$ is also exogenous, we can consistently estimate the structural parameters of the function $Q_{Y_{d}|X}(\tau )$ by minimizing the check loss for the standard quantile regression of KB_1978 and estimate those of $CTE_{Y_{d}|X}\left(\tau\right)$ by minimizing the FZ loss. This then allows for consistent estimation of the corresponding QTE and CTATE. However, in the presence of endogeneity, such an estimation procedure may be invalid and may bias the results.
We consider treatment effect estimation in a two-sided noncompliance framework in which the treatment status is potentially endogenous. Assume that there is also a binary instrumental variable $Z\in \{0,1\}$, which acts as an indicator for a random assignment of eligibility for receiving a treatment. Given noncompliance, there are four types ($T$) of individuals in the population: compliers $(T=c)$, always-takers $(T=a)$, never-takers $(T=ne)$ and defiers $(T=de)$. Always-takers will and never-takers will not be treated regardless of their eligibility to receive the treatment. Compliers will take the treatment only if they are eligible and defiers will do so only if they are not eligible. Let $D_{z}$, $z\in \{0,1\}$, denote potential treatment statuses under eligibility assignment $z$. We can then summarize the relations between type $T$ and the potential treatment statuses $D_{1}$ and $D_{0}$ as follows: (1) $D_{1}=D_{0}=1$ if and only if $T=a$; (2) $D_{1}=D_{0}=0$ if and only if $T=ne$; (3) $D_{1}=1$ and $D_{0}=0$ if and only if $T=c$; (4) $D_{1}=0$ and $D_{0}=1$ if and only if $T=de$.
Following AAI_2002, we make the following assumptions to identify the treatment effects for compliers under the two-sided noncompliance framework:
Assumptions similar to those stated above are common in the literature on the Rubin causal model with instrumental variables IA_1994, AIR_1996, Abadie_2003, IR_2015. Assumption 1.1 is analogous to an exclusion restriction requiring that, given $X$, the instrument $Z$ should not directly affect both the potential outcomes and potential treatment statuses. Thus the effect of $Z$ on outcome $Y$ can only occur through that on the observed treatment status $D$. This assumption holds when $Z$ is randomly assigned conditional on the covariates $X$. Assumptions 1.2 and 1.3 are mild and ensure that the relevance condition that $Z$ and $D$ are correlated conditional on $X$ holds. Assumption 1.3 also guarantees that $P\left(T=c|X\right)>0$ Abadie_2003 so that there is a nonzero proportion of compliers in the population. Assumption 1.4 is a monotonicity condition, implying that the potential treatment status $D_{z}$ is weakly increasing with $z$, thereby ruling out the existence of defiers. This condition is plausible in various applications. See, for example, AAI_2002 and Abadie_2003 for further discussions on these assumptions.
Under Assumption 1, AAI_2002 show that unconfounded treatment selection holds within the group of compliers:
Therefore, for the group of compliers, the distribution of potential outcome $Y_{d}$ given $X$ amounts to that of the observed outcome $Y$ given $X$ and $D=d$, which enables us to derive the causal parameters of interest for the compilers. In the present paper, we focus on the CTATE estimation for compliers, or local CTATE (LCTATE). Let $CTATE(\tau,T=c)$ and $QTE(\tau, T=c)$ denote the CTATE and QTE for compliers at the $\tau$th quantile, respectively, given $X$. If the complier quantile regression model is specified as linear-in-parameters,
it then follows from ((ref)) that the complier CTE given $D$ and $X$ is also linear in parameters such that
where \[ \alpha_{2,\tau} =\frac{1}{\tau}\int\nolimits_{0}^{\tau }\alpha _{1,u}du,\text{ } \boldsymbol{\beta}_{2,\tau } =\frac{1}{\tau}\int\nolimits_{0}^{\tau }\boldsymbol{\beta}_{1,u}du. \]
From above, it can be seen that $CTATE(\tau,T=c)=\alpha_{2,\tau}$ and $QTE(\tau, T=c)=\alpha_{1,\tau}$. In ((ref)), we are mainly concerned with the scalar parameter $\alpha_{2,\tau}$, which delivers the $\tau$th quantile level CTATE for compliers. However, this model is not directly applicable for estimation because the group of compliers are generally not observed in the data. To address this issue, we follow the weighting scheme of AAI_2002 and Abadie_2003, who show that under Assumptions 1, given any function $h(Y,D,X)$ with a finite mean, the following result holds:
where
Exploiting this result, we can estimate the parameters in ((ref)) by minimizing a sample analog of the numerator term of ((ref)) with the function $h$ being a loss function for the estimation of the CTE. For this purpose, we propose using the FZ loss FZ_2016, which leads to a weighted empirical FZ loss minimization-based estimation problem. Before formally setting forth the FZ loss in the next section, we note that the weight $K\left(D,Z,X\right)$ may not be non-negative for all realizations of $\left(D,Z,X\right)$, thus rendering the estimation problem computationally difficult. Following AAI_2002, we replace $K\left(D,Z,X\right)$ with
It is straightforward to see that
In addition, under Assumption 1, we have that $\bar{K}\left(Y,D,X\right) = P\left(T=c|Y,D,X\right)$ so that the projected weight $\bar{K}\left(Y,D,X\right)$ is almost surely non-negative.
For practical implementation, we plug in consistent estimators of the unknown components $\pi(X)$ and $E\left(Z|Y,D,X\right)$ of ((ref)) to estimate $\bar{K}\left(Y,D,X\right)$. Let $\hat{\bar{K}}_{i}:=\hat{\bar{K}}(Y_{i},D_{i},X_{i})$ denote such a plug-in estimator evaluated at the $i$th observation. We will show in Section 3 that $\hat{\bar{K}}$ is consistent for $\bar{K}$ under certain regularity conditions. However, the estimated weight $\hat{\bar{K}}_{i}$ might not be guaranteed to lie within the interval $(0,1)$ for every observation in the finite samples. If $\hat{\bar{K}}_{i}>1$ or $\hat{\bar{K}}_{i}<0$, we will then further truncate it by resetting its value to be $1$ or 0 accordingly.
Before we proceed to parameter estimation, it is worth comparing the setting of our paper to that of the instrumental variable quantile regression (IVQR) framework of CH_2005, CH_2006. In the IVQR setting, the outcome $Y$ is assumed to be strictly increasing in a structural unobservable that can correlate with the treatment status $D$ but is uniformly distributed conditional on the covariates $X$ and instruments $Z$. Such a setting allows for identification of the quantiles of potential outcomes $Q_{Y_{d}|X}(\tau)$ for $d\in\{0,1\}$ as well as the resulting QTE and CTATE, which are conditional on $X$ but unconditional over the agent's compliance status. By contrast, in our framework the QTE and CTATE are only identified for compliers. As the agent's compliance decision can be endogenous, the QTE and CTATE unconditionally over the agent's compliance status are generally different from their local counterparts conditional on the group of compilers.
There are other notable differences between the IVQR setting and the causal framework of our paper. The IVQR leaves the dimensionality of unobserved heterogeneity in the treatment selection equation unrestricted. However, in the outcome equation, all unobserved heterogeneities should be summarized by a scalar unobservable that needs to satisfy the rank similarity condition, a restriction on variations of this unobservable across different treatment statuses. By contrast, our setting does not restrict dimensionality of the unobservables in the outcome equation but its monotonicity assumption (Assumption 1.4) essentially restricts the unobservables in the treatment selection equation to being a scalar. Thus the IVQR and local QTE models are nonnested and may complement each other in empirical analysis Melly_2017.
Despite the differences between the IVQR and local QTE, Wuthrich_2020 show that the QTE in the IVQR amounts to that for compliers at adjusted quantile levels under certain regularity conditions. This finding has interesting implications for exploring the connection between the two estimation frameworks. For example, if the QTE for compliers is constant across quantiles, then the CTATE for compliers and the IVQR-based QTE are all constant across quantiles. Moreover, the CTATE for compliers and the IVQR-based QTE will all be of the same sign if the QTE for complier is positive (or negative) across all quantiles.
The FZ loss FZ_2016 is a class of consistent loss functions for eliciting both the quantile and CTE of a random variable, meaning that these two statistical functionals can be jointly identified by minimizing the expectation of the FZ loss. Interestingly, unlike the quantile, which can be elicited with some consistent loss functions (say, the check loss), there does not exist a consistent loss function solely for eliciting the CTE FZ_2016.
Let $\mathcal{F}$ be a class of distribution functions on real numbers $\mathbb{R}$ with finite first moments and unique $\tau$-quantiles. For a random variable $Y$ following a distribution $F\in\mathcal{F}$, let $Q_{Y}\left(\tau\right)$ denote the $\tau$-quantile and $CTE_{Y}\left(\tau\right)$ denote the corresponding CTE. FZ_2016 show that
where
is the FZ loss. Here $LQ_{\tau}\left(q,y\right):=\tau^{-1}\max\left(q-y,0\right)-q$, $\tau\in\left(0,1\right)$ is the quantile level, $(q,e,y)\in\mathbb{R}^{3}$ and $e\leq q$. $G_{1}\left(.\right)$, $G_{2}\left(.\right)$ and $\eta\left(.\right)$ are functions of real numbers and $G_{2}^{\prime}\left(.\right)$ is the subgradient of $G_{2}\left(.\right)$. It is required that $1_{\left[-\infty,q\right)}G_{1}\left(.\right)$ be $F$ integrable for all $q\in\mathbb{R}$ and all $F\in\mathcal{F}$, and $G_{2}^{\prime}\left(.\right)$ and $\eta(.)$ be $F$ integrable for all $F\in\mathcal{F}$. In addition, $G_{1}\left(.\right)$ should be increasing and $G_{2}\left(.\right)$ should be increasing and convex. If $G_{2}\left(.\right)$ is both strictly increasing and strictly convex, the minimizer in ((ref)) is unique and the loss function $FZ_{\tau}\left(q,e,y\right)$ is said to be strictly consistent\footnote{Let $L(x,y)$ denote a loss function for obtaining a statistical functional of a random variable. In our case, $L=F_{\tau}^{sp}$ and $x=(q,e)$. Let $\mathcal{F}$ denote a class of distribution functions and $F$ be an element in $\mathcal{F}$. Let $\lambda:\mathcal{F}\mapsto \mathbb{S}$ denote a statistical functional which maps $F\in \mathcal{F}$ to a set $\mathbb{S}$, which in our case amounts to the set $\mathbb{R}^{2}$. $L(x,y)$ is consistent for the statistical functional $\lambda(F)$ if $ E_{F}\left[L\left(\lambda(F),Y\right)\right]\leq E_{F}\left[L\left(x,Y\right)\right]$ for all $F\in\mathcal{F}$ and all $x\in\mathbb{S}$ and a random variable $Y$ following the distribution $F$. If a loss function is consistent and $E_{F}\left[L\left(\lambda(F),Y\right)\right]=E_{F}\left[L\left(x,Y\right)\right]$ implies $x=\lambda\left(F\right)$, the loss function is said to be strictly consistent.} for the $\tau$-quantile and the corresponding CTE of all $F\in\mathcal{F}$ FZ_2016.
Let $\boldsymbol{\theta}_{1,\tau}=(\alpha_{1,\tau},\boldsymbol{\beta}_{1,\tau}^\top)^\top$ and $\boldsymbol{\theta}_{2,\tau}=(\alpha_{2,\tau},\boldsymbol{\beta}_{2,\tau}^\top)^\top$ be the vectors that collect the true parameters of models ((ref)) and ((ref)), respectively. Using the weighted loss minimization approach of Section 2.2, we can estimate $\boldsymbol{\theta}_{1,\tau}$ and $\boldsymbol{\theta}_{2,\tau}$ by minimizing the sample analog of the left-hand side of ((ref)), with the function $h$ being replaced by the FZ loss ((ref)). Specifically, let $\hat{\boldsymbol{\theta}}_{1,\tau}$ and $\hat{\boldsymbol{\theta}}_{2,\tau}$ denote the estimators of $\boldsymbol{\theta}_{1,\tau}$ and $\boldsymbol{\theta}_{2,\tau}$. Assume that the data consist of a sample of $n$ observations $\left( Y_{i},D_{i},X_{i},Z_{i}\right), i=1,\ldots,n$. The estimators $\hat{\boldsymbol{\theta}}_{1,\tau}$ and $\hat{\boldsymbol{\theta}}_{2,\tau}$ can then be computed as the solution to the problem
where $q_{i}(\boldsymbol{\theta}_{1})=q(D_{i},X_{i},\boldsymbol{\theta}_{1}):= \alpha _{1}D_{i}+X_{i}^\top\boldsymbol{\beta}_{1}$ and $e_{i}(\boldsymbol{\theta}_{2})=e(D_{i},X_{i},\boldsymbol{\theta}_{1}):=\alpha _{2}D_{i}+X_{i}^\top\boldsymbol{\beta}_{2}$ are regression functions of ((ref)) and ((ref)) evaluated at parameter values $\boldsymbol{\theta}_{1}=(\alpha_{1},\boldsymbol{\beta}_{1}^\top)^\top$ and $\boldsymbol{\theta}_{2}=(\alpha_{2},\boldsymbol{\beta}_{2}^\top)^\top$, and $\tilde{K}_{i}=\min\left\{\max\left\{\hat{\bar{K}}_{i},0\right\},1\right\}$ is an estimate of $\bar{K}\left(Y_{i},D_{i},X_{i}\right)$ truncated to take value in the interval $[0,1]$.
To implement the estimation practically, we must specify the functional forms of $G_{1}\left(.\right)$, $G_{2}\left(.\right)$ and $\eta\left(.\right)$. In the present paper, we set $G_{1}\left(t\right)=0$, $G_{2}\left(t\right)=\eta\left(t\right)=\ln\left(1+\exp\left(t\right)\right)$ (the softplus function) for $t\in\mathbb{R}$, which enables us to solve the problem ((ref)) through a simple iterative computation algorithm. Under these settings, the loss function $FZ_{\tau}\left(q,e,y\right)$ becomes
The loss function ((ref)) is defined for all $\left(q,e,y\right)\in\mathbb{R}^{3}$. As $G_{2}(.)$ is the softplus function, it is straightforward to verify that $FZ_{\tau}^{sp}\left(q,e,y\right)$ is a strictly consistent loss function. Moreover, because $FZ_{\tau}(q,e,y)\geq0$ whenever $\eta(y)=\tau G_{1}(y)+G_{2}(y)$ DB_2019, it follows that $FZ_{\tau}^{sp}\left(q,e,y\right)$ is a non-negative function as well.
For our specifications, $G_{1}(t)=0$ is the most frequently used specification in practice FZ_2016,DB_2019,PFC_2019. Previous studies also have found that such a choice of $G_{1}(t)$ is a good candidate to make the FZ loss fulfill the positively homogeneous property, which means that the ordering of the losses will be unaffected by changing the unit of measurement of the data. Notice that although $G_{1}(q)=G_{1}(y)=0$, the check loss is still implicitly kept in $LQ_{\tau}(q,y)$ in ((ref)) and therefore estimating the quantile function is still viable through $FZ_{\tau}^{sp}$. As for $G_{2}(t)=\ln(1+\exp(t))$, our setting makes its first and second order derivatives bounded for all $t\in \mathbb{R}$. This property helps us to establish the relevant asymptotic results and asymptotic covariance matrix of the estimators in a parsimonious and less restrictive manner (see Section 3). In addition, this setting does not restrict the sign of the CTE estimate, which is crucial for applications in microeconometrics. In Appendix A.10, we provide a more detailed discussion on issues of the specifications. We also show that using alternative specifications for $G_{1}(t)$ and $G_{2}(t)$ does not seem to affect our empirical results.
With the FZ loss specification ((ref)), we solve the following minimization problem to obtain estimates of $\boldsymbol{\theta}_{1,\tau}$ and $\boldsymbol{\theta}_{2,\tau}$:
However, joint minimization over $\left(\boldsymbol{\theta}_{1},\boldsymbol{\theta}_{2}\right)$ in ((ref)) could be computationally nontrivial, as the function $LQ_{\tau}\left(q,y\right)$ involves the term $\max\left(q-y,0\right)$, which has a kink and is thus not everywhere differentiable. PFC_2019 adopted a smoothed approximation of the objective function in their FZ loss minimization problem. They solved the smoothed problem to obtain an approximate solution, which was then used as an initial estimate in a non-gradient based numerical procedure for the joint minimization of the original non-smoothed objective function in the estimation problem. Here we propose an alternative computational approach, which can be efficiently and stably implemented to solve the problem ((ref)).
Our computational algorithm builds on the following iterative scheme. Let $\hat{\boldsymbol{\theta}}_{1}^{(k)}$ and $\hat{\boldsymbol{\theta}}_{2}^{(k)}$ denote the computed estimates of $\boldsymbol{\theta}_{1}$ and $\boldsymbol{\theta}_{2}$ at the $k$-th iteration step. Note that, fixing the value of $\boldsymbol{\theta}_{1}$ at $\hat{\boldsymbol{\theta}}_{1}^{(k)}$, minimization of the objective function of ((ref)) with respect to $\boldsymbol{\theta}_{2}$ is a smooth optimization problem, which can be easily solved using commonly used numerical optimization solvers (e.g., the optim function of R). Next, fixing the value of $\boldsymbol{\theta}_{2}$ at $\hat{\boldsymbol{\theta}}_{2}^{(k)}$, minimization of the objective function of ((ref)) with respect to $\boldsymbol{\theta}_{1}$ reduces to the following weighted quantile regression estimation problem
where $c_{\tau}(u) = (\tau - 1\{u<0\})u$ is the check loss and the weight
is nonnegative. To see this, note that fixing $\boldsymbol{\theta}_{2}$ at $\hat{\boldsymbol{\theta}}_{2}^{(k)}$, the only term involving with $\boldsymbol{\theta}_{1}$ in ((ref)) is $\hat{\xi}_{i}^{(k)}LQ_{\tau}\left(q_{i}(\boldsymbol{\theta}_{1}),Y_{i}\right)$. Furthermore, it can be shown that $\hat{\xi}_{i}^{(k)}LQ_{\tau}\left(q_{i}(\boldsymbol{\theta}_{1}),Y_{i}\right) \propto \hat{\xi}_{i}^{(k)}c_{\tau}(Y_{i}-q_{i}(\boldsymbol{\theta}_{1}))$. Problem ((ref)) can be solved efficiently using well-known algorithms for quantile regression estimation (e.g., the quantreg package of R).
The proposed iterated scheme is an example of alternating optimization (or block coordinate descent algorithm). This type of optimization procedure is designed to replace the joint optimization of a multivariate function over all variables, which is sometimes difficult to solve, with a sequence of easily solved optimizations involving grouped subsets of the variables. Note that in practice, $\hat{\boldsymbol\theta}_{1}^{(1)}$ used for the first iteration step can be obtained from directly solving a weighted quantile regression estimation problem with weight $\tilde{K}_{i}$. Using the arguments above, our estimators $\hat{\boldsymbol{\theta}}_{1,\tau}$ and $\hat{\boldsymbol{\theta}}_{2,\tau}$ are thus computed as the solution upon the convergence of the iterative algorithm.
From ((ref)), a more straightforward approach to estimate the CTATE for compliers is to integrate estimates of the QTE for compliers from zero to a specific quantile level $\tau$. Suppose ((ref)) holds and let
denote such an integrated-QTE based estimator for estimating the CTATE for compliers at quantile level $\tau\in (0,1)$, where $\hat{\alpha}_{1,u}$ is an estimate of $\alpha_{1,u}$, the QTE for compliers at quantile level $u$. The integrated-QTE based estimator is derived from the difference between two integrated-quantile based estimators (e.g., LW_2018) for estimating the CTE of $Y_d$ at quantile level $\tau$. One advantage of ((ref)) is its ability to accommodate various models of QTE, such as those for compliers under the two-sided non-compliance framework AAI_2002, or the instrumental variable quantile regression (IVQR) framework of CH_2004, CH_2005, CH_2006, CH_2008.
Our proposed estimator for estimating the CTATE requires a specification of the functions $G_1$ and $G_2$ in the FZ loss. Nonetheless, our proposed estimator using the FZ loss still has several advantages over the integrated-QTE based estimator. One of them is that it is easier to derive theoretical properties of the proposed estimator. For example, to establish “pointwise” consistency of $\widehat{\text{IntQ}}^{\star}(\tau)$ for estimating the CTATE at $\tau\in (0,1)$, since this estimator integrates QTE estimates $\hat{\alpha}_{1,u}$ over $u\in (0,\tau)$, typically one would have to derive uniform consistency of $\hat{\alpha}_{1,u}$ over $u\in (0,\tau)$. However, as shown in Theorem 1 in Section 3, such uniform approximation property is not required to establish the pointwise consistency of our proposed estimator.
Furthermore, since the integration domain in ((ref)) includes low quantile indices that fall around the neighborhood of zero, conducting inference based on the estimator $\widehat{\text{IntQ}}^{\star}(\tau)$ requires addressing the statistical behavior of the QTE estimators evaluated at extreme quantiles. The literature on extremal quantile estimation and inference indicates that the finite sample distribution of the quantile regression estimator evaluated at extremal quantile levels cannot be accurately approximated by the normal distribution Chernozhukov_2005, CF_2011. Consequently, the asymptotic analysis of $\widehat{\text{IntQ}}^{\star}(\tau)$ is nonstandard as the process of estimated QTEs over $u\in (0,\tau)$ does not generally converge weakly to a Gaussian process.
In practice, one might consider using a “trimmed” version of ((ref)) to circumvent the issues associated with estimation at extreme quantiles, such as
where $\underline{\tau}$ is a small constant and $0<\underline{\tau}<\tau$. It is evident that ((ref)) will incur a truncation bias in the estimation of the CTATE ($\text{IntQ}^{\star}(\tau)$) at quantile level $\tau$, but it is hoped that this bias will not be excessively large if the trimming constant $\underline{\tau}$ is set to a very small value. However, as demonstrated in simulation results in Appendix A.9, even for a very small $\underline{\tau}$, the resulting bias can remain substantial.
Finally, the integrated-QTE based estimator ((ref)) relies on the computation of $\hat{\alpha}_{1,u}$ over a finely distributed collection of quantile indices for numerical integration. The computational burden of the QTE-based approach can be substantially reduced by employing a more sophisticated numerical integration algorithm than the brute-force one of using the estimated QTEs evaluated at many different quantiles. However, differences in computational costs between the two competing estimators could still occur as the complexity of the regression models and the sample size increase.
We conclude Section 2 by remarking here on one related causal parameters that build upon the CTATE. Define inter-quantile expectation (IQE) of a potential outcome $Y_{d}$ between quantile levels $\tau^{\prime}$ and $\tau$ with $0<\tau^{\prime}<\tau<1$ as
Using the notion of IQE, we can define the inter-quantile average treatment effect (IQATE), which is the difference between inter-quantile expectations of two potential outcomes:
Through ((ref)), the IQATE can be viewed as an aggregate of quantile treatment effects over a range of quantile levels and would thus be useful for providing a summary of heterogeneous treatment effects locally over a specified quantile range. Under Assumption 1 and model ((ref)), we can deduce from ((ref)) that the IQATE for compliers in the binary treatment causal framework of Section 2.2 reduces to \[ \frac{\tau\alpha_{2,\tau}-\tau^{\prime}\alpha_{2,\tau^{\prime}}}{\tau-\tau^{\prime}}, \] which can be readily estimated using the proposed estimator for estimating CTATE for compliers.
In this section we establish asymptotic properties of the estimators $\hat{\boldsymbol{\theta}}_{1,\tau}$ and $\hat{\boldsymbol{\theta}}_{2,\tau}$. We assume that the probability $\pi\left(X\right)$ in ((ref)) depends on some parameters $\boldsymbol{\gamma}_{0}$ and rewrite it as $\pi\left(X,\boldsymbol{\gamma}_{0}\right)$. Let $W:=\left(D, X^\top\right)^\top$, $V:=\left(Y,W\right)$ and $v_{0}\left(V\right):=E[Z|V]$. Let $\mathcal{V}$ denote the support of $V$ and $\boldsymbol{\Theta}$ be the parameter space of $\left(\boldsymbol{\theta}_{1,\tau},\boldsymbol{\theta}_{2,\tau}\right)$. Define
With ((ref)) to ((ref)), it can be seen that $K\left(D,Z,X\right)$ in ((ref)) can be expressed as $ K\left(W,Z;\boldsymbol{\gamma}_{0}\right)$, and $\bar{K}\left(Y,D,X\right)$ in ((ref)) can be expressed as $\bar{K}\left(V;v_{0},\boldsymbol{\gamma}_{0}\right)$. The truncated version of $\bar{K}\left(Y,D,X\right)$ is then given by $\tilde{K}\left(V;v_{0},\boldsymbol{\gamma}_{0}\right)$.
In the following, we derive asymptotic results for a more general setup where $Q_{Y_{i}|W_{i},T=c}\left(\tau\right)$ and $ CTE_{Y_{i}|W_{i},T=c}\left(\tau\right)$ take parametric structural forms $q_{i}\left(\boldsymbol{\theta}_{1}\right):=q\left(W_{i},\boldsymbol{\theta}_{1}\right)$ and $e_{i}\left(\boldsymbol{\theta}_{2}\right):=e\left(W_{i},\boldsymbol{\theta}_{2}\right)$, which could be nonlinear but are known up to some finite dimensional vectors of parameters. Let $\hat{\boldsymbol{\gamma}}$ and $\hat{v}$ denote some estimators of $\boldsymbol{\gamma}_{0}$ and $v_{0}$. We are interested in the parameter values $\boldsymbol{\theta}_{1,\tau}$ and $\boldsymbol{\theta}_{2,\tau}$, which are estimated by $\hat{\boldsymbol{\theta}}_{1,\tau}$ and $\hat{\boldsymbol{\theta}}_{2,\tau}$, where
and
We make the following regularity assumptions for the consistency of the parameter estimators.
Assumptions 2.1, 2.3 and 2.4 are standard. Assumption 2.2 is sufficient for the uniqueness of any $\tau$-quantile of the distribution of $Y$ given $W$ and $T=c$. Assumption 2.5 requires that the FZ loss function ((ref)) be strictly consistent for eliciting quantiles and CTEs. Assumption 2.6 is the rank condition for parameter identification. For ((ref)) and ((ref)), this condition immediately holds when, conditional on $T=c$, the support of $W$ is not contained in any proper linear subspace of $\mathbb{R}^{k}$ where $k$ denotes the dimension of $W$. Assumptions 2.5 and 2.6 ensure that the solution $\left(\boldsymbol{\theta}_{1,\tau},\boldsymbol{\theta}_{2,\tau}\right)$ is unique in the expected FZ loss minimization problem ((ref)). Assumptions 2.7 and 2.8 are imposed to establish uniform convergence of the empirical objective function in ((ref)) to its population counterpart. Assumption 2.7 is a dominance condition, which can hold under the compactness of $\boldsymbol{\Theta}$ and the aforementioned requirements for the FZ loss function, provided that $W$ has a bounded support. Assumption 2.8 hinges on the consistency of the plug-in estimators $\hat{\boldsymbol{\gamma}}$ and $\hat{v}$. With Assumptions 1 and 2, we can establish the following result.
In our numerical studies, we estimate the function $v_{0}$ using the power series estimator Newey_1997 and $\gamma_{0}$ with some parametric binary choice model. Let $W=(W_{d},W_{c})$, where $W_{d}$ and $W_{c}$ denote the discrete and continuous components of $W$ respectively. Let $\mathcal{W}_{d}$ denote the support of $W_{d}$, which takes only a finite number of possible values. The estimator $\hat{v}(V)$ takes the form
where, for each $m\in\mathcal{W}_{d}$, $\hat{v}_{m}(Y,W_{c})$ is a power series estimator of the conditional expectation $v_{0,m}(Y,W_{c}):=E[Z|Y,W_{d}=m,W_{c}]$. Let $\kappa_{m}$ denote the number of series terms in the power series approximation of $v_{0,m}$, $s_{m}$ be the order of continuous derivatives of $v_{0,m}$ and $r$ be the dimension of $W_{c}$. The next assumption allows us to verify the uniform consistency of the estimated weight $\hat{\bar{K}}(.)$.
Assumptions 3.1 and 3.2 are standard conditions for power series estimators Newey_1997. Assumptions 3.3 is also a mild condition on the smoothness of $\pi(X,\boldsymbol{\gamma})$. Assumptions 3.4 requires that the estimator $\hat{\boldsymbol{\gamma}}$ be consistent for $\boldsymbol{\gamma}_{0}$. Under Assumption 3, we have the following result.
We now provide further regularity assumptions for deriving the asymptotic distribution of the estimators.
Assumption 4.1 can be easily fulfilled for various parametric estimators in econometrics. Assumption 4.2 is a technical condition related to the local Lipschitz continuity of the FZ loss on neighborhood of the true parameter value. Assumptions 4.3 and 4.4 are for the existence of the asymptotic covariance matrix of our estimator. Assumption 4.5, which strengthens Assumption 2.8, requires that $\bar{K}(V;\hat{v},\hat{\boldsymbol{\gamma}})$ should converge uniformly at a rate faster than $n^{-1/4}$. For the power series estimator ((ref)), this assumption holds under the conditions of Lemma 1 with the growth rate of $\kappa_{m}$ being further restricted such that $\kappa_{m}^{6}/n\rightarrow0$ and $n^{1/4}\kappa_{m}^{1-s_{m}/(r+1)}\rightarrow0$. The next theorem establishes the asymptotic normality of the proposed estimator and provides the form of its asymptotic covariance matrix.
For the asymptotic covariance matrix of the estimators, the effect of the presence of the estimated parameter $\hat{\boldsymbol{\gamma}}$ is taken into account. Ignoring the effect of the estimated nuisance parameters will cause the asymptotic covariance matrix to be inconsistent, leading to invalid confidence interval constructions NM_1994. Without such an effect, the vector $\mathbf{J}_{\tau}$ will only have the first term. For estimating the asymptotic covariance matrix, we can use a plug-in estimator by replacing $\left(\boldsymbol{\theta}_{1,\tau},\boldsymbol{\theta}_{2,\tau},v_{0},\boldsymbol{\gamma}_{0}\right)$ in the formula of the asymptotic covariance matrix with their estimates $\left(\hat{\boldsymbol{\theta}}_{1,\tau},\hat{\boldsymbol{\theta}}_{2,\tau},\hat{v},\boldsymbol{\hat{\gamma}}\right)$. In Appendix A.1, we detail the estimation procedures and derive the asymptotic result of the plug-in estimator when $FZ_{\tau}^{sp}\left(q,e,y\right)$ is used and the structural forms of $q\left(W,\boldsymbol{\theta}_{1,\tau}\right)$ and $e\left(W,\boldsymbol{\theta}_{2,\tau}\right)$ are linear in parameters. The resulting estimated asymptotic covariance matrix is then used to construct the pointwise confidence bands of the estimated CTATE and QTE for compliers in the empirical application in Section 5.
The theoretical results of the proposed estimator above can be used immediately to derive the theoretical properties for estimating the IQATE in ((ref)). Again, we focus on the case when the models are linear in parameters. Divide the interval $\left(0,1\right)$ into $L+1$ subintervals with $L$ break points $\tau_{1},\ldots,\tau_{L}$, where $0<\tau_{1}<\tau_{2}\ldots\tau_{L}<1$. For $l,l^{\prime}\in\left\{ 1,\ldots,L\right\} $, and $l^{\prime}<l$, the IQATE between quantile levels $\tau_{l^{\prime}}$ and $\tau_{l}$ is \[ IQATE\left(\tau_{l},\tau_{l^{\prime}}\right)=\frac{\tau_{l}\alpha_{2,\tau_{l}}-\tau_{l^{\prime}}\alpha_{2,\tau_{l^{\prime}}}}{\tau_{l}-\tau_{l^{\prime}}}. \] If Assumptions 1, 2 and 4 hold, using Theorem 1 we can obtain the pointwise consistency:
The standard deviation of the above estimator for IQATE with $\hat{\alpha}_{2,\tau_{l}}$ and $\hat{\alpha}_{2,\tau_{l^{\prime}}}$ is given by \[\sqrt{ \frac{1}{\left(\tau_{l}-\tau_{l^{\prime}}\right)^{2}}\left[\tau_{l}^{2}Var\left(\hat{\alpha}_{2,\tau_{l}}\right)+\tau_{l^{\prime}}^{2}Var\left(\hat{\alpha}_{2,\tau_{l^{\prime}}}\right)-2\tau_{l}\tau_{l^{\prime}}Cov\left(\hat{\alpha}_{2,\tau_{l}},\hat{\alpha}_{2,\tau_{l^{\prime}}}\right)\right]}. \] If $\dim\left(X\right)=p\times1$, $Var\left(\hat{\alpha}_{2,\tau_{l}}\right)$ (and $Var\left(\hat{\alpha}_{2,\tau_{l}^{\prime}}\right)$) is the $\left(p+2\right)$th diagonal element of the matrix $\mathbf{H}_{\tau_{l}}^{-1}\boldsymbol{\Omega}_{\tau_{l}}\mathbf{H}_{\tau_{l}}^{-1}$ (and $\left(\mathbf{H}_{\tau_{l^{\prime}}}^{-1}\boldsymbol{\Omega}_{\tau_{l^{\prime}}}\mathbf{H}_{\tau_{l^{\prime}}}^{-1}\right)$), and $Cov\left(\hat{\alpha}_{2,\tau_{l}},\hat{\alpha}_{2,\tau_{l^{\prime}}}\right)$ is the $\left(p+2\right)$th diagonal element of the matrix $\mathbf{H}_{\tau_{l}}^{-1}E\left[\mathbf{J}_{\tau_{l}}\mathbf{J}_{\tau_{l^{\prime}}}^\top\right]\mathbf{H}_{\tau_{l^{\prime}}}^{-1}$, both scaled by $n^{-1}$.
Next we provide the result of uniform consistency: $ \left(\hat{\boldsymbol{\theta}}_{1,\tau},\hat{\boldsymbol{\theta}}_{2,\tau}\right)\stackrel{p}{\rightarrow}\left(\boldsymbol{\theta}_{1,\tau},\boldsymbol{\theta}_{2,\tau}\right)$ uniformly over $\tau\in\mathcal{T}\subset\left(0,1\right)$. We focus on a special case in which the loss function $FZ_{\tau}^{sp}(q,e,y)$ is used. Let \[h\left(\tau,\boldsymbol{\theta}\right):=\tau\left[ FZ_{\tau}^{sp}\left(q\left(W,\boldsymbol{\theta}_{1}\right),e\left(W,\boldsymbol{\theta}_{2}\right),Y\right)-\ln(1+\exp(Y))\right], \]where $\boldsymbol{\theta}:=\left(\boldsymbol{\theta}_{1},\boldsymbol{\theta}_{2}\right)$. Notice that minimizing $E\left[FZ_{\tau}^{sp}\left(q\left(W,\boldsymbol{\theta}_{1}\right),e\left(W,\boldsymbol{\theta}_{2}\right),Y\right)\right]$ and $E\left[h\left(\tau,\boldsymbol{\theta}\right)\right]$ w.r.t. $\boldsymbol{\theta}$ will obtain the same minimizer, as the former is just the latter scaled by a positive and finite constant $1/\tau$ and plus $E\left[\ln(1+\exp(Y))\right]$. Therefore using $h\left(\tau,\boldsymbol{\theta}\right)$ as a loss function to estimate $\boldsymbol{\theta}$ will yield the same result as using $FZ_{\tau}^{sp}\left(q\left(\boldsymbol{\theta}_{1}\right),e\left(\boldsymbol{\theta}_{2}\right),y\right)$. We will use $h\left(\tau,\boldsymbol{\theta}\right)$ as the loss function in proving the uniform consistency. The proof will rely on using the following additional assumptions and results of Newey_1991 to show that the empirical loss function uniformly converges to the true loss function over $\left(\tau,\boldsymbol{\theta}\right)\in\mathcal{T}\times\boldsymbol{\Theta}$.
Assumption 5.1 is standard. Assumption 5.2 is for the equicontinuity of the true loss function $E[\bar{K}(V;v_{0},\boldsymbol{\gamma}_{0})h(\tau,\boldsymbol{\theta})]$. Assumptions 5.3 requires the first moment of the quantity $B_{k,\tau}(V)$, $k=1,\ldots,5$ (defined in Lemma A.2 in the Appendix A.4) to be finite. This is a technical condition related to the global Lipschitz continuity of the FZ loss on the parameter space. The following theorem establishes uniform consistency of our estimator over the quantile index range $\mathcal{T}$.
In this section we examine the finite-sample performance of our proposed method in the simulations. The data $\left(Y_{i},D_{i},Z_{i},X_{i}\right)$, $i=1,\ldots,n$, where $X_{i}=\left(X_{1i},X_{2i}\right)$, for the simulation are generated according to the following design:
Here $U\left(0,1\right)$ is the uniform distribution in $[0,1]$, $Bern\left(p\right)$ is the Bernoulli distribution with parameter $p$, $\Phi\left(.\right)$ denotes the cumulative distribution function of the standard normal random variable, and
where $MVN\left(\mathbf{0},\Sigma_{\varepsilon,\vartheta}\right)$ denotes the multivariate normal distribution with mean zero and covariance matrix $\Sigma_{\varepsilon,\vartheta}$. We set $T_{i}=a$, if $D_{1i}=D_{0i}=1$; $T_{i}=c$, if $D_{1i}=1,D_{0i}=0$; and $T_{i}=ne$, if $D_{1i}=D_{0i}=0$. Under this simulation design, the proportion of compliers is approximately 50%, the proportions of always and never takers are both approximately 25%, and there is no defier. The correlation coefficient $\rho$ controls for the degree of endogeneity. When $\rho\neq0$, for $T=a,ne$ (always and never takers), the treatment status $D$ is correlated with potential outcomes $Y_{1}$ and $Y_{0}$ through $\varepsilon$ and endogeneity arises. But for $T=c$ (compliers), the condition of unconfounded treatment selection: $D\perp\left(Y_{1},Y_{0}\right)|\left(X_{1},X_{2}\right)$ holds here. Under this setting,
The QTE and CTATE for compliers are $\alpha_{1,\tau}:=b_{0}Q_{\varepsilon|X,T=c}\left(\tau\right)$ and $\alpha_{2,\tau}:=b_{0}CTE_{\varepsilon|X,T=c}\left(\tau\right)$ respectively. We set the parameters $b_{0}=1$, $b_{1}=0$ and $b_{2}=b_{3}=1$. We consider two different sample sizes $n\in\{500,3000\}$, and the simulation is iterated 1000 times. We perform simulations under the cases of $\rho=0$ (no endogeneity) and $\rho=0.5$ (with endogeneity).
We consider two alternative constructions of the projected weights in the weighted FZ loss minimization approach and assess how they affect the finite-sample performance of our proposed method. The first weight estimator is given by
where $\hat{v}\left(Y,D,X\right)$ is an estimate for $E\left[Z|Y,D,X\right]$ from a polynomial regression, and $\pi\left(X,\hat{\boldsymbol{\gamma}}\right)$ is an estimate for $P(Z=1|X)$ from a probit model with an intercept term and covariates $\left(X_{1},X_{2}\right)$. The estimate $\hat{v}\left(Y,D,X\right)$ is computed separately for the $D=1$ and $D=0$ groups, and the covariates used in the regressions include $(Y,X)$, their higher order and interaction terms. For the second weight estimator, we first calculate $K\left(D,Z,X\right)$ in ((ref)) for each observation using $\pi\left(X,\hat{\boldsymbol{\gamma}}\right)$ obtained from the same probit model used in the first estimator. We then fit a polynomial regression of the calculated $K\left(D,Z,X\right)$ on $\left(Y,D,X\right)$, their higher order and interaction terms and two additional covariates $\left(D_{0},D_{1}\right)$ which entail crucial information for classifying the types of individuals. The fitted value from the polynomial regression is used as the second projected weight estimator. By the law of iterated expectations, the identity ((ref)) still holds with the weight function $\bar{K}(Y,D,X)$ being replaced by the weight $\bar{K}_{2}(Y,D,X,D_{1},D_{0}):=E[K(D,Z,X)|Y,D,X,D_{1},D_{0}]$, which motivates our construction of the second type of projected weight. Adopting the proof of Lemma 3.2 in AAI_2002, we can see that $\bar{K}_{2}(Y,D,X,D_{1},D_{0})=P(T=c|Y,D,X,D_{1},D_{0})$. Thus working with the weight $\bar{K}_{2}$ amounts to estimating with information on classifying the types of individuals. We note that this weight estimator is infeasible in practice because we do not observe both $D_{0}$ and $D_{1}$ for each individual in the data. We compare the results using ((ref)) with those of $\bar{K}_{2}$ as the latter exploits more information and is expected to improve the performances of the resulting weighted FZ loss minimization estimators. Finally, since the weights are bounded from zero to one, if the estimated weights are greater than one or less than zero, we will shrink their values to one or zero.
Let M1 and M2 denote the weighted FZ loss minimization estimation approaches implemented using the aforementioned first and second projected weight estimators respectively. In addition, we also consider the following two benchmarks where no weighting scheme is used: (1) Estimation using all data without imposing any weight on the samples (no adjustment for endogeneity, denoted by M3); (2) Estimation using only the sample of compliers (oracle estimation, denoted by M4).
Figures (ref) to (ref) summarize the simulation results, including the biases, variances and mean squared errors (MSEs) of the CTATE and QTE estimators under settings M1 to M4. The key findings are discussed as follows. For $\rho=0$, there is no concern of endogeneity. In this case, the CTATE and QTE for compliers estimated using all the data without imposing any weight scheme (M3), have the lowest variance and MSE. Imposing weights to account for endogeneity would increase estimator variability, as can be seen from our theoretical result in Theorem 2, which shows the estimated weight contributes to variances of the proposed estimators. The estimation results using only the complier samples (M4) are associated with higher variance than those of M3. This is mainly due to the sample size effect as M4 tends to use fewer observations than M3 does in the implementation. The plots of variances for the estimated CTATE for compliers reveal a downward trend as the quantile level rises, but those for the estimated QTE for compliers indicate that the variances are larger at the lower and higher quantile levels. For the estimation bias, the simulation results are mixed and there is no clear dominant method here. The bias, variance and MSE are all improved as the sample size increases from 500 to 3000.
For $\rho=0.5$, the endogeneity issue arises. In this case, M3 results in a much higher bias and MSE than the other three methods, although it still results in a lower variance. The high MSE of M3 is mostly due to the high bias from a lack of adjustment for endogeneity. The reasons for the lower variance of M3 are the same as those in the case of $\rho=0$: all samples are used in the estimations and there is no estimated weight. Notably, M4 (the oracle estimation case) is associated with the lowest bias, because it uses only the sample of compliers. The bias curve under M2 is flat over the quantile levels, whereas that under M1 shows declining trend. It is worth noting that M1 results in a higher MSE than M2, which suggests that using potential treatment status information to estimate the weight may be helpful on improving the MSE. The performances of M1, M2 and M4 evidently improve as the sample size increased. In summary, in the case of endogeneity, although estimation without adjustment for endogeneity (M3) could still yield a lower variance, it could also result in a higher bias and thus a higher MSE. The higher variances under M1 and M2 arise from the use of the estimated weights. However, these weight schemes help to effectively mitigate the estimation bias. Overall, relative to M3, the MSE under M1 and M2 can be substantially improved under the weighting adjustment for endogeneity.
In the following we illustrate the usefulness of our proposed method through an empirical study. We estimate the CTATE for compliers using our proposed weighted FZ loss minimization approach to evaluate the effects of enrolling in Title II programs of the Job Training Partnership Act (JTPA) in the US. We use the data of adult men and women who participated in these programs between November 1987 and September 1989. These data were previously used by AAI_2002, who estimated the QTE for compliers on the earnings of the job training programs. We assume that the observations are i.i.d., as in AAI_2002, for estimation. The outcome variable $Y$ is the sum of earnings in the 30 months after the random assignment. In practice, the sum of 30-month earnings is generally viewed as a continuous variable, and we assume that it remains continuously distributed given $W=(D,X)$ and conditional on the group of compliers. The treatment variable $D$ is a binary variable for enrollment in the JTPA services (1) or not (0), and the instrumental variable $Z$ is a binary variable for being offered such services (1) or not (0). The exogenous covariates include age, which is a categorical variable, as well as a set of dummy variables: black, Hispanic, high-school graduates (including GED holders), marital status, AFDC receipt (for adult women), whether the applicant worked for at least 12 weeks in the 12 months preceding the random assignment, the original recommended service strategy (classroom, OJT/JSA, other), and whether earnings data were from the second follow-up survey. These exogenous covariates are all discrete variables, and as shown in Section 3, this is allowed in our method. The total sample size is 11,204 (5,102 for adult men and 6,102 for adult women).
Offers of the JTPA services were randomly assigned to applicants but only approximately 60% of those who were offered the services enrolled in the programs AAI_2002. This may induce the problem of endogeneity in that the treatment status may be self-selected and correlated with the potential outcomes. As the offers were randomly assigned and were considered to potentially affect an applicant's intention to participate in the program, we use offer assignment as an instrumental variable. Finally, in the data, there were still individuals who received the JTPA services but did not obtain the assignment. However, as pointed out by AAI_2002, the proportion of such violations relative to the entire samples is very low (less than 2%); therefore this has a negligible impact on our estimation.
We estimate ((ref)) and ((ref)) over a grid of quantile levels $\tau$ ranging from 0.1 to 0.9 with the grid size being 0.01 (81 grid points). We estimate $E[Z|Y,D,X]$ with a power series estimator separately for $D=1$ and $D=0$, and include the outcome variable $Y$ and its higher order terms as covariates. The offer assignment probability, $P(Z=1|X)$, is estimated using a probit model. These estimates are then used to construct the weight $\tilde{K}(.)$ in ((ref)) to account for endogeneity in the estimation problem. In Figure (ref), we present the CTATE and QTE estimates after endogenous adjustment for compliers (the parameter estimates $\alpha_{1,\tau}$ of ((ref)) and $\alpha_{2,\tau}$ of ((ref))) and the corresponding 95% pointwise confidence band (pcb), evaluated with the analytic standard errors of the estimators (see Appendix A.1), over the specified range of quantile levels. We also show the 95% bootstrap pointwise and simultaneous confidence bands (scb's) with 300 bootstrap samples. As suggested by HL_2021, we construct the bootstrap pcb using the percentile method Van_1998, which recenters the bootstrapped parameter estimator yet does not rescale it by its standard deviation. We construct two scb’s. The first one, denoted by scb-bootstrap-ns, is also constructed using the percentile method, and the second one, denoted by scb-bootstrap-qs, standardizes the bootstrapped estimator using the rescaled bootstrap quantile spread CFM_2013 to estimate the asymptotic standard error of the parameter estimator. The procedures for constructing the scb's can be found in Appendix A.5. The results without the endogenous adjustment are shown in Appendix A.6.
For both adult men and women, the CTATE estimates consistently increase across quantile levels, whereas the QTE estimates exhibit some fluctuations. Upon comparing the pcb's calculated using the analytic standard errors and the bootstrap method, the differences are small. For the case of adult men, the scb's demonstrate that both the CTATE and QTE estimates generally lack statistical significance across most quantile levels. This implies a negligible stochastic dominance relationship, indicating that the benefits of participating in the JTPA program for adult men are not conclusively supported by our data. For the case of adult women, the strength of the CTATE estimates surpasses that of men, with scb's showing statistically significant positivity at certain quantile levels. Overall, the results in Figure (ref) suggest a stronger evidence for adult women that earnings from not participating in the JTPA are second order stochastically dominated by those from participating in the JTPA, and the JTPA was beneficial for risk averse female workers who complied with the assignment of the JTPA offer.
We then divide the quantile levels into eight equal-length intervals and estimate the corresponding inter-quantile average treatment effect (IQATE) for compliers. We also compare the estimated IQATE with a naive estimator: a local average of the QTE estimates for compliers (LAVG-QTE) within the same interval of quantile levels. Figure (ref) shows the estimation results of the IQATE, the corresponding pcb's (implemented with the analytic standard errors and bootstrap), and the LAVG-QTE for earnings of adult men and women at different quantile level intervals.
For adult men, only the IQATE estimates above the quantile level of 0.6 are statistically significantly positive. It is worth noting that at the intervals of quantile levels of 0.8 and 0.9, the QTE estimates are not all statistically significantly positive under the pcb's, but the IQATE estimate is. This indicates that the IQATE estimate may provide a more coherent result when we want to evaluate a policy at quantile level intervals. For adult women, the IQATE estimates are statistically significantly positive over the eight quantile level intervals. Comparing the IQATE and LAVG-QTE estimates, for both cases, they are not very different when the quantile level is above 0.5, but at the quantile levels lower than 0.5, they rather show some mild differences in the case of adult women.
We have introduced the conditional tail average treatment effect (CTATE), defined as the difference between CTEs of potential outcomes. The CTATE is a valuable tool for policy evaluations, as it allows for capturing the heterogeneity of treatment effects over different quantiles and is useful for detecting second order stochastic dominance and for estimating the Lorenz curve. We have developed a semiparametric method using a class of consistent loss functions proposed by FZ_2016 to estimate the CTATE for compliers under endogeneity. We also have derived asymptotic properties for our proposed estimator. Our simulation results show that the proposed method works well in the presence of treatment endogeneity. We apply our estimation approach to a policy evaluation for the JTPA program participation. We find that, after adjustment for endogeneity, for the case of adult men, both the FOSD and SOSD relationships hardly held between earning distributions of those who participated in the JTPA and of those who did not. Yet, for adult women, the SOSD of earnings for the JTPA participant over those for non-participant appeared to hold. These empirical results suggest that the JTPA could be beneficial for the risk-averse female workers who complied with the assignment of the JTPA offer.
{ \setcounter{equation}{0}