Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
88,705 characters · 14 sections · 89 citation commands
Partial Mean Processes with Generated Regressors: Continuous Treatment Effects and Nonseparable Models
Partial mean with generated regressors arises in several econometric problems, such as the distribution of potential outcomes with continuous treatments and the quantile structural function in a nonseparable triangular model. This paper proposes a nonparametric estimator for the partial mean process, where the second step consists of a kernel regression on regressors that are estimated in the first step. The main contribution is a uniform expansion that characterizes in detail how the estimation error associated with the generated regressor affects the limiting distribution of the marginal integration estimator. The general results are illustrated with two examples: the generalized propensity score for a continuous treatment \citep*{HI04} and control variables in triangular models \citep*{NPV99ETA, IN09ETA}. An empirical application to the Job Corps program evaluation demonstrates the usefulness of the method. \\ \\ Keywords: Continuous treatment, partial means, nonseparable models, generated regressors, control variables \\ JEL Classification: C13, C14, C31
This paper studies nonparametric estimation of continuous treatment effects. Endogeneity is allowed in the continuous treatment variable and is corrected by including a control variable as an additional regressor.\footnote{According to M07, “A {\it control variable} is a function of observable variables such that conditioning on its value purges any statistical dependence that may exist between the observable and unobservable explanatory variables in an original model."} This approach leads to identification of the unconditional distribution of the potential outcome with continuous treatments by a partial mean, where a partial mean is defined as a marginal integration of the conditional outcome distribution over the control variable with the continuous treatments held fixed. We allow the control variable to be estimated in a preliminary step and enter the partial mean as a generated regressor. A generated regressor can be viewed as a possibly infinite-dimensional nuisance parameter in many economics examples: the control variables in triangular models \citep*{NPV99ETA, IN09ETA}, the generalized propensity score for continuous treatment \citep*{HI04}, or the propensity score for sample selection \citep*{DNV03}. The main contribution of this paper is to characterize how the estimation error of general generated regressors affects the limit properties of partial mean estimation and to provide uniform inference for the entire potential outcome distribution of continuous treatments.
Our results provide new insights on nonparametric estimation of continuous treatment effect models. Two key features of these models are non-separability in the unobservables and heterogeneity in treatment intensity effects. The proposed methods capture heterogeneous treatment intensity effects by estimating an array of distributional structural features that can be applied to a variety of economics questions. For example, when evaluating a social program, researchers might be interested in how the length of exposure to the program affects the wage distribution. The proposed method includes inference on smooth functionals of the outcome distribution process. In this example, a researcher could consider how inequality responds to the length of exposure to an anti-poverty program by tracing out the Gini coefficient of the wage distribution by time in the program. In demand analysis, we can estimate the Engel curve that describes how the distribution of household food consumption responds to an exogenous change in total expenditure. Analyzing these questions often involves generated regressors to account for the endogeneity of the continuous explanatory variable. Understanding the estimation error of the generated regressors is fundamental to perform correct inference.
We define the weighted {\it partial mean process} indexed by $(y,t)$ by
\footnote{The notation ${\bf 1}{\{A\}}$ is an indicator function for the event $A$.} The conditional expectation $F_{Y|TV}(y|t,V)$ is the conditional cumulative distribution function (CDF) of the outcome $Y$ given the continuous explanatory variables $T$ and control variables $V$. The regressor $V = v_0(\cdot)$ is a function of observables, which can be estimated in a first step as a {\it generated regressor}. In the second step, the conditional CDF $F_{Y|TV}(y|t,V)$ is estimated nonparametrically by a kernel regression using the first-step generated regressor. The third step is sample analogue of a marginal integration that averages out the regressor $V$ and the weight $W$, but fixes the value of the continuous treatment $T$ at $t$. Because the regression function $F_{Y|TV}(y|t,V)$ contains more arguments than being averaged over in the third step, this is known as a {\it partial mean} in the terminology of Newey94ET. As $T$ is continuous, the partial mean is an infinite-dimensional nonparametric object. We analyze the three-step estimator for the partial mean of the weighted conditional CDF in Eq.$\!$ ((ref)) that builds on and extends the partial mean literature. Our main contribution is a stochastic expansion that is uniform over $(y,t)$ and accounts for the estimation error of the generated regressors. The influence of the generated regressors is characterized by an integral operator on the estimation error $\hat v(\cdot) - v_0(\cdot)$. The estimator $\hat v$ can be (semi)parametric or nonparametric for a general function $v_0$, such as mean regression, density, or quantile regression.
Another set of our results is weak convergence of the partial mean process indexed by the threshold value $y$, for each $t$. By extending the results to the Hadamard-differentiable functionals of the partial mean process, we provide the limiting distribution and uniform inference method for estimating various distributional structural features and common inequality measures, such as the Gini coefficient. The mean is the {\it average structural function} or the {\it dose response function} \citep*{BP03,Flores07}. The quantile is the {\it quantile structural function} \citep*{IN09ETA}. The difference between two treatment levels is the quantile treatment effect. We show that a multiplier bootstrap method is valid for uniform inference over $y$, which enables functional hypotheses tests for the whole distribution, such as tests for stochastic dominance.
We illustrate the usefulness of our general results by two examples. The first is the generalized propensity score (GPS), defined as the conditional density function of treatment given observable characteristics $X$. Under the unconfoundedness assumption, the GPS is known to reduce dimensionality in the second-step regression \citep*{HI04}. Using the GPS as a generated regressor $V = f_{T|X}(t|X)$ or weight is now a common practice in program evaluation where the length of participation is often taken as the continuous treatment \citep*{ACW10, FFGN12ReStat, KSUZ12, GW, HHLP}. This paper is the first presentation of a complete limit theory of nonparametric kernel regression on the estimated GPS. Another causal object of interest is the treatment effect {\it for the treated}, where {\it the treated} subpopulation is defined by individuals who are currently choosing treatment value $\bar t$. The local average response and the marginal treatment effect for the treated are based on the conditional average structural function given $T = \bar t$ \citep*{AM05ETA, FHMV08ETA}. We add to the literature a new estimator of the effect for the treated that regresses on two GPSs, one for the counterfactual treatment value $t$ and one for the treated value $\bar t$, i.e., $V = (f_{T|X}(t|X), f_{T|X}(\bar t|X))'$. We illustrate the effectiveness of our methods by analyzing the Job Corps program. We estimate the causal effects of the length of participation on employment that uncovers heterogeneities in the effects of dosages of academic and vocational training. For example, we find that extending the length in Job Corps from one month to six months increases the proportion of weeks employed in the second year by 2.5% on average.
The second example is the control variable as in the triangular simultaneous equations models in NPV99ETA and IN09ETA. Our limit theory applies to nonparametrically estimate the bounds of the average and quantile structural functions. Moreover the stochastic expansion, which is uniform over $t$ and $y$, is useful when a partial mean is an intermediate step in a multi-step estimation a semiparametric setting. In such cases, the estimation error from the control variable might not be ignorable. This extension to functionals of the partial mean is a direct calculation from our stochastic expansion. For example, the compensating variation in welfare analysis studied by DB can be expressed as functionals of the average structural function.
To our best knowledge, this is the first paper that analyzes a marginal integration estimator for partial means of a kernel regression estimator, accounting for general generated regressors. There are two infinite-dimensional parameters in the partial mean: the generated regressor $V=v_0(\cdot)$ and the regression function $F_{Y|TV}(y|t, \cdot )$. The generated regressor enters the partial mean estimation through two channels: first, it is an {\it argument} of the regression function that directly enters the partial mean; second, it is a {\it regressor} that determines the functional form of the regression $F_{Y|TV}(y|t, \cdot)$. Our new uniform expansion distinguishes the {\it argument} and {\it regressor} roles of the generated regressor and contains three important {\it elements} --- (i) partial mean structure, (ii) generated regressor as an index function of observables, and (iii) projection of the weight $\mathbb{E}[W|V]$. This finding appears to be of significant practical importance by providing conditions under which the estimation error is negligible. These {\it elements} are generic for marginal integration estimators that involve a kernel regression with generated regressors, in both nonparametric or semiparametric models.
The rest of the paper is organized as follows: Section (ref) discusses our contribution and the related literature. Section (ref) introduces the setup and causal parameters of interest. We outline three-step nonparametric kernel estimation for the weighted partial mean process with generated regressors defined in Eq.$\!$ ((ref)). Section (ref) presents the main asymptotic theorems. Section (ref) illustrates the usefulness of our results by economic examples. Section (ref) presents the limit theories for treatment effects for the Hadamard-differentiable policy functionals and specifically for the quantile processes. A multiplier bootstrap method enables uniform inference. Section (ref) is the empirical application on the Job Corps program. Section (ref) concludes this paper. The proofs are in the Appendix.
Throughout this paper, let an uppercase letter $V$ be a generic $d_v$-dimensional random vector with support $\mathcal{V} = Supp(V)\subseteq \mathbb{R}^{d_v}$, where $d_v$ denotes the dimension of $V$. Denote the interior support of $V$ to be $\mathcal{V}_0$. The realized value of $V$ is denoted by a lowercase letter $v$.
The {\it argument} role of the generated regressor is played by an unknown function affecting the sampling variation of the final estimator directly and has been studied extensively in the literature; for example, PP89ETA, Andrews94, Sherman94, CLK03ETA, IL10, among many others. The challenge lies in the {\it regressor} role of the generated regressor that indirectly affects the sampling variation of the final estimator through the second-step nonparametric regression. There is a recent literature on nonparametric regression with generated regressors, for example, Song08ET,Song14JoE, Sperlich, MRS12A, MRS12B, HR13ETA, HR16, HahnLiaoRidder, EJL, among others. Most of the previous work focuses on the impact of generated regressors on the nonparametric regression and its application in semiparametric models. In contrast to the semiparametric models where the estimator integrates the regression function over {\it all} continuous regressors and hence is a full mean, our partial mean is a marginal integration over the generated regressor but evaluates the continuous regressor $T$ at a fixed value $t$.
HR13ETA are among the first to characterize the asymptotic variance of a semiparametric estimator contributed by the dual role of the generated regressors by using Newey94ETA path-derivative method. HR16 further consider the control variable estimator with non-separable errors and endogenous regressors. In contrast, we use stochastic expansion that gives sufficient conditions under which the asymptotically linear representation of the kernel-based estimator is valid. The empirical process approach allows us to express the impact of the generated regressor on the final estimator by an integral operator on the estimation error $\hat v(\cdot) - v_0(\cdot)$. So our result can be applied to general estimators for the generated regressor, such as regression or density.\footnote{ HahnLiaoRidder study two-step sieve M estimation, where the sieve estimation may involve generated regressors. Their parameters of interest are {\it known} functionals of the unknown functions estimated in both steps that do not include the partial mean/full mean. EJL use a stochastic equicontinuity argument on a full mean based on a general class of conditional mean functions. MRS12B study a class of semiparametric optimization estimators that involves nonparametric regression with generated regressors and verify sufficient conditions for the asymptotic normality developed in CLK03ETA. Besides the apparent difference to the aforementioned papers that our partial mean is infinite-dimensional, we give conditions directly on the kernel regression estimator and generated regressor. We do not assume the high-level smoothness assumption that the regression function is very smooth with respect to the regressor. We discuss more technical detail in Section (ref) in the Appendix.}
Our empirical process approach builds on the stochastic equicontinuity argument for the kernel estimator developed in MRS12A. While MRS12A contribute a detailed characterization of how the generated regressor affects the regression estimator, we focus on its impact on the partial mean of such regression. The partial mean is a function of the continuous variables $t$ and hence is estimated nonparametrically at a convergence rate slower than regular root-$n$. Deriving the weak convergence and multiplier method demands more involved technical arguments than applying the standard Donsker's Theorem as in Rothe10JoE, FP11, DHB, CFM12, for example.
This section introduces the causal objects of interest, which are identified by functionals of the partial means in Eq.$\!$ ((ref)). The partial mean is also a statistical object with broad applications, such as additive nonparametric models and differential equation solution introduced in Newey94ET. We use the potential outcome framework that is convenient for interpretation and simplifies notation. The treatment effect model is known to be equivalent to a nonseparable outcome with a general disturbance, where the outcome equation is $Y = \phi(T, \epsilon)$, e.g., IN09ETA, WC13ER. The error $\epsilon$ represents unobservable individual heterogeneity. No functional form assumption is imposed on the general disturbances $\epsilon$, like monotonicity, dimensionality, or separability. Let $Y(t) = \phi(t,\epsilon)$ denote the potential outcome corresponding to the treatment level $t$, where the randomness comes from the unobserved disturbances $\epsilon$. If the treatment value $\bar t$ is chosen by an individual $i$, then $T_i = \bar t$ and the observed outcome $Y_i =Y_i(\bar t)$ is one of the potential outcomes $\{Y_i(t) = \phi(t, \epsilon_i) \}_{t \in {\mathcal{T}}}$. For example, we observe the quantity demanded at the observed price, but cannot observe what the demand would have been given other prices. Regarding program evaluation, we observe the wage of a participant after one year in a job training program, but we wish to learn what the wage would have been if the participant had stayed in the program for two years.
The key causal object of interest is the unconditional distribution of $Y(t)$ by averaging out other covariates and unobservable heterogeneity. The cumulative distribution function (CDF) of the potential outcome $F_{Y(t)}(y) = \mathbb{E}[{\bf 1}{\{\phi(t, \epsilon) \leq y\}}] = \int {\bf 1}{\{\phi(t, \epsilon) \leq y\}} f_\epsilon(\epsilon) d\epsilon$ is the outcome distribution when the value of the treatment $T$ is fixed at $t$ and the expectation is taken with respect to the marginal density of $\epsilon$, $f_\epsilon$. In other words, it is the unconditional outcome distribution if, hypothetically, the individual had been assigned to the treatment level $t$. An array of estimands are often of interest based on these causal outcome distributions. We consider a general class of functionals $\Gamma$ on $F_{Y(t)}(\cdot)$. For example, if interest centers on the quantile treatment effect, we let $\Gamma$ be the quantile operator $\Gamma(F_{Y(t)}) = Q_\tau(Y(t)) = \inf\{y: F_{Y(t)}(y) \geq \tau\}$. The quantile treatment effect corresponding to a change from $\bar t$ to $t$ is $\Gamma(F_{Y(t)}) - \Gamma(F_{Y(\bar t)})$. When the outcome structural function $\phi$ is monotone in the error $\epsilon$, the {\it quantile structural function} $Q_\tau(Y(t))$ equals the structural function evaluated at $(t, \tau)$, $\phi(t,\tau)$, by a normalization. If interest is on the mean treatment effect, then let $\Gamma$ be the mean operator. When the outcome is separable $Y = \phi(T) + \epsilon$, the {\it average structural function} $\mathbb{E}[Y(t)]$ equals $\phi(t)$ up to an additive constant. Other inequality measures are also applicable, such as the coefficient of variation, the interquantile range, the Theil index, the Gini coefficient, the Lorenz curve.
We use a conditional independence assumption to show that the partial mean process in Eq.$\!$ ((ref)) identifies the causal objects of interest.
Under Assumption (ref),
for any $v \in Supp(V|T=t)$. Let $\mathcal{V}^*$ denote a subset of the support of $V$ conditional on $T=t$. We will give examples of particular $\mathcal{V}^*$ in the following sections. Then for any $v \in \mathcal{V}^*$, $\mathbb{E}[Y|T=t, V = v]$ is well-defined and hence $\mathbb{E}[Y(t)|V = v]$ is identified from Eq.$\!$ ((ref)). In words, the conditional mean of the potential outcome given the control variable is recovered by the mean of the observed outcome of those who had received $t$ and with the same value of the control variable. Therefore, we can identify the unconditional mean of $Y(t)$ defined on $\mathcal{V}^*$,
where $\pi(v) = {\bf 1}\{v \in \mathcal{V}^\ast\}$. This is the {\it dose response function} or {\it average structural function} for the subpopulation whose control variable takes values in $\mathcal{V}^*$.
When the common support $\mathcal{V}^*$ equals the support of $V$, i.e., $\pi(v) = 1$, the population average structural function $\mathbb{E}^*[Y(t)] = \mathbb{E}[Y(t)]$ is therefore identified. This is known as the {\it common support} or {\it overlapping assumption} in the literature, i.e., the support of $V$ conditional on $T=t$ equals to the support of $V$. The common support assumption requires everyone in the population to be able to find a match who shares the same value of the control variable $V$ and has received the counterfactual value $t$.
Nevertheless, the common support assumption is known to be strong in many applications. We do not impose this common support assumption, but estimate the causal object defined by the common support $\mathcal{V}^*$:
where the trimming function $\pi(v) \equiv {\bf 1}{\left\{\inf_{t\in \mathcal{T}^*} f_{TV}(t, v) \geq c\right\}} = {\bf 1}{\{v \in \mathcal{V^*}\}}$ with a positive constant $c$ and a range of treatment values of interest $\mathcal{T}^*$ that is an interior subsupport of $T$. The trimming function selects a common support $\mathcal{V}^*$, which is a subset of the intersection of the supports of $V$ conditional on $T=t$ for $t \in \mathcal{T}^*$,
An advantage of the trimming approach is that the identification set is valid for any positive trimming parameter $c$. Another purpose of trimming is technical: The trimming function allows us to work on a compact interior subsupport where the density functions are bounded away from zero, satisfying the smoothness Assumption (ref).\footnote{We can then use uniform convergence of the first- and second-step estimators over the range of integration that suffices for deriving the properties of the final estimator. The edge effect and the boundary problem of the local constant estimator are avoided. IN09ETA refer the edge effect to the case when the convergence rate of the estimator depends on how fast the joint density goes to zero on the boundary. As the joint density of the endogenous variable $T$ and the control variable $V$ often goes to zero at the boundary of the support of the control variable in the triangular model, averaging over the control variable over the entire support upweights the tails relative to the joint distribution. } The trimming approach has been used in similar contexts, for example, Newey94ET, IN09ETA, IchimuraTodd, FFGN12ReStat, to simplify the theoretical analysis and practical application.
The weight $W$ is the additional observed variable that is not involved in the regression function $F_{Y|TV}$. We will discuss specific estimators for the economic examples in the later sections. We find that the estimation approach outlined below has different properties depending on the details of implementation corresponding to each estimand. As a result, different asymptotic distribution results are proceeded for the different versions of this general estimation approach described below. The estimation procedure involves three steps:
From the above procedure, we see that the generated regressors $\{\hat V_i\}_{i=1}^n$ in Step 1 are used as {\it regressors} in Step 2 and used again as {\it arguments} of the regression function in Step 3. The treatments $T$ or covariates $V$ could contain discrete variables and the kernel is replaced by an indicator function, known as the frequency method. For notational convenience, discrete covariates are not allowed for. The smoothness Assumption (ref) (ii) requires that the treatment variables cannot have point masses, i.e., ${\rm Pr}(T=t) = 0$ for $t \in \mathcal{T}^\ast$.
We first present the limit theory for estimating the partial mean process when all the regressors are observed. This is for the case to identify the causal distribution of the potential outcome under unconfoundedness, where the Conditional Independence Assumption (ref) is satisfied by $V = X$ observable characteristics. This limiting distribution is also for the case when the Step 1 estimation error of the generated regressor is asymptotically ignorable. We show the partial mean estimator weakly converges to a Gaussian process indexed by $y$ for each $t$, which serves as a building block for analyzing the estimator involving generated regressors in Section (ref). Our main result is a uniform stochastic expansion of the three-step estimator, revealing the influence of estimating the generated regressors of general function form on the final estimator. The stochastic expansion is uniform over the treatment value of interest $t$ and $y$ for the CDF. We characterize the leading bias that can be useful to compute the asymptotically mean squared error optimal bandwidth and to conduct robust inference; see recent development by CCF and AK, for example. The asymptotic theorems focus on the local constant estimator. Then we present the results for the local polynomial estimators.
Following the procedure described in Section (ref), we estimate the partial mean process with observable regressors $V$ in Eq.$\!$ ((ref)) by $\hat F_{Y(t)}(y;V,W\hat \pi) = n^{-1} \sum_{i=1}^n \hat F_{Y|TV}(y|t, V_i) \ W_i\hat \pi(V_i)$, where the trimming function $\hat \pi(V_i) = {\bf 1}{\{\inf_{t\in\mathcal{T}^*}\hat f_{TV}(t, V_i) \geq c\}}$.
The convergence rate of the partial mean estimator is $\sqrt{nh^{d_t}}$ depending on $d_t$, the dimension of the continuous conditioning variables that are fixed in the partial mean. It is known that when more conditioning variables are averaged out in the partial mean, the nonparametric estimator converges at a faster rate \citep*{Newey94ET}. A {\it full mean} estimator is the case when all arguments of the regressors are averaged out. That is the case when the regressor $T$ is dropped or $T$ is discrete ($d_t = 0$) in our estimation procedure and limit theory. So the converge rate is root-$n$; for example, the average derivative in PSS89ETA and the propensity score regression estimator for multivalued treatment effects in Lee14DTT.
This section presents the asymptotic theory for nonparametric estimation of the partial mean process in Eq.$\!$ ((ref)) when the regressors $V$ are unobserved and estimated at the first step. In particular, we consider $V = v_0(T,S)$, where $S \subseteq (X,Z)$.
To distinguish the {\it regressor} and {\it argument} roles of the generated regressor in its impact on the final estimator, we introduce some notation:
where $\nabla_v$ denotes the vector differential operator with respective to $v$. $REG_{yt}$ is for using $\hat V$ as the {\it regressor} to estimate the regression function, and $ARG_{yt}$ is for using $\hat V$ as the {\it argument} of the regression function. Let $\mathcal{T}^\dagger \times \mathcal{S}^\dagger$ be an interior subsupport of $(T, S)$, defined by the trimming function $\pi_i = {\bf 1}{\{\inf_{t \in \mathcal{T}^*} f_{TV}(t, V_i) \geq c\}} = {\bf 1}{\{V_i=v_0(T_{i}, S_{i}) \in \mathcal{V}^*\}} = {\bf 1}{\{(T_{i}, S_{i}) \in \mathcal{T}^\dagger \times\mathcal{S}^\dagger\}}$. The following theorem states the main result.
The impact of the generated regressor $\Delta_{yt}(\hat v)$ is characterized by an integral operator on the estimation error $\hat v(\cdot) - v_0(\cdot)$. The Step 1 estimator $\hat v(\cdot)$ is taken as a fixed function given the sample and the integration is taken over the underlying random variables $(T, S, W)$. The integral operator enables a further derivation by plugging in a linear representation of the estimation error $\hat v - v_0$ of a specific estimator. We illustrate the stochastic expansion by kernel estimators and parametric estimators $\hat v$ in the examples in Section (ref).
We emphasize three {\it elements} of our stochastic expansion that summarize the important insight of the impact of generated regressors: (i) the partial/full mean structure of $\Delta_{yt}(\hat v)$, (ii) {\it index bias} $F_{Y|TV} - F_{Y|TS}$ and {\it index ratio} $f_{T|S}/f_{T|V}$, and (iii) the {\it projection of the weight} $\mathbb{E}[W|V]$. These {\it elements} are generic for the marginal integration estimators that involve a kernel regression with generated regressors. We discuss these elements in more detail below and illustrate by the examples in Section (ref).
The last two terms ((ref)) in the stochastic expansion from the estimation error of the Step 3 summation diminish at a root-$n$ rate. We maintain the smaller-order terms in the expansion for two reasons: First, when the partial mean is an intermediate step in a multi-step estimation procedure, our uniform expansion can be used to analyze the asymptotic property of the final estimator. In the case when the final estimator is $\sqrt{n}$-consistent, the impact of Step 3 partial mean and the impact of Step 1 generated regressors could be of first order. The second reason is to improve the asymptotic approximation in finite sample. The asymptotic variance could be estimated by the second moment of the estimated influence function in the stochastic expansion, including the terms that are asymptotically negligible.
Next we consider the Step 2 regression to be the $p$th-order local polynomial estimator. Theorem (ref) shows that the first-order asymptotic properties are the same as described in Theorems (ref) and (ref) for the local constant estimator, under an undersmoothing condition such that the leading bias is of smaller order. We apply Theorem 2 in MRS12B for bounds on the supremum norm of the approximation error and modify the corresponding regularity conditions; see the detailed discussion in Section (ref) in the Appendix.
The usefulness of the general estimation procedure and asymptotic theory is illustrated by, but not limited to, the economic examples in this section. The three {\it elements} --- partial mean, index, and projection of the weight --- critically determine the influence of estimating the generated regressors on the partial mean estimator. We allow the generated regressors to be general functions; for the examples considered in this section, the generated regressors are conditional density functions, conditional CDF, or residual from a regression. The Step 1 estimator of the generated regressors can be (semi)parametric or nonparametric, satisfying Assumption (ref) and with an asymptotic linear representation. We illustrate with kernel estimators and parametric estimators in these examples. For the parametric estimator, the weak convergence of the partial mean estimator is asymptotically equivalent to the Gaussian process as if the true generated regressor was observed. For the nonparametric estimator, we provide conditions under which the Step 1 estimation error is asymptotically ignorable. We follow the setup in Section (ref) and consider a single continuous treatment variable ($d_t =1$) and a set of treatment values of interest $\mathcal{T}^*$.
Under unconfoundedness or selection on observables assumption, the observable characteristics $X$ is a valid control variable satisfying Assumption (ref). HI04 show that regressing on the generalized propensity score (GPS) is sufficient for estimating continuous treatment effects {\it for the population}, i.e., $\mathbb{E}[Y(t) - Y(\bar t)]$ for $(t, \bar t) \in \mathcal{T}^\ast$. The key identifying assumption is that for each $t$, $Y(t)$ is independent of $T$ given $V = v_0(X) = f_{T|X}(t|X)$. We extend the results for the effects for the population in HI04 to the effects {\it for the treated} that is $\mathbb{E}[Y(t) - Y(\bar t)|T=\bar t]$ by switching the treatment value from $\bar t$ to $t$ {\it for the subpopulation with treatment value} $\bar t$. We find that to identify $\mathbb{E}[Y(t)|T=\bar t]$, we need to adjust for two generalized propensity scores --- one for the potential treatment value $t$ and one for the treated value $\bar t$. Thus the generated regressor is $V= v_0(X) = (f_{T|X}(t|X), f_{T|X}(\bar t|X))'$. We extend the kernel-based propensity score regression estimator in HIT98RES to a continuous treatment case and provide a limit theory.
We first discuss what we learn from the three elements of the stochastic expansion $\Delta_{yt}$ in Theorem (ref). (i) The GPS $f_{T|X}(t|X)$ has two regressors $(T, X)$, but $T$ is evaluated at $t$ in the estimation procedure, i.e., $f_{T|X}(t|X)$ is not a function of $T$. As a result, $\Delta_{yt}$ is a partial mean in the sense that only $X$ is integrated out and $T$ is fixed at $t$. It follows that when the GPS is estimated nonparametrically, $\Delta_{yt}$ converges at a nonparametric rate. (ii) The index ratio $f_{T|X}(t|X)/f_{T|V}(t|v_0(X))=1$ and $\nabla_v f_{T|V}(t|v)|_{v=V} = 1$ for $V = f_{T|X}(t|X)$. (iii) For the population treatment effect, $W=1$ and hence $\nabla_v \mathbb{E}[W|V=v] = 0$. For the treatment effect for the treated where $V= (f_{T|X}(t|X), f_{T|X}(\bar t|X))'$, $\mathbb{E}[W|V] = f_{T|X}(\bar t|X)/f_T(\bar t)$ and hence $\nabla_v \mathbb{E}[W|V=v] = (0,1/f_T(\bar t))'$.
Now we present the detail of the estimator for the effects for the population and its asymptotic properties. We estimate the CDF of $Y(t)$ for the subpopulation whose observable characteristics take values on a common support $\mathcal{X}^*$, using the GPS as the generated regressor $V=v_0(X) = f_{T|X}(t|X)$ for each $t \in \mathcal{T}^*$, i.e., as defined in Eq. $\!$((ref)),
where the trimming function $\pi(V) = {\bf 1}{\{ \inf_{t\in\mathcal{T}^*} f_{T|X}(t|X) \geq c \}} = {\bf 1}{\{ \inf_{t\in\mathcal{T}^*} V \geq c \}} = {\bf 1}{\{ X \in \mathcal{X}^*\}}$.\footnote{The support of $V$ conditional on $T=t$ equals the support of $X$ conditional on $T=t$ by definition. So the common support $\mathcal{V^*} = \mathcal{X}^*$ is defined the same as the case using $X$ as the control variable in Eq.$\!$ ((ref)). }
The stochastic expansion in Theorem (ref) implies the influence from estimating the GPS is expressed by an integral operator on $\hat f_{T|X}(t|\cdot) - f_{T|X}(t|\cdot)$:
We see that the influence $\Delta_{yt}$ is determined by the index bias $F_{Y|TV} - F_{Y|TX}$.
We consider two estimators for the GPS in Step 1. Step 2 and Step 3 follow the procedure described in Section (ref). We can choose the bandwidth $h \sim n^{-0.22}$ for Step 2 local linear estimator or local constant estimator.
In the following Corollary (ref)(I), we present the influence from estimating the GPS by the two estimators described above. For a nonparametric kernel estimator in Step 1(a), the influence $\Delta_{yt} = O_p((nh_1)^{-1/2})$. We also characterize the leading bias from Step 1(a). For the commonly used probit estimator in Step 1(b), the influence $\Delta_{yt} = O_p(n^{-1/2})$. Corollary (ref)(I\!I) presents the limiting property when the nonparametric estimation of the GPS is not first-order ignorable. Finally in Corollary (ref)(I\!I\!I) when the GPS is estimated parametrically or nonparametrically with larger bandwidth such that $h/h_1 \rightarrow 0$, the first-order asymptotic property is the same as if the true GPS was observed. In this case, the influence functions provided in Corollary (ref)(I) could be useful for estimating functionals of the partial means or for improving asymptotic approximation in finite sample.
Next we estimate the treatment effects for the treated. Let the generated regressor $V = v_0(X) = (f_{T|X}(t|X), f_{T|X}(\bar t|X))'$. Theorem (ref) below extends Theorems 2.1 and 3.1 in HI04.
By the same arguments in Eq.$\!$ ((ref)), Theorem (ref) implies that $\mathbb{E}[Y(t)|T=\bar t, V = (v, \bar v)] = \mathbb{E}[Y(t)|T=t, V = (v, \bar v)] = \mathbb{E}[Y |T=t, V = (v, \bar v)]$ for any $(v, \bar v) \in \mathcal{V}^* \subseteq Supp(V|T=t)\cap Supp(V|T=\bar t)$. Then we can identify the conditional mean of $Y(t)$ given $T=\bar t$, defined on $\mathcal{V}^*$, by the following,
where the trimming function $\pi(V) \equiv {\bf 1}\{ V \in \mathcal{V}^\ast\}$. This is the {\it conditional average structural function} given $T=\bar t$ or the {\it average dose response function for the treated} $\bar t$ subpopulation whose control variable takes values in $\mathcal{V}^*$. Therefore we can identify the corresponding CDF, $F^\ast_{Y(t)|T}(y|\bar t) \equiv \mathbb{E}^*\big[{\bf 1}\{Y(t) \leq y\}\big|T=\bar t\big] = F_{Y(t)}(y; V, W\pi)$, where $W = f_{T|X}(\bar t|X)/f_T(\bar t)$. We propose two partial mean estimators following the procedure in Section (ref) with an estimated weight:
Corollary (ref) and Theorem (ref) suggest that regression on the nonparametrically estimated GPS is first-order asymptotically equivalent to regressing on $X$, using the same bandwidth and kernel function for both the Step 1 and Step 2 estimators. So there is no efficiency gain in using the GPS. These results parallel the findings for the binary treatment in HIT98RES, HR13ETA, MRS12B, and the multivalued treatment effects for the treated in Lee14DTT. This is because the whole set of observables $X$ provides finer conditioning variables than its index $f_{T|X}(t|X)$ or $f_{T|X}(\bar t|X)$. When the index bias is not zero, i.e., there exists $x$ such that $F_{Y|TV}(y|t,v_{0}(x)) \neq F_{Y|TX}(y|t,x)$, the inequality between the corresponding asymptotic variances is strict. Lemma (ref) in the Appendix provides a formal and general result comparing expectations of the conditional variances given the whole set of observables $X$ and given the index $v_0(X)$, respectively.
When unconfoundedness is violated, one approach to satisfy Assumption (ref) is through control variables as in the triangular simultaneous equations model. For example, IN09ETA show the conditional distribution function of the endogenous variable given the instrumental variables is a valid control variable $V = F_{T|Z}(T|Z)$, where $Z$ is an exogenous instrumental variable. Let the bounded support of $Y$ be $[y_l, y_u]$ and the trimmed proportion $P^* \equiv {\rm Pr}(V \notin \mathcal{V}^*)$. IN09ETA identify the bounds for the average structural function $\mathbb{E}[Y(t)]$ by $\mathbb{E}^*[Y(t)] +\ y_l P^* \leq \mathbb{E}[Y(t)] \leq \mathbb{E}^*[Y(t)] +\ y_u P^*.$ For a smaller trimming threshold $c$, $P^*$ is smaller and hence the identified set is smaller. By replacing the dependent variable $Y(t)$ with ${\bf 1}{\{Y(t) \leq y\}}$, the CDF of the outcome at a hypothetical value $t$ is bounded by the following
as defined in Eq. $\!$((ref)). It follows that the lower bound of the $\tau$-quantile structural function is $y_l$ if $\tau \leq P^*$ and is ${F^{*-1}_{Y(t)}}(\tau - P^*)$ if $\tau > P^*$. The upper bound of the $\tau$-quantile structural function is ${F^{*-1}_{Y(t)}}(\tau)$ if $\tau < 1-P^*$ and is $y_u$ if $\tau \geq 1-P^*$. The partial mean estimator for the bounds of the average and quantile structural functions is described in Section (ref) with $W=1$. And $\hat P^* = 1-n^{-1}\sum_{i=1}^n \hat \pi(\hat V_i)$.
We first discuss what we learn from the three elements of the stochastic expansion in Theorem (ref). The exclusion assumption of the instrumental variable implies (ii) the index bias is zero $F_{Y|TZ}(y|t,Z) = F_{Y|TV}(y|t,v_0(t, Z))$. So Theorem (ref) implies the influence of the generated regressors for the {\it regressor} role is simplified to $REG_{yt}(z,v) = - \frac{\partial}{\partial v} F_{Y|TV}(y|t, v) \frac{f_{T|Z}(t|z)}{ f_{T|V}(t| v)}$. The integral operator in $\Delta_{yt}(\hat v)$ gives (i) a full mean structure, where the regressor $Z$ is integrated out and $T$ is in the dependent variable. It implies that the influence of the nonparametrically estimated control variable converges at a root-$n$ rate. Our stochastic expansion characterizes the terms that are asymptotically negligible for the partial mean estimator and could be used to improve finite-sample inference. Moreover, it is the building block when the average structural function is an intermediate ingredient of a more complicated estimation procedure, such as VV13, BL14, and HR16. In such cases, the final estimator are functionals of the partial mean and the influence of the estimated control variable might not be asymptotic ignorable.
In an earlier working paper of Lee, we derive a complete inference theory that accounts for the estimation error of the control variables. To conserve space in this paper, we present the detailed results in a separate paper. In particular, we consider three estimators of the control variable in Step 1. (I) a local constant estimator of the CDF $F_{T|Z}(T|Z)$ in IN09ETA and the residual $T-\mathbb{E}[T|Z]$ in NPV99ETA, where $\mathbb{E}[T|Z]$ is estimated by (I\!I) a local constant estimator and (I\!I\!I) an ordinary least squares estimator. Our results contribute to the literature on inference in triangular models, for example, NPV99ETA, SuUllah, IN09ETA, MRS12A, Stouli, among others.\footnote{ IN09ETA obtain a convergence rate of a power series estimator for the average structural function. NPV99ETA use a two-step series estimator to exploit their additive structural function and characterize the estimation error of the control variable in the asymptotic variance. For estimating the triangular model in NPV99ETA, SuUllah and MRS12A control the sampling variation of the reduced form residual to be first-order ignorable by choosing the tuning parameters. In contrast, the full mean structure of our stochastic expansion implies the sampling variation from Step 1 nonparametric estimation is of order $n^{-1/2}$ that is first-order ignorable without further restricting the tuning parameters. Stouli estimate semiparametric nonseparable triangular models. }
Often the objects of ultimate interest are policy effects or inequality measures. Such objects can be expressed as functionals of the potential outcome distributions identified by the partial mean $F_{Y(t)}(y;V, W\pi)$ and estimated in previous sections. The key to the distribution theory for a class of smooth functionals is the functional delta method for Hadamard-differentiable functionals. The results are illustrated by the mean and quantile operators. In this section, we let $\theta_t(y) = F_{Y(t)}(y;V, W\pi)$ by suppressing $(V, W\pi)$ in the notation for brevity. The corresponding asymptotic theorem derived in previous sections provides the influence function and weak convergence: denoting as $\sqrt{nh^{d_t}}\big( \hat \theta_t - \theta_t \big) = n^{-1/2} \sum_{i=1}^n \psi_{tin} + o_p(1)$ and converges weakly to a Gaussian process $\mathbb{G}_t$.
These Hadamard-differentiable functionals can be highly nonlinear functionals of the CDF, but admit a linear functional derivative. Weak convergence of the estimators will be implied by the functional delta method in empirical process theory. Assumption (ref) is a high-level assumption that could impose restrictions or smoothness on the distribution functions of potential outcomes. In particular, when $\Gamma$ is the $\tau$-quantile operator on $\theta_t(y) = F_{Y(t)}(y;V, W\pi)$, $\Gamma$ is a generalized inverse $\theta_t^{-1}: (0,1) \rightarrow \mathcal{Y}$ given by $\theta_t^{-1}(\tau)= \inf\{y: \theta_t(y) \geq \tau\}$. Then Assumption (ref) means $\theta_t(y)$ is continuously differentiable at the $\tau$-quantile, with the derivative being strictly positive and bounded over a compact neighborhood. Additional assumptions might be needed for different policy functionals. For example, B07JoE gives regularity conditions for Hadamard-differentiability of Lorenz and Gini functionals.
The above result gives the policy/inequality treatment effects of shifting the treatment from $\bar t$ to $t$, $\Gamma\big(\theta_t\big) - \Gamma\big(\theta_{\bar t}\big)$. The estimators of the distributional features at different treatment levels $t$ and $\bar t$, $\Gamma\big(\theta_t\big)$ and $ \Gamma\big(\theta_{\bar t}\big)$, are asymptotically uncorrelated.
The mean for the CDF $\theta_t$ is $\Gamma(\theta_t) = \int_{\mathcal Y} u d\theta_t(u)$, which has the Hadamard derivative $\Gamma'(\theta) = \int u d\theta(u)$. Then the estimator is $\int y\ d\hat \theta_t(y)$ by replacing the dependent variable ${\bf 1}{\{Y \leq y\}}$ with $Y$ in the estimation procedure described in Section (ref). Alternatively, we can use the transformed outcome in Remark (ref). Theorems (ref) and (ref) provide the asymptotic theory of estimating the mean $\mathbb{E}^*[Y(t)]$ by simply replacing $F_{Y|TV}(y|t, V_i)$ with $\mathbb{E}[Y|T=t, V= V_i]$ in the influence function $\psi_{tin}(y;V, W\pi)$, $ARG$, and $REG$ given in Theorem (ref).
The unconditional quantile function is inverted directly from the unconditional CDF. For the quantile process $\{Q_\tau: \tau \in (0,1)\}$ of the CDF $\theta_t$, $Q_\tau \equiv \inf \{y : \theta_t(y) \geq \tau\}$. The following corollary gives the asymptotic theory of estimating unconditional quantile function of $Y(t)$ for the whole population assuming unconfoundedness and using control variables.
The asymptotic theorems in the previous sections can be used to calculate point-wise confidence intervals. The influence function can be estimated by replacing unknown functions with consistent estimators. Then the covariance matrix can be estimated by the sample variance of the estimated influence functions. Our expansions in Theorem (ref) characterize the second-order terms, which could be included in the estimated influence function to improve asymptotic approximation. Alternatively, the covariance matrix can be estimated by a plug-in method that is a sample analogue with consistently estimated unknown functions. The procedure is standard and omitted for brevity.
Bandwidth selection is essential to implement the nonparametric estimation procedure. SuUllah propose a plug-in method by minimizing the asymptotic integrated mean squared error. The stochastic expansion in EJL is uniform over the bandwidth which allows the use of a data-driven bandwidth choice procedure. Although a bandwidth selection procedure is beyond the scope of this paper, a rule of thumb method satisfying the bandwidth Assumptions in this paper can be implemented in practice and our empirical application in Section (ref).
Besides point-wise inference, we might be interested in testing a hypothesis involving a policy on the whole distribution: constant effect or stochastic dominance. We suggest using a multiplier method to simulate the empirical processes defined in Theorem (ref). The multiplier method has been used in DHB and DH14JoE to simulate a distribution process. It is easy to perform asymptotically valid inference on distributional features defined by the Hadamard-differentiable functionals. Let $\{U_i\}_{i=1}^n$ be a sequence of $i.i.d.$ random variables with mean zero and variance one, for example, $\mathcal{N}(0,1)$, independent of the data. We assume that there exists a uniformly consistent estimator $\hat \psi_{tin}$ of the influence function $\psi_{tin}$ that is monotone in $y$ by using second-order kernel or a monotone transformation.\footnote{ This low-level assumption on $\hat \psi_{tin}$ is sufficient for the general high-level assumptions for the validity of the multiplier bootstrap in HsuMB, for example. } The following theorem shows that $n^{-1/2} \sum_{i=1}^n U_i \hat \psi_{tin}(\cdot)$ simulates the asymptotic distribution of the estimator.
We illustrate our method by evaluating the Job Corps program in the United States conducted in the mid-1990s. The largest publicly founded job training program targets disadvantaged youth. The participants are exposed to different numbers of actual hours of academic and vocational training. The participants' labor market outcomes may differ if they accumulate different amounts of human capital acquired through different lengths of exposure. We estimate the average/quantile dose response functions to investigate the relationship between employment and the length of exposure to academic and vocational training. We also estimate the dose response functions for the treated $\bar t$, i.e., for the participants who have received $\bar t$ hours of training. Therefore, we are able to estimate the quantile treatment effect for the treated and contribute to the literature that have been focusing on the binary treatment effects of Job Corps by employing a binary indicator of participation and the average continuous treatment effects for the population. As our analysis builds on FFGN12ReStat and HHLP, we refer the readers to the reference therein for the detail of Job Corps.
\paragraph{Data} We use the same dataset in HHLP. We consider the outcome variable ($Y$) to be the proportion of weeks employed in the second year following the program assignment. The continuous treatment variable ($T$) is the total hours spent in academic and vocational training in the first year. The Conditional Independence Assumption (ref) means that selection into different levels of the treatment is random, conditional on a rich set of observed covariates, denoted by $X$. The identifying Assumption (ref) is indirectly assessed in FFGN12ReStat. In addition to socio-economic characteristics at base line, $X$ includes expectations about Job Corps and interaction with the recruiters that are predictive for the duration in Job Corps. Our sample consists of 4,024 individuals who completed at least 40 hours (one week) of academic and vocational training. Table (ref) provides descriptive statistics.
\paragraph{Calculations} To calculate the average and quantile dose response functions (DRF) for the population and for the treated $\bar t$, we follow the semiparametric method described in Section (ref). The details are as follows. Construct grid points of hours in training $\mathcal{T}^\dagger \equiv \{80, 120,..., 2000\}$ and consider $t, \bar t \in \mathcal{T}^\dagger$.
To estimate the $\tau$-quantile DRF, replace the dependent variable with ${\bf 1}{\{Y \leq y\}}$ in the above procedure and solve the inverse function of the CDF estimate. Specifically, the $\tau$-quantile DRF for the population is estimated by $\hat Q_\tau(Y(t)) = \hat F^{-1}_{Y(t)}(\tau)$, where $\hat F_{Y(t)}(y) = n^{-1}\sum_{i=1}^n \hat{\mathbb{E}}[{\bf 1}{\{Y \leq y\}}|T=t, \hat V = \hat V_i]$ in Step 3(a). The $\tau$-quantile DRF for the treated $\bar t$ is estimated by $\hat Q_\tau(Y(t)|T=\bar t) = \hat F^{-1}_{Y(t)|T}(\tau|\bar t)$, where $\hat F_{Y(t)|T}(y|\bar t) = n^{-1}\sum_{i=1}^n \hat{\mathbb{E}}[{\bf 1}{\{Y \leq y\}}|T=t, \hat V = \hat V_i] \hat W_i$ in Step 3(b). In Step 4, the derivative $f_{Y(t)}(y)$ in the influence function derived in Corollary (ref) is estimated using the same procedure for $\hat F_{Y(t)}(y)$ by changing the Step 2 local linear regression to a kernel conditional density estimation $\hat f_{Y|T\hat V}(y|t, v)$.
\paragraph{Results} Figure (ref) presents the estimates of the average DRFs $\mathbb{E}[Y(t)]$ for the full sample and the subsamples of different gender and race. The estimates suggest an inverted-U relationship between the employment and the length of participation for the full sample, the female sample, and the black sample. For example, the average treatment effect of extending the length from one month to six months is about $2.5\%$ for the full sample, based on the estimate $\hat{\mathbb{E}}[Y(960)] - \hat{\mathbb{E}}[Y(160)] = 2.5$ with standard error $0.61$. We also see significant differences in the estimates across demographic groups. In particular, the average employment in the second year for blacks reaches its maximum at around 1000 hours (six mouths) in training. For example, for blacks, the average treatment effect on employment of extending the length from one month to six months is about 2.43%, based on the estimate $\widehat{\mathbb{E}}[Y(960)]-\widehat{\mathbb{E}}[Y(160)] = 2.43$ with standard error $0.74$. In contrast, the average employment for whites decreases over a longer duration in the program, e.g., more than about three months. The average treatment effect of extending the length from three months to eleven months is about $-3.24\%$ for whites, based on the estimate $\widehat{\mathbb{E}}[Y(1760)]-\widehat{\mathbb{E}}[Y(480)] = -3.24$ with standard error $1.95$. The average employment for males does not significantly change over the lengths of exposure. The average employment for females reaches its maximum at around six mouths in the program. For example, the average treatment effect of extending the length from one month to six months is about 1.54% for females, based on the estimate $\widehat{\mathbb{E}}[Y(960)]-\widehat{\mathbb{E}}[Y(160)] = 1.54$ with standard error $0.88$.
We also compute the ordinary least squares (OLS) of employment on a cubic function of the length of exposure and $X$. The OLS estimates present the same inverted-U relationship for all subsamples that could misspecify and over estimate the effects. Our semiparametric estimator is more flexible relative to OLS to capture heterogeneities in the effects of length of exposure to academic and vocational training.
Figure (ref) presents the estimates of the quantile DRFs $Q_{\tau}(Y(t))$ for $\tau = 0.5, 0.75$, for the full sample and the subsamples of different gender and race. Since the 25th percentile of $Y$ is zero and our asymptotic theory is established for the interior compact support of $Y$, we focus on the median and the 75th percentile. The median estimates show similar patterns with the average estimates presented in Figure (ref). For example, the median treatment effect of extending the length from one month to six months is about $7.69\%$ for the full sample, based on the estimate $\hat Q_{.5}(Y(960)) - \hat Q_{.5}(Y(160)) = 7.69$ with standard error $1.32$. The $75\%$-quantile DRF seems to decrease over a loner duration in Job Corps across demographic groups. For example, the $75\%$-quantile treatment effect of extending the length from one month to eleven months is about $-5.77\%$ for the full sample, based on the estimate $\hat Q_{.75}(Y(1760)) - \hat Q_{.75}(Y(160)) = -5.77$ with standard error $1.66$. For blacks, the average employment is larger than the median employment, suggesting that the distribution of employment might be skewed to the right. This is in contrary to whites, whose employment distribution might be skewed to the left.
The estimates for the average and quantile DRFs for the treated $\bar t = 400, 1000, 1760$ (i.e., around 25th, 50th, 75th percentiles of $T$, respectively) are not significantly different from the estimates for the population. Thus we do not present the results for the DRFs for the treated.
We demonstrate that our analysis uncovers heterogeneities in the effects of training in Job Corps along the different lengths of exposure. Job Corps staff may suggest participants to take additional or less training. The decreasing or inverted-U relationship might indicate the lock-in effect, due to the lack of labor market experience in the second year for participants with a longer enrollment. FFGN12ReStat have studied the lock-in effect by considering outcome that puts participants on an equal footing in terms of the time they have been in the labor market after exiting the program. To illustrate our methods and conserve space in this paper, we separate more comprehensive analysis using other labor market outcomes in another project.
We derive a stochastic expansion showing how the presence of generated regressors affects the limiting behavior of the three-step nonparametric estimator of the partial mean process defined in Eq.$\!$ ((ref)). We allow the generated regressors to be general functions and estimated by a (semi)parametric or nonparametric method. The influence of estimating the generated regressors has three important elements: partial mean structure, index, and projection of the weight. These elements provide insights on conditions under which the estimation error is ignorable. We study continuous treatment effects in nonseparable models. We provide an inference method for the bounds of the average and quantile structural functions by a fixed trimming approach. The endogeneity of the continuous treatment is corrected by control variables or the generalized propensity score.
Our stochastic expansion accounting for the estimation error of the generated regressor is uniform over the treatment value $t$ and the distributional threshold value $y$. Therefore, these results can be applied to more complicated estimation procedure where the partial mean is an intermediate step; for example, test for stochastic dominance in LLW09ETA, the average compensating variation in DB and BL14, test for overidentifying restriction in HR16, and the transformation model in VV13.