Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
104,274 characters · 18 sections · 63 citation commands
Nonparametric Estimation of Truncated Conditional Expectation Functions
A truncated sample mean is the mean calculated after discarding some of the highest and/or lowest values in a sample. Such quantities, which estimate the corresponding truncated expectations, are used in a wide range of economic applications. In studies of inequality, income dispersion can be summarized by reporting the mean income in different quintiles of its distribution, i.e., the mean income of the 20% of households with the lowest income, followed by the mean income of households between the 20th and 40th percentile of the income distribution, etc.\ Semega2020. In finance, the expected shortfall denotes the expected value of a certain proportion, e.g.\ 5%, of top losses. It is a widely-used risk measure, which informs about the performance of a portfolio of assets in the worst-case scenarios Chen2008. Truncated means are also used in settings with contaminated data, where the sharp bounds on the true expected outcome are obtained by considering the extreme scenarios in which the contaminated data points have the highest or the lowest outcomes, and by trimming the respective tails of the outcome distribution horowitz1995identification. This partial identification approach has been adapted to several impact evaluation settings to address sample selection problems; see, e.g., ZhangRubin2003, Lee2009, chen2015bounds.
In all the above examples, the analysis can be enriched by incorporating covariates. First, the anatomy of income inequality can be better understood when analyzed conditionally on characteristics such as age or work experience. Second, an estimator of the expected shortfall can be more informative if it takes into account covariates, such as past returns. Third, in impact evaluation, the heterogeneity of treatment effects can be explored based on individuals' characteristics. Furthermore, gerard2020bounds apply the trimming approach of horowitz1995identification to regression discontinuity designs with a manipulated running variable, which necessarily involve conditioning on a covariate.
In this paper, I propose a novel, nonparametric estimator of truncated conditional expectation functions. As in the above-mentioned applications, I consider setups where the outcome variable needs to be truncated above or below certain quantiles of its conditional distribution. For ease of exposition, I focus on one-sided truncation. I consider a nonparametric setting with a continuous outcome variable, denoted by $Y$, and a vector of continuous covariates, denoted by $X$.\footnote{If the covariates take on only a small number of distinct values, then the truncated conditional expectation function can be estimated using sample truncated means binned by covariate values.} For a quantile level $\eta \in (0,1)$ and $x$ in the support of $X$, let $Q(\eta,x)$ be the conditional $\eta$-quantile of $Y$ given $X=x$. The object of interest is the following function:
I refer to $\eta$ in the above definition as the truncation quantile level. It might be chosen by the analyst, in which case it is a fixed, known number, but in some applications the truncation quantile level has to be estimated from the data. The considered setting is nonparametric, meaning that only smoothness restrictions on the functions $m(\eta,x)$ and $Q(\eta,x)$ are imposed.
In this estimation problem, the function $Q(\eta,\,\cdot\,)$ is a nuisance parameter. If it was known, then based on a sample $\{(X_i,Y_i)\}_{i=1}^n$ from the distribution of $(X,Y)$, one could estimate $m(\eta,x)$ using standard nonparametric regression techniques, e.g., kernel estimators, applied to the sample restricted to observations with $Y_i \leq Q(\eta,X_i)$. Alternatively, motivated by the equivalent representation of the estimand as:
one could run a nonparametric regression with $\frac{1}{\eta}Y_i \mathds{1}\{Y_i \leq Q(\eta,X_i)\}$ as the outcome variable. Feasible versions of these two estimators, however, require estimating the function $Q(\eta, \,\cdot\,)$ in the first stage. This additional estimation step may affect the properties of the resulting estimators in a potentially complicated manner.
In order to alleviate the impact of the first-stage estimation error on the final estimator, I propose a modification of the latter approach using a conditional moment equation that is Neyman-orthogonal to the conditional quantile function neyman1979c. Specifically, the proposed estimation approach is based on the following representation of the estimand:
Compared to (ref), the conditional moment in (ref) contains an additional term that is mean-zero conditional on $X$.\footnote{The conditional moment in (ref) is the quantity of interest when the outcome variable has mass points, but, as shown in this paper, there are reasons to consider this formula even with a continuous outcome variable.} Its inclusion renders the whole expression insensitive to small perturbations of $Q(\eta,\,\cdot\,)$ in the sense that its derivative with respect to the the conditional quantile evaluated at the truth equals zero,
Such orthogonal, or locally-robust, conditional moments feature prominently in the literature in setups where a nuisance parameter has to be estimated in the first stage belloni2017program, chernozhukov2018double. The orthogonality property ensures that estimation of the nuisance parameter has no first-order effect on the asymptotic distribution of the final estimator.
Based on the conditional moment in equation (ref), my proposed estimator is constructed in two steps using local linear methods fan1996local. In the first stage, I estimate the local linear approximation of the function $Q(\eta,\,\cdot\,)$. In the second stage, I run a local linear regression with a generated outcome variable corresponding to the expression under the conditional expectation in equation (ref).\footnote{Based on the local linear methods, one can also construct estimators of truncated conditional expectations functions motivated by the expressions in (ref) and (ref). I discuss them in Online Appendix (ref).} The estimator is easy to implement, and the bandwidths for the two local linear regressions can be selected as in standard nonparametric regressions.
This paper contains two main theoretical results. First, I show that the proposed estimator is asymptotically equivalent to its infeasible analog using the true conditional quantile function. Given this result, the asymptotic properties follow from the standard theory of local linear estimation. The proposed estimator has good bias properties, and it is straightforward to adapt existing inference methods to do inference on truncated conditional expectation functions. Second, I study the asymptotic properties of my estimator when the truncation quantile level is estimated from the data. Under a high-level assumption on $\widehat{\eta}$, I derive an expansion of the proposed estimator evaluated at $\widehat{\eta}$ about the estimator evaluated at the true value $\eta$. This expansion can be used on a case-by-case basis to derive the asymptotic distribution for specific estimators $\widehat\eta$.
I apply the proposed estimator in two empirical settings. First, I estimate bounds on the local average treatment effect in regression discontinuity designs with a manipulated running variable gerard2020bounds. Second, I estimate bounds on the conditional wage effect of a job training program Lee2009. These bounds involve truncated conditional expectation functions with truncation quantile levels that need to be estimated from the data.
\paragraph*{Related Literature.} Nonparametric estimation of truncated conditional expectation functions has been extensively studied in the context of the conditional expected shortfall estimation. scaillet2005nonparametric, Cai2008, and Kato2012 propose estimators based on first-stage estimates of the conditional cumulative distribution function (c.d.f.) of the outcome variable. This estimation strategy, however, is not well-suited for estimation at points on the boundary of the support of the conditioning variables. The Nadaraya-Watson estimator of the conditional c.d.f., employed by scaillet2005nonparametric, exhibits the so-called boundary effects in that its bias is of larger order at the boundary than in the interior.\footnote{Estimation of a conditional c.d.f. can be cast as a regression problem with outcome variable $\mathds{1}\{Y_i \leq y\}$.} Cai2008 and Kato2012 use the weighted Nadaraya-Watson estimator, which is asymptotically equivalent to the local linear estimator at interior points, but, unlike the local linear estimator, it is guaranteed to yield a proper c.d.f. The weighted Nadaraya-Watson estimator, however, is not defined for boundary points. In contrast, my proposed approach is well-suited for estimation at boundary points.
linton2013estimation propose an estimator based on the orthogonal conditional moment equation in (ref). Their analysis, however, applies specifically to setups where the conditional variance of the outcome variable is infinite, which results in the first-stage local polynomial quantile estimator converging faster than the final estimator. The proof of linton2013estimation does not apply to models with finite variance of the outcome variable considered in this paper, where the first stage and the final estimator have the same rates of convergence. Their estimator is also more computationally intensive as it requires estimating a separate local polynomial quantile regression for each observation used in the second stage.
Various ways of estimating truncated conditional expectation functions have also been proposed in parametric settings. koenker1978regression, Ruppert1980, and jurevckova1984regression consider generalizations of truncated means to linear models. In the first stage, they estimate quantile regressions, and in the second stage they run a regression on a sample truncated according to the first-stage estimates. In an independent work, Barendse2020 also uses a generated outcome variable based on the orthogonal moment equation. He additionally considers efficient weighting, analogous to, possibly nonlinear, weighted least squares. dimitriadis2019joint develop a joint quantile and expected shortfall estimation framework and find estimators that can be more efficient than the simple two-stage procedure described above. The efficiency gains of dimitriadis2019joint and Barendse2020, however, are specific to parametric models.
In the above-cited papers, it is assumed that the truncation quantile level is chosen by the analyst. A setting with estimated conditional truncation quantile levels and possibly continuous covariates is studied by semenova2020better. She exploits a moment that is similar to (ref), but it includes additional terms that render the expression orthogonal also to the truncation quantile level.\footnote{This property is achieved using a specific conditional moment defining the truncation quantile level.} Her focus, however, is on integrated truncated conditional expectations, and she does not provide conditional estimates. Truncated means with estimated trimming proportions have also been studied in the unconditional case, e.g., by shorack1974random and Lee2009.
\paragraph*{Outline of the Paper.} The remainder of this paper is structured as follows. In Section (ref), I formally introduce the proposed estimator. Its asymptotic properties are studied in Section (ref). In Section (ref), I discuss inference. I present a Monte Carlo study in Section (ref). In Section (ref), I consider two empirical applications: (i) sharp regression discontinuity designs with a manipulated running variable and (ii) randomized experiments with sample selection. Section (ref) concludes.
In this section, I formally introduce the proposed estimator. To simplify the exposition, $X$ is assumed to be univariate. A natural extension for the multivariate case is presented in Online Appendix (ref). I consider estimation at a selected covariate value $x_0$. The truncation quantile level, in turn, might be known or might have to be estimated from the data.
In the first stage, I estimate the conditional $\eta$-quantile function $Q(\eta,\,\cdot\,)$. For the second-stage estimator, it suffices if $Q(\eta,\,\cdot\,)$ is estimated well for covariate values close to $x_0$. The level and slope of the function $Q(\eta,\,\cdot\,)$ at $x_0$ are estimated in a local linear quantile regression as
where $\rho_\eta(v)=v(\eta - \mathds{1}\{v \leq 0\})$ is the `check' function, $k(\cdot)$ is a kernel function, $a$ is a bandwidth, and $k_a(v)=k(v/a)/a$. Based on these estimates, for $x$ in the estimation window relevant for the second stage, $Q(\eta,x)$ is estimated with its implied local linear approximation:
In the second stage, I run a local linear regression with a generated outcome variable corresponding to the expression in (ref) to estimate the truncated conditional expectation $m(\eta,x_0)$:
where $e_1=(0,1)^\top$, $h$ is another bandwidth, and $$ \psi_i(\eta,q )= \frac{1}{\eta} \left( Y_i\mathds{1}\{ Y_i \leq q \} - q(\mathds{1}\{Y_i \leq q\}-\eta)\right). $$ If $\eta$ is not known, but an estimate $\widehat \eta$ is available, I estimate $m(\eta,x_0)$ as $\widehat{m}(\widehat\eta, x_0; a, h)$.
In this section, I introduce the assumptions and study the asymptotic properties of the proposed estimator. I use the following notation. I put $\partial_x^k m(\eta,x_0)= \frac{\partial^k}{\partial x^k}m(\eta,x)|_{x=x_0}$ and $\partial_x^k Q(\eta,x_0)= \frac{\partial^k}{\partial x^k}Q(\eta,x)|_{x=x_0}$. For positive sequences $b_n$ and $c_n$, I write $b_n \prec c_n$ if $b_n / c_n \to 0$, and $b_n \asymp c_n$ if $C_1 b_n \leq c_n \leq C_2 b_n$ for some positive constants $C_1$ and $C_2.$
I consider estimation based on independent and identically distributed (i.i.d.) data. This modeling assumption is appropriate for the microeconometric applications considered in this paper.\footnote{The asymptotic analysis could be extended to allow for dependent data satisfying the $\alpha$-mixing condition under assumptions similar to those imposed by masry1997local for the standard nonparametric regression.}
I follow the classic literature on local polynomial modeling and assume that the covariate is continuous. The density of $X$ is denoted by $f_X(x)$. The conditional distribution function of $Y$ given $X$ is denoted by $F_{Y|X}(y|x)$, and the corresponding conditional density by $f_{Y|X}(y|x)$. Subsequent assumptions involve smoothness requirements for the functions $Q(\eta,\,\cdot\,)$ and $m(\eta,\,\cdot\,)$. I adopt the following convention. For a point on the left (right) boundary of $\mathcal{X}$, I define the derivative with respect to the covariate value as the right (left) derivative at that point.
Assumption (ref) comprises standard conditions for the asymptotic analysis of the local linear quantile estimator. The smoothness assumption on $Q(\eta,x)$ is used to control the order of the bias introduced by approximating the possibly nonlinear function $Q(\eta,x)$ with its first-order Taylor expansion in $x$. The restrictions on the density $f_X(x)$ ensure that with high probability there are observations around the estimation point. The restrictions on the conditional density $f_{Y|X}(y|x)$ ensure that the conditional $\eta$-quantile function can be precisely estimated. Assumption (ref) implies also smoothness of the coefficients in the local linear approximation of the function $Q(u,\,\cdot\,)$ for quantile levels close to $\eta$, which is exploited in the analysis with an estimated truncation quantile level.
Assumption (ref) is a natural adaptation of the standard conditions for the local linear estimator in the nonparametric mean regression for estimating truncated conditional expectations. Even if the function $Q(\eta, \,\cdot\,)$ was known, a continuous second-order derivative of $m(\eta,x)$ with respect to $x$ would be required to characterize the leading bias introduced by approximating the function $m(\eta,\,\cdot\,)$ with its first-order Taylor expansion. Parts (b) and (c) are needed to obtain asymptotic normality of the proposed estimator.
The restrictions on the kernel are standard. The requirements on the bandwidths are necessary for ensuring consistency in both stages. In a preliminary study of the proposed estimator, the convergence rates of the two bandwidths are not linked, but I impose further restrictions in Theorems (ref) and (ref). All results cover an important special case where $a\asymp h$.
In this section, I study the asymptotic properties of the proposed estimator when the truncation quantile level is known. The key technical result is stated in Lemma (ref). It shows that the estimator $\widehat{m}$ is asymptotically equivalent to its infeasible analog using the true conditional quantile function.
The remainder $R(\eta,x_0;a,h)$ is driven by the estimation error from the first stage on the interval $\mathcal{X}(x_0,h)\equiv [x_0-h,x_0+h] \cap \mathcal{X}$, which is relevant for the second-stage estimator. There are two sources of this estimation error. First, the function $Q(\eta, \,\cdot\,)$ is replaced with its local linear approximation, which results in an error of order $O(h^2)$. Second, the intercept and slope of this approximation are estimated at rates $O_p(a^2+(an)^{-1/2})$ and $O_p(a+(a^3n)^{-1/2})$, respectively.\footnote{In fact, these are the only properties of the first-stage estimator required in the proof of Lemma (ref).} As a result, the estimated conditional quantile function satisfies
If $h(nh)^{-1/3}\prec a$, then $w_n\to 0$, and $R(\eta,x_0;a,h)$ is of order smaller than $O_p(w_n)$. This low sensitivity to the first-stage estimation error is obtained by construction, owing to the use of an orthogonal moment.
Lemma (ref) holds regardless of whether the variance of the outcome variable is finite or infinite. If Assumption (ref) holds in addition to the assumptions of Lemma (ref), then asymptotic normality of $\widehat m$ follows from the standard theory of local linear estimation. If the variance of the outcome variable is infinite, then the asymptotic distribution of $\widehat m$ can be obtained under alternative assumptions following the steps of linton2013estimation. I focus on the former case.
The asymptotic distribution presented in Theorem (ref) involves typical kernel constants, which differ depending on whether $x_0$ lies in the interior or on the boundary of the support of $X$, but this dependence is left implicit. Let $\mu_2=\int v^2 k(v)dv$, $\kappa_0= \int k(v)^2dv $, $\bar\mu= (\bar{\mu}_2^2 - \bar{\mu}_1 \bar{\mu}_{3})/(\bar{\mu}_2\bar{\mu}_0-\bar{\mu}_1^2)$, and $\bar\kappa= \int_0^\infty(k(v)(\bar{\mu}_1 v - \bar{\mu}_2))^2dv/ (\bar{\mu}_2\bar{\mu}_0-\bar{\mu}_1^2)^2 $, where $\bar{\mu}_{j}= \int_{0}^{\infty}v^j k(v)dv$. If $x_0$ lies in the interior of $\mathcal{X}$, I put $\mu=\mu_2$ and $\kappa= \kappa_0 $. If $x_0$ lies on the boundary of $\mathcal{X}$, I put $\mu=\bar\mu$ and $\kappa= \bar\kappa$.
As in the standard nonparametric regression, the leading bias is proportional to the second derivative of the function that is being estimated. The variance is fully analogous to the variance of the unconditional truncated mean. The additional conditions imposed on the bandwidths ensure that the remainder $R(\eta,x_0;a,h)$ is of order $o_p(h^2+(nh)^{-1/2})$. These conditions admit certain degrees of both under- and oversmoothing in the first stage relative to the second stage. For example, if $h \asymp n^{-1/5}$, then I require that $n^{-1/3}\prec a \prec n^{-1/10}$. Subject to these restrictions, the choice of the first-stage bandwidth does not affect the first-order asymptotic distribution of $\widehat m$. In practice, the two bandwidths can be set equal. This choice yields the minimal rate of the remainder $R(\eta,x_0;a,h)$.
In some applications, the truncation quantile level of interest has to be estimated from the data. In this section, I study the properties of the proposed estimator evaluated at an estimated truncation quantile level $\widehat \eta$ under a high-level condition on $\widehat \eta$.
Theorem (ref) provides an expansion of the estimator with an estimated truncation quantile level about the estimator using the true quantile level. To keep the exposition transparent, I restrict the analysis to bandwidths such that $a \asymp h$. The estimator $\widehat \eta$ is only required to converge at a rate not slower than the estimator $\widehat{m}(\eta,x_0;a,h)$ does.
The coefficient on the term $\widehat \eta - \eta$ equals the first derivative of the estimand with respect to the pre-estimated parameter, which is typical for such two-step estimation problems. It is also in line with the results of shorack1974random and Lee2009, who study the unconditional truncated mean with random trimming proportions. The expansion in Theorem (ref) can be used on a case-by-case basis to derive the asymptotic distribution of $\widehat{m}( \widehat{\eta}, x_0;a,h)$ for specific estimators $\widehat{\eta}$. Two such examples are discussed in Section (ref). I note that in Theorem (ref), it is essential that $\eta<1$, imposed in Assumption (ref). Otherwise, if $Y$ has unbounded support, the derivative $\partial_{\eta} m(\eta,x_0)$ is infinite, and the above expansion is not valid.
The asymptotic results in Section (ref) suggest that the bandwidth can be selected and inference can be conducted following methods developed for the standard nonparametric regression, ignoring the fact that the conditional quantile function is estimated in the first stage.
The bandwidth can be chosen, e.g., so as to minimize the asymptotic mean squared error, defined as $AMSE_n(h) = B(\eta,x_0)^2 h^4 + V(\eta,x_0)/(nh)$, where $B(\eta,x_0) = \frac{1}{2} \mu\, \partial_x^2 m(\eta, x_0)$. The optimal bandwidth is then given by $h_{\text{opt}}=\left(V(\eta,x_0)/(4B(\eta,x_0)^2)\right)^{1/5}n^{-1/5}$. It can be estimated following procedures analogous to those proposed by imbens2012optimal and calonico2014robust. To implement these bandwidth selectors, under additional assumptions, one can estimate $\partial^2_x m(\eta,x_0)$ using the local quadratic version of the estimator $\widehat m$, discussed in Online Appendix (ref), and estimate the asymptotic variance based on the second-stage residuals.
Given a bandwidth $h$, the asymptotic distribution in Theorem (ref) forms the basis for conducting statistical inference.\footnote{As discussed in Section (ref), the estimator is relatively insensitive to the choice of the first-stage bandwidth. In practice, one can set $a=h$.} Constructing a confidence interval (CI) requires accounting for the bias, which can be done by adapting any of the three following approaches. The first, classic approach is called undersmoothing (US). It relies on choosing a `small' bandwidth that ensures that the bias is asymptotically negligible. If $h \prec n^{-1/5}$, or equivalently $nh^5 \to 0$, then the bias is of smaller order than the standard error. As a result, an asymptotically valid $1-\alpha$ CI can be formed as
where $z_u$ is the $u$-quantile of the standard normal distribution and $\widehat{se}(h)$ is some consistent standard error. The two further approaches allow for bandwidths of order $n^{-1/5}$, such as the AMSE-optimal bandwidth.
The second approach is analogous to the robust bias corrections proposed by calonico2014robust. It involves subtracting an estimate of the leading bias term and accounting for the additional variation in the bias-corrected estimator when forming a CI. The CI takes the form as in (ref), except that a bias-corrected estimator and an adjusted standard error are used.
The third approach is motivated by the `bias-aware' approach of armstrong2020simple, who propose `honest' CIs that account for the largest possible bias under restrictions on the smoothness of the function that is being estimated. Suppose that $|\partial_x^2m(\eta,x_0)|$ is bounded by some known constant $M$. Then the leading bias term is bounded in absolute value by $\frac{1}{2}|\mu|Mh^2$, and an asymptotically valid $1-\alpha$ confidence interval can be formed as
where $\widehat{r}(h)=\frac{1}{2}|\mu|Mh^2/\widehat{se}(h)$ and $\textrm{cv}_{1-\alpha}(t)$ is the $1-\alpha$ quantile of the folded normal distribution $|\mathcal{N}(t,1)|$.\footnote{I do not discuss coverage properties uniform in the data generating processes, which would require ensuring that the remainder in Lemma (ref) is uniformly small. } One can also account for the maximal bias of the infeasible estimator $\widetilde m$ conditional on the realizations of the covariate. The bandwidth can be also chosen so as to minimize the worst-case mean squared error or the length of the CI. Implementation of bandwidth selectors and of the CIs requires imposing a bound on $\partial_x^2m(\eta,x)$. See armstrong2020simple and noack2021bias for discussions of the choice of the smoothness constant in the standard nonparametric regression.
In this section, I present simulation evidence for two claims. First, I show that the feasible estimator $\widehat{m}$ is close to the infeasible estimator $\widetilde{m}$ in terms of the mean squared difference. Second, I show that inference based on $\widehat{m}$ performs almost identically as inference based on the infeasible estimator $\widetilde{m}$. For concreteness, in this simulation study, I use the third approach discussed in Section (ref), which exploits a bound on $\partial_x^2m(\eta,x)$. In its implementation, I account for the exact worst-case bias of the infeasible estimator $\widetilde{m}$ conditional on the realizations of the covariate.
I generate data from a location-scale model of the form
where $X$ is uniformly distributed on $[-1,1]$ and $\varepsilon \sim \mathcal{N}(0,1)$. I consider three specifications for the conditional expectation function, which were used by armstrong2020simple in their Monte Carlo study comparing different inference methods. Let
where $s(x)=\max\{x,0\}^2 $. These functions are illustrated in Figure (ref). Their second derivatives are bounded in absolute value by $M=2$. I consider homoskedastic and hetersokedastic residuals, induced by functions $\sigma_1(x)=0.5$ and $\sigma_2(x)=0.5+x$, respectively.
Due to normality of the residuals, the truncated conditional expectation functions have a simple, closed-form expression. It holds that
where $\phi(\cdot)$ is the density and $q_\eta$ is the $\eta$-quantile of the standard normal distribution, respectively. With homoskedastic residuals, the truncated conditional expectation functions have the same shape as the respective conditional expectation functions, but they are shifted downwards. With heteroskedastic residuals, the slopes change as well, but this type of heteroskedasticity does not affect the curvature. Figure (ref) illustrates that for $\eta=0.8$ and $m(x)=m_1(x)$. Other cases are analogous.
In all simulations, the sample size is $n=1,000$, and the number of replications is $S=10,000$. The truncated conditional expectation functions are estimated at $x_0=0$ and three quantile levels, $\eta \in \{0.2, 0.5, 0.8\}$. I use the triangular kernel and the EHW variance estimator.
In Table (ref), I report the root mean squared error (RMSE) of the infeasible estimator $\widetilde{m}$ and the feasible estimator $\widehat{m}$, as well as the root mean squared error difference between the two. The estimators are evaluated with the RMSE-optimal bandwidth chosen for the infeasible estimator using the bandwidth selector of armstrong2020simple employing the true smoothness constant ($M=2$). In all cases, the difference between the infeasible and feasible estimators is small compared to their mean squared errors.\footnote{This qualitative conclusion remains the same when using the true bound multiplied or divided by two.} Moreover, the results are very similar in the homoskedastic and heteroskedastic settings, which shows that the estimator adapts to different slopes of the conditional quantile and truncated expectation functions very well. Depending on the specific data generating process, the feasible estimator may perform slightly better or worse than the infeasible estimator, which is due to the higher-order remainder terms in Lemma (ref).
In Table (ref), I present results regarding the bandwidth choice, empirical coverage, and length of 95% CIs. Here, I also use the true smoothness constant ($M=2$). The bandwidth selector for the feasible estimator chooses virtually the same bandwidth as would be chosen for the infeasible estimator $\widetilde m$, and the coverage is nearly identical. I note that even for the infeasible estimator, the CI based on the true smoothness constant can have coverage below the nominal confidence level despite correctly accounting for maximal bias. The reason for that is that although $Y$ is conditionally normally distributed, the outcome variable $\psi(\eta,Q(\eta,X))$ is not. The non-normality is more pronounced for lower truncation quantile levels. In Online Appendix (ref), I discuss a rule of thumb for choosing the smoothness constant that performs well in this simulation setting.
In this section, the theoretical results from Section (ref) are applied to two empirical settings: (i) sharp regression discontinuity designs with a manipulated running variable and (ii) randomized experiments with sample selection. Both applications involve evaluation of truncated conditional expectations at estimated truncation quantile levels that are given by the ratio of two densities and two conditional expectations, respectively.
gerard2020bounds study regression discontinuity (RD) designs with a manipulated running variable. They develop a complex estimation approach applicable to fuzzy RD designs, which encompass sharp RD designs as a special case. Their estimation routine involves numerical integration and inference is based on a bootstrap procedure. I study an approach tailored specifically to sharp RD designs that is easier to implement.
In a sharp RD design, units receive a treatment if and only if a special covariate, the running variable, exceeds a fixed cutoff value. If the distribution of units' potential outcomes varies smoothly with the running variable around the cutoff, then the (local to the cutoff) average treatment effect is identified by the difference in average outcomes of the treated and untreated units whose realization of the running variable is just to the right or just to the left of the cutoff, respectively lee2010regression. The key identifying assumption, however, is often questionable if the running variable is not exogenously determined.
To allow for violations of the smoothness assumption, gerard2020bounds develop a framework where there are two unobservable types of units: always-assigned units, for which the realization of the running variable is always to the right of the cutoff, and hence they are assigned the treatment; and potentially-assigned units, whose density of the running variable is smooth around the cutoff, and hence they satisfy the standard assumptions of an RD design. gerard2020bounds show that the average treatment effect for the subpopulation of potentially-assigned units at the cutoff, denoted by $\Gamma$, is partially identified. The bounds are derived as follows. Under their behavioral model, the share of always-assigned units among the units just to the right of the cutoff, denoted by $\tau$, is identified by the discontinuity in the density of the running variable, denoted by $f_X$, at the cutoff as
where $x_0$ is the cutoff value.\footnote{For a generic function $g(\cdot)$, I put $g(x_0^+)=\lim_{x\to x_0^+}g(x)$ and $g(x_0^-)=\lim_{x\to x_0^-}g(x)$.} Given $\tau$, the sharp bounds on $\Gamma$ are obtained by considering the `extreme' scenarios in which the always-assigned units have the lowest or the highest outcomes among the units just to the right of the cutoff. The resulting lower and upper bound are given by:
I discuss the main ingredients of the bounds estimator and its asymptotic properties. The details are given in Appendix (ref). The bounds $\Gamma^L$ and $\Gamma^U$ involve truncated conditional expectation functions, which I estimate using the estimator $\widehat{m}$ developed in this paper.\footnote{Estimation with truncation from below can be performed using the procedure developed for estimation with truncation from above by taking the negative of the estimator applied to the data $\{X_i, -Y_i\}^{n}_{i=1}$.} Since $\tau$ is the proportion of truncated data, the quantile level $\eta$ in the previous sections corresponds to $1-\tau$, i.e. $\eta$ is the proportion of potentially-assigned units just to the right of the cutoff. The first step is to estimate $\tau$. The density limits can be estimated using estimators such as the linear smoother of the histogram cheng1997boundary, mccrary2008manipulation, the linear smoother of the empirical density function jones1993simple, LejeuneSarda1992, or the local quadratic smoother of the empirical distribution function cattaneo2020simple.
Under regularity conditions, the resulting estimator of the truncation quantile level, $\widehat{\eta}=1-\widehat{\tau}$, satisfies the high-level assumption of Theorem (ref). Moreover, since $\widehat{\eta}$ depends only on the running variable, it is conditionally uncorrelated with the estimators of the truncated conditional expectations with known $\eta$, which simplifies the asymptotic variance formula. The conditional expectation just to the left of the cutoff, $\mathbb{E}[Y|X=x_0^-]$, can be estimated using a standard local linear estimator. The estimators of the bounds have an asymptotically normal distribution, which can be used to form confidence intervals for the partially identified treatment effect.
I evaluate the procedure that I propose by implementing it for the empirical application of gerard2020bounds.\footnote{The authors kindly implemented my procedure on their restricted-use data for comparison purposes.} They investigate the effect of unemployment insurance (UI) benefits on the formal reemployment in Brazil. They exploit the rule that a worker involuntarily laid off from a private-sector firm is eligible for the UI benefit only if there was at least 16 months between the date of her layoff and the date of the last layoff after which she applied for and drew UI benefits. This rule creates a discontinuity in the eligibility for UI benefits, which is reflected in a 70pp increase in the actual take-up of UI benefits. In the following, I focus on an intention-to-treat analysis, where the eligibility for UI benefits is the treatment, and the outcome of interest is the duration without a formal job after the layoff.
Despite the 16-month rule being rather arbitrary, gerard2020bounds point out the following ways in which violations of the standard RD assumptions may arise in this setup. Some workers may provoke their layoffs or ask their employers to report their quit as involuntary once they become eligible for a UI benefit. Other workers may have managed to delay their layoff to a date when they were eligible for the UI benefit. All theses workers are always-assigned units in the manipulation framework outlined in the previous subsection.
Figure (ref) reproduces the graphical evidence for this RD design. The running variable is the difference in days between the layoff date and the eligibility date, so that the cutoff is at 0. In the left panel, I present the density of the running variable. The share of always-assigned units is estimated to be 6.4%, which is relatively well separated from zero. This is essential for the good quality of the normal approximation of the asymptotic distribution of $\widehat{\tau}$. In the right panel, the dots represent the average outcome by day (of all observations). There is a marked jump in the mean duration without a formal job at the cutoff. I note that a substantial share, about 12--14%, of duration outcomes is censored at 24 months. This, however, does not require any adjustment in my estimation and inference procedure.
Following gerard2020bounds, I conduct two types of analysis. First, I estimate bounds on $\Gamma$ using an estimated proportion of the always-assigned units to the right of the cutoff. Second, I conduct a sensitivity analysis, where I report bounds for different levels of potential manipulation. I report my results along with the original estimates of gerard2020bounds. Their estimator is based on a local linear estimator of the conditional c.d.f., and they conduct inference via bootstrap. For comparability with gerard2020bounds, all estimators use a 30-day bandwidth, and the confidence intervals are formally justified by undersmoothing.
In Table (ref), I present estimates of the bounds and the 95% confidence intervals for $\Gamma$ with estimated $\tau$. As a reference point, the point estimate ignoring the possibility of manipulation indicates that the eligibility for UI benefits increases the duration of unemployment by about 62 days. When accounting for manipulation, however, the estimated identified set spans the range from 31 to 81 days.
In the second part of the analysis, a certain hypothetical, fixed degree of manipulation in the data is presumed. The results are presented in Figure (ref). At $\tau=0$, the results correspond to the `no manipulation case', where the treatment effect is point identified. The bounds become wider as the presumed degree of manipulation increases. The vertical black line marks the estimated proportion of always-assigned units among all units just to the right of the cutoff.
The results are nearly identical when using the procedure of gerard2020bounds and mine. This similarity, however, is specific to this dataset, where the conditional quantile functions at the truncation quantile levels are flat. I show in Appendix (ref) that compared to my estimator, approaches based on first-stage estimates of the conditional c.d.f. have an additional bias term when the conditional quantile function has a nonzero slope.
Lee2009 studies the effect of a job training program on wage rates. In his analysis, he uses covariates to narrow down the bounds on the unconditional effect semenova2020better. The conditional treatment effects, however, may be of interest in their own right.
In a randomized experiment, units are randomly split into the treatment and control groups. Treatment effects are then typically identified by comparing units in the two treatment arms. In some settings, however, such comparisons are invalid due to sample selection. For example, job training affects not only the wage rates but also the employment status. As a result, individuals whose wages are observed in the treatment and control groups are not comparable even if the treatment was random assigned.
For such settings, Lee2009 derives bounds on the wage effect for the subpopulation of always-observed individuals, i.e. those who would work regardless of whether they obtained the treatment. In the first step, he identifies the proportion of individuals whose employment status is affected by the treatment status. By random assignment to the program, this proportion is given by the difference in the employment rates in the treatment and control group. If the training program weakly encourages to work, then the bounds on the wage rates of the always-observed in the treatment group are obtained by considering the extreme scenarios in which the always-observed individuals have the highest or the lowest wage rates among the employed.\footnote{If the treatment discourages from working, then the control group would need to be truncated.} This reasoning holds unconditionally as well as conditionally on covariates.
To state these bounds formally, let $D$ be the treatment indicator and $S$ the employment indicator. Further, let $X$ be some additional covariate. The conditional proportion of individuals among the employed in the treatment group who are employed if and only if they are treated is identified as
The sharp lower and upper bounds on the local average treatment effect on wage rates are given by Lee2009
where $Q_{D=1,S=1}(u,x)$ denotes the $u$-quantile of $Y$ conditional on $D=1$, $S=1$, and $X=x$.
I discuss the main ingredients of the bounds estimator. The details are given in Appendix (ref). The conditional probabilities $\mathbb{P}[S=1|D=d,X=x]$ can be estimated using a local linear estimator with $S_i$ as the outcome and $X_i$ as a regressor, run on the sample restricted to observations with $D_i=d$ for $d \in \{0,1\}$. Under regularity conditions, the resulting estimator $\widehat{\eta}=1-\widehat{p}(x_0)$ satisfies the high-level assumption of Theorem (ref). The truncated conditional expectations in the definition of $\Delta^L(x)$ and $\Delta^U(x)$ can be estimated using the estimator proposed in this paper and the conditional expectation function in the control group can be estimated using the standard local linear estimator. Restricting the samples based on the values of indicators $S_i$ and $D_i$ does not cause any complications in the asymptotic analysis. The estimators of the bounds have an asymptotically normal distribution, which can be used to form confidence intervals for the partially identified treatment effect.
I evaluate the effect of the job training offered under the Job Corps program in the United States. I use data from the National Job Corps Study conducted in mid 90s. I follow Lee2009 closely in terms of the sample definition. The individuals who applied to the program were followed for four years after random assignment. There are 3599 individuals in the control group and 5546 in the treatment group, giving a total of 9145 observations. I investigate the effect on wage rates four years after the random assignment, conditioning on the usual weekly earnings at the most recent job reported at the baseline.
The results are presented in Figure (ref). The bandwidth is selected based on smoothness constants calibrated through the procedure described in Online Appendix (ref). The point estimates indicate that the treatment encourages taking up employment. The bounds on the treatment effect on wage rates are relatively flat for low weakly earnings at the baseline, where they are very similar to the unconditional estimates of Lee2009. I note that there is a mass point in the distribution of the covariate at zero, but this does not invalidate the results.
I propose a nonparametric estimator of truncated conditional expectation functions based on an orthogonal conditional moment and local linear methods. When the truncation quantile level is known, I show that the proposed estimator is asymptotically equivalent to its infeasible analog that uses the true conditional quantile function, and I find its asymptotic distribution. I also consider estimation with an estimated truncation quantile level. The proposed estimator is applied in two empirical settings: sharp regression discontinuity designs with a manipulated running variable and randomized experiments with sample selection.