EconBase
← Back to paper

Testing for Unobserved Heterogeneous Treatment Effects with Observational Data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

60,397 characters · 11 sections · 90 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Testing for Unobserved Heterogeneous Treatment Effects with Observational Data

\def\spacingset#1{ {#1}} \spacingset{1}

\thispagestyle{empty}

abstractUnobserved heterogeneous treatment effects have been emphasized in the recent policy evaluation literature heckman2005structural. This paper proposes a nonparametric test for unobserved heterogeneous treatment effects in a treatment effect model with a binary treatment assignment, allowing for individuals' self-selection to the treatment. Under the standard local average treatment effects assumptions, i.e., the no defiers condition, we derive testable model restrictions for the hypothesis of unobserved heterogeneous treatment effects. Also, we show that if the treatment outcomes satisfy a monotonicity assumption, these model restrictions are also sufficient. Then, we propose a modified Kolmogorov-Smirnov-type test which is consistent and simple to implement. Monte Carlo simulations show that our test performs well in finite samples. For illustration, we apply our test to study heterogeneous treatment effects of the Job Training Partnership Act on earnings and the impacts of fertility on family income, where the null hypothesis of homogeneous treatment effects gets rejected in the second case but fails to be rejected in the first application.

{\it Keywords:} endogenous treatment assignment, local average treatment effects, nonseparable model, unobserved heterogeneous treatment effects

{$^*$Author of correspondece: [email removed]. \\ Acknowledgment: The authors are grateful to the co-editor Yoon-Jae Whang, three anonymous referees, Jason Abrevaya, Qi Li, Xiaojun Song and Quang Vuong for valuable comments and suggestions on previous versions of the paper. Yu-Chin Hsu gratefully acknowledges the research support from Ministry of Science and Technology of Taiwan (MOST2628-H-001-007, MOST110-2634-F-002-045), Academia Sinica Investigator Award of Academia Sinica (AS-IA-110-H01), and Center for Research in Econometric Theory and Applications (110L9002) from the Featured Areas Research Center Program within the framework of the Higher Education Sprout Project by the Ministry of Education of Taiwan. Ta-Cheng Huang is indebted to Qi Li for his continued inspiration, guidance and support. Haiqing Xu would like to dedicate this paper to the memory of Professor Halbert White. }

\spacingset{1.2}

Introduction

\fontsize{11}{16pt}\selectfont Heterogeneous treatment effects due to unobserved latent variables have been emphasized in the policy evaluation literature. See e.g., IA_1994_ECMA, heckman1997making, heckman2001policy, heckman2005structural, abadie2002instrumental, abadie2002bootstrap,abadie2003semiparametric, blundell2003endogeneity , matzkin2003nonparametric, Chesher_2003_ECMA,Chesher_2005_ECMA, CH_2005_ECMA, florens2008identification, imbens2009identification, frolich2013unconditional, d2015identification, and torgovitsky2015identification, among many others. In the empirical study of treatment effects using observational data, the interpretation of the widely used instrumental variable (IV) estimation relies on the key assumption that after we control for covariates, treatment effects are homogeneous across individuals. In the presence of unobserved heterogeneous treatment effects, the standard IV approach only estimates the local average treatment effects (LATE), rather than the average treatment effects (ATE)\replaced[id=TC]{; see}{. See} IA_1994_ECMA and imbens2010better.

In this paper, we develop a nonparametric test for the (unobserved) heterogeneous treatment effects. We model the unobserved heterogeneous treatment effects by a nonparametric and nonseparable model, i.e., the error terms are not additively separable from the treatment indicator. Together with the endogeneity issue introduced due to individuals' self-selection to treatment, it is well known in the literature that identification and estimation of nonseparable models are challenging. On the other hand, the homogeneous treatment effects assumption substantially simplifies econometrics analysis of treatment effects, since it implies that ATE is the same as LATE, after controlling for observed heterogeneity (i.e., covariates). For instance, angrist1991does use a two-stage least squares approach to estimate treatment effects of compulsory schooling on earnings. Therefore, if ATE is the main object of research interest, by testing for heterogeneous treatment effects, our method can assess whether the complicated nonparametric and nonseparable treatment effect model is more appropriate (than e.g., the two-stage least squares approach) for a program evaluation assignment.

Though important, there are only a handful of papers on testing for such unobserved heterogeneity.\footnote{There exists a substantial literature for testing observed heterogeneity, i.e., whether (conditional) average treatment effects vary across different subpopulations defined by observed covariates. For example, see heckman1997making, crump2008nonparametric, chang2015nonparametric, abrevaya2015estimating, athey2016recursive, hsu2017consistent, and lee2017doubly, among many others.} In the context of ideal social experiment data, i.e., lack of endogeneity, heckman1997making develop a lower bound for the variance of heterogeneous treatment effects, thereby providing a test for whether the data are consistent with the identical treatment effects model. Moreover, Hoderlein2009on discuss specification tests for endogeneity as well as unobserved heterogeneity in nonseparable triangular models. Recently, LW_2014_JoE and STU_2014_ER establish nonparametric tests for unobserved heterogeneous treatment effects under the unconfoundedness assumption. In particular, LW_2014_JoE test unobserved heterogeneity in treatments effects via testing an equivalent independence condition on observables. Another closely related paper is by HSU_2010_JoE, who test the absence of self-selection on the gain to treatment in the generalized Roy model framework, allowing for unobserved heterogeneous treatment effects. Furthermore, our paper is also related to heckman2005structural, who suggest an approach to test heterogeneity of the {\it marginal treatment effects} (MTE). Our test object of interest, however, focuses on whether there exists individual-level unobserved heterogeneity in treatment effects, rather than group-level (defined by a margin) unobserved heterogeneity, i.e., whether MTE varies across margins.

Motivated by LW_2014_JoE, we show that in the presence of endogeneity, model restrictions arising from the homogeneous treatment effects hypothesis can also be characterized by a set of independence conditions that involve LATE. These testable implications are related to important literature on testing whether the distributions of potential outcomes are affected by the treatment. In the LATE framework, abadie2002bootstrap considers the null hypothesis of the equality between distribution functions of the potential outcomes among the compliers in treatment and control groups, and also first-order and second-order stochastic dominance of the two distribution functions. lee2009nonparametric and chang2015nonparametric generalize abadie2002bootstrap's test for conditional distributional effects by allowing for observed treatment effects heterogeneity.\footnote{See also e.g., jun2016treatment and hsu2017consistent for further extensions, and references therein.} Our test problem differs from that literature in that we investigate whether the distribution function of the treatment group is an (unknown) constant shift from the control group's distribution. Specifically, the equality hypothesis on the two distribution functions is a special case of our test implication.

Nonparametric tests for conditional independence restrictions have been well studied in different contexts. See, e.g., andrews1997conditional, dauxois1998nonlinear, su2007consistent, su2008nonparametric, su2014testing, huang2010testing, bouezmarni2014nonparametric, hoderlein2012nonparametric, linton2014testing, and huang2015flexible, among many others. When one tests independence restrictions of variables that are nonparametrically constructed, a key technical issue arises in the case, in particular, in which the nonparametric components are functions of continuous covariates LW_2014_JoE. Motivated by stinchcombe1998consistent, we modify the classic Kolmogorov-Smirnov tests by using the primitive function of CDF's. Such a modification is novel and plays a key role in our approach. Moreover, we establish the asymptotic properties of the proposed tests under the null and alternative hypotheses.

The remainder of this paper is organized as follows. In Section (ref), we introduce the model and derive testable model restrictions. Section (ref) discusses our test statistics and their asymptotic results. We distinguish whether the covariates include continuous variables. In Section (ref), we conduct Monte Carlo experiments to study the finite-sample performance of the proposed test. Section (ref) illustrates our testing approach by two empirical applications. All proofs are collected in the Appendix.

Model and Testable Restrictions

We consider the following nonseparable treatment effect model:

equation[equation omitted — 47 chars of source]

where $Y \in \mathrm{R}$ is the outcome variable, $D \in \{0,1\}$ denotes the treatment status, $X\in\mathrm R^{d_X}$ is a vector of covariates, $\epsilon$ is an unobserved random disturbance of general form (e.g., without invoking any restriction on the dimensionality of $\epsilon$), and $g$ is an unknown but smooth function defined on $\{0,1\}\times \mathcal S_{X\epsilon}$.\footnote{For a generic random vector $A$, $\mathcal{S}_A$ denotes the support of $A$.} In particular, the treatment variable $D$ is allowed to be correlated with $\epsilon$ to allow for selection to the treatment; see, e.g., heckman1997making. To deal with endogeneity, we introduce a binary instrumental variable $Z \in \{0,1\}$. Throughout the paper, we use uppercase letters to denote random variables, and their corresponding lowercase letters to stand for realizations of random variables.

As is motivated in the seminal paper by matzkin2003nonparametric, the non-additivity of the structural relationship $g$ in $\epsilon$ captures the idea of unobserved heterogeneous treatment effects in that the individual treatment effect, $g(1,X,\epsilon)-g(0,X,\epsilon)$, would depend on the unobserved individual heterogeneity $\epsilon$, even after controlling for covariates $X$. Therefore, we have the following proposition.

propSuppose (ref) holds, then the homogeneous treatment effects hypothesis, i.e., for some measurable function $\delta(\cdot):\mathcal{S}_X \mapsto \mathrm{R}$, \begin{equation} \mathcal{H}_{0}:\ g(1,X,\cdot) - g(0,X,\cdot)=\delta(X) \end{equation} holds if and only if the structural relationship $g$ is additively separable in $\epsilon$ (w.r.t. $D$), i.e., \begin{equation} g(D,X, \epsilon) = m(D,X) + \nu(X, \epsilon), \end{equation}where $m: \mathcal{S}_{DX} \mapsto \mathrm R$ and $\nu:\mathcal{S}_{X\epsilon} \mapsto \mathrm R$.

Proposition (ref) directly follows LW_2014_JoE. Note that if (ref) holds, $\delta(x) = m(1,x)-m(0,x)$ in (ref), which is the homogenous individual treatment effects across individuals with covariates $X=x$.

The key insight in LW_2014_JoE is that they further show \deleted[id=TC]{that}the equivalence between the additive separability hypothesis and a conditional independence restriction on observables. In the presence of treatment endogeneity, we derive a similar result. Let $p(x,z) = \mathbf{P} (D=1|X=x,Z=z)$ be the {\it propensity score} for each $x\in\mathcal S_X$ and $z \in \{0,1\}$.

assmpSuppose $Z \perp \!\!\!\perp \epsilon|X$. For all $x \in \mathcal{S}_X$, $\mathbf{P}(Z=1|X=x)$ is bounded away from zero and one, and $p(x,0) \neq p(x,1)$.

Assumption (ref) is standard in \added[id=TC]{the} literature\replaced[id=TC]{ and}{, which} requires the instrumental variable $Z$ to be (conditionally) exogenous and relevant. See, e.g., IA_1994_ECMA and CH_2005_ECMA. Throughout, we maintain Assumption (ref). Moreover, let $\mu(x,z) = \mathbf{E}(Y|X=x,Z=z)$. Under $\mathcal{H}_0$ and Assumption (ref), we have

align*[align* omitted — 186 chars of source]

In the above system of equations, we treat $\mathbf{E}\left[g(0,X,\epsilon)|X=x\right]$ and $\delta(x)$ as two unknowns. Solving the equations, we then identify LATE $\delta(x)$ as follows:

equation[equation omitted — 135 chars of source]

See IA_1994_ECMA for the LATE interpretation of (ref). Note that $\delta(x)$ is well defined given $p(x,0)\neq p(x,1)$ under Assumption 1, and identified as well directly from the data regardless \added[id=TC]{of} the monotonicity of the selection.

Let $W \equiv Y + (1-D) \times \frac{\text{Cov}(Y,Z|X)}{\text{Cov}(D,Z|X)}$. Under the null hypothesis $\mathcal{H}_0$, we have \[ W = g(D,X,\epsilon) + (1-D) \times [g(1,X,\epsilon)-g(0,X,\epsilon)]=g(1,X,\epsilon) \] which implies that $W$ is conditionally independent of $Z$ given $X$ under Assumption (ref). Therefore, we obtain the following lemma.

lemSuppose (ref) and Assumption (ref) hold. Then, $\mathcal{H}_0$ implies that $W \perp \!\!\! \perp Z | X$. On the other hand, if $W \perp \!\!\! \perp Z | X$, then the observed data can be rationalized by a structure that satisfies $\mathcal{H}_0$.

Lemma (ref) shows that the conditional independence condition is all the testable restrictions of $\mathcal{H}_0$, i.e., it is sharp in the sense of Definition 1 of HLS2019. Regarding the first part of Lemma (ref), intuitively, if treatment effects are homogeneous, we can estimate them by the IV method, and further construct potential outcomes that are independent of the instrumental variable.\footnote{Note that one could also define $W^a=Y -D\times \delta(X)$, which \replaced[id=TC]{is equal}{equals} to $g(0,X,\epsilon)$ under Assumption (ref).} Note that the conditional independence condition in Lemma (ref) can be equivalently rewritten as

multline*[multline* omitted — 207 chars of source]

provided \deleted[id=TC]{by}$p(X,1)\neq p(X,0)$ almost surely. Under the additional monotonicity condition on the selection, both sides in the above equation can be interpreted as the conditional distribution of “potential outcomes” given the compliers group in IR_1997_ReStud.

According to Lemma (ref), rejecting $W \perp \!\!\! \perp Z | X$ suggests the presence of unobserved heterogeneous treatment effects. It is worth pointing out, however, that the other direction of the above statement is also true under additional assumptions given below. These additional assumptions have been widely used for obtaining identification of quantile treatment effects, and LATE in the IV literature IA_1994_ECMA,CH_2005_ECMA.

assmp[Single-index error term] There exists a measurable function $\tilde g: \mathcal{S}_{DX}\times \mathrm R\mapsto \mathrm{R}$ and $\nu:\mathcal{S}_{X\epsilon} \mapsto \mathrm {R}$ such that \[ g(D,X,\epsilon) = \tilde g(D,X,\nu(X,\epsilon)). \] Moreover, $\tilde g(d,x,\cdot)$ is strictly increasing in the scalar-valued index $\nu$.

Assumption 2 imposes the monotonicity of the single-index error term, of which various simplified assumptions have also been made in the literature for identification and estimation of nonseparable functions. For instance, among many others, matzkin2003nonparametric and Chesher_2003_ECMA assume that the structural function $g$ is strictly increasing in the scalar-valued error term $\epsilon$. Note that Assumption 2 holds under the null hypothesis $\mathcal H_0$ because (ref) would hold under $\mathcal H_0$. Assumption 2 narrows down the space of alternatives such that the model restrictions derived in Lemma (ref) are also sufficient to distinguish the null and alternative hypotheses.

assmp[Monotone selection] The selection to the treatment is given by \begin{equation} D = \1 \left[ \theta(X, Z) - \eta \geq 0 \right], \end{equation} where $\theta$ is an unknown function, and $\eta \in \mathrm {R}$ is an error term satisfying $Z \perp \!\!\! \perp (\epsilon, \eta)|X$.

IA_1994_ECMA first introduce the monotone selection assumption, which is essentially the “no defier” condition. Moreover, Vytlacil_2002_ECMA shows that such a monotonicity condition is observationally equivalent to the weak monotonicity of (ref) in the error term $\eta$.

For any $x\in \mathcal{S}_X$, let $\mathcal C_x \equiv \{ \eta \in \mathrm{R}: \min\{\theta(x,0),\theta(x,1)\}<\eta \leq \max\{\theta(x,0),\theta(x,1)\}\}$. Note that $\mathcal C_x$ is called the “complier group” if $p(x,0)<p(x,1)$ IA_1994_ECMA

assmpThe support of $g(d,x,\epsilon)$ given $X=x$ and the complier group $\mathcal C_x$ is equal to the support of $g(d,x,\epsilon)$ given $X=x$, i.e., $\mathcal{S}_{g(d,x,\epsilon) | X=x, \eta \in \mathcal C_x} = \mathcal{S}_{g(d,x,\epsilon)|X=x}$.

Assumption (ref) is a support condition, first introduced by vuong2014counterfactual as the effectiveness of the IV. It implies that $\mathcal{S}_{g(d,x,\epsilon) | X=x, \eta \in \mathcal C_x} = \mathcal{S}_{Y|D=d,X=x}$.\footnote{To see this, note that $\mathcal{S}_{g(d,x,\epsilon) | X=x, \eta \in \mathcal C_x} \subseteq \mathcal{S}_{g(d,x,\epsilon)|D=d,X=x}\subseteq \mathcal{S}_{g(d,x,\epsilon)|X=x}$.} Note that the distribution of $g(d,x,\epsilon)$ given $X=x$ and $\eta\in\mathcal C_x$ can be identified; see, e.g., IR_1997_ReStud. Thus, Assumption 4 is testable. Specifically, for all $t\in\mathrm R$,

multline*[multline* omitted — 206 chars of source]

from which we can identify the support $\mathcal S_{g(d,x,\epsilon)|X=x,\eta\in\mathcal C_x}$.

Assumption (ref) allows one to use the data to address questions involving counterfactuals of outcomes of the “always-takers” and the “never-takers” groups. It is possible to provide sufficient primitive conditions for Assumption (ref). For instance, if one assumes $\mathcal{S}_{\epsilon | X=x, \eta \in \mathcal C_x} = \mathcal{S}_{\epsilon|X=x}$, or even a stronger condition that $(\epsilon,\eta)$ has a rectangular support conditional on $X=x$, then Assumption (ref) holds.

thmSuppose that (ref) and Assumptions (ref)-(ref) hold. Then $\mathcal{H}_0$ holds if and only if $W \perp \!\!\! \perp Z|X$.

Recall the definition of $W$;Theorem (ref) shows that testing the null hypothesis $\mathcal{H}_0$ should just rely on the information from the compliers group. It is worth pointing out that Theorem (ref) is related to LW_2014_JoE, who show that $\mathcal{H}_0$ holds if and only if $Y-\mathbf{E} (Y|D,X) \perp \!\!\!\perp D|X$ under the unconfoundedness condition (i.e., $D\perp \!\!\!\perp \epsilon|X$) and Assumption (ref).

It should also be noted that Assumption 2 is a crucial condition for the equivalence between the null hypothesis $\mathcal{H}_0$ and the testable model restrictions $W\bot Z|X$. CH_2005_ECMA show that this assumption is observationally equivalent to the rank similarity assumption. In the current literature, Assumptions 2 (or the rank similarity assumption) has been widely used for identifying heterogeneous treatment effects. See e.g.\ CH_2005_ECMA and vuong2014counterfactual.

We also note that throughout, we maintain the validity of the instrument, i.e., Assumption (ref). If this assumption is questionable, then our test should be more carefully interpreted as a joint test of Assumption (ref) and the homogeneity of treatment effects.

Consistent tests

Based on Theorem (ref), we now propose tests for unobserved treatment effect heterogeneity via testing the conditional independence restriction. Because $Z$ is binary, the conditional independence restriction in Theorem (ref) is equivalent to \[ F_{W|XZ}(\cdot|x,0)=F_{W|XZ}(\cdot|x,1), \ \forall \ x\in\mathcal S_X. \] Note that the variable $W$ needs to be nonparametrically constructed from the data. In the following discussion, we distinguish the cases where the covariates $X$ are continuous random variables because the continuous-covariates case is more difficult to deal with due to the nonparametric function $\delta(\cdot)$ in the construction of $W$.

Case 1: Discrete Covariates

We first discuss the case where $X$ takes only a finite number of values. Let $\{(Y_i,D_i,X_i, Z_i): i\leq n\}$ be a random sample of $(Y,D,X,Z)$. By Theorem (ref), we test $\mathcal{H}_0$ via the following model restrictions: \[ F_{W|XZ}(\cdot\ |x,0) = F_{W|XZ}(\cdot\ |x,1), \ \forall \ x\in\mathcal S_X, \]where $W = Y + (1-D)\times \delta(X)$ is generated from the observables.

For a generic $k$-dimensional random vector $(A_1,\cdots, A_k)$, we let $\1_{A_1\cdots A_k}(a_1,\cdots,a_k) \equiv \1(A_1=a_1, \cdots, A_k=a_k)$ and $\1(\cdot)$ be the indicator function. We estimate $\delta(X_i)$ as follows \[ \hat{\delta}(x) = \frac{ \sum_{ i=1}^n Y_i \1_{X_iZ_i}(x,1) \times \sum_{ i=1}^n \1_{X_i}(x) - \sum_{ i=1}^n Y_i \1_{X_i}(x) \times \sum_{ i=1}^n \1_{X_iZ_i}(x,1) }{ \sum_{ i=1}^n D_i \1_{X_iZ_i}(x,1) \times \sum_{ i=1}^n \1_{X_i}(x) - \sum_{ i=1}^n D_i \1_{X_i}(x) \times \sum_{ i=1}^n \1_{X_iZ_i}(x,1) }, \] and \replaced[id=TC]{further let}{let further} $\widehat W_i = Y_i + (1-D_i)\times \hat\delta(X_i)$. We now define our test statistic: \[ \widehat{\mathcal{T}}_n = \sup_{(w,x) \in \mathcal{S}_{WX}} \sqrt{n}\ \left| \widehat{F}_{{W}|XZ}(w|x,0) - \widehat{F}_{{W}|XZ}(w|x,1) \right|, \] where $\widehat{F}_{{W}|XZ}(w|x,z) = \frac{\sum_{i=1}^n \1(\widehat{W}_i \leq w) Z_i \1_{X_i}(x)}{\sum_{i=1}^n Z_i \1_{X_i}(x)}$.

Next, we establish the asymptotic properties of the test \replaced[id=TC]{statistic}{statistics} $\widehat{\mathcal{T}}_n $. Let \[ f_{WD|XZ}(w,d|x,z) \equiv f_{W|DXZ}(w|d,x,z) \times \mathbf{P}(D=d|X=x,Z=z), \] and \[ \kappa(w,x)\equiv - \frac{f_{WD|XZ}(w,0|x,1)-f_{WD|XZ}(w, 0|x,0)}{p(x,1)-p(x,0)}. \] Note that, under Assumptions (ref) and 3, $\kappa(w,x) \geq 0$ since it becomes the conditional density of $g(0,x,\epsilon)$ given the complier group and $X=x$.

assmpAssume that \begin{enumerate}[(i)] • $X$ is discrete and takes a finite number of values and $P(X=x,Z=z)>0$ for all $(x,z)\in\mathcal{S}_{XZ}$; • $p(x,1)-p(x,0)>0$ for all $x\in\mathcal{S}_X$; • $W$ has a compact support and $\partial f_{WD|XZ}(w,0|x,z)/\partial w$ is bounded above uniformly over $(w,x,z)\in\mathcal{S}_{WXZ}$. \end{enumerate}

Moreover, let

align[align omitted — 519 chars of source]

We now derive the asymptotic behavior of the test statistic.

thmSuppose that (ref) and Assumptions (ref)-(ref) hold. Then, under $\mathcal{H}_0$, \[ \widehat{\mathcal{T}}_n \overset{d}{\rightarrow} \sup_{(w,x) \in \mathcal{S}_{WX}} | \mathcal{Z}(w,x) |, \] where $\mathcal{Z}(\cdot,\cdot)$ is a mean-zero Gaussian process with a covariance kernel: \[ \text{Cov} \left[ \mathcal{Z}(w,x), \mathcal{Z}(w',x') \right] = \mathrm E \left[ (\psi_{wx} + \phi_{wx}) (\psi_{w'x'} + \phi_{w'x'})\right], \ \forall (w,x),(w',x') \in \mathcal{S}_{WX}. \] Moreover, under $\mathcal{H}_1$, we have \[ n^{-\frac{1}{2}}\widehat{\mathcal T}_n\overset{p}{\rightarrow} \sup_{(w,x) \in \mathcal{S}_{WX}} \left\vert {F}_{{W}|XZ}(w|x,0) - { F}_{{W}|XZ}(w|x,1) \right\vert >0. \]

In Appendix, we show that the influence function for $\sqrt{n}[\widehat{F}_{{W}|XZ}(w|x,0) - \widehat{F}_{{W}|XZ}(w|x,1)-{F}_{{W}|XZ}(w|x,0) + {F}_{{W}|XZ}(w|x,1)]$ is $\psi_{wx} + \phi_{wx}$, in which $\psi_{wx}$ is the influence function when $\delta(x)$ is known, and $\phi_{wx}$ accounts for the estimation effect of estimating $\delta(x)$. By Theorem (ref), our test is one-sided: reject $\mathcal{H}_0$ at significance level $\alpha$ if and only if $\widehat{\mathcal T}_n> c_{\alpha}$, where $c_\alpha$ is the $(1-\alpha)$-th quantile of $\sup_{(w,x) \in \mathcal{S}_{WX}} |\mathcal{Z}(w,x)|$.

Since the asymptotic distribution of $\sup_{(w,x) \in \mathcal{S}_{WX}} |\mathcal{Z}(w,x)|$ is non-pivotal and complicated, we apply the multiplier bootstrap method to approximate the entire process to construct the critical value. See, e.g., van1996weak, delgado2001significance, donald2003ecma, and DonaldHsu2014. Specifically, we simulate a sequence of i.i.d.\ pseudo-random variables $\{U_i: i=1,\cdots,n\}$ that is independent of the random sample path $\{(Y_i,X_i,D_i,Z_i): i=1,2,\cdots\}$ with $\mathrm E[U]=0$, $\mathrm E[U^2]=1$, and $\mathrm E[U^4] < +\infty$. Then, we obtain the following simulated empirical process: \[ \widehat{\mathcal{Z}}^u(w,x) = \frac{1}{\sqrt{n}} \sum_{i=1}^n U_{i} \times ( \hat{\psi}_{wx,i} + \hat{\phi}_{wx,i}), \] where $\hat{\psi}_{wx,i} + \hat{\phi}_{wx,i}$ is the estimated influence function. Namely,

align*[align* omitted — 865 chars of source]

where $\widehat{\mathbf{P}}(X=x,Z = z)$ and $\hat{\kappa}(w,x) = -\frac{\hat{f}_{{W}D|XZ}(w,0|x,1)-\hat{f}_{{W}D|XZ}(w, 0|x,0)}{\hat{p}(x,1) - \hat{p}(x,0)}$ are uniformly consistent nonparametric estimators for ${\mathbf{P}}(X=x,Z = z)$ and ${\kappa}(w,x)$, respectively. For a given significance level $\alpha$, the critical value $\hat{c}_{n}(\alpha)$ is obtained as the $(1-\alpha)$-quantile of the simulated distribution of $\sup_{(w,x) \in \mathcal{S}_{WX}} \big| \widehat{\mathcal{Z}}^u(w,x)\big|$.

Now, we give additional conditions for the validity of the multiplier bootstrap critical value.

assmpAssume that $\{U_i: i=1,\cdots,n\}$ is a sequence of i.i.d.\ pseudo random variables that is independent of the random sample path $\{(Y_i,X_i,D_i,Z_i): i=1,2,\cdots\}$ with $\mathrm E[U]=0$, $\mathrm E[U^2]=1$, and $\mathrm E[U^4] <\infty$.

In simulations and empirical studies, we set $U_i$'s as standard normals so that Assumption (ref) is satisfied.

assmpAssume that for $z=0,1$, \begin{enumerate}[(i)] • $\sup_{x\in\mathcal{S}_{X}}|\widehat{\mathbf{P}}(X=x,Z = z)-\mathbf{P}(X=x,Z = z) |\stackrel{p}{\rightarrow} 0$; • $\sup_{x\in\mathcal{S}_{X}}|\hat{p}(x,z)-p(x,z) |\stackrel{p}{\rightarrow} 0$; • $\hat{f}_{{W}D|XZ}(w,0|x,z)$ is continuous in $w$ for all $x\in\mathcal{S}_{X}$ and \\ $\sup_{(w,x)\in\mathcal{S}_{WX}} |\hat{f}_{{W}D|XZ}(w,0|x,z)-f_{{W}D|XZ}(w,0|x,z)|\stackrel{p}{\rightarrow} 0$; • $\sup_{x\in\mathcal{S}_{X}}|\hat{\delta}(x)-{\delta}(x)|\stackrel{p}{\rightarrow} 0$. \end{enumerate}
thmSuppose Assumptions (ref)-(ref) hold. Then \begin{enumerate}[(a)] • under $H_0$, $\lim_{n\rightarrow \infty}P(\widehat{\mathcal{T}}_n \geq \hat{c}_{n}(\alpha))=\alpha$; • under $H_1$, $\lim_{n\rightarrow \infty}P(\widehat{\mathcal{T}}_n \geq \hat{c}_{n}(\alpha))=1$. \end{enumerate}

Theorem (ref) shows the size and power of our test for the discrete case. The proof of Theorem (ref) follows standard arguments once we establish the validity of the multiplier bootstrap for the processes in that $\widehat{\mathcal{Z}}^u(\cdot,\cdot)\Rightarrow \mathcal{Z}(\cdot,\cdot)$ conditional on the sample path $\{(Y_i,X_i,D_i,Z_i): i=1,2,\cdots\}$ with probability approaching one. Assumptions (ref) contains high-level conditions and we provide estimators and give low-level conditions in Appendix. Please see the discussion after the proof of Theorem (ref).

By Assumption (ref), $X$ is assumed to be a discrete random variable taking a finite number of values. In this case, we literally conduct the test by sample splitting in that we test the equality of conditional distributions over all subpopulations defined by covariate value $x$. As a result, it is straightforward to extend our test to the case of discrete random vector $X$. Therefore, we omit the details for brevity.

Case 2: Continuous Covariates

We now consider the case where $X$ is a vector of continuous covariates. To extend the empirical process argument used in the proof of Theorem (ref) to this case, we propose a modified Kolmogorov-Smirnov test statistic. Such a modification allows us to account for the estimation effects from the generated variable $W$ which is constructed from the unknown function $\delta(\cdot)$ as an infinite-dimensional parameter.

Let $\lambda(t)= -t \times \mathrm 1 (t\leq 0)$ and $\Pi(w|x,z)= \mathrm E [\lambda(W-w)|X=x, Z=z]$. Note that $\lambda(\cdot)$ is continuous and has a directional derivative. By \deleted[id=TC]{the}definition, $\Pi(\cdot |x,z) $ is the primitive function of the $F_{W|XZ}(\cdot|x,z)$, i.e., \[ \frac{\partial }{\partial w} \Pi(w|x,z) = F_{W|XZ}(w|x,z). \] Thus, the model restriction $W \perp \!\!\!\perp Z | X$ can be characterized as follows \[ \frac{\partial }{\partial w}\Pi(w|x,0) =\frac{\partial }{\partial w}\Pi(w|x,1), \ \ \forall (w,x)\in\mathcal{S}_{WX}, \] which holds if and only if $\Pi(w|x,0) =\Pi(w|x,1)$ for all $(w,x)\in\mathcal{S}_{WX}$.\footnote{The equivalence holds by the fact that for a continuous function $f(t)$, $f(t)=0$ for all $t\in[0,1]$ if and only if $\int_0^tf(s)ds =0$ for all $t\in[0,1]$.}

In terms of the probability distribution of $W$ given $(X,Z) = (x,z)$, both $F_{W|XZ}(\cdot|x,z)$ and $\Pi(\cdot|x,z)$ contain the exact same amount of information. The latter, however, allows us to derive a test statistic and establish its limiting distribution when $W$ has to be nonparametrically generated. When covariates $X$ are continuously distributed, the generated sample $\{\widehat{W}_i: i \le n\}$ involves the nonparametric component $\hat{\delta}(X_i)$. We exploit the smoothness of $\lambda(\cdot)$ and show that this first-stage estimation error can be further aggregated out at the $\sqrt{n}$-rate in our test statistic defined on $\{\widehat{W}_i: i \le n\}$. It should be noted that when covariates $X$ are discrete as discussed in the last subsection, we can also apply a similar testing procedure via testing $\Pi(\cdot|x,1) = \Pi(\cdot|x,0)$. Moreover, we assume that $\mathcal{S}_W$ is compact for expositional simplicity.

We denote $f_{XZ}(x,z)\equiv f_{X|Z}(x|z)\times \mathbf{P}(Z=z)$. We also let $\1^*_{A}(a) \equiv \1(A \leq a)$. For $z\in\{0,1\}$, let $z'=1-z$ and \[ G(w,x, z)=\mathbf{E} \left[\lambda (W-w)\1^*_{X}(x) \1_Z(z) f_{XZ}(X,z') \right]. \] Motivated by stinchcombe1998consistent, we rewrite the above conditional expectation restrictions by the following unconditional ones:

equation[equation omitted — 94 chars of source]

To see the equivalence, first note that \[ G(w,x, z)=\mathbf{E} \left[\lambda (W-w) \1(X\leq x) f_{X|Z}(X|z') |Z=z\right] \mathbf{P}(Z=0) \mathbf{P}(Z=1). \] Moreover, by the law of iterated expectation, \[ \frac{\partial}{\partial x} \mathbf{E}[ \lambda(W-w) \1 (X\leq x) f_{X|Z}(X|z')| Z=z] = \Pi(w|x,z) f_{X|Z}(x|0)f_{X|Z}(x|1). \] Therefore, we obtain the conditional expectation restrictions as the derivatives of (ref). Note that the estimation of $G(w,x,z)$ avoids any denominator issues, which thereafter simplifies our asymptotic analysis.

For a random variable $A$ and a value $a$, let $K_{A,h}(a) \equiv K((A-a)/h)/h$ where $K$ and $h$ are a kernel function and a smoothing bandwidth, respectively. For an $d_X$-dimensional random vector $A$, we let $K_{A,h}(a) = \Pi_{j=1}^{d_X} K_{A_j,h}(a_j)=h^{-d_X} K((A-a)/h)$. We estimate $\delta(X_i)$ by \[ \hat{\delta}(X_i)=\frac{\sum_{j\neq i} Y_j Z_j K_{X_j,h}(X_i) \sum_{j\neq i} K_{X_j,h}(X_i) - \sum_{j\neq i} Y_j K_{X_j,h}(X_i) \sum_{j\neq i} Z_j K_{X_j,h}(X_i)}{\sum_{j\neq i} D_j Z_j K_{X_j,h}(X_i) \sum_{j\neq i} K_{X_j,h}(X_i) -\sum_{j\neq i} D_j K_{X_j,h}(X_i) \sum_{j\neq i} Z_j K_{X_j,h}(X_i)}. \] Note that in this paper, we consider a kernel estimator for the nonparametric components of the test. To avoid the boundary issue, we follow the literature to trim the support of $X$. To be specific, we will assume that the support of covariates $X$ is a Cartesian product of compact intervals, $\mathcal{S}_X=\prod_{j=1}^{d_X}[x_{\ell j}, x_{uj}]$ and $\mathcal{S}_X^\xi=\prod_{j=1}^{d_X}[x_{\ell j}+\xi, x_{uj}-\xi]$ where $\xi>0$ is a small positive number. Moreover, let $\widehat W_i = Y_i + (1-D_i)\times \hat\delta(X_i)$ and

align*[align* omitted — 251 chars of source]

Thus, we define our test statistic as follows: \[ \widehat{\mathcal{T}}^{c}_{n} = \sup_{w\in\mathcal{S}_W,x\in\mathcal{S}_X^\xi}\ \sqrt{n} \left| \widehat{G}(w, x, 0) - \widehat{G}(w, x, 1) \right|. \] In \added[id=TC]{the} above definition, the supports of $W$ and $X$ are assumed to be known for simplicity. In practice, this assumption can be relaxed by using consistent set estimators $\widehat{\mathcal{S}}_W$ and $\widehat{\mathcal{S}}_X$.

We show that the proposed test statistic $\widehat{\mathcal T}^c_n$ converges in distribution at the regular parametric rate under the null. The key step of our proof is to show that

equation[equation omitted — 162 chars of source]

where $\widetilde G(w,x,z)=\frac{1}{n}\sum_{i=1}^n \1(X_i\in\mathcal{S}_X^\xi) (\widehat{W}_i-w) \1(W_i \leq w) \1^{*}_{X_i}(x)\1_{Z_i}(z) \hat f_{XZ}(X_i,z')$. The above result requires that the nonparametric elements in the estimation of $\hat{\delta}(\cdot)$ should converge to the corresponding true values uniformly at a rate faster than $n^{-1/4}$.

assmpAssume that \begin{enumerate}[(i)] • the support of the $d_X$-dimensional covariates $X$ is a Cartesian product of compact intervals, $\mathcal{S}_X=\prod_{j=1}^{d_X}[x_{\ell j}, x_{uj}]$; • For $z=0,1$, $ \inf_{x\in\mathcal{S}_{X}} f_{X|Z}(x|z)>0$, $\sup_{x\in\mathcal S_{X}} f_{X|Z}(x|z)<\infty$, and $\inf_{x\in\mathcal S_X}|p(x,1)-p(x,0)|>0$. \end{enumerate}
assmpFor $z=0,1$, $f_{X|Z}(x|z)$, $p(x,z)$ and $\mu(x,z)$ are continuous in $x\in \mathcal S_{X}$.
assmpFor some $\iota>\frac{1}{4}$, we have $h\rightarrow 0$ and $n^\iota /\sqrt {nh^{d_X}}\rightarrow 0$ as $n\rightarrow \infty$. Moreover, the first-stage estimators satisfy \added[id=TC]{the condition} that for $z=0,1$ \begin{align*} &\sup_{x\in\mathcal{S}_X^\xi}\Big|\mathbf{E} \Big[\frac{1}{n}\sum_{j=1}^n \1_{Z_j}(z) K_{X_j,h}(x)\Big]-f_{XZ}(x,z)\Big|= O(n^{-\iota}),\\ &\sup_{x\in\mathcal{S}_X^\xi}\Big|\mathbf{E} \Big[\frac{1}{n}\sum_{j=1}^n D_j \1_{Z_j}(z) K_{X_j,h}(x) \Big]-p(x,z) f_{XZ}(x,z)\Big|= O(n^{-\iota}),\\ &\sup_{x\in\mathcal{S}_X^\xi}\Big|\mathbf{E}\Big[\frac{1}{n}\sum_{j=1}^n Y_j \1_{Z_j}(z) K_{X_j,h}(x)\Big]-\mathbf{E} (Y|X=x,Z=z) f_{XZ}(x,z)\Big|= O(n^{-\iota}). \end{align*}

Assumptions (ref) and (ref) are standard in the nonparametric estimation literature. Assumption (ref) is a high-level condition that requires the nonparametric estimation bias \replaced[id=TC]{to diminish}{diminishes} uniformly at a rate faster than $n^{1/4}$. Such a condition on the bias term can be satisfied under additional primitive conditions on the kernel function and the bandwidth respectively, $K(\cdot)$ and $h$, as well as the smoothness of the underlying structural functions. See, e.g., PaganUllah1999.

lemSuppose that Assumptions (ref)-(ref) and (ref)-(ref) hold. Then, (ref) holds for $z =0,1$.

By Lemma (ref), it suffices to establish the limiting distribution of $\widetilde G(w, x, 1)-\widetilde G(w, x, 0)$ for the asymptotic properties of our test statistics. Note that in the definition of $\widetilde G(w,x,z)$, there \replaced[id=TC]{are}{contains} no nonparametric elements estimated in the indicator function.

To establish asymptotic properties for our test, we make the following assumption.

assmpAssume that for $z=0,1$, \begin{align*} \sup_{x\in\mathcal{S}_X^{\xi}}\big|\mathbf{E} [\hat \delta(x)]-\delta(x)\big|=o(n^{-1/2}), \sup_{x\in\mathcal{S}_X^{\xi}}\big|\mathbf{E}[ \hat f_{XZ}(x,z)]-f_{XZ}(x,z)\big|=o(n^{-1/2}). \end{align*}

Assumption (ref) strengthens Assumption (ref) by requiring the bias term in the first-stage nonparametric estimation to be smaller than $o_p(n^{-1/2})$, which can be established by using higher-order kernels powell1989semiparametric.

Next, we establish the asymptotic properties of the test statistic $\widehat{\mathcal{T}}^c_n $. Let $F^*_{WD|XZ}(w,d|x,z) \equiv F_{W|DXZ}(w|d,x,z) \times \mathbf{P} (D=d|X=x,Z=z)$ and \[ \kappa^c(w,x)= - \frac{F^*_{WD|XZ}(w,0|x,1)-F^*_{WD|XZ}(w, 0|x,0)}{p(x,1)-p(x,0)}. \] Moreover, we define

align[align omitted — 622 chars of source]
thmSuppose that Assumptions (ref)-(ref) and (ref)-(ref) hold. Then, under $\mathcal{H}_0$, \begin{align*} & \widehat{\mathcal{T}}^c_n\overset{d}{\rightarrow }\sup_{w\in\mathcal{S}_W,x\in\mathcal{S}_X^{\xi}} |\mathcal Z^c(w,x)| \end{align*} where $\mathcal Z^c(\cdot,\cdot)$ is a mean-zero Gaussian process with covariance kernel \[ \text{Cov} \left[ \mathcal{Z}^c(w,x), \mathcal{Z}^c(w',x') \right] = \mathbf{E} \left[(\psi^{c}_{wx} + \phi^{c}_{wx}) (\psi^{c}_{w'x'} + \phi^{c}_{w'x'})\right], \ \forall w,w'\in\mathcal{W},~x,x'\in\mathcal{S}_X^{\xi}. \] Moreover, under $\mathcal{H}_1$, we have \[ n^{-\frac{1}{2}}\widehat{\mathcal{T}}^c_n \overset{p}{\rightarrow} \ \sup_{w\in\mathcal{S}_W,x\in\mathcal{S}_X^{\xi}}\ |G(w,x,0)-G(w,x,1)|>0. \]

We also show in Appendix that the influence function for $\sqrt{n}(\widehat{G}(w, x, 0) - \widehat{G}(w, x, 1)-{G}(w, x, 0) + {G}(w, x, 1))$ is $\psi^c_{wx} + \phi^c_{wx}$, in which $\psi^c_{wx}$ is the influence function when $\delta(x)$ and $f_{XZ}(x,z)$ are known, and $\phi^c_{wx}$ accounts for the estimation effect of $\delta(x)$ and $f_{XZ}(x,z)$. Note that the $\1(X\in\mathcal{S}_X^\xi)$ term in the influence function accounts for the fact that we consider a trimmed support of $X$.

Let the simulated empirical process for the continuous case be \[ \widehat{\mathcal{Z}}^{c,u}(w,x) = \frac{1}{\sqrt{n}} \sum_{i=1}^n U_{i} \times ( \hat{\psi}^c_{wx,i} + \hat{\phi}^c_{wx,i}), \] where $\hat{\psi}^c_{wx,i} + \hat{\phi}^c_{wx,i}$ is the estimated influence function such that

align*[align* omitted — 828 chars of source]

where $\hat{f}_{XZ}(X_i,z)$, $\widehat{\mathbf{E}}[W|X_i, Z=z]$, $\widehat{F}_{W|DXZ}(w|0,X_i,z)$ and $\hat{p}(X_i,z)$ are uniformly consistent nonparametric estimators for ${f}_{XZ}(X,z)$, ${\mathbf{E}}[W|X,Z=z]$, ${F}_{W|DXZ}(w|0,X,z)$ and ${p}(X,z)$, respectively. For a given significance level $\alpha$, the critical value $\hat{c}^c_{n}(\alpha)$ is obtained as the $(1-\alpha)$-quantile of the simulated distribution of $\sup_{w\in\mathcal{S}_W,x\in\mathcal{S}_X^{\xi}} \big| \widehat{\mathcal{Z}}^{u,c}(w,x)\big|$. We would reject $\mathcal{H}_0$ at significance level $\alpha$ when $\widehat{\mathcal{T}}^c_n> \hat{c}^{c}_{\alpha}$.

Now, we give additional conditions for the validity of the multiplier bootstrap critical value.

assmpAssume that for $z=0,1$, \begin{enumerate}[(i)] • $\sup_{x\in\mathcal{S}^\xi_X}|\hat{f}_{XZ}(x,z)-{f}_{XZ}(x,z)|\stackrel{p}{\rightarrow} 0$; • $\sup_{x\in\mathcal{S}^\xi_X}|\widehat{\mathbf{E}}[W|X=x]-{\mathbf{E}}[W|X=x]|\stackrel{p}{\rightarrow} 0$; • $\sup_{x\in\mathcal{S}^\xi_X}|\hat{p}(x,z)-{p}(x,z) |\stackrel{p}{\rightarrow} 0$; • $\sup_{x\in\mathcal{S}^\xi_X, w\in\mathcal{S}_W}|\widehat{F}_{W|DXZ}(w|0,x,z)-F_{W|DXZ}(w|0,x,z)| \stackrel{p}{\rightarrow} 0$, and $\widehat{F}_{W|DXZ}(w|0,x,z)$ is non-decreasing in $w$ for all $x$ and $z$; • $\sup_{x\in\mathcal{S}^\xi_X}|\hat{\delta}(x)-\delta(x)|\stackrel{p}{\rightarrow} 0$. \end{enumerate}
thmSuppose Assumptions (ref)-(ref), (ref), and (ref)-(ref) hold. Then \begin{enumerate}[(a)] • under $H_0$, $\lim_{n\rightarrow \infty}P(\widehat{\mathcal{T}}^c_n \geq \hat{c}^c_{n}(\alpha))=\alpha$; • under $H_1$, $\lim_{n\rightarrow \infty}P(\widehat{\mathcal{T}}^c_n \geq \hat{c}^c_{n}(\alpha))=1$. \end{enumerate}

Theorem (ref) shows the size and power of our test for the continuous case. The proof of Theorem (ref) is similar to that of the discrete case. Note that Assumptions (ref), (ref) and (ref) are high-level conditions and we provide estimators and give low-level conditions in Appendix. Please see the discussion after the proof of Theorem (ref).

We can extend our test to cases where the covariate vector contains both discrete and continuous variables as in our second empirical study. We leave the details to Appendix (ref).

Computational Issue

When the dimension of the covariates $X$ is large, there will be a computational issue. For example, suppose that $X=(X_1,X_2,X_3)$ contains three continuous variables. If we use 100 grids in $W$ and 100 grids in each element of $X$, it will result in $100^4$ points of $(w,x)$ when we calculate the test statistic and the simulated critical value. Therefore, the computation burden can be too heavy to be practical when the number of grids we use in each dimension is large, or the dimension of covariates is large. We suggest conducting the test based on the test statistic calculated by all combinations of two covariates only to reduce the burden. Specifically, we calculate the test statistic based on the supremum of $100^3$ grids in $(W,X_1, X_2)$, $(W,X_1, X_3)$ and $(W,X_2, X_3)$, respectively, so that we just need to use $3\cdot 100^3$ grids to calculate the test statistic and the simulated critical value. The procedure is similar to the one recommended by andrews2013inference.

Monte Carlo Simulations

In this section, we investigate the finite sample performance of our tests with a simulation study. The data are simulated as follows:

align*[align* omitted — 197 chars of source]

where both $\epsilon_1$ and $\epsilon_2$ conform to a uniform distribution on $[-2,2]$, $\rho = 0.7$, and $Z\sim Bernoulli(p)$ with $p=0.25,0.5$ and $0.75$ respectively.\footnote{Note that we also try different values for the correlation coefficient, and all the results are qualitatively similar. Additional Monte Carlo simulation results are available upon request.} For simplicity, $X,Z$ and $(\epsilon,\eta)$ are mutually independent. Moreover, $X$ is uniformly distributed on $\{1,2,3,4\}$ and on $[0,1]$ in the discrete covariates and the continuous covariates case, respectively. Furthermore, $\gamma\in[0,1]$ describes the degree of unobserved heterogeneous treatment effects in our specification. In particular, $\mathcal{H}_0$ holds if and only if $\gamma=1$. Intuitively, the smaller $\gamma$ is, the more power we expect from our tests. To investigate the size and power of our tests, we choose $\gamma\in\{1,0.75, 0.5\}$.

We consider sample size $n=1000,2000,4000$, a nominal level of $\alpha=5\%$, and $1000$ Monte Carlo repetitions. Given $X_i = x$, to compute the suprema of the simulated stochastic processes, we use $n/20$ grids on the support of $[\min(\widehat{W}_{i:X_i=x}),\max(\widehat{W}_{i:X_i=x})]$. Moreover, we use $1000$ multiplier bootstrap samples to simulate the $p$-values. Regarding the estimation of $\kappa(w,x)$, we choose the second-order Epanechnikov kernel function with the bandwidth $h_x = c \cdot\text{std}(\widehat{W}_{i:X_i=x})\cdot n^{-1/4.5}$, and we set $c \in \{ 1.7, 2, 2.34, 2.6\}$ to study the sensitivity of the test to the bandwidth.

Table (ref) reports rejection probabilities of our simulations in the discrete-covariates case under the null hypothesis (i.e., $\gamma=1$) and alternative hypotheses (i.e., $\gamma=0.75,0.5$). From Panel A, the level of our test is fairly well behaved: It gets closer to the nominal level as the sample size increases\added[id=TC]{,} and the rejection probabilities are not sensitive to the constant $c$ for the bandwidth choice. Panels B and C show that the power of the test is reasonable. In particular, when $\gamma$ is closer to $1$, it is more difficult to detect such a “local” alternative. Therefore, we obtain relatively \replaced[id=TC]{low}{small} power even when \added[id=TC]{the} sample size reaches $n=2000$ in Panel B. For relatively “small” sample size, e.g., $n=1000$, our results show that our test performs better with a larger bandwidth choice. Moreover, when $p$ (i.e., the probability of $Z=1$) is 0.5, all the results for size and power dominate the other two cases with $p=0.25,0.75$, which is expected by our asymptotic theory.

Next, we evaluate the performance of our tests in the case where the covariates $X$ are continuous. To compute the suprema, we calculate the test statistic by using $n/20$ grid points in the support $[\min_{i=1}^n(X_i),\max_{i=1}^n(X_i)]$, as well as in the support $[Q_{\widehat{W}_i}(0.05),$ $Q_{\widehat{W}_i}(0.95)]$, where $Q_{\widehat{W}_i}(\cdot)$ is the quantile function of $\widehat{W}_i$. In the estimation of $\delta(X_i)$ and $G(w,x)$, we choose the fourth-order Epanechnikov kernel function with the bandwidth $h_n = c \times \text{std}(X_i) \times n^{-1/3}$ to reduce the bias. Also, we study the sensitivity of the test to the bandwidth with $c \in \{ 1.7, 2, 2.34, 2.6\}$. Table (ref) reports the size and power properties of our test, which are qualitatively similar to those in the discrete-covariates case.

Empirical Applications

\deleted[id=TC]{The}Effect of Job Training Program on Earnings

We now apply our tests to study the effects of \replaced[id=TC]{a}{the} job training program on earnings, i.e., the National Job Training Partnership Act (JTPA), commissioned by the Department of Labor \added[id=TC]{of the U.S.} This program \replaced[id=TC]{funded}{began funding} training from 1983 to \added[id=TC]{the} late 1990's to increase employment and earnings for participants. The major component of JTPA aims to support training for the economically disadvantaged. The effects of JTPA training programs on earnings have also been studied by e.g., heckman1997making and abadie2002instrumental under a general framework allowing for unobserved heterogeneous treatment effects. The data \replaced[id=TC]{are}{is} publicly available at \url{ https://upjohn.org/node/952}.

Our sample consists of 11,204 observations from the JTPA, a survey dataset from over $20,000$ adults and out-of-school youths who applied for JTPA in $16$ local areas across the country between 1987 and 1989.\footnote{JTPA services are provided at 649 sites, which might not be randomly chosen. For a given site, the applicants were randomly selected for the JTPA dataset. } Each participant was randomly assigned to either a program group or a control group (1 out of 3 on average). Members of the program group \replaced[id=TC]{were}{are} eligible to participate \added[id=TC]{in} JTPA services, including classroom training, on-the-job training or job search assistance, and other services, while members of control group \replaced[id=TC]{were}{are} not eligible for JTPA services for 18 months. Following the literature bloom1997benefits, we use the program eligibility as an instrumental variable for the endogenous individual participation decision.

The outcome variable is individual earnings, measured by the sum of earnings in the 30-month period following the offer. The observed covariates include a set of dummies for \replaced[id=TC]{race, high-school graduate,}{races, high-school graduates,} and marriage, whether the applicant worked at least 12 weeks in the 12 months preceding random assignment, and also five age-group dummies (22-24, 25-29, 30-35, 36-44, and 45-54), among others. Descriptive statistics can be found in Table (ref). For simplicity, we group all applicants into three age categories (22-29, 30-35, and 36 and above), and pool all non-White applicants as minority applicants.

To implement the test, given $X_i=x$, we use the second-order Epanechnikov kernel and set the smoothing parameter to $2.34 \times \text{std}(\widehat{W}_{i:X_i=x}) \times n^{-1/4.5}$ when we estimate ${\kappa}(x,w)$. For the critical value, we use $10,000$ multiplier bootstrap samples and search for the suprema by using $5,000$ grid points. The $p$-value of our test is $0.5732$. Therefore, the null hypothesis (i.e., no unobserved heterogenous treatment effects) cannot be rejected at the $10\%$ significance level. Our results are robust to the size of bootstrap samples, the number of grid points, and the choices of bandwidth. Note that our results are consistent with abadie2002instrumental, who estimate quantile treatment effects under a linear specification. In particular, one cannot reject the null hypothesis that quantile treatment effects are invariant across different quantile levels.\footnote{ We obtain a pointwise 95% confidence interval for each of the quantile treatment effects from Table III of abadie2002instrumental and find that these confidence intervals overlap. We can conclude no evidence against homogeneous treatment effects because a joint confidence band is, in general, wider than a pointwise confidence band. }

The Impact of Fertility on Family Income

The second empirical illustration considers the heterogeneous impacts of children on parents' labor supply and income. Recently, frolich2013unconditional studied the heterogeneous effects of fertility on family income within the general LATE framework. To deal with the endogeneity of fertility decisions, rosenzweig1980testing, angrist1998children, bronars1994economic and 10.2307/146376, among many others, suggest \replaced[id=TC]{using}{to use the} twin births as an instrumental variable.

Our data use the 1% and 5% Census Public Use Micro Sample (PUMS) from 1990 and 2000 censuses, consisting of 602,767 and 573,437 observations, respectively. The data \replaced[id=TC]{are}{is} publicly available at \url{https://www.census.gov/main/www/pums.html}. Similar to frolich2013unconditional, our sample is restricted to 21- to 35-year-old married mothers with at least one child, since we use twin birth as an instrument for fertility. The outcome variable of interest is the family's annual labor income.\footnote{It includes wages, salary, armed forces pay, commissions, tips, piece-rate payments, cash bonuses earned before deductions were made for taxes, bonds, pensions, union dues, etc. See frolich2013unconditional for more details.} The treatment variable is a dummy variable that takes the value $1$ to indicate when a mother has two or more children. The instrumental variable is also a dummy variable and it equals $1$ if the first birth is a twin. The covariates include mother's age, race, and educational level. Some covariates, i.e., age and years in education, are treated as continuous variables. Summary statistics can be found in Table (ref).

Similar to the previous empirical illustration, we use the fourth-order Epanechnikov kernel and set the smoothing parameter to $2.34 \times \mbox{std}(X_i) \times n^{-1/3}$ when we estimate the influence function. For the critical value, we use $5000$ bootstrapped samples and search for the suprema by using $1000$ grids for each of the \replaced[id=TC]{supports of both $W$ and $X$'s.}{support of $W$ and $X$'s.} The bandwidths are selected in the same manner as those in the JTPA case. The $p$-values of our tests are $0.0013$ and $0.0007$ for the 1990 and 2000 censuses, respectively. These results suggest that the null hypothesis, i.e., homogeneous treatment effects, should be rejected at all usual significance levels.

Extensions

When $Z$ takes multiple values rather than being binary, one could extend our approach of testing for unobserved heterogeneous treatment effects. Namely, let $W = Y + (1-D)\delta(X)$ where $\delta(x) = \frac{\text{Cov}(Y,Z|X=x)}{\text{Cov}(D,Z|X=x)}$. Then we test $\mathcal{H}_0$ by testing $W \perp \!\!\! \perp Z | X$. Since $Z$ takes more than binary values in its support, this model restriction can be equivalently written as \[ F_{W|XZ}(\cdot|x,z) = F_{W|X}(\cdot|x),\ \forall\ (x,z) \in \mathcal{S}_{XZ} \] or \[ \Pi(\cdot|x,z) = \mathbf{E}(\lambda(W-\cdot)|X=x), \forall\ (x,z) \in \mathcal{S}_{WXZ} \] depending on whether covariates $X$ or instruments $Z$ contain any continuously distributed components or not.

Such a test, however, does not exploit model restrictions arising from multiplicity of $Z$. For instance, suppose that $\mathcal{S}_Z = \{0,1,2\}$. Under $H_0$ and Assumption 1, we have \[ \frac{\mathbf{E}(Y|X=x,Z=z)-\mathbf{E}(Y|X=x,Z=z)}{p(x,z) - \mathbf{E}(D|X=x)} = \delta(x),\ \forall x. \] As a matter of fact, our test does not exploit such a model restriction.

Our analysis naturally extends the case where the treatment variable $D$ takes multiple values. For illustration, suppose $\mathcal{S}_D = \{0,1,2\}$. Under the homogeneous treatment effects hypothesis, denote $\delta_1(x) \equiv g(1,x,\cdot) - g(0,x,\cdot)$ and $\delta_2(x) \equiv g(2,x,\cdot) - g(0,x,\cdot)$. For $d=1,2$, let $p_d(X,Z) = P(D=d|X,Z)$ and $W_d \equiv \delta_d(X) + Y - \sum_{d'=1}^2 \1 (D=d') \times \delta_{d'}(X)$. Note that under $H_0$, i.e., $g(d,x,\cdot) - g(0,x,\cdot) = \delta_d(x)$, we have \[ W_d = g(d,X,\epsilon). \] By a similar argument, we test for unobserved heterogeneous treatment effects by testing $W_d \perp \!\!\! \perp Z|X$ for $d=1,2$. To complete our analysis, it suffices to establish the identification of $\delta_d(x)$. Note that under $\mathcal{H}_0$ and Assumption 1, we have \[ \mathbf{E}(Y|X=x,Z=z) = \mathbf{E}[g(0,X,\epsilon)|X=x] + \delta_1(x) \times p_1(x,z) + \delta_2(x) \times p_2(x,z),\ \forall z. \] Therefore, $\delta_d$ is identified if $\{ ( p_1(x,z),p_2(x,z) )': z \in \mathcal{S}_{Z|X=x}\}$ has the full rank. Note that such a rank condition requires $\mathcal{S}_{Z|X=x}$ to contain at least three values.