Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
67,443 characters · 15 sections · 52 citation commands
Testing for homogeneous treatment effects in linear and nonparametric instrumental variable models
KEYWORDS: Nonparametric Instrumental variables, Hypothesis testing, Homogeneous treatment effects.
\spacingset{1.4}
We are interested in the effect of a (possibly continuous) random treatment $Z$ (with support $\mathcal {Z}\subset\mathbb{R}^{p}$) on a scalar outcome outcome $Y$. Let $Y(z)$ be the scalar potential outcome of the outcome $Y$ under treatment status $z\in\mathcal{Z}$. The potential outcomes are not observed, but $Y=Y(Z)$ is part of the data. We write $$Y(z)=\varphi(z)+U(z),$$ where $U(z)$ is an error term such that ${\mathbb{E}}[U(z)]=0$. The goal is to identify the structural function $\varphi$, which allows to obtain the average treatment effect of changing treatment level from $z\in\mathcal{Z}$ to $z'\in\mathcal{Z}$, i.e. ${\mathbb{E}}[Y(z')-Y(z)]=\varphi(z')-\varphi(z).$
We say that $Z$ is endogenous when the structural function $\varphi$ is not characterized by the distribution of $Y$ conditional on $Z$. Even in this setting, it is possible to identifiy $\varphi$ thanks to an instrumental variable $W$ with support $\mathcal{W}\subset{\mathbb{R}}^q$. Loosely speaking, an instrumental variable satisfies two conditions. First, it is (in some sense specified in the paper) independent of $\{U(z)\}_z$. Second, it is sufficiently related to $Z$.
Our goal is to test the following hypothesis $$H_0: \text{ $U(z) =U$ for all $z\in \mathcal{Z}$}.$$ We call this assumption “homogeneous treatment effects” since it stipulates that the treatment effect $Y(z’)-Y(z)$ is a deterministic variable. Under $H_0$, we have $Y=\varphi(Z)+U$, which corresponds to the standard nonparametric instrumental variable (NPIV) model, see newey2003instrumental,horowitz2007nonparametric,darolles2011nonparametric,chen2012estimation and the book li2007nonparametric, among others.
There are two reasons why one would be interested in testing $H_0$. First, under the instrumental variable conditions, the aforementioned works in the instrumental variable literature show that the structural function $\varphi$ is identified if $H_0$ holds. When $H_0$ is not satisfied, instrumental variable estimators can be more biased than naive estimators ignoring confounding since the error term $U(Z)$ of the model can be dependent of $W$ (through $Z$). In this case, the techniques developed in the nonparametric instrumental variable literature outlined above should not be used. Second, if $H_0$ is true, the treatment effects are constant (i.e. $Y(z’)-Y(z)$ is a deterministic variable). This allows to estimate all counterfactuals for each observation in the dataset (individualized predictions) using an estimator of $\varphi$. It also means that the treatment does not distort the ranks of potential outcomes, which has implications in terms of inequalities.
The contributions of the paper are as follows. We propose two tests for $H_0$. The tests rely on a scalar covariate $X$ such that $(\{U(z)\}_z,X) $ is independent of one component $W_k$ of the instrument $W$. Importantly, the covariable $X$ is allowed to be dependent of $\{U(z)\}_z$ (even conditional on $W$). The first test relies on the assumption that the potential outcomes are linear in the regressors. In a first step, it estimates $U$ as the residual of a two stage least-squares regression. Then, it tests that the estimator of $U$ is uncorrelated with $X(W_k-{\mathbb{E}}[W_k])$, which is an implication of homogeneous treatment effects under our assumptions. The second test is nonparametric. In this case, the error term $U$ is estimated by the residual of a Tikhonov regression. The second step of the test is then the same as that of the linear test. In both cases, the treatment can be discrete or continuous. We study the asymptotic distribution of the test statistic and the power of the test. We discuss how to choose $X$: to maximize the likelihood that the conditions on $X$ are satisfied , we argue in the paper that $X$ should be chosen so as to be independent of $W$. Finally, the empirical performance of the tests is assessed through simulations and illustrated thanks to applications on returns to schooling and price elasticity of demand in a fish market.\\
Related literature Our paper is related to a large literature on testing for (various implications of) homogeneous treatment effects: see koenker2002inference, crump2008nonparametric, chernozhukov2005subsampling, ding2016randomization,hsu2017consistent, goldman2018comparing,chung2021permutation, sant2021nonparametric and dai2022u among others. Apart from sant2021nonparametric, none of these papers considers the case where there is endogeneity. sant2021nonparametric only studies the setting where both the treatment and the instrument are binary and a monotonicity assumption such as in the literature on local average treatment effects holds (see angrist1996identification). Instead, our paper allows for confounding and continuous treatments.
Another related literature is that on the hypothesis of rank invariance (see chernozhukov2005iv). This assumption is the counterpart of our homogeneous treatment effects condition in nonseparable instrumental variable models (in contrast with the present paper which studies separable models). chernozhukov2005iv notes that identification of $\varphi$ can be obtained under a weaker (but less interpretable) form of rank invariance called rank similarity. We could define an assumption analogous to rank similarity in our case of separable models. Our tests would also work for this assumption. We chose to focus solely on homogeneous treatment effects to simplify the exposition. Moreover, dong2018testing, frandsen2018testing and kim2022testing develop tests of the rank invariance assumption in nonseparable instrumental variable models in the case where the treatment and the instrument are binary. The case with continuous treatment is much more challenging because regularization is needed for nonparametric estimation. Note also that dong2018testing and kim2022testing rely on a monotonicity assumption from local average treatment effects literature (see angrist1996identification), while our approach does not require it. The test developed in frandsen2018testing bears similarities with our tests since they are based on equivalent single restrictions. However, our models and the first steps of our tests are different. We provide a more in depth analysis of the underlying assumptions of the method and study it in a different context with models for average treatment effects and allow for continuous treatments.
Outline Section (ref) presents the linear test along with theory, simulations, and an application to returns to schooling. Then, in Section (ref), we introduce the nonparametric test, discuss its asymptotic properties, evaluate its finite sample performance through numerical experiments, and illustrate the test with an application to demand estimation.
\setcounter{equation}{0}
In this section, we develop theory for a test assuming that the mapping $\varphi$ is linear and the instrument $W$ is uncorrelated with $\{U(z)\}_z$, i.e. there exists $\beta\in{\mathbb{R}}^p$ such that $$Y(z)=z^\top\beta+U(z),\ {\mathbb{E}}[WU(z)]=0,$$ for all $z\in\mathcal{Z}$. We also impose the usual relevance assumption for instrumental variables in the linear model, i.e. ${\mathbb{E}}[WZ^\top]$ has rank equal to $p$. The aim is to recover $\beta$, since the average treatment effect of changing treatment level from $z\in\mathcal{Z}$ to $z'\in\mathcal{Z}$ can be expressed as ${\mathbb{E}}[Y(z')-Y(z)]=(z'-z)^\top\beta.$
Let the symbol $\protect\mathpalette{\protect\independenT}{\perp}$ stand for statistical independence. To test for $H_0$, we leverage a scalar covariate $X$ satisfying the following condition:
This assumption allows $X$ and $\{U(z)\}_z$ to be dependent (even conditional on $W$). Hence $X$ does not need to be exogenous or satisfy an exclusion restriction. In particular, $X$ is not a second instrument.
We argue that Assumption (ref) is very likely to hold for some $X$ if $W_k$ is a good instrument. Indeed, if $W_k$ is independent of the unobserved heterogeneity of the model, there is every reason to believe that it is also independent of some observable variables. Note that Assumption (ref) is equivalent to the standard instrumental variable condition $\{U(z)\}_z\protect\mathpalette{\protect\independenT}{\perp} W_k$ and the additional condition $X\protect\mathpalette{\protect\independenT}{\perp} W_k|\{U(z)\}_z.$ Hence, the condition we add to the standard instrumental variable literature is really $X\protect\mathpalette{\protect\independenT}{\perp} W_k|\{U(z)\}_z.$
The tests of $H_0$ that we propose depend crucially on Assumption (ref). Contrarily to $H_0$, hypothesis (ref) can be heuristically justified. The econometrician should try to select a covariable $X$ which is likely to be (jointly with $\{U(z)\}_z$) independent of $W_k$. In applications, we recommend to pick a variable $X$ independent of $W_k$ since this should make Assumption (ref) more likely to hold as $X\protect\mathpalette{\protect\independenT}{\perp} W_k$ is a necessary condition for Assumption (ref). The independence of $X$ and $W_k$ can be assessed by a statistical test of independence such as $\chi^2$ or Kolmogorov-Smirnov independence tests. Our empirical applications illustrate possible choices of $X$.
Note that Assumption (ref) could hold for several components of $W$ and a multidimensional $X$. The extension of the test would be straightforward in this case. We focus on the scalar case to simplify the exposition and because the validity of this assumption for a single component of $W$ is less restrictive.\\
Let us now define the population analog of the two-stage least squares estimator (henceforth, TSLS) in this context: $$\beta^{TSLS} =\left[\Gamma^\top {\mathbb{E}}[WW^\top]\Gamma\right]^{-1}\Gamma^\top {\mathbb{E}}[WY],$$ where $\Gamma ={\mathbb{E}}[WW^\top]^{-1}{\mathbb{E}}[WZ^\top].$ In this case, under $H_0$, $\beta$ can be estimated by TSLS, i.e. $\beta^{TSLS}=\beta$. However, when $H_0$ does not hold, then the bias of the TSLS estimator can be larger than that of the ordinary least squares (henceforth, OLS) estimator as illustrated by the following example.\\
Example 1. We study the case of a randomized experiment with two-sided noncompliance and monotonicity (see angrist1996identification). Let $W=(1,W_2)^\top$, where $W_2$ is a Bernoulli random variable with ${\mathbb{P}}(W_2=1)=1/2$ and $Z=(1,Z_2)^\top$, with $Z_2= W_21\left\{\epsilon\ge 1/2\right\}+(1-W_2)1\left\{1/2\le \epsilon<3/4\right\}$, where $\epsilon\protect\mathpalette{\protect\independenT}{\perp} W$ follows a uniform distribution on the interval $[0,1]$. Note that ${\mathbb{P}}( Z_2=1|W_2=0)=1/4$ and ${\mathbb{P}}( Z_2=1| W_2=1)=1/2$. For some $\alpha>0$, we also impose $\beta_1=0$ and $\beta_2=\alpha/4$ (i.e. $Y(1,z_2)=(\alpha/4)z_2+ U(Z)$) and $U(1,z_2)=z_2(1\left\{\epsilon\ge 3/4\right\}-(1/4))\alpha$ for $z_2=0,1$). This definition ensures that $E[U(z)]=0,$ for $z=0,1$. This model satisfies all the usual instrumental variable assumptions except $H_0$. The average treatment effect over the whole population is $$\Delta ={\mathbb{E}}[Y(1,1)-Y(1,0)]=\frac{\alpha}{4}.$$ We show in the supplementary material that the population analog of the OLS estimator of $\beta_2$ is given by $$\beta^{OLS}_2= {\mathbb{E}}[Y|Z_2=1]-{\mathbb{E}}[Y|Z_2=0]=\frac{\alpha}{3}.$$ Instead, in the context of the present example, it is known (angrist1996identification) that $\beta^{TSLS}_2$ is equal to the average treatment effects on the population of compliers. The compliers are the subjects who change treatment status with the instrument, or, equivalently, with $\epsilon\ge 3/4$. As a result, it holds that $\beta_2^{TSLS}=\alpha$. Hence, the bias of the TSLS estimator is larger than that of the OLS estimator for $\Delta$ and can even go to infinity as $\alpha\to\infty$.
Let $U^{TSLS}= Y-Z^\top\beta^{TSLS}$. Under the null hypothesis $H_0$, $U(z)$ will not depend on the treatment level $z$ and $U^{TSLS}=U$, which yields $(U^{TSLS},X)\protect\mathpalette{\protect\independenT}{\perp} W_k$ (by Assumption 2.1). This last independence condition implies that
This is the testable implication of $H_0$ that we use to construct the test. In particular, our test estimates the moment (ref) and checks if the latter is close to $0$. Remark that the role of the covariable $X$ appears clearly in (ref). Indeed, when $W$ contains an intercept, the equality
always holds by definition of $U^{TSLS}$, regardless of the validity of $H_0$. Hence, a test based on (ref) would have no power. The role of $X$ is therefore to give power to our test.
A few remarks are in order. First, there are multiple reasons why we choose to focus on testing the moment condition (ref) instead of any of the other implications of $H_0$. This is because, under our assumptions, (ref) is equivalent to saying that the coefficients of the linear projection of $W_k-{\mathbb{E}}[W_k]$ onto a constant, $U^{TSLS}$, $X$, and the product $XU^{TSLS}$ are all null. Indeed, the latter coefficients are “proportional" to $\mathbb{E}[(W_{k}-{\mathbb{E}}[W_k])(1,U^{TSLS},X,U^{TSLS} X)]$, with $\mathbb{E}[(W_{k}-{\mathbb{E}}[W_k])(1,U^{TSLS})]=0$ by construction and $\mathbb{E}[(W_{k}-{\mathbb{E}}[W_k])X]=0$ by Assumption (ref), so that the coefficients are null if and only if (ref) holds (under Assumption (ref)). We therefore focus on (ref) because of this nice interpretation.
Second, we acknowledge that instead of focusing on (ref) we could directly test the independence condition $(U^{TSLS},X)\protect\mathpalette{\protect\independenT}{\perp} W_k$. However, we prefer to focus on (ref) for three main reasons. First, as we show in the next section, our test based on (ref) is very simple to implement, and hence it will be attractive to practitioners. Differently, building a test for the independence condition $ (U^{TSLS},X)\protect\mathpalette{\protect\independenT}{\perp} W_k$ would require performing a multi-step estimation, where in a first step we estimate the error $U^{TSLS}$, in a second step we should estimate the cumulative distribution functions of $W_k$ and $(U^{TSLS},X)$, and in a third step we should compare these cumulative distribution functions on the basis of a certain distance. A test of this type would be less attractive to implement in practice. Second, as we show in Section (ref), a test based on the moment condition ((ref)) will have power against a wide range of alternatives. We also remark that focusing on the moment condition in ((ref)) is not necessarily a limitation: while we do not have consistency against all fixed alternatives, we focus the power towards the specific direction of the moment (ref). Instead, a test comparing the nonparametric cumulative distribution functions of $(U^{TSLS},X)$ and $W$ would spread the power along all directions, and may have lower power in the direction of our moment (ref) and, hence, in the direction of the linear projection of $W_k-{\mathbb{E}}[W_k]$ onto a constant, $U^{TSLS}$, $X$ and the product $XU^{TSLS}$, as discussed above.
Third, remark that when $X$ is binary (support equal to $\{0,1\}$), (ref) becomes ${\mathbb{E}}[U^{TSLS}(W_k-{\mathbb{E}}[W_k])|X=1]=0$. Hence, in this case our test checks if $U^{TSLS}$ and $W_k$ are uncorrelated (that is $W_k$ is exogenous) in the subpopulations for which $X=0$ and $X=1$.
Let us formally outline the test. Consider an i.i.d. sample $\{Y_i,Z_i,W_i,X_i\}_{i=1}^n$ generated from the model of Section (ref). Let $\widehat{\beta}^{TSLS}$ be the TSLS estimator in this sample, given by $$\widehat\beta^{TSLS}=\left[\widehat \Gamma^\top \left(\frac1n\sum_{i=1}^n W_iW_i^\top\right)\widehat \Gamma\right]^{-1}\widehat \Gamma^\top \left(\frac1n\sum_{i=1}^nW_iY_i\right),$$ where $\widehat \Gamma =\left(\frac1n \sum_{i=1}^nW_iW_i^\top\right)^{-1}\frac1n\sum_{i=1}^n W_iZ_i^\top.$ An estimator for $U_i$ is therefore $$\widehat U_i=Y_i-Z_i^\top \widehat\beta^{TSLS}.$$ Let $\overline{W}_k=n^{-1}\sum_{i=1}^n W_{ki}$. The test statistic is then $$T_n=\frac{1}{\sqrt{n}} \sum_{i=1}^n \widehat U_iX_i(W_{ki}-\overline{W}_k).$$ We use bootstrap to obtain the p-value of the test, since the variance of the statistic has a tedious expression and variance estimators based on analytical formulas tend to perform poorly in the presence of heteroscedasticity (see e.g. mackinnon1985some). The procedure to compute the p-value is as follows.
If the p-value so obtained is smaller than $\alpha$ (the nominal size of the test), the null hypothesis is rejected at the $\alpha$ nominal level. Notice that in the above procedure we are computing a { symmetric} p-value. An alternative procedure is to compute an { equal-tailed} p-value, but in our simulations the symmetric one has a satisfying performance.
Informally, when $H_0$ does not hold, $U^{TSLS}$ will, in general, not be equal to some $U(z)$ for a fixed $z\in\mathcal{Z}$. Therefore, in this case, there is no reason that $ (U^{TSLS},X)\protect\mathpalette{\protect\independenT}{\perp} W_k$, and hence, that (ref) holds. This gives power to our test. Going further, we argue in this section that it is overwhelmingly likely that (ref) is wrong when $H_0$ does not hold. \color{black} The test will have asymptotic power equal to $1$ under alternative hypotheses for which
Since $U^{TSLS}=Y-Z^\top\beta^{TSLS}=U(Z)+Z^\top(\beta-\beta^{TSLS})$, Equation (ref) is equivalent to
We see that there are two reasons why the test would have power:
We argue that (1) and (2) are likely to hold in applications. Indeed, statement (1) is probable since $W$ and $Z$ are correlated and $U(Z)$ depends on $Z$. The fact that $E[X(W_k-E[W_k]) Z^\top]\ne 0$ is likely under the assumption that $E[ZW^\top]$ has rank $p$. When $H_0$ does not hold, $\beta$ and $\beta^{TSLS}$ should be different, they would be equal only under very specific values of $U(Z)$.
Notice that it would be possible for (1) and (2) to hold while the test does not have power. This happens when the two terms in (ref) compensate each other. This case, however, requires very specific data generating processes.
Overall, although there are cases where the test does not have power, the above discussion suggests that the test has power against a wide range of alternative hypotheses. The following example illustrates this claim.\\
Example 2. Let $(W,E,X)^\top$ be a $3\times 1$ Gaussian vector with mean zero and variance equal to the identity matrix. Let also $Z=W+WE+\rho WX$, $U(Z) = Z(E+\rho X)$, where $\rho \in{\mathbb{R}}$, and $\beta= 0$ so that $Y=U(Z)=Z(E+\rho X)$. In this case, we have
Moreover, ${\mathbb{E}}[ZW]={\mathbb{E}}[W^2+W^2E+\rho W^2X]=1$. In the present context, it is well-known that $\beta^{TSLS}={\mathbb{E}}[YW]/{\mathbb{E}}[ZW]$, which yields $\beta^{TSLS}=1+\rho^2$. We have $U^{TSLS} =Y-Z\beta^{TSLS}= Z\left\{(E+\rho X)-(1+\rho^2)\right\}$. Hence, we get
As a result, the test has no power only when $X$ and $U$ are independent, i.e. $\rho =0$, which is a degenerate case. \\
In this previous parametric example, we see that the alternatives under which the test has no power have measure equal to 0 for the uniform measure. The example also shows the role of $X$ in providing power to the test, since the latter does not have power only when $X$ and $(Z,\{U(z)\}_z)$ are independent.
In this section, we state the asymptotic properties of the test statistic. We make the following assumption, which ensures the convergence of the TSLS estimator.
We have the following theorem.
The influence function representation of the test statistic is given in Lemma 1.1 in the supplementary material. The asymptotic variance of the test statistic can be derived from this result. Note that the fact that $|T_n|/\sqrt{n}\to C\ne 0$ when (ref) holds implies that the power of the test goes to $1$ as $n$ goes to $\infty$ under alternative hypotheses satisfying (ref).
We study the following data generating process. Let $(W_2, E, X, Z_3)$ follow a standard $4$-dimensional Gaussian distribution. We define $Z_2=W_2+E+X$, $Z=(1,Z_2,Z_3)^\top$, $W=(1,W_2,Z_3)^\top$. The variable $Z_2$ suffers from confounding and is instrumented by $W_2$, while the variable $Z_3$ is exogenous. We let $U(z)=(1+\rho z_2)(E+X),$ for all $z\in\{1\}\times {\mathbb{R}}^2,$ where $\rho\in {\mathbb{R}}$. The null hypothesis $H_0$ holds when $\rho =0$. The outcome is $Y=U(Z)$ so that the causal regression function of interest is $\varphi= 0$. We set $k=2$, i.e. we use the second component of $W$ to compute the test statistic $T_n=n^{-1/2} \sum_{i=1}^n \widehat U_iX_i(W_{2i}-\overline{W}_2).$
We generate data with sample sizes $n\in\{100,200, 500, 1000\}$. We investigate the empirical size of the test when $\rho =0$ (Table (ref)). The results are averages over $10000$ replications using $B=1000$ bootstrap resamples. We also study the empirical power of the test when $\rho=-1, -0.9,\dots,1$, with a thousand replications and bootstrap resamples (See Figure (ref)). The empirical size of the test is almost nominal even for low sample sizes. The power of the test increases as the deviation from the null ($|\rho|$), or the sample size, become larger.
In this section, we test for {homogeneous treatment effects} in a famous application on returns to schooling from card1993using. The database is available in the R package ivreg and is an extract from the 1976 National Longitudinal Survey (NLS) of young men. Here, $Y=\log(wage)$, $Z=(\boldsymbol{1},education, experience, experience^2, smsa, south, ethnicity)^\top$, where $\boldsymbol{1}$ denotes a constant, { education} and work $experience$ are measured in years, and $smsa, south, ethnicity$ are controls whose definition can be found in the ivreg package. { Education} and { experience} are endogenous since they depend on individual’s ability which is unobserved. The instrument is $W=(\boldsymbol{1}, nearcollege, age, age^2, smsa, south, ethnicity)^\top$, where $nearcollege$ is an indicator whose value is equal to $1$ when the individual grew up near an accredited four-year college. The validity of the instrument is documented in card1993using.
This corresponds to one of the specifications in the original paper of card1993using. The variable $X$ is an indicator for being married. It is unlikely that being married is (jointly with the unobserved heterogeneity of the model) dependent from growing up near a four-year college. In fact, a $\chi^2$ test does not reject the hypothesis of independence of $W_2=nearcollege$ and $X=married$ (p-value equal to $0.55$). Hence, we can reasonably assume that $(\{U(z)\}_z,X)\protect\mathpalette{\protect\independenT}{\perp} W_2$ for all $z\in\mathcal{Z}$, which corresponds to Assumption (ref). Using 10,000 bootstrap resamplings, the p-value of the test is equal to 0.0965, so that the null hypothesis of {homogeneous treatment effects} can be (borderline) rejected at the 10% level. This result cautions against interpreting the estimates of card1993using as causal effects.
\setcounter{equation}{0} In this section we extend the linear framework to consider a nonparametric treatment function. We assume that $Z$ and $W$ are unidimensional to simplify the exposition (it is rare to apply nonparametric procedures with multidimensional variables). We continue supposing that Assumption (ref) holds and we keep the notation previously introduced. We write $$Y(z)=\varphi(z)+U(z),\ {\mathbb{E}}[U(z)|W]=0, \text{ for all }z\in\mathcal{Z}\, ,$$ where the conditional mean independence $\mathbb{E}[U(z)|W]=0$ is a direct consequence of Assumption (ref). The functional form of $\varphi$ is unknown. Let $L^2(Z)$ be the set of functions that are square integrable with respect to the distribution of $Z$. We maintain the following assumption which is also standard in the NPIV literature.
The first part of the assumption is the completeness condition standard in the NPIV literature. It means that $W$ is sufficiently associated with $Z$. Under the null hypothesis $H_0$, such a condition is necessary and sufficient for identifying $\varphi$ (see, among others, carrasco_chapter_2007, darolles2011nonparametric, newey_instrumental_2003, dhaultfoeuille_completeness_2011). The second part states that the NPIV model in the above equation ($\mathbb{E}[Y|W]=\mathbb{E}[\widetilde \varphi(Z)|W]$) is well specified. We will discuss this assumption in Section (ref). \\
Recall that the hypothesis that we want to test is $$H_0: \text{ $U(z) =U$ for all $z\in \mathcal{Z}$}.$$ Under the null hypothesis of {homogeneous treatment effects} $\mathbb{E}[U(Z)|W]=\mathbb{E}[U(z)|W]={\mathbb{E}}[U(z)]=0$, so the treatment function $\varphi$ must satisfy the following (functional) equation in $\phi\in L^2(Z)$
Under the completeness condition (Assumption (ref) (a)) the solution to the above equation is { unique} and can be consistently estimated (we introduce an estimator below). Since under the null hypothesis of {homogeneous treatment effects} the treatment function $\varphi$ will satisfy the above equation, $\varphi$ will be identified as its unique solution. Thus, under the null hypothesis of {homogeneous treatment effects} we can consistently estimate the nonparametric treatment function $\varphi$. Differently, when $H_0$ does not hold, the solution to Equation ((ref)) will be different, in general, from the treatment function of interest $\varphi$. This solution, denoted with $\widetilde{\varphi}$, is introduced in Assumption (ref) (b). So, if $H_0$ is not true, although we will be able to estimate consistently the solution to Equation ((ref)), we will not in general obtain a consistent estimator of $\varphi$. It is therefore crucial to test for the {homogeneous treatment effects} hypothesis.
Let $\widetilde{U}(Z):=Y(Z)-\widetilde{\varphi}(Z)$ and $\widetilde{\varphi}$ denote the unique solution to Equation (ref) (which exists by Assumption (ref) (ii)). Under the null of homogeneous treatment effects $\widetilde{U}(Z)=U$ and $(X,\widetilde{U}(Z))\protect\mathpalette{\protect\independenT}{\perp} W$. Thus, by arguing similarly as in the previous section, we test
Remark that, in the NPIV model, instead of conducting a test based on the moment condition in ((ref)), we could build a test to directly check that $(\widetilde{U}(Z),X)\protect\mathpalette{\protect\independenT}{\perp} W$. However, we choose to focus on the single moment condition in ((ref)) for several reasons. First, our test based on a single moment restriction has power under a wide range of alternatives. Indeed, a power analysis similar to that in Section (ref) could be carried out in the present nonparametric case. Some examples of alternative hypotheses under which the test has power are given in Section S2.2 of the supplementary material. Second, as argued in Section (ref), our test avoids to spread power across multiple directions. Third, focusing on a single restriction allows to propose a test which is simple to implement. In contrast, building a test for the condition $(\widetilde{U}(Z),X)\protect\mathpalette{\protect\independenT}{\perp} W$ would be theoretically intricate. In fact, it would require to estimate the error $\widetilde{U}(Z)$ in a first step and then to estimate the cumulative distribution functions of $W$ and $(\widetilde{U}(Z),W)$ in a second step. It is likely that in this case we would need an $n^{-1/4}$ convergence rate for the first-step NPIV estimator of $\widetilde{\varphi}$, as it is required in multi-step estimations with a nonlinear second step, see CLV. This would require a slow convergence rate of the regularization parameter to allow the variance of the first-step NPIV estimator to converge fast enough to zero, see Assumption (ref) ahead. At the same time, we would need the regularization bias of $\widetilde \varphi$ to disappear fast enough, and this would require that the regularization parameter converges quickly to zero, see Assumption (ref) ahead. Therefore, there would be some tensions between these two conditions and the construction of the test would become intricate. In particular, the rates of convergences obtained with the Tikhonov regularization scheme would not allow to obtain a consistent nonparametric test. A more complicated first-step estimator would be required.
Let $\widehat{\varphi}$ be an estimator for $\varphi$ (below we define our estimator), and let us define the estimated residuals $\widehat{U}_i=Y_i-\widehat{\varphi}(Z_i)$, for $i=1,\ldots,n$. Our statistic is the empirical counterpart of the above moment
where $\overline{W}=n^{-1}\sum_{i=1}^nW_i$ is the empirical mean of $W$. In the following section, we explain how to compute the above statistic and implement the test based on it. In Section (ref), we present the assumptions and the results for the validity of the test. Next, in Section (ref), we provide evidence about the finite sample performance of the test based on $S_n$. Finally, we illustrate the procedure by an application to demand estimation in Section (ref).
In this section, we outline the computation of the statistic $S_n$ and implement the test. To this end, we need (a) to compute $\widehat{\varphi}(Z_i)$ for $i=1,\ldots,n$, and (b) to obtain the p-value necessary for testing. \\
Computation of $\widehat{\varphi}$. Let $\pi$ and $\tau$ be positive functions on $\mathbb{R}$. We denote with $L^2_\pi(\mathbb{R})$ the space of square integrable functions with respect to $\pi$, and let $L^{2}_\tau({\mathbb{R}})$ be similarly defined. The functions $\pi$ and $\tau$ are introduced for the aim of generality and for technical reasons. In simulations they are set to Gaussian density functions. Let us briefly explain their roles. First, from a technical point of view it would be ideal to work with the spaces $L^2(Z)$ and $L^2(W)$, i.e. the spaces of square integrable functions with respect to the distribution of $Z$ and $W$. However, since we do not know such distributions, it is convenient to replace $L^2(Z)$ and $L^2(W)$ with $L^2_{\tau}(\mathbb{R})$ and $L^2_{\pi}(\mathbb{R})$ and work with the latter spaces. Second, by working with $L^2_{\tau}(\mathbb{R})$ and $L^2_{\pi}(\mathbb{R})$ we can obtain that the estimator of the dual of $A$ (defined below) is actually the dual of the estimator of $A$. This is an important property for establishing the asymptotic normality of our statistic $S_n$ under $H_0$.
Let $f_W$ stand for the density of $W$ with respect to the Lebesgue measure, and assume that $\tau(w)>0$ whenever $f_W(w)>0$. Then, Equation ((ref)) is equivalent to
To build an estimator for $\widetilde{\varphi}$ we start from the above integral equation. Let us denote by $f_{Z W}$ the joint density of $(Z,W)$ with respect to the Lebesgue measure and introduce $L^2_{\pi \otimes \tau}({\mathbb{R}}^{2})$, which is the set of functions from ${\mathbb{R}}^2$ to ${\mathbb{R}}$ square integrable with respect to the product measure $\pi \otimes \tau$ (the measure generated by the product density $\pi \cdot \tau$). By assuming that $f_{Z W}/(\pi\,\tau)\in L^2_{\pi \otimes \tau}({\mathbb{R}}^{2})$, we define $A: L^2_\pi({\mathbb{R}})\mapsto L^2_\tau({\mathbb{R}})$ as the operator $$(A\phi)(\cdot)=\int_{ }\phi(z) f_{ZW}(z,\cdot)dz\,\frac{1}{\tau(\cdot)}.$$ Its Hilbert adjoint $A^*:L^2_\tau({\mathbb{R}})\mapsto L^2_\pi({\mathbb{R}})$ is given by $$(A^*\psi)(\cdot)=\int_{ }\psi(w)f_{Z W}(\cdot,w)\,dw\,\frac{1}{\pi(\cdot)}.$$ We also define the “left hand side" of the integral equation ((ref)) as $$r(\cdot)=\mathbb{E}[Y|W=\cdot]\frac{f_W(\cdot)}{\tau(\cdot)}.$$ We estimate $A$ and $A^*$ by replacing $f_{ZW}$ with its kernel estimator $$\widehat{f}_{ZW}(z,w) =\frac{1}{n\, h_Z\,h_W}\sum_{i=1}^n K_Z\left(\frac{Z_i-z}{h_Z}\right)K_W\left(\frac{W_i-w}{h_W}\right),$$ where $K_Z$ and $K_W$ are two kernel functions while $(h_Z,h_W)$ denote two bandwidths converging to zero as the sample size increases. The mapping $r$ is instead estimated by
To estimate $\widetilde{\varphi}$, i.e. the solution to Equation ((ref)), we use a Tikhonov regularization scheme $$\widehat \varphi =\left(\lambda I+\widehat A^*\widehat A\right)^{-1}\widehat A^* \widehat r,$$ where $I$ stands for the identity operator, while $\lambda>0$ denotes the Tikhonov regularization parameter. See, e.g., carrasco_chapter_2007 and darolles2011nonparametric.
To compute $\widehat{\varphi}$ we need to select two bandwidths $(h_Z,h_W)$ and a regularization parameter $\lambda$. We will explain below how to do it. For the moment, let us outline the computation of $\widehat{\varphi}$ in the previous display for given $(h_Z,h_W,\lambda)$. Although $\widehat{\varphi}$ seems to have an abstract expression, its computation is straightforward as it reduces to matrix products. To see this, we first approximate $\widehat{A}$ and $\widehat{A}^*$ by using a common bias computation similarly as in centorrino_additive_2017 :
The above approximations hold for $h_Z\, ,\, h_W\, \rightarrow 0$. Let $M_Z$ be the $n\times n$ matrix having on the $i$th row and $j$th column the element $K((Z_i-Z_j)/h_Z)/[\pi(Z_i) n h_Z]$ and let $M_W$ be the $n\times n$ matrix having on the $i$th row and $j$th column the element $K((W_i-W_j)/h_W)/[\tau(W_i) n h_W]$. Let $\overrightarrow{\widehat \varphi}:=(\widehat{\varphi}(Z_1),\ldots ,\widehat{\varphi}(Z_n))^\top$. For a generic function $\psi$ of $W$ let $\overrightarrow{\widehat{A}^*\psi}:=((\widehat{A}^*\psi)(W_1),\ldots ,(\widehat{A}^*\psi)(W_n))^\top$. Let $\overrightarrow{\widehat r}$ and $\overrightarrow{\widehat A \phi}$ (for a generic function $\phi$) be similarly defined. By the expression of $\widehat{\varphi}$ and the above approximations we get
so that
up to an approximation error. Thus, the computation of $\widehat{\varphi}$ at the data points reduces to a simple matrix computation and the residuals can be easily computed as $\widehat{U}_i=Y_i-\widehat \varphi(Z_i)$ with $i=1,\ldots,n$.
Let us turn now to the selection of the smoothing parameters $(h_Z,h_W)$ and the regularization parameter $\lambda$. The bandwidths $h_Z$ and $h_W$ are selected by the Silverman Rule of Thumb, so that $h_Z=\widehat{\sigma}_Z\,n^{-1/5}$ and $h_W=\widehat{\sigma}_W\,n^{-1/5}$, where $\widehat{\sigma}_Z$ and $\widehat{\sigma}_W$ denote the sample standard deviations of $Z$ and $W$. Finally, to select the Tikhonov regularization parameter, we use the Cross-Validation method developed in centorrino_additive_2017. Thus,
where $(\widehat A \widehat{\varphi}^{\lambda})_{-i}(W_i):=\left[ \widehat{A}_{-i}(\lambda I + \widehat{A}^*\widehat{A})^{-1}\widehat{A}^*_{-i}\widehat{r}\right](W_i)$. Here $\widehat{A}_{-i}$ and $\widehat{A}^*_{-i}$ denote the “leave-one-out" versions of $\widehat{A}$ and $\widehat{A}^*$ that use the entire sample except the $i$th observation. By the approximations of $\widehat{A}$ and $\widehat{A}^*$ previously outlined, the vector $\{(\widehat A \widehat{\varphi}^{\lambda})_{-i}(W_i):\ i=1,\ldots,n\}$ is computed as
where $diag(M_W)$ denotes the diagonal matrix having the same main diagonal as $M_W$, and $diag(M_Z)$ is similarly defined. The objective function in ((ref)) has a U-shaped form as a function of $\lambda$ and can be minimized in a simple way by numerical methods, see centorrino_additive_2017.
For the sake of clarity, we sum up the steps needed for computing $S_n$ as follows:
Implementation of the test. To implement the test, we are just left with the computation of the p-value. As we show in the next section, under the null hypothesis the statistic $S_n$ is asymptotically normal, but the asymptotic covariance has an intricate expression. So, to obtain the p-value necessary for testing we rely on the pairwise bootstrap. The validity of the pairwise bootstrap is confirmed by our simulations and is not surprising given that we show asymptotic normality. In fact, under different conditions, CLV show the validity of the bootstrap for a { real-valued} estimator based on a first step { infinite-dimensional} estimator, as is $\widehat\varphi$ in this paper. The procedure goes as follows:
In this section, we outline the assumptions under which we obtain the asymptotic properties of the test statistics. The following definition introduces some regularity features that the joint density $f_{ZW}$ and the NPIV function $\widetilde{\varphi}$ must satisfy. \\
Let $f_Z$ denote the density of $Z$ with respect to the Lebesgue measure and let us define $$g(z):=\mathbb{E}[(W-\mathbb{E}[W]) X|Z=z]f_Z(z)\, .$$
Assumption (ref) is a common regularity condition allowing for several degrees of integrability and differentiability florens_instrumental_2012. Under the above assumption $\widehat{A}:L^2_\pi ({\mathbb{R}})\mapsto L^2_\tau ({\mathbb{R}})$, $\widehat{A}^*: L^2_\tau ({\mathbb{R}}) \mapsto L^2_\pi ({\mathbb{R}})$, and $\widehat{r}\in L^2_\tau ({\mathbb{R}})$. Notice that $\widehat{A}^*$ is actually the Hilbert adjoint of $\widehat{A}$. This aspect is used multiple times in the proofs.
The kernel orders in Assumption (ref) are assumed to be equal only for notational simplicity. To simplify our theoretical exposition, we further assume that $h_W=h_Z$ and denote each bandwidth by $h$. However, our proofs also hold when such smoothing parameters are set to different values (and the kernels have different orders).
The order $\rho$ of the kernels $K_Z$ and $K_W$, the bandwidth $h$ and the regularization parameter $\lambda$ satisfy the following assumption.
Such conditions allow us to control the regularization bias and the variance of $\widehat{\varphi}$ appearing in the expansion of $S_n$. In particular, the condition on $n h^2 \lambda^{3/2}$ allows controlling the "variance term". The conditions on $n \lambda^2$ and $h^\rho \lambda^{-3/4}$ allow us to show that the regularization bias that appears in the expansion of $S_n$ is negligible. This bias is due to the ill-posed nature of the inverse problem.
Let us denote with $\left<\cdot,\cdot \right>$ the inner product of either $L^2_\pi(\mathbb{R})$ or $L^2_\tau(\mathbb{R})$, and let $\|\cdot\|$ be the norm induced by such an inner product. The specific inner product or norm we refer to will be clear at each time from the context. Let $(\lambda_j,\varphi_j,\psi_j)_j$ be the Singular Value decomposition of the operator $A$, where $(\lambda_j)_j$ is the sequence of singular values in $\mathbb{R}$, $(\varphi_j)_j$ is an orthonormal sequence in $L^2_\pi(\mathbb{R})$, and $(\psi_j)_j$ is an orthonormal sequence in $L^2_\tau(\mathbb{R})$ (see kress_linear_1999). The following assumptions introduces the usual source conditions.
Source conditions are standard in the NPIV literature (see carrasco_chapter_2007 or darolles2011nonparametric). In general, the source condition is imposed only on $\widetilde{\varphi}$ when the interest is on estimating $\widetilde{\varphi}$. Here, we need a source condition on both $\widetilde{\varphi}$ and $g$ to establish the asymptotic properties of the statistic based on the residuals from the nonparametric regression. To this end, we also need some regularity conditions on the estimator $\widehat{\varphi}$. So, let $N_{[\cdot]}(\epsilon,\Phi,||\cdot||_{\mathbb{P}})$ be the bracketing number of size $\epsilon$ of a class of functions $\Phi$, where $||\cdot||_{\mathbb{P}}$ denotes the $L_2$ norm with respect to the probability law of $Z$, see Van der Vaart (1998).
The above condition is a high-level assumptions allowing us to handle an empirical process in the expansion of our statistic. It essentially introduces a Donsker feature for the function $\widehat{\varphi}$. Such an assumption has been used among others by rothe_semiparametric_2009, escanciano_identification_2016, and mammen_semiparametric_2016. Sufficient conditions for it can be found in vaart_asymptotic_1998 or vaart_weak_1996. We have also derived alternative proofs based on sample splitting or cross-fitting that avoid such a high-level condition. We have run simulations with cross-fitting and sample-splitting, but we did not notice any improvement in terms of finite sample performance. Thus, we choose to keep the above assumption and avoid sample splitting (or cross-fitting) for a simple implementation of the test. Alternatively such a high-level condition could be avoided by using a Sobolev penalized method for estimating $\varphi$. This however would produce much longer proofs and complicate the implementation of the test.
Let ${\mathbb{P}}_n$ be the empirical mean operator. We have the following theorem.
The proof of the Theorem is in Section S2.1 of the supplementary material. Let us briefly comment on the asymptotic expansion (influence-function representation) of the statistic. The first term on the right hand side is the version of our statistic that we could use had we observed the error $U$ and $\mathbb{E}[W]$. The second term arises because the expectation $\mathbb{E}[W]$ is unobserved and is replaced by its estimated counterpart. Similarly, the third term arises because the error $U$ is unobserved and to estimate it we replace the true function $\varphi$ with its estimator $\widehat{\varphi}$. More in detail, such a term originates from a (nonparametric) bias involving the difference $\widehat{\varphi}-\varphi$ that appears in the expansion of $S_n$. This inflates the asymptotic variance of the statistic, as compared to the case where $U$ is known, and hence represents the price to pay for not observing the error term. To handle this bias term, we rely on decompositions from darolles2011nonparametric and babii2017unobservables.\\ The expansion in Theorem (ref) allows us to obtain the asymptotic normality of our statistic under the null hypothesis. To this end, notice that the first term has zero expectation under the null of {homogeneous treatment effects}. The second term has clearly a null expectation. Finally, since $\mathbb{E}[U|W]=0$, the last term also has zero expectation. Therefore, a standard Central Limit Theorem implies that $S_n \overset{d}{\to} \mathcal{N}(0,\Psi)$, where $\Psi$ is a covariance matrix defined by the influence function representation in Theorem (ref). Since such a covariance matrix has an intricate expression, it would be unconvenient to estimate it in practice. Thus, we suggest to bootstrap the statistic $S_n$ by the pairwise scheme to obtain the p-values of the test. As for the linear test, the fact that $|S_n|/\sqrt{n}\to C\ne 0$ when $\mathbb{E}[\widetilde{U}X(W-\mathbb{E}[W])]\neq 0$ holds implies that the power of the test goes to $1$ as $n$ goes to $\infty$ under alternative hypotheses satisfying $\mathbb{E}[\widetilde{U}X(W-\mathbb{E}[W])]\neq 0$.
Notice that we are keeping Assumption (ref) both under the null of {homogeneous treatment effects} and under the alternative. We have chosen to do this to simplify the exposition. However, it is possible that if $\mathbb{E}[\widetilde U X (W-\mathbb{E}[W])]\neq 0$ (and {treatment effects are not homogeneous}), the NPIV model might be misspecified and/or completeness might not hold. In such a case we would need to modify the proof about the power of the test. Let us briefly discuss these modifications. First, when the model is misspecified and completeness does not hold, it is possible to show that the Tikhonov regularized estimator $\widehat{\varphi}$ converges to $\widetilde{\varphi}^{\perp}$ that is the element of $\mathcal{N}(A)^{\perp}$ (the orthogonal complement of the null space of $A$) that solves $\min_{\phi\in \mathcal{N}(A)^{\perp}}\|r-A\phi\|$. In this case, we could define $\widetilde{U}$ as $Y-\widetilde{\varphi}^\perp (Z)$ and modify the statement of Theorem (ref) accordingly.
We set $\varphi(z):=\left(1+\exp(-z)\right)^{-2};\ X\sim\mathcal{N}(-0.5,1);\ W\sim\mathcal{N}(0,1);\ V\sim\mathcal{N}(0,1);\ Z=0.4 W + 0.2 V; \ U(z)=(V+X)\,(1+\gamma z)\text{ for all } z\in{\mathbb{R}},$ where $X,W,V$ are mutually independent. The outcome $Y$ is equal to $\varphi(Z)+U(Z)$. When $\gamma=0$ we are under the null of homogeneous treatment effects, while for $\gamma\neq 0$ we are under the alternative hypothesis. So, $\gamma$ represents the magnitude of the departure from the null.
To check the robustness of our test with respect to the choices of the smoothing and regularization parameters, we let them vary around benchmark values. So, we set $(h_Z,h_W)=C_h (h^*_Z,h_W^*)$ and $\lambda=C_{\lambda} \lambda^*$, where $(h_Z^*,h_W^*)=(\text{sd}(Z)n^{-1/5}, \text{sd}(W)n^{-1/5})$, where $\text{sd}(Z)$ (resp. $\text{sd}(W)$) is the empirical standard deviation of $Z$ (resp. $W$), while $\lambda^*$ is the regularization parameter selected according to the Cross-Validation method, as seen in Section (ref).. Note that $(h_Z^*,h_W^*)$ are the bandwidths set according to a version of Silverman's Rule of Thumb (silverman1986density). $C_h$ and $C_{\lambda}$ are fixed constants. We run simulations for $C_h,C_{\lambda}\in\{0.5,1,2\}$ and we consider the modest sample sizes of $n=100,250,500$. As a kernel we use the standard Gaussian density. We set $\pi$ to the Gaussian density with mean equal to the sample mean of $Z$ and variance equal to twice the sample variance of $Z$. Similarly, the measure $\tau$ is set to the Gaussian density with mean equal to the sample mean of $W$ and variance equal to twice the sample variance of $W$. florens_instrumental_2012 used a similar strategy to select $\pi$ and $\tau$ in their simulation setting. Intuitively, since $\pi$ and $\tau$ have to be two densities and at the same time appear at the denominators in $\widehat{A}$ and $\widehat{A}^*$, we want them (a) to integrate to 1 and (b) to converge towards zero sufficiently slowly on the tails. To speed up computations, we use the warp-speed method by giacomini_warp-speed_2013: for each Monte Carlo iteration we draw a single bootstrap sample, and we use the whole set of bootstrap statistics to compute the bootstrap p-values associated with each original statistic. We perform a very large number of Monte Carlos iterations equal to 10,000.
The results under the null hypothesis of {homogeneous treatment effects} ($\gamma=0$) are reported in Table (ref). The tests are implemented at 5 and 10 percent nominal levels. The nominal sizes of each of the tests are in bold. The result show that the test is reasonably stable with respect to the choices of $\lambda$, $h_Z$, and $h_W$ and controls well the size under the null hypothesis. Also, as expected, the error in the rejection probability (i.e. the difference between the empirical rejection proportions and the nominal size of the test) tends to become smaller as the sample size increases.
To check the power properties, we run simulations under the alternative hypothesis for different values of $\gamma$. The results are shown in Figure (ref). We only report results for the benchmark values of $h_Z$, $h_W$, and $\lambda$ and for the nominal size of 5 percent. For the other choices of $h_Z$, $h_W$, and $\lambda$ and for the 10 percent nominal level the results are qualitatively similar. The test shows good power under the alternative hypothesis. As the departure from the null increases, i.e. $\gamma$ gets further from 0, the test rejects the null hypothesis with an increasing frequency. Also, such a rejection frequency increases with the sample size at every value of $\gamma$.
To sum up, these simulations show that (a) for the moderate sample sizes of $n=100,250,500$ the tests display a good performance both in terms of size control and in terms of power, and (b) the results are stable with respect to the choices of the smoothing and regularization parameters.
In this section we apply the test of {homogeneous treatment effects} to a real dataset. Note that the application of returns to schooling is not well-suited to the nonparametric test because, in this application, the endogenous variable (education) has a support much larger than the instrument (college proximity) which is binary.
Instead, we focus on estimating a fish demand equation.This describes the relationship between the demanded quantity of fish and its price. It is common practice in econometrics to assume that prices are endogenous and estimate demand equations by instrumental variable regressions.
The data we use come from graddy2006markets and contain daily observations at the New York Fulton fish market about the log of the price, the log of the sold quantity, an indicator for each day of the week, and the wind speed registered in the previous day. The dataset can be downloaded at “\url{https://users.ox.ac.uk/ nuff0078/EconometricModeling/}”. The total number of observations is 111.
We study a demand equation. The outcome $Y$ is equal to the logarithm of the quantity sold and $Z$ is the logarithm of the market price. Market price is likely to be endogenous since it also depends on expected demand, which itself also affects the quantity. Following the literature ( graddy2006markets), we choose a variable related to weather as an instrument, so we set $W$ equal to the wind speed recorded on the day corresponding to the observation. Such a variable is viewed as sufficiently correlated with the price ($Z$), since the weather affects the ability to fish and is at the same time exogenous with respect to the errors $\{U(z)\}_z$ because it is probably unrelated to demand shocks related to human factors.
The choice of the covariable $X$ is crucial. We set $X$ equal to the indicator for week days (Monday to Thursday). Such a variable should not be correlated with the wind speed since the weather does not depend on the day of the week. In fact, a Kolmogorov-Smirnov test does not reject the hypothesis of independence between $X$ and $W$ (p-value of $0.17$). Hence, we can reasonably assume that $(\{U(z)\}_z,X)\protect\mathpalette{\protect\independenT}{\perp} W_k,$ which corresponds to Assumption (ref).
We perform the nonparametric test using the cross-validated penalty $\lambda^*$ and bandwidth $(h_Z^*,h_W^*)=(\text{sd}(Z)n^{-1/5}, \text{sd}(W)n^{-1/5})$, where $\text{sd}(Z)$ (resp. $\text{sd}(W)$) is the empirical standard deviation of $Z$ (resp. $W$) and $n$ is the sample size. This is a version of Silverman's Rule of Thumb (silverman1986density). We use $1,000$ bootstrap resamplings. The p-value of the test is $0.0362$. Hence, we reject the null hypothesis of {homogeneous treatment effects} at the 5 percent nominal level. So, we conclude that the causal intepretation of the separable NPIV regression should be doubted for this specific example. This is a reason to use a different approach such as instrumental variable quantile regression (chernozhukov2008instrumental) which allows heterogeneity of treatment effects by quantiles.
All the proofs along with some examples regarding the analysis of the power of the nonparametric test can be found in the online supplement.