Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
99,112 characters · 16 sections · 179 citation commands
Instrument Validity for Heterogeneous Causal Effects
\onehalfspacing
Keywords: {Instrument validity, heterogeneous causal effects, power improvement, extended continuous mapping theorem, extended delta method}
JEL Classification: {C10, C12, C14, C26}
The local average treatment effect (LATE) framework, introduced by the seminal works of imbens1994identification and angrist1996identification, is a commonly used approach in studies of instrumental variable (IV) models with treatment effect heterogeneity. The local quantile treatment effect (LQTE) is a concept similar to LATE. While LATE shows the treatment effect on the mean of the outcome, LQTE is more informative in regard to the effect on the outcome distribution.\footnote{See, for example, studies of LQTE in abadie2002bootstrap, ananat2008effect, cawley2012medical, frolich2013unconditional, and eren2014benefits.} These causal effect models rely on several strong and sometimes controversial assumptions of IV validity: 1) The instrument should not affect the outcome directly; 2) it should be as good as random assignment; and 3) it affects the treatment in monotone fashion. Violations of these conditions can generally lead to unidentification and inconsistent estimation of treatment effects. Relevant surveys and discussion of this can be found in angrist2008mostly, angrist2014mastering, imbens2014instrumental, imbens2015causal, koenker2017handbook, melly2017local, and huber2018local. Since the plausibility of the analyses of such models depends on IV validity, economics research has developed methods to examine these conditions based on testable implications.
Cases where the IV validity conditions may be violated can be found in empirical applications. For example, the college proximity was used as an instrument of education attainment in the study of NBERw4483. If the education level is treated as a binary variable (four-year college degree), the validity of the college proximity is rejected by the test of kitagawa2015test when no conditioning covariates are added in the model. mogstad2021causal considered tuition and college proximity as multiple instruments for college attendance. They showed that, if a homogeneity condition does not hold for individuals, the validity of the multiple instruments will be violated. The quarter of birth instrument used in angrist1991does is questionable because the exclusion restriction may not hold due to seasonal birth patterns bound1995problems,buckles2013season. The monotonicity condition of IV validity fails in the selection with two-way flows example in lee2018identifying.
kitagawa2015test was the first paper to propose a test of IV validity in heterogeneous causal effect models with a binary treatment based on the testable implications in the literature. It was the first to show the sharpness of these testable implications. Their test, constructed using a bootstrap method, was shown to be asymptotically uniformly size controlled and consistent. mourifie2016testing reformulated the testable implications used in kitagawa2015test as conditional inequalities. They then showed that these inequalities could be tested in the intersection bounds framework of chernozhukov2013intersection using the Stata package provided by chernozhukov2014implementing. The present paper provides a general framework for testing such IV validity assumptions. The proposed test can be applied in more general settings in which the treatment variable can be multivalued ordered or unordered\footnote{Studies of LATE with binary treatments can be found in angrist1990lifetime, angrist1991does, and vytlacil2002independence. Those with multivalued treatments can be found in angrist1995two, angrist1995split, and vytlacil2006ordered. Identification of causal effects in unordered choice (treatment) models can be found in heckman2006understanding, heckman2007econometric, heckman2008instrumental, and heckman2018unordered.} and the outcome variable can be unbounded\footnote{See reed2001pareto,reed2003pareto and toda2012double for the approximation of income distributions by members of the double Pareto parametric family.}. Also, the proposed test achieves power improvement by solving a technical issue and employing a novel bootstrap approach. huber2015testing derived a testable implication for a weaker LATE identifying condition, that is, that the potential outcomes are mean independent of instruments, conditional on each selection type.\footnote{The condition of potential outcomes being mean independent of instruments is not sufficient if we are concerned with distributional features of a complier's potential outcomes, such as the quantile treatment effects for compliers; see abadie2002instrumental for details.} The focus of the present paper is on full statistical independence of potential outcomes and instruments.
A modified variance-weighted Kolmogorov--Smirnov (KS) test statistic is employed in our test. As mentioned by kitagawa2015test, variance-weighted KS statistics have been widely applied in the literature on conditional moment inequalities, such as in andrews2013inference, armstrong2014weighted, armstrong2016multiscale, and chetverikov2017adaptive. More general KS statistics can be found in the stochastic dominance testing literature, such as in abadie2002bootstrap, barrett2003consistent, horvath2006testing, linton2010improved, barrett2014consistent, and donald2016improving. To investigate the asymptotic properties of the proposed test, we introduce $L^r$ ($r\in\mathbb{N}$) spaces with which a series of fundamental results are established, such as the compactness of particular function spaces, the Glivenko--Cantelli and the Donsker results, and so on. Based on these results, we obtain the asymptotic behavior of the test statistic.\footnote{See further discussion before Theorem (ref).} The asymptotic properties of the proposed test are established accordingly.
There are two major complications in deriving and approximating the asymptotic distribution of the test statistic under null. First, the test statistic involves a nonsmooth (nondifferentiable) map of unknown parameters (underlying probability distributions), and the delta method fails to work. We provide an extended continuous mapping theorem and an extended delta method, which might be of independent interest, to overcome this difficulty. By showing that the conditions of the extended delta method are satisfied under several weak assumptions, we establish the null asymptotic distribution of the test statistic. Second, since the null asymptotic distribution involves a nonlinear function, the standard bootstrap method may fail to approximate this distribution consistently. {Discussion of this issue can be found in dumbgen1993nondifferentiable, andrews2000inconsistency, hirano2012impossibility, hansen2017regression, hong2014numerical, and fang2014inference.} To achieve a consistent approximation, we extend the bootstrap approach proposed by fang2014inference\footnote{Other applications of this bootstrap method can be found in beare2015nonparametric, Beare2016global, Seo2016tests, Beare2015improved, and Beare2017improved. A similar bootstrap approach can be found in hong2014numerical.} and provide a valid bootstrap critical value. The test is found to be asymptotically size controlled and consistent. Evidence that the test performs well on finite samples is provided via simulations.
We now introduce the following notation, which will be used throughout the paper. We let $\leadsto$ denote Hoffmann--J\o rgensen weak convergence in a metric space. For a set $\mathbb{D}$, denote the space of bounded functions on $\mathbb{D}$ by $\ell^{\infty}(\mathbb{D})$: $ \ell^{\infty}\left( \mathbb{D}\right) =\left\{ f:\mathbb{D}\rightarrow \mathbb{R} :\left\Vert f\right\Vert _{\infty}<\infty\right\}$, { where } $\left\Vert f\right\Vert _{\infty}=\sup_{x\in\mathbb{D}}\left\vert f\left( x\right) \right\vert $. If $\mathbb{D}$ is a topological space, let $C\left( \mathbb{D}\right) $ denote the set of continuous functions on $\mathbb{D}$: $ C\left( \mathbb{D}\right) =\left\{ f:\mathbb{D}\rightarrow \mathbb{R} :f\text{ is continuous}\right\} $. Let $(\Omega,\mathcal{A},\mathbb{P})$ be a probability space on which all random elements are well defined. Let $\mathcal{B}_{\mathbb{R}^m}$ denote the Borel $\sigma$-algebra on $\mathbb{R}^m$ for all $m\in\mathbb{N}$. We use $\hat{Q}$ and $\hat{Q}^B$ to denote the empirical probability measure and the bootstrap empirical probability measure of each probability measure $Q$, respectively.
To formally introduce the topic of interest, we first consider the heterogeneous causal effect model of imbens1994identification. Let $Y\in\mathbb{R}$ be the observable outcome variable, and let $D\in\left\{ 0,1\right\} $ be the observable treatment variable, where $D=1$ indicates that an individual receives treatment. Let $Z\in\left\{ 0,1\right\} $ be a binary instrumental variable. Let $Y_{dz} \in\mathbb{R}$ be the potential outcome variable\footnote{See rubin1974 and splawa1990application for further discussion of the potential outcomes.} for $D=d$ and $Z=z$, where $d,z\in\{0,1\}$. Similarly, let $D_{z}$ be the potential treatment variable for $Z=z$. The instrument validity assumption for binary treatment and binary IV is formalized as follows.
Assumption (ref) is from imbens1997estimating, but it does not require strict instrument monotonicity. In this paper, we are not concerned with the strict monotonicity assumption, which is also known as the instrument relevance assumption.\footnote{As mentioned by kitagawa2015test, the instrument relevance assumption can be assessed by inferring the coefficient in the first-stage regression of $D$ onto $Z$.}
For all Borel sets $B$ and $C$, we follow kitagawa2015test and define probability measures as follows:\footnote{For simplicity of notation, we implicitly assume that $(Y,D,Z)$ is $(\mathcal{A},\mathcal{B}_{\mathbb{R}^3})$-measurable.}
Under Assumption (ref)(i), we can define a potential outcome variable $Y_{d}$ such that $Y_{d}=Y_{d0}=Y_{d1}$ almost surely. imbens1997estimating showed that for every Borel set $B$,
To see why (ref) is true, we can write
where the second equality follows from Assumptions (ref)(i) and (ref)(ii) and the third equality follows from Assumption (ref)(iii). Similar reasoning yields the second equation in (ref). Since the probabilities in (ref) are nonnegative, we obtain the testable implication of Assumption (ref) in balke1997bounds, imbens1997estimating, and heckman2005structural: For all $B\in\mathcal{B}_{ \mathbb{R} }$,
To understand (ref) graphically, suppose that $Y$ is a continuous variable and that $p_z\left( y,d\right) $ is the derivative of the function $P_z\left( (-\infty,y],\{d\}\right)$ with respect to $y$ for all $d,z\in\{0,1\}$. The following graphs show a case where (ref) holds.
The first inequality in (ref) is shown in Figure (ref), where the derivative $p_1\left( y,1\right) $ is greater than $p_0\left( y,1\right) $ everywhere. The second inequality in (ref) is shown in Figure (ref), where the derivative $p_0\left( y,0\right) $ is greater than $p_1\left( y,0\right) $ everywhere. Additional graphical examples can be found in kitagawa2015test.
Section (ref) discussed the case where the treatment and the instrument are both binary. In many applications, $D$ and $Z$ can be multivalued. See, for example, angrist1995two, where the treatment variable is the number of years of schooling completed by a student and can take more than two values. Now suppose that $D\in\mathcal{D}=\left\{ d_{1},\ldots,d_J\right\} $\footnote{The support $\mathcal{D}$ can be generalized to the case where $\mathcal{D}=\left\{d_{1},d_{2},\ldots\right\} $. See details in the supplementary appendix.} and $Z\in\mathcal{Z}=\left\{ z_{1},\ldots,z_{K}\right\} $. We let $d_{\max}$ be the maximum value of $D$, and $d_{\min}$ the minimum value of $D$. Suppose there exist potential variables $Y_{dz}$ for $d\in\mathcal{D}$ and $z\in\mathcal{Z}$, and $D_{z}$ for $z\in \mathcal{Z}$. The IV validity assumption for multivalued treatment $D$ and multivalued instrument $Z$ is then formalized as follows.
Assumption (ref) is similar to Assumptions 1 and 2 of angrist1995two. Theorems 1 and 2 of angrist1995two showed that a weighted average of $K$ average causal responses can be identified under Assumption (ref). Since we allow multivalued $Z$, the monotonicity assumption needs to hold for each pair $(D_{z_{k}},D_{z_{k+1}})$. The next lemma establishes a testable implication of Assumption (ref).
Lemma (ref) generalized testable implication (ref) to the case where the treatment and the instrument can both be multivalued. The testable implication (first-order stochastic dominance) discussed by angrist1995two for Assumption (ref) is equivalent to (ref). Clearly, if $D$ and $Z$ are both binary as assumed in Section (ref), with $d_{\max}=1$ and $d_{\min}=0$, then (ref) is equivalent to (ref) and (ref) is implied by (ref). liu2020two proposed testable implication (ref) for the case where $\mathcal{D}=\{0,1,2\}$ and $\mathcal{Z}=\{0,1,2\}$, and (ref) is also implied by (ref) in this case. Thus, (ref) and (ref) together can be viewed as a generalized form of their condition.
Studies of identification of causal effects in unordered choice (treatment) models can be found in heckman2006understanding, heckman2007econometric, and heckman2008instrumental. heckman2018unordered showed that the assumptions\footnote{See heckman2018unordered for a discussion of these assumptions.} in the preceding literature could be relaxed, and they defined a new monotonicity condition for the identification of causal effects in such models. We follow heckman2018unordered and suppose that the support $\mathcal{D}$ of $D$ is an unordered set with $\mathcal{D=}\left\{d_{1},\ldots,d_{J}\right\} $ and that the support $\mathcal{ Z}$ of $Z$ with $\mathcal{Z}=\{z_1,\ldots,z_K\}$ can be {unordered} as well. The unordered monotonicity condition proposed by heckman2018unordered is as follows (Assumption A-3 of heckman2018unordered).
It is worth noting that in Assumption (ref), $D$ is allowed to be a vector random element. In the case where $D,Z\in\{0,1\}$, Assumption (ref) is equivalent to the assumption that $1\left\{ D_{1}=1\right\} \geq1\left\{ D_{0}=1\right\}$ almost surely or $ 1\left\{ D_{1}=1\right\} \leq1\left\{ D_{0}=1\right\} $ almost surely. In practice, we often assume a specific direction in the assumption, such as $1\left\{ D_{1}=1\right\} \geq1\left\{ D_{0}=1\right\}$ almost surely, which is equivalent to $D_{1}\geq D_{0}$ almost surely in Assumption (ref)(iii). With the specific direction, we can prespecify a set $\mathcal{C}\subset\mathcal{D}\times\mathcal{Z}\times\mathcal{Z}$ and assume that $1\left\{ D_{z^{\prime}}=d\right\} \le 1\left\{ D_{z}=d\right\}$ almost surely for all $(d,z,z^{\prime})\in\mathcal{C}$. {For example, in the above case where $D,Z\in\{0,1\}$ and $1\left\{ D_{1}=1\right\} \geq1\left\{ D_{0}=1\right\}$ almost surely, we let $\mathcal{C}=\{(0,0,1), (1,1,0)\}$.} With this monotonicity condition of specified direction, we introduce the IV validity assumption for unordered treatment.\footnote{The test proposed in this paper can be extended for Assumption (ref) in which the direction is not specified. See details in Appendix (ref).}
Under this assumption, we can define $Y_{d}$ such that $Y_d=Y_{dz}$ almost surely for all $z$, and hence
for all Borel sets $B$ and all $\left( d,z,z^{\prime}\right) \in\mathcal{C}$.
As shown in kitagawa2015test and mourifie2016testing, the testable implication in (ref) is sharp. When the treatment or the instrument is multivalued (ordered or unordered), the cases could be complicated. kedagni2020generalized considered testing the joint assumptions of instrument exclusion and statistical independence, which are parts of (and different from) Assumption (ref) and Assumption (ref). The exclusion condition of kedagni2020generalized is the same as that in the present paper (Assumption (ref)(i) and Assumption (ref)(i)). The statistical independence condition of kedagni2020generalized (the instrument $Z$ is jointly independent of $(Y_{d_1},\ldots, Y_{d_J})$) is weaker than (and implied by) the random assignment condition in the present paper (Assumption (ref)(ii) and Assumption (ref)(ii)). Thus, the underlying assumptions tested by kedagni2020generalized are weaker than those tested by the present paper.
kedagni2020generalized provided sharp testable implications (the generalized instrumental inequalities) for the joint assumptions of instrument exclusion and statistical independence. Consider a simple case where the outcome $Y\in \{0,1\}$, the treatment $D\in\mathcal{D}$ is multivalued, and the instrument $Z\in\mathcal{Z}$ is also multivalued. Suppose that the exclusion condition and the statistical independence condition hold. We can then define $Y_d=Y_{dz_1}=\cdots=Y_{dz_K}$ for every $d\in\mathcal{D}$. For each $y\in\{0,1\}$, every $d\in\mathcal{D}$, and every $z\in\mathcal{Z}$, we have that
which implies that
For all $y_1,\ldots,y_J\in\{0,1\}$,
It then follows that
Next, for every $j$ and every $y_j\in\{0,1\}$,
With (ref), we have that
where
The inequalities in (ref)--(ref) are the testable restrictions derived by kedagni2020generalized, which are different from the proposed testable implications in (ref)--(ref). kedagni2020generalized suggest using the approach of chernozhukov2013intersection to test the restrictions in (ref)--(ref).\footnote{See Section 5 of kedagni2020generalized.}
Though the underlying assumptions tested by kedagni2020generalized are weaker than those tested by the present paper, no evidence has been found that, in general, the testable implications of kedagni2020generalized are weaker than (or implied by) those proposed by the present paper. Thus, to the best of our knowledge, the testable implications in kedagni2020generalized and those in the present paper could be complementary to each other. That is, the proposed testable restrictions may not be sharp for the IV validity Assumptions (ref) and (ref), and the proposed test may not be testing all possible restrictions. In practice, we suggest that users first apply the method of chernozhukov2013intersection to test the restrictions in kedagni2020generalized, and then apply the proposed method to test the joint IV validity assumptions in the present paper. In this way, the test results could be more informative about which part of the IV validity assumptions may fail.
Another interesting question is that if we combine the inequalities of kedagni2020generalized and those of the present paper together, are they sharp for the IV validity assumptions? To show this, we may draw on the sharpness results of kitagawa2015test, mourifie2016testing, and kedagni2020generalized. However, this would not be straightforward because we now allow both the treatment and the instrument to be multivalued (ordered or unordered), and the IV validity assumptions involve more conditions (random assignment and monotonicity). Since this technical complication may be beyond the main context of the present paper, we leave it for future study as an independent topic.
To highlight the idea, we first introduce the test for the case where the treatment is multivalued ordered, with support $\mathcal{D=}\left\{d_{1},\ldots,d_J\right\} $. The unordered treatment case will be discussed as an extension in Section (ref). Appendix (ref) in the appendix extends the proposed test for the cases where conditioning covariates may be present. Also, we let $Z$ be multivalued with support $\mathcal{Z}=\{z_1,\ldots,z_K\}$. The test is constructed based on the testable implication given in (ref) and (ref). Without loss of generality, we assume that $d_{\min}=0$ and $d_{\max}=1$. In practice, we can always normalize $d_{\min}$ and $d_{\max}$ to $0$ and $1$, respectively. Then (ref) and (ref) are equivalent to
for all $k$ with $1\le k\le K-1$, all closed intervals $B$ in $\mathbb{R}$, each $d\in\{0,1\}$, and all $C=(-\infty,c]$ with $c\in\mathbb{R}$. Here, (ref) and (ref) originally require (ref) to hold for all Borel sets $B$. Similar to {Lemma B.7} of kitagawa2015test, {we can} show (by applying Lemma C1 of andrews2013inference) that (ref) holding for all closed intervals $B$ is equivalent to (ref) holding for all Borel sets $B$.
By definition, for all $B,C\in\mathcal{B}_{\mathbb{R}}$ and all $k$ with $1\le k\le K$, $ \mathbb{P}\left( Y\in B,D\in C|Z=z_{k}\right)={\mathbb{P}\left( Y\in B,D\in C,Z=z_{k}\right) }/{\mathbb{P}\left( Z=z_{k}\right) } $. We now define function spaces
Let $\mathcal{P}$ denote the set of probability measures on $(\mathbb{R}^{3},\mathcal{B}_{\mathbb{R}^{3}})$. We use an i.i.d.\ sample $\{\left( Y_i,D_i,Z_i \right)\}_{i=1}^{n} $ which is distributed according to some probability distribution $Q$ in $\mathcal{P}$, that is, that the measure $Q(G)=\mathbb{P}((Y_i,D_i,Z_i)\in G)$ for all $G\in\mathcal{B}_{\mathbb{R}^3}$, to construct a test for the testable implication given in (ref) and (ref) (or in (ref)). For every $Q\in\mathcal{P}$ and every measurable function $v$, by an abuse of notation we define
Define, by convention (see, for example, folland2013real), that
For each $Q\in\mathcal{P}$, the closure of $\mathcal{H}$ in $L^2(Q)$ is equal to $\bar{\mathcal{H}}$ (Lemma (ref)). For every $Q\in\mathcal{P}$ and every $\left( h,g\right) \in {\bar{\mathcal{H}}\times\mathcal{G}}$ with $g=(g_{1},g_{2})$, define
With (ref), $\phi_Q$ is always well defined. Then the null hypothesis equivalent to (ref) is
if the underlying distribution of the data is $Q$. Since $Q(v)$ is continuous on $L^2(Q)$, (ref) is equivalent to $\sup_{\left( h,g\right) \in {\bar{\mathcal{H}}\times\mathcal{ G}}}\phi_Q\left( h,g\right)\le 0$. The alternative hypothesis is naturally set to
Define the sample analogue of $\phi_Q$ by
where $\hat{Q}$ denotes the empirical probability measure of $Q$ such that for every measurable function $v$,
and $\{( Y_{i},D_{i},Z_{i})\}_{i=1}^n$ is the i.i.d.\ sample distributed according to $Q$.
The goal of this section is to construct a test for the $H_0$ in (ref). To evaluate the ability of the test to provide size control, we consider a “local” sequence of probability distributions $\{P_n\}_{n=1}^{\infty}\subset\mathcal{P}$ under which the testable implication is true and $P_n$ converges to some probability measure $P\in\mathcal{P}$. We introduce the next two assumptions to formalize the above settings.
Assumptions (ref) and (ref) assume an i.i.d.\ sample whose distribution $P_n$ is allowed to {change} as $n$ increases, and to converge to some probability distribution $P$ as defined in (3.10.10) of van1996weak. The local analysis of fang2014inference considered the case where the value of the underlying parameter may be close to a point at which the map involved in the test statistic is only directionally differentiable (not fully differentiable). A similar convergent distribution sequence was introduced to show the local size control of their test.\footnote{See Examples 2.1 and 2.2 of fang2014inference.} As will be shown later, our test statistic involves a nondifferentiable (neither fully nor directionally differentiable) map. We follow fang2014inference and assume such a convergent distribution sequence to show the local size control of our test.
Clearly, ${\mathcal{H}}\times\mathcal{ G}\subset L^2(P)\times(L^2(P)\times L^2(P))$. Under Assumption (ref), define a metric $\rho_{P}$ on $L^2(P)\times(L^2(P)\times L^2(P))$ by
for all $\left( h,g\right) ,\left( h^{\prime},g^{\prime}\right) \in L^2(P)\times(L^2(P)\times L^2(P))$ with $g=(g_1,g_2)$ and $g^{\prime}=(g_1^{\prime},g_2^{\prime})$. By Lemma (ref), the closure of $\mathcal{H}\times\mathcal{G}$ in $L^2(P)\times(L^2(P)\times L^2(P))$ under $\rho_{P}$ is equal to ${\bar{\mathcal{H}}\times\mathcal{G}}$, where $\bar{\mathcal{H}}$ is defined in (ref). Define
where $\hat{P}_n$ is the empirical probability measure of $P_n$ defined as in (ref). Under Assumption (ref), we mainly consider the nontrivial case where $\Lambda(P)>0$. Also, for every $Q\in\mathcal{P}$, define
for all $(h,g)\in \bar{\mathcal{H}}\times\mathcal{G}$ with $g=(g_1,g_2)$, where $Q^m(v)=[Q(v)]^m$ for all $m\in\mathbb{N}$ and all measurable $v$.
Lemma (ref) provides the pointwise ($P$ is fixed) asymptotic distribution of $\sqrt{T_n}( \hat{\phi}_{P_n}-\phi_P)$ as $P_n$ converges to $P$ under Assumption (ref). We note that the pointwise asymptotic distribution of $\sqrt{T_n}(\hat{\phi}_{P_n}-\phi_P)$ is different from the asymptotic distribution of $\sqrt{T_n}(\hat{\phi}_{P_n}-\phi_{P_n})$ which can be obtained by Theorem 3.10.12 of van1996weak. The weak convergence of $\sqrt{T_n}(\hat{\phi}_{Q}-\phi_{Q})$ uniform in $Q$ may be obtained under different assumptions following the notion of van1996weak. We derive the pointwise asymptotic distribution in order to obtain the null asymptotic distribution of the test statistic using the proposed extended delta method. See the discussion after Theorem (ref). Lemma (ref) also provides the asymptotic variance of $\sqrt{T_n}( \hat{\phi}_{P_n}-\phi_P)$, which is uniformly bounded by $1$ for all $K>1$. We used the quantity $\sqrt{T_n}$ instead of $\sqrt{n}$ to establish the asymptotic distribution in order to achieve a parameter-free bound for the asymptotic variance as shown in (ref). The quantity $T_{n}$ is asymptotically equivalent to $n$ in the sense that $T_{n}/n\rightarrow\prod_{k=1}^{K}\mathbb{P}\left( Z=z_{k}\right) $ in probability. If we use $\sqrt{n}$, the bound of the asymptotic variance may involve the underlying parameter $P$. In the binary instrument case where $Z\in\{0,1\}$, we let $m_0=\sum_{i=1}^{n}1\left\{ Z_{i}=0\right\} $ and $m_1=\sum_{i=1}^{n}1\left\{ Z_{i}=1\right\} $. It then follows that $T_{n}=m_0 m_1/n$ which is used in the test of kitagawa2015test. Suppose instead $Z\in\{0,1,2\}$, and we let $m_z=\sum_{i=1}^{n}1\left\{ Z_{i}=z\right\} $ for $z\in\{0,1,2\}$. Then $T_{n}=m_0 m_1 m_2/n^2$.
The bound in (ref) will be useful when we construct the test statistic. By (ref), for every $\left( h,g\right) \in\bar{\mathcal{H}}\mathcal{\times G}$ with $g=\left( g_{1},g_{2}\right) $, define the sample analogue of $\sigma_P^{2}\left( h,g\right) $ by
Note that for each $h\in\bar{\mathcal{H}}$ and each $g_{l}\in\mathcal{G}_{K}$, if $\hat{P}_{n}(g_{l})=0$ then $\hat{P}_{n}(h\cdot g_{l})=0$. By (ref), $\hat{\sigma}_{P_n}^2$ is well defined. Similar to (ref), we can find a bound for $\hat{\sigma}_{P_n}$ for every finite sample. It can be shown that for all $(h,g)$,
Clearly, the bounds for $\sigma_P$ and $\hat{\sigma}_{P_n}$ will decrease as $K$ increases.
We may extend the idea of kitagawa2015test and construct the test statistic to be
for some positive number (trimming parameter) $\xi$. Here, $\xi$ plays two roles: (1) Since $\hat{\sigma}_{P_n}$ can be zero, $\xi$ bounds the denominator away from zero;\footnote{In practice, when the sample size is small, it is possible that we only have a small number of observations for $Z=z_k$ for some $k$. In this case, we can use $\sqrt{n}$ instead of $\sqrt{T_n}$ to construct the test statistic. We then use (ref) to find an empirical bound for $\hat{\sigma}_{P_n}$, and use this bound to determine the values of $\xi$. We could also redefine the instrument $Z$, in some cases, such that we have more observations for each possible value of the redefined instrument. For example, we may define the new instrument $\tilde{Z}=1\{Z\ge z_0\}$ for some $z_0$. However, this will change the definitions of all types of individuals (always takers, compliers, defiers, and never takers). In this case, we need to guarantee that the instrument used in the empirical analysis and the instrument used in the test are the same. The test result for $\tilde{Z}$ may be false for $Z$.} (2) as shown in the Monte Carlo studies of kitagawa2015test and the present paper, different values of $\xi$, from small (close to 0) to large (close to 1), may lead to different powers of the test for the same data generating process (DGP), which could be close to 0. kitagawa2015test suggests that if there is no prior knowledge available about a likely alternative, the default choice of $\xi$ could be set to $0.07$ according to the simulation studies for the binary treatment and binary instrument case. They also suggest that users report test results using different values of $\xi$. This paper constructs the test statistic in a way that, loosely speaking, computes the weighted average of the test statistics in (ref) over $\xi$.\footnote{In this way, we can avoid repeating the test using the same data set but different values of $\xi$ and making a decision based on all these results. The potential issue of multiple comparisons can be prevented accordingly.} If we put all the weight on one particular value of $\xi$, the test statistic degenerates to the test statistic in (ref).
Let $\Xi$ be a predetermined closed subset of $[0,1]$ such that $0\not\in\Xi$. The set $\Xi$ contains all the values of $\xi$ used for constructing the test statistic. Only one of the values greater than (or equal to) the bound in Lemma (ref), say $1$, needs to be included in $\Xi$. The test statistic in (ref) reduces to the unweighted KS statistic when $\xi=1$. Let $\nu$ be a positive measure on $\Xi$.
Now we set the test statistic to
The measure $\nu$ could be a Dirac measure centered at some fixed $\xi\in\Xi$. This is equivalent to using a particular value for the trimming parameter to construct the test statistic as in (ref). Or $\nu$ could be a discrete or continuous probability measure that assigns probabilities to the elements of $\Xi$. This is equivalent to using a weighted average of the test statistics in (ref) over $\xi$. By using (ref), we take into account the fact that the values of $\xi$ may influence the power of the test, and we can also avoid the multiple testing issue. See the discussion in Section (ref) about the computational simplification of the test statistic in (ref). Define
Since $1_{\left\{ a\right\} \times\left\{ 0\right\} \times\mathbb{R}},-1_{\left\{ a\right\} \times\left\{ 1\right\} \times\mathbb{R}} \in{\mathcal{H}}$ for all $a\in\mathbb{R}$, $\Psi_{{\mathcal{H}}\times\mathcal{G}}$ and $\Psi_{\bar{\mathcal{H}}\times\mathcal{G}}$ are not empty.
In the following theorem, we establish the asymptotic distribution of the test statistic under null. We note that the $L^r$ ($r\in\mathbb{N}$) spaces play an important role in deriving this asymptotic distribution. For example, we show that $\bar{\mathcal{H}}$ is compact in $L^2(Q)$ for every $Q\in\mathcal{P}$ and $\bar{\mathcal{H}}\times\mathcal{G}$ is compact in $L^2(P)\times(L^2(P)\times L^2(P))$ under $\rho_P$ (constructed based on the $L^2$ norm). We obtain the Glivenko--Cantelli and the Donsker results using the $L^1$ and the $L^2$ norms. We also show that the weak limit $\mathbb{G}$ of $\sqrt{T_n}(\hat{\phi}_{P_n}-\phi_P)$ in Lemma (ref) has a continuous path under $\rho_P$.\footnote{See Appendix (ref) for more detailed results.} The weak convergence in Theorem (ref) is established by using these fundamental results.
Theorem (ref) provides the pointwise ($P$ is fixed) asymptotic distribution of the test statistic if the $H_0$ in (ref) is true with $Q=P_n$ for all $n$.\footnote{More precisely, the weak convergence in (ref) is under $P_n$.} To find this asymptotic distribution, we employed the pointwise weak convergence in Lemma (ref) and the extended delta method provided in Appendix (ref). Because of a nondifferentiability issue, the existing delta methods fail to work in establishing the weak convergence in (ref). In Appendix (ref), we provide an extended continuous mapping theorem and an extended delta method elaborated by Theorems (ref) and (ref), respectively, to deal with this technical issue. See further discussion in Remark (ref). Theorem (ref) can be viewed as an extension of Theorem 1.11.1 of van1996weak, and Theorem (ref) can be viewed as an extension of Theorem 3.9.5 of van1996weak and of Theorem 2.1 of fang2014inference. For simplicity of notation, we let
where $\mathbb{G}_0$ is some random element such that $ \mathbb{G}=\mathbb{G}_0 + \Lambda(P)^{1/2}\mathcal{L}'_P(Q_0)$,\footnote{See more details in the proof of Theorem (ref).} where $Q_0(v)=P(vv_0)$ for all suitable $v$, and for all $(h,g)\in\bar{\mathcal{H}}\times\mathcal{G}$,
It can be shown that $\mathcal{L}'_P(Q_0)\le 0$ on $\Psi_{\bar{\mathcal{H}}\times\mathcal{G}}$ under $H_0$, and thus $\mathbb{G}\le\mathbb{G}_0$. Following the proof of Theorem (ref), we can show that
It then follows that $\mathbb{T}\le\mathbb{T}_0$. When $P_n$ is fixed at some $P$ for all $n$, then $v_0=0$ and $\mathbb{G}_0=\mathbb{G}$.
It was shown that the asymptotic distribution in (ref) involves the set $\Psi_{{\mathcal{H}}\times\mathcal{G}}$ that depends on the underlying probability measure $P$. Therefore, we need to find a “valid” estimator $\widehat{\Psi_{{\mathcal{H}}\times\mathcal{G}}}$ for $\Psi_{{\mathcal{H}}\times\mathcal{G}}$ in order to consistently approximate the asymptotic distribution. By the definition of $\Psi_{{\mathcal{H}}\times\mathcal{G}}$ in (ref), we construct $\widehat{\Psi_{{\mathcal{H}}\times\mathcal{G}}}$ by
with $\tau_{n}\rightarrow\infty$ and $\tau_{n}/\sqrt{n}\rightarrow0$ as $n\rightarrow\infty$, where $\xi_0$ is a small positive number. We suggest using $\xi_0=0.001$ in practice.\footnote{It can be shown that ${\widehat{\Psi_{{\mathcal{H}}\times\mathcal{G}}}}$ can also be used to approximate the asymptotic distribution when $\mathcal{D}=\{d_1,d_2,\ldots\}$. See (ref).} This is a method similar to that which is used in Beare2015improved and Beare2017improved to estimate contact sets in independent contexts. {See linton2010improved and lee2018testing for further discussion of estimation of contact sets.} Each $(h,g)$ is included in $\widehat{\Psi_{{\mathcal{H}}\times\mathcal{G}}}$ if $\sqrt{T_n}|\hat{\phi}_{P_n}(h,g)|$ is no more than $\tau_n$ estimated standard deviations from zero. As mentioned by Beare2017improved, we can effectively use pointwise confidence intervals to select points in this way.
We implement the test in the following sequence of steps:
The following theorem presents the asymptotic properties of the proposed test. Under Assumption (ref), Theorem (ref)(i) provides the local size control of the test. As discussed in fang2014inference, the asymptotic distribution of the test statistic may discontinuously depend on the parameter of interest, if the map involved in the test statistic is not fully differentiable. However, the finite sample distribution of the test statistic often continuously depends on the parameter of interest. imbens2004confidence emphasize that this discrepancy may cause poor finite sample properties of the test. As suggested by fang2014inference, a local analysis can help better approximate the finite sample properties of the test when the parameter of interest is close to a point at which the map is not fully differentiable. Our test statistic involves a nondifferentiable map, and Theorem (ref)(i) provides evidence for the good finite sample size property of the test.
It is implied by Theorem 11.1 of davydov1998local that in (i) of Theorem (ref), the CDF of $\mathbb{T}_0$ is differentiable and has a positive derivative everywhere except at countably many points in its support, provided that $\mathbb{T}_0\neq0$. If $\mathbb{T}_0=0$ at null configurations, our test statistic converges to zero in probability and so does the critical value. Theorem (ref) does not show clearly how the rejection rate of the test will behave asymptotically in this case. As discussed in Beare2017improved, this is a common theoretical limitation for irregular testing problems. Tests based on the machinery of fang2014inference, and also those based on generalized moment selection andrews2010inference,andrews2013inference, may encounter this issue. One practical resolution is to replace the bootstrap critical value $\hat{c}_{1-\alpha}$ with $\max\{\hat{c}_{1-\alpha},\eta\}$ or $\hat{c}_{1-\alpha}+\eta$, where $\eta$ is some small positive constant. See, for instance, donald2016improving. Simulation results in Table (ref) showed that the empirical rejection rates of our test with $\eta=0$ ($\tau_n=2$) are well controlled by the nominal significance level when $\mathbb{T}_0=0$ under null configurations.\footnote{In Section (ref), the value of $\tau_n$ is chosen to be $2$.}
Theorem (ref)(i) shows that the test is locally size controlled for every convergent distribution sequence satisfying the null. The convergent probability distributions $\{P_n\}$ depend on $n$, that is, the underlying distribution $P_n$ of the data can be different for every $n$. As $n\to\infty$, $P_n$ (satisfying the null) converges to $P$ under Assumption (ref). Theorem (ref) provides the pointwise ($P$ is fixed) asymptotic distribution of the test statistic $TS_n$ for this convergent sequence of probability distributions $P_n\to P$. With this pointwise asymptotic distribution, we then obtain the local size control along such a probability distribution sequence: $\lim_{n\to\infty}\mathbb{P}(TS_n >\hat{c}_{1-\alpha})\le\alpha$. When both $D$ and $Z$ are binary, kitagawa2015test and the present paper consider testing the same null and alternative hypotheses. kitagawa2015test obtains the uniform size control for their test under different conditions. That is, $\limsup_{n\to\infty}\sup_{Q\in\mathcal{P}_0}\mathbb{P}(TS_n >\hat{c}^K_{1-\alpha})\le\alpha$, where $\mathcal{P}_0$ denotes the set of probability distributions in $\mathcal{P}$ that satisfy $H_0$, and the superscript “$K$” denotes the critical value of kitagawa2015test (the test statistic of kitagawa2015test is equivalent to that in the present paper when both $D$ and $Z$ are binary). Since Theorem (ref)(i) assumes $P_n\in\mathcal{P}_0$, clearly we have that for every $P_n$,
which indicates that the uniform size control of kitagawa2015test implies local size control. In general, without additional assumptions, the local size control of the proposed test does not directly imply the uniform size control over the class of data generating processes in the null.
In this section, we consider the special case where the treatment $D$ and the instrument $Z$ are both binary. We show how to achieve power improvement over the test of kitagawa2015test based on the results of kitagawa2015test and those from Section (ref). As shown in Section (ref), the null hypothesis for the testable implications consists of a set of inequalities. kitagawa2015test used an upper bound on the asymptotic distribution of the test statistic under null to construct the bootstrap critical value. The upper bound is identical to the asymptotic distribution when all the inequalities in the null are binding. Therefore, their test could be conservative. The present paper establishes the asymptotic distribution of the test statistic under null. We then construct the critical value based on this asymptotic distribution, rather than on an upper bound, and therefore improve the power of the test.
Let $z_1=0$, $z_2=1$, $d_{1}=0$, and $d_{2}=1$. The test statistic in (ref) is now numerically equal to the one constructed by kitagawa2015test if we let $\nu$ be a Dirac measure. Recall that the instrument is allowed to be multivalued under the constructions in Section (ref).\footnote{For the case where the treatment is binary and the instrument is multivalued, kitagawa2015test constructed the test statistic by first computing the normalized differences of two empirical probability measures between neighboring pairs of values of instruments (ordered according to the propensity score), and then taking the maximum value of all these differences. Since these differences can be mutually correlated, it would not be straightforward to obtain the asymptotic distribution of their test statistic and approximate its null distribution by bootstrap.}
We consider a simple case where $P_n=P$ for all $n$ and the $H_0$ in (ref) is true with $Q=P$. As introduced in Section (ref), we follow kitagawa2015test and define probability measures
for all $B,C\in\mathcal{B}_{\mathbb{R}}$. Now we define
and write $P_d(f)=\int f \,\mathrm{d}P_d$ for all measurable $f$ and each $d\in\{0,1\}$. kitagawa2015test showed that their critical value converged to the $1-\alpha$ quantile of the distribution $\sup_{f\in\mathcal{F}_b}\mathbb{G}_{H}(f)/(\xi\vee\sigma_{H}(f))$, where $H=\lambda P_1+(1-\lambda)P_0$, $\lambda=\mathbb{P}(Z=1)$, $\mathbb{G}_H$ is an $H$-Brownian bridge, and $\sigma_H(f)$ is the standard deviation of $\mathbb{G}_H(f)$, that is, $\sigma_H^2(f)=H(f^2)-H^2(f)$. Let $\mathcal{F}_b^{\ast}=\left\{f\in\mathcal{ F}_b:P_0(f)=P_1(f)\right\}$. Then it is easy to show that $H(f)=P_0(f)=P_1(f)$ for all $f\in\mathcal{ F}_b^{\ast}$. Let $\nu$ be a Dirac measure centered at some $\xi$. It can be shown that
where $\mathbb{T}$ is the asymptotic distribution of the test statistic in (ref) and “$\overset{L}{=}$” means equivalence in distribution.
kitagawa2015test constructed a pooled-data bootstrap approximation for the Gaussian process ${\mathbb{G}_H}/({\xi\vee\sigma_H})$, denoted by ${\mathbb{G}_H^{B}}/({\xi\vee\sigma_H^B})$, and then computed the bootstrap test statistic by $\sup_{f\in\mathcal{ F}_b}{\mathbb{G}_H^B(f)}/({\xi\vee\sigma_H^B(f)})$. This bootstrap statistic approximates the distribution of $\sup_{f\in\mathcal{ F}_b} {\mathbb{G}_H(f)}/({\xi\vee\sigma_H(f)})$. For the case where $D$ and $Z$ are both binary, we suggest modifying the test in Section (ref) to achieve power improvement over the test of kitagawa2015test.\footnote{The modification may also be applied to the case where $D$ is multivalued and $Z$ is binary.} Specifically, we first estimate $\mathcal{F}_b^{\ast}$ by a subset of $\mathcal{F}_b$, denoted by $\widehat{\mathcal{F}_b^{\ast}}$, in a way similar to (ref). Then we follow the bootstrap approach of kitagawa2015test to construct ${\mathbb{G}_H^B}$ and ${\sigma_H^B}$, and construct the bootstrap test statistic by $\sup_{f\in\widehat{\mathcal{ F}_b^{\ast}}}{\mathbb{G}_H^B(f)}/({\xi\vee\sigma_H^B(f)})$. Clearly, the proposed critical value is always smaller than that of kitagawa2015test because $\widehat{\mathcal{F}_b^{\ast}}\subset\mathcal{F}_b$. It can also be shown that our critical value converges to the $1-\alpha$ quantile of $\sup_{f\in\mathcal{ F}_b^{\ast}}{\mathbb{G}_H(f)}/({\xi\vee\sigma_H(f)})$ (equivalently, $\mathbb{T}$) under $H_0$. Since the test statistic in (ref) is numerically equivalent to that of kitagawa2015test, this shows that the power of the test can be improved by the use of our approach. This improvement is against all alternatives according to the construction of the critical value. See the simulation evidence in Appendix (ref).
The test of mourifie2016testing for the inequalities in (ref) employed the intersection bounds framework of chernozhukov2013intersection. As shown in Proposition 1 of mourifie2016testing,\footnote{See also Theorem 6 of chernozhukov2013intersection.} the limiting rejection rate under null is equal to the nominal significance level $\alpha$ only when all the inequalities in the null are binding, and below the nominal level elsewhere in the null. This result is similar to that of kitagawa2015test, because only when all the inequalities are binding, the (contact) set $\mathcal{F}_b^{\ast}$ is equal to $\mathcal{F}_b$. If we are at a point in the null where the inequalities are not all binding, then the tests of kitagawa2015test and mourifie2016testing would have limiting rejection rates below the nominal level, and thus lack power against nearby points in the alternative. Theorem (ref) in the present paper shows that the proposed test can achieve the nominal level over a larger region in the null, where the inequalities in the testable implication could not all be binding, thereby improving power.
With testable implication (ref), we define the function space
For every probability measure $Q$ with (ref), we define $\phi_Q$ by $ \phi_{Q}\left( h,g\right) ={Q\left( h\cdot g_2 \right) }/{Q\left( g_2 \right) }-{Q\left( h\cdot g_1\right) }/{Q\left( g_1 \right) } $ for every $(h,g)\in\mathcal{H}\times\mathcal{G}$ with $g=(g_1,g_2)$. Testable implication (ref) is equivalent to the $H_0$ in \[ H_0: \sup_{(h,g)\in\mathcal{H}\times\mathcal{G}}\phi_{Q}\left( h,g\right) \le 0 \text{ and } H_1: \sup_{(h,g)\in\mathcal{H}\times\mathcal{G}}\phi_{Q}\left( h,g\right) > 0 \] if $Q$ is the underlying probability distribution of the data. Then we can follow the test procedure in Section (ref) to conduct the test with the function space $\mathcal{H}\times\mathcal{G}$ defined in (ref).
We first designed Monte Carlo simulations for the case where $D$ and $Z$ are both multivalued random variables such that $D\in\{0,1,2\}$ and $Z\in\{0,1,2\}$. Additional Monte Carlo studies can be found in Appendix (ref). Each simulation consisted of $1000$ Monte Carlo iterations and $1000$ bootstrap iterations. To expedite the simulation, we employed the warp-speed method of giacomini2013warp. The nominal significance level $\alpha$ was set to $0.05$. As shown in (ref) and (ref), $\sigma_P^2$ and $\hat{\sigma}_{P_n}^2$ are bounded by $(1/2)\cdot (K-1)^{-(K-1)}$, where $K=3$ in our setting. The simulations constructed in this section are similar to those in kitagawa2015test. In each simulation, the measure $\nu$ was set to be a Dirac measure $\delta_{\xi}$ centered at one of the following values of $\xi$: $0.07$, $0.1$, $0.13$, $0.16$, $0.19$, $0.22$, $0.25$, $0.28$, $0.3$, and $1$, or to be a probability measure $\bar{\nu}_{\xi}$ that assigns equal probabilities (weights) to the values of $\xi$ listed above. Four values of $\xi$ were used in the simulations of kitagawa2015test: $0.07$, $0.22$, $0.3$, and $1$, where $0.07\approx\sqrt{0.005(1-0.005)}$, $0.22\approx\sqrt{0.05(1-0.05)}$, and $0.3=\sqrt{0.1(1-0.1)}$. As shown in (ref), for every $(h,g)\in\bar{\mathcal{H}}\times\mathcal{G}$ with $g=(g_1,g_2)$,
The values of $\xi\in\{0.07,0.22,0.3\}$ take the form of $\sqrt{\pi(1-\pi)}$ where $\pi\in\{0.005,0.05,0.1\}$. As discussed in kitagawa2015test, $\pi$ can be interpreted as that if both ${ \hat{P}_n\left( h^2\cdot g_{1}\right) }/{\hat{P}_n\left( g_{1}\right) }$ and ${ \hat{P}_n\left( h^2\cdot g_{2}\right) }/{\hat{P}_n\left( g_{2}\right) }$ are less than $\pi$, then the weight becomes the inverse of $\xi$ instead of the inverse of the estimated standard deviation. As $\pi$ gets larger, less weight is put on $\hat{\phi}_{P_n}$ for smaller probability events, and vice versa. In the following simulations, we chose the values of $\xi$ following the choice of kitagawa2015test. In empirical practice, application-based simulations can be applied to choose $\Xi$ and $\nu$, which is illustrated in Section (ref).
When calculating the supremum in the test statistic $TS_n$ in (ref), we followed the numerical computation approach used by kitagawa2015test. Specifically, we calculated the supremum using only the closed intervals $B$ with the values of $\{Y_i\}_{i=1}^{n}$ observed in the data as the endpoints, that is, $B=[a,b]$ with $a,b\in\{Y_1,Y_2,\ldots,Y_n\}$ and $a\le b$. It is not hard to show that the test statistic calculated in this way is equal to that in (ref). We also used such closed intervals to calculate the bootstrap test statistic $TS^B_n$ in (ref). From all such intervals, we found those that satisfy the inequality in (ref) and used them to calculate the supremum of $ {\sqrt{T_n^B}( \hat{\phi}_{P_n}^{B}-\hat{\phi}_{P_n}) }/\max\{\xi,{\hat{\sigma}_{P_n}^{B}}\}$ for each $\xi$ listed above.
The first set of simulations was designed to investigate the size of the test and the selection of the tuning parameter. As shown in (ref), the estimate $\widehat{\Psi_{{\mathcal{H}}\times\mathcal{G}}}$ involves a tuning parameter $\tau_n$ with $\tau_n\to\infty$ and $\tau_n/\sqrt{n}\to0$ as $n\to\infty$. In practice, we need to use a particular value of $\tau_n$ for each sample size $n$. For this set of simulations, we set $n$ to $3000$ and $\tau_n$ to $0.1$, $0.5$, $1$, $2$, $3$, $4$, and $\infty$. For $\tau_n=\infty$, $\widehat{\Psi _{{\mathcal{H}}\times\mathcal{G}}}={{\mathcal{H}} \times\mathcal{G}}$ and the test is conservative. We compared the rejection rates obtained using each of these values of $\tau_n$ and decided which value would be a good option for sample sizes close to $3000$. We let $U\sim\mathrm{Unif}(0,1)$, $V\sim\mathrm{Unif}(0,1)$, $N_0\sim \mathrm{N}(0,1)$, $N_1\sim \mathrm{N}(1,1)$, $N_2\sim \mathrm{N}(2,1)$, $Z=2\times 1\{U \le 0.5\}+1\{0.5<U \le 0.7\}$ ($\mathbb{P}(Z=2)=0.5$), $D_z=2\times 1\{V\le 0.33\}+ 1\{0.33 <V\le 0.66\}$ for $z=0,1,2$, $D=\sum_{z=0}^2 1\{Z=z\}\times D_z$, and $Y=\sum_{d=0}^{2}1\{D=d\}\times N_d$. All the variables $U$, $V$, $N_0$, $N_1$, and $N_2$ are mutually independent. Clearly, Assumption (ref) holds in this case with $z_1=0$, $z_2=1$, and $z_3=2$.
Table (ref) shows the results of the simulations. The rejection rates were influenced by the values of $\tau_n$ and $\xi$. For each measure $\nu$, a smaller $\tau_n$ yields greater rejection rates, because a smaller $\tau_n$ leads to a smaller critical value according to (ref). For $\tau_n=2$, all the rejection rates were close to those for $\tau_n=\infty$ (the conservative case). Similar to the pattern of the results shown in kitagawa2015test, some rejection rates for $\tau_n=2$ with $\delta_{\xi}$ centered at particular values of $\xi$ were slightly upwardly biased compared to the nominal size. Overall, however, the results showed good performance of the test in terms of size control. When sample sizes are less than or close to $3000$, we suggest using $\tau_n=2$ in practice to achieve good size control without a significant power loss. When the sample size increases, $\tau_n$ should be increased accordingly. It is also worth noting that when we used the measure $\bar{\nu}_{\xi}$, the rejection rates could be well controlled by the nominal significance level. Thus if we have no additional information about the choice of $\xi$, $\bar{\nu}_{\xi}$ can be a default choice for us.
The second set of simulations was designed to investigate the power of the test. Six data generating processes (DGPs) in total were considered, and Assumption (ref) did not hold with $z_1=0$, $z_2=1$, and $z_3=2$. Sample sizes were set to $n=200$, $600$, $1000$, $1100$, and $2000$. The probability $\mathbb{P}(Z=2)=r_n$, with $r_n=1/2$, $1/6$, $1/2$, $1/11$, and $1/2$ for the corresponding sample sizes. We set $\tau_n$ to $2$, as suggested in the preceding set of simulations. DGPs (1)--(4) are the cases where (ref) was violated and (ref) was not violated, and DGPs (5) and (6) are the cases where both (ref) and (ref) were violated. We let $U\sim\mathrm{Unif}(0,1)$, $V\sim\mathrm{Unif}(0,1)$, $W\sim\mathrm{Unif}(0,1)$, and $Z=2\times 1\{U \le r_n\}+1\{r_n<U \le r_n+0.2\}$.
For DGPs (1)--(4), we let $D_z=2\times 1\{V\le 0.45\}+ 1\{0.45 <V\le 0.55\}$ for $z=0,1,2$, $D=\sum_{z=0}^2 1\{Z=z\}\times D_z$, $N_{00}\sim \mathrm{N}(0,1)$, $N_{10}\sim \mathrm{N}(0,1)$, and $N_{dz}\sim \mathrm{N}(0,1)$ for $d=0,1,2$ and $z=1,2$.
For DGPs (5) and (6), we let $N_0\sim \mathrm{N}(0,1)$, $N_1\sim \mathrm{N}(1,1)$, and $N_2\sim \mathrm{N}(2,1)$.
All the variables $U$, $V$, $N_{00}$, $N_{10}$, $N_{20}$, $N_{01}$, $N_{11}$, $N_{21}$, $N_{02}$, $N_{12}$, $N_{22}$, $N_0$, $N_1$, and $N_2$ were set to be mutually independent. We briefly explain how DGPs (1)--(4) violate (ref), which is shown graphically in Figure (ref). We let $p_z(y,d)$ be the derivative of $\mathbb{P}(Y\in(-\infty,y],D=d|Z=z)$ with respect to $y$ for all $d,z\in\{0,1,2\}$. Similar to Figure (ref), if (ref) were true, then we would have $p_0(y,2) \le p_1(y,2) \le p_2(y,2)$ everywhere. For DGPs (1)--(4), $p_1(y,2)=p_2(y,2)$ held for all $y$, but $p_0(y,2) \le p_1(y,2)$ did not hold on some range of $\mathbb{R}$. DGPs (5) and (6) are the cases where the monotonicity assumption did not hold and both (ref) and (ref) were violated.
Table (ref) shows the rejection rates under DGPs (1)--(6), that is, the power of the test. For each DGP and each measure $\nu$, the rejection rate increased as the sample size $n$ was increased. The results for $\nu=\bar{\nu}_{\xi}$ showed that if we have no information about the choice of $\xi$, using the weighted average of the statistics over $\xi$ is a desirable option. When $n>200$, the rejection rates for using $\nu=\bar{\nu}_{\xi}$ were at a relatively high level compared to the results for using a Dirac measure.
We revisit one empirical example discussed by kitagawa2015test to show the performance of the proposed test in practice. The example is from NBERw4483, who used college proximity as an instrument of years of schooling to study the causal link between education and earnings. The data are from the Young Men Cohort of the National Longitudinal Survey. In the original study of NBERw4483, the educational level $D$ is a multivalued treatment variable, while kitagawa2015test treated it as a binary treatment variable $T$ with $T=1\{D\ge 16\}$. The results of the test of kitagawa2015test showed that the instrument was not valid when no covariates were controlled.
We use the originally defined treatment variable $D$ to reconduct the test. Specifically, the treatment $D$ is education attainment observed in 1976 (the variable “ed76”), the instrument $Z$ is whether an individual grew up near a 4-year college (the variable “nearc4”), and the outcome is log wage observed in 1976 (the variable “lwage76”) in the data set. The available sample size is 3010. We follow the setup in Section (ref) with $\mathcal{D}=\{1,2,\ldots,18\}$ and $\mathcal{Z}=\{0,1\}$. The instrument $Z=1$ implies that an individual grew up near a 4-year college. {Table} (ref) shows the $p$-values obtained from our test using each measure $\nu$. From these results, we conclude that we do not reject the validity of instrument $Z$. In Section (ref), we show more results by using application-based simulations to choose $\Xi$ and $\nu$. The results are similar to those in Table (ref).
The testable implication used by kitagawa2015test for binary $T$ is that
for all closed intervals $B$. The inequalities in (ref) are equivalent to the following for all closed intervals $B$:
which are different from those in the testable implication given in (ref) and (ref) and are not implied by Assumption (ref). Thus a valid instrument $Z$ for multivalued $D$ which satisfies the testable implication given in (ref) and (ref) may not satisfy the inequalities in (ref), that is, $Z$ may not remain valid for binary (or coarsened) $T$. This provides a possible explanation for why we accept $Z$ but kitagawa2015test rejected it.
In this paper, we provided a general framework for testing instrument validity in heterogeneous causal effect models. We generalized the testable implications of the instrument validity assumptions in the literature, and based on them we proposed a nonparametric bootstrap test. An extended continuous mapping theorem and an extended delta method were provided to establish the asymptotic distribution of the test statistic, which may be of independent interest. The proposed test can be applied in more general settings and may achieve power improvement.
\setcounter{equation}{0}