Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
141,854 characters · 24 sections · 167 citation commands
Testing Instrument Validity with Covariates
Keywords: Program evaluation, local average treatment effect, marginal treatment effect.
Instrumental variable (IV) methods are widely used in empirical economics. The credibility of IV estimates relies on the validity of the identifying assumptions, which are often subject to debate. In IV estimation, conditioning covariates are frequently used to enhance the credibility of the identifying assumptions or improve the precision of estimates. For example, when using proximity to college as an instrument to estimate the return to schooling, researchers control for location characteristics which may otherwise induce dependence between the instrument and the outcome. hubermellace15Testing, Kitagawa2015, and mourifiewan propose statistical tests for instrument validity in the heterogeneous treatment effect model of imbensangrist. However, the state of the art literature lacks high-power tests for instrument validity that remain computationally feasible for more than a small number of covariates.
The first contribution of this paper is to propose an easy-to-implement test procedure for instrumental validity in the heterogeneous treatment effect model that can accommodate a moderate to large number of conditioning covariates. Building on the semiparametric approach of CHV11, we specify that conditioning covariates are independent of unobserved hetergoneity and affect potential outcomes linearly. Under these restrictions, the conditional expectation of the outcome given conditioning covaraites and the instrument follows a partially linear model that depends non-parametrically on the probability of receiving treatment (propensity score). This allows the conditioning covariates to be partialed out and partial residuals obtained. Testable implications for the instrumental variable identifying assumptions which hold without conditioning on covariates can be stated for the subdensities of these partial residuals. With partial residuals in hand, testing can be performed without reference to covariate values. In comparison to the test procedure with conditioning covariates proposed in Kitagawa2015, this greatly reduces the computational burden. Increasing the number of conditioning covariates increases only the complexity of estimating the propensity score and partially linear model, not the number of implications to be tested.
Two testable implications are available: index sufficiency and nesting inequalities. Index sufficiency states that the partial residual subdensities depend on the instrument through the propensity score only. The nesting inequalities are monotonicity inequalities for the joint distributions of the partial residuals and treatment status. They state that, for two distinct values of the propensity score, the distribution of partial residuals for treated units with the higher propensity score nests that of units with the lower propensity score, and the converse is true for untreated units. The two testable implications complement each other: index sufficiency holds the propensity score fixed and compares subdensities for different values of the instrument, while the nesting inequalities compare subdensities for different propensity scores. To exploit this complementarity, we propose jointly testing the two.
As the subdensity nesting inequalities are ordered with respect to the propensity score, a natural approach to testing them would be to coarsen the propensity score and test the monotonicity of the subdensities conditional on this coarsened propensity score. When the propensity score is strongly correlated with the conditioning covariates, this approach can result in a test with low power. In such cases, classifying observations by propensity score will homogenise the distribution of instrument values across propensity score bins, which can mask violations of instrument validity.
Our second contribution is to propose a novel approach to testing the nesting inequalities. We suggest testing instead monotonicity of the conditional partial residual subdensities given the instrument. This approach ensures that observations with different instrument values are kept separate. We show that, when the identifying assumptions are satisfied, the nesting inequalities hold over instrument values if the conditional distribution of propensity scores is monotonic in the instrument in the sense of first-order stochastic dominance. If first-order stochastic dominance does not hold for the instrument, we propose trimming the sample and testing the nesting inequalities for a sub-sample where first-order stochastic dominance is satisfied. We refer to this process as distillation as it distills the sample down to the observations best able to detect violations.
We show analytically and numerically the benefits of implementing distillation in terms of the ability to detect violation of instrument validity conditional on covariates. Specifically, we present a class of data generating processes for which distillation dominates conditioning on the coarsened propensity score in the following sense. A violation of the nesting inequalities conditional on the coarsened propensity score implies a larger violation of the nesting inequalities with a distilled sample, but not vice versa. That is, a violation of the nesting inequalities with a distilled sample does not imply the existence of a violation of the nesting inequalities conditional on the coarsened propensity score. Consequently, there exist data generating processes where instrument validity fails but the violation can only be detected by the nesting inequalities when performing distillation.
We perform Monte Carlo exercises to evaluate the size and power of the proposed test procedure and compare it to two alternatives: the Kitagawa2015 test that does not control for covariates, and the test of nesting inequalities using a coarsened propensity score. We consider processes where the instrument is independent of the covariates, and processes where it is not. In the first case, an instrument that is valid conditional on the covariates will also be valid without conditioning on covariates, so the size of the Kitagawa2015 test will not be distorted. In the second, size will only be controlled if the conditioning covariates are appropriately controlled for. We find that the size of the proposed test is controlled in all cases. For processes where the Kitagawa2015 test has the correct size, we find our test procedure has comparable power. For all processes, it has higher power than the test of nesting inequalities using a coarsened propensity score. We also investigate the contributions of the index sufficiency and nesting inequality tests to the power of the joint test. These results suggest that index sufficiency has most power when the distributions of propensity scores for different instrument values are similar, so that there are observations with different instrument values and a similar propensity score. On the other hand, for the processes considered, rejection probabilities when testing the nesting inequalities do not depend strongly on the similarity of the propensity score distributions across instrument values. This is consistent with the two testable implications being complementary.
To demonstrate the feasbility of the test procedure and illustrate how it can be used to resolve questions of instrument credibility, we present three applications. The first is Card1993. This instrument has been tested in hubermellace15Testing, Kitagawa2015, mourifiewan and sun2020 but, for computational reasons, in each case only a subset of the conditioning covariates of the original specification is used. The proposed test procedure allows all conditioning covariates to be included and remains computationally feasibile. Second, we apply the test to the same-sex instrument of angristevans1998. The validity of this instrument was been questioned in RosenzweigWolpin2000 and its validity has been formally tested by huber2015 and mourifiewan. As with Card1993, these studies use only a subset of the conditioning covariates, while our test procedure allows us to include all conditioning covariates. Finally, we turn to identifying strategies that rely on geographical variation in the timing of changes in compulsory schooling laws. Using data for America, stephensyang2014 presents evidence that changes in education laws are correlated with other unobserved state-level changes, which suggests that this instrument is not valid. We test its validity in the case of the UK using the data of oreopoulos2006.
The remainder of the paper proceeds as follows. The related literature is reviewed in Section (ref) below. Section 2 introduces the semiparametric testable implications. Section 3 describes the test procedure. Section 4 presents Monte Carlo exercises. Section 5 discusses our three empirical applications. Section 6 concludes.
This paper contributes to the active literature on the identification and estimation of the local average treatment effect (LATE) model of angristetal and imbensangrist and the marginal treatment effect model (MTE) of heckmanvytlacil99, heckmanvytlacil01, heckman_vytlacil_2001, heckmanvytlacil05. This literature has recently been reviewed by MogstadTorgovitsky2018. A testable implication of the LATE identifying restrictions was derived in balkepearl, imbensrubin and heckmanvytlacil05. Kitagawa2015 shows that this testable implication is sharp, and proposes a test procedure using a Kolmogorov-Smirnov statistic. mourifiewan reformulates this testable implication as a conditional moment inequality, and proposes a test based on the intersection bounds framework of chernozhukovleerosen2013. It is straightforward to generalise this testable implication to incorporate conditioning covariates, but assessing it nonparametrically in a finite sample is limited by the capacity of nonparametric methods to handle multiple conditioning covariates.
To facilitate the estimation of marginal treatment effects while controlling for multiple conditioning covariates, CHV11 further imposes that potential outcomes are linear in the conditioning covariates, unobserved heterogeneities enter additively, and these unobservables are statistically independent of the covariates and instrument. The test procedure presented in this paper corresponds to a test for the validity of the CHV11 identifying assumptions. MaestasMullenStrand2013, Eisenhaueretal2015, KlineWalters2016, Cornelissenetal2018, FelfeLalive2018, Bhulleretal2020, and CKST2022 among others apply the approach of CHV11 to recover marginal treatment effects. Our test procedure can be applied in any of these contexts to assess the validity of the identifying assumptions. Furthermore, even if the validity of the instrument is undisputed, the test procedure can still be used to detect errors in the functional form specification for the relationship between potential outcomes and conditioning covariates. brinchmogstadwiswall further develops the approach of CHV11, and show how marginal treatment effects can be identified in a variety of settings. mogstadetal discusses how to extrapolate marginal treatment effects to set identify policy relevant treatment effects.
Several alternative tests of the LATE identifying assumptions have been proposed. hubermellace15Testing considers a weaker set of identifying restrictions than Kitagawa2015, derives a testable implication, and proposes several test procedures. LaffersMellace2017 shows that the testable implication of hubermellace15Testing is sharp for these identifying assumptions. sun2020 derives testable implications when the treatment is an ordered or unordered multi-valued variable, and proposes a power improving refinement of the Kitagawa2015 test procedure. araietal2022 extends the Kitagawa2015 test to fuzzy regression discontinuity designs. farbmacherguberklaassen2022 proposes a test procedure that uses random forests and classification and regression trees to find violations of the monotonicity and exclusion assumptions. kedagnimourifie derives a testable implication of the instrument independence assumption in the form of a set of inequalities, reformulates these inequalities as a set of conditional expectations, and develops a test procedure using the framework of chernozhukovleerosen2013. machadoetal discusses tests for instrument validity in the case of a binary outcome which responds monotonically with respect to the treatment. frandsenetal2023 develops a test of the identifying assumptions of research designs which exploit random assignment of judges. The closest paper to ours is MaoSantAnna20, which also uses the MTE framework of CHV11, and derives index sufficiency and the nesting inequalities. This paper focuses on testing the nesting inequalities, and does not jointly test index sufficiency. The nesting inequalities are tested by considering many possible discretisations of the propensity score. We differ in that we jointly test index sufficiency and the nesting inequalities. In addition, for the nesting inequalities, rather than examining subdensities indexed by the propensity score, we propose constructing a distilled sample so that the inequalities hold for subdensities indexed by the instrument, and show the advantages of this approach.
We contribute to several empirical literatures by testing the validity of commonly used instruments and identification strategies. The college proximity instrument of Card1993 has also been employed in CameronTaber2004 to estimate the effects of borrowing constraints on education, CHV11 to estimate the marginal treatment effect of education, and Heckmanetal2018 to estimate the effects of education on labour market and health outcomes. The same-sex instrument of angristevans1998 has been used widely to identify the effects of family size. BlackDeveruxSalvanes2010,ConleyGlauber2006, AngristLavySchlosser2010 and Beckeretal2010 use this instrument to empirically investigate the existence of a quantity-quality trade-off for children, while CrucesGaliani2007 employs it to investigate the relationship between family size and maternal labour supply. Other than oreopoulos2006, geographical variation in schooling laws has been employed by acemogluangrist2000 to estimate human capital externalities, lochnermoretti2004 to estimate the effect of education on imprisonment rate, and ClarkRoyer2013 to estimate the effects of education on life expectancy.
The data is a random sample $\left( Y,D,X,Z\right) $, where $Y$ is an observed outcome continuously distributed on $\mathbb{R}$, $D\in \left\{ 1,0\right\} $ is an observed binary treatment status indicating treated $(D=1)$ or non-treated $\left( D=0\right) $, $X\in \mathcal{X}\subset \mathbb{R}^{k_{X}}$ is a vector of individual pretreatment observable covariates, and $Z\in \mathcal{Z}\subset \mathbb{R} ^{k_{Z}}$ is a vector of instrumental variables. For the moment, we impose no restrictions on the instrument. \ Let $\left\{ Y_{dz}\right\} _{d\in \left\{ 1,0\right\} ,z\in \mathcal{Z}}$ be the set of potential outcomes indexed by treatment status $d\in \left\{ 1,0\right\} $ and instrument value $z\in \mathcal{Z}$. The observed outcome $Y$ can be written as
We write a selection equation of unconstrained form as
We interpret this selection equation as follows. With $U_{D}$ fixed at $u_{D}$, $D_{xz}=1\{v(x,z,u_{D})\geq 0\}$ generates the counterfactual selection response for an individual when their pre-treatment characteristics $X$ and instrument $Z$ are exogenously set at $x$ and $z$. \\
The following three assumptions identify the marginal treatment effect (MTE):
MTE Identification Assumptions (Heckman and Vytlacil (2005)):
The instrument exclusion restriction (A1) precludes the instrument from having a direct causal impact on the outcome for any member of the population. With this restriction imposed, an individual's potential outcomes are reduced to a pair indexed only by their treatment status, i.e, $Y_{1}$ denotes their potential outcome with treatment and $Y_{0}$ their potential outcome without treatment. The observed outcome $Y$ can then be written as $Y=Y_{1}D+Y_{0}(1-D)$.
Assumption (A2) states that, conditional on $X$, the instrument $Z$ is assigned without reference to underlying potential outcomes and unobserved heterogeneity in selection response. This corresponds to the conventional assumption of instrument exogeneity. In the heterogeneous treatment effect model, however, it should be noted that instrument exogeneity takes the form of joint statistical independence of potential outcomes and selection heterogeneities, rather than the zero-correlation exogeneity restriction. Assumption (A3) states that a hypothetical change in the instrument from $z^{\prime }$ to $z$ can induce some individuals to opt in to treatment or some individuals to opt out of treatment, but never both simultaneously. Equivalently, in the terminology of imbensangrist, there is a sub-population of compliers associated with this hypothetical change, but no sub-population of defiers. Instrument monotonicity of the form (A3) allows for multiple instrumental variables. Mogstadetal2020 argue that, in the presence of multiple instruments, (A3) implies severe restrictions on individual selection responses. Our test procedure allows for multiple instruments and can be used to assess a necessary testable implication of (A3) jointly with the exclusion and random assignment restrictions. Section (ref) discusses in detail the implications of multiple instruments for instrument validity testing.
Following heckmanvytlacil01, heckmanvytlacil05, assumptions (A1) - (A3) imply a heterogeneous treatment effect model of the following form:
with
Here, $E[Y_{1}|X=x]=m_{1}(x)$ and $E[Y_{0}|X=x]=m_{0}(x)$ are regression equations of potential outcomes $(Y_{1},Y_{0})$ on the control covariates, $\tilde{U}_{1}$ and $\tilde{U}_{0}$ are unobserved heterogeneities in the potential outcomes $(Y_{1},Y_{0})$, $p(x,z) = \Pr (D=1|X=x,Z=z)$ is the propensity score, and $V$ is uniformly distributed on $[0,1]$ and is independent of $(X,Z)$. \
The MTE identifying assumptions introduced above involve counterfactual variables that are never jointly observed for the same individual, so cannot be directly verified using the distribution of observables. However, necessary testable implications have been established. These conditions relate to the distribution of observables, so may be examined empirically. They consist of single index sufficiency and nesting inequalities. We present a proof in Appendix (ref) for completeness.
Condition (i), index sufficiency, restricts the conditional distribution of $(Y,D)$ given $(X,Z)$ to depend on $Z$ through the propensity score $p(X,Z)$ only. That is, if there are multiple values of $Z$, say $z \neq z' \in \mathcal{Z}$, that yield the same value of the propensity score given $X=x$, the conditional distributions of $(Y,D)$ given $(X=x,Z=z)$ and $(X=x,Z=z')$ must be identical. Condition (ii) provides distributional monotonicity inequalities among the conditional distributions of $(Y,D)$ given $(p,x)$ with respect to $p = p(X,Z)$. This corresponds to the testable monotonicity shown by heckmanvytlacil05. The difference between the two sides of ((ref)) can be expressed as
Individuals whose selection heterogeneity $U_{D}$ falls in $(p^{\prime },p]$ can be viewed as compliers conditional on $X=x$. ((ref)) imposes non-negativity of the treated outcome probability density function for these conditional compliers. Detecting any violation of the conditions of Proposition 1 allows the joint restrictions (A1) - (A3) to be refuted. Neither of these testable implications restrict the support of the instrument $Z$, i.e., $Z$ can be discrete, continuous, or multi-dimensional.
The two testable implications shown in Proposition (ref) are distinct and assess different aspects of the distribution of observables. Holding $X$ fixed, index sufficiency compares conditional distributions of $(Y,D)$ given $(p(X,Z),X)$ at a fixed value of the propensity score, while the nesting inequalities compare conditional distributions across different values of the propensity score. Consequently, there are scenarios in which one testable implication is useful for assessing instrument validity but the other is not. For instance, if $Z$ is a scalar and $p(x,z)$ is strictly monotonic in $z$ at every $x$, instrument validity cannot be assessed through index sufficiency since their is no variation in the instrument conditional on $(p(X,Z),X)$, whereas it can be assessed using the nesting inequalities. Conversely, when the propensity score does not vary with the instrument conditional on $X$ (i.e. the instrument is irrelevant), index sufficiency does have content for assessing instrument validity while the nesting inequalities do not. As such, the two testable implications complement each other and joint assessment of them is desirable.
When the dimension of $X$ is not small, nonparametrically testing the implications listed in Proposition (ref) is challenging. Kitagawa2015 proposes testing the inequalities of Proposition (ref) (ii) with a Kolmogorov-Smirnov type test statistic where a supremum is taken over a large class of instrument functions defined on the product space of $Y$ and $X$. Similar to optimization in the empirical risk minimizing classification problem, the computational complexity of searching for this supremum increases with the dimension of $X$. Due to this computational barrier, the finite sample power of this test is unknown, and it has rarely been implemented in practice. Furthermore, accommodating a continuous instrument in this procedure remains an open problem.
We develop a test for instrument validity which circumvents the implementation issues that arise due to conditioning covariates. This test jointly assesses index sufficiency and the nesting inequalities. In addition, the test procedure can accommodate continuous instruments.
To achieve this, we impose the semiparametric restrictions proposed by CHV11, which have since been used in many empirical studies of marginal treatment effects.
In addition, strengthen the exogeneity restriction (A2) to:
(A4) restricts the conditional mean potential outcomes to depend linearly on the covariates. (A5) implies that the mean zero unobserved terms $\left( \tilde{U}_{1},\tilde{U} _{0}\right) $ are independent of the covariates. (A4) and (A5) jointly imply that covariates affect the potential outcomes only through their conditional means. Note that this assumption does not imply that the treatment effect is homogeneous, since the individual causal effect $Y_{1}-Y_{0}$ remains random even conditional on $X$. Note also that, while only conditional mean treatment effects may depend on $X$, the distribution of treatment effects remains otherwise unconstrained.
In general, under (A1) - (A3), the MTE at $X=x$ and $V=p\in \mathcal{P}$ is nonparametrically identified by
CHV11 shows that imposing the parametric functional forms of (A4) together with strengthening the exogeneity assumption of (A2) to (A5), yields the following partially linear regression equation for the observed outcome $Y$:
where $\phi \left( \cdot \right) $ is a unknown function of the propensity score, which absorbs the intercept term $p\alpha _{1}+\left( 1-p\right) \alpha _{0}$. As a result, the MTE at $X=x$ and $V=p$ is identified by
Since the only nonparametric component in this regression equation is $\phi \left( \cdot \right)$, which is a function of the scalar-valued propensity score, the computational complexity of estimation does not explode as the dimension of $X$ increases. CHV11 argues that this is a practical advantage of the functional form specification (A4) and the stronger instrument exogeneity assumption (A5).
Since the partial linear regression ((ref)) identifies $\left( \theta _{1},\theta _{0} \right)$, it is possible to compute
for every treated observation, and
for every untreated observation.
The joint restrictions (A1), (A3), (A4) and (A5) simplify not only estimation of MTE, but also the testable implications of MTE identifying assumption. To see how, consider the distribution of $U_{1}$ conditional on $\left( D=1,p\left( X,Z\right) =p,X=x\right) $. \ Under (A1), (A3) and (A4),
and similarly,
That is, for $d=0,1$, the distribution of the residual, $U_{d}=Y_{d}-X^{\prime }\theta _{d}$, conditional on $\left( D=d,p\left( X,Z\right) ,X\right)$ depends only on $p\left( X,Z\right)$. Exploiting this single index sufficiency for the distributions of $U_1$ and $U_0$, we obtain the following necessary testable implications for the joint restrictions (A1), (A3), (A4) and (A5).
We refer to the joint restrictions of (A1), (A3), (A4), and (A5) as instrument validity and base our test on the testable implications of Proposition (ref). Compared to Proposition (ref), Proposition (ref) replaces the outcome variable $Y$ with the outcome residuals $U$, and removes the requirement to condition on $X$. Index sufficiency is strengthened such that the distribution of $(U,D)$ depends on $(X,Z)$ only through the propensity score $p(X,Z)$. The nesting inequalities are strengthened to imply distributional monotonicity with respect to changes in $p(X,Z)$ regardless of the underlying variation in $(X,Z)$. These testable implications are also shown in MaoSantAnna20.
As the testable implications of Proposition (ref) follow from assumptions (A1), (A3), (A4) and (A5), any test based on them is a joint test of these assumptions. That is, in addition to the conventional exclusion, random assignment and monotonicity, we will be testing the functional form assumption of (A4) and strong exogenienty assumption of (A5). Violation of testable implications can be due to misspecification in the functional form ((ref)) or the restricted heterogeneity of MTEs with respect to the observable characteristics, i.e., additive separability of MTEs in ((ref)). In addition, when the propensity score is estimated, misspecification of the propensity score function may also lead to a rejection of the null.
As index sufficiency should hold for all $(X,Z)$ it leads to a large number of inequalities, and testing all of them may be impractical. In this paper our focus is on testing the validity of the instrument, assuming that the strong exogeneity of $X$ holds. As such, we focus on testing that the conditional joint distribution of $(U, D)$ given the propensity score does not vary with $Z$.
In this section we show how the conditional distribution of the propensity score given $Z$ affects the usefulness of each of the testable implications characterized in Proposition (ref). We use the simple case of a single binary instrument to illustrate this point.
Index sufficiency ((ref)) can be tested by examining whether the difference
is zero.
For a data generating process to have empirical content to asses index sufficiency, there must exist some $p$ where observations with both $Z=0$ and $Z=1$ are available. This occurs if we can find covariate values $x \neq x'$ such that $p(x,1) = p(x^{\prime},0)$. If no such $p$ exists, then for all $p$ either $\Pr(p(X,Z) = p, Z=0) = 0$ or $\Pr(p(X,Z) = p, Z=1) = 0$ and ((ref)) cannot be calculated. Hence, the practical relevance of ((ref)) is determined by the degree of overlap between the propensity score distributions conditional on $Z=0$ and $Z=1$.
Now consider the nesting inequalities ((ref)). These can be viewed as analogous to those tested in Kitagawa2015 with the observed outcome $Y$ replaced by the outcome residual $U$ and the instrument replaced by the propensity score $p(X,Z)$. If the propensity score were coarsened, e.g., by binning, the test procedure with no conditioning covariates and a discrete instrument of Kitagawa2015 could be applied. Although this approach would attain an asymptotically valid test size,\footnote{Assuming the estimation errors in $U$ and $p(X,Z)$ asymptotically vanish} there are circumstances where conditioning on the propensity score has a detrimental effect on power.
Maintain the simplifying assumption of a single binary instrument. Assume that the conditioning covariates $X$ are statistically independent of the unobservables $(\tilde{U}_1, \tilde{U}_0, V)$,as required by (A5). Furthermore, assume that they are also independent of $Z$. Consider the following simple violation of the exclusion restriction (A1) where the residuals of the potential outcome regression for $D = 1$ depend on $Z$
with $\nu \neq 0$. As $U_1$ is a sum of quantities that are independent of $Z$, it will be independent of $Z$.
If $\Pr(Z=1|p(X,Z)=p)$ is bounded away from 0 and 1, the additive separability of the selection equation and the independence of $X$ leads to the following decomposition:
The first equality states that, for a given $p$ where $\Pr(Z=1|p(X,Z)=p)$ is bounded away from 0 and 1, we have a mixture of observations with $Z=0$ and $Z=1$ with conditioning covariates implicitly adjusting to hold the propensity score constant. The second equality uses that the event $(U \in A, D = 1)$ is equivalent to $(U_1 \in A, V \leq p)$. To see the third, note that conditioning on $(p(X,Z) = p, Z=z)$ is equivalent to conditioning on $Z = z$ and $X \in \{x|p(X = x, Z = z) = p\}$. $X$ is independent of $V$ and $U_1$, so $X \in \{x|p(X = x, Z = z) = p\}$ can be dropped from the conditioning set.
Assume that the violation of the exclusion restriction assumption leads there to be some $A \subset \mathcal{Y}$ and $p^{\prime} < p$ where the following inequality holds:
Notice that if the instrument were valid the reverse inequality would hold.
To see how ((ref)) can drive a violation of the first nesting inequality ((ref)), apply the decomposition (ref) to the difference between the left and right side of ((ref))
The violation ((ref)) appears in the second line of ((ref)) multiplied by $(\Pr(Z=1|p) - \Pr(Z=1|p')$. The remaining terms in the third line are all negative. In order for ((ref)) to result in a violation of the nesting inequality (i.e.,for ((ref)) to be positive), it is desirable that $\Pr(Z=1|p) - \Pr(Z=1|p')$ is of large magnitude.
This magnitude depends on the strength of the relationship between the covariates $X$ and the propensity score. If variation in the propensity scores is driven mainly by $X$, $\Pr(Z=1|p)$ will change little in $p$, and the magnitude of $(\Pr(Z=1|p) - \Pr(Z=1|p')$ will be small. This will result in limited detectability of violations of instrument validity through the nesting inequality ((ref)).
In the Monte Carlo studies in Section (ref), we show numerically that tests of the nesting inequalities that directly examine the monotonicity of $\Pr(U \in A, D = 1 | p)$ in $p$ suffer from low power if variation in the propensity score is driven by the conditioning covariates $X$.
To improve power when testing nesting inequalities, we consider conditioning on $Z$ rather than conditioning on the propensity score. Consider the difference between the probabilities of $\{U \in A,D=1 \}$ conditional on $Z$,
where $f(p|Z=z)$ denotes the conditional probability density (or probability mass function) of the propensity score $p(X,Z)$ given $Z=z$.
Observing a positive sign for ((ref)) can imply three possibilities:
Hence, ((ref)) is informative for detecting violations of instrument validity only if first-order stochastic ordering holds between the distributions of $p(X,Z)|Z=0$ and $p(X,Z)|Z=1$. First-order stochastic dominance would hold, for instance, if the propensity score were given by a probit function $\Phi(X'\delta + \gamma Z)$ with $\gamma > 0$, and $X$ and $Z$ were uncorrelated. However, first-order stochastic dominance is not implied by (A1), (A3), (A4) and (A5). If $\gamma > 0$, $\delta > 0$ and $X$ and $Z$ are negatively correlated, first-order stochastic dominance can fail even if (A1) - (A5) hold. In general, a positive value of ((ref)) cannot be taken as evidence against instrument validity.
Case 3 renders a positive sign of ((ref)) inconclusive about violations of instrument validity. To rule out Case 3, first-order stochastic dominance can be checked as part of the test procedure. In cases where it does not hold, we consider trimming the sample so that the conditional distribution of the propensity score $p(X,Z)$ given $Z$ is stochastically monotonic in $Z$. We then assess the sign of an inequality analogous to ((ref)) using the potentially trimmed sample. We refer to this process as distillation.
Given an instrument, $Z \in \mathcal{Z} \subset \mathbb{R}$, we define a binary indicator $S_1 \in \{ 0, 1 \}$ such that $S_1=1$ indicates that an observation is included in the sample used for testing nesting inequalities, and $S_1=0$ indicates that the observation is trimmed and discarded. The distribution of $S_1 \in \{0, 1\}$ can depend on $(X,Z)$, but we require $S_1$ to be independent of $(U_1,U_0,V)$.
Under the assumptions of Proposition (ref), index sufficiency of the propensity score for the distribution of $(U_d,V)$ implies
where $P_{p|Z S_1}$ is the conditional distribution of the propensity score $p(X,Z)$ given $Z=z$ and $S_1=1$. The third equality follows by Assumption (A5) and $S_1$ being independent of $(U_1,U_0,V)$. Since the integrand in ((ref)) is monotonic in $p$ by Proposition (ref), $\Pr(U \in A, D=1|Z=z,S_1=1)$ is monotonically increasing in $z$ if $P_{p|ZS_1}(\cdot|z,S_1=1)$ is monotonic in $z$ in terms of first-order stochastic dominance. Similarly, the stochastic monotonicity of $P_{p|ZS_1}(\cdot|z,S_1=1)$ in $z$ implies
is monotonically decreasing in $z$.
We now formally define distillation
As discussed above, the nesting inequalities shown in Proposition (ref) remain valid even when we replace the propensity score $p(X,Z)$ with the instrument $Z$ and restrict the sample to a subsample with $S_1=1$. We hence obtain the next proposition.
Proposition (ref) shows that a testable implication exists in terms of the subdensities conditional on $Z$ provided that the sample is appropriately trimmed. However, it does not show any advantages of this approach relative to a direct test of the nesting inequalities in Proposition (ref) with a coarsened propensity score. In the proposition below, we characterise a class of processes and a restriction on trimming such that distillation will outperform a test of the nesting inequalities for observations binned by propensity score values in terms of detecting violations of instrument validity.
A proof is presented in Appendix (ref)
Regarding the construction of $S_1$, condition (ii) requires that all observations with $Z = 0$ and $p \in P^{-}$ are retained, and all observation with $Z = 1$ and $p \in P^{+}$ are retained. On the other hand, no restriction is placed on the trimming of observations with $p \in P^{+}$ and $Z = 0$ or observations with $p \in P^{-}$ and $Z = 1$. That is, first order stochastic dominance should be attained solely through trimming observations with $Z = 0$ and a high propensity score, and observations with $Z = 1$ and a low propensity score. In Section (ref) we present an algorithm that satisfies these conditions.
The inequality (ref) can always be satisfied by setting $Z = 1$ for the instrument value with a higher proportion of observations in $P^{+}$. Condition (iv) imposes restrictions on the dependence between the partial residuals and $Z$. For example, condition (iv) follows if $Z$ affects the density of partial residuals, and this relationship is independent of $X$ i.e. the partial residuals admit the representation
The first statement of Proposition (ref) says that, for any $A$ where a nesting inequality calculated using propensity score bins is positive, the nesting inequality calculated using a distilled sample will also be positive and at least as large in magnitude. That is, any violation of instrument validity that can be detected by a comparison of the data across the propensity score bins can also be detected by a comparison of the data across different instrument values, and the magnitude of the violation can be larger when comparing across instrument values. The second statement says that, for some $A$ where the nesting inequality calculated using a distilled sample is positive, the nesting inequality calculated with propensity score bins is not positive. In other words, at such $A$ the violation becomes undetectable if we compare across propensity score bins rather than instrument values. The second statement compares the nesting inequalities and distilled inequalities at common $A$, while the third statement compares their maxima over $A$. The third statement says that there exist data generating processes where instrument validity does not hold, but this violation can only be detected by comparing across instrument values. This proposition characterises a class of data generating processes where distillation is guaranteed to outperform the coarsened propensity score approach in terms of detecting invalid instruments. It thus provides theoretical support for implementing distillation.
Proposition (ref) implies that we can more easily detect violations of instrument validity by testing the nesting inequalities with a distilled sample. To verify the practical relevance of this gain, we perform a numerical exercise using the following process
$X$ is assumed to consist of $k_X \geq 3$ covariates. We assume $\nu \neq 0$, so the instrument has a direct causal effect on the outcome. In addition, we impose $\alpha > 0$ and $\delta_0^{\prime}\delta_0 = \delta_1^{\prime}\delta_1 = \delta^{\prime}\delta > 0$.
We consider the asymptotic setting, with both propensity scores and the parameters $\theta_1$ and $\theta_0$ estimated. We assume the propensity scores are estimated consistently. Given the above specification, the distribution of propensity scores conditional on $Z = 1$ will first-order stochastically dominate the distribution conditional on $Z = 0$. Conditions (i) and (ii) of Proposition (ref) can then be satisfied by including the entire sample, and condition (iii) is immediate.
Let $\hat{\theta}_1^{*} $ and $\hat{\theta}_0^{*}$ be the probability limits of the estimates of $\theta_1$ and $\theta_0$. We assume these estimates are obtained semiparametrically.\footnote{Specifically, we assume these estimates are obtained as described in Section (ref).} The partial residuals are
In Appendix (ref) we show that condition (iv) of Proposition (ref) follows immediately if these estimates are consistent i.e. $\hat{\theta}_1^{*} = \theta_1$ and $\hat{\theta}_0^{*} = \theta_0$. However, as the exclusion restriction is violated, in general $\hat{\theta}_1^{*} \neq \theta_1$ and $\; \hat{\theta}_0^{*} \neq \theta_0$. Proposition (ref) in Appendix (ref) shows that the effect of the asymptotic bias on the partial residual subdensities can be summarised by two constants which depend on the parameters $\nu, \alpha$ and $\delta^{\prime}\delta$ introduced above, and $\rho_{\delta}$
Furthermore, $\nu$ enters only as a scalar multiplier. Hence, the numerical exercise considers a class of processes indexed by values of $\delta^{\prime}\delta$ and $\rho_{\delta}$. Proposition (ref) further shows that the effect of the bias is the same for the density of $(U, D = 1)$ conditional on $Z = 1$ and the density of $(U, D = 0)$ conditional on $Z = 0$, and the density of $(U, D = 0)$ conditional on $Z = 1$ and the density of $(U, D = 1)$ conditional on $Z = 0$. As the conditional densities for $(U, D = 0)$ thus mirror those for $(U, D = 1)$, we present results for the $(U, D = 1)$ densities only.
The numerical exercise is as follows. $\mu_1$ is set to 0.3, $\mu_0$ to 0, and the elements of $\Sigma$ are drawn randomly. We set $\alpha = 0.3$, $\nu = 0.2$, and loop over values of $\rho_{\delta}$ and $\delta^{\prime}\delta$. We define $P^{-}$ to include propensity scores below the median and $P^{+}$ to include propensity scores above the median. At each set of values, we calculate the violation of the nesting inequality with observations split into bins by the median propensity score as
and the violation of the nesting inequality with a distilled sample as
where $\mathcal{C}(\mathcal{Y})$ is the set of connected intervals. We then compare these two violations. Figure (ref) plots the difference between the violation of the nesting inequality with a distilled sample and the nesting inequality with propensity score bins. For most parameter values, the violation with a distilled sample is larger, with the gain largest for high values of $\delta^{\prime}\delta$ and $\rho_{\delta}$ close to zero. Figure (ref) plots the region of the parameter space where the violation with a distilled sample is larger than the violation with propensity score bins, and the region where Statement 3 of Proposition (ref) holds. That is, the violation can only be detected by using a distilled sample. Testing with a distilled sample outperforms coarsening the propensity score for the majority of parameter values, and the region where Statement 3 holds is large.
As in Proposition 5, consider data such that $p \in P^{-} \cup P^{+}$. Let $i$ index observations. Sort observations $i = 1,...,N$ into ascending order by propensity score, so that observation $i = 1$ has the lowest propensity score and observation $i=N$ the highest. Split ties by ordering observations with $Z=0$ before observations with $Z=1$. Let $S_{1,i}$ be the sample inclusion indicator for observation $i$. First-order stochastic dominance holds when
These constraints restrict the conditional propensity score CDF for $Z=1$ to never exceed the conditional CDF for $Z=0$.
It is immediate from (ref) that
That is, the observation in the trimmed sample with the lowest propensity score must have $Z=0$ and the observation with the highest propensity score must have $Z=1$. In all the applications considered in Section (ref), no further trimming is necessary. Assume that we begin with a sample that has been trimmed in this manner, so that $Z_i=0$ for $i = 1$ and $Z_i=1$ for $i = N$. Let $n_0$ be the number of observations with $Z = 0$ and $n_1$ be the number with $Z=1$. Define $\Delta_j$ to be the difference between the conditional propensity score distributions for $Z = 1$ and $Z = 0$ at point $j$
The algorithm is as follows
The algorithm begins by including all possible observations. In step 2, it checks if first-order stochastic dominance holds with all possible observations included. If stochastic dominance would not hold, we proceed to steps 3 and 4. Step 3 considers observations with $p \in P^{-}$. In step 3a, the algorithm calculates for each index $j$ the number of observations with $i \leq j$ and $Z_i = 1$ that need to be trimmed for there to be no violation of stochastic dominance at $p(X_j, Z_j)$. Taking the maximum gives the overall number of observations that must be trimmed, $d_1$, and the index by which this trimming must take place, $j^-$. Step 3b then selects the observations to be trimmed. Looping over $j$ from 1 to $j^-$, an observation is trimmed whenever its inclusion would lead the trimmed propensity score distribution for $Z = 1$ to cross the propensity score distribution for $Z = 0$. Step 4 is analogous to step 3, but instead considers observations with $p \in P^{+}$. Step 4a calculates the number of observations with $i \geq j$ and $Z_i = 1$ that need to be trimmed for there to be no violation of stochastic dominance at $p(X_j, Z_j)$. Taking the maximum of this gives a number of observations to trim, $d_0$, and an index above which this trimming must take place $j^+$. In step 4b, an observation is trimmed whenever its inclusion would lead the trimmed complementary propensity score distribution for $Z = 0$ to cross the trimmed complementary distribution for $Z = 0$.
Figure (ref) presents an example with artifical data. Panel \subref{fig:subsample simple cdfs} plots the empirical distributions of the propensity scores conditional on $Z$. First-order stochastic dominance is violated as the distribution for $Z = 1$ lies above the distribution for $Z = 0$ for some $p$. Panel \subref{fig:subsample simple d1} plots $d_{1j}$ for each $j \in J^{-}$. In step 3a, $d_1$ and $j^{-}$ are found by taking the maximum. In this case, $d_1 = 6$. Panel \subref{fig:subsample simple P^-} shows step 3b. An observation is trimmed whenever the trimmed propensity score distribution for $Z = 1$ would rise above the distribution for $Z = 0$. Panel \subref{fig:subsample simple d0} plots $d_{0j}$ for each $j \in J^{+}$. In step 4a, $d_0$ and $j^+$ are found by taking the maximum. In this case $d_0 = 11$. Panel \subref{fig:subsample simple P^+} shows step 4b. Here an observation is trimmed whenever the trimmed complementary propensity score distribution for $Z = 0$ would rise above the trimmed complementary distribution for $Z = 1$. Panel \subref{fig:subsample simple trimmed cdfs} shows the resulting trimmed distributions, where first-order stochastic dominance holds as required by distillation.
This algorithm satisfies Definition (ref) as first order stochastic dominance is guaranteed to hold for the trimmed sample, and trimming is performed based on propensity score values only. In addition, as all observations with $Z_i = 0$ and $p \in P^{-}$ or $Z_i = 1$ and $p \in P^{+}$ are retained
which is sufficient to satisfy condition (ii) of Proposition (ref).
In general there will be many trimmed samples that satisfy Definition (ref) and condition (ii) of Proposition (ref). To select a particular sample, in this algorithm we choose the observations that are trimmed in order to make the propensity score distributions for $Z = 0$ and $Z = 1$ as similar as possible. This can be motivated heuristically as follows: variation in subdensities driven by propensity score can mask variation in the subdensities driven by $Z$. Hence, we wish to compare subdensities for similar propensity score distribution.
As the second step of the algorithm does not take into account any subsequent trimming of observations with $Z_i = 0$, it may trim more observations with $Z_i = 1$ than is necessary to attain stochastic dominance. Appendix (ref) details how the algorithm may be modified to avoid unnecessary trimming. The additional step essentially determines $d_0$ and $d_1$ by iterating using the algorithm above, so satisfies the conditions of Definition (ref) and condition (ii) of Proposition (ref) in the same manner.
We set the null hypothesis for our test to be the joint restrictions of index sufficiency ((ref)) and the nesting inequalities with a distilled sample ((ref)) shown in Proposition (ref). This section presents the construction of our test statistics and an implementation procedure with a bootstrap algorithm for computing p-values. Our exposition of the test here is restricted to a binary instrument $Z \in \{1,0\}$, while the support of covariates $\mathcal{X}$ is unconstrained. Appendix (ref) presents a test procedure for a single multi-valued discrete instrument.
Let $\lambda = \Pr(Z=1) \in (0,1)$ and $(Z,S_1) \in \{0,1 \} \times \{0,1 \}$ be a binary instrument and a sample inclusion indicator as defined in Definition (ref). With binary $Z$, we can rewrite the nesting inequalities with a distilled sample ((ref)) and index sufficiency of Proposition (ref) as the following moment inequalities and equalities: for every measurable subset $A \subset \mathcal{Y}$,
where $S_2$ is a sample inclusion indicator based on the value of the inverse probability weighting term $\Pr(Z=1| p(X_i, Z_i))$. For example, $S_2 = 1$ if $\Pr(Z=1| p(X_i, Z_i)) \in [ 0.05, 0.95]$ and $S_2 = 0$ otherwise. Trimming in this manner avoids dividing by near zero probabilities which regularizes the statistical behaviour of the sample analogues of these expectations.
The inequalities ((ref)) and ((ref)) express the nesting inequalities of Proposition (ref) in terms of the conditional expectations given $(Z,S_1=1)$. Using Bayes rule, the moment equality hypotheses ((ref)) and ((ref)) express the index sufficiency restrictions of Proposition (ref) in terms of the conditional expectation given $(Z,S_2=1)$. Trimming the sample through $S_2$ does not affect validity of the index sufficiency restrictions insofar as $S_2=1$ depends only on the value of the propensity score.
Denote the conditional distributions of $(U,D,X)$ given $Z=1$ and $Z=0$ by $P$ and $Q$, respectively, and the corresponding empirical distributions by $P_{n_1}$ and $Q_{n_0}$, where $n_1$ and $n_0$ are the number of observations with $Z=1$ and $Z=0$, respectively. Following the notational conventions of empirical process theory (e.g., vandervaart1996), we denote the expectation of a function of $(U,D,X,S_1)$, $f$, with respect to a generic distribution, $P$, by $Pf$ and the sample analogue (sample average) by $P_{n_1}f$. Similarly, for a function $g$ of $(U,D,X,S_2)$, the expectation with respect to a generic distribution $P$ and its sample analogue are denoted by $Pg$ and $P_{n_1}g$, respectively.
Let $\mathcal{C}(\mathcal{Y})$ be the class of closed and connected intervals in $\mathcal{Y}$. Define
Consider the following test statistic,
Statistic $T$ is a variance-weighted Kolmogorov-Smirnov test statistic that jointly tests ((ref)) - ((ref)) by searching for a maximal violation of their variance-weighted sample analogues. Here, the denominators of the terms appearing in the max operator are consistent estimators for the asymptotic standard deviations of the numerators. $\xi$ is a user-specified trimming constant which serves to bound the denominator away from zero. Similar to the test of Kitagawa2015, searching over the closed and connected intervals $\mathcal{C}(\mathcal{Y})$ suffices for detecting violations in any measurable set in $\mathcal{Y}$.
Keeping the standard deviation estimators fixed, we want to resample $T_1(\cdot, \cdot)$ and $T_2(\cdot, \cdot)$ from a distribution such that the null hypothesis holds and their variance-covariance structure is preserved. Specifically, we consider a multiplier bootstrap: let $(\hat{M}_1, \dots, \hat{M}_{n})$ be iid random bootstrap multipliers such that they are independent of the original sample and satisfy $E(\hat{M}_i) = 0$ and $Var(\hat{M}_i) = 1$. A bootstrap analogue of $T$ under a least favorable null can be constructed as
where, for the random variable $(a_i: i=1, \dots, n)$, we define
An implementation of the test is as follows.
The test procedure described above involves semiparametric estimation of a partially linear model where each conditioning covariate and the outcome are non-parametrically regressed on the propensity score. While non-parametric estimation can be computationally taxing, this step is only performed once and the main determinant of the computational burden is the sample size rather than the number of conditioning covariates. In cases where this non-parametric estimation step is infeasible, the test procedure can still be implemented by specifying a parametric functional form for $\phi(p)$, such as a quadratic polynomial in the propensity score.
The test proposed here does not directly test instrument validity in the context of linear 2SLS. This is because the set of identifying assumptions considered in Proposition (ref) is distinct from the set of conditions that allows us to interpret the linear two-stage least square (2SLS) estimand as a weighted average of LATEs/MTEs with positive weights. Specifically, causal interpretability of linear 2SLS requires only the weaker exogeneity condition (A2), not the strong exogeneity of (A5), and does not require the functional form specification for the potential outcome equations of (A4). On the other hand, as shown in Abadie_2003, Kolesar_2013, sloczynski2021, and Blandhol_etal_2022, causal interpretability of 2SLS relies crucially on the linearity of $E(Z|X)$ in the covariate vector $X$, whereas the identifying assumptions here considered place no restrictions on $E(Z|X)$.
However, we believe our test is useful as a specification check when one reports linear 2SLS estimates for the following reasons. First, non-rejection in our test means the data do not reject strong exogeneity (A5), implying that the data also do not contradict weak exogeneity (A2). With credible evidence for the linearity of $E(Z|X)$, non-rejection of our test can be used to support causal interpretability of linear 2SLS. Second, with specifications of the potential outcome equations and the propensity score, our test is useful for checking which variables should be included as controlling covariates. Hence, p-values of our test can also suggest which set of covariates should be included in the linear 2SLS estimation.
The test procedure can accommodate continuous instruments as follows. The initial estimation of a partially linear model can be performed using the continuous instrument. With estimated partial residuals in hand, the instrument can then be discretized, and the test procedure for a multi-valued discrete instrument of Appendix (ref) applied.
The discussion to this point has considered only a scalar $Z$, here we consider the case where $Z$ is a vector of multiple instruments. Propositions (ref) and (ref) hold irrespective of the number of instruments. Hence, both testable implications remain available. However, a test of the nesting inequalities using a coarsened propensity score would be subject to the same concerns regarding power as the single instrument case.
Consider the case of two binary instruments $Z = (Z_1, Z_2)$, $Z_1 \in \{0,1\}$, $Z_2 \in \{0,1\}$. When testing nesting inequalities and index sufficiency, $Z$ can be treated as a single instrument taking four values. That is, the estimation step is performed taking $Z$ to be vector, but when testing we map from $(0,0), (0,1), (1,0), (1,1)$ to four values of a single categorical variable and implement the test for a multi-valued instrument described in Appendix (ref). In cases where this is not viable, one approach is to aggregate the multiple instruments into a single index, and then treat this index as single multivalued instrument. Suppose the propensity score has a generalised linear form with a probit link
where $\varphi(\cdot)$ reflects the researchers assumptions about the relationship between the instruments and $Z$ and $\rho$ is a vector of coefficients. In this case, we view $\varphi(Z)$ as an aggregated index of instruments, and perform the test as if it was a single instrument.
Our testing approach can be extended to the setting of Mogstadetal2019, Mogstadetal2020 where the montonicity assumption imbensangrist is replaced by partial monotonicity. Mogstadetal2019, Mogstadetal2020 show that inference using a single instrument remains valid as long as the remaining instruments are controlled for. In the context of our test procedure, the estimation of the propensity score and residuals would be performed using the entire set of instruments, but the final step would be performed separately for each instrument. That is, index sufficiency and the nesting inequalities are tested for every instrument, with a distilled sample used in the test of nesting inequalities. A Bonferroni correction can be applied to P-values to correct for multiple tests.
We perform two sets of Monte Carlo exercises. The first evaluates the size and power of the joint test of nesting inequalities with a distilled sample and index sufficiency, and compares it to the Kitagawa2015 test which does not control for covariates and the test of the nesting inequalities using a coarsened propensity score. The second looks separately at the power of the tests of index sufficiency and nesting inequalities with a distilled sample.
All exercises use
with diagonal entries of $\Sigma$ set to 1 and off-diagonal entries to 0.3. Elements of $\theta$ are drawn from the uniform distribution with bounds $(-1,1)$. The distributions for $\gamma$ and $\delta$ are described below. $\alpha_0$, $\alpha_1$, and the specification for $Y_1$ differ for the size and power processes. When checking size we consider a valid a but irrelevant instrument
Here, conditional on the propensity score, the joint distribution of $(U, D)$ conditional on $Z$ is identical for $Z=0$ and $Z=1$. Four specifications are used to check power. They share
where $\Phi^{-1}()$ is the standard normal quantile function. The specifications for $\tilde{U}_1$ are
where $\mu_l$ takes values $(-1,-0.5,0,0.5,1)$ with probabilities $(0.15, 0.2, 0.3, 0.2, 0.15)$. Figure (ref) plots the densities of $(\tilde{U}_1,D = 1|Z)$ for each of the four processes for one draw of parameters. In the case of DGP1 the distribution for $Z=0$ is shifted, with DGP2 it has thicker tails, with DGP3 it is more concentrated, and with DGP4 there are violations at each value of $\mu_l$. In all cases, for a given propensity score, the distribution of $(\tilde{U}_1, D=1)$ conditional on $Z$ does not coincide for $Z=0$ and $Z=1$.
We consider three values of the sample size $N$: 200, 500, and 1000. For each process and sample size, we generate 1,000 samples. When comparing test procedures, the same simulated samples are used for each test procedure. The proposed test procedure requires estimating propensity scores, a partially linear model for outcomes, and $\Pr(Z=1|p(X,Z))$. Propensity scores are estimated by probit with the correct specification supplied. For the nonparametric regressions involved in the estimation of the partially linear model we use local linear regressions. For the estimation of $\Pr(Z=1|p(X,Z))$, we use local constant regressions. In both cases bandwidths are chosen by least squares cross validation. The partially linear model is estimated under the assumption that the instrument is valid, so the specification is
In the case DGP1-DGP4, this specification is incorrect, so estimates will be biased.\footnote{For the size process, this specification is correct. However, as the instrument is irrelevant, the model is not identified. In particular, it can be shown that asymptotically there is perfect multicollinearity among the columns of $X - E[X|p]$. In the finite samples used in these Monte Carlo exercises it is still possible to calculate an estimate of $\theta$. This estimate reflects finite sample noise, so the associated partial residuals can be used to check size.} We set the sample inclusion indicator for the index sufficiency test to $S_2 = \mathbbm{1} \{ \Pr(Z=1|p(X,Z)) \in [0.05, 0.95]\}$. Results are reported for two values of the trimming parameter $\xi$: $\sqrt{0.05 \times 0.95} \approx 0.21$ and 0.3. Results for additional values of $\xi$ are available upon request. P-values are calculated using 500 bootstrap samples.
Here we evaluate the size and power of the proposed test procedure, and compare it to the Kitagawa2015 test that does not control for covariates and the test of nesting inequalities with a binarised propensity score.
For this exercise, the elements of $\delta$ are drawn from the uniform distribution with bounds $(-1,1)$. For $\gamma$, we consider two cases. The first sets all elements of $\gamma$ to 0. $X$ and $Z$ are then independent, so an instrument that is valid conditional on the covariates will also be valid without conditioning on covariates and failure to control appropriately for $X$ should not affect size. In the second case each element of $\gamma$ is drawn from the uniform distribution with bounds $(-1,1)$. Here $X$ and $Z$ will not, in general, be independent, and the instrument can only be valid conditional on $X$.
Tables (ref) and (ref) present rejection rates. Turning first to the proposed test procedure, rejection rates for the size process are below nominal size irrespective of whether $X$ and $Z$ are independent. For the four power DGPs, with the exception of the smaller sample sizes for DGP4, all rejection rates are well above the nominal size. Similar rejection rates are obtained irrespective of whether $X$ and $Z$ are independent.
Next consider the test with a binarised propensity score. This test performs well in terms of size: rejection rates for the size processes are below nominal size. However, consistent with the results of Proposition (ref), there is a considerable loss in power compared to the proposed test procedure. Rejection rates above nominal size are obtained only for DGP1. For all other DGPs, rejection rates are low and change little as the sample size increases.
Figures (ref) and (ref) illustrate why the test with a binarized propensity score lacks power. To construct these figures, for each of the four power DGPs, we generate a single draw for $\theta$ and $\delta$. To simplify, $\gamma$ is set to the zero vector. We approximate the asymptotic estimates of $\theta$, $\hat{\theta}^{\star}$, by generating a sample of a million observations and regressing $Y - E[Y|p]$ on $X - E[X|p]$. Partial residuals are then calculated as $U = Y - X^{\prime}\hat{\theta}^{\star}$. Figure (ref) plots the densities of $(U, D = 1)$ conditional on $Z = 0$ and $Z = 1$, whereas Figure (ref) plots the subdensities conditional on the binarised propensity score $Z^{*}$. In Figure (ref), when we condition on $Z$, violations of the null are apparent in all cases, but in Figure (ref) violations are either small or have vanished entirely. Each value of $Z^{*}$ mixes observations with $Z=0$ and $Z=1$. The subdensities in Figure (ref) are thus a mixture of those in Figure (ref), which masks differences in shape. Furthermore, by construction, all observations with $Z^{*}=0$ have a lower propensity score than observations with $Z^{*}=1$, so the $D=1$ subdensity for $Z^{*}=0$ is guaranteed to have less total mass than the subdensity for $Z^{*}=1$.
The Kitagawa2015 test, which does not control for covariates, performs reasonably well when the instrument is independent of the covariates. The size of the test is controlled, but power is lower than the proposed test procedure. This is for two reasons. First, the Kitagawa2015 test essentially tests the nesting inequalities only, and some violations are more easily detected by index sufficiency. Second, failure to control for covariates adds noise to the outcome and which makes violations of the nesting inequalities more difficult to detect. For a single draw of $\theta$ and $\delta$, and with $\gamma$ set to zero, Figure (ref) plots the densities of $(Y, D=1)$ conditional on $Z = 0$ and $Z = 1$ for each DGP. These differ from the densities of Figure (ref) as the effect of the covariates on the outcome has not been partialed out. Compared to the densities in Figure (ref), violations are either reduced or have vanished entirely. As expected, when $Z$ and $X$ are not independent, this test procedure has rejection rates above nominal size regardless of whether the instrument is valid conditional on covariates.
Our second exercise compares the nesting inequality and index sufficiency tests. Section 2 argued that the relative power of testing nesting inequalities and index sufficiency will depend on similarity between propensity score distributions conditional on $Z$. To examine this more formally, we compare rejection rates foe the nesting inequalities and index sufficiency in two cases. Both set all elements of $\gamma$ to 0, so that the instrument and covariates are independent. In the first, the elements of $\delta$ are drawn from the uniform distribution with bounds $(-1,1)$, as in section (ref). In the second, the elements of $\delta$ are drawn from the uniform distribution with bounds $(-0.1,0.1)$. This second case limits the influence of covariates on the propensity score, which in turn limits the overlap of the conditional propensity score distributions for $Z = 0$ and $Z = 1$. Figure (ref) plots the density of propensity scores conditional on $Z = 0$ and $Z = 1$ for different values of $\delta$. As the elements of $\delta$ grow larger in magnitude, the densities overlap more. Figure (ref) plots the $\Pr(Z=1|p(X,Z))$ function associated with each value of $\delta$. For the smallest $\delta$ this function is close to a step function but, as the elements of $\delta$ grow in magnitude, it becomes flatter.
Table (ref) presents rejection rates. When the elements of $\delta$ are drawn from the uniform distribution with bounds $(-1,1)$, index sufficiency has higher rejection rates. When the elements of $\delta$ are drawn from the uniform distribution with bounds $(-0.1,0.1)$, index sufficiency loses power in all cases. Nesting inequalities now have considerably higher power for DGP1, slightly higher power for DGP2 and DGP3, while index sufficiency still exhibits more power for DGP4. In all cases, the overall rejection rate is close to the maximum of the rejection rates for the nesting inequalities and index sufficiency. This suggests that testing both testable implications jointly dominates testing only one of them.
In this section we apply the proposed test procedure to the instrumental variables of Card1993, angristevans1998, and oreopoulos2006. In all cases we maintain the outcome, treatment, instrument and conditioning covariates of the original design, and estimate a potential outcome model of the form described in Section 2. For the propensity score, we estimate a probit with the instrument, conditioning covariates and interactions between the two as controls. For the outcome, we allow coefficients on the covariates to depend on the treatment. In the case of Card1993 and angristevans1998, the partially linear model is estimated as described in step 2 of Section (ref), with local linear regressions used to estimate $E[X|p]$ and $E[Y|p]$. $\Pr(Z = 1|p)$ is estimated by local constant regressions. Bandwidths are selected by least squares cross validation. For oreopoulos2006, the partially linear model is estimated by assuming that the nonparametric component $\phi(p)$ is a cubic polynomial in $p$, and $\Pr(Z = 1|p)$ is estimated by probit.
Table (ref) presents p-values. A p-value below 0.05 corresponds to a rejection of the null hypothesis at the 5% level of significance. p-values are reported for a range of values of the trimming parameter $\xi$. For each application, Table (ref) shows the number of observations in the data set, and the number of observations included in the trimmed samples used in the tests of nesting inequalities with a distilled sample and index sufficiency.
Card1993 estimates returns to education. The outcome $Y$ is the logarithm of weekly earnings, $D$ is years of education, and $Z$ is growing up in a local labour market with an accredited four year college. The headline specification uses four ordered discrete, three categorical and six binary conditioning covariates. These include potential experience, parental education, and residence in a standard metropolitan statistical area (SMSA). The sample consists of 3,010 observations.
The validity of this instrument has previously been tested by hubermellace15Testing, Kitagawa2015, mourifiewan and sun2020. However, for computational reasons, at most only a subset of the conditioning covariates have been used. sun2020 does not control for covariates, hubermellace15Testing conditions on two covariates, mourifiewan three, and Kitagawa2015 five. In addition, all of the conditioning covariates that have been used are binary. The proposed test procedure allows us to include all conditioning covariates.
As the treatment is not a binary variable, we binarize it. The binarized treatment is defined to be 1 if years of education is at least 16. angristimbens1995 and marshall_2016 show that coarsening the treatment variable in this way can lead the instrument to violate the exclusion restriction through coarsening bias.\footnote{Section 5 of sun2020 discusses coarsening bias in the context of testing for instrument validity.} This occurs when the instrument induces changes in the underlying continuous treatment which are not captured by the discretised treatment. We verify that, conditional on the binarised treatment, the distribution of the continuous treatment is similar for both values of the instrument. Therefore, the binarized treatment does appear to capture changes in the underlying continuous treatment.
Panel \subref{fig:Card outcome} of Figure (ref) plots joint $(Y,D)$ subdensities for $Z=0$ and $Z=1$. Kitagawa2015 finds that, without conditioning on covariates, the null of instrument validity can be rejected. This violation can be seen in the $Y(0)$ subdensities. Figure (ref) plots the conditional distributions and smoothed conditional densities for the estimated propensity score. The distribution for $Z=1$ lies below the distribution for $Z=0$, which indicates that first-order stochastic dominance holds, and there is considerable overlap between the densities. As show in Table (ref), no observations are removed for either the test of nesting inequalities with a distilled sample or the test of index sufficiency. Panel \subref{fig:Card nesting} of Figure (ref) plots the subdensities of estimated partial residuals used in the test of nesting inequalities. After controlling for covariates, the violation in the $D=0$ subdensities is no longer evident. Panel \subref{fig:Card index} plots the reweighted densities used to test index sufficiency. These are similar to those in panel \subref{fig:Card nesting}. p-values are presented in the top panel of (ref). p-values are lower for the index sufficiency test than the test of nesting inequalities with a distilled sample, but the null is not rejected at any conventional level of significance.
angristevans1998 estimate the effect of family size on parental labour supply using the same-sex instrument, which takes a value of one when the two eldest children are the same sex and zero otherwise. Their justification for this instrument is that the sex of children is randomly determined and that there is a parental preference for mixed sex children, so parents whose first two children are the same sex are more likely to have a third. The validity of this instrument was subsequently challenged in RosenzweigWolpin2000. Using a simple model of labour supply choice, they show that additional restrictions on preferences are needed for this instrument to satisfy the exclusion restriction.
In response to this debate, the validity of the instrument has been tested by huber2015 and mourifiewan. In both cases, the angristevans1998 sample is split into subgroups based on values of a subset of the covariates, and then instrument validity is tested for each subgroup. huber2015 defines 66 subgroups, and finds that the hubermellace15Testing test rejects the null in a small number of them, but the Kitagawa2015 test never rejects the null. mourifiewan reports no rejections at the 10% level of significance in 24 subgroups. Neither huber2015 or mourifiewan uses the full set of covariates from the original specification of angristevans1998, andsome of the variables that are used are coarsened.
We use the replication files for angristevans1998 to construct the sample and variables. We focus on the 1980 Census sample of married mothers. This consists of married mothers aged between 21 and 35 who have at least two children and were at least 15 years old at the time of their first birth. The sample contains 254,652 observations. To keep estimation of the semiparametric model feasible, we use a randomly chosen 1 / 20 subsample of this data, which results in 12,732 observations. $Z$ is the same-sex instrument. $D$ is an indicator set to 1 for mothers with more than two children. angristevans1998 presents results for several measures of labour supply. Here we set $Y$ to be the logarithm of annual family labour income. The headline specification controls for demographic characteristics of the mother, the sex of the first child, and the sex of the second child.
Panel \subref{fig:Angrist Evans outcome} of Figure (ref) plots subdensities for the outcomes without controlling for covariates. In this case there is no visual evidence of violation of the instrument validity. angristevans1998 shows that the same-sex instrument almost exactly balances covariate values, which suggests that the instrument and covariates are independent. As discussed in Section 4, when the instrument and covariates are independent, failure to control for covariates can mask violations of instrument validity.
Figure (ref) plots the conditional distributions and smoothed densities for the estimated propensity scores. There is no obvious violation of first-order stochastic dominance, and the propensity score densities overlap. One observation is removed when constructing a distilled sample for the test of nesting inequalities. However, as seen in Table (ref), only a small fraction of the observations are used in the test of index sufficiency. Since angristevans1998 uses only a small number of discrete covariates, there are few estimated propensity score values where there is a mixture of observations with $Z=0$ and $Z=1$. Hence, the estimate of $\Pr(Z=1|p(X,Z))$ is either 0 or 1 for the majority of the sample, and these observations are removed for the test of index sufficiency. Panels \subref{fig:Angrist Evans nesting} and \subref{fig:Angrist Evans index} of Figure (ref) plot smoothed subdensities for the samples used in the nesting inequalities with a distilled sample and index sufficiency tests. There is no visual evidence of a violation of the nesting inequalities. For index sufficiency, both the $U(0)$ and $U(1)$ subdensities for $Z=0$ and $Z=1$ appear to differ. However, the results in Table (ref) show that this difference is not statistically significant.
Changes in schooling laws have been used as an instrument to estimate the effects of education for a variety of outcomes. acemogluangrist2000 uses this instrument to estimate human capital externalities, lochnermoretti2004 estimates effects on crime, LlerasMuney2005 effects on health, and oreopoulos2006 the returns to education. These designs exploit variation in the timing of changes in compulsory schooling laws across locations. Conditional on covariates, such as location and time fixed effects, this variation is assumed to be unrelated to other changes. stephensyang2014 presents evidence that this assumption does not hold in data for US states and, in particular, that changes in compulsory schooling laws are correlated with changes in school quality. brunello2013testing develops a test of uncorrelatedness between changes in compulsory schooling laws and school quality. They apply this test to European data and find that they cannot reject the null of uncorrelatedness. BolzernHuber2017 tests the validity of the compulsory schooling law instrument using the hubermellace15Testing and Kitagawa2015 tests with data from seven European countries, and does not reject the null of instrument validity.
We apply our test to the data of oreopoulos2006. The instrument is the legal minimum school leaving age for the UK. This was increased from 14 to 15 in 1947 in Great Britain and in 1957 in Northern Ireland. The dataset consists of individuals aged between 16 and 65. It is constructed by combining UK and Northern Ireland household surveys for the years 1984 through to 2006.\footnote{The published paper oreopoulos2006 used data from the UK General Household Survey for 1983-1998, and data from the Northern Ireland Continuous Household Survey for 1985-1998. A corrigendum with subsequent editions of each survey incorporated into the data was later published online. We use the revised data.} $Y$ is log real earnings, $D$ is age of leaving full-time education, and $Z$ is a dummy for the legal minimum leaving age faced at age 14 being 15. Conditioning covariates are year aged 14, current age, sex, survey year and residence in Northern Ireland. As with Card1993, the treatment variable must be binarized. Here we define the treatment to be one if the age of leaving education is at least 15. We investigate the possibility that this induces coarsening bias, but the binarised treatment seems to capture well the variation in the underlying treatment induced by the instrument.
The data consists of 82,908 observations which are aggregated into 30,587 cells. Each cell consists of individuals with common values of the treatment, instrument and conditioning covariates. The outcome variable is averaged at the cell level, and each cell is assigned a weight equal to the number of individuals it contains. The identifying assumptions stated in Section 2.2 were not for data that has been aggregated in this way. However, the aggregation procedure preserves all variables apart from the outcome, which implies that the propensity score is also preserved. The identifying assumptions may then be simply restated in terms of the averaged outcome, and the same testable implications derived.\footnote{Note, however, that failure to reject the LATE assumptions for the aggregated data does not imply that they would not be rejected for the unaggregated data and vice versa.}
Panel \subref{fig:Oreopoulos outcome} of Figure (ref) plots subdensities for the outcome variable. It is immediately noticeable that there are very few observations with $Z=1$ and $D=0$. Figure (ref) plots the conditional distributions and densities for the propensity score. First-order stochastic dominance clearly holds, but the distributions barely overlap\footnote{The minimum estimated propensity score for a cell with Z = 1 is 0.678, and the estimated maximum propensity score with Z = 0 is 0.683. The overlap contains 8 cells (21 observations) with Z = 0 and 12 cells (94 observations) with Z = 1.}. The estimated $\Pr(Z=1|p(X,Z))$ is either 1 or 0 for all cells. The entire sample is used in the test of nesting inequalities with a distilled sample, but there are no observations for which to perform the test of index sufficiency. Table (ref) therefore present p-values for the test of nesting inequalities with a distilled sample only. We find that the null is rejected at even the 1% level of significance. The violation is in the lower tail of the $U(1)$ subdensities, which are plotted in panel \subref{fig:Oreopoulos nesting} of Figure (ref). The subdensity of $U(1)$ for $Z=0$ does not lie entirely within the support of the subdensity for $Z=1$.
This paper develops a novel test for instrument validity in the LATE/MTE framework that can accommodate conditioning covariates. We follow CHV11 in assuming a linear relationship between potential outcomes and conditioning covariates and that unobserved heterogeneities are independent of the instrument and covariates.. Expected outcomes then follow a partially linear model, depending linearly on conditioning covariates and non-parametrically on the propensity score. With this functional form, partial residuals can be obtained by partialing out the conditioning covariates. We derive testable implication for the subdensities of these partial residuals: index sufficiency and the nesting inequalities. We propose a test procedure consisting of an initial step where the partially linear model is estimated and partial residuals obtained, followed by a joint test of index sufficiency and the nesting inequalities. Crucially, as conditioning covariates are partialed out, they increase only the complexity of estimating the propensity score and the partially linear model. In this way, the test procedure overcomes the computational hurdles of implementing tests such as Kitagawa2015, hubermellace15Testing and mourifiewan in the presence of conditioning covariates.
The testable implications are complementary. Index sufficiency holds the propensity score fixed and compares subdensities for different values of the instrument, whereas the nesting inequalities are distributional monotonicity relations ordered with respect to the propensity score. A straightforward approach to testing the nesting inequalities is to coarsen the propensity score and check the inequalities for subdensities indexed by this coarsened variable. However, in cases where the propensity score is strongly correlated with the conditioning covaraites, the distribution of instrument values can become homogenised across propensity score bins, leading to a lack of power. We showed that the nesting inequalities hold with the original instrument in place of the propensity score if the conditional distribution of the propensity score given the instrument is stochastically monotonic in the sense of first-order stochastic dominance. Since first-order stochastic dominance is not implied by the identifying assumptions, we proposed testing the nesting inequalities using a trimmed sample where first-order stochastic dominance does hold. We refer to this process as distillation, and present an algorithm for constructing distilled samples. We show analytically a class of data generating processeses where distillation is guaranteed to outperform a test of the nesting inequalities over propensity score bins in terms of detecting violations of instrument validity. Monte Carlo exercises confirm this finding, with consistent gains in power when using the proposed test procedure in place of testing the nesting inequalities with a coarsened propensity score.
We believe that there are two areas which future research could address. First, formal characterisation of the size and power of this test, incorporating uncertainty over estimation of the partial residuals and the distillation process, is left for future research. Second, as noted in Section 2, in general there will be many possible distilled samples that satisfy the first-order stochastic dominance condition. In Proposition (ref), we showed a restriction on trimming that is guaranteed to improve detection of violations of instrument validity for a class of data generating processes, and we provided an algorithm that satisfies this restriction. However, other possibilities exist, and the optimal choice of distillation process is an open question.