Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
203,569 characters · 33 sections · 106 citation commands
Salvaging Falsified Instrumental Variable Models
JEL classification: C14; C18; C21; C26; C51
Keywords: Instrumental Variables, Nonparametric Identification, Partial Identification, Sensitivity Analysis
\onehalfspacing
Many models used in empirical research are falsifiable, in the sense that there exists a population distribution of the observable data which is inconsistent with the model. We equivalently call these overidentified or refutable models.\footnote{Keep in mind that falsifiability is a property of a model, not a parameter. It is nonetheless common for researchers to discuss “overidentified parameters.” These are parameters whose identified sets are either empty (when the model is falsified) or a singleton (when the model is not falsified). Thus such parameters are point identified when the model is not falsified. To avoid confusion between properties of a parameter and properties of a model, we only use the terms falsifiable and refutable when describing models.} With finite samples, researchers often use specification tests to check whether their baseline model is refuted. A well known example is the overidentifying restrictions test in linear instrumental variable models (AndersonRubin1949, Sargan1958, Hansen1982). Abstracting from sampling uncertainty\footnote{We study population level identification and falsification in this paper. We briefly discuss finite sample estimation and inference in sections (ref) and (ref), but this is not the focus of the paper.}, the population versions of such specification tests have a persistent problem: What should researchers do when their baseline model is refuted?
One option is to ignore the finding and report some point estimand, like the 2SLS estimand in an overidentified linear instrumental variable model. There are two problems with this approach. First, it is often justified by appealing to large sample sizes. For example, Nevo2001 refers to this justification on page 325: “It is well known that with a large enough sample a chi-squared test will reject essentially any model.” If we are willing to relax assumptions, however, this common justification is not true: There always exist assumptions which are sufficiently weak that they are not refuted. Second, such point estimands can be highly misleading. For example, HeckmanHotz1989 compare estimates from observational data with those from experimental data. In their data, they show that the observational estimates from models which fail specification tests are much farther from the experimental estimates than observational estimates from models which pass specification tests. Essentially, parameters like the 2SLS estimand are motivated by arguments that they deliver a causal effect of interest under a set of baseline assumptions. When those assumptions are known to be false, there is no reason for the corresponding estimand to be close to the causal effect of interest. In appendix (ref) we discuss five other responses to baseline falsification from the literature. We review the related literature in appendix (ref).
Instead of ignoring findings from specification tests, we provide four constructive ways for researchers to salvage a falsified baseline model. First, researchers can measure the extent of falsification. To do this, we consider continuous relaxations of the baseline assumptions of concern. We then define the falsification frontier: The smallest relaxations of the baseline model which are not refuted. This frontier provides a quantitative measure of the extent of falsification. Second, researchers can present the identified set for the parameter of interest under the assumption that the true model lies somewhere on this frontier. We call this the falsification adaptive set (FAS). This set collapses to the baseline identified set or point estimand when the baseline model is not refuted. When the baseline model is refuted, this set expands to include all parameter values consistent with the data and a model which is relaxed just enough to make it non-refuted. This set generalizes the standard baseline estimand to account for possible falsification. Importantly, researchers do not need to select or calibrate sensitivity parameters to compute the falsification adaptive set. While this set is agnostic about the direction of relaxation, our third suggestion is that researchers present the identified set for a specific minimally relaxed model. Finally, as a further sensitivity analysis, researchers can present identified sets for points beyond the falsification frontier. We discuss a complementary Bayesian approach in appendix (ref).
To illustrate these four constructive ways to salvage a falsified baseline model, we study a classic source of overidentification: observation of several instrumental variables. We do this in two different models. The first is the classical constant coefficients linear model with multiple instruments. This model imposes homogeneous treatment effects, but allows for continuous treatments. We generalize the usual overidentifying restrictions by allowing the instruments to have some direct effect on outcomes. We use this result to characterize the falsification frontier, which trades off bounds on the magnitudes of these direct instrument effects. We then characterize the identified set along the falsification frontier. This leads to a particularly simple closed form expression for the falsification adaptive set, depending only on the value of a handful of 2SLS regression coefficients. We also show how to use these identified sets to do sensitivity analysis for non-falsified models beyond the frontier.
It is well known that the classical overidentifying conditions may not hold when treatment effects are heterogeneous, even if all instrument exogeneity and exclusion restrictions hold. We therefore also study a second model, which allows for heterogeneous treatment effects and multiple instruments. In this model we focus on binary treatments. We consider both binary and continuous outcomes. Although this model allows for heterogeneous treatment effects, it is nonetheless falsifiable. We relax statistical independence between each instrument and potential outcomes using a latent propensity score distance from our previous work, MastenPoirier2017. Under these relaxations, we derive the identified set for the marginal distributions of potential outcomes. We show how to quickly compute the falsification frontier and the falsification adaptive set using convex optimization. We then discuss how to use these identified sets to do sensitivity analysis regardless of whether the baseline model is refuted.
We show how to use our results in four previously published empirical studies: \citet*[The Review of Economic Studies]{DurantonMorrowTurner2014}, \citet*[The Quarterly Journal of Economics]{AlesinaGiulianoNunn2013}, \citet*[The American Economic Review]{AcemogluJohnsonRobinson2001}, and Nevo2001. Each paper reports 2SLS estimates using multiple instrumental variables and each paper discusses concerns about instrument validity. All four papers run overidentification tests, which sometimes fail. Even when they do not fail, the authors sometimes express a concern that this could simply be due to low sample size. We show that the falsification adaptive set is an informative complement to these traditional tests: Rather than focusing on null hypothesis significance testing, the FAS summarizes the range of estimates obtained from alternative models which which are not falsified by the data. Thus the FAS reflects the model uncertainty that arises from a falsified baseline model: Relying on different instruments to different degrees yields different results. The FAS gives this range of results.
Finally, it is important to distinguish between two kinds of falsifiable models. The first kind is falsified because of an assumption made directly on observed random variables. For example, suppose we observe a scalar random variable. Consider the model which assumes that this variable is normally distributed. At the population level, we can simply check whether the observed variable is actually normally distributed. If not, the model is refuted. When the model is refuted, it can be salvaged by removing the normality assumption. A massive literature in statistics and econometrics on semi- and non-parametrics is largely concerned with salvaging refuted models by relaxing parametric assumptions on observed random variables. For example, see the survey by Spanos2018. The second kind of model is falsified because of an assumption made on unobserved random variables, or unobserved structural parameters. Salvaging these models is delicate because, even at the population level, the data themselves do not tell us the correct alternative assumptions. We study this second kind of falsifiable model in this paper.
In this section we consider a general falsifiable model. We use this model to precisely define our four recommended responses to baseline falsification. In particular, we formally define the falsification frontier and falsification adaptive set. In sections (ref) and (ref) we illustrate these general concepts in two specific instrumental variable models. We discuss some interpretation issues in section (ref). In section (ref) we discuss the relationship between the falsification frontier and the breakdown frontier we studied in our previous work (MastenPoirier2017BF). We discuss estimation and inference in section (ref).
Let $W$ be a vector of observed random variables. Let $\mathcal{F}$ denote the set of all cdfs on the support of $W$, $\operatorname*{supp}(W)$. A model is a set of underlying structural parameters which generate the observed distribution $F_W$ and restrictions on those structural parameters. This definition of a model suffices for our purposes. See section 2 of Matzkin2007, for example, for a more formal definition. A given model $\mathcal{M}$ is falsifiable if there are some observed distributions $F_W$ which could not have been generated by the model. If such a cdf is observed, we say the model is falsified (equivalently, refuted). For a given model, let $\mathcal{F}_\text{f}$ denote the set of cdfs $F_W$ which falsify the model. Let $\mathcal{F}_\text{nf}$ denote the set of cdfs $F_W$ which do not falsify the model. Falsifiable models are often said to have testable implications. We prefer the terms `falsifiable' or `refutable', to distinguish this population level feature from issues arising in statistical hypothesis testing.\footnote{Informally, it is possible that for every non-falsified distribution of observables there exists a falsified distribution of observables which is arbitrarily close. Thus hypothesis tests in finite samples have a difficult time distinguishing the two cases, even though we can distinguish them at the population level. In this case we say the model is falsifiable but not testable. CanaySantosShaikh2013 study an example of this; also see Freyberger2017.}
Suppose we begin with a falsifiable baseline model $\mathcal{M}(0_L)$. Let $\mathcal{F}_\text{nf}(0_L)$ denote the set of joint distributions of the data which are not falsified by this model. Hence $\mathcal{F}_\text{nf}(0_L)$ is a strict subset of $\mathcal{F}$. $L$ denotes the number of assumptions which we believe may be the reason the model is falsifiable. For each assumption $\ell \in \{1,\ldots,L\}$, we define a class of assumptions indexed by a parameter $\delta_\ell$ such that the assumption is imposed for $\delta_\ell = 0$, the assumption is not imposed for $\delta_\ell = 1$, and the assumption is partially imposed for $\delta_\ell \in (0,1)$.\footnote{More generally, the range of $\delta$ could be $[0,\delta_\text{max}]$ for some value $\delta_\text{max} \geq 0$, possibly $+\infty$.} These assumptions must be nested in the sense that for $\delta_\ell' \geq \delta_\ell$, assumption $\delta_\ell'$ is weaker than assumption $\delta_\ell$. Let $\mathcal{M}(\delta)$ denote the model which imposes assumptions $\delta = (\delta_1,\ldots,\delta_L)$. Let $\mathcal{F}_\text{nf}(\delta)$ denote the set of joint distributions of the data which are not falsified by this model. Suppose further that the model $\mathcal{M}(1_L)$ which does not impose any of the $L$ assumptions is not falsifiable.
Recall that $F_W$ denotes the observed distribution of the data. Suppose $F_W \notin \mathcal{F}_\text{nf}(0_L)$, so that the baseline model is falsified. We first consider the case where we only relax a single assumption.
That is, the falsification point is the smallest relaxation of the baseline assumption such that the model is not falsified by the observed data $F_W$. For any $\delta < \delta^*$, the model is falsified. For any $\delta > \delta^*$, the model is not falsified. Under mild conditions, we can show that the set $ \{ \delta \in [0,1] : F_W \in \mathcal{F}_\text{nf}(\delta) \}$ is closed, and hence the model is not falsified at $\delta^*$.
The falsification frontier is the multi-dimensional version of the falsification point. To formally define it, partition $[0,1]^L$ into the set of all assumptions which are falsified, \[ \mathcal{D}_\text{f} = \{ \delta \in [0,1]^L : F_W \notin \mathcal{F}_\text{nf}(\delta) \} \] and the set of all assumptions which are not falsified, \[ \mathcal{D}_\text{nf} = \{ \delta \in [0,1]^L : F_W \in \mathcal{F}_\text{nf}(\delta) \}. \]
That is, the falsification frontier is the set of assumptions which are not falsified, but if strengthened in any component, leads to a refuted model. In the one dimensional case, this corresponds to the smallest relaxation $\delta$ such that the model is not refuted. For example, suppose $L=2$. For each $\delta_1 \in [0,1]$ define \[ \delta_2^*(\delta_1) = \inf \{ \delta_2 \in [0,1] : F_W \in \mathcal{F}_\text{nf}(\delta_1,\delta_2) \}. \] This function $\delta_2^*(\cdot)$ plots the falsification frontier in the two-dimensional case. Figure (ref) shows an example of such a function. Our baseline model is refuted and we have picked two different assumptions to focus on as possible explanations. The horizontal axis measures the relaxation of the first assumption. The vertical axis measures the relaxation of the second assumption. The origin at the lower left represents the baseline model. The point at the top right of the box represents the model where neither of the two baseline assumptions are imposed. First, look at the top left and lower right corners. We see that if we completely drop one assumption while maintaining the other, the model is no longer falsified. However, we can learn more by studying the points at which the model becomes non-refutable---that is, by examining the falsification frontier. This is shown as the boundary between the two regions. Consider its vertical and horizontal intercepts. Supposing the units are comparable, then the model continues to be refuted even for large relaxations of assumption 1, while maintaining assumption 2. Conversely, the model is not refuted if we allow moderate relaxations of assumption 2, while maintaining assumption 1. Looking along the frontier, strengthening assumption 2 by some amount requires relaxing assumption 1 by even more, if we want to avoid falsification. Thus the falsification frontier gives the trade-off between different assumptions' ability to falsify the model.
In this section we discuss two well-known points that are nevertheless important to keep in mind when discussing falsification. First, although a falsified model cannot be true, a non-falsified model is not necessarily true. This asymmetry underlies Popper's Popper1934 philosophy of falsification. In the classical linear instrumental variables model of section (ref), this is discussed in KadaneAnderson1977, Newey1985, Breusch1986, Small2007, ParenteSantosSilva2012, and Guggenberger2012. They noted that the classical overidentifying conditions are really constraints that the instruments are consistent with each other. Thus it is possible that the instrument exogeneity assumptions fail in such a way that all instruments yield the same biased estimand for the true parameter. In this case, the baseline model is not refuted even though it is false. This is often explained by saying that the classical overidentification test is “not consistent against all alternatives.” This is why the falsification frontier is not a partition of true and false assumptions, but rather of falsified and non-falsified assumptions. We discuss this point further in appendix (ref).
Second, when the baseline model is falsified, it could be due to any combination of our baseline assumptions. Importantly, this is true regardless of the shape of the falsification frontier. The falsification frontier does not tell us why the baseline assumptions were refuted. For example, consider figure (ref). This figure does not tell us that `assumption 1 is probably the reason the baseline model was refuted'. Instead, it tells us about the robustness of falsification to certain relaxations from the baseline assumptions. It tells us that, maintaining assumption 2, we must substantially relax assumption 1 in order to prevent falsification.
In section (ref), we formalized our first recommendation: Measure the extent of falsification. When the falsification frontier is far from the origin in all directions, researchers may wish to stop there. In other cases, researchers may want to proceed and present identified sets for non-falsified relaxations of the baseline model.
Let $\Theta_I(\delta)$ denote the identified set for a parameter of interest $\theta \in \Theta$, given the model which imposes the assumptions $\delta$. When $\delta \in \mathcal{D}_\text{f}$, $\delta$ is below the falsification frontier. In this case, the identified set $\Theta_I(\delta)$ is empty. When $\delta \in \mathcal{D}_\text{nf}$, $\delta$ is on or above the falsification frontier. In this case, the identified set $\Theta_I(\delta)$ is nonempty.
The set of $\delta$'s on the falsification frontier are minimally non-falsified, in the sense that they are the closest to the baseline model while still not leading to a falsified model. These $\delta$'s, along with the $\delta$'s beyond the frontier, are consistent with the data. As discussed in section (ref), the data alone cannot tell us which of these models are true. It only tells us that $\delta$'s below the falsification frontier are not true. Indeed, the same remark is true when the baseline model is not falsified. Yet in that case, it is common to present identified sets or point estimands under the baseline model. To generalize this practice to the case where the baseline model is falsified, consider the following definition.
The falsification adaptive set is the identified set for the parameter of interest under the assumption that the true model lies somewhere on the frontier. When the baseline model is not refuted, this set collapses to the $\Theta_I(0_L)$, the baseline identified set (which may be a singleton). This is what researchers typically report when their baseline model is not refuted. When the baseline model is refuted, however, the falsification adaptive set expands to account for uncertainty about which assumption along the frontier is true. Hence this set generalizes the standard baseline estimand to account for possible falsification. We recommend that researchers report this set.
In some cases, researchers may have a prior belief over the relative role of the various assumptions in falsification. In this case, researchers may also want to present $\Theta_I(\delta)$ for a specific $\delta$ on the falsification frontier. For example, consider figure (ref). Two natural relaxations of the baseline model are the horizontal intercept $(\delta_1^*,0)$ and the vertical intercept $(0,\delta_2^*)$. These correspond to fully maintaining one assumption while sufficiently relaxing the other. The corresponding identified sets are $\Theta_I(\delta_1^*,0)$ and $\Theta_I(0,\delta_2^*)$. Presenting sets like these for specific points on the frontier is our third recommendation.
Although the points on the falsification frontier are consistent with the data, this does not mean one of them is true. For this reason, the falsification adaptive set---although it represents uncertainty due to the nature of the misspecification---is nonetheless an optimistic set. It assumes that, although the model is misspecified, it is only minimally misspecified. Since this may not be true, our fourth and final recommendation is that researchers present identified sets $\Theta_I(\delta)$ for points $\delta$ beyond the falsification frontier, as a sensitivity analysis. We discuss this recommendation in the next subsection.
The widespread practice in existing empirical work is to first present results from a usually optimistic baseline model---which typically point identifies a parameter of interest---and then to present results from various sensitivity analyses. Our recommendations directly generalize this standard practice to falsified baseline models: First present the falsification adaptive set and then present identified sets for further relaxations.
Our fourth recommendation is to present identified sets $\Theta_I(\delta)$ for $\delta$'s beyond the falsification frontier. To guide this analysis, researchers can use the breakdown frontier concept which we previously studied in MastenPoirier2017BF. In this subsection we briefly summarize this concept and relate it to the falsification frontier. Consider the model $\mathcal{M}(\delta)$ from section (ref). As above, let $\Theta_I(\delta) \subseteq \Theta$ be the identified set for some parameter of interest $\theta \in \Theta$. Let $\mathcal{C} \subseteq \Theta$. Suppose we are interested in the conclusion that $\theta \in \mathcal{C}$. For example, if our parameter is the average treatment effect, then we may be interested in the conclusion that the average treatment effect is positive, $\text{ATE} \geq 0$. Then $\mathcal{C} = [0,\infty)$. Define the robust region \[ \text{RR} = \{ \delta \in [0,1]^L : \Theta_I(\delta) \subseteq \mathcal{C}, \Theta_I(\delta) \neq \emptyset \}. \] The conclusion of interest holds for all assumptions in this set. That is, for all assumptions such that all elements of the identified set are also elements of $\mathcal{C}$. We also restrict the set to assumptions such that the model is not refuted. The breakdown frontier is the boundary between the robust region and its complement: \[ \text{BF} = \overline{\text{RR}} \cap \overline{\text{RR}^c}. \] In the two dimensional case, we can write the breakdown frontier as the function \[ \text{BF}(\delta_1) = \sup \{ \delta_2 \in [0,1] : \Theta_I(\delta_1,\delta_2) \subseteq \mathcal{C}, \Theta_I(\delta_1,\delta_2) \neq \emptyset \}. \]
It is possible for the robust region to be empty. This happens when the baseline model is falsified, but as soon as the model is not falsified, the identified set $\Theta_I(\delta)$ contains values of $\theta$ outside of $\mathcal{C}$. When the robust region is not empty, so that the breakdown frontier is not empty, the breakdown frontier will always be weakly larger than the falsification frontier. This follows since the identified set is empty for all points below the falsification frontier.
Figure (ref) depicts an analysis combining both the falsification frontier and the breakdown frontier. Here the robust region is a band between the falsification frontier and the breakdown frontier. It is the region of assumptions which are weak enough that the model is not falsified, but strong enough that our conclusion of interest holds.
A key difference between the breakdown and falsification frontiers is that the falsification frontier does not depend on (a) the parameter of interest or (b) the conclusion of interest. For a given class of relaxations from the baseline assumptions, there is only one falsification frontier. In contrast, there are many different breakdown frontiers, depending on the parameter and conclusion of interest. This difference is important since there have been extensive debates in the literature regarding what parameters researchers should study (for example, see Imbens2010 and HeckmanUrzua2010). Thus one's opinions on this issue do not affect the problem of measuring the extent of falsification.
In this paper we focus on population level analysis. In this section we briefly discuss how to implement our recommendations with finite sample data. First, to measure the extent of falsification, one can estimate and do inference on the falsification frontier. The basic idea is similar to our analysis of inference on breakdown frontiers (MastenPoirier2017BF). For the breakdown frontier, we recommended constructing lower confidence bands. For the falsification frontier, we recommend constructing upper confidence bands, since this provides an outer confidence set for the area under the falsification frontier---the set of models which are refuted. When doing a combined analysis, we recommend constructing an inner confidence set for the robust region; that is, for the area above the falsification frontier but below the breakdown frontier. This will require constructing two-sided confidence bands, whereas one only needs to construct one-sided confidence bands when considering each frontier separately.
When there are closed form expressions for the identified sets $\Theta_I(\delta)$---like in our analysis below---the falsification adaptive set can be estimated by sample analog: \[ \widehat{\text{FAS}} = \bigcup_{\delta \in \widehat{\text{FF}}} \widehat{\Theta}_I(\delta). \] In the linear instrumental variable model of section (ref), we provide a particularly simple characterization of the falsification adaptive set which does not require pre-estimation of the falsification frontier. In that case, estimation and inference is straightforward. Recall that, by definition, the falsification adaptive set simplifies to the standard estimand when the baseline model is not refuted. Thus there is no need to do a specification pre-test---researchers can just always present the estimated falsification adaptive set with accompanying confidence intervals. However, note that in finite samples the estimated falsification adaptive set will generally always be an interval. This situation is similar to HaileTamer2003, who note that their population bounds collapse to a single point in a special case, although in finite samples their estimator still generally gives an interval.
In this section and section (ref), we illustrate our recommendations using two variations of instrumental variable models. While many kinds of refutable assumptions have been considered in the literature, we focus on the classical case where variation from two or more instruments is used to falsify a model. In this section we consider the classical linear model with homogeneous treatment effects and continuous outcomes. In section (ref) we consider a model with heterogeneous treatment effects.
We thus begin with the classical constant coefficient model
where $Y(x,z)$ are potential outcomes defined for values $(x,z) \in \ensuremath{\mathbb{R}}^{K+L}$ and $U$ is an unobserved random variable. $\beta$ is an unknown constant $K$-vector. $\gamma$ is an unknown constant $L$-vector. Let $X$ be an observed $K$-vector of endogenous variables. Throughout we suppose $X$ does not contain a constant. Hence $U$ absorbs any nonzero constant intercept. Let $Z$ be an observed $L$-vector of potentially invalid instruments. For simplicity we have omitted any additional known exogenous covariates $W$ in equation (ref). In appendix (ref) we show how to easily include them via partialling out.
We observe the outcome $Y = Y(X,Z)$. Thus our equation for observed outcomes is \[ Y = X'\beta + Z'\gamma + U. \] Equation (ref) imposes homogeneous treatment effects since the difference between two potential outcomes \[ Y(x_1,z) - Y(x_0,z) = (x_1 - x_0)' \beta \] is constant across the population. We consider a model with heterogeneous treatment effects in section (ref).
In addition to this homogeneous treatment effect assumption, we maintain the following relevance and sufficient variation assumptions throughout this section.
A(ref) implies the order condition $L \geq K$. When there is just one endogenous variable ($K=1$), A(ref) only requires $\operatorname*{cov}(X,Z_\ell) \neq 0$ for at least one instrument. Other instruments may have zero correlation. In this case, these other instruments provide additional falsifying power. We discuss this further below. A(ref) simply requires that the instruments are not degenerate and are not affine combinations of each other.
The classical model imposes two more assumptions:
A(ref)--A(ref) imply that the coefficient vector $\beta$ is point identified and equals the two-stage least squares (2SLS) estimand. Furthermore, these assumptions imply well-known overidentifying conditions. The following proposition gives these conditions when there is just a single endogenous variable.
When all instruments are relevant, so that $\operatorname*{cov}(X,Z_\ell) \neq 0$ for all $\ell \in \{1,\ldots,L\}$, equation (ref) can be written as \[ \frac{\operatorname*{cov}(Y,Z_m)}{\operatorname*{cov}(X,Z_m)} = \frac{\operatorname*{cov}(Y,Z_\ell)}{\operatorname*{cov}(X,Z_\ell)}. \] That is, the linear IV estimand must be the same for all instruments $Z_\ell$. This result is the basis for the classical test of overidentifying restrictions (AndersonRubin1949, Sargan1958, Hansen1982). Suppose the distribution of $(Y,X,Z)$ is such that the model is refuted. This happens when at least one of our model assumptions fails: (a) homogeneous treatment effects, (b) linearity in $X$, (c) instrument exogeneity, or (d) instrument exclusion.
In this section we maintain the homogeneous treatment effects assumption and focus on instrument exclusion or exogeneity as an explanation for refutation. We consider heterogeneous treatment effects in section (ref). We also maintain linearity of potential outcomes in $X$. In principle our analysis can be extended to allow for relaxations of the linearity assumption, but we leave this to future work. Note that linearity holds automatically when $X$ is binary.
We thus focus on failure of (c) instrument exogeneity or (d) instrument exclusion as reasons for refuting the baseline model. These are two different substantive assumptions. Mathematically, however, we can use the same analysis to relax both assumptions. For simplicity, here we formally maintain the exogeneity assumption A(ref) and focus on failure of the exclusion assumption A(ref). In appendix (ref) we explain how to use our results to relax either or both assumptions.
We relax exclusion as follows.
This kind of relaxation of the baseline instrumental variable assumptions was previously considered by Small2007 and ConleyHansenRossi2012; also see AngristKrueger1994 and BoundJaegerBaker1995. For continuous instruments one may want to standardize the instruments to have variance one, to make the magnitudes of the components in $\delta = (\delta_1,\ldots,\delta_L)$ comparable. We discuss interpretation of $\delta_\ell$ further in section 3.6 of MastenPoirierFFarxiv; our empirical analysis in section (ref) does not require us to select or calibrate this parameter.
Under partial exclusion, the instruments may have a direct causal effect on outcomes. For sufficiently small values of the components in $\delta$, the model may nonetheless continue to be refuted. For sufficiently large values, however, the model will not be refuted. To characterize the falsification frontier, we begin by deriving the identified set for $\beta$ as a function of $\delta$.
The identified set $\mathcal{B}(\delta)$ depends on the data via two terms: \[\label{eq:defOfPsiPi} \underset{(L \times 1)}{\psi} \equiv \operatorname*{var}(Z)^{-1}\operatorname*{cov}(Z,Y) \qquad \text{and} \qquad \underset{(L \times K)}{\Pi} \equiv \operatorname*{var}(Z)^{-1}\operatorname*{cov}(Z,X). \] $\psi$ is the reduced form regression of $Y$ on $Z$. $\Pi$ is the first stage of $X$ on $Z$. If we demeaned $(Y,X,Z)$ then we would have $\psi = \ensuremath{\mathbb{E}}(ZZ')^{-1} \ensuremath{\mathbb{E}}(ZY)$ and $\Pi = \ensuremath{\mathbb{E}}(ZZ')^{-1} \ensuremath{\mathbb{E}}(ZX')$. Theorem (ref) shows that the identified set is the intersection of $L$ pairs of parallel half-spaces in $\ensuremath{\mathbb{R}}^K$. When $\delta = 0$, this identified set becomes the intersection of $L$ hyperplanes in $\ensuremath{\mathbb{R}}^K$. In this case, $\beta$ is point identified when \[ \operatorname*{cov}(Z,Y) = \operatorname*{cov}(Z,X) b \] for a unique $b \in \ensuremath{\mathbb{R}}^K$. If $\operatorname*{cov}(Z,Y) \neq \operatorname*{cov}(Z,X)b$ for all $b \in \ensuremath{\mathbb{R}}^K$, then the baseline model $\delta = 0$ is refuted.
Increasing the components of $\delta$ leads to a weakly larger identified set. Furthermore, there always exists a $\delta$ with large enough components so that $\mathcal{B}(\delta)$ is nonempty. We characterize the set of such $\delta$ below. Before that, we show that the identified set can be written as simple intersection bounds when there is a single endogenous variable.
To interpret this result, first consider an instrument $Z_\ell$ with a zero first stage coefficient, $\pi_\ell = 0$. If $Z_\ell$ has a sufficiently large covariance with the outcome, so that $\psi_\ell \pm \delta_\ell$ does not contain zero, then the model is refuted. Furthermore, in this case falsification can be solely attributed to the assumption that $| \gamma_\ell | \leq \delta_\ell$. This is similar to what is sometimes called the `zero first-stage test' (for example, see Slichter2014 and the references therein). When this covariance with the outcome is sufficiently small, however, $Z_\ell$ unsurprisingly has no falsifying or identifying power for $\beta$.
Next consider a relevant instrument $Z_\ell$; so $\pi_\ell \neq 0$. To interpret corollary (ref) in this case, we use the following lemma.
This lemma shows that $\psi_\ell / \pi_\ell$ is the population 2SLS coefficient on $X$ using $Z_\ell$ as the excluded instrument and $Z_{-\ell}$ as controls. Thus the identified set $\mathcal{B}(\delta)$ is the intersection of intervals around these 2SLS coefficients using one relevant instrument at a time and controlling for the rest.
Finally, consider the baseline case where $\delta = 0$. Corollary (ref) implies that $\mathcal{B}(0)$ is nonempty if and only if \[ \frac{\psi_m}{\pi_m} = \frac{\psi_\ell}{\pi_\ell} \] for any $m, \ell \in \{ 1,\ldots, L \}$ with $\pi_m, \pi_\ell \neq 0$. Moreover, in this case $\mathcal{B}(0)$ is a singleton equal to this common value. In this case---when the baseline model is not refuted---we also have
for all $\ell \in \{1,\ldots,L\}$. That is, $\psi_\ell / \pi_\ell$ equals the population 2SLS coefficient on $X$ using $Z_\ell$ as an instrument and not including $Z_{-\ell}$ as controls. This equality of single instrument 2SLS coefficients with and without controls for the other instruments is an alternative characterization of the classic overidentifying conditions from proposition (ref).
So far we have characterized the identified set for $\beta$ given a fixed value of $\delta$, the upper bound on the violation of the exclusion restriction. We now consider the possibility that this identified set is empty when $\delta = 0$, so that the baseline model is falsified. Our next result characterizes the falsification frontier, the minimal set of $\delta$'s which lead to a non-empty identified set. Here we focus on the single endogenous regressor case. We extend this result to multiple endogenous regressors in section (ref).
In the proof we show that this set satisfies our definition (ref) of the falsification frontier. Specifically: Any $\delta \in \text{FF}$ maps to a non-empty identified set, and strengthening any of the assumptions for a given $\delta \in \text{FF}$ leads to an empty identified set.
If we knew a priori that some instruments are valid, then we could obtain the falsification frontier for relaxing the potentially invalid instrument assumptions by simply setting $\delta_\ell = 0$ for the valid instruments $\ell$.
To better understand proposition (ref), consider the two instrument case. In this case, the following corollary characterizes the falsification frontier.
Thus, in the two instrument case, the falsification frontier is simply the line with horizontal intercept $(\delta_1^*,0)$ and vertical intercept $(0,\delta_2^*)$, where \[ \delta_1^* = \left| \frac{\psi_1}{\pi_1} - \frac{\psi_2}{\pi_2} \right| | \pi_1 | \qquad \text{and} \qquad \delta_2^* = \left| \frac{\psi_1}{\pi_1} - \frac{\psi_2}{\pi_2} \right| | \pi_2 | . \] The classical overidentifying restrictions (proposition (ref)) state that the baseline model is refuted if and only if $| \psi_1/\pi_1 - \psi_2/\pi_2 |$ is nonzero. A key aspect of our analysis is that not all refuted models are equivalent. Specifically, two dgps can have the same value of this difference but have different falsification frontiers.
Figure (ref) illustrates this point. It shows three example falsification frontiers, corresponding to three different distributions of $(Y,X,Z_1,Z_2)$. All three dgps have \[ \left| \frac{\psi_1}{\pi_1} - \frac{\psi_2}{\pi_2} \right| = 4. \] The relative strength of the two instruments differs across the dgps, however. Specifically, we set $| \pi_1 / \pi_2 |$ to be either $0.5$, $1$, or $2$. That is, the middle dgp has two instruments with exactly the same first stage coefficient. Hence there is a one-to-one trade off in the falsification frontier. If we weaken $Z_1$ relative to $Z_2$, however, the falsification frontier rotates inwards as the horizontal intercept $\delta_1^*$ gets smaller. Conversely, if we strengthen $Z_1$ relative to $Z_2$, the falsification frontier rotates outwards, as the horizontal intercept $\delta_1^*$ gets larger. That is, a false baseline model is harder to save by relaxing instrument exclusion for a strong instrument, relative to a weak one.
Thus far we have characterized the identified set for $\beta$ given any assumption $\delta \in \ensuremath{\mathbb{R}}_{\geq 0}^L$ as well as the falsification frontier. Next we characterize the falsification adaptive set, which is the identified set for $\beta$ under the assumption that one of the points on the falsification frontier is true. Here we again focus on the single endogenous regressor case. We generalize to multiple endogenous regressors in section (ref).
We first sketch the proof of this result and then discuss its implications. It follows from two main steps: First, the identified set $\mathcal{B}(\delta)$ is a singleton for any $\delta \in \text{FF}$ (see lemma (ref) in the appendix). Second, each of these singleton sets corresponds to an element in the interval on the right hand side of equation (ref) (follows using proposition (ref)). Thus we obtain the entire interval by taking the union over all of these singletons.
One of our main recommendations is that researchers report the falsification adaptive set. Theorem (ref) shows that, in the classical linear model we consider here, this set has an exceptionally simple form. Most importantly, no $\delta$'s appear on the right hand side of equation (ref). This implies that we can obtain the falsification adaptive set without pre-computing the falsification frontier or selecting any sensitivity parameters. Furthermore, it is very simple to compute, since it just requires running $L$ different 2SLS regressions. In appendix (ref) we show that the baseline 2SLS estimand using all instruments does not have to be inside the falsification adaptive set. In fact, the baseline 2SLS estimand can be arbitrarily far from the FAS.
In finite samples, researchers can present sample analogs or bias-corrected versions of equation (ref), along with corresponding confidence sets (e.g., ChernozhukovLeeRosen2013). We further discuss finite sample practice in section (ref).
In this model, we can also immediately see how this set adapts to falsification of the baseline model. When the baseline model is not false, $\psi_m / \pi_m = \psi_\ell / \pi_\ell$ for all $m, \ell \in \{1,\ldots,L\}$ with nonzero $\pi_m$ and $\pi_\ell$. In this case, the falsification adaptive set collapses to the singleton equal to the common value. This is the same point estimand researchers would usually present when their baseline model is not refuted. As the baseline model becomes more refuted, the values of $\psi_\ell/\pi_\ell$ become more different, and the falsification adaptive set expands. Thus the size of this set reflects the severity of baseline falsification.
To show how all of the concepts we have discussed fit together, we return to the earlier two instrument numerical illustration discussed earlier. Figure (ref) extends figure (ref) by adding a third dimension: Values of $\beta$. Specifically, each column in figure (ref) is a different dgp. Within a column, the top and bottom row are just different viewing angles. For each plot, the horizontal plane shows different values of $(\delta_1,\delta_2)$. The solid line in that plane shows the falsification frontier, as in figure (ref). The vertical axis shows the identified set $\mathcal{B}(\delta)$ as a function of $\delta$. For $\delta$ values sufficiently close to zero, this identified set is empty. But for $\delta$'s beyond the falsification frontier, this set is nonempty. For any point $\delta$ on the falsification frontier, this set is a singleton. But the set grows as both $\delta_1$ and $\delta_2$ increase. Finally, the falsification adaptive set is shown as the solid line segment on the vertical axis. It is the union over over the identified sets $\mathcal{B}(\delta)$ as $\delta$ varies over the falsification frontier. Since $\mathcal{B}(\delta_1^*,0) = \{ \psi_2/\pi_2 \}$, $\mathcal{B}(0,\delta_2^*) = \{ \psi_1/\pi_1 \}$, and $\mathcal{B}(\delta)$ varies monotonically as we traverse the falsification frontier, the falsification adaptive set is simply the interval between the points $\psi_1/\pi_1$ and $\psi_2/\pi_2$. In this illustration, these points are $1$ and $5$ for all three dgps and hence the falsification adaptive set is the interval $[1,5]$.
With more than two instruments, we cannot draw the space of $\delta$'s and $\beta$ at the same time. Nonetheless, theorem (ref) shows that we can compute the falsification adaptive set by simply taking the convex hull of the $L$ different 2SLS estimands $\psi_\ell / \pi_\ell$.
In section (ref) we discussed estimation and inference in general. Here we discuss the specific case given in theorem (ref). This characterization of the falsification adaptive set requires that we first screen for weak instruments. It is not clear how to best do this. We present a first pass approach, but leave a detailed analysis to future work. Let \[ \mathcal{L} = \{ \ell \in \{1,\ldots,L \} : \pi_\ell \neq 0 \} \] be the set of indices corresponding to relevant instruments. Estimate this set by \[ \widehat{\mathcal{L}} = \{ \ell \in \{1,\ldots,L \} : F_\ell \geq C_n \}. \] $F_\ell$ is the first stage $F$-statistic when considering $Z_\ell$ as an instrument and $Z_{-\ell}$ as controls. $C_n$ is a cutoff that converges to zero as the sample size grows. Specifying the cutoff to shrink ensures that asymptotically we only discard instruments whose coefficients are exactly zero, so that $\widehat{\mathcal{L}}$ is consistent for $\mathcal{L}$.
We then estimate the falsification adaptive set by \[ \widehat{\text{FAS}} = \left[\min_{\ell \in \widehat{\mathcal{L}}} \hspace{-1.5mm} \widehat{\phantom{\big(}\frac{\psi_\ell}{\pi_\ell}\phantom{\big)}} \hspace{-1.5mm} , \ \max_{\ell \in \widehat{\mathcal{L}}} \hspace{-1.5mm} \widehat{\phantom{\big(}\frac{\psi_\ell}{\pi_\ell}\phantom{\big)}} \hspace{-1.5mm} \right] \] where $\widehat{\psi_\ell / \pi_\ell}$ is the estimated 2SLS coefficient on $X$ using $Z_\ell$ as the excluded instrument and $Z_{-\ell}$ as controls. We use this estimator in our empirical analysis of section (ref). There we use $C_n = 10$ as our default, although we sometimes consider other cutoffs, or a sequence of cutoffs.
Beyond presenting the falsification frontier and the falsification adaptive set, researchers may want to also present $\mathcal{B}(\delta)$ for specific choice of model relaxation. We consider the case where the researcher specifies the direction of $\delta$ a priori and then chooses the smallest magnitude $\| \delta \|$ such that $\mathcal{B}(\delta) \neq \emptyset$.
Let $\delta = m \cdot d$ for $m \in \ensuremath{\mathbb{R}}$ and known $d = (d_1,\ldots,d_L)' \in (0,\infty)^L$. For example, when all the instruments are continuous, one reasonable direction may be $d = (\sqrt{ \operatorname*{var}(Z_1)}, \ldots, \sqrt{ \operatorname*{var}(Z_L)} )'$. Given this fixed direction $d$, our assumptions are now parameterized by a single scalar, $m$. The following proposition shows that there is a unique falsification point $m^*$.
The following corollary characterizes the identified set at the point on the falsification frontier in the direction $d$.
Theorem (ref) characterizes the identified set for the vector of coefficients on the endogenous variables, as a function of the exclusion restriction relaxation. Our subsequent characterizations of the falsification frontier and the falsification adaptive set, however, restricted attention to the case with just one endogenous variable---see proposition (ref) and theorem (ref). In this section, we extend these two results to the general case with $K \geq 1$ endogenous variables. These results can be used in at least three different cases: (1) If there are multiple `basic' endogenous variables. For example, in our empirical application in section (ref), one could allow both price and advertising spending to be endogenous. (2) If the outcome equation has interactions of one basic endogenous variable with covariates. (3) If the outcome equation is nonlinear in one basic endogenous variable. For example, it could be a quadratic function.
Recall our notation for the reduced form and first stage regressions: \[ \underset{(L \times 1)}{\psi} \equiv \operatorname*{var}(Z)^{-1}\operatorname*{cov}(Z,Y) \qquad \text{and} \qquad \underset{(L \times K)}{\Pi} \equiv \operatorname*{var}(Z)^{-1}\operatorname*{cov}(Z,X). \] $L$ denotes the number of instruments while $K$ denotes the number of endogenous variables. Let $\pi_\ell'$ denote the $\ell$th row of the matrix $\Pi$. When $K = 1$, this is a scalar and $\pi_\ell' = \pi_\ell$.
When there is a single endogenous variable, theorem (ref) shows that the falsification adaptive set is the interval \[ \left[\min_{\ell=1,\ldots,L:\pi_\ell \neq 0} \frac{\psi_\ell}{\pi_\ell}, \ \max_{\ell=1,\ldots,L:\pi_\ell \neq 0} \frac{\psi_\ell}{\pi_\ell}\right]. \] This interval can be interpreted as follows: Pick any set of $K$ instruments. Impose the exclusion restriction for those instruments. For the remaining instruments, completely drop the exclusion restriction. This yields a non-refutable model that is weaker than the original model which imposed exclusion for all of the instruments. But this weaker model is still strong enough that $\beta$ is point identified (ignoring the possibility of irrelevant instruments, which we discuss below). Moreover, in this weaker model $\beta$ equals the population 2SLS coefficient on $X$ using the chosen $K$ instruments as the excluded instruments and the remaining instruments as controls. Collect all of these different 2SLS estimands, as we vary which set of $K$ instruments we pick. Then take their convex hull.
When $K=1$, this convex hull is precisely the interval above: We pick a single relevant instrument to exclude, and include the others as controls. This gives us a single just-identified 2SLS point estimand. By cycling through which instrument we exclude, we obtain a variety of different point estimands. Since $K=1$, these estimands are scalars. So their convex hull is simply the interval between the smallest and largest values.
When $K > 1$ and $L = K + 1$, the falsification adaptive set is precisely the convex hull of just-identified 2SLS estimands that we've just described. For $L > K + 1$, it is somewhat more complicated. We discuss these general results next.
In the $K=1$ case we allowed for irrelevant instruments. For $K>1$ we focus on the case where all instruments are relevant for simplicity. Specifically, we impose assumption A(ref)$^\prime$ below. To state this assumption, we consider submatrices of $\Pi$. Let $\mathcal{L} \subseteq \{ 1,\ldots, L \}$. Let $\Pi_\mathcal{L}$ be the $| \mathcal{L} | \times K$ submatrix of $\Pi$ formed by removing all rows $\ell \notin \mathcal{L}$.
A(ref).1$^\prime$ is equivalent to the following:
For example, suppose $L = K+1$. Then dropping the exclusion restriction for any single instrument returns us to a non-refutable model. Here we simply assume that $\beta$ is still point identified regardless of which exclusion restriction we choose to drop. Finally, note that A(ref).1$^\prime$ implies our earlier relevance assumption A(ref). A1.2$^\prime$ means that there does not exist a hyperplane that passes through all of the $\pi_\ell$ vectors. It is equivalent to linear independence of $(\pi_L - \pi_1,\ldots, \pi_2 - \pi_1)$.
Let
denote the polytope defined by the convex hull of the set of just-identified 2SLS estimands.
To illustrate theorem (ref), consider the two endogenous variables ($K=2$) and three instruments ($L=3$) case. Since there are two endogenous variables, there are two coefficients we're interested in. Consider the left plot in figure (ref). This plot shows possible values $(b_1, b_2)$ for these two coefficients. The exclusion restriction from instrument $\ell$ imposes a single linear constraint $\psi_\ell = \pi_\ell' b$. These constraints are simply lines in $\ensuremath{\mathbb{R}}^2$. Since there are three instruments, there are three constraints. When these three lines do not intersect at a common point, the baseline model is refuted. This case is shown in the figure. Suppose we drop the exclusion restriction for instrument 1. Then two linear constraints remain, $\beta$ is point identified, and it equals the intersection point $\beta_{\{2,3\}}^\textsc{2sls}$ in the lower right of the figure. A similar interpretation applies to the other two intersection points $\beta_{\{1,3\}}^\textsc{2sls}$ and $\beta_{\{1,2\}}^\textsc{2sls}$. The falsification adaptive set is then simply the convex hull of these three points in $\ensuremath{\mathbb{R}}^2$, which is shown as the shaded triangular region.
Next consider the case with $K > 1$ and $L > K + 1$. We handle this case by reducing it to the case we just studied where $L = K + 1$. Let \[ \mathcal{P}_{\mathcal{L}} = \text{conv} \big( \big\{ \beta_{\mathcal{L} \setminus \{ \ell \}}^\textsc{2sls} : \ell \in \mathcal{L} \big\} \big). \] In the following proposition, we consider subsets of indices $\mathcal{L} \subseteq \{ 1,\ldots, L \}$ such that $| \mathcal{L} | = K + 1$. Thus $| \mathcal{L} \setminus \{ \ell \} | = K$. So $\beta_{\mathcal{L} \setminus \{ \ell \}}^\textsc{2sls}$ is a just-identified 2SLS estimand. We form $K+1$ different estimands by cycling through which exclusion restriction to drop. We then take their convex hull. This set is similar to the set defined in equation (ref). The key difference is that here we only look at the estimands which use all but one index from a reference set $\mathcal{L}$.
We use Farkas' lemma and Caratheodory's theorem from convex analysis to prove these results. As in the previous cases, $\mathcal{P}$ can be computed by just running a variety of 2SLS regressions. Next note that, although each $\mathcal{P}_\mathcal{L}$ is convex, their union generally is not. Nonetheless, we are often only interested in linear functionals of the coefficient vector $\beta$. For example, we often care about just one component of $\beta$. The following corollary shows that the falsification adaptive set for a linear functional of $\beta$ again has a simple form. For this result, recall the definition of $\text{FAS}^*$ from equation (ref).
This result shows that we can simply cycle through all possible just identified models, compute the corresponding 2SLS estimand, take the convex hull, and project it onto one component to get the FAS for that component.
The right plot in figure (ref) illustrates the $L > K+1$ case. Here we have $K=2$ and $L=4$. There are 6 different just-identified 2SLS estimands. The falsification adaptive set is no longer a convex set. Nonetheless, the projection of the convex hull of all just-identified 2SLS estimands onto the first component still gives the falsification adaptive set for $\beta_1$. Moreover, this projection can be computed by simply taking the largest and smallest estimated values of $\beta_1$, among the just-identified 2SLS estimands. In this sense, corollary (ref) directly generalizes theorem (ref) to the case $K > 1$.
We conclude this subsection by briefly discussing estimation and inference. Corollary (ref) shows that the FAS for a linear functional of $\beta$ has an intersection bounds form. As in the $K=1$ case, researchers can present sample analog estimators or bias-corrected versions, along with corresponding confidence sets (see ChernozhukovLeeRosen2013). Our discussion of weak instruments in section (ref) applies here as well. More generally, researchers may want to do estimation and inference on the FAS for the entire vector $\beta$. Theorem (ref) shows that this set is defined by a finite set of unconditional linear moment inequalities. Note that here we have omitted covariates. Depending on how covariates enter the model, the FAS may continue to depend only on unconditional linear moment inequalities, or it may also depend on conditional linear moment inequalities. See appendix (ref) for details. While there is a large literature on inference in general moment inequality models---see CanayShaikh2017 and Molinari2019 for surveys---there is now a growing literature on the linear case. This includes HsiehShiShum2017, AndrewsRothPakes2019, ChoRussell2019, and Gafarov2019. We conjecture that some of these results can be applied to do inference on the FAS for $\beta$, the FAS for subvectors of $\beta$ with two or more components, and the FAS for nonlinear functionals of $\beta$. We leave a full analysis for future work, however.
{0.03em} {0.03em}
In this section we apply our results from section (ref) to four different empirical studies: \citet*[The Review of Economic Studies]{DurantonMorrowTurner2014}, \citet*[The Quarterly Journal of Economics]{AlesinaGiulianoNunn2013}, \citet*[The American Economic Review]{AcemogluJohnsonRobinson2001}, and Nevo2001. Each paper reports 2SLS estimates using multiple instrumental variables and each paper discusses concerns about instrument validity. In particular, all four papers run overidentification tests, which sometimes fail. Even when they do not fail, the authors sometimes express a concern that this could simply be due to low sample size. To address these concerns, we estimate falsification adaptive sets and compare them with the results reported in the original papers. We argue that the FAS is an informative complement to traditional overidentification test $p$-values: Rather than focusing on the null hypothesis that the instruments are consistent with each other, the FAS summarizes the range of estimates obtained from alternative models which which are not refuted by the data.
DurantonMorrowTurner2014 study the relationship between roads and trade. Specifically, they consider a dataset of 66 regions (`cities') in the United States. Their treatment variable is the log number of kilometers of interstate highways within a city, in 2007. This variable directly affects the cost of leaving a city, and therefore the cost of exporting from a city: It is easier to export from a city with many kilometers of interstate highways passing through it. Their outcome variable is a measure of how much that city exports. They consider two different ways of measuring exports: Weight (in tons) and value (in dollars). We focus on the weight measure for brevity. We discuss the value measure in appendix (ref). They begin by estimating a gravity equation relating the weight of a city's exports to other cities with the highway distance between those cities, both measured in 2007. This equation includes a fixed effect for the exporting city. The estimate of this fixed effect is their main outcome variable. They call this variable the “propensity to export weight.” Thus their main goal is to estimate the causal effect of within city highways on the propensity to export weight.
We cannot learn this causal effect by simply regressing the propensity to export weight on within city highways since there is a classic simultaneity problem. We expect that building highways within the city will boost exports. But high export cities may also build more highways to facilitate their existing exports. The authors solve this problem by instrumenting for the number of kilometers of within city highways. They consider three different instruments:
We first summarize their arguments for validity of these instruments. We then review their analysis and present our new results.
Instrument relevance can be directly assessed from the first stage results in the data. We discuss this later. Still, there are a priori reasons to believe the instruments will be correlated with treatment. Building railroad tracks requires leveling ground. So does building roads. Thus old unused railroad tracks can be easily converted to highways. Exploration routes were easy to travel on foot, horseback, or wagon. Such routes are also likely to be good for cars. Finally, most of the highways on the 1947 plan were ultimately built, although additional highways not on the plan were built as well.
Next consider instrument exogeneity and exclusion. Exogeneity concerns the presence of omitted variables that affect the instrument and the outcome, or of simultaneity between the instrument and the outcome. Exclusion concerns the direct causal effect of the instrument on outcomes. We consider each instrument in turn:
Overall, the authors raise concerns about validity of all three instruments. Although they address these concerns with various controls, these controls may still not perfectly fix failures of exogeneity, exclusion, or both. Hence the authors lean on overidentification, stating that
With this motivation, we next present the results.
First consider table (ref). In this and all other tables, the non-highlighted parts reproduce results from the original paper. The highlighted parts are new computations which we have added. Panel A reproduces columns 1--4 of table 5 in DurantonMorrowTurner2014. These are their main results. In particular, they are interested in the coefficient on log highway km, the log number of highway kilometers within the city. This coefficient represents their estimate of the causal effect of roads on trade. Here it is estimated by 2SLS, using railroads, exploration, and plan as instruments. As mentioned above, we are concerned about possible invalidity of the instruments. Thus they implement the standard test of overidentifying restrictions. At conventional sizes, it easily passes in the longest specification, fails in the second specification, and marginally passes in the first specification. Also note that these specifications do not include all of the controls discussed above; the authors include those in separate analyses, which we discuss later (our table (ref)).
We add the falsification adaptive set to these baseline results. This is the last row of panel A. There are two things to notice: First, except for the last specification, none of the 2SLS estimates are within the FAS. We knew this was possible, given our theoretical results in appendix (ref) (which relate the 2SLS estimand to the FAS). Second, the FAS magnitudes are all generally smaller than the 2SLS point estimates.
To better understand how we computed the FAS, and how to interpret it, next consider table (ref). Columns 1--3 include the same baseline controls as column 3 in table (ref) while columns 4--6 include the same baseline controls as column 4 in table (ref). The only difference is that we no longer use all three variables (plan, railroad, exploration) as instruments. Instead, in panel A, we use only one of these variables as an instrument and we ignore the other two variables. Panel A reproduces columns 4--6 from table 6 in DurantonMorrowTurner2014. The authors used these results as their main robustness check. They argue that the three estimates 0.38, 0.64, and 0.34 from columns 4--6, panel A, table (ref) are consistent with their baseline estimates of 0.47 and 0.39 from columns 3 and 4, panel A, table (ref).
However, as we have discussed, omitting an invalid instrument can lead to omitted variable bias. In this application we are concerned that some of the instruments may be invalid. Thus the alternative models of interest are those where one of the instruments is valid but the others are not. When computing results in these alternative models, the invalid instruments should be included as controls. Panel B shows these results. Here we use one instrument while controlling for the other two. For example, in column 1 we use plan as an instrument and control for railroad and exploration.
For brevity, here we only describe the results in columns 4--6. These results use the full baseline specification. Consider column 5, panel B. This result uses railroad as an instrument, controlling for plan and highway. Unlike the uncontrolled result from panel A, railroad is a very weak instrument. Hence we ignore the result using railroad alone, as discussed in section (ref). Next consider column 4. Here we use plan as the instrument, controlling for the other two. Despite these controls, plan is still a strong instrument. The estimated effect 0.18 is roughly half as large as the estimate from panel A, 0.38. It is also no longer statistically significant at any conventional level. Next consider column 6. Here we use exploration as the instrument, controlling for the other two. Exploration continues to be a strong instrument with these controls. The estimated effect 0.42 in panel B is similar to the effect from panel A, 0.34. It is no longer statistically significant, however.
Putting these coefficient estimates together gives us the estimated FAS, $[0.18, 0.42]$. The endpoints of this set correspond to point estimates from alternative models which maintain validity of only one instrument at a time. The interior of this set corresponds to alternative models which relax validity of all instruments at once, but just enough to avoid falsification. Thus the FAS reflects model uncertainty: Relying on different instruments to different degrees yields different results. This range of results is given by the FAS.
In panel B of table (ref) we found that railroad is a weak instrument when controlling for the other two, and hence it yields the largest point estimates. Given this finding, one may also wonder how removing railroads as an instrument affects the baseline analysis. This is shown in panel B of table (ref). All of the coefficients on log highway km are smaller, to the point that they are no longer statistically significant for all but the shortest specification. Moreover, the standard overidentification tests are now all easily passed. (Note that these tests are only comparing results using plan and exploration as instruments.) However, the coefficients on railroads are statistically significant for all but the fourth column. This suggests that the full baseline model using all three instruments could be rejected, and also explains the source of the relatively small overidentification test $p$-values in panel A.
Thus far we have focused on the baseline results, which do not include all of the possible control variables discussed earlier. Table (ref) shows results with these controls. We begin with the full set of baseline control variables, as used in column 4 of table (ref). We then add just one control. Each row corresponds to a different control. The last row shows the results that add all controls at once. Unlike the main baseline result, column 4 of table (ref), here we only use one instrument at a time. Columns 1--3 use a single instrument, without controlling for the other two. These results reproduce appendix table 6 of DurantonMorrowTurner2014. Based on these results, the authors argue that
They also argue that using one instrument at a time is an “even more demanding exercise” than examining the effect of additional controls when using all three instruments (not shown here; see their appendix table 5). As we've discussed, however, omitting the invalid instruments may cause omitted variable bias. So in columns 4--6, we replicate columns 1--3, except now controlling for the other two instruments.
There are three main differences between the results with the instrument controls and those without. First, the railroad instrument is again very weak, leading to large coefficients. This informs our understanding of the results from columns 1--3, since there we observed that the coefficients in column 2 are always larger than those in columns 1 and 3, and are often substantially larger. Second, none of the results are statistically significant at conventional levels. Finally, the coefficients using plan as the instrument all become smaller once the other instruments are controlled for (column 4 versus column 1), while the coefficients using exploration as the instrument all become larger once the other instruments are controlled for (column 6 versus column 3). Thus, ignoring the results using railroads, the overall range of point estimates is larger. This is reflected in the falsification adaptive sets, which are presented in the final column.
Overall, there are two main conclusions from our analysis: First, the evidence suggests that the railroad instrument is the most questionable, and should be used only as a control. Thus the estimates in panel B of table (ref) are arguably the most appropriate baseline results. Second, there is substantially more uncertainty in the magnitude of the causal effect of roads on trade than suggested by the original results of DurantonMorrowTurner2014. This is reflected in the various falsification adaptive sets we present. In particular, the FAS for the longest specification is $[0.18, 0.69]$; see table (ref). Accounting for sampling uncertainty only increases this range. That said, these results do not change the overall qualitative conclusions of the paper: All points in the estimated FAS for the longest specification are still positive, suggesting that the number of within city highways appears to positively affect propensity to export weight.
AlesinaGiulianoNunn2013 study the relationship between traditional agricultural practices and modern gender norms. Specifically, following Boserup1970 they distinguish between two kinds of traditional agriculture: shifting cultivation, which “is labor intensive and uses handheld tools like the hoe and the digging stick,” and plough cultivation, which “is much more capital intensive, using the plough to prepare the soil.” They note that “unlike the hoe or digging stick, the plough requires significant upper body strength, grip strength, and bursts of power.” Hence “when plough agriculture is practiced, men have an advantage in farming relative to women.” Consequently, in societies that used plough agriculture, “men tended to work outside the home in the fields, while women specialized in activities within the home.” Boserup hypothesized that, as AlesinaGiulianoNunn2013 put it, “this division of labor then generated norms about the appropriate role of women in society” which persist today.
AlesinaGiulianoNunn2013 consider a variety of datasets and methods to test this hypothesis. We focus on their cross country instrumental variable analysis. Their dataset includes 160 countries. Their treatment variable is the “estimated proportion of citizens with ancestors that traditionally used the plough in pre-industrial agriculture.” They explain how they construct this variable in section 3. For this analysis they consider three different outcome variables: The female labor force participation rate in 2000, the share of a country's firms with some female ownership (measured using data between 2005--2011), and the share of political positions held by women in 2000. Their appendix A2 gives further details on the definition of these variables. They also consider a composite outcome variable, called the average effect size (AES). This is constructed by first dividing each outcome variable by its standard deviation and then averaging the three variables.
Any correlation between treatment and outcomes is not necessarily causal for a few reasons. First, there may be reverse causality: Countries with less equal gender norms may have actively adopted innovations like the plough which conformed with these norms. Second, there may be omitted variables: Areas that were economically more developed may have been more likely to both adopt the plough and to have more equal gender norms today. This would tend to mask any causal effect of the plough on gender norms. To address these concerns, the authors use an instrumental variable approach.
The authors use properties of countries' ancestral land geography as instruments. Specifically, Pryor1985 classifies crops as either plough positive or plough negative. Plough positive crops benefit from the use of the plough while plough negative crops generally do not. The authors focus on cereal crops only, since they are similar except for how much they benefit from the plough. Given this classification, they construct instruments as follows: Consider a large piece of land. First measure the area of land that is suitable for growing crops in general. Pick a single plough positive cereal. Within this area of overall suitability, compute the area of land suitable for growing this specific plough positive cereal. Do this for each plough positive cereal. Note that some land is suitable for growing multiple crops, so there will be overlap in these areas. So take the average area across all plough positive cereals. Divide by the area that is suitable for growing crops overall. That is their measure of plough positive share. Next do the same for plough negative crops. Thus we have a procedure for defining suitability of a single piece of land. Using this procedure, the authors define the country level variable plough positive environment as the “average fraction of ancestral land that was suitable for growing [plough positive cereals] divided by the fraction that was suitable for any crops.” They use the same procedure as in their definition of the treatment variable, described in their section 3. They define the country level variable plough negative environment analogously.
Their main concern with validity of these instruments is failure of exogeneity due to geographical variables which affect the instruments and also modern gender norms. In their baseline results (table 8) they include a variety of geographical controls to account for this. They consider more extensive geographical controls in appendix table A14. They consider many other controls in appendix table A15. Finally, they also use the presence of two instruments to conduct overidentification tests.
Table (ref) reproduces their main instrumental variable results. As in our analysis of DurantonMorrowTurner2014, the non-highlighted parts reproduce results from the original paper. The highlighted parts are new computations which we have added. First focus on their results. The overidentification test fails at the conventional 5% level for columns 1 and 2, as well as for the average effect size, columns 7 and 8.\footnote{The authors report overidentification test $p$-values of 0.31 in column 4 and 0.06 in column 8. These differ slightly from the $p$-values we report in table (ref). The reason is that we use heteroskedasticity robust standard errors for all calculations, including columns 4 and 8. The authors used homoskedastic standard errors for their columns 4 and 8 overidentification test (but not for the standard errors reported under their point estimates).} The authors do not comment on this finding. To provide a constructive response to this rejection, we compute the falsification adaptive set. To do this, we must first check instrument relevance. As shown in panel B, plough positive environment satisfies the relevance assumption. Plough negative environment, however, does not. The authors commented on this, since they did present the first stage coefficients on both instruments, although they did not present the separate $F$-test statistics as we do. Despite this relevance failure, the authors used both instruments in their main results.
In panel A, we present the estimated falsification adaptive set. Since there are only two instruments, one of which fails relevance, this set is simply a singleton. Specifically, this is the point estimate from the model using plough positive environment as an instrument, controlling for plough negative environment. Overall, the authors' qualitative findings all continue to hold using the FAS: All point estimates are still negative and statistically significant at conventional levels. Quantitatively, however, the FAS point estimates all have smaller magnitudes than the original baseline findings, except for columns 5 and 6. Excluding those columns, the point estimates are an average of 22% smaller. Finally, note that the columns where the estimates changed the least---columns 5 and 6---are those with the largest overidentification test $p$-values. Likewise, the columns where the estimates changed the most---columns 1 and 2---are the ones with the smallest overidentification test $p$-values. This is to be expected, given how the overidentification test is constructed. Our point is that the FAS is an informative complement to these $p$-values, since it directly summarizes the range of estimates corresponding to non-refuted alternative models. As in our analysis of DurantonMorrowTurner2014, here it also highlighted the weakness of the plough negative instrument, suggesting that the FAS point estimates we present in panel A of table (ref) are more appropriate baselines than those originally presented by the authors.
\newcolumntype{K}[1]{>{\arraybackslash}p{#1}}
AcemogluJohnsonRobinson2001 study the relationship between institutions and economic development. Since this paper is exceptionally well known, here we only briefly summarize it and then present our results. They consider a cross section of about 60 countries. Their treatment variable is the average protection against expropriation risk between 1985 and 1995, a measure of the strength of property rights in the country. Their outcome variable is log GDP per capita in 1995. Their primary instrument measures mortality rates faced by early settlers in the country. Their main results (table 4) use just this one instrument. In table 8, however, they consider five more instruments. We focus on that analysis.
They use the additional instruments as follows: First, they only ever consider pairs of instruments. As they say in footnote 29, they do this to avoid the many instruments problem. We follow them and also only consider two instruments at a time. For each of the five additional instruments, they estimate the coefficient on the treatment variable using either (a) that instrument alone (their panel A) or (b) that instrument alone but also including settler mortality as a control variable (their panel D). They test overidentification by using a Hausman test to compare these two coefficient estimates (their panel C). They also present the estimated coefficient on settler mortality when it is included as a control, noting that it is not statistically indistinguishable from zero at conventional levels (their panel D).
In our table (ref) we replicate and extend those results. Whereas AcemogluJohnsonRobinson2001 focused on using these additional instruments to perform overidentification tests, we instead focus on the point estimates themselves. Thus in our panel A we present a new set of baseline results, which use two instruments at a time. AcemogluJohnsonRobinson2001 did not present these results. Here we include the classical overidentification test $p$-value; the authors had instead presented a Hausman test result. Columns 1--5 present results with no control variables (except for 3 and 4, which include a single control: years since independence) while columns 6--10 add a control for latitude.
The overidentification tests do not reject at conventional levels in each column. Nonetheless, it is informative to ask: What estimates would we obtain from alternative non-refutable models? The last line of panel A answers this question by presenting the FAS. This set is computed from the estimated coefficients on the treatment variable in panels B and C. The authors had originally computed one of the endpoints of this set, although they focused on the coefficient on the controlled instrument. Here we compute the other endpoint of the FAS. First note that none of the instruments are particularly strong. This is a well known problem in this specific empirical application (for example, see pages 3072--3073 of albouy2012). Since we have not formally studied inference on the FAS, we leave the question of how to accurately measure sampling uncertainty in this application to future work. Here we will simply focus on the point estimates. For these estimates, we use a small $F$-statistic cut-off for inclusion in the FAS, so that the FAS is an interval for all specifications. For every specification, the estimated FAS always lies above zero. Hence the FAS supports the qualitative conclusion of this paper: Increasing property rights causes an increase in GDP. That said, for some of the specifications the length of the FAS is quite large, indicating substantial variation in the estimated magnitude of this effect.
As in our previous examples, the authors used multiple instruments to assess validity of the exogeneity and exclusion restrictions. They did this by focusing on a variety of statistical tests of overidentification. Here we illustrated the value of using the falsification adaptive set to go beyond the binary pass/fail outcome of a specification test and to instead summarize the substantive variation in estimates obtained from alternative non-refutable models.
In our final example, we study a classical application of instrumental variables: estimation of consumer demand. In a series of papers Aviv Nevo estimated the demand for ready-to-eat cereal as a first step to answering several important economic questions: How can we predict the welfare effects of a merger (Nevo2000)? How can we test for collusion (Nevo2001)? How can we construct price indices that account for new products and quality changes (Nevo2003)? Answering these questions requires reliable estimates of price elasticities of demand. If one's baseline model of consumer demand is refuted, however, it is not clear whether the estimated elasticities from this known-to-be misspecified model are relevant for these questions. For this reason, we recommend computing falsification adaptive sets. These sets show the range of estimates consistent with non-refuted models.
Since all three of the papers Nevo2000, Nevo2001, Nevo2003 use essentially the same data and demand estimation, we focus on Nevo2001. First we briefly summarize his analysis, focusing on falsification. We then discuss our results. Here we assume familiarity with standard discrete choice demand models. See Nevo2011 for a survey.
Nevo begins by estimating a variety of logit demand models; see his table 5. In 9 out of the 10 specifications, the overidentification test rejects the model. Commenting on this, he says “it is unclear whether the large number of observations is the reason for the rejection or that the IV's are not valid” (page 325). He notes that the most flexible specification, column 10, passes the test. This specification added city fixed effects, and hence he concludes that it is important to flexibly control for demographics. In our results below, however, the specification with city fixed effects is strongly rejected.
In logit demand models estimated using overidentified 2SLS, the model can be refuted for two distinct reasons: (1) The instruments are inconsistent with each other or (2) the logit functional form is incorrect. The second issue is well known and is commonly addressed by using more flexible models of demand, like the random coefficients logit model. Nevo presents estimates from this model in table 6. Using these estimates he notes that the logit model “is easily rejected” (footnote 28 on page 332).
The random coefficients logit model itself, however, is falsifiable. Nevo does not report the overidentification test result for his full model. He does report the GMM objective function value in table 6, but it is not clear if this objective function uses the optimal weighting matrix, as required to compute the $J$-test statistic. If it does, and assuming the statistic has not already been scaled by the sample size, then the overidentification test based on the reported value easily rejects the random coefficients logit model as well. Regardless, we also estimated the random coefficients logit model using our data (not reported here), and it is easily rejected.
As with the baseline logit model, this could be either because (1) the instruments are invalid or (2) the random coefficients logit functional form is incorrect. One response is to use even more flexible models of demand. For recent work along these lines, see Compiani2018 and TebaldiTorgovitskyYang2019. An alternative approach is to maintain the random coefficients logit functional form, but relax the exogeneity and exclusion restrictions on the instruments. Or one could relax both assumptions simultaneously. Here we focus only on the instrument assumptions. Moreover, we also focus only on the baseline logit demand model. Our characterization of the falsification adaptive set (theorem (ref)) can be extended to the random coefficients logit model, but since this is a substantial extension we leave it to separate work.
Finally, recall that Berry1994 shows how the nested logit model can be written as a linear instrumental variable model with multiple endogenous variables. This result combined with our multiple endogenous variable results in section (ref) allows us to immediately study the nested logit model with endogenous price. Moreover, nested logit has been used as the baseline model in previous studies of the demand for ready-to-eat cereal---see CotterillHaller1997. We omit a full empirical study of the nested logit baseline for brevity, however.
Unlike our previous three applications, here the original data is not publicly available. We instead construct a new dataset, following the details of Nevo2001 as closely as possible. We consider a panel dataset of 53 cities between 2012 and 2016. Nevo's data covered 1988 to 1992. We obtained quarterly data on price and quantity of each cereal sold in these cities from Nielsen scanner data. We focus on the top 25 cereal brands, shown in table (ref). Table (ref) shows the volume market shares (pounds sold) for the top 6 firms, for the last quarter of each year. Comparing this to table 1 from Nevo2001, we see that firm market shares are relatively stable over time. The industry continues to be highly concentrated, with the top 3 firms having around 70% of the market and the top 6 having around 84%. The market share of private label cereals has increased somewhat, ranging from 12 to 15% in our data, whereas it ranged from 3 to 8% between 1988 and 1992.
{
}
We obtained quarterly advertising spending from Kantar Media. Table (ref) shows summary statistics for advertising, along with price and market shares. Compared to Nevo's table 4, we see that prices are generally substantially lower, with less variation over time. Advertising spending has increased substantially.
{
}
We gathered nutrition data for each cereal by examining cereal boxes. For each city we obtained demographic information by sampling individuals from the March Current Population Survey. Table (ref) provides summary statistics for nutrition and demographics.
Nevo considered two kinds of instruments:
Table (ref) shows our main results. We consider four specifications. All specifications include advertising and time fixed effects. Column 1 includes product characteristics as controls. Columns 2--4 replace these with brand dummies. Column 3 adds demographic controls: log of median income, log of median age, and median household size. Column 4 replaces this with city fixed effects. Panel A shows the OLS results. Panel B shows 2SLS results using the per-quarter Hausman instruments while panel C uses the aggregated Hausman instruments.\footnote{In panel A: Column 1 corresponds to column 1 of Nevo's table 5, Column 2 to his column 2, and column 3 to his column 3. He did not present the specification in our panel A, column 4. In panel B: Column 3 corresponds to Nevo's column 9, and column 4 to his column 10. He did not present the specifications in our panel B, columns 1 and 2. He also did not present any of the specifications in panel C.}
First consider panel A. In the first three columns, the OLS point estimates are much smaller in magnitude than the 2SLS point estimates. In column 4 the OLS and 2SLS estimates are quite similar. However, notice that the overidentification test easily rejects all specifications (columns 1--4, panels B and C). This suggests that the instruments may not be valid. There has in fact been substantial debate over the validity of these instruments. See Bresnahan's comment on Hausman1996.
To address this concern, we next estimate the falsification adaptive set. As discussed above, we only do this for panel C. We consider two different cut-offs for the definition of a weak instrument: $F \geq 10$ and $F \geq 5$ (see section (ref)). With the stricter cut-off, only one of the three 2SLS estimates is not weak, and thus the estimated FAS is a singleton. All the point estimates are larger in magnitude than the corresponding baseline estimates. The coefficients are also more stable across specifications. With the weaker cut-off, the FAS is an interval for columns 2--4. The length of this interval shrinks substantially as we move from column 2 to 3 to 4. Consider the FAS in column 4: $[-58.96, -17.97]$. All elements of this interval are larger in magnitude than the estimate $-16.23$ from the refuted baseline model. They are also larger than the OLS estimate of $-14.97$. In the logit model the price elasticities are proportional to the coefficient on price. So the lower bound on the FAS in column 4 suggests that the price elasticities could be as much as about three times larger than those suggested by the refuted baseline estimate. Thus ignoring the result of the overidentification test and using the baseline estimates anyway could lead to a large under-estimate of these elasticities. The falsification adaptive set is a constructive and informative tool researchers can use when their baseline model is rejected.
We conclude this section by briefly comparing our analysis with that of NevoRosen2012, who also study discrete choice demand models. They derive identified sets under a discrete relaxation of the classical instrument exogeneity and exclusion restrictions. They do not discuss the general problem of salvaging a falsified model. In particular, their identified sets can be empty. They do not discuss what researchers should do in this case. Furthermore, we relax instrument validity in a different and non-nested way. Hence our analysis is complementary to theirs. Using our data, we estimated the Nevo and Rosen identified set for the coefficient on price. Under their assumptions 3 and 4, and for the specification in column 4 of table (ref), this set is $(-\infty, -16.727]$. This set is unbounded from below and strictly contains the estimated FAS $[-58.96, -17.97]$. In their proposition 5 they discuss an additional assumption which can sometimes be used to obtain a bounded identified set. They only consider the two instrument case, and hence we cannot apply their result here. Moreover, this result requires choosing a sensitivity parameter $\gamma$ a priori. By definition, the FAS does not require choosing any sensitivity parameters.
Even when both instrument exogeneity and exclusion assumptions hold, the classical overidentifying restrictions (ref) may fail when the homogeneous treatment effects assumption fails. Hence we next consider models with heterogeneous treatment effects. In section (ref) we consider binary outcomes while in section (ref) we consider continuous outcomes. As earlier, we omit additional covariates for simplicity.
Unlike the baseline model in section (ref) using zero-correlation conditions, the baseline model we study here using statistical independence assumptions is falsifiable even with a single instrument. Thus we begin with the case where a single binary instrument is available. This allows us to explain the main ideas and results while keeping the notation simple. After this, we consider the general case where multiple discrete instruments are available.
Let $X \in \{ 0, 1 \}$ denote the observed treatment variable. Let $Y_1, Y_0 \in \{0,1\}$ denote binary potential outcomes. We observe
Let $Z \in \{0,1 \}$ denote an observed instrument. As mentioned above, we consider multiple discretely supported instruments later. Thus we observe the joint distribution of $(Y,X,Z)$. Consider the following assumption.
\begin{het1Assump}[Instrument exogeneity] $Z \mathbin{ \mathpalette{\@indep}{} } Y_1$ and $Z \mathbin{ \mathpalette{\@indep}{} } Y_0$. \end{het1Assump}
In this model, Manski1990 derived the identified set for $\ensuremath{\mathbb{P}}(Y_x=1)$ for $x \in \{0,1\}$ as well as the identified set for the average treatment effect \[ \text{ATE} = \ensuremath{\mathbb{P}}(Y_1=1) - \ensuremath{\mathbb{P}}(Y_0=1) \] under B(ref).\footnote{Manski's Manski1990 analysis considered a general case which does not require outcomes, treatment, or instruments to be binary. In this general setting, he used a mean independence assumption. When outcomes are binary, mean independence of $Y_x$ from $Z$ is equivalent to statistical independence of $Y_x$ and $Z$.} BalkePearl1997 showed these identified sets can be empty, and hence that this model is falsifiable. Consequently, researchers whose model is falsified face the same question we discussed earlier: What should they do? We answer this question using the four recommendations provided in section (ref). To do this, we must first define continuous relaxations of the instrument exogeneity assumption B(ref). We do this using the following concept from MastenPoirier2017.
When $c=0$, $c$-dependence is equivalent to $Z \mathbin{ \mathpalette{\@indep}{} } Y_x$. When $c \geq \max\{ \ensuremath{\mathbb{P}}(Z=1), \ensuremath{\mathbb{P}}(Z=0) \}$, $c$-dependence does not constrain the stochastic relationship between $Z$ and $Y_x$. $c \in (0,1)$ partially constrains the stochastic relationship between $Z$ and $Y_x$. Specifically, while it allows the probability of $Z=1$ given $Y_x=y_x$ to vary with $y_x$, it must be within a band around the unconditional probability: \[ \ensuremath{\mathbb{P}}(Z=1 \mid Y_x=y_x) \in \big[ \ensuremath{\mathbb{P}}(Z=1) - c, \ \ensuremath{\mathbb{P}}(Z=1) + c \big] \cap [0,1] \] for all $y_x \in \operatorname*{supp}(Y_x)$. MastenPoirier2017 give additional discussion of how to interpret $c$-dependence.
Thus we relax B(ref) as follows.
Under this relaxation, we will derive identified sets for various parameters of interest. We use these identified sets to characterize the falsification point and the falsification adaptive set. We can also use these identified sets to do sensitivity analysis beyond the falsification point.
Let $p_Z = \ensuremath{\mathbb{P}}(Z=1)$. The following assumption rules out trivial cases.
\begin{het1Assump} Suppose $p_Z \in (0,1)$. Suppose $\ensuremath{\mathbb{P}}(X=1 \mid Z=z) \in (0,1)$ for $z \in \{0,1\}$. \end{het1Assump}
First we set up some notation. Then we state and discuss the main theorem. For $z\in\{0,1\}$, define
Note that $k_z(0) =1$, $k_z(1) = 0$, and $k_z(\cdot)$ is weakly decreasing on $c \in [0,1]$. Using this function, define the set
which depends only on $c$ and the marginal distribution of $Z$. For $x \in \{0,1\}$, define
which depends on the joint distribution of $(Y,X,Z)$. Let \[ \Theta_x(c) = \mathcal{D}(c) \cap \mathcal{H}_x \] denote the intersection of these two sets.
The identified sets in theorem (ref) are defined via two kinds of constraints: the set $\mathcal{D}(c)$ and the sets $\mathcal{H}_x$ for $x\in\{0,1\}$. The set $\mathcal{H}_x$ is simply the identified set for \[ \big( \ensuremath{\mathbb{P}}(Y_x=1 \mid Z=0), \ \ensuremath{\mathbb{P}}(Y_x=1 \mid Z=1) \big) \] without imposing any assumptions on the stochastic relationship between $Y_x$ and $Z$, as derived by Manski1990. This set can be thought of as the intersection of two pairs of parallel half-spaces. The shaded boxes in figure (ref) show one example of this set.
The set $\mathcal{D}(c)$ imposes the $c$-dependence constraint. For example, if $c=0$, it imposes that \[ \ensuremath{\mathbb{P}}(Y_x=1 \mid Z=0) = \ensuremath{\mathbb{P}}(Y_x=1 \mid Z=1) \] for $x \in \{0,1\}$. This is simply the statistical independence assumption B(ref) studied by Manski1990. Thus the identified sets in theorem (ref) simplify to those derived by Manski when $c = 0$. As $c$ gets larger, we no longer require the values $\ensuremath{\mathbb{P}}(Y_x=1 \mid Z=z)$ for $z \in \{0,1\}$ to exactly equal each other, but we do require them to be sufficiently close, as defined by the set $\mathcal{D}(c)$. Specifically, it constrains the pairs $(\ensuremath{\mathbb{P}}(Y_x=1 \mid Z=0), \ensuremath{\mathbb{P}}(Y_x=1 \mid Z=1))$ to lie inside a diamond-shaped parallelogram, as shown in figure (ref). As $c$ gets larger, this parallelogram gets larger, eventually approaching the entire square $[0,1]^2$ as $c \rightarrow \max \{ p_Z, 1- p_Z \}$. Thus, for $c \geq \max \{ p_Z, 1 - p_Z \}$, theorem (ref) simplifies to the no assumption bounds derived by Manski1990. Thus theorem (ref) continuously connects the no assumption bounds with the bounds under the statistical independence assumption B(ref).
While the no assumption bounds ($c \geq \max \{p_Z, 1-p_Z \}$) are never empty, the bounds under statistical independence ($c=0$) can be empty, and hence the baseline statistical independence assumption can be falsified. This happens when, for some $x \in \{0,1\}$, the no assumption bounds $\mathcal{H}_x$ have an empty intersection with the statistical independence constraint set $\mathcal{D}(0)$. Graphically, this happens when the box defined by the no assumption bounds does not intersect the 45-degree line. This is shown in figure (ref). The falsification point is simply the smallest $c$ such that the parallelogram defined by $\mathcal{D}(c)$ has a nonempty intersection with the no assumption bounds $\mathcal{H}_x$ for each $x \in \{0,1\}$. This intersection is illustrated in the middle plot of figure (ref).
We next state some useful properties of the identified set given in theorem (ref).
The value $c^*$ is the falsification point. In the simple model with a single binary instrument it is possible to give a closed form expression for the falsification point, but we omit it for brevity. For part 2, recall that a polytope in $[0,1]^4$ is a set that can be written as the intersection of finitely many half-spaces. We use continuity of the correspondence $\phi$ below to show that bounds on functionals like ATE are continuous in $c$.
We next show how to use theorem (ref) to get identified sets for the counterfactual probabilities $\ensuremath{\mathbb{P}}(Y_x=1)$ and for the average treatment effect. By the law of total probability, \[ \ensuremath{\mathbb{P}}(Y_x=1) = \ensuremath{\mathbb{P}}(Y_x=1 \mid Z=0) \ensuremath{\mathbb{P}}(Z=0) + \ensuremath{\mathbb{P}}(Y_x=1 \mid Z=1) \ensuremath{\mathbb{P}}(Z=1). \] The weights $\ensuremath{\mathbb{P}}(Z=0)$ and $\ensuremath{\mathbb{P}}(Z=1)$ are point identified. The joint identified set for $\ensuremath{\mathbb{P}}(Y_x=1 \mid Z=0)$ and $\ensuremath{\mathbb{P}}(Y_x=1 \mid Z=1)$ is given by $\Theta_x(c)$. Thus we can simply minimize and maximize the above convex combination over this set to obtain the identified set for $\ensuremath{\mathbb{P}}(Y_x=1)$. Hence we define \[ \overline{P}_x(c) = \sup_{(a_0,a_1) \in \Theta_x(c)} \big( a_0 \ensuremath{\mathbb{P}}(Z=0) + a_1 \ensuremath{\mathbb{P}}(Z=1) \big) \] and \[ \underline{P}_x(c) = \inf_{(a_0,a_1) \in \Theta_x(c)} \big( a_0 \ensuremath{\mathbb{P}}(Z=0) + a_1 \ensuremath{\mathbb{P}}(Z=1) \big). \] These are both finite dimensional linear programs and hence can be computed easily.
The falsification adaptive set for $\ensuremath{\mathbb{P}}(Y_x=1)$ is $[\underline{P}_x(c^*),\overline{P}_x(c^*)]$. This will be a singleton for one $x \in \{0,1\}$ and will typically be an interval with a nonempty interior for the other value of $x$. This can be seen in figure (ref). In the middle plot, $\ensuremath{\mathbb{P}}(Y_x=1)$ is point identified since $\Theta_x(c^*)$ is a singleton. There is an analogous plot, however, for $\ensuremath{\mathbb{P}}(Y_{1-x}=1 \mid Z=z)$. In this plot, the box $\mathcal{H}_x$ will typically not also have a singleton intersection with the diamond $\mathcal{D}(c^*)$. Thus $\ensuremath{\mathbb{P}}(Y_{1-x}=1)$ will typically be partially identified. This discussion implies that ATE will typically be partially identified at the falsification point. That is: The falsification adaptive set for ATE, $[\underline{\text{ATE}}(c^*), \overline{\text{ATE}}(c^*)]$, will generally be an interval with a nonempty interior.
We next consider the case where there are multiple discrete instruments. We relax each instrument exogeneity condition separately. We derive identified sets for a given relaxation of the instrument exogeneity conditions. As before, this allows us to derive the falsification frontier and the falsification adaptive set.
Suppose we observe $L \geq 1$ instruments, $Z = (Z_1,\ldots,Z_L) \in \ensuremath{\mathbb{R}}^L$. Let $\operatorname*{supp}(Z_\ell) = \{ z_\ell^1, \ldots, z_\ell^{J_\ell} \}$ for some constant $J_\ell$. For notational simplicity we only consider the case $J_\ell = J$ for all $\ell \in \{1,\ldots,L \}$. As in theorem (ref), we will derive the joint identified set for the probabilities \[ \ensuremath{\mathbb{P}}(Y_x=1 \mid Z_\ell = z_\ell^j) \] for $x \in \{0,1\}$, $\ell \in \{1,\ldots,L\}$, and $j \in \{1,\ldots,J \}$. For each $x \in \{0,1\}$, there are $LJ$ conditional probabilities. As in the single binary instrument case, we make the following assumption to rule out trivial cases.
Let $Z_{-\ell}$ denote the vector of instruments without $Z_\ell$. $Z$ has $J^L$ support points while $Z_{-\ell}$ has $J^{L-1}$ support points. Let $\operatorname*{supp}(Z_{-\ell}) = \{ z_{-\ell}^1,\ldots, z_{-\ell}^{J^{L-1}} \}$. By the law of total probability,
The terms \[ \ensuremath{\mathbb{P}}(Y = 1,X=x \mid Z_\ell = z_\ell^j) \qquad \text{and} \qquad \frac{\ensuremath{\mathbb{P}}(X=1-x, Z_\ell = z_\ell^j, Z_{-\ell} = z_{-\ell}^k)}{\ensuremath{\mathbb{P}}(Z_\ell = z_\ell^j)} \] are observed directly in the data. In contrast, the probabilities \[ \ensuremath{\mathbb{P}}(Y_x = 1 \mid X=1-x, Z_\ell = z_\ell^j, Z_{-\ell} =z_{-\ell}^k) \] are not observed. We do know that they must lie in the interval $[0,1]$, however. For each $x \in \{0,1\}$, there are $J^L$ of these unknown probabilities. Thus for each $x \in \{0,1\}$ the $LJ$ probabilities $\ensuremath{\mathbb{P}}(Y_x=1 \mid Z_\ell = z_\ell^j)$ are determined by a system of $LJ$ known linear functions of $J^L$ unknowns which are all bounded between 0 and 1. Denote this set by
where $\textbf{b}_x \in \ensuremath{\mathbb{R}}^{LJ}$ is defined by \[ [\textbf{b}_x]_{(\ell-1)J + j} = \ensuremath{\mathbb{P}}(Y=1,X=x \mid Z_\ell = z_\ell^j) \] for $\ell \in\{1,\ldots,L\}$, $j\in \{1,\ldots,J\}$ and $\textbf{A}_x \in \ensuremath{\mathbb{R}}^{LJ \times J^{L}}$ is defined by
for $\ell \in\{1,\ldots,L\}$, $j\in \{1,\ldots,J\}$, $m \in \{1,\ldots J^{L}\}$. Here we use $z^m$ to denote the $m$th element of the support of $Z$.
Although this notation is a bit complicated, it is simply the vector description of equation (ref) above: The elements of $\textbf{b}_x$ are the intercepts, while the non-zero elements of the matrix $\textbf{A}_x$ are elements of the form \[ \ensuremath{\mathbb{P}}(X=1-x, Z_{-\ell} = z_{-\ell}^k \mid Z_\ell = z_\ell^j) \] for some $\ell \in \{1,\ldots,L\}$, $j \in \{1,\ldots,J\}$, and $k \in \{1,\ldots, J^{L-1} \}$.
As in the single instrument case, this set $\mathcal{H}_x$ is the identified set for the joint vector of $\ensuremath{\mathbb{P}}(Y_x=1 \mid Z_\ell = z_\ell^j)$ probabilities with no assumption on the dependence between $Y_x$ and each instrument $Z_\ell$. We next consider constraints on this dependence. As before, we relax independence using $c$-dependence. Specifically, we make the following assumption.
This assumption generalizes $c$-dependence to allow for multiple discrete instruments. As in the single binary instrument case, $c_\ell = 0$ implies $Y_x \mathbin{ \mathpalette{\@indep}{} } Z_\ell$, while $c_\ell > 0$ allows for some dependence. Importantly, we allow $c_\ell$ to vary across instruments. Let $c = (c_1,\ldots,c_L)$.
As shown in theorem (ref) below, the set \[ \mathcal{D}(c) = \prod_{\ell = 1}^L \mathcal{D}_\ell(c_\ell) \] where
characterizes the constraints that B(ref)$^{\prime \prime}$ imposes on the probabilities $\ensuremath{\mathbb{P}}(Y_x=1 \mid Z_\ell = z_\ell^j)$. This set $\mathcal{D}(c)$ generalizes the parallelogram defined by equation (ref). As in the single binary instrument case, it is defined by a finite number of linear inequalities.
Let \[ \Theta_x(c) = \mathcal{D}(c) \cap \mathcal{H}_x \] denote the intersection of these generalized diamond and box sets.
This result is a direct generalization of theorem (ref). All of the discussion and interpretation there applies here as well. One new issue arises: The baseline model, $c = (0,\ldots,0)$, is falsifiable even if none of the instruments are falsifiable themselves. That is, consider the model that imposes $Y_x \mathbin{ \mathpalette{\@indep}{} } Z_\ell$ but no constraints on the statistical dependence between the other instruments $Z_{-\ell}$ and $Y_x$. We can obtain this model by setting $c = \iota - e_\ell$ where $\iota$ is a vector of 1's and $e_\ell$ is the unit vector. Suppose this model is not falsified, regardless of which instrument $\ell$ we choose to impose exogeneity on. Despite this, it is nonetheless possible that the baseline model $c = (0,\ldots,0)$ is falsified. In this case the instruments are incompatible with each other.
As in proposition (ref) in the single binary instrument case, we next state some useful properties of the identified set given in theorem (ref).
The set $\mathcal{C}$ is the set of all instrument partial exogeneity assumptions which do not refute the model. The complement of this set is the set of all instrument partial exogeneity assumptions which are refuted. The falsification frontier in this case is defined as follows. Let \[ S_{< c} = \{ \tilde{c} \in [0,1]^L : \tilde{c} < c \} \] be the set of points which are smaller than $c$. Then
This is the set of all points $c$ such that any points smaller than $c$ are not in $\mathcal{C}$. This frontier is identified from the data since $\mathcal{C}$ is defined implicitly by $\{ \Theta_0(c) \times \Theta_1(c): c\in [0,1]^L\}$, which is point identified.
We next show how to use theorem (ref) to get identified sets for counterfactual probabilities $\ensuremath{\mathbb{P}}(Y_x=1)$ and for the average treatment effect. Define \[ \overline{P}_x(c) = \sup_{\textbf{a} \in \Theta_x(c)} \; \sum_{j = 1}^J \sum_{\ell = 1}^L a_{(\ell-1)J + j}\ensuremath{\mathbb{P}}(Z_\ell = z_\ell^j) \] and \[ \underline{P}_x(c) = \inf_{\textbf{a} \in \Theta_x(c)} \; \sum_{j = 1}^J \sum_{\ell = 1}^L a_{(\ell-1)J + j}\ensuremath{\mathbb{P}}(Z_\ell = z_\ell^j). \] As in the single binary instrument case, these are both finite dimensional linear programs and hence can be computed easily.
In the single binary instrument case, there is a simple closed form expression for the falsification point $c^*$. Consequently, computing the falsification adaptive set for ATE can be done easily by using linear programming to compute $\overline{P}_x(c^*)$ and $\underline{P}_x(c^*)$. In the multiple discrete instrument case, there is not a simple closed form expression for the falsification frontier. Nonetheless, it can be computed by a sequence of convex optimization problems. Specifically, the sets $\mathcal{D}(c)$ and $\mathcal{H}_x$ are polytopes in $\ensuremath{\mathbb{R}}^{LJ}$. Define the `distance' between these polytopes as \[ \text{dist}(\mathcal{D}(c), \mathcal{H}_x) = \inf_{(v_1,v_2) \in \mathcal{D}(c) \times \mathcal{H}_x} \; \| v_1 - v_2 \|. \] The model is refuted if and only if $\text{dist}(\mathcal{D}(c), \mathcal{H}_x) > 0$ for some $x \in \{0,1 \}$. For any $c$, the distance between these polytopes can be computed via convex optimization. Thus one can do a grid search over $c$ and collect all values where the distance is zero for both $x \in \{0,1\}$, and hence the model is not refuted. This is a numerical approximation to the set $\mathcal{C}$. We can then use this set to approximate the falsification frontier.
Given this approximate falsification frontier, we can compute the falsification adaptive set for ATE as follows. For any $c$, we can compute $\overline{P}_x(c)$ and $\underline{P}_x(c)$ via linear programming, since the objective function is linear and the constraint set is a polytope in $\ensuremath{\mathbb{R}}^{LJ}$. Thus we can simply compute these bounds for all $c \in \text{FF}$. The union of those bounds is the falsification adaptive set. Identified sets for ATE at values of $c$ beyond the falsification frontier can be computed similarly.
In section (ref) we considered binary outcomes, binary treatment, and multiple discrete instruments. In this case, the joint distribution of the data is characterized by a finite dimensional vector. It is well known that with discretely distributed data identified sets can often be computed using linear programming; see our literature review in appendix (ref). There are two advantages, however, to deriving analytical results like those of theorems (ref) and (ref). First, they can explain the role of identifying assumptions, as illustrated in figures (ref) and (ref). Second, they allow us to generalize the results to cases where brute force computation is costly or impossible. In this section, we show that the analytical results we derived under binary outcomes generalize to continuous outcomes. This leads us to a relatively simple and feasible approach for computing identified sets, falsification frontiers, and the falsification adaptive set under relaxations of instrument exogeneity with continuous outcomes.
We begin by assuming that outcomes are continuously distributed.
\begin{het2Assump} For any $x, x' \in \{0,1\}$ and $z \in \{0,1\}^L$, the distribution of $Y_x \mid X=x', Z=z$ is continuous with respect to the Lebesgue measure. \end{het2Assump}
Assumption C(ref) supposes that, conditional on the treatment and instruments, potential outcomes are continuously distributed. It implies that, conditional on the treatment and instruments, observed outcomes are also continuously distributed. We also suppose there are $L$ binary instruments $Z = (Z_1,\ldots,Z_L)$. We can allow for discrete instruments as in section (ref), but we only consider binary instruments to simplify the notation.
Let $\mathcal{P}(A)$ denote the set of densities (with respect to the Lebesgue measure) with support contained in $A$. This set can be written as \[ \mathcal{P}(A) = \left\{ p: \int_A p(y) \; dy = 1, \ p(y) \geq 0 \text{ for all } y \in \ensuremath{\mathbb{R}} \right\}. \] We begin by considering the baseline case where the instruments are exogenous. When there is just a single instrument, this setting was studied in Kitagawa2009. The following result is essentially his proposition 3.1.
The assumption that $\operatorname*{supp}(Y_x) \subseteq \mathcal{Y}_x$ is not restrictive since we can let $\mathcal{Y}_x = \ensuremath{\mathbb{R}}$, but this notation allows us to impose bounds on the support if they are known.
We generalize proposition (ref) in two ways. First, we consider an arbitrary number $L$ of binary instruments. Second, we relax instrument exogeneity B(ref) to instrument partial exogeneity B(ref)$^{\prime \prime}$. As discussed above, our proof strategy generalizes the result in section (ref). Independence of each instrument with each potential outcome requires that \[ f_{Y_x|Z_\ell}(y \mid 1) = f_{Y_x|Z_\ell}(y \mid 0) \] for all $y \in \ensuremath{\mathbb{R}}$ and $\ell \in \{1, \ldots, L \}$. This constraint is analogous to the diagonal constraint in the left plots of figures (ref) and (ref). The $c$-dependence assumption B(ref)$^{\prime \prime}$ does not require the densities $f_{Y_x | Z_\ell}(\cdot \mid z)$ to be the same for all $z \in \{0,1\}$, but it bounds the variation across $z$ by a magnitude determined by the vector $c$ of sensitivity parameters. We make this constraint precise below.
Analogously to section (ref), we will derive the joint identified set for the conditional densities \[ \left\{f_{Y_x|Z_\ell}(\cdot \mid z): x \in \{0,1\}, z \in\{0,1\}, \ell \in \{ 1, \ldots, L \} \right\} \] under instrument partial exogeneity. We define the following notation for these densities. For all $\ell \in \{1, \ldots, L\}$ and $x\in\{0,1\}$, let \[ \mathbf{f}_{Y_x|Z_\ell} =
\in \mathcal{P}(\operatorname*{supp}(Y_x \mid Z_\ell = 0))\times \mathcal{P}(\operatorname*{supp}(Y_x \mid Z_\ell = 1)) \] and \[ (\mathbf{f}_{Y_x|Z_\ell})_{all \ell} =
\in \prod_{\ell=1}^L \Big( \mathcal{P}(\operatorname*{supp}(Y_x \mid Z_\ell = 0))\times \mathcal{P}(\operatorname*{supp}(Y_x \mid Z_\ell = 1))\Big). \] Thus we will derive the joint identified set for the vectors $(\mathbf{f}_{Y_0|Z_\ell})_{\text{all } \ell}$ and $(\mathbf{f}_{Y_1|Z_\ell})_{\text{all } \ell}$.
This identified set has a structure analogous to that in section (ref). Thus we first define the continuous outcome generalization of the set $\mathcal{D}(c)$ defined in equations (ref) and (ref). To do this, define \[ k_z^\ell(c_\ell) = \frac{\ensuremath{\mathbb{P}}(Z_\ell=z)\max\{\ensuremath{\mathbb{P}}(Z_\ell=1-z) - c_\ell,0\}}{\ensuremath{\mathbb{P}}(Z_\ell = 1-z)\min\{\ensuremath{\mathbb{P}}(Z_\ell=z)+c_\ell,1\}} \] for $\ell \in \{1, \ldots, L \}$ and $z \in \{0,1 \}$. This function generalizes equation (ref). For each $x \in \{0,1\}$, let $\mathcal{Y}_x$ be a known set. We assume this set contains the support of $Y_x \mid X,Z$ for all values of the conditioning variables. Define
for $\ell \in \{1, \ldots, L \}$ and $x \in \{0,1\}$. This set generalizes equations (ref) and (ref). It depends on $x \in \{0,1\}$ only via the possibility that $\mathcal{Y}_x$ might depend on $x$. Finally, take the product to get \[ \mathcal{D}_x(c) = \prod_{\ell = 1}^L \mathcal{D}_{x,\ell}(c_\ell) \] We will show that $\mathcal{D}_x(c)$ is the set of densities $(\mathbf{f}_{Y_x|Z_\ell})_{\text{all } \ell}$ consistent with $c$-dependence (B(ref)$^{\prime \prime}$). Note that setting $c_\ell =0$ implies that $f_{Y_x|Z_\ell} = f_{Y_x}$, as required by the baseline model.
Next we generalize the set $\mathcal{H}_x$ defined in equations (ref) and (ref). With continuous outcomes this will be the set of proper densities $(\mathbf{f}_{Y_x|Z_\ell})_{\text{all } \ell}$ consistent with the observed data. Let $\mathbf{A}_x$ be the $2L \times 2^L$ matrix such that for $\ell \in \{1, \ldots, L\}$ and $m \in \{ 1,\ldots,2^L \}$: \[ [\mathbf{A}_x]_{(\ell-1)2+1,m} = \frac{\ensuremath{\mathbb{P}}(X=1-x, Z = z^m)}{\ensuremath{\mathbb{P}}(Z_\ell = 0)} \ensuremath{\mathbbm{1}} \big( [z^m]_\ell = 0 \big) \] and \[ [\mathbf{A}_x]_{(\ell-1)2+2,m} = \frac{\ensuremath{\mathbb{P}}(X=1-x, Z = z^m)}{\ensuremath{\mathbb{P}}(Z_\ell = 1)} \ensuremath{\mathbbm{1}} \big( [z^m]_\ell = 1 \big). \] Here we use $z^m$ to denote the $m$th element of the support of $Z$. Let \[ f_{Y,X|Z_\ell}(\cdot,x) =
\qquad and \qquad (\mathbf{f}_{Y,X|Z_\ell}(\cdot,x) )_{all \ell} =
. \] These are vectors of observed densities of outcomes and treatment conditional on different instruments. Define \[ \mathcal{H}_x = \left\{ \mathbf{f}_x \in \mathcal{P}(\mathcal{Y}_x)^{2L} : \mathbf{f}_x = (\mathbf{f}_{Y,X|Z_\ell}(\cdot,x))_{\text{all } \ell} + \mathbf{A}_x \mathbf{q} \ \text{ for some } \mathbf{q} \in \mathcal{P}(\mathcal{Y}_x)^{2^L} \right\}. \] This set generalizes equation (ref). Finally, let \[ \Theta_x(c) = \mathcal{D}_x(c) \cap \mathcal{H}_x. \]
As in proposition (ref), knowledge of the set $\mathcal{Y}_x$ is not restrictive since we can let $\mathcal{Y}_x = \ensuremath{\mathbb{R}}$, but this notation allows us to impose bounds on its support if they are known. We could also generalize this set to allow it to vary with $x'$ and $z^m$, at the cost of additional notation.
The identified set $\Theta_0(c) \times \Theta_1(c)$ is a convex set of functions defined by linear equality and inequality constraints. It weakly expands as components of $c$ increase. It nests the identified set under the baseline independence assumption ($c=0$) and the identified set under no assumptions on the dependence between potential outcomes and instruments ($c = \iota_L$, an $L$-vector of ones).
For a given $c$, this identified set can be empty, which implies the model at that $c$ is falsified. Define \[ \mathcal{C} = \{c \in [0,1]^L: \Theta_0(c) \times \Theta_1(c) \neq \emptyset\}. \] This is the set of all partial exogeneity assumptions which are not refuted. It has the property that $c \in \mathcal{C}$ and $c' \geq c$ implies $c'\in\mathcal{C}$. Using this set, we can compute the falsification frontier as in equation (ref).
As in section (ref), we can obtain the identified set for functionals of the densities $f_{Y_x \mid Z_\ell}$ via convex optimization. For example, the identified set for $(f_{Y_0}(\cdot), f_{Y_1}(\cdot))$ is $I_0(c) \times I_1(c)$ where
Given this identified set, suppose we are interested in a linear functional \[ \Gamma(f_1,f_0) = \int_{\mathcal{Y}_1} \omega_1(y_1) f_1(y_1) \; dy_1 + \int_{\mathcal{Y}_0} \omega_0(y_0) f_0(y_0) \; dy_0 \] where $\omega_1$ and $\omega_0$ are known weight functions. For example, with $\omega_1(y) = y$ and $\omega_0(y) = -y$, $\Gamma(f_1,f_0)$ equals the average treatment effect. Let \[ \overline{\Gamma}(c) = \sup_{f_1 \in I_1(c)} \int_{\mathcal{Y}_1} \omega_1(y_1)f_1(y_1) \; dy_1 + \sup_{f_0 \in I_0(c)}\int_{\mathcal{Y}_0} \omega_0(y_0)f_0(y_0) \; dy_0 \] and \[ \underline{\Gamma}(c) = \inf_{f_1 \in I_1(c)} \int_{\mathcal{Y}_1} \omega_1(y_1)f_1(y_1) \; dy_1 + \inf_{f_0 \in I_0(c)}\int_{\mathcal{Y}_0} \omega_0(y_0)f_0(y_0) \; dy_0. \] Then $[\underline{\Gamma}(c), \overline{\Gamma}(c)]$ is the closure of the identified set for $\Gamma(f_{Y_1},f_{Y_0})$. This follows by convexity of the identified set in theorem (ref). These bounds are weakly increasing in each component of $c$. We discuss computation next.
The identified set $\Theta_0(c) \times \Theta_1(c)$ is infinite dimensional, since it is a set of nonparametric densities. Suppose we are interested in linear functionals $\Gamma(f_1,f_0)$. Even though the corresponding identified set---when it is nonempty---is an interval, we have characterized it via optimization over the infinite dimensional spaces $\Theta_x(c)$. Here we discuss one approach to computing these identified sets in practice. When the baseline model is falsified, this will also allow us to compute approximations to the falsification frontier and the falsification adaptive set. Specifically, we will approximate the infinite dimensional space of densities of interest by a finite dimensional sieve space. Similar approximations of identified sets have been used in MogstadSantosTorgovitsky2016, for example. Alternatively, it is likely that the computational approach developed in ChristensenConnault2019 could be adapted to our setting. Unlike the sieve based approach we consider below, the dimension of their optimization problem does not depend on the precision of the density approximation. We leave the application of their approach to our problem to future work.
For simplicity, let $\mathcal{Y}_x = [0,1]$ for $x \in \{0,1\}$. This restriction can be relaxed by transforming the outcome variable to have support in the unit interval. We also assume that for any $x, x' \in \{0,1\}$ and $z \in \{0,1\}^L$ the density $f_{Y_x|X,Z}(\cdot \mid x',z)$ is continuous on $[0,1]$. Let $\widetilde{\mathcal{P}}([0,1])$ denote this set of continuous pdfs. Let
where $\mathcal{A}_M \subseteq \ensuremath{\mathbb{R}}^M$ and $\{ b_m(\cdot) \}$ are known basis functions. We assume that every element of $\widetilde{\mathcal{P}}([0,1])$ can be approximated by a sequence of elements in $\mathcal{F}_{M}$ as $M\rightarrow\infty$. For example, $\mathcal{F}^M$ could be a set of Bernstein polynomials, which can uniformly approximate any continuous function on $[0,1]$. For Bernstein polynomials, \[ \mathcal{A}_M = \left\{ \textbf{a}^M : a_m \geq 0, \sum_{m=0}^M a_m = 1\right\}. \] Hence $\mathcal{A}_M$ is a closed, bounded, and convex subset of $\ensuremath{\mathbb{R}}^M$.
To approximate $\Theta_x(c) = \mathcal{D}_x(c) \cap \mathcal{H}_x$, we approximate each of the sets $\mathcal{D}_x(c)$ and $\mathcal{H}_x$. We first consider a finite dimensional approximation to the set $\mathcal{D}_x(c)$. Since $\mathcal{Y}_x = [0,1]$ does not depend on $x$, we let $\mathcal{D}(c) = \mathcal{D}_x(c)$. Let
Define the set
Then \[ \mathcal{D}_\ell^M(c_\ell) = \left\{(f(\cdot \mid 0),f(\cdot \mid 1)) = \left( \textbf{a}^M(0)'\textbf{b}^M(\cdot), \textbf{a}^M(1)'\textbf{b}^M(\cdot)\right): \left( \textbf{a}^M(0),\textbf{a}^M(1) \right) \in \mathcal{A}^M_\ell(c_\ell)\right\}. \] The set $\mathcal{A}_\ell^M(c_\ell)$ is closed and convex since it is defined by a finite number of weak inequalities. There are a continuum of inequalities, however. In practice, we also approximate $\mathcal{A}_\ell^M(c_\ell)$ as follows. Pick a grid of points $\{y_1,\ldots,y_N\} \subseteq [0,1]$ that becomes dense in $[0,1]$ as $N \rightarrow \infty$. Then let $\mathcal{A}_\ell^{M,N}(c_\ell)$ be the set of coefficients $\textbf{a}^M(0)$ and $\textbf{a}^M(1)$ satisfying \[ \left( \textbf{a}^M(0)k_0^\ell(c_\ell) - \textbf{a}^M(1) \right)' \textbf{b}^M(y_n)\leq 0 \] and \[ \left( \textbf{a}^M(1)k_1^\ell(c_\ell) - \textbf{a}^M(0) \right)' \textbf{b}^M(y_n)\leq 0 \] for all $n \in \{1 ,\ldots,N \}$. This is a finite set of linear inequalities. We keep the $N$ implicit in the notation from here on.
Approximate the overall set $\mathcal{D}(c)$ by the product \[ \mathcal{D}^M(c) = \prod_{\ell = 1}^L \mathcal{D}^M_\ell(c_\ell). \] This set is characterized by \[ \mathcal{A}^M(c) = \prod_{\ell = 1}^L \mathcal{A}^M_\ell(c_\ell), \] a closed, convex, and bounded set in $\ensuremath{\mathbb{R}}^{2ML}$.
We similarly approximate $\mathcal{H}_x$ by a finite dimensional set. Let \[ \mathbf{f}_{Y|X,Z_\ell}(\cdot \mid x) =
\qquad and \qquad (\mathbf{f}_{Y|X,Z}(\cdot \mid x) )_{all \ell} =
. \] These conditional pdfs are point identified directly from the data. For each $z \in \{0,1\}$, we approximate $f_{Y|X,Z_\ell}(\cdot \mid x,z)$ by
where $\{ b_m(\cdot) \}$ are known basis functions. When these basis functions are orthonormal, so that \[ \int_0^1 b_m(y) b_n(y) \; dy= \ensuremath{\mathbbm{1}}(m=n), \] this approximation can be obtained choosing \[ \xi_\ell(x,z)_m = \int_0^1 f_{Y|X,Z_\ell}(y \mid x,z)b_m(y) \; dy. \] More concretely, if we use a Bernstein polynomial approximation, then \[ \xi_\ell(x,z)_m = f_{Y|X,Z_\ell} \left( \frac{m}{M} \mid x,z \right). \] Here we have approximated the conditional distribution of $Y \mid X,Z$. But $\mathcal{H}_x$ is defined in terms of distributions of $(Y,X) \mid Z$. We can write the vector $(\textbf{f}_{Y,X|Z}(y,x))_{\text{all } \ell}$ as \[ (\textbf{f}_{Y,X|Z}(y,x))_{\text{all } \ell} = \textbf{B}_x (\textbf{f}_{Y|X,Z}(y \mid x))_{\text{all } \ell} \] where \[ \mathbf{B}_x =
. \] Hence its approximation can be written as \[ (\textbf{f}_{Y,X|Z}(y,x))_{\text{all } \ell} = \textbf{B}_x \Xi^M \textbf{b}^M(y) \] where \[ \Xi^M =
. \] Next, approximate $\mathcal{P}(\mathcal{Y}_x)^{2^L}$ by \[ \left\{ \textbf{f}(\cdot) = \Gamma \textbf{b}^M(\cdot): \gamma_1^M, \ldots, \gamma_{2^L}^M \in \mathcal{A}^M\right\} \] where \[ \Gamma =
\] is a $2^L \times M$ matrix of coefficients. Thus we approximate $\mathcal{H}_x$ by
If $\mathcal{A}_M$ is closed, convex, and bounded, the affine mapping \[ \textbf{B}_x \Xi^M + \textbf{A}_x \mathcal{A}_M^{2^L} \] will also be closed, convex, and bounded.
Thus we approximate the identified set for $((\mathbf{f}_{Y_0|Z_\ell})_{\text{all } \ell}, (\mathbf{f}_{Y_1|Z_\ell})_{\text{all } \ell})$ by \[ \Theta_0^M(c) \times \Theta_1^M(c) = \big( \mathcal{D}^M(c) \cap \mathcal{H}^M_0 \big) \times \big( \mathcal{D}^M(c) \cap \mathcal{H}^M_1 \big) \] where \[ \mathcal{D}^M(c) \cap \mathcal{H}^M_x = \left\{ \textbf{f}(y) = \textbf{E}^M \textbf{b}^M(y): \textbf{E}^M \in \mathcal{A}^M(c) \cap (\textbf{B}_x \Xi^M + \textbf{A}_x \mathcal{A}_M^{2^L}) \right\}. \] This approximate identified set is empty---and hence the approximate model is falsified---if the set \[ \mathcal{A}^M(c) \cap (\textbf{B}_x \Xi^M + \textbf{A}_x \mathcal{A}_M^{2^L}) \] is empty for either $x=0$ or $x=1$. These two sets are finite dimensional polytopes. Hence it is computationally tractable to compute the distance between them, \[ \text{dist}(\mathcal{A}^M(c), \ \textbf{B}_x \Xi^M + \textbf{A}_x \mathcal{A}_M^{2^L}) = \inf_{(\textbf{v}_1, \textbf{v}_2) \in \mathcal{A}^M(c) \times (\textbf{B}_x \Xi^M + \textbf{A}_x \mathcal{A}_M^{2^L}) }\|\textbf{v}_1 - \textbf{v}_2 \|, \] for any given $c$. Since they are closed, convex, and bounded, the approximate model will be non-refuted if and only if \[ \text{dist}(\mathcal{A}^M(c), \ \textbf{B}_x \Xi^M + \textbf{A}_x \mathcal{A}_M^{2^L}) = 0. \] This characterization can be used to compute an approximate falsification frontier as well as an approximate falsification adaptive set.
Although we omit a full analysis, we expect the identified sets $\Theta_0^M(c) \times \Theta_1^M(c)$ will converge to $\Theta_0(c) \times \Theta_1(c)$ as $M \rightarrow \infty$ under suitable regularity conditions. Consequently, the approximate falsification frontier and falsification adaptive sets should also converge as $M \rightarrow \infty$. Finally, the value of $M$ can be chosen large enough so that the approximate objects of interest change by less than a preset tolerance for further increases in $M$.
In this paper we outlined a systematic and constructive answer to the question “What should researchers do when their baseline model is refuted?” Our answer focuses on what can be learned from falsified models, rather than treating falsification as a nuisance to be ignored or as a fatal flaw which dooms a study.
We gave four recommendations. First: Measure the extent of falsification. Do this by defining continuous relaxations of the key identifying assumptions of interest, and then relaxing these assumptions until the model is no longer falsified. We call the set of points at which this happens the falsification frontier. Second: Present the falsification adaptive set, the identified set for the parameter of interest under the assumption that the true model lies somewhere on the falsification frontier. This second recommendation is a generalization of standard practice for non-refuted baseline models. Moreover, it does not require the researcher to select or calibrate sensitivity parameters. Third: Present the identified set for selected points of interest on the falsification frontier. Fourth: Present identified sets for points beyond the falsification frontier, as a further sensitivity analysis.
We illustrated these four recommendations in two different overidentified instrumental variable models. The first model imposes homogeneous treatment effects but allows for continuous treatments, while the second model allows for heterogeneous treatment effects but focuses on binary treatments. In both models, multiple instruments are observed and the key identifying assumptions are exclusion and exogeneity of each instrument. We considered continuous relaxations of these assumptions. We then characterized the falsification frontier, the falsification adaptive set, and identified sets for points beyond the frontier. In the homogeneous treatment effect model, the falsification adaptive set has a particularly simple closed form expression, depending only on the value of a handful of 2SLS regression coefficients. In the heterogeneous treatment effect model, the falsification adaptive set can be quickly computed using convex optimization.
We showed how to use our results in four substantively different empirical applications. There we emphasized that that falsification adaptive set is an informative complement to traditional overidentification test $p$-values: It directly summarizes the range of estimates corresponding to non-falsified alternative models. We also emphasized the importance of controlling for the possibly invalid instruments when considering alternative models.
In this paper we outlined our general approach and then applied it to the classical linear instrumental variable model. The falsification frontier and falsification adaptive sets, however, must generally be derived separately for each baseline model and class of relaxations from those baseline assumptions. Thus in future work we plan to do this kind of analysis for different baseline models and different classes of relaxations. Here we focused on instrumental variable models. Other possible applications include various kinds of placebo tests, like those discussed in HeckmanHotz1989. It may also be possible to derive results for a general class of models, by applying the results of ChesherRosen2017 or Torgovitsky2018, for example. Such extensions are important since---just like the choice of the baseline assumptions---the choice of relaxation matters. Hence comparing empirical results for different relaxations will be informative.
Finally, in this paper we focused on population level analysis. We briefly discussed estimation and inference in sections (ref) and (ref), but a full analysis of this remains for future work.
\singlespacing