Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
35,885 characters · 8 sections · 6 citation commands
Non-testability of instrument validity under continuous treatments
Since their introduction in Appendix B of \citet*{wright1928tariff}, instrumental variables have become the main tool for identifying causal effects in empirical settings with endogeneity arising from selection on latent variables. Applications range from randomized controlled trials with imperfect compliance to classical estimation of supply- and demand systems in economics, see \citet*{stock2003retrospectives} for a historical analysis and \citet*{imbens2015causal} for a general reference.
An instrument for an endogenous treatment captures most of the effect of the treatment on the outcome while being exogenous itself. A random variable hence needs to fulfill two criteria for being an instrument in a hidden variable model: it needs to (i) have an influence on the treatment and (ii) be valid, i.e. independent of the latent terms in the model and without a direct influence on the outcome. Figure (ref) presents a schematic of the instrumental variable model in terms of directed acyclic graphs in the sense of \citet*{pearl1995testability}. Validity, captured by missing arrows between $Z$ and $U$ as well as $Z$ and $Y$ in Figure (ref), is arguably the main requirement for a random variable to be an instrument.
The question of whether instrument validity is testable has been of interest since the introduction of instrumental variables. In this context, testability is understood in the theoretical sense of existence of restrictions on the data-generating process induced by the model. In other words, if there exist data-generating processes of the observable variables $Y$, $X$, and $Z$ which cannot be replicated by the model structure in Figure (ref), then the model is theoretically testable, i.e. falsifiable. If, on the other hand, the model is too general in the sense that it places no restrictions on the observable data-generating process, it is not testable.
\citet*{pearl1995testability} was the first to explicitly address questions of instrument validity in a general setting. He derived an “instrumental inequality”, a necessary theoretical condition for an instrument to be valid. This inequality requires the endogenous treatment $X$ to be discrete and led Pearl to conjecture that testability is not possible when $X$ is continuous. Since then, several results concerning the testability of instrument validity have been derived. \citet*{manski2003partial} arrives at the same instrumental inequality in the missing data context. \citet*{kitagawa2015test} derives a test when the outcome is continuous and treatment and instrument are binary, also testing monotonicity of the instrument. \citet*{wang2017falsification} derive practically useful tests of instrument validity in the binary case. \citet*{kedagni2018sharp} show the necessity and sufficiency of Pearl's conjecture in the case where all variables are binary and augment Pearl's inequality in the case where $Z$ is discrete, $Y$ is general, and $X$ is binary. \citet*{jiang2020measurement} derive sharp bounds for the binary model under the assumption of measurement error. Finally, \citet*{bonet2001instrumentality} provides a proof of Pearl's conjecture in the special case where the outcome and the instrument are discrete.
In spite of these advancements, the question concerning the testability of the validity of instruments in general instrumental variable models has remained open in the case of continuous endogenous variables, which is a common setting in applied research: from sensitivity curves and dose-response functions in clinical trials to equilibrium models in economics. This note addresses this issue by providing a proof of Pearl's conjecture \citep*{pearl1995testability} in the most general setting. It shows that an instrumental variable model without any structural restrictions on the relations between the observable variables is too general for inducing testable implications for instrument validity when the endogenous variable is continuous. On a more positive note, we argue that weak assumptions like continuity or monotonicity between the observables re-establish theoretical testability. This provides a first na\"ive answer to the other open questions in \citet*{pearl1995testability}, asking if differentiability or monotonicity in the relation between the observable variables re-establishes theoretical testability.
For the purposes of this note, it is convenient to represent the instrumental variable model depicted in Figure (ref) as a structural model. The structural form of the model is \citep*{pearl1995testability}
where $Y$ is the outcome variable of interest, $X$ is the endogenous treatment, $Z$ is the potential instrument, and $U$ is the latent confounder. We use the terms in the singular by referring to the outcome, treatment, and instrument, even though the variables can be of arbitrary dimension. Throughout, $Z\protect\mathpalette{\protect\independenT}{\perp} U$ means that $Z$ is independent of $U$, i.e. $P_{Z,U}(A\times B) = P_{Z}(A)P_{U}(B)$ for any set of Borel sets $A$ and $B$, where $P_{Z,U}$ denotes the joint distribution of $Z$ and $U$. The latent variable $U$ captures the individual heterogeneity, i.e. all unobserved but relevant variables. The treatment $X$ is endogenous in the sense that it depends on $U$. The functions $h$ and $g$ are unknown and completely unrestricted. Potential covariates of interest can be straightforwardly included in the model by conditioning on them. We also allow for the pathological case that $X$ is continuous and independent of $Z$, in which case the relevance criterion of an instrument would be violated. \citet*{pearl1995testability} already proved that instrument validity is not testable in this special case. This also shows that the weak instrument case, i.e. the case where $X$ and $Z$ are weakly correlated, is included in our setting and has no effect on the result. Model (ref) encodes the validity of the instrument $Z$ by (i) the fact that $h$ cannot be written as a function of $Z$ (which corresponds to the missing arrow between $Z$ and $Y$ in Figure (ref)) and (ii) the independence restriction $Z\protect\mathpalette{\protect\independenT}{\perp} U$ (which corresponds to the missing arrow between $Z$ and $U$ in Figure (ref)). Full independence is required because the functions $h$ and $g$ are completely unrestricted.
Both assumptions have interpretations in the context of a double-blind clinical trial \citep*{pearl1995testability}. Here, the utilization of placebos guarantees that the treatment assignment $Z$ does not have a direct influence on the outcome process except through the actual treatment taken ($X$). In addition, randomization of the treatment ensures independence of $Z$. When randomization is not available, which is the case in observational studies for instance, instrument validity is equivalent to the independence restriction under model (ref).
This section contains the statement of Pearl's conjecture and an outline of the proof. The complete proof is relegated to the supplementary material. We prove the conjecture under the most general setting by considering a non-atomic conditional law $P_{X|Z=z}$ on Polish spaces, i.e. complete separable metric spaces. A measure $P$ on a space $\mathcal{X}$ is non-atomic if for every Borel set $A\subset\mathcal{X}$ with $P(A)>0$ there exists a Borel set $B\subset A$ with $P(A)>P(B)>0$. This level of generality allows us to consider the setting where all variables can be infinite dimensional, which could be useful in settings where dynamic considerations play a role.
In the following, calligraphic letters denote general sets. For instance, $\mathcal{X}_z$ denotes the support of $P_{X|Z=z}$ for fixed $z\in\mathcal{Z}$, i.e. the smallest closed set such that $P_{X|Z=z}(\mathcal{X}_z)=1$. All supports can be of different dimensions without affecting the result. A small letter in the function $g(\cdot,\cdot)$ denotes the realization of the corresponding random variable. For instance, $g(z,U)$ denotes the map transporting the law $P_U$ to the law $P_{X|Z=z}$ for the realization $z$.
With these preparations, we can state Pearl's conjecture.
The above statement is more technical than the wording of Pearl's original conjecture; in particular, Pearl stated that “if [$X$] is continuous, then every joint density $f_{Y,X|Z=z}$ can be generated by the instrumental process defined in [model (ref)]”. Since we work in an instrumental variable model, the relevant probability measures are $P_{Y,X|Z=z}$ and $P_{X|Z=z}$. Theorem (ref) is also stronger than Pearl's original conjecture in that it shows that testability cannot be re-established by simply assuming $g(Z,U)$ is invertible in $U$, an assumption that is sometimes made in the literature (e.g. \citeauthor*{dette2016testing} dette2016testing, Assumption 1).
To understand why the above theorem implies that instrument validity is not testable, note that the observable distribution which gives us correct information on our causal inference problem is $P_{Y,X|Z}$. In particular, the observable $P_{Y|X}$ is biased because of the endogeneity problem between $Y$ and $X$: we are interested in the unobservable counterfactual $P_{Y(x)}$, which is the conditional distribution of $Y$ given that we fix $X=x$ exogenously. But if a model can produce any possible data-generating process in the form of $P_{Y,X|Z}$, then no observable distribution can induce a testable implication on the model as mentioned in the introduction.
To connect our proof with the conjecture in \citet*{pearl1995testability}, we call $g(Z,U)$ a generator.
One-to-one generators are useful because of the following lemma, which allows us to reduce the proof of the conjecture to the first stage. This lemma is stated and proved in \citet*{pearl1995testability}, but we also provide a proof using our notation.
Restrictions on the dimension of $U$ cannot render Lemma (ref) incorrect, which follows by a classical argument about isomorphisms between Polish spaces \citep*[Theorem 9.2.2]{bogachev2007measure2}. Also note that by using Lemma 1 we do not make any assumptions on the distribution of $Y$, so that we can allow for general distributions of $Y$ without changing the result.
The idea of the proof of Theorem (ref) is hence to construct for every non-atomic $P_{X|Z}$ a first stage $X=g(Z,U)$ which is invertible in $Z$. We show even more, by arguing that one can always find a $g$ that is invertible in both $Z$ and $U$. The following paragraphs describe the main idea.
The intuition for why Pearl's conjecture is true is that for non-atomic $P_{X|Z=z}$, one can always find a continuum of different injective functions $g(z,\cdot)$ on $A_z$ and $A_u$ since every open subset contains a continuum of values. This reasoning fails in the discrete setting. We give a simple example in the next section. This is the idea for why Pearl's conjecture is true even though recent results \citep*{kedagni2018sharp} show that increasingly many restrictions are placed on the model (ref) with increasing cardinality of $X$: in the continuum limit, these restrictions become vacuous.
To introduce the underlying intuition of the proof, we focus on the case where the law of $Z$, $P_Z$, is nonatomic. The formal proof in the appendix allows for general $P_Z$ with potentially countably many atoms. The challenge is to find a one-to-one generator for any possible non-atomic $P_{X|Z}$. This level of generality requires us to work in the abstract, i.e. without specifying functional forms for $P_{X|Z=z}$ or $g$.
The fundamental idea which helps us achieve this is to regard the function $g(z,u)$ as a continuous two-dimensional array, as depicted in Figure (ref). The vertical axis in Figure (ref) indexes the continuum values of $Z$ and the horizontal axis indexes the values of $X$. Since we want our approach to work for any non-atomic $P_{X|Z=z}$, we need to work with abstract (Borel-) subsets $A_x$, $A_u$, and $A_z$ for which we assume that $x=g(z,u)$ is not one-to-one in $z$, i.e. $g(z_i,u)=g(z_j,u)$ and is onto. In the supplement, we show that we can reduce the proof of the conjecture to the sets $A_z$, $A_u$, and $A_x$ without loss of generality. The set $A_u$ is implicitly depicted by the colour coding in Figure (ref), which is easiest understood vertically: the colour coding of one infinitesimal column describes for which values $z\in A_z$ the function $g(z,u)$ maps the same $u$ to the given value $x$. For instance, the left panel of Figure (ref) is completely monochromatic on each vertical strip, meaning that for each $x\in A_x$ all $g(z,u)$ are identical for all $z\in A_z$. Different colours or shadings imply that different parts of $A_u$ are mapped to parts of $A_x$.
The goal is to make this abstract $g$ invertible in $Z$ (and incidentally $U$) on all of $A_z$. Our approach is to change $g$ iteratively by a specific Cantor scheme \citep*[Definition 6.1]{kechris1995classical} for both $A_z$ and $A_u$: the idea is to split the set $A_z$ into two disjoint subsets $A_z^{1}$ and $A_z^2$ of the same size, i.e. $P_Z(A_z^1)=P_Z(A_z^2) = \frac{1}{2}P_Z(A_z)$. This is possible because $P_Z$ is non-atomic. Since we are allowed to choose the distribution of $U$ for this proof, we assume that $P_U$ is the uniform distribution on the unit interval. We also split $A_u$ into two disjoint subsets $A_u^1$ and $A_u^2$ of the same size. This is captured in the first panel of Figure (ref).
Now here is the key to the proof: since $P_{X|Z=z}$ is non-atomic by assumption, we can also set up a Cantor scheme here and split up $A_x$ into two disjoint subsets $A_x^1$, $A_x^2$ of the same size, i.e. $P_{X|Z=z}(A_x^1)=P_{X|Z=z}(A_x^2)$ for all $z$. This is the main requirement for the construction. We now define the switching map $T^{(1)}_z:A_x\to A_x$: for $z\in A_z^2$ we let $T^{(1)}_z$ be the identity, i.e. mapping $A_x^j$ to itself, $j\in\{1,2\}$; for $z\in A^1_z$, we define $T^{(1)}_z$ to be \[T_z^{(1)}g(z,A^1_{u}) = g(z,A^2_{u})\qquad\text{and}\qquad T_z^{(1)}g(z,A^2_{u}) = g(z,A^1_{u}),\] i.e. switching $A_x^1$ and $A_x^2$. This is depicted in panel 2 of Figure (ref). By construction, it now holds that \[T^{(1)}_{z_i}g(z_i,A_u^j)\neq T^{(1)}_{z_j}g(z_j,A_u^j), \quad j\in\{1,2\},\quad\text{for all $z_i\in A_z^1$ and $z_j\in A_z^2$,}\] which is one step closer to making $T^{(1)}_zg(z,\cdot)$ the one-to-one generator.
We proceed iteratively and obtain and split the first of the subsequent subsets, i.e. $A_z^1$ into disjoint subsets $A_z^{11}, A_z^{12}$ of the same size, as captured in panels 3 and 4 of Figure (ref). Defining the switching map $T^{(2)}_z$ as the identity on $A_z^{(12)}$ and as \[T_z^{(2)}g(z,A^{\iota\wedge1}_{u}) = g(z,A^{\iota\wedge2}_{u})\qquad\text{and}\qquad T_z^{(2)}g(z,A^{\iota\wedge2}_{u}) = g(z,A^{\iota\wedge1}_{u})\] for $\iota\in\{1,2\}$, where $\iota\wedge 1$ means appending $1$ to the string $\iota$. This implies that \[T^{(2)}_{z_i}g(z_i,A_u^k)\neq T^{(2)}_{z_j}g(z_j,A_u^k), \quad k\in\{11,12,21,22\},\quad\text{for all $z_i\in A_z^{11}$ and $z_j\in A_z^{12}$.}\]
We keep this process up iteratively, i.e. constructing a classical Cantor scheme for $A_z, A_x$, and $A_u$, for which we have to perform this switching $T^{(n)}_z$ for every subset $A_z^{\iota}$, $\iota\in\{1,2\}^n$, $n\in\mathbb{N}$, we encounter in this process. As $n\to\infty$, this will lead the one-to-one generator as the sets $A_u^\iota$ and $A_z^\iota$ shrink to unique single points $u$ and $z$ for each $\iota\in 2^n$. The final panel in Figure (ref) as $n\to\infty$ hence has a different colour in each infinitesimal pixel of each infinitesimal vertical strip. We can even guarantee that each infinitesimal pixel of each infinitesimal horizontal strip has a different colour, hence guaranteeing that $g$ is invertible in both $Z$ and $U$. The proof in the supplementary material contains all details.
This approach leads to a $g$ which in general does not satisfy standard regularity assumptions like smoothness or monotonicity. This is an indication that we can reinstate theoretical testability of the model if we make functional form assumptions on the $g(z,u)$, in terms of the relationship between $X$ and $Z$. We provide an argument for this in the next section.
The verification of Pearl's conjecture is a seemingly counterintuitive result given the positive testability results derived in \citet*{pearl1995testability}, \citet*{manski2003partial}, \citet*{kitagawa2015test}, and \citet*{kedagni2018sharp}, among others. The intuition for the correctness of Pearl's conjecture lies in the complexity of the admissible models in the continuous case compared to the discrete case. \citet*{pearl1995testability} already gave an intuitive explanation for why the conjecture should be true, and we can now complement this intuition from a more rigorous perspective.
All tests of instrument validity use the idea that if model (ref) is correct and $Z\protect\mathpalette{\protect\independenT}{\perp} U$, then a change in $Z$ should not change the outcome $Y$ too drastically without changing $X$, as the latter mediates the influence of $Z$ on $Y$. As noted already by Pearl, this idea is related to Bell's inequality from quantum physics (bell2004speakable bell2004speakable, clauser1969proposed clauser1969proposed): in both settings the inequalities derive restrictions on the observable data-generating process $P_{Y,X|Z}$ which cannot be replicated by a model with a latent variable capturing the unobservable heterogeneity $U$ of the system.
For instance, consider the simple setting from the remark in \citet*{pearl1995testability}, where $X$ and $Z$ each take three values, $x_1$, $x_2$, and $x_3$ as well as $z_1$, $z_2$, $z_3$, and let $U$ be uniformly distributed on the unit interval. Moreover, assume that
In this case, the function $g(z,u)$ constructed for the first stage in the proof of Theorem (ref) cannot exist. No matter how we partition the unit interval for $U$, the fact that $P_{X|Z=z_1}(x_1)+P_{X|Z=z_2}(x_1)>1$ always implies that there is some Borel set $A_{u}\subset[0,1]$ of probability \[P_{U}(A_{u})=P_{X|Z=z_1}(x_1)+P_{X|Z=z_2}(x_1)-1=\varepsilon_{u}>0,\] which gets mapped to $x_1$ for both $z_1$ and $z_2$, because $g(z,\cdot)$ preserves measure by construction. In the continuous setting, this example cannot hold because every Borel set of positive probability contains a continuum of points by definition, so that there is never a single point $x_1$ satisfying the above property.
It is helpful to interpret the function $g(z,u)$ for fixed $u$ as a response profile of the unobservable unit $u$ for any given action $z$, i.e. one path $X_z(u)\equiv g(z,u)$ of a counterfactual stochastic process $X_z$. This idea was already implicit in \citet*{pearl1995testability} and \citet*{angrist1996identification} in the binary setting. Condition (ref) then implies that the response $x$ is the same for the actions $z_1$ and $z_2$ for almost all units $u\in A_u$. Similarly, one has a response profile $Y_x(u)\equiv h(x,u)$ for the second stage. Together they imply a joint response profile \[(Y,X)_z(u)\equiv \left\{Y_{X_z(u)}(u), X_z(u)\right\}\equiv \{h(g(z,u),u), g(z,u)\},\] which takes the simple form because of the exclusion restriction. Importantly, the exclusion restriction implies that the marginal process $Y_{X_z(u)}(u)$ only depends on the position $x$ of $X_z(u)$ for given $z$ and not the whole path. The law of this joint profile $(Y,X)_z(u)$ is the one that needs to be compared to the observable data-generating-process $P_{Y,X|Z=z}$.
Now if $g(z,u)$ is constant for some $z_1$ and $z_2$ and a given $A_u$ of positive probability as in (ref), it can happen that the observable $P_{Y,X|Z=z}$ requires a joint response profile $(Y,X)_z(u)$ in which the marginal response $Y_{X_{z}(u)}(u)$ needs to change between $z_1$ and $z_2$ for all $u\in A_u$, i.e. \[Y_{X_{z_1}(u)}(u) \neq Y_{X_{z_2}(u)}(u).\] However, (ref) implies that $X_{z_1}(u)=X_{z_1}(u)$ for these $u\in A_u$, which is a contradiction due to the fact that $Y_{X_z}$ only depends on the position $X_z$ for $z$ by the exclusion restriction. This implies immediately that the model (ref) cannot replicate this specific $P_{Y,X|Z=z}$. Here, we can recognize the interplay of the exclusion restriction, which implies the form of the response profile $Y_{X_z(u)}(u)$, and the independence $Z\protect\mathpalette{\protect\independenT}{\perp} U$. At least one of the two has to be violated in the case of (ref), which implies a testable restriction on the model.
The above argumentation is not possible in the setting where $P_{X|Z=z}$ is non-atomic without further assumptions. The reason is that an equation like (ref) cannot hold in the continuous setting, as we condition on sets of measure zero. Hence, the observable $P_{X|Z=z}$ can never introduce a case where the response profile $X_{z}(u)$ is constant over some set $A_z\subset\mathcal{Z}$ and for some $u\in A_u$. This is the intuition of the proof of Pearl's conjecture as captured in Figure (ref). The observable data-generating-process $P_{X|Z=z}$ could still be very erratic in principle, but the proof of Pearl's conjecture shows that we can always find an erratic enough response profile $X_z(u)\equiv g(z,u)$ which allows us to replicate it. However, for each $u$ the constructed paths $X_z(u)\equiv g(z,u)$ in the proof of Pearl's conjecture do not satisfy regularity properties like continuity or monotonicity in general.
Therefore, one can re-establish testability of model (ref) even in the continuous case if one is willing to make weak structural assumptions on the paths $X_z(u)$. In particular, by the above reasoning one can immediately obtain testable implications if one assumes that $X_z(u)$ is constant on some set $A_z\in\mathcal{Z}$ of positive probability for some set $A_u$ as above. This assumption is rather artificial. However, using the same reasoning, one can straightforwardly show that a continuity assumption on $X_z$ and $Y_x$, based on the Kolmogorov-Chentsov theorem \citep*[corollary 14.9]{kallenberg2006foundations} for instance, is already enough to introduce testable implications in principle. Intuitively, the joint response profile $(Y,X)_z(u)$ must be continuous if both $Y_x(u)$ and $X_z(u)$ are continuous, and there exist data-generating processes $P_{Y,X|Z=z}$ which would induce a response profile with jumps. A similar reasoning can be made rigorous for monotone response profiles. This gives an informal positive answer to the other open questions in \citet*{pearl1995testability} about re-establishing testability through structural assumptions like differentiability and monotonicity. More work needs to be done to analyze stronger functional form assumptions on $g$ and $h$ which make the model testable in practice, not just theoretically.
Lastly, this article may also have implications for local hidden variable theories \citep*{bell2004speakable}. It shows that even though a Bell-type inequality does not exist in our continuous setting ou1992realization, one has to allow for general models without functional form restrictions in order to arrive at this conclusion.