EconBase
← Back to paper

Controlling for Latent Confounding with Triple Proxies

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

56,339 characters · 11 sections · 34 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Controlling for Latent Confounding with Triple Proxies

abstractWe present new results for nonparametric identification of causal effects using noisy proxies for unobserved confounders. Our approach builds on the results of Hu2008 who tackle the problem of general measurement error. We call this the `triple proxy' approach because it requires three proxies that are jointly independent conditional on unobservables. We consider three different choices for the third proxy: it may be an outcome, a vector of treatments, or a collection of auxiliary variables. We compare to an alternative identification strategy introduced by Miao2018a in which causal effects are identified using two conditionally independent proxies. We refer to this as the `double proxy' approach. The triple proxy approach identifies objects that are not identified by the double proxy approach, including some that capture the variation in average treatment effects between strata of the unobservables. Moreover, the conditional independence assumptions in the double and triple proxy approaches are non-nested.

A key challenge for causal inference is to credibly adjust for all confounding factors -variables that impact both treatments and outcomes. Unfortunately, important confounders may not be recorded in available data and researchers must make do with noisy and biased proxies for these factors.\footnote{We use the term `proxy' very generally to refer to observables that are statistically associated with some unobservables.} Methods that adjust for observed confounding like inverse propensity score re-weighting or nonparametric regression typically fail to recover causal effects when some of the confounders are replaced with noisy proxies. This is a problem of potentially non-classical measurement error, in which the mismeasured variables are controls.

For concreteness, suppose we wish to assess the effect of an educational intervention on a student's high school GPA. We adjust for observed pre-treatment characteristics like family background and age of the student. In addition, we would like to control for academic aptitude which could confound treatments and outcomes. While we do not observe aptitude directly, we may have access to test scores which are noisy and possibly biased proxies for academic ability. Simply controlling for test scores and pre-treatment covariates will generally not recover a causal effect.

In this work we provide new results on the nonparametric identification of causal effects using proxies for unobserved confounders. In order to deal with the mismeasured confounding we apply results from Hu2008 (hereon HS). HS has been applied to numerous problems in empirical economics, notably to the analysis of skill-formation in Cunha. HS identify the joint distribution of the observables and underlying mismeasured variables under non-classical measurement error. In order to achieve this, HS require three vectors of proxies that are jointly independent conditional on the latent mismeasured variables. We show that under appropriate conditions either the outcome or vector of treatments can serve as the third proxy. Due to the use of three vectors of proxies we refer to identification based on HS as the `triple proxy' approach.

Crucially, in our setting only confounders are mismeasured. We are interested in effects of treatments, which are measured correctly. This allows us to drop a key assumption of HS, namely that one vector of proxies is a mean- or median-unbiased signal for the unobservables. Without this condition we can only identify distributions involving the latent confounders up to an unknown one-to-one transformation of these factors. However, many causal estimands are invariant to such transformations. For example, the average treatment effect or the effect of treatment on the treated. Therefore, we are able to point identify these objects without an unbiasedness condition. The same is true of some objects that quantify the degree of heterogeneity in treatment effects between different strata of the latent confounders, e.g., the variance of the conditional average treatment effect.

Depending on the choice of the third proxy (either the outcome, treatments, or a vector of auxiliary variables), identification of causal effects may require a two-step approach. First we identify distributions involving the latent confounders up to an invertible transformation of these factors using results from HS. We then identify objects of interest from a linear integral equation that involves an object obtained in the first step. A closely related two-step strategy was previously explored in the context of regression discontinuity design in Rokkanen2015.

We show that by carefully selecting the third proxy and possibly controlling for treatments or outcomes in the HS step, our approach can accommodate a wide variety of causal relationships between proxies, treatments, and outcomes. This is important because it can be difficult to rule out a priori the possibility that say, a pre-treatment proxy determines treatment, or a post-treatment proxy is affected by treatment. In order to thoroughly and systematically assess the sets of modeling assumptions compatible with our approach, we make extensive use of Directed Acyclic Graphs (DAGs). Our identification results are based on conditional independence restrictions, however the DAGs (or more precisely, the nonparametric structural models associated with each DAG) serve as intuitive primitive conditions for these independence assumptions.

We compare our results to those of a closely related literature on nonparametric identification using two conditionally independent proxies for latent confounders. We refer to this as the `double proxy' approach. The double proxy approach constitutes a large and recent literature, key papers include Miao2018a, Deaner2021, Kallus2021, and Singh2020. A comparative advantage of the triple proxy approach is that it can identify heterogeneity in causal effects between strata of the latent factors, this is not identified by the double proxy approach. We show that the exclusion restrictions (and consequent conditional independence assumptions) of the double and triple proxy approaches are non-nested. That is, there are conditions under which the double proxy approach is applicable but not the triple proxy approach, and vice versa. Thus our work expands the settings in which one can credibly identify causal effects using proxies for unobserved confounders.

While causal effects of the confounders are not of primary interest in our analysis, an advantage of the triple proxy over the double proxy approach is that under additional restrictions it can be adapted straightforwardly to identify these effects. In a recent work, Freyberger2021 shows that if the mean- or median-unbiasedness condition of HS is replaced with a related monotonicity condition, then one can identify objects involving quantiles of the latent variables. We apply a similar strategy to Freyberger2021 in our setting and thus identify the causal effect of shifting the latent confounders between quantiles.

In Section 1 of the paper we provide a motivating example for our analysis. Section 2 contains a brief summary of results in Hu2008. Section 3 establishes identification using various choices of the third proxy (treatment, outcomes, or auxiliary variables) and discusses the requisite exclusion restrictions. Section 4 shows how, under additional conditions we can extend our results to identify a richer set of objects including causal effects of the latent variables themselves. Section 5 concludes. The appendix contains proofs as well as additional results that show we can achieve partial identification (and possibly point identification) of causal effects under weaker exclusion restrictions so long as a certain rank invariance condition holds.

Notation and Technicalities

Throughout we use upper case letters to denote random variables while the corresponding lower case letters denote values of these random variables. If $A$, $B$, $C$, and $D$ are random variables and $A$ and $B$ admit a joint probability density function conditional on $C$ and $D$, then $f_{AB|CD}(a,b|c,d)$ denotes such a density evaluated at $A=a$, $B=b$, $C=c$ and $D=d$.

If $Y$ is an outcome variable and $X$ a treatment, then $Y(x)$ is the potential outcome from a counterfactual level $x$ of $X$. Throughout we implicitly assume that $Y(X)=Y$, a condition sometimes known as `consistency'.

As a technical note, statements about probability densities of random variables should be interpreted to hold for some density compatible with the joint probability measure of these variables. For example, the statement `$f_{AB}(a,b)=f_A(a)f_B(b),\,\forall a,b$' should be understood to mean that `$A$ and $B$ admit a joint density $f_{AB}$ and marginal densities $f_A$ and $f_B$ so that $f_{AB}(a,b)=f_A(a)f_B(b),\forall a,b$'.

A Motivating Example

Suppose we are interested in the causal effect of an educational intervention $X$ on a student's GPA at the end of high-school $Y$. Whether or not a student receives the intervention is determined by the student's teachers, parents, and perhaps the student herself. These actors base their decision, at least in part, on their private assessments of the student's academic aptitude.

In this setting, academic aptitude (at the time treatment status is decided) is an unobserved confounder. It affects the decision to treat the student and it has an effect on high-school GPA, regardless of treatment. The researcher has access to some test scores that reflect academic ability, but which do not measure it perfectly. Note that our analysis does not require aptitude be one-dimensional, rather it can be understood as a finite-dimensional vector.

In sum, test scores are noisy and possibly biased measurements of an unobserved confounder (academic aptitude). The need to account for the mismeasurement of ability arises in numerous empirical applications, for example in Griliches1972, Fruehwirth2016, and Deaner2021.

Of course, there may be other potential confounders, for example the student's socio-economic characteristics, which are correctly measured in the data. We abstract away from this by implicitly assuming that the researcher has already conditioned on these factors (i.e., our analysis applies within each stratum of these covariates).

Identification is complicated by the fact that test scores can directly cause or be caused by treatments and outcomes. If the educational intervention affects a student's academic progress, then it presumably affects the scores on tests taken after the intervention. If a test is taken prior to the intervention, then it may determine eligibility for treatment. A test score could directly enter into the GPA calculation or may be used to decide some feature of the student's education other than the intervention, and thus it may have an affect on GPA that is not mediated by treatment.

Ruling out causal relationships of this kind requires detailed institutional knowledge. In this work we show that the triple proxy approach allows for causal relationships between the proxies, treatments, and outcomes that are incompatible with the double proxy approach, and vice versa. Thus there may be settings in which institutional knowledge is compatible with one of the two approaches but not the other.

In Figure 1 we present Directed Acyclic graphs (DAGs) that encode possible causal relationships between academic ability, the treatment (an educational intervention), the outcome (GPA at the end of high-school), and two or three sets of test scores.

Each DAG encodes a set of exclusion restrictions in a Non-Parametric Structural Equations Model (NPSEM) of the kind in Pearl2009. Each NPSEM implies a set of conditional independence restrictions required for identification of causal effects using either the double or triple proxy approach.

figure[figure omitted — 2,588 chars of source]

The causal diagram in Figure 1.a implies the conditional independence restrictions required by both the double and triple proxy approaches. In this DAG one set of scores are from `early tests' taken prior to the decision to treat and some are from `late tests' taken after treatment is administered.

The DAG in Figure 1.a encodes an assumption that there is no direct effect of the tests on high-school GPA, nor on treatment. This effectively rules out the possibility that the test scores are used to determine any important aspects of a student's education. In Deaner2021 this is justified by the fact that the test scores are only observed by researchers who have no input into the students' education.

The graph in Figure 1.a allows the educational intervention to affect the scores on post-treatment tests, this is important because if the educational intervention affects academic performance then it likely affects future test scores.

The graph in Figure 1.b is adapted from Miao2018a. In this case one set of early tests can determine treatment. The other early tests cannot affect treatment but could impact some other aspect of the student's education and thus affect the outcome.

Figure 1.b is compatible with the double proxy approach but not the triple proxy approach. The triple proxy approach employs Hu2008, which requires three proxies that are independent conditional on unobservables. Under Figure 1.b no three of the four observables are guaranteed to be jointly independent conditional on ability. Nor are there three observables that are independent conditional on ability and whichever observable is left over.

Figure 1.c is compatible with the triple proxy approach but not the double proxy approach. In this case, one set of scores are from tests taken after high-school graduation. High-school GPA could affect say, college attendance and thus later test scores. The treatment may affect these late test scores so long as this is mediated by academic progress in high-school as measured by GPA.

The model in Figure 1.d allows for identification using the triple proxy approach but not double proxies. This is because two of the three available proxies have direct causal links to both treatments and outcomes, which rules out their use in the double proxy approach. In this model one set of early tests can determine treatment and may have some other impact on the student's education and thus affect GPA directly. Treatment can impact the late test scores, and these scores may determine the course of a student's later education and thus affect GPA.

Identification in Hu and Schennach (2008)

HS prove identification of a nonparametric factor model. We restate their results below. Let $W$ be an unobserved, possibly vector-valued latent factor with support $\mathcal{W}$. Let $V$, $Z$, and $C$ be observable random vectors with respective supports $\mathcal{V}$, $\mathcal{Z}$, and $\mathcal{C}$. The following conditions are from Hu and Schennach (2008).

\theoremstyle{definition} \newtheorem*{AHS1}{HS Assumption 1} \begin{AHS1} $V$, $Z$, $W$, and $C$ admit a bounded, non-zero density with respect to the product measure of the Lebesgue measure on $\mathcal{V}\times\mathcal{Z}\times\mathcal{W}$ and some dominating measure $\mu$ on $\mathcal{C}$. All marginal and conditional densities are also bounded. \end{AHS1}

\theoremstyle{definition} \newtheorem*{AHS2}{HS Assumption 2} \begin{AHS2} $V$, $Z$, and $C$ are jointly independent conditional on $W$. Formally, $V\perp \!\!\! \perp(Z,C)|W$ and $Z\perp \!\!\! \perp C|W$. \end{AHS2}

\theoremstyle{definition} \newtheorem*{AHS3}{HS Assumption 3} \begin{AHS3} For any bounded function $\delta$ in $ L_1(\mathcal{W})$: \[\int_\mathcal{W} f_{V|W}(V|w)\delta(w)dw\overset{a.s.}{=}0\implies \delta(W)\overset{a.s.}{=}0\] and the same holds with $V$ replaced by $Z$. \end{AHS3}

\theoremstyle{definition} \newtheorem*{AHS4}{HS Assumption 4} \begin{AHS4} For any $w_1,w_2\in\mathcal{W}$ if $w_1\neq w_2$ then $P\big(f_{C|W}(C|w_1)\neq f_{C|W}(C|w_2)\big)>0$. \end{AHS4}

HS Assumption 1 ensures some bounded densities exist. HS Assumption 2 states that $V$, $Z$, and $C$ are jointly independent conditional on $W$.

If the marginal densities $f_W$ , $f_V$, and $f_Z$ are bounded below away from zero over $\mathcal{W}$, $\mathcal{V}$, and $\mathcal{Z}$ respectively, then Assumption 3 is equivalent to two bounded completeness conditions. Namely, bounded completeness of $W$ for $V$, and bounded completeness of $W$ for $Z$. Note that Assumption 3 differs slightly from the corresponding condition in Hu2008, the version we use here is employed in the Handbook of Econometrics treatment of HS (see Schennach2020).

Statistical completeness conditions are used to identify Nonparametric Instrumental Variables (NPIV) models of the kind in Newey2003 and Ai2003. Thus condition 3 states that $V$ and $Z$ are both relevant instruments for $W$ in the sense of NPIV.

Assumptions 4 is a relatively weak condition on the association between $C$ and $W$. Hu2008 note that this assumption is weaker than imposing HS Assumption 3 on $C$. In words it states that any change in $W$ must induce some change in the conditional distribution of $C$. This condition can hold even if $C$ is a binary random variable and $W$ is a continuous random vector.

\theoremstyle{plain} \newtheorem*{HST}{HS Theorem (Hu and Schennach (2008))}

HSTUnder HS Assumptions 1 and 2 the following equality holds: \begin{equation} f_{ZC|V}(z,c|v)=\int_\mathcal{W} f_{C|W}(c|w)f_{W|V}(w|v)f_{Z|W}(z|w)dw \end{equation} Moreover, under HS Assumptions 1-4, $f_{W|V}$, $f_{Z|W}$, and $f_{W|C}$ are identified from the above up to a reordering of $W$.

To formalize what we mean by `identified up to reorderings', suppose some other conditional densities $\tilde{f}_{W|V}$, $\tilde{f}_{Z|W}$, and $\tilde{f}_{W|C}$ satisfy Assumption 1-4 and ((ref)):

align*[align* omitted — 110 chars of source]

Then there exists an bijective function $\varphi:\mathcal{W}\to\mathcal{W}$ so that $\tilde{f}_{W|V}(w|v)=f_{\varphi(W)|V}(w|v)$, $\tilde{f}_{Z|W}(z|w)=f_{Z|\varphi(W)}(z|w)$, and $\tilde{f}_{C|W}(c|w)=f_{C|\varphi(W)}(c|w)$.

Theorem 1 identifies conditional densities up to a reordering of $W$. HS pin down a single ordering using an additional assumption given below. In the case of mismeasured control variables, causal effects of interest are often invariant to reordering of $W$, and so we do not require this condition. However, we revisit it in Section 4.

\theoremstyle{definition} \newtheorem*{AHS5}{HS Assumption 5} \begin{AHS5} There is a known functional $M$ so that $M[f_{Z|W}(\cdot|w)]=w,\,\forall w\in\mathcal{W}$. \end{AHS5}

If the functional $M$ returns the mean of the distribution in its argument then the assumption states that $Z$ is mean-unbiased for $W$. If $M$ returns the median, then the assumption states that $Z$ is median-unbiased for $W$. It is implicit in the assumption that the dimensions of $W$ and $Z$ are the same.

Identifying Causal Effects

To identify causal effects with mismeasured controls, we suppose that two vectors of proxies $V$ and $Z$ are available. The third proxy $C$ will be either the outcome $Y$, treatment $X$, or some additional observables. These choices are appropriate under different sets of exclusion restrictions.

Objects of Interest

As discussed in the previous section, without a condition like HS Assumption 5, we can only identify objects involving $W$ up to a reordering of this variable. Fortunately, causal effects of the treatment $X$ are typically invariant to reordering of $W$, and so we can point identify many causal objects without invoking such an assumption. These include objects that capture the degree of heterogeneity in treatment response among groups with different values of $W$. Such objects are not identified using the double proxy approach.

In order to point identify causal objects, we first identify the joint distributions of potential outcomes $Y(x)$, latent variables $W$, and realized treatments $X$, for $\mu_X$-almost all $x$, up to a reordering of $W$.\footnote{$\mu_X$ and $\mu_Y$ are a probability laws that dominates those of $X$ and $Y$ respectively. We allow for both discrete and continuous treatments and likewise for the outcome.}

Let $f_{Y(x)WX}$ be a joint density of $Y(x)$, $W$, and $X$. The marginal density of potential outcomes and the density conditional on the realized treatment are given by the expressions below. These are invariant to reordering of $W$ and so can be point identified even when we cannot recover the ordering of $W$.

align[align omitted — 174 chars of source]

From the above we can further obtain say, average treatment effects and the effect of treatment on the treated, as well as quantile treatment effects.\footnote{By `quantile treatment effects' we mean the differences in quantiles of potential outcomes under different treatments. These may differ from the quantiles of the treatment effect.} With continuous treatments we can identify average partial effects under a suitable differentiability condition.

In addition, from $f_{Y(x)WX}$ we can recover objects that capture the heterogeneity in treatment response between groups with different values of the latent variables $W$. Suppose $X$ is binary, then the conditional average treatment effect $\beta(W)\equiv E[Y(1)-Y(0)|W]$, is given by: \[ \beta(w)=\int_\mathcal{Y} y f_{Y(1)|W}(y|w)d\mu_Y-\int_{\mathcal{Y}} y f_{Y(0)|W}(y|w)d\mu_Y \] Where we implicitly assume the mean exists and is finite. While $\beta(w)$ itself is not invariant to reorderings of $W$, its distribution is invariant. Thus we can identify say, the variance of $\beta(W)$. The cumulative distribution function (CDF) of $\beta(W)$ conditional on $X=x$, and the marginal CDF are given below:

align[align omitted — 180 chars of source]

Neither the marginal nor conditional distribution of $\beta(W)$ are identified by the double proxy approach.

Similarly, when treatment is continuous, and the derivatives of potential outcomes exist and are bounded, we can obtain the conditional average partial effect $\pi(W)\equiv E[\frac{\partial}{\partial x} Y(x)|W]$ up to a reordering of $W$ as follows: \[ \pi(W)=\frac{\partial}{\partial x}\int_{\mathcal{Y}} y F_{Y(x)|W}(dy|W) \] The conditional and marginal CDFs of the conditional average partial effect are then given by:

align*[align* omitted — 147 chars of source]

As a technical caveat, when treatments are continuous we can only identify the densities $f_{Y(x)WX}$ for $\mu_X$-almost all $x$ (as opposed to all $x$), up to a reordering of $W$. In practice this is likely to be of little consequence because for estimation one typically assumes the support of $X$ is rectangular and objects of interest like $E[Y(x)]$ are continuous in $x$. In this case, if $E[Y(x)]$ is identified almost everywhere then it is in fact identified everywhere.

Outcome Proxies

We first apply the results in Section 2 to identify $f_{Y|WX}$ and $f_{WX}$ up to a reordering of $W$ in a first stage. The outcome $Y$ acts as the proxies $C$. We apply the results in Section 2 after conditioning on $X$, i.e., within each stratum of the treatment. Having identified $f_{Y|WX}$ and $f_{WX}$ we recover $f_{Y(x)WX}$ up to a reordering. We can then point identify causal objects like ((ref)) which are invariant to reordering of $W$.

In this context we assume the existence of bounded densities akin to HS Assumption 1 in the previous section. We also assume that at each $x$, the potential outcome function $Y(\cdot)$ is almost surely continuous. This assumption is trivially true when treatment is discrete and allows us to avoid some technical issues that arise with continuous treatments. Let $\mathcal{V}$, $\mathcal{Z}$, $\mathcal{W}$, $\mathcal{Y}$, and $\mathcal{X}$ be the supports of $V$, $Z$, $W$, $Y$, and $X$ respectively.

\theoremstyle{definition} \newtheorem*{A1}{Assumption 1} \begin{A1} $V$, $Z$, $W$, $Y$, and $X$ admit a bounded, non-zero density with respect to the product of the Lebesgue measure on $\mathcal{V}\times\mathcal{Z}\times\mathcal{W}$, some dominating measure $\mu_Y$ on $\mathcal{Y}$, and a dominating measure $\mu_X$ on $\mathcal{X}$. All marginal and conditional densities are also bounded. \end{A1}

Figure 2 contains three alternative causal graphs. These graphs encode exclusion restrictions on an NPSEM that imply a set of conditional independence restrictions which we use for identification.

figure[figure omitted — 1,499 chars of source]

The causal graphs in Figures 2.a and 2.b suggest that $V$ is a pre-treatment variable and allow $V$ to affect the treatment $X$. The graph in 2.c suggests $V$ is a post-treatment and could be affected by treatment.

The graphs preclude $X$ affecting $Z$ or vice versa. This is most credible when $Z$ is a pre-treatment variable (and thus cannot be affected by treatment), which is not used to decide treatment.

Crucially, neither the proxies $Z$ nor $V$ may directly affect the outcome $Y$. Moreover, $Z$ must not directly affect $V$ nor vice versa.

The graphs in Figure 2 imply a set of conditional independence restrictions given in Proposition 1. The proposition can be verified straight-forwardly using the tools in Pearl2009. Our identification results directly assume the conditional independence restrictions in the conclusion of Proposition 1. Thus the graphs in Figure 2 can be understood to represent primitive conditions for these conditional independence restrictions.

\theoremstyle{plain} \newtheorem*{P1}{Proposition 1} \begin{P1} The NPSEMs associated with the causal graphs in Figure 2 all imply the following conditional independence restrictions:

i. $Y\perp \!\!\! \perp(V,Z)|(W,X)$, ii. $V\perp \!\!\! \perp Z|(W,X)$, iii. $Z\perp \!\!\! \perp X|W$, and iv. $Y(x)\perp \!\!\! \perp (X,V)|W$ \end{P1}

The conditional independence restrictions in Proposition 1 are stronger than those required for the double proxy approach. In particular, the double proxy approach requires conditions ii., iii., and iv. but not condition i.

In this setting we use the Assumption 2 below in place of HS Assumption 3.

\theoremstyle{definition} \newtheorem*{A2}{Assumption 2} \begin{A2} For each $x\in\mathcal{X}$ and any bounded function $\delta$ in $ L_1(\mathcal{W})$: \[\int_\mathcal{W} f_{V|WX}(V|w,x)\delta(w)dw\overset{a.s.}{=}0,\,\implies \delta(W)\overset{a.s.}{=}0\] and the same holds with $V$ replaced by $Z$. \end{A2}

Assumption 2 differs from HS Assumption 3 in that it must hold within each stratum of the treatment $X$. If the conditional densities $f_{W|X}$, $f_{V|X}$, and $f_{Z|X}$ are all bounded below away from zero and have bounded supports, then Assumption 2 is equivalent to the completeness conditions in Deaner2021.

Finally, Assumption 3 below plays the role of HS Assumption 4. Note that this condition allows for the possibility that the outcome $Y$ is binary, even if $W$ is a continuous random vector. \theoremstyle{definition} \newtheorem*{A3}{Assumption 3} \begin{A3} For all $x\in\mathcal{X}$ and any $w_1,w_2\in\mathcal{W}$, if $w_1\neq w_2$ then: \[P\big(f_{Y|WX}(Y|w_1,x)\neq f_{Y|WX}(Y|w_2,x)\big)>0\] \end{A3}

We now identify causal effects under the conditions above.

\theoremstyle{plain} \newtheorem*{T1}{Theorem 3.1 (Outcome Proxies)} \begin{T1} Suppose Assumptions 1-3 and conclusions i., ii., and iii. of Proposition 1 and hold. Then $f_{Y|WX}$, $f_{Z|W}$, and $f_{W|VX}$ (and thus $f_{WX}$) are identified up to a reordering of $W$ from the equation below:

equation*[equation* omitted — 97 chars of source]

Under conclusion iv. of Proposition 1, for $\mu_X$-almost all $x_1$:

equation[equation omitted — 80 chars of source]

And so $f_{Y(x)WX}$ is identified for $\mu_X$-almost all $x$, up to a reordering of $W$. Causal objects can then be point identified by say, ((ref)) or ((ref)). \end{T1}

Without Conclusion iii. of Proposition 1, we could still apply Hu2008 to identify $f_{Y|WX}(\cdot|\cdot,x)$, $f_{Z|WX}(\cdot|\cdot,x)$, and $f_{W|VX}(\cdot|\cdot,x)$ up to a reordering of $W$ for each $x$ in the support of $X$. However, the reordering of $W$ could differ between the values of $x$. We revisit this possibility in Appendix A and show that partial identification (and possibly point identification) can be achieved without condition iii. under a rank invariance assumption.

Treatment Proxies

We now consider the case in which the third proxy $C$, is the vector of treatments $X$. In this case identification proceeds in two stages. In a first stage we use results from HS to identify conditional distributions involving $W$ up to reordering. In a second step, the conditional distribution of potential outcomes is identified via a linear integral equation. This two-step approach is similar to one employed in Rokkanen2015.

In this case, the causal diagrams below are sufficient for the conditional independence restrictions under which we establish identification.

figure[figure omitted — 1,498 chars of source]

The diagrams in Figure 3 suggest that $V$ is determined prior to the outcome, and the graphs allow $V$ to directly affect the outcome. However, $V$ and $Z$ cannot directly affect, or be directly affected by, the treatment $X$. This contrasts with the double proxy case, which allows one of the two proxies to be directly causally connected to the treatment.

\newtheorem*{P2}{Proposition 2} \begin{P2} The NPSEMs associated with the causal graphs in Figure 3 imply the following conditional independence restrictions:

i. $V\perp \!\!\! \perp(X,Z)|W$, ii. $X\perp \!\!\! \perp Z|W$, iii. $Y\perp \!\!\! \perp Z|(W,X)$, and iv. $Y(x)\perp \!\!\! \perp (X,Z)|W$ \end{P2}

Conditions i., and iv. in Proposition 2 are those required for the double proxy approach. However, the double proxy approach does not require condition ii.

Strictly speaking, condition iii. is not required for the double proxy approach. However, if condition iv. is strengthened slightly to $Y(\cdot)\perp \!\!\! \perp (Z,X)|W$ this implies condition iii.

Assumption 4 replaces HS Assumption 4. Note that Assumption 4 may hold even if the treatment $X$ is binary.

\theoremstyle{definition} \newtheorem*{A4}{Assumption 4} \begin{A4} For any $w_1,w_2\in\mathcal{W}$ if $w_1\neq w_2$ then $P\big(f_{X|W}(X|w_1)\neq f_{X|W}(X|w_2)\big)>0$. \end{A4}

\theoremstyle{plain} \newtheorem*{T2}{Theorem 3.2 (Treatment Proxies)} \begin{T2} Suppose conclusions i. and ii. of Proposition 2, HS Assumption 3, and Assumptions 1 and 4 hold. Then $f_{Z|W}$ is identified up to a reordering of $W$ from the equation below:

equation*[equation* omitted — 89 chars of source]

If conclusion iii. of Proposition 2 also holds, $f_{YWX}$ is then identified up to reorderings of $W$ from the integral equation below.

equation*[equation* omitted — 79 chars of source]

Under conclusion iv. of Proposition 2 we recover $f_{Y(x)WX}$ for $\mu_X$-almost all $x$, up to a reordering of $W$ from ((ref)). Causal objects can then be point identified by say, ((ref)) or ((ref)). \end{T2}

Outcome-Conditional Treatment Proxies

We again consider the case in which the third proxy $C$, is the vector of treatments $X$. However, we apply Hu2008 within each stratum of the outcome. This allows for the possibility that the outcome directly affects one of the proxies $V$, which is generally incompatible with the double proxy approach.

figure[figure omitted — 1,097 chars of source]

The diagrams in Figure 4 differ from those in Figure 3 in that $V$ is a post-outcome variable and can be impacted directly by the outcome. Note that $V$ is a post treatment variable but $X$ must not affect $V$ directly.

Recall the test score example with the outcome $Y$ measuring GPA in the final year of high-school. Suppose the tests in $V$ taken a year after high-school graduation. GPA may affect college attendance which could in turn impact scores on the college-age tests $V$. The educational intervention $X$ can influence post-high school test scores, so long as the effect of the intervention is mediated by academic achievement over high school as measured by final GPA.

\newtheorem*{P3}{Proposition 3} \begin{P3} The NPSEMs associated with the causal graphs in Figure 4 imply the following conditional independence restrictions:

i. $V\perp \!\!\! \perp(X,Z)|(W,Y)$, ii. $X\perp \!\!\! \perp Z|(W,Y)$, iii. $Y\perp \!\!\! \perp Z|W$, and iv., $Y(x)\perp \!\!\! \perp X|W$. \end{P3}

The conditions in Proposition 3 are insufficient for the double proxy approach. The double proxy approach requires that $V\perp \!\!\! \perp(X,Z)|W$ which generally rules out $V$ having a direct causal effect on $Y$ (unless we were to assume $X$ has no causal effect on $Y$ which defeats the purpose of our analysis). Conversely, the double proxy approach does not require any independence between $X$ and $Z$, conditional or otherwise.

In this setting we need HS Assumption 3 to apply within each stratum of the outcome.

\theoremstyle{definition} \newtheorem*{A5}{Assumption 5} \begin{A5} For each $y\in\mathcal{Y}$ and any bounded function $\delta$ in $ L_1(\mathcal{W})$: \[\int_\mathcal{W} f_{V|WY}(V|w,y)\delta(w)dw\overset{a.s.}{=}0,\,\implies \delta(W)\overset{a.s.}{=}0\] and the same holds with $V$ replaced by $Z$. \end{A5}

Finally, we need HS Assumption 4 to hold for the treatment proxy within each stratum of $Y$.

\theoremstyle{definition} \newtheorem*{A6}{Assumption 6} \begin{A6} For all $y\in\mathcal{Y}$ and any $w_1,w_2\in\mathcal{W}$, if $w_1\neq w_2$ then: \[P\big(f_{X|WY}(X|w_1,y)\neq f_{X|WY}(X|w_2,y)\big)>0\] \end{A6}

\theoremstyle{plain} \newtheorem*{T3}{Theorem 3.3 (Conditional Treatment Proxies)} \begin{T3}

Suppose conclusion i., ii., and iii, of Proposition 3 and Assumptions 1, 5, and 6 hold. Then $f_{X|WY}$, $f_{Z|W}$, and $f_{W|VY}$ are identified up to a reordering of $W$ from the equation below:

equation*[equation* omitted — 102 chars of source]

$f_{YWX}$ can be written in terms of $f_{X|WY}$, $f_{W|VY}$, and $f_{VY}$. If conclusion iv. of Proposition 3 we recover $f_{Y(x)WX}$ for $\mu_X$-almost all $x$, up to a reordering of $W$ from ((ref)). Causal objects can then be point identified by say, ((ref)) or ((ref)).

\end{T3}

Auxiliary Proxies

Finally, we consider the case in which the third proxy $C$ is a vector of auxiliary variable (as opposed to $X$ or $Y$). In this case we apply the results from Section 2 within each stratum of the treatments.

\theoremstyle{definition} \newtheorem*{A7}{Assumption 7} \begin{A7} $V$, $Z$, $W$, $Y$, $X$, and $C$ admit a bounded, non-zero density with respect to the product of the Lebesgue measure on $\mathcal{V}\times\mathcal{Z}\times\mathcal{W}$, some dominating measure $\mu_Y$ on $\mathcal{Y}$, a dominating measure $\mu_X$ on $\mathcal{X}$, and a dominating measure $\mu_C$ on $\mathcal{C}$. All marginal and conditional densities are also bounded. \end{A7}

figure[figure omitted — 1,944 chars of source]

The causal graphs in Figure 5 provide the key exclusion restrictions in this setting. The strongest restrictions in Figure 5 are on $Z$. $Z$ cannot directly cause or be caused by treatment or outcome.

Suppose $C$ is a post-treatment proxy and $V$ is a pre-treatment proxy, and both are determined prior to the outcome $Y$. In addition, let us assume that $W$ is determined prior to all the observables, with the possible exception of $V$. Then $V$ and $C$ can be caused by all variables determined prior to them other than $Z$ and can determine all variables determined after them other than $Z$.

The exclusion restrictions in Figure 5 are not sufficient for the independence restrictions of the double proxy approach. In the double proxy approach each of the two proxies must be independent of either the treatment or outcome. This effectively rules out any proxies being directly causally related to both $X$ and $Y$. Figure 5 allows $C$ and $V$ to be causally related to both $X$ and $Y$ and so neither can act as proxy in the double proxy case.

\newtheorem*{P4}{Proposition 4} \begin{P4} The NPSEMs associated with the causal graphs in Figure 5 imply the following conditional independence restrictions:

i. $C\perp \!\!\! \perp(V,Z)|(W,X)$, ii. $V\perp \!\!\! \perp Z|(W,X)$, iii. $X\perp \!\!\! \perp Z|W$, iv. $Y\perp \!\!\! \perp Z|(W,V,X)$, and v. $Y(x)\perp \!\!\! \perp X|(W,V)$. \end{P4}

In this setting Assumption 8 replaces HS Assumption 4.

\theoremstyle{definition} \newtheorem*{A8}{Assumption 8} \begin{A8} For all $x\in\mathcal{X}$ and any $w_1,w_2\in\mathcal{W}$, if $w_1\neq w_2$ then: \[P\big(f_{C|WX}(C|w_1,x)\neq f_{C|WX}(C|w_2,x)\big)>0\] \end{A8}

\theoremstyle{plain} \newtheorem*{T4}{Theorem 3.4 (Auxiliary Proxies)} \begin{T4} Suppose conclusions i., ii., and iii. of Proposition 4 and Assumptions 2, 7, and 8 hold. Then $f_{Z|W}$ is identified up to a reordering of $W$ from the equation below: \[ f_{CZ|VX}(c,z|v,x)=\int_\mathcal{W} f_{C|WX}(c|w,x)f_{W|VX}(w|v,x)f_{Z|W}(z|w)dw \] In addition, if conclusion iv. of Proposition 4 holds, $f_{YWVX}$ is identified (up to a reordering of $W$) from the linear integral equation below: \[ f_{YZVX}(y,z,v,x)=\int_{\mathcal{W}}f_{YWVX}(y,w,v,x)f_{Z|W}(z|w)dw \] Finally, if conclusion v. of Proposition 4 holds then for $\mu_X$-almost all $x_1$: \[ f_{Y(x_1)WX}(y,w,x_2) =\int_{\mathcal{V}} f_{Y|WVX}(y|w,v,x_1)f_{VWX}(v,w,x_2)dv \] Thus $f_{Y(x)WX}$ is identified for $\mu_X$-almost all $x$, up to a reordering of $W$. Causal objects can then be point identified by say, ((ref)) or ((ref)). \end{T4}

Confounder Effects

The results in previous section identify causal effects of the treatment $X$, which are invariant to reordering of $W$. This allows us to avoid HS Assumption 5. However, if we do impose this condition then we can point identify a richer set of objects including causal effects of the latent confounders.

Inspired by a result in Freyberger2021, we provide a weaker condition than HS Assumption 5 that is sufficient to point identify objects involving the quantile rank of each coordinate of $W$. For example, suppose $W$ is a scalar that represents a student's skill at mathematics. Under a weaker condition than HS Assumption 5 we can identify the causal effect of increasing math skill from the 25-th to the 50-th percentile.

Assumption 9 below is weaker than HS Assumption 5. Rather than assume $M[f_{Z|W}(\cdot|w)]$ is equal to $w$, we require only that it is an unknown increasing function of $w$. The condition is similar to, but distinct from, the monotonicity assumption in Freyberger2021 which instead requires that $Z$ can be written as a strictly monotone function of $W$ and some independent noise.

\theoremstyle{definition} \newtheorem*{AN2}{Assumption 9} \begin{AN2} There is a known functional $M$ so that for some (unknown) function $\phi$, $M[f_{Z|W}(\cdot|w)]=\phi(w),\,\forall w\in\mathcal{W}$. $\phi(w)$ has the same length as $W$, each coordinate of $\phi(w)$ depends only on the corresponding coordinate of $w$, and each coordinate is strictly increasing in the corresponding coordinate of $w$. \end{AN2}

The assumptions above allows us to point identify objects that involve the quantile rank of each coordinate of $W$. Let $Q_W$ be the coordinate-wise quantile function of $W$. That is, for a vector $\tau\in[0,1]^{d_W}$, $Q_W(\tau)$ is the vector whose $k$-th coordinate is the $\tau_k$ quantile of the $k$-th coordinate of $W$, where $d_W$ is the dimension of $W$ and $\tau_k$ is the $k$-th coordinate of $\tau$. Assumption 9 allows us to recover say, the density of $Y(x)$ conditional on $W=Q_W(\tau)$ for a known $\tau$.

\theoremstyle{plain} \newtheorem*{T41}{Theorem 4.1} \begin{T41} Suppose the conditions of any of Theorems 3.1-3.4 hold so that we identify densities $\tilde{f}_{Z|W}$, $\tilde{f}_{Y(x)X|W}$, and $\tilde{f}_{W}$, that are equal to $f_{Z|W}$, $f_{Y(x)X|W}$, and $\tilde{f}_{W}$ up to a reordering of $W$. Let $\tilde{W}$ be a random variable with density $\tilde{f}_{W}$ and let $\alpha(w)=M[\tilde{f}_{Z|W}(\cdot|w)]$.

i. Under HS Assumption 5 we have $f_W(w)=f_{\alpha(\tilde{W})}(w)$, and: \[ f_{Y(x_1)X|W}(y,x_2|w)=\tilde{f}_{Y(x_1)X|W}(y,x_2|\alpha^{-1}(w)) \]

ii. Under Assumption 9: \[ f_{Y(x_1)X|W}(y,x_2|Q_{W}(\tau))=\tilde{f}_{Y(x_1)X|W}(y,x_2|q(\tau)) \] Where $q(\tau)=\alpha^{-1}({Q}_{\alpha(\tilde{W})}(\tau))$ for ${Q}_{\alpha(\tilde{W})}$ the coordinate-wise quantile function of $\alpha(\tilde{W})$. \end{T41}

Suppose treatment is binary and potential outcomes have finite absolute first moments. Under the conditions of Theorem 4.1 and HS Assumption 5 we can identify say, the average treatment effect within a given stratum $w$ of $W$:\footnote{Strictly speaking we identify this object for almost all $w$.} \[ E[Y(1)-Y(0)|W=w] \] While Assumption 9 is insufficient to identify the object above, Theorem 4.1.ii shows this condition it does allow us to identify say, the average treatment effect within the stratum of $W$ with quantile rank $\tau$: \[ E[Y(1)-Y(0)|W=Q_W(\tau)] \] For example, we can identify the average treatment effect for individuals with ability in the $25$th percentile.

We can use similar arguments to identify causal effects of the confounders $W$ themselves. Let $Y(x,w)$ denote the potential outcome under a counterfactual in which $X$ and $W$ are respectively set to values $x$ and $w$. Proposition 5 provides conditional independence restrictions involving $Y(x,w)$ which follow from the causal graphs in the previous section. These independence conditions enable us to identify causal effects of the latent factors themselves.

\newtheorem*{P5}{Proposition 5} \begin{P5} The NPSEMs associated with all of the causal graphs in Figures 2, 3.a, 3.b, and 4 imply i. $Y(x,w)\perp \!\!\! \perp (X,W)$. The NPSEMs associated with the graphs in Figure 3.c and Figure 6 imply ii. $Y(x,w)\perp \!\!\! \perp (X,W)|V$. \end{P5}

Under the conditions of any of the theorems in Section 3, and one of the conclusions of Proposition 5, we can identify the joint distribution of $Y(x,w)$ and $X$ up to a reordering of $W$.

\newtheorem*{L1}{Lemma 4.1} \begin{L1} Suppose the conditions of any of Theorems 1, 2, 3 and conclusion i. of Proposition 5 holds, or the conditions of Theorem 4 and conclusion ii. of Proposition 5 hold. Then for $\mu_X$-almost all $x$ and almost all $w$, $f_{Y(x,w)X}$, $f_{Z|W}$, and $f_W$ are identified up to a reordering of $W$. \end{L1}

\theoremstyle{plain} \newtheorem*{T42}{Theorem 4.2} \begin{T42} Suppose the conditions of any of Lemma 4.1 hold so that we identify densities $\tilde{f}_{Z|W}$, $\tilde{f}_{Y(x,w)X}$, and $\tilde{f}_W$ that are equal to $f_{Z|W}$, $f_{Y(x,w)X}$, and ${f}_W$ up to a reordering of $W$. Define $\alpha$ and $q$ as in Theorem 4.1.

i. Under HS Assumption 5: \[ f_{Y(x_1,w)X}(y,x_2)=\tilde{f}_{Y(x_1,\alpha^{-1}(w))X}(y,x_2) \]

ii. Under Assumption 9: \[ f_{Y(x_1,Q_{W}(\tau))X}(y,x_1)=\tilde{f}_{Y(x,q(\tau))X}(y,x_2) \] \end{T42}

If $Y(x,w)$ has bounded derivatives then Theorem 4.2.i allows us to identify say, the average partial effect of an increase in $W$: \[E[\frac{\partial}{\partial w}Y(x,w)]\]

In the context of the educational example, the above isolates the contribution of academic aptitude to a student's GPA. Theorem 4.2.ii identifies the average partial effect of an increase in the quantile rank of $W$: \[E\big[\frac{\partial}{\partial \tau}Y\big(x,Q_W(\tau)\big)\big]\]

The object above is likely to be of only limited use when designing an optimal intervention on $W$. This is because it is uninformative about the size of a change in $W$ required to achieve a given causal impact. However, if $W$ measures ability then we can never hope to design such an intervention and the quantile rank of $W$ may be more readily interpretable than the level of $W$ (see Freyberger2021 for discussion).

Conclusion and Further Comparison with Double Proxies

In this work we establish identification of causal objects using the `triple proxy' approach. We show that there are sets of exclusion restrictions under which we can establish identification (or partial identification) using the triple proxy approach, but not using double proxies. Conversely, there are exclusion restrictions that support the double proxy but not the triple proxy approach. In some settings the exclusion restrictions may allow for both approaches, in this case a comparison of the merits of the two strategies is more nuanced.

One advantage of the triple proxy approach is that it identifies some objects not identified using the double proxy approach. In particular, objects that measure the degree of heterogeneity in treatment effects between strata of the latent variables. Moreover, under some additional conditions, we can adapt our approach to identify causal effects of the latent variables themselves.

An important distinction is that the double proxy approach identifies causal effects from equations involving only observables. This avoids the need to directly specify any modeling assumptions on distributions of the latent factors and simplifies estimation. However, we may wish to impose a priori constraints on the densities of latent factors either to improve precision, or to test these conditions. The double proxy approach precludes this. By contrast, the triple proxy approach is built on equations involving densities of the latent confounders and so we could impose such constraints in estimation.

Simple non-parametric estimators of causal effects are available for the double proxy approach. For example, Deaner2021 suggests an estimator that is similar to sieve two-stage least squares. We leave nonparametric estimation using the triple proxy approach as an open problem. Hu2008 suggest a sieve maximum likelihood method which one could apply to the HS step in our identification results. However, for some choices of the third proxy and conditioning variables, our identification strategy involves a second integral equation. One may be able to apply NPIV methods to estimate a solution to this equation. In the meantime the nonparametric identification results presented here may act as motivation for a parametric estimation strategy or estimation of a discretized version of the model.