Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
57,145 characters · 13 sections · 43 citation commands
Nonparametric Identification and Estimation of Causal Effects on Latent Outcomes
\singlespacing
\thispagestyle{empty}
\doparttoc \faketableofcontents
\setcounter{page}{1}
\doublespacing
Randomized experiments are widely regarded as the most credible design for estimating causal effects. Yet in many substantive applications, the outcome of real interest is not directly observed. Researchers often care about latent constructs such as ideology, state capacity, political trust, social capital, mental health, cognitive ability, or human capital. These quantities are theoretically central, but they are not observed directly. Instead, researchers measure them indirectly through survey items, tests, administrative indicators, behavioral traces, or other proxies. Although the experimental literature has made enormous progress on identification and estimation under random assignment, it has paid far less attention to a basic and pervasive problem: how to conduct causal inference when the outcome itself is latent and must be inferred from multiple imperfect measurements. Existing treatments of experiments typically proceed as if the observed outcome were the outcome of interest, leaving unresolved how one should identify and estimate treatment effects when several noisy measurements are available and none is privileged a priori.
The problem is pervasive across the social sciences. A study of ideology may rely on several roll-call measures or survey batteries; a study of cognitive ability may use different test items; a study of mental health may combine multiple screening instruments; a study of state capacity may draw on several administrative indicators. In each case, the outcome of interest is not any single observed variable, but an unobserved construct that the measurements are intended to capture. In SI (ref), we present a set of representative articles published in political science over the past five years. Faced with multiple indicators, researchers typically aggregate them in some way before conducting causal analysis. Common practices include simple averages, principal components analysis (PCA), inverse covariance weighting (ICW) anderson2008multiple, item response theory model (IRT) stoetzer2022causal, and inverse regression analysis (IRA) zhang2025inverse. These approaches are often useful descriptively, but they are not generally designed to study the average causal effect of treatment on the latent outcome itself.
This paper argues that causal inference with latent outcomes poses distinctive comparability challenges that existing approaches do not adequately resolve. The core difficulty is that latent outcomes have no intrinsic metric. Their empirical meaning depends on how they are measured. Once this is recognized, two overlooked challenges come into view.
The first is a study noncomparability challenge: when two studies measure the same latent construct using different sets of indicators, standard dimension-reduction methods generally do not ensure that the resulting low-dimensional outcomes represent the same quantity. Even if the true average latent treatment effect is identical across studies, the estimated effects may differ simply because one measurement differs across the studies. A difference in estimated treatment effects may therefore reflect differences in measurement rather than differences in causal effects. This undermines the accumulation of knowledge across studies.
The second is a measurement noncomparability challenge within a study. Different measurements may relate to the same latent outcome in different ways. Some measurements may be more sensitive to changes in the latent variable than others; some may be nonlinear transformations of the latent outcome; others may emphasize different regions of the latent distribution. As a result, the observed measurements are not directly commensurable. Current methods either impose strong model specifications (such as IRT models or linear models in SEM) or remain agnostic to the latent structure (such as PCA and ICW). As a result, they are either not robust to model misspecification or inefficient due to ignoring the common latent outcome. A satisfactory framework must therefore solve two challenges at once: it must make different measurements comparable within a study, and it must make causal estimands comparable across studies.
This paper develops a general nonparametric framework for causal inference with latent outcomes that clarifies these issues and provides a constructive solution. Our solution is design-based and centers on two ideas: a benchmark measurement and a measurement bridge function. We assume that across studies there exists at least one common measurement that can serve as a benchmark. We then define, for each additional measurement, a nonparametric bridge function that maps that measurement onto the benchmark in expectation conditional on the latent outcome. In this way, heterogeneous measurements are first transformed so that they carry the same latent information in expectation, and only then are they combined for downstream causal analysis. We show that the bridge function can be characterized and identified within a nonparametric instrumental variables framework. The key insight is that we do not need to rely on external instruments; instead, within the experimental setting, treatment assignment, covariates, and additional measurements can all serve as valid instrumental variables.
This approach has several attractive features. First, it does not require the researcher to impose a specific parametric latent-variable model. The relationship between each observed measurement and the latent outcome may be linear or nonlinear and may remain unknown. What matters is whether one can recover a bridge function that aligns the measurement to the benchmark. Second, because the framework begins from the causal estimand rather than from a dimension-reduction algorithm, it clarifies what the resulting estimate means. The goal is not merely to construct a low-variance index, but to identify and estimate the average latent treatment effect.
The framework also yields substantive guidance for empirical practice. It shows that measurement is not merely a nuisance that can be handled after the fact; it is part of the definition of the causal estimand. It therefore matters how researchers design measurements, which measurements they choose to share across studies, and whether the chosen indicators preserve enough variation in the latent outcome to support identification. In particular, the framework highlights the value of designing measurements that are informative about the latent variable and of including at least one benchmark measurement that can anchor comparisons across settings. These design considerations are especially important when researchers hope to compare experimental findings across populations, interventions, or research teams.
\paragraph{Related Literature.} We contribute to the literature on design-based causal inference with latent outcomes. Although the statistical literature on experimental design and analysis has grown markedly in recent years, the issue of imperfect measurement of latent outcomes has not been systematically discussed, even in otherwise comprehensive textbooks such as angrist2009mostly, pearl2009causality, gerber2012field, imbens2015causal, and hernan2020causal. Recently, stoetzer2022causal and fu2025causal have developed causal inference frameworks for latent outcomes within the potential outcomes framework. They clarify identification assumptions and propose estimators based on different parametric models, including hierarchical item response models zhou2019hierarchical and linear models, respectively. Our approach generalizes these frameworks by adopting a fully nonparametric specification. When latent outcomes are constructed from nonstructural data, the causal inference literature has primarily focused on addressing “learning-induced interference” landy2025causal,egami2022make.
Our identification strategy parallels recent bridge-function approaches in the proximal causal inference literature, where auxiliary variables are used to recover information about unobserved quantities miao2018identifying,miao2024confounding,tchetgen2020introduction. Our contribution is to extend this logic to causal inference with latent outcomes. We propose using measurement bridge functions to align imperfect measurements to a common benchmark, so that downstream causal analysis targets a comparable latent estimand. The bridge function is identified through nonparametric instrumental variables, which constitutes an ill-posed inverse problem and poses challenges for estimation newey2003instrumental,darolles2011nonparametric,horowitz2011applied,chen2025adaptive. The problem becomes more tractable when the target parameter is a linear functional of the nonparametric function chen2026thin,severini2012efficiency,santos2011instrumental,zhang2023proximal. blundell2007semi was among the first to use bounded completeness for the identification of NPIV models and to develop efficient estimation methods for finite-dimensional parameters. Our estimation approach builds primarily on bennett2025inference.
Our paper also relates to a large literature on the measurement of latent variables. A central insight in this literature is that latent constructs have no intrinsic metric and must be defined through their relationships with observed indicators; they are typically measured with error. Classical work in SEM formalizes this idea by specifying a measurement model that links observed outcomes to an unobserved latent variable, typically through factor loadings and measurement errors bagozzi1977structural,sorbom1981structural,kano2001structural. Item response theory (IRT) is also widely used in the social sciences embretson2013item,zhou2019hierarchical,lord2008statistical. In contrast to these explicit modeling approaches, when multiple measurements are available, ICW is a dominant method for efficient hypothesis testing anderson1988structural. Another commonly used approach is PCA hotelling1933analysis. Multiple measurements and factor models are also broadly discussed in the measurement error literature (see the review by schennach2022measurement). We argue that when there is a well-defined latent outcome, our method can yield more informative estimates.
Consider a sample of $n$ units drawn from a superpopulation. Let $Z_i$ denote the binary treatment assignment for unit $i$. The outcome of interest is latent: $\eta_i^1$ is the latent outcome that would be realized if unit $i$ is assigned to receive the treatment, and $\eta_i^0$ is the latent outcome that would be realized in the absence of treatment. We assume that the stable unit treatment value assumption (SUTVA) holds rubin1980randomization.\footnote{The SUTVA assumption is a standard requirement for any study of cause and effect; in this context, it implies that potential outcomes remain stable regardless of which subjects receive treatment, ruling out, for example, spillovers between subjects. Formally, (1) $\eta_i(z_i,z_{-i})=\eta_i(z_i,z'_{-i}) \forall z_{-i}, z'_{-i}$; (2) If $Z_i=z$, then $\eta_i=\eta_i(z)$, $\forall i$ and $z \in Z$.} These latent variables are not directly observed. Examples include ideology and state capacity in political science, preferences and human capital in economics, mental health or cognitive ability in psychology, and socioeconomic status and social capital in sociology. These constructs are conceptually important and constitute the true outcomes of interest.
In order to study them, researchers typically measure latent outcomes by treating them as unobserved constructs and linking them to observable indicators. Formally, researchers design measurement devices, which are distributions indexed by $\eta_i$, $P_j(\cdot|\eta_i)$, to measure the latent outcome. Therefore, each observed measure $Y_{ij}, j=1,2,...,J$, is a random draw from $P_j(\cdot|\eta_i)$. If the distribution is degenerate, the measure is error-free; otherwise, the measurement is imperfect, and we focus on this latter case. A simple and widely used classical specification is $Y_{ij}= \lambda_j \eta_i + \epsilon_{ij}$, where $\lambda_j$ is the factor loading that captures how the latent variable systematically affects the measurement, and $\epsilon_{ij}$ is the measurement error with mean zero, $\mathbb{E}[\epsilon_{ij}]=0$; for example, it may follow a normal distribution $N(0,\sigma^2_{\epsilon_j})$. In this case $Y_{ij}\sim P_j(\cdot|\eta_i)=N(\lambda_j\eta_i,\sigma^2_{\epsilon_j})$. Another widely used model is item response theory: $P(Y_{ij}=1|\eta_i)=\frac{1}{1+exp(-a_j(\eta-b_j))}$, where $\eta_i$ is interpreted as latent ability, and the parameters $a_j$ and $b_j$ capture item discrimination and difficulty, respectively. In our framework, we do not impose a specific functional form; instead, the measurements can have any unknown relationship with the latent outcome.
We are interested in the causal effect of treatment on the latent variable, which stoetzer2022causal and fu2025causal call the latent treatment effect (LTE). The individual-level LTE is defined as $\tau_i=\eta^1_{i}-\eta^0_{i}$. Because it is impossible to observe both potential latent outcomes simultaneously, we focus on the average latent treatment effect (ALTE): $\tau=\mathbb{E}[\eta^1_{i}-\eta^0_{i}]$.
Figure (ref) illustrates the key elements of the framework. Treatment $Z_i$ affects the latent outcome $\eta_i$. The latent outcome can be written as $\eta_i= \mathbb{E}\eta_i^0 + \tau Z_i + \zeta_i$, where $\zeta_i$ captures how individual treatment effects deviate from the average: $\zeta_i=\eta_i^0 - \mathbb{E}\eta^0_i+Z_i[(\eta_i^1-\mathbb{E}\eta^1_i)-(\eta_i^0-\mathbb{E}\eta^0_i)]$. This term has mean zero by construction, $\mathbb{E}\zeta_i=0$. The measurements $Y_{ij}$ are connected to the same latent outcome. As long as the measurement is not a deterministic function of the latent outcome, we expect the presence of measurement error $\epsilon_{ij}$. Throughout the paper, we assume that each measure $Y_{ij}$ is not independent of $\eta_i$. In other words, the observed measurements indeed capture some aspects of the latent outcome. How to design informative measurement devices is itself an important topic; however, in this study, we primarily focus on the downstream question of how to identify and estimate causal effects on latent outcomes.
For ease of exposition, we omit $i$ when it does not create confusion.
The latent outcome is not directly observed and has no intrinsic metric. It depends heavily on the observed measurements. For example, unobservable mathematical ability is primarily reflected in observed test scores, while unobservable ideology is primarily reflected in observed roll-call votes. This dependence on measurement creates important, yet long-overlooked, challenges in studying causal effects on latent outcomes. In this section, we illustrate two noncomparability challenges arising from the nature of latent outcomes.
Latent outcomes must be measured. It is common for researchers to use different approaches to measure the same latent variable due to differing theoretical considerations or feasibility constraints. There is rarely consensus on a single criterion for measuring a latent outcome. However, the use of different measurements creates a serious challenge: even when researchers seek to identify average treatment effects on the same latent outcome across studies, the resulting estimates are often not comparable. We refer to this issue as the Study Noncomparability Challenge.
Consider two studies that target the same latent outcome variable $\eta$. Suppose two research teams apply largely the same set of measurements, except that one measurement differs. Let $(Y_1,Y_2,...,Y_J)$ denote the measurements used in the first study and $(Y_1,Y_2,\ldots,Y'_J)$ those used in the second study, where $Y_J\sim P_j(\cdot\mid\eta_i)$ while $Y'_J\sim \tilde{P}_j(\cdot\mid\eta_i)$, and the two distributions are assumed to differ.
Given multiple measurements, researchers typically apply dimension-reduction techniques, including simple averaging, PCA, ICW, IRT, IRA. We denote these methods collectively as a mapping $M: \mathbb{R}^J \rightarrow \mathbb{R}$, which produces a single variable $\tilde{Y}_A=M(Y_1,\ldots,Y_J)$ and $\tilde{Y}_B=M(Y_1,\ldots,Y'_J)$ for each individual in studies A and B, respectively. The resulting low-dimensional variables ($\tilde{Y}_A$ and $\tilde{Y}_B$) are then treated as outcomes in downstream causal inference in the two studies.
The Study Noncomparability Challenge arises because the mapping $M$ does not guarantee that the same low-dimensional variable is generated when the underlying measurements differ. Because we focus on ALTE, the relevant form of the Study Noncomparability Challenge is noncomparability at the level of expectations: $\mathbb{E}[\tilde{Y}_A|\eta_i] \neq \mathbb{E}[\tilde{Y}_B|\eta_i]$. Even though $\tilde{Y}_A$ and $\tilde{Y}_B$ are intended to capture the same latent outcome $\eta$, they may in fact encode different information.
Suppose two studies investigate the same treatment $Z_i$. Even if the true ALTE, $\mathbb{E}[\eta^{1}-\eta^0]$, is identical across the two studies, the difference in the $J^{th}$ measurement can lead to different estimated average causal effects: $\mathbb{E}[\tilde{Y}_A^1-\tilde{Y}_A^0] \neq E[\tilde{Y}^1_B-\tilde{Y}^0_B]$. If the treatments differ across studies, the resulting average causal effects become even less comparable and may be misleading. For example, if the estimated treatment effect in study A is larger than in study B, this difference may simply reflect the use of measurement procedures that do not account for discrepancies in the $J^{th}$ measurement, even when the true effects are identical or even reversed. This phenomenon is problematic because the estimated causal effect depends on the measurement choices rather than the underlying causal quantity. From the perspective of scientific inquiry, this undermines the accumulation of knowledge across studies.
The root cause of this problem is that latent variables have no intrinsic metric—their meaning depends on the measurements used to define them. Existing methods typically standardize and combine these measurements in an ad hoc manner, which in turn alters the interpretation of the latent outcome depending on the scaling and composition of the observed variables.
In addition to study-level noncomparability, there is another often overlooked issue: even within a single study, different measurements may not be comparable to one another. In practice, researchers frequently employ multiple measurements for an abstract latent outcome because they believe that different measures capture distinct aspects of the latent variable. For example, mathematical ability can be assessed using algebra, geometry, and analysis questions, as these are thought to capture different dimensions of mathematical ability. Alternatively, one may view mathematical ability as being defined by multiple subcomponents. We call this the Measurement Noncomparability Challenge: $\mathbb{E}[Y_{j}|\eta] \neq \mathbb{E}[Y_{k}|\eta]$ for $j \neq k$.
There are two broad responses. First, researchers explicitly model the relationship between the latent outcome and the observed measurements. For instance, in a standard structural equation model (SEM), researchers assume a linear model $Y_{j}= \lambda_j \eta + \epsilon_{j}$, where the factor loadings $\lambda_j$ are allowed to differ across measurements. Similarly, in the IRT model, $P(Y_{ij}=1|\eta_i)=\frac{1}{1+exp(-a_j(\eta-b_j))}$, each item has its own discrimination parameter $a_j$ and difficulty parameter $b_j$. The goal is to recover these finite-dimensional parameters and thereby identify the latent outcome. A key concern, however, is that the model may be misspecified, and therefore, the result is not robust.
The second strand of the literature does not impose a parametric measurement model. Instead, it does not attempt to recover the latent outcome directly, but rather constructs a lower-dimensional variable based on a chosen criterion. For example, PCA aims to obtain a weighted average of the measurements that maximizes variance among all possible linear combinations, known as the first principal component score, which is then used as the outcome variable. In contrast, ICW combines measurements so as to minimize variance. Because it is used as an outcome variable, lower variance implies higher statistical power. However, neither approach is designed specifically for causal inference: PCA is primarily used for dimension reduction, while ICW is typically motivated by hypothesis testing rather than the estimation of causal effects. When researchers posit a latent structure—namely, that the measurements are intended to capture the same latent outcome, as illustrated in Figure (ref)—ignoring this information would incur unnecessary efficiency loss.
A satisfactory estimator must therefore address two challenges simultaneously: it must render different measurements comparable within a study while remaining robust to model misspecification and preserving information, and it must ensure that causal estimands are comparable across studies.
First, because our object of interest is the ALTE and the multiple measurements are designed to capture the same latent outcome, it is desirable to extract a common latent signal from these measurements. Ignoring the latent structure would result in a loss of useful information. Moreover, because the latent outcome is measurement-dependent, ignoring the latent structure may exacerbate the study noncomparability challenge. However, measurement noncomparability suggests that we require a bridge function that transforms different measurements so that they convey the same information, at least in expectation, before extracting the common latent signal. Formally, we require a measurement bridge function $\varphi_j$ for measurement $Y_j$, so that $\mathbb{E}[\varphi_j(Y_j) \mid \eta] = \mathbb{E}[Y_1 \mid \eta]$. Therefore, conditional on latent outcome $\eta$, all measurements on expectation are equal to $Y_1$. $Y_1$ here is chosen arbitrarily.
Next, to address the study noncomparability challenge, the dimension-reduction mapping $M$ must preserve the common information shared across different sets of measurements. This is a strong requirement, as it must hold for all possible measurement sets, while the latent outcome is, in principle, defined through multiple measurements. Rather than imposing strong restrictions on the mapping $M$, we propose a design-based approach. Specifically, we require the existence of a common measurement across different studies. Without loss of generality, we denote this as the first measurement, $Y_1$. This variable serves as a bridge linking different sets of measurements.
The following flow chart (ref) summarizes our proposed method. First, we select a benchmark variable, denoted by $Y_{1}$. This measurement serves to link different sets of measurements. Next, we identify a measurement bridge function $\varphi_j$ that transforms other measurements $Y_j$, $j \neq 1$, such that
This bridge function maps each measurement onto the reference measurement $Y_{1}$ so that all measurements contain the same information in expectation. Because $Y_j$ and $Y_1$ are functions of the same latent variable $\eta$, it is natural to expect that such a bridge function exists; we will formalize this in the next section. If such a bridge function does not exist, it is generally impossible to extract the same information from different measurements.
Because the latent outcome has no intrinsic scale or unit, we interpret it using the same metric of $Y_1$. In other words, we treat $Y_1$ and the nonparametrically transformed measurements as outcome variables. We then use them for downstream causal analysis. For example, we can combine the transformed measurements through a weighted average, $\tilde{Y} = \sum_{j=1}^J \omega_j \varphi_j(Y_j)$, with weights satisfying $\sum_{j=1}^J \omega_j = 1$. The optimal weights that minimize variance can be used. We then treat $\tilde{Y}$ as the new observed outcome to estimate average treatment effects.
Under this method, all measurements within a single study contain the same information in expectation and can therefore be jointly used for causal inference. Moreover, consider two studies, A and B, that target the same latent variable $\eta$ but use different sets of measurements. As long as they share at least one common variable—say $Y_1$—the resulting transformed variables $\varphi_j(Y_j)$ in both studies are mean-comparable.
The idea of using one measurement as a benchmark, developed in the 1960s and 1970s, has largely been overlooked in recent work. In the SEM framework, $Y_{j}= \lambda_j \eta + \epsilon_{j}$, and because the latent outcome has no intrinsic scale, identification typically relies on normalization. There are generally two approaches. One is to impose that $\eta$ has mean zero and unit variance; the other is to set $\lambda_1=1$. When $\lambda_1=1$, this assigns the scale of the first measurement to the latent outcome. This normalization is without loss of generality because the latent outcome has no intrinsic scale. Under this normalization, the latent outcome is interpreted on the same scale as the first measurement. Although this idea is well known in the SEM literature, current practice typically adopts the first normalization, which can generate noncomparability challenges. fu2025causal illustrates this issue and proposes an optimal estimator under the linear model specification, called (W)egihted (S)caled (I)ndex, WSI. We generalize their estimator to the nonparametric setting. Because the measurements are (N)onparametrically (S)caled to form a comparable (I)ndex, we refer to this approach as NSI.
For the proposed method to be meaningful—so that $Y_1$ captures the latent variable $\eta$—we can impose the following centering assumption.
This assumption implies that the measurement provides an unbiased representation of the latent outcome. Many methods have been proposed to obtain unbiased latent variable from biased measurement vishnubhatla2026proxy. If multiple centering measurements exist, the choice of benchmark variable is inconsequential. We emphasize that only one centering measurement is required, and we allow the remaining measurements to have arbitrary (possibly nonlinear) relationships with the latent variable $\eta$. For these non-centering measurements, we use a bridge function $\varphi_j$ to transform them so that they are centered in the same way as $Y_1$.
To address the challenge of noncomparability, we propose NSI using a benchmark variable and measurement bridge functions. This section introduces the identification assumptions and estimation techniques.
Different measurements may have different relationships with the same latent outcome. A measurement bridge function is therefore required to recover a common representation of the latent outcome. Under what conditions does such a bridge function exist? As we will show later, this problem can be formulated as a nonparametric instrumental variables (NPIV) problem.
It is convenient to approach this problem from a functional analysis perspective. Define the Hilbert spaces $H_1 = L^2(F(Y_{ij}))$ and $H_2 = L^2(F(\eta_i))$, with inner product $\langle h_1, h_2 \rangle = \mathbb{E}[h_1 h_2]$, which denote the spaces of all square-integrable functions with respect to the cumulative distribution function $F$, respectively. Let $K$ be a conditional expectation operator mapping $H_1$ to $H_2$, defined by $K\varphi=\mathbb{E}[\varphi(Y_{ij})|\eta_i]$, and define $r(\eta_i) = \mathbb{E}[Y_{i1} \mid \eta_i]$. Then the condition $\mathbb{E}[\varphi_j(Y_{ij}) \mid \eta_i] = \mathbb{E}[Y_{i1} \mid \eta_i]$ can be written compactly as $K\varphi = r$. This equation is known as a Fredholm integral equation of the first kind.
From functional analysis, the measurement bridge function $\varphi$ exists if and only if $r$ lies in the range of the operator $K$. In general, this condition is not easy to verify. The common sufficient condition that has practical implications is the completeness condition.
This assumption first implies that $\eta_i$ and $Y_{ij}$ are not independent. Intuitively, it requires that the measurement $Y_{ij}$ captures the full information (variability) of the latent outcome $\eta_i$. It accommodates both categorical and continuous variables. For example, if $\eta_i$ is discrete with $k$ categories and $Y_{ij}$ has $m$ categories, then we must have $m \ge k$. To see this, $$ \underbrace{
}_{A} \underbrace{
}_{g} =
$$ Completeness ($Ag=0 \Rightarrow g=0$) requires that the matrix $A$ has full rank, which can only hold if $m \ge k$. Many models, such as those in the exponential family, satisfy completeness.
This issue also arises in the proximal causal inference literature. Let $(\mu_n,\varphi_n,\psi_n)$ denote the singular system of $K$. The following result, adapted from miao2018identifying, shows that completeness implies the existence of $\varphi$.
Completeness has important implications for the design of measurements. For this reason, we do not recommend using measurements that lose information (such as discretized or categorical outcomes). Instead, we encourage researchers to design measures that plausibly exhibit a monotonic relationship with the underlying latent variable of interest, as this provides a simple way to satisfy the completeness assumption.
Given that the measurement bridge function exists, we must ensure that it is identified. As in the proximal causal inference literature miao2024confounding, identifying $\varphi_j$ requires auxiliary information. In general, we need additional variables that serve as instrumental variables; then $\varphi_j$ can be identified through a standard NPIV approach.
From the proposition, we note that candidate instrumental variables should first satisfy mean independence conditional on the latent outcome. One sufficient (though not necessary) condition is that the instrumental variables are independent of the measurement errors in both $Y_{i1}$ and $Y_{ij}$. The second requirement, completeness, is the nonparametric analogue of the order condition in the parametric setting darolles2011nonparametric,horowitz2011applied,newey2003instrumental. If $W_i$ and $Y_{ij}$ are discrete with finite support, then the support of $W_i$ must be at least as rich as that of $Y_{ij}$. In our context, candidate instrumental variables include treatment assignment $Z_i$, measurements other than $Y_{i1}$ and $Y_{ij}$, and pre-treatment covariates $X_i$.
Once $\varphi_j$ is identified, we obtain transformed measurements $\tilde{Y}_j = \varphi_j(Y_{ij})$. In this case, we have multiple outcome variables, so that the ALTE is over-identified. There are several ways to utilize these multiple measurements. For example, we can combine the $J$ transformed measurements to form a single “observed” outcome variable $\tilde{Y}_i=\sum_{j=1}^J \omega_j \tilde{Y}_j$, where the weights satisfy $\sum_{j=1}^J \omega_j=1$. To achieve minimal variance, we propose using inverse-variance weights, $\omega_j=\frac{\Sigma^{-1}I_J}{I^{'}_J\Sigma^{-1}I_J}$, where $\Sigma$ denotes the variance-covariance matrix of the transformed outcome. With the optimally weighted outcome $\tilde{Y}_i$, we can then conduct standard causal inference. The average LTE is identified as $\mathbb{E}[\tilde{Y}_i \mid Z = 1] - \mathbb{E}[\tilde{Y}_i \mid Z = 0]$. Alternatively, a more convenient approach is to use optimal GMM to estimate the ALTE.
To estimate our target parameter, the ALTE, we first need to estimate the nonparametric measurement bridge functions. This is an NPIV problem (i.e., a conditional moment restriction with endogeneity). Once these bridge functions are estimated, we can estimate the ALTE through over-identified unconditional moment restrictions. In practice, the nuisance function defined by NPIV may be weakly identified. bennett2025inference refer to this issue as follows: "That is, the conditional moment restrictions can be severely ill-posed as well as admit multiple solutions (if our Proposition (ref) fails), so that the nuisance functions are not uniquely identified and are not stable as underlying distributions vary." An important observation is that our primary target parameter is not the measurement bridge function itself; rather, our interest lies in the ALTE, which is a linear functional of the nuisance bridge function. Even if the bridge function is weakly identified, the target parameter can still be strongly identified, allowing for asymptotically normal estimation at the $\sqrt{n}$ rate severini2012efficiency,santos2011instrumental.
We therefore adopt the estimation framework proposed by bennett2025inference. They develop methods for estimation and inference of continuous linear functionals of nuisance functions defined by conditional moment restrictions, particularly in settings with an NPIV first stage. We sketch the main idea here and provide full details in the SI (ref).
Define $s(Z_i)=\frac{Z_i}{\pi}-\frac{1-Z_i}{1-\pi}$, where $\pi = \Pr(Z_i = 1)$, as the Horvitz–Thompson transform. Then, our target parameter, the ALTE, can be expressed as a continuous linear functional of the measurement bridge functions:
Following bennett2025inference, we assume that the ALTE is strongly identified in the sense that the following assumption holds.
This assumption essentially impose on the smoothness of ALTE $\tau=\mathbb{E}[s(Z_i) \phi_j(Y_j)]$ with respect to the NPIV conditional expectation operator $K$. Similar to the DML literature chernozhukov2018double, the key step is to construct a Neyman orthogonal score for the ALTE. The Neyman orthogonal score is constructed by introducing a debiasing nuisance function $q^*_j(W)=\mathbb{E}[\xi_0(Y_j)\mid W]$. Then, the score for the ALTE is given by $ \psi_j=s(Z_i) \phi_j(Y_j)+q_j(W)(Y_1-\phi_j(Y_j))-\tau $. There are three parameters in the score that must be estimated. To avoid overfitting bias, all nuisance functions are estimated via cross-fitting. Let the sample be split into $K$ folds $I_1,\ldots,I_K$. For each fold $k$, we estimate the bridge function using a penalized minimax criterion: $$ \hat\phi^{(-k)}_j = \arg\min_{\phi\in\Phi_n} \sup_{q\in\mathcal{Q}_n} \mathbb{E}_{n,-k} \Bigl[ \bigl(\phi(Y_j)-Y_1\bigr)q(W) -\frac{1}{2}q(W)^2 + \mu_n \phi_j(Y_j)^2\Bigr]-\gamma^q_n||q||^2_Q+\gamma^\phi_n ||\phi_j||^2_\Phi $$ where $\mathbb{E}_{n,-k}$ denotes the empirical average over the observations in $\mathcal{I}^c_k$, $\|\cdot\|_\Phi$ and $|q||^2_Q$ are penalty norms, and $\mu_{n},\gamma^{q}_n,\gamma^{\phi}_n$ are regularization hyperparameters. $\hat{\xi}^{-k}_j$ and $\hat{q}^{-k}_j$ are estimated similarly.
Once the nuisance parameters are estimated, we estimate the ALTE $\tau$ using GMM, as we have $J$ moment conditions. Define the cross-fitted moments \[ \hat m_{ij} =
\] where $k(i)$ denotes the fold containing observation $i$. Let $ \bar{\hat m}_n = \frac{1}{n}\sum_{i=1}^n( \hat m_{i1}, \dots, \hat m_{iJ})^T$. Then, the overidentified GMM estimator is \[ \hat\tau_{\mathrm{GMM}} = \mathop{\mathrm{arg\,min}}_{\tau\in\Theta} \bigl(\bar{\hat m}_n-\tau\mathbf{1}_J\bigr)^\top \hat\Omega_n^{-1} \bigl(\bar{\hat m}_n-\tau\mathbf{1}_J\bigr), \] where $\hat\Omega_n$ is a consistent estimator of the weighting matrix. The optimal weighting matrix is $\Omega=Var(\psi_i)$, where $\psi_i=(\psi_{i1},...,\psi_{iJ})$ is the stacked score vector.
bennett2025inference show that, for each $j$, under regularity assumptions, first-stage estimation does not affect the second stage, $$ \sqrt{n}\left(\frac{1}{n}\sum_{i=1}^n \hat m_{ij} - \tau_0\right) = \frac{1}{\sqrt{n}}\sum_{i=1}^n \psi_{ij} + o_p(1), \qquad j=2,\ldots,J. $$ It follows that, for any fixed $J<\infty$, $\sqrt{n}(\bar{\hat m}_n-\tau\mathbf{1}_J) \rightarrow_d N(0,\Omega)$, and $\hat{\tau}_{\mathrm{GMM}}$ is also asymptotically normal. We provide detailed estimation procedures with control variables in the SI (ref). All implementation details are available in the accompanying R package.
We simulate two studies in which both studies share the same latent untreated outcome and the same latent treatment effect, so the true average latent treatment effect is common across studies. In each replication, we draw a scalar covariate $x_i$, set $\eta_{0i}=0.9x_i+0.3x_i^2+u_i$, and let the treatment effect be $\tau_i=0.8+0.3x_i$, so that $\eta_{1i}=\eta_{0i}+\tau_i$. The target parameter is therefore the common ALTE, which is approximately $0.80$ in the simulation. Study 1 observes three measurements $(Y_1,Y_2,Y_3)$, where $Y_1$ is the common benchmark and $Y_2,Y_3$ are noisy nonlinear functions of the latent outcome (to be precise, it is a polynomial function with quadratic and cubic terms). Study 2 keeps the same benchmark $Y_1$ but replaces the other two measurements with $(Y_2',Y_3')$, which are differently coded nonlinear transformations of the same latent variable. Thus the latent causal effect is identical across studies, but the measurement system differs. Because the outcome is continuous (IRT model is not applicable), we compare PCA, ICW, WSI, and our NSI estimator.
In $1000$ Monte Carlo simulations with sample size $n=800$ per study, we compute the difference in ALTE across two studies and conduct a Wald test of equality between the two estimates. The results are reported in Table (ref). The average cross-study gap is $0.256$ for PCA and $0.366$ for ICW. PCA rejects the true null hypothesis of equal ALTE in $24.7\%$ of replications, while ICW rejects it in $100\%$ of replications. As expected, because PCA and ICW are measurement-dependent methods, the resulting estimates are also measurement-dependent and therefore not comparable, even when the latent outcome and ALTE are identical across studies.
WSI, in principle, can avoid the noncomparability problem, but it imposes a linearity assumption on the bridge function. The average cross-study gap decreases to $0.072$, and the rejection rate for equal ALTE is only $1.2\%$. This represents a substantial improvement over PCA and ICW. NSI relaxes the parametric assumption on the bridge function. It achieves the lowest cross-study gap and the lowest rejection rate among all methods.
It is also worth noting that nonparametric estimators are typically unstable, requires larger sample size, and rely heavily on assumptions about the underlying function space governing the relationship between measurements and the latent outcome. In practice, researchers rarely employ measurements with highly nonlinear relationships to the latent outcome. Therefore, in practice, it is often advisable to begin with linear methods, WSI, as advocated by fu2025causal.
To illustrate the proposed estimator, we revisit the field experiment studied in kalla2020reducing, which examined whether door-to-door canvassing changes attitudes toward undocumented immigrants. We use this application because it contains multiple outcome measures that are intended to capture a common latent construct and because the design includes rich experimental variation that is useful for identification. Specifically, the experiment includes two treatment indicators---a full treatment and an abbreviated treatment---and multiple post-treatment outcome measures concerning immigration attitudes and policy views. fu2025causal use this application to illustrate the WSI estimator under a linear measurement model. Here we use the same application to illustrate our more general nonparametric bridge-function approach.
The empirical setting is especially useful for our purposes for three reasons. First, as shown in the figure (ref), the study contains two distinct post-treatment outcome scales, one summarizing respondents' attitudes toward undocumented immigrants ($Y_1$) and the other summarizing their views on immigration-related public policy ($Y_2)$. These two scales are designed to measure the same underlying latent disposition, but they need not be linearly related to one another. Second, the experiment includes two randomized treatment indicators, which provide valid sources of exogenous variation for identifying bridge functions. Third, the data also include a pre-treatment covariate, constructed from baseline attitudes. This variable is predictive of the latent outcome and can be incorporated both to improve precision and to enrich the instrument set used in the first stage. More details about the indices can be found in SI (ref).
Let $Y_1$ denote the benchmark outcome and $Y_2$ denote the auxiliary measurement. We estimate the bridge function mapping $Y_2$ into the scale of $Y_1$ using three alternative basis constructions: a polynomial series basis, a random-forest basis, and an RKHS basis. For each basis, we then form the debiased Bennett estimator of the treatment effects on the latent outcome. To facilitate comparison with the earlier linear approach, we also report the WSI estimator. They show that the measurements follow a linear pattern, and therefore, we may expect the estimates under nonparametric and linear assumption to be similar.
Table (ref) reports the results. We use the two treatment assignment and baseline covariate as IVs in the first stage to estimate the bridge function. Panel A presents specifications in which the latent outcome is regressed on the two treatment indicators only. Panel B adds the baseline covariate to the second-stage regression. Across all three nonparametric implementations, the estimated effect of the full treatment is positive and statistically distinguishable from zero at conventional levels in most specifications. The point estimates are also quite stable across basis choices, ranging from $0.418$ to $0.444$ without covariate adjustment and from $0.398$ to $0.425$ with covariate adjustment. The abbreviated treatment has a smaller estimated effect and is not statistically distinguishable from zero in either panel. These findings mirror the substantive conclusion in fu2025causal: the full canvassing intervention produces a meaningful shift in the latent immigration-attitude outcome, whereas the abbreviated treatment does not.
The comparison with WSI is also informative. In Panel A, the nonparametric estimates of the full-treatment effect are somewhat larger than the WSI estimate, though all are similar in magnitude. In Panel B, after adjusting for the baseline covariate, the nonparametric and WSI estimates remain close, but the WSI estimator is considerably more precise. This pattern is sensible. The WSI estimator imposes a linear measurement structure and is therefore more efficient when that approximation is adequate, whereas the Bennett estimator relaxes linearity and protects against measurement noncomparability at the cost of additional variance. In this application, the close correspondence between the two sets of point estimates suggests that the linear approximation used by WSI is not badly misspecified, while the nonparametric estimates show that the substantive conclusion does not depend on imposing linear bridge functions.
Overall, this application illustrates the practical value of the proposed approach. The nonparametric estimator allows researchers to carry over the basic design logic of linear model (e.g. WSI estimator) to settings where the relationship between outcome measures may be nonlinear and unknown. At the same time, the empirical results show that, in this application, the more flexible estimator delivers conclusions that are substantively consistent with the linear WSI benchmark. This combination of robustness and interpretability is especially valuable in applications where multiple imperfect measurements are available but their functional relationship to the latent outcome is uncertain.
Causal inference in experiments is often framed as a problem of treatment assignment, identification, and estimation, taking the outcome as given. This paper argues that when the outcome of interest is latent, measurement is itself part of the causal inference problem. Because latent outcomes have no intrinsic metric, treatment effects on them are only meaningful relative to a measurement system. Once this point is recognized, two distinct comparability problems arise. Across studies, different sets of indicators may lead standard methods to target different empirical quantities even when the underlying latent treatment effect is the same. Within a study, different indicators may have different and possibly nonlinear relationships with the same latent outcome, so they are not directly commensurable.
We develop a general nonparametric framework that addresses both problems. Our approach uses a benchmark measurement and measurement bridge functions to align indicators to a common scale in expectation. This produces transformed outcomes that are comparable within studies and, when a common benchmark is shared, comparable across studies as well. The framework is design-based, allows the relationship between measurements and the latent outcome to remain unknown and nonlinear, and connects identification to a nonparametric instrumental variables problem that can be solved using treatment assignment, covariates, and additional measurements as instruments. It therefore shifts attention from ad hoc dimension reduction toward identification and estimation of the average latent treatment effect itself.
More broadly, the paper suggests that measurement design should be treated as part of experimental design. Researchers who wish to compare results across studies should, when possible, include at least one common benchmark measurement. Researchers who collect multiple indicators within a study should consider whether those indicators preserve enough information about the latent outcome to support bridge-function identification.
Several limitations and extensions remain for future work. The nonparametric approach relies on completeness-type conditions and can be less stable in finite samples than simpler linear procedures. Future research could investigate weaker identification strategies, improved regularization and efficiency in finite samples, and extensions to settings with richer treatment regimes, longitudinal measurements, or interference. Even with these limitations, the central message is clear: when outcomes are latent, causal inference cannot proceed as if measurement were secondary. Measurement choices shape the empirical meaning of the estimand, the comparability of results, and the credibility of substantive conclusions. By placing measurement at the center of design-based causal inference, this paper aims to provide a foundation for more interpretable, comparable, and robust experimental research on latent outcomes.