Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Causal Inference for Experiments with Latent Outcomes: Key Results and Their Implications for Design and Analysis
\singlespacing
abstractHow should researchers analyze randomized experiments in which the main outcome is latent and measured in multiple ways but each measure contains some degree of error? We first identify a critical study-specific noncomparability problem in existing methods for handling multiple measurements, which often rely on strong modeling assumptions or arbitrary standardization. Such approaches render the resulting estimands noncomparable across studies. To address the problem, we describe design-based approaches that enable researchers to identify causal parameters of interest, suggest ways that experimental designs can be augmented so as to make assumptions more credible, and discuss empirical tests of key assumptions. We show that when experimental researchers invest appropriately in multiple outcome measures, an optimally weighted scaled index of these measures enables researchers to obtain efficient and interpretable estimates of causal parameters by applying standard regression. An empirical application illustrates the gains in precision and robustness that multiple outcome measures can provide.
\thispagestyle{empty}
\doparttoc
\faketableofcontents
\setcounter{page}{1}
\doublespacing
Introduction
bibunitSocial scientists often seek to estimate the average causal effect of interventions and increasingly rely on experimental designs to do so convincingly. The statistical literature on experimental design and analysis has grown markedly in recent years.
However, one topic of special concern to social scientists has largely escaped attention, even from otherwise comprehensive textbooks: imperfect measurement of experimental outcomes. angrist2009mostly, gerber2012field, and imbens2015causal offer no sustained formal treatment of outcome measurement or how to analyze experiments in which outcomes are measured in more than one way. The models that these textbooks present implicitly assume that the observed outcome is the true underlying potential outcomes of interest.
In the social sciences, however, there is often slippage between the outcomes of interest and the proxy outcome measures at hand. Constructs such as economic inequality, press freedom, political violence, corruption, and post-materialism are just a few examples of latent variables that are thought to be measured imperfectly, despite the sustained efforts of researchers. To illustrate how frequently this issue arises, SI (ref)
lists 14 representative articles
published in the American Political Science Review during the past five years.
Each article uses multiple measures to gauge a latent outcome, but statistical practice varies widely: regression using additive indices constructed from standardized measures, regression using indices based on indices derived from principal components analysis or from inverse covariance algorithms, or nonlinear estimators rooted in item response theory. Given multiple imperfect measurements, how should researchers identify and estimate average treatment effects on the latent outcome?
Our paper builds on recent attempts to formalize the identification and estimation challenges that may arise when latent outcomes are measured with error. Like stoetzer2022causal, we use potential outcomes notation to unify our discussion of experimental design and outcome measurement. We make four contributions, each of which has important implications for research design and estimation.
First, we identify and address a critical study-specific noncomparability problem in existing methods for handling multiple measurements: principal components analysis (PCA), inverse covariance weighting (ICW) anderson2008multiple, item response theory model (IRT) stoetzer2022causal, and inverse regression analysis (IRA) (zhang2025inverse). The estimands in these approaches are often distorted by the estimation techniques themselves, such that even when researchers aim to identify the average treatment effects on the same latent outcome across different studies, the resulting estimates are not comparable. This drawback poses a serious obstacle to any sustained research program.
Consider, for example, a scenario in which researchers design a large experiment to estimate the treatment effect on political attitudes using several survey-based measures. With any of the aforementioned methods, they would obtain an estimate. Meanwhile, another research team might investigate the same question using slightly different measures and obtain another estimate. However, these results would not be comparable and would not, in fact, target the same latent construct of political attitude. The key reason is that a latent variable has no intrinsic metric -- its meaning is entirely determined by the measurements used to define it. Widely used methods arbitrarily standardize and combine these measurements, which in turn alters the interpretation of the latent outcome depending on the scale of the observed variables.
Building on a largely overlooked literature dating back to at least costner1971utilizing and alwin1974causal, we propose an efficient design-based method for identifying latent treatment effects without standardizing the latent outcome in ways that render its units of measurement study-specific. As a result, estimated effects are more directly comparable across studies.
Second, adapting ideas from prior work by bagozzi1977structural, sorbom1981structural, and kano2001structural to a potential outcomes framework, we propose an optimally weighted scaled index (WSI) estimator for the experimental intervention’s average treatment effect on a latent outcome. As explained in section (ref), it involves two key steps: (1) scaling the outcomes to address the study-specific noncomparability problem and (2) optimally weighting the scaled outcomes.
As demonstrated through a series of simulations, the WSI achieves higher statistical power than existing alternatives (PCA, ICW, IRT, IRA). More importantly, under optimal weighting, it attains similar efficiency to traditional structural equation modeling (SEM) estimators, which typically use maximum likelihood estimation based on strong normality assumptions.
WSI also facilitates visualization of estimated causal effects and randomization inference for hypothesis testing.
Third, we show analytically and through simulation how multiple measures of a latent outcome can improve the precision with which this average treatment effect is estimated in section (ref). When conducting experiments, researchers often face a key design dilemma: given additional resources, should they increase the sample size or collect more measurements per unit? Our framework provides a useful decision rule, showing that the optimal choice depends on the reliability of the outcome measures. When reliability is low, the marginal gain from collecting an additional measurement is relatively high relative to the gains from increasing the number of subjects. In contrast, when reliability is high, the marginal gain from additional measurements diminishes, and increasing the sample size becomes preferable. These results highlight an underappreciated practical trade-off between gathering more observations and gathering additional outcome measures.
Fourth, although our method (like any method) rests on assumptions that are not directly testable, the credibility of these assumptions hinges on design choices. This stands in contrast to existing approaches, which typically separate the design and estimation stages. For instance, we show how certain data collection strategies can help satisfy what would otherwise be unrealistic measurement assumptions.\footnote{Our approach to data collection also reduces reliance on constant treatment effects among all subjects, such as the hierarchical Item Response Theory (IRT) model proposed by stoetzer2022causal. We offer design recommendations that allow the analyst to stay within a simple modeling framework that accommodates heterogeneous treatment effects.}
We are by no means the first to note that “invalid” measurements threaten unbiased inference or that redundant outcome measurements can improve precision; our contribution is to bring a coherent analytic framework to the discussion of experimental research design and outcome measurement so as to clarify the value of research designs that allocate resources to the collection of multiple measurements of a latent outcome. We also show that the latent outcome models may be viewed as nested alternatives to seemingly unrelated regression (SUR), allowing researchers to assess empirically whether the constraints imposed by the latent outcome model are consistent with the data. If not, researchers can fall back on the more agnostic SUR approach in which each individual outcome measure is regressed on the treatment.
This paper is structured as follows. We begin by presenting a potential outcomes model that defines the target of inference and allows for some degree of slippage between the latent outcome we seek to measure and the observed measures at our disposal. Next, we show formally the conditions under which the average treatment effect on the posited latent outcome is identified. When latent outcomes are measured with random error, this average treatment effect may be estimated consistently, but precision is enhanced when the researcher gathers multiple outcome measures. When measurement errors are systematic -- especially when mismeasurement operates differently for subjects assigned to treatment or control -- methods that are typically used to analyze experiments may no longer yield unbiased estimates. We show how identification problems may be diagnosed empirically and addressed by diversifying the portfolio of measures and adopting an estimator that leverages the additional measures in a theoretically informed way. A simulated example illustrates the mechanics of this approach. In order to show its relevance to applied empirical work, we reanalyze the kalla2020reducing experiment, which features a diverse array of outcome measures that improve precision and facilitate robustness checks.
\section{Framework}
Consider a sample of $n$ units drawn from a superpopulation.\footnote{The finite population result is similar to the superpopulation case, except that variance estimation is conservative, in keeping with the standard causal inference literature imbens2015causal.} Let $Z_i$ denote the treatment assignment for unit $i$. Generalizing the potential outcomes framework, we assume two potential latent variables: $\eta_i^1$ is the latent outcome that would be realized for subject $i$ if the treatment were administered, and $\eta_i^0$ is the latent outcome that would be realized for subject $i$ in the absence of treatment. We assume the stable unit treatment value assumption, which implies that potential outcomes remain stable regardless of which subjects receive treatment rubin1980randomization.
These latent variables are not directly observed; think of them instead as abstract concepts, such as authoritarianism or corruption, that may be measured with some degree of error. Suppose that there are $J$ observed measures of the latent outcome for each subject. Let the $j^{th}$ outcome measure for unit $i$ be $Y_{ij}= \lambda_j \eta_i + \epsilon_{ij} = \lambda_j[Z_i\eta^1_{i} + (1-Z_i)\eta^0_{i}] + \epsilon_{ij}$, where $\epsilon_{ij}$ is the measurement error with mean zero $\mathbb{E}[\epsilon_{ij}]=0$.\footnote{This assumption can be relaxed to any constant since outcomes can be mean-centered.} More generally, we can allow the variance of the measurement errors to differ between the treatment and control groups by defining two separate measurement error terms $\epsilon_{ij}=Z_i\epsilon^1_{ij} + (1-Z_i)\epsilon^0_{ij}$,
with $\mathbb{E}[\epsilon^1_{ij}]=\mathbb{E}[\epsilon^0_{ij}]=0$. In the superpopulation framework, $\{Z_i,\eta^1_i,\eta^0_i, \epsilon^1_{ij},\epsilon^0_{ij}\}_{i=1}^n$ is assumed to be an i.i.d. sample from a population distribution.
Figure (ref) illustrates the relationship among the observed and unobserved variables.
The observed variables are depicted inside squares, and unobserved variables are depicted inside circles. The latent outcome $\eta_i$ can be represented as $\eta_i= \mathbb{E}\eta_i^0 + \tau Z_i + \zeta_i$, where $\zeta_i$ is the idiosyncratic disturbance: $\zeta_i=\eta_i^0 - \mathbb{E}\eta^0_i+Z_i[(\eta_i^1-\mathbb{E}\eta^1_i)-(\eta_i^0-\mathbb{E}\eta^0_i)]$. This term has mean zero by construction: $\mathbb{E}\zeta_i=0$.
We are interested in the causal effect of treatment on the latent variable, which stoetzer2022causal call the latent treatment effect (LTE). Define the individual-level LTE as $\tau_i=\eta^1_{i}-\eta^0_{i}$. Because the individual-level LTE is unidentified, we focus on the average latent treatment effect (ALTE) defined as $\mathbb{E}[\eta^1_{i}-\eta^0_{i}]$.
\begin{figure}[!h]
\caption{Graphical Depiction of an Experimental Design in which a Latent Outcome is Measured Linearly by Three Outcomes, Each Measured with Error.}
\end{figure}
Our identification strategy is rooted in a set of assumptions about the experimental design and the manner in which the latent variables are measured:
\begin{assumption}[Causal Framework for a Latent Outcome Variable]
We assume $\forall j=1,2,...,J$, and $i=1,2,...,n$,
A. Valid Measurement: $\lambda_j \neq 0$.
B. $Z_i$ is randomly assigned: $Var(Z_i) \neq 0$ and $\{\eta^0_i,\eta^1_i\} \perp Z_i$
C. $Z_i$ is excludable: $\{\epsilon^0_{ij},\epsilon^1_{ij} \}\perp Z_i$
\end{assumption}
The valid measurement assumption is straightforward: an outcome measure $Y_{ij}$ should contain information about the latent variable beyond pure noise. The second assumption requires that the treatment $Z_i$ has at least two distinct values and is randomly assigned to units, implying that it is independent of the potential latent outcomes. The assumption of excludability implies that each outcome's measurement errors are independent of the treatment. One potential scenario in which this assumption fails arises when the measured outcomes are affected by the treatment through channels other than the latent outcome of interest. We discuss this scenario and possible mitigation strategies in Appendix (ref).
This setup presupposes that the relationship between the latent variable $\eta_i$ and the observed measure $Y_{ij}$ is linear. We would offer four comments about linearity. First, our framework and results invoke linearity as a special case of a nonparametric identification strategy. (We discuss nonparametric identification and estimation in Section (ref).) Nonparametric identification typically requires an assumption that the measurements capture sufficient variability in the latent outcome $\eta_i$. This condition implies that discrete measurements with a small number of categories are generally insufficient. We argue that indexes
constructed from multiple linear measures provide one of the simplest and most effective ways to capture the variability about the latent outcome.
Second, it is more natural for researchers to design linear measurements than nonlinear ones. For example, test scores are typically constructed to increase with an individual’s latent ability, and respondents with more extreme latent attitudes are more likely to select extreme options on a seven-point Likert scale. It is rare for a designed measurement to exhibit a highly nonlinear relationship with the latent outcome unless the measured outcome is distributed in a truncated manner, with many observations falling into the highest or lowest categories. More granular measures tend to be more plausibly linear.
Third, beyond design considerations, linearity can also be evaluated empirically. If the linearity assumption holds, the measurements should satisfy $Y_{ij}=\lambda_j Y_{i1} + \epsilon_{ij} \; \forall j \neq 1$ (see lemma (ref)). Thus, evidence of nonlinearity among measurements would lead us to reject the linearity assumption. As we will show in the kalla2020reducing application below, the measures appear to exhibit the expected linear relationship. Were an outcome measure to exhibit signs of nonlinearity, the assumption of linearity can be relaxed for this problematic measure in ways that restore the desirable properties of the estimation approach, as illustrated in SI (ref).
Fourth, when outcomes are binary or ordinal, and we are not using nonparametric methods, we can still use a linear model to obtain a linear approximation; the problem, as we note below, is that assessing how well the model fits the data becomes more model-dependent. For this reason, our design recommendation is to gather outcomes in ways that satisfy the linearity assumption rather than shoehorn discrete outcome measures into a linear modeling framework. The formal results and corresponding strategies are discussed in the section (ref) and SI (ref).
\subsection{Solving the Study-specific Noncomparability Problem}
As emphasized in the introduction, because the latent variable $\eta_i$ has no inherent scale, arbitrarily standardizing measurements and reducing their dimensions -- as in methods such as PCA, ICW and IRT -- renders the latent outcome estimand study-specific. Even if two studies focus on the same latent outcome, any difference in the measurements will lead these methods to produce noncomparable results. As loehlin1998latent points out, such standardization can lead researchers to mistakenly conclude that two estimated ALTEs differ when, in fact, they are the same. This poses a serious problem for the accumulation of knowledge.
Consider two very large studies in which the intervention exerts exactly the same average treatment effect on a latent outcome measured by the same indicators $Y_{ij}$. Standardization would produce two quite different estimates of the ALTE if the measurement error variances were much larger in the first study than in the second. (Indeed, the standardized estimates would differ in this example even though a simple regression of $Y_{i1}$ on $Z_i$ in each sample would, in expectation, yield identical estimates.) Standardization may even frustrate efforts to recover causal parameters from simulated data. In their assessment of several commonly-used methods of creating standardized indexes as experimental outcomes --- simple indexes, weighted indexes, and principal components factor scores --- blair2023research conclude that \emph{none} of them recovers a known parameter in repeated simulations.\footnote{In the SI (ref), we show that both of the estimators we propose below successfully recover the ALTE in their simulations (up to scale factor, depending on which outcome measure is used to set the metric). Moreover, the estimated standard error from our approach closely matches the empirical standard error from their simulation.}
Given these limitations, we propose an “unstandardized” approach. Because the latent variable $\eta_i$ lacks an intrinsic scale, we can, without loss of generality, set $\lambda_1=1$ so that we can interpret $\eta_i$ using the same metric as $Y_1$. For example, if $\eta_i$ represents a latent distance, and $Y_{i1}$ is scaled in terms of kilometers while $Y_{i2}$ is scaled in terms of miles, setting $\lambda_1=1$ means that effects on $\eta_i$ are scaled in terms of kilometers bollen1989structural. (The choice of which observed outcome to use for scaling purposes is arbitrary and has no effect on the statistical significance of the estimated ALTE.\footnote{Beware of the fact that SEM package in Stata uses its estimated standard errors to calculate $p$-values, but these standard errors are calculated based on numerical derivatives that are subject to scale-specific error. To obtain scale invariant $p$-values, conduct a likelihood ratio test comparing the log-likelihood of the fitted model to the log-likelihood of the restricted model, which constrains the average latent treatment effect to be zero.})
As a result, as long as one measurement is shared across studies, we can treat that measure as $Y_{i1}$. The latent treatment effect then targets the same latent outcome, allowing for direct comparison of results across studies. In short, maintaining interpretable scaling units rather than standardizing simplifies the identification problem and preserves the comparability of results across experimental replications. The remaining question is how to combine $Y_{i1}$ with other measurements that have different relationships with the latent outcome due to different $\lambda_j$ values. We address this issue in the next section.
\section{Identification of Measurement Parameters}
We seek to identify the average causal effect of $Z_i$ on the latent variable $\eta_i$. However, $\eta_i$ is not directly observed, and even when setting $\lambda_1=1$ we can only partially measure $\eta_i$ given the unknown scaling parameters $\lambda_j$ and unobserved measurement errors $\epsilon_{ij}$. If we knew all the $\lambda_j$, we could rescale the observed measures, for example, by $\frac{Y_{ij}}{\lambda_j}=\eta_i+\frac{1}{\lambda_j}\epsilon_{ij}$, to approximate the latent variable $\eta_i$ using the same units as $Y_{i1}$.
The question is how to identify $\lambda_j$. As demonstrated in the following proposition, identification requires a third variable to serve as an instrumental variable. What is perhaps surprising is that this third variable can be either the treatment $Z_i$ or other measurements $Y_{ij}$ under the relatively mild conditions in Assumption (ref).\footnote{When pre-treatment covariates are available, they can also serve as instrumental variables, as discussed in Section 6.1.}
\begin{proposition}[Identification of the measurement scaling parameters]
Suppose Assumption (ref) holds. Then
in the above causal framework given $\lambda_1=1$, $\lambda_j$ is identified either by
(1) $\lambda_j=\frac{Cov(Z_i,Y_{ij})}{Cov(Z_i,Y_{i1})}$ \text{ if } $\mathbb{E}[\eta_i^1-\eta_i^0]\neq 0$ or by
(2) $\lambda_j=\frac{Cov(Y_{ik},Y_{ij})}{Cov(Y_{ik},Y_{i1})}\; \forall k \neq 1 \text{ and } k \neq j$, if $Cov(\epsilon_{ik}, \epsilon_{i1})=Cov(\epsilon_{ij}, \epsilon_{ik})=0$, $Var[\eta_i]\neq 0$, and $\{\eta_i^0,\eta_i^1\} \perp \{\epsilon^1_{ij},\epsilon^0_{ij}\}$.
\end{proposition}
\begin{proof}
The proof is in the SI (ref).
\end{proof}
Proposition (ref) suggests that $\lambda_j$ can be identified via an instrumental variables approach. We can use either the treatment $Z_i$ or other outcome measures $Y_{ij}$ as instrumental variables when certain assumptions hold. Specifically, $Z_i$ must be a “relevant” instrumental variable whose average treatment effect on $\eta_i$ is nonzero (so that the denominator of the ratio converges to a nonzero value as the sample size increases). The excludability assumption of $Z_i$ as an instrument follows from Assumption (ref)B, because it is independent of the measurement errors. Similarly, $Y_{ik}$ is a valid instrumental variable if the measurement errors are uncorrelated with one another, and the limiting covariance between $Y_{ij}$ and $Y_{i1}$ is nonzero. Even when the measurement errors are correlated, $Z_i$ remains a valid instrument so long as the treatment is unrelated to these errors of measurement.
To illustrate the reasoning behind the instrumental variables approach, consider one strategy for identifying $\lambda_2$. From Lemma (ref) in (ref), linearity between each measurement and the latent outcome implies linearity among the different measurements. In particular,
$$
Y_{i2} = \lambda_2 Y_{i1} + (\epsilon_{i2}-\lambda_2\epsilon_{i1})
$$
Because $Y_{i1}$ is correlated with the error $\epsilon_{i1}$, $\lambda_2$ is not directly identified, and we cannot simply regress $Y_{2i}$ on $\eta_i$ because the latter is unobserved. Again, instrumental variables regression provides a consistent estimator. Given the excludability of the randomly assigned treatment variable $Z_i$ (Assumption (ref)C), we can use $Z_i$ as an instrumental variable for $Y_{i1}$.
In the proof, we show that if the average treatment effect is nonzero ($\mathbb{E}\tau_i \neq 0$), then $Cov(Z_i,Y_{i1})\neq 0$. In that case, $Z_i$ is a “relevant” predictor of $Y_{i1}$. Moreover, given that $Z_i$ is randomly assigned and thus independent of measurement errors, $Z_i$ is a valid instrumental variable for $Y_{i1}$. Therefore, we can use $\frac{\widehat{Cov}(Z_i,Y_{i2})}{\widehat{Cov}(Z_i, Y_{i1})}$ to estimate $\lambda_2$. The same approach shows that $Y_{i3}$ can also serve as an instrumental variable for $Y_{i1}$ when we stipulate that the measurement errors are unrelated to the latent potential outcomes. This assumption implies, for example, that higher values of the latent potential outcomes are no more likely to coincide with higher values of the measurement errors than lower values of the latent potential outcomes.
What about cases such as Figure (ref), where the experiment includes \emph{both} a randomly assigned treatment and three measures of the latent variable? In such cases, we have more than one plug-in estimator for $\lambda_2$ and $\lambda_3$, and therefore the unknown scaling parameters are said to be overidentified. Overidentification is helpful in two ways. First, we can combine the plug-in estimators to form a more efficient estimator, using method of moments or maximum likelihood. The former has the advantage of making weaker distributional assumptions about the observed or unobserved variables in model, while the latter imposes the assumption of multivariate normality. In practice, the two estimators tend to produce similar estimates in large samples when the observed variables are distributed symmetrically olsson2000performance,browne1984asymptotically,yuan2005nonequivalence, boomsma2001robustness.
The second advantage of overidentification is that it allows us to test whether the data accord with the posited model. When two or more plug-in estimators render markedly different estimates, that is a sign that at least one modeling assumption is untenable. For example, when $Z_i$ is used as the instrumental variable for identifying $\lambda_2$, we need not invoke any assumptions about the uncorrelated measurement errors (i.e., the assumption that $\epsilon_{i1}$, $\epsilon_{i2}$, and $\epsilon_{i3}$ are uncorrelated with one another), whereas when we identify $\lambda_2$ using $Y_{i3}$ as the instrumental variable, we do invoke this assumption when asserting that the limiting covariances among the errors are zero. If the two IV estimates differ markedly, that is a sign that uncorrelated measurement errors may be an untenable assumption.
\section{Estimation of the Average Latent Treatment Effect}
Once the $\lambda_j$ are identified, researchers have many options for estimating the average treatment effect on the latent variable. These options fall into two broad categories. The first category involves forming a weighted average of the observed outcome measures to approximate the latent variable, albeit with some residual error. This index-building approach has the advantage of allowing the analyst to use conventional methods, such as regression, to estimate the ALTE. An alternative approach is to estimate both the measurement parameters and the ALTE simultaneously as part of a system of linear equations. This approach is typically referred to as structural equation modeling or SEM. Although SEM is often associated with full-information estimators that assume multivariate normality joreskog1970general, the same models can be rooted in more agnostic assumptions and estimated by method-of-moments kline2023principles.
This section discusses these two estimation approaches -- weighted scale indices (WSI) and maximum likelihood -- using simulated data to assess their ability to recover known treatment effects and estimate sampling variability. Fortunately, both approaches tend to perform well and generate similar results. The WSI has advantages in terms of transparency, as it allows the analyst to display regression results visually using individual-level data cook2009regression; maximum likelihood has the advantage of providing a unified framework for estimating treatment effects and using overidentification to test modeling assumptions.
\subsection{Difference-in-means based on a Weighted Scaled Index (WSI)}
The simplest and most transparent estimator is the difference-in-means estimator using the weighted scaled index as the outcome. For each individual $i$, we first \textit{scale} each outcome measure using $\lambda_j$ to address the study-specific noncomparability problem. We then construct an outcome index measure by creating a \textit{weighted} average of the various outcome measures. Taking this weighted average to be the outcome, we use difference-in-means to estimate the causal effect. Formally, for each individual $i$, let $\tilde{Y}_{ij}=\frac{1}{\lambda_j}Y_{ij}$, and $\tilde{Y}_i=\sum_{j=1}^J \omega_j \tilde{Y}_{ij}$, where $\omega_j$ is the weight and $\sum_{j=1}^J \omega_j=1$. The most elementary weighting scheme is $\frac{1}{J}$. That is, the new outcome measure $\tilde{Y}_i$ is a simple average of all measurements for individual $i$. Although simple averages are widely used in practice ansolabehere2008strength, they are not optimal from the standpoint of estimating the ALTE as precisely as possible. A more efficient approach is to use inverse-variance weighting. Because we divide each measure by $\lambda_j$, the true variance of the measurement error is $\frac{\sigma^2(\epsilon^z_{\cdot j})}{\lambda^2_j}$; this leads to the optimal weight $\omega^*_j = \frac{\lambda^2_j/\sigma^2(\epsilon^z_{\cdot j})}{\sum_{j=1}^k \lambda^2_j/\sigma^2(\epsilon^z_{\cdot j})}$.
This optimal weighting scheme presented above assumes that measurement errors are independent; when errors are correlated, the optimal weighting algorithm becomes more complex.\footnote{In general, the weight is derived by minimizing the variance of the weighted average in treatment and control groups. For example, given each group, we minimize $\sum_{j=1}^J \omega_j \tilde{Y}_{ij}$ subject to the constraint that $1'w=1$. Define $\Sigma$ as the variance-covariance matrix of the measurement errors. By using the Lagrange multiplier method, we get the weight $w=(1'\Sigma^{-1}1)^{-1}(1'\Sigma)$, where $\Sigma$ is the variance and covariance matrix. If measurement errors are independent, the optimal weight reduces to $\omega^*_j = \frac{\lambda^2_j/\sigma^2(\epsilon^z_{\cdot j})}{\sum_{j=1}^k \lambda^2_j/\sigma^2(\epsilon^z_{\cdot j})}$.} When the variance of the measurement error is assumed to be the same across the treatment and control groups, we can efficiently estimate a common $\sigma^2$ using the pooled data.
Using the WSI, a researcher applies conventional methods for estimating the ALTE. We begin by considering the difference-in-means estimator:
$$
\begin{aligned}
\hat{\tau}&=\frac{1}{n_1}\sum_{i=1}^{n}Z_i\tilde{Y}_i - \frac{1}{n_0}\sum_{i=1}^{n}(1-Z_i)\tilde{Y}_i\\
&=\frac{1}{n_1}\sum_{i=1}^{n}Z_i [\sum_{j=1}^k \omega_j\tilde{Y}_{ij}]- \frac{1}{n_0}\sum_{i=1}^{n}(1-Z_i)[\sum_{j=1}^k \omega_j \tilde{Y}_{ij}]
\end{aligned}
$$
Suppose we know the true $\lambda_j$. Then, it is straightforward to see that the oracle estimator is unbiased. When $\lambda_j$ is unknown, we can still use a consistent estimator of these scaling parameters to obtain a consistent difference-in-means estimator.
\begin{proposition}
Suppose assumption (ref) holds and all moments are finite. Consider a weighted difference-in-means estimator $\hat{\tau}=\frac{1}{n_1}\sum_{i=1}^{n}Z_i\tilde{Y}_i - \frac{1}{n_0}\sum_{i=1}^{n}(1-Z_i)\tilde{Y}_i$.
(1) If $\lambda_j$ is known, then $\hat{\tau}$ is unbiased.
(2) If $\lambda_j$ is estimated consistently under proposition (ref), then the weighted difference-in-means estimator is consistent.
\end{proposition}
\begin{proof}
The proof is in the SI (ref).
\end{proof}
In practice, regression is widely used to estimate the ALTE, especially when covariates are used to improve precision. In SI (ref), we show that the above difference-in-means estimator is equivalent to a “stacked” regression.
Let us now consider the variance of this estimator. If the $\lambda_j$ scaling factors are known based on the measurement properties observed in prior studies (and $\omega$ is pre-specified), then we can apply the conventional Neyman variance estimator, as discussed in SI (ref). If the $\lambda_j$ and/or $\omega_j$ are estimated from the data, we have to account for the estimation uncertainty. The standard approach, as with the \textit{inverse probability weighting estimator}, uses generalized method of moments (GMM), as discussed in SI (ref).
\subsection{Structural Equation Modeling}
An alternative approach is to model all of the parameters in Figure (ref) as a system of linear equations, one equation for each of the endogenous variables in the causal graph. Figure (ref), for example, implies one equation for the latent outcome variable ($\eta_i$) and three equations for each of the observed outcome measures. After imposing the scaling metric $\lambda_1=1$, we are left with a total of 8 free parameters: the variance of the randomized treatment ($\phi$), the average treatment effect of $Z_i$ on $\eta_i$, the variance of the unobserved causes of the latent outcome variable ($\psi$), two scaling parameters ($\lambda_2$ and $\lambda_3$), and three variances of the measurement errors associated with $Y_{i1}$, $Y_{i2}$, $Y_{i3}$. As we show in SI (ref), these 8 parameters may be used to predict the 10 observed variances and covariances among the four measured variables ($Z_i$,$Y_{i1}$,$Y_{i2}$,$Y_{i3}$). Since some of these parameters are overidentified, the estimator we use may affect our empirical estimates because different estimators use different objective functions when gauging the fit between the estimated parameters and the observed variance-covariance matrix. As noted above, the most commonly used estimator is maximum likelihood under the assumption of multivariate normality, but other estimators that make weaker distributional assumptions are available and tend to perform well when the number of observations is large relative to the number of parameters yuan2005nonequivalence.
\subsection{Comparison to Current Methods}
Our WSI method consistently estimates latent treatment effects that are comparable across studies. Commonly used approaches for handling multiple measurements include principal components analysis and inverse covariance weighting. Table in SI (ref) lists 14 recent papers published in the \emph{American Political Science Review} that employ these two methods. However, unlike SEM (which is rarely used to analyze experiments), these methods were not originally designed to estimate average latent treatment effects; inverse covariance weighting was designed to facilitate hypothesis testing, and principal components analysis was designed to reduce the dimensionality of a collection of correlated measures. We compare these methods with WSI and demonstrate that our approach achieves additional statistical precision.
Figure (ref) compares the power of five alternative methods for constructing outcome measures.\footnote{stoetzer2022causal propose a hierarchical item response theory (hIRT) model to estimate the latent treatment effect. Their model has two components. First, the latent variable $\eta_i$ is assumed to be a linear function of the treatment: $\eta_i = \gamma_0 + \gamma_1 Z_i + \epsilon_i$, where $\epsilon_i$ is normally distributed with a constant variance $\sigma^2$. In this equation, $\gamma_1$ is the average treatment effect of interest. Next, the observed outcomes $Y_{ij}$ are generated by a characteristic function of the item: $\mathbb{P}[Y_{ij}=h | \eta_i]=P_{ijh}(\eta_i;\alpha_{ijh},\beta_{ij})$, where $\alpha_{ijh}$ and $\beta_{ij}$ are the item difficulty and item discrimination parameters for item $j$ and answering $h$. As stoetzer2022causal note (e.g. page 28), this model makes a set of parametric assumptions, including constant treatment effects across subjects and a particular item characteristic function. Our approach relaxes some of these assumptions. We allow $Z_i$ to have heterogeneous treatment effects. We also remain agnostic about the non-linear transformation function. Because the IRT model is not suitable in the presence of heterogeneous treatment effects (HTE), we do not report power calculations. More discussion may be found in SI (ref).} WSI and SEM estimators prove to have the greatest power (they are almost indistinguishable in the figure). Both methods are clearly superior to an equally weighted scaled index, because optimal weights increase the signal-noise ratio. The WSI and SEM also tend to be more powerful than PCA and ICW. Ironically, those dominated methods are the most frequently used approaches in recently published work. Detailed discussion of the simulations may be found in Appendix (ref).
\begin{figure}[!h]
\caption{Power analysis: WSI, Equal weighting, SEM, ICW, and PCA}
\end{figure}
\textbf{t-ratio and noise in the outcome variable.} Why does suboptimal measurement of experimental outcomes affect statistical power? Recall the regression function in the framework: $\eta_i = \mathbb{E}[\eta_i^0] + \tau Z_i + \xi_i$. For simplicity, suppose $\lambda_j = 1 ; \forall j$. Recall that the weighted (scaled) outcome is $\tilde{Y}_i = \sum_{j=1}^J \omega_j Y_{ij}$. Therefore, the observed regression $\tilde{Y}_{i}=\mathbb{E}[\eta_i^0] + \tau Z_i + (\tilde{\xi}_i+\epsilon'_i)$, where $\tilde{\xi}_i=\sum_{j=1}^J \omega_j \xi_{ij}$ and $\tilde{\epsilon}_i=\sum_{j=1}^J \omega_j \epsilon_{ij}$.
The $t$-ratio for the treatment variable is $t=\frac{\hat{\tau}}{\hat{se}(\hat{\tau})}$, where the denominator $\hat{se}(\hat{\tau})$ depends on the estimated variance of the error term: $\hat{se}(\hat{\tau})=\sqrt{\frac{\hat{\sigma}^2_{\tilde{\xi}}+\hat{\sigma}^2_{\tilde{\epsilon}}}{\sum_{i}(Z_i-\overline{Z}_i)}}$. From this expression, we see that when the noise in the outcome measure is large, the $t$-ratio becomes smaller.
\section{Design Trade-off: More Outcome Measures or More Subjects?}
One implication of the preceding section is that, as broockman2017design argue, researchers could improve the precision of their experiments by investing in additional outcome measures. Under what conditions is such an investment warranted?
The super-population variance of the (oracle) estimator is
\begin{equation}
Var(\hat{\tau}) =\frac{Var[\tilde{Y}^1_i]}{n_1}+\frac{Var[\tilde{Y}^0_i]}{n_0}
\end{equation}
This formula shows that researchers can decrease this variance by adding more measurements, which decreases the variance in the numerator, or by adding more observations, which increases the denominator. We examine both approaches and explore the optimal allocation under a budget constraint.
\subsection{Advantages of Adding More Outcome Measurements}
The variances of the difference-in-means or regression estimators depend mainly on the variance of the weighted averages $\tilde{Y}^1_i$ and $\tilde{Y}^0_i$. Assuming $Cov(\epsilon_{ij},\epsilon_{ik})=0$ and ignoring the estimation variance of $\hat{\lambda}$ on the grounds that it is known based on existing studies, under optimal inverse-variance weighting $\omega^*_j$, the variance of $\tilde{Y}^1_i$ is $\sigma^2(\eta^1_i) + \frac{1}{\sum_{j=1}^J \lambda^2_j/\sigma^2(\epsilon_{\cdot j})}$, where the latter part is due to (optimally weighted) measurement error.
The variance formula shows that the inclusion of a highly reliable outcome measure (i.e., a measure with a high ratio of trait variance to error variance) may substantially increase the precision with which the ALTE is estimated. However, the marginal gains from each additional measure depend on the reliability of the measures already contained in the additive index. Fortunately, with optimal weighting, in expectation more measurements never increase the variance of the estimated ALTE because unreliable measures are down-weighted accordingly.\footnote{It is important to note that this desirable feature of adding outcome measures may not hold without optimal weighting. See more discussion in (ref). }
\subsection{Gathering More Observations}
In the super-population framework, the variance of our WSI estimator is shown in Equation (1),
the numerators are the variances of weighted average potential outcomes in the super-population. Adding more observations affects the denominator. Under a balanced design, when one adds $j$ more observations in each group, the variance will be $[\frac{2Var[\tilde{Y}^1_i]}{n+2j}+\frac{2Var[\tilde{Y}^0_i]}{n+2j}] / [\frac{2Var[\tilde{Y}^1_i]}{n}+\frac{2Var[\tilde{Y}^0_i]}{n}] =\frac{n}{n+2j}$ of the variance when the sample size is $n$; in other words, the percentage reduction in variance is $\frac{2j}{n+2j}$. Therefore, researchers can add $\frac{n}{4}$ subjects to the treatment and control groups to decrease the variance by $\frac{1}{3}$.
\subsection{Trade-off under a budget constraint}
We can apply this framework to a more general budget allocation problem. Suppose that researchers have a budget $B > 0$ and the cost of adding one more measure is $c_m$ and the cost of adding one more observation is $c_o$. Consider the optimal allocation of budget for the sample size $n$ and the number of items $J$. This is a typical constrained optimization problem
where objective is to minimize the sampling variance given the budget constraint $c_mJ+c_0 n\leq B$. Optimization may be used to find the best budget allocation between additional outcome measures and additional respondents.
Now, let us consider a more concrete scenario. Suppose researchers have an extra $\$5000$ budget. They may decide to recruit more respondents and/or add more measures. What is the optimal decision? To be consistent with the previous simulation, we assume the current sample size is $n=500$, and there is only one measure. The marginal cost of the measure is $c_m=\$1000$ and the marginal cost of each respondent is $c_o=\$10$. Let us consider two scenarios: the outcomes are measured with high reliability ($0.75$) or low reliability ($0.4$).\footnote{To be specific, we set $Var(\eta^1)=Var(\eta^0)=1$. For low reliability, the variance of the measurement error is $1.701$ so that $\frac{Var(\eta)}{Var(Y)}\approx 0.4$. For high reliability, the variance of the measurement error is $0.379$ so that $\frac{Var(\eta)}{Var(Y)}\approx 0.75$. The objective function is to maximize the variance gain, or equivalently, $\min_{k,j} \frac{-2k[Var(\eta^1)+Var(\eta^0)]}{500(500+k)}+\frac{-4(k+500j+jk)}{500(1+j)\Sigma(500+k)}$, where $j$ is the number of additional measures and even number $k$ represents the additional respondents, and $\Sigma:=\frac{\lambda^2}{\sigma^2(\epsilon)}$. The constraint is $1000j + 10k \leq 5000$.}
For high reliability case, the optimal solution is to add one more measure and recruit 400 more respondents. However, for the low reliability case, the optimal solution is to add two more measures and recruit 300 more respondents. In other words, although the extra measures are low quality, they are an especially good investment. This somewhat surprising result is formally explained in SI(ref); but it can be intuitively understood from the simulation in Figure (ref). When measurement error is large (low reliability), the variance reduction is relatively greater.
\begin{figure}[!h]
\caption{\textbf{Simulation Illustrating Variance Reduction Due to Collection of Additional Outcome Measures.} The horizontal line represents the number of measures and the vertical line is the estimated variance of the optimal weighting estimator. The variance reduction is larger if the reliability is lower. A detailed explanation may be found in SI (ref).}
\end{figure}
Two other considerations also arise when weighing the advantages of collecting additional outcome measures. The first is that additional observable manifestations of the latent outcome make for more overidentifying restrictions on the key scaling parameters. Each additional measure gives us more ways to assess the scaling properties of all of the outcome measures at hand. This fact, in turn, allows us to assess the goodness-of-fit of a posited multi-equation model. The second consideration is that the more outcome measures one gathers, the more likely one is to discover the inadequacies of one or more of the measures. One symptom of a measurement problem is poor fit between the predicted and actual variance-covariance matrix; another is fluctuation of the estimated ALTE when certain outcome measures are included or excluded. Ideally, additional outcome measures improve precision and robustness; however, the inclusion of flawed measures may undercut these benefits. Therefore, when gauging the trade-off between allocating resources to subjects or outcome measurements, one should bear in mind that the value of additional outcome measures hinges on their validity and reliability.
\section{Strategies for Improving Precision and Assessing Robustness}
\subsection{Covariates}
The analysis of experiments with latent outcomes may benefit in three ways from the inclusion of covariates, defined as variables measured prior to random assignment. First, covariates may be used to identify measurement parameters, such as the $\lambda_j$ in Figure (ref). Indeed, when it comes to identifying scaling parameters, covariates play the same identifying role as randomly assigned interventions, and the assumptions required are similar: the covariate must be statistically independent of the measurement errors but correlated with the structural disturbance term ($\zeta_i$). In other words, the covariate must predict $Y_{i1}$ to some extent so that the instrumental variables estimator is defined. Second, covariates that are strongly prognostic of outcomes will improve the precision with which the average treatment effects are estimated. This property of covariate adjustment is in keeping with conventional experimental analysis, where outcomes are modeled directly without reference to latent variables. Third, because covariates contribute to the identification of measurement parameters, covariates allow the researcher to use overidentification to assess the goodness-of-fit of the posited multi-equation model. In sum, although covariates are not strictly necessary for the identification or estimation of measurement parameters, they may nevertheless play a useful role, especially when they are predictive of outcomes.
\subsection{Nonparametric Identification and Linearity}
A key assumption of our approach is the linear relationship between the latent outcome $\eta_i$ and its observed measurements $Y_{ij}$. As emphasized in previous sections, this is fundamentally a design-based assumption: researchers are responsible for designing appropriate measures of the latent construct. The simplest and most natural way to measure a latent outcome is by positing a linear relationship -- for instance, test scores should increase with latent ability. It is difficult to justify a measurement that has a highly nonlinear relationship with the latent outcome, especially when a simpler linear measure is available.
Our framework, in fact, allows for more flexible nonparametric extensions. We establish the formal nonparametric results—both for technical rigor and completeness—in a separate paper, and sketch the main ideas here. Instead of assuming a linear relationship, each measurement $Y_{ij}$ can be viewed as a random draw from some distribution indexed by the latent variable $\eta_i$, denoted $P_j(\cdot|\eta_i)$. The randomness naturally implies the presence of measurement errors. Our linear specification, $Y_{ij}= \lambda_j \eta_i + \epsilon_{ij}$ is a special case in which the expectation of $Y_{ij}$ is linear in $\eta_i$.
To address the study-specific noncomparability problem (as in the linear case), we need to rescale all measurements relative to $Y_{i1}$. Specifically, we require a measurement bridge function $h_j$ for each measurement $j\neq 1$ such that $\mathbb{E}[h_j(Y_{ij})|\eta_i]=\mathbb{E}[Y_{i1}|\eta_i]$. This bridge function maps each measurement onto the reference measurement $Y_{i1}$. In the linear case, the bridge function is unique and takes the form $h_j(Y_{ij}) = \frac{1}{\lambda_j} Y_{ij}$, such that $\mathbb{E} [\frac{1}{\lambda_j} Y_{ij} \mid \eta_i] = \mathbb{E}[Y_{i1} \mid \eta_i]$.
One sufficient condition of the existence of such a measurement bridge function, $h_j$, is completeness assumption.\footnote{The completeness assumption and results may be found in the double negative control literature miao2018confounding,tchetgen2020introduction. Formally, the completeness assumption states that $\mathbf{E}[g(\eta_i)|Y_{ij}]=0$ almost surely if and only if $g(\eta_i)=0$ almost surely.} Intuitively, completeness requires that the measurement $Y_{ij}$ captures the full information (variability) of the latent outcome $\eta_i$. For example, if $\eta_i$ is discrete with three categories, then $Y_{ij}$ must also take on at least three distinct values. In the typical case in which $\eta_i$ is continuous, we do not recommend using measurements that lose information (such as discretized or categorical outcomes). Instead, we encourage researchers to design measures that plausibly bear a linear relationship to the underlying latent variable of interest, which is the simplest way to satisfy the completeness assumption.
As in the proximal causal inference literature miao2018identifying,miao2018confounding, identifying $h_j$ requires auxiliary information from either $Z$ or another measurement $Y_{ik}$ (for $k \neq 1 \text{ and } j$). For example, by taking the expectation of both sides of the bridge equation with respect to $p(\eta_i|Z_i)$, we obtain $\mathbb{E}[h_j(Y_{ij})|Z_i]=\mathbb{E}[Y_{i1}|Z_i]$.
In the linear case, with the bridge function $h_j(Y_{ij})=\frac{1}{\lambda_j} Y_{ij}$, we obtain
$Y_{ij} = \lambda_j Y_{i1} + (\epsilon_{ij}-\lambda_2\epsilon_{i1})$ and use $Z_i$ as an instrumental variable to identify $\lambda_j$, as shown in section (ref). In more general settings, $h_j$ can be estimated using nonparametric methods.
\begin{example}[Binary measurements]
Suppose researchers measure a latent outcome through one continuous measurement $Y_{i1}$ and two binary measurements $Y_{ij}, j=2,3$. Binary measures are an instructive case because they cannot be linear functions of the latent variable.
Recall, $h_j$ is the solution satisfying
$\mathbb{E}[h_j(Y_{ij})|Z_i]=\mathbb{E}[Y_{i1}|Z_i]$. Because $Y_{ij},j=2,3$ only takes two value, for $Z_i=1$, we obtain
$$
\begin{aligned}
\mathbb{E}[Y_{i1}|Z_i=1]&= \mathbb{E}[h_j(Y_{ij})|Z_i] \\
&= h_j(Y_{ij}) \mathbb{P}[Y_{ij}=1|Z_i=1] + h_j(Y_{ij}) \mathbb{P}[Y_{ij}=0|Z_i=1] \\
&= h_j^1 \mathbb{P}[Y_{ij}=1|Z_i=1] + h_j^0 \mathbb{P}[Y_{ij}=0|Z_i=1].
\end{aligned}
$$
Here $h_j^1$ and $h_j^0$ are two unknown parameters. Similarly, for $Z_i=0$, we obtain
$$
\mathbb{E}[Y_{i1}|Z_i=0]=h_j^1 \mathbb{P}[Y_{ij}=1|Z_i=0]+ h_j^0 \mathbb{P}[Y_{ij}=0|Z_i=0].
$$
We have two equations and two unknowns, and $h_j^1$ and $h_j^0$ are identified under NPIV assumptions.
To estimate the ALTE, we first obtain transformed outcomes $\tilde{Y}_{i2}=h_2(Y_{i2})$ and $\tilde{Y}_{i3}=h_3(Y_{i3})$. In this binary case, we can write as $\tilde{Y}_{i2}=Y_{i2}h_2^1+(1-Y_{i2})h_2^0$ and $\tilde{Y}_{i3}=Y_{i3}h_3^1+(1-Y_{i3})h_3^0$. The final combined outcome variable is $\tilde{Y}_i=\omega_1 Y_{i1}+\sum_{j=2}^3\omega_j \tilde{Y}_{ij}$. With this new outcome variable, we can estimate ALTE as usual, for example $\mathbb{E}[\tilde{Y}_i|Z_i=1]-\mathbb{E}[\tilde{Y}_i|Z_i=0]$ \footnote{In practice, we need to use additional instruments; otherwise, there is no efficiency gain in this example, because the sample moments are exactly the same for the three outcomes given treatment status.}.
To illustrate this identification result, we simulate $\eta_i^0 = 0$ and $\eta_i^1 \sim N(1,1)$. The outcome $Y_{i1}$ is then constructed as $Y_{i1} = \eta_i + e_{i1}$. For the binary measurements $Y_{i2}$ and $Y_{i3}$, we first generate latent continuous variables $\xi_{i2} = 0.5\eta_i + e_{2i}$ and $\xi_{i3} = 1.5\eta_i + e_i$, and then create binary measurements based on whether the latent variable $\xi$ is above or below its mean: $Y_{i2} = 1[\xi_{i2} < \overline{\xi_2}]$ and $Y_{i3} = 1[\xi_{i3} < \overline{\xi_3}]$. Figure (ref) shows that nonparametric WSI correctly estimates the true ALTE of 1. In contrast, if we ignore the nonlinearity and instead use linear WSI or SEM, the estimates are clearly biased. The implication is that a nonparametric approach allows the researcher more latitude in estimating the ALTE when some of the outcomes are binary indicators.
\begin{figure}
\caption{Binary measurements and Nonparametric Estimation}
\end{figure}
\end{example}
Although nonparametric methods provide less model-dependent results, they inevitably come with larger estimation error. We therefore suggest that researchers begin by testing for linearity. Lemma (ref) shows that linearity between the measured outcomes and the latent outcome implies linearity among the different measured outcomes. Therefore, we can test whether $Y_{ij}$ and $Y_{ik}$ follow a linear functional form. The literature offers multiple specification tests for this purpose. For example, the Rainbow test checks whether linearity holds in the central portion of the data. utts1982rainbow; the Regression Equation Specification Error Test examines whether the coefficients on added polynomial terms are jointly zero ramsey1969tests.
More generally, we can compare the linear specification to any nonparametric alternative using methods proposed by hardle1993comparing, zheng1996consistent, hsiao2007consistent, among others.
As noted above, linearity is untenable when latent outcomes are measured by survey questions with binary response options. One work-around is to use more granular response options, such as presenting an agree-disagree question with response options ranging from “strongly agree” to “strongly disagree.”\footnote{However, this approach may still yield a lopsided response distribution that would not plausibly be modeled as a linear function of the latent variable.} In the extreme case where only a few discrete measurements are available, a more convincing way to satisfy the linearity condition is to create an additive index composed of several discrete measures, in much the same way that standardized tests measure math ability by counting the number of correct answers to many specific multiple-choice questions.\footnote{Modern variants of scholastic aptitude tests adaptively calibrate the difficulty of the questions based on the answers to prior questions, and the same cost-saving principle can be applied to survey measures of political knowledge and other topics montgomery2022adaptive.}
This approach, which dates back to the early days of psychometric assessment, may require pilot testing (perhaps in a nonexperimental context or during pre-treatment baseline measurement) in order to devise a series of discrete measures that, when summed, produce granular distributions with few observations in each tail. For technical discussion, please see SI (ref).
\begin{figure}[!h]
\caption{\textbf{Assessing linearity between $\eta$ and an additive index created by adding binary items together}. We create an additive index $v_1$ by summing up 5, 10, 15, and 20 binary responses from IRT models. Data points have been jittered slightly for clarity. The rug plots on both axes denote the distribution of data. The vertical axis shows the additive index created by adding all binary variables in the simulation. The horizontal axis is the true latent variable ($\eta$) used in the IRT model.}
\end{figure}
To illustrate how additive indices help satisfy the assumed linear mapping between latent and observed outcomes, we performed a series of simulations. In each simulation, binary responses were generated using an IRT model, and two outcome indices were then created by summing half of the binary items together.\footnote{We use simTrt function in the psych R package to generate the data. To be specific, the discrimination parameter $a \sim Uniform (.5, 1.5)$, item difficulties $b \sim N(0, 2)$, and the guessing asymptote is $c=0.2$.} The number of items used to create each index increases from one simulation to the next, as indicated by the horizontal axis labels of Figure (ref). The sequence of graph panes shows that as we sum more binary items, the LOESS fitted line -- which allows a nonlinear relationship if the data suggest it -- between the observed measures and the true $\eta_i$ used in the simulation becomes increasingly linear. Figure (ref) illustrates an extreme case where all component items are binary. In SI (ref), we also show the simulation results for ordinal items. Intuitively, if the component items have more granular response options, fewer items will be needed to form an additive index that meets the linearity condition.\footnote{In practice, $\eta$ is not observed. Instead, we could show the increasingly linear properties between two observed measures. Figure (ref), which does so, conveys the same point: as the number of binary items used to create an index increases, the more linear the apparent relationship.} On the other hand, building an index from survey items that are subject to the same measurement errors due to similar question wording, order, and response format slows the rate at which linearity is achieved. As a thought experiment, consider the limiting case in which two measures elicit exactly the same responses from each subject; in this case, an index based on both measures is no more granular than each measure considered separately. The design implication is clear: if possible, researchers should measure the latent variable in different ways so as to reduce the correlation between errors of measurement andrews1984construct.
\subsection{Testing Measurement Equivalence across Experimental Groups}
Our basic framework assumes that responses from subjects in the treatment and control groups have the same measurement structure. In other words, the translation from $\eta_i$ to the $Y_{ij}$ is the same regardless of the subject's experimental condition. This assumption may not hold in practice, for example, if there is something about the treatment that changes the way that subjects interpret outcome questions.\footnote{The intuition extends to non-survey outcomes as well. If outcomes in the control group are measured in miles but outcomes in the treatment group are measured in kilometers, any apparent mean difference may be misleading.} Different scaling parameters jeopardize the exclusion restriction, which holds that the treatment itself is the only systematic reason that expected outcomes in the treatment and control groups may differ. Fortunately, the measurement equivalence assumption can be assessed empirically. In order to make the test as general as possible, we relax the implicit assumption that the observed outcome is centered on the latent variable. A more flexible model is that $Y_{ij} = \lambda^Z_j g(\eta_i) + \alpha_z + \epsilon_{ij}$, where $\alpha_z$ denotes the so-called “structural mean” bagozzi1989use for experimental group $z$. Unlike the model presented in section (ref), this one expresses the treatment effect as a shift in latent means, while at the same time allowing for the possibility that the measurement scaling parameters $\lambda^Z_j$
may differ for the treatment and control groups. The more flexible model may be tested as a nested alternative to the model that assumes the same measurement parameters apply to both experimental groups, as illustrated in the empirical example below.
\subsection{Comparison to Seemingly Unrelated Regression (SUR)}
Many experiments are designed with a clear intention to gather multiple measures of a given latent outcome, in which case the modeling framework described here is an appropriate starting point. However, other experiments gather multiple outcomes that are not necessarily viewed as redundant measures of a given latent factor. For example, field experiments on conditional cash transfers examine outcomes such as whether children in the household are enrolled in school, vaccinated, and free from signs of malnutrition. One could view such outcomes as manifestations of an underlying factor (child well-being), but one could instead consider them to be three distinct outcomes. How might we evaluate the adequacy of the two modeling approaches?
\begin{figure}[!h]
\caption{Graphical Depiction of a Linear Measurement Process with Additive Measurement Errors (left) and a Seemingly Unrelated Regression Model (right)}
\end{figure}
Fortunately, the two models may be expressed as nested alternatives. Figure (ref) shows two causal graphs side-by-side. The left pane depicts the now-familiar model in which three outcome measures are linear manifestations of a latent outcome. The right pane depicts the seemingly unrelated regression model zellner1962efficient in which the same three outcome measures are modeled as distinct outcomes.\footnote{The inverse covariance weighting method proposed by anderson2008multiple is closely connected to the SUR model because it does not posit a latent variable measured with error. This method down-weights outcomes that are correlated with one another in order to generate a unified outcome measure with minimum variance. Simulations in Figure (ref) show that the inverse covariance weighting method is less powerful than SEM when outcomes measures truly tap a latent outcome with error.
} Whereas the latent variable model has 8 free parameters, the SUR model has 10 because it does not constrain the causal path from $Z_i$ to the $Y_{ik}$ to flow through $\eta_i$. Therefore, two degrees of freedom may be used to assess the extent to which the more parsimonious latent variable model adequately fits the observed variance-covariance matrix. Rejection of the null hypothesis calls into question one or more assumptions of the latent variable model in favor of the more agnostic SUR model. Conversely, failure to reject the null hypothesis suggests that the causal pathways implied by the latent variable model are approximately the same as the three distinct average treatment effects depicted in the SUR model (i.e., $\beta \lambda_k \approx \beta_k$). On theoretical grounds, the latent variable model may not be applicable to every experiment that measures multiple outcomes. But when the latent variable model is plausible for a given application, the fact that its parameterization has testable implications means that researchers need not rely solely on theoretical intuition when choosing between the latent variable model and the SUR model.
\section{Application}
Like stoetzer2022causal, we draw our empirical example from the kalla2020reducing field experiment designed to assess the effects of two different door-to-door canvassing interventions on subjects' views concerning immigrants. Kalla and Broockman's Study 1 shows that canvassing conversations that feature a “non-judgmental exchange of narratives” produce persistent changes in subjects' attitudes about immigrants, whereas otherwise similar conversations that omit this persuasive strategy seem to produce weaker effects. One attractive feature of this study is the fact that it devoted considerable attention and resources to outcome measurement. As summarized in Table (ref), five questions gauged respondents' attitudes towards undocumented immigrants. A separate battery of three questions measured respondents' views about policy proposals concerning expediting paths to citizenship and reducing threats of deportation.
The two treatment conditions (full treatment $Z_1$ and abbreviated treatment $Z_2$) in conjunction with two primary outcome scales (attitudes $Y_1$ and policy views $Y_2$) and a pre-treatment covariate give us five observed variables. If we model these data by assuming that the attitude index and the policy views index both measure a common underlying factor, the two parameters of central interest are the average treatment effect of full treatment on the latent outcome and the average treatment effect of abbreviated treatment on the latent outcome. The latent variable model is depicted in the left pane of figure (ref).
Let us review and evaluate the assumptions under which both causal parameters in the latent variable model are identified: the treatments are randomly assigned and therefore independent of latent potential outcomes; the treatments affect measured outcomes only insofar as they influence the latent outcome (the exclusion restriction, which implies that the measurement errors for each observed outcome ($\epsilon_{ij}$) are independent of the interventions); the stable unit treatment value assumption (SUTVA), which implies no spillovers between subjects; at least one of the two treatments has a nonzero ATE on the latent outcome, or the covariate is predictive of outcomes (but not the measurement errors); and observed outcomes are linear functions of the latent outcome. Next, let us evaluate the plausibility of these assumptions.
\begin{figure}[!h]
\caption{Graphical Depiction of kalla2020reducing with Additive Measurement Errors (left) and a Seemingly Unrelated Regression Model (right).}
\end{figure}
In this application, random assignment is satisfied by the experimental design and remains tenable in the wake of attrition, which seems to be unrelated to treatment assignment. The exclusion restriction implies that the interventions do not affect the outcome measures except insofar as they affect the underlying factor stoetzer2022causal. This assumption could be jeopardized if subjects drew the connection between the treatment they received and the survey, but this study used an unobtrusive measurement strategy that did not alert participants to the connection between the series of surveys they completed and their incidental visit from a canvasser. SUTVA violations due to spillover effects seem unlikely to have a material effect on outcomes here given the separation among subjects from different households. The treatments and covariates are clearly predictive of outcomes, so the scaling parameters are identified.
As for the linearity assumption, one of the study's strengths is its use of multi-item outcome measures. These measures gauge outcomes with sufficient granularity that we can visually inspect the scatterplot between $Y_1$ and $Y_2$. As shown in (ref), fitting a flexible LOESS curve through this plot reveals a near-linear pattern that coincides with the fitted least-squares regression line. A linear mapping from the latent outcome to each of the observed outcome measures seems plausible in this application.\footnote{Tables A.2 and A.4 may be used to calculate the proportion of observed variance in the outcomes that is attributable to the latent variable $\eta_i$. For attitudes, this proportion is 0.74, and for policy views it is 0.83.}
With these identifying assumptions, consistent estimation is straightforward.
If we declare the scale of the latent outcome variable to be in the same units as the attitude index (i.e., $\lambda_1 = 1$), the ATE of the full treatment on the latent factor can be estimated consistently using the approach described in Section (ref).\footnote{Because this application features two randomly assigned interventions ($Z_{i1}$ and $Z_{i2}$) that are correlated, consistent estimates are obtained using multiple regression rather than bivariate regression.}
The ALTE of the abbreviated treatment is consistently estimated in an analogous fashion. The fact that we have more information than we need to identify certain parameters such as $\lambda_2$
means that can (1) compare alternative estimates of each parameter and (2) leverage the overidentifying information to obtain more precise estimates. Comparing alternative estimates is one way to assess model fit; when two or more estimates of the same parameter diverge, the implication is that one or more underlying assumptions are incompatible with the data.
Estimates of the two interventions' average treatment effects on the latent outcome are presented in Table (ref). The first two columns present estimates obtained using maximum likelihood, and columns (3) and (4) report WSI estimates.\footnote{When creating the weighted average, we estimate the scaling parameter using method-of-moments. The optimal weight $\omega^*_j = \frac{\hat{\lambda}^2_j/\hat{\sigma}^2(\epsilon_{\cdot j})}{\sum_{j=1}^k \hat{\lambda}^2_j/\hat{\sigma}^2(\epsilon_{\cdot j})}$, where the variance of measurement error is calculated from MLE.} Columns (5) and (6) report SEM estimates, this time using a GLS estimator rather than MLE and using bootstrapped standard errors. Each estimator is presented using two specifications, one without covariate adjustment and one adjusting for respondents' baseline attitudes about immigrants (See SI (ref)). Comparing columns (1) and (3) shows that MLE and WSI generate almost identical estimates.
Comparing columns (2) and (4) again shows that the two estimators render similar estimates. Covariate adjustment greatly improves precision and clearly indicates that the full treatment was effective in changing subjects' opinions. The MLE covariate-adjusted estimates also suggest that the full treatment was more effective than the abbreviated treatment, with a differential effect of 0.341 (SE=0.141). The MLE estimates and standard errors are almost identical to those obtained using GLS and bootstrapped standard errors.
\begin{table}[!htbp]
\begin{threeparttable}
\caption{Estimation Results: MLE, Weighted Index, and GLS}
\begin{tabular}{@{\extracolsep{5pt}}lcccccccc}
\\[-1.8ex]\hline
\hline
\\[-1.8ex] & \multicolumn{6}{c} \\
& {MLE} & {MLE} & {WSI} & {WSI} & {GLS} & {GLS} \\
\\[-1.8ex] & (1) & (2) & (3) & (4) & (5) & (6) \\
\hline \\[-1.8ex]
treat (full) & {0.384*} & {0.431***} & {0.384*} & {0.430***} & {0.384*}&{0.431***} \\
& (0.213) &(0.099) & (0.196) & (0.099) & (0.213) & (0.096)\\
treat (brief) & {0.224} & {0.090} & {0.225} & {0.090} & {0.224} & {0.090} \\
& (0.206) & (0.101) & (0.200) & (0.101) &(0.215) & (0.102)\\
baseline covariate & & {0.662***} & & {0.662***} & & {0.662***}\\
& & (0.011) & & (0.009) & & (0.012)\\
$\lambda_2$ & 1.653 & 1.549 & 1.652 & 1.549 & 1.653 &1.549 \\
\hline \\[-1.8ex]
$\chi^2$ $p-value$ & 0.923 & 0.978 & & & 0.923 &0.978 \\
Adjusted $R^2$ & & & 0.001 & 0.762 & & \\
Degree of freedom &1 & 2 & & & & \\
\hline
\hline \\
\end{tabular}
\begin{tablenotes}
• {\it Note}: Sample size $N=1,578$. $\lambda_2$ in columns 1, 2, 5, and 6 are estimated by SEM; $\lambda_2$ in columns 3 and 4 are estimated by 2SLS using the two treatments and, where applicable, covariates as instrumental variables. For GLS, the standard errors are computed from 10,000 bootstraps. The MLE models are estimated by R package lavaan, and the GLS models are estimated by R package systemfit. *** $p<0.01$, ** $p<0.05$, and * $p<0.1$.
\end{tablenotes}
\end{threeparttable}
\end{table}
Next, we compare the results from Table (ref) to estimates obtained using the seemingly unrelated regression model, which makes no assumptions about why the two outcomes may be correlated.
Table (ref) presents two different SUR specifications, with and without covariate adjustment. The first column, without covariate adjustment, reports two separate OLS regressions in which attitudes ($Y_1$) and policy views ($Y_2$) are regressed on the two treatment indicators. The second column repeats these two regressions, this time including the baseline covariate. Although the estimated treatment effects look different in Table (ref) and Table (ref), some simple manipulations show that the two tables tell very similar stories. The estimated effects on $Y_{i1}$ are scaled identically in the two tables, and the two estimates are quite close.
The estimated effects on $Y_{i2}$, however, differ because the latent variable model applies that scaling factor $\widehat{\lambda}_2=1.653$ when estimating the ALTE in column (1) of Table (ref). If we divide the SUR estimate by this estimate of $\lambda_2$, we come very close to the ALTE estimate reported in Table (ref). The same holds for all of the coefficients reported in the two tables, implying that the modeling constraints of the latent variable model fit are quite compatible with the agnostic SUR model, which by construction fits the observed variance-covariance matrix perfectly. A more formal test of the fit of the latent variable model against the fit of the SUR model comes to the same conclusion: the $p$-value of the likelihood-ratio test is $0.980$, so we have no basis to reject the adequacy of the (more parsimonious) latent variable model.
\begin{table}[!htbp]
\caption{Estimation Results: Seemingly Unrelated Regression}
\begin{center}
\begin{tabular}{l c c }
\\[-1.8ex]\hline
\hline
& \multicolumn{2}{c} \\
& Model 1 & Model 2 \\
\\[-1.8ex] \hline
equation 1: treat (full) & $0.383$ & $0.418^{***}$ \\
& $(0.214)$ & $(0.120)$ \\
equation 1: treat (brief) & $0.231$ & $0.088$ \\
& $(0.218)$ & $(0.122)$ \\
equation 2: treat (full) & $0.635$ & $0.689^{***}$ \\
& $(0.335)$ & $(0.192)$ \\
equation 2: treat (brief) & $0.364$ & $0.143$ \\
& $(0.342)$ & $(0.196)$ \\
equation 1: baseline covariate & & $0.662^{***}$ \\
& & $(0.011)$ \\
equation 2: baseline covariate & & $1.025^{***}$ \\
& & $(0.018)$ \\
\hline
eq1: R$^2$ & $0.002$ & $0.688$ \\
eq2: R$^2$ & $0.002$ & $0.673$ \\
eq1: Adj. R$^2$ & $0.001$ & $0.687$ \\
eq2: Adj. R$^2$ & $0.001$ & $0.673$ \\
\hline
\hline
\end{tabular}
\begin{tablenotes}
• {\it Note}: In Model 1, two equations represent two regression models: eq1 is attitudes $\sim$ treat (full) + treat (brief) and eq2 is policy views $\sim$ treat (full) + treat (brief). In Model 2, we add the baseline covariate. *** $p<0.01$, ** $p<0.05$, and * $p<0.1$.
\end{tablenotes}
\end{center}
\end{table}
Our final specification check is to assess whether the measurement properties of the outcome measures operate symmetrically for all three experimental groups. The null hypothesis is that the measures operate in the same way across the three assigned groups; the alternative model is that the measurement parameters -- the scaling factors and the measurement error variances -- differ across groups. A total of 6 degrees of freedom differentiate the two models. The $p$-value of the model comparison is 0.171, suggesting that there may be some concern about differential measurement across groups, but the gap in terms of goodness-of-fit is not decisive.\footnote{In their analysis of differential item functioning when applying a hierarchical IRT model to these data, stoetzer2022causal found similarly equivocal results.}
\section{Conclusion}
This paper has shown how experiments with redundant outcome measures may be analyzed using latent outcome models while retaining the agnosticism of a design-based framework. We identify a key problem of current methods for multiple measurements -- their estimands are often study-specific, hindering the accumulation of scientific knowledge. In contrast, our WSI estimator preserves comparability across experimental replications while at the same time offering other advantages, such as improved statistical precision.
Although our nonparametric identification approach is agnostic about whether the observed measures are linearly related to the latent outcome, in practice linearity is an attractive modeling assumption.
Whether linearity is a plausible assumption for a given application depends on what is being measured and how; crucially, linearity may be addressed during the design stage of an experiment and assessed empirically. For example, outcome measures may be constructed from batteries of survey questions, and the resulting indices may be sufficiently granular to warrant a linear modeling framework. In such cases, identification and estimation of causal effects become straightforward. In other words, an investment in outcome measurement pays dividends when it comes time to analyze the experimental results with less reliance on restrictive parametric assumptions, such as those invoked by an IRT model stoetzer2022causal.
A further advantage of our framework is that it allows the researcher to make use of statistical tools that are routinely used in experimental analysis, ordinary least-squares regression and instrumental variables regression. Instrumental variables regression may be used to estimate the scaling parameters that link the latent outcome to the observed outcomes, which in turn may be used to estimate the error variances associated with each observed outcome measure. With estimates of these scaling parameters and error variances in hand, a researcher may easily create an optimally weighted average of the observed measures of the latent outcome that has an interpretable scale. This WSI may be regressed on the treatments and covariates in the usual way, and the relationship between interventions and the latent outcome are easily visualized using scatterplots. Alternatively, structural equation models may be used to estimate the scaling parameters, error variances, and causal parameters in a single statistical procedure. Both approaches produce consistent estimates of the average treatment effect on the latent outcome.
Another important message of this paper is the value of overidentification via additional outcome measures, treatments, or covariates. These layers of redundancy both facilitate robust estimation of scaling parameters and create opportunities for goodness-of-fit tests that can help assess the empirical adequacy of the posited latent variable model. Failure to pass such tests implies that future studies need to invest in new or improved outcome measures.
Latent variable models are optional but potentially helpful. When an outcome cannot be observed directly, researchers should strive to design experiments that measure this latent outcome in multiple ways. For decades, statisticians have pointed out that one way to increase the power of an experiment is to measure the outcome more reliably. The algebra in section (ref) reiterates this point, showing the gains in precision that occur with each additional measure of a latent outcome. When maximizing the precision with which the average treatment effect on the latent outcome is estimated, it sometimes makes more sense to invest in additional outcome measures rather than additional subjects.
This paper has focused on the challenge of measuring a single latent outcome, but the framework may be expanded to designs in which two or more latent outcomes are posited (See the SI (ref) for an example). The development of valid and reliable measures of multiple latent constructs has a long history in psychometric research anderson1988structural,campbell1959convergent. The details of this type of investigation go beyond the scope of this paper but typically involve an iterative process of proposing measures that seem to tap into a theoretically defined construct (face validity), properly distinguish subjects that are known to differ on the latent dimension of interest (construct validity), and show the expected patterns of high and low correlations with measures of other constructs (convergent and discriminant validity) chan2014standards. The systematic study of measurement involves a combination of theoretical and empirical steps that may be conducted outside the confines of an experiment but potentially provides valuable insights for experimental applications.
\putbib