Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
97,972 characters · 21 sections · 61 citation commands
Program Evaluation with Remotely Sensed Outcomes
\def\spacingset#1{ {#1}} \spacingset{1}
\spacingset{1.5}
While traditional program evaluations rely on surveys to measure impact, important economic outcomes such as living standards and environmental quality may be costly or infeasible to collect at scale. As a consequence, researchers increasingly estimate treatment effects on economic outcomes using remotely sensed variables (RSVs). Examples include night lights as a measure of local economic activity chen2011using, henderson2012measuring, asher2021development, roofing material as a measure of housing quality marx2019there, michaels2021planning, huang2021using, mobile phone transactions as a measure of local wealth or consumption blumenstock2015predicting, aiken2025estimating, and satellite images as a measure of pollution currie2023caused, deforestation jayachandran2017cash, assunccao2023optimal, fires jack2022money, balboni2024origins, flooding chen2017validating, patel2024floods, and local poverty jean2016combining.\footnote{Figure (ref) illustrates the rapid rise of papers published in economics journals and general interest science journals using RSVs. Of those published in economics journals, we find that 90% use high-dimensional satellite images as RSVs, and 40% use the RSVs as the main outcome in their empirical analyses. See also burke2021using and jack2023remotesensing for recent reviews on the empirical uses of RSVs.} Our research question is how researchers should rigorously estimate treatment effects from remotely sensed outcomes.
A recurring empirical practice appears in about 50% of papers in general interest economics journals from 2015-2024 that use remotely sensed outcomes.\footnote{We survey AEA journals, Econometrica, Journal of Political Economy, Quarterly Journal of Economics, and Review of Economic Studies. The other 50% of papers use a similar logic, but without an explicit formula for data combination.} This common practice predicts the economic outcome from the RSV and then uses the predicted outcome in lieu of a true outcome measurement in an experimental sample. Researchers often form such predictions using an auxiliary, observational sample collected in some other context, which contains the RSV and linked outcome measurements. The predictor is typically complex, e.g., a deep learning algorithm with unknown statistical properties.
We show that this intuitive method can produce arbitrarily biased treatment effect estimates---including flipped signs---when the RSV is a post-outcome variable. Using the predicted outcome in lieu of the true outcome implicitly uses the RSV as a surrogate that mediates between the treatment and the outcome prentice1989surrogate, athey2024surrogate. However, in many empirical applications using RSVs, the opposite is more plausible: the treatment affects the outcome, and both may affect the RSV. The bias is fundamentally due to this reversal; it is present even without machine learning.
As an example, consider a binary outcome indicating whether a plot of land has been burned balboni2024origins, jack2022money and an RSV summarizing the color saturation in a satellite image. Fires cause changes in satellite images, but not vice versa; the color saturation of satellite images is a post-outcome variable. Common practice predicts fires in the experimental sample using a machine learning algorithm trained on labeled satellite images from an observational sample, and then computes the difference in predicted outcomes between treated and control units in the experimental sample. Because the RSV is post-outcome, this method's estimand combines two quantities: the desired effect of the treatment on the outcome, and the correlation between the RSV and the outcome. In the extreme case where the RSV fails to predict the outcome---when an ideal method should report infinite standard errors---this popular method will instead report a precise estimate of zero, regardless of the true treatment effect. In general, this method may flip the sign of the true treatment effect.
Our main contribution is a novel formula to nonparametrically identify treatment effects using RSVs by combining (i) an experimental sample where the outcome is missing, and (ii) an observational sample in which the treatment is non-randomized and possibly missing. Our key assumption formalizes the logic underlying the examples above: the conditional distribution of the RSV given the outcome and the treatment is stable across both samples. Consequently, the relationship between the RSV and the outcome can be learned from the observational sample and transported to the experimental sample. If the treatment is missing in the observational sample, then we need an additional assumption to restore identification: the treatment only affects the RSV through the outcome.
Our main identifying assumptions are jointly testable and lend themselves to simple diagnostics, allowing researchers to assess their plausibility in applications. We propose a diagnostic to evaluate whether an RSV is sufficiently relevant for an economic outcome to justify its use in program evaluation. Our framework also naturally extends to quasi-experiments, including instrumental variables and difference-in-differences designs, which we present in Appendix (ref).
Our secondary contribution is to characterize the representation of the RSV that achieves valid, precise, and robust $n^{-1/2}$ inference on the treatment effect. Because modern remote sensing typically involves unstructured data and complex machine learning, we derive valid inference without rate conditions and without complexity restrictions on RSV-based predictions. Valid inference only requires that (i) a learned RSV representation has some limit, and (ii) the limit predicts the outcome of interest, for which we provide a diagnostic. This enables researchers to use complex deep learning algorithms and still conduct valid inference on treatment effects.
Specifically, we build on our main identification result to establish the connection between modern remote sensing and classical conditional moments chamberlain1987asymptotic,newey1993efficient. This allows us to derive an expression for a simple representation of the RSV that maximizes precision from the conditional moments we derived. We find that three predictions are necessary for efficient downstream causal inference based on RSVs: predictions of the outcome, the treatment, and the sample indicator given the RSV. By contrast, common practice only predicts the outcome given the RSV. Moreover, $n^{-1/2}$ inference remains valid even if all three predictions are misspecified, which is a strong form of robustness. We provide the remoteoutcome R package to implement our method.
Finally, we conduct a semi-synthetic exercise, calibrated to an existing field experiment in India, which we merge with existing satellite images. Following muralidharan2016building, muralidharan2023general, we study the effect of Smartcards, a biometrically authenticated payments infrastructure, on village-level poverty measures. We use the geographic coordinates of each village to extract nighttime luminosity and high-dimensional, pre-trained embeddings of satellite images. Despite using outcomes for only half of the units, we recover the treatment effect of interest with the same precision as an unbiased regression method that has access to outcomes for all units. By contrast, common practice can have positive or negative bias for the treatment effect. Our method enables large savings in survey costs and therefore opens up possibilities for program evaluation in cost-constrained environments.
Our framework differs from existing approaches to auxiliary variables and data combination along two key dimensions: (i) the causal direction between the auxiliary variable and the outcome, and (ii) the data requirements.
In the surrogacy framework prentice1989surrogate, athey2024surrogate, kallus2024role, the auxiliary variable (surrogate) is a mediator between the treatment and the outcome, whereas the RSV is a post-outcome variable in our framework. We show that misusing a post-outcome RSV as a surrogate leads to arbitrary bias in treatment effects. The negative control literature extends the surrogacy framework to address unobserved confounding ghassami2022combining,imbens2024long, yet such extensions face the same limitation. See Remark (ref) for details.
Compared to the vast literature on data combination cross2002regressions, RidderMoffitt(07), bareinboim2016causal, d2024partially and nonclassical measurement error ChenHongNekipelov(11),schennach2020mismeasured, we place what appears to be a different key assumption. Several influential works handle measurement error in moment condition models by using auxiliary data and assuming that the conditional distribution of the variable of interest, given the imperfect measurement, is stable across samples chen2005measurement,chen2008semiparametric,GrahamPintoEgel2016, akin to the surrogacy framework. Our key identifying assumption is the opposite: the conditional distribution of the imperfect measurement, given the variable of interest, is stable across samples. This is natural when the imperfect measurement is caused by the outcome (as with satellite images of poverty or environmental outcomes) rather than merely correlated with it. Other works focus on mismeasured or missing covariates in a sample where the outcome is observed fan2014identifying,battaglia2024inference,d2024partially,d2024linear. By contrast, in our setting, the outcome is missing yet a post-outcome variable is present. Both of these distinctions lead to a novel identifying formula.
The main difference between our RSV framework and the prediction-powered inference (PPI) framework angelopoulos2023prediction,lu2025regressioncoefficientestimationremote,kluger2025prediction is also along these lines: the PPI framework uses machine learning predictions as surrogates ji2025predictions. Another difference concerns data availability. In our terminology, the PPI approach would require the researcher to observe the treatment, outcome, and RSV for a random subsample of experimental units. The data requirements in other works are similar to those in the PPI literature fong2021machine,allon2023machine,gordon2023remote,egami2024using, carlson2025unifying. By contrast, we allow the researcher to observe no outcomes for any experimental units, enabling inference in settings where PPI-style techniques cannot be applied.
Our results do not require a correctly specified generative model of how treatments and outcomes affect RSVs, which may be prone to misspecification when the RSV is a satellite image or some other type of unstructured data. Several previous works propose methods based on generative modeling that do require correct specification gentzkow2019measuring,alix2023remotely,proctor2023parameter. Similarly, methods for causal inference on outcomes that are latent concepts require a correctly specified generative model egami2022make,knox2022testing,stoetzer2024causal.
The researcher observes units in two samples, indicated by the variable $S\in \{e, o\}$: an experimental sample ($S=e$) and an observational sample ($S=o$).
Within the experimental sample $(S=e)$, we observe pre-treatment covariates $X \in \mathcal{X}$ and a binary treatment $D \in \{0, 1\}$. However, the outcome $Y\in \mathcal{Y}$ is missing.\footnote{For ease of exposition, we focus on the case where the outcome $Y$ is completely missing in the experimental sample. Remark (ref) gives the extension where $Y$ is only partially missing in the experimental sample.} In its place, we have access to a remotely sensed outcome variable $R\in\mathcal{R}$. We typically think of $R$ as high-dimensional (e.g., unstructured data such as satellite images), but it could be low-dimensional (e.g., the output of some pre-trained machine learning algorithm). The researcher would like to use the remotely sensed variable (RSV) as an imperfect measurement of the outcome in the experimental sample, without placing parametric assumptions on their relationship.
The causal parameter of interest is the effect of the treatment $D$ on the outcome $Y$ in the experimental sample. Though the outcome $Y$ is unobserved in the experimental sample, we may still define its potential outcomes $Y(d)$ and our parameter of interest.\footnote{This definition of potential outcomes has no spillovers. We discuss spillovers in Appendix (ref).}
Without further assumptions, point identification for this causal parameter is impossible horowitz1995identification. Even if $R$ can predict $Y$ with great accuracy, if the prediction is at all imperfect, then an assumption is necessary. Therefore, we place additional structure on the problem, inspired by recent empirical work in environmental and development economics.
A popular practice in environmental and development economics is to use an auxiliary dataset of outcomes and RSVs, e.g., of labeled satellite images. We refer to the auxiliary dataset as the observational sample $(S=o)$. For these units, we observe baseline covariates $X$, the outcome $Y$, and the remotely sensed variable $R$. We may or may not observe the treatment $D$. If we do, we denote it by $D\in\{0,1\}$ and refer to this scenario as having “complete” cases. If we do not, or if treatment is deterministic in the observational study, we set $D=0$ for all units in the observational study and refer to this latter scenario as having “incomplete” cases.\footnote{Incomplete cases in the observational sample may refer to three scenarios. First, the treatment status may be missing. Second, the treatment status may be present, and all observational units are untreated, hence $D=0$. Third, the treatment status may be present, and all observational units are treated. Relabeling treatment values gives $1-D=0$.} As we discuss below, incomplete cases will require stronger assumptions.
Table (ref) summarizes the setting. Each unit is characterized by the random vector $\left( S, X, D, Y(0), Y(1), R \right)$, which we assume to be independent and identically distributed.\footnote{Independence is not used to derive our main identification argument.} For units in the experimental sample ($S = e$), we observe $(X, D, R)$; for units in the observational sample ($S = o$), we observe $(X, D, Y,R)$ in complete cases or $(X,Y,R)$ in incomplete cases.
We formalize this causal setting via three assumptions. Our identifying assumptions allow $(X,Y,R)$ to be discrete or continuous. For readability, we slightly abuse notation: for a random variable $W$, we use the symbol $f_W(\cdot \mid ...)$ to refer to its (conditional) probability mass function if $W$ is discrete, or its (conditional) density if $W$ is continuous.\footnote{Formally, $f_W(\cdot \mid ...)$ is our symbol for the Radon-Nikodym derivative.}
In many empirical applications involving RSVs, such as Examples (ref) and (ref), Assumption (ref) is satisfied by design: experimental units are chosen as aggregates without spillovers, e.g., villages, and the treatment is randomly assigned for these experimental units. We defer to Appendix (ref) the study of quasi-experimental settings, such as instrumental variables and difference-in-differences.
Under Assumption (ref), if we were to observe the outcome in the experimental sample, the ATE could be identified using standard arguments. However, the outcome is not observed in the experiment; instead, we have an RSV.
We resolve this measurement issue by leveraging the observational sample. Intuitively, the idea is to learn the relationship between the RSV and the outcome of interest in the observational sample and to “transport” it to the experimental sample.
Assumption (ref)(i) is the main assumption of our framework: the conditional distribution of the remotely sensed variable $R$, given $(X, D, Y)$, is stable across the experimental and observational samples. This allows us to “transport” the measurement error distribution from the observational sample to the experimental sample. Importantly, this condition does not require stability of the underlying treatment effects, which may differ across samples. See Remark (ref) for comparisons between Assumption (ref)(i) in the RSV model and assumptions in alternative models.
Returning to our two leading examples, Assumption (ref)(i) requires that the conditional distribution of tree cover pixels $R$, given environmental outcomes $Y$ and interventions $D$ (as well as other pre-treatment covariates), is stable across the experimental and observational samples. Analogously, it requires that the conditional distribution of the satellite image $R$, given village-level poverty $Y$ and the anti-poverty program $D$ (as well as other pre-treatment covariates), is stable across the experimental and observational samples.
Our main assumption is empirically plausible, as illustrated by Figure (ref). We use data from an anti-poverty program in India, where outcomes are observed. If Assumption (ref) holds, then the conditional densities of the RSV given the outcome and treatment should be the same across the experimental and observational samples. They appear to coincide in this empirical setting.\footnote{When $X=\varnothing$, Assumption (ref) imposes four equalities: $f_R(R \mid S=e,D=d,Y=y)=f_R(R \mid S=o,D=d,Y=y)$ for $d\in\{0,1\}$ and $y\in\{0,1\}$. Each equality can be evaluated with a diagnostic plot if outcome data are available. For example, using the experimental and observational units satisfying $D=0$ and $Y=0$ in Figure (ref), we can visualize whether the density of $R\mid S=e, D=0, Y=0$ aligns with the density of $R \mid S=o, D=0, Y=0$ in Figure (ref). Since $R$ is high-dimensional, we simplify the visualization by comparing the densities of its first principal component.}
The remaining aspects of Assumption (ref) are weak regularity conditions. Assumption (ref)(ii) requires that the outcome in the observational sample has a common (or larger) support than the outcome in the experiment. Assumption (ref)(iii) ensures that the RSV distribution does not degenerate for any stratum. Assumption (ref)(iv) requires that we observe some data from both the experimental and observational samples.
Assumptions (ref) and (ref) imply identification when the observational sample has complete cases. When the observational sample has incomplete cases, we require a further assumption. In other words, if the treatment is missing or deterministic in the observational sample, then a further restriction is necessary for point identification.
Assumption (ref) imposes only one of two conditions.
Assumption (ref)(i) implies that we have access to complete cases, i.e., some observations of $(X,D,Y,R)$ where $D$ has variation. Within the observational sample, the treatment is observed and varies, although it does not need to be randomized and may suffer from unobserved confounding. Whenever we have complete cases, under Assumption (ref)(i), no further causal assumptions are needed. In particular, the treatment $D$ may have a direct effect on the remotely sensed variable $R$. In Example (ref), this would allow the environmental program to affect satellite images both indirectly, i.e., via crop burning, and directly, e.g., via visible investments in farm equipment.
If Assumption (ref)(i) is violated, then we have no complete cases, i.e., no observations of $(X,D,Y,R)$ where $D$ is variable. Without joint observations of the outcome and treatment, a further restriction is needed. Assumption (ref)(ii) fills this gap, requiring that the treatment $D$ affects the remotely sensed variable $R$ only via its effect on the outcome $Y$. In Example (ref), we may be comfortable assuming that the PES contract has no direct effect on the specific infrared band used to measure charred soil in satellite images. Assumption (ref)(ii) may also become more plausible when $Y$ is a vector of outcomes. Several outcomes may approximate all mechanisms through which the treatment affects the RSV. For exposition, we focus on scalar outcomes in the main text and we generalize to vector outcomes in Appendix (ref).
Together, Assumption (ref)(i) and Assumption (ref)(ii) imply that $\left( S, D \right) \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} R \mid X, Y$. In the next section, we will show that Assumptions (ref)(i) and (ref)(ii) are jointly testable, even when no outcome is observed from the experimental sample (Remark (ref)).
Figure (ref) illustrates our identifying assumptions as a causal graph. The treatment affects the outcome, which in turn affects the RSV. By Assumption (ref), the sample indicator does not change the conditional distribution of the RSV given the treatment and outcome. Depending on which version of Assumption (ref) is imposed, the treatment may have a direct effect on the RSV, as illustrated by the dotted line. Table (ref) summarizes the implications of our main assumptions.
In this causal setting, a commonly used procedure may lead to causal estimates with arbitrary bias. Motivated by this negative result, we prove a positive one: we nonparametrically identify the causal parameter by combining the experimental and observational samples differently.
To streamline notation, we initially focus on the setting where Assumption (ref)(ii) holds, then return to the setting where Assumption (ref)(i) holds at the end of this section. For exposition, we will also assume that the outcome is binary, i.e., $\mathcal{Y} = \{0,1\}$. Finally, we assume that the support of $R$ is at least as large as the support of $Y$ (it may be binary, discrete, or continuous). Appendices (ref) and (ref) extend our results to discrete and continuous outcomes, respectively.
In empirical research, it is common to use RSVs in two steps: (i) researchers train a predictor of the outcome $Y$ from the remotely sensed variable $R$ in the observational sample, and (ii) the predictor is applied to the experimental sample, where its predictions are used as surrogate outcomes to estimate treatment effects. While intuitive, this empirical strategy can lead to arbitrary bias for the ATE in the experimental sample.
Suppose there are no pre-treatment covariates for simplicity. The widely used two-step estimation procedure implicitly targets the estimand $\widetilde{\theta} = \widetilde{\mu}(1) - \widetilde{\mu}(0)$, where $\widetilde{\mu}(d) := \E\left\{ \E(Y \mid R, S = o) \mid D = d, S = e \right\}$ for $d \in \{0, 1\}$. Within this expression, the first step estimates the conditional expectation function $\E(Y \mid R=r, S = o)$. The second step evaluates and averages this function over the treated and untreated subgroups in the experimental sample.
As a starting point, if the RSV fails to predict the outcome, i.e., if $\E(Y|R,S=o)=\E(Y|S=o)$, then the implicit target $\widetilde{\theta}$ is zero regardless of the true treatment effect. Common practice would return a precise estimate of zero, even though the RSV provides no information about treatment effects in this case.
More importantly, even if the RSV does predict the outcome, the implicit target $\widetilde{\theta}$ can incur bias with arbitrary sign for the ATE in the experimental sample.
Proposition (ref) derives the bias of current empirical practice, which uses the RSV as a surrogate outcome. Under Assumptions (ref) and (ref)(ii), the conditional distribution of the RSV $f_R\left(R \mid Y, S = o\right)$ is stable across the experimental and observational samples. Existing practice attempts to transport the predictions $f_Y(Y \mid R, S = o)$ into the experimental sample. By Bayes' rule, this induces a bias due to differences in the marginal distributions of the outcomes and RSVs across the samples. Importantly, the bias of common practice arises even if units are randomly allocated into the experimental and observational samples.\footnote{See Appendix (ref) for a formal verification of this claim.}
We revisit an experiment conducted by jack2022money that studied whether payments for ecosystem services (PES) contracts can incentivize farmers to reduce crop burning. We measure the average treatment effect $\theta$ of being offered a PES contract $D \in \{0, 1\}$ on the likelihood that a farmer burns their fields $Y \in \{0, 1\}$. Here, $Y=1$ means not burning, so a positive $\theta$ means environmental benefit.
It is costly to measure whether the crop residue on a particular field has been burned, requiring a surveyor to make frequent visits to rural fields. Therefore, it is natural to turn to a remotely sensed variable $R$ for the outcome $Y$. To construct such a remotely sensed variable, jack2022money link surveyor-collected measurements of crop burning to satellite-based spectral indices and then train a supervised learning algorithm to predict whether these fields have been burned. Here, $R \in \{0, 1\}$ is a classifier for whether a field has not been burned, which applies a threshold rule to a machine learning prediction of the probability that a field has not been burned.\footnote{jack2022money construct two binary RSVs for crop burning by applying two alternative threshold rules to the estimated probability a field has not been burned. We use their “max accuracy” RSV in the main text, and we report analogous results in Table (ref) using the authors' “balanced accuracy” RSV.}
Our causal assumptions are plausible in this setting. Because burning crops would alter satellite images, but altering satellite images would not burn crops, $R$ is a post-outcome variable rather than a pre-outcome variable. Since access to PES contracts was randomized at the village level, Assumption (ref) is satisfied by design. Since the authors conducted randomized spot checks, surveying the outcome $Y$ for randomly selected fields, Assumption (ref) is also satisfied by design if we define the observational sample $S = o$ as the fields in which these random spot checks occurred. Finally, since specific infrared bands are used to form $R$, it is plausible that PES contracts affect these specific infrared bands only via crop burning, so Assumption (ref)(ii) is reasonable. Figure (ref) illustrates the two samples used in our re-analysis.
We implement the common empirical practice: we use the RSV as a surrogate for crop burning. Column (1) of Table (ref) estimates $\widetilde{\theta}$. Offering any PES contract to farmers appears to reduce crop burning by 7.9%.\footnote{We modify the specification of jack2022money in two ways. (i) While they distinguish between two types of PES contracts, we define the treatment as whether any PES contract was offered. (ii) They analyze the effects of PES contracts by defining a farmer-level outcome, whereas we analyze effects at the field level.}
However, by Proposition (ref), $\widetilde{\theta}$ is typically biased for $\theta$. To quantify the magnitude of its bias, we estimate $\beta := \E(R \mid Y = 1) - \E(R \mid Y= 0)$ using information collected through random spot checks of the fields. Under Assumption (ref)(ii), an algebraic argument in Appendix (ref) yields $\widetilde{\theta}= \beta \theta$ in this setting.\footnote{Compared to Remark (ref)(ii), here we have $\E(Y|R,S=o)=R$, which implies $\tilde{\beta}=1$ and hence $\widetilde{\theta}=\beta \theta_0$.} Column (2) of Table (ref) reports the estimate of $\beta$. It suggests that $\widetilde{\theta}$ understates the treatment effect of being offered any PES contract by approximately 47%. This motivates us to ask: what should empirical researchers do instead?
Our main result nonparametrically identifies the ATE in the experimental sample without a parametric model that restricts the distribution or dimensionality of the RSV. We derive a novel formula for combining the experimental and observational samples.
To begin, we express the RSV distribution in the experimental sample as an affine transformation of the potential outcomes we wish to identify. The slope and intercept depend on how the RSV relates to the outcome, and they are identified by the observational sample.
By Lemma (ref), we recover the ATE in the experimental sample by combining (i) how the RSV varies with the treatment in the experimental sample with (ii) how the RSV varies with the outcome in the observational sample. This combination leverages our key assumption: stability of the RSV across samples.
While Lemma (ref) identifies the ATE in the experimental sample, it suggests a challenging estimation problem: it involves the conditional distribution of the high-dimensional RSV. One path forward is to develop a complex parametric model, e.g., a generative model of the satellite image distribution conditional upon poverty measurements or environmental outcomes. Even if such a generative model could be developed, it would be prone to misspecification.
We follow a different path that does not require a generative model for $R$. We use Bayes' rule to rewrite Lemma (ref) as a conditional moment equation. This transformation avoids estimation of the RSV's conditional distribution and recovers a classic econometric estimation problem.
Let $\theta(x) := \E\{Y(1) - Y(0) \mid S = e, X =x\}=\mu(1,x)-\mu(0,x)$ denote the conditional average treatment effect in the experimental sample.
Theorem (ref) identifies the treatment effect as the solution to a set of conditional moment equalities. The conditional average treatment effect $\theta(x)$ balances treatment variation from the experimental sample $\Delta^{e}(x)$ and outcome variation from the observational sample $\Delta^o(x)$, so that their projections onto the remotely sensed variable $R$ match. Consequently, we can leverage a celebrated literature on conditional moment equalities chamberlain1987asymptotic, newey1993efficient for estimation and inference.
For intuition on Theorem (ref), consider the case without covariates. Then $\Delta^e$ is simply a scalar computed from the experimental sample, where the treatment is observed. Similarly, $\Delta^o$ is simply a scalar computed from the observational sample, where the outcome is observed. Theorem (ref) shows that $\theta_0=\frac{\E(\Delta^e|R)}{\E(\Delta^o|R)}.$ This ratio appears to be a new formula for data combination, where the numerator and denominator are from different samples. The numerator reflects how both the treatment and the outcome affect the RSV, so we must divide it by the effect of the outcome on the RSV, in order to isolate the desired causal parameter.
For identification, Theorem (ref) implies that we may introduce representations of the RSV. Such representations can be arbitrary, as long as they predict outcome variation.
Once again, for intuition on Corollary (ref), consider the scenario without covariates. In this case, $H(R)$ is simply a scalar representation of the RSV. Corollary (ref) shows that $\theta_0=\frac{\E\{\Delta^e H(R)\}}{\E\{\Delta^o H(R)\}}$. Any representation $H(R)$ is valid as long as the representation is predictive: $\E\{\Delta^o H(R)\}\neq 0$. A weak RSV test asks whether $\E\{\Delta^o H(R)\} \approx 0$, which can be tested using existing frameworks for weak instruments. In the same spirit, a joint test of our identifying assumptions asks whether two representations $H(R)$ and $H'(R)$ give similar estimates. Dissimilar estimates would be evidence against our assumptions.
By Corollary (ref), many representations of the RSV provide identification. However, naive choices may be inefficient, producing needlessly large standard errors. Section (ref) asks: what representation of the RSV should be chosen for precise, downstream causal inference? Our answer develops a connection between “representation learning” johannemann2019sufficient, vafa2024estimatingwagedisparitiesusing and classical results for conditional moment equalities.
So far, we have focused on the setting of Assumption (ref)(ii), which allows incomplete cases but disallows direct effects of the treatment on the RSV. Our main result also holds for the setting of Assumption (ref)(i), which requires complete cases yet allows direct effects of the treatment on the RSV. Recall that $\mu(d,x) := \E\{Y(d) \mid S = e, X = x\}$ is the conditional average potential outcome in the experimental sample.
Once again, the causal estimand reconciles treatment variation from the experimental sample with outcome variation from the observational sample, so that their projections onto the remotely sensed variable $R$ match. Theorem (ref) allows direct effects; the key assumption in our framework is stability in Assumption (ref). Corollary (ref) and Remark (ref) extend accordingly.
As a secondary contribution, we demonstrate that our identification result provides guidance on how to choose a representation of the RSV for valid, precise, and robust inference on the causal parameter. We point out and interpret the connection between modern remote sensing and classical conditional moment analysis, which allows us to use standard techniques. We derive valid $n^{-1/2}$ inference without rate conditions and without complexity restrictions on the researcher's RSV-based predictions, justifying the use of complex deep learning algorithms that may be misspecified.
For clarity, we focus on the simple case in which the outcome is binary, there are no pretreatment covariates, and Assumption (ref)(ii) holds. We discuss the more general case with discrete outcomes and discrete or continuous covariates in Appendix (ref).
Without covariates, our main identification result in Theorem (ref) simplifies. The treatment and outcome variation are simply scalars:
Since $(S,D,Y)$ are binary, the denominators can be estimated by simple counts. The treatment effect is identified as $ \theta=\frac{\E(\Delta^e|R)}{\E(\Delta^o|R)}$. Therefore, $\theta=\frac{\E\{\Delta^e H(R)\}}{\E\{\Delta^o H(R)\}} $ for any representation $H(R)$ that is predictive of outcome variation in the sense that $\E\{\Delta^o H(R)\}\neq0$. While any predictive representation can be used for inference, different choices have different efficiency properties. We now ask, which representation should researchers use in practice?
By rearranging our identifying formula, we highlight the connection to classical ideas in econometrics and thereby derive an answer. Specifically, the treatment effect $\theta$ may be viewed as the coefficient in the regression model
where the residual $\epsilon:=\Delta^e-\theta \Delta^o$ is mean zero conditional on $R$ by Theorem (ref). Results in chamberlain1987asymptotic and newey1993efficient demonstrate that the representation that achieves efficiency within a class of models satisfying (ref) is $H^*(R)=\frac{\E(\Delta^o \mid R)}{\sigma^2(\theta, R)}$, where $\sigma^2(\theta, R) := \E\{(\Delta^e - \Delta^o \theta )^2 \mid R\}$.
For downstream causal inference, we find that the efficient representation of the high-dimensional remotely sensed variable $R\in\mathcal{R}$ is a simple, continuous scalar $H^*(R)\in \R$. The formula for $H^*(R)$ is a conditional expectation divided by a conditional variance.
This connection has a direct consequence for empirical practice: to improve efficiency with the RSV, we should not only predict the outcome $Y$ using observational data, but also predict the treatment $D$ using experimental data, and predict the sample indicator $S$ using all data. The first prediction is part of common practice, but the second and third are not. By using all three predictions, we use the RSV more efficiently.
We confirm this connection by further developing the expression for $H^*(R)$: the numerator $\E(\Delta^o \mid R)$ contains $\Pr( Y = 1, S = o|R)$ while the denominator $\sigma^2(\theta,R)$ contains $\Pr( D = 1, S = e|R)$; see Lemma (ref) below. Finally, by the definition of conditional probability, $\Pr( Y = 1, S = o|R)=\Pr( Y=1 | S = o,R)\Pr(S=o|R)$ and $\Pr( D = 1, S = e|R)=\Pr( D=1 | S = e,R)\Pr(S=e|R)$.
For simplicity, we describe our inferential procedure using sample splitting with $\tr$ and $\te$ folds, though our proofs in Appendix (ref) allow for cross-fitting with any fixed number of folds. We first state our inferential procedure at a high level before filling in the details:
The overall structure is familiar angrist1999jackknife,chernozhukov2018double, although here we will show that no rate condition is required on $\widehat{H}(R)$. To simplify the algorithm statement, we introduce some notation for simple events and an algebraic identity.
Let $\E_{\tr}(\cdot)=\frac{1}{|\tr|}\sum_{i\in \tr}(\cdot)$ and $\E_{\te}(\cdot)=\frac{1}{|\te|}\sum_{i\in \te}(\cdot)$. Slightly abusing notation, the counting operation, $\#_{\textsc{event}}$, can be $\E_{\tr}(1_{\textsc{event}})$ in the second step or $\E_{\te}(1_{\textsc{event}})$ in the third step, which will be clear from the context.
See Appendix (ref) for explicit computations of $\widehat{\E}(\Delta^e|R)$, $\widehat{\E}(\Delta^o|R)$, and $\widehat{\sigma}^2(\widehat{\theta}_{\textsc{init}},R)$ in terms of the marginal probabilities, predictors, and initial estimate.
Sample splitting may be eliminated under complexity restrictions that tolerate simple machine learning procedures. See e.g., chernozhukov2020adversarial for a recent summary.
In empirical research, the RSV is typically unstructured and high-dimensional, e.g., a satellite image. The prediction is often conducted by a complex machine learning algorithm, e.g., a deep convolutional neural network. For this realistic setting, rates of convergence are often unknown. In other words, we have no reason to believe that $\textsc{pred}_Y(R)$ converges to $\Pr(Y=1|S=o,R)$, nor that $\widehat{H}(R)$ converges to $H^*(R)$. Even carefully crafted architectures positing a generative model would typically be misspecified.
For this reason, we place a weaker regularity condition: the predictions, and hence the representation estimator, have some probability limit; they may be misspecified.
Assumption (ref) does not require any complexity restriction, nor any rate of convergence. The moment restriction in (ref) is infinite-order Neyman orthogonal mackey2018orthogonal,chen2020mostly, so Algorithm (ref) enjoys the best of both worlds: no complexity restriction chernozhukov2018double,chernozhukov2023simple, and no rate requirement chamberlain1987asymptotic,newey1993efficient. Even if all three of its predictions are misspecified, Algorithm (ref) still delivers valid $n^{-1/2}$ inference, which is a strong form of robustness to misspecification.
In the following statements, we refer to $\Pr(D=1,S=e)$, $\Pr(D=0,S=e)$, $\Pr(Y=1,S=o)$, and $\Pr(Y=0,S=o)$ as the marginal probabilities.
When marginal probabilities are known, the asymptotic variance is standard (Proposition (ref)). In this case, (ref) matches the conditional moment of chamberlain1987asymptotic, so the semiparametric efficiency bound is known.
When marginal probabilities are unknown, the asymptotic variance is weighted by them (Proposition (ref)). The asymptotic variance in Proposition (ref) differs from the one in Proposition (ref) because estimation of the marginal probabilities introduces an additional estimation error of order $n^{-1/2}$ that we must take into account for asymptotic inference. In this case, it is simpler to use a bootstrap than an analytic variance estimator.
Theorem (ref) allows direct effects of the treatment on the remotely sensed variable, as long as treatment is non-missing and non-deterministic in the observational sample. We have already derived the conditional moment equations. Estimation and inference are similar.
To empirically validate our method, we conduct three semi-synthetic exercises that are increasingly realistic. We use real RSV distributions, together with
Across the exercises, we use data from an experiment analyzed in muralidharan2016building,muralidharan2023general, illustrated in Figure (ref). The authors collect data in an experimental sample and an observational sample of villages in Andhra Pradesh, India. The treatment $D$ is the early introduction of Smartcards, which are a biometrically authenticated payments infrastructure. The outcome $Y$ is a measure of village-level poverty in administrative data from 2012 to 2013 (after the experiment was deployed). See Appendix (ref) for details on the experiment.
We use the geographic coordinates of each village to extract a remotely sensed variable $R$, from free, open-access databases. The RSV concatenates nighttime luminosity measures from 2012 to 2020 (a vector in $\mathbb{R}^{50}$) asher2021development and satellite images from 2019 (a high-dimensional, pre-trained embedding vector in $\mathbb{R}^{4000}$) rolf2021generalizable. These variables have been extensively validated as predictors of poverty hsw2009,jean2016combining,stoeffler2016reaching,michaels2021planning,huang2021using,sherman2023global. Our framework allows us to test their relevance for program evaluation, similar to a “first stage” exercise in instrumental variable analysis. In our semi-synthetic exercises, we will evaluate how well we can conduct program evaluation using this real RSV.
In terms of our identification framework, the semi-synthetic settings have incomplete cases (Assumption (ref)(ii)): the treatment is missing in the observational sample. The settings also have some experimental outcomes (Remark (ref)). We will be explicit about sample definitions below. Appendix (ref) contains a complete description of the data.
We impose all of our assumptions in a calibrated, synthetic data-generating process (DGP). First, we simulate the binary treatment $D$ as a fair coin toss (Assumption (ref)). If $D=0$, we simulate the binary outcome $Y$ as a weighted coin toss, calibrated to the empirical probability of $Y=1$ among untreated experimental units in the real data. If $D=1$, we simulate $Y$ as a different weighted coin toss, calibrated to the empirical probability of $Y=1$ among treated experimental units in the real data.
In this baseline version of the DGP, the synthetic treatment effect is calibrated to the real one: $-0.07$. We consider additional synthetic treatment effect values $\theta=-0.07+\tau$ by augmenting the probability of $Y=1$ when $D=1$ for alternative values of $\tau$.
Next, we draw RSV values from the real data. If $Y=0$, we draw $R$ from the empirical distribution of $R|Y=0$ in the real experimental data. Likewise, we draw $R$ when $Y=1$. This imposes no direct effects (Assumption (ref)(ii)). For computational feasibility in this exercise, we use luminosity and only the initial $1,000$ satellite image features as our RSV.
Finally, to mimic a setting with missing outcomes, we delete $Y$ if $D=1$. In terms of Remark (ref), we set $\tilde{S}=e$ when $D=1$ and $\tilde{S}=\{e,o\}$ when $D=0$.\footnote{For simplicity, we do not have $\tilde{S}=o$. This could be easily done: if $D=0$, toss a coin to determine whether $\tilde{S}=\{e,o\}$ or $\tilde{S}=o$. With or without $\tilde{S}=o$, a direct extension of Theorem (ref) applies: within the identifying formula given there, simply replace $S=e$ with $e\in\tilde{S}$ in $\Delta^e$, and similarly replace $S=o$ with $o\in\tilde{S}$ in $\Delta^o$.} This imposes stability across samples (Assumption (ref)). Recent methods that use machine learning predictions as outcomes, e.g., prediction-powered inference angelopoulos2023prediction, provide no guidance for this DGP, because no treated unit has an observed outcome.
Figures (ref) and (ref) demonstrate that our method outperforms common practice across treatment effect values and sample sizes. We consider treatment effect values $\theta=-0.07+\tau$ with $\tau\in \{0,0.1, \ldots, 0.5\}$ and sample sizes $n\in\{1000,2000,3000\}$. By varying these aspects of the synthetic DGP, we evaluate relative performance for different signal-to-noise ratios.
We compare the bias and root mean square error of the two methods. Whereas the bias of our method is always small and vanishing with sample size, the bias of the common practice is typically very large and constant across sample sizes. Common practice has positive or negative bias for the treatment effect, consistent with Proposition (ref). While the variance of our method is similar to the variance of the common practice, our method's large improvement in bias translates into similar improvement in mean square error when the sample size is large enough.
This exercise has clear consequences for empirical practice. If the goal is to conduct valid inference on the treatment effect, common practice can be misleading due to large bias that does not vanish as the sample size increases. It can return the wrong sign, as characterized by Proposition (ref). By contrast, our method reports unbiased estimates, with valid asymptotic inference.
While our first exercise involves synthetic treatment effects, our second exercise involves real treatment effects. We now ask: compared to an unbiased benchmark, how does our method perform? The benchmark is the difference-in-means estimate and confidence interval an economist would obtain if they could observe all treatments and outcomes in the experiment. Our method gives an estimate and confidence interval using treatments and RSVs in an experiment, and outcomes and RSVs in an observational sample.\footnote{Following Remark (ref), we also observe some outcomes in the experiment, as described below.}
We conduct this exercise with three possible poverty measurements: (i) is the village in the bottom quartile of villages for per capita consumption, (ii) does a village have only low income households, and (iii) does a village have only low and middle income households. While (i) is a measure of average consumption, (ii) and (iii) describe the income distribution using formal categories from Indian administrative data; see Appendix (ref) for details.\footnote{We used the low consumption outcome (i) in the first semi-synthetic exercise.}
In this exercise, we now use the full RSV: the $50$ luminosity measures, and the full $4000$-dimensional embedding of satellite images.
In a semi-synthetic way, we classify villages into the samples described in Assumption (ref)(ii) and Remark (ref). We classify the real observational villages as $\tilde{S}=o$, for which we observe $(Y,R)$. Among the real experimental villages, we randomly classify some as $\tilde{S}=e$, for which we observe $(D,R)$, and some as $\tilde{S}=\{e,o\}$, for which we observe $(D,Y,R)$. In other words, for real experimental villages, we randomly delete half of their outcomes.
This exercise satisfies the key assumptions of our framework. The real treatment variable was randomized in the real experiment (Assumption (ref)). Stability of the RSV conditional distribution is plausible, as demonstrated in Figure (ref) (Assumption (ref)). Finally, because the real data have incomplete cases, we must argue that the treatment only affects the RSV via the outcome (Assumption (ref)(ii)). As supporting evidence, Figures (ref) and (ref) show that the distribution of $R|Y=y,D=1$ visually matches the distribution of $R|Y=y,D=0$.
Another requirement in our framework is that the RSV is relevant: $\E\{H(R) \Delta^o\} \neq 0$ in Corollary (ref). Intuitively, the representation of the RSV $H(R)$ used for inference should be correlated with outcome variation $\Delta^o$. This requirement is testable; for our learned representation in Algorithm (ref), we can test whether the empirical analogue $\E_n\{\widehat{H}(R) \widehat{\Delta}^o\}$ is significantly nonzero. Figure (ref) confirms that our RSV is relevant.
We further interpret and compare our learned representation of the satellite image in Appendix (ref). Our optimal representation is based on three predictions: the outcome, treatment, and sample indicator given the RSV. By contrast, common practice only uses the prediction of the outcome given the RSV. Interestingly, our optimal representation allows for extrapolation, i.e. negative weights on villages, whereas common practice does not.
Figure (ref) (“synthetic samples”) shows that our method recovers the treatment effect estimated by the unbiased benchmark in this exercise. For each poverty outcome, our treatment effects have the same signs and similar magnitudes as the benchmark, i.e. as if the economist could observe all treatments and outcomes in the experiment. These effects are consistent in sign with the findings in muralidharan2023general: early adoption of Smartcards reduces poverty. Compared to the benchmark, our confidence intervals are similar, and sometimes shorter; efficiently using the additional information in RSVs can boost statistical power.\footnote{Assumption (ref)(ii) is an additional restriction that may improve asymptotic precision. Relative magnitudes of standard errors may also reflect finite sample estimation error.}
These findings have a practical implication: our method may allow economists to significantly reduce survey costs. In this exercise, we mimic what would happen if an economist paid surveyors to collect outcomes for half of the experimental villages, and relied on free satellite images for the other half. We can calculate the savings, compared to paying surveyors to collect outcomes for all of the experimental villages. If surveying each individual in a village conservatively costs $\$0.50$, our results suggest savings of $\$3.6$ million.\footnote{Table (ref) summarizes the number of experimental villages and the populations in experimental villages, which we use to calculate savings. This calculation is highly conservative; viviano2022policy find that phone surveys in Pakistan cost $\$7$ per individual, rather than $\$0.50.$}
The third exercise is identical to the second exercise, except that we classify villages in a more realistic way. As before, we classify the real observational villages as $\tilde{S}=o$, for which we observe $(Y,R)$. Among the real experimental villages, we now classify the treated ones as $\tilde{S}=e$, for which we observe $(D,R)$, and the untreated ones as $\tilde{S}=\{e,o\}$, for which we observe $(D,Y,R)$. In other words, for real experimental villages, we systematically delete the outcomes of all treated villages. Methods based on missingness-at-random provide no guidance in this setting.
As before, this exercise appears to satisfy our identifying assumptions. The real treatment variable was randomized in the real experiment (Assumption (ref)). The conditional distribution of the RSV appears stable in Figure (ref) (Assumption (ref)). There do not appear to be direct effects in Figures (ref) and (ref) (Assumption (ref)(ii)). The RSV is relevant in Figure (ref) (Corollary (ref)).
In this more challenging exercise, we find qualitatively similar results. Our method recovers comparable estimates and confidence intervals to the benchmark in Figure (ref) (“real samples”). Our learned representations allow efficient extrapolation in Figure (ref).
We conclude this discussion with two final remarks on the interpretation of our semi-synthetic exercises.
Common empirical practice can be highly biased when the remotely sensed variable is post-outcome, i.e., when the RSV is caused by the outcome of interest. For example, changes in satellite images are caused by fires or variation in local income, but not vice versa. Theoretically, we demonstrate that this bias can have arbitrary sign. Empirically, we find that common practice may attenuate the treatment effect in a real environmental application, and it may have positive or negative bias in a semi-synthetic development application.
However, the core intuition underlying empirical work is powerful: the conditional distribution of the RSV given the outcome and treatment should be stable across samples. By formalizing this intuition, we develop a framework that nonparametrically identifies treatment effects and efficiently uses RSVs for program evaluation.
We conclude with some concrete recommendations for researchers conducting program evaluation with remotely sensed outcomes. We organize our recommendations into (i) implementation steps and (ii) diagnostic tests.
To apply our framework, researchers should follow three steps when designing studies with remotely sensed outcomes.
Step 1: Construct a credible observational sample. Build an auxiliary, observational sample that contains RSVs with linked outcomes. The key identifying assumption is stability: the conditional distribution of the RSV given the covariates, treatment, and outcome should be stable across the experimental and observational samples. Consequently, the researcher's most important design decision is how to construct an observational sample in a way that makes the stability assumption defensible.
Stability is most credible if the researcher randomly selects units for the experimental sample and observational sample prior to conducting the experiment. When random selection is infeasible, researchers should construct an observational sample that is similar to the experimental sample along a few dimensions. (i) Select units that are geographically close to the experimental sample. (ii) Select units with similar baseline covariates, e.g. rural-urban compositions, to the experimental sample. (iii) Select units that align with the timing of the experimental sample. That is, in the observational sample, collect outcomes measured in the same period as the experiment's endline, and collect RSVs measured at the same time that RSVs will be collected in the experiment. If exact alignment is not possible, then match the time difference so that the temporal relationship might be preserved.
Step 2: Estimate three predictions from RSVs. Efficient use of RSVs for downstream causal inference requires three predictions: the outcome, treatment, and sample indicator given the RSV. Common practice uses only the first prediction, in a way that sacrifices not only accuracy but also precision. In other words, by using all three predictions, our method improves not only bias but also variance. Researchers can construct these predictions using complex deep learning algorithms with unknown statistical properties.
Step 3: Apply robust inference algorithm. Algorithm (ref) aggregates the three predictions into an efficient scalar representation of the high-dimensional RSV (where efficiency is within the class of models satisfying the conditional moments we derive). This aggregation reduces the dimension while preserving information relevant for causal inference. Critically, Algorithm (ref) is robust to misspecification; it provides valid, $n^{-1/2}$ inference even if all three predictions are misspecified. It only requires that the learned representation converges at any rate to any limit that predicts outcome variation. This robustness property justifies the use of deep neural networks, which are widely implemented by empirical researchers in this literature, while providing the safeguard of a principled inferential framework.
In addition to the point estimate and confidence interval of Algorithm (ref), researchers should report the following diagnostics to defend the credibility of their RSV.
Diagnostic 1: RSV relevance test. Just as weak instruments threaten inference in instrumental variable models, weak RSVs---those that poorly predict the outcome---threaten inference in our framework. Researchers should assess whether the correlation between their estimated representation and the outcome is significantly different than zero. Figure (ref) illustrates this diagnostic, showing that satellite images are strongly related to village-level poverty. When relevance is weak, researchers should either improve RSV measurements (e.g., by using higher resolution images or additional data sources) or acknowledge that inference is inconclusive.
Diagnostic 2: Joint test of identifying assumptions. Our identifying assumptions are jointly testable. If two different RSV representations yield significantly different treatment effect estimates, then reject the null hypothesis that our identifying assumptions jointly hold: stability (Assumption (ref)) and no direct effects (Assumption (ref)(ii)). This is essentially a specification test. Researchers should verify that their results are robust to alternative ways of representing the RSV, such as using different machine learning architectures or different subsets of RSV features.
Diagnostic 3: RSV density plots if possible. When outcomes are observed for some experimental units, researchers can directly visualize whether stability holds by comparing conditional densities. Specifically, plot the RSV density in the experimental sample against the RSV density in the observational sample, for each combination of treatment and outcome values. If these densities align, then stability is plausible, as in Figure (ref). Similarly, researchers can assess the no-direct-effects assumption (Assumption (ref)(ii)) by comparing the RSV density of treated units against the RSV density of untreated units, for each outcome value. If treatment only affects the RSV via the outcome, these densities should coincide, as in Figures (ref) and (ref).
By following these steps, researchers can conduct rigorous program evaluation while substantially reducing costs. In our semi-synthetic exercise, we recover treatment effects and confidence intervals comparable to those obtained with direct outcome measurements, despite using satellite images for half the sample.
Our framework generalizes beyond simple randomized experiments. Appendix (ref) extends our identification results to quasi-experiments (e.g., instrumental variables and difference-in-differences). The key insight remains the same: stability in how the RSV relates to other variables across samples allows us to transport this relationship to the sample where the outcome is missing.
Our framework poses new questions. Future research may study a new experimental design problem: how to simultaneously design treatment assignment, outcome collection, and sensor deployment to maximize statistical power subject to budget constraints. Our main identification result provides the foundation to set up this problem. Our main result may inform not only the use of satellite images, but also the design of cheap, noisy surveys. Finally, future work may extend our inference guarantees to accommodate high-dimensional outcomes.
As remote sensing expands the frontier of data availability in economics, our framework offers practical principles for its use in program evaluation. This paper corrects a widespread but biased practice and provides concrete recommendations on how to incorporate satellite images and other remotely sensed data into causal inference. These new data sources and new methods broaden access to rigorous impact evaluation in cost-constrained environments.
\spacingset{1} \DeclareRobustCommand{\VAN}[2]{#2} \DeclareRobustCommand{\VAN}[2]{#1}
\spacingset{1.5}