EconBase
← Back to paper

Program Evaluation with Remotely Sensed Outcomes

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

97,972 characters · 21 sections · 61 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Program Evaluation with Remotely Sensed Outcomes

\def\spacingset#1{ {#1}} \spacingset{1}

abstractEconomists often estimate treatment effects in experiments using remotely sensed variables (RSVs), e.g., satellite images or mobile phone activity, in place of directly measured economic outcomes. A common practice is to use an observational sample to train a predictor of the economic outcome from the RSV, and then use these predictions as the outcomes in the experiment. We show that this method is biased whenever the RSV is a post-outcome variable, meaning that variation in the economic outcome causes variation in the RSV. For example, changes in poverty or environmental quality cause changes in satellite images, but not vice versa. As our main result, we nonparametrically identify the treatment effect by formalizing the intuition underlying common practice: the conditional distribution of the RSV given the outcome and treatment is stable across samples. Our identifying formula reveals that efficient inference requires predictions of three quantities from the RSV---the outcome, treatment, and sample indicator---whereas common practice only predicts the outcome. Valid inference does not require any rate conditions on RSV predictions, justifying the use of complex deep learning algorithms with unknown statistical properties. We reanalyze the effect of an anti-poverty program in India using satellite images. {\bf Keywords:} Causal inference, data fusion, experiments, satellite images, machine learning.

\spacingset{1.5}

Introduction

While traditional program evaluations rely on surveys to measure impact, important economic outcomes such as living standards and environmental quality may be costly or infeasible to collect at scale. As a consequence, researchers increasingly estimate treatment effects on economic outcomes using remotely sensed variables (RSVs). Examples include night lights as a measure of local economic activity chen2011using, henderson2012measuring, asher2021development, roofing material as a measure of housing quality marx2019there, michaels2021planning, huang2021using, mobile phone transactions as a measure of local wealth or consumption blumenstock2015predicting, aiken2025estimating, and satellite images as a measure of pollution currie2023caused, deforestation jayachandran2017cash, assunccao2023optimal, fires jack2022money, balboni2024origins, flooding chen2017validating, patel2024floods, and local poverty jean2016combining.\footnote{Figure (ref) illustrates the rapid rise of papers published in economics journals and general interest science journals using RSVs. Of those published in economics journals, we find that 90% use high-dimensional satellite images as RSVs, and 40% use the RSVs as the main outcome in their empirical analyses. See also burke2021using and jack2023remotesensing for recent reviews on the empirical uses of RSVs.} Our research question is how researchers should rigorously estimate treatment effects from remotely sensed outcomes.

A recurring empirical practice appears in about 50% of papers in general interest economics journals from 2015-2024 that use remotely sensed outcomes.\footnote{We survey AEA journals, Econometrica, Journal of Political Economy, Quarterly Journal of Economics, and Review of Economic Studies. The other 50% of papers use a similar logic, but without an explicit formula for data combination.} This common practice predicts the economic outcome from the RSV and then uses the predicted outcome in lieu of a true outcome measurement in an experimental sample. Researchers often form such predictions using an auxiliary, observational sample collected in some other context, which contains the RSV and linked outcome measurements. The predictor is typically complex, e.g., a deep learning algorithm with unknown statistical properties.

We show that this intuitive method can produce arbitrarily biased treatment effect estimates---including flipped signs---when the RSV is a post-outcome variable. Using the predicted outcome in lieu of the true outcome implicitly uses the RSV as a surrogate that mediates between the treatment and the outcome prentice1989surrogate, athey2024surrogate. However, in many empirical applications using RSVs, the opposite is more plausible: the treatment affects the outcome, and both may affect the RSV. The bias is fundamentally due to this reversal; it is present even without machine learning.

As an example, consider a binary outcome indicating whether a plot of land has been burned balboni2024origins, jack2022money and an RSV summarizing the color saturation in a satellite image. Fires cause changes in satellite images, but not vice versa; the color saturation of satellite images is a post-outcome variable. Common practice predicts fires in the experimental sample using a machine learning algorithm trained on labeled satellite images from an observational sample, and then computes the difference in predicted outcomes between treated and control units in the experimental sample. Because the RSV is post-outcome, this method's estimand combines two quantities: the desired effect of the treatment on the outcome, and the correlation between the RSV and the outcome. In the extreme case where the RSV fails to predict the outcome---when an ideal method should report infinite standard errors---this popular method will instead report a precise estimate of zero, regardless of the true treatment effect. In general, this method may flip the sign of the true treatment effect.

Our main contribution is a novel formula to nonparametrically identify treatment effects using RSVs by combining (i) an experimental sample where the outcome is missing, and (ii) an observational sample in which the treatment is non-randomized and possibly missing. Our key assumption formalizes the logic underlying the examples above: the conditional distribution of the RSV given the outcome and the treatment is stable across both samples. Consequently, the relationship between the RSV and the outcome can be learned from the observational sample and transported to the experimental sample. If the treatment is missing in the observational sample, then we need an additional assumption to restore identification: the treatment only affects the RSV through the outcome.

Our main identifying assumptions are jointly testable and lend themselves to simple diagnostics, allowing researchers to assess their plausibility in applications. We propose a diagnostic to evaluate whether an RSV is sufficiently relevant for an economic outcome to justify its use in program evaluation. Our framework also naturally extends to quasi-experiments, including instrumental variables and difference-in-differences designs, which we present in Appendix (ref).

Our secondary contribution is to characterize the representation of the RSV that achieves valid, precise, and robust $n^{-1/2}$ inference on the treatment effect. Because modern remote sensing typically involves unstructured data and complex machine learning, we derive valid inference without rate conditions and without complexity restrictions on RSV-based predictions. Valid inference only requires that (i) a learned RSV representation has some limit, and (ii) the limit predicts the outcome of interest, for which we provide a diagnostic. This enables researchers to use complex deep learning algorithms and still conduct valid inference on treatment effects.

Specifically, we build on our main identification result to establish the connection between modern remote sensing and classical conditional moments chamberlain1987asymptotic,newey1993efficient. This allows us to derive an expression for a simple representation of the RSV that maximizes precision from the conditional moments we derived. We find that three predictions are necessary for efficient downstream causal inference based on RSVs: predictions of the outcome, the treatment, and the sample indicator given the RSV. By contrast, common practice only predicts the outcome given the RSV. Moreover, $n^{-1/2}$ inference remains valid even if all three predictions are misspecified, which is a strong form of robustness. We provide the remoteoutcome R package to implement our method.

Finally, we conduct a semi-synthetic exercise, calibrated to an existing field experiment in India, which we merge with existing satellite images. Following muralidharan2016building, muralidharan2023general, we study the effect of Smartcards, a biometrically authenticated payments infrastructure, on village-level poverty measures. We use the geographic coordinates of each village to extract nighttime luminosity and high-dimensional, pre-trained embeddings of satellite images. Despite using outcomes for only half of the units, we recover the treatment effect of interest with the same precision as an unbiased regression method that has access to outcomes for all units. By contrast, common practice can have positive or negative bias for the treatment effect. Our method enables large savings in survey costs and therefore opens up possibilities for program evaluation in cost-constrained environments.

Related Work

Our framework differs from existing approaches to auxiliary variables and data combination along two key dimensions: (i) the causal direction between the auxiliary variable and the outcome, and (ii) the data requirements.

In the surrogacy framework prentice1989surrogate, athey2024surrogate, kallus2024role, the auxiliary variable (surrogate) is a mediator between the treatment and the outcome, whereas the RSV is a post-outcome variable in our framework. We show that misusing a post-outcome RSV as a surrogate leads to arbitrary bias in treatment effects. The negative control literature extends the surrogacy framework to address unobserved confounding ghassami2022combining,imbens2024long, yet such extensions face the same limitation. See Remark (ref) for details.

Compared to the vast literature on data combination cross2002regressions, RidderMoffitt(07), bareinboim2016causal, d2024partially and nonclassical measurement error ChenHongNekipelov(11),schennach2020mismeasured, we place what appears to be a different key assumption. Several influential works handle measurement error in moment condition models by using auxiliary data and assuming that the conditional distribution of the variable of interest, given the imperfect measurement, is stable across samples chen2005measurement,chen2008semiparametric,GrahamPintoEgel2016, akin to the surrogacy framework. Our key identifying assumption is the opposite: the conditional distribution of the imperfect measurement, given the variable of interest, is stable across samples. This is natural when the imperfect measurement is caused by the outcome (as with satellite images of poverty or environmental outcomes) rather than merely correlated with it. Other works focus on mismeasured or missing covariates in a sample where the outcome is observed fan2014identifying,battaglia2024inference,d2024partially,d2024linear. By contrast, in our setting, the outcome is missing yet a post-outcome variable is present. Both of these distinctions lead to a novel identifying formula.

The main difference between our RSV framework and the prediction-powered inference (PPI) framework angelopoulos2023prediction,lu2025regressioncoefficientestimationremote,kluger2025prediction is also along these lines: the PPI framework uses machine learning predictions as surrogates ji2025predictions. Another difference concerns data availability. In our terminology, the PPI approach would require the researcher to observe the treatment, outcome, and RSV for a random subsample of experimental units. The data requirements in other works are similar to those in the PPI literature fong2021machine,allon2023machine,gordon2023remote,egami2024using, carlson2025unifying. By contrast, we allow the researcher to observe no outcomes for any experimental units, enabling inference in settings where PPI-style techniques cannot be applied.

Our results do not require a correctly specified generative model of how treatments and outcomes affect RSVs, which may be prone to misspecification when the RSV is a satellite image or some other type of unstructured data. Several previous works propose methods based on generative modeling that do require correct specification gentzkow2019measuring,alix2023remotely,proctor2023parameter. Similarly, methods for causal inference on outcomes that are latent concepts require a correctly specified generative model egami2022make,knox2022testing,stoetzer2024causal.

Model and Assumptions

Goal: Identification using Remotely Sensed Outcomes

The researcher observes units in two samples, indicated by the variable $S\in \{e, o\}$: an experimental sample ($S=e$) and an observational sample ($S=o$).

Within the experimental sample $(S=e)$, we observe pre-treatment covariates $X \in \mathcal{X}$ and a binary treatment $D \in \{0, 1\}$. However, the outcome $Y\in \mathcal{Y}$ is missing.\footnote{For ease of exposition, we focus on the case where the outcome $Y$ is completely missing in the experimental sample. Remark (ref) gives the extension where $Y$ is only partially missing in the experimental sample.} In its place, we have access to a remotely sensed outcome variable $R\in\mathcal{R}$. We typically think of $R$ as high-dimensional (e.g., unstructured data such as satellite images), but it could be low-dimensional (e.g., the output of some pre-trained machine learning algorithm). The researcher would like to use the remotely sensed variable (RSV) as an imperfect measurement of the outcome in the experimental sample, without placing parametric assumptions on their relationship.

The causal parameter of interest is the effect of the treatment $D$ on the outcome $Y$ in the experimental sample. Though the outcome $Y$ is unobserved in the experimental sample, we may still define its potential outcomes $Y(d)$ and our parameter of interest.\footnote{This definition of potential outcomes has no spillovers. We discuss spillovers in Appendix (ref).}

definition[Causal parameter] The average treatment effect (ATE) in the experimental sample is $\theta := \mu(1) - \mu(0)$, where $\mu(d): = \E\{Y(d) \mid S = e\}$.

Without further assumptions, point identification for this causal parameter is impossible horowitz1995identification. Even if $R$ can predict $Y$ with great accuracy, if the prediction is at all imperfect, then an assumption is necessary. Therefore, we place additional structure on the problem, inspired by recent empirical work in environmental and development economics.

A popular practice in environmental and development economics is to use an auxiliary dataset of outcomes and RSVs, e.g., of labeled satellite images. We refer to the auxiliary dataset as the observational sample $(S=o)$. For these units, we observe baseline covariates $X$, the outcome $Y$, and the remotely sensed variable $R$. We may or may not observe the treatment $D$. If we do, we denote it by $D\in\{0,1\}$ and refer to this scenario as having “complete” cases. If we do not, or if treatment is deterministic in the observational study, we set $D=0$ for all units in the observational study and refer to this latter scenario as having “incomplete” cases.\footnote{Incomplete cases in the observational sample may refer to three scenarios. First, the treatment status may be missing. Second, the treatment status may be present, and all observational units are untreated, hence $D=0$. Third, the treatment status may be present, and all observational units are treated. Relabeling treatment values gives $1-D=0$.} As we discuss below, incomplete cases will require stronger assumptions.

Table (ref) summarizes the setting. Each unit is characterized by the random vector $\left( S, X, D, Y(0), Y(1), R \right)$, which we assume to be independent and identically distributed.\footnote{Independence is not used to derive our main identification argument.} For units in the experimental sample ($S = e$), we observe $(X, D, R)$; for units in the observational sample ($S = o$), we observe $(X, D, Y,R)$ in complete cases or $(X,Y,R)$ in incomplete cases.

table[table omitted — 770 chars of source]
example[Environmental impacts] Consider a randomized experiment that offers cash payments to households in order to incentivize environmental conservation (i.e., “payments for ecosystem services” or PES). Access to PES contracts is often randomized at the village level. We would like to measure whether access to PES contracts $D$ reduces harmful environmental behaviors $Y$, such as deforestation jayachandran2017cash or crop burning jack2022money. In a separate observational sample, we link satellite images $R$ to direct measurements $Y$ of deforestation or crop burning hansen2013high, walker2022detecting. While it is expensive to hire surveyors to record measurements of tree cover or crop management practices in rural areas, it is inexpensive to collect satellite images. We investigate how to combine these data sources and thereby identify the effect of the PES contracts in the experimental sample. $\blacktriangle$
figure[figure omitted — 1,003 chars of source]
example[Household poverty] Consider a randomized experiment evaluating an anti-poverty program, such as an unconditional cash transfer egger2022general or biometrically authenticated payment muralidharan2023general. Treatment is often randomized at the village level. We would like to study the effect of the anti-poverty program $D$ on village-level poverty $Y$. In a separate observational sample, we link satellite images $R$ to census statistics on village-level poverty $Y$. It is well documented that poverty can be predicted from satellite images, with some error jean2016combining, rolf2021generalizable. For example, huang2021using use deep learning methods to predict household wealth in Kenya from roof quality. While it is expensive to collect poverty measures through in-person surveys in the experimental sample, it is inexpensive to collect satellite images. We identify the effect of the anti-poverty program in the experimental sample. Figure (ref) illustrates an example of incomplete cases, using data from an evaluation of an anti-poverty program in India, where we apply our method in Section (ref). $\blacktriangle$

Main Assumption: Stability

We formalize this causal setting via three assumptions. Our identifying assumptions allow $(X,Y,R)$ to be discrete or continuous. For readability, we slightly abuse notation: for a random variable $W$, we use the symbol $f_W(\cdot \mid ...)$ to refer to its (conditional) probability mass function if $W$ is discrete, or its (conditional) density if $W$ is continuous.\footnote{Formally, $f_W(\cdot \mid ...)$ is our symbol for the Radon-Nikodym derivative.}

assumption[Experimental unconfoundedness] Suppose the following: \begin{enumerate}[label=\roman*.] • SUTVA: $Y = D Y(1) + (1 - D) Y(0)$ almost surely. • Randomization: $D \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} \left\{ Y(0), Y(1) \right\} \mid X, S = e$. • Overlap: $\Pr(D = 1 \mid X, S = e) $ is bounded away from zero and one almost surely. \end{enumerate}

In many empirical applications involving RSVs, such as Examples (ref) and (ref), Assumption (ref) is satisfied by design: experimental units are chosen as aggregates without spillovers, e.g., villages, and the treatment is randomly assigned for these experimental units. We defer to Appendix (ref) the study of quasi-experimental settings, such as instrumental variables and difference-in-differences.

Under Assumption (ref), if we were to observe the outcome in the experimental sample, the ATE could be identified using standard arguments. However, the outcome is not observed in the experiment; instead, we have an RSV.

We resolve this measurement issue by leveraging the observational sample. Intuitively, the idea is to learn the relationship between the RSV and the outcome of interest in the observational sample and to “transport” it to the experimental sample.

assumption[Stability of the remotely sensed variable] Suppose the following: \begin{enumerate}[label=\roman*.] • Stability: $S \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} R \mid X, D, Y$. • Common support: for some outcome support $\mathcal{Y}$, $\Pr(Y \in \mathcal{Y} \mid S = e, X) = 1$ almost surely, and $f_Y(y \mid S = o, X)$ is bounded away from zero almost surely for all $y \in \mathcal{Y}$. • Coverage: $f_R( r \mid S, X, D)$ is bounded away from zero almost surely for all $r \in \mathcal{R}$. • Two samples: $\Pr(S = e|X)$ is bounded away from zero and one almost surely. \end{enumerate}

Assumption (ref)(i) is the main assumption of our framework: the conditional distribution of the remotely sensed variable $R$, given $(X, D, Y)$, is stable across the experimental and observational samples. This allows us to “transport” the measurement error distribution from the observational sample to the experimental sample. Importantly, this condition does not require stability of the underlying treatment effects, which may differ across samples. See Remark (ref) for comparisons between Assumption (ref)(i) in the RSV model and assumptions in alternative models.

Returning to our two leading examples, Assumption (ref)(i) requires that the conditional distribution of tree cover pixels $R$, given environmental outcomes $Y$ and interventions $D$ (as well as other pre-treatment covariates), is stable across the experimental and observational samples. Analogously, it requires that the conditional distribution of the satellite image $R$, given village-level poverty $Y$ and the anti-poverty program $D$ (as well as other pre-treatment covariates), is stable across the experimental and observational samples.

figure[figure omitted — 2,040 chars of source]

Our main assumption is empirically plausible, as illustrated by Figure (ref). We use data from an anti-poverty program in India, where outcomes are observed. If Assumption (ref) holds, then the conditional densities of the RSV given the outcome and treatment should be the same across the experimental and observational samples. They appear to coincide in this empirical setting.\footnote{When $X=\varnothing$, Assumption (ref) imposes four equalities: $f_R(R \mid S=e,D=d,Y=y)=f_R(R \mid S=o,D=d,Y=y)$ for $d\in\{0,1\}$ and $y\in\{0,1\}$. Each equality can be evaluated with a diagnostic plot if outcome data are available. For example, using the experimental and observational units satisfying $D=0$ and $Y=0$ in Figure (ref), we can visualize whether the density of $R\mid S=e, D=0, Y=0$ aligns with the density of $R \mid S=o, D=0, Y=0$ in Figure (ref). Since $R$ is high-dimensional, we simplify the visualization by comparing the densities of its first principal component.}

The remaining aspects of Assumption (ref) are weak regularity conditions. Assumption (ref)(ii) requires that the outcome in the observational sample has a common (or larger) support than the outcome in the experiment. Assumption (ref)(iii) ensures that the RSV distribution does not degenerate for any stratum. Assumption (ref)(iv) requires that we observe some data from both the experimental and observational samples.

Assumptions (ref) and (ref) imply identification when the observational sample has complete cases. When the observational sample has incomplete cases, we require a further assumption. In other words, if the treatment is missing or deterministic in the observational sample, then a further restriction is necessary for point identification.

assumption[Observational completeness] Suppose that either condition holds: \begin{enumerate}[label=\roman*.] • Complete cases: $\Pr(D = 1 \mid S = o, X)$ is bounded away from zero and one almost surely; • No direct effect: $D \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} R \mid X, Y$. \end{enumerate}

Assumption (ref) imposes only one of two conditions.

Assumption (ref)(i) implies that we have access to complete cases, i.e., some observations of $(X,D,Y,R)$ where $D$ has variation. Within the observational sample, the treatment is observed and varies, although it does not need to be randomized and may suffer from unobserved confounding. Whenever we have complete cases, under Assumption (ref)(i), no further causal assumptions are needed. In particular, the treatment $D$ may have a direct effect on the remotely sensed variable $R$. In Example (ref), this would allow the environmental program to affect satellite images both indirectly, i.e., via crop burning, and directly, e.g., via visible investments in farm equipment.

If Assumption (ref)(i) is violated, then we have no complete cases, i.e., no observations of $(X,D,Y,R)$ where $D$ is variable. Without joint observations of the outcome and treatment, a further restriction is needed. Assumption (ref)(ii) fills this gap, requiring that the treatment $D$ affects the remotely sensed variable $R$ only via its effect on the outcome $Y$. In Example (ref), we may be comfortable assuming that the PES contract has no direct effect on the specific infrared band used to measure charred soil in satellite images. Assumption (ref)(ii) may also become more plausible when $Y$ is a vector of outcomes. Several outcomes may approximate all mechanisms through which the treatment affects the RSV. For exposition, we focus on scalar outcomes in the main text and we generalize to vector outcomes in Appendix (ref).

Together, Assumption (ref)(i) and Assumption (ref)(ii) imply that $\left( S, D \right) \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} R \mid X, Y$. In the next section, we will show that Assumptions (ref)(i) and (ref)(ii) are jointly testable, even when no outcome is observed from the experimental sample (Remark (ref)).

figure[figure omitted — 1,016 chars of source]

Figure (ref) illustrates our identifying assumptions as a causal graph. The treatment affects the outcome, which in turn affects the RSV. By Assumption (ref), the sample indicator does not change the conditional distribution of the RSV given the treatment and outcome. Depending on which version of Assumption (ref) is imposed, the treatment may have a direct effect on the RSV, as illustrated by the dotted line. Table (ref) summarizes the implications of our main assumptions.

Main Result: Identification

In this causal setting, a commonly used procedure may lead to causal estimates with arbitrary bias. Motivated by this negative result, we prove a positive one: we nonparametrically identify the causal parameter by combining the experimental and observational samples differently.

To streamline notation, we initially focus on the setting where Assumption (ref)(ii) holds, then return to the setting where Assumption (ref)(i) holds at the end of this section. For exposition, we will also assume that the outcome is binary, i.e., $\mathcal{Y} = \{0,1\}$. Finally, we assume that the support of $R$ is at least as large as the support of $Y$ (it may be binary, discrete, or continuous). Appendices (ref) and (ref) extend our results to discrete and continuous outcomes, respectively.

Current Practice may have Positive or Negative Bias

In empirical research, it is common to use RSVs in two steps: (i) researchers train a predictor of the outcome $Y$ from the remotely sensed variable $R$ in the observational sample, and (ii) the predictor is applied to the experimental sample, where its predictions are used as surrogate outcomes to estimate treatment effects. While intuitive, this empirical strategy can lead to arbitrary bias for the ATE in the experimental sample.

Suppose there are no pre-treatment covariates for simplicity. The widely used two-step estimation procedure implicitly targets the estimand $\widetilde{\theta} = \widetilde{\mu}(1) - \widetilde{\mu}(0)$, where $\widetilde{\mu}(d) := \E\left\{ \E(Y \mid R, S = o) \mid D = d, S = e \right\}$ for $d \in \{0, 1\}$. Within this expression, the first step estimates the conditional expectation function $\E(Y \mid R=r, S = o)$. The second step evaluates and averages this function over the treated and untreated subgroups in the experimental sample.

As a starting point, if the RSV fails to predict the outcome, i.e., if $\E(Y|R,S=o)=\E(Y|S=o)$, then the implicit target $\widetilde{\theta}$ is zero regardless of the true treatment effect. Common practice would return a precise estimate of zero, even though the RSV provides no information about treatment effects in this case.

More importantly, even if the RSV does predict the outcome, the implicit target $\widetilde{\theta}$ can incur bias with arbitrary sign for the ATE in the experimental sample.

proposition[Bias of common practice] Suppose Assumptions (ref), (ref), and (ref)(ii) hold with $X=\varnothing$. Suppose $\Pr(D = 1 \mid S = o) = 0$ and $S \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} (Y, R) \mid D$. Then the following hold. \begin{enumerate}[label=\roman*.] • The bias is $ \widetilde{\theta} - \theta = \mu(1)\int \{w(r) - 1\}f_R(r|Y = 1, S = e)\mathrm{d}r$, where $w(r) = \frac{\Pr\{Y(0) = 1 \mid S = e\} f_R(r \mid D = 1)}{\Pr\{Y(1) = 1 \mid S = e\} f_R( r \mid D = 0)}. $ • There exists a data-generating process satisfying the above restrictions with $\widetilde{\theta} - \theta > 0$, and a different data-generating process satisfying the above restrictions with $\widetilde{\theta} - \theta < 0$. \end{enumerate}

Proposition (ref) derives the bias of current empirical practice, which uses the RSV as a surrogate outcome. Under Assumptions (ref) and (ref)(ii), the conditional distribution of the RSV $f_R\left(R \mid Y, S = o\right)$ is stable across the experimental and observational samples. Existing practice attempts to transport the predictions $f_Y(Y \mid R, S = o)$ into the experimental sample. By Bayes' rule, this induces a bias due to differences in the marginal distributions of the outcomes and RSVs across the samples. Importantly, the bias of common practice arises even if units are randomly allocated into the experimental and observational samples.\footnote{See Appendix (ref) for a formal verification of this claim.}

remark[Bias summary] To gain insight into the source of bias, we consider three nested cases. (i) In the extreme “irrelevance” case, the RSV fails to predict the outcome. We would like an estimator whose variance is infinity. However, $\E(Y|R,S=o)$ is a constant function of $R$, so $\widetilde{\theta}=0$; the common practice always reports zero. (ii) In the simple linear case, something similar occurs. Suppose the data-generating process satisfies $\E(Y|R,S=o)=\tilde{\beta}_0+\tilde{\beta}R$ and $\E(R|Y,D,S=e)=\beta_0+\beta Y$. Combining these expressions, $\widetilde{\theta}=\tilde{\beta}\beta \theta_0$; the common practice suffers from attenuation bias (see Appendix (ref) for further details). (iii) In the general case, where the relationships among variables are nonlinear, Proposition (ref) shows that the bias of common practice can be positive or negative.
remark[Comparison to the surrogacy framework] Proposition (ref) provides a direct comparison to the surrogacy framework. Within the surrogacy framework, the implicit target of empirical practice $\widetilde{\theta}$ recovers the causal parameter $\theta$ if $(D, S) \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} Y \mid R, X$, i.e., if surrogacy and surrogate compatibility are satisfied prentice1989surrogate, athey2024surrogate. These assumptions contrast our Assumptions (ref) and (ref). The surrogacy assumptions state that the surrogate fully mediates the effect of the treatment on the outcome, and so the surrogate is pre-outcome. Assumptions (ref) and (ref) imply the opposite: the outcome mediates the effect of the treatment on the RSV, either partially (Assumption (ref)(i)) or fully (Assumption (ref)(ii)); the RSV is post-outcome. Figure (ref) illustrates the difference via a causal graph. Below, we argue that the RSV model is more plausible in environmental and development applications with satellite images. The surrogacy framework can be viewed as a version of the moment condition model with auxiliary data studied by e.g., chen2005measurement, chen2008semiparametric, GrahamPintoEgel2016. In our terminology, those works require that the conditional distribution of the outcome given the RSV and treatment is stable across samples. By contrast, we require that the conditional distribution of the RSV given the outcome and treatment is stable across samples. Figure (ref) suggests that our assumption is plausible in a real development application with satellite images. The surrogacy model, and related models which lead to the estimand $\widetilde{\theta}$, have been extended to include negative controls ghassami2022combining,imbens2024long. As generalizations of the surrogacy model, they suffer from the same drawback formalized in Proposition (ref). Similar to the RSV framework, the single negative control framework park2024single has one auxiliary variable. The frameworks have two key differences. Unlike the single negative control model, we allow the treatment $D$ to affect both the outcome $Y$ and the auxiliary variable $R$ when Assumption (ref)(i) holds. When no complete cases exist, i.e. the setting covered by Assumption (ref)(ii), the single negative control model provides no guidance of how to proceed.

Example: Current Practice may Underestimate Environmental Impacts

We revisit an experiment conducted by jack2022money that studied whether payments for ecosystem services (PES) contracts can incentivize farmers to reduce crop burning. We measure the average treatment effect $\theta$ of being offered a PES contract $D \in \{0, 1\}$ on the likelihood that a farmer burns their fields $Y \in \{0, 1\}$. Here, $Y=1$ means not burning, so a positive $\theta$ means environmental benefit.

It is costly to measure whether the crop residue on a particular field has been burned, requiring a surveyor to make frequent visits to rural fields. Therefore, it is natural to turn to a remotely sensed variable $R$ for the outcome $Y$. To construct such a remotely sensed variable, jack2022money link surveyor-collected measurements of crop burning to satellite-based spectral indices and then train a supervised learning algorithm to predict whether these fields have been burned. Here, $R \in \{0, 1\}$ is a classifier for whether a field has not been burned, which applies a threshold rule to a machine learning prediction of the probability that a field has not been burned.\footnote{jack2022money construct two binary RSVs for crop burning by applying two alternative threshold rules to the estimated probability a field has not been burned. We use their “max accuracy” RSV in the main text, and we report analogous results in Table (ref) using the authors' “balanced accuracy” RSV.}

Our causal assumptions are plausible in this setting. Because burning crops would alter satellite images, but altering satellite images would not burn crops, $R$ is a post-outcome variable rather than a pre-outcome variable. Since access to PES contracts was randomized at the village level, Assumption (ref) is satisfied by design. Since the authors conducted randomized spot checks, surveying the outcome $Y$ for randomly selected fields, Assumption (ref) is also satisfied by design if we define the observational sample $S = o$ as the fields in which these random spot checks occurred. Finally, since specific infrared bands are used to form $R$, it is plausible that PES contracts affect these specific infrared bands only via crop burning, so Assumption (ref)(ii) is reasonable. Figure (ref) illustrates the two samples used in our re-analysis.

table[table omitted — 889 chars of source]

We implement the common empirical practice: we use the RSV as a surrogate for crop burning. Column (1) of Table (ref) estimates $\widetilde{\theta}$. Offering any PES contract to farmers appears to reduce crop burning by 7.9%.\footnote{We modify the specification of jack2022money in two ways. (i) While they distinguish between two types of PES contracts, we define the treatment as whether any PES contract was offered. (ii) They analyze the effects of PES contracts by defining a farmer-level outcome, whereas we analyze effects at the field level.}

However, by Proposition (ref), $\widetilde{\theta}$ is typically biased for $\theta$. To quantify the magnitude of its bias, we estimate $\beta := \E(R \mid Y = 1) - \E(R \mid Y= 0)$ using information collected through random spot checks of the fields. Under Assumption (ref)(ii), an algebraic argument in Appendix (ref) yields $\widetilde{\theta}= \beta \theta$ in this setting.\footnote{Compared to Remark (ref)(ii), here we have $\E(Y|R,S=o)=R$, which implies $\tilde{\beta}=1$ and hence $\widetilde{\theta}=\beta \theta_0$.} Column (2) of Table (ref) reports the estimate of $\beta$. It suggests that $\widetilde{\theta}$ understates the treatment effect of being offered any PES contract by approximately 47%. This motivates us to ask: what should empirical researchers do instead?

Main Result: An Identification Formula

Our main result nonparametrically identifies the ATE in the experimental sample without a parametric model that restricts the distribution or dimensionality of the RSV. We derive a novel formula for combining the experimental and observational samples.

To begin, we express the RSV distribution in the experimental sample as an affine transformation of the potential outcomes we wish to identify. The slope and intercept depend on how the RSV relates to the outcome, and they are identified by the observational sample.

lemma[Identification as generative model] Suppose Assumptions (ref), (ref), and (ref)(ii) hold. Then, for any $d \in \{0, 1\},$ $x \in \mathcal{X}$, and $r \in \mathcal{R}$, $ \delta^{e}_{d}(r,x)=\left\{ \delta_{1}^o(r,x) - \delta_{0}^o(r,x) \right\}\mu(d,x)+ \delta^o_{0}(r,x), $ where $\mu(d,x) := \E\{Y(d) \mid S = e, X = x\}$ is the conditional average potential outcome in the experimental sample. Here, $\delta_{d}^{e}(r,x) := f_R( r \mid S = e, X = x, D = d)$ and $\delta_{y}^{o}(r,x) := f_R( r \mid S = o, X = x, Y = y)$ are RSV conditional densities in the experimental and observational samples, respectively.

By Lemma (ref), we recover the ATE in the experimental sample by combining (i) how the RSV varies with the treatment in the experimental sample with (ii) how the RSV varies with the outcome in the observational sample. This combination leverages our key assumption: stability of the RSV across samples.

While Lemma (ref) identifies the ATE in the experimental sample, it suggests a challenging estimation problem: it involves the conditional distribution of the high-dimensional RSV. One path forward is to develop a complex parametric model, e.g., a generative model of the satellite image distribution conditional upon poverty measurements or environmental outcomes. Even if such a generative model could be developed, it would be prone to misspecification.

We follow a different path that does not require a generative model for $R$. We use Bayes' rule to rewrite Lemma (ref) as a conditional moment equation. This transformation avoids estimation of the RSV's conditional distribution and recovers a classic econometric estimation problem.

Let $\theta(x) := \E\{Y(1) - Y(0) \mid S = e, X =x\}=\mu(1,x)-\mu(0,x)$ denote the conditional average treatment effect in the experimental sample.

theorem[Identification as conditional moment] Under the conditions of Lemma (ref), $$ \text{for any $x \in \mathcal{X}$,} \quad \E\{\Delta^{e}(x) - \Delta^o(x) \theta(x) \mid X = x, R\} = 0 \quad \text{almost surely, where} $$ $\Delta^{e}(x) := \frac{1\{ D = 1, S = e \}}{\Pr( D = 1, S = e \mid X = x )} - \frac{1\{ D = 0, S = e \}}{\Pr( D = 0, S = e \mid X = x)}$ and $\Delta^{o}(x) := \frac{1\{ Y = 1, S = o \}}{\Pr( Y = 1, S = o \mid X = x )} - \frac{1\{ Y = 0, S = o \}}{\Pr( Y = 0, S = o \mid X = x )}$.

Theorem (ref) identifies the treatment effect as the solution to a set of conditional moment equalities. The conditional average treatment effect $\theta(x)$ balances treatment variation from the experimental sample $\Delta^{e}(x)$ and outcome variation from the observational sample $\Delta^o(x)$, so that their projections onto the remotely sensed variable $R$ match. Consequently, we can leverage a celebrated literature on conditional moment equalities chamberlain1987asymptotic, newey1993efficient for estimation and inference.

For intuition on Theorem (ref), consider the case without covariates. Then $\Delta^e$ is simply a scalar computed from the experimental sample, where the treatment is observed. Similarly, $\Delta^o$ is simply a scalar computed from the observational sample, where the outcome is observed. Theorem (ref) shows that $\theta_0=\frac{\E(\Delta^e|R)}{\E(\Delta^o|R)}.$ This ratio appears to be a new formula for data combination, where the numerator and denominator are from different samples. The numerator reflects how both the treatment and the outcome affect the RSV, so we must divide it by the effect of the outcome on the RSV, in order to isolate the desired causal parameter.

For identification, Theorem (ref) implies that we may introduce representations of the RSV. Such representations can be arbitrary, as long as they predict outcome variation.

corollary[Identification as representation] Under Lemma (ref)'s conditions, $\theta(x) = \frac{\E\{ H(x, R) \Delta^{e}(x) \mid X = x\}}{\E\{ H(x, R) \Delta^o(x) \mid X = x\}}$ for any representation $H(x,R)$ with $\E\{H(x, R) \Delta^o(x) \mid X = x\} \neq 0$.
remark[Testable implication] Our main identifying assumptions are jointly testable. By Corollary (ref), any predictive representation $H(x,R)$ identifies $\theta(x)$. If different representations yield significantly different estimates, then we can reject our identifying assumptions. Future work may develop bounds under violations of our identifying assumptions.

Once again, for intuition on Corollary (ref), consider the scenario without covariates. In this case, $H(R)$ is simply a scalar representation of the RSV. Corollary (ref) shows that $\theta_0=\frac{\E\{\Delta^e H(R)\}}{\E\{\Delta^o H(R)\}}$. Any representation $H(R)$ is valid as long as the representation is predictive: $\E\{\Delta^o H(R)\}\neq 0$. A weak RSV test asks whether $\E\{\Delta^o H(R)\} \approx 0$, which can be tested using existing frameworks for weak instruments. In the same spirit, a joint test of our identifying assumptions asks whether two representations $H(R)$ and $H'(R)$ give similar estimates. Dissimilar estimates would be evidence against our assumptions.

By Corollary (ref), many representations of the RSV provide identification. However, naive choices may be inefficient, producing needlessly large standard errors. Section (ref) asks: what representation of the RSV should be chosen for precise, downstream causal inference? Our answer develops a connection between “representation learning” johannemann2019sufficient, vafa2024estimatingwagedisparitiesusing and classical results for conditional moment equalities.

So far, we have focused on the setting of Assumption (ref)(ii), which allows incomplete cases but disallows direct effects of the treatment on the RSV. Our main result also holds for the setting of Assumption (ref)(i), which requires complete cases yet allows direct effects of the treatment on the RSV. Recall that $\mu(d,x) := \E\{Y(d) \mid S = e, X = x\}$ is the conditional average potential outcome in the experimental sample.

theorem[Identification as conditional moment with direct effects] Suppose Assumptions (ref), (ref), and (ref)(i) hold. Then, for any $d \in \{0, 1\}$ and any $x \in \mathcal{X}$, $$ \E\{\widetilde{\Delta}^e(d,x) - \widetilde{\Delta}^o(d,x) \mu(d,x) \mid R, X = x\} = 0 \quad \text{almost surely, where} $$ $\widetilde{\Delta}^e(d,x) := \frac{1\{ D = d, S = e \}}{ \Pr(D = d, S = e \mid X = x) } - \frac{1\{ Y = 0, D = d, S = o \}}{ \Pr( Y = 0, D = d, S = o \mid X = x ) }$, and $\widetilde{\Delta}^o(d,x) := \frac{1\{ Y = 1, D = d, S = o \}}{ \Pr(Y = 1, D = d, S = o \mid X = x) } - \frac{1\{ Y = 0, D = d, S = o \}}{ \Pr(Y = 0, D = d, S = o \mid X = x) }$.

Once again, the causal estimand reconciles treatment variation from the experimental sample with outcome variation from the observational sample, so that their projections onto the remotely sensed variable $R$ match. Theorem (ref) allows direct effects; the key assumption in our framework is stability in Assumption (ref). Corollary (ref) and Remark (ref) extend accordingly.

remark[Some experimental outcomes] In some empirical applications, researchers may additionally collect outcomes for a small subsample of the experimental sample. This information can be directly incorporated into our procedure for estimation and inference. For this extension, define the extended sampling indicator as $\tilde{S} \in \{\{e, o \},e, o \}$. Here, $\tilde{S} = \{e, o \}$ indicates that a unit is experimental, and we have $Y$ for this unit. $\tilde{S} = e$ indicates that a unit is experimental but we do not have $Y$. Finally, $\tilde{S} = o$ indicates that a unit is observational. To apply our results, replace the expression $S = e$ with $e \in \tilde{S}$, and replace the expression $S = o$ with $o \in \tilde{S}$.
remark[Multi-valued outcomes] Our identification, estimation, and inference results generalize to discrete and continuous outcomes. Appendix (ref) extends our identification result to discrete outcomes. We generalize the conditional moment equations. Estimation and inference remain essentially the same, under a minimum rank condition that requires the RSV to predict each outcome value well. Appendix (ref) extends our results to continuous outcomes.

Estimation and Inference

As a secondary contribution, we demonstrate that our identification result provides guidance on how to choose a representation of the RSV for valid, precise, and robust inference on the causal parameter. We point out and interpret the connection between modern remote sensing and classical conditional moment analysis, which allows us to use standard techniques. We derive valid $n^{-1/2}$ inference without rate conditions and without complexity restrictions on the researcher's RSV-based predictions, justifying the use of complex deep learning algorithms that may be misspecified.

For clarity, we focus on the simple case in which the outcome is binary, there are no pretreatment covariates, and Assumption (ref)(ii) holds. We discuss the more general case with discrete outcomes and discrete or continuous covariates in Appendix (ref).

Choice of Representation for Program Evaluation

Without covariates, our main identification result in Theorem (ref) simplifies. The treatment and outcome variation are simply scalars:

equation[equation omitted — 266 chars of source]

Since $(S,D,Y)$ are binary, the denominators can be estimated by simple counts. The treatment effect is identified as $ \theta=\frac{\E(\Delta^e|R)}{\E(\Delta^o|R)}$. Therefore, $\theta=\frac{\E\{\Delta^e H(R)\}}{\E\{\Delta^o H(R)\}} $ for any representation $H(R)$ that is predictive of outcome variation in the sense that $\E\{\Delta^o H(R)\}\neq0$. While any predictive representation can be used for inference, different choices have different efficiency properties. We now ask, which representation should researchers use in practice?

By rearranging our identifying formula, we highlight the connection to classical ideas in econometrics and thereby derive an answer. Specifically, the treatment effect $\theta$ may be viewed as the coefficient in the regression model

equation[equation omitted — 111 chars of source]

where the residual $\epsilon:=\Delta^e-\theta \Delta^o$ is mean zero conditional on $R$ by Theorem (ref). Results in chamberlain1987asymptotic and newey1993efficient demonstrate that the representation that achieves efficiency within a class of models satisfying (ref) is $H^*(R)=\frac{\E(\Delta^o \mid R)}{\sigma^2(\theta, R)}$, where $\sigma^2(\theta, R) := \E\{(\Delta^e - \Delta^o \theta )^2 \mid R\}$.

For downstream causal inference, we find that the efficient representation of the high-dimensional remotely sensed variable $R\in\mathcal{R}$ is a simple, continuous scalar $H^*(R)\in \R$. The formula for $H^*(R)$ is a conditional expectation divided by a conditional variance.

This connection has a direct consequence for empirical practice: to improve efficiency with the RSV, we should not only predict the outcome $Y$ using observational data, but also predict the treatment $D$ using experimental data, and predict the sample indicator $S$ using all data. The first prediction is part of common practice, but the second and third are not. By using all three predictions, we use the RSV more efficiently.

We confirm this connection by further developing the expression for $H^*(R)$: the numerator $\E(\Delta^o \mid R)$ contains $\Pr( Y = 1, S = o|R)$ while the denominator $\sigma^2(\theta,R)$ contains $\Pr( D = 1, S = e|R)$; see Lemma (ref) below. Finally, by the definition of conditional probability, $\Pr( Y = 1, S = o|R)=\Pr( Y=1 | S = o,R)\Pr(S=o|R)$ and $\Pr( D = 1, S = e|R)=\Pr( D=1 | S = e,R)\Pr(S=e|R)$.

Inference with Learned Representations

For simplicity, we describe our inferential procedure using sample splitting with $\tr$ and $\te$ folds, though our proofs in Appendix (ref) allow for cross-fitting with any fixed number of folds. We first state our inferential procedure at a high level before filling in the details:

enumerate• Divide the sample into $\tr$ and $\te$ folds. • Learn the representation on $\tr$: $\widehat{H}(R)$. \begin{itemize} • Train predictors of $Y$, $D$, and $S$: $\textsc{pred}_Y(R)$, $\textsc{pred}_D(R)$, $\textsc{pred}_S(R)$. • Construct an initial estimate, by regressing $\widehat{\E}(\Delta^e|R)$ on $\widehat{\E}(\Delta^o|R)$: $\widehat{\theta}_{\textsc{init}}$. • Combine these into the representation: $\widehat{H}(R)$. \end{itemize} • Construct a causal estimate on $\te$: $\widehat{\theta}$.

The overall structure is familiar angrist1999jackknife,chernozhukov2018double, although here we will show that no rate condition is required on $\widehat{H}(R)$. To simplify the algorithm statement, we introduce some notation for simple events and an algebraic identity.

Let $\E_{\tr}(\cdot)=\frac{1}{|\tr|}\sum_{i\in \tr}(\cdot)$ and $\E_{\te}(\cdot)=\frac{1}{|\te|}\sum_{i\in \te}(\cdot)$. Slightly abusing notation, the counting operation, $\#_{\textsc{event}}$, can be $\E_{\tr}(1_{\textsc{event}})$ in the second step or $\E_{\te}(1_{\textsc{event}})$ in the third step, which will be clear from the context.

lemma$ (\Delta^e - \Delta^o \theta )^2= \frac{1\{ D = 1, S = e \}}{\Pr( D = 1, S = e )^2}+\frac{1\{ D = 0, S = e \}}{\Pr( D = 0, S = e )^2} + \theta^2\left[\frac{1\{ Y = 1, S = o \}}{\Pr( Y = 1, S = o )^2}+\frac{1\{ Y = 0, S = o \}}{\Pr( Y = 0, S = o )^2}\right]. $
algorithm[algorithm omitted — 2,018 chars of source]

See Appendix (ref) for explicit computations of $\widehat{\E}(\Delta^e|R)$, $\widehat{\E}(\Delta^o|R)$, and $\widehat{\sigma}^2(\widehat{\theta}_{\textsc{init}},R)$ in terms of the marginal probabilities, predictors, and initial estimate.

Sample splitting may be eliminated under complexity restrictions that tolerate simple machine learning procedures. See e.g., chernozhukov2020adversarial for a recent summary.

Robustness to Misspecification

In empirical research, the RSV is typically unstructured and high-dimensional, e.g., a satellite image. The prediction is often conducted by a complex machine learning algorithm, e.g., a deep convolutional neural network. For this realistic setting, rates of convergence are often unknown. In other words, we have no reason to believe that $\textsc{pred}_Y(R)$ converges to $\Pr(Y=1|S=o,R)$, nor that $\widehat{H}(R)$ converges to $H^*(R)$. Even carefully crafted architectures positing a generative model would typically be misspecified.

For this reason, we place a weaker regularity condition: the predictions, and hence the representation estimator, have some probability limit; they may be misspecified.

assumption[Limit] The learned representation has some mean square limit: $\E_R[\{\widehat{H}(R)-\widetilde{H}(R)\}^2]=o_p(1)$, where $\E\{\widetilde{H}(R)^2\}$ is finite, and possibly $\widetilde{H}(R)\neq H^*(R)$. This limit is correlated with outcome variation: $\E\{\widetilde{H}(R) \Delta^o\}$ is bounded away from zero.

Assumption (ref) does not require any complexity restriction, nor any rate of convergence. The moment restriction in (ref) is infinite-order Neyman orthogonal mackey2018orthogonal,chen2020mostly, so Algorithm (ref) enjoys the best of both worlds: no complexity restriction chernozhukov2018double,chernozhukov2023simple, and no rate requirement chamberlain1987asymptotic,newey1993efficient. Even if all three of its predictions are misspecified, Algorithm (ref) still delivers valid $n^{-1/2}$ inference, which is a strong form of robustness to misspecification.

In the following statements, we refer to $\Pr(D=1,S=e)$, $\Pr(D=0,S=e)$, $\Pr(Y=1,S=o)$, and $\Pr(Y=0,S=o)$ as the marginal probabilities.

proposition[Inference with known counts] Suppose Theorem (ref)'s conditions and Assumption (ref) hold. If the marginal probabilities are known and bounded away from zero, then $n^{1/2}(\hat{\theta}-\theta)\rightsquigarrow \mathcal{N}\left(0, \frac{\E\{(\Delta^e-\theta\Delta^o)^2 \widetilde{H}(R)^2\}}{[\E\{\Delta^o\widetilde{H}(R)\}]^2}\right)$. Moreover, if $\widetilde{H}(R)=H^*(R)$, then $\hat{\theta}$ is semiparametrically efficient for $\theta$ satisfying (ref) with marginal probabilities known to the researcher.
proposition[Inference with unknown counts] Suppose Theorem (ref)'s conditions and Assumption (ref) hold. If the marginal probabilities and their counting estimators are bounded away from zero, then $n^{1/2}(\hat{\theta}-\theta)\rightsquigarrow \mathcal{N}\left(0, \frac{V}{[\E\{\Delta^o\widetilde{H}(R)\}]^2}\right)$, where $V$ is defined in Lemma (ref).

When marginal probabilities are known, the asymptotic variance is standard (Proposition (ref)). In this case, (ref) matches the conditional moment of chamberlain1987asymptotic, so the semiparametric efficiency bound is known.

When marginal probabilities are unknown, the asymptotic variance is weighted by them (Proposition (ref)). The asymptotic variance in Proposition (ref) differs from the one in Proposition (ref) because estimation of the marginal probabilities introduces an additional estimation error of order $n^{-1/2}$ that we must take into account for asymptotic inference. In this case, it is simpler to use a bootstrap than an analytic variance estimator.

Theorem (ref) allows direct effects of the treatment on the remotely sensed variable, as long as treatment is non-missing and non-deterministic in the observational sample. We have already derived the conditional moment equations. Estimation and inference are similar.

remark[Inference with covariates] For clarity of exposition, this section has focused on the setting without covariates $X$. When $X$ is discrete with a finite support, estimation and inference are straightforward: apply Algorithm (ref) within each covariate stratum, then average over $X$. See Appendix (ref). When $X$ is continuous with a low dimension, there are two main modifications: (i) estimate the conditional probabilities $\Pr(D=d, S=e \mid X=x)$ and $\Pr(Y=y, S=o \mid X=x)$ as smooth functions of $x$; (ii) estimate the conditional average treatment effect $\theta(x)$ via local regression or series methods, before averaging over $X$. Appendix (ref) discusses this setting. Crucially, the main robustness property---that $n^{-1/2}$ inference requires no rate condition on the learned representation $\widehat{H}(x,R)$---continues to hold in both settings. This justifies the use of complex machine learning algorithms for RSV representations.

Program Evaluation using Satellite Images

To empirically validate our method, we conduct three semi-synthetic exercises that are increasingly realistic. We use real RSV distributions, together with

itemize• synthetic treatment effects and synthetic sample definitions; • real treatment effects and synthetic sample definitions; or • real treatment effects and real sample definitions.

Across the exercises, we use data from an experiment analyzed in muralidharan2016building,muralidharan2023general, illustrated in Figure (ref). The authors collect data in an experimental sample and an observational sample of villages in Andhra Pradesh, India. The treatment $D$ is the early introduction of Smartcards, which are a biometrically authenticated payments infrastructure. The outcome $Y$ is a measure of village-level poverty in administrative data from 2012 to 2013 (after the experiment was deployed). See Appendix (ref) for details on the experiment.

We use the geographic coordinates of each village to extract a remotely sensed variable $R$, from free, open-access databases. The RSV concatenates nighttime luminosity measures from 2012 to 2020 (a vector in $\mathbb{R}^{50}$) asher2021development and satellite images from 2019 (a high-dimensional, pre-trained embedding vector in $\mathbb{R}^{4000}$) rolf2021generalizable. These variables have been extensively validated as predictors of poverty hsw2009,jean2016combining,stoeffler2016reaching,michaels2021planning,huang2021using,sherman2023global. Our framework allows us to test their relevance for program evaluation, similar to a “first stage” exercise in instrumental variable analysis. In our semi-synthetic exercises, we will evaluate how well we can conduct program evaluation using this real RSV.

In terms of our identification framework, the semi-synthetic settings have incomplete cases (Assumption (ref)(ii)): the treatment is missing in the observational sample. The settings also have some experimental outcomes (Remark (ref)). We will be explicit about sample definitions below. Appendix (ref) contains a complete description of the data.

Our Method Outperforms Current Practice across Effect Sizes and Sample Sizes

We impose all of our assumptions in a calibrated, synthetic data-generating process (DGP). First, we simulate the binary treatment $D$ as a fair coin toss (Assumption (ref)). If $D=0$, we simulate the binary outcome $Y$ as a weighted coin toss, calibrated to the empirical probability of $Y=1$ among untreated experimental units in the real data. If $D=1$, we simulate $Y$ as a different weighted coin toss, calibrated to the empirical probability of $Y=1$ among treated experimental units in the real data.

In this baseline version of the DGP, the synthetic treatment effect is calibrated to the real one: $-0.07$. We consider additional synthetic treatment effect values $\theta=-0.07+\tau$ by augmenting the probability of $Y=1$ when $D=1$ for alternative values of $\tau$.

Next, we draw RSV values from the real data. If $Y=0$, we draw $R$ from the empirical distribution of $R|Y=0$ in the real experimental data. Likewise, we draw $R$ when $Y=1$. This imposes no direct effects (Assumption (ref)(ii)). For computational feasibility in this exercise, we use luminosity and only the initial $1,000$ satellite image features as our RSV.

Finally, to mimic a setting with missing outcomes, we delete $Y$ if $D=1$. In terms of Remark (ref), we set $\tilde{S}=e$ when $D=1$ and $\tilde{S}=\{e,o\}$ when $D=0$.\footnote{For simplicity, we do not have $\tilde{S}=o$. This could be easily done: if $D=0$, toss a coin to determine whether $\tilde{S}=\{e,o\}$ or $\tilde{S}=o$. With or without $\tilde{S}=o$, a direct extension of Theorem (ref) applies: within the identifying formula given there, simply replace $S=e$ with $e\in\tilde{S}$ in $\Delta^e$, and similarly replace $S=o$ with $o\in\tilde{S}$ in $\Delta^o$.} This imposes stability across samples (Assumption (ref)). Recent methods that use machine learning predictions as outcomes, e.g., prediction-powered inference angelopoulos2023prediction, provide no guidance for this DGP, because no treated unit has an observed outcome.

figure[figure omitted — 1,019 chars of source]
figure[figure omitted — 1,029 chars of source]

Figures (ref) and (ref) demonstrate that our method outperforms common practice across treatment effect values and sample sizes. We consider treatment effect values $\theta=-0.07+\tau$ with $\tau\in \{0,0.1, \ldots, 0.5\}$ and sample sizes $n\in\{1000,2000,3000\}$. By varying these aspects of the synthetic DGP, we evaluate relative performance for different signal-to-noise ratios.

We compare the bias and root mean square error of the two methods. Whereas the bias of our method is always small and vanishing with sample size, the bias of the common practice is typically very large and constant across sample sizes. Common practice has positive or negative bias for the treatment effect, consistent with Proposition (ref). While the variance of our method is similar to the variance of the common practice, our method's large improvement in bias translates into similar improvement in mean square error when the sample size is large enough.

This exercise has clear consequences for empirical practice. If the goal is to conduct valid inference on the treatment effect, common practice can be misleading due to large bias that does not vanish as the sample size increases. It can return the wrong sign, as characterized by Proposition (ref). By contrast, our method reports unbiased estimates, with valid asymptotic inference.

Our Method Recovers the True Effect with Randomly Missing Outcomes

While our first exercise involves synthetic treatment effects, our second exercise involves real treatment effects. We now ask: compared to an unbiased benchmark, how does our method perform? The benchmark is the difference-in-means estimate and confidence interval an economist would obtain if they could observe all treatments and outcomes in the experiment. Our method gives an estimate and confidence interval using treatments and RSVs in an experiment, and outcomes and RSVs in an observational sample.\footnote{Following Remark (ref), we also observe some outcomes in the experiment, as described below.}

We conduct this exercise with three possible poverty measurements: (i) is the village in the bottom quartile of villages for per capita consumption, (ii) does a village have only low income households, and (iii) does a village have only low and middle income households. While (i) is a measure of average consumption, (ii) and (iii) describe the income distribution using formal categories from Indian administrative data; see Appendix (ref) for details.\footnote{We used the low consumption outcome (i) in the first semi-synthetic exercise.}

In this exercise, we now use the full RSV: the $50$ luminosity measures, and the full $4000$-dimensional embedding of satellite images.

In a semi-synthetic way, we classify villages into the samples described in Assumption (ref)(ii) and Remark (ref). We classify the real observational villages as $\tilde{S}=o$, for which we observe $(Y,R)$. Among the real experimental villages, we randomly classify some as $\tilde{S}=e$, for which we observe $(D,R)$, and some as $\tilde{S}=\{e,o\}$, for which we observe $(D,Y,R)$. In other words, for real experimental villages, we randomly delete half of their outcomes.

This exercise satisfies the key assumptions of our framework. The real treatment variable was randomized in the real experiment (Assumption (ref)). Stability of the RSV conditional distribution is plausible, as demonstrated in Figure (ref) (Assumption (ref)). Finally, because the real data have incomplete cases, we must argue that the treatment only affects the RSV via the outcome (Assumption (ref)(ii)). As supporting evidence, Figures (ref) and (ref) show that the distribution of $R|Y=y,D=1$ visually matches the distribution of $R|Y=y,D=0$.

Another requirement in our framework is that the RSV is relevant: $\E\{H(R) \Delta^o\} \neq 0$ in Corollary (ref). Intuitively, the representation of the RSV $H(R)$ used for inference should be correlated with outcome variation $\Delta^o$. This requirement is testable; for our learned representation in Algorithm (ref), we can test whether the empirical analogue $\E_n\{\widehat{H}(R) \widehat{\Delta}^o\}$ is significantly nonzero. Figure (ref) confirms that our RSV is relevant.

figure[figure omitted — 1,348 chars of source]

We further interpret and compare our learned representation of the satellite image in Appendix (ref). Our optimal representation is based on three predictions: the outcome, treatment, and sample indicator given the RSV. By contrast, common practice only uses the prediction of the outcome given the RSV. Interestingly, our optimal representation allows for extrapolation, i.e. negative weights on villages, whereas common practice does not.

figure[figure omitted — 1,649 chars of source]

Figure (ref) (“synthetic samples”) shows that our method recovers the treatment effect estimated by the unbiased benchmark in this exercise. For each poverty outcome, our treatment effects have the same signs and similar magnitudes as the benchmark, i.e. as if the economist could observe all treatments and outcomes in the experiment. These effects are consistent in sign with the findings in muralidharan2023general: early adoption of Smartcards reduces poverty. Compared to the benchmark, our confidence intervals are similar, and sometimes shorter; efficiently using the additional information in RSVs can boost statistical power.\footnote{Assumption (ref)(ii) is an additional restriction that may improve asymptotic precision. Relative magnitudes of standard errors may also reflect finite sample estimation error.}

These findings have a practical implication: our method may allow economists to significantly reduce survey costs. In this exercise, we mimic what would happen if an economist paid surveyors to collect outcomes for half of the experimental villages, and relied on free satellite images for the other half. We can calculate the savings, compared to paying surveyors to collect outcomes for all of the experimental villages. If surveying each individual in a village conservatively costs $\$0.50$, our results suggest savings of $\$3.6$ million.\footnote{Table (ref) summarizes the number of experimental villages and the populations in experimental villages, which we use to calculate savings. This calculation is highly conservative; viviano2022policy find that phone surveys in Pakistan cost $\$7$ per individual, rather than $\$0.50.$}

Our Method Recovers the True Effect in a Realistic Setting

The third exercise is identical to the second exercise, except that we classify villages in a more realistic way. As before, we classify the real observational villages as $\tilde{S}=o$, for which we observe $(Y,R)$. Among the real experimental villages, we now classify the treated ones as $\tilde{S}=e$, for which we observe $(D,R)$, and the untreated ones as $\tilde{S}=\{e,o\}$, for which we observe $(D,Y,R)$. In other words, for real experimental villages, we systematically delete the outcomes of all treated villages. Methods based on missingness-at-random provide no guidance in this setting.

As before, this exercise appears to satisfy our identifying assumptions. The real treatment variable was randomized in the real experiment (Assumption (ref)). The conditional distribution of the RSV appears stable in Figure (ref) (Assumption (ref)). There do not appear to be direct effects in Figures (ref) and (ref) (Assumption (ref)(ii)). The RSV is relevant in Figure (ref) (Corollary (ref)).

In this more challenging exercise, we find qualitatively similar results. Our method recovers comparable estimates and confidence intervals to the benchmark in Figure (ref) (“real samples”). Our learned representations allow efficient extrapolation in Figure (ref).

We conclude this discussion with two final remarks on the interpretation of our semi-synthetic exercises.

remark[Spillovers] In the first exercise, we compare our method to the true value of the treatment effect in Definition (ref). There are no spillovers by design. In the second and third exercises, we compare our method to a natural benchmark: a difference-in-means estimate an economist would obtain if they could observe all treatments and outcomes in the experiment. If there are no spillovers, then we are comparing our method to an unbiased estimate of the treatment effect in Definition (ref). If there are spillovers, then we are comparing to an unbiased estimate of a reasonable estimand: a difference-in-means of expected outcomes, indexing by each unit's treatment status and averaging over other all other units. In Appendix (ref), we consider spillovers more carefully, by introducing additional notation. We discuss how our method can be adapted to estimate the global average treatment effect, a common parameter of interest in models with spillovers.
remark[Timing] Outcomes were measured at the end of the experiment in 2013, so we study the effect on poverty in 2013. Meanwhile, RSVs were measured through 2020. Appendix (ref) confirms that it is valid to use RSVs measured after the outcomes. The auxiliary sample links outcomes from 2013 with RSVs through 2020, so the estimand and estimator remain unchanged despite these differences in timing.

Recommendations for Practice

Common empirical practice can be highly biased when the remotely sensed variable is post-outcome, i.e., when the RSV is caused by the outcome of interest. For example, changes in satellite images are caused by fires or variation in local income, but not vice versa. Theoretically, we demonstrate that this bias can have arbitrary sign. Empirically, we find that common practice may attenuate the treatment effect in a real environmental application, and it may have positive or negative bias in a semi-synthetic development application.

However, the core intuition underlying empirical work is powerful: the conditional distribution of the RSV given the outcome and treatment should be stable across samples. By formalizing this intuition, we develop a framework that nonparametrically identifies treatment effects and efficiently uses RSVs for program evaluation.

We conclude with some concrete recommendations for researchers conducting program evaluation with remotely sensed outcomes. We organize our recommendations into (i) implementation steps and (ii) diagnostic tests.

Three Steps for Robust Inference with RSVs

To apply our framework, researchers should follow three steps when designing studies with remotely sensed outcomes.

Step 1: Construct a credible observational sample. Build an auxiliary, observational sample that contains RSVs with linked outcomes. The key identifying assumption is stability: the conditional distribution of the RSV given the covariates, treatment, and outcome should be stable across the experimental and observational samples. Consequently, the researcher's most important design decision is how to construct an observational sample in a way that makes the stability assumption defensible.

Stability is most credible if the researcher randomly selects units for the experimental sample and observational sample prior to conducting the experiment. When random selection is infeasible, researchers should construct an observational sample that is similar to the experimental sample along a few dimensions. (i) Select units that are geographically close to the experimental sample. (ii) Select units with similar baseline covariates, e.g. rural-urban compositions, to the experimental sample. (iii) Select units that align with the timing of the experimental sample. That is, in the observational sample, collect outcomes measured in the same period as the experiment's endline, and collect RSVs measured at the same time that RSVs will be collected in the experiment. If exact alignment is not possible, then match the time difference so that the temporal relationship might be preserved.

Step 2: Estimate three predictions from RSVs. Efficient use of RSVs for downstream causal inference requires three predictions: the outcome, treatment, and sample indicator given the RSV. Common practice uses only the first prediction, in a way that sacrifices not only accuracy but also precision. In other words, by using all three predictions, our method improves not only bias but also variance. Researchers can construct these predictions using complex deep learning algorithms with unknown statistical properties.

Step 3: Apply robust inference algorithm. Algorithm (ref) aggregates the three predictions into an efficient scalar representation of the high-dimensional RSV (where efficiency is within the class of models satisfying the conditional moments we derive). This aggregation reduces the dimension while preserving information relevant for causal inference. Critically, Algorithm (ref) is robust to misspecification; it provides valid, $n^{-1/2}$ inference even if all three predictions are misspecified. It only requires that the learned representation converges at any rate to any limit that predicts outcome variation. This robustness property justifies the use of deep neural networks, which are widely implemented by empirical researchers in this literature, while providing the safeguard of a principled inferential framework.

Three Diagnostics for RSVs

In addition to the point estimate and confidence interval of Algorithm (ref), researchers should report the following diagnostics to defend the credibility of their RSV.

Diagnostic 1: RSV relevance test. Just as weak instruments threaten inference in instrumental variable models, weak RSVs---those that poorly predict the outcome---threaten inference in our framework. Researchers should assess whether the correlation between their estimated representation and the outcome is significantly different than zero. Figure (ref) illustrates this diagnostic, showing that satellite images are strongly related to village-level poverty. When relevance is weak, researchers should either improve RSV measurements (e.g., by using higher resolution images or additional data sources) or acknowledge that inference is inconclusive.

Diagnostic 2: Joint test of identifying assumptions. Our identifying assumptions are jointly testable. If two different RSV representations yield significantly different treatment effect estimates, then reject the null hypothesis that our identifying assumptions jointly hold: stability (Assumption (ref)) and no direct effects (Assumption (ref)(ii)). This is essentially a specification test. Researchers should verify that their results are robust to alternative ways of representing the RSV, such as using different machine learning architectures or different subsets of RSV features.

Diagnostic 3: RSV density plots if possible. When outcomes are observed for some experimental units, researchers can directly visualize whether stability holds by comparing conditional densities. Specifically, plot the RSV density in the experimental sample against the RSV density in the observational sample, for each combination of treatment and outcome values. If these densities align, then stability is plausible, as in Figure (ref). Similarly, researchers can assess the no-direct-effects assumption (Assumption (ref)(ii)) by comparing the RSV density of treated units against the RSV density of untreated units, for each outcome value. If treatment only affects the RSV via the outcome, these densities should coincide, as in Figures (ref) and (ref).

Broader Implications

By following these steps, researchers can conduct rigorous program evaluation while substantially reducing costs. In our semi-synthetic exercise, we recover treatment effects and confidence intervals comparable to those obtained with direct outcome measurements, despite using satellite images for half the sample.

Our framework generalizes beyond simple randomized experiments. Appendix (ref) extends our identification results to quasi-experiments (e.g., instrumental variables and difference-in-differences). The key insight remains the same: stability in how the RSV relates to other variables across samples allows us to transport this relationship to the sample where the outcome is missing.

Our framework poses new questions. Future research may study a new experimental design problem: how to simultaneously design treatment assignment, outcome collection, and sensor deployment to maximize statistical power subject to budget constraints. Our main identification result provides the foundation to set up this problem. Our main result may inform not only the use of satellite images, but also the design of cheap, noisy surveys. Finally, future work may extend our inference guarantees to accommodate high-dimensional outcomes.

As remote sensing expands the frontier of data availability in economics, our framework offers practical principles for its use in program evaluation. This paper corrects a widespread but biased practice and provides concrete recommendations on how to incorporate satellite images and other remotely sensed data into causal inference. These new data sources and new methods broaden access to rigorous impact evaluation in cost-constrained environments.

\spacingset{1} \DeclareRobustCommand{\VAN}[2]{#2} \DeclareRobustCommand{\VAN}[2]{#1}

\spacingset{1.5}