EconBase
← Back to paper

Generalizability with ignorance in mind: learning what we do (not) know for archetypes discovery

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

100,203 characters · 11 sections · 48 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Generalizability with ignorance in mind: learning what we do (not) know for archetypes discovery

\spacingset{1.45}

abstractWhen studying policy interventions, researchers often pursue two goals: i) identifying for whom the program has the largest effects (heterogeneity) and ii) determining whether those patterns of treatment effects have predictive power across environments (generalizability). We develop a framework to learn when and how to partition observations into groups of individual and environmental characterstics within which treatment effects are predictively stable, and when instead extrapolation is unwarranted and further evidence is needed. Our procedure determines in which contexts effects are generalizable and when, instead, researchers should admit ignorance and collect more data. We provide a decision-theoretic foundation, derive finite-sample regret guarantees, and establish asymptotic inference results. We illustrate the benefits of our approach by reanalyzing a multifaceted anti-poverty program across six countries.

Introduction

Given the rise of experimental and quasiexperimental methods in social science and access to increasingly rich data, researchers can now measure the treatment effects of policy interventions in larger, more representative populations and across diverse contexts. In many cases, the policy maker seeks to understand where and for whom to scale promising interventions and when more data or pilot experiments are necessary. At the same time, social scientists are often interested in model discovery to infer economic behaviors from the data (e.g., a “law of motion" or “story" of how agents behave in an environment). Both sets of goals require an understanding of: i) the extent to which patterns in data from a specific environment can be generalized to other contexts borenstein2021introduction; and ii) patterns of heterogeneity in treatment effects based on a potentially high-dimensional set of observable characteristics.\footnote{We can interpret the study of heterogeneity both for applications in meta-analysis, where researchers have access to multiple studies (e.g., with heterogeneous site-specific characteristics or research teams), or for applications in treatment effect heterogeneity with a single experiment.}

In this paper, we present an econometric framework and a set of empirical tools for the joint task of predicting effect heterogeneity and assessing generalizability across environments. Our goal is to understand whether there are systematic groups of observable characteristics (archetypes) predictive for others through a statistical or economic model. Implicit in this goal is an equally important second aspect: we want to detect those contexts that are uninformative for the construction of the archetypes and, therefore, for which we are unable to claim generalizability. That is, rather than drawing conclusions about treatment effects in all environments in the data as standard approaches do,\footnote{This would force pooling information across possibly highly heterogeneous environments, leading to misleading conclusions when different “forces" explain economic phenomena in disparate contexts.} we identify which aspects of the data cannot be pooled together to inform where additional evidence is needed. We refer to the group of observations that may exhibit a lack of generalizability as basin of ignorance.

As an illustrative example, consider the multifaceted “Graduation program” studied experimentally in six countries by banerjee2015multifaceted. The program's goal is to lift individuals out of extreme poverty through income generation, and it typically includes a large asset transfer (e.g., cows), training, savings accounts, and short-term cash transfers. This is a context where heterogeneity and generalizability are of first-order importance. Ex ante, it is unclear which types of person may react the most to the program (e.g., by age, relative wealth, marital status) as well as how market conditions might affect the program success (e.g., through credit access, labor demand, or supply chains for dairy products).\footnote{While a researcher could explicitly model each of these forces and incorporate them structurally into estimation, we think this is practically difficult for a few reasons. First, some of these factors may not be directly observable (for example, risk preferences). Second, different mechanisms might be at play in different places, and may be unknown ex-ante.} To navigate potentially high-dimensional heterogeneity, the researcher needs to understand generalizability, or the extent to which, say widows in Pakistan, can (or cannot) inform our understanding of say young job seekers in Peru. Moreover, if the estimates for some individuals and contexts contain sufficient noise and exhibit very different effects from others, we might not be comfortable making any inference about them without collecting more data.

This paper introduces a general framework where a researcher has access to data from a number of environments (either within site e.g., villages or cross sites e.g., countries) that include individual outcomes observed after an intervention, environmental and individual characteristics. For each individual in the study, using the data collected so far, the researcher can estimate (predict) treatment effects conditional on observable characteristics. However, unlike existing methods, here the researcher has the option to abstain from making a prediction, admit ignorance, and recommend collecting more experimental evidence at a given cost. The optimization problem balances two objectives: predicting effects using a given statistical or economic model and recommending where to collect new data to build better predictions.

We provide two equivalent interpretations. From a decision-theoretic perspective, here generalizability quantifies whether the researcher would rather rely on existing evidence instead of collecting more data at a given cost. This approach is also equivalent to a Bayesian decision maker who can decide where to elicit more evidence by imposing common priors (i.e., a statistical model) only over an ex ante unknown subset of the data, and allowing for arbitrary heterogeneity on the remaining observations.

Our approach stands in contrast to existing procedures for meta-analysis and effect heterogeneity, which tend to force a statistical or economic model across all contexts observed in the data. Examples include Bayesian hierarchical models (that through common priors form posteriors across all environments) and frequentist methods that similarly use sparsity or smoothness restrictions and do not leave scope for exploration. Here, we take the view that generalizing across contexts is an epistemic act: it is often unclear when and why all contexts should be informative to others.

A way to understand this optimization problem is through the following trade-off: given a set of possible prediction functions for treatment effects (e.g., smooth functions), we would like to maximize the number of units for which we form a prediction (claim generalizability) but also minimize the prediction error on such individuals. If the true data-generating process is complex (e.g., nonsmooth), this trade-off would require abstaining from making predictions for some of the units. In its dual formulation, our objective criterion maximizes the number of individuals for which we form a prediction under a constraint on the largest prediction error that we can tolerate.

Given that for some observations researchers might admit ignorance, the prediction we form for the remaining ones should not pool information across all observations. We refer to these as generalizability-aware predictions: these are predictions that jointly optimize over the assignment of observations to the basin of ignorance and archetypes.

Using an available (pilot) study, we construct estimators in two steps. First, for each (small) group $x$ of the observable individual-level and environmental characteristics, we form unbiased but possibly noisy estimates of the conditional average effect (CATE) and its variance. Second, we assign each of these groups to either an archetype or the basin of ignorance. Assignment to the basin of ignorance incurs a fixed cost. The estimated cost for groups comprising an archetype is instead equal to the approximation error of the statistical model, estimated by taking the squared prediction error, and subtracting the within sampling variation at $x$.

We justify our approach through a set of theoretical guarantees. We focus on regret, i.e., the difference in terms of the researcher's loss function between the best set of predictions with no estimation error and our estimator. Without imposing distributional assumptions other than standard moment restrictions, we show that regret converges to zero at a fast (parametric) rate in the size of the study. This is possible by assuming and leveraging the independence (but not identical distributions) of each observation together with geometric restrictions on the prediction function class and basin of ignorance. Such guarantees require novel derivations to jointly control the supremum of an empirical process obtained from a prediction and classification function class. In addition, we provide guarantees for inference to, e.g., test whether heterogeneity is constant in certain characteristics, and derive computational properties.

We apply our method to the multifaceted Graduation program and observe large positive effects on an index of outcomes for individuals with low baseline consumption and assets and smaller effects on households with moderate levels of baseline consumption. The method places the richest households in the basin of ignorance. In contrast, forcing pooling across all individuals would lead to significant increases in estimation error, and misguided conclusions for sub-populations with higher level of baseline consumption or assets. A set of simulations calibrated to our empirical application demonstrate up to fifty percent improvement reductions in prediction error over the generalizable set, when our method is compared to shrinkage (empirical Bayes) procedures and forest-based methods, and even when only $4\%$ of observations in the basin of ignorance exhibit large and unpredictable heterogeneity. These results illustrate the importance of the basin of ignorance both for detecting where we lack sufficient evidence and also for improving robustness where effects do generalize.

We connect with the literature on meta-analysis and machine learning-based heterogeneity methods, which are increasingly prevalent in applied work.\footnote{For example, recent meta-analyses tackle topics including deworming, cash transfers, education interventions, the link between democracy and growth, and tests of Allport's contact hypothesis croke2024meta,angrist2023implementation,crosta2024unconditional,doucouliagos2008democracy,paluck2019contact. A related empirical literature has also emerged focusing on policy design and targeting banerjee2021selecting,haushofer2022targeting.} In nesting generalizability and effect heterogeneity within the same framework, we hope that our method can be practically useful for a wide range of applications. In each of these domains meager2022aggregating,spiess2023finding, chernozhukov2018generic, wager2018estimation, venkateswaran2024robustly, bonhomme2015grouped, ishihara2021evidence, menzel2023transfer, adjaho2022externally, manski2004statistical, athey2021policy, kitagawa2018should, existing literature has focused on producing estimates of treatment effect heterogeneity (or making treatment decisions) for any context in the population of interest. Our innovation with respect to all such references is the possibility for the researcher to abstain from making predictions (learning where not to pool observations and instead elicit more evidence).

Specifically, the concept of ignorance introduced here allows typical assumptions imposed by the treatment effect heterogeneity literature wager2018estimation, chernozhukov2018generic, bonhomme2015grouped to hold only locally for a (ex-ante unknown) subset of the data, as opposed to hold globally in the data as assumed by this literature, therefore making such methods more robust in practice. Similarly, existing methods that account for statistical noise to maximize power or via shrinkage spiess2023finding, meager2019understanding do not allow units that, even absent estimation error, cannot be correctly predicted due to misspecification. Importantly, such misspecification can also pollute predictions on the remaining units. As we highlight further in Section (ref), similar differences apply more broadly to typical Bayesian hierarchical models (BHMs).

We connect to the robust statistics literature huber2011robust,garcia1999robustness. Here, instead of positing ex ante a (robust) loss function, which can be difficult to choose in practice, we embed the estimation of the non-generalizable set in a formal decision problem. Our approach of assigning observations to the basin of ignorance therefore can tackle the sensitivity of point estimates to deleting few observations, which has been shown to be prevalent in applied work broderick2020automatic. Our decision-theoretic motivation that combines statistical modeling with exploration and our (regret) guarantees are also novel.

Other studies of generalizability have focused on quantifying heterogeneity for a given prediction function when there is no opportunity of further experimentation. See, for example, deeb2019clustering, bisbee2017local, andrews2022transfer, and manski2020toward. Another body of work models heterogeneity to inform experimental design gechter2024selecting, olea2024externally in the absence of empirical evidence. Our contribution lies between these two phases of research: we use existing data to inform future experimentation, but also to produce counterfactual predictions when accurate. This justifies our approach, which learns where we lack sufficient evidence from the data, trading of its costs and benefits.

Finally, this paper builds to our knowledge the first connection between classification with rejection options in machine learning chow1957optimum, chow1970optimum, cortes2016learning, franc2023optimal, and more broadly shafer1992dempster's theory to the literature on treatment effect heterogeneity. Rejection options allow binary classifiers to abstain from making a prediction, focusing on unconstrained decisions; recent work on regression assumes correct model specification or exchangeability assumptions denis2020regression, sokol2024conformalized. None of these references studies generalizability or effect heterogeneity. Here we consider a more general joint classification and regression problem, with non-vanishing misspecification error and non-exchangeability (with in addition possible constraints on the estimators' class). This motivates a different class of estimators that compare between and within variation of treatment effects estimates. It also requires novel guarantees on regret and a novel decision-theoretic foundation that connects the rejection option to future experimentation.

A framework for generalizability

Consider a settings where individuals may be organized into many (very small) groups. Such groups may contain the cross-product of individual-level characteristics and experimental-level characteristics such as the site or country of the experiment. Formally, individuals are organized into many observable types $x$, where $x \in \mathcal{X}$ and $\mathcal{X}$ is discrete but possibly high-dimensional (i.e., $\mathcal{X}$ can grow proportionally with the sample size). Researchers are interested in studying a given estimand for group $x$, which we refer to as property $\phi(x) \in \mathbb{R}$, such as conditional average treatment effect for a given outcome. (In Appendix (ref) we also allow for multiple outcomes/properties.) In practice, we only observe a noisy (pilot) study. We introduce our main framework absent of sampling uncertainty in this section, and return to sampling uncertainty in the following section.

Ignorance and generalizability-aware predictions

In principle, the function $\phi: \mathcal{X} \mapsto \mathbb{R}$ can be highly complex. Such complexity may encode heterogeneity across characteristics, contexts, etc. Researchers' goal is to summarize $\phi$ with a simpler approximation function $\bar{\phi}(x), \bar{\phi} \in \mathcal{F}$, where $\mathcal{F}$ encodes economic, communication or statistical constraints. For example, researchers may want to summarize heterogeneity into a finite number of groups chernozhukov2018generic, athey2016recursive.

However, approximating $\phi(\cdot)$ with some simpler function $\bar{\phi}$ has two drawbacks: (i) it can lead to poor approximations for some observations $x \in \mathcal{X}$ where $\bar{\phi}$ may perform poorly (e.g., outliers); (ii) such units with large heterogeneity can pollute the choice $\bar{\phi}$ and increase prediction errors for the remaining units (see e.g., Section (ref)).

Admitting ignorance Motivated by these considerations, we introduce a framework where researchers may either make a prediction using an approximation function $\bar{\phi}$ or abstain at a given opportunity or economic cost that we define as $\sigma^2$. Conceptually, here $\sigma^2$ denotes the cost of collecting further evidence in a given context. (All our results extend to $\sigma^2$ being a function of $x$.)

Specifically, define $\pi(x) \in \{0,1\}$, a binary decision denoting whether the researchers make a prediction as a function of $x$. The researcher incurs a loss\footnote{The loss function captures the researcher's objective function. For instance, when $L(\bar{\phi}, \phi) = (\bar{\phi} - \phi)^2$, our leading example throughout, the objective $(\bar{\phi} - \phi)^2$ defines the difference in accuracy from using a aggregator. When instead $\phi$ denotes a welfare effect, $L(\bar{\phi}, \phi) = \phi 1\{\phi \ge 0\} - \phi 1\{\bar{\phi} \ge 0\}$ denotes the welfare regret of taking an action using $\bar{\phi}$ instead of $\phi$. }

equation[equation omitted — 249 chars of source]

Component (ii) is our first key innovation: we consider a scenario in which the researcher makes a prediction $\bar{\phi}$ (e.g., a posterior mean obtained from previous experiments) or can abstain, and recommend collect further evidence about $\phi(x)$.

Whenever $\sigma^2 \rightarrow \infty$, there is no scope for ignorance and new research. This is the underlying assumption of all existing estimators for heterogeneity, but undesiderable when researcher have the possibility to inform where further evidence is needed.

This formulation reflects an important idea: the cost of making a poor prediction—especially by pooling over unrelated groups—can outweigh the opportunity cost of withholding prediction. Conceptually, errors from pooling observations when we should not can be epistemically misleading, suggesting generalizability where none exists. Collecting the loss across observations, we define the researcher's reward $$

alignedW_\phi(\pi;\sigma, \bar{\phi}) &= -\mathcal{R}_\phi(\pi;\bar{\phi}) - \sigma^2 \Big(1 - \bar{N}(\pi)\Big),

$$ where $\mathcal{R}$ denotes an approximation error from making predictions and $\bar{N}$ the average number of units for which researchers do not abstain from making a prediction,

equation[equation omitted — 234 chars of source]

We think of $p(x)$ as a target types' distribution.

\paragraph{Generalizability-aware predictions} Once we give to researchers the possibility of elicit further evidence, the construction of the prediction function may also change. Our next key innovation is to jointly build predictions taking into account ignorance.

Even with no statistical noise for $\bar{\phi}$, existing estimators for heterogeneity do not allow for ignorance bonhomme2015grouped, wager2018estimation, chernozhukov2018generic. This can make the choice of $\bar{\phi}$ sensitive to (possibly few) units that fail to be well approximated by some prediction function $\bar{\phi} \in \mathcal{F}$. Returning to our example of grouping heterogeneity into a few groups, it might be that the construction of such groups is sensitive to a few units in the population. What we would like to do, instead, is to build “good predictions" only for those subgroups for which effects can be generalized and claim ignorance otherwise. As we show in the next subsection, ignorance here connects to epistemic ambiguity about treatment effects.

defn[Generalizability aware predictions and basin of ignorance] For given policy spaces $\pi \in \Pi$, and function class $\mathcal{F}$ containing functions $\bar{\phi}: \mathcal{X} \mapsto \mathbb{R}$ define the generalizability aware predictions as \begin{equation} \begin{aligned} \Big(\pi^\star, \bar{\phi}^{\star}\Big) \in \mathrm{arg} \max_{\pi \in \Pi} \max_{\bar{\phi} \in \mathcal{F}} W_\phi(\pi; \sigma, \bar{\phi}). \end{aligned} \end{equation} We define the basin of ignorance as the set $\mathcal{A} = \Big\{x \in \mathcal{X}: \pi^\star(x) = 0\Big\}$ and the set of generalizable archetypes as its complement $\mathcal{X} \setminus \mathcal{A}$. We refer to $1/\sigma^2$ as resolution.

We define generalizability-aware predictions as those that maximize reward over both the choice of the basin of ignorance and the prediction space. We refer to $1/\sigma^2$ as model resolution given its tight connection to the approximation error we are willing to tolerate (Remark (ref)). Finally, note that the function class $\mathcal{F}_\pi$ may also depend on $\pi$, implicit here for notational convenience. An illustration is in Figure (ref).

\paragraph{Summary of the decision problem} The decision problem goes as follows:

itemize• For a given prediction function $\bar{\phi}$ that aims to approximate $\phi$, researchers either predict effects with $\bar{\phi}(x)$, or abstain and admit ignorance at a cost $\sigma^2$. Given a pre-specified partition of $\mathcal{X}$, $\Pi$, this decision is defined as $$ \small \begin{aligned} \pi: \mathcal{X} \mapsto \{0,1\}, \quad \pi(x) = \begin{cases} 1 & \text{ if } \text{make prediction with } \bar{\phi} \\ 0& \text{ admit ignorance} \end{cases}, \quad \pi \in \Pi. \end{aligned} $$ • For a given type $x$, the researcher pays an expected cost $ L(\bar{\phi}(x), \phi(x)) \pi(x) + \sigma^2 (1 - \pi(x)). $ The reward $W_\phi(\pi;\sigma, \bar{\phi})$ aggregates over individuals with known weights $p(x)$. • Researchers optimize jointly $\pi \in \Pi, \bar{\phi} \in \mathcal{F}$. We think of $\Pi$ and $\mathcal{F}$ having bounded complexity, encoding communication or economic constraints (Assumption (ref)).\footnote{See for example kitagawa2018should,venkateswaran2024robustly. }
rem[Choosing $\sigma^2$ in practice] A simple interpretation of $\sigma^2$ is through the lens of duality theory. From dual theory, we can typically find a constant $\lambda_\sigma$ such that maximizing reward is equivalent to \begin{equation} \begin{aligned} \max_\pi \bar{N}(\pi), such that \mathcal{R}_\phi(\pi, \bar{\phi}) \le \lambda_\sigma. \end{aligned} \end{equation} The optimization corresponds to maximizing the probability over which a prediction is made, under the constraint that the approximation is sufficiently small. Researchers can equivalently choose $\lambda$ in lieu of $\sigma^2$ (e.g., $20\%$ the error of using a common mean): these capture preferences towards the largest approximation error we can tolerate. An equivalent formulation is to minimize $\mathcal{R}_\phi$ under a lower bound on $\bar{N}(\pi)$. This encodes preferences for abstaining only for a small fraction of the population. In our application, we illustrate how reporting results with several values of $\sigma^2$ is beneficial for decision-making. \qed
exmp[Connections to physical sciences] Consider a physicist with basic knowledge of Newtonian mechanics (and therefore drag) but no knowledge of electromagnetism. The physicist wants to study the acceleration of objects dropped down tubes of different materials in different laboratories. Here $x$ indexes combinations of the (i) object's size, (ii) mass, (iii) tube's material, etc. Most objects accelerate downward at 9.8 m/s$^2$; drag takes effect with cross-sectional surface area when the lab is filled with denser gas. However, something striking happens for $x = (\cdot, \text{magnet}) \times (\cdot, \text{metal})$, even in vaccuums. Magnets inside some metal tubes show zero acceleration. (In fact, we now know that the motion of a magnet into a conductive non-magnetic metal induces an upward electromagnetic force (Lenz's law, via Eddy currents)). In our framework, the magnet-in-metal case is assigned to the basin of ignorance: its behavior is too different to pool with the rest. This will encourage the researcher to explore this phenomenon further without being able to form, from existing data, a coherent theory that does not include electromagnetism. But conventional techniques force air resistance and electromagnetic forces to pool, which of course is unnatural. Economics only complicates the problem that emerges even with basic physics. Suppose, we are interested in building a useful (not necessarily “true") model to predict or interpret the effect of the multifaceted program in banerjee2015multifaceted, conducted across multiple countries. Here, we can think of different $x$ as observable characteristics of individuals in different countries. Researchers may posit a (potentially large) set of ex ante “reasonable” statistical or economic models of how individuals may react to the intervention. However, given the complexity of this intervention, it is unlikely that simple and interpretable models can summarize all possible mechanisms; at the same time, it would be inappropriate to pool contexts where different microfoundational stories are at play. The researcher instead would like to learn what the (small) number of tractable models are that have predictive power (e.g., for decision-making or model discovery) and where tractable models instead fail to explain the data, motivating collecting further evidence.
figure[figure omitted — 1,223 chars of source]

A decision-theoretic interpretation

We pause here and provide a decision-theoretic foundation when to goal is to learn treatment effects under a squared loss function. Researchers construct from a pilot study precise estimates $\bar{\phi} \in \mathcal{F}$. Here $\mathcal{F}$ encodes communication, economic or statistical constraints. Researchers can instead recommend to construct a possibly noisy but (approximately) unbiased $\phi^{new}(x)$, by e.g., collecting new evidence. For instance, $\phi^{new}$ may define a non-parametric estimator from a new experiment. For simplicity, let each $\bar{\phi} \in \mathcal{F}$ have no statistical noise, which holds (asymptotically) under complexity restrictions on $\mathcal{F}$. We return to settings with statistical noise in the next section.

assResearchers can report $\Big(\pi \bar{\phi}, \pi\Big)$ for some $\pi \in \Pi, \bar{\phi} \in \mathcal{F}$, with $\mathcal{F}, \Pi$ encoding modeling or communications constraints. Whenever $\pi(x) = 1$, an audience form a prediction $\bar{\phi}(x)$ about $\phi(x)$. Whenever $\pi(x) = 0$, an audience collects new evidence and form an unbiased but noisy prediction $\phi^{new}(x)$ about context $x$ with $\mathbb{E}[\phi^{new}(x)] = \phi(x)$ and $\mathbb{V}(\phi^{new}(x)) = \sigma^2$.

Intuitively, the researcher can shape the prediction (belief) of an audience by either extrapolating effect $\phi(x)$ in context $x$ with a simple function $\bar{\phi}(x)$ or recommending collecting new evidence.

By letting $x \sim p$, the risk under a squared loss function is defined as

equation[equation omitted — 183 chars of source]

Intuitively, the risk defines the expected prediction error from either relying on existing evidence, as opposed to asking for additional one.

prop[Interpretation of $\sigma^2$] Let Assumption (ref) hold, consider a squared loss function $L(\cdot)$. Then for any $\pi \in \Pi, \bar{\phi} \in \mathcal{F}$, $ \mathcal{L}_\phi(\bar{\phi}, \pi) = - W_\phi(\pi;\sigma, \bar{\phi}). $
proofSee Appendix (ref)

Proposition (ref) illustrates the equivalent interpretation of $\sigma^2$ as the noise when collecting new evidence from context $x$, as opposed of relying on extrapolation through some $\bar{\phi} \in \mathcal{F}$, that may encode an economic or statistical model.

\paragraph{Connection with misspecification and Bayesian models} It is instructive to compare our method to shrinkage methods and canonical Bayesian Hierarchical Models (BHMs) in particular which are are the dominant tool in meta-analyses rubin1981estimation,gelman2006prior,meager2022aggregating,crostaetal2024cash, gechter2024selecting. To understand this connection it is useful to impose a simple prior assumption although this is not used for our subsequent results other than Corollary (ref). Specifically, suppose we can write for some $\bar{\phi}^\star \in \mathcal{F}, \pi^\star \in \Pi$,

equation[equation omitted — 231 chars of source]

Here, Equation (ref) states that we can find a function $\bar{\phi}^\star$ in a restricted function class which is correctly specified locally for some contexts $x$. In the remaining contexts, $\eta^2$ characterizes the degree of misspecification, as $\phi(x)\neq \bar{\phi}^\star(x)$.

Define the posterior expectation for some $\bar{\phi}^\star, \pi^\star$ and $\phi^{new}$ as

equation[equation omitted — 342 chars of source]

That is, once an audience collects additional evidence $\phi^{new}$, $\eta^2$ defines how much the audience will rely on the precise prediction $\bar{\phi}^\star$ as opposed to new evidence.

cor[Risk under Bayesian audience] Suppose Assumption (ref) hold and consider a prior as in Equation (ref), with corresponding posterior expectation in Equation (ref). Then $$ -W_\phi(\bar{\phi}^\star;\sigma, \pi^\star) = \lim_{\eta \rightarrow \infty} \sum_x p(x) \mathbb{E}\left[\Big(\phi(x) - \mathbb{E}_\eta[\phi(x) | \bar{\phi}^\star, \phi^{new}] \Big)^2 \Big| \phi\right]. $$

Corollary (ref) illustrates the identity between the minimum risk under Assumption (ref) and the risk of a Bayesian audience with an uninformative prior over the basin of ignorance. Here $\eta^2 \rightarrow \infty$ precisely defines ignorance: for some contexts $x$, the (possibly best) predictor $\bar{\phi}^\star$ within the class $\mathcal{F}$ can incur an arbitrary large error.

To compare with standard BHMs, note that the typical BHM takes the form $\hat{\phi}(x) \sim \mathcal{N}(\phi(x), \gamma^2), \phi(x) \sim \mathcal{N}(\bar{\phi}(x), \eta^2), \eta^2 < \infty$, where $\hat{\phi}(x)$ is a pilot and noisy estimate of $\phi(x)$. Here, we think of $\bar{\phi}$ as a simple function, such as a mean after controlling for observable low dimensional covariates or also obtained from mixture models.\footnote{For simplicity, we can treat here $\bar{\phi}$ as known, but in practice that can be replaced by precise estimates as e.g., for Empirical Bayes.} Effectively, the Bayesian model shrinks all observations towards the simple function $\bar{\phi}(x)$. This shrinkage becomes more prevalent as the pilot noise $\gamma^2$ is larger.

Intuitively, the Bayesian hierarchical models does not allow for classifying observations into a basin of ignorance pooling information only outside of it.

This approach (and more broadly BHMs with possibly different parametrizations) makes undesirable assumptions in our context. It forces predictions across units without leaving scope for future experimentation. This differs from our chosen loss function that accounts for the possibility of collecting new evidence $\phi^{new}$. In addition, it may contaminate real, identifiable archetypes with ill-fitting data, by pooling information across sub-populations which may exhibit arbitrary heterogeneity. This amounts of reporting a function $\bar{\phi}$ constructed using information from all (instead of some) $x$. Instead, here we want to learn when (and how) to pool information together, and when instead we should admit ignorance to guide future research.

Estimation using existing evidence

In this section we introduce sampling uncertainty to build our prediction functions $\bar{\phi}$. We construct estimators obtained from a (pilot) study of $n$ individuals. Specifically, researchers observe $n$ individuals organized through discrete set $\mathcal{X}$ possibly growing with $n$ (i.e., $\mathcal{X}$ can be an implicit function of $n$). Each individual $i$ is associated with covariates $X_{i} \in \mathcal{X}$ characterizing their type. Throughout our analysis, we will condition on $X = (X_1, \cdots, X_n)$.

ass[Existing data] Researchers observe for each $x \in \mathcal{X}$, a pair $ \Big(\hat{\phi}(x), \hat{\eta}(x)^2\Big) \sim_{i.n.i.d.} \mathcal{D}_{x}, $ independent across $x$, with $\mathcal{D}_{x}$ possibly unknown, such that $$ \small \begin{aligned} \mathbb{E}[\hat{\phi}(x)] & = \phi(x), \quad \mathbb{E}[\hat{\eta}(x)^2] & = \eta(x)^2, \quad \mathbb{E}[\hat{\phi}(x)^2] - \phi(x)^2 & = \eta(x)^2. \end{aligned} $$ For all $x$, $|\phi(x)| \le K, \eta(x)^2 \le \bar{\eta}^2$ for some possibly unknown constants $K, \bar{\eta}^2 < \infty$.

Assumption (ref) states that for each type $x$, we observe an unbiased (but possibly noisy/inconsistent) estimate of its mean and variance, assuming at least two observations for each value of $x$.

Randomness in $\hat{\phi}(x)$ may be driven by randomness in the sampling and treatment assignment in the experiment. Sampling uncertainty for $\hat{\phi}(x)$ (as in abadie2020sampling) occurs when only a small fraction of individuals with covariates $x$ is observed.

The variance $\eta(x)^2$ is uniformly bounded, ruling out settings where we observe no observation for type $x$. Therefore, our focus here is on studying generalizability between types $x$ for which we have a pilot study. For example $x$ in our application denote individuals with different baseline assets and consumption, marital status, age and education observed in Peru, India, Pakistan, Honduras, Ghana and Ethiopia where a pilot experiment was conducted. We are interested in generalizability between these six countries. Notably, $\hat{\phi}(x)$ does not need to be consistent for $\phi(x)$.

Finally, we assume independence, but this can be relaxed to local dependence by combining the results we derive in the current paper with techniques in viviano2024policy.

For a given prediction $\bar{\phi}(x)$, and $\pi(x)$, we form an estimate for the approximation error

equation[equation omitted — 274 chars of source]

Intuitively, we measure the distance of the estimated property from its prediction and subtracts the (within) variation of the group property. Here, we subtract the estimator's variance to avoid that the quadratic loss would be otherwise biased for its population loss. We construct the empirical reward as

equation[equation omitted — 166 chars of source]

Given a function space $\bar{\phi} \in \mathcal{F}$, we can then form data dependent $\hat{\pi}$ and data-dependent predictions within the basin of ignorance $\hat{\phi}^\star$ by solving $$ \Big(\hat{\pi}, \hat{\phi}^\star\Big) \in \mathrm{arg}\max_{\pi \in \Pi, \bar{\phi} \in \mathcal{F}} \hat{W}(\pi; \sigma, \bar{\phi}). $$

exmpConsider an experiment with randomized independent treatments $D_i \in \{0,1\}$ and outcomes $Y_i= D_i Y_i (1) + (1 - D_i)Y_i(0)$ where $Y(1), Y(0)$ denote potential outcomes. Define $P(D_i = 1|X_{i} = x) = o(x), s(x) = |i:X_{i} = x|$, with $s(x) \ge 2$ and \begin{equation} \begin{aligned} \hat{\phi}(x) = \frac{\sum_{i:X_i = x} \tilde{Y}_i}{s(x)}, \quad \tilde{Y}_i = \frac{D_i Y_i}{o(X_{i})} - \frac{(1 - D_i)Y_i}{1 - o(X_{i})}, \end{aligned} \end{equation} the outcome reweighted by the inverse propensity score. One unbiased estimator of the variance of $\hat{\phi}(x)$ is $\hat{\eta}(x)^2 = \frac{\sum_{i:X_i = x} (\tilde{Y}_i - \hat{\phi}(x))^2}{s(x)(s(x) - 1)}$.\footnote{We discuss alternative estimators in Appendix (ref). } \qed

Generalizability with discrete archetypes

We propose predictions that first group units into subgroups, and then form generalizability-aware predictions for such subgroups.\footnote{These are common prediction functions, see bonhomme2015grouped, wager2018estimation, venkateswaran2024robustly. The focus on these is interpretability and easy of communication, see also Remark (ref).} Our main assumption is that the policy and prediction space have bounded complexity, measured through its VC-dimension.\footnote{The VC dimension denotes the cardinality of the largest set of points that the function can shatter. Intuitively, it defines the largest sample size for which the model specification has enough “degrees of freedom" to perfectly rationalize every possible pattern across those observations -- an intuitive measure of the class’s capacity (and potential to over-fit) in finite samples, standard in the analysis of algorithms devroye2013probabilistic.} We do not require distributional assumptions other than moment conditions.

We start by posing a set of partitions $\mathcal{G}$ of the space $\mathcal{X}$, an input of the researcher. This is the set of partitions that the researcher is willing to report to a policy-maker. Here $\mathcal{G}$ may entail, for example, ruling out partitions that divide the space of observable characteristics discontinuously to enhance interpretability, or other restrictions motivated by economic theories. We group individuals into (at most) $G$ groups, so that we obtain functions $\alpha: \mathcal{X} \mapsto \mathbb{R}, \alpha \in \mathcal{G}$, and define $$

aligned\alpha(x) \in \{1, \cdots, G\}, \quad \alpha \in \mathcal{G}.

$$ Here, the function $\alpha(x)$ defines the group or partition assigned to $x$. Without loss, we let the first group correspond to the basin of ignorance, so that

equation[equation omitted — 164 chars of source]
ass[Grouping function] Suppose that $\Pi$ is as in Equation (ref) and $\alpha \in \mathcal{G}$ is a given set of possible partitions of $\mathcal{X}$, with \begin{itemize} • Each $\alpha(x), \alpha \in \mathcal{G}$ takes (at most) $G$ possible different values; • $\Pi$ has a bounded VC-dimension $\mathrm{VC}(\Pi) < \infty$; • For each $\alpha \in \mathcal{G}$, for each $g > 1$, $\sum_{x \in \mathcal{X}} 1\{\alpha(x) = g\}$ either equals to zero or is greater than $\underline{\kappa} |\mathcal{X}|$, for some constant $\underline{\kappa} > 0$. \end{itemize} For a given partition $\alpha$, consider predictions $$ \small \begin{aligned} \bar{\phi} \in \mathcal{F}_\alpha, \quad \mathcal{F}_\alpha = \Big\{\phi: \phi(x) = \phi(x') \text{ if } \alpha(x) = \alpha(x')\Big\}. \end{aligned} $$

Assumption (ref) considers settings where individuals are partitioned into (at most) $G$ groups. The choice of the grouping can be arbitrary, as long as it lies in a pre-specified set $\mathcal{G}$ satisfying conditions (A)-(C). Condition (A) states that there are at most $G$ groups. The restriction on $G$ group is often imposed in practice to enhance interpretability and inherits robustness properties under discrete archetypes venkateswaran2024robustly. Condition (B) requires that the complexity of the basin of ignorance, measured through its VC-dimension, is finite. This is attained by many common partitions. For example, it is attained for trees, maximum score functions zhou2023offline, kitagawa2018should, mbakop2021model, as well as for interval partitions of the real line (and assumed here since $|\mathcal{X}|$ grows with $n$). See Figure (ref) for an example. Finally Condition (C) states each group outside the basin of ignorance (i.e., $\alpha(x) > 1$) must contain sufficiently many units in the population. This restriction is natural, since, for example a group with a single individual could not be defined as part of the generalizable set. These complexity constraints reflect a commitment to interpretability: our bounded complexity class ensures that generalizations arise from tractable and communicable groupings.

For a given $\alpha \in \mathcal{G}$, we construct estimated groups' means in the same group of $x'$

equation[equation omitted — 163 chars of source]

corresponding to the (weighted) sample mean within group $\alpha(x')$. Finally, we estimate $$ \hat{\alpha} \in \mathrm{arg} \max_{\alpha \in \mathcal{G}} \hat{W}(\pi^\alpha;\sigma, \hat{\phi}_\alpha^\star), \quad \hat{\pi}^\star(x) = 1\Big\{\hat{\alpha}(x) \neq 1\Big\}. $$

ass[Moment conditions] Let the following hold \begin{itemize} • Suppose in addition that for all $(x, x')$, and for any constant $u' \in (0,1]$, and possibly unknown constant $M_{u'} < \infty$ $$ \small \begin{aligned} \max\Big\{& \mathbb{E}\Big[\Big|\hat{f}_d(x, x')]\Big|^3 \Big], \mathbb{E}\Big[\Big|\hat{f}_d(x, x')\Big|^{2 - 2u'} \Big] \Big\} \le M_{u'}, \quad d \in \{1,2\} \end{aligned} $$ where $\hat{f}_1(x, x') = \hat{\phi}(x) \hat{\phi}(x') - \mathbb{E}[\hat{\phi}(x) \hat{\phi}(x')]$ and $\hat{f}_2(x, x') = \hat{\eta}(x) \hat{\eta}(x') - \mathbb{E}[\hat{\eta}(x) \hat{\eta}(x')]$. • The covariates' target distribution $p(x)$ satisfies $p(x) \in [\frac{\underline{p}}{|\mathcal{X}|},\frac{\bar{p}}{|\mathcal{X}|}]$ for some $\underline{p} \in (0,1], 1 \le \bar{p} < \infty$. \end{itemize}

Condition (A) is a simple moment condition. It requires that the sixth moments of $\hat{\phi},\hat{\eta}$ are uniformly bounded. This is attained for sub-exponential (and sub-gaussian) random variables. Note that here we do not require that $\hat{\phi}, \hat{\eta}$ concentrate around their mean (they can have a non-vanishing variance), in which case $M$ can be an arbitrary positive constant (e.g., we can take $u' = 1$ and $M$ is a constant larger than one). This is our leading example, as we think of $|\mathcal{X}|$ as high dimensional. However, when these functions concentrate around their expectation, we expect the constant $M$ to be close to zero, and to capture the concentration behavior of such functions. In this case concentration depends through their $2 - 2u'$ moment, where $u'$ is positive but arbitrary small. Condition (B) states that the target distribution over covariates' has sufficiently many individuals for each $x$.

We study the regret of our proposed procedure, a standard notion of optimality in the literature, see manski2004statistical, kitagawa2018should. By Proposition (ref) the regret measures the distance of the risk under our estimator from the smallest researcher's risk for given $\mathcal{G}$, therefore characterizing the performance of our procedure.

thm[Finite sample regret guarantees] Let Assumptions (ref), (ref), (ref) hold. Then for any $u' \in (0,1]$ $$ \mathbb{E}\Big[\max_{\alpha \in \mathcal{G}, \bar{\phi} \in \mathcal{F}_\alpha} W_{\phi}(\pi^\alpha;\sigma, \bar{\phi}) - W_{\phi}(\hat{\pi}^\star; \sigma, \hat{\phi}_{\hat{\alpha}^\star}^\star) \Big| \phi \Big] \le \frac{\bar{C}G}{u'} \sqrt{\frac{(M_{u'} + \bar{\eta}^2) \mathrm{VC}(\Pi)}{|\mathcal{X}|}}, $$ where the expectation is conditional on the true properties $\{\phi(x)\}_{x \in \mathcal{X}}$, $\bar{C}$ is a finite constant such that $\bar{C} \le \frac{c_0 K \bar{p}^2 }{ \delta \underline{p} \underline{\kappa}}$ for a universal constant $c_0<\infty$.
proofSee Appendix (ref).

Theorem (ref) establishes (frequentist) regret guarantees of the proposed plug-in estimator. The guarantees are valid for any $|\mathcal{X}|, n$. It only requires that Assumptions (ref) (our restriction on the class of predictions $\mathcal{G}$) and (ref), (ref) (independence and moment conditions) hold, but no assumptions on the data-generating process or $\phi(x)$.

The regret exhibits a fast rate of convergence that depends on the number of types $|\mathcal{X}|$. The regret also depends on the complexity of the class of predictions, through $\mathrm{VC}(\Pi)$, a measure of complexity of $\mathcal{G}$ as discussed below Assumption (ref), and $G$.

Finally, the regret depends on the large deviations of the estimated reward. Such large deviations are captured through the bounds on the higher-order moments of recentered random variables $\hat{\phi}, \hat{\eta}$ through $M$, and the variance $\bar{\eta}^2$. The constant $\bar{C}$ capture large deviations that mostly depend on overlap restrictions.

Whenever $\hat{\phi}, \hat{\eta}$ have non-vanishing variance, the rate is the minimax rate found in different contexts for policy learning, e.g., kitagawa2018should, athey2021policy, with in our case $|\mathcal{X}|$ in lieu of the sample size. When $\hat{\phi}, \hat{\eta}$ also concentrates at say rate $\bar{n}_{|\mathcal{X}|}^{-1/2}$ each, for some $\bar{n}$, the rate is of order $\frac{1}{\sqrt{|\mathcal{X}| \bar{n}_{|\mathcal{X}|}^{1 - 2u'}}}$.

Notions of generalizability-aware predictions are novel to the literature, and, as a result, the derivations of Theorem (ref) use novel techniques compared to existing literature. The main challenge is to control jointly the estimation error from the group-means and the adversarial error from the class of partitions $\mathcal{G}$ by studying properties of the supremum of an empirical process generated by $\mathcal{G}$.

rem[Larger and growing function class] Our main innovation here is to combine the construction of prediction functions with the task of generalizability. One could consider more general function classes $\mathcal{F}$, such as $ \mathcal{F}_\alpha = \Big\{\phi: \phi(x) = \beta_{\alpha(x)}^\top x\Big\} $ allowing for group-level linear regressions. Or similarly, one could consider a function class $\mathcal{F}$ that does not use discrete partitions. That is, the concept of archetype can be general and allow for more flexible prediction functions. The cost of increasing the complexity lies in higher estimation error and weaker interpretability. Regret bounds in this cases would depend on uniform deviations of the estimated prediction function from its population counterpart. Similarly, one can use a function class whose complexity grows with $n$ (e.g., $G_n$ is a function of $n$). Since our results are finite sample results, these continue to hold as the VC-complexity is indexed by the sample size. \qed

Inference and optimization

Next, we complement our regret guarantees with a theory of inference. Denote $$ \mathcal{G}^\star \subseteq \mathcal{G}, \quad \mathcal{G}^\star = \Big\{\alpha \in \mathcal{G}: \sup_{\alpha' \in \mathcal{G}, \bar{\phi} \in \mathcal{F}_{\alpha'}} W_{\phi}(\pi^{\alpha'}; \sigma, \bar{\phi}) = \sup_{\bar{\phi} \in \mathcal{F}_{\alpha}} W_{\phi}(\pi^{\alpha}; \sigma, \bar{\phi})\Big\}, $$ the set of partitions that achieve the largest reward.

For a given subset of partitions $\mathcal{G'}$, we would like to test the null hypothesis $\mathcal{G}' \subseteq \mathcal{G}^\star$. For instance, $\mathcal{G}'$ may contain partitions that only use some but not all covariates. To answer this question, consider first the simpler problem of testing, for a given partition $\alpha$, $ H_0: \alpha \in \mathcal{G}^\star $ (so that effectively $\mathcal{G}'$ is a singleton). We will return to the case where $\mathcal{G}'$ is not a singleton at the end of the discussion. To do so, take $\hat{\alpha}^o$ an arbitrary partition independent of estimates $\hat{\phi}(x), \hat{\eta}(x)$, estimated out-of-sample.

defn[Out-of-sample partition $\hat{\alpha}^o$] Suppose that for all $x \in \mathcal{X}$, we are given independent copies of $\hat{\phi}(x), \hat{\eta}(x)$, denoted $\hat{\phi}^o(x), \hat{\eta}^o(x)$. Suppose that such copies also satisfy Assumption (ref) with $\hat{\phi}^o(x), \hat{\eta}^o(x)$ in lieu of $\hat{\phi}(x), \hat{\eta}(x)$. Such copies can be constructed using a simple sample splitting technique, for which half of the observations for each $x$ are used to construct $\hat{\phi}(x), \hat{\eta}(x)$ and the other half are used to construct $\hat{\phi}^o(x), \hat{\eta}^o(x)$. Using $\hat{\phi}^o, \hat{\eta}^o$ only, we can construct an (out-of-sample) estimated reward function $\hat{W}^o(\pi^\alpha;\sigma, \hat{\phi}_\alpha^{\star o})$, as for $\hat{W}$ but with $\hat{\phi}^o, \hat{\eta}^o$ in lieu of $\hat{\phi}, \hat{\eta}$ and where $\hat{\phi}_\alpha^{\star o}$ denote the group-means as in Equation (ref) using out-of-sample estimates $\hat{\phi}^o(x)$ in lieu of $\hat{\phi}(x)$. Define $ \hat{\alpha}^o \in \mathrm{arg} \max_{\alpha \in \mathcal{G}} \hat{W}^o(\pi^\alpha;\sigma, \hat{\phi}_\alpha^{\star o}) $ the estimated partition $\hat{\alpha}^o$ out-of-sample. \qed

We then proceed to build a test-statistic using in-sample observations $(\hat{\phi}, \hat{\eta}$ in Assumption (ref)). In particular, for a given partition $\alpha$, we build a test statistic

equation[equation omitted — 251 chars of source]

where $\hat{\phi}_{\hat{\alpha}^o}^\star, \hat{\phi}_\alpha^\star$ denote the estimated means for grouping $\hat{\alpha}^o, \alpha$, respectively as in Equation (ref) (using in-sample units). That is, given the out-of-sample partition $\hat{\alpha}^o$ we then proceed to estimate the reward using in-sample observations. \\ Variance of the test-statistic Before proceeding, define for $\bar{\phi}_{\alpha}^\star(x) = \frac{\sum_{x':\alpha(x) = \alpha(x')} p(x') \phi(x')}{\sum_{x':\alpha(x) = \alpha(x')} p(x')}$,

equation[equation omitted — 503 chars of source]

Appendix Lemma (ref) shows that $v^2$ corresponds to the asymptotic variance of the test-statistic. Because $\mathbb{V}(Y_x|\hat{\alpha}^o)$ is not necessarily identified, we will use an upper bound

equation[equation omitted — 338 chars of source]

which can be consistently estimated using the sample analog.\footnote{Formally, we can consistently estimate $\tilde{v}$ with the estimator

equation[equation omitted — 229 chars of source]

where $Y_{x}(\cdot)$ is as in (ref) with $\bar{\phi}^\star$ replaced by $\hat{\phi}^\star$ in (ref).}

In Appendix Lemma (ref) we show that $\tilde{v}(\alpha, \hat{\alpha}) = \mathcal{O}(1)$ (and therefore also $v(\alpha, \hat{\alpha}) = \mathcal{O}(1)$), i.e., the rate of convergence of $\hat{T}$ is at least of order $1/\sqrt{|\mathcal{X}|}$. \\ Inference We construct a test $ t_\gamma(\alpha) = 1\Big\{ \sqrt{|\mathcal{X}|} \hat{T}_\alpha(\hat{\alpha}^o) > \Phi^{-1}(1 - \gamma) \tilde{v}(\alpha, \hat{\alpha}^o) \Big\} $ an implicit function of $\hat{\alpha}^o$, where $\Phi(\cdot)$ is the Gaussian CDF.

thm[Inference] Let $\hat{\alpha}^o$ be independent of $(\hat{\phi}(x), \hat{\eta}(x)), x \in \mathcal{X}$. Let Assumptions (ref), (ref), (ref) hold. Suppose in addition that $v(\alpha, \hat{\alpha}^o) > l$ for some positive constant $l > 0$ (i.e., it is non-degenerate). Then for any $\alpha \in \mathcal{G}^\star$ (i.e., under $H_0$) $$ \small \begin{aligned} \lim_{|\mathcal{X}| \rightarrow \infty} \mathbb{E}[t_\gamma(\alpha) | \phi] \le \gamma. \end{aligned} $$ In addition, suppose that $\hat{\alpha}^o$ is estimated as the out-of-sample maximizer of $\hat{W}^o$ in Definition (ref) and $\gamma > 0$. Then for any $\alpha$ such that $\sup_{\alpha' \in \mathcal{G}} W_{\phi}(\pi^\star; \bar{\phi}^\star_{\alpha'}) - W_{\phi}(\pi^{\alpha}; \bar{\phi}_\alpha^\star) > J$ for some fixed constant $J > 0$ $$ \small \begin{aligned} \lim_{|\mathcal{X}| \rightarrow \infty} \mathbb{E}[t_\gamma(\alpha)| \phi] = 1. \end{aligned} $$
proofSee Appendix (ref).

Theorem (ref) establishes two results. First, our proposed procedure controls size. Second, our procedure asymptotically discards partitions whose reward is strictly dominated by a positive factor. Here, we condition on $\phi$ to highlight that these are frequentist hypothesis testing guarantees.

The theorem focuses on partitions $(\alpha, \hat{\alpha}^o)$ for which the variance of the test-statistic is non-degenerate, that is $v^2(\alpha, \hat{\alpha}^o)$ is bounded away from zero. This implies that $\alpha$ is different from $\hat{\alpha}^o$, and requires that $\hat{\phi}$ and $\hat{\eta}$ have a variance bounded from below (i.e., we have a finite number of units for each value of $x$). One could consider alternative scenarios where $v^2$ converges to zero at a given rate (e.g., when the size of each group $x$ is also growing), which we omit for brevity.

\paragraph{Estimating sets of partitions} We can directly extend Theorem (ref) to conduct inference on a given subset $\mathcal{G}' \subset \mathcal{G}$. We formally show this in Appendix (ref). The idea is to conduct separate testing on each $\alpha \in \mathcal{G}'$, with an appropriate correction for multiple testing and return a data-dependent set $\hat{\mathcal{G}} \subseteq \mathcal{G}'$. Algorithm (ref) returns an estimated set $\hat{\mathcal{G}}$ that prunes $\mathcal{G}'$ (a given subset of partitions of interest) from those partition that are not in $\mathcal{G}^\star$ with high probabiliy. In Appendix (ref) we show that the estimated set $\hat{\mathcal{G}}$ in Algorithm (ref) contains $\mathcal{G}' \subset \mathcal{G}^\star$ with high probability and asymptotically discards sub-optimal partitions $\alpha \not \in \mathcal{G}^\star$ (under restrictions in Theorem (ref)).

For example, suppose we consider a class of trees $\mathcal{G}'$ that can use all covariates except for the first entry of $x$. Algorithm (ref) can test whether we can find an optimal partition without using such a covariate.

algorithm[algorithm omitted — 2,670 chars of source]
rem[Alternative upper bounds] The upper bound in Equation (ref) is chosen to minimize $ \min_f \sum_{x} p(x)^2 \mathbb{E}\Big[\Big(Y_{x} - f\Big)^2\Big| \hat{\alpha}^o\Big] $ with the minimizer $f^* = \frac{1}{\sum_{x} p(x)^2} \sum_{x} p(x)^2 \mathbb{E}[Y_{x}|\hat{\alpha}^o]$.\footnote{This is a valid upper bound because $\mathbb{E}[(Y_{x} - \mathbb{E}[Y_{x}|\hat{\alpha}^o])^2|\hat{\alpha}^o] \le \mathbb{E}[(Y_{x} - f_{x})^2|\hat{\alpha}^o] $ for any deterministic $f_{x}$, here chosen constant across $(x)$.} One could also choose $f$ more flexibly, for example allowing $f_{\hat{\alpha}^o(x)}$ to be a function of $\hat{\alpha}^o(x)$, so that $f_g^* = \frac{1}{\sum_{x:\hat{\alpha}^o(x) = g} p(x)^2} \sum_{x: \hat{\alpha}^o(x) = g} p(x)^2 \mathbb{E}[Y_{x}|\hat{\alpha}^o]$. As for Equation (ref), this approach also provides us with a (tighter) upper bound.\footnote{Whenever instead we do have access to (asymptotically) independent copies $Y_{x}, Y_{x}^o$ it is possible to estimate consistently $v^2$ instead of relying on an upper bound. In this case, we can form an estimate of $v^2$, by taking (since $\mathbb{E}[Y_{x} Y_{x}^o] = \mathbb{E}[Y_{x}]^2$) $ |\mathcal{X}| \sum_{x} p(x)^2 \Big(\frac{Y_{x}^2 + (Y_{x}^o)^2}{2} - Y_{x} Y_{x}^o\Big). $} \qed

Optimization

In this section, we discuss the implementation of our method focusing on settings where $\mathcal{G}$ denotes a class of trees (with $G$ groups/labels), while deferring formal details (including regret guarantees and computational complexity) to Appendix (ref). Tree-based methods typically satisfy the complexity restriction in Assumption (ref), see zhou2023offline. They inherit an interpretable representation and impose natural constraints.

To map the setting with tree-based method to our framework, suppose we can organize types $x$ into a vector each $\tilde{x} \in \mathbb{R}^{r}$ with $r$ columns (implicit a function of $x$).

defn[$L$-depth tree] A $L$-depth tree is a tree with $L - 1$ layers consisting of branch nodes, and the $L^{th}$ layer with leaf nodes. In each branch node $l$, we consider one variable over which to do a split, denoted as $j(l) \in \{1, \cdots, r\}$ and the value of such a split $b(l)$. Units with $\tilde{x}^{j(l)} < b(l)$ are assigned to left-node of the next leaf, and the units to the right-node. Each node forms a path, with the leaf nodes defining a final grouping of units $x$. We consider at most $S$ possible splits (values of $b(l)$).

Recall that in our notation $\alpha(x) = 1$ denotes the basin of ignorance and $\alpha(x) > 1$ denotes the generalizable set. Within the generalizable set, we can then form at most $G - 1$ partitions. Here, $S$ denotes the number of splits at each node, which is an input of the researcher (e.g., the number of support points of the covariates).

We would like to be flexible in the construction of the basin of ignorance. Intuitively, the units $x, x'$ can be part of the basin of ignorance if they are very different in observables $x, x'$. The idea proceeds as follows. We construct a set of trees of depth at most $L$. Each leaf node in each tree can (i) either be part of the basin of ignorance, i.e., $\alpha(x) = 1$, or (ii) be an archetype, i.e., $\alpha(x) = g > 1$. This implies that we can be flexible in how to construct the basin of ignorance where two groups of observations, even with different $x, x'$ can be part of it. The depth $L$ controls with how much “granularity" we are willing to detect units in the basin of ignorance. Higher depth implies that we are able to form the basin of ignorance as the union of very small groups of units. The lower depth improves the interpretability in the construction of the basin of ignorance. (See Remark (ref) for settings where researchers may be more agnostic about $L$.) An illustration is provided in Figure (ref).

defn[Partition $\alpha \in \mathcal{G}$ through trees] The partition consists of a depth $L$ tree. Each leaf node is either assigned a label of one or zero. If it is assigned a one then this implies that $\alpha(x) = 1$ for each element in the leaf node (i.e., $(x)$ is in the basin of ignorance). If it is assigned a zero, then this implies that $\alpha(x) > 1$. The leaf nodes for which $\alpha(x) > 1$, each is assigned to a different archetype $\alpha(x) = g > 1$, with at most $G - 1$ many archetypes.

For any tree of depth $L = \log_2(G - 1)$, the number of archetypes is at most $G - 1$. For any tree with $L > \log_2(G - 1)$, only $G - 1$ of the leaf nodes can be archetypes, and the remaining ones must be part of the basin of ignorance.

Consider first the case where $L \le \log_2(G - 1)$. The exact solution to this problem is provided in Algorithm (ref) (Appendix (ref)): after growing a tree of depth $L$, in each final branch of the tree, it searches for the split (variable and value of such a variable) that maximizes reward within that branch. It then proceeds recursively.\footnote{For instance, consider a depth $L = 1$ tree. Then the algorithm runs over all combinations of variables and values, and finds the optimal split. For each (possibly empty) group obtained from this split, it asks separately, whether the reward generated by each group if this group were to form an archetype exceeds the reward generated by this same group if the group were assigned to the basin of ignorance. If it does, it forms an archetype using such a group, otherwise it assigns the group to the basin of ignorance. It then sums the reward over the two groups and repeat recursively.} The recursive structure makes the algorithm simple to implement. Because the tree can decide at the branch level whether to assign groups of observations to the basin of ignorance or not, its complexity is of order $\mathcal{O}(|\mathcal{X}|^L S^L r^L)$, polynomial in the dimension $r$ and number of observations $|\mathcal{X}|$. This is formalized in Appendix Proposition (ref).

If instead we consider higher-depth trees but a small number of archetypes, so that $G < 2^L + 1$, computations become harder: assigning a branch to the basin of ignorance requires comparing the loss functions across all possible trees. To solve this problem, we propose a greedy Algorithm (ref). The algorithm has the same computational complexity as Algorithm (ref) and returns the optimum up to a known optimization error. This error is informative about its regret guarantees formalized in Appendix (ref).

rem[Algorithms that do not specify $L$] Here, the depth $L$ controls the complexity of the basin of ignorance. It is possible to not specify the depth $L$, and instead specify alternative constraints on the basin of ignorance, as long as these constraints implicitly impose a maximum tree depth $L^*$. In these cases, one could grow run Algorithm (ref) with depth $L^*$ and discard trees that do not meet the given constraints. \qed

Empirical application and numerical studies

In this section, we illustrate the properties of our method by re-analyzing the six experimental evaluations of a multifaceted antipoverty (“Graduation”) program, first described in banerjee2015multifaceted. The core intervention consists of providing a bundle of asset transfer, consumption support, training, and access to financial and health services. The specific implementation was adjusted to each of the six local contexts (Ethiopia, Ghana, Honduras, India, Pakistan, and Peru). The goal is to give poor households the tools to generate a sustained improvement in living standards. Across all six pilot experiments, researchers enrolled 10,495 households spanning more than 500 villages. The randomization was conducted at the individual (household) level for three countries and village level in the remaining three, and approximately half of subjects were randomly assigned to treatment and half to control.

banerjee2015multifaceted conclude that this “big push” program has large and robust impacts after pooling across experimental sites, despite the fact that the experimental sites “span three continents, and different cultures, market access and structures, religions, subsistence activities, and overlap with government safety net programs.” Specifically, they show that the program had positive effects on total consumption, an index measuring food security and an index measuring total assets.

We illustrate the properties of our procedure focusing on individual direct (conditional) treatment effects on these three outcomes one year after the intervention.

This is a natural setting where heterogeneity could matter substantially across a few a priori unknown groups. In particular, some of the literature has pointed out that the efficacy of “big push" as the one in this experiment may crucially depend on whether individuals are facing a poverty trap and can be moved into a new steady state.\footnote{See balboni2022people for related evidence from Bangladesh.} To do so, not only individuals need to be sufficiently poor, but also the treatment needs to be sufficiently effective to move individuals out of the poverty trap. The efficacy of treatment can interact with individual and environmental characteristics.

We standardize the outcomes to have variance one as in banerjee2015six. We use as covariates $x$ the country (experiment), baseline outcomes (total consumption, the food security and asset index measured at baseline), the total amount of individual loan measured at baseline and whether other individuals were treated in the same village (to capture heterogeneity due to possible spillovers). Because each observation corresponds to a different value of covariates, we have $|\mathcal{X}| = n$ as we discuss below.\footnote{In Appendix (ref) (Figure (ref)), we also report effects when we consider binary outcomes corresponding of whether the effect is positive.}

To illustrate the properties of generalizability-aware predictions, we estimate the conditional average treatment effects using Generalizability Aware trees (G-Aware for short) with at most four archetypes, and consider different tree structures that allow for more flexibility when detecting the basin of ignorance. We vary the cost claiming ignorance ($\sigma^2$), and illustrate that not allowing for a basin of ignorance may misguide the study of effect heterogeneity. In particular, we show that failing to account for ignorance can form misguided counterfactual predictions not only for those individuals whose effect may not be predictable, but also for the other units in the sample. At the end of this section, we complement our findings with a set of calibrated simulations.

Empirical analysis

\paragraph{Estimation of $\hat{\phi}$ and $\hat{\eta}$} For each individual $i$ in each country we construct an unbiased measure of its conditional average treatment effect using $\hat{\phi}(X_i) = \tilde{Y}_i$ with $\tilde{Y}_i$ in Equation (ref). This corresponds to unit $i$'s individual outcome (in a given country), appropriately weighted by the inverse probabability weights; the propensity score corresponds to the empirical probability of treatment in each country. This allows us to form individual-level $\hat{\phi}(x)$ unbiased for $\phi(x)$ with no assumptions on its heterogeneity structure. Given that each individual has effectively possibly similar but different values of $x$, $|\mathcal{X}|$ corresponds to the overall sample size of about 10,100 observations after removing the few observations for which covariates are missing.

We estimate the variance $\hat{\eta}(x)^2$ via a linear regression with Lasso and cross validation within each country $e$, therefore assuming a sparse variance heteroskedasticity within each country. This approach facilitates our analysis, although other (nonparametric / kernel) estimators for the variance that do not rely on sparsity of the estimators' variance are possible and formally discussed in Appendix (ref).\footnote{In practice, we observe substantial homoskedasticity in the estimated variance and estimates are robust as we directly impose homoskedasticity within each country.}

\paragraph{Estimation of G-Aware Tree (with multiple outcomes)} We estimate the generalizability aware tree with Algorithm (ref). We consider three different outcomes when estimating the G-Aware Tree. With multiple outcomes, the archetype structure (groups) is the same across properties, whereas the predictions are different for each property (as we formalize in Appendix (ref)). We consider two different types of tree: (i) a depth-three tree, where therefore there is flexibility in the construction of the basin of ignorance and archetypes; (ii) a simpler depth-two tree, where each leaf node can identify either an archetype or the basin of ignorance. We find similar results between (i) and (ii) as we further discuss below and in Appendix (ref). We consider as tuning parameters in the HelperTree Algorithm (ref) a minimum number of elements in a leaf node equal to twenty and number of splits at each variable equal to five.

\paragraph{Choice of $\sigma^2$} We estimate a G-Aware tree for different $\sigma^2 \in \{0.5, 1.5, 2.5, \cdots, 5.5\}$. We study the impact of $\sigma^2$ through its impact on the share of observation assigned to the basin of ignorance and the prediction error, both in Figure (ref). Specifically, given the raw prediction error in predicting the outcome,

equation[equation omitted — 164 chars of source]

where $\tilde{Y}_{i}$ is the reweighted outcome as in Equation (ref), we report the average error across the three outcomes of interest.

The share of observations in the basin of ignorance for a depth-three tree varies between about $25\%$ of the overall sample for $\sigma^2 = 0.5$ to $0\%$ for $\sigma^2 = 5.5$, corresponding to a standard regression tree. The error in Equation (ref) is increasing in $\sigma^2$. The G-Aware tree achieves a large (up to more than $30\%$) prediction improvement compared to the tree that does not allow for a basin of ignorance ($\sigma^2 = 5.5$) at the cost of abstaining from making a prediction for at most $30\%$ of the units.

Our preferred specification for a tree with depth tree is $\sigma^2 = 1.5$ as this corresponds to about $25\%$ of individuals (a small but non-negligible number) classified in the basin of ignorance (and for a simpler depth-two tree $\sigma^2 = 2.5$ as we discuss below). In practice, we recommend reporting the results on different values of $\sigma^2$, alongside plots as in Figure (ref) to be able to balance prediction error and ignorance for choosing $\sigma^2$.

\paragraph{Main specification} We first report results for a more complex three ($G = 4$ and depth-three tree). This allows us to study settings where we allow for flexibility in the construction of the basin of ignorance. Of the four archetypes, two of these archetypes have almost identical average predictions of the outcomes and, therefore, are merged into a single archetype (see Appendix Figure (ref) and Appendix (ref) for a more comprehensive discussion and analysis). This suggests that the effective number of low-dimensional archetypes is small (three), whereas the remaining observations are assigned to the basin of ignorance.

Under our preferred specifications for $\sigma^2 = 1.5$ and a depth-three tree, we are unable to say anything for richer individuals (Figure (ref)). However, we observe large positive effects on individuals with the lowest consumption, smaller but economically meaningful effects on individuals with fairly low consumption, and medium levels of assets, close to zero effects for individuals with medium level of baseline consumption. This is illustrated in Figure (ref) where we report the median values of baseline consumption and baseline asset index for each archetype we find.

Figure (ref) reports the composition of each archetype and the basin of ignorance in the different countries. We observe an overarching archetype in all countries except Peru and India, corresponding to positive effects on standardized results on average equal to $9\%$. A second archetype is in Peru (about $30\%$ of observations in Peru), with close to zero effects. A third archetype is in India (about $50\%$ of observations in India) with the largest effects. The size of the basin of ignorance oscillates between $10$ and $50\%$ of observations across the six countries. Heterogeneity by country may be driven by several factors, one of which is the different composition of the archetypes found in different countries. In particular, the individuals of the archetype found in Peru exhibit a higher level of baseline consumption and those selected by the archetype in India have the lowest levels of consumption, as Figure (ref) shows.

In conclusion, most of the units can be grouped in very few (three) archetypes and exhibit substantial homogeneity. On the other hand, this homogeneity fails as we also consider richer individuals at baseline. In particular, units corresponding to those with higher level of consumption and assets cannot sensibly form an archetype.

Some of the individuals in the basin of ignorance include those with the highest consumption and smallest asset stocks. We may expect that only a few individuals may fit this category, and some of them may be recorded in this category because of measurement error (e.g., issues with data entry). Pooling their outcomes with other units may therefore pollute estimation of the underlying model. Our method automatically detects such units from the data and assign them to a basin of ignorance. In doing so, this can be viewed as a way to trim out outliers that would otherwise drive the entire results of an empirical analysis, a common issue in empirical practice broderick2020automatic.

\paragraph{Generalizability with a simpler tree and robustness} To investigate the robustness of our results, we investigate heterogeneity when we consider a simpler tree of depth two and $G = 4$. Given the simpler structure, we allow for a larger $\sigma^2$ (error of the underlying model), choosing $\sigma^2 = 2.5$, although results are qualitatively similar for smaller choices of $\sigma^2$. The simpler tree finds two archetypes and assigns the remaining units into the basin of ignorance. Similar to before, we observe large positive effects on individuals with fairly low consumption, small but positive effects on individuals with fairly low assets, and medium levels of consumption. This is illustrated in Figure (ref), and Appendix Figure (ref).

\paragraph{Comparison with trees without ignorance} How important is to allow for ignorance in this application? We report the estimated tree using the same variables as the G-Aware tree, but forcing $\sigma^2 = \infty$, whose predictions are colored blue in Figures (ref) (and Appendix Figure (ref)) as a function of the baseline index and the consumption level. Once we eliminate the possibility of a basin of ignorance, the estimated tree appears differently. The tree exhibits heterogeneity for individuals with somewhat similar baseline consumption levels (effects ranging between $35\%$ and $8\%$, and oscillate non-monotonically). This may suggest some instability of regression methods that do not account for arbitrary heterogeneity.\footnote{A similar phenomenum is also illustrated in Appendix Figure (ref).} We draw similar conclusions as we consider a simpler tree of depth two in Figure (ref), where lacking a basin of ignorance (panel at the bottom) leads to non-monotonic predictions in baseline asset levels.

To further investigate this point, Figure (ref) collects the prediction errors as in Equation (ref) for four different subgroups of observations below and above the median baseline log-consumption and assets levels. We consider both the Generalizability-Aware Tree of depth tree and the corresponding Tree with no ignorance ($\sigma^2 = \infty$), for which we report both the prediction error on all units in each subgroup, and those units classified by the G-Aware tree into the basin of ignorance. For this figure, trees are estimated via five-fold cross-fitting as described below Figure (ref) to construct valid confidence intervals (with clustered standard errors as in banerjee2015multifaceted).

The right-hand side plot of Figure (ref) shows that units with higher-baseline assets or higher baseline consumption are more likely to be classified into the basin of ignorance. About $10\%$ of units with low consumption but high assets are classified into the basin of ignorance, $20\%$ of those with high consumption but low assets, and $60\%$ with both. This is consistent with our findings above. The left-hand-side plot of Figure (ref) shows that the prediction error on the basin of ignorance can be economically and statistical significantly larger than the prediction error on the remaining units.

figure[figure omitted — 599 chars of source]
figure[figure omitted — 919 chars of source]
figure[figure omitted — 443 chars of source]
figure[figure omitted — 1,491 chars of source]

\paragraph{Prediction error on generalizable set: calibrated numerical studies} To complement our empirical findings, we provide a calibrated numerical study, focusing on the simple tree structure of depth-two and a small basin of ignorance ($\sigma^2 = 2.5$). The tree is in Figure (ref). We show that even when the basin of ignorance accounts for a small portion of observations this may pollute predictions on the generalizable set too. We consider as the target outcome the average outcome of the three outcome measures considered in our main application.

The estimated tree in Figure (ref) has two regions corresponding to the basin of ignorance, one for a small subset of observations in Peru and the other (larger) outside Peru. Effects ouside Peru are assumed to be homogeneous, forcing the basin of ignorance to be part of the first archetype. For these regions, the outcome for each archetype is drawn from a Normal distribution with variance one and centered around the effects estimated by the G-Aware tree.

However, we simulate treatment effects $\phi(x)$ as arbitrary heterogeneous in Peru and drawn from a Cauchy distribution with given scale parameter between $0.1$ and $3$. This setup mimics setting with heterogeneity arising from a small set of observations, corresponding to only $4\%$ of the total sample size. Conditional on the treatment effects, outcomes are drawn from a Gaussian distribution with variance one (therefore $\hat{\phi}(x) | \phi(x)$ is centered around $\phi(x)$ and has finite moments conditional on $\phi(x)$). Treatment assignments are drawn from a Bernoulli distribution, and for simplicity, we impose homoskedasticity of the outcomes' variance.

figure[figure omitted — 993 chars of source]

We compare the performance of the G-Aware tree as we vary $\sigma^2 \in \{0.1, 0.5, 1, 1.5, 2\}$, to the same estimator that forces no basin of ignorance ($\sigma^2 = 100$), a standard regression tree of depth two, Generalized Random Forest with default options from the package of athey2019generalized and two versions of Empirical Bayes procedures. Empirical Bayes first estimates the conditional mean using a standard regression tree. It then assumes that each observation is drawn from a Gaussian distribution centered around the conditional mean predicted by the estimated regression tree. We use two versions of the Empirical Bayes, either by using the correct variance of the outcome, or by using the empirical estimate of the variance.

In the top-panel of Figure (ref) we report the prediction error in logarithmic scale of the best competitor (conditional on $\phi(x)$), the tree without ignorance and the Generalizable aware tree (worst case for $\sigma^2 \le 2$). Each prediciton error is averaged over 100 replications. Importantly, the error reported is over the generalizable set. We report the error as a function of the scale parameter that controls for the degree of heterogeneity in the basin of ignorance. The error is relative to the smallest error of the simple tree. Whenever heterogeneity is small, our method is comparable to those of our competitors. However, as soon as the scale parameter is $0.5$ or larger, our method presents substantial improvements over the predicted set, up-to fifty percent smaller than the best competitor and eighty percent smaller than a simple tree.

Figure (ref) (bottom-panel) illustrates the behavior of the G-Aware tree as we vary $\sigma^2$. Whenever heterogeneity is high, our method immediately detects the basin of ignorance. When, instead, the degree of heterogeneity is small and $\sigma^2$ is also sufficiently small, the procedure collapses to a simple regression tree as we may expect. That is, the G-Aware tree is able to perfectly classify observations in the basin of ignorance, bringing its prediction error close to zero. This is in stark contrast to our competitors that are particularly sensitive to such outliers, even if these only form $4\%$ of the sample.

In summary, even when only $4\%$ of observations may present arbitrary heterogeneity, common estimators may produce large mean-squared errors up to $50\%$ times larger than the proposed procedure on the generalizable set.

figure[figure omitted — 1,431 chars of source]

\paragraph{Final policy implications} This analysis shows that the effects are the largest on individuals with low consumption and assets. Effects instead are ambiguous and possibly arbitrarily heterogeneous on richer individuals. Therefore, a policy-maker interested in expanding the program to the population outside the ultra-poor should collect more data about the efficacy of the program on individuals with higher consumption and assets. This conclusion differs from what we would have concluded ignoring ignorance, which would have claimed large efficacy for ultra-poor individuals as well as for some individuals with higher baseline consumption. Simulation results illustrate the benefits of accounting for the basin of ignorance to improve stability also over units where effects are generalizable.

Discussion and some practical lessons

The growing availability of experiments across different environments (and with heterogeneous individuals) has motivated a large literature on effect heterogeneity. Estimators in this literature typically aim to learn treatment effects by pooling information across individuals through, e.g., shrinkage or sparsity restrictions. This paper instead focuses on the task of learning when (and how) information from different individuals can be pooled together and when it cannot. To that end, we provide a framework to study generalizability and introduce a class of prediction functions that jointly estimate when and how to form predictions across different observable characteristics and environments. We give the researcher the option to admit ignorance at a given (opportunity) cost. We provide a decision-theoretic foundation of this problem, derive strong finite sample regret guarantees, asymptotic theory for inference and discuss numerical properties of the procedure. An application analyzing a multifaceted program by banerjee2015multifaceted illustrates the benefits of our approach.

The results of the paper provide practical guidance for an applied researcher interested in treatment effect heterogeneity within a single study, meta-analysis across studies, and model discovery. We study a regime where researchers do not have strong priors on (i) which covariates matter and, most importantly, (ii) when and whether the set of models posed by the researchers is predictive of treatment effects observed in the data. Therefore, our method can be used both to inform where to collect further evidence (e.g., relevant for meta-analyses) and to detect anomalies in the data, which is relevant to inform model discovery. Our method applies well beyond looking at environment-by-agent characteristic heterogeneity in the sense that one can interpret the environment much more broadly. For instance, it also provides a vocabulary to study heterogeneity in research teams, methods, or implementation features. For example, one could use our method to study when effects observed in field experiments are predictive of similar interventions in lab experiments and vice-versa, relevant in behavioral (and development) economics kagel2020handbook.

We leave the reader with many open directions for future work. First, implementing our method may often require harmonizing both outcomes and covariates across studies, and we need better methods to process the data even if variables collected by different researchers are not directly comparable. Second, if we seek to learn about mechanisms, rather than simply form predictions, the variables predictive of heterogeneity might not be the exact variables that drive the economic phenomena but rather predictive proxies. This opens the questions of how to combine model selection with our current framework, something we discuss further in Appendix (ref), where we introduce generalizability-aware ensamble methods. Third, there are likely deeper implications of our method for how to design future experiments. Specifically, once we learn which observations form the basin of ignorance, there may be ways to prioritize where (and for which units) to run the next experiment. This raises the question of how to combine our method in a dynamic research process, where researchers may sequentially collect data to maximize the production of knowledge, while leveraging techniques for site selection similar to olea2024externally, gechter2024selecting.