Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
71,566 characters · 15 sections · 53 citation commands
Better Measurement or Larger Samples? Data Collection for Policy Learning with Unobserved Heterogeneity
{ \thispagestyle{empty}
}
Governments and institutions increasingly rely on individualized treatment rules to allocate interventions in heterogeneous populations. From targeting cash transfers to assigning job training, the goal is to identify subgroups that benefit most from a given policy, based on observable characteristics. Recent advances in policy learning formalize this task as the problem of estimating assignment rules that maximize expected welfare, using experimental or observational data kitagawa_who_2018,athey_policy_2021.
A large body of empirical and theoretical research highlights that individuals' responses to treatments may depend not only on covariates such as age or income, but also on latent characteristics such as motivation, prior experience, or ability.\footnote{e.g. heckman_policy-relevant_2001, heckman_structural_2005.} In structural econometric settings, these unobservables are often modeled through fixed effects or individual-specific components, which can be estimated under repeated observations or panel structures.\footnote{e.g. Wooldridge_2005, shoser_2020.} Alternatively, applied researchers measure proxies of the unobserved factors and consider treatment effect variation along their values. For example, performance indicators have been used as proxies for workers' skill level to assess the impact of new technologies on workers' productivity;\footnote{e.g. qje_AI.} community ratings and psychometric measures of business skills have been used to target resources to high-growth microentrepreneurs in developing countries.\footnote{e.g. hussam_2022, bryan_2024.}
As a result, a policymaker interested in maximizing social welfare may decide to assign policies based on the estimated values of these relevant latent traits. This decision problem raises two questions.
First, under what conditions is leveraging such a source of information to assign treatments welfare-improving? \\ To shed light on this question, I show that the proxy's measurement error propagates into the decision problem. Therefore, for its inclusion to improve worst-case performance, the variation in treatment effects explained by the underlying latent factor must outweigh (i) the additional estimation error introduced and (ii) the increase in policy space complexity. To study this trade-off formally, I derive rate-sharp regret bounds for rules that ignore unobserved heterogeneity (Covariate-Based rules) and rules that acknowledge its presence by including the estimate, or proxy ($\hat{a}$-Augmented rules). This comparison is delivered by a simple theoretical innovation. I define regret as the expected welfare loss of any estimated rule relative to an oracle that observes the true latent factor. This provides a common benchmark across policy classes and makes the comparison between the two classes of interest meaningful.
Because the proxy's estimation error affects the policy's worst-case performance, the policymaker may consider to invest in its precision, for instance refining measurement, designing incentive-compatible elicitation mechanisms, or collecting richer datasets to train predictive models.\footnote{Examples include (i.) acquiring satellite images at higher resolution satellite_data, or repeating measurement hussam_2022; (ii.) designing a Becker-DeGroot-Marschak meachanism BDM_method; (iii.) collecting data along the long dimension of a panel dataset when estimating $A_i$ with a fixed or random effects model.} However, under a finite budget, such an investment implies a smaller sample size to learn the optimal policy, leading to a higher welfare loss due to the increase in policy space complexity.
This tension raises the second question. How much should the policymaker invest in the proxy's precision relative to sample size to maximize the policy's performance? \\ I study the design of data collections for policy learning when the policymaker faces a fixed budget. I show that when latent heterogeneity in treatment effects and returns to investment in the proxy's precision are sufficiently high, it is optimal to devote resources to the measurement (or estimation) of the latent factor. By contrast, when it is too costly to improve on the proxy's precision, or its relevance is limited, it is optimal to allocate the budget to enlarging the policy-learning sample and to rely on treatment rules based only on standard covariates. I leverage the regret bounds to derive the threshold conditions that separate these cases yielding a sufficient condition for minimax optimal budget allocation.
In line with the econometric literature in policy learning, I adopt the minimax approach to provide theoretical guarantees on regret manski_statistical_2004, and derive optimal data collection plans epanomeritakis2025, breza2025. To provide practical guidance for applied researchers that do not adopt a minimax perspective, I also propose two sample-splitting procedures that can be implemented in a given empirical setting to provide evidence on: (i) the ranking of treatment rules that ignore or incorporate unobserved heterogeneity; and (ii) how to scale up data collections optimally by allocating resources between measuring (or estimating) the proxy and increasing the sample used to learn the policy.
I apply these new procedures to the context studied in hussam_2022. The authors conduct a cash transfer randomized controlled trial in rural India and present a new proxy for micro-entrepreneurs' business skills based on the rankings entrepreneurs give to each other. They call this measure community rankings. The main result they report is that community rankings improve targeting of cash transfers. First, I confirm the original result by showing that it indeed increases average welfare by $3\%$, and reduces the probability of producing welfare losses by a third compared to scaling up the intervention using only covariates. The proxy was based on the average assessment of five separate rankers. This feature allows me to report two other key findings. First, I ignore the collection cost and show that the welfare gain would have been substantially smaller, had the number of rankers been lower while keeping fixed sample size. Second, I pretend the data from the study were used as a pilot to guide bigger data collections and I estimate the optimal allocation of finite budgets between the number of rankers and sample size of the RCT. I show that for limited budgets it is optimal to select two rankers instead of five in favor of sample size.
The rest of the paper is organized as follows. In section (ref), I review the related literature and describe the main contribution of this paper; in section (ref), I introduce the formal setting, definitions, and main assumptions; in section (ref), I derive the regret bounds for Covariate-Based, and $\hat{a}$-Augmented rules; in section (ref), I study the data collection problem; in section (ref), I present the empirical application; section (ref) concludes.
This paper contributes to the literature on policy learning connecting new regret bounds for policy rules that include or ignore unobserved heterogeneity in treatment effects to the design of minimax optimal data collection plans. This connection is made possible by a new definition of regret that fixes as a benchmark for all classes an oracle that directly observes the latent factor and has complete knowledge of the causal structure underlying the data. This simple theoretical innovation allows one to derive non-trivial rate-sharp regret bounds for both classes and to reduce the data collection problem to a tractable budget allocation problem between competing objectives. These theoretical results come with practical, data-driven procedures to rank policy rules that ignore or incorporate unobserved heterogeneity and estimate the optimal allocation of budget between measuring (or estimating) latent factors more accurately and increasing sample size. To the best of my knowledge, this is the first paper that combines (i) the policy learning problem when policy-relevant variables are estimated or observed with error, with (ii) the resulting trade-offs involved in designing data collections.
The problem of learning optimal treatment assignment rules has attracted attention in economics, statistics, and machine learning. Foundational work by manski_statistical_2004 framed the problem of treatment choice as an empirical risk minimization problem, considering regret as a key evaluation metric. kitagawa_who_2018 formalized empirical welfare maximization as a framework for optimizing treatment rules with controlled complexity, deriving minimax regret bounds for policy classes with finite complexity. athey_policy_2021 extended this framework by focusing on observational studies. More recent contributions explore extensions beyond the standard approach: viviano_fair_2024 and kitagawa_equality-minded_2021 formalize notions of fairness and equality in policy learning; viviano_policy_2024 studies treatment assignment under network interference; kitagawa2025 studies the case in which the set of covariates that is relevant in explaining treatment effects heterogeneity is wider than the set used for targeting. One closely related paper is mbakop_model_2021. It proposes the Penalized Welfare Maximization (PWM) framework, which addresses model selection in treatment choice by penalizing policy complexity. The main similarity relates to the formulation of the problem: both papers consider the problem of optimally selecting the set of policy-relevant variables. However, PWM's guarantees would not apply trivially to the context of unobserved heterogeneity, as it does not explicitly consider noise propagation in the decision problem, which is the main focus of the present work. Moreover, they do not frame the data collection problem or study the trade-offs involved in it.
The econometric literature has long recognized that treatment effect heterogeneity often arises from unobserved factors. Seminal work by heckman_policy-relevant_2001,heckman_structural_2005 introduced the concept of essential heterogeneity and the marginal treatment effect (MTE), showing how unobserved traits influence both treatment selection and gains. This framework highlights that ignoring latent heterogeneity can bias causal inference and limit the effectiveness of policy rules. Building on these insights, a large body of work has focused on the identification and estimation of treatment effects under limited exogeneity.\footnote{For instance, abadie_instrumental_2002, and chernozhukov_iv_2005 develop IV-based methods for estimating heterogeneous effects, while frolich_unconditional_2013, and dhaultfoeuille_identification_2015 extend these approaches to continuous treatments and nonparametric settings.} Recent work has begun to explore policy learning under unobserved confounding. kallus_confounding-robust_2018 proposes minimax regret bounds that hedge against hidden bias, while cui_semiparametric_2021 adapts instrumental variables methods to estimate optimal treatment rules. Proximal causal inference approaches tchetgen_introduction_2024 use proxies to adjust for unobserved confounders. This paper takes a different perspective. I show that even when standard identification issues from unobserved heterogeneity, such as differential compliance, selection into treatment assignment, or spillovers, are not present, an important theoretical trade-off emerges from the fact that relevant unobserved traits need to be estimated or measured with error. Finally, none of these papers studies the data collection problem.
This paper is also related to the long standing econometric literature on measurement error BOUND20013705, decon_repeated_2008, BATTISTIN2014707, deaner_2023, rahul_corrupted_data. Foundational surveys such as nonlinear_review_2011 and review_mesurement_error review how mismeasured covariates affect identification and estimation in linear and nonlinear models, and emphasize the role of auxiliary information such as validation data, repeated measurements, and instrumental variables. In particular, the model considered in the present paper can be classified as weakly-classical measurement error model review_mesurement_error as I allow for the proxy to be biased and for the distribution of measurement error to vary with covariates, but I do not allow for it to vary with the true value of the latent factor. The present paper differs from this literature in both object and question. Rather than studying the bias induced by measurement error in the estimation of treatment effects, I take identification of welfare-relevant treatment effects as given and analyze how proxy noise propagates into the policy-learning decision problem. This shifts attention from identification and bias correction to regret bounds, minimax dominance across policy classes, and the resulting trade-off between improving proxy precision and increasing the sample size used to learn the policy.
Finally, this paper contributes to an emerging econometric literature on data collection problems and experimental design manski_2017, gechter2024selectingexperimentalsitesexternal, epanomeritakis2025,breza2025 by formalizing the problem of designing data collection plans tailored to the problem of learning optimal policies when unobserved heterogeneity is policy-relevant.
\paragraph{Data Generating Process.} Consider the random vector $(X_i,A_i)$, with $(x,a) \in \mathcal{X}\times\mathcal{A}$, $\mathcal{X} \subseteq \mathbb{R}^d$, and $\mathcal{A} \subseteq \mathbb{R}$.
Define $\mathcal{D} = \{0,1\}$ a binary treatment and $D_{i} \in \mathcal{D}$ the treatment indicator. Consider the outcome $Y_i$ and denote with $(Y_i(0), Y_i(1)) \sim P_Y$ the potential outcomes in case $D_{i} = 0$ or $1$ respectively. Define the treatment effect $\tau_i:=Y_i(1)- Y_i(0)$. Denote with $Y_i = (1-D_i)\cdot Y_i(0) + D_i\cdot Y_i(1)$ the observed potential outcome. We observe one realization of $(Y_i, X_i,D_i)\sim_{\text{i.i.d.}}P_{Y,X,D}\in \mathcal{P}$ for all $i \in S_n$ where $S_n$ is a random sample of $n$ units.
We also observe a proxy, or estimate of $A_i$, $\hat{A}_i$ that takes values $\hat{a} \in \hat{\mathcal{A}}$. This can be a direct measurement with error, or a data-dependent estimate.
\paragraph{Policy Rules and Policy Classes.} A policy rule is a function that maps a general set of characteristics $Z_i$ into the target set: $G: \mathcal{Z} \rightarrow \{0,1\}$.
I define as Covariate-Based (CB) the rules that consider only the values of observed covariates to identify targets: \( G(x): \mathcal{X}\rightarrow \{0,1\} \), as $a$-Augmented (${a}$-CB for later reference) rules, the rules that also include unobserved variables: \( G(x,a) : \mathcal{X} \times \mathcal{A} \rightarrow \{0,1\} \), and as feasible $a$-Augmented rules ($\hat{a}$-CB for later reference), rules that leverage observed covariates and estimates of unobserved variables: \( G(x,\hat{a}) : \mathcal{X} \times \hat{\mathcal{A}} \rightarrow \{0,1\} \). Let \( \mathcal{G}_x \), \( \mathcal{G}_{x,a} \), and \( \mathcal{G}_{x,\hat{a}} \) denote the respective policy classes defined as collections of rules. I indicate with $\mathcal{G}_z = \{G(z)\}$ the class of policy rules that belong to any of the three types described above. I denote with $v_z = \text{VC}(\mathcal{G}_z)$ the VC-dimension of the class $\mathcal{G}_z$.
We restrict our attention to the classes of parametric policies defined as:
where $\theta \in \Theta_z$ and $s: \mathcal{Z} \rightarrow \mathbb{R}$.
Moreover, define the conditional average treatment effect function:
And the first best rule:
\paragraph{Welfare.} Population welfare is defined as:
The best-in-class rule is defined as the rule that directly maximizes population welfare. Formally,
We cannot solve this problem directly because we observe only a random sample of the population of interest and we lack knowledge of the causal law underlying $(Y_i(0),Y_i(1))$. Therefore, following kitagawa_who_2018, we rely on its empirical analog and estimate the empirical optimal rule:
where $e(Z_i)$ is the propensity score given $Z_i$.
We evaluate the performance of estimated treatment rules in comparison with an oracle that observes both the values of $X_i$ and $A_i$:
The main assumptions can be divided into assumptions on the data generating process (Assumption (ref)), on the generating process of $\hat{A}_i$ (Assumption (ref)), and on the policy space (Assumption (ref)).
Assumption (ref).i implies that both potential outcomes, and thus treatment effects, are uniformly bounded in absolute value by $M$. Boundedness is a standard condition in the statistical learning literature as it enables the use of uniform concentration inequalities hoeffding_probability_1963,van_der_vaart_weak_2023. Assumption (ref).ii characterizes a quasi-experimental environment in which treatment assignment is independent of potential outcomes and $\hat{A}_i$ conditional on observed covariates. Moreover, the potential outcome of each unit $i$ depends only on their own treatment status, and propensity scores are known. Finally, Assumption (ref).iii is standard in the causal inference literature and guarantees that all units have a positive probability of receiving either treatment or control.
Assumption (ref).1 imposes that $\hat{A}_i$ is produced by a measurement with error. In particular, it imposes additive separability between noise and signal. Assumption (ref).2 imposes that the measurement error is random conditional on covariates. As a whole, Assumption (ref) allows $\hat{A}_i$ to be biased and its error's distribution to vary across covariate values, while requiring the measurement error to be independent of the true values, conditional on the covariates. In Appendix (ref) I extend Assumption (ref) for the case where $\hat{A}_i$ is estimated from external data, rather than measured with error.
Assumption (ref).1 restricts the complexity of the policy class by ensuring that it cannot shatter arbitrarily large sets. The use of VC-dimension as a complexity measure in policy learning was introduced in kitagawa_who_2018, and has been widely adopted by the subsequent literature. Assumption (ref).2 requires the policy class to be flexible enough to contain the true CATE function, or the CATE function to be simple enough to be contained inside the policy class. This assumption is the most restrictive in the set considered. Note that it is only needed to simplify the regret bound for covariate based rules which otherwise would carry an additional term that cannot be bounded non-trivially. I defer a more detailed discussion to the results section and Appendix (ref). Assumption (ref).3 rules out degenerate distributions that place all the probability mass close to the region where the score function $s_\theta(X_i,{A}_i)$ is equal to zero. Assumption (ref).4 rules out score functions that are not Lipschitz continuous.
In this section, I present the regret bounds for Covariate-Based (section (ref)), and $\hat{a}$-Augmented (section (ref)) rules. In section (ref) I illustrate the minimax comparison.
The formal proof is reported in Appendix (ref). Theorem (ref) introduces a bound on the regret for CB rules arising from (i) completely ignoring the source of unobserved heterogeneity, (ii) the lack of complete knowledge on the counterfactual outcomes. The bound in Eq. (ref) decomposes regret into a statistical error term diminishing with sample size that equals the bound in Theorem 2.1 kitagawa_who_2018, and an approximation error term due to (i) ignoring unobserved heterogeneity ($\bar{\sigma}_{\tau|x}$, upper bounded by $\sigma_0$) and (ii) considering an assignment rule that is less flexible compared to the CATE ($\Delta(s,\Theta_x)$). Note that, under Assumption (ref).2, this third term equals zero.
The formal proof is reported in Appendix (ref). Theorem (ref) establishes that the regret of Covariate-Based policy rules is bounded below by the sum of a statistical term of order $\sqrt{v^\theta_x/n}$ and an approximation term proportional to the residual variation in treatment effects unexplained by observed covariates, ${\sigma}_0$. Combined with the upper bound in Theorem (ref), this result implies that the regret bound for Covariate-Based rules is minimax sharp up to constants over the class $\mathcal P(\sigma_0)$.
Define the root Mean Squared Errror (rMSE) of $\hat{A}_i$ as:
The formal proof is reported in Appendix (ref). Theorem (ref) introduces a bound on the regret for $\hat{a}$-CB rules arising from (i) not observing the unobserved factor $A_i$, and (ii) the lack of complete knowledge on the counterfactual outcomes. This bound is composed of the bound proposed by kitagawa_who_2018 plus a constant that depends on the class of rules $\mathcal{G}^\theta_{x,\hat{a}}$ through the Lipschitz and margin constants (see Assumption (ref)) times the root MSE of $\hat{A}_i$. The proof is composed of the following steps. First, regret can be decomposed into the sum of the distance between an oracle that observes $A_i$ (full information) and an oracle that observes $\hat{A}_i$ (partial information), and the distance between the latter and the feasible rule. Because of Assumptions (ref), (ref), and (ref).1, this second term can be bounded by the bound in Theorem 2.2 kitagawa_who_2018. Because of Assumption (ref).3, the first term can be bounded by the probability of disagreement between the two oracles scaled by $M$. Because of Assumption (ref).4 such probability can be bounded by a multiple of the expected absolute difference between $\hat{A}_i$ and $A_i$, which in turn can be bounded by the (upper bound of the) rMSE of $\hat{A}_i$.
The formal proof is reported in Appendix (ref). Theorem (ref) shows that the regret of $\hat{a}$-Augmented policy rules is bounded below by the sum of a statistical term of order $\sqrt{v^\theta_{x,\hat{a}}/n}$ and an irreducible estimation-error term proportional to the estimation error in the proxy, $\rho$. Combined with the upper bound in Theorem (ref), this result establishes that the regret bound for $\hat{a}$-Augmented rules is minimax sharp up to constants over the class $\mathcal{P}(\rho)$. In particular, even with infinite data, imperfect observation of the latent factor $A_i$ induces a non-vanishing welfare loss whenever $\rho>0$, reflecting a fundamental limit to the gains from incorporating noisy estimates of unobserved heterogeneity into policy learning.
In Appendix (ref), I extend Assumptions (ref) and (ref) to allow for $\hat{A}_i$ to be produced as a data-dependent estimate. I show that, in case $\hat{A}_i = \hat{f}(X_i)$ where $\hat{f}$ is learned in an independent sample $S_m$ and then applied to $S_n$, the same results in Theorems (ref) and (ref) apply conditional on $S_m$.
Corollary (ref) shows that, when latent heterogeneity in treatment effects exceeds the sum of (i) the increase in policy space complexity due to adding $\hat{A}_i$ to the decision problem, and (ii) the probability of disagreement between the oracles with full and partial information of $A_i$ rescaled by $M$, then, accounting for unobserved heterogeneity through $\hat{A}_i$ when learning the optimal policy minimax-dominates ignoring it. The result follows by combining Theorem (ref) and Theorem (ref) on the common class \(\mathcal P(\sigma_0,\rho)\). Indeed, the lower-bound construction used in the proof of Theorem (ref) can be extended to \(\mathcal P(\sigma_0,\rho)\) by augmenting it with any proxy process satisfying Assumption (ref) and \(\mathrm{rMSE}(\hat A_i)\le \rho\). Since Covariate-Based rules do not depend on \(\hat A_i\), this extension leaves the regret unchanged, and hence
On the other hand, by Theorem (ref), for every \(P\in\mathcal P(\sigma_0,\rho)\),
and therefore
The condition above then implies the claimed minimax ordering.
In this section, I leverage the regret bounds derived in section (ref) to study how a policymaker should design data collection before learning policies. I consider the precision of the proxy $\hat{A}_i$ and the available sample size for learning policies as the outcome of ex ante design choices. On the one hand, collecting richer information on the latent factor, for instance by administering longitudinal surveys, collecting repeated measurements, or increasing training sample size for statistical models can improve the precision of $\hat{A}_i$. On the other hand, these same resources could be used to increase the sample size available for learning policies, for instance by running a larger field experiment, or acquiring a larger observational dataset. This creates a resource allocation problem between two competing objectives: reducing the measurement error in the proxy and reducing the statistical error in the estimated policy.
To formalize this problem, I introduce an information index $t \in \mathcal{T}$ that maps into the rMSE of $\hat{A}_i$. Higher values of $t$ correspond to richer information and therefore to more precise measurements of $A_i$.
Assumption (ref) requires that there exists a function that maps the information index into the rMSE of $\hat{A}_i$ and that such function is non-increasing. Define $\hat{A}_i(t)$ as the measurement or estimate of $A_i$ under information $t$.
I now define the policymaker's design problem. The policymaker jointly chooses the information level $t$ and the sample size $n$ before learning the policy. The objective is to minimize worst-case regret subject to a finite budget. The policy can either ignore the proxy and rely only on observed covariates, or incorporate the proxy estimated at information level $t$.
Definition (ref) makes explicit that the policymaker faces two margins of choice. The first concerns whether to use a proxy for the latent factor at all, and if so with what level of precision. The second concerns how many observations to collect for learning the optimal policy. The budget constraint captures the idea that improving one dimension necessarily crowds out investment in the other.
In general, the results in section (ref) do not rewrite the minimax problem in Definition (ref) exactly. Rather, they provide regret bounds that can be used to derive sufficient conditions for minimax dominance across feasible designs.
Let $n(B_0,t):= \max\{n \in \mathbb N : c(t,n) \leq B_0\}$ denote the feasible sample size at budget level $B_0$ and information index $t$. The feasible sample size under the Covariate-Based design is $n_{CB}=n(B_0, 0)$, while the feasible sample size under the augmented design with information level $t$ is $n_A(t)=n(B_0, t)$. Given the results in Theorems (ref)--(ref), define:
By Theorems (ref) and (ref), $\overline{R}_{CB}(B_0)$ and $\overline{R}_{A}(t,B_0)$ are valid upper bounds on the minimax regret of the feasible Covariate-Based and augmented designs. By Theorems (ref) and (ref), together with the same extension argument used in Corollary (ref), $\underline{R}_{CB}(B_0)$ and $\underline{R}_{A}(t,B_0)$ are valid lower bounds.
Let $V_{CB}(B_0)$ denote the minimax regret of the feasible Covariate-Based design, and let $V_A(t,B_0)$ denote the minimax regret of the feasible augmented design with information level $t$.
Assume the policymaker must have some prior information on the severity of the approximation error incurred by ignoring latent heterogeneity.
Assumption (ref) does not require point identification of the unexplained heterogeneity in treatment effects. It only requires an upper bound that can be interpreted as prior or contextual knowledge about the empirical relevance of latent heterogeneity.
The next proposition uses these bounds to provide sufficient conditions for minimax dominance across feasible designs.
Proposition (ref) provides sufficient conditions under which the regret bounds from Theorems (ref)--(ref) are informative enough to certify minimax dominance across feasible designs. Part 1 delivers a pairwise comparison between the Covariate-Based design and any augmented design indexed by $t$. Part 2 can then be used to rank augmented designs among themselves and identify a minimax-dominant information level whenever the comparison with the Covariate-Based design does not certify dominance in its favor. The formal proof is reported in Appendix (ref).
In this section, I introduce two procedures to (i) rank policy rules that ignore or incorporate unobserved heterogeneity and (ii) estimate the optimal allocation of budget between the tasks of measuring (or estimating) latent factors and estimating policies. I conduct three empirical exercises and deliver new insights on the data from hussam_2022.
The authors study the effect of providing a cash grant to micro-entrepreneurs on their profits with a randomized controlled trial in rural India. They introduce a new proxy to measure entrepreneurs' business skills, an unobserved dimension identified from previous literature as policy-relevant for targeting interventions apt to stimulate economic development. This proxy is based on the ranking that groups of five entrepreneurs give each other across different outcomes. The authors name this proxy community rankings and claim as their main result that it can help target high-growth micro-entrepreneurs.
The study by hussam_2022 provides a good setting of application for two main reasons. First, the applied research question, whether targeting based on a proxy of a policy-relevant unobserved characteristic is welfare-improving, is strongly aligned with the theoretical investigation of the present paper. Second, the way the proxy is measured, through the average of five repeated measurements, allows me to study how the performance of policy recommendations varies with the precision of community rankings, and estimate the optimal allocation of budget between larger experiments and higher number of measurements.
In the first exercise, I confirm qualitatively the main result from hussam_2022 and provide new estimates for the magnitude of the welfare gains. I show that targeting resources along the values of community rankings increases average welfare by $5\%$, and reduces by two thirds the probability of producing welfare losses (harm rate for later reference) as compared to scaling up the intervention by random assignment. This gain reduces to $3\%$ welfare increase and half harm rate reduction, when compared to covariate based rules.
In the second exercise, I leverage the fact that the proxy was based on the average measurement of five separate rankers to show that, keeping sample size fixed, the gains of targeting based on community ranking increase with the number of rankers.
In the third exercise, I impose a budget constraint and estimate the optimal number of measurements and sample sizes for different budgets. As new insights, I show that (i) even for limited budgets, it is never optimal to ignore the heterogeneity induced by business skills; (ii) when budget is limited, it is optimal to collect fewer measurements, in favor of a larger sample size; (iii) for high budgets, it is optimal to collect as many measurements as possible.
The trial was conducted in the city of Amravati, India, between 2016 and 2018. It was designed to assess whether local community members possess predictive information about heterogeneity in entrepreneurial returns and can be useful to improve the targeting of cash grants.
The sample consists of 1,345 micro-entrepreneurs operating informal businesses in retail and services. First, participants were assigned to peer groups of five or six based on geographic proximity. Within these groups, individuals were asked to rank their peers on future business outcomes, including future profits and marginal returns to capital. The main measure of community ranking used in the paper is the average fraction of peers who ranked a given entrepreneur in the top quartile across the different outcomes for which rankings were elicited. One-third of the sample was then randomly assigned to receive an unconditional cash grant of 6,000 INR (roughly \$100).
The available data include a set of characteristics collected at baseline and after the treatment. I consider as outcome variable the profits realized 2 months after the intervention.
In this section, I rank CB and $\hat{a}$-Augmented rules and quantify the welfare gains from incorporating community rankings into treatment assignment.
\paragraph{Policy Classes.} I consider the following Covariate-Based rule:
where $X_{i,1}$ is age and $X_{i,2}$ is education in years. Age and education are both identified as policy-relevant dimensions by hussam_2022 and previous literature. The $\hat{a}$-CB rule is then defined as:
where $\hat{A}_i$ is community ranking.
Finally, I also consider a benchmark random rule $G_\text{rand}$ that assigns the treatment at random. To evaluate the performance of each rule, I first randomly split the sample into an estimating and test set; then, I use the estimating set to estimate the rules $G(X_i)$, $G(X_i,\hat{A}_i)$ that solve the respective maximization problems; finally, I compute the empirical welfare generated by each estimated rule in the test sample. I leverage the randomness of the sample split to recover the distribution of out-of-sample empirical welfare over $B=2000$ different draws of the estimating and test set data. The sample splitting procedure is illustrated in Figure (ref). I illustrate the evaluation algorithm in Algorithm (ref).
The test set empirical welfare of a given rule $G_z$ is computed as:
where $Y_i$ denotes profits the micro-entrepreneur made in the 60 days following the intervention.
In Figure (ref), I report the cumulative distribution of welfare over the different draws of the estimating and test sets. In column 1 of Table (ref), I report the empirical cdf of welfare evaluated at the status quo, the harm rate. It measures the probability that a given rule generates a welfare lower than the status quo. In columns $(2)-(4)$ of Table (ref), I report the average pairwise difference in test welfare between different rules.
First, all non-random rules dominate the random rule. Therefore, if a government were to scale up this intervention, scaling it without targeting would not be optimal. Second, the CB rule is stochastically dominated by augmented rules. Therefore, as claimed in hussam_2022, using community ranking as a targeting variable produces a welfare gain. In particular, $\hat{a}$-CB rules achieve an average welfare $247\$$ ($5\%$) higher than random rules and $182\$$ ($4%$) higher than CB rules. Finally, targeting using community rankings reduces the harm rate by a half compared to random rules and a third compared to CB rules. This means that, had the policymaker scaled up the cash transfer intervention learning the optimal $\hat{a}$-CB rule from a sample of the size of the training set, the probability of that being harmful over the distribution of the estimating sample is reduced by a half (third) compared to random (CB) rules.
In this section, I provide empirical evidence for the theoretical prediction embedded in Theorem (ref): welfare gains from $\hat{a}$-CB rules decrease as proxy noise increases. Recall that community rankings are defined as the average fraction of peers who rank a given entrepreneur in the top quartile, elicited from four or five separate rankers.\footnote{Only 37 entrepreneurs have five rankers.} This feature of the experimental design allows me to vary proxy precision by restricting the number of rankers used to construct it.
Fixing the sample size to the full dataset, I define $\hat{a}_j$ as the community ranking proxy constructed from $j \in \{1,\ldots,5\}$ randomly selected rankers. Higher values of $j$ correspond to more precise measurements of the latent business skill, with $\hat{a}_5$ coinciding with the full proxy analyzed in the previous subsection.\footnote{Refer to Example (ref) for a formal justification.} In Figure (ref), I compare the original measure with $\hat{a}_j$ and show that for $j=5$ the two measures coincide, while as $j$ decreases $\hat{a}_j$ gets scattered around the original proxy.
Table (ref) reports, for each value of $j$, the average welfare gain of the $\hat{a}_j$-CB rule relative to three benchmarks: the status quo, i.e., treating no one (column 1); random assignment (column 2); and the CB rule (column 3). To avoid ranker-specific effects, the average is computed over the sample splits $B=2000$ and over $R=30$ random selections of $j$ rankers.
The welfare gain of $\hat{a}_j$-CB over random assignment and CB rules is positive and increases monotonically in $j$ for $j\in [1,4]$. This pattern mimicks the upper bound in Theorem (ref): as $\operatorname{rMSE}(\hat{A}_j)$ shrinks, the noise-related term in the $\hat{a}$-CB regret bound falls, narrowing the gap relative to the oracle and widening the welfare advantage over rules that ignore latent heterogeneity altogether. One puzzling result is that welfare gains slightly decrease at $j=5$. This pattern may be explained by the lack of statistical power due to the small size of the sample, expecially considering that only 37 entrepreneurs in the sample have 5 separate non-self rankers.
In this section, I consider the problem of estimating the optimal data collection plan. The key design margin in this setting is the number of peer rankings used to construct the proxy for business skill. I use this feature of the data to study how welfare changes as the policymaker trades off the amount of information used to measure the latent trait against the sample size used to learn the optimal policy.
Formally, let $t \in \{0,1,2,3,4,5\}$ denote the number of non-self rankers used to construct the proxy. The case $t=0$ corresponds to a design in which no ranking information is collected and the policymaker relies only on Covariate-Based rules. For $t>0$, I randomly draw $t$ rankers among those available for each entrepreneur, and compute the average ranking across the selected rankers. Higher values of $t$ correspond to richer information and therefore to more precise measurements of the latent trait.
To introduce the budget constraint, suppose that collecting one observation for policy learning costs $c_n$, while each additional ranking used to construct the proxy costs $c_t$ for each unit. Then, for a given budget $B_0$, the feasible sample size satisfies
Therefore, increasing $t$ improves the precision of the proxy but reduces the number of observations that can be used to learn the policy.
I evaluate this trade-off over a grid of budgets $B_0 \in \{600,800,...,2000\}$, setting $c_n=0.75$ and $c_t=0.25$. For each budget and each value of $t$, I estimate the welfare generated by the feasible design using repeated sample splitting. At the beginning of each repetition, I draw a common test sample and a common training pool from the main analysis data. When $t=0$, I draw up to $n(t,B_0)$ observations from the training pool and estimate the Covariate-Based rectangular rule defined above. When $t>0$, I first generate a random proxy $\hat{A}_i(t)$ by selecting $t$ rankers for each entrepreneur. I then draw up to $n(t,B_0)$ observations from the resulting training pool. On this feasible sample, I estimate both the Covariate-Based rule and the augmented rule. As in the ranking exercise, I also consider a benchmark random rule $G_{\text{rand}}$. The algorithm is described formally in Algorithm (ref).
I then evaluate the out-of-sample welfare generated by each estimated rule in the corresponding test sample. Repeating this procedure over $B=200$ sample splits and, for each $t>0$, over $R=30$ random realizations of the proxy allows one to recover the average welfare associated with each feasible design. Finally, within each budget level, I define the optimal design as the one that yields the highest average welfare. This procedure allows me to trace the welfare frontier over feasible designs and to estimate how the optimal allocation between the number of measurements $t$ and the policy-learning sample size $n$ changes with the available budget.
Table (ref) and Figure (ref) report the main findings. Three results stand out.
First, the optimal design always includes proxy measurements: $t^* \geq 2$ for every budget level considered. Even at the tightest budget (\$600), allocating resources to community rankings, despite reducing the policy-learning sample from 794 to 480 observations, yields a welfare gain of $\$100$ $(+2\%)$ over the CB rule computed with maximum sample. Ignoring latent heterogeneity is suboptimal even when measurement is costly.
Second, the optimal number of rankers increases with the budget. At low budgets $(\$600-\$800)$, $t^*=2$: the marginal cost of additional rankers crowds out too much sample size, so fewer measurements and a larger sample are preferred. From $\$1,000$ onwards, the constraint relaxes and $t^*=4$ becomes optimal, combining higher proxy precision with a feasible sample size.
Third, the design saturates: from $\$1,400$, $n^*$ reaches the sample cap (793 observations) and further budget increases yield no additional welfare gain. The welfare frontier flattens, with gains stabilizing at $+\$179$ $(+4\%)$ over the CB benchmark.
In Figures (ref) and (ref) we report the results for different cost functions. In Figure (ref) we consider the case where the cost of collecting one more measurement is higher than the cost of collecting one more experimental unit. In this case, it is optimal to colelct two measurements, for any budget. In Figure (ref) we consider the case where the two marginal costs are equal. In this case, the conclusions are closer to the main specification.
Standard policy learning studies the performance of treatment assignment rules based on observable characteristics. A large body of empirical work has established that latent traits, such as ability, motivation, or business skills, are of first-order importance in understanding treatment effect heterogeneity. Incorporating these traits into assignment rules comes with two costs: (i) measurement error propagates into the welfare criterion and (ii) the complexity of the policy class increases.
I study this trade-off formally deriving rate-sharp regret bounds for Covariate-Based and $\hat{a}$-Augmented rules, showing that the proxy's inclusion improves worst-case performance only when the treatment effect variation explained by the latent factor outweighs the combined costs of noise propagation and policy space complexity. A new definition of regret, relative to an oracle that directly observes $A_i$, provides a common benchmark that makes this derivation tractable and the comparison meaningful.
Moreover, I frame the allocation problem between improving measurement precision and enlarging the policy-learning sample. I derive the conditions that separate the two regimes, yielding a sufficient condition for the minimax optimal allocation of resources, and propose sample-splitting procedures to implement these findings empirically.
In an application to hussam_2022, I show that incorporating community rankings improves average welfare by $4\%$ and halves the probability of generating welfare losses relative to Covariate-Based rules. Moreover, I show that ignoring latent heterogeneity is not optimal, even under tight budget constraints, and that the optimal number of rankers increases with the available budget.