Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
150,500 characters · 19 sections · 105 citation commands
Policy design in experiments with unknown interference
\if00 {
} \fi
\if10 { \ \\
\ \\
} \fi
{\it Keywords:} Experimental Design, Spillovers, Welfare Maximization, Causal Inference. \\ {\it JEL Codes:} C31, C54, C90.
\onehalfspacing
One of the goals of a government or NGO is to estimate the welfare-maximizing policy. Network interference is often a challenge: treating an individual may also generate spillovers and affect the design of the optimal policy. For instance, approximately 40% of experimental papers published in the “top-five” economic journals in 2020 mention spillover effects as a possible threat when estimating the effect of the program.\footnote{This is based on the authors' calculation. The top-five economic journals are American Economic Review, Econometrica, Journal of Political Economy, Quarterly Journal of Economics, Review of Economic Studies.} Since budget constraints often bind, researchers have become increasingly interested in experimental designs for choosing the treatment rule (policy) that maximizes welfare. However, when it comes to experiments with spillovers, standard approaches are geared towards the estimation of treatment effects. Estimation of treatment effects, on its own, is not sufficient for welfare maximization.\footnote{Examples of treatment effects are the direct effects of the treatment and the overall effect, i.e., the effect if we treat all individuals, compared with treating none. For welfare maximization, none of these estimands are sufficient. The direct effect ignores spillovers, whereas the optimal rule may only treat some but not all individuals because of treatment costs or constraints.} For example, when designing information campaigns, information may have the largest direct effect on people living in remote areas but generate the smallest spillovers. This trade-off has significant policy implications when treating each individual is costly or infeasible.
This paper studies experimental designs in the presence of interference when the goal is welfare maximization. The main difficulty in these settings is that spillovers can be challenging to measure: when spillovers occur through an unobserved network, for example, collecting network information can be very costly because it may require enumerating all individuals and their connections in the population breza2017using. We, therefore, focus on a setting with limited information on the interference mechanism, formalized by assuming units are organized into a small (finite) number of large clusters, such as schools, districts, or regions, and interact through an unobserved network (in unknown ways) within each cluster. In a development study, we may expect that treatments generate spillovers to those living in the same or nearby villages, but spillovers are negligible between different regions egger2019general.\footnote{A finite number of clusters allows researchers to be agnostic on spillovers between different villages and only requires (approximate) independence between a few regions. Namely, the number of individuals who interact between different regions is “small" relative to the number of individuals in a region leung2023network.} We propose the first experimental design to estimate welfare-maximizing treatment rules in such contexts with unobserved spillovers.
This paper makes two main contributions. First, we introduce a design where researchers randomize treatments and collect outcomes once (single-wave experiment) with two goals in mind: (i) to test whether one or more treatment allocation rules, such as the one currently implemented by the policymaker, maximize welfare; and (ii) to estimate how one can improve welfare with a (small) change to allocation rules. The experimental design is based on a simple idea. With a small number of clusters, we do not have enough information to estimate the welfare-maximizing treatment rule precisely. However, if we take two clusters and assign treatments in each cluster independently with slightly different (locally perturbated) probabilities, we can estimate the marginal effect of a change in the treatment assignment rule, which we refer to as marginal policy effect (MPE). In the development study example above, the MPE defines the marginal effect of treating more people in remote areas, taking spillover effects into account.\footnote{The MPE is the derivative of welfare with respect to the policy's parameters, taking spillovers into account, different from what is known in observational studies as the marginal treatment effect carneiro2010evaluating, which instead depends on the individual selection into treatment mechanism.} Using the MPE, we introduce a practical test for whether a welfare-improving treatment allocation rule exists. The MPE indicates the direction for a welfare improvement, and the test provides evidence on whether conducting additional experiments to estimate a welfare-improving treatment allocation is worthwhile.
Using a small (finite) number of clusters, the experiment pairs clusters and randomizes treatments independently within clusters, with local perturbations to treatment probabilities within each pair. The difference in treatment probabilities balances the bias and variance of a difference-in-differences estimator. We show that the estimator for each pair converges to the marginal effect as the cluster's size increases, and we derive properties for inference with finitely many clusters. Importantly, the experiment separately estimates the direct, spillover and welfare effects -- often of independent interest -- by pooling observations across all pairs.
As a second contribution, we offer an adaptive (i.e., multiple-wave) experiment to estimate welfare-maximizing allocation rules. The goal is to adaptively randomize treatments to estimate the welfare-maximizing policy while improving participants' welfare, desirable in (large-scale) experiments muralidharan2017experimentation. Our design guarantees tight small-sample bounds for both the (i) out-of-sample regret, i.e., the difference between the maximum welfare and the welfare evaluated at the estimated policy deployed on a new population, and the (ii) in-sample regret, i.e., the regret of the experiment participants.
The experimental design groups clusters into pairs, using as many pairs as the number of iterations (or more); every iteration, it randomizes treatments in a cluster and perturbs the treatment probability within each pair; finally, it updates policies sequentially, using the information on the marginal effects from a different pair via gradient descent. Because of repeated sampling, conditional on the past, the estimated marginal effect may present a bias due to serial dependence and interference, different from standard adaptive (batch) experiments. We introduce a novel algorithm that avoids this bias through sequential updates.
We investigate the theoretical properties of the method. A corollary of the small-sample guarantees is that the out-of-sample regret converges at a faster-than-parametric rate in the number of clusters and iterations and, similarly, the in-sample regret. No regret guarantees in previous literature are tailored to unobserved interference. Existing results with $i.i.d.$ data, treating clusters as sampled observations, would instead imply a slower convergence in the number of clusters.\footnote{Here, the average out-of-sample regret converges at a rate $1/T$, where $T$ is the number of iterations and proportional to the number of clusters, and at a rate $\log(T)/T$ for the in-sample regret. For the out-of-sample regret, we derive an exponential rate $\exp(-c_0 T)$, for a positive constant $c_0$ under additional restrictions (see Section (ref)). KitagawaTetenov_EMCA2018, shamir2013complexity establish distribution-free lower bounds of order $1/\sqrt{n}$ for treatment choice and continuous stochastic bandits, respectively. Optimization connects to bandits of flaxman2004online, agarwal2010optimal, which, however, provide slower rates for high-probability bounds (see also Section (ref)). wager2019experimenting provide rates of order $1/T$ for in-sample regret but leverage an explicit model for market interactions with asymptotically independent individuals. Here, we do not impose assumptions on the interference mechanism and consider a different setup with partial interference and finitely many clusters. } We achieve a faster rate by (a) exploiting within-cluster variation in assignments and between clusters' local perturbations; (b) deriving concentration within each cluster; (c) assuming and leveraging decreasing marginal effects of increasing neighbors' treatment probability. Fast convergence rates in the number of (large) clusters are particularly interesting when researchers have limited knowledge about interference and can partition units only into a few (approximately) independent clusters.
What is the benefit (and cost) of designing policies without network data? As an additional contribution, Section (ref) characterizes the welfare value of collecting network data. We consider experiments with network spillovers occurring through a sufficiently dense network, and separable direct and spillover effects. We bound the difference between the maximum welfare achievable for any policy that uses network information and the welfare of the policy that does not use network data. This bound depends only on the direct treatment effect minus the cost of treatment. This can be identified in single-wave experiments without network data and provides novel results to guide practitioners on the value of network data.
We then turn to the implementation of the experimental design. In collaboration with Precision Development (PxD), an NGO providing agronomy advice in developing countries, we implemented a large-scale experiment with over 250,000 farmers to test some of the method's properties with two-wave experiments. The experiment provided geo-localized (county-level) weather forecasts to farmers in rural Pakistan to improve agronomy activities, as farmers often lack geolocalized forecasts (available forecasts are typically at the state instead of the county level). Spillover effects are relevant in this application: in a survey conducted by PxD, $80\%$ of surveyed individuals said they shared weather information with other farmers. The experiment consisted of two consecutive waves. Each wave was designed ex-ante to implement our perturbation design as in Section (ref). We use variation between counties to learn marginal effects and spillovers when treating $50\%$ of the individuals in the first wave and $70\%$ in the second wave (the experiment also included some variants discussed in Section (ref)). Using high-frequency survey data merged with daily weather information, we show that farmers improve their beliefs about one-day ahead weather forecasts, and the program generates spillovers. We observe positive marginal policy effects over the first wave and close to zero marginal effects over the second wave, suggesting that treating $70\%$ suffices to maximize information diffusion. By using information about the marginal effect from our experiment, we can reduce the costs of the program by one million US dollars/year once implemented at scale in Pakistan. We complement our findings with simulations, calibrated to experiments on information cai2015social and cash-transfers alatas2012targeting.
Throughout the text, we assume that the maximum degree of dependence grows at an appropriate slower rate than the cluster size; covariates and potential outcomes are identically distributed between clusters; treatment effects do not carry over in time. In the Appendix, we relax these assumptions and study three extensions: (a) experimental design with a global interference mechanism; (b) matching clusters via distributional embeddings with covariates drawn from cluster-specific distributions; and (c) experimental design with dynamic treatment effects, and propose a novel experimental design in this setting. Practitioners may refer to Section (ref) for more discussion about the applicability of our methods.
We contribute to the literature on single-wave experiments, where existing network experiments include clustered experiments and saturation designs baird2018optimal. References with observed networks include basse2018model, viviano2020experimental among others. For the analysis of the bias of average treatment effect estimators with interference, see also basse2016analyzing, johari2020experimental, and imai2009essential. Additional references are bai2019optimality, tabord2018stratification with $i.i.d.$ data. These authors study experimental designs for inference on treatment effects but not inference on welfare-maximizing policies. Different from the above references, we propose a design to identify the marginal policy effect under interference, used for hypothesis testing and welfare maximization. The focus on marginal policy effects connects to the literature on optimal taxation chetty2009sufficient, which differs from our setting by considering observational studies with independent units.
With multiple-wave experiments, we introduce a framework for adaptive experimentation with unknown interference. We connect to the recent literature on adaptive exploration bubeck2012regret, kasy2019adaptive, and the one on derivative-free stochastic optimization, dating back to kiefer1952stochastic, and flaxman2004online, kleinberg2005nearly, shamir2013complexity, agarwal2010optimal, among others. These references do not study the problem of interference (and inference). Here, we leverage between-cluster perturbations and within-cluster concentration to obtain fast rates of regret in high probability (see Section (ref) for a comprehensive discussion). wager2019experimenting study price estimation in a single market munro2021treatment. They assume infinitely many individuals and an explicit model for market prices under which agents are asymptotically independent. As noted by the authors, the structural assumptions imposed in their paper do not allow for spillovers on a network (i.e., individuals may depend arbitrarily on neighbors' assignments). Our setting differs because individuals are organized into finitely many independent clusters here, where unobserved (network) spillovers may occur. These differences motivate (i) our design, which exploits two-level randomization at the cluster and individual level instead of individual-level randomization, and (ii) cluster-level perturbations. From a theoretical perspective, dependence and repeated sampling induce novel challenges studied in this paper.
We relate to inference under interference and draw from hudgens2008toward for definitions of potential outcomes. aronow2017estimating, manski2013identification, leung2019treatment, goldsmith2013social, li2020random assume an observed network, while vazquez2017identification, ibragimov2010t consider clusters among others. savje2017average study inference of the direct effect only. Discussion about relevant estimands in these frameworks can also be found in more recent work by hu2021average. None of these study welfare maximization or experimental design, different from this paper.
More broadly, we connect to the treatment choice literature on estimation manski2004, KitagawaTetenov_EMCA2018, athey2017efficient, stoye2009minimax, mbakop2016model, kitagawa2020should, sasaki2020welfare, viviano2019policy, and inference andrews2019inference, rai2018statistical, armstrong2015inference, kasy2016partial, hadad2019confidence, hirano2020asymptotic. This literature considers an existing experiment instead of experimental designs, and has not studied policy design with unobserved interference. Here, we leverage an adaptive procedure to maximize out-of-sample and participants' welfare. We broadly relate also to the literature on targeting on networks bloch2019centrality, banerjee2013diffusion, akbarpour2018just, which mainly focuses on particular models of interactions in a single observed network -- different from here, where we leverage clusters' variations; the one on peer-group composition graham2010measuring, the one on inference with externalities bhattacharya2013estimating, and pioneering work on vaccination campaigns manski2010vaccination, manski2017mandating. None of these study experimental designs.
We consider a setting with $K$ clusters, where $K$ is an even number. We assume each cluster has $N$ individuals, whereas the framework directly extends to clusters of different but proportional sizes. Observables and unobservables are jointly independent between clusters but not necessarily within clusters, as often assumed in economic applications abadie2017should. Each cluster $k$ is associated with a vector of outcomes, treatments, and covariates. These are $ Y_{i,t}^{(k)} \in \mathcal{Y}, D_{i,t}^{(k)} \in \{0,1\}, X_i^{(k)} \in \mathcal{X} \subseteq \mathbb{R}^{L}, $ respectively. Here, $(Y_{i,t}^{(k)}, D_{i,t}^{(k)})$ denote the outcome and treatment assignment of individual $i$ at time $t$ in cluster $k$, respectively, $X_i^{(k)}$ are time-invariant (baseline) covariates. For each period $t$, researchers observe a random subsample, $$
$$ where $n$ defines the sample size of observations from each cluster and is proportional to the cluster size for expositional convenience. There are $T$ periods. Although units sampled each period may or may not be the same, with abuse of notation, we index sampled units $i \in \{1, \cdots, n\}$. We denote $Y_{i,t}^{(k)}(\mathbf{d}_1^{(k)}, \cdots, \mathbf{d}_t^{(k)}), \mathbf{d}_s^{(k)} \in \{0,1\}^N, s \le t$ the potential outcome of individual $i$ in cluster $k$ at time $t$, as a function of the treatments of all other units in the same cluster. The definition of potential outcomes implicitly imposes no cross-interference between clusters and no anticipation, standard in the literature \citep[e.g.][]{athey2018design}. We will refer to $Y^{(k)}(\cdot)$ as the \textit{potential} outcome functions of all units in cluster $k$.
Whenever we provide asymptotic analyses, we let $N$ grow through a sequence of data-generating processes and let $K$ be fixed. Here, $n$ is proportional to $N$ for expositional convenience. We take a super-population perspective where potential outcomes $Y^{(k)}(\cdot)$ are random variables. The super-population perspective can also be interpreted as assuming that finite $K$ clusters are drawn from a super-population (see Remark (ref) and Section (ref)).
We focus on a parametric class of policies (treatment rules) indexed by some parameter $\beta$, $$
$$ a map that prescribes the individual treatment probability based on covariates. Here, $\mathcal{B}$ is a compact parameter space, and $\pi(x, \beta)$ is twice differentiable in $\beta$. The experiment assigns treatments independently based on $\pi(\cdot)$, and time/cluster-specific parameters $\beta_{k,t}$. Motivated by empirical practice \citep[e.g.,][]{baird2018optimal}, we focus on two-stage experiments where, given the parameter $\beta_{k,t}$ in cluster $k$ at time $t$, treatments are assigned independently.
Assumption (ref) defines a treatment rule in experiments. Treatments are assigned independently based on covariates and time and cluster-specific parameters $\beta_{k,t}$. The assignment in Assumption (ref) is easy to implement: it can be implemented in an online fashion (i.e., sequentially across units) and does not use information about the outcomes' dependence structure, which justifies its choice; also, it generalizes assignments in saturation designs studied for inference on treatment effects baird2018optimal. An example is treating individuals with equal probability akbarpour2018just, i.e., $ \pi(\cdot; \beta) = \beta \in [0,1]$. We can also target treatments, i.e., $ \pi(x;\beta) = \beta_x, $ indicating the treatment probability for $X_i^{(k)} = x$ (with $\mathcal{X}$ discrete). The parameters $\beta_{k,t}$ must be exogenous with respect to potential outcomes in the same cluster to guarantee unconfoundedness, as in standard RCTs. It is possible to let $\beta_{k,t}$ depend on observable clusters' characteristics as discussed in Appendix (ref), and omitted here for brevity. With an adaptive experiment, Assumption (ref) holds for the design presented in Section (ref) (see Remark (ref)).
We defer to Section (ref), studying more complex assignments with dependent treatments.
Throughout the main text, whenever we write $\pi(\cdot; \beta)$, omitting the subscripts $(k,t)$, we refer to a generic exogenous (i.e., not data dependent) vector of parameters $\beta$. We define $\mathbb{E}_\beta[\cdot]$ as the expectation taken over the distribution of treatments assigned according to $\pi(\cdot; \beta)$.
Assumption (ref) (i) states that effects do not carry over in time, as often assumed in studies on experiments kasy2019adaptive, athey2018design (this is only required for the adaptive but not the single wave experiment). It also states that covariates $X_i$ have the same distribution across clusters. Appendix (ref) presents extensions with dynamics, and Appendix (ref) with covariates drawn from different distributions.
Assumption (ref) (ii) imposes restrictions on the expectation of the (potential) outcome, also integrating over the distribution of the other units' assignments. The first component in Equation (ref) is the conditional expectation given the individual covariates and the parameter $\beta_{k,t}$, unconditional on other units' assignments and unobservables (i.e., potential outcome function). The dependence of $m(\cdot)$ with $\beta_{k,t}$ captures spillover effects because treatments' distribution depends on $\beta_{k,t}$. The second components are separable fixed effects.
Whereas treatment effects may exhibit individual-level heterogeneity (see Example (ref) and discussion therein), treatments do not interact with clusters' fixed effects, imposing homogeneity of treatment effects across different clusters. Homogeneity restrictions across clusters is common in many applications cai2015social, miguel2004worms, crepon2013labor, duflo2023chat.
Assumption (ref) (ii) also states that outcomes depend with at most $\gamma_N$ many other outcomes in the same cluster (conditional on the assignment mechanism $\beta_{k,t}$). Here, $\gamma_N$ provides an interpretable restriction on the dependence structure. As we show in Section (ref), in our leading application of network spillovers, $\gamma_N^{1/2}$ defines the largest number of connections of a given individual and, therefore, restrictions on $\gamma_N$ imposes restrictions on the maximum degree. From a theoretical perspective, we require forms of weak dependence within each cluster, motivated by clusters being large regions; in some cases, we can allow for settings where $\gamma_N$ can grow arbitrarily with $N$, see Remark (ref) for a discussion.
We defer to Section (ref) a discussion on the applicability of our assumptions and to Appendix (ref) numerous extensions, including settings with observed heterogeneity.
Welfare defines the expected outcome had treatments been assigned with policy $\pi(\cdot, \beta)$. We do not include fixed effects in the definition of welfare without loss, since such effects are separable. The expectation is taken over treatment assignments, covariates, and potential outcomes. We interpret $y(x,\beta)$, the outcome net of costs and incorporate the costs in the outcome function, as often assumed KitagawaTetenov_EMCA2018. We define the welfare-optimal policy and the marginal effect (under differentiability in Assumption (ref))
The marginal effect defines the derivative of the welfare with respect to the vector of parameters $\beta$. Finally, we define the direct and marginal spillover effects, respectively as $$
$$ The direct effect denotes the effect of the treatment, keeping constant the neighbors' treatment probability, and the marginal spillover effect $S(\cdot)$, the marginal effect of changing neighbors' treatment probabilities, keeping constant the individual treatment. We can write
The MPE $M(\beta)$ depends on the weighted direct (D) marginal spillover (S) effects. Equation (ref) follows in the spirit of decompositions in hudgens2008toward.\footnote{We also note that in more recent work, hu2021average motivate targeting as causal estimand the average indirect effect, different from $S(\cdot)$ with heterogeneous assignments. graham2010measuring present peer effects' decompositions in the different contexts of peer groups' formation. } As we show, the marginal effect is key to improving (maximizing) welfare with only a few clusters. We conclude with examples of welfare functions in the presence of network spillovers, our leading example, and defer to Section (ref) general models with such spillovers.
Ideally, we would like to leverage variation from a single-wave experiment to estimate treatment rules as in KitagawaTetenov_EMCA2018, athey2017efficient, rai2018statistical. Two constraints here make this infeasible: researchers (i) do not know the spillover mechanism in each cluster (e.g., do not have access to network data in the presence of network spillovers); (ii) researchers only have access to a limited (finite) number of (approximately) independent clusters (e.g., small villages cannot be directly used as clusters because spillovers may also propagate between small villages). Because of (i), we cannot estimate the spillover effects on each individual from a single cluster; because of (ii), we cannot consider each cluster as a sampled observation. Instead, we leverage restrictions on the heterogeneity across clusters and limited dependence within each cluster to show that we can use two clusters to consistently estimate the marginal effect $M(\beta)$, at given $\beta$.
As an illustrative example, consider a policymaker who must allocate treatments to half of the population. Consider two household types, $X_i \in \{0, 1\}$, with $P(X_i = 1) = 1/2$, e.g., those living in urban and more remote areas. The policymaker assigns treatments $D_{i,t} | X_i = x \sim \mathrm{Bern}(\pi(x, \beta) )$, where $\pi(x, \beta) = x \beta + (1 - x) ( 1 -\beta)$ is the treatment probability for $x \in \{0,1\}$ that by construction incorporates the budget constraint. Different treatment probabilities for people in remote areas produce different welfare effects. Figure (ref) presents an illustration calibrated to data from alatas2012targeting, alatas2016network.\footnote{Figure (ref) serves as a simple illustration. We estimate a function heterogeneous in the distance of the household's village from the district's center. We use information from approximately $400$ observations, whose $80\%$ or more neighbors are observed. We let $X_i \in \{0,1\}, X_i = 1$ if the household is farther from the district's center than the median household, and estimate a quadratic model, with treatment denoting a cash transfer and the outcome denoting the individual satisfaction with the program.} Spillovers may exhibit decreasing marginal effects, and assigning all treatments to individuals in remote areas is sub-optimal. In addition, because we do not know the spillover mechanism, using variation from a single-wave experiment is insufficient to estimate $\beta^*$.
Instead, we show that with only two clusters, we can estimate the marginal effect for:
Given the marginal effect, we can present to the policy-maker how we can improve policies through incremental updates to the baseline intervention, only using a few clusters. In addition, we can test whether the line's slope in Figure (ref) is zero (with one or two-sided tests), suggesting evidence of whether the current policy is welfare-optimal. (Note that, as in standard hypothesis testing setups, rejection can be informative, while failure of rejection is informative only with well-powered studies, i.e., sufficiently large clusters' size $n$.)
\paragraph{Single wave} We proceed to construct estimators of the marginal effect. We start from Equation (ref). The direct effect (D) can be identified from a single cluster, taking the difference between treated and untreated outcomes. However, the spillover effect (S) cannot be identified from a single cluster. We instead exploit variations between two clusters. We take two clusters, such as two regions. We collect baseline ($t = 0$) outcomes and covariates; we then randomize treatments with slightly different probabilities between the regions. In the first region, we treat individuals in remote areas ($X_i = 1$) with probability $\beta + \eta_n$. Here, $\eta_n$ is a small deterministic number (local perturbation). The remaining individuals are treated with probability $1 - \beta - \eta_n$. In the second region, we treat individuals in remote areas with probability $\beta - \eta_n$, and the remaining ones with probability $1 - \beta + \eta_n$.
As shown in Figure (ref), we can estimate welfare for two different but similar treatment probabilities; the line's slope between the points is approximately equal to the marginal effect. That is, for a suitable choice of $\eta_n$ (see Theorem (ref)), a consistent marginal effect's estimator is
where $\bar{Y}_t^{(h)}$ is the outcomes' sample average in cluster $h$ at time $t$, $Y_{i,0}$ is the baseline outcome with no experiment in place yet, and $(k, k+1)$ index the two clusters. The above estimator is a difference-in-differences; we subtract baseline outcomes due to fixed effects. In Section (ref), we present a test for $M(\beta) = 0$ using a few clusters' pairs. We also discuss estimation and inference on direct and marginal spillover effects, see Table (ref). A by-product of our design is that it does not require large deviations from the baseline intervention between different regions (large deviations can be expensive or difficult to justify to the general public).
\paragraph{Multi-wave} Using the marginal effect, we then propose and study the following sequential experiment: (1) we pair clusters and organize pairs in a circle as in Figure (ref); (2) every step $t$, we estimate the marginal effect within each pair; (3) using the estimated marginal effect from the subsequent pair on the circle, we update the policy in a given clusters' pair.
The sequential updating rule guarantees that the policy achieves an optimum, either global with a (quasi)concave objective or local optimum otherwise. Step (3) is key to overcoming a bias that, as we show in Section (ref), would otherwise arise here due to repeated sampling, while it maximizes the number of clusters that we can use in the experiment.
We measure the method's performance based on the out-of-sample and in-sample regret, respectively defined for an estimated policy $\hat{\beta}$ and sequence of policies $\{\beta_{k,t}\}_{k=1, t=1}^{K, T}$ in the experiment, $ W(\beta^*) - W(\hat{\beta})$, and $\max_{k \in \{1, \cdots, K\}} \frac{1}{T} \sum_{t=1}^T \Big[W(\beta^*) - W(\beta_{k,t})\Big]. $
We pause here to discuss our main assumptions and their applicability.
Our approach leverages two main assumptions: (i) treatments do not generate heterogeneous effects in expectation across clusters; (ii) outcomes have limited (weak) dependence within each cluster. Here, (i) guarantees that potential outcomes' expectations are comparable across different clusters. Condition (ii) guarantees that we can estimate marginal effects even with only two clusters. With unobserved heterogeneity and/or arbitrary dependence, we could not learn optimal policies (and marginal effects) with a few clusters.
The potential outcome model is consistent with models used in many applications, such as spillovers for agronomy advice duflo2023chat, and others cai2015social, miguel2004worms, crepon2013labor. All of these papers consider specifications with homogeneous effects across clusters. Researchers may test for homogeneity by comparing the average baseline covariates across different clusters. An example is in Table (ref), where we show substantial homogeneity in our empirical application. In the presence of heterogeneity, however, we recommend appropriately balancing clusters (see Appendix (ref)).
We impose forms of weak dependence within clusters, mostly (but not necessarily) captured through restrictions on how $\gamma_N$ grows with $N$. This is motivated by clusters being large regions as in our application, where, we may expect, individuals interact only with a subset of individuals in the region de2018identifying. For example, in settings where we observe network data cai2015social, individuals tend to connect with a few individuals within and between villages but not between different regions.
Two additional assumptions we will use with multiple waves of randomization are welfare (quasi)concavity and no carry-over (dynamics) in effects. Examples of concavity are Example (ref), where neighbors' effects induce decreasing marginal effects (see Figure (ref) using data from cai2015social), or settings with negative externalities in Example (ref). Concavity fails when spillovers occur only after “enough" individuals have received the treatment, for which we provide theoretical guarantees in Appendix (ref), under strict-quasi concavity. Under failure of (quasi)concavity, our proposed method will return a local instead of global optimum. See Assumption (ref) and discussion therein.
No carry-overs is a common assumption in (adaptive) experiments kasy2019adaptive, athey2018design, and in applications duflo2023chat, cai2015social. In practice, carryovers do not occur if either each period $t$ is sufficiently far in time from the previous period or if the intervention only has short-term effects on the outcome. We encourage researchers to appropriately choose the time window $t$ and the outcome to guarantee that no dynamics occur. For example, in our application, the treatment (providing weather forecast for the upcoming few days) affects our main target outcome, i.e., a proxy for one-day ahead predictions of weather, but, as we show in Appendix (ref), it does not affect forecasts in the upcoming weeks. See athey2018design for a discussion on carry-overs and Appendix (ref) for an extension with dynamics.
We conclude with micro-foundation of Assumption (ref) in contexts with network spillovers, our leading application. Practitioners may skip this subsection and refer to Section (ref) directly. Suppose individuals are connected with other individuals through an unobserved and cluster-specific adjacency matrix $A^{(k)}$. Individuals can form a link with an (unknown) subset of individuals in each cluster. Nodes in each cluster are spaced under some latent space lubold2020identifying and can interact with at most the $\gamma_N^{1/2}$ closest nodes under the latent space. We say $1\{i_k \leftrightarrow j_k\} = 1$ if individual $i$ can interact with $j$ in cluster $k$. Conditional on $1\{i_k \leftrightarrow j_k\}$,
for an arbitrary and unknown function $l(\cdot)$ and unobservables $U_i^{(k)}$. Whether two individuals interact depends on (i) whether they are close enough within a certain latent space (captured by $1\{i_k \leftrightarrow j_k \}$); (ii) their covariates and unobserved individual heterogeneity (i.e., $X_i, U_i$), which capture homophily. Equation (ref) also states that covariates are $i.i.d.$ unconditionally on $A^{(k)}$, but not necessarily conditionally. Figure (ref) provides an illustration. Here, we condition on the indicators $1\{i_k \leftrightarrow j_k\}$ (which can differ across clusters) to control the network's maximum degree, but we do not condition on the network $A^{(k)}$. We can interpret such indicators as exogenously drawn from some arbitrary distribution.\footnote{Formally, $\mathcal{I}_k \sim \mathcal{P}_k, \quad (X_i^{(k)}, U_i^{(k)}) | \mathcal{I}_k \sim_{i.i.d.} F_{U|X} F_X, \quad A_{i,j}^{(k)} = l\left(X_i^{(k)}, X_j^{(k)}, U_i^{(k)}, U_j^{(k)}\right)1\{i_k \leftrightarrow j_k \}$, where $\mathcal{I}_k$ is the matrix of such indicators in cluster $k$ and $\mathcal{P}_k$ is a cluster-specific distribution left unspecified.} Equation (ref) states that the distribution of covariates and unobservables is the same across different clusters ($F_X, F_{X|U}$ do not depend on the cluster's identity). It implies that the clusters' networks are drawn from the same distribution. We now provide a micro-foundation to our model.
Condition (A) states the following: before being born, each individual may interact with $\gamma_N^{1/2}$ many other individuals (i.e., maximum degree). After birth, the individual's gender, income, and parental status determine her type and the distribution of her and her potential connections' edges.\footnote{See jackson1996strategic, li2020random for pairwise interactions. Extensions where the networks also depend on non-separable shocks $\omega_{i,j}$ are possible, as discussed in previous versions of this draft.} Condition (B) states that potential outcomes depend on neighbors' assignments, observables, and unobservables. Heterogeneity in spillovers occurs arbitrarily through neighbors' observables and unobservables $(D_j, U_j, X_j)$. Such variables can interact with each other, allowing for observed and unobserved heterogeneity in direct and spillover effects (i.e., $r(\cdot)$ is invariant to permutations of the entries of $A_{i, \cdot}^{(k)}$, $r(\cdot)$ is not invariant in neighbors' observables and unobservables). Whereas treatments may exhibit individual-level heterogeneity, treatments do not interact with clusters' fixed effects.
The proof is in Appendix (ref). Proposition (ref) motivates Assumption (ref) in our leading example with network spillovers.
Next, we present the single-wave experiment in Algorithm (ref), a summary of the main estimators, and a brief description of the tests in Table (ref). Define the vector
\paragraph{Algorithm description} Algorithm (ref) presents the design. The algorithm pairs clusters into $G$ pairs. It estimates the marginal effect within each pair by inducing local perturbations $\eta_n$. It then aggregates information across pairs to construct a test statistic.
For the sake of brevity, throughout the main text, we allow for arbitrary pairs in the design of Algorithm (ref). Without loss, we index clusters such that each pair contains two consecutive clusters $\{k, k+1\}$ with $k$ being an odd number. Pairing clusters may occur based on observed heterogeneity, omitted for brevity and formalized in Appendix (ref).
\paragraph{Null hypothesis and inference} Let $\beta^* \in \mathcal{B}$ be an interior point. If $W(\beta) = W(\beta^*)$, then
The above implication is at the core of the proposed approach. We can test whether $p_1$ arbitrary entries of the marginal effect are equal to zero. Rejection implies a lack of global optimality. For expositional convenience, we consider $p_1 = 1$ only as in our application. In Appendix (ref), we show how the proposed method generalizes to $p_1 > 1$. We may also consider one sided tests $M^{(j)}(\beta) \le 0$; for example, for $\pi(x, \beta) = \beta_x$ (with $\mathcal{X}$ discrete), the one-sided test is informative for whether treatment probabilities for individuals with $x = j$ should be increased (without assuming that $\beta^*$ is in the interior). The critical value for the test for $H_0$ in (ref) is obtained by permuting the sign of each pair's estimated marginal effect in the spirit of canay2017randomization, and recomputing the test statistic in Equation (ref) across the different permutations. Corollary (ref) and Appendix (ref) present a formalization.
Finally, we recommend researchers to report $\bar{M}_n(\beta)$ (Equation (ref)) in their results -- the average estimated marginal effect across clusters' pairs. Section (ref) (and Appendix (ref) for $p > 1$) shows that $\bar{M}_n(\beta)$ consistently estimate $M^{(1)}(\beta)$ as $n \rightarrow \infty$, $G$ is finite.
\paragraph{Other effects identified by the experiment} Algorithm (ref) also allows us to estimate the direct effect of the treatment, the (marginal) spillover effect separately, and the welfare respectively under Assumption (ref) below $$
$$ The direct effect is the treatment effect, keeping fixed the neighbors' treatment probability. $S_1(\cdot)$, the spillover effect, is the marginal effect of a small change in the first entry of $\beta$ (e.g., the neighbors' treatment probability), keeping fixed individual treatment status. Our framework also extends to estimating $S_j(\cdot)$ for arbitrary entries of $\beta$ as in Appendix \ref{sec:pilot_general}. For a given pair of clusters $(k, k+1)$, we estimate
The estimator pools observations between the two clusters and takes a difference between treated and control units within each cluster, divided by the probability of treatments as in horvitz1952generalization. We average direct effects across clusters' pairs to obtain a single estimate $\bar{\Delta}_n = \frac{1}{G} \sum_{g} \widehat{\Delta}_g(\beta)$. The indirect effect is estimated as follows:
The estimator takes a weighted difference between the two clusters' control units. Researchers can report the between-pairs average $\bar{S}_n(0, \beta) = \frac{1}{G} \sum_{g} \widehat{S}_g(0, \beta)$ (and similarly $\hat{S}(1, \beta)$ for treated units), which captures spillovers on the control units.
Researchers may also be interested in estimating welfare effects at a given $\beta$, $W(\beta)$, pooling information across clusters, using as an estimator $ \bar{W}_n(\beta) = \frac{1}{K} \sum_{k=1}^K \Big[\bar{Y}_1^{(k)} - \bar{Y}_0^{(k)}\Big]. $
Inference on each of these estimands can be conducted through permutation tests, see Table (ref). Theorem (ref) provides guarantees such that the bias arising from pooling for the direct and welfare effect is negligible for inference.
Next, we discuss the multi-wave experiment. For illustrative purposes, we provide the algorithm for the one-dimensional case $p = 1$, in Algorithm (ref), that is, when $\beta \in \mathcal{B} = [\mathcal{B}_1, \mathcal{B}_2]$ is a scalar. In Remark (ref) and formally in Appendix (ref), we provide the complete algorithm for the $p$-dimensional case. Let $\hat{M}_{k, t}$ be as in Equation (ref) for $k$ odd.
The algorithm pairs clusters (here two consecutive clusters form a pair) and initializes clusters at the same starting value $\beta_0$, $\check{\beta}_1^1 = \cdots = \check{\beta}_K^1 = \beta_0$. At $t = 0$, it randomizes treatments independently using the same starting value $\beta_0$ for all clusters. Here, $\beta_0$ is chosen exogenously, e.g., it is the current policy in place. Over each iteration $t$, we assign treatments based on $\beta_{k,t}$ for cluster $k$ at time $t$, which equals the parameter $\check{\beta}_k^t$ obtained from a previous iteration plus a positive (negative) perturbation $\eta_n$ in the first (second) cluster in a pair. The local perturbation follows similarly to what is discussed in the single-wave experiment. Also, by construction, $\check{\beta}_k^t$ is the same for a given pair $(k, k+1)$, where $k$ is odd. We choose $\check{\beta}_k^{t + 1}$ via sequential cross-fitting: we wrap clusters in a circle and update the parameter in a pair of clusters $(k, k+1)$ using information from the subsequent pair (see Figure (ref)). The algorithm runs over $T$ periods and returns $ \hat{\beta}^* = \frac{1}{K} \sum_{k = 1}^K \check{\beta}_{k}^{T + 1} . $ Choosing the average is motivated by the theoretical properties of gradient descent, although other statistics are also possible.
The proof is in Appendix (ref). Lemma (ref) shows that the parameters used in the experiment are independent of potential outcomes and covariates in the same cluster. Namely, the sequential cross-fitting breaks the dependence due to repeated sampling, which would otherwise confound the experiment. The main distinction from most of the previous literature on adaptive experiments kasy2019adaptive, wager2019experimenting, hadad2019confidence, zhang2020inference is that in all such references repeated sampling does not occur, and batches are independent each period. Here, instead, clusters are dependent over each period, motivating our sequential estimation procedure. Also, note that existing cross-fitting procedures that would instead use all pairs except the current pair of individual $i$ for a policy update would also have a confounding bias whenever $T > 2$ (see Appendix (ref)).
Next, we turn to the theoretical guarantees to study properties of the design. Practitioners only interested in the implementation of the experiment may skip this section.
Assumption (ref) imposes smoothness and boundedness restrictions. These restrictions hold for a large set of linear and non-linear functions, assuming that $\mathcal{X}$ is compact. Boundedness is often imposed in the literature KitagawaTetenov_EMCA2018.
The proof is in Appendix (ref). Theorem (ref) shows one can consistently estimate the marginal effects with two large clusters. Consistency depends on the degree of dependence among potential outcomes (which also depends on neighbors' treatments). Once we interpret $\gamma_N^{1/2}$ as the maximum degree of a network (see Example (ref)), the convergence rate depends on the minimum between the maximum degree of the network, which is proportional to $\gamma_N^{1/2}$, and the covariances among unobservables, captured by $\rho_n$. The theorem also illustrates the trade-off in the choice of the deviation parameter $\eta_n$: a larger parameter $\eta_n$ decreases the variance, but it increases the bias (motivating our rule of thumb in Appendix (ref)).
Assumption (ref) imposes standard moment bounds and a lower bound on the variance of the estimator. In particular, Assumption (ref) states that the variance does not converge to zero at a rate faster than $1/n$. To gain further intuition, note that
Assumption (ref) is stating that $\rho_n \ge 1$, i.e., $\rho_n$ does not converge to zero. This requires that the negative covariance components (if any) do not outweigh the variances in Equation (ref), holding with no or positive correlations and guarantees that the variance is not zero.
The proof is in Appendix (ref). Theorem (ref) guarantees asymptotic normality. The theorem assumes that $\gamma_N$ grows at a slower rate than the sample size of order $N^{1/4}$ (and hence $n^{1/4}$ because $n$ is proportional to $N$). This condition is stronger than what is required for consistency only.\footnote{We conjecture that weaker restrictions on the degree are possible. We leave their study to future research. } Given Theorem (ref), it is possible to conduct inference on the null in Equation (ref) by using either a $t$-student distribution for critical values as in ibragimov2010t (see Theorem (ref)), or using randomization tests in canay2017randomization.
To our knowledge, this set of results is the first for inference on welfare-maximizing policies with unknown interference. We conclude with a study on the estimated direct, spillover, and welfare effects.
The proof is in Appendix (ref). The bias of the estimated direct effect is asymptotically negligible at a rate faster than the parametric rate $n^{-1/2}$ when pooling observations from different clusters. Our main insight here is that, with pairing and perturbations of opposite signs, the first-order bias cancels out. Here, $\eta_n = o(n^{-1/4})$ is consistent with requirements in previous theorems. Given that the bias is asymptotically negligible, we can use existing results for inference on the direct effect savje2017average. For completeness, we show consistency in Corollary (ref) in the Appendix. Inference on the marginal spillover effects follows similarly to inference on the marginal effect, and omitted for brevity.
Next, we derive theoretical properties of the adaptive experiment. Theoretical results are for the general $p$-dimensional case ($p$ is finite). Let $\check{T} = T/p$. We assume the following.
Condition (A) states that unobservables have sub-Gaussian tails (attained by bounded random variables); (B) assumes that the number of clusters is at least twice the number of waves, which guarantees that Lemma (ref) (unconfoundedness) holds.
An example is Example (ref), where neighbors' effects induce decreasing marginal effects, and the treatment may present some costs, see real-world data example in Figure (ref). Strong concavity also arises in linear models with negative externalities, see Example (ref). Assumption (ref) fails when spillovers occur only after that “enough" individuals have received the treatment. To accommodate this setting, we relax Assumption (ref) in Appendix (ref), allowing for a strictly quasi-concave objective that is best suited for these settings. Settings where Assumption (ref) fails are those where also the spillover mechanism (e.g., the network) changes with the intervention, left to future research. In these cases, the proposed method returns a local optimum. When using multiple starting values of our adaptive algorithm, we only require concavity locally to each starting value.
The proof is in Appendix (ref). Theorem (ref) provides a bound on the distance between the estimated policy and the optimal one. The bound depends only on $T$ (and not $n$) because $n$ is assumed to be sufficiently larger than $T$.
The proof is in Appendix (ref). The corollary formalizes the out-of-sample regret bound for $K = 2 (T/p + 1)$. Also, the rate in $K$ does not depend on $p$, as $n \rightarrow \infty$. This is different from grid-search procedures, where the rate in $K$ would be exponentially slower in $p$. Researchers may wonder whether the procedure is “harmless” also on the in-sample units.
The proof is in Appendix (ref). Theorem (ref) guarantees that the cumulative welfare in each cluster $k$, incurred by deploying the current policy $\check{\beta}_k^w$ at wave $w$ (recall that in the general $p$-dimensional case we have $\check{T}$ many waves), converges to the largest achievable welfare at a rate $\log(T)/T$, also for those units participating in the experiment.\footnote{By a first-order Taylor expansion, a corollary is that the bound also holds for $\check{\beta}_k^w \pm \eta_n$ up to an additional factor which scales to zero at rate $\eta_n$ (and therefore negligible under the conditions imposed on $n$).} This result guarantees that the proposed design controls the regret on the experiment participants. This is a useful property that would not be attained, for example, by grid-search procedures for $W(\beta)$ (see Appendix (ref)). We conclude with an exponential convergence rate of the out-of-sample (but not in-sample) regret with a different learning rate.
The proof is in Appendix (ref). The main restriction is that the sample size grows exponentially in the number of iterations (instead of polynomially). The theorem leverages properties of the gradient descent under strong concavity and smoothness bubeck2012regret. Fast rates for the out-of-sample regret are achieved under an appropriate choice of the learning rate that leverages the smoothness of the objective function. The choice of a learning rate invariant in the iteration $t$ requires a sample size exponential in $T$. This differs from the choice of a learning rate as $1/t$ in Theorem (ref), where the adaptive learning rate enables controlling the cumulative error polynomially in $n$. To our knowledge, these regret guarantees are the first under unknown (and partial) interference.
We now contrast the above results with past literature. In the online optimization literature, the rate $1/T$ is common for convex optimization, assuming independent units duchi2018minimax. Here, because of interference, we leverage between-clusters perturbations. Also, we do not have direct access to the gradient, and related optimization procedures are those in the literature on zero-th order optimization kiefer1952stochastic. flaxman2004online, agarwal2010optimal in particular are related to our approach, where regret can converge at rate $O(1/T)$ in expectation only, whereas high-probability bounds are $1/\sqrt{T}$ agarwal2010optimal. Here, we exploit within-cluster concentration and between clusters' variation to control for large deviations of the estimated gradients and obtain faster rates for high-probability bounds. This approach also allows us to extend out-of-sample guarantees beyond global strong concavity (assumed in the above references) in Appendix (ref). In our derivations, the perturbation parameter depends on the sample size, differently from the references above, and the idea of sequential estimation is novel due to repeated sampling. wager2019experimenting derive $1/T$ regret guarantees in the different settings of market pricing, as $n \rightarrow \infty$, with independent units and samples each wave. Our results do not impose independence or modeling assumptions other than partial interference. viviano2019policy considers a single network, with observed neighbors of experiment participants, instead of a sequential experiment. He imposes geometric (VC) restrictions on the policy and solves a mixed-integer linear program. Here, we introduce an adaptive experiment and we do not require network information, using network concentration not studied in previous works.
These differences require a different set of techniques for derivations. The proof of the theorem (i) uses concentration arguments for locally dependent graphs janson2004large; (ii) uses the within-cluster and between-clusters variation for consistent estimation of the marginal effect, together with the cluster pairing; (iii) it uses a recursive argument to bound the cumulative error obtained through the estimation and sequential cross-fitting.
Here, we ask how $\beta^*$ compares with the policy that assigns treatments without restrictions on the policy function, and provide useful bounds on the value of collecting network data. We focus on a setting with network spillovers, where $A$ denotes the unobserved adjacency matrix as in Example (ref), and omit the super-script $k$ because the argument applies to any cluster. We study
with $\mathcal{F}$ as the set of all conditional distribution of the vector $D \in \{0,1\}^N$, given network $A$ and the covariates of all observations $X$ as defined in Section (ref). Equation (ref) denotes the difference between the expected outcomes, evaluated at the global optimum over all possible assignments (with $A, X$ observed), and the welfare evaluated at $\beta^*$ (without observing $A$).
Assumption (ref) states that researchers assign treatments based on finitely many observable types as in manski2004, graham2010measuring. Each type $x \in \mathcal{X}$ is assigned a different probability $\beta_x$, which can take any value between zero and one. Assumption (ref) also states that conditional on individual's type $(X_i, U_i)$, any other unobserved type $U_j$ can form a connection with individual $i$ with some positive probability, provided that $i$ and $j$ are connected under the latent space representation (recall Equation (ref)). This condition is consistent with the model in Example (ref) (and restrictions on $\gamma_N$), because the assumption states that the expected minimum degree is bounded from below by $\underline{\kappa} \gamma_N^{1/2}$, which is smaller than the maximum degree $\gamma_N^{1/2}$. The second restriction is on the potential outcomes. Let
where $0/0 = 0$. Here, $\Delta(\cdot)$ is the direct treatment effect, and $v(\cdot)$ is the cost of the treatment; $s(\cdot)$ captures the spillover effects. Spillovers depend on the fraction of treated neighbors and are heterogeneous in the neighbors' types, with no interactions with direct effects.
The proof is in Appendix (ref). Theorem (ref) bounds the welfare difference by the expected direct effects minus costs. If direct effects are small compared with the treatment costs, such a difference is negligible (for any spillover effects). The bound is identified without network data under separability of direct and spillover effects. The theorem assumes that the maximum degree converges to infinity, but it may converge at a slower rate than $N$, consistent with our conditions in previous theorems. This result is novel in the context of the literature on targeting networked individuals and provides a formal characterization of the value of collecting network information.\footnote{We note akbarpour2018just study network value from the different angle of network diffusion: for a class of network formation models and diffusion mechanisms, the authors show that random seeding is approximately optimal as researchers treat a few more individuals. The main differences are that here (i) we do not study the problem from the perspective of network diffusion but instead focus on an exogenous interference mechanism with heterogeneity; (ii) we provide an upper bound in terms of the direct treatment effect, leveraging a different model and theory. Different from akbarpour2018just, the upper bound does not state that we should treat $\epsilon$-more individuals (since we consider a different model of spillovers). } Theorem (ref) does not state that spillovers are not relevant ($\beta^*$ depends on the spillovers). Instead, it states that one can compute best policies, without knowledge of the network in settings where direct effects are small.
One can estimate the bound by taking an absolute difference between the treated and control units for different individual types, and average across types. In Example (ref), the bound equals $\phi_1$ (the direct treatment effect) minus the cost of implementing the treatment.
Next, we present a large-scale experiment where we implemented our single wave experiment over two consecutive experimentation waves. We use each wave to illustrate properties of the single wave experiment. We also use the second wave to the welfare gains of our experiment. Finally, we present simulations with many waves calibrated to existing experiments.
We now describe the main steps for experiment implementation. See Table (ref) for a summary.
\paragraph{Treatment $D$} The experiment was implemented through Precision Development (PxD), an NGO that provides farmers with phone-based agricultural advisory services. Farmers often lack access to geo-localized weather forecasts, and digital delivery offers solutions to address this challenge fabregas2019realizing. Prior to the experiment, only 45% of cotton growers reported consistent access to weather information, usually via radio or television. About 86% of cotton growers indicated that weather information helps plan agricultural activities (\url{https://precisiondev.org/weather-forecasting-product-for-punjab-pakistan/}). In addition, those farmers with access to weather forecasts only access forecasts produced at the district level, a higher administrative unit that typically includes 3-4 tehsils (tehsils are administrative units equivalent to US counties).
In partnership with a private forecast provider, Precision Development developed calibrated (geo-localized) weather forecast information localized at the tehsil level. The treatment consists of calling farmers to provide weather forecasts via robocalls, meant to improve farmers' ability to take measures in their plots. The experiment was randomized at large scale across approximately 400,000 farmers. We expected the experiment to generate spillovers. In a survey, $80\%$ of the respondents said they actively shared weather information with other farmers, providing suggestive evidence of spillovers.
\paragraph{Target outcome $Y$} We study the effect of the treatment on farmers' ability to predict short-run weather. This is relevant in these applications: correctly predicting weather improves efficiency in the use of resources by, for example, using irrigation or pesticides more efficiently and better invest, see burlig2024long.
As our main data source, we use repeated high-frequency (daily) cross-sectional survey data collected from June to October 2022. To measure farmers' weather forecasts, we ask farmers: “What do you expect will be the maximum (minimum) temperature in your area tomorrow?". We merge this information with PxD forecast weather the day after the survey interview with the specific farmer. We measure the absolute difference between the farmer's predicted maximum (and minimum) temperature and those predicted by PxD forecasts. To combine beliefs about maximum and minimum temperature, we construct a statistical index as described in viviano2021should which serves as our main outcome.
Temperature variables define incorrect beliefs, i.e., negative treatment effects indicate when that farmer's prediction is closer to the PxD forecast or actual temperature. Predicted temperature is a convenient proxy for farmers' one-day ahead weather perceptions since (i) it is less volatile than precipitation grenci2001world; (ii) it does not exhibit effect dynamics/time heterogeneity (see Appendix (ref)); (iii) it is relatively stable within a tehsil. Also, PxD forecast is a good proxy for real temperature. Table (ref) below shows that PxD forecasts and real weather are strongly (and statistically significant) positively correlated.
The survey was run over approximately $6,000$ farmers, stratified across tehsils and individual treatment status, of which we have approximately $1,000$ respondents for our main outcome. We check for balance on take up rates on many dimensions, see Section (ref).
\paragraph{Clusters} In total, 40 tehsils were exposed to experimental variation. Of these, 25 are exposed to our main experiment/design (with in total 287,000 farmers), whereas the remaining 15 are exposed to a different design. Figure (ref) illustrates the region in Pakistan exposed to experimental variation and the sample size within each district (not all tehsils in a district are in the experiment). Tehsils have from 5,000 to 20,000 farmers in the program. We consider a tehsil a cluster. The assumption is that spillovers between different tehsils are negligible, here justified by the fact that tehsils denote large geographic areas, and forecasts are geo-localized at the tehsil level. In contrast to some prior work banerjee2013diffusion, our design allows for spillovers across villages in the same tehsil.
\paragraph{Policy $\beta$} Our policy of interest is choosing how many people to treat. Each treatment costs $0.29\$$ per farmer/year. As shown below, learning whether one can maximize information diffusion without treating all individuals in the population is relevant for decision making once the experiment is implemented at large scale in Pakistan.
\paragraph{Two wave design} We deployed the local perturbation design presented in Section (ref) over two consecutive waves: We induced perturbations around $\beta = 50\%$ in the first wave, where $\beta$ denotes the share of treated individuals, and in the second wave, we induced perturbations around $\beta = 70\%$. The first wave started in April 2022, during which approximately half of the population was exposed to treatment. The second wave started in August 2022 when we increased the total number of treated individuals across all clusters. This increase was planned ex-ante by the NGO's, since the NGO wanted to reach a larger number of treated units by the end of the intervention. To do so, we used a sequential design and induced local perturbation over each wave following our design in Section (ref). This allows us to learn the marginal effects in each wave. In addition, since we find positive marginal effects over forecast accuracy in the first experimentation wave, and close to zero marginal effects in the second wave, the two waves will be helpful to estimate counterfactual welfare benefits of learning marginal effects through a sequential experiment.
\paragraph{Details about first wave and choice of $\eta_n$} The first experimentation wave allows us to learn the marginal effect around $\beta = 50\%$. Over the first experimental wave, we randomly draw a group of twelve tehsils (“Negative Perturbation/Medium Saturation") to have an average treatment probability across tehsils in this group of $\beta = 0.4$, hence inducing a negative perturbation $\eta_n = 10\%$. The choice of the perturbation should depend on power considerations, as, in principle, we may also be interested in more refined marginal effects, at the expense of lower power. To study here trade-offs in the choice of the perturbation parameter $\eta_n$, we select the Negative Perturbation group to have $\beta = 0.4$ on average, with half of the (randomly selected) clusters in the Negative Perturbation group having exactly $\beta = 0.35$ and half of the clusters with $\beta = 0.45$. We repeat the same with a “Positive Perturbation/High Saturation" group with approximately $\beta = 0.6$ on average (and, similarly as before, with six tehsils in this group having $\beta = 0.55$ and seven $\beta = 0.65$). This gives us two (nested) perturbation designs. First, we obtain a better-powered perturbation design (which we refer to as our main design) with a total of 25 clusters and perturbations around $\beta = 0.5$, with perturbations equal to $\eta_n = 10\%$ on average. The second design induces within group perturbation of smaller order $5\%$, which allows us to also learn marginal effects at two more values $\beta \in \{40, 60\}\%$, with half of the clusters. A key intuition is that, by pooling clusters around smaller perturbations, all of our theoretical results directly apply to the main design, up to a small bias (see e.g., Remark (ref)). We report results from the main (better powered) design with $\beta = 50\%, \eta_n = 10\%$ on average; we show that for smaller choice of $\eta_n$ estimates can be under-powered, see Appendix Table (ref). We recommend the choice of two nested designs to avoid under-powered studies.
\paragraph{Details about second wave} The second wave experiment allows us to learn marginal effects at $\beta = 0.7$. Over the second wave (August - October), the “Negative Perturbation" group was exposed to a larger treatment probability $\beta = 0.6$ and the “Positive Perturbation" group was exposed to a treatment probability $\beta = 0.8$. Therefore, over the second experimentation wave, we have two groups with treatment probabilities $\beta = 70\% \pm \eta_n, \eta_n = 10\%$.\footnote{Over the second wave, we also perturbed by 0.05 the probability of treatment for different types of farmers, those below and above the median response rate in the first round, keeping the overall treatment probability constant. This latter perturbation enables estimating heterogeneous treatment effects, omitted from the main analysis for brevity and discussed in Appendix (ref).}
\paragraph{Roadmap of the main design} Our main design identifies the marginal effects, the marginal spillover effects, direct effects, and welfare effects at $\beta = 50\%$ in the first wave and at $\beta = 70\%$ over the second wave. Table (ref) illustrates which effects are identified by the experiment and reports the main robustness checks in the Appendix. It also compares our experiment to a standard saturation experiment that chooses $\beta \in \{50, 70\}\%$ without using local perturbations, and which, therefore identifies a smaller set of parameters.
Next, we study marginal effects on beliefs about PxD forecasts (i.e., whether the farmer's prediction agrees with PxD forecast), illustrating properties of our design on our main outcome (forecast temperature). We assume no cluster fixed effects because of lack of baseline outcomes. Although this is a strong assumption, it is motivated by balance across clusters on pre-treatment observables (Table (ref)). In practice, we recommend to collect baseline outcomes when possible and, as in this case, when infeasible, to check for balance on observable covariates between clusters exposed to different treatment probabilities. Appendix Table (ref) provides results for response rates for which we observe baseline outcomes.
\paragraph{Estimated Marginal Effects} Figure (ref) plots the estimated marginal effects in the main design, i.e., for $\beta \in \{50, 70\}\%$ (with $\eta_n = 10\%$ on average). The figure also reports the estimated welfare at each point $\beta \in \{0.4, 0.6, 0.8\}$. We observe decreasing marginal effects when moving from $\beta = 0.5$ to $\beta = 0.7$.
Table (ref) shows that the marginal effect is large and statistically significant at $\beta = 50\%$, preserve sign but is smaller and non-significant at $\beta = 70\%$. P-values are computed via randomization inference for one sided test, formally described in Appendix (ref). This result is suggestive that treating $50\%$ of the population is sub-optimal, whereas treating $70\%$ of the individuals is close to be optimal. Therefore, our design allows us not only to learn the value of welfare around $\beta \in \{50, 70\}\%$ but also its corresponding marginal effects. Marginal effects can be useful to understand whether we should increase treatment probabilities to improve welfare. Table (ref) reports direct and marginal spillover effects. In particular, we observe marginal effects are mostly driven by large and significant marginal spillover effects at $\beta = 50\%$ (i.e., marginal effects of increasing the friends' treatment probability), whereas marginal spillover effects are close to zero at $\beta = 70\%$.
\paragraph{Welfare gains and welfare comparison with standard saturation design} Using the two experimental waves, we can estimate the welfare improvement of an adaptive experiment that, in the first wave, estimates the marginal effects at $\beta = 50\%$, and in the second wave estimates marginal effects at $\beta = 70\%$. We contrast our design with a typical saturation experiment or grid search method would predict in Figure (ref): a saturation experiment treating $\{0, 50\%, 100\%\}$ sinclair2012detecting of the individuals would not able to identify decreasing marginal effects near $70\%$, and similarly for other choices of treatment probabilities. This is because a standard saturation design would not induce local perturbations. Such a saturation design would recommend all individuals to be assigned to treatment. Our experiment uses information about the marginal effects to identify the optimum near $70\%$. Figure (ref) reports the relative improvement from the one wave experiment to the second wave experiment where $70\%$ of individuals are treated. Increasing number of treated units from $50\%$ to $70\%$ of individuals leads to statistically significant increase in welfare.
A saturation experiment that would recommend treating all individuals would lead to small improvements: We can use as a conservative estimate (upper bound) of welfare at $\beta = 100\%$, its Taylor approximation at $\beta = 70\%$, $W(0.7) + 0.3 M(0.7)$ (this is a conservative estimate because we might most likely expect decreasing marginal effects, as supported by Table (ref)). Despite using a conservative upper bound, predicted improvement when treating all units in the population are small relative to only treating $70\%$ (equal to $8\%$) and non-significant. Treating only $70\%$ of the individuals instead of $100\%$ would save approximately 0.29\$ per farmer/year. This is economically significant if we consider a policy implemented on all farmers in Pakistan (approximately ten millions), saving one million US dollars/year.
We conclude with a brief overview of additional analyses and balance checks in the Appendix.
\paragraph{Balance} We use auxiliary data about farmers' baseline characteristics for all farmers enrolled with PxD in the main experiment (more than 287,000 farmers) to test for homogeneity in covariates between different clusters, a relevant assumption in our framework. Namely, given that our framework requires homogeneity across clusters, we test for homogeneity of covariates between different clusters using information from all individuals in the experiment. Appendix Table (ref) reports the sample means across observable baseline covariates from program administrative data (each covariate is described below Table (ref)). We test for differences in covariates between clusters exposed to different treatment probabilities.\footnote{When estimating marginal effects, it is easy to show that our framework only requires homogeneity restrictions between groups of clusters used to estimate the marginal effects (e.g., the group of clusters in different treatment exposures), but not necessarily between individual clusters having the same exposures.} The relevant null hypothesis is that the expected value of each covariate in Table (ref) in each cluster is the same across all clusters. We construct these tests via randomization inference formally described in Appendix (ref). These tests are informative of whether such groups are comparable and are conducted with a large sample size ($n \approx 10,000$ on average in each tehsil). We observe similar estimates across all covariates. The smallest p-value is $0.21$, the median is above $0.5$, suggesting lack of tehsil-level heterogeneity.
In Appendix Tables (ref), (ref), (ref) we also report balance table on response rates (both among all surveyed individuals and between respondents and non-respondents individuals), where results show substantial balance in relevant baseline characteristics.
\paragraph{Treatment take-up and accuracy} The treatment group received approximately three times more frequent calls than the control group by design -- where the control group's calls were about other activities of the NGO. Appendix Table (ref) shows that the larger number of calls does not negatively affect response rates. Treated individuals present higher (and statistically significant at the $1\%$ level) response rates per call, engaging more with calls.
Table (ref) shows that forecast and real precipitation and temperature are strongly positively correlated, motivating our main focus on farmers' beliefs about PxD forecasts: PxD predicted and real weather follow very similar patterns, but beliefs about PxD forecasts are less noisy.
\paragraph{Parametric regression estimates} Our design allows for standard regression methods. We illustrate this in Tables (ref). Table (ref) reports regression estimates of farmers' incorrect beliefs about temperature and rain with respect to forecast rain from PxD, for which we find mostly significant spillover effects. For parametric regression estimates we can use information from all clusters, including the lower saturation group, after appropriately controlling for the treatment probability, since also in this group treatment are randomized.
\paragraph{Dynamics and additional outcomes} Appendix Table (ref) illustrates lack of dynamics on our primary outcome (temperature forecasts). In Table (ref), we also illustrates effects on other outcomes. We collect information about predicted rain, asking “Do you think it will rain in your area tomorrow?" We use a binary indicator indicating whether the farmers incorrectly predict no rain and, instead, it rains or vice versa (or replies “I do not know"). As shown in Table (ref), we do not consider rain as the main welfare proxy because, different from temperature, this may exhibit treatment effect heterogeneity over time, since the experiment spans seasons of different rain intensity (dry and monsoon seasons). Finally, we use survey information about farming activities to show effects on these in Appendix (ref).
To evaluate the performance of our design with many waves, we calibrate simulations to data from cai2015social and alatas2012targeting, alatas2016network, while making simplifying assumptions whenever necessary. As in our application, we let $\beta$ denote the treatment probability and $\eta_n = 10\%$.\footnote{Here $10\%$ is consistent with the rule of thumb for $\eta_n \approx \sqrt{\sigma^2/c} n^{-1/3}$ (Appendix (ref)), where $\sigma^2$ is the outcomes' variance and $c$ is the objective's curvature, which would prescribe values between $7\%$ and $12\%$ as we vary $n$. In the online supplement, we report results as we vary $\eta_n$ (Figure (ref)).} In the first calibration, the outcome is insurance adoption, and the treatment is whether an individual received an intensive information session. In the second calibration, the treatment is whether a household received a cash transfer, and the outcome is program satisfaction. The experiment of cai2015social contains multiple arms. Here, we only focus on the treatment effects of intensive information sessions, pooling the remaining arms together for simplicity. The experiment of alatas2012targeting contains different arms assigned at the village level, as well as information on cash transfers assigned at the household level. Here, we study the effect of cash transfers only and control for village-level treatments when estimating the parameters of interest.
In each cluster $k$, we generate
where $c$ is the cost of the treatment. We consider two sets of parameters $\Big(\phi_0, \phi_1, \phi_2, \phi_3, \sigma^2\Big)$ calibrated to data from cai2015social and alatas2012targeting, alatas2016network respectively. We obtain information on neighbors' treatment directly from data from cai2015social. For the second application, we merge data from alatas2012targeting, and alatas2016network, and use information from approximately 100 observations whose neighbors' treatments are all observable to estimate the parameters.\footnote{This approach introduces a sampling bias in the estimation procedure, which we ignore for simplicity, given that our goal is not the analysis of the original experiment but only calibrating numerical studies.} For either application, we estimate a linear model as in Equation (ref), also controlling for additional covariates to guarantee the unconfoundedness of the treatment.\footnote{For cai2015social the covariates are gender, age, rice area, literacy level, a coefficient that captures the risk aversion, the baseline disaster probability, education, and a dummy containing information on whether the individual has one to five friends. For alatas2012targeting, we control for the education level, village-level treatments, i.e., how individuals have been targeted in a village (i.e., via a proxy variable for income, a community-based method, or a hybrid), the size of the village, the consumption level, the ranking of the individual poverty level, the gender, marital status, household size, the quality of the roof and top.} For simplicity, we consider as cost of treatment $c = \phi_1$, i.e., the opportunity cost of allocating the treatment to a population of disconnected individuals.
We generate $K$ clusters, each with $N = 600$ units, and sample $n \in \{200, 400, 600\}$. We generate a geometric network $ A_{i,j} = 1\Big\{||U_i - U_j||_1 \le 2 \rho/\sqrt{N}\Big\}, U_i \sim_{i.i.d.} \mathcal{N}(0, I_2), $ where the parameter $\rho$ governs the density of the network. The geometric formation process and the $1/\sqrt{N}$ follow similarly to simulations in leung2019treatment. We report results for $\rho = 2$ here, while results are robust as we increase $\rho$ (see Appendix (ref)). Throughout the analysis, without loss, we report welfare divided by its maximum $W(\beta^*)$ (i.e., $W(\beta^*) = 1$), and we subtract the intercept $\phi_0$.
In Appendix (ref), we study the performance of the one-wave experiment. We show that the proposed test controls size uniformly across specifications and present desirable properties for power. Here, we present simulations for the multi-wave experiment. In the adaptive experiment, we choose the learning rate $10\%/\sqrt{t}$ with gradient norm rescaling as Remark (ref).\footnote{This choice guarantees that for each iteration, we only vary treatment probabilities by at most $10\%$, and the size of the variation is decreasing over each iteration, as for the learning rate under strong concavity without norm rescaling. This choice is preferable to $10\%/\sqrt{T}$ because it allows for larger steps in the initial iterations. A valid alternative is $10\%/t$. The latter case has a practical drawback: updates become very small after a few iterations. Comparisons for different learning rates are in the online supplement (Fig (ref)). } Since the model does not allow for time-varying fixed effects, we estimate marginal effects without baseline outcomes. For the multi-wave experiment, we initialize parameters at a small treatment probability $\beta = 0.2$ (here the optimum is around $60\%$).
We let $T \in \{5, 10, 15, 20\}$. In Table (ref), we report the welfare improvement of the proposed method with respect to a grid search method that samples observations from an equally spaced grid between $[0.1, 0.9]$ with a size equal to the number of clusters (i.e., $2T$). We consider the best competitor between the one that maximizes the estimated welfare obtained from a correctly specified quadratic function and the one that chooses the treatment with the largest value within the grid. For both the competing methods, but not for the proposed procedure, we divide the outcomes' variance $\sigma^2$ by $T$, simulating settings where researchers may sample outcomes $T$ times (hence outcomes with a lower variance) from each cluster before estimating treatment effects, and obtaining more precise information. The panel at the top of Table (ref) reports the out-of-sample welfare improvement. The improvement is positive, and up to three percentage points for targeting information and up to sixty percentage points for targeting cash transfers. Improvements are generally larger for larger $T$. The panel at the bottom of Table (ref) reports positive and large improvements for the in-sample welfare across all the designs, worst-case across clusters. For the worst-case regret, we fix the number of clusters to $K = 40$ for the proposed method and study the properties as a function of the number of iterations. The improvements are twice as large for targeting information and thirty percentage points larger for targeting cash transfers. These are often increasing in $T$ with a few exceptions since uniform concentration may deteriorate for large $T$ and small $n$ as we consider the worst-case welfare across clusters.
In the online Appendices (ref), (ref), we report results across many other specifications of the network, policy functions, and choice of different parameters and different starting values (e.g., also when $\beta$ is initialized near the optimum).
This paper makes two main contributions. First, it introduces a single-wave experimental design to estimate the marginal effect of the policy and test for policy optimality. The experiment also enables identifying and estimating treatment effects, which can be of independent interest. Second, it introduces an adaptive experiment to maximize welfare. We derive asymptotic properties for inference and provide a set of guarantees on the in-sample and out-of-sample regret. We illustrate the benefits of the method in a large-scale field experiment on information diffusion. Our empirical application shows that using the marginal effect can be informative for decision-making even with few (two) waves.
This work opens new questions also from a theoretical perspective. We leave to future research the study of properties of the estimators when (i) clusters are not fully disconnected, in the spirit of leung2023network; (ii) clusters need to be estimated, similarly to graph-clustering procedures; (iii) clusters present different distributions, as we discuss in Appendix (ref). Similarly, studying the properties of the proposed method, as the degree of interference is proportional to the sample size, is an interesting direction. This is theoretically possible, as illustrated in Theorem (ref), and we leave its comprehensive analysis to future research. Finally, an open question is how to estimate policies when the network is only partially observed breza2017using, manresa2013estimating, and how to measure costs and benefits of collecting network data, on which Section (ref) provides novel directions for future research.