EconBase
← Back to paper

Policy design in experiments with unknown interference

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

150,487 characters

Policy design in experiments with unknown interference



\if00
{
  \title{\vspace{-13mm} \bf Policy design in experiments with unknown interference\footnote{This paper is a revised version of the first author's job market paper, superseeeding previous versions that can be found at \url{https://arxiv.org/abs/2011.08174}. We are especially grateful to Graham Elliott, James Fowler, Paul Niehaus, Yixiao Sun, and Kaspar W\"uthrich for their continuous advice and support, and Karun Adusumilli, Isaiah Andrews, Tim Armstrong, Peter Aronow, Susan Athey,  Emily Breza, Ivan Canay, Raj Chetty, Arun Chandrasekhar, Tim Christensen, Aureo de Paula, Tjeerd de Vries, Matt Goldman, Peter Hull, Guido Imbens, Brian Karrer, Max Kasy, Larry Katz, Michal Kolesar, Toru Kitagawa, Pat Kline, Michael Kremer, Amanda Kowalski, Michael Leung, Xinwei Ma, Elena Manresa, Craig Mcintosh, Konrad Menzel, Karthik Muralidharan, Ariel Pakes,  Jack Porter, Max Tabord-Meehan, Ulrich Mueller, Gautam Rao, Jonathan Roth, Cyrus Samii, Pedro Sant'Anna, Fredrik S\"avje, Azeem Shaikh, Jesse Shapiro, Pietro Spini, Elie Tamer, Alex Tetenov, Ye Wang, Andrei Zeleneev, and participants at numerous seminars for helpful comments. We particularly thank the teams at Precision Development and the Center for Economic Research in Pakistan, especially Jagori Chatterjee, Khushbakht Jamal, and Adeel Shafqat.  The experiment is registered at \url{https://www.socialscienceregistry.org/trials/9945}. The experiment has received IRB approval from Stanford University, where DV was affiliated as a postdoctoral fellow during the data collection, and the University of Chicago. We thank the Chae Family Economics Research Fund, the International Fund for Agricultural Development (IFAD), the NBER Social Learning Fund, and the Development Innovation Lab for generous funding support. Helena Franco, Benjamin Zeisberg, Angelina Zhang provided excellent research assistance.
  All mistakes are our own.}}
  \author{Davide Viviano\footnote{Department of Economics, Harvard University. Correspondence: [email removed].} $\quad$ Jess Rudder\footnote{Department of Economics, University of Chicago. Email: [email removed]. }}
    \date{First Version: November, 2020; \\ This Version: May, 2024
   }

  \maketitle
} \fi

\if10
{
  \bigskip
  \bigskip
  \bigskip
  \
 \\

 \
 \\


  \medskip
} \fi

\vspace{-5mm}
 \begin{abstract}
\noindent



This paper studies experimental designs for estimation and inference on policies with spillover effects. Units are organized into a \textit{finite} number of large clusters and interact in unknown ways within each cluster. First, we introduce a single-wave experiment that, by varying the randomization across cluster pairs, estimates the marginal effect of a change in treatment probabilities, taking spillover effects into account. Using the marginal effect, we propose a test for policy optimality.
Second, we design a multiple-wave experiment to estimate welfare-maximizing treatment rules. We provide strong theoretical guarantees and an implementation in a large-scale field experiment.






























\end{abstract}

\noindent
{\it Keywords:} Experimental Design, Spillovers, Welfare Maximization, Causal Inference. \\
{\it JEL Codes:} C31, C54, C90.
\vfill

\newpage
 \section{Introduction}
 \onehalfspacing





One of the goals of a government or NGO is to estimate the welfare-maximizing policy. Network interference is often a challenge: treating an individual may also generate spillovers and affect the design of the optimal policy. For instance, approximately 40\% of experimental papers published in the ``top-five” economic journals in 2020 mention spillover effects as a possible threat when estimating the effect of the program.\footnote{This is based on the authors' calculation. The top-five economic journals are \textit{American Economic Review}, \textit{Econometrica}, \textit{Journal of Political Economy}, \textit{Quarterly Journal of Economics}, \textit{Review of Economic Studies}.} Since budget constraints often bind, researchers have become increasingly interested in experimental designs for choosing the treatment rule (policy) that maximizes welfare. However, when it comes to experiments with spillovers, standard approaches are geared towards the estimation of treatment effects. Estimation of treatment effects, on its own, is not sufficient for welfare maximization.\footnote{Examples of  treatment effects are the direct effects of the treatment and the overall effect, i.e., the effect if we treat all individuals, compared with treating \textit{none}. For welfare maximization, none of these estimands are sufficient. The direct effect ignores spillovers, whereas the optimal rule may only treat some but not all individuals because of treatment costs or constraints.} For example, when designing information campaigns, information may have the largest direct effect on people living in remote areas but generate the smallest spillovers.
This trade-off has significant policy implications when treating each individual is costly or infeasible.






This paper studies experimental designs in the presence of interference when the goal is welfare maximization.
The main difficulty in these settings is that spillovers can be challenging to measure: when spillovers occur through an unobserved network, for example, collecting network information can be very costly because it may require enumerating all individuals and their connections in the population \citep[see][for a discussion]{breza2017using}. We, therefore, focus on a setting with limited information on the interference mechanism, formalized by assuming units are organized into a \textit{small} (finite) number of large clusters, such as schools, districts, or regions, and interact through an unobserved network (in unknown ways) within each cluster. In a development study, we may expect that treatments generate spillovers to those living in the same or nearby villages, but spillovers are negligible between different regions \citep[e.g.,][]{egger2019general}.\footnote{A finite number of clusters allows researchers to be agnostic on spillovers between different villages and only requires (approximate) independence between a few regions. Namely, the number of individuals who interact between different regions is ``small" relative to the number of individuals in a region \citep[][]{leung2023network}.} We propose the first experimental design to estimate \textit{welfare-maximizing} treatment rules in such contexts with unobserved spillovers.




















This paper makes two main contributions. First, we introduce a design where
researchers randomize treatments and collect outcomes once (\textit{single-wave experiment}) with two goals in mind: (i) to test whether one
or more treatment allocation rules, such as the one currently implemented by the policymaker, maximize welfare; and (ii) to estimate how
one can improve welfare with a (small) change to
allocation rules. The experimental design is based on a simple idea.
With a small number of clusters, we do not have enough information to estimate the welfare-maximizing treatment rule precisely.
However, if we take \textit{two} clusters and assign treatments in each cluster independently with slightly different (locally perturbated) probabilities, we can estimate the marginal effect of a change in the treatment assignment rule, which we refer to as marginal \textit{policy} effect (MPE).
In the development study example above, the MPE defines the marginal effect of treating more people in remote areas, taking spillover effects into account.\footnote{The MPE is the derivative of welfare with respect to the policy's parameters, taking spillovers into account, different from what is known in observational studies as the marginal treatment effect \citep{carneiro2010evaluating}, which instead depends on the individual selection into treatment mechanism.}
Using the MPE, we introduce a practical test for whether a welfare-improving treatment allocation rule exists. The MPE indicates the \textit{direction} for a welfare improvement, and the test provides evidence on whether
conducting additional experiments to estimate a welfare-improving treatment allocation is worthwhile.














































Using a small (finite) number of clusters, the experiment \textit{pairs} clusters and randomizes treatments independently within clusters, with local perturbations to treatment probabilities within each pair.  The difference in treatment probabilities balances the bias and variance of a difference-in-differences estimator. We show that the estimator for each pair converges to the marginal effect as the cluster's size increases, and we derive properties for inference with finitely many clusters.
Importantly, the experiment separately estimates the direct, spillover and welfare effects -- often of independent interest -- by \textit{pooling} observations across all pairs.








































As a second contribution, we offer an adaptive (i.e., \textit{multiple-wave}) experiment to estimate welfare-maximizing allocation rules. The goal is to adaptively randomize treatments to estimate the welfare-maximizing policy while improving participants' welfare, desirable in (large-scale) experiments \citep[][]{muralidharan2017experimentation}.
Our design guarantees tight small-sample bounds for \textit{both} the (i) out-of-sample regret, i.e., the difference between the maximum welfare and the welfare evaluated at the estimated policy deployed on a new population, and the (ii) in-sample regret, i.e., the regret of the experiment participants.


The experimental design groups clusters into pairs, using as many pairs as the number of iterations (or more); every iteration, it randomizes treatments in a cluster and perturbs the treatment probability within each pair; finally, it updates policies sequentially, using the information on the marginal effects from a different pair via gradient descent. Because of repeated sampling, conditional on the past, the estimated marginal effect
may present a bias due to serial dependence and interference, different from standard adaptive (batch) experiments. We introduce a novel algorithm that avoids this bias through sequential updates.













We investigate the theoretical properties of the method. A corollary of the small-sample guarantees is that the out-of-sample regret converges at a faster-than-parametric rate in the number of clusters and iterations and, similarly, the in-sample regret. No regret guarantees in previous literature are tailored to unobserved interference. Existing results with $i.i.d.$ data, treating clusters as sampled observations, would instead imply a slower convergence in the number of clusters.\footnote{Here, the average out-of-sample regret converges at a rate $1/T$, where $T$ is the number of iterations and proportional to the number of clusters, and at a rate $\log(T)/T$ for the in-sample regret. For the out-of-sample regret, we derive an exponential rate $\exp(-c_0 T)$, for a positive constant $c_0$ under additional restrictions (see Section \ref{sec:theory}).
\cite{KitagawaTetenov_EMCA2018}, \cite{shamir2013complexity} establish distribution-free lower bounds of order $1/\sqrt{n}$ for treatment choice and continuous stochastic bandits, respectively. Optimization connects to bandits of \cite{flaxman2004online, agarwal2010optimal}, which, however, provide slower rates for high-probability bounds (see also Section \ref{sec:theory}). \cite{wager2019experimenting} provide rates of order $1/T$ for in-sample regret but leverage an explicit model for market interactions with asymptotically independent individuals. Here, we do not impose assumptions on the interference mechanism and consider a different setup with partial interference and finitely many clusters.  } We achieve a faster rate by (a) exploiting \textit{within}-cluster variation in assignments and \textit{between} clusters' local perturbations; (b) deriving concentration within each cluster; (c) assuming and leveraging decreasing marginal effects of increasing neighbors' treatment probability. Fast convergence rates in the number of (large) clusters are particularly interesting when researchers have limited knowledge about interference and can partition units only into a few (approximately) independent clusters.




What is the benefit (and cost) of designing policies without network data? As an additional contribution,
Section \ref{sec:comp} characterizes the welfare value of collecting network data. We consider experiments with network spillovers occurring through a sufficiently dense network, and separable direct and spillover effects. We bound the difference between the maximum welfare achievable for \textit{any} policy that uses network information and the welfare of the policy that does not use network data. This bound depends only on the direct treatment effect minus the cost of treatment. This can be identified in single-wave experiments \textit{without} network data and provides novel results to guide practitioners on the value of network data.














We then turn to the implementation of the experimental design. In collaboration with Precision Development (PxD), an NGO providing agronomy advice in developing countries, we implemented a large-scale experiment with over 250,000 farmers to test some of the method's properties with two-wave experiments. The experiment provided geo-localized (county-level) weather forecasts to farmers in rural Pakistan to improve agronomy activities, as farmers often lack geolocalized forecasts (available forecasts are typically at the state instead of the county level). Spillover effects are relevant in this application: in a survey conducted by PxD, $80\%$ of surveyed individuals said they shared weather information with other farmers. The experiment consisted of two consecutive waves. Each wave was designed ex-ante to implement our perturbation design as in Section \ref{sec:jh}. We use variation between counties to learn marginal effects and spillovers when treating $50\%$ of the individuals in the first wave and $70\%$ in the second wave (the experiment also included some variants discussed in Section \ref{sec:field}).
Using high-frequency survey data merged with daily weather information, we show that farmers improve their beliefs about one-day ahead weather forecasts, and the program generates spillovers. We observe positive marginal policy effects over the first wave and close to zero marginal effects over the second wave, suggesting that treating $70\%$ suffices to maximize information diffusion. By using information about the marginal effect from our experiment, we can reduce the costs of the program by one million US dollars/year once implemented at scale in Pakistan. We complement our findings with simulations, calibrated to experiments on information \citep{cai2015social} and cash-transfers \citep{alatas2012targeting}.






 Throughout the text, we assume that the maximum degree of dependence grows at an appropriate slower rate than the cluster size;  covariates and potential outcomes are identically distributed between clusters; treatment effects do not carry over in time.
In the Appendix, we relax these assumptions and study three extensions: (a) experimental design with a global interference mechanism; (b) matching clusters via distributional embeddings with covariates drawn from cluster-specific distributions; and (c) experimental design with \textit{dynamic} treatment effects, and propose a novel experimental design in this setting. Practitioners may refer to Section \ref{sec:why_assumptions} for more discussion about the applicability of our methods.

















































































































































We contribute to the literature on \textit{single-wave} experiments, where existing network experiments include clustered experiments and saturation designs  \citep{baird2018optimal}. References with observed networks include  \cite{basse2018model},
 \cite{viviano2020experimental} among others. For the analysis of the bias of average treatment effect estimators with interference, see also \cite{basse2016analyzing}, \cite{johari2020experimental}, and \cite{imai2009essential}.
  Additional references are \cite{bai2019optimality, tabord2018stratification} with $i.i.d.$ data. These authors study experimental designs for inference on treatment effects but not inference on welfare-maximizing policies. Different from the above references, we propose a design to identify the \textit{marginal} policy effect under interference, used for hypothesis testing and welfare maximization. The focus on marginal policy effects connects to the literature on optimal taxation \citep{chetty2009sufficient}, which differs from our setting by considering observational studies with independent units.




With \textit{multiple-wave} experiments, we introduce a framework for adaptive experimentation with unknown interference.
We connect to the recent literature on adaptive exploration \citep[e.g.,][among others]{bubeck2012regret, kasy2019adaptive}, and the one on derivative-free stochastic optimization, dating back to \cite{kiefer1952stochastic}, and \cite{flaxman2004online, kleinberg2005nearly, shamir2013complexity, agarwal2010optimal}, among others.
These references do not study the problem of interference (and inference).
Here, we leverage between-cluster perturbations and within-cluster concentration to obtain fast rates of regret in high probability (see Section \ref{sec:theory} for a comprehensive discussion).
\cite{wager2019experimenting} study price estimation in a single market \citep[and similarly][in subsequent work to ours]{munro2021treatment}.  They assume infinitely many individuals and an explicit model for market prices under which agents are asymptotically independent. As noted by the authors, the structural assumptions imposed in their paper do not allow for spillovers on a network (i.e., individuals may depend arbitrarily on neighbors' assignments). Our setting differs because individuals are organized into finitely many independent clusters here, where unobserved (network) spillovers may occur. These differences motivate (i) our design, which exploits two-level randomization at the cluster and individual level instead of individual-level randomization, and (ii) cluster-level perturbations.
From a theoretical perspective, dependence and repeated sampling induce novel challenges studied in this paper.










We relate to inference under interference and draw from \cite{hudgens2008toward} for definitions of potential outcomes. \cite{aronow2017estimating, manski2013identification, leung2019treatment, goldsmith2013social, li2020random} assume an observed network, while \cite{vazquez2017identification}, \cite{ibragimov2010t} consider clusters among others. \cite{savje2017average} study inference of the direct effect only. Discussion about relevant estimands in these frameworks can also be found in more recent work by \cite{hu2021average}. None of these study welfare maximization or experimental design, different from this paper.




More broadly, we connect to the treatment choice literature on estimation \cite{manski2004, KitagawaTetenov_EMCA2018, athey2017efficient, stoye2009minimax, mbakop2016model, kitagawa2020should, sasaki2020welfare, viviano2019policy}, and inference \cite{andrews2019inference, rai2018statistical, armstrong2015inference, kasy2016partial, hadad2019confidence, hirano2020asymptotic}. This literature considers an existing experiment instead of experimental designs, and has not studied policy design with unobserved interference. Here, we leverage an adaptive procedure to maximize out-of-sample and participants' welfare. We broadly relate also to the literature on targeting on networks \citep[e.g.,][]{bloch2019centrality, banerjee2013diffusion, akbarpour2018just}, which mainly focuses on particular models of interactions in a single observed network -- different from here, where we leverage clusters' variations; the one on peer-group composition \citep{graham2010measuring}, the one on inference with externalities \citep[e.g.,][]{bhattacharya2013estimating}, and pioneering work on vaccination campaigns \citep{manski2010vaccination, manski2017mandating}. None of these study experimental designs.







\vspace{-3.5mm}

\section{Setup and overview} \label{sec:2nd}






We consider a setting with $K$ clusters, where $K$ is an even number. We assume each cluster has $N$ individuals, whereas the framework directly extends to clusters of different but proportional sizes.
Observables and unobservables are jointly independent between clusters but not necessarily within clusters, as often assumed in economic applications \citep[e.g.,][see Remark \ref{rem:dependent_clusters} for discussion]{abadie2017should}. Each cluster $k$ is associated with a vector of outcomes, treatments, and covariates. These are
$
Y_{i,t}^{(k)} \in \mathcal{Y}, D_{i,t}^{(k)} \in \{0,1\},  X_i^{(k)} \in \mathcal{X} \subseteq \mathbb{R}^{L},
$ respectively.
Here, $(Y_{i,t}^{(k)}, D_{i,t}^{(k)})$ denote the outcome and treatment assignment of individual $i$ at time $t$ in cluster $k$, respectively, $X_i^{(k)}$ are time-invariant (baseline) covariates.
For each period $t$, researchers observe a random subsample,
 $$
 \small
 \begin{aligned}
 \Big(Y_{i,t}^{(k)}, X_i^{(k)}, D_{i,t}^{(k)}\Big)_{i=1}^n,\quad  n = \lambda N, \quad \lambda \in (0,1],
 \end{aligned}
 $$
where $n$ defines the sample size of observations from each cluster and is proportional to the cluster size for expositional convenience. There are $T$ periods.
Although units sampled each period may or may not be the same, with abuse of notation, we index sampled units $i \in \{1, \cdots, n\}$. We denote $Y_{i,t}^{(k)}(\mathbf{d}_1^{(k)}, \cdots, \mathbf{d}_t^{(k)}), \mathbf{d}_s^{(k)} \in \{0,1\}^N, s \le t$ the potential outcome of individual $i$ in cluster $k$ at time $t$, as a function of the treatments of all other units in the same cluster. The definition of potential outcomes implicitly imposes no cross-interference between clusters and no anticipation, standard in the literature \citep[e.g.][]{athey2018design}. We will refer to $Y^{(k)}(\cdot)$ as the \textit{potential} outcome functions of all units in cluster $k$.

 Whenever we provide asymptotic analyses, we let $N$ grow through a sequence of data-generating processes and let $K$ be fixed. Here, $n$ is proportional to $N$ for expositional convenience. We take a super-population perspective where potential outcomes $Y^{(k)}(\cdot)$ are random variables. The super-population perspective can also be interpreted as assuming that finite $K$ clusters are drawn from a super-population (see Remark  \ref{rem:model_based} and Section \ref{sec:1a}).








\vspace{-3mm}

\subsection{Outcomes, policy choice and welfare maximization} \label{sec:1b}


We focus on a parametric class of policies (treatment rules) indexed by some parameter $\beta$,
$$
\small
\begin{aligned}
\pi(\cdot;\beta): \mathcal{X} \mapsto [0,1], \quad \beta \in \mathcal{B},
\end{aligned}
$$
a map that prescribes the individual treatment probability based on covariates. Here, $\mathcal{B}$ is a compact parameter space, and $\pi(x, \beta)$ is twice differentiable in $\beta$. The experiment assigns treatments independently based on $\pi(\cdot)$, and time/cluster-specific parameters $\beta_{k,t}$.
Motivated by empirical practice \citep[e.g.,][]{baird2018optimal}, we focus on two-stage experiments where, given the parameter $\beta_{k,t}$ in cluster $k$ at time $t$, treatments are assigned independently.





\begin{ass}[Treatment assignments in the experiment] \label{defn:bernoulli} For  $\beta_{k,t} \perp \Big(X^{(k)},Y^{(k)}(\cdot) \Big)$,
$$
\small
\begin{aligned}
D_{i,t}^{(k)} | X^{(k)}, Y^{(k)}(\cdot), \beta_{k,t} \sim_{i.n.i.d.} \mathrm{Bern}\Big(\pi(X_i^{(k)};\beta_{k,t})\Big),
\end{aligned}
$$
which, for brevity of notation, we refer to as $D_{i,t}^{(k)} | X^{(k)}, Y^{(k)}(\cdot), \beta_{k,t} \sim \pi(X_i^{(k)}, \beta_{k,t})$, where $i.n.i.d.$ indicates independently and not identically distributed.
\end{ass}

Assumption \ref{defn:bernoulli} defines a treatment rule in experiments. Treatments are assigned independently based on covariates and time and cluster-specific parameters $\beta_{k,t}$.
The assignment in Assumption \ref{defn:bernoulli} is easy to implement: it can be implemented in an online fashion (i.e., sequentially across units)  and does not use information about the outcomes' dependence structure, which justifies its choice; also, it generalizes assignments in saturation designs studied for inference on treatment effects \citep{baird2018optimal}. An example is treating individuals with equal probability \citep{akbarpour2018just}, i.e.,
 $
\pi(\cdot; \beta) = \beta \in [0,1]$.
We can also \textit{target} treatments, i.e., $
\pi(x;\beta) = \beta_x,
$
indicating the treatment probability for $X_i^{(k)} = x$ (with $\mathcal{X}$ discrete).
The parameters $\beta_{k,t}$ must be exogenous with respect to potential outcomes in the same cluster to guarantee unconfoundedness, as in standard RCTs. It is possible to let $\beta_{k,t}$ depend on observable clusters' characteristics as discussed in Appendix \ref{app:het}, and omitted here for brevity.
With an adaptive experiment, Assumption \ref{defn:bernoulli} holds for the design presented in Section \ref{sec:main_design} (see Remark \ref{rem:aaa}).

We defer to Section \ref{sec:comp}, studying more complex assignments with dependent treatments.





Throughout the main text, whenever we write $\pi(\cdot; \beta)$, omitting the subscripts $(k,t)$, we refer to a generic exogenous (i.e., not data dependent) vector of parameters $\beta$. We define $\mathbb{E}_\beta[\cdot]$ as the expectation taken over the distribution of treatments assigned according to $\pi(\cdot; \beta)$.



















































\begin{ass}[Data generating process] \label{ass:ass_0}
Suppose that for any $(i,t,k)$,
\begin{itemize}
\item[(i)] $Y_{i,t}^{(k)}(\mathbf{d}_1^{(k)}, \cdots, \mathbf{d}_t^{(k)})$ is constant in $\mathbf{d}_1^{(k)}, \cdots, \mathbf{d}_{t-1}^{(k)}$, and $X_i^{(k)} \sim F_X$ for (unknown) $F_X$;
\item[(ii)] Under an assignment in Assumption \ref{defn:bernoulli} with parameter $\beta_{k,t}$, the following holds:
\begin{equation} \label{eqn:main_Y0}
\small
\begin{aligned}
\mathbb{E}_{\beta_{k,t}}\Big[Y_{i,t}^{(k)} | D_{i,t}^{(k)} = d , X_i^{(k)} = x\Big] = m(d, x, \beta_{k,t}) + \alpha_t + \tau_k, \quad Y_{i,t}^{(k)} = Y_{i,t}^{(k)}(D_1^{(k)}, \cdots, D_t^{(k)})
\end{aligned}
\end{equation}
for some (unknown) function $m(\cdot)$, and fixed effects $\alpha_t, \tau_k$, where the expectation is also taken over the potential outcome function $Y_{i,t}^{(k)}(\cdot)$.  In addition, let $Y_{i,t}^{(k)} \perp \{Y_{j,t}^{(k)}\}_{j \not \in \mathcal{I}_i^{(k)}} | \beta_{k,t}$ for  a set of indexes $\mathcal{I}_i^{(k)}$ with cardinality $|\mathcal{I}_i^{(k)}| \le 2 \gamma_N$, for some $\gamma_N \ge 1$.
 \end{itemize}
\end{ass}


Assumption \ref{ass:ass_0} (i) states that effects do not carry over in time, as often assumed in studies on experiments \citep{kasy2019adaptive, athey2018design} (this is only required for the adaptive but not the single wave experiment). It also states that covariates $X_i$ have the same distribution across clusters. Appendix \ref{sec:dynamics_main} presents extensions with dynamics, and Appendix \ref{app:het} with covariates drawn from different distributions.


   Assumption \ref{ass:ass_0} (ii) imposes restrictions on the expectation of the (potential) outcome, also integrating over the distribution of the other units' assignments. The first component in Equation \eqref{eqn:main_Y0} is the conditional expectation given the individual covariates and the parameter $\beta_{k,t}$, \textit{unconditional} on other units' assignments and unobservables (i.e., potential outcome function). The dependence of $m(\cdot)$ with $\beta_{k,t}$ captures spillover effects because treatments' distribution depends on $\beta_{k,t}$.
The second components are separable fixed effects.

Whereas treatment effects may exhibit individual-level heterogeneity (see Example \ref{exmp:microfoundation} and discussion therein), treatments do not interact with clusters' fixed effects, imposing homogeneity of treatment effects across different clusters. Homogeneity restrictions across clusters is common in many applications \citep[e.g.,][]{cai2015social, miguel2004worms, crepon2013labor, duflo2023chat}.

 Assumption \ref{ass:ass_0} (ii) also states that outcomes depend with at most $\gamma_N$ many other outcomes in the same cluster (conditional on the assignment mechanism $\beta_{k,t}$). Here, $\gamma_N$ provides an interpretable restriction on the dependence structure.
As we show in Section \ref{sec:1a}, in our leading application of network spillovers, $\gamma_N^{1/2}$ defines the largest number of connections of a given individual and, therefore, restrictions on $\gamma_N$ imposes restrictions on the maximum degree. From a theoretical perspective, we require forms of weak dependence within each cluster, motivated by clusters being large regions; in some cases, we can allow for settings where $\gamma_N$ can grow arbitrarily with $N$, see Remark \ref{rem:global} for a discussion.

We defer to Section \ref{sec:why_assumptions} a discussion on the applicability of our assumptions and to Appendix \ref{sec:extensions1} numerous extensions, including settings with observed heterogeneity.





\begin{defn}[Welfare] \label{defn:welf1} For treatments as in Assumption \ref{defn:bernoulli} with $\beta$ parameter, let welfare be
 $W(\beta) =  \int y(x,\beta) dF_X(x),
$
where $y(x, \beta) = \pi(x;,\beta)  m(1, x, \beta) + (1 - \pi(x;\beta)) m(0, x, \beta)$.
 \end{defn}

Welfare defines the expected outcome had treatments been assigned with policy $\pi(\cdot, \beta)$. We do not include fixed effects in the definition of welfare without loss, since such effects are separable. The expectation is taken over treatment assignments, covariates, and potential outcomes. We interpret $y(x,\beta)$, the outcome \textit{net of costs} and incorporate the costs in the outcome function, as often assumed \citep{KitagawaTetenov_EMCA2018}. We define the welfare-optimal policy and the marginal effect (under differentiability in Assumption \ref{ass:regularity_basic})
 \begin{equation}  \label{eqn:estimand}
 \small
 \begin{aligned}
 \beta^* \in \mathrm{arg} \sup_{\beta \in \mathcal{B}} W(\beta), \quad M(\beta) = \frac{\partial W(\beta)}{\partial \beta}.
 \end{aligned}
 \end{equation}

The marginal effect defines the derivative of the welfare with respect to the vector of parameters $\beta$. Finally, we define the direct and marginal spillover effects, respectively as
$$
\small
\begin{aligned}
\Delta(x, \beta) = m(1, x, \beta) - m(0, x, \beta), \quad S(d, x, \beta) = \frac{\partial m(d, x, \beta)}{\partial \beta}, \quad d \in \{0, 1\}, x \in \mathcal{X}, \beta \in \mathcal{B}.
\end{aligned}
$$
The direct effect denotes the effect of the treatment, keeping constant the neighbors' treatment probability, and the marginal spillover effect $S(\cdot)$, the marginal effect of changing neighbors' treatment probabilities, keeping constant the individual treatment. We can write
 \begin{equation} \label{eqn:marginal}
 \small
 M(\beta) = \int  \Big[\underbrace{\pi(x;\beta) S(1, x, \beta) + (1 - \pi(x; \beta))  S(0, x, \beta)}_{(S)} +  \underbrace{\frac{\partial \pi(x;\beta)}{\partial \beta} \Delta(x, \beta)}_{(D)}\Big] dF_X(x).
 \end{equation}
The MPE $M(\beta)$ depends on the weighted direct (D) marginal spillover (S) effects. Equation \eqref{eqn:marginal} follows in the spirit of decompositions in  \cite{hudgens2008toward}.\footnote{We also note that in more recent work, \cite{hu2021average} motivate targeting as causal estimand the average indirect effect, different from $S(\cdot)$ with heterogeneous assignments.  \cite{graham2010measuring} present peer effects' decompositions in the different contexts of peer groups' formation.
}
As we show, the marginal effect is key to improving (maximizing) welfare with only a \textit{few} clusters. We conclude with examples of welfare functions in the presence of network spillovers, our leading example, and defer to Section \ref{sec:1a} general models with such spillovers.









































\begin{exmp}[Positive externalities with decreasing returns from neighbors' treatments] \label{exmp:main} Let $D_{i,t}^{(k)} \sim_{i.i.d.} \mathrm{Bern}(\beta)$, $\mathcal{N}_i$ the set of friends (neighbors) of individual $i$, and
\begin{equation} \label{eqn:ex}
\small
 \begin{aligned}
 Y_{i, t} &=  \alpha_t + D_{i, t} \phi_1 + \frac{\sum_{j \in \mathcal{N}_i} D_{j,t}^{(k)} }{|\mathcal{N}_i|} \phi_2  - \left(\frac{\sum_{j \in \mathcal{N}_i} D_{j,t}}{|\mathcal{N}_i|} \right)^2 \phi_3 + \nu_{i,t}, \quad \mathbb{E}[\nu_{i,t}] = 0.
 \end{aligned}
 \end{equation}
Equation \eqref{eqn:ex} states that outcomes depend on the individual treatment, and the percentage of treated neighbors. Let the number of friends $|\mathcal{N}_i| \sim \mathcal{D}_N$ for some unknown $\mathcal{D}_N$. With some algebra, taking expectations, for $X_i = 1$, and letting $c$ denote the cost of treatment
$$
\small
\begin{aligned}
y(1, \beta) =  \beta (\phi_1 - c) + \beta \phi_2  - \beta \phi_3 \iota - \beta^2 \phi_3 (1 - \iota), \quad \iota = \mathbb{E}[1/|\mathcal{N}_i|].
\end{aligned}
$$
See Appendix Figure \ref{fig:objectives_analysis} calibrated to data from \cite{cai2015social}, and  \cite{alatas2012targeting}.
\end{exmp}



\begin{exmp}[Negative externalities] \label{exmp:main2} Let $D_{i,t}^{(k)} \sim_{i.i.d.} \mathrm{Bern}(\beta)$,
\begin{equation} \label{eqn:ex2}
\small
 \begin{aligned}
 Y_{i, t} &=  \alpha_t + D_{i, t} \phi_1 - \frac{\sum_{j \in \mathcal{N}_i} D_{j,t}^{(k)} }{|\mathcal{N}_i|} \phi_2  - D_{i,t} \frac{\sum_{j \in \mathcal{N}_i} D_{j,t}^{(k)} }{|\mathcal{N}_i|} \phi_3 + \nu_{i,t}, \quad \mathbb{E}[\nu_{i,t}] = 0.
 \end{aligned}
 \end{equation}
Equation \eqref{eqn:ex} states that outcomes depend on the individual treatment, treatments may generate negative externalities for positive $\phi_1, \phi_2, \phi_3$ \citep[such as in labor markets,][]{crepon2013labor}. It follows for $X_i = 1$,  $c$ the cost of treatment,
$
y(1, \beta) =  \beta (\phi_1 - \phi_2 - c)  - \beta^2 \phi_3.
$   \qed
\end{exmp}


\begin{rem}[Non-separable fixed effects] It is possible to extend our framework to settings with non separable fixed effects in time and cluster identity $\alpha_{k, t}$, assuming that spillovers only occur either on the treated or control units. We provide details in Appendix \ref{sec:non_separable}. \qed
\end{rem}


\begin{rem}[Dependent clusters] \label{rem:dependent_clusters} In some applications, clusters may only be approximately independent. In this case, we would require that between-clusters correlations are \textit{asymptotically} negligible at an appropriate fast rate.  \qed
\end{rem}


\begin{rem}[Global interference] \label{rem:global} Although some of our results impose restrictions on how $\gamma_N$ grows with $N$, Theorem \ref{thm:const1} and Appendix \ref{sec:global} present extensions for which no restrictions on $\gamma_N$ are imposed. Theorem \ref{thm:const1} shows that for consistency we only require that the correlation between potential outcomes (but not necessarily the maximum degree of dependence) decays at an appropriate slow rate \citep[see][for a discussion on weak dependence]{leung2023network}. \qed
\end{rem}







 \vspace{-2mm}

\subsection{Method's overview: What can we learn with a few clusters?} \label{sec:overview}







Ideally, we would like to leverage variation from a single-wave experiment to estimate treatment rules as in \cite{KitagawaTetenov_EMCA2018, athey2017efficient, rai2018statistical}. Two constraints here make this infeasible: researchers (i) do not know the spillover mechanism in each cluster (e.g., do not have access to network data in the presence of network spillovers); (ii) researchers only have access to a limited (finite) number of (approximately) independent clusters (e.g., small villages cannot be directly used as clusters because spillovers may also propagate between small villages). Because of (i), we cannot estimate the spillover effects on each individual from a single cluster; because of (ii), we cannot consider each cluster as a sampled observation. Instead, we leverage restrictions on the heterogeneity across clusters and limited dependence within each cluster to show that we can use two clusters to consistently estimate the marginal effect $M(\beta)$, at given $\beta$.

As an illustrative example, consider a policymaker who must allocate treatments to \textit{half} of the population. Consider two household types, $X_i \in \{0, 1\}$, with $P(X_i = 1) = 1/2$, e.g., those living in urban and more remote areas. The policymaker assigns treatments $D_{i,t} | X_i = x \sim \mathrm{Bern}(\pi(x, \beta) )$, where $\pi(x, \beta) = x \beta + (1 - x) ( 1 -\beta)$ is the treatment probability for $x \in \{0,1\}$ that by construction incorporates the budget constraint. Different treatment probabilities for people in remote areas produce different welfare effects. Figure \ref{fig:ex1} presents an illustration calibrated to data from \cite{alatas2012targeting, alatas2016network}.\footnote{Figure \ref{fig:ex1} serves as a simple illustration. We estimate a function heterogeneous in the distance of the household's village from the district's center. We use information from approximately $400$ observations, whose $80\%$ or more neighbors are observed. We let $X_i \in \{0,1\}, X_i = 1$ if the household is farther from the district's center than the median household, and estimate a quadratic model, with treatment denoting a cash transfer and the outcome denoting the individual satisfaction with the program.} Spillovers may exhibit decreasing marginal effects, and assigning all treatments to individuals in remote areas is sub-optimal. In addition, because we do not know the spillover mechanism, using variation from a single-wave experiment is insufficient to estimate $\beta^*$.





Instead, we show that with only \textit{two clusters}, we can estimate the marginal effect for:
 \begin{itemize}
\item[(a)] \textit{Policy updating}: estimate the welfare-improving \textit{direction} (increase or decrease $\beta$);
\item[(b)] \textit{Hypothesis testing}: assuming $\beta^*$ is an interior point,
 $
M(\beta) \neq 0$, implies $ W(\beta) \neq W(\beta^*).
 $
 \end{itemize}

Given the marginal effect, we can present to the policy-maker how we can improve policies through \textit{incremental} updates to the baseline intervention, only using a few clusters.
In addition, we can test whether the line's slope in Figure \ref{fig:ex1} is zero (with one or two-sided tests), suggesting evidence of whether the current policy is welfare-optimal. (Note that, as in standard hypothesis testing setups, rejection can be informative, while failure of rejection is informative only with well-powered studies, i.e., sufficiently large clusters' size $n$.)






\paragraph{Single wave} We proceed to construct estimators of the marginal effect. We start from Equation \eqref{eqn:marginal}.
The direct effect (D) can be identified from a single cluster, taking the difference between treated and untreated outcomes. However, the spillover effect (S) cannot be identified from a single cluster. We instead exploit variations between two clusters. We take two clusters, such as two regions. We collect \textit{baseline} ($t = 0$) outcomes and covariates; we then randomize treatments with slightly different probabilities between the regions. In the first region, we treat individuals in remote areas ($X_i = 1$) with probability $\beta + \eta_n$. Here, $\eta_n$ is a small deterministic number (local perturbation). The remaining individuals are treated with probability $1 - \beta - \eta_n$. In the second region, we treat individuals in remote areas with probability $\beta - \eta_n$, and the remaining ones with probability $1 - \beta + \eta_n$.



As shown in Figure \ref{fig:ex1}, we can estimate welfare for two different but similar treatment probabilities; the line's slope between the points is approximately equal to the marginal effect. That is, for a suitable choice of $\eta_n$ (see Theorem \ref{thm:const1}), a consistent marginal effect's estimator is
\begin{equation} \label{eqn:gradient_main}
\small
\begin{aligned}
 \widehat{M}_{(k, k+1)}(\beta)  = \frac{1}{2 \eta_n }  \Big[\bar{Y}_1^{(k)} - \bar{Y}_0^{(k)}\Big] - \frac{1}{2 \eta_n } \Big[\bar{Y}_1^{(k + 1)} - \bar{Y}_0^{(k + 1)}\Big],
 \end{aligned}
\end{equation}
where $\bar{Y}_t^{(h)}$ is the outcomes' sample average in cluster $h$ at time $t$, $Y_{i,0}$ is the baseline outcome with no experiment in place yet, and $(k, k+1)$ index the two clusters. The above estimator is a difference-in-differences; we subtract baseline outcomes due to fixed effects.
In Section \ref{sec:jh}, we present a test for $M(\beta) = 0$ using a few clusters' pairs. We also discuss estimation and inference on direct and marginal spillover effects, see Table \ref{tab:methods1}. A by-product of our design is that it does not require large deviations from the baseline intervention between different regions (large deviations can be expensive or difficult to justify to the general public).






\vspace{-4mm}
\paragraph{Multi-wave} Using the marginal effect, we then propose and study the following sequential experiment: (1) we pair clusters and organize pairs in a circle as in Figure \ref{fig:clusters}; (2) every step $t$, we estimate the marginal effect within each pair; (3) using the estimated marginal effect from the subsequent pair on the circle, we update the policy in a given clusters' pair.

The sequential updating rule guarantees that the policy achieves an optimum, either global with a (quasi)concave objective or local optimum otherwise.
Step (3) is key to overcoming a bias that, as we show in Section \ref{sec:main_design}, would otherwise arise here due to repeated sampling, while it maximizes the number of clusters that we can use in the experiment.

We measure the method's performance based on the out-of-sample and in-sample regret, respectively defined for an estimated policy $\hat{\beta}$ and \textit{sequence} of policies $\{\beta_{k,t}\}_{k=1, t=1}^{K, T}$ in the experiment,
 $
 W(\beta^*) - W(\hat{\beta})$, and $\max_{k \in \{1, \cdots, K\}}   \frac{1}{T} \sum_{t=1}^T \Big[W(\beta^*) - W(\beta_{k,t})\Big].
 $



\begin{rem}[Single wave: A free lunch] \label{rem:free_lunch}
Empirical approaches often choose few (e.g., two) treatment probabilities ($\beta_1, \beta_2$), and assign multiple clusters to \textit{each} of these probabilities \citep[see the examples in][]{baird2018optimal, egger2019general}. Within each cluster, researchers randomize treatments as in Assumption \ref{defn:bernoulli}.
Researchers then estimate the contrast $m(d, \beta_1) - m(d, \beta_2), d \in \{0,1\}$, with simple differences in means estimators. Instead, in these settings, this paper recommends using the randomization device in Equation \eqref{eqn:perturbation} for each probability $(\beta_1,\beta_2)$, to estimate (i) the contrast $m(d, \beta_1) - m(d, \beta_2)$ with the same precision as in the original experiment, and (ii) the marginal effects $M(\beta_1), M(\beta_2)$. This is possible by inducing perturbations \textit{around} $(\beta_1, \beta_2)$, and pooling observations around $(\beta_1, \beta_2)$ when estimating $m(d, \beta_1), m(d, \beta_2)$. Section \ref{sec:theory_exp} shows that this approach induces a bias asymptotically negligible for inference on the contrasts. Because of pooling, it uses the same number of observations of standard saturation experiments to estimate $m(d, \beta_1) - m(d, \beta_2)$ without decreasing the estimator's variance. In addition, the proposed experiment allows estimating the marginal effects $M(\beta_1), M(\beta_2)$ that standard designs do not identify. See Table \ref{tab:advantages_pert2} for an illustration in the context of our application. \qed
\end{rem}


 \begin{rem}[Multiple waves: Alternatives for policy choice]
An alternative approach for estimating $\beta^*$ is to first estimate the function $y(\cdot)$ by assigning different treatment probabilities $\beta$ to different clusters, and then extrapolating the entire response function $y(\cdot)$. However, for a generic $p$-dimensional $\beta$, the out-of-sample regret is either sensitive to the model used for extrapolation or suffers a curse of dimensionality (e.g., when a grid search is employed). Second, this alternative approach does not control the \textit{in-sample} regret: it must incur significant in-sample welfare loss to estimate $y(\cdot)$. Appendix \ref{app:non_adaptive} presents a formalization.

An important insight is that, under restrictions on the within-cluster correlation, learning the MPE only requires one cluster pair. Using the MPE, we can (i) guarantee fast convergence rates of the regret, and (ii) provide direct guidance for decision-making as shown in our empirical application.
\qed
\end{rem}




\begin{figure}[!h]
\vspace{-2mm}
\center
\includegraphics[scale=0.35]{./figures/figure_sec2-eps-converted-to.pdf}
\caption{Example of experimental design, fixing the overall fraction of treated population to be half, and choosing between two types of individuals to treat (those in remote and non-remote regions).  The left panel is a single-wave experiment with two clusters. In the first cluster, we assign the policy colored in green, and the second cluster colored in brown. The right panel is a two-wave experiment. We use a pair of clusters to estimate the marginal effect and update the policy for a different pair.} \label{fig:ex1}
\end{figure}

\vspace{-8mm}

\subsection{Main assumptions and applicability of the method} \label{sec:why_assumptions}

We pause here to discuss our main assumptions and their applicability.

Our approach leverages two main assumptions: (i) treatments do not generate heterogeneous effects \textit{in expectation} across clusters; (ii) outcomes have limited (weak) dependence within each cluster. Here, (i) guarantees that potential outcomes' expectations are comparable across different clusters. Condition (ii) guarantees that we can estimate marginal effects even with only two clusters. With \textit{unobserved} heterogeneity and/or arbitrary dependence, we could not learn optimal policies (and marginal effects) with a few clusters.

The potential outcome model is consistent with models used in many applications, such as spillovers for agronomy advice \citep{duflo2023chat}, and others \citep[][]{cai2015social, miguel2004worms, crepon2013labor}. All of these papers consider specifications with homogeneous effects across clusters. Researchers may test for homogeneity by comparing the average baseline covariates across different clusters. An example is in Table \ref{tab:summaries}, where we show substantial homogeneity in our empirical application. In the presence of heterogeneity, however, we recommend appropriately balancing clusters (see Appendix \ref{app:het}).

We impose forms of weak dependence within clusters, mostly (but not necessarily) captured through restrictions on how $\gamma_N$ grows with $N$. This is motivated by clusters being large regions as in our application, where, we may expect, individuals interact only with a subset of individuals in the region \citep[see Example \ref{exmp:microfoundation} or][]{de2018identifying}. For example, in settings where we observe network data \citep{cai2015social}, individuals tend to connect with a few individuals within and between villages but not between different regions.


Two additional assumptions we will use with multiple waves of randomization are welfare (quasi)concavity and no carry-over (dynamics) in effects. Examples of concavity are Example \ref{exmp:main}, where neighbors' effects induce decreasing marginal effects (see Figure \ref{fig:objectives_analysis} using data from \cite{cai2015social}), or settings with negative externalities in Example \ref{exmp:main2}. Concavity fails when spillovers occur only after ``enough" individuals have received the treatment, for which we provide theoretical guarantees in Appendix \ref{sec:quasi_concavity}, under strict-quasi concavity. Under failure of (quasi)concavity, our proposed method will return a local instead of global optimum. See Assumption \ref{ass:strong_concavity} and discussion therein.

No carry-overs is a common assumption in (adaptive) experiments \citep[e.g.][]{kasy2019adaptive, athey2018design}, and in applications \citep[e.g.][]{duflo2023chat, cai2015social}.
In practice, carryovers do not occur if either each period $t$ is sufficiently far in time from the previous period or if the intervention only has short-term effects on the outcome. We encourage researchers to appropriately choose the time window $t$ and the outcome to guarantee that no dynamics occur. For example, in our application, the treatment (providing weather forecast for the upcoming few days) affects our main target outcome, i.e., a proxy for one-day ahead predictions of weather, but, as we show in Appendix \ref{app:more_experiment}, it does not affect forecasts in the upcoming weeks. See  \cite{athey2018design} for a discussion on carry-overs  and
Appendix \ref{sec:dynamics_main} for an extension with dynamics.





\begin{rem}[Super-population] \label{rem:model_based} We adopt a super-population perspective. This is useful due to unobserved spillovers and a finite number of clusters; the randomness in potential outcomes captures uncertainty over the spillover mechanism (as in Example \ref{exmp:microfoundation}) and allows us to control the dependence within a cluster. The focus on also maximizing welfare on \textit{new} clusters naturally requires restrictions on the potential outcomes' (repeated) sampling.
\qed
\end{rem}





























\subsection{Micro-foundation with network spillovers} \label{sec:1a}


We conclude with micro-foundation of Assumption \ref{ass:ass_0} in contexts with network spillovers, our leading application. Practitioners may skip this subsection and refer to Section \ref{sec:designs} directly. Suppose individuals are connected with other individuals through an unobserved and cluster-specific adjacency matrix $A^{(k)}$. Individuals can form a link with an (unknown) subset of individuals in each cluster. Nodes in each cluster are spaced under some latent space \citep{lubold2020identifying} and can interact with at most the $\gamma_N^{1/2}$ closest nodes under the latent space. We say $1\{i_k \leftrightarrow j_k\} = 1$ if individual $i$ can interact with $j$ in cluster $k$.
Conditional on $1\{i_k \leftrightarrow j_k\}$,
\begin{equation} \label{eqn:network}
\small
\begin{aligned}
&  (X_i^{(k)}, U_i^{(k)}) \sim_{i.i.d.}  F_X F_{U|X}, \quad A_{i,j}^{(k)} = l\Big(X_i^{(k)}, X_j^{(k)}, U_i^{(k)}, U_j^{(k)}\Big)1\{i_k \leftrightarrow j_k \}, \quad l:\mathcal{X}^2 \times \mathcal{U}^2 \mapsto [0,1],
\end{aligned}
\end{equation}
for an arbitrary and unknown function $l(\cdot)$ and unobservables $U_i^{(k)}$. Whether two individuals interact depends on (i) whether they are close enough within a certain latent space (captured by $1\{i_k \leftrightarrow j_k \}$); (ii) their covariates and unobserved individual heterogeneity (i.e., $X_i, U_i$), which capture homophily. Equation \eqref{eqn:network} also states that covariates are $i.i.d.$ unconditionally on $A^{(k)}$, but not necessarily conditionally.
Figure \ref{fig:network} provides an illustration.
Here, we condition on the indicators $1\{i_k \leftrightarrow j_k\}$ (which can differ across clusters) to control the network's maximum degree, but we do not condition on the network $A^{(k)}$. We can interpret such indicators as exogenously drawn from some arbitrary distribution.\footnote{Formally, $\mathcal{I}_k \sim \mathcal{P}_k, \quad (X_i^{(k)}, U_i^{(k)}) | \mathcal{I}_k \sim_{i.i.d.} F_{U|X} F_X, \quad A_{i,j}^{(k)} = l\left(X_i^{(k)}, X_j^{(k)}, U_i^{(k)}, U_j^{(k)}\right)1\{i_k \leftrightarrow j_k \}$, where $\mathcal{I}_k$ is the matrix of such indicators in cluster $k$ and $\mathcal{P}_k$ is a cluster-specific distribution left unspecified.}
Equation \eqref{eqn:network} states that the distribution of covariates and unobservables is the same across different clusters ($F_X, F_{X|U}$ do not depend on the cluster's identity). It implies that the clusters' networks are drawn from the same distribution. We now provide a micro-foundation to our model.








\begin{exmp}[Microfoundation with network model] \label{exmp:microfoundation} Consider the following restrictions:
\begin{itemize}
\item[(A)]  For $i \in \{1, \cdots, N\}, k \in \{1, \cdots, K\}$, let Equation \eqref{eqn:network} hold given the indicators $1\{i_k \leftrightarrow j_k\}$, for some unknown $l(\cdot)$; in addition, $\sum_{j = 1}^N 1\{i_k \leftrightarrow j_k\} = \gamma_N^{1/2}$.
\item[(B)] Suppose that for any $i,t,k, \mathbf{d}_s^{(k)} \in \{0,1\}^N, s \le t$
$$
\small
\begin{aligned}
Y_{i,t}^{(k)}(\mathbf{d}_1^{(k)}, \cdots, \mathbf{d}_{t}^{(k)}) = r\Big( \mathbf{d}_{i,t}^{(k)}, \mathbf{d}_{\mathcal{N}_i^{(k)},t}^{(k)}, X_i^{(k)},  X_{\mathcal{N}_i^{(k)}}^{(k)}, U_i, U_{\mathcal{N}_i^{(k)}}, A_{i, \cdot}^{(k)}, |\mathcal{N}_i^{(k)}|, \nu_{i,t}^{(k)}\Big) + \tau_k + \alpha_{t}
\end{aligned}
$$
where $\mathcal{N}_i^{(k)} = \{j: A_{i,j}^{(k)} > 0\}$, for some unknown $r(\cdot)$, symmetric in the argument $A_{i, \cdot}^{(k)}$ (but not necessarily in $(\mathbf{d}_{\mathcal{N}_{i}^{(k)},t}, X_{\mathcal{N}_i^{(k)}}, U_{\mathcal{N}_i^{(k)}})$), stationary (but possibly serially dependent) unobservables $\nu_{i,\cdot}^{(k)} | X^{(k)}, U^{(k)} \sim_{i.i.d.} P_{\nu}$, fixed effects $\tau_k, \alpha_t$.
\end{itemize}
\end{exmp}



Condition (A) states the following: before being born, each individual may interact with $\gamma_N^{1/2}$ many other individuals (i.e., maximum degree). After birth, the individual's gender, income, and parental status determine her type and the distribution of her and her potential connections' edges.\footnote{See \cite{jackson1996strategic}, \cite{li2020random} for pairwise interactions. Extensions where the networks also depend on non-separable shocks $\omega_{i,j}$ are possible, as discussed in previous versions of this draft.} Condition (B) states that potential outcomes depend on neighbors' assignments, observables, and unobservables. \textit{Heterogeneity} in spillovers occurs arbitrarily through neighbors' observables and unobservables $(D_j, U_j, X_j)$. Such variables can interact with each other, allowing for observed and unobserved heterogeneity in direct and spillover effects (i.e., $r(\cdot)$ is invariant to permutations of the entries of $A_{i, \cdot}^{(k)}$, $r(\cdot)$ is \textit{not} invariant in neighbors' observables and unobservables). Whereas treatments may exhibit individual-level heterogeneity, treatments do not interact with clusters' fixed effects.



\begin{prop}[Microfoundation with network spillovers] \label{lem:lem0} Consider treatments assigned as in Assumption \ref{defn:bernoulli}. Let (A) and (B) in Example \ref{exmp:microfoundation} hold. Then Assumption \ref{ass:ass_0} holds.
\end{prop}

The proof is in Appendix \ref{sec:lem_main}. Proposition \ref{lem:lem0} motivates Assumption \ref{ass:ass_0} in our leading example with network spillovers.






 \begin{figure}[!ht]
 \centering
 \vspace{-7mm}
    \begin{tikzpicture}




\coordinate (1) at (-10,1);
\coordinate (2) at (-2,1);
\coordinate (3) at (-2,5);
\coordinate (4) at (-10,5);



  \node[draw, circle] (a) at (-6, 3.2) {};
  \node[draw, circle] (b) at (-6, 4.1) {};
  \node[draw, circle] (c) at (-7, 3.5) {};
 \node[draw, circle] (d) at (-6.8, 2.5) {};
  \node[draw, circle] (e) at (-5.3, 2.5) {};
  \node[draw, circle] (f) at (-5, 3.5) {};
 \node[circle] (h) at (-5.1, 4.4) {};
 \node[circle] (i) at (-4.3, 2.9) {};
  \node[circle] (l) at (-4.3, 4) {};
   \node[circle] (m) at (-6.9, 4.4) {};
   \node[circle] (n) at (-7.6, 3) {};
   \node[circle] (o) at (-7.4, 2) {};
   \node[circle] (p) at (-7.7, 3.9) {};
     \node[circle] (q) at (-6, 4.8) {};
       \node[ circle] (r) at (-6, 1.9) {};
       \node[circle] (s) at (-4.8, 1.9) {};

 \node[circle] (g) at (-3,3) {$\longrightarrow$};
\node[circle] (g) at (3,3) {$\longrightarrow$};

 \node[circle] (g) at (-3,3) {$\longrightarrow$};

 \node[circle] (g) at (-6,1.5) {$\text{Possible connections}$};
  \node[circle] (g) at (0,1.5) {$\text{Types' assignment}$};
  \node[circle] (g) at (6,1.5) {$\text{Network formation}$};

  \node[draw, fill = green, circle] (aaa) at (0, 3.2) {};
  \node[draw, fill = green, circle] (bbb) at (0, 4.1) {};
  \node[draw, fill = blue, circle] (ccc) at (-1, 3.5) {};
 \node[draw, fill = blue, circle] (ddd) at (-0.8, 2.5) {};
  \node[draw, fill = blue, circle] (eee) at (0.7, 2.5) {};
  \node[draw, fill = blue, circle] (fff) at (1, 3.5) {};
 \node[circle] (hhh) at (0.9, 4.4) {};
 \node[circle] (iii) at (1.7, 2.9) {};
  \node[circle] (lll) at (1.7, 4) {};
   \node[circle] (mmm) at (-0.9, 4.4) {};
   \node[circle] (nnn) at (-1.6, 3) {};
   \node[circle] (ooo) at (-1.4, 2) {};
   \node[circle] (ppp) at (-1.7, 3.9) {};
     \node[circle] (qqq) at (0, 4.8) {};
       \node[ circle] (rrr) at (0, 1.9) {};
       \node[circle] (sss) at (1.2, 1.9) {};



  \node[draw, fill = green, circle] (aa) at (6, 3.2) {};
  \node[draw,fill = green,  circle] (bb) at (6, 4.1) {};
  \node[draw, fill = blue, circle] (cc) at (5, 3.5) {};
 \node[draw,  fill = blue,circle] (dd) at (5.2, 2.5) {};
  \node[draw, fill = blue, circle] (ee) at (6.7, 2.5) {};
  \node[draw,  fill = blue, circle] (ff) at (7, 3.5) {};
 \node[circle] (hh) at (6.9, 4.4) {};
 \node[circle] (ii) at (7.7, 2.9) {};
  \node[circle] (ll) at (7.7, 4) {};
   \node[circle] (mm) at (5.1, 4.4) {};
   \node[circle] (nn) at (4.4, 3) {};
   \node[circle] (oo) at (4.6, 2) {};
   \node[circle] (pp) at (4.3, 3.9) {};
     \node[circle] (qq) at (6, 4.8) {};
       \node[ circle] (rr) at (6, 1.9) {};
       \node[circle] (ss) at (7.2, 1.9) {};


    \draw[-, dashed] (a) edge (b) (a) edge (c) (a) edge (d) (a) edge (e) (a) edge (f);
    \draw[-] (aa) edge (bb) (aa) edge (ee);

 \draw[-, dashed] (b) edge (c) (c) edge (d) (d) edge (e) (e) edge (f) (f) edge (b);
 \draw[-, dashed] (f) edge (h) (f) edge (i) (e) edge (i) (b) edge (h) (f) edge (l);
\draw[-, dashed] (m) edge (c) (n) edge (c) (p) edge (c) (n) edge (d) (o) edge (d) (r) edge (d) (r) edge (e) (s) edge (e) (b) edge (q) (b) edge (m);

    \draw[-, dashed] (aaa) edge (bbb) (aaa) edge (ccc) (aaa) edge (ddd) (aaa) edge (eee) (aaa) edge (fff);

 \draw[-, dashed] (bbb) edge (ccc) (ccc) edge (ddd) (ddd) edge (eee) (eee) edge (fff) (fff) edge (bbb);
 \draw[-, dashed] (fff) edge (hhh) (fff) edge (iii) (eee) edge (iii) (bbb) edge (hhh) (fff) edge (lll);
\draw[-, dashed] (mmm) edge (ccc) (nnn) edge (ccc) (ppp) edge (ccc) (nnn) edge (ddd) (ooo) edge (ddd) (rrr) edge (ddd) (rrr) edge (eee) (sss) edge (eee) (bbb) edge (qqq) (bbb) edge (mmm);

 \draw[-] (bb) edge (cc) (dd) edge (ee) (ee) edge (ff) (ff) edge (bb);
 \draw[-] (ff) edge (ii)  (bb) edge (hh) (ff) edge (ll);
\draw[-] (mm) edge (cc) (nn) edge (cc) (pp) edge (cc);






    \end{tikzpicture}
    \vspace{-18mm}
  \caption{Example of the network formation model, with $\gamma_N = 5$.  Individuals are assigned different types, which may or may not be observed by the researcher (corresponding to different colors). Individuals interact based on their types and form links among the possible connections. The possible connections and the realized adjacency matrix remain unobserved.  }  \label{fig:network}
\end{figure}




\vspace{-4mm}

\section{Experimental designs} \label{sec:designs}




\subsection{Single-wave experiment: estimation and inference}  \label{sec:jh}

Next, we present the single-wave experiment in Algorithm \ref{alg:my_pilot}, a summary of the main estimators, and a brief description of the tests in Table \ref{tab:methods1}. Define the vector
\begin{equation} \label{eqn:ej}
\small
\begin{aligned}
\underline{e}_j  =
\begin{bmatrix}
    0, \cdots, 0, 1, 0, \cdots, 0
\end{bmatrix}, \text{ where }  \underline{e}_j \in \{0,1\}^p, \text{ and } \underline{e}_j^{(j)} = 1.
\end{aligned}
\end{equation}









\vspace{-4mm}
\paragraph{Algorithm description} Algorithm \ref{alg:my_pilot} presents the design. The algorithm pairs clusters into $G$ pairs. It estimates the marginal effect within each pair by inducing local perturbations $\eta_n$. It then aggregates information across pairs to construct a test statistic.

 For the sake of brevity, throughout the main text, we allow for arbitrary pairs in the design of Algorithm \ref{alg:my_pilot}. Without loss, we index clusters such that each pair contains two consecutive clusters $\{k, k+1\}$ with $k$ being an odd number.
  Pairing clusters may occur based on observed heterogeneity, omitted for brevity and formalized in Appendix \ref{app:het}.




\vspace{-4mm}

\paragraph{Null hypothesis and inference} Let $\beta^* \in \mathcal{B}$ be an interior point. If $W(\beta) = W(\beta^*)$, then
\begin{equation} \label{eqn:h0}
\small
\begin{aligned}
H_0: M^{(j)}(\beta) = 0, \quad \forall j \in \{1, \cdots, p_1\},  p_1 \le p.
\end{aligned}
 \end{equation}

The above implication is at the core of the proposed approach.
We can test whether $p_1$ arbitrary entries of the marginal effect are equal to zero. Rejection implies a lack of global optimality. For expositional convenience, we consider $p_1 = 1$ only as in our application. In Appendix \ref{sec:pilot_general}, we show how the proposed method generalizes to $p_1 > 1$. We may also consider one sided tests $M^{(j)}(\beta) \le 0$; for example, for $\pi(x, \beta) = \beta_x$ (with $\mathcal{X}$ discrete), the one-sided test is informative for whether treatment probabilities for individuals with $x = j$ should be increased (without assuming that $\beta^*$ is in the interior). The critical value for the test for $H_0$ in \eqref{eqn:h0} is obtained by permuting the sign of each pair's estimated marginal effect in the spirit of \cite{canay2017randomization}, and recomputing the test statistic in Equation \eqref{eqn:test_stat} across the different permutations. Corollary \ref{thm:inference4b} and Appendix \ref{sec:perm_tests} present a formalization.

Finally, we recommend researchers to report $\bar{M}_n(\beta)$ (Equation \ref{eqn:test_stat}) in their results -- the average estimated marginal effect across clusters' pairs. Section \ref{sec:theory_exp} (and Appendix \ref{sec:pilot_general} for $p > 1$) shows that $\bar{M}_n(\beta)$ consistently estimate $M^{(1)}(\beta)$ as $n \rightarrow \infty$, $G$ is finite.




















\begin{table}[!htbp] \centering
  \scalebox{0.65}{
\begin{tabular}{@{\extracolsep{5pt}} ccccc}
\\[-1.8ex]\hline
\hline \\[-1.8ex]
Estimand & Estimand's Description & Estimator  & Estimator's Description & Randomization Inference \\
\hline \\[-1.8ex]
$M(\beta)$  & Marginal effect & $\bar{M}_n(\beta) = \frac{1}{G} \sum_{g=1}^G \widehat{M}_g(\beta)$  & DiD estimators  & Permute signs of $\widehat{M}_g(\beta)$ \\
& & $\widehat{M}_g(\beta)$ as in Eq \eqref{eqn:gradient_main} & from each clusters' pairs $g$ & for each clusters' pair $g$ \\
 \\
$\Delta(\beta)$ & Direct effect & $\bar{\Delta}_n = \frac{1}{G} \sum_g \hat{\Delta}_g(\beta)$ &  Pooled IPW estimators & Permute sign of estimated effect \\
& & $\hat{\Delta}_g$ as in Eq \eqref{eqn:direct_estimated}  & from each cluster & for each cluster $k$\\  \\
$S(0,\beta)$ & Marginal spillovers & $\bar{S}_n(0, \beta) = \frac{1}{G} \sum_g \widehat{S}_g(0, \beta)$ & DiD + IPW estimators & Permute sign of $\widehat{S}_g(0, \beta)$ \\
& on controls & $\widehat{S}_g(0, \beta)$ as in Eq \eqref{eqn:estimated_spill} & from each clusters' pairs $g$ & for each clusters' pair $g$ \\
 \\
 $S(1,\beta)$ & Marginal spillovers & $\bar{S}_n(1, \beta) = \frac{1}{G} \sum_g \widehat{S}_g(1, \beta)$ & DiD + IPW estimator & Permute sign of  $\widehat{S}_g(1, \beta)$\\
& on treated & & from each clusters' pairs $g$  & for each clusters' pair $g$ \\ \\
$W(\beta)$ & Welfare at $\beta$ & $\bar{W}_n(\beta) = \frac{1}{K} \sum_{k=1}^K \Big[\bar{Y}_1^{(k)} - \bar{Y}_0^{(k)}\Big]$ & Pooled pre-post difference & Permute sign for each cluster\\
\hline \\[-1.8ex]
\end{tabular}
}
 \caption[Caption for LOF]{Estimands and estimators from single wave experiment with treatment probability $\beta$. Inference procedure is formally described in Corollary \ref{thm:inference4b} (and Appendix \ref{sec:perm_tests})}
  \label{tab:methods1}
\end{table}


\begin{algorithm} [!ht]   \caption{One-wave experiment for inference with $p_1 = 1$}\label{alg:my_pilot}
    \begin{algorithmic}[1]
    \Require Value $\beta \in \mathbb{R}^p$ (exogenous), $K$ clusters, constant $\bar{C}$, size $\alpha$;
    
     \stepcounter{algsubstate}
     \Statex {\footnotesize\alph{algsubstate}:}\space Organize clusters into $G = K/2$ pairs with consecutive indexes $\{k, k+1\}$;

    
     \stepcounter{algsubstate}
     \Statex {\footnotesize\alph{algsubstate}:}\space $t = 0$ (baseline): either nobody receives treatments or treatments are assigned with $\pi(\cdot;\beta)$ (either case is allowed).
    \begin{algsubstates}
        
     \stepcounter{algsubstate}
     \Statex {\footnotesize\alph{algsubstate}:}\space Experimenters collect baseline outcomes: for $n$ units in each cluster observe $Y_{i,0}^{(h)}, X_i^{(h)}, h  \in \{1, \cdots, K\}$.
        \end{algsubstates}

    
     \stepcounter{algsubstate}
     \Statex {\footnotesize\alph{algsubstate}:}\space $t = 1$: experiment starts
    \begin{algsubstates}
        
     \stepcounter{algsubstate}
     \Statex {\footnotesize\alph{algsubstate}:}\space  For each pair $g = \{k, k+1\}$, randomize
   \begin{equation} \label{eqn:perturbation}
     \small
     \begin{aligned}
    D_{i,1}^{(k)} | \beta, X_i^{(k)} = x \sim \begin{cases}
    & \mathrm{Bern}(\pi(x, \beta + \eta_n\underline{e}_1)) \text{ if } h = k  \\
    & \mathrm{Bern}(\pi(x, \beta - \eta_n\underline{e}_1)) \text{ if } h = k+1
    \end{cases} , \quad \bar{C} n^{-1/2} < \eta_n < \bar{C} n^{-1/4},
    \end{aligned}
    \end{equation}

        
     \stepcounter{algsubstate}
     \Statex {\footnotesize\alph{algsubstate}:}\space For $n$ units in each cluster $h$ observe $Y_{i,1}^{(h)}$;
        
     \stepcounter{algsubstate}
     \Statex {\footnotesize\alph{algsubstate}:}\space Estimate the marginal effect as in Equation \eqref{eqn:marginal}.
      \end{algsubstates}

 
     \stepcounter{algsubstate}
     \Statex {\footnotesize\alph{algsubstate}:}\space Construct  the t-statistic to test $H_0$ in Equation \eqref{eqn:h0} (with $j = 1$)
\begin{equation} \label{eqn:test_stat}
\small
\begin{aligned}
 \mathcal{T}_n = \frac{\sqrt{G} \bar{M}_n(\beta)}{\sqrt{(G - 1)^{-1}\sum_{g} (\widehat{M}_g(\beta) - \bar{M}_n(\beta))^2}}, \quad \bar{M}_n(\beta) = \frac{1}{G} \sum_g \widehat{M}_g(\beta);
 \end{aligned}
\end{equation}
here, $\widehat{M}_g$ is the marginal effect estimated in pair $g$ as in Equation \eqref{eqn:gradient_main}.

     \stepcounter{algsubstate}
     \Statex {\footnotesize\alph{algsubstate}:}\space Construct tests $1\Big\{|\mathcal{T}_n| > \mathrm{cv}_{G}(\alpha) \Big\}$ with size $\alpha$, with critical values obtained by permuting the sign of the estimated marginal effect as described in Corollary \ref{thm:inference4b} (and Appendix \ref{sec:perm_tests}).


         \end{algorithmic}
\end{algorithm}


\paragraph{Other effects identified by the experiment} Algorithm \ref{alg:my_pilot} also allows us to estimate the direct effect of the treatment, the (marginal) spillover effect separately, and the welfare respectively under Assumption \ref{ass:regularity_basic} below
$$
\small
\begin{aligned}
\Delta(\beta) = \int \left[m(1, x, \beta) - m(0, x, \beta)\right] dF_X(x), \quad S_1(d, \beta)= \int \frac{\partial m(d, x, \beta)}{\partial \beta^{(1)}} dF_X(x), \quad W(\beta).
\end{aligned}
$$
The direct effect is the treatment effect, keeping fixed the neighbors' treatment probability. $S_1(\cdot)$, the spillover effect, is the marginal effect of a small change in the first entry of $\beta$ (e.g., the neighbors' treatment probability), keeping fixed individual treatment status. Our framework also extends to estimating $S_j(\cdot)$ for arbitrary entries of $\beta$ as in Appendix \ref{sec:pilot_general}. For a given pair of clusters $(k, k+1)$, we estimate
\begin{equation} \label{eqn:direct_estimated}
\small
\begin{aligned}
\widehat{\Delta}_{(k, k+1)}(\beta) & =  \frac{1}{2 n} \sum_{h \in \{k, k+1\}} \sum_{i=1}^n \left[\frac{D_{i,1}^{(h)} Y_{i,1}^{(h)}}{\pi(X_i^{(h)}, \beta + \eta_n v_h \underline{e}_1)} - \frac{(1 - D_{i,1}^{(h)}) Y_{i,1}^{(h)}}{1 - \pi(X_i^{(h)}, \beta + \eta_n v_h \underline{e}_1)}\right], \quad v_h = \begin{cases}  1 \text{ if } h = k \\
-1 \text{ if } h = k + 1.
\end{cases}
\end{aligned}
\end{equation}
The estimator pools observations between the two clusters and takes a difference between treated and control units within each cluster, divided by the probability of treatments as in \cite{horvitz1952generalization}. We average direct effects across clusters' pairs to obtain a single estimate $\bar{\Delta}_n = \frac{1}{G} \sum_{g} \widehat{\Delta}_g(\beta)$. The indirect effect is estimated as follows:
\begin{equation} \label{eqn:estimated_spill}
\small
\begin{aligned}
\widehat{S}_{(k, k+1)}(0, \beta) = \frac{1}{2 n} \sum_{h \in \{k, k+1\}} \frac{v_h}{\eta_n} \sum_{i=1}^n  \left[ \frac{Y_{i,1}^{(h)} (1 - D_{i,1}^{(h)})}{1 - \pi(X_i^{(h)}, \beta + v_h \eta_n \underline{e}_1)} - \bar{Y}_0^{(h)}\right].
\end{aligned}
\end{equation}
The estimator takes a weighted difference between the two clusters' control units. Researchers can report the between-pairs average $\bar{S}_n(0, \beta) = \frac{1}{G} \sum_{g} \widehat{S}_g(0, \beta)$ (and similarly $\hat{S}(1, \beta)$ for treated units), which captures spillovers on the control units.

Researchers may also be interested in estimating welfare effects at a given $\beta$, $W(\beta)$, \textit{pooling} information across clusters, using as an estimator
$
\bar{W}_n(\beta) = \frac{1}{K} \sum_{k=1}^K \Big[\bar{Y}_1^{(k)} - \bar{Y}_0^{(k)}\Big].
$


Inference on each of these estimands can be conducted through permutation tests, see Table \ref{tab:methods1}.  Theorem \ref{thm:bias_direct} provides guarantees such that the bias arising from pooling for the direct and welfare effect is negligible for inference.






\begin{rem}[Choice of $\eta_n$] \label{rem:eta1} The choice of the perturbation $\eta_n$ must balance the bias and variance of the estimator as discussed in Theorem \ref{thm:const1}.   Appendix \ref{app:rule_thumb} provides a rule of thumb. It is also possible to choose different levels of perturbations in the same design by assigning a group of clusters to a treatment probability $\beta + \eta_n$, a different group to $\beta -\eta_n$; within each group, then repeat this same procedure, inducing more minor perturbation of order $\beta + \eta_n \pm \eta_n', \eta_n' = o(\eta_n)$, and similarly for the second group. These two nested designs do not affect our theoretical results for inference on $M(\beta)$, as long as $\eta_n' = o(\eta_n)$. Choosing different levels of perturbations may allow learning a larger set of marginal effects (both at $\beta$, and at  $\beta \pm \eta_n$), while avoiding under-powered studies for the main effect $M(\beta)$.  \qed
\end{rem}






\subsection{Multi-wave experiment: welfare maximization} \label{sec:main_design}


Next, we discuss the multi-wave experiment.
For illustrative purposes, we provide the algorithm for the one-dimensional case $p = 1$, in Algorithm \ref{alg:adaptive}, that is, when $\beta \in \mathcal{B} = [\mathcal{B}_1, \mathcal{B}_2]$ is a scalar. In Remark \ref{rem:general} and formally in Appendix \ref{app:algorithms}, we provide the complete algorithm for the $p$-dimensional case. Let $\hat{M}_{k, t}$ be as in Equation \eqref{eqn:defn1} for $k$ odd.



  \begin{algorithm} [!h]   \caption{Multiple-wave experiment with $\beta$ scalar}\label{alg:adaptive}
    \begin{algorithmic}[1]
    \Require Starting value $\beta_0$, $K$ clusters, $T + 1$ periods, constant $\bar{C}$.
    
     \stepcounter{algsubstate}
     \Statex {\footnotesize\alph{algsubstate}:}\space   Create pairs of clusters $\{k, k+1\}, k \in \{1, 3, \cdots, K-1\}$;
    
     \stepcounter{algsubstate}
     \Statex {\footnotesize\alph{algsubstate}:}\space $t = 0$ (initialization):
    \begin{algsubstates}
        
     \stepcounter{algsubstate}
     \Statex {\footnotesize\alph{algsubstate}:}\space  Assign treatments as
    $
    D_{i,0}^{(h)} | X_i^{(h)} = x \sim \mathrm{Bern}(\pi(x, \beta_0)) \text{ for all } h \in \{1, \cdots, K\}.
    $
     
     \stepcounter{algsubstate}
     \Statex {\footnotesize\alph{algsubstate}:}\space For $n$ units in each cluster observe $Y_{i,0}^{(h)}, h \in \{1, \cdots, K\}$; initalize $\widehat{M}_{k,t} = 0$, $\check{\beta}_{k}^0 = \beta_0$.
        \end{algsubstates}
    \While{$1 \le t \leq T$}
    \begin{algsubstates}
    
     \stepcounter{algsubstate}
     \Statex {\footnotesize\alph{algsubstate}:}\space Define
    $$
    \small
    \begin{aligned}
    \check{\beta}_{h}^t = \begin{cases} & P_{\mathcal{B}_1, \mathcal{B}_2 - \eta_n}\Big[\check{\beta}_{h}^{t-1} + \alpha_{h + 2, t}\widehat{M}_{h + 2,t - 1}\Big], \quad h \in \{1, \cdots, K - 2\},  \\
    & P_{\mathcal{B}_1, \mathcal{B}_2 - \eta_n}\Big[\check{\beta}_{h}^{t-1} + \alpha_{1, t} \widehat{M}_{1,t - 1}\Big], \quad \quad \quad h \in \{K - 1, K\};
    \end{cases}
    \end{aligned}
    $$
    where $\alpha_{k,t}$ is the learning rate $P_{a, b}(x) = \mathrm{arg} \min_{x' \in [a, b]^p} ||x - x'||^2$.
        
     \stepcounter{algsubstate}
     \Statex {\footnotesize\alph{algsubstate}:}\space  Assign treatments as (for  $\bar{C} n^{-1/2}< \eta_n < \bar{C} n^{-1/4}$)
   \begin{equation} \label{eqn:rand_adaptive}
   \small \begin{aligned}
    D_{i,t}^{(h)} | X_i^{(h)} = x \sim
    \mathrm{Bern}(\pi(x, \beta_{h,t})), \quad  \beta_{h,t} = \begin{cases} & \check{\beta}_{h}^t  + \eta_n \text{ if } h \text{ is odd}  \\
    & \check{\beta}_{h}^t  - \eta_n \text{ if } h \text{ is even}
    \end{cases}
    \end{aligned}
    \end{equation}
        
     \stepcounter{algsubstate}
     \Statex {\footnotesize\alph{algsubstate}:}\space For $n$ units in each cluster $h \in \{1, \cdots, K\}$ observe $Y_{i,t}^{(h)}$;
        
     \stepcounter{algsubstate}
     \Statex {\footnotesize\alph{algsubstate}:}\space For each pair $\{k, k+1\}$, estimate
      \begin{equation} \label{eqn:defn1}
      \small
      \begin{aligned}
      \hat{M}_{k,t} =
      \hat{M}_{k + 1,t} = \frac{1}{2 \eta_n} \Big[\bar{Y}_t^{(k)} - \bar{Y}_0^{(k)}\Big] -  \frac{1}{2 \eta_n} \Big[\bar{Y}_t^{(k + 1)} - \bar{Y}_0^{(k +1)}\Big].
      \end{aligned}
      \end{equation}
        \EndWhile
      \end{algsubstates}
 
     \stepcounter{algsubstate}
     \Statex {\footnotesize\alph{algsubstate}:}\space Return
 $
 \hat{\beta}^* = \frac{1}{K} \sum_{k = 1}^K \check{\beta}_{k}^T
 $

         \end{algorithmic}
\end{algorithm}

 The algorithm pairs clusters (here two consecutive clusters form a pair) and initializes clusters at the same starting value $\beta_0$, $\check{\beta}_1^1 = \cdots = \check{\beta}_K^1 = \beta_0$. At $t = 0$, it randomizes treatments independently using the same starting value $\beta_0$ for all clusters.
 Here, $\beta_0$ is chosen exogenously, e.g., it is the current policy in place. Over each iteration $t$, we assign treatments based on $\beta_{k,t}$  for cluster $k$ at time $t$, which equals the parameter $\check{\beta}_k^t$ obtained from a previous iteration plus a positive (negative) perturbation $\eta_n$ in the first (second) cluster in a pair. The local perturbation follows similarly to what is discussed in the single-wave experiment. Also, by construction, $\check{\beta}_k^t$ is the same for a given pair $(k, k+1)$, where $k$ is odd. We choose $\check{\beta}_k^{t + 1}$ via \textit{sequential cross-fitting}: we wrap clusters in a \textit{circle} and update the parameter in a pair of clusters $(k, k+1)$ using information from the subsequent pair (see Figure \ref{fig:clusters}). The algorithm runs over $T$ periods and returns
$
\hat{\beta}^* = \frac{1}{K} \sum_{k = 1}^K \check{\beta}_{k}^{T + 1} .
$
Choosing the average is motivated by the theoretical properties of gradient descent, although other statistics are also possible.










\begin{lem}[Unconfoundedness] \label{lem:1a} Let $T/p +  1\le K/2$. Consider the experimental design in Algorithm \ref{alg:adaptive_complete} for generic $p$-dimensions (and Algorithm \ref{alg:adaptive} for $p = 1$). Then, for any $k$,
$$
\small
\begin{aligned}
\Big(\beta_{k,1}, \cdots, \beta_{k, T} \Big) \perp \Big\{Y_{i,t}^{(k)}(\mathbf{d}), X_i^{(k)}, \mathbf{d} \in \{0,1\}^N\Big\}_{i \in \{1, \cdots, N\}, t \le T}.
\end{aligned}
$$
\end{lem}


The proof is in Appendix \ref{sec:lemma_cf}.
Lemma \ref{lem:1a} shows that the parameters used in the experiment are independent of potential outcomes and covariates in the same cluster. Namely, the sequential cross-fitting breaks the dependence due to repeated sampling, which would otherwise confound the experiment. The main distinction from most of the previous literature on adaptive experiments \citep[e.g.][]{kasy2019adaptive, wager2019experimenting, hadad2019confidence, zhang2020inference} is that in all such references repeated sampling does not occur, and batches are independent each period. Here, instead, clusters are dependent over each period, motivating our sequential estimation procedure. Also, note that existing cross-fitting procedures that would instead use all pairs except the current pair of individual $i$ for a policy update would also have a confounding bias whenever $T > 2$ (see Appendix \ref{sec:lemma_cf}).





















\begin{figure}[!h]
\centering
    \begin{tikzpicture}[scale = 1.16]
    \node[draw, black,ultra thick, inner sep=0pt,
  text width=12mm,
  align=center,   circle] (h) at (-3,-2) {$Y_{i,t-1}^{(k)}$};
    \node[draw, black,ultra thick,  inner sep=0pt,
  text width=12mm,
  align=center, circle] (d) at (-3,-4) {$\beta_{k,t}$};
    \node[draw, black,ultra thick,inner sep=0pt,
  text width=12mm,
  align=center, circle] (b) at (1,0) {$\nu_{i,t}^{(k)}$};

    \node[] (e) at (-8,-5) {Policy on a new population};
   \node[] (e) at (-1,-5) {Experiment with repeated sampling};
    \node[draw, black,ultra thick, inner sep=0pt,
  text width=12mm,
  align=center, circle] (e) at (1,-2) {$Y_{i,t}^{(k)}$};
 \node[draw, black,ultra thick,  inner sep=0pt,
  text width=12mm,
  align=center, circle] (f) at (1,-4) {$D_{i,t}^{(k)}$};
 \node[draw, black,ultra thick,  inner sep=0pt,
  text width=12mm,
  align=center, circle] (g) at (-3,0) {$\nu_{i,t - 1}^{(k)}$};

    \node[draw, black,ultra thick, inner sep=0pt,
  text width=12mm,
  align=center,   circle] (hh) at (-10,-2) {$Y_{i,t-1}^{(k)}$};
    \node[draw, black,ultra thick,  inner sep=0pt,
  text width=12mm,
  align=center, circle] (dd) at (-10,-4) {$\beta^*$};
    \node[draw, black,ultra thick,inner sep=0pt,
  text width=12mm,
  align=center, circle] (bb) at (-6,0) {$\nu_{i,t}^{(k)}$};

    \node[draw, black,ultra thick, inner sep=0pt,
  text width=12mm,
  align=center, circle] (ee) at (-6,-2) {$Y_{i,t}^{(k)}$};
 \node[draw, black,ultra thick,  inner sep=0pt,
  text width=12mm,
  align=center, circle] (ff) at (-6,-4) {$D_{i,t}^{(k)}$};
 \node[draw, black,ultra thick,  inner sep=0pt,
  text width=12mm,
  align=center, circle] (gg) at (-10,0) {$\nu_{i,t - 1}^{(k)}$};




    \draw[->, -triangle 90]    (g) edge (h)  (h) edge (d)
    (d) edge (f) (f) edge (e) (b) edge (e) (g) edge (b);

   \draw[->, -triangle 90] (dd) edge (hh)   (gg) edge (hh)
    (dd) edge (ff) (ff) edge (ee) (bb) edge (ee) (gg) edge (bb);

    \end{tikzpicture}
    \caption{The left panel shows the dependence structure when a static policy is implemented on a new population (I omit $D_{i,t-1}^{(k)}$ for expositional convenience), where $\nu_{i,t}$ denote unobservable characteristics. The right panel shows the dependence structure of a sequential experiment that uses the same units for policy updates over subsequent periods with \textit{repeated} sampling.
} \label{fig:dag}
    \end{figure}





\begin{figure}[!ht]
\vspace{-3mm}
\centering
\begin{tikzpicture}[
node distance = 6mm and 6mm,
  start chain = going right,
    mw/.style = {minimum width=#1},
  list/.style = {rectangle split, rectangle split parts=1,
                 rectangle split horizontal, draw,
                 align=center,
                 text width=7mm,
                 minimum height=9mm,
                 inner sep=1mm, on chain},
          > = stealth,
          ]

  \node[list] (A) {\nodepart[mw=7mm]{two}  };
  \node[list] (B) {\nodepart[mw=7mm]{two}  };
  \node[list] (C) {\nodepart[mw=7mm]{two} 3};

  \node[list, below=of A] (D) {\nodepart[mw=7mm]{two}  };
  \node[list] (E) {\nodepart[mw=7mm]{two}  };
  \node[list] (F) {\nodepart[mw=7mm]{two} 3};
 \node[circle] (g) at (4.5,-0.8) {$\longrightarrow$};


    \node[list, fill = cyan] (A) at (5, 0) {\nodepart[mw=7mm]{two}  };
  \node[list, fill = cyan] (B) {\nodepart[mw=7mm]{two}  };
  \node[list, fill = cyan] (C) {\nodepart[mw=7mm]{two} 3};

  \node[list, below=of A, fill = gray] (D) {\nodepart[mw=7mm]{two}  };
  \node[list, fill = gray] (E) {\nodepart[mw=7mm]{two}  };
  \node[list, fill = gray] (F) {\nodepart[mw=7mm]{two} 3};

   \node[circle] (g) at (10.5,-0.8) {$\longrightarrow$};
  \draw[-] (A) edge (D) (B) edge (E) (C) edge (F);
  \node[list, fill = cyan] (A) at (11, 0) {\nodepart[mw=7mm]{two}  };
  \node[list, fill = cyan] (B) {\nodepart[mw=7mm]{two}  };
  \node[list, fill = cyan] (C) {\nodepart[mw=7mm]{two} 3};

  \node[list, below=of A, fill = gray] (D) {\nodepart[mw=7mm]{two}  };
  \node[list, fill = gray] (E) {\nodepart[mw=7mm]{two}  };
  \node[list, fill = gray] (F) {\nodepart[mw=7mm]{two} 3};

  \draw[-] (A) edge (D) (B) edge (E) (C) edge (F);

  \path[->,  -triangle 90] (A) edge [bend right = 300]  (B);
   \path[->,  -triangle 90] (B) edge [bend right = 300]  (C);
    \path[->,  -triangle 90] (F) edge [bend right = -90]  (D);
\end{tikzpicture}
\vspace{-4mm}
\caption{Sequential cross-fitting method. Clusters (rectangles) are paired. Within each pair, researchers assign different treatment probabilities to clusters with different colors. Finally, the policy in each pair is updated using information from the consecutive pair. Note that because $K > 2 T$, the algorithm never ``circles back" to the initial pair.
 } \label{fig:clusters}
\end{figure}



\begin{rem}[$p$-dimensional case: Algorithm \ref{alg:adaptive_complete}]  \label{rem:general}
The algorithm for the $p$-dimensional case follows similarly to the uni-dimensional case with a minor change: we consider $T/p$ many \textit{waves}/iterations, each consisting of $p$ periods. Within each wave $w$, every period, we perturb a single coordinate of $\check{\beta}_k^w$, compute the marginal effect for that coordinate, and repeat over all coordinates $j \in \{1, \cdots p\}$ before making the next policy update to select $\check{\beta}_k^{w+1}$. \qed
\end{rem}

\begin{rem}[Learning rate] \label{rem:learning_rate} We are now left to discuss how ``large" the step size should be: if the marginal effect is positive, by how much should we increase the treatment probability? Assuming strong concavity of the objective function, the learning rate $\alpha_{k,t}$ should be of order $1/t$ (e.g., $J/t$). When $\beta$ denotes a treatment probability a natural choice is $J \in [10\%, 20\%]$.  A more robust choice with moderate or large $T$ (see Theorem \ref{thm:rate}) is
\begin{equation} \label{eqn:gradient}
\small
\begin{aligned}
\alpha_{k,t} = \begin{cases}
&  \frac{J}{T^{1/2 - v/2} ||\hat{M}_{k,t}||} \text{ if } ||\hat{M}_{k,t}||_2^2 >  \frac{\kappa}{\check{T}^{1 - v}} -  \epsilon_n, \\
&0\text{ otherwise}
\end{cases} ,
\end{aligned}
\end{equation}
for a positive $\epsilon_n$, $\epsilon_n \rightarrow 0$, and small constants $1 \ge v, J, \kappa > 0$.\footnote{Formally, we let   $\epsilon_n$ be proportional to $\sqrt{\frac{\gamma_N}{\eta_n^2 n}} + \eta_n$. See Theorem \ref{thm:rate} for more details. }
Here, the learning rate divides the estimated marginal effect by its norm (known as gradient norm rescaling, \citealt{hazan2015beyond}) and guarantees control of the out-of-sample regret under strict quasi-concavity. This choice is appealing because it guarantees comparable step sizes between different clusters.  \qed
\end{rem}


\begin{rem}[Why sequential cross-fitting?] \label{rem:aaa}
Next, we illustrate the source of bias if the sequential cross-fitting was not employed. Every period, the researcher can only identify the expected outcome of $Y_{i,t}^{(k)}$ conditional on the parameter $\beta_{k,t}$, namely
$
\widetilde{W}(\beta_{k,t}) = \mathbb{E}_{\beta_{k,t}}[Y_{i,t}^{(k)} | \beta_{k,t}].
$
If $\beta_{k,t}$ were chosen exogenously, based on information from a different cluster, $\mathbb{E}_{\beta_{k,t}}[Y_{i,t}^{(k)} | \beta_{k,t}] = \mathbb{E}_{\beta_{k,t}}[Y_{i,t}^{(k)}] = W(\beta_{k,t})$, where $W(\beta_{k,t})$ defines the expected welfare once we deploy the policy $\beta_{k,t}$ on a new population. However, the equality conditional and unconditional on $\beta_{k,t}$ does not occur when $\beta_{k,t}$ is estimated using information on $Y_{i,t-1}^{(k)}$. Consider the example where the outcome depends on some auto-correlated unobservables $\nu_{i,t}$ and treatment assignments in Figure \ref{fig:dag}.
The \textit{dependence} structure of Figure \ref{fig:dag} implies:
$
W(\beta_{k,t}) = \mathbb{E}_{\beta_{k,t}}[Y_{i,t}^{(k)}] \neq \mathbb{E}_{\beta_{k,t}}[Y_{i,t}^{(k)} | \beta_{k,t}] = \widetilde{W}(\beta_{k,t}),
$
if $\beta_{k,t}$ depends on covariates and unobservables previous outcomes (and so on unobservables $\nu_{i,t}^{(k)}$) in cluster $k$.
Here, $W(\beta_{k,t}) $ captures the estimand of interest. Instead, $\widetilde{W}(\beta_{k,t})$ denotes what we can identify.
The proposed algorithm breaks such dependence and guarantees unconfounded experimentation. \qed
\end{rem}

\vspace{-4mm}
\section{Theoretical guarantees}

Next, we turn to the theoretical guarantees to study properties of the design. Practitioners only interested in the implementation of the experiment may skip this section.

\subsection{Single wave experiment: consistency and inference} \label{sec:theory_exp}



\begin{ass}[Regularity 1] \label{ass:regularity_basic} Suppose that for all $x \in\mathcal{X}, d \in \{0,1\}$, $\pi(x, \beta)$, and $m(d, x, \beta)$ are uniformly bounded and twice differentiable with bounded derivatives.
\end{ass}

Assumption \ref{ass:regularity_basic} imposes smoothness and boundedness restrictions. These restrictions hold for a large set of linear and non-linear functions, assuming that $\mathcal{X}$ is compact. Boundedness is often imposed in the literature \citep[e.g.,][]{KitagawaTetenov_EMCA2018}.





\begin{thm}[Marginal effects] \label{thm:const1} Suppose that $Y_{i,t}^{(k)}$ is sub-Gaussian. Let  Assumptions \ref{ass:ass_0}, \ref{ass:regularity_basic} hold. Let $\mathrm{Var}(\sqrt{n} \hat{M}_{(k, k+1)}(\beta)) \le \tilde{C}_{k , k+1}\rho_n$, for arbitrary $\rho_n$ and constant $\tilde{C}_{k, k+1}$. Then, with probability at least $1 - \delta$, for any $\delta \in (0,1)$, for a finite constant $c_0 < \infty$ independent of $(n, N, \gamma_N, K, \beta)$,
$$
\small
\begin{aligned}
 \Big|\hat{M}_{(k, k+1)}(\beta) - M^{(1)}(\beta)\Big| \le c_0 \Big(\eta_n +  \min\Big\{\sqrt{\frac{\gamma_N \log(\gamma_N/\delta)}{n \eta_n^2}},  \sqrt{\frac{\tilde{C}_{k, k+1} \rho_n}{n \eta_n^2 \delta}}\Big\}  \Big),
 \end{aligned}
$$
where $\hat{M}_{(k, k+1)}$ is estimated as in Algorithm \ref{alg:my_pilot}.

For $\gamma_N \log(\gamma_N)/N^{1/3} = o(1), \eta_n = n^{-1/3}, \hat{M}_{(k, k+1)}(\beta) \rightarrow_p M^{(1)}(\beta), \bar{M}_n \rightarrow_p M^{(1)}(\beta)$.
\end{thm}

The proof is in Appendix \ref{sec:proof1}.
Theorem \ref{thm:const1} shows one can consistently estimate the marginal effects with two large clusters. Consistency depends on the degree of dependence among potential outcomes (which also depends on neighbors' treatments).
Once we interpret $\gamma_N^{1/2}$ as the maximum degree of a network (see Example \ref{exmp:microfoundation}),
the convergence rate depends on the \textit{minimum} between the maximum degree of the network, which is proportional to $\gamma_N^{1/2}$, and the covariances among unobservables, captured by $\rho_n$. The theorem also illustrates the trade-off in the choice of the deviation parameter $\eta_n$: a larger parameter $\eta_n$ decreases the variance, but it increases the bias (motivating our rule of thumb in Appendix \ref{app:rule_thumb}).









\begin{ass}[Regularity 2] \label{ass:var} Assume that for treatments as assigned in Algorithm \ref{alg:my_pilot}, for all $k \in \{1, \cdots, K\}$, $Y_{i,t}^{(k)}$ has a bounded fourth moment, and for some $\bar{C}_k > 0$, $\rho_n \ge 1$,
\begin{equation} \label{eqn:variance}
\small
\begin{aligned}
\mathrm{Var}\left(\sqrt{n}\Big[\bar{Y}_1^{(k)} - \bar{Y}_0^{(k)}\Big]\right) = \bar{C}_k \rho_n.
\end{aligned}
\end{equation}

\end{ass}




Assumption \ref{ass:var} imposes standard moment bounds and a \textit{lower bound} on the variance of the estimator. In particular, Assumption \ref{ass:var} states that the variance does not converge to zero at a rate faster than $1/n$. To gain further intuition, note that
\begin{equation} \label{eqn:aaaaa}
\small
\begin{aligned}
\bar{C}_k \rho_n = \frac{1}{n} \sum_{i=1}^n \mathrm{Var}\left(Y_{i, 1}^{(k)} - Y_{i,0}^{(k)}\right) +  \frac{1}{n} \sum_{i, j, j\neq i} \mathrm{Cov}\left(Y_{i,1}^{(k)} - Y_{i, 0}^{(k)}, Y_{j,1}^{(k)} - Y_{j, 0}^{(k)}\right).
\end{aligned}
\end{equation}
Assumption \ref{ass:var} is stating that $\rho_n \ge 1$, i.e., $\rho_n$ does not converge to zero. This requires that the negative covariance components (if any) do not outweigh the variances in Equation \eqref{eqn:aaaaa}, holding with no or positive correlations and guarantees that the variance is not zero.




  \begin{thm} \label{thm:inference1}  Let  Assumptions \ref{ass:ass_0}, \ref{ass:regularity_basic}, \ref{ass:var} hold. Let $n^{1/4} \eta_n = o(1), \gamma_N/N^{1/4} = o(1)$, $K < \infty$. Then, for each pair $(k, k+1)$, for $\widehat{M}_{(k, k+1)}$ estimated as in Algorithm \ref{alg:my_pilot},
$$
\small
\begin{aligned}
&\mathrm{Var}\Big(\widehat{M}_{(k, k+1)}\Big)^{-1/2} \Big(\widehat{M}_{(k, k+1)} - M^{(1)}(\beta)\Big)  \rightarrow_d \mathcal{N}(0,1).
\end{aligned}
$$
  \end{thm}

  The proof is in Appendix \ref{sec:inference_thm}.
  Theorem \ref{thm:inference1} guarantees asymptotic normality. The theorem assumes that $\gamma_N$ grows at a slower rate than the sample size of order $N^{1/4}$ (and hence $n^{1/4}$ because $n$ is proportional to $N$). This condition is stronger than what is required for consistency only.\footnote{We conjecture that weaker restrictions on the degree are possible. We leave their study to future research. } Given Theorem \ref{thm:inference1}, it is possible to conduct inference on the null in Equation \eqref{eqn:h0} by using either a $t$-student distribution for critical values as in \cite{ibragimov2010t} (see Theorem \ref{thm:inference4}), or using randomization tests in \cite{canay2017randomization}.








\begin{cor}[Randomization tests] \label{thm:inference4b} Let the conditions in Theorem \ref{thm:inference1} hold. For any $\alpha \in (0,1)$,  $\lim_{n \rightarrow \infty} P\Big(|\mathcal{T}_n| \le \mathrm{cv}_{K/2}^P(\alpha)\Big| H_0\Big) = 1 - \alpha$, where $\mathrm{cv}_{K/2}^P(\alpha)$ is a $(1 - \alpha)^{th}$ quantile of t-statistics computed from all permutations over the pairs' sign as described in Appendix \ref{sec:perm_tests}, and $H_0$ is as in Equation \eqref{eqn:h0}.
\end{cor}



To our knowledge, this set of results is the first for inference on welfare-maximizing policies with unknown interference. We conclude with a study on the estimated direct, spillover, and welfare effects.


\begin{thm}[Asymptotically neglegible bias of treatment effects] \label{thm:bias_direct} Let  Assumptions \ref{ass:ass_0}, \ref{ass:regularity_basic} hold, and $\eta_n = o(n^{-1/4})$. Then, $\mathbb{E}\left[\bar{\Delta}_n(\beta)\right] = \Delta(\beta) + o(n^{-1/2})$, where the second term does not depend on $K$. Similarly,
$\mathbb{E}\left[\bar{W}_n(\beta)\right] = W(\beta) + o(n^{-1/2})$, where the second term does not depend on $K$.  In addition, for all pairs $(k, k+1)$,
$
\mathbb{E}\Big[\widehat{S}_{(k, k+1)}(0, \beta) \Big] = S_1(0, \beta) + \mathcal{O}(\eta_n).
$
\end{thm}

The proof is in Appendix \ref{sec:proof3}.
The bias of the estimated direct effect is asymptotically negligible at a rate faster than the parametric rate $n^{-1/2}$ when \textit{pooling} observations from different clusters.  Our main insight here is that, with pairing and perturbations of opposite signs, the first-order bias cancels out.
Here, $\eta_n = o(n^{-1/4})$ is consistent with requirements in previous theorems. Given that the bias is asymptotically negligible, we can use existing results for inference on the direct effect \citep[e.g.,][who study inference on the direct effect without perturbations]{savje2017average}. For completeness, we show consistency in Corollary \ref{cor:3} in the Appendix. Inference on the marginal spillover effects follows similarly to inference on the marginal effect, and omitted for brevity.





























\vspace{-3mm}

\subsection{Multiwave experiment: policy optimization} \label{sec:theory}

Next, we derive theoretical properties of the adaptive experiment.
Theoretical results are for the general $p$-dimensional case ($p$ is finite). Let $\check{T} = T/p$. We assume the following.


\begin{ass}\label{ass:bounded}  Let (A) $Y_{i,t}^{(k)}$ be sub-Gaussian; and  (B) $K \ge 2(T/p + 1)$.
 \end{ass}


Condition (A) states that unobservables have sub-Gaussian tails (attained by bounded random variables); (B) assumes that the number of clusters is at least twice the number of waves, which guarantees that Lemma \ref{lem:1a} (unconfoundedness) holds.




\begin{ass}[Strong concavity] \label{ass:strong_concavity} Assume $W(\beta)$ is $\sigma$-strongly concave, for some $\sigma > 0$ (i.e., $W(\beta)$'s Hessian is strictly negative definite).
\end{ass}

An example is Example \ref{exmp:main}, where neighbors' effects induce decreasing marginal effects, and the treatment may present some costs, see  real-world data example in Figure \ref{fig:ex1}. Strong concavity also arises in linear models with negative externalities, see Example \ref{exmp:main2}. Assumption \ref{ass:strong_concavity} fails when spillovers occur only after that ``enough" individuals have received the treatment. To accommodate this setting, we relax Assumption \ref{ass:strong_concavity} in Appendix \ref{sec:quasi_concavity}, allowing for a strictly quasi-concave objective that is best suited for these settings. Settings where Assumption \ref{ass:strong_concavity} fails are those where also the spillover mechanism (e.g., the network) changes with the intervention, left to future research. In these cases, the proposed method returns a local optimum. When using multiple starting values of our adaptive algorithm, we only require concavity \textit{locally} to each starting value.








\begin{thm} \label{thm:rate2b}   Let  Assumptions \ref{ass:ass_0}, \ref{ass:regularity_basic}, \ref{ass:bounded}, \ref{ass:strong_concavity} hold. Take a small $1/4 > \xi > 0$, $\alpha_{k,w} = J/w$ for a finite $J \ge 1/\sigma$.   Let $n^{1/4 - \xi} \ge C \sqrt{p \log(n) \gamma_N T^{B p} \log(KT)}$, $\eta_n = 1/n^{1/4 + \xi}$, for finite constants $B, C > 0$. Then, with probability at least $1 - 1/n$, for a constant $\bar{C}' < \infty$, independent of $(p, n, N, K, T)$,
$
||\beta^* - \hat{\beta}^*||^2 \le \frac{p \bar{C}'}{\check{T}}.
$
\end{thm}

The proof is in Appendix \ref{sec:proof5}.
Theorem \ref{thm:rate2b} provides a bound on the distance between the estimated policy and the optimal one. The bound depends only on $T$ (and not $n$) because $n$ is assumed to be sufficiently larger than $T$.


\begin{cor} \label{cor:4}  Let the conditions in Theorem \ref{thm:rate2b} hold, and $K = 2(T/p + 1)$. With probability at least $1 - 1/n$,
$
W(\beta^*) - W(\hat{\beta}^*) \le \frac{p C'}{K},
$
for a constant $C' < \infty$ independent of $(p, n, N, K, T)$.
\end{cor}

The proof is in Appendix \ref{sec:corollary}.
The corollary formalizes the out-of-sample regret bound for $K = 2 (T/p + 1)$.  Also, the rate in $K$ does not depend on $p$,  as $n \rightarrow \infty$. This is different from grid-search procedures, where the rate in $K$ would be exponentially slower in $p$. Researchers may wonder whether the procedure is ``harmless'' also on the in-sample units.

\begin{thm}[In-sample regret] \label{thm:rate3} Let the conditions in Theorem \ref{thm:rate2b} hold. Then, with probability at least $1 - 1/n$, for a constant $c < \infty$ independent of $(p, n, N, K, T)$,
$$
\small
\begin{aligned}
\max_{k \in \{1,\cdots, K\}} \frac{1}{\check{T}} \sum_{w=1}^{\check{T}} \Big[W(\beta^*) - W(\check{\beta}_k^{w})\Big] \le  c \frac{p \log(\check{T})}{\check{T}}.
\end{aligned}
$$
\end{thm}

The proof is in Appendix \ref{sec:proof6}. Theorem \ref{thm:rate3} guarantees that the cumulative welfare in \textit{each} cluster $k$, incurred by deploying the current policy $\check{\beta}_k^w$ at wave $w$ (recall that in the general $p$-dimensional case we have $\check{T}$ many waves), converges to the largest achievable welfare at a rate $\log(T)/T$, also for those units participating in the experiment.\footnote{By a first-order Taylor expansion, a corollary is that the bound also holds for $\check{\beta}_k^w \pm \eta_n$ up to an additional factor which scales to zero at rate $\eta_n$ (and therefore negligible under the conditions imposed on $n$).} This result guarantees that the proposed design controls the regret on the experiment participants. This is a useful property that would not be attained, for example, by grid-search procedures for $W(\beta)$ (see Appendix \ref{app:non_adaptive}). We conclude with an  \textit{exponential} convergence rate of the out-of-sample (but not in-sample) regret with a different learning rate.



\begin{thm}[Out-of-sample regret with larger sample size] \label{thm:rate2bb}   Let  Assumptions \ref{ass:ass_0}, \ref{ass:regularity_basic}, \ref{ass:bounded}, \ref{ass:strong_concavity} hold, with $W(\beta)$ being $\tau$-smooth, and $K = 2T + 2$. Take a small $1/4 > \xi > 0$, $\alpha_{k,w} = 1/\tau$.   Let $n^{1/4 - \xi} \ge C \sqrt{p \log(n) \gamma_N e^{T B p} \log(KT)}$, $\eta_n = 1/n^{1/4 + \xi}$, for finite constants $B, C > 0$. Then, with probability at least $1 - 1/n$, for constants $0 < c_0, c_0'< \infty$, independent of $(n, N, K, T)$,
$$
\small
\begin{aligned}
W(\beta^*) - W(\hat{\beta}^*) \le c_0 \exp(- c_0' K).
\end{aligned}
$$
\end{thm}

The proof is in Appendix \ref{sec:exponential2}. The main restriction is that the sample size grows \textit{exponentially} in the number of iterations (instead of polynomially).  The theorem leverages properties of the gradient descent under strong concavity and smoothness \citep{bubeck2012regret}. Fast rates for the out-of-sample regret are achieved under an appropriate choice of the learning rate that leverages the smoothness of the objective function.  The choice of a learning rate invariant in the iteration $t$ requires a sample size exponential in $T$. This differs from the choice of a learning rate as $1/t$ in Theorem \ref{thm:rate2b}, where the adaptive learning rate enables controlling the cumulative error polynomially in $n$. To our knowledge, these regret guarantees are the first under unknown (and partial) interference.

We now contrast the above results with past literature. In the online optimization literature, the rate $1/T$ is common for convex optimization, assuming independent units \citep[see][for out-of-sample regret rates]{duchi2018minimax}. Here, because of interference, we leverage between-clusters perturbations. Also, we do not have direct access to the gradient, and related optimization procedures are those in the literature on zero-th order optimization \citep{kiefer1952stochastic}. \cite{flaxman2004online, agarwal2010optimal} in particular are related to our approach, where regret can converge at rate $O(1/T)$ in expectation only, whereas high-probability bounds are $1/\sqrt{T}$ \citep[see Theorem 6 in][and the discussion below]{agarwal2010optimal}.  Here, we exploit within-cluster concentration and between clusters' variation to control for large deviations of the estimated gradients and obtain faster rates for high-probability bounds. This approach also allows us to extend out-of-sample guarantees beyond global strong concavity (assumed in the above references) in Appendix \ref{sec:quasi_concavity}. In our derivations, the perturbation parameter depends on the sample size, differently from the references above, and the idea of sequential estimation is novel due to repeated sampling.
 \cite{wager2019experimenting} derive $1/T$ regret guarantees in the different settings of market pricing, as $n \rightarrow \infty$, with independent units and samples each wave.
Our results do not impose independence or modeling assumptions other than partial interference.  \cite{viviano2019policy} considers a single network, with \textit{observed} neighbors of experiment participants, instead of a sequential experiment. He imposes geometric (VC) restrictions on the policy and solves a mixed-integer linear program. Here, we introduce an adaptive experiment and we do not require network information, using network concentration not studied in previous works.




These differences require a different set of techniques for derivations.  The proof of the theorem (i) uses concentration arguments for locally dependent graphs \citep{janson2004large}; (ii) uses the within-cluster and between-clusters variation for consistent estimation of the marginal effect, together with the cluster pairing; (iii) it uses a recursive argument to bound the cumulative error obtained through the estimation and sequential cross-fitting.
















\vspace{-3mm}

\section{Computing the value of collecting network data} \label{sec:comp}

Here, we ask
how $\beta^*$ compares with the policy that assigns treatments without restrictions on the policy function, and provide useful bounds on the value of collecting network data. We focus on a setting with network spillovers, where $A$ denotes the unobserved adjacency matrix as in Example \ref{exmp:microfoundation}, and omit the super-script $k$ because the argument applies to any cluster. We study
\begin{equation} \label{eqn:difference_global}
\small
\begin{aligned}
W_N^* - W(\beta^*),  \quad W_N^* = \sup_{\mathcal{P}_N(\cdot) \in \mathcal{F}} \frac{1}{N} \sum_{i=1}^N \mathbb{E}\Big[\mathbb{E}_{D \sim \mathcal{P}_N(A, X)}[Y_{i,t} | A, X]\Big]
\end{aligned}
\end{equation}
with $\mathcal{F}$ as the set of \textit{all} conditional distribution of the vector $D \in \{0,1\}^N$, given network $A$ and the covariates of all observations $X$ as defined in Section \ref{sec:1a}.
Equation \eqref{eqn:difference_global} denotes the difference between the expected outcomes, evaluated at the global optimum over all possible assignments (with $A, X$ observed), and the welfare evaluated at $\beta^*$ (without observing $A$).
\begin{ass}[Discrete parameter space, assignment, and minimum degree] \label{ass:discrete_space} Consider a network model as in Example \ref{exmp:microfoundation}. Assume that $X_i \in \mathcal{X}, \mathcal{X} = \{1, \cdots, |\mathcal{X}|\}, |\mathcal{X}| < \infty, P(X = x) > \bar{\kappa} > 0$ $\forall x \in \mathcal{X}$. Let $\pi(x, \beta) = \beta_x$, and $\mathcal{B} = [0,1]^{|\mathcal{X}|}$. Let $\inf_{x, x', u'} \int l(x, u, x', u') dF_{U| X = x}(u) \ge \underline{\kappa}$, for some $\underline{\kappa}, \bar{\kappa} \in (0, 1]$, with $l(\cdot)$ defined in Equation \eqref{eqn:network}.
\end{ass}
Assumption \ref{ass:discrete_space} states that researchers assign treatments based on finitely many observable types as in \cite{manski2004}, \cite{graham2010measuring}. Each type $x \in \mathcal{X}$ is assigned a different probability $\beta_x$, which can take any value between zero and one. Assumption \ref{ass:discrete_space} also states that conditional on individual's type $(X_i, U_i)$, any other unobserved type $U_j$ can form a connection with individual $i$ with some positive probability, provided that $i$ and $j$ are connected under the latent space representation (recall Equation \ref{eqn:network}). This condition is consistent with the model in Example \ref{exmp:microfoundation} (and restrictions on $\gamma_N$), because the assumption states that the expected minimum degree is bounded from below by $\underline{\kappa} \gamma_N^{1/2}$, which is smaller than the maximum degree $\gamma_N^{1/2}$.  The second restriction is on the potential outcomes. Let
\begin{equation} \label{eqn:outcome_optimality}
\small
\begin{aligned}
Y_{i,t}(\mathbf{d}_t) & = \Big[\Delta(X_i) - v(X_i)\Big] \mathbf{d}_{i,t} + \mathcal{S}_{i,t}(\mathbf{d}_t) + \nu_{i,t}, \quad \mathbb{E}[\nu_{i,t} | X, A] = 0 \\
\mathcal{S}_{i,t}(\mathbf{d}_t) &=  s\Big(\frac{\sum_{j = 1}^n A_{i,j} \mathbf{d}_{j,t} 1\{X_j = 1\}}{\sum_{j = 1}^n A_{i,j} 1\{X_j = 1\}}, \cdots, \frac{\sum_{j = 1}^n A_{i,j} \mathbf{d}_{j,t} 1\{X_j = |\mathcal{X}|\}}{\sum_{j = 1}^n A_{i,j} 1\{X_j = |\mathcal{X}|\}}\Big),
\end{aligned}
\end{equation}
where $0/0 = 0$.
Here, $\Delta(\cdot)$ is the direct treatment effect, and $v(\cdot)$ is the cost of the treatment; $s(\cdot)$ captures the spillover effects. Spillovers depend on the fraction of treated neighbors and are heterogeneous in the neighbors' types, with no interactions with direct effects.









\begin{thm} \label{thm:optimality_global} Consider a model in Example \ref{exmp:microfoundation}. Let Equation \eqref{eqn:outcome_optimality} hold, with $s(\cdot)$ twice differentiable with bounded derivatives.  Suppose that Assumption  \ref{ass:discrete_space} hold. Then, with $W_N^*$ as in Equation \eqref{eqn:difference_global},
$
\lim_{N, \gamma_N \rightarrow \infty} \Big\{W_N^* - W(\beta^*)\Big\} \le  \mathbb{E}\Big[|\Delta(X) - v(X)|\Big].
$

\end{thm}

The proof is in Appendix \ref{sec:proof7}.
Theorem \ref{thm:optimality_global} bounds the welfare difference by the expected direct effects minus costs. If direct effects are small compared with the treatment costs, such a difference is negligible (for any spillover effects). The bound is identified \textit{without} network data under separability of direct and spillover effects. The theorem assumes that the maximum degree converges to infinity, but it may converge at a slower rate than $N$, consistent with our conditions in previous theorems. This result is novel in the context of the literature on targeting networked individuals and provides a formal characterization of the \textit{value} of collecting network information.\footnote{We note \cite{akbarpour2018just} study network value from the different angle of network diffusion: for a class of network formation models and diffusion mechanisms, the authors show that random seeding is approximately optimal as researchers treat a few more individuals. The main differences are that here (i) we do not study the problem from the perspective of network diffusion but instead focus on an exogenous interference mechanism with heterogeneity; (ii) we provide an upper bound in terms of the direct treatment effect, leveraging a different model and theory. Different from \cite{akbarpour2018just}, the upper bound does not state that we should treat $\epsilon$-more individuals (since we consider a different model of spillovers). }
Theorem \ref{thm:optimality_global} does \textit{not} state that spillovers are not relevant ($\beta^*$ depends on the spillovers). Instead, it states that one can compute best policies, without knowledge of the network in settings where direct effects are small.

One can estimate the bound by taking an absolute difference between the treated and control units for different individual types, and average across types.
In Example \ref{exmp:main}, the bound equals $\phi_1$ (the direct treatment effect) minus the cost of implementing the treatment.

\begin{cor} \label{cor:c_e} Let the conditions in Theorem \ref{thm:optimality_global} hold. Let $C_e$ be the cost of collecting network information per individual (with total cost for observing the network $A$ equal to $N C_e$). Then, $\lim_{N, \gamma_N \rightarrow \infty} W_N^* - W(\beta^*) - C_e \le 0$, if $C_e \ge \mathbb{E}\Big[|\Delta(X) - v(X)|\Big]$.
\end{cor}


\vspace{-5mm}


\section{Field experiment and calibrated numerical studies} \label{sec:field}

Next, we present a large-scale experiment where we implemented our single wave experiment over two consecutive experimentation waves. We use each wave to illustrate properties of the single wave experiment. We also use the second wave to the welfare gains of our experiment.  Finally, we present simulations with many waves calibrated to existing experiments.

\vspace{-3mm}


\subsection{Experimental design}

We now describe the main steps for experiment implementation. See Table \ref{tab:main_summaries} for a summary.

\vspace{-4mm}

\paragraph{Treatment $D$} The experiment was implemented through Precision Development (PxD), an NGO that provides farmers with phone-based agricultural advisory services. Farmers often lack access to geo-localized weather forecasts, and digital delivery offers solutions to address this challenge \citep{fabregas2019realizing}. Prior to the experiment, only 45\% of cotton growers reported consistent access to weather information, usually via radio or television. About 86\% of cotton growers indicated that weather information helps plan agricultural activities (\url{https://precisiondev.org/weather-forecasting-product-for-punjab-pakistan/}). In addition, those farmers with access to weather forecasts only access forecasts produced at the district level, a higher administrative unit that typically includes 3-4 tehsils  (tehsils are administrative units equivalent to US counties).

In partnership with a private forecast provider, Precision Development developed calibrated (geo-localized) weather forecast information localized at the tehsil level. The treatment consists of calling farmers to provide weather forecasts via robocalls, meant to improve farmers' ability to take measures in their plots. The experiment was randomized at large scale across approximately 400,000 farmers. We expected the experiment to generate spillovers. In a survey, $80\%$ of the respondents said they actively shared weather information with other farmers, providing suggestive evidence of spillovers.

\vspace{-4mm}


\paragraph{Target outcome $Y$} We study the effect of the treatment on farmers' ability to predict short-run weather. This is relevant in these applications: correctly predicting weather improves efficiency in the use of resources by, for example, using irrigation or pesticides more efficiently and better invest, see \cite{burlig2024long}.

As our main data source, we use repeated high-frequency (daily) cross-sectional survey data collected from June to October 2022. To measure farmers' weather forecasts, we ask farmers: ``What do you expect will be the maximum (minimum) temperature in your area
tomorrow?". We merge this information with PxD forecast weather the day after the survey interview with the specific farmer. We measure the absolute difference between the farmer's predicted maximum (and minimum) temperature and those predicted by PxD forecasts.
  To combine beliefs about maximum and minimum temperature, we construct a statistical index as described in \cite{viviano2021should} which serves as our main outcome.


 Temperature variables define \textit{incorrect} beliefs, i.e.,  negative treatment effects indicate when that farmer's prediction is closer to the PxD forecast or actual temperature.  Predicted temperature is a convenient proxy for farmers' one-day ahead weather perceptions since (i) it is less volatile than precipitation \citep{grenci2001world}; (ii) it does not exhibit effect dynamics/time heterogeneity (see Appendix \ref{sec:activities}); (iii) it is relatively stable within a tehsil.  Also, PxD forecast is a good proxy for real temperature.
 Table \ref{tab:forecast} below shows that PxD forecasts and real weather are strongly (and statistically significant) positively correlated.

The survey was  run over approximately $6,000$ farmers, stratified across tehsils and individual treatment status, of which we have approximately $1,000$ respondents for our main outcome. We check for balance on take up rates on many dimensions, see Section \ref{Sec:more_exp1}.



\vspace{-4mm}




\paragraph{Clusters} In total, 40 tehsils were exposed to experimental variation. Of these, 25 are exposed to our main experiment/design (with in total 287,000 farmers), whereas the remaining 15 are exposed to a different design.
Figure \ref{fig:sample_size} illustrates the region in Pakistan exposed to experimental variation and the sample size within each district (not all tehsils in a district are in the experiment). \textit{Tehsils} have from 5,000 to 20,000 farmers in the program.
We consider a tehsil a cluster. The assumption is that spillovers between different tehsils are negligible, here justified by the fact that tehsils denote large geographic areas, and forecasts are geo-localized at the tehsil level. In contrast to some prior work  \citep[e.g.,][]{banerjee2013diffusion}, our design allows for spillovers across villages in the same tehsil.








\begin{figure}[!htp]
\centering
\includegraphics[scale = 0.6, trim={0 1cm 0 1cm},clip]{./figures/sample_size1.pdf}
\includegraphics[scale = 0.5, trim={0 0.3cm 0 0.8cm},clip]{./figures/sample_size2.pdf}
\caption{Pakistan's map, organized in districts (each district contains multiple tehsils). Gray regions indicate areas selected for the experiment. Next to each district, we report the total sample size obtained from the tehsils in the experiment in the given district.}  \label{fig:sample_size}
\end{figure}






\vspace{-4mm}

 \paragraph{Policy $\beta$} Our policy of interest is choosing how many people to treat. Each treatment costs $0.29\$$ per farmer/year. As shown below, learning whether one can maximize information diffusion without treating all individuals in the population is relevant for decision making once the experiment is implemented at large scale in Pakistan.

 \vspace{-4mm}

 \paragraph{Two wave design} We deployed the \textit{local perturbation design} presented in Section \ref{sec:jh} over two consecutive waves: We induced perturbations around $\beta = 50\%$ in the first wave, where $\beta$ denotes the share of treated individuals, and in the second wave, we induced perturbations around $\beta = 70\%$. The first wave started in April 2022, during which approximately half of the population was exposed to treatment. The second wave started in August 2022 when we increased the total number of treated individuals across all clusters. This increase was planned ex-ante by the NGO's, since the NGO wanted to reach a larger number of treated units by the end of the intervention. To do so, we used a sequential design and induced local perturbation over each wave following our design in Section \ref{sec:jh}. This allows us to learn the marginal effects in each wave. In addition, since we find positive marginal effects over forecast accuracy in the first experimentation wave, and close to zero marginal effects in the second wave, the two waves will be helpful to estimate counterfactual welfare benefits of learning marginal effects through a sequential experiment.

 \vspace{-4mm}
\paragraph{Details about first wave and choice of $\eta_n$}   The first experimentation wave allows us to learn the marginal effect around $\beta = 50\%$.
Over the first experimental wave, we randomly draw a group of twelve tehsils (``Negative Perturbation/Medium Saturation") to have an average treatment probability across tehsils in this group of $\beta = 0.4$, hence inducing a negative perturbation $\eta_n = 10\%$. The choice of the perturbation should depend on power considerations, as, in principle, we may also be interested in more refined marginal effects, at the expense of lower power. To study here trade-offs in the choice of the perturbation parameter $\eta_n$, we select the Negative Perturbation group to have $\beta = 0.4$ \textit{on average}, with half of the (randomly selected) clusters in the Negative Perturbation group having exactly $\beta = 0.35$ and half of the clusters with $\beta = 0.45$.
We repeat the same with a ``Positive Perturbation/High Saturation" group with approximately $\beta = 0.6$ on average (and, similarly as before, with six tehsils in this group having $\beta = 0.55$ and seven $\beta = 0.65$). This gives us two (nested) perturbation designs. First, we obtain a better-powered perturbation design (which we refer to as our \textit{main design}) with a total of 25 clusters and perturbations around $\beta = 0.5$, with perturbations equal to $\eta_n = 10\%$ on average. The second design induces \textit{within} group perturbation of smaller order $5\%$, which allows us to also learn marginal effects at two more values $\beta \in \{40, 60\}\%$, with half of the clusters. A key intuition is that, by pooling clusters around smaller perturbations, all of our theoretical results directly apply to the main design, up to a small bias (see e.g., Remark \ref{rem:eta1}). We report results from the main (better powered) design with $\beta = 50\%, \eta_n = 10\%$ on average; we show that for smaller choice of $\eta_n$ estimates can be under-powered, see Appendix Table \ref{tab:marginal_effects_ma}. We recommend the choice of two nested designs to avoid under-powered studies.


\vspace{-4mm}


\paragraph{Details about second wave} The second wave experiment allows us to learn marginal effects at $\beta = 0.7$.
Over the second wave (August - October), the ``Negative Perturbation" group was exposed to a larger treatment probability $\beta = 0.6$ and the ``Positive Perturbation" group was exposed to a treatment probability $\beta = 0.8$. Therefore, over the second experimentation wave, we have two groups with treatment probabilities $\beta = 70\% \pm \eta_n, \eta_n = 10\%$.\footnote{Over the second wave, we also perturbed by 0.05 the probability of treatment for different types of
farmers, those below and above the median response rate in the first round, keeping the overall treatment
probability constant. This latter perturbation enables estimating heterogeneous treatment effects, omitted
from the main analysis for brevity and discussed in Appendix \ref{app:more_experiment}.}



\begin{table}[!ht]
\centering
\scalebox{0.6}{\begin{tabular}{@{\extracolsep{5pt}} l|cc}
\\[-1.8ex]\hline
\hline \\[-1.8ex]
 &  Wave 1  & Wave 2 \\
\hline \\[-1.8ex]
 Treatment $D$ &  One-day ahead geo-localized weather forecast  &  One-day ahead geo-localized weather forecast \\
 & & \\
 Outcome $Y$ &  Farmer's one-day ahead correct forecast (temp) & Farmer's one-day ahead correct forecast (temp) \\
 & & \\
  Policy $\beta$ &  Share of treated farmers  & Share of treated farmers \\
  & & \\
  Choice of $\beta$ & $50\%$ & $70\%$ \\
  & &  \\
  Perturbation main exp $\eta_n$ & $10\%$ & $10\%$ \\
  & & \\
  Total \# of clusters & 25 & 25 \\
  & & \\
  Total \# of farmers in main experiment & 287,487 & 287,487 \\
  & & \\
 Total \# surveyed individuals in main exp/ & 247 & 633 \\
  & & \\
  \hline \\
  Estimated marginal effect [p-value] & -3.39$^{**}$ [0.03] & -0.93 [0.24] \\
  & & \\
 Mechanism  & Large marginal spillover effects & Close to zero marginal spillover effects \\
    & & \\
  Cost treatment farmer/year  & $0.29\$$ farmer/year & $0.29\$$ farmer/year \\
  & & \\
  Policy implication & Increase share of treated individuals & Treat $\sim 70\%$ of farmers \\ & &  (save 1,000,000\$/year once implemented at scale)
\end{tabular}
}
  \caption{Illustration of how our theoretical framework maps to this experiment. P-value is for one sided test computed via randomization inference. Marginal spillovers denote the marginal effect of increasing friends' treatment probabilities of the control units. }
  \label{tab:main_summaries}
\end{table}








\vspace{-4mm}

\paragraph{Roadmap of the main design} Our main design identifies the marginal effects, the marginal spillover effects, direct effects, and welfare effects at $\beta = 50\%$ in the first wave and at $\beta = 70\%$ over the second wave. Table \ref{tab:advantages_pert2} illustrates which effects are identified by the experiment and reports the main robustness checks in the Appendix. It also compares our experiment to a standard saturation experiment that chooses $\beta \in \{50, 70\}\%$ without using local perturbations, and which, therefore identifies a smaller set of parameters.





\begin{table}[!ht]
\centering
\scalebox{0.6}{\begin{tabular}{@{\extracolsep{5pt}} l|ccc}
\\[-1.8ex]\hline
\hline \\[-1.8ex]
\textbf{Main estimates:} for $\beta \in \{50, 70\}\%$ &  Our Perturbation Design    & Standard Saturation Design & Estimation w/ perturbation design  \\
& & \\
Effect $W(\beta)$  & $\checkmark$ (Figure \ref{fig:effects})  & $\checkmark$ & Pooling around $\beta$ \\
 Direct Effect $\Delta(\beta)$ & $\checkmark$ (Table \ref{tab:marginal_effects_my_main})   & $\checkmark$ & Pooling around $\beta$ \\
Marginal Effect $M(\beta)$ &   $\checkmark$ (Table \ref{tab:marginal_effects_my_main})    &$\bm{\times}$ & Comparison positive/negative perturbation
 \\
  Marginal Spillover $S(\beta)$ &   $\checkmark$ (Table \ref{tab:marginal_effects_my_main})   & $\bm{\times}$  & Comparison positive/negative perturbation
 \\
 & & \\
 \hline \\[-1.8ex]
 \textbf{Additional estimates/robustness}  & \\
 & & \\
 Regression estimates  & $\checkmark$ (Table \ref{tab:beliefs_forecast2})  &  $\checkmark$ \\
 Balance table for cluster heterogeneity  & $\checkmark$ (Table \ref{tab:summaries}) & $\checkmark$\\
 Balance table on surveyed individuals   & $\checkmark$ (Tables  \ref{tab:summaries_response_rates}, \ref{tab:summaries_response_rates_v2}, \ref{tab:table_non_respondents}) & $\checkmark$\\
 Check for dynamic effects  & $\checkmark$ (Table \ref{tab:dynamics}) & $\checkmark$\\
 Check for treatment efficacy  & $\checkmark$ (Tables \ref{tab:preliminary_treatment2}, \ref{tab:forecast}) & $\checkmark$ \\
 More refined marginal effects & $\checkmark$ (Table \ref{tab:marginal_effects_ma})   & $\bm{\times}$
\end{tabular}
}
  \caption{Comparisons between two designs: our perturbation design that induces local perturbations around two treatment probabilities $\beta_1 = 50\% \pm \eta_n$, and $\beta_2 = 70\% \pm \eta_n$, as in our experiment in Section \ref{sec:field} and a standard saturation design that chooses $\beta = 50\%$ and $\beta_2 = 70\%$ (without local perturbation). First column indicates the identified effects at $\beta \in \{50, 70\}\%$ (up-to a bias negligible for inference) of the proposed perturbation design. The second column indicates which effects are (check) and are not (cross) identified from a saturation experiment with two treatment probabilities exactly equal to $\beta \in\{50, 70\}\%$. }
  \label{tab:advantages_pert2}
\end{table}



\begin{table}[!htp] \centering

\scalebox{0.65}{\begin{tabular}{@{\extracolsep{5pt}} lcccc}
\\[-1.8ex]\hline
\hline \\[-1.8ex]
Group of Tehsils &  Number of Farmers & Number of Tehsils & Average $\beta$ (Wave 1) & Average $\beta$ (Wave 2) \\
\hline \\[-1.8ex]
Negative Perturbation (Medium Saturation)  & 137 729 & 12 & $50\% - \eta_n = 40\%$  & $70\% - \eta_n = 60\%$\\
& & & \\
Positive Perturbation (High Saturation) &  149 758 & 13 & $50\% + \eta_n = 60\%$  & $70\% + \eta_n = 80\%$
 \\ & & &  \\
Low Saturation, not following main experiment & 111 300 & 10 & $11\%$ & $25\%$  \\
\hline \\[-1.8ex]
\end{tabular}
}
  \caption{Statistics of the experiment. $\beta$ indicates the average treatment probability across each group of tehsils. For lower saturation, we assigned different probabilities to each tehsil. }
  \label{tab:saturations}
\end{table}

\begin{rem}[Additional saturation group] The experiment also encompasses a third group of tehsils, ``Low Saturation", with a \textit{different} design, which assigns tehsil-specific perturbations to treatment probabilities, with $\beta = 0.11$ on average over the first wave and $\beta = 0.25$ over the second wave, but without inducing perturbations as in the other groups.\footnote{For the low saturation group, we follow a different design and assign tehsil-specific treatment probabilities with, on average, $0.11$ treatment probability. We vary such probabilities between tehsils as a function of the overall rural population in a tehsil, fixing the share of the rural population receiving the treatment.} Each group of tehsils was stratified across districts. We use the Negative and Positive Perturbation groups to compute the marginal effects since these groups closely follow Section \ref{sec:jh}.   \qed
\end{rem}


 \vspace{-4mm}
























\vspace{-3mm}

\subsection{Main results: marginal and welfare effects}



Next, we study marginal effects
on beliefs about PxD forecasts (i.e., whether the farmer's prediction agrees with PxD forecast), illustrating properties of our design on our main outcome (forecast temperature). We assume no cluster fixed effects because of lack of baseline outcomes. Although this is a strong assumption, it is motivated by balance across clusters on pre-treatment observables (Table \ref{tab:summaries}). In practice, we recommend to collect baseline outcomes when possible and, as in this case, when infeasible, to check for balance on observable covariates between clusters exposed to different treatment probabilities. Appendix Table \ref{tab:main_experiment3} provides results for response rates for which we observe baseline outcomes.


\vspace{-4mm}
\paragraph{Estimated Marginal Effects} Figure \ref{fig:effects} plots the estimated marginal effects in the main design, i.e., for $\beta \in \{50, 70\}\%$ (with $\eta_n = 10\%$ on average). The figure also reports the estimated welfare at each point $\beta \in \{0.4, 0.6, 0.8\}$. We observe decreasing marginal effects when moving from $\beta = 0.5$ to $\beta = 0.7$.

Table \ref{tab:marginal_effects_my_main} shows that the marginal effect is large and statistically significant at $\beta = 50\%$, preserve sign but is smaller and non-significant at $\beta = 70\%$. P-values are computed via randomization inference for one sided test, formally described in Appendix \ref{sec:perm_tests}.  This result is suggestive that treating $50\%$ of the population is sub-optimal, whereas treating $70\%$ of the individuals is close to be optimal. Therefore, our design allows us not only to learn the value of welfare around $\beta \in \{50, 70\}\%$ but also its corresponding marginal effects. Marginal effects can be useful to understand whether we should increase treatment probabilities to improve welfare. Table \ref{tab:marginal_effects_my_main} reports direct and marginal spillover effects. In particular, we observe marginal effects are mostly driven by large and significant marginal spillover effects at $\beta = 50\%$ (i.e., marginal effects of increasing the friends' treatment probability), whereas marginal spillover effects are close to zero at $\beta = 70\%$.






\vspace{-4mm}

\paragraph{Welfare gains and welfare comparison with standard saturation design} Using the two experimental waves, we can estimate the welfare improvement of an adaptive experiment that, in the first wave, estimates the marginal effects at $\beta = 50\%$, and in the second wave estimates marginal effects at $\beta = 70\%$. We contrast our design with a typical saturation experiment or grid search method would predict in Figure \ref{fig:final_benefit}: a saturation experiment treating $\{0, 50\%, 100\%\}$ \citep{sinclair2012detecting}  of the individuals would not able to identify decreasing marginal effects near $70\%$, and similarly for other choices of treatment probabilities. This is because a standard saturation design would not induce local perturbations. Such a saturation design would recommend \textit{all} individuals to be assigned to treatment. Our experiment uses information about the marginal effects to identify the optimum near $70\%$.
Figure \ref{fig:final_benefit} reports the relative improvement from the one wave experiment to the second wave experiment where $70\%$ of individuals are treated. Increasing number of treated units from $50\%$ to $70\%$ of individuals leads to statistically significant increase in welfare.

A saturation experiment that would recommend treating all individuals would lead to small improvements: We can use as a \textit{conservative} estimate (upper bound) of welfare at $\beta = 100\%$, its Taylor approximation at $\beta = 70\%$, $W(0.7) + 0.3 M(0.7)$ (this is a conservative estimate because we might most likely expect decreasing marginal effects, as supported by Table \ref{tab:marginal_effects_my_main}). Despite using a conservative upper bound, predicted improvement when treating all units in the population are small relative to only treating $70\%$ (equal to $8\%$) and non-significant.
Treating only $70\%$ of the individuals instead of $100\%$ would save approximately 0.29\$ per farmer/year. This is economically significant if we consider a policy implemented on all farmers in Pakistan (approximately ten millions), saving one million US dollars/year.




\begin{minipage}{\textwidth}
  \begin{minipage}[b]{0.49\textwidth}
    \centering
   \includegraphics[scale=0.4]{./figures/temperature_effects}
    \captionof{figure}{Difference between the farmer's predicted temperature and PxD's temperature forecast for the day after the interview. The larger dots report the estimated effects at $\beta = 50\%, \beta = 70\%$, from the first and second wave. The lines report the estimated marginal effects, and the smaller dots the effect estimated at $\beta \in \{40, 60, 80\}\%$ over the first wave (first line) and second wave (second line).} \label{fig:effects}
  \end{minipage}
  \hfill
  \begin{minipage}[b]{0.49\textwidth}
    \centering
   \scalebox{0.6}{\begin{tabular}{@{\extracolsep{5pt}}lcc}
\\[-1.8ex]\hline
\hline \\[-1.8ex]
Incorrect beliefs about  & \multicolumn{2}{c}{PxD forecast Temperature}  \\
\\[-1.8ex] & $\beta = 50\%$ (Wave 1) & $\beta = 70\%$ (Wave 2)\\
\hline \\[-1.8ex]
Marginal Effect & -3.39$^{**}$ & -0.93  \\
p-value & [0.03] & [0.24] \\
& & \\
Direct Effect & -0.94$^{**}$ & -0.53 \\
p-value & [0.04] & [0.19]  \\
& & \\
Marginal Spillovers on Treated & 1.69 & -1.91  \\
p-value & [0.27] & [0.19]  \\
& &  \\
Marginal Spillovers on Controls & -6.40$^{**}$ & 0.77  \\
p-value & [0.03] & [0.42]  \\
  & & \\
\hline \\[-1.8ex]
Observations & 247 & 633   \\
\hline \\[-1.8ex]
\textit{Note:}  & \multicolumn{2}{r}{$^{*}$p$<$0.1; $^{**}$p$<$0.05; $^{***}$p$<$0.01}
\end{tabular} }
      \captionof{table}{Estimated effects over first and second wave from the main design. P-values are computed via randomization inference for one sided tests.} \label{tab:marginal_effects_my_main}
    \end{minipage}
  \end{minipage}







\begin{figure}[!ht]
    \centering
   \includegraphics[scale = 0.5]{figures/bar_plot.eps}
    \caption{Benefits of a sequential experiment using predicted temperature as a proxy for welfare. The light blue column reports the percentage change in average forecast accuracy  generated by a policy recommendation using the proposed adaptive experiment, either with one-wave experiment (first column), or two waves (third column), or using a standard saturation experiment with probabilities $\{0, 0.5, 1\}$ (last column). The second, fourth, and sixth columns report the cost of the intervention that would be recommended by each of these experiments relative to the cost of treating everybody in the population. The policy-maker using only the first wave experiment deploys a policy that treats $50\%$ of the individuals, with corresponding costs of the intervention equal to $50\%$ of the total costs relative to treating everybody. The second wave experiment identifies positive marginal effects at $70\%$ and recommends treating around $70\%$ of the individuals. The forecast accuracy increases, as well as the costs of the total intervention. The standard saturation experiment does not identify marginal effects and recommends treating $100\%$ of the individuals. The error bars report $10\%$ confidence intervals over the improvement from the first to the second wave and from the two wave to treating everybody in the population (what Saturation/Grid would suggest). These are obtained via randomization inference on the gradient at $\beta = 0.5$ and $\beta = 0.7$, respectively, and using a first-order Taylor approximation to the welfare around $0.7$ to obtain a conservative estimate of the effect at $\beta = 100\%$.     }
    \label{fig:final_benefit}
\end{figure}




\newpage




\subsection{Additional analyses: balance and regression estimates} \label{Sec:more_exp1}

We conclude with a brief overview of additional analyses and balance checks in the Appendix.


\vspace{-4mm}

\paragraph{Balance} We use auxiliary data about farmers' baseline characteristics for \textit{all} farmers enrolled with PxD in the main experiment (more than 287,000 farmers) to test for homogeneity in covariates between different clusters, a relevant assumption in our framework. Namely, given that our framework requires homogeneity across clusters, we test for homogeneity of covariates between different clusters using information from all individuals in the experiment.
Appendix Table \ref{tab:summaries} reports the sample means across observable baseline covariates from program administrative data (each covariate is described below Table \ref{tab:summaries}).
We test for differences in covariates between clusters exposed to different treatment probabilities.\footnote{When estimating marginal effects, it is easy to show that our framework only requires homogeneity restrictions between groups of clusters used to estimate the marginal effects (e.g., the group of clusters in different treatment exposures), but not necessarily between individual clusters having the same exposures.}    The relevant null hypothesis is that the expected value of each covariate in Table \ref{tab:summaries} in each cluster is the same across all clusters. We construct these tests via randomization inference formally described in Appendix \ref{sec:perm_tests}.
These tests are informative of whether such groups are comparable and are conducted with a large sample size ($n \approx 10,000$ on average in each tehsil). We observe similar estimates across all covariates. The smallest p-value is $0.21$, the median is above $0.5$, suggesting lack of  tehsil-level heterogeneity.

In Appendix Tables \ref{tab:summaries_response_rates}, \ref{tab:summaries_response_rates_v2}, \ref{tab:table_non_respondents} we also report balance table on response rates (both among all surveyed individuals and between respondents and non-respondents individuals), where results show substantial balance in relevant baseline characteristics.

\vspace{-4mm}
\paragraph{Treatment take-up and accuracy} The treatment group received approximately three times more frequent calls than the control group by design -- where the control group's calls were about other activities of the NGO. Appendix Table \ref{tab:preliminary_treatment2} shows that the larger number of calls does \textit{not} negatively affect response rates. Treated individuals present higher (and statistically significant at the $1\%$ level) response rates per call, engaging more with calls.

Table \ref{tab:forecast} shows that forecast and real precipitation and temperature are strongly positively correlated, motivating our main focus on farmers' beliefs about PxD forecasts: PxD predicted and real weather follow very similar patterns, but beliefs about PxD forecasts are less noisy.

\begin{table}[!htbp] \centering

\scalebox{0.6}{\begin{tabular}{@{\extracolsep{5pt}}lccc}
\\[-1.8ex]\hline
\hline \\[-1.8ex]
 & \multicolumn{3}{c}{\textit{Dependent variable:}} \\
\cline{2-4}
\\[-1.8ex] & \multicolumn{3}{c}{ }
\\[-1.8ex] & Real Precipitation & Real Temperature Max & Correct Rain Forecast\\
\hline \\[-1.8ex]
Forecast Precipitation & 0.675$^{***}$ &  &  \\
  & (0.020) &  &  \\
Forecast Temperature Max &  & 0.914$^{***}$ &  \\
  &  & (0.029) &  \\
 Constant & 1.585$^{***}$ & 0.274 & 0.786$^{***}$ \\
  & (0.069) & (1.112) & (0.005) \\
\hline \\[-1.8ex]
\hline
\hline \\[-1.8ex]
\textit{Note:}  & \multicolumn{3}{r}{$^{*}$p$<$0.1; $^{**}$p$<$0.05; $^{***}$p$<$0.01} \\
\end{tabular} }
  \caption{Forecast vs real weather in 2022. Sample size equal to 22230. The first column uses precipitation as a continuous variable and the last column regresses the indicator of whether the forecast of whether it will rain correctly predicts whether it rains.   In parenthesis standard errors clustered at the tehsil level.}
  \label{tab:forecast}
\end{table}


\vspace{-4mm}
\paragraph{Parametric regression estimates} Our design allows for standard regression methods. We illustrate this in Tables \ref{tab:beliefs_forecast2}. Table \ref{tab:beliefs_forecast2} reports regression estimates of farmers' incorrect beliefs about temperature and rain with respect to forecast rain from PxD, for which we find mostly significant spillover effects. For parametric regression estimates we can use information from all clusters, including the lower saturation group, after appropriately controlling for the treatment probability, since also in this group treatment are randomized.

\vspace{-4mm}

\paragraph{Dynamics and additional outcomes} Appendix Table \ref{tab:dynamics} illustrates lack of dynamics on our primary outcome (temperature forecasts). In Table \ref{tab:dynamics}, we also illustrates effects on other outcomes. We collect information about predicted rain, asking ``Do you think it will rain in your
area tomorrow?" We use a binary indicator indicating whether the farmers incorrectly predict no rain and, instead, it rains or vice versa (or replies ``I do not know"). As shown in Table \ref{tab:dynamics}, we do not consider rain as the main welfare proxy because, different from temperature, this may exhibit treatment effect heterogeneity over time, since the experiment spans seasons of different rain intensity (dry and monsoon seasons). Finally, we use survey information about farming activities to show effects on these in Appendix \ref{app:more_experiment}.














\subsection{Calibrated numerical studies} \label{sec:experiment}






To evaluate the performance of our design with many waves, we calibrate simulations to data from \cite{cai2015social} and \cite{alatas2012targeting, alatas2016network}, while making simplifying assumptions whenever necessary. As in our application, we let $\beta$ denote the treatment probability and $\eta_n = 10\%$.\footnote{Here $10\%$ is consistent with the rule of thumb for $\eta_n \approx \sqrt{\sigma^2/c} n^{-1/3}$ (Appendix \ref{app:rule_thumb}), where $\sigma^2$ is the outcomes' variance and $c$ is the objective's curvature, which would prescribe values between $7\%$ and $12\%$ as we vary $n$. In the online supplement, we report results as we vary $\eta_n$ (Figure \ref{fig:comparison_eta}).} In the first calibration, the outcome is insurance adoption, and the treatment is whether an individual received an intensive information session. In the second calibration, the treatment is whether a household received a cash transfer, and the outcome is program satisfaction. The experiment of \cite{cai2015social} contains multiple arms. Here, we only focus on the treatment effects of intensive information sessions, pooling the remaining arms together for simplicity. The experiment of \cite{alatas2012targeting} contains different arms assigned at the village level, as well as information on cash transfers assigned at the household level. Here, we study the effect of cash transfers only and control for village-level treatments when estimating the parameters of interest.


In each cluster $k$, we generate
\begin{equation} \label{eqn:calibrated_model}
\small
\begin{aligned}
Y_{i,t} = \phi_0 + \phi_1 D_{i,t} + \phi_2 S_{i,t} + \phi_3 S_{i,t}^{2} - c D_{i,t} + \eta_{i,t}, \quad S_{i,t} = \frac{\sum_{j\neq i} A_{i,j} D_{i,t} }{ \sum_{j \neq i} A_{i,j}}, \eta_{i,t} \sim_{i.i.d.} \mathcal{N}(0, \sigma^2),
\end{aligned}
\end{equation}
where $c$ is the cost of the treatment.
We consider two sets of parameters $\Big(\phi_0, \phi_1, \phi_2, \phi_3, \sigma^2\Big)$ calibrated to data from \cite{cai2015social} and \cite{alatas2012targeting, alatas2016network} respectively. We obtain information on neighbors' treatment directly from data from  \cite{cai2015social}. For the second application, we merge data from  \cite{alatas2012targeting}, and \cite{alatas2016network}, and use information from approximately 100 observations whose neighbors' treatments are all observable to estimate the parameters.\footnote{This approach introduces a sampling bias in the estimation procedure, which we ignore for simplicity, given that our goal is not the analysis of the original experiment but only calibrating numerical studies.} For either application, we estimate a linear model as in Equation \eqref{eqn:calibrated_model}, also controlling for additional covariates to guarantee the unconfoundedness of the treatment.\footnote{For \cite{cai2015social} the covariates are gender, age, rice area, literacy level, a coefficient that captures the risk aversion, the baseline disaster probability, education, and a dummy containing information on whether the individual has one to five friends. For \cite{alatas2012targeting}, we control for the education level, village-level treatments, i.e., how individuals have been targeted in a village (i.e., via a proxy variable for income, a community-based method, or a hybrid), the size of the village, the consumption level, the ranking of the individual poverty level, the gender, marital status, household size, the quality of the roof and top.}  For simplicity, we consider as cost of treatment $c = \phi_1$, i.e., the opportunity cost of allocating the treatment to a population of disconnected individuals.





 We generate $K$ clusters, each with $N = 600$ units, and sample $n \in \{200, 400, 600\}$.
We generate a geometric network
$
A_{i,j} = 1\Big\{||U_i - U_j||_1 \le  2 \rho/\sqrt{N}\Big\},  U_i \sim_{i.i.d.} \mathcal{N}(0, I_2),
$
where the parameter $\rho$ governs the density of the network. The geometric formation process and the $1/\sqrt{N}$ follow similarly to simulations in \cite{leung2019treatment}. We report results for $\rho = 2$ here, while results are robust as we increase $\rho$ (see Appendix \ref{app:more_sim}). Throughout the analysis, without loss, we report welfare divided by its maximum $W(\beta^*)$ (i.e., $W(\beta^*) = 1$), and we subtract the intercept $\phi_0$.








In Appendix \ref{sec:experiment2}, we study the performance of the one-wave experiment. We show that the proposed test controls size uniformly across specifications and present desirable properties for power. Here, we present simulations for the multi-wave experiment. In the adaptive experiment, we choose the learning rate $10\%/\sqrt{t}$ with gradient norm rescaling as Remark \ref{rem:learning_rate}.\footnote{This choice guarantees that for each iteration, we only vary treatment probabilities by at most $10\%$, and the size of the variation is decreasing over each iteration, as for the learning rate under strong concavity without norm rescaling. This choice is preferable to $10\%/\sqrt{T}$ because it allows for larger steps in the initial iterations. A valid alternative is $10\%/t$. The latter case has a practical drawback: updates become very small after a few iterations. Comparisons for different learning rates are in the online supplement (Fig \ref{fig:learning_rate_comparisons}). } Since the model does not allow for time-varying fixed effects, we estimate marginal effects without baseline outcomes. For the multi-wave experiment, we initialize parameters at a small treatment probability $\beta = 0.2$ (here the optimum is around $60\%$).


 We let $T \in \{5, 10, 15, 20\}$. In Table \ref{tab:cov02}, we report the welfare improvement of the proposed method with respect to a grid search method that samples observations from an equally spaced grid between $[0.1, 0.9]$ with a size equal to the number of clusters (i.e., $2T$). We consider the best competitor between the one that maximizes the estimated welfare obtained from a correctly specified quadratic function and the one that chooses the treatment with the largest value within the grid. For both the competing methods, but not for the proposed procedure, we divide the outcomes' variance $\sigma^2$ by $T$, simulating settings where researchers may sample outcomes $T$ times (hence outcomes with a \textit{lower} variance) from each cluster before estimating treatment effects, and obtaining \textit{more precise} information.
The panel at the top of Table \ref{tab:cov02} reports the out-of-sample welfare improvement. The improvement is positive, and up to three percentage points for targeting information and up to sixty percentage points for targeting cash transfers. Improvements are generally larger for larger $T$. The panel at the bottom of Table \ref{tab:cov02} reports positive and large improvements for the in-sample welfare across all the designs, worst-case across clusters. For the worst-case regret, we fix the number of clusters to $K = 40$ for the proposed method and study the properties as a function of the number of iterations.  The improvements are twice as large for targeting information and thirty percentage points larger for targeting cash transfers. These are often increasing in $T$ with a few exceptions since uniform concentration may deteriorate for large $T$ and small $n$ as we consider the worst-case welfare across clusters.

In the online Appendices  \ref{sec:aa1}, \ref{sec:aa2}, we report results across many other specifications of the network, policy functions, and choice of different parameters and different starting values (e.g., also when $\beta$ is initialized near the optimum).










\begin{table}[!htp]\centering
\caption{Multiple-wave experiment. $200$ replications. The relative improvement in welfare with respect to the best competitor for $\rho = 2$. The panel at the top reports the out-of-sample regret, and the one at the bottom the worst-case in-sample regret.
}    \label{tab:cov02}
\renewcommand{#1}{1.3}
\scalebox{0.8}{\begin{tabular}{@{}lrrrrrrrrrrrrrr@{}}\toprule
& \multicolumn{6}{c}{Information} & \ & \multicolumn{6}{c}{Cash Transfer} \\
\cmidrule{2-7}  \cmidrule{9-14}
$T = $ &  5 & 10  & 15 & 20&  &   && & 5 & 10  & 15 & 20 & &  \\
\Xhline{.8pt}
$n=200$ & $0.057$ & $0.135$ & $0.297$ & $0.212$ &  &   && & $0.232$ & $0.243$ & $0.264$ & $0.287$ \\
$n=400$ & $0.226$ & $0.209$ & $0.355$ & $0.346$ &  &   && & $0.243$ & $0.274$ & $0.321$ & $0.335$ \\
 $n=600$ & $0.299$ & $0.281$ & $0.344$ & $0.492$ &  &   && & $0.261$ & $0.313$ & $0.343$ & $0.360$ \\
 \Xhline{.8pt}
$n=200$ & $0.621$ & $0.731$ & $0.736$ & $0.752$ &  &   && &  $0.247$ & $0.279$ & $0.300$ & $0.320$ \\
$n=400$ & $0.652$ & $0.745$ & $0.874$ & $0.898$ &  &   && & $0.266$ & $0.306$ & $0.343$ & $0.352$ \\
 $n=600$ & $0.646$ & $0.801$ & $0.942$ & $1.125$ &  &   && & $0.294$ & $0.360$ & $0.387$ & $0.387$ \\
 \bottomrule
\end{tabular}
}
\end{table}






\vspace{-8mm}

\section{Conclusions} \label{sec:conclusions}



This paper makes two main contributions. First, it introduces a single-wave experimental design to estimate the marginal effect of the policy and test for policy optimality. The experiment also enables identifying and estimating treatment effects, which can be of independent interest.
Second, it introduces an adaptive experiment to maximize welfare. We derive asymptotic properties for inference and provide a set of guarantees on the in-sample and out-of-sample regret. We illustrate the benefits of the method in a large-scale field experiment on information diffusion.  Our empirical application shows that using the marginal effect can be informative for decision-making even with few (two) waves.



This work opens new questions also from a theoretical perspective. We leave to future research the study of properties of the estimators when (i) clusters are not fully disconnected, in the spirit of   \cite{leung2023network};  (ii) clusters need to be estimated, similarly to graph-clustering procedures; (iii)  clusters present different distributions, as we discuss in Appendix \ref{app:het}. Similarly, studying the properties of the proposed method, as the degree of interference is proportional to the sample size, is an interesting direction. This is theoretically possible, as illustrated in Theorem \ref{thm:const1}, and we leave its comprehensive analysis to future research. Finally, an open question is how to estimate policies when the network is only partially observed \citep[e.g.,][]{breza2017using, manresa2013estimating}, and how to measure costs and benefits of collecting network data, on which Section \ref{sec:comp} provides novel directions for future research.









\vspace{-5mm}


\bibliography{my_bib2}
\vspace{-2mm}
\bibliographystyle{chicago}
















\newpage