EconBase
← Back to paper

Estimating treatment-effect heterogeneity across sites, in multi-site randomized experiments with few units per site

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

119,229 characters · 0 sections · 60 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Estimating treatment-effect heterogeneity across sites, in multi-site randomized experiments with few units per site.

abstractIn multi-site randomized trials with many sites and few randomization units per site, an Empirical-Bayes estimator can be used to estimate the variance of the treatment effect across sites. When this estimator indicates that treatment effects do vary, we propose estimators of the coefficients from regressions of site-level effects on site-level characteristics that are unobserved but can be unbiasedly estimated, such as sites' average outcome without treatment, or site-specific treatment effects on mediator variables. In experiments with imperfect compliance, we show that the sign of the correlation between local average treatment effects (LATEs) and site-level characteristics is identified, and we propose a partly testable assumption under which the variance of LATEs is identified. We use our results to revisit behaghel2014private, who study the effect of counseling programs on job seekers' job-finding rate, in 200 job placement agencies in France. We find considerable treatment-effect heterogeneity, both for intention to treat and LATE effects, and the treatment effect is negatively correlated with sites' job-finding rate without treatment.

\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Introduction}

\paragraph{Motivation.} From 2014 to 2016, “AEJ: Applied Economics” published 12 multi-site RCTs with treated and control units within each site, thus making it possible to estimate the treatment effect in each site. Typically, those RCTs are conducted in dozens, and sometimes hundreds, of different neighborhoods, villages, or regions, but they have a small number of randomization units per site. Few of these 12 papers investigate the treatment-effect's heterogeneity across sites.\footnote{One paper estimates the treatment-effect's variance across sites, and two more estimate the average treatment effect separately for different subgroups of geographical locations.} This paper provides novel estimators that researchers can use to estimate and predict that heterogeneity. Doing so, we hope to help generalize this type of heterogeneity analyses, as we believe that they can lead to useful insights. If one finds that the treatment effect is not heterogeneous across sites, this suggests that the RCT results may have some external validity, and might also apply to sites not included in the RCT. If on the other hand effects are heterogeneous, finding predictors of the sites' effects can provide suggestive evidence of the mechanisms underlying the treatment's effect. For instance, in a job-search counseling RCT, it can be interesting to study whether sites that have the largest effects on the job-finding rate are also the sites that have the largest effect on job-seekers' search effort, as a “predictive mediation analysis” of whether the job-finding effect can be “explained” by the job-search effect. Finding predictors of the sites' effects can also improve the program"s targeting, and under additional assumptions this can help predict the effect in sites not included in the RCT hotz2005predicting.

\paragraph{Set-up.} We consider an RCT stratified at the site level. We allow for imperfect compliance with treatment assignment, and consider both the heterogeneity of intention-to-treat effects (ITTs) and local-average-treatment-effects (LATEs) across sites. We assume that each site has at least two treated and two control units, so that $\text{ITT}_s$, the ITT effect of site $s$, can be unbiasedly estimated, using an estimator $\widehat{\text{ITT}}_s$ whose variance can also be unbiasedly estimated. Finally, in our asymptotic analysis, we assume that the number of randomization units in each site $n_s$ is fixed, while the number of sites $S$ goes to infinity, hereafter referred to as a “large $S$ small $n_s$” sequence. A common rule of thumb is that asymptotic approximations start being reliable when the index supposed to go to infinity exceeds 40 angrist2008mostly.\footnote{Of course, this rule of thumb is not always reliable, and researchers with more than 40 but less than, say, 100 sites in their RCT may want to conduct simulations taylored to their data to verify the coverage of the asymptotic confidence intervals we propose.} Under this rule of thumb, our “large $S$ small $n_s$”approximation is well suited to the multi-site RCTs in our survey: 10 out of 12 have at least 40 sites, while the median number of units per site is 12.5.

\paragraph{Estimating the variance of ITTs across sites.} As is well-known, to non-parametrically estimate the ITTs' variance across sites, one can use the Empirical Bayes (EB) variance estimator morris1983parametric. In a multi-site RCT, the EB estimator is equal to the variance of $\widehat{\text{ITT}}_s$ across sites, minus the average of robust variance estimators of the $\widehat{\text{ITT}}_s$ estimators.

\paragraph{Predicting site-specific ITT effects.} Our target parameter is $\bold{\beta}^{\text{ITT}}_{X}(\lambda)$, the coefficient from a ridge regression hoerl1970ridge of the site-specific ITTs on $\bold{X}_s$, a vector of predictors, with hyper-parameter $\lambda$. OLS is a special case of ridge, with $\lambda=0$. Ridge regressions can lead to more precisely estimated coefficients than OLS when the number of regressors is not negligible with respect to the sample size. This might be the case in multi-site RCTs, where one typically has a few dozens to a few hundreds of sites. Importantly, some elements of $\bold{X}_s$ might be unobserved variables that can be unbiasedly estimated. For instance, one may want to regress sites' ITTs on sites' outcomes without a treatment offer, to assess if treatment offers reduce or increase inequalities across sites. One could also be interested in regressing the ITTs for the main outcome variable on sites' ITTs for mediator variables, like in the job finding/job search example. To estimate $\bold{\beta}^{\text{ITT}}_{X}(\lambda)$, one cannot simply regress the estimated ITTs on the estimated covariates $\widehat{\bold{X}}_s$, due to the measurement error in the dependent and independent variables. However, this measurement error can be accounted for, as in an RCT one can unbiasedly estimate the variance of $(\widehat{\text{ITT}}_s,\widehat{\bold{X}}_s).$ We show that the resulting estimator $\widehat{\bold{\beta}}^{\text{ITT}}_{X}(\lambda)$ is asymptotically normal, and we provide an estimator of its asymptotic variance.

\paragraph{Predicting and estimating LATEs’ heterogeneity.} We start by showing that the sign of the correlation between the LATEs and any site-level characteristic is identified. This result can for instance be used to estimate the sign of the correlation between sites' FSs and LATEs, which could be useful to test if there is Roy selection across sites, whereby sites with the largest FSs are also those with the largest LATEs. Turning to the variance of LATEs, walters2015inputs has shown that a naive EB estimator using site-specific 2SLS estimators as building blocks is often negative and therefore uninformative on the LATEs' variance, because sites with first-stages (FSs) close to zero have large variances. Moreover, as the site-specific LATE estimators and their variance estimators are not unbiased, that estimator may not be consistent in the “large $S$ small $n_s$” sequence we consider. To bypass this issue, we provide two assumptions under which the LATEs' variance can be written as a function of sites' ITTs and FSs, and can thus be estimated leveraging only ITTs and FSs estimators. Our first assumption requires that sites' FSs and LATEs be independent. This is a strong assumption, that rules out Roy selection, but which is partly testable as the sign of the correlation between FSs and LATEs is identified. Our second assumption requires that the relationship between sites' FSs and LATEs is linear, and that LATEs' skewness is equal to zero.

\paragraph{Estimation of effect heterogeneity across strata in stratified RCTs.} Replacing the word “site” by the word “stratum” in all that precedes, our estimators can readily be used to estimate and predict effect heterogeneity across strata, in any stratified RCT with at least two treated and two control units per stratum. On the other hand, our estimators are not applicable to paired RCTs, which may lead researchers to prefer instead a design with, say, strata of four. In a finely stratified RCT, if one is ready to assume that the treatment effect does not vary within each stratum, the variance of treatment effects across strata is equal to the variance across randomization units. Then, our estimators can offer an alternative to methods directly taylored to study effect heterogeneity across units wager2018estimation. Investigating the pros and cons of both approaches may be an interesting question for future research.

\paragraph{Application.} We use our results to revisit behaghel2014private, who conducted an RCT to study the effect of intensive counseling programs on job seekers' employment, in more than 200 local public employment offices in France. The goal of their study is to compare the effectiveness of publicly- and privately-provided counseling. Accordingly, in each site job seekers are randomly assigned to either the control group, or to a program ran by the public employment service, or to a program ran by a private provider. This yields a fairly unique setting, where in each site, we can estimate the effect of two similar programs, ran by different providers. We leverage this feature to assess if the heterogeneity in programs' effects across sites is due to heterogeneity in providers' effectiveness. We find that while both programs increase job seekers' job finding rate by around 2 percentage points, the standard deviation of the ITT effects across sites is equal to 381% of the ITT estimate for the public program, and to 448% of the ITT estimate for the private one. Assuming that site-specific ITTs follow a normal distribution, the public and private programs respectively have a negative effect in 40% and 41% of the sites. We also find that the ITTs of the public and private programs are strongly positively correlated, thus suggesting that effects' heterogeneity is not entirely driven by providers’ effects. Surprisingly, sites' ITT effects are not significantly correlated with their FS effects. On the other hand, ITT effects are strongly negatively correlated with sites' average job-finding rate without treatment. We decompose sites' job-finding rate without treatment into a prediction based on their job-seekers' characteristics and a residual, and find that in a regression of their ITTs on these two variables, only the residual has some predictive power. Thus, the programs seem to be more effective in less tight local labor markets, and to increase their effectiveness, one could target them to the sites where earlier cohorts of job seekers had the lowest job finding rate. Turning to LATEs, we cannot reject the null that FSs and LATEs are uncorrelated, which is interesting in and of itself, and lends credibility to our first assumption to estimate the variance of the LATEs. Under that assumption, we estimate that the standard deviation of the effects across sites is equal to 364% of the LATE estimate for the public program, and to 432% for the private one.

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}*{Related literature and contributions}

\paragraph{Predicting site-specific ITT effects.} kline2022systemic is a fairly rare example of a multi-site RCT systematically investigating effect heterogeneity across sites (companies in their setting). In their Section 10, they use estimated site-specific ITTs as an explanatory variable in OLS regressions, using Bayesian shrinkage to account for measurement error. We show that measurement error can be accounted for non-parametrically. Deriving the asymptotic distribution of our estimators is also straightforward, another advantage with respect to regressions using posteriors from Bayesian shrinkage deeb2021framework. In the multi-site RCT literature, the most closely related paper is raudenbush2015learning, who discuss the estimation of the covariance between sites' ITTs and their average outcome without treatment (see their Equation (18)), without specifying explicitly how to unbiasedly estimate the variables' measurement error. In the teacher value-added (VA) literature, rose2022effects use teachers' estimated VA as an explanatory variable in OLS regressions. Building upon kline2020leave, they propose ideas similar to ours to account for measurement error. However, estimators of the variance of the measurement error differ in multi-site RCTs and in VA models, and are not numerically equivalent after some relabelling as we show in our application. Long before our and those papers, deaton1985panel had proposed to use repeated cross-sections to estimate a cohort-level panel, and use estimators of the variance of the cohort-level averages to account for measurement error when those averages are used as explanatory variables in regressions. Overall, our contribution is to slightly extend a result from li2017 to propose an unbiased estimator of the variance of $(\widehat{\text{ITT}}_s,\widehat{\bold{X}}_s)$, use that estimator to propose an estimator of $\bold{\beta}^{\text{ITT}}_{X}(\lambda)$, and derive the asymptotic distribution of $\widehat{\bold{\beta}}^{\text{ITT}}_{X}(\lambda)$. Another related paper is menzel2023transfer, who proposes functional-data methods to predict sites' effects based on observed covariates. Instead, our primary focus is on using unobserved variables that can be unbiasedly estimated to predict sites' effects. Relatedly, a vast literature studies meta-regressions, namely regressions of study-specific effects on moderators stanley2012meta. This literature mostly considers moderators that do not need to be estimated.

\paragraph{Estimating and predicting LATEs’ heterogeneity.} Other papers have tried to bypass the issue that a naive EB estimator cannot be used to estimate the variance of LATEs in “large $S$ small $n_s$” multi-site RCTs. walters2015inputs estimates a parametric random-coefficient model, while adusumilli2024heterogeneity estimate a parametric grouped-random-effect model. Instead, we pursue the complementary route of estimating that variance under non-parametric assumptions.

\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Set-up}

\paragraph{Completely randomized experiment, with at least two units assigned to treatment and control per site.} We consider a stratified RCT conducted in a fixed, finite population of $S$ sites. Site $s$ has $n_s$ units, and let $n=\sum_{s=1}^S n_s$ denote the total number of units in the RCT. Let $Z_{is}$ be an indicator for whether unit $i$ in site $s$ is assigned to treatment. $\textbf{Z}_s$ stacks all assignment indicators in site $s$.

hypFor all $s$, there exists $n_{1s}\in \{2,...,n_s-2\}$ such that for every $(z_1,...,z_{n_s})\in \{0,1\}^{n_s}$ such that $z_{1}+...+z_{n_s}=n_{1s}$, $P(\textbf{Z}_s=(z_1,...,z_{n_s}))=\frac{1}{{n_s \choose n_{1s}}}$.

\paragraph{Potential treatments, outcomes, and mediators.} For all $(i,s)\in\{1,...,n_s\}\times\{1,...,S\}$, the potential treatments of unit $i$ in site $s$ without and with assignment to treatment are denoted $D_{is}(0)$ and $D_{is}(1)$. Similarly, their potential outcomes without and with treatment are denoted $Y_{is}(0)$ and $Y_{is}(1)$.\footnote{This notation implicitly assumes that assignment to treatment has no direct effect on the outcome, the so-called exclusion restriction, see angrist1996identification.} Furthermore, we let $\bold{M}_{is}(0)$ denote a vector stacking the values of $m$ intermediate outcomes, or mediators, without treatment, while $\bold{M}_{is}(1)$ denotes the values of the mediators with treatment. Then to simplify notation let us introduce “reduced-form” potential outcome and mediators, that are functions of the assignment to treatment: $Y^r_{is}(0)=Y_{is}(D_{is}(0))$, $Y^r_{is}(1)=Y_{is}(D_{is}(1))$, $\bold{M}^r_{is}(0)=\bold{M}_{is}(D_{is}(0))$, and $\bold{M}^r_{is}(1)=\bold{M}_{is}(D_{is}(1))$. Finally, let $D_{is}=Z_{is}D_{is}(1)+(1-Z_{is})D_{is}(0)$, $Y_{is}=Z_{is}Y^r_{is}(1)+(1-Z_{is})Y^r_{is}(0)$, and $\bold{M}_{is}=Z_{is}\bold{M}^r_{is}(1)+(1-Z_{is})\bold{M}^r_{is}(0)$ denote the units' observed treatment, outcome, and mediators. We assume that potential treatments, outcomes, and mediators are independent and identically distributed (iid) in each site, independent of the treatment assignment in each site, and that potential treatments, outcomes, and mediators, as well as assignments, are independent across sites.

hyp\begin{enumerate} • For all $s$, the vectors $(D_{is}(0),D_{is}(1),Y_{is}(0),Y_{is}(1),\bold{M}_{is}(0),\bold{M}_{is}(1))$ are independent and identically distributed across $i$. • For all $s$, $(D_{is}(0),D_{is}(1),Y_{is}(0),Y_{is}(1),\bold{M}_{is}(0),\bold{M}_{is}(1))_{i\in \{1,...,n_s\}}\perp \!\!\! \perp \textbf{Z}_s$. • The random vectors $((D_{is}(0),D_{is}(1),Y_{is}(0),Y_{is}(1),\bold{M}_{is}(0),\bold{M}_{is}(1))_{i\in \{1,...,n_s\}},\textbf{Z}_s)$ are mutually independent across $s$. \end{enumerate}

Assumption (ref) for instance holds if in each site, the units included in the experiment are randomly drawn from a larger population. When units are not effectively drawn from a larger population, one can assume that such sampling took place. Then, all effects below apply to this hypothetical larger population, rather than to the study sample only. Assuming random sampling is convenient to avoid the well-known issue that in RCTs conducted in convenience samples, the variance of treatment-effect estimators is not identified Neyman1923. As potential treatments, outcomes, and mediators are assumed to be iid in each site, for all $s$ let $(D_{s}(0),D_{s}(1),Y_{s}(0),Y_{s}(1),\bold{M}_{s}(0),\bold{M}_{s}(1))$ denote a vector with the same probability distribution as $(D_{is}(0),D_{is}(1),Y_{is}(0),Y_{is}(1),\bold{M}_{is}(0),\bold{M}_{is}(1))$.

\paragraph{First-stage and intention-to-treat effects.} For all $s$ let

equation*[equation* omitted — 50 chars of source]

denote the first-stage (FS) effect in site $s$, and let

equation*[equation* omitted — 51 chars of source]

be a weighted average of the FSs across sites, for some non-negative and non-stochastic weights $w_s$ that sum to one. With $w_s=n_s/n$, $\text{FS}$ is the FS effect across units. With $w_s=1/S$, $\text{FS}$ is the FS effect across sites.\footnote{If the analysis is at a more disaggregated level than randomization units (e.g. the randomization is at the village level and stratified at the region level, but the analysis is at the villager level), $w_s$ could be proportional to the number of observations in site $s$.} Similarly, for all $s$ let

equation*[equation* omitted — 55 chars of source]

denote the intention-to-treat effect in site $s$, and let

equation*[equation* omitted — 54 chars of source]

Finally, for all $s$ let

equation*[equation* omitted — 85 chars of source]

denote the intention-to-treat effects on the mediators in site $s$, and let

equation*[equation* omitted — 86 chars of source]

\paragraph{Local average treatment effects.} As in Imbens1994, we assume that monotonicity holds and that the first-stage is strictly positive:

hypFor all $s$ $D_{s}(1)\ge D_{s}(0),$ and $\text{FS}>0.$

Then, for all $s$ such that $\text{FS}_s>0$, let

equation*[equation* omitted — 64 chars of source]

denote the local average treatment effect (LATE) in site $s$, and let

equation[equation omitted — 142 chars of source]

where the second equality follows from the definitions of $\text{ITT}$ and $\text{LATE}_s$.

\paragraph{FS, ITT, and LATE estimators.} For all $s$, let $n_{0s}=n_s-n_{1s}$ denote the number of untreated units in site $s$. For any generic variable $x_{is}$ defined for all $i\in\{1,...,n_s\}$ and $s\in\{1,...,S\}$, let $\overline{x}_s=\frac{1}{n_s}\sum_{i=1}^{n_s}x_{is}$ denote the average of $x_{is}$ in site $s$, let $\overline{x}_{1s}=\frac{1}{n_{1s}}\sum_{i=1}^{n_{s}}Z_{is}x_{is}$ and $\overline{x}_{0s}=\frac{1}{n_{0s}}\sum_{i=1}^{n_{s}}(1-Z_{is})x_{is}$ respectively denote the average of $x_{is}$ among the treated and untreated units in site $s$, and let $\overline{x}=\frac{1}{S}\sum_{s=1}\overline{x}_s$ denote the average of $x_{s}$ across sites. Then, let $\tilde{w}_s=Sw_s$ denote the weights re-scaled by the number of sites. For example, if $w_s=\frac{1}{S}$ then $\tilde{w}_s=1$ and if $w_s=\frac{n_s}{n}$ $\tilde{w}_s=\frac{n_s}{\overline{n}}$ where $\overline{n}$ is the average number of units per site. Finally, let

eqnarray[eqnarray omitted — 360 chars of source]

respectively denote the FS, ITTs, and LATE estimators in site $s$, and let

eqnarray[eqnarray omitted — 483 chars of source]

respectively denote the FS, ITTs, and LATE estimators across sites. Under Assumptions (ref) and (ref), $\widehat{\text{FS}}_s$, $\widehat{\text{ITT}}_s$, and $\widehat{\text{\bf{ITT}}}_{\text{M},s}$ are unbiased, so $\widehat{\text{FS}}$, $\widehat{\text{ITT}}$, and $\widehat{\text{\bf{ITT}}}_{\text{M}}$ are also unbiased.

\paragraph{Robust site-specific variance estimators.} For all $s\in\{1,...,S\}$, for any variable $x_{is}$ defined for every $i\in\{1,...,n_s\}$, let $r^2_{x,s}=\frac{1}{n_s-1}\sum_{i=1}^{n_s}(x_{is}-\overline{x}_s)^2$ denote the variance of $x_{is}$ in site $s$, and let $r^2_{x,1,s}=\frac{1}{n_{1s}-1}\sum_{i=1}^{n_{s}}Z_{is}(x_{is}-\overline{x}_{1s})^2$ and $r^2_{x,0,s}=\frac{1}{n_{0s}-1}\sum_{i=1}^{n_{s}}(1-Z_{is})(x_{is}-\overline{x}_{0s})^2$ respectively denote the variance of $x_{is}$ among the treated and untreated units in site $s$. Then let,

eqnarray[eqnarray omitted — 158 chars of source]

denote the robust estimator of the variance of $\widehat{\text{ITT}}_s$ eicker1963asymptotic,huber1967behavior,white1980heteroskedasticity. As is well-known imbens2015, under Assumptions (ref) and (ref),

equation[equation omitted — 155 chars of source]

Similarly,

eqnarray*[eqnarray* omitted — 130 chars of source]

is unbiased for $V\left(\widehat{\text{FS}}_s\right).$

\paragraph{Variances across sites.} As many of our target parameters are variances or covariances of vectors of real numbers across sites, we introduce a dedicated notation. Let $A^T$ denote the transpose of a matrix $A$. For any site-specific $K\times 1$ vector of real numbers $(\bold{U}_s)_{s\in \{1,...,S\}}$, let

equation*[equation* omitted — 182 chars of source]

denote the weighted variance matrix of those vectors across sites.

\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Application: the effects of publicly- and privately-provided counseling for job seekers.}

\paragraph{Study design and data.} behaghel2014private conduct a large-scale RCT, in 216 local Public Employment Service (PES) offices in France, to compare the public and private provision of counseling to job seekers. During their first interview at the local PES office, 43,977 job seekers are randomly assigned to one of three groups. The first group is a control group, where they receive the standard services provided by the PES. The second group is assigned to an intensive counseling program provided by the PES, and the third is assigned to an intensive counseling program provided by a private provider. Our framework is applicable to this RCT, with local public employment offices as sites and job seekers as randomization units. A first slight difference is that each unemployed has two assignment variables $Z_{1,is}$ and $Z_{2,is}$, respectively equal to one if they are assigned to the PES-operated and to the privately-operated program. This difference is immaterial for our results. For instance, if one is interested in the heterogeneous effects of the PES-provided program, in the estimators defined below one lets $Z_{is}$ stand for $Z_{1,is}$, and one drops job seekers assigned to the privately-operated program from the sample.\footnote{In particular, it follows from Theorem 3 in li2017 that the formulas we use below for the variances of treated versus control comparisons still apply to RCTs with more than two treatments.} A second slight difference is that for the private program, 12 offices have less than two treated or two control units: they have to be dropped from our analysis. For the public program, 16 offices have to be dropped for the same reason. Compliance with randomized assignment is imperfect. While almost no job seekers unassigned to the counseling programs gets access to them, only 32% (resp. 43%) of job seekers assigned to the public (resp. private) counseling program took it up. The outcome we consider is an indicator for holding any employment 6 months after randomization, one of the three main employment outcomes considered by the authors. Results are similar if we consider the authors' two other outcomes.

\paragraph{Study's strengths and weaknesses for our purposes.} Unfortunately, the authors' data set does not contain mediators, such as measures of workers' job-search effort, thus precluding us from conducting “predictive mediation” analyses. Moreover, as randomization takes place within local-labor markets, the programs may generate displacement effects, and their ITTs are partial rather than general equilibrium effects. On the other hand, this study exhibits a rare feature: in each site we can estimate the effect of two similar programs ran by different providers. This will help us assess if effects' heterogeneity is due to heterogeneity in providers' effectiveness.

\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Estimating and predicting ITTs' and FSs' heterogeneity.}

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Estimating the variance of ITTs and FSs across sites.}

\paragraph{Target parameters.} In this section, our target parameter is $\sigma^2\left[\text{ITT}\right]$, the variance of the ITTs across sites. The variance of the FS effects and the variances of the ITT effects on the mediators can be estimated similarly.

\paragraph{Estimating $\sigma^2\left[\text{ITT}\right]$ using an Empirical Bayes estimator.} Let $$\widehat{\sigma}^2\left[\text{ITT}\right]=\sum_{s=1}^S w_s \left[\left(\widehat{\text{ITT}}_s-\widehat{\text{ITT}}\right)^2-\widehat{V}_{rob}\left(\widehat{\text{ITT}}_s\right)\right].$$ $\widehat{\sigma}^2\left[\text{ITT}\right]$ is the standard Empirical Bayes (EB) variance estimator morris1983parametric, applied to multi-site RCTs. In RCTs stratified at a finer level than the sites, the variance of ITTs across sites can still be estimated by replacing, in the definition of $\widehat{\sigma}^2\left[\text{ITT}\right]$, $\widehat{V}_{rob}\left(\widehat{\text{ITT}}_s\right)$ by a weighted sum of the robust variance estimators across the strata of site $s$.

\paragraph{Asymptotic distribution of the EB estimator.} Let $\phi_{s,1}=\tilde{w}_s\left[\left(\widehat{\text{ITT}}_s-\text{ITT}\right)^2-\widehat{V}_{rob}\left(\widehat{\text{ITT}}_s\right)\right]$.

hypSufficient conditions under which $\widehat{\sigma}^2\left[\text{ITT}\right]$ is asymptotically normal. \begin{enumerate} • The sequences $\left(\tilde{w}_{s}\widehat{\text{ITT}}_s\right)_{s\geq 1}$ and $\left(\phi_{s,1}\right)_{s\geq 1}$ satisfy the Lyapunov condition. • For all $s$, $\tilde{w}_s<N$ for some $N>0$ and $N<+\infty$. • $\text{ITT}$, $\frac{1}{S}\sum_{s=1}^{S} V\left(\phi_{s,1}\right)$, $\frac{1}{S}\sum_{s=1}^{S} E(\phi_{s,1})$, $\frac{1}{S}\sum_{s=1}^{S}E(\phi^2_{s,1})$ converge towards finite limits when $S\rightarrow \infty$. \end{enumerate}

Point 1 of Assumption (ref) requires that one can apply the Lyapunov central limit theorem to $\widehat{\text{ITT}}$ and to an infeasible version of $\widehat{\sigma}^2\left[\text{ITT}\right]$ where $\widehat{\text{ITT}}$ is replaced by $\text{ITT}$. Point 2 of Assumption (ref) requires that the rescaled weights for each site be bounded. Finally, Point 3 requires that certain deterministic averages have finite limits. Under Assumption (ref), let

align*[align* omitted — 145 chars of source]

and let $\widehat{\phi}_{s,1}=\tilde{w}_s\left[\left(\widehat{\text{ITT}}_s-\widehat{\text{ITT}}\right)^2-\widehat{V}_{rob}\left(\widehat{\text{ITT}}_s\right)\right]$ and $$\widehat{V}_{\sigma^2\left[\text{ITT}\right]}=\frac{1}{S}\sum_{s=1}^{S}\left[\widehat{\phi}_{s,1}- \overline{\widehat{\phi}_1} \right]^2.$$

thmIf Assumptions (ref), (ref), and (ref) hold, \begin{align*} \sqrt{S}\left(\widehat{\sigma}^2\left[ITT\right]-\sigma^2\left[ITT\right]\right) \overset{d}{\longrightarrow} N(0,V_{\sigma^2\left[ITT\right]}), \end{align*} and $\widehat{V}_{\sigma^2\left[\text{ITT}\right]}\overset{\mathbb{P}}{\longrightarrow}\overline{v}$, where $\overline{v}$ is a real number larger than $V_{\sigma^2\left[\text{ITT}\right]}$ defined in the proof.

Theorem (ref) shows that in the “large $S$ fixed $n_s$” asymptotic sequence we consider, $\widehat{\sigma}^2\left[\text{ITT}\right]$ is asymptotically normal for $\sigma^2\left[\text{ITT}\right]$, and $\widehat{V}_{\sigma^2\left[\text{ITT}\right]}$ is a conservative estimator of its asymptotic variance. Thus, Theorem (ref) can be used to obtain conservative confidence intervals for $\sigma^2\left[\text{ITT}\right]$. The conservativeness of $\widehat{V}_{\sigma^2\left[\text{ITT}\right]}$ is due to the fact we assume that the $S$ sites we observe are a fixed population. If one were to assume instead that the $S$ sites are a random sample from a super-population of sites, we conjecture that $\widehat{V}_{\sigma^2\left[\text{ITT}\right]}$ would not be conservative anymore.

\paragraph{Application: the variance across sites of the ITT effects of publicly- and privately-provided counseling.} In Table (ref), we start by estimating the ITT effect of each treatment. On average across all sites, both programs increase job seekers' employment rate after six months by around two percentage points (pp).\footnote{Effects very slightly differ from those in the paper, owing to the slightly different estimation sample.} However, this hides very substantial heterogeneity across sites. $\widehat{\sigma}^2\left[\text{ITT}\right]$ is large and significantly different from zero for both programs. $\sqrt{\widehat{\sigma}^2\left[\text{ITT}\right]}/\widehat{\text{ITT}}=$ 381% for the public program, and 448% for the private one. This is a very substantial amount of treatment effect heterogeneity. For instance, assuming for illustrative purposes that site-specific ITTs follow a truncated normal,\footnote{The outcome is binary so ITTs have to belong to $[-1,1]$.} where the underlying untruncated distribution has a mean equal to $\widehat{\text{ITT}}$ and a standard deviation equal to $\sqrt{\widehat{\sigma}^2\left[\text{ITT}\right]}$, the public program has a negative effect in 40% of the sites, while the private program has a negative effect in 41% of them. We also re-estimate the variance of the ITT effects of the public program using the estimator of kline2020leave, in the special case described in their Example 2 with a single binary regressor, in which case the target parameter coincides with $\sigma^2\left[\text{ITT}\right]$. Doing so, we obtain an estimator around 20% smaller than our estimator, thus showing that the two approaches do not coincide after some relabeling.\footnote{In our calculations, we divided $\tilde{z}_i$ by $T_g$ in their covariance representation equation page 1868, as we interpreted the missingness of $T_g$ as a typo. Without that change, their estimator is 50 times smaller than ours.}

table[table omitted — 1,712 chars of source]

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Predicting site-specific ITT and FS effects}

\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont}{Theory}

\paragraph{Target parameter.} Let $\bold{X}_s$ denote a $K\times 1$ vector of site-level variables, which we want to use to predict sites' ITTs. $\bold{X}_s$ may include observed variables, like some baseline covariates of site $s$. $\bold{X}_s$ may also include unobserved variables that have to be estimated. Let $\mu(\bold{X})=\sum_{s=1}^Sw_s \bold{X}_s$, and let $\bold{I}_K$ denote the $K\times K$ identity matrix. Assuming that $\sigma^2[\bold{X}]+\lambda \bold{I}_K$ is invertible, our target is $$\bold{\beta}^{\text{ITT}}_{X}(\lambda) \equiv\left(\sigma^2[\bold{X}]+\lambda \bold{I}_K\right)^{-1}\left(\sum_{s=1}^S w_s\left(\bold{X}_s-\mu(\bold{X})\right)\left(\text{ITT}_s-\text{ITT}\right)\right),$$ the coefficients on $\bold{X}_s $ in a Ridge regression of the demeaned $\text{ITT}_s$ on the demeaned $\bold{X}_s$, weighted by $w_s$, and with hyper-parameter $\lambda$. $\bold{\beta}^{\text{ITT}}_{X}(0)$ is a standard OLS regression coefficient, denoted $\bold{\beta}^{\text{ITT}}_{X}$. When $\lambda=0$, an auxiliary target is the R-squared of the OLS regression, $$\text{R}^{\text{ITT}}_{X}\equiv\frac{\left(\bold{\beta}^{\text{ITT}}_{X}\right)^T\sigma^2[\bold{X}]\bold{\beta}^{\text{ITT}}_{X}}{\sigma^2[\text{ITT}]}.$$

\paragraph{Connection with regressions of unit-specific effects on unit-specific predictors.} Let $\bold{\beta}^{\text{ITT}_{i}}_{X_i}$ denote the coefficient from an (infeasible) regression of unit-specific ITT effects on unit-specific predictors $\bold{X}_{i,s}$, whose average in site $s$ is equal to $\bold{X}_{s}$. When $\bold{X}_{i,s}$ is of dimension one, it follows from the law of total covariance that in general $\bold{\beta}^{\text{ITT}_{i}}_{X_i}\ne \bold{\beta}^{\text{ITT}}_{X}$, and the coefficients could even be of a different sign, a version of the so-called ecological inference problem. Thus, regressions of site-specific ITTs on site-specific covariates can in general not be used to infer the coefficients from regressions of unit-specific ITTs on unit-specific covariates. A first exception is when $\bold{X}_{i,s}$ does not vary within sites, in which case $\bold{\beta}^{\text{ITT}_{i}}_{X_i}= \bold{\beta}^{\text{ITT}}_{X}$. A second exception is when the unit-specific ITT effects do not vary within sites, in which case the coefficients are of the same sign and $|\bold{\beta}^{\text{ITT}_{i}}_{X_i}|\leq |\bold{\beta}^{\text{ITT}}_{X}|:$ the site-level coefficient is always further away from zero than the unit-level one. When the estimators in this paper are applied not to a multi-site RCT, but to a finely stratified RCT, say with strata of four, where the stratification is based on predictors of $\bold{X}_{i,s}$ or of the unit-specific ITTs, it might be reasonable to assume that $\bold{X}_{i,s}$ or the unit-specific ITT effects do not vary within strata.

\paragraph{Leading examples of unobserved variables one might want to include in $\bold{X}_s$.} We have four leading examples in mind of potentially interesting unobserved variables one might want to include in $\bold{X}_s$. The first one is $\text{FS}_s$, the first-stage effect in site $s$. For instance, one can use the regression of $\text{ITT}_s$ on $\text{FS}_s$ to test the null that LATEs do not vary across sites: this null holds if and only if the regression's intercept is equal to zero while its R-squared is equal to one, an equivalence already noted by walters2015inputs though the chi-squared test therein is not applicable to the small $n_s$ applications we consider. The second unobserved variable one might want to include in $\bold{X}_s$ is $E(Y_s^r(0))$, the average outcome in the control group. Regressing $\text{ITT}_s$ on $E(Y_s^r(0))$ is a way to assess if ITTs are larger or lower in sites with the lowest control outcomes, to assess if treatment offers reduce or increase inequalities across sites. The third one is $\text{\bf{ITT}}_{\text{M},s}$, the site-specific ITT effects on mediator variables. Regressing $\text{ITT}_s$ on $\text{\bf{ITT}}_{\text{M},s}$ is a way to do “predictive mediation” analysis, by assessing if sites with large effects on the mediators also tend to have large effects on the final outcome. Of course, this type of mediation analysis remains predictive and not causal: larger effects in sites with larger mediator effects could be due to omitted variables rather than the mediator themselves. The fourth one is $\text{ITT}_{2,s}$, the site-specific ITT effect of a second assignment variable $Z_{2,is}$, as in our empirical application where in each site job seekers can be randomly assigned to two treatments. When the two treatments are similar interventions delivered by different providers, regressing $\text{ITT}_{s}$ on $\text{ITT}_{2,s}$ can be a way to suggestively test if the heterogeneity in $\text{ITT}_{s}$ is due to provider effects. When the two treatments are different interventions, regressing $\text{ITT}_{s}$ on $\text{ITT}_{2,s}$ can be a way to assess if targeting should be intervention specific.

\paragraph{Unbiased estimators of $\bold{X}_s$.} As explained above, $\bold{X}_s$ may include unobserved variables, that need to be estimated. Then, we assume that we have an unbiased estimator of $\bold{X}_s$, denoted $\widehat{\bold{X}}_s$, that is a function of $((D_{is}(0),D_{is}(1),Y_{is}(0),Y_{is}(1),\bold{M}_{is}(0),\bold{M}_{is}(1))_{i\in \{1,...,n_s\}},\textbf{Z}_s)$ and known real numbers. Of course, for all coordinates $X_{k,s}$ of $\bold{X}_s$, that are observed and do not need to be estimated, $\widehat{X}_{k,s}=X_{k,s},$ so $\widehat{X}_{k,s}$ is non-stochastic. We let $\widehat{\mu}(\bold{X})=\sum_{s=1}^Sw_s \widehat{\bold{X}}_s$. Letting $\widehat{X}_{k,s}$ denote the $k$th coordinate of $\widehat{\bold{X}}_s$, we assume that for all $k\in \{1,...,K\}$ we also have unbiased estimators of $\text{Cov}\left(\widehat{X}_{k,s},\widehat{\text{ITT}}_s\right)$, denoted $\widehat{\text{Cov}}\left(\widehat{X}_{k,s},\widehat{\text{ITT}}_s\right)$, and we let $\widehat{\text{Cov}}\left(\widehat{\bold{X}}_{s},\widehat{\text{ITT}}_s\right)$ denote a vector stacking those estimators. Finally, we assume that we have an unbiased estimator of $V\left(\widehat{\bold{X}}_{s}\right)$, denoted $\widehat{V}\left(\widehat{\bold{X}}_{s}\right)$. Those conditions are satisfied in our four leading examples. For all $s\in\{1,...,S\}$, for any variables $q_{is}$ and $x_{is}$, $c_{q,x,s}=\frac{1}{n_s-1}\sum_{i=1}^{n_s}(q_{is}-\overline{q}_s)(x_{is}-\overline{x}_s)$ denotes the covariance between $q_{is}$ and $x_{is}$ in site $s$, and $c_{q,x,1,s}=\frac{1}{n_{1s}-1}\sum_{i=1}^{n_{s}}Z_{is}(q_{is}-\overline{q}_{1s})(x_{is}-\overline{x}_{1s})$ and $c_{q,x,0,s}=\frac{1}{n_{0s}-1}\sum_{i=1}^{n_{s}}(1-Z_{is})(q_{is}-\overline{q}_{1s})(x_{is}-\overline{x}_{0s})$ denote the covariance between $q_{is}$ and $x_{is}$ among treated and untreated units in site $s$.

lemIf Assumptions (ref) and (ref) hold, \begin{enumerate} • $E\left(\widehat{\text{FS}}_s\right)=\text{FS}_s$, $E\left(\frac{c_{D,Y,0,s}}{n_{0,s}}+\frac{c_{D,Y,1,s}}{n_{1,s}}\right)=\text{Cov}\left(\widehat{\text{FS}}_{s},\widehat{\text{ITT}}_s\right)$, and $E\left(\frac{r^2_{D,0,s}}{n_{0,s}}+\frac{r^2_{D,1,s}}{n_{1,s}}\right)=V\left(\widehat{\text{FS}}_s\right)$. • $E\left(\overline{Y}_{0s}\right)=E(Y_s^r(0))$, $E\left(-\frac{r^2_{Y,0,s}}{n_{0,s}}\right)=\text{Cov}\left(\overline{Y}_{0s},\widehat{\text{ITT}}_s\right)$, and $E\left(\frac{r^2_{Y,0,s}}{n_{0,s}}\right)=V\left(\overline{Y}_{0s}\right)$. • $E\left(\widehat{\text{\bf{ITT}}}_{\text{M},s}\right)=\text{\bf{ITT}}_{\text{M},s}$, for all $k\in \{1,...,K\}$ $E\left(\frac{c_{M_k,Y,0,s}}{n_{0,s}}+\frac{c_{M_k,Y,1,s}}{n_{1,s}}\right)=\text{Cov}\left(\widehat{\text{ITT}}_{\text{M}_k,s},\widehat{\text{ITT}}_s\right)$, and for all $(k,k')\in \{1,...,K\}^2$ $E\left(\frac{c_{M_k,M_{k'},0,s}}{n_{0,s}}+\frac{c_{M_k,M_{k'},1,s}}{n_{1,s}}\right)=\text{Cov}\left(\widehat{\text{ITT}}_{\text{M}_k,s},\widehat{\text{ITT}}_{\text{M}_{k'},s}\right)$, and $E\left(\frac{r^2_{\text{M}_{k},0,s}}{n_{0,s}}+\frac{r^2_{\text{M}_{k},1,s}}{n_{1,s}}\right)=V\left(\widehat{\text{ITT}}_{\text{M}_k,s}\right)$ • Letting $n_{2,s}$ denote the number of units assigned to the second treatment in site $s$, and $r^2_{Y,2,s}$ denote the outcome variance across those units, $E\left(\widehat{\text{ITT}}_{2,s}\right)=\text{ITT}_{2,s}$, $E\left(-\frac{r^2_{Y,0,s}}{n_{0,s}}\right)=\text{Cov}\left(\widehat{\text{ITT}}_s,\widehat{\text{ITT}}_{2,s}\right)$, and $E\left(\frac{r^2_{Y,0,s}}{n_{0,s}}+\frac{r^2_{Y,2,s}}{n_{2,s}}\right)=V\left(\widehat{\text{ITT}}_{2,s}\right)$. \end{enumerate}

Lemma (ref) follows from Theorem 3 in li2017, who derive, conditional on potential outcomes, the variance of the vector of ITT estimators on several outcomes, in a potentially multi-armed RCT.

\paragraph{Estimator of $\bold{\beta}^{\text{ITT}}_{X}(\lambda)$.} We let $$\widehat{\bold{\beta}}^{\text{ITT}}_{X}(\lambda)=\left(\sigma^2\left[\widehat{\bold{X}}\right]-\sum_{s=1}^S w_s\widehat{V}\left(\widehat{\bold{X}}_{s}\right)+\lambda \bold{I}_K\right)^{-1}\left(\sum_{s=1}^S w_s\left(\left(\widehat{\bold{X}}_s-\widehat{\mu}(\bold{X})\right)\left(\widehat{\text{ITT}}_s-\widehat{\text{ITT}}\right)-\widehat{\text{Cov}}\left(\widehat{\bold{X}}_{s},\widehat{\text{ITT}}_s\right)\right)\right).$$ Similarly, when $\lambda=0$, we let $$\widehat{\text{R}}^{\text{ITT}}_{X}=\frac{\left(\widehat{\bold{\beta}}^{\text{ITT}}_{X}\right)^T\widehat{\sigma}^2[\bold{X}]\widehat{\bold{\beta}}^{\text{ITT}}_{X}}{\widehat{\sigma}^2[\text{ITT}]}$$ denote the estimator of $\text{R}^{\text{ITT}}_{X}$.

\paragraph{Intuition for the estimator.} Without the terms involving $\widehat{V}\left(\widehat{\bold{X}}_{s}\right)$ and $\widehat{\text{Cov}}\left(\widehat{\bold{X}}_{s},\widehat{\text{ITT}}_s\right)$, $\widehat{\bold{\beta}}^{\text{ITT}}_{X}(\lambda)$ would just be the coefficient on $\widehat{\bold{X}}_s$ in a naive Ridge regression of the demeaned $\widehat{\text{ITT}}_s$ on the demeaned $\widehat{\bold{X}}_s$. Due to measurement error, the naive regression suffers from a standard attenuation bias, biasing the coefficient towards zero. As the dependent variable is also measured with error, the naive regression can also suffer from an additional bias, whose direction is unknown, if the measurement error in $\widehat{\bold{X}}_s$ is correlated to that in $\widehat{\text{ITT}}_s$. In multi-site RCTs, correcting for those two biases is easy, as one can unbiasedly estimate the variance of $\widehat{\bold{X}}_s$ and its covariance with $\widehat{\text{ITT}}_s$. This is exactly the role of the terms involving $\widehat{V}\left(\widehat{\bold{X}}_{s}\right)$ and $\widehat{\text{Cov}}\left(\widehat{\bold{X}}_{s},\widehat{\text{ITT}}_s\right).$

\paragraph{Consistency and asymptotic normality.} Let

align*[align* omitted — 706 chars of source]
align*[align* omitted — 546 chars of source]

and let $V_{\bold{\beta}^{\text{ITT}}_{X}(\lambda)}$ denote the limit of $\frac{1}{S}\sum_{s=1}^S V\left(\phi_{s,4}\right)$, which is assumed to exist in Assumption (ref) in the Appendix.

thmSuppose that Assumptions (ref) and (ref) hold, and that the technical conditions in Assumption (ref) in the Appendix hold. Then, $$\widehat{\bold{\beta}}^{\text{ITT}}_{X}(\lambda)-\bold{\beta}^{\text{ITT}}_{X}(\lambda) \overset{\mathbb{P}}{\longrightarrow}0,$$ and \begin{align*} \sqrt{S}\left(\widehat{\bold{\beta}}^{ITT}_{X}(\lambda)-\bold{\beta}^{ITT}_{X}(\lambda)\right) \overset{d}{\longrightarrow} N(0,V_{\bold{\beta}^{ITT}_{X}(\lambda)}). \end{align*}

Let

align*[align* omitted — 676 chars of source]

We conjecture that using similar steps as in the proof of Theorem (ref), one can show that $\widehat{V}_{\bold{\beta}^{\text{ITT}}_{X}(\lambda)}$, the sample variance of $\widehat{\phi}_{s,4}$, is a conservative estimator of $V_{\bold{\beta}^{\text{ITT}}_{X}(\lambda)}$.\footnote{For a vector, a conservative variance estimator means that for any $K\times 1$ vector of real numbers $\theta$, $\theta' \widehat{V}_{\bold{\beta}^{\text{ITT}}_{X}(\lambda)} \theta$ converges to a limit weakly larger than that of $\theta' V_{\bold{\beta}^{\text{ITT}}_{X}(\lambda)} \theta$.}

\paragraph{Choice of hyper-parameter.} golub1979generalized propose to use a generalized cross-validation (GCV) method to choose $\lambda$. Applying their Equation (1.4) to our multi-site RCT setting, rewriting explicitly the inner product in the numerator and using the linearity and cyclicality of the trace operator to rewrite the denominator, GCV amounts to using $\lambda^*$, the minimizer of

equation[equation omitted — 283 chars of source]

where $\text{Tr}(.)$ denotes the trace operator. (ref) makes it clear that for any $\lambda$, $V(\lambda)$ can be consistently estimated, replacing $\sigma^2[ITT]$, $B$, $A(\lambda)$, and $\sigma^2[\bold{X}]$ by their estimators. Accordingly, we propose to use $\widehat{\lambda}^*$, the minimizer of $\widehat{V}(\lambda)$. While it should be feasible to derive the asymptotic variance of $\widehat{\bold{\beta}}^{\text{ITT}}_{X}\left(\widehat{\lambda}^*\right)$ using standard results from M-estimation, for now we rely on the bootstrap.

\paragraph{Estimating a LASSO regression coefficient?} A natural question is whether one could also estimate the coefficients from a LASSO regression santosa1986linear,tibshirani1996regression of the ITTs on $\bold{X}_s$. With respect to Ridge, LASSO sets the coefficients of the least significant predictors to zero, thus yielding a more-interpretable vector of coefficients with a small number of non-zero entries. loh2011high and sorensen2015measurement propose a regularized-corrected LASSO estimator, when independent variables are measured with error. In our setting, their estimator amounts to minimizing

equation[equation omitted — 263 chars of source]

with respect to $b$, where $||b||_1$ is the $L^1$ norm of $b$. This loss function does not account for the measurement error in the dependent variable, which could maybe be achieved by minimizing

equation[equation omitted — 374 chars of source]

instead.\footnote{Note that with $\lambda=0$, the minimizer of (ref) is the OLS estimator $\widehat{\bold{\beta}}^{\text{ITT}}_{X}$.} To our knowledge, LASSO regressions with measurement error in both the dependent and independent variables have not been studied yet. Accounting for measurement error in the dependent variables alone is not trivial and is still an active area of research datta2017cocolasso, as the loss function in (ref) is non-convex when the number of regressors is strictly larger than the number of observations loh2011high. Overall, the extension to LASSO regressions is not a straightforward one.

\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont}{Application: predicting site-specific ITT effects of the publicly- and privately-provided counseling programs.}

\paragraph{The ITTs of the public and private programs are positively correlated.} Table (ref) reports several univariate OLS regressions of sites' ITT effects on predictors. In Panel A Column (1), we find a strong positive correlation between the ITTs of the public and private programs, with an estimated R2 of almost 0.3. In each site, the two programs are delivered by different providers. Therefore, this suggests that the heterogeneity in sites' ITT effects is unlikely to be entirely driven by providers' effects. In another regression not shown in the table, we find an even stronger positive correlation between the FSs of the public and private programs, with an estimated R2 of 0.6.

\paragraph{Sites' FSs do not predict their ITTs.} In Column (2), we regress sites' ITTs on their FSs. While FSs varies across sites (sd = 11.4pp for the public program, 11.8pp for the private program, see Table (ref) below), FSs are not significantly correlated with ITTs.

\paragraph{Sites' average outcome without treatment strongly predict their ITTs.} In Column (3), we regress sites' ITTs on $E(Y_s^r(0))$, their average outcome without a treatment offer. As less than 5% of control job seekers receive one of the two treatments, $E(Y_s^r(0))$ is essentially sites' outcome without treatment. The estimated standard deviation of the control group's job finding rate is quite large (13.3pp), and the ITTs of both programs are negatively correlated with that variable. For the private program, the regression's estimated R2 is almost 0.5.

\paragraph{The correlation between ITTs and $E(Y_s^r(0))$ is not due to heterogeneous job-seeker characteristics across sites.} In Table (ref) we regress ITTs on sites’ predicted job finding rate without treatment given their job seekers’ characteristics, and the residual of that prediction. To predict the job finding rate without treatment, we follow behaghel2014private and estimate a job-seeker level logistic regression of whether they find a job on 43 job-seeker level variables, measuring their educational levels, their prior work and unemployment history, their demographics, and their reservation wage. Many variables are statistically significant, and the regression's pseudo-R2 is equal to 0.08.\footnote{Using a LASSO logistic regression instead yields extremely similar predictions. Similarly, adding site fixed effects to estimate the covariates' coefficients, thus ensuring that those coefficients are only estimated out of variation between workers within sites, also yields extremely similar predictions. That last regression has to be estimated with OLS, to avoid an incidental parameter problem.} While the predicted job finding rate varies across sites (sd = 3.5pp), it is not significantly correlated to ITTs, unlike the residual. Thus, the correlation between ITTs and the job finding rate without treatment is not due to heterogeneous job-seeker characteristics across sites. In their Table 8, behaghel2014private show that the private program is less effective among jobseekers' with a higher predicted job finding rate. While the predicted job finding rate predicts heterogeneous effects at the individual level, our analysis shows that the average of that variable at the site level does not predict the site's effect, thus exemplifying the so-called ecological inference problem.

\paragraph{The local unemployment rate does not predict sites' ITTs, but it predicts $E(Y_s^r(0))$.} In Column (4) of Table (ref), we regress ITTs on sites' local unemployment rate, which we could retrieve for all but one site.\footnote{Specifically, we matched the data of behaghel2014private to a dataset produced by the French National Office of Statistics, with unemployment rates at the city level in 2007, the year when the RCT was conducted.} While the unemployment rate varies across sites (sd = 4.4pp), it is not correlated with sites' ITTs. This may seem to contradict the results in Column (3), but Table (ref) shows that while the control-group job finding rate is negatively and significantly correlated with the local unemployment rate, the correlation between the two variables is not perfect (R2=$0.10$ in the private program sample, and R2=$0.13$ in the public program sample). Then, the local unemployment rate may be an imperfect proxy of the labor market conditions faced by the job seekers eligible for this RCT, namely those at high risk of long-term employment.

\paragraph{Using the correlation between ITTs and $E(Y_s^r(0))$ to improve the targeting of the program?} The strong negative correlation between ITTs and $E(Y_s^r(0))$ may be used to better target the program. While $E(Y_s^r(0))$ is not observed ex-ante, one could use, as a proxy for $E(Y_s^r(0))$, the job finding rate of an earlier cohort of job seekers in each site, restricting attention to job seekers that would have been eligible for the program if the program had been available when their unemployment spell started. Moreover, finding predictors of $E(Y_s^r(0))$ may be easier than finding predictors of the ITTs, as $E(Y_s^r(0))$ is estimated with less error than the ITTs athey2023machine.

\paragraph{Comparing our regression coefficients $\widehat{\bold{\beta}}^{\text{ITT}}_{X}$ to naive ones.} At the bottom of each column of Table (ref), we show coefficients from naive OLS regressions, that do not account for the measurement error in the variables. When the explanatory variable is estimated (Columns (1), (2), and (3)), the coefficient of the naive regression differs from $\widehat{\bold{\beta}}^{\text{ITT}}_{X}$, and its standard error is much smaller. When the characteristic is not estimated (Column (4)), the naive regression leads to the same coefficient and a very slightly different standard error.

table[table omitted — 2,035 chars of source]
table[table omitted — 1,512 chars of source]
table[table omitted — 745 chars of source]
comment\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont}{Other perhaps useful things.} For the Crepon et al application for which we just applied for the data: the RCT has around 50 strata of 5 PES. In each strata, one PES such that 100% of people treated, one PES such that 75% treated, etc all the way to 0%. We could look at the variance of effects across strata, which is a lower bound of the variance of effects across PES. Great advantage of doing so: for each strata, we can have an unbiased estimator of $E(Y_s(1,1)-Y_s(0,0))$, the average effect of assigning everyone versus no one to treatment in the PES. We just compare the fully treated and fully untreated PES. Difficulty: unbiasedly estimating the variance of that estimator, as we have only one treated and one control unit per strata. Proposal: to estimate the variance of $Y_s(1,1)$, we use the variance between average outcome in the fully treated PES and the average outcome across the 75% treated in the 75% treated PES. Calculations attached, to be double checked, suggest that this yields upward biased estimator (and therefore downward biased estimator of TE variance across strata), if for instance $Y_s(1,1)-Y_s(1,0.75)$ is uncorrelated with, or negatively correlated with, $Y_s(1,1)$. For the Crepon et al 2015 paper, they write: “On average, two pairs per branch were kept for the evaluation. In each pair, one village was randomly assigned to treatment, and the other to control. In total, 81 pairs belonging to 47 branches were included in the evaluation.” If branches either have one or two pairs, then we have 34 branches with 2 pairs, and 13 branches with 1 pair. We could regroup the 13 branches with 1 pair into 5 pairs and 1 triple of most closely located branches, to reach exactly 40 groups, with at least two treated and two control units. Now the issue is that in each group, the 2 treated and 2 controls are not chosent at random out of the 4, so we do not have that variance of outcome across the two treated is unbiased for the variance of the treated outcome in the group. Let $\overline{Y}_{p,s}(z)$ and $\overline{Y^2}_{p,s}(z)$ denote the average potential outcome and squared PO with assignment $z$ in pair $p$ of site $s$, and let $Y^z_{p,s}$ denote the outcome of unit with assignment $z$ in pair $p$ of site $s$. As the DOF corrected variance of two numbers is 1/2 times the square of their difference, we have that $$r^2_{Y,z,s}=1/2\left(Y^z_{1,s}-Y^z_{2,s}\right)^2.$$ \begin{align*} &E\left(1/2\left(Y_{1,s}^z-Y_{2,s}^z\right)^2|PO\right)\\ =&1/2E\left(\left(Y_{1,s}^z\right)^2|PO\right)+1/2E\left(\left(Y_{2,s}^z\right)^2|PO\right)-E\left(Y_{1,s}^z|PO\right)E\left(Y_{2,s}^z|PO\right)\\ =&1/2\overline{Y^2}_{1,s}(z)+1/2\overline{Y^2}_{2,s}(z)-\overline{Y}_{1,s}(z)\overline{Y}_{2,s}(z)\\ =&1/2\left(\overline{Y^2}_{1,s}(z)-\left(\overline{Y}_{1,s}(z)\right)^2\right)+1/2\left(\overline{Y^2}_{2,s}(z)-\left(\overline{Y}_{2,s}(z)\right)^2\right)+1/2\left(\overline{Y}_{1,s}(z)-\overline{Y}_{2,s}(z)\right)^2\\ =&1/2r^2_{Y(1),1,s}+1/2r^2_{Y(1),2,s}+r^2_{\overline{Y}(1),s} \end{align*} First equality uses independence of $Y_{1,s}^z$ and $Y_{2,s}^z$. In last equality, $r^2_{Y(1),p,s}:$ non-DOF adjusted variance of $Y_{ip,s}(1)$ in pair $p$ of site $s$, while $r^2_{\overline{Y}(1),s}$, DOF adjusted variance of $\overline{Y}_{p,s}(z)$ across the two pairs. \begin{align*} &r^2_{Y(1),s}\\ =&4/3\left(1/4 \sum_{p=1}^2\sum_{i=1}^2\left(Y_{i,p,s}(1)-\overline{Y}_{s}(1)\right)^2\right)\\ =&4/3 \left(1/2r^2_{Y(1),1,s}+1/2r^2_{Y(1),2,s}+1/2 r^2_{\overline{Y}(1),s}\right), \end{align*} where last equality follows from law of total variance / within-between variance decomposition. Therefore, $$E\left(1/2\left(Y_{1,s}^z-Y_{2,s}^z\right)^2|PO\right)\geq r^2_{Y(1),s}$$ if and only if $$1/2r^2_{Y(1),1,s}+1/2r^2_{Y(1),2,s}\leq r^2_{\overline{Y}(1),s}.$$ Fairly plausible? average of within-pair variances should be smaller than two times the between-pair variance (2 because $r^2_{\overline{Y}(1),s}$ DOF adjusted while $r^2_{Y(1),p,s}$ are not) ? If not plausible, we can just consider $4/3r^2_{Y,z,s}$ as our variance estimator (though may become too conservative).

\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Estimating and predicting LATEs' heterogeneity.}

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Estimating the covariance between the LATEs and a covariate.}

Let $X_s$ denote a site-specific variable, that is either observed or can be unbiasedly estimated. In this section, our target parameter is

equation*[equation* omitted — 136 chars of source]

a weighted covariance between the LATEs and $X_s$, where the weight assigned to site $s$ corresponds to the weight assigned to that site in $\text{LATE}$ (see (ref)). Let also $\bold{\beta}^{\text{FS}}_X$ denote the analogue of $\bold{\beta}^{\text{ITT}}_X$, but for a regression of $\text{FS}_s$ on $X_s$.

thmSuppose that Assumptions (ref)- (ref) hold. Then, \begin{align*} \sigma\left[LATE,X\right]=\frac{\sigma^2\left[X\right]}{FS}\left(\bold{\beta}^{ITT}_{X}-LATE\times \bold{\beta}^{\text{FS}}_{X}\right). \end{align*}

As a covariance is unnormalized, its magnitude is hard to interpret. Normalizing $\sigma\left[\text{LATE},\text{X}\right]$ would require identifying the variance of LATEs, which, as we will soon see, can be achieved at the expense of imposing an additional assumption. Yet, Theorem (ref) already shows that without imposing any additional assumption, the sign of the correlation between $X_s$ and the LATEs is identified, and is equal to the sign of $\bold{\beta}^{\text{ITT}}_{X}-\text{LATE}\times \bold{\beta}^{\text{FS}}_{X}$. A case of particular interest is when $X_s=\text{FS}_s$: knowing the sign of the correlation between LATEs and FSs may be useful to assess is there is Roy selection into treatment across sites, whereby sites where takeup is the largest are also the sites where compliers' gains from treatment are the largest roy1951some. In this special case, the sign of the correlation is just equal to the sign of $\bold{\beta}^{\text{ITT}}_{\text{FS}}-\text{LATE}$. Table (ref) shows that in our application, one cannot reject that LATEs and first-stages are uncorrelated, be it for the private or the public program.

table[table omitted — 1,059 chars of source]

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Estimating the variance of LATEs.}

\paragraph{Target parameter.} In this section, our target parameter is

equation*[equation* omitted — 133 chars of source]

a weighted variance of LATEs, where the weight assigned to site $s$ again corresponds to the weight assigned to that site in $\text{LATE}$.\footnote{With a slight abuse of notation, we keep the same $\sigma^2\left[.\right]$ notation as in the previous section, despite the difference in the weights.}

\paragraph{Studying LATEs' heterogeneity when FSs are homogeneous.} If $\sigma^2\left[\text{FS}\right]=0$, then $\text{LATE}_s=\text{ITT}_s/\text{FS}$, so $\sigma^2\left[\text{LATE}\right]=\sigma^2\left[\text{ITT}\right]/\text{FS}^2$, and one can just use $\widehat{\sigma}^2\left[\text{ITT}\right]/\widehat{\text{FS}}^2$ to estimate $\sigma^2\left[\text{LATE}\right].$ However, there are applications where FSs are heterogeneous across sites, and our empirical application is a good example. Table (ref) shows that in behaghel2014private, first-stage effects vary across sites, both for the public and for the private program. The estimated standard deviation of FS effects is around 11.4pp for the public program, namely 33% of the average FS effect of the public program, and around 11.8pp for the private program, namely 29% of is average FS effect. In the remainder of this section, we assume that $\sigma^2\left[\text{FS}\right]>0.$

table[table omitted — 1,474 chars of source]

\paragraph{Identification of $\sigma^2\left[\text{LATE}\right]$.} Let $\overline{FS^2}=\sum_{s=1}^Sw_s \text{FS}_s^2$ denote the average of the squared first-stages. Let $(\lambda_0,\lambda_1)$ denote the coefficients on $(1,\text{LATE}_s)$, in a regression of $\text{FS}_s$ on $(1,\text{LATE}_s)$, weighted by $\frac{w_s \text{FS}_s}{\text{FS}}$: $$(\lambda_0,\lambda_1)=\text{argmin}_{l_0,l_1}\sum_{s=1}^S\frac{w_s \text{FS}_s}{\text{FS}}\left(\text{FS}_s-l_0-l_1\text{LATE}_s\right)^2.$$ It follows from standard least-square algebra that

align[align omitted — 260 chars of source]

Let $\text{U}_s=\text{FS}_s-(\lambda_0+\lambda_1\text{LATE}_s)$ denote the residual from the regression. We consider the following assumption.

hyp$\sum_{s=1}^{S}\frac{w_s \text{FS}_s}{\text{FS}}\text{U}_s \text{LATE}_s^2=0$, and $\lambda_1=0$ or $\sum_{s=1}^{S}\frac{w_s \text{FS}_s}{\text{FS}}[\text{LATE}_s-\text{LATE}]^3=0$.

A sufficient condition for $\sum_{s=1}^{S}\frac{w_s \text{FS}_s}{\text{FS}}\text{U}_s \text{LATE}_s^2=0$ to hold is that $\lambda_2$, the coefficient on $\text{LATE}_s^2$ in a regression of $\text{FS}_s$ on $(1,\text{LATE}_s,\text{LATE}_s^2)$ weighted by $\frac{w_s \text{FS}_s}{\text{FS}}$, is equal to zero, meaning that the relationship between $\text{FS}_s$ and $\text{LATE}_s$ is linear. Then, Assumption (ref) either requires that sites' FSs and LATEs be uncorrelated ($\lambda_1=0$), or that the weighted skewness of LATEs be equal to zero. Theorem (ref) implies that $\lambda_1=0$ is fully testable.

thmIf Assumption (ref) holds, then \begin{align*} \sigma^2\left[LATE\right]=\frac{\sum_{s=1}^{S}w_s(ITT_s-FS_s\times LATE)^2}{\overline{FS^2}}. \end{align*}

\paragraph{Estimation of $\sigma^2\left[\text{LATE}\right]$.} Let

align*[align* omitted — 56 chars of source]

As $\sum_{s=1}^S w_s \nu_s=0$, the numerator of $\sigma^2\left[\text{LATE}\right]$ in Theorem (ref) is equal to the variance of $\nu_s$ across sites. Then, we will show that an EB variance estimator with outcome variable

align*[align* omitted — 74 chars of source]

converges to the same limit as $\sum_{s=1}^{S}w_s(\text{ITT}_s-\text{FS}_s\times \text{LATE})^2$. Turning to the denominator, as $$E\left(\widehat{\text{FS}}_s^2-\widehat{V}_{rob}\left(\widehat{\text{FS}}_s\right)\right)=E\left(\widehat{\text{FS}}_s^2\right)-V\left(\widehat{\text{FS}}_s\right)=\text{FS}_s^2,$$ we will show that $$\sum_{s=1}^{S}w_s\left(\widehat{\text{FS}}_s^2-\widehat{V}_{rob}\left(\widehat{\text{FS}}_s\right)\right)$$ converges to the same limit as that of $\sum_{s=1}^{S}w_s\text{FS}_s^2$. Finally, taking the ratio of these two estimators will yield a consistent estimator of $\sigma^2\left[\text{LATE}\right]$. More formally, let

align*[align* omitted — 440 chars of source]

Let $$\phi_{s,5}=\frac{\tilde{w}_s\tilde{\nu}_s}{\text{FS}},$$ and let $$\phi_{s,6}=\frac{\tilde{w}_s\left((\tilde{\nu}_s)^2-\widehat{V}_{rob}\left(\tilde{\nu}_s\right)\right)-2(C_1+C_2)\phi_{s,5}-\tilde{w}_s\left(\widehat{\text{FS}}_s^2-\widehat{V}_{rob}\left(\widehat{\text{FS}}_s\right)\right)C_3}{C_4},$$ where $C_1$, $C_2$, $C_3$, and $C_4$ respectively denote the limits of $\frac{1}{S} \sum_{s=1}^{S}\tilde{w}_sE(\widehat{\text{FS}}_s\tilde{\nu}_s)$,\\ $\frac{1}{S} \sum_{s=1}^{S}\tilde{w}_sE\left(\frac{\text{LATE}\times r^2_{D,1,s}-c_{D,Y,1,s}}{n_{1s}}+\frac{\text{LATE}\times r^2_{D,0,s}-c_{D,Y,0,s}}{n_{0s}}\right)$, $\sigma^2[\text{LATE}]$, and $\frac{1}{S} \sum_{s=1}^{S}\tilde{w}_s \text{FS}_s^2$, which are assumed to exist in Assumption (ref) below. Let $V_{\sigma^2[\text{LATE}]}$ denote the limit of $\frac{1}{S}\sum_{s=1}^S V\left(\phi_{s,6}\right)$, which is also assumed to exist below. Finally, let

align*[align* omitted — 310 chars of source]
hyp\begin{enumerate} • The sequence $(\phi_{s,6})_{s\geq 1}$ satisfies the Lyapunov condition. • The limits of the following sequences exist: i) $\frac{1}{S}\sum_{s=1}^S \tilde{w}_s FS^2_s$; ii) $\frac{1}{S}\sum_{s=1}^S \tilde{w}_s E(\widehat{\text{FS}}_s\tilde{\nu}_s)$; iii) $\sigma^2[\text{LATE}]$; iv) $\frac{1}{S} \sum_{s=1}^{S}\tilde{w}_sE\left(\frac{\text{LATE}\times r^2_{D,1,s}-c_{D,Y,1,s}}{n_{1s}}+\frac{\text{LATE}\times r^2_{D,0,s}-c_{D,Y,0,s}}{n_{0s}}\right)$; v) $\frac{1}{S}\sum_{s=1}^S V\left(\phi_{s,6}\right)$. • $\underset{S\rightarrow +\infty}{\lim}\frac{1}{S}\sum_{s=1}^S\tilde{w}_sFS^2_s>0$. \end{enumerate}
thmSuppose that Assumptions (ref)-(ref) hold. Then, \begin{align*} \sqrt{S}(\widehat{\sigma}^2[LATE]-\sigma^2[LATE])\overset{d}{\longrightarrow} N(0,V_{\sigma^2[LATE]}). \end{align*}

We conjecture that using similar steps as in the proof of Theorem (ref), one can show that the sample variance of $\widehat{\phi}_{s,6}$, a variable where all the population quantities in $\phi_{s,6}$ are replaced by their sample equivalents, converges to a limit weakly larger than $V_{\sigma^2[\text{LATE}]}$, and can thus be used as a conservative variance estimator.

\paragraph{Application: estimating the variance of the LATEs of the publicly- and privately-provided counseling programs.} In Table (ref), we estimate the variance of LATEs across sites, under Assumption (ref). Our variance estimators are statistically significant for both programs. Our estimate of LATEs' standard deviation across sites is equal to 364% of the LATE estimate for the public program, and to 432% of the LATE estimate for the private one.

table[table omitted — 1,263 chars of source]
comment\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Predicting site-specific LATEs} \paragraph{Target parameter.} Let $x_s$ be a scalar and binary site-level covariate. Letting $S_{x,1}=\sum_{s=1}^Sx_sw_s$ and $S_{x,0}=\sum_{s=1}^S(1-x_s)w_s$, our target parameter is $$\bold{\beta}^{\text{LATE}}_{x} \equiv \frac{1}{S_{x,1}}\sum_{s=1}^S x_s w_s\text{LATE}_s-\frac{1}{S_{x,0}}\sum_{s=1}^S (1-x_s)w_s\text{LATE}_s,$$ the difference between the average LATEs of sites with $x_s=1$ and $x_s=0$. Let $\text{LATE}_{x,1}$ and $\text{LATE}_{x,0}$ respectively denote the average LATE across sites with $x_s=1$ and $x_s=0$. \paragraph{Identification.} Assume that for any functions $f$ and $g$, \begin{align} \frac{1}{S_{x,0}}\sum_{s=1}^{S}(1-x_s)w_sf(LATE_s)\times g(FS_s)=&\left(\frac{1}{S_{x,0}}\sum_{s=1}^{S}(1-x_s)w_sf(LATE_s)\right)\times \left(\frac{1}{S_{x,0}}\sum_{s=1}^{S}(1-x_s)w_sg(FS_s)\right)\nonumber\\ \frac{1}{S_{x,1}}\sum_{s=1}^{S}x_sw_sf(LATE_s)\times g(FS_s)=&\left(\frac{1}{S_{x,1}}\sum_{s=1}^{S}x_sw_sf(\text{LATE}_s)\right)\times \left(\frac{1}{S_{x,1}}\sum_{s=1}^{S}x_sw_sg(\text{FS}_s)\right), \end{align} meaning that sites' FSs and LATEs are independent within the subsample of sites with $x_s=1$ and within the subsample of sites with $x_s=0$, a conditional version of Assumption (ref). \begin{thm} If (ref) holds, then \begin{align*} \bold{\beta}^{\text{LATE}}_{x}=\frac{\frac{1}{S_{x,1}}\sum_{s=1}^S x_s \left(\text{ITT}_s-\text{FS}_s\text{LATE}_{x,0}\right)-\frac{1}{S_{x,0}}\sum_{s=1}^S (1-x_s)\left(\text{ITT}_s-\text{FS}_s\text{LATE}_{x,1}\right)}{\text{FS}}. \end{align*} \end{thm} To consistently estimate $\bold{\beta}^{\text{LATE}}_{x}$, one can just replace $\text{ITT}_s$, $\text{FS}_s$, $\text{LATE}_{x,0}$, $\text{LATE}_{x,1}$ and $\text{FS}$ by their estimators in the previous display.

\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Conclusion}

In multi-site randomized controlled trials, with a large number of sites but few randomization units per site, an Empirical-Bayes (EB) estimator can be used to estimate the variance of the treatment effect across sites. We propose a consistent estimator of the coefficient from a ridge regression of site-level effects on site-level characteristics that are unobserved but can be unbiasedly estimated, such as sites' average outcome without treatment, or site-specific treatment effects on mediator variables. For instance, in a multi-site job-search counseling RCT, it can be interesting to study whether sites that have the largest effects on job-seekers' job finding rate are also the sites that have the largest effect on their search effort, as a “predictive mediation analysis” of whether the job-finding effect can be “explained” by the job-search effect. In experiments with imperfect compliance, we also propose a non-parametric and partly testable assumption under which the variance of local average treatment effects (LATEs) across sites can be estimated. We revisit behaghel2014private, who study the effect of counseling programs on job seekers job-finding rate, in more than 200 job placement agencies in France. We find considerable treatment-effect heterogeneity, both for intention to treat and LATE effects, and the treatment effect is negatively correlated with sites' job-finding rate without treatment.

center[center omitted — 38 chars of source]

\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Proofs}

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Proof of Theorem (ref)}

Asymptotic normality.

Let $$\tilde{\sigma}^2\left[\text{ITT}\right]=\sum_{s=1}^S w_s \left[\left(\widehat{\text{ITT}}_s-\text{ITT}\right)^2-\widehat{V}_{rob}\left(\widehat{\text{ITT}}_s\right)\right].$$

align[align omitted — 715 chars of source]

The fourth equality follows from the fact $\widehat{\text{ITT}}$ is unbiased for $\text{ITT}$, from applying the law of large numbers in Lemma 1 of liu1988bootstrap to the sequence of independent and bounded random variables $\tilde{w}_{s}\widehat{\text{ITT}}_s$, and from Point 3 of Assumption (ref). The fifth equality follows from applying the Lyapunov CLT to $\left(\tilde{w}_{s}\widehat{\text{ITT}}_s\right)_{s\geq 1}$. Then, as

align*[align* omitted — 493 chars of source]
align[align omitted — 194 chars of source]

The result follows from (ref) and (ref), from applying the Lyapunov CLT to $\left(\phi_{s,1}\right)_{s\geq 1}$, and from the Slutsky lemma.

Asymptotically conservative variance estimator.

Let $$\widehat{V}^I_{bound}=\frac{1}{S}\sum_{s=1}^{S}\left[\phi_{s,1}- \overline{\phi}_1 \right]^2.$$

align[align omitted — 382 chars of source]

Let $(x,y,z)\mapsto g(x,y,z)=\tilde{w}_{s}\left[\left(x-y\right)^2-z\right].$ \\ $\phi_{s,1}=g\left(\widehat{\text{ITT}}_s,\text{ITT},\widehat{V}_{rob}\left(\widehat{\text{ITT}}_s\right)\right)$, and $\widehat{\phi}_{s,1}=g\left(\widehat{\text{ITT}}_s,\widehat{\text{ITT}},\widehat{V}_{rob}\left(\widehat{\text{ITT}}_s\right)\right)$. Under Points 1 and 2 of Assumption (ref), $(\widehat{\text{ITT}}_s,\text{ITT},\widehat{V}_{rob}\left(\widehat{\text{ITT}}_s\right))$ belongs to a compact subset $\Theta$ of $\mathbb{R}^3$, and as $g$ is continuously differentiable, there exists a real number $C$ such that $\left|\frac{\partial g}{\partial y}(x,y,z)\right|\leq C$ for all $(x,y,z)\in \Theta$.

align*[align* omitted — 563 chars of source]

The first inequality follows from the triangle inequality, the equality follows from the mean value theorem. Then, as $\widehat{\text{ITT}}-\text{ITT}=o_P(1)$, the previous display implies that

equation[equation omitted — 121 chars of source]

One can use similar steps to show that

equation[equation omitted — 125 chars of source]

Finally, it follows from (ref)-(ref), the fact that under Assumptions (ref) and (ref) $\frac{1}{S}\sum_{s=1}^{S}\phi_{s,1}\overset{\mathbb{P}}{\longrightarrow}\underset{S\rightarrow+\infty}{\lim}\frac{1}{S}\sum_{s=1}^{S}E\left(\phi_{s,1}\right)$, and the continuous mapping theorem, that

align[align omitted — 117 chars of source]

Finally, under Assumptions (ref) and (ref), $$\widehat{V}^I_{bound}\overset{\mathbb{P}}{\longrightarrow}\overline{v}\equiv\underset{S\rightarrow+\infty}{\lim}\frac{1}{S}\sum_{s=1}^{S}E\left(\phi_{s,1}^2\right)-\left(\underset{S\rightarrow+\infty}{\lim}\frac{1}{S}\sum_{s=1}^{S}E\left(\phi_{s,1}\right)\right)^2\geq V_{\sigma^2\left[\text{ITT}\right]},$$ where the inequality follows by convexity of $x\mapsto x^2$. The result follows from (ref) and the previous display.

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Proof of Lemma (ref)}

Proof of Point 1

The first and last equalities are well-known results. The proof of the second one is similar to the proof of the second and third equalities in Point 3 below.

Proof of Point 2

$E\left(\overline{Y}_{0s}\right)=E(Y_s^r(0))$ is a well-known result. Conditional on $(Y^r_{is}(0))_{i\in \{1,...,n_s\}}$, the only source of randomness in $\overline{Y}_{0s}$ is the random sampling, without replacement, of $n_{0,s}$ units out of $n_s$ assigned to the control group. Then, as is well-known, $$V\left(\overline{Y}_{0s}|(Y^r_{is}(0))_{i\in \{1,...,n_s\}}\right)=r^2_{Y_s^r(0),s}\left(\frac{1}{n_{0,s}}-\frac{1}{n_{s}}\right).$$ Then, from the law of total variance and the fact that $E\left(r^2_{Y_s^r(0),s}\right)=V(Y^r_s(0))$, it follows that

equation[equation omitted — 106 chars of source]

Then,

align[align omitted — 389 chars of source]

The first equality follows from the fact that for any random variables $A$ and $B$, $V(A-B)=V(A)+V(B)-2\text{Cov}(A,B)$. The second equality follows from (ref), an equivalent equality for $V\left(\overline{Y}_{1s}\right)$, and the fact that under Assumptions (ref) and (ref), $V\left(\widehat{\text{ITT}}_s\right)=\frac{V(Y^r_s(0))}{n_{0,s}}+\frac{V(Y^r_s(1))}{n_{1,s}}$ imbens2015. (ref) directly implies that

align[align omitted — 140 chars of source]

Finally, the result follows from (ref), (ref), and the fact that under Assumptions (ref) and (ref), $r^2_{Y,0,s}$ is unbiased for $V(Y^r_s(0))$.

Proof of Point 3

$E\left(\widehat{\text{\bf{ITT}}}_{\text{M},s}\right)=\text{\bf{ITT}}_{\text{M},s}$ is a well-known result. We only prove that $E\left(\frac{c_{M_k,Y,0,s}}{n_{0,s}}+\frac{c_{M_k,Y,1,s}}{n_{1,s}}\right)=\text{Cov}\left(\widehat{\text{ITT}}_{\text{M}_k,s},\widehat{\text{ITT}}_s\right)$, the proof that $E\left(\frac{c_{M_k,M_{k'},0,s}}{n_{0,s}}+\frac{c_{M_k,M_{k'},1,s}}{n_{1,s}}\right)=\text{Cov}\left(\widehat{\text{ITT}}_{\text{M}_k,s},\widehat{\text{ITT}}_{\text{M}_{k'},s}\right)$ is similar. Let $\mathcal{T}_s=(Y^r_{is}(0),Y^r_{is}(1),M^r_{k,is}(0),M^r_{k,is}(1))_{i\in \{1,...,n_s\}}$. Under Assumptions (ref) and (ref), we can apply Theorem 3 in li2017 conditional on $\mathcal{T}_s$, to show that

align[align omitted — 281 chars of source]

Then,

align[align omitted — 1,103 chars of source]

The first equality follows from the law of total covariance. The second equality follows from (ref), and the fact that $\widehat{\text{ITT}}_{\text{M}_k,s}$ and $\widehat{\text{ITT}}_{s}$ are conditionally unbiased for the sample ITT effects on the outcome and the mediator. The third equality follows from the fact that the vectors $(Y^r_{is}(0),Y^r_{is}(1),M^r_{k,is}(0),M^r_{k,is}(1))$ are iid across $i$. The result follows from the previous display, and the fact that under Assumptions (ref) and (ref), $c_{M_k,Y,0,s}$ and $c_{M_k,Y,1,s}$ are respectively unbiased for $\text{Cov}(M^r_{k,s}(0),Y^r_s(0))$, and $\text{Cov}(M^r_{k,s}(1),Y^r_s(1))$.

Proof of Point 4

The proof follows from similar arguments as the proofs of Points 1 to 3, and from the fact that Theorem 3 in li2017 implies that standard variance formulas in two-arm RCTs still apply to multi-arm RCTs.

hyp\begin{enumerate} • There exists real numbers $M_0$ and $M_1$ such that $|\widehat{\bold{X}}_s|\leq M_0$ and $\tilde{w}_s\leq M_1$, and the sequence $(\phi_{s,4})_{s\geq 1}$ satisfies the Lyapunov condition. • The limits of the following sequences, when $S\rightarrow +\infty$, exist: \begin{enumerate} • $\sum_{s=1}^S w_s\bold{X}_s\bold{X}_s^T$$\mu(\bold{X})$$\sum_{s=1}^S w_s\bold{X}_s\text{ITT}_s$$1/S\sum_{s=1}^S V(\phi_{s,4})$. \end{enumerate} \end{enumerate}

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Proof of Theorem (ref)}

Proof of consistency.

We have

align[align omitted — 251 chars of source]

and

align[align omitted — 515 chars of source]

Moreover,

align[align omitted — 230 chars of source]

The first equality follows from the fact $\widehat{V}\left(\widehat{\bold{X}}_{s}\right)$ is unbiased for $V\left(\widehat{\bold{X}}_{s}\right)=E\left(\widehat{\bold{X}}_s\widehat{\bold{X}}_s^T\right)-E\left(\widehat{\bold{X}}_s\right)E\left(\widehat{\bold{X}}_s^T\right)$. The second equality follows from the fact $\widehat{\bold{X}}_s$ is unbiased.

Similarly,

align[align omitted — 262 chars of source]

The first equality follows from the fact $\widehat{\text{Cov}}\left(\widehat{\bold{X}}_{s},\widehat{\text{ITT}}_s\right)$ is unbiased for $\text{Cov}\left(\widehat{\bold{X}}_{s},\widehat{\text{ITT}}_s\right)=E\left(\widehat{\bold{X}}_s\widehat{\text{ITT}}_s\right)-E\left(\widehat{\bold{X}}_s\right)E\left(\widehat{\text{ITT}}_s\right)$. The second equality follows from the fact $\widehat{\bold{X}}_s$ and $\widehat{\text{ITT}}_s$ are unbiased.

Finally, the result follows from (ref)-(ref), the fact that $\widehat{\bold{X}}_s$ and the normalized weights $\tilde{w}_s$ are bounded, the fact that random variables are independent across sites, the law of large numbers for independent variables in Lemma 1 of liu1988bootstrap, Point 2 of Assumption (ref), and repeated uses of the continuous mapping theorem.

Proof of asymptotic normality.

Let

align*[align* omitted — 441 chars of source]

As $E\left(\sum_{s=1}^S w_s\left(\widehat{\bold{X}}_s-\mu(\bold{X})\right)\right)=0$, it follows from a Taylor expansion that

equation[equation omitted — 192 chars of source]

Similarly,

equation[equation omitted — 103 chars of source]

Using the same arguments as in the proof of Theorem (ref), one can show that $A(\lambda)=\frac{1}{S}\sum_{s=1}^S E(\phi_{s,2})$. Combined with (ref), this implies that

equation[equation omitted — 179 chars of source]

Similarly, one can show that

equation[equation omitted — 162 chars of source]

Finally, using the fact that

equation[equation omitted — 322 chars of source]

it follows from (ref) and (ref) that $$\sqrt{S}\left(\widehat{\bold{\beta}}^{\text{ITT}}_{X}(\lambda)-\bold{\beta}^{\text{ITT}}_{X}(\lambda)\right)=\frac{1}{\sqrt{S}}\sum_{s=1}^S \left(\phi_{s,4}-E(\phi_{s,4})\right)+o_P(1).$$ The result follows from applying the Lyapunov CLT to $\left(\phi_{s,4}\right)_{s\geq 1}$, and from the Slutsky lemma.

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Proof of Theorem (ref)}

By (ref), $$\sum_{s=1}^S\frac{w_s\text{FS}_s}{\text{FS}}(\text{LATE}_s-\text{LATE})=0.$$ Therefore,

align*[align* omitted — 629 chars of source]

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Proof of Theorem (ref)}

By construction, $\sum_{s=1}^{S}\frac{w_s \text{FS}_s}{\text{FS}}\text{U}_s=\sum_{s=1}^{S}\frac{w_s \text{FS}_s}{\text{FS}}\text{U}_s\text{LATE}_s=0$. Therefore, under Assumption (ref),

equation[equation omitted — 143 chars of source]

Then,

align[align omitted — 418 chars of source]

where the last equality follows from (ref). Now, if $\lambda_1=0$, it directly follows from (ref) that

align*[align* omitted — 170 chars of source]

thus proving the result. If $\lambda_1\ne 0$ but the skewness of the LATEs is equal to zero,

align*[align* omitted — 721 chars of source]

thus proving the result.

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Proof of Theorem (ref)}

It follows from, e.g., (A28) in de2018fuzzy and the fact that $\frac{1}{S}\sum_{s=1}^S E(\phi_{s,5})=0$ that

equation[equation omitted — 154 chars of source]

As the variables $\phi_{s,5}$ are independent and bounded, it then follows from the law of large numbers in Lemma 1 of liu1988bootstrap that

equation[equation omitted — 96 chars of source]

Then, letting $\tilde{\nu}_s(x)=\widehat{\text{ITT}}_s-x\times \widehat{\text{FS}}_s$,

align*[align* omitted — 842 chars of source]

where the second and third equalities follow from the mean-value theorem, for some $\tilde{\text{LATE}}_s$ included between $\text{LATE}$ and $\widehat{\text{LATE}}$, and for some $\bar{\text{LATE}}_s$ included between $\text{LATE}$ and $\tilde{\text{LATE}}_s$. As $\frac{\partial \left(\tilde{\nu}_s^2\right)}{\partial x}(x)= -2\widehat{\text{FS}}_s\left(\widehat{\text{ITT}}_s-\widehat{\text{FS}}_sx\right)$ and $\frac{\partial^2 \left(\tilde{\nu}_s^2\right)}{\partial x^2}(x)= 2\widehat{\text{FS}}_s^2$,

align*[align* omitted — 428 chars of source]

where the last equality follows from (ref), from applying the law of large numbers in Lemma 1 of liu1988bootstrap to the sequence of independent and bounded random variables $\tilde{w}_s\widehat{\text{FS}}_s^2$, and from Point 2i) of Assumption (ref). Therefore,

align[align omitted — 580 chars of source]

The second equality follows from applying the law of large numbers in Lemma 1 of liu1988bootstrap to the sequence of independent and bounded random variables $\tilde{w}_s\widehat{\text{FS}}_s\tilde{\nu}_s$ and from Point 2ii) of Assumption (ref). The third equality follows from (ref).

Similarly, let

align*[align* omitted — 443 chars of source]

One has

align*[align* omitted — 275 chars of source]

Then, using arguments similar to those used to show (ref),

align[align omitted — 917 chars of source]

Then, it follows from (ref) and (ref) that

align[align omitted — 356 chars of source]

Let $$\tilde{\sigma}^2[\text{LATE}]=\frac{\frac{1}{S} \sum_{s=1}^{S}\left(\tilde{w}_s\left(\tilde{\nu}_s)^2-\widehat{V}_{rob}\left(\tilde{\nu}_s\right)\right)-2(C_1+C_2)\phi_{s,5}\right)}{\frac{1}{S} \sum_{s=1}^{S}\tilde{w}_s\left(\widehat{\text{FS}}_s^2-\widehat{V}_{rob}\left(\widehat{\text{FS}}_s\right)\right)}.$$ It follows from, e.g., (A28) in de2018fuzzy, and from the fact that $\frac{1}{S}\sum_{s=1}^S E(\phi_{s,5})=0$, that

align[align omitted — 198 chars of source]

Then, it follows from (ref), (ref) and Point 3 of Assumption (ref) that

align[align omitted — 200 chars of source]

The result follows from applying the Lyapunov CLT to $\left(\phi_{s,6}\right)_{s\geq 1}$, and from the Slutsky lemma.

comment\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Proof of Theorem (ref)} \begin{align*} &\frac{1}{S_{x,1}}\sum_{s=1}^S x_sw_s \left(ITT_s-FS_sLATE_{x,0}\right)-\frac{1}{S_{x,0}}\sum_{s=1}^S (1-x_s)w_s\left(ITT_s-FS_sLATE_{x,1}\right)\\ =&\frac{1}{S_{x,1}}\sum_{s=1}^S x_sw_s \text{FS}_s\left(\text{LATE}_s-\text{LATE}_{x,0}\right)-\frac{1}{S_{x,0}}\sum_{s=1}^S (1-x_s)w_s\text{FS}_s\left(\text{LATE}_s-\text{LATE}_{x,1}\right)\\ =&\left(\frac{1}{S_{x,1}}\sum_{s=1}^S x_sw_s \text{FS}_s\right)\left(\frac{1}{S_{x,1}}\sum_{s=1}^S x_sw_s\left(\text{LATE}_s-\text{LATE}_{x,0}\right)\right)\\ -&\left(\frac{1}{S_{x,0}}\sum_{s=1}^S (1-x_s)w_s\text{FS}_s\right)\left(\frac{1}{S_{x,0}}\sum_{s=1}^S (1-x_s)w_s\left(\text{LATE}_s-\text{LATE}_{x,1}\right)\right)\\ =&\bold{\beta}^{\text{LATE}}_{x}\text{FS}. \end{align*}
landscape\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Survey of Multi-Site RCTs} \begin{table}[H] \caption{Multi-site RCTs in AEJ: Applied Economics 2014-2016 } \resizebox{\paperwidth}{!}{ \begin{tabular}{ c c c c } Title & Units of Observation & Units of Randomization & Sites \\ \hline Keeping It Simple: Financial Literacy and Rules of Thumb & Individual Clients & 1,193 Individual Clients & 107 Barrio\\ Improving Educational Quality through Enhancing Community Participation: Results from a Randomized Field Experiment in Indonesia& Students & 520 Schools & 44 Subdistricts \\ The Demand for Medical Male Circumcision & Individuals & 1,634 Individuals & 29 Enumeration Areas \\ Should Aid Reward Performance? Evidence from a Field Experiment on Health and Education in Indonesia & Individuals & 300 Kecamatan & 20 Kabupaten \\ Private and Public Provision of Counseling to Job Seekers: Evidence from a Large Controlled Experiment & Individuals & 43,977 Individuals & 216 Employment Offices\\ Estimating the Impact of Microcredit on Those Who Take It Up: Evidence from a Randomized Experiment in Morocco & Households & Villages (81 pairs) & 47 Branches \\ Microcredit Impacts: Evidence from a Randomized Microcredit Program Placement Experiment by Compartamos Banco & Households & 250 Geographic Clusters & Superclusters of 4 Adjacent Clusters\\ The Impacts of Microcredit: Evidence from Bosnia and Herzegovina & Individuals & 1,196 Individuals & 282 City/Towns or 14 Branches\\ Social Networks and the Decision to Insure & Households & 5,300 Households & 185 Villages \\ Inputs in the Production of Early Childhood Human Capital: Evidence from Head Start & Individuals & 4,442 Individuals & 353 Head Start Centers\\ The Returns to Microenterprise Support among the Ultrapoor: A Field Experiment in Postwar Uganda\footnote{Phase 2 Experiment} & Individuals & 904 Individuals & 60 Villages\\ The Impact of High School Financial Education: Evidence from a Large-Scale Evaluation in Brazil & Student & 892 Schools (in matched pairs) & Municipalities\\ \hline \end{tabular} } \begin{minipage}{21.0cm} \scriptsize{"The Returns to Microenterprise Support among the Ultrapoor: A Field Experiment in Postwar Uganda" corresponds to the Phase 2 experiment.\\ "Social Networks and the Decision to Insure" corresponds to the household level randomization and analysis. } \end{minipage} \end{table}