The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
27,782 characters
Optimal Experimental Design and Estimation when Potential Outcomes are Bounded
\maketitle
\begin{abstract}
\noindent I study the optimal design and analysis of randomized experiments for estimating finite-population average treatment effects when potential outcomes are known to be bounded, as with binary outcomes. Among all assignment mechanisms and a broad class of affine estimators, worst-case mean-squared error (MSE) is minimized by independent random assignment and an unconventional regression of the support-midpoint-centered outcome on the recentered treatment, with no intercept. This contrasts with the usual prescription of balanced complete randomization and difference-in-means estimation: when outcomes are bounded, randomness in the realized treatment share is informative. The worst-case gain over full-sample complete randomization is asymptotically small, but gains can be first-order relative to other designs: complete within-pair randomization and pair-fixed-effect regression have twice the worst-case MSE. I extend the result to allow for arbitrary estimators. Independent random assignment remains optimal, and the generally-nonlinear optimal estimator can meaningfully reduce worst-case MSE.
\end{abstract}
\newpage
\onehalfspacing
\section{Introduction}
Complete randomization in experiments---where a fixed number of units are treated---is often viewed as an unambiguous efficiency improvement over independent random assignment. By fixing the treated share, it eliminates chance imbalances while preserving uniform randomness over unit labels. Recent minimax results formalize this intuition, selecting complete randomization and difference-in-means estimation of average treatment effects (ATEs) with unrestricted potential outcomes \citep{bai2023why,kallus2021optimality}.
This paper shows that the formal optimality of complete randomization depends on the researcher not having an \emph{a priori} bound on potential outcomes---as one has with binary or otherwise limited-support outcomes. I show that, given outcome bounds, complete randomization may no longer be minimax. Instead, worst-case mean-squared error (MSE) is minimized by independent random assignment---with a random treated share. Intuitively, when outcomes are unbounded complete randomization is preferred because variation in the treated share can be exploited by an adversary: making the level of potential outcomes arbitrarily large or small can blow up the estimator with even a small imbalance. By anchoring the level of potential outcomes, ex ante bounds prevent this and make a random treated share useful for ATE estimation. In particular, independent randomization avoids any pairwise correlation between treatment assignments which an adversary can exploit through the heterogeneity (rather than the level) of potential outcomes.
I first show this reversal over a large class of affine estimators satisfying a natural condition, which ensures equivariance to midpoint-preserving shifts in treatment effects.\footnote{More precisely, the equivariance restriction says that increasing every treatment effect by some $t$ while holding each potential-outcome midpoint fixed must increase the estimate by exactly $t$.} This class includes simple or weighted difference-in-means, ordinary or weighted least squares with predetermined controls or fixed effects, and many other estimators. I show independent randomization is minimax-optimal with an unconventional estimator: a regression of the support-midpoint-centered outcome on the recentered treatment with no intercept. Treatment recentering ensures design-based identification \citep{borusyakHull2023nonrandom}, while midpoint-centering the outcome makes the random treated share variation informative. I further show complete randomization is strictly suboptimal, regardless of the affine estimator.
The worst-case MSE reduction of the optimal design and affine estimator is asymptotically small relative to full-sample complete randomization and the usual difference-in-means estimator, but it can be large relative to other popular strategies. Indeed, I show the procedure of paired randomization and pair-fixed-effect regression has twice the worst-case MSE by eliminating variation in treated shares within each pair. While the minimax analysis may be overly pessimistic for this case, as it doesn't incorporate \emph{a priori} information on the possible similarity of potential outcomes within strata, it formalizes the sense in which repeated balance restrictions can be costly when the strata are not prognostically informative.
Finally, I consider the general minimax problem which allows for any design and measurable estimator. I show that independent randomization remains minimax-optimal, while the optimal estimator is generally nonlinear.\footnote{I show that complete randomization is suboptimal in small samples and conjecture this holds generally. } This estimator is characterized by a finite-dimensional convex program with quadratic constraints, and has a least-favorable-prior representation analogous to \cite{hodges1982minimax}. It generally improves on the unbiased affine estimator by adding bias and reducing variance. Numerically, I find that using this optimal estimator decreases worst-case MSE by 15-30\% for moderate sample sizes, although this benefit declines asymptotically.
This paper builds on several related literatures. From statistics, a closely related literature studies optimal finite-population sampling and estimation of means. Early results study different restrictions on the population or the estimator class. \cite{godambe1955unified} shows that, for broad sampling designs and unrestricted population values, there is generally no unbiased linear estimator that uniformly minimizes variance. \cite{godambeJoshi1965admissibility} establish the admissibility of the Horvitz–Thompson estimator within the class of design-unbiased estimators. \cite{bickelLehmann1981minimax} impose a bound on finite-population dispersion and show that the sample mean is minimax under simple random sampling. Most directly related is \cite{hodges1982minimax}, who show that if all finite-population values lie in a known interval then the minimax estimator under simple random sampling shrinks the sample mean toward the interval midpoint; moreover, simple random sampling paired with this estimator is minimax among all sampling and estimation strategies.\footnote{Other work develops minimax sampling theory for alternative parameter spaces and estimator classes, including \cite{joshi1979best}, \cite{chengLi1983minimax}, \cite{gabler1988conditional,gabler1990minimax}, and \cite{stenger1989asymptotic}.} This paper adopts a similar logic for ATE estimation.\footnote{\citet[][Lemma 4.1 and Proposition 4.2]{harshawEtAl2024balancing} provide a close antecedent. Fixing the Horvitz–Thompson estimator and restricting to designs with marginal treatment probability one-half, they show that worst-case MSE over an $\ell_2$ bound is minimized by independent randomization. This paper's Proposition \ref{prop1} strengthens this result by allowing arbitrary assignment mechanisms and optimizing jointly over a large set of affine estimators. Theorem \ref{theorem1} further generalizes by optimizing over all measurable estimators. }
The closest contemporary paper is \cite{aronowLopatto2026minimax}, who consider finite-population totals when each unit’s outcome has a known and potentially unit-specific support interval. Holding marginal inclusion probabilities fixed, they derive a sharp lower bound on the maximum mean squared error of any design-unbiased estimator and show that it is attained precisely when the inclusion indicators are pairwise independent, in which case the optimal estimator is a midpoint-differenced Horvitz–Thompson estimator. They also jointly optimize the sampling design and estimator subject to an expected-sample-size constraint. My problem has the same bounded-support and midpoint-adjustment logic, but differs in two substantive ways. First, I study a causal setting where treatment assignment yields a nonstandard “one-from-each-pair” sampling from treated and untreated potential outcomes. Second, and more importantly, I impose no design-unbiasedness restriction and ultimately optimize over all measurable estimators.\footnote{A related robust-design literature asks how randomization protects against misspecification or unknown dependence rather than bounded support. \cite{wu1981robustness} connects randomization to worst-case mean squared error protection in comparative experiments, while \cite{bickelHerzberg1979robustness} study designs robust to serially correlated errors. These papers provide an additional precedent for interpreting independent randomization as protection against an adversarial outcome configuration, although the settings are quite different.}
This paper more directly contributes to the literature on optimal assignments and estimators for average treatment effects. Classic Neyman allocation fixes assigned shares to minimize the variance of differences in means given arm-specific outcome variances. Modern minimax results yield decision-theoretic foundations for related prescriptions. \cite{kallus2021optimality} shows complete randomization is minimax when the relevant conditional-mean class is permutation symmetric. \cite{bai2023why} studies joint optimization of assignment and linear estimation, obtaining difference-in-means at the Neyman allocation under his parameter class. My results are best viewed as complementing these analyses which leave outcome location unrestricted.\footnote{My minimax analysis under bounds is similar in spirit to \cite{deChaisemartin2024trading}, who assumes bounded strata-specific ATEs and derives the minimax linear combination of stratum-specific unbiased estimators. } I show that a known outcome support can be exploited with a random treated share; conditional on the share, however, my optimal design is exactly the complete-randomization designs considered in earlier work.\footnote{Other work uses baseline information to improve on complete randomization: \cite{bai2022optimality} proves optimality of appropriately constructed matched pairs among stratified designs that treat every unit with probability one half; \cite{hahnHiranoKarlan2011adaptive} and \cite{tabordMeehan2023stratification} use earlier experimental waves to choose treatment propensities and strata, and \cite{harshawEtAl2024balancing} construct a design that explicitly trades covariate balance against worst-case robustness. \cite{kallus2018optimal} shows that complete randomization is minimax without structure linking baselines to potential outcomes.}
Lastly, this paper builds on a recent literature on estimation with recentered estimators. \cite{borusyakHull2023nonrandom} establish recentering for design-based identification of formula treatments, combining as-if-random assignments with other predetermined variables, while \cite{borusyakHull2026optimal} characterize efficient formula instruments for a given design. \cite{borusyakHullMunro2026robust} jointly optimize the design and recentered instrument under a minimax approximate-variance criterion. A motivating example in \cite{borusyakHullMunro2026robust} shows that independent---rather than complete---randomization is minimax-optimal when the treatment formula is the assignment itself, using an asymptotic variance approximation and a restricted class of well-behaved recentered estimators. This paper shows a similar result holds for finite-sample MSE over an unrestricted class of estimators.
The rest of this paper is organized as follows. The next section shows the initial result for restricted affine estimators. Section \ref{sec:general} then considers the general problem. Section \ref{sec:conclusion} concludes. All proofs are given in the appendix.
\section{Restricted Affine Estimators}\label{sec:affine}
Consider a fixed population of $N$ units. For each unit $i$, let $Y_i(0)$ and $Y_i(1)$ be fixed potential outcomes and let $D_i\in\{0,1\}$ denote treatment assignment. A researcher knows, prior to assignment, that potential outcomes are bounded: for known $L<U$,
\begin{align*}
Y_i(0),Y_i(1) \in [L,U],\hspace{0.3cm} i=1,\dots,N.
\end{align*}
The parameter of interest is the finite-population ATE:
\begin{align*}
\beta = \frac{1}{N}\sum_{i=1}^N (Y_i(1)-Y_i(0)).
\end{align*}
To estimate $\beta$, the researcher first chooses a \emph{design} $\delta\in \Delta (\{0,1\}^N)$, where $\Delta(\cdot)$ denotes the simplex, along with an estimator. Treatment assignments are drawn from $\delta$ and determine outcomes $Y_i=Y_i(0)(1-D_i)+Y_i(1)D_i$. The researcher then applies the estimator to $(D_i,Y_i)_{i=1}^N$, yielding an estimate $\hat\beta$. Here we restrict to affine estimators, of the form:
\begin{align*}
\hat\beta = a(D)+b(D)^\prime Y,
\end{align*}
where $D$ and $Y$ are $N\times 1$ vectors of the assignments and outcomes, $a(\cdot)$ is a fixed function, and $b(\cdot)$ is an $N\times 1$ vector of fixed functions. The goal of the researcher is to pick the design and estimator to minimize worst-case MSE over the unknown potential outcomes. Formally, for $Y(\cdot)=(Y_i(0),Y_i(1))_{i=1}^N$, they solve:
\begin{align}
\inf _{(\delta,\hat\beta)}\sup_{Y(\cdot)\in [L,U]^{2N}}\mathbb E_\delta \left[(\hat\beta - \beta)^2\right].\label{eq:objective}
\end{align}
We first analyze this problem for the large set of affine estimators that satisfy a natural equivariance condition:
\begin{align}
b(D)^\prime \left(D-\frac{1}{2}\mathbf{1}\right)=1,\hspace{0.3cm} \delta\text{-almost surely}\label{eq:equivariance}
\end{align}
\noindent To see why this condition is natural, note that we can rewrite the outcome equation as:
\begin{align*}
Y_i = \frac{Y_i(1)+Y_i(0)}{2}+(Y_i(1)-Y_i(0))\left(D_i-\frac{1}{2}\right).
\end{align*}
Thus, under equation \eqref{eq:equivariance}, shifting all treatment effects by some constant $t$ while keeping the location (i.e. midpoint) of potential outcomes fixed increases the estimate by exactly $t$:
\begin{align*}
\hat\beta^{new} &= a(D) + \sum_ib(D)_i \left(\frac{Y_i(1)+Y_i(0)}{2}+(Y_i(1)-Y_i(0)+t)\left(D_i-\frac{1}{2}\right)\right) \\
&= a(D)+b(D)^\prime Y + b(D)^\prime\left(D-\frac{1}{2}\mathbf{1}\right)\times t = \hat\beta + t.
\end{align*}
In this sense, equation \eqref{eq:equivariance} ensures the estimator is equivariant to midpoint-preserving treatment effect shifts. One can verify that it is satisfied for many common estimators---including simple or weighted difference-in-means, ordinary or weighted least squares with an intercept and possibly other predetermined controls or fixed effects, and matching estimators that match on predetermined strata---with the analogous equivariance holding for many nonlinear procedures such as differences of medians or trimmed means.\footnote{The general condition can be stated, for estimators $\hat\beta(D,Y)$, as $\hat\beta(D,Y+t(D-\frac{1}{2}\mathbf{1}))=\hat\beta(D,Y)+t$.}
The first result shows independent randomization is optimal for this class of affine estimators, with an unconventional estimator:
\begin{proposition}\label{prop1}
Let $M=\frac{U+L}{2}$. Restricting to affine estimators satisfying equation \eqref{eq:equivariance},
\begin{align*}
\inf _{(\delta,\hat\beta)}\sup_{Y(\cdot)\in [L,U]^{2N}}\mathbb E_\delta \left[(\hat\beta - \beta)^2\right] = \frac{(U-L)^2}{N}\equiv V^*.
\end{align*}
A minimax procedure is independent random assignment, $D_i\stackrel{iid}{\sim}Bernoulli(0.5)$, with the unbiased estimator
\begin{align}
\hat \beta^* = \frac{\sum_i (D_i-\frac{1}{2})(Y_i-M)}{\sum_i (D_i-\frac{1}{2})^2} = \frac{2}{N}\sum_{i=1}^N (2D_i-1)(Y_i-M).\label{eq:betastar}
\end{align}
\end{proposition}
\noindent The appendix proof is straightforward. It shows that, for any design and any estimator satisfying \eqref{eq:equivariance}, the adversary can generate an MSE of at least $(U-L)^2/N$ by placing both potential outcomes at one of the two support endpoints: i.e., selecting a $B_i\in\{L,U\}$ for each unit $i$ and setting $Y_i(0)=Y_i(1)=B_i$. If treatment assignments are correlated across units, the adversary can choose the endpoint pattern to align with that dependence and increase MSE above the lower bound. Independent random assignment eliminates this possibility and attains the $V^*$ bound when paired with the optimal estimator $\hat\beta^*$.
The $\hat\beta^*$ estimator is nonstandard. Equation \eqref{eq:betastar} shows it arises from an ordinary least squares regression of the support-midpoint-centered outcome $Y_i-M$ on the recentered treatment $D_i-1/2$, with no intercept. Equivalently, it is the Horvitz–Thompson ATE estimator applied to the centered outcome.\footnote{The second expression in \eqref{eq:betastar} can also be written $\hat\beta^*=\frac{1}{N}\sum_i \left(\frac{D_i(Y_i-M)}{\pi}-\frac{(1-D_i)(Y_i-M)}{1-\pi}\right)$ for $\pi=1/2$.} While omitting an intercept from a regression or using unnormalized Horvitz-Thompson weights is often undesirable, here it is exactly what allows the estimator to attain the minimax bound under independent randomization, by avoiding the slight negative correlation in observation weights induced by demeaning. Recentering the treatment by its known propensity score of $1/2$ ensures unbiasedness, while centering the outcome at the support midpoint anchors outcomes, leaving only midpoint deviations that cannot be exploited by the adversary under independent randomization.
A closely related result shows that complete randomization is strictly suboptimal:
\begin{proposition}\label{prop2}
Let $N$ be even and restrict to affine estimators satisfying equation \eqref{eq:equivariance}. Then if $\delta$ is chosen to be balanced complete randomization,
\begin{align*}
\inf _{\hat\beta}\sup_{Y(\cdot)\in [L,U]^{2N}}\mathbb E_\delta \left[(\hat\beta - \beta)^2\right] \ge \frac{(U-L)^2}{N-1} =\frac{N}{N-1} V^*.
\end{align*}
Moreover, the difference-in-means estimator attains this bound.
\end{proposition}
\noindent The optimality of difference-in-means for complete randomization is unsurprising; the new insight is that this conventional estimator-design pair has strictly higher worst-case MSE than the $V^*$ bound. Intuitively, complete randomization generates a slight negative correlation between treatment assignment pairs which the adversary can exploit by making exactly half of the $B_i$ equal to $L$ and half equal to $U$. Averaging over such endpoint patterns makes MSE proportional to the squared magnitude of the difference-in-means outcome weights $b(D)$, with the usual $N/(N-1)$ finite-population factor from sampling without replacement.
The $N/(N-1)$ factor inflating worst-case MSE with complete randomization is, of course, small for even moderately large populations; repeated exact-balance restrictions, however, can produce a first-order loss. To illustrate, consider paired randomization with $J$ predetermined pairs and the conventional pair-fixed-effect regression estimator. Here there are $J$ exact-balance restrictions with exactly one unit selected for treatment in each pair, and assignments are perfectly negatively correlated within pairs. The adversary can exploit this restriction by choosing a no-effect schedule in which one unit’s potential outcomes equal $L$ and the other’s equal $U$ in every pair. Each pair then contributes an independent estimation error of magnitude $U-L$, yielding worst-case MSE that is twice that of $V^*$.
\begin{proposition}\label{pair_prop}
Suppose $N = 2J$ units are partitioned into $J$ pairs, indexed by $(j, 1)$ and $(j, 2)$. Restrict to affine estimators satisfying \eqref{eq:equivariance} and set $\delta$ to paired randomization with exactly one unit in each pair assigned with probability $1/2$, independently across pairs. Then:
\begin{align*}
\inf_{\hat\beta}\sup_{Y(\cdot)\in [L,U]^{2N}}\mathbb E_\delta \left[(\hat\beta - \beta)^2\right] \ge \frac{(U-L)^2}{J} =2 V^*.
\end{align*}
Moreover, the pair-fixed-effect regression estimator attains this bound.
\end{proposition}
The minimax analysis is likely too pessimistic for the expected MSE of such a paired randomization strategy, as it does not incorporate any \emph{a priori} information on the similarity of potential outcomes within pairs. In practice, stratification is typically based on such information and can significantly improve MSE (see, e.g., Bai 2022; Tabord-Meehan 2023). In such settings, Proposition \ref{pair_prop} can be viewed as a formalization of the idea that stratification can come at a high MSE cost when the strata are not overly informative about potential outcomes, such that highly heterogeneous potential outcomes could coexist within strata.
\section{Optimal Estimation}\label{sec:general}
We now consider the general minimax problem \eqref{eq:objective} without restricting the class of estimators. The first main result shows that independent randomization remains optimal, while the optimal estimator is generally nonlinear:
\begin{theorem}\label{theorem1}
For nonnegative integers $k,m$ with $k+m\leq N$, define $\theta_{k,m}=\frac{2k+m}{N}$ and
\begin{align*}
p_{k,m}(x)=2^{-m}\binom{m}{x-k},
\quad x=0,\ldots,N,
\end{align*}
where $\binom{m}{j}=0$ when $j\notin\{0,\ldots,m\}$. Let $\mathcal K_N$ be the set of all such $(k,m)$ and define
\begin{align}\label{eq:problem}
\kappa_N
=
\min_{d=(d_0,\ldots,d_N)\in[0,2]^{N+1}}
\max_{(k,m)\in\mathcal K_N}
\sum_{x=0}^N
p_{k,m}(x)(d_x-\theta_{k,m})^2.
\end{align}
Then, over all measurable estimators,
\begin{align*}
\inf _{(\delta,\hat\beta)}\sup_{Y(\cdot)\in [L,U]^{2N}}\mathbb E_\delta \left[(\hat\beta - \beta)^2\right] = (U-L)^2 \kappa_N.
\end{align*}
A minimax procedure is independent random assignment, $D_i\stackrel{iid}{\sim}Bernoulli(0.5)$, with the estimator
\begin{align*}
\hat\beta^*_{NL} = \hat\beta^* + (U-L)\sum_{x=0}^N\left(d_x^*-\frac{2x}{N}\right)p_x(Z)
\end{align*}
where $\hat\beta^*$ is the optimal affine estimator in Proposition \ref{prop1}, $d^*=(d_0^*,\dots,d_N^*)$ is any solution to \eqref{eq:problem}, $p_x(z)=Pr(\sum_i B_i=x)$ for $B_i\stackrel{ind}{\sim}Bernoulli(z_i)$, and $Z=(Z_i)_{i=1}^N$ for
\begin{align*}
Z_i=\frac{1}{2}+\frac{(2D_i - 1)(Y_i-\frac{U+L}{2})}{U-L}\in[0,1].
\end{align*}
\end{theorem}
\noindent The appendix proof follows in the same spirit of \cite{hodges1982minimax}; as in their bounded finite-population sampling problem, the minimax estimator $\hat\beta^*_{NL}$ is shown to be a Bayes rule under a least-favorable prior supported on the potential outcome endpoints. This estimator takes the form of the previous restricted-affine $\hat\beta^*$ plus a term that is nonlinear in $Y$. This term generally makes $\hat\beta^*_{NL}$ biased while reducing worst-case MSE.
To understand this new estimator, it is useful to specialize it to the previous affine class while relaxing the equivariance restriction \eqref{eq:equivariance}. The next result shows that a ``shrunk'' version of $\hat\beta^*$ is then optimal (along with independent randomization):
\begin{corollary}\label{corollary1}
Over all affine estimators, without imposing \eqref{eq:equivariance}, minimax MSE is $(U-L)^2/(\sqrt{N}+1)^2$. A minimax procedure is independent random assignment with the estimator
\begin{align*}
\hat\beta^*_{NL} = \frac{\sqrt{N}}{\sqrt{N}+1}\hat\beta^*
\end{align*}
Moreover, $\kappa_N<(\sqrt{N}+1)^{-2}$ when $N>2$.
\end{corollary}
\noindent The first part of this corollary shows that shrinking $\hat\beta^*$ towards zero---generating some bias and violating the equivariance restriction---reduces worst-case MSE by a factor of $N/(\sqrt{N}+1)^2$. The second part shows that further expanding to nonlinear estimators is useful for any nontrivial number of units.
Computing $\hat\beta^*_{NL}$ is relatively straightforward, as $\eqref{eq:problem}$ is a finite-dimensional convex program with quadratic constraints that depends only on the population size $N$. A direct solver is practical for populations in low thousands; for larger $N$, constraint generation gives certified lower and upper bounds. Numerically, I find a relatively small share of constraints bind.\footnote{E.g., the number of binding or nearly binding constraints is in the low tens when $N$ is in the hundreds.}
Figure 1 shows the potential benefit in using $\hat\beta^*_{NL}$ instead of $\hat\beta^*$. Specifically, it plots $\kappa_N N$, the worst-case MSE from using $\hat\beta_{NL}$ as a fraction of the worst-case MSE from using $\hat\beta^*$ (both with independent randomization) on a grid up to $N=2,500$. The relative worst-case MSE reduction is substantial for moderate sample sizes and declines smoothly with $N$; for example, it is around 30\% when $N=100$ and falls to around 15\% for $N=1,000$. I conjecture it converges to zero as $N\rightarrow\infty$, as with the gain in using $\hat\beta^*$ and independent randomization relative to balanced complete randomization and difference-in-means estimation.\footnote{One may wonder whether complete randomization remains strictly suboptimal with general estimators. I have verified this numerically in small samples and conjecture it is true, but as of now have not proved it.} The figure also plots the corresponding relative worst-case MSE of the best affine estimator under independent randomization, $N/(\sqrt{N}+1)^2$, which has a similar shape.
\begin{figure}[t]
\centering
\caption{Nonlinear minimax risk}
\includegraphics[width=0.75\textwidth]{kappaN.pdf}
\label{fig:minimax_risk}
\begin{minipage}{0.9\textwidth}
\footnotesize
\textit{Notes:} The solid blue line plots $N\kappa_N$ against the actual population size $N$, where $\kappa_N$ is the minimax MSE from Theorem \ref{theorem1} divided by $(U-L)^2$. Thus, it reports the worst-case MSE of the unrestricted minimax estimator as a fraction of the worst-case MSE of the equivariant affine
estimator $\widehat\beta^*$ from Proposition \ref{prop1} (both under independent
randomization). Values solve program \ref{eq:problem} on the grid $N=j^2$, $j=1,\ldots,50$. The dashed black line plots the corresponding relative worst-case MSE of the best affine estimator under independent randomization, $N/(\sqrt{N}+1)^2$.
\end{minipage}
\end{figure}
\section{Conclusion}\label{sec:conclusion}
This paper shows a curious advantage of independent randomization, relative to complete or otherwise stratified randomization, when estimating ATEs for bounded potential outcomes. Exact balance of treated shares can be suboptimal once the level of outcomes is anchored, because the estimator variance can no longer be made arbitrarily large by a random treated share multiplying some arbitrarily high or low outcomes. Independent randomization is generally preferred by avoiding negative correlations in treatment assignments which can otherwise blow up MSE. While such correlations tend to be small for complete randomization, they can be first-order if balance is imposed repeatedly, such as in paired randomization.
In practice, complete, paired, or more elaborate randomization schemes are likely to dominate simple independent randomization when based on informative observables. Similarly, more elaborate regressions that incorporate informative controls or fixed effects are likely to dominate the unconventional support-midpoint-centered estimator in the restricted affine class and likely also the minimax-optimal estimator $\hat\beta^*_{NL}$. Extending this paper's minimax problem to additional assumptions on the homogeneity of potential outcomes within strata, or other prognostic structure, is a natural direction for future work.
\singlespacing
\bibliographystyle{aer}
\bibliography{causalHL}