EconBase
← Back to paper

Kernel regression analysis of tie-breaker designs

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

83,574 characters

Kernel regression analysis of tie-breaker designs



\begin{frontmatter}

 \title{Kernel regression analysis of tie-breaker designs \thanks{This version of the paper is published in the \textit{Electronic Journal of Statistics} with open access (DOI: doi.org/10.1214/23-EJS2102). There are formatting differences between the arXiv and journal versions.}}
 \runtitle{Kernel regression analysis of tie-breaker designs}

 \begin{aug}
\author{\fnms{Dan M.} \snm{Kluger}\ead[label=e1]{[email removed]}}
\and
\author{\fnms{Art B.} \snm{Owen}\ead[label=e2]{[email removed]}}

\address{Department of Statistics, Stanford University, Stanford CA, 94305 \\
\printead{e1,e2}}


 \end{aug}

\runauthor{Kluger and Owen}

\maketitle

\begin{abstract}
Tie-breaker experimental designs are hybrids of Randomized Controlled Trials (RCTs) and Regression Discontinuity Designs (RDDs) in which subjects with moderate scores are placed in an RCT while subjects with extreme scores are deterministically assigned to the treatment or control group. In settings where it is unfair or uneconomical to deny the treatment to the more deserving recipients, the tie-breaker design (TBD) trades off the practical advantages of the RDD with the statistical advantages of the RCT. The practical costs of the randomization in TBDs can be hard to quantify in generality, while the statistical benefits conferred by randomization in TBDs have only been studied under linear and quadratic models. In this paper, we discuss and quantify the statistical benefits of TBDs without using parametric modelling assumptions. If the goal is estimation of the average treatment effect or the treatment effect at more than one score value, the statistical benefits of using a TBD over an RDD are apparent. If the goal is nonparametric estimation of the mean treatment effect at merely one score value, we prove that about 2.8 times more subjects are needed for an RDD in order to achieve the same asymptotic mean squared error.  We further demonstrate using both theoretical results and simulations from the \cite{AngristLavy} classroom size dataset, that larger experimental radii choices for the TBD lead to greater statistical efficiency.





\end{abstract}

\begin{keyword}[class=MSC]
\kwd[Primary ]{62K99}
\kwd[; secondary ]{62G20}
\kwd{62G08}
\end{keyword}

\begin{keyword}
\kwd{causal inference}
\kwd{experimental design}
\kwd{hybrid experiments}
\kwd{local linear regression}
\kwd{regression discontinuity designs}
\end{keyword}
\tableofcontents

\end{frontmatter}


\section{Introduction} \label{sec:Intro}

In this paper we study a nonparametric regression approach
to tie-breaker studies.
In the settings of tie-breaker studies, there is a costly
treatment while the control is inexpensive or even free. In addition, an investigator can decide how to allocate the costly treatment using a priority ordering on the subjects. The priority ordering could be based on how deserving of the treatment each subject is, or based on how strongly each subject is expected to respond to the treatment.
Examples include offering
scholarships or school placement to students \citep{Angrist2020Nebraska}, offering
a drug rehabilitation program to people
of varying needs \citep{CappelleriXX},
or assigning interventions to reduce
risk factors for child abuse and neglect \citep{kran:2022}.
In these settings a randomized controlled trial
(RCT) is inappropriate because it is extremely
inefficient economically, or even ethically
questionable.
The natural, even automatic, approach to settings like
this is to rank the subjects $i=1,\dots,N$
according to their value of a running variable $x_i$
and assign the treatment to only those subjects with
the highest values of $x_i$. For simplicity one can
assume that the number of subjects to treat is
fixed and that the treatment is offered to subject
$i$ if and only if $x_i> t$ for some threshold $t$.

The problem with a deterministic treatment based on
$x_i$ is that it complicates causal inference of the
effect of the treatment.  One can use regression discontinuity
analysis \citep{this:camp:1960,cattaneo2022regression} but
the regression discontinuity design (RDD) is known to
give treatment effect estimates a very high variance.
See, for example, \cite{Gelman_dont_do_high_order_polynomial}.
In a parametric regression model the treatment and running
variable are highly correlated, making for an inefficient
design \citep{gold:1972,jacob2012practical}.
In a tie-breaker design (TBD), the subjects are given the
treatment with a probability that increases with $x_i$.
The top ranked subjects get the treatment with probability one,
the bottom ranked subjects do not get the treatment and subjects
in between are randomly assigned to either treatment or control.
The tie-breaker design interpolates between two extremes: the RCT
and the RDD, trading off statistical efficiency with the short term economic value
of aligning treatment to the running variable.

We believe that there are many more good uses for TBDs.  Many companies interact electronically
with their customers and partners.
Perks, such as service upgrades, can
easily be assigned with some randomization. Because the treatment
is costly, it is important to
evaluate the treatment efficacy later.
This provides a strong motivation to introduce
randomization.
When the perk is simply a gift to some subjects there
is less ethical concern over whether it goes to the
most loyal customers or introduces some randomization.
We also expect that tie-breakers will be useful
in evaluating governmental programs such as
the one in \cite{kran:2022}, as well as educational programs such as those in \cite{Angrist2020Nebraska}.

TBDs have been primarily studied as experimental
designs using parametric regression modeling assumptions.  While the design literature
focuses on parametric models, the RDD literature primarily uses nonparametric
regression methods.  In this paper, we quantify the statistical gains to
be obtained by conducting a TBD instead of an RDD using nonparametric regression.

Our main theoretical contributions are as follows.
We study a kernel
weighted local linear regression with a slope
and intercept for both treated and control
subjects and bandwidth $h$.
The RDD can consistently estimate
the treatment effect only at $x=t$, and so
we focus our comparison at that point. We find
an expression for the optimal bandwidth
for estimating the treatment effect at $x=t$ under the TBD.
We then compare the optimal mean squared
error at $x=t$ for the two designs.
For the popular triangular kernel, a TBD
reduces the
asymptotic mean squared error (AMSE) by
a factor of about $2.27$ compared to an RDD
of the same sample size, $N$.  For other popular
kernels, the AMSE is reduced by a slightly greater
factor.  In this setting, the AMSE decreases
proportionally to $N^{-4/5}$, and using $N$
points in an RDD is comparable to using only
$0.36N$ points in a TBD.
The asymptotic analysis has a bandwidth
that converges to zero.
Since this convergence
is at the very slow $N^{-1/5}$ rate, we cannot
assume that in practice $h$ will be small enough such that subjects without randomized treatment are discarded. Therefore, we also compare the designs when $h$ is fixed and large enough to include nonnrandomized subjects in the regression.  In this setting, the
efficiency ratio, which we define as the relative variance of treatment effect estimators under the two designs, can be as large as four. For a fixed bandwidth and for the triangular and boxcar kernels, we also find that the efficiency ratio is monotone non-decreasing in the proportion of subjects who are given a randomized treatment assignment.

A further advantage of the TBD is that
it can give consistent nonparametric estimates
of the treatment effect for any value of $x$
in the randomization window.
It can also be used to estimate the average
causal effect over that window. These additional advantages are described more explicitly in Section \ref{sec:CausalParamsAndProblem}.




An outline of this paper is as follows.
Section~\ref{sec:lit} reviews the literature on
tie-breakers as well as the much larger literature
on RDDs.  In Section~\ref{sec:CausalParamsAndProblem}, we define a causal parameter of interest that can be used to compare the TBD to the RDD. We also introduce the causal identification assumptions needed and the local linear regressions used for estimating that parameter. In Section~\ref{sec:bandwidth_shrink_asymptotics}, we compare the mean squared error (MSE) in asymptotically optimal estimation of our causal parameter of interest under an RDD to that under a TBD. In that asymptotic setting, the optimal bandwidth decreases at the slow $O(N^{-1/5})$ rate and then the local linear regression is eventually supported entirely in the experimental region of the TBD. In Section~\ref{sec:AsymApprox_via_integrals}, we investigate another regime where the bandwidth $h$ is fixed and is assumed to be larger than the radius $\Delta$ of the experimental region. For this setting, deferring an investigation of the bias to Appendix \ref{sec:MagicBandwidth}, we study the variance of our estimator as a function of $\Delta$ and find the efficiency ratio to be monotone in $\Delta$ for the triangular and boxcar kernels.   Section~\ref{sec:Data_application}
shows how one can compute efficiency ratios empirically using one's
actual assignment variable levels, focusing on the Israeli classroom size data from \cite{AngristLavy} as an example.  The curves of the empirical efficiency ratios are quite similar to the ones obtained theoretically. Section~\ref{sec:Conclusion} presents a discussion. Appendix \ref{sec:GenPneqHalf} extends our results to TBDs in which each subject in the experimental group is given the treatment with probability $p \neq 1/2$. Appendices \ref{sec:proof_of_TBD_MSE}, \ref{sec:relMSETheoremProof}, \ref{sec:TSefficiencyCalc}, and \ref{sec:MonotoneEffTS} contain some of our proofs.

\section{Literature review}\label{sec:lit}

Here we survey the small TBD literature
and some recent developments in the
much larger RDD literature.
We also note connections to the experimental
design literature.
Most of the TBD literature has focused
on global parametric models.  Those are also
the dominant model for experimental design.
The TBD is usually compared to the RDD, for which nonparametric models are the norm.


\subsection{Regression discontinuity methods}

Here we present some concepts from the regression
discontinuity literature drawing heavily on
\cite{cattaneo2022regression}.
We begin with a setting where there is a running
variable $x_i$ (also called a score or index)
for subject $i=1,\dots,N$.
Subjects with $x_i>t$ are given
the treatment and others get the control.
The treatment levels are typically $T_i\in\{0,1\}$
with $T_i=0$ being the control.  For the TBD
setting it is more convenient to use $Z_i\in\{-1,1\}$
with $Z_i=-1$ indicating the control.
The potential outcomes for subject $i$
are $Y_{i+}$ if treated and $Y_{i-}$ for control.

There are two main approaches to RDD in causal
inference, continuity-based and local
randomization-based.  The continuity-based approach
assumes that the mean response
for treated subjects is continuous in $x$
as is that for control subjects.  If the mean response
for all subjects shows a discontinuity at $x=t$, then
the magnitude of this discontinuity is defined to
be the causal effect of treatment on subjects at $x=t$,
and one then considers how to estimate that effect.
The version from \cite{hahn2001identification}
has IID tuples $(x_i,Y_{i+},Y_{i-})$ where
$Y_i$ equals $Y_{i+}$ for $x_i>t$ and $Y_i$
equals $Y_{i-}$ otherwise.
The treatment effect is then
$$
\tau = \lim_{x\downarrow t}\mu_+(x)
-\lim_{x\uparrow t}\mu_-(x)
$$
where $\mu_{\pm} = \mathbb{E}(Y_{\pm}\!\mid\! X=x)$.
We will work primarily
with a superpopulation setting
where the subjects in the study are sampled
from a joint distribution. \cite{cattaneo2022regression} discuss this
setting along with some other
settings that focus on causal inference
for the given subjects.

The local randomization approach from \cite{cattaneo2015randomization}
assumes the existence of a window $\mathcal{W}=[t-h,t+h]$
such that for $x\in \mathcal{W}$ the treatment variable is
`as good as randomized'.
In the local randomization approach
we assume {\bf a)} that the joint distribution of
the $Z_i$ for $x_i\in\mathcal{W}$ is known, and,
{\bf b)} that the potential outcomes
$(Y_{i+},Y_{i-})$ are independent of $x_i$.
In particular, both mean responses must
be constant functions of $x\in\mathcal{W}$.
A variant of local randomization
has the treatment based on
a threshold of $x$ where $x$ is a noisy
version of a latent variable $u$
\citep{eckl:igna:wage:2020} and where we have
outside information under which the
probability of treatment given $u$ is known.

Both frameworks have challenges.
The obvious difficulty with local randomization
is choosing the window $\mathcal{W}$ (or knowing
the treatment probability given $u$).
A smaller window provides a smaller dataset to use while making the
window larger will normally increase
the discrepancy between
the model and the ground truth.
The TBD can be viewed as a strategy to impose
by design the first assumption in the local
randomization approach while not making the
second assumption.

The challenge in the continuity framework is
in estimating the necessary limits.  In a parametric
model, those limits are estimated from all the
data but have a bias due to lack of fit of the
parametric model.  As a result, nonparametric
regression methods based on local polynomial
models are favored.
The challenge there is that one must choose
a bandwidth $h$, analogously to the window
size from the local randomization
framework.  Because the mean responses are only
locally polynomial we must contend with
a bias-variance tradeoff in estimating the limits.

Our theoretical and numerical results compare
the TBD to the RDD in the continuity framework.
We think that this is the more likely alternative
analysis for our motivating applications if
randomization had not been used, because
the local randomization assumptions
do not seem natural in those applications.
We focus on the
accuracy of point estimation.  There is also a large
literature on constructing confidence intervals
around the RDD estimate (see \cite{cattaneo2022regression}).
We describe some of those concepts in this paper,
but we do not develop confidence intervals
for the TBD due to space constraints.

There are many different settings where treatments
depend in a discontinuous way on $x$.
In a sharp design, the treatment is $Z_i=1$
if and only if $x_i>t$.
In a fuzzy design, the assignment to treatment
or control might not perfectly match
$x_i>t$ versus $x_i\leqslant t$, for reasons
beyond the control of the investigator.
For instance there may be subjects that
do not comply with their assigned treatment.
A related issue is that some subjects might be
able to manipulate their value of the running
variable in order to get (or avoid) the treatment.
\cite{rosenman2019optimized} discuss some ways
to counter that problem.
In the settings we consider, the investigator
has control of the treatment and so we study
the sharp design.  We also do not address issues of subject compliance, as we suppose that the effect of the treatment assignment or the intent to treat is of sufficient interest to the investigator.

The RDD setting has been extended well beyond
the simple framework described above.
There are versions with treatments at more than
two levels as well as versions with
continuous treatments.
The cutoff can be defined in terms of a vector
of covariates yielding a discontinuity set of
dimension one less than the vector has.
The treatment discontinuity could be defined
by geographical boundaries.
There are multi-cutoff
settings where subject $i$ gets the treatment
when $x_i>t_i$.  There are models where it is
the derivative of $\mathbb{E}(Y\!\mid\! x,Z=1)-\mathbb{E}(Y\!\mid\! x,Z=-1)$
that has a step discontinuity at $t$.
For discussion and references to the variants
above, see \cite{cattaneo2022regression}.

\subsection{Experimental design}

While we study TBDs in comparison to
regression discontinuity,
they can also be considered within an
experimental design framework such as the
covariate-dependent designs considered by
\cite{metelkina2017information}. That paper
emphasizes sequential problems,
and like most of the design literature it
works primarily with parametric
models.
They find conditions where the treatment
policies converge to optimal deterministic
functions of a covariate vector.
The TBD does not use
deterministic allocations which is an advantage
if the response distribution is subject to change
between experiments.

Experimental design, especially in a sequential
setting is closely related to bandit methods.
We are motivated by problems where the responses
$Y_i$ arrive too slowly for bandit methods to
be suitable.  In a business setting, the responses
may arrive after a year or calendar quarter while
the effect of a scholarship on graduation rates can
only be seen years later.

\subsection{Tie-breakers}

The simplest tie-breaker design replaces the threshold $t$ by
two thresholds $t\pm\Delta$.
Subjects with $x_i>t+\Delta$ get the treatment,
subjects with $x_i<t-\Delta$ get the control
and other subjects are randomized to
either treatment or control.
The simplest choice has
\begin{equation}\label{eq:threelevel}
\Pr(Z_i=1\!\mid\! x_i)
=\begin{cases}
0, & x_i <t-\Delta,\\
\frac12, &|x_i-t|\leqslant\Delta\\
1, & x_i >t+\Delta.
\end{cases}
\end{equation}
\cite{camp:1969} describes the $\Delta=0$ version of this design.
Some subjects are exactly at the
threshold $t$ and then randomization breaks the
ties among them.
\cite{boru:1975} considers positive
values of $\Delta$ such that differences
in the running variable among subjects
with $|x-t|\leqslant \Delta$ are essentially
arbitrary because $x$ is an imperfect measure.
\cite{abdu:etal:2022} study the New York
school system that
breaks ties among applicants by lottery, or
standardized test, or audition, depending on
the program.  We only consider randomized
tie-breaking.

\cite{gold:1972} considers a simple two line regression model that in our
notation is
\begin{align}\label{eq:ovmodel}
Y_i & = \beta_1 +\beta_2x_i+\beta_3Z_i+\beta_4x_iZ_i+\varepsilon_i
\end{align}
for IID errors $\varepsilon_i$ with mean zero
and variance $\sigma^2$.
He finds that an RDD estimates these
coefficients with a variance that is
asymptotically $\pi/(\pi-2)\approx2.75$ times as large
as it would be under an RCT.  His setting
has Gaussian $x_i$ and $\beta_4=0$. \cite{jacob2012practical} generalizes the above model to polynomials of
degree two or three in $x$ with or
without interactions between $x$
and $Z$. Their Table 6 shows that an RCT is 4 times as efficient as the RDD for
a uniformly distributed running variable
with $t$ at the midpoint of its range.
They also provide similar efficiency estimates
for other polynomial models for both uniform
and Gaussian $x$ and include settings
where $t$ is not at the median of the
distribution of $x$.

The above comparisons of RCTs to RDDs do not
include tie-breakers.  \cite{CDP94_TBD_power3Delta}
compare small, medium and large randomization
windows in which 20\%, 35\% and 50\%, respectively,
of the subjects get a randomized treatment.
They tabulate the sample sizes needed to attain
a certain level of statistical power for three treatment effect sizes
in these TBDs as well as in an RDD and in an RCT.  All
designs had half of the subjects getting the
treatment. The running variable $x$ was normally
distributed.  The model was~\eqref{eq:ovmodel}
with $\beta_4=0$, making the treatment effect
constant. The power calculations were done
by Monte Carlo sampling. The required sample
sizes became smaller with increased
randomization at any level of power and
effect size.

\cite{owen:vari:2020} work out the
asymptotic variance of $\hat\beta$ in the
model~\eqref{eq:ovmodel} as a function of
$\Delta$.  They consider both $\mathbb{U}[-1,1]$
and $\mathcal{N}(0,1)$ distributions for $x$ and
a threshold $t$ at the median of $x$'s distribution.
The estimated treatment effect is
$2(\hat\beta_3+\hat\beta_4x)$ and they find
for uniform $x$ that this estimate has
asymptotic variance proportional to
$16(1+3x^2)/[1+3\Delta^2(2-\Delta^2)]$
where $\Delta=0$ describes the RDD and
$\Delta=1$ is the RCT. This decreases
monotonically in $\Delta$ while increasing
monotonically in $|x|$.

They also consider the opportunity cost
of experimentation compared to the RDD.
For $x_i\sim\mathbb{U}[-1,1]$ the expected
value of $\sum_{i=1}^NY_i$ is approximately
$(\beta_1 + \beta_4(1-\Delta^2)/2) N$. If
larger $Y_i$ are better and $\beta_4>0$ then
the opportunity cost
grows proportionally to $\beta_4\Delta^2$.
They discuss how one might trade off this
opportunity cost against statistical efficiency.

The TBD has so far been analyzed for simpler
methods than the RDD has.  This can be understood
by comparing their workflows.  In a TBD we
measure $x_i$, then sample $Z_i$ and then
some time later observe $Y_i$. For an RDD we
usually get $(x_i,Z_i,Y_i)$ all at once.
The investigator planning a TBD only has $x_i$,
must decide how to assign the $Z_i$,
and may not know what model will be fit later,
and then chooses some specific model to design for.
When one studies the TBD theoretically, one does
not even have the $x_i$ and then it is natural to
assume a distribution for them.
The TBD is prospective while the
RDD is retrospective.

When vectors $\boldsymbol{x}_i$ of covariates
are available, \citet[Section 8]{owen:vari:2020}
describe how to investigate numerically
the efficiency of a TBD that fits a regression
model on some collection of features of $\boldsymbol{x}_i$
that interact with $Z_i$ where the treatment
window is based on a linear combination of $\boldsymbol{x}_i$.

\cite{TimArt22} study multiple regression for
a tie-breaker in a regression model
$Y_i = \boldsymbol{x}_i^\mathsf{T}\beta + Z_i\boldsymbol{x}_i^\mathsf{T}\gamma+\varepsilon_i$
with $\Pr(Z_i=1)=p_i\in[0,1]$.
They study a prospective $D$-optimality criterion
that maximizes the determinant of $\mathbb{E}(\mathcal{X}^\mathsf{T}\mathcal{X})$,
where $\mathcal{X}$ is the (random) design matrix built from
$\boldsymbol{x}_i$ and $Z_i$.
For any known $\boldsymbol{x}_i$, the finite sample optimal $p_i$
can be computed by convex optimization.
For as yet unobserved $\boldsymbol{x}_i$ the prospective $D$-optimality
criterion averages over both random $Z_i$ and random
$\boldsymbol{x}_i$ from an assumed distribution for $\boldsymbol{x}_i$.
For random $\boldsymbol{x}_i$, they study
a three level tie-breaker with running variable
$\boldsymbol{x}_i^\mathsf{T}\eta$ and treatment probabilities $0$, $0.5$
and $1$.

\cite{owen:vari:2020} consider replacing the
simple trichotomy~\eqref{eq:threelevel} by various
sliding scales where $\Pr(Z_i=1 \!\mid\! x_i )$ is a monotone
function of $x_i$. They find no advantage to such
alternatives when $x_i$ has a symmetric distribution
about $t$ and half the subjects are treated.
\cite{HarrisonArt22} revisit that problem for the
two line model and find optimal designs for
general $x_i$ distributions and general fractions
of treated subjects without assuming that half
of the subjects will be treated.
These optimal designs can greatly improve
upon the design defined in \eqref{eq:threelevel}.
They still have $\Pr(Z_i=1 \!\mid\! x_i)$ as piecewise
constant functions of $x_i$.
If we impose monotonicity $\Pr(Z_i=1 \!\mid\! x_i )\geqslant\Pr(Z_{i'}=1 \!\mid\! x_{i'})$
whenever $x_i\geqslant x_{i'}$, then only two treatment
probability levels are needed.



A limitation of previous comparisons between RDDs and TBDs is that they all assume parametric regression
models for the response $Y_i$.  We compare them using local linear
regression.
For simplicity, we restrict our attention to the three level version of the TBD in~\eqref{eq:threelevel}.

Our model assumes an additive error
on top of smooth functions of $x$ for the
data points where $|x-t|>\Delta$.
When $h>\Delta$, the causal estimate we consider merges deterministic
and randomized treatment allocations and then
cannot be analyzed in a potential outcomes framework.
It is common in causal inference to ignore
such data.  For instance, a rule of thumb
in \cite{crump2009dealing} is to omit data
where the treatment probability is outside
$[0.1,0.9]$.
Asymptotically, $h<\Delta$ and then a potential
outcomes analysis is available. Otherwise, to stay within the potential outcomes framework, one
must choose between ignoring some data and
using the additive error model like we do.


\section{Causal estimand and problem formulation} \label{sec:CausalParamsAndProblem}

Throughout the text we will compare the TBD to the RDD. In our comparison, we will define $t$ to be the putative RDD threshold and $\Delta$ to be the experimental radius, and we consider allocation of the treatments to the $N$ subjects according to the 3-level tie-breaker design~\eqref{eq:threelevel}.


Next we discuss the estimands of interest.
For each subject we consider the assignment variable $X\in\mathbb{R}$, the treatment $Z\in\{-1,1\}$ and two potential outcomes:
$Y_+=Y(Z=1)$ and $Y_-=Y(Z=-1)$. Defining \begin{equation}\label{eq:muPlusMinus_def}
    \mu_+(x) \equiv \mathbb{E} \big(Y_+ \!\mid\! X=x \big)
\quad\text{and}\quad
\mu_-(x) \equiv  \mathbb{E} \big(Y_- \!\mid\! X= x \big),
\end{equation} the treatment effect at $X=x$ is
$$\tau(x) =\mathbb{E}( Y_+ -Y_-\!\mid\! X=x) = \mu_+(x)-\mu_-(x).$$
If the investigator chooses an RDD with a  threshold at $t$, then under certain regularity conditions, the causal estimand $$\tau_{\mathrm{thresh}} \equiv
\tau(t)$$ can be consistently estimated. In particular, we assume IID samples (or sufficiently weak dependence between the samples) and that:
\begin{compactenum}[(i)]
    \item The density $f(\cdot)$ of the assignment variable $X$ is continuous at $t$ with $f(t)>0$.
    \item The conditional mean functions $\mu_+$ and $\mu_-$ in \eqref{eq:muPlusMinus_def} have at least 3 continuous derivatives in an open neighborhood of $t$.
    \item The conditional variance functions $\sigma_{\pm}^2(x) \equiv \text{Var}\big(Y_{\pm} \!\mid\! X=x \big)$ are both bounded in a neighborhood of $t$ and continuous at $t$.
\end{compactenum}
Under these conditions, $\tau_{\mathrm{thresh}}$ can be consistently estimated by local linear regression with $O_p(N^{-2/5})$ errors \citep{ImbensKalyanaraman_optimalBW}. On the other hand, the conditions above do not suffice to let an RDD consistently estimate $\tau(x)$ for any $x\ne t$.

If an investigator runs a TBD with $\Delta>0$, then assumptions like those above replacing $t$ by $x$ allow consistent estimation of $\tau(x)$ for any $x\in(t-\Delta,t+\Delta)$.
Furthermore, as long as $\mathrm{var}(Y_{\pm})<\infty$,
$$\tau_{\text{ATE}}(\Delta) \equiv
\mathbb{E}( \tau(X) \!\mid\! t-\Delta <X<t+\Delta)
$$
can be consistently estimated with error $O_p(N^{-1/2})$ in a TBD without requiring assumptions (i), (ii), and (iii).

The discussion above leaves open the possibility that an RDD could be better than a TBD
when estimating $\tau_{\mathrm{thresh}}$. Therefore, for the remainder of the paper our primary focus will be on showing that even if the only goal is estimating $\tau_{\mathrm{thresh}}$, it is still beneficial to run a TBD rather than an RDD. Our other focus will be to show that when the only goal is to estimate $\tau_{\mathrm{thresh}}$, it is beneficial to pick a larger $\Delta$ in the experimental design stage when the option is available. Picking a larger $\Delta$ has other benefits as well such as making $\tau(x)$ identifiable for more values of the assignment variable and making $\tau_{\text{ATE}}(\Delta)$ more representative of the overall population and easier to estimate.  Naturally, there are
non-statistical reasons to keep $\Delta$ smaller.


\subsection*{Local linear estimation}


In keeping with current RDD practice, we suppose that {under an RDD} $\tau_{\mathrm{thresh}}$ will be estimated with local linear regression. In particular, we assume that a parameter vector $\beta$ defined by
 \begin{align}\label{eq:kernelcriterion}
 \hat{\beta} = \operatorname*{arg\,min}\limits_{\beta \in \mathbb{R}^4 }
 \sum\limits_{i=1}^N K \Big( \frac{x_i - t}{h} \Big)\big( Y_i -  (\beta_1 + \beta_2x_i+\beta_3 Z_i +\beta_4 x_i Z_i)  \big)^2
\end{align} will be fit for some symmetric kernel function $K(\cdot)\geqslant0$ and bandwidth parameter $h>0$, and that $\tau_{\mathrm{thresh}}$ will be estimated with \begin{equation}\label{eq:def_hat_tauthresh}
\hat{\tau}_{\mathrm{thresh}}=2\hat{\beta}_3+2\hat{\beta}_4 t.\end{equation}
While this formulation of estimating $\hat{\tau}_{\mathrm{thresh}}$ may be less familiar than the approach of fitting separate local linear regressions for the treatment and control groups, it is easy to check that the two formulations yield the same estimator.

Throughout the paper, we suppose that under a TBD, $\hat{\tau}_{\mathrm{thresh}}$ will also be estimated using local linear regression according to \eqref{eq:kernelcriterion} and \eqref{eq:def_hat_tauthresh}.
We do not use the same bandwidth $h$ for the TBD and RDD.
Indeed, in the next section we see that the optimal bandwidth choice (in terms of AMSE) is different for the two designs.

Because kernels with unbounded support are not typically used in RDD analysis (\cite{cattaneo2022regression}),  we  only consider kernels with bounded support. We assume without loss of generality, that the kernel is supported on $[-1,1]$.
We have a special interest in a uniform (boxcar) kernel $K_{\mathrm{BC}}(x)=1_{|x|\leqslant 1}$ because it is a popular kernel choice and is a local version of the regression model~\eqref{eq:ovmodel}. We are also interested in a triangular spike kernel $K_{\mathrm{TS}}(x)=(1-|x|)_+$ where $z_+=\max(0,z)$. This kernel was shown by \cite{cheng1997automatic} to optimize a bias-variance tradeoff for extrapolation from $x_i>t$ to $\mathbb{E}(Y\!\mid\! x=t)$ and has been advocated for RDD analysis by \cite{ImbensKalyanaraman_optimalBW} and \cite{calonico2014robust} among others.



The local linear regression estimator from~\eqref{eq:def_hat_tauthresh} has a bias and variance that both depend on the bandwidth $h$. Larger $h$ typically bring greater bias because the true regression is not precisely linear over a region centered on $t$.  Smaller $h$ bring greater variance because then fewer data points are in the regression. \cite{ImbensKalyanaraman_optimalBW} develop a method for choosing the bandwidth $h$ that is asymptotically mean squared optimal for the RDD. In the next section, we compare the AMSE of the TBD with that of the RDD, when each of them has their asymptotically optimal bandwidth choice.


In this paper, we focus on the accuracy of the estimated treatment
effect.  The RDD literature includes several papers devoted
to the construction of confidence intervals.
There it is necessary to account for the bias in a local polynomial
regression.  A simple approach is to choose $h$ to undersmooth the
regression function, resulting
in a bias of lower order than the standard error, and this simplifies
confidence interval construction.  Undersmoothing, however, brings less
accuracy \citep{calonico2014robust}.  See \cite{calo:catt:farr:2019}
for a discussion of bandwidth choices to optimize estimation, or
optimize confidence interval construction, or to get robust
(asymptotically valid) confidence intervals using the bandwidth
that is optimal for estimation.

\section{Asymptotic mean square optimal error}\label{sec:bandwidth_shrink_asymptotics}

In this section, we demonstrate the advantage of the TBD over the RDD
when each design's bandwidth is chosen to
minimize the AMSE in the estimation of $\tau_{\mathrm{thresh}}=\mu_+(t)-\mu_-(t)$. Following \cite{ImbensKalyanaraman_optimalBW}, we assume the following more general regularity conditions for estimating the causal effect at $X=t$: \begin{compactenum}[(i)]
    \item The triples $\big(X_i,Y_{i+}, Y_{i-}\big)$ for $i=1,\dots,N$ are IID.
    \item The distribution of $X_i$ has density $f(\cdot)$, which is continuously differentiable at $t$ with $f(t)>0$.
    \item Conditional means $\mu_{\pm}(\cdot)$ both have at least three continuous derivatives in an open neighborhood of $t$, with the $k$'th derivatives at $t$ denoted $\mu_\pm^{(k)}(t)$.
    \item The kernel $K(\cdot)$ is nonnegative, symmetric, bounded, has support $[-1,1]$, is continuous on its support, and is strictly positive somewhere.
    \item The conditional variances $\sigma_\pm^2(x) \equiv \mathrm{var} (Y_{i \pm} \!\mid\! X_i=x )$
    are both bounded in an open neighborhood of $t$ and are continuous and strictly positive at $t$.
    \item $\mu_+^{(2)}(t) \neq \mu_-^{(2)}(t)$.
\end{compactenum}


Under an RDD,  $Z_i=1$ if $X_i> t$
and is $-1$ otherwise,
so these assumptions imply Assumptions 3.1--3.6 that \cite{ImbensKalyanaraman_optimalBW} make for an RDD. To allow for analysis in the TBD setting, our assumptions
(i)--(vi) are slightly
stronger than those in \cite{ImbensKalyanaraman_optimalBW}.
For example, unlike in our Assumption (iii), \cite{ImbensKalyanaraman_optimalBW} make no assumptions on $\mu_+(\cdot)$ in the interval $(-\infty,t)$ or on $\mu_-(\cdot)$ in the interval $(t, \infty)$.
 Regarding assumption (vi), \cite{ImbensKalyanaraman_optimalBW} also consider the case where $\mu_+^{(2)}(t) = \mu_-^{(2)}(t)$ and show that in this case, their proposed method of estimating $\tau_{\mathrm{thresh}}$ has error $O_p(N^{-3/7})$ rather than $O_p(N^{-2/5})$. We do not consider the case where $\mu_+^{(2)}(t) = \mu_-^{(2)}(t)$ in detail for the TBD as the result should be similar to that for the RDD and is of less interest for our head-to-head comparison of TBD with RDD.

Because our assumptions (i)--(vi) imply Assumptions 3.1--3.6 in \cite{ImbensKalyanaraman_optimalBW} for an RDD, if we let
$$\tilde{\nu}_j \equiv \int_0^{\infty} u^j K(u) \,\mathrm{d} u  \quad \text{ and } \quad \tilde{\pi}_j \equiv \int_0^{\infty} u^j K^2(u) \,\mathrm{d} u$$
for $j \in \mathbb{N}$, and let \begin{align}\label{eq:defctilde}
\tilde{C}_1 \equiv \frac{1}{4} \Big( \frac{\tilde{\nu}_2^2 - \tilde{\nu}_1 \tilde{\nu}_3}{\tilde{\nu}_0 \tilde{\nu}_2 - \tilde{\nu}_1^2} \Big)^2  \quad \text{ and } \quad \tilde{C}_2 \equiv \frac{\tilde{\nu}_2^2 \tilde{\pi}_0 -2 \tilde{\nu}_1 \tilde{\nu}_2 \tilde{\pi}_1+\tilde{\nu}_1^2 \tilde{\pi}_2}{( \tilde{\nu}_2 \tilde{\nu}_0 - \tilde{\nu}_1^2)^2},
\end{align}
and define
\begin{equation}\label{eq:AMSE_RDD_def}
    \text{AMSE}_{\text{RDD}}(h,N) \equiv \tilde{C}_1 \big( \mu_+^{(2)}(t) -\mu_-^{(2)}(t) \big)^2 h^4+ \frac{\tilde{C}_2}{Nh} \Big( \frac{\sigma_+^2(t)+\sigma_-^2(t)}{f(t)} \Big),
\end{equation}
then Lemma 3.1 of \cite{ImbensKalyanaraman_optimalBW} holds. We reproduce the statement of this lemma below.


 \begin{lemma}
 \label{lemma:MSE_RDD_IK} Under Assumptions (i)--(vi), if an \emph{RDD} determines the treatment assignment and both $h \to 0$ and $Nh \to \infty$ as the number of samples $N \to \infty$, then the mean squared error in estimating $\tau_{\emph{thresh}}$ is given by \begin{equation}\label{eq:asymp_MSE_RDD}
     \emph{MSE}_{\emph{RDD}}(h,N)= \emph{AMSE}_{\emph{RDD}}(h,N) + o_p\Big( h^4 + \frac{1}{Nh} \Big),
 \end{equation}
 and the asymptotically optimal bandwidth, defined by $\operatorname*{arg\,min}_h \emph{AMSE}_{\emph{RDD}}(h,N)$ is given by \begin{equation}\label{eq:hopt_RDD}  h_{\emph{opt,RDD}}(N)=  \Big( \frac{\tilde{C}_2}{4 \tilde{C}_1} \Big)^{1/5}  \Big( \frac{\sigma_+^2(t) +\sigma_-^2(t)}{f(t) \big( \mu_+^{(2)}(t)- \mu_-^{(2)}(t)\big)^2} \Big)^{1/5}  N^{-1/5}.
 \end{equation}
 \end{lemma}
\begin{proof}
\citet[Lemma 3.1]{ImbensKalyanaraman_optimalBW}.
\end{proof}

Because we wish to compare the RDD to the tie-breaker design, we derive a similar result for the asymptotic MSE for the tie-breaker design. The TBD counterparts to the RDD quantities above are
\begin{equation}\label{eq:def_nuAndPi_newer}
    \nu_j \equiv \int_{-\infty}^{\infty} u^j K(u) \,\mathrm{d} u  \quad \text{ and } \quad \pi_j \equiv \int_{-\infty}^{\infty} u^j K^2(u) \,\mathrm{d} u
\end{equation}
for $j \in \mathbb{N}$,
\begin{equation}\label{eq:C_def}C_1 \equiv \frac{1}{4} \Big( \frac{\nu_2^2 - \nu_1 \nu_3}{\nu_0 \nu_2 - \nu_1^2} \Big)^2  \quad \text{ and } \quad C_2 \equiv  \frac{2\nu_2^2 \pi_0 -4 \nu_1 \nu_2 \pi_1+2\nu_1^2 \pi_2}{( \nu_2 \nu_0 - \nu_1^2)^2},
\end{equation}
and
\begin{equation}\label{eq:AMSE_TBD_def}
    \text{AMSE}_{\text{TBD}}(h,N) \equiv C_1 \big(  \mu_+^{(2)}(t) -\mu_-^{(2)}(t) \big)^2 h^4+ \frac{C_2}{Nh} \Big( \frac{  \sigma_+^2(t)+\sigma_-^2(t)}{f(t)} \Big).
\end{equation}




 \begin{lemma}\label{lemma:MSE_TBD} Under Assumptions (i)--(vi), if a \emph{TBD} with a fixed experimental radius $\Delta >0$ determines the treatment assignment and both $h \to 0$ and $Nh \to \infty$ as the number of samples $N \to \infty$,  then the mean squared error in estimating $\tau_{\emph{thresh}}$ is given by \begin{equation}\label{eq:asymp_MSE_TBD}
     \emph{MSE}_{\emph{TBD}}(h,N)= \emph{AMSE}_{\emph{TBD}}(h,N) + o_p\Big( h^4 + \frac{1}{Nh} \Big),
 \end{equation} and the asymptotically optimal bandwidth, defined by $\operatorname*{arg\,min}_h \emph{AMSE}_{\emph{TBD}}(h,N)$ is \begin{equation}\label{eq:hopt_TBD}  h_{\emph{opt,TBD}}(N)=  \Big( \frac{C_2}{4 C_1} \Big)^{1/5}  \Big( \frac{  \sigma_+^2(t) +\sigma_-^2(t)}{f(t) \big( \mu_+^{(2)}(t)- \mu_-^{(2)}(t)\big)^2} \Big)^{1/5}  N^{-1/5}.  \end{equation}
 \end{lemma}

 \begin{proof}
 See Appendix \ref{sec:proof_of_TBD_MSE}.
 \end{proof}

 The proof of this lemma is very similar to the proof of Lemma 3.1 in \cite{ImbensKalyanaraman_optimalBW}, from their appendix. Instead of pointing to their proof and noting the parts of their proof that differ in the tie-breaker design setting, we write out the proof of Lemma~\ref{lemma:MSE_TBD} in Appendix~\ref{sec:proof_of_TBD_MSE} to ensure there are no subtle issues with using their proof in the tie-breaker design setting.

{The leading order MSE formulas are derived by evaluating and summing the leading order terms for both the squared-bias and the variance.} In formulas \eqref{eq:AMSE_RDD_def} and \eqref{eq:AMSE_TBD_def} for the leading order MSE, the first term gives the leading order squared-bias while the second term gives the leading order variance. See formulas \eqref{eq:as_bias_TBD} and \eqref{eq:TBD_as_var_formula} for explicit calculations of the leading order bias and variance in the TBD case, and see the formulas for `B' and `V' in the appendix of \cite{ImbensKalyanaraman_optimalBW} for explicit calculations of these quantities in the RDD case. It is not surprising that the formulas for the leading order squared-bias, variance and MSE are different for the two design types because for an RDD, estimation of $\tau_{\mathrm{thresh}}$ involves estimation of the mean functions at a boundary point, whereas for a TBD, estimation of $\tau_{\mathrm{thresh}}$ involves estimation of mean functions at an interior point.

 In Figure \ref{fig:bias_var_tradeoff}, we plug in scalar multiples of the optimal bandwidth for the RDD given in Lemma \ref{lemma:MSE_RDD_IK} to the first and second terms of formulas \eqref{eq:AMSE_RDD_def} and \eqref{eq:AMSE_TBD_def} to visualize the trade-off for the leading order squared-bias and variance in a tie-breaker design compared to a regression discontinuity design. The formulas simplify when defining the quantity \begin{equation}\label{eq:alpha_definition}
     \alpha \equiv \frac{5}{4} \big(  \mu_+^{(2)}(t) - \mu_-^{(2)}(t) \big)^{2/5} \Big( \frac{  \sigma_+^2(t) + \sigma_-^2(t) }{  f(t)}     \Big)^{4/5},
 \end{equation} which does not depend on $h$, $N$ or the kernel choice.


\begin{figure}[t]
\centering
\includegraphics[width=1 \hsize]{bias_variance_tradeoff.png}
\caption{\label{fig:bias_var_tradeoff} A comparison of the asymptotic bias-variance tradeoff for the regression discontinuity design (dotted lines) versus for the tie-breaker design (solid lines). The x-axes are in units of asymptotically MSE optimal bandwidth for RDD given at \eqref{eq:hopt_RDD} while the y-axes are the leading order terms in units of $\alpha N^{-4/5}$ where $N$ is the sample size and $\alpha$ is a constant given in \eqref{eq:alpha_definition} that depends on properties of the joint distribution of $(X,Y,Z)$ in a neighborhood of the cutoff. }
\end{figure}


In practice, the optimal bandwidth is not known and must be estimated. For both the RDD and the TBD, the optimal bandwidth depends on the quantity \begin{equation}\label{eq:def_gamma}
    \gamma \equiv \Big( \frac{\sigma_+^2(t) +\sigma_-^2(t)}{f(t) \big( \mu_+^{(2)}(t)- \mu_-^{(2)}(t)\big)^2} \Big)^{1/5}
\end{equation} which must be estimated from the observed data. We consider the regularized estimator for $\gamma$ of \cite{ImbensKalyanaraman_optimalBW}. We take the estimated optimal bandwidth $\hat{h}_{\text{opt}}$ proposed in their Section 4.2 and set $\hat{\gamma}_{\text{RDD}}= (4 \tilde{C}_1/\tilde{C_2} )^{1/5}\hat{h}_{\text{opt}} N^{1/5}$. It can be seen from the proof of Theorem 4.1 in \cite{ImbensKalyanaraman_optimalBW} that under assumptions (i)--(vi), $\hat{\gamma}_{\text{RDD}} \xrightarrow{p} \gamma$.
In the TBD case, we know a consistent estimator of $\gamma$ exists. For example, if we let $\hat{\gamma}_{\text{TBD,naive}}$ be an estimator of $\gamma$ that is constructed similarly to $\hat{\gamma}_{\text{RDD}}$ using only the subset of the data which looks like an RDD, $\hat{\gamma}_{\text{TBD,naive}} \xrightarrow{p} \gamma$. Of course such an estimator of $\gamma$ is inefficient; in practice one should instead use an estimator of $\gamma$ that does not throw out all of the control samples for which $x>t$ and all of the treated samples for which $x<t$. For our theoretical comparison of TBDs with RDDs, we are not concerned with the actual form of $\hat{\gamma}_{\text{TBD}}$ as long as it is consistent. Therefore, in the TBD case we will let $\hat{\gamma}_{\text{TBD}}$ be any estimator that satisfies $\hat{\gamma}_{\text{TBD}} \xrightarrow{p} \gamma$. We make a few remarks about estimation of $\gamma$ in the TBD setting in the discussion section.


To compare the AMSE for the RDD versus the TBD, we will assume that if the investigator were to run an RDD and were seeking mean squared optimal estimation of $\tau_{\mathrm{thresh}}$, they would ultimately use the bandwidth \begin{equation}\label{eq:hopt_RDD_est}
    \hat{h}_{\text{opt,RDD}}(N) = \Big( \frac{\tilde{C}_2}{4 \tilde{C}_1} \Big)^{1/5}  \hat{\gamma}_{\text{RDD}}  N^{-1/5},
\end{equation}
where $\hat{\gamma}_{\text{RDD}}$ is the consistent estimator for $\gamma$ described above and $\tilde{C}_j$ are defined at~\eqref{eq:defctilde}.
We will also assume that if the investigator were to run a TBD seeking mean squared optimal estimation of $\tau_{\mathrm{thresh}}$, they would ultimately use the bandwidth \begin{equation}\label{eq:hopt_TBD_est}
    \hat{h}_{\text{opt,TBD}}(N) = \Big( \frac{C_2}{4 C_1} \Big)^{1/5}  \hat{\gamma}_{\text{TBD}}  N^{-1/5},
\end{equation} where $\hat{\gamma}_{\text{TBD}}$ is any consistent estimator of $\gamma$ and $C_j$ are defined at \eqref{eq:C_def}.

The following theorem compares the RDD with $N$ points to a TBD
with $\theta N$ points for some $\theta>0$. We will use the value of $\theta$ that provides equal MSEs for estimation of $\tau_{\mathrm{thresh}}$ as a metric to compare the two designs.

\begin{theorem}\label{theorem:MSE_RDD_versu_TBD}
Let $\theta>0$ be a constant. Under assumptions (i)--(vi), as $N \to \infty$
\begin{equation}\label{eq:MSE_rat}
    \frac{\emph{MSE}_{\emph{RDD}}\big(\hat{h}_{\emph{opt,RDD}}(N),N \big)}{\emph{MSE}_{\emph{TBD}}\big(\hat{h}_{\emph{opt,TBD}}(\theta N), \theta N \big)} \xrightarrow{p} \theta^{4/5} \Big( \frac{\tilde{C}_1 \tilde{C}_2^4}{C_1 C_2^4}\Big)^{1/5}
\end{equation}
holds for any tie-breaker design
of the form \eqref{eq:threelevel} with $\Delta>0$.
\end{theorem}

\begin{proof}
See Appendix~\ref{sec:relMSETheoremProof}.
\end{proof}

Theorem~\ref{theorem:MSE_RDD_versu_TBD} uses the
assumption that $\Pr(Z_i=1 \!\mid\! x_i )=1/2$ for $x_i$ in
the randomization window.
If $\sigma^2_+(t)\ne\sigma^2_-(t)$, then we might
prefer to offer the treatment with probability
$p\ne1/2$.
In Appendix \ref{sec:GenPneqHalf}, we study
a treatment probability $p\in(0,1)$.
When $p = \sigma_+(t)/(\sigma_+(t) +\sigma_-(t))$, the asymptotic MSE is minimized, though
an investigator would also want to account for the
cost of the treatment.  If one chooses $p$
using poor prior estimates of $\sigma_\pm$ it is
possible that the resulting TBD will have
a higher asymptotic MSE than the RDD.  However, for any of the kernels
in Table~\ref{table:Kernels}, one can
protect against that by choosing
$p\in[0.18,0.82]$.

\subsection*{Asymptotic MSE comparison for some specific kernels}


We now use Theorem~\ref{theorem:MSE_RDD_versu_TBD} to compare the MSE in estimating $\tau_{\mathrm{thresh}}$ for the RDD versus the TBD, under optimal bandwidth choices for various kernels of interest. See Table \ref{table:Kernels}. If an investigator is deliberating between an RDD with $N$ samples versus conducting a TBD (for a fixed $\Delta>0$) with $N$ samples, and either experimental design is to be analyzed with the asymptotically optimal bandwidth choice for the prespecified kernel,
then the ratio of the MSEs will converge in probability to $\big((\tilde{C}_1 \tilde{C}_2^4)/(C_1 C_2^4)\big)^{1/5}$ as $N \to \infty$. Using formulas \eqref{eq:defctilde} and \eqref{eq:C_def}, the fourth column of Table \ref{table:Kernels} gives the value of the quantity $\big((\tilde{C}_1 \tilde{C}_2^4)/(C_1 C_2^4)\big)^{1/5}$ rounded to 2 decimal places. For the boxcar and triangular kernels respectively, this quantity is precisely $64^{1/5}$ and $60.46618^{1/5}$ without rounding.

It is also interesting to consider the quantity given by \begin{equation}\label{eq:DefthetaStar}
    \theta_* = \frac{C_1^{1/4}C_2}{\tilde{C}_1^{1/4} \tilde{C}_2}.
\end{equation}
As a result of Theorem \ref{theorem:MSE_RDD_versu_TBD}, an experimental designer deciding to use a TBD rather than an RDD would only need to collect $\theta_*$ times as many samples in order to achieve the same asymptotic MSE in estimating $\tau_{\mathrm{thresh}}$.




\begin{table}[!t]
\caption{An asymptotic comparison of the regression discontinuity designs with tie-breaker designs in Kernel regression-based estimation of $\tau_{\mathrm{thresh}}$. The fourth column gives the quantity $\big((\tilde{C}_1 \tilde{C}_2^4)/(C_1 C_2^4)\big)^{1/5}$, which is computed using \eqref{eq:defctilde} and \eqref{eq:C_def} and rounded to 2 decimal places. The last column gives the quantity $\theta_*$ given by \eqref{eq:DefthetaStar} rounded to 3 decimal places.}
\label{table:Kernels}
\begin{center}
\begin{tabular}{ l l c c c }
\toprule
 Kernel & Function $K(u)$ & Support & Relative AMSE & $\theta_*$  \\
\midrule
 Boxcar & $1/2$ & $[-1,1]$  & 2.30 & 0.354 \\ [1ex]

Triangular & $(1-\vert u \vert)_+$ & $[-1,1]$  & 2.27 &  0.359 \\  [1ex]

 Epanechnikov & $ \frac{3}{4}(1- u^2)_+$ & $[-1,1]$ & 2.31 & 0.351 \\ [1ex]

 Quartic & $ \frac{15}{16} (1-u^2)_+^2$  & $[-1,1]$ & 2.29 & 0.355  \\ [1ex]

 Triweight & $ \frac{35}{32} (1-u^2)_+^3$  & $[-1,1]$ & 2.28 & 0.357  \\ [1ex]


  Tricube & $ \frac{70}{81} (1-\vert u\vert^3 )_+^3$ & $[-1,1]$ & 2.31 & 0.352  \\ [1ex]


 Cosine & $\frac{\pi}{4} \cos \big( \frac{\pi}{2} u \big)$  & $[-1,1]$ & 2.31 & 0.352 \\
 \bottomrule
\end{tabular}
\end{center}
\end{table}

Table \ref{table:Kernels} shows that the kernel choice has a remarkably small impact on the relative benefit of using a TBD rather than an RDD to estimate $\tau_{\mathrm{thresh}}$. It is well known in the usual kernel smoothing setting that there is little difference in performance among the widely used kernels. See \cite{wand:jone:1994}.
 If $\tau_{\mathrm{thresh}}$ is to be estimated with local linear regression using one of the seven popular kernel choices exhibited in Table \ref{table:Kernels}, then the RDD has an asymptotic MSE that is about 2.3 times as large as that of the TBD, and the TBD will require 64 to 65 percent fewer samples than the RDD in order to achieve the same asymptotic MSE.

We think that a version of Theorem~\ref{theorem:MSE_RDD_versu_TBD}
will hold also for unbounded kernels such as the $\mathcal{N}(0,1)$ density under
reasonable but stronger regularity conditions on $f(\cdot)$, $\mu_\pm(\cdot)$
and $\sigma_\pm(\cdot)$. We do not develop such a result
as \cite{cattaneo2022regression} state that kernels with unbounded
support are not used in RDD analysis.



\section{
Variance comparisons at fixed $h> \Delta$} \label{sec:AsymApprox_via_integrals}


The AMSE comparison in Section \ref{sec:bandwidth_shrink_asymptotics} depends upon the optimal TBD bandwidth, $\hat{h}_{\text{opt,TBD}}$, eventually becoming smaller than the
positive experimental radius $\Delta$.
However, the optimal $h$ converges to zero only at the very
slow rate $N^{-1/5}$.  Furthermore, the constant in
that rate includes the factor
$|\mu_+^{(2)}(t)-\mu_-^{(2)}(t)|^{-2/5}$
which could be very large.
We believe that in many applied settings the optimal
value of $h$ will not be smaller than $\Delta$.
Then $\Delta/h$ is not necessarily
within the support of the kernel and $Y$ values
data from outside the experimental region are included in the local linear regression.

In this section, we complement the
prior analysis with one where $h$ is fixed
and larger than $\Delta$.
We assume a symmetric
kernel function that is Lipschitz continuous on its support.

The kernel regression estimate of
$\hat\tau_{\mathrm{thresh}}$ has a leading bias of $O(h^2)$.
In the regime where the bandwidth is bigger than $\Delta$, mean squared optimality analysis for estimating $\tau_{\mathrm{thresh}}$ is complicated by the fact that for the TBD there will often exist an $h> \Delta$ such that the constant in this $O(h^2)$ term vanishes.
Remarkably, such a bandwidth depends only on the experimental radius $\Delta$ and the kernel $K$. It does not depend on $\mu_+$, $\mu_-$, $f$, or $N$. In Appendix \ref{sec:MagicBandwidth}, we prove that under certain regularity conditions on $f$, $\mu_{\pm}$, and $K$, a bandwidth $h$ that solves $\nu_2^2 = 4 \int_{\Delta/h}^{\infty} uK(u) \,\mathrm{d} u  \int_{\Delta/h}^{\infty} u^3K(u) \,\mathrm{d} u$ removes the leading order bias, and moreover, such a solution exists. See Table \ref{table:MagicBandwidth} for numerical solutions of this equation for the kernel choices considered previously. We find that for these kernel choices, the bandwidth removing the leading order bias ranges from approximately $3.13 \Delta$ for the Boxcar kernel to approximately $4.84 \Delta$ for the Triweight kernel. We caution investigators against picking this bandwidth
because it does not shrink with $N$. It
could place too little weight on reducing variance for small $N$ and the third order bias term will be $O(1)$.

Due to the existence of a fixed bandwidth bigger than $\Delta$ that removes the leading order bias of $\hat{\tau}_{\mathrm{thresh}}$, analysis of bias and mean squared optimality using second order Taylor expansions of $\mu_{\pm}(\cdot)$ would be misleading. Hence, we do not conduct an analysis similar to that seen in Section \ref{sec:bandwidth_shrink_asymptotics} for the regime where $h>\Delta$.
For that regime, we instead restrict our attention to the variance in estimating $\tau_{\mathrm{thresh}}$ at a fixed bandwidth $h$.

The variance of the local linear estimator $\hat{\tau}_{\mathrm{thresh}}$ given in \eqref{eq:kernelcriterion} and \eqref{eq:def_hat_tauthresh} can be computed as follows. The design matrix for the regression is $\mathcal{X}\in\mathbb{R}^{N\times 4}$ with $i$'th
row $(1,x_i,Z_i,x_iZ_i)$.  The response is $\mathcal{Y} = (Y_1,\dots,Y_N)^\mathsf{T}$.
For simplicity we assume $\mathrm{var}(\mathcal{Y} \!\mid\! \mathcal{X})=\sigma^2 I_{N}$ and without loss of generality we assume $t=0$. The kernel weights are $K(x_i/h)$, and we
let $\mathcal{W}=\mathcal{W}(h)\in\mathbb{R}^{N\times N}=\mathrm{diag}(K(x_i/h))$. Then
\begin{align}\label{eq:betahat}
\hat\beta=\hat\beta(\Delta)
=(\mathcal{X}^\mathsf{T}\mathcal{W}\mathcal{X})^{-1}\mathcal{X}^\mathsf{T}\mathcal{W}\mathcal{Y},
\end{align}
and under the assumption that $\mathrm{var}(\mathcal{Y} \!\mid\! \mathcal{X})=\sigma^2 I_{N}$ we have
\begin{align}\label{eq:varbetahat}
\mathrm{var}(\hat\beta\!\mid\! \mathcal{X};\Delta) = (\mathcal{X}^\mathsf{T}\mathcal{W}\mathcal{X})^{-1}\mathcal{X}^\mathsf{T}\mathcal{W}^2\mathcal{X}(\mathcal{X}^\mathsf{T}\mathcal{W}\mathcal{X})^{-1}\sigma^2.
\end{align}
Formula~\eqref{eq:betahat}  for $\hat\beta$ matches the familiar generalized least
squares formula for the case where $\mathrm{var}(\mathcal{Y}\!\mid\! \mathcal{X})=\mathcal{W}\sigma^2$.  Here $\mathcal{W}$
arises from weights that are not of inverse variance type and hence the formula for $\mathrm{var}(\hat\beta\!\mid\!\mathcal{X};\Delta)$ involves a $\mathcal{W}^2$ factor and less cancellation than we might have expected. The boxcar kernel is special because then $K(x_i/h)\in\{0,1\}$ equals its own square. In that case $\mathrm{var}(\hat\beta\!\mid\!\mathcal{X};\Delta)=(\mathcal{X}^\mathsf{T}\mathcal{W}\mathcal{X})^{-1}\sigma^2$. The estimator is $\hat{\tau}_{\mathrm{thresh}} = 2\hat{\beta}_3$. Therefore, we study $\mathrm{var}(\hat\beta_3\!\mid\!\mathcal{X};\Delta)$ under a tie-breaker design as $(\mathrm{var}(\hat\beta\!\mid\!\mathcal{X};\Delta))_{3,3}$ using the expression in~\eqref{eq:varbetahat}.



At the stage where the experiment is being designed and $\Delta$ is being chosen, the investigator does not have much information about $\mathcal{X}\in\mathbb{R}^{N\times 4}$ but we will later see, quite a bit is known about the quantity $N \times \mathrm{var}(\hat\beta_3\!\mid\!\mathcal{X};\Delta)/\sigma^2$. For $x_i$ from a real dataset, we see in Section \ref{sec:Data_application} (e.g. Figure \ref{fig:MC_boxcar_plots_AL}) that this quantity does not vary much for different simulations of the random treatment assignments $(Z_i)_{i=1}^N$. To get theoretical insight, we turn our attention to the uniformly spaced setting with $x_i=(2i-N-1)/N$ to develop tractable theoretical results. We give an asymptotic justification
for this assumption using results from
\cite{fan1996local} in Section~\ref{sec:Data_application}. This rank transformation is also used in \cite{owen:vari:2020}.



For $x_i = (2i-N-1)/N$, the matrices $\mathcal{X}^\mathsf{T}\mathcal{W}\mathcal{X}/N$ and $\mathcal{X}^\mathsf{T}\mathcal{W}^2\mathcal{X}/N$ contain elements that can be approximated by integrals of the form
\begin{align}
\mathcal{I}_{}^{rst}=\mathcal{I}^{rst}(\Delta,h,K)\equiv\frac12\int_{-1}^1x^r\mathbb{E}(Z^s\!\mid\! x;\Delta)K\Bigl(\frac{x}h\Bigr)^t\,\mathrm{d} x
\end{align}
for integer exponents $r$, $s$ and $t$.
Our expressions will simplify somewhat because $Z^2=1$ making every $\mathcal{I}^{r,2,t}=\mathcal{I}^{r,0,t}$ and also because both $x$ and $\mathbb{E}(Z\!\mid\! x;\Delta)$ are antisymmetric functions of $x$ making them orthogonal to $K(x/h)$ which we have assumed to be symmetric.
The error in those moment approximations is
$O_p(N^{-1/2})$ if the $Z_i$ are independent
random variables. The error can be much less with
other sampling schemes. For instance, we could use stratified sampling, forming pairs of subjects $(i,i+1)$ in the experimental region and randomly setting $Z_i=\pm1$ and $Z_{i+1}=-Z_i$.
We will use $\approx$ to describe approximations that are $O_p(N^{-1/2})$ or better.

Applying first $Z^2=1$ and then using symmetry and
anti-symmetry
\begin{align*}
\frac1N\mathcal{X}^\mathsf{T}\mathcal{W}\mathcal{X} &\approx
\kbordermatrix{ & 1 & x & z & xz \\
1 &\mathcal{I}^{001} & \mathcal{I}^{101} & \mathcal{I}^{011} &\mathcal{I}^{111}\\[1ex]
x &\mathcal{I}^{101} & \mathcal{I}^{201} & \mathcal{I}^{111}&\mathcal{I}^{211}\\[1ex]
z &\mathcal{I}^{011} & \mathcal{I}^{111} & \mathcal{I}^{021} & \mathcal{I}^{121}\\[1ex]
xz & \mathcal{I}^{111} &\mathcal{I}^{211} & \mathcal{I}^{121} & \mathcal{I}^{221}
 }
&=
\begin{bmatrix}
\mathcal{I}^{001} & 0 & 0 &\mathcal{I}^{111}\\[1ex]
0 & \mathcal{I}^{201} & \mathcal{I}^{111}&0\\[1ex]
0& \mathcal{I}^{111} & \mathcal{I}^{001} & 0\\[1ex]
\mathcal{I}^{111} & 0 &0 & \mathcal{I}^{201}
\end{bmatrix}.
\end{align*}
Because $K^2(\cdot)$ is also a symmetric function we also get
$$
\frac1N\mathcal{X}^\mathsf{T}\mathcal{W}^2\mathcal{X} \approx
\begin{bmatrix}
\mathcal{I}^{002} & 0 & 0 &\mathcal{I}^{112}\\[1ex]
0 & \mathcal{I}^{202} & \mathcal{I}^{112}&0\\[1ex]
0& \mathcal{I}^{112} & \mathcal{I}^{002} & 0\\[1ex]
\mathcal{I}^{112} & 0 &0 & \mathcal{I}^{202}
\end{bmatrix}.
$$

From all of the symmetries involved in the 32 components of these two matrices, we need to consider at most six distinct integrals.
We rewrite those matrices, beginning with
 \begin{align}\label{eq:xtwxbyn} \frac{1}{N}  \mathcal{X}^T \mathcal{W}  \mathcal{X} \approx \begin{bmatrix} \kappa_0 & 0 & 0 & \phi(\Delta) \\ 0 & \kappa_2 & \phi(\Delta) & 0 \\ 0 & \phi (\Delta) & \kappa_0 & 0 \\ \phi(\Delta) & 0 & 0 & \kappa_2 \end{bmatrix}
\end{align}
where
\begin{align}\label{eq:defnu}
\begin{split}\kappa_0&
=\frac12\int_{-1}^{1}K\Bigl(\frac{x}h\Bigr)\,\mathrm{d} x,\quad
\kappa_2
=\frac12\int_{-1}^{1}x^2K\Bigl(\frac{x}h\Bigr)\,\mathrm{d} x,\quad\text{and}\\
\phi(\Delta)&
=\frac12\int_{-1}^{-\Delta}(-x)K\Bigl(\frac{x}h\Bigr)\,\mathrm{d} x +\frac12\int_{\Delta}^{1}xK\Bigl(\frac{x}h\Bigr)\,\mathrm{d} x=\int_{\Delta}^1xK\Bigl(\frac{x}h\Bigr)\,\mathrm{d} x.
\end{split}
\end{align}
Note that $\kappa_0$ and $\kappa_2$ may depend on $h$ but they do not depend on $\Delta$.
A similar argument shows that
\begin{align}\label{eq:xtwwxbyn}   \frac{1}{N}  \mathcal{X}^T \mathcal{W}^2  \mathcal{X} \approx \begin{bmatrix} \lambda_0 & 0 & 0 & \psi(\Delta) \\ 0 & \lambda_2 & \psi(\Delta) & 0 \\ 0 & \psi (\Delta) & \lambda_0 & 0 \\ \psi(\Delta) & 0 & 0 & \lambda_2 \end{bmatrix} \end{align}
for
\begin{align}\label{eq:defpi}
\begin{split}\lambda_0
&=\frac12\int_{-1}^{1}K^2\Bigl(\frac{x}h\Bigr)\,\mathrm{d} x,\quad
\lambda_2=\frac12\int_{-1}^{1}x^2K^2\Bigl(\frac{x}h\Bigr)\,\mathrm{d} x,\quad\text{and}\\
\psi(\Delta)&=\int_{\Delta}^1xK^2\Bigl(\frac{x}h\Bigr)\,\mathrm{d} x.\end{split}
\end{align}
Now we are ready to describe the asymptotic
variance of $\hat\beta_3$.
\begin{theorem}\label{thm:varbetahat3}
Let $x_i= (2i-N-1)/N$, select $Z_i\in\{-1,1\}$ by the tie-breaker
equation~\eqref{eq:threelevel} with $t=0$. Let $Y_i$ be uncorrelated random
variables with common variance $\sigma^2$, conditionally on
$\mathcal{X}=( (1,x_1,Z_1,x_1Z_1),\cdots,(1,x_N,Z_N,x_NZ_N))$. Next, for a symmetric kernel
$K(\cdot)\geqslant0$ that is Lipschitz continuous on its support
and a bandwidth
$h>0$, let $\hat\beta$ be estimated by the kernel weighted
regression~\eqref{eq:kernelcriterion}. Then
\begin{align}\label{eq:varbetahat3}
N\mathrm{var}(\hat\beta_3\!\mid\!\mathcal{X};\Delta) =
\frac{\sigma^2 \big(\kappa_2^2 \lambda_0-2 \kappa_2 \phi(\Delta) \psi(\Delta) +\lambda_2 \phi^2(\Delta) \big)}{ \big(\kappa_0 \kappa_2 - \phi^2(\Delta)  \big)^{2}}
+O_p\Bigl(\frac1{\sqrt{N}}\Bigr),
\end{align}
where $\kappa_0$, $\kappa_2$ and $\phi(\Delta)$ are defined in~\eqref{eq:defnu} and $\lambda_0$, $\lambda_2$ and $\psi(\Delta)$ are defined in~\eqref{eq:defpi}.
\end{theorem}
\begin{proof}
Reordering the components of $\beta$ we find after substituting equations~\eqref{eq:xtwxbyn} and~\eqref{eq:xtwwxbyn} into~\eqref{eq:varbetahat} that $\sqrt{N}(\hat\beta_1,\hat\beta_4,\hat\beta_2,\hat\beta_3)$ has variance
\begin{align*}
\begin{pmatrix}
\kappa_0 &\phi & 0 & 0\\
\phi & \kappa_2 & 0 & 0\\
0 & 0 & \kappa_0 & \phi\\
0 & 0 & \phi & \kappa_2
\end{pmatrix}^{\!\!-1}
\!\!\!
\begin{pmatrix}
\lambda_0 &\psi & 0 & 0\\
\psi & \lambda_2 & 0 & 0\\
0 & 0 & \lambda_0 & \psi\\
0 & 0 & \psi & \lambda_2
\end{pmatrix}
\!\!
\begin{pmatrix}
\kappa_0 &\phi & 0 & 0\\
\phi & \kappa_2 & 0 & 0\\
0 & 0 & \kappa_0 & \phi\\
0 & 0 & \phi & \kappa_2
\end{pmatrix}^{\!\!-1}\!\!\sigma^2+O_p\Bigl(\frac1{\sqrt{N}}\Bigr).
\end{align*}
Now~\eqref{eq:varbetahat3} follows directly by matrix inversion and multiplication.
\end{proof}


Our Lipschitz condition on the kernel $K$ is present for technical reasons. Using a formulation with fixed and discrete $x_i=(2i-N-1)/N$, this condition allows us to obtain the same error rate of $O_p(N^{-1/2})$ as would be obtained using the formulation with random $x_i \stackrel{\text{IID}}{\sim} \mathbb{U}[-1,1]$. Without the Lipschitz condition, an adversarially chosen kernel $K(\cdot)$ might have point discontinuities at every rational multiple of the bandwidth $h$. We remark that our Lipschitz condition can be loosened to a $1/2$-Hölder continuity condition, with details available from the first author upon request.

The variance formula in Theorem~\ref{thm:varbetahat3}
does not require the linear model~\eqref{eq:ovmodel}
to hold.
When it does not hold there will generally be some bias
where $\mathbb{E}(2\hat\beta_3\!\mid\!\mathcal{X};\Delta)\ne \mu_+(0)-\mu_-(0)$.
We suppose that the user will choose an $h$ to appropriately navigate the bias variance tradeoff, but that step takes place after the outcomes $Y_i$ are observed, which are not available when $\Delta$ is chosen, so we resort to comparing the variance for any choice of $h$.


We are primarily interested in comparing the asymptotic variance of $\hat{\tau}_0=2\hat\beta_3$ for various  choices of $\Delta$. We especially want to compare the efficiency of tie-breaker designs with $\Delta>0$ to the RDD with $\Delta=0$. To do this we consider the efficiency ratio
\begin{align}\label{eq:defasyeffN}
\mathrm{Eff}^{(N)}(\Delta) \equiv  \frac{\mathrm{var}(\hat{\tau}_{\mathrm{thresh}}\!\mid\!\mathcal{X}; 0)}{\mathrm{var}(\hat{\tau}_{\mathrm{thresh}} \!\mid\! \mathcal{X}; \Delta)}=
\frac{\mathrm{var}(\hat{\beta}_3\!\mid\!\mathcal{X}; 0)}{\mathrm{var}(\hat{\beta}_3 \!\mid\! \mathcal{X}; \Delta)}.
\end{align}

Using Theorem~\ref{thm:varbetahat3}, $\mathrm{Eff}^{(N)}(\Delta)$ converges in
probability to the asymptotic efficiency ratio
\begin{align}\label{eq:general_efficiency_ratio} \mathrm{Eff}(\Delta) =  \frac{\big(\kappa_2^2 \lambda_0-2 \kappa_2 \phi(0) \psi(0) +\lambda_2 \phi^2(0) \big) \big(\kappa_0 \kappa_2 - \phi^2(\Delta)  \big)^{2} }{ \big(\kappa_2^2 \lambda_0-2 \kappa_2 \phi(\Delta) \psi(\Delta) +\lambda_2 \phi^2(\Delta) \big) \big(\kappa_0 \kappa_2 - \phi^2(0)  \big)^{2}}  \end{align}
using quantities that we defined at~\eqref{eq:defnu} and~\eqref{eq:defpi}.

\subsection*{Efficiency with boxcar and triangular kernels}



In this subsection we present the efficiency ratios under the conditions of Theorem~\ref{thm:varbetahat3} for the two kernels of greatest interest: the boxcar kernel and the triangular kernel. We work with $x_i = (2i-N-1)/N$ throughout this subsection.

For the boxcar kernel
$K_{\mathrm{BC}}(u) = 1_{|u|\leqslant 1}$,
we can assume without loss of generality that $h\leqslant 1$ because there are no data with $|x_i-t|=|x_i|>1$, and then any $h>1$ will give the same estimate as $h=1$. We find for this kernel that
\begin{align}\label{eq:boxcarpieces}\kappa_0=\lambda_0= h,\quad \kappa_2=\lambda_2=\frac{h^3}{3},\quad\text{and}\quad \phi(\Delta)=\psi(\Delta)=\frac{(h^2- \Delta^2)_+}{2}.
\end{align}
Using some foresight, we define the local tie-breaker constant $\delta=\Delta/h$. This is the fraction of the local regression region in which the treatment was assigned at random.

\begin{proposition}\label{prop:boxcarefficiency}
Under the conditions of Theorem~\ref{thm:varbetahat3} and using the boxcar kernel $K_{\mathrm{BC}}$, the asymptotic efficiency ratio of the tie-breaker design is
\begin{align}\label{eq:boxcarefficiency}
\mathrm{Eff}_{\mathrm{BC}} = 1+6\delta^2-3\delta^4
\end{align}
for $\delta = \Delta/h\leqslant1$. If $\delta>1$, then $\mathrm{Eff}_{\mathrm{BC}}=4$.
\end{proposition}
\begin{proof}
Because many quantities from~\eqref{eq:boxcarpieces} are identical, substituting them into \eqref{eq:general_efficiency_ratio} produces numerous simplifications that yield
\begin{align*}
\mathrm{Eff}_{\mathrm{BC}} &= \frac{\kappa_0\kappa_2-\phi^2(\Delta)}{\kappa_0\kappa_2-\phi^2(0)}
=\frac{\frac{h^4}3-\frac{(h^2-\Delta^2)_+^2}4}{\frac{h^4}3-\frac{h^4}4}
=4-3(1-\delta^2)_+^2.
\end{align*}
For $0\leqslant\delta<1$ formula~\eqref{eq:boxcarefficiency}  follows from expanding the quadratic while for $\delta>1$ the positive part term vanishes.
\end{proof}

Choosing $h=1$ makes the local regression a global one. We then get the same efficiency ratio as in equation (6) from \cite{owen:vari:2020}. By taking derivatives it is easy to show that the efficiency ratio in~\eqref{eq:boxcarefficiency} is strictly increasing as the local amount of experimentation $\delta$ varies over the interval $0<\delta< 1$.
Figure~\ref{fig:Unif_Efficiecy_plots} plots $\mathrm{Eff}_{\mathrm{BC}}$ versus $\delta$.


The triangular spike kernel
$K_{\mathrm{TS}}(x) = (1 -\vert x \vert )_+$
(triangular kernel for short) is more complicated than the boxcar kernel because for it,
$K^2$ is not proportional to $K$.
Once again, we assume that $h \in [0,1]$. For this kernel we compute
$$\kappa_0= \frac{h}{2},\quad \kappa_2 = \frac{h^3}{12},
\quad \lambda_0 = \frac{h}{3},\quad\text{and}\quad \lambda_2=\frac{h^3}{30}$$
and then using $\delta = \Delta/h$, we get
$$\phi(\Delta)= \frac{h^2}{6} (1- 3 \delta^2+ 2 \delta^3)\quad\text{and}\quad\psi(\Delta)= \frac{h^2}{12} (1- 6 \delta^2+ 8 \delta^3 - 3 \delta^4).$$


\begin{proposition}\label{prop:triangleefficiency}
Under the conditions of Theorem~\ref{thm:varbetahat3} and using the triangular kernel $K_{\mathrm{TS}}$, the asymptotic efficiency of the tie-breaker design is
\begin{align}\label{eq:triangleefficiency}
\mathrm{Eff}_{\mathrm{TS}} =
\frac{ 2\bigl(3-2(1-3\delta^2+2\delta^3)^2\bigr)^2  }
{ 5-5(1-3\delta^2+2\delta^3)(1-6\delta^2+8\delta^3-3\delta^4)+2(1-3\delta^2+2\delta^3)^2 }
\end{align}
for $\delta = \Delta/h\leqslant1$.
\end{proposition}
\begin{proof}
This follows from plugging in the values of $\kappa_0$, $\kappa_2$, $\lambda_0$, $\lambda_2$, $\phi(\Delta)$, and $\psi(\Delta)$ for the triangular kernel into \eqref{eq:general_efficiency_ratio}. See Appendix \ref{sec:TSefficiencyCalc} for the explicit calculations.
\end{proof}

The second panel in Figure~\ref{fig:Unif_Efficiecy_plots}
shows $\mathrm{Eff}_{\mathrm{TS}}$ versus the local experiment size $\delta$. The efficiency curve has a similar monotone increasing shape as we saw for the boxcar kernel.  The maximum efficiency ratio, at $\delta=1$, is $18/5=3.6$ instead of $4$. The efficiency ratio is a rational function of $\delta$ with a numerator of degree $12$ and a denominator of degree $7$.  It is strictly increasing on the interval $0<\delta<1$, though the proof is lengthy enough to move to the Appendix.
\begin{proposition}\label{prop:itsmonotone}
The derivative of $\mathrm{Eff}_{\mathrm{TS}}$ with respect to $\delta$ is positive for $0<\delta<1$.
\end{proposition}
\begin{proof}
See Appendix \ref{sec:MonotoneEffTS}.
\end{proof}

\begin{figure}[t!]
    \centering
        \begin{subfigure}[b]{6 cm}
            \centering
            \includegraphics[width=6 cm]{EfficiencyBoxcarUnif3.png}
            \label{fig:boxcarERUnifPlot}
        \end{subfigure}
        \begin{subfigure}[b]{6 cm }
            \centering
            \includegraphics[width=6 cm]{EfficiencyTriangularUnif3.png}
            \label{fig:triangularERUnifPlot}
        \end{subfigure}
        \caption[]
        {\label{fig:Unif_Efficiecy_plots}
        The left panel shows the efficiency ratio of the tie-breaker design for uniform $x_i$ and the boxcar kernel as a function of $\delta=\Delta/h$.  The right panel shows this efficiency ratio for the triangular kernel.
        }
\end{figure}




\section{Classroom size data} \label{sec:Data_application}

We explored the efficiency ratio for the tie-breaker design for $x_i$ with a uniform distribution.  While that can be arranged by using ranks, in other situations we might prefer to use the original value of a running variable and those might not be uniformly distributed.
We show how to do this using a dataset from \cite{AngristLavy} on classroom sizes.

\cite{AngristLavy} studied the causal effect of classroom size on test performance of elementary school students in Israel. In Israel,
the Maimonides rule mandates that elementary school classes cannot exceed 40 students. If a school has 41 students enrolled in a particular grade that grade must be split into two classes. Note that grades that have 40 or fewer enrolled students are allowed to split into multiple classes and that grades with slightly more than 40 students occasionally violate the Maimonides rule and do not split into multiple classes. Despite this, we can consider this a setting for RDD where the treatment variable is whether or not the school is legally mandated to split a particular grade into smaller classes.

The dataset, published on the Harvard Dataverse \citep{ALHarvard2}, has verbal and math scores for 3rd, 4th and 5th graders across Israel. We chose to focus exclusively on 4th grade verbal scores as our response variable and 4th grade enrollments as our assignment variable because \cite{AngristLavy} suggest that a slightly significant effect of the treatment on 4th grade verbal scores exists.
Even though the data were not
generated by a tie-breaker we can still compute the
relative efficiency that a tie-breaker design would have had.

To simplify the analysis, we removed all schools that either had more than 80 students or more than two 4th-grade classes from the dataset. We further removed all schools that had NA entries for either class size or verbal scores, leaving $N=711$ schools in our filtered dataset. See Figure \ref{fig:hist_of_enrollments} for a visualization of the distribution of the 4th grade enrollments and Figure \ref{fig:RDD_fits_AL} for visualizations of the local linear regression based-RDD on this dataset using boxcar and triangular kernels.  We use the bandwidths $h_{\mathrm{IK}}$ given by the \cite{ImbensKalyanaraman_optimalBW} procedure, which were computed using that paper's MATLAB code.
The apparent benefit from smaller classrooms is positive but small and it turns
out, not statistically significant in this analysis.
The 95\% confidence interval (assuming homoscedastic errors) for the effect size at the boundary of the local linear regression-based RDD was $(-1.5,9.2)$ when a boxcar kernel with bandwidth $h_{\mathrm{IK},\mathrm{BC}} = 7.09$ was used. The 95\% confidence interval for the effect size at boundary of this RDD was $(-2.4, 9.4)$ when a triangular kernel with bandwidth $h_{\mathrm{IK},\mathrm{TS}} = 9.02$ was used.

\begin{figure}[t]
\centering
\includegraphics[width=0.75 \hsize]{hist_of_enrollments.png}
\caption{\label{fig:hist_of_enrollments} A histogram of 4th grade enrollments for our filtered dataset (with schools exceeding 80 4th grade students or three 4th grade classes removed).}
\end{figure}


\begin{figure}[t]
    \centering
        \begin{subfigure}[b]{6 cm}
            \centering
            \includegraphics[width=6 cm]{RDD_boxcarAL.png}
            \label{fig:RDD_boxcarAL}
        \end{subfigure}
        \hfill
        \begin{subfigure}[b]{6 cm}
            \centering
            \includegraphics[width=6 cm]{RDD_triangularAL.png}
            \label{fig:RDD_triangularAL}
        \end{subfigure}

        \caption[]
        {RDD fit to the 4th grader verbal scores from the \cite{ALHarvard2} dataset when using a boxcar kernel (left) and a triangular kernel (right). For these two fits, the bandwidths $h_{\mathrm{IK},\mathrm{BC}}$ and $h_{\mathrm{IK},\mathrm{TS}}$ were chosen as in \cite{ImbensKalyanaraman_optimalBW}. }
        \label{fig:RDD_fits_AL}
\end{figure}

Next we illustrate how an investigator can estimate the efficiency ratio of tie-breaker designs as a function of $\Delta$ on sample values of the assignment variable.
First we translate the data, replacing $x_i$ by $x_i-40.5$ to move the threshold from $t=40.5$ to $t=0$.
Next, for each $\Delta$ of interest we use $1000$ Monte Carlo samples to estimate $\mathrm{var}(\hat\beta_3\!\mid\!\mathcal{X};\Delta)$ and also $\mathrm{var}(\hat\beta_3\!\mid\!\mathcal{X} ;0)$, both up to a constant $\sigma^2$.  That gives us $1000$ efficiency ratios $\mathrm{Eff}^{(N)}(\Delta)=\mathrm{var}(\hat\beta_3\!\mid\!\mathcal{X};0)/\mathrm{var}(\hat\beta_3\!\mid\!\mathcal{X};\Delta)$ for each $\Delta$.
In each of our $1000$ samples, we simulate random assignments $Z_i$ for a tie-breaker design at the given experimental radius $\Delta$. The random assignments are stratified: in each consecutive pair of classroom sizes in the experimental region, one was randomly chosen to have $Z=1$ and the other got $Z=-1$.
The $x_i$ and the random $Z_i$ let us compute the matrices $\mathcal{X}$ and $\mathcal{W}$ defined in the beginning of Section \ref{sec:AsymApprox_via_integrals}, from which we compute a non-asymptotic $\mathrm{var}(\hat{\beta}_3\!\mid\!\mathcal{X};\Delta)$ using \eqref{eq:varbetahat}.
We do not simulate any $Y_i$ values because efficiency only depends on $\mathcal{X}$, and
we are retaining the bandwidths from the \cite{ImbensKalyanaraman_optimalBW} procedure on the original data.  A more detailed simulation randomizing
the bandwidth choice is out of scope. Our simulations demonstrate that the TBD is more efficient at each fixed $h$, so we expect that it will also be more efficient at a randomly chosen $h$. There could be exceptions if the bandwidth is adversarially correlated with the estimation errors but we do not think
that is likely.

Figure \ref{fig:MC_boxcar_plots_AL} shows
boxplots of $1000$ simulated $\mathrm{Eff}^{(N)}(\Delta)$ values
for various choices of $\Delta \in \mathbb{N}$ to plot the full efficiency curve. It is clear from Figure~\ref{fig:MC_boxcar_plots_AL} that with stratified allocations the efficiency is very reproducible.
Figure \ref{fig:Efficiency_plt_many_bandwidths} shows results for different bandwidths, ranging from $h_{\mathrm{IK}}/2$ to $3h_{\mathrm{IK}}/2$.  Because the efficiencies are so reproducible given the bandwidth, we just plot curves of the mean and standard deviations of estimated $\mathrm{Eff}$ values. For both the boxcar and triangular kernels, we see that the tie-breaker design is reproducibly more efficient than the RDD and the effect increases as $\delta =\Delta/h$ increases for all $h$ we studied. The efficiency curves for this dataset under various bandwidth choices look similar to the theoretical efficiency curves derived in Section \ref{sec:AsymApprox_via_integrals} for the case of a uniform assignment variable.

\begin{figure}[t!]
    \centering
        \begin{subfigure}[b]{6 cm}
            \centering
            \includegraphics[width=6 cm]{BoxplotALEfficiency_withTheoryBoxcar.png}
            \label{fig:boxplot_boxcarAL}
        \end{subfigure}
        \hfill
        \begin{subfigure}[b]{6 cm}
            \centering
            \includegraphics[width=6 cm]{BoxplotALEfficiency_withTheoryTriangular.png}
            \label{fig:boxplot_triangularAL}
        \end{subfigure}

        \caption[]
        {Boxplots of the Monte-Carlo efficiency ratio estimates for various values of $\Delta \in \mathbb{N}$ when using a boxcar kernel (left) and a triangular kernel (right). For both kernels we used bandwidths from Figure \ref{fig:RDD_fits_AL}: $h_{\mathrm{IK},\mathrm{BC}}=7.09$ and $h_{\mathrm{IK},\mathrm{TS}}=9.02$. The dashed lines give the theoretical efficiency curves under the assumption of a uniform assignment variable, given by Propositions \ref{prop:boxcarefficiency} (left) and \ref{prop:triangleefficiency} (right).
    \label{fig:MC_boxcar_plots_AL}}
\end{figure}


\begin{figure}[t]
    \centering
        \begin{subfigure}[b]{6 cm}
            \centering
            \includegraphics[width=6 cm]{ExpectedEff_boxcar_AL2.png}
            \label{fig:ExpectedEff_boxcar_AL2}
        \end{subfigure}
        \hfill
        \begin{subfigure}[b]{6 cm}
            \centering
            \includegraphics[width=6 cm]{SD_Eff_AL_boxcar2.png}
            \label{fig:SD_Eff_AL_boxcar2}
        \end{subfigure}
        \begin{subfigure}[b]{6 cm}
            \centering
            \includegraphics[width=6 cm]{ExpectedEff_triangular_AL2.png}
            \label{fig:ExpectedEff_triangular_AL2}
        \end{subfigure}
        \hfill
        \begin{subfigure}[b]{6 cm}
            \centering
            \includegraphics[width=6 cm]{SD_Eff_AL_triangular2.png}
            \label{fig:SD_Eff_AL_triangular2}
        \end{subfigure}
        \caption[]
        {Monte-Carlo estimates of the expected value (left) and standard deviation (right) of $\mathrm{Eff}^{(N)}(\Delta)$ versus $\Delta/h$ for the \cite{ALHarvard2} dataset of 4th grader verbal scores. For these plots a boxcar kernel (top) and a triangular kernel (bottom) were used. The bandwidths plotted are scalar multiples of $h_{\mathrm{IK},\mathrm{BC}}$ and $h_{\mathrm{IK},\mathrm{TS}}$ from the procedure of \cite{ImbensKalyanaraman_optimalBW}. The legend for the plots on the right is the same as for the plots on the left.
        These curves are not smooth because, to avoid redundancy, only points that corresponded to integer values of $\Delta$ were used. The thick dashed lines give the theoretical efficiency curves $\mathrm{Eff}_{\mathrm{BC}}$ (top left) and $\mathrm{Eff}_{\mathrm{TS}}$ (bottom left) derived in Section~\ref{sec:AsymApprox_via_integrals}.}
        \label{fig:Efficiency_plt_many_bandwidths}
\end{figure}


For a further discussion of the Maimonides rule, see
\cite{angr:etal:2019}.  They consider different data
sets and also investigate the possibility that the
class sizes are sometimes manipulated to be above
the threshold triggering a classroom split.

\subsection*{Comparison with theoretical results for uniform assignment variable}
Our theoretical analysis in Section \ref{sec:AsymApprox_via_integrals} is for a uniformly spaced assignment variable. We can offer one explanation for why the empirical efficiencies on non-uniformly distributed data look so similar to the theoretical ones for uniformly distributed data (see the left panels in Figure \ref{fig:Efficiency_plt_many_bandwidths}). The explanation uses some results about non-parametric regression from \citet[Table 2.1]{fan1996local}. Nonparametric regression estimates $\hat\mu(t)$ typically have an asymptotic variance where the leading term is proportional to $1/f(t)$ where $f$ is the probability density of the $x_i$. This arises because the local sample size is asymptotically proportional to $f(t)$. Hence, when considering nonuniform distributions, the $1/f(t)$ factors in the leading order variance terms will cancel out when computing the efficiency ratios. Some of the nonparametric regression estimators, such as the Nadaraya-Watson estimator, have a lead term in their bias that depends on the derivative $f'(t)$, and while $f'(t)=0$ for uniformly distributed data, it is not zero in general. Kernel weighted least squares methods (with symmetric $K(\cdot)$) do not have a dependency on $f'(t)$ in their bias. There is a curvature bias from $\mu''(t)$ but that is not related to the sampling distribution of the $x_i$. The lead terms in bias and variance for local linear regressions do not distinguish between distributions with the same value of $f(t)$ but different $f'(t)$. Thus the effects of non-uniformity of $X$ are asymptotically negligible.

\section{Discussion} \label{sec:Conclusion}

If an investigator is able to implement a 3-level tie-breaker design with any experimental radius $\Delta \in (0,\Delta_{\text{max}})$, our results show that the TBD has considerable statistical advantages over the RDD.

The most obvious advantage is that the TBD allows estimation of multiple causal parameters of interest including the average treatment effect over subjects with $x \in (t- \Delta,t+\Delta)$ as well as the expected treatment effect at any particular $x \in (t- \Delta,t+\Delta)$. The former is estimable at a faster rate and with fewer assumptions, whereas the latter may still be of interest for choosing a future policy threshold. Meanwhile, the RDD only allows estimation of $\tau_{\mathrm{thresh}}$, the expected treatment effect at $x=t$.

Even if the only goal is estimation of $\tau_{\mathrm{thresh}}$, our results indicate a statistical advantage to running a TBD rather than an RDD and an advantage to picking a larger experimental radius $\Delta \in (0,\Delta_{\text{max}})$. As seen in Section \ref{sec:bandwidth_shrink_asymptotics}, to achieve the same asymptotic MSE in mean squared optimal estimation of $\tau_{\mathrm{thresh}}$, a TBD would require roughly 64 percent fewer samples than would be needed for an RDD. Moreover, the asymptotic advantage for a TBD is largely driven by its lower variance (Figure \ref{fig:bias_var_tradeoff}).
Hence, if the convenient, but controversial, method of undersmoothing to construct asymptotically valid confidence intervals for $\tau_{\mathrm{thresh}}$ is used instead of more nearly optimal approaches, the TBD would exhibit even greater advantages over the RDD. We point readers to the introduction of \cite{Calonico_dont_undersmooth} for an overview of the history of undersmoothing, and \cite{calo:catt:farr:2019} for a modern approach to constructing confidence intervals that has better coverage properties than undersmoothing has.


In terms of the statistical advantages of picking a larger $\Delta$, \cite{owen:vari:2020} found an efficiency advantage for the tie-breaker in a global regression, wherein the estimation variance decreased monotonically in $\Delta$. We provide a comparable finding for the now more standard local linear regression approach: for any fixed bandwidth $h$, we see a theoretical efficiency that increases with the amount $\Delta$ of experimentation. We have not investigated the effect of $\Delta$ on the subsequent choice of $h$ when $\hat{h}_{\text{opt,TBD}} > \Delta$, although one candidate choice is an $h> \Delta$ that removes the leading order bias term, which we derived in Appendix~\ref{sec:MagicBandwidth}.

There is room for an improved estimator of $\gamma$
in the TBD context which uses data from both treatments on both
sides of the threshold $t$.  We leave this for further work.
A critical ingredient is the estimation of $\mu_{\pm}^{(2)}(t)$.
Compared to the method in \cite{ImbensKalyanaraman_optimalBW},
one could use a bandwidth tuned for an internal point $t$ instead
of one tuned for an endpoint.  Also the curvature estimates
in \cite{ImbensKalyanaraman_optimalBW}
use local quadratic regressions while \citet[p 63]{fan1996local}
suggest using local cubic regressions for curvature estimation at an interior point.


\begin{acks}[Acknowledgments]
This work was supported by the U.S.\ National Science
Foundation under grants IIS-1837931
and DMS-2152780 and by Stanford University's SGF and SIGF fellowships.
We thank Hal Varian and Harrison Li for commenting on the paper as well as Steve Marron and Wolfgang H\"ardle for
some discussions about nonparametric regression.
We also thank anonymous reviewers for comments that
led us to improve the paper.
\end{acks}

\bibliographystyle{imsart-nameyear.bst}
\bibliography{TieBreakerWriteUp}