The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
176,096 characters
Robust Empirical Bayes Confidence Intervals
\maketitle
\begin{abstract}
We construct robust \acp{EBCI} in a normal means problem. The intervals are
centered at the usual linear empirical Bayes estimator, but use a critical
value accounting for shrinkage. Parametric \acp{EBCI} that assume a normal
distribution for the means \citep{morris83} may substantially undercover when
this assumption is violated. In contrast, our \acp{EBCI} control coverage
regardless of the means distribution, while remaining close in length to the
parametric \acp{EBCI} when the means are indeed Gaussian. If the means are
treated as fixed, our \acp{EBCI} have an average coverage guarantee: the
coverage probability is at least $1-\alpha$ on average across the $n$
\acp{EBCI} for each of the means. Our empirical application considers the
effects of U.S.\ neighborhoods on intergenerational mobility.\\[1ex]
\textbf{Keywords:} average coverage, empirical Bayes, confidence interval, shrinkage\\
\textbf{JEL codes:} C11, C14, C18
\end{abstract}
\clearpage
\section{Introduction}\label{sec:introduction}
Empirical researchers in economics are often interested in estimating effects
for many individuals or units, such as estimating teacher quality for teachers
in a given geographic area. In such problems, it is common to shrink unbiased
but noisy preliminary estimates of these effects toward baseline values, say the
average effect for teachers with the same experience. In addition to estimating
teacher quality \citep{KaSt08,JaLe08,cfr14i}, shrinkage techniques are used in a
wide range of applications including estimating school quality \citep{ahpw17},
hospital quality \citep{hull20}, the effects of neighborhoods on
intergenerational mobility \citep{chetty_impacts_2018}, and patient risk scores
across regional health care markets \citep{fghw17}.
The shrinkage estimators used in these applications can be motivated by an
\ac{EB} approach. One imposes a working assumption that the individual effects
are drawn from a normal distribution (or, more generally, a known family of
distributions). The \ac{MSE} optimal point estimator then has the form of a
Bayesian posterior mean, treating this distribution as a prior distribution.
Rather than specifying the unknown parameters in the prior distribution ex ante,
the \ac{EB} estimator replaces them with consistent estimates, just as in random
effects models. This approach is attractive because one does not need to assume
that the effects are in fact normally distributed, or even take a ``Bayesian''
or ``random effects'' view: the \ac{EB} estimators have lower
\ac{MSE} (averaged across units) than the unshrunk unbiased estimators, even
when the individual effects are treated as nonrandom \citep{JaSt61}.
In spite of the popularity of \ac{EB} methods, it is currently not known how to
provide uncertainty assessments to accompany the point estimates without
imposing strong parametric assumptions on the effect distribution. Indeed,
\citet[p. 116]{Hansen2016} describes inference in shrinkage settings as an open
problem in econometrics. The natural \ac{EB} version of a \ac{CI} takes the form
of a Bayesian credible interval, again using the postulated effect distribution
as a prior \citep{morris83}. If the distribution is correctly specified, this
\emph{parametric \acf{EBCI}} will cover 95\%, say, of the true effect
parameters, under repeated sampling of the observed data \emph{and} of the
effect parameters. We refer to this notion of coverage as ``\ac{EB} coverage'',
following the terminology in \citet{morris83}. Unfortunately, we show that, in
the context of a normal means model, the parametric \ac{EBCI} with nominal level
95\% can have actual \ac{EB} coverage as low as 74\% for certain non-normal
effect distributions. The potential undercoverage is increasing in the degree of
shrinkage, and we derive a simple ``rule of thumb'' for gauging the potential
coverage distortion.
To allow easy uncertainty assessment in \ac{EB} applications that is reliable
irrespective of the degree of shrinkage, we construct novel \emph{robust
\acp{EBCI}} that take a simple form and control \ac{EB} coverage
\emph{regardless} of the true effect distribution. Our baseline model is an
(approximate) normal means problem $Y_i \sim N(\theta_{i}, \sigma_i^2)$,
$i=1, \dotsc, n$. In applications, $Y_i$ represents a preliminary estimate of
the effect $\theta_i$ for unit $i$. Like the parametric \ac{EBCI} that assumes a
normal distribution for $\theta_i$, the robust \ac{EBCI} we propose is centered
at the normality-based \ac{EB} point estimate $\hat{\theta}_i$ that shrinks
$Y_{i}$ toward some baseline value, but it uses a larger critical value to
account for bias due to shrinkage.\footnote{\label{fn:software} Our methods are implemented in the Stata package \texttt{ebreg}, R package \texttt{ebci}, and Matlab package \texttt{ebci\_matlab}, which are available at SSC, CRAN, and GitHub, respectively.} \ac{EB} coverage is controlled in the class of all
distributions for $\theta_{i}$ that satisfy certain moment bounds, which we
estimate consistently from the data (similarly to the parametric \ac{EBCI},
which uses the second moment). We show that the baseline implementation of our
robust \ac{EBCI} is ``adaptive'': its length is close to that
of the parametric \ac{EBCI} when the $\theta_{i}$'s are in fact normally
distributed. Thus, little efficiency is lost from using the robust \ac{EBCI} in
place of the non-robust parametric one.\footnote{\label{fn:nonnormal} If the
$\theta_{i}$'s are not normally distributed, our robust \acp{EBCI} are valid but may leave room for greater efficiency improvement, as we discuss in \Cref{sec:compar_optim}.}
In addition to controlling \ac{EB} coverage, the robust \acp{EBCI} with level
$1-\alpha$ have a frequentist \emph{average coverage} property: If the means
$\theta_{1}, \dotsc, \theta_{n}$ are treated as \emph{fixed}, the coverage
probability, averaged across the $n$ parameters $\theta_i$, is at least
$1-\alpha$. In fact, under mild conditions, at least a fraction $1-\alpha$ of the $n$ \acp{EBCI} will contain their respective parameters (with high probability as $n \to\infty$). This weakening of the usual requirement of coverage for \emph{each}
parameter $\theta_i$ allows our robust \ac{EBCI} to be shorter than the usual
\ac{CI} centered at the unshrunk estimate $Y_{i}$, and often substantially
so.\footnote{\label{fn:impossibility}Relaxing the usual notion of coverage in some way is necessary to
obtain intervals that reflect the efficiency improvement of the empirical
Bayes approach. In particular, the results in \citet{pratt61} imply that for
\acp{CI} with coverage 95\%, one cannot achieve expected length improvements
greater than 15\% relative to the usual unshrunk \acp{CI}, even if one happens
to optimize length for the true parameter vector
$(\theta_{1}, \dotsc, \theta_{n})$. See, for example, Corollary 3.3 in
\citet{armstrong_optimal_2018} and the discussion following it.} Intuitively,
the average coverage criterion only requires us to guard against the
\emph{average} coverage distortion induced by the biases of the individual
shrinkage estimators $\hat{\theta}_i$, and the data is quite informative about
whether \emph{most} of these biases are large, even though individual biases are
difficult to estimate. To complement the frequentist properties, our \acp{EBCI}
can be viewed as Bayesian credible sets that are robust to the prior on
$\theta_i$, in terms of \emph{ex ante} coverage.
The average coverage criterion has the same motivation as the usual frequentist
justification of the \ac{EB} \emph{point estimator}: the \ac{EB} point estimator
achieves lower \ac{MSE} on average across units at the expense of potentially
worse performance for some individual units \citep[see, for example,][Ch.
1.3]{Efron2010}. Thus, researchers who use \ac{EB} estimators instead of the
unshrunk $Y_{i}$'s prioritize favorable group performance over protecting
individual performance. Our average coverage intervals make an analogous
tradeoff: they guarantee coverage and achieve short length on average across
units at the expense of giving up on a coverage guarantee for every individual
unit. We examine this tradeoff in more detail in \Cref{sec:compar}.
We caution, however, that the average coverage criterion is typically
inappropriate in applications where shrinkage point estimation is unattractive. This includes settings where one is interested in the magnitude or
the identity of the largest $\theta_{i}$, or the true effect for the largest
observed $Y_{i}$ (as in, for example, \citealp{HuFi19}, or
\citealp{Andrews2021}).\footnote{\label{fn:cutoff_diverge}As we show in
\Cref{sec:cover_selection}, our methods do extend to settings where we keep a
subset of units $i$ that exceed a given cutoff. However, we do not allow this
cutoff to diverge with the sample size, such as when one focuses on the unit $i$ with the single largest
observed $Y_{i}$.} It also includes settings where a particular effect, say
$\theta_{1}$, is of primary interest, or, more generally, settings where the
effects are not exchangeable, and their ordering is relevant \citep{GrRi19}. Our methods are also not applicable if one is
interested in functionals of the random effects \emph{distribution} (as in
\citealp{bonhomme2020}, or \citealp{ignatiadis2019}), rather than in the effects
themselves.
Finally, the justification for our methods is asymptotic in the number of
parameters $n$. In our Monte Carlo simulations, we find that our \acp{EBCI} have
close to nominal coverage over a range of \acp{DGP} once $n$ is greater than
$100$.
We illustrate our results by computing \acp{EBCI} for the causal effects of
growing up in different U.S. neighborhoods (specifically commuting zones) on
intergenerational mobility. We follow \citet{chetty_impacts_2018}, who apply
\ac{EB} shrinkage to initial fixed effects estimates. Depending on the
specification, we find that the robust \acp{EBCI} are on average 12--25\% as
long as the unshrunk \acp{CI}.
Our underlying ideas extend to other linear and non-linear shrinkage settings
with possibly non-Gaussian data. For example, our techniques allow for the
construction of robust \acp{EBCI} that contain (nonlinear) soft thresholding
estimators, as well as average coverage confidence bands for nonparametric
regression functions.
The average coverage criterion was originally introduced in the literature on
nonparametric regression (\citealp{Wahba1983}; \citealp{Nychka1988};
\citealp[Ch. 5.8]{Wasserman2006}). \citet{Cai2015} construct adaptive average
coverage confidence bands. These procedures are challenging to implement in our
\ac{EB} setting, and lack a clear finite-sample justification, unlike our
procedure. \citet{Liu2019} construct forecast intervals in a dynamic panel data
model that guarantee average coverage in a Bayesian sense (for a fixed prior).
We discuss alternative approaches to inference in \ac{EB} settings in
\Cref{sec:compar}.
The rest of this paper is organized as follows. \Cref{sec:simple-example}
illustrates our methods in the context of a simple homoskedastic Gaussian model.
\Cref{sec:pract-impl} presents our recommended baseline procedure and discusses
practical implementation issues. \Cref{sec:general-results} presents our main
results on the coverage and efficiency of the robust \ac{EBCI}, and on the
coverage distortions of the parametric \ac{EBCI}; we also verify the
finite-sample coverage accuracy of the robust \ac{EBCI} through extensive
simulations. \Cref{sec:compar} compares our \ac{EBCI} with other inference
approaches. \Cref{sec:extensions} discusses extensions of the basic framework.
\Cref{sec:empir-appl} contains an empirical application to inference on
neighborhood effects.
\Cref{sec:moment_estimates,sec:comput,coverage_results_sec_append} give details
on finite-sample corrections, computational details, and formal asymptotic
coverage results. The Online Supplement contains proofs as well as further
technical results. Applied readers are encouraged to focus on
\Cref{sec:simple-example,sec:pract-impl,sec:empir-appl}.
\section{Simple example}\label{sec:simple-example}
This section illustrates the construction of the robust \acp{EBCI} that we
propose in a simplified setting with no covariates and with known, homoskedastic
errors. \Cref{sec:pract-impl} relaxes these restrictions, and discusses other
empirically relevant extensions of the basic framework, as well as
implementation issues.
We observe $n$ estimates $Y_{i}$ of elements of the parameter vector
$\theta=(\theta_{1}, \dotsc, \theta_{n})'$. Each estimate is normally
distributed with common, known variance $\sigma^{2}$,
\begin{equation}\label{eq:homoskedastic-normal-means}
Y_{i}\mid \theta \sim N(\theta_{i}, \sigma^{2}), \qquad i=1, \dotsc, n.
\end{equation}
In many applications, the $Y_{i}$'s arise as preliminary least squares estimates
of the parameters $\theta_{i}$. For instance, they may correspond to fixed
effect estimates of teacher or school value added, neighborhood effects, or firm
and worker effects. In such cases, $Y_{i}$ will only be \emph{approximately}
normal in large samples by the \ac{CLT}; we take this explicitly into account in
the theory in \Cref{coverage_results_sec_append}.
A popular approach to estimation that substantially improves upon the raw
estimator $\hat{\theta}_{i}=Y_{i}$ under the compound \ac{MSE}
$\sum_{i=1}^{n}E[(\hat{\theta}_{i}-\theta_{i})^{2}]$ is based on \acf{EB}
shrinkage. In particular, suppose that the $\theta_{i}$'s are themselves
normally distributed,
\begin{equation}\label{eq:normal-distr}
\theta_{i}\sim N(0, \mu_{2}).
\end{equation}
Our discussion below applies if \Cref{eq:normal-distr} is viewed as a subjective
Bayesian prior distribution for a single parameter $\theta_i$, but for
concreteness we will think of \Cref{eq:normal-distr} as a ``random effects''
sampling distribution for the $n$ mean parameters $\theta_1, \dotsc, \theta_n$.
Under \Cref{eq:normal-distr}, it is optimal to estimate $\theta_i$ using the
posterior mean $\hat{\theta}_i=w_{EB}Y_i$, where
$w_{EB}=1-\sigma^{2}/(\sigma^{2}+\mu_{2})$. To avoid having to specify the
variance $\mu_{2}$, the \ac{EB} approach treats it as an unknown parameter, and
replaces the marginal precision of $Y_{i}$, $1/(\sigma^{2}+\mu_{2})$, with a
method of moments estimate $n/\sum_{i=1}^{n}Y_{i}^{2}$, or the
degrees-of-freedom adjusted estimate $(n-2)/\sum_{i=1}^{n}Y_{i}^{2}$. The latter
leads to the classic estimator of \citet{JaSt61},
$\hat{w}_{EB}=1-\sigma^2(n-2)/\sum_{i=1}^{n}Y_{i}^{2}$.
One can also use~\Cref{eq:normal-distr} to construct \acp{CI} for the
$\theta_{i}$'s. In particular, since the marginal distribution of
$w_{EB}Y_{i}-\theta_{i}$ is normal with mean zero and variance
$(1-w_{EB})^{2}\mu_{2}+w_{EB}^{2}\sigma^{2}=w_{EB}\sigma^{2}$, this leads to the
$1-\alpha$ \ac{CI}
\begin{equation}\label{eq:parametric-ebci}
w_{EB}Y_{i}\pm z_{1-\alpha/2}w_{EB}^{1/2}\sigma,
\end{equation}
where $z_{\alpha}$ is the $\alpha$ quantile of the standard normal distribution.
Since the form of the interval is motivated by the parametric
assumption~\eqref{eq:normal-distr}, we refer to it as a parametric \ac{EBCI}.
With $\mu_{2}$ unknown, one can replace $w_{EB}$ by
$\hat{w}_{EB}$.\footnote{Alternatively, to account for estimation error in
$\hat{w}_{EB}$, \citet{morris83} suggests adjusting the variance estimate
$\hat{w}_{EB}\sigma^{2}$ to
$\hat{w}_{EB}\sigma^{2}+2Y_{i}^{2}(1-\hat{w}_{EB})^{2}/(n-2)$. The adjustment
does not matter asymptotically.} This is asymptotically equivalent
to~\eqref{eq:parametric-ebci} as $n\to\infty$.
The coverage of the parametric \ac{EBCI} in~\eqref{eq:parametric-ebci} is
$1-\alpha$ under repeated sampling of $(Y_{i}, \theta_{i})$ according to
\Cref{eq:homoskedastic-normal-means,eq:normal-distr}. To distinguish this notion
of coverage from the case with fixed $\theta$, we refer to coverage under
repeated sampling of $(Y_{i}, \theta_{i})$ as ``empirical Bayes coverage''. This
follows the definition of an \acf{EBCI} in \citet[Eq.\ 3.6]{morris83} and
\citet[Ch. 3.5]{CaLo00}. Unfortunately, this coverage property relies heavily on
the parametric assumption~\eqref{eq:normal-distr}. We show in
\Cref{sec:param-ebci-cover} that the actual \ac{EB} coverage of the nominal
$1-\alpha$ parametric \ac{EBCI} can be as low as $1-1/\max\{z_{1-\alpha/2},1\}$
for certain non-normal distributions of $\theta_{i}$ with variance $\mu_{2}$;
for 95\% \acp{EBCI}, this evaluates to 74\%. This contrasts with existing
results on estimation: although the \ac{EB} estimator is motivated by the
parametric assumption~\eqref{eq:normal-distr}, it performs well even if this
assumption is dropped, with low \ac{MSE} even if we treat $\theta$ as fixed.
This paper constructs an \ac{EBCI} with a similar robustness property: the
interval will be close in length to the parametric \ac{EBCI} when
\Cref{eq:normal-distr} holds, but its \ac{EB} coverage is at least $1-\alpha$
without any parametric assumptions on the distribution of $\theta_{i}$. To
describe the construction, suppose that all that is known is that $\theta_{i}$
is sampled from a distribution with second moment given by $\mu_{2}$ (in
practice, we can replace $\mu_2$ by the consistent estimate
$n^{-1}\sum_{i=1}^{n}Y^{2}_{i}-\sigma^{2}$). Conditional on $\theta_{i}$, the
estimator $w_{EB}Y_{i}$ has bias $(w_{EB}-1)\theta_{i}$ and variance
$w_{EB}^{2}\sigma^{2}$, so that the $t$-statistic
$(w_{EB}Y_{i}-\theta_{i})/w_{EB}\sigma$ is normally distributed with mean
$b_{i}=(1-1/w_{EB})\theta_{i}/\sigma$ and variance $1$. Therefore, if we use a
critical value $\chi$, the non-coverage of the \ac{CI}
$w_{EB}Y_{i}\pm \chi w_{EB}\sigma$, conditional on $\theta_{i}$, will be given
by the probability
\begin{equation}\label{eq:rb}
r(b_{i}, \chi)=P(\abs{Z-b_{i}}\geq \chi\mid
\theta_{i})=\Phi(-\chi-b_{i})+\Phi(-\chi+b_{i}),
\end{equation}
where $Z$ denotes a standard normal random variable, and $\Phi$ denotes its cdf. Thus, by iterated expectations, under repeated sampling of
$\theta_{i}$, the non-coverage is bounded by
\begin{equation}\label{eq:non-coverage-bound}
\rho(\sigma^{2}/\mu_{2}, \chi)=\sup_{F}E_{F}[r(b, \chi)]\quad
\text{s.t.}\quad E_{F}[b^{2}]=\frac{(1-1/w_{EB})^{2}}{\sigma^{2}}\mu_{2}
=\frac{\sigma^{2}}{\mu_{2}},
\end{equation}
where $E_{F}$ denotes expectation under $b\sim F$. Although this is an
infinite-dimensional optimization problem over the space of distributions, it
turns out that it admits a simple closed-form solution, which we give in
\Cref{theorem:bound-second-moment} in \Cref{sec:comput}. Moreover, because the
optimization is a linear program, it can be solved even in the more general
settings of applied relevance that we consider in \Cref{sec:pract-impl}.
Set $\chi=\operatorname{cva}_{\alpha}(\sigma^{2}/\mu_{2})$, where
$\operatorname{cva}_{\alpha}(t)=\rho^{-1}(t, \alpha)$, and the inverse is with respect to
the second argument. Then the resulting interval
\begin{equation}\label{eq:robust-ebci}
w_{EB}Y_{i}\pm \operatorname{cva}_{\alpha}(\sigma^{2}/\mu_{2})w_{EB}\sigma
\end{equation}
will maintain coverage $1-\alpha$ among all distributions of $\theta_{i}$ with
$E[\theta_{i}^{2}]=\mu_{2}$ (recall that we estimate $\mu_{2}$ consistently from
the data). For this reason, we refer to it as a robust \ac{EBCI}.
\Cref{fig:cva05} in \Cref{sec:baseline-model} gives a plot of the critical values for $\alpha=0.05$. We show
in \Cref{sec:efficiency} below that by also imposing a constraint on the fourth
moment of $\theta_{i}$, in addition to the second moment constraint, one can
construct a robust \ac{EBCI} that ``adapts'' to the Gaussian case in the sense
that its
length will be close to that of the parametric \ac{EBCI} in \Cref{eq:parametric-ebci} if these moment constraints are compatible with a normal distribution.
Instead of considering \ac{EB} coverage, one may alternatively wish to assess
uncertainty associated with the estimates $\hat{\theta}_i=w_{EB}Y_{i}$ when $\theta$ is treated
as fixed. In this case, the \ac{EBCI} in \Cref{eq:robust-ebci} has an average
coverage guarantee that
\begin{equation} \label{eq:simple-aci} \frac{1}{n}\sum_{i=1}^{n} P\big(\theta_i
\in [w_{EB}Y_{i}\pm \operatorname{cva}_{\alpha}(\sigma^{2}/\mu_{2})w_{EB}\sigma] \;\big|\;
\theta \big)\geq 1-\alpha,
\end{equation}
provided that the moment constraint can be interpreted as a constraint on the
empirical second moment on the $\theta_{i}$'s,
$n^{-1}\sum_{i=1}^{n}\theta_{i}^{2}=\mu_{2}$. In other words, if we condition on
$\theta$, then the coverage is at least $1-\alpha$ on average across the $n$
\acp{EBCI} for $\theta_{1}, \dotsc, \theta_{n}$. To see this, note that the
average non-coverage of the intervals is bounded
by~\eqref{eq:non-coverage-bound}, except that the supremum is only taken over
possible empirical distributions for $\theta_{1}, \dotsc, \theta_{n}$ satisfying
the moment constraint. Since this supremum is necessarily smaller than
$\rho(\sigma^{2}/\mu_{2}, \chi)$, it follows that the average coverage is at
least $1-\alpha$. In fact, if the $Y_i$'s exhibit limited dependence across $i$, a stronger property holds: the probability that at least a fraction $1-\alpha$ of the $n$ \acp{EBCI} contain their respective true parameters converges to 1 as $n \to \infty$, cf.\ \Cref{rem:ac_alt_def_remark} below.
The usual \acp{CI} $Y_{i}\pm z_{1-\alpha/2}\sigma$ also of course achieve
average coverage $1-\alpha$. The robust \ac{EBCI} in \Cref{eq:robust-ebci} will,
however, be shorter, especially when $\mu_{2}$ is small relative to
$\sigma^{2}$---see \Cref{fig:efficiency_unshrunk} below. The reduction in length
is achieved by weakening the requirement that each \ac{CI} covers its true
parameter $1-\alpha$ percent of the time to the requirement that the coverage probability equal
$1-\alpha$ on average across the \acp{CI}.
It may seem surprising that we can construct a narrower \ac{CI} by centering it
at the shrinkage estimates $w_{EB}Y_{i}$. The intuition for this is that the
shrinkage reduces the variability of the estimates, at the expense of
introducing bias in the estimates. The bias necessitates the use of a larger
critical value $\operatorname{cva}_{\alpha}(\sigma^{2}/\mu_{2})$. Because under the average
coverage criterion we only need to control the bias \emph{on average} across
$i$, rather than for each individual $\theta_{i}$, this increase in the critical
value is smaller than the reduction in the standard error.
\section{Practical implementation}\label{sec:pract-impl}
We now describe how to compute a robust \ac{EBCI} that allows for
heteroskedasticity, shrinks towards more general regression estimates rather
than towards zero, and exploits higher moments of the bias to yield a narrower
interval. In \Cref{sec:baseline-model}, we describe the empirical Bayes model
that motivates our baseline approach. \Cref{sec:baseline-implementation}
describes the practical implementation of our baseline approach.
\subsection{Motivating model and robust EBCI}\label{sec:baseline-model}
In applied settings, the unshrunk estimates $Y_i$ will typically have
heteroskedastic variances. Furthermore, rather than shrinking towards zero, it
is common to shrink toward an estimate of $\theta_i$ based on some covariates
$X_i$, such as a regression estimate $X_i'\hat{\delta}$. We now describe how to
adapt the ideas in \Cref{sec:simple-example} to such settings.
Consider a generalization of the model in \Cref{eq:homoskedastic-normal-means}
that allows for heteroskedasticity and covariates,
\begin{equation} \label{eq:hierarch_y}
Y_{i}\mid \theta_{i}, X_{i}, \sigma_{i} \sim N(\theta_{i},
\sigma_{i}^2), \qquad i=1,\dotsc, n.
\end{equation}
The covariate vector $X_i$ may contain just the intercept, and it may also
contain (functions of) $\sigma_{i}$. To construct an \ac{EB} estimator of
$\theta_{i}$, consider the working assumption that the sampling distribution of
the $\theta_i$'s is conditionally normal:
\begin{equation} \label{eq:hierarch_theta} \theta_i\mid X_{i}, \sigma_{i} \sim
N(\mu_{1,i}, \mu_{2}), \quad \text{where} \quad \mu_{1,i}=X_i'\delta.
\end{equation}
The hierarchical model~\eqref{eq:hierarch_y}--\eqref{eq:hierarch_theta} leads to
the Bayes estimate $\hat{\theta}_{i}=\mu_{1, i}+w_{EB, i}(Y_i-\mu_{1, i})$,
where $w_{EB, i}=\frac{\mu_{2}}{\mu_{2}+\sigma_{i}^{2}}$. This estimate shrinks
the unrestricted estimate $Y_{i}$ of $\theta_{i}$ toward
$\mu_{1,i}=X_{i}'\delta$. In contrast to~\eqref{eq:hierarch_y}, the normality
assumption~\eqref{eq:hierarch_theta} typically cannot be justified simply by
appealing to the \ac{CLT}; the linearity of the conditional mean
$\mu_{1,i}=X_i'\delta$ may also be suspect. Our robust \ac{EBCI} will therefore
be constructed so that it achieves valid \ac{EB} coverage even if
assumption~\eqref{eq:hierarch_theta} fails. To obtain a narrow robust \ac{EBCI},
we augment the second moment restriction used to compute the critical value in
\Cref{eq:non-coverage-bound} with restrictions on higher moments of the bias of
$\hat{\theta}_{i}$. In our baseline specification, we add a restriction on the
fourth moment.
In particular, we replace assumption~\eqref{eq:hierarch_theta} with the much
weaker requirement that the conditional second moment and kurtosis of
$\varepsilon_i=\theta_{i}-X_{i}'\delta$ do not depend on $(X_{i}, \sigma_{i})$:
\begin{align}\label{eq:moment_independence}
E[(\theta_i-X_i'\delta)^{2} \mid X_{i}, \sigma_i]&=\mu_{2}, &
E[(\theta_i-X_i'\delta)^{4} \mid X_{i}, \sigma_i]/\mu^{2}_{2}&=\kappa,
\end{align}
where $\delta$ is defined as the probability limit of the regression estimate
$\hat{\delta}$.\footnote{Our framework can be modified to let $(X_i, \sigma_i)$
be fixed, in which case $\delta$ depends on $n$. See the discussion following
\Cref{thm:coverage_baseline} below.} We discuss this requirement further in
\Cref{rem:conditional_coverage} below, and we relax it in \Cref{rem:nonparam}
below.
We now apply analysis analogous to that in \Cref{sec:simple-example}. Let us
suppose for simplicity that $\delta$, $\mu_{2}$, $\kappa$, and $\sigma_{i}$ are
known; we relax this assumption in \Cref{sec:baseline-implementation} below, and
in the theory in \Cref{sec:general-results}. Denote the conditional bias of
$\hat{\theta}_i$ normalized by the standard error by
$b_{i} = (w_{EB, i}-1)\varepsilon_i/(w_{EB, i}\sigma_i) =
-\sigma_{i}\varepsilon_{i}/\mu_{2}$. Under repeated sampling of $\theta_i$, the
non-coverage of the \ac{CI} $\hat{\theta}_i \pm \chi w_{EB, i}\sigma$,
conditional on $(X_{i}, \sigma_{i})$, depends on the distribution of the
normalized bias $b_{i}$, as in \Cref{sec:simple-example}. Given the moments
$\mu_{2}$ and $\kappa$, the \emph{maximal} non-coverage is given by
\begin{equation}\label{eq:fourth_moment_bound}
\rho(m_{2,i}, \kappa, \chi)=\sup_{F}E_{F}[r(b, \chi)]\quad
\text{s.t.}\quad E_{F}[b^{2}]=m_{2,i}, \, E_{F}[b^{4}]=\kappa m_{2,i}^2,
\end{equation}
where $b$ is distributed according to the distribution $F$. Here
$m_{2,i} = E[b_{i}^{2} \mid X_{i}, \sigma_{i}] =\sigma^{2}_{i}/\mu_{2}$. Observe
that the kurtosis of $b_{i}$ matches that of $\varepsilon_{i}$.
\Cref{sec:comput} shows that the infinite-dimensional linear
program~\eqref{eq:fourth_moment_bound} can be reduced to two nested
\emph{univariate} optimizations. We also show that the least favorable
distribution---the distribution $F$
maximizing~\eqref{eq:fourth_moment_bound}---is a discrete distribution with up
to 4 support points (see \Cref{remark:lf-distro}).
Define the critical value
$\operatorname{cva}_{\alpha}(m_{2,i}, \kappa)=\rho^{-1}(m_{2,i}, \kappa, \alpha)$, where the
inverse is in the last argument. \Cref{fig:cva05} plots this function for
$\alpha=0.05$ and selected values of $\kappa$. This leads to the robust
\ac{EBCI}
\begin{equation}\label{eq:conditional_ci}
\hat{\theta}_{i} \pm \operatorname{cva}_{\alpha}(m_{2, i}, \kappa)w_{EB, i}\sigma_{i},
\end{equation}
which, by construction, has coverage at least $1-\alpha$ under repeated sampling
of $(Y_i, \theta_i)$, conditional on $(X_{i}, \sigma_{i})$, so long as
\Cref{eq:moment_independence} holds; it is not required
that~\eqref{eq:hierarch_theta} holds. Note that both the critical value and the
\ac{CI} length are increasing in $\sigma_{i}$.
\begin{figure}[t]
\centering{
\begin{tikzpicture}[x=1pt,y=1pt]
\definecolor{fillColor}{RGB}{255,255,255}
\path[use as bounding box,fill=fillColor,fill opacity=0.00] (0,0) rectangle (411.94,231.26);
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{255,255,255}
\definecolor{fillColor}{RGB}{255,255,255}
\path[draw=drawColor,line width= 0.6pt,line join=round,line cap=round,fill=fillColor] ( 0.00, 0.00) rectangle (411.94,231.26);
\end{scope}
\begin{scope}
\path[clip] ( 37.55, 30.40) rectangle (406.94,226.26);
\definecolor{fillColor}{RGB}{255,255,255}
\path[fill=fillColor] ( 37.55, 30.40) rectangle (406.94,226.26);
\definecolor{drawColor}{RGB}{208,208,208}
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55, 30.40) --
(406.94, 30.40);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55, 67.31) --
(406.94, 67.31);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55,102.79) --
(406.94,102.79);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55,138.28) --
(406.94,138.28);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55,173.77) --
(406.94,173.77);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55,209.26) --
(406.94,209.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55, 30.40) --
( 37.55,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (111.58, 30.40) --
(111.58,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (185.60, 30.40) --
(185.60,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (259.63, 30.40) --
(259.63,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (333.65, 30.40) --
(333.65,226.26);
\definecolor{drawColor}{RGB}{81,81,81}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 37.55, 30.40) --
( 70.45, 47.92) --
(103.35, 70.18) --
(136.25, 95.88) --
(169.15,120.34) --
(202.05,142.62) --
(234.95,163.10) --
(267.85,182.15) --
(300.75,200.03) --
(333.65,216.94);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 2pt off 2pt ,line join=round] ( 37.55, 30.40) --
( 70.45, 43.35) --
(103.35, 52.74) --
(136.25, 60.21) --
(169.15, 66.53) --
(202.05, 72.12) --
(234.95, 77.17) --
(267.85, 81.81) --
(300.75, 86.13) --
(333.65, 90.19);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 2pt ,line join=round] ( 37.55, 30.40) --
( 70.45, 44.69) --
(103.35, 57.58) --
(136.25, 69.29) --
(169.15, 79.82) --
(202.05, 89.51) --
(234.95, 98.77) --
(267.85,107.71) --
(300.75,116.37) --
(333.65,124.78);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55, 30.40) --
( 70.45, 44.44) --
(103.35, 56.44) --
(136.25, 67.09) --
(169.15, 76.77) --
(202.05, 85.70) --
(234.95, 94.03) --
(267.85,101.87) --
(300.75,109.30) --
(333.65,116.37);
\end{scope}
\begin{scope}
\path[clip] ( 37.55, 30.40) rectangle (406.94,226.26);
\definecolor{drawColor}{RGB}{0,0,0}
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 1.00] at (333.65, 86.75) {$\mathrm{cva}_{0.05}(m_{2}, 1)$};
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 1.00] at (333.65,112.93) {$\mathrm{cva}_{P, 0.05}(m_{2})$};
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 1.00] at (333.65,122.32) {$\mathrm{cva}_{0.05}(m_{2}, 3)$};
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 1.00] at (333.65,213.49) {$\mathrm{cva}_{0.05}(m_{2}, \infty)$};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 37.55, 30.40) --
( 37.55,226.26);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05, 27.64) {1.96};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05, 64.55) {3};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05,100.04) {4};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05,135.53) {5};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05,171.02) {6};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05,206.50) {7};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05, 30.40) --
( 37.55, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05, 67.31) --
( 37.55, 67.31);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05,102.79) --
( 37.55,102.79);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05,138.28) --
( 37.55,138.28);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05,173.77) --
( 37.55,173.77);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05,209.26) --
( 37.55,209.26);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 37.55, 30.40) --
(406.94, 30.40);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 37.55, 27.90) --
( 37.55, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (111.58, 27.90) --
(111.58, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (185.60, 27.90) --
(185.60, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (259.63, 27.90) --
(259.63, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (333.65, 27.90) --
(333.65, 30.40);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 37.55, 20.39) {0};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (111.58, 20.39) {1};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (185.60, 20.39) {2};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (259.63, 20.39) {3};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (333.65, 20.39) {4};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (222.24, 6.94) {$m_{2}$};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 11.89,128.33) {Critical value};
\end{scope}
\end{tikzpicture}
}
\caption{Function $\operatorname{cva}_{\alpha}(m_{2}, \kappa)$ for $\alpha=0.05$ and
selected values of $\kappa$. The function $\operatorname{cva}_{\alpha}(m_{2})$, defined in
\Cref{sec:simple-example}, that only imposes a constraint on the second
moment, corresponds to $\operatorname{cva}_{\alpha}(m_{2}, \infty)$. The function
$\operatorname{cva}_{P, \alpha}(m_{2})=z_{1-\alpha/2}\sqrt{1+m_{2}}$ corresponds to the
critical value under the assumption that $\theta_{i}$ is normally
distributed.}\label{fig:cva05}
\end{figure}
\subsection{Baseline implementation}\label{sec:baseline-implementation}
Our baseline implementation of the robust \ac{EBCI} plugs in consistent
estimates of the unknown quantities in \Cref{eq:conditional_ci}, based on the
data $\{Y_{i}, X_{i}, \hat{\sigma}_{i}\}_{i=1}^{n}$, where $\hat{\sigma}_{i}$ is a consistent estimate of $\sigma_{i}$ (such as the standard error of the
preliminary estimate $Y_{i}$), and $X_{i}$ is a vector of covariates that are
thought to help predict $\theta_{i}$.
\begin{enumerate}
\item\label{item:param-estimates} Regress $Y_i$ on $X_i$ to obtain the fitted
values $X_i'\hat\delta$, with
$\hat\delta=(\sum_{i=1}^{n}\omega_i X_{i}X_{i}')^{-1}\sum_{i=1}^n\omega_{i}
X_{i}Y_{i}$ denoting the weighted least squares estimate with precision
weights $\omega_i$. Two natural choices are setting
$\omega_{i}=\hat\sigma_i^{-2}$, or setting $\omega_{i}=1/n$ for unweighted
estimates; see \Cref{sec:weighting} for further discussion. Let
$\hat\mu_{2}=\max\left\{\frac{\sum_{i=1}^n
\omega_i(\hat{\varepsilon}_i^2-\hat\sigma_i^2)}{\sum_{i=1}^n\omega_i},
\frac{2\sum_{i=1}^n\omega_i^2\hat\sigma_i^4}{\sum_{i=1}^n\omega_i \cdot
\sum_{i=1}^n \omega_i\hat\sigma_i^{2}} \right\}$, and
$\hat{\kappa}=\max\left\{\frac{\sum_{i=1}^{n}
\omega_i(\hat{\varepsilon}_{i}^{4}-6\hat{\sigma}_i^2\hat{\varepsilon}_{i}^{2}
+3\hat{\sigma}_i^4)}{\hat{\mu}_{2}^{2}\sum_{i=1}^n\omega_i}, 1 + \frac{32
\sum_{i=1}^n
\omega_i^2\hat\sigma_i^{8}}{\hat\mu_2^2\sum_{i=1}^n\omega_i\cdot
\sum_{i=1}^n\omega_i\hat\sigma_i^4} \right\}$, where
$\hat\varepsilon_i=Y_i-X_i'\hat\delta$.
\item\label{item:1} Form the \ac{EB} estimate
\begin{equation*}
\hat\theta_i= X_i'\hat\delta + \hat{w}_{EB, i}(Y_i-X_i'\hat\delta),
\quad \text{where} \quad
\hat{w}_{EB, i} = \frac{\hat\mu_{2}}{\hat{\mu}_{2} + \hat{\sigma}_{i}^{2}}.
\end{equation*}
\item Compute the critical value
$\operatorname{cva}_{\alpha}(\hat\sigma_i^{2}/\hat\mu_{2}, \hat{\kappa})$ defined
below~\Cref{eq:fourth_moment_bound}.
\item Report the robust \ac{EBCI}
\begin{equation}\label{eq:ebci_under_indep}
\hat\theta_i \pm \operatorname{cva}_{\alpha}(\hat\sigma_i^{2}/\hat{\mu}_{2},
\hat{\kappa})\hat{w}_{EB, i}\hat\sigma_{i}.
\end{equation}
\end{enumerate}
We provide fast and stable software packages that automate these
steps (see \cref{fn:software}). We now discuss the assumptions needed for
validity of the robust \ac{EBCI}.
\begin{remark}[Conditional \ac{EB} coverage and moment
independence]\label{rem:conditional_coverage}
A potential concern about \ac{EB} coverage in a heteroskedastic setting is
that in order to reduce the length of the \ac{CI} on average, one could choose to overcover
parameters $\theta_{i}$ with small $\sigma_{i}$ and undercover parameters
$\theta_{i}$ with large $\sigma_{i}$. Our robust \ac{EBCI} ensures that this
does not happen by requiring \ac{EB} coverage to hold conditional on
$(X_{i}, \sigma_{i})$. This also avoids analogous coverage concerns as a
result of the value of $X_{i}$.
The key to ensuring this property is assumption~\eqref{eq:moment_independence}
that the conditional second moment and kurtosis of
$\varepsilon_{i}=\theta_{i}-X_{i}'\delta$ do not depend on
$(X_{i}, \sigma_{i})$. Conditional moment independence assumptions of this
form are common in the literature. For instance, it is imposed in the analysis
of neighborhood effects in \citet{chetty_impacts_2018} (their approach
requires independence of the second moment), which is the basis for our
empirical application in \Cref{sec:empir-appl}. Nonetheless, such conditions
may be strong in some settings, as argued by \citet{xie_sure_2012} in the
context of \ac{EB} point estimation. In \Cref{rem:nonparam} below, we drop
condition~\eqref{eq:moment_independence} entirely by replacing $\hat\mu_{2}$
and $\hat\kappa$ with nonparametric estimates of these conditional moments;
alternatively, one could relax it by using a flexible parametric
specification.\footnote{\label{fn:t-stat_shrinkage}Another way to drop
condition~\eqref{eq:moment_independence} is to base shrinkage on the
$t$-statistics $Y_i/\sigma_i$, applying the baseline implementation above
with $Y_i/\hat\sigma_i$ in place of $Y_i$ and 1 in place of
$\hat{\sigma}_i$. Then the homoskedastic analysis in
\Cref{sec:simple-example} applies, leading to valid \acp{EBCI} without any
assumptions about independence of the moments. See Remark 3.8 and Appendix
D.1 in \citet{akp20v2} for further discussion.}
\end{remark}
\begin{remark}[Nonparametric moment estimates]\label{rem:nonparam}
As a robustness check to guard against failure of the moment independence
assumption~\eqref{eq:moment_independence}, one may replace the critical value
in~\Cref{eq:ebci_under_indep} with
$\operatorname{cva}_{\alpha}((1-1/\hat{w}_{EB, i})^{2}\hat{\mu}_{2i}/\hat{\sigma}^{2}_{i},
\hat{\kappa}_{i})$, where $\hat{\mu}_{2i}$ and $\hat{\kappa}_{i}$ are
consistent nonparametric estimates of
$\mu_{2i}=E[(\theta_{i}-X_{i}'\delta)^{2}\mid X_{i}, \sigma_{i}]$ and
$\kappa_{i}=E[(\theta_{i}-X_{i}'\delta)^{4}\mid X_{i},
\sigma_{i}]/\mu_{2i}^{2}$. The resulting \ac{CI} will be asymptotically
equivalent to the \ac{CI} in the baseline implementation if
\Cref{eq:moment_independence} holds, but it will achieve valid \ac{EB} coverage
even if this assumption fails. In our empirical application, we use
nearest-neighbor estimates, as described in \Cref{sec:finite_n}. As a simple
diagnostic to gauge how much the second moment of $\theta_i-X_i'\delta$ varies with
$(X_{i}, \sigma_{i})$, one can report the $R^{2}$ gain in predicting
$\hat{\varepsilon}_{i}^{2}-\hat{\sigma}^{2}_{i}$ using $\hat{\mu}_{2i}$ rather
than the baseline estimate $\hat{\mu}_{2}$, as we illustrate in our empirical
application.
\end{remark}
\begin{remark}[Average coverage and non-independent
sampling]\label{rem:average_coverage_independence}
We show in \Cref{sec:general-results} that the robust \ac{EBCI} satisfies an
average coverage criterion of the form~\eqref{eq:simple-aci} when the
parameters $\theta=(\theta_{1}, \dotsc, \theta_{n}$) are considered fixed, in
addition to achieving valid \ac{EB} coverage when the $\theta_i$'s are viewed
as random draws from some underlying distribution. To guarantee average
coverage or \ac{EB} coverage, we do not need to assume that the $Y_i$'s and
$\theta_i$'s are drawn independently across $i$. This is because the average
coverage and \ac{EB} coverage criteria only depend on the marginal
distribution of $(Y_{i}, \theta_{i})$, not the joint distribution. Indeed, in
deriving the infeasible \ac{CI} in \Cref{eq:conditional_ci}, we made no
assumptions about the dependence structure of $(Y_{i}, \theta_{i})$ across
$i$. Consequently, to guarantee asymptotic coverage of the feasible interval
in \Cref{eq:ebci_under_indep} as $n\to\infty$, we only need to ensure that the
estimates $\hat{\mu}_2, \hat{\kappa}, \hat{\delta}, \hat{\sigma}_i$
are consistent for $\mu_{2}, \kappa, \delta, \sigma_i$, which is the case
under many forms of weak dependence or clustering. Furthermore, our baseline
implementation above does not require the researcher to take an explicit stand
on the dependence of the data; for example, in the case of clustering, the
researcher does not need to take an explicit stand on how the clusters are
defined.
\end{remark}
\begin{remark}[Estimating moments of the distribution of
$\theta_{i}$]\label{rem:estimating_moments}
The estimators $\hat{\mu}_{2}$ and $\hat{\kappa}$ in
step~\ref{item:param-estimates} of our baseline implementation above are based
on the moment conditions
$E[(Y_{i}-X_{i}'\delta)^{2}-\sigma^{2}_{i}\mid X_{i}, \sigma_{i}]=\mu_{2}$ and
$E[(Y_{i}-X_{i}'\delta)^{4}+3\sigma^{4}_{i}-6\sigma^{2}_{i}(Y_{i}-X_{i}'\delta)^{2}
\mid X_{i}, \sigma_{i}]=\kappa \mu_{2}^{2}$, replacing population expectations
by weighted sample averages. In addition, to avoid small-sample coverage
issues when $\mu_2$ and $\kappa$ are near their theoretical lower bounds of 0
and 1, respectively, these estimates incorporate truncation on $\hat\mu_2$ and
$\hat\kappa$. These truncated estimates approximate the Bayesian posterior means under a flat prior on $\mu_2$ and $\kappa$, as in
\citet{morris83pebci,morris83}. Although the resulting \acp{EBCI} do not
directly account for estimation uncertainty in $\mu_{2}$ and $\kappa$, we
verify their small-sample coverage accuracy via extensive simulations in
\Cref{sec:sim}. \Cref{sec:finite_n} discusses the choice of the moment
estimates, as well as other ways of performing truncation.
\end{remark}
\begin{remark}[Using higher moments and other forms of shrinkage]\label{rem:higher_moments}
In addition to using the second and fourth moment of bias, one may
augment~\eqref{eq:fourth_moment_bound} with restrictions on higher moments of
the bias in order to further tighten the critical value. In
\Cref{sec:efficiency}, we show that using other moments in addition to the
second and fourth moment does not substantially decrease the critical value in
the case where $\theta_i$ is normally distributed. Thus, the CI in our
baseline implementation is robust to failure of the normality
assumption~\eqref{eq:hierarch_theta}, while being near-optimal when this
assumption does hold. \Cref{sec:efficiency} also shows that further efficiency
gains are possible if one uses the linear estimator
$\tilde{\theta}_{i}=\mu_{1,i}+w_{i}(Y_i-\mu_{1,i})$ with the shrinkage
coefficient $w_{i}$ chosen to optimize \ac{CI} length, instead of using the
\ac{MSE}-optimal shrinkage $w_{EB, i}$. For efficiency under a non-normal
distribution of $\theta_{i}$, one needs to consider non-linear shrinkage; we
discuss this extension in \Cref{sec:general_shrinkage}.
\end{remark}
\section{Main results}\label{sec:general-results}
This section provides formal statements of the coverage properties of the
\acp{CI} presented in \Cref{sec:simple-example,sec:pract-impl}. Furthermore, we
show that the \acp{CI} presented in \Cref{sec:simple-example,sec:pract-impl} are
highly efficient when the mean parameters are in fact normally distributed.
Next, we calculate the maximal coverage distortion of the parametric \ac{EBCI},
and derive a rule of thumb for gauging the potential coverage distortion.
Finally, we present a comprehensive simulation study of the finite-sample
performance of the robust \ac{EBCI}. Applied readers interested primarily in
implementation issues may skip ahead to the empirical application in
\Cref{sec:empir-appl}.
\subsection{Coverage under baseline implementation}\label{sec:cover-under-basel}
In order to state the formal result, let us first carefully define the notions
of coverage that we consider. Consider intervals $CI_{1}, \dotsc, CI_{n}$ for
elements of the parameter vector $\theta=(\theta_1, \dotsc, \theta_n)'$. The
probability measure $P$ denotes the joint distribution of $\theta$ and
$CI_1, \dotsc, CI_n$. Following \citet[Eq.\ 3.6]{morris83} and \citet[Ch.
3.5]{CaLo00}, we say that the interval $CI_i$ is an (asymptotic) $1-\alpha$
\acf{EBCI} if
\begin{equation}\label{eq:ebci_def}
\liminf_{n\to\infty} P(\theta_i\in CI_i) \ge 1-\alpha.
\end{equation}
We say that the intervals $CI_i$ are (asymptotic) $1-\alpha$ \acp{ACI} under the
parameter sequence $\theta_1, \dotsc, \theta_n$ if
\begin{equation}\label{eq:aci_def}
\liminf_{n\to\infty} \frac{1}{n}\sum_{i=1}^n P(\theta_i\in CI_i\mid \theta) \ge 1-\alpha.
\end{equation}
The average coverage property~\eqref{eq:aci_def} is a property of the
distribution of the data conditional on $\theta$ and therefore does not require
that we view the $\theta_i$'s as random (as in a Bayesian or ``random effects''
analysis). To maintain consistent notation, we nonetheless use the
conditional notation $P(\cdot\mid \theta)$ when considering average coverage.
See \Cref{coverage_results_sec_append} for a formulation with
$\theta$ treated as nonrandom.
Observe that under the exchangeability condition that
$P(\theta_i\in CI_i)=P(\theta_j\in CI_j)$ for all $i, j$, if the \ac{ACI}
property~\eqref{eq:aci_def} holds almost surely, then the \ac{EBCI}
property~\eqref{eq:ebci_def} holds, since then
\begin{equation*}
P(\theta_{i}\in CI_{i})= \frac{1}{n}\sum_{j=1}^n P(\theta_{j}\in CI_{j})
\ge 1-\alpha+o(1)\quad\text{for all $i$.}
\end{equation*}
We now provide coverage results for the baseline implementation described in
\Cref{sec:baseline-implementation}. To keep the statements in the main text as
simple as possible, we (i) maintain the assumption that the unshrunk estimates
$Y_i$ follow an exact normal distribution conditional on the parameter
$\theta_i$, (ii) state the results only for the homoskedastic case where the
variance $\sigma_i$ of the unshrunk estimate $Y_i$ does not vary across $i$, and
(iii) consider only unconditional coverage statements of the
form~\eqref{eq:ebci_def} and~\eqref{eq:aci_def}. In
\Cref{coverage_results_sec_append}, we allow the estimates $Y_i$ to be only
approximately normally distributed and allow $\sigma_i$ to vary, and we verify
that our assumptions hold in a linear fixed effects panel data model. We also
formalize the statements about conditional coverage made in
\Cref{rem:conditional_coverage}.
\begin{theorem}\label{thm:coverage_baseline}
Suppose $Y_i\mid \theta\sim N(\theta_i, \sigma^2)$. Let
$\mu_{j, n}=\frac{1}{n}\sum_{i=1}^n (\theta_i-X_i'\delta)^j$ and let
$\kappa_n=\mu_{4,n}/\mu_{2,n}^2$. Suppose the sequence
$\theta=\theta_1, \dotsc, \theta_n$ and the conditional distribution
$P(\cdot\mid \theta)$ satisfy the following conditions with probability one:
\begin{enumerate}
\item $\mu_{2,n}\to \mu_{2}$ and $\mu_{4,n}/\mu_{2,n}^2\to \kappa$ for some
$\mu_{2}\in (0,\infty)$ and $\kappa\in (1,\infty)$.
\item Conditional on $\theta$,
$(\hat\delta, \hat\sigma, \hat\mu_{2}, \hat\kappa)$ converges in probability
to $(\delta, \sigma, \mu_{2}, \kappa)$.
\end{enumerate}
Then the \acp{CI} in \Cref{eq:ebci_under_indep} with $\hat\sigma_i=\hat\sigma$
satisfy the \ac{ACI} property~\eqref{eq:aci_def} with probability one.
Furthermore, if $\theta_1, \dotsc, \theta_n$ follow an exchangeable
distribution and the estimators $\hat\delta$, $\hat\sigma$, $\hat\mu_{2}$ and
$\hat\kappa$ are exchangeable functions of the data
$(X_1', Y_1)', \dotsc, (X_n', Y_n)'$, then these \acp{CI} satisfy the \ac{EB}
coverage property~\eqref{eq:ebci_def}.
\end{theorem}
\Cref{thm:coverage_baseline} follows immediately from
\Cref{indep_shrinkage_coverage_thm} in \Cref{coverage_results_sec_append}. In
order to cover both the \ac{EB} coverage condition~\eqref{eq:ebci_def} and the
average coverage condition~\eqref{eq:aci_def}, \Cref{thm:coverage_baseline}
considers a random sequence of parameters $\theta_1, \dotsc, \theta_n$, and
shows average coverage conditional on these parameters. See
\Cref{coverage_results_sec_append} for a formulation with $\theta$ treated as
nonrandom.
The condition on the moments $\mu_2$ and $\kappa$ avoids degenerate cases such
as when $\mu_{2}=0$, in which case the \ac{EB} point estimator $\hat{\theta}_i$
shrinks each preliminary estimate $Y_i$ all the way to $X_i'\hat{\delta}$. Note
also that the theorem does not require that $\hat{\delta}$ be the \ac{OLS}
estimate in a regression of $Y_{i}$ onto $X_{i}$, and that $\delta$ be the
population analog; one can define $\delta$ in other ways, the theorem only
requires that $\hat{\delta}$ be a consistent estimate of it. The definition of
$\delta$ does, however, affect the plausibility of the moment independence
assumption in \Cref{eq:moment_independence} needed for conditional coverage
results stated in
\Cref{coverage_results_sec_append}.\footnote{\label{fn:effect_on_width}The
specification of $\mu_{1i}=X_{i}'\delta$ also affects the \ac{EBCI} width
through its effect on $\mu_{2}$ and $\kappa$.}
\begin{remark}\label{rem:ac_alt_def_remark}
As shown in \Cref{coverage_results_sec_append}, if \acp{CI} satisfy the average coverage condition~\eqref{eq:aci_def}
given $\theta_{1}, \dotsc, \theta_{n}$, they will typically also satisfy the stronger
condition
\begin{equation}\label{eq:alt_aci_def}
\frac{1}{n}\sum_{i=1}^n \1{\theta_i\in CI_i} \ge 1-\alpha + o_{P(\cdot\mid \theta)}(1),
\end{equation}
where $o_{P(\cdot\mid \theta)}(1)$ denotes a sequence that converges in
probability to zero conditional on $\theta$ (\Cref{eq:alt_aci_def} implies
\Cref{eq:aci_def} since the left-hand side is uniformly bounded). That is, at
least a fraction $1-\alpha$ of the $n$ \acp{CI} contain their respective true
parameters, asymptotically. This is analogous to the result that for
estimation, the difference between the squared error
$\frac{1}{n}\sum_{i=1}^n(\hat\theta_i-\theta_i)^2$ and the \ac{MSE}
$\frac{1}{n}\sum_{i=1}^{n} E[(\hat\theta_i-\theta_i)^2\mid \theta]$ typically
converges to zero.
\end{remark}
\subsection{Relative efficiency}\label{sec:efficiency}
The robust \ac{EBCI} in \Cref{eq:conditional_ci}, unlike the parametric
\ac{EBCI} $\hat{\theta}_{i}\pm z_{1-\alpha/2}\sigma_{i}\sqrt{w_{EB, i}}$, does
not rely on the normality assumption in \Cref{eq:hierarch_theta} for its
validity. We now show that this robustness does not come at a high cost in terms
of efficiency: if the normality assumption~\eqref{eq:hierarch_theta} in fact
holds, the efficiency loss is limited unless the signal-to-noise ratio
$\mu_{2}/\sigma^{2}_{i}$ is very small.
There are two reasons for the inefficiency of the robust \ac{EBCI}. First, the
robust \ac{EBCI} only makes use of the second and fourth moment of the
conditional distribution of $\theta_{i}-X_{i}'\delta$, rather than its full
distribution. Second, if we only have knowledge of these two moments, it is no
longer optimal to center the \ac{EBCI} at the estimator $\hat{\theta}_{i}$: one
may need to consider other, perhaps non-linear, shrinkage estimators, as we do
below in \Cref{sec:general_shrinkage}.
We decompose the sources of inefficiency by studying the relative length of the
robust \ac{EBCI} relative to the \ac{EBCI} that picks the amount of shrinkage
optimally. For the latter, we maintain
assumption~\eqref{eq:moment_independence}, and consider a more general class of
estimators $\tilde{\theta}(w_{i})=\mu_{1, i}+w_{i}(Y_{i}-\mu_{1, i})$. For
tractability, we focus on fixed-length \acp{CI} based on linear
shrinkage estimators, but allow the amount of shrinkage $w_{i}$ to be optimally
determined. The normalized bias of $\tilde{\theta}(w_{i})$ is given by
$b_{i}=(1/w_{i}-1)\varepsilon_{i}/\sigma_{i}$, which leads to the \ac{EBCI}
\begin{equation*}
\mu_{1, i}+w_{i}(Y_{i}-\mu_{1, i})\pm
\operatorname{cva}_{\alpha}((1-1/w_{i})^{2}\mu_{2}/\sigma^{2}_{i}, \kappa)w_{i}\sigma_{i}.
\end{equation*}
The half-length of this \ac{EBCI},
$\operatorname{cva}_{\alpha}((1-1/w_{i})^{2}\mu_{2}/\sigma^{2}_{i}, \kappa)w_{i}\sigma_{i}$,
can be numerically minimized as a function of $w_{i}$ to find the \ac{EBCI}
length-optimal shrinkage. Denote the minimizer by
$w_{opt}(\mu_{2}/\sigma^{2}_{i}, \kappa, \alpha)$. Like $w_{EB, i}$, the optimal
shrinkage depends on $\mu_{2}$ and $\sigma_{i}^{2}$ only through the
signal-to-noise ratio $\mu_{2}/\sigma^{2}_{i}$. Numerically evaluating the
minimizer shows that $w_{opt}(\cdot, \kappa, \alpha)\geq w_{EB, i}$ for
$\kappa\geq 3$ and $\alpha\in\{0.05,0.1\}$. The resulting \ac{EBCI} is optimal
among all fixed-length \acp{EBCI} centered at linear estimators
under~\eqref{eq:moment_independence}, and we call it the optimal robust
\ac{EBCI}.\footnote{\label{fn:optimal_robust}Since the optimal robust \ac{EBCI}
is always shorter than the robust \ac{EBCI} in~\Cref{eq:conditional_ci}, the former is
preferable on efficiency grounds. It may not contain the \ac{MSE}-optimal point estimator $\hat{\theta}_i$, however.}
\begin{figure}[t]
\centering{
\begin{tikzpicture}[x=1pt,y=1pt]
\definecolor{fillColor}{RGB}{255,255,255}
\path[use as bounding box,fill=fillColor,fill opacity=0.00] (0,0) rectangle (411.94,231.26);
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{255,255,255}
\definecolor{fillColor}{RGB}{255,255,255}
\path[draw=drawColor,line width= 0.6pt,line join=round,line cap=round,fill=fillColor] ( 0.00, 0.00) rectangle (411.94,231.26);
\end{scope}
\begin{scope}
\path[clip] ( 37.55, 30.40) rectangle (406.94,226.26);
\definecolor{fillColor}{RGB}{255,255,255}
\path[fill=fillColor] ( 37.55, 30.40) rectangle (406.94,226.26);
\definecolor{drawColor}{RGB}{208,208,208}
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55, 30.40) --
(406.94, 30.40);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55, 68.21) --
(406.94, 68.21);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55,106.02) --
(406.94,106.02);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55,143.83) --
(406.94,143.83);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55,181.65) --
(406.94,181.65);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55,219.46) --
(406.94,219.46);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55, 30.40) --
( 37.55,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (129.90, 30.40) --
(129.90,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (222.24, 30.40) --
(222.24,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (314.59, 30.40) --
(314.59,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (406.94, 30.40) --
(406.94,226.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 37.92,122.99) --
( 41.64,107.13) --
( 45.37, 99.49) --
( 49.09, 94.00) --
( 52.81, 89.62) --
( 56.54, 85.95) --
( 60.26, 82.77) --
( 63.99, 79.97) --
( 67.71, 77.46) --
( 71.43, 75.19) --
( 75.16, 73.10) --
( 78.88, 71.19) --
( 82.60, 69.41) --
( 86.33, 67.75) --
( 90.05, 66.21) --
( 93.78, 64.75) --
( 97.50, 63.38) --
(101.22, 62.08) --
(104.95, 60.86) --
(108.67, 59.69) --
(112.39, 58.58) --
(116.12, 57.52) --
(119.84, 56.52) --
(123.57, 55.55) --
(127.29, 54.63) --
(131.01, 53.74) --
(134.74, 52.89) --
(138.46, 52.07) --
(142.18, 51.29) --
(145.91, 50.53) --
(149.63, 49.80) --
(153.36, 49.10) --
(157.08, 48.42) --
(160.80, 47.77) --
(164.53, 47.14) --
(168.25, 46.52) --
(171.97, 45.93) --
(175.70, 45.36) --
(179.42, 44.81) --
(183.15, 44.27) --
(186.87, 43.75) --
(190.59, 43.25) --
(194.32, 42.77) --
(198.04, 42.29) --
(201.76, 41.84) --
(205.49, 41.39) --
(209.21, 40.96) --
(212.94, 40.54) --
(216.66, 40.14) --
(220.38, 39.75) --
(224.11, 39.36) --
(227.83, 38.99) --
(231.55, 38.63) --
(235.28, 38.29) --
(239.00, 37.95) --
(242.73, 37.62) --
(246.45, 37.30) --
(250.17, 36.99) --
(253.90, 36.69) --
(257.62, 36.40) --
(261.34, 36.12) --
(265.07, 35.84) --
(268.79, 35.58) --
(272.51, 35.32) --
(276.24, 35.07) --
(279.96, 34.83) --
(283.69, 34.60) --
(287.41, 34.37) --
(291.13, 34.15) --
(294.86, 33.94) --
(298.58, 33.74) --
(302.30, 33.54) --
(306.03, 33.35) --
(309.75, 33.17) --
(313.48, 32.99) --
(317.20, 32.82) --
(320.92, 32.66) --
(324.65, 32.50) --
(328.37, 32.35) --
(332.09, 32.20) --
(335.82, 32.06) --
(339.54, 31.93) --
(343.27, 31.80) --
(346.99, 31.67) --
(350.71, 31.56) --
(354.44, 31.45) --
(358.16, 31.34) --
(361.88, 31.24) --
(365.61, 31.14) --
(369.33, 31.05) --
(373.06, 30.97) --
(376.78, 30.89) --
(380.50, 30.81) --
(384.23, 30.74) --
(387.95, 30.67) --
(391.67, 30.61) --
(395.40, 30.55) --
(399.12, 30.50) --
(402.85, 30.45) --
(406.57, 30.40);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 2pt off 2pt ,line join=round] ( 37.92, 53.69) --
( 41.64, 49.10) --
( 45.37, 46.89) --
( 49.09, 45.31) --
( 52.81, 44.07) --
( 56.54, 43.03) --
( 60.26, 42.15) --
( 63.99, 41.39) --
( 67.71, 40.71) --
( 71.43, 40.11) --
( 75.16, 39.57) --
( 78.88, 39.08) --
( 82.60, 38.64) --
( 86.33, 38.23) --
( 90.05, 37.86) --
( 93.78, 37.52) --
( 97.50, 37.21) --
(101.22, 36.92) --
(104.95, 36.66) --
(108.67, 36.41) --
(112.39, 36.19) --
(116.12, 35.98) --
(119.84, 35.78) --
(123.57, 35.60) --
(127.29, 35.44) --
(131.01, 35.28) --
(134.74, 35.14) --
(138.46, 35.01) --
(142.18, 34.89) --
(145.91, 34.77) --
(149.63, 34.66) --
(153.36, 34.53) --
(157.08, 34.40) --
(160.80, 34.26) --
(164.53, 34.12) --
(168.25, 33.97) --
(171.97, 33.83) --
(175.70, 33.69) --
(179.42, 33.55) --
(183.15, 33.41) --
(186.87, 33.28) --
(190.59, 33.14) --
(194.32, 33.01) --
(198.04, 32.89) --
(201.76, 32.76) --
(205.49, 32.64) --
(209.21, 32.53) --
(212.94, 32.41) --
(216.66, 32.30) --
(220.38, 32.20) --
(224.11, 32.09) --
(227.83, 32.00) --
(231.55, 31.90) --
(235.28, 31.81) --
(239.00, 31.72) --
(242.73, 31.63) --
(246.45, 31.55) --
(250.17, 31.47) --
(253.90, 31.40) --
(257.62, 31.33) --
(261.34, 31.26) --
(265.07, 31.20) --
(268.79, 31.13) --
(272.51, 31.07) --
(276.24, 31.02) --
(279.96, 30.97) --
(283.69, 30.92) --
(287.41, 30.87) --
(291.13, 30.83) --
(294.86, 30.79) --
(298.58, 30.75) --
(302.30, 30.71) --
(306.03, 30.68) --
(309.75, 30.65) --
(313.48, 30.62) --
(317.20, 30.59) --
(320.92, 30.57) --
(324.65, 30.55) --
(328.37, 30.53) --
(332.09, 30.51) --
(335.82, 30.49) --
(339.54, 30.48) --
(343.27, 30.46) --
(346.99, 30.45) --
(350.71, 30.44) --
(354.44, 30.43) --
(358.16, 30.43) --
(361.88, 30.42) --
(365.61, 30.41) --
(369.33, 30.41) --
(373.06, 30.41) --
(376.78, 30.40) --
(380.50, 30.40) --
(384.23, 30.40) --
(387.95, 30.40) --
(391.67, 30.40) --
(395.40, 30.40) --
(399.12, 30.40) --
(402.85, 30.40) --
(406.57, 30.40);
\definecolor{drawColor}{RGB}{81,81,81}
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 2pt ,line join=round] ( 37.92,216.94) --
( 41.64,202.21) --
( 45.37,194.32) --
( 49.09,188.11) --
( 52.81,182.78) --
( 56.54,177.98) --
( 60.26,173.58) --
( 63.99,169.45) --
( 67.71,165.56) --
( 71.43,161.84) --
( 75.16,158.27) --
( 78.88,154.83) --
( 82.60,151.50) --
( 86.33,148.26) --
( 90.05,145.11) --
( 93.78,142.03) --
( 97.50,139.01) --
(101.22,136.05) --
(104.95,133.14) --
(108.67,130.28) --
(112.39,127.46) --
(116.12,124.69) --
(119.84,121.95) --
(123.57,119.24) --
(127.29,116.57) --
(131.01,113.92) --
(134.74,111.31) --
(138.46,108.71) --
(142.18,106.14) --
(145.91,103.60) --
(149.63,101.07) --
(153.36, 98.56) --
(157.08, 96.08) --
(160.80, 93.61) --
(164.53, 91.16) --
(168.25, 88.72) --
(171.97, 86.31) --
(175.70, 83.92) --
(179.42, 81.54) --
(183.15, 79.19) --
(186.87, 76.86) --
(190.59, 74.57) --
(194.32, 72.31) --
(198.04, 70.08) --
(201.76, 67.90) --
(205.49, 65.77) --
(209.21, 63.70) --
(212.94, 61.70) --
(216.66, 59.76) --
(220.38, 57.89) --
(224.11, 56.11) --
(227.83, 54.41) --
(231.55, 52.79) --
(235.28, 51.26) --
(239.00, 49.81) --
(242.73, 48.45) --
(246.45, 47.17) --
(250.17, 45.97) --
(253.90, 44.84) --
(257.62, 43.79) --
(261.34, 42.81) --
(265.07, 41.90) --
(268.79, 41.04) --
(272.51, 40.24) --
(276.24, 39.50) --
(279.96, 38.81) --
(283.69, 38.17) --
(287.41, 37.57) --
(291.13, 37.01) --
(294.86, 36.49) --
(298.58, 36.00) --
(302.30, 35.55) --
(306.03, 35.13) --
(309.75, 34.74) --
(313.48, 34.38) --
(317.20, 34.04) --
(320.92, 33.72) --
(324.65, 33.43) --
(328.37, 33.15) --
(332.09, 32.90) --
(335.82, 32.66) --
(339.54, 32.44) --
(343.27, 32.24) --
(346.99, 32.05) --
(350.71, 31.87) --
(354.44, 31.71) --
(358.16, 31.56) --
(361.88, 31.42) --
(365.61, 31.29) --
(369.33, 31.17) --
(373.06, 31.06) --
(376.78, 30.96) --
(380.50, 30.87) --
(384.23, 30.78) --
(387.95, 30.70) --
(391.67, 30.63) --
(395.40, 30.56) --
(399.12, 30.50) --
(402.85, 30.45) --
(406.57, 30.40);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.92, 79.65) --
( 41.64, 69.33) --
( 45.37, 64.53) --
( 49.09, 61.05) --
( 52.81, 58.25) --
( 56.54, 55.88) --
( 60.26, 53.80) --
( 63.99, 51.95) --
( 67.71, 50.29) --
( 71.43, 48.77) --
( 75.16, 47.38) --
( 78.88, 46.11) --
( 82.60, 44.93) --
( 86.33, 43.84) --
( 90.05, 42.85) --
( 93.78, 41.93) --
( 97.50, 41.09) --
(101.22, 40.33) --
(104.95, 39.63) --
(108.67, 39.00) --
(112.39, 38.43) --
(116.12, 37.92) --
(119.84, 37.45) --
(123.57, 37.04) --
(127.29, 36.66) --
(131.01, 36.33) --
(134.74, 36.03) --
(138.46, 35.77) --
(142.18, 35.53) --
(145.91, 35.32) --
(149.63, 35.13) --
(153.36, 34.96) --
(157.08, 34.81) --
(160.80, 34.68) --
(164.53, 34.56) --
(168.25, 34.45) --
(171.97, 34.34) --
(175.70, 34.23) --
(179.42, 34.11) --
(183.15, 33.98) --
(186.87, 33.85) --
(190.59, 33.72) --
(194.32, 33.58) --
(198.04, 33.45) --
(201.76, 33.31) --
(205.49, 33.17) --
(209.21, 33.04) --
(212.94, 32.90) --
(216.66, 32.77) --
(220.38, 32.64) --
(224.11, 32.51) --
(227.83, 32.39) --
(231.55, 32.26) --
(235.28, 32.15) --
(239.00, 32.03) --
(242.73, 31.92) --
(246.45, 31.82) --
(250.17, 31.71) --
(253.90, 31.62) --
(257.62, 31.52) --
(261.34, 31.44) --
(265.07, 31.35) --
(268.79, 31.27) --
(272.51, 31.20) --
(276.24, 31.13) --
(279.96, 31.06) --
(283.69, 31.00) --
(287.41, 30.94) --
(291.13, 30.89) --
(294.86, 30.84) --
(298.58, 30.79) --
(302.30, 30.75) --
(306.03, 30.71) --
(309.75, 30.67) --
(313.48, 30.64) --
(317.20, 30.61) --
(320.92, 30.58) --
(324.65, 30.56) --
(328.37, 30.53) --
(332.09, 30.51) --
(335.82, 30.50) --
(339.54, 30.48) --
(343.27, 30.47) --
(346.99, 30.45) --
(350.71, 30.44) --
(354.44, 30.43) --
(358.16, 30.43) --
(361.88, 30.42) --
(365.61, 30.41) --
(369.33, 30.41) --
(373.06, 30.41) --
(376.78, 30.40) --
(380.50, 30.40) --
(384.23, 30.40) --
(387.95, 30.40) --
(391.67, 30.40) --
(395.40, 30.40) --
(399.12, 30.40) --
(402.85, 30.40) --
(406.57, 30.40);
\definecolor{drawColor}{RGB}{144,144,144}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 65.78, 50.24) {Opt, $\kappa=3$};
\definecolor{drawColor}{RGB}{81,81,81}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 66.13, 76.20) {Rob, $\kappa=3$};
\definecolor{drawColor}{RGB}{144,144,144}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 68.27,119.55) {Opt, $\kappa=\infty$};
\definecolor{drawColor}{RGB}{81,81,81}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 68.59,213.49) {Rob, $\kappa=\infty$};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 37.55, 30.40) --
( 37.55,226.26);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05, 27.64) {1.00};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05, 65.45) {1.25};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05,103.27) {1.50};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05,141.08) {1.75};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05,178.89) {2.00};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05,216.70) {2.25};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05, 30.40) --
( 37.55, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05, 68.21) --
( 37.55, 68.21);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05,106.02) --
( 37.55,106.02);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05,143.83) --
( 37.55,143.83);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05,181.65) --
( 37.55,181.65);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05,219.46) --
( 37.55,219.46);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 37.55, 30.40) --
(406.94, 30.40);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 37.55, 27.90) --
( 37.55, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (129.90, 27.90) --
(129.90, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (222.24, 27.90) --
(222.24, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (314.59, 27.90) --
(314.59, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (406.94, 27.90) --
(406.94, 30.40);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 37.55, 20.39) {0};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (129.90, 20.39) {0.25};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (222.24, 20.39) {0.5};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (314.59, 20.39) {0.75};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (406.94, 20.39) {1};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (222.24, 6.94) {$w_{EB, i}=\mu_{2}/(\mu_{2}+\sigma^2_i)$};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 11.89,128.33) {Relative length};
\end{scope}
\end{tikzpicture}
}
\caption{Relative efficiency of robust \ac{EBCI} (Rob) and optimal robust
\ac{EBCI} (Opt) relative to the normal benchmark, for $\alpha=0.05$. The
figure plots ratios of Rob length,
$2\operatorname{cva}_{\alpha}(\sigma_{i}^{2}/\mu_{2}, \kappa) \cdot \sigma_{i}
{\mu_{2}/(\mu_{2}+\sigma_{i}^{2})}$, and Opt length,
$2\operatorname{cva}_{\alpha}((1-1/w_{opt}(\mu_{2}/\sigma_{i}^{2}, \kappa,
\alpha))^{2}\mu_{2}/\sigma_{i}^{2}, \kappa)\cdot \sigma_{i}
w_{opt}(\mu_{2}/\sigma_{i}^{2}, \kappa, \alpha)$, relative to the parametric
\ac{EBCI} length
$2z_{1-\alpha/2}\sqrt{\mu_{2}/(\mu_{2}+\sigma_{i}^{2})}\sigma_{i}$ as a
function of the shrinkage factor
$w_{EB, i}=\mu_{2}/(\mu_{2}+\sigma_{i}^{2})$, which maps the signal-to-noise
ratio $\mu_{2}/\sigma_{i}^{2}$ to the interval
$[0,1]$.}\label{fig:ebci_efficiency}
\end{figure}
\Cref{fig:ebci_efficiency} plots the ratio of lengths of the optimal robust
\ac{EBCI} and robust \ac{EBCI} relative to the parametric \ac{EBCI}, for
$\alpha=0.05$. The figure shows that to maintain efficiency relative to the
normal benchmark, it is important to impose the fourth moment constraint. If
this constraint is imposed, the efficiency loss of the robust \ac{EBCI} is
modest unless the signal-to-noise ratio is very small: if $w_{EB, i}\geq 0.1$
(which is equivalent to $\mu_{2}/\sigma^{2}_{i}\geq 1/9$), the efficiency loss
is at most $11.4\%$ for $\alpha=0.05$; up to half of the efficiency loss is due
to not using the optimal shrinkage. For $\alpha=0.1$ (not plotted), the results
are very similar; in particular, if $w_{EB, i}\geq 0.1$, the efficiency loss is
at most $12.9\%$.
\begin{figure}[t]
\centering{
\begin{tikzpicture}[x=1pt,y=1pt]
\definecolor{fillColor}{RGB}{255,255,255}
\path[use as bounding box,fill=fillColor,fill opacity=0.00] (0,0) rectangle (411.94,231.26);
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{255,255,255}
\definecolor{fillColor}{RGB}{255,255,255}
\path[draw=drawColor,line width= 0.6pt,line join=round,line cap=round,fill=fillColor] ( 0.00, 0.00) rectangle (411.94,231.26);
\end{scope}
\begin{scope}
\path[clip] ( 37.55, 30.40) rectangle (406.94,226.26);
\definecolor{fillColor}{RGB}{255,255,255}
\path[fill=fillColor] ( 37.55, 30.40) rectangle (406.94,226.26);
\definecolor{drawColor}{RGB}{208,208,208}
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55, 30.40) --
(406.94, 30.40);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55, 79.36) --
(406.94, 79.36);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55,128.33) --
(406.94,128.33);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55,177.30) --
(406.94,177.30);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55,226.26) --
(406.94,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55, 30.40) --
( 37.55,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (129.90, 30.40) --
(129.90,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (222.24, 30.40) --
(222.24,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (314.59, 30.40) --
(314.59,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (406.94, 30.40) --
(406.94,226.26);
\definecolor{drawColor}{RGB}{81,81,81}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 37.92, 44.23) --
( 41.64, 74.44) --
( 45.37, 89.77) --
( 49.09,101.12) --
( 52.81,110.33) --
( 56.54,118.14) --
( 60.26,124.94) --
( 63.99,130.97) --
( 67.71,136.38) --
( 71.43,141.27) --
( 75.16,145.73) --
( 78.88,149.82) --
( 82.60,153.57) --
( 86.33,157.04) --
( 90.05,160.24) --
( 93.78,163.21) --
( 97.50,165.97) --
(101.22,168.52) --
(104.95,170.89) --
(108.67,173.10) --
(112.39,175.15) --
(116.12,177.04) --
(119.84,178.80) --
(123.57,180.43) --
(127.29,181.94) --
(131.01,183.33) --
(134.74,184.61) --
(138.46,185.78) --
(142.18,186.85) --
(145.91,187.82) --
(149.63,188.70) --
(153.36,189.49) --
(157.08,190.20) --
(160.80,190.82) --
(164.53,191.37) --
(168.25,191.84) --
(171.97,192.23) --
(175.70,192.56) --
(179.42,192.83) --
(183.15,193.04) --
(186.87,193.19) --
(190.59,193.29) --
(194.32,193.35) --
(198.04,193.38) --
(201.76,193.38) --
(205.49,193.35) --
(209.21,193.32) --
(212.94,193.29) --
(216.66,193.26) --
(220.38,193.25) --
(224.11,193.26) --
(227.83,193.29) --
(231.55,193.36) --
(235.28,193.46) --
(239.00,193.61) --
(242.73,193.80) --
(246.45,194.03) --
(250.17,194.30) --
(253.90,194.61) --
(257.62,194.97) --
(261.34,195.37) --
(265.07,195.80) --
(268.79,196.28) --
(272.51,196.78) --
(276.24,197.32) --
(279.96,197.90) --
(283.69,198.50) --
(287.41,199.12) --
(291.13,199.78) --
(294.86,200.45) --
(298.58,201.15) --
(302.30,201.87) --
(306.03,202.61) --
(309.75,203.36) --
(313.48,204.13) --
(317.20,204.92) --
(320.92,205.72) --
(324.65,206.53) --
(328.37,207.36) --
(332.09,208.19) --
(335.82,209.04) --
(339.54,209.89) --
(343.27,210.76) --
(346.99,211.63) --
(350.71,212.50) --
(354.44,213.39) --
(358.16,214.28) --
(361.88,215.17) --
(365.61,216.07) --
(369.33,216.98) --
(373.06,217.88) --
(376.78,218.80) --
(380.50,219.71) --
(384.23,220.63) --
(387.95,221.55) --
(391.67,222.47) --
(395.40,223.39) --
(399.12,224.32) --
(402.85,225.25) --
(406.57,226.17);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 37.92, 41.96) --
( 41.64, 66.74) --
( 45.37, 79.09) --
( 49.09, 88.13) --
( 52.81, 95.42) --
( 56.54,101.56) --
( 60.26,106.88) --
( 63.99,111.57) --
( 67.71,115.76) --
( 71.43,119.54) --
( 75.16,122.97) --
( 78.88,126.11) --
( 82.60,128.99) --
( 86.33,131.64) --
( 90.05,134.08) --
( 93.78,136.35) --
( 97.50,138.45) --
(101.22,140.40) --
(104.95,142.22) --
(108.67,143.90) --
(112.39,145.47) --
(116.12,146.94) --
(119.84,148.30) --
(123.57,149.58) --
(127.29,150.77) --
(131.01,151.88) --
(134.74,152.93) --
(138.46,153.91) --
(142.18,154.85) --
(145.91,155.74) --
(149.63,156.59) --
(153.36,157.41) --
(157.08,158.22) --
(160.80,159.01) --
(164.53,159.79) --
(168.25,160.58) --
(171.97,161.37) --
(175.70,162.17) --
(179.42,162.98) --
(183.15,163.80) --
(186.87,164.64) --
(190.59,165.50) --
(194.32,166.38) --
(198.04,167.27) --
(201.76,168.18) --
(205.49,169.11) --
(209.21,170.06) --
(212.94,171.02) --
(216.66,172.00) --
(220.38,172.99) --
(224.11,174.00) --
(227.83,175.01) --
(231.55,176.04) --
(235.28,177.08) --
(239.00,178.13) --
(242.73,179.18) --
(246.45,180.25) --
(250.17,181.32) --
(253.90,182.39) --
(257.62,183.47) --
(261.34,184.56) --
(265.07,185.65) --
(268.79,186.74) --
(272.51,187.84) --
(276.24,188.94) --
(279.96,190.04) --
(283.69,191.14) --
(287.41,192.24) --
(291.13,193.34) --
(294.86,194.44) --
(298.58,195.55) --
(302.30,196.65) --
(306.03,197.75) --
(309.75,198.85) --
(313.48,199.95) --
(317.20,201.04) --
(320.92,202.14) --
(324.65,203.23) --
(328.37,204.32) --
(332.09,205.40) --
(335.82,206.49) --
(339.54,207.56) --
(343.27,208.64) --
(346.99,209.71) --
(350.71,210.78) --
(354.44,211.84) --
(358.16,212.90) --
(361.88,213.96) --
(365.61,215.01) --
(369.33,216.05) --
(373.06,217.09) --
(376.78,218.12) --
(380.50,219.14) --
(384.23,220.16) --
(387.95,221.17) --
(391.67,222.18) --
(395.40,223.18) --
(399.12,224.18) --
(402.85,225.18) --
(406.57,226.17);
\path[] (124.20,146.94) -- (117.36,146.94);
\definecolor{drawColor}{RGB}{81,81,81}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (141.97,173.60) {$\alpha=0.05$};
\definecolor{drawColor}{RGB}{144,144,144}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (141.97,143.49) {$\alpha=0.1$};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 37.55, 30.40) --
( 37.55,226.26);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05, 27.64) {0.00};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05, 76.61) {0.25};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05,125.58) {0.50};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05,174.54) {0.75};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05,223.51) {1.00};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05, 30.40) --
( 37.55, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05, 79.36) --
( 37.55, 79.36);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05,128.33) --
( 37.55,128.33);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05,177.30) --
( 37.55,177.30);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05,226.26) --
( 37.55,226.26);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 37.55, 30.40) --
(406.94, 30.40);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 37.55, 27.90) --
( 37.55, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (129.90, 27.90) --
(129.90, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (222.24, 27.90) --
(222.24, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (314.59, 27.90) --
(314.59, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (406.94, 27.90) --
(406.94, 30.40);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 37.55, 20.39) {0};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (129.90, 20.39) {0.25};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (222.24, 20.39) {0.5};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (314.59, 20.39) {0.75};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (406.94, 20.39) {1};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (222.24, 6.94) {$w_{EB, i}=\mu_{2}/(\mu_{2}+\sigma^2_i)$};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 11.89,128.33) {Relative length};
\end{scope}
\end{tikzpicture}
}
\caption{Efficiency of robust \ac{EBCI}
$\hat{\theta}_{i}\pm \operatorname{cva}_{\alpha}(\sigma_i^{2}/\mu_{2}, \kappa=\infty)
\cdot \sigma {\mu_{2}/(\mu_{2}+\sigma_i^{2})}$ relative to the unshrunk
\ac{CI} $Y_{i}\pm z_{1-\alpha/2}\sigma_i$. The figure plots the ratio of the
length of the robust \ac{EBCI} relative to the unshrunk \ac{CI} as a
function of the shrinkage factor
$w_{EB, i} =
\mu_{2}/(\mu_{2}+\sigma_{i}^{2})$.}\label{fig:efficiency_unshrunk}
\end{figure}
When the signal-to-noise ratio is very small, so that $w_{EB, i}<0.1$, the
efficiency loss of the robust \ac{EBCI} is higher (up to 39\% for $\alpha=0.05$
or $0.1$). Using the optimal robust \ac{EBCI} ensures that the efficiency loss
is below 20\%, irrespective of the signal-to-noise ratio. On the other hand,
when the signal-to-noise ratio is small, any of these \acp{CI} will be
significantly tighter than the unshrunk \ac{CI}
$Y_{i}\pm z_{1-\alpha/2}\sigma_{i}$. To illustrate this point,
\Cref{fig:efficiency_unshrunk} plots the efficiency of the robust \ac{EBCI} that
imposes the second moment constraint only, relative to this unshrunk \ac{CI}. It
can be seen from the figure that shrinkage methods allow us to tighten the
\ac{CI} by 44\% or more when $\mu_{2}/\sigma^{2}_{i}\leq 0.1$.
\subsection{Undercoverage of parametric EBCI}\label{sec:param-ebci-cover}
The parametric \ac{EBCI}
$\hat\theta_i \pm z_{1-\alpha/2}w_{EB, i}^{1/2}\sigma_{i}$ is an
\ac{EB} version of a Bayesian credible interval that
treats~\eqref{eq:hierarch_theta} as a prior. We now assess its potential
undercoverage when \Cref{eq:hierarch_theta} is violated.
Given knowledge of only the second moment $\mu_{2}$ of
$\varepsilon_i=Y_i-X_i'\delta$, the maximal undercoverage of this interval is
given by
\begin{equation} \label{eq:eb_param_noncov}
\rho(1/w_{EB, i}-1, z_{1-\alpha/2}/\sqrt{w_{EB, i}}),
\end{equation}
since $w_{EB, i}= \mu_{2}/(\mu_{2}+\sigma_i^2)$. Here $\rho$ is the non-coverage
function defined in \Cref{eq:non-coverage-bound}. \Cref{fig:noncov_param} plots
the maximal non-coverage probability as a function of $w_{EB, i}$, for
significance levels $\alpha=0.05$ and $\alpha=0.10$. The figure suggests a
simple ``rule of thumb'': if $w_{EB, i} \geq 0.3$, the maximal coverage
distortion is less than 5 percentage points for these values of $\alpha$.
The following lemma confirms that the maximal non-coverage is decreasing in
$w_{EB, i}$, as suggested by the figure. It also gives an expression for the
maximal non-coverage across all values of $w_{EB, i}$ (which is achieved in the
limit $w_{EB, i} \to 0$).
\begin{lemma}\label{thm:eb_param}
The non-coverage probability~\eqref{eq:eb_param_noncov} of the parametric \ac{EBCI}
is weakly decreasing as a
function of $w_{EB, i}$, with the supremum given by $1/\max\{z_{1-\alpha}^2,1\}$.
\end{lemma}
The maximal non-coverage probability $1/\max\{z_{1-\alpha/2}^2,1\}$ equals
$0.260$ for $\alpha=0.05$ and $0.370$ for $\alpha=0.10$. For
$\alpha>2\Phi(-1)\approx 0.317$, the maximal non-coverage probability is 1.
\begin{figure}[tp]
\centering{
\begin{tikzpicture}[x=1pt,y=1pt]
\definecolor{fillColor}{RGB}{255,255,255}
\path[use as bounding box,fill=fillColor,fill opacity=0.00] (0,0) rectangle (411.94,231.26);
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{255,255,255}
\definecolor{fillColor}{RGB}{255,255,255}
\path[draw=drawColor,line width= 0.6pt,line join=round,line cap=round,fill=fillColor] ( 0.00, 0.00) rectangle (411.94,231.26);
\end{scope}
\begin{scope}
\path[clip] ( 37.55, 30.40) rectangle (406.94,226.26);
\definecolor{fillColor}{RGB}{255,255,255}
\path[fill=fillColor] ( 37.55, 30.40) rectangle (406.94,226.26);
\definecolor{drawColor}{RGB}{208,208,208}
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55, 39.30) --
(406.94, 39.30);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55, 68.35) --
(406.94, 68.35);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55, 97.39) --
(406.94, 97.39);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55,126.44) --
(406.94,126.44);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55,155.49) --
(406.94,155.49);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55,184.53) --
(406.94,184.53);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55,213.58) --
(406.94,213.58);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 37.55, 30.40) --
( 37.55,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (129.90, 30.40) --
(129.90,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (222.24, 30.40) --
(222.24,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (314.59, 30.40) --
(314.59,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (406.94, 30.40) --
(406.94,226.26);
\definecolor{drawColor}{RGB}{81,81,81}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 37.59,156.87) --
( 37.73,152.28) --
( 37.92,149.21) --
( 39.77,136.56) --
( 41.63,130.27) --
( 43.48,125.76) --
( 45.34,122.16) --
( 47.19,119.13) --
( 49.05,116.48) --
( 50.90,114.13) --
( 52.75,112.01) --
( 54.61,110.06) --
( 56.46,108.25) --
( 58.32,106.57) --
( 60.17,105.00) --
( 62.03,103.51) --
( 63.88,102.10) --
( 65.73,100.76) --
( 67.59, 99.49) --
( 69.44, 98.26) --
( 71.30, 97.09) --
( 73.15, 95.96) --
( 75.01, 94.87) --
( 76.86, 93.82) --
( 78.72, 92.80) --
( 80.57, 91.81) --
( 82.42, 90.85) --
( 84.28, 89.92) --
( 86.13, 89.01) --
( 87.99, 88.13) --
( 89.84, 87.27) --
( 91.70, 86.43) --
( 93.55, 85.60) --
( 95.40, 84.80) --
( 97.26, 84.01) --
( 99.11, 83.24) --
(100.97, 82.49) --
(102.82, 81.75) --
(104.68, 81.02) --
(106.53, 80.30) --
(108.39, 79.60) --
(110.24, 78.91) --
(112.09, 78.23) --
(113.95, 77.57) --
(115.80, 76.91) --
(117.66, 76.26) --
(119.51, 75.63) --
(121.37, 75.00) --
(123.22, 74.38) --
(125.07, 73.77) --
(126.93, 73.17) --
(128.78, 72.58) --
(130.64, 72.00) --
(132.49, 71.42) --
(134.35, 70.86) --
(136.20, 70.30) --
(138.06, 69.74) --
(139.91, 69.20) --
(141.76, 68.66) --
(143.62, 68.13) --
(145.47, 67.61) --
(147.33, 67.09) --
(149.18, 66.58) --
(151.04, 66.07) --
(152.89, 65.58) --
(154.74, 65.09) --
(156.60, 64.60) --
(158.45, 64.13) --
(160.31, 63.66) --
(162.16, 63.19) --
(164.02, 62.73) --
(165.87, 62.28) --
(167.73, 61.83) --
(169.58, 61.39) --
(171.43, 60.96) --
(173.29, 60.53) --
(175.14, 60.11) --
(177.00, 59.69) --
(178.85, 59.28) --
(180.71, 58.88) --
(182.56, 58.48) --
(184.41, 58.08) --
(186.27, 57.70) --
(188.12, 57.31) --
(189.98, 56.94) --
(191.83, 56.57) --
(193.69, 56.20) --
(195.54, 55.84) --
(197.40, 55.49) --
(199.25, 55.14) --
(201.10, 54.79) --
(202.96, 54.45) --
(204.81, 54.12) --
(206.67, 53.79) --
(208.52, 53.47) --
(210.38, 53.15) --
(212.23, 52.84) --
(214.08, 52.53) --
(215.94, 52.23) --
(217.79, 51.93) --
(219.65, 51.63) --
(221.50, 51.35) --
(223.36, 51.06) --
(225.21, 50.78) --
(227.07, 50.51) --
(228.92, 50.24) --
(230.77, 49.97) --
(232.63, 49.71) --
(234.48, 49.46) --
(236.34, 49.21) --
(238.19, 48.96) --
(240.05, 48.72) --
(241.90, 48.48) --
(243.75, 48.24) --
(245.61, 48.01) --
(247.46, 47.79) --
(249.32, 47.57) --
(251.17, 47.35) --
(253.03, 47.14) --
(254.88, 46.93) --
(256.74, 46.72) --
(258.59, 46.52) --
(260.44, 46.32) --
(262.30, 46.13) --
(264.15, 45.94) --
(266.01, 45.75) --
(267.86, 45.57) --
(269.72, 45.39) --
(271.57, 45.22) --
(273.42, 45.05) --
(275.28, 44.88) --
(277.13, 44.71) --
(278.99, 44.55) --
(280.84, 44.39) --
(282.70, 44.24) --
(284.55, 44.09) --
(286.40, 43.94) --
(288.26, 43.79) --
(290.11, 43.65) --
(291.97, 43.51) --
(293.82, 43.38) --
(295.68, 43.24) --
(297.53, 43.11) --
(299.39, 42.99) --
(301.24, 42.86) --
(303.09, 42.74) --
(304.95, 42.62) --
(306.80, 42.51) --
(308.66, 42.39) --
(310.51, 42.28) --
(312.37, 42.18) --
(314.22, 42.07) --
(316.07, 41.97) --
(317.93, 41.87) --
(319.78, 41.77) --
(321.64, 41.67) --
(323.49, 41.58) --
(325.35, 41.49) --
(327.20, 41.40) --
(329.06, 41.32) --
(330.91, 41.23) --
(332.76, 41.15) --
(334.62, 41.07) --
(336.47, 41.00) --
(338.33, 40.92) --
(340.18, 40.85) --
(342.04, 40.78) --
(343.89, 40.71) --
(345.74, 40.64) --
(347.60, 40.57) --
(349.45, 40.51) --
(351.31, 40.45) --
(353.16, 40.39) --
(355.02, 40.33) --
(356.87, 40.27) --
(358.73, 40.22) --
(360.58, 40.17) --
(362.43, 40.12) --
(364.29, 40.07) --
(366.14, 40.02) --
(368.00, 39.97) --
(369.85, 39.93) --
(371.71, 39.88) --
(373.56, 39.84) --
(375.41, 39.80) --
(377.27, 39.76) --
(379.12, 39.72) --
(380.98, 39.69) --
(382.83, 39.65) --
(384.69, 39.62) --
(386.54, 39.58) --
(388.40, 39.55) --
(390.25, 39.52) --
(392.10, 39.49) --
(393.96, 39.47) --
(395.81, 39.44) --
(397.67, 39.41) --
(399.52, 39.39) --
(401.38, 39.36) --
(403.23, 39.34) --
(405.08, 39.32) --
(406.94, 39.30);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 37.59,217.36) --
( 37.73,209.89) --
( 37.92,204.96) --
( 39.77,185.08) --
( 41.63,175.47) --
( 43.48,168.68) --
( 45.34,163.32) --
( 47.19,158.84) --
( 49.05,154.98) --
( 50.90,151.57) --
( 52.75,148.50) --
( 54.61,145.70) --
( 56.46,143.13) --
( 58.32,140.75) --
( 60.17,138.52) --
( 62.03,136.43) --
( 63.88,134.45) --
( 65.73,132.59) --
( 67.59,130.81) --
( 69.44,129.11) --
( 71.30,127.49) --
( 73.15,125.93) --
( 75.01,124.43) --
( 76.86,122.99) --
( 78.72,121.60) --
( 80.57,120.26) --
( 82.42,118.95) --
( 84.28,117.69) --
( 86.13,116.47) --
( 87.99,115.28) --
( 89.84,114.12) --
( 91.70,112.99) --
( 93.55,111.89) --
( 95.40,110.82) --
( 97.26,109.77) --
( 99.11,108.75) --
(100.97,107.75) --
(102.82,106.78) --
(104.68,105.82) --
(106.53,104.89) --
(108.39,103.97) --
(110.24,103.08) --
(112.09,102.20) --
(113.95,101.35) --
(115.80,100.51) --
(117.66, 99.68) --
(119.51, 98.88) --
(121.37, 98.09) --
(123.22, 97.31) --
(125.07, 96.56) --
(126.93, 95.81) --
(128.78, 95.09) --
(130.64, 94.37) --
(132.49, 93.68) --
(134.35, 92.99) --
(136.20, 92.32) --
(138.06, 91.67) --
(139.91, 91.03) --
(141.76, 90.40) --
(143.62, 89.78) --
(145.47, 89.18) --
(147.33, 88.59) --
(149.18, 88.01) --
(151.04, 87.45) --
(152.89, 86.90) --
(154.74, 86.36) --
(156.60, 85.83) --
(158.45, 85.31) --
(160.31, 84.81) --
(162.16, 84.31) --
(164.02, 83.83) --
(165.87, 83.36) --
(167.73, 82.90) --
(169.58, 82.45) --
(171.43, 82.01) --
(173.29, 81.58) --
(175.14, 81.16) --
(177.00, 80.75) --
(178.85, 80.36) --
(180.71, 79.97) --
(182.56, 79.59) --
(184.41, 79.22) --
(186.27, 78.86) --
(188.12, 78.51) --
(189.98, 78.16) --
(191.83, 77.83) --
(193.69, 77.50) --
(195.54, 77.19) --
(197.40, 76.88) --
(199.25, 76.58) --
(201.10, 76.29) --
(202.96, 76.00) --
(204.81, 75.72) --
(206.67, 75.46) --
(208.52, 75.19) --
(210.38, 74.94) --
(212.23, 74.69) --
(214.08, 74.45) --
(215.94, 74.22) --
(217.79, 73.99) --
(219.65, 73.77) --
(221.50, 73.56) --
(223.36, 73.35) --
(225.21, 73.15) --
(227.07, 72.96) --
(228.92, 72.77) --
(230.77, 72.58) --
(232.63, 72.41) --
(234.48, 72.23) --
(236.34, 72.07) --
(238.19, 71.91) --
(240.05, 71.75) --
(241.90, 71.60) --
(243.75, 71.45) --
(245.61, 71.31) --
(247.46, 71.18) --
(249.32, 71.04) --
(251.17, 70.92) --
(253.03, 70.80) --
(254.88, 70.68) --
(256.74, 70.56) --
(258.59, 70.45) --
(260.44, 70.35) --
(262.30, 70.25) --
(264.15, 70.15) --
(266.01, 70.05) --
(267.86, 69.96) --
(269.72, 69.88) --
(271.57, 69.79) --
(273.42, 69.71) --
(275.28, 69.64) --
(277.13, 69.56) --
(278.99, 69.49) --
(280.84, 69.42) --
(282.70, 69.36) --
(284.55, 69.30) --
(286.40, 69.24) --
(288.26, 69.18) --
(290.11, 69.13) --
(291.97, 69.08) --
(293.82, 69.03) --
(295.68, 68.98) --
(297.53, 68.94) --
(299.39, 68.90) --
(301.24, 68.86) --
(303.09, 68.82) --
(304.95, 68.78) --
(306.80, 68.75) --
(308.66, 68.72) --
(310.51, 68.69) --
(312.37, 68.66) --
(314.22, 68.64) --
(316.07, 68.61) --
(317.93, 68.59) --
(319.78, 68.57) --
(321.64, 68.55) --
(323.49, 68.53) --
(325.35, 68.51) --
(327.20, 68.50) --
(329.06, 68.48) --
(330.91, 68.47) --
(332.76, 68.46) --
(334.62, 68.45) --
(336.47, 68.44) --
(338.33, 68.43) --
(340.18, 68.42) --
(342.04, 68.41) --
(343.89, 68.41) --
(345.74, 68.40) --
(347.60, 68.39) --
(349.45, 68.39) --
(351.31, 68.39) --
(353.16, 68.38) --
(355.02, 68.38) --
(356.87, 68.38) --
(358.73, 68.38) --
(360.58, 68.37) --
(362.43, 68.37) --
(364.29, 68.37) --
(366.14, 68.37) --
(368.00, 68.37) --
(369.85, 68.37) --
(371.71, 68.37) --
(373.56, 68.37) --
(375.41, 68.37) --
(377.27, 68.37) --
(379.12, 68.36) --
(380.98, 68.36) --
(382.83, 68.36) --
(384.69, 68.36) --
(386.54, 68.36) --
(388.40, 68.36) --
(390.25, 68.35) --
(392.10, 68.35) --
(393.96, 68.35) --
(395.81, 68.35) --
(397.67, 68.35) --
(399.52, 68.35) --
(401.38, 68.35) --
(403.23, 68.35) --
(405.08, 68.35) --
(406.94, 68.35);
\definecolor{drawColor}{RGB}{81,81,81}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (275.28, 48.58) {$\alpha=0.05$};
\definecolor{drawColor}{RGB}{144,144,144}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (275.28, 73.31) {$\alpha=0.1$};
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (148.37, 30.40) -- (148.37,226.26);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 37.55, 30.40) --
( 37.55,226.26);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05, 36.54) {0.05};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05, 65.59) {0.10};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05, 94.64) {0.15};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05,123.68) {0.20};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05,152.73) {0.25};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05,181.78) {0.30};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 33.05,210.83) {0.35};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05, 39.30) --
( 37.55, 39.30);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05, 68.35) --
( 37.55, 68.35);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05, 97.39) --
( 37.55, 97.39);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05,126.44) --
( 37.55,126.44);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05,155.49) --
( 37.55,155.49);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05,184.53) --
( 37.55,184.53);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 35.05,213.58) --
( 37.55,213.58);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 37.55, 30.40) --
(406.94, 30.40);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 37.55, 27.90) --
( 37.55, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (129.90, 27.90) --
(129.90, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (222.24, 27.90) --
(222.24, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (314.59, 27.90) --
(314.59, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (406.94, 27.90) --
(406.94, 30.40);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 37.55, 20.39) {0};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (129.90, 20.39) {0.25};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (222.24, 20.39) {0.5};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (314.59, 20.39) {0.75};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (406.94, 20.39) {1};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (222.24, 6.94) {$w_{EB, i}=\mu_{2}/(\mu_{2}+\sigma^2_i)$};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 11.89,128.33) {Max. non-coverage probability};
\end{scope}
\end{tikzpicture}
}
\caption{Maximal non-coverage probability of parametric \ac{EBCI},
$\alpha \in \lbrace 0.05, 0.10 \rbrace$. The vertical line marks the ``rule of
thumb'' value $w_{EB, i}=0.3$, above which the maximal coverage distortion is
less than 5 percentage points for these two values of
$\alpha$.}\label{fig:noncov_param}
\end{figure}
If we additionally impose knowledge of the kurtosis of $\varepsilon_i$, the
maximal non-coverage of the parametric \ac{EBCI} can be similarly computed using
\Cref{eq:fourth_moment_bound}, as illustrated in the application in
\Cref{sec:empir-appl}.
\subsection{Monte Carlo simulations}\label{sec:sim}
Here we show through simulations that the robust \ac{EBCI} achieves accurate average coverage in finite samples.
\subsubsection{Design}
The \ac{DGP} is a simple linear fixed effects panel data model. We first draw
$\theta_i$, $i=1,\dotsc, n$, i.i.d.\ from a random effects distribution specified
below. Then we simulate panel data from the model
\begin{equation*}
W_{it} = \theta_i + U_{it}, \quad i=1, \dotsc, n, \quad t=1, \dotsc, T,
\end{equation*}
where the errors $U_{it}$ are mean zero and i.i.d.\ across $(i, t)$ and
independent of the $\theta_i$'s. The unshrunk estimator of $\theta_i$ is the
sample average of $W_{it}$ for unit $i$, with standard error obtained from the
usual unbiased variance estimator:
\begin{equation*}
Y_i = \frac{1}{T}\sum_{t=1}^T W_{it},
\quad \hat{\sigma}_i = \sqrt{\frac{1}{T(T-1)}\sum_{t=1}^T (W_{it}-Y_i)^2}.
\end{equation*}
We draw $U_{it}$ from one of two distributions: (1) a normal distribution and
(2) a (shifted) chi-squared distribution with 3 degrees of freedom. In case (1),
$Y_i$ is exactly normal conditional on $\theta_i$, but $\hat{\sigma}_i^2$ does
not exactly equal $\operatorname{var}(Y_i \mid \theta_i)$ for finite $T$. In case (2), $Y_i$
is non-normal and positively skewed (conditional on $\theta_i$) for finite $T$.
We consider six random
effects distributions for $\theta_i$ (see \Cref*{sec:sim-details} for detailed
definitions):
\begin{enumerate*}[label=(\roman*)]
\item normal (kurtosis $\kappa=3$);
\item scaled chi-squared with 1 degree of freedom ($\kappa=15$);
\item two-point distribution ($\kappa \approx 8.11$);
\item three-point distribution ($\kappa=2$);
\item the least favorable distribution for the robust \ac{EBCI} that exploits
only second moments ($\kappa$ depends on $\mu_{2}$, see \Cref{sec:comput});
and
\item the least favorable distribution for the parametric \ac{EBCI}.
\end{enumerate*}
Given $T$, we scale the $\theta_i$ distribution to match one of four signal-to-noise ratios
$\mu_2/\operatorname{var}(Y_i \mid \theta_i) \in \lbrace 0.1,0.5,1,2 \rbrace$, for a total of
$6 \times 4 = 24$ \acp{DGP} for each distribution of $U_{it}$. We shrink towards
the grand mean ($X_i=1$ for all $i$). We construct the robust \acp{EBCI}
following the baseline implementation in~\Cref{sec:baseline-implementation}
(with $\omega_{i}=1/n$), as well as a version that does not impose constraints
on the kurtosis.
As $T \to \infty$, we recover the idealized setting in
\Cref{sec:simple-example}, with $(Y_{i}-\theta_{i})/\sqrt{\operatorname{var}(Y_i \mid \theta_i)}$
converging in distribution to a standard normal (conditional on $\theta_i$), and
$\hat{\sigma}_i^2/\operatorname{var}(Y_i \mid \theta_i)$ converging in probability to 1, for
each $i$.
\subsubsection{Results}
\Cref{tab:sim_panel_normal} shows that the 95\% robust \acp{EBCI} achieve good
average coverage when the panel errors $U_{it}$ are normally distributed. This
is true for all \acp{DGP}, panel dimensions $n$ and $T$, and whether we exploit
one or both of the (estimated) moments $\mu_2$ and $\kappa$. When the time
dimension $T$ equals 10, the maximal coverage distortion across all \acp{DGP} and all
cross-sectional dimensions $n \in \lbrace 100,200,500\rbrace$ is 3.2 percentage
points. For $T \geq 20$, the coverage distortion of the robust \acp{EBCI} is
always below 2.1 percentage points.
\begin{table}[tp]
\centering
\begin{threeparttable}
\caption{Monte Carlo simulation results, panel data with normal
errors.}\label{tab:sim_panel_normal}
\begin{tabular*}{0.95\linewidth}{@{\extracolsep{\fill}}@{}l cccc cccc cccc@{}}
& \multicolumn{4}{@{}c}{Robust, $\mu_2$ only}
& \multicolumn{4}{@{}c}{Robust, $\mu_2$ \& $\kappa$}
& \multicolumn{4}{@{}c}{Parametric} \\
\cmidrule(rl){2-5}\cmidrule(rl){6-9}\cmidrule(rl){10-13}
$T$&10&20&$\infty$&ora&10&20&$\infty$&ora&10&20&$\infty$&ora\\
\midrule
\multicolumn{13}{@{}l}{Panel A\@{}: Average coverage (\%), minimum across 24 \acsp{DGP}}\\
\cmidrule(r){1-9}
$n=100$ & 92.1 & 93.7 & 94.0 & 95.0 & 91.8 & 93.2 & 93.2 & 94.6 & 79.2 & 79.7 & 79.3 & 86.9 \\
$n=200$ & 91.9 & 93.4 & 92.9 & 95.0 & 91.8 & 93.3 & 92.9 & 94.8 & 80.7 & 80.3 & 81.0 & 86.3 \\
$n=500$ & 91.9 & 93.6 & 94.8 & 95.0 & 91.9 & 93.5 & 94.3 & 94.9 & 84.2 & 85.1 & 85.1 & 85.6
\\
\multicolumn{13}{@{}l}{Panel B\@{}: Relative average length, average across 24 \acsp{DGP}}\\
\cmidrule(r){1-9}
$n=100$ & 1.09 & 1.10 & 1.11 & 1.16 & 1.03 & 1.02 & 1.02 & 1.00 & 0.81 & 0.82 & 0.83 & 0.86 \\
$n=200$ & 1.09 & 1.10 & 1.12 & 1.16 & 1.02 & 1.02 & 1.01 & 1.00 & 0.81 & 0.82 & 0.84 & 0.86 \\
$n=500$ & 1.10 & 1.11 & 1.13 & 1.16 & 1.04 & 1.03 & 1.01 & 1.00 & 0.82 & 0.83 & 0.84 & 0.86
\end{tabular*}
\begin{tablenotes}
\item \emph{Notes:} Normally distributed errors. Nominal average confidence
level $1-\alpha=95\%$. All \ac{EBCI} procedures use baseline estimate of
$\hat{\mu}_2$ and (if applicable) $\hat{\kappa}$, except columns labeled
``ora'', which use oracle values of $\mu_2$ and $\kappa$. Columns
$T=\infty$ and ``ora'' use oracle standard errors $\sigma_i$. For each
\acs{DGP}, ``average coverage'' and ``average length'' refer to averages
across units $i=1, \dotsc, n$ and across 2,000 Monte Carlo repetitions.
Average \acs{CI} length is measured relative to the robust \ac{EBCI} that
exploits the oracle values of $\mu_2$, $\kappa$, and $\sigma_i$ (but not
of the grand mean $\delta=E[\theta]$).
\end{tablenotes}
\end{threeparttable} \\
\bigskip\bigskip
\begin{threeparttable}
\caption{Monte Carlo simulation results, panel data with chi-squared errors.}\label{tab:sim_panel_chi2}
\begin{tabular*}{0.95\linewidth}{@{\extracolsep{\fill}}@{}l cccc cccc cccc@{}}
& \multicolumn{4}{@{}c}{Robust, $\mu_2$ only} & \multicolumn{4}{@{}c}{Robust, $\mu_2$ \& $\kappa$} & \multicolumn{4}{@{}c}{Parametric} \\
\cmidrule(rl){2-5}\cmidrule(rl){6-9}\cmidrule(rl){10-13}
$T$&10&20&50&ora&10&20&50&ora&10&20&50&ora\\
\midrule
\multicolumn{13}{@{}l}{Panel A\@{}: Average coverage (\%), minimum across 24 \acsp{DGP}}\\
\cmidrule(r){1-9}
$n=100$ & 87.9 & 90.9 & 93.1 & 95.0 & 87.8 & 90.8 & 92.6 & 94.7 & 79.9 & 79.3 & 79.3 & 87.0 \\
$n=200$ & 87.9 & 90.8 & 93.0 & 94.9 & 87.8 & 90.8 & 92.8 & 94.8 & 77.8 & 79.8 & 80.3 & 86.2 \\
$n=500$ & 87.8 & 90.8 & 93.0 & 95.0 & 87.8 & 90.7 & 92.9 & 94.9 & 82.0 & 84.1 & 84.8 & 85.6
\\
\multicolumn{13}{@{}l}{Panel B\@{}: Relative average length, average across 24 \acsp{DGP}}\\
\cmidrule(r){1-9}
$n=100$ & 1.05 & 1.08 & 1.10 & 1.16 & 1.01 & 1.02 & 1.02 & 1.00 & 0.79 & 0.81 & 0.82 & 0.86 \\
$n=200$ & 1.04 & 1.08 & 1.10 & 1.16 & 0.99 & 1.00 & 1.00 & 1.00 & 0.78 & 0.81 & 0.82 & 0.86 \\
$n=500$ & 1.05 & 1.09 & 1.11 & 1.16 & 0.99 & 1.00 & 1.00 & 1.00 & 0.79 & 0.82 & 0.83 & 0.86
\end{tabular*}
\begin{tablenotes}
\item \emph{Notes:} Chi-squared distributed errors. See caption for \Cref{tab:sim_panel_normal}. Results for $T=\infty$ are by definition the same as in \Cref{tab:sim_panel_normal}.
\end{tablenotes}
\end{threeparttable}
\end{table}
\Cref{tab:sim_panel_chi2} shows that coverage distortions are somewhat larger
when the panel errors $U_{it}$ are chi-squared distributed and $T$ is small. The
robust \acp{EBCI} undercover by up to 7.2 percentage points when $T=10$ due to
the pronounced non-normality of $Y_i$ given $\theta_i$. However, the distortion
is at most 4.3 percentage points when $T=20$, and at most 2.4 percentage points
when $T \geq 50$. The coverage distortion due to non-normality when $T$ is small
is similar to the coverage distortion of the usual unshrunk \ac{CI} (not
reported).
Importantly, in all cases considered in
\Cref{tab:sim_panel_normal,tab:sim_panel_chi2}, the worst-case coverage
distortion of the parametric \ac{EBCI} substantially exceeds that of the
corresponding robust \acp{EBCI}, sometimes by more than 10 percentage points.
Nevertheless, the cost of robustness in terms of extra \ac{CI} length is modest
and consistent with the theoretical results in \Cref{sec:efficiency}.
Both the estimation of the standard errors $\sigma_i$ and the estimation of the
moments $\mu_2$ and $\kappa$ contribute to the finite-sample coverage
distortions. The ``ora'' columns in \Cref{tab:sim_panel_normal} exploit the
oracle (true) values of $\mu_2$, $\kappa$, and $\sigma_i=\sqrt{\operatorname{var}(Y_i \mid \theta_i)}$, while the
$T=\infty$ columns use oracle standard errors but not oracle moments. By
comparing these columns, we see that estimation of $\mu_2$ and $\kappa$ is
responsible for modest coverage distortions when $n=100$ or $200$. However,
estimation of the standard errors $\sigma_i$ also contributes to the
distortions, as can be seen by comparing the $T=10$ and $T=\infty$ columns.
In \Cref*{sec:sim-heterosk} we show that the robust
\ac{EBCI} also has good coverage in a heteroskedastic design calibrated to the
empirical application in \Cref{sec:empir-appl} below.
\section{Comparison with other approaches}\label{sec:compar}
Here we compare our \ac{EBCI} procedure with other approaches to confidence
interval construction in the normal means model. We also discuss other related
inference problems.
\subsection{Average coverage vs.\ alternative coverage concepts}\label{sec:compar_avg_cov}
The average coverage requirement in \Cref{eq:aci_def} is less stringent than the
usual (pointwise) notion of frequentist coverage that
$P(\theta_i\in CI_i\mid \theta) \geq 1-\alpha$ for all $i$. An even stronger
coverage requirement is that of simultaneous coverage:
$P(\forall i\colon \theta_i\in CI_{i}\mid \theta) \geq 1-\alpha$. As outlined
in~\cref{fn:impossibility}, under the pointwise coverage criterion, one cannot
achieve substantial reductions in length relative to the unshrunk \ac{CI}. Under
the simultaneous coverage criterion, it is likewise impossible to substantially
improve upon the usual sup-$t$ confidence band based on the unshrunk estimates
\citep{Cai2015}. Thus, undercoverage for some $\theta_{i}$'s must be tolerated
if one wants to use shrinkage to improve \ac{CI} length.
The fact that our \acp{EBCI} achieve improvements in average length at the
expense of undercovering for certain units $i$ is analogous to well-known
properties of \ac{EB} point estimators. We now show that the units $i$ for which
our \ac{EBCI} undercovers are quantitatively similar to the units for which the
shrinkage estimator $\hat{\theta}_{i}$ has higher \ac{MSE} than the unshrunk
estimator $Y_i$. Let $\varepsilon_{i}=\theta_{i}-X_{i}'\delta$ be the ``shrinkage error'' defined in
\Cref{sec:baseline-model}. The pointwise coverage of our \ac{EBCI} is decreasing in the normalized shrinkage error $\abs{\varepsilon_{i}}/\sqrt{\mu_2}$, for a
fixed signal-to-noise ratio $\mu_2/\sigma_i^2$.\footnote{The pointwise coverage (conditional on
$X_{i}$) equals
$1-r(\sqrt{1/w_{EB, i}-1}\cdot \abs{\varepsilon_{i}}/\sqrt{\mu_{2}},
\operatorname{cva}_{\alpha}(1/w_{EB, i}-1,\kappa))$, with $r$ defined in \Cref{eq:rb} and $w_{EB, i}=\mu_{2}/(\mu_{2}+\sigma_{i}^{2})$.} Hence, the units $i$ for which our
\ac{EBCI} undercovers are those whose covariate-predicted value $X_i'\delta$
fails to approximate their true effect $\theta_i$ well. The \ac{MSE} of the
shrinkage estimator (for an individual unit $i$), normalized by the \ac{MSE}
of the unshrunk estimator, is
similarly increasing in $\abs{\varepsilon_{i}}/\sqrt{\mu_2}$.\footnote{The ratio of \acp{MSE} equals
$E[(\hat{\theta}_i-\theta_i)^2 \mid \theta_i, X_i]/\sigma_{i}^{2}=w_{EB,
i}^{2}+(1-w_{EB, i}) w_{EB, i}\cdot \abs{\varepsilon_{i}}/\sqrt{\mu_2}$.}
\begin{figure}[t]
\centering
\tikzset{font=\small}
\begin{tikzpicture}[x=1pt,y=1pt]
\definecolor{fillColor}{RGB}{255,255,255}
\path[use as bounding box,fill=fillColor,fill opacity=0.00] (0,0) rectangle (411.94,231.26);
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{255,255,255}
\definecolor{fillColor}{RGB}{255,255,255}
\path[draw=drawColor,line width= 0.6pt,line join=round,line cap=round,fill=fillColor] ( 0.00, 0.00) rectangle (411.94,231.26);
\end{scope}
\begin{scope}
\path[clip] ( 27.33, 30.40) rectangle (406.94,226.26);
\definecolor{fillColor}{RGB}{255,255,255}
\path[fill=fillColor] ( 27.33, 30.40) rectangle (406.94,226.26);
\definecolor{drawColor}{RGB}{208,208,208}
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 27.33, 30.40) --
(406.94, 30.40);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 27.33, 69.57) --
(406.94, 69.57);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 27.33,108.74) --
(406.94,108.74);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 27.33,147.92) --
(406.94,147.92);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 27.33,187.09) --
(406.94,187.09);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 27.33,226.26) --
(406.94,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 27.33, 30.40) --
( 27.33,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (122.23, 30.40) --
(122.23,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (217.13, 30.40) --
(217.13,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (312.04, 30.40) --
(312.04,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (406.94, 30.40) --
(406.94,226.26);
\definecolor{drawColor}{RGB}{81,81,81}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 31.13,189.27) --
( 31.47,188.61) --
( 31.82,187.98) --
( 32.16,187.38) --
( 32.51,186.79) --
( 32.85,186.23) --
( 33.20,185.69) --
( 33.54,185.17) --
( 33.89,184.66) --
( 34.23,184.17) --
( 34.58,183.69) --
( 34.92,183.22) --
( 35.27,182.76) --
( 35.61,182.31) --
( 35.96,181.88) --
( 36.30,181.45) --
( 36.65,181.03) --
( 36.99,180.62) --
( 37.34,180.22) --
( 37.68,179.82) --
( 38.03,179.43) --
( 38.37,179.05) --
( 38.72,178.68) --
( 39.06,178.31) --
( 39.41,177.95) --
( 39.75,177.59) --
( 40.10,177.24) --
( 40.44,176.89) --
( 40.79,176.55) --
( 41.13,176.21) --
( 41.48,175.87) --
( 41.82,175.54) --
( 42.17,175.22) --
( 42.51,174.90) --
( 42.86,174.58) --
( 43.20,174.27) --
( 43.55,173.96) --
( 43.89,173.65) --
( 44.24,173.35) --
( 44.59,173.04) --
( 44.93,172.75) --
( 45.28,172.45) --
( 45.62,172.16) --
( 45.97,171.87) --
( 46.31,171.59) --
( 46.66,171.30) --
( 47.00,171.02) --
( 47.35,170.75) --
( 47.69,170.47) --
( 48.04,170.20) --
( 48.38,169.93) --
( 48.73,169.66) --
( 49.07,169.39) --
( 49.42,169.13) --
( 49.76,168.86) --
( 50.11,168.60) --
( 50.45,168.35) --
( 50.80,168.09) --
( 51.14,167.83) --
( 51.49,167.58) --
( 51.83,167.33) --
( 52.18,167.08) --
( 52.52,166.83) --
( 52.87,166.59) --
( 53.21,166.34) --
( 53.56,166.10) --
( 53.90,165.86) --
( 54.25,165.62) --
( 54.59,165.38) --
( 54.94,165.15) --
( 55.28,164.91) --
( 55.63,164.68) --
( 55.97,164.45) --
( 56.32,164.22) --
( 56.66,163.99) --
( 57.01,163.76) --
( 57.35,163.53) --
( 57.70,163.31) --
( 58.04,163.08) --
( 58.39,162.86) --
( 58.73,162.64) --
( 59.08,162.41) --
( 59.42,162.19) --
( 59.77,161.98) --
( 60.11,161.76) --
( 60.46,161.54) --
( 60.80,161.33) --
( 61.15,161.11) --
( 61.49,160.90) --
( 61.84,160.69) --
( 62.19,160.47) --
( 62.53,160.26) --
( 62.88,160.05) --
( 63.22,159.85) --
( 63.57,159.64) --
( 63.91,159.43) --
( 64.26,159.23) --
( 64.60,159.02) --
( 64.95,158.82) --
( 65.29,158.61) --
( 65.29,158.61) --
( 68.74,156.63) --
( 72.19,154.71) --
( 75.63,152.86) --
( 79.08,151.06) --
( 82.53,149.31) --
( 85.97,147.60) --
( 89.42,145.93) --
( 92.87,144.29) --
( 96.32,142.68) --
( 99.76,141.10) --
(103.21,139.54) --
(106.66,138.01) --
(110.10,136.49) --
(113.55,134.99) --
(117.00,133.50) --
(120.45,132.03) --
(123.89,130.57) --
(127.34,129.13) --
(130.79,127.69) --
(134.23,126.26) --
(137.68,124.83) --
(141.13,123.42) --
(144.58,122.00) --
(148.02,120.60) --
(151.47,119.19) --
(154.92,117.79) --
(158.36,116.39) --
(161.81,115.00) --
(165.26,113.60) --
(168.71,112.21) --
(172.15,110.82) --
(175.60,109.43) --
(179.05,108.05) --
(182.49,106.67) --
(185.94,105.30) --
(189.39,103.93) --
(192.84,102.58) --
(196.28,101.24) --
(199.73, 99.92) --
(203.18, 98.62) --
(206.62, 97.34) --
(210.07, 96.10) --
(213.52, 94.88) --
(216.97, 93.69) --
(220.41, 92.54) --
(223.86, 91.43) --
(227.31, 90.37) --
(230.75, 89.34) --
(234.20, 88.35) --
(237.65, 87.41) --
(241.10, 86.51) --
(244.54, 85.65) --
(247.99, 84.83) --
(251.44, 84.05) --
(254.88, 83.31) --
(258.33, 82.61) --
(261.78, 81.94) --
(265.23, 81.31) --
(268.67, 80.71) --
(272.12, 80.14) --
(275.57, 79.60) --
(279.01, 79.09) --
(282.46, 78.60) --
(285.91, 78.14) --
(289.36, 77.70) --
(292.80, 77.29) --
(296.25, 76.90) --
(299.70, 76.52) --
(303.14, 76.17) --
(306.59, 75.83) --
(310.04, 75.51) --
(313.49, 75.21) --
(316.93, 74.92) --
(320.38, 74.65) --
(323.83, 74.38) --
(327.27, 74.14) --
(330.72, 73.90) --
(334.17, 73.67) --
(337.62, 73.46) --
(341.06, 73.25) --
(344.51, 73.06) --
(347.96, 72.87) --
(351.40, 72.69) --
(354.85, 72.53) --
(358.30, 72.36) --
(361.75, 72.21) --
(365.19, 72.06) --
(368.64, 71.92) --
(372.09, 71.79) --
(375.53, 71.66) --
(378.98, 71.54) --
(382.43, 71.42) --
(385.88, 71.31) --
(389.32, 71.21) --
(392.77, 71.11) --
(396.22, 71.01) --
(399.67, 70.92) --
(403.11, 70.83) --
(406.56, 70.75);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 2pt off 2pt ,line join=round] ( 43.20,225.92) --
( 43.55,223.92) --
( 43.89,221.97) --
( 44.24,220.09) --
( 44.59,218.27) --
( 44.93,216.50) --
( 45.28,214.78) --
( 45.62,213.11) --
( 45.97,211.49) --
( 46.31,209.91) --
( 46.66,208.38) --
( 47.00,206.89) --
( 47.35,205.43) --
( 47.69,204.02) --
( 48.04,202.64) --
( 48.38,201.30) --
( 48.73,199.99) --
( 49.07,198.71) --
( 49.42,197.46) --
( 49.76,196.24) --
( 50.11,195.05) --
( 50.45,193.89) --
( 50.80,192.75) --
( 51.14,191.64) --
( 51.49,190.55) --
( 51.83,189.49) --
( 52.18,188.44) --
( 52.52,187.43) --
( 52.87,186.43) --
( 53.21,185.45) --
( 53.56,184.49) --
( 53.90,183.55) --
( 54.25,182.63) --
( 54.59,181.73) --
( 54.94,180.85) --
( 55.28,179.98) --
( 55.63,179.12) --
( 55.97,178.29) --
( 56.32,177.47) --
( 56.66,176.66) --
( 57.01,175.87) --
( 57.35,175.09) --
( 57.70,174.33) --
( 58.04,173.58) --
( 58.39,172.84) --
( 58.73,172.12) --
( 59.08,171.40) --
( 59.42,170.70) --
( 59.77,170.01) --
( 60.11,169.33) --
( 60.46,168.66) --
( 60.80,168.01) --
( 61.15,167.36) --
( 61.49,166.72) --
( 61.84,166.10) --
( 62.19,165.48) --
( 62.53,164.87) --
( 62.88,164.27) --
( 63.22,163.68) --
( 63.57,163.10) --
( 63.91,162.53) --
( 64.26,161.97) --
( 64.60,161.41) --
( 64.95,160.86) --
( 65.29,160.32) --
( 65.29,160.32) --
( 68.74,155.31) --
( 72.19,150.90) --
( 75.63,146.99) --
( 79.08,143.50) --
( 82.53,140.34) --
( 85.97,137.49) --
( 89.42,134.88) --
( 92.87,132.49) --
( 96.32,130.29) --
( 99.76,128.26) --
(103.21,126.37) --
(106.66,124.62) --
(110.10,122.98) --
(113.55,121.45) --
(117.00,120.01) --
(120.45,118.66) --
(123.89,117.39) --
(127.34,116.18) --
(130.79,115.04) --
(134.23,113.96) --
(137.68,112.94) --
(141.13,111.97) --
(144.58,111.04) --
(148.02,110.15) --
(151.47,109.31) --
(154.92,108.50) --
(158.36,107.73) --
(161.81,106.99) --
(165.26,106.28) --
(168.71,105.60) --
(172.15,104.94) --
(175.60,104.31) --
(179.05,103.71) --
(182.49,103.12) --
(185.94,102.56) --
(189.39,102.01) --
(192.84,101.49) --
(196.28,100.98) --
(199.73,100.49) --
(203.18,100.02) --
(206.62, 99.56) --
(210.07, 99.12) --
(213.52, 98.68) --
(216.97, 98.27) --
(220.41, 97.86) --
(223.86, 97.47) --
(227.31, 97.09) --
(230.75, 96.72) --
(234.20, 96.35) --
(237.65, 96.00) --
(241.10, 95.66) --
(244.54, 95.33) --
(247.99, 95.01) --
(251.44, 94.69) --
(254.88, 94.38) --
(258.33, 94.09) --
(261.78, 93.79) --
(265.23, 93.51) --
(268.67, 93.23) --
(272.12, 92.96) --
(275.57, 92.70) --
(279.01, 92.44) --
(282.46, 92.18) --
(285.91, 91.94) --
(289.36, 91.70) --
(292.80, 91.46) --
(296.25, 91.23) --
(299.70, 91.00) --
(303.14, 90.78) --
(306.59, 90.57) --
(310.04, 90.36) --
(313.49, 90.15) --
(316.93, 89.94) --
(320.38, 89.75) --
(323.83, 89.55) --
(327.27, 89.36) --
(330.72, 89.17) --
(334.17, 88.99) --
(337.62, 88.81) --
(341.06, 88.63) --
(344.51, 88.46) --
(347.96, 88.29) --
(351.40, 88.12) --
(354.85, 87.96) --
(358.30, 87.80) --
(361.75, 87.64) --
(365.19, 87.48) --
(368.64, 87.33) --
(372.09, 87.18) --
(375.53, 87.03) --
(378.98, 86.89) --
(382.43, 86.74) --
(385.88, 86.60) --
(389.32, 86.47) --
(392.77, 86.33) --
(396.22, 86.20) --
(399.67, 86.07) --
(403.11, 85.94) --
(406.56, 85.81);
\definecolor{drawColor}{RGB}{0,0,0}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (144.57,131.52) {coverage};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (323.83, 93.21) {MSE};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 27.33, 30.40) --
( 27.33,226.26);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 22.83, 27.64) {0};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 22.83, 66.81) {1};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 22.83,105.99) {2};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 22.83,145.16) {3};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 22.83,184.34) {4};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 22.83,223.51) {5};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 24.83, 30.40) --
( 27.33, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 24.83, 69.57) --
( 27.33, 69.57);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 24.83,108.74) --
( 27.33,108.74);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 24.83,147.92) --
( 27.33,147.92);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 24.83,187.09) --
( 27.33,187.09);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 24.83,226.26) --
( 27.33,226.26);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 27.33, 30.40) --
(406.94, 30.40);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 27.33, 27.90) --
( 27.33, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (122.23, 27.90) --
(122.23, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (217.13, 27.90) --
(217.13, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (312.04, 27.90) --
(312.04, 30.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (406.94, 27.90) --
(406.94, 30.40);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at ( 27.33, 20.39) {0};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (122.23, 20.39) {0.25};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (217.13, 20.39) {0.5};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (312.04, 20.39) {0.75};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.80] at (406.94, 20.39) {1};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (217.13, 6.94) {$w_{EB, i}=\mu_{2}/(\mu_{2}+\sigma^2_i)$};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 11.89,128.33) {$|\varepsilon_{i}|/\sqrt{\mu_{2}}$};
\end{scope}
\end{tikzpicture}
\caption{Value of $\abs{\varepsilon_{i}}/\sqrt{\mu_{2}}$, as a function of
$w_{EB, i}$, such that the \ac{MSE} of the shrinkage point estimator equals
that of the unshrunk estimator (MSE), and such that the coverage of the
robust \ac{EBCI} with $\kappa=\infty$ equals the nominal average coverage
$1-\alpha$ (coverage), for $\alpha=0.05$.}\label{fig:cond_coverage}
\end{figure}
\Cref{fig:cond_coverage} shows that the knife-edge value of
$\abs{\varepsilon_{i}}/\sqrt{\mu_2}$ for which the pointwise coverage of our
\ac{EBCI} equals $1-\alpha$ is quantitatively close to the value of
$\abs{\varepsilon_i}/\sqrt{\mu_2}$ for which the MSE of the shrinkage estimator
equals that of the unshrunk estimator. In other words, to the extent that one
worries about undercoverage for certain types of $\theta_i$ values, one should
simultaneously worry about the relative performance of the shrinkage point
estimator for those same values.
We stress that the pointwise coverage depends on the unobservable shrinkage
error $\varepsilon_{i}$, which cannot be gauged directly from the observables
$(Y_{i}, X_{i})$. If one wishes to avoid systematic differences in coverage
across units $i$ with different genders, say (i.e.,\ one is worried that
$\varepsilon_{i}$ correlates with gender) one can simply add gender to the set
of covariates $X_{i}$: the baseline procedure in
\Cref{sec:baseline-implementation} ensures control of average coverage
conditional on the covariates $X_i$. In \Cref{sec:cover_selection}, we show how
to adapt our \acp{EBCI} to settings where one focuses the analysis on a subset
of units $i$ based on the values of their unshrunk estimates $Y_i$ (e.g.,
keeping only the estimates that exceed a given threshold).
From a Bayesian point of view, our robust \ac{EBCI} can be viewed as an
uncertainty interval that is robust to the choice of prior distribution in the
\emph{unconditional} gamma-minimax sense: the coverage probability of this
\ac{CI} is at least $1-\alpha$ when averaged over the distribution of the data
and over the prior distribution for $\theta_{i}$, for any prior distribution
that satisfies the moment bounds. This follows directly from the derivations in
\Cref{sec:simple-example}, reinterpreting the random effects distribution for
$\theta_i$ as a prior distribution. In contrast, \emph{conditional}
gamma-minimax credible intervals, discussed recently by \citet[p.
6]{Kitagawa2019}, are too stringent in our setting. This notion requires that
the posterior credibility of the interval be at least $1-\alpha$ regardless of
the choice of prior, in any data sample, which would require reporting the
entire parameter space (up to the moment bounds).
\subsection{Finite-sample vs.\ asymptotic coverage}\label{sec:compar_finite}
Our procedures are asymptotically valid as $n\to\infty$, as proved in
\Cref{sec:cover-under-basel}. These asymptotics do not capture the impact of
estimation error in the ``hyper-parameters'' $\hat{\sigma}_i$, $\hat{\delta}$,
$\hat{\mu}_2$, and $\hat{\kappa}$, or the impact of lack of exact normality of
the $Y_{i}$'s, on the finite-sample performance of the \acp{EBCI}. As detailed
in \Cref{sec:baseline-implementation} and \Cref{sec:moment_estimates}, we do
apply a finite-sample adjustment to the moments $\hat{\mu}_2$ and
$\hat{\kappa}$, which is motivated by the same heuristic arguments that
\citet{morris83pebci,morris83} uses to motivate finite-sample adjustments to the
parametric \ac{EBCI}.\footnote{\label{fn:bootstrap}An alternative approach would be to adapt the
bootstrap adjustment proposed by \citet[Ch. 3.5.3]{CaLo00} in the context of
parametric \ac{EBCI} construction \citep[see also][]{efron2019}. As with the
\citet{morris83pebci,morris83} adjustment, we are not aware of a formal result
justifying it.} The promising simulation results in \Cref{sec:sim}
notwithstanding, these adjustments do not ensure exact average coverage control
in finite samples.\footnote{\label{fn:bonferroni}One could account for hyperparameter uncertainty by
computing the critical value
$\sup_{\tilde{\sigma}_i, \tilde{\mu}_2,\tilde{\kappa} \in \hat{\mathcal{C}}_i}
\operatorname{cva}_{\alpha}(\tilde{\sigma}_i^{2}/\tilde{\mu}_{2}, \tilde{\kappa})$ over an
initial confidence set $\hat{\mathcal{C}}_i$ for the hyper-parameters, coupled
with a Bonferroni adjustment of the confidence level $1-\alpha$. This approach
appears to be highly conservative in practice.}
Our results are thus analogous to standard results on coverage of Eicker-Huber-White
\acp{CI} in cross-sectional \ac{OLS}: asymptotic
validity follows by consistency of the \ac{OLS} variance estimate and asymptotic normality of the outcomes, while adjustments to account for finite-sample issues (such as the HC2 or HC3 variance estimators studied in \citealp{MaWh85}) are justified
heuristically. Deriving \acp{EBCI} with finite-sample coverage guarantees is an
interesting problem that we leave for future research; the problem appears to be
challenging even in the context of constructing parametric \acp{EBCI}.
\subsection{Local vs.\ global optimality}\label{sec:compar_optim}
Our \acp{EBCI} are designed to provide uncertainty assessments to accompany
linear shrinkage estimates that, as the Introduction argues, have been popular
in applied work. Our procedure's global validity, as well as local near-optimality when the $\theta_{i}$'s are normal (cf. \Cref{sec:efficiency}), is analogous to
Eicker-Huber-White \acp{CI} for \ac{OLS} estimators: these \acp{CI} are optimal
under normal homoskedastic regression errors, but remain valid when this
assumption is dropped.
Similar to the Eicker-Huber-White \acp{CI}, our \acp{EBCI} are not globally
efficient: when the $\theta_{i}$'s are not Gaussian, it is generally inefficient to restrict attention to \acp{CI} that are centered at a linear point estimator and have fixed width. While we expect our \acp{EBCI} to remain
near-efficient under mild departures from normality, substantial efficiency
gains may be possible if the effect size distribution is, for example, heavy-tailed or bimodal.\footnote{\label{fn:conservative}Indeed, if the true effect
distribution puts mass $1/2$ on $\theta_{i}=K$ and $\theta_{i}=-K$, then, as
$K$ gets large, our \acp{EBCI} become arbitrarily conservative relative to an
oracle that reports the highest posterior density set under this prior.}
\Cref{sec:general_shrinkage} shows how our method can be adapted to construct
\acp{EBCI} that are locally near-optimal under non-normal baseline priors using
non-linear shrinkage, such as soft thresholding. Since the distribution of
$\theta_{i}$ is nonparametrically identified under the normal
model~\eqref{eq:hierarch_y}, it is in principle possible to construct \acp{EBCI}
that are globally efficient using nonparametric methods. In the context of the
homoskedastic model with no covariates in \Cref{eq:homoskedastic-normal-means},
various approaches to nonparametric point estimation of the $\theta_{i}$'s have
been proposed, including kernels \citep{BrGr09}, splines \citep{efron2019}, or
nonparametric maximum likelihood \citep{KiWo56,JiZh09,KoMi14}. An interesting
problem for future research is to adapt these methods to \ac{EBCI} construction,
while ensuring asymptotic validity, good finite-sample performance, and allowing
for covariates, heteroskedasticity, and possible dependence across $i$.
\subsection{Other inference problems}\label{sec:compar_other_inf}
A number of alternative inference procedures have been proposed in the context
of the normal means model. \citet{Efron2015} develops a formula for the
frequentist standard error of \ac{EB} estimators, but this cannot be used to
construct \acp{CI} without a corresponding estimate of the bias. There is a
substantial literature on shrinkage confidence balls, i.e., confidence sets of
the form $\{\theta\colon\sum_{i=1}^{n} (\theta_i-\hat\theta_i)^2\le \hat c\}$
(see \citealp{casella_shrinkage_2012}, for a review). While theoretically
interesting, these sets can be difficult to visualize and report in
practice.\footnote{Confidence balls can be translated into average coverage
intervals using Chebyshev's inequality \citep[see][Ch. 5.8]{Wasserman2006}.
However, such intervals are very conservative compared to the ones we
construct.}
Finally, while we focus on \ac{CI} length in our relative efficiency
comparisons, our approach can be fruitfully applied when the goal of \ac{CI}
construction is to discern non-null effects, rather than to construct short
\acp{CI}. In particular, suppose one forms a test of the null hypothesis
$H_{0,i}:\theta_i=\theta_0$ for some null value $\theta_0$ by rejecting when
$\theta_0\notin CI_i$, where $CI_i$ is our robust \ac{EBCI} given
in~\eqref{eq:conditional_ci}. In \Cref*{sec:power-details}, we show that the
test based on our \ac{EBCI} has higher average power than the usual $z$-test
based on the unshrunk estimate when $X_i'\delta$ (the regression line towards
which we shrink) is far enough from the null value $\theta_0$, and that these
power gains can be substantial. Furthermore, such tests can be combined with
corrections from the multiple testing literature to form procedures that
asymptotically control the \ac{FDR}, a commonly used criterion for multiple
testing.\footnote{In particular, \citet{storey_direct_2002} shows that the
\citet{benjamini_controlling_1995} procedure asymptotically controls the
\ac{FDR} so long as the $p$-values do not exhibit too much statistical
dependence and the proportion of rejected null hypotheses does not converge
too quickly to zero. While \citet{storey_direct_2002} assumes that the
uncorrected tests control size in the classical sense, the argument goes
through essentially unchanged so long as the tests invert \acp{CI} that
satisfy (\ref{eq:alt_aci_def}), which holds so long as the \acp{CI} do not
exhibit too much statistical dependence, as discussed in
\Cref{rem:ac_alt_def_remark}. We note, however, that this does not hold for
modifications of the \citet{benjamini_controlling_1995} procedure that use
initial estimates of the proportion of true null hypotheses.}
\section{Extensions}\label{sec:extensions}
We now discuss two extensions of our method: adapting our intervals to general,
possibly non-linear shrinkage, and constructing intervals that achieve coverage
conditional on $Y_{i}$ falling into a pre-specified interval.
\subsection{General shrinkage}\label{sec:general_shrinkage}
Our method can be generalized to cover general, possibly non-linear shrinkage
based on possibly non-Gaussian data. Let
$\mathcal{S}(y; \chi, \tilde{X}_{i})\subseteq \mathbb{R}$ be a family of
candidate confidence sets for a parameter $\theta_{i}$, which depends on the
data $Y_{i}=y$, a tuning parameter $\chi\in\mathbb{R}$ to be selected below, and
covariates $\tilde{X}_{i}$ (that include any known nuisance parameters) that we
treat as fixed. We assume that $\mathcal{S}$ is increasing in $\chi$, in the
sense of set containment, and that the non-coverage probability conditional on
$\theta$ satisfies
\begin{equation}\label{eq:conditional_coverage_general}
P(\theta_{i}\not\in
\mathcal{S}(Y_{i}; \chi, \tilde{X}_{i}) \mid \theta, \tilde{X}^{(n)})=
\tilde{r}(a_{i}, \chi),
\end{equation}
where $a_{i}$ is some function of $\theta_i$,
$\tilde{X}^{(n)}=(\tilde{X}_{1}, \dotsc, \tilde{X}_{n})$, and $\tilde r$ is a
known function (perhaps computed numerically or through simulation). Similarly
to linear shrinkage in the normal means model,
\Cref{eq:conditional_coverage_general} may only hold approximately if the set
$\mathcal{S}$ depends on estimated parameters (such as standard error estimates
or tuning parameters), or if we use a large-sample approximation to the
distribution of $Y_{i}$. We assume that $a_{i}$ satisfies the moment constraints
$E_{F}[g(a_{i})\mid \tilde{X}^{(n)}]=m$, where $g$ is a $p$-vector of moment
functions, and the expectation is over the conditional distribution $F$ of
$a_{i}$ conditional on $\tilde{X}^{(n)}$.\footnote{The moment functions $g$ need
not be simple moments, and could incorporate constraints used for selection of
hyper-parameters, such as constraints on the marginal data distribution or, if
an unbiased risk criterion is used, the constraint that the derivative of the
risk equals zero at the selected prior hyper-parameters.} To guarantee \ac{EB}
coverage, we compute the maximal non-coverage
\begin{equation}\label{eq:rho_general}
\rho_{g}(m, \chi)=\sup_{F}E_{F}[\tilde r(a, \chi)], \qquad E_{F}[g(a)]=m,
\end{equation}
analogously to \Cref{eq:fourth_moment_bound}. This is a linear program, which
can be computed numerically to a high degree of precision even with several
constraints; see \Cref{sec:comput} for details. Given an estimate $\hat m$ of
the moment vector $m$, we form a robust \ac{EBCI} as
\begin{equation}\label{eq:ebci_general}
\mathcal{S}(Y_{i};\hat \chi, \tilde{X}_{i}),
\quad\text{where}\quad
\hat\chi = \inf\{\chi\colon \rho_{g}(\hat{m}, \chi)\leq \alpha\}.
\end{equation}
\begin{example}[Linear shrinkage in the normal model]\label{example:main_case}
The setting in \Cref{sec:baseline-model} obtains if we set
$\tilde{X}_{i}=(X_{i}, \sigma_{i})$ and
$\mathcal{S}(y;\chi, \tilde{X}_{i})=\{(1-w_{EB, i})X_{i}'\delta+w_{EB,
i}Y_{i}\pm \chi w_{EB, i}\sigma_{i}\}$. Here $a_{i}$ is given by the
normalized bias $b_{i}=(1/w_{EB, i}-1)(\theta_{i}-X_{i}'\delta)/\sigma_{i}$,
and the function $\tilde r$ is given by the function $r(b, \chi)$ defined
in~\eqref{eq:rb}. Our baseline implementation uses constraints on the second
and fourth moments, $g(a_{i})=(a_{i}^{2}, a_{i}^{4})$.
\end{example}
\begin{example}[Nonlinear soft thresholding]\label{example:soft_thresholding}
Consider for simplicity the homoskedastic normal model
$Y_i \mid \theta_i \sim N(\theta_i, \sigma^2)$ without covariates. A popular
alternative to linear estimators is the soft thresholding estimator
$\hat{\theta}_{ST, i} = \text{sign}(Y_i)\max\lbrace
\abs{Y_i}-\sqrt{2\sigma^2/\mu_2},0\rbrace$ \citep[e.g.][]{Abadie2019}. It equals the
posterior mode corresponding to a baseline Laplace prior with second moment
$\mu_2$, which has density
$\pi_0(\theta)=\frac{1}{\sqrt{2\mu_{2}}} \exp(-\abs{\theta}\sqrt{2/\mu_{2}})$
\citep[Example 2.5]{johnstone19}. To construct a robust \ac{EBCI} that always
contains the soft thresholding estimator, we calibrate the corresponding
highest posterior density set:
\begin{equation}\label{eq:hpd_st}
\mathcal{S}(Y_i; \chi) = \left\lbrace t \in \mathbb{R} \colon \log
\frac{\sigma^{-1}\phi((Y_i- t)/\sigma)\pi_0(t)
}{\int_{-\infty}^\infty \sigma^{-1}\phi((Y_i-
\tilde{\theta})/\sigma)\pi_0(\tilde{\theta})\, d\tilde{\theta}} + \chi \geq
0 \right\rbrace,
\end{equation}
where $\phi$ is the standard normal density. This set is available in closed
form and takes the form of an interval (see \Cref*{sec:appendix_softthresh}).
Here $a_{i}=\theta_{i}$, and the function $\tilde r(a, \chi)$
in~\eqref{eq:conditional_coverage_general} can be computed via numerical
integration.
In contrast to the \acp{EBCI} in \Cref{example:main_case} (which
may be viewed as calibrating the highest posterior density set under a
\emph{normal} prior), the Laplace prior $\pi_{0}$ leads to nonlinear
shrinkage and an \ac{EBCI} whose length depends on the data $Y_i$. This
reflects the suboptimality of linear shrinkage and fixed-length intervals under
the Laplace prior.
In \Cref*{sec:appendix_softthresh}, we show that the resulting robust
\ac{EBCI} that imposes the constraint $E[\theta_{i}^{2}]=\mu_{2}$ not only has
robust \ac{EB} coverage (by definition), it also achieves substantial expected
length improvements when the $\theta_i$'s are in fact Laplace distributed. For
$\alpha=0.05$ and $\mu_2/\sigma^2 \leq 0.2$, the expected length under the
Laplace distribution of the soft thresholding \ac{EBCI} is at least 49\%
smaller than the length of the unshrunk CI\@. This exceeds the length
reduction achieved by the linear robust \ac{EBCI} shown in
\Cref{fig:efficiency_unshrunk}.
\end{example}
\begin{example}[Poisson shrinkage]\label{example:poisson}
\Cref*{sec:appendix_poisson} constructs a robust \ac{EBCI} for the rate
parameter $\theta_i$ in a Poisson model
$Y_i \mid \theta_i \sim \text{Poisson}(\theta_i)$. This example demonstrates
that our general approach does not require normality of the data.
\end{example}
\begin{example}[Linear estimators in other settings]\label{example:general_linear}
While our focus has been on \ac{EB} shrinkage, our approach applies to other
settings in which an estimator $\hat\theta_i$ is approximately normally
distributed with non-negligible bias. In particular, suppose
$(\hat\theta_i-\theta_i)/\operatorname{se}_i$ is distributed $N(a_i,1)$, where $\operatorname{se}_i$ is
the standard deviation of the estimate $\hat\theta_i$, which for simplicity we
take to be known. This holds whenever $\hat\theta_i$ is a linear function of
jointly normal observations $W_{1}, \dotsc, W_{N}$, i.e.,
$ \hat{\theta}_{i}=\sum_{j=1}^{N} k_{ij}W_{j}$ for some deterministic weights $k_{ij}$.
Examples include series, kernel, or local polynomial estimators in a
nonparametric regression with fixed covariates and normal errors. We can
construct a confidence interval for $\theta_{i}$ as
$\hat\theta_i\pm \chi \cdot \operatorname{se}_i$, in which case
\Cref{eq:conditional_coverage_general} holds with $\tilde{r}=r$ given in
\Cref{eq:rb}. It follows from \Cref{high_level_conditional_coverage_thm} in
\Cref{coverage_results_sec_append} that if the moment constraints $m$ on the
normalized bias in~\Cref{eq:rho_general} are replaced by consistent estimates,
the resulting robust \ac{EBCI} will satisfy the average coverage
property~\eqref{eq:aci_def} in large samples. We leave a full treatment of
these applications for future research.
\end{example}
\subsection{Coverage after selection}\label{sec:cover_selection}
In some applications, researchers may be primarily interested in parameters
corresponding to those units $i$ whose initial estimates $Y_{i}$ fall in a given
interval $[\iota_1,\iota_2]$, where
$-\infty \leq \iota_1 < \iota_2 \leq \infty$. For example, in a teacher value
added application, we may only be interested in the ability $\theta_i$ of those
teachers $i$ whose fixed effect estimates $Y_i$ are positive, corresponding to
setting $\iota_1=0$ and $\iota_2=\infty$. Because of the selection on outcomes,
na\"{i}vely applying our baseline \ac{EBCI} procedure to the selected sample
$\lbrace i\colon Y_i \in [\iota_1,\iota_2]\rbrace$ does not yield the desired
average coverage across the selected units $i$. We now show how to correct for
the selection bias in the simple homoskedastic model
$Y_i \mid \theta_i \sim N(\theta_i, \sigma^2)$ without covariates from
\Cref{sec:simple-example} (reintroducing the extra model features in
\Cref{sec:baseline-model} only complicates notation).
We seek a critical value $\chi$ such that the average coverage of the \ac{CI}
$[\hat{\theta}_i \pm \chi w_{EB} \sigma]$ is at least $1-\alpha$
\emph{conditional} on the sample selection, i.e.,
\begin{equation}\label{eq:conditional_ebci_def}
P(\theta_i \in \hat{\theta}_i \pm \chi w_{EB} \sigma \mid Y_i \in [\iota_1,\iota_2]) \geq 1-\alpha
\end{equation}
under repeated sampling of $(Y_i, \theta_i)$, regardless of the distribution for
$\theta_i$ (we maintain focus on linear shrinkage for simplicity, but our
approach extends to nonlinear shrinkage using the ideas in
\Cref{sec:general_shrinkage}). Straightforward calculations show that the
non-coverage, conditional on $\theta_{i}$ and on selection, equals
\begin{multline*}
\tilde{r}_{\iota_1,\iota_2}(\theta_i, \chi) = P(\theta_i \notin \hat{\theta}_i \pm \chi w_{EB} \sigma \mid Y_i \in [\iota_1,\iota_2], \theta_i) \\
= \min\left\lbrace 1-\frac{\Phi(\min\lbrace
\chi-b_{i}, (\iota_2-\theta_i)/\sigma\rbrace)-\Phi(\max\lbrace
-\chi-b_{i}, (\iota_1-\theta_i)/\sigma\rbrace)}{\Phi((\iota_2-\theta_i)/\sigma)-\Phi((\iota_1-\theta_i)/\sigma)},
1\right\rbrace,
\end{multline*}
where $b_{i}=(1-1/w_{EB})\theta_{i}/\sigma$ as in \Cref{sec:simple-example}.
Among all distributions for $\theta_i$ consistent with the conditional moment
$\tilde{\mu}_{2,\iota_1,\iota_2}=E\left[\theta_i^2 \mid Y_i \in
[\iota_1,\iota_2]\right]$, the worst-case non-coverage probability,
conditional on selection, is given by
\begin{equation*}
\tilde{\rho}_{\iota_1,\iota_2}(\tilde{\mu}_{2,\iota_1,\iota_2}, \chi) \equiv
\sup_{F} E_{F}[\tilde{r}_{\iota_1,\iota_2}(\theta_i, \chi)] \quad \text{s.t.} \quad
E_{F}[\theta_i^2]=\tilde{\mu}_{2,\iota_1,\iota_2},
\end{equation*}
where $E_{F}$ denotes expectation under $\theta_{i}\sim F$. This is an
infinite-dimensional linear program that can be solved numerically to a high
degree of accuracy, cf.\ \Cref{sec:comput}. To achieve robust conditional
coverage, we solve numerically for the $\chi$ such that
$\tilde{\rho}_{\iota_1,\iota_2}(\tilde{\mu}_{2,\iota_1,\iota_2}, \chi) =
\alpha$.
We can estimate the conditional second moment $\tilde{\mu}_{2,\iota_1,\iota_2}$
as follows. Denote the log marginal density of $Y_i$ by
$\ell(y) \equiv \log \int \phi(y-\theta)\, d\Gamma_0(\theta)$, where $\Gamma_0$
is the true distribution of $\theta_i$. Tweedie's formulas
\citep[e.g.][Eq.~(26)]{efron2019} imply
\begin{equation}\label{eq:cond_cov_moment}
\tilde{\mu}_{2,\iota_1,\iota_2} = E\left[\theta_i^2 \mid Y_i
\in [\iota_1,\iota_2]\right] = 1 +
E\left[(Y_i + \ell'(Y_i))^2 + \ell''(Y_i) \;\big|\; Y_i \in [\iota_1,\iota_2] \right].
\end{equation}
Let $\hat{\ell}(y)$ be a kernel estimate of the log marginal density function of
the data $Y_1, \dotsc, Y_n$. Then the estimate
\begin{equation*}
\widehat{\tilde{\mu}}_{2,\iota_1,\iota_2} \equiv 1 +
\frac{\sum_{i \colon Y_i \in [\iota_1,\iota_2]}
\lbrace (Y_i + \hat{\ell}'(Y_i))^2 + \hat{\ell}''(Y_i)\rbrace}{{\#}\lbrace i \colon Y_i \in [\iota_1,\iota_2] \rbrace}
\end{equation*}
will be consistent as $n\to\infty$ for $\tilde{\mu}_{2,\iota_1,\iota_2}$
in~\eqref{eq:cond_cov_moment} under mild regularity conditions.
The criterion~\eqref{eq:conditional_ebci_def} can be viewed as the \ac{EB}
analogue of the criterion
$P(\theta_i \in CI_i \mid Y_i \in [\iota_1,\iota_2], \theta) \geq 1-\alpha$,
which requires frequentist coverage conditional on the event
$\{Y_i\in [\iota_1,\iota_2]\}$. The latter criterion has been considered in the recent ``selective inference'' literature
\citep{benjamini_false_2005,lee_exact_2013,HuFi19,Andrews2021}. In contrast to this literature, we cannot allow $\iota_1$ to be given by the maximum of the
initial estimates \citep[as in][]{Andrews2021}, as we require $\iota_{1}$ and $\iota_{2}$ to converge in probability to distinct nonrandom limits. On the other hand, weakening the
notion of frequentist coverage to \ac{EB} (or average) coverage allows for
improvements in the length of the intervals, similar to the analysis in \Cref{sec:efficiency} in the absence of selection.
\section{Empirical application}\label{sec:empir-appl}
We illustrate our methods using the data and model in
\citet{chetty_impacts_2018}, who are interested in the effect of neighborhoods
on intergenerational mobility.
\subsection{Framework}
We follow \citet{chetty_impacts_2018} in using two definitions of a
``neighborhood effect'' $\theta_{i}$. The first focuses on effects for children
growing up in low-income families, and defines $\theta_{i}$ as the effect of
spending an additional year of childhood in \ac{CZ} $i$ on children's rank in
the income distribution at age 26, for children with parents at the 25th
percentile of the national income distribution. The second definition is
analogous, except it focuses on children growing up in high-income families, and
consequently conditions on children with parents at the 75th percentile.
\citet{chetty_impacts_2018} argue that these definitions approximately capture
the mean rank effects for children in below-median and above-median income
families. Using de-identified tax returns for all children born between 1980 and
1986 who move across \acp{CZ} exactly once as children,
\citet{chetty_impacts_2018} exploit variation in the age at which children move
between \acp{CZ} to obtain preliminary fixed effect estimates $Y_{i}$ of
$\theta_{i}$.
Since these preliminary estimates are measured with noise, to predict $\theta_{i}$, \citet{chetty_impacts_2018} shrink $Y_{i}$ towards
average outcomes of permanent residents of \ac{CZ} $i$ (children with parents at
the same percentile of the income distribution who spent all of their childhood
in the \ac{CZ}). To give a sense of the accuracy of these forecasts,
\citet{chetty_impacts_2018} report estimates of their unconditional \ac{MSE}
(i.e.,\ treating $\theta_{i}$ as random), under
the implicit assumption that the moment independence assumption in
\Cref{eq:moment_independence} holds. Here we complement their analysis by
constructing robust \acp{EBCI} associated with these forecasts.
Our sample consists of 595 U.S. \acp{CZ}, with population over 25,000 in the
2000 census: this is the sample for which \citet{chetty_impacts_2018} report
baseline estimates $Y_{i}$ of the effects $\theta_{i}$.
These baseline estimates are normalized so that their population-weighted mean
is zero. We may therefore interpret $\theta_{i}$ as the effect relative to an
``average'' \ac{CZ}. We follow the baseline implementation from
\Cref{sec:baseline-implementation} with standard errors $\hat{\sigma}_{i}$
reported by \citet{chetty_impacts_2018}, and covariates $X_{i}$ corresponding to
a constant and the average outcomes for permanent residents. In line with the
original analysis, we use precision weights $\omega_i=1/\hat{\sigma}^{2}_{i}$
when constructing the estimates $\hat{\delta}$, $\hat{\mu}_{2}$ and
$\hat{\kappa}$.
\subsection{Results}
\begin{table}[p]
\centering
\begin{threeparttable}
\caption{Statistics for 90\% \acp{EBCI} for neighborhood
effects.}\label{tab:ch}
\begin{tabular*}{0.9\linewidth}{@{\extracolsep{\fill}}@{}l@{}lrrrr@{}}
&& \multicolumn{2}{c}{Baseline} & \multicolumn{2}{c}{Nonparametric} \\
\cmidrule(rl){3-4} \cmidrule(rl){5-6}
&& \multicolumn{1}{r}{(1)} & \multicolumn{1}{r}{(2)}
& \multicolumn{1}{r}{(3)} & \multicolumn{1}{r}{(4)}\\
\phantom{a}&Percentile & 25th & 75th & 25th & 75th \\
\midrule
\multicolumn{2}{@{}l}{Panel A\@{}: Summary statistics}\\
\cmidrule{1-2}
& $E[\sqrt{\mu_{2,i}}]$ & 0.079& 0.044& 0.076& 0.042\\
& $E[\kappa_{i}]$ & 778.5& 5948.6& 1624.9& 43009.9\\
& $E[\mu_{2i}/\sigma_i^2]$ & 0.142& 0.040& 0.139& 0.072\\
& $\hat{\delta}_{\text{intercept}}$ & $-1.441$& $-2.162$& $-1.441$& $-2.162$\\
& $\hat{\delta}_{\text{perm.\ resident}}$ & 0.032& 0.038& 0.032& 0.038\\
& $E[w_{EB, i}]$ & 0.093& 0.033& 0.093& 0.033\\
& $E[w_{opt, i}]$ & 0.191& 0.100& 0.191& 0.100\\
& $E[\text{non-cov of parametric EBCI}_i]$& 0.227& 0.278& 0.210& 0.292\\
\multicolumn{2}{@{}l}{Panel B\@: $E[\text{half-length}_i]$}\\
\cmidrule{1-2}
& Robust EBCI & 0.195& 0.122& 0.186& 0.116 \\
& Optimal robust EBCI & 0.149& 0.090& 0.145& 0.094 \\
& Parametric EBCI & 0.123& 0.070& 0.123& 0.070 \\
& Unshrunk CI & 0.786& 0.993& 0.786& 0.993 \\[1ex]
\multicolumn{2}{@{}l}{Panel C\@: Efficiency relative to robust EBCI}\\
\cmidrule{1-2}
& Optimal robust EBCI & 1.312& 1.352& 1.289& 1.238 \\
& Parametric EBCI & 1.582& 1.731& 1.509& 1.648 \\
& Unshrunk CI & 0.248& 0.123& 0.237& 0.117 \\
\end{tabular*}
\begin{tablenotes}
\item \emph{Notes:} Columns (1) and (2) correspond to shrinking $Y_{i}$ as
in the baseline implementation that imposes \Cref{eq:moment_independence},
so that $\mu_{2i}=E[(\theta_{i}-X_{i}'\delta)^{2}\mid X_{i}, \sigma_{i}]$
and
$\kappa_{i}=E[(\theta_{i}-X_{i}'\delta)^{4}\mid X_{i},
\sigma_{i}]/\mu_{2i}^{2}$ do not vary with $i$. Columns (3) and (4) use
nonparametric estimates of $\mu_{2i}$ and $\kappa_{i}$, using the nearest
neighbor estimator described in \Cref{sec:finite_n}. The number of nearest
neighbors $J=422$ (column (3)) and $J=525$ (column (4)) is selected using
cross-validation. For all columns,
$\hat{\delta}=(\hat{\delta}_{\text{intercept}}, \hat{\delta}_{\text{perm.\
resident}})$ is computed by regressing $Y_{i}$ onto a constant and
outcomes for permanent residents. ``Optimal Robust EBCI'' refers to a
robust EBCI based on length-optimal shrinkage $w_{opt, i}$, described in
\Cref{sec:efficiency}. ``$E[\text{non-cov of parametric EBCI}_i]$'':
average of maximal non-coverage probability of parametric EBCI, given the
estimated moments.
\end{tablenotes}
\end{threeparttable}
\end{table}
Columns (1) and (2) in \Cref{tab:ch} summarize the main estimation and
efficiency results. The shrinkage magnitude and relative efficiency results are
similar for children with parents at the 25th and 75th percentiles of the income
distribution. In both columns, the estimate of the kurtosis $\kappa$ is large
enough so that it does not affect the critical values or the form of the optimal
shrinkage: specifications that only impose constraints on the second moment
yield identical results.\footnote{The truncation in the $\hat{\kappa}$ formula
in our baseline algorithm in \Cref{sec:baseline-implementation} binds in
columns (1) and (2), although the non-truncated estimates 345.3 and 5024.9 are
similarly large; using these non-truncated estimates yields identical
results.} In line with this finding, a density plot of the $t$-statistics
(reported as Figure S2 in \citet{akp20v2}) exhibits a fat lower tail. As a
robustness check, columns (3) and (4) show that results based on nonparametric
moment estimates (see \Cref{rem:nonparam,sec:finite_n}) are very similar to our baseline specification. Indeed, the $R^{2}$ gain in predicting
$\hat{\varepsilon}_{i}^{2}-\hat{\sigma}^{2}_{i}$ using $\hat{\mu}_{2i}$ is less
than $0.001$ in both specifications, indicating that there is little evidence in
the data against the moment independence assumption.
The baseline robust 90\% \acp{EBCI} are 75.2--87.7\% shorter than the usual
unshrunk \acp{CI} $Y_{i}\pm z_{1-\alpha/2}\hat{\sigma}_{i}$. To interpret these
gains in dollar terms, for children with parents at the 25th percentile of the
income distribution, a percentile gain corresponds to an annual income gain of
\$818 \citep[p.~1183]{chetty_impacts_2018}. Thus, the average half-length of the
baseline robust \acp{EBCI} in column (1) implies \acp{CI} of the form
$\pm \$160$ on average, while the unshrunk \acp{CI} are of the form $\pm \$643$
on average. These large gains are a consequence of a low signal-to-noise ratio
$\mu_{2}/\sigma_{i}^{2}$ in this application. Because the shrinkage
magnitude is so large on average, the tail behavior of the bias matters, and
since the kurtosis estimates suggests these tails are fat, it is important to
use the robust critical value: the parametric \ac{EBCI} exhibits average
potential size distortions of 12.7--17.8 percentage points. Indeed, for over 90\% of the \acp{CI} in
the specifications in columns (1) and (2), the shrinkage coefficient $w_{EB, i}$
falls below the ``rule of thumb'' threshold of 0.3 derived in
\Cref{sec:param-ebci-cover}.
\begin{figure}[t]
\centering
\tikzset{font=\small}
\begin{tikzpicture}[x=1pt,y=1pt]
\definecolor{fillColor}{RGB}{255,255,255}
\path[use as bounding box,fill=fillColor,fill opacity=0.00] (0,0) rectangle (411.94,231.26);
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{255,255,255}
\definecolor{fillColor}{RGB}{255,255,255}
\path[draw=drawColor,line width= 0.6pt,line join=round,line cap=round,fill=fillColor] ( 0.00, 0.00) rectangle (411.94,231.26);
\end{scope}
\begin{scope}
\path[clip] ( 31.38, 27.75) rectangle (406.94,226.26);
\definecolor{fillColor}{RGB}{255,255,255}
\path[fill=fillColor] ( 31.38, 27.75) rectangle (406.94,226.26);
\definecolor{drawColor}{RGB}{208,208,208}
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 31.38, 29.60) --
(406.94, 29.60);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 31.38, 68.84) --
(406.94, 68.84);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 31.38,108.09) --
(406.94,108.09);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 31.38,147.33) --
(406.94,147.33);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 31.38,186.57) --
(406.94,186.57);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 31.38,225.82) --
(406.94,225.82);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] ( 38.08, 27.75) --
( 38.08,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (122.65, 27.75) --
(122.65,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (207.22, 27.75) --
(207.22,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (291.79, 27.75) --
(291.79,226.26);
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 4pt off 4pt ,line join=round] (376.36, 27.75) --
(376.36,226.26);
\definecolor{fillColor}{RGB}{144,144,144}
\path[fill=fillColor] (114.11,127.40) circle ( 1.96);
\path[fill=fillColor] (345.93,173.61) circle ( 1.96);
\path[fill=fillColor] (242.82, 69.39) circle ( 1.96);
\path[fill=fillColor] ( 52.68,114.68) circle ( 1.96);
\path[fill=fillColor] (302.62,112.45) circle ( 1.96);
\path[fill=fillColor] (385.64,106.19) circle ( 1.96);
\path[fill=fillColor] (365.54,150.20) circle ( 1.96);
\path[fill=fillColor] (250.11,154.03) circle ( 1.96);
\path[fill=fillColor] (293.02,153.44) circle ( 1.96);
\path[fill=fillColor] (208.17, 92.46) circle ( 1.96);
\path[fill=fillColor] (126.54, 81.97) circle ( 1.96);
\path[fill=fillColor] (117.34, 96.43) circle ( 1.96);
\definecolor{fillColor}{RGB}{81,81,81}
\path[fill=fillColor] (112.14,108.66) --
(116.07,108.66) --
(116.07,112.58) --
(112.14,112.58) --
cycle;
\path[fill=fillColor] (343.97,114.92) --
(347.89,114.92) --
(347.89,118.85) --
(343.97,118.85) --
cycle;
\path[fill=fillColor] (240.86,105.04) --
(244.78,105.04) --
(244.78,108.96) --
(240.86,108.96) --
cycle;
\path[fill=fillColor] ( 50.72,105.88) --
( 54.65,105.88) --
( 54.65,109.81) --
( 50.72,109.81) --
cycle;
\path[fill=fillColor] (300.66,110.51) --
(304.58,110.51) --
(304.58,114.44) --
(300.66,114.44) --
cycle;
\path[fill=fillColor] (383.68,112.37) --
(387.60,112.37) --
(387.60,116.30) --
(383.68,116.30) --
cycle;
\path[fill=fillColor] (363.58,113.83) --
(367.50,113.83) --
(367.50,117.75) --
(363.58,117.75) --
cycle;
\path[fill=fillColor] (248.15,110.56) --
(252.07,110.56) --
(252.07,114.48) --
(248.15,114.48) --
cycle;
\path[fill=fillColor] (291.06,111.94) --
(294.98,111.94) --
(294.98,115.86) --
(291.06,115.86) --
cycle;
\path[fill=fillColor] (206.21,104.92) --
(210.13,104.92) --
(210.13,108.85) --
(206.21,108.85) --
cycle;
\path[fill=fillColor] (124.57, 98.34) --
(128.50, 98.34) --
(128.50,102.27) --
(124.57,102.27) --
cycle;
\path[fill=fillColor] (115.37, 96.99) --
(119.30, 96.99) --
(119.30,100.91) --
(115.37,100.91) --
cycle;
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] (109.88,148.78) --
(118.33,148.78);
\path[draw=drawColor,line width= 0.6pt,line join=round] (114.11,148.78) --
(114.11,106.03);
\path[draw=drawColor,line width= 0.6pt,line join=round] (109.88,106.03) --
(118.33,106.03);
\path[draw=drawColor,line width= 0.6pt,line join=round] (341.70,217.24) --
(350.16,217.24);
\path[draw=drawColor,line width= 0.6pt,line join=round] (345.93,217.24) --
(345.93,129.99);
\path[draw=drawColor,line width= 0.6pt,line join=round] (341.70,129.99) --
(350.16,129.99);
\path[draw=drawColor,line width= 0.6pt,line join=round] (238.59,102.01) --
(247.05,102.01);
\path[draw=drawColor,line width= 0.6pt,line join=round] (242.82,102.01) --
(242.82, 36.77);
\path[draw=drawColor,line width= 0.6pt,line join=round] (238.59, 36.77) --
(247.05, 36.77);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 48.46,130.35) --
( 56.91,130.35);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 52.68,130.35) --
( 52.68, 99.00);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 48.46, 99.00) --
( 56.91, 99.00);
\path[draw=drawColor,line width= 0.6pt,line join=round] (298.39,153.22) --
(306.85,153.22);
\path[draw=drawColor,line width= 0.6pt,line join=round] (302.62,153.22) --
(302.62, 71.68);
\path[draw=drawColor,line width= 0.6pt,line join=round] (298.39, 71.68) --
(306.85, 71.68);
\path[draw=drawColor,line width= 0.6pt,line join=round] (381.41,142.54) --
(389.87,142.54);
\path[draw=drawColor,line width= 0.6pt,line join=round] (385.64,142.54) --
(385.64, 69.84);
\path[draw=drawColor,line width= 0.6pt,line join=round] (381.41, 69.84) --
(389.87, 69.84);
\path[draw=drawColor,line width= 0.6pt,line join=round] (361.31,200.42) --
(369.77,200.42);
\path[draw=drawColor,line width= 0.6pt,line join=round] (365.54,200.42) --
(365.54, 99.97);
\path[draw=drawColor,line width= 0.6pt,line join=round] (361.31, 99.97) --
(369.77, 99.97);
\path[draw=drawColor,line width= 0.6pt,line join=round] (245.88,205.64) --
(254.34,205.64);
\path[draw=drawColor,line width= 0.6pt,line join=round] (250.11,205.64) --
(250.11,102.42);
\path[draw=drawColor,line width= 0.6pt,line join=round] (245.88,102.42) --
(254.34,102.42);
\path[draw=drawColor,line width= 0.6pt,line join=round] (288.79,202.48) --
(297.25,202.48);
\path[draw=drawColor,line width= 0.6pt,line join=round] (293.02,202.48) --
(293.02,104.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (288.79,104.40) --
(297.25,104.40);
\path[draw=drawColor,line width= 0.6pt,line join=round] (203.94,115.82) --
(212.40,115.82);
\path[draw=drawColor,line width= 0.6pt,line join=round] (208.17,115.82) --
(208.17, 69.09);
\path[draw=drawColor,line width= 0.6pt,line join=round] (203.94, 69.09) --
(212.40, 69.09);
\path[draw=drawColor,line width= 0.6pt,line join=round] (122.31, 98.63) --
(130.76, 98.63);
\path[draw=drawColor,line width= 0.6pt,line join=round] (126.54, 98.63) --
(126.54, 65.31);
\path[draw=drawColor,line width= 0.6pt,line join=round] (122.31, 65.31) --
(130.76, 65.31);
\path[draw=drawColor,line width= 0.6pt,line join=round] (113.11,102.17) --
(121.56,102.17);
\path[draw=drawColor,line width= 0.6pt,line join=round] (117.34,102.17) --
(117.34, 90.70);
\path[draw=drawColor,line width= 0.6pt,line join=round] (113.11, 90.70) --
(121.56, 90.70);
\definecolor{drawColor}{RGB}{81,81,81}
\path[draw=drawColor,line width= 0.6pt,line join=round] (109.88,122.88) --
(118.33,122.88);
\path[draw=drawColor,line width= 0.6pt,line join=round] (114.11,122.88) --
(114.11, 98.36);
\path[draw=drawColor,line width= 0.6pt,line join=round] (109.88, 98.36) --
(118.33, 98.36);
\path[draw=drawColor,line width= 0.6pt,line join=round] (341.70,132.79) --
(350.16,132.79);
\path[draw=drawColor,line width= 0.6pt,line join=round] (345.93,132.79) --
(345.93,100.98);
\path[draw=drawColor,line width= 0.6pt,line join=round] (341.70,100.98) --
(350.16,100.98);
\path[draw=drawColor,line width= 0.6pt,line join=round] (238.59,121.69) --
(247.05,121.69);
\path[draw=drawColor,line width= 0.6pt,line join=round] (242.82,121.69) --
(242.82, 92.31);
\path[draw=drawColor,line width= 0.6pt,line join=round] (238.59, 92.31) --
(247.05, 92.31);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 48.46,117.90) --
( 56.91,117.90);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 52.68,117.90) --
( 52.68, 97.79);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 48.46, 97.79) --
( 56.91, 97.79);
\path[draw=drawColor,line width= 0.6pt,line join=round] (298.39,128.13) --
(306.85,128.13);
\path[draw=drawColor,line width= 0.6pt,line join=round] (302.62,128.13) --
(302.62, 96.83);
\path[draw=drawColor,line width= 0.6pt,line join=round] (298.39, 96.83) --
(306.85, 96.83);
\path[draw=drawColor,line width= 0.6pt,line join=round] (381.41,129.52) --
(389.87,129.52);
\path[draw=drawColor,line width= 0.6pt,line join=round] (385.64,129.52) --
(385.64, 99.16);
\path[draw=drawColor,line width= 0.6pt,line join=round] (381.41, 99.16) --
(389.87, 99.16);
\path[draw=drawColor,line width= 0.6pt,line join=round] (361.31,132.17) --
(369.77,132.17);
\path[draw=drawColor,line width= 0.6pt,line join=round] (365.54,132.17) --
(365.54, 99.41);
\path[draw=drawColor,line width= 0.6pt,line join=round] (361.31, 99.41) --
(369.77, 99.41);
\path[draw=drawColor,line width= 0.6pt,line join=round] (245.88,128.98) --
(254.34,128.98);
\path[draw=drawColor,line width= 0.6pt,line join=round] (250.11,128.98) --
(250.11, 96.05);
\path[draw=drawColor,line width= 0.6pt,line join=round] (245.88, 96.05) --
(254.34, 96.05);
\path[draw=drawColor,line width= 0.6pt,line join=round] (288.79,130.21) --
(297.25,130.21);
\path[draw=drawColor,line width= 0.6pt,line join=round] (293.02,130.21) --
(293.02, 97.60);
\path[draw=drawColor,line width= 0.6pt,line join=round] (288.79, 97.60) --
(297.25, 97.60);
\path[draw=drawColor,line width= 0.6pt,line join=round] (203.94,119.72) --
(212.40,119.72);
\path[draw=drawColor,line width= 0.6pt,line join=round] (208.17,119.72) --
(208.17, 94.05);
\path[draw=drawColor,line width= 0.6pt,line join=round] (203.94, 94.05) --
(212.40, 94.05);
\path[draw=drawColor,line width= 0.6pt,line join=round] (122.31,110.81) --
(130.76,110.81);
\path[draw=drawColor,line width= 0.6pt,line join=round] (126.54,110.81) --
(126.54, 89.80);
\path[draw=drawColor,line width= 0.6pt,line join=round] (122.31, 89.80) --
(130.76, 89.80);
\path[draw=drawColor,line width= 0.6pt,line join=round] (113.11,103.95) --
(121.56,103.95);
\path[draw=drawColor,line width= 0.6pt,line join=round] (117.34,103.95) --
(117.34, 93.94);
\path[draw=drawColor,line width= 0.6pt,line join=round] (113.11, 93.94) --
(121.56, 93.94);
\definecolor{drawColor}{gray}{0.50}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.85] at (242.82, 30.76) {Union};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.85] at (385.64, 59.05) {Olean};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.85] at (208.17, 58.30) {Albany};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.85] at (126.53, 54.52) {Poughkeepsie};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.85] at (117.34, 79.91) {New York};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.85] at (114.10,153.69) {Syracuse};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.85] at (326.00,217.37) {Oneonta};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.85] at ( 52.68,135.26) {Buffalo};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.85] at (302.62,158.13) {Elmira};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.85] at (365.54,205.33) {Watertown};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.85] at (232.50,210.55) {Plattsburgh};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.85] at (282.45,207.39) {Amsterdam};
\definecolor{drawColor}{RGB}{81,81,81}
\path[draw=drawColor,line width= 0.6pt,dash pattern=on 1pt off 3pt ,line join=round] (-344.17, 93.00) -- (782.49,126.93);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 31.38, 27.75) --
( 31.38,226.26);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.50] at ( 26.88, 27.88) {-1.0};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.50] at ( 26.88, 67.12) {-0.5};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.50] at ( 26.88,106.36) {0.0};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.50] at ( 26.88,145.61) {0.5};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.50] at ( 26.88,184.85) {1.0};
\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 0.50] at ( 26.88,224.10) {1.5};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 28.88, 29.60) --
( 31.38, 29.60);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 28.88, 68.84) --
( 31.38, 68.84);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 28.88,108.09) --
( 31.38,108.09);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 28.88,147.33) --
( 31.38,147.33);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 28.88,186.57) --
( 31.38,186.57);
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 28.88,225.82) --
( 31.38,225.82);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 31.38, 27.75) --
(406.94, 27.75);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{144,144,144}
\path[draw=drawColor,line width= 0.6pt,line join=round] ( 38.08, 25.25) --
( 38.08, 27.75);
\path[draw=drawColor,line width= 0.6pt,line join=round] (122.65, 25.25) --
(122.65, 27.75);
\path[draw=drawColor,line width= 0.6pt,line join=round] (207.22, 25.25) --
(207.22, 27.75);
\path[draw=drawColor,line width= 0.6pt,line join=round] (291.79, 25.25) --
(291.79, 27.75);
\path[draw=drawColor,line width= 0.6pt,line join=round] (376.36, 25.25) --
(376.36, 27.75);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.50] at ( 38.08, 19.80) {43};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.50] at (122.65, 19.80) {44};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.50] at (207.22, 19.80) {45};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.50] at (291.79, 19.80) {46};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 0.50] at (376.36, 19.80) {47};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (219.16, 6.94) {Mean rank of children of perm.~residents at $p=25$};
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (411.94,231.26);
\definecolor{drawColor}{RGB}{68,68,68}
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 11.89,127.01) {Effect of 1 year of exposure on income rank};
\end{scope}
\end{tikzpicture}
\caption{Neighborhood effects for New York and 90\% robust \acp{EBCI} for
children with parents at the $p=25$ percentile of the national income
distribution, plotted against mean outcomes of permanent residents. Gray
lines correspond to \acp{CI} based on unshrunk estimates represented by
circles, and black lines correspond to robust \acp{EBCI} based on \ac{EB}
estimates represented by squares that shrink towards a dotted regression
line based on permanent residents' outcomes. Baseline implementation as in
\Cref{sec:baseline-implementation}.}\label{fig:ch_cz25NY}
\end{figure}
To visualize these results, \Cref{fig:ch_cz25NY} plots the unshrunk 90\%
\acp{CI} based on the preliminary estimates, as well as robust \acp{EBCI} based
on \ac{EB} estimates for cities in the state of New York for children with
parents at the 25th percentile. While the \acp{EBCI} for large \acp{CZ} like New
York City or Buffalo are similar to the unshrunk \acp{CI}, they are much
tighter for smaller \acp{CZ} like Plattsburgh or Watertown, with point estimates
that shrink the preliminary estimates $Y_{i}$ most of the way toward the regression
line $X_{i}'\hat{\delta}$.
In summary, shrinkage allows us to considerably tighten the \acp{CI} based on
preliminary estimates. This is true even though the \acp{CI} effectively only use second moment constraints---imposing kurtosis constraints does not affect
the critical values in this application.