EconBase
← Back to paper

Confidence intervals for intentionally biased estimators

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

73,397 characters · 15 sections · 73 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Confidence intervals for intentionally biased estimators

\parindent=1.5em

\doublespacing

abstractWe propose and study three confidence intervals (CIs) centered at an estimator that is intentionally biased to reduce mean squared error. The first CI simply uses an unbiased estimator's standard error; compared to centering at the unbiased estimator, this CI has higher coverage probability for confidence levels above $91.7\%$, even if the biased and unbiased estimators have equal mean squared error. The second CI trades some of this “excess” coverage for shorter length. The third CI is centered at a convex combination of the two estimators to further reduce length. Practically, these CIs apply broadly and are simple to compute. Keywords: bias--variance tradeoff, coverage probability, mean squared error, smoothing JEL classification: C13 Disclosure statement: we (the authors) have no competing interests to declare.

\doublespacing

Introduction

We propose simple confidence intervals based on estimators that intentionally introduce bias in order to reduce mean squared error (MSE). Such estimators have a long history and continued popularity. For example, the classic averaging/shrinkage approach of JamesStein1961 was applied to simultaneous equation models by Maasoumi1978, which in turn inspired the recent asymptotic risk dominance results of Hansen2017 for a closely related 2SLS/OLS averaging estimator, while others have considered Stein-like GMM averaging with misspecified moments ChengLiaoShi2019,DiTraglia2016,Liu2022. As another example, the $L_2$ penalty of ridge regression HoerlKennard1970 has been modified to get lasso Tibshirani1996, bridge KnightFu2000, and SCAD FanLi2001, with additional optimality results by Zou2006 and ChetverikovEtAl2021, among others. A third category includes smoothed estimators. Even aside from nonparametric estimators that increase bias in order to minimize MSE, examples include Horowitz1992's (Horowitz1992) smoothed version of the binary choice maximum score estimator Manski1975 and GroeneboomEtAl2010's (GroeneboomEtAl2010) smoothed version of the nonparametric maximum likelihood estimator of the current status model GroeneboomWellner1992. Another example is KaplanSun2017's (KaplanSun2017) smoothed estimator of instrumental variables quantile regression ChernozhukovHansen2005,ChernozhukovHansen2006 and quantile regression KoenkerBassett1978; even stronger theoretical results for smoothed quantile regression are developed by FernandesEtAl2021 and HeEtAl2023, among others. As detailed in (ref), although our results do not apply universally to all these examples, they apply to several, with smoothed quantile regression as the example we detail throughout the paper.

In finite samples, ignoring the bias is problematic for the default confidence interval (CI) centered at the biased estimator and using its own standard error. Such a CI always “undercovers,” meaning coverage probability is below the desired confidence level. Sometimes asymptotic arguments are made that the bias goes to zero fast enough that the coverage probability increases toward $1-\alpha$ in the limit, but this does not fully address the finite-sample reality.

In this paper, we propose two CIs that are centered at the intentionally biased (lower MSE) estimator but attain at least the desired coverage probability, and a third CI centered at a convex combination of the unbiased and biased estimators. The first CI uses the unbiased estimator's standard error, while the second additionally uses the biased estimator's standard error in order to trade some “excess” coverage probability for shorter length. Both CIs have advantages over the other common approaches at the $95\%$ confidence level. Compared to using the biased estimator's CI mentioned above, our CIs do not suffer from undercoverage. Compared to CIs based on bias-correction, our CIs do not require any model-specific knowledge or bias estimation. Bias-corrected CIs work well in certain settings, but generally the bias can depend on parameters that are very challenging to estimate, like high-order derivatives of conditional densities of unobserved error terms as in (10) of KaplanSun2017 for smoothed instrumental variables quantile regression and as in Theorem 1 of FernandesEtAl2021 for smoothed quantile regression, and properly accounting for the extra uncertainty from bias estimation can make CIs longer than necessary, as noted by ArmstrongKolesar2020 in the context of regression discontinuity. That is, the relative practical convenience and statistical properties of our CIs compared to bias-corrected CIs depends on the specific model and setting. Compared to centering at the unbiased estimator, both our CIs have higher coverage probability, and our second CI has shorter length. By using the correlation between the unbiased and intentionally biased estimators, CI length can be further reduced by basing the CI on a convex combination of the two estimators.

Our first results focus on the coverage probability (CP) of our first CI compared with the benchmark CI that is centered at the unbiased estimator in addition to using its standard error. Our CI has the same length by construction, yet at the conventional $95\%$ confidence level, our CI has higher CP than the benchmark, even if the estimators' MSEs are equal (so the biased estimator is not actually better). Higher CP means fewer coverage errors (CI fails to contain true value). Although typically the goal is to minimize length subject to an upper bound on the coverage error rate, given a fixed CI length we prefer a lower error rate, or equivalently a higher CP. More generally, given equal MSE, our CP is higher than the benchmark whenever the confidence level is at least $91.7\%$, regardless of the magnitude of bias, and our worst-case CP is $90.00\%$ (rounded) at the $90\%$ confidence level, too. When the biased estimator has lower MSE, then our CP is even higher. However, at lower confidence levels, our CI's undercoverage can be significant, and CP can be near zero for confidence levels below $68.3\%$ when the bias is large. That is, our CI's unambiguous advantage over the benchmark CI at conventional confidence levels is not trivial because it disappears at lower confidence levels.

The results for our second CI show how it can further achieve shorter length than the benchmark. Essentially, it calibrates the critical value to achieve exact coverage probability if the two estimators have the same MSE, given the ratio of the two estimators' standard errors. Equivalently, “same MSE” can be interpreted as using the upper bound for bias like in ArmstrongKolesar2021, who in turn build upon the fixed-length CI results of Donoho1994. From our first results, at confidence level $95\%$, we know that if the biased estimator's standard error is strictly smaller than the unbiased estimator's standard error, then our first CI has CP strictly above $95\%$, even in the worst case of equal MSE. Thus, we can shorten the CI's length and still achieve at least $95\%$ CP, and even higher if the biased estimator's MSE is strictly below the unbiased estimator's MSE. This is what our second CI does. Compared to our first CI, one practical downside is the additional reliance on the biased estimator's standard error, but if such an estimate is readily available, then we recommend our second CI in practice because it has correct coverage while being more precise.

If the setting is strengthened to joint normality with know correlation in addition to known variances, then this second CI can be further shortened by using a convex combination of the unbiased and biased estimators. Our second CI is the special case with all weight on the biased estimator; by searching over all possible convex combinations, we can find the shortest possible CI. The benefit of this is highest when the correlation is lower and the original unbiased and biased estimators' variances are similar.

The work most similar to ours is from ArmstrongKolesar2021. The limiting experiment in their (13), which is an asymptotic version of the setting in their Section 2, is similar to our strongest setting with joint normality and known correlation. Indeed, our third CI based on a linear combination of the unbiased and biased estimators is inspired by their use of a linear estimator. However, our work also differs from theirs in several ways. First, our first CI relies only on the unbiased estimator's standard error, which is often easy and fast to compute, and even our second CI does not require knowledge of the correlation. Second, our bias bound is implied by the intentionally biased estimator's lower MSE, rather than require separate knowledge. Third, our setting always has an unbiased estimator (and thus valid CI) available; we show how at conventional confidence levels we can improve upon this, even though at other confidence levels this is not generally true. Of course, ArmstrongKolesar2021 provide results that apply to general settings different than ours, like generalized method of moments.

After describing the setup in (ref), in (ref) we characterize properties of our first proposed CI in the case when the biased estimator's MSE is identical to the unbiased estimator's MSE. We consider CP as a function of bias as well as the nominal confidence level. We characterize the cases for which our CI has higher CP than the CI centered at the unbiased estimator. Then in (ref) we show that if the biased estimator has strictly lower MSE, the CP becomes even higher, specifically if we reduce the bias while keeping the variance fixed. (ref) describes our second CI and establishes its properties. (ref) extends the second CI to an optimal convex combination of the unbiased and biased estimators. (ref) contain simulation and empirical results.

\paragraph{Notation and abbreviations} For notation, $\mathrm{N}(\mu , \sigma^2)$ is the normal distribution with mean $\mu$ and variance $\sigma^2$, and $\Phi(\cdot)$ and $\phi(\cdot)$ are respectively the CDF and PDF of the standard normal distribution $\mathrm{N}(0,1)$, whose $p$-quantile is denoted $z_p$; the sine, cosine, tangent, and secant functions are respectively $\sin(\cdot)$, $\cos(\cdot)$, $\tan(\cdot)$, and $\sec(\cdot)$. Acronyms used include those for confidence interval (CI), coverage probability (CP), cumulative distribution function (CDF), generalized method of moments (GMM), mean squared error (MSE), ordinary least squares (OLS), smoothed quantile regression (SQR), probability density function (PDF), quantile regression (QR), smoothed quantile regression (SQR), smoothly clipped absoluted deviation (SCAD), and two-stage least squares (2SLS).

Setup

We describe the estimators and their sampling distributions in (ref), some possible confidence intervals in (ref), and how well our setting applies to specific estimators in (ref).

Estimators and distributions

Given a scalar parameter $\theta$, consider the two estimators

equation[equation omitted — 165 chars of source]

where $s_1$ and $s_2$ are the respective standard errors and $b_2$ is the bias of $\hat\theta_2$. Estimator $\hat{\theta}_2$ is intentionally biased to reduce MSE:

equation[equation omitted — 119 chars of source]

The normal distributions in (ref) represent an asymptotic approximation. Our (ref) is very similar to settings used in other papers; for example, it is a weaker version of the main example setting on page 7 of ArmstrongEtAl2023, who instead study optimal estimation and more strongly assume joint normality with known covariance, and it is similar to the limiting experiment in (13) of ArmstrongKolesar2021. The goal of (ref) is simply to capture bias in a meaningful way, as opposed to asymptotic approximations in which the bias completely disappears. For example, given some $r>0$ and $q>0$, imagine $n^r(\hat\theta_2-\theta-n^{-q}\tilde{b}_2)\xrightarrow{d}\mathrm{N}(0,\tilde{s}_2^2)$, so approximately $\hat\theta_2\sim\mathrm{N}(\theta+n^{-q}\tilde{b}_2, \tilde{s}_2^2/n^{2r})$. Given a value of sample size $n$, our (ref) uses $b_2=n^{-q}\tilde{b}_2$ and $s_2^2=\tilde{s}_2^2/n^{2r}$. Some papers argue that asymptotically we can ignore the bias if $q>r$ so that the order $n^{-q}$ bias is of smaller order of magnitude than the $n^{-r}$ standard deviation, but we want capture the effect of bias that can be important in finite samples, while also allowing zero bias ($b_2=0$) as a special case. For $\hat\theta_1$, the finite-sample bias does not need to be exactly zero as long as (ref) provides a good approximation, meaning the bias of $\hat\theta_1$ is negligible compared to $b_2$, $s_1$, and $s_2$, like for the quantile regression estimator $\hat\theta_1$ in our simulation ((ref)). If $\hat\theta_2$ is a smoothed estimator, then $\hat\theta_1$ could be a highly under-smoothed estimator, although optimal bandwidths in such cases are beyond our scope.

The scalar parameter $\theta$ can be a summary of an underlying vector-valued or function-valued parameter as long as (ref) holds, but confidence sets or bands for non-scalar $\theta$ are beyond our scope.

In (ref), we treat some parameters as known and some as unknown. Viewing (ref) as an asymptotic approximation, “known” means “can be estimated consistently.” We always assume the unbiased estimator's standard error $s_1$ is known and that the biased estimator $\hat\theta_2$ can be computed. The first CI we propose does not require any further knowledge. Our CI in (ref) further requires $s_2$, and our CI in (ref) requires further strengthening (ref) to joint normality with known correlation. Although analytic forms of $s_2$ (or $s_1$) may be difficult to estimate, often bootstrap or subsampling PolitisRomano1994a can be used. We never require $b_2$ because estimating the bias is often infeasible or unreliable.

Confidence intervals

With confidence level $1-\alpha$, consider the following two-sided CIs:

equation[equation omitted — 298 chars of source]

where $z_{1-\alpha/2}$ is the $(1-\alpha/2)$-quantile of the standard normal distribution. The benchmark $\textrm{CI}_1$ is the usual CI using only the unbiased estimator. We focus on $\textrm{CI}_2$, which has not been proposed or studied previously. We include $\textrm{CI}_3$ only for completeness; it has not been proposed and is not good. As noted earlier, $\textrm{CI}_4$ has been proposed and justified by assuming the bias is asymptotically negligible.

Among the CIs in (ref), we propose using $\textrm{CI}_2$. To the best of our knowledge, this has not been proposed before, perhaps due to the counterintuitive mixing of one estimator for centering and another estimator's standard error. Our results characterize the coverage probability advantage of using $\textrm{CI}_2$ over the benchmark $\textrm{CI}_1$.

Unlike $\textrm{CI}_1$ and $\textrm{CI}_2$, both $\textrm{CI}_3$ and $\textrm{CI}_4$ always undercover. The $\textrm{CI}_3$ is centered at $\hat\theta_1$ like $\textrm{CI}_1$ but is shorter because $s_2<s_1$; because $\textrm{CI}_1$ has exact $1-\alpha$ coverage probability, $\textrm{CI}_3$ has less than $1-\alpha$ CP. Also, $\textrm{CI}_4$ covers $\operatorname{E}(\hat{\theta}_2)$ with probability $1-\alpha$, but $\operatorname{E}(\hat{\theta}_2)=\theta+b_2\ne\theta$, so it has less than $1-\alpha$ CP for $\theta$.

Additionally, in (ref), we propose CIs that convert some of the excess coverage probability of $\textrm{CI}_2$ into shorter length.

Applications

Our setting is a reasonable approximation for certain MSE-reducing estimators but not others, as we describe below.

Our original research question was simply: is $\textrm{CI}_2$ valid for smoothed instrumental variables quantile regression KaplanSun2017,sivqr? In that context, $\hat\theta_2$ is generally easier and faster to compute than $\hat\theta_1$, in addition to having lower MSE KaplanSun2017, although $s_1$ is also easy to compute ChernozhukovHansen2006. Both $\hat\theta_1$ and $\hat\theta_2$ are asymptotically normal. \footnote{ There are multiple algorithms for computing $\hat\theta_1$, but they all aim to solve for the same unsmoothed solution $\hat\theta_1$ (up to some smaller-order terms) and are consequently all asymptotically normal; for example, see ChernozhukovHansen2006, ChenLee2018, and KaidoWuthrich2021. } Our results suggest that for confidence levels above around $90\%$, $\textrm{CI}_2$ is not only valid but even better than $\textrm{CI}_1$. Further, our $\textrm{CI}_5$ ((ref)) can provide an even shorter valid CI, and $\textrm{CI}_6$ ((ref)) yet shorter.

Our results also apply well to the related setting of smoothed quantile regression (QR). There too, smoothing the indicator function in the moment conditions (estimating equations) introduces bias but reduces variance, resulting in an overall MSE reduction. Both the unsmoothed QR estimator and the smoothed QR estimator are asymptotically normal; for example, see KoenkerBassett1978, AngristEtAl2006, FernandesEtAl2021, and HeEtAl2023. Asymptotic MSE-optimal smoothing bandwidths are given by KaplanSun2017, FernandesEtAl2021, and HeEtAl2023. Although the difference is generally smaller than with instrumental variables QR, there can be settings (large sample size and/or number of regressors) where the unsmoothed QR estimator is significantly slower than smoothed QR, in which case $\textrm{CI}_1$ is not convenient; for example, see Figure 1(b) of HeEtAl2023. Our simulations in (ref) illustrate the finite-sample benefits of our proposed CIs in the context of smoothed QR.

With some caution, our results can be applied to other smoothed estimators even when the original unsmoothed estimator is not asymptotically normal. For example, consider the maximum score estimator Manski1975 or the nonparametric maximum likelihood estimator in the “current status” failure time model GroeneboomWellner1992, neither of which is asymptotically normal. Smoothing attains asymptotic normality in each case. Further, compared to using a very small smoothing bandwidth that incurs very little bias, increasing the bandwidth increases bias while reducing MSE, up to a point. In such cases, our $\hat\theta_1$ is a highly under-smoothed (small bandwidth, small bias) estimator, while our $\hat\theta_2$ is the estimator with the (approximate) MSE-optimal bandwidth. For example, from the results of Horowitz1992, Theorem 2(c) provides the asymptotic MSE-optimal bandwidth, which results in the asymptotically normal (but biased) distribution in Theorem 2(b) represented by our $\hat\theta_2$, while an undersmoothed bandwidth sets his $\lambda=0$ and removes the bias in Theorem 2(b), thus serving reasonably as our $\hat\theta_1$. Similarly, from GroeneboomEtAl2010, Theorem 3.5 (or Thm. 3.6 or Cor. 3.7, or Thm. 4.2, Thm. 4.3, or Cor. 4.4) provides the (biased) asymptotically normal distribution with the asymptotic MSE-optimal bandwidth rate for our $\hat\theta_2$, while undersmoothing with bandwidth going to zero faster than $n^{-1/5}$ (or $n^{-1/7}$) reduces the order of bias more and can serve as our $\hat\theta_1$.

With yet more caution, our results can apply to other estimators that reduce MSE in many but not all cases. For example, using a fixed data-generating process (pointwise asymptotics), the adaptive lasso and SCAD are asymptotically normal and as efficient as the infeasible “oracle” estimator knowing the true model (hence lower MSE than the corresponding unpenalized estimator); see Theorems 2 and 4 of Zou2006 and Theorem 2 of FanLi2001. However, among others, LeebPotscher2008b show that such oracle arguments do not hold uniformly (i.e., under all drifting sequences of data-generating processes), which in finite samples translates to portions of the parameter space where MSE is not lower than ordinary least squares; see their Theorem 2.1 and Section 3. More generally, there may be biased estimators that reduce MSE under certain conditions but not others. In practice, if such conditions seem plausible enough to use the intentionally biased estimator, then it seems reasonable to construct CIs under the same conditions. In such cases, while keeping $\hat\theta_1$ and $\textrm{CI}_1$ as a robustness check, the main results could show $\hat\theta_2$ with one of our proposed CIs.

Finally, our results are suggestive for Stein-like averaging estimators but leave gaps to fill in future work. Many modern averaging estimators like Hansen2017 or ChengLiaoShi2019 take a weighted average of a conservative estimator (like our $\hat\theta_1$) and an aggressive estimator that is biased but usually has lower variance (though not necessarily lower MSE). Usually the two estimators are jointly asymptotically normal. The infeasible oracle averaging estimator using the infeasible optimal weight (which is usually a constant) is thus itself asymptotically normal and fits our assumptions about $\hat\theta_2$. However, unlike in other contexts, the weight cannot be estimated consistently, so it is random even asymptotically. A random-weighted average of two normals is not normal, so this does not fit our $\hat\theta_2$; for example, see (9) of Hansen2017 or Lemma 4.2(a) of ChengLiaoShi2019. For weights near zero (or one), the resulting weighted average is nearly normal, and our conclusions may apply more generally if the distribution is “close enough” to normal (as seen in simulations like Figure 2 of Liu2015 for the plug-in averaging estimator), but formally extending our results to the general case of averaging estimators remains an important gap for future work to fill.

Coverage probability comparison: equal MSE

This section studies the CP of our proposed $\textrm{CI}_2$ when the biased estimator's MSE equals the unbiased estimator's MSE,

equation[equation omitted — 117 chars of source]

For intuition, consider the extreme case with the maximum possible bias satisfying (ref), $b_2=s_1$ and $s_2=0$. Note $\textrm{CI}_3$ and $\textrm{CI}_4$ both have length zero and thus $0\%$ CP. Because $s_2=0$, $\hat\theta_2=\theta+b_2=\theta+s_1$ is non-random, so $\textrm{CI}_2$ simplifies to

equation*[equation* omitted — 93 chars of source]

This $\textrm{CI}_2$ is also non-random and thus has either $0\%$ or $100\%$ CP. Specifically, $\textrm{CI}_2$ includes the true $\theta$ if and only if $s_1-z_{1-\alpha/2}s_1\le0$, which simplifies to $z_{1-\alpha/2}\ge1$ or approximately $1-\alpha\ge0.683$. Thus, in this special case, $\textrm{CI}_2$ is better than $\textrm{CI}_1$ if the confidence level is above $68.3\%$, but worse if below $68.3\%$. We will more generally characterize when $\textrm{CI}_2$ is better than $\textrm{CI}_1$, in terms of both the bias $b_2$ and the confidence level $1-\alpha$.

More generally, to satisfy the equal MSE in (ref), let

equation[equation omitted — 108 chars of source]

where $\sin(\cdot)$ and $\cos(\cdot)$ are the sine and cosine functions. Intuitively, $t$ is a transformed version of the bias: $b_2$ is increasing in $t$, with $b_2=0$ when $t=0$, up to $b_2=s_1$ when $t=\pi/2$. The standard deviation $s_2$ moves in the opposite direction, decreasing in $t$ from $s_2=s_1$ at $t=0$ down to $s_2=0$ at $t=\pi/2$. The earlier special case implicitly set $t=\pi/2$ to get $s_2=0$ and $b_2=s_1$.

The coverage probability of $\textrm{CI}_2$ depends on $t$ and the critical value $z$:

align[align omitted — 1,120 chars of source]

where $\Phi(\cdot)$ is the standard normal CDF, $\tan(\cdot)$ is the tangent function, and $\sec(\cdot)=1/\cos(\cdot)$ is the secant function. Note $\textrm{CP}(\cdot,z)$ is an even function because

align[align omitted — 325 chars of source]

using the facts that $\sec(-t)=\sec(t)$, $\tan(-t)=-\tan(t)$, and $\Phi(x)=1-\Phi(-x)$. Thus, we let $t>0$ without loss of generality.

figure[figure omitted — 169 chars of source]

(ref) shows $\textrm{CI}_2$'s coverage probability as a function of $t$, for several fixed $z$. Specifically, each plot shows $\textrm{CP}(t, z_{1-\alpha/2})$ over $0\le t\le\pi/2$, with the different plots showing confidence levels $1-\alpha=0.99,0.95,0.90,0.81,0.69,0.68$. The horizontal line in each plot shows $1-\alpha$, which is the coverage probability of the benchmark $\textrm{CI}_1$, as well as the value of $\textrm{CP}(0,z_{1-\alpha/2})$ because $t=0$ means $b_2=0$ and $s_2=s_1$, so $\hat\theta_2\overset{d}{=}\hat\theta_1$. At the usual confidence level $95\%$ (and $99\%$), the coverage probability of $\textrm{CI}_2$ is always higher than the nominal $1-\alpha$, which is also the CP of $\textrm{CI}_1$. At confidence level $90\%$, $\textrm{CI}_2$ has a worst-case $0.89995$ CP, technically below the nominal $1-\alpha$ but visually and practically indistinguishable; at other $t$, CP exceeds the nominal $1-\alpha=0.9$. Numerical analysis finds that $1-\alpha\ge0.917$ is the threshold for $\textrm{CI}_2$ having at least $1-\alpha$ CP uniformly over $t\in[0,\pi/2]$. At uncommon confidence level $81\%$, the worst-case CP is $0.8007$, and CP exceeds $1-\alpha$ for larger $t$ closer to $\pi/2$. At the unusual confidence levels $69\%$ and $68\%$, there can be severe undercoverage, especially for $68\%$ because $z_{1-\alpha/2}<1$; this reflects the results from the earlier special case with $t=\pi/2$, in which CP was $100\%$ if $z_{1-\alpha/2}\ge1$ (confidence level above $68.3\%$) but $0\%$ if $z_{1-\alpha/2}<1$ (confidence level below $68.3\%$). In sum, $\textrm{CI}_2$ is uniformly better than $\textrm{CI}_1$ for confidence levels $95\%$ and $99\%$, and practically uniformly better at confidence level $90\%$, but can be much worse with unusually low confidence levels.

(ref) summarizes these results.

theoremConsider the setting of (ref). The coverage probability of $\textrm{CI}_1$ equals the confidence level $1-\alpha$. For $1-\alpha\ge0.917$, $\textrm{CI}_2$ has coverage probability strictly above $1-\alpha$ when $b_2\ne0$. For $1-\alpha=0.9$ and any $b_2$, $\textrm{CI}_2$ has coverage probability of at least $90.00\%$ (rounded), increasing to $100\%$ as $\lvertb_2\rvert\to s_1$.
proofThe result for $\textrm{CI}_1$ is well known. The coverage probability of $\textrm{CI}_2$ is derived in (ref) as $\Phi(z\sec(t)-\tan(t))-\Phi(-z\sec(t)-\tan(t))$, in terms of the parameterization in (ref). Numerical study of this function leads to the stated results, including the worst-case CP with $1-\alpha=0.9$ being $0.899953$, which rounds to $0.9000$.

Coverage probability comparison: lower MSE

Our finding that $\textrm{CI}_2$ is better than $\textrm{CI}_1$ for the most common confidence levels extends from the special case of $\textrm{MSE}(\hat\theta_2)=\textrm{MSE}(\hat\theta_1)$ to the general case of $\textrm{MSE}(\hat\theta_2) \le \textrm{MSE}(\hat\theta_1)$. If $\hat\theta_2$ has strictly lower MSE than $\hat\theta_1$, then we can imagine a hypothetical estimator $\hat\theta_B$ ($B$ for “bound”) with the same variance but higher bias than $\hat\theta_2$, such that $\hat\theta_B$ has the same MSE as $\hat\theta_1$. As shown in (ref), adding bias like this decreases CP, yet from (ref) we know the CI centered at $\hat\theta_B$ still has higher CP than the benchmark at conventional confidence levels. Thus, allowing $\hat\theta_2$ to have lower MSE than $\hat\theta_1$ further strengthens the advantages of our proposed $\textrm{CI}_2$ compared to the benchmark $\textrm{CI}_1$.

corollaryIn (ref), the same results hold after relaxing the equal-MSE condition in (ref) to the weakly-lower-MSE condition in (ref).
proofConsider the hypothetical estimator $\hat\theta_B$ ($B$ for “bound”) with the same variance as $\hat\theta_2$ but larger magnitude bias $b_B$ such that its MSE equals that of the unbiased $\hat\theta_1$: $b_B=\pm\sqrt{s_1^2-s_2^2}$. From (ref), the CP of $\textrm{CI}_2$ is at least that of the CI $\hat\theta_B\pm z_{1-\alpha/2}s_1$, to which the results of (ref) apply.
lemmaLet $\hat\theta_B\sim\mathrm{N}(\theta+b_B,s_2^2)$ and $\hat\theta_2\sim\mathrm{N}(\theta+b_2,s_2^2)$ with $\lvertb_2\rvert<\lvertb_B\rvert$. For any $s_1>0$ and $z>0$, $\textrm{CI}_2$ ($\hat\theta_2\pm z s_1$) has higher coverage probability than $\hat\theta_B\pm z s_1$: \begin{equation*} \operatorname{P}(\hat\theta_2 - z s_1 \le\theta\le \hat\theta_2 + z s_1) > \operatorname{P}(\hat\theta_B - z s_1 \le\theta\le \hat\theta_B + z s_1) . \end{equation*}
proofThe CP for each estimator can be written in terms of the standard normal CDF $\Phi(\cdot)$, and then (ref) can be applied. First, \begin{align*} \operatorname{P}(\hat\theta_2 - z s_1 \le \theta \le \hat\theta_2 + z s_1) &= \operatorname{P}(\theta - z s_1 \le \hat\theta_2 \le \theta + z s_1) \&= \operatorname{P}\biggl(\frac{-z s_1 - b_2}{s_2} \le \overbrace{\frac{\hat\theta_2-\theta-b_2}{s_2}}^{\equiv Z_1\sim\mathrm{N}(0,1)} \le \frac{z s_1 - b_2}{s_2} \biggr) \&= \Phi\biggl(\frac{ z s_1 - b_2}{s_2}\biggr) -\Phi\biggl(\frac{-z s_1 - b_2}{s_2}\biggr) ,\\ \operatorname{P}(\hat\theta_B - z s_1 \le \theta \le \hat\theta_B + z s_1) &= \operatorname{P}(\theta - z s_1 \le \hat\theta_B \le \theta + z s_1) \&= \operatorname{P}\biggl(\frac{-z s_1 - b_B}{s_2} \le \overbrace{\frac{\hat\theta_B-\theta-b_B}{s_2}}^{\equiv Z_2\sim\mathrm{N}(0,1)} \le \frac{z s_1 - b_B}{s_2} \biggr) \&= \Phi\biggl(\frac{ z s_1 - b_B}{s_2}\biggr) -\Phi\biggl(\frac{-z s_1 - b_B}{s_2}\biggr) . \end{align*} The result follows by applying (ref) with $d\equiv z s_1 / s_2$, $a\equiv -b_2/s_2$, and $b\equiv -b_B/s_2$, noting that the (ref) condition $\lverta\rvert<\lvertb\rvert$ is satisfied because $\lvertb_2\rvert<\lvertb_B\rvert$.

Shorter CI based on bias bound

Given the estimator sampling distributions and the MSE bound in (ref), we can bound the magnitude of the bias. Then, similar in spirit to the strategy of ArmstrongKolesar2021, we construct a symmetric two-sided CI to have exact coverage when the bias attains the upper bound. By (ref), the coverage is even better with smaller magnitude bias, so this CI has at least $1-\alpha$ coverage probability for all levels of bias satisfying the MSE bound. Note that this relationship between CP and bias is opposite that of $\textrm{CI}_2$: here, CP is highest when $b_2=0$, decreasing to $1-\alpha$ as $\lvertb_2\rvert$ increases to the bound, whereas for $\textrm{CI}_2$ the CP is $1-\alpha$ at $b_2=0$ and (for conventional confidence levels) increases toward $100\%$ as $\lvertb_2\rvert$ increases.

This CI trades some of the higher CP of $\textrm{CI}_2$ for shorter length. For example, at confidence level $95\%$ when $\textrm{CI}_2$ has CP strictly above $95\%$ for any non-zero bias $b_2\ne0$, this alternative CI is strictly shorter than $\textrm{CI}_2$ while maintaining at least $95\%$ CP. This is generally preferable to $\textrm{CI}_2$. The only additional requirement is the reliance on $s_2$, the biased estimator's standard error. Usually a consistent estimator is available, corresponding to known $s_2$ in our setting, but if not then $\textrm{CI}_2$ can still be used.

CI definition and properties

Formally, the new CI is defined as follows. To facilitate comparison, it is defined the same as $\textrm{CI}_2$ but replacing $z_{1-\alpha/2}$ with $\tilde{z}_{1-\alpha/2}$, defined implicitly using the $\textrm{CP}(t,z)$ function from (ref):

equation[equation omitted — 181 chars of source]

This $\tilde{z}_{1-\alpha/2}$ sets CP to $1-\alpha$ when $\lvertb_2\rvert$ satisfies $b_2^2+s_2^2=s_1^2$ (equal MSE), as detailed in the proof of (ref).

theoremGiven the sampling distributions in (ref) with known $s_1$ and $s_2$ but unknown $b_2$, for $\textrm{CI}_5$ defined in (ref), coverage probability is at least $1-\alpha$ for any $b_2$ satisfying the MSE inequality in (ref): \[ \inf_{-\sqrt{s_1^2-s_2^2} \le b_2 \le \sqrt{s_1^2-s_2^2}} \operatorname{P}(\theta\in\textrm{CI}_5) = 1-\alpha , \] with the minimum $1-\alpha$ attained when $b_2=\pm\sqrt{s_1^2-s_2^2}$.
proofConsider $\textrm{CI}_5$ with a generic critical value $\tilde{z}$, $\hat\theta_2\pm\tilde{z}s_1$. Because this has the same structure as $\textrm{CI}_2$, its CP is the same as derived in (ref), only replacing $z$ with $\tilde{z}$: \begin{equation} \operatorname{P}(\hat\theta_2-\tilde{z}s_1 \le \theta \le \hat\theta_2+\tilde{z}s_1) = \Phi(( \tilde{z}s_1-b_2)/s_2) -\Phi((-\tilde{z}s_1-b_2)/s_2) . \end{equation} This is an even function of $b_2$ because the value does not change when replacing $b_2$ with $-b_2$: using $\Phi(x)=1-\Phi(-x)$, \begin{align}\notag &\Phi(( \tilde{z}s_1-(-b_2))/s_2) -\Phi((-\tilde{z}s_1-(-b_2))/s_2) \&= \notag 1-\Phi((-\tilde{z}s_1-b_2)/s_2) -[1-\Phi(( \tilde{z}s_1-b_2)/s_2)] \&= \Phi(( \tilde{z}s_1-b_2)/s_2) -\Phi((-\tilde{z}s_1-b_2)/s_2) . \end{align} Below, we show that the choice $\tilde{z}=\tilde{z}_{1-\alpha/2}$ defined in (ref) sets this CP equal to $1-\alpha$ in the bounding case of equal MSE. Equal MSE means $b_2^2+s_2^2=s_1^2$, which is the same as using $t=\cos^{-1}(s_2/s_1)$ with $s_2=s_1\cos(t)$ and $b_2=s_1\sin(t)$ as in (ref), where using $b_2>0$ is without loss of generality because CP is an even function of $b_2$ as shown in (ref). That is, given $b_2^2+s_2^2=s_1^2$ and $\tilde{z}=\tilde{z}_{1-\alpha/2}$ from (ref), (ref) implies \begin{align*} \operatorname{P}(\theta\inCI_5) = \Phi(( \tilde{z}_{1-\alpha/2}s_1-b_2)/s_2) -\Phi((-\tilde{z}_{1-\alpha/2}s_1-b_2)/s_2) = CP\bigl( \cos^{-1}(s_2/s_1), \tilde{z}_{1-\alpha/2} \bigr) , \end{align*} which by (ref) equals $1-\alpha$. Applying (ref) with $b_B=\pm\sqrt{s_1^2-s_2^2}$ and $z=\tilde{z}_{1-\alpha/2}$, this $1-\alpha$ CP is the lower bound when $\lvertb_2\rvert<\lvertb_B\rvert$ given the same $s_1$ and $s_2$.

In terms of length, $\textrm{CI}_5$ is shorter than $\textrm{CI}_2$ essentially whenever $\textrm{CI}_2$'s CP is above $1-\alpha$, which at the $95\%$ confidence level is always. That is, compared to $\textrm{CI}_2$, $\textrm{CI}_5$ trades some coverage probability for shorter length, which is usually desired as long as CP remains at least $1-\alpha$, which (ref) says it does. The cases where $\textrm{CI}_5$ is longer are those in which $\textrm{CI}_2$ may undercover, which are lower than the conventionally used confidence levels anyway.

theoremGiven the sampling distributions in (ref) with known $s_1$ and $s_2$ (with $s_2<s_1$) but unknown $b_2$ satisfying the MSE inequality in (ref), $\textrm{CI}_5$ defined in (ref) is strictly shorter than $\textrm{CI}_2$ defined in (ref) for all confidence levels above $91.7\%$. For confidence level $90\%$, $\textrm{CI}_5$ is no more than $0.014\%$ longer than $\textrm{CI}_2$, and can be significantly shorter.
proofGiven their similar structures, $\textrm{CI}_5$ is shorter than $\textrm{CI}_2$ if and only if $\tilde{z}_{1-\alpha/2}<z_{1-\alpha/2}$. Consider the implicit definition of $\tilde{z}_{1-\alpha/2}$ in (ref). Given that the $\textrm{CP}(\cdot,\cdot)$ function is increasing in its second argument (wider CI has higher CP), if we plug in $z_{1-\alpha/2}$ and get a CP value above $1-\alpha$, then we know $\tilde{z}_{1-\alpha/2}<z_{1-\alpha/2}$. According to (ref), for $1-\alpha\ge0.917$, $\textrm{CP}(t,z_{1-\alpha/2})>1-\alpha$ for all $0<t\le\pi/2$. Thus, for $1-\alpha\ge0.917$, $\textrm{CP}\bigl(\cos^{-1}(s_2/s_1),z_{1-\alpha/2}\bigr)>1-\alpha$ for all $s_2<s_1$, implying $\tilde{z}_{1-\alpha/2}<z_{1-\alpha/2}$ and thus $\textrm{CI}_5$ being shorter than $\textrm{CI}_2$. For $1-\alpha=0.9$, numerical analysis find the function $\textrm{CP}(t,z_{1-\alpha/2})$ is minimized at $t=0.359$ with value $0.899953$; given that worst-case $t$, $0.9=\textrm{CP}(0.359,1.6451)$, so the solution is $\tilde{z}_{1-\alpha/2}=1.6451$ rather than $z_{1-\alpha/2}=1.6449$, and (without rounding first) $\tilde{z}_{1-\alpha/2}/z_{1-\alpha/2}=1.0001369$, or $0.014\%$ larger.
figure[figure omitted — 235 chars of source]

(ref) shows the length advantage of $\textrm{CI}_5$ over $\textrm{CI}_2$ (which has the same length as $\textrm{CI}_1$) as a function of $s_2/s_1$ at confidence level $1-\alpha=0.95$. If $s_2$ is near $s_1$, then the lengths are very similar. If $s_2$ is a small fraction of $s_1$, then $\textrm{CI}_5$ can be below $60\%$ the length of $\textrm{CI}_2$. Even for intermediate values of $s_2/s_1$, $\textrm{CI}_5$ offers a meaningful advantage in length. The yet shorter $\textrm{CI}_6$ is derived in (ref).

Example

Consider the properties of $\textrm{CI}_5$ and $\textrm{CI}_2$ in the following example with confidence level $95\%$. Let $s_1=1$ and $s_2=s_1/2$. From (ref), $\textrm{CI}_2 \equiv \hat\theta_2\pm z_{1-\alpha/2}s_1$, using only $s_1$ and not $s_2$. From (ref), using both $s_1$ and $s_2$, $\textrm{CI}_5 \equiv \hat\theta_2\pm \tilde{z}_{1-\alpha/2}s_1$ with $\tilde{z}_{0.975}$ solving $0.95=\textrm{CP}(\cos^{-1}(s_2/s_1),\tilde{z}_{0.975})$, yielding $\tilde{z}_{0.975}=1.69$.

Regardless of the bias $b_2$, $\textrm{CI}_5$ is around $14\%$ shorter than $\textrm{CI}_2$ (which is the same length as $\textrm{CI}_1$) because $\tilde{z}_{0.975}/z_{0.975}\approx0.86$.

The CP of each CI further depends on the bias $b_2$. For example, consider $b_2=s_1/2$, so the biased estimator has strictly lower MSE than the unbiased estimator: \[ \textrm{MSE}(\hat\theta_2) = b_2^2+s_2^2 = s_1^2/2<s_1^2 = \textrm{MSE}(\hat\theta_1) . \] Using (ref), the CP of $\textrm{CI}_2$ is

align*[align* omitted — 262 chars of source]

This is better than $\textrm{CI}_1$ because $\textrm{CI}_2$ is the same length and has only $0.2\%$ coverage error rate, instead of $5\%$ error rate. The CP of $\textrm{CI}_5$ uses the same formula as for $\textrm{CI}_2$ above but with the lower $\tilde{z}_{0.975}$:

align*[align* omitted — 80 chars of source]

Despite being $14\%$ shorter than $\textrm{CI}_2$, $\textrm{CI}_5$ still has CP well above $1-\alpha=0.95$ because $\textrm{MSE}(\hat\theta_2)$ is well below $\textrm{MSE}(\hat\theta_1)$. In this case, it seems well worth trading the $99.8\%$ CP of $\textrm{CI}_2$ for the slightly lower $99.1\%$ CP of $\textrm{CI}_5$ in return for a $14\%$ shorter interval.

If instead there is the largest possible bias of $b_2=\sqrt{3}s_1/2$, which implies $\textrm{MSE}(\hat\theta_2)=\textrm{MSE}(\hat\theta_1)$, then the CP of both intervals is smaller. The CP of $\textrm{CI}_2$ is \[ \Phi( 2z_{0.975}-\sqrt{3}) -\Phi(-2z_{0.975}-\sqrt{3}) = 0.986 , \] and the CP of $\textrm{CI}_5$ drops to the nominal $95\%$ level:

equation*[equation* omitted — 96 chars of source]

In that case, $\textrm{CI}_5$ trades the entire “excess” CP of $\textrm{CI}_2$ in return for the $14\%$ shorter length.

Shortest CI based on convex combination estimator

If we strengthen (ref) to joint normality with known correlation

equation[equation omitted — 387 chars of source]

then further improvement may be possible by taking a convex combination of $\hat\theta_1$ and $\hat\theta_2$,

equation[equation omitted — 111 chars of source]

Specifically, our $\textrm{CI}_6$ defined in (ref) is centered at this $\hat\theta_{3w}$, with the known value $w$ chosen by the user to minimize the corresponding CI's length subject to the desired worst-case coverage probability. As seen below, this optimal $w$ depends on the standard errors $s_1$ and $s_2$ (and $\rho$). We do not allow more general linear combinations with $w>1$ or $w<0$ because then the MSE of the linear combination may be higher than that of $\hat\theta_1$; similarly, for general linear combination weights that do not sum to $1$, the MSE of the linear combination may be (arbitrarily) higher than that of $\hat\theta_1$ without additional assumptions of bounds on $\theta$.

This setting is similar to Section 4.1 of ArmstrongKolesar2021. \footnote{Thanks to an anonymous reviewer for making this connection and suggesting we consider convex combinations.} Our (ref) are similar to their (13) and corresponding linear estimator, except our bound on bias $b_2$ comes from the assumed MSE improvement of $\hat\theta_2$, and our convex combination weight is not required to satisfy any other condition like their (5).

Practically, the strengthening to joint normality is rarely restrictive, but the additional reliance on the correlation $\rho$ may incur extra computation time and estimation error.

The strategy here is to show that the convex combination $\hat\theta_{3w}$ is normally distributed with known variance, so we can apply the same strategy as for $\textrm{CI}_5$. One difference is that we know more about the maximum bias, because we know $b_2^2\le s_1^2-s_2^2$ and that the bias of $\hat\theta_{3w}$ is $wb_2$ for known $w$. This provides a tighter bias bound than simply using $(wb_2)^2\le s_1^2 - s_{3w}^2$ from the MSE inequality.

The new CI is derived as follows. As formally stated and proved in (ref), the CI $\hat\theta_{3w}\pm s_1z$ with general critical value $z$ has coverage probability

equation[equation omitted — 282 chars of source]

using the parameterization $b_2=s_1\sin(t)$ and $s_2=s_1\cos(t)$ from (ref). Similar to the definition of $\tilde{z}_{1-\alpha/2}$ in (ref), define $\tilde{z}_{w,1-\alpha/2}$ to solve

equation[equation omitted — 130 chars of source]

Define

equation[equation omitted — 188 chars of source]

That is, $\tilde{z}_{w,1-\alpha/2}$ is calibrated to achieve the desired coverage probability for any $w$, and $w^*$ is then chosen to minimize the CI's length.

theoremGiven (ref) and the definitions in (ref), the coverage probability of $\textrm{CI}_6$ is at least $1-\alpha$, and the length of $\textrm{CI}_6$ is less than or equal to the length of $\textrm{CI}_5$.
proofBy construction, $\tilde{z}_{w,1-\alpha/2}$ in (ref) sets the coverage probability equal to $1-\alpha$ exactly in the equal-MSE case. By (ref) (replacing $s_2$ with $s_{3w}$, and $b_2$ with $wb_2$, and $b_B=\pm w\sqrt{s_1^2-s_2^2}$), if $\textrm{MSE}(\hat\theta_2)<\textrm{MSE}(\hat\theta_1)$ strictly, then the CP is even higher. The previous $\textrm{CI}_5$ is the special case with $w=1$; given that $w^*$ minimizes length over $0\le w\le1$, $\textrm{CI}_6$ is shorter than $\textrm{CI}_5$ when $w^*<1$ and has equal length when $w^*=1$.

(ref) (from earlier) shows how the benefit of $\textrm{CI}_6$ depends on the correlation $\rho=\operatorname{Corr}(\hat\theta_1,\hat\theta_2)$ and the ratio $s_2/s_1$. If $\rho$ is near $1$, then $\hat\theta_1$ contains little information not already contained in $\hat\theta_2$, so taking a convex combination cannot improve much over simply using $\hat\theta_2$. If $s_2/s_1$ is small, then $\textrm{CI}_5$ already has much shorter length than $\textrm{CI}_2$, so the CI-optimal convex combination will put most or all of the weight on $\hat\theta_2$ anyway; that is $w^*\approx1$, so $\textrm{CI}_6\approx\textrm{CI}_5$. Conversely, if $s_2/s_1$ is closer to one and the correlation $\rho$ is farther below one, then using $\hat\theta_{3w}$ can shorten the CI length considerably. For example, with $\rho=0.1$ and $s_2/s_1=1$, the length of $\textrm{CI}_6$ is only $74\%$ of the length of $\textrm{CI}_5$.

Simulation

The following simulation illustrates the finite-sample properties of our proposed CIs in the context of smoothed quantile regression (SQR). We use the R code for the smoothed instrumental variables quantile regression estimator of KaplanSun2017, which inclues SQR as a special case (when the instrument vector equals the regressor vector), using their automatic plug-in version of the bandwidth formula in their Proposition 2 whose rate minimizes the asymptotic mean squared error (see also their Section 5 including Theorem 7). Additional theoretical results for SQR are provided by FernandesEtAl2021, including allowing a stochastic bandwidth sequence (Theorem 5(S)) and uniformity over ranges of the quantile index, and HeEtAl2023 further extend SQR to a growing number of regressors. Code in R R.core to replicate our simulation results is available on the first author's website. \footnote{https://kaplandm.github.io/}

Within each simulation replication, we do the following. First, data are generated by

equation[equation omitted — 258 chars of source]

with $\beta_0=1$ and $\theta=2$. In this case, the slope of the conditional $\tau$-quantile function equals $\theta$ for all $0<\tau<1$. Second, we use nonparametric pairs bootstrap to estimate the standard errors $s_1$ and $s_2$ from (ref) and the correlation $\rho$ from (ref). Within each bootstrap replication, we compute estimator $\hat\theta_1$ using \lstinline{rq()} from the \lstinline{quantreg} package R.quantreg and compute estimator $\hat\theta_2$ using the code based on KaplanSun2017; then, using all bootstrapped estimates, we compute the standard deviations and correlation. Third, with $1-\alpha=0.95$, we compute $\textrm{CI}_1$ and $\textrm{CI}_2$ from (ref), $\textrm{CI}_5$ from (ref), and $\textrm{CI}_6$ from (ref). As a crude alternative to help avoid undercoverage due to estimation error in $\rho$, we compute $\textrm{CI}_6^s$ using value $(1+\rho)/2$.

table[table omitted — 3,063 chars of source]

(ref) shows the simulated coverage probability and median length of each CI, with the following patterns. First, as expected, among our four proposed CIs, both CP and median length decrease from left to right: $\textrm{CI}_2$ is highest (for both CP and length), then $\textrm{CI}_5$, then $\textrm{CI}_6^s$, and finally $\textrm{CI}_6$ has lowest CP and length. That is, the CIs progressively trade more of $\textrm{CI}_2$'s “extra” CP (above $1-\alpha$) for shorter length. Second, compared to $\textrm{CI}_1$, $\textrm{CI}_2$ has the same length (by construction) and higher CP. That is, $\textrm{CI}_2$ makes fewer coverage errors than $\textrm{CI}_1$, without sacrificing length. Third, compared to $\textrm{CI}_1$, $\textrm{CI}_5$ always has higher CP as well as always having shorter length. That is, it successfully trades some of $\textrm{CI}_2$'s extra CP for shorter length. Fourth, compared to $\textrm{CI}_1$, $\textrm{CI}_6^s$ always has weakly higher CP while (by construction) reducing length even more than $\textrm{CI}_5$. Fifth, compared to $\textrm{CI}_1$, $\textrm{CI}_6$ has weakly higher CP in all but two cases while always achieving the shortest length among all the CIs. Despite having lower CP than $\textrm{CI}_1$ in two cases (and undercovering in one of those), $\textrm{CI}_6$ fixes $\textrm{CI}_1$'s undercoverage in three other cases and increases CP at least $0.5$ percentage points in yet four other cases. That is, although $\textrm{CI}_6$ does not achieve uniformly better CP than the benchmark $\textrm{CI}_1$, arguably it still improves CP overall, even while reducing length by 5--10%.

Empirical illustration

Like the simulation, our empirical illustration also uses smoothed quantile regression, also with replication code on the first author's website. \footnote{https://kaplandm.github.io/} We use the 1980 and 1990 Census data from AngristEtAl2006, and like their Figure 2A we focus on the coefficient on years of schooling in a log earnings regression that also includes experience and its square; see Section 4 of AngristEtAl2006 for details about the data. They use data for U.S.-born men aged 40--49, and we further restrict to Black men. Like them, we multiply the schooling coefficient by $100$ to get the approximate return to schooling as a percent. As in the simulation, we use confidence level $1-\alpha=0.95$ and $399$ bootstrap replications.

table[table omitted — 1,780 chars of source]

(ref) reports $\textrm{CI}_1$, $\textrm{CI}_5$, and $\textrm{CI}_6$, for both Census years and quantile levels $\tau\in\{0.1,0.5,0.9\}$. Briefly, although not relevant to our methodology, we note some similarities and differences with Figure 2A of AngristEtAl2006, whose sample includes Black men but is dominated by white men. Both our results show the same upward shift in schooling coefficients from 1980 to 1990, but their results show the $\tau=0.9$ schooling coefficient increasing much faster than (and in 1990 well exceeding) the $\tau\in\{0.1,0.5\}$ coefficients, whereas our results for Black men show a consistently lower return to schooling at the $0.9$-quantile level than at lower quantiles.

(ref) also reports the length of our $\textrm{CI}_6$ as a percent of the benchmark $\textrm{CI}_1$, with modest but consistent improvements of $3$--$7$% that are similar to the $5$--$10$% improvements from the simulation. Unfortunately, as usual, we cannot know the true coverage in empirical examples, so we cannot see the other benefit of increased coverage probability shown in the theory and simulations, where $\textrm{CI}_5$ always had higher coverage probability than $\textrm{CI}_1$, and $\textrm{CI}_6$ almost always.

Conclusion

In practice, given an intentionally biased estimator that has at least as good mean squared error as the corresponding unbiased estimator, as a default choice we recommend $\textrm{CI}_5$, our CI centered at the biased estimator using the unbiased estimator's standard error and with critical value calibrated to be shorter than a CI using the unbiased estimator's standard error. This CI is easy to compute, avoids the undercoverage from using the biased estimator's standard error, and for conventional confidence levels is better than the benchmark of using on the unbiased estimator and its standard error, having both shorter length and higher coverage probability. In some cases our $\textrm{CI}_6$ is even shorter by using the correlation to solve for the convex combination estimator that minimizes CI length while maintaining coverage. Our results are not model-specific, so they apply not only to all the estimators cited in the introduction but also any future estimators that introduce bias in order to reduce mean squared error. If the biased estimator's standard error is not reliably estimated, then using the usual critical value (and the unbiased estimator's standard error) also improves upon other confidence intervals and is even easier to compute in practice.

In future work, it would be valuable to extend these results more precisely to non-normal distributions, specifically those arising in averaging estimators with random weights like in Hansen2017 and ChengLiaoShi2019. It would also be interesting to consider confidence sets for vector-valued parameters, as in papers going back to Stein1962, as well as uniform confidence bands based on function-valued estimators (such as nonparametric regression or smoothed CDF estimators) with asymptotic Gaussian process limits: under what conditions on the bias function $b(\cdot)$ and covariance functions $s_1(\cdot,\cdot)$ and $s_2(\cdot,\cdot)$ would a uniform confidence band have higher coverage probability when centered at the (more) biased estimator?

Acknowledgments

We thank Editor Esfandiar Maasoumi and the anonymous associate editor and referees for their help improving this paper, as well as Alyssa Carlson for feedback on multiple drafts.

\singlespacing

thebibliography{36} \bibitem[{Angrist, Chernozhukov, and Fern\'{a}ndez-Val(2006)}]{AngristEtAl2006} Angrist, Joshua, Victor Chernozhukov, and Iv\'{a}n Fern\'{a}ndez-Val. 2006. \newblock “Quantile Regression under Misspecification, with an Application to the {U.S.} Wage Structure.” \newblock Econometrica 74 (2):539--563. \newblock URL https://doi.org/10.1111/j.1468-0262.2006.00671.x. \bibitem[{Armstrong, Kline, and Sun(2023)}]{ArmstrongEtAl2023} Armstrong, Timothy B., Patrick Kline, and Liyang Sun. 2023. \newblock “Adapting to Misspecification.” \newblock Working paper available at https://arxiv.org/abs/2305.14265. \bibitem[{Armstrong and Koles{\'a}r(2020)}]{ArmstrongKolesar2020} Armstrong, Timothy B. and Michal Koles{\'a}r. 2020. \newblock “Simple and honest confidence intervals in nonparametric regression.” \newblock Quantitative Economics 11 (1):1--39. \newblock URL https://doi.org/10.3982/QE1199. \bibitem[{Armstrong and Koles{\'a}r(2021)}]{ArmstrongKolesar2021} ---------. 2021. \newblock “Sensitivity analysis using approximate moment condition models.” \newblock Quantitative Economics 12 (1):77--108. \newblock URL \texttt{https://doi.org/10.3982/QE1609}. \bibitem[{Chen and Lee(2018)}]{ChenLee2018} Chen, Le-Yu and Sokbae Lee. 2018. \newblock “Exact computation of {GMM} estimators for instrumental variable quantile regression models.” \newblock \emph{Journal of Applied Econometrics} 33 (4):553--567. \newblock URL \texttt{https://doi.org/10.1002/jae.2619}. \bibitem[{Cheng, Liao, and Shi(2019)}]{ChengLiaoShi2019} Cheng, Xu, Zhipeng Liao, and Ruoyao Shi. 2019. \newblock “On Uniform Asymptotic Risk of Averaging {GMM} Estimators.” \newblock \emph{Quantitative Economics} 10 (3):931--979. \newblock URL \texttt{https://doi.org/10.3982/QE711}. \bibitem[{Chernozhukov and Hansen(2005)}]{ChernozhukovHansen2005} Chernozhukov, Victor and Christian Hansen. 2005. \newblock “An {IV} Model of Quantile Treatment Effects.” \newblock \emph{Econometrica} 73 (1):245--261. \newblock URL \texttt{https://www.jstor.org/stable/3598944}. \bibitem[{Chernozhukov and Hansen(2006)}]{ChernozhukovHansen2006} ---------. 2006. \newblock “Instrumental quantile regression inference for structural and treatment effect models.” \newblock \emph{Journal of Econometrics} 132 (2):491--525. \newblock URL \texttt{https://doi.org/10.1016/j.jeconom.2005.02.009}. \bibitem[{Chetverikov, Liao, and Chernozhukov(2021)}]{ChetverikovEtAl2021} Chetverikov, Denis, Zhipeng Liao, and Victor Chernozhukov. 2021. \newblock “On Cross-Validated Lasso in High Dimensions.” \newblock \emph{Annals of Statistics} 49 (3):1300--1317. \newblock URL \texttt{https://doi.org/10.1214/20-AOS2000}. \bibitem[{DiTraglia(2016)}]{DiTraglia2016} DiTraglia, Francis J. 2016. \newblock “Using invalid instruments on purpose: Focused moment selection and averaging for {GMM}.” \newblock \emph{Journal of Econometrics} 195 (2):187--208. \newblock URL \texttt{https://doi.org/10.1016/j.jeconom.2016.07.006}. \bibitem[{Donoho(1994)}]{Donoho1994} Donoho, David L. 1994. \newblock “Statistical Estimation and Optimal Recovery.” \newblock \emph{Annals of Statistics} 22 (1):238--270. \newblock URL \texttt{https://doi.org/10.1214/aos/1176325367}. \bibitem[{Fan and Li(2001)}]{FanLi2001} Fan, Jianqing and Runze Li. 2001. \newblock “Variable Selection via Nonconcave Penalized Likelihood and its Oracle Properties.” \newblock \emph{Journal of the American Statistical Association} 96 (456):1348--1360. \newblock URL \texttt{https://doi.org/10.1198/016214501753382273}. \bibitem[{Fernandes, Guerre, and Horta(2021)}]{FernandesEtAl2021} Fernandes, Marcelo, Emmanuel Guerre, and Eduardo Horta. 2021. \newblock “Smoothing Quantile Regressions.” \newblock \emph{Journal of Business & Economic Statistics} 39 (1):338--357. \newblock URL \texttt{https://doi.org/10.1080/07350015.2019.1660177}. \bibitem[{Groeneboom, Jongbloed, and Witte(2010)}]{GroeneboomEtAl2010} Groeneboom, Piet, Geurt Jongbloed, and Birgit I. Witte. 2010. \newblock “Maximum smoothed likelihood estimation and smoothed maximum likelihood estimation in the current status model.” \newblock \emph{Annals of Statistics} 38 (1):352--387. \newblock URL \texttt{https://doi.org/10.1214/09-AOS721}. \bibitem[{Groeneboom and Wellner(1992)}]{GroeneboomWellner1992} Groeneboom, Piet and Jon A. Wellner. 1992. \newblock \emph{Information Bounds and Nonparametric Maximum Likelihood Estimation}. \newblock Birkh{\"a}user. \bibitem[{Hansen(2017)}]{Hansen2017} Hansen, Bruce E. 2017. \newblock “A {Stein}-Like {2SLS} Estimator.” \newblock \emph{Econometric Reviews} 36 (6--9):840--852. \newblock URL \texttt{https://doi.org/10.1080/07474938.2017.1307579}. \bibitem[{He et al.(2023)He, Pan, Tan, and Zhou}]{HeEtAl2023} He, Xuming, Xiaoou Pan, Kean Ming Tan, and Wen-Xin Zhou. 2023. \newblock “Smoothed quantile regression with large-scale inference.” \newblock \emph{Journal of Econometrics} 232 (2):367--388. \newblock URL \texttt{https://doi.org/10.1016/j.jeconom.2021.07.010}. \bibitem[{Hoerl and Kennard(1970)}]{HoerlKennard1970} Hoerl, Arthur E. and Robert W. Kennard. 1970. \newblock “Ridge Regression: Biased Estimation for Nonorthogonal Problems.” \newblock \emph{Technometrics} 12 (1):55--67. \newblock URL \texttt{https://www.jstor.org/stable/1267351}. \bibitem[{Horowitz(1992)}]{Horowitz1992} Horowitz, Joel L. 1992. \newblock “A smoothed maximum score estimator for the binary response model.” \newblock \emph{Econometrica} 60 (3):505--531. \newblock URL \texttt{https://www.jstor.org/stable/2951582}. \bibitem[{James and Stein(1961)}]{JamesStein1961} James, W. and Charles Stein. 1961. \newblock “Estimation with Quadratic Loss.” \newblock In \emph{Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability}, vol. 1. Berkeley, CA: University of California Press, 361--379. \newblock URL \texttt{https://projecteuclid.org/euclid.bsmsp/1200512173}. \bibitem[{Kaido and W{\"u}thrich(2021)}]{KaidoWuthrich2021} Kaido, Hiroaki and Kaspar W{\"u}thrich. 2021. \newblock “Decentralization estimators for instrumental variable quantile regression models.” \newblock \emph{Quantitative Economics} 12 (2):443--475. \newblock URL \texttt{https://doi.org/10.3982/QE1440}. \bibitem[{Kaplan(2022)}]{sivqr} Kaplan, David M. 2022. \newblock “Smoothed instrumental variables quantile regression.” \newblock \emph{Stata Journal} 22 (2):379--403. \newblock URL \texttt{https://doi.org/10.1177/1536867X221106404}. \bibitem[{Kaplan and Sun(2017)}]{KaplanSun2017} Kaplan, David M. and Yixiao Sun. 2017. \newblock “Smoothed Estimating Equations for Instrumental Variables Quantile Regression.” \newblock \emph{Econometric Theory} 33 (1):105--157. \newblock URL \texttt{https://doi.org/10.1017/S0266466615000407}. \bibitem[{Knight and Fu(2000)}]{KnightFu2000} Knight, Keith and Wenjiang Fu. 2000. \newblock “Asymptotics for Lasso-Type Estimators.” \newblock \emph{Annals of Statistics} 28 (5):1356--1378. \newblock URL \texttt{https://www.jstor.org/stable/2674097}. \bibitem[{Koenker(2023)}]{R.quantreg} Koenker, Roger. 2023. \newblock \emph{quantreg: Quantile Regression}. \newblock URL \texttt{https://CRAN.R-project.org/package=quantreg}. \newblock R package version 5.96. \bibitem[{Koenker and Bassett(1978)}]{KoenkerBassett1978} Koenker, Roger and Gilbert Bassett, Jr. 1978. \newblock “Regression Quantiles.” \newblock \emph{Econometrica} 46 (1):33--50. \newblock URL \texttt{https://www.jstor.org/stable/1913643}. \bibitem[{Leeb and P{\"o}tscher(2008)}]{LeebPotscher2008b} Leeb, Hannes and Benedikt M. P{\"o}tscher. 2008. \newblock “Sparse estimators and the oracle property, or the return of {Hodges'} estimator.” \newblock \emph{Journal of Econometrics} 142 (1):201--211. \newblock URL \texttt{https://doi.org/10.1016/j.jeconom.2007.05.017}. \bibitem[{Liu(2015)}]{Liu2015} Liu, Chu-An. 2015. \newblock “Distribution theory of the least squares averaging estimator.” \newblock \emph{Journal of Econometrics} 186 (1):142--159. \newblock URL \texttt{https://doi.org/10.1016/j.jeconom.2014.07.002}. \bibitem[{Liu(2022)}]{Liu2022} Liu, Xin. 2022. \newblock “Averaging estimation for instrumental variables quantile regression.” \newblock Working paper available at \texttt{https://xinliu16.github.io/}. \bibitem[{Maasoumi(1978)}]{Maasoumi1978} Maasoumi, Esfandiar. 1978. \newblock “A Modified {Stein-like} Estimator for the Reduced Form Coefficients of Simultaneous Equations.” \newblock \emph{Econometrica} 46 (3):695--703. \newblock URL \texttt{https://www.jstor.org/stable/1914241}. \bibitem[{Manski(1975)}]{Manski1975} Manski, Charles F. 1975. \newblock “Maximum score estimation of the stochastic utility model of choice.” \newblock \emph{Journal of Econometrics} 3 (3):205--228. \newblock URL \texttt{https://doi.org/10.1016/0304-4076(75)90032-9}. \bibitem[{Politis and Romano(1994)}]{PolitisRomano1994a} Politis, Dimitris N. and Joseph P. Romano. 1994. \newblock “Large sample confidence regions based on subsamples under minimal assumptions.” \newblock \emph{Annals of Statistics} 22 (4):2031--2050. \bibitem[{{R Core Team}(2023)}]{R.core} {R Core Team}. 2023. \newblock \emph{R: A Language and Environment for Statistical Computing}. \newblock R Foundation for Statistical Computing, Vienna, Austria. \newblock URL \texttt{https://www.R-project.org}. \bibitem[{Stein(1962)}]{Stein1962} Stein, C. M. 1962. \newblock “Confidence Sets for the Mean of a Multivariate Normal Distribution.” \newblock \emph{Journal of the Royal Statistical Society: Series B} 24 (2):265--285. \newblock URL \texttt{https://www.jstor.org/stable/2984224}. \bibitem[{Tibshirani(1996)}]{Tibshirani1996} Tibshirani, Robert J. 1996. \newblock “Regression shrinkage and selection via the lasso.” \newblock \emph{Journal of the Royal Statistical Society: Series B} 58 (1):267--288. \newblock URL \texttt{www.jstor.org/stable/2346178}. \bibitem[{Zou(2006)}]{Zou2006} Zou, Hui. 2006. \newblock “The Adaptive Lasso and Its Oracle Properties.” \newblock \emph{Journal of the American Statistical Association} 101 (476):1418--1429. \newblock URL \texttt{https://doi.org/10.1198/016214506000000735}.