Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
100,269 characters · 16 sections · 84 citation commands
Local Asymptotic Power of Honest Confidence Intervals
\noindentKeywords: honest inference; bias-aware confidence intervals; local asymptotic power; dead zone; partial identification.\\ \noindentJEL classification: C12, C14, C23.
A growing body of applied work reports confidence intervals that are deliberately conservative to untestable biases, where an estimate is widened by a worst-case bound on its bias so that coverage holds uniformly over a class of data-generating processes the data cannot rule out. These honest or bias-aware intervals are now standard for difference-in-differences under possible violations of parallel trends rambachan2023more, instrumental variables under possible exclusion failures conley2012plausibly, regression discontinuity armstrong2018optimal,armstrong2020simple, and panels with potentially weak factors armstrong2022robust. This paper asks what such intervals cost in power, and shows the cost can be total. Indeed, at the usual parametric rate, an honest interval may reject no local alternative with probability approaching one. The loss is not a flaw of any one procedure; a minimax argument shows it is intrinsic to honesty against an untestable bias. The same argument cuts the other way: the standard bias-aware interval is rate-optimal among honest procedures, and the directional critical value of armstrong2018optimal attains the sharp honest power frontier. The results therefore price the honesty requirement itself, rather than fault the constructions that meet it. Taking a local-alternative approach makes the size--power trade-off of conservative tests precise, and the duality between the length of a confidence interval and the power of its dual test is itself classical pratt1961length.
The organising idea is that this power loss is not, in essence, about conservatism or bias-awareness, but about interval width that does not shrink. A level-$\alpha$ confidence interval whose width fails to vanish fast enough relative to the alternatives traces out a set of positive volume, and against any alternative in the interior of that set the dual test is powerless. Only at the boundary can one-sided power survive. Partial identification is the limiting case, in which the width converges to a fixed identified set that does not shrink at all imbens2004confidence,stoye2009more. Honest inference under an untestable bias is the same phenomenon at a slower rate, the bound playing the role of the identified-set width. Seen this way, the three regimes developed below are one statement: local power degenerates exactly when the suitably rescaled interval width has positive limiting volume. Conservatism against an untestable bias is the leading economic instance, and the one developed in detail.
Assume the analyst possesses an estimator $\hat\beta$, and upper bound $\sup_\theta |\hat{B}|$, such that,
where $\sup_\theta |\hat{B}|$ is a function of untestable user-defined parameters $\theta$ and may be much larger than the true bias, $\hat{B}$. The term $O_p(n^{-\epsilon})$ represents error from noise, e.g. in a parametric model $\epsilon = 1/2$. Confidence intervals built from such a worst-case bias bound are the bias-aware intervals of armstrong2018optimal,armstrong2020simple and imbens2019optimized. That the bound is governed by $\theta$, which cannot be learned from the data and so must be supplied by the analyst, is fundamental rather than a deficiency of any particular estimator low1997nonparametric,bahadur1956nonexistence.
For $\hat{B} = o_p\left(n^{-\epsilon}\right)$, there is negligible asymptotic bias. For $\hat{B} \gtrsim c\cdot n^{-\epsilon}$, non-negligible asymptotic bias persists. Note, the analyst only has $\sup_\theta |\hat{B}|$, not $\hat B$, so cannot use this for debiasing. Also note, the bound $\sup_\theta |\hat{B}|$ may be very conservative - indeed this framework admits estimators with negligible true asymptotic bias, $\hat B$, but non-negligible bounds $\sup_\theta |\hat{B}|$.
The object that organises everything is the test's dead zone: the set of true effects so close to the hypothesised value that an honest test cannot tell them apart, rejecting them no more often than its nominal size. A dead zone is the unavoidable price of the bound. Because the worst-case $\sup_\theta|\hat B|$ lets the bias move the centre of $\hat\beta$ by as much as $\pm\sup_\theta|\hat B|$, a true effect and the null produce the same data whenever they lie within twice that bound of each other, exactly so in the leading applied settings, such that no honest test can separate them (Figure (ref)). The half-width of this zone, twice the bias bound, is the paper's headline disclosure.
A two-number example makes the accounting concrete. Let the bound be $\sup_\theta|\hat B|=1$, suppose the realised bias also equals one, and ignore sampling noise, so $\hat\beta=\beta_0+1$. The honest interval $\hat\beta\pm1=[\beta_0,\beta_0+2]$ covers the truth exactly at its edge and retains every value out to twice the bound above the truth: its far half defends not the world that generated the data but the observationally identical twin with bias $-1$, in which the truth would be $\beta_0+2$ (Figure (ref)(b)). Here the width is well spent, with bias genuinely at the bound, the full allowance is needed for coverage. The genuine cost appears when the bound is conservative, sharpest at $\hat B=0$: the instrument in fact valid, the trend truly parallel. Then $\hat\beta=\beta_0$, yet the interval $[\beta_0-1,\beta_0+1]$ still retains every value within one bound of the truth, spending the full width $2\sup_\theta|\hat B|$ where an oracle who knew $\hat B=0$ would spend only the sampling width $2z_{1-\alpha/2}\,se(\hat\beta)$. Because the bias is untestable, the saving can never be identified or harvested: the analyst pays for the worst admissible world in every world, and the dead zone is that payment expressed in power rather than length.
Whether the zone matters asymptotically turns on how fast the bound shrinks relative to the local alternatives. If the bound shrinks faster, the dead zone closes and conservatism is asymptotically free; if the two shrink at the same rate, it persists at a fixed, explicit width and the power loss is bounded (Theorem (ref), Proposition (ref)); and if the bound shrinks more slowly, typical in the parametric case, the dead zone engulfs every local alternative and local power is zero (Theorems (ref)--(ref)). A minimax argument shows the zone is intrinsic to honesty itself, not to any one interval (Theorem (ref)). The same picture recurs across the four settings: the factor-model, difference-in-differences, and instrumental-variables cases are developed now, and regression discontinuity, the matched-rate case, from Section (ref) onward.
Rejection rates for Example (ref), the debiased estimator in armstrong2022robust, in Figure (ref) are taken from simulations in Section (ref). The CIs conservative to weak factors have absolute upper bounds on bias of order $\bar{B} = O_p(\max\{N,T\}/n)$, where $n = NT$. Take $T = c\cdot N^\tau$ for $c\in (0,1)$ and $\tau \in [0,1]$, i.e. $N = \max\{N,T\}$, such that $\sqrt{n}$ multiplied by bias is of order $\sqrt{n}\bar{B} = O_p(N^{(1 - \tau)/2})$. Then, for $\tau = 1$, i.e. $N$ and $T$ proportional to each other, scaled sup bias is $\sqrt{n}\bar B = O_p(1)$, i.e. may be bounded by a constant but is not growing asymptotically. This still leads to finite sample power loss, but with CIs converging still at the parametric rate. However, when $\tau \neq 1$, scaled asymptotic sup bias is asymptotically divergent, and the test based on bias-aware confidence intervals retains no local asymptotic power when local power is defined at the $\sqrt{n}$ rate. This power loss is without gain when true bias is asymptotically negligible, e.g. under a strong true factor model.
The same structure underlies two settings common in applied work.
The confidence intervals studied here belong to the literature on honest and bias-aware inference, in which an interval is widened by a worst-case bound on the estimator's bias so that coverage holds uniformly over a class of data-generating processes. The feature central to this paper is that the asymptotic width of such intervals is governed by a user-specified bound that indexes this class and is, in general, not consistently estimable from the data.
In nonparametric regression this bound is a smoothness or magnitude constant. donoho1994statistical characterises fixed-length intervals for linear functionals through the modulus of continuity over a convex parameter space, and li1989honest shows that honest intervals must account for the bias that smoothing induces. armstrong2018optimal and armstrong2020simple develop the modern bias-aware construction, replacing the usual critical value with one that incorporates the worst-case bias. Applications and extensions include regression discontinuity with a discrete running variable kolesar2018inference, finite-sample minimax regression discontinuity under a bound on the second derivative imbens2019optimized, and inference under a bound on the magnitude of control coefficients armstrong2020biasaware,noack2024bias. The bias-aware intervals of armstrong2022robust studied in Example (ref) are of exactly this form, with the role of the user-specified bound played by the assumed number of weak factors.
That the governing bound cannot be learned from the data is fundamental rather than incidental. low1997nonparametric and cailow2004adaptation establish that, within standard smoothness classes, one cannot adapt the length of an honest interval to the unknown regularity of the underlying function while maintaining uniform coverage. The length is pinned down by the worst case, echoing the classical impossibility result of bahadur1956nonexistence. An early form of this floor appears in the one-dimensional subproblem at the heart of our limit experiment: in the bounded normal mean, the minimax half-length of a fixed-length $1-\alpha$ interval, affine or nonlinear, equals the bias bound itself whenever the bound does not exceed $z_{1-\alpha}$ noise units donoho1994statistical. This is a length statement at the scale of one bias budget; its power counterpart, at twice that scale and uniform over all honest procedures, is the subject of Section (ref). Conservatism against an untestable bias therefore carries an unavoidable cost in length, and hence, as emphasised here, in power.
A parallel strand applies the same logic to inference that is robust to model misspecification, where the analyst specifies how far the model may depart from a baseline. conley2012plausibly relax the instrument exclusion restriction by a user-chosen amount and widen intervals accordingly; armstrong2021sensitivity conduct sensitivity analysis in approximate moment-condition models under a bound on local misspecification; and andrews2017measuring and bonhomme2022minimizing study the sensitivity of estimates to misspecified moments. rambachan2023more construct honest intervals for difference-in-differences that remain valid under bounded violations of parallel trends, and explicitly characterise the power of the resulting tests as a function of the assumed bound, a design motivated in part by the low power of conventional pre-trend tests roth2022pretest.
The trade-off between interval length and the power of the associated test is classical. pratt1961length relates the expected length of a confidence interval to the power of the dual family of tests. Not all of these procedures face the asymptotic power properties studied herein, and several make the trade-off explicit li1989honest,low1997nonparametric,rambachan2023more,noack2024bias. The contribution of this paper is to make the trade-off precise in the bias-aware setting, where the relevant length is not a fixed multiple of the standard error but a user-specified bias bound whose asymptotic order may exceed the estimator's usual convergence rate. The weak-factor example is convenient precisely because the power trade-off is a direct function of how conservative the analyst chooses to be, namely the number of factors specified as weak.
This section begins with a simple and instructional result regarding power of the dual test under bias-aware intervals. This establishes the breadth of the paper's contribution. Subsection (ref) establishes the power loss is an inherent feature of the honest requirement rather than the particular procedure proposed by bias-aware intervals. That is, the power loss is unavoidable once honesty under undetectable bias is deemed strictly necessary.
Assume that $\bar B := \sup_\theta |\hat B|$ is bounded. Notation $\Theta_p $ implies an exact asymptotic bound: $A_n = \Theta_p(B_n) \iff A_n = O_p(B_n)$ and $A_n^{-1} = O_p(B_n^{-1})$.
Assumption (ref) upper bounds the supremum over asymptotic bias to be constant, and admits settings with non-negligible asymptotic bias if $\delta < \epsilon$.
Local alternative hypotheses are defined per newey1994large as follows:
Throughout the power analysis $\beta_0$ is the null value and the data are generated under it, $\beta=\beta_0$: the null is correctly specified. It is then asked how well the dual test excludes local perturbations of this truth, the local alternatives $\beta_0+\xi\,n^{-\rho}$, rejecting such a value when it falls outside confidence interval $\mathcal B_n$. The local power is $\Pr_{\beta_0}\{\,\beta_0+\xi\,n^{-\rho}\notin\mathcal B_n\,\}$, and zero local power means $\mathcal B_n$ covers these perturbations with probability approaching one. Since the dead zone has the same width about any value, the applications of Section (ref) report it about the null of no effect simply by convention. Fixed alternative hypotheses are admitted in Definition (ref) by setting $\rho = 0$.
Theorem (ref) states an immediate implication of scaling up confidence intervals at a rate faster than the worst-case bias shrinks: the intervals diverge. Its hypothesis is satisfied by a range of honest procedures in current use armstrong2022robust,rambachan2023more,conley2012plausibly, so the power loss is inherited wholesale rather than by any one construction. Section (ref) shows that the loss it induces is intrinsic to honesty under user defined biases.
The results of this subsection share one mechanism. After rescaling, the local power function of any interval is the non-coverage function of its limiting rescaled form. This is the organising claim of the introduction: the limit of the rescaled width determines local power.
Theorems (ref)--(ref) below are instances, differing only in the limit $(L,U)$ of the bias-aware interval: $(-\infty,+\infty)$ in the divergent regime $\delta<\rho$ (zero power); the constants $\big((c-1)\bar b,(c+1)\bar b\big)$ at the matched rate with negligible noise (a step-function power); and Gaussian endpoints when the noise survives. Proposition (ref) is the same statement with the identified set as the limit.
Theorem (ref) states that whenever $\hat B = o_p(\bar B)$, local alternatives that are asymptotically small with respect to the bias bound, $\bar B$, are not rejected with probability approaching one (wpa1). Whilst the set of local alternatives, $\Xi_n$, can grow with sample size, hence be unbounded in the limit, the rate this can grow is potentially very slow. Thus, it is easier to interpret this as a compact set.
Theorem (ref) shows zero local power under a possibly slack bound, $\hat B = o_p(\bar B)$. In the sequel, Theorem (ref) shows that this power loss persists even when the sign of the bias is correct and the bound is tight, provided the alternatives shrink faster than the bias bound, i.e. $\delta<\rho$. Knowing the sign of bias is of course outside the capabilities of bias-aware confidence intervals, since then a simple debias could be performed. It is nonetheless a useful exercise to study power properties had these been pinned down in the confidence interval widenings.
Theorem (ref) isolates the role of the bound's tightness. When $\delta<\rho$ the scaled bound $n^{\rho}\bar B$ diverges, so even a bound that exactly equals the bias in the limit with known sign cannot restore two-sided power. The test is asymptotically powerless against local alternatives on the same side as the bias, and detects only opposite-signed alternatives. The best that pinning down bias can do is provide one-sided power. For any $c\in(-1,1)$ all local power is lost, exactly as in Theorem (ref).
The matched-rate case $\delta=\rho$, where the alternative shrinks at precisely the rate of the bias bound, is the knife edge between the divergent regime above and the standard regime $\delta>\rho$. Theorem (ref) pins down the resulting non-degenerate power loss.
Theorem (ref) displays the cost of conservatism. When $\rho=\delta$ the bound neither washes out the alternative ($\delta<\rho$) nor becomes negligible ($\delta>\rho$), and contributes a constant $2\bar b$ to width. When $\rho<\epsilon$ the sampling error is of smaller order, so the limiting power is a step function. Alternatives in $[(c-1)\bar b,(c+1)\bar b]$ are never rejected, while every alternative outside is rejected wpa1. This window has half-width $\bar b$ and depends on the realised bias $c$, which is untestable; its union over $c\in[-1,1]$ is $[-2\bar b,2\bar b]$, the worst-case dead zone of Section (ref).
Conservatism thus leads to several scenarios. Widths exactly twice the limiting bound, against which the test has no power, and setting $\bar b=0$ recovers the oracle that rejects every $\xi\neq0$. At the rate $\rho=\epsilon$ the sampling error survives, and the power curve is smooth. It is the usual Gaussian power function, with the rejection threshold inflated from $z_{1-\alpha/2}\sigma$ to $\bar b+z_{1-\alpha/2}\sigma$ and recentered by the bias $c\bar b$. The added $\bar b$ is the exact, finite power cost of bias-aware conservatism.
The results above concern a single procedure, i.e. the bias-aware interval $\mathcal{B}_n$ and the test it induces. It is shown here that every test that is honest faces the same three-regime power behaviour, so the loss documented in Section (ref) is a property of the inference problem rather than of only the bias-aware confidence interval widening.
Let $P_{\beta,\theta}^{(n)}$ denote the data-generating process with target coefficient $\beta$ and untestable nuisance $\theta\in\mathcal{C}_n$, and let $\mathcal{P}_n = \{P_{\beta,\theta}^{(n)}: \beta\in\mathbb{R},\, \theta\in\mathcal{C}_n\}$ be the class over which coverage is required. A test is a measurable $\phi_n\in\{0,1\}$, rejecting $H_0:\beta=\beta_0$ when $\phi_n=1$.
Definition (ref) is the hypothesis test counterpart of uniform coverage. Recall that an interval $\mathcal{B}_n$ is honest over $\{\mathcal{P}_n\}$ at level $1-\alpha$ if it covers each process's own parameter value uniformly,
Pairing $\mathcal{B}_n$ with its dual tests $\phi_n(b):=\mathbbm{1}\{b\notin\mathcal{B}_n\}$, covering the true value is non-rejection of the true value. For every process, $\Pr_{P_{\beta,\theta}}\{\beta\in\mathcal{B}_n\}=1-\Pr_{P_{\beta,\theta}}\{\phi_n(\beta)=1\}$. Taking the infimum and using $\inf(1-x)=1-\sup x$ together with $\liminf_n(1-s_n)=1-\limsup_n s_n$, condition (ref) is equivalent to uniform asymptotic size control of the dual tests over the whole class,
Retaining only the term $\beta=\beta_0$ in (ref), which can only lower the supremum, yields Definition (ref) for the dual test $\phi_n(\beta_0)=\mathbbm{1}\{\beta_0\notin\mathcal{B}_n\}$, and equally the size control of the dual test $\phi_n(b)=\mathbbm{1}\{b\notin\mathcal{B}_n\}$ at any value $b$. Thus honest coverage (ref) implies Definition (ref). The proof of Theorem (ref) invokes this size control at the tested perturbation $A_n$, coverage of $A_n$ under the confounding law whose parameter is $A_n$, which any honest interval provides, so the bound applies to every honest interval.
The primitive driving the bound is that the class can confound a shift in $\beta$ with a change in the nuisance. This is the information-theoretic dual of the worst-case bias $\bar B$, where $\bar B$ measures the largest shift in $\hat\beta$ the class can induce. The condition below measures the largest shift in $\beta$ the class can hide in the likelihood. It is a genuine assumption on the information in the class, and is not implied by Assumption (ref), which restricts only the bias of $\hat\beta$. Define the total variation distance,
Assumption (ref) says that a coefficient shift of up to twice the worst-case bias can be absorbed by the class, with detectability governed by the fraction $\Delta/2\bar B$ of that reach used. The factor of two is Figure (ref)(b): the adversary spends the bias in both laws, pushing the observable centre of the true law up by $\bar B$ and that of the confounding law down by $\bar B$, so the two laws approach each other from both sides. The observable centres are $\beta_0+m_0$ and $\beta_0+\Delta+m_1$ with $|m_0|,|m_1|\le\bar B$, and they coincide only if
with equality at $m_0=+\bar B$, $m_1=-\bar B$. Beyond $2\bar B$ the centres cannot meet, and the separation between the laws is forced positive. A shift that is a vanishing fraction of $2\bar B$ is asymptotically undetectable ($\tau(0)=0$), while a shift equal to the full reach remains only borderline detectable, $\tau(1)<1-\alpha$, so that even a full-reach shift leaves the power of any honest test bounded away from one. The restriction to $[0,1-\alpha)$ is costless in the regular case: Section (ref) shows that when the confounding is measured through the modulus, the Neyman--Pearson analogue of this bound has $\tau(u)=\Phi(s(u)-z_{1-\alpha})-\alpha<1-\alpha$ because $\Phi<1$. The calibration in terms of the ratio $\Delta/2\bar B$, rather than $\Delta n^{\epsilon}$, is exactly the statement that the worst-case bias corresponds to $O(1)$ units of statistical separation, i.e. the modulus-of-continuity relation $\bar B \asymp \omega(n^{-1/2})$ of donoho1994statistical. The function $\tau$ bundles the shape of the modulus with the two-point affinity and need not be specified beyond its qualitative properties for the rate conclusions.
In the two leading location-type settings the confounding is not merely bounded but exact: the class contains observationally identical laws whose targets differ by up to twice the bias bound, so Assumption (ref) holds with $\tau\equiv0$ and the dead zone follows with no regularity conditions at all.
The dead zone in these settings is thus a statement about observational equivalence, not about asymptotic approximation: within $2\bar b$ of the truth there are laws the data literally cannot tell apart. Equivalently, $2\bar B$ is the diameter of the local identified set.
The bound in part (ii) of Theorem (ref) is sharp under regularity, and the sharp value is an explicit Gaussian power function with a power “dead zone”. Consider the noise-surviving matched rate $\delta=\rho=\epsilon$ and localise $\beta=\beta_0+h\,n^{-\epsilon}$ where $h\in\mathbb{R}$ is the local coefficient of Definition (ref) treated as a free coordinate that indexes the experiment. The perturbation of interest is the point $h=\xi$.
Assumption (ref) is the exact-confounding form of Assumption (ref). In the limit a shift in $\beta$ of up to $\bar B$ is matched by a nuisance shift, so the data cannot separate the two, and honesty reduces to size control over the contaminated null mean, $\sup_{|\nu|\le\bar b}\Pr_{N(\nu,\sigma^2)}\{\phi=1\}\le\alpha$. In the location settings of Lemma (ref) it holds exactly: the confounding laws coincide, so the estimator carries all the usable information and the sufficiency clause is innocuous. In genuinely nonparametric settings such as regression discontinuity it does not hold for the full data: the confounding there is approximate rather than exact and the separation $s(u)$ of Section (ref) is strictly positive, so the estimator is not asymptotically sufficient, and a directional honest test tuned to a specific perturbation can retain some power inside the window. In those settings Proposition (ref) describes the experiment generated by the estimator, hence every procedure in the bias-aware (affine) class, while the frontier over all honest tests is the smooth $\Phi(s(\xi)-z_{1-\alpha})$ of Proposition (ref).
In the same limit the bias-aware interval excludes the perturbation $\xi$ with worst-case power $\Phi(s(\xi)-z_{1-\alpha/2})$. This reproduces the dead zone $|\xi|\le 2\bar b$ exactly and matches the frontier $\pi_\alpha$ up to the fixed critical-value gap $z_{1-\alpha/2}-z_{1-\alpha}$, which the directional value $\mathrm{cv}_\alpha$ of armstrong2018optimal removes (there $\mathrm{cv}_\alpha(t)-t\to z_{1-\alpha}$).
The worst case over the nuisance is the operative reading precisely because the bias is untestable. Under any single law with bias $c\bar b$ the undetected window is $[(c-1)\bar b,(c+1)\bar b]$ by Theorem (ref)(i), which has half-width $\bar b$ about the bias-shifted centre, which is what the shaded regions of the simulations realise. The dead zone is the union of these windows over the admissible biases,
the least-favourable bias being chosen per perturbation ($m_0=\operatorname{sign}(\xi)\,\bar b$): a perturbation such as $\xi=\tfrac32\bar b$ is excluded wpa1 under a zero-bias law, yet never excluded under any bias $c\ge\tfrac12$; since the analyst cannot know $c$, only the union can be promised. By Proposition (ref), at the matched rate the dead zone of every honest interval contains the window $|\xi|\le 2\bar b$, and for the bias-aware interval with $\rho<\epsilon$ it is exactly $[-2\bar b,2\bar b]$. In the units of the estimand this is a half-width $2\bar B$ (as $\bar b=\lim_n n^{\delta}\bar B$), the diameter of the local identified set.
The scope follows Lemma (ref): in the exactly confounded settings the statement binds every honest procedure with no further conditions, while in the approximately confounded settings it binds the bias-aware class, with the all-tests frontier given by Proposition (ref). Reporting this single number alongside the interval is the disclosure recommended in Section (ref). Figure (ref) makes the per-law/union distinction visible in a single plot. Sweeping the realised bias over its admissible range produces the family of sliding windows, whose lower envelope, the empirical counterpart of the $\inf_\theta$ in Definition (ref), is flat on exactly $\pm2\bar b$. In contrast, the average over a uniformly drawn bias first reaches one at the zone's edge: a randomly drawn bias locates $2\bar b$, but only the worst case is flat, because the dead zone is an infimum, not a mean.
This is the dead zone of Figure (ref), now made sharp. The frontier $\pi_\alpha$ is flat at $\alpha$ throughout $|\xi|\le 2\bar b$, and its half-width is $2\bar b$ rather than $\bar b$ precisely because the least-favourable bias shifts the centre by $\bar b$ under the null and again under the alternative (Figure (ref)(b)).
Corollary (ref) establishes that the bias-aware interval is not needlessly wasteful. The rate at which power degrades is the fastest any honest procedure can achieve. The comparison is sharpest at the matched rate $\delta=\rho$, where Proposition (ref) gives the achievable frontier and Theorem (ref) the interval's actual power. By Proposition (ref) the gap between them is exactly the critical-value gap $z_{1-\alpha/2}-z_{1-\alpha}$, the slack a directional critical value recovers without violating coverage. This is the asymptotic counterpart of the finite-sample finding of armstrong2022robust that the bias-aware critical value cannot be reduced by more than roughly a factor of two without sacrificing coverage.
In the weak-factor setting of Example (ref), Assumption (ref) is the formal content of weak-factor non-detectability. A factor at the boundary of detection is statistically invisible at the available sample size onatski2012, so a coefficient shift aligned with its loading is absorbed into the nuisance, realising a difference in $\beta$ up to $\bar B$ with vanishing detectability and $\tau$ given by the Gaussian divergence. Therefore the lower bound applies straightforwardly to Example (ref).
Proposition (ref) is the local-power counterpart of the length non-adaptivity of low1997nonparametric and cailow2004adaptation. They show that an honest interval cannot be shorter than the worst-case modulus of continuity, whereas here it is shown that its dual test cannot have local power inside the resulting modulus-width window. The two-point modulus construction of donoho1994statistical is common to both, and the constant that lower-bounds their length is the $\bar b$ that bounds our power. By the length-power duality of pratt1961length, an interval that cannot beat the worst-case modulus induces a test that cannot detect within it.
The mechanism is in fact more general than honest inference, and is most transparent through partial identification. Say $\beta$ is partially identified with identified set $\mathrm{ID}(P)$, the values of $\beta$ consistent with the data law $P$ over the maintained class, and recall the dual test excludes a value when that value leaves the interval.
Proposition (ref) applies the paper's regimes under one principle. Local power degenerates whenever a procedure's interval width fails to shrink fast enough relative to the alternatives. Partial identification is the extreme since the width converges to a set that does not shrink at all, and honest inference under a fixed bound ($\delta=0$) is exactly this case. The intermediate regimes $0<\delta\le\rho$ are “partial identification at the local scale”, since after rescaling by $n^{\rho}$ the honest half-width $n^{\rho}\bar B=\Theta(n^{\rho-\delta})$ does not vanish, so the rescaled identified set has positive volume and the same degeneracy applies, by Proposition (ref) with $\beta_0$ interior (zero power) or at the boundary (one-sided power, recovering Theorem (ref)). The power loss is thus a feature of any procedure that endorses confidence intervals with slowly converging width, of which conservative honest inference is one instance imbens2004confidence,stoye2009more. The bound in part (i) is $\alpha$ rather than zero, and this is sharp. Honesty caps the exclusion power in excess of size, not the exclusion probability itself: a dual test that exhausts its size at $A_n$ still excludes $A_n$, under the truth $\beta_0$, with probability $\alpha+o(1)$. A particular honest interval may exclude perturbations with vanishing probability, but none can guarantee excluding one with more than $\alpha$, which is exactly the local power that degenerates.
Taken together, Sections (ref)-(ref) show that the bounds on honest local power coincide in rate and transition at $\delta=\rho$. The power loss of conservative confidence intervals is not a deficiency of particular constructions, but intrinsic to the uniform coverage requirement over a class that the data cannot discipline. Against a class wide enough to contain the bias, no honest procedure recovers the local power that conservatism forgoes. The only way to regain power is to relax the strict honesty requirement, for instance by accepting an undersized test that retains power low1997nonparametric,rambachan2023more.
The Gaussian limits in Section (ref) rest on standard primitives. The matched-rate normal approximation of Theorem (ref)(ii) and Proposition (ref) holds whenever $\hat\beta$ is asymptotically linear: $\hat\beta-\beta-\hat B=\sum_i w_{ni}u_i+o_p(n^{-\epsilon})$; and the array $\{w_{ni}u_i\}$ satisfies the Lindeberg condition with $n^{2\epsilon}\sum_i w_{ni}^2\,\mathrm{E}u_i^2\to\sigma^2$; the variance estimator is consistent, $n^{\epsilon}se(\hat\beta)\xrightarrow{p}\sigma$; and the bias is stable, $n^{\delta}\hat B\xrightarrow{p}c\bar b$, $n^{\delta}\bar B\to\bar b$. For $\epsilon=1/2$ this is the usual Lyapunov central limit theorem. For $\epsilon<1/2$ it is the standard Lindeberg condition for kernel estimators with effective sample size $n^{2\epsilon}$, and holds for the estimators of the designs in the simulations of Section (ref).
The function $\tau$ of Assumption (ref) is not a free primitive but a single scalar implied by the model. It shows how statistically different two laws within the same class of DGPs must be under a shift in the parameter of interest.
The four settings are instances, distinguished by the leading behaviour of $s$ at the origin. Where $s$ has a linear term, its coefficient $s_\star$ in $s(u)\approx s_\star u$ is the honest analogue of the classical Pitman slope that summarises local power vandervaart1998asymptotic; in general the leading exponent matters as much as the coefficient:
The endpoints are the theory's two extremes. Plausibly-exogenous IV and honest DiD have $s\equiv0$ on $[0,1)$: by Lemma (ref) a target shift of up to $2\bar B$ is matched exactly within the class, which is the perfect-confounding form of Assumption (ref) and the flat-dead-zone, pure zero-power case. Regression discontinuity has $s$ strictly increasing: the confounding is approximate, its cost from the H\"older modulus, so the frontier $\Phi(s(\xi)-z_{1-\alpha})$ rises smoothly rather than staying flat. There is no exactly flat dead zone over the full class, only the (bounded) resolution limit of the bias-aware procedures in use. The exponent is worth a remark: for $|f''|\le M$ the modulus is $\omega(\varepsilon)\asymp\varepsilon^{4/5}$, and the calibration $\bar B\asymp\omega(n^{-1/2})$ gives $s(u)\asymp u^{5/4}$. Read literally as a derivative at the origin, then, $s'(0)=0$ in every row. A concave modulus $\omega\asymp\varepsilon^{r}$ with $r<1$ always yields the leading power $1/r>1$, so the rows are distinguished by leading behaviour, not by a common Taylor coefficient, and the RD frontier lifts off $\alpha$ only slowly near the origin, its first-order term missing.
Where the linearisation $s(u)\approx s_\star u$ is used, $s_\star$ is the slope at the operating scale, the modulus' slope at the noise level, a secant rather than the origin derivative. The factor model is the one non-convex instance, where the modulus machinery does not apply. A factor at the boundary of detection onatski2012 makes a loading-aligned shift undetectable, which gives the qualitative content $s(0)=0$ that Theorem (ref)(i) needs but not a closed-form leading term, which is not claimed here. So $\tau$, the dead zone, and the interval length are three readings of the one object $s$. In classical terms the rows deform the textbook power envelope vandervaart1998asymptotic in three distinct ways: exact confounding preserves its slope and translates it by $2\bar b$; regression discontinuity replaces the parametric slope by the modulus' slope at the noise level; and the factor row flattens slope and curvature alike at the origin.
This section illustrates the power loss of Theorems (ref)--(ref) across five distinct honest-inference problems. Each design pairs a conservative, honest procedure with a conventional competitor that ignores the worst-case bias; the difference-in-differences design adds a third, intermediate trend-adjusted procedure, described there. In every case the data are generated at a fixed true parameter value, a level-$\alpha$ confidence interval is constructed once per Monte Carlo replication, and the dual test is obtained by inversion, rejecting any value that falls outside the interval. With the data generated at the true value $\beta_0$, sweeping the local alternative $A_n=\beta_0+\xi\,n^{-\rho}$ and averaging the rejection indicator $\mathbbm 1\{A_n\notin\mathcal B_n\}$ over replications traces the local-power curve as a function of $\xi$. Throughout, $\alpha=0.05$ and $z_{1-\alpha/2}$ denotes the standard normal quantile.
The five designs span the regimes of Section (ref). The interactive fixed-effects, difference-in-differences, and plausibly-exogenous designs place the bound at a constant (or sample-size-inflated) order relative to the $n^{-1/2}$ sampling rate, illustrating the zero-power collapse of Theorem (ref). The regression-discontinuity design is genuinely nonparametric, with sampling error and worst-case bias of the same order, illustrating the non-degenerate power loss of Theorem (ref). The linear-functional design contrasts a parametric estimator with an honest one that converges at a slower nonparametric rate, so the honest interval's $n^{1/2}$-rescaled width diverges and its rejection probability is asymptotically flat in $\xi$: the dual test cannot discriminate among $n^{-1/2}$ alternatives (Lemma (ref)). In the three bound-dominated designs the honest procedure attains size arbitrarily close to zero; in the regression-discontinuity and linear-functional designs, where the worst-case bias and the standard error are of the same order, size is conservative but bounded away from zero. In every case the formal price is the documented loss of local power.\footnote{Replication code for all five designs is available from the author; the empirical applications of Section (ref) are reproduced by the script described there.}
The leading example is the panel regression with interactive fixed effects of armstrong2022robust. With $n=NT$, data are generated as
with $\lambda_i,f_t\sim\mathrm{i.i.d.}\,N(0,1)$, the product $\lambda_i f_t$ normalised to unit variance, $\varepsilon_{it},\eta_{it}\sim\mathrm{i.i.d.}\,N(0,1)$, $\beta=1$, and $T=\lfloor N^{3/5}\rfloor$. The scalar $c_n$ controls the strength of the factor that contaminates the regressor: $c_n=n^{-1/3}$ yields a weak factor (Figure (ref)) and $c_n=1$ a strong factor (Figure (ref)). Local alternatives are taken at the parametric rate, $A_n=\beta\pm\xi/\sqrt{n}$.
\paragraph{Estimators (Figures (ref)-(ref)).} The conventional factor estimator is the iterated least-squares estimator of Bai2009, with interval $\hat\beta\pm z_{1-\alpha/2}\,\mathrm{se}(\hat\beta)$. The honest bias-aware interval is that of armstrong2022robust, $\{\hat\beta_{ba}\pm[\bar B+z_{1-\alpha/2}\,\mathrm{se}(\hat\beta_{ba})]\}$, where the user supplies the number of factors treated as weak (here one) and $\bar B$ is the implied worst-case bias bound. The analyst cannot detect whether the factor is weak, and so must guard against it. Under the weak factor the factor estimator's interval undercovers and its dual test over-rejects the truth, while under the strong factor the bias-aware interval is conservative even though the true bias is negligible.
The second design is the honest difference-in-differences setting of rambachan2023more. An event study with three pre-treatment periods and one post-treatment period is summarised by its coefficient vector $\hat\beta=(\hat\beta_{\mathrm{pre}}^\top,\hat\beta_{\mathrm{post}})^\top\in\mathbb{R}^4$, generated directly as a Gaussian draw
where $\beta_{\mathrm{post}}=0$ is the true post-period effect and the entries of $\mu$ encode a linear differential trend $\delta_t=s\cdot t$ over event time $t\in\{-3,-2,-1,1\}$, plus a second-difference kink $w$ at $t=1$, the reference period $t=0$ being normalised out. The scalars $s\ge 0$ and $w\ge0$ index the strength of the linear pre-trend and of the deviation from linearity, and the conventional standard error is of order $N^{-1/2}$. The target is the post-period effect, $\theta=\ell^\top\hat\beta_{\mathrm{post}}$ with $\ell=1$, whose true value remains $\beta_{\mathrm{post}}$ for every $(s,w)$. Local alternatives are taken at $A_n=\beta_{\mathrm{post}}\pm\xi/\sqrt N$, and the sample sizes are $N\in\{100,500,2500\}$.
Setting $s=w=0$ recovers exact parallel trends, in which the honest interval is needlessly conservative. For $s>0$ with $w=0$ there is a linear pretrend, and the pre-treatment coefficients trend linearly and the post-period coefficient is contaminated by $\delta_1=s$. Because a linear trend has zero second differences, it lies in $\Delta^{SD}(M)$ for every $M\ge 0$ (here set to $M =0.2$). The honest interval extrapolates the estimated pre-trend and continues to cover the true effect $\beta_{\mathrm{post}}$, whereas the conventional event-study interval ignores the pre-periods and centres on the biased post coefficient $\beta_{\mathrm{post}}+s$, so its coverage of the truth collapses as $N$ grows. Note the honest interval is not the event-study interval widened about a common centre, as in the plausibly-exogenous design below: the extrapolation recentres it at the truth, so the two non-rejection windows separate, and with $s=M$ the event-study window collapses onto the upper edge of the honest window. This is the upper endpoint of the local identified set $[-M,M]$.
The design therefore includes a third, intermediate procedure that realises the recentring without the widening: the trend-adjusted interval, the fixed-length interval over $\Delta^{SD}(0)$, which extrapolates the estimated linear pre-trend with no allowance for curvature. This is the parametric linear adjustment common in applied work dobkin2018economic,bhuller2013broadband, corresponding to $M=0$ in the class of rambachan2023more. Under an exactly linear trend it is correctly centred and nominally sized and retains full local power. It claims the power the honest interval forgoes by asserting the untestable $w=0$, exactly as the event study asserts $\delta=0$, the same wager, one derivative up. The kinked configuration ($s=0$, $w=M$) prices that wager. It is the least-favourable path of Lemma (ref)(ii), flat pre-periods with $\delta_1=M$, on which the extrapolation of the flat pre-trend is zero, so the conventional and trend-adjusted intervals are both miscentred by $M$ and reject the truth with probability approaching one, while the honest interval, operating at its designed worst case, spends exactly its size $\alpha$ at the truth: its undetected window is the one-sided $[0,2M]$, the $c=1$ sliding window of Theorem (ref)(i). Figure (ref) reports the strong pre-trend ($s=0.2$, $w=0$), the kinked violation ($s=0$, $w=0.2$), and no pre-trend ($s=w=0$).
\paragraph{Estimators (Figure (ref)).} Terms $M$ and $\Delta^{SD}(M)$ in the following are defined in rambachan2023more. The conventional event-study interval is the standard confidence set that imposes parallel trends ($\delta=0$), $\hat\theta\pm z_{1-\alpha/2}\,\mathrm{se}(\hat\theta)$. The trend-adjusted interval is the fixed-length interval over $\Delta^{SD}(0)$, computed with the same machinery at $M=0$. The honest interval is the fixed-length confidence interval over the smoothness class $\Delta^{SD}(M)$, which bounds the second differences of the differential trend by a user-chosen $M$ (here $M=0.2$). The honest interval is centred at the extrapolation-corrected estimator and widened by the worst-case post-period bias that a violation within $\Delta^{SD}(M)$ can induce; as $N$ grows with $M$ fixed, this constant-order term dominates the shrinking standard error.
The third design is the honest sharp regression-discontinuity setting of armstrong2018optimal, armstrong2020simple, with the Monte Carlo design of their supplemental. Running variable $x_i\sim\mathrm{Uniform}[-1,1]$, $\varepsilon_i\sim\mathrm{i.i.d.}\,N(0,\sigma^2)$ with $\sigma^2=0.1295$, and
where $\tau=0$ is the true jump and $f(x)=C\,\operatorname{sign}(x)\,x^2$ ($C=1$) is the least-favourable odd quadratic, saturating $|f''|=2C$ everywhere, scaled by an amplitude $a\in[0,1]$. The sample sizes are $n\in\{500,2500,12500\}$ and, the estimator being nonparametric, local alternatives are taken at the rate $A_n=\tau\pm\xi\,n^{-2/5}$.
\paragraph{Estimators (Figure (ref)).} Both intervals use the triangular-kernel local-linear estimate $\hat\tau$ at a common bandwidth $h$ (H\"older class $\{|f''|\le M\}$, $M=2C$). The conventional interval $\hat\tau\pm z_{1-\alpha/2}\,\mathrm{se}(\hat\tau)$ drops the smoothing bias, whereas the honest interval replaces $z_{1-\alpha/2}$ by the bias-aware value $\mathrm{cv}_\alpha(\overline{\mathrm{bias}}/\mathrm{se})$, with $\overline{\mathrm{bias}}$ the worst-case bias over the class. At the optimal bandwidth the bias-to-standard-error ratio $\overline{\mathrm{bias}}/\mathrm{se}$ is a bounded constant ($\approx 0.5$), so the bias-aware critical value is only modestly above $z_{1-\alpha/2}$ armstrong2020simple. RD is therefore the matched-rate case of Theorem (ref) ($\delta=\rho=\epsilon=2/5$), in which conservatism costs only a bounded amount. When $a=1$ the curvature saturates the bound, so the conventional interval under-covers (coverage $\approx 0.92$; honest needed), while when $a=0$ the design is linear, the conventional interval is correct, and the honest interval is only $\approx 11\%$ wider. The much larger under-coverage documented by armstrong2018optimal arises with data-driven bandwidths that mis-estimate curvature; here only the honest/conventional margin at a common bandwidth is explored, for clarity.
The fourth design is the stylised problem of donoho1994statistical, and unlike the others the honest procedure here converges at a slower rate than the conventional one. Design points $x_i\sim\mathrm{Uniform}[0,1]$, $i=1,\dots,n$, generate
with $\sigma=1$ and target the boundary value $f(0)=\theta=0$. The term $a\,C\min(x,h)$ is a Lipschitz wedge of amplitude $a\in[0,1]$ confined to the window $[0,h]$, of width equal to the honest bandwidth $h\sim n^{-1/3}$ defined below. Sample sizes are $n\in\{500,2500,12500\}$ and local alternatives are taken at the parametric rate $A_n=\theta\pm\xi/\sqrt n$.
\paragraph{Estimators (Figure (ref)).} The conventional parametric estimator assumes the approximately linear model of donoho1994statistical holds exactly ($f$ affine). It is the OLS intercept $\hat\theta_{ols}$ with interval $\hat\theta_{ols}\pm z_{1-\alpha/2}\,\mathrm{se}(\hat\theta_{ols})$ and $\mathrm{se}$ of order $(n^{-1/2})$. The honest estimator only assumes $f$ lies in the Lipschitz class $\{|f(x)-f(x')|\le C|x-x'|\}$ ($C=2$). It is the one-sided local average over $[0,h]$ at the MSE-optimal bandwidth $h=(\sigma^2/(C^2 n))^{1/3}$, widened by the modulus-of-continuity bias $C\,h$, $\hat\theta_{loc}\pm(C\,h+z_{1-\alpha/2}\,\mathrm{se}(\hat\theta_{loc}))$, with the slower nonparametric rate $\mathrm{se}$ of order $(n^{-1/3})$. When $a=0$ the function is flat and honesty is not needed, and the parametric interval is valid and sharper, the honest interval needlessly wide. When $a>0$ the OLS line extrapolates the bulk level $a\,C\,h$ and undercovers the boundary value $f(0)$, while the honest interval stays valid. Crucially the wedge is untestable: confined to $[0,h]$ it perturbs only $\sim nh=n^{2/3}$ observations at magnitude $\sim n^{-1/3}$, so no specification test detects it. Indeed a $y\sim x+x^2$ test keeps size $0.05$ for every $n$, and the conservatism cannot be sidestepped by first testing for non-linearity, unlike a global quadratic whose curvature is $\sqrt n$-estimable. In both cases the honest interval shrinks only at $n^{-1/3}$, so against $n^{-1/2}$ alternatives its rejection probability is asymptotically flat in $\xi$ (Lemma (ref)): no local discrimination, precisely the price of robustness.
The final design is the plausibly-exogenous instrumental-variables setting of conley2012plausibly. With a single instrument $z_i$, endogenous regressor $x_i$, and structural outcome $y_i$,
where $v_i,\nu_i\sim\mathrm{i.i.d.}\,N(0,1)$, first-stage strength $\pi=1$, endogeneity $\rho=0.6$, and structural coefficient $\beta=1$. The analyst relaxes the exclusion restriction to $|\gamma|\le\bar\gamma$ with $\bar\gamma=0.30$. The knob is the true exclusion violation, $\gamma\in\{0.3,\ 0.05,\ 0\}$: a strongly endogenous, weakly endogenous, and exogenous instrument, respectively (first-stage strength is held at $\pi=1$ throughout). Under the strongly endogenous instrument ($\gamma=0.3$, the violation at the bound) the conventional IV estimator is miscentred by $\gamma/\pi=0.3$ and honesty is needed; under the exogenous instrument ($\gamma=0$) the exclusion restriction in fact holds, the instrumental-variables analogue of the strong-factor design, and the honest interval is conservative despite negligible true bias. Sample sizes are $n\in\{100,500,2500\}$ and local alternatives are at the parametric rate $A_n=\beta\pm\xi/\sqrt n$.
\paragraph{Estimators (Figure (ref)).} Conventional standard IV imposes $\gamma=0$, $\hat\beta_{IV}=(z^\top y)/(z^\top x)$, with interval $\hat\beta_{IV}\pm z_{1-\alpha/2}\,\mathrm{se}(\hat\beta_{IV})$. The honest interval is the union of confidence intervals of conley2012plausibly: for each admissible $g\in[-\bar\gamma,\bar\gamma]$ one forms the IV interval for $\beta$ from the netted outcome $y-g z$, $\hat\beta(g)=z^\top(y-g z)/(z^\top x)$, and takes the union $\big[\min_{|g|\le\bar\gamma}\mathrm{lb}(g),\ \max_{|g|\le\bar\gamma}\mathrm{ub}(g)\big]$. The union widens the interval by an amount of order $\bar\gamma\,(z^\top z)/(z^\top x)\approx\bar\gamma/\pi$, a constant that dominates the $O(n^{-1/2})$ sampling error as $n$ grows.
The paper's practical recommendation is to report the power “dead zone” alongside bias-aware intervals. This is the set of effect sizes the chosen bound forgoes the ability to detect. This section makes the recommendation concrete on two published studies, chosen so that the exercise is easily reproducible.\footnote{The script Applications.R reproduces both numbers and figures from the replication data bundled in the RDHonest and HonestDiD packages; no data need be downloaded.} The two sit in the paper's two non-degenerate regimes: bounded and cheap at the matched rate, decisive and non-vanishing when the bound dominates.
In both applications the relevant null is no effect, $\beta_0=0$, and the dead zone is the set of true effects whose detection cannot be guaranteed. That is, the zone symmetrically about $\beta_0=0$ that an honest test rejects no more often than its size. The dead-zone half-width $2\bar B$ is a property of the bound, not of the null: it is the same window around any hypothesised value, so it is drawn at the economically natural $\beta_0=0$. At the matched rate (as in RD) it is a fixed band of half-width $2\bar B$. When the bound dominates (as in DiD) the band grows with the bound.
Consider the lee2008randomized regression-discontinuity estimate of the U.S. House incumbency advantage, using the bias-aware procedure of armstrong2018optimal, armstrong2020simple with the data-driven smoothness constant of its default rule of thumb. The point estimate is a $5.85$ percentage-point incumbency advantage. The conventional interval, which drops the smoothing bias, is $(3.17,8.53)$; the honest interval, which inflates the critical value from $z_{1-\alpha/2}=1.96$ to the bias-aware $\mathrm{cv}_\alpha=2.31$, is $(2.69,9.01)$, only $18\%$ wider. This is the matched-rate case $\delta=\rho=\epsilon=2/5$ of Theorem (ref). The worst-case bias and the standard error are of the same order ($\bar B/se\approx0.65$), so conservatism costs only a bounded amount. The worst-case bias is $\bar B=0.89$, so the dead zone of the reported procedure has half-width $2\bar B=1.78$ percentage points. Here the estimate exceeds the dead-zone half-width by a factor of more than three (Figure (ref)), so the finding is robust to the conservatism, and disclosing the dead zone confirms rather than overturns it.
Consider the benzarti2019who event study of a large French VAT cut, asking whether the reduction was captured as firm profits, with the honest difference-in-differences procedure of rambachan2023more. Imposing exact parallel trends, the first post-reform year ($2009$) shows a sharply significant profit effect of $0.196$ with conventional interval $(0.159,0.233)$ ($t\approx10$). This is the bound-dominated regime: under the relative-magnitudes restriction $\Delta^{RM}(\bar M)$, which allows the post-period violation of parallel trends to be at most a multiple $\bar M$ of the largest pre-period violation, the honest interval widens with $\bar M$ and does not shrink with the sample. The lower limit reaches zero at a breakdown of $\bar M\approx1.75$ (Figure (ref)). rambachan2023more analyse this application with the same machinery and report a closely related sensitivity analysis of the profit effect; the incremental content here is the dead-zone reading of that exercise: the breakdown $\bar M$ is the point at which the estimate enters the dead zone of the no-effect test. The conventionally significant effect becomes undetectable once one concedes a post-period trend deviation roughly twice the largest one seen before treatment. Disclosure of the dead zone in this application makes explicit that the ability to reject a zero effect rests on the bound the analyst is prepared to defend, i.e. on the prescribed tolerance for pre-trend violations.
The two applications bracket the message. In regression discontinuity the dead zone is a bounded, modest object and the headline survives it; in difference-in-differences the dead zone is governed entirely by an untestable bound, and can be wide enough that even a very large $t$-statistic is not, by itself, decisive. In both cases the dead-zone half-width is a single number, in the units of the estimand, that the interval's endpoints alone do not convey.
This paper establishes a local-alternative power trade-off intrinsic to conservative inference. At the usual parametric rate, an honest interval may have no local power, and the loss is shared by every honest procedure, not just bias-aware constructions. The paper is closed with some practical discussion.
\paragraph{Disclosure of dead zone.} Consider comparing the bias bound $\bar B$ with the standard error. $\bar B=\sup_\theta|\hat B|$ is chosen by the analyst, so the comparison merely re-expresses how conservative a bound is. The power cost is not a hidden risk to be detected but a deterministic consequence of a stated belief about unobservable biases, i.e. a monotone function of that preference. What this analysis recommends is disclosure: report the dead-zone half-width $2\bar b$, in the units of the estimand, alongside any bias-aware interval, since it states exactly which alternatives the chosen bound forgoes the ability to detect. This is information the interval's endpoints alone do not convey.
\paragraph{Honesty versus power.} When the bound dominates standard errors, which is the typical parametric-rate case with zero honest local power, and the analyst nonetheless needs power, the only escape is to relax strict honesty. For instance, to defend a tighter bound on substantive grounds, or to report an undersized test that retains local power low1997nonparametric,rambachan2023more. Which is preferable is a genuine modelling choice that the size--power frontier here makes explicit rather than resolves. Conservatism buys uniform coverage at a power cost that is, at the parametric rate, total. Quantifying when the bound can be credibly tightened, and characterising the optimal undersized test are natural next steps for research.
{2pt}