EconBase
← Back to paper

Bias-Aware Inference in Fuzzy Regression Discontinuity Designs

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

153,500 characters · 20 sections · 43 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Bias-Aware Inference in Fuzzy Regression Discontinuity DesignsFirst Version: June 11, 2019. This Version: . We would like to thank, Tim Armstrong, Marinho Bertanha, Yingying Dong, Keisuke Hirano, Michal Koles\'ar, the anonymous referees and numerous seminar participants for their helpful comments and suggestions. The authors gratefully acknowledge financial support by the European Research Council (ERC) through grant SH1-77202. Contact information: Claudia Noack, Nuffield College and Department of Economics, University of Oxford, email: [email removed], website: http://claudianoack.github.io. Christoph Rothe, Department of Economics, University of Mannheim, 68131 Mannheim, Germany, email: [email removed], website: http://www.christophrothe.net.

\pagestyle{plain}

\newtheorem{theorem}{Theorem} \newtheorem{definition}{Definition} \newtheorem{lemma}{Lemma} \newtheorem{assumption}{Assumption} \newtheorem{assumptionLL}{Assumption}

\theoremstyle{definition} \newtheorem{example}{Example} \newtheorem{remark}{Remark}

{1pt}

abstractWe propose new confidence sets (CSs) for the regression discontinuity parameter in fuzzy designs. Our CSs are based on local linear regression, and are bias-aware, in the sense that they take possible bias explicitly into account. Their construction shares similarities with that of Anderson-Rubin CSs in exactly identified instrumental variable models, and thereby avoids issues with “delta method” approximations that underlie most commonly used existing inference methods for fuzzy regression discontinuity analysis. Our CSs are asymptotically equivalent to existing procedures in canonical settings with strong identification and a continuous running variable. However, due to their particular construction they are also valid under a wide range of empirically relevant conditions in which existing methods can fail, such as setups with discrete running variables, donut designs, and weak identification.

\onehalfspacing

Introduction

Regression discontinuity designs can deliver credible identification of treatment effects from observational data in settings where the probability of receiving the treatment changes discontinuously with a running variable at some known threshold value. Such designs are called sharp (SRD) if the probability changes from zero to one, and fuzzy (FRD) otherwise. With both types of designs, methods based on local linear regression are widely used in empirical practice for estimation and inference hahn2001identification.

The confidence intervals (CIs) typically reported in empirical FRD studies are obtained by applying techniques for handling smoothing bias, such as robust bias correction calonico2014robust or bias-aware critical values armstrong2018optimal, to a delta method (DM) approximation of the FRD estimator. Such DM CIs can be unreliable in practice, however, if the running variable is not continuously distributed with full support around the cutoff, or the jump in treatment probabilities is “small”. These limitations are important because empirical researchers often face running variables like test scores or class sizes that take only a limited number of distinct values, “donut designs” that exclude units close to the cutoff to increase the credibility of causal estimates, or weakly identified setups where treatment assignment only has a moderate impact on treatment probabilities.\footnote{To illustrate the scope of the issue, we surveyed the articles published between 2015 and 2021 in the “Top 5” economics journals. We found 20 papers that used a fuzzy regression discontinuity design as one of their main empirical specification; 9 of which had an “irregular support” (in the sense of having less than 100 support points within the bandwidth window on each side of the cutoff), and one was potentially affected by weak identification (in the sense that the reported first stage estimate differed by less than three standard errors from zero).}

In this paper, we propose a new class of FRD confidence sets (CSs) that are not subject to these issues. The idea is to apply a bias-aware approach, which takes possible finite sample biases into account, to a particular local linear SRD estimator that is conceptually similar to an Anderson-Rubin statistic in a linear IV model staiger1997instrumental,feir2016weak. We show that the resulting CSs are “honest” in the sense of li1989honest, meaning that they have correct asymptotic coverage uniformly over a class of conditional expectation functions of outcomes and treatments with bounded second derivatives, irrespective of the distribution of the running variable or the strength of identification. We also show that our CSs are asymptotically equivalent to bias-aware DM CIs in settings with a continuous running variable and strong identification.

Regression discontinuity methods that explicitly take possible bias into account have been shown to have favorable theoretical and practical properties, for instance, by armstrong2018optimal, armstrong2018simple, kolesar2018discrete and imbens2019optimized. The CSs in this paper complement these methods by providing reliable FRD inference under potentially challenging circumstances without sacrificing efficiency in the canonical setup. Our approach is related to that of feir2016weak, who also consider Anderson-Rubin-type statistics in FRD designs with potentially small jumps in treatment probabilities, but differs in that it allows for discrete (or otherwise irregularly supported) running variables, takes potential bias explicitly into account, and includes a method for choosing bandwidths in practice.

Setup and Preliminaries

Let $Y_i\in\mathbb{R}$ be the outcome, $T_i\in\{0,1\}$ the actual treatment status, $Z_i\in\{0,1\}$ the assigned treatment, and $X_i\in\mathbb{R}$ the running variable of the $i$th unit in a random sample of size $n$ from a large population. Treatment is assigned if the running variable falls above a known cutoff that we normalize to zero, so that $Z_i = \mathbf{1}\{X_i \geq 0\}$. The parameter of interest is $\theta = \tau_Y/\tau_T,$ where for a generic random variable $W_i$ we write $\mu_W(x)=\mathbb{E}(W_i|X_i=x)$ for is its conditional expectation function given the running variable; $\mu_{W+} =\lim_{x\downarrow 0}\mu_W(x) $ and $\mu_{W-} =\lim_{x\uparrow 0}\mu_W(x)$ for the right and left limit at the cutoff; and $\tau_W =\mu_{W+}-\mu_{W-}$ for the corresponding jump.\footnote{We write $f_{+} =\lim_{x\downarrow 0}f(x) $ and $f_{-} =\lim_{x\uparrow 0}f(x)$ for generic functions $f$ throughout the paper.} In a potential outcomes framework with certain continuity and monotonicity conditions hahn2001identification, the parameter $\theta$ has a causal interpretation as the local average treatment effect among units at the cutoff whose treatment decision is affected by the assignment rule.

Our goal is to construct powerful confidence sets $\mathcal{C}^\alpha$ that cover the parameter $\theta$ in large samples with at least probability $1-\alpha$, uniformly over $(\mu_Y,\mu_T)$ in some function class $\mathcal{F}$ that embodies shape restrictions imposed by the analyst:\footnote{Note that we leave the dependence of the probability measure $\mathbb{P}$ and the parameter $\theta$ on $\mu_Y$ and $\mu_T$ implicit in our notation. Each function pair $(\mu_Y,\mu_T)$ corresponds to a single distribution of $(Y,T,X,Z)=(\mu_Y(X)+\epsilon_M,\mathbf{1}\{\mu_T(X)\geq\epsilon_T\},X,Z)$, where $(\epsilon_M,\epsilon_T)$ is some fixed random vector. }

align[align omitted — 140 chars of source]

Following li1989honest, we refer to such CSs as honest with respect to $\mathcal{F}$. This is a stronger requirement than correct pointwise asymptotic coverage:

align[align omitted — 160 chars of source]

In particular, under (ref) we can always find a sample size $n$ such that the coverage probability of $\mathcal{C}^\alpha$ is not below $1-\alpha$ by more than an arbitrarily small amount for every $(\mu_Y, \mu_T)\in\mathcal{F}$. Under (ref) there is no such guarantee, and even in very large samples the coverage probability of $\mathcal{C}^\alpha$ could be poor for some $(\mu_Y, \mu_T)\in\mathcal{F}$. Since we do not know in advance which function pair is the correct one, honesty as in (ref) is necessary for good finite sample coverage of $\mathcal{C}^\alpha$ across data generating processes.

As in imbens2019optimized or armstrong2018simple, we specify $\mathcal{F}$ as a smoothness class. Specifically, let $\mathcal{F}_H(B)=\{f_1(x)\mathbf{1}\{x\geq 0\}- f_0(x)\mathbf{1}\{x< 0\}: \|f_w''\|_{\infty} \leq B, w=0,1\}$ be the Hölder-type class of real functions that are potentially discontinuous at zero, are twice differentiable on either side of the threshold, and whose second derivatives are uniformly bounded by some constant $B>0$; and let $\mathcal{F}_H^\delta(B)=\{f\in\mathcal{F}_H(B): |f_+ - f_-|>\delta\}$ be a similar class of functions whose jump at zero is larger than some $\delta\geq 0$. We then assume that there are constants $B_Y$ and $B_T$, whose choice we discuss in Section (ref), such that

align[align omitted — 115 chars of source]

As $\mathcal{F}$ is a Cartesian product, this rules out cross-restrictions between $\mu_Y$ and $\mu_T$. Note that we impose $\mu_T\in \mathcal{F}_H^0(B_T)$, and thus $\tau_T\neq 0$, only to ensure that $\theta=\tau_Y/\tau_T$ is well-defined. We explicitly allow $\tau_T$ to be arbitrarily close to zero.

If the running variable is discrete, or more generally such that there are gaps in its support, condition (ref) is understood to mean that there exists a function pair $(\mu_Y,\mu_T) \in \mathcal{F}$ such that $(\mu_Y(X_i),\mu_T(X_i))=(\mathbb{E}(Y_i|X_i),\mathbb{E}(T_i|X_i))$ with probability one kolesar2018discrete. With this interpretation, the parameter $\theta$ is generally partially identified:

multline*[multline* omitted — 200 chars of source]

where the identified set $\Theta_I$ is either (i) a closed interval $[a_1,a_2]$ with $a_1\leq a_2$; (ii) the union of two disjoint half-lines, $(-\infty,a_1]\cup[a_2,\infty)$ with $a_1<0<a_2$; (iii) the entire real line; or, as a knife-edge case (iv) a half-line $[a_1,\infty)$ or $(-\infty,-a_1]$, with $a_1>0$.\footnote{This holds because the range of $(m_{Y+}-m_{Y-},m_{T+}-m_{T-})$ over $(m_Y,m_T)\in \mathcal{F}$ is a Cartesian product of two intervals $I_Y \times I_T$. The four cases then obtain depending on which of these two intervals contain zero, possibly as a boundary value. } The classical point identification result when the support of $X_i$ contains an open neighborhood around the cutoff is then simply a special case of (i) with $\Theta_I$ a singleton. Our goal is to construct CSs for $\theta$ that have correct uniform asymptotic coverage under both point and partially identification imbens2004confidence, without applied researchers having to decide which of the two notions of identification more accurately applies to their specific setting.

Bias-Aware Anderson-Rubin-Type Confidence Sets

We argue in Section (ref) that conventional CIs, based on local linear regressions fan1996local and “delta method” (DM) arguments, can potentially break down in a number of practically relevant setups, including ones with discrete running variables or weak identification. Our proposed approach is still based on local linear regression, but avoids these issues through a construction similar to that of anderson1949estimation for inference in exactly identified linear IV models. It also takes possible bias from local linear smoothing explicitly into account. We hence refer to our CSs as bias-aware AR CSs.

To describe the approach, we write $\widehat\tau_W(h)$ for the local linear estimator of the jump $\tau_W$ in the conditional expectation of some generic random variable $W_i$ given the running variable $X_i$ at the cutoff:

align[align omitted — 204 chars of source]

Here $K(\cdot)$ is a kernel function, $h>0$ is a bandwidth, $e_1 = (1,0,0,0)^\top$ is the first unit vector, and the $w_i(h)$ are weights, given explicitly in Appendix (ref), that only depend on the data through the realizations $\mathcal{X}_n = (X_1,\ldots,X_n)^\top$ of the running variable. We refer to estimators of the form in (ref) as SRD-type estimators of $\tau_W$ in the following.

The natural point estimator of $\theta$ is $\widehat\theta(h) = \widehat\tau_Y(h)/\widehat\tau_T(h)$, but we will base inference on different statistics. Define the auxiliary parameters $\tau_{M}(c)=\tau_{Y}- c\tau_{T}$ for $c\in\mathbb{R}$, and note that $\tau_{M}(c) = \mu_{M+}(c)-\mu_{M-}(c)$, with $\mu_{M}(x,c) =\mathbb{E}(M_i(c)|X_i=x)$ and $M_i(c) = Y_i - cT_i$. We then consider the SRD-type estimator $\widehat\tau_{M}(h,c)=\sum_{i=1}^n w_{i}(h) M_i(c)$, and exploit the properties of the regression weights $w_i(h)$ to write its conditional bias $b_{M}(h,c)=\mathbb{E}(\widehat\tau_{M}(h,c)|\mathcal{X}_n)-\tau_{M}(c)$ and conditional variance $s^2_{M}(h,c) =\mathbb{V}(\widehat\tau_{M}(h,c)|\mathcal{X}_n)$ given $\mathcal{X}_n$ as

align*[align* omitted — 151 chars of source]

respectively, with $\sigma_{M,i}^2(c) = \mathbb{V}(M_i(c)|X_i)$ the conditional variance of $M_i(c)$ given $X_i$.\footnote{To keep the notation simple, the estimator $\widehat\tau_{M}(h,c)=\widehat\tau_{Y}(h)- c\widehat\tau_{T}(h)$ uses the same bandwidth on each side of the cutoff, and also the same bandwidth for estimating $\tau_Y$ and $\tau_T$. It is straightforward to accommodate more general bandwidth choices; see Online Appendix (ref) for details.}

The bias depends on $(\mu_Y,\mu_T)$ through the transformation $\mu_M = \mu_Y - c \mu_T$ only, and we have that $\mu_M \in \mathcal{F}_H(B_Y+|c| B_T)$ by (ref). As in armstrong2018simple, for any value of the bandwidth $h$ we can thus explicitly bound $b_{M}(h,c)$ in absolute value over $\mathcal{F}$:

align*[align* omitted — 173 chars of source]

The supremum is achieved by the “worst case” pair of piecewise quadratic conditional expectation function with second derivatives equal to $(B_Y \textnormal{sign}(x), - B_T \textnormal{sign}(x))$ over $x\in[-h,h]$.\footnote{ Note that this bound may not be sharp if no such pair of piecewise quadratic functions is a feasible candidate for $(\mu_Y,\mu_T)$. For example, there is no function $\mu_T$ with $\mu_T''(x)= B_T\textnormal{sign}(x)$ and $\mu_T(x)\in [0,1]$ for all $x\in[-h,h]$ if $h> (2 / B_T)^{1/2}$. Still, the bias bound is valid in such cases.}

For every $c\in\mathbb{R}$ we can then construct an infeasible bias-aware CI for $\tau_{M}(c)$ as $$C_M^{\alpha}(h,c) =\left[\widehat\tau_{M}(h,c) \pm\textnormal{cv}_{1-\alpha}( r_{M}(h,c)) s_{M}(h,c)\right],$$ where $r_{M}(h,c) = \overline{b}_{M}(h,c)/s_{M}(h,c)$ is the “worst case” bias to standard deviation ratio, and $\textnormal{cv}_{1-\alpha}(r)$ is the $(1-\alpha)$-quantile of “folded” normal distribution $|N(r,1)|$. armstrong2018optimal,armstrong2018simple show that such CIs are honest with respect to $\mathcal{F}_H(B_W)$ irrespective of the distribution of the running variable, have correct asymptotic coverage $1-\alpha$ at the “worst case” conditional expectations, are valid for wide ranges of bandwidths, and are highly efficient for SRD inference if the running variable is continuous.

The bandwidth that minimizes this CI's asymptotic length is $$h_M(c) = \operatorname*{argmin}_{h} \textnormal{cv}_{1-\alpha}( r_{M}(h,c)) s_{M}(h,c).$$ We assume that this optimal bandwidth is unique; and the minimization in its definition is understood to be carried out over the set of bandwidths for which the involved objects are well-defined. An efficient but infeasible bias-aware AR CS for $\theta$ is then given by the set of all $c\in\mathbb{R}$ for which the auxiliary CI $C_M^{\alpha}(h_M(c),c)$ contains the value zero:

align[align omitted — 176 chars of source]

Our proposed bias-aware AR CSs are feasible versions of (ref) based on a standard error $\widehat s_{M}(h,c)= \sum_{i=1}^nw_i(h)^2\widehat\sigma_{M,i}^2(c)$ and some estimate $\widehat{h}_M(c)$ of the optimal bandwidth:

align[align omitted — 232 chars of source]

with $\widehat r_{M}(h,c) = \overline{b}_{M}(h,c)/\widehat s_{M}(h,c)$. Both the standard error and bandwidth estimator can be implemented in different ways, and our theoretical analysis below therefore only imposes some weak “high level” conditions. We propose a specific standard error based on nearest-neighbor linear regression estimates $\widehat\sigma_{M,i}^2(c)$ of $\sigma_{M,i}^2(c)$ in Section (ref); and a feasible bandwidth that combines a plug-in construction with a safeguard against certain small sample distortions in Section (ref).

Theoretical Properties

Coverage

Our main theoretical result is that $ \mathcal{C}_{\textnormal{ar}}^{\alpha}$ is an honest CS for $\theta$ with respect to $\mathcal{F}$, in the sense of (ref), under the following rather weak conditions.

assumption(i) The data $\{(Y_i,T_i,X_i),i=1,\ldots,n\}$ are an i.i.d. sample; (ii) $\mathbb{E}( (Y_i - \mathbb{E}(Y_i|X_i))^q| X_i=x)$ exists and is bounded uniformly over $x\in\textrm{supp}(X_i)$ and $(\mu_Y,\mu_T)\in\mathcal{F}$ for some $q>2$; (iii) $\mathbb{V} (Y_i|X_i=x)$ is bounded away from zero uniformly over $x\in\textrm{supp}(X_i)$ and $(\mu_Y,\mu_T)\in\mathcal{F}$; and $\textrm{Cov}(Y_i, T_i|X_i=x)^2/(\mathbb{V} (Y_i|X_i=x) \mathbb{V} (T_i|X_i=x))$ is bounded away from one uniformly over $x\in\textrm{supp}(X_i)\cup\{x: \mathbb{V} (T_i|X_i=x)>0 \}$ and $(\mu_Y,\mu_T)\in\mathcal{F}$; (iv) the kernel function $K$ is a continuous, unimodal, symmetric density function that is equal to zero outside some compact set, say $[-1,1]$.

Assumption (ref) collects mostly standard conditions from the literature on local linear regression. Part (i) could be weakened to allow for certain forms of dependent sampling, such as cluster sampling. Parts (ii) and (iii) ensure that $\mathbb{V}(M_i(c)|X_i=x)$ is bounded away from zero for all $c\in\mathbb{R}$ and allow for the special case of a SRD design. Part (iv) is satisfied by most kernel functions commonly used in applied RD analysis, such as the triangular or the Epanechnikov kernels.

assumptionThe following holds uniformly over $(\mu_Y,\mu_T)\in\mathcal{F}$: (i) $\widehat{h}_M(c)=h_M(c)(1 +o_P(1))$; and (ii) $\widehat{s}_{M}( \widehat{h}_M(c),c) = s_{M}( h_{M}(c),c)(1+o_P(1))$.

Part (i) of Assumption (ref) states that the empirical bandwidth is consistent for the infeasible optimal one, and part (ii) states that the empirical standard error is consistent for the true standard deviation at the infeasible optimal bandwidth. We discuss specific implementations in Sections (ref) and (ref).

assumptionLLThe support of the running variable $X_i$ is finite and symmetric, in the sense that it is of the form $\{\pm x_1, \ldots, \pm x_k\}$, for positive constants $(x_1,\ldots,x_k$) over some open neighborhood of the cutoff.
assumptionLLThe running variable $X_i$ is continuously distributed with continuous density $f_X$ that is bounded and bounded away from zero over an open neighborhood of the cutoff.

Assumptions (ref)--(ref) describe RD setups with discrete and continuously distributed running variables, respectively. These settings are meant to be exemplary and are considered because they allow explicit characterization of the bandwidth $h_M(c)$. Note that the symmetry of the support in Assumption (ref) is for notational convenience only. Discrete running variables with asymmetric support can easily be accommodated by using a different bandwidth on each side of the cutoff, as described in Online Appendix (ref). In Appendix (ref) we also consider an alternative asymptotic framework for the discrete case.

In Lemma (ref) in the Appendix, we show that our assumptions have two main implications: (i) using an estimate of the optimal bandwidth instead of its population version has a small impact, in some appropriate sense, on the quantities involved in the construction of our CS; (ii) the magnitude of each weight $w_i(h_{M}(c))$ is small relative to the others' in large samples, in the sense that $w_\textnormal{ratio}(h) \equiv \max_{j=1,\ldots,n} w_j(h)^2/\sum_{i=1}^n w_i(h)^2=o_P(1),$ so that a CLT applies to an appropriately standardized version of the estimator of $\tau_{M}(c)$. This yields the following formal result.

theoremSuppose that Assumptions (ref)--(ref) and either (ref) or (ref) hold. Then $\mathcal{C}_{\textnormal{ar}}^{\alpha}$ is honest with respect to $\mathcal{F}$ in the sense of (ref).

Shape

Because our CS $\mathcal{C}_{\textnormal{ar}}^{\alpha}$ is defined through an inversion argument, it is interesting to study its shape. A simple sufficient condition for $\mathcal{C}_{\textnormal{ar}}^{\alpha}$ to be non-empty is that the bandwidth $\widehat{h}_M(c)$ is continuous in $c$, but beyond that it is difficult to make general statements. To see why, recall that $c\in \mathcal{C}_{\textnormal{ar}}^{\alpha}$ if and only if

align*[align* omitted — 159 chars of source]

The above quantities depend on $c$ directly, but also indirectly through $\widehat h_{M}(c)$. While the former dependence is rather simple in structure, the latter introduces complicated nonlinearities that make it impossible to give a simple analytical result regarding the shape of our CS. Such a result is possible, however, for a version that uses a fixed bandwidth.

theoremLet $ \mathcal{C}_{\textnormal{ar}}^{\alpha}(h)$ be a version of $\mathcal{C}_{\textnormal{ar}}^{\alpha}$ that uses a bandwidth $h$ that does not depend on $c$. Then either $\; \mathcal{C}_{\textnormal{ar}}^{\alpha}(h)=[a_1,a_2]$, or $\; \mathcal{C}_{\textnormal{ar}}^{\alpha}(h)=(-\infty,a_1]\cup[a_2,\infty)$, or $\; \mathcal{C}_{\textnormal{ar}}^{\alpha}(h)=(-\infty,\infty)$, or $\; \mathcal{C}_{\textnormal{ar}}^{\alpha}(h)=[a_1,\infty)$ or $\; \mathcal{C}_{\textnormal{ar}}^{\alpha}(h)=(-\infty,a_1]$, for some constants $a_1<a_2$.

The result mirrors the discussion at the end of Section (ref). It suggests that our actual CS should take one of these general shapes as long as $\widehat h_{M}(c)$ does not vary “too much” with $c$. We found this to be the case in every simulation run and every empirical analysis that we conducted in the context of this paper. The last two cases in Theorem (ref), in which $\mathcal{C}_{\textnormal{ar}}^{\alpha}(h)$ is a half-line, are also “knife-edge” cases: they only occur if one of the boundaries of a bias-aware CI for $\tau_T$ is exactly equal to zero, and are thus largely irrelevant for empirical practice.

Comparison with Delta Method Inference

Method and Limitations

The CIs commonly reported in empirical FRD studies are based on a linearization or “delta method” (DM) argument. It starts by noting that, after centering, the FRD point estimator $\widehat\theta(h) = \widehat\tau_Y(h)/\widehat\tau_T(h)$ can be written as the sum of an SRD-type estimator $\widehat\tau_U(h)$ with unobserved dependent variable $U_i$, and a remainder $\widehat\rho(h)$:

align*[align* omitted — 378 chars of source]

with $\widehat\tau_T^*(h)$ an intermediate value between $\tau_T$ and $\widehat\tau_T(h)$. One then imposes conditions under which $\widehat\rho(h)$ is asymptotically negligible relative to $\widehat\tau_U(h)$, and forms a CI for $\theta$ by applying some method for SRD inference to $\widehat\tau_U(h)$, which differ mainly in how they handle potential bias. Such DM CIs are proposed, for example, by calonico2014robust and armstrong2018simple in combination with robust bias correction and a bias-aware approach, respectively.\footnote{In empirical papers, FRD estimates are sometimes obtained through the two-stage least squares regression $Y_i = \theta T_i + \beta_+ X_iZ_i + \beta_-X_i(1-Z_i) + \varepsilon_{i}$ with $Z_i$ as an instrument for $T_i$, using only data in some window around the cutoff. This is numerically equivalent to a ratio of local linear regressions with a uniform kernel, and the resulting CI is thus of the DM type hahn2001identification,imbens2008regression.} As $U_i$ is unobserved, any such method must also be made feasible by using an estimate $\widehat U_i$ in which $\tau_Y$ and $\tau_T$ are replaced by suitable preliminary estimators.

Obvious downsides of such constructions include that they only control the bias of a first-order approximation of $\widehat\theta(h)$, and not the bias of $\widehat\theta(h)$ itself; and that replacing $U_i$ with an estimate $\widehat U_i$ introduces additional uncertainty, which is asymptotically second-order and hence generally unaccounted for in practice. In FRD designs such CIs are therefore generally subject to additional finite-sample distortions, relative to SRD designs.

A more principal issue with DM CIs is that a central condition for their validity, namely that $\widehat\rho(h)$ is asymptotically negligible relative to $\widehat\tau_U(h)$, is not innocuous. In particular, this condition is not compatible with a discrete running variable, or more generally one with support gaps around the cutoff. This is because $\tau_T$ and $\tau_Y$ are generally only partially identified in this case, and hence cannot be consistently estimated; see Section (ref). The term $\widehat\rho(h)$ then generally has a non-zero probability limit, and cannot be ignored for the purpose of inference on $\theta$. This issue occurs irrespective of the method chosen to control the bias of $\widehat\tau_U(h)$, including bias-aware inference. Because running variables with discrete or irregular support are ubiquitous in practice, this is an important limitation.

Another issue with DM CIs is that the conditions for their uniform validity rule out weakly identified settings with $\tau_T$ close to zero. This problem occurs even if the running variable is continuously distributed, and with any method chosen to control the bias of $\widehat\tau_U(h)$, including bias-aware inference. This is because for any DM CI to be honest with respect to $\mathcal{F}$, the term $\widehat\rho(h)$ must be of smaller order than $\widehat\tau_U(h)$ not only at the “true” function pair $(\mu_Y,\mu_T)$, but uniformly over all $(\mu_Y,\mu_T)\in\mathcal{F}$. But since $\tau_T$ can be arbitrarily close to zero over $(\mu_Y,\mu_T)\in\mathcal{F}$, we have that $\sup_{(\mu_Y,\mu_T)\in\mathcal{F}}| \widehat\rho(h)| = \infty$, which means that DM CIs can be unreliable in such settings.\footnote{feir2016weak also point out coverage issues of DM CIs under weak identification, but use different types of arguments. Specifically, they show that DM CIs based on infeasible “undersmoothing” bandwidths do not have correct asymptotic coverage under pointwise (with respect to the involved conditional expectation functions) asymptotics if $\tau_T$ tends to zero at an appropriate rate related to that of the bandwidth. They also show that undersmoothing AR CSs can have correct poinwise asymptotic coverage in this case.}

An Equivalence Result

armstrong2018simple study bias-aware DM CIs under conditions for which such DM CIs are asymptotically valid. These include Assumption (ref), which implies that is $X_i$ continuously distributed, and that $(\mu_Y,\mu_T) \in \mathcal{F}_H(B_Y) \times \mathcal{F}_H^\delta(B_T) \equiv \mathcal{F}^\delta$ for some $\delta>0$, which means that $\tau_T$ is well-separated from zero. They show that bias-aware DM CIs are honest with respect to $\mathcal{F}^\delta$ in this case, and also near-optimal, in the sense that no other method can substantially improve upon their length in large samples. The next theorem shows that our bias-aware AR CSs are as efficient as their DM counterparts in such settings for which DM CIs are specifically designed.

To avoid introducing additional high-level assumptions about the implementation details of bias-aware DM CIs we consider an infeasible version $\mathcal{C}_{\Delta}^{\alpha}$, formally defined in (ref), and compare it to its infeasible counterpart $\mathcal{C}_{*}^{\alpha}$ in our setup. Equal efficiency is established in the sense that both CSs have the same local asymptotic coverage for a drifting parameter within a $O(n^{-2/5})$ neighborhood of $\theta$. Such neighborhoods are appropriate to consider as the length of $\mathcal{C}_{\Delta}^{\alpha}$ is $O_P(n^{-2/5})$ uniformly over $\mathcal{F}^\delta$.

theoremSuppose that Assumptions (ref)--(ref) and (ref) hold, and put $\theta^{(n)}=\theta + \kappa n^{-2/5}$ for some constant $\kappa$. Then $$\limsup_{n\to\infty} \sup_{(\mu_Y,\mu_T)\in\mathcal{F}^\delta}\left| \mathbb{P}\left(\theta^{(n)}\in \mathcal{C}_{*}^{\alpha}\right) - \mathbb{P}\left(\theta^{(n)}\in \mathcal{C}_{\Delta}^{\alpha}\right)\right|=0.$$

This result parallels the well-known finding that there is no loss of efficiency when using the AR approach in exactly identified IV models relative to one based on a conventional $t$-test andrews2019weak. It is not an obvious corollary, however, as there are, for example, no analogues to the bandwidth and the smoothing bias in such IV models.

Implementation Details and Extensions

Standard Errors

Natural standard errors for $\widehat\tau_M(h,c)$ are of the form $\widehat s_{M}(h,c) = (\sum_{i=1}^n w_i(h)^2 \widehat\sigma_{M,i}^2(c))^{1/2}$, with $\widehat\sigma_{M,i}^2(c)$ some estimate of $\sigma_{M,i}^2(c)$. Setting $\widehat\sigma_{M,i}^2(c)$ to the squared difference between the outcome of unit $i$ and the average outcome among its nearest neighbors in terms of the running variable abadie2006large,abadie2014inference is commonly recommended in the RD literature calonico2014robust. However, this nearest-neighbor standard error is actually not uniformly consistent over $\mathcal{F}$ because the leading bias of $\widehat\sigma_{M,i}^2(c)$ is proportional to the first derivative of $\mu_M(\cdot,c)$ at $X_i$, which is unbounded over $\mathcal{F}$. We therefore propose a novel procedure that replaces the local sample average with a local best linear predictor. This modification makes the bias of $\widehat\sigma_{M,i}^2(c)$ proportional to the second derivative of $\mu_M(\cdot,c)$ at $X_i$, which is bounded in absolute value over $\mathcal{F}$ by $B_Y+|c| B_T$.

We propose a version that explicitly allows for ties among the realizations of the running variable. For $R$ a small integer, denote the rank of $|X_j -X_i|$ among the elements of the set $\{|X_{s} -X_i|: s\in\{1,\ldots,n\}\setminus\{i\}, X_{s} X_i > 0\}$ by $r(j,i)$, let $\mathcal{R}_i$ be the set of indices such that $r(j,i) \leq Q_i$, where $Q_i$ is the smallest integer such that $\mathcal{R}_i$ contains at least $R$ elements, and let $R_i$ be the resulting cardinality of $\mathcal{R}_i$.\footnote{Note that if every realization of $X_i$ is unique, then $R=Q_i=R_i$, and $\mathcal{R}_i$ is the set of unit $i$'s $R$ nearest neighbors' indices; but with ties in the data $R_i$ could be greater than $R$. We use $R=5$ in our simulations and the empirical application. } The estimator $\widehat{\sigma}^2_{M,i}(c)$ is then defined as the scaled squared difference between $M_i(c)$ and its best linear predictor given its $R_i$ nearest neighbors:

align*[align* omitted — 433 chars of source]

Here $\widetilde{X}_i = (1,X_i)^\top$ if the running variable takes at least two distinct values among the $R_i$ nearest neighbors of unit $i$, and $\widetilde{X}_i = 1$ otherwise. The scaling term $H_i$ ensures that $\widehat{\sigma}^2_{M,i}(c)$ is approximately unbiased in large samples. The next result, which we prove in Online Appendix (ref), shows that our new standard error is indeed uniformly consistent under general conditions. We recommend its use not just for our CS, but more generally for bias-aware inference methods that work with bounds on second derivatives.

theoremSuppose that Assumption (ref), Assumption (ref)(i), and either Assumption (ref) or Assumption (ref) are satisfied; that $\mathbb{V} (Y_i|X_i=x)$, $\mathbb{V} (T_i|X_i=x)$ and $\textrm{Cov}(Y_i, T_i|X_i=x)$ are Lipschitz continuous on each side of the cutoff uniformly over $x\in\textrm{supp}(X_i)$ and $(\mu_Y,\mu_T)\in\mathcal{F}$; and that $\mathbb{E}( (Y_i - \mathbb{E}(Y_i|X_i))^4| X_i=x)$ is uniformly bounded over $x\in\mathbb{R}$ and $(\mu_Y,\mu_T)\in\mathcal{F}$. Then Assumption (ref)(ii) holds for the standard error described in this subsection.

Bandwidth Choice

An obvious candidate for a feasible bandwidth is the empirical analogue of $h_M(c)$, which minimizes the length of the auxiliary CI in Section 4: $$\widehat{h}_M^* (c) = \operatorname*{argmin}_{h} \textnormal{cv}_{1-\alpha}(\widehat r_{M}(h,c)) \widehat s_{M}(h,c).$$ While this choice is generally attractive, it could lead to coverage distortions if $B_Y + |c| B_T$ is very large relative to sampling uncertainty. To see why, recall from the discussion at the end of Section (ref) that $\widehat\tau_{M}(h,c) =\sum_{i=1}^n w_{i}(h) M_i(c)$ is asymptotically normal if $w_{\textnormal{ratio}}(h)\equiv \max_{j=1,\ldots,n} w_j(h)^2/\sum_{i=1}^n w_i(h)^2=o_P(1)$. Normality should thus be a “good” finite-sample approximation if $w_{\textnormal{ratio}}(h)$ is “close” to zero (this reasoning also follows from a Berry-Esseen-type result). However, if $B_Y + |c| B_T$ (and thus the worst-case bias) is large, then $\widehat h_{M}^*(c)$ is typically small. The weights $w_i(\widehat h_{M}^*(c))$ then concentrate on few observations close to the cutoff, $w_{\textnormal{ratio}}(\widehat h_{M}^*(c))$ is large, and CLT approximations can be inaccurate as $\widehat\tau_{M}(\widehat h_{M}^*(c),c)$ then effectively behaves like a sample average of a small number of observations.

To address this issue, we propose imposing a lower bound on the bandwidth, chosen such that the value of $w_{\textnormal{ratio}}(h)$ remains below some reasonable threshold constant $\eta>0$, which we set to $\eta=.075$ in our simulations and empirical application.\footnote{To motivate this choice, suppose that $\mathcal{X}_n = \{\pm .02, \pm .04,\ldots,\pm 1\}$, that $K(t)=(1-|t|)\mathbf{1}\{|t|<1\}$ is the triangular kernel, and that $h=1$. Then $\widehat\tau_M(h,c)$ is a weighted least squares estimator that gives positive weight to 50 observations on each side of cutoff, and $w_{\textnormal{ratio}}(h)\approx .075$. } $$\widehat h_{M}(c) = \max\left\{\widehat h_{M}^*(c), h_{\min}(\eta)\right\}, \quad h_{\min}(\eta) = \min\left\{h: w_{\textnormal{ratio}}(h) < \eta\right\}.$$

Under standard conditions like Assumption (ref) or (ref) the lower bound on the bandwidth clearly never binds asymptotically, but imposing it can improve the finite-sample coverage of our CSs: as $\widehat{h}_M(c) \geq \widehat{h}_M^*(c)$, our construction trades off a possible increase in finite-sample bias against normality being a better finite-sample approximation. This improves coverage because our CSs explicitly account for the exact bias, but cannot capture deviations from normality. This idea can also be used for SRD inference, and more generally in all settings where finite-sample accuracy of inference faces a similar “bias vs.\ normality” trade-off. For example, armstrong2018finite use our approach in the context of inference on average treatment effects under unconfoundedness with limited overlap.

Choosing Smoothness Bounds

In order to compute $\mathcal{C}_{\textnormal{ar}}^{\alpha}$, researchers needs to choose the smoothness bounds $B_Y$ and $B_T$. Such bounds cannot be estimated consistently without imposing strong additional assumptions; and without choosing such bounds it is generally not possible to conduct inference on $\theta$ that is both valid and informative, even in large samples low1997nonparametric,armstrong2018optimal,bertanha2016impossible. Methods that seem to such choices still require restrictions on smoothness implicitly to attain approximately correct CI coverage in practice armstrong2018simple.\footnote{For example, an undersmoothing SRD CI can only be expected to have approximately correct coverage if the bias of the local linear estimator is “small” relative to its standard error. This can only be expected if the underlying function is “close” to linear, which is equivalent to its maximum second derivative being “close” to zero. A researcher that considers an undersmoothing SRD CI to be reliable has thus implicitly imposed a smoothness bound. An analogous argument applies to robust bias correction kamat2018nonparametric. } Explicitly specifying $B_Y$ and $B_T$ makes it transparent on which assumptions the inferential statements are based.

Roughly speaking, “small” values of $B_Y$ and $B_T$ amount to the assumption that the respective functions are “close” to linear on either side of the cutoff, whereas larger values allow the functions to be increasingly “curved”. The choice should be guided by subject knowledge but is arguably difficult in empirical practice, where there will be no single objectively correct value. In line with the previous literature, we hence recommend considering a range of plausible values as a form of sensitivity analysis. We also recommend estimating lowers bound $\widehat{B}_{Y,\textnormal{low}}$ and $\widehat{B}_{T,\textnormal{low}}$, and to compute one-sided CIs for $B_Y$ and $B_T$, respectively, via the methods proposed in armstrong2018optimal and kolesar2018discrete, to guard against overly optimistic choices.

Two heuristic “rules of thumb” (ROT) for determining plausible values in practice have also been considered in the literature. Both are based on fitting global polynomial specifications $\widetilde \mu_{Y,k}$ and $\widetilde \mu_{T,k}$ of order $k$ on either side of the cutoff by conventional least squares. armstrong2018simple use fourth-order polynomials, and propose the ROT $\widehat B_{Y,\textnormal{ROT1}}= \sup_{x\in\mathcal{X}}|\widetilde \mu_{Y,4}''(x)|$ and $\widehat B_{T,\textnormal{ROT1}}= \sup_{x\in\mathcal{X}}|\widetilde \mu_{T,4}''(x)|$, where $\mathcal{X}$ denotes the support of the running variable. imbens2019optimized consider a ROT in which the maximal curvature implied by a quadratic fit is multiplied by some moderate factor, say 2, yielding $\widehat B_{Y,\textnormal{\textnormal{ROT2}}}=2\sup_{x\in\mathcal{X}}|\widetilde \mu_{Y,2}''(x)|$ and $\widehat B_{T,\textnormal{\textnormal{ROT2}}}=2\sup_{x\in\mathcal{X}}|\widetilde \mu_{T,2}''(x)|$.

Such rules of thumb can provide useful first guidance, but should be complemented with other approaches in a sensitivity analysis. We strongly recommend to always check the fit of the respective polynomial specification, and to dismiss the ROT value if the fit is obviously poor. In Online Appendix (ref), we argue that in “roughly quadratic” settings the fourth-order polynomial specification that underlies ROT1 tends to produce quite erratic over-fits of the data that can lead to vast over-estimates of the true smoothness bounds, and corresponding CSs with poor statistical power. ROT2, on the other hand, tends to produce more reasonable values in many such setups. This pattern is also visible in our simulations.

Regression Kink Designs

In Online Appendix (ref), we present an extension of our approach to CSs for the ratio $\theta^{(v)}\equiv\tau_Y^{(v)} /\tau_T^{(v)} \equiv (\mu_{Y+}^{(v)}- \mu_{Y-}^{(v)})/(\mu_{T+}^{(v)}- \mu_{T-}^{(v)})$ of the jumps in the $v$th-order derivatives of two generic conditional expectation functions $(\mu_Y,\mu_T)$ at some threshold value. This setup covers the important Fuzzy Regression Kink Design card2015inference, where the parameter of interest is the ratio $\theta^{(1)}$ of two jumps in first derivatives.

We again form bias-aware CIs for an auxiliary parameter $\tau^{(v)}_M(c) = \tau^{(v)}_M - c\tau^{(v)}_T$, now based on $p$th-order local polynomial regression (where $p\geq v$ and typically $p=v+1$), and collect all values of $c$ for which such CIs contain zero to create an AR CS for $\theta^{(v)}$. The construction is largely analogous to that described in Section (ref), with the main difference concerning the bias bound. Specifically, we derive the apparently novel result that if $\mu_Y$ and $\mu_T$ are both $(p+1)$ times continuously differentiable on either side of the threshold, with derivatives of order $(p+1)$ uniformly bounded by $B_Y$ and $B_T$, respectively, and $\widehat\tau_{M,p}^{(v)}(h,c)=\sum_{i=1}^n w_{vp,i}(h)M_i(c)$ is the local $p$th order polynomial estimator of $\tau^{(v)}_M(c)$ with bandwidth $h$, the conditional bias $\mathbb{E}(\widehat\tau_{M,p}^{(v)}(h,c)|\mathcal{X}_n)-\tau_{M,p}^{(v)}(c)$ is absolutely bounded by

align*[align* omitted — 150 chars of source]

uniformly over the $(\mu_Y,\mu_T)$; see Online Appendix (ref) for details.

Numerical Illustrations

Empirical Application

figure[figure omitted — 30,600 chars of source]

In this subsection, we illustrate our methods revisiting data from battistin2009retirement, who study the effects of retirement on consumption in Italy. The data are a sample of $n=30,006$ individuals, obtained by combining several waves of the Bank of Italy Survey on Household Income and Wealth (SHIW) for the period 1993-–2004. We take the natural logarithm of total household spending as the outcome, retirement as the treatment, and years of age from the formal retirement eligibility threshold, which is normalized to zero, as the running variable. The running variable is thus discrete, but still has somewhat rich support. Figure (ref) shows the average of log consumption and the empirical proportion of retired individuals in the data as a function of the running variable.

We then compute our bias-aware AR CSs with both rules of thumb ROT1 and ROT2 to choose the smoothness bounds. The resulting CSs turn out to be intervals. We compare them to bias-aware DM CIs that use the same smoothness bounds, to robust bias correction DM CIs, to DM CIs based on global linear regression with separate intercepts and slopes on each side of the cutoff, and to DM CIs based on a global 4th order polynomial regression with a dummy for retirement eligibility.\footnote{In Online Appendix (ref), we also report results for variants of these CSs that use the optimized RD estimator of imbens2019optimized instead of local linear regression.} We report the results in Table (ref) in the form “midpoint $\pm$ half-length” to make comparing differences in the CIs' location and length easier.

table[table omitted — 903 chars of source]

We see that bias-aware AR CSs can differ meaningfully from their bias-aware DM CI counterparts in terms of both length and location, even if the same smoothness bounds are used. Our preferred rule ROT2 produces markedly smaller smoothness bounds than ROT1, which is reflected in the shorter CSs. Robust bias correction yields DM CIs that are qualitatively closer to those obtained under ROT1 by bias-aware methods. The two global parametric methods are the only ones that yield CIs that do not cover zero, but of course this does not account for the model misspecification bias apparent from Figure (ref).

Simulations

In this subsection, we compare the practical performance of our bias-aware AR CS to that of alternative procedures through simulations. We consider a number of data generating processes calibrated to the data from battistin2009retirement with varying curvature of the conditional expectation functions, richness of the running variable’s support, and strength of identification.

Data Generating Processes

We first create three versions of each of the two CEFs of outcomes and treatment, shown in Figure (ref). Specifically, for $W_i$ equal to either $T_i$ or $Y_i$, we create the $s$th CEF version $\mu_{W,s}(x)$ by fitting a second order spline with four knots on each side of the cutoff, that is,

align*[align* omitted — 280 chars of source]

via least squares to the data from battistin2009retirement, with $\lfloor \cdot \rfloor = \max(0, \cdot)$. Here our first version uses the knot points $v_j\in\{0, 10, 20, 30\}$; the second version uses the knot points $v_{j} \in\{0, 5, 15, 30\}$, creating greater curvature near the cutoff; while the third version also uses the knot points $v_{j} \in\{0, 5, 15, 30\}$ and additionally fixes the intercept parameters $\beta_{0,W+}$ and $\beta_{0,W-}$ to generate very high curvature near the cutoff.\footnote{For the conditional treatment probability, our least squares fits also impose the constraint that the function is increasing, and that it is equal to zero and one below and above the lowest and highest support point, respectively.} The magnitude of the jump in the conditional treatment probability, and hence the strength of identification, also decreases across versions. We then consider all combinations of these CEFs, but for brevity only report results for the four combinations of $\mu_{T,j}$ and $\mu_{Y,k}$ with $(j,k)\in\{(1,1),(2,2),(2,3),(3,2)\}$. See Online Appendix (ref) for the remaining results. In the following we refer to these settings as having either low, moderate or high CEF curvature.

figure[figure omitted — 25,269 chars of source]

For each combination of CEFs, we also consider four different distributions for the running variable $X_i$: the “mildly discrete” empirical distribution of the data from battistin2009retirement; a continuous distribution obtained by adding uniformly distributed $U(0,1)$ noise to a draw from the empirical distribution; and two more “coarsely discrete” distributions obtained by rounding draws from the empirical distribution to the integers $\{\pm 1, \pm(1+d), \pm(1+2d), \dots \}$ for $d\in \{3,6\}$. Finally, for any draw of $X_i$ we draw treatment status $T_i$ from a Bernoulli distribution with mean $\mu_{T,j}(X_i)$ and outcome $Y_i$ from a normal distribution with mean $\mu_{T,k}(X_i)$ and variance $\sigma^2(X_i)$, for $(j,k)\in\{1,3\}\times\{1,2,3\}$ and $\sigma^2(x)$ the sampling variance of $Y_i$ among units with running variable equal to the original draw of the running variable from the empirical distribution in the data.

Methods

We study the performance of eight different implementations of AR CSs in our simulations: (i) our bias-aware CS, using the respective true smoothness bounds $B_Y$ and $B_T$; (ii) our bias-aware CS, using twice the true $B_Y$ and $B_T$; (iii) our bias-aware CS, using half the true $B_Y$ and $B_T$; (iv) our bias-aware CS, using ROT1 estimates of $B_Y$ and $B_T$; (v) our bias-aware CS, using ROT2 estimates of $B_Y$ and $B_T$; (vi) a naive CS that ignores bias, using an estimate of the “pointwise-MSE optimal” bandwidth imbens2012optimal; (vii) an undersmoothing CS, using $n^{-1/20}$ times the estimated IK bandwidth; and (viii) a robust bias correction CS, using local quadratic regression to estimate the bias, and estimated IK bandwidths. In addition, we also consider the performance of eight different DM CIs using the just-mentioned approaches to handling bias.\footnote{ Computations are carried out with the statistical software R. All bias-aware CSs are computed using our own software, which builds on the package RDHonest. All other CSs are computed using functions from the package rdrobust. A triangular kernel is used in all cases. We note that the IK bandwidth estimates computed by rdrobust are sometimes too small for the respective CSs to be well-defined if the running variable is discrete. In those cases, we manually set the main bandwidth such that positive weights are given to three support points on each side of the cutoff (for the bias correction bandwidth we use four support points). In Online Appendix (ref), we also report results for variants of our CSs in which local linear regression is replaced with the optimized RD estimator of imbens2019optimized that are based on the package optrdd.}

Results

Table (ref) shows the simulated coverage rates of the various CSs under the sixteen different DGPs we consider in our simulations (four combinations of CEFs times four running variable distributions). We first discuss results for AR CSs, shown in the left panel. With the true smoothness bounds, the coverage rates of our bias-aware CSs are close to and mostly slightly above the nominal level, irrespective of running variable distribution, curvature of the unknown functions, and identification strength. The slight overcoverage occurs because the function $\mu_Y(x) - \theta\mu_T(x)$ is not exactly quadratic in either setting, and thus the bias does not achieve its worst-case value. Using twice the true bounds increases simulated coverage as expected, while half the true value results in meaningful undercoverage in some settings

Using one of the ROTs for the smoothness bounds leads to potentially severe distortions in some settings, which highlights the need to investigate the fit of the respective underlying global polynomial approximation in practice (cf.\ Online Appendix (ref)). Combining a naive approach, undersmoothing, or robust bias correction with an AR construction leads to CSs with undercoverage that is modest in some DGPs we consider, but can be substantial especially for those with more coarse running variable support, stronger curvature of the CEFs, and weak identification (high treatment CEF curvature).

Turning to results for DM CIs in the right panel of Table (ref), we see that combining a bias-aware approach with this construction does not lead to CIs with correct coverage in all settings even when using the true smoothness bound. This is because bias-aware DM CIs only control the bias of a first-order approximation of the estimator on which they are based. Coverage distortions are particularly severe in settings of strong curvature, and they further amplify in settings with weak identification. Using the ROT choices for the smoothness bounds leads to further distortions in some cases. The coverage of DM CIs that use the naive approach, undersmoothing, or robust bias correction is distorted in most settings, and the distortions generally become more severe with a more coarse support, in settings with a higher curvature and it further amplifies in settings of weak identification (high treatment CEF curvature).

\afterpage{

landscape\begin{table} \caption{Simulated CS coverage (%)} \resizebox*{1.35\textwidth}{!}{ \begin{threeparttable} \begin{tabular}{c@{\hskip 0.5in}rrrrrrrr@{\hskip 0.5in}rrrrrrrr} \toprule & \multicolumn{8}{c}{Anderson-Rubin} &\multicolumn{8}{c}{Delta Method} \\ \cmidrule[0.4pt](lr{0.3in}){2-9} \cmidrule[0.4pt](lr{0.1in}){10-17} & \multicolumn{5}{c}{Bias-Aware} && &&\multicolumn{5}{c}{Bias-Aware} &&&\\ \cmidrule{2-6} \cmidrule{10-14} Support & TC & TC$\times .5$ & TC$\times 2$ & ROT1 & ROT2 & Naive & US & RBC & TC & TC$\times .5$ & TC$\times 2$ & ROT1 & ROT2 & Naive & US & RBC \\ \midrule \multicolumn{13}{l}{Setting 1 - Low outcome CEF curvature, low treatment CEF curvature }&&&&\\ Baseline & 96.7 & 91.0 & 98.4 & 97.9 & 92.4 & 88.7 & 93.7 & 94.6 & 97.9 & 93.7 & 99.3 & 98.9 & 93.9 & 94.2 & 96.0 & 94.2 \\ Continuous & 96.6 & 92.7 & 97.4 & 97.1 & 94.3 & 89.6 & 94.0 & 94.7 & 97.2 & 94.4 & 98.0 & 97.8 & 95.1 & 94.7 & 95.4 & 94.7 \\ $\{\pm 1, \pm 4, \dots \}$ & 95.9 & 92.6 & 98.4 & 97.8 & 94.4 & 88.1 & 93.8 & 94.6 & 96.7 & 93.8 & 99.3 & 98.7 & 95.0 & 94.0 & 89.3 & 94.6 \\ $\{\pm 1, \pm 7, \dots \}$ & 97.8 & 91.6 & 100.0 & 99.6 & 95.5 & 78.0 & 84.8 & 93.2 & 98.6 & 92.3 & 100.0 & 99.8 & 96.0 & 85.1 & 77.5 & 93.1 \\ \multicolumn{13}{l}{\phantom{a}}&&&&\\ \multicolumn{13}{l}{Setting 2 - Moderate outcome CEF curvature, moderate treatment CEF curvature }&&&&\\ Baseline & 97.2 & 88.7 & 99.2 & 91.9 & 36.4 & 47.6 & 80.9 & 58.6 & 91.1 & 74.0 & 98.3 & 81.9 & 24.7 & 73.3 & 93.0 & 72.9 \\ Continuous & 96.4 & 92.2 & 97.3 & 94.3 & 59.8 & 68.6 & 87.6 & 78.2 & 93.7 & 84.7 & 96.5 & 89.3 & 46.6 & 84.9 & 94.5 & 86.3 \\ $\{\pm 1, \pm 4, \dots \}$ & 98.3 & 88.8 & 100 & 92.1 & 63.5 & 27.5 & 71.1 & 45.6 & 92.7 & 81.8 & 99.0 & 86.0 & 46.6 & 70.7 & 65.9 & 71.3 \\ $\{\pm 1, \pm 7, \dots \}$ & 100 & 79.9 & 100 & 92.1 & 26.2 & 1.3 & 6.5 & 0.3 & 95.4 & 56.0 & 100 & 78.2 & 19.8 & 6.0 & 5.5 & 4.7 \\ \multicolumn{13}{l}{\phantom{a}}&&&&\\ \multicolumn{13}{l}{Setting 3 - High outcome CEF curvature, moderate treatment CEF curvature }&&&&\\ Baseline & 96.5 & 80.9 & 99.3 & 72.2 & 0.0 & 84.2 & 83.1 & 92.9 & 82.3 & 48.8 & 93.9 & 33.0 & 0.0 & 14.6 & 69.1 & 26.0 \\ Continuous & 95.8 & 90.1 & 97.1 & 90.5 & 1.7 & 89.4 & 94.1 & 94.3 & 89.5 & 75.9 & 92.1 & 77.3 & 0.0 & 48.6 & 85.6 & 66.0 \\ $\{\pm 1, \pm 4, \dots \}$ & 98.6 & 52.1 & 100 & 55.0 & 6.1 & 4.9 & 5.0 & 84.0 & 78.7 & 25.5 & 100 & 27.9 & 0.0 & 1.9 & 1.9 & 5.9 \\ $\{\pm 1, \pm 7, \dots \}$ & 100 & 2.7 & 100 & 6.6 & 0.0 & 0.0 & 2.4 & 0.0 & 61.0 & 0.0 & 100 & 0.1 & 0.0 & 0.0 & 0.0 & 0.0 \\ \multicolumn{13}{l}{\phantom{a}}&&&&\\ \multicolumn{13}{l}{\textbf{Setting 4} - \textit{Moderate outcome CEF curvature, high treatment CEF curvature} }&&&&\\ Baseline & 96.8 & 82.1 & 99.4 & 73.0 & 0.0 & 85.9 & 84.2 & 93.4 & 23.5 & 0.1 & 58.5 & 1.2 & 0.0 & 14.2 & 55.0 & 16.6 \\ Continuous & 95.8 & 90.1 & 97.0 & 90.9 & 0.8 & 89.7 & 94.0 & 94.4 & 50.8 & 2.1 & 71.3 & 16.3 & 0.0 & 53.0 & 72.8 & 58.2 \\ $\{\pm 1, \pm 4, \dots \}$ & 99.0 & 55.9 & 100 & 60.5 & 6.6 & 5.3 & 5.1 & 79.5 & 12.8 & 3.7 & 40.6 & 5.3 & 0.0 & 0.9 & 0.8 & 3.9 \\ $\{\pm 1, \pm 7, \dots \}$ & 100 & 5.2 & 100 & 17.6 & 0.0 & 0.0 & 2.0 & 0.0 & 0.0 & 0.0 & 3.1 & 0.0 & 0.0 & 0.0 & 1.0 & 0.0 \\ \bottomrule \end{tabular} \begin{tablenotes} • \textit{Notes:} Results based on 50,000 Monte Carlo draws for a nominal confidence level of 95%. Columns show results for bias aware approach with true constants (TC), two times true constants (TC$\times2$), half true constants (TC$\times.5$), and with rule of thumb estimates (ROT1) and (ROT2); naive approach that ignores bias (Naive); undersmoothing (US); and robust bias correction (RBC). \end{tablenotes} \end{threeparttable}} \end{table}

}

\afterpage{

landscape\begin{figure}[!t] \begin{minipage}{.65\textwidth} (a) Bias-aware Anderson-Rubin CSs \resizebox{\linewidth}{!}{ \begin{tikzpicture}[x=1pt,y=1pt] \definecolor{fillColor}{RGB}{255,255,255} \path[use as bounding box,fill=fillColor,fill opacity=0.00] (0,0) rectangle (361.35,252.94); \begin{scope} \path[clip] ( 49.20, 61.20) rectangle (336.15,203.75); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.48) -- ( 15.55, 66.48) -- ( 45.07, 66.49) -- ( 56.42, 66.50) -- ( 67.78, 66.54) -- ( 79.13, 66.76) -- ( 90.48, 67.34) -- (101.84, 69.08) -- (113.19, 73.24) -- (124.55, 82.15) -- (135.90, 98.39) -- (147.26,121.36) -- (158.61,146.96) -- (169.97,167.83) -- (181.32,180.15) -- (192.68,183.50) -- (204.03,179.99) -- (215.38,169.25) -- (226.74,150.77) -- (238.09,127.19) -- (249.45,104.25) -- (260.80, 86.51) -- (272.16, 75.73) -- (283.51, 70.24) -- (294.87, 67.89) -- (306.22, 66.98) -- (317.57, 66.65) -- (328.93, 66.51) -- (340.28, 66.48) -- (369.80, 66.48) -- (414.09, 66.48); \end{scope} \begin{scope} \path[clip] ( 0.00, 0.00) rectangle (361.35,252.94); \definecolor{drawColor}{RGB}{0,0,0} \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (192.68, 15.60) {Distance to theta}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 10.80,132.47) {Simulated Coverage Probability}; \end{scope} \begin{scope} \path[clip] ( 49.20, 61.20) rectangle (336.15,203.75); \definecolor{drawColor}{RGB}{27,158,119} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.48) -- ( 15.55, 66.48) -- ( 45.07, 66.48) -- ( 56.42, 66.48) -- ( 67.78, 66.48) -- ( 79.13, 66.50) -- ( 90.48, 66.56) -- (101.84, 66.97) -- (113.19, 68.62) -- (124.55, 73.70) -- (135.90, 86.99) -- (147.26,111.41) -- (158.61,142.93) -- (169.97,169.16) -- (181.32,182.06) -- (192.68,182.31) -- (204.03,172.48) -- (215.38,151.49) -- (226.74,123.01) -- (238.09, 96.42) -- (249.45, 78.91) -- (260.80, 70.54) -- (272.16, 67.64) -- (283.51, 66.77) -- (294.87, 66.53) -- (306.22, 66.48) -- (317.57, 66.48) -- (328.93, 66.48) -- (340.28, 66.48) -- (369.80, 66.48) -- (414.09, 66.48); \definecolor{drawColor}{RGB}{217,95,2} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.48) -- ( 15.55, 66.49) -- ( 45.07, 66.62) -- ( 56.42, 66.84) -- ( 67.78, 67.29) -- ( 79.13, 68.35) -- ( 90.48, 70.49) -- (101.84, 74.69) -- (113.19, 82.63) -- (124.55, 95.07) -- (135.90,112.74) -- (147.26,133.51) -- (158.61,153.73) -- (169.97,169.98) -- (181.32,179.70) -- (192.68,183.42) -- (204.03,182.12) -- (215.38,176.00) -- (226.74,164.26) -- (238.09,147.33) -- (249.45,127.37) -- (260.80,107.99) -- (272.16, 92.21) -- (283.51, 80.82) -- (294.87, 74.00) -- (306.22, 70.26) -- (317.57, 68.25) -- (328.93, 67.29) -- (340.28, 66.82) -- (369.80, 66.52) -- (414.09, 66.48); \definecolor{drawColor}{RGB}{117,112,179} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.48) -- ( 15.55, 66.49) -- ( 45.07, 66.61) -- ( 56.42, 66.78) -- ( 67.78, 67.17) -- ( 79.13, 68.02) -- ( 90.48, 69.87) -- (101.84, 73.51) -- (113.19, 80.65) -- (124.55, 92.72) -- (135.90,110.36) -- (147.26,132.18) -- (158.61,153.30) -- (169.97,170.12) -- (181.32,179.85) -- (192.68,183.26) -- (204.03,181.55) -- (215.38,174.59) -- (226.74,161.59) -- (238.09,143.87) -- (249.45,123.88) -- (260.80,105.34) -- (272.16, 90.58) -- (283.51, 80.24) -- (294.87, 73.85) -- (306.22, 70.36) -- (317.57, 68.41) -- (328.93, 67.42) -- (340.28, 66.93) -- (369.80, 66.56) -- (414.09, 66.48); \definecolor{drawColor}{RGB}{230,171,2} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.48) -- ( 15.55, 66.48) -- ( 45.07, 66.48) -- ( 56.42, 66.48) -- ( 67.78, 66.48) -- ( 79.13, 66.50) -- ( 90.48, 66.55) -- (101.84, 66.96) -- (113.19, 68.72) -- (124.55, 74.34) -- (135.90, 88.61) -- (147.26,113.44) -- (158.61,143.93) -- (169.97,168.30) -- (181.32,181.00) -- (192.68,182.68) -- (204.03,175.50) -- (215.38,157.91) -- (226.74,131.67) -- (238.09,104.21) -- (249.45, 84.40) -- (260.80, 73.08) -- (272.16, 68.56) -- (283.51, 67.06) -- (294.87, 66.59) -- (306.22, 66.50) -- (317.57, 66.48) -- (328.93, 66.48) -- (340.28, 66.48) -- (369.80, 66.48) -- (414.09, 66.48); \end{scope} \begin{scope} \path[clip] ( 0.00, 0.00) rectangle (361.35,252.94); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20, 61.20) -- (336.15, 61.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 81.97, 61.20) -- ( 81.97, 55.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (137.32, 61.20) -- (137.32, 55.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (192.68, 61.20) -- (192.68, 55.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (248.03, 61.20) -- (248.03, 55.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (303.38, 61.20) -- (303.38, 55.20); \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 81.97, 39.60) {-0.5}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (137.32, 39.60) {-0.25}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (192.68, 39.60) {0}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (248.03, 39.60) {0.25}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (303.38, 39.60) {0.5}; \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20, 61.20) -- ( 49.20,203.75); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20, 66.48) -- ( 43.20, 66.48); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20, 90.48) -- ( 43.20, 90.48); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20,114.47) -- ( 43.20,114.47); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20,138.47) -- ( 43.20,138.47); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20,162.47) -- ( 43.20,162.47); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20,186.47) -- ( 43.20,186.47); \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20, 63.04) {0}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20, 87.03) {0.2}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20,111.03) {0.4}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20,135.03) {0.6}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20,159.03) {0.8}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20,183.02) {1}; \end{scope} \begin{scope} \path[clip] ( 49.20, 61.20) rectangle (336.15,203.75); \definecolor{drawColor}{RGB}{169,169,169} \path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] ( 49.20,180.47) -- (336.15,180.47); \path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] (200.00, 61.20) -- (200.00,203.75); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,169.67) -- (283.40,169.67); \definecolor{drawColor}{RGB}{27,158,119} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,158.87) -- (283.40,158.87); \definecolor{drawColor}{RGB}{217,95,2} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,148.07) -- (283.40,148.07); \definecolor{drawColor}{RGB}{117,112,179} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,137.27) -- (283.40,137.27); \definecolor{drawColor}{RGB}{230,171,2} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,126.47) -- (283.40,126.47); \definecolor{drawColor}{RGB}{0,0,0} \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,166.57) {TC-AR}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,155.77) {TC x 0.5}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,144.97) {TC x 2}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,134.17) {ROT 1}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,123.37) {ROT 2}; \end{scope} \end{tikzpicture} } \end{minipage} \begin{minipage}{.65\textwidth} (b) Other Anderson-Rubin CSs \resizebox{\linewidth}{!}{ \begin{tikzpicture}[x=1pt,y=1pt] \definecolor{fillColor}{RGB}{255,255,255} \path[use as bounding box,fill=fillColor,fill opacity=0.00] (0,0) rectangle (361.35,252.94); \begin{scope} \path[clip] ( 49.20, 61.20) rectangle (336.15,203.75); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.48) -- ( 15.55, 66.48) -- ( 45.07, 66.49) -- ( 56.42, 66.50) -- ( 67.78, 66.54) -- ( 79.13, 66.76) -- ( 90.48, 67.34) -- (101.84, 69.08) -- (113.19, 73.24) -- (124.55, 82.15) -- (135.90, 98.39) -- (147.26,121.36) -- (158.61,146.96) -- (169.97,167.83) -- (181.32,180.15) -- (192.68,183.50) -- (204.03,179.99) -- (215.38,169.25) -- (226.74,150.77) -- (238.09,127.19) -- (249.45,104.25) -- (260.80, 86.51) -- (272.16, 75.73) -- (283.51, 70.24) -- (294.87, 67.89) -- (306.22, 66.98) -- (317.57, 66.65) -- (328.93, 66.51) -- (340.28, 66.48) -- (369.80, 66.48) -- (414.09, 66.48); \end{scope} \begin{scope} \path[clip] ( 0.00, 0.00) rectangle (361.35,252.94); \definecolor{drawColor}{RGB}{0,0,0} \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (192.68, 15.60) {Parameter Value}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 10.80,132.47) {Simulated Coverage Probability}; \end{scope} \begin{scope} \path[clip] ( 49.20, 61.20) rectangle (336.15,203.75); \definecolor{drawColor}{RGB}{66,134,244} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.48) -- ( 15.55, 66.48) -- ( 45.07, 66.49) -- ( 56.42, 66.49) -- ( 67.78, 66.50) -- ( 79.13, 66.52) -- ( 90.48, 66.57) -- (101.84, 66.73) -- (113.19, 67.07) -- (124.55, 68.47) -- (135.90, 74.95) -- (147.26, 95.17) -- (158.61,127.91) -- (169.97,157.34) -- (181.32,174.18) -- (192.68,177.51) -- (204.03,169.71) -- (215.38,153.04) -- (226.74,131.75) -- (238.09,110.47) -- (249.45, 93.07) -- (260.80, 80.90) -- (272.16, 73.85) -- (283.51, 70.08) -- (294.87, 68.25) -- (306.22, 67.33) -- (317.57, 66.92) -- (328.93, 66.69) -- (340.28, 66.58) -- (369.80, 66.51) -- (414.09, 66.48); \definecolor{drawColor}{RGB}{102,166,30} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.50) -- ( 15.55, 66.53) -- ( 45.07, 66.66) -- ( 56.42, 66.76) -- ( 67.78, 66.91) -- ( 79.13, 67.19) -- ( 90.48, 67.58) -- (101.84, 68.29) -- (113.19, 70.13) -- (124.55, 75.67) -- (135.90, 90.27) -- (147.26,114.56) -- (158.61,141.84) -- (169.97,163.01) -- (181.32,175.34) -- (192.68,179.75) -- (204.03,177.54) -- (215.38,169.52) -- (226.74,156.97) -- (238.09,141.74) -- (249.45,125.58) -- (260.80,110.36) -- (272.16, 97.45) -- (283.51, 87.32) -- (294.87, 79.92) -- (306.22, 75.08) -- (317.57, 71.83) -- (328.93, 69.84) -- (340.28, 68.63) -- (369.80, 67.13) -- (414.09, 66.62); \definecolor{drawColor}{RGB}{166,118,29} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.48) -- ( 15.55, 66.48) -- ( 45.07, 66.50) -- ( 56.42, 66.52) -- ( 67.78, 66.57) -- ( 79.13, 66.69) -- ( 90.48, 66.87) -- (101.84, 67.25) -- (113.19, 68.10) -- (124.55, 70.19) -- (135.90, 76.32) -- (147.26, 93.00) -- (158.61,120.55) -- (169.97,149.75) -- (181.32,169.76) -- (192.68,179.00) -- (204.03,178.70) -- (215.38,170.16) -- (226.74,154.44) -- (238.09,134.77) -- (249.45,114.27) -- (260.80, 97.32) -- (272.16, 84.70) -- (283.51, 76.69) -- (294.87, 71.92) -- (306.22, 69.43) -- (317.57, 68.02) -- (328.93, 67.31) -- (340.28, 66.93) -- (369.80, 66.57) -- (414.09, 66.49); \end{scope} \begin{scope} \path[clip] ( 0.00, 0.00) rectangle (361.35,252.94); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20, 61.20) -- (336.15, 61.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 81.97, 61.20) -- ( 81.97, 55.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (137.32, 61.20) -- (137.32, 55.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (192.68, 61.20) -- (192.68, 55.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (248.03, 61.20) -- (248.03, 55.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (303.38, 61.20) -- (303.38, 55.20); \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 81.97, 39.60) {-0.5}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (137.32, 39.60) {-0.25}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (192.68, 39.60) {0}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (248.03, 39.60) {0.25}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (303.38, 39.60) {0.5}; \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20, 61.20) -- ( 49.20,203.75); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20, 66.48) -- ( 43.20, 66.48); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20, 90.48) -- ( 43.20, 90.48); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20,114.47) -- ( 43.20,114.47); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20,138.47) -- ( 43.20,138.47); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20,162.47) -- ( 43.20,162.47); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20,186.47) -- ( 43.20,186.47); \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20, 63.04) {0}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20, 87.03) {0.2}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20,111.03) {0.4}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20,135.03) {0.6}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20,159.03) {0.8}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20,183.02) {1}; \end{scope} \begin{scope} \path[clip] ( 49.20, 61.20) rectangle (336.15,203.75); \definecolor{drawColor}{RGB}{169,169,169} \path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] ( 49.20,180.47) -- (336.15,180.47); \path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] (200.00, 61.20) -- (200.00,203.75); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,169.67) -- (283.40,169.67); \definecolor{drawColor}{RGB}{66,134,244} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,158.87) -- (283.40,158.87); \definecolor{drawColor}{RGB}{102,166,30} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,148.07) -- (283.40,148.07); \definecolor{drawColor}{RGB}{166,118,29} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,137.27) -- (283.40,137.27); \definecolor{drawColor}{RGB}{0,0,0} \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,166.57) {TC-AR}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,155.77) {Naive}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,144.97) {US}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,134.17) {RBC}; \end{scope} \end{tikzpicture} } \end{minipage} \begin{minipage}{.65\textwidth} (c) Bias-Aware Delta Method CI \resizebox{\linewidth}{!}{ \begin{tikzpicture}[x=1pt,y=1pt] \definecolor{fillColor}{RGB}{255,255,255} \path[use as bounding box,fill=fillColor,fill opacity=0.00] (0,0) rectangle (361.35,252.94); \begin{scope} \path[clip] ( 49.20, 61.20) rectangle (336.15,203.75); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.48) -- ( 15.55, 66.48) -- ( 45.07, 66.49) -- ( 56.42, 66.50) -- ( 67.78, 66.54) -- ( 79.13, 66.76) -- ( 90.48, 67.34) -- (101.84, 69.08) -- (113.19, 73.24) -- (124.55, 82.15) -- (135.90, 98.39) -- (147.26,121.36) -- (158.61,146.96) -- (169.97,167.83) -- (181.32,180.15) -- (192.68,183.50) -- (204.03,179.99) -- (215.38,169.25) -- (226.74,150.77) -- (238.09,127.19) -- (249.45,104.25) -- (260.80, 86.51) -- (272.16, 75.73) -- (283.51, 70.24) -- (294.87, 67.89) -- (306.22, 66.98) -- (317.57, 66.65) -- (328.93, 66.51) -- (340.28, 66.48) -- (369.80, 66.48) -- (414.09, 66.48); \end{scope} \begin{scope} \path[clip] ( 0.00, 0.00) rectangle (361.35,252.94); \definecolor{drawColor}{RGB}{0,0,0} \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (192.68, 15.60) {Parameter Value}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 10.80,132.47) {Simulated Coverage Probability}; \end{scope} \begin{scope} \path[clip] ( 49.20, 61.20) rectangle (336.15,203.75); \definecolor{drawColor}{gray}{0.40} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.48) -- ( 15.55, 66.48) -- ( 45.07, 66.48) -- ( 56.42, 66.48) -- ( 67.78, 66.50) -- ( 79.13, 66.54) -- ( 90.48, 66.79) -- (101.84, 67.72) -- (113.19, 70.43) -- (124.55, 77.17) -- (135.90, 91.96) -- (147.26,115.47) -- (158.61,144.99) -- (169.97,168.26) -- (181.32,180.54) -- (192.68,183.98) -- (204.03,180.99) -- (215.38,169.06) -- (226.74,144.51) -- (238.09,114.60) -- (249.45, 92.46) -- (260.80, 77.95) -- (272.16, 70.77) -- (283.51, 67.92) -- (294.87, 66.91) -- (306.22, 66.58) -- (317.57, 66.49) -- (328.93, 66.48) -- (340.28, 66.48) -- (369.80, 66.48) -- (414.09, 66.48); \definecolor{drawColor}{RGB}{27,158,119} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.48) -- ( 15.55, 66.48) -- ( 45.07, 66.48) -- ( 56.42, 66.48) -- ( 67.78, 66.48) -- ( 79.13, 66.48) -- ( 90.48, 66.51) -- (101.84, 66.69) -- (113.19, 67.65) -- (124.55, 71.41) -- (135.90, 82.86) -- (147.26,107.69) -- (158.61,142.14) -- (169.97,169.89) -- (181.32,182.28) -- (192.68,183.50) -- (204.03,174.52) -- (215.38,148.31) -- (226.74,108.47) -- (238.09, 84.72) -- (249.45, 72.24) -- (260.80, 67.92) -- (272.16, 66.76) -- (283.51, 66.50) -- (294.87, 66.48) -- (306.22, 66.48) -- (317.57, 66.48) -- (328.93, 66.48) -- (340.28, 66.48) -- (369.80, 66.48) -- (414.09, 66.48); \definecolor{drawColor}{RGB}{217,95,2} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.48) -- ( 15.55, 66.48) -- ( 45.07, 66.50) -- ( 56.42, 66.54) -- ( 67.78, 66.67) -- ( 79.13, 67.09) -- ( 90.48, 68.19) -- (101.84, 70.83) -- (113.19, 76.28) -- (124.55, 86.68) -- (135.90,103.76) -- (147.26,127.03) -- (158.61,152.16) -- (169.97,170.75) -- (181.32,180.68) -- (192.68,184.12) -- (204.03,182.99) -- (215.38,176.55) -- (226.74,162.06) -- (238.09,138.86) -- (249.45,115.09) -- (260.80, 95.67) -- (272.16, 81.87) -- (283.51, 73.69) -- (294.87, 69.64) -- (306.22, 67.73) -- (317.57, 66.95) -- (328.93, 66.64) -- (340.28, 66.53) -- (369.80, 66.48) -- (414.09, 66.48); \definecolor{drawColor}{RGB}{117,112,179} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.48) -- ( 15.55, 66.48) -- ( 45.07, 66.50) -- ( 56.42, 66.55) -- ( 67.78, 66.67) -- ( 79.13, 66.99) -- ( 90.48, 67.96) -- (101.84, 70.16) -- (113.19, 75.00) -- (124.55, 84.75) -- (135.90,101.71) -- (147.26,125.53) -- (158.61,151.67) -- (169.97,170.79) -- (181.32,180.74) -- (192.68,184.04) -- (204.03,182.60) -- (215.38,175.03) -- (226.74,158.44) -- (238.09,134.66) -- (249.45,112.25) -- (260.80, 94.49) -- (272.16, 82.05) -- (283.51, 74.11) -- (294.87, 70.08) -- (306.22, 68.13) -- (317.57, 67.15) -- (328.93, 66.76) -- (340.28, 66.59) -- (369.80, 66.48) -- (414.09, 66.48); \definecolor{drawColor}{RGB}{230,171,2} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.48) -- ( 15.55, 66.48) -- ( 45.07, 66.48) -- ( 56.42, 66.48) -- ( 67.78, 66.48) -- ( 79.13, 66.49) -- ( 90.48, 66.52) -- (101.84, 66.72) -- (113.19, 67.93) -- (124.55, 72.42) -- (135.90, 85.35) -- (147.26,110.45) -- (158.61,142.83) -- (169.97,168.50) -- (181.32,181.24) -- (192.68,183.37) -- (204.03,176.50) -- (215.38,156.93) -- (226.74,125.59) -- (238.09, 97.34) -- (249.45, 79.38) -- (260.80, 70.56) -- (272.16, 67.60) -- (283.51, 66.74) -- (294.87, 66.52) -- (306.22, 66.48) -- (317.57, 66.48) -- (328.93, 66.48) -- (340.28, 66.48) -- (369.80, 66.48) -- (414.09, 66.48); \end{scope} \begin{scope} \path[clip] ( 0.00, 0.00) rectangle (361.35,252.94); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20, 61.20) -- (336.15, 61.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 81.97, 61.20) -- ( 81.97, 55.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (137.32, 61.20) -- (137.32, 55.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (192.68, 61.20) -- (192.68, 55.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (248.03, 61.20) -- (248.03, 55.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (303.38, 61.20) -- (303.38, 55.20); \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 81.97, 39.60) {-0.5}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (137.32, 39.60) {-0.25}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (192.68, 39.60) {0}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (248.03, 39.60) {0.25}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (303.38, 39.60) {0.5}; \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20, 61.20) -- ( 49.20,203.75); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20, 66.48) -- ( 43.20, 66.48); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20, 90.48) -- ( 43.20, 90.48); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20,114.47) -- ( 43.20,114.47); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20,138.47) -- ( 43.20,138.47); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20,162.47) -- ( 43.20,162.47); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20,186.47) -- ( 43.20,186.47); \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20, 63.04) {0}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20, 87.03) {0.2}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20,111.03) {0.4}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20,135.03) {0.6}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20,159.03) {0.8}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20,183.02) {1}; \end{scope} \begin{scope} \path[clip] ( 49.20, 61.20) rectangle (336.15,203.75); \definecolor{drawColor}{RGB}{169,169,169} \path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] ( 49.20,180.47) -- (336.15,180.47); \path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] (200.00, 61.20) -- (200.00,203.75); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,169.67) -- (283.40,169.67); \definecolor{drawColor}{gray}{0.40} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,158.87) -- (283.40,158.87); \definecolor{drawColor}{RGB}{27,158,119} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,148.07) -- (283.40,148.07); \definecolor{drawColor}{RGB}{217,95,2} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,137.27) -- (283.40,137.27); \definecolor{drawColor}{RGB}{117,112,179} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,126.47) -- (283.40,126.47); \definecolor{drawColor}{RGB}{230,171,2} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,115.67) -- (283.40,115.67); \definecolor{drawColor}{RGB}{0,0,0} \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,166.57) {TC-AR}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,155.77) {TC-DM}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,144.97) {TC x 0.5}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,134.17) {TC x 2}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,123.37) {ROT 1}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,112.57) {ROT 2}; \end{scope} \end{tikzpicture} } \end{minipage} \begin{minipage}{.65\textwidth} (d) Other Delta Method CIs \resizebox{\linewidth}{!}{ \begin{tikzpicture}[x=1pt,y=1pt] \definecolor{fillColor}{RGB}{255,255,255} \path[use as bounding box,fill=fillColor,fill opacity=0.00] (0,0) rectangle (361.35,252.94); \begin{scope} \path[clip] ( 49.20, 61.20) rectangle (336.15,203.75); \definecolor{drawColor}{gray}{0.40} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.48) -- ( 15.55, 66.48) -- ( 45.07, 66.48) -- ( 56.42, 66.48) -- ( 67.78, 66.50) -- ( 79.13, 66.54) -- ( 90.48, 66.79) -- (101.84, 67.72) -- (113.19, 70.43) -- (124.55, 77.17) -- (135.90, 91.96) -- (147.26,115.47) -- (158.61,144.99) -- (169.97,168.26) -- (181.32,180.54) -- (192.68,183.98) -- (204.03,180.99) -- (215.38,169.06) -- (226.74,144.51) -- (238.09,114.60) -- (249.45, 92.46) -- (260.80, 77.95) -- (272.16, 70.77) -- (283.51, 67.92) -- (294.87, 66.91) -- (306.22, 66.58) -- (317.57, 66.49) -- (328.93, 66.48) -- (340.28, 66.48) -- (369.80, 66.48) -- (414.09, 66.48); \end{scope} \begin{scope} \path[clip] ( 0.00, 0.00) rectangle (361.35,252.94); \definecolor{drawColor}{RGB}{0,0,0} \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (192.68, 15.60) {Parameter Value}; \node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 10.80,132.47) {Simulated Coverage Probability}; \end{scope} \begin{scope} \path[clip] ( 49.20, 61.20) rectangle (336.15,203.75); \definecolor{drawColor}{RGB}{66,134,244} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.48) -- ( 15.55, 66.48) -- ( 45.07, 66.49) -- ( 56.42, 66.50) -- ( 67.78, 66.54) -- ( 79.13, 66.67) -- ( 90.48, 67.04) -- (101.84, 68.20) -- (113.19, 71.32) -- (124.55, 78.10) -- (135.90, 91.32) -- (147.26,112.22) -- (158.61,137.25) -- (169.97,159.62) -- (181.32,174.78) -- (192.68,180.55) -- (204.03,177.79) -- (215.38,166.58) -- (226.74,147.01) -- (238.09,122.99) -- (249.45,100.26) -- (260.80, 83.72) -- (272.16, 74.16) -- (283.51, 69.73) -- (294.87, 67.74) -- (306.22, 66.94) -- (317.57, 66.63) -- (328.93, 66.53) -- (340.28, 66.49) -- (369.80, 66.48) -- (414.09, 66.48); \definecolor{drawColor}{RGB}{102,166,30} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.49) -- ( 15.55, 66.55) -- ( 45.07, 66.95) -- ( 56.42, 67.43) -- ( 67.78, 68.28) -- ( 79.13, 70.05) -- ( 90.48, 73.20) -- (101.84, 78.75) -- (113.19, 87.44) -- (124.55,100.30) -- (135.90,116.51) -- (147.26,134.83) -- (158.61,152.29) -- (169.97,166.73) -- (181.32,176.18) -- (192.68,180.66) -- (204.03,180.44) -- (215.38,175.50) -- (226.74,165.78) -- (238.09,151.64) -- (249.45,134.33) -- (260.80,116.28) -- (272.16,100.47) -- (283.51, 87.80) -- (294.87, 79.23) -- (306.22, 73.55) -- (317.57, 70.36) -- (328.93, 68.61) -- (340.28, 67.58) -- (369.80, 66.67) -- (414.09, 66.50); \definecolor{drawColor}{RGB}{166,118,29} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (-28.74, 66.48) -- ( 15.55, 66.48) -- ( 45.07, 66.50) -- ( 56.42, 66.53) -- ( 67.78, 66.60) -- ( 79.13, 66.90) -- ( 90.48, 67.57) -- (101.84, 69.26) -- (113.19, 72.91) -- (124.55, 80.30) -- (135.90, 93.31) -- (147.26,112.16) -- (158.61,134.36) -- (169.97,155.25) -- (181.32,170.59) -- (192.68,178.72) -- (204.03,179.13) -- (215.38,172.41) -- (226.74,158.36) -- (238.09,138.58) -- (249.45,116.72) -- (260.80, 97.26) -- (272.16, 83.15) -- (283.51, 74.49) -- (294.87, 70.17) -- (306.22, 68.05) -- (317.57, 67.13) -- (328.93, 66.74) -- (340.28, 66.56) -- (369.80, 66.48) -- (414.09, 66.48); \end{scope} \begin{scope} \path[clip] ( 0.00, 0.00) rectangle (361.35,252.94); \definecolor{drawColor}{RGB}{0,0,0} \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20, 61.20) -- (336.15, 61.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 81.97, 61.20) -- ( 81.97, 55.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (137.32, 61.20) -- (137.32, 55.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (192.68, 61.20) -- (192.68, 55.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (248.03, 61.20) -- (248.03, 55.20); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (303.38, 61.20) -- (303.38, 55.20); \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 81.97, 39.60) {-0.5}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (137.32, 39.60) {-0.25}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (192.68, 39.60) {0}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (248.03, 39.60) {0.25}; \node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (303.38, 39.60) {0.5}; \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20, 61.20) -- ( 49.20,203.75); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20, 66.48) -- ( 43.20, 66.48); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20, 90.48) -- ( 43.20, 90.48); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20,114.47) -- ( 43.20,114.47); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20,138.47) -- ( 43.20,138.47); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20,162.47) -- ( 43.20,162.47); \path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 49.20,186.47) -- ( 43.20,186.47); \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20, 63.04) {0}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20, 87.03) {0.2}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20,111.03) {0.4}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20,135.03) {0.6}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20,159.03) {0.8}; \node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 37.20,183.02) {1}; \end{scope} \begin{scope} \path[clip] ( 49.20, 61.20) rectangle (336.15,203.75); \definecolor{drawColor}{RGB}{169,169,169} \path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] ( 49.20,180.47) -- (336.15,180.47); \path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] (200.00, 61.20) -- (200.00,203.75); \definecolor{drawColor}{gray}{0.40} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,169.67) -- (283.40,169.67); \definecolor{drawColor}{RGB}{66,134,244} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,158.87) -- (283.40,158.87); \definecolor{drawColor}{RGB}{102,166,30} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,148.07) -- (283.40,148.07); \definecolor{drawColor}{RGB}{166,118,29} \path[draw=drawColor,line width= 0.8pt,line join=round,line cap=round] (267.20,137.27) -- (283.40,137.27); \definecolor{drawColor}{RGB}{0,0,0} \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,166.57) {TC-DM}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,155.77) {Naive}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,144.97) {US}; \node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.90] at (291.50,134.17) {RBC}; \end{scope} \end{tikzpicture} } \end{minipage} \caption{Simulated coverage rates of various values of parameter values and for different types of confidence sets. Based on the Setting 1 and a continuous running variable as described in the main text. Bias aware approach with true constants (TC (ref); as reference function in all graphs), two times true constants (TC$\times2$), 0.5 times true constants (TC$\times.5$), and with rule of thumb smoothness bounds (ROT1) and (ROT2); robust bias correction (RBC); naive approach that ignores bias (Naive); and undersmoothing (US).} \end{figure}

}

To show that our bias-aware AR CSs not only have good coverage properties but also yield comparatively powerful inference, we simulate the rates at which the various CSs we consider cover parameter values other than the true one. We report the results for one exemplary setting (low curvature of outcome and treatment CEF, continuously distributed running variable) in Figure (ref).\footnote{We focus on this setting because the coverage of the true parameter is reasonably close to the nominal level for all procedures, and thus a comparison of coverage rates at “non-true" parameter values is meaningful across CSs. Analogous plots for other DGPs are available from the authors.} To avoid having all 16 coverage curves in one plot, we split the results into four panels: the five bias-aware AR CSs in (a), the three other AR CSs in (b), the five bias-aware DM CIs in (c), and the three other DM CIs in (d). Panels (b)--(d) also show the curve for our bias-aware AR CS with the true constants to have a common point of reference.

Panel (a) then shows that the coverage rate of bias-aware AR CSs drops very quickly to zero away from the true parameter. Panels (b)--(d) show that the coverage of bias-aware AR CSs away from the true parameter is also below that of most competing procedures over almost all the parameter space, with the exception of those which exhibit a meaningful distortion at the true parameter value and are therefore not suitable for a direct comparison. This confirms that the accurate coverage of our CSs in settings with discrete running variables and weak identification does not come at the expense of statistical power in a canonical setup, for which most competing CSs are specifically constructed.