EconBase
← Back to paper

Short and Simple Confidence Intervals when the Directions of Some Effects are Known

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

95,084 characters · 12 sections · 49 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Short and Simple Confidence Intervals when the Directions of Some Effects are Known

abstractWe provide adaptive confidence intervals on a parameter of interest in the presence of nuisance parameters when some of the nuisance parameters have known signs. The confidence intervals are adaptive in the sense that they tend to be short at and near the points where the nuisance parameters are equal to zero. We focus our results primarily on the practical problem of inference on a coefficient of interest in the linear regression model when it is unclear whether or not it is necessary to include a subset of control variables whose partial effects on the dependent variable have known directions (signs). Our confidence intervals are trivial to compute and can provide significant length reductions relative to standard confidence intervals in cases for which the control variables do not have large effects. At the same time, they entail minimal length increases at any parameter values. We prove that our confidence intervals are asymptotically valid uniformly over the parameter space and illustrate their length properties in an empirical application to a factorial design field experiment and a Monte Carlo study calibrated to the empirical application. Keywords: confidence intervals, adaptive inference, uniform inference, sign restrictions, boundary problems

\thispagestyle{empty} \setcounter{page}{0}

Introduction

Consider the common empirical setting for which a researcher is interested in estimating the causal effect of one variable on another via a linear regression in the presence of one or more observed potential control variables. The researcher believes that a regression including this full set of controls should not suffer from omitted variables bias but is uncertain whether it is necessary to include them all to overcome this bias. To obtain more informative inference, the researcher would prefer not to include controls unnecessarily and knows that if some of these controls indeed influence the outcome variable, it must be in a known positive or negative direction. Indeed, typical heuristic explanations for the potential inclusion of a control variable to mitigate omitted variables bias involve a known “direction” for the effect of the omitted variable on the outcome of interest. In this paper, we develop confidence intervals (CIs) with desirable properties for these types of settings.

More specifically, we develop CIs for a parameter of interest in the presence of nuisance parameters with a known sign. Our CIs are designed to have uniformly correct (asymptotic) coverage and desirable length properties across the entire parameter space while becoming particularly short when these nuisance parameters are small or zero. In the regression context, this latter property is motivated by common practical situations for which the researcher believes the regression coefficients on a subset of control variables with known partial effect directions are likely to be small or zero. In general, our CIs can be used for inference on a parameter in any well-behaved finite-dimensional model with a large-sample normally distributed estimator when some nuisance parameters are restricted above or below by zero (possibly after a location shift). This includes regression models estimated by ordinary, generalized and two stage least squares as well as models with bounded parameter spaces such as (G)ARCH Bollerslev:86 and random coefficient models BLP:95,Andrews:99. Even though the standard “constrained” estimator is not normally distributed in large samples when the true parameter vector is at (or close to) the boundary of the parameter space Andrews:99, there often exists a “quasi-unconstrained” estimator that is Ket18. While noting this generality, we mainly focus on regression models for ease of exposition.

To construct our CIs, we use the fact that knowledge of the signs of control variable coefficients, in addition to a standard consistent estimator of the covariance matrix of the underlying coefficient estimates, can be used to determine the sign of the corresponding omitted variables biases incurred by omitting the corresponding control variables. In turn, standard one-sided CIs for the coefficient of interest based on regressions that omit some of these control variables maintain correct coverage. We show that a particular form of these latter CIs is expected excess length-optimal (among affine CIs\textemdash see Proposition (ref) for details) when the corresponding control coefficients are equal to zero. It also has low expected excess length when the control coefficients are close to zero but its expected excess length grows without bound as the control coefficients grow larger. On the other hand, we show that standard one-sided CIs based upon the regression including all controls have the minimal maximum expected excess length (among affine CIs) over the parameter space that imposes the sign of the control coefficients.\footnote{In fact, we show that these two types of CIs are optimal at each quantile of the excess length distribution greater than one minus their nominal coverage probabilities. See Proposition (ref) below for details.} They also have correct coverage and expected excess length that does not depend upon the true values of the control coefficients. We propose adaptive one-sided CIs that utilize the strengths of both of these types of CIs by intersecting them. We make use of the same logic for constructing two-sided CIs essentially by intersecting our lower- and upper- one-sided CIs.

In particular, we propose a computationally trivial method to find the subset of controls that is able to produce the largest expected (excess) length reductions when using this intersection principle. In addition, the restricted parameter space implies that the coverage of these intersected CIs is lowest at its boundary. This feature allows us to provide the user a simple means to compute the smallest CI endpoints that yield correct coverage uniformly across the parameter space via response surface regression output, rather than using a conservative Bonferroni correction. Using our reported response surface regression coefficients, the user can immediately compute these CI endpoints as a function of one or two empirical correlation parameters, depending upon whether they are forming a one- or two-sided CI. A Stata package available in the SSC archive automatically computes the CIs we propose.\footnote{The package name is “ssci”. Corresponding Matlab code is available on the authors' webpage.}

We show that our proposed CIs are uniformly asymptotically valid and characterize their length properties. The latter depend upon the correlation structure of the underlying data and the true values of the unknown control coefficients. For extreme values of correlation between the estimators of the coefficient of interest and sign-restricted controls, the expected (excess) length of our CIs can be close to 100% smaller than that of standard CIs based upon the regression including all controls. For correlation values more likely to be encountered in practice, these expected (excess) length reductions can still exceed 30% for commonly used confidence levels. On the other hand, for a confidence level of 95%, for example, our proposed two-sided CIs cannot be more than 2.28% longer than the corresponding standard CI for any realization of the data and the expected excess length of our one-sided CIs cannot be more than 3% longer than that of the corresponding standard CI.

A leading example of where our proposed CIs should prove useful is in the context of factorial (or “cross-cutting”) designs in field experiments. Take, for example, the $2 \times 2$ factorial design where two treatments are administered independently such that there are three treatment arms, the two “main” treatments (separately) and the combination of the two, and a control arm. The corresponding treatment effects can be consistently estimated by OLS using the “long regression”, i.e., the regression of the outcome variable on a constant and three dummy variables, one for each treatment arm.\footnote{Recently, Muralidharan19 have highlighted the importance of using the long regression, as opposed to the “short” regression, i.e., the regression of the outcome variable on a constant and a dummy for the treatment of interest, to avoid omitted variable biases and accompanying size distortions.} In many cases, researchers have prior knowledge about the signs of the main treatment effects. For example, ethics boards for research grants are unlikely to fund experiments unless they are very likely to entail non-negative average treatment effects. Moreover, experimenters often conduct pilot studies prior to conducting full scale experiments in part to confirm their prior beliefs about the direction of average treatment effects.\footnote{For example, possible negative effects of a treatment may be closely monitored during the pilot phase and, if realized, even lead to early termination of the experiment. Note also that imposing the absence of any negative effects on individual participants (that could be associated with the treatment) is stronger than necessary, because our CIs only require sign restrictions on average (treatment) effects in this context.} Such prior knowledge can then be used to obtain sign restrictions on the corresponding regression coefficients. In this context, our proposed CIs will be “short” if the estimated effects of the main treatments are small and/or have the “wrong” sign, which is likely to occur if the unknown population effects are small or zero. To illustrate the potential usefulness of our CIs in the context of factorial designs, we revisit blattman2017 who study the effect of “therapy” and “cash” on violent and criminal behavior in Liberia using a $2 \times 2$ factorial design. Indeed, for some of the treatment effects under study, we find our proposed CIs to be up to 36% shorter than the corresponding standard CIs.

Relationship with the Literature

Several results in the statistics and econometrics literatures provide bounds on the ability for CIs to simultaneously maintain uniformly correct coverage over a class of data-generating processes (DGPs) while adapting to a given subclass. Here we develop CIs with this very goal in mind: our CIs maintain correct coverage for the parameter of interest uniformly across the parameter space for the nuisance parameters while becoming shorter when these nuisance parameters are equal to zero. Although most of this literature is devoted to nonparametric methods (e.g., Low97; CL04), the recent work of AK18 has produced similar implications for parametric models like those in the asymptotic versions of the problems we study. Indeed, AKK20 provide bounds on the ability to shorten CIs while maintaining correct coverage for regression coefficients at points for which potential control coefficients are zero. However, all of the aforementioned results rely upon an assumption of symmetry about zero for the underlying parameter space (among others). Because we are interested in problems with sign-restricted nuisance parameters, the underlying parameter space is asymmetric and these results do not apply, allowing for us to achieve the goal of constructing CIs that become significantly shorter at empirically-relevant parameter values.

Depending on the application, it may, of course, be possible that a researcher has prior knowledge on the magnitude of the control variables' coefficients rather than their sign. In this case, the recent work by AKK20 can be employed LM21. Indeed, Muralidharan19 study and suggest (among others) the CI proposed by AKK20 as a means to improve over standard CIs in the context of factorial designs. In particular, they argue that researchers may, depending on the application, be willing to assume prior knowledge of the maximum (absolute) value of an “interaction effect”, e.g., the effect of providing two treatments jointly minus the sum of the two main treatment effects. Here, we provide complementary results to be applied in settings for which it is natural for researchers to know the direction, rather than the magnitude, of control variables' coefficients.

This is certainly not the first paper to produce CIs that adapt to subclasses of DGPs while retaining uniform control of coverage probability. Several authors have provided such adaptive CIs for various smoothness classes and shape constraints in the nonparametric literature. See, e.g., CL04, CLX13, Arm15, KK20 and KK20b. Given our focus on finite-dimensional models, we are not concerned with the rate of convergence adaptation in this literature but rather finite-sample length adaptation for CIs. Nevertheless, our CIs share some similarities with some of the CIs in this literature. Like the ones we propose, the CIs of CL04, KK20 and KK20b are obtained by intersecting CIs that are optimal under different subclasses of DGPs. Within this literature, KK20 and KK20b are probably the closest studies to ours as they focus on nonparametric regression models with coordinate-wise monotone regression functions. In addition, both KK20 and KK20b provide a means of shortening adaptive nonparametric CIs relative to simple Bonferroni corrections in a similar spirit to our CI endpoints computed from response surface regression output. However, in contrast to the existing literature on adaptive CIs for nonparametric models, we prove our CIs are uniformly asymptotically valid without assuming Gaussian disturbances or fixed regressors.

Finally, our work is related to the literature on uniform inference when nuisance parameters may be at or near a boundary, e.g., AG09, McC17 and Ket18. While CIs with uniform asymptotic validity could in principle be computed by inverting the tests in this literature, this is often computationally prohibitive, especially when the nuisance parameter exceeds one or two dimensions. Similarly, inverting weighted average power maximizing tests such as those of MM13 or EMW15 is computationally intractable for most realistic applications. In contrast, our CIs are direct and trivial to compute since they do not rely on test inversion. Moreover, our CIs are designed to have length properties that are desirable from a practical perspective without requiring the user to specify weights or tuning parameters to optimize over.

Outline of Paper

The remainder of this paper is organized as follows. Section (ref) imparts the basic intuition of our CI constructions in a stylized asymptotic version of the inference problem we consider before providing computationally trivial algorithms for constructing the one- and two-sided CIs we propose in the general asymptotic setting. Section (ref) then shows how our CIs are constructed in practical finite-sample applications and provides theoretical results establishing their uniform asymptotic validity across a wide variety of applications. In Section (ref), we illustrate the usefulness of our CIs in an empirical application of inference on treatment effects in a factorial design field experiment while Section (ref) examines their finite-sample properties in a simulation study calibrated to the empirical application. Appendix (ref) provides the mathematical proofs of our theoretical results and Appendix (ref) specifies a parameter space for the standard linear regression model that satisfies the requirements for some of our theoretical results. Appendix (ref) contains additional tables referenced in the text while the online supplemental appendix provides details on the numerical computations underlying some of the results in this paper.

Throughout this paper, we use the following notational conventions. For any two column vectors $a$ and $b$, we sometimes write $(a, b)$ instead of $(a', b')'$ and let $a \geq b$ denote the element-by-element inequality. Let $\mathbb{R}_{+} = [0,\infty)$, $\mathbb{R}_{+,\infty}=\mathbb{R}_{+}\cup\{\infty\}$, $\mathbb{R}_{\infty}=\mathbb{R}\cup\{\infty\}\cup\{-\infty\}$ and $z_{\xi}$ denote the $\xi^{th}$ quantile of the standard normal distribution. For a square matrix $A$, $\operatorname{Diag}(A)$ denotes the diagonal matrix with the same diagonal entries as $A$ and $\lambda_{\min}(A)$ and $\lambda_{\max}(A)$ denote its smallest and largest eigenvalues.

commentFor scalar $\beta$ and an unrestricted parameter space for $\delta$, EMW12 note that the usual $t$-test using $Y_{\beta}$ that ignores the $Y_{\delta}$ component is essentially optimal from a testing point of view in the sense that the one-sided $t$-test is uniformly most powerful and the two-sided $t$-test is uniformly most powerful unbiased. Similarly, AK15 provide bounds on adaptivity that apply to confidence intervals for scalar $\beta$ in model (ref) when the parameter space for $(\beta,\delta^{\prime})^{\prime}$ is convex and symmetric about the zero vector. These results essentially imply that one can hardly shorten the (excess) length of a confidence interval for $\beta$ over the minimax optimal confidence interval under these conditions on the parameter space if one wishes to control coverage probability uniformly over the parameter space. For inference on scalar $\beta$, the minimax confidence intervals are formed by inverting the usual $t$-tests. Together, these complementary results suggest that there is very little scope for improving inference (e.g., shortening confidence intervals or increasing the power of tests) on a scalar $\beta$ in these problems if one is not willing to restrict the parameter space. However, at least in the context of the linear regression model in the presence of one or more potential controls, it would indeed seem natural for the practitioner to work with a restricted parameter space for $\delta$ that is not symmetric about the zero vector. Indeed, most heuristic explanations for the potential inclusion of a control variable to mitigate omitted variable bias involves a plausible “direction” for the effect of the omitted variable on the outcome of interest. This translates into sign restrictions on the elements of $\delta$. The goal of this research is to then develop inference methods for the Gaussian shift limit experiment with a constrained parameter space that are both (i) (approximately) optimal (in some sense) and (ii) easy to compute. In principle, the approximately weighted average power optimal tests of MM13 and EMW15 could be applied to this testing problem. However, these tests are (i) very difficult, if not impossible, to compute when the dimension of $\delta$ exceeds one or two and (ii) require the user to choose a weight function over the parameter space of the alternative hypothesis. Some would argue that this latter choice is arbitrary and/or difficult to interpret. Moreover, if one wished to form a confidence interval by inverting such tests, it is not clear that this would be computationally feasible even for one-dimensional $\delta$. Two more approaches to the testing problem under study are given by the CLR and GLR approaches of Ket18. The CLR test is similar and easy to compute, even for higher dimensional $\delta$. One may invert this test to form confidence intervals. Ket18 shows that the GLR test has desirable power properties when the dimension of $\delta$ is equal to one. However, for higher-dimensional $\delta$, the GLR test becomes difficult to compute and may have worse power properties since it relies upon the use of a least favorable CV.

Normal Means Large Sample Problem

LeCam's Limits of Experiments Theory provides that inference on the parameter of a well-behaved model is equivalent to inference on the mean of a Gaussian random vector with known variance matrix in large samples. This powerful result incorporates regression models, instrumental variables models, maximum likelihood models and models estimated by the generalized method of moments.\footnote{This result holds under assumptions ensuring the model is well-behaved, amounting to the existence of an asymptotically normally distributed estimator for the (finite-dimensional) parameter vector in our context. For example, in the context of instrumental variables models or models estimated by the generalized method of moments, this does not allow for weak instruments or other forms of weak identification. As alluded to in the Introduction, one can use the results of Ket18 to obtain this limit experiment even for models that may not be defined outside the parameter space, such as the random coefficients logit BLP:95 and (G)ARCH models for which variance parameters must be non-negative. Indeed, such models provide other natural applications for the CIs we introduce in this paper.} See Chapter 9 of vdV98 and Chapter 13 of LR05 for textbook treatments of this theory. Since the variance matrix is known in this setting, each element of the Gaussian random vector can be scale-normalized so that the large sample inference problem reduces to inference on the mean vector $h$ from a single observation $Y\overset{d}\sim \mathcal{N}(h,\Omega)$, where $\Omega$ is a known correlation matrix.

It is often the case in econometric applications that the researcher is interested in constructing a CI for a scalar parameter of interest in the presence of nuisance parameters. In addition, the researcher often has knowledge about the sign of the nuisance parameters. For example, when performing inference on a single coefficient in the linear regression model when “control” variables may be included in the regression to mitigate potential omitted variable bias, the researcher often knows the direction of the partial effects of some of the controls from economic theory or logical reasoning. In the large sample problem, this corresponds to conducting inference on a scalar $\beta$ from a single observation

equation[equation omitted — 195 chars of source]

where $\Omega$ is a known positive-definite correlation matrix and $\delta$ is a finite-dimensional nuisance parameter whose elements are known to be greater than or equal to zero.\footnote{The restriction $\delta\geq 0$ is without loss of generality because parameters without sign restrictions may be dropped from the analysis in the limiting problem and limiting Gaussian random variables corresponding to parameters restricted to be greater/less than or equal to a known number may be linearly transformed to conform to (ref).}

In many contexts, it is natural for the researcher to desire a CI with the following properties: (i) correct coverage $1-\alpha$ (coverage of at least $1-\alpha$) across the entire $\delta\geq 0$ parameter space, (ii) good length properties across the entire $\delta\geq 0$ parameter space and (iii) shortness when $\delta$ is equal or close to zero. For example, if it is not obvious whether a regressor should enter as a control variable or not, it is sensible to desire an especially short CI when the unknown population regression coefficient is equal to or near zero (reflecting the researcher's uncertainty about whether it is an important variable) while maintaining correct coverage and decent length no matter the coefficient's magnitude. In this section, we provide CI constructions for the large sample problem with this very goal in mind. We begin by describing the intuition for the CIs in the simplest version of the problem and subsequently provide general formulations for both one- and two-sided CIs.

Basic Intuition

To communicate the basic intuition for our CIs, we specialize the large sample problem (ref) to the case for which $\delta$ is one-dimensional and the correlation between $Y_\beta$ and $Y_{\delta}$ is positive:

equation*[equation* omitted — 235 chars of source]

where $\rho> 0$, $\beta$ is unrestricted, and $\delta \geq 0$. Consider the formation of an upper one-sided CI for $\beta$ with the goal of satisfying properties (i)--(iii) above. To illustrate the tension between properties (ii) and (iii), note that the standard CI that ignores the information in $Y_{\delta}$, i.e., $$CI_u(Y_{\beta})=[Y_{\beta}-z_{1-\alpha},\infty),$$ satisfies (ii) but not (iii) since its expected excess length is always simply equal to $z_{1-\alpha}$.\footnote{Expected excess length of an upper one-sided CI for $\beta$ is defined as $E[\beta - \text{lb}]$, where lb denotes the lower bound of the CI.} On a more technical level, we show in Proposition (ref)(i) below that this CI achieves the minimal maximum excess length quantile for all quantiles larger than $\alpha$ across the $\delta\geq 0$ parameter space. That is, the standard CI is minimax for the problem we are interested in. On the other hand, Proposition (ref)(ii) below shows that the CI that is excess length-optimal for all excess length quantiles larger than $\alpha$ when $\delta$ is known to equal zero is equal to \[\widetilde{CI}_u(Y_{\beta},\rho Y_{\delta})=\left[Y_\beta-\rho Y_\delta-\sqrt{1-\rho^2}z_{1-\alpha},\infty\right).\] This CI satisfies (iii) but not (ii) since its expected excess length is equal to $\rho \delta +\sqrt{1-\rho^2}z_{1-\alpha}$, which diverges as $\delta \to \infty$.\footnote{Both CIs satisfy (i) in this context since we have assumed $\rho> 0$, see equation (ref).}

In order to attain property (iii) but not at the expense of property (ii), we propose CIs with length performance designed to adapt to the data. Consider intersecting the two CIs $CI_u(Y_{\beta})$ and $\widetilde{CI}_u(Y_{\beta},\rho Y_{\delta})$ to simultaneously retain property (ii) of the former and property (iii) of the latter:

align*[align* omitted — 350 chars of source]

for some $\gamma\in(0,\alpha)$. Note that $\widehat{CI}_u\left(Y_{\beta},\rho Y_{\delta};z_{1-\alpha+\gamma},\sqrt{1-\rho^2}z_{1-\gamma}\right)$ maintains correct coverage probability over the parameter space:

align[align omitted — 600 chars of source]

for all $(\beta,\delta)\in\mathbb{R}\times \mathbb{R}_+$, where the first inequality follows from the Bonferroni inequality and the second inequality uses the fact that

align*[align* omitted — 312 chars of source]

with $$\tilde{Z}_\rho=Y_{\beta}-\rho Y_{\delta}-(\beta-\rho\delta)\overset{d}\sim \mathcal{N}(0,{1-\rho^2}),$$ where the inequality uses the fact that $\rho\delta \geq0$.

Since $\widehat{CI}_u\left(Y_{\beta},\rho Y_{\delta};z_{1-\alpha+\gamma},\sqrt{1-\rho^2}z_{1-\gamma}\right)$ makes use of a multiplicity correction based upon the Bonferroni bound, for similar reasons used to motivate the adjusted Bonferroni critical values of McC17, it is possible to decrease the excess length of $\widehat{CI}_u\left(Y_{\beta},\rho Y_{\delta};z_{1-\alpha+\gamma},\sqrt{1-\rho^2}z_{1-\gamma}\right)$ while retaining uniform control of coverage probability. In particular, fix $\gamma\in(0,\alpha)$ and find the constant $c^*\in [0,\sqrt{1-\rho^2}z_{1-\gamma}]$ that solves

equation[equation omitted — 106 chars of source]

in $c$, where

equation*[equation* omitted — 213 chars of source]

The CI $\widehat{CI}_u(Y_{\beta},\rho Y_{\delta};z_{1-\alpha+\gamma},c^*)$ is contained in $\widehat{CI}_u\left(Y_{\beta},\rho Y_{\delta};z_{1-\alpha+\gamma},\sqrt{1-\rho^2}z_{1-\gamma}\right)$ and maintains correct coverage probability over the parameter space:

align*[align* omitted — 455 chars of source]

for all $(\beta,\delta)\in\mathbb{R}\times\mathbb{R}_+$, where the second equality follows from the fact that $(Y_{\beta},Y_{\delta})\overset{d}\sim(\beta,\delta)+(Z_1,Z_2)$ and the inequality again uses the fact that $\rho\delta\geq 0$. The problem (ref) is computationally straightforward and can, for example, be solved by means of Monte Carlo simulations. A heuristic approach to choosing the tuning parameter $\gamma$ makes use of similar reasoning to that used to compute the adjusted Bonferroni critical values in McC17: a “small” $\gamma$ such as $\gamma=\alpha/10$ yields only slightly higher expected excess length when $\delta$ is “large” but significantly lower expected excess length when $\delta$ is “small”.

Finally, it is interesting to note that $\widehat{CI}_u(Y_{\beta},\rho Y_{\delta};z_{1-\alpha+\gamma},c^*)$ can be viewed as a CI that results from a model selection procedure designed for inference. In the context of the regression model example, we can view the model selection procedure as follows:

enumerate• If $Y_{\delta}>(z_{1-\alpha+\gamma}-c^*)/\rho$, construct the CI for $\beta$ from the “full” regression using the critical value $z_{1-\alpha+\gamma}$. • If $Y_{\delta}\leq(z_{1-\alpha+\gamma}-c^*)/\rho$, construct the CI for $\beta$ from the “short” regression using the critical value $c^*$.

The model selection pretest rule $Y_{\delta}>(z_{1-\alpha+\gamma}-c^*)/\rho$ is analogous to using a $t$-test as a pretest but with a nonstandard critical value that incorporates both the two-step nature of the inference procedure as well as the dependence between $Y_{\beta}$ and $Y_{\delta}$. Note that as $\rho\rightarrow 1$, this nonstandard pretest approaches a standard $t$-test pretest. Unlike standard model selection procedures, this procedure is designed for inference in the sense that (i) it uniformly controls coverage probability by directly incorporating the model selection uncertainty in its construction and (ii) it is designed to yield low excess length rather than a different notion of risk (such as mean-squared error).\footnote{Though some recent post-selection inference procedures (e.g., BCH14,McC17) uniformly control coverage probability/size, the selection procedures used in their construction are not designed to yield CIs with desirable length properties.}

One-Sided Confidence Intervals

In this section, we focus on forming analogous adaptive one-sided CIs but now allowing $\delta\geq 0$ to be multidimensional so that the large sample problem corresponds to (ref), where

equation[equation omitted — 169 chars of source]

Without loss of generality, we focus on upper one-sided CIs for $\beta$ since lower one-sided CIs may be attained analogously upon multiplying $Y_\beta$ by negative one. The optimal $(1-\alpha)$-level upper one-sided CI for $\beta$ when $\delta=0$ is equal to \[ \left[Y_\beta - \Omega_{\beta\delta} \Omega_{\delta\delta}^{-1} Y_\delta - z_{1-\alpha}\sqrt{1-\Omega_{\beta\delta} \Omega_{\delta\delta}^{-1}\Omega_{\delta\beta} },\infty \right). \] The CI that intersects this CI with the standard CI for $\beta$ that ignores the information in $Y_\delta$ will not maintain coverage in general. More specifically, the argument in (ref) for showing correct coverage only generalizes when all of the elements of $\Omega_{\beta\delta} \Omega_{\delta\delta}^{-1}$ are non-negative. In the case that this condition does not hold, we can still find adaptive CIs with potential length improvements by “dropping” elements of $Y_{\delta}$ from consideration. The following algorithm is designed to do just that while maintaining particularly low excess length when $\delta$ is equal or close to zero.

For $\gamma\in(0,\alpha)$, consider the function $c:[0,1)\rightarrow [0,z_{1-\gamma}]$ such that

equation[equation omitted — 102 chars of source]

where

equation*[equation* omitted — 182 chars of source]

The following result ensures that $c:[0,1)\rightarrow [0,z_{1-\gamma}]$ is well-defined and continuous.

propositionFor $\alpha\in(0,1/2)$, $c:[0,1)\rightarrow[0,z_{1-\gamma}]$ as defined in (ref) exists and is continuous.

Note that $c(0)=z_{1-\alpha}$. Let $Y_{\delta}^{(s)}$ denote an arbitrary subvector of $Y_{\delta}$, including the empty one, with

equation*[equation* omitted — 323 chars of source]

where by convention $\delta^{(s)}$, $\Omega_{\beta\delta^{(s)}}$ and $\Omega_{\delta^{(s)}\delta^{(s)}}$ (as well as $\Omega_{\beta\delta^{(s)}} \Omega_{\delta^{(s)}\delta^{(s)}}^{-1}$ and $\Omega_{\beta\delta^{(s)}} \Omega_{\delta^{(s)}\delta^{(s)}}^{-1}\Omega_{\delta^{(s)}\beta}$) are set equal to zero when $Y_{\delta}^{(s)} = \emptyset$.

\begin{one sided} Amongst all subvectors of $Y_{\delta}$ (including the empty one) such that the elements of $\Omega_{\beta\delta^{(s)}} \Omega_{\delta^{(s)}\delta^{(s)}}^{-1}$ are non-negative, find the subvector $Y_{\delta}^{(s^*)}$ such that the expected excess length of \[ \widehat{CI}_u(Y_{\beta},\Omega_{\beta\delta^{(s)}}\Omega_{\delta^{(s)}\delta^{(s)}}^{-1}Y_{\delta}^{(s)};z_{1-\alpha+\gamma},c\left(\Omega_{\beta\delta^{(s)}} \Omega_{\delta^{(s)}\delta^{(s)}}^{-1}\Omega_{\delta^{(s)}\beta}\right)) \] at $\delta = 0$ is minimized at $s=s^*$. Then, construct

equation[equation omitted — 287 chars of source]

\end{one sided}

The goal of this algorithm is to generate short CIs when the user is agnostic about which elements of $\delta$ are more likely to be (close to) zero. Figure (ref) shows the expected excess length of $\widehat{CI}_u(Z_1,\tilde Z_2,z_{1-\alpha+\gamma},c(\omega))$ as a function of $\omega$, for $\alpha \in \{0.1, 0.05, 0.01\}$ and our recommended value of $\gamma = \alpha/10$.\footnote{Expected excess length is obtained numerically on the following grid of values: $\omega\in\{0,0.001,0.002,\dots,0.999\}$. See the online supplemental appendix for details.} This expected excess length is strictly decreasing in $\omega$ (at least for the considered choices of $\alpha$ and $\gamma$). Since it does not depend upon $\beta$, this implies that the expected excess length of $\widehat{CI}_u(Y_{\beta},\Omega_{\beta\delta^{(s)}}\Omega_{\delta^{(s)}\delta^{(s)}}^{-1}Y_{\delta}^{(s)};z_{1-\alpha+\gamma},c\left(\Omega_{\beta\delta^{(s)}} \Omega_{\delta^{(s)}\delta^{(s)}}^{-1}\Omega_{\delta^{(s)}\beta}\right))$ evaluated at $\delta=0$ is smallest for the subvector $Y_\delta^{(s)}$ that maximizes $\Omega_{\beta\delta^{(s)}} \Omega_{\delta^{(s)}\delta^{(s)}}^{-1}\Omega_{\delta^{(s)}\beta}$, leading us to the following simplified algorithm.

figure[figure omitted — 336 chars of source]

\begin{one sided star} Amongst all subvectors of $Y_{\delta}$ such that the elements of $\Omega_{\beta\delta^{(s)}} \Omega_{\delta^{(s)}\delta^{(s)}}^{-1}$ are non-negative, find the subvector $Y_{\delta}^{(s^*)}$ such that $\Omega_{\beta\delta^{(s)}} \Omega_{\delta^{(s)}\delta^{(s)}}^{-1}\Omega_{\delta^{(s)}\beta}$ is maximized at $s=s^*$. Then, construct

equation[equation omitted — 338 chars of source]

\end{one sided star}

It is worth noting that (i) $c(\omega)$ for $\omega\in(0,1)$ is very simple to compute via Monte Carlo simulation, while $c(0) = z_{1-\alpha}$, and (ii) Algorithm One-Sided* only requires one to evaluate the function $c(\cdot)$ at the single point $\Omega_{\beta\delta^{(s^*)}} \Omega_{\delta^{(s^*)}\delta^{(s^*)}}^{-1}\Omega_{\delta^{(s^*)}\beta}$. Therefore, the algorithm carries very low computational cost. We also note that $c\left(\Omega_{\beta\delta^{(s^*)}} \Omega_{\delta^{(s^*)}\delta^{(s^*)}}^{-1}\Omega_{\delta^{(s^*)}\beta}\right)$ can always be replaced by $\sqrt{1-\Omega_{\beta\delta^{(s^*)}} \Omega_{\delta^{(s^*)}\delta^{(s^*)}}^{-1}\Omega_{\delta^{(s^*)}\beta}}z_{1-\gamma}$ in the algorithm to yield a CI with correct coverage but worse excess length (in analogy with the CI using the Bonferroni correction in the previous section).

In addition to the numerical justification for the selection of the subvector $Y_{\delta}^{(s^*)}$ to form $ \widehat{CI}_u^*(Y_{\beta},Y_\delta,\Omega)$, Algorithm One-Sided* is also justified on theoretical grounds as it entails the intersection of CIs that are optimal over two different subclasses of DGPs like those of e.g., CL04, KK20 and KK20b. More specifically, the following proposition applies a general result of AK18 to the current inference setting to formalize what we mean by “optimal” here.

propositionFor inference on $\beta$ in (ref), the following statements hold for $\alpha\in (0,1)$: (i) among all upper one-sided CIs with coverage of at least $(1-\alpha)$ for all $\delta\geq 0$, the CI that minimizes all maximum excess length quantiles over the $\delta\geq 0$ parameter space at quantile levels greater than $\alpha$ is equal to \[[Y_\beta-z_{1-\alpha},\infty),\] (ii) among all upper one-sided CIs with coverage of at least $(1-\alpha)$ for all $\delta\geq 0$, the CI that minimizes all excess length quantiles at the point $\delta=0$ and quantile levels greater than $\alpha$ is equal to \[ \left[Y_\beta - \Omega_{\beta\delta^{(s^*)}} \Omega_{\delta^{(s^*)}\delta^{(s^*)}}^{-1} Y_{\delta^{(s^*)}} - z_{1-\alpha}\sqrt{1-\Omega_{\beta\delta^{(s^*)}} \Omega_{\delta^{(s^*)}\delta^{(s^*)}}^{-1}\Omega_{\delta^{(s^*)}\beta} },\infty \right). \]

In particular, this proposition implies that for $\alpha < 1/2$ the CIs in (i) and (ii) above minimize maximum median excess length across $\delta\geq 0$ and median excess length at $\delta=0$, respectively. Since median and expected excess length coincide for CIs that are affine in the data ($Y$), we also have that the CIs in (i) and (ii) minimize maximum expected excess length across $\delta\geq 0$ and expected excess length at $\delta=0$ among all affine CIs, respectively, if $\alpha < 1/2$. Algorithm One-Sided* computes a CI that intersects two optimal CIs of these forms while making a non-conservative multiplicity correction that improves upon the conservative Bonferroni adjustement.

As can be seen from Figure (ref), extreme values of $\Omega_{\beta\delta^{(s^*)}} \Omega_{\delta^{(s^*)}\delta^{(s^*)}}^{-1}\Omega_{\delta^{(s^*)}\beta}$ such as 0.999 can lead to expected excess lengths of our one-sided CI close to zero, entailing expected excess length reductions of nearly 100% relative to the expected excess length of the standard CI. At more empirically-relevant values of $\Omega_{\beta\delta^{(s^*)}} \Omega_{\delta^{(s^*)}\delta^{(s^*)}}^{-1}\Omega_{\delta^{(s^*)}\beta}$, such as say 0.7, Figure (ref) still implies expected excess length reductions of more than 30% for $\alpha = 0.05$. On the other hand, the expected excess length of our one-sided CI is bounded above by $z_{1-\alpha+\gamma}=1.695$ for $\alpha=0.05$, $\gamma=\alpha/10$ and any value of $\Omega_{\beta\delta^{(s^*)}} \Omega_{\delta^{(s^*)}\delta^{(s^*)}}^{-1}\Omega_{\delta^{(s^*)}\beta}$. This implies that the expected excess length of our recommended CI relative to the standard one-sided CI is bounded above by $z_{1-\alpha+\gamma}/z_{1-\alpha}=1.695/1.645\approx 1.03$ for $\alpha=0.05$. That is the expected excess length increase of our recommended CI relative to the standard CI is bounded above by roughly 3% for $\alpha=0.05$ and any value of $\Omega_{\beta\delta^{(s^*)}} \Omega_{\delta^{(s^*)}\delta^{(s^*)}}^{-1}\Omega_{\delta^{(s^*)}\beta}$.\footnote{For $\alpha$ equal to 0.01 and 0.1 (and $\gamma=\alpha/10$), the expected excess length of our one-sided CI is bounded above by 2.366 and 1.341 and the expected excess length of the standard one-sided CI is equal to 2.326 and 1.282, respectively. This implies that the expected excess length increase of our CI relative to the standard CI is bounded above by 1.7% and 4.6%, for $\alpha=0.01$ and $\alpha=0.1$ respectively.}

figure[figure omitted — 662 chars of source]

Figure (ref) plots the expected excess length of our one-sided CI ($\widehat{CI}_u(\cdot)$), the minimax standard one-sided CI (${CI}_u(\cdot)$) and the CI that is optimal when $\delta=0$ ($\widetilde{CI}_u(\cdot)$) as a function of $\delta$, for $\alpha = 0.05$ (and $\gamma = \alpha/10$) and several values of $\omega=\Omega_{\beta\delta^{(s^*)}} \Omega_{\delta^{(s^*)}\delta^{(s^*)}}^{-1}\Omega_{\delta^{(s^*)}\beta}$. Similarly to Figure (ref), Figure (ref) shows that the gains in expected excess length of our one-sided CI compared to the standard one-sided CI are more pronounced for larger values of $\omega$. Furthermore, Figure (ref) illustrates the adaptive nature of our one-sided CI: at the endpoints of the parameter space, $\delta = 0$ and $\delta = \infty$, its expected excess length approaches those of the optimal CI at $\delta=0$ and the minimax standard one-sided CI.

A different choice of $\gamma$ from $\alpha/10$ would entail different tradeoffs for our CI over the $\delta\geq 0$ parameter space. For example, a larger choice of $\gamma$ would yield lower expected excess length in a neighborhood of $\delta=0$ by bringing it “closer” to the optimal CI at $\delta=0$. Conversely, such a choice for $\gamma$ would yield higher expected excess length at large values of $\delta$. In fact, the user of our CI could choose $\gamma$ according to how much of an increase in expected excess length they are willing to tolerate relative to the standard CI at large values of $\delta$ since the ratio of the expected excess length of our CI relative to the standard CI is bounded above by $z_{1-\alpha+\gamma}/z_{1-\alpha}$. For example, the choice of $\gamma=\alpha/2$ would entail an expected excess length increase relative to the standard CI bounded above by about 19% for $\alpha=0.05$ since $z_{1-\alpha+\gamma}/z_{1-\alpha}=1.96/1.645\approx 1.19$. Analogous implications for the choice of $\gamma$ apply to our two-sided CI constructions below.

table[table omitted — 586 chars of source]

In order to make practical implementation computationally trivial for the user, requiring no Monte Carlo simulation, we approximate $c(\omega)$ via a polynomial response surface regression. Table (ref) provides the estimated coefficients for a 6$^\text{th}$ order polynomial approximation of $c(\omega)$ for $\alpha \in \{0.01,0.05,0.1\}$ and $\gamma = \alpha/10$. For each value of $\alpha$, the $R^2$ is greater than 0.999.\footnote{The regression is performed on the same grid that underlies Figure (ref), i.e., $\omega \in \{0,0.001,0.002,\dots,0.999\}$.} Nevertheless, using the response surface approximation of $c(\omega)$ in the construction of our one-sided CI can lead to small (asymptotic) coverage distortions: A grid search over $\{0,0.001,0.002,\dots,0.999\}$ reveals a minimum coverage probability of 94.80% when $\alpha = 0.05$.\footnote{The corresponding values for $\alpha = 0.01$ and $\alpha = 0.1$ are 98.94% and 89.68%, respectively.} However in finite-sample practical applications, the coverage distortions resulting from using the response surface approximation tend to be smaller than those induced by using the large-sample normal approximation to conduct inference. Evidence of this can be found in the simulation results of Section (ref) for which the finite-sample coverage distortions of our CIs that are implemented using the response surface approximation are very close to those of the standard CI that does not rely on such an approximation. Thus from a practical perspective, the small size distortions arising from the response surface approximation are insignificant.

Two-Sided Confidence Intervals

In this section, we focus on forming analogous adaptive two-sided CIs allowing $\delta\geq 0$ to be multidimensional in the large sample problem characterized by (ref) and (ref). These two-sided CIs use the same basic logic as the one-sided CIs of the previous section but work to shorten each side of the CI separately while maintaining correct coverage. Unlike our one-sided CIs that use a single critical value $c\left(\Omega_{\beta\delta^{(s^*)}} \Omega_{\delta^{(s^*)}\delta^{(s^*)}}^{-1}\Omega_{\delta^{(s^*)}\beta}\right)$, two-sided CIs are formed using two, one corresponding to the upper bound of the interval and the other corresponding to the lower bound. We address this additional degree of freedom by choosing these upper and lower critical values to minimize expected length at $\delta=0$, subject to the constraint that the resulting two-sided CI maintains correct coverage across the $\delta\geq 0$ parameter space. The following algorithms provide the details.

Let

equation[equation omitted — 298 chars of source]

and $\widetilde{\mathcal{C}}=\{(c_u,\tilde\omega)\in\mathbb{R}_{\infty}\times \bar{\mathcal{S}}:c_u\in [\underline{c_u}(\tilde\omega),\infty]\}$, where $\bar{\mathcal{S}} = {\mathcal{S}} \cup \{ (x,y,z)\in \mathbb{R}^3: x \in [0,1), y = z = 0\} \cup \{ (x,y,z)\in \mathbb{R}^3: y \in (0,1), x = z = 0\}$ with $\mathcal{S}=\{(x,y,z)\in \mathbb{R}^3:x,y\in (0,1),-z^2+2xyz + xy - x^2y - xy^2 > 0\}$,\footnote{In terms of arguments $(x,y,z)$, the definition of $\mathcal{S}$ is equivalent to the positive definiteness of the matrix $\left(

array[array omitted — 55 chars of source]

\right).$} where $c_u:\bar{\mathcal{S}}\rightarrow \mathbb{R}$ is implicitly defined by

equation[equation omitted — 165 chars of source]

For $\gamma\in(0,\alpha)$, consider the function $ \tilde c:\widetilde{\mathcal{C}}\rightarrow \mathbb{R}_\infty$ implicitly defined by

equation[equation omitted — 180 chars of source]

at points $(c_u,\tilde\omega)\in\widetilde{\mathcal{C}}$ for which $\omega_{12},\omega_{13}\neq 0$. The domain $\widetilde{\mathcal{C}}$ of $\tilde c(\cdot)$ is defined in terms of the lower bound $\underline{c_u}(\tilde\omega)$ on $c_u$ in (ref) so that for any given $\tilde\omega$, the solution to (ref) exists. More specifically, the lower bound $\underline{c_u}(\tilde\omega)$ rules out $c_u$ values that are too small to admit a solution to (ref). Next, for $(c_u,\tilde\omega)\in\widetilde{\mathcal{C}}$ with $\omega_{12}= 0$, define $\tilde{c}(c_u,\tilde\omega)=\lim_{\bar\omega_{12}\rightarrow 0}\tilde{c}(c_u,\bar\omega_{12},\omega_{13},\omega_{23})$ and for $(c_u,\tilde\omega)\in\widetilde{\mathcal{C}}$ with $\omega_{13}= 0$, define $\tilde{c}(c_u,\tilde\omega)=\lim_{\bar\omega_{13}\rightarrow 0}\tilde{c}(c_u,\omega_{12},\bar\omega_{13},\omega_{23})$.\footnote{The limits in these definitions exist by the continuity of $\tilde{c}(c_u,\tilde\omega)$ at all $(c_u,\tilde\omega)\in \widetilde{\mathcal{C}}$ with $\omega_{12},\omega_{13}\neq 0$. See Lemma (ref) in the Appendix. We define $\tilde{c}(c_u,\tilde\omega)$ at $\omega_{12}=0$ and $\omega_{13}=0$ in terms of limits because multiple values of $\tilde{c}(c_u,\tilde\omega)$ satisfy (ref) when $\omega_{12}=0$ and we wish to treat $\omega_{12}$ and $\omega_{13}$ symmetrically in light of Proposition (ref) below.} Finally, define the correspondence $ \tilde c_u:\bar{\mathcal{S}}\rightrightarrows \mathbb{R}$ as

equation[equation omitted — 267 chars of source]

Note that $ \tilde c_u(0)=z_{1-\alpha/2}$. The following proposition ensures that $ \tilde c_u:\bar{\mathcal{S}}\rightrightarrows \mathbb{R}$ is well-defined and possesses some desirable properties.\footnote{In our numerical work, we have found the solution to (ref) to be a singleton and $\tilde c_u$ to be a continuous function when $\omega_{12},\omega_{13}\neq 0$.}

propositionFor any $\tilde\omega\in\bar{\mathcal{S}}$, $\tilde c_{u}(\tilde\omega)\subset \mathbb{R}_{\infty}$ defined in (ref) is non-empty and compact and $\tilde c_{u}:\bar{\mathcal{S}}\rightrightarrows\mathbb{R}_\infty$ is upper hemicontinuous.

Now, let $Y_{\delta}^{(s_1)}$ and $Y_{\delta}^{(s_2)}$ denote two arbitrary (possibly empty) subvectors of $Y_{\delta}$ with

equation*[equation* omitted — 573 chars of source]

where by convention, $\delta^{(s_1)}$ ($\delta^{(s_2)}$), $\Omega_{\beta\delta^{(s_1)}}$ ($\Omega_{\beta\delta^{(s_2)}}$), $\Omega_{\delta^{(s_1)}\delta^{(s_1)}}$ ($\Omega_{\delta^{(s_2)}\delta^{(s_2)}}$), and $\Omega_{\delta^{(s_1)}\delta^{(s_2)}}$, as well as $\Omega_{\beta\delta^{(s_1)}} \Omega_{\delta^{(s_1)}\delta^{(s_1)}}^{-1}$ ($\Omega_{\beta\delta^{(s_2)}} \Omega_{\delta^{(s_2)}\delta^{(s_2)}}^{-1}$), $\Omega_{\beta\delta^{(s_1)}} \Omega_{\delta^{(s_1)}\delta^{(s_1)}}^{-1}\Omega_{\delta^{(s_1)}\beta}$ ($\Omega_{\beta\delta^{(s_2)}} \Omega_{\delta^{(s_2)}\delta^{(s_2)}}^{-1}\Omega_{\delta^{(s_2)}\beta}$), and $\Omega_{\beta \delta^{(s_1)}} \Omega_{ \delta^{(s_1)} \delta^{(s_1)}}^{-1} \Omega_{ \delta^{(s_1)} \delta^{(s_2)}} \Omega_{ \delta^{(s_2)} \delta^{(s_2)}}^{-1}\Omega_{ \delta^{(s_2)}\beta}$, are set equal to zero when $Y_{\delta}^{(s_1)} = \emptyset$ ($Y_{\delta}^{(s_2)} = \emptyset$). Define

align*[align* omitted — 706 chars of source]

where $c_u(\tilde\omega)\in \tilde c_u(\tilde\omega)$, $c_\ell(\tilde\omega)= \tilde c(c_u(\tilde\omega),\tilde\omega)$ and {\[\tilde\Omega^{(s_1,s_2)}=({\Omega_{\beta \delta^{(s_1)}} \Omega_{ \delta^{(s_1)} \delta^{(s_1)}}^{-1}\Omega_{ \delta^{(s_1)}\beta}},{\Omega_{\beta \delta^{(s_2)}} \Omega_{ \delta^{(s_2)} \delta^{(s_2)}}^{-1}\Omega_{ \delta^{(s_2)}\beta}},\Omega_{\beta \delta^{(s_1)}} \Omega_{ \delta^{(s_1)} \delta^{(s_1)}}^{-1} \Omega_{ \delta^{(s_1)} \delta^{(s_2)}} \Omega_{ \delta^{(s_2)} \delta^{(s_2)}}^{-1}\Omega_{ \delta^{(s_2)}\beta}).\] } For any given pair of subvectors $Y_\delta^{(s_1)}$ and $Y_\delta^{(s_2)}$ such that all elements of $\Omega_{\beta \delta^{(s_1)}} \Omega_{ \delta^{(s_1)} \delta^{(s_1)}}^{-1}$ are non-negative and all elements of $\Omega_{\beta \delta^{(s_2)}} \Omega_{ \delta^{(s_2)} \delta^{(s_2)}}^{-1}$ are non-positive, $c_\ell(\tilde\Omega^{(s_1,s_2)})$ and $c_u(\tilde\Omega^{(s_1,s_2)})$ minimize the expected length at $\delta=0$ of CIs of the form \[\widehat{CI}_t(Y_{\beta},\Omega_{\beta\delta^{(s_1)}} \Omega_{\delta^{(s_1)}\delta^{(s_1)}}^{-1} Y_\delta^{(s_1)},\Omega_{\beta\delta^{(s_2)}} \Omega_{\delta^{(s_2)}\delta^{(s_2)}}^{-1} Y_\delta^{(s_2)};z_{1-(\alpha-\gamma)/2},c_\ell,c_u)\] amongst all $(c_\ell,c_u)$ values that have valid coverage for all $\delta\geq 0$. To impart intuition, note that the two subvectors $Y_\delta^{(s_1)}$ and $Y_\delta^{(s_2)}$ play separate roles in our two-sided CI construction: the lower (upper) bound of the CI is a function of $Y_\delta^{(s_1)}$ ($Y_\delta^{(s_2)}$) so that small values of $Y_\delta^{(s_1)}$ ($Y_\delta^{(s_2)}$) serve to increase (decrease) the lower (upper) bound of $\widehat{CI}_t(\cdot)$.

\begin{two sided} Amongst all pairs of subvectors of $Y_{\delta}$ (including the empty ones) such that all elements of $\Omega_{\beta \delta^{(s_1)}} \Omega_{ \delta^{(s_1)} \delta^{(s_1)}}^{-1}$ are non-negative and all elements of $\Omega_{\beta \delta^{(s_2)}} \Omega_{ \delta^{(s_2)} \delta^{(s_2)}}^{-1}$ are non-positive, find the subvector pair $Y_{\delta}^{(s_1^*)}$ and $Y_{\delta}^{(s_2^*)}$ such that the expected length of \[ \widehat{CI}_t(Y_{\beta},\Omega_{\beta\delta^{(s_1)}} \Omega_{\delta^{(s_1)}\delta^{(s_1)}}^{-1} Y_\delta^{(s_1)},\Omega_{\beta\delta^{(s_2)}} \Omega_{\delta^{(s_2)}\delta^{(s_2)}}^{-1} Y_\delta^{(s_2)};z_{1-(\alpha-\gamma)/2},c_\ell(\tilde\Omega^{(s_1,s_2)}),c_u(\tilde\Omega^{(s_1,s_2)})) \] at $\delta = 0$ is minimized at $s_1=s_1^*$ and $s_2=s_2^*$. Then, construct

equation[equation omitted — 353 chars of source]

\end{two sided}

figure[figure omitted — 243 chars of source]

Figure (ref) shows the fitted surface of a 6$^\text{th}$ order polynomial regression of the expected length of $\widehat{CI}_t(Z_1,\tilde Z_2, \tilde Z_3;z_{1-(\alpha-\gamma)/2},c_\ell(\tilde\omega),c_u(\tilde \omega))$ on $\omega_{12}$ and $\omega_{13}$ alone, for $\alpha = 0.05$ and $\gamma = \alpha/10$. The values of $\tilde \omega$ on which this regression is based are given by $\bar{\Omega} = \bar{\mathcal{S}} \cap \mathcal{G}^2 \times -\mathcal{G} \cup \mathcal{G} \cup \{-0.99,-0.98,\dots,0.99\}$, where

equation[equation omitted — 122 chars of source]

and $-\mathcal{G} = \{g : -g \in \mathcal{G}\}$. The corresponding $R^2$ is greater than 0.999, implying that the expected length is nearly invariant to $\omega_{23}$. Similarly, the maximum difference between the largest and the smallest expected length over the set $\bar{\Omega}$ for any given $(\omega_{12}, \omega_{13})$ is equal to 0.0289. Note also that the fitted expected length in Figure (ref) is strictly decreasing in $\omega_{12}$ and $\omega_{13}$. Since the expected length does not depend upon $\beta$, this implies that the expected length of $\widehat{CI}_t(Y_{\beta},\Omega_{\beta\delta^{(s_1)}} \Omega_{\delta^{(s_1)}\delta^{(s_1)}}^{-1} Y_\delta^{(s_1)},\Omega_{\beta\delta^{(s_2)}} \Omega_{\delta^{(s_2)}\delta^{(s_2)}}^{-1} Y_\delta^{(s_2)};z_{1-(\alpha-\gamma)/2},c_\ell(\tilde\Omega^{(s_1^*,s_2^*)}),c_u(\tilde\Omega^{(s_1,s_2)}))$ evaluated at $\delta=0$ is approximately smallest for the subvectors $Y_\delta^{(s_1)}$ and $Y_\delta^{(s_2)}$ that maximize $\Omega_{\beta\delta^{(s_1)}} \Omega_{\delta^{(s_1)}\delta^{(s_1)}}^{-1}\Omega_{\delta^{(s)}\beta}$ and $\Omega_{\beta\delta^{(s_2)}} \Omega_{\delta^{(s)}\delta^{(s_2)}}^{-1}\Omega_{\delta^{(s_2)}\beta}$, motivating the following simplified algorithm.

\begin{two sided star} Amongst all pairs of subvectors of $Y_{\delta}$ (including the empty ones) such that all elements of $\Omega_{\beta \delta^{(s_1)}} \Omega_{ \delta^{(s_1)} \delta^{(s_1)}}^{-1}$ are non-negative and all elements of $\Omega_{\beta \delta^{(s_2)}} \Omega_{ \delta^{(s_2)} \delta^{(s_2)}}^{-1}$ are non-positive, find the subvector pair $Y_{\delta}^{(s_1^*)}$ and $Y_{\delta}^{(s_2^*)}$ such that ${\Omega_{\beta \delta^{(s_1)}} \Omega_{ \delta^{(s_1)} \delta^{(s_1)}}^{-1}\Omega_{ \delta^{(s_1)}\beta}}$ and ${\Omega_{\beta \delta^{(s_2)}} \Omega_{ \delta^{(s_2)} \delta^{(s_2)}}^{-1}\Omega_{ \delta^{(s_2)}\beta}}$ are maximized at $s_1=s_1^*$ and $s_2=s_2^*$. Then, construct

gather[gather omitted — 414 chars of source]

\end{two sided star}

Similarly to the one-sided CI case above, Figure (ref) shows that the expected length of our two-sided CI can be very small for extreme values of $\Omega_{\beta\delta^{(s_1^*)}} \Omega_{\delta^{(s_1^*)}\delta^{(s_1^*)}}^{-1}\Omega_{\delta^{(s_1^*)}\beta}$ and $\Omega_{\beta\delta^{(s_2^*)}} \Omega_{\delta^{(s_2^*)}\delta^{(s_2^*)}}^{-1}\Omega_{\delta^{(s_2^*)}\beta}$. At the same time, for any realization of the data, the realized length of our two-sided CI cannot exceed $2\times z_{1-(\alpha-\gamma)/2}=4.009$ for $\alpha=0.05$ and $\gamma=\alpha/10$. This implies that the length increase of our recommended CI cannot exceed 2.28% relative to the fixed length $2\times z_{1-\alpha/2}=3.92$ of the standard two-sided CI.\footnote{For $\alpha$ equal to 0.01 and 0.1, the length of our two-sided CI cannot exceed 5.224 and 3.391 and the length of the standard two-sided CI is equal to 5.152 and 3.290, respectively. This implies a maximum length increase of 1.41% and 3.07%, respectively.}

Table (ref) provides the estimated coefficients for a 6$^\text{th}$ order polynomial approximation of $c_u(\tilde \omega)$ for $\alpha = 0.05$ and $\gamma = \alpha/10$ in terms of $\omega_{12}$ and $\omega_{13}$ alone. Tables (ref) and (ref) in Appendix (ref) provide the corresponding coefficients for $\alpha = 0.01$ and $\alpha = 0.1$ (and $\gamma = \alpha/10$), respectively. For all three values of $\alpha$, the corresponding $R^2$ is greater than 0.999.\footnote{The regressions are performed on the same grid that underlies Figure (ref), i.e., $\omega \in \bar{\Omega}$.} This is remarkable, as it implies that $c_u(\cdot)$ and $c_\ell(\cdot)$ are nearly invariant to $\omega_{23}$. In fact, the maximum (asymptotic) size distortions that result from relying on the polynomial approximation of $c_u(\cdot)$ and $c_\ell(\cdot)$ in the construction of our two-sided CI are very similar to those that we found for relying on the polynomial approximation in the construction of our one-sided CI.\footnote{For $\alpha$ equal to 0.01, 0.05, and 0.1, a grid search over $\bar{\Omega}$ reveals minimum coverage probabilities of 98.96%, 94.78%, and 89.60%, respectively.} We therefore also expect minimal size distortions from employing the polynomial approximations of $c_u(\cdot)$ and $c_\ell(\cdot)$ in practice, which is again corroborated by the finite sample coverage probabilities found in Section (ref).

table[table omitted — 1,084 chars of source]

The following proposition also enables one to directly compute an approximation to $c_{\ell}(\tilde\omega)$ from Table (ref) (and Tables (ref) and (ref)) by simply reversing the roles of $\omega_{12}$ and $\omega_{13}$ in the computation of $c_u(\tilde\omega)$.

proposition$c_\ell(\omega_{13},\omega_{12},\omega_{23})=c_u(\omega_{12},\omega_{13},\omega_{23})$.

Finite Sample Problem of Restricted Nuisance Parameters

Consider inference on a scalar parameter of interest $b\in\mathbb{R}$ in a well-behaved model with a vector nuisance parameter $d\in \mathbb{R}_+^k$ for some $k\geq 1$ that is known to have all elements greater than or equal to zero. For a standard parameter estimator $(\hat b,\hat d^{\prime})^{\prime}$, as the number of observations $n$ in the sample grows, standard assumptions imply\footnote{For simplicity of notation, we suppress the dependence of certain finite-sample quantities on the sample size $n$ until Section 3.2.}

equation[equation omitted — 286 chars of source]

where $\Sigma$ is a consistently estimable covariance matrix. Note that this setting accommodates regression models, instrumental variables models, maximum likelihood models and models estimated by the generalized method of moments under standard assumptions when some nuisance parameters are known to be greater or less than a given bound via simple reparameterization of the nuisance parameters. For example, say that the researcher knows from economic theory that the nuisance parameter $\tilde d$ is less than or equal to $\tilde c$ for some known constant $\tilde c\in\mathbb{R}$. Then $d=-(\tilde d-\tilde c)$ is the simple reparameterization that fits this setting.

exampleOne of the most common examples that fits this setting is inference on a regression coefficient of interest $b$ in the standard linear regression model for observations $i=1,\ldots,n$ \[y_i=bz_i+x_i^{\prime}d+w_i^{\prime}c+\varepsilon_i,\] where $y_i$ is the dependent variable, $z_i$ is the scalar regressor of interest, $x_i\in\mathbb{R}^{\mathcal D_x}$ are control variables with known positive partial effects $d\geq 0$ on $y_i$, $w_i\in\mathbb{R}^{\mathcal D_w}$ are control variables with unrestricted partial effects $c$ and $\varepsilon_i$ is the error term. The ordinary least squares estimator $(\hat b,\hat d^{\prime})^{\prime}$ satisfies (ref) under standard assumptions on the linear regression model.

Implementation

For a consistent covariance matrix estimator $\widehat\Sigma$, (ref) suggests the following large-sample distributional approximation consistent with (ref) and (ref):

equation[equation omitted — 283 chars of source]

where $\beta=\sqrt{n}b/\sqrt{\Sigma_{bb}}$, $\delta=\operatorname{Diag}(\Sigma_{dd})^{-1/2}\sqrt{n}d$ and $\Omega=\operatorname{Diag}(\Sigma)^{-1/2}\Sigma\operatorname{Diag}(\Sigma)^{-1/2}$. Note, however, that $\Omega$ is not typically known in practice but can be consistently estimated by $\widehat\Omega=\operatorname{Diag}(\widehat\Sigma)^{-1/2}\widehat\Sigma\operatorname{Diag}(\widehat\Sigma)^{-1/2}$. Let $\hat{\delta}^{(s)}=\operatorname{Diag}(\widehat\Sigma_{dd}^{(s)})^{-1/2}\sqrt{n}\hat{d}^{(s)}$, $\hat s^*$ denote the subset of the set of indices $\{1,\ldots,k\}$ that maximizes $\widehat\Omega_{bd^{(s)}}\widehat\Omega_{d^{(s)}d^{(s)}}^{-1}\widehat\Omega_{d^{(s)}b}$ amongst all subsets of indices $s\subseteq\{1,\ldots,k\}$ such that the elements of $\widehat\Omega_{bd^{(s)}}\widehat\Omega_{d^{(s)}d^{(s)}}^{-1}$ are non-negative and $(\hat s_1^*,\hat s_2^*)$ denote the subsets of the set of indices $\{1,\ldots,k\}$ that maximize ${\widehat\Omega_{\beta \delta^{(s_1)}} \widehat\Omega_{ \delta^{(s_1)} \delta^{(s_1)}}^{-1}\widehat\Omega_{ \delta^{(s_1)}\beta}}$ and ${\widehat\Omega_{\beta \delta^{(s_2)}} \widehat\Omega_{ \delta^{(s_2)} \delta^{(s_2)}}^{-1}\widehat\Omega_{ \delta^{(s_2)}\beta}}$ amongst all subsets of indices $s_1,s_2\subseteq\{1,\ldots,k\}$ such that the elements of $\widehat\Omega_{\beta \delta^{(s_1)}} \widehat\Omega_{ \delta^{(s_1)} \delta^{(s_1)}}^{-1}$ are non-negative and the elements of $\widehat\Omega_{\beta \delta^{(s_2)}} \widehat\Omega_{ \delta^{(s_2)} \delta^{(s_2)}}^{-1}$ are non-positive.

The distributional approximation in (ref) and the availability of the consistent estimator $\widehat\Omega$ suggest that we can use

gather[gather omitted — 639 chars of source]

and

align[align omitted — 921 chars of source]

as upper one-sided and two-sided CIs for the parameter $b$, where $\widehat{CI}_u^*(\cdot)$ and $\widehat{CI}_t^*(\cdot)$ are defined in Algorithms One-Sided* and Two-Sided*. The theoretical results of the following section formally confirm that these CIs attain uniformly correct asymptotic coverage under weak conditions.

Asymptotic Properties

We now present theoretical results ensuring the uniformly correct asymptotic coverage of both the one- and two-sided finite-sample CIs defined in (ref) and (ref), as well as a uniform upper bound on their asymptotic coverage, under a set of widely-applicable sufficient conditions on the parameter space. In particular, let the parameter $\lambda$ index the true distribution of the observations used to construct the CIs and decompose $\lambda$ as follows: $\lambda=(b,d,\Sigma,F)$, where $b$ is the scalar parameter of interest, $d$ is the nuisance parameter known to have all elements greater than zero, $\Sigma$ is the asymptotic variance corresponding to the parameter estimator $(\hat b_n,\hat d_n^{\prime})^{\prime}$ used by the researcher and $F$ is a (potentially) infinite-dimensional parameter that, along with $(b,d)$, determines the distribution of the observed data. We assume that we have a consistent estimator $\widehat\Sigma_n$ of $\Sigma$ at our disposal.

The parameter space $\Lambda$ for $\lambda$ is defined to include parameters $\lambda=(b,d,\Sigma,F)$ such that for some finite $\kappa>0$, the following conditions hold:

(i) $b\in\mathbb{R}$ and $d\in\mathbb{R}_+^k$ for some positive integer $k$;

(ii) $\Sigma\in \Phi$, $\lambda_{\min}(\Sigma)\geq \kappa$ and $\lambda_{\max}(\Sigma)\leq \kappa^{-1}$, where $\Phi$ denotes the set of all positive definite covariance matrices.

In addition, under any sequence of parameters $\{\lambda_{n,\mathfrak{b},\mathfrak{d},\Sigma^*}=(b_{n,\mathfrak{b}},d_{n,\mathfrak{d}},\Sigma_{n,\Sigma^*},F_{n,\mathfrak{b},\mathfrak{d},\Sigma^*}):n\geq 1\}$ in $\Lambda$ such that

gather[gather omitted — 200 chars of source]

for some $(\mathfrak{b},\mathfrak{d},\Sigma^*)\in \mathbb{R}_{\infty}\times \mathbb{R}_{+,\infty}^k\times \Phi$, the following remaining conditions hold:

(iii) $\widehat\Sigma_n$ exists and $\lambda_{\min}(\widehat\Sigma_n)>0$ with probability 1 for all $n\geq 1$ and $\widehat\Sigma_n\overset{p}\longrightarrow \Sigma^*$;

(iv) $\sqrt{n}(\hat b_n-b_{n,\mathfrak{b}},\hat d_n^{\prime}-d_{n,\mathfrak{d}}^{\prime})^{\prime}\overset{d}\longrightarrow \mathcal{N}(0,\Sigma^*)$;

(v) for any sequence $\{\lambda_{n,\mathfrak{b},\mathfrak{d},\Sigma^*}\}$ in $\Lambda$ and any subsequence $\{s_n:n\geq 1\}$ of $\{n:n\geq 1\}$ for which (ref)--(ref) hold along the subsequence, conditions (iii)--(iv) also hold along the subsequence.

In conjunction with a particular model, parameter estimator $(\hat b_n,\hat d_n^{\prime})^{\prime}$ and covariance matrix estimator $\widehat\Sigma_n$, this definition of the parameter space $\Lambda$ effectively serves as a set of high-level assumptions on the underlying DGP. More specifically, (i)--(ii) are standard parameter space assumptions while (iii)--(v) can typically be verified under standard dependence and moment conditions on the underlying data via laws of large numbers and central limit theorems. We refer the interested reader to Appendix (ref) for details in the context of the standard linear regression model.

With the relevant parameter space defined, we may now state the main theoretical result of this paper that establishes lower and upper bounds on the uniform asymptotic coverage probability of the CIs we propose.

theoremFor $\alpha\in (0,1/2)$ and $\gamma \in (0,\alpha)$, \[\liminf_{n\rightarrow\infty}\text{ }\inf_{\lambda\in\Lambda}P_\lambda\left(b\in CI_{\cdot,n}(\hat b_n,\hat d_n;\widehat\Sigma_n)\right)\geq 1-\alpha\] and \[\limsup_{n\rightarrow\infty} \text{ }\sup_{\lambda\in\Lambda}P_\lambda\left(b\in CI_{\cdot,n}(\hat b_n,\hat d_n;\widehat\Sigma_n)\right)\leq 1-\alpha+\gamma,\] where $CI_{\cdot,n}(\cdot)$ is equal to either $CI_{u,n}(\cdot)$ or $CI_{t,n}(\cdot)$.

These results show not only that our proposed CIs have correct asymptotic coverage in a strong sense but also that by choosing $\gamma$ to be “small” reduces how conservative the CIs can be. However, there is a tradeoff in the choice of $\gamma$: although a smaller $\gamma$ leads to CIs that are closer to being similar across the parameter space, it also allows for less length gains when the elements of $d$ are close or equal to zero.

Empirical Application of Sign-Restricted Regression

For our proposed CIs to to be able to improve upon the length of standard CIs in the standard linear regression context, the researcher must know the sign of at least one of the control variables' coefficients and the estimator of the coefficient of interest must be (asymptotically) correlated with the estimator of the sign-restricted control variables' coefficients. Both conditions are often satisfied in the context of treatment effect regressions for cross-cutting/factorial designs in field experiments. Take, for example, the 2$\times$2 factorial design:

equation[equation omitted — 110 chars of source]

where $E[u|T_1,T_2] = 0$ and $T_1$ and $T_2$ denote two independent, randomly assigned treatments with $T_i \in \{0,1\}$ for $i \in \{1,2\}$. Here, $\alpha_1$ and $\alpha_2$ are the treatment effects of $T_1$ and $T_2$ “relative to a business-as-usual counterfactual” \citep*{Muralidharan19} and $\alpha_3$ is the “interaction effect”, i.e., the treatment effect of jointly providing both treatments minus the sum of the treatment effects of $T_1$ and $T_2$.\footnote{Using the potential outcomes notation, where $Y_{t_1,t_2}$ is the potential outcome of $Y$ when $T_1 = t_1$ and $T_2 = t_2$, the three treatment effects can be written as $\alpha_1 = E[Y_{1,0} - Y_{0,0}]$, $\alpha_2 = E[Y_{0,1} - Y_{0,0}]$, and $\alpha_3 = E[Y_{1,1} - Y_{0,0}] - (E[Y_{1,0} - Y_{0,0}] + E[Y_{0,1} - Y_{0,0}])$.} If $Y$ is a “positive” outcome, it is often reasonable to assume that $\alpha_1 \geq 0$ and $\alpha_2 \geq 0$. For example, a research ethics committee is unlikely to clear an experimental design if this is not the case. Furthermore, the OLS estimators of the three treatment effects are likely to be highly correlated in this setting. For example, if each treatment is assigned with probability 1/2 and the error term $u$ is conditionally homoskedastic, then the asymptotic correlation matrix of $\sqrt{n}(\hat{\alpha}_1,\hat{\alpha}_2,\hat{\alpha}_3)'$ is given by \[ \left[

array[array omitted — 101 chars of source]

\right]. \]

Any of the three treatment effects may be of interest and, under the assumption that $\alpha_1 \geq 0$ and $\alpha_2 \geq 0$, it is reasonable to be interested in upper one-sided CIs for $\alpha_1$ and $\alpha_2$ and a two-sided CI for $\alpha_3$. The above correlation structure implies that our upper one-sided CIs for ${\alpha}_1$ and ${\alpha}_2$ have the potential to improve upon the length of standard upper one-sided CIs. Similarly, our two-sided CI for $\alpha_3$ has the potential to improve upon the length of the standard two-sided CI through a smaller upper bound.

Sometimes researchers are interested in the following alternative specification of the above regression:

equation[equation omitted — 145 chars of source]

where $\alpha^*_3 = \alpha_3 - \alpha_1 - \alpha_2$ is the effect of “both” treatments provided jointly, relative to a business-as-usual counterfactual.\footnote{That is $\alpha_3^* = E[Y_{1,1} - Y_{0,0}]$.} This regression again results in high correlation between OLS estimators: under the same conditions as in the example above, the asymptotic correlation matrix of $\sqrt{n}(\hat{\alpha}_1,\hat{\alpha}_2,\hat{\alpha}^*_3)'$ is given by \[ \left[

array[array omitted — 77 chars of source]

\right]. \] In this case, our two-sided CI for $\alpha_3^*$ has the potential to improve upon the length of the standard two-sided CI through a larger lower bound.

To illustrate the usefulness of our proposed CIs, we apply them in the context of a field experiment where a 2$\times$2 factorial design was used. In particular, we revisit blattman2017 (BJS) who recruited 999 poor young men in Liberia who exhibited “high rates of violence, crime, and other antisocial behaviors” to participate in an experiment. The two treatments are “therapy”, an eight-week program of group cognitive behavior therapy, and “cash”, a \$200 grant corresponding to roughly three months' wages. In simple terms, the main research question is whether “therapy” and “cash” can help reduce violent, criminal, and other antisocial behaviors. The hypothesized channels are improved noncognitive skills such as self-control (“therapy”) and an increase in legal work (“cash”). BJS conducted two follow-up surveys, the first 2--5 weeks and the second 12--13 months after the intervention to elicit “short-term” and “long-term” impacts, respectively.

table[table omitted — 675 chars of source]

Table (ref) reproduces the results concerning the treatments' long-term impact on a summary index of antisocial behaviors (times minus one) (cf.\ the first row of Panel B of Table 2 in BJS). The table includes one of the main findings of BJS: while the two treatments do not have statistically significant long-term effects in isolation, they do have a joint positive long-term effect on the index of antisocial behaviors. Column 1 ($\hat{\beta}$) shows the OLS point estimates for “therapy” (T), “cash” (C), “both” (B), and “interaction” (I) as defined above and column 2 (SE) reports the corresponding (heteroskedasticity-robust) standard errors. Note that BJS only consider the specification given in equation (ref), i.e., they only estimate the effect of “both” treatments and not the “interaction” effect.\footnote{In fact, BJS consider the specification given in equation (ref) augmented by a set of additional controls. For the purpose of this analysis, we take the signs of these additional controls as unknown. See BJS for more information on the additional controls.} Column 3 (SSCI\textemdash Simple and Short Confidence Interval) shows our proposed CIs for $\alpha = 0.05$, which are upper one-sided for T and C and two-sided for B and I, when assuming that the treatment effects of “therapy” and “cash” are a priori known to be nonnegative. They are constructed using Algorithms One-Sided$^*$ and Two-Sided$^*$ in combination with the response surface approximations.\footnote{The estimated (asymptotic) correlation matrices for the estimator of the effects of i) T, C, and B and ii) T, C, and I are given by \[ \left[

array[array omitted — 119 chars of source]

\right] and \left[

array[array omitted — 117 chars of source]

\right], \] respectively. We augmented the corresponding regressions by the same set of controls as BJS, cf.\ footnote (ref).} Column 5 (SCI\textemdash Standard Confidence Interval) shows the corresponding standard CIs. Columns 4 and 6 (both (E)L) give the (“excess”) lengths of SSCI and SCI, where the “excess” length of one-sided CIs here is computed as the difference between $\hat b$ and the CI's lower bound.\footnote{We write “excess” in quotes to emphasize the fact that this is not equal to the true excess length that cannot be computed here in the absence of knowledge of the true value of the regression coefficients.} Column 7 (Ratio) computes the ratio of the (“excess”) length of SSCI relative to SCI. We find that, while our proposed CI is marginally longer than the standard CI for C\textemdash its “excess” length reaching the bound on expected excess length increase of $\sim 3\%$, it is much shorter for T, B, and I.

Calibrated Simulations for Sign-Restricted Regression

To illustrate the finite-sample properties of our proposed CIs, we perform a Monte Carlo study calibrated to the BJS factorial design regression of the previous section. In particular, we create 10,000 bootstrap samples by drawing with replacement from the sample of $n = 947$ men underlying the regression results in Table (ref). In each bootstrap sample, we estimate the regressions (ref) and (ref). Since the expected value of the treatment effect of “cash” under the empirical distribution is equal to the point estimate in the original sample, -0.1316, it is outside of the sign-restricted parameter space $\alpha_2\geq 0$. We therefore recenter the estimates of the treatment effect of “cash” over the bootstrap samples to have mean zero (by adding 0.1316). For each bootstrap sample, we construct our proposed CIs, using Algorithms One-Sided$^*$ and Two-Sided$^*$ in combination with the response surface approximations, standard CIs, the (excess) length of each CI and whether they cover the true parameter value, i.e., the corresponding (re-centered) point estimate in the original sample. All CIs are constructed using standard heteroskedasticity-robust variance-covariance matrix estimators computed within each bootstrap sample. Since the empirical distribution from which the bootstrap samples are drawn is not normally distributed, this simulation exercise captures the effect on CI coverage of departures from the large sample normal means problem of Section (ref).

comment{\color{red} \begin{table}[h!] \begin{center} \caption{Monte Carlo results} \begin{tabular}{cc|rrrrrr} \hline \hline & & \multicolumn{1}{c}{T} & \multicolumn{1}{c}{C} & \multicolumn{1}{c}{B} & \multicolumn{1}{c}{I} & \multicolumn{1}{c}{B0} & \multicolumn{1}{c}{I0} \\ \hline \multirow{2}{*}{CP} & SSCI & 94.81 & 94.99 & 95.24 & 95.09 & 94.81 & 94.45 \\ & SCI & 94.81 & 94.61& 94.84 & 94.70 & 94.84 & 94.70 \\ \hline \multirow{3}{*}{E(E)L} & SSCI & 0.14 & 0.16 & 0.35 & 0.44 & 0.33 & 0.41 \\ & SCI & 0.15 & 0.16 & 0.36 & 0.50 & 0.36 & 0.50 \\ & Ratio & 92.24 & 100.70 &97.25 & 88.72 & 91.64 & 81.36 \\ \hline \end{tabular} \end{center} \end{table} }
table[table omitted — 727 chars of source]

Table (ref) reports the coverage probability (CP) computed across bootstrap realizations of our proposed CIs and of the standard CIs for all four treatment effects, T, C, B, and I. Table (ref) also reports the expected (excess) length (E(E)L) of these CIs across the bootstrap realizations. In addition to the above DGP, we also consider a modification where the true value of the treatment effect of “therapy” is set equal to zero (by subtracting the point estimate in the original sample, 0.0829, from the corresponding estimates in the bootstrap iterations). The corresponding results for the effect of “both” treatments and the “interaction” effect are given in the last two columns, B0 and I0.

We observe that our proposed CIs have good finite sample coverage, comparable to that of the standard CIs, with little coverage distortion despite the non-normally distributed data. In terms of expected (excess) length, most of our proposed CIs offer sizeable improvements over standard CIs, with expected (excess) length improvements of up to nearly 18% for this particular data calibration.

comment\section{Finite-Sample Properties for Sign-Restricted Regression} The following table shows the coverage probability (CP), the expected (excess) length (E(E)P) of the proposed confidence intervals, and the ratio of the latter over the E(E)P of the corresponding standard confidence intervals. Here, “TO” stands for “Therapy only”, “CO” stands for “Cash only”, “B” for “Both”, and “I” for “Interaction”, see Section (ref) for more details. Table (ref) is calibrated to the empirical application. As we are imposing that the coefficients on TO and CO are non-negative, we take as true mean of our asymptotically normal random vector the projection of the point estimates onto the corresponding parameter space (with respect to the norm $\| \lambda \| = (\lambda'\hat{\Omega}^{-1}\lambda)^{1/2}$), i.e., $\hat{y} = \operatorname{argmin}_{y \in \mathbb{R}_+^2\times\mathbb{R}} (\hat{\theta} - y)'\hat{\Omega}^{-1}(\hat{\theta} - y)$. When our third coefficient estimate corresponds to the variable “Both”, we have \[ \hat{\theta} = \left[ \begin{array}{r} 0.8921 \\ -1.3586 \\ 2.7945 \end{array} \right], \ \hat{\Omega} = \left[ \begin{array}{rrr} 1.0000 & 0.5238 & 0.6104 \\ 0.5238 & 1.0000 & 0.5543 \\ 0.6104 & 0.5543 & 1.0000 \end{array} \right], \text{ and } \hat{y} = \left[ \begin{array}{r} 1.6037 \\ 0 \\ 3.5476 \end{array} \right]. \] When instead our third coefficient estimate corresponds to the variable “Interaction”, we have \[ \hat{\theta} = \left[ \begin{array}{r} 0.8921 \\ -1.3586 \\ 2.3544 \end{array} \right], \ \hat{\Omega} = \left[ \begin{array}{rrr} 1.0000 & 0.5238 & -0.7154 \\ 0.5238 & 1.0000 & -0.7699 \\ -0.7154 & -0.7699 & 1.0000 \end{array} \right], \text{ and } \hat{y} = \left[ \begin{array}{r} 1.6037 \\ 0 \\ 1.3085 \end{array} \right]. \] For TO and CO we consider upper one-sided CIs and for B and I we consider two-sided CIs. {\color{red} \begin{table}[h!] \begin{center} \caption{Finite-Sample results (with projection)} \begin{tabular}{c|rrrrrrrrrr} \hline \hline & \multicolumn{1}{c}{TO} & \multicolumn{1}{c}{CO} & \multicolumn{1}{c}{B$^*$} & \multicolumn{1}{c}{B} & \multicolumn{1}{c}{B$^*$0} & \multicolumn{1}{c}{B0} & \multicolumn{1}{c}{I$^*$} & \multicolumn{1}{c}{I} & \multicolumn{1}{c}{I$^*$0} & \multicolumn{1}{c}{I0} \\ \hline CP & 95.00 & 95.50 & 95.49 & 95.50 & 94.97 & 95.00 & 95.52 & 95.52 & 94.96 & 95.00\\ E(E)L & 1.52 & 1.69 & 3.90 & 3.91 & 3.58 & 3.58 & 3.65 & 3.65 & 3.18 & 3.19 \\ Ratio & 92.43 & 102.55 & 99.55 & 99.62 & 91.25 & 91.45 & 93.01 & 93.17 & 81.17 & 81.44\\ \hline \end{tabular} \end{center} \end{table} } {\color{red} \begin{table}[h!] \begin{center} \caption{Finite-Sample results (with “truncation”)} \begin{tabular}{c|rrrrrrrrrr} \hline \hline & \multicolumn{1}{c}{TO} & \multicolumn{1}{c}{CO} & \multicolumn{1}{c}{B$^*$} & \multicolumn{1}{c}{B} & \multicolumn{1}{c}{B$^*$0} & \multicolumn{1}{c}{B0} & \multicolumn{1}{c}{I$^*$} & \multicolumn{1}{c}{I} & \multicolumn{1}{c}{I$^*$0} & \multicolumn{1}{c}{I0} \\ \hline CP & 95.00 & 95.46 & 95.43 & 95.44 & 94.97 & 95.00 & 95.48 & 95.48 & 94.96 & 95.00\\ E(E)L & 1.52 & 1.65 & 3.79 & 3.80 & 3.58 & 3.58 & 3.46 & 3.47 & 3.18 & 3.19 \\ Ratio & 92.43 & 100.56 &96.78 & 96.90 & 91.25 & 91.45 & 88.38 & 88.60 & 81.17 & 81.44\\ \hline \end{tabular} \end{center} \end{table} } The superscript $^*$ denotes that the Algorithm Two-Sided$^*$ was used as opposed to the Algorithm Two-Sided. And the “0” signifies that $\hat{y}_{1:2}$ was replaced by $0_2$.