Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
35,161 characters · 5 sections · 49 citation commands
A Simple, Short, but Never-Empty Confidence Interval for Partially Identified Parameters
\onehalfspacing
Inference under partial identification is by now the subject of a broad literature.\footnote{See Manski2003 for an early monograph, Tamer10 for a historical introductions, and CS17 and MolinariHOE for recent surveys that extensively cover inference.} Only recently did attention turn to the following concern: If a partially identified model is misspecified, this may manifest in either an empty or --and arguably worse-- in a misleadingly small confidence region. That is, misspecified inference can be spuriously precise.
The reason is that most confidence regions used in partial identification invert tests of $H_0:\theta \in \Theta_I$; here, $\theta$ is a parameter and $\Theta_I$ is the identified set. If $H_0$ is rejected at every $\theta$, the confidence region is empty. If $H_0$ is barely not rejected at a few parameter values, the confidence region may be very small. This issue is empirically relevant. For example, an empty sample analog of $\Theta_I$ occurs in Haushofer, whose inquiry sparked the present research and whose data are reanalyzed below.
The literature on this issue is still young. PT11 provide an early diagnosis. KW13 propose a notion of pseudotrue identified set and an estimator thereof. MolinariHOE explains the issue in detail and highlights it as important area for further investigation. The most thorough treatment is by AndrewsKwon19, who emphasize the issue's importance and provide a general inference method that avoids spurious precision and ensures coverage of a pseudotrue identified set.
The present paper is in the spirit of AndrewsKwon19. I focus on the simple but empirically salient case of a scalar parameter with upper and lower bounds whose estimators are jointly asymptotically normal. That is, I revisit the setting of IM04 and Stoye09. For this setting, I propose a confidence interval with the following features:
For target coverage of $95\%$ and for the special case of uncorrelated estimators, e.g. in this paper's empirical application, the confidence interval can be verbally defined as follows:
While this paper generally proposes a somewhat less “cute" procedure with broader applicability, this specialized finding is probably the most striking part.\footnote{Full disclaimer: I discovered it by simulation and initially assumed a bug.} Neither of the above two intervals is valid by itself; it is just that their coverage events are correlated in exactly the right way.
Section (ref) develops the proposal more formally and gives an intuition for why it works, though proofs are relegated to the Appendix. Section (ref) provides a numerical illustration and Section (ref) an application to the data that motivated this research. Section (ref) concludes.
While the interpretation of what follows is inference on a scalar parameter $\theta$, the only assumption is that one has well-behaved estimators of two other parameter values.
The motivation is that the researcher estimates an identified set $\Theta_I \equiv [\theta_L,\theta_U]$ containing a true parameter value $\theta$. Assumption (ref) is unrestrictive if, as in the empirical application, $(\hat{\theta}_L,\hat{\theta}_U)$ are smooth functions of sample moments. It is unlikely to hold for intersection bounds AS13,CLR13 and will hold for bounds that result from projecting a higher-dimensional identified set BCS17,KMS19, including components of partially identified vectors, only in benign cases.
The obvious estimator of $\Theta_I$ is $[\hat{\theta}_L,\hat{\theta}_U]$, but defining a confidence interval is delicate. Following IM04, the literature mostly focuses on confidence intervals that (asymptotically) contain the true parameter value with prespecified probability $(1-\alpha)$, irrespective of its location in $\Theta_I$, i.e. confidence intervals that control $\inf_{\theta \in \Theta_I} \Pr(\theta \in CI)$. Finding such intervals is subtle because the nature of the testing problem qualitatively depends on the length $\Delta \equiv \theta_U-\theta_L$ of $\Theta_I$. Heuristically, this problem is one-sided if $\Delta$ is “large" and two-sided if it is “short," i.e. near point identification. Ascertaining which case obtains is subject to difficulties reminiscient of post-model selection inference LP05 and parameter-on-the-boundary issues Andrews00.
The literature on how to circumvent this issue is by now considerable. Most approaches invert a test, that is, they report all values of $\theta$ for which $H_0:\theta \in \Theta_I$ was not rejected. Any such confidence set can be empty; in this paper's settings, that will happen if $\hat{\theta}_L$ is much larger than $\hat{\theta}_U$, where the meaning of “much" varies across papers. This feature can be advertised as an embedded specification test but may not be wanted.\footnote{That was the sales pitch in Stoye09, but not all referees were sold on it. The embedded specification test is analyzed in more detail by AS10.} Arguably even more problematic is that, if the model is misspecified, a test inversion confidence interval can be short, suggesting precision when the true issue is misspecification. A specification test will not resolve this: In this paper's setting, the best-practice such test BCS15 just reports whether the test inversion interval is empty.\footnote{This equivalence does not generalize, but AndrewsKwon19 show that in “slightly misspecified" parameter regimes, spuriously precise inference generally coexists with low power of specification tests.}
Addressing this concern requires a notion of coverage for the case of misspecification, i.e. if $\theta_L>\theta_U$. Following AndrewsKwon19, define the pseudotrue identified set
This definition is natural because $\Theta_I^*=\arg\min_\theta \max\{(\theta-\theta_U)/\sigma_U,(\theta_L-\theta)/\sigma_L,0\}$; thus, $\Theta_I^*$ is the estimand implied by the frequent choice of $\max\{(\theta-\hat{\theta}_U)/\hat{\sigma}_U,(\hat{\theta}_L-\theta)/\hat{\sigma}_L,0\}$ as test statistic. Note also that $\Theta_I^*$ is never empty and that $\Theta_I^*=\Theta_I$ whenever $\Theta_I \neq \emptyset$.
The revised notion of validity of a confidence interval is as follows:
Forcing coverage of $\theta^*$ will ensure that the interval is nonempty and also that it is statistically interpretable as targeting $\Theta_I^*$. An obvious caveat is that, as with the related literature going back to White82, the coverage target's substantive relevance may not be clear if the model is in fact misspecified. As AndrewsKwon19 elaborate, this has to be traded off against concerns with spurious precision.
While the coverage notion exactly mimics AndrewsKwon19, the confidence interval will be quite different. It goes “back to basics" in that, like early entries in the literature IM04,Stoye09, it essentially just adds a certain number of standard errors to estimated bounds. An advantage is computational and conceptual simplicity; with test inversion intervals, critical values generally depend on $\theta$ even in this simple setting and therefore must be computed many times. However, the main motivation is that the new interval performs well. Its heuristic definition is as follows:
The new confidence interval is obviously never empty; indeed, its length cannot drop below $2\hat{\sigma}^*\Phi^{-1}(1-\alpha/2)$. Its formal definition and theoretical justification are as follows.
Expression (ref) is numerically evaluated for different values of $\rho$ and target coverages in Table (ref). In particular, simulation with very high accuracy suggests that $\hat{c}$ is just the one-sided critical value for $\rho$ up to at least $.8$; it then gradually increases toward the two-sided critical value, which is easily seen to solve (ref) for $\hat{\rho}=1$.\footnote{The table was generated by gridding and using $B=4000000$ simulations. This is feasible on a run-of-the-mill netbook. The relevant simulation error is the coverage error at the suggested $\hat{c}$. Given $B$, it will be much smaller than what is routinely accepted in simulation-based, e.g. bootstrap, inference. For $\rho \leq .8$, further simulations establish to high accuracy that coverage is first increasing and then decreasing in $\Delta$ and minimized as $\Delta \to \infty$, the same feature that is analytically proved for $\rho=0$ and which justifies $\hat{c}=\Phi^{-1}(1-\alpha)$.}
The proof of Theorem (ref) contains three steps. First, it is relatively routine to show that $CI_{MA}$ would be valid if, in line with the heuristic definition, expression (ref) explicitly took the infimum also over values of $(\sigma_L,\sigma_U)$ as well as $\theta \in \Theta_I$. In a second step, we can concentrate out all of these. In particular, one can restrict attention to one of $\theta=\theta_L$ or $\theta=\theta_U$; expression (ref) arbitrarily chooses the latter. This finding is not obvious: For given $\Delta$, coverage is not equally minimized at the interval's endpoints; it is only that the corresponding infima over $\Delta \in [0,\infty)$ are the same. As final flourish in this step, it turns out that asymptotic coverage at $\theta_U$ depends on $(\Delta,\sigma_L,\sigma_U)$ only through $\Delta/\sigma_L$. For the purpose of evaluating worst-case coverage over $\Delta \geq 0$, we can therefore set both standard deviations to $1$.
The final, and by far most delicate, step is that if $\rho=0$, coverage is provably minimized as $\Delta \to \infty$, justifying use of the one-sided critical value $\hat{c}=\Phi^{-1}(1-\alpha)$. To appreciate this claim, consider again the two components of $CI_{MA}$ in (ref). For $\alpha=.05$, the left-hand interval's coverage for either $\theta_L$ or $\theta_U$ may be as low as $.9$ if $\Delta=0$ and approach $.95$ from below as $\Delta \to \infty$. The right-hand interval's coverage of these values is $.95$ at $\Delta=0$ (where both coincide with $\theta^*$) but rapidly decreases to $0$ as $\Delta$ increases. That these effects aggregate to coverage uniformly above $.95$ is far from obvious and heavily relies on specific features of the bivariate Normal distribution.
Numerically, the final step extends to moderate $\rho$ (see again Table (ref)), and the proof uses conservative bounds. Some analytic result of higher generality might, therefore, be available. However, for large positive $\rho$, coverage is minimized at small positive $\Delta$. Therefore, if $\rho$ is unknown, estimating it cannot be avoided. In particular, in view of Table (ref), a pre-test for “small enough" $\rho$ would be counterproductive: Since $\hat{c}$ as a function of $\rho$ is mostly completely flat, one would be unlikely to recover the inferential cost (in the sense of Remark (ref)) of the pre-test.
Figures (ref) and (ref) compare $CI_{MA}$ with a test inversion interval $CI_{TI}$ that arguably reflects the state of the established literature.\footnote{The interval closely follows RSW14; other established methods AS10,AB12,Bugni10,Canay10 would inform similar constructions. As of writing of this manuscript, at least two rather distinct (from the preceding and from each other) proposals are in the pipeline APR19,CS19. Both invert a test and can be empty; APR19 also has a tuning parameter. They are compared in CS19. A comparison of all these approaches in simple examples might be worthwhile.} It inverts a test of $H_0:\theta \leq \theta_U,\theta \geq \theta_L$ by taking the maximum (studentized) violation as test statistic, i.e. the same test statistic that generally implies $\Theta_I^*$ as pseudotrue identified set. The critical value is based on a pre-test --specifically, a one-sided $(.1\alpha)$-Wald test-- that potentially discards one of the inequality constraints as nonbinding. Depending on the pre-test's result, the critical value is then either a simple one-sided critical value or computed by a simulation that takes $\rho$ into account. In either case, the second-stage test is of size $.9\alpha$, so that the pre-test is accounted for by Bonferroni correction. The resulting test is inverted, and the critical value is recomputed, as $\theta$ changes, making the interval considerably shorter than early entries in the literature IM04,Stoye09. Compared to $CI_{MA}$, test inversion adds orders of magnitude of computational cost, though at a very low absolute level. I abstract from asymptotic approximation by drawing estimators straight from limiting distributions and taking $(\sigma_L,\sigma_U,\rho)$ to be known. Interval length $\Delta$ is denominated in estimator standard errors because $\sqrt{n}\sigma_L=\sqrt{n}\sigma_U=1$ throughout.
The comparison is extended into the misspecified range by letting $\Delta$ take on negative values. The test inversion interval obviously undercovers in that range. To clarify comparisons, I also compute $CI_{TI} \cup CI_{\theta^*}$. Recall that $CI_{MA}$ can be loosely intuited as refining this construction by adjusting the critical value to account for union-taking. Nominal coverage is $95\%$ throughout.
Figure (ref) illustrates the results for $\rho=0$ (top panels) and $\rho=.7$ (bottom panels); Figure (ref) extends the exercise to $\rho=-.7$ and finally to $\rho=.95$. The last case is arguably contrived but serves to illustrate that $\Delta\to \infty$ is not always least favorable. By the same token, this is the only case in which $\hat{c}>\Phi^{-1}(.95)$.
With one caveat discussed below, the figures suggest dominating performance of $CI_{MA}$: It is shorter, and this is also reflected in more precise size control and thereby more power of the implied test. The advantage is especially apparent for small positive $\Delta$. What happens here is that the correction provided by $CI_{\theta^*}$ allows $CI_{MA}$ to transition to just adding $1.64$ standard errors considerably more quickly than a pre-test could justify. Indeed, for $\rho\leq .4$, this transition occurs at a negative estimated interval length $\hat{\Delta}$; that is, $CI_{MA}$ just adds $1.64$ standard errors to bounds estimates whenever these are ordered in the expected way. The slight advantage of $CI_{MA}$ for large $\Delta$ reflects that $CI_{TI}$ accounts for a pre-test.
One might wonder how AndrewsKwon19 would perform in the example. While the exact answer depends on choice of multiple tuning parameters, some qualitative considerations are as follows. Their interval starts from $CI_{TI}$ and expands it in order to avoid spurious precision.\footnote{AndrewsKwon19 implement $CI_{TI}$ through AS10 but point out that RSW14 could be used instead. The difference will be small in the present setting.} As a result, it will be bounded from below in both length and coverage by the blue curves in Figures (ref) and (ref). In an initial refinement, AndrewsKwon19 form the union between $CI_{TI}$ and a never-empty confidence interval. Their preferred confidence interval does this only if an additional pre-test fails to reject misspecification. While this mitigates the effect of expanding $CI_{TI}$, the final confidence interval still contains $CI_{TI}$ and considerably exceeds it for small positive $\Delta$ (see their Section 8.1, whose setting resembles the present one). This will obviously be reflected in its statistical performance. Conversely, an intriguing feature of $CI_{MA}$ is that it “spends" the “coverage capital" gained from ensuring nonemptiness by being shorter than $CI_{TI}$ for interesting values of $\Delta$. In fairness to AndrewsKwon19, it appears far from obvious how to implement such a feature in their much more general setting.
The advantage of $CI_{MA}$ fades out, and even reverses, in the special case where $\rho\to 1$ but not $\Delta \to 0$. In that limit, $\hat{c}$ will converge to the two-sided critical value, whereas a pre-test will eventually recommend a one-sided test. While such scenarios can obviously be simulated, they arguably are contrived. The possibility of high $\rho$ and correspondingly precise estimation of $\Delta$ is empirically relevant, but it corresponds to the superefficiency case discussed in Remark (ref) and therefore to small $\Delta$ as well as to a case distinction that can be decided in pre-data analysis. Also, one could in principle fix this issue by layering a pre-test on top of $CI_{MA}$; however, as general advice in this matter, I stand by Remark (ref).
\Citet{Haushofer} estimate upper and lower bounds on behavioral parameters from different treatments in a between-subjects design, meaning that estimators are uncorrelated. At the same time, bounds can and did in fact invert, triggering an inquiry by the authors that led to the present paper.
Table (ref) displays estimated bounds, $CI_{MA}$, and $CI_{TI}$ for selected instances of the “weak bounds" data. This refers to a baseline setting before inducing experimenter demand. For more details, I refer to Haushofer, particularly Figure 1 and corresponding explanations. The last column divides the length of $CI_{MA}$ by the length of $CI_{TI}$, subtracting $\max\{\hat{\Delta},0\}$ from both. Both intervals make full use of $\rho=0$ being known.
The comparison is between $CI_{MA}$ and $CI_{TI}$; obviously, $CI_{TI}\cup CI_{\theta^*}$ would be larger than $CI_{TI}$. The data include one case (*) where bound estimators are inverted and where ex post, $CI_{MA}=CI_{\theta^*}$.\footnote{This case would not have led any specification test to reject the model, even before taking multiple hypothesis testing into account.} There are also two cases (**) of short estimated intervals (relative to standard errors), i.e. of near point identification. Because $CI_{MA}$ cannot be empty, one might have conjectured it to be the longer one in these cases. In fact, it is noticeably shorter in all of them -- the effect of “spending coverage capital" from the nonemptiness correction dominates. In all other cases, both intervals effectively add $1.64$ standard errors.\footnote{In those cases, the small differences favoring $CI_{MA}$ reflect Bonferroni adjustment for pre-tests, i.e. the specifics of RSW14. In cases where $[\theta_L,\theta_U]$ is obviously “long," researchers will in practice be tempted to claim an asymptotic pre-test and just use $1.64$ standard errors.}
For a simple, but empirically relevant, partial identification problem, I propose a confidence interval that has competitive size control and length including in the misspecified case, while being extremely easy to compute. The most striking finding is that in many cases, a seemingly crude fix to a nominal $90\%$ confidence interval ensures $95\%$ coverage at little cost in terms of interval length and with practically zero computation. Simulations are encouraging, and the confidence interval improves on current best practice in application to recent lab experiments.
The approach is complementary to AndrewsKwon19, from whom I take the broad motivation as well as the novel coverage requirement. Of course, their approach applies far beyond the present paper's simple setting. On the other hand, it has several tuning parameters and expands a conventional confidence interval, whereas the present proposal is tuning parameter free and compensates for expanding the conventional interval by reducing its standalone nominal coverage. A question of obvious interest, but also beyond my current reach, is whether this last feature can be usefully generalized. As it stands, the present proposal is limited to a specific setting but appears both practical and powerful when that setting obtains.