EconBase
← Back to paper

Can we have it all? Non-asymptotically valid and asymptotically exact confidence intervals for expectations and linear regressions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

106,045 characters · 23 sections · 39 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Can we have it all? Non-asymptotically valid and asymptotically exact confidence intervals for expectations and linear regressions

abstractWe contribute to bridging the gap between large- and finite-sample inference by studying confidence sets (CSs) that are both non-asymptotically valid and asymptotically exact uniformly (NAVAE) over semi-parametric statistical models. NAVAE CSs are not easily obtained; for instance, we show they do not exist over the set of Bernoulli distributions. We first derive a generic sufficient condition: NAVAE CSs are available as soon as uniform asymptotically exact CSs are. Second, building on that connection, we construct closed-form NAVAE confidence intervals (CIs) in two standard settings -- scalar expectations and linear combinations of OLS coefficients -- under moment conditions only. For expectations, our sole requirement is a bounded kurtosis. In the OLS case, our moment constraints accommodate heteroskedasticity and weak exogeneity of the regressors. Under those conditions, we enlarge the Central Limit Theorem-based CIs, which are asymptotically exact, to ensure non-asymptotic guarantees. Those modifications vanish asymptotically so that our CIs coincide with the classical ones in the limit. We illustrate the potential and limitations of our approach through a simulation study. Keywords: non-asymptotic inference; efficient confidence intervals; heteroskedastic linear models. MSC Classification: 62G15; 62J05.

Introduction

Mathematical statistics results are often divided into two branches: non-asymptotic (also called finite-sample) and asymptotic results. In this way, most confidence sets are typically designed to have either good asymptotic properties (based on limiting results, usually the Central Limit Theorem) or good non-asymptotic properties (based on concentration inequalities). In this paper, we try to bridge this gap and construct confidence sets -- more precisely, intervals -- that are NAVAE, i.e., both Non-Asymptotically Valid (their coverage probability is at least their nominal level for any sample size \(n\)) and Asymptotically Exact uniformly over the statistical model (their coverage probability uniformly tends to the nominal level as $n \to \infty$).

Non-asymptotically exact confidence sets (defined by their coverage probability being exactly equal to their nominal level for any sample size) are automatically NAVAE uniformly on the corresponding parameter set. However, they can be constructed only in very particular cases, such as the classical Student confidence interval for the expectation of a normal distribution. Even in seemingly simple cases, constructing uniform NAVAE confidence sets can prove difficult or sometimes even impossible, as displayed in the following result on Bernoulli distributions.

propLet \(\mathfrak{Ber}\) be the class of all Bernoulli distributions with parameter \({p \in (0,1)}\). For any \(\alpha \in (0,1)\), there is no confidence set for $p$ that is both non-asymptotically valid and asymptotically exact uniformly over $\mathfrak{Ber}$ at nominal level \(1 - \alpha\).

Confidence sets and their properties are formally defined in Section (ref) and a proof of this proposition is given in Appendix (ref). Note that it is not a consequence of the impossibility results by bahadur1956 or by bertanha2016impossible. It is also unrelated to dufour1997some's impossibility results, which apply to locally almost unidentified parameters; this is not the case here since there is no identification issue.

Despite the negative result of Proposition (ref), we construct in this paper NAVAE confidence intervals uniformly over infinite-dimensional classes of distributions for two specific parameters:

itemize• the expectation of a (potentially non-normal) distribution, discussed in Section (ref); • a scalar parameter of the form ${u}'\beta_0$, where ${u}$ is a known vector and $\beta_{0}$ corresponds to the coefficients of a linear regression, discussed in Section (ref). Individual coefficients fall in that framework.

To build such confidence intervals, we only impose moment restrictions on our class of distributions. In particular, we do not use any parametric assumption, be it in the expectation or in the linear regression case. In the latter setup, our moment conditions encompass very general regression designs: regressors are only assumed to be orthogonal to error terms, and heteroskedasticity is allowed. Those moment conditions are necessary to avoid the impossibility results of bahadur1956 and enable us to leverage the asymptotic normality of the sample mean/OLS estimator. That property is crucial in our construction. Indeed, our confidence intervals (CIs) are built from asymptotically negligible modifications of the standard intervals based on the Central Limit Theorem (CLT). Those modifications rely on bounds on Edgeworth expansions (see, e.g., esseen1945,cramer1962, and the introduction in derumigny2022explicitBE for additional references) to control the distance of the distribution of a standardized empirical mean to its Gaussian limit and obtain finite-sample guarantees (NAV part). Uniformly over the class delineated by our moment conditions, our CIs further coincide asymptotically with the standard ones based on the CLT and are thus uniformly asymptotically exact (AE part). Besides, that coincidence implies that our intervals are asymptotically of minimal width among NAVAE ones according to the efficiency criterion of romano2000.

Our confidence intervals for the expectation are related to the ones proposed by romano2000 and austern2022efficient. Both papers primarily focus on efficiency instead of uniform asymptotic exactness. Furthermore, they assume that the random variables are bounded or of bounded variation around their expectation. In contrast, the confidence intervals we propose are NAVAE uniformly over a set of distributions that includes distributions with unbounded support (even after recentering at the expectation). The recent paper by waudby2024estimating proposes new confidence intervals that also rely on the boundedness assumption. hall1995 construct confidence intervals that are asymptotically exact uniformly over a certain infinite-dimensional class of distributions, but their CIs are not shown to be non-asymptotically valid.

In the linear regression setting, gossner2013finite build two non-asymptotic inference methods, valid under bounded outcome variables and allowing for heteroskedastic errors. There are several limitations, though, either theoretical or practical: the first method is not exact asymptotically, and neither the first nor the second yield closed-form confidence sets, which may entail cumbersome computations. dHaultfoeuille2024robust and pouliot2024 propose non-asymptotically exact confidence sets by inverting permutation-based tests. In dHaultfoeuille2024robust, non-asymptotic exactness relies on an independence assumption between error terms and the regressors of interest, conditional on the remaining regressors, while in pouliot2024, the imposed condition is that the joint distribution of regressors and errors is invariant to a permutation of the errors. There are two main limitations to these approaches: the maintained conditions restrict possible heteroskedasticity patterns, and, unlike ours, the proposed confidence sets are not of closed form. On the other hand, these methods are shown to be asymptotically exact under assumptions close to ours. In the statistics literature, kuchibhotla2024hulc recently proposed a method to build NAVAE confidence sets on functionals in a broad class of statistical models. While these results are very general and proved under mild moment assumptions on the data distribution, the proposed method is not efficient in linear regression models. This means that it delivers confidence intervals that are not of the smallest possible width asymptotically. In a general \(M\)-estimation framework, brunel2019nonasymptotic derive non-asymptotic confidence intervals with closed-form expressions on individual components of the coefficients' vector. The main drawback of their result is a lack of asymptotic exactness. All in all, to the best of our knowledge, our procedure is the first to meet the following five criteria at the same time in linear regression:

enumerate• non-asymptotic validity, • uniform asymptotic exactness, • efficiency \`a la romano2000, • allowing for arbitrary heteroskedasticity, • having a closed-form expression.

The rest of the paper is organized as follows. Section (ref) introduces some notation and a number of quality measures for confidence sets. In this section, we also present and prove some generic impossibility/possibility results on constructing NAVAE confidence sets. Section (ref) discusses how to construct NAVAE confidence intervals for an expectation and illustrates the major trade-offs and strategies in this simpler case. Some additional material on this question is provided in Appendix (ref). Our CIs for linear regressions are defined and shown to be NAVAE in Section (ref), and we further study the rate of convergence of their coverage towards the nominal level in that section. We discuss the practical implementation of our methods in Section (ref) together with some simulations. The rest of the proofs and useful intermediary lemmas are reported in Appendices (ref), (ref), (ref), and (ref). All confidence intervals introduced in this paper are implemented in the open-source R package NAVAECI NAVAECI.

Notation and quality measures for confidence sets

Notation

For a random variable $D$, we denote by $P_{D}$ its distribution and by \(\mathrm{support}(D)\) or \(\mathrm{support}(P_{D})\) its support. Similarly, \(P_{D, \, U}\) denotes the joint distribution of a pair of random variables \((D, U)\). For a parameter $\theta$ associated with a given statistical model, \(P_\theta\) denotes a distribution indexed by that parameter. Remember that \(P_{D, \, U} = P_D \otimes P_U\) means that $D$ and $U$ are independent. For any set $\mathcal{D}$, $\mathcal{P}(\mathcal{D})$ denotes the set of all probability distributions on $\mathcal{D}$. For any real vector \(u = (u_1, \ldots, u_d)\), \(\left\|u\right\| := (u_1^2 + \ldots + u_d^2)^{1/2}\) denotes its Euclidean or \(\ell_2\)-norm. For matrices, we consider the operator norm induced by the Euclidean norm: for any real matrix \(M\), \(\left\|M\right\|\) denotes its spectral norm (also known as the operator 2-norm), namely the square root of the largest eigenvalue of \(M'M\), with \(M'\) the transpose of \(M\). We also denote by \(\lambda_{\min}(M)\) (respectively \(\lambda_{\max}(M)\)) the smallest (resp. largest) eigenvalue of \(M\), \({ \mathrm{tr} }(M)\) its trace, and \(M^\dagger\) the Moore–Penrose pseudo-inverse of a square matrix \(M\). \({\textnormal{vec}}(\cdot)\) denotes the vectorization of a matrix, that is, if \(M\) is a \(m \times n\) matrix, \({\textnormal{vec}}(M)\) is the \(mn \times 1\) column vector obtained by stacking the columns of the matrix \(M\) on top of one another. For any positive integer \(p\), \(\mathbb{I}_p\) denotes the \(p \times p\) identity matrix. For any distribution $P$ and real number ${\tau \in (0,1)}$, $q_P(\tau)$ denotes the quantile at order $\tau$ of the distribution. Since we remain in an independent and identically distributed (i.i.d.) setup throughout the article, we sometimes drop the subscript \(i\) of random variables to lighten notations. In other words, if \(D_1, \ldots, D_n\) are \(n\) i.i.d. random variables, \(D\) without subscript denotes a generic random variable with the same distribution. Moreover, for a statistical model $(P_\theta)_{\theta \in \Theta}$ we denote by \({\mathbb{P}_{\theta}}\) (respectively \({\mathbb{E}_{\theta}}\), ${\mathbb{V}_{\theta}}$, and so on) the probability (respectively expectation, variance, etc.) with respect to the joint distribution \(P_\theta^{\otimes n}\).

Quality measures for confidence sets

We start by formally defining several attributes of confidence sets that characterize their quality. We do so in a general framework that encompasses linear models. Let $(\mathcal{X}, \mathcal{B})$ be a measurable space, and assume that we observe ${n \in \mathbb{N}^{*}}$ i.i.d. $\mathcal{X}$-valued random elements $(\xi_i)_{i=1}^n$ from a common probability space $(\Omega, \mathcal{A}, P_\theta)$, where $P_\theta$ is a distribution from a given statistical model $(P_\theta)_{\theta \in \Theta}$.

remIn most applications $\mathcal{X} = \mathbb{R}^d$ for some $d \in \mathbb{N}^{*}$ and in many cases, the parameter space $\Theta$ can be written as a product space $\Theta = \mathbb{R}^p \times \Theta_2$, where $\Theta_2$ is a topological space and \({p \in \mathbb{N}^{*}}\). This arises, for instance, in semi-parametric models, which are ubiquitous in econometrics. A model is said to be semi-parametric when the distribution of the data is characterized by both a finite-dimensional parameter $\theta_1$, which is the one of interest, and an infinite-dimensional “nuisance” parameter $\theta_2$, which may belong, for example, to some set of probability distributions. In linear regression models, $\Theta_2$ is a non-parametric set of joint distributions for regressors and error terms. In such a case, the target $T(\theta_1, \theta_2)$ usually only depends on $\theta_1$, for example, if we want to estimate one of the coefficients of a linear regression.

We seek to construct a confidence set for a real-valued function $T(\cdot)$ of the parameter $\theta$. Loosely speaking, a confidence set is a subset of $\mathbb{R}$ computable from the data that aims to contain \(T(\theta)\) with prescribed probability \({1 - \alpha}\), known as its confidence or nominal level, for some user-chosen \({\alpha \in (0,1)}\). Formally, we can define a method $\textnormal{CS}(\, \cdot \,)$ of constructing confidence sets in several equivalent ways:

enumerate• for every $n \in \mathbb{N}^{*}$, $\textnormal{CS}(\, \cdot \,)$ is a function from $(0, 1) \times \mathcal{X}^n$ to the set of Borel subsets $\mathcal{B}(\mathbb{R})$; • $\textnormal{CS}(\, \cdot \,)$ is a function from $(0, 1) \times \bigsqcup_{n = 1}^{+\infty} \mathcal{X}^n$ to $\mathcal{B}(\mathbb{R})$; • $\textnormal{CS}(\, \cdot \,)$ is a function from $(0, 1) \times \mathbb{N}^{*} \times \mathcal{X}^{\mathbb{N}^{*}}$ to $\mathcal{B}(\mathbb{R})$ where $\textnormal{CS}(1 - \alpha, n, (\xi_i)_{i \in \mathbb{N}^{*}})$ only depends on the first $n$-th entries of $(\xi_i)_{i \in \mathbb{N}^{*}}$;

such that $\{\omega \in \Omega: \, T(\theta) \in \textnormal{CS}(1 - \alpha, (\xi_i(\omega))_{i=1}^n)\}$ is measurable for every $\alpha \in (0, 1)$, \(n \in \mathbb{N}^{*}\), and $\theta \in \Theta$. This is the minimal requirement to be able to define the coverage probabilities ${\mathbb{P}_{\theta}}\big(\text{CS}(1 - \alpha, (\xi_i)_{i=1}^n) \ni T(\theta) \big)$. For a given nominal level, we then write a confidence set as $\textnormal{CS}(1 - \alpha, (\xi_i)_{i=1}^n)$, simplified to $\textnormal{CS}(1-\alpha,n)$. For brevity, we also use the notation $\textnormal{CS}(1-\alpha,n)$ to denote the sequence of confidence sets \((\textnormal{CS}(1 - \alpha, (\xi_i)_{i=1}^n))_{n \in \mathbb{N}^{*}}\). Confidence intervals (CIs) can be seen as a particular case of confidence sets (CSs) that are intervals.

Several criteria exist to assess the quality of $\textnormal{CS}(1-\alpha,n)$ for a given $\alpha$. $\textnormal{CS}(1-\alpha,n)$ is said to be asymptotically valid pointwise over $\Theta$ at level $1 - \alpha$ if

equation[equation omitted — 202 chars of source]

$\textnormal{CS}(1-\alpha,n)$ is said to be asymptotically exact pointwise over $\Theta$ at level $1 - \alpha$ if

equation[equation omitted — 194 chars of source]

A stronger asymptotic criterion exists: $\textnormal{CS}(1-\alpha,n)$ is said to be asymptotically valid uniformly over $\Theta$ at level $1 - \alpha$ if

equation[equation omitted — 194 chars of source]

and $\textnormal{CS}(1-\alpha,n)$ is said to be asymptotically exact uniformly over $\Theta$ at level $1 - \alpha$ if

equation[equation omitted — 207 chars of source]

We say that $\textnormal{CS}(1-\alpha,n)$ is non-asymptotically valid over $\Theta$ at level $1 - \alpha$ if

equation[equation omitted — 215 chars of source]

This property evolves into non-asymptotic exactness over $\Theta$ at level $1 - \alpha$ if the following stronger condition holds

equation[equation omitted — 209 chars of source]

For any of those definitions, when the level $1 - \alpha$ is not specified, it means that the method of constructing confidence sets satisfies the corresponding property for all \(\alpha \in (0,1)\). For instance, we say that $\textnormal{CS}(\, \cdot \,)$ is asymptotically exact pointwise over $\Theta$ if \(\textnormal{CS}(1-\alpha,n)\) is so at all levels \(1 - \alpha\).

These definitions follow usual conventions and enable us to compare the quality of competing CSs along several dimensions. Note that these concepts are not exclusive: Figure (ref) summarizes the implications between all those quality measures. For instance, exactness always implies validity. On the other hand, CSs that are valid but not exact are called conservative since their coverage probability is strictly larger than the targeted nominal level.

figure[figure omitted — 571 chars of source]

Unlike non-asymptotic properties, asymptotic ones can be further characterized by their pointwise or uniform nature. Asymptotic uniform validity relates to the notion of “honesty.” Compared to pointwise guarantees, it can be argued as more reliable regarding CSs' finite-sample performance (ArmstrongKolesarQE2020SimpleHonnestCI).

Property (ref) (pointwise AE) is what applied econometricians implicitly have in mind when they rely on the asymptotic normality of an estimator to conduct inference. This property is usually achievable over a non-parametric set $\Theta$. In models where (ref) holds, it is often possible to define a large subset $\widetilde{\Theta} \subset \Theta$ on which (ref) (uniform AE) is verified; see e.g. kasy2019. Regarding finite-sample inference, Property (ref) (the NAV part of NAVAE) has been shown to hold in many models. These results are predominantly found in the mathematical statistics literature (see brunel2019nonasymptotic for a recent illustration). These are powerful findings. Yet, the resulting CSs are usually asymptotically conservative, uniformly and even pointwise; in other words, they are not AE, hence not NAVAE confident sets. Finally, the strongest notion (ref), namely non-asymptotic exactness (which, a fortiori, implies the NAVAE property uniformly over the statistical model), can be obtained at the cost of placing fairly strong restrictions on $P_{\theta}$, for instance by imposing that $P_{\theta}$ belong to a parametric family. Overall, it appears challenging to build NAVAE confidence sets uniformly over a non-parametric set of distributions. We continue this discussion in Appendix (ref), focusing on inference on an expectation.

The implications that are not displayed in Figure (ref) are not satisfied. This is straightforward to see since there exist confidence sets that are, for example, asymptotically valid pointwise but not non-asymptotically valid. Nevertheless, if we are given confidence sets that are asymptotically valid or exact uniformly over some parameter set $\Theta$, it is always possible to construct confidence sets that are non-asymptotically valid over $\Theta$. This result is presented in the following proposition.

prop[Obtaining non-asymptotic validity from asymptotic uniform validity or exactnessn] \begin{enumerate} • If there exists an asymptotically valid confidence set \(\textnormal{CS}(1-\alpha,n)\) uniformly over $\Theta$ at level \(1 - \alpha \in (0,1)\), then, for every $\widetilde\alpha > \alpha$, there exists a non-asymptotically valid confidence set \(\widetilde{\textnormal{CS}}({1-\widetilde\alpha, n})\) over $\Theta$ at level \(1-\widetilde{\alpha}\). Furthermore, if $\textnormal{CS}({1-\alpha, n})$ is almost surely different from \(T(\Theta)\) for $n$ large enough, so is $\widetilde{\textnormal{CS}}({1-\widetilde\alpha, n})$. • If there exists a method $\textnormal{CS}(\, \cdot \,)$ of constructing confidence sets that is asymptotically exact uniformly over $\Theta$, then, there exists a method \(\widetilde{\textnormal{CS}}(\, \cdot \,)\) of constructing confidence sets that is both non-asymptotically valid and asymptotically exact (NAVAE) uniformly over $\Theta$. \end{enumerate}
proof[Proof of Proposition (ref)] (i) Let $\widetilde\alpha > \alpha$ and $f(n) := \inf_{\theta\in\Theta} {\mathbb{P}_{\theta}} \big(\textnormal{CS}(1-\alpha,n)\ni T(\theta) \big)$, for any positive integer \(n\). By definition, $\liminf_{n\to+\infty} f(n) \geq 1-\alpha$. As \(1 - \widetilde{\alpha} < 1 - \alpha\), there exist $n^* \in \mathbb{N}^{*}$ such that, for every $n \geq n^*$, $f(n) \geq 1 - \widetilde\alpha$. Then, we define $\widetilde{\textnormal{CS}} (1-\widetilde\alpha, n) := \textnormal{CS}(1-\alpha, n)$ for $n \geq n^*$ and $\widetilde{\textnormal{CS}}(1-\widetilde\alpha, n) := T(\Theta)$ for $n < n^*$. Note that $\widetilde{\textnormal{CS}}(1-\widetilde\alpha, n)$ is indeed non-asymptotically valid over \(\Theta\) at level \(1 - \widetilde{\alpha}\). The second part of (i) is a direct consequence of the construction of $\widetilde{\textnormal{CS}}(1-\widetilde\alpha, n)$. (ii) Let $f(\alpha, n) := \sup_{\theta\in\Theta} \left| {\mathbb{P}_{\theta}}\big(\textnormal{CS}(1-\alpha,n)\ni T(\theta) \big) - (1-\alpha) \right|$ for any \(n \in \mathbb{N}^{*}\) and \(\alpha \in (0,1)\). By definition, $\lim_{n\to+\infty} f(\alpha, n) = 0$ for every fixed $\alpha \in (0, 1)$. Fix \(\alpha \in (0,1)\). For every integer $k \geq 2$, let $\alpha_k := \alpha \times (1 - 1/k)$ and $n_k$ such that, for all $n \geq n_k$, $f(\alpha_k, n) \leq \alpha / k$. Then, we construct an increasing integer sequence \((n'_k)_{k \geq 2}\) by $n'_2 := n_2$ and $n'_k := 1 + \max(n'_{k-1}, n_k)$ recursively. Thus, we can define $\widetilde{\textnormal{CS}}(1-\alpha, n) := \textnormal{CS}(1-\alpha_k, n)$ for $n$ such that $n'_k \leq n < n'_{k+1}$ and $\widetilde{\textnormal{CS}}(1-\alpha, n) := T(\Theta)$ for $n$ such that $n < n_2$. We now check the properties of \(\widetilde{\textnormal{CS}}(1-\alpha, n)\). First, for \(n\) large enough, \begin{align*} \sup_{\theta\in\Theta} \Big| {\mathbb{P}_{\theta}}\big(\widetilde{CS}(1-\alpha, n) \ni T(\theta) \big) &- (1-\alpha) \Big| = \sup_{\theta\in\Theta} \left| {\mathbb{P}_{\theta}}\big(CS(1-\alpha_k, n) \ni T(\theta) \big) - (1-\alpha) \right| \\ &\leq \sup_{\theta\in\Theta} \left| {\mathbb{P}_{\theta}}\big(CS(1-\alpha_k, n) \ni T(\theta) \big) - (1-\alpha_k) \right| + \alpha / k \\ &= f(\alpha_k, n) + \alpha / k \leq 2 \alpha / k, \end{align*} so $\widetilde{\textnormal{CS}}(1-\alpha, n)$ is asymptotically exact uniformly over $\Theta$ because \(k \to +\infty\) when \(n \to +\infty\). Second, $\inf_{\theta\in\Theta} {\mathbb{P}_{\theta}} \big(\widetilde{\textnormal{CS}}(1-\alpha, n) \ni T(\theta) \big) \geq 1 - \alpha_k - \alpha / k = 1 - \alpha$, and $\widetilde{\textnormal{CS}}(1-\alpha, n)$ is indeed non-asymptotically valid over $\Theta$. Those properties hold for every \(\alpha \in 0,1)\), yielding the desired NAVAE property for the method \(\widetilde{\textnormal{CS}}(\,\cdot\,)\) of constructing confidence sets.
remInspecting the previous proof shows that a sufficient condition for the existence of a NAVAE confidence set uniformly over $\Theta$ at a given level $1-\alpha$ is the existence of a sequence $(\alpha_k)_{k \geq 1} \in (0, 1)$ such that (1) $\alpha_k \to \alpha$ as $k \to \infty$; (2) $\forall k \in \mathbb{N}^{*}, \, \alpha_k < \alpha$; (3) an asymptotically exact confidence set $\textnormal{CS}(1-\alpha_k, n)$ uniformly over $\Theta$ exists at each level $1 - \alpha_k$, \(k \in \mathbb{N}^{*}\).

As a direct consequence of Proposition (ref)(i), we can state the subsequent corollary.

corIf there exists a method $\textnormal{CS}(\, \cdot \,)$ of constructing confidence sets that is asymptotically valid uniformly over $\Theta$, then, there exists a method \(\widetilde{\textnormal{CS}}(\, \cdot \,)\) of constructing confidence sets that is non-asymptotically valid over $\Theta$.

Note that the impossibility claim made on Bernoulli distributions in Proposition (ref) can actually be strengthened with the help of Proposition (ref): not only NAVAE confidence sets do not exist over $\mathfrak{Ber}$, but even finding a method of constructing confidence sets that is asymptotically exact uniformly over $\mathfrak{Ber}$ is impossible.

corFor any method \(\textnormal{CS}(\, \cdot \,)\) of constructing confidence sets, the set of real numbers \begin{equation*} \big\{ \alpha \in (0, 1): \, (CS(1-\alpha,n))_{n \geq 1} is asymptotically exact uniformly over \mathfrak{Ber} at level 1 - \alpha \big\} \end{equation*} has an empty interior. Consequently, there is no method of constructing confidence sets that is asymptotically exact uniformly over \(\mathfrak{Ber}\).
proofIf the interior of this set was not empty, we could find an interval $[\alpha_-, \alpha_+]$ included inside. Therefore, by Remark (ref), we would obtain the existence of a NAVAE confidence set uniformly over \(\mathfrak{Ber}\) at nominal level \(1 - \alpha_+\). But this is not possible by Proposition (ref).

We conjecture that the set mentioned in Corollary (ref) is empty: for any $\alpha \in (0, 1)$, there is no confidence set that is asymptotically exact uniformly over $\mathfrak{Ber}$ at level \(1 - \alpha\). Extending Corollary (ref) in this direction would be non-trivial and is left for future research.

The “automatic” construction of a NAVAE confidence set $\widetilde{\textnormal{CS}}(1-\alpha, n)$ as given in the proof of Proposition (ref)(ii) is mainly interesting from a theoretical point-of-view since the approach is difficult to apply in practice as such (the quantities \(f(\alpha_k,n)\) are not available in general). Still, if an upper bound \(g(\cdot, \cdot)\) on \(f(\cdot, \cdot)\) is available, then one can construct the suitably enlarged confidence sets introduced in the proof of Proposition (ref)(ii), simply using \(g\) instead of \(f\). The set constructed from $g$ remains non-asymptotically valid, and it is asymptotically exact uniformly over $\Theta$ if and only if $\lim_{n \to +\infty} g(\alpha, n) = 0$ for all $\alpha$. Inspired by this construction, we use bounds on the distance to the asymptotic normality for sample means/OLS estimators in the rest of the paper to suitably enlarge CLT-based confidence intervals.

Among NAVAE CIs, it is sometimes possible to ask for more: in particular, we could aim for efficient intervals \`a la romano2000 (see their Theorem 2.1(i) and Remark 3). We briefly summarize their results using our notations. A non-asymptotically valid confidence interval $[U_n, L_n]$ is said to be efficient if its width is asymptotically minimal among all non-asymptotically valid confidence intervals. Under some regularity conditions, any efficient confidence interval of a parameter $T(\theta)$ estimated by some suitable $\widehat{T(\theta)}$ must be of the form

equation[equation omitted — 184 chars of source]

where $\tau(\theta)^2$ is the corresponding asymptotic variance. In other words, an efficient CI must be an asymptotically negligible enlargement of a CLT-based confidence interval. As shown in the next two sections, all the CIs we build are exactly of the form (ref), which implies they are efficient in addition to being NAVAE.

NAVAE confidence intervals for expectations

In this section, we are interested in presenting a simple case: conducting inference on the expectation of a real random variable that admits (at least) a non-zero finite variance, that is, with the formalization of the previous section,

equation*[equation* omitted — 189 chars of source]

where the parameter is the pair combining an expectation and the distribution of the corresponding centered variable (which is a nuisance parameter). In this way, we can decompose any random variable $\xi$ following $P_\theta$ where $\theta = (\theta_1, P)$ by $\xi := \theta_1 + V$ where $V \sim P$. Thus, with our formalization, $p=1$ and $T(\theta_1, P) = \theta_1$. Note that, in this sense, $\Theta$ is in bijection with the set of univariate distributions with finite but non-zero variance, and we will use this identification in the following. We assume the variance to be non-zero to remove pathological cases.

In this canonical scenario, we illustrate the challenges of building NAVAE confidence intervals that “converge” asymptotically to standard CIs based on the CLT. Henceforth, we fix a desired nominal level \(1 - \alpha \in (0,1)\).

Setting and motivation

Let $ \widehat{\sigma}^2:= n^{-1} \sum_{i=1}^n(\xi_i-\overline{\xi}_n)^2$, with $\overline{\xi}_n := n^{-1} \sum_{i=1}^n \xi_i$. We know by the Central Limit Theorem (CLT) and Slutsky's lemma that

equation[equation omitted — 241 chars of source]

is asymptotically exact pointwise over $\Theta$ at level \(1 - \alpha\). Among practitioners, it is the most common way of conducting inference on an expectation and, therefore, a natural candidate on which to base a uniformly NAVAE confidence interval.

Our goal is to propose an asymptotically negligible modification of \({\mathrm{CI}_{\mathrm{CLT}}(1-\alpha,n)}\) that is nonetheless sufficient to make the resulting confidence interval NAVAE over a large (i.e., non-parametric) subset of $\Theta$. As suggested in Section (ref), an appealing idea consists in controlling non-asymptotically the distance between the unknown distribution of the sample mean and its Gaussian limiting distribution. To do so, we leverage Berry-Esseen inequalities (berry1941, esseen1942) or Edgeworth expansions as these tools “quantify” the CLT (by controlling, for any sample size, some distance between the distribution of a normalized sum and the standard Gaussian distribution \(\mathcal{N}(0, 1)\)). It is key to note that placing additional restrictions on $\Theta$ to reach our goal is not superfluous: indeed, we have $\mathfrak{Ber} \subset \Theta$ so that no method \(\textnormal{CS}(\, \cdot \,)\) of constructing confidence sets can be NAVAE uniformly over $\Theta$.

We now introduce some additional notation to develop these ideas. Let $\sigma(\theta) := \sqrt{{\mathbb{V}_{\theta}}(\xi)}$, ${\lambda_3(\theta) := {\mathbb{E}_{\theta}}\big[(\xi-{\mathbb{E}_{\theta}}[\xi])^3/\sigma(\theta)^3\big]}$, and $K_4(\theta) := {\mathbb{E}_{\theta}}\big[ (\xi-{\mathbb{E}_{\theta}}[\xi])^4/\sigma(\theta)^4\big]$ for any parameter \(\theta \in \Theta\). Finally, $S_n$ denotes the standardized sum ${\sum_{i=1}^n (\xi_i - {\mathbb{E}_{\theta}}[\xi]) \, / \, (\sigma(\theta) \sqrt{n})}$. Berry-Esseen (BE) inequalities control the uniform distance between the cumulative distribution function (c.d.f) of the empirical mean (properly centered and standardized) and the c.d.f \(\Phi\) of its limit \(\mathcal{N}(0,1)\) distribution (below \(\varphi\) denotes the p.d.f of a standard Normal distribution),

equation[equation omitted — 169 chars of source]

Edgeworth expansions (EE) are refinements adjusting for the presence of non-asymptotic skewness. They control the following uniform distance

equation[equation omitted — 232 chars of source]

What kind of additional constraints can we place on the parameter set to derive an explicit bound $\delta_n$ on ${\Delta_{n,\mathrm{B}}}$ or ${\Delta_{n,\mathrm{E}}}$? Existing results require a control on some moments of the distribution, beyond moments of order two\footnote{ Recall that restricting to subsets $\widetilde{\Theta}$ of $\Theta$ that impose more than $\mathbb{E}_\theta[\xi^2]<+\infty$ is necessary to build NAVAE CSs that are uniform over $\widetilde{\Theta}$. If one is willing to focus on non-parametric subsets $\widetilde{\Theta}$ of $\Theta$, the following is shown in romano2004: ${\mathrm{CI}_{\mathrm{CLT}}(1-\alpha,n)}$ is uniformly asymptotically exact over $\widetilde{\Theta}$ if and only if $\lim_{\lambda\to\infty}\sup_{\theta\in\widetilde{\Theta}}{\mathbb{E}_{\theta}}\!\left[\frac{(\xi - {\mathbb{E}_{\theta}}[\xi])^2}{\sigma(\theta)^2}\mathds{1}\!\left\{\frac{|\xi - {\mathbb{E}_{\theta}}[\xi]|}{\sigma(\theta)} > \lambda\right\}\right] = 0$ \((\ast)\). Given Proposition (ref), it is thus possible to derive NAVAE CSs on such parameter sets. This explains why NAVAE CSs can be constructed on $\Theta_{\mathrm{EE}}$, since the previous condition \((\ast)\) is satisfied for \(\widetilde{\Theta} = \Theta_{\mathrm{EE}}\). }: the standardized $(2+\mu)$-th absolute moment for BE, the fourth one for EE. For the sake of brevity, we only consider the smaller class of distributions \(\Theta_{\mathrm{EE}}:= \{ \theta \in \Theta : K_4(\theta) \leq K \}\) for some known positive constant $K$, which allows to leverage the strength of BE and EE bounds at the same time. Note that EE-based inequalities are more complex but yield (asymptotically) tighter bounds than BE; see Remark (ref) below and derumigny2022explicitBE for a detailed comparison between these inequalities.

Equipped with a bound $\delta_n$ on ${\Delta_{n,\mathrm{B}}}$ or ${\Delta_{n,\mathrm{E}}}$, we can now explain how to construct NAVAE confidence intervals for \(\theta_1 = {\mathbb{E}_{\theta}}[\xi]\). It is enlightening to start with the simplified case of a distribution with known variance to give the main intuitions behind our construction.

Known variance

To simplify exposition, let us assume that $\delta_n$ is a bound on ${\Delta_{n,\mathrm{B}}}$ for a moment (we show below that the same result holds if \(\delta_n\) is a bound on \({\Delta_{n,\mathrm{E}}}\) instead). Using the BE inequality and simple computations, we obtain the following: for any positive real number \(x\),

equation[equation omitted — 279 chars of source]

Then, for any given \(\alpha \in (0,1)\), setting \(x\) such that the right-hand side of Equation (ref) is equal to \(\alpha\) and considering the complementary event give

equation[equation omitted — 276 chars of source]

whenever this “modified Gaussian quantile” is well-defined, namely when $1 - \alpha/2 + \delta_n < 1$. When that condition is not met, we can still claim that $\theta_1 \in \mathbb{R}$ with probability one (therefore at least $1-\alpha$) by definition of $\Theta$. Consequently, in the case of a known variance \(\sigma_{\textnormal{known}}^2\), the confidence interval

equation*[equation* omitted — 366 chars of source]

is non-asymptotically valid over $\Theta_{\mathrm{EE}}^{\textnormal{known}} := \{ \theta \in \Theta_{\mathrm{EE}} : {\mathbb{V}_{\theta}}(\xi) = \sigma_{\textnormal{known}}^2 \}$.

The first inequality in Equation (ref) also holds when replacing ${\Delta_{n,\mathrm{B}}}$ by ${\Delta_{n,\mathrm{E}}}$ since the term $\frac{\lambda_{3}(\theta)}{6\sqrt{n}}(1-x^2)\varphi(x)$ is symmetric and vanishes. Therefore, either the EE or the BE construction yields a non-asymptotically valid CI: a known bound \(\delta_n\) on the minimum of \({\Delta_{n,\mathrm{B}}}\) and \({\Delta_{n,\mathrm{E}}}\) is enough in the definition of \({\mathrm{CI}_{\mathrm{\sigma_{\textnormal{known}}}}(1-\alpha,n)}\). Proposition (ref) summarizes this discussion (see Appendix (ref) for its proof).

propLet \(n \geq 1\) and $\delta_n \geq \sup_{\theta \in \Theta_{\mathrm{EE}}} \min\{ {\Delta_{n,\mathrm{E}}}(\theta) , {\Delta_{n,\mathrm{B}}}(\theta) \}$. Then, \newline ${\mathrm{CI}_{\mathrm{\sigma_{\textnormal{known}}}}(1-\alpha,n)}$ is non-asymptotically valid over $\Theta_{\mathrm{EE}}^{\textnormal{known}}$ at level \(1 - \alpha\).

In Remark (ref), we show that there exist bounds \(\delta_n\) decreasing to \(0\) when \(n\) goes to infinity for our class $\Theta_{\mathrm{EE}}$. Hence, for any \(\alpha \in (0,1)\), for \(n\) large enough, there is enough information in the sample for our confidence interval \({\mathrm{CI}_{\mathrm{\sigma_{\textnormal{known}}}}(1-\alpha,n)}\) to be informative in the sense of being strictly included in the whole real line \(\mathbb{R}\). Since $\delta_n$ is deterministic and converges to zero, we remark that, for any $\theta$ in $\Theta^{\textnormal{known}} := \{ \theta \in \Theta : {\mathbb{V}_{\theta}}(\xi) = \sigma_{\textnormal{known}}^2 \}$,

equation[equation omitted — 307 chars of source]

Therefore, \({\mathrm{CI}_{\mathrm{\sigma_{\textnormal{known}}}}(1-\alpha,n)}\) behaves like a confidence interval based on the asymptotic normality of the sample mean and is thus pointwise asymptotically exact over $\Theta^{\textnormal{known}}$. In fact, our CI can be shown to be uniformly asymptotically exact over $\Theta_{\mathrm{EE}}^{\textnormal{known}}$, as stated in Proposition (ref) and proved in Appendix (ref).

propLet $(\delta_n)$ be a sequence such that ${\delta_n \geq \sup_{\theta \in \Theta_{\mathrm{EE}}} \min\{ {\Delta_{n,\mathrm{E}}}(\theta) , {\Delta_{n,\mathrm{B}}}(\theta) \}}$ and $\delta_n \to 0$. Then \begin{itemize} • ${\mathrm{CI}_{\mathrm{\sigma_{\textnormal{known}}}}(1-\alpha,n)}$ is asymptotically exact pointwise over $\Theta^{\textnormal{known}}$ at level \(1 - \alpha\); • ${\mathrm{CI}_{\mathrm{\sigma_{\textnormal{known}}}}(1-\alpha,n)}$ is asymptotically exact uniformly over \(\Theta_{\mathrm{EE}}^{\textnormal{known}}\) at level \(1 - \alpha\). \end{itemize}

Propositions (ref) and (ref)(ii) ensure ${\mathrm{CI}_{\mathrm{\sigma_{\textnormal{known}}}}(1-\alpha,n)}$ is NAVAE uniformly over $\Theta_{\mathrm{EE}}^{\textnormal{known}}$. Combined with Equation (ref), this result implies that our confidence interval is of the form (ref) and therefore efficient.

rem[Possible choices for $\delta_n$] Proposition (ref) and subsequent results require a known bound \(\delta_n\), on either \({\Delta_{n,\mathrm{B}}}(\theta)\) or \({\Delta_{n,\mathrm{E}}}(\theta)\). Furthermore, the smaller the bound \(\delta_n\), the shorter the resulting CI. Relying on shevtsova2013 and the inequality \(\mathbb{E}[|\xi|^3] \leq (\mathbb{E}[\xi^4])^{3/4}\), we deduce that \(\delta_{1,n} := 0.4690 K^{3/4} / \sqrt{n}\) is a valid bound on \({\Delta_{n,\mathrm{B}}}(\theta)\) over \(\Theta_{\mathrm{EE}}\). derumigny2022explicitBE provides \(\delta_{2,n} := 0.1995 (K^{3/4} + 1) / \sqrt{n} + O(n^{-1}) \) as a bound on \({\Delta_{n,\mathrm{E}}}(\theta)\) over \(\Theta_{\mathrm{EE}}\), with an explicit remainder term implemented in the R package BoundEdgeworth (BoundEdgeworth). Consequently, Proposition (ref) holds with \(\delta_n := \min(\delta_{1,n}, \delta_{2,n})\). Although \(\delta_{1,n}\) and \(\delta_{2,n}\) decrease to \(0\) as \(n\) goes to infinity at the same rate, \(\delta_{2,n}\) is asymptotically smaller than \(\delta_{1,n}\). For \(n\) large enough, our confidence intervals thus rely on an Edgeworth expansion. Under additional restrictions (that rule out discrete distributions), derumigny2022explicitBE show that the quantity \(\delta_{3,n} := (0.195 K + 0.01465 K^{3/2}) / n + O(n^{-5/4})\) upper bounds \({\Delta_{n,\mathrm{E}}}(\theta)\), allowing to build CIs that are one order of magnitude shorter than Berry-Esseen-based ones (using the previously mentioned package). See Remark (ref) for discussion about the choice of $K$ in practice.

Unknown variance

When the variance is unknown, a construction in the spirit of ${\mathrm{CI}_{\mathrm{\sigma_{\textnormal{known}}}}(1-\alpha,n)}$ remains possible at the cost of controlling the variance estimation error. The process we use is twofold. First, the ratio between the “oracle” variance estimator \(\widehat{\sigma}^2_0 := n^{-1} \sum_{i=1}^n (\xi_i - \theta_1)^2\) and its limit $\sigma^2(\theta)$ is tightly controlled from below using Theorem 2.19 in pena2008self. That theorem allows us to write, for every $a > 1$ and $\theta \in \Theta_{\mathrm{EE}}$,

equation*[equation* omitted — 215 chars of source]

From now on, we introduce a sequence \((a_n)\) of tuning parameters. Indeed, to build CIs with good asymptotic properties, we need $a_n$ to depend on $n$ as specified in Proposition (ref). Second, we control \(\sqrt{n}(\overline{\xi}_n - \theta_1)/\sigma(\theta)\) similarly to what was done when building ${\mathrm{CI}_{\mathrm{\sigma_{\textnormal{known}}}}(1-\alpha,n)}$.

Combining the two steps, we obtain the following confidence interval \({\mathrm{CI}_{\mathrm{ \widehat{\sigma}}}(1-\alpha,n)}:=\)

equation*[equation* omitted — 414 chars of source]

where

equation*[equation* omitted — 201 chars of source]

Besides the data \((\xi_i)_{i=1}^n\) and the level \(1 - \alpha\), \({\mathrm{CI}_{\mathrm{ \widehat{\sigma}}}(1-\alpha,n)}\) depends on three quantities: \(\delta_n\), \(K\), and the tuning parameter \(a_n\). \({\mathrm{CI}_{\mathrm{ \widehat{\sigma}}}(1-\alpha,n)}\) can be shown to be non-asymptotically valid over \(\Theta_{\mathrm{EE}}\). Proposition (ref) formalizes our finite-sample results on \({\mathrm{CI}_{\mathrm{ \widehat{\sigma}}}(1-\alpha,n)}\). Its proof can be found in Appendix (ref).

propLet \(n \geq 1\), $a_n \in (1,+\infty)$, and $\delta_n \geq \sup_{\theta \in \Theta_{\mathrm{EE}}} \min\{ {\Delta_{n,\mathrm{E}}}(\theta) , {\Delta_{n,\mathrm{B}}}(\theta) \}$. Then, ${\mathrm{CI}_{\mathrm{ \widehat{\sigma}}}(1-\alpha,n)}$ is non-asymptotically valid over \( \Theta_{\mathrm{EE}}\) at level \(1 - \alpha\).

In line with ${\mathrm{CI}_{\mathrm{\sigma_{\textnormal{known}}}}(1-\alpha,n)}$, ${\mathrm{CI}_{\mathrm{ \widehat{\sigma}}}(1-\alpha,n)}$ can be shown to be asymptotically close to ${\mathrm{CI}_{\mathrm{CLT}}(1-\alpha,n)}$. In fact, if $a_n$ is such that $a_n \to 1$ and $n(1-1/a_n)^2 \to +\infty$, we can see that for any positive value of \(K\) and for any \(\theta \in \Theta\), \(\nu_n^{\textrm{Var}} \to 0\) as \(n \to +\infty\). This implies

equation[equation omitted — 356 chars of source]

where, in this representation, the value \(K\) is hidden in the \(o_{P_\theta} \big(1/\sqrt{n} \big)\) term. Consequently, as is the case for \({\mathrm{CI}_{\mathrm{CLT}}(1-\alpha,n)}\), \({\mathrm{CI}_{\mathrm{ \widehat{\sigma}}}(1-\alpha,n)}\) is asymptotically exact pointwise over the whole parameter space \(\Theta\). As stated in the next proposition, ${\mathrm{CI}_{\mathrm{ \widehat{\sigma}}}(1-\alpha,n)}$ is also asymptotically exact uniformly over $\Theta_{\mathrm{EE}}$. The proof can be found in Section (ref).

propLet $(\delta_n)$ be a sequence such that ${\delta_n \geq \sup_{\theta \in \Theta_{\mathrm{EE}}} \min\{ {\Delta_{n,\mathrm{E}}}(\theta) , {\Delta_{n,\mathrm{B}}}(\theta) \}}$ and $\delta_n \to 0$. Assume that $b_n := a_n - 1$ is such that $b_n \to 0$ and \(b_n \sqrt{n} \to +\infty\). Then \begin{itemize} • ${\mathrm{CI}_{\mathrm{ \widehat{\sigma}}}(1-\alpha,n)}$ is asymptotically exact pointwise over $\Theta$ at level \(1 - \alpha\); • ${\mathrm{CI}_{\mathrm{ \widehat{\sigma}}}(1-\alpha,n)}$ is asymptotically exact uniformly over \(\Theta_{\mathrm{EE}}\) at level \(1 - \alpha\). \end{itemize}

Propositions (ref), (ref)(ii), and Equation (ref) ensure ${\mathrm{CI}_{\mathrm{ \widehat{\sigma}}}(1-\alpha,n)}$ is NAVAE uniformly over $\Theta_{\mathrm{EE}}$ and efficient since it is of the form (ref).

Similar to \({\mathrm{CI}_{\mathrm{\sigma_{\textnormal{known}}}}(1-\alpha,n)}\), our interval might be non-informative, namely equal to \(\mathbb{R}\). The tuning parameter \(a_n\) must satisfy the constraint

equation[equation omitted — 139 chars of source]

for the interval to be informative. For $n$ small, there might not be any value $a_n \in ( 1, +\infty)$ for which the constraint holds. However, it will be the case for \(n\) large enough as shown in Proposition (ref) (proved in Appendix (ref)). Note that as soon as there is one solution, we know that the set of possible $a_n$ forms an open interval.

propLet $\alpha \in (0, 1/2)$, $K > 0$. \begin{enumerate} • Let $n, \delta_n > 0$. The subset of values of $a_n$ of $(1, +\infty)$ satisfying the constraint (ref) is an open interval $I_n$ (potentially empty). If $I_n$ is not empty, then it must be of the form $(a_{1,n}, a_{2,n})$, with $1 < a_{1,n}$ and $a_{2,n} < +\infty$. If moreover $(\delta_n)$ is a decreasing sequence, then the sequence $I_n$ is increasing (with respect to the inclusion order). Furthermore $(a_{1,n})$ and $(a_{2,n})$ are respectively decreasing and increasing sequences, with respective limits $1$ and $+\infty$. • Let $a_n = 1 + b_n \in (1, +\infty)$ with $b_n \to 0$ and \(b_n \sqrt{n} \to +\infty\). Assume that for $n$ large enough, $1 - \dfrac{\alpha}{2} + \delta_n$ is bounded above by a constant strictly smaller than $1$. Then, for $n$ large enough, the constraint (ref) is satisfied. \end{enumerate}
remWhen \(I_n\) is not empty, a natural choice for \(a_n\) is to minimize the width of our confidence interval, which can be done numerically. Such an automatic selection rule for \(a_n\) is appealing as it avoids having to choose manually a value for this parameter and yields the shortest CI. Furthermore, for a fixed bound \(K\) (as opposed to plug-in, see Section (ref)), that optimal \(a_n\) is not data-dependent. Table (ref) in Section (ref) compares our optimized choice to the ad hoc alternative \(a_n = 1 + n^{-1/5}\), which satisfies the rate requirement of Proposition (ref). We find that these competing choices lead to comparable inference results in simulations.

In the simple case of an expectation, we have thus managed to construct closed-form and non-randomized NAVAE confidence intervals uniformly over a non-parametric class of distributions delineated by moment conditions only. As highlighted in the explanatory case of a known variance, the key behind that construction is to enlarge \({\mathrm{CI}_{\mathrm{CLT}}(1-\alpha,n)}\) properly through asymptotically negligible modifications based on Berry-Esseen inequalities or Edgeworth expansions. In Section (ref), we move on to linear regressions. Following the same principle, we derive confidence intervals that enjoy the same appealing theoretical properties.

NAVAE confidence intervals for linear regressions' coefficients

Model formulation

For the sake of completeness and precision, we first state the statistical model and its assumptions.

hyp[Linear model] We observe $n$ independent replications $(X_1,Y_1), \dots, $ $(X_n,Y_n)$ following the distribution $P_{X,Y},$ where $X$ is an explanatory random vector of dimension $p$ and $Y$ is an outcome real random variable such that, for some real random variable \(\varepsilon\) and vector \(\beta_{0}\): \begin{align} &Y = X' \beta_{0} + \varepsilon \quad and \quad P_{X,\varepsilon} \in \mathcal{P}_{X,\varepsilon}, where \\ &\mathcal{P}_{X,\varepsilon} := \Big\{ P_{X,\varepsilon} \in \mathcal{P}(\mathbb{R}^{p+1}) : \mathbb{E}[X\varepsilon] = 0, \, \mathbb{E}\big[\|X\|^4\big] < +\infty, \, \lambda_{\min}(\mathbb{E}[XX']) > 0, \, \nonumber \\ & \lambda_{\min}\!\left(\mathbb{E}\!\left[XX'\varepsilon^2\right]\right) > 0, \, \lambda_{\max}\!\left(\mathbb{E}\!\left[XX'\varepsilon^2\right]\right) < +\infty \Big\}. \nonumber \end{align}

The parameter set of the associated statistical model is \(\Theta := \left\{ \theta = (\beta_{0}, P_{X,\varepsilon}) \in \mathbb{R}^p \times \mathcal{P}_{X,\varepsilon} \right\} \). In this model, $P_{\theta}$ denotes a distribution of $(X, Y)$ indexed by $\theta \in \Theta$. For brevity, we implicitly omit the index $\theta$ in the expectation operator $\mathbb{E}$. In what follows, we consider several subsets of $\Theta$ characterized by additional restrictions that enable us to build several confidence sets. We will investigate and compare their properties in line with the discussion conducted in Section (ref).

Assumption (ref) sets a basic linear regression model. \(\Theta\) is indeed the largest parameter set compatible with usual economic assumptions and minimal statistical conditions. The condition \(\mathbb{E}[X\varepsilon] = 0\) corresponds to the (weak) exogeneity of covariates, i.e., the orthogonality condition of the linear projection of \(Y\) on \(X\). It is implied by the (strong) exogeneity assumption \(\mathbb{E}[\varepsilon | X] = 0\), but does not require the conditional expectation of \(Y\) given \(X\) to be linear. The other moment conditions allow for heteroskedasticity while ensuring the asymptotic normality of the Ordinary Least Squares (OLS) estimator of $\beta_{0}$:

equation*[equation* omitted — 157 chars of source]

More precisely, under Assumption (ref), a finite second-order moment, \(\mathbb{E}\big[ \|X\|^2 \big] < +\infty\), is sufficient for that result. However, a finite fourth-order moment is necessary to consistently estimate the asymptotic variance of \(\widehat{\beta}\) by its empirical counterpart and perform inference on \(\beta_{0}\) in practice. That is why we incorporate this additional requirement directly into our basic statistical model. Besides, remark that the condition \( \lambda_{\min}(\mathbb{E}(XX')) > 0 \) is equivalent to the invertibility of the matrix \( \mathbb{E}[XX'] \). Thus, under Assumption (ref), \((n^{-1} \sum_{i=1}^n X_i X_i')^{\dagger} = (n^{-1} \sum_{i=1}^n X_i X_i')^{-1}\) with probability approaching one as the sample size \(n\) goes to infinity.

For a given known vector ${u}$ of $\mathbb{R}^p$, our goal is to build a confidence interval for a linear functional of the form ${u}' \beta_{0}$. It encompasses CIs for each individual component of $\beta_{0}$ (taking for ${u}$ the canonical vectors) and also differences of coefficients that appear when investigating the relative impact of two covariates. We consider henceforth an arbitrary vector ${u} \in \mathbb{R}^p \setminus \{0_{\mathbb{R}^p}\}$.

Properties of the usual confidence interval

As mentioned in the introduction, the standard way to proceed is to construct a CI centered at the estimator ${u}' \widehat{\beta}$ relying on the asymptotic normality of the OLS estimator:

equation[equation omitted — 249 chars of source]

where

equation*[equation* omitted — 239 chars of source]

is the standard estimator of the asymptotic variance \(V := \mathbb{E}[XX']^{-1} \mathbb{E}[XX'\varepsilon^2] \mathbb{E}[XX']^{-1}\) of $\widehat{\beta}$, and \(\widehat{\varepsilon}_i := Y_i-X_i'\widehat{\beta}\) is the residual for the $i$-th observation.

The pros and cons of $\textnormal{CI}_{u}^{\textnormal{Asymp}}(1-\alpha,n)$ are well-understood and very close to those of ${\mathrm{CI}_{\mathrm{CLT}}(1-\alpha,n)}$ that were detailed in Section (ref). By applications of the Law of Large Numbers, the CLT, and Slutsky's lemma, $\textnormal{CI}_{u}^{\textnormal{Asymp}}(1-\alpha,n)$ is known to be asymptotically exact pointwise over $\Theta$. Besides, following kasy2019, $\textnormal{CI}_{u}^{\textnormal{Asymp}}(1-\alpha,n)$ can be strengthened to be asymptotically exact uniformly at level \(1 - \alpha\) over

equation*[equation* omitted — 242 chars of source]

with $0 < m \leq M < +\infty$.

Furthermore, $\textnormal{CI}_{u}^{\textnormal{Asymp}}(1-\alpha,n)$ becomes non-asymptotically exact at level \(1 - \alpha\) over

equation*[equation* omitted — 193 chars of source]

provided ${ \widehat{V}}$ is replaced with the (homoscedastic) estimator \( (n-p)^{-1} \sum_{i=1}^n \widehat{\varepsilon}_i^{\,2} \left( n^{-1} \sum_{i=1}^n X_i X_i'\right)^{\dagger}\) and the quantile ${ q_{\mathcal{N}(0,1)} }(1 - \alpha/2)$ with the quantile of a Student distribution with ${n - p}$ degrees of freedom.

However, assuming that the true distribution belongs to $\Theta_{\mathrm{Gauss}}$ -- so as to achieve such finite-sample properties -- is often considered too restrictive in practice: Gaussian error terms impede skewed or heavier-tail shocks, and independence between $\varepsilon$ and $X$ rules out heteroskedasticity. In what follows, we propose NAVAE CIs without relying on such independence or parametric assumptions. This corresponds to the same objectives as in Section (ref) now for the case of linear regressions; they will be met using the same tools (Berry-Esseen-type inequalities or bounds on Edgeworth expansions) in this different setup.

Presentation of our confidence interval

Intuitions and assumptions

Remark that

equation*[equation* omitted — 182 chars of source]

but the second term in the previous sum cannot be written as an empirical mean of i.i.d. variables, rendering it harder to analyze theoretically. Under Assumption (ref), the quantity \((n^{-1} \sum_{i=1}^n X_i X_i')^{\dagger}\) converges in probability to \(\mathbb{E}[XX']^{-1}\). Therefore, we obtain the following linearization

equation[equation omitted — 219 chars of source]

and the error of the estimator can thus be approximated by an average of $n$ i.i.d. random variables distributed as $\xi := {u}' \, \mathbb{E}[XX']^{-1} X \varepsilon$.

Like in Section (ref), our CIs cannot be non-asymptotically valid on $\Theta$ itself because this space is simply too large. Therefore, we now introduce the subset $\Theta_{\mathrm{EE}}$ of $\Theta$ described in the following Assumption (ref) in order to obtain non-asymptotic guarantees. To state our conditions, we introduce the rotated version of $X$, defined as $\widetilde{X}:= \mathbb{E}(XX')^{-1/2} X$.

hyp[Bounds on DGP] Let $\lambda_{\mathrm{reg}}$, $K_{\mathrm{reg}}$, \(K_{\varepsilon}\), and \(K_{\xi}\) be some positive constants. We define $\Theta_{\mathrm{EE}}$ to be the set of parameters $\theta = (\beta_{0}, P_{X, \varepsilon}) \in \Theta$ such that the joint distribution $P_{X, \varepsilon}$ satisfies: \begin{enumerate} • \( \lambda_{\min}(\mathbb{E}[XX']) \geq \lambda_{\mathrm{reg}} \); • \(\mathbb{E}\!\left[ \big\| {\textnormal{vec}} \big(\widetilde{X}\widetilde{X}' - \mathbb{I}_p \big) \big\|^2 \right] \leq K_{\mathrm{reg}}\); • \(\mathbb{E}\!\left[ \big\|\widetilde{X}\varepsilon\big\|^4 \right] \leq K_{\varepsilon}\); • \( \mathbb{E}\!\left[\xi^4\right] / \, \mathbb{E}\!\left[\xi^2\right]^2 \leq K_{\xi}\). \end{enumerate}

Assumption (ref) defines a broad non-parametric class of distributions delineated by the different constants \(\lambda_{\mathrm{reg}}\), \(K_{\mathrm{reg}}\), \(K_{\varepsilon}\), and \(K_{\xi}\). These constants appear explicitly in the construction of our CIs, as did the constant \(K\) in the previous section for the case of an expectation. The user needs to specify their values in practice, which should be done with care. We elaborate on these choices in Section (ref). Relying on explicit constants may seem restrictive compared to standard asymptotic inference. However, outside of some specific parametric models, using such types of bounds or other known quantities (like the median-bias of the estimator in kuchibhotla2024hulc) is unavoidable to obtain non-asymptotic properties.

Overall, the different parts of Assumption (ref) strengthen the moment conditions of the basic linear model $\Theta$. Part (i) rules out \(\mathbb{E}[XX']\) matrices arbitrarily close to being singular, an unfavorable situation in which $\beta_{0}$ is not identified. Part (ii) helps control the concentration of \( n^{-1} \sum_{i=1}^n X_i X_i' \) and ensures, together with part (i), that it is invertible with large probability for every (large enough) sample size. Part (iii) complements (ii) to control the linearization (ref) of the OLS estimator. Part (iii) is implied by (ref)(i) and the simpler (but stricter) constraint \(\mathbb{E}\!\left[ \big\|X \varepsilon\big\|^4 \right] \leq C\) for some constant $C > 0$. Part (iv) allows to bound the Edgeworth expansion of the distribution of $n^{-1} \sum_{i = 1}^n \xi_i$ and enables to derive a tight control on the distance between $n^{-1} \sum_{i = 1}^n \xi_i^2$ and its expectation, which happens to be the asymptotic variance associated with $\sqrt{n}u'(\widehat{\beta}-\beta_{0})$.

Parts (i) to (iii) of the previous assumption are critical in the construction of our CI as they enable us to ensure that with large probability

equation*[equation* omitted — 181 chars of source]

i.e., the centered and scaled OLS estimator gets “close” to a self-normalized sum of centered i.i.d. random variables on a large-probability event. More precisely, they first guarantee that the linearization

equation*[equation* omitted — 158 chars of source]

holds with probability at least \(1-2\gamma\), for any \(\gamma > 0\), where

align*[align* omitted — 205 chars of source]

is an explicit bound on the linearization error term and \(\widetilde{\gamma} := \sqrt{K_{\mathrm{reg}}/(n\gamma)}\). Second, with probability at least \(1-2\gamma\) too, they enable us to prove that

equation*[equation* omitted — 132 chars of source]

where

align*[align* omitted — 862 chars of source]

and \(S\) is a shorthand notation for $n^{-1} \sum_{i=1}^n X_i X_i'$.

Remark that \(n^{-1/2}\sum_{i=1}^n\xi_i\,\big/ \sqrt{n^{-1}\sum_{i=1}^n \xi_i^2}\) is the ratio of a sample mean (centered at the true expectation, here equal to 0) and the square-root of its corresponding oracle variance (see Section (ref) for a definition). This quantity can thus be managed using the method introduced in Section (ref), which relies on Berry-Esseen or Edgeworth-type controls and an exponential deviation inequality between $n^{-1}\sum_{i=1}^n \xi_i^2$ and its limit. This is where Assumption (ref).(iv) comes into play, with \(K_{\xi}\) playing the same role as \(K\) in the case of an expectation.

Formal definition

For any tuning parameters $\omega_n \in (0, 1)$ and $a_n \in (1, +\infty)$, we define the “modified Gaussian quantile”

equation*[equation* omitted — 157 chars of source]

which depend on some perturbation terms $\nu_n^{\mathrm{Edg}}$ and $\nu_n^{\mathrm{Approx}}$ defined as

align*[align* omitted — 297 chars of source]

where $\delta_n \geq \sup_{\theta \in \Theta_{\mathrm{EE}}} \min \! \left\{ {\Delta_{n,\mathrm{B}}}(\theta) \, , \, {\Delta_{n,\mathrm{E}}}(\theta) \right\}$. This condition is similar to the one used in Section (ref) in the framework of an expectation. Remember that \({\Delta_{n,\mathrm{B}}}(\theta)\) and \({\Delta_{n,\mathrm{E}}}(\theta)\) are defined in Equations (ref) and (ref) with \(\xi\) here equal to \({u}' \, \mathbb{E}[XX']^{-1} X \varepsilon\). Therefore, the choices of \(\delta_n\) introduced in Remark (ref) can be used in the current setting as well, replacing \(K_{\xi}\) with \(K\).

To ensure an informative CI is feasible, we impose that $n$ be larger than \\ $n_0 := \max \{n \in \mathbb{N}^{*}: n \leq 2 K_{\mathrm{reg}} / (\omega_n \alpha) \text{ or } \nu_n^{\mathrm{Edg}} \geq \alpha / 2 \}$. All in all, our confidence interval is centered at ${u}' \widehat{\beta}$ (for $n > n_0$) and defined by

align*[align* omitted — 521 chars of source]

Note that $ \textnormal{CI}_{u}^{\textnormal{Edg}}(1-\alpha,n)$ depends on two tuning parameters, \(\omega_n\) and \(a_n\), the quantity $\delta_n$, and the four bounds $K_{\xi}, K_{\mathrm{reg}}, K_{\varepsilon}, \lambda_{\mathrm{reg}}$ that delineate the set \(\Theta_{\mathrm{EE}}\) and appear in Assumption (ref). We do not indicate this dependence to lighten notations. This interval is similar to \(\textnormal{CI}_{u}^{\textnormal{Asymp}}(1-\alpha,n)\) in Equation (ref) with the addition of $\left\|{u}\right\|^2R_{n,\mathrm{var}}(\omega\alpha/2)$ in the variance term and the modified Gaussian quantile $Q_n^{\mathrm{Edg}}$ which depends on $\nu_n^{\mathrm{Edg}}$ and $\nu_n^{\mathrm{Approx}}$. The term $\nu_n^{\mathrm{Approx}}$ is a random quantity\footnote{ Strictly speaking, the denominator of \(\nu_n^{\mathrm{Approx}}\) could be null in some pathological situations; nevertheless, this does not impact our confidence interval since this term simplifies. } since $R_{n,\mathrm{var}}$ and ${ \widehat{V}}$ depend on the sample. Therefore, unlike \({ q_{\mathcal{N}(0,1)} } \big( 1-\alpha/2 \big)\), $Q_n^{\mathrm{Edg}}$ is a random quantity, introducing another randomness source into the width of $ \textnormal{CI}_{u}^{\textnormal{Edg}}(1-\alpha,n)$.

Properties of our confidence interval

In this section, we state several results on the asymptotic and non-asymptotic properties of our confidence interval $ \textnormal{CI}_{u}^{\textnormal{Edg}}(1-\alpha,n)$. We start by a result on the non-asymptotic validity, proved in Section (ref).

thmLet \(n \geq 1\), $\delta_n \geq \sup_{\theta \in \Theta_{\mathrm{EE}}} \min\{ {\Delta_{n,\mathrm{E}}}(\theta) , {\Delta_{n,\mathrm{B}}}(\theta) \}$, $a_n \in (1, +\infty)$, and \(\omega_n \in (0, 1)\). Then, \( \textnormal{CI}_{u}^{\textnormal{Edg}}(1-\alpha,n)\) is non-asymptotically valid over \(\Theta_{\mathrm{EE}}\) at level \(1-\alpha\).

This theorem shows it is possible to build a CI that is non-asymptotically valid over a large class of data-generating processes (here $\Theta_{\mathrm{EE}}$) without imposing independence between $X$ and $\varepsilon$, nor parametric restrictions on $\varepsilon$. What is the behavior of our CI as $n$ goes to infinity? Under appropriate restrictions on $a_n$, $\omega_n$ and $\delta_n$, it is shown in Section (ref) that, for any positive constants \(K_{\xi}\), \(K_{\mathrm{reg}}\), \(K_{\varepsilon}\), and \(\lambda_{\mathrm{reg}}\), and for any \(\theta \in \Theta\),

align[align omitted — 280 chars of source]

with the values $K_{\xi}, K_{\mathrm{reg}}, K_{\varepsilon}, \lambda_{\mathrm{reg}}$ hidden in the $o_{P_\theta}(1)$ terms in this representation. In other words, our interval coincides at the limit \(n \to \infty\) with the standard interval obtained from the CLT and therefore is asymptotically exact pointwise over the whole parameter set \(\Theta\). Theorem (ref) (proved in Section (ref)) formalizes the latter result.

thmLet $(\delta_n)$ be such that ${\delta_n \geq \sup_{\theta \in \Theta_{\mathrm{EE}}} \min\{ {\Delta_{n,\mathrm{E}}}(\theta) , {\Delta_{n,\mathrm{B}}}(\theta) \}}$ and $\delta_n \to 0$. Assume that \(\omega_n \to 0\), \(\omega_n n^{2/3} \to +\infty\), and that \(b_n := a_n - 1\) is such that \(b_n \to 0\) and \(b_n \sqrt{n} \to +\infty\). Then, for every $(K_{\xi}, K_{\mathrm{reg}}, K_{\varepsilon}, \lambda_{\mathrm{reg}}) \in (0, +\infty)^4$, $ \textnormal{CI}_{u}^{\textnormal{Edg}}(1-\alpha,n)$ is asymptotically exact pointwise over $\Theta$ at level \(1-\alpha\).

To obtain uniform asymptotic exactness, we place ourselves on a subset of \(\Theta_{\mathrm{EE}}\) defined in the following Assumption (ref) while choosing $a_n$, $\omega_n$ and $\delta_n$ as in Theorem (ref). The result is stated in Proposition (ref) below.

hyp[Bounds on DGP (continued)] Let \(\lambda_{\varepsilon} > 0\), $\rho \geq 0$, and \(K_{X} > 0\) be some constants. We define $\Theta_{\mathrm{EE}}^{\mathrm{strict}}$ to be the set of parameters \(\theta = (\beta_{0}, P_{X, \varepsilon}) \in \Theta_{\mathrm{EE}}\) such that the joint distribution $P_{X, \varepsilon}$ satisfies: \begin{enumerate} • \(\lambda_{\min}(\mathbb{E}[XX'\varepsilon^2]) \geq \lambda_{\varepsilon}\); • \(\mathbb{E}\!\left[ \| X \|^{4(1+\rho)}\right] \leq K_{X}\). \end{enumerate}

Assumption (ref)(i) implies in particular a bound on the kurtosis of $K_{\xi}$ on $\Theta_{\mathrm{EE}}^{\mathrm{strict}}$. Therefore, this gives a potentially stricter bound on $K_{\xi}$ than what was previously assumed on $\Theta_{\mathrm{EE}}$.

propAssume that the sequences $(a_n)$, $(\omega_n)$ and $(\delta_n)$ satisfy the conditions of Theorem (ref). Then, $ \textnormal{CI}_{u}^{\textnormal{Edg}}(1-\alpha,n)$ is asymptotically exact uniformly over $\Theta_{\mathrm{EE}}^{\mathrm{strict}}$ at level \(1-\alpha\).

Combining Theorem (ref), Proposition (ref), and Equation (ref) guarantees that $ \textnormal{CI}_{u}^{\textnormal{Edg}}(1-\alpha,n)$ be NAVAE uniformly over $\Theta_{\mathrm{EE}}^{\mathrm{strict}}$ and efficient (since it is of the form (ref)).

To achieve asymptotic uniform exactness, we restrict the parameter space from $\Theta_{\mathrm{EE}}$ (where non-asymptotic validity holds) to a smaller subset, $\Theta_{\mathrm{EE}}^{\mathrm{strict}}$, by strengthening two moment conditions. Part (i) requires that \(\lambda_{\min}(\mathbb{E}[XX'\varepsilon^2])\) is not only positive, but equal to or larger than the constant \(\lambda_{\varepsilon}\). Part (ii) reinforces \(\mathbb{E}[\left\|X\right\|^4] < +\infty\) in two directions. First, it specifies an explicit upper bound on moments of $P_X$. Second, it enables (if \(\rho > 0\)) to consider higher order moments, which yields faster rates (see Theorem (ref)). Overall, \(\Theta_{\mathrm{EE}}^{\mathrm{strict}} \subsetneq \Theta_{\mathrm{EE}} \subsetneq \Theta\), and the properties of our interval under those different assumptions are summarized in Table (ref) below.

table[table omitted — 1,173 chars of source]

Theorem (ref) and Proposition (ref) specify adequate rates for the choice of \(a_n\) and \(\omega_n\), although they do not provide explicit values for those tuning parameters. As in the case of the expectation, we could consider minimizing the width of the interval (when it is informative) to choose them. However, the situation is more intricate in the OLS case\footnote{Since the width of \( \textnormal{CI}_{u}^{\textnormal{Edg}}(1-\alpha,n)\) is stochastic, and may not even admit an expectation.}. Therefore, we follow another path here by focusing on the coverage probabilities.

Proposition (ref) is actually a consequence of the following theorem that gives a precise bound on the worst-case distance between the coverage probability of our CI and its nominal level. It quantifies the error of the uniform asymptotic exactness property.

thm[Non-asymptotic bound on uniform exactness with rates] Assume that the sequences $(a_n)$, $(\omega_n)$ and $(\delta_n)$ satisfy the conditions of Theorem (ref). Then, there exist $C>0$ and $n_* > 1$ (both depending on $\alpha$) such that, for every $n \geq n_*$, we have \begin{align*} \sup_{\theta \in \Theta_{\mathrm{EE}}^{\mathrm{strict}}} {\mathbb{P}_{\theta}} \Big( {u}' \beta_{0} \in CI_{u}^{Edg}(1-\alpha,n) \Big) &\leq 1 - \alpha + C\left\{\sqrt{b_n} + \nu_n^{\mathrm{Edg}} + \mathfrak{e}_n \right\}, \end{align*} where $\mathfrak{e}_n :=$ \begin{align*} \left\{ \begin{array}{ll} \dfrac{1}{n^{1/4} \omega_n^{3/8}} & if \rho = 0, \\[1.5em] \min\!\left\{ \! \dfrac{1}{n^{1/4} \omega_n^{3/8}} \, , \, \dfrac{1}{(n\omega_n)^{1/4}} \! + \! \dfrac{1}{\sqrt{n}\omega_n^{3/4}} \! + \! \ln(n)\left( \dfrac{\mathds{1}\{\rho \leq 1\}}{n^{\rho}} \! + \! \dfrac{\mathds{1}\{\rho > 1\}}{n^{(1+\rho)/2}} \right) \! \right\} & if \rho \in (0 , + \infty), \\[1.5em] \dfrac{1}{(n\omega_n)^{1/4}} + \dfrac{1}{\sqrt{n}\omega_n^{3/4}} & if \rho = + \infty. \end{array} \right. \end{align*}

Combining this result with the non-asymptotic validity of our confidence interval (Theorem (ref)), we directly obtain that

align*[align* omitted — 291 chars of source]

for the same constant $C > 0$ and $n \geq n_*$. We now specialize Theorem (ref) to obtain the best rates for explicit values of the tuning parameters of the form $\omega_n = n^{-a}$ and $b_n = n^{-b}$.

propAssume that the sequence $(\delta_n)$ satisfies the conditions of Theorem (ref). Let $\rho \in [0, +\infty]$. If $\omega_n = n^{-r(\rho)}$ and $b_n = n^{- b}$ for any $b \in [2/5, 1/2)$, then \begin{align*} \sup_{\theta \in \Theta_{\mathrm{EE}}^{\mathrm{strict}}} {\mathbb{P}_{\theta}} \Big( {u}' \beta_{0} \in CI_{u}^{Edg}(1-\alpha,n) \Big) &\leq 1 - \alpha + C n^{- r(\rho)}, \end{align*} for some constant $C > 0$, where \begin{align*} r(\rho) := \frac{2}{11} \mathds{1}\{\rho < 2/11\} + \rho \mathds{1}\{2/11 \leq \rho \leq 1/5\} + \frac{1}{5} \mathds{1}\{1/5 < \rho\}. \end{align*}

This proposition is proved in Section (ref). Interestingly, additional constraints on the moments of $\| X \|$ from $4$ to $4 \times (1 + 2/11) = 52 / 11 \approx 4.72$ does not seem to improve the rate. Between the $4.72$-th moment and the $4 \times (1 + 1/5) = 24 / 5 = 4.75$-th moment, the rate $r(\rho)$ is the identity function. Beyond the $4.75$-th moment, the rate is fixed and not improved by boundedness of additional moments. Overall, the effect of $\rho$ is quite mild since with four moments the rate is close to $n^{-0.1819}$ whereas it attains $n^{-0.20}$ when all moments are bounded.

Practical considerations and simulation study

R code for our simulations is available at\newline \url{https://github.com/AlexisDerumigny/Reproducibility-NAVAE_CI_LinModels}.

Plug-in

Be it for expectations or the coefficients of a linear regression, our assumptions impose lower or upper bounds on moments (or functions thereof) of the distribution \(P_{\xi}\) (Section (ref)) and \(P_{X, \, \varepsilon}\) (Section (ref)). Remember that those bounds are required to compute our CIs. While a priori choices for those bounds are natural in specific cases (see Remark (ref) below), it often remains difficult to form intuition about their values or how to choose them. Instead, a natural idea is to replace these (unknown) bounds with estimates of the corresponding moments. In the case of expectations, the confidence interval depends only on one bound \(K\) (see Section (ref)), on the kurtosis of the distribution of the observations, which can be approximated by

itemize\(\widehat{K} := n^{-1} \sum_{i=1}^n (\xi_i - \overline{\xi}_{n})^4 \, / \, \big[n^{-1} \sum_{i=1}^n (\xi_i - \overline{\xi}_{n})^2\big]^2 \).

For linear regressions, several bounds are involved (see Assumption (ref)) that can be approximated by

itemize\(\widehat\lambda_{\mathrm{reg}} := \lambda_{\min}(n^{-1} \sum_{i=1}^n X_i X_i'))\), • \(\widehatK_{\mathrm{reg}} := n^{-1} \sum_{i=1}^n \Big\| {\textnormal{vec}} \Big( \widehat{\widetilde{X}_i} \widehat{\widetilde{X}'_i} - \mathbb{I}_p \Big) \Big\|^2\) with \(\widehat{\widetilde{X}_i} := \left(\big( n^{-1} \sum_{j=1}^{n} X_j X_j' \big)^{\dagger}\right)^{1/2} X_i\), • \(\widehatK_{\varepsilon} := n^{-1} \sum_{i=1}^n \Big\|\widehat{\widetilde{X}_i} \widehat{\varepsilon}_i \Big\|^4\) with \(\widehat{\varepsilon}_i := Y_i-X_i'\widehat{\beta}\), • \(\widehat K_{\xi} := n^{-1} \sum_{i=1}^n \widehat{\xi_i}^4 \,/\, \Big( n^{-1} \sum_{i=1}^n \widehat{\xi_i}^2 \Big)^2\) with \(\widehat\xi_i := {u}'\big( n^{-1} \sum_{j=1}^{n} X_j X_j' \big)^{\dagger} X_i \widehat{\varepsilon}_i\).

Using plug-in estimates rather than deterministic bounds has one major drawback from a theoretical point of view: the resulting CI is no longer non-asymptotically valid. Nevertheless, in both cases (Sections (ref) and (ref)), it is still pointwise asymptotically exact over \(\Theta_{\mathrm{EE}}\). More generally, with a similar reasoning as that of Theorem (ref) (where the choice of the bounds \((K_{\xi}, K_{\mathrm{reg}}, K_{\varepsilon}, \lambda_{\mathrm{reg}})\) does not matter), plug-in CIs remain asymptotically exact pointwise as soon as the previous plug-in approximations have finite non-zero limits in probability. In a sense, the use of plug-in in our approach can be compared to the approximation error faced in practice when one uses an inference procedure based on simulations such as that of cherno2009 or diciccio2017robust. To further control the impact of sampling uncertainty when using a plug-in version of our interval, a practical possibility would be to multiply the plug-in estimates by \((1 + M/\sqrt{n})\) for some positive constant \(M\).

rem[Some possibilities to overcome plug-in] Choosing reasonable values for \(K\) and \(K_{\xi}\) without resorting to a plug-in strategy turns out to be possible. A large class of univariate distributions exhibits a bound of at most $9$ on the kurtosis: Normal, Laplace, asymmetric Laplace, Logistic, Uniform, Student with at least five degrees of freedom, two-point symmetric mixtures of Normals, Gumbel, hyperbolic secant, and skewed Normal. This class includes both symmetric and asymmetric distributions, some of which only have a few number of finite moments (Student distributions with few degrees of freedom). We investigate the impact of the choice \(K = 9\) for expectations (respectively, \(K_{\xi} = 9\) for linear regressions) as an alternative to the plug-in approach in the following simulation section.

Simulations for Section (ref): inference on an expectation

Framework

This section presents some simulation results on our confidence interval for an expectation. We consider an i.i.d. sample from an Exponential distribution with expectation set to \(1\), the targeted parameter \(\theta_1\) here. Compared to Normal distributions, remember that Exponential distributions are skewed (with a skewness coefficient equal to \(2\)) and display fatter tails with a kurtosis of \(9\). In that sense, we consider the boundary case (while remaining correctly specified) when setting a bound \(K = 9\) on the kurtosis as discussed in Remark (ref). We present the results for that choice \(K = 9\) and for a plug-in version of our CI using the empirical kurtosis \(\widehat{K}\) instead.

In addition to the bound \(K\), \({\mathrm{CI}_{\mathrm{\sigma_{\textnormal{known}}}}(1-\alpha,n)}\) depends on \(\delta_n\) and \({\mathrm{CI}_{\mathrm{ \widehat{\sigma}}}(1-\alpha,n)}\) depends additionally on \(a_n\). We follow Remark (ref) by choosing \(\delta_n\) as the minimum between \(\delta_{1,n}\) proposed by shevtsova2013 and \(\delta_{2,n}\) by derumigny2022explicitBE. We thus do not take advantage of the continuity of the Exponential distribution (see \(\delta_{3,n}\) and the comparisons made in derumigny2022explicitBE for further details). We set \(a_n = 1 + n^{-1/5}\) satisfying the requirements of Proposition (ref) for the asymptotic exactness of our CI. Choosing that tuning parameter is a tradeoff between exiting the \(\mathbb{R}\) regime earlier (meaning, for smaller \(n\) or smaller \(\alpha\)) and the precision of our CI (meaning, its width). For instance, with a smaller power of \(n\), say, \(a_n = 1 + n^{-1/10}\), the minimal \(\alpha\) that exits the \(\mathbb{R}\) regime with \(K = 9\) and \(n = 1,000\) is \(15.6\)% instead of 26.1% (see Table (ref) below), but the resulting CI's width is larger.

About the minimum level \texorpdfstring{${ \alpha_{\min_{}} }$}{alpha min}

We focus here on the case of an unknown variance. Table (ref) reports ${ \alpha_{\min_{}} }$, defined as the minimal \(\alpha\) for which our CI is informative, that is, satisfies \({\mathrm{CI}_{\mathrm{ \widehat{\sigma}}}(1-\alpha,n)} \subsetneq \mathbb{R}\). Note that ${ \alpha_{\min_{}} }$ depends on $K$, $a_n$ and $\delta_n$. For instance, for \(K = 9\), \(a_n = 1 + n^{-1/5}\) and \(\delta_n\) chosen as above, a sample with \(n = 5,000\) observations allows to compute an informative \({\mathrm{CI}_{\mathrm{ \widehat{\sigma}}}(1-\alpha,n)}\) at level 90% but not at 95% (since \({ \alpha_{\min_{}} } \approx 7\)%). Intuitively, given the tools we use to build our CI and the values of \(K\), \(a_n\), and \(\delta_n\), a 95% nominal level is too strong a requirement to guarantee non-asymptotic validity for that sample size. With \(K = \widehat{K}\), the condition to exit the \(\mathbb{R}\) regime becomes random and so does ${ \alpha_{\min_{}} }$. We report the mean and median of \({ \alpha_{\min_{}} }\) obtained over \(M = 20,000\) Monte-Carlo repetitions. As expected, for large enough sample sizes, using \(K = 9\) or \(K = \widehat{K}\) leads to indistinguishable results as \(\widehat{K}\) converges to the kurtosis of an Exponential (equal to \(9\)).

table[table omitted — 987 chars of source]

Coverage performance

We now focus on the choice \(K = 9\) and consider coverage performance for \(\alpha = 0.10\). We present the results by comparing our CIs with the classical one derived from the Central Limit Theorem and Slutsky's lemma, \({\mathrm{CI}_{\mathrm{CLT}}(1-\alpha,n)}\). Table (ref) reports the empirical coverage over \(M = 20,000\) Monte-Carlo repetitions. As expected from Propositions (ref) and (ref), the coverage of our CI is always greater than the nominal level \(1 - \alpha\), 90% here, and decreases to the nominal level when the sample size increases.

In the unknown variance case, we compare two versions of our CI: one with the a priori choice \(a_n = 1 + n^{-1/5}\) -- which satisfies the requirements of Proposition (ref) -- and one with the optimized choice of \(a_n\) discussed in Remark (ref). The latter choice logically reduces coverage closer to the targeted nominal level, although the improvement remains limited.

Finally, the third row presents the coverage of \({\mathrm{CI}_{\mathrm{\sigma_{\textnormal{known}}}}(1-\alpha,n)}\). As expected, the coverage of this oracle CI lies in between those of the standard CLT-based CI and of our feasible one since \({\mathrm{CI}_{\mathrm{\sigma_{\textnormal{known}}}}(1-\alpha,n)}\) bypasses the additional step of replacing the variance by a consistent estimator.

table[table omitted — 1,010 chars of source]

Width of the confidence intervals

Figure (ref) compares the confidence intervals \({\mathrm{CI}_{\mathrm{\sigma_{\textnormal{known}}}}(1-\alpha,n)}\), \({\mathrm{CI}_{\mathrm{ \widehat{\sigma}}}(1-\alpha,n)}\) and \({\mathrm{CI}_{\mathrm{CLT}}(1-\alpha,n)}\) regarding their width using the bound \(K = 9\) on the kurtosis and the optimized tuning parameter \(a_n\) when the variance is unknown.

Panel (a) reports the average width of the intervals as a function of the sample size \(n\) over \(M = 20,000\) Monte-Carlo repetitions (the absolute widths of the latter two CIs are data-dependent, hence stochastic, through \(\widehat{\sigma}\)).

Panel (b) shows the relative width of our CIs with respect to the usual CLT-based CI. When the variance is assumed to be known, we also use that information for the CLT-based interval and compare \({\mathrm{CI}_{\mathrm{\sigma_{\textnormal{known}}}}(1-\alpha,n)}\) to the analogue of \({\mathrm{CI}_{\mathrm{CLT}}(1-\alpha,n)}\) with a known variance, replacing \(\sigma_{\textnormal{known}}\) with \( \widehat{\sigma}\) in Equation (ref). The ratio of their widths (blue line in Panel (b)) is equal to \({ q_{\mathcal{N}(0,1)} } \big( 1 - \alpha/2 + \delta_n \big) \, / \; { q_{\mathcal{N}(0,1)} } \big( 1 - \alpha/2 \big)\) and is not data-dependent (only \(n\) and $K$ matter, through \(\delta_n\)). When the variance is unknown, the width ratio between \({\mathrm{CI}_{\mathrm{ \widehat{\sigma}}}(1-\alpha,n)}\) and \({\mathrm{CI}_{\mathrm{CLT}}(1-\alpha,n)}\) (red line) is equal to \(C_n { q_{\mathcal{N}(0,1)} } \big( 1 - \alpha/2 + \delta_n + \nu_n^{\textrm{Var}}/2 \big) \, / \; { q_{\mathcal{N}(0,1)} } \big( 1 - \alpha/2 \big)\), and is also not data-dependent as \( \widehat{\sigma}\) cancels out.

The convergence to \(1\) of the relative widths illustrates that our CI coincides in the limit with the classical one obtained from the CLT. For information, the inflexion point (at \(n = \) 38,707) corresponds to the switch from \(\delta_{1, n}\) to \(\delta_{2, n}\), that is, from Berry-Esseen inequalities to Edgeworth Expansions to obtain \(\delta_n\) (see Remark (ref)).

figure[figure omitted — 1,283 chars of source]

Simulations for Section (ref): inference on linear regression' coefficients

We now consider a simulation study on linear regressions. We generate i.i.d. samples from the model \(Y = 2 + 1 X_1 - 3 X_2 + \varepsilon\) where \((X_1, X_2)\) follows a centered bivariate Normal distribution, with respective variances 1 and 2 and a correlation set to 0.5, and \(\varepsilon\) drawn from a Gumbel distribution parameterized such that the error terms' expectation is null and their conditional variance is equal to \((X_1 + X_2)^2\), implying heteroskedasticity. Gumbel distributions are skewed and have heavier tails than Normal ones. We perform inference on \(\beta_{0,{2}}\), that is, we choose \({u} = (0,1,0)'\).

For \(\delta_n\), we choose again the minimum between \(\delta_{1, n}\) and \(\delta_{2, n}\) (see Remark (ref)). While our CIs for expectations require only one bound \(K\) on the kurtosis, several bounds are needed to compute \( \textnormal{CI}_{u}^{\textnormal{Edg}}(1-\alpha,n)\). As explained in Section (ref), while the choice \(K_{\xi} = 9\) is sensible for the kurtosis of the influence function, fixing values for \(\lambda_{\mathrm{reg}}\), \(K_{\mathrm{reg}}\), and \(K_{\varepsilon}\) is not straightforward. In the simulations, we stick to the choice \(K_{\xi}=9\) and resort to plug-in versions of the other bounds. For the tuning parameters, we set \(b_n = 20 \times n^{-2/5}\) and \(\omega_n = n^{-1/5}\) which complies with the assumptions underlying Propositions (ref) and (ref).

Remember that our interval is informative (not equal to \(\mathbb{R}\)) when the sample size is such that \(n > 2 K_{\mathrm{reg}} / (\omega_n \alpha)\) and \(\nu_n^{\mathrm{Edg}} < \alpha / 2\). With fixed bounds, both conditions are non-stochastic. In our simulations, the former becomes stochastic as we use the plug-in \(\widehat{K_{\mathrm{reg}}}\). However, in our setting and for \(\alpha = 0.10\), only the latter condition, \(\nu_n^{\mathrm{Edg}} < \alpha / 2\), happens to be binding and is satisfied for \(n \geq\) 3,656.

Figure (ref) reports the absolute and relative width of \( \textnormal{CI}_{u}^{\textnormal{Edg}}(1-\alpha,n)\) with respect to \(\textnormal{CI}_{u}^{\textnormal{Asymp}}(1-\alpha,n)\) for \(n\) ranging from 5,000 to 200,000; more precisely, we report their average over \(M =\) 20,000 Monte-Carlo repetitions. Our CI remains far wider than the one based on the CLT, even for large sample sizes. This can stem from two possible sources: either a lack of tightness of our intervals or a substantial gap between the shortest possible NAVAE CI and \(\textnormal{CI}_{u}^{\textnormal{Asymp}}(1-\alpha,n)\) over the class of distributions we consider. Finding the shortest non-asymptotically valid confidence interval is a hard task and this question is left for future research.

figure[figure omitted — 1,181 chars of source]