Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
91,700 characters · 12 sections · 60 citation commands
Explicit non-asymptotic bounds for the distance to the first-order Edgeworth expansion
Keywords: Berry-Esseen bound, Edgeworth expansion, normal approximation, central limit theorem, non-asymptotic tests.
MSC Classification: 62E17; 60F05; 62F03.
As the number of observations $n$ in a statistical experiment goes to infinity, many statistics of interest have the property to converge weakly to a $\mathcal{N}(0,1)$ distribution, once adequately centered and scaled, see, e.g., Chapter 5 of vanderVaart2000 for a thorough introduction. Hence, when little is known on the distribution of a statistic for a fixed sample size, a classical approach to conduct inference on the parameters of the statistical model amounts to approximating that distribution by its tractable Gaussian limit. A recurring theme in statistics and probability is thus to quantify the distance between those two distributions for a given $n$.
In this article, we present some refined results in the canonical case of a standardized sum of independent random variables. We consider independent but not necessarily identically distributed random variables to encompass a broader range of applications. For instance, certain bootstrap schemes such as the multiplier ones (see Chapter 9 in vanderVaartWellner1996 or Chapter 10 in Kosorok2006) boil down to studying a sequence of mutually independent not necessarily identically distributed (i.n.i.d.) random variables conditional on the initial sample.
More formally, let ${(X_i)_{i=1, \dots, n}}$ be a sequence of i.n.i.d. random variables satisfying for every ${i \in \{1,...,n\}}$, ${\mathbb{E}[X_i] = 0}$ and ${\gamma_i := \mathbb{E}[ X_i^4 ] < +\infty}$. We also define the standard deviation $B_n$ of the sum of the $X_i$'s, i.e., ${B_n := \sqrt{\sum_{i=1}^n\mathbb{E}[X_i^2]}},$ so that the standardized sum can be written as ${S_n := \sum_{i=1}^n X_i/B_n}$. Finally, we define the average individual standard deviation ${\overline{B}_n := B_n/\sqrt{n}}$ and the average standardized third raw moment ${\lambda_{3,n} := \frac{1}{n}\sum_{i=1}^n\mathbb{E}[X_i^3]/\overline{B}_n^3}$. The main results of this article are of the form
where $\Phi$ is the cumulative distribution function of a standard Gaussian random variable, $\varphi$ its density function and $\delta_n$ is a positive sequence that depends on the first four moments of $(X_i)_{i=1,\dots,n}$ and tends to zero under some regularity conditions. In the following, we use the notation $G_n(x) := \Phi(x) + \lambda_{3,n} (6\sqrt{n})^{-1} (1-x^2)\varphi(x)$.
The quantity $G_n(x)$ is usually called the one-term Edgeworth expansion of $\mathbb{P}\left(S_n \leq x\right)$, hence the letter E in the notation $\Delta_{n,\text{E}}$. Controlling the uniform distance between $\mathbb{P}\left(S_n \leq \cdot \right)$ and $G_n(\cdot)$ has a long tradition in statistics and probability, see for instance esseen1945 and the books by cramer1962 and bhattacharyarao1976. As early as in the work of esseen1945, it was acknowledged that in independent and identically distributed (i.i.d.) cases, $\Delta_{n,\text{E}}$ was of the order $n^{-1/2}$ in general and of the order $n^{-1}$ if $(X_i)_{i=1,\dots,n}$ has a nonzero continuous component. These results were then extended in a wide variety of directions, often in connection with bootstrap procedures, see for instance hall1992bootstrap and lahiri2003resampling for the dependent case.
A one-term Edgeworth expansion can be seen as a refinement of the so-called Berry-Esseen inequality (berry1941, esseen1942) which goal is to bound
The refinement stems from the fact that in $\Delta_{n,\text{E}},$ the distance between $\mathbb{P}\left(S_n \leq \cdot \right)$ and $\Phi$ is adjusted for the presence of non-asymptotic skewness in the distribution of $S_n$. Contrary to the literature on Edgeworth expansions, there is a substantial amount of work devoted to explicit constants in the Berry-Esseen inequality and its extensions, see, e.g., bentkus1996, bentkus2003, pinelis2016, chernozhukov2017, raivc2018multivariate, raic2019. The sharpest known result in the i.n.i.d. univariate framework is due to shevtsova2013, which shows that for every ${n \in \mathbb{N}^{*}}$, if ${\mathbb{E}[|X_i|^3]<+\infty}$ for every ${i \in \{1,...,n\}}$, then $\Delta_{n,\text{B}} \leq 0.5583 \, K_{3,n} / \sqrt{n}$ where $K_{p,n} := n^{-1} \sum_{i=1}^n\mathbb{E}[|X_i|^p]/(\overline{B}_n)^p$, for ${p \in \mathbb{N}^{*}}$, denotes the average standardized $p$-th absolute moment. $K_{p,n}$ measures tail thickness, with $K_{2,n}$ normalized to 1 and $K_{4,n}$ the kurtosis. An analogous result is given in shevtsova2013 under the i.i.d. assumption where $0.5583 $ is replaced with $0.4690$. A close lower bound is due to esseen1956: there exists a distribution such that $\Delta_{n,\text{B}} = (C_B/\sqrt{n}) \left(n^{-1} \sum_{i=1}^n\mathbb{E}[|X_i|^3]/\overline{B}_n^3\right)$ with ${C_B \approx 0.4098}$. Another line of research applies Edgeworth expansions in order to get a bound on $\Delta_{n,\text{B}}$ that contains higher-order terms, see adell2008, boutsikas2011 and zhilova2020new.
Despite the breadth of those theoretical advances, there remain some limits to take full advantage of those results even in simple statistical applications, for instance, when conducting inference on the expectation of a real random variable.\footnote{In this article, we only give results for standardized sums of random variables, i.e., sums that are rescaled by their standard deviation. In practice, the variance is unknown and has to be replaced with some empirical counterpart, leading to what is usually called a self-normalized sum. This is an important question in practice that we leave aside for future research. There exist numerous results on self-normalized sums in the fields of Edgeworth expansions and Berry-Esseen inequalities (hall1987, delapena2009). However, the practical limitations of existing results that we point out in our work still prevail.} If we focus on Berry-Esseen inequalities, we show in Section (ref) shows that even the sharpest upper bound to date on $\Delta_{n,\text{B}}$ can be uninformative when conducting inference on an expectation even for $n$ larger than 59,000. Therefore, it is natural to wonder whether bounds derived from a one-term Edgeworth expansion could be tighter in moderately large samples (such as a few thousands). In the i.i.d. case and under some smoothness conditions, senatov2011 obtains such improved bounds. To our knowledge, the question is nevertheless still open in the i.n.i.d. setup, as well as in the general setup when no condition on the characteristic function is assumed. In particular, most articles that present results of the form of (ref) do not provide a fully explicit value for $\delta_n$, that is, $\delta_n$ is defined up to some “universal” but unknown constant, see for instance cramer1962 and bentkus1997, among others.
In this article, we derive novel inequalities of the form of (ref) that aim to be relevant in practical applications. Such “user-friendly” bounds seek to achieve two goals. First, we provide explicit values for $\delta_n$, which are implemented in the new R package BoundEdgeworth packageBoundEdgeWorth using the function Bound_EE1 (the function Bound_BE provides a bound on \(\Delta_{n,\text{B}}\)). Second, the bounds $\delta_n$ should be small enough to be informative even with small (${n \approx}$ hundreds) to moderate (${n \approx}$ thousands) sample sizes. We obtain these bounds in an i.i.d. setting and in a more general i.n.i.d. case only assuming finite fourth moments.
We give improved bounds on $\Delta_{n,\text{E}}$ under some regularity assumptions on the tail behavior of the characteristic function $f_{S_n}$ of $S_n$. Such conditions are related to the continuity of the distribution of $S_n$ and the differentiability of the corresponding density (with respect to Lebesgue's measure). These are well-known conditions required for the Edgeworth expansion to be a good approximation of ${\mathbb{P}(S_n\leq \cdot \,)}$ with fast rates. Our main results are summed up in Table (ref).
In the rest of this section, we introduce notation used in the rest of the paper. Section (ref) presents our bounds on $\Delta_{n,\text{E}}$ under moment conditions only in i.n.i.d. or i.i.d. settings. In Section (ref), we develop tighter bounds under regularity assumptions on the characteristic function of $S_n$. They rely on an alternative control of \(\Delta_{n,\text{E}}\) that involves the integral of \(f_{S_n}\), enabling us to use additional regularity assumptions on the tails of that function. In Section (ref), we discuss practical aspects related to our bounds: how to choose or estimate the moments of the distribution of \(S_n\) involved in order to compute our bounds. We also perform numerical comparisons between our and existing bounds for some particular distributions (Student and Gamma).\footnote{ The code to replicate our results is available in the Github repository \newline \url{https://github.com/AlexisDerumigny/Reproducibility-BoundsDistanceEdgeworth}. } In Section (ref), we apply our results to analyze several aspects of one-sided tests based on the normal approximation of a sample mean. In particular, based on our bounds, we propose a new method to compute sufficient sample sizes for experimental design with given effect size to be detected and nominal power. All proofs are postponed in the appendix. The proofs of the main results are gathered in Appendix (ref), relying on the computations of Appendix (ref). Useful lemmas are given in Appendix (ref).
Additional notation. $\vee$ (resp. $\wedge$) denotes the maximum (resp. minimum) operator. For a random variable $X$, we denote its probability distribution by $P_X$. For a distribution $P$, let $f_P$ denote its characteristic function; similarly, for a random variable $X$, we denote by $f_X$ its characteristic function. We recall that $f_{\mathcal{N}(0,1)}(t)=e^{-t^2/2}$. We denote the (extended) lower incomplete Gamma function by $\gamma(a, x) := \int_0^x |u|^{a-1} e^{-u} du$ (for $a > 0$ and $x \in \mathbb{R}$), the upper incomplete Gamma function by $\Gamma(a,x) := \int_x^{+\infty} u^{a-1} e^{-u} du$ (for $a \geq 0$ and $x > 0$) and the standard gamma function by $\Gamma(a) := \Gamma(a,0) = \int_0^{+\infty} u^{a-1} e^{-u} du$ (for $a > 0$). For two sequences $(a_n),$ $(b_n),$ we write $a_n = O(b_n)$ whenever there exists $C>0$ such that ${a_n \leq C b_n}$; $a_n = o(b_n)$ whenever $a_n / b_n \to 0$; and $a_n \asymp b_n$ whenever $a_n = O(b_n)$ and $b_n = O(a_n)$. We denote by $\chi_1$ the constant $\chi_1 := \sup_{x>0} x^{-3} |\cos(x)-1+x^2/2| \approx 0.099$ shevtsova2010, and by $\theta_1^*$ the unique root in $(0,2\pi)$ of the equation $\theta^2+2\theta\sin(\theta)+6(\cos(\theta)-1)=0$. We also define $t_1^* := \theta_1^* / (2\pi) \approx 0.64$ shevtsova2010. For every ${i \in \mathbb{N}^{*}}$, we define the individual standard deviation ${\sigma_{i} := \sqrt{\mathbb{E}[X_i^2]}}$. Henceforth, we reason for a fixed arbitrary sample size ${n \in \mathbb{N}^{*}}$. Densities and continuous distributions are always assumed implicitly to be with respect to Lebesgue's measure.
For clarity, we define below the concept of an explicit expression. In the rest of the article, the goal is to find bounds on $\Delta_{n,\text{E}}$ that are explicit expressions in the sense of Definition (ref).
We start by introducing two versions of our basic assumptions on the distribution of the variables $(X_i)_{i=1, \dots, n}$.
Assumption (ref) corresponds to the classical i.i.d. sampling with finite fourth moment while Assumption (ref) is its generalization in the i.n.i.d. framework. Those two assumptions primarily ensure that enough moments of $(X_i)_{i=1,\dots,n}$ exist to build a non-asymptotic upper bound on $\Delta_{n,\text{E}}.$ In some applications, such as the bootstrap, it is required to consider an array of random variables $(X_{i,n})_{i=1,\dots,n}$ instead of a sequence. For example, efron1979's nonparametric bootstrap procedure consists in drawing $n$ elements in the random sample $(X_{1,n},...,X_{n,n})$ with replacement. Conditional on $(X_{i,n})_{i=1,\dots,n},$ the $n$ values drawn with replacement can be seen as a sequence of $n$ i.i.d. random variables with distribution $\frac{1}{n}\sum_{i=1}^n\delta_{\{X_{i,n}\}}$, denoting by $\delta_{\{a\}}$ the Dirac measure at a given point ${a \in \mathbb{R}}$. Our results encompass these situations directly. Nonetheless, we do not use the array terminology here as our results hold non-asymptotically, i.e., for any fixed sample size $n$.
To state our first theorem, remember that ${\overline{B}_n := (1 / \sqrt{n}) \sqrt{\sum_{i=1}^n\sigma_{i}^2}}$, for $p \in \mathbb{N}^{*}$, $K_{p,n} := n^{-1} \sum_{i=1}^n\mathbb{E}[|X_i|^p]/ \overline{B}_n^p$, and let us introduce \(\widetilde{K}_{3,n} := K_{3,n} + \frac{1}{n}\sum_{i=1}^n\mathbb{E}|X_i|\sigma_{i}^2 / \overline{B}_n^3\), $\Delta:= (1 - 4 \chi_1 - \sqrt{K_{4,n}/n}) / 2$, and the terms $r_{1,n}^{\textnormal{inid,skew}}$, $r_{1,n}^{\textnormal{inid,noskew}}$, $r_{1,n}^{\textnormal{iid,skew}}$ and $r_{1,n}^{\textnormal{iid,noskew}}$.
These remainder terms are defined by:
and
where
The following theorem is proved in Sections (ref) (“i.n.i.d.” case) and (ref) (“i.i.d.” case).
{\color{black}
}
Note that it is possible to replace $\widetilde{K}_{3,n}$ by the simpler upper bound $2 K_{3,n}$ under Assumption (ref) (respectively by $K_{3,n}+1$ under Assumption (ref)). This theorem displays a bound of order $n^{-1/2}$ on $\Delta_{n,\text{E}}$ {\color{black}in the regime where $K_{4,n}$ is bounded by a fixed constant.} The rate $n^{-1/2}$ cannot be improved when only assuming moment conditions on $(X_i)_{i=1,\dots,n}$ (esseen1945, cramer1962). Another nice aspect of those bounds is their dependence on $\lambda_{3,n}$. For many classes of distributions, $\lambda_{3,n}$ can, in fact, be exactly zero. This is the case if for every $i = 1,\dots,n$, $X_i$ has a non-skewed distribution, such as any distribution that is symmetric around its expectation. More generally, $|\lambda_{3,n}|$ can be substantially smaller than $K_{3,n}$, decreasing the related terms.
As mentioned in the Introduction, we are not aware of explicit bounds on $\Delta_{n,\text{E}}$ under moment conditions only. It is thus difficult to assess how our bounds compare to the literature. On the other hand, there exist well-established bounds on $\Delta_{n,\text{B}}$. Using Theorem (ref), the bound $(1-x^2)\varphi(x) / 6 \leq \varphi(0)/6 \leq 0.0665 $ for all $x \in \mathbb{R}$, and applying the triangle inequality, we can control \(\Delta_{n,\text{B}}\) as well. More precisely, for every $n \geq 3$, we have
Under Assumption (ref), \(\widetilde{K}_{3,n} \leq 2 K_{3,n} \). Combined with the refined inequality $|\lambda_{3,n}| \leq 0.621K_{3,n}$ pinelis2011relations, we can derive a simpler bound that involves only $K_{3,n}$
The bound $\Delta_{n,\text{B}} \leq 0.4403 K_{3,n} / \sqrt{n} + O(n^{-1})$ is already tighter than the sharpest known Berry-Esseen inequality in the i.n.i.d. framework, $\Delta_{n,\text{B}} \leq 0.5583 K_{3,n}/\sqrt{n}$, as soon as the remainder term \(O(n^{-1}\) is smaller than the difference $0.118 K_{3,n}/\sqrt{n}$. This bound is also tighter than the sharpest known Berry-Esseen inequality in the i.i.d. case, $\Delta_{n,\text{B}} \leq 0.4690 K_{3,n}/\sqrt{n}$, up to a \(O(n^{-1}\) term. We recall that the sharpest existing bounds shevtsova2013 only require a finite third moment while we use further regularity in the form of a finite fourth moment. We refer to Example (ref) and Figure (ref) for a numerical comparison, showing improvements for $n$ of the order of a few thousands. The most striking improvement is obtained in the unskewed case when ${\mathbb{E}[X_i^3] = 0}$ for every integer $i$. In this case, Theorem (ref) and the inequality ${\widetilde{K}_{3,n} \leq 2 K_{3,n}}$ yield ${\Delta_{n,\text{B}} \leq 0.3990 K_{3,n} / \sqrt{n} + O(n^{-1})}$. Note that this result does not contradict esseen1956's lower bound ${0.4098 K_{3,n} / \sqrt{n}}$ as the distribution he constructs does not satisfy $\mathbb{E}[X_i^3] = 0$ for every $i$.
Under Assumption (ref), $\widetilde{K}_{3,n} \leq K_{3,n}+1$ and we can combine this with (ref) and the inequality $|\lambda_{3,n}| \leq 0.621K_{3,n}$, so that we obtain
As in the i.n.i.d. case discussed above, the numerical constant in front of $K_{3,n}$ in the leading term is smaller than the lower bound constant ${C_B:=0.4098}$ derived in esseen1956. The point is addressed in detail in shevtsova2012, where the author explains that the constant coming from esseen1956 cannot be improved only if one seeks control of $\Delta_{n,\text{B}}$ with a leading term of the form $c_1 K_{3,n}/\sqrt{n}$ for some $c_1 > 0$. In contrast, our bound on $\Delta_{n,\text{B}}$ exhibits a leading term of the form $(c_1 K_{3,n} + c_2)/\sqrt{n}$ for positive constants $c_1$ and $c_2$.
In this section, we derive tighter bounds on $\Delta_{n,\text{E}}$ under additional regularity conditions on the tail behavior of the characteristic function of $S_n$. They follow from Theorem (ref), which provides an alternative upper bound on $\Delta_{n,\text{E}}$ that involves the tail behavior of $f_{S_n}$. To state this theorem, let us introduce the terms $r_{2,n}^{\textnormal{inid,skew}}$, $r_{2,n}^{\textnormal{inid,noskew}}$, $r_{2,n}^{\textnormal{iid,skew}}$ and $r_{2,n}^{\textnormal{iid,noskew}}$ {\color{black}
and
}
Recall also that $t_1^* \approx 0.64$ and let $a_n := 2t_1^*\pi\sqrt{n}/\widetilde{K}_{3,n} \wedge 16\pi^3n^2/\widetilde{K}_{3,n}^4$ and $b_n := 16\pi^4n^2/\widetilde{K}_{3,n}^4$. In practice, even for fairly small \(n\), \(a_n\) is equal to \(2t_1^*\pi\sqrt{n}/\widetilde{K}_{3,n}\).
{\color{black}
}
This theorem is proved in Section (ref) under Assumption (ref) (resp. in Section (ref) under Assumption (ref)). The first term contains quantities that were already present in the term of order $1/n$ in the bound of Theorem (ref): $0.195K_{4,n}$ and $0.038 \lambda_{3,n}^2$. On the contrary, the other terms are encompassed in the integral term and in the remainder. Indeed, a careful reading of the proofs (see notably Section (ref) that outlines the structure of the proofs of all theorems) shows that the leading term $0.1995 \, \widetilde{K}_{3,n} / \sqrt{n}$ in the bound (ref) comes from choosing a free tuning parameter $T$ of the order of $\sqrt{n}$. Here, we make another choice for $T$ such that this term is now negligible. The cost of this change of $T$ is the introduction of the integral term involving $f_{S_n}$. The leading term of the bound thus depends on the tail behavior of $f_{S_n}$.
Note that the result is obtained under the same conditions as Theorem (ref), namely under moment conditions only. Nonetheless, it is mainly interesting combined with some assumptions on \(f_{S_n}\) over the interval \([a_n, b_n]\), otherwise we do not have an explicit control on the integral term involving \(f_{S_n}\). In the rest of this section, we present two possible assumptions on \(f_{S_n}\) that yield such a control.
As a first regularity condition on \(f_{S_n}\), we can assume a polynomial rate decrease. Corollary (ref) presents the resulting bound in the i.n.i.d. case. In fact, a similar condition could be invoked with i.i.d. data by requesting a polynomial decrease of the characteristic function of \(X_n / \sigma_n\). However, we present in the next paragraph milder assumptions in the i.i.d. case that remain sufficient to obtain an explicit control of the tails of $f_{S_n}$.
Besides moment conditions, Corollary (ref) requires a uniform control of $f_{S_n}$ outside the interval $(-a_n,a_n)$. When $\widetilde{K}_{3,n} = o(\sqrt{n})$, $a_n$ goes to infinity. In this case, the condition is a tail control of the characteristic function of $S_n$ in a neighborhood of infinity, thus making the condition weaker to impose.
Placing restrictions on the tails of $f_{S_n}$ is not very common in statistical applications. However, this notion is closely related to the smoothness of the underlying distribution of $S_n$. Proposition (ref) in the Appendix (which builds upon classical results such as ushakov2011) shows that the tail condition on $f_{S_n}$ is satisfied with $p \geq 1$ whenever $P_{S_n}$ has a density $g_{S_n}$ that is $p-1$ times differentiable and such that its $(p-1)$-th derivative is of bounded variation with total variation $V_n := \mathrm{Vari}[g_{S_n}^{(p-1)}]$ uniformly bounded in $n$. In such situations, we can take $C_0 = 1 \vee \sup_{n \in \mathbb{N}^{*}}V_n$.
Although Corollary (ref) is valid for every positive $p$, it is only an improvement on the results of the previous section under the stricter condition ${p > 1}$, a situation in which $P_{S_n}$ admits a density with respect to Lebesgue's measure (second part of Proposition (ref)). In particular when $p=2$, $a_n^{-p}$ is exactly {\color{black}proportional to} $n^{-1}$ and we obtain
for every $n \geq 3$. When ${p > 2}$, $a_n^{-p}$ becomes negligible compared to $n^{-5/4}$ so that
Combining these bounds on $\Delta_{n,\text{E}}$ with the expression of the Edgeworth expansion translates into upper bounds on $\Delta_{n,\text{B}}$ of the form
As soon as the previous \(O(n^{-1})\) term gets smaller than $0.0413 K_{3,n}/\sqrt{n}$, the bound on $\Delta_{n,\text{B}}$ becomes much better than $0.5583 K_{3,n}/\sqrt{n}$ or $0.4690 K_{3,n}/\sqrt{n}$. This can happen even for sample sizes $n$ of the order of a few thousands, assuming that $K_{3,n}$ and $K_{4,n}$ are reasonable (e.g. $K_{4,n} \leq 9$). When $\mathbb{E}[X_i^3]=0$ for every $i = 1,\dots,n$, we remark that $\Delta_{n,\text{B}} = \Delta_{n,\text{E}}$, meaning that we obtain a bound on $\Delta_{n,\text{B}}$ of order $n^{-1}$.
We confirm these rates through a numerical application in Example (ref) for the specific choices $C_0 = 1$ and $p = 2$. These choices are satisfied for common distributions such as the Laplace distribution (for which these values of $C_0$ and $p$ are sharp) and the Gaussian distribution. This actually opens the way for another restriction on the tails of $f_{S_n}$: we could impose $|f_{S_n}(t)| \leq \max_{1 \leq r \leq M}|\rho_r(t)|$ for all $|t| \geq a_n$ and for $(\rho_r)_{r=1, \dots, M}$ a family of known characteristic functions. This second suggestion boils down to a semiparametric assumption on $P_{S_n}$: $f_{S_n}$ is assumed to be controlled in a neighborhood of $\pm\infty$ by the behavior of at least one of the $M$ characteristic functions $(\rho_r)_{r=1, \dots, M},$ but $f_{S_n}$ need not be exactly one of those $M$ characteristic functions. This semiparametric restriction becomes less and less stringent as $n$ increases since we need to control $f_{S_n}$ on a region that vanishes as $n$ goes to infinity. Since $S_n$ is centered and of variance $1$ by definition, the choice of possible $\rho_r$ is naturally restricted to the set of characteristic functions that correspond to such standardized distributions.
We state a second corollary that deals with the i.i.d. framework. We define the following quantity $\kappa_n := \sup_{t: \, |t| \geq a_n/\sqrt{n}} |f_{X_n/\sigma_n}(t)|$ and let $c_n := b_n/a_n$. Under Assumption (ref), we remark that $\sup_{t: \, |t|\geq a_n} |f_{S_n}(t)| = \kappa_n^n.$
Note that for any given $s > 0$ and any random variable $Z$, $\sup_{t: |t| \geq s} |f_Z(t)| = 1$ if and only if $P_{Z}$ is a lattice distribution, i.e., concentrated on a set of the form $\{a + nh, n \in \mathbb{Z} \}$ ushakov2011. Therefore, $\kappa_n < 1$ as soon as the distribution is not lattice, which is the case for any distribution with an absolute continuous component.
In Corollary (ref), the first term on the right-hand side of the inequality as well as $r_{2,n}$ are unchanged compared to Theorem (ref) and Corollary (ref). The second term on the right-hand side of the inequality, $(1.0253/\pi)\kappa_n^n\log(c_n),$ corresponds to an upper bound on the integral term of Equation (ref) in Theorem (ref). Imposing $K_{4,n} \leq K_4$, we can only claim that $1.0253 \, \kappa_n^n\log(c_n)/\pi = O(\kappa_n^n\log n)$, which does not provide an explicit rate on $\Delta_{n,\text{E}}$. If we also assume $\sup_{n\geq 3}\kappa_n<1$ then we can write
and
When is the assumption $\sup_{n\geq 3}\kappa_n < 1$ reasonable? First, it always holds in the i.i.d. setting with a distribution of the \((X_i)_{i = 1, \ldots, n}\) independent of $n$ and continuous. By definition of $a_n$ and by the fact that $\widetilde{K}_{3,n} \geq 1$, $a_n/\sqrt{n}$ is larger than $2 t_1^* \pi$ for $n$ large enough. Consequently, $\kappa_n$ is upper bounded by $\kappa := \sup_{t: \, |t| \geq 2t_1^*\pi} |f_{X_1/\sigma_1}(t)|$ for $n$ large enough. In this case, if $P_{X_1/\sigma_1}$ has an absolutely continuous component, \(\kappa < 1\). For smaller \(n\), we use the fact that \(\kappa_n < 1\) for every \(n\) as explained right after Corollary (ref). The value of $\kappa$ depends on the distribution $P_{X_1/\sigma_1}$. The closer to one $\kappa$ gets, the less regular $P_{X_1/\sigma_1}$ is, in the sense that the latter becomes hardly distinguishable from a lattice distribution.
Second, we could impose that the characteristic function $f_{X_n/\sigma_n}$ be controlled by some finite family of known characteristic functions $\rho_1, \dots, \rho_M$ (independent of $n$) beyond $a_n / \sqrt{n}$. This follows the suggestion mentioned after Corollary (ref), except that we now obtain an exponential upper bound instead of a polynomial one. Indeed, for $n$ large enough, $\kappa_n \leq \kappa := \sup_{t:|t|\geq 2t_1^*\pi}\max_{1\leq m \leq M}|\rho_m(t)|$ and $\kappa < 1$ provided that $(\rho_m)_{m=1,\dots,M}$ are characteristic functions of continuous distributions.
In Example (ref), we plot our bounds on $\Delta_{n,\text{B}}$ by imposing the restriction $\kappa_n \leq 0.99$ which we argue is a very reasonable choice. To justify this claim, we compare our restriction to the value of $\kappa_n$ we would get if $X_n/\sigma_n$ were standard Laplace, a distribution whose characteristic function has much fatter tails than the standard Gaussian or Logistic for instance. In fact, if we were to compute $\sup_{t:|t|\geq 2t_1^*\pi}|\rho(t)|$ with $\rho$ the characteristic function of a standard Laplace distribution, we would get $\kappa_n < 0.11$. Despite our fairly conservative bound on $\kappa_n$, we witness considerable improvements of our bounds compared to those given in Section (ref).
{\color{black} As seen in the previous examples, explicit values or bounds on some functionals of $P_{S_n}$ are required to compute our non-asymptotic bounds on a standardized sample mean. This phenomenon is not unique to our bounds, and arises for any Berry-Esseen- or Edgeworth-type bounds. A value or a bound on \(K_{3,n}\) is indeed required to compute existing Berry-Esseen bounds as in the seminal works of berry1941 and esseen1942 and its recent improvement (e.g. shevtsova2013). Similar to us, recent extensions to these bounds proposed in adell2008, boutsikas2011 and zhilova2020new also depend on several (potentially unknown) moments of the distributions.
Under moment conditions only, the main term and remainder \(r_{1,n}\) of Theorem (ref) solely depend on $\lambda_{3,n}$, \(K_{3,n}\) or \(\widetilde{K}_{3,n}\), and \(K_{4,n}\). As a matter of fact, a bound on \(K_{4,n}\) is sufficient to control all those quantities: pinelis2011relations ensures $|\lambda_{3,n}| \leq 0.621 K_{3,n}$, and a convexity argument yields $K_{3,n} \leq K_{4,n}^{3/4}$ (and remember that \(\widetilde{K}_{3,n}\) is lower than \(2 K_{3,n}\) in the i.n.i.d. case and \(K_{3,n} + 1\) in the i.i.d. case). Having access to a bound on \(K_{4,n}\) is thus crucial to compute our bounds in practice.
First, in some situations, one may rely on auxiliary information about the distribution. In the i.i.d. case in particular, we note that imposing the bound ${K_{4,n} \leq 9}$ allows for a wide family of distributions used in practice: any Gaussian, Gumbel, Laplace, Uniform, or Logistic distribution satisfies it, as well as any Student with at least 5 degrees of freedom, any Gamma or Weibull with shape parameter at least 1. In this case, remember that \(K_{4,n}\) is the kurtosis of \(X_n\), a natural and well-studied feature of a distribution.
In the i.n.i.d. case, $K_{4,n}$ can be rewritten as a weighted average of individual kurtosis. In that respect, the bound ${K_{4,n} \leq 9}$ indicates that, on average, the individual kurtosis are lower than \(9\).
Second, if a bound on \(K_{4,n}\) is not available, a “plug-in” approach remains applicable. The idea is to estimate the moments $\lambda_{3,n}$, $K_{3,n}$ and $K_{4,n}$ by their empirical counterparts in the data (method of moments estimation), and then compute \(\delta_n\) by replacing the unknown needed quantities with those estimates. We acknowledge that this type of “plug-in” approach is only approximately valid, although somewhat unavoidable when bounds on the unknown moments are not given to the researcher.
In addition to the dependence on these moment bounds, Theorem (ref) involves the integral $\int_{a_n}^{b_n} |f_{S_n}(t)| / t \, dt$ that depends on the a priori unknown characteristic function of $S_n$. The application of the resulting Corollaries (ref) and (ref) requires a control on the tail of this characteristic function through the quantities \(C_0\) and \(p\) in the i.n.i.d. case (respectively \(\kappa_n\) in the i.i.d. case), which can be given using expert knowledge of the regularity of the density of $S_n$, as discussed in Section (ref). } It is also possible to estimate the integral directly, for instance using the empirical characteristic function ushakov2011.
{\color{black}
To give a better sense of the accuracy of our results, we perform a comparison between our bounds on \(x \mapsto \mathbb{P}(S_n \leq x)\) and the existing ones shevtsova2013. Indeed, a control \(\delta_n\) on \(\Delta_{n,\text{E}}\) (respectively \(\Delta_{n,\text{B}}\)) naturally yields upper and lower brackets on \(\mathbb{P}(S_n \leq x)\) of the form \(\left[\Phi(x) + \lambda_{3,n} / (6 \sqrt{n}) \times (1 - x^2) \varphi(x)\right] \pm \delta_n\) (respectively \(\Phi(x) \pm \delta_n\)), for any real \(x\). We plot those upper and lower brackets in the i.i.d. framework for three distinct distributions: Student distributions with 5 (Figure (ref)) or 8 (Figure (ref)) degrees of freedom and an Exponential distribution with expectation equal to 1, re-centered to fall in our framework (Figure (ref)). These three distributions are continuous with respect to Lebesgue's measure which allows us to resort to our sharpest i.i.d. bounds, namely those presented in Corollary (ref) (compared to Figures (ref) and (ref), we only report those improved bounds here). On the contrary, remember that the existing bounds shevtsova2013 assume finite third-order moments only; hence, they do not leverage the additional information about skewness and regularity of the considered distributions.
The bound \(\delta_n\) depends on various features of the distribution of \(S_n\). In line with Example (ref), we set \(\kappa = 0.99\), which happens to be a conservative choice with those distributions as \(\kappa = 0.42\) for a Student(df = 8), 0.54 for a Student(df = 5), and 0.63 for the Exponential distributions we consider. In the following comparisons, we focus on the impact of the unknown moments \(K_{4,n}\), \(K_{3,n}\), and \(\lambda_{3,n}\) on the accuracy of our bounds. }
{\color{black}
The Student distributions illustrate the unskewed case, where our bounds use the information \(\lambda_{3,n} = 0\). Figures (ref) and (ref) report several bounds contrasting the suggested practical choice \(K_{4,n} \leq 9\), to deal with the fact that moments are unknown, with the “oracle” bounds where we use the true values of \(\lambda_{3,n}\), \(K_{3,n}\), and \(K_{4,n}\) (computed or approximated by Monte-Carlo). As a comparison, we also report two versions of the existing bound: a “practical” one using \(K_{3,n} \leq K_{4,n}^{3/4} \leq 9^{3/4}\), and an “oracle” version using the true value of \(K_{3,n}\). The kurtosis of a Student distribution is equal to \( 3 + 6 / (\textnormal{df } - 4)\) with \(\textnormal{df} > 4\) its degree of freedom. Therefore, for any Student with at least 5 degrees of freedom, the upper bound \(K_{4,n} \leq 9\) is valid, but all the more conservative as \(\textnormal{df}\) is large. We consider two different values of \(\textnormal{df}\) to assess the loss of accuracy of our bounds when the discrepancy between the actual \(K_{4,n}\) and our suggested default choice of \(9\) increases.
In Figure (ref), we choose \(\textnormal{df} = 8\) so that the true value is \(K_{4,n} = 4.5\) and the proposed bound \(K_{4,n} \leq 9\) is thus conservative. On the contrary, in Figure (ref), because \(\textnormal{df} = 5\), the true value of \(K_{4,n}\) is equal to the suggested choice of \(9\), which becomes sharp. In that respect, it is a more favorable situation. Nonetheless, remark that there remains a difference between the “practical” and “oracle” versions of our bounds: the latter uses the true value of \(K_{3,n}\) (here, approximately equal to \(1.8\)) while the former controls \(K_{3,n}\) by \(9^{3/4} \approx 5.2\).
The Exponential distribution displayed in Figure (ref) illustrates our bounds for a skewed distribution. We choose an Exponential distribution with expectation equal to 1. This distribution has a kurtosis \(K_{4,n} = 9\) so that the main difference with Figure (ref) can be expected to stem from the presence of skewness. In line with the Student case, we report two versions of Shevtsova's bounds and ours, a practical version which uses only the information $K_{4,n} \leq 9$ and an “oracle” one based on knowledge of $\lambda_{3,n}$, $K_{3,n}$ and $K_{4,n}$. We recall that \(\Delta_{n,\text{B}} \neq \Delta_{n,\text{E}}\) when \(\lambda_{3,n} \neq 0\). What is more, the existing bounds (plotted in red) are bounds on \(\Delta_{n,\text{B}}\) whereas ours (in green) originate from a control of \(\Delta_{n,\text{E}}\).
The “oracle” version can be interpreted as a noise-free implementation of the plug-in approach. We remark that oracle versions of existing bounds and ours are twice as accurate as their counterparts which rely on \(K_{4,n} \leq 9\). These oracle bounds use by definition the true values of the moments, and therefore correspond to the most favorable case, in the sense of the tightness of the bounds.
}
We now examine some implications of our theoretical results for the non-asymptotic validity of one-sided statistical tests based on the Gaussian approximation of the distribution of a sample mean using i.i.d. data.
Let $(Y_i)_{i=1, \dots, n}$ be an i.i.d. sequence of random variable with expectation $\mu$, known variance $\sigma^2$ and finite fourth moment with ${K_4 : = \mathbb{E}\left[(Y_n-\mu)^4\right]/\sigma^4}$ the kurtosis of the distribution of $Y_n$. We want to conduct a test of the null hypothesis ${H_0 : \mu \leq \mu_0}$, for some fixed real number $\mu_0$, against the alternative ${H_1 : \mu > \mu_0}$ with a type I error at most $\alpha \in (0,1)$, and ideally equal to $\alpha$. The classical approach to this problem {\color{black}(Gauss test)} amounts to comparing ${S_n = \sum_{i=1}^n X_i / \sqrt{n}}$, where $X_i := (Y_i - \mu_0) / \sigma$, with the $1-\alpha$ quantile of the $\mathcal{N}(0,1)$ distribution, $q_{\mathcal{N}(0,1)}(1-\alpha)$, and reject $H_0$ if $S_n$ is larger. {\color{black}We study this Gauss test in the general non-asymptotic framework without imposing Gaussianity of the data distribution, and we control the difference with respect to normality using the bounds developed in the previous sections.}
{\color{black}
In certain fields such as medicine or economics, researchers routinely set up experiments that seek to answer a specific question on an explained variable \(Y\). The number of individuals included in the experiment has to be carefully justified as large-scale analyses are very costly. This is typically done through the construction of a so-called “pre-analysis plan” which presents the sample size needed to detect a given effect with a pre-specified testing power $\beta \in (0,1)$. In the Gauss test setting considered here, the researcher determines the effect of interest by fixing a particular alternative hypothesis $H_{1, \eta}: \mu = \mu_0 + \sigma \eta$ (with $\mu>\mu_0$). The quantity \(\eta := (\mu - \mu_0) / \sigma\) is a positive number called the effect size that indicates how far away (in terms of standard deviations) the alternative hypothesis is, compared to the null hypothesis $H_0: \mu \leq \mu_0$. Remark that in our framework, $H_{1, \eta}$ is formally the set of all distributions with mean $\mu$, variance $\sigma^2$, that satisfy our additional moment and regularity conditions. $H_{1, \eta}$ can be seen as a nonparametric class of distributions at a fixed distance $\eta$ of the null hypothesis.
Researchers usually rely on an asymptotic normal approximation to infer the sample size needed to detect a given effect at power $\beta$. Our results allow us to bypass this asymptotic approximation and to propose a procedure to choose the sample size $n$ of the experiment such that
for any distribution belonging to the alternative hypothesis space. Any $n$ that satisfies Equation (ref) for all distributions in the alternative hypothesis is called a (non-asymptotic) sufficient sample size for the effect size $\eta$ at power $\beta$.
Observe that
where $X_i := (Y_i - \mu) / \sigma$ are centered with mean $0$ and variance $1$ and $x_n := q_{\mathcal{N}(0,1)}(1-\alpha) - \eta \sqrt{n}$. We remind the reader that the general result from Theorem (ref) or Corollary (ref) implies the following upper and lower bounds for every \(x \in \mathbb{R}\) and \(n \geq 3\),
where $\delta_n$ is the corresponding bound on $\Delta_{n,\text{E}}$. From Equation (ref), we thus obtain
Therefore,
As a consequence, the sample size $n = n_{\eta, \beta}$ defined as the solution of the following equation
is a non-asymptotic sufficient sample size. Note that the same reasoning can be also applied if we only impose an upper bound on $\lambda_{3,n}$. In particular, if we only know $K_{4,n}$, we can use the bound $0.621 K_{4,n}^{3/4}$ and then a sufficient sample size $n$ can be found as the solution to
Numerical applications can be found in Table (ref) which displays the computed sample sizes for different choices of effect sizes $\eta$ and of power $\beta$. In this experiment, we choose $K_{4,n} \leq 9$ and $\kappa \leq 0.99$, as before. We can observe that, as expected, $n_{\eta, \beta}$ increases with $\beta$ and decreases with $\eta$. For $\eta$ large enough, $n_{\eta, \beta}$ becomes approximately constant in $\eta$ as Equation (ref) simplifies to $1 - \delta_n = \beta.$ Conversely, it is also possible to use directly Equation (ref) to compute the power for different effects and sample sizes. The results are displayed in Table (ref).
}
{\color{black} As explained below, the non-asymptotic bounds introduced in Sections (ref) and (ref) can be used to evaluate the actual (for a finite sample size) level of our one-sided test of interest.
Recall that Berry-Esseen-type inequalities aim to bound \(\Delta_{n,\text{B}}\), defined in Equation (ref), the uniform distance between \(\mathbb{P}(S_n \leq \cdot)\) and \(\Phi(\cdot)\). In particular, for a nominal level \(\alpha\), we thus have
where the probability operator is to be understood under any data-generating process such that \(\mu = \mu_0\), to be as close as possible to the alternative hypothesis $H_1$. Either “classical” Berry-Esseen inequalities or ours obtained through an Edgeworth expansion provide bounds on \(\Delta_{n,\text{B}}\) (see the different bounds displayed in Examples (ref) and (ref) in the i.i.d. case). In this context, a bound on \(\Delta_{n,\text{B}}\) is said to uninformative when it is larger than \(\alpha\). Indeed, in that case, we cannot exclude that $\mathbb{P}\!\left( S_n \leq q_{\mathcal{N}(0,1)}(1-\alpha) \right)$ is arbitrarily close to 1, or equivalently, that the probability to reject $H_0$ is arbitrarily close to $0$, and therefore that the test is arbitrarily conservative (type I error arbitrarily smaller than the nominal level \(\alpha\)). We denote by $n_{\max}(\alpha)$ the largest sample size $n$ for which the bound is uninformative. Intuitively, $n_{\max}(\alpha)$ indicates the sample size above which the asymptotic normal approximation to the distribution of \(S_n\) becomes sensible under the assumptions used to bound \(\Delta_{n,\text{B}}\). Indeed, $n_{\max}(\alpha)$ is specific to the bound \(\delta_n\) used, which itself depends on various features of the distribution: number of finite moments, (lack of) skewness, regularity, etc. Table (ref) reports the value of $n_{\max}(\alpha)$ for different Berry-Esseen bounds and usual nominal levels \(\alpha \in \{0.10, 0.05, 0.01\}\).
For each bound, \(n_{\max}(\alpha)\) is decreasing in \(\alpha\). For \(\alpha = 0.01\) in particular, the situation deteriorates strikingly except in the most favorable case of a regular and unskewed distribution. With our bounds, the presence or absence of skewness strongly influences \(n_{\max}(\alpha)\). We also remark that imposing the additional regularity assumption introduced in Section (ref) significantly lowers \(n_{\max}(\alpha)\). }
We explain now that our non-asymptotic bounds on the Edgeworth expansion can be used to detect whether the test is conservative or liberal. This goes one step further than merely checking whether it is arbitrarily conservative or not. Equation (ref) shows that $\mathbb{P}(S_n\leq x)$ belongs to the interval \[ \mathcal{I}_{n,x} := \left[ \Phi(x) + \lambda_{3,n} (1-x^2)\varphi(x) / (6\sqrt{n}) \pm \delta_n \right], \] which is not centered at $\Phi(x)$ whenever $\lambda_{3,n} \neq 0$ and $x \neq \pm \, 1$. The length of the interval does not depend on $x$ and shrinks at speed $\delta_n$. On the contrary, its location depends on $x$. For given nonzero skewness \(\lambda_{3,n}\) and sample size \(n\), the middle point of $\mathcal{I}_{n,x}$ is all the more shifted away from the asymptotic approximation $\Phi(x)$ as \( (1-x^2)\varphi(x)\) is large in absolute value. The function \(x \mapsto (1-x^2)\varphi(x)\) has global maximum at $x=0$ and minima at the points $x \approx \pm \, 1.73$. Consequently, irrespective of $n$, the largest gaps between $\mathbb{P}(S_n\leq x)$ and $\Phi(x)$ may be expected around $x=0$ or $x = \pm \, 1.73$. $\Phi(x)$ could even lie outside $\mathcal{I}_{n,x}$, in which case $\mathbb{P}( S_n \leq x )$ has to be either strictly smaller or larger than $\Phi(x)$. More precisely, \(\mathbb{P}( S_n \leq x )\) is all the further from its normal approximation \(\Phi(x)\) as the skewness \(\lambda_{3,n}\) is large in absolute value; whether \(\mathbb{P}( S_n \leq x )\) is strictly smaller or larger than \(\Phi(x)\) depends on the sign of \(1 - x^2\) as developed in Table (ref).
These observations allow us to quantify possible non-asymptotic distortions between the nominal level and actual rejection rate of the one-sided test we consider. Let us set $x = q_{\mathcal{N}(0,1)}(1-\alpha)$ (henceforth denoted $q_{1-\alpha}$ to lighten notation), which implies that $\Phi(x) = 1- \alpha$. Here, we focus solely on the case $|q_{1-\alpha}|>1$ to encompass all tests with nominal level $\alpha \leq 0.15$, thus in particular the conventional levels 10%, 5%, and 1%. When $\lambda_{3,n} > 6\sqrt{n}\delta_n/\big((q_{1-\alpha}^2-1)\varphi(q_{1-\alpha})\big)$, we conclude that $\mathbb{P}\left( S_n \leq q_{1-\alpha} \right) < 1-\alpha$. Since the event \( \{ S_n \leq q_{1-\alpha} \} \) is the complement of the rejection region, the probability of rejecting $H_0$ under the null exceeds \(\alpha\); in other words, the test cannot guarantee its stated control \(\alpha\) on the type I error and is said liberal. Conversely, when $\lambda_{3,n} < 6\sqrt{n}\delta_n/\big((1-q_{1-\alpha}^2)\varphi(q_{1-\alpha})\big)$, the probability $\mathbb{P}\left( S_n \leq q_{1-\alpha} \right)$ has to be larger than $1-\alpha$; equivalently, the probability to reject under the null is below \(\alpha\) so that the test is conservative.
The distortion can also be seen in terms of p-values. In the unilateral test we consider, the p-value is \({pval := 1 - \mathbb{P}(S_n \leq s_n)}\) with $s_n$ the observed value of $S_n$ in the sample. In contrast, the approximated p-value is ${\widetilde{pval} := 1 - \Phi(s_n)}$. Setting $x = s_n$ in Equation (ref) yields
Therefore,
In line with the explanations preceding Table (ref), \(\widetilde{pval}\) is strictly smaller or larger than \(pval\) when the skewness is sufficiently large in absolute value relative to \(\delta_n\). Indeed, if \(\lambda_{3,n} \neq 0\), the interval from Equation (ref) that contains the true p-value \(pval\) is not centered at the approximated p-value \(\widetilde{pval}\). Under additional regularity assumptions (see Corollary (ref) in the i.i.d. case), the remainder term \(\delta_n = O(n^{-1})\) whereas the “bias” term involving \(\lambda_{3,n}\) vanishes at rate \(n^{-1/2}\). As a result, the interval locates closer to \(\widetilde{pval}\) as $n$ increases and its width shrinks to zero at an even faster rate.
Finally, we stress that such distortions regarding rejection rates and p-values are specific to one-sided tests. For bilateral or two-sided tests, the skewness of the distribution enters symmetrically in the approximation error and cancels out thanks to the parity of \( x \mapsto (1 - x^2) \phi(x) \).