EconBase
← Back to paper

On the construction of confidence intervals for ratios of expectations

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

80,036 characters · 18 sections · 18 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

On the construction of confidence intervals for ratios of expectations

titlepage\thispagestyle{empty} \begin{abstract} In econometrics, many parameters of interest can be written as ratios of expectations. The main approach to construct confidence intervals for such parameters is the delta method. However, this asymptotic procedure yields intervals that may not be relevant for small sample sizes or, more generally, in a sequence-of-model framework that allows the expectation in the denominator to decrease to $0$ with the sample size. In this setting, we prove a generalization of the delta method for ratios of expectations and the consistency of the nonparametric percentile bootstrap. We also investigate finite-sample inference and show a partial impossibility result: nonasymptotic uniform confidence intervals can be built for ratios of expectations but not at every level. Based on this, we propose an easy-to-compute index to appraise the reliability of the intervals based on the delta method. Simulations and an application illustrate our results and the practical usefulness of our rule of thumb. Keywords: delta method, confidence regions, uniformly valid inference, sequence of models, nonparametric percentile bootstrap. MSC: Primary 62F25; secondary 62F40, 62P20. JEL: C18, C19. \end{abstract}

Introduction

In applied econometrics, the prevalent method for constructing confidence intervals (CIs) is asymptotic: the theoretical guarantees for most CIs used in practice hold only when the number of observations tends to infinity. For a large class of parameters, the construction of asymptotic CIs also relies on the delta method. In this paper, we focus on parameters that can be expressed as ratios of expectations for which the delta method is a standard procedure to conduct inference. The objective is twofold: study the behavior of the delta method and other confidence intervals in some difficult settings and provide tools to detect cases in which the delta method may behave poorly.

Many popular parameters in economics take the form of ratios of expectations. Typical examples are conditional expectations since any conditional expectation with a discrete conditioning variable, or a conditioning event, can be written as a ratio of unconditional expectations. For instance, assume that we observe an independent and identically distributed (i.i.d.) sample of individuals indexed by $i \in \{1,\ldots,n\}$ with $W_i$ the wage of an individual and $D_i$ an indicator equal to $1$ whenever individual $i$ belongs to some treatment group, say a training program; $0$ otherwise. Suppose you are interested in the average wage of participants in the program. We have $\mathbbm{E}\left[W \mid D=1\right] = \mathbbm{E}\left[W D\right] /\,\mathbbm{E}\left[D\right]$ as $D$ is binary.

Most confidence intervals used in practice are based on asymptotic justifications, hence possible concerns as regards their finite-sample reliability. For ratios of expectations, we document this issue on simulations (see Section (ref)). One of our findings is that the coverage of the CIs based on the delta method happens to be far below their nominal level, even for large sample sizes, when the expectation in the denominator is close to $0$.\footnote{The definitions of coverage and other fundamental properties of confidence intervals are recalled in Appendix (ref) with the conventions that we use.} For some scenarios, these asymptotic CIs require above 100,000 observations to get reasonably close to their nominal level. Yet, denominators close to $0$ are not unusual in practice. Coming back to the treatment/wage example, a small denominator would correspond to a binary treatment with a low participation rate.

In order to deal with that issue, we consider sequences of models, namely we authorize the distribution of the observations to change with the sample size. This framework enables to formalize in an asymptotic way the idea of a denominator close to $0$. Indeed, in a standard asymptotic viewpoint, with the expectation in the denominator different from $0$, all parameters are fixed and well-defined. Hence, $n$ always grows large enough so that empirical means are close to their expectations and the CIs based on the delta method are valid. In other words, the signal that we want to estimate is constant while the noise goes to $0$, and therefore the problem vanishes in this asymptotic perspective. We would like to model more difficult cases, in which the signal can go to $0$ as well. This is precisely what the sequence-of-model set-up allows.\footnote{This can also rationalize the practice of applied social researchers (see Example (ref)). The heuristic idea is that researchers can consider narrower effects as the data gets richer.} This is similar to some frameworks that have been developed for weak instrumental variables (IV), see notably staiger1997,stock2005testing,andrews2018_weak.

In this literature, another approach does not consider sequences of models but designs “robust” procedures that allow to be exactly in the problematic case, namely a null covariance between the instrument and the endogenous regressor (see anderson1949estimation). In this case, the parameter of interest is unidentified. In contrast with the weak IV framework, it is worth noting that for ratios in general the parameter of interest is not even defined when the denominator is exactly equal to $0$. As a consequence, such an approach seems difficult to extend to our problem.

In our setting, it is unclear, even asymptotically, what the properties of the CIs based on the delta method are when the expectation in the denominator tends to $0$. We show that usual CIs can fail and the limiting law of ${\widehat{\theta}_n - \theta_n}$ may not be Gaussian anymore, denoting by $\theta_n$ the ratio of expectations and $\widehat{\theta}_n$ its empirical counterpart. In some cases, the difference ${\widehat{\theta}_n - \theta_n}$ may actually have a Cauchy limit, as can be found in the weak IV literature.

We show in this sequence-of-model framework that confidence intervals provided by the nonparametric percentile bootstrap have the same asymptotic properties as the ones obtained with the delta method. Simulations support that claim and even suggest the former have better coverage than the latter in finite samples.

Even in standard settings with a fixed but small denominator, simulations document that asymptotic-based CIs may require very large sample sizes to attain their nominal level. This suggests to study more in details nonasymptotic inference. More precisely, we construct finite-sample CIs, extending old-established concentration inequalities for means to ratios of means. Concentration inequalities for the mean refer to upper bounds on the probability that an empirical mean departs from its expectation more than a given threshold. Such inequalities permit to construct confidence intervals valid for any sample size and for large classes of probability distributions (see in particular boucheron2013concentration). To our knowledge, there is no such result for ratios. We consider distributions within a class characterized by a lower bound on the first moment for the denominator variable, and an upper bound on the second moment for both the numerator and denominator variables.\footnote{We refer to this setting as the “Bienaymé-Chebyshev” (BC) case. In Appendix (ref), we present similar results for distributions whose supports are bounded (“Hoeffding” case).}

One additional result highlights there exists a critical confidence level, above which it is not possible to construct nonasymptotic CIs, uniformly valid on such classes, and that are almost surely bounded under every distribution of those classes. More precisely, we exhibit explicit upper and lower bounds on this critical confidence level: the former is a threshold above which we show it is impossible to construct such CIs; the latter is a threshold below which we show how to construct them.

These ideas closely relate to some impossibility results as regards the construction of confidence intervals. A large share of the research effort has concentrated on the problem of constructing confidence intervals for expectations. In an early contribution, bahadur1956 show that, when $\mathcal{P}$ is the set of all distributions on the real line with finite expectation, the parameter of interest $\theta(P)$ is the expectation with respect to a distribution ${P \in \mathcal{P}}$ and $\Theta=\mathbb{R}$, a confidence interval built from an i.i.d. sample of ${n \in \mathbbm{N}^{*}}$ observations that has uniform coverage ${1-\alpha}$ over $\mathcal{P}$ must contain any real number with probability at least ${1-\alpha}$. Broadly speaking, any confidence interval must have infinite length with positive probability for every $P\in\mathcal{P}$ to ensure a coverage of ${1-\alpha}$.

Stronger results can be derived when one further restricts $\mathcal{P}$ or $\Theta$. When $\mathcal{P}$ is taken to be the set of all distributions on the real line with variance uniformly bounded by a finite constant, it is possible to show (using the Bienaym\'e-Chebyshev inequality) that for every $n \in \mathbbm{N}^{*}$ and every $\alpha\in(0,1)$, there exists a confidence interval that is almost surely bounded under every $P\in\mathcal{P}$ and has coverage $1-\alpha$. In this case, the obtained CIs have the advantage that their length shrinks to $0$ at the optimal rate $1/\sqrt{n}$. But on the downside, they are not of size $1-\alpha$, even asymptotically, except for some extreme distributions. This means that they tend to be conservative in practice.

A strand of the literature has also investigated more complex problems in which $\theta(P)$ is not restricted to being an expectation. For general parameters, dufour1997 derives a generalization of bahadur1956. An implication of the results in dufour1997 is the existence of an impossibility theorem for ratios of expectations. Let $P$ be a distribution on $\mathbb{R}^2$ with marginals $P_X$ and $P_Y$. If $\theta(P)=\mathbbm{E}_{P_X}\left[X\right] / \, \mathbbm{E}_{P_Y}\left[Y\right]$, then for every $\alpha\in(0,1)$, it is impossible to build nontrivial CIs of coverage $1-\alpha$ when $\mathcal{P}$ is the set of all distributions on $\mathbb{R}^2$ with finite second moments and $\Theta = \left\{\theta=\mathbbm{E}_{P_X}\left[X\right] / \, \mathbbm{E}_{P_Y}\left[Y\right]:\left(\mathbbm{E}_{P_X}\left[X\right],\mathbbm{E}_{P_Y}\left[Y\right]\right)\in\mathbb{R}\times\mathbb{R}^*\right\}$. As will be explained below, this impossibility result disappears as soon as $\mathcal{P}$ is chosen such that $\left|\mathbbm{E}_{P_Y}\left[Y\right]\right|$ is bounded away from $0$ uniformly over $\mathcal{P}$. Interestingly, the impossibility breaks down only partly in the sense that there remains an upper bound on confidence levels (that depends on $n$) above which it is impossible to build nontrivial CIs.

Other interesting results can be found in romano2000 and pinelis2016. romano2000 construct nonasymptotic valid confidence intervals that happen to be also asymptotically optimal. However, they only consider expectations. pinelis2016 study smooth functions of a vector of means and give bounds on the distance between the distribution of the normalized and centered estimator and its Gaussian limiting distribution. Nonetheless, the authors do not link their results to the construction of confidence intervals.

In the light of that existing literature, our nonasymptotic findings can be interpreted as a partial impossibility result. Indeed, even if we assume a known positive lower bound on the expectation in the denominator, the limitation on the attainable coverage of our nonasymptotic CIs remains. That point complements dufour1997: for a given sample size $n$, interesting CIs can be built but not at every confidence level. By contrast, provided the expectation in the denominator is not null, the delta method gives CIs at every confidence level, but their coverage is only asymptotic.

To bridge this gap, we suggest a rule of thumb to assess the reliability of the delta method for ratios of expectations in finite samples. The heuristic idea is simply, for a given sample, to compute an estimator of the lower bound on the above-mentioned critical confidence level. This lower bound can be seen as a conservative value for the unknown critical level, which is a necessary criterion to conduct valid inference in finite samples uniformly over a given class of distributions. Hence, for any desired level higher than this bound, the CIs based on the delta method cannot reach this desired uniform level in finite samples. We illustrate the empirical usefulness of that rule of thumb on simulations and with an application to gender wage disparities in France for the years 2010-2017.

The rest of the paper is organized as follows. Section (ref) details our framework and assumptions. In Section (ref), we illustrate the weaknesses of the CIs based on the delta method with a denominator “close to 0” on simulations and detail the asymptotic behavior of the delta method and of the nonparametric percentile bootstrap in our sequence-of-model setting. Section (ref) is devoted to the construction of nonasymptotic confidence intervals and presents a lower bound on the aforementioned critical confidence level. In Section (ref), we derive an upper bound on the critical confidence level as well as a lower bound on the length of nonasymptotic CIs. This section also includes the description of a practical index to gauge the soundness of the CIs based on the delta method in finite samples. Section (ref) present simulations and an application to a real dataset to illustrate our methods. Section (ref) concludes. General definitions about confidence intervals are recalled in Appendix (ref). The proofs of all results are postponed to Appendix (ref). Additional results under an alternative set of assumptions (“Hoeffding” case) are detailed in Appendix (ref). Appendix (ref) presents supplementary simulations.

Our framework

Throughout the paper, for any random variable $U$ and $n$ i.i.d. replications $(U_{1,n},\ldots,U_{n,n})$, we denote by $\overline{U}_n$ the empirical mean of $U$, that is ${n^{-1} \sum_{i=1}^n U_{i,n}}$. Assumption (ref) defines our sequence-of-model framework and provides the basic requirements to state our asymptotic results.

hypFor every ${n \in \mathbbm{N}^{*}}$, we observe a sample $(X_{i,n}, Y_{i,n})_{i=1,\ldots,n} \buildrel {\text{i.i.d.}} \over \sim P_{X,Y,n}$, where $P_{X,Y,n}$ is a given distribution on $\mathbb{R}^2$ that satisfies $\mathbbm{E}[Y_{1,n}] > 0$, $\mathbbm{E}[X_{1,n}^2] < + \infty$, and $\mathbbm{E}[Y_{1,n}^2] < + \infty$.

Remark that $n$ indexes both the distribution $P_{X,Y,n}$ of the observations in this model and the number of observations $n$. This encompasses the standard i.i.d. set-up if the distribution does not change with $n$: for every ${n \in \mathbbm{N}^{*}}$, ${P_{X,Y,n} = P_{X,Y}}$ for some given distribution $P_{X,Y}$. As we assume the existence of a finite expectation, we can consider ${\mathbbm{E}[Y_{1,n}] \geq 0}$ without loss of generality.\footnote{Otherwise, we simply replace $Y_{i,n}$ by its opposite $-Y_{i,n}$.} In order to have properly defined ratios of interest, we need to assume away a null denominator, namely suppose that for every ${n \in \mathbbm{N}^{*}}$, ${\mathbbm{E}[Y_{1,n}] > 0}$.

example[Sequences of models and the practice of applied researchers] \newline Researcher may look at the average value of a variable $A_{i,n}$ of interest in a subgroup of the data. Subgroups could be defined as the intersections of, say, time, geographical area, gender, age, income brackets and so on. As the number of observations $n$ grows, it is possible to consider subgroups $g_n$ that become thinner and thinner (intersection of more and more variables for instance). This practice could be modelled as estimating $\theta_n := \mathbbm{E}\left[A_{i,n} \mid G_{i,n} = 1\right] = \mathbbm{E}\left[A_{i,n} G_{i,n}\right] / \, \mathbbm{P}\left(G_{i,n}=1\right)$ where $G_{i,n}$ is a binary variable that is equal to 1 if an individual $i$ belongs to the subgroup $g_n$. This corresponds to our framework denoting $X_{i,n} := A_{i,n} \times G_{i,n}$ and $Y_{i,n} := G_{i,n}$.

To derive our nonasymptotic results, Assumption (ref) has to be strengthened.

hypFor every ${n \in \mathbbm{N}^{*}}$, there exist positive finite constants $l_{Y,n}$, $u_{X,n}$, and $u_{Y,n}$ such that (i) ${\mathbbm{E}[Y_{1,n}] \geq l_{Y,n} > 0}$, (ii) $\mathbbm{E}[X_{1,n}^2] \leq u_{X,n}$ and $\,\mathbbm{E}[Y_{1,n}^2] \leq u_{Y,n}$.

Note that in practice, the value of the constants $l_{Y,n}$, $u_{X,n}$, and $u_{Y,n}$ may not be available for practitioners. This is the reason why, in Section (ref), we propose heuristic methods that palliate the lack of knowledge of those constants.

The first part of the assumption bounds the expectation of $Y_{1,n}$ away from $0$ while the second states that the second moments of $X_{1,n}$ and $Y_{1,n}$ are bounded. These are necessary to derive nonasymptotic CIs with maintained coverage uniformly over a class of distributions and that are not trivial. Otherwise, if ${l_{Y,n}=0}$ or in the absence of the upper bounds $u_{X,n}$ and $u_{Y,n}$, the impossibility theorem of dufour1997 applies and prevents from constructing nontrivial CIs for any confidence level. In a way, given this result, Assumption (ref) can be seen as close to the minimal hypothesis that allows for the possibility of nontrivial confidence intervals with finite-sample guarantees for ratios of expectations. Furthermore, the sequence-of-model framework allows $l_{Y,n}$ to decrease to $0$, which enables us to study limiting cases close to but different from the problematic case ${l_{Y,n}=0}$.

This set-up, where Assumptions (ref) and (ref) hold, is named the BC case since it is possible under these assumptions to construct nonasymptotic CIs using the Bienaym\'e-Chebyshev inequality. In Appendix (ref), we present an adapted version of our results under the assumption that $X_{1,n}$ and $Y_{1,n}$ have a bounded support instead of bounded second moments; a setting we call the Hoeffding case.

To sum up, Assumptions (ref) and (ref) define a set $\mathcal{P}$ of distributions for some constants $l_{Y,n}$, $u_{X,n}$ and $u_{Y,n}$. For a distribution $P_{X,Y,n}$ in $\mathcal{P}$, the parameter of interest $\theta(P_{X,Y,n})$ is denoted ${\theta_n := \mathbbm{E}[X_{1,n}] /\,\mathbbm{E}[Y_{1,n}]}$ with values in $\mathbb{R}$. To estimate this parameter, we consider its empirical counterpart ${\widehat{\theta}_n := \overline{X}_n /\,\overline{Y}_n}$. We seek to construct confidence intervals $C_{n,\alpha}$ for ${\theta_n}$ with nominal level ${1-\alpha}$ based on this estimator.

In practice, it is possible that ${\overline{Y}_n = 0}$ and it may even happen with a strictly positive probability for non-continuous distributions of $Y$. The estimator $\widehat{\theta}_n$ does not exist for such samples. In such a case, it is difficult to construct meaningful confidence intervals. Different conventions are possible:

itemize• We could choose to define ${C_{n,\alpha} = \mathbb{R}}$. This entails that $\theta_n$ belongs to $C_{n,\alpha}$ by construction. We believe that such a choice would artificially improve the coverage of $C_{n,\alpha}$ as it induces that the higher $\mathbbm{P}(\overline{Y}_n = 0)$, the better the interval in terms of coverage. • We could choose ${C_{n,\alpha} = \emptyset}$. The hypothesis ${\theta_n = \theta_0}$ would then be rejected for every ${\theta_0 \in \mathbb{R}}$ using the duality between tests and confidence intervals. We would also like to avoid this situation because it may not be reasonable to always reject for the mere reason that $\theta_n$ cannot be estimated in the sample. • Other choices are possible, for example ${C_{n,\alpha} = \{0\}}$, but they do not seem sensible either since there is no reason to select only $0$ in our confidence interval, especially if ${\overline{X}_n \neq 0}$.

For these considerations, we choose to let $C_{n,\alpha}$ undefined whenever ${\overline{Y}_n = 0}$, following the convention that ratios $x/0$ are undefined for any real $x$.\footnote{When facing ${\overline{Y}_n = 0}$, applied researchers may use other estimators. For instance, one could consider sub-samples (possibly several and combine them in some way) of the data for which the empirical mean in the denominator differs from $0$. Nevertheless, the construction of satisfactory estimators in this case lies beyond the scope of this paper.} In practice, when given a realization $\omega \in \Omega$ and a real $a \in \mathbb{R}$, we either know that $a$ belongs to $C_{n, \alpha}(\omega)$, or we know that $a$ does not belong to $C_{n, \alpha}(\omega)$, or $C_{n, \alpha}(\omega)$ is undefined. As a consequence, we have the decomposition $\Omega = \{\omega: a \in C_{n, \alpha}(\omega) \} \sqcup \{\omega: a \notin C_{n, \alpha}(\omega) \} \sqcup \{\omega: C_{n, \alpha}(\omega) \text{ undefined}\},$ where $\sqcup$ denotes the disjoint union of sets. This means that $\mathbbm{P} \{ a \in C_{n, \alpha} \} + \mathbbm{P} \{ a \notin C_{n, \alpha} \} + \mathbbm{P} \{ C_{n, \alpha} \text{ undefined}\} = 1$.

Limitations of the delta method: when are asymptotic confidence intervals valid?

In practice, for a sample of size $n$, the coverage of asymptotic CIs may be well below their nominal level ${1-\alpha}$. Intuitively, this phenomenon should be driven by “problematic” distributions in $\mathcal{P}$ in the following sense: when the true distribution $P$ is close to the boundary of the class $\mathcal{P}$, the probability $c(n,P) := \mathbbm{P}_{P^{\otimes n}}\left(C_{n,\alpha} \ni \theta(P)\right)$ may be much smaller than ${1-\alpha}$.\footnote{Recall that in the nonasymptotic approach, the coverage of any given confidence interval $C_{n,\alpha}$ is defined as the infimum of $c(n,P)$ for $P$ ranging over the studied class $\mathcal{P}$ of distributions.}

In Section (ref), with $C_{n,\alpha}$ the confidence interval based on the delta method, we illustrate on simulations that $c(n,P)$ can fail to match ${1-\alpha}$ when the expectation in the denominator is fixed close to $0$. In other words, it may require a very large number of observations to make reasonable the asymptotic approximation. In Section (ref), we investigate a more serious issue: in the sequence-of-model framework, we let the expectation in the denominator not only be small but converge to $0$ as $n$ increases. We show on simulations that depending on the speed at which the denominator goes to $0$, $c(n,P)$ can either converge to the nominal level (more or less quickly) or even not converge at all to this target. This sheds light on a partial failure of the delta method when the denominator goes to $0$ that we derive formally in Section (ref). Finally, in Section (ref), we show the asymptotic consistency of the nonparametric percentile bootstrap (also known as Efron's percentile bootstrap) in this sequence-of-model framework.

Asymptotic approximation takes time to hold

In this subsection, we consider the i.i.d. case.\footnote{For every $n \in \mathbbm{N}^{*}$, $P_{X,Y,n}$ is identical, hence denoted $P_{X,Y}$. To simplify notations, we also denote by $(X,Y)$ a random vector following $P_{X,Y}$.} Under Assumption (ref), asymptotic confidence intervals are easily obtained combining the multivariate central limit theorem (CLT) and the delta method:

equation[equation omitted — 259 chars of source]

where ${\Sigma = \mathbbm{V}[X]/ \,\mathbbm{E}[Y]^2 + \mathbbm{E}[X]^2\mathbbm{V}[Y] / \,\mathbbm{E}[Y]^4 - 2 \, \mathbbm{C}ov\left[X,Y\right] \mathbbm{E}[X] / \,\mathbbm{E}[Y]^3}$ and in practice is replaced by a consistent estimate (Slutsky's lemma).

To assess the quality of the CI based on (ref), we compute its $c(n,P)$ using simulations for different sample sizes $n$ and distributions $P$ and compare it to the nominal level. By definition, the pointwise coverage $c(n,P)$ forms an upper bound on the uniform coverage. In our simulations, we choose the level ${1-\alpha = 95\%}$. For different sample sizes $n$ and values of $\mathbbm{E}[Y]$, we draw $M=$ 5,000 i.i.d. samples of size $n$ following $\mathcal{N}(1,1) \otimes \mathcal{N}(\mathbbm{E}[Y],1)$. We compute $c(n,P)$ for the interval based on the delta method for every pair $(n, \, \mathbbm{E}[Y])$ using the 5,000 replications. The expectation $\mathbbm{E}[Y]$ ranges from $0.01$ (the denominator is close to $0$) to $0.75$ (the denominator is far from $0$). Figure (ref) sums up the results. For every $n$, it turns out that the closer $\mathbbm{E}[Y]$ to $0$, the smaller the $c(n,P)$ of the delta method. When $\mathbbm{E}[Y]=0.01$, we observe that $c(n,P)$ gets close to the nominal level only for $n$ above 300,000. Additional simulations indicate that the phenomenon is robust across different choices of the distribution $P_{X,Y}$ (see Section (ref)).

figure[figure omitted — 575 chars of source]

Asymptotic results may not hold in the sequence-of-model framework

Unlike the result displayed in (ref), it is unclear how $\sqrt{n}\left(\overline{X}_n/\,\overline{Y}_n-\mathbbm{E}[X]/\,\mathbbm{E}[Y]\right)$ behaves asymptotically when we consider sequences of models such that the expectation in the denominator tends to $0$ as $n$ increases. For a given specification, Figure (ref) shows the $c(n,P)$ of the CIs based on the delta method when ${\mathbbm{E}[Y_{1,n}] = Cn^{-b}}$ where $C$ is set to $0.025$ and $b$ varies. For a speed ${b \geq 1/2}$ (i.e. faster than the usual rate of the CLT), the pointwise coverage $c(n,P)$ of the asymptotic CIs obtained by (ref) is not good in the sense that it is far lower than the nominal level $1-\alpha$ and it does not converge to the latter. Our simulations even suggest that the coverage tends to $0$ for $b>1/2$. For $b<1/2$, the upper bound $c(n,P)$ on the coverage of the delta method seems to tend to $1-\alpha$. Yet, in line with Figure (ref), the validity of the asymptotic approximation requires very large sample sizes.

figure[figure omitted — 616 chars of source]

At this stage, Figure (ref) presents some evidence that the CIs based on the delta method need to be adapted for sequences of models and that the rate of decrease toward $0$ of the expectation $\mathbbm{E}[Y_{1,n}]$ matters. The next subsection details formal results in this set-up.

Extension of the delta method for ratios of expectations in the sequence-of-model framework

We are interested in the asymptotic distribution, as $n$ tends to infinity, of the real random variable $S_n := \sqrt{n} \left({\overline{X}_n}/\,{\overline{Y}_n} - {\mathbbm{E}[X_{1,n}]}/\,{\mathbbm{E}[Y_{1,n}]} \right)$. The following theorem states the asymptotic behavior of $S_n$ according to the comparison of $\mathbbm{V}[Y_{1,n}] \, /\sqrt{n}$ and $\mathbbm{E}[Y_{1,n}]$ under a multivariate Lyapunov condition. It is proved in Section (ref).

We show that in some cases $|S_n| \underset{n\to +\infty}{\overset{a.s.}{\longrightarrow}} + \infty$. It is then impossible to state the limiting distribution $S_n$ in the traditional sense. Despite that, we can still get a more precise result looking at the subsequent terms in the asymptotic expansion of $S_n$. Such an asymptotic expansion is complicated to state, especially in our sequence-of-model framework, since the distributions $P_{X,Y,n}$ change with $n$ without any link from one to the next. To overcome this problem, we consider equivalents in distribution of $S_n$ in the following sense. We say that two sequences of random variables $S_n$ and $T_n$ are equivalent in distribution if there exist a probability space $\tilde \Omega$ and two sequences of random variables $\tilde S_n, \tilde T_n$ such that $\forall n \in \mathbbm{N}^{*}$, $S_n \stackrel{d}{=} \tilde S_n$ and $T_n \stackrel{d}{=} \tilde T_n$, and $\tilde S_n$ is equivalent to $\tilde T_n$ almost surely as $n \to \infty$. This means that for almost every $\tilde \omega \in \tilde \Omega$, $\tilde S_n(\tilde \omega)$ is equivalent to $\tilde T_n(\tilde \omega)$ (considered as deterministic sequences of real numbers). This notion enables to formalize the link between $S_n$ and a simpler expression $T_n$.

thmLet Assumption (ref) hold and (i) $\mathbbm{V}[(\gamma_{X,n} X_{1,n} \, , \, \gamma_{Y,n} Y_{1,n})] \to V$ as ${n \to \infty}$ for some positive sequences $\{\gamma_{X,n}\}_{n \in \mathbbm{N}^{*}}$ and $\{\gamma_{Y,n}\}_{n \in \mathbbm{N}^{*}}$ where $V$ is a definite positive $2 \times 2$ matrix, (ii) $\sup_{n\in\mathbbm{N}^{*}}\mathbbm{E}\left[|X_{1,n}|^3 \gamma_{X,n}^3 + |Y_{1,n}|^3 \gamma_{Y,n}^3\right] < +\infty$, and (iii) $\mathbbm{P}(\overline{Y}_n = 0) \to 0$ as $n \to \infty$. Denote the signal-to-noise-ratio by ${\text{SNR}_n := \mathbbm{E}[Y_{1,n}] / (V_{2,2}^{1/2} n^{-1/2} \gamma_{Y,n}^{-1})}$. Then, the sequence of random variables $S_n := \sqrt{n} \left(\overline{X}_n/\,\overline{Y}_n - \mathbbm{E}[X_{1,n}]/\,\mathbbm{E}[Y_{1,n}] \right)$ satisfies as ${n \to \infty}$: \begin{enumerate} • If ${\text{SNR}_n \to + \infty}$, then $S_n$ is equivalent in distribution to: \begin{equation*} \frac{\sqrt{n} \gamma_{X,n} (\overline{X}_n - \mathbbm{E}[X_{1,n}])}{\mathbbm{E}[Y_{1,n}] \gamma_{X,n}} - \frac{\sqrt{n} \gamma_{Y,n} (\overline{Y}_n - \mathbbm{E}[Y_{1,n}]) \mathbbm{E}[X_{1,n}]}{\mathbbm{E}[Y_{1,n}]^2 \gamma_{Y,n}}. \end{equation*} • If there exists a finite constant ${C \neq 0}$ such that ${\text{SNR}_n \to C}$, then $S_n$ is equivalent in distribution to: \begin{align*} n \gamma_{Y,n} \mathbbm{E}[X_{1,n}] &\Bigg( \frac{1}{C + \sqrt{n} \gamma_{Y,n} (\overline{Y}_n - \mathbbm{E}[Y_{1,n}])} - \frac{1}{C} \Bigg) \quad \\ & \qquad \qquad + \frac{n \gamma_{X,n} (\overline{X}_n - \mathbbm{E}[X_{1,n}]) \times \gamma_{Y,n}} {\big(C + \sqrt{n} \gamma_{Y,n} (\overline{Y}_n - \mathbbm{E}[Y_{1,n}]) \big) \times \gamma_{X,n}}. \end{align*} • If ${\text{SNR}_n \to 0}$, then $S_n$ is equivalent in distribution to: \begin{equation*} \sqrt{n} \left( \frac{\sqrt{n} \gamma_{X,n} (\overline{X}_n - \mathbbm{E}[X_{1,n}])} {\sqrt{n} \gamma_{Y,n} (\overline{Y}_n - \mathbbm{E}[Y_{1,n}])} \times \frac{\gamma_{Y,n}}{\gamma_{X,n}} - \frac{\mathbbm{E}[X_{1,n}]}{\mathbbm{E}[Y_{1,n}]} \right). \end{equation*} \end{enumerate}
figure[figure omitted — 1,630 chars of source]

Theorem (ref) can thus be interpreted as a generalization of the result given by the CLT and the delta method for ratios of expectations. The sequence-of-model framework allows both the expectation and the variance in the denominator to tend to $0$. In particular, this happens whenever $Y_{i,n}$ follows a Bernoulli distribution with a parameter $p_n$ tending to $0$, as detailed in Example (ref). For instance, when we estimate a conditional expectation with a discrete conditioning variable or a conditioning event, the denominator is an average of indicator variables that follow a Bernoulli distribution. Figure (ref) and its companion table highlight the different asymptotic regimes depending on the behaviors of $\{\mathbbm{E}[X_{1,n}]\}_{n \in \mathbbm{N}^{*}}$, $\{\mathbbm{E}[Y_{1,n}]\}_{n \in \mathbbm{N}^{*}}$, $\{\gamma_{X,n}\}_{n \in \mathbbm{N}^{*}}$ and $\{\gamma_{Y,n}\}_{n \in \mathbbm{N}^{*}}$.

The main takeaway of Theorem (ref) is that when $\mathbbm{E}[X_{1,n}]=C_1/n^a$, $\mathbbm{E}[Y_{1,n}]=C_2/n^b$ and $\mathbbm{V}[Y] = C_3 / n^{b'}$ for some constants $C_1, C_2, C_3 \neq 0$, and $b<1/2+b'$, $S_n$ properly renormalized by $n$ to some power still converges in distribution to a Normal random variable. This can be explained using the signal-to-noise ratio (SNR) defined in Theorem (ref). Indeed, in this first case, the $\text{SNR}_n$ tends to $+\infty$: the signal in the denominator (that is the expectation of $Y_{1,n}$) is asymptotically bigger than the noise (which is $1/(\gamma_{Y,n} n^{1/2})$ up to a constant factor). Asymptotic inference based on the Normal approximation remains valid, even if the length of such confidence intervals may not decrease with the sample size $n$.

In all other cases, when the noise dominates in the denominator, $S_n$ converges weakly to a non-Gaussian distribution, in some cases to a generalized Cauchy distribution with parameters that depend on the data generating process (up to a normalization of some power of $n$). By construction, when the noise dominates, we do not have much information and thus may not be able to conduct inference in these settings. This echoes the impossibility results presented in Section (ref), notably Remark (ref). In the next section, we provide another method for constructing confidence intervals using the nonparametric percentile bootstrap.

exampleWhen $Y_{1,n}$ follows a Bernoulli distribution with parameter $p_n$ in $(0,1)$, we are always in the first case of Theorem (ref), meaning that its expectation $p_n$ is always larger than the noise $\sqrt{p_n (1-p_n)/n}$. This latter formula is obtained by remarking that the standard deviation of $Y_{i,n}$ is $\sqrt{p_n (1-p_n)}$ so that $\gamma_{Y,n} = 1/\sqrt{p_n (1-p_n)}$. However, in order to satisfy the constraint $\mathbbm{P}(\overline{Y}_n = 0) \to 0$, we have to impose that $n p_n \to + \infty$. Therefore, when $p_n = n^{-b}$, confidence intervals based on the delta method will be pointwise consistent if $b < 1$.

Validity of the nonparametric bootstrap for sequences of models

In this part, we construct confidence intervals for ratios of expectations using Efron's percentile bootstrap. This technique relies on the nonparametric bootstrap resampling scheme that we now recall. We fix a number $B > 0$ of bootstrap replications. For a given initial sample $(X_{i,n}, Y_{i,n}), i=1, \dots, n$, and a given integer $b$ smaller than $B$, we define the bootstrapped sample $(X_{i,n}^{(b)}, Y_{i,n}^{(b)}), i=1, \dots, n$, which is obtained by $n$ i.i.d. resampling from the initial sample, i.e. with replacement. Let $\overline{X}_n^{(b)} := n^{-1} \sum_{i=1}^n X_{i,n}^{(b)}$ be the empirical mean of the numerator in the $b$-th bootstrapped sample (resp. $\overline{Y}_n^{(b)}$ for the denominator).

Then, Efron's percentile bootstrap, also known as the nonparametric percentile bootstrap, consists in using the quantiles of the bootstrapped distribution conditional on the data to conduct inference. More precisely, for every $\tau \in (0,1)$, let $q_\tau^{boot}$ denote the quantile at level $\tau$ of $\overline{X}_n^{(1)} / \, \overline{Y}_n^{(1)}$, which is estimated in practice by the empirical quantile at level $\tau$ of the bootstrapped statistics $\big( \overline{X}_n^{(b)} / \, \overline{Y}_n^{(b)}\big)_{b = 1,\ldots,B}$. For a given nominal level ${1-\alpha \in (0,1)}$, the confidence interval we consider is defined as $C_{n,\alpha}^{boot} := \big[ q_{\alpha/2}^{boot} \, , \, q_{1-\alpha/2}^{boot} \big]$. The following theorem states the consistency of this interval. It is proved in Section (ref).

thmLet Assumption (ref) hold and (i) $\mathbbm{V}[(\gamma_{X,n} X_{1,n} \, , \, \gamma_{Y,n} Y_{1,n})] \to V$ as ${n \to \infty}$ for some positive sequences $\{\gamma_{X,n}\}_{n \in \mathbbm{N}^{*}}$ and $\{\gamma_{Y,n}\}_{n \in \mathbbm{N}^{*}}$ where $V$ is a definite positive $2 \times 2$ matrix, (ii) ${\sup_{n \in \mathbbm{N}^{*}} \mathbbm{E} \Big[ (\gamma_{X,n} X_{1,n})^{4+\delta} + (\gamma_{Y,n} Y_{1,n})^{4+\delta} \Big] < + \infty}$ for some $\delta > 0$, (iii) $\mathbbm{P}(\overline{Y}_n = 0) \to 0$ as $n \to \infty$, and (iv) $\mathbbm{P}(\overline{Y}_n^{(1)} = 0) \to 0$ as $n \to \infty$. Denote the signal-to-noise-ratio by ${\text{SNR}_n := \mathbbm{E}[Y_{1,n}] / (V_{2,2}^{1/2} n^{-1/2} \gamma_{Y,n}^{-1})}$. If $\text{SNR}_n \to +\infty$, then for every ${\alpha \in (0,1)}$, the confidence interval $C_{n,\alpha}^{boot}$ is pointwise consistent at level ${1-\alpha}$, viz. $\mathbbm{P} \big( C_{n,\alpha}^{boot} \ni \mathbbm{E}[X_{1,n}]/\,\mathbbm{E}[Y_{1,n}] \big) \to 1 - \alpha \text{ as } n \to \infty.$

The assumption $\mathbbm{P}(\overline{Y}_n^{(1)} = 0) \to 0$ is satisfied for a large set of cases, for instance when the variables $Y_{i,n}$ are continuous or when they follow a Bernoulli distribution with a parameter decreasing to $0$ not too fast (see Example (ref) below).

figure[figure omitted — 1,146 chars of source]

Note that the moment condition of order $4+\delta$ is nearly sharp. Indeed, the proofs require the strong law of large numbers for ${n^{-1}\sum_{i=1}^n X_{1,n}^2}$ and ${n^{-1}\sum_{i=1}^n Y_{1,n}^2}$. As we are dealing with a triangular array of random variables, Theorem 3.1 of gut1992complete shows that moments of order at least $4$ are necessary, even in the simpler case where the distribution $P_{X,Y,n}$ does not depend on $n$.

example[Example (ref) continued] When $Y_{1,n}$ follows a Bernoulli distribution with parameter $p_n = 1/n^b$ for a given $b > 0$, the condition $\mathbbm{P}(\overline{Y}_n^{(1)} = 0) \to 0$ is satisfied when $b < 1$. We refer the reader to Section (ref) for a proof of this claim.

In practice, even if the theoretical results of the delta method and of the bootstrap are valid under nearly the same set of assumptions, we observe in the simulations in Figure (ref) a gap between their pointwise coverage.\footnote{Additional simulations comparing the two types of asymptotic confidence intervals are presented in Appendix (ref).} This fact appears even when $P_{X,Y,n}$ does not depend on $n$ (i.e. $b=0)$. Nonetheless, the coverage gap between these two methods shrinks as $n$ increases provided ${b<0.5}$. In the sequence of models where the denominator decreases slowly (i.e. $b=0.25$) in Figure (ref), the bootstrap's coverage is much higher than the one of the delta method. Therefore, the CI provided by the nonparametric percentile bootstrap may be an interesting alternative compared to the delta method when conducting inference with a given sample. This is all the more so as the mean in the denominator is close to $0$ (in Figure (ref), of the size of $n^{-0.25}/10$ for a variance normalized to $1$) and the number of observations is moderately large (a few thousands here).

Construction of nonasymptotic confidence intervals for ratios of expectations

To construct nonasymptotic confidence intervals, we rely on the possibility to ensure that with large probability (i) $\overline{X}_n$ is close to $\mathbbm{E}[X_{1,n}]$, and (ii) $\overline{Y}_n$ is both close to $\mathbbm{E}[Y_{1,n}]$ and bounded away from 0. Under Assumptions (ref) and (ref), the Bienaym\'e-Chebyshev inequality can be applied to obtain (i) and (ii). On the other hand, without further restrictions, we are only able to build nonasymptotic CIs at nominal levels that are not too close to $1$ (see Section (ref)).

This limitation does not arise with nonasymptotic confidence intervals for expectations. In that sense, we can say that building nonasymptotic CIs for ratios of expectations is more demanding. Intuitively, the extra difficulty of the latter task comes from the need to ensure (ii). To stress that point, we show in the next subsection that when $\overline{Y}_n$ is bounded away from $0$ and positive almost surely, we can build nonasymptotic CIs at every nominal level.

An easy case: the support of the denominator is well-separated from \texorpdfstring{$0$}{0}

We present a simple framework in which it is possible to build nonasymptotic CIs, valid for every $n\in\mathbbm{N}^{*}$, and with coverage $1-\alpha$ for every $\alpha\in(0,1)$. To do so, we restrict further the set $\mathcal{P}$ of admissible distributions with the following assumption.

hypFor every ${n \in \mathbbm{N}^{*}}$, there exists a positive finite constant $a_{Y,n}$ such that $Y_{1,n}\geqa_{Y,n}$ almost surely.

Under Assumption (ref), for every $n \in \mathbbm{N}^{*}$, $\overline{Y}_n \geq a_{Y,n} > 0$ almost surely under every distribution in $\mathcal{P}$ and $\overline{Y}_n^{-1}$ is bounded from above. This assumption obviously rules out binary $\{0,1\}$ random variables in the denominator of the ratio, which can be quite restrictive in practice. Under this assumption, the following theorem gives a concentration inequality for our ratio of expectations. It is proved in Section (ref).

thmLet Assumptions (ref), (ref) and (ref) hold. For every $n\in\mathbbm{N}^{*}$, $\varepsilon > 0$, we have \begin{align*} \sup_{P \in \mathcal{P}} \mathbbm{P}_{P^{\otimes n}} \Bigg( \bigg| \frac{\overline{X}_n}{\overline{Y}_n} - \frac{\mathbbm{E}[X_{1,n}]}{\mathbbm{E}[Y_{1,n}]} \bigg| & > \frac{\big(\varepsilon+\sqrt{u_{X,n}}\big)\varepsilon}{\aYnl_{Y,n}} + \frac{\varepsilon}{l_{Y,n}} \Bigg) \leq \frac{u_{X,n}}{n\varepsilon^2}+\frac{u_{Y,n}-l_{Y,n}^2}{n\varepsilon^2}. \end{align*} As a consequence, $\inf_{P \in \mathcal{P}} \mathbbm{P}_{P^{\otimes n}} \Big( \mathbbm{E}[X_{1,n}] / \, \mathbbm{E}[Y_{1,n}] \in \left[\, \overline{X}_n /\,\overline{Y}_n \pm t \right] \Big) \geq 1 - \alpha$, with the choice $$t:= \frac{1}{l_{Y,n}}\sqrt{\frac{u_{X,n}+u_{Y,n}-l_{Y,n}^2}{n\alpha}}\left(1+\frac{1}{a_{Y,n}}\left\{\sqrt{\frac{u_{X,n}+u_{Y,n}-l_{Y,n}^2}{n\alpha}}+\sqrt{u_{X,n}}\right\}\right),$$ for every $\alpha \in (0,1)$.

The theorem shows that it is possible to construct nonasymptotic CIs for ratios of expectations, with guaranteed coverage at every confidence level, that are almost surely bounded under every distribution in $\mathcal{P}$ characterized by Assumptions (ref), (ref) and (ref). In Section (ref), we give an analogous result that only requires Assumptions (ref) and (ref) to hold, so that it encompasses the case of $\{0,1\}$-valued denominators. However, the cost to pay will be an upper bound on the achievable coverage of the confidence intervals.

General case: no assumption on the support of the denominator

We seek to build nontrivial nonasymptotic CIs under Assumptions (ref) and (ref) only. Under Assumption (ref), ${\mathbbm{E}[Y_{1,n}]\neq 0}$, so that there is no issue in considering the fraction $\mathbbm{E}[X_{1,n}]/\,\mathbbm{E}[Y_{1,n}]$. However, without Assumption (ref), $\left\{\overline{Y}_n=0\right\}$ has positive probability in general so that $\overline{X}_n/\,\overline{Y}_n$ is well-defined with probability less than one. Note that when $P_{Y,n}$ is continuous with respect to Lebesgue's measure, there is no issue in defining $\overline{X}_n/\,\overline{Y}_n$ anymore since the event $\left\{\overline{Y}_n = 0\right\}$ has probability zero. This is not an easier case from a theoretical point of view though since, without more restrictions, $\overline{Y}_n$ can still be arbitrarily close to $0$ with positive probability.

thmLet Assumptions (ref) and (ref) hold. For every $n\in\mathbbm{N}^{*}$, $\varepsilon > 0, \tilde \varepsilon \in (0,1)$, we have \begin{align*} \sup_{P \in \mathcal{P}} \mathbbm{P}_{P^{\otimes n}} \Bigg( \bigg| \frac{\overline{X}_n}{\overline{Y}_n} - \frac{\mathbbm{E}[X_{1,n}]}{\mathbbm{E}[Y_{1,n}]} \bigg| > \bigg( \frac{\big(\sqrt{u_{X,n}}+\varepsilon\big) \tilde \varepsilon}{(1 - \tilde \varepsilon)^2} + \varepsilon \bigg) \frac{1}{l_{Y,n}} \Bigg) \leq \frac{u_{X,n}}{n\varepsilon^2}+\frac{u_{Y,n}-l_{Y,n}^2}{n\tilde\varepsilon^2l_{Y,n}^2}. \end{align*} As a consequence, $\inf_{P \in \mathcal{P}} \mathbbm{P}_{P^{\otimes n}} \Big( \mathbbm{E}[X_{1,n}] / \, \mathbbm{E}[Y_{1,n}] \in \left[\, \overline{X}_n /\,\overline{Y}_n \pm t \right] \Big) \geq 1-\alpha$, with the choice \begin{align*} &t= \frac{1}{l_{Y,n}}\left(\frac{\left(\sqrt{u_{X,n}}+\sqrt{2u_{X,n}/(n\alpha)}\right)\sqrt{2(u_{Y,n}-l_{Y,n}^2)/(n\alphal_{Y,n}^2)}}{\left(1-\sqrt{2(u_{Y,n}-l_{Y,n}^2)/(n\alpha l_{Y,n}^2})\right)^2}+\sqrt{\frac{2u_{X,n}}{n\alpha}}\right), \end{align*} for every $\alpha > \overline{\alpha}_n := \frac{2(u_{Y,n}-l_{Y,n}^2)}{nl_{Y,n}^2}$.\footnote{Equivalently, it means that for a given $\alpha$, the above choice of $t$ is valid for every integer $n > \overline{n}_\alpha :=2(u_{Y,n}-l_{Y,n}^2)/(\alphal_{Y,n}^2)$.}

This theorem is proved in Section (ref). It states that when $l_{Y,n} > 0$, it is possible to build valid nonasymptotic CIs with finite length up to the confidence level $1-\overline{\alpha}_n$. This is a more positive result than dufour1997 which states that it is not possible to build nontrivial nonasymptotic CIs when $l_{Y,n}$ is taken equal to 0, no matter the confidence level. Note that Theorem (ref) is not an impossibility theorem since it only claims that considering confidence levels smaller than ${1-\overline{\alpha}_n}$ is sufficient to build nontrivial CIs under Assumptions (ref) and (ref). The remaining question is to find out whether it is necessary to focus on confidence levels that do not exceed a certain threshold under Assumptions (ref) and (ref). We answer this in Section (ref).

Theorem (ref) has two other interesting consequences: for every confidence level up to $1-\overline{\alpha}_n$, a nonasymptotic interval of the form $\left[\overline{X}_n/\,\overline{Y}_n\pm \tilde{t}\right]$ with $\tilde{t}>t$ has coverage ${1-\alpha}$ but is unnecessarily conservative. Moreover, if the data generating process does not depend on $n$ (i.e. in the standard i.i.d. set-up), the length of the confidence interval shrinks at the optimal rate $1/\sqrt{n}$ for every fixed $\alpha$. Note that the coefficient $2$ in the definition of $\overline{\alpha}_n$ defined above can be reduced to any number ${w > 1}$, at the expense of increasing the length of the confidence interval (this length actually tends to infinity when $w$ tends to $1$).

Nonasymptotic CIs: impossibility results and practical guidelines

In this section, we prove two impossibility results: a maximum confidence level above which it is impossible to build nontrivial nonasymptotic CIs and a necessary lower bound on the length of nonasymptotic CIs.

An upper bound on testable confidence levels

propLet $\mathcal{P}$ be the class of all distributions satisfying Assumptions (ref) and (ref) and $\underline{\alpha}_n := \big(1-l_{Y,n}^2/u_{Y,n}\big)^n$. For every $n\in\mathbbm{N}^{*}$ and every $\alpha \in \left(0, \underline{\alpha}_n \right)$, if $l_{Y,n}^2/u_{Y,n} < 1$, there is no finite $t>0$ such that $\left[\overline{X}_n/\,\overline{Y}_n\pm t\right]$ has coverage $1-\alpha$ over $\mathcal{P}$.

This theorem asserts that confidence intervals of the form $\left[\overline{X}_n/\,\overline{Y}_n\pm t\right]$ with coverage higher than ${1-\underline{\alpha}_n}$ under Assumptions (ref) and (ref) are not defined (or are of infinite length) with positive probability for at least one distribution in $\mathcal{P}$. This is due to the fact that $\underline{\alpha}_n$ is a lower bound on $\mathbbm{P}(\overline{Y}_n = 0)$ over all distributions in $\mathcal{P}$.

Remark that when $u_{Y,n}/l_{Y,n}^2=1$, there is no impossibility result anymore: assume that $u_{Y,n}/l_{Y,n}^2=1$ and let $Q$ be a distribution on $\mathbb{R}^2$ that satisfies Assumptions (ref) and (ref). Let $(X_{i,n},Y_{i,n})_{i=1}^n\buildrel {\text{i.i.d.}} \over \sim Q$. We have that $\mathbbm{V}[Y_{1,n}]=0$, which implies that $Y_{1,n}=\mathbbm{E}[Y_{1,n}]$ almost surely. Assumption (ref) further ensures that $Y_{1,n}\neq 0$ almost surely. Consequently, the results of Section (ref) apply and allow us to conclude that under Assumptions (ref), (ref) and $u_{Y,n}/l_{Y,n}^2=1$, it is possible to build nontrivial nonasymptotic CIs at every confidence level. Indeed, in that case, we are in fact only estimating a simple mean, and therefore there is no constraint on $\alpha$.

Proposition (ref) is actually a corollary of the more general Theorem (ref). It states it is impossible to construct confidence intervals that contain $\overline{X}_n/\,\overline{Y}_n$ almost surely and are almost surely bounded over $\mathcal{P}$ with coverage greater than ${1 - \underline{\alpha}_n}$. It is proved in Section (ref).

thmLet $\mathcal{P}$ be the class of all distributions satisfying Assumptions (ref) and (ref). Let $n\in\mathbb{N}^*$, and a random set $I_n$ that contains $\overline{X}_n/\,\overline{Y}_n$ almost surely whenever it is defined and is undefined if $\overline{Y}_n=0$. Then $\sup_{P \in \mathcal{P}} \mathbbm{P}_{P^{\otimes n}} \big( I_n \, undefined \big) \geq \underline{\alpha}_n.$

Combining Theorems (ref) and (ref), we conclude that there exists some critical level $1-\alpha_n^c$ belonging to the interval $[1-\overline{\alpha}_n,1-\underline{\alpha}_n]$ such that it is impossible to build nontrivial nonasymptotic confidence intervals if and only if their nominal level is above $1-\alpha_n^c$. Finally, it is worth remarking that with a sample of size $n$, the CIs based on the delta method with a nominal level $1-\alpha>1-\alpha_n^c$ cannot have coverage ${1-\alpha}$ uniformly over $\mathcal{P}$ as such CIs verify the condition of Theorem (ref).

Figure (ref) below shows the critical level and its bounds obtained in our nonasymptotic results.

figure[figure omitted — 2,391 chars of source]
remIn the same spirit as in Theorem (ref), we consider a modified version of the signal-to-noise ratio defined by $\widetilde{\text{SNR}}_n := l_{Y,n}/(u_{Y,n}^{1/2} n^{-1/2})$. When $\widetilde{\text{SNR}}_n \to + \infty$ $($resp. $0)$ as $n \to \infty$, $\underline{\alpha}_n$ and $\overline{\alpha}_n$ tend to $0$ $($resp. $+\infty)$. When we have enough information ${(\widetilde{\text{SNR}}_n \to + \infty)}$, the critical level ${1-\alpha_n^c}$ tends to $1$. Therefore, for every ${\alpha \in (0,1)}$, nonasymptotic confidence intervals can be constructed at every level for $n$ large enough. On the contrary, when ${\widetilde{\text{SNR}}_n \to 0}$, the critical level ${1-\alpha_n^c}$ tends to $0$, which means that it is impossible to construct uniformly valid CIs for $n$ large enough. Finally, when ${\widetilde{\text{SNR}}_n \to C}$ for a positive constant $C$, a critical level remains as in the nonasymptotic case since $\underline{\alpha}_n \to \exp(-C)$.

A lower bound on the length of nonasymptotic confidence intervals

The following theorem is an extension of catoni2012challenging[Proposition 6.2] to ratios. It is proved in Section (ref).

thmFor every integer $n\geq 7$, $\alpha\in \big( 0, 1\,\wedge\, n/\big(l_{Y,n}+\sqrt{u_{Y,n}-l_{Y,n}^2}\big)^2\big)$, and $\xi < 1$ there exists a distribution $Q$ on $\mathbb{R}^2$ that satisfies Assumptions (ref) and (ref) such that for $\left(X_{i,n},Y_{i,n}\right)_{i=1}^n\overset{i.i.d}\sim~Q$, we have \begin{align*} &\mathbbm{P}_{Q^{\otimes n}} \left(\left| \frac{\overline{X}_n}{\overline{Y}_n} - \frac{\mathbbm{E}[X_{1,n}]}{\mathbbm{E}[Y_{1,n}]} \right| > \xi \sqrt{\frac{v_n}{3n\alpha}} \right) >\alpha, \end{align*} where $v_n:= u_{X,n} / \big(l_{Y,n}+\sqrt{u_{Y,n}-l_{Y,n}^2}\big)^2$.

With this theorem, we can claim that CIs of the form $\left[\overline{X}_n/\,\overline{Y}_n\pm t\right]$ cannot have uniform coverage $1-\alpha$, for every $\alpha\in\big(0,1\wedge n/\big(l_{Y,n}+\sqrt{u_{Y,n}-l_{Y,n}^2}\big)^2\big)$, under Assumptions (ref) and (ref) if they are shorter than $\sqrt{v_n / (3n\alpha)}$. By a careful inspection of the proof (see Lemma (ref)), we can in fact replace the value $3$ in the theorem by any number strictly larger than $e=\exp(1)$, at the price of assuming $n \geq n_0$ for $n_0$ large enough. It is interesting to note that the distributions $Q$ that are built in the proof of the theorem are on the boundary of $\mathcal{P}$ in the sense that they satisfy $\mathbbm{E}[X_{1,n}^2]=u_{X,n}$, $\mathbbm{E}[Y_{1,n}]=l_{Y,n}$ and $\mathbbm{E}[Y_{1,n}^2]=u_{Y,n}$.

Practical methods and plug-in estimators

Nonasymptotic confidence intervals and the thresholds $\overline{\alpha}_n$ and $\overline{n}_\alpha$ based on Theorem (ref) rely on Assumptions (ref) and (ref). In practice, building such CIs or computing those thresholds require the knowledge of the constants $l_{Y,n}$, $u_{X,n}$ and $u_{Y,n}$ that determine the class of distributions we consider.\footnote{Actually, the computation of $\overline{\alpha}_n$ and $\overline{n}_\alpha$ only require the knowledge of $l_{Y,n}$ and $u_{Y,n}$.} Therefore, we need to state some values for those constants. Note that constructing nontrivial and nonasymptotic CIs that overcome the limitations of having to choose some a priori class of distributions is not possible. Indeed, we would get back to bahadur1956 and dufour1997 type impossibility results.

How to choose $l_{Y,n}$, $u_{X,n}$ and $u_{Y,n}$ depends on the specific application. Sometimes, stating values can be sensible if researchers do have control or expert knowledge of the variables. Resuming an example started in the introduction, if the variable in the denominator is an indicator of being treated in the setting of a Randomized Controlled Trial, researchers can have intuitions about reasonable values for the lower and upper bounds of the probability of being treated.

The unknown constants are upper and lower bounds on moments that characterize the class $\mathcal{P}$. As such, they can never be recovered from the data since observations are by construction drawn from a single distribution ${P \in \mathcal{P}}$. Under i.i.d. sampling, sample means converge to their corresponding theoretical moments, provided the latter are finite. Hence, without prior information, a plug-in strategy has to be used which consists in: (i) using the moments of a single distribution instead of the bounds on the class, (ii) estimating those moments with their empirical counterparts. As a consequence, this approach is valid pointwise only and not uniformly over $\mathcal{P}$ anymore. Furthermore, it is only asymptotically justified. On the other hand, for any sample provided ${\overline{Y}_n \neq 0}$, this plug-in strategy enables us to construct our CIs and the quantity $\overline{n}_\alpha$ (or $\overline{\alpha}_n$), which can be a useful rule of thumb as explained below. We stick to that principle in our simulations and application.

For a given level ${1-\alpha}$ and a class of distributions satisfying Assumptions (ref) and (ref), $\overline{n}_\alpha$ is the minimal sample size required to construct our nonasymptotic CIs. In other words, for a sample size ${n < \overline{n}_\alpha}$, the data is not rich enough to construct the nonasymptotic CIs of Theorem (ref) at this level. Heuristically, the comparison of $\overline{n}_\alpha$ and $n$ can be used as a rule of thumb to assess whether the coverage of the CIs based on the delta method matches their nominal level.\footnote{Equivalently, we could compare $\overline{\alpha}_n$ and $\alpha$. As a rule of thumb, $\overline{\alpha}_n$ can be seen as the lowest $\alpha$ (hence the highest nominal level $1-\alpha$) for which the asymptotic CIs based on the delta method are reliable given the sample size $n$.} Several simulations tend to confirm the practical interest of that rule of thumb as $\overline{n}_\alpha$ turns out to be very close to the sample size above which the gap between the coverage of the asymptotic CIs based on the delta method and their nominal level becomes negligible. (see Section (ref) and Appendix (ref)).

Numerical applications

Simulations

This section presents simulations that support the use of $\overline{n}_\alpha$, or equivalently $\overline{\alpha}_n$, as a rule of thumb to inspect the reliability of the asymptotic confidence intervals from the delta method.

In Figure (ref), a nominal level ${1-\alpha}$ is fixed and we show the $c(n,P)$ of the CIs based on the delta method as a function of the sample size $n$, as well as $\overline{n}_\alpha$ derived in Theorem (ref). It happens that the coverage converges toward its nominal level for sample sizes around $\overline{n}_\alpha$, which supports $\overline{n}_\alpha$ as a rule of thumb of interest in practice.\footnote{This fact holds across various specifications (see additional simulations in Appendix (ref)).} In Figure (ref), a sample size is fixed and we show the coverage for different nominal levels, as well as the quantity $\overline{\alpha}_n$. It is the converse of Figure (ref) in that sense. In this simulation, $\overline{\alpha}_n$ turns out to fall close to the lowest $\alpha$ (hence highest $1-\alpha$) for which the coverage of the CIs based on the delta method attains their nominal level.

figure[figure omitted — 1,041 chars of source]
figure[figure omitted — 1,164 chars of source]

All in all, Figures (ref) and (ref) and additional simulations advocate the use of $\overline{n}_\alpha$ derived in Theorem (ref) (or conversely $\overline{\alpha}_n$) as a rule of thumb to appraise the dependability of the CIs obtained with the delta method for ratios of expectations.

Application to real data

We illustrate our methods with an application related to gender wage disparities. The application resumes our canonical example of conditional expectations since we estimate the proportion of women within wage brackets that are defined as having a wage higher than a given threshold. We use $n =$ 204,246 observations from the French Labor Survey data between 2010 and 2017.\footnote{Enquête Emploi en continu (version FPR) -- 2010-2017, INSEE [producteur], ADISP [diffuseur].}

Let $W$ be a real random variable that indicates the wage of an employee (expressed in euros per month) and $F$ an indicator variable equal to $1$ if the employee is a woman and $0$ otherwise. For a given threshold wage $w_0$, the parameter of interest is ${\mathbbm{E}[F \mid W \geq w_0]}$. It can be written as a ratio of expectations with $X = F\,\mathds{1}\{W \geq w_0\} = \mathds{1}\{F = 1, W \geq w_0\}$ in the numerator and $Y = \mathds{1}\{W \geq w_0\}$ in the denominator. As we consider higher thresholds $w_0$, the expectation in the denominator gets closer to $0$. As an illustration, out of $n=$ 204,246 observations, $355$ individuals have monthly wages higher than 10,000 euros (which corresponds to a mean in the denominator equal to $0.0017$); $44$ individuals above 20,000 ($\overline{Y}_n = 2.2 \times 10^{-4}$); and only $17$ above 30,000 ($\overline{Y}_n = 8.3 \times 10^{-5}$).\footnote{To give a sense of the wage distribution, note that the empirical quantiles of $W$ at orders 90%; 95%; 99%; and 99.99% are respectively: 2,989; 3,728; 6,000; and 26,024.}

figure[figure omitted — 674 chars of source]

For various thresholds $w_0$, Figure (ref) presents the estimate $\widehat{\theta}_n$ and two 95%-nominal-level confidence intervals for the parameter ${\mathbbm{E}[F \mid W \geq w_0]}$: the one based on the delta method (see Section (ref)) and the one using Efron's percentile bootstrap (see Section (ref)). With higher thresholds, the expectation in the denominator is closer to $0$ which results in wider confidence intervals. For very high thresholds, the CIs become hardly informative. In particular, the lower end of the interval based on the delta method is negative whereas the parameter of interest belongs to $[0,1]$ by construction.

The dashed vertical line relates to our rule of thumb introduced in Section (ref). More precisely, given the level ${1-\alpha=0.95}$, for each threshold $w_0$, we compute the plug-in counterpart of $\overline{n}_\alpha$ defined in Theorem (ref): $2\left(n^{-1}\sum_{i=1}^n Y_i^2 - \overline{Y}_n^{2}\right) / \big(\alpha \overline{Y}_n^{2} \big)$. Given that $Y$ is a binary variable, the latter quantity is increasing with $w_0$ and exceeds $n$ at some threshold represented by the dashed vertical line (here a little above 20,000). Consequently, for higher thresholds, our rule of thumb suggests that the confidence intervals obtained with the delta method might undercover as the expectation in the denominator is “too close to $0$” relative to the number of observations. Actually, in the application, it is around this vertical line that the two CIs start to differ. In particular, the upper end of Efron's percentile confidence interval becomes larger than the upper end of the interval based on the delta method.

Conclusion

This paper studies the construction of confidence intervals for ratios of expectations, which are frequent parameters of interest in applied econometrics.

The most common method to do so is asymptotic and yields CIs based on the asymptotic normality of the empirical means that estimate the numerator and the denominator combined with the delta method. We document on simulations that the coverage of the confidence intervals based on the delta method may fall short of their nominal level when the expectation in the denominator is close to $0$, even with fairly large sample size.

To further study the reliability of those CIs, we use a sequence-of-model framework, analogous to what a strand of the weak IV literature does. Indeed, it enables to consider limiting cases, namely here denominators tending to $0$. In the weak IV case, the equivalent is to move closer to a null covariance between the endogenous regressor and the instrument. At the limit, the coefficient of interest is not identified. Our problem differs since the parameter is not even defined in the problematic case of a null denominator. This issue underlies the impossibility type results presented in the paper.

First, in an asymptotic perspective, the possibility of a denominator arbitrarily close to $0$ explains why we need a sufficiently slow rate of convergence of the expectation in the denominator to $0$ to conduct meaningful inference. More precisely, our main asymptotic results basically show that the CIs based on the delta method are valid, as well as those obtained by Efron's percentile bootstrap, when this speed is lower than $1/\sqrt{n}$ (the standard speed of the CLT). Furthermore, on simulations, Efron's percentile bootstrap CIs reach their nominal level sooner (namely for smaller sample sizes) than the CIs based on the delta method. It suggests that beyond the sequence-of-model rationalization, when confronted in practice to a mean in the denominator close to $0$ relative to the size of the sample at hand, Efron's percentile bootstrap CIs may be more trustworthy than the delta method's ones.

Obviously, those cases where the coverage of the CIs based on the delta method can be well below their nominal level do not self-signal to practitioners. This is why the second part of the paper proposes a rule of thumb to detect those cases and thus assess the dependability of the asymptotic CIs based on the delta method on finite samples. This index is based on the construction of nonasymptotic confidence intervals and on impossibility results that stem from the problematic null denominator case.

In substance, even if we bound away from $0$ the expectation in the denominator, there remains a partial impossibility result. Indeed, we show that there exists a critical nominal level above which the coverage of any nonasymptotic confidence interval that is undefined when $\overline{Y}_n = 0$ cannot uniformly attain its target level. More precisely, we derive explicit upper and lower bounds on this critical level as a function of the characteristics of the considered class of distributions. Then, the heuristic of our rule of thumb consists in estimating by plug-in a lower bound on this critical level (or equivalently, for a given level, an upper bound on the minimal required sample size). The resulting index can thus be computed immediately on any sample. In addition to its theoretical foundations, various simulations and an application to real data attest the practical usefulness of this rule of thumb.

This paper can be seen as a first step towards nonasymptotic inference in econometric models where the issue of close-to-zero denominators arises. Notable examples may include weak IV, Wald ratios, and difference-in-difference estimands.