Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
The Hellinger Bounds on the Kullback--Leibler Divergence and the Bernstein Norm
abstractThe Kullback--Leibler divergence, the Kullback--Leibler variation, and the Bernstein “norm” are used to quantify discrepancies among probability distributions in likelihood models such as nonparametric maximum likelihood and nonparametric Bayes.
They are closely related to the Hellinger distance, which is often easier to work with.
Consequently, it is of interest to characterize conditions under which the Hellinger distance serves as an upper bound for these measures.
This article characterizes a necessary and sufficient condition for each of the discrepancy measures to be bounded by the Hellinger distance.
It accommodates unbounded likelihood ratios and generalizes all previously known results.
We then apply it to relax the regularity condition for the sieve maximum likelihood estimator.
Introduction
Controlling the size of a function class is a central step in nonparametric statistics.
The Kullback--Leibler divergence, the Kullback--Leibler variation, and the Bernstein “norm” are standard measures of discrepancy in minimum contrast estimation vw2023 and nonparametric Bayesian analysis gv2017.
These quantities are closely related to the Hellinger distance, which endows the space of probability distributions with a convenient Hilbert space structure.
A well\hypknown fact is that the Kullback--Leibler divergence bounds the Hellinger distance from above---an inequality often used, for example, to establish identification in maximum likelihood estimation v1998.
The reverse inequality generally fails to hold.
For complexity control of function classes, however, the converse is typically the more wanted direction.
This asymmetry explains why general posterior contraction theorems involve two distinct types of neighborhoods gv2017.
Hence, it is natural to seek sufficient conditions under which the Hellinger distance serves as an upper bound for the three discrepancy measures.
b1983 and bm1998 showed that the Kullback--Leibler divergence is bounded by the Hellinger distance when the likelihood ratio is uniformly bounded.
\Citet[p. 327]{vw1996} and ggv2000 showed that the Bernstein “norm” of the log-likelihood ratio is bounded under the same condition.
gv2017 provided the corresponding bound on the Kullback--Leibler variation.
The uniform boundedness condition, however, is arguably restrictive---particularly in models involving unbounded random variables.
Even the canonical normal location model violates this requirement.
Some attempts have been made to relax this condition.
ws1995 showed that the Kullback--Leibler divergence and variation are bounded by the Hellinger distance when a local moment of the likelihood ratio is bounded by the Hellinger distance.
This result is general enough to accommodate many models with unbounded likelihood ratios but is somewhat underused in later literature; e.g., ggv2000 reformulated it in a way that the multiples on the Hellinger distance diverge as the Hellinger distance tends to zero.
When the multiples diverge, the subsequent statistical application may lead to a compromised rate of convergence or contraction.
For the Bernstein “norm”, ggv2000 proved an analogous bound under a similar condition for which the multiple diverges.
kmp2023 showed that the Bernstein “norm” of half the log-likelihood ratio is bounded by the Hellinger distance under the condition that the likelihood ratio has a finite local moment conditional on the likelihood ratio exceeding a threshold.
This was the first to accommodate unbounded likelihood ratios for a sharp Hellinger bound on the Bernstein “norm”.
kr2023 showed that the Kullback--Leibler divergence and variation are bounded under the same condition.
In this article, we identify the necessary and sufficient condition for each of the discrepancy measures to be bounded by the Hellinger distance.
We verify that the aforementioned sufficient conditions imply our necessary and sufficient conditions.
We also study the relationship between existing conditions in the literature.
Finally, we apply the results to obtain the rate of convergence of a nonparametric sieve maximum likelihood estimator under a relaxed regularity requirement, which permits the likelihood ratio to diverge faster than previously possible.
The paper is organized as follows.
(ref) introduces notation and definitions.
(ref) develops the necessary and sufficient condition for each of the discrepancy measures to be bounded by the Hellinger distance.
(ref) compares the existing conditions and the new ones.
(ref) applies the results to establish a new result of the convergence rate for a sieve maximum likelihood estimator.
(ref) concludes.
Definition
We work with probability measures defined on a measurable space $(\mathcal{X},\mathcal{A})$.
A probability measure is denoted with a capital letter (e.g., $P$), and the corresponding density (Radon--Nikodym derivative with respect to some $\sigma$\hypfinite dominating measure) with a lowercase letter (e.g., $p$).
All integrals considered in this paper do not depend on the particular choice of the dominating measure; hence it is made implicit.
For example, $\int(p-q)$ means $\int_{\mathcal{X}}(p(x)-q(x))d\mu(x)$, where $\mu$ is any $\sigma$\hypfinite measure dominating both $P$ and $Q$.
Expectations are often written using operator notation:
\[
P f\coloneqq\mathbb{E}_{X\sim P}[f(X)]=\int_{\mathcal{X}}f(x)dP(x).
\]
We also write $P(A)$ to denote the probability of an event $A\in\mathcal{A}$ under $P$.
Hence, for a measurable function $f$, $P(f)$ means $\mathbb{E}_{X\sim P}[f(X)]$ while $P(f\geq 0)$ refers to $\mathbb{E}_{X\sim P}[\mathbbm{1}\{f(X)\geq 0\}]$.
The notation $P(f\mid f\geq 0)$ denotes the conditional expectation $\mathbb{E}_{X\sim P}[f(X)\mid f(X)\geq 0]$, which we define to be $0$ if $P(f\geq 0)=0$.
The discrepancy measures of interest are defined as follows.
defn[Hellinger distance]
The {\em Hellinger distance} between two probability measures $P$ and $Q$ is
\footnote{Some authors include $1/2$ inside the integral b1983,bm1998.}
\[
h(p,q)\coloneqq\Bigl[\int(\sqrt{p}-\sqrt{q})^2\Bigr]^{1/2}
=\Bigl[\int_{\mathcal{X}}\bigl(\sqrt{p(x)}-\sqrt{q(x)}\bigr)^2 d\mu(x)\Bigr]^{1/2}.
\]
defn[Kullback--Leibler divergence and variation]
The {\em Kullback--Leibler divergence} of $Q$ from $P$ is
\[
K(p\mathrel{\Vert} q)\coloneqq P\log\frac{p}{q}=\int_{\mathcal{X}}\biggl(\log\frac{p(x)}{q(x)}\biggr)dP(x).
\]
For $k>1$, the {\em $k$th order Kullback--Leibler variation} of $Q$ from $P$ is
\[
V_k(p\mathrel{\Vert} q)\coloneqq P\biggl|\log\frac{p}{q}\biggr|^k=\int_{\mathcal{X}}\biggl|\log\frac{p(x)}{q(x)}\biggr|^k dP(x),
\]
and the {\em $k$th order centered Kullback--Leibler variation} of $Q$ from $P$ is
\[
V_{k,0}(p\mathrel{\Vert} q)\coloneqq P\biggl|\log\frac{p}{q}-K(p\mathrel{\Vert} q)\biggr|^k=\int_{\mathcal{X}}\biggl|\log\frac{p(x)}{q(x)}-P\log\frac{p}{q}\biggr|^k dP(x).
\]
If $P(A)=0$ and $Q(A)>0$, the event $A$ is ignored in these integrals.
Meanwhile, if $P(A)>0$ and $Q(A)=0$, we may understand the integrals to be infinity.
The following Bernstein “norm” was introduced in vw1996 to apply Bernstein's inequality for a maximal inequality to obtain the rate of convergence of an $M$-estimator.
defn[Bernstein “norm”]
The {\em Bernstein “norm”} of a measurable function $f$ is
\[
\lVert f\rVert_{P,B}\coloneqq\sqrt{2P(e^{\lvert f\rvert}-1-\lvert f\rvert)}=\Bigl[2\int_{\mathcal{X}}\bigl(e^{\lvert f(x)\rvert}-1-\lvert f(x)\rvert\bigr)dP(x)\Bigr]^{1/2}.
\]
This is not a true norm, as it neither is homogeneous nor satisfies the triangle inequality vw2023, but satisfies the so-called Riesz property: $\lvert f\rvert\leq\lvert g\rvert$ implies $\lVert f\rVert\leq\lVert g\rVert$ vw2023.
The idea is to control exponential deviation of a function.
Since $x^2\leq 2(e^{\lvert x\rvert}-1-\lvert x\rvert)$ for every $x\in\mathbb{R}$, we have $\lVert f\rVert_2\leq\lVert f\rVert_{P,B}$.
Moreover, as the exponential function grows faster than any polynomial, this dominates all $L^p$ norms for $p\geq 2$ up to a constant depending only on $p$.
The Bernstein “norm” naturally relates to Bernstein's inequality vw2023 and is particularly useful for minimum contrast estimation vw2023.
We also make use of a “norm” that is equivalent to the Bernstein “norm” but is much more convenient when dealing with log-likelihood ratios.
Since for every $x\in\mathbb{R}$,
equation[equation omitted — 117 chars of source]
we have
\[
\lVert f\rVert_2\leq\sqrt{P(e^f+e^{-f}-2)}\leq\lVert f\rVert_{P,B}\leq\sqrt{2 P(e^f+e^{-f}-2)}.
\]
Therefore, the “norm” $\sqrt{P(e^f+e^{-f}-2)}$ is equivalent to the Bernstein “norm”.
However, this goes well with log-likelihood ratios as it comes with various identities
equation[equation omitted — 160 chars of source]
which are particularly useful when $f$ is a log-likelihood ratio.
Finally, we occasionally use notation $x\vee y\coloneqq\max\{x,y\}$ and $x\wedge y\coloneqq\min\{x,y\}$.
The Hellinger Bounds
Bernstein “Norm”
\Citet[Chapter 3.4]{vw1996} introduced the Bernstein “norm” to derive a type of maximal inequality---a bound on the supremum of the empirical process over a class of functions.
Apparently, the empirical process evaluated at a function is homogeneous in the sense that bounding $\lVert\mathbb{G}_n\rVert_{\mathcal{F}}=\sup_{f\in\mathcal{F}}\lvert\sqrt{n}(\mathbb{P}_n-P)f\rvert$ is equivalent to bounding $2\lVert\mathbb{G}_n\rVert_{2^{-1}\mathcal{F}}=2\sup_{f\in\mathcal{F}}\lvert\sqrt{n}(\mathbb{P}_n-P)(f/2)\rvert$.
However, the Bernstein “norm” is not homogeneous in that we always have $\lVert f\rVert_{P_0,B}\geq 2\lVert 2^{-1}f\rVert_{P_0,B}$ where the RHS can even be finite when the LHS is not.
Therefore, bounding the halved Bernstein “norm” is always at least as easy as bounding the original Bernstein “norm”.
This fact was mentioned in vw1996 and was used by kmp2023 and kr2023 to prove the maximal inequality for a nonparametric classifier.
In this spirit, we present the necessary and sufficient condition for bounding the arbitrary {\em fractional} Bernstein “norm”---$\lVert\delta f\rVert_{P_0,B}$ for any $\delta\in(0,1]$---when $f$ is a log-likelihood ratio.
In practice, it suffices to meet the condition for just one $\delta\in(0,1]$ to secure a maximal inequality.
Let $\delta\in(0,1]$.
The necessary and sufficient condition for the fractional Bernstein “norm” of the log-likelihood ratio to be bounded by the Hellinger distance,
equation[equation omitted — 109 chars of source]
is found to be
equation[equation omitted — 154 chars of source]
Precisely, the following theorem shows that when $P_0([\frac{p_0}{p}]^\delta\mathbbm{1}\{\frac{p_0}{p}>4\})\leq M h(p_0,p)^2$ holds, we have $\lVert\delta\log\frac{p_0}{p}\rVert_{P_0,B}^2\leq(18\delta+2 M)h(p_0,p)^2$, along with the other direction with a different multiple.
It also shows that the Kullback--Leibler divergence and second- or higher-order Kullback--Leibler variation are bounded by the Hellinger distance.
These additional bounds are sharp in $h$ but not in the multiples.
They are enough when we assume ((ref)), but if we wish to allow the multiples to diverge or if we do not need the Bernstein “norm”, we may instead want to use (ref) below.
thm[Bernstein “norm”; $\text{(\ref{asm:BN})}\Leftrightarrow\text{(\ref{eq:BN})}$]
For arbitrary probability measures $P_0$ and $P$ and $\delta\in(0,1]$,
\[
\biggl(1-\frac{1}{4^\delta}\biggr)^2 P_0\biggl(\frac{p_0^\delta}{p^\delta}\mathbbm{1}\biggl\{\frac{p_0}{p}>4\biggr\}\biggr)
\leq \biggl\|\delta\log\frac{p_0}{p}\biggr\|_{P_0,B}^2
\leq 18\delta h(p_0,p)^2+2 P_0\biggl(\frac{p_0^\delta}{p^\delta}\mathbbm{1}\biggl\{\frac{p_0}{p}>4\biggr\}\biggr).
\]
Moreover, we have
\begin{enumerate}[(i)]
• $h(p_0,p)^2\leq K(p_0\mathrel{\Vert} p)\leq 3 h(p_0,p)^2+\delta^{-1}P_0([\frac{p_0}{p}]^\delta\mathbbm{1}\{\frac{p_0}{p}>4\})$,
• $2^{-k}V_{k,0}(p_0\mathrel{\Vert} p)\leq V_k(p_0\mathrel{\Vert} p)\leq 2^{-1}\Gamma(k+1)\delta^{-k}\lVert\delta\log\tfrac{p_0}{p}\rVert_{P_0,B}^2$
for every real $k\geq 2$.
\end{enumerate}
remFor $k=2$, $\operatorname{Var}(X)\leq\mathbb{E}[X^2]$ gives a better bound $V_{2,0}(p_0\mathrel{\Vert} p)\leq V_2(p_0\mathrel{\Vert} p)$.
proof$\text{(\ref{eq:BN})}\Rightarrow\text{(\ref{asm:BN})}$.
We first show that $(\sqrt{x^\delta}-1)^2\leq\delta(\sqrt{x}-1)^2$ for $x\geq\frac{1}{4}$ and $0<\delta\leq 1$.
Consider $x\geq 1$, so both $(\sqrt{x^\delta}-1)^2$ and $(\sqrt{x}-1)^2$ are increasing and they coincide at $x=1$.
The derivatives are, respectively,
\[
\delta\tfrac{1}{x^{1-\delta}}-\tfrac{\delta}{\sqrt{x}}\tfrac{1}{x^{\frac{1-\delta}{2}}}, \qquad
\delta-\tfrac{\delta}{\sqrt{x}}.
\]
Since $1\leq x^{\frac{1-\delta}{2}}\leq x^{1-\delta}$, we see that
\[
\delta-\tfrac{\delta}{\sqrt{x}}\geq\bigl(\delta-\tfrac{\delta}{\sqrt{x}}\bigr)\tfrac{1}{x^{1-\delta}}\geq\delta\tfrac{1}{x^{1-\delta}}-\tfrac{\delta}{\sqrt{x}}\tfrac{1}{x^{\frac{1-\delta}{2}}},
\]
so $(\sqrt{x}-1)^2$ grows faster than $(\sqrt{x^\delta}-1)^2$.
Thus, we have $(\sqrt{x^\delta}-1)^2\leq\delta(\sqrt{x}-1)^2$.
Next, consider $\frac{1}{4}\leq x<1$ so both are decreasing.
As long as $(\sqrt{x}-1)^2$ decreases faster, we have the desired inequality.
That is, we want
\[
\Bigl(\delta\tfrac{1}{x^{1-\delta}}-\tfrac{\delta}{\sqrt{x}}\tfrac{1}{x^{\frac{1-\delta}{2}}}\Bigr)-\bigl(\delta-\tfrac{\delta}{\sqrt{x}}\bigr)
=\Bigl[\delta\Bigl(\tfrac{1}{x^{\frac{1-\delta}{2}}}+1\Bigr)-\tfrac{\delta}{\sqrt{x}}\Bigr]\Bigl(\tfrac{1}{x^{\frac{1-\delta}{2}}}-1\Bigr)\geq 0.
\]
This is equivalent to
\[
\Bigl(\tfrac{1}{x^{\frac{1-\delta}{2}}}+1\Bigr)-\tfrac{1}{\sqrt{x}}\geq 0
\iff
\delta\geq\tfrac{\log(1-\sqrt{x})}{\log\sqrt{x}}.
\]
The RHS is increasing and equal to one at $x=\frac{1}{4}$.
Since $\delta\leq 1$, this holds on $\frac{1}{4}\leq x<1$ and so does our desired inequality. Thus, $(\sqrt{x^\delta}-1)^2\leq\delta(\sqrt{x}-1)^2$ for $x\geq\frac{1}{4}$.
Now we bound the fractional Bernstein “norm”.
Since $x^\delta+\frac{1}{x^\delta}-2<x^\delta$ for $x>4$,
\begin{align*}
\bigl\|\delta\log\tfrac{p_0}{p}\bigr\|_{P_0,B}^2&\leq 2 P_0\bigl(\bigl[\tfrac{p_0}{p}\bigr]^\delta+\bigl[\tfrac{p}{p_0}\bigr]^\delta-2\bigr) \tag{by ((ref))} \\
&\leq 2 P_0\bigl(\bigl[\tfrac{p}{p_0}\bigr]^{\delta/2}-1\bigr)^2\bigl(1+\bigl[\tfrac{p_0}{p}\bigr]^{\delta/2}\bigr)^2\mathbbm{1}\bigl\{\tfrac{p_0}{p}\leq 4\bigr\}
+2 P_0\bigl[\tfrac{p_0}{p}\bigr]^\delta\mathbbm{1}\bigl\{\tfrac{p_0}{p}>4\bigr\} \tag{by ((ref))} \\
&\leq 2 P_0\delta\bigl(\sqrt{\tfrac{p}{p_0}}-1\bigr)^2 (1+2)^2+2 P_0\bigl[\tfrac{p_0}{p}\bigr]^\delta\mathbbm{1}\bigl\{\tfrac{p_0}{p}>4\bigr\} \tag{paragraphs above} \\
&\leq 18\delta h(p_0,p)^2+2 P_0\bigl[\tfrac{p_0}{p}\bigr]^\delta\mathbbm{1}\bigl\{\tfrac{p_0}{p}>4\bigr\}.
\end{align*}
$\text{(\ref{eq:BN})}\Leftarrow\text{(\ref{asm:BN})}$.
Since $(1-\frac{1}{4^\delta})^2 x^\delta<(1-\frac{1}{x^\delta})^2 x^\delta=x^\delta+\frac{1}{x^\delta}-2$ for $x>4$, by ((ref)),
\[
\bigl(1-\tfrac{1}{4^\delta}\bigr)^2 P_0\bigl[\tfrac{p_0}{p}\bigr]^\delta\mathbbm{1}\bigl\{\tfrac{p_0}{p}>4\bigr\}
\leq P_0\bigl(\bigl[\tfrac{p_0}{p}\bigr]^\delta+\bigl[\tfrac{p}{p_0}\bigr]^\delta-2\bigr)\mathbbm{1}\bigl\{\tfrac{p_0}{p}>4\bigr\}
\leq \bigl\|\delta\log\tfrac{p_0}{p}\bigr\|_{P_0,B}^2.
\]
((ref))
The lower bound follows from $(\sqrt{x}-1)^2\leq x-1-\log x$ as
\[
h(p_0,p)^2=P_0\bigl(\sqrt{\tfrac{p}{p_0}}-1\bigr)^2+P(p_0=0)
\leq P_0\bigl(\tfrac{p}{p_0}-1-\log\tfrac{p}{p_0}\bigr)+P(p_0=0)=K(p_0\mathrel{\Vert} p).
\]
For the upper bound, since $x-1-\log x\leq 3(\sqrt{x}-1)^2$ for $x\geq\frac{1}{4}$, we can write
\begin{align*}
K(p_0\mathrel{\Vert} p)&=P_0\bigl(\tfrac{p}{p_0}-1-\log\tfrac{p}{p_0}\bigr)+P(p_0=0)\\
&\leq 3 P_0\bigl(\sqrt{\tfrac{p}{p_0}}-1\bigr)^2\mathbbm{1}\bigl\{\tfrac{p}{p_0}\geq\tfrac{1}{4}\bigr\}+P(p_0=0)+P_0\bigl(\tfrac{p}{p_0}-1-\log\tfrac{p}{p_0}\bigr)\mathbbm{1}\bigl\{\tfrac{p_0}{p}>4\bigr\}\\
&\leq 3 h(p_0,p)^2+P_0\log\tfrac{p_0}{p}\mathbbm{1}\bigl\{\tfrac{p_0}{p}>4\bigr\}.
\end{align*}
Finally,
\(
P_0\log\tfrac{p_0}{p}\mathbbm{1}\{\tfrac{p_0}{p}>4\}
=\delta^{-1}P_0\delta\log\tfrac{p_0}{p}\mathbbm{1}\{\tfrac{p_0}{p}>4\}
\leq\delta^{-1}P_0[\tfrac{p_0}{p}]^\delta\mathbbm{1}\{\tfrac{p_0}{p}>4\}
\).
((ref))
For the first inequality, observe that for $k\geq 1$,
\begin{align*}
V_{k,0}(p_0\mathrel{\Vert} p)^{1/k}&=
\bigl(P_0\bigl|\log\tfrac{p_0}{p}-P_0\log\tfrac{p_0}{p}\bigr|^k\bigr)^{1/k}\\
&\leq\bigl(P_0\bigl|\log\tfrac{p_0}{p}\bigr|^k\bigr)^{1/k}+\bigl|P_0\log\tfrac{p_0}{p}\bigr| \tag{triangle inequality} \\
&\leq\bigl(P_0\bigl|\log\tfrac{p_0}{p}\bigr|^k\bigr)^{1/k}+\bigl(P_0\bigl|\log\tfrac{p_0}{p}\bigr|^k\bigr)^{1/k} \tag{Jensen's inequality} \\
&=2 V_k(p_0\mathrel{\Vert} p)^{1/k}.
\end{align*}
For the second inequality, it suffices to show that $x^k/\Gamma(k+1)\leq e^x-1-x$ for $k\geq 2$ and $x\geq 0$, since then by letting $x=\lvert\delta\log\frac{p_0}{p}\rvert$ we obtain
\[
P_0\bigl|\log\tfrac{p_0}{p}\bigr|^k=\delta^{-k} P_0\bigl|\delta\log\tfrac{p_0}{p}\bigr|^k\leq 2^{-1}\Gamma(k+1)\delta^{-k}\bigl\|\delta\log\tfrac{p_0}{p}\bigr\|_{P_0,B}^2.
\]
By the definition of the gamma function, for arbitrary $t\geq 0$,
\[
\Gamma(k-1)=\int_0^\infty z^{k-2}e^{-z}dz
\geq\int_t^\infty z^{k-2}e^{-z}dz
\geq t^{k-2}\int_t^\infty e^{-z}dz=t^{k-2}e^{-t}.
\]
Thus, we deduce that
\[
\tfrac{x^k}{\Gamma(k+1)}=\int_0^x\int_0^y\tfrac{t^{k-2}}{\Gamma(k-1)}dt dy
\leq\int_0^x\int_0^y e^t dt dy=e^x-1-x.
\]
This completes the proof.
In condition ((ref)) and throughout the paper, the threshold of four appears frequently.
This choice is somewhat arbitrary---it can be any fixed number strictly above $1$.
We choose a square number as we deal with many square roots.
Note that ((ref)) is {\em not} equivalent to the one without the cutoff,
\(
P_0([\frac{p_0}{p}]^\delta)\lesssim h(p_0,p)^2
\);
since $\frac{p_0}{p}$ approaches one as $h(p_0,p)^2\to 0$, the LHS never vanishes.
A sufficient but not necessary condition for ((ref)) is $P_0(\lvert\sqrt{\tfrac{p_0}{p}}-1\rvert^{2\delta})\lesssim h(p_0,p)^2$.
\footnote{It is necessary when $\delta=1$. Then, it means that the “reverse” Hellinger distance is of the same order as the “forward” Hellinger distance.}
Kullback--Leibler Divergence and Variation
The next question is when the Kullback--Leibler divergence and variation are bounded by the Hellinger distance, that is, for $k\geq 2$,
align[align omitted — 157 chars of source]
We establish that the necessary and sufficient condition for each, respectively, is
align[align omitted — 293 chars of source]
Since the integrand of the divergence has positive and negative parts that cancel each other, the necessity of ((ref)) is not trivial.
This implies that the second set of bounds given in gv2017 cannot be improved.
Also note that the higher the $k$, the stronger the condition.
The third point of the next theorem tells that the multiples grow at different rates.
For example, if $P_0([\log\tfrac{p_0}{p}]^2\mathbbm{1}\{\tfrac{p_0}{p}>4\})\leq M h(p_0,p)^2$, we have $P_0([\log\tfrac{p_0}{p}]\mathbbm{1}\{\tfrac{p_0}{p}>4\})\leq 4\sqrt{M}h(p_0,p)^2$.
thm[Kullback--Leibler divergence and variation]
For arbitrary probability measures $P_0$ and $P$, the following hold.
\begin{enumerate}[(i)]
• $\text{(\ref{asm:KLD})}\Leftrightarrow\text{(\ref{eq:KLD})}$:
\[
\frac{1}{3}P_0\biggl(\biggl[\log\frac{p_0}{p}\biggr]\mathbbm{1}\biggl\{\frac{p_0}{p}>4\biggr\}\biggr)\leq K(p_0\mathrel{\Vert} p)\leq 3 h(p_0,p)^2+P_0\biggl(\biggl[\log\frac{p_0}{p}\biggr]\mathbbm{1}\biggl\{\frac{p_0}{p}>4\biggr\}\biggr).
\]
• $\text{(\ref{asm:KLV})}\Leftrightarrow\text{(\ref{eq:KLV})}$:
For every real $k\geq 2$,
\begin{multline*}
P_0\biggl(\biggl[\log\frac{p_0}{p}\biggr]^k\mathbbm{1}\biggl\{\frac{p_0}{p}>4\biggr\}\biggr)\leq V_k(p_0\mathrel{\Vert} p)
\leq 4\biggl([2(\log 4)^{k-2}]\vee\biggl[\frac{k}{e}\biggr]^k\biggr)h(p_0,p)^2\\
+P_0\biggl(\biggl[\log\frac{p_0}{p}\biggr]^k\mathbbm{1}\biggl\{\frac{p_0}{p}>4\biggr\}\biggr).
\end{multline*}
• $\text{(\ref{asm:KLV})}\Rightarrow\text{(\ref{asm:KLD})}$:
For every real $k'\geq k>0$,
\[
P_0\biggl(\biggl[\log\frac{p_0}{p}\biggr]^k\mathbbm{1}\biggl\{\frac{p_0}{p}>4\biggr\}\biggr)
\leq 4 h(p_0,p)^{2(1-\frac{k}{k'})}\biggl[P_0\biggl(\biggl[\log\frac{p_0}{p}\biggr]^{k'}\mathbbm{1}\biggl\{\frac{p_0}{p}>4\biggr\}\biggr)\biggr]^{\frac{k}{k'}}.
\]
\end{enumerate}
proof((ref))
$\text{(\ref{asm:KLD})}\Rightarrow\text{(\ref{eq:KLD})}$.
It is shown in the proof of (ref) ((ref)).
$\text{(\ref{asm:KLD})}\Leftarrow\text{(\ref{eq:KLD})}$.
Since $\log\frac{1}{x}<3(x-1-\log x)$ for $0<x<\frac{1}{4}$,
\[
\tfrac{1}{3}P_0\log\tfrac{p_0}{p}\mathbbm{1}\bigl\{\tfrac{p_0}{p}>4\bigr\}
\leq P_0\bigl(\tfrac{p}{p_0}-1-\log\tfrac{p}{p_0}\bigr)\leq P_0\log\tfrac{p_0}{p}.
\]
((ref))
$\text{(\ref{asm:KLV})}\Rightarrow\text{(\ref{eq:KLV})}$.
Note that $(\log x)^2\leq 8(\sqrt{x}-1)^2$ for $x\geq\frac{1}{4}$.
Hence, for $\frac{1}{4}\leq x\leq 4$ and $k\geq 2$, we have
\(
\lvert\log x\rvert^k\leq(\log 4)^{k-2}(\log x)^2\leq 8(\log 4)^{k-2}(\sqrt{x}-1)^2
\).
Now, we want $C_k$ such that
\[
(\log x)^k\leq C_k(\sqrt{x}-1)^2
\]
for $x>4$.
This is equivalent to bounding
\(
\sup_{x>4}\frac{(\log x)^k}{(\sqrt{x}-1)^2}
\).
Since
\(
\frac{(\log x)^k}{(\sqrt{x}-1)^2}<4\frac{(\log x)^k}{x}
\)
for $x>4$, we have
\[
C_k\leq 4\sup_{x>4}\frac{(\log x)^k}{x}.
\]
This is attained at $x=e^k$, so
\(
C_k\leq 4(\tfrac{k}{e})^k
\).
Then, the bound follows with $x=\frac{p}{p_0}$.
$\text{(\ref{asm:KLV})}\Leftarrow\text{(\ref{eq:KLV})}$.
Trivially, $P_0(\log\frac{p_0}{p})^k\mathbbm{1}\{\frac{p_0}{p}>4\}\leq P_0\lvert\log\frac{p_0}{p}\rvert^k$.
((ref))
$\text{(\ref{asm:KLV})}\Rightarrow\text{(\ref{asm:KLD})}$.
It suffices to consider $k'>k$.
Since $\frac{1}{2}<1-\sqrt{x}\leq 1$ if $\frac{1}{x}>4$,
\begin{align*}
P_0\bigl(\log\tfrac{p_0}{p}\bigr)^k\mathbbm{1}\bigl\{\tfrac{p_0}{p}>4\bigr\}
&\leq 4 P_0\bigl(\log\tfrac{p_0}{p}\bigr)^k\mathbbm{1}\bigl\{\tfrac{p_0}{p}>4\bigr\}\bigl(1-\sqrt{\tfrac{p}{p_0}}\bigr)^{2\frac{k'-k}{k'}} \\
&\leq 4\bigl[P_0\bigl(\log\tfrac{p_0}{p}\bigr)^{k'}\mathbbm{1}\bigl\{\tfrac{p_0}{p}>4\bigr\}\bigr]^{\frac{k}{k'}}\bigl[P_0\bigl(1-\sqrt{\tfrac{p}{p_0}}\bigr)^2\bigr]^{\frac{k'-k}{k'}}
\end{align*}
by H\"older's inequality with $p=\frac{k'}{k}>1$ and $q=\frac{k'}{k'-k}>1$.
remIf we want to bound the variation for $1\leq k<2$, we need to impose an assumption to control not only the event $\{\frac{p_0}{p}>4\}$ but also $\{\frac{p_0}{p}\approx 1\}$.
Comparison
We compare condition ((ref)) with four other conditions.
The first one is the {\em uniform boundedness condition}:
equation[equation omitted — 147 chars of source]
used in bm1998, ggv2000, and gv2017.
The second is the condition imposed in ws1995: for some $\delta\in(0,1]$ and $M<\infty$,
equation[equation omitted — 170 chars of source]
which is equivalent to ((ref)) but looks weaker as the threshold is made to diverge as $\delta\to 0$.
The third is the {\em finite moment condition}:
equation[equation omitted — 82 chars of source]
employed in ggv2000 and gv2017.
The fourth is the {\em finite conditional moment condition}:
equation[equation omitted — 153 chars of source]
introduced in kmp2023 and kr2023.
In particular, we show that the following relationship holds:
\[
((ref))\implies((ref))\implies
array[array omitted — 203 chars of source]
\implies
\tikzmarknode{M}{((ref), $\delta<1$)}
\]
tikzpicture[tikzpicture omitted — 229 chars of source]
Comparison among Different $\delta$'s
For $\delta\leq\delta'$, we obviously have
\(
P_0([\frac{p_0}{p}]^\delta\mathbbm{1}\{\frac{p_0}{p}>4\})\leq P_0([\frac{p_0}{p}]^{\delta'}\mathbbm{1}\{\frac{p_0}{p}>4\})
\),
so ((ref)) for $\delta'$ implies ((ref)) for $\delta$.
Not only that, we can further show that the multiples grow at different rates.
For example, the following theorem implies that if $P_0(\frac{p_0}{p}\mathbbm{1}\{\frac{p_0}{p}>4\})\leq M h(p_0,p)^2$, then $P_0(\sqrt{\tfrac{p_0}{p}}\mathbbm{1}\{\frac{p_0}{p}>4\})\leq 4\sqrt{M}h(p_0,p)^2$.
prop[$\text{(\ref{asm:BN}, $\delta'$)}\Rightarrow\text{(\ref{asm:BN}, $\delta$)}$]
For arbitrary probability measures $P_0$ and $P$ and $0<\delta\leq\delta'$,
\[
P_0\biggl(\biggl[\frac{p_0}{p}\biggr]^\delta\mathbbm{1}\biggl\{\frac{p_0}{p}>4\biggr\}\biggr)\leq 4 h(p_0,p)^{2(1-\frac{\delta}{\delta'})} \biggl[P_0\biggl(\biggl[\frac{p_0}{p}\biggr]^{\delta'}\mathbbm{1}\biggl\{\frac{p_0}{p}>4\biggr\}\biggr)\biggr]^{\frac{\delta}{\delta'}}.
\]
That is, if ((ref)) holds for $\delta'$, then ((ref)) holds for $\delta$ with a multiple that grows slower if grows at all.
proofIt mirrors (ref) ((ref)).
Replace $\log\tfrac{p_0}{p}$, $k$, and $k'$ with $\tfrac{p_0}{p}$, $\delta$, and $\delta'$.
exa[$\text{(\ref{asm:BN}, $\delta'$)}\nLeftarrow\text{(\ref{asm:BN}, $\delta$)}$]
Let $p_0(x)=\mathbbm{1}\{0<x<1\}$ be the uniform density over $(0,1)$, and consider $p(x)=2 x\mathbbm{1}\{0<x<1\}$.
Then, it is easy to verify that ((ref)) for $\delta=1$ fails but ((ref)) for $\delta=1/2$ holds.
Condition in ws1995
While ((ref)) is equivalent to ((ref)), ws1995 used it to bound the Kullback--Leibler measures rather than the Bernstein “norm”.
We verify that ((ref)) is indeed a sufficient condition for ((ref)) and ((ref)).
The next proposition recovers the multiples of the same orders as ws1995 when combined with (ref) for $k=1,2$.
prop[$\text{(\ref{asm:D})}\Rightarrow\text{(\ref{asm:KLD}, \ref{asm:KLV})}$]
If ((ref)) holds for $\delta\in(0,1]$ and $M<\infty$, then for every real $k>0$,
\[
P_0\biggl(\biggl[\log\frac{p_0}{p}\biggr]^k\mathbbm{1}\biggl\{\frac{p_0}{p}>4\biggr\}\biggr)\leq\delta^{-k}\biggl[4+\frac{e}{(\sqrt{e}-1)^2}(k\vee\log M)^k\biggr]h(p_0,p)^2.
\]
proofFor $0<\delta\leq 1$ and $e^{-1/\delta}\leq x<\frac{1}{4}$, we have
\(
(\log\tfrac{1}{x})^k\leq\tfrac{1}{\delta^k}=\tfrac{4}{\delta^k}(\tfrac{1}{\sqrt{4}}-1)^2\leq\tfrac{4}{\delta^k}(\sqrt{x}-1)^2
\),
noting that for $(\log 4)^{-1}<\delta\leq 1$, there is no $x$ that satisfies $e^{-1/\delta}\leq x<\frac{1}{4}$ so it is vacuously true.
Therefore,
\[
P_0\bigl(\log\tfrac{p_0}{p}\bigr)^k\mathbbm{1}\bigl\{\tfrac{p_0}{p}>4\bigr\}\leq P_0\bigl(\log\tfrac{p_0}{p}\bigr)^k\mathbbm{1}\bigl\{\tfrac{p_0}{p}>e^{\frac{1}{\delta}}\bigr\}+\tfrac{4}{\delta^k}P_0\bigl(\sqrt{\tfrac{p}{p_0}}-1\bigr)^2.
\]
Observe that for $k>0$, $\eta>0$, and $B>0$,
\begin{equation}
\inf_{0<r\leq\eta}\tfrac{1}{r^k}B^{r}=\begin{cases}\frac{e^k}{k^k}(\log B)^k&if $B>e^{\frac{k}{\eta}}$,\\ \frac{1}{\eta^k}B^{\eta}&if $B\leq e^{\frac{k}{\eta}}$.\end{cases}
\end{equation}
Hence, for $0<r\leq 1$ and $0<x<e^{-1/\delta}$, we have
\(
\tfrac{k^k}{e^k}\tfrac{1}{r^k\delta^k}(\tfrac{1}{x})^{r\delta}
\geq\tfrac{k^k}{e^k}\inf_{0<r\delta}\tfrac{1}{r^k\delta^k}(\tfrac{1}{x})^{r\delta}
=(\log\tfrac{1}{x})^k
\).
Now we see that
\begin{align*}
P_0\bigl(\log\tfrac{p_0}{p}\bigr)^k\mathbbm{1}\bigl\{\tfrac{p_0}{p}>e^{\frac{1}{\delta}}\bigr\}&\leq\tfrac{k^k}{e^k}\tfrac{1}{r^k\delta^k}P_0(\tfrac{p_0}{p})^{r\delta}\mathbbm{1}\bigl\{\tfrac{p_0}{p}>e^{\frac{1}{\delta}}\bigr\}\\
&\leq\tfrac{k^k}{e^k}\tfrac{1}{r^k\delta^k}P_0(\tfrac{p_0}{p})^{r\delta}\mathbbm{1}\bigl\{\tfrac{p_0}{p}>e^{\frac{1}{\delta}}\bigr\}(1-e^{-1/2})^{-2}\bigl(1-\sqrt{\tfrac{p}{p_0}}\bigr)^{2(1-r)} \tag{for $\delta\leq 1$, $r\leq 1$} \\
&\leq\tfrac{k^k}{e^k}\tfrac{1}{r^k\delta^k}\tfrac{1}{(1-e^{-1/2})^2}\bigl[P_0\bigl(\tfrac{p_0}{p}\bigr)^{\delta}\mathbbm{1}\bigl\{\tfrac{p_0}{p}>e^{\frac{1}{\delta}}\bigr\}\bigr]^r[h(p_0,p)^2]^{1-r}\tag{H\"older's inequality for $p=\frac{1}{r}\geq 1$, $q=\frac{1}{1-r}>1$}\\
&\leq\tfrac{k^k}{e^k}\tfrac{1}{r^k\delta^k}\tfrac{e}{(\sqrt{e}-1)^2}M^r h(p_0,p)^2. \tag{by ((ref))}
\end{align*}
Since this holds for every $0<r\leq 1$, use ((ref)) again to obtain the bound.
Finite Conditional Moment Condition
Condition ((ref)) was introduced in kmp2023 and kr2023 and was the first to accommodate unbounded likelihood ratios for the sharp Hellinger bound on the Bernstein “norm”.
However, it is stronger than ((ref)) in the following sense:
enumerate[(a)]
• For fixed $P_0$ and $P$, ((ref), $\delta=1$) and ((ref)) are equivalent.
• A universal constant in ((ref), $\delta=1$) does not imply a universal constant in ((ref)).
These are consequences of the following proposition and example.
prop[$\text{(\ref{asm:local})}\Rightarrow\text{(\ref{asm:BN}, $\delta=1$)}$]
The following hold.
\begin{enumerate}[(i)]
• If ((ref)) holds, then
\[
P_0\biggl(\frac{p_0}{p}\mathbbm{1}\biggl\{\frac{p_0}{p}>4\biggr\}\biggr)\leq(2 M+1)^2 h(p_0,p)^2.
\]
• If ((ref)) holds with $\delta=1$, then ((ref)) holds but $M$ can be arbitrarily large.
\end{enumerate}
proof((ref))
Note that the infimum in ((ref)) is attained at some finite $c\leq 1\vee M$.
Denote $C=[1+\frac{1}{2c}]^2$.
Since $C>1$,
\begin{multline*}
h(p_0,p)^2
\geq\int(\sqrt{p_0}-\sqrt{p})^2\mathbbm{1}\bigl\{\tfrac{p_0}{p}\geq C\bigr\}
\geq\int\bigl(\sqrt{p_0}-\sqrt{\tfrac{p_0}{C}}\bigr)^2\mathbbm{1}\bigl\{\tfrac{p_0}{p}\geq C\bigr\} \\
=\bigl(1-\tfrac{1}{\sqrt{C}}\bigr)^2 P_0\bigl(\tfrac{p_0}{p}\geq C\bigr)
=\tfrac{1}{(2c+1)^2}P_0\bigl(\tfrac{p_0}{p}\geq C\bigr).
\end{multline*}
This implies
\[
P_0\tfrac{p_0}{p}\mathbbm{1}\bigl\{\tfrac{p_0}{p}>4\bigr\}
\leq P_0\tfrac{p_0}{p}\mathbbm{1}\bigl\{\tfrac{p_0}{p}\geq C\bigr\}
=P_0\bigl(\tfrac{p_0}{p}\bigm|\tfrac{p_0}{p}\geq C\bigr)P_0\bigl(\tfrac{p_0}{p}\geq C\bigr)
\leq\tfrac{(2c+1)^2}{c}M h(p_0,p)^2.
\]
Use $1\leq c\leq 1\vee M$ to complete the proof.
((ref))
Trivially, if ((ref)) holds for $\delta=1$, then ((ref)) holds with some $M<\infty$.
(ref) demonstrates that $M$ can be arbitrarily large.
exa[$\text{(\ref{asm:local})}\nLeftarrow\text{(\ref{asm:BN}, $\delta=1$)}$]
Let $p_0$ be the uniform density over $(0,1)$.
For $\theta\in[0,1/4)$, consider the model
\[
p_\theta(x)=\begin{cases}\theta&0<x<\theta^2,\\1-\theta&\theta^2\leq x<1-\theta,\\\frac{1-\theta^3-(1-\theta)(1-\theta-\theta^2)}{\theta}&1-\theta\leq x<1.\end{cases}
\]
Then, $h(p_0,p_\theta)^2$ is approximately linear at $\theta=0$ with slope $3-2\sqrt{2}$ and $P_0(\frac{p_0}{p_\theta}\mathbbm{1}\{\frac{p_0}{p_\theta}>4\})=\theta$.
Therefore, ((ref)) is satisfied for $\delta=1$ and the multiple $M=(3-2\sqrt{2})^{-1}$.
Meanwhile, if we pick $c$ such that $\frac{1}{1-\theta}<[1+\frac{1}{2c}]^2$, then
\footnote{Since $c\geq 1$ and $\theta<1/4$, $[1+\frac{1}{2c}]^2<\frac{1}{\theta}$ is granted.}
\[
P_0\biggl(\frac{p_0}{p_\theta}\biggm|\frac{p_0}{p_\theta}\geq\biggl[1+\frac{1}{2c}\biggr]^2\biggr)=\frac{1}{\theta}\operatorname*{\mathchoice{
\,\longrightarrow\,}{
\rightarrow}{
\rightarrow}{
\rightarrow}
}\infty \qquad \text{as}\quad\theta\to 0.
\]
If we pick $c$ such that $\frac{1}{1-\theta}\geq[1+\frac{1}{2c}]^2$, then
\[
c P_0\biggl(\frac{p_0}{p_\theta}\biggm|\frac{p_0}{p_\theta}\geq\biggl[1+\frac{1}{2c}\biggr]^2\biggr)\geq\frac{1}{2}\frac{\sqrt{1-\theta}}{1-\sqrt{1-\theta}}\cdot\frac{\frac{1}{\theta}\theta^2+\frac{1}{1-\theta}(1-\theta-\theta^2)}{1-\theta}\operatorname*{\mathchoice{
\,\longrightarrow\,}{
\rightarrow}{
\rightarrow}{
\rightarrow}
}\infty.
\]
Therefore, $M$ in ((ref)) can be arbitrarily large.
This gain in generality is due to the localization by an unconditional expectation.
When we want to achieve $P_0(\sqrt{\frac{p_0}{p}}-1)^2\lesssim h(p_0,p)^2$, it does not have to be that
\[
P_0\biggl(\sqrt{\frac{p_0}{p}}-1\biggr)^2\mathbbm{1}\biggl\{\frac{p_0}{p}>C\biggr\}\lesssim\int(\sqrt{p_0}-\sqrt{p})^2\mathbbm{1}\biggl\{\frac{p_0}{p}>C\biggr\}
\]
event by event.
Condition ((ref)) requires that this hold for the local unfavorable event $1<C\leq\frac{9}{4}$, while condition ((ref)) allows for the possibility that
\[
P_0\biggl(\sqrt{\frac{p_0}{p}}-1\biggr)^2\mathbbm{1}\biggl\{\frac{p_0}{p}>4\biggr\}\lesssim\int(\sqrt{p_0}-\sqrt{p})^2\mathbbm{1}\biggl\{\frac{p_0}{p}\leq 4\biggr\}.
\]
Uniform Boundedness Condition
Condition ((ref)), which was used in bm1998, ggv2000, and gv2017, requires that the likelihood ratio be uniformly bounded on the entire support.
This condition is arguably restrictive, especially when the underlying random variables are unbounded, since then $p_0$ and $p$ can take arbitrarily small values.
It trivially implies the finite conditional moment, hence all other conditions.
prop[$\text{(\ref{asm:bounded})}\Rightarrow\text{(\ref{asm:local})}$]
If ((ref)) holds, then ((ref)) holds for $M\leq\lVert\frac{p_0}{p}\rVert_\infty$.
proof$M=\inf_{c\geq 1}c P_0(\frac{p_0}{p}\mid\frac{p_0}{p}\geq[1+\frac{1}{2c}]^2)\leq P_0(\frac{p_0}{p}\mid\frac{p_0}{p}\geq\frac{9}{4})\leq\lVert\frac{p_0}{p}\rVert_\infty$.
This example shows that the normal location model satisfies ((ref)) but not ((ref)).
exa[$\text{(\ref{asm:bounded})}\nLeftarrow\text{(\ref{asm:local})}$; Normal location model]
Let $\mathcal{P}=\{P_\theta=N(\theta,1):\theta\in\mathbb{R}\}$ and $P_0=N(0,1)$.
Then, we have
\[
\frac{p_0(x)}{p_\theta(x)}=\exp\biggl(-\theta x+\frac{\theta^2}{2}\biggr).
\]
Hence, $\lVert\frac{p_0}{p_\theta}\rVert_\infty$ is infinity for $\theta\neq 0$, failing ((ref)).
Meanwhile,
\[
\inf_{c\geq 1}c P_0\biggl(\frac{p_0}{p_\theta}\biggm|\frac{p_0}{p_\theta}\geq\biggl[1+\frac{1}{2c}\biggr]^2\biggr)
\leq P_0\biggl(\frac{p_0}{p_\theta}\biggm|\frac{p_0}{p_\theta}\geq e\biggr)
=\frac{e^{\theta^2}\Phi(\frac{\lvert\theta\rvert}{2}-\frac{1}{\lvert\theta\rvert}+\lvert\theta\rvert)}{\Phi(\frac{\lvert\theta\rvert}{2}-\frac{1}{\lvert\theta\rvert})}.
\]
This is finite for every $\theta$ (it approaches $1$ as $\theta\to 0$).
Therefore, ((ref)) holds on every bounded set of $\theta$.
Finite Moment Condition
Condition ((ref)) was employed in ggv2000 and gv2017 to derive Hellinger bounds whose multiples diverge as the Hellinger distance converges to zero.
Conditions ((ref)) and ((ref), $\delta<1$) do not imply each other, while each of them is implied by ((ref), $\delta=1$).
exa[$\text{(\ref{asm:moment})}\nLeftarrow\text{(\ref{asm:BN}, $\delta<1$)}$]
In (ref), we have $P_0(\frac{p_0}{p})=\infty$ while ((ref), $\delta=1/2$) holds.
Thus, ((ref)) is not even necessary for the Hellinger dominance.
The next example shows that ((ref), $\delta<1$) is not implied by ((ref)), and hence neither is ((ref), $\delta=1$) implied by ((ref)).
exa[$\text{(\ref{asm:moment})}\nRightarrow\text{(\ref{asm:BN}, $\delta<1$)}$]
First, we show that $\sup_p P_0(\frac{p_0}{p})<\infty$ does not imply ((ref)) for $\delta<1$ with a universal constant $M<\infty$.
Let $p_0(x)=\mathbbm{1}\{0<x<1\}$ be the uniform density over $(0,1)$, and consider the model
\[
p_\theta(x)=\begin{cases}\theta&0<x\leq\theta,\\1+\theta&\theta<x<1,\end{cases}
\]
indexed by $\theta\in[0,1/4)$.
Hence, $p_\theta$ equals $p_0$ when $\theta=0$.
This model satisfies ((ref)) since
\[
P_0\biggl(\frac{p_0}{p_\theta}\biggr)=1+\frac{1-\theta}{1+\theta}\leq 2.
\]
Meanwhile, we have
\begin{gather*}
P_0\biggl(\sqrt{\frac{p_0}{p_\theta}}\mathbbm{1}\biggl\{\frac{p_0}{p_\theta}>4\biggr\}\biggr)=\sqrt{\theta},\\
h(p_0,p_\theta)^2=\theta(1-\sqrt{\theta})^2+(1+\theta)(1-\sqrt{1+\theta})^2\approx\theta.
\end{gather*}
Thus, ((ref), $\delta=1/2$) fails along $\theta\to 0$.
We can further verify that Hellinger dominance fails as $\theta\to 0$.
We have already shown that ((ref), $\delta<1$) is implied by ((ref), $\delta=1$) in (ref), so what remains to be proved is that ((ref)) is implied by ((ref), $\delta=1$).
prop[$\text{(\ref{asm:BN}, $\delta=1$)}\Rightarrow\text{(\ref{asm:moment})}$]
If ((ref)) holds for $\delta=1$, then ((ref)) holds.
proofBy the Cauchy--Schwarz inequality,
\begin{align*}
P_0\tfrac{p_0}{p}
&=P_0\tfrac{p_0}{p}\mathbbm{1}\bigl\{\tfrac{p_0}{p}>4\bigr\}+P_0\bigl(1-\sqrt{\tfrac{p}{p_0}}\bigr)\sqrt{\tfrac{p_0}{p}}\bigl(\sqrt{\tfrac{p_0}{p}}+1\bigr)\mathbbm{1}\bigl\{\tfrac{p_0}{p}\leq 4\bigr\}+P_0\bigl(\tfrac{p_0}{p}\leq 4\bigr)\\
&\leq P_0\tfrac{p_0}{p}\mathbbm{1}\bigl\{\tfrac{p_0}{p}>4\bigr\}+6 \sqrt{P_0\bigl(1-\sqrt{\tfrac{p}{p_0}}\bigr)^2}+1
\leq P_0\tfrac{p_0}{p}\mathbbm{1}\bigl\{\tfrac{p_0}{p}>4\bigr\}+6 h(p_0,p)+1,
\end{align*}
which is finite under ((ref), $\delta=1$) since $h(p_0,p)^2\leq 2$ by construction.
Application to Maximum Likelihood Estimation
We now apply (ref) to relax the bounded likelihood ratio condition in nonparametric maximum likelihood estimation, namely vw2023.
To control the complexity of a class of functions, we use the bracketing integral.
defn[Bracketing number and integral]
Let $d$ be a premetric on real\hypvalued functions that is compatible with pointwise partial ordering gv2017, that is, (i) $d(f,f)=0$, (ii) $d(f,g)=d(g,f)\geq 0$, and (iii) $d(l,u)=\sup\{d(f,g):l\leq f,g\leq u\}$ for every $f,g,l,u$.
A pair of functions $[l,u]$ is called an {\em $\varepsilon$\hypbracket} if $\ell\leq u$ and $d(l,u)<\varepsilon$.
The {\em bracketing number} $N_{[]}(\varepsilon,\mathcal{F},d)$ is the minimum number of $\varepsilon$\hypbrackets needed to cover a set of functions $\mathcal{F}$.
\footnote{The bracketing functions need not come from $\mathcal{F}$ but are confined to the function space defined by $d$. This means that if $d$ is (derived from) a norm, the brackets need to have finite norms vw2023.}
The {\em nonstandardized bracketing integral} is defined by
\[
\tilde{J}_{[]}(\delta,\mathcal{F},d)\coloneqq\int_0^\delta\sqrt{1+\log N_{[]}(\varepsilon,\mathcal{F},d)}d\varepsilon.
\]
The bracketing integral is an increasing and concave function in $\delta$; hence, for $c\geq 1$, we have $\tilde{J}_{[]}(c\delta,\mathcal{F},d)\leq c\tilde{J}_{[]}(\delta,\mathcal{F},d)$.
The next theorem generalizes the sieve maximum likelihood theorem of vw2023 to models with unbounded likelihood ratios.
\footnote{Another similar result is gv2017.
This is a special case of vw2023 with an additional assumption that $p_0\in\mathcal{P}_n$, so one can always take $p_n=p_0$ in the theorem statement.}
Let $\mathbb{P}_n$ denote the empirical measure of an i.i.d. sample $X_1,X_2,\dots,X_n$.
thm[Rate of convergence of sieve MLE]
Let $X_1,X_2,\dots$ be an independent sequence from a probability distribution $P_0$.
Let $\mathcal{P}_n$ be a sequence of arbitrary sets of probability distributions, and denote by $\mathcal{P}_{n,\delta}=\{p\in\mathcal{P}_n:h(p_0,p)\leq\delta\}$ the $\delta$\hypneighborhood of $p_0$ with respect to the Hellinger distance.
Suppose there exist sequences $p_n\in\mathcal{P}_n$ and $\delta_n\geq 0$ and $M\in[0,\infty)$ that satisfy the following three conditions:
\begin{gather}
h(p_0,p_n)\lesssim\delta_n, \\
P_0\biggl(\biggl[\log\frac{p_0}{p_n}\biggr]^2\mathbbm{1}\biggl\{\frac{p_0}{p_n}>4\biggr\}\biggr)\leq M \delta_n^2, \\
\tilde{J}_{[]}(\delta_n,\mathcal{P}_{n,\delta_n},h)\leq\delta_n^2\sqrt{n}.
\end{gather}
Then, the approximate maximizer $\hat{p}_n\in\mathcal{P}_n$ of the likelihood $p\mapsto\prod_{i=1}^n p(X_i)$ in the sense that
\begin{equation}
\mathbb{P}_n\log\hat{p}_n\geq\mathbb{P}_n\log p_n-O_P(\delta_n^2)
\end{equation}
satisfies $h(p_0,\hat{p}_n)=O_P(\delta_n\vee n^{-1/2})$.
remCondition ((ref)) quantifies the sieve's approximation power in terms of the Hellinger distance, which is equivalent to the Kullback--Leibler divergence thanks to condition ((ref)) and (ref).
A sufficient condition for ((ref)) is that $p_0/p_n$ is uniformly bounded.
Condition ((ref)) controls the local entropy of the sieve, for which a sufficient condition is the restriction of the global entropy, $\tilde{J}_{[]}(\delta_n,\mathcal{P}_n,h)\leq\delta_n^2\sqrt{n}$.
Condition ((ref)) requires that the optimization algorithm does as good as $p_n\in\mathcal{P}_n$, so a sufficient condition is to pin down the global maximizer within a tolerance, $\mathbb{P}_n\log\hat{p}_n\geq\sup_{p\in\mathcal{P}_n}\mathbb{P}_n\log p-O_P(\delta_n^2)$.
With these sufficient conditions, (ref) reduces to vw2023.
remIt is straightforward to replace $M$ in ((ref)) with $M_n\to\infty$ and obtain the rate $\delta_n\sqrt{M_n}$.
Then, replace also $\delta_n$ in ((ref)) and ((ref)) by $\delta_n\sqrt{M_n}$.
Note that even when ((ref)) holds, this $M_n$ grows logarithmically slower than the multiple that arises from vw2023.
\footnote{To see this, use similar arguments as those in (ref).}
proofTo exploit the trick introduced by bm1993 vw2023, let $m_p^q\coloneqq\log\frac{p+q}{2p}$.
We will apply vw2023 to prove this theorem, where the mapping between their notation (LHS) and ours (RHS) is
\begin{align*}
\theta&=p,\\
\theta_{n,0}&=p_0,\\
\theta_n&=p_n,\\
\Theta_n&=\mathcal{P}_n,\\
d_n(\theta,\theta_{n,0})&=h(p_0,p),\\
\underaccent{\bar}{\delta}_n&=h(p_0,p_n),\\
\delta_n&=2(\sqrt{6}+\sqrt{3})\underaccent{\bar}{\delta}_n,\\
\mathbb{M}_n(\theta)&=(4+2\sqrt{2})^2\mathbb{P}_n m_{p_0}^p,\\
M_n(\theta)&=(4+2\sqrt{2})^2 P_0 m_{p_0}^p,\\
\phi_n(\delta)&=an appropriate majorant of \tilde{J}_{[]}(\delta,\mathcal{P}_{n,\delta},h)\Bigl[1+\tfrac{\tilde{J}_{[]}(\delta,\mathcal{P}_{n,\delta},h)}{\delta^2\sqrt{n}}\Bigr].
\end{align*}
A key departure from vw2023 is that we do not let $\theta_{n,0}=p_n$ but set $\theta_{n,0}=p_0$.
This saves us unnecessary complications that arise from dealing with “misspecification” of $\mathcal{P}_n$.
We begin by introducing a useful inequality for later use.
Since $(1-\frac{1}{\sqrt{2}})^2(\sqrt{p_0}-\sqrt{p})^2\leq(\sqrt{p_0}-\sqrt{\frac{p_0+p}{2}})^2\leq\frac{1}{2}(\sqrt{p_0}-\sqrt{p})^2$, we have
\begin{equation}
\bigl(1-\tfrac{1}{\sqrt{2}}\bigr)^2 h(p_0,p)^2\leq h\bigl(p_0,\tfrac{p_0+p}{2}\bigr)^2\leq\tfrac{1}{2}h(p_0,p)^2.
\end{equation}
Second, since $0\leq\frac{2p_0}{p_0+p}\leq 2$, when we compare $\frac{p_0+p}{2}$ against $p$, the LHS of ((ref)) is zero (so the multiple is zero).
Now, we verify each condition of vw2023.
The first condition is that for every $n$ and $\delta>\underaccent{\bar}{\delta}_n$,
\[
\sup_{p\in\mathcal{P}_n:\frac{\delta}{2}<h(p_0,p)\leq\delta}(4+2\sqrt{2})^2\bigl(P_0 m_{p_0}^p-P_0 m_{p_0}^{p_0}\bigr)\leq-\delta^2.
\]
This immediately follows from (ref) ((ref)) and ((ref)) since
\begin{multline*}
P_0 m_{p_0}^p-P_0 m_{p_0}^{p_0}=-P_0\log\tfrac{2p_0}{p_0+p}\\
\leq-h\bigl(p_0,\tfrac{p_0+p}{2}\bigr)^2
\leq-\bigl(1-\tfrac{1}{\sqrt{2}}\bigr)^2 h(p_0,p)^2
\leq-\bigl(1-\tfrac{1}{\sqrt{2}}\bigr)^2\tfrac{\delta^2}{4}
=-\tfrac{\delta^2}{(4+2\sqrt{2})^2}.
\end{multline*}
The second condition to establish is
\[
\mathbb{E}^*\sup_{p\in\mathcal{P}_n:h(p_0,p)\leq\delta}(4+2\sqrt{2})^2\sqrt{n}\bigl|(\mathbb{P}_n-P_0)m_{p_0}^p-(\mathbb{P}_n-P_0)m_{p_0}^{p_0}\bigr|\lesssim\phi_n(\delta).
\]
Note that (ref) and ((ref)) yield
\[
\lVert m_{p_0}^p\rVert_{P_0,B}=\bigl\|\log\tfrac{2p_0}{p_0+p}\bigr\|_{P_0,B}\leq\sqrt{18}h\bigl(p_0,\tfrac{p_0+p}{2}\bigr)\leq 3\delta.
\]
Thus, in light of vw2023, we have
\begin{multline*}
\mathbb{E}^*\sup_{p\in\mathcal{P}_n:h(p_0,p)\leq\delta}\sqrt{n}\bigl|(\mathbb{P}_n-P_0)m_{p_0}^p-(\mathbb{P}_n-P_0)m_{p_0}^{p_0}\bigr|\\
\lesssim\tilde{J}_{[]}(3\delta,\mathcal{M}_{n,\delta},\lVert\cdot\rVert_{P_0,B})\Bigl[1+\tfrac{\tilde{J}_{[]}(3\delta,\mathcal{M}_{n,\delta},\lVert\cdot\rVert_{P_0,B})}{9\delta^2\sqrt{n}}\Bigr],
\end{multline*}
where $\mathcal{M}_{n,\delta}=\{m_{p_0}^p:p\in\mathcal{P}_n,h(p_0,p)\leq\delta\}$.
Let $[\ell,u]$ be an $\varepsilon$\hypbracket in $\mathcal{P}$ with respect to $h$.
Since $u\geq\ell$ and $e^{\lvert x\rvert}-1-\lvert x\rvert\leq 2(e^{x/2}-1)^2$ for $x\geq 0$, we have
\[
\lVert m_{p_0}^u-m_{p_0}^\ell\bigr\|_{P_0,B}^2
\leq 4 P_0\bigl(\sqrt{\tfrac{p_0+u}{p_0+\ell}}-1\bigr)^2
\leq 4\int\bigl(\sqrt{p_0+u}-\sqrt{p_0+\ell}\bigr)^2\leq 4 h(u,\ell)^2.
\]
Thus, $[m_{p_0}^\ell,m_{p_0}^u]$ makes a $2\varepsilon$\hypbracket in $\mathcal{M}$ with respect to the Bernstein “norm”, so
\(
N_{[]}(2\varepsilon,\mathcal{M}_{n,\delta},\lVert\cdot\rVert_{P_0,B})\leq N_{[]}(\varepsilon,\mathcal{P}_{n,\delta},h)
\).
This implies
\begin{align*}
\tilde{J}_{[]}(3\delta,\mathcal{M}_{n,\delta},\lVert\cdot\rVert_{P_0,B})
&=\int_0^{3\delta}\sqrt{1+\log N_{[]}(\varepsilon,\mathcal{M}_{n,\delta},\lVert\cdot\rVert_{P_0,B})}d\varepsilon \\
&\leq\int_0^{3\delta}\sqrt{1+\log N_{[]}(\tfrac{\varepsilon}{2},\mathcal{P}_{n,\delta},h)}d\varepsilon \\
&=2\int_0^{\frac{3}{2}\delta}\sqrt{1+\log N_{[]}(\varepsilon,\mathcal{P}_{n,\delta},h)}d\varepsilon \\
&=2\tilde{J}_{[]}(\tfrac{3}{2}\delta,\mathcal{P}_{n,\delta},h)
\leq 3\tilde{J}_{[]}(\delta,\mathcal{P}_{n,\delta},h).
\end{align*}
Thus, we obtain
\[
\mathbb{E}^*\sup_{p\in\mathcal{P}_{n,\delta}}\sqrt{n}\bigl|(\mathbb{P}_n-P_0)m_{p_0}^p-(\mathbb{P}_n-P_0)m_{p_0}^{p_0}\bigr|
\lesssim\tilde{J}_{[]}(\delta,\mathcal{P}_{n,\delta},h)\Bigl[1+\tfrac{\tilde{J}_{[]}(\delta,\mathcal{P}_{n,\delta},h)}{\delta^2\sqrt{n}}\Bigr].
\]
Since $\tilde{J}_{[]}$ is increasing and concave in $\delta$, the RHS has a majorant $\phi_n(\delta)$ that is increasing in $\delta\geq\underaccent{\bar}{\delta}_n$ and, if divided by $\delta^\alpha$ for $1<\alpha<2$, is decreasing in $\delta$.
We now check the properties of $p_n$ and $\delta_n$ required by vw2023.
By ((ref)), we have
\[
\tilde{J}_{[]}(\delta_n,\mathcal{P}_{n,\delta_n},h)\Bigl[1+\tfrac{\tilde{J}_{[]}(\delta_n,\mathcal{P}_{n,\delta_n},h)}{\delta_n^2\sqrt{n}}\Bigr]\leq 2\delta_n^2\sqrt{n}.
\]
Therefore, $\phi_n$ can be made to satisfy $\phi_n(\delta_n)\leq\delta_n^2\sqrt{n}$.
Next, we show that
\[
\delta_n^2\geq(4+2\sqrt{2})^2\bigl(P_0 m_{p_0}^{p_0}-P_0 m_{p_0}^{p_n}\bigr).
\]
This can be verified using (ref) ((ref)) and ((ref)),
\[
P_0 m_{p_0}^{p_0}-P_0 m_{p_0}^{p_n}
=P_0\log\tfrac{2p_0}{p_0+p_n}
\leq 3 h\bigl(p_0,\tfrac{p_0+p_n}{2}\bigr)^2
\leq \tfrac{3}{2} h(p_0,p_n)^2.
\]
Hence, by letting $\delta_n=(4+2\sqrt{2})\sqrt{3/2}\underaccent{\bar}{\delta}_n=2(\sqrt{6}+\sqrt{3})\underaccent{\bar}{\delta}_n$, we have $\delta_n^2\geq(4+2\sqrt{2})^2(P_0 m_{p_0}^{p_0}-P_0 m_{p_0}^{p_n})$ as well as $\delta_n\geq\underaccent{\bar}{\delta}_n$.
Finally, it remains to show that
\(
\mathbb{P}_n m_{p_0}^{\hat{p}_n}\geq\mathbb{P}_n m_{p_0}^{p_n}-O_P(\delta_n^2)
\).
By the convexity of the logarithm,
\[
2\mathbb{P}_n m_{p_0}^{\hat{p}_n}
=2\mathbb{P}_n\log\tfrac{p_0+\hat{p}_n}{2p_0}
\geq\mathbb{P}_n\log\tfrac{p_0}{p_0}+\mathbb{P}_n\log\tfrac{\hat{p}_n}{p_0}
\geq\mathbb{P}_n\log\tfrac{p_n}{p_0}+\mathbb{P}_n\log\tfrac{\hat{p}_n}{p_n}.
\]
The second term is bounded from below by $-O_P(\delta_n^2)$ by ((ref)).
The first term can be decomposed as
\(
\mathbb{P}_n\log\tfrac{p_n}{p_0}=-(\mathbb{P}_n-P_0)\log\tfrac{p_0}{p_n}-P_0\log\tfrac{p_0}{p_n}
\).
Combined together, it suffices to show that
\[
-\tfrac{1}{2}(\mathbb{P}_n-P_0)\log\tfrac{p_0}{p_n}-\tfrac{1}{2}P_0\log\tfrac{p_0}{p_n}+(\mathbb{P}_n-P_0)\log\tfrac{2p_0}{p_0+p_n}+P_0\log\tfrac{2p_0}{p_0+p_n}\geq-O_P(\delta_n^2).
\]
For the nonrandom terms, we apply (ref) ((ref)) and ((ref)) for $k=2$, ((ref)), ((ref)), and ((ref)) to obtain
\begin{align*}
-\tfrac{1}{2}P_0\log\tfrac{p_0}{p_n}+P_0\log\tfrac{2p_0}{p_0+p_n}&\geq-\tfrac{1}{2}\bigl(3+4\sqrt{M}\bigr)\delta_n^2+h\bigl(p_0,\tfrac{p_0+p_n}{2}\bigr)^2\\
&\geq\bigl[-\tfrac{3}{2}-2\sqrt{M}+\bigl(1-\tfrac{1}{\sqrt{2}}\bigr)^2\bigr]\delta_n^2\geq-O_P(\delta_n^2).
\end{align*}
To bound the remaining random terms,
observe that, since $\{X_i\}$ is a random sample,
\[
\operatorname{Var}\bigl((\mathbb{P}_n-P_0)\log\tfrac{p_0}{p_n}\bigr)=\tfrac{1}{n}V_{2,0}(p_0\mathrel{\Vert} p_n)\leq\tfrac{1}{n}V_2(p_0\mathrel{\Vert} p_n)\leq\tfrac{8+M}{n}\delta_n^2,
\]
where the last inequality follows from (ref) ((ref)) for $k=2$, ((ref)), and ((ref)).
Thus,
\[
(\mathbb{P}_n-P_0)\log\tfrac{p_0}{p_n}=O_P\bigl(\tfrac{\delta_n}{\sqrt{n}}\bigr).
\]
(The fact that this is not $O_P(\delta_n^2)$ makes the rate of convergence at least as slow as $n^{-1/2}$.)
The same argument
implies $(\mathbb{P}_n-P_0)\log\tfrac{2p_0}{p_0+p_n}=O_P(\delta_n/\sqrt{n})$.
We have thus verified all assumptions of vw2023.
The new regularity condition imposed in (ref) is ((ref)).
This is of course weaker than uniform boundedness, but note also that it is based on ((ref)), not on ((ref)), even though the proof uses the Bernstein “norm”.
This is thanks to the trick introduced by bm1993 and to setting $\theta_{n,0}=p_0$, unlike $\theta_{n,0}=p_n$ in vw2023.
\footnote{It is possible to further relax ((ref)) by using a more general weak law of large numbers such as g1992w when bounding $(\mathbb{P}_n-P_0)\log\frac{p_0}{p_n}$.}
The next example contrasts (ref) with vw2023.
To keep it simple, we contrive a “sieve” version of the normal location model that is covered by the former but not by the latter.
exa[Sieve normal location model]
Let $\mathcal{P}=\{P_\theta=N(\theta,1):\theta\in\mathbb{R}\}$ and $P_0=N(0,1)$.
Consider a sieve $\mathcal{P}_n=\{p_\theta\in\mathcal{P}:\lvert\theta\rvert\geq 1/\sqrt{n}\}$, intentionally excluding $p_0$.
As discussed in (ref), this model does not have bounded likelihood ratios and hence falls outside the scope of vw2023.
To establish the parametric rate, we set $\theta_n=1/\sqrt{n}$, $p_n=p_{\theta_n}$, and $\delta_n\sim 1/\sqrt{n}$, and check the assumptions of (ref).
Observe that for the normal location model,
\[
h(p_{\theta_1},p_{\theta_2})^2=2-2 e^{-\frac{(\theta_1-\theta_2)^2}{8}}.
\]
Thus, we have $h(p_0,p_n)=O(n^{-1/2})$, and ((ref)) is satisfied.
In (ref), it is verified that the normal location model satisfies ((ref)), and since ((ref)) implies ((ref)), it satisfies ((ref)).
Since ((ref)) is satisfied by the maximum likelihood estimator, it remains to show ((ref)).
For $\delta_n\sim n^{-1/2}$, it boils down to
\[
\tilde{J}_{[]}(\delta,\mathcal{P}_{n,\delta},h)=O(\delta),
\]
or equivalently, that $N_{[]}(\delta,\mathcal{P}_{n,\delta},h)$ is uniformly bounded over $\delta>0$.
The closed\hypform expression of the Hellinger distance implies that $h(p_0,p_\theta)$ is
first\hyporder approximated by $\lvert\theta\rvert/2$ around $\theta=0$.
Therefore, in the neighborhood of $\theta=0$, we may interchangeably use $\lvert\theta\rvert\leq\delta$ and $h(p_0,p_\theta)\leq\delta/2$.
Let $[\ell,u]\subset\mathbb{R}$ be a bracket in $\mathbb{R}$ and consider a bracket $[p_L,p_U]$ in $\mathcal{P}$ of the form
\[
p_L(x)=\inf_{\theta\in[\ell,u]}p_\theta(x)
=\begin{cases}
p_u(x)&x<\frac{u-\ell}{2},\\
p_\ell(x)&x\geq\frac{u-\ell}{2},
\end{cases}
\quad
p_U(x)=\sup_{\theta\in[\ell,u]}p_\theta(x)
=\begin{cases}
p_\ell(x)&x<\ell,\\
p_0(0)&x\in[\ell,u],\\
p_u(x)&x>u.
\end{cases}
\]
Note that this bracket contains every distribution $p_\theta$ for $\ell\leq\theta\leq u$.
Since $N_{[]}(\delta,\{\theta\in\mathbb{R}:\lvert\theta\rvert\leq\delta\},\lvert\cdot\rvert)$ is independent of $\delta$ (namely, $1$), if we show that $h(p_U,p_L)=O(u-\ell)$, then $N_{[]}(\delta,\mathcal{P}_{n,\delta},h)$ can be bounded by a constant independent of $\delta$, establishing ((ref)).
Observe that
\begin{align*}
h(p_U,p_L)^2
&=\int_{-\infty}^\infty\bigl(\sqrt{p_u}-\sqrt{p_\ell}\bigr)^2+\int_\ell^u\bigl[\bigl(\sqrt{p_0(0)}-\sqrt{p_L}\bigr)^2-\bigl(\sqrt{p_U}-\sqrt{p_L}\bigr)^2\bigr]\\
&\leq h(p_u,p_\ell)^2+\int_\ell^u\bigl(\sqrt{p_0(0)}-\sqrt{p_u(\ell)}\bigr)^2
=h(p_u,p_\ell)^2+o((u-\ell)^2).
\end{align*}
Thus, we have shown that $h(p_U,p_L)=O(u-\ell)$.
Therefore, (ref) implies that the maximum likelihood estimator converges at rate $n^{-1/2}$.
Although this example is deliberately simple, it demonstrates the capacity of (ref) to deliver sharp rates in models that escaped previously discussed conditions.
Additional applications of (ref) arise in the study of posterior contraction rates in nonparametric Bayesian inference. As noted by gv2017, once the equivalence between the Hellinger distance and the Kullback--Leibler divergence and variation distance is established, gv2017 can be reformulated solely in terms of the Hellinger distance. A further application is the relaxation of the boundedness assumption in gv2017. Finally, (ref) may also help connect minimax convergence rates under different loss functions; see, for example, that similar inequalities are being used in b1983, bbm1999, and yb1999.
Conclusion
We established sharp Hellinger dominance for likelihood-based discrepancy measures under minimal moment conditions.
In particular, we developed the necessary and sufficient conditions for the Hellinger bounds over the fractional Bernstein “norm” of the log-likelihood ratio ((ref)), the Kullback--Leibler divergence ((ref) ((ref))), and the Kullback--Leibler variation ((ref) ((ref))).
They accommodate unbounded likelihood ratios and generalize all known results.
In all cases, it boils down to controlling the behavior of the integrands on an unfavorable event $\{\frac{p_0}{p}>4\}$.
We then compared the sufficient conditions in the literature with our necessary and sufficient conditions.
It was shown that the following implications hold.
\[
((ref))\implies((ref))\implies
array[array omitted — 172 chars of source]
\iff
\tikzmarknode{D2}{((ref))}
\implies
array[array omitted — 177 chars of source]
\implies
array[array omitted — 177 chars of source]
\]
The minimal assumption that implies the Hellinger dominance on all three discrepancy measures is ((ref)) with some $\delta\in(0,1]$, and the minimal assumption for the Kullback--Leibler divergence and variation is ((ref)).
Next, we applied our results to relax the bounded likelihood ratio condition in nonparametric sieve maximum likelihood estimation.
(ref) introduces a new regularity condition ((ref)), which allows for unbounded likelihood ratios without compromising the convergence rate.
The example based on the normal location model demonstrated the potential usefulness of this generalization.
Other possible applications include posterior contraction rates in nonparametric Bayesian inference and minimax convergence rates in nonparametric estimation.
\addcontentsline{toc}{section}{\refname}