EconBase
← Back to paper

On the Asymptotic Inadmissibility of Double Machine Learning Estimators Under Structure-Agnostic Models

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

70,013 characters

On the Asymptotic Inadmissibility of Double Machine Learning Estimators Under Structure-Agnostic Models



\title{On the Asymptotic Inadmissibility of Double Machine Learning Estimators Under
Structure-Agnostic Models}

\author[1]{Lin Liu\thanks{\href{[email removed]}{[email removed]}.}\orcidlink{0000-0002-9883-7962}}

\author[2]{Rajarshi Mukherjee\thanks{\href{[email removed]}{[email removed]}.}\orcidlink{0000-0002-5761-8958}}

\author[3]{James M. Robins\thanks{\href{[email removed]}{[email removed]}. R Mukherjee's research is supported by NSF CAREER Award 8529216-01. All authors are grateful for the hospitality of the Isaac Newton Institute of Mathematical Sciences at the University of Cambridge during the completion of this work.}\orcidlink{0000-0001-6609-209X}}

\affil[1]{Institute of Natural Sciences, MOE-LSC, School of Mathematical Sciences, SJTU-Yale Joint Center for Biostatistics and Data Science, Shanghai Jiao Tong University}

\affil[2]{Department of Biostatistics, Harvard T. H. Chan School of Public Health}

\affil[3]{Department of Epidemiology and Department of Biostatistics, Harvard T. H. Chan School of Public Health}

\date{\today}
\maketitle

\vspace{-2em}

\begin{abstract}
Structure-agnostic (SA) models introduced by \citet{balakrishnan2026fundamental} aim to reflect the general lack of knowledge of structural assumptions on data-generating laws such as smoothness or sparsity in practice. Roughly speaking, SA models restrict the observed-data generating law to be in some $r_{n}$-neighborhood of (black-box machine learning) estimates, treated as given and fixed, where $r_{n}$ encodes the convergence rates of the estimates to the truth. Under SA models, \citet{balakrishnan2026fundamental} show that the popular Double Machine Learning (DML) estimators for three functionals, the quadratic functional in the Gaussian sequence model, the quadratic density integral functional and the expected conditional covariance, are minimax. However, minimax estimators may be inadmissible. In this paper, we show that, for the first two of the three functionals, the DML estimator is asymptotically inadmissible under the SA model. In particular, we show that these two functionals fall into a class of functionals, which we refer to as the \emph{monotone bias class}. For this class, we exhibit second-order ($U$-statistic) estimators, which asymptotically dominate DML estimators, under the SA model. These second-order estimators are empirical higher-order influence function (HOIF) estimators introduced in \citet{liu2017semiparametric}. Furthermore, the empirical HOIF estimator, like the DML estimator, is minimax for the third functional (the expected conditional covariance), although neither asymptotically dominates the other. Finally, we compare the SA model with the assumption-lean model of \citet{liu2020nearly, liu2024assumption} that imposes no assumptions beyond the trivial and empirically untestable hypothesis that the bias of any estimator, including DML estimators and empirical HOIF estimators, may be of order~$1$. As a consequence, under our assumption-lean model, a Wald confidence interval centered at a DML estimator may under-cover. \citet{liu2024assumption} introduced a class of valid tests that can falsify, for functionals in the \emph{monotone bias class}, the hypothesis that a DML-estimator-centered confidence interval covers the truth at its nominal level or greater. However, our tests are not consistent under the assumption-lean model, because no consistent tests exist \citep{robins1997toward}. Furthermore, for any functional with the mixed bias property of \citet{rotnitzky2021characterization}, such as the expected conditional covariance or the average treatment effect \citep{jin2025structure}, the above falsification tests can falsify the hypothesis of \emph{rate-double-robustness}.
\end{abstract}

\textbf{Keywords: } Foundations of Statistics, Higher-Order Influence Functions, Structure-Agnostic Models, Assumption-Lean Inference, Minimaxity, (In)admissibility

\affil[1]{Institute of Natural Sciences, MOE-LSC, School of Mathematical Sciences, SJTU-Yale Joint Center for Biostatistics and Data Science, Shanghai Jiao Tong University}

\affil[2]{Department of Biostatistics, Harvard T. H. Chan School of Public Health}

\affil[3]{Department of Epidemiology and Department of Biostatistics, Harvard T. H. Chan School of Public Health}

\onehalfspacing
\allowdisplaybreaks
\nopagebreak


\section{Introduction}
\label{sec:intro}

In scientific disciplines such as epidemiology, clinical medicine and economics, one of the most important statistical tasks is to infer from the observed data a low-dimensional, smooth functional $\psi(\theta)$ of the underlying data-generating law $\mathbb{P}_{\theta}$ posited to belong to a statistical model denoted by
\begin{align*}
\mathcal{P} \equiv \mathcal{P}(\Theta) \coloneqq \left\{  \mathbb{P}_{\theta}: \theta \in \Theta \right\}.
\end{align*}
Here we parameterize the data generating law by $\theta\in\Theta$. Without an essential loss of generality, we take $\psi\equiv\psi(\theta)\in\mathbb{R}$. Throughout this paper, we let $n$ denote the sample size and use $\psi$ and
$\theta$ to denote the true values, which should cause no confusion.

Many examples of $\psi$, such as the average treatment effect under
ignorability, are of substantive interest in practice. To avoid model
misspecification bias, it is natural to take $\Theta$ to be high- or even
infinite-dimensional and estimate $\theta$ nonparametrically by kernels or
series in classical statistics. In terms of $\psi$, it has become common
practice to construct the so-called Double Machine Learning (DML) estimator
$\widehat{\psi}_{1,n}\equiv\widehat{\psi}_{1,n}(\widehat{\theta})$ based on the
first-order influence function of $\psi$
\citep{newey1990semiparametric, scharfstein1999adjusting, ai2003efficient, chernozhukov2018double, shi2026and}.
Owing to the curse-of-dimensionality, uniformly consistent estimators exist
neither for $\theta$ nor for $\psi$ without any additional assumptions on
$\Theta$
\citep{stone1980optimal, stone1982optimal, ritov1990achieving, robins1997toward}.
For this reason, structural assumptions, traditionally in terms of smoothness
or sparsity, are often imposed on $\Theta$ to obtain uniformly consistent
estimators of $\psi$ that converge to $\psi$ at parametric rates.



Recently, \citet{balakrishnan2026fundamental} introduced the \emph{structure-agnostic} (SA) models, a new class of submodels of $\mathcal{P}$ that do not impose traditional structural assumptions on $\Theta$ such as smoothness or sparsity. The SA model, denoted as $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})$, is parameterized by a pair of indices $(\widehat{\theta},r_{n})$, where $\widehat{\theta}$ is an initial estimator of $\theta$ treated as fixed and independent of the randomness of the data, and $r_{n}$ indicates convergence rates that are nonincreasing functions of $n$. Concretely, suppose that $\theta=(\theta
_{1},\cdots,\theta_{J})^{\top}$ and $r_{n}=(r_{n,1},\cdots,r_{n,J})^{\top}$
have $J$ components. Then $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})$ is
generally defined as\footnote{In certain problems, we have $\theta_{j_{1}}
\equiv\theta_{j_{2}}$ and $r_{n, j_{1}} \equiv r_{n, j_{2}}$. As is standard
in the literature, in such a case, we use the same estimator $\widehat{\theta
}_{j_{1}} \equiv\widehat{\theta}_{j_{2}}$ in computing $\widehat{\psi}_{1, n}$. We
will discuss the implications of this choice in the examples in
Sections~\ref{sec:gsm}--~\ref{sec:ecc}; specifically, see
Remarks~\ref{rem:cf1} and~\ref{rem:cf2}.}
\begin{equation}
\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})\coloneqq\left\{  \mathbb{P}
_{\theta}\in\mathcal{P}:\Vert\widehat{\theta}_{j}-\theta_{j}\Vert^{2} \leq
r_{n,j},\,j=1,\cdots,J\right\} .\label{SA model}
\end{equation}
In other words, $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta}, r_{n})$ contains the
subset of all possible $\mathbb{P}_{\theta}$'s such that each component
$\theta_{j}$ is contained in the corresponding $\sqrt{r_{n, j}}$-$\Vert\cdot\Vert$-neighborhood of a given initial estimator $\widehat{\theta}_{j}$. For all the concrete examples in this paper (see Sections~\ref{sec:gsm}--\ref{sec:ecc}), we effectively take $r_{n, j}$ to be some large
enough constant $R^{\ast}$ for $j > 2$ so assumptions are only imposed over at
most two components of $\theta$.

The SA model $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})$ has several
notable features. First, it does not impose explicit complexity reducing
structural assumptions such as smoothness or sparsity\footnote{Even if the
true smoothness or sparsity class were known, due to the theory-practice gap
\citep{adcock2021gap, xu2022deepmed, chen2024causal}, when $\theta$ is
estimated by modern deep neural networks, the SA model is still relevant
because the smoothness/sparsity assumption alone fails to determine the
properties of $\widehat{\theta}$ or $\widehat{\psi}_{1,n}$.}.
Second, \citet{balakrishnan2026fundamental} showed that the first-order DML
estimator $\widehat{\psi}_{1,n}$ of $\psi$ attains optimal convergence rates in
the minimax sense under the SA model $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},
r_{n})$, for several concrete examples of $\psi$, including the quadratic
functional in the Gaussian sequence model, the quadratic density integral
functional (with the extra condition $r_{n}^{2} \gtrsim n^{-1}$ for these two
examples; see Theorem~\ref{thm:high level} for explanation), and the expected
conditional covariance. More recent follow-up papers
\citep{jin2025structure, bonvini2024doubly, jin2025sharp, gu2026optimally, gu2025open} establish the
minimaxity of $\widehat{\psi}_{1,n}$ under the SA model for the average treatment
effect, the average treatment effect on the treated, and other related
parameters. Furthermore, the estimator $\widehat{\psi}_{1, n}$ is the same and
minimax in $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})$ for all values of
$r_{n}$. The minimaxity of $\widehat{\psi}_{1, n}$ under the SA model
$\mathcal{P}_{\mathrm{SA}}(\widehat{\theta}, r_{n})$ has been used by some
analysts to justify the use of current practice in (bio)statistics and econometrics.

\subsection*{Our Contribution}

Our main technical contribution of this paper is related to the
decision-theoretic properties of estimators, tracing back to the classical
work of Abraham Wald
\citep{wald1941principles, wald1945statistical, wald1947essentially}. Wald is
renowned as the inventor of the minimax principle \citep{wald1945statistical},
arguably the most popular theoretical paradigm used to measure the quality of
an estimator by modern (bio)statisticians and econometricians
\citep{brown1994minimaxity, andrews2021model, adusumilli2026sample} and
adopted in \citet{balakrishnan2026fundamental}.

However, certain minimax estimators may be inadmissible. One celebrated
example of the difference between minimaxity and (in)admissibility is Stein's
paradox, asserting that the maximum likelihood estimator (MLE), although
minimax, is everywhere dominated in mean squared error (MSE) for every sample
size $n$ by the James-Stein (JS) estimator in the many-normal-means model of
dimension at least three
\citep{stein1956inadmissibility, james1961estimation, brown1971admissible}.
Hence, the MLE is inadmissible. In this paper, we will show that, under
$\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})$, there exists a class of
parameters $\psi$, which we refer to as the \emph{monotone bias class}, for
which higher-order influence function (HOIF) estimators
\citep{robins2008higher, robins2016technical, liu2017semiparametric} dominate the first-order DML
estimator $\widehat{\psi}_{1,n}$ in the large-$n$ limit whenever $\prod_{j =
1}^{J} r_{n, j} \gtrsim n^{-1}$ (see Theorem~\ref{thm:high level} for the
actual statement). That is, we show that in the $\mathcal{P}_{\mathrm{SA}
}(\widehat{\theta},r_{n})$ model, the mimimax DML estimator is asymptotically
inadmissible, when estimating a parameter in the \emph{monotone bias class}.
To make the above claims precise, we next define: \emph{asymptotic
(in)admissibility} and the \emph{monotone bias class} of functionals. We shall
see that two out of the three functionals studied in
\citet{balakrishnan2026fundamental} are in the \emph{monotone bias class}.

\begin{definition}
[Asymptotic (in)admissibility]\label{def:admissible} An estimator sequence
$\psi_{n}$ of $\psi(\theta)$ indexed by $n$ is said to be \emph{asymptotically
inadmissible} in scaled MSE loss under model $\mathcal{P}_{\mathrm{SA}}
(\widehat{\theta}, r_{n})$ if there exists another estimator sequence $\psi
_{n}^{\prime}$ of $\psi(\theta)$ also indexed by $n$ such that
\[
\sup_{\theta\in\mathcal{P}_{\mathrm{SA}} (\widehat{\theta}, r_{n})} \limsup_{n
\rightarrow\infty} \frac{\mathsf{mse} (\psi_{n}^{\prime})-\mathsf{mse}
(\psi_{n})}{\mathsf{mse} (\psi_{n})} \leq0 \text{ and } \inf_{\theta
\in\mathcal{P}_{\mathrm{SA}}(\widehat{\theta}, r_{n})} \limsup_{n \rightarrow
\infty} \frac{\mathsf{mse}(\psi_{n}^{\prime}) - \mathsf{mse} (\psi_{n}
)}{\mathsf{mse} (\psi_{n})} < 0,
\]
where $\mathsf{mse} (\cdot) \equiv\mathsf{mse}_{\theta} (\cdot)
\coloneqq \mathsf{E}_{\theta} (\cdot-\psi(\theta))^{2}$. If otherwise, we say
that $\psi_{n}$ is \emph{asymptotically admissible}.
\end{definition}

In the above definition, the difference in MSEs is scaled. Because the MSE of
a reasonable estimator should decay to zero under the large $n$ limit,
obtaining nontrivial results requires an appropriate scaling. Here, we choose
the MSE of $\psi_{n}$ as a natural scaling factor. We refer the interested
readers to Appendix~\ref{app:scale} for a more elaborate discussion on the
choice of the denominator.



The following definition of the \emph{monotone bias class} is different from,
but as explained below, is essentially equivalent to that in our previous work \citep{liu2020nearly}.

\begin{definition}
\label{def:mbc} Let $\mathsf{bias} (\widehat{\psi}_{1, n})$ and $\mathsf{var}
(\widehat{\psi}_{1, n})$ be, respectively, the bias and variance of the
first-order DML estimator $\widehat{\psi}_{1, n}$ of $\psi$. $\psi$ is said to be
in the \emph{monotone bias class} if the following hold:

\begin{enumerate}
[label = (\arabic*)]

\item For any $r_{n}$, $\vert\mathsf{bias} (\widehat{\psi}_{1, n}) \vert
\lesssim\prod_{j = 1}^{J} r_{n, j}$, there always exists another estimator,
denoted by $\widehat{\psi}_{2, n}$, such that $\vert\mathsf{bias} (\widehat{\psi}_{2,
n}) \vert/ \vert\mathsf{bias} (\widehat{\psi}_{1, n}) \vert- 1 \leq0$, and the inequality becomes strict (asymptotically) at some law in $\mathcal{P}
_{\mathrm{SA}} (\widehat{\theta}, r_{n})$;

\item The variances of $\widehat{\psi}_{1,n}$ and $\widehat{\psi}_{2,n}$ satisfy the
following condition: there exists a constant $v>0$ such that $\mathsf{var}
(\widehat{\psi}_{1,n}) = v/n$ and
\begin{equation}
\begin{split}
\limsup_{n\rightarrow\infty}\frac{\mathsf{var}(\widehat{\psi}_{2,n})}
{\mathsf{var}(\widehat{\psi}_{1,n})}=\left\{
\begin{array}
[c]{ll}
1 & \text{if } \dfrac{|\mathsf{bias}(\widehat{\psi}_{1,n})|-|\mathsf{bias}
(\widehat{\psi}_{2,n})|}{v}=o(1),\\
\delta & \text{if } \dfrac{|\mathsf{bias}(\widehat{\psi}_{1,n})|-|\mathsf{bias}
(\widehat{\psi}_{2,n})|}{v} \gtrsim1,
\end{array}
\right. \label{var bound}
\end{split}
\end{equation}
for some constant $\delta\neq1$ but possibly $\delta>1$.

\end{enumerate}

Here the constants $v$ and $\delta$ can depend on the data generating
distribution $\mathbb{P}_{\theta}$.
\end{definition}

We are now ready to state the following general theorem regarding the
asymptotic inadmissibility of the DML estimator $\widehat{\psi}_{1, n}$. The proof
is given following a few remarks. \setcounter{theorem}{-1}

\begin{theorem}
\label{thm:high level} For $\psi$ belonging to the \emph{monotone bias class},
under the SA model $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})$, there
exists another estimator $\widehat{\psi}_{2,n}$ such that (i) $\widehat{\psi}_{2,n}$
is asymptotically not greater in scaled MSE than $\widehat{\psi}_{1,n}$ over
$\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})$ for any $r_{n}$, and (ii)
$\widehat{\psi}_{2,n}$ is asymptotically strictly smaller than $\widehat{\psi}_{1,n}$
in scaled MSE at some law in $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})$
whenever $\prod_{j=1}^{J}r_{n,j}\gtrsim n^{-1}$. Hence, $\widehat{\psi}_{1, n}$ is
asymptotically inadmissible if and only if $\prod_{j=1}^{J}r_{n,j}\gtrsim
n^{-1}$.
\end{theorem}

It should be noted that Theorem~\ref{thm:high level} does not discuss
minimaxity of either estimator. As we shall see, for two examples in the
\emph{monotone bias class} -- the quadratic functional in the Gaussian
sequence model in Section~\ref{sec:gsm} and the quadratic density integral
functional in Section~\ref{sec:quad}, the minimaxity of $\widehat{\psi}_{1, n}$ or
$\widehat{\psi}_{2, n}$ requires an additional condition $\sqrt{\prod_{j = 1}^{J}
r_{n, j}} \gtrsim n^{-1}$. As explained in
\citet{balakrishnan2026fundamental}, when $\sqrt{\prod_{j = 1}^{J} r_{n, j}} \ll n^{-1}$, the so-called plug-in estimators dominate both $\widehat{\psi}_{1, n}$ and $\widehat{\psi}_{2, n}$ in these two examples, because the plug-in estimators have zero variance and squared bias $\sqrt{\prod_{j = 1}^{J} r_{n, j}} \ll n^{-1}$, while both $\widehat{\psi}_{1, n}$ and $\widehat{\psi}_{2, n}$ have variances of order $n^{-1}$. However, as also noted by \citet{balakrishnan2026fundamental}, $\sqrt{\prod_{j = 1}^{J} r_{n, j}} \ll n^{-1}$ generally does not hold if the sample size used to estimate $\widehat{\theta}$ is of the same order as $n$. Therefore, there is essentially no loss of generality if we exclude the case where $\sqrt{\prod_{j = 1}^{J} r_{n, j}} \ll n^{-1}$ holds, as we do in Sections~\ref{sec:gsm} and \ref{sec:quad}, in which case both $\widehat{\psi}_{1, n}$ and $\widehat{\psi}_{2, n}$ are minimax rate-optimal.

\begin{remark}
\label{rem:high}
In all the examples covered in this paper and in
\citet{balakrishnan2026fundamental}, $|\mathsf{bias}(\widehat{\psi}_{1,n})|^{2}$
is upper bounded by $\prod_{j=1}^{J} r_{n,j}$ times a constant. In view of
Theorem~\ref{thm:high level}, when $\psi$ is in the \emph{monotone bias
class}, $\widehat{\psi}_{1,n}$ is \emph{asymptotically inadmissible} and is
asymptotically dominated by $\widehat{\psi}_{2,n}$ when $|\mathsf{bias}(\widehat{\psi
}_{1,n})|\gtrsim n^{-1/2}$. When $|\mathsf{bias} (\widehat{\psi}_{1,n})| \ll
n^{-1/2}$, $\widehat{\psi}_{2,n}$ is asymptotically never worse than $\widehat{\psi}_{1,n}$ and both are minimax. In particular, when DML estimators $\widehat{\psi
}_{1, n}$ are deployed in practice, it is often implicitly assumed that
$\prod_{j = 1}^{J} r_{n, j} \ll n^{-1}$ holds, a condition often referred to
as the \emph{rate-double-robustness} when $J = 2$. Under this condition,
neither estimator dominates the other; further, $|\mathsf{bias} (\widehat{\psi
}_{1,n})| \ll n^{-1 / 2}$ automatically holds, and thus both $\widehat{\psi}_{1,
n}$ and $\widehat{\psi}_{2, n}$ can be used to construct an asymptotically valid
Wald confidence interval (CI) of length $O (n^{-1 / 2})$, which is common
practice, although other non-Wald CI constructions are also under rapid
development \citep{zheng2025perturbed}.
\end{remark}

Since in practice $r_{n}$ is unknown and the possibility that it is of order $1$ cannot be empirically excluded \citep{robins1997toward, ritov2014bayesian}, we consider the following assumption-lean model.
\begin{definition}
\label{def:al}
Given a constant $R^{\ast} > 0$, we define the model $\mathcal{P}_{\rm AL} (\widehat{\theta}) \equiv \mathcal{P}_{\rm SA} (\widehat{\theta}, r_{n} = (R^{\ast}, \cdots, R^{\ast}))$ as the assumption-lean model \citep{liu2024assumption}.
\end{definition}

Under the assumption-lean model, for parameters in the \emph{monotone bias class}, it follows from Theorem~\ref{thm:high level} that $\widehat{\psi}_{2,n}$
dominates $\widehat{\psi}_{1,n}$ asymptotically, because $\prod_{j=1}^{J}
r_{n,j}=1\gg n^{-1}$, and thus $\widehat{\psi}_{1,n}$ is asymptotically
inadmissible.
In fact, as discussed in Section~\ref{sec:sa} below, we can say more. Given a
parameter $\psi$ in the \emph{monotone bias class}, we can construct an
asymptotically level-$\alpha$ falsification test of the null hypothesis
$\mathcal{H}_{0}$: $\mathsf{bias}(\widehat{\psi}_{1,n})\ll n^{-1/2}$ that, when
$\mathcal{H}_{0}$ is rejected, provides empirical evidence for the alternative hypothesis $|\mathsf{bias}(\widehat{\psi}_{1,n})|\gtrsim n^{-1/2}$ and thus also empirical evidence that the Wald CI centered on $\widehat{\psi}_{1,n}$ under-covers even in large samples \citep{liu2020nearly}. However, by the aforementioned results of \citet{robins1997toward} and \citet{ritov2014bayesian}, any such test must be inconsistent under the assumption-lean model $\mathcal{P}_{\mathrm{AL}} (\widehat{\theta})$. Hence, failure to reject does not provide evidence for or against the null hypothesis $\mathcal{H}_{0}$ even asymptotically. We return to this issue in the concluding section of the paper (Section~\ref{sec:sa}).

\begin{remark}
\label{rem:intuition}
It will be clear in later sections that all examples studied by \citet{balakrishnan2026fundamental} are in the \emph{monotone bias class} defined in Definition~\ref{def:mbc}, except for the expected conditional covariance (see Section~\ref{sec:ecc}). The two parts of conditions in Definition~\ref{def:mbc} need further elaboration. Part (1) states that the bias of $\widehat{\psi}_{2, n}$ is never greater than but sometimes strictly smaller than that of $\widehat{\psi}_{1, n}$. Part (2) says that the difference between the variances of $\widehat{\psi}_{1, n}$ and $\widehat{\psi}_{2, n}$ is negligible (of order $o (1 / n)$) if the bias reduction of $\widehat{\psi}_{2, n}$ is negligible (of order $o (1)$). The original definition of the \emph{monotone bias class} in \citet{liu2020nearly} contains only part (1) but part (2) was implicit. As described later, the alternative estimator $\widehat{\psi}_{2, n}$ is a second-order $U$-statistic (heretofore referred to as the second-order estimators for simplicity), constructed via the theory of HOIFs \citep{robins2008higher, robins2016technical, liu2017semiparametric}. Such second-order estimators satisfy both Parts (1) and (2). The difference between the variances of $\widehat{\psi}_{2, n}$ and $\widehat{\psi}_{1, n}$ can be bounded as follows, as is the case in all our examples in later sections:
\[
\mathsf{var}(\widehat{\psi}_{2,n})-\mathsf{var}(\widehat{\psi}_{1,n})\lesssim\frac
{k}{n^{2}}+\frac{|\mathsf{bias}(\widehat{\psi}_{1,n})|-|\mathsf{bias}(\widehat{\psi
}_{2,n})|}{n}+\frac{v^{1/2}\{|\mathsf{bias}(\widehat{\psi}_{1,n})|-|\mathsf{bias}
(\widehat{\psi}_{2,n})|\}^{1/2}}{n},
\]
where $k$ is a tuning parameter chosen so that $k=o(n)$. This will be
demonstrated in the proofs of Theorem~\ref{thm:quad gsm}
--Theorem~\ref{thm:ecv} in Appendix~\ref{app:proofs}. It is then not difficult
to verify that Part (2) of Definition~\ref{def:mbc} holds. In terms of Part
(1), the second-order estimator $\widehat{\psi}_{2,n}$ can be viewed as debiasing
$\widehat{\psi}_{1,n}$ by estimating a part of $\mathsf{bias}(\widehat{\psi}_{1,n})$;
also see comments after Lemma~\ref{lem:quad gsm 2}, \ref{lem:quad2}, and
\ref{lem:ecc2}, and Theorem~\ref{thm:ecv}. For functionals outside the
\emph{monotone bias class} such as the expected conditional covariance
functional covered in Section~\ref{sec:ecc}, $\widehat{\psi}_{2,n}$ is still rate-optimal but may have a bias exceed that of $\widehat{\psi}_{1,n}$ under certain laws in $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})$ .
\end{remark}

With these ingredients, the proof of Theorem~\ref{thm:high level} is almost
immediate by elementary calculations, so we record the proof here.

\begin{proof}[Proof of Theorem~\ref{thm:high level}]
The result follows directly from Definition~\ref{def:mbc}. To see this, we first decompose the scaled MSE difference as
\begin{align*}
\frac{\mathsf{mse} (\widehat{\psi}_{2, n}) - \mathsf{mse} (\widehat{\psi}_{1, n})}{\mathsf{mse} (\widehat{\psi}_{1, n})} = - \frac{\mathsf{bias} (\widehat{\psi}_{1, n})^{2} - \mathsf{bias}^{2} (\widehat{\psi}_{2, n})}{\mathsf{bias} (\widehat{\psi}_{1, n})^{2} + \mathsf{var} (\widehat{\psi}_{1, n})} + \frac{\mathsf{var} (\widehat{\psi}_{2, n}) - \mathsf{var} (\widehat{\psi}_{1, n})}{\mathsf{bias} (\widehat{\psi}_{1, n})^{2} + \mathsf{var} (\widehat{\psi}_{1, n})} \coloneqq T_{1, n} + T_{2, n}.
\end{align*}
When $\dfrac{|\mathsf{bias} (\widehat{\psi}_{1, n})| - |\mathsf{bias} (\widehat{\psi}_{2, n})|}{n \cdot \mathsf{var} (\widehat{\psi}_{1, n})} = o (1)$ holds, by \eqref{var bound}, we always have:
\begin{align*}
\limsup_{n \rightarrow \infty} T_{2, n} = \limsup_{n \rightarrow \infty} \frac{\mathsf{var} (\widehat{\psi}_{2, n}) / \mathsf{var} (\widehat{\psi}_{1, n}) - 1}{\mathsf{bias} (\widehat{\psi}_{1, n})^{2} / \mathsf{var} (\widehat{\psi}_{1, n}) + 1} = 0.
\end{align*}
Since $T_{1, n}$ is always non-positive, we have $\limsup_{n \rightarrow \infty} T_{1, n} + T_{2, n} \leq 0$. In words, when the bias reduction is sufficiently small, by (2) of Definition~\ref{def:mbc}, the asymptotic variance of the second-order estimator is not different from that of the first-order DML estimator. On the contrary, we now suppose that $|\mathsf{bias} (\widehat{\psi}_{1, n})| - |\mathsf{bias} (\widehat{\psi}_{2, n})| \gtrsim n \cdot \mathsf{var} (\widehat{\psi}_{1, n})$, and hence also $|\mathsf{bias} (\widehat{\psi}_{1, n})| \gtrsim n \cdot \mathsf{var} (\widehat{\psi}_{1, n})$. We must also have
\begin{align*}
\mathsf{bias} (\widehat{\psi}_{1, n})^{2} - \mathsf{bias} (\widehat{\psi}_{2, n})^{2} \gg \mathsf{var} (\widehat{\psi}_{2, n}) - \mathsf{var} (\widehat{\psi}_{1, n}),
\end{align*}
and by the non-positivity of $T_{1, n}$, we have $\limsup_{n \rightarrow \infty} T_{1, n} + T_{2, n} \leq 0$. When $|\mathsf{bias} (\widehat{\psi}_{1, n})|^{2} \lesssim \prod_{j = 1}^{J} r_{n, j} \ll n^{-1}$, the denominator in the scaled MSE difference is dominated by $\mathsf{var} (\widehat{\psi}_{1, n}) = v / n$. The numerator can only be of smaller order than the denominator, rendering $\limsup_{n \rightarrow \infty} T_{1, n} + T_{2, n} = 0$. When $\prod_{j = 1}^{J} r_{n, j} \gtrsim n^{-1}$, there must exist a law such that $\mathsf{bias} (\widehat{\psi}_{1, n}) \gtrsim n^{-1 / 2}$. By Definition~\ref{def:mbc}, there must exist a distribution for which $\limsup_{n \rightarrow \infty} T_{1, n} < 0$, which completes the proof.
\end{proof}


\section*{Notation and Organization}

Throughout this paper, we always let $C > 0$ denote a sufficiently large
constant independent of the sample size $n$. For any functions mentioned in
the paper, they are understood to be squared integrable with respect to the
Lebesgue measure. $\mathbb{U}_{n, m} [\cdot]$ denotes a $m$-th order
$U$-statistic operator. Given a collection of $k$ different functions $\bar
{f}_{k} = (f_{1}, \cdots, f_{k})^{\top}$, we let $\mathsf{\Pi}_{\mathbb{P}} (\cdot \mid \bar{f}_{k})$ denote the operator of $L^{2} (\mathbb{P})$-projection
onto the linear span of $\bar{f}_{k}$ and $\Vert\cdot\Vert_{2, \mathbb{P}}$
denote the $L^{2} (\mathbb{P})$-norm. If $\mathbb{P}$ is the Lebesgue measure, we omit the subscript and write $\mathsf{\Pi} (\cdot \mid \bar{f}_{k})$ and $\Vert\cdot\Vert_{2}$ for short. $\Vert\cdot\Vert_{2}$ also denotes the
$\ell^{2}$-norm of a vector. We denote the population Gram matrix of $\bar
{f}_{k}$ under the distribution $\mathbb{P}$ as $\Sigma_{\mathbb{P}, \bar
{f}_{k}} \coloneqq \mathsf{E} [\bar{f}_{k} (X) \bar{f}_{k} (X)^{\top}]$. When
it is clear from the context, we omit the dependence in the subscript on
$\mathbb{P}$ or $\bar{f}_{k}$ or both.

The remainder of this paper makes Theorem~\ref{thm:high level} concrete.
Sections~\ref{sec:gsm}--\ref{sec:ecc} cover four examples of $\psi$, one of
which does not belong to the \emph{monotone bias class}.
Theorem~\ref{thm:high level} will then be specialized for the three examples
in the \emph{monotone bias class}. Section~\ref{sec:sa} concludes the paper by
making some additional comments on the relevance of the SA model to
practitioners who are more interested in uncertainty quantification or
statistical inference. Proofs are deferred to the Appendix.

\section{Quadratic Functional in the Gaussian Sequence Model}

\label{sec:gsm}

As in \citet{balakrishnan2026fundamental}, we observe data drawn from the
infinite Gaussian sequence model:
\begin{align}
\label{gsm}Y_{i} = \theta_{i} + \varepsilon_{i}, i = 1, 2, \cdots
\end{align}
where $\{\varepsilon_{i}, i = 1, 2, \cdots\} \overset{\mathrm{i.i.d.}}{\sim}
\mathrm{N} (0, n^{-1})$. Let $\theta\coloneqq \{\theta_{i}, i = 1, 2,
\cdots\}$ and $Y \coloneqq \{Y_{i}, i = 1, 2, \cdots\}$. We are interested in
learning about the quadratic functional
\begin{align}
\label{gsm:quad}Q (\theta) \coloneqq \Vert\theta\Vert_{2}^{2} \equiv\sum_{i =
1}^{\infty} \theta_{i}^{2}.
\end{align}


The SA model corresponding to $\psi(\theta)$ is defined by
\citet{balakrishnan2026fundamental} as:
\begin{align}
\label{quad gsm model}\mathcal{P}_{\mathrm{SA}} ((\widehat{\theta}, \widehat{\theta}),
(r_{n}, r_{n})) \coloneqq \left\{  \theta: \Vert\widehat{\theta} - \theta\Vert
_{2}^{2} \leq r_{n} \right\} ,
\end{align}
where $\widehat{\theta}$ is some initial estimator of $\theta$. It is noteworthy
that we deliberately write $\widehat{\theta}$ and $r_{n}$ twice in the notation
$\mathcal{P}_{\mathrm{SA}} ((\widehat{\theta}, \widehat{\theta}), (r_{n}, r_{n}))$ to
emphasize that we take $J = 2$ and $\prod_{j = 1}^{J} r_{n, j} = r_{n}^{2}$ in
this case. In the sequel, however, we write $\mathcal{P}_{\mathrm{SA}}
(\widehat{\theta}, r_{n})$ instead in the text to simplify the notation. We adopt
a similar convention for the quadratic density integral functional in
Section~\ref{sec:quad} and the expected conditional variance in
Section~\ref{sec:ecc}.

\citet{balakrishnan2026fundamental} obtained the following results.

\begin{lemma}
\label{lem:quad gsm} The following hold. When $r_{n} \gtrsim n^{-1}$,
\begin{align*}
\mathfrak{R}_{n} (Q; \mathcal{P}_{\mathrm{SA}} (\widehat{\theta}, r_{n}))
\coloneqq \inf_{\widehat{Q}} \sup_{\theta\in\mathcal{P}_{\mathrm{SA}} (\widehat
{\theta}, r_{n})} \mathsf{E}_{\theta} [(\widehat{Q} - Q (\theta))^{2}] \gtrsim
r_{n}^{2} + \frac{\Vert\widehat{\theta} \Vert_{2}^{2}}{n}.
\end{align*}
This lower bound is attained by the first-order estimator $\widehat{Q}_{1, n}$,
defined as
\begin{align*}
\widehat{Q}_{1, n} \coloneqq 2 \langle Y, \widehat{\theta} \rangle- \Vert\widehat{\theta}
\Vert_{2}^{2}.
\end{align*}
The bias, variance, and mean squared error (MSE) of $\widehat{Q}_{1, n}$ have the
following forms:
\begin{align*}
\mathsf{bias} (\widehat{Q}_{1, n}) = - \Vert\widehat{\theta} - \theta\Vert_{2}^{2},
\mathsf{var} (\widehat{Q}_{1, n}) = \frac{4}{n} \Vert\widehat{\theta} \Vert_{2}^{2},
\text{ and } \mathsf{mse} (\widehat{Q}_{1, n}) = \Vert\widehat{\theta} - \theta
\Vert_{2}^{4} + \frac{4}{n} \Vert\widehat{\theta} \Vert_{2}^{2}.
\end{align*}

\end{lemma}

The lower and upper bounds can be found in Theorem~1, Part~1 and Theorem~2,
Part~1 of \citet{balakrishnan2026fundamental}, respectively. These results,
taken together, prove the rate optimality of $\widehat{Q}_{1, n}$ in the minimax sense.

To show the asymptotic inadmissibility of the minimax estimator $\widehat{Q}
_{1,n}$, we need to exhibit a different estimator that improves upon $\widehat
{Q}_{1,n}$. To this end, we adopt the following second-order estimator
appeared in \cite{robins2006adaptive}:
\begin{align*}
\widehat{Q}_{2,n}(k)  &  \coloneqq\sum_{i=1}^{k}Y_{i,1}Y_{i,2}+\sum_{i=k+1}
^{\infty}\left(  2Y_{i}\widehat{\theta}_{i}-\widehat{\theta}_{i}^{2}\right) \\
&  \equiv\widehat{Q}_{1,n}+\sum_{i=1}^{k}\left(  Y_{i,1}Y_{i,2}-2Y_{i}\widehat{\theta
}_{i}+\widehat{\theta}_{i}^{2}\right) ,
\end{align*}
where $Y_{i,1}\coloneqq Y_{i}+\Phi^{-1}(U_{i})/\sqrt{n}$, $Y_{i,2}
\coloneqq Y_{i}-\Phi^{-1}(U_{i})/\sqrt{n}$, $\Phi$ is the standard normal
cumulative distribution function and $U_{i}$'s are independent uniform random
variables over $[0,1]$. Here $Y_{i,1}
\mathpalette{\protect\independenT}{\perp}Y_{i,2}$. The difference between
$\widehat{Q}_{2,n}(k)$ and $\widehat{Q}_{1,n}$ is an unbiased estimator of
$\Vert \mathsf{\Pi}_{k}(\theta-\widehat{\theta}) \Vert_{2}^{2}$, where $\mathsf{\Pi}_{k}(\cdot)$ denotes the projection onto the first $k$ coordinates of the
input infinite-dimensional vector, with $\mathsf{\Pi}_{k}^{\perp} (\cdot)$
naturally meaning the projection onto the $(k+1)$-th coordinate and onward.

The following lemma characterizes the bias, variance, and mean squared error
of $\widehat{Q}_{2, n} (k)$. The proof can be found in
Appendix~\ref{app:lem:quad gsm 2}.

\begin{lemma}
\label{lem:quad gsm 2} The bias and variance of $\widehat{Q}_{2, n} (k)$ read as:
\begin{align*}
\mathsf{bias} (\widehat{Q}_{2, n} (k))  &  = - \Vert\mathsf{\Pi}_{k}^{\perp}
(\widehat{\theta} - \theta) \Vert_{2}^{2} \lesssim r_{n},\\
\mathsf{var} (\widehat{Q}_{2, n} (k))  &  = \frac{4}{n} \Vert\widehat{\theta}
\Vert_{2}^{2} + \frac{4 k}{n^{2}} + \frac{4}{n} \Vert\mathsf{\Pi}_{k}
(\widehat{\theta} - \theta) \Vert_{2}^{2} - \frac{4}{n} \langle\mathsf{\Pi}_{k}
\widehat{\theta}, \mathsf{\Pi}_{k} (\widehat{\theta} - \theta) \rangle\lesssim\frac
{1}{n},
\end{align*}
where the inequalities hold for $\theta\in\mathcal{P}_{\mathrm{SA}}
(\widehat{\theta}, r_{n})$.
\end{lemma}

We note that $\Vert\mathsf{\Pi}_{k} (\widehat{\theta} - \theta) \Vert_{2}^{2}
\equiv\mathsf{bias} (\widehat{Q}_{1, n}) - \mathsf{bias} (\widehat{Q}_{2, n} (k))$ so
$\widehat{Q}_{2, n} (k)$ corrects the bias of $\widehat{Q}_{1, n}$ by estimating a
lower bound of $\mathsf{bias} (\widehat{Q}_{1, n}) \lesssim r_{n}$. By
Lemma~\ref{lem:quad gsm 2}, $Q (\theta)$ belongs to the \emph{monotone bias
class}. Comparing $\mathsf{mse} (\widehat{Q}_{2, n} (k))$ and $\mathsf{mse}
(\widehat{Q}_{1, n})$ in the asymptotic sense, we obtain the first main
statistical result of this paper. There always exists a distribution in
$\mathcal{P}_{\mathrm{SA}} (\widehat{\theta}, r_{n})$ such that
Definition~\ref{def:mbc}(1) holds. To see this, consider the case where
$\theta= \widehat{\theta} + r_{n}^{1 / 2} \upsilon$, where $\Vert\upsilon\Vert_{2}
= 1$ and the coordinates of $\upsilon$ from $k + 1$ onward are all zeros. With
this choice, $\mathsf{bias} (\widehat{Q}_{2, n} (\bar{\phi}_{k})) = 0$. The rest
of the proof can be found in Appendix~\ref{app:quad gsm}.

\begin{theorem}
\label{thm:quad gsm} Under Model $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},
r_{n})$ with $r_{n} \gtrsim n^{-1}$, $\widehat{Q}_{2,n}(k)$ is
\emph{asymptotically minimax} and the following hold as long as $k$ is chosen
such that $k = o(n\Vert\widehat{\theta} \Vert_{2}^{2})$.
\begin{align*}
&  \sup_{\theta\in\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})}
\limsup_{n\rightarrow\infty}\frac{\mathsf{mse}(\widehat{Q}_{2,n}(k))-\mathsf{mse}
(\widehat{Q}_{1,n})}{\mathsf{mse} (\widehat{Q}_{1,n})}\leq0,\text{ and when $r_{n}^{2}
\gtrsim n^{-1}$}\\
&  \inf_{\theta\in\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})}
\limsup_{n\rightarrow\infty}\frac{\mathsf{mse}(\widehat{Q}_{2,n}(k))-\mathsf{mse}
(\widehat{Q}_{1,n})}{\mathsf{mse}(\widehat{Q}_{1,n})}<0.
\end{align*}
Thus, by Definition~\ref{def:admissible}, the first-order DML estimator
$\widehat{Q}_{1,n}$ is \emph{asymptotically inadmissible} when $r_{n}^{2} \gtrsim
n^{-1}$. The same conclusions hold when we replace the SA model $\mathcal{P}
_{\mathrm{SA}} (\widehat{\theta}, r_{n})$ with the assumption-lean model
$\mathcal{P}_{\mathrm{AL}} (\widehat{\theta})$ and drop the assumptions on $r_{n}$.
\end{theorem}

Echoing the comment right after Theorem~\ref{thm:high level}, for the
minimaxity of $\widehat{Q}_{1, n}$ or $\widehat{Q}_{2, n} (k)$, we need $r_{n} \gtrsim
n^{-1}$. When $r_{n} \ll n^{-1}$, the so-called plug-in estimator $\widehat
{Q}_{\mathrm{pi}} \coloneqq \Vert\widehat{\theta} \Vert^{2}$ has zero variance and
squared bias of order $r_{n}$, thus dominating both $\widehat{Q}_{1, n}$ and
$\widehat{Q}_{2, n} (k)$ when $\Vert\widehat{\theta} \Vert^{2}$ is of order 1.
However, $r_{n} \ll n^{-1}$, or equivalently $\Vert\widehat{\theta} - \theta
\Vert\ll n^{-1 / 2}$, is generally difficult to hold if the sample used to
compute $\widehat{\theta}$ is of size similar to $n$, as such a condition says
that we can estimate the possibly infinite-dimensional $\theta$ at a rate much
faster than the parametric rate. A similar discussion also applies to the
quadratic density integral functional to be discussed next.

\section{Quadratic Density Integral Functional}

\label{sec:quad}

The second example is about estimating the quadratic density integral
functional of the probability density function $f$ of $X$ based on $n$ i.i.d.
observations $\{X_{i} \in[0, 1]^{d}\}_{i = 1}^{n} \sim f$:
\begin{equation}
\label{quad}\psi(f) \coloneqq \int f (x)^{2} \mathrm{d} x.
\end{equation}


The SA model corresponding to $\psi(f)$ is defined by
\citet{balakrishnan2026fundamental} as:
\begin{align}
\label{quad model}\mathcal{P}_{\mathrm{SA}} ((\widehat{f}, \widehat{f}), (r_{n},
r_{n})) \coloneqq \left\{  f: \Vert\widehat{f} - f \Vert_{2}^{2} \leq r_{n}, \int
f (x) \mathrm{d} x = 1, f \geq0, \Vert\widehat{f} \Vert_{\infty} \leq C, \Vert f
\Vert_{\infty} \leq C \right\}  ,
\end{align}
where $\widehat{f}$ is some initial estimator of $f$ computed from a separate
independent sample treated as fixed. Similar to the case in
Section~\ref{sec:gsm}, we take $J = 2$ and $\prod_{j = 1}^{J} r_{n, j} =
r_{n}^{2}$, and write $\mathcal{P}_{\mathrm{SA}} (\widehat{f}, r_{n})$ instead.

\begin{remark}
\label{rem:cf1} As mentioned in footnote 1, we use the same estimator $\widehat{f}
\equiv\widehat{f}_{1} \equiv\widehat{f}_{2}$ to compute $\widehat{\psi}_{1, n}$, which is
the standard DML estimator commonly employed in the literature
\citep{chernozhukov2018double} but excludes more refined estimators with $f$
estimated by separate samples studied in \citet{newey2018cross, mcgrath2026nuisance, mcclean2026double}.
\end{remark}

The following lemma, paraphrasing the results of
\citet{balakrishnan2026fundamental}, summarizes the lower and upper bounds of
the error rate of estimating $\psi(f)$ under $\mathcal{P}_{\mathrm{SA}}
(\widehat{f}, r_{n})$.

\begin{lemma}
\label{lem:quad} The following hold. When $r_{n} \gtrsim n^{-1}$,
\begin{align*}
\mathfrak{R}_{n} (\psi; \mathcal{P}_{\mathrm{SA}} (\widehat{f}, r_{n}))
\coloneqq \inf_{\widehat{\psi}} \sup_{f \in\mathcal{P}_{\mathrm{SA}} (\widehat{f},
r_{n})} \mathsf{E}_{f} [(\widehat{\psi} - \psi(f))^{2}] \gtrsim r_{n}^{2} +
\frac{1}{n} \left(  \Vert\widehat{f} \Vert_{3}^{3} - \Vert\widehat{f} \Vert_{2}^{4}
\right) .
\end{align*}
This lower bound is attained by the first-order estimator $\widehat{\psi}_{1, n}$,
defined as
\begin{align*}
\widehat{\psi}_{1, n} \coloneqq \frac{2}{n} \sum_{i = 1}^{n} \widehat{f} (X_{i}) -
\int\widehat{f} (x)^{2} \mathrm{d} x.
\end{align*}
The bias, variance and MSE of $\widehat{\psi}_{1, n}$ have the following forms:
\begin{align*}
\mathsf{bias} (\widehat{\psi}_{1, n})  &  = - \int(\widehat{f} (x) - f (x))^{2}
\mathrm{d} x \equiv- \Vert\widehat{f} - f \Vert_{2}^{2},\\
\mathsf{var} (\widehat{\psi}_{1, n})  &  = \frac{4}{n} \mathsf{var} (\widehat{f} (X))
\equiv\frac{4}{n} \left\{  \int\widehat{f} (x)^{2} f (x) \mathrm{d} x - \left(
\int\widehat{f} (x) f (x) \mathrm{d} x \right)  ^{2} \right\}  , \text{ and}\\
\mathsf{mse} (\widehat{\psi}_{1, n})  &  = \Vert\widehat{f} - f \Vert_{2}^{4} +
\frac{4}{n} \mathsf{var} (\widehat{f} (X)).
\end{align*}

\end{lemma}

The lower and upper bounds can be found in Theorem~1, Part~2 and Theorem~2,
Part~2 of \citet{balakrishnan2026fundamental}, respectively. These results,
taken together, prove the optimality of $\widehat{\psi}_{1, n}$ in the minimax sense.

To show the asymptotic inadmissibility of the minimax estimator $\widehat{\psi
}_{1, n}$, when $r_{n}^{2} \gtrsim n^{-1}$ or equivalently $r_{n}^{1 / 2}
\gtrsim n^{-1 / 4}$, we exhibit a different estimator that improves on
$\widehat{\psi}_{1, n}$. To this end, let $\bar{\phi}_{k} \coloneqq (\phi_{1},
\cdots, \phi_{k})^{\top}$ be a $k$-dimensional orthonormal basis with respect
to the Lebesgue measure \citep{chen2007large}. We then construct the following
second-order $U$-statistic estimator:
\begin{align*}
\widehat{\psi}_{2, n} (\bar{\phi}_{k})  &  \coloneqq \widehat{\psi}_{1, n} +
\mathbb{U}_{n, 2} \left[  \left(  \bar{\phi}_{k} (X_{1}) - \int\bar{\phi}_{k}
(x) \widehat{f} (x) \mathrm{d} x \right)  ^{\top} \left(  \bar{\phi}_{k} (X_{2}) -
\int\bar{\phi}_{k} (x) \widehat{f} (x) \mathrm{d} x \right)  \right] \\
&  \equiv\mathbb{U}_{n, 1} [2 (\widehat{f} (X) - \mathsf{\Pi}[\widehat{f} \mid \bar{\phi}_{k}] (X))] + \mathbb{U}_{n, 2} [\bar{\phi}_{k} (X_{1})^{\top} \bar{\phi}_{k}
(X_{2})] - \int\left(  \widehat{f} (x)^{2} - \mathsf{\Pi}[\widehat{f} \mid \bar{\phi}
_{k}] (x)^{2} \right)  \mathrm{d} x.
\end{align*}


\begin{remark}
\label{rem:quad2} Expert readers shall realize that $\widehat{\psi}_{2,n}
(\bar{\phi}_{k})$ debiases $\widehat{\psi}_{1,n}$ by subtracting an unbiased
estimator of a part of its bias, based on HOIFs. The part of the bias of
$\widehat{\psi}_{1,n}$ to be estimated is determined by the choice of $\bar{\phi
}_{k}$. We mention in passing that the falsification test of
\citet{liu2020nearly} mentioned earlier is essentially based on the statistic
$\widehat{\psi}_{2,n}(\bar{\phi}_{k})-\widehat{\psi}_{1,n}$. Similar tests or
estimators have also been considered in instrumental variable or proximal
causal inference settings \citep{breunig2024adaptive, liu2024assumption}.
\end{remark}

Let $\eta\coloneqq \int\bar{\phi}_{k} (x) f (x) \mathrm{d} x$ and $\widehat{\eta}
\coloneqq \int\bar{\phi}_{k} (x) \widehat{f} (x) \mathrm{d} x$. We also make the
following assumption on $\Sigma$.

\begin{assumption}
\label{as:Sigma quad} $\Sigma$ is assumed to have bounded spectra.
\end{assumption}

We now state the following lemma. The proof can be found in
Appendix~\ref{app:lem:quad2}.

\begin{lemma}
\label{lem:quad2} The bias and variance of $\widehat{\psi}_{2, n} (\bar{\phi}
_{k})$ read as:
\begin{align*}
\mathsf{bias} (\widehat{\psi}_{2, n} (\bar{\phi}_{k})) =  &  - \int(\widehat{f} (x) -
f (x))^{2} \mathrm{d} x + \int\mathsf{\Pi}[\widehat{f} - f \mid \bar{\phi}_{k}]
(x)^{2} \mathrm{d} x \equiv- \Vert\widehat{f} - f \Vert_{2}^{2} + \Vert
\mathsf{\Pi}[\widehat{f} - f \mid \bar{\phi}_{k}] \Vert_{2}^{2}\\
\equiv &  - \Vert\widehat{f} - f \Vert_{2}^{2} + \Vert\widehat{\eta} - \eta\Vert
_{2}^{2} \equiv- \Vert\mathsf{\Pi}^{\perp} [\widehat{f} - f \mid \bar{\phi}_{k}]
\Vert_{2}^{2} \lesssim r_{n},\\
\mathsf{var} (\widehat{\psi}_{2, n} (\bar{\phi}_{k})) =  &  \ \frac{4}{n}
\mathsf{var} [\widehat{f} (X)] + \frac{8}{n} \left(  \int\bar{\phi}_{k} (x) f (x)
\widehat{f} (x) \mathrm{d} x - \widehat{\eta} \right)  ^{\top} (\widehat{\eta} - \eta)\\
&  + \frac{2}{n (n - 1)} \left\{
\begin{array}
[c]{c}
\mathsf{Tr} (\Sigma^{2}) - 4 \widehat{\eta}^{\top} \Sigma\eta+ 2 \widehat{\eta}^{\top}
\Sigma\widehat{\eta} + 2 \widehat{\eta}^{\top} \widehat{\eta} \cdot\eta^{\top} \eta\\
+ \, 2 (\widehat{\eta}^{\top} \eta)^{2} - 4 \widehat{\eta}^{\top} \widehat{\eta} \cdot
\widehat{\eta}^{\top} \eta+ (\widehat{\eta}^{\top} \widehat{\eta})^{2}
\end{array}
\right\} \\
\leq &  \ \frac{4}{n} \mathsf{var} [\widehat{f} (X)] + \frac{C}{n} \Vert
\mathsf{\Pi} [\widehat{f} - f \mid \bar{\phi}_{k}] \Vert_{2} + \frac{C k}{n^{2}}
\lesssim\frac{1}{n},
\end{align*}
where the inequalities hold for $f \in\mathcal{P}_{\mathrm{SA}} (\widehat{f},
r_{n})$ and under Assumption~\ref{as:Sigma quad}.
\end{lemma}

We note that $\Vert\mathsf{\Pi}[\widehat{f} - f \mid \bar{\phi}_{k}] \Vert_{2}^{2}
\equiv\mathsf{bias} (\widehat{\psi}_{1, n}) - \mathsf{bias} (\widehat{\psi}_{2, n}
(\bar{\phi}_{k}))$ so $\widehat{\psi}_{2, n} (\bar{\phi}_{k})$ corrects the bias
of $\widehat{\psi}_{1, n}$ by estimating a lower bound of $\mathsf{bias}
(\widehat{\psi}_{1, n}) \lesssim r_{n}$. By Lemma~\ref{lem:quad gsm}, $\psi(f)$
belongs to the \emph{monotone bias class}. The following theorem therefore
instantiates Theorem~\ref{thm:high level} for the quadratic density integral
functional $\psi(f)$. There always exists a distribution in $\mathcal{P}
_{\mathrm{SA}} (\widehat{f}, r_{n})$ such that Definition~\ref{def:mbc}(1) holds.
To see this, consider the case where $f = \widehat{f} + r_{n}^{1 / 2} \beta^{\top}
\bar{\phi}_{k}$ with $\Vert\beta\Vert_{2} = 1$, for which $\mathsf{bias}
(\widehat{\psi}_{2, n} (\bar{\phi}_{k})) = 0$. The rest of the proof can be found
in Appendix~\ref{app:quad}.

\begin{theorem}
\label{thm:quad} Under model $\mathcal{P}_{\mathrm{SA}}(\widehat{f},r_{n})$ and
Assumption~\ref{as:Sigma quad}, $\widehat{\psi}_{2,n}(\bar{\phi}_{k})$ is
\emph{asymptotically minimax} and the following hold as long as $k$ is chosen
such that $k=o(n\mathsf{var}[\widehat{f}(X)])$:
\begin{align*}
&  \sup_{f\in\mathcal{P}_{\mathrm{SA}}(\widehat{f},r_{n})}\limsup_{n\rightarrow
\infty}\frac{\mathsf{mse}(\widehat{\psi}_{2,n}(\bar{\phi}_{k}))-\mathsf{mse}
(\widehat{\psi}_{1,n})}{\mathsf{mse}(\widehat{\psi}_{1,n})}\leq0,\text{ and when
$r_{n}^{2} \gtrsim n^{-1}$}\\
&  \inf_{f\in\mathcal{P}_{\mathrm{SA}}(\widehat{f},r_{n})}\limsup_{n\rightarrow
\infty}\frac{\mathsf{mse}(\widehat{\psi}_{2,n}(\bar{\phi}_{k}))-\mathsf{mse}
(\widehat{\psi}_{1,n})}{\mathsf{mse}(\widehat{\psi}_{1,n})}<0.
\end{align*}
Thus, by Definition~\ref{def:admissible}, the first-order DML estimator
$\widehat{\psi}_{1,n}$ is \emph{asymptotically inadmissible} when $r_{n}^{2}
\gtrsim n^{-1}$. The same conclusions hold when we replace the SA model
$\mathcal{P}_{\mathrm{SA}} (\widehat{f}, r_{n})$ with the assumption-lean model
$\mathcal{P}_{\mathrm{AL}} (\widehat{f})$ and drop the assumptions on $r_{n}$.
\end{theorem}

\section{Expected Conditional Covariance}

\label{sec:ecc}

All the functionals that we have analyzed so far fall within the
\emph{monotone bias class}.
In this section, we turn to the Expected Conditional Covariance (ECC)
functional, defined as
\begin{align*}
\psi(a, b) \coloneqq \mathsf{E} [\mathsf{cov} (A, Y \mid X)] \equiv\mathsf{E} [(A
- a (X)) (Y - b (X))],
\end{align*}
where $X \in[0, 1]^{d}$ denotes the baseline covariates, $A, Y \in\mathbb{R}$
are two types of responses, $a (\cdot) \coloneqq \mathsf{E} (A \mid X = \cdot)$
and $b (\cdot) \coloneqq \mathsf{E} (Y \mid X = \cdot)$. The ECC functional, as
extensively discussed in \citet{liu2020nearly}, is not in the \emph{monotone
bias class}. Therefore, not surprisingly, we can no longer conclude the
asymptotic inadmissibility of the first-order DML estimator $\widehat{\psi}_{1,
n}$ for $\psi(a, b)$. Specifically, based on $n$ i.i.d. observations $\{X_{i},
A_{i}, Y_{i}\}_{i = 1}^{n}$, $\widehat{\psi}_{1, n}$ reads as:
\begin{align*}
\widehat{\psi}_{1, n} \coloneqq \frac{1}{n} \sum_{i = 1}^{n} (A_{i} - \widehat{a}
(X_{i})) (Y_{i} - \widehat{b} (X_{i})).
\end{align*}


As usual, before presenting our new results, we first summarize the
statistical properties and minimaxity of $\widehat{\psi}_{1, n}$ obtained in
\citet{balakrishnan2026fundamental} under the SA model defined by
\citet{balakrishnan2026fundamental} for $\psi(a, b)$:
\begin{equation}
\label{ecc}\mathcal{P}_{\mathrm{SA}} ((\widehat{a}, \widehat{b}), (r_{n}, s_{n}))
\coloneqq \left\{  (a, b): \Vert a - \widehat{a} \Vert_{2, \mathbb{P}}^{2}
\lesssim r_{n}, \Vert b - \widehat{b} \Vert_{2, \mathbb{P}}^{2} \lesssim s_{n}
\right\}  .
\end{equation}
We let $p$ denote the marginal density of $X$, which, for simplicity, is
assumed to be $\mathrm{Unif} ([0, 1]^{d})$.

\begin{lemma}
\label{lem:ecc} The following hold.
\begin{align*}
\mathfrak{R}_{n} (\psi; \mathcal{P}_{\mathrm{SA}} ((\widehat{a}, \widehat{b}), (r_{n},
s_{n}))) \coloneqq \inf_{\widehat{\psi}} \sup_{(a, b) \in\mathcal{P}_{\mathrm{SA}}
((\widehat{a}, \widehat{b}), (r_{n}, s_{n}))} \mathsf{E}_{a, b} [(\widehat{\psi} - \psi(a,
b))^{2}] \gtrsim r_{n} \cdot s_{n} + \frac{1}{n}.
\end{align*}
This lower bound is attained by the first-order estimator $\widehat{\psi}_{1, n}$.
The bias, variance, and MSE of $\widehat{\psi}_{1, n}$ have the following forms:
\begin{align*}
\mathsf{bias} (\widehat{\psi}_{1, n})  &  = - \langle a - \widehat{a}, b - \widehat{b}
\rangle_{\mathbb{P}},\\
\mathsf{var} (\widehat{\psi}_{1, n})  &  = \frac{1}{n} \left\{  \mathsf{E} [(A -
\widehat{a} (X))^{2} (Y - \widehat{b} (X))^{2}] - \mathsf{E}^{2} [(A - \widehat{a} (X)) (Y
- \widehat{b} (X))] \right\}  , \text{ and}\\
\mathsf{mse} (\widehat{\psi}_{1, n})  &  = \langle a - \widehat{a}, b - \widehat{b}
\rangle_{\mathbb{P}}^{2} + \frac{1}{n} \left\{  \mathsf{E} [(A - \widehat{a}
(X))^{2} (Y - \widehat{b} (X))^{2}] - \mathsf{E}^{2} [(A - \widehat{a} (X)) (Y -
\widehat{b} (X))] \right\}  .
\end{align*}

\end{lemma}

To construct the second-order estimator, we similarly find a $k$-dimensional
dictionary $\bar{\phi}_{k}$ and denote $\Sigma\coloneqq \mathsf{E} [\bar{\phi
}_{k} (X)^{\otimes2}]$. In practice, one needs to estimate $\Sigma$ from data.
We make the following assumptions on $\Sigma$ and its estimator.

\begin{assumption}
\label{as:Sigma} $\Sigma$ is assumed to have bounded spectra and there exists
an estimator $\widehat{\Sigma}$ of $\Sigma$ such that $\widehat{\Sigma}$ also has
bounded spectra and $\Vert\widehat{\Sigma} - \Sigma\Vert_{\mathrm{op}} = o (1)$,
where $\Vert\cdot\Vert_{\mathrm{op}}$ denotes the matrix operator norm.
Without loss of generality, we take $\Sigma= \Sigma^{-1} = \mathrm{I}$.
\end{assumption}

We then construct the following second-order estimator for $\psi(a, b)$.
\begin{align*}
&  \widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma}) \coloneqq \widehat{\psi}_{1,
n} + \widehat{U}_{n, 2} (\bar{\phi}_{k}; \widehat{\Sigma}), \text{ where }\\
&  \widehat{U}_{n, 2} (\bar{\phi}_{k}; \widehat{\Sigma}) \coloneqq \mathbb{U}_{n, 2}
\left[  (A_{1} - \widehat{a} (X_{1})) \bar{\phi}_{k} (X_{1})^{\top} \widehat{\Sigma
}^{-1} \bar{\phi}_{k} (X_{2}) (Y_{2} - \widehat{b} (X_{2})) \right]  .
\end{align*}


We further introduce some short-hand notation for ease of exposition:
\begin{align*}
&  \widehat{\varepsilon}_{a} \coloneqq A - \widehat{a} (X), \widehat{\varepsilon}_{b}
\coloneqq Y - \widehat{b} (X),\\
&  \alpha\coloneqq \mathsf{E} [(a (X) - \widehat{a} (X)) \bar{\phi}_{k} (X)],
\beta\coloneqq [(b (X) - \widehat{b} (X)) \bar{\phi}_{k} (X)],\\
&  \Sigma_{a, a} \coloneqq \mathsf{E} [(A - \widehat{a} (X))^{2} \bar{\phi}_{k}
(X) \bar{\phi}_{k} (X)^{\top}], \Sigma_{b, b} \coloneqq \mathsf{E} [(Y -
\widehat{b} (X))^{2} \bar{\phi}_{k} (X) \bar{\phi}_{k} (X)^{\top}],\\
&  \text{and } \Sigma_{a, b} \coloneqq \mathsf{E} [(A - \widehat{a} (X)) (Y -
\widehat{b} (X)) \bar{\phi}_{k} (X) \bar{\phi}_{k} (X)^{\top}].
\end{align*}


We are now ready to state the following lemma. The proof is by direct
calculations and can be found in Appendix~\ref{app:lem:ecc2}.

\begin{lemma}
\label{lem:ecc2} The bias and variance of $\widehat{\psi}_{2, n} (\bar{\phi}_{k};
\widehat{\Sigma})$ read as:
\begin{align*}
\mathsf{bias} (\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma}))  &  = -
\langle\mathsf{\Pi}^{\perp} [\widehat{a} - a \mid \bar{\phi}_{k}], \mathsf{\Pi
}^{\perp} [\widehat{b} - b \mid \bar{\phi}_{k}] \rangle_{\mathbb{P}} + \alpha^{\top}
(\widehat{\Sigma}^{-1} - \mathrm{I}) \beta\lesssim r_{n}^{1 / 2} \cdot s_{n}^{1 /
2},\\
\mathsf{var} (\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma}))  &  =
\mathsf{var} (\widehat{\psi}_{1, n}) + \mathsf{var} (\widehat{U}_{n, 2} (\bar{\phi
}_{k}; \widehat{\Sigma})) + 2 \mathsf{cov} (\widehat{\psi}_{1, n}, \widehat{U}_{n, 2}
(\bar{\phi}_{k}; \widehat{\Sigma})) \lesssim\frac{1}{n},
\end{align*}
where
\begin{align*}
\mathsf{var} (\widehat{\psi}_{1, n}) =  &  \ \frac{1}{n} \left\{  \mathsf{E}
[\widehat{\varepsilon}_{a}^{2} \widehat{\varepsilon}_{b}^{2}] - \mathsf{E}^{2}
[\widehat{\varepsilon}_{a} \widehat{\varepsilon}_{b}] \right\}  ,\\
\mathsf{var} (\widehat{U}_{n, 2} (\bar{\phi}_{k}; \widehat{\Sigma})) =  &  \ \frac
{1}{n (n - 1)} \mathsf{Tr} \left\{  \Sigma_{a, a} \widehat{\Sigma}^{-1} \Sigma_{b,
b} \widehat{\Sigma}^{-1} + (\Sigma_{a, b} \widehat{\Sigma}^{-1})^{2} \right\} \\
&  + \frac{n - 2}{n (n - 1)} \left(  \alpha^{\top} \widehat{\Sigma}^{-1}
\Sigma_{b, b} \widehat{\Sigma}^{-1} \alpha+ \beta^{\top} \widehat{\Sigma}^{-1}
\Sigma_{a, a} \widehat{\Sigma}^{-1} \beta+ 2 \alpha^{\top} \widehat{\Sigma}^{-1}
\Sigma_{a, b} \widehat{\Sigma}^{-1} \beta\right) \\
&  - \frac{2 (2 n - 3)}{n (n - 1)} (\alpha^{\top} \widehat{\Sigma}^{-1} \beta
)^{2},\\
\leq &  \ \frac{C k}{n^{2}} + \frac{C}{n} \left\{  \Vert\mathsf{\Pi}[\widehat{a} -
a \mid \bar{\phi}_{k}] \Vert_{2, \mathbb{P}}^{2} + \Vert\mathsf{\Pi}[\widehat{b} - b
| \bar{\phi}_{k}] \Vert_{2, \mathbb{P}}^{2} \right\}  \lesssim\frac{1}{n}\\
2 \mathsf{cov} (\widehat{\psi}_{1, n}, \widehat{U}_{n, 2} (\bar{\phi}_{k}; \widehat
{\Sigma})) =  &  \ \frac{2}{n} \left\{  \mathsf{E} [\widehat{\varepsilon}_{a}
\widehat{\varepsilon}_{b}^{2} \bar{\phi}_{k} (X)^{\top}] \widehat{\Sigma}^{-1} \alpha+
\mathsf{E} [\widehat{\varepsilon}_{a}^{2} \widehat{\varepsilon}_{b} \bar{\phi}_{k}
(X)^{\top}] \widehat{\Sigma}^{-1} \beta- 2 \mathsf{E} [\widehat{\varepsilon}_{a}
\widehat{\varepsilon}_{b}] \alpha^{\top} \widehat{\Sigma}^{-1} \beta\right\} \\
\leq &  \ \frac{C}{n} \left\{  \Vert\mathsf{\Pi}[\widehat{a} - a \mid \bar{\phi}_{k}]
\Vert_{2, \mathbb{P}} + \Vert\mathsf{\Pi}[\widehat{b} - b \mid \bar{\phi}_{k}]
\Vert_{2, \mathbb{P}} \right\}  \lesssim\frac{1}{n}.
\end{align*}
The inequalities hold for $(a, b) \in\mathcal{P}_{\mathrm{SA}} ((\widehat{a},
\widehat{b}), (r_{n}, s_{n}))$ under Assumption~\ref{as:Sigma}.
\end{lemma}

It is not difficult to also see that $\widehat{\psi}_{2, n} (\bar{\phi}_{k};
\widehat{\Sigma})$ corrects the bias of $\widehat{\psi}_{1, n}$ by estimating a lower
bound of $\mathsf{bias} (\widehat{\psi}_{1, n}) \lesssim r_{n}^{1 / 2} \cdot
s_{n}^{1 / 2}$.

\begin{remark}
\label{rem:Sigma_hat} In \citet{liu2020nearly} and
\citet{liu2017semiparametric}, we have shown that when $\widehat{\Sigma}$ is the
sample Gram matrix estimator $\Vert\widehat{\Sigma} - \mathrm{I} \Vert
_{\mathrm{op}} = \sqrt{k \log k / n} = o (1)$ when $k = o (n / \log^{2} n)$
when the sample used to compute $\widehat{\Sigma}$ is also of size $n$
\citep{tropp2015introduction}. When we know $\Sigma= \mathrm{I}$,
$\mathsf{bias} (\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \mathrm{I}))$ is reduced to
$- \langle\mathsf{\Pi}^{\perp} [\widehat{a} - a \mid \bar{\phi}_{k}], \mathsf{\Pi
}^{\perp} [\widehat{b} - b \mid \bar{\phi}_{k}] \rangle_{\mathbb{P}} \lesssim
r_{n}^{1 / 2} \cdot s_{n}^{1 / 2}$ because there is no extra bias due to
estimating $\Sigma$. However, even if we estimate $\Sigma$ by $\widehat{\Sigma}$,
the extra bias incurred is of the form
\begin{align*}
\alpha^{\top} (\widehat{\Sigma}^{-1} - \mathrm{I}) \beta\lesssim\Vert\mathsf{\Pi
}[\widehat{a} - a \mid \bar{\phi}_{k}] \Vert_{2, \mathbb{P}} \Vert\mathsf{\Pi}
[\widehat{b} - a \mid \bar{\phi}_{k}] \Vert_{2, \mathbb{P}} \Vert\widehat{\Sigma}^{-1} -
\mathrm{I} \Vert_{\mathrm{op}} = o (r_{n}^{1 / 2} \cdot s_{n}^{1 / 2}).
\end{align*}
Thus as long as we have a consistent estimator of $\Sigma$, the second-order
estimator is still asymptotically minimax under the SA model.
\end{remark}

\begin{theorem}
\label{thm:ecc} Under Model $\mathcal{P}_{\mathrm{SA}} ((\widehat{a}, \widehat{b}),
(r_{n}, s_{n}))$, both $\widehat{\psi}_{1, n}$ and $\widehat{\psi}_{2, n} (\bar{\phi
}_{k}; \widehat{\Sigma})$ are asymptotically minimax, as long as $\Sigma$ and
$\widehat{\Sigma}$ satisfy Assumption~\ref{as:Sigma}. The same conclusions hold
when we replace the SA model $\mathcal{P}_{\mathrm{SA}} ((\widehat{a}, \widehat{b}),
(r_{n}, s_{n}))$ with the assumption-lean model $\mathcal{P}_{\mathrm{AL}}
((\widehat{a}, \widehat{b}))$.
\end{theorem}

\begin{proof}
The minimaxity of $\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma})$ can be concluded using the orders of its bias and variance shown in Lemma~\ref{lem:ecc2}.
\end{proof}


\begin{remark}
\label{rem:ate} Similar statements to those in Theorem~\ref{thm:ecc} hold for
the average treatment effect and the average treatment effect on the treated.
For instance, for the treatment specific mean, this can be seen by replacing
the notation $a, b, \widehat{\varepsilon}_{a}, \widehat{\varepsilon}_{b}$ by the
following instead:
\begin{align*}
&  a (\cdot) = 1 / \mathsf{E} [A \mid X = \cdot], b (\cdot) = \mathsf{E} [Y \mid X =
\cdot, A = 1], \widehat{\varepsilon}_{a} = A \widehat{a} (X) - 1, \widehat{\varepsilon
}_{b} = A (Y - \widehat{b} (X)).
\end{align*}
The dictionary $\bar{\phi}_{k}$ will also be weighted by the treatment
indicator $A \bar{\phi}_{k}$. The minimaxity of the first-order DML estimators
of these two functionals has been shown in \citet{jin2025structure}.
\end{remark}

\begin{remark}
\label{rem:dr} As indicated after Definition~\ref{def:al}, we will discuss in
Section~\ref{sec:sa} that the higher-order generalization of the second-order
estimators can be used to falsify the null hypothesis $\mathcal{H}
_{0}:\mathsf{bias}(\widehat{\psi}_{1,n})\ll n^{-1/2}$ when $\psi$ belongs to the
\emph{monotone bias class}. If $\psi$ is the expected conditional covariance
or the treatment specific mean parameter mentioned in Remark~\ref{rem:ate},
$\psi$ belongs to the so-called mixed-bias class
\citep{rotnitzky2021characterization} but not the \emph{monotone bias class}.
Here, we cannot claim the asymptotic inadmissibility of the DML estimator
$\widehat{\psi}_{1,n}$ of $\psi$ and similarly we cannot directly falsify
$\mathcal{H}_{0}:\mathsf{bias}(\widehat{\psi}_{1,n})\ll n^{-1/2}$. Nonetheless, we
can empirically falsify the rate-double-robustness of $\widehat{\psi}_{1,n}$
\citep{liu2024assumption}, where rate-double-robustness refers to the
assumption $r_{n}^{1/2}\cdot s_{n}^{1/2}=o(n^{-1/2})$ for both the expected
conditional covariance or the treatment specific mean parameter.
\end{remark}

\subsection*{Specializing to the Expected Conditional Variance}

A special case of the ECC functional--the expected conditional variance (abbreviated as the ECV functional) $\psi(a) \equiv\psi(a, a)$, however, belongs to the \emph{monotone bias class}, when $A = Y$ with probability 1. Here, the corresponding DML estimator is $\widehat{\psi}_{1, n} \coloneqq n^{-1} \sum_{i = 1}^{n} (A_{i} - \widehat{a} (X_{i}))^{2}$ and the corresponding SA model is defined as
\begin{align*}
\mathcal{P}_{\mathrm{SA}} (\widehat{a}, r_{n}) \equiv \mathcal{P}_{\mathrm{SA}} ((\widehat{a}, \widehat{a}), (r_{n}, r_{n})) \coloneqq \Big\{ a: \Vert \widehat{a} - a \Vert_{2, \mathbb{P}}^{2} \leq r_{n} \Big\}.
\end{align*}


\begin{remark}
\label{rem:cf2} Similar to Remark~\ref{rem:cf1}, we use the same estimator
$\widehat{a} \equiv\widehat{a}_{1} \equiv\widehat{a}_{2}$ to compute $\widehat{\psi}_{1, n}$,
again excluding the estimators studied in \citet{newey2018cross, mcgrath2026nuisance, mcclean2026double}.
\end{remark}

Analogously, the second-order estimator for $\psi(a)$ takes the following
form:
\begin{align*}
&  \widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma}) \coloneqq \widehat{\psi}_{1,
n} + \widehat{U}_{n, 2} (\bar{\phi}_{k}; \widehat{\Sigma}), \text{ where }\\
&  \widehat{U}_{n, 2} (\bar{\phi}_{k}; \widehat{\Sigma}) \coloneqq \mathbb{U}_{n, 2}
\left[  (A_{1} - \widehat{a} (X_{1})) \bar{\phi}_{k} (X_{1})^{\top} \widehat{\Sigma
}^{-1} \bar{\phi}_{k} (X_{2}) (A_{2} - \widehat{a} (X_{2})) \right].
\end{align*}
Lemma~\ref{lem:ecc} and Lemma~\ref{lem:ecc2} immediately imply the two
corollaries below for the ECV functional $\psi(a)$.

\begin{corollary}
\label{cor:ecc} The following hold.
\begin{align*}
\mathfrak{R}_{n} (\psi; \mathcal{P}_{\mathrm{SA}} (\widehat{a}, r_{n}))
\coloneqq \inf_{\widehat{\psi}} \sup_{a \in\mathcal{P}_{\mathrm{SA}} (\widehat{a},
r_{n})} \mathsf{E}_{a} [(\widehat{\psi} - \psi(a))^{2}] \gtrsim r_{n}^{2} +
\frac{1}{n}.
\end{align*}
This lower bound is attained by the first-order estimator $\widehat{\psi}_{1, n}$.
The bias, variance, and MSE of $\widehat{\psi}_{1, n}$ have the following forms:
\begin{align*}
\mathsf{bias} (\widehat{\psi}_{1, n})  &  = - \Vert a - \widehat{a} \Vert_{2,
\mathbb{P}}^{2} = - \, \Vert\mathsf{\Pi}^{\perp} [\widehat{a} - a \mid \bar{\phi}
_{k}] \Vert_{2, \mathbb{P}}^{2} - \alpha^{\top} \alpha,\\
\mathsf{var} (\widehat{\psi}_{1, n})  &  = \frac{1}{n} \left\{  \mathsf{E}
[\widehat{\varepsilon}_{a}^{4}] - \mathsf{E}^{2} [\widehat{\varepsilon}_{a}^{2}]
\right\}  , \text{ and}\\
\mathsf{mse} (\widehat{\psi}_{1, n})  &  = \Vert a - \widehat{a} \Vert_{2, \mathbb{P}
}^{4} + \frac{1}{n} \left\{  \mathsf{E} [\widehat{\varepsilon}_{a}^{4}] -
\mathsf{E}^{2} [\widehat{\varepsilon}_{a}^{2}] \right\}  .
\end{align*}

\end{corollary}

\begin{corollary}
\label{cor:ecv2} The bias and variance of $\widehat{\psi}_{2, n} (\bar{\phi}_{k};
\widehat{\Sigma})$ read as:
\begin{align*}
\mathsf{bias} (\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma}))  &  = -
\Vert\mathsf{\Pi}^{\perp} [\widehat{a} - a \mid \bar{\phi}_{k}] \Vert_{2, \mathbb{P}
}^{2} + \alpha^{\top} (\widehat{\Sigma}^{-1} - \mathrm{I}) \alpha\lesssim r_{n},\\
\mathsf{var} (\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma}))  &  =
\mathsf{var} (\widehat{\psi}_{1, n}) + \mathsf{var} (\widehat{U}_{n, 2} (\bar{\phi
}_{k}; \widehat{\Sigma})) + 2 \mathsf{cov} (\widehat{\psi}_{1, n}, \widehat{U}_{n, 2}
(\bar{\phi}_{k}; \widehat{\Sigma})) \lesssim\frac{1}{n},
\end{align*}
where
\begin{align*}
\mathsf{var} (\widehat{\psi}_{1, n}) =  &  \ \frac{1}{n} \mathsf{var}
(\widehat{\varepsilon}_{a}^{2}),\\
\mathsf{var} (\widehat{U}_{n, 2} (\bar{\phi}_{k}; \widehat{\Sigma})) =  &  \ \frac
{2}{n (n - 1)} \mathsf{Tr} \left\{  (\Sigma_{a, a} \widehat{\Sigma}^{-1})^{2}
\right\}  + \frac{4 n - 8}{n (n - 1)} \left(  \alpha^{\top} \widehat{\Sigma}^{-1}
\Sigma_{a, a} \widehat{\Sigma}^{-1} \alpha\right) \\
&  - \frac{4 n - 6}{n (n - 1)} (\alpha^{\top} \widehat{\Sigma}^{-1} \alpha)^{2}\\
\leq &  \ \frac{C k}{n^{2}} + \frac{C}{n} \Vert\mathsf{\Pi}[\widehat{a} - a \mid \bar{\phi}_{k}] \Vert_{2, \mathbb{P}}^{2} \lesssim\frac{1}{n},\\
2 \mathsf{cov} (\widehat{\psi}_{1, n}, \widehat{U}_{n, 2} (\bar{\phi}_{k}; \widehat
{\Sigma})) =  &  \ \frac{4}{n} \left\{  \mathsf{E} [\widehat{\varepsilon}_{a}^{3}
\bar{\phi}_{k} (X)^{\top}] \widehat{\Sigma}^{-1} \alpha- \mathsf{E} [\widehat
{\varepsilon}_{a}^{2}] \alpha^{\top} \widehat{\Sigma}^{-1} \alpha\right\} \\
\leq &  \ \frac{C}{n} \Vert\mathsf{\Pi}[\widehat{a} - a \mid \bar{\phi}_{k}]
\Vert_{2, \mathbb{P}} \lesssim\frac{1}{n}.
\end{align*}
The inequalities hold for $a \in\mathcal{P}_{\mathrm{SA}} (\widehat{a}, r_{n})$
under Assumption~\ref{as:Sigma}.
\end{corollary}

By Corollary~\ref{cor:ecv2}, $\psi(a)$ belongs to the \emph{monotone bias
class}. By piecing together the above two corollaries, we obtain the final
theoretical result of this paper. There always exists a distribution in
$\mathcal{P}_{\mathrm{SA}} (\widehat{a}, r_{n})$ such that
Definition~\ref{def:mbc}(1) holds. To see this, consider the case where $a =
\widehat{a} + r_{n}^{1 / 2} \beta^{\top} \bar{\phi}_{k}$ with $\Vert\beta\Vert_{2}
= 1$, for which $\mathsf{bias} (\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat
{\Sigma})) = \alpha^{\top} (\widehat{\Sigma}^{-1} - \mathrm{I}) \alpha
\ll\mathsf{bias} (\widehat{\psi}_{1, n}) = \alpha^{\top} \alpha$. The rest of the
proof can be found in Appendix~\ref{app:ecv}.

\begin{theorem}
\label{thm:ecv} Under Model $\mathcal{P}_{\mathrm{SA}} (\widehat{a}, r_{n})$ and
Assumption~\ref{as:Sigma}, $\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma})$
is \emph{asymptotically minimax} and the following hold as long as $k$ is
chosen such that $k = o (n \mathsf{var} (\widehat{\varepsilon}_{a}^{2}))$:
\begin{align*}
&  \sup_{a \in\mathcal{P}_{\mathrm{SA}} (\widehat{a}, r_{n})} \limsup_{n
\rightarrow\infty} \frac{\mathsf{mse} (\widehat{\psi}_{2, n} (\bar{\phi}_{k};
\widehat{\Sigma})) - \mathsf{mse} (\widehat{\psi}_{1, n})}{\mathsf{mse} (\widehat{\psi
}_{1, n})} \leq0, \text{ and when $r_{n}^{2} \gtrsim n^{-1}$}\\
&  \inf_{a \in\mathcal{P}_{\mathrm{SA}} (\widehat{a}, r_{n})} \limsup_{n
\rightarrow\infty} \frac{\mathsf{mse} (\widehat{\psi}_{2, n} (\bar{\phi}_{k};
\widehat{\Sigma})) - \mathsf{mse} (\widehat{\psi}_{1, n})}{\mathsf{mse} (\widehat{\psi
}_{1, n})} < 0.
\end{align*}
Thus, by Definition~\ref{def:admissible}, the first-order DML estimator
$\widehat{\psi}_{1, n}$ is \emph{asymptotically inadmissible} when $r_{n}^{2}
\gtrsim n^{-1}$. The same conclusions hold when we replace the SA model
$\mathcal{P}_{\mathrm{SA}} (\widehat{a}, r_{n})$ with the assumption-lean model
$\mathcal{P}_{\mathrm{AL}} (\widehat{a})$ and drop the assumption on $r_{n}$.
\end{theorem}

\begin{remark}
\label{rem:Sigma_hat var} We note that $\mathsf{bias} (\widehat{\psi}_{2, n}
(\bar{\phi}_{k}; \widehat{\Sigma}))^{2} - \mathsf{bias} (\widehat{\psi}_{1, n})^{2}
\asymp- \Vert\mathsf{\Pi}[\widehat{a} - a \mid \bar{\phi}_{k}] \Vert_{2, \mathbb{P}
}^{2} (1 - \Vert\widehat{\Sigma}^{-1} - \mathrm{I} \Vert_{\mathrm{op}}) \asymp-
\Vert\mathsf{\Pi} [\widehat{a} - a \mid \bar{\phi}_{k}] \Vert_{2, \mathbb{P}}^{2}$ by
Assumption~\ref{as:Sigma}. Thus, asymptotically, the second-order estimator
still has smaller bias than the first-order DML estimator $\widehat{\psi}_{1, n}$,
and the impact of estimating $\Sigma$ is asymptotically negligible. In
addition, $\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma})$ corrects the bias
of $\widehat{\psi}_{1, n}$ by estimating a lower bound of $\mathsf{bias}
(\widehat{\psi}_{1, n}) \lesssim r_{n}$.
\end{remark}

\section{Concluding Remarks}
\label{sec:sa}

The SA model introduced in \citet{balakrishnan2026fundamental} is a mathematically appealing construct that has inspired follow-up work \citep{jin2025structure, bonvini2024doubly, jin2025normal, jin2025sharp, gu2026optimally, gu2025open}, including our current paper. The assumption-lean model $\mathcal{P}_{\mathrm{AL}}(\widehat{\theta})$ defined in Definition~\ref{def:al}, is aligned with the goal of understanding what can be learned from a model that makes almost no assumptions. As discussed earlier, in terms of point estimation, both the first-order DML estimator $\widehat{\psi}_{1,n}$ and our second-order estimator $\widehat{\psi}_{2,n}$ remain minimax with rate $O (1)$ in the assumption-lean model $\mathcal{P}_{\mathrm{AL}}(\widehat{\theta})$ for all the parameters studied in \citet{balakrishnan2026fundamental}; for $\psi$ in \emph{monotone bias class}, our $\widehat{\psi}_{2,n}$ continues to asymptotically dominate $\widehat{\psi}_{1,n}$ in the scaled MSE.

However, statisticians care about uncertainty quantification or inference as much as or even more than point estimation. Neither the (asymptotic) minimaxity/inadmissibility of $\widehat{\psi}_{1,n}$ nor the minimaxity of $\widehat{\psi}_{2,n}$ in the assumption-lean model $\mathcal{P}_{\mathrm{AL}}(\widehat{\theta})$ offer any guidance on how to quantify uncertainty, absent further knowledge of $\widehat{\theta}$ or $\Theta$. The above argument is not new. Before \citet{balakrishnan2026fundamental}, we considered inference on $\psi$ under the assumption-lean model $\mathcal{P}_{\mathrm{AL}} (\widehat{\theta})$ in \citet{liu2020nearly} and \citet{liu2024assumption}. The former paper was discussed by the authors of
\citet{balakrishnan2026fundamental}; see \citet{kennedy2020discussion} and \citet{liu2020rejoinder}.
Since no uniformly consistent estimators of $\psi$ exist in model
$\mathcal{P}_{\mathrm{AL}}(\widehat{\theta})$
\citep{ritov1990achieving, robins1997toward, ritov2014bayesian}, we, instead,
developed valid falsification tests of the following null hypothesis for
$\psi$ in the \emph{monotone bias class}:

\begin{quote}
$\mathcal{H}_{0}$: \emph{The bias of the first-order DML estimator $\widehat{\psi
}_{1, n}$ of $\psi$ is sufficiently small such that a standard Wald CI
centered at $\widehat{\psi}_{1,n}$ has nominal coverage asymptotically.}
\end{quote}

The proposed tests are only falsification tests because, although valid under $\mathcal{H}_{0}$, they will have no power under many alternatives to $\mathcal{H}_{0}$. However, when a test rejects the null, it provides empirical evidence that the bias of $\widehat{\psi}_{1,n}$ is too large for the Wald CI to deliver valid inference. The test statistics used are based on the same second-order estimators that we have analyzed in this paper or their higher-order extensions \citep{robins2008higher, robins2016technical, liu2017semiparametric}. For $\psi$ belonging to the so-called mixed-bias classes (which includes the expected conditional covariance analyzed above) \citep{rotnitzky2021characterization}, in \citet{liu2020nearly} and \citet{liu2024assumption}, we showed that these tests are no longer valid under $\mathcal{H}_{0}$. However, these tests remain valid falsification tests of the so-called rate-double-robustness property, as defined in Remark~\ref{rem:high} or Remark~\ref{rem:dr}. We note that the rate-double-robustness implies that $\mathcal{H}_{0}$ is true. For this reason, complexity-reducing assumptions strong enough to imply rate-double-robustness are often made by investigators to justify the validity of their Wald CIs centering $\widehat{\psi}_{1,n}$. In our view, unlike the minimaxity of $\widehat{\psi}_{1,n}$ or of $\widehat{\psi}_{2,n}$, these falsification tests provide further empirical information even in the assumption-lean model $\mathcal{P}_{\mathrm{AL}}(\widehat{\theta})$, whenever they reject and thus can be of value to domain scientists for whom inference is important.






\bibliographystyle{plainnat}
\bibliography{Master.bib}


\newpage