EconBase
← Back to paper

On the Asymptotic Inadmissibility of Double Machine Learning Estimators Under Structure-Agnostic Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

70,051 characters · 8 sections · 71 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

On the Asymptotic Inadmissibility of Double Machine Learning Estimators Under Structure-Agnostic Models

\affil[1]{Institute of Natural Sciences, MOE-LSC, School of Mathematical Sciences, SJTU-Yale Joint Center for Biostatistics and Data Science, Shanghai Jiao Tong University}

\affil[2]{Department of Biostatistics, Harvard T. H. Chan School of Public Health}

\affil[3]{Department of Epidemiology and Department of Biostatistics, Harvard T. H. Chan School of Public Health}

abstractStructure-agnostic (SA) models introduced by balakrishnan2026fundamental aim to reflect the general lack of knowledge of structural assumptions on data-generating laws such as smoothness or sparsity in practice. Roughly speaking, SA models restrict the observed-data generating law to be in some $r_{n}$-neighborhood of (black-box machine learning) estimates, treated as given and fixed, where $r_{n}$ encodes the convergence rates of the estimates to the truth. Under SA models, balakrishnan2026fundamental show that the popular Double Machine Learning (DML) estimators for three functionals, the quadratic functional in the Gaussian sequence model, the quadratic density integral functional and the expected conditional covariance, are minimax. However, minimax estimators may be inadmissible. In this paper, we show that, for the first two of the three functionals, the DML estimator is asymptotically inadmissible under the SA model. In particular, we show that these two functionals fall into a class of functionals, which we refer to as the monotone bias class. For this class, we exhibit second-order ($U$-statistic) estimators, which asymptotically dominate DML estimators, under the SA model. These second-order estimators are empirical higher-order influence function (HOIF) estimators introduced in liu2017semiparametric. Furthermore, the empirical HOIF estimator, like the DML estimator, is minimax for the third functional (the expected conditional covariance), although neither asymptotically dominates the other. Finally, we compare the SA model with the assumption-lean model of liu2020nearly, liu2024assumption that imposes no assumptions beyond the trivial and empirically untestable hypothesis that the bias of any estimator, including DML estimators and empirical HOIF estimators, may be of order $1$. As a consequence, under our assumption-lean model, a Wald confidence interval centered at a DML estimator may under-cover. liu2024assumption introduced a class of valid tests that can falsify, for functionals in the monotone bias class, the hypothesis that a DML-estimator-centered confidence interval covers the truth at its nominal level or greater. However, our tests are not consistent under the assumption-lean model, because no consistent tests exist robins1997toward. Furthermore, for any functional with the mixed bias property of rotnitzky2021characterization, such as the expected conditional covariance or the average treatment effect jin2025structure, the above falsification tests can falsify the hypothesis of rate-double-robustness.

Keywords: Foundations of Statistics, Higher-Order Influence Functions, Structure-Agnostic Models, Assumption-Lean Inference, Minimaxity, (In)admissibility

\affil[1]{Institute of Natural Sciences, MOE-LSC, School of Mathematical Sciences, SJTU-Yale Joint Center for Biostatistics and Data Science, Shanghai Jiao Tong University}

\affil[2]{Department of Biostatistics, Harvard T. H. Chan School of Public Health}

\affil[3]{Department of Epidemiology and Department of Biostatistics, Harvard T. H. Chan School of Public Health}

\onehalfspacing \allowdisplaybreaks \nopagebreak

Introduction

In scientific disciplines such as epidemiology, clinical medicine and economics, one of the most important statistical tasks is to infer from the observed data a low-dimensional, smooth functional $\psi(\theta)$ of the underlying data-generating law $\mathbb{P}_{\theta}$ posited to belong to a statistical model denoted by

align*[align* omitted — 120 chars of source]

Here we parameterize the data generating law by $\theta\in\Theta$. Without an essential loss of generality, we take $\psi\equiv\psi(\theta)\in\mathbb{R}$. Throughout this paper, we let $n$ denote the sample size and use $\psi$ and $\theta$ to denote the true values, which should cause no confusion.

Many examples of $\psi$, such as the average treatment effect under ignorability, are of substantive interest in practice. To avoid model misspecification bias, it is natural to take $\Theta$ to be high- or even infinite-dimensional and estimate $\theta$ nonparametrically by kernels or series in classical statistics. In terms of $\psi$, it has become common practice to construct the so-called Double Machine Learning (DML) estimator $\widehat{\psi}_{1,n}\equiv\widehat{\psi}_{1,n}(\widehat{\theta})$ based on the first-order influence function of $\psi$ newey1990semiparametric, scharfstein1999adjusting, ai2003efficient, chernozhukov2018double, shi2026and. Owing to the curse-of-dimensionality, uniformly consistent estimators exist neither for $\theta$ nor for $\psi$ without any additional assumptions on $\Theta$ stone1980optimal, stone1982optimal, ritov1990achieving, robins1997toward. For this reason, structural assumptions, traditionally in terms of smoothness or sparsity, are often imposed on $\Theta$ to obtain uniformly consistent estimators of $\psi$ that converge to $\psi$ at parametric rates.

Recently, balakrishnan2026fundamental introduced the structure-agnostic (SA) models, a new class of submodels of $\mathcal{P}$ that do not impose traditional structural assumptions on $\Theta$ such as smoothness or sparsity. The SA model, denoted as $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})$, is parameterized by a pair of indices $(\widehat{\theta},r_{n})$, where $\widehat{\theta}$ is an initial estimator of $\theta$ treated as fixed and independent of the randomness of the data, and $r_{n}$ indicates convergence rates that are nonincreasing functions of $n$. Concretely, suppose that $\theta=(\theta _{1},\cdots,\theta_{J})^{\top}$ and $r_{n}=(r_{n,1},\cdots,r_{n,J})^{\top}$ have $J$ components. Then $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})$ is generally defined as\footnote{In certain problems, we have $\theta_{j_{1}} \equiv\theta_{j_{2}}$ and $r_{n, j_{1}} \equiv r_{n, j_{2}}$. As is standard in the literature, in such a case, we use the same estimator $\widehat{\theta }_{j_{1}} \equiv\widehat{\theta}_{j_{2}}$ in computing $\widehat{\psi}_{1, n}$. We will discuss the implications of this choice in the examples in Sections (ref)-- (ref); specifically, see Remarks (ref) and (ref).}

equation[equation omitted — 217 chars of source]

In other words, $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta}, r_{n})$ contains the subset of all possible $\mathbb{P}_{\theta}$'s such that each component $\theta_{j}$ is contained in the corresponding $\sqrt{r_{n, j}}$-$\Vert\cdot\Vert$-neighborhood of a given initial estimator $\widehat{\theta}_{j}$. For all the concrete examples in this paper (see Sections (ref)--(ref)), we effectively take $r_{n, j}$ to be some large enough constant $R^{\ast}$ for $j > 2$ so assumptions are only imposed over at most two components of $\theta$.

The SA model $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})$ has several notable features. First, it does not impose explicit complexity reducing structural assumptions such as smoothness or sparsity\footnote{Even if the true smoothness or sparsity class were known, due to the theory-practice gap adcock2021gap, xu2022deepmed, chen2024causal, when $\theta$ is estimated by modern deep neural networks, the SA model is still relevant because the smoothness/sparsity assumption alone fails to determine the properties of $\widehat{\theta}$ or $\widehat{\psi}_{1,n}$.}. Second, balakrishnan2026fundamental showed that the first-order DML estimator $\widehat{\psi}_{1,n}$ of $\psi$ attains optimal convergence rates in the minimax sense under the SA model $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta}, r_{n})$, for several concrete examples of $\psi$, including the quadratic functional in the Gaussian sequence model, the quadratic density integral functional (with the extra condition $r_{n}^{2} \gtrsim n^{-1}$ for these two examples; see Theorem (ref) for explanation), and the expected conditional covariance. More recent follow-up papers jin2025structure, bonvini2024doubly, jin2025sharp, gu2026optimally, gu2025open establish the minimaxity of $\widehat{\psi}_{1,n}$ under the SA model for the average treatment effect, the average treatment effect on the treated, and other related parameters. Furthermore, the estimator $\widehat{\psi}_{1, n}$ is the same and minimax in $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})$ for all values of $r_{n}$. The minimaxity of $\widehat{\psi}_{1, n}$ under the SA model $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta}, r_{n})$ has been used by some analysts to justify the use of current practice in (bio)statistics and econometrics.

Our Contribution

Our main technical contribution of this paper is related to the decision-theoretic properties of estimators, tracing back to the classical work of Abraham Wald wald1941principles, wald1945statistical, wald1947essentially. Wald is renowned as the inventor of the minimax principle wald1945statistical, arguably the most popular theoretical paradigm used to measure the quality of an estimator by modern (bio)statisticians and econometricians brown1994minimaxity, andrews2021model, adusumilli2026sample and adopted in balakrishnan2026fundamental.

However, certain minimax estimators may be inadmissible. One celebrated example of the difference between minimaxity and (in)admissibility is Stein's paradox, asserting that the maximum likelihood estimator (MLE), although minimax, is everywhere dominated in mean squared error (MSE) for every sample size $n$ by the James-Stein (JS) estimator in the many-normal-means model of dimension at least three stein1956inadmissibility, james1961estimation, brown1971admissible. Hence, the MLE is inadmissible. In this paper, we will show that, under $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})$, there exists a class of parameters $\psi$, which we refer to as the monotone bias class, for which higher-order influence function (HOIF) estimators robins2008higher, robins2016technical, liu2017semiparametric dominate the first-order DML estimator $\widehat{\psi}_{1,n}$ in the large-$n$ limit whenever $\prod_{j = 1}^{J} r_{n, j} \gtrsim n^{-1}$ (see Theorem (ref) for the actual statement). That is, we show that in the $\mathcal{P}_{\mathrm{SA} }(\widehat{\theta},r_{n})$ model, the mimimax DML estimator is asymptotically inadmissible, when estimating a parameter in the monotone bias class. To make the above claims precise, we next define: asymptotic (in)admissibility and the monotone bias class of functionals. We shall see that two out of the three functionals studied in balakrishnan2026fundamental are in the monotone bias class.

definition[Asymptotic (in)admissibility] An estimator sequence $\psi_{n}$ of $\psi(\theta)$ indexed by $n$ is said to be asymptotically inadmissible in scaled MSE loss under model $\mathcal{P}_{\mathrm{SA}} (\widehat{\theta}, r_{n})$ if there exists another estimator sequence $\psi _{n}^{\prime}$ of $\psi(\theta)$ also indexed by $n$ such that \[ \sup_{\theta\in\mathcal{P}_{\mathrm{SA}} (\widehat{\theta}, r_{n})} \limsup_{n \rightarrow\infty} \frac{\mathsf{mse} (\psi_{n}^{\prime})-\mathsf{mse} (\psi_{n})}{\mathsf{mse} (\psi_{n})} \leq0 \text{ and } \inf_{\theta \in\mathcal{P}_{\mathrm{SA}}(\widehat{\theta}, r_{n})} \limsup_{n \rightarrow \infty} \frac{\mathsf{mse}(\psi_{n}^{\prime}) - \mathsf{mse} (\psi_{n} )}{\mathsf{mse} (\psi_{n})} < 0, \] where $\mathsf{mse} (\cdot) \equiv\mathsf{mse}_{\theta} (\cdot) \coloneqq \mathsf{E}_{\theta} (\cdot-\psi(\theta))^{2}$. If otherwise, we say that $\psi_{n}$ is asymptotically admissible.

In the above definition, the difference in MSEs is scaled. Because the MSE of a reasonable estimator should decay to zero under the large $n$ limit, obtaining nontrivial results requires an appropriate scaling. Here, we choose the MSE of $\psi_{n}$ as a natural scaling factor. We refer the interested readers to Appendix (ref) for a more elaborate discussion on the choice of the denominator.

The following definition of the monotone bias class is different from, but as explained below, is essentially equivalent to that in our previous work liu2020nearly.

definitionLet $\mathsf{bias} (\widehat{\psi}_{1, n})$ and $\mathsf{var} (\widehat{\psi}_{1, n})$ be, respectively, the bias and variance of the first-order DML estimator $\widehat{\psi}_{1, n}$ of $\psi$. $\psi$ is said to be in the monotone bias class if the following hold: \begin{enumerate} [label = (\arabic*)] • For any $r_{n}$, $\vert\mathsf{bias} (\widehat{\psi}_{1, n}) \vert \lesssim\prod_{j = 1}^{J} r_{n, j}$, there always exists another estimator, denoted by $\widehat{\psi}_{2, n}$, such that $\vert\mathsf{bias} (\widehat{\psi}_{2, n}) \vert/ \vert\mathsf{bias} (\widehat{\psi}_{1, n}) \vert- 1 \leq0$, and the inequality becomes strict (asymptotically) at some law in $\mathcal{P} _{\mathrm{SA}} (\widehat{\theta}, r_{n})$; • The variances of $\widehat{\psi}_{1,n}$ and $\widehat{\psi}_{2,n}$ satisfy the following condition: there exists a constant $v>0$ such that $\mathsf{var} (\widehat{\psi}_{1,n}) = v/n$ and \begin{equation} \begin{split} \limsup_{n\rightarrow\infty}\frac{\mathsf{var}(\widehat{\psi}_{2,n})} {\mathsf{var}(\widehat{\psi}_{1,n})}=\left\{ \begin{array} [c]{ll} 1 & if \dfrac{|\mathsf{bias}(\widehat{\psi}_{1,n})|-|\mathsf{bias} (\widehat{\psi}_{2,n})|}{v}=o(1),\\ \delta & if \dfrac{|\mathsf{bias}(\widehat{\psi}_{1,n})|-|\mathsf{bias} (\widehat{\psi}_{2,n})|}{v} \gtrsim1, \end{array} \right. \end{split} \end{equation} for some constant $\delta\neq1$ but possibly $\delta>1$. \end{enumerate} Here the constants $v$ and $\delta$ can depend on the data generating distribution $\mathbb{P}_{\theta}$.

We are now ready to state the following general theorem regarding the asymptotic inadmissibility of the DML estimator $\widehat{\psi}_{1, n}$. The proof is given following a few remarks. \setcounter{theorem}{-1}

theoremFor $\psi$ belonging to the monotone bias class, under the SA model $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})$, there exists another estimator $\widehat{\psi}_{2,n}$ such that (i) $\widehat{\psi}_{2,n}$ is asymptotically not greater in scaled MSE than $\widehat{\psi}_{1,n}$ over $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})$ for any $r_{n}$, and (ii) $\widehat{\psi}_{2,n}$ is asymptotically strictly smaller than $\widehat{\psi}_{1,n}$ in scaled MSE at some law in $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})$ whenever $\prod_{j=1}^{J}r_{n,j}\gtrsim n^{-1}$. Hence, $\widehat{\psi}_{1, n}$ is asymptotically inadmissible if and only if $\prod_{j=1}^{J}r_{n,j}\gtrsim n^{-1}$.

It should be noted that Theorem (ref) does not discuss minimaxity of either estimator. As we shall see, for two examples in the monotone bias class -- the quadratic functional in the Gaussian sequence model in Section (ref) and the quadratic density integral functional in Section (ref), the minimaxity of $\widehat{\psi}_{1, n}$ or $\widehat{\psi}_{2, n}$ requires an additional condition $\sqrt{\prod_{j = 1}^{J} r_{n, j}} \gtrsim n^{-1}$. As explained in balakrishnan2026fundamental, when $\sqrt{\prod_{j = 1}^{J} r_{n, j}} \ll n^{-1}$, the so-called plug-in estimators dominate both $\widehat{\psi}_{1, n}$ and $\widehat{\psi}_{2, n}$ in these two examples, because the plug-in estimators have zero variance and squared bias $\sqrt{\prod_{j = 1}^{J} r_{n, j}} \ll n^{-1}$, while both $\widehat{\psi}_{1, n}$ and $\widehat{\psi}_{2, n}$ have variances of order $n^{-1}$. However, as also noted by balakrishnan2026fundamental, $\sqrt{\prod_{j = 1}^{J} r_{n, j}} \ll n^{-1}$ generally does not hold if the sample size used to estimate $\widehat{\theta}$ is of the same order as $n$. Therefore, there is essentially no loss of generality if we exclude the case where $\sqrt{\prod_{j = 1}^{J} r_{n, j}} \ll n^{-1}$ holds, as we do in Sections (ref) and (ref), in which case both $\widehat{\psi}_{1, n}$ and $\widehat{\psi}_{2, n}$ are minimax rate-optimal.

remarkIn all the examples covered in this paper and in balakrishnan2026fundamental, $|\mathsf{bias}(\widehat{\psi}_{1,n})|^{2}$ is upper bounded by $\prod_{j=1}^{J} r_{n,j}$ times a constant. In view of Theorem (ref), when $\psi$ is in the monotone bias class, $\widehat{\psi}_{1,n}$ is asymptotically inadmissible and is asymptotically dominated by $\widehat{\psi}_{2,n}$ when $|\mathsf{bias}(\widehat{\psi }_{1,n})|\gtrsim n^{-1/2}$. When $|\mathsf{bias} (\widehat{\psi}_{1,n})| \ll n^{-1/2}$, $\widehat{\psi}_{2,n}$ is asymptotically never worse than $\widehat{\psi}_{1,n}$ and both are minimax. In particular, when DML estimators $\widehat{\psi }_{1, n}$ are deployed in practice, it is often implicitly assumed that $\prod_{j = 1}^{J} r_{n, j} \ll n^{-1}$ holds, a condition often referred to as the rate-double-robustness when $J = 2$. Under this condition, neither estimator dominates the other; further, $|\mathsf{bias} (\widehat{\psi }_{1,n})| \ll n^{-1 / 2}$ automatically holds, and thus both $\widehat{\psi}_{1, n}$ and $\widehat{\psi}_{2, n}$ can be used to construct an asymptotically valid Wald confidence interval (CI) of length $O (n^{-1 / 2})$, which is common practice, although other non-Wald CI constructions are also under rapid development zheng2025perturbed.

Since in practice $r_{n}$ is unknown and the possibility that it is of order $1$ cannot be empirically excluded robins1997toward, ritov2014bayesian, we consider the following assumption-lean model.

definitionGiven a constant $R^{\ast} > 0$, we define the model $\mathcal{P}_{\rm AL} (\widehat{\theta}) \equiv \mathcal{P}_{\rm SA} (\widehat{\theta}, r_{n} = (R^{\ast}, \cdots, R^{\ast}))$ as the assumption-lean model liu2024assumption.

Under the assumption-lean model, for parameters in the monotone bias class, it follows from Theorem (ref) that $\widehat{\psi}_{2,n}$ dominates $\widehat{\psi}_{1,n}$ asymptotically, because $\prod_{j=1}^{J} r_{n,j}=1\gg n^{-1}$, and thus $\widehat{\psi}_{1,n}$ is asymptotically inadmissible. In fact, as discussed in Section (ref) below, we can say more. Given a parameter $\psi$ in the monotone bias class, we can construct an asymptotically level-$\alpha$ falsification test of the null hypothesis $\mathcal{H}_{0}$: $\mathsf{bias}(\widehat{\psi}_{1,n})\ll n^{-1/2}$ that, when $\mathcal{H}_{0}$ is rejected, provides empirical evidence for the alternative hypothesis $|\mathsf{bias}(\widehat{\psi}_{1,n})|\gtrsim n^{-1/2}$ and thus also empirical evidence that the Wald CI centered on $\widehat{\psi}_{1,n}$ under-covers even in large samples liu2020nearly. However, by the aforementioned results of robins1997toward and ritov2014bayesian, any such test must be inconsistent under the assumption-lean model $\mathcal{P}_{\mathrm{AL}} (\widehat{\theta})$. Hence, failure to reject does not provide evidence for or against the null hypothesis $\mathcal{H}_{0}$ even asymptotically. We return to this issue in the concluding section of the paper (Section (ref)).

remarkIt will be clear in later sections that all examples studied by balakrishnan2026fundamental are in the monotone bias class defined in Definition (ref), except for the expected conditional covariance (see Section (ref)). The two parts of conditions in Definition (ref) need further elaboration. Part (1) states that the bias of $\widehat{\psi}_{2, n}$ is never greater than but sometimes strictly smaller than that of $\widehat{\psi}_{1, n}$. Part (2) says that the difference between the variances of $\widehat{\psi}_{1, n}$ and $\widehat{\psi}_{2, n}$ is negligible (of order $o (1 / n)$) if the bias reduction of $\widehat{\psi}_{2, n}$ is negligible (of order $o (1)$). The original definition of the monotone bias class in liu2020nearly contains only part (1) but part (2) was implicit. As described later, the alternative estimator $\widehat{\psi}_{2, n}$ is a second-order $U$-statistic (heretofore referred to as the second-order estimators for simplicity), constructed via the theory of HOIFs robins2008higher, robins2016technical, liu2017semiparametric. Such second-order estimators satisfy both Parts (1) and (2). The difference between the variances of $\widehat{\psi}_{2, n}$ and $\widehat{\psi}_{1, n}$ can be bounded as follows, as is the case in all our examples in later sections: \[ \mathsf{var}(\widehat{\psi}_{2,n})-\mathsf{var}(\widehat{\psi}_{1,n})\lesssim\frac {k}{n^{2}}+\frac{|\mathsf{bias}(\widehat{\psi}_{1,n})|-|\mathsf{bias}(\widehat{\psi }_{2,n})|}{n}+\frac{v^{1/2}\{|\mathsf{bias}(\widehat{\psi}_{1,n})|-|\mathsf{bias} (\widehat{\psi}_{2,n})|\}^{1/2}}{n}, \] where $k$ is a tuning parameter chosen so that $k=o(n)$. This will be demonstrated in the proofs of Theorem (ref) --Theorem (ref) in Appendix (ref). It is then not difficult to verify that Part (2) of Definition (ref) holds. In terms of Part (1), the second-order estimator $\widehat{\psi}_{2,n}$ can be viewed as debiasing $\widehat{\psi}_{1,n}$ by estimating a part of $\mathsf{bias}(\widehat{\psi}_{1,n})$; also see comments after Lemma (ref), (ref), and (ref), and Theorem (ref). For functionals outside the monotone bias class such as the expected conditional covariance functional covered in Section (ref), $\widehat{\psi}_{2,n}$ is still rate-optimal but may have a bias exceed that of $\widehat{\psi}_{1,n}$ under certain laws in $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})$ .

With these ingredients, the proof of Theorem (ref) is almost immediate by elementary calculations, so we record the proof here.

proof[Proof of Theorem (ref)] The result follows directly from Definition (ref). To see this, we first decompose the scaled MSE difference as \begin{align*} \frac{\mathsf{mse} (\widehat{\psi}_{2, n}) - \mathsf{mse} (\widehat{\psi}_{1, n})}{\mathsf{mse} (\widehat{\psi}_{1, n})} = - \frac{\mathsf{bias} (\widehat{\psi}_{1, n})^{2} - \mathsf{bias}^{2} (\widehat{\psi}_{2, n})}{\mathsf{bias} (\widehat{\psi}_{1, n})^{2} + \mathsf{var} (\widehat{\psi}_{1, n})} + \frac{\mathsf{var} (\widehat{\psi}_{2, n}) - \mathsf{var} (\widehat{\psi}_{1, n})}{\mathsf{bias} (\widehat{\psi}_{1, n})^{2} + \mathsf{var} (\widehat{\psi}_{1, n})} \coloneqq T_{1, n} + T_{2, n}. \end{align*} When $\dfrac{|\mathsf{bias} (\widehat{\psi}_{1, n})| - |\mathsf{bias} (\widehat{\psi}_{2, n})|}{n \cdot \mathsf{var} (\widehat{\psi}_{1, n})} = o (1)$ holds, by (ref), we always have: \begin{align*} \limsup_{n \rightarrow \infty} T_{2, n} = \limsup_{n \rightarrow \infty} \frac{\mathsf{var} (\widehat{\psi}_{2, n}) / \mathsf{var} (\widehat{\psi}_{1, n}) - 1}{\mathsf{bias} (\widehat{\psi}_{1, n})^{2} / \mathsf{var} (\widehat{\psi}_{1, n}) + 1} = 0. \end{align*} Since $T_{1, n}$ is always non-positive, we have $\limsup_{n \rightarrow \infty} T_{1, n} + T_{2, n} \leq 0$. In words, when the bias reduction is sufficiently small, by (2) of Definition (ref), the asymptotic variance of the second-order estimator is not different from that of the first-order DML estimator. On the contrary, we now suppose that $|\mathsf{bias} (\widehat{\psi}_{1, n})| - |\mathsf{bias} (\widehat{\psi}_{2, n})| \gtrsim n \cdot \mathsf{var} (\widehat{\psi}_{1, n})$, and hence also $|\mathsf{bias} (\widehat{\psi}_{1, n})| \gtrsim n \cdot \mathsf{var} (\widehat{\psi}_{1, n})$. We must also have \begin{align*} \mathsf{bias} (\widehat{\psi}_{1, n})^{2} - \mathsf{bias} (\widehat{\psi}_{2, n})^{2} \gg \mathsf{var} (\widehat{\psi}_{2, n}) - \mathsf{var} (\widehat{\psi}_{1, n}), \end{align*} and by the non-positivity of $T_{1, n}$, we have $\limsup_{n \rightarrow \infty} T_{1, n} + T_{2, n} \leq 0$. When $|\mathsf{bias} (\widehat{\psi}_{1, n})|^{2} \lesssim \prod_{j = 1}^{J} r_{n, j} \ll n^{-1}$, the denominator in the scaled MSE difference is dominated by $\mathsf{var} (\widehat{\psi}_{1, n}) = v / n$. The numerator can only be of smaller order than the denominator, rendering $\limsup_{n \rightarrow \infty} T_{1, n} + T_{2, n} = 0$. When $\prod_{j = 1}^{J} r_{n, j} \gtrsim n^{-1}$, there must exist a law such that $\mathsf{bias} (\widehat{\psi}_{1, n}) \gtrsim n^{-1 / 2}$. By Definition (ref), there must exist a distribution for which $\limsup_{n \rightarrow \infty} T_{1, n} < 0$, which completes the proof.

Notation and Organization

Throughout this paper, we always let $C > 0$ denote a sufficiently large constant independent of the sample size $n$. For any functions mentioned in the paper, they are understood to be squared integrable with respect to the Lebesgue measure. $\mathbb{U}_{n, m} [\cdot]$ denotes a $m$-th order $U$-statistic operator. Given a collection of $k$ different functions $\bar {f}_{k} = (f_{1}, \cdots, f_{k})^{\top}$, we let $\mathsf{\Pi}_{\mathbb{P}} (\cdot \mid \bar{f}_{k})$ denote the operator of $L^{2} (\mathbb{P})$-projection onto the linear span of $\bar{f}_{k}$ and $\Vert\cdot\Vert_{2, \mathbb{P}}$ denote the $L^{2} (\mathbb{P})$-norm. If $\mathbb{P}$ is the Lebesgue measure, we omit the subscript and write $\mathsf{\Pi} (\cdot \mid \bar{f}_{k})$ and $\Vert\cdot\Vert_{2}$ for short. $\Vert\cdot\Vert_{2}$ also denotes the $\ell^{2}$-norm of a vector. We denote the population Gram matrix of $\bar {f}_{k}$ under the distribution $\mathbb{P}$ as $\Sigma_{\mathbb{P}, \bar {f}_{k}} \coloneqq \mathsf{E} [\bar{f}_{k} (X) \bar{f}_{k} (X)^{\top}]$. When it is clear from the context, we omit the dependence in the subscript on $\mathbb{P}$ or $\bar{f}_{k}$ or both.

The remainder of this paper makes Theorem (ref) concrete. Sections (ref)--(ref) cover four examples of $\psi$, one of which does not belong to the monotone bias class. Theorem (ref) will then be specialized for the three examples in the monotone bias class. Section (ref) concludes the paper by making some additional comments on the relevance of the SA model to practitioners who are more interested in uncertainty quantification or statistical inference. Proofs are deferred to the Appendix.

Quadratic Functional in the Gaussian Sequence Model

As in balakrishnan2026fundamental, we observe data drawn from the infinite Gaussian sequence model:

align[align omitted — 78 chars of source]

where $\{\varepsilon_{i}, i = 1, 2, \cdots\} \overset{\mathrm{i.i.d.}}{\sim} \mathrm{N} (0, n^{-1})$. Let $\theta\coloneqq \{\theta_{i}, i = 1, 2, \cdots\}$ and $Y \coloneqq \{Y_{i}, i = 1, 2, \cdots\}$. We are interested in learning about the quadratic functional

align[align omitted — 118 chars of source]

The SA model corresponding to $\psi(\theta)$ is defined by balakrishnan2026fundamental as:

align[align omitted — 209 chars of source]

where $\widehat{\theta}$ is some initial estimator of $\theta$. It is noteworthy that we deliberately write $\widehat{\theta}$ and $r_{n}$ twice in the notation $\mathcal{P}_{\mathrm{SA}} ((\widehat{\theta}, \widehat{\theta}), (r_{n}, r_{n}))$ to emphasize that we take $J = 2$ and $\prod_{j = 1}^{J} r_{n, j} = r_{n}^{2}$ in this case. In the sequel, however, we write $\mathcal{P}_{\mathrm{SA}} (\widehat{\theta}, r_{n})$ instead in the text to simplify the notation. We adopt a similar convention for the quadratic density integral functional in Section (ref) and the expected conditional variance in Section (ref).

balakrishnan2026fundamental obtained the following results.

lemmaThe following hold. When $r_{n} \gtrsim n^{-1}$, \begin{align*} \mathfrak{R}_{n} (Q; \mathcal{P}_{\mathrm{SA}} (\widehat{\theta}, r_{n})) \coloneqq \inf_{\widehat{Q}} \sup_{\theta\in\mathcal{P}_{\mathrm{SA}} (\widehat {\theta}, r_{n})} \mathsf{E}_{\theta} [(\widehat{Q} - Q (\theta))^{2}] \gtrsim r_{n}^{2} + \frac{\Vert\widehat{\theta} \Vert_{2}^{2}}{n}. \end{align*} This lower bound is attained by the first-order estimator $\widehat{Q}_{1, n}$, defined as \begin{align*} \widehat{Q}_{1, n} \coloneqq 2 \langle Y, \widehat{\theta} \rangle- \Vert\widehat{\theta} \Vert_{2}^{2}. \end{align*} The bias, variance, and mean squared error (MSE) of $\widehat{Q}_{1, n}$ have the following forms: \begin{align*} \mathsf{bias} (\widehat{Q}_{1, n}) = - \Vert\widehat{\theta} - \theta\Vert_{2}^{2}, \mathsf{var} (\widehat{Q}_{1, n}) = \frac{4}{n} \Vert\widehat{\theta} \Vert_{2}^{2}, and \mathsf{mse} (\widehat{Q}_{1, n}) = \Vert\widehat{\theta} - \theta \Vert_{2}^{4} + \frac{4}{n} \Vert\widehat{\theta} \Vert_{2}^{2}. \end{align*}

The lower and upper bounds can be found in Theorem 1, Part 1 and Theorem 2, Part 1 of balakrishnan2026fundamental, respectively. These results, taken together, prove the rate optimality of $\widehat{Q}_{1, n}$ in the minimax sense.

To show the asymptotic inadmissibility of the minimax estimator $\widehat{Q} _{1,n}$, we need to exhibit a different estimator that improves upon $\widehat {Q}_{1,n}$. To this end, we adopt the following second-order estimator appeared in robins2006adaptive:

align*[align* omitted — 294 chars of source]

where $Y_{i,1}\coloneqq Y_{i}+\Phi^{-1}(U_{i})/\sqrt{n}$, $Y_{i,2} \coloneqq Y_{i}-\Phi^{-1}(U_{i})/\sqrt{n}$, $\Phi$ is the standard normal cumulative distribution function and $U_{i}$'s are independent uniform random variables over $[0,1]$. Here $Y_{i,1} \mathpalette{\protect\independenT}{\perp}Y_{i,2}$. The difference between $\widehat{Q}_{2,n}(k)$ and $\widehat{Q}_{1,n}$ is an unbiased estimator of $\Vert \mathsf{\Pi}_{k}(\theta-\widehat{\theta}) \Vert_{2}^{2}$, where $\mathsf{\Pi}_{k}(\cdot)$ denotes the projection onto the first $k$ coordinates of the input infinite-dimensional vector, with $\mathsf{\Pi}_{k}^{\perp} (\cdot)$ naturally meaning the projection onto the $(k+1)$-th coordinate and onward.

The following lemma characterizes the bias, variance, and mean squared error of $\widehat{Q}_{2, n} (k)$. The proof can be found in Appendix (ref).

lemmaThe bias and variance of $\widehat{Q}_{2, n} (k)$ read as: \begin{align*} \mathsf{bias} (\widehat{Q}_{2, n} (k)) & = - \Vert\mathsf{\Pi}_{k}^{\perp} (\widehat{\theta} - \theta) \Vert_{2}^{2} \lesssim r_{n},\\ \mathsf{var} (\widehat{Q}_{2, n} (k)) & = \frac{4}{n} \Vert\widehat{\theta} \Vert_{2}^{2} + \frac{4 k}{n^{2}} + \frac{4}{n} \Vert\mathsf{\Pi}_{k} (\widehat{\theta} - \theta) \Vert_{2}^{2} - \frac{4}{n} \langle\mathsf{\Pi}_{k} \widehat{\theta}, \mathsf{\Pi}_{k} (\widehat{\theta} - \theta) \rangle\lesssim\frac {1}{n}, \end{align*} where the inequalities hold for $\theta\in\mathcal{P}_{\mathrm{SA}} (\widehat{\theta}, r_{n})$.

We note that $\Vert\mathsf{\Pi}_{k} (\widehat{\theta} - \theta) \Vert_{2}^{2} \equiv\mathsf{bias} (\widehat{Q}_{1, n}) - \mathsf{bias} (\widehat{Q}_{2, n} (k))$ so $\widehat{Q}_{2, n} (k)$ corrects the bias of $\widehat{Q}_{1, n}$ by estimating a lower bound of $\mathsf{bias} (\widehat{Q}_{1, n}) \lesssim r_{n}$. By Lemma (ref), $Q (\theta)$ belongs to the monotone bias class. Comparing $\mathsf{mse} (\widehat{Q}_{2, n} (k))$ and $\mathsf{mse} (\widehat{Q}_{1, n})$ in the asymptotic sense, we obtain the first main statistical result of this paper. There always exists a distribution in $\mathcal{P}_{\mathrm{SA}} (\widehat{\theta}, r_{n})$ such that Definition (ref)(1) holds. To see this, consider the case where $\theta= \widehat{\theta} + r_{n}^{1 / 2} \upsilon$, where $\Vert\upsilon\Vert_{2} = 1$ and the coordinates of $\upsilon$ from $k + 1$ onward are all zeros. With this choice, $\mathsf{bias} (\widehat{Q}_{2, n} (\bar{\phi}_{k})) = 0$. The rest of the proof can be found in Appendix (ref).

theoremUnder Model $\mathcal{P}_{\mathrm{SA}}(\widehat{\theta}, r_{n})$ with $r_{n} \gtrsim n^{-1}$, $\widehat{Q}_{2,n}(k)$ is asymptotically minimax and the following hold as long as $k$ is chosen such that $k = o(n\Vert\widehat{\theta} \Vert_{2}^{2})$. \begin{align*} & \sup_{\theta\in\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})} \limsup_{n\rightarrow\infty}\frac{\mathsf{mse}(\widehat{Q}_{2,n}(k))-\mathsf{mse} (\widehat{Q}_{1,n})}{\mathsf{mse} (\widehat{Q}_{1,n})}\leq0, and when $r_{n}^{2} \gtrsim n^{-1}$\\ & \inf_{\theta\in\mathcal{P}_{\mathrm{SA}}(\widehat{\theta},r_{n})} \limsup_{n\rightarrow\infty}\frac{\mathsf{mse}(\widehat{Q}_{2,n}(k))-\mathsf{mse} (\widehat{Q}_{1,n})}{\mathsf{mse}(\widehat{Q}_{1,n})}<0. \end{align*} Thus, by Definition (ref), the first-order DML estimator $\widehat{Q}_{1,n}$ is asymptotically inadmissible when $r_{n}^{2} \gtrsim n^{-1}$. The same conclusions hold when we replace the SA model $\mathcal{P} _{\mathrm{SA}} (\widehat{\theta}, r_{n})$ with the assumption-lean model $\mathcal{P}_{\mathrm{AL}} (\widehat{\theta})$ and drop the assumptions on $r_{n}$.

Echoing the comment right after Theorem (ref), for the minimaxity of $\widehat{Q}_{1, n}$ or $\widehat{Q}_{2, n} (k)$, we need $r_{n} \gtrsim n^{-1}$. When $r_{n} \ll n^{-1}$, the so-called plug-in estimator $\widehat {Q}_{\mathrm{pi}} \coloneqq \Vert\widehat{\theta} \Vert^{2}$ has zero variance and squared bias of order $r_{n}$, thus dominating both $\widehat{Q}_{1, n}$ and $\widehat{Q}_{2, n} (k)$ when $\Vert\widehat{\theta} \Vert^{2}$ is of order 1. However, $r_{n} \ll n^{-1}$, or equivalently $\Vert\widehat{\theta} - \theta \Vert\ll n^{-1 / 2}$, is generally difficult to hold if the sample used to compute $\widehat{\theta}$ is of size similar to $n$, as such a condition says that we can estimate the possibly infinite-dimensional $\theta$ at a rate much faster than the parametric rate. A similar discussion also applies to the quadratic density integral functional to be discussed next.

Quadratic Density Integral Functional

The second example is about estimating the quadratic density integral functional of the probability density function $f$ of $X$ based on $n$ i.i.d. observations $\{X_{i} \in[0, 1]^{d}\}_{i = 1}^{n} \sim f$:

equation[equation omitted — 74 chars of source]

The SA model corresponding to $\psi(f)$ is defined by balakrishnan2026fundamental as:

align[align omitted — 290 chars of source]

where $\widehat{f}$ is some initial estimator of $f$ computed from a separate independent sample treated as fixed. Similar to the case in Section (ref), we take $J = 2$ and $\prod_{j = 1}^{J} r_{n, j} = r_{n}^{2}$, and write $\mathcal{P}_{\mathrm{SA}} (\widehat{f}, r_{n})$ instead.

remarkAs mentioned in footnote 1, we use the same estimator $\widehat{f} \equiv\widehat{f}_{1} \equiv\widehat{f}_{2}$ to compute $\widehat{\psi}_{1, n}$, which is the standard DML estimator commonly employed in the literature chernozhukov2018double but excludes more refined estimators with $f$ estimated by separate samples studied in newey2018cross, mcgrath2026nuisance, mcclean2026double.

The following lemma, paraphrasing the results of balakrishnan2026fundamental, summarizes the lower and upper bounds of the error rate of estimating $\psi(f)$ under $\mathcal{P}_{\mathrm{SA}} (\widehat{f}, r_{n})$.

lemmaThe following hold. When $r_{n} \gtrsim n^{-1}$, \begin{align*} \mathfrak{R}_{n} (\psi; \mathcal{P}_{\mathrm{SA}} (\widehat{f}, r_{n})) \coloneqq \inf_{\widehat{\psi}} \sup_{f \in\mathcal{P}_{\mathrm{SA}} (\widehat{f}, r_{n})} \mathsf{E}_{f} [(\widehat{\psi} - \psi(f))^{2}] \gtrsim r_{n}^{2} + \frac{1}{n} \left( \Vert\widehat{f} \Vert_{3}^{3} - \Vert\widehat{f} \Vert_{2}^{4} \right) . \end{align*} This lower bound is attained by the first-order estimator $\widehat{\psi}_{1, n}$, defined as \begin{align*} \widehat{\psi}_{1, n} \coloneqq \frac{2}{n} \sum_{i = 1}^{n} \widehat{f} (X_{i}) - \int\widehat{f} (x)^{2} \mathrm{d} x. \end{align*} The bias, variance and MSE of $\widehat{\psi}_{1, n}$ have the following forms: \begin{align*} \mathsf{bias} (\widehat{\psi}_{1, n}) & = - \int(\widehat{f} (x) - f (x))^{2} \mathrm{d} x \equiv- \Vert\widehat{f} - f \Vert_{2}^{2},\\ \mathsf{var} (\widehat{\psi}_{1, n}) & = \frac{4}{n} \mathsf{var} (\widehat{f} (X)) \equiv\frac{4}{n} \left\{ \int\widehat{f} (x)^{2} f (x) \mathrm{d} x - \left( \int\widehat{f} (x) f (x) \mathrm{d} x \right) ^{2} \right\} , and\\ \mathsf{mse} (\widehat{\psi}_{1, n}) & = \Vert\widehat{f} - f \Vert_{2}^{4} + \frac{4}{n} \mathsf{var} (\widehat{f} (X)). \end{align*}

The lower and upper bounds can be found in Theorem 1, Part 2 and Theorem 2, Part 2 of balakrishnan2026fundamental, respectively. These results, taken together, prove the optimality of $\widehat{\psi}_{1, n}$ in the minimax sense.

To show the asymptotic inadmissibility of the minimax estimator $\widehat{\psi }_{1, n}$, when $r_{n}^{2} \gtrsim n^{-1}$ or equivalently $r_{n}^{1 / 2} \gtrsim n^{-1 / 4}$, we exhibit a different estimator that improves on $\widehat{\psi}_{1, n}$. To this end, let $\bar{\phi}_{k} \coloneqq (\phi_{1}, \cdots, \phi_{k})^{\top}$ be a $k$-dimensional orthonormal basis with respect to the Lebesgue measure chen2007large. We then construct the following second-order $U$-statistic estimator:

align*[align* omitted — 613 chars of source]
remarkExpert readers shall realize that $\widehat{\psi}_{2,n} (\bar{\phi}_{k})$ debiases $\widehat{\psi}_{1,n}$ by subtracting an unbiased estimator of a part of its bias, based on HOIFs. The part of the bias of $\widehat{\psi}_{1,n}$ to be estimated is determined by the choice of $\bar{\phi }_{k}$. We mention in passing that the falsification test of liu2020nearly mentioned earlier is essentially based on the statistic $\widehat{\psi}_{2,n}(\bar{\phi}_{k})-\widehat{\psi}_{1,n}$. Similar tests or estimators have also been considered in instrumental variable or proximal causal inference settings breunig2024adaptive, liu2024assumption.

Let $\eta\coloneqq \int\bar{\phi}_{k} (x) f (x) \mathrm{d} x$ and $\widehat{\eta} \coloneqq \int\bar{\phi}_{k} (x) \widehat{f} (x) \mathrm{d} x$. We also make the following assumption on $\Sigma$.

assumption$\Sigma$ is assumed to have bounded spectra.

We now state the following lemma. The proof can be found in Appendix (ref).

lemmaThe bias and variance of $\widehat{\psi}_{2, n} (\bar{\phi} _{k})$ read as: \begin{align*} \mathsf{bias} (\widehat{\psi}_{2, n} (\bar{\phi}_{k})) = & - \int(\widehat{f} (x) - f (x))^{2} \mathrm{d} x + \int\mathsf{\Pi}[\widehat{f} - f \mid \bar{\phi}_{k}] (x)^{2} \mathrm{d} x \equiv- \Vert\widehat{f} - f \Vert_{2}^{2} + \Vert \mathsf{\Pi}[\widehat{f} - f \mid \bar{\phi}_{k}] \Vert_{2}^{2}\\ \equiv & - \Vert\widehat{f} - f \Vert_{2}^{2} + \Vert\widehat{\eta} - \eta\Vert _{2}^{2} \equiv- \Vert\mathsf{\Pi}^{\perp} [\widehat{f} - f \mid \bar{\phi}_{k}] \Vert_{2}^{2} \lesssim r_{n},\\ \mathsf{var} (\widehat{\psi}_{2, n} (\bar{\phi}_{k})) = & \ \frac{4}{n} \mathsf{var} [\widehat{f} (X)] + \frac{8}{n} \left( \int\bar{\phi}_{k} (x) f (x) \widehat{f} (x) \mathrm{d} x - \widehat{\eta} \right) ^{\top} (\widehat{\eta} - \eta)\\ & + \frac{2}{n (n - 1)} \left\{ \begin{array} [c]{c} \mathsf{Tr} (\Sigma^{2}) - 4 \widehat{\eta}^{\top} \Sigma\eta+ 2 \widehat{\eta}^{\top} \Sigma\widehat{\eta} + 2 \widehat{\eta}^{\top} \widehat{\eta} \cdot\eta^{\top} \eta\\ + \, 2 (\widehat{\eta}^{\top} \eta)^{2} - 4 \widehat{\eta}^{\top} \widehat{\eta} \cdot \widehat{\eta}^{\top} \eta+ (\widehat{\eta}^{\top} \widehat{\eta})^{2} \end{array} \right\} \\ \leq & \ \frac{4}{n} \mathsf{var} [\widehat{f} (X)] + \frac{C}{n} \Vert \mathsf{\Pi} [\widehat{f} - f \mid \bar{\phi}_{k}] \Vert_{2} + \frac{C k}{n^{2}} \lesssim\frac{1}{n}, \end{align*} where the inequalities hold for $f \in\mathcal{P}_{\mathrm{SA}} (\widehat{f}, r_{n})$ and under Assumption (ref).

We note that $\Vert\mathsf{\Pi}[\widehat{f} - f \mid \bar{\phi}_{k}] \Vert_{2}^{2} \equiv\mathsf{bias} (\widehat{\psi}_{1, n}) - \mathsf{bias} (\widehat{\psi}_{2, n} (\bar{\phi}_{k}))$ so $\widehat{\psi}_{2, n} (\bar{\phi}_{k})$ corrects the bias of $\widehat{\psi}_{1, n}$ by estimating a lower bound of $\mathsf{bias} (\widehat{\psi}_{1, n}) \lesssim r_{n}$. By Lemma (ref), $\psi(f)$ belongs to the monotone bias class. The following theorem therefore instantiates Theorem (ref) for the quadratic density integral functional $\psi(f)$. There always exists a distribution in $\mathcal{P} _{\mathrm{SA}} (\widehat{f}, r_{n})$ such that Definition (ref)(1) holds. To see this, consider the case where $f = \widehat{f} + r_{n}^{1 / 2} \beta^{\top} \bar{\phi}_{k}$ with $\Vert\beta\Vert_{2} = 1$, for which $\mathsf{bias} (\widehat{\psi}_{2, n} (\bar{\phi}_{k})) = 0$. The rest of the proof can be found in Appendix (ref).

theoremUnder model $\mathcal{P}_{\mathrm{SA}}(\widehat{f},r_{n})$ and Assumption (ref), $\widehat{\psi}_{2,n}(\bar{\phi}_{k})$ is asymptotically minimax and the following hold as long as $k$ is chosen such that $k=o(n\mathsf{var}[\widehat{f}(X)])$: \begin{align*} & \sup_{f\in\mathcal{P}_{\mathrm{SA}}(\widehat{f},r_{n})}\limsup_{n\rightarrow \infty}\frac{\mathsf{mse}(\widehat{\psi}_{2,n}(\bar{\phi}_{k}))-\mathsf{mse} (\widehat{\psi}_{1,n})}{\mathsf{mse}(\widehat{\psi}_{1,n})}\leq0, and when $r_{n}^{2} \gtrsim n^{-1}$\\ & \inf_{f\in\mathcal{P}_{\mathrm{SA}}(\widehat{f},r_{n})}\limsup_{n\rightarrow \infty}\frac{\mathsf{mse}(\widehat{\psi}_{2,n}(\bar{\phi}_{k}))-\mathsf{mse} (\widehat{\psi}_{1,n})}{\mathsf{mse}(\widehat{\psi}_{1,n})}<0. \end{align*} Thus, by Definition (ref), the first-order DML estimator $\widehat{\psi}_{1,n}$ is asymptotically inadmissible when $r_{n}^{2} \gtrsim n^{-1}$. The same conclusions hold when we replace the SA model $\mathcal{P}_{\mathrm{SA}} (\widehat{f}, r_{n})$ with the assumption-lean model $\mathcal{P}_{\mathrm{AL}} (\widehat{f})$ and drop the assumptions on $r_{n}$.

Expected Conditional Covariance

All the functionals that we have analyzed so far fall within the monotone bias class. In this section, we turn to the Expected Conditional Covariance (ECC) functional, defined as

align*[align* omitted — 118 chars of source]

where $X \in[0, 1]^{d}$ denotes the baseline covariates, $A, Y \in\mathbb{R}$ are two types of responses, $a (\cdot) \coloneqq \mathsf{E} (A \mid X = \cdot)$ and $b (\cdot) \coloneqq \mathsf{E} (Y \mid X = \cdot)$. The ECC functional, as extensively discussed in liu2020nearly, is not in the monotone bias class. Therefore, not surprisingly, we can no longer conclude the asymptotic inadmissibility of the first-order DML estimator $\widehat{\psi}_{1, n}$ for $\psi(a, b)$. Specifically, based on $n$ i.i.d. observations $\{X_{i}, A_{i}, Y_{i}\}_{i = 1}^{n}$, $\widehat{\psi}_{1, n}$ reads as:

align*[align* omitted — 135 chars of source]

As usual, before presenting our new results, we first summarize the statistical properties and minimaxity of $\widehat{\psi}_{1, n}$ obtained in balakrishnan2026fundamental under the SA model defined by balakrishnan2026fundamental for $\psi(a, b)$:

equation[equation omitted — 263 chars of source]

We let $p$ denote the marginal density of $X$, which, for simplicity, is assumed to be $\mathrm{Unif} ([0, 1]^{d})$.

lemmaThe following hold. \begin{align*} \mathfrak{R}_{n} (\psi; \mathcal{P}_{\mathrm{SA}} ((\widehat{a}, \widehat{b}), (r_{n}, s_{n}))) \coloneqq \inf_{\widehat{\psi}} \sup_{(a, b) \in\mathcal{P}_{\mathrm{SA}} ((\widehat{a}, \widehat{b}), (r_{n}, s_{n}))} \mathsf{E}_{a, b} [(\widehat{\psi} - \psi(a, b))^{2}] \gtrsim r_{n} \cdot s_{n} + \frac{1}{n}. \end{align*} This lower bound is attained by the first-order estimator $\widehat{\psi}_{1, n}$. The bias, variance, and MSE of $\widehat{\psi}_{1, n}$ have the following forms: \begin{align*} \mathsf{bias} (\widehat{\psi}_{1, n}) & = - \langle a - \widehat{a}, b - \widehat{b} \rangle_{\mathbb{P}},\\ \mathsf{var} (\widehat{\psi}_{1, n}) & = \frac{1}{n} \left\{ \mathsf{E} [(A - \widehat{a} (X))^{2} (Y - \widehat{b} (X))^{2}] - \mathsf{E}^{2} [(A - \widehat{a} (X)) (Y - \widehat{b} (X))] \right\} , and\\ \mathsf{mse} (\widehat{\psi}_{1, n}) & = \langle a - \widehat{a}, b - \widehat{b} \rangle_{\mathbb{P}}^{2} + \frac{1}{n} \left\{ \mathsf{E} [(A - \widehat{a} (X))^{2} (Y - \widehat{b} (X))^{2}] - \mathsf{E}^{2} [(A - \widehat{a} (X)) (Y - \widehat{b} (X))] \right\} . \end{align*}

To construct the second-order estimator, we similarly find a $k$-dimensional dictionary $\bar{\phi}_{k}$ and denote $\Sigma\coloneqq \mathsf{E} [\bar{\phi }_{k} (X)^{\otimes2}]$. In practice, one needs to estimate $\Sigma$ from data. We make the following assumptions on $\Sigma$ and its estimator.

assumption$\Sigma$ is assumed to have bounded spectra and there exists an estimator $\widehat{\Sigma}$ of $\Sigma$ such that $\widehat{\Sigma}$ also has bounded spectra and $\Vert\widehat{\Sigma} - \Sigma\Vert_{\mathrm{op}} = o (1)$, where $\Vert\cdot\Vert_{\mathrm{op}}$ denotes the matrix operator norm. Without loss of generality, we take $\Sigma= \Sigma^{-1} = \mathrm{I}$.

We then construct the following second-order estimator for $\psi(a, b)$.

align*[align* omitted — 419 chars of source]

We further introduce some short-hand notation for ease of exposition:

align*[align* omitted — 640 chars of source]

We are now ready to state the following lemma. The proof is by direct calculations and can be found in Appendix (ref).

lemmaThe bias and variance of $\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma})$ read as: \begin{align*} \mathsf{bias} (\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma})) & = - \langle\mathsf{\Pi}^{\perp} [\widehat{a} - a \mid \bar{\phi}_{k}], \mathsf{\Pi }^{\perp} [\widehat{b} - b \mid \bar{\phi}_{k}] \rangle_{\mathbb{P}} + \alpha^{\top} (\widehat{\Sigma}^{-1} - \mathrm{I}) \beta\lesssim r_{n}^{1 / 2} \cdot s_{n}^{1 / 2},\\ \mathsf{var} (\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma})) & = \mathsf{var} (\widehat{\psi}_{1, n}) + \mathsf{var} (\widehat{U}_{n, 2} (\bar{\phi }_{k}; \widehat{\Sigma})) + 2 \mathsf{cov} (\widehat{\psi}_{1, n}, \widehat{U}_{n, 2} (\bar{\phi}_{k}; \widehat{\Sigma})) \lesssim\frac{1}{n}, \end{align*} where \begin{align*} \mathsf{var} (\widehat{\psi}_{1, n}) = & \ \frac{1}{n} \left\{ \mathsf{E} [\widehat{\varepsilon}_{a}^{2} \widehat{\varepsilon}_{b}^{2}] - \mathsf{E}^{2} [\widehat{\varepsilon}_{a} \widehat{\varepsilon}_{b}] \right\} ,\\ \mathsf{var} (\widehat{U}_{n, 2} (\bar{\phi}_{k}; \widehat{\Sigma})) = & \ \frac {1}{n (n - 1)} \mathsf{Tr} \left\{ \Sigma_{a, a} \widehat{\Sigma}^{-1} \Sigma_{b, b} \widehat{\Sigma}^{-1} + (\Sigma_{a, b} \widehat{\Sigma}^{-1})^{2} \right\} \\ & + \frac{n - 2}{n (n - 1)} \left( \alpha^{\top} \widehat{\Sigma}^{-1} \Sigma_{b, b} \widehat{\Sigma}^{-1} \alpha+ \beta^{\top} \widehat{\Sigma}^{-1} \Sigma_{a, a} \widehat{\Sigma}^{-1} \beta+ 2 \alpha^{\top} \widehat{\Sigma}^{-1} \Sigma_{a, b} \widehat{\Sigma}^{-1} \beta\right) \\ & - \frac{2 (2 n - 3)}{n (n - 1)} (\alpha^{\top} \widehat{\Sigma}^{-1} \beta )^{2},\\ \leq & \ \frac{C k}{n^{2}} + \frac{C}{n} \left\{ \Vert\mathsf{\Pi}[\widehat{a} - a \mid \bar{\phi}_{k}] \Vert_{2, \mathbb{P}}^{2} + \Vert\mathsf{\Pi}[\widehat{b} - b | \bar{\phi}_{k}] \Vert_{2, \mathbb{P}}^{2} \right\} \lesssim\frac{1}{n}\\ 2 \mathsf{cov} (\widehat{\psi}_{1, n}, \widehat{U}_{n, 2} (\bar{\phi}_{k}; \widehat {\Sigma})) = & \ \frac{2}{n} \left\{ \mathsf{E} [\widehat{\varepsilon}_{a} \widehat{\varepsilon}_{b}^{2} \bar{\phi}_{k} (X)^{\top}] \widehat{\Sigma}^{-1} \alpha+ \mathsf{E} [\widehat{\varepsilon}_{a}^{2} \widehat{\varepsilon}_{b} \bar{\phi}_{k} (X)^{\top}] \widehat{\Sigma}^{-1} \beta- 2 \mathsf{E} [\widehat{\varepsilon}_{a} \widehat{\varepsilon}_{b}] \alpha^{\top} \widehat{\Sigma}^{-1} \beta\right\} \\ \leq & \ \frac{C}{n} \left\{ \Vert\mathsf{\Pi}[\widehat{a} - a \mid \bar{\phi}_{k}] \Vert_{2, \mathbb{P}} + \Vert\mathsf{\Pi}[\widehat{b} - b \mid \bar{\phi}_{k}] \Vert_{2, \mathbb{P}} \right\} \lesssim\frac{1}{n}. \end{align*} The inequalities hold for $(a, b) \in\mathcal{P}_{\mathrm{SA}} ((\widehat{a}, \widehat{b}), (r_{n}, s_{n}))$ under Assumption (ref).

It is not difficult to also see that $\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma})$ corrects the bias of $\widehat{\psi}_{1, n}$ by estimating a lower bound of $\mathsf{bias} (\widehat{\psi}_{1, n}) \lesssim r_{n}^{1 / 2} \cdot s_{n}^{1 / 2}$.

remarkIn liu2020nearly and liu2017semiparametric, we have shown that when $\widehat{\Sigma}$ is the sample Gram matrix estimator $\Vert\widehat{\Sigma} - \mathrm{I} \Vert _{\mathrm{op}} = \sqrt{k \log k / n} = o (1)$ when $k = o (n / \log^{2} n)$ when the sample used to compute $\widehat{\Sigma}$ is also of size $n$ tropp2015introduction. When we know $\Sigma= \mathrm{I}$, $\mathsf{bias} (\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \mathrm{I}))$ is reduced to $- \langle\mathsf{\Pi}^{\perp} [\widehat{a} - a \mid \bar{\phi}_{k}], \mathsf{\Pi }^{\perp} [\widehat{b} - b \mid \bar{\phi}_{k}] \rangle_{\mathbb{P}} \lesssim r_{n}^{1 / 2} \cdot s_{n}^{1 / 2}$ because there is no extra bias due to estimating $\Sigma$. However, even if we estimate $\Sigma$ by $\widehat{\Sigma}$, the extra bias incurred is of the form \begin{align*} \alpha^{\top} (\widehat{\Sigma}^{-1} - \mathrm{I}) \beta\lesssim\Vert\mathsf{\Pi }[\widehat{a} - a \mid \bar{\phi}_{k}] \Vert_{2, \mathbb{P}} \Vert\mathsf{\Pi} [\widehat{b} - a \mid \bar{\phi}_{k}] \Vert_{2, \mathbb{P}} \Vert\widehat{\Sigma}^{-1} - \mathrm{I} \Vert_{\mathrm{op}} = o (r_{n}^{1 / 2} \cdot s_{n}^{1 / 2}). \end{align*} Thus as long as we have a consistent estimator of $\Sigma$, the second-order estimator is still asymptotically minimax under the SA model.
theoremUnder Model $\mathcal{P}_{\mathrm{SA}} ((\widehat{a}, \widehat{b}), (r_{n}, s_{n}))$, both $\widehat{\psi}_{1, n}$ and $\widehat{\psi}_{2, n} (\bar{\phi }_{k}; \widehat{\Sigma})$ are asymptotically minimax, as long as $\Sigma$ and $\widehat{\Sigma}$ satisfy Assumption (ref). The same conclusions hold when we replace the SA model $\mathcal{P}_{\mathrm{SA}} ((\widehat{a}, \widehat{b}), (r_{n}, s_{n}))$ with the assumption-lean model $\mathcal{P}_{\mathrm{AL}} ((\widehat{a}, \widehat{b}))$.
proofThe minimaxity of $\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma})$ can be concluded using the orders of its bias and variance shown in Lemma (ref).
remarkSimilar statements to those in Theorem (ref) hold for the average treatment effect and the average treatment effect on the treated. For instance, for the treatment specific mean, this can be seen by replacing the notation $a, b, \widehat{\varepsilon}_{a}, \widehat{\varepsilon}_{b}$ by the following instead: \begin{align*} & a (\cdot) = 1 / \mathsf{E} [A \mid X = \cdot], b (\cdot) = \mathsf{E} [Y \mid X = \cdot, A = 1], \widehat{\varepsilon}_{a} = A \widehat{a} (X) - 1, \widehat{\varepsilon }_{b} = A (Y - \widehat{b} (X)). \end{align*} The dictionary $\bar{\phi}_{k}$ will also be weighted by the treatment indicator $A \bar{\phi}_{k}$. The minimaxity of the first-order DML estimators of these two functionals has been shown in jin2025structure.
remarkAs indicated after Definition (ref), we will discuss in Section (ref) that the higher-order generalization of the second-order estimators can be used to falsify the null hypothesis $\mathcal{H} _{0}:\mathsf{bias}(\widehat{\psi}_{1,n})\ll n^{-1/2}$ when $\psi$ belongs to the monotone bias class. If $\psi$ is the expected conditional covariance or the treatment specific mean parameter mentioned in Remark (ref), $\psi$ belongs to the so-called mixed-bias class rotnitzky2021characterization but not the monotone bias class. Here, we cannot claim the asymptotic inadmissibility of the DML estimator $\widehat{\psi}_{1,n}$ of $\psi$ and similarly we cannot directly falsify $\mathcal{H}_{0}:\mathsf{bias}(\widehat{\psi}_{1,n})\ll n^{-1/2}$. Nonetheless, we can empirically falsify the rate-double-robustness of $\widehat{\psi}_{1,n}$ liu2024assumption, where rate-double-robustness refers to the assumption $r_{n}^{1/2}\cdot s_{n}^{1/2}=o(n^{-1/2})$ for both the expected conditional covariance or the treatment specific mean parameter.

Specializing to the Expected Conditional Variance

A special case of the ECC functional--the expected conditional variance (abbreviated as the ECV functional) $\psi(a) \equiv\psi(a, a)$, however, belongs to the monotone bias class, when $A = Y$ with probability 1. Here, the corresponding DML estimator is $\widehat{\psi}_{1, n} \coloneqq n^{-1} \sum_{i = 1}^{n} (A_{i} - \widehat{a} (X_{i}))^{2}$ and the corresponding SA model is defined as

align*[align* omitted — 225 chars of source]
remarkSimilar to Remark (ref), we use the same estimator $\widehat{a} \equiv\widehat{a}_{1} \equiv\widehat{a}_{2}$ to compute $\widehat{\psi}_{1, n}$, again excluding the estimators studied in newey2018cross, mcgrath2026nuisance, mcclean2026double.

Analogously, the second-order estimator for $\psi(a)$ takes the following form:

align*[align* omitted — 417 chars of source]

Lemma (ref) and Lemma (ref) immediately imply the two corollaries below for the ECV functional $\psi(a)$.

corollaryThe following hold. \begin{align*} \mathfrak{R}_{n} (\psi; \mathcal{P}_{\mathrm{SA}} (\widehat{a}, r_{n})) \coloneqq \inf_{\widehat{\psi}} \sup_{a \in\mathcal{P}_{\mathrm{SA}} (\widehat{a}, r_{n})} \mathsf{E}_{a} [(\widehat{\psi} - \psi(a))^{2}] \gtrsim r_{n}^{2} + \frac{1}{n}. \end{align*} This lower bound is attained by the first-order estimator $\widehat{\psi}_{1, n}$. The bias, variance, and MSE of $\widehat{\psi}_{1, n}$ have the following forms: \begin{align*} \mathsf{bias} (\widehat{\psi}_{1, n}) & = - \Vert a - \widehat{a} \Vert_{2, \mathbb{P}}^{2} = - \, \Vert\mathsf{\Pi}^{\perp} [\widehat{a} - a \mid \bar{\phi} _{k}] \Vert_{2, \mathbb{P}}^{2} - \alpha^{\top} \alpha,\\ \mathsf{var} (\widehat{\psi}_{1, n}) & = \frac{1}{n} \left\{ \mathsf{E} [\widehat{\varepsilon}_{a}^{4}] - \mathsf{E}^{2} [\widehat{\varepsilon}_{a}^{2}] \right\} , and\\ \mathsf{mse} (\widehat{\psi}_{1, n}) & = \Vert a - \widehat{a} \Vert_{2, \mathbb{P} }^{4} + \frac{1}{n} \left\{ \mathsf{E} [\widehat{\varepsilon}_{a}^{4}] - \mathsf{E}^{2} [\widehat{\varepsilon}_{a}^{2}] \right\} . \end{align*}
corollaryThe bias and variance of $\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma})$ read as: \begin{align*} \mathsf{bias} (\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma})) & = - \Vert\mathsf{\Pi}^{\perp} [\widehat{a} - a \mid \bar{\phi}_{k}] \Vert_{2, \mathbb{P} }^{2} + \alpha^{\top} (\widehat{\Sigma}^{-1} - \mathrm{I}) \alpha\lesssim r_{n},\\ \mathsf{var} (\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma})) & = \mathsf{var} (\widehat{\psi}_{1, n}) + \mathsf{var} (\widehat{U}_{n, 2} (\bar{\phi }_{k}; \widehat{\Sigma})) + 2 \mathsf{cov} (\widehat{\psi}_{1, n}, \widehat{U}_{n, 2} (\bar{\phi}_{k}; \widehat{\Sigma})) \lesssim\frac{1}{n}, \end{align*} where \begin{align*} \mathsf{var} (\widehat{\psi}_{1, n}) = & \ \frac{1}{n} \mathsf{var} (\widehat{\varepsilon}_{a}^{2}),\\ \mathsf{var} (\widehat{U}_{n, 2} (\bar{\phi}_{k}; \widehat{\Sigma})) = & \ \frac {2}{n (n - 1)} \mathsf{Tr} \left\{ (\Sigma_{a, a} \widehat{\Sigma}^{-1})^{2} \right\} + \frac{4 n - 8}{n (n - 1)} \left( \alpha^{\top} \widehat{\Sigma}^{-1} \Sigma_{a, a} \widehat{\Sigma}^{-1} \alpha\right) \\ & - \frac{4 n - 6}{n (n - 1)} (\alpha^{\top} \widehat{\Sigma}^{-1} \alpha)^{2}\\ \leq & \ \frac{C k}{n^{2}} + \frac{C}{n} \Vert\mathsf{\Pi}[\widehat{a} - a \mid \bar{\phi}_{k}] \Vert_{2, \mathbb{P}}^{2} \lesssim\frac{1}{n},\\ 2 \mathsf{cov} (\widehat{\psi}_{1, n}, \widehat{U}_{n, 2} (\bar{\phi}_{k}; \widehat {\Sigma})) = & \ \frac{4}{n} \left\{ \mathsf{E} [\widehat{\varepsilon}_{a}^{3} \bar{\phi}_{k} (X)^{\top}] \widehat{\Sigma}^{-1} \alpha- \mathsf{E} [\widehat {\varepsilon}_{a}^{2}] \alpha^{\top} \widehat{\Sigma}^{-1} \alpha\right\} \\ \leq & \ \frac{C}{n} \Vert\mathsf{\Pi}[\widehat{a} - a \mid \bar{\phi}_{k}] \Vert_{2, \mathbb{P}} \lesssim\frac{1}{n}. \end{align*} The inequalities hold for $a \in\mathcal{P}_{\mathrm{SA}} (\widehat{a}, r_{n})$ under Assumption (ref).

By Corollary (ref), $\psi(a)$ belongs to the monotone bias class. By piecing together the above two corollaries, we obtain the final theoretical result of this paper. There always exists a distribution in $\mathcal{P}_{\mathrm{SA}} (\widehat{a}, r_{n})$ such that Definition (ref)(1) holds. To see this, consider the case where $a = \widehat{a} + r_{n}^{1 / 2} \beta^{\top} \bar{\phi}_{k}$ with $\Vert\beta\Vert_{2} = 1$, for which $\mathsf{bias} (\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat {\Sigma})) = \alpha^{\top} (\widehat{\Sigma}^{-1} - \mathrm{I}) \alpha \ll\mathsf{bias} (\widehat{\psi}_{1, n}) = \alpha^{\top} \alpha$. The rest of the proof can be found in Appendix (ref).

theoremUnder Model $\mathcal{P}_{\mathrm{SA}} (\widehat{a}, r_{n})$ and Assumption (ref), $\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma})$ is asymptotically minimax and the following hold as long as $k$ is chosen such that $k = o (n \mathsf{var} (\widehat{\varepsilon}_{a}^{2}))$: \begin{align*} & \sup_{a \in\mathcal{P}_{\mathrm{SA}} (\widehat{a}, r_{n})} \limsup_{n \rightarrow\infty} \frac{\mathsf{mse} (\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma})) - \mathsf{mse} (\widehat{\psi}_{1, n})}{\mathsf{mse} (\widehat{\psi }_{1, n})} \leq0, and when $r_{n}^{2} \gtrsim n^{-1}$\\ & \inf_{a \in\mathcal{P}_{\mathrm{SA}} (\widehat{a}, r_{n})} \limsup_{n \rightarrow\infty} \frac{\mathsf{mse} (\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma})) - \mathsf{mse} (\widehat{\psi}_{1, n})}{\mathsf{mse} (\widehat{\psi }_{1, n})} < 0. \end{align*} Thus, by Definition (ref), the first-order DML estimator $\widehat{\psi}_{1, n}$ is asymptotically inadmissible when $r_{n}^{2} \gtrsim n^{-1}$. The same conclusions hold when we replace the SA model $\mathcal{P}_{\mathrm{SA}} (\widehat{a}, r_{n})$ with the assumption-lean model $\mathcal{P}_{\mathrm{AL}} (\widehat{a})$ and drop the assumption on $r_{n}$.
remarkWe note that $\mathsf{bias} (\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma}))^{2} - \mathsf{bias} (\widehat{\psi}_{1, n})^{2} \asymp- \Vert\mathsf{\Pi}[\widehat{a} - a \mid \bar{\phi}_{k}] \Vert_{2, \mathbb{P} }^{2} (1 - \Vert\widehat{\Sigma}^{-1} - \mathrm{I} \Vert_{\mathrm{op}}) \asymp- \Vert\mathsf{\Pi} [\widehat{a} - a \mid \bar{\phi}_{k}] \Vert_{2, \mathbb{P}}^{2}$ by Assumption (ref). Thus, asymptotically, the second-order estimator still has smaller bias than the first-order DML estimator $\widehat{\psi}_{1, n}$, and the impact of estimating $\Sigma$ is asymptotically negligible. In addition, $\widehat{\psi}_{2, n} (\bar{\phi}_{k}; \widehat{\Sigma})$ corrects the bias of $\widehat{\psi}_{1, n}$ by estimating a lower bound of $\mathsf{bias} (\widehat{\psi}_{1, n}) \lesssim r_{n}$.

Concluding Remarks

The SA model introduced in balakrishnan2026fundamental is a mathematically appealing construct that has inspired follow-up work jin2025structure, bonvini2024doubly, jin2025normal, jin2025sharp, gu2026optimally, gu2025open, including our current paper. The assumption-lean model $\mathcal{P}_{\mathrm{AL}}(\widehat{\theta})$ defined in Definition (ref), is aligned with the goal of understanding what can be learned from a model that makes almost no assumptions. As discussed earlier, in terms of point estimation, both the first-order DML estimator $\widehat{\psi}_{1,n}$ and our second-order estimator $\widehat{\psi}_{2,n}$ remain minimax with rate $O (1)$ in the assumption-lean model $\mathcal{P}_{\mathrm{AL}}(\widehat{\theta})$ for all the parameters studied in balakrishnan2026fundamental; for $\psi$ in monotone bias class, our $\widehat{\psi}_{2,n}$ continues to asymptotically dominate $\widehat{\psi}_{1,n}$ in the scaled MSE.

However, statisticians care about uncertainty quantification or inference as much as or even more than point estimation. Neither the (asymptotic) minimaxity/inadmissibility of $\widehat{\psi}_{1,n}$ nor the minimaxity of $\widehat{\psi}_{2,n}$ in the assumption-lean model $\mathcal{P}_{\mathrm{AL}}(\widehat{\theta})$ offer any guidance on how to quantify uncertainty, absent further knowledge of $\widehat{\theta}$ or $\Theta$. The above argument is not new. Before balakrishnan2026fundamental, we considered inference on $\psi$ under the assumption-lean model $\mathcal{P}_{\mathrm{AL}} (\widehat{\theta})$ in liu2020nearly and liu2024assumption. The former paper was discussed by the authors of balakrishnan2026fundamental; see kennedy2020discussion and liu2020rejoinder. Since no uniformly consistent estimators of $\psi$ exist in model $\mathcal{P}_{\mathrm{AL}}(\widehat{\theta})$ ritov1990achieving, robins1997toward, ritov2014bayesian, we, instead, developed valid falsification tests of the following null hypothesis for $\psi$ in the monotone bias class:

quote$\mathcal{H}_{0}$: The bias of the first-order DML estimator $\widehat{\psi }_{1, n}$ of $\psi$ is sufficiently small such that a standard Wald CI centered at $\widehat{\psi}_{1,n}$ has nominal coverage asymptotically.

The proposed tests are only falsification tests because, although valid under $\mathcal{H}_{0}$, they will have no power under many alternatives to $\mathcal{H}_{0}$. However, when a test rejects the null, it provides empirical evidence that the bias of $\widehat{\psi}_{1,n}$ is too large for the Wald CI to deliver valid inference. The test statistics used are based on the same second-order estimators that we have analyzed in this paper or their higher-order extensions robins2008higher, robins2016technical, liu2017semiparametric. For $\psi$ belonging to the so-called mixed-bias classes (which includes the expected conditional covariance analyzed above) rotnitzky2021characterization, in liu2020nearly and liu2024assumption, we showed that these tests are no longer valid under $\mathcal{H}_{0}$. However, these tests remain valid falsification tests of the so-called rate-double-robustness property, as defined in Remark (ref) or Remark (ref). We note that the rate-double-robustness implies that $\mathcal{H}_{0}$ is true. For this reason, complexity-reducing assumptions strong enough to imply rate-double-robustness are often made by investigators to justify the validity of their Wald CIs centering $\widehat{\psi}_{1,n}$. In our view, unlike the minimaxity of $\widehat{\psi}_{1,n}$ or of $\widehat{\psi}_{2,n}$, these falsification tests provide further empirical information even in the assumption-lean model $\mathcal{P}_{\mathrm{AL}}(\widehat{\theta})$, whenever they reject and thus can be of value to domain scientists for whom inference is important.