Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
113,859 characters · 25 sections · 61 citation commands
Sharp Structure-Agnostic Lower Bounds for General Linear Functional Estimation
\onehalfspacing
Let $\{O_t\}_{t=1}^N$ be i.i.d.\ training samples from an unknown distribution $P_0$ on ${\mathcal{O}}={\mathcal{Z}}\times{\mathcal{W}}$, where each $O=(Z,W)$ contains covariates $Z$ and an outcome $W$, and let $\{Z_i\}_{i=1}^N$ be i.i.d.\ target covariate samples from an unknown distribution $Q_0$ on ${\mathcal{Z}}$. We consider the problem of estimating a linear functional of a regression-type nuisance learned under the training law $P_0$ but evaluated under the target law $Q_0$:
where for any fixed $z$ the mapping $\gamma\mapsto m_1(z,\gamma)$ is linear, and the nuisance function $\gamma(z;P)$ is the solution to a generalized regression problem under the training law:
where $\mu_Z$ is some known measure on ${\mathcal{Z}}=\mathrm{supp}(Z)$.
This formulation covers many causal and predictive estimation tasks and has found important applications in numerous disciplines such as economics hirano2003efficient,imbens2004nonparametric, education oreopoulos2006estimating, epidemiology little2000causal,wood2008empirical, and political science mayer2011does.
Estimating the ATE is one of the central problems in causal inference. In view of its practical importance, a large body of work is devoted to developing statistically efficient estimators for the ATE based on regression robins1994estimation,robins1995analysis,imbens2003mean, matching heckman1998matching,rosenbaum1989optimal,abadie2006large, and propensity scores rosenbaum1983central,hirano2003efficient, as well as their combinations. Beyond ATE, influence-function-based methods have been developed for a range of related estimands, including selection/conditioning targets such as ATT, policy learning objectives, weighted average derivatives, and covariate-shift/data-fusion targets; see, for example, athey2021policy,newey1993efficiency,powell1989semiparametric,sugiyama2007covariate,reddi2015doubly and references therein.
Statistical limits for estimating treatment-effect-type parameters are studied in robins2009semiparametric,balakrishnan2019hypothesis,kennedy2022minimax,robins2008higher, typically under H\"older-smoothness assumptions on the nuisances. When the nonparametric components of the data-generating process are estimable at a fast enough rate (typically $n^{-1/4}$), semiparametric efficiency newey1994asymptotic provides optimal variance constants multiplying the leading rate. Finally, bradic2019minimax characterizes minimax conditions for root-$n$ estimability, albeit under strong linearity restrictions and constant effects. These works crucially rely on structural assumptions on the underlying function classes, which enables tight rates but can be cumbersome to deploy in practice when the relevant structure is unknown or violated.
Since the nuisance function $\gamma$ in (ref) is unknown and may have complex structures, and since the dimension $K$ of the covariates $X$ can be large relative to the number of data $n$ in many applications, it is extremely suitable to apply modern machine learning (ML) methods for the non-parametric, flexible and adaptive estimation of these nuisance functions, including penalized linear regression methods belloni2014pivotal,van2014asymptotically,chernozhukov2022automatic,zou2005regularization, random forest methods breiman2001random,hastie2009random,biau2008consistency,wager2015adaptive,syrgkanis2020estimation, gradient boosted forests friedman2001greedy,buhlmann2003boosting,zhang2005boosting and neural networks schmidt2020nonparametric,farrell2021deep, as well as ensemble and model selection approaches that combine all the above using out-of-sample cross-validation metrics wolpert1992stacked,zhang1993model,freund1997decision,van2007super,dvzeroski2004combining,sill2009feature,wegkamp2003model,Arlot2010,chetverikov2021cross. However, ML methods typically require some forms of regularization to avoid overfitting, which can potentially make the resulting estimator severely biased.
A principled way to combine flexible nuisance estimation with accurate estimation of a low-dimensional target is to use orthogonal (Neyman-orthogonal) estimating equations derived from influence-function theory robins1995analysis,robins1995semiparametric. Double/debiased machine learning (DML) is a prominent and widely used implementation of this idea chernozhukov2017double,chernozhukov2018double,athey2021policy,chernozhukov2022automatic: one estimates the nuisances (often via cross-fitting) and then evaluates an orthogonal score whose first-order sensitivity to nuisance estimation errors vanishes at the truth. As a consequence, the estimation error admits a decomposition of the schematic form \[ \text{(parametric noise)}\;+\;\text{(higher-order remainder depending on nuisance errors)}, \] where the leading remainder is typically proportional to a product of nuisance estimation errors (and, for some generalized regression targets, may also include a squared term); see chernozhukov2018double and the references therein.
In the special case with no covariate shift ($Q_0=P_{0,Z}$) and ordinary least-squares regression ($\ell(o,\gamma)=(w-\gamma)^2/2$, so the score $\rho(o,\gamma)=\gamma-w$ is affine), write $\gamma_0(z):=\gamma(z;P_0)$ (which equals $\mathbb{E}[W\mid Z=z]$ for squared loss). Assume that the linear functional $h\mapsto \mathbb{E}_{Q_0}[m_1(Z,h)]$ is continuous on $L^2(P_{0,Z})$. Equivalently, there exists a weight $\nu_m(\cdot;P_0,Q_0)\in L^2(P_{0,Z})$ such that \[ \mathbb{E}_{Q_0}\big[m_1(Z,h)\big]=\mathbb{E}_{P_{0,Z}}\big[h(Z)\,\nu_m(Z;P_0,Q_0)\big]\qquad\text{for all }h\in L^2(P_{0,Z}). \] Since $\rho(o,\gamma)=\gamma-w$ has derivative $1$ in its regression argument, we define the orthogonal weight \[ \alpha_0(z):=\alpha(z;P_0,Q_0):=-\nu_m(z;P_0,Q_0). \] The corresponding orthogonal score is $\psi(O;\gamma,\alpha):=m_1(Z,\gamma)+\alpha(Z)\rho(O,\gamma(Z))$, and the cross-fitted estimator can be written in the augmented plug-in form \[ \hat\chi = \frac{1}{n}\sum_{i=1}^n m_1\big(Z_i,\hat\gamma\big) \;+\; \frac{1}{n}\sum_{i=1}^n \hat\alpha(Z_i)\,\big(\hat\gamma(Z_i)-W_i\big), \] with sample splitting/cross-fitting to ensure independence between the evaluation sample and the first-stage fits. Moreover, letting $\mathbb{P}_n$ denote the empirical measure of $\{O_i\}_{i=1}^n$, a standard orthogonality expansion gives (up to negligible empirical-process terms controlled by cross-fitting) \[ \hat\chi-\chi(P_0,Q_0) = (\mathbb{P}_n-P_0)\psi(O;\gamma_0,\alpha_0) \;+\; \mathbb{E}_{P_0}\Big[\big(\hat\alpha(Z)-\alpha_0(Z)\big)\big(\hat\gamma(Z)-\gamma_0(Z)\big)\Big], \] so $\hat\chi-\chi(P_0,Q_0)={\mathcal{O}}_P(n^{-1/2}+{\epsilon}_{n,\gamma}{\epsilon}_{n,\alpha})$ under the mean-squared error constraints imposed below. This approach also generalizes to the setting with covariate shift and generalized regression, as shown in chernozhukov2023automatic and chernozhukov2021automatic respectively. The generalized approach will be discussed in details in Section (ref).
Motivated by the wide adoption and use of black-box adaptive estimation methods polley2019package,ledell2020h2o,wang2021flaml,karmaker2021automl for these non-parametric components of the data generating process, as well as their superior empirical performance bach2024hyperparameter, we will examine the statistical optimality of the aforementioned procedure within the structure-agnostic minimax framework that was recently introduced in balakrishnan2023fundamental. In particular, the only assumption that we will be making about our data generating process is that we have access to estimates $\hat{\gamma}$ and $\hat{\alpha}$ that achieve some statistical error rate, as measured by the mean-squared error, i.e. $\|\hat{\gamma}(Z)-\gamma(Z;P_0)\|_{P_{0,Z},2}\leq {\epsilon}_{n,\gamma}$ and $\|\hat{\alpha}(Z)-\alpha(Z;P_0,Q_0)\|_{P_{0,Z},2}\leq {\epsilon}_{n,\alpha}$, where for any function $v:\ensuremath{{\cal X}} \to 3$, we denote $\|v(X)\|_{P_X,2}:=\sqrt{\mathbb{E}[v(X)^2]}$. Having access to such estimates for these two non-parametric components and imposing the aforementioned estimation error constraints on the data generating process, we resolve the optimal statistical rate achievable by any estimation algorithm for the parameters of interest.
The structure-agnostic framework is particularly appealing as it essentially restricts any estimation approach to only use non-parametric regression estimates as a black-box and not tailor the estimation strategy to particular structural assumptions about the regression function or the propensity. These further structural assumptions can many times be brittle and violated in practice, rendering the tailored estimation strategy invalid or low-performing. Hence, the structure-agnostic statistical lower bound framework has the benefit that it yields lower bounds that can be matched by estimation procedures that are easy to deploy and robust.
\paragraph{Contributions and main message.} Our main contribution is a general, sharp structure-agnostic lower bound theory for a broad class of functionals of the form (ref), where the nuisance $\gamma(\cdot;P)$ is defined as the solution to a (generalized) regression problem and the functional is linear in $\gamma$. The class includes the average treatment effect and a range of causal and policy estimands that admit influence-function-based orthogonal scores. Under assumptions that we verify for a collection of examples, our results identify two regimes:
In both regimes we provide matching upper bounds via first-order debiasing/DML estimators. Consequently, without additional structural information beyond mean-squared error guarantees for nuisance estimation, one cannot improve the dependence on nuisance errors beyond what is achieved by DML.
For general non-parametric functional estimation, it has been known for decades that if the function possesses certain smoothness properties, then higher-order debiasing schemes can be designed that lead to improved error rates bickel1988estimating,birge1995estimation. Specifically, first-order debiasing methods are suboptimal even when the nuisance function estimators are minimax optimal. Estimators based on higher-order debiasing have also been proposed and analyzed for functionals that arise in causal inference problems robins2008higher,van2014higher,robins2017minimax,liu2017semiparametric,kennedy2022minimax. However, none of these approaches enjoy the structure-agnostic property that we explicitly impose in our minimax framework.
We then instantiate the general theory for a broad range of estimands, including the average treatment effect (ATE), average treatment effect on the treated (ATT), expected conditional covariance (ECC), weighted average derivative (WAD), distribution shift (DS), average policy effect (APE), log-odds-difference (LOD) and expected derivatives of conditional quantiles (EQD).
Prior to this work, optimal structure-agnostic error rates are not well understood, except for a few specific problem instances. balakrishnan2023fundamental was the first to establish sharp structure-agnostic lower bounds, but their proof techniques only apply to inner product functionals like ECC (see additional discussions in Section (ref)). Later, jin2024structure established similar results separately for ATE and ATT. This paper is a generalized version of jin2024structure and subsumes the results therein. Another recent work jin2025its considered structure-agnostic estimation in a partial linear outcome model and an in-depth discussion of their results can be found in Remark (ref).
\paragraph{Technical contributions.} The main technical contribution is a general lower-bound principle that applies uniformly across a broad class of statistical estimands, including targets that involve generalized regression that fall outside the mixed-bias regime. Our proof relies on a number of novel technical ideas, as we explain next.
Our lower bounds are proved via the method of fuzzy hypotheses, reducing estimation to testing between carefully constructed mixtures. The core difficulty is to build composite null and alternative hypotheses that (i) remain within the prescribed structure-agnostic nuisance neighborhood and (ii) induce separation in the target functional of the desired order, while keeping the two mixtures close in Hellinger distance. To achieve this we introduce a two-step sequential perturbation construction that decouples feasibility (staying inside the nuisance neighborhood) from separation (moving the target functional). A key ingredient is a geometric partitioning/“pairing” argument (based on ham-sandwich-type results) that lets us place localized perturbations while enforcing the exact invariances required by our lower-bound theorem. A complete overview of our proofs can be found in Section (ref).
We write $O=(Z,W)\in{\mathcal{O}}={\mathcal{Z}}\times{\mathcal{W}}$ for a generic observation. We observe i.i.d.\ training samples $\{O_t\}_{t=1}^N\sim P_0$ on ${\mathcal{O}}$ and i.i.d.\ target covariates $\{Z_i\}_{i=1}^N\sim Q_0$ on ${\mathcal{Z}}$; more generally we write $P$ and $Q$ for candidate training and target laws. The nuisance $\gamma(\cdot;P)$ is a function on ${\mathcal{Z}}$ (typically in $L^2(\mu_Z)$), and the target functional has the form $\chi(P,Q)=\mathbb{E}_Q[m_1(Z,\gamma(\cdot;P))]$ with $m_1(z,\cdot)$ linear, as in (ref)--(ref). We use subscripts to denote marginals: if $P$ is a distribution on a product space, we write $P_Z$ (resp.\ $P_X$) for the marginal law of $Z$ (resp.\ $X$). We write $\mathbb{E}_P[\cdot]$ for expectation under $P$ (and similarly $\mathbb{E}_Q[\cdot]$), and we use $\mathbb{P}_n$ for the empirical measure of an i.i.d.\ sample of size $n$ when this is convenient.
For any function $f:3^n\mapsto3^k$ and distribution $P$ over $3^n$, we define its $L^r(P)$ norm as $\left\Vertf\right\Vert_{P,r}=\big(\int \left\Vertf(x)\right\Vert^r\,\textup{d} P(x)\big)^{1/r}$ for $r\in(0,\infty)$, and $\left\Vertf\right\Vert_{P,\infty}=\text{ess sup}\{\left\Vertf(X)\right\Vert:X\sim P\}$. When the distribution is clear from context we also write $\left\Vertf\right\Vert_r$. For deterministic sequences $(a_n)_{n\ge 1}$ and $(b_n)_{n\ge 1}$ we write $a_n={\mathcal{O}}(b_n)$ if there exists $C>0$ such that $|a_n|\le C|b_n|$ for all $n$, and $a_n=\Omega(b_n)$ if there exists $c>0$ such that $|a_n|\ge c|b_n|$ for all $n$. For random variables, $X_n={\mathcal{O}}_P(b_n)$ means $X_n/b_n$ is bounded in probability. We write $L^r(P)$ for the corresponding function space $\{f:\left\Vertf\right\Vert_{P,r}<\infty\}$.
Throughout this paper we fix $\sigma$-finite reference measures $\mu_Z$ on ${\mathcal{Z}}$ and $\mu_W$ on ${\mathcal{W}}$, and write $\mu=\mu_Z\otimes\mu_W$ on ${\mathcal{O}}={\mathcal{Z}}\times{\mathcal{W}}$ (often the uniform measure on its support). Our theory applies to probability measures that are absolutely continuous with respect to these reference measures. For $P\ll\mu$ we write $p=\textup{d} P/\textup{d}\mu$ for the density and $p_Z(z):=\int p(z,w)\,\textup{d}\mu_W(w)$ (also denoted $p(z,\cdot)$) for the $Z$-marginal density. For $Q\ll\mu_Z$ we write $q=\textup{d} Q/\textup{d}\mu_Z$ for its density. For any two distributions $P_1,P_2\ll\mu$ with densities $p_1,p_2$ and common support ${\mathcal{O}}$, we define their $L_\infty$ distance by $d_{\mu,\infty}(P_1,P_2)=\text{ess sup}_{o\in{\mathcal{O}}}|p_1(o)-p_2(o)|$.
We define the directional derivative of a functional $\chi(P,Q)$ at $(P,Q)$ in the direction of a joint perturbation pair $(H,K)$ (when it exists) as \[ \chi_{(P,Q)}'(P,Q)[H,K] := \left.\frac{\textup{d}}{\textup{d} t}\right|_{t=0}\chi(P+tH,Q+tK), \] where $H$ is a finite signed measure on ${\mathcal{O}}$ and $K$ is a finite signed measure on ${\mathcal{Z}}$. Similarly, we define the mixed second derivative in directions $(H_0,K_0)$ and $(H_1,K_1)$ by \[ \chi''(P,Q)[(H_0,K_0),(H_1,K_1)] := \left.\frac{\textup{d}^2}{\textup{d} t\,\textup{d} s}\right|_{t=s=0}\chi(P+tH_0+sH_1,Q+tK_0+sK_1), \] and \[ \chi''(P,Q)[(H,K)]:=\chi''(P,Q)[(H,K),(H,K)]. \] When the functional of interest is of the form $\Psi(U,P,Q)$ where $U$ is an additional parameter, we use $\Psi_P'(U_0,P_0,Q_0)[H]$, $\Psi_Q'(U_0,P_0,Q_0)[K]$ and their second-order analogues to denote partial distributional derivatives.
The main contribution of this paper is general lower bounds on the estimation error for functionals of the form (ref), in the case where no structural priors are available. In this subsection, we provide an overview of these results before stating the formal results and assumptions.
Given estimates $\hat{\gamma}, \hat{\alpha}$ of $\gamma$ and $\alpha$ (defined later in Section (ref), where $\alpha$ is some transformation of $m$ for ATE) and some specified error bounds ${\epsilon}_{n,\gamma}$ and ${\epsilon}_{n,\alpha}$, the set of all plausible ground-truth data distributions $P$ consists of those with nuisance functions $\gamma(Z;P)$ and $\alpha(Z;P)$ satisfying
Any estimator $\hat{\theta}$ can be viewed as a (possibly random) mapping from the observed data $\{O_i\}_{i=1}^n$ to $\mathbb{R}$. For any distribution $P$, when $\{O_i\}_{i=1}^n$ are i.i.d. samples from $P$, the estimator $\hat{\theta}$ induces a distribution of estimates on $\mathbb{R}$. Let $\xi\in(0,1)$ be a pre-specified tolerance probability. By comparing this distribution with the true parameter $\theta(P)$, we can measure the quality of the estimator $\hat{\theta}$ via the $(1-\xi)$-quantile of $|\hat{\theta}-\theta(P)|$. The worst-case error of $\hat{\theta}$ is then naturally defined as the supremum of this quantile over all possible $P$ satisfying the nuisance constraint in (ref). Our main result can be summarized as follows:
Our proof of the lower bounds uses the method of fuzzy hypotheses, which reduces our estimation problem to testing between a pair of mixtures of hypotheses. While such methods are widely adopted in establishing lower bounds for non-parametric functional estimation problems tsybakov2008introduction,robins2009semiparametric,kennedy2022minimax,balakrishnan2023fundamental, we introduce a novel two-step sequential perturbation technique to construct the null and alternative hypotheses with the desired properties. The two perturbation steps are asymmetric in general, and interchanging them would lead to two different types of optimal rates. We elaborate on this technique in Section (ref). Due to the more complicated relationship between the estimand and the data distribution, existing constructions of composite hypotheses robins2009semiparametric,kennedy2022minimax,balakrishnan2023fundamental do not apply to our setting, as we explain next.
In balakrishnan2023fundamental, the authors investigate the estimation problem of three functionals: quadratic functionals in Gaussian sequence models, quadratic integral functionals, and the expected conditional covariance. They establish their lower bound by reducing it to a related hypothesis testing problem. The testing error is then lower-bounded by constructing priors (mixtures) over the composite null and alternative hypotheses. The priors they construct are based on adding or subtracting “bumps” on top of a fixed hypothesis in a symmetric manner, which is a standard proof strategy for functional estimation problems ingster1994minimax,robins2009semiparametric,arias2018remember,balakrishnan2019hypothesis. The reason why the proof strategy of balakrishnan2023fundamental fails for ATE and most other functionals is that the functional relationships between the nuisance parameters and these target parameters take significantly different forms. Specifically, the target parameters that balakrishnan2023fundamental investigates are all of the form
where $f,g$ are unknown nuisance parameters that lie in some Hilbert space ${\mathcal{H}}$. To be concrete, consider the example of the expected conditional covariance $\theta^{\textsc{Cov}}$. Let $\mu_0(x) = \mathbb{E}\left[ Y \mid X=x\right]$, then we have that
where $p_X$ is the marginal density of $X$. The first term, $\mathbb{E}[DY]$, can be estimated at the standard ${\mathcal{O}}(n^{-1/2})$ rate, so it suffices to estimate the second term, which is exactly in the form of (ref). However, the ATE functional does not take this inner product form. Instead, it is of the form \opt{opt-arxiv}{
} \opt{opt-or}{
} Stepping outside of the realm of inner product functionals is the major challenge in extending existing approaches of establishing lower bounds to ATE and other relevant functionals, and is our main technical innovation.
Before going into full generality, we first revisit the ATE example to build intuition for first-order debiasing and the structure-agnostic viewpoint. In the standard setting, we observe $O=(X,D,Y)$, where $X$ is a high-dimensional covariate vector, $D\in \{0,1\}$ is a binary treatment, and $Y\in3$ is an outcome. Let $Y(1)$ and $Y(0)$ denote the potential outcomes under each treatment level. The average treatment effect (ATE) is defined as
We consider the case when all potential confounders $X\in \ensuremath{{\cal X}} \subseteq 3^K$ of the treatment and outcome are observed, a setting that has received substantial attention in the causal inference literature. In particular, we will make the widely used assumption of conditional ignorability:
We assume that we are given data that consist of samples of the tuple of random variables $(X, D, Y)$, that satisfy the basic consistency property
Without loss of generality, the data generating process obeys the regression equations:
where $U,V$ are noise variables. Note that when the outcome $Y$ is also binary, then the non-parametric functions $g_0$ and $m_0$, as well as the marginal probability law of the covariates $X$, fully determine the likelihood of the observed data.
Under conditional ignorability, consistency and the overlap assumption that both treatment values are probable conditional on $X$, i.e., $m_0(X)\in [c, 1-c]$ almost surely, for some $c>0$, it is well known that the ATE is identified by the statistical estimands:
This is the no-shift specialization ($Q_0=P_{0,Z}$) of (ref), with squared-loss regression for $\gamma$.
If we have access to a nuisance estimate $\hat{g}$, a straightforward approach is to plug it into (ref) and replace the expectation with a sample average. However, this approach makes the estimation accuracy of the target parameter highly susceptible to errors in the outcome regression nuisance function, which can be large due to high dimensionality, regularization, and model selection. Moreover, the function spaces over which these estimators operate might be complex and do not necessarily satisfy commonly invoked Donsker conditions dudley2014uniform.
To mitigate this dependence on the outcome regression model and to lift restrictions on the nuisance estimation algorithm beyond root-mean-squared-error (RMSE) accuracy, first-order debiasing/DML uses sample splitting and an orthogonal score. For the ATE, this yields a sample-splitting variant of the well-known doubly robust estimator; see, for example, robins1995analysis,chernozhukov2018double,foster2023orthogonal: \opt{opt-arxiv}{
} \opt{opt-or}{
} where $\hat{g}, \hat{m}$ are estimates of $g_0$ and $m_0$ respectively.
Even though the $n^{1/4}$ requirement can be achieved by a broad range of machine learning methods bickel2009simultaneous,belloni2011l1,belloni2013least,chen1999improved,wager2018estimation,athey2019generalized (under assumptions), it can often be violated in practice. Even when this requirement is violated, a small modification of the arguments in chernozhukov2018double,foster2023orthogonal can be used to show that $\hat{\theta}^{\textsc{ATE}} - \theta^{\textsc{ATE}} = {\mathcal{O}}_P\left( {\epsilon}_{n,m}{\epsilon}_{n,g} + n^{-1/2}\right)$, under the assumption that \opt{opt-arxiv}{
} \opt{opt-or}{
}
The formal statement of this result is presented below. The proof is in Section (ref).
Theorem (ref) highlights an important practical benefit of the doubly robust estimator: its accuracy depends only on the root-mean-squared error (RMSE) rates of the nuisance estimators, with no explicit structural assumptions on the nuisance classes. This is in stark contrast with higher-order debiasing schemes, which can lead to improved error rates bickel1988estimating,birge1995estimation,robins2008higher,van2014higher,robins2017minimax,liu2017semiparametric,kennedy2022minimax under smoothness assumptions but no longer apply when these assumptions are violated.
To establish the matching lower bound, we restrict ourselves to the case of binary outcomes:
Given that the black-box nuisance function estimators satisfy (ref), we define the following constraint set \opt{opt-arxiv}{
} \opt{opt-or}{
} where ${\epsilon}_{n,m},{\epsilon}_{n,g} = o(1)\quad (n\to +\infty).$ Note that introducing Assumption (ref) and constraints on $P_X$ in (ref) only strengthens the lower bound that we are going to prove, since they provide additional information on the ground-truth model. Moreover, the constraints $0\leq m(x), g(d,x)\leq 1$ naturally holds due to the fact that both the treatment and outcome variables are binary.
We then define the minimax $(1-\xi)$-quantile risk of estimating $\theta^{\textsc{ATE}}$ over a function space ${\mathcal{F}}$ as $$\mathfrak{M}_{n,\xi}^{\textsc{ATE}}\left({\mathcal{F}}\right) = \inf_{\hat{\theta}:\left({\mathcal{X}}\times{\mathcal{D}}\times{\mathcal{Y}}\right)^n\mapsto3} \sup_{(m^*,g^*)\in{\mathcal{F}}} {\mathcal{Q}}_{P_{m^*,g^*},1-\xi}\left( |\hat{\theta}-\theta^{\textsc{ATE}}| \right),$$ where ${\mathcal{Q}}_{P,\gamma}(X)=\inf\left\{x\in3: P[X\leq x]\geq\gamma\right\}$ denotes the quantile function of a random variable $X$ with distribution $P$, and $P_{m^*,g^*}$ is the joint distribution of $(X,D,Y)$ which is uniquely determined by the functions $m^*$ and $g^*$. Specifically, let $\mu$ be the uniform distribution on ${\mathcal{X}}\times{\mathcal{D}}\times{\mathcal{Y}}=[0,1]^K\times\{0,1\}\times\{0,1\}$, then the density $p_{m^*,g^*}=\textup{d} P_{m^*,g^*}/{\textup{d}\mu}$ can be expressed as $p_{m^*,g^*}(x,d,y) = m^*(x)^d(1-m^*(x))^{1-d}g^*(d,x)^y(1-g^*(d,x))^{1-y}.$
By definition, $\mathfrak{M}_{n,\xi}^{\textsc{ATE}}\left({\mathcal{F}}\right) \geq \rho$ would imply that for any estimator $\hat{\theta}$ of ATE, there must exist some $(m^*,g^*)\in{\mathcal{F}}$, such that under the induced data distribution, the probability of $\hat{\theta}$ having estimation error $\geq\rho$ is at least $1-\xi$. This provides a stronger form of lower bound compared with the minimax expected risk defined in balakrishnan2023fundamental, in the sense that the lower bound $\mathfrak{M}_{n,\xi}^{\textsc{ATE}}\left({\mathcal{F}}\right) \geq \rho$ implies a lower bound $(1-\xi)\rho$ of the minimax expected risk, but the converse does not necessarily hold.
The main objective of this section is to derive lower bounds for $\mathfrak{M}_{n,\xi}^{\textsc{ATE}}\left({\mathcal{M}}_1(\hat{P};{\epsilon}_{n,m},{\epsilon}_{n,g})\right)$ in terms of ${\epsilon}_{n,m},{\epsilon}_{n,g}$ and $n$. To derive our lower bound, we also need to assume that the estimators $\hat{m}(x): [0,1]^K\mapsto[0,1]$ and $\hat{g}(d,x):\{0,1\}\times[0,1]^K\mapsto[0,1]$ are bounded away from $0$ and $1$.
The assumption that $c \leq \hat{m}(x)\leq 1-c$ is common in deriving upper bounds for the error induced by debiased estimators. On the other hand, the assumption that $c\leq \hat{g}(d,x)\leq 1-c$ is typically not needed for deriving upper bounds, but it is also made in prior works for proving lower bounds of estimating the expected conditional covariance $\mathbb{E}\left[\mathrm{Cov}(D,Y)\mid X \right]$ robins2009semiparametric,balakrishnan2023fundamental.
Now we are ready to state our main results for ATE.
The proof can be found in Section (ref). As discussed in Section (ref), it relies on a fundamentally different construction of fuzzy hypotheses compared with the lower bound proof of ECC in balakrishnan2023fundamental.
In this section, we present the generic debiased estimator and its error bound (Theorem (ref)) in the general setting described in Section (ref). We then isolate the special “affine-score” case, in which the quadratic term ${\epsilon}_{N,\gamma}^2$ disappears and the error becomes purely doubly robust. This affine case coincides with the mixed bias property discussed in Section (ref) and eventually leads to a different optimal error rate, as we will see in Theorem (ref).
Assume that the linear functional $\gamma\mapsto \mathbb{E}_{Q}[m_1(Z,\gamma)]$ is continuous on $L^2(P_Z)$. Equivalently, there exists a function $\nu_m(\cdot;P,Q)\in L^2(P_Z)$ such that for any $\gamma\in L^2(P_Z)$,
We think of $\nu_m(\cdot;P,Q)$ as the “cross-population Riesz weight” representing the linear functional $\gamma\mapsto \mathbb{E}_{Q}[m_1(Z,\gamma)]$ under the $L^2(P_Z)$ inner product. This identity is the direct analogue of the key representer condition used in the DML literature chernozhukov2018double and still works in the presence of covariate shift chernozhukov2023automatic,
\paragraph{Generalized regression for $\gamma$ and the score $\rho$.} The nuisance $\gamma(\cdot;P)$ is defined via generalized regression, i.e.\ as a pointwise minimizer of an expected loss (ref) under the training law $P$. By first-order optimality,
where the score $\rho$ is the derivative of the loss in the regression direction, $\rho(o,\gamma)=\frac{\textup{d}}{\textup{d} a}\ell(o,\gamma+a)\big|_{a=0}$.
\paragraph{The weighted Riesz identity and the auxiliary nuisance $\alpha$.} Assuming the derivative $\nu_\rho(z;P)$ defined below exists and is nonzero, we define the auxiliary nuisance \opt{opt-arxiv}{
} \opt{opt-or}{
} The definition of $\alpha$ is chosen so that the first-order sensitivity of the debiased estimator to $\gamma$-perturbations cancels. Note that, unlike in the single-distribution setting, $\alpha(\cdot;P,Q)$ depends on both the training law $P$ (through $\nu_\rho$) and the target law $Q$ (through $\nu_m$).
\paragraph{The first-order debiased estimator.} Given black-box estimators $\hat{\gamma},\hat{\alpha}$ for $\gamma(\cdot;P_0)$ and $\alpha(\cdot;P_0,Q_0)$, the (debiased / orthogonal) estimator is
In practice $\hat\gamma,\hat\alpha$ are obtained by sample splitting / cross-fitting; we suppress these details since our focus is on the error scaling in $({\epsilon}_{N,\gamma},{\epsilon}_{N,\alpha})$. The following theorem provides an upper bound on this scaling and the proof can be found in Section (ref).
Theorem (ref) shows that first-order debiasing yields a structure-agnostic error bound: up to the sampling term $N^{-1/2}$, the dominant contribution is either the product ${\epsilon}_{N,\gamma}{\epsilon}_{N,\alpha}$ (the doubly robust rate) or, in the presence of curvature in the score, the additional ${\epsilon}_{N,\gamma}^2$ term. Our main theorems show that these rates are not artifacts of the proof technique. Rather, they are actually minimax optimal in terms of the nuisance estimation errors.
Theorem (ref) distinguishes two regimes: an affine-score regime with doubly robust rate ${\epsilon}_{N,\gamma}{\epsilon}_{N,\alpha}$, and a general regime with an extra ${\epsilon}_{N,\gamma}^2$ term. We now explain the structural reason behind the affine-score regime. In this case the target functional admits a second linear representation in terms of the auxiliary nuisance $\alpha$. This is the mixed bias property of rotnitzky2021characterization, and it is exactly the condition under which the quadratic term disappears in both upper and lower bounds.
Suppose that $\rho(o,\gamma)$ is affine in $\gamma$, i.e.\ there exist measurable functions $\rho_0,\rho_1:{\mathcal{O}}\to3$ such that
Then the conditional first-order condition (ref) becomes $\mathbb{E}_P[\rho_0(O)\mid Z] + \mathbb{E}_P[\rho_1(O)\mid Z]\,\gamma(Z;P)=0$, hence \[ \gamma(z;P)=-\frac{\mathbb{E}_P[\rho_0(O)\mid Z=z]}{\mathbb{E}_P[\rho_1(O)\mid Z=z]} \qquad\text{and}\qquad \nu_{\rho}(z;P)=\mathbb{E}_P[\rho_1(O)\mid Z=z], \] provided the denominator is nonzero.
Since $\alpha(z;P,Q)=-\nu_m(z;P,Q)/\nu_{\rho}(z;P)$, we have $\nu_m(z;P,Q)=-\alpha(z;P,Q)\nu_{\rho}(z;P)$ and therefore
where the last step uses that $\alpha(Z;P,Q)$ is $Z$-measurable. Thus $\chi(P,Q)$ can also be written as the expectation of a linear functional of $\alpha$ (under the training law):
This is the mixed bias property. In this regime the score has zero curvature in the $\gamma$ direction, which is why the ${\epsilon}_{N,\gamma}^2$ term vanishes in Theorem (ref). Our lower-bound Theorem (ref) shows that the remaining doubly robust rate ${\epsilon}_{N,\gamma}{\epsilon}_{N,\alpha}$ is minimax optimal.
Our goal is to understand the best possible (minimax) accuracy for estimating a semiparametric functional (ref) when the underlying data-generating mechanisms are only partially identified through first-stage nuisance estimates. As in the ATE analysis, we work in a deliberately structure-agnostic fashion: we treat the first-stage learners as black boxes, and we quantify their quality only through $L^2$-type error tolerances.
In this section we formalize the general lower bound framework by specifying (i) the risk criterion and the data-generating experiment, (ii) the uncertainty set of distribution pairs compatible with given nuisance-error tolerances, and (iii) the regularity and curvature conditions under which our lower-bound constructions operate.
\paragraph{Target estimand.} Recall that we consider semiparametric functionals of the form
where $P$ denotes a training distribution for a generic observation $O=(Z,W)\in{\mathcal{O}}$ and $Q$ denotes a (possibly different) target distribution for covariates $Z\in{\mathcal{Z}}$.\footnote{Throughout this section we take ${\mathcal{Z}}={\mathcal{Z}}_1\times{\mathcal{Z}}_2$ and ${\mathcal{O}}={\mathcal{Z}}\times{\mathcal{W}}$ as in Assumption (ref).} We emphasize that while the experiment provides separate samples from $P$ and $Q$, the underlying model class ${\mathcal{P}}_0$ may impose coupling constraints between them. Important special cases include: (i) no covariate shift, where $Q$ equals the $Z$-marginal of $P$ (often written informally as “$Q=P$”), and (ii) selection/conditioning operators such as ATT, where $Q$ is a conditional distribution derived from $P$, as shown in Example (ref). Throughout, we work under the simplifying assumption that the training and target sample sizes are equal, and we write ${\epsilon}_{N,\gamma}$ and ${\epsilon}_{N,\alpha}$ for the nuisance-error tolerances of $\gamma$ and $\alpha$, respectively.
\paragraph{Risk criterion.} For $\xi\in(0,1)$ and any collection $\mathcal{P}$ of candidate pairs $(P,Q)$, we define the minimax $(1-\xi)$-quantile risk as
where $Q_{R,1-\xi}(\cdot)$ denotes the $(1-\xi)$-quantile under $R$. Quantile risk avoids imposing tail assumptions on $\hat\chi-\chi(P,Q)$ and is convenient for the fuzzy-hypothesis arguments used in the proofs. We use the same letter $N$ for both the training and the target sample sizes.
As in our upper bound analysis, we treat first-stage nuisance estimators as black boxes and quantify their quality only through $L^2$-type error bounds. Let $\hat\gamma:{\mathcal{Z}}\to\mathbb{R}$ and $\hat\alpha:{\mathcal{Z}}\to\mathbb{R}$ denote (possibly data-dependent) estimators of the nuisance functions $\gamma(\cdot;P)$ and $\alpha(\cdot;P,Q)$. Given error tolerances ${\epsilon}_{N,\gamma}$ and ${\epsilon}_{N,\alpha}$, define the (data-dependent) collection of admissible distributional pairs by
where $P_Z$ is the $Z$-marginal of the training law $P$.
As before, we lower bound the minimax risk by anchoring at a single feasible pair. In general, the first-stage nuisance estimates $\hat h$ (think $(\hat\gamma,\hat\alpha)$) need not be exactly induced by any feasible model pair. Lemma (ref) shows that for lower bounds it is enough to work in an anchored neighborhood around a feasible pair $(\hat P,\hat Q)$ whose induced nuisances are within the same error tolerances. This reduction is illustrated in Figure (ref).
In what follows we fix such an anchor pair $(\hat P,\hat Q)$ and focus on the risk over the anchored class ${\mathcal{M}}((\hat P,\hat Q);{\epsilon}_{N,\gamma},{\epsilon}_{N,\alpha})$. For notational simplicity we also set \[ \hat\gamma(\cdot):=\gamma(\cdot;\hat P), \qquad \hat\alpha(\cdot):=\alpha(\cdot;\hat P,\hat Q), \] so that the nuisance constraints are centered at the anchor.
We now state the main conditions needed for establishing our general lower bounds.
Assumption (ref) requires that the anchor training and target laws admit densities (with respect to the chosen dominating measures) that are uniformly bounded above and away from zero.
Assumption (ref) requires that $({\mathcal{Z}}_1,\mu_{Z_1})$ admits enough “degrees of freedom”. A sufficient condition in our applications is that ${\mathcal{Z}}_1\subset\mathbb{R}^d$ and $\mu_{Z_1}$ has a density on a set with non-empty interior (e.g., one may take $f_k$ as coordinate monomials up to the required order).
Assumption (ref)(1) formalizes that the training nuisance $\gamma$ is a “conditional” functional: for each fixed $Z_1=z_1$, the function $z_2\mapsto \gamma(z_1,z_2;P)$ depends on $P$ only through the slice $p_{z_1}(z_2,w)=p(z_1,z_2,w)$. The same is true for $\alpha$, except that in the covariate-shift setting $\alpha$ may also depend on the corresponding target slice $q_{z_1}(z_2)=q(z_1,z_2)$. This setup covers many familiar nuisances (e.g., outcome regressions and propensities) that are computed by conditioning on part of the covariates.
The role of ${\mathcal{M}}$ is to encode any regularity restrictions needed to make the nuisances well-defined (e.g., overlap or boundedness of denominators). Importantly, ${\mathcal{M}}$ is a local constraint: it must hold slice-by-slice in $z_1$.
We now formally define distributional perturbations that are allowed in our setting, where the model class ${\mathcal{P}}_0$ may impose coupling constraints between the training law $P$ and the target law $Q$ (e.g.\ $Q=P_Z$ in the no-shift case, or conditioning/selection operators such as the ATT in Example (ref)). A training perturbation $G$ is a finite signed measure on ${\mathcal{O}}$ with $G({\mathcal{O}})=0$ and $G\ll \mu$, and we write $P+tG$ for the signed measure with density $p+t g$ with respect to $\mu$ whenever $g=\textup{d} G/\textup{d}\mu$ exists. A target perturbation $K$ is a finite signed measure on ${\mathcal{Z}}$ with $K({\mathcal{Z}})=0$ and $K\ll \mu_Z$, and we write $Q+tK$ analogously.
In coupled models, it is generally not possible to perturb $P$ while keeping $Q$ fixed. Accordingly, all directional derivatives in the lower bound will be taken along feasible joint perturbations $(G,K)$.
Assumption (ref) is a local smoothness condition: in a neighborhood of the anchor pair, the maps $P\mapsto \gamma(\cdot;P)$ and $(P,Q)\mapsto \alpha(\cdot;P,Q)$ admit second-order expansions along feasible perturbation paths, with remainders that are uniformly controlled. The same type of control is imposed directly on the target functional $\chi(P,Q)$. Intuitively, this means that small distributional changes lead to small and predictable changes in the nuisances and the estimand (up to quadratic error), rather than producing discontinuous jumps. Note that $\gamma$ depends only on the training law, while $\alpha$ may depend on both the training and target laws through the cross-population Riesz object.
We finally state the key assumption needed for our main results.
Assumption (ref) requires the existence, at the anchor, of feasible perturbation directions that keep the key nuisances fixed along small paths (invariance), while the target functional still exhibits a nonzero mixed second-order response when combining these directions (non-degenerate curvature). In addition, it posits the existence of a feasible direction along which the estimand varies to first order.
The following proposition shows that this assumption is naturally satisfied in a canonical non-shift setting. The proof can be found in Section (ref).
Finally, we define the anchored candidate class we will lower bound.
We now state the minimax lower bounds for the covariate shift functional (ref) over the anchored, structure-agnostic uncertainty set (ref). Recall that we observe an i.i.d.\ training sample of size $N$ from $P$ and an independent i.i.d.\ target sample of size $N$ from $Q$.
Compared with Theorem (ref), the lower bound in Theorem (ref) contains the additional term ${\epsilon}_{N,\gamma}^2$. This term is the price of curvature: it is driven by the second-order response of the functional along $\gamma$-directions that are compatible with the nuisance-error constraints. When $\rho$ is affine in $\gamma$, this curvature vanishes and one necessarily has $\chi''(\hat{P})[(H_0,L_0),(H_0,L_0)]=0$ (Proposition (ref)), which is why the affine-score regime is covered separately by Theorem (ref). Theorem (ref) matches the generic upper bound in Theorem (ref) and shows that, in the absence of the mixed-bias structure, achieving pure double robustness is fundamentally impossible in a structure-agnostic sense.
\paragraph{Overview.} The proof has two main components: (i) a fuzzy-hypothesis lower bound that reduces minimax risk to constructing a family of local alternatives with small pairwise divergences, and (ii) a second-order Taylor analysis showing that the parameter separation between null and alternatives scales as ${\epsilon}_{N,\gamma}{\epsilon}_{N,\alpha}$ (and, in the non-affine case, ${\epsilon}_{N,\gamma}^2$) while staying inside the uncertainty set.
\paragraph{Method of fuzzy hypotheses.} We use the method of fuzzy hypotheses robins2009semiparametric,kennedy2022minimax,balakrishnan2023fundamental. One constructs a null distribution (the anchor $\hat P$) and a carefully designed mixture of alternatives $\{P_{\lambda}\}$ such that:
\paragraph{Two-step perturbations and second-order separation.} The alternatives $P_{\lambda}$ are built as two-step local perturbations of $\hat P$. The first step moves along an invariant direction for either $\gamma$ or $\alpha$ (Assumption (ref)(1)), so that the corresponding nuisance remains exactly unchanged and the alternative stays inside the uncertainty set. The second step is chosen to (i) satisfy the remaining nuisance-error constraint and (ii) create a nontrivial second-order change in $\chi$ (Assumption (ref)(2)). The sizes of the two steps are calibrated as \[ \text{(first step)}\asymp \max\{{\epsilon}_{N,\gamma},{\epsilon}_{N,\alpha}\}, \qquad \text{(second step)}\asymp \min\{{\epsilon}_{N,\gamma},{\epsilon}_{N,\alpha}\}, \] so that the resulting parameter shift is of order ${\epsilon}_{N,\gamma}{\epsilon}_{N,\alpha}$ when the score is affine.
\paragraph{Why the ${\epsilon}_{N,\gamma}^2$ term appears in the non-affine case?} If $\rho$ is not affine, the conditional score has curvature captured by $\upsilon_{\rho}$ in Assumption (ref). In this case, even when we use a $\gamma$-invariant first step, the Taylor expansion of $\chi(P_{\lambda})$ can contain a pure ${\epsilon}_{N,\gamma}^2$ term (formalized via the condition $\chi''(\hat P)[H_0,H_0]\neq 0$). This is precisely the mechanism behind Theorem (ref).
\paragraph{The role of the mixed-bias property.} When $\rho$ is affine, the curvature term disappears and one can symmetrically control the second-order expansion so that the leading term is always ${\epsilon}_{N,\gamma}{\epsilon}_{N,\alpha}$. This corresponds to the mixed-bias representation discussed in Section (ref) and yields Theorem (ref).
In this section, we revisit several widely studied structural parameter estimation problems in the literature, and deduce structure-agnostic lower bounds as corollaries of Theorem (ref) and Theorem (ref).
We first show how Theorem (ref) can be derived as a special case of Theorem (ref). Recall that in the ATE case, we are given observational data $\{(X_i,D_i,Y_i)\}_{i=1}^n$, where $X_i\in3^K$ is the covariate, $D_i$ is a binary treatment variable and $Y_i=Y_i(D_i)$ is the corresponding binary outcome. The covariate space ${\mathcal{X}} = \mathrm{supp}(X)$ can be either continuous or discrete. Let $P_0$ be the ground-truth observation distribution, recall that under the conditional ignorability assumption, the ATE is identified as
Now we show how Theorem (ref) directly implies a lower bound for structure-agnostic estimation of ATE. Let the base measure $\mu$ be the uniform distribution on ${\mathcal{X}}\times{\mathcal{D}}\times{\mathcal{Y}}$, ${\mathcal{Z}} = {\mathcal{Z}}_1 = {\mathcal{X}}\times{\mathcal{D}}$, ${\mathcal{W}} = {\mathcal{Y}}$. For any distribution $P\ll \mu$, its density can be written as $p(x,d,y) := p_X(x) \pi(x;P)^d \big(1-\pi(x;P)\big)^{1-d} g(d,x;P)^y \big(1-g(d,x;P)\big)^{1-y}, $ where $\pi(x;P)=\mathbb{E}_P[D\mid X=x]$ and $g(d,x;P)=\mathbb{E}_P[Y\mid X=x, D=d]$ are the nuisance functions under the distribution $P$. We then have the following theorem:
The proof is in Section (ref).
Theorem (ref) states that the minimax structure-agnostic rate for such type of distributions is lower-bounded by ${\epsilon}_{n,\gamma}{\epsilon}_{n,\alpha}+1/{\sqrt{n}}$ up to a constant factor. Compared with Theorem (ref), the only difference is that here, the nuisances are chosen as $\gamma$ and $\alpha$ rather than $\gamma$ and $\pi$. However, they are equivalent up to a constant factor depending only on $c$ since \opt{opt-arxiv}{
} \opt{opt-or}{$\|\alpha(D,X;\hat{P})-\alpha(D,X;P)\|_{P,2}^2 = \mathbb{E}\big[ \big(\frac{1}{\pi(X;P)\pi(X;\hat{P})^2}+\frac{1}{(1-\pi(X;P))(1-\pi(X;\hat{P}))^2}\big) \big(\pi(X;P)-\pi(X;\hat{P})\big)^2 \big]$} so that under Assumption (2) in Theorem (ref), we have $\sqrt{2}\|\pi(X;P)-\pi(X;\hat{P})\|_{P,2}\leq \|\alpha(D,X;\hat{P})-\alpha(D,X;P)\|_{P,2} \leq \sqrt{2/c^{3}}\|\pi(X;P)-\pi(X;\hat{P})\|_{P,2}$. Hence we reproduce the lower bound for ATE directly by applying Theorem (ref).
We next derive a structure-agnostic lower bound for the average treatment effect on the treated (ATT). Let $P_0$ denote the observational distribution of $O=(X,D,Y)$, where $X\in{\mathcal{X}}$ is a vector of covariates, $D\in\{0,1\}$ is a binary treatment indicator, and $Y=Y(D)\in\{0,1\}$ is the observed outcome. Under conditional ignorability and overlap, the ATT is identified by
To connect (ref) to our general covariate shift functional (ref), define the training law $P$ as the joint law of $(X,D,Y)$ and define the target law $Q$ as the distribution of $X$ among treated units, i.e.\ $Q(\cdot)=P(X\in\cdot\mid D=1)$.\footnote{Equivalently, one may allow $Q$ to be any covariate distribution of interest and interpret $\mathbb{E}_{P_0}[g_0(X)\mid D=1]$ as $\mathbb{E}_{X\sim Q}[g_0(X)]$; the coupled choice $Q=P(\cdot\mid D=1)$ is the canonical ATT instance.} Set $Z=X$ and $W=(D,Y)$ so that $O=(Z,W)$. Let $\gamma(\cdot;P)$ be the control regression $x\mapsto g_0(x)$, which can be characterized as the unique solution to the conditional moment restriction
Define
so that $\theta_{\textsc{ATT}}(P_0)=\mathbb{E}_{P_0}[Y\mid D=1]-\chi_{\textsc{ATT}}(P_0,Q_0)$ with $Q_0=P_0(\cdot\mid D=1)$. Since $\mathbb{E}_{P_0}[Y\mid D=1]$ is a regular (parametric-rate) functional of $P_0$, the structure-agnostic difficulty of ATT is governed by $\chi_{\textsc{ATT}}$.
The proof is in Section (ref).
When the treatment variable is continuous, the weighted average derivative (WAD) is a commonly considered parameter of interest in econometrics with applications in index models and demand analysis hardle1991empirical,powell1989semiparametric,newey1993efficiency,imbens2009identification; see also chernozhukov2021automatic and rotnitzky2021characterization for formal definitions. WAD can naturally be interpreted as the continuous version of the ATE. Given observational data $\{(X_i,D_i,Y_i)\}_{i=1}^n$ where $X_i\in3^K$ is the covariate, $D_i$ is a real-valued treatment variable and $Y_i=Y_i(D_i)$ is the corresponding binary outcome. Define $g(x,d;P) = \mathbb{E}_P[Y\mid X=x, D=d]$. Suppose that $g$ is ${\mathcal{C}}^1$ in $d$, then we are interested in estimating
where $\omega$ is a known probability density function (PDF). Assuming that $\omega$ is continuously differentiable and has support on $(0,1)$, integration by parts implies that \opt{opt-arxiv}{
} \opt{opt-or}{
} where $s(u)=-\omega(u)^{-1}\omega'(u)$ and $U$ is a random variable independent of $O=(X,D,Y)$. The following theorem provides a lower bound for structure-agnostic estimation of WAD.
The proof is in Section (ref).
Assume that $D\in[0,1]$. We consider the average policy effect as in stock1989nonparametric: \[ \chi_{\textsc{APE}}(P)=\mathbb{E}\big[g\{X,\tau(D);P\}-g(X,D;P)\big], \] where $\tau:[0,1]\to[0,1]$ is a known counterfactual transformation. Throughout this example we assume that $\tau$ is a $C^1$-bijection of $[0,1]$ onto itself and that there exist constants $0<\underline\tau\le \overline\tau<\infty$ such that \[ \underline\tau\le |\tau'(d)|\le \overline\tau,\qquad \forall d\in[0,1], \] so that $\tau^{-1}$ is well-defined and Lipschitz. This estimand fits our framework by taking $Z_1=X$, $Z_2=D$, $W=Y$ and \[ m_1(o,h)=h\{x,\tau(d)\}-h(x,d),\qquad \rho(o,\gamma)=y-\gamma(x,d). \] The associated Riesz representer depends on the conditional density of $\tau(D)$ given $X$, which can be obtained by a change of variables.
The proof is in Section (ref).
We consider the DGP given in (ref), and the goal is to estimate the expected conditional covariance (ECC), which is defined as
In robins2009semiparametric, the authors derive minimax rates for estimating ECC under Holder-smoothness assumptions on the nuisance functions $m,g$ and the CATE function. balakrishnan2023fundamental considers a structure-agnostic setting as our paper and shows that the minimax rate scales as ${\epsilon}_{n,\gamma}{\epsilon}_{n,\alpha}+1/{\sqrt{n}}$. It is worth noticing that this rate applies to the fully nonparametric regression model (ref). One may wonder, however, if this rate is still optimal in a partial linear outcome model, i.e., if one additionally assumes that the treatment effect is constant in $X$, namely
Let \[ g(x;P):=\mathbb P_P(T=1\mid X=x), \qquad q(x;P):=\mathbb E_P[Y\mid X=x]. \] Then the ECC functional can be written as the residual-covariance numerator
The next theorem shows that even under the PLM restriction --- which is a strict submodel of (ref) --- one cannot improve on the doubly robust rate in a structure-agnostic neighborhood. In this sense, it is strictly stronger than the ECC lower bound of balakrishnan2023fundamental, which do not impose the partial linear assumption.
Fix an anchor PLM distribution $\hat P$ such that $X\sim\mathrm{Unif}([0,1]^K)$ and $T,Y\in\{0,1\}$. For radii ${\epsilon}_{n,g},{\epsilon}_{n,q}>0$, define the PLM-restricted anchored neighborhood \[ {\mathcal{M}}_{\mathrm{PLM}}(\hat P;{\epsilon}_{n,g},{\epsilon}_{n,q}) := \Big\{P\in\mathcal P_{\mathrm{PLM}}:\ \|g(X;P)-g(X;\hat P)\|_{P,2}\le {\epsilon}_{n,g},\ \|q(X;P)-q(X;\hat P)\|_{P,2}\le {\epsilon}_{n,q}\Big\}, \] where $\mathcal P_{\mathrm{PLM}}$ denotes the set of all distributions satisfying (ref) with $X\sim\mathrm{Unif}([0,1]^K)$.
The proof is in Section (ref). It constructs a PLM-preserving perturbation family directly at the density level (starting from $\hat p(x,t,y)$), verifies the “perturbation invariance” and nondegenerate mixed second derivative conditions of Assumption (ref), and then applies Theorem (ref).
Let $\{O_i\}_{i=1}^n$ be i.i.d. observations, where $O_i=(X_i,Y_i)$, $X_i\in{\mathcal{X}}$ and $Y_i\in{\mathcal{Y}}=\{0,1\}$. (For the lower bound it suffices to consider the binary-outcome case.) Let $F_1,F_2$ be two known distributions on ${\mathcal{X}}$ with densities $f_1,f_2$ w.r.t. $\mu_{{\mathcal{X}}}$. Consider the distribution-shift estimand \[ \chi_{\textsc{DS}}(P)=\int_{{\mathcal{X}}}\gamma(x;P)\,\textup{d}(F_2-F_1)(x),\qquad \gamma(x;P)=\mathbb{E}_P[Y\mid X=x]. \] This estimand fits our framework by taking $Z_1=X$, $Z_2=\varnothing$, $W=Y$, with \[ m_1(o,h)=\int_{{\mathcal{X}}}h(x)\{f_2(x)-f_1(x)\}\,\textup{d}\mu_{{\mathcal{X}}}(x),\qquad \rho(o,\gamma)=y-\gamma(x). \]
The proof is in Section (ref).
Let $O=(X,D,Y)\in {\mathcal{O}}:={\mathcal{X}}\times\{0,1\}\times\{0,1\}$ where $X\in{\mathcal{X}}:=[0,1]^K$, $D\in\{0,1\}$ is a binary treatment indicator, and $Y\in\{0,1\}$ is a binary response. We set $Z_1:=X$, $Z_2:=D$, $W:=Y$, and $Z:=(Z_1,Z_2)=(X,D)$. Let \[ g(d,x;P):=\mathbb{E}_P[Y\mid D=d,X=x],\qquad \pi(x;P):=\mathbb{E}_P[D\mid X=x], \] and define the log-odds regression function \[ \gamma(d,x;P):=\log\left(\frac{g(d,x;P)}{1-g(d,x;P)}\right),\qquad (d,x)\in\{0,1\}\times{\mathcal{X}}. \] The log-odds-difference estimand is \[ \chi_{\mathrm{LOD}}(P):=\mathbb{E}_P\big[\gamma(1,X;P)-\gamma(0,X;P)\big]. \] Let $\Lambda(t):=(1+\exp(-t))^{-1}$ denote the logistic link and consider the generalized regression score \[ \rho(o,\gamma):=\frac{y-\Lambda(\gamma)}{\Lambda(\gamma)\{1-\Lambda(\gamma)\}},\qquad o=(x,d,y). \] Then $\mathbb{E}_P[\rho\{O,\gamma(Z;P)\}\mid Z]=0$ and $\Lambda\{\gamma(d,x;P)\}=g(d,x;P)$. Moreover, one can verify that $\nu_\rho(z;P)\equiv -1$ and $\upsilon_\rho(z;P)=1-2g(z;P)$, where $g(z;P):=\mathbb{E}_P[Y\mid Z=z]$. Finally, the Riesz representer of $h\mapsto\mathbb{E}_P\{h(1,X)-h(0,X)\}$ is \[ \nu_m(z;P)=\frac{d}{\pi(x;P)}-\frac{1-d}{1-\pi(x;P)}, \] and hence $\alpha(z;P)=-\nu_m(z;P)/\nu_\rho(z;P)=\nu_m(z;P)$.
The proof is in Section (ref).
Suppose that we have i.i.d. data $\{(X_i,Y_i)\}_{i=1}^n$ from some distribution $P$ where $X_i\in{\mathcal{X}}\subseteq3^K$ and $Y_i\in{\mathcal{Y}}\subseteq3$. Suppose that $X_i=(X_{i,-1},X_{i,1})\in{\mathcal{X}}_{-1}\times{\mathcal{X}}_{1}$ where $X_{i,1}$ is a scalar variable of interest and $X_{i,-1}$ are the remaining control variables. For some fixed $q\in(0,1)$, let $\gamma(x;P)={\mathcal{Q}}_{q}(P(Y\mid X=x))$ be the $q$-th quantile of the conditional distribution of $Y$ given $X=x$, we are interested in estimating a weighted average of the partial derivative of $\gamma(x;P)$ in the direction of $x_1$
where $w(x)$ is a known weight function. First-order debiasing estimators of $\chi_{\textsc{EQD}}$ have been constructed in previous works chernozhukov2022automatic,sasaki2022unconditional. To present our lower bound for estimating $\chi_{\textsc{EQD}}$, we need a few more regularity assumptions:
The proof is in Section (ref).
This paper develops sharp structure-agnostic minimax lower bounds for a broad class of semiparametric functionals built from (generalized) regression nuisances. Assuming only $L^2$ error rates for black-box nuisance estimates, our main theorems identify two regimes: an affine-score / mixed-bias regime in which the optimal error is of order ${\epsilon}_{n,\gamma}{\epsilon}_{n,\alpha}+n^{-1/2}$ (matching the doubly robust rate), and a more general non-affine regime in which an additional term ${\epsilon}_{n,\gamma}^2$ is unavoidable. These lower bounds match the generic first-order debiasing upper bounds and therefore imply that these methods are unimprovable without additional modeling structure.
Technically, the lower bounds are proved via the method of fuzzy hypotheses, reducing estimation to testing between carefully constructed mixtures of local alternatives. The key new ingredient is a two-step sequential perturbation scheme that decouples feasibility (staying inside the anchored nuisance neighborhood) from separation (creating the desired second-order change in the target functional), together with a ham-sandwich-style partitioning argument that enforces exact invariances needed to “hide” perturbations from the nuisance constraints. We verify the required conditions for a collection of canonical examples, illustrating how the abstract theory specializes to concrete causal, policy, and quantile-based targets.
Several directions are suggested by these results. First, it would be valuable to broaden the class of functionals for which a unified optimality theory can be established. Second, our analysis is minimax by design; a natural next step is to disentangle the roles of approximation and stochastic errors as suggested in pmlr-v291-gu25b. Third, extending the framework beyond i.i.d.\ sampling (e.g.\ dependent data, clustering, distribution shift with weak overlap, or heavy-tailed outcomes) and connecting these lower bounds more directly to finite-sample inference remain important open problems.