Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
90,035 characters · 10 sections · 65 citation commands
Source Condition Double Robust Inference on Functionals of Inverse Problems
Many important problems in social and biomedical sciences can be formulated as the estimation of linear functionals of unknown functions that are defined as solutions to linear inverse problems. Examples include nonparametric instrumental variable (IV) regression problems newey2013nonparametric,ai2012semiparametric,ai2003efficient,chen2015sieve, missing-not-at-random problems d2010new,miao2015identification,LiMiao2022, causal inference with unmeasured confounders in the presence of proxy variables (a.k.a. proximal causal inference) miao2018a,cui2020semiparametric,deaner2018proxy,kallus2021causal, partially linear regression problems with endogenous regressors chen2021robust,bennett2022inference,ai2007estimation, and off-policy evaluation in confounded contextual bandits or partially observable Markov decision processes tennenholtz2020off,shen2022optimal,bennett2021proximal,Miao2022OffPolicy.
All these problems are encompassed by the following general statistical estimation problem: we are given data that contain samples of the random variable $W$ and our parameter of interest is defined as:
where $h \mapsto \widetilde m(W;h)$ is a known linear functional of $h$. Given two $W$-measurable variables $X$ and $Z$, the function $h_0$ is defined as a solution to a linear inverse problem:
where $h_0$ lies in a closed linear sub-space $\bar{\mathcal{H}}$ of the space $L_2(X)$ of square integrable functions of $X$, $\mathcal{T}: \bar{\mathcal{H}} \mapsto \bar{\ensuremath{{\cal Q}}}$ is a projected conditional expectation operator, i.e.,\xspace, $\mathcal{T} h=\Pi_{\bar{\ensuremath{{\cal Q}}}} \operatorname{\mathbb{E}}[h(X)\mid Z=\cdot]$, where $\bar{\ensuremath{{\cal Q}}}$ is a closed linear sub-space of $L_2(Z)$ and $\Pi_{\bar{\ensuremath{{\cal Q}}}}$ denotes the mean-squared projection on the space $\bar{\ensuremath{{\cal Q}}}$, and $r_0\in \bar{\ensuremath{{\cal Q}}}$ is the Riesz representer of another known linear functional $q \mapsto m(W;q)$, i.e.,\xspace:
The prototypical case is when $\bar{\mathcal{H}}=L_2(X), \bar{\ensuremath{{\cal Q}}}=L_2(Z)$ and the moment $m$ is of the simple form $m(W;f)=Y f(Z)$ for some $W$-measurable variable $Y$, in which case $r_0=\operatorname{\mathbb{E}}[Y\mid Z]$ and $h_0$ corresponds to the solution of a non-parametric instrumental variable (IV) regression problem, i.e.,\xspace, $\operatorname{\mathbb{E}}[Y - h(X)\mid Z]=0$.
Despite the significance of this problem, asymptotically normal inference for the parameter of interest $\theta_0$ presents considerable challenges, particularly when the nuisance function $h_0$ is weakly identified. In particular, since $h_0$ is the solution to an inverse problem, many qualitative attributes of $h_0$ could be smoothened out or distorted by the linear operator $\mathcal{T}$, to the point that we would need very many samples to recover these attributes well, or even to the point that these attributes are irrecoverable even in the limit of infinite samples, i.e.,\xspace, the function $h_0$ is not uniquely identified by (ref). These problems are typically referred to in the literature as the ill-posedness of the inverse problem carrasco2007linear,cavalier2011inverse. A large line of work assumes unique identification in the limit and further imposes quantitative bounds on measures of ill-posedness, so as to establish estimation rates for the function $h_0$. However, it is known that the uniqueness assumption is easily violated in practical scenarios newey2003instrumental,andrews2005inference,santos2012inference,kallus2021causal.
We instead focus on the minimum norm solution to the inverse problem, which removes the need for a uniquely identified $h_0$; even when the inverse problem has multiple solutions, the minimum norm solution is necessarily unique. Most importantly, under reasonable assumptions, the parameter $\theta_0$ of interest is invariant to the chosen solution of the inverse problem and therefore focusing on the minimum norm solution $h_0$ is without loss of generality. More concretely, let $a_0 \in \bar{\mathcal{H}}$ denote the Riesz representer of the linear functional $h\mapsto \operatorname{\mathbb{E}}[\widetilde{m}(W;h)]$:
Moreover, let $\mathcal{T}^*:\bar{\ensuremath{{\cal Q}}}\mapsto \bar{\mathcal{H}}$ denote the adjoint operator of $\mathcal{T}$, which corresponds to the projected conditional expectation operator $\mathcal{T}^* q=\Pi_{\bar{\mathcal{H}}}\operatorname{\mathbb{E}}[q(Z)\mid X=\cdot]$. As was already shown in prior work of severini2006some,severini2012efficiency,bennett2022inference, if the dual inverse problem $\mathcal{T}^* q = a_0$ admits any solution $q_0 \in \bar{\ensuremath{{\cal Q}}} \subseteq L_2(Z)$, then $\theta_0=\operatorname{\mathbb{E}}[\widetilde{m}(W;h_0)]$ takes the same value for any $h_0\in \bar{\mathcal{H}}$ solving (ref), irrespective of what solution we use. Intuitively, existence of a solution $q_0$ is an assumption that the Riesz representer $a_0$ lies primarily on the higher order spectrum of the eigendecomposition of the operator $\mathcal{T}$. Even though $h_0$ is not uniquely identified, the parameter $\theta_0$ is the projection of $h_0$ on the higher order spectrum, and this projection is uniquely identified.
When such a solution $q_0$ exists, the parameter $\theta_0$ also admits a doubly robust representation:
which satisfies the mixed bias property:
This formula lends itself to a natural estimation strategy: estimate $\hat{h}$ and $\hat{q}$ on a separate sample and then estimate $\theta_0$ in a plug-in manner, by taking the empirical analogue of the doubly robust representation formula:
Prior work of chernozhukov2023simple, bennett2022inference shows that this estimate is root-$n$ asymptotically normal when
and both $\|\hat{h}-h_0\|_{L_2}=o_p(1)$ and $\|\hat{q}-q_0\|_{L_2}=o_p(1)$. This observation implies that obtaining estimators with sufficient convergence guarantees for either $ \|\mathcal{T} (\hat h-h_0)\|_{L_2}\|\hat q-q_0\|_{L_2}$ or $\|\hat h-h_0\|_{L_2}\|\mathcal{T}^{*}(\hat q-q_0)\|_{L_2} $ is adequate for achieving asymptotically normal inference. Note that the estimation metric $\|\mathcal{T} (\hat{h}-h_0)\|_{L_2}=\sqrt{\operatorname{\mathbb{E}}\left[\left(\Pi_{\bar{\ensuremath{{\cal Q}}}}\operatorname{\mathbb{E}}\left[\hat{h}(X)-h_0(X)\mid Z\right]\right)^2\right]}$, which we refer to as the weak metric, can be much smaller than $\|\hat{h}-h_0\|_{L_2}=\sqrt{\operatorname{\mathbb{E}}[(\hat{h}(X)-h_0(X))^2]}$, which we refer to as the strong metric. Estimation of $h_0$ with respect to the weak metric does not typically suffer from the ill-posedness of the inverse problem that defines $h_0$. Similarly for estimating $q_0$ with respect to its corresponding weak metric $\|\mathcal{T}^{*}(\hat q-q_0)\|_{L_2}$. Thus for $\hat{\theta}$ to be root-$n$ asymptotically normal, it suffices to estimate only one of the two nuisance functions $h_0, q_0$ at a sufficiently fast rate with respect to the strong metric and the other with respect to the weak metric. Moreover, we always need to ensure consistency for both functions with respect to the strong metric, but without any rate.
\noindentMain contribution. The main goal of our work is to establish novel estimators for $\hat{h}$ and $\hat{q}$ that allow for general non-linear function spaces and which provide guarantees on the quantity $\ensuremath{{\cal B}}_n$ as a function of measures of ill-posedness of the primal and dual inverse problems that define $h_0$ and $q_0$, respectively, in the absence of unique identification.
Our main result is an ill-posedness doubly robust estimator: we will give a single estimation algorithm for $\hat{h}$ and $\hat{q}$ such that if one of the two inverse problems is sufficiently well-posed then the parameter estimate $\hat{\theta}$ is root-$n$ asymptotically normal. Crucially the estimation algorithm does not need to know which inverse problem is the well-posed one. Moreover, our estimation algorithm adapts to the degree of well-posedness of the most well-posed of the two inverse problems. For instance, as the largest of the two degrees goes to infinity, our requirements for asymptotic normality converge to the ones that correspond to the case when $h_0$ corresponds to the solution of a simple regression problem. This is the first such double robustness result, with respect to ill-posedness, in the literature.
We measure the ill/well-posedness of the inverse problems using the source condition. Unlike other measures proposed in the literature, the source condition is an appropriate measure even in the absence of unique identification. To describe the source condition, let us restrict for the moment to linear operators that admit a countable singular value decomposition
where $\sigma_1\geq\sigma_2\geq\ldots$ are the singular values; $\mathcal{T}$ and form an orthonormal basis of $\bar{\ensuremath{{\cal Q}}}$, i.e.,\xspace, $\operatorname{\mathbb{E}}[u_i(Z) u_j(Z)] = 1\{i=j\}$; and $v_i: X\to \mathbb{R}$ are the right eigenfunctions of $\mathcal{T}$ and also form an orthonormal basis of $\bar{\mathcal{H}}$, i.e.,\xspace, $\operatorname{\mathbb{E}}[v_i(X) v_j(X)]=1\{i=j\}$. The $\beta$-source condition on (ref) states that the minimum norm solution $h_0$ is primarily supported on the lower part of the spectrum of the eigendecomposition of $\mathcal{T}$:
Note this implies that for any $m$, $\sum_{i=m}^\infty \langle h_0, v_i\rangle_{L_2}^2 \lesssim \sigma_m^{2\beta}$ (note also that the minimum norm solution has zero inner product with eigenfunctions for which $\sigma_i=0$). Thus the parameter $\beta$ in the source condition controls the amount of support that $h_0$ is allowed to have on the tail of the spectrum. As $\beta$ goes to infinity, the function $h_0$ behaves as being supported only on a finite set of eigenfunctions and the ill-posedness problem vanishes. More generally, $\beta$ controls the degree of ill-posedness and larger $\beta$ means that the inverse problem is more well-posed.
The source condition is well defined even when the linear operator does not admit a singular value decomposition and, in its general form, requires the minimum norm solution $h_0$ to satisfy
Let $\beta_h, \beta_q$ denote the degree of well-posedness of the primal and dual inverse problems that define $h_0$ and $q_0$ respectively, and let $\beta = \max\{\beta_h,\beta_q\}$, denote the largest of the two degrees, i.e.,\xspace, the degree of the most well-posed of the two inverse problems. Moreover, we will impose the inductive biases that minimum norm solutions to the inverse problems and regularized variants of them belong to the smaller function classes $\mathcal{H}\subseteq \bar{\mathcal{H}}, \ensuremath{{\cal Q}}\subseteq \bar{\ensuremath{{\cal Q}}}$. Let $\delta_n$ denote the statistical complexity of appropriately defined function classes related to $\mathcal{H}, \ensuremath{{\cal Q}}$ and function spaces $\mathcal{F}, \ensuremath{{\cal G}}$ that encompass their composition with the linear operators, i.e.,\xspace, $\bar{\ensuremath{{\cal Q}}} \supseteq \mathcal{F} \supseteq \mathcal{T} \circ (h_0 - \mathcal{H})=\{\mathcal{T} (h_0 - h): h\in \mathcal{H}\}$ and $\bar{\mathcal{H}} \supseteq \ensuremath{{\cal G}} \supseteq \mathcal{T}^*\circ (q_0 - \ensuremath{{\cal Q}}) = \{\mathcal{T}^* (q_0 - q): q\in \ensuremath{{\cal Q}}\}$. We will measure statistical complexity using the well-established notion of the critical radius, defined via the means of localized Rademacher complexities of the corresponding classes.
Our main technical result is the development of an estimation algorithm for $\hat{h},\hat{q}$, such that the resulting parameter estimate $\hat{\theta}$ is root-$n$ asymptotically normal if:
Notably, as $\beta\to \infty$, we get that we require $\delta_n = o(n^{-1/4})$, which is the requirement when $h_0$ is the solution to a regression, or equivalently a conditional expectation, problem chernozhukov2017double. Moreover, for $\beta \in [1, 3]$, the requirement is $\delta_n=o(n^{-1/3})$, matching the prior work of bennett2022inference, which applied only when $\beta=1$ and this degree of well-posedness was assumed to be satisfied by the inverse problem that defines $q_0$. Finally, even for severely ill-posed problems, where $\beta\geq \epsilon>0$ for some small $\epsilon$, the requirement is $\delta_n=o(n^{-1/2 + \kappa})$ for some $\kappa>0$, which is satisfied, for instance, for VC-subgraph classes.
If we know which of the two inverse problems is more well-posed, then we show that we can further weaken the requirement to:
Notably the loss due to not knowing which inverse problem is more well posed is minimal and primarily occurs when $\beta\in [2,3]$. Importantly, there is no loss for moderately ill-posed problems when $\beta\leq 1$, which is arguably the most practically relevant case. Moreover, if we make the further assumption that the function spaces are smooth enough that the projected $L_2$ norm is related to the projected $L_\infty$ norm, i.e.,\xspace, $\|T h\|_{L_\infty}=O\left(\|Th\|_{L_2}^\gamma\right)$, and similarly for $q$, then this loss can be further ameliorated. For instance, as $\gamma\to 1$ (a property satisfied for instance by Reproducing Kernel Hilbert Spaces with an exponential eigendecay, which is the case for the Gaussian kernel), then there is no loss to not knowing which inverse function is more well-posed and the requirement is always of the even weaker form:
\noindentMain techniques. Our result is enabled by several novel contributions of independent interest. Our first goal is the development of estimation algorithms for a function defined via a linear inverse problem that satisfies a known $\beta$-source condition. Such algorithms can be applied both for the primal and for the dual inverse problems. We describe our results here in the context of the primal inverse problem $\mathcal{T} h=r_0$.
Our overall goal is an estimation algorithm that produces an estimate $\hat{h}$, such that irrespective of whether the source condition holds, it guarantees a fast convergence rate of $O(\delta_n^2)$, with respect to the weak metric. Moreover, when the source condition does hold it also guarantees a statistical rate of $O(\delta_n^{\kappa(\beta_h)})$, for some exponent $\kappa(\beta_h)$ with respect to the strong metric. Note that if we manage to construct such an estimation algorithm, then if we apply this algorithm for both the primal and the dual inverse problems, then we will be guaranteeing that $\ensuremath{{\cal B}}_n=O\left(\delta_n^{2 + \kappa(\beta)}\right)$, where $\beta=\max\{\beta_h,\beta_q\}$, which would then lead to the required condition on $\delta_n$.
Achieving a fast learning rate with respect to the weak metric has already been established in prior work of dikkala2020minimax, with the use of an adversarial estimation strategy, which was further extended in bennett2022inference to ensure consistency with respect to the strong metric, via the means of Tikhonov regularization, for the following estimator:
Moreover, prior work of LiaoLuofeng2020PENE established strong metric rates for this Tikhonov regularized adversarial estimator, for smooth function classes with neural network function approximation. However, the rate in LiaoLuofeng2020PENE is sub-optimal, does not adapt to the critical radius of arbitrary function spaces, and does not adapt to large values of $\beta$ (only to $\beta\leq 1$). Our first result is a fast strong metric rate result for the Tikhonov regularized estimator for $\beta \leq 1$. Our second result is a fast rate result for an iterated version of the Tikhonov regularized estimator, that uses prior iteration estimates to center the regularization appropriately and which adapts to large values of $\beta$, leading to strong metric rates of $O(\delta_n^2)$ as $\beta \to \infty$. These two results are novel in the literature on estimation of linear inverse problems under a source condition and are of independent interest.
Finally, we show that the desired simultaneous guarantee can be ensured via a constrained Tikhonov regularized adversarial estimator. In particular, our estimator first solves the un-regularized objective to find a solution that guarantees a fast weak metric rate. Subsequently, it solves the regularized objective within the sub-space of solutions that also achieve a small un-regularized risk, compared to the un-regularized solution. We show that this estimator simultaneously enjoys both guarantees: a weak metric rate of $\delta_n^2$, without requiring a $\beta$-source condition and a strong metric rate of $\delta_n^{2\max\left\{\frac{\min(\beta,1)}{1+\min(\beta,1)}, \frac{\beta-1}{\beta+1}\right\}}$, when the $\beta$-source condition holds. This theorem enables our main double robustness result.
We first discuss related work specifically focusing on nonparametric IV regression functions, i.e.,\xspace, solutions to $\operatorname{\mathbb{E}}[Y-h (X)\mid Z]=0$. Later, we delve into related work that focuses on estimating functionals $\theta_0$ of nuisance functions that are defined as linear inverse problems, including nonparametric IV regression functions.
\noindentNonparametric IV regression. Instrumental variable estimation has garnered significant interest as a particular subset within the realm of inverse problems, as exemplified by the comprehensive investigations carrasco2007linear,cavalier2011inverse, newey2013nonparametric. Nonparametric instrumental variable estimation encounters considerable challenges due to its ill-posed nature, even when the operator $\mathcal{T}$ and response $r_0$ are known. The ill-posedness is often characterized by the presence of one or more of the following aspects: (1) the absence of solutions, (2) the existence of multiple solutions, and (3) the discontinuity of the pseudo-inverse of $\mathcal{T}$. To tackle these challenges, various regularization techniques have been proposed, such as imposing compactness on the solution space newey2003instrumental and employing Tikhonov regularization carrasco2007linear. In practical settings where $\mathcal{T}$ and $r_0$ are unknown, a range of estimators has been proposed in the literature, including series-based estimators ai2003efficient,hall2005nonparametric,blundell2007semi,chen2011rate,darolles2011nonparametric,chen2012estimation,florens2011identification,chen2021robust, kernel-based estimators hall2005nonparametric,horowitz2007asymptotic, RKHS-based estimators singh2019kernel,muandet2020dual,bennett2020variational and high-dimensional linear estimators under sparsity gautier2011high,gautier2018high,gautier2022fast.
Recently, there has been an increasing interest in applying general function approximation techniques, such as deep neural networks and random forests, to instrumental variable problems in a unified manner dikkala2020minimax,lewis2018adversarial,bennett2019deep,zhang2020maximum,dikkala2020minimax,bennett2020variational. However, the guarantees of most current methods remain unclear when solutions are not unique. Notable exceptions include the works of LiaoLuofeng2020PENE,bennett2023minimax,which provide finite-sample convergence rate guarantees even when solutions may not be unique.
In the closely related work of LiaoLuofeng2020PENE, they establish $L_2$ convergence by connecting minimax optimization with Tikhonov regularization under the assumption of the source condition. Specifically, when the number of iterations is limited to one in our method, their method coincides with ours in solving inverse problems. However, our paper presents two important contributions beyond their work. Firstly, we introduce a new iterative procedure that achieves a fast convergence rate under high-order source conditions with $\beta \geq 2$. Secondly, even when the number of iterations is limited to one, our paper achieves a faster convergence rate than theirs due to improved analysis. We note that their paper also makes its own contribution by specifically considering scenarios where function classes are neural networks and providing a theoretical analysis that takes into account the optimization procedure.
Another closely related work bennett2023minimax proposes a method that treats IV regression as a constrained optimization problem. While their method does not require the “closedness assumption,” which implies smoothness of the operator $\mathcal{T}$, compared to our work, the guaranteed convergence rate in their paper is slower than ours because their method cannot effectively exploit potentially high-order source conditions with $\beta \geq 2$. Additionally, their method does not provide any guarantees under the source condition with $\beta \geq 2$.
We note that there are a number of alternative approaches for integrating machine learning into instrumental variable estimation hartford2017deep,yu2018deep,xu2020learning,liu2020deep,kato2021learning,lu2021machine. However, to the best of our knowledge, these approaches do not offer an $L_2$ convergence rate guarantee in the absence of the assumption of uniqueness.
We are given access to a set of independent and identically distributed observations $\{X_i,Z_i,W_i\}_{i=1}^n$ drawn from the distribution of the random variables $X, Z, W$. Our ultimate objective is to estimate the parameter $\theta_0 = \operatorname{\mathbb{E}}[\widetilde{m}(W; h_0)]$, as presented in (ref), and construct a valid confidence interval around it.
Throughout this work, we assume there exists a solution to the linear inverse problem given by (ref). (Numbered assumptions are assumed to hold throughout the paper.)
Moreover, it will be convenient to express the constraints that identify $h_0$ in a combined manner, which can be derived by simple algebra from (ref) and the properties of mean-squared projections onto closed linear spaces:\footnote{For any $q\in \bar{\ensuremath{{\cal Q}}}$, $r\in L_2(Z)$ it follows from properties of projections on closed linear spaces that $\langle r, q\rangle_{L_2(Z)} = \langle \Pi_{\bar{\ensuremath{{\cal Q}}}}r, q\rangle_{L_2(Z)}$. Hence: $\operatorname{\mathbb{E}}[h_0(X) q(Z)] = \operatorname{\mathbb{E}}[\operatorname{\mathbb{E}}[h_0(X)\mid Z]\, q(Z)] = \operatorname{\mathbb{E}}[(\mathcal{T} h_0)(Z)\, q(Z)] = \operatorname{\mathbb{E}}[r_0(Z) q(Z)] = \operatorname{\mathbb{E}}[m(W;q)]$.}
In general, the solution mentioned above may not be unique. For this reason we aim to estimate the least $L_2$-norm solution, i.e.,\xspace, we define $h_0$ as
This least norm solution always exists uniquely, as shown in Lemma 1 of bennett2023minimax, and is frequently employed as a target in the literature when solutions are not unique santos2011instrumental,florens2011identification. Notably, as we elaborate in the next section, for many linear functionals, the specific choice of the solution to the linear inverse problem is irrelevant and all solutions lead to the same value for the parameter $\theta_0$.
Our setting encompasses many well-studied problems in econometrics and statistics. We present here two illustrative examples. Other examples that fall in our framework include partially linear IV and proximal causal inference models, missing-not-at-random data with shadow variables, and offline policy evaluation in partially observable MDPs.
\noindentNotation and preliminary definitions. Before delving into the main technical exposition we need to introduce some technical notation. Throughout the paper, whenever we use a generic norm of a function $\|h\|$, we will be referring to the $L_2$-norm with respect to the distribution of the input of the function, i.e.,\xspace,
We will also be using the shorthand notation $\operatorname{\mathbb{E}}_n[\cdot]$ for the empirical average, e.g., \xspace, $\operatorname{\mathbb{E}}_n[X] = \frac{1}{n}\sum_{i=1}^n X_i$. For any set $A$, we denote the closure of $A$ by $\ensuremath{\mathrm{cl}}(A)$. For any function space $\mathcal{F}$, containing functions that are uniformly and absolutely bounded by $1$, we will be using the critical radius as the measure of statistical complexity (c.f. wainwright2019high for a more detailed exposition). To define the critical radius we first define the localized Rademacher complexity:
The star hull of a function space is defined as $\ensuremath{\text{star}}(\mathcal{F})=\{\gamma f: f\in \mathcal{F}, \gamma\in [0,1]\}$. The critical radius $\delta_n$ of $\ensuremath{\text{star}}(\mathcal{F})$ is the smallest positive solution to the inequality:
Throughout we will be assuming a uniform absolute bound of $1$ for all random variables and functions. This can be lifted to any finite bound $b$ by rescaling.
We begin by noting that any linear functional $\widetilde{m}$ of $h_0$ admits a doubly robust representation:
where $q_0$ solves a dual inverse problem of the same nature as $h_0$, but with respect to functional $\widetilde{m}$ instead of $m$. More specifically:
where $\mathcal{T}^*:\bar{\ensuremath{{\cal Q}}}\mapsto \bar{\mathcal{H}}$ is the adjoint operator to $\mathcal{T}:\bar{\mathcal{H}}\mapsto \bar{\ensuremath{{\cal Q}}}$. The adjoint operator is defined as:
If $\bar{\mathcal{H}}= L_2({\mathcal X})$, then $\mathcal{T}^*$ simplifies to a conditional expectation $T^*q = \operatorname{\mathbb{E}}[q(Z)\mid X=\cdot]$. Note that the conditions that define $q_0$ can be expressed in a combined manner, similar to Equation (ref):
We assume the existence of a solution $q_0\in \bar{\ensuremath{{\cal Q}}}$ to the inverse problem defined in Equation (ref). If the inverse problem in Equation (ref) has multiple solutions, then we will again denote by $q_0$ the minimum $L_2$-norm solution to the inverse problem.
The reader might wonder why we need this assumption to estimate $\theta_0$ on top of the existence of solutions in the primal inverse problem. As was shown in severini2012efficiency, when $\bar{\mathcal{H}}=L_2(X)$ and $\bar{\ensuremath{{\cal Q}}}=L_2(Z)$, the assumption $a_0 \in \mathcal{R}(\mathcal{T}^*)$ is a necessary condition for the $\sqrt{n}$-estimability of the parameter $\theta_0$. In this sense the assumption is unavoidable for root-$n$ asymptotic normality.
Now, we are ready to state the mixed bias property of (ref).
Based on this crucial lemma, we can then apply the general machinery of Neyman orthogonality and debiased machine learning chernozhukov2017double to arrive at a corollary that shows how and when one can deduce root-$n$ consistency and asymptotic normality of the estimate $\hat{\theta}$ of $\theta_0$ presented in Equation (ref) (see also chernozhukov2023simple for a finite sample version in a slightly simpler setting):
If we want to apply the latter corollary, it remains to show how we can estimate $\hat{h},\hat{q}$ in a manner such that:
and such that $\|\hat{h}-h_0\|=o_p(1)$ and $\|\hat{q}-h_0\|=o_p(1)$. By the definition of $\mathcal{T}$, applying a tower law, a Cauchy-Schwarz inequality and orthogonality of projections on linear spaces, we have
Similarly, by the definition of $\mathcal{T}^*$ and the orthogonality of projections, we also have that:
and therefore a sufficient condition for Equation (ref) to hold is:
Note that both $\hat{h}$ and $\hat{q}$ are instances of the same estimation problem. In particular, both statistical problems can be defined as finding a function $h\in \bar{\mathcal{H}}$ that satisfies a linear inverse problem $\mathcal{T} h=r_0$, where $r_0$ is the Riesz reprsenter of a linear functional and $\mathcal{T}$ is a projected conditional expectation operator. Thus we will describe how one can solve the primal problem of estimating $\hat{h}$ and the exact same analysis can be applied to the dual problem where the operator $\mathcal{T}$ is replaced by the dual $\mathcal{T}^*$, the function space $\bar{\mathcal{H}}$ is replaced by $\bar{\ensuremath{{\cal Q}}}$, and the linear functional $m(W;q)$ is replaced by the linear functional $\widetilde{m}(W;h)$.
We will provide estimation rates as a function of inductive biases further imposed on the solutions $h_0$ and $q_0$. In particular, we will consider the following realizability assumptions:
The function classes $\mathcal{H}, \ensuremath{{\cal Q}}, \mathcal{F}, \ensuremath{{\cal G}}$, are meant to be classes of bounded statistical complexity, such as for instance norm-constrained high-dimensional linear models, Reproducing Kernel Hilbert Spaces, or neural network classes. This assumption can be relaxed to allow for approximation errors in all inclusion statements, at the expense of additive such error bounds in all our theorems. We omit such an extension for simplicity of exposition.
Given access to $n$ samples $\{X_i, Z_i, W_i\}_{i=1}^n$, our objective is to find an estimator $\hat h$ that converges to the minimum norm solution $h_0$ to Equation (ref). Our goal is to provide estimation rates with respect to both the strong and weak metrics defined below:
The weak metric is “weak” in the sense that it is smaller than the strong metric and is a pseudo metric. Specifically, $\|\mathcal{T}(\hat{h}-h_0)\|^2= \|\mathcal{T}(\hat{h}-h'_0)\|^2$ whenever $Th_0'=r_0$ even if $h'_0\neq h_0$.
Within the literature on nonparametric instrumental variable (IV) regression, a commonly employed approach focuses on optimizing empirical versions of the weak metric criterion, also known as the minimum distance criterion. This optimization process is carried out over simple hypothesis spaces denoted as $\mathcal{H}_n$ of increasing complexity. These hypothesis spaces, often referred to as sieves, aim to approximate the function $h_0$ while enforcing some uniform control over the ratio between the strong and weak metrics across the entire class, also known as the measure of ill-posedness of the inverse problems:
This measure, as indicated in previous works dikkala2020minimax,chen2012estimation, reflects the degree of ill-posedness in inverse problems. However, in cases where unique identification is absent this measure can easily diverge to infinity. If the sieves $\mathcal{H}_n$ are taken so that eventually they uniformly approximate the function space $\mathcal{H}$, then note that there will exist a different solution $h_0'\in \mathcal{H}$ to the inverse problem, that is $\epsilon$ approximated by a function $h_{\epsilon}$ in $\mathcal{H}_n$. Thus we will have that $\|h_0-h_0'\| \to \gamma > 0$ and $\|\mathcal{T}(h_0-h_{\epsilon})\|\leq \epsilon + \|\mathcal{T}(h_0-h_0')\| = \epsilon\to 0$.
An alternative approach that remains effective even in the absence of the uniqueness assumption involves imposing the $\beta$-source condition on the minimum norm solution $h_0$. Under the $\beta$-source condition, it becomes possible to achieve desirable rates on the strong metric by introducing an additional $L_2$ regularization penalty to the minimum distance criterion. This regularization technique is also known as Tikhonov regularization Carrasco2007. In this study, we adopt the latter approach.
\noindentThe Source Condition. We begin by introducing the $\beta$-source condition, which is commonly used in the literature on inverse problems Carrasco2007 and captures mathematically how strongly the function $h_0$ is identified by the data that we observed.
Essentially, this assumption is claiming
Intuitively, when the parameter $\beta$ is large, the function $h_0$ exhibits greater smoothness, in the sense that the $L_2$-inner product of $h_0$ with eigenfunctions that have smaller eigenvalues relative to the operator $\mathcal{T}$ tend to decay faster. A more concrete interpretation of the assumption when the operator $\mathcal{T}$ is compact and admits a singular value decomposition is given in (ref).
Inspired by the combined moment constraints in Equation (ref) (and similarly Equation (ref) for $q_0$), we will consider an adversarial population criterion for the estimation of $h_0$:
where $\mathcal{F} \subseteq \bar{\ensuremath{{\cal Q}}}$, as defined in Assumption (ref). By the definition of the function $h_0$, we have $\operatorname{\mathbb{E}}[m(W;f)]=\operatorname{\mathbb{E}}[h_0(X)\, f(Z)]$. Hence, the above population criterion is equivalent to:
Moreover, by an application of the tower law of expectations and the properties of projections onto the closed linear space $\bar{\ensuremath{{\cal Q}}}$ and since $\mathcal{F}\subseteq \bar{\ensuremath{{\cal Q}}}$, the criterion is also equivalent to:
Since $\mathcal{F}$ contains functions of the form $\mathcal{T} (h_0 - h)$ for any $h\in \mathcal{H}$ (with $\mathcal{H}$ as in Assumption (ref)):
and the population criterion is equivalent to the weak metric, for any $h\in \mathcal{H}$:
In other words, by minimizing the adversarial population criterion, we are essentially minimizing the weak metric distance to $h_0$.
\noindentTikhonov Regularized Adversarial Estimator. As a first step, we consider a Tikhonov regularized adversarial estimator. The regularized population criterion and the corresponding regularized population solution is defined as:
We define the Tikhonov regularized adversarial estimator, as the solution to an empirical analogue the adversarial formulation of the regularized population criterion, replacing also $\bar{\mathcal{H}}$ with $\mathcal{H}$ (as defined in Assumption (ref)):
An essential implication of the source condition is the ability to control the regularization bias caused by the Tikhonov regularization term, denoted as $h_* - h_0$, as a function of both $\beta$ and the regularization strength $\lambda$ as follows.
The first inequality is widely recognized (c.f. cavalier2011inverse). In our analysis we also incorporate the second inequality, which bounds the bias in terms of the weak metric.
Having established control over the bias, we are now poised to formally demonstrate the previously mentioned estimation rates for the Tikhonov regularized estimator presented in (ref). In the following theorem, we make use of the concept of critical radius, a standard measure for quantifying convergence in nonparametric regression wainwright2019high. In particular, when the function classes under consideration are VC classes, the critical radius is determined to be $\delta_n = O(n^{-1/2})$. It is worth noting that the critical radius can be readily computed for various function classes such as Sobolev balls, Holder balls, and sparse linear models (for detailed calculations see dikkala2020minimax,wainwright2019high,foster2019orthogonal).
Hereafter, we present the implications of the aforementioned theorem. Firstly, assumptions (b) and (c) are standard, as utilized in dikkala2020minimax,LiaoLuofeng2020PENE. In cases where these assumptions are violated, additional calculations allow us to obtain results with extra bias terms, which quantify the extent of the violation. Secondly, if we optimally choose $\lambda\sim \delta_n^{\frac{2}{1 + \min(\beta,1)}}$, so as to minimize the strong metric upper bound in Theorem (ref), then we simultaneously get the rates:
For special cases of the setting we cover here, these rates are equivalent to the current state-of-the-art rates, in terms of the weak metric as presented in dikkala2020minimax, and the strong metric as shown in bennett2023minimax. Moreover, unlike these prior works, our analysis offers simultaneously state-of-the-art guarantees for both metrics. Consequently, our analysis shows that for the one-step Tikhonov regularized adversarial estimator there is no trade-off in terms of which metric to prioritize (up to multiplicative constants). We now provide a more detailed comparison of our results with existing ones.
\noindentComparison with LiaoLuofeng2020PENE. The authors proposed the same estimator and focus on $\mathcal{H}$ and $\mathcal{F}$ that are neural networks classes. By naively following their analysis and extending it to cases involving general function classes, we obtain the following result:
Consequently, for $\beta\leq 2$, the resulting rate is of $\|\hat h-h_0 \|$ is $\delta_n^{\beta/2(1+\beta)}$, while for $\beta\geq 2$, the corresponding rate is $\delta_n^{1/3}$. As depicted in Figure (ref), this rate is slower compared to the one we obtained. The looser analysis in their work stems from their failure to fully exploit the strong convexity of the loss function (ref) with respect to $h$ in terms of the strong metric. In our analysis, we refine the bound by employing the localization technique, which entails bounding the empirical process term using the errors $\|\hat h-h_0\|$ and $\|\mathcal{T}(\hat h-h_0)\|$.
\noindentComparison with bennett2023minimax. They propose an estimator that achieves $\| \hat h-h_0 \|^2=O(\delta_n)$ when $\beta\geq 1$. Although their estimator can relax assumption (c) by replacing it with a realizability assumption of the form “there exists $w_0 \in \mathcal{H}$ such that $h_0 = \mathcal{T}^* w_0$”, it remains unclear whether their estimator can achieve $\|\mathcal{T}(\hat h-h_0)\|=O(\delta_n^2)$. Furthermore, their estimator does not provide any guarantee when $\beta < 1$.
\noindentComparison with dikkala2020minimax. They derive an estimator that achieves $\|\mathcal{T}(\hat h-h_0)\|^2=O(\delta_n^2)$. While their result does not necessitate the source condition, their estimator lacks any guarantee regarding the convergence in terms of the strong metric $\|\hat h-h_0\|$, especially when the primal inverse problem does not have a unique solution.
One limitation of the result in the previous section is its lack of adaptability to the degree of ill-posedness in the inverse problem, particularly for larger values of $\beta$ corresponding to milder problems. Ideally, as $\beta\to\infty$, it is expected that the strong metric rates would converge to $O(\delta_n^2)$. Indeed, when the operator $\mathcal{T}$ coincides with the identity operator, leading to $\beta=\infty$, nonparametric IV regression reduces to standard nonparametric regression. In such cases, achieving an $O(\delta_n^2)$ rate is well-known wainwright2019high. The observed convergence rate saturation after some small value of $\beta$ is a recognized drawback of Tikhonov regularization Carrasco2007. To address this limitation, we propose an iterated Tikhonov regularized adversarial estimator. At each iteration $t$ (commencing with $h_{*,0}=\hat{h}_0=0$) \footnote{Technically, we can set any function as $\hat{h}_0$.}, we consider the population criterion:
and the corresponding empirical criterion:
We will show that by choosing $t$ and $\lambda$ appropriately then with high probability, as long as $n$ is larger than some constant then this iterated estimator achieves a strong metric rate of $\approx \delta_n^{2 \frac{\beta}{\beta+2}}$ that adapts to large values of $\beta$, i.e.,\xspace, becomes faster as $\beta\to \infty$ and converges to $\delta_n^2$ in the limit.
The limitation observed in the previous analysis stems from the inability of the bias in \prettyref{lem:bias-tikhonov} to adapt to higher values of $\beta$. However, the iterative nature of Tikhonov regularization possesses a crucial characteristic that mitigates this issue. It is demonstrated that the regularization bias of the $t$-th population iterate, denoted as $h_{*,t} - h_0$, is significantly smaller than that of the non-iterated one, particularly when $\beta$ is large.
With such an improved bias control at hand we can prove our claimed guarantee for the iterated tikhonov regularizedd estimator. The proof is based on starkly different approach to controlling the variance part, than what we used in Theorem (ref). This alternative approach is required, so as to avoid compounding of bias terms over the $t$ iterates.
If one uses the crude bound of $\max_{k\in [t-1]} \|\mathcal{T}(h_{*,k}-h_0)\|_{L_{\infty}}^2=O(1)$, then this theorem yields strong and weak metric rates of:
If we let $\beta_t = \min\{\beta, 2t\}$ and choose a regularization strength of $\lambda\sim \delta_n^{\frac{1}{\beta_t+2}}$, then we get rates:
If we choose $t=\lceil \min\{\beta/2, \log\log(1/\delta_n)\}\rceil$ then we get a rate of:
Hence for any constant $\beta$, as $n$ grows, eventually $\log\log(1/\delta_n)\geq \beta$ and we get the rate of $O\left(\delta_n^{2 \frac{\beta}{\beta+2}}\right)$. This rate can also be achieved, even if $\beta$ is allowed to grow with $n$ in the asymptotics, as long as it grows slower than $\log\log(1/\delta_n)$. If $\delta_n \sim n^{-\alpha}$ for some $\alpha>0$, then we note that if we take $t= \lceil \min\{\beta/2, \sqrt{\log(1/\delta_n)}\}\rceil$, then $16^{\sqrt{\log(1/\delta_n)}}=o(n^{\epsilon})$ for any $\epsilon>0$, thus we again conclude the rate $\widetilde{O}\left(\delta_n^{2 \frac{\beta_t}{\beta_t+2}}\right)$. This result allows us to claim a rate of $O\left(\delta_n^{2 \frac{\beta}{\beta+2}}\right)$ even if $\beta$ is growing with $n$ in the asymptotics as long as it grows slower than $\sqrt{\log(1/\delta_n)}$, which is considerably faster then $\log\log(1/\delta_n)$.
In summary, as depicted in Figure (ref), when $\beta \geq 2$ and the sample size $n$ is sufficiently large, the iterated algorithm exhibits an improved rate in terms of the strong metric compared to the non-iterated version and the existing work LiaoLuofeng2020PENE, where the rate saturates after $\beta=2$. Notably, as $\beta$ tends to infinity, the rate of the iterated algorithm approaches the fast rate $O(\delta_n^2)$. However, such hyperparameter settings tuned for the strong metric do not yield a fast rate in terms of the weak metric, but rather a rate of $O\left(\delta_n^{2\frac{1+\beta}{2+\beta}}\right)$. Hence, there is a trade-off in the choice of estimation hyperparameters, dependent on which metric one is targeting.
As was already noted, the problem that defines $q_0$ is of exactly the same nature as the primary inverse problem that defines $h_0$. The only difference is that the moment is $\widetilde{m}$ instead of the moment $m$ that defines $h_0$, the linear operator is the adjoint $\mathcal{T}^*$ of the operator that defines $h_0$ and the roles of $X$ and $Z$ are reversed, i.e.,\xspace, $X$ becomes the “exogenous” or “conditioning” variable and $Z$ the “endogenous”. Hence, the Tikhonov Regularized adversarial estimator for $q_0$ becomes:
where $\ensuremath{{\cal Q}}$ and $\ensuremath{{\cal G}}$ are defined in Assumption (ref). We can similarly define the iterated version. Moreover, we can obtain convergence guarantees for $\hat q$ and its iterated version via \prettyref{thm:adv-l2} and \prettyref{thm:adv-l2-iter}. We will assume that both the primal and dual inverse problems satisfy a source condition for some positive $\beta$ (a very benign assumption, i.e.,\xspace, arbitrary weak source condition).
The condition for the dual problem states that $q_0$ should be primarily supported on the lower spectrum of the operator $T^*$, equivalently, $a_0$ should be primarily supported on the lower spectrum of the operator $T$ (see Appendix (ref) for a more detailed discussion).
We present sufficient conditions for guaranteeing asymptotically normal inference on $\hat{\theta}$. We will denote with $\beta=\max\{\beta_h, \beta_q\}$, to be the largest of the two source condition parameters (i.e.,\xspace, the degree of the most well-posed inverse problem). First, we examine the case where we know which of the two inverse problems satisfies the $\beta$-source condition and without loss of generality we will assume that to be dual inverse problem that defines $q_0$, while the primal inverse problem that defines $h_0$ satisfies an arbitrary weak source condition. These roles can be interchanged without difficulty. Subsequently, we consider a more challenging scenario where we do not know which one satisfies the non-vacuous source condition.
Suppose that $\beta=\beta_q \geq \beta_h > 0$. The complementary case follows identically, interchanging the roles of $h$ and $q$. Theorem (ref) implies that by choosing $\lambda \sim\delta_n^{\frac{2}{1+\min(\beta_h, 1)}}$, w.p. $1-O(\zeta)$:
Thus we guarantee fast rates of $O(\delta_n^2)$ for the weak metric and consistency with respect to the strong metric; as required by the asymptotic normality Corollary (ref). Here $\delta_n$ is an upper bound on the critical radius of the function classes $\ensuremath{\text{star}}(\mathcal{H}\cdot \mathcal{F})$, $\ensuremath{\text{star}}(m \circ \mathcal{F})$, $\ensuremath{\text{star}}(\mathcal{F})$, $\ensuremath{\text{star}}(\mathcal{H} - \mathcal{H})$. Moreover, since $q_0$ satisfies a $\beta_q$-source condition, we can apply the best of Theorem (ref) or Theorem (ref) with an optimal choice of regularization hyperparameter and number of iterations, to get:
where $\delta_n$ is an upper bound on the critical radius of the function classes $\ensuremath{\text{star}}(\ensuremath{{\cal Q}}\cdot \ensuremath{{\cal G}})$, $\ensuremath{\text{star}}(\widetilde{m} \circ \ensuremath{{\cal G}})$, $\ensuremath{\text{star}}(\ensuremath{{\cal G}})$, $\ensuremath{\text{star}}(\ensuremath{{\cal Q}} - \ensuremath{{\cal Q}})$.
Combining the two aforementioned observations, we can construct estimators $\hat{h}, \hat{q}$ for both $h_0$ and $q_0$ to be used in the context of Corollary (ref), that ensure asymptotic normality if the following rate condition is satisfied:
We can conclude an analogous statement by flipping the role of $h_0$ and $q_0$. The size requirement of the function classes varies depending on the source condition. More specifically, we need
This rate is illustrated in the blue line in (ref). For $1\leq \beta\leq 2$, it is sufficient to have $\delta_n=n^{-1/3}$, which still permits non-parametric function classes. In the extreme case where $\beta \to 0$, the condition converges to $\delta_n=n^{-1/2}$, which corresponds to a parametric rate (i.e.,\xspace, function classes are VC classes). Therefore, our theorem accommodates scenarios where both $h_0$ and $q_0$ exhibit severe ill-posedness. Conversely, in the limit as $\beta\to \infty$, the condition requires $\delta_n=n^{-1/4}$, which is the requirement typically encountered when dealing with functionals of functions defined through regression problems rather than inverse problems (c.f. chernozhukov2017double).
\noindentComparison with bennett2022inference. To provide a comparison, we examine our condition alongside recent related work by bennett2022inference. In bennett2022inference, in their study, under the assumption of $\beta=1$, they demonstrate that asymptotically normal inference is feasible as long as
The corresponding required rate for $\delta_n$ is illustrated in the blue line in (ref). This aligns with our conclusion when instantiating $\beta=1$ in (ref). However, our condition is significantly more general. Firstly, when $\beta<1$, our condition in (ref) still allows for the possibility of asymptotically normal inference. Specifically, as $\beta \to 0$, our requirement converges to $\delta_n = o(n^{-1/2})$. In contrast, bennett2022inference do not provide any guarantees when $\beta<1$. Secondly, when $\beta>1$, we can leverage the potentially larger size of the function classes as $\beta$ increases. Specifically, as $\beta \to \infty$, our requirement becomes $\delta_n^4 = o(n^{-1/4})$. This adaptability to $\beta$ is not obtained in bennett2022inference.
The current results do not guarantee double robustness in terms of source conditions. We have presented an estimator that enables asymptotically normal inference when $h_0$ satisfies the $\beta$-source condition and $q_0$ satisfies an arbitrary weak source condition. Similarly, we have established a corresponding statement in a scenario when we interchange their role. In this section, we aim to develop an estimator that achieves double robustness, allowing for asymptotically normal inference in both scenarios simultaneously.
To achieve this objective, we propose the following refined methods. For convenience, let
Then we introduce the constrained function classes:
where $\mu_{h,n}$ and $\mu_{q,n}$ are parameters that will be specified later. After defining these version spaces, we follow the same procedure as before, but now using $\widetilde \mathcal{H}$ and $\widetilde \ensuremath{{\cal Q}}$ instead of $\mathcal{H}$ and $\ensuremath{{\cal Q}}$, respectively. Here, we construct these version spaces such that with high probability: (1) all the functions in $\widetilde \mathcal{H}$ are sufficiently close to $h_0$ in terms of weak metric, even in the absence of the source condition, and (2) $h_{*} \in \widetilde \mathcal{H}$ under the $\beta$-source condition (similarly for $\widetilde \ensuremath{{\cal Q}}$). The first property plays a crucial role in achieving source double robustness, while the second property allows us to apply (ref) to $\widetilde \mathcal{H}$ instead of $\mathcal{H}$ and ensure a sufficiently fast rate in terms of strong metric under the $\beta$-source condition.
Formally, the resulting estimators with version spaces possess the following properties. An analogous conclusion can be also derived for $\hat q$. We first consider the non-iterated version.
Importantly, the first statement concerning the weak metric does not rely on the source condition. Utilizing this theorem, let us consider a scenario where either $h_0$ or $q_0$ satisfies the $\beta$-source condition, but we are uncertain about which one. In such cases, by setting $\mu_n = c (\delta^2_n + \lambda^{\min(\beta+1,2)})$, asymptotically normal inference remains feasible under the condition that
By optimally setting $\lambda\sim \delta_n^{\frac{2}{1+\min(\beta, 1)}}$, the above becomes
Moreover, since $\beta>0$, the above setting of $\lambda$ also ensures that consistency with respect to the strong metric at some arbitrarily slow rate continues to hold, even when the $\beta$-source condition does not hold for the function but only an $\epsilon$ source condition holds for an arbitrarily small $\epsilon>0$ (see also discussion in the previous section). This result is noteworthy since, in the previous section using Theorem (ref), we required prior knowledge regarding which one between $h_0$ and $q_0$ satisfies the source condition.
Next, we consider the iterated version to leverage a higher value of $\beta\geq 1$. At each iteration, we optimize over the constrained spaces ${\widetilde \mathcal{H}_t, \widetilde \ensuremath{{\cal Q}}_t}$, defined similar to $\widetilde\mathcal{H},\widetilde\ensuremath{{\cal Q}}$, but with an iterate dependent upper bound $\mu_{n,t}$ instead of $\mu_n$. By replacing $\mathcal{H}$ and $\ensuremath{{\cal Q}}$ with $\widetilde \mathcal{H}_t$ and $\widetilde \ensuremath{{\cal Q}}_t$ respectively at each iteration, we obtain estimators $\hat h_t$ and $\hat q_t$. Here is the property of the estimators.
Again, in contrast to (ref), the first statement does not impose any source condition. Using this theorem, let us consider a scenario where either $h_0$ or $q_0$ satisfies the $\beta$-source condition, but we are uncertain about which one. In this case, asymptotically normal inference is possible as long as
By optimizing $\lambda\sim \delta_n^{\frac{2}{\beta+1}}$, we can conclude that when $(\beta+1)/2 \leq t\leq c' \log\log(1/\delta_n)$, the asymptotically normal inference for $\hat{\theta}$ is possible as long as
Moreover, for $\beta\geq 1$, this setting of $\lambda$ is $o(\delta_n)$ and therefore always ensures also consistency with respect to the strong metric, at some arbitrarily slow rate continues to hold, even when the $\beta$-source condition does not hold for the function but only an $\epsilon$ source condition holds for an arbitrarily small $\epsilon>0$. Interestingly, this conclusion suggests that in the scenario of source double robustness, our focus should be on achieving a fast rate for the weak metric, even if it comes at the expense of the strong metric. \footnote{Recall, by optimizing $\lambda$ just for the strong metric and for $\gamma=0$, we could achieve $\delta_n^{2\beta/(2+\beta)}$, which is faster than $ \delta^{\frac{2(\beta-1)}{\beta+1} }_n$.}
To summarize, by combining the best of the two product rate conditions from the two Corollaries and choosing the best of the two conditions, based on the known value of $\beta$, we get the size requirement on the function classes:
Note that the smoothness assumption always holds when $\gamma=0$, and therefore, in the absence of any smoothness condition the condition is derived from the latter formula with $\gamma=0$.