EconBase
← Back to paper

Kernel Ridge Riesz Representers: Generalization, Mis-specification, and the Counterfactual Effective Dimension

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

90,001 characters · 19 sections · 83 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Kernel Ridge Riesz Representers: Generalization, Mis-specification, and the Counterfactual Effective Dimension

\def\spacingset#1{ {#1}} \spacingset{1}

abstractKernel balancing weights provide confidence intervals for average treatment effects, based on the idea of balancing covariates for the treated group and untreated group in feature space, often with ridge regularization. Previous works on the classical kernel ridge balancing weights have certain limitations: (i) not articulating generalization error for the balancing weights, (ii) typically requiring correct specification of features, and (iii) justifying Gaussian approximation for only average effects. I interpret kernel balancing weights as kernel ridge Riesz representers (KRRR) and address these limitations via a new characterization of the counterfactual effective dimension. KRRR is an exact generalization of kernel ridge regression and kernel ridge balancing weights. I prove strong properties similar to kernel ridge regression: population $L_2$ rates controlling generalization error, and a standalone closed form solution that can interpolate. The framework relaxes the stringent assumption that the underlying regression model is correctly specified by the features. It extends Gaussian approximation beyond average effects to heterogeneous effects, justifying confidence sets for causal functions. I use KRRR to quantify uncertainty for heterogeneous treatment effects, by age, of 401(k) eligibility on assets.

{\it Keywords:} causal inference, Gaussian approximation, heterogeneous treatment effect, reproducing kernel Hilbert space (RKHS), semiparametric efficiency

\spacingset{1.65}

Introduction

Kernel balancing weights adapt kernel methods to quantify uncertainty on average effects, e.g. the average treatment effect when the treatment is randomly assigned conditional on covariates.\footnote{Kernel ridge balancing weights also apply to confidence intervals for the average treatment effects on the treated and untreated subpopulations, as well as subpopulations defined in terms of a discrete covariate taking a finite number of values. This paper justifies confidence sets for causal functions.} The central idea is to balance the feature-mapped covariates in the treated group with those in the untreated group, subject to regularization. The weights that achieve such balance, when applied to the observed outcomes, estimate the average treatment effect and appear in its confidence interval. These weights are widely used in social science via popular R packages, and they are classical in statistics speckman1979minimax.

The motivation for kernel balancing weights is to inherit the simplicity and flexibility of kernel methods. However, not all of the familiar properties of, say, kernel ridge regression have been fully articulated for classical kernel ridge balancing weights, e.g. population $L_2$ rates of convergence for generalization error, a simple standalone closed form solution, and robustness to mis-specification of features. Such properties would lead to strong semiparametric guarantees for causal parameters via well-known “targeting” and “debiasing” arguments, which tolerate some types of mis-specification and which provide Gaussian approximation for causal functions, justifying confidence sets with nominal coverage.

This paper's main contribution is a population $L_2$ rate controlling the generalization error of the classical kernel ridge balancing weights. The rate coincides with the lower bound of regression in the reproducing kernel Hilbert space (RKHS) framework. It holds under mis-specification of features. A technical innovation powers the result: I define the counterfactual effective dimension, and prove it is not much greater than the actual effective dimension. This technique may be of independent interest.

As a secondary contribution, this paper provides a standalone closed form solution for kernel ridge balancing weights that appears to have been previously unknown. This standalone closed form solution can be evaluated at locations different than the training data, allowing kernel ridge balancing weights to be used in “debiased” treatment effect estimators with sample splitting. Such estimators accommodate regression models outside of the RKHS, unlike many previous works on kernel ridge balancing weights.

Conceptually, I study the kernel ridge Riesz representer (KRRR) as an exact generalization of kernel ridge regression and kernel ridge balancing weights. This generalization formalizes the sense in which the population $L_2$ rate is sometimes optimal. For treatment effects, I verify inference with rate-optimal regularization, rather than undersmoothing, of the ridge penalty. The generalization also justifies the construction of confidence sets for causal functions, e.g. heterogeneous policy effects and heterogeneous treatment effects, from kernel ridge balancing weights.\footnote{Heterogeneous treatment effects are a function of a continuous covariate; see Section (ref).}

Section (ref) situates these contributions within the context of related work. Section (ref) formalizes the key assumptions. Section (ref) describes the procedure, derives its standalone closed form, and clarifies how it extends known procedures. Section (ref) theoretically justifies the procedure with population $L_2$ rates of generalization error, which imply strong semiparametric guarantees. Section (ref) proves the rates via the counterfactual effective dimension. Section (ref) demonstrates nominal coverage in heterogeneous treatment effect simulations, with less bias than some previous estimators, then quantifies uncertainty for the heterogeneous treatment effects of 401(k) eligibility on assets as a function of age. The effect appears to be five times stronger for 45-60 year olds, compared to 30 year olds.

Related work

I extend the conceptual framework of kernel balancing weights speckman1979minimax,hazlett2020kernel,wong2018kernel,zhao2019covariate,kallus2020generalized,hirshberg2019augmented,hirshberg2019kernel,chernozhukov2020adversarial. Works in this literature often focus on kernel ridge balancing weights as a vector that possibly converges in $\ell_2$ hirshberg2019augmented,hirshberg2019kernel, whereas I will study them as evaluations of a function that converges in population $L_2$. The former is the training error for a balancing weight that is defined in-sample; see Corollary (ref). By contrast, I analyze generalization error for a balancing weight that can interpolate and that can be evaluated out-of-sample. This distinction is important for stronger semiparametric conclusions: coverage guarantees for causal functions, sample splitting, and regression models beyond the RKHS; see Remark (ref).

bruns2023augmented prove equivalences among various formulations of the classical kernel ridge balancing weights, including a minimax formulation with ridge regularization. Regardless of nomenclature, many of the salient properties of these classical balancing weights, such as population $L_2$ rates of generalization error and a standalone closed form, appear to be contributions. I recover the known, nonasymptotic balance-variance trade-off.

The properties that I prove for KRRR generalize known results for kernel ridge regression. I obtain $L_2$ rates using integral operator techniques for kernel methods, drawing on smale2007learning,caponnetto2007optimal,fischer2017sobolev, among others. See fischer2017sobolev for references and a review of integral operator techniques and empirical process techniques used to analyze RKHS estimators in various norms, many of which rely on the effective dimension. My characterization of the counterfactual effective dimension appears to be new, and it avoids auxiliary approximation assumptions. The closed form solution of KRRR echoes classical solutions kimeldorf1971some,scholkopf2001generalized.

I also build on the insight from semiparametric theory that the balancing weights for the average treatment effect are a special case of Riesz representers for causal parameters, which can be directly estimated robins2007comment,avagyan2021high,chernozhukov2018global,hirshberg2019augmented,chernozhukov2018learning,smucler2019unifying,hirshberg2019kernel,chernozhukov2020adversarial. Similar to chernozhukov2018global,chernozhukov2018learning,chernozhukov2020adversarial, I provide population $L_2$ rates. A key point of departure is that I study KRRR, which is a different estimator more closely aligned with kernel ridge regression, classical kernel ridge balancing weights, and popular R packages. Another distinction is that I place standard RKHS assumptions, which facilitates comparison with known optimality theory, including rate-optimal regularization. Finally, I avoid some auxiliary approximation assumptions of those works by defining and analyzing the counterfactual effective dimension; see Remark (ref).

The contribution of this work is not to provide new semiparametric theory. Rather, I demonstrate that KRRR has good properties similar to kernel ridge regression that were not fully articulated, and hence it is compatible with strong existing results, e.g. zheng2011cross,chernozhukov2018original,rotnitzky2021characterization,kennedy2023towards, among many others. These strong results allow for some mis-specification and rate-optimal regularization, rather than the correct specification and undersmoothing studied by e.g. hirshberg2019kernel,mou2023kernel. See e.g. hirshberg2019augmented,ben2021balancing for detailed citations on balancing weights outside of the RKHS. See e.g. van2011targeted for detailed citations on semiparametrics.

An initial draft was circulated in February 2021 singh2021debiased. Prior to the current draft, wang2021estimation extend the kernel balancing weights of wong2018kernel to heterogeneous treatment effects and prove consistency under correct specification, without formal inference, a closed form, or $L_2$ rates for balancing weights. This paper's results are complementary: formal inference, some mis-specification, a closed form, and $L_2$ rates for balancing weights, within a framework that includes other estimands.\footnote{By formal inference, I mean justification for Gaussian approximation and nominal coverage guarantees.} Subsequent works extend kernel balancing weights to confounding bridges kallus2021causal,ghassami2021minimax and study other nonlinear function spaces chernozhukov2021automatic.

Key assumptions and main results

Warm up: Balancing weights

Consider the problem of estimating the average treatment effect (ATE) $\theta_0$ from $n$ independently and identicically distributed (i.i.d.) observations of the baseline covariates $X\in\mathbb{R}^{dim(X)}$, treatment $D\in\{0,1\}$, and outcome $Y \in\mathbb{R}$. When treatment assignment is as good as random conditional on $X$, the ATE can be expressed in two ways. First, the regression formulation is $\theta_0=\mathbb{E}\{\gamma_0(1,X)-\gamma_0(0,X)\}$, where $\gamma_0(D,X)=\mathbb{E}(Y|D,X)$ is the outcome regression. Second, the balancing weight formulation is $\theta_0=\mathbb{E}\{Y\alpha_0(D,X)\}$, where $\alpha_0(D,X)=D\pi_0(X)^{-1}-(1-D)\{1-\pi_0(X)\}^{-1}$ is the balancing weight expressed as a function and $\pi_0(X)=\mathbb{E}(D|X)$ is the propensity score. One may derive the latter from the former using the law of iterated expectations and the regularity condition that the propensity score $\pi_0$ is bounded away from zero and one.

These two formulations suggest estimation approaches. One option is to estimate the regression as $\hat{\gamma}$ then the ATE as $\mathbb{E}_n\{\hat{\gamma}(1,X)-\hat{\gamma}(0,X)\}$, where $\mathbb{E}_n(\cdot)$ is the empirical mean. Another is to estimate the balancing weight as $\hat{\alpha}$ then the ATE as $\mathbb{E}_n\{Y\hat{\alpha}(D,X)\}$. One may also combine these estimators into a so-called doubly robust procedure $\mathbb{E}_n[\hat{\gamma}(1,X)-\hat{\gamma}(0,X)+\hat{\alpha}(D,X)\{Y-\hat{\gamma}(D,X)\}]$. Each approach also has a sample splitting variant, where $(\hat{\gamma},\hat{\alpha})$ are estimated from, say $n/2$ held out observations, and the empirical mean is taken with respect to the remaining $n/2$ observations. For the sample splitting variant, $\hat{\alpha}$ must be a function that can be evaluated out-of-sample.

The sample splitting variant of the doubly robust procedure is sometimes called “debiased” machine learning chernozhukov2018original, which is closely related to targeted learning van2006targeted and one step estimation pfanzagl1982lecture, and it has quite general inferential guarantees for $\theta_0$. In particular, it allows for mis-specification of $\hat{\gamma}$ or $\hat{\alpha}$, and rate-optimal regularization of $\hat{\gamma}$ and $\hat{\alpha}$, as long as population $L_2$ rates for $\hat{\gamma}$ and $\hat{\alpha}$ are known, i.e. formal characterizations of $\|\hat{\gamma}-\gamma_0\|_2$ and $\|\hat{\alpha}-\alpha_0\|_2$ where $\|W\|_2=\{\mathbb{E}(W^2)\}^{1/2}$.

Previous work proposes a kernel ridge balancing weight $\tilde{\alpha}$. My main contribution is to analyze its generalization error $\|\tilde{\alpha}-\alpha_0\|_2$, achieving the well-known lower bound for regression in the RKHS framework, without auxiliary approximation assumptions. Once a population $L_2$ rate is articulated, general semiparametric guarantees from the literature become available. These guarantees justify Gaussian approximation for causal functions and some mis-specification, unlike many previous works on kernel ridge balancing weights.

Generalizing from balancing weights to Riesz representers

The ATE is a special case of a causal parameter. This paper considers parameters of the form $\theta_0=\mathbb{E}\{m(W,\gamma_0)\}$, where $W$ concatenates observed variables, $\gamma_0(W)=\mathbb{E}(Y|W)$ is a regression, and the formula $m$ accommodates various parameters beyond the ATE. For ATE, $W=(D,X)$ and $m(w,f)=f(1,x)-f(0,x)$. I study a class defined by a regularity condition generalizing the idea that the propensity score is away from zero and one.

assumption[Mean square continuity] The functional $f\mapsto \mathbb{E}\{m(W,f)\}$ is linear. There exists some $\bar{M}<\infty$ such that $\mathbb{E}\{m(W,f)^2\}\leq \bar{M} \|f\|_2^2$ for all $f \in L_2$.

Intuitively, this condition imposes that if $f_1$ and $f_2$ are close, then their functionals $\mathbb{E}\{m(W,f_1)\}$ and $\mathbb{E}\{m(W,f_2)\}$ are close. In the ATE example, if $\pi_0(X)\in(1-\bar{\pi},\bar{\pi})$ almost surely, then $\bar{M}=2\{\bar{\pi}^{-1}+(1-\bar{\pi})^{-1}\}$.

Nonasymptotic analysis allows $\bar{M}$ to be a diverging sequence as the sample size increases. In particular, this class includes functionals of the form $f\mapsto \mathbb{E}\{m_h(W,f)\}$ where $m_h(w,f)=\ell_h(w)\check{m}(w,f)$. The weighting $\ell_h(w)$ is indexed by a vanishing bandwidth $h\downarrow 0$, which also indexes $\bar{M}=\bar{M}_h\uparrow \infty$. In this thought experiment, a sequence of semiparametric quantities, each satisfying Assumption (ref), provides a pointwise approximation to a nonparametric causal function evaluated at a point. I provide two concrete examples.

example[Heterogeneous policy effects] Here $W=X$ are regressors and $\check{m}(w,f)=f\{t(x)\}-f(x)$, where $t$ is a known mapping that serves as the counterfactual policy of interest. The weighting $\ell_h(w)=\textsc{local}\{(x_1-x_1^*)/h\}$ is a localization around the location $X_1=x_1^*$ (see Appendix (ref) for details). Suppose that the density ratio is bounded above, i.e. $\omega(W)=\textsc{density}\{t(W)\}/\textsc{density}(W)\leq\bar{\omega}$ almost surely. Suppose that the weighting function at a fixed $h$ is bounded, i.e. $|\ell_h(W)|\leq \bar{\ell}_h$ almost surely. Then $\bar{M}_h=2\bar{\ell}_h^2(\bar{\omega}+1)$.
example[Heterogeneous treatment effects] Here $W=(D,X)$ concatenates the treatment $D\in\{0,1\}$ and covariates $X$, and $\check{m}(w,f)=f(1,x)-f(0,x)$. The weighting $\ell_h(w)=\textsc{local}\{(x_1-x_1^*)/h\}$ is a localization around the location $X_1=x_1^*$ (see Appendix (ref) for details). Suppose that the propensity score is bounded away from zero and one, i.e. $\pi_0(X)\in(1-\bar{\pi},\bar{\pi})$ almost surely. Suppose that the weighting function at a fixed $h$ is bounded, i.e. $|\ell_h(W)|\leq \bar{\ell}_h$ almost surely. Then $\bar{M}_h=2\bar{\ell}_h^2\{\bar{\pi}^{-1}+(1-\bar{\pi})^{-1}\}$.

For further interpretation, suppose the the first covariate $X_1$ is continuous. In the finite sample, Example (ref) is the heterogeneous treatment effect for the subpopulation whose first covariate is local to $x_1^*$. As $h\downarrow 0$, Example (ref) converges to the heterogeneous treatment effect for the subpopulation whose first covariate value is $x_1^*$. Section (ref) gives a real world example where formal uncertainty quantification leads to economic insights.

Taking $\ell_h(w)=1$, Examples (ref) and (ref) reduce to the average policy effect and average treatment effect, respectively. The class defined by Assumption (ref) contains not only average effects but also heterogeneous effects, i.e. pointwise approximations of causal functions.

The balancing weight is a special case of a Riesz representer. It is well known that each parameter within the class defined by Assumption (ref) has both a regression formulation and a generalized balancing weight formulation.

lemma[Riesz representation; Lemma S3.1 of chernozhukov2018global] Suppose Assumption (ref) holds and $\gamma_0\in \mathcal{G} \subset L_2$. Then there exists a Riesz representer $\alpha_0\in L_2$ such that $ \mathbb{E}\{m(W,f)\}=\mathbb{E}\{\alpha_0(W)f(W)\}$ for all $f\in \mathcal{G}$. Moreover, there exists a unique minimal Riesz representer $\alpha_0^{\min}\in closure\{span(\mathcal{G})\}$ that satisfies this equation.

For the rest of the paper, I will set aside the “correct specification” assumption that $\gamma_0\in \mathcal{G}$, where $\mathcal{G}$ is a known subset of $L_2$ that can be imposed in estimation of $\hat{\gamma}$. As such, there exists a unique Riesz representer $\alpha_0=\alpha_0^{\min}$. See Remark (ref) for a list of many works that place “correct specification” style assumptions requiring $\gamma_0$ to be within the RKHS.

For a functional $f\mapsto \mathbb{E}\{m_h(W,f)\}$ where $m_h(w,f)=\ell_h(w)\check{m}(w,f)$, its Riesz representer is $\alpha_h(w)=\ell_h(w)\check{\alpha}_0(w)$, where $\check{\alpha}_0(w)$ is the Riesz representer for $f\mapsto \mathbb{E}\{\check{m}(W,f)\}$.

In summary, previous works on classical kernel ridge balancing weights provide formal inference for a limited class of causal parameters, i.e. average effects. They also typically require correct specification of $\gamma_0$ within the RKHS. I study the broader class defined in Assumption (ref), which includes heterogeneous policy effects (Example (ref)) and heterogeneous treatment effects (Example (ref)). I take the agnostic stance of $\gamma_0\in L_2$. By analyzing generalization error $\|\tilde{\alpha}-\alpha_0\|_2$, general semiparametric guarantees become immediate.

RKHS notation and key assumptions

The kernel ridge balancing weight algorithm conducts estimation in a reproducing kernel Hilbert space (RKHS) $H\subset L_2$. $H$ consists of functions of the form $f:\mathcal{W}\rightarrow\mathbb{R}$. I denote its symmetric, positive definite kernel by $k:\mathcal{W}\times\mathcal{W}\rightarrow \mathbb{R}$, and its feature map by $\phi:w\mapsto k(w,\cdot)$, so that $k(w,w')=\langle \phi(w),\phi(w') \rangle_H$ and $f(w)=\langle f,\phi(w) \rangle_H$ for $f\in H$.\footnote{Assumptions (ref) and (ref) below formalize weak regularity conditions.}

I use familiar spectral notation from kernel ridge regression analysis. Let $L:L_2\rightarrow L_2$ be the convolution operator $f\mapsto \int k(\cdot,w) f(w) \mathrm{d}\mathbb{P}(w)$. Since $L$ is self adjoint and positive definite, it has weakly decreasing countable eigenvalues $(\eta_j)$ and corresponding eigenfunctions $(\varphi_j)$. Define the space $ H^c=(f=\sum_{j=1}^{\infty}f_j\varphi_j:\sum_{j=1}^{\infty} f_j^2\eta_j^{-c} <\infty) $. In this notation, $H^0=L_2$ and $H^1=H$ when $k$ is characteristic sriperumbudur2010relation. The RKHS $H$ is the subset of $L_2$ for which higher order terms in the series $(\varphi_j)$ have a smaller contribution. By Mercer's theorem, the feature map is $ \phi(w)=\{\eta_j^{1/2} \varphi_j(w)\}^{\infty}_{j=1}$.

I also use the extended notation of chernozhukov2020adversarial, whose critical insight is that causal inference introduces counterfactual features. Since each $\varphi_j$ is in $L_2$, $m(w,\varphi_j)$ is well defined. I write the counterfactual feature map as $\phi^{(m)}(w)=\{\eta_j^{1/2} m(w,\varphi_j)\}^{\infty}_{j=1}$. The counterfactual kernel is then $k^{(m)}(w,w')=\langle \phi^{(m)}(w),\phi^{(m)}(w') \rangle_H$, and the counterfactual evaluation is $m(w,f)=\langle f, \phi^{(m)}(w)\rangle_H$ for $f\in H$. In the example of ATE, the formula is $m(w,f)=f(1,x)-f(0,x)$, so $ \phi^{(m)}(W_i)=[\eta_j^{1/2} \{\varphi_j(1,X_i)-\varphi_j(0,X_i)\}]^{\infty}_{j=1}=\phi(1,X_i)-\phi(0,X_i). $ The counterfactual feature map applied to observation $W_i$ replaces the observed treatment value $D_i$ with the counterfactual values $d=1$ and $d=0$.

I place two standard assumptions from kernel ridge regression analysis. The first, called the source condition, concerns an object I call $\alpha_H \in H$, which is the best RKHS approximation to the Riesz representer $\alpha_0\in L_2$.\footnote{For ATE, $\alpha_H\in\operatorname*{\arg\!\min}_{\alpha\in H} \|\alpha-\alpha_0\|_2$ and $\alpha_0(W)=D\pi_0(X)^{-1}-(1-D)\{1-\pi_0(X)\}^{-1}$.} This assumption helps to analyze bias. The second assumption is the spectral decay condition, which quantifies the effective dimension of the features $\phi(W)$ of the RKHS $H$, and helps to analyze variance. I denote by $\mathcal{P}(b,c)$ the class of distributions that satisfy these two assumptions.

assumption[Smoothness] Assume $\alpha_H\in H^c$ for some $c\in[1,2]$.

By definition of $\alpha_H \in H$ as the best kernel approximation to $\alpha_0 \in L_2$, $c\geq 1$ automatically holds in Assumption (ref). A value $c>1$ means that $\alpha_H$ is a particularly smooth element of $H$. The $L_2$ rates in this paper do not improve beyond $c=2$, reflecting the well known saturation effect of Tikhonov regularization bauer2007regularization.

assumption[Spectral decay] (i) If $H$ is infinite dimensional, the eigenvalues $(\eta_j)$ decay at least polynomially: $ j^b \cdot \eta_j\leq \bar{B}$ for $j\geq1$, where $b\in(1,\infty)$ and $\bar{B}>0$ is a constant. If $H$ is finite dimensional, write its dimension $J$ as $J\leq\bar{B}<\infty$ and write $b=\infty$. (ii) If $H$ is infinite dimensional, we may further impose that the eigenvalues decay exactly polynomially: $\underline{B} \leq j^b \cdot \eta_j \leq \bar{B}$ where $\underline{B},\bar{B}>0$ are constants.

For any bounded kernel $k$, $b\geq 1$ automatically holds in Assumption (ref)(i) fischer2017sobolev. A value $b>1$ means that the RKHS has a lower effective dimension in light of the data distribution. The $L_2$ rates in this paper improve all the way to the parameteric rate as $b\rightarrow \infty$.

The upper bound, Assumption (ref)(i), applies to the Mat\'ern, Gaussian, and other kernels with spectral decay that is polynomial, exponential, or “better”. The lower bound, Assumption (ref)(ii), applies to the Mat\'ern kernel but not the Gaussian kernel. The main results use Assumption (ref)(i) only; Assumption (ref)(ii) is for comparison to optimality theory.

These assumptions generalize familiar assumptions from the analysis of Sobolev spaces. For example, take $\mathcal{W}=\mathbb{R}^p$ and denote by $\mathbb{H}_2^s$ the Sobolev space with $s>p/2$ square integrable derivatives, which is an RKHS with the Mat\'ern kernel. If the RKHS used in estimation is $H=\mathbb{H}_2^s$ and the best RKHS approximation to the Riesz representer satisfies $\alpha_H\in \mathbb{H}_2^{s_0}$, then $c=s_0/s$ and $\mathbb{H}_2^{s_0}=(\mathbb{H}_2^{s})^c$. In words, $c$ quantifies how smooth $\alpha_H$ is relative to its estimator $\tilde{\alpha}\in H$. In this RKHS, $b=2s/p$. As such, $b$ quantifies how smooth the estimator $\tilde{\alpha}\in H$ is relative to the ambient dimension, i.e. its effective dimension.

Preview of main results

I study a classical kernel ridge estimator $\tilde{\alpha}$ of the Riesz representer $\alpha_0$. It turns out that $\tilde{\alpha}=[\mathbb{E}_n\{\phi(W)\otimes \phi(W)^*\}+\lambda]^{-1}\mathbb{E}_n\{\phi^{(m)}(W)\}$, where I use the outer product notation $\{\phi(W)\otimes \phi(W)^*\}(\cdot)=\phi(W)\langle \phi(W),\cdot \rangle_H$ and the ridge regularization notation $\lambda=\lambda I$ with identity $I:H\rightarrow H$. Since $\tilde{\alpha}$ involves counterfactual features $\phi^{(m)}(W)$, a notion of their effective dimension is necessary; to analyze the variance of the estimator $\tilde{\alpha}$, the counterfactual effective dimension is unavoidable.

lemma[Main lemma] Under Assumption (ref) and weak regularity conditions (Assumptions (ref) and (ref) below), the counterfactual effective dimension of $\phi^{(m)}(W)$ is upper bounded by $\bar{M}$ times the effective dimension of $\phi(W)$. See Section (ref) for details.

By Lemma (ref), a familiar regularity condition from semiparametric theory (Assumption (ref)), together with a familiar effective dimension condition from learning theory (Assumption (ref)), implies control of the counterfactual effective dimension and hence the variance of $\tilde{\alpha}$. Due to this insight, I avoid an additional approximation assumption on the counterfactual effective dimension when proving the main result; see Remark (ref).

theorem[Main theoretical result] Under Assumptions (ref), (ref) and (ref)(i) as well as weak regularity conditions (Assumptions (ref) and (ref) below), $\|\tilde{\alpha}-\alpha_H\|_2^2=O_p\{n^{-bc/(bc+1)}\}$ using regularization $\lambda=n^{-b/(bc+1)}$, when $b\in(1,\infty)$, $c\in(1,2]$, and $\bar{M}$ is fixed. This rate is optimal for some choices of $m$ when Assumption (ref)(ii) also holds. See Section (ref) for details, including results when $b=\infty$, $c=1$, or $\bar{M}_h\uparrow \infty$.
corollary[Main corollary] Under weak regularity conditions and correct specification of $\hat{\gamma}$ and $\hat{\alpha}$, perhaps outside of the RKHS, if $\bar{M}$ is fixed then $\hat{\theta}$ converges to $\theta_0$ at the rate $n^{-1/2}$ and is asymptotically normal. Under mis-specification of $\hat{\gamma}$ or $\hat{\alpha}$, $\hat{\theta}$ converges to $\theta_0$, albeit at a slower rate. For pointwise approximations of nonparametric causal functions with $\bar{M}_h\uparrow\infty$, $\hat{\theta}$ converges to $\theta_0$ at the rate $(nh)^{-1/2}$ and is asymptotically normal when the bandwidth is $h=o(1)$ and $n^{-1/2}h^{-3/2}=o(1)$. See Section (ref) and Appendix (ref).
proposition[A practical result] The estimator $\tilde{\alpha}$ has a standalone closed form solution that can be computed from $k$ and $m$, without directly evaluating the feature maps, and that can interpolate. Kernel ridge regression and kernel ridge balancing weights are special cases, and the latter have an equivalent minimax formulation. See Section (ref) for details.

Kernel ridge Riesz representers

General loss function

An important departure from many works on kernel ridge balancing weights is that I allow incorrect specification by the RKHS.\footnote{See Remark (ref) below.} In particular, I allow $\gamma_0 \not \in H$ and $\alpha_0 \not \in H$. For the scenario where $\alpha_0\not\in H$, I define the best kernel approximation $\alpha_H \in H$ to the Riesz representer $\alpha_0\in L_2$. I defer proofs for this section to Appendix (ref).

definition[Kernel approximation to Riesz representer] Let $\alpha_H\in\operatorname*{\arg\!\min}_{\alpha\in H} \|\alpha-\alpha_0\|_2$.

I place weak regularity conditions on the original spaces and on the RKHS $H$ that allow for further characterization of $\alpha_H$.

assumption[Original spaces] $W$ takes values in a Polish space $\mathcal{W}$, i.e. a separable and completely metrizable topological space. $Y$ takes values in $\mathcal{Y}\subset \mathbb{R}$.
assumption[RKHS regularity] The kernels $k$ and $k^{(m)}$ are bounded, i.e. $ \sup_{w\in\mathcal{W}}\|\phi(w)\|_H\leq \sqrt{\kappa}$ and $ \sup_{w\in\mathcal{W}}\|\phi^{(m)}(w)\|_H\leq \sqrt{\kappa^{(m)}}$. The feature maps $\phi(w)$ and $\phi^{(m)}(w)$ are measurable. The kernel $k^{(m)}$ is characteristic sriperumbudur2010relation in its components that vary.

The space $\mathcal{W}$ accommodates general data types, e.g. texts, images, and graphs. Commonly used kernels are bounded. Boundedness implies that the counterfactual kernel mean embedding $\mu^{(m)}=\mathbb{E}\{\phi^{(m)}(W)\}$ and the uncentered covariance operator $T=\mathbb{E}\{\phi(W)\otimes \phi(W)^*\}$ are well defined.\footnote{Here, I use outer product notation: $(a\otimes b^*)c=a \langle b,c\rangle_H$, where $b^*$ is the adjoint of $b$.} This counterfactual kernel mean embedding generalizes the well known kernel mean embedding smola2007hilbert. It is unconditional and does not involve the outcome. Measurability is another standard condition.

The characteristic property ensures that no counterfactual information is lost when approximating $\alpha_0$ as $\alpha_H$. In particular, it means that $\mathbb{P}\mapsto \mu^{(m)}$ is injective, so $\mu^{(m)}$ may be used as a sufficient statistic for $\alpha_H$, as shown in the following lemma.

lemma[Towards a loss for KRRR] Suppose Assumptions (ref), (ref), and (ref) hold. Then we have $ \alpha_H\in\operatorname*{\arg\!\min}_{\alpha\in H} \mathcal{L}(\alpha)$ where $\mathcal{L}(\alpha)=-2\langle \alpha, \mu^{(m)} \rangle_H+\langle \alpha, T\alpha\rangle_H $.

The sufficient statistics $\mu^{(m)}\in H$ and $T\in H\otimes H^*$ are infinite dimensional generalizations of the counterfactual moment and the covariance matrix in automatic debiased machine learning (Auto-DML) chernozhukov2018global,chernozhukov2018learning. Auto-DML uses $p$ explicit basis functions, and its sufficient statistic are a vector in $\mathbb{R}^p$ and a matrix in $\mathbb{R}^{p \times p}$. Kernel balancing weights use the countable eigenfunctions of the kernel $k$ as infinitely many implicit basis functions.

Inspired by Lemma (ref), I study a general loss for kernel ridge Riesz representers (KRRR) at the population level. As we will see below, its empirical analogue recovers the loss for kernel ridge balancing weights previously proposed for treatment effects. Let $\hat{\mu}^{(m)}=\mathbb{E}_n\{\phi^{(m)}(W)\}$ and $\hat{T}=\mathbb{E}_n\{\phi(W)\otimes \phi(W)^*\}$ be the empirical summary statistics.

definition[Loss for KRRR] Let $ \alpha_{\lambda}=\operatorname*{\arg\!\min}_{\alpha\in H} \mathcal{L}_{\lambda}(\alpha) $ where $ \mathcal{L}_{\lambda}(\alpha)=\mathcal{L}(\alpha)+\lambda \|\alpha\|^2_H. $ The estimator is its empirical analogue: $ \tilde{\alpha}=\operatorname*{\arg\!\min}_{\alpha\in H}\mathcal{L}^{n}_{\lambda}(\alpha)$ where $\mathcal{L}^{n}_{\lambda}(\alpha)=-2\langle \alpha, \hat{\mu}^{(m)} \rangle_H+\langle \alpha, \hat{T}\alpha\rangle_H+\lambda \|\alpha\|^2_H. $

The KRRR loss in Definition (ref) differs from Dantzig chernozhukov2018global, Lasso chernozhukov2018learning, and adversarial chernozhukov2020adversarial losses for Riesz representers whose generalization errors have been previously analyzed. The KRRR loss has a simple first order condition which implies a nonasymptotic balance-variance trade-off similar to those of kernel balancing weights in the literature hazlett2020kernel,wong2018kernel,zhao2019covariate,kallus2020generalized,hirshberg2019kernel.

lemma[First order condition] Suppose the conditions of Lemma (ref) hold. The first order conditions give $ T\alpha_H=\mu^{(m)}$, $\alpha_{\lambda}=(T+\lambda)^{-1}\mu^{(m)}$, and $\tilde{\alpha}=(\hat{T}+\lambda)^{-1}\hat{\mu}^{(m)}. $
corollary[Balance-variance trade-off] Suppose the conditions of Lemma (ref) hold. Then we have $\|\mathbb{E}_n\{\tilde{\alpha}(W)\phi(W)-\phi^{(m)}(W)\}\|_{H}/\|\tilde{\alpha}\|_{H}=\lambda$.

By Definition (ref), a smaller regularization value $\lambda$ translates to an estimator $\tilde{\alpha}$ with a smaller bias and a larger variance. Corollary (ref) illustrates that a smaller (normalized) imbalance of the features, $\|\mathbb{E}_n[\tilde{\alpha}(W)\phi(W)-\phi^{(m)}(W)]\|_{H}/\|\tilde{\alpha}\|_{H}$, is possible if and only if we pay the cost of a larger variance. In the finite sample analysis to follow, $\lambda\downarrow 0$ in a way that optimally navigates this trade-off (when rate lower bounds are known). In particular, we will not require undersmoothing for inference on the causal parameter.

Detailed comparisons: Ridge regression and balancing weights

Lemma (ref) has several more corollaries, which formalize the sense in which KRRR builds on and extends previous frameworks. In these corollaries, let $\tilde{\gamma}$ be the kernel ridge regression estimator. Throughout, I assume that the conditions of Lemma (ref) hold.

corollary[Special case: Kernel ridge regression] $\tilde{\alpha}=\tilde{\gamma}$ when $m(w,f)=yf(w)$.

Next, let $\tilde{\beta}\in\mathbb{R}^n$ be the kernel ridge balancing weight previously defined for average effects speckman1979minimax,kallus2020generalized,hirshberg2019augmented,hirshberg2019kernel.

definition[Loss for kernel ridge balancing weights] Let $\tilde{\beta}=\operatorname*{\arg\!\min}_{\beta\in\mathbb{R}^n} \tilde{\mathcal{L}}^{n}_{\lambda}(\beta)$ where $\tilde{\mathcal{L}}^{n}_{\lambda}(\beta)=[\sup_{f\in H:\|f\|_H\leq 1}\frac{1}{n}\sum_{i=1}^n \{\beta_i f(W_i)-m(W_i,f)\}]^2+n^{-1}\lambda \beta^\top\beta. $
corollary[Special case: Kernel ridge balancing weights; eq. 8 of bruns2023augmented] On the training data, $\tilde{\alpha}(W_i)=\tilde{\beta}_i$. Appendix D.3 of hirshberg2019kernel studies $\frac{1}{n}\sum_{i=1}^n \{\tilde{\beta}_i-\alpha_0(D_i,X_i)\}^2$, i.e. training error.

In summary, KRRR nests kernel ridge regression and kernel ridge balancing weights, unlike the adversarial Riesz representer of chernozhukov2020adversarial. The connection to kernel ridge regression foreshadows the new generalization error guarantees for KRRR.

Next, I recap known equivalences for the estimation of treatment effects speckman1979minimax,kallus2020generalized,hirshberg2019kernel and more generally for the estimation of causal parameters within the class of Assumption (ref) bruns2023augmented. With $\hat{\alpha}$ taken to be KRRR $\tilde{\alpha}$, these equivalences only hold when the regression estimator $\hat{\gamma}$ is kernel ridge regression $\tilde{\gamma}$, and when the same observations are used for $\tilde{\gamma}$, $\tilde{\alpha}$, and the causal parameter.

corollary[Equivalence in a special case; Proposition 27 of kallus2020generalized] When $\tilde{\gamma}$ is kernel ridge regression and observations are reused for $\tilde{\gamma}$, $\tilde{\alpha}$, and the causal parameter, $\mathbb{E}_n\{m(W,\tilde{\gamma})\}=\mathbb{E}_n\{Y\tilde{\alpha}(W)\}$.
corollary[Non-equivalence in general] For a generic regression estimator $\hat{\gamma}\neq\tilde{\gamma}$, we have $\mathbb{E}_n\{m(W,\hat{\gamma})\}\neq \mathbb{E}_n\{Y\tilde{\alpha}(W)\}$. For any regression estimator $\hat{\gamma}$ (including $\tilde{\gamma}$), if $\hat{\gamma}$ or $\tilde{\alpha}$ is estimated on a held out sample, then the $\mathbb{E}_n\{m(W,\hat{\gamma})\}\neq \mathbb{E}_n\{Y\tilde{\alpha}(W)\}$.

When $\gamma_0\not \in H$, alternative regression estimators lead to better guarantees for semiparametric inference; see Section (ref) below.

Standalone closed form solution

In summary, the standalone closed form for KRRR is not necessary for causal estimation in the special case of Corollary (ref). However, Corollary (ref) shows that if $\hat{\gamma}\neq \tilde{\gamma}$, or if $\hat{\gamma}$ or $\tilde{\alpha}$ is estimated on a held out sample, then a standalone closed form for KRRR that can be evaluated away from the training data is necessary for causal estimation. I derive it below using techniques of chernozhukov2020adversarial, as a secondary contribution.

Even for the special case of Corollary (ref), this paper's primary contribution is the population $L_2$ rate $\|\tilde{\alpha}-\alpha_H\|_2$ for kernel ridge balancing weights (and more generally KRRR) in Section (ref). Generalization error is essential for appealing to strong semiparametric guarantees that allow some mis-specification and justify inference on causal functions.

I use some additional notation for features. Define the operator $\Phi:H\rightarrow \mathbb{R}^n$ with $i$th component $\langle \phi(W_i),\cdot \rangle_H$ and likewise the operator $\Phi^{(m)}:H\rightarrow \mathbb{R}^n$ with $i$th component $\langle \phi^{(m)}(W_i),\cdot \rangle_H$. Concatenate $\Phi$ and $\Phi^{(m)}$ as $\Psi:H\rightarrow \mathbb{R}^{2n}$, with adjoint $\Psi^*:\mathbb{R}^{2n}\rightarrow H$.

lemma[Standalone closed form exists; c.f. Lemma 2 of chernozhukov2020adversarial] Suppose the conditions of Lemma (ref) hold. There exists some $\rho\in\mathbb{R}^{2n}$ such that $\tilde{\alpha}=\Psi^*\rho$.

Whereas kernel ridge regression has a solution in $\mathbb{R}^n$ kimeldorf1971some,scholkopf2001generalized, KRRR has a solution in $\mathbb{R}^{2n}$ because of the counterfactual features. Intuitively, causal inference involves reasoning about $n$ actual observations and $n$ counterfactual observations.\footnote{The conjecture in hirshberg2019augmented that kernel ridge balancing weights may have a representation $\rho\in \mathbb{R}^n$, such that $\tilde{\alpha}=\Phi^*\rho$, does not appear to be supported in general.}

The familiar kernel matrix $K^{(1)}=\Phi\Phi^*\in\mathbb{R}^{n\times n}$ encodes inner products between actual observations. The extended kernel matrix $K=\Psi \Psi^* \in \mathbb{R}^{2n\times 2n}$ encodes inner products between actual observations and counterfactual observations. Each entry of $K$ is computed from the kernel $k$ and formula $m$ as discussed in Section (ref). For intuition, write $$ \Psi=

bmatrix[bmatrix omitted — 31 chars of source]

,\quad K=

bmatrix[bmatrix omitted — 54 chars of source]

=

bmatrix[bmatrix omitted — 102 chars of source]

. $$

proposition[Standalone closed form for KRRR] Suppose the conditions of Lemma (ref) hold. Given observations $(W_i)$ and evaluation location $w$, as well as the formula $m$, kernel $k$, and regularization $\lambda$, the KRRR closed form is as follows. \begin{enumerate} • Calculate $K^{(j)}\in\mathbb{R}^{n\times n}$ as defined above, for $j\in\{1,2,3,4\}$. • Calculate $\Omega^{2n\times 2n}$ and $u(w),v\in\mathbb{R}^{2n}$ by $$ \Omega=\begin{bmatrix} K^{(1)}K^{(1)} & K^{(1)}K^{(2)}\\ K^{(3)}K^{(1)} & K^{(3)}K^{(2)} \end{bmatrix},\; v=\begin{bmatrix} K^{(2)} \\ K^{(4)} \end{bmatrix} \mathbbm{1}_{n},\; u_i(w)=\begin{cases} k(W_i,w) & \text{ if }i\in \{1,...,n\}\\ k_m(W_i,w)& \text{ if }i\in \{n+1,...,2n\}\end{cases} $$ where $\mathbbm{1}_{n}\in\mathbb{R}^{n}$ is a vector of ones, $k(w,w')=\langle \phi(w),\phi(w')\rangle_H$, and $k_m(w,w')=\langle \phi^{(m)}(w),\phi(w')\rangle_H$. • Set $ \tilde{\alpha}(w)=v^{\top}(\Omega+n\lambda K)^{-1} u(w). $ \end{enumerate}

The standalone closed form solution for kernel ridge balancing weights, and KRRR more generally, appears to have been previously unknown. Proposition (ref) demonstrates that it can be easily computed using only the formula $m$ and kernel $k$, without directly evaluating the feature maps $\phi(w)$ or $\phi^{(m)}(w)$. In this sense, it is a practical contribution for inference. See Appendix (ref) for its specialization to Example (ref).

Generalization error with mis-specification

Main result

I now provide the main result: population $L_2$ rates for KRRR that coincide with the optimal rate, where optimal rates are known. I study the class of Riesz representers defined by Assumption (ref), imposing familiar smoothness and spectral decay conditions for the RKHS setting defined by Assumptions (ref) and (ref). I also impose mild regularity conditions that help in the algorithm derivation, defined by Assumptions (ref) and (ref).

For clarity, I write the population $L_2$ rate for KRRR, estimated from the observations $(W_i)_{i\in[n]}$, in terms of $\mathcal{R}(\tilde{\alpha})=\mathbb{E}[\{\tilde{\alpha}(W)-\alpha_H(W)\}^2|(W_i)_{i\in[n]}]$. It is the generalization error for a test observation $W$, allowing for mis-specification.

theorem[Upper bound for KRRR] If Assumptions (ref), (ref), (ref)(i), (ref), and (ref) hold with $\bar{M}$ bounded above, then $ \lim_{\tau\rightarrow \infty} \lim\sup_{n\rightarrow\infty} \sup_{p\in\mathcal{P}(b,c)} \mathbb{P}_{(W_i)\sim p^{n}}\{\mathcal{R}(\tilde{\alpha})>\tau \cdot r_n\}=0 $, where $$ r_n=\begin{cases} n^{-1} & \text{ if }\quad b=\infty \text{ and } \lambda=n^{-\frac{1}{2}}; \\ n^{-\frac{bc}{bc+1}} & \text{ if }\quad b\in(1,\infty),\; c\in(1,2]\text{ and } \lambda=n^{-\frac{b}{bc+1}}; \\ \ln^{\frac{b}{b+1}}(n)\cdot n^{-\frac{b}{b+1}} & \text{ if }\quad b\in(1,\infty),\; c=1\text{ and } \lambda=\ln^{\frac{b}{b+1}}(n)\cdot n^{-\frac{b}{b+1}}. \end{cases} $$
remark[Causal functions] When the causal parameter has a formula of the form $m_h(w,f)=\ell_h(w)\check{m}(w,f)$, as in Example (ref), $\bar{M}_h\uparrow \infty$. In such case, I use the estimator $\hat{\alpha}_h(W)=\ell_h(W)\tilde{\alpha}(W)$, where $\tilde{\alpha}$ is KRRR for the formula $\check{m}(w,f)$. When $\tilde{\alpha}$ has the rate $r_n$, $\hat{\alpha}_h$ has the rate $h^{-2}r_n$ under weak regularity conditions. See Appendix (ref) for details.

I defer the proof to Section (ref). I defer the proofs of corollaries to Appendix (ref).

Theorem (ref) does not require Assumption (ref)(ii). Moreover, as an intermediate step, I prove a nonasymptotic bound (Proposition (ref)) without imposing Assumptions (ref) and (ref)(i).

When the RKHS is finite dimensional, i.e. $b=\infty$ in Assumption (ref), KRRR achieves the parametric rate $r_n=n^{-1}$. When the RKHS is infinite dimensional with at least polynomial spectral decay, i.e. $b<\infty$ in Assumption (ref), KRRR achieves a familiar rate $r_n=n^{-\frac{bc}{bc+1}}$. This rate depends on $c$ in Assumption (ref), which quantifies the smoothness of the best kernel approximation $\alpha_H \in H$ to the Riesz representer $\alpha_0 \in L_2$. If no additional smoothness of $\alpha_H$ is known, i.e. all we have is that $\alpha_H\in H$ by construction, then $c=1$ and the rate has an extra logarithmic factor.

The rate $r_n=n^{-\frac{bc}{bc+1}}$ generalizes the familiar rate from Sobolev analysis. For example, take $\mathcal{W}=\mathbb{R}^p$ and denote by $\mathbb{H}_2^s$ the Sobolev space with $s>p/2$ square integrable derivatives, as before. Suppose that KRRR $\tilde{\alpha}$ is estimated using the RKHS $H=\mathbb{H}_2^s$. Suppose that $\alpha_H\in \mathbb{H}_2^{s_0}$, i.e. the best kernel approximation to the Riesz representer $\alpha_0\in L_2$ has $s_0>s$ square integrable derivatives. The population $L_2$ rate for $\tilde{\alpha}$ is then $r_n=n^{-\frac{2s_0}{2s_0+p}}$. When $s_0=s$, which is a lower bound on $s_0$ by construction, $r_n=\ln(n)^{\frac{2s}{2s+p}}\cdot n^{-\frac{2s}{2s+p}}$. Theorem (ref) applies to settings beyond Sobolev spaces over Euclidean domains.

corollary[Special case: Kernel ridge regression; Theorem 1 of caponnetto2007optimal] Suppose the conditions of Theorem (ref) hold, replacing Assumption (ref) with $|Y|\leq \bar{Y}$ almost surely. Then $ \lim_{\tau\rightarrow \infty} \lim\sup_{n\rightarrow\infty} \sup_{p\in\mathcal{P}(b,c)} \mathbb{P}_{(W_i)\sim p^{n}}\{\mathcal{R}(\tilde{\gamma})>\tau \cdot r_n\}=0. $
remark[Excess risk] Some works on kernel ridge regression state results in terms of the excess risk $\mathcal{E}(\tilde{\gamma})-\inf_{\gamma\in H}\mathcal{E}(\gamma)$, where $\mathcal{E}(f)=\mathbb{E}[\{Y-f(W)\}^2]$. This is equivalent to $\mathcal{R}(\tilde{\gamma})=\mathbb{E}[\{\tilde{\gamma}(W)-\gamma_H(W)\}^2|(W_i)_{i\in[n]}]$ when $\gamma_H\in \operatorname*{\arg\!\min}_{\gamma \in H} \mathcal{E}(\gamma)$ exists, which is assumed in those works and in this one. See e.g. caponnetto2007optimal and fischer2017sobolev.

Just as KRRR generalizes kernel ridge regression, Theorem (ref) generalizes caponnetto2007optimal. The contribution is meaningful because KRRR also generalizes the classical kernel ridge balancing weights, for which generalization error was not fully articulated. In settings where lower bounds are known, the upper bounds in Theorem (ref) often achieve them.

lemma[Lower bound; Theorem 2 of caponnetto2007optimal] Suppose the conditions of Corollary (ref) hold as well as Assumption (ref)(ii), $b\in (1,\infty)$, and $c\in [1,2]$. Then $ \lim_{\tau\rightarrow 0} \lim\inf_{n\rightarrow\infty} \inf_{\hat{\gamma}} \sup_{p\in \mathcal{P}(b,c)}\mathbb{P}_{(W_i)\sim p^{n}}\{\mathcal{R}(\hat{\gamma})>\tau \cdot n^{-\frac{bc}{bc+1}}\}=1. $
corollary[Optimality] When $b\in(1,\infty)$ and $c\in(1,2]$, Theorem (ref) matches Lemma (ref). When $b\in(1,\infty)$ and $c=1$, they match up to a log factor.

Semiparametric inference beyond the RKHS

To showcase the role of Theorem (ref) in semiparametric inference, I quote a result from the targeted and debiased machine learning literature. Unlike many previous works on kernel ridge balancing weights, inference allows $\gamma_0\not\in H$. It also allows for sample splitting, which is helpful when the regression estimator $\hat{\gamma}$ is not kernel ridge regression, and is instead estimated in a function space with high entropy. Theorem (ref) is directly compatible with the quoted result, unlike previous guarantees for kernel ridge balancing weights. In this sense, KRRR provides debiased kernel methods.

To state the result, denote the doubly robust estimator with sample splitting as $\hat{\theta}$, with the well known analytic variance estimator $\hat{\sigma}^2$.

definition[Debiased machine learning] Given a sample $(Y_i,W_i)_{i\in[n]}$, partition the sample into folds $(I_{\ell})_{\ell\in [L]}$. Denote by $I_{\ell}^c$ the complement of $I_{\ell}$. \begin{enumerate} • For each fold $\ell$, estimate $\hat{\gamma}_{\ell}$ and $\hat{\alpha}_{\ell}$ from observations in $I_{\ell}^c$. • Estimate $\theta_0$ as $ \hat{\theta}=n^{-1}\sum_{\ell=1}^L\sum_{i\in I_{\ell}} [m(W_i,\hat{\gamma}_{\ell})+\hat{\alpha}_{\ell}(W_i)\{Y_i-\hat{\gamma}_{\ell}(W_i)\}] $. • Estimate $\hat{\sigma}^2=n^{-1}\sum_{\ell=1}^L\sum_{i\in I_{\ell}} [m(W_i,\hat{\gamma}_{\ell})+\hat{\alpha}_{\ell}(W_i)\{Y_i-\hat{\gamma}_{\ell}(W_i)\}-\hat{\theta}]^2 $. • Construct the $(1-a) 100$% confidence interval as $ \hat{\theta}\pm c_{a}\hat{\sigma} n^{-1/2}$, where $c_{a}$ is the $1-a/2$ quantile of the standard Gaussian. \end{enumerate}

Let $\gamma_{\mathcal{G}}$ be the best approximation to $\gamma_0$ in some space $\mathcal{G}$, which may not be an RKHS. Similarly, let $\alpha_{\mathcal{A}}$ be the best approximation to $\alpha_0$ in some space $\mathcal{A}$. Slightly abusing notation, I now let $\mathcal{R}(\hat{\gamma})=\mathbb{E}[\{\hat{\gamma}(W)-\gamma_{\mathcal{G}}(W)\}^2|(W_i)_{i\in[n]}]$ and $\mathcal{R}(\hat{\alpha})=\mathbb{E}[\{\hat{\alpha}(W)-\alpha_{\mathcal{A}}(W)\}^2|(W_i)_{i\in[n]}]$.\footnote{In Theorem (ref) above, $\mathcal{A}=H$. In Corollary (ref) above, $\mathcal{G}=H$.}

To quote the result, I introduce some additional notation. Let $\psi(W,\theta,\gamma,\alpha)=m(W,\gamma)+\alpha(W)\{Y-\gamma(W)\}-\theta$, so that $\psi_0(W)=\psi(W,\theta_0,\gamma_0,\alpha_0)$ is the influence function with moments $\sigma^2=\mathbb{E}\{\psi_0(W)^2\}$, $\zeta^3=\mathbb{E}\{\psi_0(W)^3\}$, and $\chi^4=\mathbb{E}\{\psi_0(W)^4\}$.

lemma[Inference; Corollary 1 of chernozhukov2021simple] Suppose Assumption (ref) holds. Suppose the following regularity conditions: $ \mathbb{E}[\{Y-\gamma_0(W)\}^2 \mid W]\leq \bar{\sigma}^2$, $ \|\alpha_0\|_{\infty}\leq\bar{\alpha}$, $ \|\hat{\alpha}\|_{\infty}\leq\bar{\alpha}'$, and $\mathopen{}\mathclose\bgroup\originalleft\{\mathopen{}\mathclose\bgroup\originalleft(\zeta/\sigma\aftergroup\egroup\originalright)^3+\chi^2\aftergroup\egroup\originalright\}n^{-1/2}\rightarrow0. $ Suppose the following learning rate conditions: \begin{enumerate} • $\mathopen{}\mathclose\bgroup\originalleft(\bar{M}^{1/2}+\bar{\alpha}/\sigma+\bar{\alpha}'\aftergroup\egroup\originalright)\{\mathcal{R}(\hat{\gamma})+\|\gamma_{\mathcal{G}}-\gamma_0\|_2^2\}^{1/2}=o_p(1)$; $\bar{\sigma}\{\mathcal{R}(\hat{\alpha})+\|\alpha_{\mathcal{A}}-\alpha_0\|_2^2\}^{1/2}=o_p(1)$; • $[n \{\mathcal{R}(\hat{\gamma})+\|\gamma_{\mathcal{G}}-\gamma_0\|_2^2\} \{\mathcal{R}(\hat{\alpha})+\|\alpha_{\mathcal{A}}-\alpha_0\|_2^2\}]^{1/2}/\sigma =o_p(1)$. \end{enumerate} Then $\sigma^{-1}n^{1/2}(\hat{\theta}-\theta_0)\leadsto\mathcal{N}(0,1)$ and $ \mathbb{P}\mathopen{}\mathclose\bgroup\originalleft\{\theta_0 \in \mathopen{}\mathclose\bgroup\originalleft(\hat{\theta}\pm c_a\hat{\sigma} n^{-1/2} \aftergroup\egroup\originalright)\aftergroup\egroup\originalright\}\rightarrow 1-a $.

Using weak regularity conditions, population $L_2$ rates $\mathcal{R}(\hat{\gamma})$ and $\mathcal{R}(\hat{\alpha})$, and approximation errors $\|\gamma_{\mathcal{G}}-\gamma_0\|_2$ and $\|\alpha_{\mathcal{A}}-\alpha_0\|_2$, the estimator $\hat{\theta}$ is consistent and asymptotically Gaussian at the rate $n^{-1/2}\sigma$. Its confidence interval includes $\theta_0$ with probability approaching the nominal level. Importantly, $\hat{\gamma}$ does not have to be kernel ridge regression.

remark[Mis-specification of features] Crucially, $\gamma_0$ does not need to be well specified in the RKHS $H$. Instead, $\|\gamma_{\mathcal{G}}-\gamma_0\|_2$ must vanish for some $\mathcal{G}$ that may not be $H$. Previous works on classical kernel ridge balancing weights, which provide rates similar to $n^{-1/2}\sigma$ for the causal parameter, typically require $\gamma_0\in H$; see e.g. hazlett2020kernel, wong2018kernel, zhao2019covariate, kallus2020generalized, nie2021quasi, hirshberg2019kernel, wang2021estimation and mou2023kernel. hirshberg2019augmented requires $\|\hat{\gamma}-\gamma_0\|_H=O_p(1)$ and $\mathbb{E}_n[\{\hat{\gamma}(W)-\gamma_0(W)\}^2]=o_p(1)$, which is satisfied by correctly specified kernel ridge regression, i.e. $\hat{\gamma}=\tilde{\gamma}$ and $\gamma_0\in H$ hirshberg2019augmented. To use Lemma (ref) with classical kernel ridge balancing weights, allowing $\gamma_0\not \in H$, we require a rate $\mathcal{R}(\hat{\alpha})$, which is the main result of this paper (Theorem (ref)).

As written, Lemma (ref) applies to pointwise approximations of nonparametric causal functions such as heterogeneous treatment effects (Example (ref)). In such case, $\sigma=\sigma_h \asymp h^{-1/2}$, so the rate of Gaussian approximation is $(nh)^{-1/2}$. In practice, I take $h\asymp n^{-1/5}$ so the rate becomes $n^{-2/5}$; see Section (ref). This nonparametric rate is slower than the $n^{-1/2}$ obtained for regular cases such as ATE, similar to results of e.g. van2018cv,nie2021quasi,foster2023orthogonal,kennedy2023towards and many references therein, which could have been quoted instead. The point of Lemma (ref) is not to provide new semiparametric theory but to demonstrate how the main result of this paper (Theorem (ref)) makes general semiparametric guarantees immediate. See Appendix (ref) for more details on nonparametric causal functions.\footnote{When appealing to Theorem (ref) in the regular case, the rate for $\mathcal{R}(\hat{\alpha})$ is $r_n$. When appealing to Theorem (ref) for causal functions, the rate for $\mathcal{R}(\hat{\alpha}_h)$ is $h^{-2}r_n$.}

remark[Rate-optimal ridge regularization] Lemma (ref) does not require undersmoothing of the regularization in $\hat{\gamma}$ or $\hat{\alpha}$ in order to prove inference for $\hat{\theta}$, departing from e.g. hirshberg2019kernel,mou2023kernel. Instead, it requires $L_2$ rates for $\hat{\gamma}$ and $\hat{\alpha}$, the latter of which Theorem (ref) provides. These rates are obtained with rate-optimal ridge regularization when optimality is known, as formalized in Corollary (ref). When $b<\infty $ and $c=1$, optimal ridge regularization (up to log factors) is $\lambda=n^{-\frac{b}{b+1}}$. For comparison, when $b<\infty$, hirshberg2019kernel imposes $\lambda \ll n^{-1}$, which undersmooths. The regularization in hirshberg2019augmented is $\lambda\asymp n^{-1}$, and the authors “generally recommend” $\lambda=\bar{\sigma}^2 n^{-1}$, which undersmooths. The discussion following hirshberg2019augmented considers alternative choices, e.g. $\lambda\asymp n^{-\frac{b}{2}}$ when $b<\infty$ hirshberg2019kernel which seems to undersmooth for $b>1$.\footnote{Future work may characterize whether rate-optimal ridge regularization of KRRR is compatible with hirshberg2019augmented. hirshberg2019augmented demonstrates rate-optimal regularization for H\"older spaces rather than RKHSs.} See bruns2023augmented for a comprehensive and insightful discussion when $\hat{\gamma}=\tilde{\gamma}$ and $\hat{\alpha}=\tilde{\alpha}$.

The familiar doubly robust guarantee also holds, where either $\hat{\gamma}$ or $\hat{\alpha}$ may be mis-specified yet $\hat{\theta}$ remains consistent, albeit at a slower rate than $n^{-1/2}\sigma$. In particular, either $\|\gamma_{\mathcal{G}}-\gamma_0\|_2$ or $\|\alpha_{\mathcal{A}}-\alpha_0\|_2$ may be non-vanishing. Denote the mis-specified second moment as $\sigma_{\textsc{mis}}^2=\mathbb{E}\{\psi(W,\theta_0,\gamma_{\mathcal{G}},\alpha_{\mathcal{A}})^2\}$.

proposition[Consistency under mis-specification] Suppose Assumption (ref) holds. Suppose the following regularity conditions: $ \mathbb{E}[\{Y-\gamma_{\mathcal{G}}(W)\}^2 \mid W]\leq \bar{\sigma}^2$, $ \|\hat{\alpha}\|_{\infty}\leq\bar{\alpha}'$, and $\sigma_{\textsc{mis}} n^{-1/2}\rightarrow0 $. Suppose the following learning rate conditions: $\mathopen{}\mathclose\bgroup\originalleft(\bar{M}^{1/2}+\bar{\alpha}'\aftergroup\egroup\originalright)\mathcal{R}(\hat{\gamma})^{1/2}=o_p(1)$ and $\bar{\sigma}\mathcal{R}(\hat{\alpha})^{1/2}=o_p(1)$. If either $\gamma_{\mathcal{G}}=\gamma_0$ or $\alpha_{\mathcal{A}}=\alpha_0$, then $\hat{\theta}=\theta_0+o_p(1)$.

See Appendix (ref) for the proof and the exact nonasymptotic rate of convergence, which depends on $\mathcal{R}(\hat{\gamma})$ and $\mathcal{R}(\hat{\alpha})$. Previous works on kernel ridge balancing weights typically require $\gamma_0\in H$. By contrast, Proposition (ref) tolerates nonvanishing $\|\gamma_{\mathcal{G}}-\gamma_0\|_2$ or nonvanishing $\|\alpha_{\mathcal{A}}-\alpha_0\|_2$, where neither $\mathcal{G}$ nor $\mathcal{A}$ may be $H$. To achieve this robustness to mis-specification for KRRR, we require a rate $\mathcal{R}(\hat{\alpha})$, i.e. the main result of this paper.

Again, as written, Proposition (ref) applies to pointwise approximations of nonparametric causal functions. See Appendix (ref) for details.

Proposition (ref) summarizes the nonasymptotic Proposition (ref), which is a modest refinement of the celebrated double robustness guarantee. It is merely for illustrative purposes; stronger asymptotic results are available in the literature. For inference under mis-specification, see e.g. van2014targeted,benkeser2017doubly,dukes2021doubly.

Proof via the counterfactual effective dimension

I now prove the main result, drawing on integral operator techniques for kernel ridge regression from smale2007learning,caponnetto2007optimal,fischer2017sobolev and many important references therein. The main innovation is the counterfactual effective dimension.

The steps are as follows: (i) translate the desired statement from generalization error to $H$ norm (Lemma (ref)); (ii) characterize high probability events via concentration in $H$ (Proposition (ref)); (iii) use these high probability events to prove a nonasymptotic bound in terms of standard learning theory quantities (Proposition (ref)).

Among the high probability events is one that involves the counterfactual features $\phi^{(m)}(W)$. It pertains to the variance of the KRRR estimator $\tilde{\alpha}$. To justify this high probability event, I bound the counterfactual effective dimension of $\phi^{(m)}(W)$ in terms of the actual effective dimension of $\phi(W)$ (Lemma (ref)).

The main result (Theorem (ref)) then follows from simplifying the nonasymptotic bound (Proposition (ref)) using known bounds on the learning theory quantities that hold under smoothness (Assumption (ref)) and spectral decay (Assumption (ref)) conditions. The latter controls the actual effective dimension. I defer proofs of lemmas to Appendix (ref).

lemma[From generalization error to $H$ norm; c.f. Lemma 12 of fischer2017sobolev] Suppose the conditions of Lemma (ref) hold. Then $ \mathcal{R}(\tilde{\alpha})=\|T^{\frac{1}{2}}(\tilde{\alpha}-\alpha_H)\|^2_H. $
definition[Standard learning theory quantities] Define the residual $\mathcal{A}(\lambda)=\|T^{1/2}(\alpha_{\lambda}-\alpha_H)\|_H^2$, the reconstruction error $ \mathcal{B}(\lambda)= \|\alpha_{\lambda}-\alpha_H\|^2_H$, and the actual effective dimension $\mathcal{N}(\lambda)= Tr\{(T+\lambda)^{-1}T\} $.
definition[New learning theory quantity] Define the counterfactual effective dimension $\mathcal{N}^{(m)}(\lambda)=Tr\{(T+\lambda)^{-1}T^{(m)}\}$, where $T^{(m)}=\mathbb{E}[\phi^{(m)}(W)\otimes \{\phi^{(m)}(W)\}^*]$ is the counterfactual covariance operator.
lemma[New learning theory bound] Suppose the conditions of Lemma (ref) hold. Then $ \mathcal{N}^{(m)}(\lambda)\leq \bar{M} \cdot \mathcal{N}(\lambda). $
remark[Avoiding an auxiliary approximation assumption] Lemma (ref) demonstrates, by a simple argument, that a standard assumption in semiparametric theory (Assumption (ref)) implies that the counterfactual effective dimension (Definition (ref)) is not much greater than the actual effective dimension (Definition (ref)). This connection between semiparametric theory and learning theory appears to be new, and it powers the main result. Previous works that study generalization error, for different estimators of Riesz representers, appear to place auxiliary approximation assumptions to limit the counterfactual effective dimension, e.g. chernozhukov2018global, chernozhukov2018learning, and chernozhukov2020adversarial.
definition[High probability events in $H$] Let $\|\cdot\|_{\mathcal{L}_2(H,H)}$ be the Hilbert-Schmidt norm for operators that map from $H$ to $H$. Define the events { \begin{align*} \mathcal{E}_1&=\mathopen\mathclose\bgroup\originalleft[\|(T+\lambda)^{-1}(\hat{T}-T)\|_{\mathcal{L}_2(H,H)}\leq 2\ln(6/\delta)\mathopen\mathclose\bgroup\originalleft\{\frac{2\kappa}{\lambda n}+\sqrt{\frac{\kappa\mathcal{N}(\lambda)}{\lambda n}}\aftergroup\egroup\originalright\}\aftergroup\egroup\originalright], \\ \mathcal{E}_2&=\mathopen\mathclose\bgroup\originalleft[\|(T-\hat{T})(\alpha_{\lambda}-\alpha_H)\|_H\leq 2\ln(6/\delta)\mathopen\mathclose\bgroup\originalleft\{\frac{2\kappa\sqrt{\mathcal{B}(\lambda)}}{n}+\sqrt{\frac{\kappa\mathcal{A}(\lambda)}{n}}\aftergroup\egroup\originalright\} \aftergroup\egroup\originalright], \\ \mathcal{E}_3&=\mathopen\mathclose\bgroup\originalleft[\|(T+\lambda)^{-\frac{1}{2}}(\hat{\mu}^{(m)}-\hat{T}\alpha_H)\|_H \leq 2\ln(6/\delta)\mathopen\mathclose\bgroup\originalleft\{\frac{1}{n}\sqrt{\Upsilon^2\frac{\kappa'}{\lambda}}+\sqrt{\frac{\Sigma^2\mathcal{N}(\lambda)}{n}}\aftergroup\egroup\originalright\} \aftergroup\egroup\originalright], \end{align*} } where $ \Upsilon=2(1+\sqrt{\kappa}\|\alpha_H\|_H)$, $\Sigma^2=2(\bar{M}+\kappa \|\alpha_H\|^2_H)$, and $\kappa'=\max\{\kappa,\kappa^{(m)}\}. $
proposition[High probability events in $H$] Suppose the conditions of Lemma (ref) hold. Then $ \mathbb{P}(\mathcal{E}_j^c)\leq \delta/3$ for each $j\in\{1,2,3\}. $

The proof of Proposition (ref) appeals to Lemma (ref) in order to handle $\mathcal{E}_3$, which contains the counterfactual features via the counterfactual mean embedding $\hat{\mu}^{(m)}=\mathbb{E}_n\{\phi^{(m)}(W)\}$.

proposition[Nonasymptotic bound] Suppose the conditions of Lemma (ref) hold. If $n$ is sufficiently large that $n\geq 192 \ln^2(6/\delta) \kappa' \mathcal{N}(\lambda) \lambda^{-1}$ and $\lambda\leq \|T\|_{op}, $ then with probability $1-\delta$, $ \|T^{1/2}(\tilde{\alpha}-\alpha_H)\|^2_H\leq 96 \ln^2(6/\delta) \mathopen{}\mathclose\bgroup\originalleft\{\mathcal{A}(\lambda)+\frac{\kappa^2\mathcal{B}(\lambda)}{n^2\lambda}+\frac{\kappa \mathcal{A}(\lambda)}{ n\lambda}+\frac{\kappa' \Upsilon^2}{n^2\lambda}+\frac{\Sigma^2 \mathcal{N}(\lambda)}{n}\aftergroup\egroup\originalright\}. $

The proof of Proposition (ref) appeals to Proposition (ref) and the union bound.

Simulated and real data analysis

Nominal pointwise coverage of heterogeneous effects

I present coverage simulations for heterogeneous treatment effects (Example (ref)) viewed as a causal function. To lighten notation, I let $V=X_1$ and I write the heterogeneous treatment effects as $\textsc{cate}(v)$.

I follow the simulation design of abrevaya2015estimating, where $\textsc{cate}(v)=v(1+2v)^2(v-1)^2$ is the function on which we aim to conduct inference using observations of the outcome $Y$, binary treatment $D$, and covariates $X$. The covariate of interest $V\in [-0.5,0.5]$ is continuous. I conduct inference at the three locations $v^*\in\{-0.25,0,0.25\}$, corresponding to three causal parameters: $\theta_0=\textsc{cate}(-0.25)=-0.10$, $\theta_0=\textsc{cate}(-0.25)=-0.10$, and $\theta_0=\textsc{cate}(0.25)=-0.32$. Figure (ref) visualizes $\textsc{cate}(v)$ and the three values of $\theta_0$. See Appendix (ref) for details on the data generating process.

wrapfigure[wrapfigure omitted — 292 chars of source]

I evaluate confidence sets constructed from the KRRR estimator $\tilde{\alpha}$ (Definition (ref)) in the context of debiased machine learning $(\hat{\theta},\hat{\sigma}^2)$ (Definition (ref)). I consider three variations of $(\hat{\theta},\hat{\sigma}^2)$, corresponding to different choices of the nonparametric regression estimator $\hat{\gamma}$: lasso, random forest, and neural network. Previous works on classical kernel ridge balancing weights typically disallow $\gamma_0\not\in H$, and appear not to justify inference on heterogeneous treatment effects. This use of kernel ridge balancing weights is justified by the main result of this paper.

Following chernozhukov2021simple, I consider low dimensional and high dimensional specifications. In the low dimensional specification, the estimator $\hat{\gamma}$ uses $(D,X)$ as well as their interactions. In the high dimensional specification, $\hat{\gamma}$ uses fourth order polynomials. I focus here on the latter. See Appendix (ref) for the former.

Throughout, KRRR uses the Gaussian kernel for $X$, which corresponds to using all Hermite polynomials, appropriately weighted. See Appendix (ref) for implementation details.

An important choice is how to tune the bandwidth $h$ in $\textsc{local}\{(v-v^*)/h\}$. I follow the heuristic $h=c_h\hat{\sigma}_v n^{-1/5}$ of previous work, where $\hat{\sigma}^2_v$ is the sample variance of $V$. This heuristic satisfies the theoretical requirements $h=o(1)$ and $n^{-1/2}h^{-3/2}=o(1)$, and it implies a nonparametric rate of convergence of $(nh)^{-1/2}=n^{-2/5}$. The hyperparameter $c_h$ is chosen by the analyst. To evaluate robustness to tuning, I consider $c_h\in\{0.25, 0.50, 1.00\}$.

table*[table* omitted — 1,441 chars of source]

Tables (ref) and (ref) present results across 500 simulations. The initial columns denote the choice of $v^*$ and true value $\theta_0=\textsc{cate}(v^*)$. The next column is the hyperparameter value $c_h$. I report the average point estimate, average standard error, and 95% coverage across 500 simulations for a given choice of $(v^*,\theta_0,c_h)$.

Coverage is close to nominal and quite stable when $c_h\in\{0.25, 0.5\}$, across estimators and across locations $v^*$. When $c_h$ is too large, e.g. $c_h=1.00$, there is some under coverage for the random forest and neural network for the locations $v^*\in\{-0.25,0.00\}$.

Compared to results for lasso Riesz representers (LRR) chernozhukov2018learning, KRRR considerably reduces the absolute bias in this smooth design. See chernozhukov2021simple for values analogous to those in Table (ref). In particular at the location $v^*=0.25$, for lasso regression, KRRR gives absolute bias $(0.06,0.05,0.03)$ while LRR gives absolute bias $(0.09,0.13,0.11)$; for random forest regression, KRRR gives $(0.04,0.04,0.02)$ while LRR gives $(0.05,0.08,0.07)$; and for neural network regression, KRRR gives $(0.02,0.02,0.00)$ while LRR gives $(0.07,0.07,0.06)$.

Confidence sets for heterogeneous effects of 401(k)

Having justified the pointwise coverage of KRRR for heterogeneous treatment effects, I now use the procedure to quantify uncertainty for the heterogeneous effects of 401(k) eligibility on assets by age. The empirical analysis yields meaningful economic insights.

I follow the identification strategy of poterba1995. The authors assume that when 401(k) was introduced, workers ignored whether jobs offered 401(k) plans and made employment decisions based on income and other observable job characteristics. They assume 401(k) eligibility was as good as randomly assigned, conditional on covariates.

I use data from the 1991 US Survey of Income and Program Participation, following the sample selection and variable construction of chernozhukov2004effects. The outcome $Y$ is net financial assets defined as the sum of IRA balances, 401(k) balances, checking accounts, US saving bonds, other interest-earning accounts, stocks, mutual funds, and other interest-earning assets minus nonmortgage debt. The treatment $D$ is eligibility to enroll in a 401(k) plan. The covariates $X$ are age, income, years of education, family size, marital status, two-earner status, benefit pension status, IRA participation, and home-ownership. The data include $n = 9915$ observations.

As in the simulations, I consider a high dimensional specification and various choices of the nonparametric regression estimator $\hat{\gamma}$: lasso, random forest, and neural network. I use KRRR $\tilde{\alpha}$ with the Gaussian kernel. See Appendix (ref) for the low dimensional specification.

Figures (ref) and (ref) visualize point estimates and pointwise 95% confidence sets based on KRRR, in black. I find that 401(k) eligibility has positive and statistically significant effects that vary by age. The effects appear to be small and positive for 30 year olds (about \$2,500), and generally increasing, before plateauing for 45 to 60 year olds (about \$12,500). The effect is not statistically significant for 65 year olds.

In summary, under the stated assumptions, 401(k) eligibility seems to cause middle aged individuals to save about five times more than 30 years olds. Individuals in the former subpopulation tend to be at the peak of their earning potentials, while those in the latter subpopulation are closer to the beginning of their careers. The results are robust, across variations of the regression estimator and the specification.

For comparison, I also visualize the smooth nonparametric estimator of singh2020kernel, in grey. The confidence sets of this work corroborate their estimate, and quantify uncertainty. In particular, this paper's confidence sets suggest that the apparent dip in effects for 65 year olds is not statistically significant.

figure[figure omitted — 541 chars of source]

Discussion

Causal inference introduces a counterfactual effective dimension. Under standard regularity conditions on causal parameters from semiparametric theory, e.g. a propensity score bounded away from zero and one, I show that the counterfactual effective dimension is not much greater than the actual effective dimension. Using this technique, I prove population $L_2$ rates of generalization error for kernel ridge balancing weights, similar to kernel ridge regression. These rates connect kernel ridge balancing weights with general semiparametric results which tolerate $\gamma_0\not \in H$ and justify confidence sets for causal functions.

Future work may prove a comprehensive lower bound on $L_2$ rates for Riesz representers in the RKHS framework, and evaluate whether KRRR rates are optimal in general. Future work may also study benign overfitting for KRRR, similar to kernel ridge regression, using the counterfactual effective dimension technique developed in this paper.

\spacingset{1}

\DeclareRobustCommand{\VAN}[2]{#2} \DeclareRobustCommand{\VAN}[2]{#1}

thebibliography\bibitem[Abrevaya et al., 2015]{abrevaya2015estimating} Abrevaya, J., Hsu, Y.-C., and Lieli, R. P. (2015). \newblock Estimating conditional average treatment effects. \newblock {\em Journal of Business & Economic Statistics}, 33(4):485--505. \bibitem[Avagyan and Vansteelandt, 2021]{avagyan2021high} Avagyan, V. and Vansteelandt, S. (2021). \newblock High-dimensional inference for the average treatment effect under model misspecification using penalized bias-reduced double-robust estimation. \newblock {\em Biostatistics & Epidemiology}, 6(2):221--238. \bibitem[Bauer et al., 2007]{bauer2007regularization} Bauer, F., Pereverzev, S., and Rosasco, L. (2007). \newblock On regularization algorithms in learning theory. \newblock {\em Journal of Complexity}, 23(1):52--72. \bibitem[Ben-Michael et al., 2021]{ben2021balancing} Ben-Michael, E., Feller, A., Hirshberg, D. A., and Zubizarreta, J. R. (2021). \newblock The balancing act in causal inference. \newblock arXiv:2110.14831. \bibitem[Benkeser et al., 2017]{benkeser2017doubly} Benkeser, D., Carone, M., van der Laan, M., and Gilbert, P. B. (2017). \newblock Doubly robust nonparametric inference on the average treatment effect. \newblock {\em Biometrika}, 104(4):863--880. \bibitem[Bruns-Smith et al., 2023]{bruns2023augmented} Bruns-Smith, D., Dukes, O., Feller, A., and Ogburn, E. L. (2023). \newblock Augmented balancing weights as linear regression. \newblock arXiv:2304.14545. \bibitem[Caponnetto and De Vito, 2007]{caponnetto2007optimal} Caponnetto, A. and De Vito, E. (2007). \newblock Optimal rates for the regularized least-squares algorithm. \newblock {\em Foundations of Computational Mathematics}, 7(3):331--368. \bibitem[Chernozhukov et al., 2018]{chernozhukov2018original} Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W. K., and Robins, J. M. (2018). \newblock Double/debiased machine learning for treatment and structural parameters. \newblock {\em The Econometrics Journal}, 21(1):C1--C68. \bibitem[Chernozhukov and Hansen, 2004]{chernozhukov2004effects} Chernozhukov, V. and Hansen, C. (2004). \newblock The effects of 401(k) participation on the wealth distribution: An instrumental quantile regression analysis. \newblock {\em Review of Economics and Statistics}, 86(3):735--751. \bibitem[Chernozhukov et al., 2021]{chernozhukov2021automatic} Chernozhukov, V., Newey, W. K., Quintas-Martinez, V., and Syrgkanis, V. (2021). \newblock Automatic debiased machine learning via neural nets for generalized linear regression. \newblock {\em arXiv:2104.14737}. \bibitem[Chernozhukov et al., 2022a]{chernozhukov2018learning} Chernozhukov, V., Newey, W. K., and Singh, R. (2022a). \newblock Automatic debiased machine learning of causal and structural effects. \newblock {\em Econometrica}, 90(3):967--1027. \bibitem[Chernozhukov et al., 2022b]{chernozhukov2018global} Chernozhukov, V., Newey, W. K., and Singh, R. (2022b). \newblock Debiased machine learning of global and local parameters using regularized {R}iesz representers. \newblock {\em The Econometrics Journal}, 25(3):576--601. \bibitem[Chernozhukov et al., 2023]{chernozhukov2021simple} Chernozhukov, V., Newey, W. K., and Singh, R. (2023). \newblock A simple and general debiased machine learning theorem with finite-sample guarantees. \newblock {\em Biometrika}, 110(1):257--264. \bibitem[Chernozhukov et al., 2020]{chernozhukov2020adversarial} Chernozhukov, V., Newey, W. K., Singh, R., and Syrgkanis, V. (2020). \newblock Adversarial estimation of {R}iesz representers. \newblock arXiv:2101.00009. \bibitem[Craven and Wahba, 1978]{craven1978smoothing} Craven, P. and Wahba, G. (1978). \newblock Smoothing noisy data with spline functions: Estimating the correct degree of smoothing by the method of generalized cross-validation. \newblock {\em Numerische Mathematik}, 31(4):377--403. \bibitem[De Vito and Caponnetto, 2005]{de2005risk} De Vito, E. and Caponnetto, A. (2005). \newblock Risk bounds for regularized least-squares algorithm with operator-value kernels. \newblock Report MIT-CSAIL-TR-2005-031, MIT CSAIL. \bibitem[Dukes et al., 2021]{dukes2021doubly} Dukes, O., Vansteelandt, S., and Whitney, D. (2021). \newblock On doubly robust inference for double machine learning. \newblock {\em arXiv:2107.06124}. \bibitem[Fischer and Steinwart, 2020]{fischer2017sobolev} Fischer, S. and Steinwart, I. (2020). \newblock Sobolev norm learning rates for regularized least-squares algorithms. \newblock {\em Journal of Machine Learning Research}, 21:1--38. \bibitem[Foster and Syrgkanis, 2023]{foster2023orthogonal} Foster, D. J. and Syrgkanis, V. (2023). \newblock Orthogonal statistical learning. \newblock {\em The Annals of Statistics}, 51(3):879--908. \bibitem[Ghassami et al., 2021]{ghassami2021minimax} Ghassami, A., Ying, A., Shpitser, I., and Tchetgen Tchetgen, E. (2021). \newblock Minimax kernel machine learning for a class of doubly robust functionals. \newblock {\em arXiv:2104.02929}. \bibitem[Hazlett, 2020]{hazlett2020kernel} Hazlett, C. (2020). \newblock Kernel balancing: A flexible non-parametric weighting procedure for estimating causal effects. \newblock {\em Statistica Sinica}, 30(3):1155--1189. \bibitem[Hirshberg et al., 2019]{hirshberg2019kernel} Hirshberg, D. A., Maleki, A., and Zubizarreta, J. R. (2019). \newblock Minimax linear estimation of the retargeted mean. \newblock arXiv:1901.10296. \bibitem[Hirshberg and Wager, 2021]{hirshberg2019augmented} Hirshberg, D. A. and Wager, S. (2021). \newblock Augmented minimax linear estimation. \newblock {\em Annals of Statistics}, 49(6):3206--3227. \bibitem[Kallus, 2020]{kallus2020generalized} Kallus, N. (2020). \newblock Generalized optimal matching methods for causal inference. \newblock {\em Journal of Machine Learning Research}, 21(62):1--54. \bibitem[Kallus et al., 2021]{kallus2021causal} Kallus, N., Mao, X., and Uehara, M. (2021). \newblock Causal inference under unmeasured confounding with negative controls: A minimax learning approach. \newblock {\em arXiv:2103.14029}. \bibitem[Kennedy, 2023]{kennedy2023towards} Kennedy, E. H. (2023). \newblock Towards optimal doubly robust estimation of heterogeneous causal effects. \newblock {\em Electronic Journal of Statistics}, 17(2):3008--3049. \bibitem[Kimeldorf and Wahba, 1971]{kimeldorf1971some} Kimeldorf, G. and Wahba, G. (1971). \newblock Some results on {T}chebycheffian spline functions. \newblock {\em Journal of Mathematical Analysis and Applications}, 33(1):82--95. \bibitem[Li, 1986]{li1986asymptotic} Li, K.-C. (1986). \newblock Asymptotic optimality of {CL} and generalized cross-validation in ridge regression with application to spline smoothing. \newblock {\em The Annals of Statistics}, pages 1101--1112. \bibitem[Mou et al., 2023]{mou2023kernel} Mou, W., Ding, P., Wainwright, M. J., and Bartlett, P. L. (2023). \newblock Kernel-based off-policy estimation without overlap: Instance optimality beyond semiparametric efficiency. \newblock arXiv:2301.06240. \bibitem[Nie and Wager, 2021]{nie2021quasi} Nie, X. and Wager, S. (2021). \newblock Quasi-oracle estimation of heterogeneous treatment effects. \newblock {\em Biometrika}, 108(2):299--319. \bibitem[Pfanzagl, 1982]{pfanzagl1982lecture} Pfanzagl, J. (1982). \newblock Lecture notes in statistics. \newblock {\em Contributions to a General Asymptotic Statistical Theory}, 13. \bibitem[Poterba et al., 1995]{poterba1995} Poterba, J. M., Venti, S. F., and Wise, D. A. (1995). \newblock Do 401(k) contributions crowd out other personal saving? \newblock {\em Journal of Public Economics}, 58(1):1--32. \bibitem[Robins et al., 2007]{robins2007comment} Robins, J. M., Sued, M., Lei-Gomez, Q., and Rotnitzky, A. (2007). \newblock Comment: Performance of double-robust estimators when “inverse probability weights” are highly variable. \newblock {\em Statistical Science}, 22(4):544--559. \bibitem[Rotnitzky et al., 2021]{rotnitzky2021characterization} Rotnitzky, A., Smucler, E., and Robins, J. M. (2021). \newblock Characterization of parameters with a mixed bias property. \newblock {\em Biometrika}, 108(1):231--238. \bibitem[Sch{\"o}lkopf et al., 2001]{scholkopf2001generalized} Sch{\"o}lkopf, B., Herbrich, R., and Smola, A. (2001). \newblock A generalized representer theorem. \newblock In {\em International Conference on Computational Learning Theory}, pages 416--426. Springer. \bibitem[Singh, 2021]{singh2021debiased} Singh, R. (2021). \newblock Debiased kernel methods. \newblock {\em arXiv:2102.11076}. \bibitem[Singh et al., 2024]{singh2020kernel} Singh, R., Xu, L., and Gretton, A. (2024). \newblock Kernel methods for causal functions: dose, heterogeneous and incremental response curves. \newblock {\em Biometrika}, 111(2):497--516. \bibitem[Smale and Zhou, 2007]{smale2007learning} Smale, S. and Zhou, D.-X. (2007). \newblock Learning theory estimates via integral operators and their approximations. \newblock {\em Constructive Approximation}, 26(2):153--172. \bibitem[Smola et al., 2007]{smola2007hilbert} Smola, A., Gretton, A., Song, L., and Sch{\"o}lkopf, B. (2007). \newblock A {H}ilbert space embedding for distributions. \newblock In {\em International Conference on Algorithmic Learning Theory}, pages 13--31. \bibitem[Smucler et al., 2019]{smucler2019unifying} Smucler, E., Rotnitzky, A., and Robins, J. M. (2019). \newblock A unifying approach for doubly-robust $\ell_1$ regularized estimation of causal contrasts. \newblock arXiv:1904.03737. \bibitem[Speckman, 1979]{speckman1979minimax} Speckman, P. (1979). \newblock Minimax estimates of linear functionals in a {H}ilbert space. \newblock Unpublished manuscript. \bibitem[Sriperumbudur et al., 2010]{sriperumbudur2010relation} Sriperumbudur, B., Fukumizu, K., and Lanckriet, G. (2010). \newblock On the relation between universality, characteristic kernels and {RKHS} embedding of measures. \newblock In {\em International Conference on Artificial Intelligence and Statistics}, pages 773--780. \bibitem[Steinwart and Christmann, 2008]{steinwart2008support} Steinwart, I. and Christmann, A. (2008). \newblock {\em Support Vector Machines}. \newblock Springer Science & Business Media. \bibitem[Sutherland, 2017]{sutherland2017fixing} Sutherland, D. J. (2017). \newblock Fixing an error in {C}aponnetto and de {V}ito (2007). \newblock {\em arXiv:1702.02982}. \bibitem[van der Laan, 2014]{van2014targeted} van der Laan, M. J. (2014). \newblock Targeted estimation of nuisance parameters to obtain valid statistical inference. \newblock {\em The International Journal of Biostatistics}, 10(1):29--57. \bibitem[van der Laan and Rose, 2011]{van2011targeted} van der Laan, M. J. and Rose, S. (2011). \newblock {\em Targeted Learning: Causal Inference for Observational and Experimental Data}. \newblock Springer Series in Statistics. Springer. \bibitem[van der Laan et al., 2018]{van2018cv} van der Laan, M. J., Rose, S., Bibaut, A., and Luedtke, A. R. (2018). \newblock Cv-tmle for nonpathwise differentiable target parameters. \newblock In {\em Targeted Learning in Data Science: Causal Inference for Complex Longitudinal Studies}, pages 455--481. Springer. \bibitem[van der Laan and Rubin, 2006]{van2006targeted} van der Laan, M. J. and Rubin, D. (2006). \newblock Targeted maximum likelihood learning. \newblock {\em The International Journal of Biostatistics}, 2(1). \bibitem[Wang et al., 2021]{wang2021estimation} Wang, J., Wong, R. K. W., Yang, S., and Chan, K. C. G. (2021). \newblock Estimation of partially conditional average treatment effect by hybrid kernel-covariate balancing. \newblock {\em arXiv:2103.03437}. \bibitem[Wong and Chan, 2018]{wong2018kernel} Wong, R. K. W. and Chan, K. C. G. (2018). \newblock Kernel-based covariate functional balancing for observational studies. \newblock {\em Biometrika}, 105(1):199--213. \bibitem[Zhao, 2019]{zhao2019covariate} Zhao, Q. (2019). \newblock Covariate balancing propensity score by tailored loss functions. \newblock {\em Annals of Statistics}, 47(2):965--993. \bibitem[Zheng and van der Laan, 2011]{zheng2011cross} Zheng, W. and van der Laan, M. J. (2011). \newblock Cross-validated targeted minimum-loss-based estimation. \newblock In {\em Targeted Learning: Causal Inference for Observational and Experimental Data}, pages 459--474. Springer.

\spacingset{1.65}