Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
90,001 characters · 19 sections · 83 citation commands
Kernel Ridge Riesz Representers: Generalization, Mis-specification, and the Counterfactual Effective Dimension
\def\spacingset#1{ {#1}} \spacingset{1}
{\it Keywords:} causal inference, Gaussian approximation, heterogeneous treatment effect, reproducing kernel Hilbert space (RKHS), semiparametric efficiency
\spacingset{1.65}
Kernel balancing weights adapt kernel methods to quantify uncertainty on average effects, e.g. the average treatment effect when the treatment is randomly assigned conditional on covariates.\footnote{Kernel ridge balancing weights also apply to confidence intervals for the average treatment effects on the treated and untreated subpopulations, as well as subpopulations defined in terms of a discrete covariate taking a finite number of values. This paper justifies confidence sets for causal functions.} The central idea is to balance the feature-mapped covariates in the treated group with those in the untreated group, subject to regularization. The weights that achieve such balance, when applied to the observed outcomes, estimate the average treatment effect and appear in its confidence interval. These weights are widely used in social science via popular R packages, and they are classical in statistics speckman1979minimax.
The motivation for kernel balancing weights is to inherit the simplicity and flexibility of kernel methods. However, not all of the familiar properties of, say, kernel ridge regression have been fully articulated for classical kernel ridge balancing weights, e.g. population $L_2$ rates of convergence for generalization error, a simple standalone closed form solution, and robustness to mis-specification of features. Such properties would lead to strong semiparametric guarantees for causal parameters via well-known “targeting” and “debiasing” arguments, which tolerate some types of mis-specification and which provide Gaussian approximation for causal functions, justifying confidence sets with nominal coverage.
This paper's main contribution is a population $L_2$ rate controlling the generalization error of the classical kernel ridge balancing weights. The rate coincides with the lower bound of regression in the reproducing kernel Hilbert space (RKHS) framework. It holds under mis-specification of features. A technical innovation powers the result: I define the counterfactual effective dimension, and prove it is not much greater than the actual effective dimension. This technique may be of independent interest.
As a secondary contribution, this paper provides a standalone closed form solution for kernel ridge balancing weights that appears to have been previously unknown. This standalone closed form solution can be evaluated at locations different than the training data, allowing kernel ridge balancing weights to be used in “debiased” treatment effect estimators with sample splitting. Such estimators accommodate regression models outside of the RKHS, unlike many previous works on kernel ridge balancing weights.
Conceptually, I study the kernel ridge Riesz representer (KRRR) as an exact generalization of kernel ridge regression and kernel ridge balancing weights. This generalization formalizes the sense in which the population $L_2$ rate is sometimes optimal. For treatment effects, I verify inference with rate-optimal regularization, rather than undersmoothing, of the ridge penalty. The generalization also justifies the construction of confidence sets for causal functions, e.g. heterogeneous policy effects and heterogeneous treatment effects, from kernel ridge balancing weights.\footnote{Heterogeneous treatment effects are a function of a continuous covariate; see Section (ref).}
Section (ref) situates these contributions within the context of related work. Section (ref) formalizes the key assumptions. Section (ref) describes the procedure, derives its standalone closed form, and clarifies how it extends known procedures. Section (ref) theoretically justifies the procedure with population $L_2$ rates of generalization error, which imply strong semiparametric guarantees. Section (ref) proves the rates via the counterfactual effective dimension. Section (ref) demonstrates nominal coverage in heterogeneous treatment effect simulations, with less bias than some previous estimators, then quantifies uncertainty for the heterogeneous treatment effects of 401(k) eligibility on assets as a function of age. The effect appears to be five times stronger for 45-60 year olds, compared to 30 year olds.
I extend the conceptual framework of kernel balancing weights speckman1979minimax,hazlett2020kernel,wong2018kernel,zhao2019covariate,kallus2020generalized,hirshberg2019augmented,hirshberg2019kernel,chernozhukov2020adversarial. Works in this literature often focus on kernel ridge balancing weights as a vector that possibly converges in $\ell_2$ hirshberg2019augmented,hirshberg2019kernel, whereas I will study them as evaluations of a function that converges in population $L_2$. The former is the training error for a balancing weight that is defined in-sample; see Corollary (ref). By contrast, I analyze generalization error for a balancing weight that can interpolate and that can be evaluated out-of-sample. This distinction is important for stronger semiparametric conclusions: coverage guarantees for causal functions, sample splitting, and regression models beyond the RKHS; see Remark (ref).
bruns2023augmented prove equivalences among various formulations of the classical kernel ridge balancing weights, including a minimax formulation with ridge regularization. Regardless of nomenclature, many of the salient properties of these classical balancing weights, such as population $L_2$ rates of generalization error and a standalone closed form, appear to be contributions. I recover the known, nonasymptotic balance-variance trade-off.
The properties that I prove for KRRR generalize known results for kernel ridge regression. I obtain $L_2$ rates using integral operator techniques for kernel methods, drawing on smale2007learning,caponnetto2007optimal,fischer2017sobolev, among others. See fischer2017sobolev for references and a review of integral operator techniques and empirical process techniques used to analyze RKHS estimators in various norms, many of which rely on the effective dimension. My characterization of the counterfactual effective dimension appears to be new, and it avoids auxiliary approximation assumptions. The closed form solution of KRRR echoes classical solutions kimeldorf1971some,scholkopf2001generalized.
I also build on the insight from semiparametric theory that the balancing weights for the average treatment effect are a special case of Riesz representers for causal parameters, which can be directly estimated robins2007comment,avagyan2021high,chernozhukov2018global,hirshberg2019augmented,chernozhukov2018learning,smucler2019unifying,hirshberg2019kernel,chernozhukov2020adversarial. Similar to chernozhukov2018global,chernozhukov2018learning,chernozhukov2020adversarial, I provide population $L_2$ rates. A key point of departure is that I study KRRR, which is a different estimator more closely aligned with kernel ridge regression, classical kernel ridge balancing weights, and popular R packages. Another distinction is that I place standard RKHS assumptions, which facilitates comparison with known optimality theory, including rate-optimal regularization. Finally, I avoid some auxiliary approximation assumptions of those works by defining and analyzing the counterfactual effective dimension; see Remark (ref).
The contribution of this work is not to provide new semiparametric theory. Rather, I demonstrate that KRRR has good properties similar to kernel ridge regression that were not fully articulated, and hence it is compatible with strong existing results, e.g. zheng2011cross,chernozhukov2018original,rotnitzky2021characterization,kennedy2023towards, among many others. These strong results allow for some mis-specification and rate-optimal regularization, rather than the correct specification and undersmoothing studied by e.g. hirshberg2019kernel,mou2023kernel. See e.g. hirshberg2019augmented,ben2021balancing for detailed citations on balancing weights outside of the RKHS. See e.g. van2011targeted for detailed citations on semiparametrics.
An initial draft was circulated in February 2021 singh2021debiased. Prior to the current draft, wang2021estimation extend the kernel balancing weights of wong2018kernel to heterogeneous treatment effects and prove consistency under correct specification, without formal inference, a closed form, or $L_2$ rates for balancing weights. This paper's results are complementary: formal inference, some mis-specification, a closed form, and $L_2$ rates for balancing weights, within a framework that includes other estimands.\footnote{By formal inference, I mean justification for Gaussian approximation and nominal coverage guarantees.} Subsequent works extend kernel balancing weights to confounding bridges kallus2021causal,ghassami2021minimax and study other nonlinear function spaces chernozhukov2021automatic.
Consider the problem of estimating the average treatment effect (ATE) $\theta_0$ from $n$ independently and identicically distributed (i.i.d.) observations of the baseline covariates $X\in\mathbb{R}^{dim(X)}$, treatment $D\in\{0,1\}$, and outcome $Y \in\mathbb{R}$. When treatment assignment is as good as random conditional on $X$, the ATE can be expressed in two ways. First, the regression formulation is $\theta_0=\mathbb{E}\{\gamma_0(1,X)-\gamma_0(0,X)\}$, where $\gamma_0(D,X)=\mathbb{E}(Y|D,X)$ is the outcome regression. Second, the balancing weight formulation is $\theta_0=\mathbb{E}\{Y\alpha_0(D,X)\}$, where $\alpha_0(D,X)=D\pi_0(X)^{-1}-(1-D)\{1-\pi_0(X)\}^{-1}$ is the balancing weight expressed as a function and $\pi_0(X)=\mathbb{E}(D|X)$ is the propensity score. One may derive the latter from the former using the law of iterated expectations and the regularity condition that the propensity score $\pi_0$ is bounded away from zero and one.
These two formulations suggest estimation approaches. One option is to estimate the regression as $\hat{\gamma}$ then the ATE as $\mathbb{E}_n\{\hat{\gamma}(1,X)-\hat{\gamma}(0,X)\}$, where $\mathbb{E}_n(\cdot)$ is the empirical mean. Another is to estimate the balancing weight as $\hat{\alpha}$ then the ATE as $\mathbb{E}_n\{Y\hat{\alpha}(D,X)\}$. One may also combine these estimators into a so-called doubly robust procedure $\mathbb{E}_n[\hat{\gamma}(1,X)-\hat{\gamma}(0,X)+\hat{\alpha}(D,X)\{Y-\hat{\gamma}(D,X)\}]$. Each approach also has a sample splitting variant, where $(\hat{\gamma},\hat{\alpha})$ are estimated from, say $n/2$ held out observations, and the empirical mean is taken with respect to the remaining $n/2$ observations. For the sample splitting variant, $\hat{\alpha}$ must be a function that can be evaluated out-of-sample.
The sample splitting variant of the doubly robust procedure is sometimes called “debiased” machine learning chernozhukov2018original, which is closely related to targeted learning van2006targeted and one step estimation pfanzagl1982lecture, and it has quite general inferential guarantees for $\theta_0$. In particular, it allows for mis-specification of $\hat{\gamma}$ or $\hat{\alpha}$, and rate-optimal regularization of $\hat{\gamma}$ and $\hat{\alpha}$, as long as population $L_2$ rates for $\hat{\gamma}$ and $\hat{\alpha}$ are known, i.e. formal characterizations of $\|\hat{\gamma}-\gamma_0\|_2$ and $\|\hat{\alpha}-\alpha_0\|_2$ where $\|W\|_2=\{\mathbb{E}(W^2)\}^{1/2}$.
Previous work proposes a kernel ridge balancing weight $\tilde{\alpha}$. My main contribution is to analyze its generalization error $\|\tilde{\alpha}-\alpha_0\|_2$, achieving the well-known lower bound for regression in the RKHS framework, without auxiliary approximation assumptions. Once a population $L_2$ rate is articulated, general semiparametric guarantees from the literature become available. These guarantees justify Gaussian approximation for causal functions and some mis-specification, unlike many previous works on kernel ridge balancing weights.
The ATE is a special case of a causal parameter. This paper considers parameters of the form $\theta_0=\mathbb{E}\{m(W,\gamma_0)\}$, where $W$ concatenates observed variables, $\gamma_0(W)=\mathbb{E}(Y|W)$ is a regression, and the formula $m$ accommodates various parameters beyond the ATE. For ATE, $W=(D,X)$ and $m(w,f)=f(1,x)-f(0,x)$. I study a class defined by a regularity condition generalizing the idea that the propensity score is away from zero and one.
Intuitively, this condition imposes that if $f_1$ and $f_2$ are close, then their functionals $\mathbb{E}\{m(W,f_1)\}$ and $\mathbb{E}\{m(W,f_2)\}$ are close. In the ATE example, if $\pi_0(X)\in(1-\bar{\pi},\bar{\pi})$ almost surely, then $\bar{M}=2\{\bar{\pi}^{-1}+(1-\bar{\pi})^{-1}\}$.
Nonasymptotic analysis allows $\bar{M}$ to be a diverging sequence as the sample size increases. In particular, this class includes functionals of the form $f\mapsto \mathbb{E}\{m_h(W,f)\}$ where $m_h(w,f)=\ell_h(w)\check{m}(w,f)$. The weighting $\ell_h(w)$ is indexed by a vanishing bandwidth $h\downarrow 0$, which also indexes $\bar{M}=\bar{M}_h\uparrow \infty$. In this thought experiment, a sequence of semiparametric quantities, each satisfying Assumption (ref), provides a pointwise approximation to a nonparametric causal function evaluated at a point. I provide two concrete examples.
For further interpretation, suppose the the first covariate $X_1$ is continuous. In the finite sample, Example (ref) is the heterogeneous treatment effect for the subpopulation whose first covariate is local to $x_1^*$. As $h\downarrow 0$, Example (ref) converges to the heterogeneous treatment effect for the subpopulation whose first covariate value is $x_1^*$. Section (ref) gives a real world example where formal uncertainty quantification leads to economic insights.
Taking $\ell_h(w)=1$, Examples (ref) and (ref) reduce to the average policy effect and average treatment effect, respectively. The class defined by Assumption (ref) contains not only average effects but also heterogeneous effects, i.e. pointwise approximations of causal functions.
The balancing weight is a special case of a Riesz representer. It is well known that each parameter within the class defined by Assumption (ref) has both a regression formulation and a generalized balancing weight formulation.
For the rest of the paper, I will set aside the “correct specification” assumption that $\gamma_0\in \mathcal{G}$, where $\mathcal{G}$ is a known subset of $L_2$ that can be imposed in estimation of $\hat{\gamma}$. As such, there exists a unique Riesz representer $\alpha_0=\alpha_0^{\min}$. See Remark (ref) for a list of many works that place “correct specification” style assumptions requiring $\gamma_0$ to be within the RKHS.
For a functional $f\mapsto \mathbb{E}\{m_h(W,f)\}$ where $m_h(w,f)=\ell_h(w)\check{m}(w,f)$, its Riesz representer is $\alpha_h(w)=\ell_h(w)\check{\alpha}_0(w)$, where $\check{\alpha}_0(w)$ is the Riesz representer for $f\mapsto \mathbb{E}\{\check{m}(W,f)\}$.
In summary, previous works on classical kernel ridge balancing weights provide formal inference for a limited class of causal parameters, i.e. average effects. They also typically require correct specification of $\gamma_0$ within the RKHS. I study the broader class defined in Assumption (ref), which includes heterogeneous policy effects (Example (ref)) and heterogeneous treatment effects (Example (ref)). I take the agnostic stance of $\gamma_0\in L_2$. By analyzing generalization error $\|\tilde{\alpha}-\alpha_0\|_2$, general semiparametric guarantees become immediate.
The kernel ridge balancing weight algorithm conducts estimation in a reproducing kernel Hilbert space (RKHS) $H\subset L_2$. $H$ consists of functions of the form $f:\mathcal{W}\rightarrow\mathbb{R}$. I denote its symmetric, positive definite kernel by $k:\mathcal{W}\times\mathcal{W}\rightarrow \mathbb{R}$, and its feature map by $\phi:w\mapsto k(w,\cdot)$, so that $k(w,w')=\langle \phi(w),\phi(w') \rangle_H$ and $f(w)=\langle f,\phi(w) \rangle_H$ for $f\in H$.\footnote{Assumptions (ref) and (ref) below formalize weak regularity conditions.}
I use familiar spectral notation from kernel ridge regression analysis. Let $L:L_2\rightarrow L_2$ be the convolution operator $f\mapsto \int k(\cdot,w) f(w) \mathrm{d}\mathbb{P}(w)$. Since $L$ is self adjoint and positive definite, it has weakly decreasing countable eigenvalues $(\eta_j)$ and corresponding eigenfunctions $(\varphi_j)$. Define the space $ H^c=(f=\sum_{j=1}^{\infty}f_j\varphi_j:\sum_{j=1}^{\infty} f_j^2\eta_j^{-c} <\infty) $. In this notation, $H^0=L_2$ and $H^1=H$ when $k$ is characteristic sriperumbudur2010relation. The RKHS $H$ is the subset of $L_2$ for which higher order terms in the series $(\varphi_j)$ have a smaller contribution. By Mercer's theorem, the feature map is $ \phi(w)=\{\eta_j^{1/2} \varphi_j(w)\}^{\infty}_{j=1}$.
I also use the extended notation of chernozhukov2020adversarial, whose critical insight is that causal inference introduces counterfactual features. Since each $\varphi_j$ is in $L_2$, $m(w,\varphi_j)$ is well defined. I write the counterfactual feature map as $\phi^{(m)}(w)=\{\eta_j^{1/2} m(w,\varphi_j)\}^{\infty}_{j=1}$. The counterfactual kernel is then $k^{(m)}(w,w')=\langle \phi^{(m)}(w),\phi^{(m)}(w') \rangle_H$, and the counterfactual evaluation is $m(w,f)=\langle f, \phi^{(m)}(w)\rangle_H$ for $f\in H$. In the example of ATE, the formula is $m(w,f)=f(1,x)-f(0,x)$, so $ \phi^{(m)}(W_i)=[\eta_j^{1/2} \{\varphi_j(1,X_i)-\varphi_j(0,X_i)\}]^{\infty}_{j=1}=\phi(1,X_i)-\phi(0,X_i). $ The counterfactual feature map applied to observation $W_i$ replaces the observed treatment value $D_i$ with the counterfactual values $d=1$ and $d=0$.
I place two standard assumptions from kernel ridge regression analysis. The first, called the source condition, concerns an object I call $\alpha_H \in H$, which is the best RKHS approximation to the Riesz representer $\alpha_0\in L_2$.\footnote{For ATE, $\alpha_H\in\operatorname*{\arg\!\min}_{\alpha\in H} \|\alpha-\alpha_0\|_2$ and $\alpha_0(W)=D\pi_0(X)^{-1}-(1-D)\{1-\pi_0(X)\}^{-1}$.} This assumption helps to analyze bias. The second assumption is the spectral decay condition, which quantifies the effective dimension of the features $\phi(W)$ of the RKHS $H$, and helps to analyze variance. I denote by $\mathcal{P}(b,c)$ the class of distributions that satisfy these two assumptions.
By definition of $\alpha_H \in H$ as the best kernel approximation to $\alpha_0 \in L_2$, $c\geq 1$ automatically holds in Assumption (ref). A value $c>1$ means that $\alpha_H$ is a particularly smooth element of $H$. The $L_2$ rates in this paper do not improve beyond $c=2$, reflecting the well known saturation effect of Tikhonov regularization bauer2007regularization.
For any bounded kernel $k$, $b\geq 1$ automatically holds in Assumption (ref)(i) fischer2017sobolev. A value $b>1$ means that the RKHS has a lower effective dimension in light of the data distribution. The $L_2$ rates in this paper improve all the way to the parameteric rate as $b\rightarrow \infty$.
The upper bound, Assumption (ref)(i), applies to the Mat\'ern, Gaussian, and other kernels with spectral decay that is polynomial, exponential, or “better”. The lower bound, Assumption (ref)(ii), applies to the Mat\'ern kernel but not the Gaussian kernel. The main results use Assumption (ref)(i) only; Assumption (ref)(ii) is for comparison to optimality theory.
These assumptions generalize familiar assumptions from the analysis of Sobolev spaces. For example, take $\mathcal{W}=\mathbb{R}^p$ and denote by $\mathbb{H}_2^s$ the Sobolev space with $s>p/2$ square integrable derivatives, which is an RKHS with the Mat\'ern kernel. If the RKHS used in estimation is $H=\mathbb{H}_2^s$ and the best RKHS approximation to the Riesz representer satisfies $\alpha_H\in \mathbb{H}_2^{s_0}$, then $c=s_0/s$ and $\mathbb{H}_2^{s_0}=(\mathbb{H}_2^{s})^c$. In words, $c$ quantifies how smooth $\alpha_H$ is relative to its estimator $\tilde{\alpha}\in H$. In this RKHS, $b=2s/p$. As such, $b$ quantifies how smooth the estimator $\tilde{\alpha}\in H$ is relative to the ambient dimension, i.e. its effective dimension.
I study a classical kernel ridge estimator $\tilde{\alpha}$ of the Riesz representer $\alpha_0$. It turns out that $\tilde{\alpha}=[\mathbb{E}_n\{\phi(W)\otimes \phi(W)^*\}+\lambda]^{-1}\mathbb{E}_n\{\phi^{(m)}(W)\}$, where I use the outer product notation $\{\phi(W)\otimes \phi(W)^*\}(\cdot)=\phi(W)\langle \phi(W),\cdot \rangle_H$ and the ridge regularization notation $\lambda=\lambda I$ with identity $I:H\rightarrow H$. Since $\tilde{\alpha}$ involves counterfactual features $\phi^{(m)}(W)$, a notion of their effective dimension is necessary; to analyze the variance of the estimator $\tilde{\alpha}$, the counterfactual effective dimension is unavoidable.
By Lemma (ref), a familiar regularity condition from semiparametric theory (Assumption (ref)), together with a familiar effective dimension condition from learning theory (Assumption (ref)), implies control of the counterfactual effective dimension and hence the variance of $\tilde{\alpha}$. Due to this insight, I avoid an additional approximation assumption on the counterfactual effective dimension when proving the main result; see Remark (ref).
An important departure from many works on kernel ridge balancing weights is that I allow incorrect specification by the RKHS.\footnote{See Remark (ref) below.} In particular, I allow $\gamma_0 \not \in H$ and $\alpha_0 \not \in H$. For the scenario where $\alpha_0\not\in H$, I define the best kernel approximation $\alpha_H \in H$ to the Riesz representer $\alpha_0\in L_2$. I defer proofs for this section to Appendix (ref).
I place weak regularity conditions on the original spaces and on the RKHS $H$ that allow for further characterization of $\alpha_H$.
The space $\mathcal{W}$ accommodates general data types, e.g. texts, images, and graphs. Commonly used kernels are bounded. Boundedness implies that the counterfactual kernel mean embedding $\mu^{(m)}=\mathbb{E}\{\phi^{(m)}(W)\}$ and the uncentered covariance operator $T=\mathbb{E}\{\phi(W)\otimes \phi(W)^*\}$ are well defined.\footnote{Here, I use outer product notation: $(a\otimes b^*)c=a \langle b,c\rangle_H$, where $b^*$ is the adjoint of $b$.} This counterfactual kernel mean embedding generalizes the well known kernel mean embedding smola2007hilbert. It is unconditional and does not involve the outcome. Measurability is another standard condition.
The characteristic property ensures that no counterfactual information is lost when approximating $\alpha_0$ as $\alpha_H$. In particular, it means that $\mathbb{P}\mapsto \mu^{(m)}$ is injective, so $\mu^{(m)}$ may be used as a sufficient statistic for $\alpha_H$, as shown in the following lemma.
The sufficient statistics $\mu^{(m)}\in H$ and $T\in H\otimes H^*$ are infinite dimensional generalizations of the counterfactual moment and the covariance matrix in automatic debiased machine learning (Auto-DML) chernozhukov2018global,chernozhukov2018learning. Auto-DML uses $p$ explicit basis functions, and its sufficient statistic are a vector in $\mathbb{R}^p$ and a matrix in $\mathbb{R}^{p \times p}$. Kernel balancing weights use the countable eigenfunctions of the kernel $k$ as infinitely many implicit basis functions.
Inspired by Lemma (ref), I study a general loss for kernel ridge Riesz representers (KRRR) at the population level. As we will see below, its empirical analogue recovers the loss for kernel ridge balancing weights previously proposed for treatment effects. Let $\hat{\mu}^{(m)}=\mathbb{E}_n\{\phi^{(m)}(W)\}$ and $\hat{T}=\mathbb{E}_n\{\phi(W)\otimes \phi(W)^*\}$ be the empirical summary statistics.
The KRRR loss in Definition (ref) differs from Dantzig chernozhukov2018global, Lasso chernozhukov2018learning, and adversarial chernozhukov2020adversarial losses for Riesz representers whose generalization errors have been previously analyzed. The KRRR loss has a simple first order condition which implies a nonasymptotic balance-variance trade-off similar to those of kernel balancing weights in the literature hazlett2020kernel,wong2018kernel,zhao2019covariate,kallus2020generalized,hirshberg2019kernel.
By Definition (ref), a smaller regularization value $\lambda$ translates to an estimator $\tilde{\alpha}$ with a smaller bias and a larger variance. Corollary (ref) illustrates that a smaller (normalized) imbalance of the features, $\|\mathbb{E}_n[\tilde{\alpha}(W)\phi(W)-\phi^{(m)}(W)]\|_{H}/\|\tilde{\alpha}\|_{H}$, is possible if and only if we pay the cost of a larger variance. In the finite sample analysis to follow, $\lambda\downarrow 0$ in a way that optimally navigates this trade-off (when rate lower bounds are known). In particular, we will not require undersmoothing for inference on the causal parameter.
Lemma (ref) has several more corollaries, which formalize the sense in which KRRR builds on and extends previous frameworks. In these corollaries, let $\tilde{\gamma}$ be the kernel ridge regression estimator. Throughout, I assume that the conditions of Lemma (ref) hold.
Next, let $\tilde{\beta}\in\mathbb{R}^n$ be the kernel ridge balancing weight previously defined for average effects speckman1979minimax,kallus2020generalized,hirshberg2019augmented,hirshberg2019kernel.
In summary, KRRR nests kernel ridge regression and kernel ridge balancing weights, unlike the adversarial Riesz representer of chernozhukov2020adversarial. The connection to kernel ridge regression foreshadows the new generalization error guarantees for KRRR.
Next, I recap known equivalences for the estimation of treatment effects speckman1979minimax,kallus2020generalized,hirshberg2019kernel and more generally for the estimation of causal parameters within the class of Assumption (ref) bruns2023augmented. With $\hat{\alpha}$ taken to be KRRR $\tilde{\alpha}$, these equivalences only hold when the regression estimator $\hat{\gamma}$ is kernel ridge regression $\tilde{\gamma}$, and when the same observations are used for $\tilde{\gamma}$, $\tilde{\alpha}$, and the causal parameter.
When $\gamma_0\not \in H$, alternative regression estimators lead to better guarantees for semiparametric inference; see Section (ref) below.
In summary, the standalone closed form for KRRR is not necessary for causal estimation in the special case of Corollary (ref). However, Corollary (ref) shows that if $\hat{\gamma}\neq \tilde{\gamma}$, or if $\hat{\gamma}$ or $\tilde{\alpha}$ is estimated on a held out sample, then a standalone closed form for KRRR that can be evaluated away from the training data is necessary for causal estimation. I derive it below using techniques of chernozhukov2020adversarial, as a secondary contribution.
Even for the special case of Corollary (ref), this paper's primary contribution is the population $L_2$ rate $\|\tilde{\alpha}-\alpha_H\|_2$ for kernel ridge balancing weights (and more generally KRRR) in Section (ref). Generalization error is essential for appealing to strong semiparametric guarantees that allow some mis-specification and justify inference on causal functions.
I use some additional notation for features. Define the operator $\Phi:H\rightarrow \mathbb{R}^n$ with $i$th component $\langle \phi(W_i),\cdot \rangle_H$ and likewise the operator $\Phi^{(m)}:H\rightarrow \mathbb{R}^n$ with $i$th component $\langle \phi^{(m)}(W_i),\cdot \rangle_H$. Concatenate $\Phi$ and $\Phi^{(m)}$ as $\Psi:H\rightarrow \mathbb{R}^{2n}$, with adjoint $\Psi^*:\mathbb{R}^{2n}\rightarrow H$.
Whereas kernel ridge regression has a solution in $\mathbb{R}^n$ kimeldorf1971some,scholkopf2001generalized, KRRR has a solution in $\mathbb{R}^{2n}$ because of the counterfactual features. Intuitively, causal inference involves reasoning about $n$ actual observations and $n$ counterfactual observations.\footnote{The conjecture in hirshberg2019augmented that kernel ridge balancing weights may have a representation $\rho\in \mathbb{R}^n$, such that $\tilde{\alpha}=\Phi^*\rho$, does not appear to be supported in general.}
The familiar kernel matrix $K^{(1)}=\Phi\Phi^*\in\mathbb{R}^{n\times n}$ encodes inner products between actual observations. The extended kernel matrix $K=\Psi \Psi^* \in \mathbb{R}^{2n\times 2n}$ encodes inner products between actual observations and counterfactual observations. Each entry of $K$ is computed from the kernel $k$ and formula $m$ as discussed in Section (ref). For intuition, write $$ \Psi=
,\quad K=
=
. $$
The standalone closed form solution for kernel ridge balancing weights, and KRRR more generally, appears to have been previously unknown. Proposition (ref) demonstrates that it can be easily computed using only the formula $m$ and kernel $k$, without directly evaluating the feature maps $\phi(w)$ or $\phi^{(m)}(w)$. In this sense, it is a practical contribution for inference. See Appendix (ref) for its specialization to Example (ref).
I now provide the main result: population $L_2$ rates for KRRR that coincide with the optimal rate, where optimal rates are known. I study the class of Riesz representers defined by Assumption (ref), imposing familiar smoothness and spectral decay conditions for the RKHS setting defined by Assumptions (ref) and (ref). I also impose mild regularity conditions that help in the algorithm derivation, defined by Assumptions (ref) and (ref).
For clarity, I write the population $L_2$ rate for KRRR, estimated from the observations $(W_i)_{i\in[n]}$, in terms of $\mathcal{R}(\tilde{\alpha})=\mathbb{E}[\{\tilde{\alpha}(W)-\alpha_H(W)\}^2|(W_i)_{i\in[n]}]$. It is the generalization error for a test observation $W$, allowing for mis-specification.
I defer the proof to Section (ref). I defer the proofs of corollaries to Appendix (ref).
Theorem (ref) does not require Assumption (ref)(ii). Moreover, as an intermediate step, I prove a nonasymptotic bound (Proposition (ref)) without imposing Assumptions (ref) and (ref)(i).
When the RKHS is finite dimensional, i.e. $b=\infty$ in Assumption (ref), KRRR achieves the parametric rate $r_n=n^{-1}$. When the RKHS is infinite dimensional with at least polynomial spectral decay, i.e. $b<\infty$ in Assumption (ref), KRRR achieves a familiar rate $r_n=n^{-\frac{bc}{bc+1}}$. This rate depends on $c$ in Assumption (ref), which quantifies the smoothness of the best kernel approximation $\alpha_H \in H$ to the Riesz representer $\alpha_0 \in L_2$. If no additional smoothness of $\alpha_H$ is known, i.e. all we have is that $\alpha_H\in H$ by construction, then $c=1$ and the rate has an extra logarithmic factor.
The rate $r_n=n^{-\frac{bc}{bc+1}}$ generalizes the familiar rate from Sobolev analysis. For example, take $\mathcal{W}=\mathbb{R}^p$ and denote by $\mathbb{H}_2^s$ the Sobolev space with $s>p/2$ square integrable derivatives, as before. Suppose that KRRR $\tilde{\alpha}$ is estimated using the RKHS $H=\mathbb{H}_2^s$. Suppose that $\alpha_H\in \mathbb{H}_2^{s_0}$, i.e. the best kernel approximation to the Riesz representer $\alpha_0\in L_2$ has $s_0>s$ square integrable derivatives. The population $L_2$ rate for $\tilde{\alpha}$ is then $r_n=n^{-\frac{2s_0}{2s_0+p}}$. When $s_0=s$, which is a lower bound on $s_0$ by construction, $r_n=\ln(n)^{\frac{2s}{2s+p}}\cdot n^{-\frac{2s}{2s+p}}$. Theorem (ref) applies to settings beyond Sobolev spaces over Euclidean domains.
Just as KRRR generalizes kernel ridge regression, Theorem (ref) generalizes caponnetto2007optimal. The contribution is meaningful because KRRR also generalizes the classical kernel ridge balancing weights, for which generalization error was not fully articulated. In settings where lower bounds are known, the upper bounds in Theorem (ref) often achieve them.
To showcase the role of Theorem (ref) in semiparametric inference, I quote a result from the targeted and debiased machine learning literature. Unlike many previous works on kernel ridge balancing weights, inference allows $\gamma_0\not\in H$. It also allows for sample splitting, which is helpful when the regression estimator $\hat{\gamma}$ is not kernel ridge regression, and is instead estimated in a function space with high entropy. Theorem (ref) is directly compatible with the quoted result, unlike previous guarantees for kernel ridge balancing weights. In this sense, KRRR provides debiased kernel methods.
To state the result, denote the doubly robust estimator with sample splitting as $\hat{\theta}$, with the well known analytic variance estimator $\hat{\sigma}^2$.
Let $\gamma_{\mathcal{G}}$ be the best approximation to $\gamma_0$ in some space $\mathcal{G}$, which may not be an RKHS. Similarly, let $\alpha_{\mathcal{A}}$ be the best approximation to $\alpha_0$ in some space $\mathcal{A}$. Slightly abusing notation, I now let $\mathcal{R}(\hat{\gamma})=\mathbb{E}[\{\hat{\gamma}(W)-\gamma_{\mathcal{G}}(W)\}^2|(W_i)_{i\in[n]}]$ and $\mathcal{R}(\hat{\alpha})=\mathbb{E}[\{\hat{\alpha}(W)-\alpha_{\mathcal{A}}(W)\}^2|(W_i)_{i\in[n]}]$.\footnote{In Theorem (ref) above, $\mathcal{A}=H$. In Corollary (ref) above, $\mathcal{G}=H$.}
To quote the result, I introduce some additional notation. Let $\psi(W,\theta,\gamma,\alpha)=m(W,\gamma)+\alpha(W)\{Y-\gamma(W)\}-\theta$, so that $\psi_0(W)=\psi(W,\theta_0,\gamma_0,\alpha_0)$ is the influence function with moments $\sigma^2=\mathbb{E}\{\psi_0(W)^2\}$, $\zeta^3=\mathbb{E}\{\psi_0(W)^3\}$, and $\chi^4=\mathbb{E}\{\psi_0(W)^4\}$.
Using weak regularity conditions, population $L_2$ rates $\mathcal{R}(\hat{\gamma})$ and $\mathcal{R}(\hat{\alpha})$, and approximation errors $\|\gamma_{\mathcal{G}}-\gamma_0\|_2$ and $\|\alpha_{\mathcal{A}}-\alpha_0\|_2$, the estimator $\hat{\theta}$ is consistent and asymptotically Gaussian at the rate $n^{-1/2}\sigma$. Its confidence interval includes $\theta_0$ with probability approaching the nominal level. Importantly, $\hat{\gamma}$ does not have to be kernel ridge regression.
As written, Lemma (ref) applies to pointwise approximations of nonparametric causal functions such as heterogeneous treatment effects (Example (ref)). In such case, $\sigma=\sigma_h \asymp h^{-1/2}$, so the rate of Gaussian approximation is $(nh)^{-1/2}$. In practice, I take $h\asymp n^{-1/5}$ so the rate becomes $n^{-2/5}$; see Section (ref). This nonparametric rate is slower than the $n^{-1/2}$ obtained for regular cases such as ATE, similar to results of e.g. van2018cv,nie2021quasi,foster2023orthogonal,kennedy2023towards and many references therein, which could have been quoted instead. The point of Lemma (ref) is not to provide new semiparametric theory but to demonstrate how the main result of this paper (Theorem (ref)) makes general semiparametric guarantees immediate. See Appendix (ref) for more details on nonparametric causal functions.\footnote{When appealing to Theorem (ref) in the regular case, the rate for $\mathcal{R}(\hat{\alpha})$ is $r_n$. When appealing to Theorem (ref) for causal functions, the rate for $\mathcal{R}(\hat{\alpha}_h)$ is $h^{-2}r_n$.}
The familiar doubly robust guarantee also holds, where either $\hat{\gamma}$ or $\hat{\alpha}$ may be mis-specified yet $\hat{\theta}$ remains consistent, albeit at a slower rate than $n^{-1/2}\sigma$. In particular, either $\|\gamma_{\mathcal{G}}-\gamma_0\|_2$ or $\|\alpha_{\mathcal{A}}-\alpha_0\|_2$ may be non-vanishing. Denote the mis-specified second moment as $\sigma_{\textsc{mis}}^2=\mathbb{E}\{\psi(W,\theta_0,\gamma_{\mathcal{G}},\alpha_{\mathcal{A}})^2\}$.
See Appendix (ref) for the proof and the exact nonasymptotic rate of convergence, which depends on $\mathcal{R}(\hat{\gamma})$ and $\mathcal{R}(\hat{\alpha})$. Previous works on kernel ridge balancing weights typically require $\gamma_0\in H$. By contrast, Proposition (ref) tolerates nonvanishing $\|\gamma_{\mathcal{G}}-\gamma_0\|_2$ or nonvanishing $\|\alpha_{\mathcal{A}}-\alpha_0\|_2$, where neither $\mathcal{G}$ nor $\mathcal{A}$ may be $H$. To achieve this robustness to mis-specification for KRRR, we require a rate $\mathcal{R}(\hat{\alpha})$, i.e. the main result of this paper.
Again, as written, Proposition (ref) applies to pointwise approximations of nonparametric causal functions. See Appendix (ref) for details.
Proposition (ref) summarizes the nonasymptotic Proposition (ref), which is a modest refinement of the celebrated double robustness guarantee. It is merely for illustrative purposes; stronger asymptotic results are available in the literature. For inference under mis-specification, see e.g. van2014targeted,benkeser2017doubly,dukes2021doubly.
I now prove the main result, drawing on integral operator techniques for kernel ridge regression from smale2007learning,caponnetto2007optimal,fischer2017sobolev and many important references therein. The main innovation is the counterfactual effective dimension.
The steps are as follows: (i) translate the desired statement from generalization error to $H$ norm (Lemma (ref)); (ii) characterize high probability events via concentration in $H$ (Proposition (ref)); (iii) use these high probability events to prove a nonasymptotic bound in terms of standard learning theory quantities (Proposition (ref)).
Among the high probability events is one that involves the counterfactual features $\phi^{(m)}(W)$. It pertains to the variance of the KRRR estimator $\tilde{\alpha}$. To justify this high probability event, I bound the counterfactual effective dimension of $\phi^{(m)}(W)$ in terms of the actual effective dimension of $\phi(W)$ (Lemma (ref)).
The main result (Theorem (ref)) then follows from simplifying the nonasymptotic bound (Proposition (ref)) using known bounds on the learning theory quantities that hold under smoothness (Assumption (ref)) and spectral decay (Assumption (ref)) conditions. The latter controls the actual effective dimension. I defer proofs of lemmas to Appendix (ref).
The proof of Proposition (ref) appeals to Lemma (ref) in order to handle $\mathcal{E}_3$, which contains the counterfactual features via the counterfactual mean embedding $\hat{\mu}^{(m)}=\mathbb{E}_n\{\phi^{(m)}(W)\}$.
The proof of Proposition (ref) appeals to Proposition (ref) and the union bound.
I present coverage simulations for heterogeneous treatment effects (Example (ref)) viewed as a causal function. To lighten notation, I let $V=X_1$ and I write the heterogeneous treatment effects as $\textsc{cate}(v)$.
I follow the simulation design of abrevaya2015estimating, where $\textsc{cate}(v)=v(1+2v)^2(v-1)^2$ is the function on which we aim to conduct inference using observations of the outcome $Y$, binary treatment $D$, and covariates $X$. The covariate of interest $V\in [-0.5,0.5]$ is continuous. I conduct inference at the three locations $v^*\in\{-0.25,0,0.25\}$, corresponding to three causal parameters: $\theta_0=\textsc{cate}(-0.25)=-0.10$, $\theta_0=\textsc{cate}(-0.25)=-0.10$, and $\theta_0=\textsc{cate}(0.25)=-0.32$. Figure (ref) visualizes $\textsc{cate}(v)$ and the three values of $\theta_0$. See Appendix (ref) for details on the data generating process.
I evaluate confidence sets constructed from the KRRR estimator $\tilde{\alpha}$ (Definition (ref)) in the context of debiased machine learning $(\hat{\theta},\hat{\sigma}^2)$ (Definition (ref)). I consider three variations of $(\hat{\theta},\hat{\sigma}^2)$, corresponding to different choices of the nonparametric regression estimator $\hat{\gamma}$: lasso, random forest, and neural network. Previous works on classical kernel ridge balancing weights typically disallow $\gamma_0\not\in H$, and appear not to justify inference on heterogeneous treatment effects. This use of kernel ridge balancing weights is justified by the main result of this paper.
Following chernozhukov2021simple, I consider low dimensional and high dimensional specifications. In the low dimensional specification, the estimator $\hat{\gamma}$ uses $(D,X)$ as well as their interactions. In the high dimensional specification, $\hat{\gamma}$ uses fourth order polynomials. I focus here on the latter. See Appendix (ref) for the former.
Throughout, KRRR uses the Gaussian kernel for $X$, which corresponds to using all Hermite polynomials, appropriately weighted. See Appendix (ref) for implementation details.
An important choice is how to tune the bandwidth $h$ in $\textsc{local}\{(v-v^*)/h\}$. I follow the heuristic $h=c_h\hat{\sigma}_v n^{-1/5}$ of previous work, where $\hat{\sigma}^2_v$ is the sample variance of $V$. This heuristic satisfies the theoretical requirements $h=o(1)$ and $n^{-1/2}h^{-3/2}=o(1)$, and it implies a nonparametric rate of convergence of $(nh)^{-1/2}=n^{-2/5}$. The hyperparameter $c_h$ is chosen by the analyst. To evaluate robustness to tuning, I consider $c_h\in\{0.25, 0.50, 1.00\}$.
Tables (ref) and (ref) present results across 500 simulations. The initial columns denote the choice of $v^*$ and true value $\theta_0=\textsc{cate}(v^*)$. The next column is the hyperparameter value $c_h$. I report the average point estimate, average standard error, and 95% coverage across 500 simulations for a given choice of $(v^*,\theta_0,c_h)$.
Coverage is close to nominal and quite stable when $c_h\in\{0.25, 0.5\}$, across estimators and across locations $v^*$. When $c_h$ is too large, e.g. $c_h=1.00$, there is some under coverage for the random forest and neural network for the locations $v^*\in\{-0.25,0.00\}$.
Compared to results for lasso Riesz representers (LRR) chernozhukov2018learning, KRRR considerably reduces the absolute bias in this smooth design. See chernozhukov2021simple for values analogous to those in Table (ref). In particular at the location $v^*=0.25$, for lasso regression, KRRR gives absolute bias $(0.06,0.05,0.03)$ while LRR gives absolute bias $(0.09,0.13,0.11)$; for random forest regression, KRRR gives $(0.04,0.04,0.02)$ while LRR gives $(0.05,0.08,0.07)$; and for neural network regression, KRRR gives $(0.02,0.02,0.00)$ while LRR gives $(0.07,0.07,0.06)$.
Having justified the pointwise coverage of KRRR for heterogeneous treatment effects, I now use the procedure to quantify uncertainty for the heterogeneous effects of 401(k) eligibility on assets by age. The empirical analysis yields meaningful economic insights.
I follow the identification strategy of poterba1995. The authors assume that when 401(k) was introduced, workers ignored whether jobs offered 401(k) plans and made employment decisions based on income and other observable job characteristics. They assume 401(k) eligibility was as good as randomly assigned, conditional on covariates.
I use data from the 1991 US Survey of Income and Program Participation, following the sample selection and variable construction of chernozhukov2004effects. The outcome $Y$ is net financial assets defined as the sum of IRA balances, 401(k) balances, checking accounts, US saving bonds, other interest-earning accounts, stocks, mutual funds, and other interest-earning assets minus nonmortgage debt. The treatment $D$ is eligibility to enroll in a 401(k) plan. The covariates $X$ are age, income, years of education, family size, marital status, two-earner status, benefit pension status, IRA participation, and home-ownership. The data include $n = 9915$ observations.
As in the simulations, I consider a high dimensional specification and various choices of the nonparametric regression estimator $\hat{\gamma}$: lasso, random forest, and neural network. I use KRRR $\tilde{\alpha}$ with the Gaussian kernel. See Appendix (ref) for the low dimensional specification.
Figures (ref) and (ref) visualize point estimates and pointwise 95% confidence sets based on KRRR, in black. I find that 401(k) eligibility has positive and statistically significant effects that vary by age. The effects appear to be small and positive for 30 year olds (about \$2,500), and generally increasing, before plateauing for 45 to 60 year olds (about \$12,500). The effect is not statistically significant for 65 year olds.
In summary, under the stated assumptions, 401(k) eligibility seems to cause middle aged individuals to save about five times more than 30 years olds. Individuals in the former subpopulation tend to be at the peak of their earning potentials, while those in the latter subpopulation are closer to the beginning of their careers. The results are robust, across variations of the regression estimator and the specification.
For comparison, I also visualize the smooth nonparametric estimator of singh2020kernel, in grey. The confidence sets of this work corroborate their estimate, and quantify uncertainty. In particular, this paper's confidence sets suggest that the apparent dip in effects for 65 year olds is not statistically significant.
Causal inference introduces a counterfactual effective dimension. Under standard regularity conditions on causal parameters from semiparametric theory, e.g. a propensity score bounded away from zero and one, I show that the counterfactual effective dimension is not much greater than the actual effective dimension. Using this technique, I prove population $L_2$ rates of generalization error for kernel ridge balancing weights, similar to kernel ridge regression. These rates connect kernel ridge balancing weights with general semiparametric results which tolerate $\gamma_0\not \in H$ and justify confidence sets for causal functions.
Future work may prove a comprehensive lower bound on $L_2$ rates for Riesz representers in the RKHS framework, and evaluate whether KRRR rates are optimal in general. Future work may also study benign overfitting for KRRR, similar to kernel ridge regression, using the counterfactual effective dimension technique developed in this paper.
\spacingset{1}
\DeclareRobustCommand{\VAN}[2]{#2} \DeclareRobustCommand{\VAN}[2]{#1}
\spacingset{1.65}