EconBase
← Back to paper

Generalized Kernel Ridge Regression for Causal Inference with Missing-at-Random Sample Selection

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

92,797 characters · 19 sections · 66 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Generalized Kernel Ridge Regression for Causal Inference with Missing-at-Random Sample Selection

abstractI propose kernel ridge regression estimators for nonparametric dose response curves and semiparametric treatment effects in the setting where an analyst has access to a selected sample rather than a random sample; only for select observations, the outcome is observed. I assume selection is as good as random conditional on treatment and a sufficiently rich set of observed covariates, where the covariates are allowed to cause treatment or be caused by treatment---an extension of missingness-at-random (MAR). I propose estimators of means, increments, and distributions of counterfactual outcomes with closed form solutions in terms of kernel matrix operations, allowing treatment and covariates to be discrete or continuous, and low, high, or infinite dimensional. For the continuous treatment case, I prove uniform consistency with finite sample rates. For the discrete treatment case, I prove $\sqrt{n}$ consistency, Gaussian approximation, and semiparametric efficiency. Keywords: RKHS, dose response, doubly robust, missing-at-random

Introduction

In causal inference with observational data, an analyst may have access to a selected sample rather than a random sample, from which the analyst may wish to estimate the dose response curve of a continuous treatment or the scalar treatment effect of a discrete treatment. For selected observations, the outcome is observed; otherwise it is missing. This issue is also called outcome attrition, and it may afflict even experimental data collection. I study the setting where the dataset includes a sufficiently rich set of covariates to ameliorate the problem: conditional on treatment and these covariates, selection is as good as random. I allow the covariates of the selection mechanism to either cause treatment or be caused by treatment, an extension of the canonical setting called missing-at-random (MAR). I refer to the static sample selection case when covariates cause treatment. I refer to the dynamic sample selection case when some covariates cause treatment and other covariates are caused by treatment.

For this setting of data limitations (observational study, sample selection) and compensatory data access (sufficiently many and well-placed covariates in the causal graph), I propose a family of estimators for causal inference based on kernel ridge regression. The estimators are easy to compute from closed form expressions, though treatment and covariates are allowed to be discrete or continous and low, high, or infinite dimensional. For the continuous treatment case, I propose dose response curve estimators and prove uniform consistency with finite sample rates under standard RKHS assumptions. For the discrete treatment case, I propose treatment effect estimators with confidence intervals and prove $\sqrt{n}$-consistency, Gaussian approximation, and semiparametric efficiency under standard RKHS assumptions.

I extend the framework in three ways: (i) from causal parameters of the full population to those of subpopulations and alternative populations; (ii) from means to increments and distributions of counterfactual outcomes; (iii) from static sample selection to dynamic sample selection. Tables (ref) and (ref) summarize my contributions. All estimators and guarantees are new.

table[table omitted — 2,095 chars of source]
table[table omitted — 936 chars of source]

Preview

I preview the framework with a vignette about the static case. Suppose an analyst observes a random sample in which some observations are selected ($S=1$) and others are not ($S=0$). Selected observations are complete in the sense that they include outcome $Y$, treatment $D$, and covariates $X$. Unselected observations are incomplete in the sense that they include only treatment $D$ and covariates $X$. The challenge for causal inference is that the selection mechanism may be confounded, so naive treatment effect estimation with complete cases may be misleading.

The baseline goal is to infer causal parameters of the full population, though only complete observations from the selected population are accessible. Consider the causal parameter $ \theta^{ATE}_0(d):=\mathbb{E}\{Y^{(d)}\} $, where $Y^{(d)}$ is the potential outcome given the intervention $D=d$. With machine learning, one can estimate the conditional expectation function $ \gamma_0(1,d,x)=\mathbb{E}(Y|S=1,D=d,X=x) $ because selected observations include $(Y,D,X)$. It turns out that an appropriate formula for the causal parameter is of the form $ \theta^{ATE}_0(d)=\int \gamma_0(1,d,x)\mathrm{d}\mathbb{Q} $ where $\mathbb{Q}$ is a distribution that reweights $\gamma_0$ in order to adjust for confounding of the treatment assignment and sample selection mechanisms. This distribution $\mathbb{Q}$ may be estimated from both complete and incomplete observations.

As my first contribution, I prove that nonparametric estimation of a dose response curve from a selected sample can be reduced to the inner product of generalized kernel ridge regressions. Suppose $\gamma_0$ is correctly specified as a function in an RKHS $\mathcal{H}$ over selection $S$, treatment $D$, and covariates $X$. I estimate $\hat{\gamma}$ by classic kernel ridge regression. Next, I prove that the reweighting distribution $\mathbb{Q}$ can be represented by another function $\mu$ in the RKHS. The injective mapping $\mathbb{Q}\mapsto \mu$ is called the kernel mean embedding smola2007hilbert. I estimate $\hat{\mu}$ by a procedure that involves generalized kernel ridge regression. Finally, I take the inner product of $\hat{\gamma}$ and $\hat{\mu}$, which corresponds to taking the expectation with respect to the counterfactual distribution $\mathbb{Q}$: $\hat{\theta}^{ATE}(d)=\int \hat{\gamma}(1,d,x)\mathrm{d}\hat{\mathbb{Q}}=\langle \hat{\gamma},\hat{\mu} \rangle_{\mathcal{H}}$.

As my second contribution, I prove that semiparametric inference for a treatment effect from a selected sample can be reduced to combinations of generalized kernel ridge regressions as well. In particular, confidence intervals involve Riesz representers, which can be estimated by generalized kernel ridge regressions and sequential mean embeddings. I construct confidence intervals from the Riesz representers by bias correction and sample splitting. Across settings, I appeal to RKHS geometry to articulate smoothness and spectral decay approximation assumptions for the various generalized kernel ridge regressions. These approximation assumptions generalize the standard smoothness and spectral decay assumptions in RKHS learning theory.

The structure of the paper is as follows. Section (ref) describes related work. Section (ref) presents the assumptions, which are standard in RKHS learning theory. Section (ref) presents nonparametric and semiparametric theory for the case where covariates cause treatment. Section (ref) presents nonparametric and semiparametric theory for the case where covariates may cause or be caused by treatment. Section (ref) concludes. I extend nonparametric results to counterfactual distributions in Appendix (ref).

Related work

Sample selection with continuous treatment. I express dose response curves with static and dynamic sample selection as reweightings of an underlying conditional expectation function, extending the g-formula framework in biostatistics robins1986new and the partial means framework in econometrics newey1994kernel. In the static case, the reweighting is a simple marginal or conditional distribution. In the dynamic case, the reweighting is a product of intertemporal distributions. Importantly, I reweight for both treatment and selection mechanisms using covariates, a setting sometimes called double selection with missing-at-random (MAR) data rubin1976inference. To formulate dose response curves in this way, I generalize nonparametric identification theorems of huber2012identification,huber2014treatment. Existing work on continuous treatment studies uses parametric assumptions and proposes parametric estimators heckman1979sample,hausman1979attrition or uses conditional moment and rank restrictions and proposes series estimators das2003nonparametric. A machine learning dose response estimator for double selection does not appear to exist. I pose as a question for future work how to adapt the dose response estimators to other sample selection problems such as missing-not-at-random data (MNAR) scharfstein1999adjusting,rotnitzky1998semiparametric, also called non-ignorable attrition, where auxiliary information such as instrumental variables das2003nonparametric,huber2012identification and shadow variables d2010new,miao2015identification are used instead.

Sample selection with discrete treatment. I express treatment effects with static and dynamic sample selection as functionals of an underlying conditional expectation function. In the static case, the functionals are linear; in the dynamic case, they are nonlinear. Importantly, I adjust for both treatment and selection mechanisms negi2020doubly,bia2020double using covariates in a multiply robust moment function robins1994estimation,robins1995semiparametric. A rich literature provides results adjusting only for the selection mechanism in a doubly robust moment function robins1994estimation,bang2005doubly. In the static case, I derive the doubly robust moment function for a broad class of treatment effects, extending the ATE characterization of bia2020double. For the dynamic case, I quote the multiply robust moment function bia2020double. I combine the multiply robust moments with sample splitting bickel1982adaptive, following the principles of targeted maximum likelihood estimation (TMLE) van2006targeted and debiased machine learning (DML) chernozhukov2016locally,chernozhukov2018original,chernozhukov2018global,chernozhukov2018learning. I combine generalized kernel ridge regressions to estimate nonparametric quantities without explicit density estimation. I prove new, sufficiently fast nonparametric rates to verify abstract rate conditions from chernozhukov2018original,bia2020double,chernozhukov2021simple. I pose as a question for future work how to adapt the treatment effect estimators to other sample selection problems where data are MNAR rather than MAR.

RKHS in causal inference. Existing work incorporates the RKHS into static and dynamic causal inference without sample selection. nie2017quasi propose the R learner (named for robinson1988root) to nonparametrically estimate the heterogenous treatment effect of binary treatment. The authors prove mean square error rates, which foster2019orthogonal situate within a broader theoretical framework. singh2019kernel study the nonparametric instrumental variable problem and elucidate the role of conditional expectation operators and conditional mean embeddings in causal inference. kallus2020generalized proposes kernel optimal matching, a semiparametric estimator of average treatment on the treated with binary treatment, and proves $\sqrt{n}$ consistency. hirshberg2019minimax propose a semiparametric balancing weight estimator for average treatment on the treated and its variants, and provide bias aware confidence intervals. singh2020kernel,singh2020negative,singh2021workshop propose nonparametric and semiparametric estimators with closed form solutions based on kernel ridge regression for a broad variety of static and dynamic treatment effects including the dose response curve and heterogeneous treatment effect. It appears that this project is the first to propose RKHS algorithms for causal inference with sample selection.

RKHS and sample selection. In the machine learning literature, the term sample selection is used to refer to prediction problems in which the training set and testing set are drawn from different distributions and the analyst possibly has access to unlabeled (i.e. excluding outcome) observations from the testing set. A rich variety of nonparametric reweighting techniques have been developed using RKHS methods huang2006correcting,cortes2008sample,gretton2009covariate. By contrast, the version of the sample selection problem I study is causal. The goal is to estimate nonparametric dose response curves, semiparametric treatment effects, and counterfactual distributions. To forge a connection between the present work and the problem studied in the machine learning literature, I extend my results to the setting with distribution shift: using labeled and unlabeled observations drawn from a population $\mathbb{P}$ and unlabeled observations drawn from an alternative population $\tilde{\mathbb{P}}$, the analyst seeks to estimate causal parameters for the alternative population.

Counterfactual distributions. Previous work on counterfactual distributions considers static and dynamic settings without sample selection. Many papers study distributional generalizations of average treatment effect (ATE) or average treatment on the treated (ATT) for binary treatment dinardo1996labor,firpo2007efficient,imbens2009identification,cattaneo2010efficient,chernozhukov2013inference,muandet2020counterfactual. singh2020kernel,singh2021workshop present RKHS estimators for a richer class of static and dynamic treatment effects where treatment may be discrete or continuous. It appears that this project is the first to propose nonparametric algorithms for counterfactual distributions with sample selection.

Unlike previous work, I (i) provide a framework that encompasses nonparametric and semiparametric treatment effects in static and dynamic sample selection problems; (ii) propose new estimators with closed form solutions based on a new RKHS construction; and (iii) prove uniform consistency for the continuous treatment case and $\sqrt{n}$ consistency, Gaussian approximation, and semiparametric efficiency for the discrete treatment case via original arguments that handle low, high, or infinite dimensional data.

RKHS assumptions

I summarize RKHS notation, interpretation, and assumptions from cucker2002mathematical, steinwart2008support, singh2020kernel and singh2021workshop. A reproducing kernel Hilbert space (RKHS) $\mathcal{H}$ has elements that are functions $f:\mathcal{W}\rightarrow\mathbb{R}$ where $\mathcal{W}$ is a Polish space, i.e a separable and completely metrizable topological space. Let $k:\mathcal{W}\times\mathcal{W}\rightarrow \mathbb{R}$ be a function that is continuous, symmetric, and positive definite. I call $k$ the kernel, and I call $\phi:w\mapsto k(w,\cdot)$ the feature map. The kernel is the inner product of features in the sense that $k(w,w')=\langle \phi(w),\phi(w')\rangle_{\mathcal{H}}$. The RKHS is the closure of the span of the kernel steinwart2008support.

The key approximation assumptions for the RKHS are smoothness and spectral decay. I articulate the definition of the RKHS as well as these assumptions in terms of the convolution operator in which $k$ serves as the kernel, i.e. $ L:L_{\nu}^2(\mathcal{W})\rightarrow L_{\nu}^2(\mathcal{W}),\; f\mapsto \int k(\cdot,w)f(w)\mathrm{d}\nu(w) $, where $L^2_{\nu}(\mathcal{W})$ the space of square integrable functions with respect to measure $\nu$. By the spectral theorem, I express the eigendecomposition of the convolution operator $L$ as $ Lf=\sum_{j=1}^{\infty} \lambda_j\langle \varphi_j,f \rangle_{L^2_{\nu}(\mathcal{W})}\cdot \varphi_j $ where $(\lambda_j)$ are weakly decreasing eigenvalues and $(\varphi_j)$ are orthonormal eigenfunctions that form a basis of $L_{\nu}^2(\mathcal{W})$.

First, I define the RKHS in terms of the orthonormal basis $(\varphi_j)$. For comparison, I also define the more familiar space $L^2_{\nu}(\mathcal{W})$ in terms of this orthonormal basis. For any $f,g\in L_{\nu}^2(\mathcal{W})$, write $ f=\sum_{j=1}^{\infty}a_j\varphi_j $ and $ g=\sum_{j=1}^{\infty}b_j\varphi_j $. By cucker2002mathematical,

align*[align* omitted — 402 chars of source]

In summary, the RKHS $\mathcal{H}$ can be defined as the subset of $L^2_{\nu}(\mathcal{W})$ for which higher order terms in the series $(\varphi_j)$ have a smaller contribution, subject to $\nu$ satisfying the conditions of Mercer's theorem steinwart2012mercer.

Next, I define the smoothness and spectral decay assumptions in terms of the eigenvalues $(\lambda_j)$. Denote the statistical target $f_0$ that is estimated in the RKHS $\mathcal{H}$. A smoothness assumption on $f_0$ can be formalized as

equation[equation omitted — 195 chars of source]

In words, $f_0$ is well approximated by the leading terms in the series $(\varphi_j)$. A larger value of $c$ corresponds to a smoother target $f_0$. This condition is also called the source condition. A spectral decay assumption on $f_0$ can be formalized as $\lambda_j\asymp j^{-b}$. A larger value of $b$ corresponds to a faster rate of spectral decay and therefore a lower effective dimension. Both $(b,c)$ are joint assumptions on the kernel and data distribution smale2007learning,caponnetto2007optimal,carrasco2007linear.

To interpret smoothness and spectral decay, I recap the classic example of Sobolev spaces. Let $\mathcal{W}\subset \mathbb{R}^p$. Denote by $\mathbb{H}_2^{\nu}$ the Sobolev space with $\nu>\frac{p}{2}$ derivatives that are square integrable. This space is an RKHS generated by the Mat\`ern kernel. Suppose $\mathcal{H}=\mathbb{H}_2^{\nu}$ with $\nu>\frac{p}{2}$ is chosen as the RKHS for estimation. If $f_0\in \mathbb{H}_2^{\nu_0}$, then $c=\frac{\nu_0}{\nu}$ fischer2017sobolev; $c$ quantifies the smoothness of $f_0$ relative to $\mathcal{H}$. In this Sobolev space, $b=\frac{2\nu}{p}>1$ edmunds2008function. The effective dimension is increasing in the original dimension $p$ and decreasing in the degree of smoothness $\nu$.

Finally, I state the five assumptions I place in this paper, generalizing the standard RKHS learning theory assumptions for sample selection models.

Identification. I assume selection on observables and missingness-at-random (MAR) in order to express dose response curves as reweightings of the conditional expectation function of outcome given selection, treatment, and covariates. For distribution shift settings, I assume the shift is only in the distribution of selection, treatment, and covariates.

RKHS regularity. I construct an tensor product RKHS for the conditional expectation function that facilitates the analysis of partial means. The RKHS construction requires kernels that are bounded and characteristic. For incremental responses, I require kernels that are differentiable.

Original space regularity. I allow treatment $D$ and covariates $X$ to be discrete or continuous and low, high, or infinite dimensional. I require that each of these variables is supported on a Polish space. For simplicity, I assume outcome $Y$ is bounded.

Smoothness. I estimate the conditional expectation function $\gamma_0$ by kernel ridge regression and I estimate the kernel mean embedding $\mu$ by a kernel mean or generalized kernel ridge regression. To analyze the bias from ridge regularization, I assume smoothness of each object estimated by a (generalized) kernel ridge regression.

Effective dimension. For semiparametric inference, I assume an effective dimension condition in the RKHS construction. I unconver a double spectral robustness: some kernels in the RKHS construction may not satisfy the effective dimension condition, as long as other kernels do.

The initial four assumptions suffice for uniformly consistent nonparametric estimation. I instantiate these assumptions for static sample selection (Section (ref)), dynamic sample selection (Section (ref)), and counterfactual distributions (Appendix (ref)). The fifth assumption is not necessary for uniform consistency, but is necessary for semiparametric inference. I instantiate this assumption for static sample selection (Section (ref)) and dynamic sample selection (Section (ref)).

Static sample selection

Causal parameters

I denote an observation by $(SY,S,D,X)$, which encodes both the selected case $(S=1)$ and unselected case $(S=0)$. In this section, I study a basic sample selection model in which the selection mechanism is determined by baseline covariates only. In other words, covariates may cause treatment but may not be caused by treatment. I denote the counterfactual outcome $Y^{(d)}$ given a hypothetical intervention on treatment $D=d$. I consider the enitre class of static causal parameters studied by singh2020kernel.

definition[Causal parameters] I define the following dose response curves, incremental response curves, counterfactual distributions, and treatment effects. \begin{enumerate} • $\theta_0^{ATE}(d):=\mathbb{E}[Y^{(d)}]$ is the counterfactual mean outcome given intervention $D=d$ for the entire population. • $ \theta_0^{DS}(d,\tilde{\mathbb{P}}):=\mathbb{E}_{\tilde{\mathbb{P}}}[Y^{(d)}]$ is the counterfactual mean outcome given intervention $D=d$ for an alternative population with data distribution $\tilde{\mathbb{P}}$ (elaborated in Assumptions (ref) and (ref)). • $ \theta_0^{ATT}(d,d'):=\mathbb{E}[Y^{(d')}|D=d]$ is the counterfactual mean outcome given intervention $D=d'$ for the subpopulation who actually received treatment $D=d$. • $\theta_0^{CATE}(d,v):=\mathbb{E}[Y^{(d)}|V=v]$ is the counterfactual mean outcome given intervention $D=d$ for the subpopulation with subcovariate value $V=v$. \end{enumerate} Likewise I define incremental responses, e.g. $\theta_0^{\nabla:ATE}(d):=\mathbb{E}[\nabla_d Y^{(d)}]$ and $\theta_0^{\nabla:ATT}(d,d'):=\mathbb{E}[\nabla_{d'} Y^{(d')}|D=d]$. See Appendix (ref) for counterfactual distributions.

When treatment $D$ (and subcovariate $V$) are discrete, these causal parameters are treatment effects. When treatment $D$ (or subcovariate $V$) is continuous, these causal parameters are dose response curves. For the continuous treatment case, I additionally consider incremental response curves. For both cases, I consider counterfactual distributions. My results for means, increments, and distributions of potential outcomes immediately imply results for differences thereof. See singh2020kernel for further discussion of these causal parameters.

Identification

huber2012identification,huber2014treatment articulates distribution-free sufficient conditions under which average treatment effect with discrete treatment can be measured from observations $(SY,S,D,X)$ in a selected sample. I refer to this collection of sufficient conditions as static sample selection on observables and demonstrate that it identifies the entire class of causal parameters in Definition (ref).

assumption[Static sample selection on observables] Assume \begin{enumerate} • No interference: if $D=d$ then $Y=Y^{(d)}$. • Conditional exchangeability: $\{Y^{(d)}\}\raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} D |X$ and $\{Y^{(d)}\}\raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} S|D,X$. • Overlap: if $f(x)>0$ then $f(d|x)>0$ and if $f(d,x)>0$ then $\mathbb{P}(S=1|d,x)>0$; where $f(x)$, $f(d|x)$, and $f(d,x)$, are densities. \end{enumerate}

No interference (also called the stable unit treatment value assumption) rules out network effects (also called spillovers). Conditional exchangeability states that conditional on covariates, treatment assignment is as good as random. Moreover, conditional on treatment and covariates, selection is as good as random; outcomes are missing-at-random (MAR) rubin1976inference. Overlap ensures that there is no covariate stratum $X=x$ such that treatment has a restricted support, and there is no treatment-covariate stratum $(D,X)=(d,x)$ such that selection is impossible.

For $\theta_0^{DS}$, I also make a standard assumption in transfer learning.

assumption[Static distribution shift] Assume \begin{enumerate} • $\tilde{\mathbb{P}}(Y,S,D,X)=\mathbb{P}(Y|S,D,X)\tilde{\mathbb{P}}(S,D,X)$; • $\tilde{\mathbb{P}}(S,D,X)$ is absolutely continuous with respect to $\mathbb{P}(S,D,X)$. \end{enumerate}

The difference in population distributions $\mathbb{P}$ and $\tilde{\mathbb{P}}$ is only in the distribution of selection, treatment, and covariates. Moreover, the support of $\mathbb{P}$ contains the support of $\tilde{\mathbb{P}}$. An immediate consequence is that the regression function for complete cases $\mathbb{E}[Y|S=1,D=d,X=x]$ remains the same across the different populations $\mathbb{P}$ and $\tilde{\mathbb{P}}$.

I present a general identification result below. Whereas huber2012identification,huber2014treatment identifies $\theta_0^{ATE}(d)$, I identify the entire class of causal parameters in Definition (ref), which includes many new results. I prove that each causal parameter is a reweighting of the regression using complete cases, i.e. $ \mathbb{E}(Y|S=1,D=d,X=x) $. This nuance is critical, since the data limitations of the problem setting imply that $\mathbb{E}(Y|S=1,D=d,X=x)$ is estimable while $\mathbb{E}(Y|S=0,D=d,X=x)$ is not.

remark[Regression notation] To maintain the semantics of the observation model $(SY,S,D,X)$, I formally define $\gamma_0(s,d,x):=\mathbb{E}[SY|S=s,D=d,X=x]$. Because $\gamma_0(1,d,x)=\mathbb{E}[SY|S=1,D=d,X=x]=\mathbb{E}(Y|S=1,D=d,X=x)$, I use $\mathbb{E}[SY|S=1,D=d,X=x]$ and $\mathbb{E}(Y|S=1,D=d,X=x)$ interchangeably.
theorem[Identification of causal parameters under static sample selection] If Assumption (ref) holds then \begin{enumerate} • $\theta_0^{ATE}(d)=\int \gamma_0(1,d,x)\mathrm{d}\mathbb{P}(x)$ huber2012identification,huber2014treatment. • If in addition Assumption (ref) holds, then $\theta_0^{DS}(d,\tilde{\mathbb{P}})=\int \gamma_0(1,d,x)\mathrm{d}\tilde{\mathbb{P}}(x)$. • $\theta_0^{ATT}(d,d')=\int \gamma_0(1,d',x)\mathrm{d}\mathbb{P}(x|d)$. • $\theta_0^{CATE}(d,v)=\int \gamma_0(1,d,v,x)\mathrm{d}\mathbb{P}(x|v)$. \end{enumerate} For $\theta_0^{CATE}$, I use $\gamma_0(1,d,v,x)=\mathbb{E}[Y|S=1,D=d,V=v,X=x] $. Likewise I identify incremental effects, e.g. $\theta_0^{\nabla:ATE}(d)=\int \nabla_d \gamma_0(1,d,x)\mathrm{d}\mathbb{P}(x)$. See Appendix (ref) for counterfactual distributions.

See Appendix (ref) for the proof.

Algorithm: Continuous treatment

Next, I propose nonparametric algorithm for the continuous treatment case and semiparametric algorithms for the discrete treatment case. In preparation, I propose the following RKHS construction. Define scalar valued RKHSs for selection $S$, treatment $D$, and covariates $X$. For example, the RKHS for selection is $\mathcal{H}_{\mathcal{S}}$ with feature map $\phi(s)=\phi_{\mathcal{S}}(s)$. I construct the tensor product RKHS $\mathcal{H}=\mathcal{H}_{\mathcal{S}}\otimes \mathcal{H}_{\mathcal{D}} \otimes \mathcal{H}_{\mathcal{X}}$ with feature map $\phi(s)\otimes \phi(d) \otimes \phi(x)$. In this construction, the kernel for $\mathcal{H}$ is $k(s,d,x;s',d',x')=k_\mathcal{S}(s,s')\cdot k_\mathcal{D}(d,d')\cdot k_{\mathcal{X}}(x,x')$.

For the continuous treatment case, I assume $\gamma_0\in\mathcal{H}$. Therefore by the reproducing property $ \gamma_0(s,d,x)=\langle \gamma_0, \phi(s)\otimes \phi(d)\otimes \phi(x)\rangle_{\mathcal{H}} $. Likewise for $\theta_0^{CATE}$, with the extended defintion $\mathcal{H}=\mathcal{H}_{\mathcal{S}}\otimes \mathcal{H}_{\mathcal{D}} \otimes \mathcal{H}_{\mathcal{V}}\otimes \mathcal{H}_{\mathcal{X}}$. I place regularity conditions on this RKHS construction in order to represent the causal parameters as inner products in $\mathcal{H}$. In anticipation of Section (ref) and Appendix (ref), I include conditions for a follow-up covariate RKHS $\mathcal{H}_{\mathcal{M}}$ and an outcome RKHS $\mathcal{H}_{\mathcal{Y}}$ in parentheses.

assumption[RKHS regularity conditions] Assume \begin{enumerate} • $k_{\mathcal{S}}$ is bounded. $k_{\mathcal{D}}$, $k_{\mathcal{V}}$, $k_{\mathcal{X}}$ (and $k_{\mathcal{M}}$, $k_{\mathcal{Y}}$) are continuous and bounded. Formally, $ \sup_{s\in\mathcal{S}}\|\phi(s)\|_{\mathcal{H}_{\mathcal{S}}}\leq \kappa_s$, $ \sup_{d\in\mathcal{D}}\|\phi(d)\|_{\mathcal{H}_{\mathcal{D}}}\leq \kappa_d$, $ \sup_{v\in\mathcal{V}}\|\phi(v)\|_{\mathcal{H}_{\mathcal{V}}}\leq \kappa_v$, $ \sup_{x\in\mathcal{X}}\|\phi(x)\|_{\mathcal{H}_{\mathcal{X}}}\leq \kappa_x $ \{and $ \sup_{m\in\mathcal{M}}\|\phi(m)\|_{\mathcal{H}_{\mathcal{M}}}\leq \kappa_m$, $ \sup_{y\in\mathcal{Y}}\|\phi(y)\|_{\mathcal{H}_{\mathcal{Y}}}\leq \kappa_y$\}. • $\phi(s)$, $\phi(d)$, $\phi(v)$, $\phi(x)$ \{and $\phi(m)$, $\phi(y)$\} are measurable. • $k_{\mathcal{X}}$ (and $k_{\mathcal{M}}$, $k_{\mathcal{Y}}$) are characteristic. \end{enumerate} For incremental responses, further assume $\mathcal{D}\subset \mathbb{R}$ is an open set and $\nabla_d \nabla_{d'} k_{\mathcal{D}}(d,d')$ exists and is continuous, hence $\sup_{d\in\mathcal{D}}\|\nabla_d\phi(d)\|_{\mathcal{H}}\leq \kappa_d'$. For the semiparametric case, take $k_{\mathcal{D}}(d,d')=\operatorname{\mathbbm 1}_{d=d'}$ instead.

Popular kernels are continuous and bounded, and measurability is a weak condition. The characteristic property ensures injectivity of the mean embeddings sriperumbudur2010relation, so that the encoding of the reweighting distributions will be without loss. For example, the Gaussian kernel is characteristic over a continuous domain.

theorem[Representation of dose and incremental response curves via kernel mean embeddings] Suppose the conditions of Theorem (ref) hold. Further suppose Assumption (ref) holds and $\gamma_0\in\mathcal{H}$. Then \begin{enumerate} • $\theta_0^{ATE}(d)=\langle \gamma_0, \phi(1)\otimes \phi(d)\otimes \mu_x\rangle_{\mathcal{H}} $ where $\mu_x:=\int\phi(x) \mathrm{d}\mathbb{P}(x) $. • $\theta_0^{DS}(d,\tilde{\mathbb{P}})=\langle \gamma_0, \phi(1)\otimes\phi(d)\otimes \nu_x\rangle_{\mathcal{H}} $ where $\nu_x:=\int\phi(x) \mathrm{d}\tilde{\mathbb{P}}(x) $. • $\theta_0^{ATT}(d,d')=\langle \gamma_0, \phi(1)\otimes\phi(d')\otimes \mu_x(d)\rangle_{\mathcal{H}} $ where $\mu_x(d):=\int\phi(x) \mathrm{d}\mathbb{P}(x|d)$. • $\theta_0^{CATE}(d,v)=\langle \gamma_0, \phi(1)\otimes \phi(d)\otimes \phi(v)\otimes \mu_{x}(v)\rangle_{\mathcal{H}} $ where $\mu_{x}(v):= \int \phi(x) \mathrm{d}\mathbb{P}(x|v)$. \end{enumerate} Likewise for incremental responses, e.g. $\theta_0^{\nabla:ATE}(d)=\langle \gamma_0, \phi(1)\otimes\nabla_d\phi(d)\otimes \mu_x\rangle_{\mathcal{H}} $. See Appendix (ref) for counterfactual distributions.

See Appendix (ref) for the proof. The kernel mean embedding $\mu_x:=\int\phi(x) \mathrm{d}\mathbb{P}(x)$ encodes the distribution $\mathbb{P}(x)$ as an element $\mu_x\in\mathcal{H}_{\mathcal{X}}$ smola2007hilbert. The inner product representation suggests estimators that are inner products, e.g. $\hat{\theta}^{ATE}(d)=\langle \hat{\gamma}, \phi(1)\otimes \phi(d)\otimes \hat{\mu}_x\rangle_{\mathcal{H}}$ where $\hat{\gamma}$ is a standard kernel ridge regression using the selected sample and $\hat{\mu}_x$ is an empirical mean. When a conditional distribution is encoded by a conditional mean embedding such as $\mu_x(d):=\int\phi(x) \mathrm{d}\mathbb{P}(x|d)$, I estimate $\hat{\mu}_x(d)$ by a generalized kernel ridge regression.

algorithm[algorithm omitted — 1,637 chars of source]

See Appendix (ref) for the derivation. See Theorem (ref) for theoretical values of $(\lambda,\lambda_1,\lambda_2)$ that balance bias and variance. See Appendix (ref) for a practical tuning procedure based on the closed form solution of leave-one-out cross validation to empirically balance bias and variance.

Algorithm: Discrete treatment

So far, I have presented nonparametric algorithms for the continuous treatment cases. Now I present semiparametric algorithms for the discrete treatment case. Towards this end, I characterize the doubly robust moment function for the static sample selection model, generalizing the result of bia2020double to treatment effects defined for subpopulations and alternative populations.

theorem[Doubly robust moment for static sample selection] Suppose treatment $D$ and subcovariate $V$ are discrete. If Assumption (ref) holds then each treatment effect can be expressed as \begin{align*} \theta^{(m)}_0 &=\mathbb{E}\bigg[ m(1,D,X;\gamma_0) + \alpha^{(m)}_0(S,D,X)\{SY-\gamma_0(1,D,X)\} \bigg], \end{align*} where $\gamma_0(1,D,X)=\mathbb{E}[Y|S=1,D,X]$ is the regression for complete cases, $\gamma\mapsto \mathbb{E}[m(1,D,X;\gamma)]$ is a linear functional for complete cases, and $\alpha^{(m)}_0(S,D,X)$ is the Riesz representer to the functional. In particular, \begin{enumerate} • For $\theta^{(m)}_0=\theta_0^{ATE}(d)$, $m(1,D,X;\gamma)=\gamma(1,d,X)$ and $\alpha_0^{ATE}(S,D,X)=\frac{S\cdot \mathbbm{1}_{D=d}}{\mathbb{P}(S=1|d,X)\mathbb{P}(d|X)}$ bia2020double. • If in addition Assumption (ref) holds, then for $\theta^{(m)}_0=\theta_0^{DS}(d,\tilde{\mathbb{P}})$, $m(1,D,X;\gamma)=\gamma(1,d,X)$ and $\alpha_0^{DS}(S,D,X)=\frac{S\cdot \mathbbm{1}_{D=d}}{\tilde{\mathbb{P}}(S=1|d,X)\tilde{\mathbb{P}}(d|X)}$. • For $\theta_0^{ATT}(d,d')=\frac{\theta^{(m)}_0}{\mathbb{P}(d)}$, $m(1,D,X;\gamma)=\gamma(1,d',X)\mathbbm{1}_{D=d}$ and $\alpha_0^{ATT}(S,D,X)=\frac{S\cdot \mathbbm{1}_{D=d'}\mathbb{P}(d|X)}{\mathbb{P}(S=1|d',X)\mathbb{P}(d'|X)}$. • For $\theta_0^{CATE}(d,v)=\frac{\theta^{(m)}_0}{\mathbb{P}(v)}$, $m(1,D,V,X;\gamma)=\gamma(1,d,v,X)\mathbbm{1}_{V=v}$ and $\alpha_0^{CATE}(S,D,V,X)=\frac{S\cdot \mathbbm{1}_{D=d}\mathbbm{1}_{V=v}}{\mathbb{P}(S=1|d,v,X)\mathbb{P}(d|v,X)}$. \end{enumerate}

See Appendix (ref) for the proof. Note that inference for the abstract treatment effect $\theta^{(m)}_0$ implies inference for $\theta_0^{ATT}(d,d')$ and $\theta_0^{CATE}(d,v)$ by delta method.

The equation in Theorem (ref) has three important properties. First, it consists of estimable quantities given the data limitations of the static sample selection problem. Second, it is doubly robust with respect to the nonparametric objects $(\gamma_0,\alpha^{(m)}_0)$ in the sense that it continues to hold if one of the nonparametric objects is misspecified. For example, for any $\gamma$,

align*[align* omitted — 128 chars of source]

Double robustness implies that the semiparametric guarantees below hold even if one of the nonparametric objects $(\gamma_0,\alpha^{(m)}_0)$ is not actually an element of an RKHS. Third, double robustness is with respect to $\alpha^{(m)}_0$ rather than the treatment and selection propensity scores. The technique of the kernel ridge Riesz representer permits estimation of $\alpha^{(m)}_0$ without estimation of the treatment and selection propensity scores singh2020kernel. I adapt the kernel ridge Riesz representer next.

algorithm[algorithm omitted — 2,014 chars of source]

The meta algorithm that combines the doubly robust moment function from Theorem (ref) with sample splitting is called DML for static sample selection chernozhukov2018original,bia2020double,chernozhukov2021simple. I instantiate DML for static sample selection using the new nonparametric estimator from Algorithm (ref).

algorithm[algorithm omitted — 1,147 chars of source]

Unlike bia2020double, the proposed procedure does not require explicit propensity score estimation. Note that $\hat{\alpha}^{(m)}$ is computed according to Algorithm (ref) with regularization parameter $\lambda_3$. $\hat{\gamma}$ is a standard kernel ridge regressions with regularization parameter $\lambda$. See Appendix (ref) for explicit computations. See Theorem (ref) for theoretical values of regularization parameters that balance bias and variance. See Appendix (ref) for a practical tuning procedure based on the closed form solution for leave-one-out cross validation that empirically balances bias and variance.

Guarantees: Continuous treatment

Towards formal guarantees, I place weak regularity conditions on the original spaces. In anticipation of Section (ref) and Appendix (ref), I also include conditions for the follow-up covariate space and outcome space in parentheses.

assumption[Original space regularity conditions] Assume \begin{enumerate} • $\mathcal{D}$, $\mathcal{X}$ (and $\mathcal{M}$, $\mathcal{Y}$) are Polish spaces, i.e, separable and completely metrizable topological spaces. • $Y\in\mathbb{R}$ and it is bounded, i.e. there exists $C<\infty$ such that $|Y|\leq C$ almost surely. \end{enumerate}

This weak regularity condition allows treatment and covariates to be discrete or continuous and low, high, or infinite dimensional. The assumption that $Y\in\mathbb{R}$ is bounded simplifies the analysis, but may be relaxed.

Next, I assume that the various objects estimated by kernel ridge regressions and generalized kernel ridge regressions are smooth in the sense of (ref) in Section (ref). For dose response and incremental response curves, these are the regression $\gamma_0$ and conditional mean embeddings $(\mu_{x}(d),\mu_x(v))$. The smoothness assumption for the regression is as in Section (ref).

assumption[Smoothness of regression] Assume $\gamma_0\in\mathcal{H}^c$.

For conditional mean embeddings, I follow the abstract statement of smoothness from singh2020kernel. Define the abstract conditional mean embedding $\mu_{a}(b):=\int \phi(a)\mathrm{d}\mathbb{P}(a|b)$ where $a\in\mathcal{A}_j$ and $b\in\mathcal{B}_j$. I parametrize the smoothness of $\mu_{a}(b)$, with $a\in\mathcal{A}_j$ and $b\in\mathcal{B}_j$, by $c_j$. Define the abstract conditional expectation operator $E_j:\mathcal{H}_{\mathcal{A}_j}\rightarrow\mathcal{H}_{\mathcal{B}_j}$, $f(\cdot)\mapsto \mathbb{E}\{f(A_j)|B_j=\cdot\}$. The relationship between $E_j$ and $\mu_{a}(b)$ is given by $$ \mu_{a}(b)=\int \phi(a)\mathrm{d}\mathbb{P}(a|b) =\mathbb{E}_{A_j|B_j=b} \{\phi(A_j)\} =E_j^* \{\phi(b)\},\quad a\in\mathcal{A}_j, b\in\mathcal{B}_j $$ where $E_j^*$ is the adjoint of $E_j$. Let $\mathcal{L}_2(\mathcal{H}_{\mathcal{A}_j},\mathcal{H}_{\mathcal{B}_j})$ be the space of Hilbert-Schmidt operators between $\mathcal{H}_{\mathcal{A}_j}$ and $\mathcal{H}_{\mathcal{B}_j}$ by $\mathcal{L}_2(\mathcal{H}_{\mathcal{A}_j},\mathcal{H}_{\mathcal{B}_j})$, which is also an RKHS grunewalder2013smooth,singh2019kernel. With this notation, I state the abstract statement of smoothness.

assumption[Smoothness of mean embedding] Assume $E_j\in \{\mathcal{L}_2(\mathcal{H}_{\mathcal{A}_j},\mathcal{H}_{\mathcal{B}_j})\}^{^{c_j}}$.

With these assumptions, I arrive at the first main result.

theorem[Uniform consistency of dose and incremental responses curves] Suppose Assumptions (ref), (ref), (ref), and (ref) hold. Set $(\lambda,\lambda_1,\lambda_2)=(n^{-\frac{1}{c+1}},n^{-\frac{1}{c_1+1}},n^{-\frac{1}{c_2+1}})$. \begin{enumerate} • Then $ \|\hat{\theta}^{ATE}-\theta_0^{ATE}\|_{\infty}=O_p\left(n^{-\frac{1}{2}\frac{c-1}{c+1}}\right). $ • If in addition Assumption (ref) holds, then $ \|\hat{\theta}^{DS}(\cdot,\tilde{\mathbb{P}})-\theta_0^{DS}(\cdot,\tilde{\mathbb{P}})\|_{\infty}=O_p\left( n^{-\frac{1}{2}\frac{c-1}{c+1}}+\tilde{n}^{-\frac{1}{2}}\right). $ • If in addition Assumption (ref) holds with $\mathcal{A}_1=\mathcal{X}$ and $\mathcal{B}_1=\mathcal{D}$, then $ \|\hat{\theta}^{ATT}-\theta_0^{ATT}\|_{\infty}=O_p\left(n^{-\frac{1}{2}\frac{c-1}{c+1}}+n^{-\frac{1}{2}\frac{c_1-1}{c_1+1}}\right). $ • If in addition Assumption (ref) holds with $\mathcal{A}_2=\mathcal{X}$ and $\mathcal{B}_2=\mathcal{V}$, then $ \|\hat{\theta}^{CATE}-\theta_0^{CATE}\|_{\infty}=O_p\left(n^{-\frac{1}{2}\frac{c-1}{c+1}}+n^{-\frac{1}{2}\frac{c_2-1}{c_2+1}}\right). $ \end{enumerate} Likewise for incremental responses, e.g. $ \|\hat{\theta}^{\nabla:ATE}-\theta_0^{\nabla:ATE}\|_{\infty}=O_p\left(n^{-\frac{1}{2}\frac{c-1}{c+1}}\right) $. See Appendix (ref) for counterfactual distributions.

See Appendix (ref) for the proof and exact finite sample rates. The rates are at best $n^{-\frac{1}{6}}$ when $(c,c_1,c_2)=2$. The slow rates reflect the challenge of a $\sup$ norm guarantee, which encodes caution about worst case scenarios when informing policy decisions.

Guarantees: Discrete treatment

Consider the semiparametric case where $D$ is binary. Then $\theta_0^{ATE}(d)$ is a vector in $\mathbb{R}^{2}$ and Theorem (ref) simplifies to a guarantee on the maximum element of the vector of differences $|\hat{\theta}^{ATE}(d)-\theta_0^{ATE}(d)|$. For the semiparametric case, I improve the rate from $n^{-\frac{1}{6}}$ to $n^{-\frac{1}{2}}$ by using Algorithm (ref) and imposing additional assumptions. In particular, I place additional assumptions on the regression propensity scores, Riesz representer, and corresponding RKHSs.

assumption[Bounded propensities and Riesz representer] Assume there exist $\epsilon>0$ and $\bar{\alpha}'<\infty$ such that \begin{enumerate} • $\epsilon \leq \mathbb{P}(D=1|X)\leq 1-\epsilon $ and $\epsilon \leq \mathbb{P}(S=1|D,X)\leq 1-\epsilon $ with probability one; • $\|\hat{\alpha}^{(m)}\|_{\infty}\leq \bar{\alpha}'$ with probability approaching one. \end{enumerate} Likewise for the propensity scores in the cases of $\theta_0^{DS}$ and $\theta_0^{CATE}$.

Part one of Assumption (ref) is a mild strengthening of the overlap condition in Assumption (ref). It implies boundedness of the Riesz representer $\alpha_0^{(m)}$ and mean square continuity of the corresponding functional chernozhukov2021simple. Part two can be imposed by censoring extreme evaluations of the estimator $\hat{\alpha}^{(m)}$. Note that censoring can only improve prediction quality because $\alpha^{(m)}_0$ is bounded as an implication of part one.

Finally, I assume smoothness and spectral decay assumptions for the Riesz representer as described in Section (ref).

assumption[Smoothness of Riesz representer] Assume $\alpha^{(m)}_0\in \mathcal{H}^{c_3}$.

To quantify the effective dimension of the Riesz representer kernel, recall the convolution operator notation from Section (ref): $\lambda_j(k)$ is the $j$-th eigenvalue of the convolution operator $L:f\mapsto \int k(\cdot, w)f(w)\mathrm{d}\mathbb{P}(w)$ of the kernel $k$.

assumption[Effective dimension of Riesz representer] Assume $\lambda_j(k_{\mathcal{S}}\cdot k_{\mathcal{D}} \cdot k_{\mathcal{X}})\asymp j^{-b_3}$.

The latter statement quantifies the effective dimension of $\alpha_0^{(m)} \in \mathcal{H}$. Recall from Section (ref) that Sobolev kernels possess this property. See scetbon2021spectral for additional examples. In Appendix (ref), I argue that it is sufficient for either the kernel of the Riesz representer $\alpha_0^{(m)}$ to have polynomial spectral decay or for the kernel of the regression $\gamma_0$ to have polynomial spectral decay---a kind of double spectral robustness singh2020kernel,singh2021workshop.

theorem[Semiparametic consistency, Gaussian approximation, and efficiency for static sample selection] Suppose the conditions of Theorem (ref) hold, as well as Assumptions (ref), (ref), and (ref). For a given treatment effect, denote the moments of $$ \psi^{(m)}_0(W):=m(1,D,X;\gamma_0)+\alpha^{(m)}_0(S,D,X)[SY-\gamma_0(1,D,X)]-\theta^{(m)}_0$$ by $\sigma^2=\mathbb{E}[\psi_0(W)^2]$, $\eta^3=\mathbb{E}[|\psi_0(W)|^3]$, and $\chi^4=\mathbb{E}[\psi_0(W)^4]$. Assume the regularity conditions that $\sigma^2$ is bounded away from zero and $\left\{\left(\frac{\eta}{\sigma}\right)^3+\chi^2\right\}n^{-\frac{1}{2}}\rightarrow0$. Set $\lambda=n^{-\frac{1}{2}\frac{c-1}{c+1}}$. Set $\lambda_3=n^{-\frac{1}{2}}$ if $b_3=\infty$ and $\lambda_3=n^{-\frac{b_3}{b_3c_3+1}}$ otherwise. Then for any $c,c_3\in(1,2]$ and $b_3\in(1,\infty]$ satisfying $$ \frac{b_3c_3}{b_3c_3+1}+\frac{c-1}{c+1}>1, $$ the abstract estimator $\hat{\theta}^{(m)}$ is consistent, i.e. $\hat{\theta}^{(m)}\overset{p}{\rightarrow}\theta_0^{(m)}$, and the confidence interval includes $\theta_0^{(m)}$ with probability approaching the nominal level, i.e. $$\lim_{n\rightarrow\infty} \mathbb{P}\left[\theta_0^{(m)}\in \left\{\hat{\theta}^{(m)}\pm c_a\frac{\hat{\sigma}}{\sqrt{n}}\right\}\right]=1-a.$$

See Appendix (ref) for the proof and a finite sample Gaussian approximation. Observe the tradeoff among the smoothness and effective dimension assumptions across nonparametric objects. The spectral decay of the Riesz representer $\alpha_0^{(m)}$ parametrized by $b_3$ must be sufficiently fast relative to the smoothness of the nonparametric objects $(\gamma_0,\alpha^{(m)}_0)$ parametrized by $(c,c_3)$. To see that this parameter region is nonempty, observe that if $(c,c_3)=2$ then the condition simplifies to $b_3>1$. In other words, when the various nonparametric objects are smooth, it is sufficient for the Riesz representer to belong to an RKHS with any polynomial rate of spectral decay.

Dynamic sample selection

Causal parameters

So far I have studied the static sample selection model where selection is as good as random conditional on treatment $D$ and baseline covariates $X$. A richer class of sample selection models allows for selection to be as good as random conditional on treatment $D$, baseline covariates $X$, and follow-up covariates $M$ measured after treatment. In other words I open up the possibility that, among the covariates in the selection mechanism, some are baseline covariates $X$ that may cause treatment while others are follow-up covariates $M$ that may be caused by treatment. I refer to this setting as dynamic sample selection. The dynamic sample selection case is more challenging from a statistical perspective. I study a subset of the causal parameters from Section (ref) due to this additional complexity, though in principle the same techniques should apply to the entire class of causal parameters.

Identification

bia2020double articulate distribution-free sufficient conditions under which average treatment effect can be measured from observations $(SY,S,D,X,M)$. I refer to this collection of sufficient conditions as dynamic sample selection on observables and argue that it identifies additional causal parameters from Definition (ref).

assumption[Dynamic sample selection on observables] Assume \begin{enumerate} • No interference: if $D=d$ then $Y=Y^{(d)}$. • Conditional exchangeability: $\{Y^{(d)}\}\raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} D |X$ and $\{Y^{(d)}\}\raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} S|D,X,M$. • Overlap: if $f(x)>0$ then $f(d|x)>0$ and if $f(d,x,m)>0$ then $\mathbb{P}(S=1|d,x,m)>0$; where $f(x)$, $f(d|x)$, and $f(d,x,m)$, are densities. \end{enumerate}

Observe that Assumption (ref) is a generalization of Assumption (ref). In particular, $\{Y^{(d)}\}\raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} S|D,X,M$ generalizes $\{Y^{(d)}\}\raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} S|D,X$. In this sense, the dynamic sample selection model generalizes missing-at-random (MAR) rubin1976inference. The updated overlap condition is more stringent: there is no treatment-baseline-follow up stratum $(D,X,M)=(d,x,m)$ such that selection is impossible. Similarly, I adapt Assumption (ref) to handle $\theta_0^{DS}$.

assumption[Dynamic distribution shift] Assume \begin{enumerate} • $\tilde{\mathbb{P}}(Y,S,D,X,M)=\mathbb{P}(Y|S,D,X,M)\tilde{\mathbb{P}}(S,D,X,M)$; • $\tilde{\mathbb{P}}(S,D,X,M)$ is absolutely continuous with respect to $\mathbb{P}(S,D,X,M)$. \end{enumerate}

As before, difference in population distributions $\mathbb{P}$ and $\tilde{\mathbb{P}}$ is only in the distribution of selection, treatments, and covariates. An immediate consequence is that the regression function for complete cases $\gamma_0(1,d,x,m)=\mathbb{E}(Y|S=1,D=d,X=x,M=m)$ remains the same across the different populations $\mathbb{P}$ and $\tilde{\mathbb{P}}$.

Similar to Theorem (ref), I argue that different causal parameters are reweightings of the regression using complete cases, i.e. $\gamma_0(1,d,x,m)$. This nuance is critical, since the data limitations of the problem setting imply that $\mathbb{E}(Y|S=1,D=d,X=x,M=m)$ is estimable while $\mathbb{E}(Y|S=0,D=d,X=x,M=m)$ is not. The reweightings are more complex due to dynamic sample selection. I modestly extend the argument of bia2020double to cover the additional cases of distribution shift and counterfactual distributions.

theorem[Identification of causal parameters under dynamic sample selection] If Assumption (ref) holds then \begin{enumerate} • $ \theta_0^{ATE}(d)=\int \gamma_0(1,d,x,m) \mathrm{d}\mathbb{P}(m|d,x) \mathrm{d}\mathbb{P}(x) $ bia2020double. • If in addition Assumption (ref) holds then $ \theta_0^{DS}(d,\tilde{\mathbb{P}})=\int \gamma_0(1,d,x,m) \mathrm{d}\tilde{\mathbb{P}}(m|d,x) \mathrm{d}\tilde{\mathbb{P}}(x) $. \end{enumerate} See Appendix (ref) for counterfactual distributions.

See Appendix (ref) for the proof.

Algorithm: Continuous treatment

I extend the RKHS construction from Section (ref) to accommodate the follow-up covariates $M$. In particular, I define an additional scalar valued RKHS $\mathcal{H}_{\mathcal{M}}$ and construct the tensor product RKHS $\mathcal{H}=\mathcal{H}_{\mathcal{S}}\otimes \mathcal{H}_{\mathcal{D}} \otimes \mathcal{H}_{\mathcal{X}}\otimes \mathcal{H}_{\mathcal{M}}$ with feature map $\phi(s)\otimes \phi(d) \otimes \phi(x) \times \phi(m)$. In this construction, the kernel for $\mathcal{H}$ is $k(s,d,x,m;s',d',x',m')=k_\mathcal{S}(s,s')\cdot k_\mathcal{D}(d,d')\cdot k_{\mathcal{X}}(x,x')\cdot k_{\mathcal{M}}(m,m')$.

As before, by the reproducing property, $\gamma_0\in \mathcal{H}$ implies $ \gamma_0(s,d,x,m)=\langle \gamma_0, \phi(s)\otimes \phi(d)\otimes \phi(x) \otimes \phi(m) \rangle_{\mathcal{H}} $. The regularity conditions in Assumption (ref) help to represent the causal parameters as inner products in $\mathcal{H}$.

theorem[Representation of dose response curves via kernel mean embeddings] Suppose the conditions of Theorem (ref) hold. Further suppose Assumption (ref) holds and $\gamma_0\in\mathcal{H}$. Then \begin{enumerate} • $\theta_0^{ATE}(d)=\langle \gamma_0, \phi(1)\otimes\phi(d) \otimes \int \{\phi(x)\otimes \mu_{m}(d,x)\}\mathrm{d} \mathbb{P}(x) \rangle_{\mathcal{H}}$ where $\mu_{m}(d,x):=\int \phi(m)\mathrm{d}\mathbb{P}(m|d,x)$, • $\theta_0^{DS}(d,\tilde{\mathbb{P}})=\langle \gamma_0, \phi(1)\otimes\phi(d) \otimes \int \{\phi(x)\otimes \nu_{m}(d,x)\}\mathrm{d} \tilde{\mathbb{P}}(x) \rangle_{\mathcal{H}}$ where $\nu_{m}(d,x):=\int \phi(m)\mathrm{d}\tilde{\mathbb{P}}(m|d,x)$. \end{enumerate} See Appendix (ref) for counterfactual distributions.

See Appendix (ref) for the proof. The conditional mean embedding $\mu_{m}(d,x):=\int \phi(m)\mathrm{d}\mathbb{P}(m|d,x)$ encodes the conditional distribution $\mathbb{P}(m|d,x)$. The sequential mean embedding $\int \{\phi(x)\otimes \mu_{m}(d,x)\}\mathrm{d}\mathbb{P}(x)$ encodes the counterfactual distribution of the sequence of covariates $(X,M)$ when treatment is $D=d$, adapting a technique introduced by singh2021workshop. The inner product representation suggests estimators that are inner products, e.g. $\hat{\theta}^{ATE}(d)=\langle \hat{\gamma}, \phi(1)\otimes\phi(d) \otimes \frac{1}{n}\sum_{i=1}^n\phi(X_i)\otimes \hat{\mu}_{m}(d,X_i)\} \rangle_{\mathcal{H}}$ where $\hat{\gamma}$ is a standard kernel ridge regression using the selected sample and $\hat{\mu}_{m}(d,x)$ is a generalized kernel ridge regression.

algorithm[algorithm omitted — 1,553 chars of source]

See Appendix (ref) for the derivation. See Theorem (ref) for theoretical values of $(\lambda,\lambda_4,\lambda_5)$ that balance bias and variance. See Appendix (ref) for a practical tuning procedure based on the closed form solution of leave-one-out cross validation to empirically balance bias and variance.

Algorithm: Discrete treatment

So far, I have presented nonparametric algorithms for the continuous treatment cases. Now I present semiparametric algorithms for the discrete treatment case. Towards this end, I quote the multiply robust moment function for the dynamic sample selection model parametrized to avoid density estimation.

lemma[Multiply robust moment for dynamic sample selection bia2020double] \; \\ Suppose treatment $D$ is binary, so that $d\in\{0,1\}$. If Assumption (ref) holds then \begin{align*} \theta^{ATE}_0(d) &=\mathbb{E}\bigg[ \omega_0(1,d;X) + \frac{\operatorname{\mathbbm 1}_{D=d}S}{\pi_0(d;X)\rho_0(1;d,X,M)}\{SY-\gamma_0(1,d,X,M)\}\\ &\quad + \frac{\operatorname{\mathbbm 1}_{D=d}}{\pi_0(d;X)}\left\{\gamma_0(1,d,X,M)-\omega_0(1,d;X)\right\} \bigg], \end{align*} where $\omega_0(1,d;X):=\int \gamma_0(1,d,X,m) \mathrm{d}\mathbb{P}(m|d,X)$ is the sequential mean embedding, $ \pi_0(d;X):=\mathbb{P}(d|X)$ is the treatment propensity score, and $\rho_0(1;d,X,M):=\mathbb{P}(S=1|d,X,M)$ is the selection propensity score.

The equation in Lemma (ref) has three important properties. First, it consists of estimable quantities given the data limitations of the dynamic sample selection problem. Second, it is multiply robust with respect to the nonparametric objects $(\gamma_0,\omega_0,\pi_0,\rho_0)$ in the sense that it continues to hold if one of the nonparametric objects is misspecified. For example, for any $\gamma$,

align*[align* omitted — 291 chars of source]

Multiple robustness implies that the semiparametric guarantees below hold even if one of the nonparametric objects $(\gamma_0,\omega_0,\pi_0,\rho_0)$ is not actually an element of an RKHS. Third, multiple robustness is with respect to $\omega_0$ rather than the conditional density $f(m|d,x)$. The technique of sequential mean embeddings permits estimation of $\omega_0$ without estimation of $f(m|d,x)$ singh2021workshop.

The meta algorithm that combines the multiply robust moment function from Lemma (ref) with sample splitting is called DML for dynamic sample selection chernozhukov2018original,bia2020double,chernozhukov2021simple. I instantiate DML for dynamic sample selection using the new nonparametric estimator from Algorithm (ref).

algorithm[algorithm omitted — 1,656 chars of source]

See Appendix (ref) for explicit computations. The sequential mean embedding $\hat{\omega}$ is computed according to Algorithm (ref) with regularization parameters $(\lambda,\lambda_4)$. The remaining nonparametric subroutines $(\hat{\gamma},\hat{\pi},\hat{\rho})$ are standard kernel ridge regressions with regularization parameters $(\lambda,\lambda_6,\lambda_7)$. See Theorem (ref) for theoretical values of the regularization parameters that balance bias and variance. See Appendix (ref) for a practical tuning procedure based on the closed form solution of leave-one-out cross validation to empirically balance bias and variance.

Unlike bia2020double, Algorithm (ref) does not require multiple levels of sample splitting. Crucially, I am able to prove a new, sufficiently fast rate of convergence on the sequential mean embedding $\hat{\omega}$ from Algorithm (ref) without additional sample splitting, which allows for this simplification.

Guarantees: Continuous treatment

As in Section (ref), I place regularity conditions on the original spaces, assume the regression $\gamma_0$ is smooth, and assume the conditional mean embeddings $\mu_{m}(d,x)$ and $\nu_{m}(d,x)$ are smooth in order to prove uniform consistency for the continuous treatment case.

theorem[Uniform consistency of dose and incremental response curves] Suppose Assumptions (ref), (ref), (ref), and (ref) hold. Set $(\lambda,\lambda_4,\lambda_5)$= $(n^{-\frac{1}{c+1}},n^{-\frac{1}{c_4+1}},\tilde{n}^{-\frac{1}{c_5+1}})$. \begin{enumerate} • If in addition Assumption (ref) holds with $\mathcal{A}_4=\mathcal{X}$ and $\mathcal{B}_4=\mathcal{D}\times \mathcal{X}$, then $ \|\hat{\theta}^{ATE}-\theta_0^{ATE}\|_{\infty}=O_p\left(n^{-\frac{1}{2}\frac{c-1}{c+1}}+n^{-\frac{1}{2}\frac{c_4-1}{c_4+1}}\right). $ • If in addition Assumptions (ref) and (ref) hold with $\mathcal{A}_5=\mathcal{X}$ and $\mathcal{B}_5=\mathcal{D}\times \mathcal{X}$, then $ \|\hat{\theta}^{DS}-\theta_0^{DS}\|_{\infty}=O_p\left(n^{-\frac{1}{2}\frac{c-1}{c+1}}+\tilde{n}^{-\frac{1}{2}\frac{c_5-1}{c_5+1}}\right) $. \end{enumerate} See Appendix (ref) for counterfactual distributions.

See Appendix (ref) for the proof and exact finite sample rates. As before, these rates are at best $n^{-\frac{1}{6}}$ when $(c,c_4,c_5)=2$.

Guarantees: Discrete treatment

Consider the semiparametric case where $D$ is binary. Then $\theta_0^{ATE}(d)$ is a vector in $\mathbb{R}^{2}$ and Theorem (ref) simplifies to a guarantee on the maximum element of the vector of differences $|\hat{\theta}^{ATE}(d)-\theta_0^{ATE}(d)|$. For the semiparametric case, I improve the rate from $n^{-\frac{1}{6}}$ to $n^{-\frac{1}{2}}$ by using Algorithm (ref) and imposing additional assumptions.

Since $\gamma_0(1,d,x,m)=\mathbb{E}(Y|S=1,D=d,X=x,M=m)$, I refer to $Y-\gamma_0(1,D,X,M)$ as the regression residual. I assume that the regression residual has non-degenerate variance.

assumption[Non-degenerate residual] Assume there exists $\chi>0$ such that $\inf_{d}\mathbb{E}[\{Y-\gamma_0(1,d,X,M)\}^2]\geq \chi$.

To lighten notation for the binary treatment case, let $\pi_0(x):=\pi_0(1;x)=\mathbb{P}(D=1|X=x)$ and $\rho_0(d,x,m):=\rho_0(1;d,x,m)=\mathbb{P}(S=1|D=d,X=x,M=m)$. I assume these treatment and selection propensity scores are bounded away from zero and one.

assumption[Bounded propensities] Assume propensity scores are bounded away from zero and one, i.e. there exists $\epsilon>0$ such that \begin{enumerate} • $\epsilon \leq \pi_0(X)\leq 1-\epsilon $ and $\epsilon \leq \rho_0(D,X,M)\leq 1-\epsilon $ with probability one; • $\epsilon \leq \inf_{x} \hat{\pi}(x)\leq \sup_{x} \hat{\pi}(x) \leq 1-\epsilon$ and $\epsilon \leq \inf_{d,x,m} \hat{\rho}(d,x,m)\leq \sup_{d,x,m} \hat{\rho}(d,x,m) \leq 1-\epsilon$ with probability approaching one. \end{enumerate}

Part one of Assumption (ref) is a mild strengthening of the overlap condition in Assumption (ref). Part two can be imposed by censoring extreme evaluations of the estimators $(\hat{\pi},\hat{\rho})$. Note that censoring can only improve prediction quality because $(\pi_0,\rho_0)$ are bounded away from zero and one by hypothesis.

Finally, I assume smoothness and spectral decay assumptions for the treatment and selection propensity scores as described in Section (ref).

assumption[Smoothness of propensities] Assume $\pi_0\in \mathcal{H}_{\mathcal{X}}^{c_6}$ and $\rho_0\in (\mathcal{H}_{\mathcal{D}}\otimes \mathcal{H}_{\mathcal{X}}\otimes \mathcal{H}_{\mathcal{M}})^{c_7}$.

To quantify the effective dimension of the propensity kernels, recall the convolution operator notation from Section (ref): $\lambda_j(k)$ is the $j$-th eigenvalue of the convolution operator $L:f\mapsto \int k(\cdot, w)f(w)\mathrm{d}\mathbb{P}(w)$ of the kernel $k$.

assumption[Effective dimension of propensities] Assume $\lambda_j(k_{\mathcal{X}})\asymp j^{-b_6}$ and $\lambda_j(k_{\mathcal{D}}\cdot k_{\mathcal{X}}\cdot k_{\mathcal{M}})\asymp j^{-b_7}$.

The former statement quantifies the effective dimension of $\pi_0 \in \mathcal{H}_{\mathcal{X}}$, and the latter quantifies the effective dimension of $\rho_0 \in \mathcal{H}_{\mathcal{D}}\otimes \mathcal{H}_{\mathcal{X}}\otimes \mathcal{H}_{\mathcal{M}}$. In Appendix (ref), I argue that it is sufficient for either the kernels of the propensities $(\pi_0,\rho_0)$ to have polynomial spectral decay or for the kernels of the regression and sequential mean embedding $(\gamma_0,\omega_0)$ to have polynomial spectral decay---a kind of double spectral robustness.

theorem[Semiparametic consistency, Gaussian approximation, and efficiency for dynamic sample selection] Suppose the conditions of Theorem (ref) hold, as well as Assumptions (ref), (ref), (ref), and (ref). Set $\lambda_6=n^{-\frac{1}{2}}$ if $b_6=\infty$ and $\lambda_6=n^{-\frac{b_6}{b_6c_6+1}}$ otherwise. Set $\lambda_7=n^{-\frac{1}{2}}$ if $b_7=\infty$ and $\lambda_7=n^{-\frac{b_7}{b_7c_7+1}}$ otherwise. Then for any $c,c_4,c_6,c_7\in(1,2]$ and $b_6,b_7\in(1,\infty]$ satisfying $$ \frac{\min(b_6c_6,b_7c_7)}{\min(b_6c_6,b_7c_7)+1}+\frac{\min(c,c_4)-1}{\min(c,c_4)+1}>1, $$ the estimator $\hat{\theta}^{ATE}(d)$ is consistent, i.e. $\hat{\theta}^{ATE}(d)\overset{p}{\rightarrow}\theta_0^{ATE}(d)$, and the confidence interval includes $\theta_0^{ATE}(d)$ with probability approaching the nominal level, i.e. $$\lim_{n\rightarrow\infty} \mathbb{P}\left[\theta_0^{ATE}(d)\in \left\{\hat{\theta}^{ATE}(d)\pm c_a\frac{\hat{\sigma}(d)}{\sqrt{n}}\right\}\right]=1-a.$$

See Appendix (ref) for the proof. Observe the tradeoff among the smoothness and effective dimension assumptions across nonparametric objects. The spectral decay of propensity scores $(\pi_0,\rho_0)$ parametrized by $(b_6,b_7)$ must be sufficiently fast relative to the smoothness of the various nonparametric objects $(\gamma_0,\mu_{m},\pi_0,\rho_0)$ parametrized by $(c,c_4,c_6,c_7)$. To see that this parameter region is nonempty, observe that if $(c,c_4,c_6,c_7)=2$ then the condition simplifies to $\min(b_6,b_7)>1$. In other words, when the various nonparametric objects are smooth, it is sufficient for the treatment and selection propensities to belong to RKHSs with any polynomial rate of spectral decay.

Conclusion

I propose a family of novel estimators for nonparametric estimation of dose response curves from selected samples as well as semiparametric estimation of treatment effects from selected samples. The estimators are easily implemented from closed form solutions in terms of kernel matrix operations. As a contribution to the sample selection literature, I propose simple estimators with closed form solutions for new causal estimands. As a contribution to the kernel methods literature, extend the approximation and estimation theory of kernel ridge regression to causal sample selection models. I pose as a question for future research how to extend this framework to the setting with missing-not-at-random (MNAR) data.

thebibliography{10} \bibitem{altun2006unifying} Yasemin Altun and Alex Smola. \newblock Unifying divergence minimization and statistical inference via convex duality. \newblock In {\em Conference on Computational Learning Theory}, pages 139--153. Springer, 2006. \bibitem{bach2012equivalence} Francis Bach, Simon Lacoste-Julien, and Guillaume Obozinski. \newblock On the equivalence between herding and conditional gradient algorithms. \newblock In {\em International Conference on Machine Learning}, pages 1355--1362, 2012. \bibitem{bang2005doubly} Heejung Bang and James M Robins. \newblock Doubly robust estimation in missing data and causal inference models. \newblock {\em Biometrics}, 61(4):962--973, 2005. \bibitem{bia2020double} Michela Bia, Martin Huber, and Luk{\'a}{\v{s}} Laff{\'e}rs. \newblock Double machine learning for sample selection models. \newblock {\em arXiv:2012.00745}, 2020. \bibitem{bickel1982adaptive} Peter J Bickel. \newblock On adaptive estimation. \newblock {\em The Annals of Statistics}, pages 647--671, 1982. \bibitem{caponnetto2007optimal} Andrea Caponnetto and Ernesto De Vito. \newblock Optimal rates for the regularized least-squares algorithm. \newblock {\em Foundations of Computational Mathematics}, 7(3):331--368, 2007. \bibitem{carrasco2007linear} Marine Carrasco, Jean-Pierre Florens, and Eric Renault. \newblock Linear inverse problems in structural econometrics estimation based on spectral decomposition and regularization. \newblock {\em Handbook of Econometrics}, 6:5633--5751, 2007. \bibitem{cattaneo2010efficient} Matias D Cattaneo. \newblock Efficient semiparametric estimation of multi-valued treatment effects under ignorability. \newblock {\em Journal of Econometrics}, 155(2):138--154, 2010. \bibitem{chernozhukov2018original} Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. \newblock Double/debiased machine learning for treatment and structural parameters. \newblock {\em The Econometrics Journal}, 21(1), 2018. \bibitem{chernozhukov2016locally} Victor Chernozhukov, Juan Carlos Escanciano, Hidehiko Ichimura, Whitney K Newey, and James M Robins. \newblock Locally robust semiparametric estimation. \newblock {\em arXiv:1608.00033}, 2016. \bibitem{chernozhukov2013inference} Victor Chernozhukov, Iv{\'a}n Fern{\'a}ndez-Val, and Blaise Melly. \newblock Inference on counterfactual distributions. \newblock {\em Econometrica}, 81(6):2205--2268, 2013. \bibitem{chernozhukov2018global} Victor Chernozhukov, Whitney Newey, and Rahul Singh. \newblock De-biased machine learning of global and local parameters using regularized {R}iesz representers. \newblock {\em arXiv:1802.08667}, 2018. \bibitem{chernozhukov2018learning} Victor Chernozhukov, Whitney K Newey, and Rahul Singh. \newblock Automatic debiased machine learning of causal and structural effects. \newblock {\em arXiv:1809.05224}, 2018. \bibitem{chernozhukov2021simple} Victor Chernozhukov, Whitney K Newey, and Rahul Singh. \newblock A simple and general debiased machine learning theorem with finite sample guarantees. \newblock {\em arXiv:2105.15197}, 2021. \bibitem{cortes2008sample} Corinna Cortes, Mehryar Mohri, Michael Riley, and Afshin Rostamizadeh. \newblock Sample selection bias correction theory. \newblock In {\em International Conference on Algorithmic Learning Theory}, pages 38--53. Springer, 2008. \bibitem{cucker2002mathematical} Felipe Cucker and Steve Smale. \newblock On the mathematical foundations of learning. \newblock {\em Bulletin of the American Mathematical Society}, 39(1):1--49, 2002. \bibitem{das2003nonparametric} Mitali Das, Whitney K Newey, and Francis Vella. \newblock Nonparametric estimation of sample selection models. \newblock {\em The Review of Economic Studies}, 70(1):33--58, 2003. \bibitem{dinardo1996labor} John DiNardo, Nicole M Fortin, and Thomas Lemieux. \newblock Labor market institutions and the distribution of wages, 1973-1992: A semiparametric approach. \newblock {\em Econometrica}, 64(5):1001--1044, 1996. \bibitem{d2010new} Xavier d’Haultfoeuille. \newblock A new instrumental method for dealing with endogenous selection. \newblock {\em Journal of Econometrics}, 154(1):1--15, 2010. \bibitem{edmunds2008function} David E Edmunds and Hans Triebel. \newblock {\em Function Spaces, Entropy Numbers and Differential Operators}. \newblock Cambridge University Press, 2008. \bibitem{firpo2007efficient} Sergio Firpo. \newblock Efficient semiparametric estimation of quantile treatment effects. \newblock {\em Econometrica}, 75(1):259--276, 2007. \bibitem{fischer2017sobolev} Simon Fischer and Ingo Steinwart. \newblock Sobolev norm learning rates for regularized least-squares algorithms. \newblock {\em Journal of Machine Learning Research}, 21:1--38, 2020. \bibitem{foster2019orthogonal} Dylan J Foster and Vasilis Syrgkanis. \newblock Orthogonal statistical learning. \newblock {\em arXiv:1901.09036}, 2019. \bibitem{gretton2009covariate} Arthur Gretton, Alex Smola, Jiayuan Huang, Marcel Schmittfull, Karsten Borgwardt, and Bernhard Sch{\"o}lkopf. \newblock Covariate shift by kernel mean matching. \newblock {\em Dataset Shift in Machine Learning}, 3(4):5, 2009. \bibitem{grunewalder2013smooth} Steffen Gr{\"u}new{\"a}lder, Arthur Gretton, and John Shawe-Taylor. \newblock Smooth operators. \newblock In {\em International Conference on Machine Learning}, pages 1184--1192, 2013. \bibitem{hausman1979attrition} Jerry A Hausman and David A Wise. \newblock Attrition bias in experimental and panel data: the {G}ary income maintenance experiment. \newblock {\em Econometrica}, pages 455--473, 1979. \bibitem{heckman1979sample} James J Heckman. \newblock Sample selection bias as a specification error. \newblock {\em Econometrica}, pages 153--161, 1979. \bibitem{hirshberg2019minimax} David A Hirshberg, Arian Maleki, and Jose R Zubizarreta. \newblock Minimax linear estimation of the retargeted mean. \newblock {\em arXiv:1901.10296}, 2019. \bibitem{huang2006correcting} Jiayuan Huang, Arthur Gretton, Karsten Borgwardt, Bernhard Sch{\"o}lkopf, and Alex Smola. \newblock Correcting sample selection bias by unlabeled data. \newblock {\em Advances in Neural Information Processing Systems}, 19:601--608, 2006. \bibitem{huber2012identification} Martin Huber. \newblock Identification of average treatment effects in social experiments under alternative forms of attrition. \newblock {\em Journal of Educational and Behavioral Statistics}, 37(3):443--474, 2012. \bibitem{huber2014treatment} Martin Huber. \newblock Treatment evaluation in the presence of sample selection. \newblock {\em Econometric Reviews}, 33(8):869--905, 2014. \bibitem{imbens2009identification} Guido W Imbens and Whitney K Newey. \newblock Identification and estimation of triangular simultaneous equations models without additivity. \newblock {\em Econometrica}, 77(5):1481--1512, 2009. \bibitem{kallus2020generalized} Nathan Kallus. \newblock Generalized optimal matching methods for causal inference. \newblock {\em Journal of Machine Learning Research}, 21:1--54, 2020. \bibitem{kanagawa2014recovering} Motonobu Kanagawa and Kenji Fukumizu. \newblock Recovering distributions from {G}aussian {RKHS} embeddings. \newblock In {\em Artificial Intelligence and Statistics}, pages 457--465. PMLR, 2014. \bibitem{kimeldorf1971some} George Kimeldorf and Grace Wahba. \newblock Some results on {T}chebycheffian spline functions. \newblock {\em Journal of Mathematical Analysis and Applications}, 33(1):82--95, 1971. \bibitem{miao2015identification} Wang Miao, Lan Liu, Eric Tchetgen Tchetgen, and Zhi Geng. \newblock Identification, doubly robust estimation, and semiparametric efficiency theory of nonignorable missing data with a shadow variable. \newblock {\em arXiv:1509.02556}, 2015. \bibitem{muandet2020counterfactual} Krikamol Muandet, Motonobu Kanagawa, Sorawit Saengkyongam, and Sanparith Marukatat. \newblock Counterfactual mean embeddings. \newblock {\em Journal of Machine Learning Research}, 22(162):1--71, 2021. \bibitem{negi2020doubly} Akanksha Negi. \newblock Doubly weighted m-estimation for nonrandom assignment and missing outcomes. \newblock {\em arXiv:2011.11485}, 2020. \bibitem{newey1994kernel} Whitney K Newey. \newblock Kernel estimation of partial means and a general variance estimator. \newblock {\em Econometric Theory}, pages 233--253, 1994. \bibitem{nie2017quasi} Xinkun Nie and Stefan Wager. \newblock Quasi-oracle estimation of heterogeneous treatment effects. \newblock {\em Biometrika}, 108(2):299--319, 2021. \bibitem{robins1986new} James Robins. \newblock A new approach to causal inference in mortality studies with a sustained exposure period—application to control of the healthy worker survivor effect. \newblock {\em Mathematical Modelling}, 7(9-12):1393--1512, 1986. \bibitem{robins1995semiparametric} James M Robins and Andrea Rotnitzky. \newblock Semiparametric efficiency in multivariate regression models with missing data. \newblock {\em Journal of the American Statistical Association}, 90(429):122--129, 1995. \bibitem{robins1994estimation} James M Robins, Andrea Rotnitzky, and Lue Ping Zhao. \newblock Estimation of regression coefficients when some regressors are not always observed. \newblock {\em Journal of the American Statistical Association}, 89(427):846--866, 1994. \bibitem{robinson1988root} Peter M Robinson. \newblock Root-n-consistent semiparametric regression. \newblock {\em Econometrica}, pages 931--954, 1988. \bibitem{rotnitzky1998semiparametric} Andrea Rotnitzky, James M Robins, and Daniel O Scharfstein. \newblock Semiparametric regression for repeated outcomes with nonignorable nonresponse. \newblock {\em Journal of the american statistical association}, 93(444):1321--1339, 1998. \bibitem{rubin1976inference} Donald B Rubin. \newblock Inference and missing data. \newblock {\em Biometrika}, 63(3):581--592, 1976. \bibitem{scetbon2021spectral} Meyer Scetbon and Zaid Harchaoui. \newblock A spectral analysis of dot-product kernels. \newblock In {\em International Conference on Artificial Intelligence and Statistics}, pages 3394--3402. PMLR, 2021. \bibitem{scharfstein1999adjusting} Daniel O Scharfstein, Andrea Rotnitzky, and James M Robins. \newblock Adjusting for nonignorable drop-out using semiparametric nonresponse models. \newblock {\em Journal of the American Statistical Association}, 94(448):1096--1120, 1999. \bibitem{simon2020metrizing} Carl-Johann Simon-Gabriel, Alessandro Barp, and Lester Mackey. \newblock Metrizing weak convergence with maximum mean discrepancies. \newblock {\em arXiv:2006.09268}, 2020. \bibitem{singh2020negative} Rahul Singh. \newblock Kernel methods for unobserved confounding: Negative controls, proxies, and instruments. \newblock {\em arXiv:2012.10315}, 2020. \bibitem{singh2019kernel} Rahul Singh, Maneesh Sahani, and Arthur Gretton. \newblock Kernel instrumental variable regression. \newblock In {\em Advances in Neural Information Processing Systems}, pages 4595--4607, 2019. \bibitem{singh2020kernel} Rahul Singh, Liyuan Xu, and Arthur Gretton. \newblock Reproducing kernel methods for nonparametric and semiparametric treatment effects. \newblock {\em arXiv:2010.04855}, 2020. \bibitem{singh2021workshop} Rahul Singh, Liyuan Xu, and Arthur Gretton. \newblock Kernel methods for multistage causal inference: Mediation analysis and dynamic treatment effects. \newblock {\em arXiv:2111.03950}, 2021. \bibitem{smale2007learning} Steve Smale and Ding-Xuan Zhou. \newblock Learning theory estimates via integral operators and their approximations. \newblock {\em Constructive Approximation}, 26(2):153--172, 2007. \bibitem{smola2007hilbert} Alex Smola, Arthur Gretton, Le Song, and Bernhard Sch{\"o}lkopf. \newblock A {H}ilbert space embedding for distributions. \newblock In {\em International Conference on Algorithmic Learning Theory}, pages 13--31, 2007. \bibitem{sriperumbudur2016optimal} Bharath Sriperumbudur. \newblock On the optimal estimation of probability measures in weak and strong topologies. \newblock {\em Bernoulli}, 22(3):1839--1893, 2016. \bibitem{sriperumbudur2010relation} Bharath Sriperumbudur, Kenji Fukumizu, and Gert Lanckriet. \newblock On the relation between universality, characteristic kernels and {RKHS} embedding of measures. \newblock In {\em International Conference on Artificial Intelligence and Statistics}, pages 773--780, 2010. \bibitem{steinwart2008support} Ingo Steinwart and Andreas Christmann. \newblock {\em Support Vector Machines}. \newblock Springer Science & Business Media, 2008. \bibitem{steinwart2012mercer} Ingo Steinwart and Clint Scovel. \newblock Mercer’s theorem on general domains: On the interaction between measures, kernels, and {RKHS}s. \newblock {\em Constructive Approximation}, 35(3):363--417, 2012. \bibitem{tolstikhin2017minimax} Ilya Tolstikhin, Bharath K Sriperumbudur, and Krikamol Muandet. \newblock Minimax estimation of kernel mean embeddings. \newblock {\em The Journal of Machine Learning Research}, 18(1):3002--3048, 2017. \bibitem{van2006targeted} Mark J Van Der Laan and Daniel Rubin. \newblock Targeted maximum likelihood learning. \newblock {\em The International Journal of Biostatistics}, 2(1), 2006. \bibitem{welling2009herding} Max Welling. \newblock Herding dynamical weights to learn. \newblock In {\em International Conference on Machine Learning}, pages 1121--1128, 2009.