EconBase
← Back to paper

Regularized Orthogonal Machine Learning for Nonlinear Semiparametric Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

104,610 characters · 13 sections · 104 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Regularized Orthogonal Machine Learning for Nonlinear Semiparametric Models

abstractThis paper proposes a Lasso-type estimator for a high-dimensional sparse parameter identified by a single index conditional moment restriction (CMR). In addition to this parameter, the moment function can also depend on a nuisance function, such as the propensity score or the conditional choice probability, which we estimate by modern machine learning tools. We first adjust the moment function so that the gradient of the future loss function is insensitive (formally, Neyman-orthogonal) with respect to the first-stage regularization bias, preserving the single index property. We then take the loss function to be an indefinite integral of the adjusted moment function with respect to the single index. The proposed Lasso estimator converges at the oracle rate, where the oracle knows the nuisance function and solves only the parametric problem. We demonstrate our method by estimating the short-term heterogeneous impact of Connecticut's Jobs First welfare reform experiment on women's welfare participation decision.

Introduction

Conditional moment restrictions (CMRs) often emerge as natural restrictions summarizing conditional independence, exclusion, or structural assumptions in economic models. In such a setting, a major challenge is to allow for a data-driven selection among a (very) large number of conditioning covariates. The sparsity assumption, which requires the number of relevant covariates to be small, has appeared to be an interpretable and plausible alternative. If the target parameter minimizes some population loss, a natural way to impose sparsity is to add an $\ell_1$-penalty on the target parameter to the sample loss. As discussed in EHT, this Lasso approach has substantial computational and statistical advantages over its alternatives. However, in general, the loss may be difficult to find. Focusing on a class of single index CMRs (Ichimura:93, kleinspady), we make the Lasso approach feasible by deriving the loss for an arbitrary single index moment function and establish convergence rate for the Lasso estimator.

The starting point of the analysis is a single index CMR that, in addition to target parameter, may depend on a nuisance component, such as the propensity score, the conditional choice probability, the conditional density, or alike, that can be substantially more high-dimensional than the target parameter itself. A natural approach would be to plug a machine learning/regularized estimate of the nuisance parameter into the moment function, such as random forest or neural networks. We first adjust the moment function so that the gradient of the future loss function is insensitive (i.e., Neyman-orthogonal, chernozhukov2016double) with respect to the first-stage regularization bias, preserving the single index property. We then take the loss function to be an indefinite integral of the adjusted moment function with respect to the single-index. If the original moment function is monotone in the single-index, we report the global minimum of $\ell_1$-regularized $M$-estimator loss, following Negahban. Under mild conditions, the proposed estimator converges at the oracle rate, where the oracle knows the true value of the nuisance parameter and solves only the parametric problem.

We demonstrate the utility of our method with theoretical and empirical applications. First, we introduce a partially linear logistic model with heterogeneous treatment effects and derive an orthogonal loss. Second, we derive an orthogonal loss in conditional moment models with missing data, as studied in carroll:95, carroll:91, chen:08, lee:95, stepanski:93 and static games of incomplete information (e.g. see bajari:10 and bajari:13 among others). In all these settings, we give sufficient primitive conditions on the nuisance parameters to achieve oracle convergence. In the empirical application, we study the heterogeneous Jobs First effect on welfare participation decision via a partially linear logistic model. To detect treatment effect's heterogeneity, it is essential to use the orthogonal loss rather than the non-orthogonal one.

\paragraph{Literature review. }This paper is related to three lines of research: single index models, high-dimensional sparse models, and orthogonal/debiased inference based on machine learning methods. The first line of research concerns with the estimation of single index models (manski:75, powell, manski:85, Ichimura:93, kleinspady, ichimura:hall). This work focuses on a low-dimensional target parameter that can be treated as fixed. Focusing on smooth models, we allow the parameter's dimension to grow with sample size and even exceed it. We show that the single index property is sufficient to ensure the uniform convergence of the sample moments towards its population analog over an $\ell_1$-restricted ball, extending the generalization bounds in the machine learning literature (see e.g., Shalev2014) to single index CMRs.

The second line of research establishes the finite-sample bounds for a high-dimensional sparse parameter (belloni:11, Negahban, loh:13, geer, loh2017, zhu:17, Zhu2017). Our contribution is to allow the loss function to depend on a functional nuisance parameter. In the convex case, we establish global convergence of the $\ell_1$-regularized $M$-estimator, following Negahban. One could follow a similar path to establish local convergence, building on loh:13 and loh2017, in the non-convex case. Focusing on the double robustness property, tan2017regularized and tan2018modelassisted establish convergence guarantees in high-dimensional models that are potentially misspecified. After we released the working paper version of this article (arxiv.ID 1702.06240), many methods have proposed similar $M$-estimation approaches. Focusing on missing data, chakrabortty2019high develops an $M$-estimator, relying on a classical Robins orthogonal score. We derive an orthogonal loss function for an arbitrary single index moment restriction, including Robins as a special case. The follow-up paper by foster2019orthogonal extends our result to $M$-estimators with decomposable regularizers, including $\ell_1$ penalty as a special case. An alternative regularized minimum distance approach has been proposed in belloni2018highdimensional.

The third line of research obtains a $\sqrt{N}$-consistent and asymptotically normal estimator of a low-dimensional target parameter $\theta$ in the presence of a nonparametric nuisance parameter (Neyman:1959, Neyman:1979, HardleStoker1989,bickel:93, NeweyStoker, andrews1994,Newey1994, Robins, chen:03). A statistical procedure is called Neyman-orthogonal if it is locally insensitive with respect to the estimation error of the first-stage nuisance parameter. Combining orthogonality and sample splitting, the Double Machine Learning framework of LRSP and chernozhukov2016double has derived a root-N consistent asymptotically normal estimator of the target parameter based on the first-stage machine learning estimates. Extending this work, we establish convergence rates for $\ell_1$-regularized $M$-estimators whose loss function gradient is an orthogonal moment. Next, we also contribute to the literature that derives orthogonal moments starting from non-orthogonal ones. Specifically, we construct a bias correction term for a nuisance parameter that is identified by a general conditional exogeneity restriction, covering conditional expectation (Newey1994) and conditional quantile (Newey:2018) as leading special cases. An alternative approach based on automatic debiasing has been proposed in chernozhukov2021debiased and chernozhukov2021automatic. Finally, we also contribute to the growing literature on orthogonal/doubly robust estimation based on machine learning methods (sasaki2018estimation, chiang2019multiway, sasaki2020unconditional, Chiang), in particular, heterogeneous treatment effects estimation (nie2017, CGST, OprescuWu, Lieli, ZimLech, Colangelo, CherSem). In contrast to this work, our main example features partialling out inside the argument of nonlinear link function, which, to the best of our knowledge, is completely new.

This paper is organized as follows. Section (ref) introduces our main examples and gives a non-technical overview of the results. Section (ref) formally states our results. Section (ref) derives the concrete conditions for our applications. Section (ref) gives an empirical application. Appendix (ref) generalizes Theorem (ref) to the case of extremum estimators beyond $M$-estimators. Appendix (ref) proves Theorem (ref) and Theorem (ref). Appendix (ref) verifies the conditions of Theorem (ref) for each of the three applications.

Set-Up

We start with the description of a single index conditional moment restriction (CMR) framework. The CMR takes the form

align[align omitted — 160 chars of source]

where the first argument $W \in \mathcal W \subseteq \mathrm{R}^{\text{dim} W}$ is the data vector, the second argument $t \in \mathrm{R}$ is the single index, and the third one $\gamma \in \mathrm{R}^d$ is the output of the functional nuisance parameter. The parameter of interest $\theta \in \mathrm{R}^p$ enters the moment function only via its inner product

align*[align* omitted — 41 chars of source]

where the index function $\Lambda(z,\gamma): \mathcal{Z} \bigtimes \mathrm{R}^d \rightarrow \mathrm{R}^p$ is known up to $\gamma$. In many cases (e.g., Ichimura:93, kleinspady), the index function $\Lambda(z,\gamma)$ reduces to $\Lambda(z,\gamma)=z$. Examples of the nuisance vector-function $$ g_0 = g_0(z) $$ include the propensity score, the conditional choice probability, and the regression function, or a combination of these functions. Given the CMR (ref), our goal is to find a loss function $Q(\theta, \gamma): \mathrm{R}^p \bigtimes \mathrm{R}^d \rightarrow \mathrm{R}$ so that the true parameter value $\theta_0$ obeys

equation[equation omitted — 80 chars of source]

We find the loss function in two steps. We first solve the ordinary differential equation (ODE)

align[align omitted — 124 chars of source]

We then plug $t=\Lambda(Z,\gamma)'\theta$ and $\gamma = g(Z)$ into the sketch of the loss function $\ell(w, t, \gamma)$ to obtain a population loss

align[align omitted — 120 chars of source]

whose gradient at $g=g_0$ and $\theta = \theta_0$ is

align[align omitted — 182 chars of source]

In what follows, we refer to $\ell (w, t, \gamma)$ as the loss sketch, to $Q(\theta, g)$ and $\nabla_{\theta} Q(\theta, g)$ as the population loss and the population gradient, to

align[align omitted — 158 chars of source]

and to $\nabla_{\theta} \widehat{Q}(\theta, \widehat{g})$ as the sample loss and the sample gradient. The conditional independence assumption (ref) ensures that (ref) is a valid moment equation for $\theta_0$. If $m(w, t, \gamma)$ is non-decreasing (non-increasing) in $t$, $\theta_0$ is the unique minimizer of $Q(\theta, g_0)$ ($-Q(\theta, g_0)$).

The proposed estimator $\widehat{\theta}$ has two stages. First, on the auxiliary sample, we construct an estimate $\widehat{g}$ of the nuisance parameter $g_0$, using a machine learning estimator capable of dealing with the high-dimensional covariate vector $Z$. Second, on the main sample, the target parameter's estimate is taken to be the minimizer of $\ell_1$-regularized sample loss. For a fixed vector, its $\ell_2$ norm is denoted by $\| \cdot \|_2$, the $\ell_1$ norm is denoted by $\| \cdot \|_1$, the $\ell_{\infty}$ norm is denoted by $\| \cdot \|_{\infty}$, and $\ell_0$ norm is denoted by $\| \cdot \|_{0}$.

definition[Regularized $M$-Estimator] Given $(\widehat{g}(Z_i))_{i=1}^n$ and the penalty parameter $\lambda \geqslant 0$, define \begin{align} \widehat{\theta} & =: \arg \min_{\theta \in \mathrm{R}^p} \widehat{Q}(\theta, \widehat{g}) + \lambda \| \theta \|_1. \end{align}

As discussed in Negahban, the sample gradient $\nabla_{\theta} \widehat{Q} (\theta_0, \widehat{g})$ at $\theta = \theta_0$ summarizes the noise of the problem. If the population gradient (ref) possesses the orthogonality property (Neyman:1959)

align[align omitted — 146 chars of source]

the biased first-stage estimation error $\widehat{g}(Z) - g_0(Z)$ has no first-order effect on the sample gradient. As a result, there exists a moderate penalty choice $\lambda=\lambda_{\text{mod}}$ obeying

align[align omitted — 105 chars of source]

that is sufficiently large to dominate the noise

align[align omitted — 162 chars of source]

If (ref) does not hold, the event (ref) requires an aggressive choice $\lambda=\lambda_{\text{agg}}$ obeying

align[align omitted — 105 chars of source]

To sum up, for $C$ sufficiently large, there exists an admissible penalty level $\lambda = \lambda_{\text{adm}} $ obeying (ref)

align[align omitted — 115 chars of source]

which reduces to $\lambda_{\text{mod}} $ if (ref) holds (i.e., $B_0 = 0$) and to $\lambda_{\text{agg}} $ otherwise. For an admissible penalty choice, Theorem (ref) establishes the following bounds

align[align omitted — 156 chars of source]

In particular, if $B_0 = 0$ and $g_n^2= o( \log p/n)^{1/2}$ and $\lambda =\lambda_{\text{adm}}$ obeying (ref), $\widehat{\theta}$ converges at the oracle rate, where the oracle knows the nuisance parameter $g_0$ and estimates only $\theta_0$. In what follows, if the condition (ref) holds, we refer to the $m(w, t, \gamma)$ and $\ell(w, t, \gamma)$ as the orthogonal moment and orthogonal loss, respectively.

We conclude this section by studying the partially linear logistic model. The model takes the form

align[align omitted — 131 chars of source]

where $D \in \mathbb{R}$ is a one-dimensional base treatment, $X \in \mathbb{R}^p$ is a vector of controls, $Y$ is a binary outcome, $W=(D,X,Y)$ is the data vector, and $G(t)$ is the logistic link function. The treatment variable $D$ affects the outcome $Y$ via its interactions with the controls $(1,X)$. In addition, $X$ affects $Y$ via the confounding function $f_0(X)$ that enters (ref) in an additively separable way. To make progress, most papers (see e.g., belloni2016postselection) require this function to be linear so that $\theta_0$ and the nuisance parameter can be estimated under a joint sparsity assumption. We describe below how to circumvent this bottleneck.

Inspired by robinson:88, we propose to partial out the controls inside the link function's argument

align[align omitted — 140 chars of source]

where

align[align omitted — 62 chars of source]

is the conditional expectation of the treatment and

align[align omitted — 97 chars of source]

is the conditional expectation of the link function's argument. In contrast to (ref), the nuisance parameters $p_0(x)$ and $q_0(x)$ are identified separately from $\theta_0$ and permit a wider class of approaches to estimate them. Thus, equation (ref) is a special case of (ref) with the conditioning vector $Z=(D,X)$, the moment function

align[align omitted — 79 chars of source]

the single index $t= (d - \gamma_1) \cdot (1,x)'\theta$ and the nuisance parameter $g_0(x) = \{ p_0(x),q_0(x) \}$ whose output is denoted by $\gamma = (\gamma_1, \gamma_2)$.

One may be tempted to proceed with the moment function (ref). Solving the ODE (ref) gives

align[align omitted — 120 chars of source]

The sample loss $\widehat{Q}(\theta, (\widehat{p}, \widehat{q}))$ in (ref) coincides with the negative logistic likelihood. However, the population gradient (ref) does not obey the orthogonality condition (ref)

align[align omitted — 239 chars of source]

As a result, the biased estimation error $\widehat{q}(X)-q_0(X)$ has a first-order effect on the sample gradient. Thus, the admissible penalty choice reduces to $\lambda_{\text{adm}}=\lambda_{\text{agg}}$ in (ref), which makes the estimator's rate in (ref) a slow one.

To restore orthogonality, we reweigh the moment function as in belloni2016postselection. The sketch of the new moment function is

align[align omitted — 95 chars of source]

where $\gamma=(\gamma_1,\gamma_2,\gamma_3)$ corresponds to the output of $g_0(z) = \{ p_0(x),q_0(x),V_0(d,x)\}$ and the weighting function $V_0(d,x)$ is the conditional variance

align[align omitted — 143 chars of source]

Solving the ODE (ref) gives

align[align omitted — 204 chars of source]

The sample loss coincides with the negative weighted logistic likelihood (belloni2016postselection). In contrast to (ref), the population gradient obeys the orthogonality condition (ref)

align[align omitted — 287 chars of source]

As a result, the admissible penalty choice (ref) reduces to $\lambda_{\text{adm}}=\lambda_{\text{mod}}$ in (ref), which makes the estimator's rate in (ref) a fast one.

Examples

example[Nonlinear Treatment Effects] Suppose the treatment effect $\theta_0$ is identified by the conditional moment restriction (ref), where the link function $G(\cdot): \mathbb{R} \rightarrow \mathbb{R}$ is a known monotone link function that may not necessarily be logistic. Define the weighting function as \begin{align} V_0(d,x) = G'( (d-p_0(x)) \cdot ((1,x)'\theta_0) + q_0(x)). \end{align} The loss sketch $\ell(w,t,\gamma)$ is an arbitrary solution to the ODE \begin{align} \dfrac{\partial}{\partial t}\ell(w,t,\gamma) = -\dfrac{1}{\gamma_3} \bigg(y-G( t + \gamma_2 ) \bigg), \end{align} where $\Lambda(z, \gamma) = (d - \gamma_1) \cdot (1,x)$, $\gamma = (\gamma_1, \gamma_2, \gamma_3)$ denotes the output of $g_0(z) = \{ p_0(x), q_0(x), V_0(d,x)\}$ defined in (ref), (ref) and (ref). Corollary (ref) establishes the mean square convergence rate for the Regularized $M$-Estimator based on the loss sketch defined in the ODE (ref).
remark[Linear Link] Consider Example (ref) with $G(t)=t$. The function $V_0(d,x)$ in (ref) simplifies to $$V_0(d,x)=1.$$ As a result, the nuisance parameter $g_0(z)$ simplifies to $g_0(z) = g_0(x) = \{ p_0(x), q_0(x)\}$ and $\gamma = (\gamma_1, \gamma_2)$. The moment sketch is \begin{align*} m(w,t, \gamma):= -(y - t-\gamma_2). \end{align*} The loss sketch $\ell(w,t,\gamma)$ is \begin{align*} \ell(w,t,\gamma) &= \dfrac{1}{2}(y - t - \gamma_2)^2, \end{align*} which corresponds to the least squares loss used in CGST and nie2017. The population gradient (ref) reduces to robinson:88-type score \begin{align} \nabla_{\theta} Q(\theta_0, g_0) &= \mathbb{E} (Y - (D-p_0(X)) \cdot (1,X)'\theta_0 - q_0(X)) \cdot (D-p_0(X)) \cdot (1,X) =0. \end{align} Its pathwise derivative with respect to $q$ is zero: \begin{align} &\dfrac{\partial}{\partial r} \nabla_{\theta} Q ( \theta_0, r (q - q_0) + q_0) =- \mathbb{E} (D-p_0(X)) (q (X) -q_0(X)) \cdot (1,X)' =0. \end{align}
remark[Logistic Link] Consider Example (ref) with $G(t)=(1+ \exp^{-t})^{-1}$. The moment sketch is \begin{align} m(w, t, \gamma):=\dfrac{y - G( t +\gamma_2)}{ \gamma_3 }. \end{align} The loss sketch $\ell(w,t,\gamma)$ is \begin{align*} \ell(w, t, \gamma) &= -\dfrac{1}{\gamma_3} \big( y\cdot \log\left(G\left( t+ \gamma_2 \right)\right) + (1-y)\cdot \log\left(1-G((t+ \gamma_2)\right) \big), \end{align*} which corresponds to the negative weighted logistic likelihood used in belloni2016postselection. The population gradient is \begin{align*} \nabla_{\theta} Q(\theta_0, g_0) =-\mathbb{E} \dfrac{ (Y - G ((D-p_0(X)) \cdot (1,X)'\theta_0 + q_0(X)) }{V_0(D,X)} \cdot (D-p_0(X)) \cdot (1,X), \end{align*} where $g_0(z) = (p_0(x),q_0(x), V_0(d,x))$ is as defined in (ref), (ref), (ref).
example[Missing Data] Suppose a researcher is interested in the parameter $\theta_0$ identified by a CMR: \begin{equation} \mathbb{E}[u(Y^{*}, X'\theta_0) | X=x]=0, \quad \forall x \in \mathcal{X}, \end{equation} where $Y^{*} \in \mathbb{R}$ is a partially observed outcome and $Z=X \in \mathbb{R}^p$ is a covariate vector. Let $V \in \{1,0\}$ indicate whether $Y^{*}$ is observed, $Y=V \cdot Y^{*}$ be the observed outcome, and $W=(V, X, Y)$ be the data vector. A standard way to make progress is to assume that $V$ is as good as randomly assigned conditional on $X$. Define the conditional probability of observing $Y^{*}$ as $$p_0(x) = \mathbb{E}[ V|X=x] $$ and the expectation function $h_0(x)$ as $$h_0(x) = \mathbb{E}\left[u(Y, X'\theta_0)\,|\,X=x, V=1\right].$$ The moment sketch is \begin{align} m(w, t, \gamma) = \dfrac{v\,}{\gamma_1}u(y, t) -\dfrac{\gamma_2}{\gamma_1} ( v- \gamma_1), \end{align} where $\gamma=(\gamma_1, \gamma_2)'$, $\Lambda(x, \gamma) = x$ and $g_0(x)=\{p_0(x), h_0(x)\}$. The loss sketch is \begin{equation} \ell(w, t, \gamma) = \dfrac{v\,}{\gamma_1} \ell_{pre}(w, t,\gamma_1) - \dfrac{\gamma_2}{\gamma_1} (v-\gamma_1)\, t, \end{equation} where $\ell_{\text{pre}} (w, t,\gamma_1)$ is an arbitrary solution to the ODE \begin{align} \dfrac{\partial}{\partial t} \ell_{pre}(w,t,\gamma_1) =u(y, t). \end{align} The loss (ref) is a special case of the loss (ref) proposed in Theorem (ref). Corollary (ref) establishes mean square convergence rate for the Regularized $M$-Estimator based on the loss function (ref).
remark[Quantile Regression with Missing Data] Consider Example (ref) with \begin{align} u_{\tau}(y,t) = -(1_{ \{ y \leqslant t \}} - \tau), \end{align} where $\tau \in (0,1)$ is a quantile level. The moment function (ref) identifies a quantile treatment effect parameter $\theta_0$. The loss sketch (ref) takes the form \begin{equation} \ell(w, t, \gamma) := - \left( \tau \cdot (y-t) 1_{ \{ y>t \}} + (1-\tau) \cdot (t-y) 1_{ \{ y<t \}} \right) \dfrac{v}{\gamma_1} -\dfrac{\gamma_2}{\gamma_1} \,(v-\gamma_1)\, t. \end{equation}
example[Static Games of Incomplete Information] Consider a two-player binary choice static game of incomplete information. The utility of action for player one is \begin{align} U (1) &= X'\alpha_0 + V \cdot \Delta_0+ \epsilon, \quad \mathbb{E}[ \epsilon | X] =0, \end{align} where $X \in \mathbb{R}^p$ is the covariate vector, $V \in \{1, 0\}$ is the opponent action, $\epsilon$ is mean independent private shock that follows Gumbel distribution. The target $p$-vector $\theta_0=(\alpha_0, \Delta_0)$ consists of the covariate effect $\alpha_0$ and interaction effect $\Delta_0$. The utility $U(0)$ of non-action for player one is normalized to zero. If the players' choices correspond to Bayes-Nash equilibrium, the outcome $Y$ obeys \begin{align*} Y &= 1[ X'\alpha_0 + p_0(X) \Delta_0 + \epsilon >0 ], \end{align*} where $p_0(x) = \mathbb{E}[ V|X=x]$. The moment sketch is \begin{align*} m(w,t, \gamma) = -(y -G ( t ) + \gamma_2 (v - \gamma_1)), \end{align*} where $\gamma=(\gamma_1, \gamma_2)'$, $\Lambda(x,\gamma) = (x, \gamma_1)$, and $g_0(x) = \{ p_0(x), h_0(x) \}$ for $h_0(x) = \Delta_0 G' ( x'\alpha_0 + p_0(x) \Delta_0 ).$ The loss sketch is \begin{align} \ell(w, t, \gamma ) :&= \ell_{pre}(w, t, \gamma_1)- \gamma_2\, (v - \gamma_1)\, t, \end{align} where the preliminary loss is the negative logistic likelihood \begin{align} \ell_{pre}(w, t, \gamma_1) = - y \cdot G (t) - (1-y) \cdot (1-G (t) ). \end{align} The loss (ref) is a special case of the loss (ref) proposed in Theorem (ref). Corollary (ref) establishes mean square convergence rate for the Regularized $M$-Estimator based on the loss function (ref).

Remarks

Following the sparsity bounds established in belloni2016postselection, we conjecture that the final estimator $\widehat{\theta}$ is sparse. Therefore, one can interpret the Lasso estimator $\widehat{\theta}$ as a model selector. We expect the post-Lasso-logistic based on single-selection procedure

align*[align* omitted — 191 chars of source]

to be too sensitive to moderate model selection mistakes, occurring when the non-zero coefficients of $\theta_0$ are statistically indistinguishable from zero. According to LeebPotcher, such inference procedures do not provide a Gaussian approximation that is uniform over the space of $\theta$ and $g_0$ and are not honest. Instead, we discuss the following the debiasing procedure of geer.

remark[Debiased Lasso of geer] Suppose $t \rightarrow m(w, t, \gamma)$ is strictly increasing in $t$ for any $w$ and $\gamma$. Abstracting away from any nuisance components, or, effectively, treating $g_0$ as known, the work of geer proposes a debiased Lasso estimator (eq. 18): \begin{align} \widehat{\theta}_{debiased} (g_0)= \widehat{\theta} -\widehat{\Gamma} \frac{1}{n} \sum_{i=1}^n m (W_i, \Lambda(Z_i, g_0(Z_i))'\widehat{\theta},g_0) \cdot \Lambda(Z_i, g_0(Z_i)), \end{align} where $\widehat{\theta} $ is a preliminary estimator of $\theta_0$, $\Gamma =(\nabla_{\theta \theta} Q(\theta, g_0))^{-1}$ is the population Hessian inverse, and $\widehat{\Gamma} $ is the estimator of $\Gamma$. If the matrix $\Gamma$ is sparse, the estimator $\widehat{\Gamma} $ can be constructed by nodewise regression. Unlike the post-selection estimator, $\widehat{\theta}_{\text{debiased}} (g_0)$ is Neyman-orthogonal with respect to the bias in the estimation error of $\widehat{\theta}-\theta_0$ and $ \widehat{\Gamma}-\Gamma_0$. If the sparsity indices of $\Gamma$ and the parameter $\theta_0$ are sufficiently small, geer shows that the estimator $\widehat{\theta}_{\text{debiased}} (g_0)$ is asymptotically Gaussian. We conjecture that the plug-in estimator $$\widehat{\theta}_{\text{debiased}} (\widehat{g})$$ continues to be asymptotically Gaussian. For the linear link function, the asymptotic normality of $\widehat{\theta}_{\text{debiased}} (\widehat{g})$ has been established in CGST.

Theoretical results

\paragraph{Notation. } We will use the following notation. For two sequences of random variables $a_n, b_n, n \geqslant 1: a_n \lesssim_{P} b_n$ means $a_n = O_{P} (b_n)$. For two sequences of numbers $a_n, b_n, n \geqslant 1$, $a_n \lesssim b_n$ means $a_n = O (b_n)$. Let $T = \{ j: \quad \theta_{0, j} \neq 0 \}$ be the set of coordinates of $\theta_0$ that are not equal to zero, and let $T^c$ be the complement of $T$. For a vector $\delta \in \mathbb{R}^p$, let $(\delta_T)_j = \delta_j$ for each $j \in T$ and $(\delta_T)_j = 0$ for $j \in T^c$. For a vector-valued function $g(z):\mathcal{Z} \rightarrow \mathbb{R}^{d}$, denote its $\ell_2$ and $\ell_{\infty}$ norms as

align[align omitted — 280 chars of source]

Furthermore, assume that the index function $\Lambda(z,\gamma): {\mathcal Z} \bigtimes \mathbb{R}^d \rightarrow \mathbb{R}^p$ is sufficiently smooth with respect to $\gamma$, so that the gradient $ \nabla_{\gamma} \Lambda_j(z, \gamma)$ and the Hessian $ \nabla_{\gamma \gamma} \Lambda_j(z, \gamma)$ of each coordinate $j \in \{1,2,\dots, p\}$ are well-defined. Finally, we will use the empirical process notation $${\mathbb{E}_n} f(W_i) := \dfrac{1}{n} \sum_{i=1}^n f(W_i).$$

assumption[Monotonicity in single index] The moment function $m(w, t, \gamma)$ is non-decreasing in $t$ for any $w \in \mathcal W$ and any $\gamma \in \Gamma$.

Assumption (ref) ensures that the loss sketch $\ell(w,t,\gamma)$ is a convex function of $t$ for any $w$ and $\gamma$. As a result, the sample loss $\theta \rightarrow \widehat{Q}(\theta, \widehat{g})$ defined in (ref) is a convex function of $\theta$. By convexity, on the event (ref), the vector of errors $ \nu = \widehat{\theta} - \theta_0 $ belongs to the restricted cone

align[align omitted — 139 chars of source]

as shown in Negahban (see Lemma (ref)). Define the restricted set as

align[align omitted — 137 chars of source]

Define the sample curvature as

align[align omitted — 174 chars of source]

and let the population curvature be the analog of (ref) based on the population Hessian $\nabla_{\theta \theta} Q(\theta, g_0) $ instead of $\nabla_{\theta \theta} \widehat{Q}(\theta, \widehat{g})$.

assumption[Identification] Let $\Sigma$ denote the population covariance matrix of the index function as \begin{align} \Sigma &= \mathbb{E} \Lambda(Z, g_0(Z))\Lambda(Z, g_0(Z))^T. \end{align} Assume that there exists a constant $C_{\text{min}}>0$ so that $\min \operatorname{eig} \Sigma \geqslant C_{\text{min}}$.
assumption[Bounded derivative on $\mathbb{B}$] There exists a constant $B_{\text{min}}>0$ so that the following bound holds: $$\inf_{\theta \in \mathbb{B} }\mathbb{E} \bigg[\nabla_{t} m(W, t , \gamma)\bigg|_{\gamma = g_0(Z), t=\Lambda(Z, \gamma)'\theta}\,|\,Z=z \bigg] \geqslant B_{\text{min}}>0 \quad \forall z \in \mathcal{Z}. $$

Assumptions (ref) with $C_{\text{min}}$ and (ref) with $B_{\text{min}}$ imply that the population curvature is bounded from below by $\bar{\gamma} = B_{\text{min}} \cdot C_{\text{min}}$ (see Lemma (ref)). On the event

align[align omitted — 210 chars of source]

the sample curvature (ref) is bounded from below by $\bar{\gamma}/2$ (see Lemma (ref)). Since

align*[align* omitted — 295 chars of source]

it suffices to show that each summand above is $o_P(1)$. We bound the first and the second summand in Lemmas (ref) and (ref), respectively.

Assumption (ref) ensures that the moment function is sufficiently smooth in its second and third arguments. Let $\mathcal W$ be an open bounded set containing the support of the data vector $W$. Likewise, let $\mathcal{T}$ be an open set containing the support of $\Lambda(Z,g_0(Z))'\theta_0$. Finally, let $\Gamma$ be an open bounded set containing the support of vector $g(Z)$, when $g \in {\mathcal G}_n$. In what follows, we assume that the sets $\mathcal W, \mathcal{T}, \Gamma$ do not change with $n$.

assumption[Smooth and bounded design] There exists a constant $U< \infty$ so that $\sup_{\theta \in \Theta} \| \theta \|_1 \leqslant U$ and for any vector $w\in\,\mathcal W$, number $\ t\in\,\mathcal{T}$ and vector $\gamma \in \Gamma$ the following conditions hold: \begin{align*} |m(w, t, \gamma )|, \|\nabla_{\gamma} m(w, t, \gamma )\|_{\infty},\, |\nabla_{t} m(w, t, \gamma)| \leqslant & U,\\ \|\nabla_{\gamma t} m(w, t, \gamma )\|_{\infty}, \| \nabla_{\gamma \gamma} m(w, t, \gamma )\|_{\infty} , |\nabla_{tt} m(w, t, \gamma)|\leqslant & U,\\ \|\Lambda_j(z, \gamma)\|_{\infty}, \|\nabla_{\gamma} \Lambda_j(z, \gamma)\|_{\infty}, \|\nabla_{\gamma \gamma} \Lambda_j(z, \gamma)\|_{\infty} \leqslant & U, \quad \forall j \in \{1,2,\dots, p\}. \end{align*}

Lemma (ref) establishes a convergence rate for the sample Hessian uniformly over the restricted set (ref). A gradient version of Lemma (ref) is available in the concurrent work by belloni2018highdimensional.

lemma[Uniform Convergence of Sample Hessian] Suppose Assumption (ref) holds with a constant $U$. Then, there exists a sequence $\tau_n = 4 k U^5 \sqrt{\dfrac{2\log p}{n}} + 2U^3 \sqrt{\dfrac{2\log p}{n}}=o(1)$ so that $\nabla_{\theta \theta} \widehat{Q}(\theta, g_0) $ uniformly converges to $\nabla_{\theta \theta} Q(\theta,g_0)$: \begin{align} \sup_{\theta \in \mathrm{B}} \| \nabla_{\theta \theta} \widehat{Q}(\theta, g_0) - \nabla_{\theta \theta} Q(\theta, g_0) \|_{\infty} = O_{P} (\tau_n) = o_P(1). \end{align}
proofConsider the function class $${\mathcal F}=\{Z \rightarrow \Lambda(Z, g_0(Z))'\theta, \quad \theta \in \mathbb{R}^p, \quad \|\theta\|_1\leqslant k U \}.$$ Define its Rademacher complexity $$ \mathcal{R}({\mathcal F}) := \mathbb{E} \bigg[ \sup_{\theta: \|\theta\|_1\leqslant k U} {\mathbb{E}_n} \sigma_i \Lambda(Z_i, g_0(Z_i))'\theta \bigg], $$ where $(\sigma_i)_{i=1}^{n}$ is a vector of i.i.d random variables from Rademacher distribution: ${\mathrm{P}} (\sigma_i=+1) = {\mathrm{P}} (\sigma_i=-1)=0.5$. Each function in the class $${\mathcal H} = \{ W \rightarrow \nabla_{\theta_k \theta_j} \ell (W, \Lambda(Z, g_0(Z))'\theta, g_0(Z)), \quad \theta \in \mathbb{R}^p, \|\theta\|_1\leqslant k U \}$$ is a combination of the function in the class ${\mathcal F}$ and a $U^3$-Lipshitz function. Invoking Contraction Lemma (Lemma 26.9) and Lemma 26.11 from Shalev2014, we obtain $$ \mathcal{R}({\mathcal H}) \leqslant U^3 \mathcal{R}({\mathcal F}) \leqslant U^5 \cdot k \sqrt{\frac{2\log(2p)}{n}}. $$ Finally, invoking Lemma 26.5 Shalev2014 for each pair $(k, j) \in \{1,\ldots, p\}^2$, w.p. $1-\delta/p^2$, \begin{align*} &\sup_{ \|\theta\|_1\leqslant k U} \bigg| [{\mathbb{E}_n} - \mathbb{E} ][\nabla_{\theta_k \theta_j} \ell(W_i, \Lambda(Z_i, g_0(Z_i) )' \theta_0, g_0(Z_i)) ] \bigg| \\ &\leqslant 2 \mathcal{R}({\mathcal H}) + U^3 \sqrt{\frac{2\log(2\,p^3/\delta)}{n}}\\ &\leqslant 2 k U^5 \sqrt{\frac{2\log(2p^3/\delta)}{n}} + U^3 \sqrt{\frac{2\log(2\,p^3/\delta)}{n}}. \end{align*} Union bound over $p^2$ pairs $(k, j) \in \{1,\ldots, p\}^2$ and $2 p^3 \leqslant p^4$ gives w.p. $1-\delta$ \begin{align*} &\sup_{ 1 \leqslant k,j \leqslant p} \sup_{ \|\theta\|_1\leqslant k U} \bigg| [{\mathbb{E}_n} - \mathbb{E} ][\nabla_{\theta_k \theta_j} \ell(W_i, \Lambda(Z_i, g_0(Z_i) )' \theta_0, g_0(Z_i)) ] \bigg| \\ &\leqslant 4 k U^5 \sqrt{\frac{2\log(p/\delta)}{n}} + 2U^3 \sqrt{\frac{2\log(p/\delta)}{n}} \ \end{align*}

Assumption (ref) formalizes the convergence of the nuisance parameter's estimator. It introduces a sequence of nuisance realization sets ${\mathcal G}_{n} \subseteq \mathcal{G}$ that contain the true value $g_0$ and the estimator $\widehat{g}$ with probability $1-\delta_n$. As the sample size $n$ increases, the sets ${\mathcal G}_{n}$ shrink. The shrinkage speed is measured by the rate $g_{n}$ and is referred to as the first-stage rate.

assumption[First-stage rate] There exist sequences of numbers $\delta_n = o(1)$ and $g_n = o(1)$ and a sequence of sets ${\mathcal G}_n \subseteq {\mathcal G}$ so that $\widehat{g} \in {\mathcal G}_n$ w.p. at least $1-\delta_n$ and $g_0 \in {\mathcal G}_n$. The sets shrink at the rate \begin{align*} \sup_{g \in {\mathcal G}_n} \| g - g_0 \| \leqslant g_{n}, \end{align*} where $\| \cdot \|$ is either the $\ell_{\infty}$ norm or the $\ell_2$ norm as defined in equation (ref).

Assumption (ref) is satisfied by many machine learning estimators under structural assumptions on the model in $\ell_2$ and/or $\ell_{\infty}$ norm. For example, it holds for Lasso (belloni:11, belloni2013) in linear and generalized linear models, for $L_2$-boosting in sparse models (Luo), for neural network (ChenWhite1991), and for random forest in low-dimensional (wagerwalther) and high-dimensional sparse (fastRF) models. For a broad overview of low-level primitive conditions that covers neural networks and random forest see, e.g., jeong2020robust, Appendix 1.

theorem[Regularized $M$-Estimator] Suppose that the nuisance space $\mathcal G$ is equipped with either the $\ell_{\infty}$ or the $\ell_2$ norm, defined in equation (ref). Suppose Assumptions (ref)-(ref) hold, and $k \left(g_n + U^5 \cdot k \cdot \sqrt{\dfrac{ 2\log 2p}{n}} + 3 d \cdot U^3 \, \sqrt{\frac{\log(2p)}{n}}\right)=o(1)$. Then, for $n$ and $C$ large enough and the admissible penalty $\lambda = \lambda_{\text{adm}}$ is chosen to obey (ref), Regularized $M$-Estimator obeys the bound (ref).

Theorem (ref) is our main result. It establishes the convergence rate for the Regularized $M$-estimator, covering the orthogonal (i.e., $B_0 = 0$) and the non-orthogonal (i.e., $B_0 \neq 0$) cases. In the former case, the admissible penalty choice $ \lambda_{\text{adm}}$ reduces to $ \lambda_{\text{mod}}$ in (ref), which makes the estimator's rate (ref) fast. Otherwise, $ \lambda_{\text{adm}}$ reduces to $ \lambda_{\text{agg}}$ in (ref), which makes the estimator's rate (ref) slow. In all our examples, we invoke Theorem (ref) twice: first, for the preliminary estimate $\check{\theta}$ based on a non-orthogonal CMR (ref) and, second, for the final estimate $\widehat{\theta}$ based on the orthogonal CMR.

Suppose that the moment function $m_{\text{pre}} (w, t, \gamma_1): \mathcal W \bigtimes \mathbb{R} \bigtimes \mathbb{R} \rightarrow \mathbb{R}$ and/or the index function $\Lambda(z, \gamma_1): \mathcal W \bigtimes \mathbb{R} \rightarrow \mathbb{R}^p$ depend on a one-dimensional functional nuisance parameter $p_0(z)$:

align[align omitted — 155 chars of source]

Furthermore, suppose the nuisance parameter is identified by a conditional exogeneity restriction

align[align omitted — 82 chars of source]

For example, $ R(W, p(X)) = D - p(X) $ defines the conditional expectation function as in (ref), and $ R(W,p(X)) = 1_{ \{ V \leqslant p(X) \}} - \tau$ defines the conditional $\tau$-quantile function. Starting from an arbitrary CMR (ref), Theorem (ref) derives a loss whose gradient obeys the orthogonality condition (ref).

theorem[Construction of Orthogonal Loss] Suppose equation (ref) holds. Define the adjusted moment function \begin{equation} m(w, t, \gamma) = m_{pre}(w , t, \gamma_1) - \gamma_2 \cdot \gamma_3^{-1} \cdot R(w, \gamma_1), \end{equation} where $\gamma = (\gamma_1, \gamma_2, \gamma_3)$ denotes the output of the nuisance parameter $g_0(z) = \{ p_0(z), h_0(z), I_0(z) \}$, consisting of $p_0(z)$ as defined in (ref), $h_0(z)$ and $I_0(z)$ defined as \begin{align} h_0(z)&=\mathbb{E}\bigg[ \nabla_{\gamma_1} m (W,t,\gamma_1) \bigg|_{\gamma_1=p_0(Z), t= \Lambda(Z, \gamma_1)'\theta_0} \bigg|Z=z\bigg] \\ I_0(z) &= \mathbb{E} \bigg[ \nabla_{\gamma_1} R(W,\gamma_1) \bigg|_{\gamma_1 = p_0(Z)} \bigg|Z=z \bigg]. \end{align} The loss sketch $\ell(w,t, \gamma)$ takes the form \begin{align} \ell(w, t, \gamma) &= \ell_{pre}(w,t,\gamma_1) - \gamma_2 \cdot \gamma_3^{-1} \cdot R(w, \gamma_1) \cdot t, \end{align} where $\ell_{\text{pre}}(w,t,\gamma_1)$ solves the Ordinary Differential Equation (ref) for the moment function $m_{\text{pre}} (w,t,\gamma_1)$. Then, the population loss $Q(\theta, g_0)$ defined in (ref) obeys the orthogonality condition (ref). Suppose $\sup_{w \in \mathcal W} \sup_{\gamma_1 \in \Gamma_1} |R (w, \gamma_1)| \leqslant U$, $\sup_{w \in \mathcal W} \sup_{\gamma_1 \in \Gamma_1} |\nabla_{\gamma_1} R (w, \gamma_1)| \leqslant U,$ $$\sup_{w \in \mathcal W} \sup_{\gamma_1 \in \Gamma_1} |\nabla_{\gamma_1 \gamma_1} R (w, \gamma_1)| \leqslant U.$$ If the original moment function $m_{\text{pre}}(w, t, \gamma_1)$ obeys Assumptions (ref)-(ref), the adjusted moment function (ref) obeys Assumptions (ref)-(ref).

Starting from an arbitrary CMR (ref), Theorem (ref) adjusts the moment function so that population gradient $\nabla_{\theta} Q(\theta_0, g_0)$ obeys orthogonality condition (ref). Special cases of equation (ref), such as the conditional expectation function and conditional quantile function, are available in Newey1994, LRSP and Newey:2018 for unconditional moment problems. We extend the results above to allow for an arbitrary conditional exogeneity restriction.

Applications

In all the applications below, our starting point is the CMR (ref) which does not obey the orthogonality condition (ref). Invoking either the weighting idea (Example (ref)) or Theorem (ref) (Examples (ref)-(ref)), we derive an orthogonal CMR obeying (ref). However, this CMR requires an additional preliminary estimate of $\theta_0$ on top of the nuisance parameters involved in (ref). Definition (ref) describes the suggested two-step procedure, spelling out the choices of main/auxiliary samples and the penalty parameters in each step.

definition[Two-Step Regularized $M$-estimator] Let $K=3$ denote the $3$-fold partition of the sample $\{1,2,\dots, n\}$ into $J_1, J_2, J_3$. For notational convenience, let $J_0 = J_3$ and $J_4=J_1$. For each $k \in \{0,1,2\}$, compute \begin{enumerate} • With $J_{k-1}$ as the auxiliary sample and $J_k$ as the main sample, let $\check{\theta}$ be the output of Definition (ref) with the preliminary loss $\ell_{\text{pre}}(w, t, \gamma_1)$ and the penalty parameter $\lambda_{\text{agg}}$ as in (ref). For each $i \in J_{k+1}$, the nuisance parameter $\widehat{g}(Z_i)$ is evaluated at $\widehat{g}(Z_i) =\widehat{g}_k(Z_i) = \{ \widehat{p}_{k-1}(Z_i), \check{\theta}_k \}$. • With ($J_{k-1}, J_k)$ as the auxiliary sample and $J_{k+1}$ as the main sample, let $\widehat{\theta}$ be the output of Definition (ref) with the final loss $\ell(w, t, \gamma)$ and the admissible penalty parameter $\lambda_{\text{mod}}$ obeying (ref). Report: $\widehat{\theta}$ \end{enumerate}

Nonlinear Treatment Effects

Consider Example (ref). Define the covariance matrix

equation[equation omitted — 116 chars of source]
assumption[Regularity Conditions for Nonlinear Treatment Effects] Suppose the following conditions hold. \begin{enumerate} • (Identification) There exists a constant $\gamma_{\text{TE}}>0$ so that $\min \operatorname{eig} \Sigma_{\text{TE}} \geqslant \gamma_{\text{TE}}$. • (Smooth and Bounded Design). The parameter space $ \Theta$ is bounded in $\ell_1$ norm by a constant $H_{\text{TE}} \geqslant 1$: $\sup_{\theta \in \Theta} \| \theta \|_1 \leqslant H_{\text{TE}}$ and $\sup_{w \in \mathcal W} \| w \|_{\infty} \leqslant H_{\text{TE}}$. The functions $t \rightarrow G(t), G'(t), G^{(2)}(t), G^{3}(t)$ are $U_{\text{TE}}$-bounded on the set $[- 3 \cdot H_{\text{TE}}^3,3 \cdot H_{\text{TE}}^3]$, where $U_{\text{TE}} \geqslant H_{\text{TE}}$. Furthermore, the functions $t \rightarrow G^{-1}(t), (G')^{-1}(t), (G')^{-2}(t), (G')^{-3}(t)$ are $U_{\text{TE}}$-bounded from above on the set $[- 3 \cdot H_{\text{TE}}^3, 3 \cdot H_{\text{TE}}^3]$. Finally, $t \rightarrow G'(t)$ is $\text{B}_{\text{min}}$-bounded from below on the set $[- 3 \cdot H_{\text{TE}}^3, 3 \cdot H_{\text{TE}}^3]$. • (First-Stage Rate). There exists an estimator $(\widehat{p}(x), \widehat{q}(x))$ of $(p_0(x), q_0(x))$ satisfying Assumption (ref) with $\pi_{n}+q_{n}$ rates in either $\ell_{\infty}$ or $\ell_2$ norm. • (Monotonicity). The function $G(t): \mathbb{R} \rightarrow \mathbb{R}$ is a known monotone function of $t$. \end{enumerate}
remark[Weighting function estimator] For $k \in \{0,1,2\}$ and $i \in J_{k+1}$, define the estimator $\widehat{V}(D_i,Z_i )=\widehat{V}_k(D_i,Z_i)$ \begin{align} \widehat{V}_k(D_i,X_i)=G'((D_i-\widehat{p}_k(X_i)) \cdot (1,X_i)'\check{\theta}_k + \widehat{q}_k(X_i)), \quad i \in J_{k+1}. \end{align} where preliminary loss $\ell_{\text{pre}}(w, t, \gamma_1)$ as in (ref) is used for $\check{\theta}$.
corollary[Nonlinear Treatment Effects] Suppose Assumption (ref) holds. Let the preliminary loss $\ell_{\text{pre}}(w, t, \gamma_1)$ be as in (ref) and the first-stage parameter estimate $(\widehat{\pi}(x), \widehat{q}(x))$. Let the final loss be as in (ref) and $\widehat{g} (d,x)= (\widehat{\pi}(x), \widehat{q}(x), \widehat{V}(d,x))$, where $\widehat{V}(d,x)$ as defined in (ref).Then, the statement of Theorem (ref) holds for each step of Definition (ref), where the first-step nuisance rate is $O(\pi_{n}+q_{n})$ and the second-step nuisance rate is $g_n = O \left( k \left( \pi_n + q_n+\sqrt{\frac{\log p}{n}} \right) \right)$.

Missing Data

Consider Example (ref). Define the covariance matrix

align[align omitted — 69 chars of source]

and the function $q(t, x)$ as

align*[align* omitted — 72 chars of source]
assumption[Regularity Conditions for Missing Data] Suppose the following conditions hold. \begin{enumerate} • (Identification). There exists $\gamma_{\text{MD}}>0$ such that $\min \operatorname{eig} \Sigma_{MD} \geqslant \gamma_{\text{MD}}$. Furthermore, there exists $B_{\text{min}}>0$ so that $\inf_{ t \in \mathcal{T}} \mathbb{E} [ \nabla_{t} u(Y,t) | X=x] \geqslant B_{\text{min}}$ for any $x \in \mathcal{X}$. • (Overlap Condition). There exists a constant $\bar{p}>0$ such that $\inf_{x \in \mathcal{X} } p_0(x)\geqslant \underline{p}>0$. • (Smooth and Bounded Design). The parameter space $ \Theta$ is bounded in $\ell_1$ norm by a constant $H_{\text{MD}}$: $\sup_{\theta \in \Theta} \| \theta \|_{\infty} \leqslant H_{\text{MD}}$ and $\sup_{w \in \mathcal W} \| w \|_{\infty} \leqslant H_{\text{MD}}$. The functions $u(y,t)$,$\nabla_{t} u(y,t)$, $\nabla_{tt} u(y,t)$ are $U_{\text{MD}}$-bounded in $\ell_{\infty}$ norm for all $w \in \mathcal W$ and $t \in \mathcal{T}$, where $U_{\text{MD}}\geqslant H_{\text{MD}}$. • (First-Stage Rate). There exists an estimator $(\widehat{p}(x), \widehat{q}(x,t))$ of $(p_0(x), q_0(x,t))$ satisfying Assumption (ref) with $\pi_{n,r}+q_{n,r}$ rates in either $\ell_{\infty}$ or $\ell_2$ norm, where $q_{n,r}$ rate is \begin{align*} q_{n, \infty}:= \sup_{q \in \mathcal{Q}_n} \sup_{t \in R} \sup_{x \in \mathcal{X}} |q(x,t) - q_0(x,t)|, \quad q_{n, 2}:=\sup_{q \in \mathcal{Q}_n}\sup_{t \in R} (\mathbb{E} (q(X,t) - q_0(X,t))^2)^{1/2}. \end{align*} • (Monotonicity). The function $u(y,t)$ is increasing in $t$ for any $y \in \mathbb{R}$. \end{enumerate}
remark[First-Stage Estimator of $h_0(x)$] For each $i \in J_{k+1}$, define the estimator $\widehat{h}(X_i)=\widehat{h}_k(X_i)$ \begin{align} \widehat{h}(X_i)=\widehat{q}_k(X_i'\check{\theta}_k, X_i), \end{align} where the preliminary estimator $\check{\theta}$ is based on the loss $\ell_{\text{pre}}(w, t, \gamma_1)$ as in (ref).
corollary[General Moment Problems with Missing Data] Suppose Assumption (ref) holds. Let the preliminary loss $\ell_{\text{pre}}(w, t, \gamma_1)$ be as in (ref) and the first-stage parameter estimate $\widehat{p}(x)$. Let the final loss be as in (ref) and $\widehat{g} (x)= (\widehat{p}(x), \widehat{h}(x))$, where $\widehat{h}(x)$ as defined in (ref). Then, the statement of Theorem (ref) holds for each step of Definition (ref), where the first-step nuisance rate is $O(\pi_{n})$ and the second-step nuisance rate $g_n = O \left(k\left( \pi_n +\sqrt{\frac{\log p}{n}} \right)+q_n\right)$.

Static Games of Incomplete Information

Consider Example (ref). Define the covariance matrix

align[align omitted — 87 chars of source]
assumption[Regularity Conditions for Static Games of Incomplete Information] Suppose the following conditions hold. \begin{enumerate} • (Identification). There exists $\gamma_{\text{games}}>0$ so that $\min \operatorname{eig} \Sigma_{\text{games}} \geqslant \gamma_{\text{games}}$. • (Smooth and Bounded Design). There exists a constant $H_{\text{games}} < \infty$ so that $\sup_{\theta \in \Theta} \| \theta \|_{\infty} \leqslant H_{\text{games}}$ and $\sup_{x \in \mathcal{X}} \|x \|_{\infty} \leqslant H_{\text{games}}$. In addition, $\sup_{\theta \in \Theta} \| \theta \|_1 \leqslant H_{\text{games}}$. • There exists an estimator $\widehat{p}(x)$ obeying Assumption (ref) with $\pi_n$ rates in either $\ell_{\infty}$ or $\ell_2$ norm. \end{enumerate}
remark[First-Stage Estimator of $h_0(x)$] For each $i \in J_{k+1}$, define the estimator $\widehat{h}(X_i)=\widehat{h}_k(X_i)$ \begin{align} \widehat{h}(X_i)=\widetilde{\Delta}_k \cdot G_k' ( X_i' \check{\alpha}_k + \check{\Delta}_k \cdot \widehat{p}(X_i) ). \end{align} where the preliminary estimator $\check{\theta} =(\check{\alpha}, \check{\Delta})$ is based on the loss $\ell_{\text{pre}}(w, t, \gamma_1)$ as in (ref).
corollary[Games of Incomplete Information] Suppose Assumption (ref) holds. Let the preliminary loss $\ell_{\text{pre}}(w, t, \gamma_1)$ be as in (ref) and the first-stage parameter estimate $\widehat{p}(x)$. Let the final loss be as in (ref) and $\widehat{g} (x)= (\widehat{p}(x), \widehat{h}(x))$, where $\widehat{h}(x)$ as defined in (ref). Then, the statement of Theorem (ref) holds for each step of Definition (ref), where the first-step nuisance rate is $O(\pi_{n})$ and the second-step nuisance rate $g_n = O \left(k \left( \pi_n +\sqrt{\frac{\log p}{n}} \right) \right)$.

Empirical Application

In this section, we study the short-term impact of Connecticut's Jobs First welfare reform experiment on women's labor supply and welfare participation decisions. Jobs First, a welfare-to-work assistance program, was introduced in Connecticut in the late 1990s as an alternative to the federal Aid to Families with Dependent Children (AFDC) program. Imposing revealed preference restrictions, KT's nonparametric bounds show that Jobs First induced many women to work but led some others to reduce their earnings in order to receive assistance. To gain more insight into the question, we postulate a partially linear logistic model for women's welfare participation decision and estimate heterogeneous Jobs First effects.

The data for our analysis are the same as in KT. They come from Manpower Demonstration Research Corporation (MDRC). In 1996, MDRC conducted a randomized trial that randomly selected a set of eligible female applicants into Jobs First (the treatment group), leaving the remaining females eligible for AFDC (the control group). The outcome of interest is the binary indicator that is equal to one if a woman receives any type of welfare (e.g., AFDC or food stamps) at the fourth quarter after random assignment. Baseline characteristics include constant, age, education level, quarterly history of employment, earnings, AFDC, and food stamps during $8$ quarters before random assignment (RA). We postulate a logistic specification with a partially linear index

align[align omitted — 103 chars of source]

where $D=1$ if a woman is assigned to Jobs First and $Y=1$ if a woman is on welfare. We assume that the treatment $D$ and the controls $X$ affect the outcome $Y$ via an additively separable index

equation[equation omitted — 66 chars of source]

entering the logistic link function $G(\cdot): \mathbb{R} \rightarrow \mathbb{R}$. The base treatment $D$ enters the index (ref) through its interactions with the control vector $(1,X)$. In addition, the control vector $X$ enters the index (ref) through an unknown function $f_0(x)$, which summarizes the confounding effect of the controls on $Y$. Our main object of interest is the heterogeneous treatment effects vector $$\theta_0= (\theta_{0,1}, \theta_{0,-1})',$$ where $\theta_{0,1}$ is the baseline treatment effect and $\theta_{0,-1}$ is the treatment interaction effect. We assume that, out of $20$ treatment effect interactions, only few have non-zero value, but do not know their identities.

A standard approach to this problem is to require the confounding function $f_0(x)$ to be a linear sparse function of the controls

align[align omitted — 144 chars of source]

where $B(x)$ is a vector of basis functions of $x$ and $n$ is the sample size. We take $B(x)$ to be a vector of $p_{\alpha} = 1, 600$ pairwise interactions of the controls. The joint estimate of $\theta_0$ and $\alpha_0$ is

align[align omitted — 417 chars of source]

where the penalty parameter $\lambda_{\text{direct}}$ recommended by belloni2016postselection is $$ \lambda_{\text{direct}} = \dfrac{1.1}{2\sqrt{n}} \Phi^{-1}(1-0.05/ (p+p_{\alpha}) \log (n) ) = 0.036. $$ Following tan2018modelassisted, we do not penalize the first coordinate $\theta$, imposing the sparsity assumption on the heterogeneous modification effects but not the level of the treatment effect for the baseline category.

We compare the direct estimator to the Regularized $M$-Estimator $\widehat{\theta}$, described in Section (ref). This estimator no longer requires the control function $f_0(x)$ to be sparse with respect to the chosen basis $B(x)$. Instead, we assume that the conditional probability of welfare $$ G_0(d,x)= \mathbb{E}[ Y | D=d, X=x] $$ can be well-approximated by trees. While this assumption is less interpretable, it allows $G_0(d,x)$ to have a nonlinear argument. We estimate $G_0(d,x)$ by the probability random forest as in Malley with $100$ trees and default size of leaf node. As for $q_0(x)$, we take $$ \widehat{q}(x) = \widehat{\mathbb{E}} \bigg[ \log \dfrac{\widehat{G}(d,x)}{1-\widehat{G}(d,x)} \bigg| X=x \bigg], $$ where the outer expectation function is estimated by regular random forest of Breiman. In addition, we assume that the propensity score $p_0(x)$ in equation (ref) is a sufficiently smooth function of $x$ and estimate it by simple logistic regression. The weighting function $\widehat{V}(d,x)$ is estimated as in (ref). In Definition (ref), we set $\lambda$ to be $$ \lambda=\lambda_{\text{ortho}} = \dfrac{1.1}{2\sqrt{n}} \Phi^{-1}(1-0.05/n ) = 0.031, $$ standardizing each covariate after interacting it with treatment. The Lasso estimate $\widehat{\theta}$ and its post-penalized analog post-Lasso-logistic $\widehat{\theta}_{\text{PL}}$ are reported in Table (ref), Columns (2)-(3). Finally, we report the unpenalized version of the orthogonal estimator defined as

align[align omitted — 142 chars of source]

Since the number of treatment interactions $p=20$ is less than the sample size, this estimator is well-defined. We report this estimate and the standard errors in Table (ref), Columns (4)-(5). Our method is straightforward to implement using the \url{glmnet} function in the \url{glmnet} $R$ package, setting the \url{offset} argument to $\widehat{q}(x)$, the \url{weights} argument to $\widehat{V}(d,x)$, and the \url{penalty.factor}=$(0,1,1,\dots, 1)$ to accommodate the penalty in (ref)

footnote{The package is available at \url{https://github.com/vsyrgkanis/plugin_regularized_estimation}}

.

Our empirical findings are as follows. The direct estimator (Column (1)) implies that Jobs First unambiguously increased the fraction of women on welfare for all women. In contrast, the orthogonal estimator (Column (2)) implies that women with long history of prior AFDC receipt (at least 5 months) were pushed out of the assistance. This pattern makes sense. The Jobs First (treatment group) faced a time limit of 21 months on welfare while the AFDC (control group) had no time limit. The orthogonal estimator picks up this pattern, while the direct one fails to do so. Finally, the unpenalized model in Columns (4) and (5) finds evidence in favor of treatment effect heterogeneity, but most interaction effects are not significant at $\alpha=0.05$. Therefore, orthogonal Lasso captures a more nuanced pattern of treatment effect that is left out by direct Lasso due to a substantially larger complexity of the nuisance parameter relative to the target one, failure of the sparsity assumption (ref), or both.

table[table omitted — 2,804 chars of source]

\setcounter{equation}{0} \setcounter{theorem}{0} \setcounter{lemma}{0} \setcounter{assumption}{0}

General Theory of Extremum Estimators

In this section, we introduce Regularized Extremum Estimator and establish its properties. Suppose there exists a population loss $Q(\theta, g)$ and its sample analog $\widehat{Q}(\theta, \widehat{g})$. Define the target parameter $\theta_0$ as the unique minimizer of $Q(\theta, g_0$) (i.e., (ref) holds), and the Regularized Extremum Estimator as in (ref). In contrast to (ref), $\widehat{Q}(\theta, \widehat{g})$ may not be a sample average. Finally, let the set $\mathbb{B}$ be as in (ref) and the event $\mathcal{V}$ be as in (ref).

assumption[Convexity] With probability one, for any $g$ in the set ${\mathcal G}_n$, the function $\theta \rightarrow \widehat{Q}(\theta,g)$ is a convex twice differentiable function of $\theta$ on an open convex set that contains the parameter space $\Theta$.
assumption[Restricted Strong Convexity on $\mathbb{B}$] There exists a constant $\bar{\gamma}>0$ so that the curvature of $Q(\theta, g_0)$ defined in (ref) is bounded from below by $\bar{\gamma}$: \begin{align*} \inf_{\theta \in \mathbb{B}, \nu=\theta - \theta_0 } \dfrac{\nu^T \nabla_{\theta \theta} Q(\theta, g_0) \nu} { \| \nu \|_2^2} \geqslant \bar{\gamma}. \end{align*}
assumption[Uniform Convergence on $\mathbb{B}$] There exists a sequence $\tau_n=o(1)$ so that $\nabla_{\theta \theta} \widehat{Q}(\theta, g_0) $ uniformly converges to $\nabla_{\theta \theta} Q(\theta,g_0)$: \begin{align} \sup_{\theta \in \mathbb{B}} \| \nabla_{\theta \theta} \widehat{Q}(\theta, g_0) - \nabla_{\theta \theta} Q(\theta, g_0) \|_{\infty} = O_{P} (\tau_n). \end{align}
assumption[Uniformly Lipschitz Hessian on $\mathbb{B}$] For any $g$ and $g'$ in ${\mathcal G}_n$, there exists a sequence $\xi_n=o(1)$ so that the following bound holds \begin{align} \sup_{\theta \in \mathbb{B}} \| \nabla_{\theta \theta} \widehat{Q}(\theta, g) - \nabla_{\theta \theta} \widehat{Q}(\theta, g') \|_{\infty} = O_{P}(g_n+ \xi_n). \end{align}
assumption[Convergence Rate of Empirical Gradient] For any $g\in {\mathcal G}_n$, there exists a sequence $\epsilon_n$ such that $ \|\nabla_{\theta} \widehat{Q}(\theta_0, g) - \nabla_{\theta} Q(\theta_0, g)\|_{\infty}=O_{P}(\epsilon_n)$.

We describe the influence of the estimation error $\widehat{g}-g_0$ on the gradient $\nabla_{\theta} Q(\theta_0, \cdot)$ using pathwise derivatives w.r.t the nuisance parameter $g_0$. The first-order derivative is

equation[equation omitted — 134 chars of source]

and the second-order derivative is

equation[equation omitted — 141 chars of source]
assumption[Bounded Gradient of Population Loss w.r.t. Nuisance] For any $g \in {\mathcal G}_n$, there exist constants $B_0$ and $B$ so that $$ \forall r \in [0,1): \left\| D_0 [g-g_0, \nabla_{\theta}Q(\theta_0, g_0 )] + D_{r}^2[g-g_0, \nabla_{\theta}Q(\theta_0, g_0 )] \right\|_{\infty} \leqslant B_0 \| g-g_0\| + B \|g - g_0\|^2. $$
theorem[Regularized Extremum Estimator] Suppose Assumptions (ref) and (ref)-(ref) hold with $k(\tau_{n} + g_{n} + \xi_{n})=o(1)$. For $\lambda = 2 C(\epsilon_{n} + B_0 g_n + B g_n^2)$, and for $C$ and $n$ large enough, \begin{align} \| \widehat{\theta} - \theta_0 \|_2 &\lesssim_P \sqrt{k} (\epsilon_n + B_0 g_n + g_n^2) & \| \widehat{\theta} - \theta_0 \|_1 &\lesssim_P k (\epsilon_n + B_0 g_n + g_n^2) . \end{align}

Theorem (ref) establishes the convergence rate for Regularized Extremum Estimator. Its building blocks are described in the following lemmas.

lemma[Convexity and Restricted Subspace, Negahban] If $\frac{\lambda}{2} \geqslant \|\nabla_\theta \widehat{Q}(\theta_0, \widehat{g})\|_{\infty}$ and $\theta \rightarrow Q(\theta, \widehat{g})$ is a convex function of $\theta$, then $\nu \in {\cal C}(T; 3)$, where $\nu=\widehat{\theta}-\theta_0$.
proof[Proof of Lemma (ref)] By definition of $\widehat{\theta}$, \begin{equation} \widehat{Q}(\widehat{\theta},\widehat{g}) - \widehat{Q}(\theta_0,\widehat{g}) \leqslant \lambda \left(\|\theta_0\|_1 - \|\widehat{\theta}\|_1\right). \end{equation} Additivity of $\ell_1$ norm and triangle inequality imply \begin{align*} \|\widehat{\theta}\|_1 &= \|\theta_0 + \nu_T\|_1 + \|\nu_{T^c}\|_1 \geqslant \|\theta_0\|_1 - \|\nu_T\|_1 + \|\nu_{T^c}\|_1 \\ \lambda \left(\|\theta_0\|_1 - \|\widehat{\theta}\|_1\right) &\leqslant \lambda \left(\|\nu_T\|_1 - \|\nu_{T^c}\|_1\right). \end{align*} Convexity of $\theta \rightarrow \widehat{Q}(\theta,\widehat{g})$, Cauchy-Schwarz inequality, and the choice of $\lambda$ imply: \begin{equation} \widehat{Q}(\widehat{\theta}, \widehat{g}) - \widehat{Q}(\theta_0, \widehat{g}) \geqslant \nabla_\theta \widehat{Q}(\theta_0, \widehat{g}) \cdot (\widehat{\theta}-\theta_0) \geqslant -\|\nabla_\theta \widehat{Q}(\theta_0, \widehat{g})\|_\infty \|\nu\|_1 \geqslant - \frac{\lambda}{2} \|\nu\|_1. \end{equation} Combining equations (ref) and (ref) gives: \begin{equation*} \lambda \left(\|\nu_T\|_1 - \|\nu_{T^c}\|_1\right) \geqslant \widehat{Q}(\widehat{\theta},\widehat{g}) - \widehat{Q}(\theta_0,\widehat{g}) \geqslant -\frac{\lambda}{2} \|\nu\|_1. \end{equation*} Dividing by $\lambda$ and re-arranging gives $3\|\nu_T\|_1 \geqslant \|\nu_{T^c}\|_1$.
lemma[Restricted Strong Convexity for Empirical Loss] Suppose Assumptions (ref) and (ref) hold. On the event $\mathcal{V}$, the following bound holds: \begin{align*} \inf_{ \nu \in {\cal C}(T;3)} \frac{ (\nabla_\theta \widehat{Q}(\theta_0 + \nu, \widehat{g}) - \nabla_\theta \widehat{Q}(\theta_0, \widehat{g})) \cdot \nu }{ \| \nu \|_2^2} \geqslant \bar{\gamma}/2. \end{align*}
proof[Proof of Lemma (ref)] Step 1. For some $\bar{r} \in [0,1)$ that may depend on $\widehat{g}$, mean value theorem implies \begin{align*} ( \nabla_\theta \widehat{Q}(\theta_0 + \nu, \widehat{g}) - \nabla_\theta \widehat{Q}(\theta_0, \widehat{g})) \cdot \nu &= \nu^T \cdot \nabla_{\theta \theta} \widehat{Q}(\theta_0 + \bar{r} \nu, \widehat{g}) \cdot \nu. \end{align*} Invoking Cauchy-Schwartz, definition of ${\cal C}(T; 3)$ in (ref), and assumption of the theorem gives: \begin{align*} &\sup_{\bar{r} \in [0,1), \nu \in {\cal C}(T;3)} \frac{| \nu^T \cdot ( \nabla_{\theta \theta} \widehat{Q}(\theta_0 + \bar{r} \nu, \widehat{g}) - \nabla_{\theta \theta} Q(\theta_0 + \bar{r} \nu, g_0) ) \cdot \nu |}{\|\nu\|^2_2} \\ &\leqslant \sup_{\nu \in {\mathcal C}(T;3)}\frac{\|\nu \|_1^2 \bar{\gamma}/(32k) }{\|\nu\|^2_2} \tag{Cauchy-Schwartz} \\ &\leqslant \sup_{\nu \in {\mathcal C}(T;3)} \frac{((1+3) \| \nu_T \|_1)^2 \bar{\gamma}/(32k) }{\|\nu\|^2_2} \tag{Definition of ${\cal C}(T; 3)$} \\ &\leqslant 16 k \bar{\gamma}/(32k) < \bar{\gamma}/2. \end{align*} Step 2. Triangle inequality implies \begin{align*} & \inf_{ \nu \in {\cal C}(T;3)} \frac{ \nu^T \cdot \nabla_{\theta \theta} \widehat{Q}(\theta_0 + \bar{r} \nu, \widehat{g}) \cdot \nu }{ \| \nu \|_2^2} \\ &\geqslant \inf_{ \nu \in {\cal C}(T;3)} \frac{ \nu^T \cdot \nabla_{\theta \theta} Q(\theta_0 + \bar{r} \nu, g_0) \cdot \nu }{ \| \nu \|_2^2} - \sup_{ \nu \in {\cal C}(T;3)} \frac{ | \nu^T \cdot ( \nabla_{\theta \theta} \widehat{Q}(\theta_0 + \bar{r} \nu, \widehat{g}) - \nabla_{\theta \theta} Q(\theta_0 + \bar{r} \nu, g_0) ) \cdot \nu |}{ \| \nu \|_2^2} \\ &\geqslant \bar{\gamma} - \bar{\gamma}/2 =\bar{\gamma}/2. \end{align*}
lemma[Oracle Inequality] Suppose Assumptions (ref) and (ref) hold. On the intersections of events (ref) and (ref), the following inequality holds: \begin{align} \|\widehat{\theta} - \theta_0\|_2 \leqslant& \frac{3\sqrt{k}}{\bar{\gamma} } \lambda & \|\widehat{\theta} - \theta_0\|_1 \leqslant& \frac{12 k}{\bar{\gamma} } \lambda. \end{align}
proof[Proof of Lemma (ref)] Let $\nu = \widehat{\theta} - \theta_0$. Invoking Lemmas (ref) and (ref) gives: \begin{align*} \lambda \left(\| \theta_0 \|_1 - \|\widehat{\theta}\|_1 \right) \geqslant & \nabla_\theta \widehat{Q}(\widehat{\theta}, \widehat{g}) \cdot (\widehat{\theta}-\theta_0) \tag{Optimality of $\widehat{\theta}$} \\ = & \nabla_\theta \widehat{Q}(\theta_0, \widehat{g}) \cdot (\widehat{\theta}-\theta_0) + ( \nabla_\theta \widehat{Q}(\widehat{\theta}, \widehat{g}) - \nabla_\theta \widehat{Q}(\theta_0, \widehat{g})) \cdot (\widehat{\theta}-\theta_0) \\ \geqslant & \nabla_\theta \widehat{Q}(\theta_0, \widehat{g}) \cdot (\widehat{\theta}-\theta_0) + \frac{\bar{\gamma}}{2} \|\nu\|_2^2 \tag{Lemmas (ref), (ref)}\\ \geqslant & -\|\nabla_{\theta} \widehat{Q}(\theta_0,\widehat{g})\|_{\infty} \cdot \|\nu\|_1 + \frac{\bar{\gamma}}{2} \|\nu\|_2^2 \tag{Cauchy-Schwarz}\\ \geqslant & -\frac{\lambda}{2} \cdot \|\nu\|_1 + \frac{\bar{\gamma}}{2} \|\nu\|_2^2 \tag{Assumption on $\lambda$} \end{align*} Rearranging the inequality gives \begin{align*} \frac{\bar{\gamma}}{2} \|\nu\|_2^2 \leqslant \frac{3\lambda}{2} \|\nu_T\|_1 - \frac{\lambda}{2} \|\nu_{T^c}\|_1 \leqslant \frac{3\lambda}{2} \|\nu_T\|_1 \leqslant \frac{3\lambda \sqrt{k}}{2} \|\nu\|_2. \end{align*} Dividing over by $\|\nu\|_2$ yields the theorem. By Lemma (ref), $\nu \in {\cal C}(T;3)$ and $\|\nu\|_1\leqslant 4 \sqrt{k} \|\nu\|_2$.
lemma[Influence on Oracle Gradient] Suppose Assumptions (ref), (ref) and (ref) hold. With probability $1-o(1)$, \begin{align} \|\nabla_\theta \widehat{Q}(\theta_0, \widehat{g})\|_{\infty} &= O_{P} (\epsilon_{n} + B_0 g_{n} + B g_{n}^2). \end{align}
proof[Proof of Lemma (ref)] In what follows, we condition on the event $\mathcal{E}_n$, which holds with probability $1-o(1)$. Triangle inequality implies: \begin{align} \|\nabla \widehat{Q}(\theta_0, \widehat{g})\|_{\infty} &\leqslant \| \nabla_{\theta} [Q(\theta_0, \widehat{g}) -Q(\theta_0, g_0) ] \|_{\infty} + \| \nabla_{\theta} [\widehat{Q}(\theta_0, \widehat{g}) -Q(\theta_0, \widehat{g})] \|_{\infty} \\ &= I + II \nonumber. \end{align} Conditional on the event $\mathcal{E}_n$, Assumption (ref) implies $$II=\| \nabla_{\theta} [\widehat{Q}(\theta_0, \widehat{g}) -Q(\theta_0, \widehat{g})] \|_{\infty} = O_{P} (\epsilon_n).$$ By Lemma 6.1 from chernozhukov2016double, this statement holds unconditionally. On the event $\mathcal{E}_n$, by Assumption (ref) \begin{align*} I= \| \nabla_{\theta} [Q(\theta_0, \widehat{g}) - Q(\theta_0, g_0) ] \|_{\infty} &\leqslant \sup_{g \in {\mathcal G}_n} \| \nabla_{\theta} [Q(\theta_0, g) - Q(\theta_0, g_0) ] \|_{\infty} \\ &\leqslant B_0 g_{n}+ B g_{n}^2. \end{align*} Thus, (ref) follows.
proof[Proof of Theorem (ref)] By Lemma (ref) and assumptions of the Theorem, \begin{align} \lambda/2 \geqslant \| \nabla_\theta \widehat{Q}(\theta_0, \widehat{g}) \|_{\infty} holds w.p. 1-o(1). \end{align} Decomposing the difference $\nabla_{\theta \theta} \widehat{Q}(\theta, \widehat{g} ) - \nabla_{\theta \theta} Q(\theta, g_0 )$ gives \begin{align*} \sup_{\theta \in \mathbb{B}} \| \nabla_{\theta \theta} \widehat{Q}(\theta, \widehat{g} ) - \nabla_{\theta \theta} Q(\theta, g_0 ) \|_{\infty} &\leqslant \sup_{\theta \in \mathbb{B}} \| \nabla_{\theta \theta} \widehat{Q}(\theta, \widehat{g} ) - \nabla_{\theta \theta} \widehat{Q}(\theta, g_0 ) \|_{\infty} \\ &+ \sup_{\theta \in \mathbb{B}} \| \nabla_{\theta \theta} \widehat{Q}(\theta, g_0) - \nabla_{\theta \theta} Q(\theta, g_0 ) \|_{\infty}\\ &= A+ B. \end{align*} By Assumptions (ref), (ref), (ref), on the event $\mathcal{E}_n$, \begin{align*} A &= O_P(g_n + \xi_n), \quad B = O_P(\tau_n). \end{align*} By assumption of the Theorem, $k (g_n + \xi_n + \tau_n) = o(1)$. Therefore, for $n$ large enough, \begin{align} \sup_{\theta \in \mathbb{B}} \| \nabla_{\theta \theta} \widehat{Q}(\theta, \widehat{g} ) - \nabla_{\theta \theta} Q(\theta, g_0 ) \|_{\infty} \leqslant \bar{\gamma}/(32k) holds w.p. 1-o(1). \end{align}

Proofs of Section 3

In this section, we prove Theorem (ref) as a special case of Theorem (ref). We also prove Theorem (ref). Define an event $\mathcal{E}_n := \{ \widehat{g}_{k} \in {\mathcal G}_n \quad \forall k \in [K] \}$, such that the nuisance parameter estimate $\widehat{g}_k$ belongs to the realization set ${\mathcal G}_n$ for each fold $k \in [K]$. By union bound, this event holds w.h.p. $${\mathrm{P}} (\mathcal{E}_n ) \geqslant 1- K \epsilon_n = 1-o(1).$$ For a given partition $k$ in $\{1,2, \dots, K\}$, define the partition-specific averages $${\mathbb{E}_{n,k}} f(W_i) := \dfrac{1}{n_k} \sum_{i \in J_k} f(W_i), \quad {\mathbb{G}_{n,k}} f(W_i) := \dfrac{1}{\sqrt{n_k}} \sum_{i \in J_k} [f(W_i) - \int f(w) dP (w)] .$$

lemma[Verification of Assumption (ref)] Assumption (ref) implies Assumption (ref).
proof[Proof of Lemma (ref)] If $m(w,t,\gamma)$ is non-decreasing in $t$, the loss $\ell(w,t,\gamma)$ is convex in $t$. Plugging the linear function $\theta \rightarrow t=\Lambda(z, g)'\theta$ into the loss sketch $\ell(w,t,\gamma)$ gives the convex sample loss $\widehat{Q} (\theta, g)$.
lemma[Verification of Assumption (ref)] Assumptions (ref) with $C_{\text{min}}$ and (ref) with $B_{\text{min}}$ imply Assumption (ref) with $\bar{\gamma} = B_{\text{min}} C_{\text{min}}$.
proof[Proof of Lemma (ref)] Observe that \begin{align*} &\inf_{\theta \in \mathbb{B} } \nu^\top \nabla_{\theta \theta} Q(\theta, g_0) \nu \\ =&\inf_{\theta \in \mathbb{B} } \mathbb{E} \bigg[ \nabla_t m(W, t, \gamma) \bigg|_{\gamma = g_0(Z), t=\Lambda(Z, \gamma)'\theta} | Z=z \bigg] \min \operatorname{eig} \, \Sigma \, \| \nu \|_2^2 \\ &\geqslant B_{min} C_{min} \| \nu \|_2^2. \end{align*}
lemma[Verification of Assumption (ref)] Lemma (ref) verifies Assumption (ref).
lemma[Verification of Assumption (ref)] Assumptions (ref) and (ref) with a constant $U$ implies Assumption (ref) with $\xi_n = 3 d \cdot U^3 \, \sqrt{\frac{\log(2p)}{n}}$ for $\mathbb{B}$ as in (ref) and $\| g \|_r$ as in (ref) for $r \in \{\infty,2\}$.
proof[Proof of Lemma (ref)] Step 1. Observe that for each $(v,j) \in \{1,2,\dots, p\}^2$, \begin{align} \nabla_{ \theta_v \theta_j} \ell(w, \Lambda(z, \gamma)'\theta, \gamma )=\nabla_{ t} m(w, \Lambda(z,\gamma)'\theta, \gamma) \cdot \Lambda_v(z, \gamma) \cdot \Lambda_j(z, \gamma) \end{align} has its gradient in $\gamma$ bounded by $L=3 \cdot U^3 $ in the absolute norm. Verification of Assumption (ref) for $\| g \| = \| g \|_{\infty}$. Assumptions (ref) implies Assumption (ref) with $L=U^3$. \begin{align*} &|\nabla_{\theta_v\theta_j} \widehat{Q}(\theta, g) - \nabla_{\theta_v\theta_j} \widehat{Q}(\theta, g')| \leqslant 3 \cdot U^3 \sup_{z\in {\mathcal Z}} \|g(z)-g'(z)\|_1. \end{align*} Step 2. Verification of Assumption (ref) for $\| g \| = \| g \|_2$. For each $(v,j) \in \{1,2,\dots, p\}^2$, the function above has $U$-bounded derivative, its derivative is bounded by $L=3U^3$. By McDiarmid's inequality, for any fixed $g, g'\in {\mathcal G}_n$, with probability $1-o(1)$, \begin{align*} \left|[{\mathbb{E}_{n,k}}-\mathbb{E}][\|g(Z)-g'(Z)\|_1] \right| \leqslant 3 d \cdot U^3 \, \sqrt{\frac{\log(2p)}{n_k}} = O\left( 3 d \cdot U^3 \, \sqrt{\frac{\log(2p)}{n}} \right):= O(\xi_{n} ). \end{align*}
lemma[Verification of Assumption (ref)] Assumption (ref) with a constant $U$ implies Assumption (ref) with $\epsilon_n = U^2 \cdot \sqrt{\frac{\log(2p)}{n}}$.
proof[Proof of Lemma (ref)] Conditional on the event $\mathcal{E}_n$, the estimate $g=\widehat{g}$ belongs to ${\mathcal G}_n$. Observe that $$\nabla_{\theta_j} \ell(w,\Lambda(z, \gamma)' \theta_0,\gamma)|_{\gamma = g(z)} = m(w, \Lambda(z, \gamma)' \theta_0, \gamma) \cdot \Lambda_j(z, \gamma) |_{\gamma = g(z)}, \quad j=1,2,\dots p$$ is an a.s. bounded function of data vector $w$, bounded in absolute value by $U^2$. By McDiarmid's inequality, the sample average of $j$'th function is within $U^2 \cdot \sqrt{\frac{\log(2p/\delta)}{n_k}}$ of its mean with probability at least $1-\frac{\delta}{p}$. Taking a union bound over the $p$ coordinates, we get that each coordinate is within $U^2 \cdot \sqrt{\frac{\log(2p/\delta)}{n_k}}$ from its respective mean with probability at least $1-\delta$. Therefore, $\epsilon_n$ can be taken to be $\epsilon_n = U^2 \cdot \sqrt{\frac{\log(2p)}{n_k}} = O(U^2 \cdot \sqrt{\frac{\log(2p)}{n}})$.
lemma[Verification of Assumption (ref)] Assumption (ref) with a constant $U$ implies Assumption (ref) with $B_0 = U^2$ and $B = 4U^2$.
proof[Proof of Lemma (ref) ] By Assumption (ref), for any $w, \theta$ and $g \in {\mathcal G}_n$, define the matrix $A_j \in \mathrm{R}^{d \bigtimes d}$ as $$ A_j:= A_j(w, \theta, g) = \nabla_{\gamma\gamma} \bigg[ m (w, t, \gamma)|_{\gamma = g(z), t = \Lambda(z, \gamma)'\theta} \cdot \Lambda_j(z, \gamma) \bigg]. $$ By Cauchy-Schwarz inequality, \begin{align*} | (g(Z)-g_0(Z))' A_j (g(Z)-g_0(Z)) | &\leqslant \| A_j \|_{\infty} \| (g(Z)-g_0(Z)) \|_1^2. \end{align*} Taking expectations on each side and invoking $\sup_{ 1 \leqslant j \leqslant p }\| A_j \|_{\infty} \leqslant 4 \cdot U^2 \quad \text{ a.s. }$ gives \begin{align*} D_r^2[g-g_0, \nabla_{\theta_j} Q(\theta_0, g_0)] &\leqslant \mathbb{E} | (g(Z)-g_0(Z))' A_j (g(Z)-g_0(Z)) | &\leqslant 4 \cdot U^2 \cdot \mathbb{E}[\|g(Z)-g_0(Z)\|_1^2]\\ &\leqslant 4 \cdot U^2 \, \sup_{z\in {\mathcal Z}} \|g(z)-g_0(z)\|_{1}^2\\ &\leqslant 4 \cdot U^2 \, \|g-g_0\|_{\infty}^2, \end{align*} Likewise, $\| D_0[g-g_0, \nabla_{\theta} Q(\theta_0, g_0)] \|_{\infty} \leqslant U^2 \mathbb{E}[\|g(Z)-g_0(Z)\|_1^2 $.
proof[Proof of Theorem (ref)] Theorem (ref) follows from the statement of Theorem (ref) and Lemmas (ref)-(ref).
proof[Proof of Theorem (ref)] Step 1. The validity of the adjusted CMR can be seen from \begin{align*} \mathbb{E} \bigg[ m (W, \Lambda(Z, \gamma_1)'\theta_0, \gamma)|_{\gamma=g_0(Z)} \bigg| Z=z \bigg] &= \mathbb{E} \bigg[ m_{pre} (W, \Lambda(Z, \gamma_1)'\theta_0, \gamma_1)|_{\gamma_1=p_0(Z) } \bigg| Z=z \bigg] \\ &- h_0(z) I_0^{-1}(z) \mathbb{E} \bigg[ R(W, p_0(Z)) \bigg| Z=z \bigg] =0, \end{align*} which follows from (ref) and (ref). Step 2. We verify the orthogonality condition (ref) in two steps. First, we compute the derivative of (ref) with respect to $\gamma=(\gamma_1, \gamma_2, \gamma_3)$ \begin{align*} &\mathbb{E} \bigg[ \nabla_{\gamma} m (W, \Lambda(z, \gamma_1)'\theta_0, \gamma) \cdot \Lambda(z, \gamma_1)|_{\gamma = g_0(z) } \bigg| Z=z \bigg] \\ + &\mathbb{E} \bigg[ m (W, \Lambda(z, \gamma_1)'\theta_0, \gamma) \cdot \nabla_{\gamma} \Lambda(z, \gamma_1) |_{\gamma = g_0(z)} \bigg| Z=z \bigg] \\ &= (0,0,0)' + (0,0,0)' = i + ii, \end{align*} where $i$ is shown in Step 3 and $ii$ follows from Step 1. Step 3. The partial derivatives with respect to $\gamma=(\gamma_1, \gamma_2, \gamma_3)$ are mean zero conditionally on $z$. \begin{align*} &\mathbb{E} \bigg[ \nabla_{\gamma_1} m (W, \Lambda(z, \gamma_1)'\theta_0, \gamma)|_{\gamma = g_0(z) } \bigg| Z=z \bigg] \\ &= \mathbb{E} \bigg[ \nabla_{\gamma_1} m_{pre} (W, \Lambda(z, \gamma_1)'\theta_0, \gamma_1) \bigg| Z=z \bigg] - h_0(z) I_0(z)^{-1} I_0(z) =0. \end{align*} \begin{align*} \mathbb{E} \bigg[ \nabla_{\gamma_2} m (W, \Lambda(z, \gamma_1)'\theta_0, \gamma)|_{\gamma = g_0(z) } \bigg| Z=z \bigg] = -I_0(z)^{-1} \mathbb{E}[ R(W, p_0(Z)) | Z=z] =0 \end{align*} \begin{align*} \mathbb{E} \bigg[ \nabla_{\gamma_3} m (W, \Lambda(z, \gamma_1)'\theta_0, \gamma)|_{\gamma = g_0(z) } \bigg| Z=z \bigg] = h_0(z) I_0^{-2} (z) \mathbb{E}[ R(W, p_0(Z)) | Z=z] =0 \end{align*} Step 4. Verification of Assumptions (ref)-(ref) for the moment function $m(w, t, \gamma)$ in (ref). First, observe that $$ h_0(z) I_0^{-1}(z) R(W, p_0(Z)) $$ does not depend on the single index $t$. Therefore, the moment function $m(w, t, \gamma)$ is monotone in $t$ if and only if $m_{\text{pre}}(w, t, \gamma_1)$ is monotone in $t$. Likewise, $m(w, t, \gamma)$ is single index in $t$ if and only if $m_{\text{pre}}(w, t, \gamma_1)$ is single index in $t$. Finally, $$\nabla_t m(w, t, \gamma) = \nabla_t m_{\text{pre}}(w, t, \gamma_1)$$ for any value of $w, t, \gamma$. The function $(w, \gamma) \rightarrow \gamma_2 \gamma_3^{-1} R(w, \gamma_1)$ is a smooth function of $\gamma=(\gamma_1, \gamma_2, \gamma_3)$ by assumption of the theorem, implying Assumption (ref).

Proofs of Section 4

Let $\check{\theta}$ be a preliminary estimator of $\theta_0$ converging at rate $\theta_n$. Lemma (ref) establishes the shrinkage rates of the nuisance realization sets introduced below.

enumerate• Suppose $$ R_h(z, t,\gamma_1)= \mathbb{E}[\nabla_{\gamma_1} m_{\text{pre}}(W, t,\gamma_1)|Z=z]$$ is a known function of $z, t,\gamma_1$. Define the realization set ${\mathcal H}_n={\mathcal H}_{n,r}$ for $r \in \{2, \infty\}$ \begin{align*} {\mathcal H}_n= \bigg \{ R_h(z, \Lambda(z, p(z) )'\theta, p(z) ): \| p -p_0\|_r \leqslant p_{n}, \| \theta - \theta_0\|_1 \leqslant \theta_{n} \bigg \}. \end{align*} • Suppose $$ R_I(z,\gamma_1)= \mathbb{E} [\nabla_{\gamma_1} R(W,\gamma_1)|Z=z] $$ is a known function of $z,\gamma_1$. Define the realization set $I_n=I_{n,r}$ for $r \in \{2, \infty\}$ as \begin{align*} I_n=\bigg\{ R_I(w,p(z)): \| p -p_0\|_r \leqslant \pi_n \bigg \} \end{align*} • Suppose $m_{\text{pre}} (w,t,\gamma_1)$ is a known function of $w, t,\gamma_1$ and let $$q_0(z,t,\gamma_1) = \mathbb{E}[ m_{\text{pre}}(W,t,\gamma_1) | Z=z].$$ Define the realization set as \begin{align*} {\mathcal H}_n(q_0)= \bigg \{ q(z, t, p(z)): \sup_{t \in \mathbb{R}} \sup_{\gamma \in \Gamma} \| q(\cdot, t,\gamma) - q_0(\cdot, t,\gamma) \|_r \leqslant q_{n}, \| \theta - \theta_0\|_1 \leqslant \theta_{n}, \| p - p_0 \|_r \leqslant \pi_n\bigg \} \end{align*}
lemma[Plausibility of Assumption (ref) for the adjusted CMR (ref)] Suppose Assumptions (ref)-(ref) hold with $U \geqslant 1$. \begin{enumerate} • The plug-in estimator $\widehat{h}(z)$ \begin{align*} \widehat{h}(z) := R_h(z, \Lambda(z, \widehat{p}(z))'\check{\theta}, \widehat{p}(z)) \end{align*} of $h_0(z)$ belongs to $ {\mathcal H}_n$ w.p. $1-o(1)$. The set ${\mathcal H}_n$ shrinks at rate $h_{n,r} = U^2 (\theta_n + \pi_{n,r})$ for $r= \infty$ and $h_{n,2} = \sqrt{2}U^2 (\theta_n + \pi_{n,2})$. • The plug-in estimator $\widehat{I}(z)$ \begin{align*} \widehat{I}(z)=R_I(z,\widehat{p}(z)) \end{align*} of $I_0(z)$ belongs to $I_n$ w.p. $1-o(1)$. The set $I_n$ shrinks at rate $i_{n,r} = U \pi_{n,r}$ for $r \in \{ \infty, 2 \}$. • The estimator $\widehat{h}(z)$ of $h_0(z)$ \begin{align*} \widehat{h}(z) := \widehat{q}(z, \Lambda(z,\widehat{p}(z))'\check{\theta},\widehat{p}(z)) \end{align*} belongs to the realization set $ {\mathcal H}_n(q_0)={\mathcal H}_{n,r}(q_0)$ for $r \in \{2, \infty\}$ w.p. $1-o(1)$. The set $ {\mathcal H}_n(q_0)$ shrinks at rate $h_{n,r} = U^2 (\theta_n + \pi_{n,r})+ q_{n,r}$ for $r = \infty$ and $h_{n,2} =2 (U^2 (\theta_n + \pi_{n,2})+ q_{n,2})$. \end{enumerate}
proof[Proof of Lemma (ref)] Step 1. Proof of Lemma (ref) (1). By Assumption (ref), $m_{\text{pre}} (w,t, \gamma_1)$ and $\Lambda(z, \gamma_1)$ are smooth functions of $t$ and $\gamma_1$. Therefore, $R_h(z, t, \gamma_1)$ is a smooth function of $t$ and $\gamma_1$ whose partial derivatives are $U$-bounded uniformly over all arguments. By intermediate value theorem, \begin{align*} R_h(z, t, \gamma) - R_h(z, t_0, \gamma_0) &= \nabla_{t} R_h( z, \bar{t},\gamma) \Lambda(z, \gamma)'(\theta-\theta_0) + \nabla_{\gamma} R_h(z,\Lambda(z,\bar{\gamma})'\theta_0, \bar{\gamma}) \cdot (\gamma - \gamma_0) \\ &= I_1(z) + I_2(z). \end{align*} By Cauchy Schwartz and Assumption (ref), for any $\theta: \| \theta - \theta_0\|_1 \leqslant \theta_{n}$, \begin{align*} \sup_{z \in {\mathcal Z}} | I_1(z) | &\leqslant U \sup_{z \in {\mathcal Z}} \sup_{\gamma \in \Gamma} | \Lambda(z, \gamma)'(\theta-\theta_0) | \leqslant U^2 \| \theta - \theta_0 \|_1 \leqslant U^2 \theta_n . \end{align*} Plugging $\gamma = p(z)$ and $\gamma_0 = p_0(z)$ gives the bound in $\ell_{\infty}$ norm: \begin{align*} \sup_{z \in {\mathcal Z}} |I_1(z) + I_2(z) | \leqslant U^2 \theta_n + U \pi_n \leqslant U^2 (\theta_n + \pi_n). \end{align*} Therefore, ${\mathcal H}_n$ shrinks at rate $h_n = U^2 (\theta_n + \pi_n)$ for $\pi_{n} = \pi_{n, \infty}$. \\ In $\ell_2$ norm, the bound is \begin{align*} (\mathbb{E} (I_1(Z) + I_2(Z))^2)^{1/2} \leqslant \sqrt{2} ((\mathbb{E} I_1^2(Z))^{1/2}+ (\mathbb{E} I_2^2(Z))^{1/2}) \leqslant \sqrt{2} U^2 (\theta_n + \pi_n) \end{align*} for $\pi_{n} = \pi_{n, 2}$. Step 2. Proof of Lemma (ref) (2). By intermediate value theorem, $I_n$ in $\ell_r$ norm shrinks at rate $i_n = U \pi_n$ if ${\mathcal P}_n$ shrinks at rate $\pi_{n,r}$, for $r \in \{2, \infty\}$. Step 3. Proof of Lemma (ref) (3). Observe that \begin{align*} q(z, t, p(z)) - q_0(z, t_0, p_0(z)) &= (q(z, t, p(z)) - q_0(z, t, p(z))) + (q_0(z, t, p(z)) - q_0(z, t_0, p_0(z))) \\ &= Q_1(z) + Q_2(z). \end{align*} Consider the case $r=\infty$. The bound on $Q_1(z)$ in $\ell_{\infty}$ follows from \begin{align*} \sup_{z \in {\mathcal Z} } | Q_1(z) | \leqslant \sup_{t \in \mathbb{R}} \sup_{\gamma \in \Gamma} | q(z, t, \gamma) - q_0(z, t, \gamma) | \leqslant q_n. \end{align*} Invoking Step 1 with $R_h(z, t, \gamma) = q_0(z,t,\gamma)$ gives \begin{align*} \sup_{z \in {\mathcal Z} } | Q_2(z) | \leqslant U^2 (\theta_n + \pi_n). \end{align*} Consider the case $r=2$. The bound on $Q_1(z)$ in $\ell_2$ follows from \begin{align*} (\mathbb{E} Q_1^2(Z))^{1/2} \leqslant \sup_{t \in \mathbb{R}} \sup_{\gamma \in \Gamma} (\mathbb{E} (q(Z, t, \gamma) - q_0(Z, t, \gamma))^2)^{1/2} \leqslant q_n. \end{align*} Invoking Step 1 with $R_h(z, t, \gamma) = q_0(z,t,\gamma)$ gives \begin{align*} (\mathbb{E} Q_2^2(Z))^{1/2} \leqslant \sqrt{2} U^2 (\theta_n + \pi_n). \end{align*}
proof[Proof of Corollary (ref)] Step 1. Assumption (ref) is a restatement of Assumption (ref) (ref). Assumption (ref) follows from Assumption (ref) (ref). Indeed, for any $w \in \mathcal W$ and $\nu \in {\cal C}(T;3)$ the following bound holds: \begin{align*} |(d - p_0(x)) \cdot (1,x)'(\theta_0 + \nu)+q_0(x)| &\leqslant H_{TE}^2 \| \theta_0 + \nu \|_1 + H_{TE} \leqslant 2 H_{TE}^3. \end{align*} Therefore, \begin{align*} &\inf_{\theta \in \mathbb{B}} \nabla_t \mathbb{E}[m(W,\Lambda(Z,g_0(Z))'\theta, g_0(Z))|Z=z] = \inf_{\theta \in \mathbb{B}} \frac{G' ((d - p_0(x)) \cdot (1,x)'\theta+q_0(x))}{V_0(z)} \geqslant U_{TE} B_{\text{min}}. \end{align*} Assumption (ref) follows from Assumption (ref) (ref). \textbf{ Step 2. } Verification of Assumption (ref). Let $x_0 = 1$. For each $j \in \{0,1,2,\dots, p-1\}$ the function $\Lambda_j (z, \gamma_1) = (d - \gamma_1) \cdot x_j$ is bounded by $H_{\text{TE}} ^2$ for any $w \in \mathcal W$ and $\gamma_1 \in \Gamma_1$. The first derivative $\nabla_{\gamma_1} \Lambda_j (z, \gamma_1) = x_j$ is bounded by $H_{\text{TE}}$ a.s. Finally, the second derivative $\nabla_{\gamma_1 \gamma_1} \Lambda_j (z, \gamma_1) = 0$ is zero. Next, the function $t \rightarrow G(t) - y$ is $3$-times differentiable in $t$ with derivatives bounded by $U_{\text{TE}}$. The first (second) derivative of $\gamma_3^{-1}$ is bounded by $U_{\text{TE}}$($2 U_{\text{TE}}$), respectively. Since the moment function $m(w, t, \gamma)$ is a product of $G(t + \gamma_2)-y$ and $\gamma_3^{-1}$, all the quantities in Assumption (ref) are bounded by $U=C_{U_{\text{TE}}} U_{\text{TE}}^2$ for a sufficiently large absolute constant $C_{U_{\text{TE}}}$. \textbf{ Step 3. } Verification of Assumption (ref). For $r \in \{2, \infty\}$, let $\pi_n=\pi_{n,r}$ and $q_n=q_{n,r}$. By Theorem (ref), Step 1 of Definition (ref) converges in $\ell_1$-norm with $\theta_n = k C_{\text{TE}} \left(\sqrt{\frac{\log p}{n}} + \pi_n + q_n \right)$ for a sufficiently large $C_{\text{TE}}$. By Lemma (ref)(1), the estimate $\widehat{V}(d,x)$ in (ref) converges at rate $ C_{U_{\text{TE}}} U_{\text{TE}}^2 ( \pi_n + q_n + \theta_n)$ in $\ell_r$ norm. Then, Assumption (ref) holds with $g_n = C_g k \left( \pi_n + q_n+\sqrt{\frac{\log p}{n}} \right)$ for a sufficiently large absolute constant $C_g$. \textbf{ Step 4. } Verification of orthogonality condition (ref). We show that the loss function gradient (ref) is an orthogonal moment equation with respect to perturbations of $g_0(z)= \{ p_0(x), q_0(x), V_0(d,x)\}$. \begin{align*} &\dfrac{\partial}{\partial r} \nabla_{\theta} Q ( \theta_0, r (p - p_0) + p_0) = -\mathbb{E} (D-p_0(X)) (p (X) -p_0(X)) \cdot (1,X)' \\ &- \dfrac{1}{V_0(D,X)} \mathbb{E} (Y - G ( (D-p_0(X)) \cdot ((1,X)'\theta_0) + q_0(X)) ) (p (X) -p_0(X)) \cdot (1,X)' =0 \\ &\dfrac{\partial}{\partial r} \nabla_{\theta} Q ( \theta_0, r (q - q_0) + q_0) = -\mathbb{E} (D-p_0(X)) (q (X) -q_0(X)) \cdot (1,X)' =0 \\ &\dfrac{\partial}{\partial r} \nabla_{\theta} Q ( \theta_0, r (V - V_0) + V_0) = -\dfrac{1}{V^2_0(D,X)} \mathbb{E} (Y - G ( (D-p_0(X)) \cdot ((1,X)'\theta_0) + q_0(X)) )\cdot \\ & (D-p_0(X)) (V(D,X) - V_0(D,X))\cdot (1,X)' =0. \end{align*}
proof[Proof of Corollary (ref)] The loss function (ref) in Example (ref) is a special case of (ref). By Theorem (ref), the population gradient based on (ref) obeys the orthogonality condition (ref). Step 1. Assumption (ref) holds by Assumption (ref) (ref). Assumption (ref) holds by \begin{align*} \inf_{\theta \in \mathbb{B}} \nabla_t \mathbb{E}[m(W,\Lambda(Z,g_0(Z))'\theta, g_0(Z))|Z=z] =\inf_{t \in \mathcal{T}} \nabla_t \mathbb{E} [u (Y, t)|X=x] \geqslant B_{min} \end{align*} for any $x \in \mathcal{X}$. Assumption (ref) directly follows from Assumption (ref) (ref). Step 2. Verification of Assumption (ref). Observe that $\Lambda(x, \gamma) = x$ and $\| x \|_{\infty} \leqslant H_{\text{MD}}$. Next, the function $t \rightarrow v \cdot u(y, t)$ is $3$-times differentiable in $t$ with derivatives bounded by $U_{\text{MD}}$. The second derivative of $\gamma_1^{-1}$ on $\Gamma_1$ is bounded by $6\bar{p}^{-3}$. Since the moment function $m_{\text{pre}}(w, t, \gamma_1)$ is a product of $v \cdot u(y,t)$ and $\gamma_1^{-1}$, all the quantities in Assumption (ref) are bounded by $C_{\text{MD}} (U_{\text{MD}}^2 + 6\bar{p}^{-3})$ for a sufficiently large absolute constant $C_{\text{MD}}$. Step 3. Verification of Assumption (ref). For $r \in \{2, \infty\}$, let $\pi_n=\pi_{n,r}$ and $q_n=q_{n,r}$. By Theorem (ref), Step 1 of Definition (ref) converges in $\ell_1$-norm with $\theta_n = k C_{\text{MD}} \left(\sqrt{\frac{\log p}{n}} + \pi_n \right)$ for a sufficiently large $C_{\text{MD}}$. By Lemma (ref)(3), the estimate $\widehat{h}(x)$ in Remark (ref) converges at rate $C_{\text{MD}} ( \pi_n + \theta_n) + q_n$. Thus, Assumption (ref) holds with $$g_n = C_g \left(k \left( \pi_n +\sqrt{\frac{\log p}{n}} \right)+q_n \right)$$ for a sufficiently large absolute constant $C_g$.
proof[Proof of Corollary (ref)] The loss function (ref) in Example (ref) is a special case of (ref). By Theorem (ref), the population gradient based on (ref) obeys the orthogonality condition (ref). Step 1. Assumption (ref) holds by Assumption (ref) (ref). Assumption (ref) holds by \begin{align*} \inf_{\theta \in \mathcal{B}} \nabla_t \mathbb{E}[m(W,\Lambda(Z,g_0(Z))'\theta, g_0(Z))|Z=z] = \inf_{t \in \mathcal{T}} {\mathcal L} (t) \geqslant B_{min} \end{align*} for any $x \in \mathcal{X}$. Assumption (ref) follows from the monotonicity of logistic function. Step 2. Verification of Assumption (ref). Observe that $\Lambda(x, \gamma) = (x;\gamma_1)$, $\| x \|_{\infty} \leqslant H_{\text{games}}$, and $\gamma_1 \in [0,1]$. Next, the function $t \rightarrow -(y - {\mathcal L}(t))$ is $3$-times differentiable in $t$ with derivatives bounded by some constant $U_{\text{games}}$, which can be taken $U_{\text{games}} \geqslant H_{\text{games}}$. Therefore, all the quantities in Assumption (ref) are bounded by $C_{U_{\text{games}}} U_{\text{games}}$ for a sufficiently large absolute constant $C_{U_{\text{games}}}$. Step 3. Verification of Assumption (ref). For $r \in \{2, \infty\}$, let $\pi_n=\pi_{n,r}$. By Theorem (ref), Step 1 of Definition (ref) converges in $\ell_1$-norm with $\theta_n = k C_{\text{games}} \left(\sqrt{\frac{\log p}{n}} + \pi_n \right)$ for a sufficiently large $C_{\text{games}}$. By Lemma (ref)(1), the estimate $\widehat{h}(x)$ in Remark (ref) converges at rate $C_{\text{games}} ( \pi_n + \theta_n)$. Thus, Assumption (ref) holds with $$g_n = C_g \left(k \left( \pi_n +\sqrt{\frac{\log p}{n}} \right) \right)$$ for a sufficiently large absolute constant $C_g$.