EconBase
← Back to paper

Post-Selection Inference for Generalized Linear Models with Many Controls

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

71,272 characters · 12 sections · 33 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Post-Selection Inference for Generalized Linear Models with Many Controls

abstractThis paper considers generalized linear models in the presence of many controls. We lay out a general methodology to estimate an effect of interest based on the construction of an instrument that immunize against model selection mistakes and apply it to the case of logistic binary choice model. More specifically we propose new methods for estimating and constructing confidence regions for a regression parameter of primary interest $\alpha_0$, a parameter in front of the regressor of interest, such as the treatment variable or a policy variable. These methods allow to estimate $\alpha_0$ at the root-$n$ rate when the total number $p$ of other regressors, called controls, potentially exceed the sample size $n$ using sparsity assumptions. The sparsity assumption means that there is a subset of $s<n$ controls which suffices to accurately approximate the nuisance part of the regression function. Importantly, the estimators and these resulting confidence regions are valid uniformly over $s$-sparse models satisfying $s^2\log^2 p = o(n)$ and other technical conditions. These procedures do not rely on traditional consistent model selection arguments for their validity. In fact, they are robust with respect to moderate model selection mistakes in variable selection. Under suitable conditions, the estimators are semi-parametrically efficient in the sense of attaining the semi-parametric efficiency bounds for the class of models in this paper. \\ Key words: uniformly valid inference, instruments, double selection, Neymanization, optimality, sparsity, model selection

Introduction

The literature on high-dimensional generalized linear models has experienced rapid development vdGeer,Unified2012. As in the case of linear mean regression models, a striking result of this literature is to achieve consistency when the total number of covariates $p$ is potentially much larger than the sample size $n$. The main underlying assumption for achieving consistency is sparsity, namely that the number of relevant controls is at most $s$, which is much smaller than $n$. Much of the interest focuses on $\ell_1$-penalized estimators that achieve desirable theoretical and computational properties, at least when the log-likelihood functions are concave. The theoretical properties are analogous to those of the corresponding $\ell_1$-penalized least squares estimator for linear mean regression models called Lasso T1996,BickelRitovTsybakov2009. Results include prediction error consistency, consistency of the parameter estimates in $\ell_k$-norms, variable selection consistency, and minimax-optimal rates.

Several papers have focused on high-dimensional logistic binary choice models, trying to exploit their structure in detail. $\ell_1$-penalized logistic regressions models were studied in Bunea2008, Bach2010, and Kwemou2012. Group logistic regression were studied in GroupLogistic2008 and Kwemou2012 to exploit addition sparsity patterns. Ising models were considered in Ising2010 and connections with robust $1$-bit recovery were derived in PlanVershynin2012. These works also derive rates of convergence for the coefficients, prediction error consistency, and variable selection consistency under various conditions.

This paper attacks the problem of estimation and inference on a regression coefficient of interest in generalized linear models allowing for the total number of covariates $p$ to be much larger than the sample size $n$. Specifically, we construct $\sqrt{n}$-consistent estimators and confidence regions for a parameter of interest $\alpha_0$, which measures the impact of a regressor of interest -- typically a “policy variable"-- on the regression function. Importantly, we show that the estimator is $\sqrt{n}$-consistent and the confidence regions achieve the required asymptotic coverage uniformly over many data-generating processes. We discuss the model framework for generalized linear models and then provide estimators to the logistic regression case to illustrate the results.

It is important to note that our estimation and inferential results are valid without assuming the conventional “separation condition" -- namely, without assuming that all the non-zero coefficients are sufficiently separated from zero. Although the separation condition is commonly used and might be appealing in some technometric applications (e.g. signal processing), it is often unrealistic and not credible in econometric, biometric, and many other applications. Even if applicable, it might not lead to accurate approximations of the finite sample behavior of estimation and inference procedures. Our procedures are robust to violation of the separation condition, and, thus, are robust to moderate model selection mistakes which inevitably occur in many applications (mistakes are very likely to occur when some coefficients are at the range of $O(n^{-1/2})$ which is typically not distinguishable from zero).

Our work contributes to a growing literature that avoids imposing separation conditions. In the context of instrumental regression, BellChernHans:Gauss and BellChenChernHans:nonGauss provide uniformly valid estimation and inference methods for instrumental variable models, using either post-selection or $\ell_1$-regularization methods to estimate “optimal instruments." They provide a $\sqrt{n}$-consistent, semi-parametrically efficient estimator of the main low-dimensional structural parameter. In the context of the linear mean regression model, BelloniChernozhukovHansenVolume,BelloniChernozhukovHansen2011 proposed a “double selection" approach to constructing uniformly valid estimation and inference methods, and c.h.zhang:s.zhang used one-step corrections to $\ell_1$-regularized estimators. In either case a $\sqrt{n}$-consistent, semi-parametrically efficient estimator of the low-dimensional regression parameter of interest is provided. In the the case of linear quantile regression models, BelloniChernozhukovKato2013a,BelloniChernozhukovKato2013b provide uniformly $\sqrt{n}$-consistent estimators and uniformly valid inference methods for least absolute deviations and quantile regressions. In an independent and contemporaneous work, vandeGeerBuhlmannRitov2013 propose an approach to inference in generalized linear models, based upon the one-step correction of $\ell_1$-penalized estimator, where the pieces of the corrections are estimated via (approximate) Lasso inversion of the sample information matrix; they also provide theoretical analysis under high-level conditions. The approach taken in the present paper is an independent proposal, and relies instead on either optimal instrument strategy or the double selection strategy which is related to Neyman's approach to dealing with nuisance parameters.

The aforementioned works as well as the current approach deviate substantially from the traditional approach of performing inference based upon perfect model selection results. leeb:potscher:hodges,potscher:leeb:dpe,potscher,LeebPotscher2005 have shown that such traditional/naive inference approach is not robust to violations of the separation condition, which bounds the magnitude of the non-zero coefficients away from zero. The naive post selection estimators and inference based upon them break down in the sense of failing to achieve $\sqrt{n}$-consistency and asymptotic normality when the separation condition is violated. We shall confirm the failure of such naive post selection procedures in Monte-Carlo experiments. In sharp contrast our procedure, by construction, is robust to violation of such assumptions. We shall demonstrate this via theoretical results as well as via Monte-Carlo experiments. The theoretical results hold uniformly in the class of $s$-sparse models and can be shown also to hold over approximately sparse models, using arguments similar to those used for linear mean and quantile models in BelloniChernozhukovHansen2011 and BelloniChernozhukovKato2013a,BelloniChernozhukovKato2013b.

We construct our estimators and confidence regions via three steps. The first step use post-model selection methods to estimate the nuisance part of the regression -- the part of the regression function associated to controls (i.e., non-main regressors). The second step uses post-model selection to estimate an optimal instrument. The third step suitably combines these estimates to form estimating equations that are immunized against crude estimation of the nuisance functions. Solutions of these equations lead to our proposed estimators and confidence regions. The framework allows for different methods to be used on each step leading to different estimators for generalized linear models. For the case of logit link function, we propose one estimator based upon instrumental logistic regression with optimal instrument and another estimator based upon double selection logistic regression. We verify the uniform validity of these procedures and demonstrate their good properties in a wide variety of experiments. While both implementations perform well, the double selection procedure emerged as the clear winner in these experiments. Our results and proofs reveal that many different estimators can be used as ingredient in the three steps of the algorithm, as long as a required sparsity and rates for estimating nuisance functions are achieved. For example, the first and second steps can be based not only on post-selection estimators but also on $\ell_1$-regularized estimators, while the third step can be alternatively approximated by a one-step correction from an initial value. Therefore several implementations having the same asymptotic properties are possible. We narrowed down our formal theoretical analysis to the set of procedures that exhibited the best performance in Monte-Carlo experiments (for example, Lasso methods performed worse than post-Lasso methods for estimating the nuisance parts, and one-step corrections performed worse than the exact solution of the estimating equation). One of the main results is to establish $\sqrt{n}$-consistency and asymptotic normality of estimators for generalized linear modes under high-level conditions on nuisance parameters.

Our constructions of the final estimators and confidence regions mainly make use of the post-model selection estimators in estimating the nuisance part of the regression function as well as the optimal instrument. As mentioned earlier, we focus on using selection as a means of regularization (which is necessary when $p>n$), mainly because compared to other methods of regularization, such as $\ell_1$-penalized maximum likelihood, they performed best in a wide set of experiments. In order to develop sharp results for these estimators we must control sparsity effectively. We therefore provide sparsity bounds for $\ell_1$-penalized logistic maximum likelihood estimators, which is used for selection, and also derive the rates of convergence for the post-model selection logistic maximum likelihood estimator. These results are of independent interest. In the estimation of optimal instruments, which we use as an ingredient in building the optimal estimating equation to create immunization property, we rely on post-selection least squares estimator with data dependent weights. The presence of data-dependent weights creates several interesting technical challenges. Finally, to obtain the asymptotic approximations to the estimators of regression coefficients of interest we rely on empirical process methods, using self-normalized maximal inequalities and entropy calculations that rely on the sparsity of the models selected via data-driven procedures. These proofs are of independent interests in other types of generalized linear models.

We organize the remainder of the paper as follows. In Section (ref), we present the framework for generalized linear models and the proposed estimators specialized to the logistic link function case. In Section (ref) we provide the statements of our main results on the uniform validity of the estimators and confidence regions. We present primitive conditions for the logistic case and high-level conditions for generalized linear models. Section (ref) contains a Monte-Carlo experiment. We present the proofs of these results in Appendix (ref). In Appendix (ref) we collect results on Lasso and Post-Lasso with estimated weights (Appendix (ref)) as well as results on $\ell_1$-penalized Logistic regression and post model selection Logistic regression (Appendix (ref)). In Appendix (ref) we present auxiliary inequalities.

Notation

Denote by $(\Omega,{\mathrm{P}})$ the underlying probability space. The notation ${\mathbb{E}_n}[\cdot]$ denotes the average over index $1 \leqslant i \leqslant n$, i.e., it simply abbreviates the notation $n^{-1} \sum_{i=1}^n[\cdot]$. For example, ${\mathbb{E}_n}[x_{ij}^{2}] = n^{-1} \sum_{i=1}^{n}x_{ij}^{2}$. Moreover, we use the notation $\bar {\mathrm{E}}[\cdot]={\mathbb{E}_n}[{\mathrm{E}}[\cdot]]$. For example, $\bar {\mathrm{E}}[ v_{i}^{2} ] = n^{-1}\sum_{i=1}^n{\mathrm{E}}[ v_{i}^{2} ]$. For a function $f: {\mathbb{R}} \times {\mathbb{R}} \times {\mathbb{R}}^{p} \to {\mathbb{R}}$, we write $\mathbb{G}_n (f) = n^{-1/2} \sum_{i=1}^{n} (f(y_{i},d_{i},x_{i}) - {\mathrm{E}} [ f(y_{i},d_{i},x_{i}) ])$. We denote the ${l}_1$-norm as $\|\cdot\|_1$, ${l}_2$-norm as $\|\cdot\|$, $l_\infty$-norm as $\|\cdot\|_\infty$, and the “${l}_0$-norm" as $\|\cdot\|_0$ to denote the number of non-zero components of a vector. For a sequence $(t_{i})_{i=1}^{n}$, we denote $\| t_{i} \|_{2,n} = \sqrt{ {\mathbb{E}_n} [ t_{i}^{2} ]}$. For example, for a vector $\delta \in {\mathbb{R}}^{p}$, $\| x_{i}'\delta \|_{2,n} = \sqrt{ {\mathbb{E}_n} [(x_{i}'\delta)^{2}]}$ denotes the prediction norm of $\delta$. Given a vector $\delta \in {\mathbb{R}}^p$, and a set of indices $T \subseteq \{1,\ldots,p\}$, we denote by $\delta_T \in {\mathbb{R}}^p$ the vector such that $(\delta_{T})_{j} = \delta_j$ if $j\in T$ and $(\delta_{T})_{j}=0$ if $j \notin T$. The support of $\delta$ as ${\rm support}(\delta) = \{ j \in \{1,...,p\}: \delta_j \neq 0\} $. We use the notation $(a)_+ = \max\{a,0\}$, $a \vee b = \max\{ a, b\}$, and $a \wedge b = \min\{ a , b \}$. We also use the notation $a \lesssim b$ to denote $a \leqslant c b$ for some constant $c>0$ that does not depend on $n$; and $a\lesssim_P b$ to denote $a=O_P(b)$. We assume that the quantities such as $p$, $s$, $y_{i}, d_i, x_{i}, \beta_{0}, \theta_{0}, T$ and $T_{\theta_0}$ are all dependent on the sample size $n$, and allow for the case where $p=p_{n} \to \infty$ and $s=s_{n} \to \infty$ as $n \to \infty$. We omit the dependence of these quantities on $n$ for notational convenience.

Generic Setup and Method

Consider a generalized linear regression model, where the outcome of interest $y_i$ relates to a scalar main regressor $d_i$ (e.g. a treatment or a policy variable) and $p$-dimensional controls $x_i$ through a link function ${G}$, namely for $i=1,\ldots,n$

equation[equation omitted — 110 chars of source]

Here $\alpha_0$ is the main target parameter, and $x_i'\beta_0$ is the nuisance regression function. The vector $\beta_0$ is a high-dimensional parameter which is assumed to be sparse, namely $\|\beta_0\|_0 \leqslant s$. We require $s$ to be small relative to $n$ in the sense that will be specified below, in particular $$ \frac{s^2 \log^2 (p\vee n)}{n} \to 0 $$ is required. In many settings this condition allows for the estimation of the nuisance function at the rate of $o(n^{-1/4})$.

Let $\{(y_i,d_i,x_i) \ : \ i=1,\ldots,n\}$ be a random sample, independent across $i$, obeying the model ((ref)) with $\|\beta_0\|_0\leqslant s$. We aim to perform statistical inference on the coefficient $\alpha_0$ that is robust to moderate model selection mistakes as those are unavoidable if coefficients are near zero. Our proposed methods rely (implicitly or explicitly) on an instrument $z_{0i}=z_0(d_i,x_i)$ such that:

eqnarray[eqnarray omitted — 406 chars of source]

The first and second relations provide an estimating equation for $\alpha_0$. Relation ((ref)) is key in our analysis and states that the estimating equation ((ref)) is insensitive with respect to first order perturbations of the nuisance function $x_i'\beta_0$. We call this orthogonality condition “immunity." Such immunization ideas can be traced to Neyman's approach to dealing with nuisance parameters, as we discuss in Section (ref). (Note also that because of ((ref)), there is also immunity with respect to perturbations on $z_0$.)

Our methods proceed in three steps:

itemize• The first step computes an estimate for the nuisance function $x_i'\beta_0$. • The second step estimates the instrument $z_{0i}$. • The third step combines these estimates to estimate the parameter of interest $\alpha_0$.

Estimation of nuisance functions $x_i'\beta_0$ and the instrument $z_{0i}$ has an asymptotically negligible effect, due to the “immunization properties" of the estimating equations. Several different choices for these procedures and for instruments are possible. Next we provide detailed recommendations for their choices.

In general, we construct a valid, optimal instrument based on the following decomposition for the weighted main regressor:

equation[equation omitted — 135 chars of source]

where

equation[equation omitted — 136 chars of source]

where $G'(t) = \frac{\partial}{\partial t} {G}(t)$. The optimal instrument is given by

equation[equation omitted — 56 chars of source]

Here too we shall impose a sparsity condition in ((ref)), namely that $\|\theta_0\|_0\leqslant s$. The use of sparsity in the main equation and this auxiliary equation can be generalized to approximate sparsity, with all results in this paper extending to this case, see Remark (ref).

The weights $f_i = w_i/\sigma_i$'s are used to achieve the orthogonality condition ((ref)):

equation[equation omitted — 207 chars of source]

and this condition immunizes the estimation of the main parameter $\alpha_0$ against crude estimation of the nuisance function $x_i'\beta_0$, in particular via post-selection estimators. The selection steps make unavoidable moderate model selection mistakes, which translate into vanishing estimation error, which has an asymptotic negligible effect on the estimator based on the sample analog of the equation ((ref)). The orthogonality condition ((ref)) is therefore a critical ingredient in achieving asymptotic uniform validity of the coverage of confidence regions. Among all instruments that provide such immunization, the instrument given in ((ref)) minimizes the asymptotic variance of the asymptotically normal and $\sqrt{n}$-consistent estimator based on the estimating equations ((ref)). Other valid (but sub-optimal) choices of instruments are discussed in Remark (ref). We will establish results for generalized linear models under high-level conditions.

Logistic Case and Specific Estimators

Next we apply the above principle to the case of logistic regression and propose specific implementations of estimators. In this case the link function $G$ is given by the logistic link function $${G}(t) = \exp(t)/\{1+\exp(t)\},$$ and the following simplification occurs: $w_i$ in ((ref)) equals the conditional variance of the outcome $\sigma_i^2$, namely $$w_i = \sigma_i^2={G}( d_i\alpha_0 + x_i'\beta_0 )\{1-{G}(d_i\alpha_0+x_i'\beta_0)\}, \ \ \mbox{and} \ \ f_i = \sqrt{w_i},$$ so that the decomposition ((ref)) and the optimal instrument ((ref)) become

equation[equation omitted — 205 chars of source]

We describe two estimators in Tables (ref) and (ref). In these tables we denote the (negative) log-likelihood function associated with the logistic link function as

equation[equation omitted — 159 chars of source]

Table (ref) displays an estimator based on the optimal instrument. The estimation in Step 1 is based on post-selection logistic regression where the model is selected based on $\ell_1$-penalized logistic regression. Step 2 is based on a post-selection least squares with estimated weights constructed based on Step 1. Note that Step 2 is used to construct the optimal instrument. Step 3 uses an instrumental logistic regression, with estimates of nuisance functions (control function $x_i'\beta_0$ and the instrument $z_{0i}$) obtained in Steps 1 and 2. The use of post-selection estimators in the first two steps instead of penalized estimators was motivated by a better finite sample performance in our experiments. We also provide two confidence regions for $\alpha_0$ in Table (ref). The direct confidence region ${\mathcal{CR}}_D$ is based on the asymptotic normality of the estimator $\check \alpha$. The indirect confidence region ${\mathcal{CR}}_I$ is based on the asymptotic $\chi^2(1)$ law of the statistic $nL_n(\alpha_0)$.

table[table omitted — 3,659 chars of source]

Table (ref) describes a second estimator, which builds upon the idea of the double selection method proposed in BelloniChernozhukovHansen2011 for partial linear mean regression models. The method replaces Step 3 in Table (ref) with a (weighted) logistic regression of the outcome on the main regressor as well as the union of controls selected in two selection steps -- Steps 1 and 2. (Note that the algorithm is stated for any generalized linear model in which case Step 3 is a weighted regression where the weights are given by $\widehat f_i/\widehat \sigma_i$ which equals to $1$ in the case of a logistic link function.) This approach creates an optimal instrument implicitly. In fact, inspection of the proof shows that the double selection estimator can be seen as an infinitely iterated version of the previous method. We refer to Section (ref) for further connections and discussions.

table[table omitted — 3,395 chars of source]
remark[Other Valid Instruments] An instrument $z_0$ is valid if it has the orthogonality property $$\left. \frac{\partial}{\partial\beta}{\mathrm{E}}[\{y_i-{G}(d_i\alpha_0+x_i'\beta)\}z_{0i}]\right|_{\beta=\beta_0} = {\mathrm{E}}[ w_iz_{0i} x_i]=0$$ and is non-trivial, namely $\bar {\mathrm{E}}[w_id_iz_{0i}] \neq 0$. A valid, non-trivial instrument is optimal if it minimizes the asymptotic variance of the final estimator of $\alpha_0$. The algorithm stated in Table (ref) uses the optimal instrument $z_{0i} := v_i/\sigma_i$. Estimation of this instrument requires that in Step 2 a Lasso method is applied in the weighted equation ((ref)). Since the weights $w_i$'s in the resulting weighted Lasso problem are estimated, with estimation errors depending upon the response variable $d_i$, estimation of the optimal instrument creates interesting technical challenges in the analysis of Lasso or Post-Lasso that are dealt with in the Appendix. Thus estimation of the optimal instruments poses an interesting problem in its own right. There are other valid instruments that we can rely on, but these instruments are not generally optimal. For example, a valid, yet sub-optimal choice of the instrument is $z_{0i} := (d_i-{\mathrm{E}}[d_i\mid x_i])/w_i$. The estimation of this instrument is technically simpler, and follows easily from available results. Indeed, assuming ${\mathrm{E}}[d_i\mid x_i] = x_i'\theta_d$, with $\theta_d$ sparse or approximately sparse, we can estimate $z_{0i}$ by estimating $\theta_d$ via standard Lasso of $d_i$ on $x_i$, and estimating $w_i$ using the estimates of the $\ell_1$-penalized logistic regression as in Step 1. Note that since no estimated weights are used in Lasso estimation of $\theta_d$, standard results on the Lasso estimator deliver the required properties. {\tiny \ensuremath{\blacksquare} }
remark[Alternative Implementations via Approximate Instrumental Regression] The instrumental logistic regression can be approximately implemented by a 1-Step estimator from the $\ell_1$-penalized logistic estimator $\widehat\alpha$ of the form $ \check \alpha = \widehat \alpha + ({\mathbb{E}_n}[\widehat w_i d_i \widehat z_i])^{-1} {\mathbb{E}_n}[\{y_i- {G}(d_{i} \widehat\alpha +x_i'\widehat \beta)\}\widehat z_i]$. However, we prefer the exact implementations, since the they perform better in an extensive set of Monte-Carlo experiments. {\tiny \ensuremath{\blacksquare} }
remark[Data Driven Choice of Penalty Parameters] The penalty parameters $\lambda_1$ and $\lambda_2$ as defined in the algorithms above are motivated by self-normalized moderate deviation theory. Other data-driven choices are possible but their theoretical validity is outside the scope of this paper. For example, cross validation typically underpenalize to reduce bias to obtain better estimates but tends to select a substantial larger number of variables. This suggests that cross validation to be more suitable for the algorithm based on optimal instrument than for the algorithm based on double selection. Another approach suggested in Section 4.2 of chernozhukov2013gaussian relies on new Gaussian approximation results and can be implemented via a multiplier bootstrap procedure.{\tiny \ensuremath{\blacksquare} }

Main Theoretical Results

Logistic Regression under Primitive Assumptions

In this section, we list and discuss primitive conditions that allow us to derive our results in the case of logistic regression. These conditions ensure good properties of $\ell_1$-penalized methods and the associated post-selection estimators. Fix some sequences of constants, $\delta_n\to 0$, and $\Delta_n \to 0$, and constants $0 < c < C <\infty$.

{\bf Condition L.} { (i) Let $\{(y_i,d_i,x_i):i=,\ldots,n\}$ be independent random vectors that obey the model given by ((ref)) and ((ref)) with $G$ being the logistic function. There exists $s=s_n$ such that $\|\beta_0\|_0+\|\theta_0\|_0\leqslant s$, $\|\beta_0\|+\|\theta_0\| \leqslant C$. (ii) The following moment conditions hold $\bar {\mathrm{E}}[\{(d,x')\xi\}^4]\leqslant C \|\xi\|^4$, $\bar {\mathrm{E}}[w_i\{(d,x')\xi\}^2]\geqslant c\|\xi\|^2$. We have that $\min_{j\leqslant p} \bar {\mathrm{E}}[w_ix_{ij}^2v_i^2] \geqslant c>0$ and $\max_{j\leqslant p} \bar {\mathrm{E}}[|\sqrt{w_i}x_{ij}v_i|^3]^{1/3}\log^{1/2}(p\vee n) \leqslant \delta_n n^{1/6}$. Furthermore, the conditional variance $\sigma_i^2$ satisfy $\min_{i\leqslant n} \sigma_i^2 \geqslant c>0$ with probability $1-\Delta_n$. (iii) For $K_q = {\mathrm{E}}[\max_{i\leqslant n} \|(d_i,z_{0i},x')\|_\infty^q]^{1/q}$, we have $K_1^2s^2\log^2(p\vee n) \leqslant \delta_n n$ and $K_4^4 s \log (p\vee n) \log^{3}n\leqslant \delta_n n$.}

Condition L(i) assumes independence across $i$ and the model described in Section (ref) and sparsity conditions which makes estimation possible even if $p>n$. Condition L(ii) assumes the conditional variance is bounded away from zero and imposes mild moment conditions. Condition L(iii) imposes growth requirements on the triple $(s,p,n)$ as $n$ grows. An important consequence of Condition L is to imply that submatrices of the design matrix are well behaved even though the design matrix cannot have rank $p$ if $p>n$; see RudelsonVershynin2008,RudelsonZhou2011 for detailed discussion. This ensures that $\ell_1$-penalized estimators are well behaved with suitable choices of penalty parameters under the stated sparsity assumptions.

Next we state the main inferential results of the paper. It concerns the (uniform) validity of the different confidence regions for the coefficient $\alpha_0$ based on the optimal instrument and double selection algorithms.

theorem[Robust Estimation and Inference based on the Optimal IV Estimator] Consider any triangular array of data $(y_i, d_i, x_i)_{i=1}^n$ that obeys Condition L for all $n\geqslant 1$. Then, the estimator $\check \alpha$ based on the optimal instrument, as defined in Table (ref), obeys as $n \to \infty$ $$ \Sigma_n^{-1} \sqrt{n} (\check \alpha - \alpha_0) = Z_n + o_{\mathrm{P}}(1), \ \ Z_n \rightsquigarrow N(0,1), $$ where $$ Z_n := \frac{\Sigma_n}{\sqrt{n}} \sum_{i=1}^n \{y_i - G(d_i \alpha_0 + x_i'\beta_0)\} z_{0i} \text{ and } \ \Sigma^2_n:= \bar {\mathrm{E}}[v_i^2]^{-1}. $$ Moreover, $$ nL_n(\alpha_0) = Z_n^2 + o_{\mathrm{P}}(1), \ \ Z_n^2 \rightsquigarrow \chi^2(1). $$ Finally, $\Sigma_n^2$ can be replaced by either $\widehat \Sigma^2_{1n} = \{{\mathbb{E}_n}[\widehat w_i d_i \widehat z_i]\}^{-1}{\mathbb{E}_n}[\{y_i-G(d_i\check\alpha+x_i'\widetilde\beta)\}^2\widehat z_i^2]\{{\mathbb{E}_n}[\widehat w_i d_i \widehat z_i]\}^{-1}$ or by $\widehat \Sigma_{2n}^2 = {{\mathbb{E}_n}[\widehat v_i^2]}^{-1}$ without affecting the result, i.e. $\widehat \Sigma_{1n}^2/\Sigma_n^2 = 1 + o_{{\mathrm{P}}}(1)$ and $\widehat \Sigma_{2n}^2/\Sigma_n^2 = 1 + o_{{\mathrm{P}}}(1)$.

Theorem (ref) establishes that the IV estimator $\check \alpha$ is $\sqrt{n}$-consistent and asymptotically normal. Under suitable conditions the large-sample variance coinciding with the semi-parametric efficiency bound for the partially linear logistic regression model (see Section (ref) for an additional discussion). The studentized estimator converges to the standard normal law, and the criterion function that this estimator minimized, when evaluated at the true value, converges to the standard chi-squares law with one degree of freedom. These results justify and imply the validity of the confidence regions ${\mathcal{CR}}_D$ and ${\mathcal{CR}}_I$ for $\alpha_0$ proposed in Table (ref). We note that these results are achieved despite possible model selection mistakes in Steps 1 and 2.

The following result derives similar properties for the double selection estimator described in Table (ref).

theorem[Robust Estimation and Inference based on Double Selection] Consider any triangular array of data $(y_i, d_i, x_i)_{i=1}^n$ that obeys Condition L for all $n\geqslant 1$. Then, the double selection estimator $\check \alpha$ as defined in Table (ref) obeys as $n \to \infty$ $$ \Sigma_n^{-1} \sqrt{n} (\check \alpha - \alpha_0) = Z_n + o_{\mathrm{P}}(1), \ \ Z_n \rightsquigarrow N(0,1), $$ where $$ Z_n := \frac{\Sigma_n}{\sqrt{n}} \sum_{i=1}^n (y_i - G(d_i \alpha_0 + x_i'\beta_0)) z_{0i} \text{ and } \ \Sigma^2_n:= \bar {\mathrm{E}}[v_i^2]^{-1}. $$ Moreover, $\Sigma_n^2$ can be replaced by $\widehat \Sigma^2_{1n} = \{{\mathbb{E}_n}[\check w_i d_i \widehat z_i]\}^{-1}{\mathbb{E}_n}[\{y_i-G(d_i\check\alpha+x_i'\check\beta)\}^2\widehat z_i^2]\{{\mathbb{E}_n}[\check w_i d_i \widehat z_i]\}^{-1}$ or $\widehat \Sigma^2_{2n}=\{{\mathbb{E}_n}[\check w_i (d_i,\check x_{i}')'(d_i,\check x_{i}')]\}^{-1}_{11}$ without affecting the result, i.e. $\widehat\Sigma_{1n}^2/\Sigma_n^2=1+o_{{\mathrm{P}}}(1)$ and $\widehat\Sigma_{2n}^2/\Sigma_n^2=1+o_{{\mathrm{P}}}(1)$, where $\check w_i= G(d_i\check\alpha+x_i'\check\beta)\{1-G(d_i\check\alpha+x_i'\check\beta)\}$ and $\check x_i = x_{i,{\rm support}(\check\beta)}$.

Theorems (ref) and (ref) allow for the data-generating processes to change with $n$, in particular allowing sequences of regression models, with coefficients never perfectly distinguishable from zero, i.e. models where perfect model selection is not possible. In turn, the results achieved in Theorems (ref) and (ref) are uniformly valid over a large class of sparse models. In what follows, we formalize these assertions as corollaries.

Let $\mathcal{Q}_{n}$ denote a collection of distributions $Q_n$ for the data $\{ (y_{i},d_{i}, x_i')' \}_{i=1}^{n}$ such that Condition L hold for the given $n$. This is the collection of all sparse models where the stated above sparsity conditions, moment conditions, and growth conditions hold. This collection expressly permits models to have near zero coefficients, and thus does not impose the separation conditions (which we believe to be unreasonable in many applications). For $Q_n \in \mathcal{Q}_n$, let the notation ${\mathrm{P}}_{Q_{n}}$ mean that under ${\mathrm{P}}_{Q_{n}}$, $\{ (y_{i},d_{i}, x_i')' \}_{i=1}^{n}$ is distributed according to $Q_{n}$.

corollary[Uniform $\sqrt{n}$-Rate of Consistency and Uniform Normality] Let $\mathcal{Q}_{n}$ be the collection of all distributions of $\{ (y_{i},d_{i},x_i')' \}_{i=1}^{n}$ for which Condition L is satisfied for the given $n \geqslant 1$. Then the estimator $\check \alpha$, based either on optimal instrument or double selection, is $\sqrt{n}$-consistent and asymptotically normal uniformly over $\mathcal{Q}_n$, namely $$ \lim_{n \to \infty} \sup_{Q_n \in \mathcal{Q}_n}\sup_{t\in {\mathbb{R}}} |{\mathrm{P}}_{Q_n} ( \Sigma_n^{-1} \sqrt{n} (\check \alpha - \alpha_0) \leqslant t) - {\mathrm{P}}( N(0,1) \leqslant t)| =0 $$ Moreover, the result continues to hold if $\Sigma_n^2$ is replaced by any of the estimators $\widehat \Sigma_n^2$ specified in the statements of the preceding theorems.
corollary[Uniformly Valid Confidence Regions] Let $\mathcal{Q}_{n}$ be the collection of all distributions of $\{ (y_{i},d_{i},x_i')' \}_{i=1}^{n}$ for which Condition L is satisfied for the given $n \geqslant 1$. Then confidence regions ${\mathcal{CR}} \in \{ {\mathcal{CR}}_D, {\mathcal{CR}}_I, {\mathcal{CR}}_{DS}\}$ are asymptotically valid uniformly in $n$, namely $$ \lim_{n \to \infty} \sup_{\xi \in (0,1)} \sup_{Q_n \in \mathcal{Q}_n} |{\mathrm{P}}_{Q_n} ( \alpha_0 \in {\mathcal{CR}} ) - (1- \xi)| =0 $$

All of the results are new under $s \to \infty$ and $p \to \infty$ asymptotics, and they are new even under the fixed $s$ and $p$ asymptotics. These results motivates interesting questions on the construction of confidence regions for many parameters of interest that are simultaneously valid. We refer to BelloniChernozhukovKato2013a, BCFH2013program, and BCCW2015 where simultaneous confidence regions are proposed in a variety of settings.

remark[Generalization to Approximately Sparse Models] The results can also be shown to hold, with identical conclusions, in the class of approximately sparse models, following the analysis of the partially linear mean regression model in BelloniChernozhukovHansenVolume,BelloniChernozhukovHansen2011. For example, if the model satisfies \begin{eqnarray} & {\mathrm{E}}[y_i \mid d_i,x_i ] = G(\alpha_0 d_i + x_i'\beta_0 +r_{yi}), & \\ & f_id_i = f_i x_i'\theta_0 +r_{di} + v_i, \ \ & {\mathrm{E}}[f_i v_ix_i] =0, \end{eqnarray} where $\|\beta_0\|_0\leqslant s$, $\|\theta_0\|_0\leqslant s$, and the approximation errors $r_{yi}$ and $r_{di}$ are such that \begin{equation} \sqrt{\bar {\mathrm{E}}[r^2_{yi}]} \leqslant C \sqrt{ s/n } , \ \ \ \sqrt{\bar {\mathrm{E}}[r^2_{di}]} \leqslant C \sqrt{ s/n }, \ \ and \ \ |\bar {\mathrm{E}}[f_i v_ir_{yi}]| \leqslant \delta_n n^{-1/2} \end{equation} We can show that the results in Theorems (ref) and (ref) and Corollaries (ref) and (ref) continue to hold for this approximately sparse model. This means that the results are robust with respect to moderate violations of the sparsity assumption. {\tiny \ensuremath{\blacksquare} }

Generalized Linear Models under High-Level Assumptions

In this section we establish $\sqrt{n}$-consistency and asymptotic normality for an estimator $\check \alpha$ of $\alpha_0$ associated with a generalized linear model based on high-level conditions. These high-level conditions cover a variety of different estimators including the estimators described in Tables (ref) and (ref). In what follows note that the estimated instrument $\widehat z_i=\widehat z_i(d_i,x_i)$ and the expectations below are evaluated at the given estimates.

{\bf Condition IR.} (i) The data $\{y_i,d_i,x_i\}$ independent across $i=1,\ldots,n$, satisfies ((ref)), $\sigma_i^2 = \text{Var}(y_i\mid d_i,x_i)$, $w_i = G'(d_i\alpha_0+x_i'\beta_0)$, and the link function $G$ such that $\sup_{t\in {\mathbb{R}}}|G(t)|\leqslant C$, $\sup_{t\in {\mathbb{R}}}|G'(t)|\leqslant C$ and $\sup_{t\in {\mathbb{R}}}|G''(t)|\leqslant C$. (ii) The following moment conditions hold ${\mathrm{E}}[ w_iz_{0i}x_i]=0$, $ |{\mathrm{E}}[w_id_iz_{0i}]|\geqslant c > 0$, $\bar {\mathrm{E}}[\sigma_i^2z_{0i}^2]\geqslant c > 0$, $\bar {\mathrm{E}}[ z_{0i}^2d_i^2 ] \leqslant C$, $\bar {\mathrm{E}}[\sigma_i^3z_{0i}^3]\leqslant C$, and $\bar {\mathrm{E}}[(x_i'\xi)^4] \leqslant C$ for all $\|\xi\|=1$. (iii) For some sequences $\delta_n\to 0$ and $\Delta_n\to 0$ with probability at least $1-\Delta_n$, the estimates $(\check \alpha,\widehat \beta,\widehat z)$ satisfy

equation[equation omitted — 391 chars of source]
equation[equation omitted — 287 chars of source]
equation[equation omitted — 225 chars of source]

(iv) with probability $1-\Delta_n$ we have $|\widehat w_i|\leqslant C$, $\|\widehat w_i-w_i\|_{2,n}\leqslant \delta_n$, $\|d_i(\widehat w_i-w_i)\|_{2,n}\leqslant \delta_n$, $\|d_i\widehat z_i\|_{2,n}\leqslant C$, $\|\widehat z_i - z_{0i}\|_{2,n}\leqslant \delta_n$, and $\|z_{0i}x_i'(\widehat\beta-\beta_0)\|_{2,n}\leqslant \delta_n$.

This set of high-level conditions allow us to cover several generalized models of interest. In particular Condition L and the choices of post-selection methods described in the previous section imply Condition IR. Next we formally state our main result for generalized linear models.

theoremUnder Condition IR(i,ii,iii) we have $$ \{\bar {\mathrm{E}}[\sigma_i^2z_{0i}^2]\}^{-1/2}\bar {\mathrm{E}}[w_id_iz_{0i}]\sqrt{n}(\check\alpha-\alpha_0) = \frac{ \{\bar {\mathrm{E}}[\sigma_i^2z_{0i}^2]\}^{-1/2}}{\sqrt{n}}\sum_{i=1}^n\{y_i-G(d_i\alpha_0+x_i'\beta_0)\}z_{0i} + o_{\mathrm{P}}(1) $$ $$\mbox{and} \ \ \ \{\bar {\mathrm{E}}[w_id_iz_{0i}]^{-1}\bar {\mathrm{E}}[\sigma_i^2z_{0i}^2]\bar {\mathrm{E}}[w_id_iz_{0i}]^{-1}\}^{-1/2}\sqrt{n}(\check \alpha-\alpha_0)\rightsquigarrow N(0,1).$$ Moreover, if Condition IR(iv) also holds, we have $$nL_n(\alpha_0) \rightsquigarrow \chi^2(1)$$ and the variance estimator is consistent, namely $${\mathbb{E}_n}[\widehat w_id_i\widehat z_i]^{-1}{\mathbb{E}_n}[\{y_i-{G}(d_i\check\alpha+x_i'\widehat\beta)\}^2\widehat z_i^2]{\mathbb{E}_n}[\widehat w_i d_i \widehat z_i]^{-1}=\bar {\mathrm{E}}[w_id_iz_{0i}]^{-1}\bar {\mathrm{E}}[\sigma_i^2z_{0i}^2]\bar {\mathrm{E}}[w_id_iz_{0i}]^{-1}+o_{\mathrm{P}}(1).$$

It is important to note that Theorem (ref) is applicable to various estimation methods and we believe it will be of interest even in settings not based on sparsity assumptions.

Monte Carlo

Here we provide a simulation study of the finite sample properties of the proposed estimators and confidence intervals. We compare their performance with the naive post selection estimator, which is defined by applying the logistic regression performed on the model selected by the $\ell_1$-penalized logistic regression.

Our simulations are based on the model: $$ {\mathrm{E}}[y\mid d, x] = {G}( d\alpha_0 + x'\{c_y\nu_y\} ), \ \ \ \ d = x'\{c_d\nu_d\} + \tilde v, $$ where the coefficient vectors $\nu_y$ and $\nu_d$ are set to $$

array[array omitted — 154 chars of source]

$$ $x = (1,z')'$ consists of an intercept and covariates $z \sim N(0,\Theta)$, and the error $\tilde v$ is i.i.d. as $N(0,1)$. The dimension $p$ of the covariates $x$ is $250$, and the sample size $n$ is $200$. In this setting the coefficients feature a declining pattern, with the smallest non-zero coefficients being hard to differentiate from zero for the given sample size. Therefore, we expect that the $\ell_1$-based model selectors will be making selection mistakes on variables with the smaller coefficients. (Additional simulations are provided in the Supplementary Material where we also consider an approximately sparse model for which all 250 coefficients are non-zero. Those experiments demonstrate that the results are robust with respect to moderate deviations away from exactly sparse models.)

The regressors are correlated with covariance $\Theta_{ij} = \rho^{|i-j|}$ and $\rho = 0.5$. The coefficient $c_d$ is used to control the $R^2$, denoted $R^2_d$, in the equation relating main regressor to the controls, and $c_y$ is used to control the $R^2$, denoted $R^d_y$, for the regression equation: $\tilde y - d\alpha_0 = x'\{c_y\nu_y\}+\epsilon$, where $\epsilon$ is logistic noise with unit variance. In the simulations below we will use different values of $\alpha_0$, $c_y$ and $c_d$, which induce different data-generating processes (dgps). For every replication, we draw new errors $v_{i}$'s and controls $x_i$'s. The regression functions $x'(c_y\nu_y)$ and $x'(c_d\nu_d)$ are sparse. As we vary the coefficients $c_y$ and $c_d$, we induce different amounts of “signal" strength, making it easier or harder for the Lasso-type methods to detect the controls with non-zero coefficients.

In Figure (ref) we consider a dgp with $\alpha_0=0.2$ and $R^2_d = R^2_y=0.75$, induced by setting $c_d = 1$ and $c_y = 0.75$. We performed 5000 Monte-Carlo simulations. Figure (ref) summarizes the performance and displays the distribution of the following estimators, which are centered by the true value $\alpha_0$ and studentized by their standard deviation:

itemize• Naive post selection estimator -- estimator of $\alpha_0$ based on logistic regression after the naive selection using $\ell_1$-penalized logistic regression; • Optimal instrument estimator -- estimator of $\alpha_0$ based on the instrumental logistic regression with the optimal instrument, as defined in Table (ref); • Double selection estimator -- estimator of $\alpha_0$ based on the logistic regression after double selection, as defined in Table (ref).

The optimal IV estimator and the double selection estimator have distribution approximately centered at the true value, with distribution agreeing closely with the standard normal distribution. They have low biases, low root mean squared errors, and confidence regions have rejection rates close to the nominal level of .05. This good performance is well aligned with our theoretical results that we have developed in the previous section. In sharp contrast, the distribution of the naive post selection estimator seems to deviate substantially from the normal distribution. This estimator exhibits large bias and high root mean squared error compared to the former procedures. This occurs because in this dgp, perfect selection is not achieved, and the resulting “moderate" selection mistakes create a large omitted variable bias. Thus, if we use naive post selection estimator with the standard normal distribution for constructing confidence intervals or performing hypothesis testing, we shall end up with rather misleading inference. This poor performance is well aligned with theoretical predictions of leeb:potscher:hodges,LeebPotscher2006,LeebPotscher2005 in the context of linear models.

figure[figure omitted — 983 chars of source]

We now examine the performance more systematically by varying

equation[equation omitted — 141 chars of source]

This gives us 300 different dgps. For each dgp we performed 1000 Monte-Carlo simulations. Figures (ref)-(ref) display the rejection frequencies of confidence regions with (nominal) significance level of .05 and Figure (ref) displays the root mean squared errors of the estimators of $\alpha_0$. The goal of this exercise is to verify numerically how good the uniformity claims derived in Corollaries (ref) and (ref) are, and also confirm that the previous conclusions continue to hold across a wide set of dgp.

In Figures (ref)-(ref) we consider the rejection (non-coverage) frequencies of confidence regions based on: naive post selection logistic estimator\footnote{This region is given by $\{ |\alpha-\widetilde\alpha|\leqslant \{{\mathbb{E}_n}[\widehat w_i(d_i,x_{i{\rm support}(\widetilde\beta)}')'(d_i,x_{i{\rm support}(\widetilde\beta)}')]\}^{-1/2}_{11}\Phi^{-1}(1-\xi/2)/\sqrt{n}\}$ where $(\widetilde\alpha,\widetilde\beta)$ is the naive post selection logistic estimator.}, optimal IV (${\mathcal{CR}}_D$ and ${\mathcal{CR}}_I$) and double selection (${\mathcal{CR}}_{DS}$). These figures illustrate the uniformity properties of the confidence regions based on the discussed estimators. The ideal figure would be a flat surface with the rejection frequency of the true value equal to the nominal level of $.05$. The confidence regions based on the naive post selection perform very poorly, and deviate strongly away from the ideal level of $.05$ throughout large parts of the model space (induced by ((ref))). In contrast, the confidence regions based on optimal IV and double selection seem to be substantially closer to the ideal level, which is in line with our theoretical results in Section (ref). The double selection estimator seems to outperform the estimator based on the explicit construction of the optimal instrument (e.g., the rejection rates and the RMSE for the case with $\alpha=.5$, where optimal instrument procedure tends to perform noticeably worse.) Thus, based on the theoretical results and on the Monte-Carlo results, we recommend the use of the double selection estimator over the optimal IV estimator and, certainly, over the naive post selection estimator.

figure[figure omitted — 637 chars of source]
figure[figure omitted — 645 chars of source]
figure[figure omitted — 651 chars of source]
figure[figure omitted — 978 chars of source]

Discussion

Relation between Double Selection and Optimal Instrument

In this section, we provide a more formal connection between the two proposed methods. It turns out that the construction of the double selection estimator implicitly approximates the optimal instrument $z_{0i}=v_i/\sqrt{w_i}$. This occurs because the model selection procedure in Step 2 associated with ((ref)) allows the estimator to achieve uniformity properties. To see that, using the notation in Table (ref) where $\widehat\beta$, $\widehat\theta$ and $\widetilde \theta$ are defined, let $\widehat T^* = {\rm support}(\widehat\beta)\cup {\rm support}(\widehat\theta)$ denote the variables selected in Step 1 and 2. By the first order conditions of the double selection logistic regression of Step 3 in Table (ref) we have $${\mathbb{E}_n}[ \{y_i-{G}(d_i\check\alpha + x_i'\check\beta)\} (d_i, \ x_{i\widehat T^*}')']=0$$ which creates an orthogonal relation to any linear combination of $(d_i, \ x_{i\widehat T^*}')'$. In particular, by taking the linear combination $(d_i, \ x_{i\widehat T^*}')(1,-\widetilde\theta')'=d_i - x_i'\widetilde\theta =\widehat z_i$, we have $${\mathbb{E}_n}[ \{y_i-{G}( d_i\check\alpha + x_i'\check\beta)\} \widehat z_i ]=0.$$ Therefore the double selection estimator $\check\alpha$ minimizes $$\widetilde L_n(\alpha) = \frac{\| {\mathbb{E}_n}[\{ y_i - {G}(d_i\alpha + x_i'\check\beta) \} \widehat z_i ] \|^2}{{\mathbb{E}_n}[\{ y_i - {G}(d_i\alpha + x_i'\check\beta) \}^2 \widehat z_i^2 ]}, $$ where $\widehat z_i$ is the instrument of the optimal instrument estimator which was {\it implicitly} created. Thus, the double selection estimator can be seen as an iterated version of the method based on instruments where $\widetilde \beta$ is replaced with $\check \beta$. Although their first order asymptotic properties coincide, in finite sample, the double selection method seems to obtain better estimates leading to a more robust performance.

Relation to Neyman's $C(\alpha)$ test

Next we discuss connections between the proposed approach and Neyman's $C(\alpha)$ test Neyman1959,Neyman1979 (here we draw on the discussion in BelloniChernozhukovKato2013a). For the sake of exposition, we assume the instruments are known and i.i.d. observations. As stated in ((ref)) and ((ref)) we rely on instruments satisfying the two equations: $$ {\mathrm{E}}[\{y_i-{G}(d_i\alpha_0+x_i'\beta_0)\} z_{0i}] =0 \ \ \mbox{and} \ \ \left.\frac{\partial}{\partial \beta} {\mathrm{E}}[\{y_i -{G}(d_i\alpha_0+x_i'\beta) \} z_{0i}] \right|_{\beta=\beta_0} = {\mathrm{E}}[w_i z_{0i}x_i]=0. $$ These conditions allows us to construct regular, $\sqrt{n}$-consistent estimators of $\alpha_0$, despite the fact that nonregular, non $\sqrt{n}$-consistent estimator for $\beta_0$ are being used to cope with high-dimensionality. In particular, regularized or post model selection estimators can be used as estimators of $\beta_0$. Neyman's $C(\alpha)$ test was motivated by the same idea which motivates the use of the term “Neymanization" to describe such procedure. Although there will be many instruments $z_{0i}$ that can achieve the property stated above, the choice $z_{0i}=v_i/\sigma_i$ proposed in Section (ref) is optimal as it minimizes the asymptotic variance of the resulting estimators.

Generally, valid (but not necessarily optimal) instruments can be constructed by generalizing the weighted equation ((ref)) to

equation[equation omitted — 115 chars of source]

where $f_i=f(d_i,z_i)$ is a nonnegative weight, and setting the instrument as $z_{0i} := f_i\tilde v_i/w_i$. Because of the zero-mean condition in ((ref)), and provided that $m_0 \in \mathcal{H}$, the function $m_0(x_i)$ in ((ref)) is the solution of the following weighted least squares problem

equation[equation omitted — 100 chars of source]

where $\mathcal{H}$ denotes the set of measurable functions $h$ satisfying ${\mathrm{E}}[f_i^2 h^2(x_i) ]< \infty$ for each $i$. In the current high-dimensional setting, it is assumed that $m_0(x_i)$ can be written as a sparse combination of the controls, namely $m_0(x_i)=x_i'\theta_0$ with $\|\theta_0\|_0 \leqslant s$, so that

equation[equation omitted — 126 chars of source]

This permits the use of Lasso or Post-Lasso to estimate $\theta_0$ which in turn can be used to construct an estimate of $z_{0i}$. Naturally, if the function $m_0$ satisfies different structured properties that could motivate different estimators (for example, we can use ridge estimators if the $m_0$ is “dense" with respect to $x$).

Our technical results establish that, uniformly over $\{ \alpha : \sqrt{n}|\alpha - \alpha_0| \leqslant C \}$,

equation[equation omitted — 182 chars of source]

for the estimators $\widehat\beta$ proposed in this work. This does not require $\widehat \beta$ to converge at root-$n$ rate to $\beta_0$ (which generically is not achievable in the current setting) though we do impose the sparsity condition $s^2 \log^2 (p\vee n)/n \to 0$ to guarantee that $\|\widehat \beta- \beta \| = o_{\mathrm{P}}(n^{-1/4})$. Equation ((ref)) implies that the empirical estimating equations behave as if $\beta_0$ was used instead of $\widehat\beta$. Hence, for estimation, we can use the instrumental logistic regression estimator, namely $\check \alpha$ as a minimizer of the statistic $$ n L_n(\alpha) = \| \sqrt{n} {\mathbb{E}_n}[\{y_i - {G}(d_i\alpha+x_i'\widehat \beta)\} z_{0i}]\|^2 \ \ / \ \ {\mathbb{E}_n}[\{y_i - {G}(d_i\alpha+x_i'\widehat \beta)\}^2 z_{0i}^2]. $$

From ((ref)) we have that $\theta_0= \bar {\mathrm{E}}[f^2_i x_i x_i' ]^{-}\bar {\mathrm{E}}[f^2_i d_i x_i ]$, where $A^-$ denotes a generalized inverse of $A$. Letting $\widehat\varepsilon_i(\alpha) = y_i - {G}(d_i\alpha+x_i'\widehat \beta)$ and $$ z_{0i} = f_i\tilde v_i/w_i = (f_i^2/w_i) d_i- (f_i^2/w_i) x_i' \bar {\mathrm{E}}[f^2_i x_i x_i' ]^{-}\bar {\mathrm{E}}[f^2_i d_i x_i ],$$ $n L_n(\alpha)$ can be rewritten as a (perhaps) familiar version of Neyman's $C(\alpha)$ statistic $$ n L_n(\alpha) = \frac{\| \sqrt{n} \{ {\mathbb{E}_n}[ \widehat\varepsilon_i(\alpha) (f_i^2/w_i) d_i] - {\mathbb{E}_n}[ \widehat\varepsilon_i(\alpha)(f_i^2/w_i) x_i' ] \bar {\mathrm{E}}[f^2_i x_i x_i' ]^{-}\bar {\mathrm{E}}[f^2_i d_i x_i ] \} \|^2}{{\mathbb{E}_n}[\widehat\varepsilon_i^2(\alpha) z_{0i}^2]}. $$ Thus, our IV estimator minimizes a Neyman's $C(\alpha)$ statistic for testing point hypotheses about $\alpha$. Hence our construction builds on the classical ideas of Neyman for dealing with (hard-to-estimate) nuisance parameters.

An estimator $\check \alpha$ that minimizes the criterion $n L_n$ up to a $o_{\mathrm{P}}(1)$ term satisfies $$ \tilde \Sigma_n^{-1} \sqrt{n} (\check \alpha - \alpha_0) \rightsquigarrow N(0,1), \ \ \ \tilde\Sigma^2_n = \bar {\mathrm{E}}[ w_i d_i z_{0i} ]^{-2} \bar {\mathrm{E}}[ \sigma_i^2z_{0i}^2]. $$ It is not difficult to check that using $f_i = w_i/\sigma_i$ leads to the smallest possible value $\bar {\mathrm{E}}[ v_i^2]^{-1}$ of $\tilde\Sigma^2_n$. Therefore, $z_{0i}=v_i/\sigma_i$ is the optimal instrument among all instruments that can be derived by the preceding approach. Using the optimal instrument translates into more precise estimators, smaller confidence regions, and better power for testing based on either $\check \alpha$ or $n L_n$.

Relation to Minimax Efficiency for Logistic Model

Next we consider a connection to the (local) minimax efficiency analysis from the semiparametric literature, where we follow the discussion in BelloniChernozhukovKato2013a. Our model is a special case of a partially linear logistic model, and kosorok:book derives an efficient score function for the latter: $$ S_i= \{y_i - {G}(d_i\alpha_0+x_i'\beta_0)\} \{d_i - m_0^*(x_i)\}, $$ where$$ m^*_0(x_i) = \frac{{\mathrm{E}}[w_i d_i|x_i]}{{\mathrm{E}}[w_i |x_i]}. $$ We note that $m_0^*(x_i)$ is $m_0(x_i)$ in ((ref)) induced by the weight $f_i =\sqrt{w_i}$. Thus, the efficient score function can be reexpressed as: $$ S_i = \{y_i - {G}(d_i\alpha_0+x_i'\beta_0)\} v_i/\sqrt{w_i}, $$ where $v_i$ is defined via ((ref)). Using this score leads to the same estimating equations as those constructed above using Neymanization (with an optimal instrument). It follows that the estimator based on the instrument $z_{0i}=v_i/\sqrt{w_i}$ is efficient in the local minimax sense (see Theorem 18.4 in kosorok:book), and inference about $\alpha_0$ based on this estimator provides best minimax power against local alternatives (see Theorem 18.12 in kosorok:book).

The preceding claim is formal provided that the least favorable submodels are permitted as deviations within the overall set of potential models $\mathcal{Q}_n$ (defined similarly to Corollary (ref)). Specifically, given a law $Q_n$, there should be a suitable neighborhood $\mathcal{Q}_n^{\delta}$ of $Q_n$ such that $Q_n \in \mathcal{Q}_n^{\delta} \subset \mathcal{Q}_n$. For that, we assume $m^*_0(x_i)=x_i'\theta_0$ and consider a collection of models indexed by $t = (t_1, t_2)$ satisfying:

eqnarray[eqnarray omitted — 248 chars of source]

where $\|\beta_0\|_0 + \| \theta_0\|_0 \leqslant s$ and Condition L as in Section (ref) hold. By construction, the model associated with $t=0$ generates precisely the model $Q_n$. As $t$ varies within a $\delta$-ball, we generate the set of models $\mathcal{Q}_n^{\delta}$ that contains the least favorable deviations, and which still belong to $\mathcal{Q}_n$. As shown in kosorok:book, $S_i$ is the efficient score for such parametric submodel so we cannot have a better regular estimator than the estimator whose influence function is $\Sigma_n S_i$. Because the set of models $\mathcal{Q}_n$ contains $ \mathcal{Q}_n^{\delta}$, all the formal conclusions about (local minimax) optimality of the proposed estimators hold from theorems cited above (using subsequence arguments to handle models changing with $n$).