EconBase
← Back to paper

Valid Post-Selection Inference in High-Dimensional Approximately Sparse Quantile Regression Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

144,285 characters · 23 sections · 63 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Valid Post-Selection Inference in High-Dimensional Approximately Sparse Quantile Regression Models

\def\spacingset#1{ {#1}} \spacingset{1}

\if11 {

} \fi

\if01 {

center[center omitted — 128 chars of source]

} \fi

abstractThis work proposes new inference methods for a regression coefficient of interest in a (heterogeneous) quantile regression model. We consider a high-dimensional model where the number of regressors potentially exceeds the sample size but a subset of them suffice to construct a reasonable approximation to the conditional quantile function. The proposed methods are (explicitly or implicitly) based on orthogonal score functions that protect against moderate model selection mistakes, which are often inevitable in the approximately sparse model considered in the present paper. We establish the uniform validity of the proposed confidence regions for the quantile regression coefficient. Importantly, these methods directly apply to more than one variable and a continuum of quantile indices. In addition, the performance of the proposed methods is illustrated through Monte-Carlo experiments and an empirical example, dealing with risk factors in childhood malnutrition.

{\it Keywords:} quantile regression, confidence regions post model selection, orthogonal score functions

\spacingset{1.45}

Introduction

Many applications of interest require the measurement of the distributional impact of a policy (or treatment) on the relevant outcome variable. Quantile treatment effects have emerged as an important concept for measuring such distributional impacts (see, e.g., Lehmann,koenker:book). In this work we focus on the quantile treatment effect $\alpha_\tau$ on a policy/treatment variable $d$ of an outcome of interest $y$ in the (heteroskedastic) partially linear model: $$ \tau\textrm{-quantile}(y\mid z, d) = d\alpha_\tau + g_\tau(z), $$ where $g_\tau$ is the (unknown) confounding function of the other covariates $z$ which can be well approximated by a linear combination of $p$ technical controls. Large $p$ arises due to the existence of many features (as in genomic studies and econometric applications) and/or the use of basis expansions in non-parametric approximations. When $p$ is comparable or larger than the sample size $n$, this brings forth the need to perform model selection or regularization.

We propose methods to construct estimates and confidence regions for the coefficient of interest $\alpha_\tau$ based upon robust post-selection procedures. We establish the (uniform) validity of the proposed methods in a non-parametric setting. Model selection in those settings generically leads to a moderate mistake and traditional arguments based on perfect model selection do not apply. Therefore, the proposed methods are developed to be robust to such model selection mistakes. Furthermore, they are directly applicable to construction of simultaneous confidence bands when $d$ is multivariate and a continuum of quantile indices is of interest.

Broadly speaking the main obstacle to construct confidence regions with (asymptotically) correct coverage is the estimation of the confounding function $g_\tau$ that is (generically) non-regular because of the high-dimensionality. To overcome this difficulty we construct (explicitly or implicitly) an orthogonal score function which leads to a moment condition that is immune to first-order mistakes in the estimation of the confounding function $g_\tau$. The construction is based on preliminary estimation of the confounding function and properly partialling out the confounding factors $z$ from the policy/treatment variable $d$. The former can be achieved via $\ell_1$-penalized quantile regression BC-SparseQR,kato or post-selection quantile regression based on $\ell_1$-penalized quantile regression BC-SparseQR. The latter is carried out by heteroscedastic post-Lasso T1996,BellChenChernHans:nonGauss applied to a density-weighted equation. Then we propose two estimators for $\alpha_{\tau}$ based on: (i) a moment condition based on an orthogonal score function; (ii) a density-weighted quantile regression with all the variables selected in the previous steps. The latter method is reminiscent of the “post-double selection" method proposed in BCH2011:InferenceGauss,BelloniChernozhukovHansen2011. Explicitly or implicitly the last step estimates $\alpha_\tau$ by minimizing a Neyman-type score statistic Neyman1979.\footnote{We mostly focus on selection as a means of regularization, but certainly other regularizations (e.g. the use of $\ell_1$-penalized fits per se) are possible, although they perform no better than the methods we focus on.}

Under mild moment conditions and approximate sparsity assumptions, we establish that the estimator $\check\alpha_\tau$, as defined by either method (see Algorithms (ref) and (ref) below), is root-$n$ consistent and asymptotically normal,

equation[equation omitted — 126 chars of source]

where $\rightsquigarrow$ denotes convergence in distribution; in addition the estimator $\check{\alpha}_{\tau}$ admits a (pivotal) linear representation. Hence the confidence region defined by

equation[equation omitted — 177 chars of source]

has asymptotic coverage probability of $1-\xi$ provided that the estimate $\widehat \sigma^2_n$ is consistent for $\sigma_n^2$, namely, $\widehat\sigma_n^2/\sigma_n^2 = 1+ o_{{\mathrm{P}}}(1)$. In addition, we establish that a Neyman-type score statistic $L_n(\alpha)$ is asymptotically distributed as the chi-squared distribution with one degree of freedom when evaluated at the true value $\alpha=\alpha_\tau$, namely,

equation[equation omitted — 86 chars of source]

which in turn allows the construction of another confidence region:

equation[equation omitted — 152 chars of source]

which has asymptotic coverage probability of $1-\xi$. These convergence results hold under array asymptotics, permitting the data-generating process ${\mathrm{P}} = {\mathrm{P}}_n$ to change with $n$, which implies that these convergence results hold uniformly over large classes of data-generating processes. In particular, our results do not require separation of regression coefficients away from zero (the so-called “beta-min" conditions) for their validity. Importantly, we discuss how the procedures naturally allow for construction of simultaneous confidence bands for many parameters based on a (pivotal) linear representation of the proposed estimator.

Several recent papers study the problem of constructing confidence regions after model selection while allowing $p\gg n$. In the case of linear mean regression, BCH2011:InferenceGauss proposes a double selection inference in a parametric setting with homoscedastic Gaussian errors; BelloniChernozhukovHansen2011 studies a double selection procedure in a non-parametric setting with heteroscedastic errors; c.h.zhang:s.zhang and vandeGeerBuhlmannRitov2013 propose methods based on $\ell_1$-penalized estimation combined with “one-step" correction in parametric models. Going beyond mean regression models, vandeGeerBuhlmannRitov2013 provides high level conditions for the one-step estimator applied to smooth generalized linear problems, BelloniChernozhukovKato2013a analyzes confidence regions for a parametric homoscedastic LAD regression model under primitive conditions based on the instrumental LAD regression, and BelloniChernozhukovWei2013 provides two post-selection procedures to build confidence regions for the logistic regression. None of the aforementioned papers deal with the problem of the present paper.

Although related in spirit with our previous work BelloniChernozhukovHansen2011,BelloniChernozhukovKato2013a,BelloniChernozhukovWei2013, new tools and major departures from the previous works are required. First is the need to accommodate the non-differentiability of the loss function (which translates into discontinuity of the score function) and the non-parametric setting. In particular, we establish new finite sample bounds for the prediction norm on the estimation error of $\ell_1$-penalized quantile regression in nonparametric models that extend results of BC-SparseQR,kato. Perhaps more importantly, the use of post-selection methods in order to reduce bias and improve finite sample performance requires sparsity of the estimates. Although sharp sparsity bounds for $\ell_1$-penalized methods are available for smooth loss functions, those are not available for quantile regression precisely due to the lack of differentiability. This led us to developing sparse estimates with provable guarantees by suitable truncation while preserving the good rates of convergence despite of possible additional model selection mistakes. To handle heteroscedsaticity, which is a common feature in many applications, consistent estimation of the conditional density is necessary whose analysis is new in high-dimensions. Those estimates are used as weights in the weighted Lasso estimation for the auxiliary regression (((ref)) below). In addition, we develop new finite sample bounds for Lasso with estimated weights as the zero mean condition is not assumed to hold for each observation but rather to the average across all observations. Because the estimation of the conditional density function is at a slower rate, it affects penalty choices, rates of convergence, and sparsity of the Lasso estimates.

This work and some of the papers cited above achieve an important uniformity guarantee with respect to the (unknown) values of the parameters. These uniform properties translate into more reliable finite sample performance of the proposed inference procedures because they are robust with respect to (unavoidable) model selection mistakes. There is now substantial theoretical and empirical evidence on the potential poor finite sample performance of inference methods that rely on perfect model selection when applied to models without separation from zero of the coefficients (i.e., small coefficients). Most of the criticism of these procedures are consequence of negative results established in LeebPotscher2005, leeb:potscher:hodges, and the references therein.

Notation. We work with triangular array data $\{ \omega_{i,n} : i=1,\dots,n; n=1,2,3,\dots\}$ where for each $n$, $\{ \omega_{n,i} ; i=1,\dots,n \}$ is defined on the probability space $(\Omega, \mathcal{S}, {\mathrm{P}}_n)$. Each $\omega_{i,n}= (y_{i,n}', z_{i,n}', d_{i,n}')'$ is a vector which are i.n.i.d., that is, independent across $i$ but not necessarily identically distributed. Hence all parameters that characterize the distribution of $\{\omega_{i,n} : i=1,\dots,n\}$ are implicitly indexed by ${\mathrm{P}}_n$ and thus by $n$. We omit this dependence from the notation for the sake of simplicity. We use ${\mathbb{E}_n}$ to abbreviate the notation $n^{-1}\sum_{i=1}^n$; for example, ${\mathbb{E}_n}[f] := {\mathbb{E}_n}[f(\omega_{i})] := n^{-1} \sum_{i=1}^n f(\omega_{i})$. We also use the following notation: $\bar {\mathrm{E}}[f] := {\mathrm{E}} \left [ {\mathbb{E}_n}[f] \right] = {\mathrm{E}} \left [ {\mathbb{E}_n}[f(\omega_{i})] \right ]= n^{-1} \sum_{i=1}^n {\mathrm{E}}[f(\omega_{i})]$. The $\ell_2$-norm is denoted by $\|\cdot\|$; the $\ell_0$-“norm” $\|\cdot\|_0$ denotes the number of non-zero components of a vector; and the $\ell_{\infty}$-norm $\| \cdot \|_{\infty}$ denotes the maximal absolute value in the components of a vector. Given a vector $\delta \in {\Bbb{R}}^p$ and a set of indices $T \subset \{1,\ldots,p\}$, we denote by $\delta_T \in {\Bbb{R}}^p$ the vector in which $\delta_{Tj} = \delta_j$ if $j\in T$ and $\delta_{Tj}=0$ if $j \notin T$.

Setting

For a quantile index $\tau \in (0,1)$, we consider a partially linear conditional quantile model

equation[equation omitted — 160 chars of source]

where $y_i$ is the outcome variable, $d_i$ is the policy/treatment variable, and confounding factors are represented by the variables $z_i$ which impact the equation through an unknown function $g_\tau$. We shall use a large number $p$ of technical controls $x_i=X(z_i)$ to achieve an accurate approximation to the function $g_\tau$ in ((ref)) which takes the form:

equation[equation omitted — 101 chars of source]

where $r_{{\tau} i}$ denotes an approximation error. We view $\beta_\tau$ and $r_{{\tau}}$ as nuisance parameters while the main parameter of interest is $\alpha_\tau$ which describes the impact of the treatment on the conditional quantile (i.e., quantile treatment effect).

In order to perform robust inference with respect to model selection mistakes, we construct a moment condition based on a score function that satisfies an additional orthogonality property that makes them immune to first-order changes in the value of the nuisance parameter. Letting $f_i = f_{\epsilon_i}(0\mid d_i, z_i)$ denote the conditional density at 0 of the disturbance term $\epsilon_i$ in ((ref)), the construction of the orthogonal moment condition is based on the linear projection of the regressor of interest $d_i$ weighted by $f_i$ on the $x_i$ variables weighted by $f_i$

equation[equation omitted — 131 chars of source]

where $\theta_{0\tau} \in \arg\min \bar {\mathrm{E}}[ f_i^2(d_i-x_i'\theta)^2]$. The orthogonal score function $\psi_i(\alpha):=(1\{y_i\leqslant d_i\alpha+ x_i'\beta_\tau + r_{{\tau}}\}-\tau)v_i$ leads to a moment condition to estimate $\alpha_\tau$,

equation[equation omitted — 125 chars of source]

and satisfies the following orthogonality condition with respect to first-order changes in the value of the nuisance parameters $\beta_\tau$ and $\theta_{0\tau}$:

equation[equation omitted — 422 chars of source]

In order to handle the high-dimensional setting, we assume that $\beta_\tau$ and $\theta_{0\tau}$ are approximately sparse, namely, it is possible to choose sparse vector $\beta_\tau$ and $\theta_\tau$ such that:

equation[equation omitted — 229 chars of source]

The latter equation requires that it is possible to choose the sparsity index $s$ so that the mean squared approximation error is of no larger order than the variance of the oracle estimator for estimating the coefficients in the approximation. See chen:Chapter for a detailed discussion of this notion of approximate sparsity.

Methods

The methodology based on the orthogonal score function ((ref)) can be used for construction of many different estimators that have the same first-order asymptotic properties but potentially different finite sample behaviors. In the main part of the paper we present two such procedures in detail (the discussion on additional variants can be found in Subsection (ref) of the Supplementary Appendix). Our procedures use $\ell_1$-penalized quantile regression and $\ell_1$-penalized weighted least squares as intermediate steps (we collect the recommended choices of the user-chosen parameters in Remark (ref) below). The first procedure stated in Algorithm (ref) is based on the explicit construction of the orthogonal score function.

algorithm[algorithm omitted — 1,135 chars of source]

Step 6 of Algorithm (ref) solves the empirical analog of ((ref)). We will show validity of the confidence regions for $\alpha_\tau$ defined in ((ref)) and ((ref)). We note that the truncation in Step 2 for the solution of the penalized quantile regression (provably) induces a sparse solution with the same rate of convergence as the original estimator. This is required because the post-selection methods exhibit better finite sample behaviors in our simulations.

The second algorithm is based on selecting relevant variables from equations ((ref)) and ((ref)), and running a weighted quantile regression.

algorithm[algorithm omitted — 764 chars of source]

Although the orthogonal score function is not explicitly constructed in Algorithm (ref), inspection of the proof reveals that an orthogonal score function is constructed implicitly via the optimality conditions of the weighted quantile regression in Step 4.

remark[Choices of User-Chosen Parameters] For $\gamma = 0.05/n$, we set the penalty levels for the heteroscedastic Lasso and the $\ell_1$-penalized quantile regression as \begin{equation} \lambda := 1.1\sqrt{n}2\Phi^{-1}(1-\gamma/2p) \quad and \quad \lambda_\tau := 1.1\sqrt{n\tau(1-\tau)}\Phi^{-1}(1-\gamma/2p). \end{equation} The penalty loading $\widehat \Gamma_\tau= \operatorname{diag} [ \widehat \Gamma_{\tau jj} , j =1,\dots,p]$ is a diagonal matrix defined by the following procedure: (1) Compute the post-Lasso estimator $\widetilde \theta_\tau^{0}$ based on $\lambda$ and initial values $\widehat \Gamma_{\tau jj} = \max_{1 \leqslant i\leqslant n} \|\widehat f_i x_i\|_\infty \{{\mathbb{E}_n}[\widehat f_i^2d_i^2]\}^{1/2}$. (2) Compute the residuals $\widehat v_i = \widehat f_i(d_i - x_i'\widetilde\theta_\tau^{0})$ and update \begin{equation} \widehat \Gamma_{\tau jj} = \sqrt{{\mathbb{E}_n}[\widehat f_i^2 x^2_{ij} \widehat v_i^2]}, \ j=1,\ldots,p. \end{equation} In Algorithm (ref) we have used the following parameter space for $\alpha$: \begin{equation} \mathcal{A}_\tau = \{ \alpha \in {\Bbb{R}} : |\alpha - \widetilde \alpha_\tau| \leqslant 10\{{\mathbb{E}_n}[d_i^2]\}^{-1/2} / \log n \}. \end{equation}
remark[Estimating Standard Errors] There are different possible choices of estimators for $\sigma_n$: \begin{equation} \begin{array}{l} \widehat\sigma_{1n}^2 :=\tau(1-\tau) \left ( {\mathbb{E}_n}[\widetilde v_i^2] \right )^{-1}, \quad \widehat\sigma_{2n}^2 := \tau(1-\tau) \left [ \{ {\mathbb{E}_n}[\widehat f_i^2( d_i, x_{i\check T}')'(d_i,x_{i\check T}')] \}^{-1} \right ]_{11}, \\ \widehat\sigma_{3n}^2 := \left ( {\mathbb{E}_n}[\widehat f_i d_i \widetilde v_i]\right)^{-2}{\mathbb{E}_n}[(1\{y_i\leqslant d_i\check\alpha_\tau+x_i'\check\beta_\tau\}-\tau)^2\widetilde v_i^2], \end{array} \end{equation} where $\check T = {\rm supp}(\check \beta_\tau) \cup {\rm supp}(\widehat \theta_\tau)$ is the set of controls used in the double selection quantile regression. Although all three estimates are consistent under similar regularities conditions, their finite sample behaviors might differ. Based on the small-sample performance in computational experiments, we recommend the use of $\widehat\sigma_{3n}$ for the orthogonal score estimator and $\widehat\sigma_{2n}$ for the double selection estimator.

Estimation of Conditional Density Function

The implementation of the algorithms in Section (ref) requires an estimate of the conditional density function $f_i$ which is typically unknown under heteroscedasticity. Following koenker:book, we shall use the observation that $1/ f_i = \partial Q(\tau\mid d_i, z_i)/\partial \tau$ to estimate $f_i$ where $Q(\cdot\mid d_i,z_i)$ denotes the conditional quantile function of the outcome. Let $\widehat Q(u\mid z_i,d_i)$ denote an estimate of the conditional $u$-quantile function $Q(u\mid z_i,d_i)$, based on either $\ell_1$-penalized quantile regression or an associated post-selection method, and let $h =h_n \to 0$ denote a bandwidth parameter. Then an estimator of $f_i$ can be constructed as

equation[equation omitted — 120 chars of source]

When the conditional quantile function is three times continuously differentiable, this estimator is based on the first order partial difference of the estimated conditional quantile function, and so it has the bias of order $h^2$. Under additional smoothness assumptions, an estimator that has a bias of order $h^4$ is given by {

equation[equation omitted — 227 chars of source]

}

remark[Implementation of the estimates $\widehat f_i$] There are several possible choices of tuning parameters to construct the estimates $\widehat f_i$. In particular the bandwidth choices set in the R package `quantreg' from koenker2016quantreg exhibits good empirical behavior. In our theoretical analysis we coordinate the bandwidth choice with the choice of the penalty level of the density weighted Lasso. In Subsection (ref) of the Supplementary Appendix we discuss in more detail the requirements associated with different choices for penalty level $\lambda$ and bandwidth $h$. Together with the recommendations made in Remark (ref), we suggest to construct $\widehat f_i$ as in ((ref)) with bandwidth $h:= \min\{ n^{-1/6}, \tau(1-\tau)/2\}$.

Theoretical Analysis

Regularity Conditions

In this section we provide regularity conditions that are sufficient for validity of the main estimation and inference results. In what follows, let $c,C$, and $q$ be given (fixed) constants with $c > 0, C \geqslant 1$ and $q \geqslant 4$, and let $\ell_n \uparrow \infty, \delta_n \downarrow 0$, and $\Delta_n \downarrow 0$ be given sequences of positive constants. We assume that the following condition holds for the data generating process ${\mathrm{P}} = {\mathrm{P}}_n$ for each $n$.

Condition AS(${\mathrm{P}}$). (i) Let $\{(y_i,d_i,x_i=X(z_i)) : i=1,\ldots,n\}$ be independent random vectors that obey the model described in ((ref)) and ((ref)) with $\|\theta_{0\tau}\|+\|\beta_\tau\|+|\alpha_\tau|\leqslant C$. (ii) There exists $s \geqslant 1$ and vectors $\beta_{\tau}$ and $\theta_\tau$ such that $x_i'\theta_{0\tau} = x_i' \theta_{\tau} + r_{{\theta \tau} i}, \ \|\theta_{\tau}\|_0 \leqslant s, \ \bar {\mathrm{E}} [r_{{\theta \tau} i}^2] \leqslant C s/n$, $\|\theta_{0\tau}-\theta_\tau\|_1 \leqslant s\sqrt{\log(pn)/n}$, and $g_\tau(z_i) = x_i' \beta_{\tau} + r_{{\tau} i}, \ \|\beta_\tau\|_0 \leqslant s, \ \bar {\mathrm{E}}[r_{{\tau} i}^2] \leqslant C s/n$. (iii) The conditional distribution function of $\epsilon_i$ is absolutely continuous with continuously differentiable density $f_{\epsilon_i\mid d_i, z_i}(\cdot\mid d_i,z_i)$ such that $0<\underline{f} \leqslant f_i \leqslant \sup_t f_{\epsilon_i\mid d_i, z_i}(t\mid d_i,z_i) \leqslant \bar f \leqslant C$ and $\sup_{t} |f_{\epsilon_i\mid d_i, z_i}'(t\mid d_i, z_i)|\leqslant \bar f'\leqslant C$.

Condition AS(i) imposes the setting discussed in Section (ref) in which the error term $\epsilon_i$ has zero conditional $\tau$-quantile. The approximate sparsity on the high-dimensional parameters is stated in Condition AS(ii). Condition AS(iii) is a standard assumption on the conditional density function in the quantile regression literature (see koenker:book) and the instrumental quantile regression literature (see ch:iqrWeakId). Next we summarize the moment conditions we impose.

Condition M(${\mathrm{P}}$). (i) We have $\bar {\mathrm{E}}[\{(d_{i},x_{i}')\xi\}^2] \geqslant c\|\xi\|^2$ and $\bar {\mathrm{E}}[\{(d_i,x_i')\xi\}^4] \leqslant C\|\xi\|^4$ for all $\xi \in {\Bbb{R}}^{p+1}$, $c \leqslant \min_{1 \leqslant j\leqslant p} \bar {\mathrm{E}}[ |f_ix_{ij}v_i-{\mathrm{E}}[f_ix_{ij}v_i]|^2]^{1/2} \leqslant \max_{1\leqslant j\leqslant p} \bar {\mathrm{E}}[|f_i x_{ij} v_i|^3]^{1/3} \leqslant C$. (ii) The approximation error satisfies $|\bar {\mathrm{E}}[f_iv_ir_{{\tau} i}]|\leqslant \delta_nn^{-1/2}$ and $\bar {\mathrm{E}}[(x_i'\xi)^2r_{{\tau} i}^2]\leqslant C\|\xi\|^2 \bar {\mathrm{E}}[r_{{\tau} i}^2]$ for all $\xi \in {\Bbb{R}}^{p}$. (iii) Suppose that $K_q = {\mathrm{E}}[\max_{1 \leqslant i \leqslant n}\|(d_i,v_i,x_i')' \|_\infty^q]^{1/q}$ is finite and satisfies $(K_q^2s^2 + s^3)\log^3(pn) \leqslant n\delta_n$ and $K_q^4s\log(pn)\log^3n\leqslant \delta_nn$.

Condition M(i) imposes moment conditions on the variables. Condition M(ii) imposes requirements on the approximation error. Condition M(iii) imposes growth conditions on $s$, $p$, and $n$. In particular these conditions imply that the population eigenvalues of the design matrix are bounded away from zero and from above. They ensure that sparse eigenvalues and restricted eigenvalues are well behaved which are used in the analysis of penalized estimators and sparsity properties needed for the post-selection estimator.

remark[Handling Approximately Sparse Models] To handle approximately sparse models to represent $g_\tau$ in ((ref)), we assume a near orthogonality between $r_{{\tau}}$ and $f v$, namely $\bar {\mathrm{E}}[ f_i v_i r_{{\tau} i} ] = o(n^{-1/2})$. This condition is automatically satisfied if the orthogonality condition in ((ref)) can be strengthen to ${\mathrm{E}}[f_iv_i\mid z_i]=0$. However, it can be satisfied under weaker conditions as discussed in Subsection (ref) of the Supplementary Appendix.

Our last set of conditions pertains to the estimation of the conditional density function $(f_i)_{i=1}^n$ which has a non-trivial impact on the analysis. We denote by $\mathcal{U}$ the finite set of quantile indices used in the estimation of the conditional density. Under mild regularity conditions the estimators ((ref)) and ((ref)) achieve

equation[equation omitted — 177 chars of source]

where $\bar k = 2$ for ((ref)) and $\bar k =4$ for ((ref)). Condition D summarizes sufficient conditions to account for the impact of density estimation via (post-selection) $\ell_1$-penalized quantile regression estimators.

{\bf Condition D.} For $u \in \mathcal{U}$, assume that $u\textrm{-quantile}(y_i\mid z_i, d_i) = d_i\alpha_u + x_i'\beta_u + r_{ui}$, $f_{ui}=f_{y_i\mid d_i,z_i}(d_i\alpha_u + x_i'\beta_u + r_{ui} \mid z_i,d_i)\geqslant c$ where $\bar {\mathrm{E}}[r_{u i}^2] \leqslant \delta_n n^{-1/2}$ and $|r_{ui}|\leqslant \delta_nh$ for all $i=1,\ldots,n$, and the vector $\beta_u$ satisfies $\|\beta_u\|_0\leqslant s$. (ii) For $\widetilde s_{\theta \tau} = s+\frac{ns\log (n\vee p)}{h^2\lambda^2}+ \left(\frac{nh^{\bar k}}{\lambda}\right)^2$, suppose $h^{\bar k}\sqrt{\widetilde s_{{\theta \tau}}\log(pn) }\leqslant \delta_n$, $h^{-2}K_q^2s\log(pn)\leqslant \delta_n n$, $\lambda K_q^2\sqrt{s} \leqslant \delta_n n$, $h^{-2}s\widetilde s_{\theta \tau} \log(pn) \leqslant \delta_n n$, $\lambda\sqrt{s \widetilde s_{\theta \tau} \log(pn)} \leqslant \delta_n n$, and $K_q^2\widetilde s_{{\theta \tau}}\log^2(pn)\log^3(n) \leqslant \delta_nn$.

Condition D(i) imposes the approximately sparse assumption for the $u$-conditional quantile function for quantile indices $u$ in a neighborhood of the quantile index $\tau$. Condition D(ii) provides growth conditions relating $s$, $p$, $n$, $h$ and $\lambda$. Subsection (ref) in the Supplementary Appendix discusses specific choices of penalty level $\lambda$ and of bandwidth $h$ together with the implied conditions on the triple $(s,p,n)$. In particular they imply that sparse eigenvalues of order $\widetilde s_{{\theta \tau}}$ are well behaved.

Main results

In this section we state our theoretical results. We establish the first order equivalence of the proposed estimators. We construct the estimators as defined as in Algorithm (ref) and (ref) with parameters $\lambda_\tau$ as in ((ref)), $\widehat \Gamma_\tau$ as in ((ref)), and $\mathcal{A}_\tau$ as in ((ref)). The choices of $\lambda$ and $h$ satisfy Condition D.

theoremLet $\{{\mathrm{P}}_n\}$ be a sequence of data-generating processes. Assume that conditions $\mathrm{AS} ({\mathrm{P}})$, $\mathrm{M} ({\mathrm{P}})$ and $\mathrm{D} ({\mathrm{P}})$ are satisfied with ${\mathrm{P}} = {\mathrm{P}}_n$ for each $n$. Then the orthogonal score estimator $\check \alpha_\tau^{OS}$ based on Algorithm (ref) and the double selection estimator $\check \alpha_\tau^{DS}$ based on Algorithm (ref) are first order equivalent, $\sqrt{n}(\check \alpha_\tau^{OS}-\check \alpha_\tau^{DS})=o_P(1).$ Moreover, either estimator satisfies $$ \sigma_n^{-1} \sqrt{n} (\check \alpha_\tau - \alpha_\tau) = \mathbb{U}_n(\tau) + o_P(1) \quad \mbox{and} \quad \mathbb{U}_n(\tau) \rightsquigarrow N(0,1), $$ where $\sigma^2_n= \tau(1-\tau)\bar {\mathrm{E}} [v_i^2]^{-1}$ and $\mathbb{U}_n(\tau) := \tau(1-\tau) \{ \bar {\mathrm{E}}[v_i^2]\}^{-1/2} n^{-1/2} \sum_{i=1}^n ( \tau - 1\{U_i\leqslant \tau\})v_{i}$, and $U_1, \ldots, U_n$ are i.i.d. uniform random variables on $(0,1)$ independent from $v_1,\ldots, v_n$. Furthermore, $$ nL_n(\alpha_\tau) = \mathbb{U}_n^2(\tau)+ o_P(1) \quad \mbox{and} \quad \mathbb{U}_n^2(\tau)\rightsquigarrow \chi^2(1). $$ The result continues to apply if $\sigma^2_n$ is replaced by any of the estimators in ((ref)), namely, $\widehat\sigma_{kn}/\sigma_n = 1 + o_{\mathrm{P}}(1)$ for $k=1,2,3$.

The asymptotically correct coverage of the confidence regions $\mathcal{C}_{\xi,n}$ and $\mathcal{I}_{\xi,n}$ as defined in ((ref)) and ((ref)) follows immediately. Theorem (ref) relies on post model selection estimators which in turn rely on achieving sparse estimates $\widehat \beta_\tau$ and $\widehat\theta_\tau$. The sparsity of $\widehat\theta_\tau$ is derived in Section (ref) in the Supplemental Appendix under the recommended penalty choices. The sparsity of $\widehat \beta_\tau$ is not guaranteed under the recommended choices of penalty level $\lambda_\tau$ which leads to sharp rates. We bypass that by truncating small components to zero (as in Step 2 of Algorithm (ref)) which (provably) preserves the same rate of convergence and ensures the sparsity.

In addition to the asymptotic normality, Theorem (ref) establishes that the rescaled estimation error $\sigma_n^{-1}\sqrt{n}(\check\alpha_\tau - \alpha_\tau)$ is approximately equal to the process $\mathbb{U}_n(\tau)$, which is pivotal conditional on $v_1,\ldots, v_n$. Such a property is very useful since it is easy to simulate $\mathbb{U}_n(\tau)$ conditional on $v_1,\ldots, v_n$. Thus this representation provides us with another procedure to construct confidence intervals without relying on asymptotic normality which are useful for the construction of simultaneous confidence bands; see Section (ref).

Importantly, the results in Theorem (ref) allow for the data generating process to depend on the sample size $n$ and have no requirements on the separation from zero of the coefficients. In particular these results allow for sequences of data generating processes for which perfect model selection is not possible. In turn this translates into uniformity properties over a large class of data generating processes. Next we formalize these uniform properties. We let $\mathcal{P}_{n}$ denote the collection of distributions ${\mathrm{P}}$ for the data $\{ (y_{i},d_{i}, z_i')' \}_{i=1}^{n}$ such that Conditions AS, M and D are satisfied for given $n$. This is the collection of all approximately sparse models where the above sparsity conditions, moment conditions, and growth conditions are satisfied. Note that the uniformity results for the approximately sparse and heteroscedastic case are new even under fixed $p$ asymptotics.

corollary[Uniform Validity of Confidence Regions] Let $\mathcal{P}_{n}$ be the collection of all distributions of $\{ (y_{i},d_{i},z_i')' \}_{i=1}^{n}$ for which Conditions $\mathrm{AS}$, $\mathrm{M}$, and $\mathrm{D}$ are satisfied for given $n \geqslant 1$. Then the confidence regions $\mathcal{C}_{\xi,n}$ and $\mathcal{I}_{\xi,n}$ defined based on either the orthogonal score estimator or by the double selection estimator are asymptotically uniformly valid $$ \lim_{n \to \infty} \sup_{{\mathrm{P}} \in \mathcal{P}_n} |{\mathrm{P}} ( \alpha_\tau \in \mathcal{C}_{\xi,n} ) - (1- \xi)| =0 \quad \mbox{and} \quad \lim_{n \to \infty} \sup_{{\mathrm{P}} \in \mathcal{P}_n} |{\mathrm{P}} ( \alpha_\tau \in \mathcal{I}_{\xi,n} ) - (1- \xi)| =0. $$

Simultaneous Inference over $\tau$ and Many Coefficients

In some applications we are interested on building confidence intervals that are simultaneously valid for many coefficients as well as for a range of quantile indices $\tau \in \mathcal{T}\subset (0,1)$ a fixed compact set. The proposed methods directly extend to the case of $d \in {\Bbb{R}}^{K}$ and $\tau \in \mathcal{T}$ $$ \tau\textrm{-quantile}(y\mid z, d) = \sum_{j=1}^K d_j\alpha_{\tau j} + \tilde g_\tau(z). $$ Indeed, for each $\tau \in \mathcal{T}$ and each $k=1,\ldots,K$, estimates can be obtained by applying the methods to the model ((ref)) as $$ \tau\textrm{-quantile}(y\mid z, d) = d_k\alpha_{\tau k} + g_\tau(z) \ \ \mbox{where} \ \ g_\tau(z) := \tilde g_\tau(z) + \sum_{j\neq k} d_j\alpha_{\tau j}. $$ For each $\tau \in \mathcal{T}$, Step 1 and the conditional density function $f_i$, $i=1,\ldots,n$, are the same for all $k=1,\ldots,K$. However, Steps 2 and 3 adapt to each quantile index and each coefficient of interest. The uniform validity of $\ell_1$-penalized methods for a continuum of problems (indexed by $\mathcal{T}$ in our case) has been established for quantile regression in BC-SparseQR and for least squares in BCFH2013program. The conclusions of Theorem (ref) are uniformly valid over $k=1,\ldots,K$ and $\tau \in \mathcal{T}$ (in the $\ell_\infty$-norm).

Simultaneous confidence bands are constructed by defining the following critical value $$ c^*(1-\xi) = \inf\left\{ t: {\mathrm{P}}\left( \sup_{\tau \in \mathcal{T},k=1,\ldots,K} |\mathbb{U}_n(\tau,k)| \leqslant t \mid \{d_i,z_i\}_{i=1}^n\right) \geqslant 1-\xi \right\}, $$ where the random variable $\mathbb{U}_n(\tau,k)$ is pivotal conditional on the data, namely, $$ \mathbb{U}_n(\tau,k) := \frac{\{\tau(1-\tau)\bar {\mathrm{E}}[v_{\tau k i}^2]\}^{-1/2}}{\sqrt{n}} \sum_{i=1}^n ( \tau - 1\{U_i\leqslant \tau\})v_{\tau k i}, $$ where $U_i$ are i.i.d. uniform random variables on $(0,1)$ independent from $\{d_i,z_i\}_{i=1}^n$, and $v_{\tau k i}$ is the error term in the decomposition ((ref)) for the pair $(\tau,k)$. Therefore $c^*(1-\xi)$ can be estimated since estimates of $v_{\tau k i}$ and $\sigma_{n\tau k}$, $\tau \in \mathcal{T}$ and $k=1,\ldots,K$, are available. Uniform confidence bands can be defined as $$ [ \check \alpha_{\tau k} - \sigma_{n\tau k}c^*(1-\xi)/\sqrt{n}, \check \alpha_{\tau k} + \sigma_{n\tau k}c^*(1-\xi)/\sqrt{n} ] \ \ \mbox{for} \ \ \tau \in \mathcal{T}, \ k = 1,\ldots,K. $$

Empirical Performance

Monte-Carlo Experiments

Next we provide a simulation study to assess the finite sample performance of the proposed estimators and confidence regions. We focus our discussion on the double selection estimator as defined in Algorithm (ref) which exhibits a better performance. We consider the median regression case ($\tau =1/2$) under the following data generating process:

align[align omitted — 180 chars of source]

where $\alpha_\tau = 1/2$, $\theta_{0j} = 1/j^2, j=1,\ldots,p$, $x = (1,z')'$ consists of an intercept and covariates $z \sim N(0,\Sigma)$, and the errors $\epsilon$ and $\tilde v$ are independent. The dimension $p$ of the covariates $x$ is $300$, and the sample size $n$ is $250$. The regressors are correlated with $\Sigma_{ij} = \rho^{|i-j|}$ and $\rho = 0.5$. In this case, $f_i= 1/\{\sqrt{\pi(2-\mu +\mu d^2)}\}$ so that the coefficient $\mu \in \{0, 1\}$ makes the conditional density function of $\epsilon$ homoscedastic if $\mu = 0$ and heteroscedastic if $\mu = 1$. The coefficients $c_y$ and $c_d$ are used to control the $R^2$ in the equations: $y - d\alpha_\tau = x'(c_y\nu_0) + \epsilon$ and $ d = x'(c_d\nu_0) + \tilde v$ ; we denote the values of $R^2$ in each equation by $R^2_y$ and $R_d^2$. We consider values $(R^2_y, R^2_d)$ in the set $\{0, .1, .2, \ldots, .9\}\times \{0, .1, .2, \ldots, .9\}$. Therefore we have 100 different designs and perform $500$ Monte-Carlo repetitions for each design. For each repetition we draw new vectors $x_i$'s and errors $\epsilon_i$'s and $\tilde v_i$'s.

We perform estimation of $f_i$'s via ((ref)) even in the homoscedastic case ($\mu=0$), since we do not want to rely on whether the assumption of homoscedasticity is valid or not. We use $\widehat \sigma_{2n}$ as the standard error estimate for the post double selection estimator based on Algorithm (ref). As a benchmark we consider the standard (naive) post-selection procedure that applies $\ell_1$-penalized median regression of $y$ on $d$ and $x$ to select a subset of covariates that have predictive power for $y$, and then runs median regression of $y$ on $d$ and the selected covariates, omitting the covariates that were not selected. We report the rejection frequency of the confidence intervals with the nominal coverage probability of $95\%$. Ideally we should see the rejection rate of $5\%$, the nominal level, regardless of the underlying generating process ${\mathrm{P}} \in \mathcal{P}_n$. This is the so called uniformity property or honesty property of the confidence regions (see, e.g., RomanoWolf2007, RomanoShaikh20012, and LeebPotscher2006).

In the homoscedastic case, reported on the left column of Figure (ref), we have the empirical rejection probabilities for the naive post-selection procedure on the first row. These empirical rejection probabilities deviate strongly away from the nominal level of $5\%$, demonstrating the striking lack of robustness of this standard method. This is perhaps expected due to the Monte-Carlo design having regression coefficients not well separated from zero (that is, the “beta min" condition does not hold here). In sharp contrast, we see that the proposed procedure performs substantially better, yielding empirical rejection probabilities close to the desired nominal level of $5\%$. In the right column of Figure (ref) we report the results for the heteroscedastic case ($\mu = 1$). Here too we see the striking lack of robustness of the naive post-selection procedure. We also see that the confidence region based on the post-double selection method significantly outperforms the standard method, yielding empirical rejection probabilities close to the nominal level of $5\%$.

figure[figure omitted — 803 chars of source]

Inference on Risk Factors in Childhood Malnutrition

The purpose of this section is to examine practical usefulness of the new methods and contrast them with the standard post-selection inference (that assumes perfect selection).

We will assess statistical significance of socio-economic and biological factors on children's malnutrition, providing a methodological follow up on the previous studies done by FKH2011 and Koenker2011. The measure of malnutrition is represented by the child's height, which will be our response variable $y$. The socio-economic and biological factors will be our regressors $x$, which we shall describe in more detail below. We shall estimate the conditional first decile function of the child's height given the factors (that is, we set $\tau=.1$). We would like to perform inference on the size of the impact of the various factors on the conditional decile of the child's height. The problem has material significance, so it is important to conduct statistical inference for this problem responsibly.

The data comes originally from the Demographic and Health Surveys (DHS) conducted regularly in more than 75 countries; we employ the same selected sample of 37,649 as in Koenker (2012). All children in the sample are between the ages of 0 and 5. The response variable $y$ is the child's height in centimeters. The regressors $x$ include child's age, breast feeding in months, mothers body-mass index (BMI), mother's age, mother's education, father's education, number of living children in the family, and a large number of categorical variables, with each category coded as binary (zero or one): child's gender (male or female), twin status (single or twin), the birth order (first, second, third, fourth, or fifth), the mother's employment status (employed or unemployed), mother's religion (Hindu, Muslim, Christian, Sikh, or other), mother's residence (urban or rural), family's wealth (poorest, poorer, middle, richer, richest), electricity (yes or no), radio (yes or no), television (yes or no), bicycle (yes or no), motorcycle (yes or no), and car (yes or no).

Although the number of covariates ($p=30$) is substantial, the sample size ($n=37,649$) is much larger than the number of covariates. Therefore, the dataset is very interesting from a methodological point of view, since it gives us an opportunity to compare various methods for performing inference to an “ideal" benchmark of standard inference based on the standard quantile regression estimator without any model selection. This was proven theoretically in He:Shao and in Belloni:Chern:Fern under the $p \to \infty, p^3/n \to 0$ regime. This is also the general option recommended by koenker:book and LeebPotscher2005 in the fixed $p$ regime. Note that this “ideal" option does not apply in practice when $p$ is relatively large; however it certainly applies in the present example.

We will compare the “ideal" option with two procedures. First the standard post-selection inference method. This method performs standard inference on the post-model selection estimator, “assuming" that the model selection had worked perfectly. Second the double selection estimator defined as in Algorithm (ref). (The orthogonal score estimator performs similarly so it is omitted due to space constrains.) The proposed methods do not assume perfect selection, but rather build a protection against (moderate) model selection mistakes.

We now will compare our proposal to the “ideal" benchmark and to the standard post-selection method. We report the empirical results in Table (ref). The first column reports results for the ideal option, reporting the estimates and standard errors enclosed in brackets. The second column reports results for the standard post-selection method, specifically the point estimates resulting from the post-penalized quantile regression, reporting the standard errors as if there had been no model selection. The last column report the results for the double selection estimator (point estimate and standard error). Note that the Algorithm (ref) is applied sequentially to each of the variables. Similarly, in order to provide estimates and confidence intervals for all variables using the naive approach, if the covariate was not selected by the $\ell_1$-penalized quantile regression, it was included in the post-model selection quantile regression for that variable.

center[center omitted — 3,334 chars of source]

What we see is very interesting. First of all, let us compare the “ideal" option (column 1) and the naive post-selection (column 2). The Lasso selection method removes 16 out of 30 variables, many of which are highly significant, as judged by the “ideal" option. (To judge significance we use normal approximations and critical value of 3, which allows us to maintain $5\%$ significance level after testing up to 50 hypotheses). In particular, we see that the following highly significant variables were dropped by Lasso: mother's BMI, mother's age, twin status, birth orders one and two, and indicator of the other religion. The standard post-model selection inference then makes the assumption that these are true zeros, which leads us to misleading conclusions about these effects. The standard post-selection inference then proceeds to judge the significance of other variables, in some cases deviating sharply and significantly from the “ideal" benchmark. For example, there is a sharp disagreement on magnitudes of the impact of the birth order variables and the wealth variables (for “richer" and “richest" categories). Overall, for the naive post-selection, 8 out of 30 coefficients were more than 3 standard errors away from the coefficients of the “ideal" option.

We now proceed to comparing our proposed options to the “ideal" option. We see approximate agreement in terms of magnitude, signs of coefficients, and in standard errors. In few instances, for example, for the car ownership regressor, the disagreements in magnitude may appear large, but they become insignificant once we account for the standard errors.

The main conclusion from our study is that the standard/naive post-selection inference can give misleading results, confirming our expectations and confirming predictions of LeebPotscher2005. Moreover, the proposed inference procedure is able to deliver inference of high quality, which is very much in agreement with the “ideal" benchmark.

center[center omitted — 48 chars of source]
description• The supplemental appendix contains the proofs, additional discussions (variants, approximately sparse assumption) and technical results. (pdf)

\setcounter{section}{0} \setcounter{page}{1}

center[center omitted — 157 chars of source]
quoteThe supplemental appendix contains the proofs of the main results, additional discussions and technical results. Section 1 collects the notation. Section 2 has additional discussions on variants of the proposed methods, assumptions of approximately sparse functions, minimax efficiency, and the choices of bandwidth and penalty parameters and their implications for the growth of $s$, $p$ and $n$. Section 3 provides new results for $\ell_1$-penalized quantile regression with approximation errors, Lasso with estimated weights under (weaker) aggregated zero mean condition, and for the solution of the zero of the moment condition associated with the orthogonal score function. The proof of the main result of the main text is provided in Section 4. Section 5 collects auxiliary technical inequalities used in the proofs, Section 6 provides the proofs and technical lemmas for $\ell_1$-quantile regression. Section 7 provides proofs and technical lemmas for Lasso with estimated weights. Section 8 provides the proof for the orthogonal moment condition estimation problem. Finally Section 9 provided rates of convergence for the estimates of the conditional density function.

\spacingset{1.45}

Notation

In what follows, we work with triangular array data $\{ \omega_{i,n} : i=1,\dots,n; n=1,2,3,\dots\}$ where for each $n$, $\{ \omega_{i,n} ; i=1,\dots,n \}$ is defined on the probability space $(\Omega, \mathcal{S}, {\mathrm{P}}_n)$. Each $\omega_{i,n}= (y_{i,n}', z_{i,n}', d_{i,n}')'$ is a vector which are i.n.i.d., that is, independent across $i$ but not necessarily identically distributed. Hence all parameters that characterize the distribution of $\{\omega_{i,n} : i=1,\dots,n\}$ are implicitly indexed by ${\mathrm{P}}_n$ and thus by $n$. We omit this dependence from the notation for the sake of simplicity. We use ${\mathbb{E}_n}$ to abbreviate the notation $n^{-1}\sum_{i=1}^n$; for example, ${\mathbb{E}_n}[f] := {\mathbb{E}_n}[f(\omega_{i})] := n^{-1} \sum_{i=1}^n f(\omega_{i})$. We also use the following notation: $\bar {\mathrm{E}}[f] := {\mathrm{E}} \left [ {\mathbb{E}_n}[f] \right] = {\mathrm{E}} \left [ {\mathbb{E}_n}[f(\omega_{i})] \right ]= n^{-1} \sum_{i=1}^n {\mathrm{E}}[f(\omega_{i})]$. The $\ell_2$-norm is denoted by $\|\cdot\|$; the $\ell_0$-“norm” $\|\cdot\|_0$ denotes the number of non-zero components of a vector; and the $\ell_{\infty}$-norm $\| \cdot \|_{\infty}$ denotes the maximal absolute value in the components of a vector. Given a vector $\delta \in {\Bbb{R}}^p$, and a set of indices $T \subset \{1,\ldots,p\}$, we denote by $\delta_T \in {\Bbb{R}}^p$ the vector in which $\delta_{Tj} = \delta_j$ if $j\in T$, $\delta_{Tj}=0$ if $j \notin T$. We also denote by $\delta^{(k)}$ the vector with $k$ non-zero components corresponding to $k$ of the largest components of $\delta$ in absolute value. We use the notation $(a)_+ = \max\{a,0\}$, $a \vee b = \max\{ a, b\}$, and $a \wedge b = \min\{ a , b \}$. We also use the notation $a \lesssim b$ to denote $a \leqslant c b$ for some constant $c>0$ that does not depend on $n$; and $a\lesssim_P b$ to denote $a=O_P(b)$. For an event $E$, we say that $E$ wp $\to$ 1 when $E$ occurs with probability approaching one as $n$ grows. Given a $p$-vector $b$, we denote $\text{support}(b) = \{ j \in \{1,...,p\}: b_j \neq 0\} $. We also use $\rho_\tau(t) = t(\tau - 1\{t\leqslant 0\})$ and $\varphi_\tau(t_1,t_2) = (\tau - 1\{t_1 \leqslant t_2\} )$.

Define the minimal and maximal $m$-sparse eigenvalues of a symmetric positive semidefinite matrix $M$ as

equation[equation omitted — 288 chars of source]

For notational convenience we write $\tilde x_i = (d_i,x_i')'$, $\phi_{{\rm min}}(m):= \phi_{{\rm min}}(m)[{\mathbb{E}_n} [\tilde x_i\tilde x_i']]$, and $\phi_{{\rm max}}(m):= \phi_{{\rm max}}(m)[{\mathbb{E}_n} [\tilde x_i\tilde x_i']]$.

Additional Discussions

Variants of the Proposed Algorithms

There are several different ways to implement the sequence of steps underlying the two procedures outlined in Algorithms (ref) and (ref). The estimation of the control function $g_\tau$ can be done through other regularization methods like $\ell_1$-penalized quantile regression instead of the post-$\ell_1$-penalized quantile regression. The estimation of the error term $v$ in Step 2 can be carried out with Dantzig selector, square-root Lasso or the associated post-selection method could be used instead of Lasso or post-Lasso. Solving for the zero of the moment condition induced by the orthogonal score function can be substituted by a one-step correction from the $\ell_1$-penalized quantile regression estimator $\widehat\alpha_\tau$, namely, $\check \alpha_\tau = \widehat \alpha_\tau + ({\mathbb{E}_n}[\widehat v_i^2])^{-1} {\mathbb{E}_n}[(\tau - 1\{ y_i \leqslant \widehat\alpha_\tau d_i+x_i'\widehat\beta_\tau\})\widehat v_i]$.

Other variants can be constructing alternative orthogonal score functions. This can be achieved by changing the weights in the equation ((ref)) to alternative weights, say $\tilde f_i$, that lead to different errors terms $\tilde v$ that satisfies ${\mathrm{E}}[ \tilde f x \tilde v ] = 0$. Then the orthogonal score function is constructed as $\psi_i(\alpha) = (\tau - 1\{y_i\leqslant d_i\alpha + g_\tau(z_i)\}) \tilde v_i (\tilde f_i/f_i)$. It turns out that the choice $\tilde f_i = f_i$ minimizes the asymptotic variance of the estimator of $\alpha_\tau$ based upon the empirical analog of ((ref)), among all the score functions satisfying ((ref)). An example is to set $\tilde f_i = 1$ which would lead to $\tilde v_i=d_i-{\mathrm{E}}[d_i\mid z_i]$. Although such choice leads to a less efficient estimator, the estimation of ${\mathrm{E}}[d_i\mid z_i]$ and $f_i$ can be carried out separably which can lead to weaker regularity conditions.

Handling Approximately Sparse Functions

As discussed in Remark (ref), in order to handle approximately sparse models to represent $g_\tau$ in ((ref)) an approximate orthogonality condition is assumed, namely

equation[equation omitted — 96 chars of source]

In the literature such a condition has been (implicitly) used before. For example, ((ref)) holds if the function $g_\tau$ is an exactly sparse linear combination of the covariates so that all the approximation errors are exactly zero, namely, $r_{{\tau} i} = 0, i=1,\ldots,n$. An alternative assumption in the literature that implies ((ref)) is to have ${\mathrm{E}}[f_id_i\mid z_i]=f_i\{x_i'\theta_{\tau}+r_{{\theta \tau} i}\}$, where $\theta_{\tau}$ is sparse and $r_{{\theta \tau} i}$ is suitably small, which implies orthogonality to all functions of $z_i$ since we have ${\mathrm{E}}[f_iv_i \mid z_i ] = 0$.

The high-dimensional setting makes the condition ((ref)) less restrictive as $p$ grows. Our discussion is based on the assumption that the function $g_\tau$ belongs to a well behaved class of functions. For example, when $g_\tau$ belongs to a Sobolev space $\mathcal{S}(\alpha,L)$ for some $\alpha \geqslant 1$ and $L>0$ with respect to the basis $\{x_j=P_j(z), j\geqslant 1\}$. As in TsybakovBook, a Sobolev space of functions consists of functions $g(z)=\sum_{j=1}^\infty\theta_jP_j(z)$ whose Fourier coefficients $\theta$ satisfy $$\theta \in \Theta(\alpha, L) = \left\{ \theta \in \ell^2(\mathbb{N}):

array[array omitted — 113 chars of source]

\right\}.$$ More generally, we can consider functions in a $p$-Rearranged Sobolev space $\mathcal{RS}(\alpha,p,L)$ which allow permutations in the first $p$ components as in \cite{BCW2014pivotal}. Formally, the class of functions $g(z)=\sum_{j=1}^\infty\theta_jP_j(z) $ such that $$\theta \in \Theta^{R}(\alpha,p,L) = \left\{ \theta \in \ell^2(\mathbb{N}) : \sum_{j=1}^\infty|\theta_j| < \infty, \

array[array omitted — 193 chars of source]

\right\}.$$ It follows that $\mathcal{S}(\alpha,L) \subset \mathcal{RS}(\alpha,p,L)$ and $p$-Rearranged Sobolev space reduces substantially the dependence on the ordering of the basis.

Under mild conditions, it was shown in BCW2014pivotal that for functions in $\mathcal{RS}(\alpha,p,L)$ the rate-optimal choice for the size of the support of the oracle model obeys $s\lesssim n^{1/[2\alpha+1]}$. It follows that $$

array[array omitted — 162 chars of source]

$$ However, this bound cannot guarantee converge to zero faster than $\sqrt{n}$-rate to potentially imply (\ref{ExtraMomentDisc}). Fortunately, to establish relation (\ref{ExtraMomentDisc}) one can exploit orthogonality with respect all $p$ components of $x_i$. Indeed we have $$

array[array omitted — 586 chars of source]

$$ Therefore, condition (\ref{ExtraMomentDisc}) holds if $n=o(p^{2\alpha-1})$, in particular, for any $\alpha\geqslant 1$, $n = o(p)$ suffices.

Minimax Efficiency

In this section we make some connections to the (local) minimax efficiency analysis from the semiparametric efficiency analysis. In this section for the sake of exposition we assume that $(y_i,x_i,d_i)_{i=1}^n$ are i.i.d., sparse models, $r_{{\theta \tau} i}=r_{{\tau} i} = 0$, $i=1,\ldots,n$, and the median case ($\tau = .5$). Lee2003 derives an efficient score function for the partially linear median regression model: $$ S_i= 2\varphi_\tau(y_i,d_i\alpha_\tau+x_i'\beta_\tau) f_i [d_i - m_\tau^*(z)], $$ where $m_\tau^*(x_i)$ is given by $$ m^*_\tau(x_i) = \frac{{\mathrm{E}}[f_i^2 d_i|x_i]}{{\mathrm{E}}[f^2_i |x_i]}. $$ Using the assumption $m^*_\tau(x_i)=x_i'\theta^*_\tau$ , where $\|\theta^*_\tau\|_0 \leqslant s \ll n$ is sparse, we have that $$ S_i= 2\varphi_\tau(y_i, d_i\alpha_\tau+x_i'\beta_\tau) v_i^*, $$ where $v_i^* = f_id_i - f_im^*_\tau(x_i)$ would correspond to $v_i$ in ((ref)). It follows that the estimator based on $v^*_i$ is actually efficient in the minimax sense (see Theorem 18.4 in kosorok:book), and inference about $\alpha_\tau$ based on this estimator provides best minimax power against local alternatives (see Theorem 18.12 in kosorok:book).

The claim above is formal as long as, given a law $Q_n$, the least favorable submodels are permitted as deviations that lie within the overall model. Specifically, given a law $Q_n$, we shall need to allow for a certain neighborhood $\mathcal{Q}_n^{\delta}$ of $Q_n$ such that $Q_n \in \mathcal{Q}_n^{\delta} \subset \mathcal{Q}_n$, where the overall model $\mathcal{Q}_n$ is defined similarly as before, except now permitting heteroscedasticity (or we can keep homoscedasticity $f_i = f_{\epsilon}$ to maintain formality). To allow for this we consider a collection of models indexed by a parameter $t = (t_1, t_2)$:

eqnarray[eqnarray omitted — 224 chars of source]

where $\|\beta_\tau\|_0 \vee \| \theta^*_\tau\|_0 \leqslant s/2$ and conditions as in Section (ref) hold. The case with $t=0$ generates the model $Q_n$; by varying $t$ within $\delta$-ball, we generate models $ \mathcal{Q}_n^{\delta}$, containing the least favorable deviations. By Lee2003, the efficient score for the model given above is $S_i$, so we cannot have a better regular estimator than the estimator whose influence function is $J^{-1}S_i$, where $J = {\mathrm{E}}[S_i^2]$. Since our model $\mathcal{Q}_n$ contains $ \mathcal{Q}_n^{\delta}$, all the formal conclusions about (local minimax) optimality of our estimators hold from theorems cited above (using subsequence arguments to handle models changing with $n$). Our estimators are regular, since under $\mathcal{Q}_n^t$ with $t = ( O(1/\sqrt{n}), o(1))$, their first order asymptotics do not change, as a consequence of Theorems in Section (ref). (Though our theorems actually prove more than this.)

Choice of Bandwidth $h$ and Penalty Level $\lambda$

The proof of Theorem (ref) provides a detailed analysis for generic choice of bandwidth $h$ and the penalty level $\lambda$ in Step 2 under Condition D. Here we discuss two particular choices for $\lambda$, for $\gamma = 0.05/n$

center[center omitted — 179 chars of source]

The choice (i) for $\lambda$ leads to a sparser estimators by adjusting to the slower rate of convergence of $\widehat f_i$, see ((ref)). The choice (ii) for $\lambda$ corresponds to the (standard) choice of penalty level in the literature for Lasso. Indeed, we have the following sparsity guarantees for the $\tilde \theta_{\tau}$ under each choice

center[center omitted — 220 chars of source]

In addition to the requirements in Condition M, $(K_q^2s^2+s^3)\log^3(pn)\leqslant \delta_n n$ and $K_q^4s\log(pn)\log^3n\leqslant \delta_n n$, which are independent of $\lambda$ and $h$, we have that Condition D simplifies to

center[center omitted — 632 chars of source]

For example, using the choice of $\widehat f_i$ as in ((ref)) so that $\bar k = 4$, we have that the following choice growth conditions suffice for the conditions above:

center[center omitted — 254 chars of source]

Analysis of the Estimators

This section contains the main tools used in establishing the main inferential results. The high-level conditions here are intended to be applicable in a variety of settings and they are implied by the regularities conditions provided in the previous sections. The results provided here are of independent interest (e.g. properties of Lasso under estimated weights). We establish the inferential results ((ref)) and ((ref)) in Section (ref) under high level conditions. To verify these high-level conditions we need rates of convergence for the estimated residuals $\widehat v$ and the estimated confounding function $\widehat g_\tau (z) = x'\widehat \beta_\tau$ which are established in sections (ref) and (ref) respectively. The main design condition relies on the restricted eigenvalue proposed in BickelRitovTsybakov2009, namely for $\tilde x_i = [ d_i, x_i' ]'$

equation[equation omitted — 159 chars of source]

where $\mathbf{c} = (c+1)/(c-1)$ for the slack constant $c>1$, see BickelRitovTsybakov2009. When $\mathbf{c}$ is bounded, it is well known that $\kappa_\mathbf{c}$ is bounded away from zero provided sparse eigenvalues of order larger than $s$ are well behaved, see BickelRitovTsybakov2009.

$\ell_1$-Penalized Quantile Regression

In this section for a quantile index $u\in (0,1)$, we consider the equation

equation[equation omitted — 164 chars of source]

where we observe $\{(\tilde y_i,\tilde x_i):i=1,\ldots,n\}$, which are independent across $i$. To estimate $\eta_u$ we consider the $\ell_1$-penalized $u$-quantile regression estimate $$ \widehat\eta_u \in \arg\min_{\eta} {\mathbb{E}_n}[ \rho_u(\tilde y_i - \tilde x_i'\eta)] + \frac{\lambda_u}{n}\|\eta\|_1. $$ and the associated post-model selection estimates. That is, given an estimator $\bar\eta_{uj}$

equation[equation omitted — 196 chars of source]

We will be typically concerned with $\widehat\eta_u$, thresholded versions of the $\ell_1$-penalized quantile regression, defined as $\widehat\eta_{uj}^\mu = \widehat\eta_{uj}1\{ |\widehat\eta_{uj}|\geqslant \mu/{\mathbb{E}_n}[\tilde x_{ij}^2]^{1/2}\}$.

As established in BC-SparseQR for sparse models and in kato for approximately sparse models, under the event that

equation[equation omitted — 165 chars of source]

the estimator above achieves good theoretical guarantees under mild design conditions. Although $\eta_u$ is unknown, we can set $\lambda_u$ so that the event in ((ref)) holds with high probability. In particular, the pivotal rule proposed in BC-SparseQR and generalized in kato proposes to set $\lambda_u := c n\Lambda_u(1-\gamma\mid \tilde x)$ for $c> 1$ where

equation[equation omitted — 183 chars of source]

where $U_i\sim U(0,1)$ are independent random variables conditional on $\tilde x_i$, $i=1,\ldots,n$. This quantity can be easily approximated via simulations. Below we summarize the high level conditions we require.

{\bf Condition PQR.} Let $T_u = {\rm supp}(\eta_u)$ and normalize ${\mathbb{E}_n}[\tilde x_{ij}^2]=1$, $j=1,\ldots,p$. Assume that for some $s\geqslant 1$, $\|\eta_u\|_0\leqslant s$, $\|r_{ui}\|_{2,n}\leqslant C\sqrt{s\log(p)/n}$. Further, the conditional distribution function of $\epsilon_i$ is absolutely continuous with continuously differentiable density $f_\epsilon(\cdot\mid \tilde x_i, r_{ui})$ such that $0<\underline{f} \leqslant f_i \leqslant \sup_t f_{\epsilon_i\mid \tilde x_i, r_{ui}}(t \mid \tilde x_i, r_{ui}) \leqslant \bar f$, $\sup_{t} |f_{\epsilon_i\mid \tilde x_i, r_{ui}}'( t \mid \tilde x_i, r_{ui})|<\bar f'$ for fixed constants $\underline{f}$, $\bar f$ and $\bar f'$.

Condition PQR is implied by Condition AS. The conditions on the approximation error and near orthogonality conditions follows from choosing a model $\eta_u$ that optimally balance the bias/variance trade-off. The assumption on the conditional density is standard in the quantile regression literature even with fixed $p$ case developed in koenker:book or the case of $p$ increasing slower than $n$ studied in Belloni:Chern:Fern.

Next we present bounds on the prediction norm of the $\ell_1$-penalized quantile regression estimator.

lemma[Estimation Error of $\ell_1$-Penalized Quantile Regression] Under Condition PQR, setting $\lambda_u \geqslant cn\Lambda_u(1-\gamma\mid \tilde x)$, we have with probability $1-4\gamma$ for $n$ large enough $$\|\tilde x_i'(\widehat\eta_u - \eta_u)\|_{2,n} \lesssim N:=\frac{\lambda_u\sqrt{s}}{n\kappa_{2\mathbf{c}}}+ \frac{1}{\kappa_{2\mathbf{c}}}\sqrt{\frac{s\log(p/\gamma)}{n}}$$ and $\widehat\eta_u - \eta_u \in A_u:=\Delta_{2\mathbf{c}}\cup \{ v : \|\tilde x_i'v\|_{2,n}=N, \|v\|_1 \leqslant 8C\mathbf{c} s\log(p/\gamma)/\lambda_u\}$, provided that $$ \sup_{\bar\delta\in A_u} \frac{{\mathbb{E}_n}[|r_{ui}||\tilde x_i'\bar\delta|^2]}{{\mathbb{E}_n}[|\tilde x_i'\bar \delta|^2]} +N\sup_{\bar\delta\in A_u} \frac{{\mathbb{E}_n}[|\tilde x_i'\bar\delta|^3]}{{\mathbb{E}_n}[|\tilde x_i'\bar \delta|^2]^{3/2}} \to 0.$$

Lemma (ref) establishes the rate of convergence in the prediction norm for the $\ell_1$-penalized quantile regression estimator. Exact constants are derived in the proof. The extra growth condition required for identification is mild. For instance we typically have $\lambda_u \sim \sqrt{n\log (np)}$ and for many designs of interest we have $$\inf_{\delta\in\Delta_\mathbf{c}} \|\tilde x_i'\delta\|_{2,n}^{3}/{\mathbb{E}_n}[|\tilde x_i'\delta|^3]$$ bounded away from zero (see BC-SparseQR). For more general designs we have { $$\inf_{\delta\in A_u} \frac{\|\tilde x_i'\delta\|_{2,n}^{3}}{{\mathbb{E}_n}[|\tilde x_i'\delta|^3]}\geqslant \inf_{\delta\in A_u} \frac{\|\tilde x_i'\delta\|_{2,n}}{ \|\delta\|_1\max_{i\leqslant n}\|\tilde x_i\|_\infty} \geqslant \frac{1}{\max_{i\leqslant n}\|\tilde x_i\|_\infty} \left( \frac{\kappa_{2\mathbf{c}}}{\sqrt{s}(1+\mathbf{c})} \wedge \frac{\lambda_u N}{8C\mathbf{c} s\log(p/\gamma) }\right) .$$}

lemma[Estimation Error of Post-$\ell_1$-Penalized Quantile Regression] Assume Condition PQR holds, and that the Post-$\ell_1$-penalized quantile regression is based on an arbitrary vector $\widehat\eta_u$. Let $\bar r_u \geqslant \|r_{ui}\|_{2,n}$, $\widehat s_u \geqslant | {\rm supp}(\widehat\eta_u)|$ and $\widehat Q \geqslant {\mathbb{E}_n}[\rho_u(\tilde y_i - \tilde x_i'\widehat\eta_u)]-{\mathbb{E}_n}[\rho_u(\tilde y_i-\tilde x_i'\eta_u))]$ hold with probability $1-\gamma$. Then we have for $n$ large enough, with probability $1-\gamma-\varepsilon-o(1)$ $$\|\tilde x_i'(\widetilde \eta_u - \eta_u)\|_{2,n} \lesssim \widetilde N:= \sqrt{ \frac{ (\widehat s_u + s) \log(p/\varepsilon)}{n\phi_{{\rm min}}(\widehat s_u+s)}} + \bar f \bar r_u + \widehat Q^{1/2}$$ provided that $$ \sup_{\|\bar\delta\|_0\leqslant \widehat s_u+s} \frac{{\mathbb{E}_n}[|r_{ui}||\tilde x_i'\bar\delta|^2]}{{\mathbb{E}_n}[|\tilde x_i'\bar \delta|^2]} +\widetilde N\sup_{\|\bar\delta\|_0\leqslant \widehat s_u+s} \frac{{\mathbb{E}_n}[|\tilde x_i'\bar\delta|^3]}{{\mathbb{E}_n}[|\tilde x_i'\bar \delta|^2]^{3/2}} \to 0.$$

Lemma (ref) provides the rate of convergence in the prediction norm for the post model selection estimator despite of possible imperfect model selection. In the current nonparametric setting it is unlikely for the coefficients to exhibit a large separation from zero. The rates rely on the overall quality of the selected model by $\ell_1$-penalized quantile regression and the overall number of components $\widehat s_u$. Once again the extra growth condition required for identification is mild. For more general designs we have $$\inf_{\|\delta\|_0 \leqslant \widehat s_u + s} \frac{\|\tilde x_i'\delta\|_{2,n}^{3}}{{\mathbb{E}_n}[|\tilde x_i'\delta|^3]}\geqslant \inf_{\|\delta\|_0 \leqslant \widehat s_u + s} \frac{\|\tilde x_i'\delta\|_{2,n}}{ \|\delta\|_1\max_{i\leqslant n}\|\tilde x_i\|_\infty} \geqslant \frac{\sqrt{\phi_{{\rm min}}(\widehat s_u+s)}}{\sqrt{\widehat s_u + s}\max_{i\leqslant n}\|\tilde x_i\|_\infty}.$$

Lasso with Estimated Weights

In this section we consider the equation

equation[equation omitted — 153 chars of source]

where we observe $\{(d_i,z_i, x_i=X(z_i)):i=1,\ldots,n\}$, which are independent across $i$. We do not observe $\{f_i = f_\tau(d_i, z_i)\}_{i=1}^n$ directly and only estimates $\{\widehat f_i\}_{i=1}^n$ are available. Importantly, we only require that $ \bar {\mathrm{E}}[ f_iv_ix_i ] = 0$ and not ${\mathrm{E}}[f_ix_iv_i]=0$ for every $i=1,\ldots,n$. Also, we have that $T_{\theta \tau}={\rm supp}(\theta_\tau)$ is unknown but a sparsity condition holds, namely $|T_{\theta \tau}| \leqslant s$. To estimate $\theta_{\theta \tau}$ and $v_i$, we compute

equation[equation omitted — 285 chars of source]

where $\lambda$ and $\widehat \Gamma_\tau$ are the associated penalty level and loadings specified below. A difficulty is to account for the impact of estimated weights $\widehat f_i$ while also only using $\bar {\mathrm{E}}[ f_iv_ix_i ] = 0$.

We will establish bounds on the penalty parameter $\lambda$ so that with high probability the following regularization event occurs

equation[equation omitted — 128 chars of source]

As discussed in BickelRitovTsybakov2009,BC-PostLASSO,BCW-SqLASSO, the event above allows to exploit the restricted set condition $\|\widehat\theta_{\tau T^c_{\theta \tau}}\|_1\leqslant \tilde \mathbf{c} \|\widehat\theta_{\tau T_{\theta \tau}}-\theta_\tau\|_1$ for some $\tilde\mathbf{c} >1$. Thus rates of convergence for $\widehat \theta_\tau$ and $\widehat v_i$ defined on ((ref)) can be established based on the restricted eigenvalue $ \kappa_{\tilde\mathbf{c}}$ defined in ((ref)) with $\tilde x_i = x_i$.

However, the estimation error in the estimate $\widehat f_i$ of $f_i$ could slow the rates of convergence. The following are sufficient high-level conditions. In what follows $\underline c , \bar c, C, \underline{f}, \bar{f}$ are strictly positive constants independent of $n$.

{\bf Condition WL.} {\it For the model ((ref)) suppose that:\\ (i) for $s\geqslant 1$ we have $\|\theta_\tau\|_0\leqslant s$, $\Phi^{-1}(1-\gamma/2p) \leqslant \delta_n n^{1/6},$ \\ (ii) $\underline{f} \leqslant f_i \leqslant \bar{f}$, $\underline c \leqslant {\displaystyle \min_{j\leqslant p}} \{\bar {\mathrm{E}}[|f_ix_{ij}v_i-{\mathrm{E}}[f_ix_{ij}v_i]|^2]\}^{1/2} \leqslant {\displaystyle \max_{j\leqslant p} } \{\bar {\mathrm{E}}[|f_ix_{ij}v_i|^3]\}^{1/3} \leqslant C,$\\ (iii) with probability $1-\Delta_n$ we have ${\mathbb{E}_n}[ \widehat f_i^2 r_{{\theta \tau} i}^2] \leqslant c_r^2$, { $$

array[array omitted — 516 chars of source]

$$} (iv) $\ell \widehat \Gamma_{\tau 0} \leqslant \widehat \Gamma_{\tau} \leqslant u \widehat \Gamma_{\tau 0}$, for $\widehat \Gamma_{\tau 0jj}=\{{\mathbb{E}_n}[\widehat f_i^2 x_{ij}^2v_i^2]\}^{1/2}$, $1 - \delta_n\leqslant \ell \leqslant u \leqslant C$ with prob $1-\Delta_n$.}

remarkCondition WL(i) is a standard condition on the approximation error that yields the optimal bias variance trade-off (see BC-PostLASSO) and imposes a growth restriction on $p$ relative to $n$, in particular $\log p = o(n^{1/3})$. Condition WL(ii) imposes conditions on the conditional density function and mild moment conditions which are standard in quantile regression models even with fixed dimensions, see koenker:book. Condition WL(iii) requires high-level rates of convergence for the estimate $\widehat f_i$. Several primitive moment conditions imply first requirement in Condition WL(iii). These conditions allow the use of self-normalized moderate deviation theory to control heteroscedastic non-Gaussian errors similarly to BellChenChernHans:nonGauss where there are no estimated weights. Condition WL(iv) corresponds to the asymptotically valid penalty loading in BellChenChernHans:nonGauss which is satisfied by the proposed choice $\widehat \Gamma_\tau$ in ((ref)).

Next we present results on the performance of the estimators generated by Lasso with estimated weights. In what follows, $\widehat \kappa_\mathbf{c}$ is defined with $\widehat f_i x_i$ instead of $\tilde x_i$ in ((ref)) so that $\widehat\kappa_\mathbf{c} \geqslant \kappa_\mathbf{c} \min_{i\leqslant n} \widehat f_i$.

lemma[Rates of Convergence for Lasso] Under Condition WL and setting $\lambda \geqslant 2c'\sqrt{n}\Phi^{-1}(1-\gamma/2p)$ for $c'>c>1$, we have for $n$ large enough with probability $1-\gamma-o(1)$ $$\begin{array}{l} \displaystyle \|\widehat f_i x_i'(\widehat\theta_\tau - \theta_\tau)\|_{2,n} \leqslant 2\{ c_f+ c_r\} + \frac{\lambda\sqrt{s}}{n\widehat\kappa_{\tilde \mathbf{c}}}\left(u+\frac{1}{c}\right)\\ \displaystyle \|\widehat\theta_\tau-\theta_\tau\|_1 \leqslant 2 \frac{\sqrt{s}\{ c_f+ c_r\}}{\widehat \kappa_{2\tilde\mathbf{c}}} + \frac{\lambda s }{n\widehat\kappa_{\tilde \mathbf{c}}\widehat \kappa_{2\tilde\mathbf{c}}}\left(u+\frac{1}{c}\right) + \left(1+\frac{1}{2\tilde\mathbf{c}}\right)\frac{2c\|\widehat\Gamma_{\tau 0}^{-1}\|_\infty}{\ell c-1}\frac{n}{\lambda}\{ c_f+ c_r\}^2\end{array}$$ where $\tilde \mathbf{c} = \|\widehat \Gamma_{\tau 0}^{-1}\|_\infty\|\widehat \Gamma_{\tau 0}\|_\infty(uc+1)/(\ell c - 1)$

Lemma (ref) above establishes the rate of convergence for Lasso with estimated weights. This automatically leads to bounds on the estimated residuals $\widehat v_i$ obtained with Lasso through the identity

equation[equation omitted — 181 chars of source]

The Post-Lasso estimator applies the least squares estimator to the model selected by the Lasso estimator ((ref)), $$ \widetilde\theta_\tau \in \arg\min_{\theta\in {\Bbb{R}}^p} \ \left\{ \ {\mathbb{E}_n}[\widehat f_i^2(d_i - x_i'\theta)^2] \ : \ \theta_j = 0, \ \mbox{if} \ \widehat\theta_{\tau j} = 0 \ \right\}, \ \ \mbox{set} \ \tilde v_i = \widehat f_i (d_i - x_i'\widetilde\theta_\tau).$$ It aims to remove the bias towards zero induced by the $\ell_1$-penalty function which is used to select components. Sparsity properties of the Lasso estimator $\widehat\theta_\tau$ under estimated weights follows similarly to the standard Lasso analysis derived in BellChenChernHans:nonGauss. By combining such sparsity properties and the rates in the prediction norm we can establish rates for the post-model selection estimator under estimated weights. The following result summarizes the properties of the Post-Lasso estimator.

lemma[Model Selection Properties of Lasso and Properties of Post-Lasso] Suppose that Condition WL holds, and $\kappa' \leqslant \phi_{{\rm min}}(\{s + \frac{n^2}{\lambda^2}\{c_f^2 + c_r^2\}\}/\delta_n) \leqslant \phi_{{\rm max}}(\{s + \frac{n^2}{\lambda^2}\{c_f^2 + c_r^2\}\}/\delta_n) \leqslant \kappa''$ for some positive and bounded constants $\kappa', \kappa''$. Then the data-dependent model $\widehat T_{\theta \tau}$ selected by the Lasso estimator with $\lambda \geqslant 2c'\sqrt{n}\Phi^{-1}(1-\gamma/2p)$ for $c'>c>1$, satisfies with probability $1-\gamma-o(1)$: \begin{equation} \|\widetilde \theta_\tau \|_0 = | \widehat T_{\theta \tau} | \lesssim s + \frac{n^2}{\lambda^2}\{ c_f^2 + c_r^2\} \end{equation} Moreover, the corresponding Post-Lasso estimator obeys with probability $1-\gamma-o(1)$ $$ \| x_i'(\widetilde \theta_\tau -\theta_\tau)\|_{2,n} \lesssim_P c_f + c_r + \sqrt{\frac{| \widehat T_{\theta \tau} |\log (p \vee n)}{n}} + \frac{\lambda\sqrt{s}}{n\kappa_\mathbf{c}}.$$

Moment Condition based on Orthogonal Score Function

Next we turn to analyze the estimator $\check \alpha_\tau$ obtained based on the orthogonal moment condition. In this section we assume that $$ L_n(\check \alpha_\tau) \leqslant \min_{\alpha \in \mathcal{A}_\tau} L_n(\alpha) + \delta_n n^{-1}.$$ This setting is related to the instrumental quantile regression method proposed in ch:iqrWeakId. However, in this application we need to account for the estimation of the noise $v$ that acts as the instrument which is known in the setting in ch:iqrWeakId. Condition IQR below suffices to make the impact of the estimation of instruments negligible to the first order asymptotics of the estimator $\check \alpha_\tau$. Primitive conditions that imply Condition IQR are provided and discussed in the main text.

Let $\{(y_i,d_i,z_i):i=1,\ldots,n\}$ be independent observations satisfying

equation[equation omitted — 255 chars of source]

Letting $\mathcal{D}\times\mathcal{Z}$ denote the domain of the random variables $(d,z)$, for $\tilde h = (\tilde g, \tilde {{\iota}})$, where $\tilde g$ is a function of variable $z$, and the instrument $\tilde {{\iota}}$ is a function that maps $(d,x)\mapsto \tilde {{\iota}}(d,z)$ we write $$

array[array omitted — 277 chars of source]

$$ We denote $h_0=(g_\tau,{{\iota}}_0)$ where ${{\iota}}_{0i}:=v_i=f_i(d_i-x_i'\theta_{0\tau})$. For some sequences $\delta_n\to 0$ and $\Delta_n\to 0$, we let $\overline{\mathcal{F}}$ denote a set of functions such that each element $\tilde h = (\tilde g,\tilde {{\iota}})\in \overline{\mathcal{F}}$ satisfies

equation[equation omitted — 531 chars of source]

and with probability $1-\Delta_n$ we have

equation[equation omitted — 307 chars of source]

We assume that the estimated functions $\widehat g$ and $\widehat {{\iota}}$ satisfy the following condition.

{\bf Condition IQR.} {\textit Let $\{(y_i,d_i,z_i):i=1,\ldots,n\}$ be random variables independent across $i$ satisfying ((ref)). Suppose that there are positive constants $0<c\leqslant C<\infty$ such that:\\ (i) $f_{y_i\mid d_i,z_i}(y\mid d_i,z_i) \leqslant \bar f$, $f_{y_i\mid d_i,z_i}'(y\mid d_i,z_i) \leqslant \bar f'$; $c \leqslant |\bar {\mathrm{E}}[f_id_i{{\iota}}_{0i}]|$, and $\bar {\mathrm{E}}[ {{\iota}}_{0i}^4 ] + \bar {\mathrm{E}}[ d_i^4 ] \leqslant C$;\\ (ii) $\{ \alpha : |\alpha-\alpha_\tau| \leqslant n^{-1/2}/\delta_n\} \subset \mathcal{A}_\tau$, where $\mathcal{A}_\tau$ is a (possibly random) compact interval; \\ (iii) with probability at least $1-\Delta_n$ the estimated functions $\widehat h = (\widehat g,\widehat {{\iota}}) \in \overline{\mathcal{F}}$ and

equation[equation omitted — 209 chars of source]

(iv) with probability at least $1-\Delta_n$, the estimated functions $\widehat h = (\widehat g,\widehat {{\iota}})$ satisfy $$\|\widehat {{\iota}}_i - {{\iota}}_{0i}\|_{2,n} \leqslant \delta_n \ \ \mbox{and} \ \ \|1\{|\epsilon_i|\leqslant |d_i(\alpha_\tau-\check\alpha_\tau)+g_{\tau i}-\widehat g_i|\}\|_{2,n}\leqslant \delta_n^2.$$}

lemmaUnder Condition IQR(i,ii,iii) we have $$ \bar\sigma_n^{-1}\sqrt{n}(\check \alpha_\tau-\alpha_\tau) = \mathbb{U}_n(\tau)+o_P(1), \ \ \mathbb{U}_n(\tau)\rightsquigarrow N(0,1)$$ where $\bar \sigma^2_n = \bar {\mathrm{E}}[f_id_i{{\iota}}_{0i}]^{-1}\bar {\mathrm{E}}[\tau(1-\tau){{\iota}}_{0i}^2]\bar {\mathrm{E}}[f_id_i{{\iota}}_{0i}]^{-1}$ and $$\mathbb{U}_n(\tau)=\{\bar {\mathrm{E}}[\psi_{\alpha_\tau,h_0}^2(y_i,d_i,z_i)]\}^{-1/2}\sqrt{n}{\mathbb{E}_n}[\psi_{\alpha_\tau,h_0}(y_i,d_i,z_i)].$$ Moreover, IQR(iv) also holds we have $$nL_n(\alpha_\tau) = \mathbb{U}_n(\tau)^2+o_P(1), \ \ \mathbb{U}_n(\tau)^2\rightsquigarrow \chi^2(1)$$ and the variance estimator is consistent, namely { $${\mathbb{E}_n}[\widehat f_id_i\widehat {{\iota}}_i]^{-1}{\mathbb{E}_n}[(\tau-1\{y_i\leqslant \widehat g_i+d_i\check\alpha_\tau\})^2\widehat {{\iota}}_i^2]{\mathbb{E}_n}[\widehat f_i d_i \widehat {{\iota}}_i]^{-1}\to_P\bar {\mathrm{E}}[f_id_i{{\iota}}_{0i}]^{-1}\bar {\mathrm{E}}[\tau(1-\tau){{\iota}}_{0i}^2]\bar {\mathrm{E}}[f_id_i{{\iota}}_{0i}]^{-1}.$$}

Proofs for Section (ref) of Main Text (Main Result)

{\bf Proof.}{\bf \ (Proof of Theorem (ref))} The first-order equivalence between the two estimators follows from establishing the same linear representation for each estimator. In Part 1 of the proof we consider the orthogonal score estimator. In Part 2 we consider the double selection estimator.

{\rm Part 1. Orthogonal score estimator.} We will verify Condition IQR and the result follows by Lemma (ref) applied with ${{\iota}}_{0i}=v_i=f_i(d_i-x_i'\theta_{0\tau})$ and noting that $1\{y_i \leqslant d_i\alpha_\tau + g_\tau(z_i)\} = 1\{U_i \leqslant \tau\}$ for some uniform $(0,1)$ random variable (independent of $\{d_i,z_i\}$) by the definition of the conditional quantile function.

Condition IQR(i) requires conditions on the probability density function that are assumed in Condition AS(iii). The fourth moment conditions are implied by Condition M(i) using $\xi=(1,0')'$ and $\xi=(1,-\theta_{0\tau}')'$, since $\bar {\mathrm{E}}[v_i^4]\leqslant \bar f^4\bar {\mathrm{E}}[ \{ (d_i,x_i')\xi\}^4] \leqslant C'(1+\|\theta_{0\tau}\|)^4$, and $\|\theta_{0\tau}\|\leqslant C$ assumed in Condition AS(i). Finally, by relation ((ref)), namely ${\mathrm{E}}[f_ix_iv_i]=0$, we have $\bar {\mathrm{E}}[f_id_iv_i] = \bar {\mathrm{E}}[v_i^2] \geqslant \underline f \bar {\mathrm{E}}[ (d_i-x_i'\theta_{0\tau})^2] \geqslant c\underline f\|(1,\theta_{0\tau}')'\|^2$ by Condition AS(iii) and Condition M(i).

Next we will construct the estimate for the orthogonal score function which are based on post-$\ell_1$-penalized quantile regression and post-Lasso with estimated conditional density function. We will show that with probability $1-o(1)$ the estimated nuisance parameters belong to $\overline{\mathcal{F}}$.

To establish the rates of convergence for $\widetilde\beta_\tau$, the post-$\ell_1$-penalized quantile regression based on the thresholded estimator $\widehat\beta_\tau^{\lambda_\tau}$, we proceed to provide rates of convergence for the $\ell_1$-penalized quantile regression estimator $\widehat\beta_\tau$. We will apply Lemma (ref) with $\gamma=1/n$. Condition PQR holds by Condition AS with probability $1-o(1)$ using Markov inequality and $\bar {\mathrm{E}}[r_{{\tau} i}^2]\leqslant s/n$.

Since population eigenvalues are bounded above and bounded away from zero by Condition M(i), by Lemma (ref) (where $\bar \delta_n\to 0$ under $K_q^2 Cs \log^2(1+Cs) \log (pn) \log n =o(n)$), we have that sparse eigenvalues of order $\ell_n s$ are bounded away from zero and from above with probability $1-o(1)$ for some slowly increasing function $\ell_n$. In turn, the restricted eigenvalue $\kappa_{2\mathbf{c}}$ is also bounded away from zero for bounded $\mathbf{c}$ and $n$ sufficiently large. Since $\lambda_\tau \lesssim \sqrt{n\log(p\vee n)}$, we will take $N = C\sqrt{s\log(n\vee p)/n}$ in Lemma (ref). To establish the side conditions note that

equation*[equation* omitted — 521 chars of source]

which implies that $N{\displaystyle \sup_{\delta \in A_\tau}}\frac{{\mathbb{E}_n}[|\tilde x_i'\delta|^3]}{\|\tilde x_i'\delta\|_{2,n}^{3}} \to 0$ with probability $1-o(1)$ under $K_q^2s^2\log^2(p\vee n) \leqslant \delta_n n$. Moreover, the second part of the side condition

equation*[equation* omitted — 800 chars of source]

Under $K_q^2s^2\log^2(p\vee n) \leqslant \delta_n n$, the side condition holds with probability $1-o(1)$. Therefore, by Lemma (ref) we have $\|\tilde x_i'\widehat \beta_\tau - \beta_\tau \|_{2,n} \lesssim \sqrt{ s \log(pn)/n }$ and $\|\widehat \beta_\tau - \beta_\tau \|_1 \lesssim s\sqrt{\log(pn)/n }$ with probability $1-o(1)$.

With the same probability, by Lemma (ref) with $\mu = \lambda_\tau/n$, since $\phi_{{\rm max}}(Cs)$ is uniformly bounded with probability $1-o(1)$, we have that the thresholded estimator $\widehat\beta^\mu_\tau$ satisfies the following bounds with probability $1-o(1)$: $\|\widehat \beta_\tau^\mu - \beta_\tau \|_1 \lesssim s\sqrt{\log(pn)/n }$, $\|\widehat\beta_\tau^\mu\|_0 \lesssim s$ and $\|\tilde x_i'(\widehat\beta^\mu_\tau - \beta_\tau)\|_{2,n} \lesssim \sqrt{ s \log(pn)/n }$. We use the support of $\widehat\beta^\mu_\tau$ as to construct the refitted estimator $\widetilde\beta_\tau$.

We will apply Lemma (ref). By Lemma (ref) and the rate of $\widehat\beta^\mu_\tau$, we can take $\widehat Q = C s\log(pn)/n$. Since sparse eigenvalues of order $Cs$ are bounded away from zero, we will use $\tilde N = C\sqrt{s\log(pn)/n}$ and $\varepsilon = 1/n$. Therefore, $\|x_i'(\widetilde\beta_\tau-\beta_\tau)\|_{2,n} \lesssim \sqrt{s\log(np)/n}$ with probability $1-o(1)$ provided the side conditions of the lemma hold. To verify the side conditions note that

equation*[equation* omitted — 576 chars of source]
equation*[equation* omitted — 536 chars of source]

and the side condition holds with probability $1-o(1)$ under $K_q^2 s \log^2(p\vee n) \leqslant \delta_n n$.

Next we proceed to construct the estimator for $v_i$. Note that by the same arguments above, under Condition D, we have the same rates of convergence for the post-selection (after truncation) estimators $(\widetilde \alpha_u,\widetilde \beta_u)$, $u\in \mathcal{U}$, $\|\widetilde\beta_u\|_0\lesssim C$ and $\|(\widetilde \alpha_u,\widetilde \beta_u)-(\widetilde \alpha_u,\widetilde \beta_u)\|\lesssim \sqrt{s\log (pn)/n}$, to estimate the conditional density function via ((ref)) or ((ref)). Thus with probability $1-o(1)$

equation[equation omitted — 200 chars of source]

where $\bar k$ depends on the estimator. (See relation ((ref)) below.) Note that the last relation implies that $\max_{i\leqslant n}\widehat f_i \leqslant 2\bar f$ is automatically satisfied with probability $1-o(1)$ for $n$ large enough. Let $\mathcal{U}$ denote the (finite) set of quantile indices used in the calculation of $\widehat f_i$.

Next we verify Condition WL to invoke Lemmas (ref) and (ref) with $c_r = C\sqrt{s\log(pn)/n}$ and $c_f= C\left( \{1/h\}\sqrt{s\log (n\vee p)/n}+h^{\bar k}\right)$. The sparsity condition in Condition WL(i) is implied by Condition AS(ii) and $\Phi^{-1}(1-\gamma/2p) \leqslant \delta_n n^{1/6}$ is implied by $\log (1/\gamma) \lesssim \log(p\vee n)$ and $\log^{3} p \leqslant \delta_n n$ from Condition M(iii). Condition WL(ii) follows from the assumption on the density function in Condition AS(iii) and the moment conditions in Condition M(i).

The first requirement in Condition WL(iii) holds with $c_r^2 \lesssim s\log(pn)/n$ since $$ {\mathbb{E}_n}[\widehat f_i^2r_{{\theta \tau} i}^2] \leqslant \max_{i\leqslant n}\widehat f_i {\mathbb{E}_n}[r_{{\theta \tau} i}^2] \lesssim \bar f s\log(pn)/n $$ from $\max_{i\leqslant n} \widehat f_i \leqslant C$ holding with probability $1-o(1)$, and Markov's inequality under $\bar {\mathrm{E}}[r_{{\theta \tau} i}^2]\leqslant Cs/n$ by Condition AS(ii).

The second requirement, $\max_{j\leqslant p}|({\mathbb{E}_n}-\bar {\mathrm{E}})[f_i^2x_{ij}^2v_i^2]|+|({\mathbb{E}_n}-\bar {\mathrm{E}})[\{f_ix_{ij}v_i-{\mathrm{E}}[f_ix_{ij}v_i]\}^2]| \leqslant \delta_n $ with probability $1-o(1)$, is implied by Lemma (ref) with $k=1$ and $K=\{{\mathrm{E}}[\max_{i\leqslant n} \|f_i x_iv_i\|_\infty^2]\}^{1/2} \leqslant \bar f \{{\mathrm{E}}[\max_{i\leqslant n} \|(x_i,v_i)\|_\infty^4]\}^{1/2}\leqslant \bar fK_q^{2}$, under $K_q^4 \log(pn) \log n \leqslant \delta_n n$.

The third requirement of Condition WL(iii) follows from uniform consistency in ((ref)) and the second part of Condition WL(ii) since $$ \max_{j\leqslant p}{\mathbb{E}_n}[(\widehat f_i - f_i)^2x_{ij}^2v_i^2] \leqslant \max_{i\leqslant n}\frac{|\widehat f_i-f_i|^2}{f_i^2} \left\{ \max_{j\leqslant p}({\mathbb{E}_n}-\bar {\mathrm{E}})[f_i^2x_{ij}^2v_i^2] + \max_{j\leqslant p}\bar {\mathrm{E}}[f_i^2x_{ij}^2v_i^2] \right\} \lesssim \delta_n $$ with probability $1-o(1)$ by ((ref)), the second requirement, and the bounded fourth moment assumption in Condition M(i).

To show the fourth part of Condition WL(iii), because both $\widehat f_i$ and $f_i$ are bounded away from zero and from above with probability $1-o(1)$, uniformly over $u\in \mathcal{U}$ (the finite set of quantile indices used to estimate the density), and $\widehat f_i^2-f_i^2 = (\widehat f_i-f_i)(\widehat f_i+f_i)$, it follows that with probability $1-o(1)$ $$ {\mathbb{E}_n}[(\widehat f_i^2-f_i^2)^2v_i^2/f_i^2] + {\mathbb{E}_n}[(\widehat f_i^2-f_i^2)^2v_i^2/\{\widehat f_i^2f_i^2\}] \lesssim {\mathbb{E}_n}[(\widehat f_i-f_i)^2v_i^2/f_i^2].$$ Next let $\delta_{u}=(\widetilde\alpha_u-\alpha_u,\widetilde \beta_{u}' - \beta_{u}')'$ where the estimators satisfy Condition D. We have that with probability $1-o(1)$

equation[equation omitted — 274 chars of source]

The following relations hold for all $u\in \mathcal{U}$ $$

array[array omitted — 517 chars of source]

$$ where we have $\|\delta_{u}\|\lesssim \sqrt{s\log (n\vee p) / n}$ and $\|\delta_{u}\|_0 \lesssim s$ with with probability $1-o(1)$. Then we apply Lemma \ref{thm:RV34} with $X_i=v_i\tilde x_i$. Thus, we can take $K=\{{\mathrm{E}}[\max_{i\leqslant n}\|X_i\|_\infty^2]\}^{1/2} \leqslant \{{\mathrm{E}}[\max_{i\leqslant n}\|(v_i,\tilde x_i')\|_\infty^4]\}^{1/2}\leqslant K_q^2$, and $\bar {\mathrm{E}}[(\delta'X_i)^2] \leqslant \bar {\mathrm{E}}[v_i^2(\tilde x_i'\delta)^2]\leqslant C\|\delta\|^2$ by the fourth moment assumption in Condition M(i) and $\|\theta_{0\tau}\|\leqslant C$. Therefore, {\small $$

array[array omitted — 297 chars of source]

$$} Under $K_q^4 s \log^{3}n \log (p\vee n) \leqslant \delta_n n$, with probability $1-o(1)$ we have $$ c_f^2 \lesssim \frac{s\log (n\vee p)}{h^2n} + h^{2{\bar k}}.$$

Condition WL(iv) pertains to the penalty loadings which are estimated iteratively. In the first iteration we have that the loadings are constant across components as $\varpi := \max_{i\leqslant n} \widehat f_i \max_{i\leqslant n}\|x_i\|_\infty \{{\mathbb{E}_n}[\widehat f_i^2 d_i^2]\}^2$. Thus the solution of the optimization problem is the same if we use penalty parameters $\tilde \lambda$ and $\widetilde \Gamma$ defined as $\tilde\lambda := \lambda \varpi/\max_{j\leqslant p} \{{\mathbb{E}_n}[ \widehat f_i^2x_{ij}^2v_i^2]\}^{1/2}$ and $\widetilde \Gamma_{jj} =\max_{j\leqslant p} \{{\mathbb{E}_n}[ \widehat f_i^2x_{ij}^2v_i^2]\}^{1/2}$. By construction $\widehat\Gamma_{0\tau jj} \leqslant \widetilde \Gamma_{jj} \leqslant C \widehat\Gamma_{0\tau jj}$ for some bounded $C$ with probability $1-o(1)$ as $\widehat\Gamma_{0\tau jj}$ are bounded away from zero and from above with probability $1-o(1)$. Since Condition WL holds for $(\tilde \lambda,\widetilde \Gamma)$, and $\tilde \lambda \lesssim \lambda \max_{i\leqslant n}\|x_i\|_\infty$, by Lemma (ref) we have with probability $1-o(1)$ $$

array[array omitted — 223 chars of source]

$$ The iterative choice of of penalty loadings satisfies $$

array[array omitted — 666 chars of source]

$$ uniformly in $j\leqslant p$. Thus the iterated penalty loadings are uniformly consistent with probability $1-o(1)$ and also satisfy Condition WL(iv) under $h^{-2}K_q^2s\log(pn)\leqslant \delta_n n$ and $K_q^4s\log(pn)\leqslant \delta_nn$. Therefore, in the subsequent iterations, by Lemma \ref{Thm:RateEstimatedLasso} and Lemma \ref{Thm:EstLassoPostRate} we have that the post-Lasso estimator satisfies with probability $1-o(1)$ $$

array[array omitted — 442 chars of source]

$$ where we used that $\phi_{{\rm max}}(\widetilde s_{\theta \tau}/\delta_n)\leqslant C$ implied by Condition D and Lemma \ref{thm:RV34}, and that $\lambda \geqslant \sqrt{n}\Phi^{-1}(1-\gamma/2p)\sim \sqrt{n\log (pn)}$ so that $ \sqrt{\widetilde s_{\theta \tau} \log (pn)/n} \lesssim (1/h)\sqrt{s\log (pn)/n} + h^{\bar k}.$

Next we construct a class of functions that satisfies the conditions required for $\overline{\mathcal{F}}$ that is used in Lemma (ref). Define the following class of functions.

equation[equation omitted — 554 chars of source]

where $\tilde f_i = f(d_i,z_i, \{\tilde \eta_u : u\in\mathcal{U}\}), \tilde{\tilde{f}}_i = f(d_i,z_i, \{\tilde{\tilde{\eta}}_u : u\in\mathcal{U}\}) \in \mathcal{J}$ are functions that satisfy

equation[equation omitted — 337 chars of source]

In particualr, taking $\tilde f_i := \widehat f_i \wedge 2\bar f$ where $\widehat f_i$ is defined in ((ref)) or ((ref)). (Due to uniform consistency ((ref)) the minimum with $2\bar f$ is not binding for $n$ large.) Therefore we have

equation[equation omitted — 262 chars of source]

We define the function class $\overline{\mathcal{F}}$ as $$ \overline{\mathcal{F}} = \{ ( \tilde g, \tilde v := \tilde f ( d- \tilde m ) ) : \tilde g \in \mathcal{G}, \tilde m \in \mathcal{M}, \tilde f \in \mathcal{J}\}$$ The rates of convergence and sparsity guarantees for $\widetilde\beta_\tau$ and $\widetilde\theta_\tau$, and Condition D implies that the estimates of the nuisance parameters $\widehat g_i := x_i'\widetilde\beta_\tau$ and $\widehat v_i=\widehat f_i(d_i-x_i'\widetilde\theta_{\tau})$ belongs to the proposed $ \overline{\mathcal{F}}$ with probability $1-o(1)$.

We will proceed to verify relations ((ref)) and ((ref)) hold under Condition AS, M and D. We begin with ((ref)). We have $$

array[array omitted — 856 chars of source]

$$ since $\|(1,-\theta)\|\leqslant C$ and $f_i\vee \tilde f_i \leqslant 2\bar f$, $\bar {\mathrm{E}}[(\tilde x_i'\xi)^4]\leqslant C\|\xi\|^4$, $\bar {\mathrm{E}}[(\tilde x_i'\xi)^2r_{\tau i}^2]\leqslant C\|\xi\|^2\bar {\mathrm{E}}[r_{\tau i}^2]$ and $\bar {\mathrm{E}}[r_{\tau i}^2]\lesssim s/n$ which hold by Conditions AS and M. Similarly we have $$

array[array omitted — 326 chars of source]

$$ as $|\tilde f_i - f_i| \leqslant 2\bar f$. Moreover we have that $$

array[array omitted — 488 chars of source]

$$ under (\ref{Eq:ExpDensity}), $h^{-2}s\log(pn) \leqslant \delta_n^2 n$, $h \leqslant \delta_n$, and $\lambda\sqrt{s} \leqslant \delta_n n$. The next relation follows from $$

array[array omitted — 632 chars of source]

$$ under $h^{-2}s^2\log^2(pn) \leqslant \delta_n^2 n$ and $h^{\bar k}\sqrt{s\log(pn)} \leqslant \delta_n$.

Finally, since $|\bar {\mathrm{E}}[f_iv_ir_{{\tau} i}]|\leqslant \delta_n n^{-1/2}$ from Condition M and $\bar {\mathrm{E}}[f_iv_ix_i]=0$ from ((ref)), we have $$ |\bar {\mathrm{E}}[f_iv_i\{\tilde g_i - g_{\tau i}\}]| \leqslant |\bar {\mathrm{E}}[f_iv_ix_i'(\tilde\beta - \beta_\tau)] |+|\bar {\mathrm{E}}[f_iv_ir_{\tau i}]| \leqslant \delta_n n^{-1/2}. $$

Next we verify relation ((ref)). By triangle inequality we have

equation[equation omitted — 784 chars of source]

Consider the first term of the right hand side in ((ref)). Note that $\mathcal{F}_1 := \{ (1\{y_i\leqslant d_i\alpha+g_{\tau i}\} - 1\{y_i\leqslant d_i\alpha+\tilde g_i\})v_i : \tilde g \in \mathcal{G}, |\alpha-\alpha_\tau| \leqslant \delta_n^2 \} \subset \mathcal{F}_{1a}-\mathcal{F}_{1b}$ where $\mathcal{F}_{1a}:=1\{y_i\leqslant d_i\alpha+g_{\tau i}\}v_i : |\alpha-\alpha_\tau| \leqslant \delta_n^2 \}$ is the product of a VC class of dimension 1 with the random variable $v$, and $\mathcal{F}_{1b}:=\{1\{y_i\leqslant d_i\alpha+\tilde g_{i}\}v_i : \tilde g \in \mathcal{G}, |\alpha-\alpha_\tau| \leqslant \delta_n^2 \}$ is the product of $v$ with the union of $\binom{p}{Cs}$ VC classes of dimension $Cs$. Therefore, its entropy number satisfies $N(\epsilon\|F_1\|_{Q,2},\mathcal{F}_1,\|\cdot\|_{Q,2}) \leqslant (A/\epsilon)^{C's}$ where we can take the envelope $F_1(y,d,x)=2|v|$. Since for any function $m \in \mathcal{F}_1$ we have $$

array[array omitted — 343 chars of source]

$$ from the conditional density function being uniformly bounded and the bounded fourth moment assumption in Condition M(i), by Lemma \ref{lemma:CCK} with $\sigma:= C\{s\log(pn)/n\}^{1/4}$, $a=pn$, $\|F_1\|_{P,2} = \bar {\mathrm{E}}[v_i^2]^{1/2} \leqslant C$, and $\|M\|_{P,2}\leqslant K_q$ we have with probability $1-o(1)$ $$\sup_{m\in \mathcal{F}_1}|({\mathbb{E}_n}-\bar {\mathrm{E}})[m]| \lesssim \frac{\{s\log(pn)/n\}^{1/4}}{n^{1/2}}\sqrt{s\log(pn)}+ n^{-1} s K_q \log(pn) \lesssim \delta_n n^{-1/2}$$ under $s^3\log^3(pn) \leqslant \delta_n^4 n$ and $K_q^2s^2\log^2(pn)\leqslant \delta_n^2 n$.

Next consider the second term of the right hand side in ((ref)). Note that $\mathcal{F}_2 := \{ (\tau - 1\{y_i\leqslant d_i\alpha+\tilde g_{i}\})(\tilde v_i-v_i) : (\tilde g,\tilde v) \in \overline{\mathcal{F}}, |\alpha-\alpha_\tau| \leqslant \delta_n^2 \} \subseteq \mathcal{F}_{2a} \cdot \mathcal{F}_{2b}$ where $\mathcal{F}_{2a} := \{ (\tau - 1\{y_i\leqslant d_i\alpha+\tilde g_{i}\}): \tilde g \in \mathcal{G}\}$ is a constant minus the union of $\binom{p}{Cs}$ VC classes of dimension $Cs$, and $\mathcal{F}_{2b} := \{ \tilde v_i - v_i : (\tilde g, \tilde v) \in \overline{\mathcal{F}} \}$. Note that standard entropy calculations yield $$ N(\epsilon \|F_{2a}F_{2b}\|_{Q,2}, \mathcal{F}_2 , \|\cdot\|_{Q,2}) \leqslant N(\mbox{$\frac{\epsilon}{2}$} \|F_{2a}\|_{Q,2}, \mathcal{F}_{2a} , \|\cdot\|_{Q,2}) \ N(\mbox{$\frac{\epsilon}{2}$}\|F_{2b}\|_{Q,2}, \mathcal{F}_{2b} , \|\cdot\|_{Q,2})$$ Furthermore, since $\tilde v_i - v_i = (\tilde f_i-f_i)v_i/f_i + \tilde f_ix_i'(\theta_{\tau 0}-\tilde \theta)$, we have $$\mathcal{F}_{2b} \subset \mathcal{F}_{2b}' + \mathcal{F}_{2b}'' := ( \mathcal{J} - \{f_i\} ) \cdot \{ v_i/f_i \} + \mathcal{J} \cdot (\{x_i'\theta_{\tau0}\} - \mathcal{M})$$ and $\mathcal{F}_2 \subset \mathcal{F}_{2a} \cdot \mathcal{F}_{2b}' + \mathcal{F}_{2a} \cdot \mathcal{F}_{2b}''$. By ((ref)), a covering for $\mathcal{J}$ can be constructed based on a covering for $\mathcal{B}_u:=\{ \tilde \eta_u : \|\tilde\eta_u\|_0 \leqslant Cs, \|\tilde\eta_u - \eta_u\|\leqslant C\sqrt{s\log(pn)/n}\}$ which is the union of $\binom{p}{Cs}$ sparse balls of dimension $Cs$. It follows that for the envelope $F_J:= 2\bar f \vee K_q^{-1}\|\tilde x_i\|_{\infty}$, we have $$ N(\epsilon \|F_J\|_{Q,2}, \mathcal{J},\|\cdot\|_{Q,2}) \leqslant \prod_{u\in \mathcal{U}} N\left( \mbox{$\frac{\epsilon h/4\bar f}{|\mathcal{U}|K_q \sqrt{2Cs}}$}, \mathcal{B}_u,\|\cdot\|\right)$$ For any $m \in \mathcal{F}_{2a}\cdot \mathcal{F}_{2b}'$ we have $$

array[array omitted — 315 chars of source]

$$ Therefore, by Lemma \ref{lemma:CCK} with $\sigma:= \frac{1}{h}\sqrt{s\log(pn)/n}+h^{\bar k}$, $a=pn$, $\|F_2'\|_{P,2} = \|(2\bar f\vee K_q^{-1}\|\tilde x_i\|_\infty)v_i/f_i\|_{P,2} \leqslant (1+2\bar f) K_q $, and $\|M\|_{P,2}\leqslant (1+2\bar f)K_q$ we have with probability $1-o(1)$ $$\sup_{m\in \mathcal{F}_{2a}\cdot \mathcal{F}_{2b}'}|({\mathbb{E}_n}-\bar {\mathrm{E}})(m)| \lesssim \frac{\frac{1}{h}\sqrt{s\log(pn)/n}+h^{\bar k}}{n^{1/2}}\sqrt{s\log(pn)}+ n^{-1} s K_q \log(pn) \lesssim \delta_n n^{-1/2}$$ under $h^{-2}s^2\log^2(pn) \leqslant \delta_n^2 n$, $h^{\bar{k}}\sqrt{s\log(pn)}\leqslant \delta_n$ and $K_q^2s^2\log^2(pn)\leqslant \delta_n^2 n$.

Moreover, $\mathcal{M}$ is the union of $\binom{p}{C\widetilde s_{{\theta \tau}}}$ VC subgraph of dimension $C\widetilde s_{{\theta \tau}}$, and for any $m \in \mathcal{F}_{2a}\cdot \mathcal{F}_{2b}''$ we have $$

array[array omitted — 264 chars of source]

$$ Therefore, by Lemma \ref{lemma:CCK} with $\sigma:= C\{\frac{1}{h}\sqrt{s\log(pn)/n}+h^{\bar k}+\lambda\sqrt{s}/n\}$, $a=pnK_qs$, $\|F_2''\|_{P,2} = \|(2\bar f\vee K_q^{-1}\|\tilde x_i\|_\infty)\|x_i\|_\infty\|_{P,2} \leqslant (1+2\bar f) K_q$, and $\|M\|_{P,2}\leqslant (1+2\bar f)K_q$ we have with probability $1-o(1)$ $$\sup_{m\in \mathcal{F}_{2a}\cdot \mathcal{F}_{2b}”}|({\mathbb{E}_n}-\bar {\mathrm{E}})(m)| \lesssim \frac{\frac{1}{h}\sqrt{\frac{s\log(pn)}{n}}+h^{\bar k}+ \frac{\lambda\sqrt{s}}{n}}{n^{1/2}}\sqrt{\widetilde s_{{\theta \tau}}\log(pn)}+ n^{-1} \widetilde s_{{\theta \tau}} K_q \log(pn) \lesssim \delta_n n^{-1/2}$$ under the conditions $h^{-2}s\widetilde s_{{\theta \tau}}\log^2(pn) \leqslant \delta_n^2 n$, $h^{\bar{k}}\sqrt{\widetilde s_{{\theta \tau}}\log(pn)}\leqslant \delta_n$, $\lambda\sqrt{s\widetilde s_{{\theta \tau}} \log(pn)}\leqslant \delta_n n$ and $K_q^2\widetilde s_{{\theta \tau}}^2\log^2(pn)\leqslant \delta_n^2 n$ assumed in Condition D.

Next we verify the second part of Condition IQR(iii), namely ((ref)). Note that since $\mathcal{A}_\tau \subset \{ \alpha : |\alpha - \alpha_\tau | \leqslant C / \log n + |\widetilde \alpha_\tau - \alpha_\tau| \leqslant C\sqrt{s\log(pn)/n}\}$, $\check\alpha_\tau \in \mathcal{A}_\tau$ and $s\log(pn) \leqslant \delta_n^2 n$ implies the required consistency $|\check\alpha_\tau-\alpha_\tau|\leqslant \delta_n$. To show the other relation in ((ref)), equivalent to with probability $1-o(1)$ $$| {\mathbb{E}_n}[\psi_{\check\alpha_\tau,\widehat h}(y_i,d_i,z_i)] |\leqslant \delta_n^{1/2}\ n^{-1/2},$$ note that for any $\alpha \in \mathcal{A}_\tau$ (since it implies $|\alpha-\alpha_\tau|\leqslant \delta_n$) and $\widehat h \in \overline{\mathcal{F}}$, we have with probability $1-o(1)$ $$

array[array omitted — 300 chars of source]

$$ from relations (\ref{Def:IdentityHelp}) with $\alpha$ instead of $\check\alpha_\tau$, (\ref{EqGamma20}), (\ref{BoundGamma2first}), and (\ref{BoundGamma2third}). Letting $\alpha^* = \alpha_\tau - \{\bar {\mathrm{E}}[f_id_iv_i]\}^{-1}{\mathbb{E}_n}[\psi_{\alpha_\tau, h_0}(y_i,d_i,z_i)]$, we have $| {\mathbb{E}_n}[\psi_{\alpha^*,\widehat h}(y_i,d_i,z_i)] | = O(\delta_n^{1/2} n^{-1/2})$ with probability $1-o(1)$ since $|\alpha^* - \alpha_\tau| \lesssim_P n^{-1/2}$. Thus, with probability $1-o(1)$ $$

array[array omitted — 404 chars of source]

$$ as ${\mathbb{E}_n}[\widehat v_i^2]$ is bounded away from zero with probability $1-o(1)$.

Next we verify Condition IQR(iv). The first condition follows from the uniform consistency of $\widehat f_i$ and $\max_{i\leqslant n}\|x_i\|_\infty \|\tilde \theta_\tau - \theta_{0\tau}\|_1 \lesssim_P K_q s \sqrt{\log(pn)/n}\to 0$ under $K_q^2s^2\log^2(pn)\leqslant \delta_n n$. The second condition in IQR(iv) also follows since { $$

array[array omitted — 564 chars of source]

$$} which implies the result under $K_q^2s^2\log(pn)\leqslant \delta_n n$.

The consistency of $\widehat\sigma_{1n}$ follows from $\|\widehat v_i-v_i\|_{2,n}\to_P 0$ and the moment conditions. The consistency of $\widehat\sigma_{3,n}$ follow from Lemma (ref). Next we show the consistency of $\widehat \sigma_{2n}^2=\{ {\mathbb{E}_n}[\check f_i^2 (d_i,x_{i\check T}')'(d_i,x_{i\check T }')]\}^{-1}_{11}$. Because $f_i\geqslant {\underline{f}}$, sparse eigenvalues of size $\ell_ns$ are bounded away from zero and from above with probability $1-\Delta_n$, and $\max_{i\leqslant n} |\widehat f_i - f_i| = o_P(1)$ by Condition D, we have $$ \{ {\mathbb{E}_n}[\widehat f_i^2 (d_i,x_{i\check T }')'(d_i,x_{i\check T }')]\}^{-1}_{11} = \{ {\mathbb{E}_n}[ f_i^2 (d_i,x_{i\check T }')'(d_i,x_{i\check T }')]\}^{-1}_{11} + o_P(1).$$ So that $\widehat\sigma_{2n}-\tilde \sigma_{2n}\to_P 0$ for $$\tilde \sigma_{2n}^2=\{ {\mathbb{E}_n}[ f_i^2 (d_i,x_{i\check T }')'(d_i,x_{i\check T }')]\}^{-1}_{11} = \{ {\mathbb{E}_n}[ f_i^2 d_i^2] - {\mathbb{E}_n}[ f_i^2 d_ix_{i\check T }']\{ {\mathbb{E}_n}[f_i^2 x_{i\check T }x_{i\check T }']\}^{-1}{\mathbb{E}_n}[ f_i^2 x_{i\check T } d_i] \}^{-1}.$$ Next define $\check\theta_\tau[\check T] = \{ {\mathbb{E}_n}[ f_i^2 x_{i\check T }x_{i\check T}']\}^{-1}{\mathbb{E}_n}[ f_i^2 x_{i\check T } d_i]$ which is the least squares estimator of regressing $f_i d_i$ on $f_ix_{i\check T }$. Let $\check \theta_\tau$ denote the associated $p$-dimensional vector. By definition $f_ix_i'\theta_\tau = f_id_i-f_ir_{{\theta \tau}}-v_i$, so that $$

array[array omitted — 595 chars of source]

$$ We have that $|{\mathbb{E}_n}[ v_i \{f_id_i-v_i\} ]|=|{\mathbb{E}_n}[ f_i v_ix_i'\theta_{0\tau} ]|=o_P(\delta_n)$ since $$\bar {\mathrm{E}}[ (v_i f_i x_i'\theta_{0\tau} )^2] \leqslant \bar f^2\{\bar {\mathrm{E}}[v_i^4]\bar {\mathrm{E}}[(x_i'\theta_{0\tau})^4]\}^{1/2}\leqslant C$$ and $\bar {\mathrm{E}}[ f_i v_ix_i'\theta_{0\tau}]=0$. Moreover, ${\mathbb{E}_n}[f_i d_i f_i r_{{\theta \tau} i}]\leqslant \bar f_i^2\|d_i\|_{2,n}\|r_{{\theta \tau} i}\|_{2,n}=o_P(\delta_n)$, $|{\mathbb{E}_n}[f_i d_i f_ix_i'(\check\theta-\theta_\tau)]| \leqslant \|d_i\|_{2,n}\|f_ix_i'(\check\theta_\tau-\theta_\tau )\|_{2,n} = o_P(\delta_n)$ since $|\check T|\lesssim \widetilde s_{{\theta \tau}} + s$ with probability $1-o(1)$ and ${\rm supp}(\widehat\theta_\tau)\subset \check T$.

\\ {\it Part 2. Proof of the Double Selection.} The analysis of $\widehat\theta_\tau$ and $\widehat \beta_\tau$ are identical to the corresponding analysis for Part 1. Let $\widehat T^*_\tau$ denote the set of variables used in the last step, namely $\widehat T^*_\tau = {\rm supp}(\widehat\beta_\tau^{\lambda_\tau})\cup {\rm supp}(\widehat\theta_\tau)$ where $\widehat\beta_\tau^{\lambda_\tau}$ denotes the thresholded estimator. Using the same arguments as in Part 1, we have with probability $1-o(1)$ $$|\widehat T^*_\tau| \lesssim \widehat s^*_\tau = s + \frac{ns\log p}{h^2\lambda^2}+\left( \frac{nh^{\bar k}}{\lambda} \right)^2. $$

Next we establish preliminary rates for $\check\eta_\tau := (\check\alpha_\tau,\check\beta_\tau')'$ that solves

equation[equation omitted — 144 chars of source]

where $\widehat f_i = \widehat f_i(d_i,x_i)\geqslant 0$ is a positive function of $(d_i,z_i)$. We will apply Lemma (ref) as the problem above is a (post-selection) refitted quantile regression for $(\tilde y_i,\tilde x_i)=(\widehat f_i y_i, \widehat f_i(d_i,x_i')')$. Indeed, conditional on $\{d_i,z_i\}_{i=1}^n$, quantile moment condition holds as $$

array[array omitted — 408 chars of source]

$$ Since $\max_{i\leqslant n}\widehat f_i \wedge \widehat f_i^{-1} \lesssim 1$ with probability $1-o(1)$, the required side conditions of Lemma \ref{Thm:MainTwoStepGeneric} hold as in Part 1. We can take $\bar r_\tau = C\sqrt{s\log(pn)/n}$ with probability $1-o(1)$. Finally, to provide the bound $\widehat Q$, let $\eta_\tau = (\alpha_\tau,\beta_\tau')'$ and $\widehat\eta_\tau=(\widehat\alpha_\tau,\widehat\beta^{\lambda_\tau}_\tau \ ')'$ which are $Cs$-sparse vectors. By definition ${\rm supp}(\widehat\beta_\tau) \subset \widehat T^*_\tau$ so that $${\mathbb{E}_n}[\widehat f_i\{\rho_\tau(y_i-(d_i,x_i')\check\eta_\tau) - \rho_\tau(y_i-(d_i,x_i')\eta_\tau)\}] \leqslant {\mathbb{E}_n}[\widehat f_i\{\rho_\tau(y_i-(d_i,x_i')\widehat\eta_\tau) - \rho_\tau(y_i-(d_i,x_i')\eta_\tau)\}] $$ By Lemma \ref{Lemma:UpperQ} we have with probability $1-o(1)$ $$ {\mathbb{E}_n}[\widehat f_i\{\rho_\tau(y_i-(d_i,x_i')\widehat\eta_\tau) - \rho_\tau(y_i-(d_i,x_i')\eta_\tau)\}] \leqslant \widehat Q:= C\frac{s\log (p n)}{n}.$$ Thus, with probability $1-o(1)$, Lemma \ref{Thm:MainTwoStepGeneric} implies $$ \| \widehat f_i (d_i,x_i')(\check\eta_\tau - \eta_\tau)\|_{2,n} \lesssim \sqrt{ \frac{(s + \widehat s^*_\tau)\log(pn)}{n\phi_{{\rm min}}(Cs+C\widehat s^*_\tau)}} $$

Since $s \leqslant \widehat s^*_\tau$ and $1/\phi_{{\rm min}}(Cs+C\widehat s^*_\tau)\leqslant C'$ by Condition D we have $ \|\check\eta_\tau-\eta_\tau\|\lesssim \sqrt{\frac{\widehat s^*_\tau\log p}{n}}$.

Next we construct an orthogonal score function based on the solution of the weighted quantile regression ((ref)). By the first order condition for $(\check\alpha_\tau,\check\beta_\tau)$ in ((ref)) we have for $s_i \in \partial \rho_\tau( y_i-d_i\check\alpha_\tau -x_i'\check\beta_\tau)$ that $$ {\mathbb{E}_n}\left[ s_i \widehat f_i\binom{ d_{i}}{x_{i \widehat T^*_\tau}} \right] = 0. $$ By taking linear combination of the selected covariates via $(1,-\widetilde\theta_\tau)$ and defining $\widehat v_i = \widehat f_i(d_{i}-x_{i \widehat T^*_\tau}'\widetilde\theta_\tau)$ we have $\psi_{\alpha,\widehat h}(y_i,d_i,z_i) = (\tau - 1\{y_{i} \leqslant d_i \alpha + x_{i}'\check \beta_\tau\})\widehat v_i$. Since $s_i = \tau - 1\{y_{i} \leqslant d_i \check \alpha_\tau + x_{i}'\check \beta_\tau\}$ if $y_i\neq d_i \check \alpha_\tau + x_{i}'\check \beta_\tau$, $$

array[array omitted — 448 chars of source]

$$ Since $\widehat s^*_\tau:= |\widehat T^*_\tau|\lesssim \widetilde s_{{\theta \tau}}$ with probability $1-o(1)$, and $y_i$ has no point mass, $ y_{i} = d_i \check \alpha_\tau + x_{i}'\check \beta_\tau$ for at most $C\widetilde s_{{\theta \tau}}$ indices in $\{1,\ldots,n\}$. Therefore, we have with probability $1-o(1)$

equation[equation omitted — 290 chars of source]

under $K_q^2 \widetilde s^2_{{\theta \tau}} \leqslant \delta_n n$. Moreover,

equation[equation omitted — 233 chars of source]

with probability $1-o(1)$ under $\sqrt{\widehat s^*_\tau}\|\widehat v_i- v_{i}\|_{2,n} \leqslant \delta_n$ holding with probability $1-o(1)$. Therefore, the orthogonal score function implicitly created by the double selection estimator $\check\alpha_\tau$ approximately minimizes $$ \widetilde L_n(\alpha) = \frac{| {\mathbb{E}_n}[ (\tau - 1\{y_i \leqslant d_i \alpha + x_i'\check\beta_\tau\})\widehat v_i ] |^2}{{\mathbb{E}_n}[\{ (\tau - 1\{y_i \leqslant d_i \alpha + x_i'\check\beta_\tau\})^2\widehat v_i^2 ]}, $$ and we have $\widetilde L_n(\check\alpha_\tau) \lesssim \delta_n^{1/3} n^{-1}$ with probability $1-o(1)$ by ((ref)) and ((ref)). Thus the conditions of Lemma (ref) hold and the double selection estimator has the stated linear representation.

{\bf $\square$}

Auxiliary Inequalities

In this section we collect auxiliary inequalities that we use in our analysis.

lemmaConsider $\widehat\eta_u$ and $\eta_u$ where $\|\eta_u\|_0\leqslant s$. Denote by $\widehat \eta^{\mu}_u$ the vector obtained by thresholding $\widehat\eta_u$ as follows $\widehat\eta^\mu_{uj}=\widehat \eta_{uj}1\{|\widehat \eta_{uj}|\geqslant \mu / {\mathbb{E}_n}[\tilde x_{ij}^2]^{1/2}\}$. We have that $$\begin{array}{rl} \|\widehat \eta^\mu_u - \eta_u\|_{1} & \leqslant \|\widehat \eta_u - \eta_u \|_{1}+s\mu /\min_{j\leqslant p}{\mathbb{E}_n}[\tilde x_{ij}^2]^{1/2} \\ |{\rm supp}(\widehat\eta^\mu)| & \leqslant s + \|\widehat \eta_u - \eta_u \|_{1}\max_{j\leqslant p}{\mathbb{E}_n}[\tilde x_{ij}^2]^{1/2}/\mu\\ \|\tilde x_i'(\widehat \eta_u^\mu-\eta_u)\|_{2,n} & \leqslant \|\tilde x_i'(\widehat \eta_u-\eta_u)\|_{2,n} + \sqrt{\phi_{{\rm max}}(s)} \left\{ \frac{2\sqrt{s}\mu}{\min_{j\leqslant p}{\mathbb{E}_n}[\tilde x_{ij}^2]^{1/2}} + \frac{\|\widehat\eta_u-\eta_u\|_{1}}{\sqrt{s}}\right\}\end{array}$$ where $\phi_{{\rm max}}(m)= \sup_{1\leqslant \|\theta\|_0\leqslant m}\|\tilde x_i'\theta\|_{2,n}^2/\|\theta\|^2$.

{\bf Proof.}{\bf \ (Proof of Lemma (ref))} Let $T_u = {\rm supp}(\eta_u)$. The first relation follows from the triangle inequality $$

array[array omitted — 574 chars of source]

$$

To show the second result note that $$\|\widehat\eta_u-\eta_u\|_{1} \geqslant \{|{\rm supp}(\widehat\eta_u^\mu)|-s\}\mu/\max_{j\leqslant p}{\mathbb{E}_n}[\tilde x_{ij}^2]^{1/2}.$$ Therefore, we have $|{\rm supp}(\widehat\eta_u^\lambda)|-s \leqslant \|\widehat\eta_u-\eta_u\|_{1}\max_{j\leqslant p}{\mathbb{E}_n}[\tilde x_{ij}^2]^{1/2}/\mu$ which yields the result.

To show the third bound, we start using the triangle inequality $$ \|\tilde x_i'(\widehat\eta^\mu_u-\eta_u)\|_{2,n} \leqslant \|\tilde x_i'(\widehat\eta^\mu_u-\widehat\eta_u)\|_{2,n} + \|\tilde x_i'(\widehat\eta_u-\eta_u)\|_{2,n}.$$ Without loss of generality assume that order the components is so that $|(\widehat\eta^\mu_u-\widehat\eta_u)_j|$ is decreasing. Let $T_1$ be the set of $s$ indices corresponding to the largest values of $|(\widehat\eta^\mu_u-\widehat\eta_u)_j|$. Similarly define $T_k$ as the set of $s$ indices corresponding to the largest values of $|(\widehat\eta^\mu_u-\widehat\eta_u)_j|$ outside $\cup_{m=1}^{k-1}T_m$. Therefore, $\widehat\eta^\mu_u-\widehat\eta_u=\sum_{k=1}^{\left\lceil p/s \right\rceil} (\widehat\eta^\mu_u-\widehat\eta_u)_{T_k}$. Moreover, given the monotonicity of the components, $\|(\widehat\eta^\mu_u-\widehat\eta_u)_{T_k}\|_{2,\varpi} \leqslant \|(\widehat\eta^\mu_u-\widehat\eta_u)_{T_{k-1}}\|_1/\sqrt{s}$. Then, we have $$

array[array omitted — 1,085 chars of source]

$$

where the last inequality follows from the first result and the triangle inequality.

{\bf $\square$}

The following result follows from Theorem 7.4 of delapena and the union bound.

lemma[Moderate Deviation Inequality for Maximum of a Vector] Suppose that $$ \mathcal{S}_{j} = \frac{\sum_{i=1}^n U_{ij}}{\sqrt{ \sum_{i=1}^n U^2_{ij}}},$$ where $U_{ij}$ are independent variables across $i$ with mean zero. We have that $$ {\mathrm{P}} \left( \max_{1 \leqslant j\leqslant p }|\mathcal{S}_{j}| > \Phi^{-1}(1- \gamma/2p) \right) \leqslant \gamma \left(1 + \frac{A}{\ell^3_n}\right), $$ where $A$ is an absolute constant, provided that for $\ell_n>0$ $$ 0 \leqslant \Phi^{-1}(1- \gamma/(2p)) \leqslant \frac{n^{1/6}}{\ell_n} \min_{1\leqslant j \leqslant p} M[U_j]-1, \ \ \ M[U_j] := \frac{\left( \frac{1}{n} \sum_{i=1}^n E U_{ij}^2\right)^{1/2}}{\left(\frac{1}{n} \sum_{i=1}^n E|U_{ij}^3| \right)^{1/3}}. $$
lemmaLet $X_i$, $i=1,\ldots,n$, be independent random vectors in $ {\Bbb{R}}^p$. Let $$\bar\delta_n:= 2\left( \bar C K \sqrt{k} \log(1+k) \sqrt{\log (p\vee n)} \sqrt{\log n} \right)/\sqrt{n},$$ where $K\geqslant \{{\mathrm{E}}[ \max_{1\leqslant i\leqslant n}\|X_i\|_\infty^2]\}^{1/2}$ and $\bar C$ is a universal constant. Then we have $$ {\mathrm{E}}\left[ \sup_{\|\alpha\|_0\leqslant k, \|\alpha\| =1} \left| {\mathbb{E}_n}\left[ (\alpha'X_i)^2 - {\mathrm{E}}[(\alpha'X_i)^2] \right]\right|\right] \leqslant \bar\delta_n^2 + \bar\delta_n \sup_{\|\alpha\|_0\leqslant k, \|\alpha\| =1} \sqrt{\bar {\mathrm{E}}[(\alpha'X_i)^2]}. $$

{\bf Proof.} It follows from Theorem 3.6 of RudelsonVershynin2008, see BelloniChernozhukovHansen2011 for details. {\bf $\square$}

We will also use the following result of chernozhukov2012gaussian.

lemma[Maximal Inequality] Work with the setup above. Suppose that $F\geqslant \sup_{f \in \mathcal{F}}|f|$ is a measurable envelope for $\mathcal{F}$ with $\| F\|_{P,q} < \infty$ for some $q \geqslant 2$. Let $M = \max_{i\leqslant n} F(W_i)$ and $\sigma^{2} > 0$ be any positive constant such that $\sup_{f \in \mathcal{F}} \| f \|_{P,2}^{2} \leqslant \sigma^{2} \leqslant \| F \|_{P,2}^{2}$. Suppose that there exist constants $a \geqslant e$ and $v \geqslant 1$ such that \begin{equation*} \log \sup_{Q} N(\epsilon \| F \|_{Q,2}, \mathcal{F}, \| \cdot \|_{Q,2}) \leqslant v \log (a/\epsilon), \ 0 < \epsilon \leqslant 1. \end{equation*} Then \begin{equation*} {\mathrm{E}}_P [ \sup_{f\in \mathcal{F}} | \mathbb{G}_n(f)| ] \leqslant K \left( \sqrt{v\sigma^{2} \log \left ( \frac{a \| F \|_{P,2}}{\sigma} \right ) } + \frac{v\| M \|_{P, 2}}{\sqrt{n}} \log \left ( \frac{a \| F \|_{P,2}}{\sigma} \right ) \right), \end{equation*} where $K$ is an absolute constant. Moreover, for every $t \geqslant 1$, with probability $> 1-t^{-q/2}$, \begin{multline*} \sup_{f\in \mathcal{F}} | \mathbb{G}_n(f)| \leqslant (1+\alpha) {\mathrm{E}}_P [ \sup_{f\in \mathcal{F}} | \mathbb{G}_n(f)| ] + K(q) \Big [ (\sigma + n^{-1/2} \| M \|_{P,q}) \sqrt{t} + \alpha^{-1} n^{-1/2} \| M \|_{P,2}t \Big ], \ \end{multline*} $\forall \alpha > 0$ where $K(q) > 0$ is a constant depending only on $q$. In particular, setting $a \geqslant n$ and $t = \log n$, with probability $> 1- c(\log n)^{-1}$, \begin{equation} \sup_{f\in \mathcal{F}} | \mathbb{G}_n(f)| \leqslant K(q,c) \left ( \sigma \sqrt{v \log \left ( \frac{a \| F \|_{P,2}}{\sigma} \right ) } + \frac{v \| M \|_{P,q} } {\sqrt{n}}\log \left ( \frac{a \| F \|_{P,2}}{\sigma} \right ) \right), \end{equation} where $ \| M \|_{P,q} \leqslant n^{1/q} \| F\|_{P,q}$ and $K(q,c) > 0$ is a constant depending only on $q$ and $c$.