Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
144,285 characters · 23 sections · 63 citation commands
Valid Post-Selection Inference in High-Dimensional Approximately Sparse Quantile Regression Models
\def\spacingset#1{ {#1}} \spacingset{1}
\if11 {
} \fi
\if01 {
} \fi
{\it Keywords:} quantile regression, confidence regions post model selection, orthogonal score functions
\spacingset{1.45}
Many applications of interest require the measurement of the distributional impact of a policy (or treatment) on the relevant outcome variable. Quantile treatment effects have emerged as an important concept for measuring such distributional impacts (see, e.g., Lehmann,koenker:book). In this work we focus on the quantile treatment effect $\alpha_\tau$ on a policy/treatment variable $d$ of an outcome of interest $y$ in the (heteroskedastic) partially linear model: $$ \tau\textrm{-quantile}(y\mid z, d) = d\alpha_\tau + g_\tau(z), $$ where $g_\tau$ is the (unknown) confounding function of the other covariates $z$ which can be well approximated by a linear combination of $p$ technical controls. Large $p$ arises due to the existence of many features (as in genomic studies and econometric applications) and/or the use of basis expansions in non-parametric approximations. When $p$ is comparable or larger than the sample size $n$, this brings forth the need to perform model selection or regularization.
We propose methods to construct estimates and confidence regions for the coefficient of interest $\alpha_\tau$ based upon robust post-selection procedures. We establish the (uniform) validity of the proposed methods in a non-parametric setting. Model selection in those settings generically leads to a moderate mistake and traditional arguments based on perfect model selection do not apply. Therefore, the proposed methods are developed to be robust to such model selection mistakes. Furthermore, they are directly applicable to construction of simultaneous confidence bands when $d$ is multivariate and a continuum of quantile indices is of interest.
Broadly speaking the main obstacle to construct confidence regions with (asymptotically) correct coverage is the estimation of the confounding function $g_\tau$ that is (generically) non-regular because of the high-dimensionality. To overcome this difficulty we construct (explicitly or implicitly) an orthogonal score function which leads to a moment condition that is immune to first-order mistakes in the estimation of the confounding function $g_\tau$. The construction is based on preliminary estimation of the confounding function and properly partialling out the confounding factors $z$ from the policy/treatment variable $d$. The former can be achieved via $\ell_1$-penalized quantile regression BC-SparseQR,kato or post-selection quantile regression based on $\ell_1$-penalized quantile regression BC-SparseQR. The latter is carried out by heteroscedastic post-Lasso T1996,BellChenChernHans:nonGauss applied to a density-weighted equation. Then we propose two estimators for $\alpha_{\tau}$ based on: (i) a moment condition based on an orthogonal score function; (ii) a density-weighted quantile regression with all the variables selected in the previous steps. The latter method is reminiscent of the “post-double selection" method proposed in BCH2011:InferenceGauss,BelloniChernozhukovHansen2011. Explicitly or implicitly the last step estimates $\alpha_\tau$ by minimizing a Neyman-type score statistic Neyman1979.\footnote{We mostly focus on selection as a means of regularization, but certainly other regularizations (e.g. the use of $\ell_1$-penalized fits per se) are possible, although they perform no better than the methods we focus on.}
Under mild moment conditions and approximate sparsity assumptions, we establish that the estimator $\check\alpha_\tau$, as defined by either method (see Algorithms (ref) and (ref) below), is root-$n$ consistent and asymptotically normal,
where $\rightsquigarrow$ denotes convergence in distribution; in addition the estimator $\check{\alpha}_{\tau}$ admits a (pivotal) linear representation. Hence the confidence region defined by
has asymptotic coverage probability of $1-\xi$ provided that the estimate $\widehat \sigma^2_n$ is consistent for $\sigma_n^2$, namely, $\widehat\sigma_n^2/\sigma_n^2 = 1+ o_{{\mathrm{P}}}(1)$. In addition, we establish that a Neyman-type score statistic $L_n(\alpha)$ is asymptotically distributed as the chi-squared distribution with one degree of freedom when evaluated at the true value $\alpha=\alpha_\tau$, namely,
which in turn allows the construction of another confidence region:
which has asymptotic coverage probability of $1-\xi$. These convergence results hold under array asymptotics, permitting the data-generating process ${\mathrm{P}} = {\mathrm{P}}_n$ to change with $n$, which implies that these convergence results hold uniformly over large classes of data-generating processes. In particular, our results do not require separation of regression coefficients away from zero (the so-called “beta-min" conditions) for their validity. Importantly, we discuss how the procedures naturally allow for construction of simultaneous confidence bands for many parameters based on a (pivotal) linear representation of the proposed estimator.
Several recent papers study the problem of constructing confidence regions after model selection while allowing $p\gg n$. In the case of linear mean regression, BCH2011:InferenceGauss proposes a double selection inference in a parametric setting with homoscedastic Gaussian errors; BelloniChernozhukovHansen2011 studies a double selection procedure in a non-parametric setting with heteroscedastic errors; c.h.zhang:s.zhang and vandeGeerBuhlmannRitov2013 propose methods based on $\ell_1$-penalized estimation combined with “one-step" correction in parametric models. Going beyond mean regression models, vandeGeerBuhlmannRitov2013 provides high level conditions for the one-step estimator applied to smooth generalized linear problems, BelloniChernozhukovKato2013a analyzes confidence regions for a parametric homoscedastic LAD regression model under primitive conditions based on the instrumental LAD regression, and BelloniChernozhukovWei2013 provides two post-selection procedures to build confidence regions for the logistic regression. None of the aforementioned papers deal with the problem of the present paper.
Although related in spirit with our previous work BelloniChernozhukovHansen2011,BelloniChernozhukovKato2013a,BelloniChernozhukovWei2013, new tools and major departures from the previous works are required. First is the need to accommodate the non-differentiability of the loss function (which translates into discontinuity of the score function) and the non-parametric setting. In particular, we establish new finite sample bounds for the prediction norm on the estimation error of $\ell_1$-penalized quantile regression in nonparametric models that extend results of BC-SparseQR,kato. Perhaps more importantly, the use of post-selection methods in order to reduce bias and improve finite sample performance requires sparsity of the estimates. Although sharp sparsity bounds for $\ell_1$-penalized methods are available for smooth loss functions, those are not available for quantile regression precisely due to the lack of differentiability. This led us to developing sparse estimates with provable guarantees by suitable truncation while preserving the good rates of convergence despite of possible additional model selection mistakes. To handle heteroscedsaticity, which is a common feature in many applications, consistent estimation of the conditional density is necessary whose analysis is new in high-dimensions. Those estimates are used as weights in the weighted Lasso estimation for the auxiliary regression (((ref)) below). In addition, we develop new finite sample bounds for Lasso with estimated weights as the zero mean condition is not assumed to hold for each observation but rather to the average across all observations. Because the estimation of the conditional density function is at a slower rate, it affects penalty choices, rates of convergence, and sparsity of the Lasso estimates.
This work and some of the papers cited above achieve an important uniformity guarantee with respect to the (unknown) values of the parameters. These uniform properties translate into more reliable finite sample performance of the proposed inference procedures because they are robust with respect to (unavoidable) model selection mistakes. There is now substantial theoretical and empirical evidence on the potential poor finite sample performance of inference methods that rely on perfect model selection when applied to models without separation from zero of the coefficients (i.e., small coefficients). Most of the criticism of these procedures are consequence of negative results established in LeebPotscher2005, leeb:potscher:hodges, and the references therein.
Notation. We work with triangular array data $\{ \omega_{i,n} : i=1,\dots,n; n=1,2,3,\dots\}$ where for each $n$, $\{ \omega_{n,i} ; i=1,\dots,n \}$ is defined on the probability space $(\Omega, \mathcal{S}, {\mathrm{P}}_n)$. Each $\omega_{i,n}= (y_{i,n}', z_{i,n}', d_{i,n}')'$ is a vector which are i.n.i.d., that is, independent across $i$ but not necessarily identically distributed. Hence all parameters that characterize the distribution of $\{\omega_{i,n} : i=1,\dots,n\}$ are implicitly indexed by ${\mathrm{P}}_n$ and thus by $n$. We omit this dependence from the notation for the sake of simplicity. We use ${\mathbb{E}_n}$ to abbreviate the notation $n^{-1}\sum_{i=1}^n$; for example, ${\mathbb{E}_n}[f] := {\mathbb{E}_n}[f(\omega_{i})] := n^{-1} \sum_{i=1}^n f(\omega_{i})$. We also use the following notation: $\bar {\mathrm{E}}[f] := {\mathrm{E}} \left [ {\mathbb{E}_n}[f] \right] = {\mathrm{E}} \left [ {\mathbb{E}_n}[f(\omega_{i})] \right ]= n^{-1} \sum_{i=1}^n {\mathrm{E}}[f(\omega_{i})]$. The $\ell_2$-norm is denoted by $\|\cdot\|$; the $\ell_0$-“norm” $\|\cdot\|_0$ denotes the number of non-zero components of a vector; and the $\ell_{\infty}$-norm $\| \cdot \|_{\infty}$ denotes the maximal absolute value in the components of a vector. Given a vector $\delta \in {\Bbb{R}}^p$ and a set of indices $T \subset \{1,\ldots,p\}$, we denote by $\delta_T \in {\Bbb{R}}^p$ the vector in which $\delta_{Tj} = \delta_j$ if $j\in T$ and $\delta_{Tj}=0$ if $j \notin T$.
For a quantile index $\tau \in (0,1)$, we consider a partially linear conditional quantile model
where $y_i$ is the outcome variable, $d_i$ is the policy/treatment variable, and confounding factors are represented by the variables $z_i$ which impact the equation through an unknown function $g_\tau$. We shall use a large number $p$ of technical controls $x_i=X(z_i)$ to achieve an accurate approximation to the function $g_\tau$ in ((ref)) which takes the form:
where $r_{{\tau} i}$ denotes an approximation error. We view $\beta_\tau$ and $r_{{\tau}}$ as nuisance parameters while the main parameter of interest is $\alpha_\tau$ which describes the impact of the treatment on the conditional quantile (i.e., quantile treatment effect).
In order to perform robust inference with respect to model selection mistakes, we construct a moment condition based on a score function that satisfies an additional orthogonality property that makes them immune to first-order changes in the value of the nuisance parameter. Letting $f_i = f_{\epsilon_i}(0\mid d_i, z_i)$ denote the conditional density at 0 of the disturbance term $\epsilon_i$ in ((ref)), the construction of the orthogonal moment condition is based on the linear projection of the regressor of interest $d_i$ weighted by $f_i$ on the $x_i$ variables weighted by $f_i$
where $\theta_{0\tau} \in \arg\min \bar {\mathrm{E}}[ f_i^2(d_i-x_i'\theta)^2]$. The orthogonal score function $\psi_i(\alpha):=(1\{y_i\leqslant d_i\alpha+ x_i'\beta_\tau + r_{{\tau}}\}-\tau)v_i$ leads to a moment condition to estimate $\alpha_\tau$,
and satisfies the following orthogonality condition with respect to first-order changes in the value of the nuisance parameters $\beta_\tau$ and $\theta_{0\tau}$:
In order to handle the high-dimensional setting, we assume that $\beta_\tau$ and $\theta_{0\tau}$ are approximately sparse, namely, it is possible to choose sparse vector $\beta_\tau$ and $\theta_\tau$ such that:
The latter equation requires that it is possible to choose the sparsity index $s$ so that the mean squared approximation error is of no larger order than the variance of the oracle estimator for estimating the coefficients in the approximation. See chen:Chapter for a detailed discussion of this notion of approximate sparsity.
The methodology based on the orthogonal score function ((ref)) can be used for construction of many different estimators that have the same first-order asymptotic properties but potentially different finite sample behaviors. In the main part of the paper we present two such procedures in detail (the discussion on additional variants can be found in Subsection (ref) of the Supplementary Appendix). Our procedures use $\ell_1$-penalized quantile regression and $\ell_1$-penalized weighted least squares as intermediate steps (we collect the recommended choices of the user-chosen parameters in Remark (ref) below). The first procedure stated in Algorithm (ref) is based on the explicit construction of the orthogonal score function.
Step 6 of Algorithm (ref) solves the empirical analog of ((ref)). We will show validity of the confidence regions for $\alpha_\tau$ defined in ((ref)) and ((ref)). We note that the truncation in Step 2 for the solution of the penalized quantile regression (provably) induces a sparse solution with the same rate of convergence as the original estimator. This is required because the post-selection methods exhibit better finite sample behaviors in our simulations.
The second algorithm is based on selecting relevant variables from equations ((ref)) and ((ref)), and running a weighted quantile regression.
Although the orthogonal score function is not explicitly constructed in Algorithm (ref), inspection of the proof reveals that an orthogonal score function is constructed implicitly via the optimality conditions of the weighted quantile regression in Step 4.
The implementation of the algorithms in Section (ref) requires an estimate of the conditional density function $f_i$ which is typically unknown under heteroscedasticity. Following koenker:book, we shall use the observation that $1/ f_i = \partial Q(\tau\mid d_i, z_i)/\partial \tau$ to estimate $f_i$ where $Q(\cdot\mid d_i,z_i)$ denotes the conditional quantile function of the outcome. Let $\widehat Q(u\mid z_i,d_i)$ denote an estimate of the conditional $u$-quantile function $Q(u\mid z_i,d_i)$, based on either $\ell_1$-penalized quantile regression or an associated post-selection method, and let $h =h_n \to 0$ denote a bandwidth parameter. Then an estimator of $f_i$ can be constructed as
When the conditional quantile function is three times continuously differentiable, this estimator is based on the first order partial difference of the estimated conditional quantile function, and so it has the bias of order $h^2$. Under additional smoothness assumptions, an estimator that has a bias of order $h^4$ is given by {
}
In this section we provide regularity conditions that are sufficient for validity of the main estimation and inference results. In what follows, let $c,C$, and $q$ be given (fixed) constants with $c > 0, C \geqslant 1$ and $q \geqslant 4$, and let $\ell_n \uparrow \infty, \delta_n \downarrow 0$, and $\Delta_n \downarrow 0$ be given sequences of positive constants. We assume that the following condition holds for the data generating process ${\mathrm{P}} = {\mathrm{P}}_n$ for each $n$.
Condition AS(${\mathrm{P}}$). (i) Let $\{(y_i,d_i,x_i=X(z_i)) : i=1,\ldots,n\}$ be independent random vectors that obey the model described in ((ref)) and ((ref)) with $\|\theta_{0\tau}\|+\|\beta_\tau\|+|\alpha_\tau|\leqslant C$. (ii) There exists $s \geqslant 1$ and vectors $\beta_{\tau}$ and $\theta_\tau$ such that $x_i'\theta_{0\tau} = x_i' \theta_{\tau} + r_{{\theta \tau} i}, \ \|\theta_{\tau}\|_0 \leqslant s, \ \bar {\mathrm{E}} [r_{{\theta \tau} i}^2] \leqslant C s/n$, $\|\theta_{0\tau}-\theta_\tau\|_1 \leqslant s\sqrt{\log(pn)/n}$, and $g_\tau(z_i) = x_i' \beta_{\tau} + r_{{\tau} i}, \ \|\beta_\tau\|_0 \leqslant s, \ \bar {\mathrm{E}}[r_{{\tau} i}^2] \leqslant C s/n$. (iii) The conditional distribution function of $\epsilon_i$ is absolutely continuous with continuously differentiable density $f_{\epsilon_i\mid d_i, z_i}(\cdot\mid d_i,z_i)$ such that $0<\underline{f} \leqslant f_i \leqslant \sup_t f_{\epsilon_i\mid d_i, z_i}(t\mid d_i,z_i) \leqslant \bar f \leqslant C$ and $\sup_{t} |f_{\epsilon_i\mid d_i, z_i}'(t\mid d_i, z_i)|\leqslant \bar f'\leqslant C$.
Condition AS(i) imposes the setting discussed in Section (ref) in which the error term $\epsilon_i$ has zero conditional $\tau$-quantile. The approximate sparsity on the high-dimensional parameters is stated in Condition AS(ii). Condition AS(iii) is a standard assumption on the conditional density function in the quantile regression literature (see koenker:book) and the instrumental quantile regression literature (see ch:iqrWeakId). Next we summarize the moment conditions we impose.
Condition M(${\mathrm{P}}$). (i) We have $\bar {\mathrm{E}}[\{(d_{i},x_{i}')\xi\}^2] \geqslant c\|\xi\|^2$ and $\bar {\mathrm{E}}[\{(d_i,x_i')\xi\}^4] \leqslant C\|\xi\|^4$ for all $\xi \in {\Bbb{R}}^{p+1}$, $c \leqslant \min_{1 \leqslant j\leqslant p} \bar {\mathrm{E}}[ |f_ix_{ij}v_i-{\mathrm{E}}[f_ix_{ij}v_i]|^2]^{1/2} \leqslant \max_{1\leqslant j\leqslant p} \bar {\mathrm{E}}[|f_i x_{ij} v_i|^3]^{1/3} \leqslant C$. (ii) The approximation error satisfies $|\bar {\mathrm{E}}[f_iv_ir_{{\tau} i}]|\leqslant \delta_nn^{-1/2}$ and $\bar {\mathrm{E}}[(x_i'\xi)^2r_{{\tau} i}^2]\leqslant C\|\xi\|^2 \bar {\mathrm{E}}[r_{{\tau} i}^2]$ for all $\xi \in {\Bbb{R}}^{p}$. (iii) Suppose that $K_q = {\mathrm{E}}[\max_{1 \leqslant i \leqslant n}\|(d_i,v_i,x_i')' \|_\infty^q]^{1/q}$ is finite and satisfies $(K_q^2s^2 + s^3)\log^3(pn) \leqslant n\delta_n$ and $K_q^4s\log(pn)\log^3n\leqslant \delta_nn$.
Condition M(i) imposes moment conditions on the variables. Condition M(ii) imposes requirements on the approximation error. Condition M(iii) imposes growth conditions on $s$, $p$, and $n$. In particular these conditions imply that the population eigenvalues of the design matrix are bounded away from zero and from above. They ensure that sparse eigenvalues and restricted eigenvalues are well behaved which are used in the analysis of penalized estimators and sparsity properties needed for the post-selection estimator.
Our last set of conditions pertains to the estimation of the conditional density function $(f_i)_{i=1}^n$ which has a non-trivial impact on the analysis. We denote by $\mathcal{U}$ the finite set of quantile indices used in the estimation of the conditional density. Under mild regularity conditions the estimators ((ref)) and ((ref)) achieve
where $\bar k = 2$ for ((ref)) and $\bar k =4$ for ((ref)). Condition D summarizes sufficient conditions to account for the impact of density estimation via (post-selection) $\ell_1$-penalized quantile regression estimators.
{\bf Condition D.} For $u \in \mathcal{U}$, assume that $u\textrm{-quantile}(y_i\mid z_i, d_i) = d_i\alpha_u + x_i'\beta_u + r_{ui}$, $f_{ui}=f_{y_i\mid d_i,z_i}(d_i\alpha_u + x_i'\beta_u + r_{ui} \mid z_i,d_i)\geqslant c$ where $\bar {\mathrm{E}}[r_{u i}^2] \leqslant \delta_n n^{-1/2}$ and $|r_{ui}|\leqslant \delta_nh$ for all $i=1,\ldots,n$, and the vector $\beta_u$ satisfies $\|\beta_u\|_0\leqslant s$. (ii) For $\widetilde s_{\theta \tau} = s+\frac{ns\log (n\vee p)}{h^2\lambda^2}+ \left(\frac{nh^{\bar k}}{\lambda}\right)^2$, suppose $h^{\bar k}\sqrt{\widetilde s_{{\theta \tau}}\log(pn) }\leqslant \delta_n$, $h^{-2}K_q^2s\log(pn)\leqslant \delta_n n$, $\lambda K_q^2\sqrt{s} \leqslant \delta_n n$, $h^{-2}s\widetilde s_{\theta \tau} \log(pn) \leqslant \delta_n n$, $\lambda\sqrt{s \widetilde s_{\theta \tau} \log(pn)} \leqslant \delta_n n$, and $K_q^2\widetilde s_{{\theta \tau}}\log^2(pn)\log^3(n) \leqslant \delta_nn$.
Condition D(i) imposes the approximately sparse assumption for the $u$-conditional quantile function for quantile indices $u$ in a neighborhood of the quantile index $\tau$. Condition D(ii) provides growth conditions relating $s$, $p$, $n$, $h$ and $\lambda$. Subsection (ref) in the Supplementary Appendix discusses specific choices of penalty level $\lambda$ and of bandwidth $h$ together with the implied conditions on the triple $(s,p,n)$. In particular they imply that sparse eigenvalues of order $\widetilde s_{{\theta \tau}}$ are well behaved.
In this section we state our theoretical results. We establish the first order equivalence of the proposed estimators. We construct the estimators as defined as in Algorithm (ref) and (ref) with parameters $\lambda_\tau$ as in ((ref)), $\widehat \Gamma_\tau$ as in ((ref)), and $\mathcal{A}_\tau$ as in ((ref)). The choices of $\lambda$ and $h$ satisfy Condition D.
The asymptotically correct coverage of the confidence regions $\mathcal{C}_{\xi,n}$ and $\mathcal{I}_{\xi,n}$ as defined in ((ref)) and ((ref)) follows immediately. Theorem (ref) relies on post model selection estimators which in turn rely on achieving sparse estimates $\widehat \beta_\tau$ and $\widehat\theta_\tau$. The sparsity of $\widehat\theta_\tau$ is derived in Section (ref) in the Supplemental Appendix under the recommended penalty choices. The sparsity of $\widehat \beta_\tau$ is not guaranteed under the recommended choices of penalty level $\lambda_\tau$ which leads to sharp rates. We bypass that by truncating small components to zero (as in Step 2 of Algorithm (ref)) which (provably) preserves the same rate of convergence and ensures the sparsity.
In addition to the asymptotic normality, Theorem (ref) establishes that the rescaled estimation error $\sigma_n^{-1}\sqrt{n}(\check\alpha_\tau - \alpha_\tau)$ is approximately equal to the process $\mathbb{U}_n(\tau)$, which is pivotal conditional on $v_1,\ldots, v_n$. Such a property is very useful since it is easy to simulate $\mathbb{U}_n(\tau)$ conditional on $v_1,\ldots, v_n$. Thus this representation provides us with another procedure to construct confidence intervals without relying on asymptotic normality which are useful for the construction of simultaneous confidence bands; see Section (ref).
Importantly, the results in Theorem (ref) allow for the data generating process to depend on the sample size $n$ and have no requirements on the separation from zero of the coefficients. In particular these results allow for sequences of data generating processes for which perfect model selection is not possible. In turn this translates into uniformity properties over a large class of data generating processes. Next we formalize these uniform properties. We let $\mathcal{P}_{n}$ denote the collection of distributions ${\mathrm{P}}$ for the data $\{ (y_{i},d_{i}, z_i')' \}_{i=1}^{n}$ such that Conditions AS, M and D are satisfied for given $n$. This is the collection of all approximately sparse models where the above sparsity conditions, moment conditions, and growth conditions are satisfied. Note that the uniformity results for the approximately sparse and heteroscedastic case are new even under fixed $p$ asymptotics.
In some applications we are interested on building confidence intervals that are simultaneously valid for many coefficients as well as for a range of quantile indices $\tau \in \mathcal{T}\subset (0,1)$ a fixed compact set. The proposed methods directly extend to the case of $d \in {\Bbb{R}}^{K}$ and $\tau \in \mathcal{T}$ $$ \tau\textrm{-quantile}(y\mid z, d) = \sum_{j=1}^K d_j\alpha_{\tau j} + \tilde g_\tau(z). $$ Indeed, for each $\tau \in \mathcal{T}$ and each $k=1,\ldots,K$, estimates can be obtained by applying the methods to the model ((ref)) as $$ \tau\textrm{-quantile}(y\mid z, d) = d_k\alpha_{\tau k} + g_\tau(z) \ \ \mbox{where} \ \ g_\tau(z) := \tilde g_\tau(z) + \sum_{j\neq k} d_j\alpha_{\tau j}. $$ For each $\tau \in \mathcal{T}$, Step 1 and the conditional density function $f_i$, $i=1,\ldots,n$, are the same for all $k=1,\ldots,K$. However, Steps 2 and 3 adapt to each quantile index and each coefficient of interest. The uniform validity of $\ell_1$-penalized methods for a continuum of problems (indexed by $\mathcal{T}$ in our case) has been established for quantile regression in BC-SparseQR and for least squares in BCFH2013program. The conclusions of Theorem (ref) are uniformly valid over $k=1,\ldots,K$ and $\tau \in \mathcal{T}$ (in the $\ell_\infty$-norm).
Simultaneous confidence bands are constructed by defining the following critical value $$ c^*(1-\xi) = \inf\left\{ t: {\mathrm{P}}\left( \sup_{\tau \in \mathcal{T},k=1,\ldots,K} |\mathbb{U}_n(\tau,k)| \leqslant t \mid \{d_i,z_i\}_{i=1}^n\right) \geqslant 1-\xi \right\}, $$ where the random variable $\mathbb{U}_n(\tau,k)$ is pivotal conditional on the data, namely, $$ \mathbb{U}_n(\tau,k) := \frac{\{\tau(1-\tau)\bar {\mathrm{E}}[v_{\tau k i}^2]\}^{-1/2}}{\sqrt{n}} \sum_{i=1}^n ( \tau - 1\{U_i\leqslant \tau\})v_{\tau k i}, $$ where $U_i$ are i.i.d. uniform random variables on $(0,1)$ independent from $\{d_i,z_i\}_{i=1}^n$, and $v_{\tau k i}$ is the error term in the decomposition ((ref)) for the pair $(\tau,k)$. Therefore $c^*(1-\xi)$ can be estimated since estimates of $v_{\tau k i}$ and $\sigma_{n\tau k}$, $\tau \in \mathcal{T}$ and $k=1,\ldots,K$, are available. Uniform confidence bands can be defined as $$ [ \check \alpha_{\tau k} - \sigma_{n\tau k}c^*(1-\xi)/\sqrt{n}, \check \alpha_{\tau k} + \sigma_{n\tau k}c^*(1-\xi)/\sqrt{n} ] \ \ \mbox{for} \ \ \tau \in \mathcal{T}, \ k = 1,\ldots,K. $$
Next we provide a simulation study to assess the finite sample performance of the proposed estimators and confidence regions. We focus our discussion on the double selection estimator as defined in Algorithm (ref) which exhibits a better performance. We consider the median regression case ($\tau =1/2$) under the following data generating process:
where $\alpha_\tau = 1/2$, $\theta_{0j} = 1/j^2, j=1,\ldots,p$, $x = (1,z')'$ consists of an intercept and covariates $z \sim N(0,\Sigma)$, and the errors $\epsilon$ and $\tilde v$ are independent. The dimension $p$ of the covariates $x$ is $300$, and the sample size $n$ is $250$. The regressors are correlated with $\Sigma_{ij} = \rho^{|i-j|}$ and $\rho = 0.5$. In this case, $f_i= 1/\{\sqrt{\pi(2-\mu +\mu d^2)}\}$ so that the coefficient $\mu \in \{0, 1\}$ makes the conditional density function of $\epsilon$ homoscedastic if $\mu = 0$ and heteroscedastic if $\mu = 1$. The coefficients $c_y$ and $c_d$ are used to control the $R^2$ in the equations: $y - d\alpha_\tau = x'(c_y\nu_0) + \epsilon$ and $ d = x'(c_d\nu_0) + \tilde v$ ; we denote the values of $R^2$ in each equation by $R^2_y$ and $R_d^2$. We consider values $(R^2_y, R^2_d)$ in the set $\{0, .1, .2, \ldots, .9\}\times \{0, .1, .2, \ldots, .9\}$. Therefore we have 100 different designs and perform $500$ Monte-Carlo repetitions for each design. For each repetition we draw new vectors $x_i$'s and errors $\epsilon_i$'s and $\tilde v_i$'s.
We perform estimation of $f_i$'s via ((ref)) even in the homoscedastic case ($\mu=0$), since we do not want to rely on whether the assumption of homoscedasticity is valid or not. We use $\widehat \sigma_{2n}$ as the standard error estimate for the post double selection estimator based on Algorithm (ref). As a benchmark we consider the standard (naive) post-selection procedure that applies $\ell_1$-penalized median regression of $y$ on $d$ and $x$ to select a subset of covariates that have predictive power for $y$, and then runs median regression of $y$ on $d$ and the selected covariates, omitting the covariates that were not selected. We report the rejection frequency of the confidence intervals with the nominal coverage probability of $95\%$. Ideally we should see the rejection rate of $5\%$, the nominal level, regardless of the underlying generating process ${\mathrm{P}} \in \mathcal{P}_n$. This is the so called uniformity property or honesty property of the confidence regions (see, e.g., RomanoWolf2007, RomanoShaikh20012, and LeebPotscher2006).
In the homoscedastic case, reported on the left column of Figure (ref), we have the empirical rejection probabilities for the naive post-selection procedure on the first row. These empirical rejection probabilities deviate strongly away from the nominal level of $5\%$, demonstrating the striking lack of robustness of this standard method. This is perhaps expected due to the Monte-Carlo design having regression coefficients not well separated from zero (that is, the “beta min" condition does not hold here). In sharp contrast, we see that the proposed procedure performs substantially better, yielding empirical rejection probabilities close to the desired nominal level of $5\%$. In the right column of Figure (ref) we report the results for the heteroscedastic case ($\mu = 1$). Here too we see the striking lack of robustness of the naive post-selection procedure. We also see that the confidence region based on the post-double selection method significantly outperforms the standard method, yielding empirical rejection probabilities close to the nominal level of $5\%$.
The purpose of this section is to examine practical usefulness of the new methods and contrast them with the standard post-selection inference (that assumes perfect selection).
We will assess statistical significance of socio-economic and biological factors on children's malnutrition, providing a methodological follow up on the previous studies done by FKH2011 and Koenker2011. The measure of malnutrition is represented by the child's height, which will be our response variable $y$. The socio-economic and biological factors will be our regressors $x$, which we shall describe in more detail below. We shall estimate the conditional first decile function of the child's height given the factors (that is, we set $\tau=.1$). We would like to perform inference on the size of the impact of the various factors on the conditional decile of the child's height. The problem has material significance, so it is important to conduct statistical inference for this problem responsibly.
The data comes originally from the Demographic and Health Surveys (DHS) conducted regularly in more than 75 countries; we employ the same selected sample of 37,649 as in Koenker (2012). All children in the sample are between the ages of 0 and 5. The response variable $y$ is the child's height in centimeters. The regressors $x$ include child's age, breast feeding in months, mothers body-mass index (BMI), mother's age, mother's education, father's education, number of living children in the family, and a large number of categorical variables, with each category coded as binary (zero or one): child's gender (male or female), twin status (single or twin), the birth order (first, second, third, fourth, or fifth), the mother's employment status (employed or unemployed), mother's religion (Hindu, Muslim, Christian, Sikh, or other), mother's residence (urban or rural), family's wealth (poorest, poorer, middle, richer, richest), electricity (yes or no), radio (yes or no), television (yes or no), bicycle (yes or no), motorcycle (yes or no), and car (yes or no).
Although the number of covariates ($p=30$) is substantial, the sample size ($n=37,649$) is much larger than the number of covariates. Therefore, the dataset is very interesting from a methodological point of view, since it gives us an opportunity to compare various methods for performing inference to an “ideal" benchmark of standard inference based on the standard quantile regression estimator without any model selection. This was proven theoretically in He:Shao and in Belloni:Chern:Fern under the $p \to \infty, p^3/n \to 0$ regime. This is also the general option recommended by koenker:book and LeebPotscher2005 in the fixed $p$ regime. Note that this “ideal" option does not apply in practice when $p$ is relatively large; however it certainly applies in the present example.
We will compare the “ideal" option with two procedures. First the standard post-selection inference method. This method performs standard inference on the post-model selection estimator, “assuming" that the model selection had worked perfectly. Second the double selection estimator defined as in Algorithm (ref). (The orthogonal score estimator performs similarly so it is omitted due to space constrains.) The proposed methods do not assume perfect selection, but rather build a protection against (moderate) model selection mistakes.
We now will compare our proposal to the “ideal" benchmark and to the standard post-selection method. We report the empirical results in Table (ref). The first column reports results for the ideal option, reporting the estimates and standard errors enclosed in brackets. The second column reports results for the standard post-selection method, specifically the point estimates resulting from the post-penalized quantile regression, reporting the standard errors as if there had been no model selection. The last column report the results for the double selection estimator (point estimate and standard error). Note that the Algorithm (ref) is applied sequentially to each of the variables. Similarly, in order to provide estimates and confidence intervals for all variables using the naive approach, if the covariate was not selected by the $\ell_1$-penalized quantile regression, it was included in the post-model selection quantile regression for that variable.
What we see is very interesting. First of all, let us compare the “ideal" option (column 1) and the naive post-selection (column 2). The Lasso selection method removes 16 out of 30 variables, many of which are highly significant, as judged by the “ideal" option. (To judge significance we use normal approximations and critical value of 3, which allows us to maintain $5\%$ significance level after testing up to 50 hypotheses). In particular, we see that the following highly significant variables were dropped by Lasso: mother's BMI, mother's age, twin status, birth orders one and two, and indicator of the other religion. The standard post-model selection inference then makes the assumption that these are true zeros, which leads us to misleading conclusions about these effects. The standard post-selection inference then proceeds to judge the significance of other variables, in some cases deviating sharply and significantly from the “ideal" benchmark. For example, there is a sharp disagreement on magnitudes of the impact of the birth order variables and the wealth variables (for “richer" and “richest" categories). Overall, for the naive post-selection, 8 out of 30 coefficients were more than 3 standard errors away from the coefficients of the “ideal" option.
We now proceed to comparing our proposed options to the “ideal" option. We see approximate agreement in terms of magnitude, signs of coefficients, and in standard errors. In few instances, for example, for the car ownership regressor, the disagreements in magnitude may appear large, but they become insignificant once we account for the standard errors.
The main conclusion from our study is that the standard/naive post-selection inference can give misleading results, confirming our expectations and confirming predictions of LeebPotscher2005. Moreover, the proposed inference procedure is able to deliver inference of high quality, which is very much in agreement with the “ideal" benchmark.
\setcounter{section}{0} \setcounter{page}{1}
\spacingset{1.45}
In what follows, we work with triangular array data $\{ \omega_{i,n} : i=1,\dots,n; n=1,2,3,\dots\}$ where for each $n$, $\{ \omega_{i,n} ; i=1,\dots,n \}$ is defined on the probability space $(\Omega, \mathcal{S}, {\mathrm{P}}_n)$. Each $\omega_{i,n}= (y_{i,n}', z_{i,n}', d_{i,n}')'$ is a vector which are i.n.i.d., that is, independent across $i$ but not necessarily identically distributed. Hence all parameters that characterize the distribution of $\{\omega_{i,n} : i=1,\dots,n\}$ are implicitly indexed by ${\mathrm{P}}_n$ and thus by $n$. We omit this dependence from the notation for the sake of simplicity. We use ${\mathbb{E}_n}$ to abbreviate the notation $n^{-1}\sum_{i=1}^n$; for example, ${\mathbb{E}_n}[f] := {\mathbb{E}_n}[f(\omega_{i})] := n^{-1} \sum_{i=1}^n f(\omega_{i})$. We also use the following notation: $\bar {\mathrm{E}}[f] := {\mathrm{E}} \left [ {\mathbb{E}_n}[f] \right] = {\mathrm{E}} \left [ {\mathbb{E}_n}[f(\omega_{i})] \right ]= n^{-1} \sum_{i=1}^n {\mathrm{E}}[f(\omega_{i})]$. The $\ell_2$-norm is denoted by $\|\cdot\|$; the $\ell_0$-“norm” $\|\cdot\|_0$ denotes the number of non-zero components of a vector; and the $\ell_{\infty}$-norm $\| \cdot \|_{\infty}$ denotes the maximal absolute value in the components of a vector. Given a vector $\delta \in {\Bbb{R}}^p$, and a set of indices $T \subset \{1,\ldots,p\}$, we denote by $\delta_T \in {\Bbb{R}}^p$ the vector in which $\delta_{Tj} = \delta_j$ if $j\in T$, $\delta_{Tj}=0$ if $j \notin T$. We also denote by $\delta^{(k)}$ the vector with $k$ non-zero components corresponding to $k$ of the largest components of $\delta$ in absolute value. We use the notation $(a)_+ = \max\{a,0\}$, $a \vee b = \max\{ a, b\}$, and $a \wedge b = \min\{ a , b \}$. We also use the notation $a \lesssim b$ to denote $a \leqslant c b$ for some constant $c>0$ that does not depend on $n$; and $a\lesssim_P b$ to denote $a=O_P(b)$. For an event $E$, we say that $E$ wp $\to$ 1 when $E$ occurs with probability approaching one as $n$ grows. Given a $p$-vector $b$, we denote $\text{support}(b) = \{ j \in \{1,...,p\}: b_j \neq 0\} $. We also use $\rho_\tau(t) = t(\tau - 1\{t\leqslant 0\})$ and $\varphi_\tau(t_1,t_2) = (\tau - 1\{t_1 \leqslant t_2\} )$.
Define the minimal and maximal $m$-sparse eigenvalues of a symmetric positive semidefinite matrix $M$ as
For notational convenience we write $\tilde x_i = (d_i,x_i')'$, $\phi_{{\rm min}}(m):= \phi_{{\rm min}}(m)[{\mathbb{E}_n} [\tilde x_i\tilde x_i']]$, and $\phi_{{\rm max}}(m):= \phi_{{\rm max}}(m)[{\mathbb{E}_n} [\tilde x_i\tilde x_i']]$.
There are several different ways to implement the sequence of steps underlying the two procedures outlined in Algorithms (ref) and (ref). The estimation of the control function $g_\tau$ can be done through other regularization methods like $\ell_1$-penalized quantile regression instead of the post-$\ell_1$-penalized quantile regression. The estimation of the error term $v$ in Step 2 can be carried out with Dantzig selector, square-root Lasso or the associated post-selection method could be used instead of Lasso or post-Lasso. Solving for the zero of the moment condition induced by the orthogonal score function can be substituted by a one-step correction from the $\ell_1$-penalized quantile regression estimator $\widehat\alpha_\tau$, namely, $\check \alpha_\tau = \widehat \alpha_\tau + ({\mathbb{E}_n}[\widehat v_i^2])^{-1} {\mathbb{E}_n}[(\tau - 1\{ y_i \leqslant \widehat\alpha_\tau d_i+x_i'\widehat\beta_\tau\})\widehat v_i]$.
Other variants can be constructing alternative orthogonal score functions. This can be achieved by changing the weights in the equation ((ref)) to alternative weights, say $\tilde f_i$, that lead to different errors terms $\tilde v$ that satisfies ${\mathrm{E}}[ \tilde f x \tilde v ] = 0$. Then the orthogonal score function is constructed as $\psi_i(\alpha) = (\tau - 1\{y_i\leqslant d_i\alpha + g_\tau(z_i)\}) \tilde v_i (\tilde f_i/f_i)$. It turns out that the choice $\tilde f_i = f_i$ minimizes the asymptotic variance of the estimator of $\alpha_\tau$ based upon the empirical analog of ((ref)), among all the score functions satisfying ((ref)). An example is to set $\tilde f_i = 1$ which would lead to $\tilde v_i=d_i-{\mathrm{E}}[d_i\mid z_i]$. Although such choice leads to a less efficient estimator, the estimation of ${\mathrm{E}}[d_i\mid z_i]$ and $f_i$ can be carried out separably which can lead to weaker regularity conditions.
As discussed in Remark (ref), in order to handle approximately sparse models to represent $g_\tau$ in ((ref)) an approximate orthogonality condition is assumed, namely
In the literature such a condition has been (implicitly) used before. For example, ((ref)) holds if the function $g_\tau$ is an exactly sparse linear combination of the covariates so that all the approximation errors are exactly zero, namely, $r_{{\tau} i} = 0, i=1,\ldots,n$. An alternative assumption in the literature that implies ((ref)) is to have ${\mathrm{E}}[f_id_i\mid z_i]=f_i\{x_i'\theta_{\tau}+r_{{\theta \tau} i}\}$, where $\theta_{\tau}$ is sparse and $r_{{\theta \tau} i}$ is suitably small, which implies orthogonality to all functions of $z_i$ since we have ${\mathrm{E}}[f_iv_i \mid z_i ] = 0$.
The high-dimensional setting makes the condition ((ref)) less restrictive as $p$ grows. Our discussion is based on the assumption that the function $g_\tau$ belongs to a well behaved class of functions. For example, when $g_\tau$ belongs to a Sobolev space $\mathcal{S}(\alpha,L)$ for some $\alpha \geqslant 1$ and $L>0$ with respect to the basis $\{x_j=P_j(z), j\geqslant 1\}$. As in TsybakovBook, a Sobolev space of functions consists of functions $g(z)=\sum_{j=1}^\infty\theta_jP_j(z)$ whose Fourier coefficients $\theta$ satisfy $$\theta \in \Theta(\alpha, L) = \left\{ \theta \in \ell^2(\mathbb{N}):
\right\}.$$ More generally, we can consider functions in a $p$-Rearranged Sobolev space $\mathcal{RS}(\alpha,p,L)$ which allow permutations in the first $p$ components as in \cite{BCW2014pivotal}. Formally, the class of functions $g(z)=\sum_{j=1}^\infty\theta_jP_j(z) $ such that $$\theta \in \Theta^{R}(\alpha,p,L) = \left\{ \theta \in \ell^2(\mathbb{N}) : \sum_{j=1}^\infty|\theta_j| < \infty, \
\right\}.$$ It follows that $\mathcal{S}(\alpha,L) \subset \mathcal{RS}(\alpha,p,L)$ and $p$-Rearranged Sobolev space reduces substantially the dependence on the ordering of the basis.
Under mild conditions, it was shown in BCW2014pivotal that for functions in $\mathcal{RS}(\alpha,p,L)$ the rate-optimal choice for the size of the support of the oracle model obeys $s\lesssim n^{1/[2\alpha+1]}$. It follows that $$
$$ However, this bound cannot guarantee converge to zero faster than $\sqrt{n}$-rate to potentially imply (\ref{ExtraMomentDisc}). Fortunately, to establish relation (\ref{ExtraMomentDisc}) one can exploit orthogonality with respect all $p$ components of $x_i$. Indeed we have $$
$$ Therefore, condition (\ref{ExtraMomentDisc}) holds if $n=o(p^{2\alpha-1})$, in particular, for any $\alpha\geqslant 1$, $n = o(p)$ suffices.
In this section we make some connections to the (local) minimax efficiency analysis from the semiparametric efficiency analysis. In this section for the sake of exposition we assume that $(y_i,x_i,d_i)_{i=1}^n$ are i.i.d., sparse models, $r_{{\theta \tau} i}=r_{{\tau} i} = 0$, $i=1,\ldots,n$, and the median case ($\tau = .5$). Lee2003 derives an efficient score function for the partially linear median regression model: $$ S_i= 2\varphi_\tau(y_i,d_i\alpha_\tau+x_i'\beta_\tau) f_i [d_i - m_\tau^*(z)], $$ where $m_\tau^*(x_i)$ is given by $$ m^*_\tau(x_i) = \frac{{\mathrm{E}}[f_i^2 d_i|x_i]}{{\mathrm{E}}[f^2_i |x_i]}. $$ Using the assumption $m^*_\tau(x_i)=x_i'\theta^*_\tau$ , where $\|\theta^*_\tau\|_0 \leqslant s \ll n$ is sparse, we have that $$ S_i= 2\varphi_\tau(y_i, d_i\alpha_\tau+x_i'\beta_\tau) v_i^*, $$ where $v_i^* = f_id_i - f_im^*_\tau(x_i)$ would correspond to $v_i$ in ((ref)). It follows that the estimator based on $v^*_i$ is actually efficient in the minimax sense (see Theorem 18.4 in kosorok:book), and inference about $\alpha_\tau$ based on this estimator provides best minimax power against local alternatives (see Theorem 18.12 in kosorok:book).
The claim above is formal as long as, given a law $Q_n$, the least favorable submodels are permitted as deviations that lie within the overall model. Specifically, given a law $Q_n$, we shall need to allow for a certain neighborhood $\mathcal{Q}_n^{\delta}$ of $Q_n$ such that $Q_n \in \mathcal{Q}_n^{\delta} \subset \mathcal{Q}_n$, where the overall model $\mathcal{Q}_n$ is defined similarly as before, except now permitting heteroscedasticity (or we can keep homoscedasticity $f_i = f_{\epsilon}$ to maintain formality). To allow for this we consider a collection of models indexed by a parameter $t = (t_1, t_2)$:
where $\|\beta_\tau\|_0 \vee \| \theta^*_\tau\|_0 \leqslant s/2$ and conditions as in Section (ref) hold. The case with $t=0$ generates the model $Q_n$; by varying $t$ within $\delta$-ball, we generate models $ \mathcal{Q}_n^{\delta}$, containing the least favorable deviations. By Lee2003, the efficient score for the model given above is $S_i$, so we cannot have a better regular estimator than the estimator whose influence function is $J^{-1}S_i$, where $J = {\mathrm{E}}[S_i^2]$. Since our model $\mathcal{Q}_n$ contains $ \mathcal{Q}_n^{\delta}$, all the formal conclusions about (local minimax) optimality of our estimators hold from theorems cited above (using subsequence arguments to handle models changing with $n$). Our estimators are regular, since under $\mathcal{Q}_n^t$ with $t = ( O(1/\sqrt{n}), o(1))$, their first order asymptotics do not change, as a consequence of Theorems in Section (ref). (Though our theorems actually prove more than this.)
The proof of Theorem (ref) provides a detailed analysis for generic choice of bandwidth $h$ and the penalty level $\lambda$ in Step 2 under Condition D. Here we discuss two particular choices for $\lambda$, for $\gamma = 0.05/n$
The choice (i) for $\lambda$ leads to a sparser estimators by adjusting to the slower rate of convergence of $\widehat f_i$, see ((ref)). The choice (ii) for $\lambda$ corresponds to the (standard) choice of penalty level in the literature for Lasso. Indeed, we have the following sparsity guarantees for the $\tilde \theta_{\tau}$ under each choice
In addition to the requirements in Condition M, $(K_q^2s^2+s^3)\log^3(pn)\leqslant \delta_n n$ and $K_q^4s\log(pn)\log^3n\leqslant \delta_n n$, which are independent of $\lambda$ and $h$, we have that Condition D simplifies to
For example, using the choice of $\widehat f_i$ as in ((ref)) so that $\bar k = 4$, we have that the following choice growth conditions suffice for the conditions above:
This section contains the main tools used in establishing the main inferential results. The high-level conditions here are intended to be applicable in a variety of settings and they are implied by the regularities conditions provided in the previous sections. The results provided here are of independent interest (e.g. properties of Lasso under estimated weights). We establish the inferential results ((ref)) and ((ref)) in Section (ref) under high level conditions. To verify these high-level conditions we need rates of convergence for the estimated residuals $\widehat v$ and the estimated confounding function $\widehat g_\tau (z) = x'\widehat \beta_\tau$ which are established in sections (ref) and (ref) respectively. The main design condition relies on the restricted eigenvalue proposed in BickelRitovTsybakov2009, namely for $\tilde x_i = [ d_i, x_i' ]'$
where $\mathbf{c} = (c+1)/(c-1)$ for the slack constant $c>1$, see BickelRitovTsybakov2009. When $\mathbf{c}$ is bounded, it is well known that $\kappa_\mathbf{c}$ is bounded away from zero provided sparse eigenvalues of order larger than $s$ are well behaved, see BickelRitovTsybakov2009.
In this section for a quantile index $u\in (0,1)$, we consider the equation
where we observe $\{(\tilde y_i,\tilde x_i):i=1,\ldots,n\}$, which are independent across $i$. To estimate $\eta_u$ we consider the $\ell_1$-penalized $u$-quantile regression estimate $$ \widehat\eta_u \in \arg\min_{\eta} {\mathbb{E}_n}[ \rho_u(\tilde y_i - \tilde x_i'\eta)] + \frac{\lambda_u}{n}\|\eta\|_1. $$ and the associated post-model selection estimates. That is, given an estimator $\bar\eta_{uj}$
We will be typically concerned with $\widehat\eta_u$, thresholded versions of the $\ell_1$-penalized quantile regression, defined as $\widehat\eta_{uj}^\mu = \widehat\eta_{uj}1\{ |\widehat\eta_{uj}|\geqslant \mu/{\mathbb{E}_n}[\tilde x_{ij}^2]^{1/2}\}$.
As established in BC-SparseQR for sparse models and in kato for approximately sparse models, under the event that
the estimator above achieves good theoretical guarantees under mild design conditions. Although $\eta_u$ is unknown, we can set $\lambda_u$ so that the event in ((ref)) holds with high probability. In particular, the pivotal rule proposed in BC-SparseQR and generalized in kato proposes to set $\lambda_u := c n\Lambda_u(1-\gamma\mid \tilde x)$ for $c> 1$ where
where $U_i\sim U(0,1)$ are independent random variables conditional on $\tilde x_i$, $i=1,\ldots,n$. This quantity can be easily approximated via simulations. Below we summarize the high level conditions we require.
{\bf Condition PQR.} Let $T_u = {\rm supp}(\eta_u)$ and normalize ${\mathbb{E}_n}[\tilde x_{ij}^2]=1$, $j=1,\ldots,p$. Assume that for some $s\geqslant 1$, $\|\eta_u\|_0\leqslant s$, $\|r_{ui}\|_{2,n}\leqslant C\sqrt{s\log(p)/n}$. Further, the conditional distribution function of $\epsilon_i$ is absolutely continuous with continuously differentiable density $f_\epsilon(\cdot\mid \tilde x_i, r_{ui})$ such that $0<\underline{f} \leqslant f_i \leqslant \sup_t f_{\epsilon_i\mid \tilde x_i, r_{ui}}(t \mid \tilde x_i, r_{ui}) \leqslant \bar f$, $\sup_{t} |f_{\epsilon_i\mid \tilde x_i, r_{ui}}'( t \mid \tilde x_i, r_{ui})|<\bar f'$ for fixed constants $\underline{f}$, $\bar f$ and $\bar f'$.
Condition PQR is implied by Condition AS. The conditions on the approximation error and near orthogonality conditions follows from choosing a model $\eta_u$ that optimally balance the bias/variance trade-off. The assumption on the conditional density is standard in the quantile regression literature even with fixed $p$ case developed in koenker:book or the case of $p$ increasing slower than $n$ studied in Belloni:Chern:Fern.
Next we present bounds on the prediction norm of the $\ell_1$-penalized quantile regression estimator.
Lemma (ref) establishes the rate of convergence in the prediction norm for the $\ell_1$-penalized quantile regression estimator. Exact constants are derived in the proof. The extra growth condition required for identification is mild. For instance we typically have $\lambda_u \sim \sqrt{n\log (np)}$ and for many designs of interest we have $$\inf_{\delta\in\Delta_\mathbf{c}} \|\tilde x_i'\delta\|_{2,n}^{3}/{\mathbb{E}_n}[|\tilde x_i'\delta|^3]$$ bounded away from zero (see BC-SparseQR). For more general designs we have { $$\inf_{\delta\in A_u} \frac{\|\tilde x_i'\delta\|_{2,n}^{3}}{{\mathbb{E}_n}[|\tilde x_i'\delta|^3]}\geqslant \inf_{\delta\in A_u} \frac{\|\tilde x_i'\delta\|_{2,n}}{ \|\delta\|_1\max_{i\leqslant n}\|\tilde x_i\|_\infty} \geqslant \frac{1}{\max_{i\leqslant n}\|\tilde x_i\|_\infty} \left( \frac{\kappa_{2\mathbf{c}}}{\sqrt{s}(1+\mathbf{c})} \wedge \frac{\lambda_u N}{8C\mathbf{c} s\log(p/\gamma) }\right) .$$}
Lemma (ref) provides the rate of convergence in the prediction norm for the post model selection estimator despite of possible imperfect model selection. In the current nonparametric setting it is unlikely for the coefficients to exhibit a large separation from zero. The rates rely on the overall quality of the selected model by $\ell_1$-penalized quantile regression and the overall number of components $\widehat s_u$. Once again the extra growth condition required for identification is mild. For more general designs we have $$\inf_{\|\delta\|_0 \leqslant \widehat s_u + s} \frac{\|\tilde x_i'\delta\|_{2,n}^{3}}{{\mathbb{E}_n}[|\tilde x_i'\delta|^3]}\geqslant \inf_{\|\delta\|_0 \leqslant \widehat s_u + s} \frac{\|\tilde x_i'\delta\|_{2,n}}{ \|\delta\|_1\max_{i\leqslant n}\|\tilde x_i\|_\infty} \geqslant \frac{\sqrt{\phi_{{\rm min}}(\widehat s_u+s)}}{\sqrt{\widehat s_u + s}\max_{i\leqslant n}\|\tilde x_i\|_\infty}.$$
In this section we consider the equation
where we observe $\{(d_i,z_i, x_i=X(z_i)):i=1,\ldots,n\}$, which are independent across $i$. We do not observe $\{f_i = f_\tau(d_i, z_i)\}_{i=1}^n$ directly and only estimates $\{\widehat f_i\}_{i=1}^n$ are available. Importantly, we only require that $ \bar {\mathrm{E}}[ f_iv_ix_i ] = 0$ and not ${\mathrm{E}}[f_ix_iv_i]=0$ for every $i=1,\ldots,n$. Also, we have that $T_{\theta \tau}={\rm supp}(\theta_\tau)$ is unknown but a sparsity condition holds, namely $|T_{\theta \tau}| \leqslant s$. To estimate $\theta_{\theta \tau}$ and $v_i$, we compute
where $\lambda$ and $\widehat \Gamma_\tau$ are the associated penalty level and loadings specified below. A difficulty is to account for the impact of estimated weights $\widehat f_i$ while also only using $\bar {\mathrm{E}}[ f_iv_ix_i ] = 0$.
We will establish bounds on the penalty parameter $\lambda$ so that with high probability the following regularization event occurs
As discussed in BickelRitovTsybakov2009,BC-PostLASSO,BCW-SqLASSO, the event above allows to exploit the restricted set condition $\|\widehat\theta_{\tau T^c_{\theta \tau}}\|_1\leqslant \tilde \mathbf{c} \|\widehat\theta_{\tau T_{\theta \tau}}-\theta_\tau\|_1$ for some $\tilde\mathbf{c} >1$. Thus rates of convergence for $\widehat \theta_\tau$ and $\widehat v_i$ defined on ((ref)) can be established based on the restricted eigenvalue $ \kappa_{\tilde\mathbf{c}}$ defined in ((ref)) with $\tilde x_i = x_i$.
However, the estimation error in the estimate $\widehat f_i$ of $f_i$ could slow the rates of convergence. The following are sufficient high-level conditions. In what follows $\underline c , \bar c, C, \underline{f}, \bar{f}$ are strictly positive constants independent of $n$.
{\bf Condition WL.} {\it For the model ((ref)) suppose that:\\ (i) for $s\geqslant 1$ we have $\|\theta_\tau\|_0\leqslant s$, $\Phi^{-1}(1-\gamma/2p) \leqslant \delta_n n^{1/6},$ \\ (ii) $\underline{f} \leqslant f_i \leqslant \bar{f}$, $\underline c \leqslant {\displaystyle \min_{j\leqslant p}} \{\bar {\mathrm{E}}[|f_ix_{ij}v_i-{\mathrm{E}}[f_ix_{ij}v_i]|^2]\}^{1/2} \leqslant {\displaystyle \max_{j\leqslant p} } \{\bar {\mathrm{E}}[|f_ix_{ij}v_i|^3]\}^{1/3} \leqslant C,$\\ (iii) with probability $1-\Delta_n$ we have ${\mathbb{E}_n}[ \widehat f_i^2 r_{{\theta \tau} i}^2] \leqslant c_r^2$, { $$
$$} (iv) $\ell \widehat \Gamma_{\tau 0} \leqslant \widehat \Gamma_{\tau} \leqslant u \widehat \Gamma_{\tau 0}$, for $\widehat \Gamma_{\tau 0jj}=\{{\mathbb{E}_n}[\widehat f_i^2 x_{ij}^2v_i^2]\}^{1/2}$, $1 - \delta_n\leqslant \ell \leqslant u \leqslant C$ with prob $1-\Delta_n$.}
Next we present results on the performance of the estimators generated by Lasso with estimated weights. In what follows, $\widehat \kappa_\mathbf{c}$ is defined with $\widehat f_i x_i$ instead of $\tilde x_i$ in ((ref)) so that $\widehat\kappa_\mathbf{c} \geqslant \kappa_\mathbf{c} \min_{i\leqslant n} \widehat f_i$.
Lemma (ref) above establishes the rate of convergence for Lasso with estimated weights. This automatically leads to bounds on the estimated residuals $\widehat v_i$ obtained with Lasso through the identity
The Post-Lasso estimator applies the least squares estimator to the model selected by the Lasso estimator ((ref)), $$ \widetilde\theta_\tau \in \arg\min_{\theta\in {\Bbb{R}}^p} \ \left\{ \ {\mathbb{E}_n}[\widehat f_i^2(d_i - x_i'\theta)^2] \ : \ \theta_j = 0, \ \mbox{if} \ \widehat\theta_{\tau j} = 0 \ \right\}, \ \ \mbox{set} \ \tilde v_i = \widehat f_i (d_i - x_i'\widetilde\theta_\tau).$$ It aims to remove the bias towards zero induced by the $\ell_1$-penalty function which is used to select components. Sparsity properties of the Lasso estimator $\widehat\theta_\tau$ under estimated weights follows similarly to the standard Lasso analysis derived in BellChenChernHans:nonGauss. By combining such sparsity properties and the rates in the prediction norm we can establish rates for the post-model selection estimator under estimated weights. The following result summarizes the properties of the Post-Lasso estimator.
Next we turn to analyze the estimator $\check \alpha_\tau$ obtained based on the orthogonal moment condition. In this section we assume that $$ L_n(\check \alpha_\tau) \leqslant \min_{\alpha \in \mathcal{A}_\tau} L_n(\alpha) + \delta_n n^{-1}.$$ This setting is related to the instrumental quantile regression method proposed in ch:iqrWeakId. However, in this application we need to account for the estimation of the noise $v$ that acts as the instrument which is known in the setting in ch:iqrWeakId. Condition IQR below suffices to make the impact of the estimation of instruments negligible to the first order asymptotics of the estimator $\check \alpha_\tau$. Primitive conditions that imply Condition IQR are provided and discussed in the main text.
Let $\{(y_i,d_i,z_i):i=1,\ldots,n\}$ be independent observations satisfying
Letting $\mathcal{D}\times\mathcal{Z}$ denote the domain of the random variables $(d,z)$, for $\tilde h = (\tilde g, \tilde {{\iota}})$, where $\tilde g$ is a function of variable $z$, and the instrument $\tilde {{\iota}}$ is a function that maps $(d,x)\mapsto \tilde {{\iota}}(d,z)$ we write $$
$$ We denote $h_0=(g_\tau,{{\iota}}_0)$ where ${{\iota}}_{0i}:=v_i=f_i(d_i-x_i'\theta_{0\tau})$. For some sequences $\delta_n\to 0$ and $\Delta_n\to 0$, we let $\overline{\mathcal{F}}$ denote a set of functions such that each element $\tilde h = (\tilde g,\tilde {{\iota}})\in \overline{\mathcal{F}}$ satisfies
and with probability $1-\Delta_n$ we have
We assume that the estimated functions $\widehat g$ and $\widehat {{\iota}}$ satisfy the following condition.
{\bf Condition IQR.} {\textit Let $\{(y_i,d_i,z_i):i=1,\ldots,n\}$ be random variables independent across $i$ satisfying ((ref)). Suppose that there are positive constants $0<c\leqslant C<\infty$ such that:\\ (i) $f_{y_i\mid d_i,z_i}(y\mid d_i,z_i) \leqslant \bar f$, $f_{y_i\mid d_i,z_i}'(y\mid d_i,z_i) \leqslant \bar f'$; $c \leqslant |\bar {\mathrm{E}}[f_id_i{{\iota}}_{0i}]|$, and $\bar {\mathrm{E}}[ {{\iota}}_{0i}^4 ] + \bar {\mathrm{E}}[ d_i^4 ] \leqslant C$;\\ (ii) $\{ \alpha : |\alpha-\alpha_\tau| \leqslant n^{-1/2}/\delta_n\} \subset \mathcal{A}_\tau$, where $\mathcal{A}_\tau$ is a (possibly random) compact interval; \\ (iii) with probability at least $1-\Delta_n$ the estimated functions $\widehat h = (\widehat g,\widehat {{\iota}}) \in \overline{\mathcal{F}}$ and
(iv) with probability at least $1-\Delta_n$, the estimated functions $\widehat h = (\widehat g,\widehat {{\iota}})$ satisfy $$\|\widehat {{\iota}}_i - {{\iota}}_{0i}\|_{2,n} \leqslant \delta_n \ \ \mbox{and} \ \ \|1\{|\epsilon_i|\leqslant |d_i(\alpha_\tau-\check\alpha_\tau)+g_{\tau i}-\widehat g_i|\}\|_{2,n}\leqslant \delta_n^2.$$}
{\bf Proof.}{\bf \ (Proof of Theorem (ref))} The first-order equivalence between the two estimators follows from establishing the same linear representation for each estimator. In Part 1 of the proof we consider the orthogonal score estimator. In Part 2 we consider the double selection estimator.
{\rm Part 1. Orthogonal score estimator.} We will verify Condition IQR and the result follows by Lemma (ref) applied with ${{\iota}}_{0i}=v_i=f_i(d_i-x_i'\theta_{0\tau})$ and noting that $1\{y_i \leqslant d_i\alpha_\tau + g_\tau(z_i)\} = 1\{U_i \leqslant \tau\}$ for some uniform $(0,1)$ random variable (independent of $\{d_i,z_i\}$) by the definition of the conditional quantile function.
Condition IQR(i) requires conditions on the probability density function that are assumed in Condition AS(iii). The fourth moment conditions are implied by Condition M(i) using $\xi=(1,0')'$ and $\xi=(1,-\theta_{0\tau}')'$, since $\bar {\mathrm{E}}[v_i^4]\leqslant \bar f^4\bar {\mathrm{E}}[ \{ (d_i,x_i')\xi\}^4] \leqslant C'(1+\|\theta_{0\tau}\|)^4$, and $\|\theta_{0\tau}\|\leqslant C$ assumed in Condition AS(i). Finally, by relation ((ref)), namely ${\mathrm{E}}[f_ix_iv_i]=0$, we have $\bar {\mathrm{E}}[f_id_iv_i] = \bar {\mathrm{E}}[v_i^2] \geqslant \underline f \bar {\mathrm{E}}[ (d_i-x_i'\theta_{0\tau})^2] \geqslant c\underline f\|(1,\theta_{0\tau}')'\|^2$ by Condition AS(iii) and Condition M(i).
Next we will construct the estimate for the orthogonal score function which are based on post-$\ell_1$-penalized quantile regression and post-Lasso with estimated conditional density function. We will show that with probability $1-o(1)$ the estimated nuisance parameters belong to $\overline{\mathcal{F}}$.
To establish the rates of convergence for $\widetilde\beta_\tau$, the post-$\ell_1$-penalized quantile regression based on the thresholded estimator $\widehat\beta_\tau^{\lambda_\tau}$, we proceed to provide rates of convergence for the $\ell_1$-penalized quantile regression estimator $\widehat\beta_\tau$. We will apply Lemma (ref) with $\gamma=1/n$. Condition PQR holds by Condition AS with probability $1-o(1)$ using Markov inequality and $\bar {\mathrm{E}}[r_{{\tau} i}^2]\leqslant s/n$.
Since population eigenvalues are bounded above and bounded away from zero by Condition M(i), by Lemma (ref) (where $\bar \delta_n\to 0$ under $K_q^2 Cs \log^2(1+Cs) \log (pn) \log n =o(n)$), we have that sparse eigenvalues of order $\ell_n s$ are bounded away from zero and from above with probability $1-o(1)$ for some slowly increasing function $\ell_n$. In turn, the restricted eigenvalue $\kappa_{2\mathbf{c}}$ is also bounded away from zero for bounded $\mathbf{c}$ and $n$ sufficiently large. Since $\lambda_\tau \lesssim \sqrt{n\log(p\vee n)}$, we will take $N = C\sqrt{s\log(n\vee p)/n}$ in Lemma (ref). To establish the side conditions note that
which implies that $N{\displaystyle \sup_{\delta \in A_\tau}}\frac{{\mathbb{E}_n}[|\tilde x_i'\delta|^3]}{\|\tilde x_i'\delta\|_{2,n}^{3}} \to 0$ with probability $1-o(1)$ under $K_q^2s^2\log^2(p\vee n) \leqslant \delta_n n$. Moreover, the second part of the side condition
Under $K_q^2s^2\log^2(p\vee n) \leqslant \delta_n n$, the side condition holds with probability $1-o(1)$. Therefore, by Lemma (ref) we have $\|\tilde x_i'\widehat \beta_\tau - \beta_\tau \|_{2,n} \lesssim \sqrt{ s \log(pn)/n }$ and $\|\widehat \beta_\tau - \beta_\tau \|_1 \lesssim s\sqrt{\log(pn)/n }$ with probability $1-o(1)$.
With the same probability, by Lemma (ref) with $\mu = \lambda_\tau/n$, since $\phi_{{\rm max}}(Cs)$ is uniformly bounded with probability $1-o(1)$, we have that the thresholded estimator $\widehat\beta^\mu_\tau$ satisfies the following bounds with probability $1-o(1)$: $\|\widehat \beta_\tau^\mu - \beta_\tau \|_1 \lesssim s\sqrt{\log(pn)/n }$, $\|\widehat\beta_\tau^\mu\|_0 \lesssim s$ and $\|\tilde x_i'(\widehat\beta^\mu_\tau - \beta_\tau)\|_{2,n} \lesssim \sqrt{ s \log(pn)/n }$. We use the support of $\widehat\beta^\mu_\tau$ as to construct the refitted estimator $\widetilde\beta_\tau$.
We will apply Lemma (ref). By Lemma (ref) and the rate of $\widehat\beta^\mu_\tau$, we can take $\widehat Q = C s\log(pn)/n$. Since sparse eigenvalues of order $Cs$ are bounded away from zero, we will use $\tilde N = C\sqrt{s\log(pn)/n}$ and $\varepsilon = 1/n$. Therefore, $\|x_i'(\widetilde\beta_\tau-\beta_\tau)\|_{2,n} \lesssim \sqrt{s\log(np)/n}$ with probability $1-o(1)$ provided the side conditions of the lemma hold. To verify the side conditions note that
and the side condition holds with probability $1-o(1)$ under $K_q^2 s \log^2(p\vee n) \leqslant \delta_n n$.
Next we proceed to construct the estimator for $v_i$. Note that by the same arguments above, under Condition D, we have the same rates of convergence for the post-selection (after truncation) estimators $(\widetilde \alpha_u,\widetilde \beta_u)$, $u\in \mathcal{U}$, $\|\widetilde\beta_u\|_0\lesssim C$ and $\|(\widetilde \alpha_u,\widetilde \beta_u)-(\widetilde \alpha_u,\widetilde \beta_u)\|\lesssim \sqrt{s\log (pn)/n}$, to estimate the conditional density function via ((ref)) or ((ref)). Thus with probability $1-o(1)$
where $\bar k$ depends on the estimator. (See relation ((ref)) below.) Note that the last relation implies that $\max_{i\leqslant n}\widehat f_i \leqslant 2\bar f$ is automatically satisfied with probability $1-o(1)$ for $n$ large enough. Let $\mathcal{U}$ denote the (finite) set of quantile indices used in the calculation of $\widehat f_i$.
Next we verify Condition WL to invoke Lemmas (ref) and (ref) with $c_r = C\sqrt{s\log(pn)/n}$ and $c_f= C\left( \{1/h\}\sqrt{s\log (n\vee p)/n}+h^{\bar k}\right)$. The sparsity condition in Condition WL(i) is implied by Condition AS(ii) and $\Phi^{-1}(1-\gamma/2p) \leqslant \delta_n n^{1/6}$ is implied by $\log (1/\gamma) \lesssim \log(p\vee n)$ and $\log^{3} p \leqslant \delta_n n$ from Condition M(iii). Condition WL(ii) follows from the assumption on the density function in Condition AS(iii) and the moment conditions in Condition M(i).
The first requirement in Condition WL(iii) holds with $c_r^2 \lesssim s\log(pn)/n$ since $$ {\mathbb{E}_n}[\widehat f_i^2r_{{\theta \tau} i}^2] \leqslant \max_{i\leqslant n}\widehat f_i {\mathbb{E}_n}[r_{{\theta \tau} i}^2] \lesssim \bar f s\log(pn)/n $$ from $\max_{i\leqslant n} \widehat f_i \leqslant C$ holding with probability $1-o(1)$, and Markov's inequality under $\bar {\mathrm{E}}[r_{{\theta \tau} i}^2]\leqslant Cs/n$ by Condition AS(ii).
The second requirement, $\max_{j\leqslant p}|({\mathbb{E}_n}-\bar {\mathrm{E}})[f_i^2x_{ij}^2v_i^2]|+|({\mathbb{E}_n}-\bar {\mathrm{E}})[\{f_ix_{ij}v_i-{\mathrm{E}}[f_ix_{ij}v_i]\}^2]| \leqslant \delta_n $ with probability $1-o(1)$, is implied by Lemma (ref) with $k=1$ and $K=\{{\mathrm{E}}[\max_{i\leqslant n} \|f_i x_iv_i\|_\infty^2]\}^{1/2} \leqslant \bar f \{{\mathrm{E}}[\max_{i\leqslant n} \|(x_i,v_i)\|_\infty^4]\}^{1/2}\leqslant \bar fK_q^{2}$, under $K_q^4 \log(pn) \log n \leqslant \delta_n n$.
The third requirement of Condition WL(iii) follows from uniform consistency in ((ref)) and the second part of Condition WL(ii) since $$ \max_{j\leqslant p}{\mathbb{E}_n}[(\widehat f_i - f_i)^2x_{ij}^2v_i^2] \leqslant \max_{i\leqslant n}\frac{|\widehat f_i-f_i|^2}{f_i^2} \left\{ \max_{j\leqslant p}({\mathbb{E}_n}-\bar {\mathrm{E}})[f_i^2x_{ij}^2v_i^2] + \max_{j\leqslant p}\bar {\mathrm{E}}[f_i^2x_{ij}^2v_i^2] \right\} \lesssim \delta_n $$ with probability $1-o(1)$ by ((ref)), the second requirement, and the bounded fourth moment assumption in Condition M(i).
To show the fourth part of Condition WL(iii), because both $\widehat f_i$ and $f_i$ are bounded away from zero and from above with probability $1-o(1)$, uniformly over $u\in \mathcal{U}$ (the finite set of quantile indices used to estimate the density), and $\widehat f_i^2-f_i^2 = (\widehat f_i-f_i)(\widehat f_i+f_i)$, it follows that with probability $1-o(1)$ $$ {\mathbb{E}_n}[(\widehat f_i^2-f_i^2)^2v_i^2/f_i^2] + {\mathbb{E}_n}[(\widehat f_i^2-f_i^2)^2v_i^2/\{\widehat f_i^2f_i^2\}] \lesssim {\mathbb{E}_n}[(\widehat f_i-f_i)^2v_i^2/f_i^2].$$ Next let $\delta_{u}=(\widetilde\alpha_u-\alpha_u,\widetilde \beta_{u}' - \beta_{u}')'$ where the estimators satisfy Condition D. We have that with probability $1-o(1)$
The following relations hold for all $u\in \mathcal{U}$ $$
$$ where we have $\|\delta_{u}\|\lesssim \sqrt{s\log (n\vee p) / n}$ and $\|\delta_{u}\|_0 \lesssim s$ with with probability $1-o(1)$. Then we apply Lemma \ref{thm:RV34} with $X_i=v_i\tilde x_i$. Thus, we can take $K=\{{\mathrm{E}}[\max_{i\leqslant n}\|X_i\|_\infty^2]\}^{1/2} \leqslant \{{\mathrm{E}}[\max_{i\leqslant n}\|(v_i,\tilde x_i')\|_\infty^4]\}^{1/2}\leqslant K_q^2$, and $\bar {\mathrm{E}}[(\delta'X_i)^2] \leqslant \bar {\mathrm{E}}[v_i^2(\tilde x_i'\delta)^2]\leqslant C\|\delta\|^2$ by the fourth moment assumption in Condition M(i) and $\|\theta_{0\tau}\|\leqslant C$. Therefore, {\small $$
$$} Under $K_q^4 s \log^{3}n \log (p\vee n) \leqslant \delta_n n$, with probability $1-o(1)$ we have $$ c_f^2 \lesssim \frac{s\log (n\vee p)}{h^2n} + h^{2{\bar k}}.$$
Condition WL(iv) pertains to the penalty loadings which are estimated iteratively. In the first iteration we have that the loadings are constant across components as $\varpi := \max_{i\leqslant n} \widehat f_i \max_{i\leqslant n}\|x_i\|_\infty \{{\mathbb{E}_n}[\widehat f_i^2 d_i^2]\}^2$. Thus the solution of the optimization problem is the same if we use penalty parameters $\tilde \lambda$ and $\widetilde \Gamma$ defined as $\tilde\lambda := \lambda \varpi/\max_{j\leqslant p} \{{\mathbb{E}_n}[ \widehat f_i^2x_{ij}^2v_i^2]\}^{1/2}$ and $\widetilde \Gamma_{jj} =\max_{j\leqslant p} \{{\mathbb{E}_n}[ \widehat f_i^2x_{ij}^2v_i^2]\}^{1/2}$. By construction $\widehat\Gamma_{0\tau jj} \leqslant \widetilde \Gamma_{jj} \leqslant C \widehat\Gamma_{0\tau jj}$ for some bounded $C$ with probability $1-o(1)$ as $\widehat\Gamma_{0\tau jj}$ are bounded away from zero and from above with probability $1-o(1)$. Since Condition WL holds for $(\tilde \lambda,\widetilde \Gamma)$, and $\tilde \lambda \lesssim \lambda \max_{i\leqslant n}\|x_i\|_\infty$, by Lemma (ref) we have with probability $1-o(1)$ $$
$$ The iterative choice of of penalty loadings satisfies $$
$$ uniformly in $j\leqslant p$. Thus the iterated penalty loadings are uniformly consistent with probability $1-o(1)$ and also satisfy Condition WL(iv) under $h^{-2}K_q^2s\log(pn)\leqslant \delta_n n$ and $K_q^4s\log(pn)\leqslant \delta_nn$. Therefore, in the subsequent iterations, by Lemma \ref{Thm:RateEstimatedLasso} and Lemma \ref{Thm:EstLassoPostRate} we have that the post-Lasso estimator satisfies with probability $1-o(1)$ $$
$$ where we used that $\phi_{{\rm max}}(\widetilde s_{\theta \tau}/\delta_n)\leqslant C$ implied by Condition D and Lemma \ref{thm:RV34}, and that $\lambda \geqslant \sqrt{n}\Phi^{-1}(1-\gamma/2p)\sim \sqrt{n\log (pn)}$ so that $ \sqrt{\widetilde s_{\theta \tau} \log (pn)/n} \lesssim (1/h)\sqrt{s\log (pn)/n} + h^{\bar k}.$
Next we construct a class of functions that satisfies the conditions required for $\overline{\mathcal{F}}$ that is used in Lemma (ref). Define the following class of functions.
where $\tilde f_i = f(d_i,z_i, \{\tilde \eta_u : u\in\mathcal{U}\}), \tilde{\tilde{f}}_i = f(d_i,z_i, \{\tilde{\tilde{\eta}}_u : u\in\mathcal{U}\}) \in \mathcal{J}$ are functions that satisfy
In particualr, taking $\tilde f_i := \widehat f_i \wedge 2\bar f$ where $\widehat f_i$ is defined in ((ref)) or ((ref)). (Due to uniform consistency ((ref)) the minimum with $2\bar f$ is not binding for $n$ large.) Therefore we have
We define the function class $\overline{\mathcal{F}}$ as $$ \overline{\mathcal{F}} = \{ ( \tilde g, \tilde v := \tilde f ( d- \tilde m ) ) : \tilde g \in \mathcal{G}, \tilde m \in \mathcal{M}, \tilde f \in \mathcal{J}\}$$ The rates of convergence and sparsity guarantees for $\widetilde\beta_\tau$ and $\widetilde\theta_\tau$, and Condition D implies that the estimates of the nuisance parameters $\widehat g_i := x_i'\widetilde\beta_\tau$ and $\widehat v_i=\widehat f_i(d_i-x_i'\widetilde\theta_{\tau})$ belongs to the proposed $ \overline{\mathcal{F}}$ with probability $1-o(1)$.
We will proceed to verify relations ((ref)) and ((ref)) hold under Condition AS, M and D. We begin with ((ref)). We have $$
$$ since $\|(1,-\theta)\|\leqslant C$ and $f_i\vee \tilde f_i \leqslant 2\bar f$, $\bar {\mathrm{E}}[(\tilde x_i'\xi)^4]\leqslant C\|\xi\|^4$, $\bar {\mathrm{E}}[(\tilde x_i'\xi)^2r_{\tau i}^2]\leqslant C\|\xi\|^2\bar {\mathrm{E}}[r_{\tau i}^2]$ and $\bar {\mathrm{E}}[r_{\tau i}^2]\lesssim s/n$ which hold by Conditions AS and M. Similarly we have $$
$$ as $|\tilde f_i - f_i| \leqslant 2\bar f$. Moreover we have that $$
$$ under (\ref{Eq:ExpDensity}), $h^{-2}s\log(pn) \leqslant \delta_n^2 n$, $h \leqslant \delta_n$, and $\lambda\sqrt{s} \leqslant \delta_n n$. The next relation follows from $$
$$ under $h^{-2}s^2\log^2(pn) \leqslant \delta_n^2 n$ and $h^{\bar k}\sqrt{s\log(pn)} \leqslant \delta_n$.
Finally, since $|\bar {\mathrm{E}}[f_iv_ir_{{\tau} i}]|\leqslant \delta_n n^{-1/2}$ from Condition M and $\bar {\mathrm{E}}[f_iv_ix_i]=0$ from ((ref)), we have $$ |\bar {\mathrm{E}}[f_iv_i\{\tilde g_i - g_{\tau i}\}]| \leqslant |\bar {\mathrm{E}}[f_iv_ix_i'(\tilde\beta - \beta_\tau)] |+|\bar {\mathrm{E}}[f_iv_ir_{\tau i}]| \leqslant \delta_n n^{-1/2}. $$
Next we verify relation ((ref)). By triangle inequality we have
Consider the first term of the right hand side in ((ref)). Note that $\mathcal{F}_1 := \{ (1\{y_i\leqslant d_i\alpha+g_{\tau i}\} - 1\{y_i\leqslant d_i\alpha+\tilde g_i\})v_i : \tilde g \in \mathcal{G}, |\alpha-\alpha_\tau| \leqslant \delta_n^2 \} \subset \mathcal{F}_{1a}-\mathcal{F}_{1b}$ where $\mathcal{F}_{1a}:=1\{y_i\leqslant d_i\alpha+g_{\tau i}\}v_i : |\alpha-\alpha_\tau| \leqslant \delta_n^2 \}$ is the product of a VC class of dimension 1 with the random variable $v$, and $\mathcal{F}_{1b}:=\{1\{y_i\leqslant d_i\alpha+\tilde g_{i}\}v_i : \tilde g \in \mathcal{G}, |\alpha-\alpha_\tau| \leqslant \delta_n^2 \}$ is the product of $v$ with the union of $\binom{p}{Cs}$ VC classes of dimension $Cs$. Therefore, its entropy number satisfies $N(\epsilon\|F_1\|_{Q,2},\mathcal{F}_1,\|\cdot\|_{Q,2}) \leqslant (A/\epsilon)^{C's}$ where we can take the envelope $F_1(y,d,x)=2|v|$. Since for any function $m \in \mathcal{F}_1$ we have $$
$$ from the conditional density function being uniformly bounded and the bounded fourth moment assumption in Condition M(i), by Lemma \ref{lemma:CCK} with $\sigma:= C\{s\log(pn)/n\}^{1/4}$, $a=pn$, $\|F_1\|_{P,2} = \bar {\mathrm{E}}[v_i^2]^{1/2} \leqslant C$, and $\|M\|_{P,2}\leqslant K_q$ we have with probability $1-o(1)$ $$\sup_{m\in \mathcal{F}_1}|({\mathbb{E}_n}-\bar {\mathrm{E}})[m]| \lesssim \frac{\{s\log(pn)/n\}^{1/4}}{n^{1/2}}\sqrt{s\log(pn)}+ n^{-1} s K_q \log(pn) \lesssim \delta_n n^{-1/2}$$ under $s^3\log^3(pn) \leqslant \delta_n^4 n$ and $K_q^2s^2\log^2(pn)\leqslant \delta_n^2 n$.
Next consider the second term of the right hand side in ((ref)). Note that $\mathcal{F}_2 := \{ (\tau - 1\{y_i\leqslant d_i\alpha+\tilde g_{i}\})(\tilde v_i-v_i) : (\tilde g,\tilde v) \in \overline{\mathcal{F}}, |\alpha-\alpha_\tau| \leqslant \delta_n^2 \} \subseteq \mathcal{F}_{2a} \cdot \mathcal{F}_{2b}$ where $\mathcal{F}_{2a} := \{ (\tau - 1\{y_i\leqslant d_i\alpha+\tilde g_{i}\}): \tilde g \in \mathcal{G}\}$ is a constant minus the union of $\binom{p}{Cs}$ VC classes of dimension $Cs$, and $\mathcal{F}_{2b} := \{ \tilde v_i - v_i : (\tilde g, \tilde v) \in \overline{\mathcal{F}} \}$. Note that standard entropy calculations yield $$ N(\epsilon \|F_{2a}F_{2b}\|_{Q,2}, \mathcal{F}_2 , \|\cdot\|_{Q,2}) \leqslant N(\mbox{$\frac{\epsilon}{2}$} \|F_{2a}\|_{Q,2}, \mathcal{F}_{2a} , \|\cdot\|_{Q,2}) \ N(\mbox{$\frac{\epsilon}{2}$}\|F_{2b}\|_{Q,2}, \mathcal{F}_{2b} , \|\cdot\|_{Q,2})$$ Furthermore, since $\tilde v_i - v_i = (\tilde f_i-f_i)v_i/f_i + \tilde f_ix_i'(\theta_{\tau 0}-\tilde \theta)$, we have $$\mathcal{F}_{2b} \subset \mathcal{F}_{2b}' + \mathcal{F}_{2b}'' := ( \mathcal{J} - \{f_i\} ) \cdot \{ v_i/f_i \} + \mathcal{J} \cdot (\{x_i'\theta_{\tau0}\} - \mathcal{M})$$ and $\mathcal{F}_2 \subset \mathcal{F}_{2a} \cdot \mathcal{F}_{2b}' + \mathcal{F}_{2a} \cdot \mathcal{F}_{2b}''$. By ((ref)), a covering for $\mathcal{J}$ can be constructed based on a covering for $\mathcal{B}_u:=\{ \tilde \eta_u : \|\tilde\eta_u\|_0 \leqslant Cs, \|\tilde\eta_u - \eta_u\|\leqslant C\sqrt{s\log(pn)/n}\}$ which is the union of $\binom{p}{Cs}$ sparse balls of dimension $Cs$. It follows that for the envelope $F_J:= 2\bar f \vee K_q^{-1}\|\tilde x_i\|_{\infty}$, we have $$ N(\epsilon \|F_J\|_{Q,2}, \mathcal{J},\|\cdot\|_{Q,2}) \leqslant \prod_{u\in \mathcal{U}} N\left( \mbox{$\frac{\epsilon h/4\bar f}{|\mathcal{U}|K_q \sqrt{2Cs}}$}, \mathcal{B}_u,\|\cdot\|\right)$$ For any $m \in \mathcal{F}_{2a}\cdot \mathcal{F}_{2b}'$ we have $$
$$ Therefore, by Lemma \ref{lemma:CCK} with $\sigma:= \frac{1}{h}\sqrt{s\log(pn)/n}+h^{\bar k}$, $a=pn$, $\|F_2'\|_{P,2} = \|(2\bar f\vee K_q^{-1}\|\tilde x_i\|_\infty)v_i/f_i\|_{P,2} \leqslant (1+2\bar f) K_q $, and $\|M\|_{P,2}\leqslant (1+2\bar f)K_q$ we have with probability $1-o(1)$ $$\sup_{m\in \mathcal{F}_{2a}\cdot \mathcal{F}_{2b}'}|({\mathbb{E}_n}-\bar {\mathrm{E}})(m)| \lesssim \frac{\frac{1}{h}\sqrt{s\log(pn)/n}+h^{\bar k}}{n^{1/2}}\sqrt{s\log(pn)}+ n^{-1} s K_q \log(pn) \lesssim \delta_n n^{-1/2}$$ under $h^{-2}s^2\log^2(pn) \leqslant \delta_n^2 n$, $h^{\bar{k}}\sqrt{s\log(pn)}\leqslant \delta_n$ and $K_q^2s^2\log^2(pn)\leqslant \delta_n^2 n$.
Moreover, $\mathcal{M}$ is the union of $\binom{p}{C\widetilde s_{{\theta \tau}}}$ VC subgraph of dimension $C\widetilde s_{{\theta \tau}}$, and for any $m \in \mathcal{F}_{2a}\cdot \mathcal{F}_{2b}''$ we have $$
$$ Therefore, by Lemma \ref{lemma:CCK} with $\sigma:= C\{\frac{1}{h}\sqrt{s\log(pn)/n}+h^{\bar k}+\lambda\sqrt{s}/n\}$, $a=pnK_qs$, $\|F_2''\|_{P,2} = \|(2\bar f\vee K_q^{-1}\|\tilde x_i\|_\infty)\|x_i\|_\infty\|_{P,2} \leqslant (1+2\bar f) K_q$, and $\|M\|_{P,2}\leqslant (1+2\bar f)K_q$ we have with probability $1-o(1)$ $$\sup_{m\in \mathcal{F}_{2a}\cdot \mathcal{F}_{2b}”}|({\mathbb{E}_n}-\bar {\mathrm{E}})(m)| \lesssim \frac{\frac{1}{h}\sqrt{\frac{s\log(pn)}{n}}+h^{\bar k}+ \frac{\lambda\sqrt{s}}{n}}{n^{1/2}}\sqrt{\widetilde s_{{\theta \tau}}\log(pn)}+ n^{-1} \widetilde s_{{\theta \tau}} K_q \log(pn) \lesssim \delta_n n^{-1/2}$$ under the conditions $h^{-2}s\widetilde s_{{\theta \tau}}\log^2(pn) \leqslant \delta_n^2 n$, $h^{\bar{k}}\sqrt{\widetilde s_{{\theta \tau}}\log(pn)}\leqslant \delta_n$, $\lambda\sqrt{s\widetilde s_{{\theta \tau}} \log(pn)}\leqslant \delta_n n$ and $K_q^2\widetilde s_{{\theta \tau}}^2\log^2(pn)\leqslant \delta_n^2 n$ assumed in Condition D.
Next we verify the second part of Condition IQR(iii), namely ((ref)). Note that since $\mathcal{A}_\tau \subset \{ \alpha : |\alpha - \alpha_\tau | \leqslant C / \log n + |\widetilde \alpha_\tau - \alpha_\tau| \leqslant C\sqrt{s\log(pn)/n}\}$, $\check\alpha_\tau \in \mathcal{A}_\tau$ and $s\log(pn) \leqslant \delta_n^2 n$ implies the required consistency $|\check\alpha_\tau-\alpha_\tau|\leqslant \delta_n$. To show the other relation in ((ref)), equivalent to with probability $1-o(1)$ $$| {\mathbb{E}_n}[\psi_{\check\alpha_\tau,\widehat h}(y_i,d_i,z_i)] |\leqslant \delta_n^{1/2}\ n^{-1/2},$$ note that for any $\alpha \in \mathcal{A}_\tau$ (since it implies $|\alpha-\alpha_\tau|\leqslant \delta_n$) and $\widehat h \in \overline{\mathcal{F}}$, we have with probability $1-o(1)$ $$
$$ from relations (\ref{Def:IdentityHelp}) with $\alpha$ instead of $\check\alpha_\tau$, (\ref{EqGamma20}), (\ref{BoundGamma2first}), and (\ref{BoundGamma2third}). Letting $\alpha^* = \alpha_\tau - \{\bar {\mathrm{E}}[f_id_iv_i]\}^{-1}{\mathbb{E}_n}[\psi_{\alpha_\tau, h_0}(y_i,d_i,z_i)]$, we have $| {\mathbb{E}_n}[\psi_{\alpha^*,\widehat h}(y_i,d_i,z_i)] | = O(\delta_n^{1/2} n^{-1/2})$ with probability $1-o(1)$ since $|\alpha^* - \alpha_\tau| \lesssim_P n^{-1/2}$. Thus, with probability $1-o(1)$ $$
$$ as ${\mathbb{E}_n}[\widehat v_i^2]$ is bounded away from zero with probability $1-o(1)$.
Next we verify Condition IQR(iv). The first condition follows from the uniform consistency of $\widehat f_i$ and $\max_{i\leqslant n}\|x_i\|_\infty \|\tilde \theta_\tau - \theta_{0\tau}\|_1 \lesssim_P K_q s \sqrt{\log(pn)/n}\to 0$ under $K_q^2s^2\log^2(pn)\leqslant \delta_n n$. The second condition in IQR(iv) also follows since { $$
$$} which implies the result under $K_q^2s^2\log(pn)\leqslant \delta_n n$.
The consistency of $\widehat\sigma_{1n}$ follows from $\|\widehat v_i-v_i\|_{2,n}\to_P 0$ and the moment conditions. The consistency of $\widehat\sigma_{3,n}$ follow from Lemma (ref). Next we show the consistency of $\widehat \sigma_{2n}^2=\{ {\mathbb{E}_n}[\check f_i^2 (d_i,x_{i\check T}')'(d_i,x_{i\check T }')]\}^{-1}_{11}$. Because $f_i\geqslant {\underline{f}}$, sparse eigenvalues of size $\ell_ns$ are bounded away from zero and from above with probability $1-\Delta_n$, and $\max_{i\leqslant n} |\widehat f_i - f_i| = o_P(1)$ by Condition D, we have $$ \{ {\mathbb{E}_n}[\widehat f_i^2 (d_i,x_{i\check T }')'(d_i,x_{i\check T }')]\}^{-1}_{11} = \{ {\mathbb{E}_n}[ f_i^2 (d_i,x_{i\check T }')'(d_i,x_{i\check T }')]\}^{-1}_{11} + o_P(1).$$ So that $\widehat\sigma_{2n}-\tilde \sigma_{2n}\to_P 0$ for $$\tilde \sigma_{2n}^2=\{ {\mathbb{E}_n}[ f_i^2 (d_i,x_{i\check T }')'(d_i,x_{i\check T }')]\}^{-1}_{11} = \{ {\mathbb{E}_n}[ f_i^2 d_i^2] - {\mathbb{E}_n}[ f_i^2 d_ix_{i\check T }']\{ {\mathbb{E}_n}[f_i^2 x_{i\check T }x_{i\check T }']\}^{-1}{\mathbb{E}_n}[ f_i^2 x_{i\check T } d_i] \}^{-1}.$$ Next define $\check\theta_\tau[\check T] = \{ {\mathbb{E}_n}[ f_i^2 x_{i\check T }x_{i\check T}']\}^{-1}{\mathbb{E}_n}[ f_i^2 x_{i\check T } d_i]$ which is the least squares estimator of regressing $f_i d_i$ on $f_ix_{i\check T }$. Let $\check \theta_\tau$ denote the associated $p$-dimensional vector. By definition $f_ix_i'\theta_\tau = f_id_i-f_ir_{{\theta \tau}}-v_i$, so that $$
$$ We have that $|{\mathbb{E}_n}[ v_i \{f_id_i-v_i\} ]|=|{\mathbb{E}_n}[ f_i v_ix_i'\theta_{0\tau} ]|=o_P(\delta_n)$ since $$\bar {\mathrm{E}}[ (v_i f_i x_i'\theta_{0\tau} )^2] \leqslant \bar f^2\{\bar {\mathrm{E}}[v_i^4]\bar {\mathrm{E}}[(x_i'\theta_{0\tau})^4]\}^{1/2}\leqslant C$$ and $\bar {\mathrm{E}}[ f_i v_ix_i'\theta_{0\tau}]=0$. Moreover, ${\mathbb{E}_n}[f_i d_i f_i r_{{\theta \tau} i}]\leqslant \bar f_i^2\|d_i\|_{2,n}\|r_{{\theta \tau} i}\|_{2,n}=o_P(\delta_n)$, $|{\mathbb{E}_n}[f_i d_i f_ix_i'(\check\theta-\theta_\tau)]| \leqslant \|d_i\|_{2,n}\|f_ix_i'(\check\theta_\tau-\theta_\tau )\|_{2,n} = o_P(\delta_n)$ since $|\check T|\lesssim \widetilde s_{{\theta \tau}} + s$ with probability $1-o(1)$ and ${\rm supp}(\widehat\theta_\tau)\subset \check T$.
\\ {\it Part 2. Proof of the Double Selection.} The analysis of $\widehat\theta_\tau$ and $\widehat \beta_\tau$ are identical to the corresponding analysis for Part 1. Let $\widehat T^*_\tau$ denote the set of variables used in the last step, namely $\widehat T^*_\tau = {\rm supp}(\widehat\beta_\tau^{\lambda_\tau})\cup {\rm supp}(\widehat\theta_\tau)$ where $\widehat\beta_\tau^{\lambda_\tau}$ denotes the thresholded estimator. Using the same arguments as in Part 1, we have with probability $1-o(1)$ $$|\widehat T^*_\tau| \lesssim \widehat s^*_\tau = s + \frac{ns\log p}{h^2\lambda^2}+\left( \frac{nh^{\bar k}}{\lambda} \right)^2. $$
Next we establish preliminary rates for $\check\eta_\tau := (\check\alpha_\tau,\check\beta_\tau')'$ that solves
where $\widehat f_i = \widehat f_i(d_i,x_i)\geqslant 0$ is a positive function of $(d_i,z_i)$. We will apply Lemma (ref) as the problem above is a (post-selection) refitted quantile regression for $(\tilde y_i,\tilde x_i)=(\widehat f_i y_i, \widehat f_i(d_i,x_i')')$. Indeed, conditional on $\{d_i,z_i\}_{i=1}^n$, quantile moment condition holds as $$
$$ Since $\max_{i\leqslant n}\widehat f_i \wedge \widehat f_i^{-1} \lesssim 1$ with probability $1-o(1)$, the required side conditions of Lemma \ref{Thm:MainTwoStepGeneric} hold as in Part 1. We can take $\bar r_\tau = C\sqrt{s\log(pn)/n}$ with probability $1-o(1)$. Finally, to provide the bound $\widehat Q$, let $\eta_\tau = (\alpha_\tau,\beta_\tau')'$ and $\widehat\eta_\tau=(\widehat\alpha_\tau,\widehat\beta^{\lambda_\tau}_\tau \ ')'$ which are $Cs$-sparse vectors. By definition ${\rm supp}(\widehat\beta_\tau) \subset \widehat T^*_\tau$ so that $${\mathbb{E}_n}[\widehat f_i\{\rho_\tau(y_i-(d_i,x_i')\check\eta_\tau) - \rho_\tau(y_i-(d_i,x_i')\eta_\tau)\}] \leqslant {\mathbb{E}_n}[\widehat f_i\{\rho_\tau(y_i-(d_i,x_i')\widehat\eta_\tau) - \rho_\tau(y_i-(d_i,x_i')\eta_\tau)\}] $$ By Lemma \ref{Lemma:UpperQ} we have with probability $1-o(1)$ $$ {\mathbb{E}_n}[\widehat f_i\{\rho_\tau(y_i-(d_i,x_i')\widehat\eta_\tau) - \rho_\tau(y_i-(d_i,x_i')\eta_\tau)\}] \leqslant \widehat Q:= C\frac{s\log (p n)}{n}.$$ Thus, with probability $1-o(1)$, Lemma \ref{Thm:MainTwoStepGeneric} implies $$ \| \widehat f_i (d_i,x_i')(\check\eta_\tau - \eta_\tau)\|_{2,n} \lesssim \sqrt{ \frac{(s + \widehat s^*_\tau)\log(pn)}{n\phi_{{\rm min}}(Cs+C\widehat s^*_\tau)}} $$
Since $s \leqslant \widehat s^*_\tau$ and $1/\phi_{{\rm min}}(Cs+C\widehat s^*_\tau)\leqslant C'$ by Condition D we have $ \|\check\eta_\tau-\eta_\tau\|\lesssim \sqrt{\frac{\widehat s^*_\tau\log p}{n}}$.
Next we construct an orthogonal score function based on the solution of the weighted quantile regression ((ref)). By the first order condition for $(\check\alpha_\tau,\check\beta_\tau)$ in ((ref)) we have for $s_i \in \partial \rho_\tau( y_i-d_i\check\alpha_\tau -x_i'\check\beta_\tau)$ that $$ {\mathbb{E}_n}\left[ s_i \widehat f_i\binom{ d_{i}}{x_{i \widehat T^*_\tau}} \right] = 0. $$ By taking linear combination of the selected covariates via $(1,-\widetilde\theta_\tau)$ and defining $\widehat v_i = \widehat f_i(d_{i}-x_{i \widehat T^*_\tau}'\widetilde\theta_\tau)$ we have $\psi_{\alpha,\widehat h}(y_i,d_i,z_i) = (\tau - 1\{y_{i} \leqslant d_i \alpha + x_{i}'\check \beta_\tau\})\widehat v_i$. Since $s_i = \tau - 1\{y_{i} \leqslant d_i \check \alpha_\tau + x_{i}'\check \beta_\tau\}$ if $y_i\neq d_i \check \alpha_\tau + x_{i}'\check \beta_\tau$, $$
$$ Since $\widehat s^*_\tau:= |\widehat T^*_\tau|\lesssim \widetilde s_{{\theta \tau}}$ with probability $1-o(1)$, and $y_i$ has no point mass, $ y_{i} = d_i \check \alpha_\tau + x_{i}'\check \beta_\tau$ for at most $C\widetilde s_{{\theta \tau}}$ indices in $\{1,\ldots,n\}$. Therefore, we have with probability $1-o(1)$
under $K_q^2 \widetilde s^2_{{\theta \tau}} \leqslant \delta_n n$. Moreover,
with probability $1-o(1)$ under $\sqrt{\widehat s^*_\tau}\|\widehat v_i- v_{i}\|_{2,n} \leqslant \delta_n$ holding with probability $1-o(1)$. Therefore, the orthogonal score function implicitly created by the double selection estimator $\check\alpha_\tau$ approximately minimizes $$ \widetilde L_n(\alpha) = \frac{| {\mathbb{E}_n}[ (\tau - 1\{y_i \leqslant d_i \alpha + x_i'\check\beta_\tau\})\widehat v_i ] |^2}{{\mathbb{E}_n}[\{ (\tau - 1\{y_i \leqslant d_i \alpha + x_i'\check\beta_\tau\})^2\widehat v_i^2 ]}, $$ and we have $\widetilde L_n(\check\alpha_\tau) \lesssim \delta_n^{1/3} n^{-1}$ with probability $1-o(1)$ by ((ref)) and ((ref)). Thus the conditions of Lemma (ref) hold and the double selection estimator has the stated linear representation.
{\bf $\square$}
In this section we collect auxiliary inequalities that we use in our analysis.
{\bf Proof.}{\bf \ (Proof of Lemma (ref))} Let $T_u = {\rm supp}(\eta_u)$. The first relation follows from the triangle inequality $$
$$
To show the second result note that $$\|\widehat\eta_u-\eta_u\|_{1} \geqslant \{|{\rm supp}(\widehat\eta_u^\mu)|-s\}\mu/\max_{j\leqslant p}{\mathbb{E}_n}[\tilde x_{ij}^2]^{1/2}.$$ Therefore, we have $|{\rm supp}(\widehat\eta_u^\lambda)|-s \leqslant \|\widehat\eta_u-\eta_u\|_{1}\max_{j\leqslant p}{\mathbb{E}_n}[\tilde x_{ij}^2]^{1/2}/\mu$ which yields the result.
To show the third bound, we start using the triangle inequality $$ \|\tilde x_i'(\widehat\eta^\mu_u-\eta_u)\|_{2,n} \leqslant \|\tilde x_i'(\widehat\eta^\mu_u-\widehat\eta_u)\|_{2,n} + \|\tilde x_i'(\widehat\eta_u-\eta_u)\|_{2,n}.$$ Without loss of generality assume that order the components is so that $|(\widehat\eta^\mu_u-\widehat\eta_u)_j|$ is decreasing. Let $T_1$ be the set of $s$ indices corresponding to the largest values of $|(\widehat\eta^\mu_u-\widehat\eta_u)_j|$. Similarly define $T_k$ as the set of $s$ indices corresponding to the largest values of $|(\widehat\eta^\mu_u-\widehat\eta_u)_j|$ outside $\cup_{m=1}^{k-1}T_m$. Therefore, $\widehat\eta^\mu_u-\widehat\eta_u=\sum_{k=1}^{\left\lceil p/s \right\rceil} (\widehat\eta^\mu_u-\widehat\eta_u)_{T_k}$. Moreover, given the monotonicity of the components, $\|(\widehat\eta^\mu_u-\widehat\eta_u)_{T_k}\|_{2,\varpi} \leqslant \|(\widehat\eta^\mu_u-\widehat\eta_u)_{T_{k-1}}\|_1/\sqrt{s}$. Then, we have $$
$$
where the last inequality follows from the first result and the triangle inequality.
{\bf $\square$}
The following result follows from Theorem 7.4 of delapena and the union bound.
{\bf Proof.} It follows from Theorem 3.6 of RudelsonVershynin2008, see BelloniChernozhukovHansen2011 for details. {\bf $\square$}
We will also use the following result of chernozhukov2012gaussian.