Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Quantile Graphical Models: Prediction and Conditional Independence with Applications to Systemic Risk
abstractWe propose two types of Quantile Graphical Models (QGMs) --- Conditional Independence Quantile Graphical Models (CIQGMs) and Prediction Quantile Graphical Models (PQGMs). CIQGMs characterize the conditional independence of distributions by evaluating the distributional dependence structure at each quantile index. As such, CIQGMs can be used for validation of the graph structure in the causal graphical models (pearl2009causality, robins1986new, heckman2015causal). One main advantage of these models is that we can apply them to large collections of variables driven by non-Gaussian and non-separable shocks. PQGMs characterize the statistical dependencies through the graphs of the best linear predictors under asymmetric loss functions. PQGMs make weaker assumptions than CIQGMs as they allow for misspecification. Because of QGMs’ ability to handle large collections of variables and focus on specific parts of the distributions, we could apply them to quantify tail interdependence. The resulting tail risk network can be used for measuring systemic risk contributions that help make inroads in understanding international financial contagion and dependence structures of returns under downside market movements.
We develop estimation and inference methods for QGMs focusing on the high-dimensional case, where the number of variables in the graph is large compared to the number of observations. For CIQGMs, these methods and results include valid simultaneous choices of penalty functions, uniform rates of convergence, and confidence regions that are simultaneously valid. We also derive analogous results for PQGMs, which include new results for penalized quantile regressions in high-dimensional settings to handle misspecification, many controls, and a continuum of additional conditioning events.
{\it Key Words}: High-dimensional graphs, conditional independence, prediction, inference, nonlinear
correlation, tail risk network, systemic risk, downside movement
Introduction
Co-movements, dependence and influence between variables are fundamental in economics and finance for decision and policy making as well as prediction. To this end, we propose Quantile Graphical Models (QGMs) as a modeling framework and consider their usefulness in three main applications. First, empirical auction models often rely on independent private values or affiliated private values, detecting collusion in these auctions is a form of testing conditional independence. Examples of studying entry, market power, or collusion can be found in bajari2003deciding, porter2005detecting, harrington2008detecting. Second, QGMs can be used to identify and measure systemic tail risk. The recent bank and sovereign crisis in the US and Europe have also boosted the interest in the important role of network spill-over effects in contagion and shaping systemic risk (Acemoglu2010, Acemoglu2013, elliott2014financial, hansen2014challenges). Many measures of systemic risk focus on spill-overs fit naturally in our QGM setting (Adrian2011, Andersen2013, hautsch2014forecasting, hautsch2015financial, hardle2016tenet). We apply these insights to re-evaluate international financial contagion in volatilities (Claessens2001).\footnote{Works on economic and financial networks include Billio2010, bonaldi2015empirical, kastl2017recent. We refer to de2017econometrics for an excellent review on the econometrics literature on networks.} Third, QGMs can measure dependence between stock returns for hedging strategies. In financial management settings, risk quantification is crucial, and advanced hedging decisions are typically focused on the tail of the distribution of stock returns rather than the mean. Moreover, such strategies aiming to reduce risk are critical precisely during the market downside movement. Empirical evidence (Ang2006,Ang2002,Patton2004) points to the non-Gaussianity of the distribution of stock returns, especially during market downturns. Therefore, it is also instructive to understand how dependence (and policy impact) would change as the downside movement of the market becomes more extreme. The proposed QGMs are flexible enough to cover all these cases.
QGMs can be viewed as part of graphical models which have been successively applied to estimate and visualize relationship (maathuis2018handbook,heckman2015causal). Graphical models are widely used in machine learning, statistical learning, and social science to model the statistical dependence among the components a $d$-dimensional random vector $X_{V}$, in the form of a graph or network $G=(V, E)$. Here $V$ is the node set contains the labels of the components and $E$ is the edge set represents unknown statistical relationships that need to be estimated, thus poses novel problems of statistical inference. In the case of Gaussian Graphical Models (GGMs), which assuming $X_{V}$ are jointly Gaussian distributed, the conditional independence structure is completely characterized by the support of the inverse of the covariance matrix of $X_{V}$. Notably in this case, the same graph will characterize conditional independence and the best linear prediction. However, in non-Gaussian settings, not only it is harder to characterize conditional independence, there are no reasons for the same graph to characterize both conditional independence and best linear prediction.
QGMs provide an alternative route to learn conditional independence and prediction under asymmetric loss functions which is appealing in non-Gaussian settings. As in non-Gaussian cases these notions do not coincide and there are needs for different estimation approaches. We propose two different QGMs to handle different types of applications. First, we propose Conditional Independence Quantile Graphical Models (CIQGMs) to characterize the conditional independence of distributions through evaluating the distributional dependence structure at each quantile index. Second, we propose Prediction Quantile Graphical Models (PQGMs) in which predictive relationship is the main focus. Note, QGMs also enable us to focus on specific parts of the distributions of variables, which play an important role in applications like financial contagion and measuring systemic risk contributions where extreme events are the main interests for practitioners.
CIQGMs can be used for validation of the graph structure in the causal graphical models (pearl2009causality, robins1986new, heckman2015causal). Conditional independence has a long history in statistical models with consequences towards parameter identification, causal inference, prediction sufficiency, and many others, see dawid1979conditional. CIQGMs aim to characterize conditional independence via the conditional quantile functions. In such models, we consider a flexible specification that can approximate well the conditional quantile functions (up to a vanishing approximation error). In turn, this allows detecting which variables have a strong or near zero impact on others which can then be used to provide guidance on conditional independence.
PQGMs focus on the prediction of a variable based on linear combinations of other variables (a reduced form relation) under asymmetric losses. An important motivation for proposing PQGMs is to allow for misspecification as the conditional quantile function is typically non-linear in non-Gaussian settings. The linear specification is widely used in practice despite possible misspecification which motivates an analysis for accomodating these issues. We characterize the uniform prediction properties under a family of asymmetric loss functions, which this family enables practitioners to investigate different parts of the tail distribution. Other papers investigated the impact of misspecification on quantile functions are abadie2014inference, angrist2006quantile, knight2008asymptotics, lee2009efficiency. Our analysis also contributes to the high-dimensional quantile regression by allowing non-vanishing misspecification.
Broadly speaking, QGMs enhance our understanding of statistical dependence among $X_V$. For example, for each quantile index $\tau$, they provide visualization of the dependence via graphs whose edges represent conditional (quantile) relationships. Given that for each specific quantile index $\tau$ we will obtain one such graph, we could have a graph process indexed by $\tau \in (0,1)$. The structure represented by a $\tau$-quantile graph represents a local relation and can be valuable in cases where tail interdependence might be of special interest.\footnote{This is similar to the contrast between quantile regression and linear regression, where the latter provides information only on the conditional mean, while the former can provide a more complete description of the distribution of the outcome.} The graph process induced by QGMs has several important features. First, a $\tau$-quantile graph enables different values of edge strength in different directions. This is important because for undirected networks, the distinction is unclear. Second, QGMs can capture the tail interdependence through estimating at a high or low quantile index. \footnote{Examples of high or low quantile index can be $\tau = 0.95$ or $\tau = 0.05$ respectively. The analysis extends to the case of $s^3\log^5p = o(n \tau (1-\tau))$. The case of $n \tau = C$, even in the fixed dimension case, leads to a substantially different analysis and limiting distributions, as shown in chernozhukov2005extremal. Here $s$ is the sparsity parameter, $p$ is the dimension of conditional variables, $n$ is the sample size, $C$ is a constant.} Third, QGMs can capture the asymmetric dependence structure at different
quantiles, which can be particularly useful in empirical applications (e.g., stock market dependence, exchange rate dependence). By considering
all the quantiles at once we can characterize conditional independence structure for a set of variables that are not jointly Gaussian.
We also provide and study the estimation procedures that allow us to learn QGMs from the observed data. Our techniques are geared for covering high-dimensional settings where the size of the model is potentially larger than the sample size. These techniques are based on $\ell_{1}$-penalized quantile regression and Neyman orthogonal equations. For CIQGMs, under
mild regularity conditions, we provide rates of convergence and edge properties of the estimated graph that hold uniformly over a large class of data generating processes. We provide simultaneously valid confidence regions (post-selection) for the coefficients of the CIQGM that are uniformly valid, despite of possible model selection mistakes. Based on proper thresholding, recovery of CIQGMs patterns is possible when coefficients are well separated from zero which parallel the results for graph recovery in the Gaussian case.\footnote{Similar to graph recovery in the Gaussian
case such exact recovery is subject to the lack of uniformity validity critiques of
Leeb and P\"otscher leeb2008can.} For PQGMs, we provide an estimator that achieves an adaptive rate of convergence, which might differ under different conditioning events. Therefore we contribute to the recent active literature on simultaneous valid confidence regions post-model selection, BCCH2012sparse,BCH2014inference,BCFH2013program,Farrell:JMP,CanerKock:HDCI,CHS:PnP GRD2014,zhang2014confidence,belloni2015uniform,BCK2013robustQR,jankova2015confidence,ning2017general,wang2016inference;
in particular, the penalty choices and theoretical results are uniformly valid and adaptive to the relevant conditioning events.
Although we build upon the quantile regression literature (Koenker2005,BCK2013robustQR), we derive new results for penalized quantile regression in high dimensional settings that are uniformly valid, robust to small coefficients (e.g. allowing for model selection mistakes), allow possibly non-vanishing misspecification, many controls and a continuum of additional conditioning events. These results contribute to a growing literature that relies on quantile based models to characterize the data generating process. zheng2015globally considers a globally adaptive quantile regression model, establishes oracle properties, and improved rates of convergence for the high-dimensional case. Screening procedures based on moment conditions motivated by the quantile models have been proposed and analyzed in he2013quantile, wu2015conditional in the high-dimensional case. joe1997multivariate considers tail dependence defined via conditional probabilities in a low dimensional setting.
Finally, we view QGMs as complementary to a large body of works on GGMs (Dempster1972,Lauritzen1996,Cox1996,Edwards2000,Drton2004,Drton2007,Drton2008, meinshausen2006high, Yuan2007,Banerjee2008,Friedman2008, Yuan2010,Cai2011,Liu2012a,Sun2012,Liu2012b,chiong2017estimation). Our work is also complementary to other works trying to relax the joint Gaussian assumption. Liu2009,Liu2012,Xue2012 work with so-called nonparanormal models or semiparametric Gaussian copula models, i.e., the variables follow a joint Gaussian distribution after monotone transformations. Ravikumar2011, peng2009partial work with Sub-Gaussian data, which restricted fatness of the tails. ravikumar2010high, Xue2012a, loh2012structure work with discrete-valued random variable, and few types of exponential families. yang2015graphical provides results for M-estimators for a subclass of exponential family graphical models. QGMs allow for different sets of distributions.
The rest of the paper is organized as follows. Section (ref) provides main motivating examples. Section (ref) lays out the foundation of the conceptual framework of QGMs. Section (ref) contains estimators for QGMs while Section (ref) contains the theoretical guarantees of the estimators. Section (ref) provides an empirical application of QGMs to measure systemic risk contribution. Finally, the appendix contains proofs, simulations, and implementation details of the estimators.
{\bf Notation}. For an integer $k$, we let $[k]:=\{1,\ldots,k\}$ denote the set of integers from $1$ to $k$. For a random variable $X$ we denote by $\mathcal{X}$ its support. We use the notation $a\vee b=\max\{a,b\}$ and $a\wedge b=\min\{a,b\}$. We use $\|v\|_{p}$ to denote the $p$-norm of a vector $v$. In particular, the $\ell_2$-norm is denoted by $\Vert \cdot \Vert$; the $\ell_{0}$-“norm” $\Vert\cdot\Vert_{0}$ denotes the number of non-zero components. Given a vector $\delta\in\mathbb{R}^{p}$, and a set of indices $T\subset\{1,...,d\}$, we denote by $\delta_{T}$ the vector in which $\delta_{Tj}=\delta_{j}$ if $j\in T$, $\delta_{Tj}=0$ if $j\notin T$. We use $\mathbb{E}_{n}$ to abbreviate the notation $n^{-1}\sum_{i=1}^n$; for example, $\mathbb{E}_{n}[f]:=\mathbb{E}_{n}[f(\omega_i)]:=n^{-1}\sum_{i=1}^nf(\omega_i)$.
Motivating Examples
Screening for Collusion
One application of CIQGM is in the empirical auction literature on the examination of entry or bidding prices patterns to detect coordinating groups. As shown in bajari2003deciding, firms’ bids, after controlling for all information about costs, are jointly independent under the competitive model and lack of independence is taken as evidence consistent with collusion. Although, collusion is only one alternative explanation, testing conditional independence can be viewed as a screening device to determine whether further investigation is warranted.
Mathematically, denote $Y_{a,t}$ as the amount bid by firm $a$ on project $t$, and $Z_{a,t}$ as covariates observed in dataset. We define $X_{a,t}$ to be residuals after projecting out $Z_{a,t}$. bajari2003deciding test whether $X_{a}$ is independent of $X_{b}$, using Fisher's Z-transformation of the coefficient of correlation between $X_{a}$ and $X_{b}$. For two Gaussian random variables, this is equivalent to test pairwise independence. In our terminology, this corresponds to an edge $(a,b)$ not being contained in the graph if and only if
equation[equation omitted — 33 chars of source]
Pairwise independence, however, needs not imply joint independence. CIQGM can also handle joint independence in the non-Gaussian setting. This is because CIQGM works with the following case
equation[equation omitted — 70 chars of source]
namely an edge $(a,b)$ is not
contained in the graph if and only if $X_{b}$ and $X_{a}$ are independent conditional on all remaining
variables $X_{V\backslash\{a,b\}}=\{X_{k};k\in V\backslash\{a,b\}\}$, via the equivalence between conditional probabilities and conditional quantiles (details can be found in Section (ref)).
Systemic Tail Risk
Measuring systemic risk taking into account tail risk network spillover effects is another application of QGMs, e.g. our framework complements to the systemic risk measure CoVaR Adrian2011 which ignore tail risk dependencies induced by the underlying financial network structure. Our framework also allows for large scale networks.
Traditional tail risk measures, such as Value of Risk (VaR), focus on the loss of an individual institution. CoVaR attempts to measure the VaR of the whole financial system or a particular financial institution
by conditioning on another being in distress. Formally, Adrian2011 define institution $b$'s CoVaR at level $\tau$ conditional
on a particular outcome of institution $a$, as the value of $CoVaR_{\tau}^{b\vert a}$
that solves
equation[equation omitted — 100 chars of source]
for some event $ \mathbb{C}(X_{a}) $ based on $X_{a}$. A special case is $\mathbb{C}(X_{a})=\{ X_a = VaR_{\tau}^{a}\}$ which, as interpreted by Adrian2011, means with probability $\tau$ institution $b$ is in trouble given that institution $a$ is in trouble.
QGMs can work with the case
equation[equation omitted — 137 chars of source]
with the main difference here is the conditioning events (or variables), i.e. from $\mathbb{C}(X_{a})$ to $\mathbb{C}(X_{a}, X_{V\backslash\{a,b\}})$ with the latter could be high dimensional. Hence, our QGMs take into account risk spillovers from other institutions driving the CoVaR. The identified risk spillovers between all financial institutions constitute a financial tail risk network which, as shown later, can be used for measuring institutions' systemic tail risk contributions. In summary, QGMs can take into account the system-wide network spillover effects via incorporating tail network spillover effects into risk measuring, thus relate systemic risk to tail spillover effects from individual institutions to the whole system.
Another related definitions of tail risk, $\Delta CoVaR$, is defined as the change in the VaR of the whole financial system conditional on a institution being under distress relative to its median state. In terms of estimation, replacing covariate $X_{a}$ by the difference between its $\tau$-th quantile (denoted as $VaR_{\tau}^{a}$), and its median (denoted as $VaR_{50\%}^{a}$), yields
equation[equation omitted — 106 chars of source]
where $\widehat{\beta}_a^{b}(\tau)$ comes from pairwise quantile regression of $X_b$ on $X_a$, and
equation[equation omitted — 147 chars of source]
where ${\check{\beta}}^{b}(\tau)$
is estimated via Algorithm (ref) or (ref), inference procedures are based on Corollary (ref).
After learning a tail risk network, we can use our new network-cooperated $\Delta CoVaR$ to measure the systemic risk contribution of each institution. The systemic risk contribution of institution $a$ can be measured by it "to" or "from" degree, see Andersen2013; or by other network centrality measures, see jackson2010social, Kolaczyk2009, newman2010networks.
To-degrees measure contributions of individual institutions to the overall risk
of systemic network events, for institution $a$ is defined as $\delta_{a}^{to} = \sum_{k}\Delta CoVaR_{\tau}^{k\vert a,V\backslash\{a,k\}}$. From-degrees measure exposure of individual institutions to systemic shocks from
the network, for institution $a$ it is defined as $\delta_{a}^{from} = \sum_{k}\Delta CoVaR_{\tau}^{a\vert k,V\backslash\{a,k\}}$. The net contribution of institution $a$ is defined as net-$\Delta CoVaR^{a}=\mbox{\ensuremath{\delta}}_{a}^{to}-\delta_{a}^{from}$.
In Section (ref), we revisit the analysis of international financial contagion through the volatility spillovers perspective. We visualize the tail risk interdependence via PQGMs as they allow for heteroskedasticity and asymmetric responses, can be used to model nonlinear tail interdependence, and to visualize potential asymmetric changes in conditional correlations. The estimated contagion network taking into account global interconnectedness is important for Eurozone financial regulators who want to identify globally systemically important EU countries, or for global financial portofolio diversification. Our the systemic tail risk analysis tools mentioned in previous paragraphs, can help with achieve those goals.
Stock Returns Under Market Downside Movements
Hedging decisions rely on the dependence of various stocks returns. Moreover, hedging is even more relevant during market downside movements, which motivates us to understand interdependence conditional on those events. Stock returns are in general non-Gaussian in those settings, as shown in the empirical finance literature, e.g. Ang2002, Longin2001,Patton2006. We can parameterize the downside movements by using a random variable $W$, which could be the market index, and conditional on the event $\Omega_\varpi = \{W \leqslant \varpi\}$. This allows us to define a $\varpi$-conditional-CIQGM as $G^I(\tau,\varpi)=(V,E^I(\tau,\varpi))$ and a $\varpi$-conditional-PQGM as $G^P(\tau,\varpi)=(V,E^P(\tau,\varpi))$, for each $\varpi \in \mathcal{W}$. We might be interest in a fixed $\varpi$ or on a family of values $\varpi\in(-\bar \varpi, 0]$. The latter induces $\mathcal{W} = \{\Omega_\varpi=\{W\leqslant \varpi\} : \varpi\in(-\bar \varpi, 0] \}$.
Figure (ref) provides an example using $ \mathcal{W}$-Conditional QGM with
\sloppy$W \text{= \{Market Index Returns\}}$, $\varpi$ as the $\tau_m$-th quantile of the
market index returns, and $\tau_m = \{0.15, 0.5, 0.75, 0.9\}$. We obtain daily stock returns from CRSP and use S&P 500 as the market index. The full sample consists of 2769 observations for 86 stocks from Jan 2, 2003 to December 31, 2013. The total number of stocks is 86 due to data availability. We define market movement as when the market index returns
are below a pre-specified level (e.g. $\tau_m$-th quantile), hence conditioning on a particular $\varpi$ corresponds to consider the subsample based on whether the corresponding date's market return is less equal to the $\tau_m$-th quantile of the
market index returns. The results show higher interdependence under market downside moments, pose different hedging decisions.
figure[figure omitted — 1,104 chars of source]
Quantile Graphical Models
In this section we describe quantile graphical models associated with a $d$-dimensional random vector $X_V$ where the set $V=[d]=\{1,\ldots,d\}$ denotes the labels of the components. These models aim to provide a description of the dependence between the random variables in $X_V$. In particular, these models induce graphs that allow for visualizing dependence structures. Nonetheless, because of the non-Gaussianity, we consider two fundamentally distinct models (one geared towards conditional independence and one geared towards prediction).
Conditional Independence Quantile Graphical Models
Conditional independence graphs have been used to provide visualization and insight on the dependence structure between random variables. Each node of the graph is associated with a component of $X_V$. We denote the conditional independence graph as $G^I=(V,E^I)$ where $G^I$ is an undirected graph with vertex set $V$ and edge set $E^I$ which is represented by an adjacency matrix ($E^I_{a,b}=1$ if the edge $(a,b)\in G^I$, and $E^I_{a,b}=0$ otherwise). An edge $(a,b)$ is not
contained in the graph if and only if
equation[equation omitted — 89 chars of source]
namely $X_{b}$ and $X_{a}$ are independent conditional on all remaining
variables $X_{V\backslash\{a,b\}}=\{X_{k};k\in V\backslash\{a,b\}\}$.
remark[Conditional Independence Under Gaussianity] In the case that $X_V$ is jointly Gaussian distributed,
$X_V\sim N(0,{\Sigma})$ with $\Sigma$ as
the covariance matrix of $X_V$, the conditional independence
structure between two components is determined by the inverse of the covariance
matrix, i.e. the precision matrix $\Theta=\Sigma^{-1}$. It follows that the non-zero elements in the precision matrix corresponds
to the non-zero coefficients of the associated (high dimensional) mean
regression. The family of Gaussian distributions with this property
is known as a Gauss-Markov random field with respect to the graph
$G$. This observation has motivated a large literature Lauritzen1996 and interesting extensions that allow for transformations of Gaussian variables Liu2009,Liu2012.
In order to achieve a tractable concept for non-Gaussian settings, we use that
((ref))
occurs if and only if
equation[equation omitted — 193 chars of source]
In turn, by the equivalence between conditional probabilities and conditional quantiles to characterize a random variable, we have that ((ref))
occurs if and only if
equation[equation omitted — 264 chars of source]
For a quantile index $\tau \in (0,1)$, the $\tau$-quantile conditional independence graph is
a directed graph $G^I(\tau)=(V,E^I(\tau))$ with vertex set $V$ and edge set
$E^I(\tau)$. An edge $(a,b)$ is not contained in the edge set $E^I(\tau)$ if and only if
equation[equation omitted — 231 chars of source]
By the equivalence between ((ref)) and ((ref)), the union of $\tau$-quantile graphs over $\tau\in(0,1)$ represents
the conditional independence structure of $X$, namely $E^I=\cup_{\tau \in (0,1)} E^I(\tau)$. We also consider a relaxation of ((ref)). For a set of quantile indices $\mathcal{T}\subset (0,1)$, we say that
equation[equation omitted — 100 chars of source]
$X_{a}$ and $X_{b}$ are $\mathcal{T}$-conditionally independent given $X_{V\backslash\{a,b\}}$, if ((ref)) holds for all $\tau \in \mathcal{T}$. Thus, we have that ((ref)) implies ((ref)).
We define the $\mathcal{T}$-quantile graph as $G^I(\mathcal{T})=(V,E^I(\mathcal{T}))$ where
$$ E^I(\mathcal{T})= \cup_{\tau \in \mathcal{T}}E^I(\tau).$$
Although the conditional independence concept relates to all quantile indices, the quantile characterization described above also lends itself to quantile specific impacts which can be of independent interest.\footnote{For example, we might be interested in some extreme events which typically correspond to crises in financial systems.}
Prediction Quantile Graphical Models
Prediction Quantile Graphical Models (PQGMs) are motivated by prediction accuracy under an asymmetric loss function (instead of conditional independence as in Section (ref)). More precisely, for each $a\in V$, we are interested in predicting $X_a$ based on linear combinations of the remaining variables, $X_{V\backslash\{a\}}$, where accuracy is measured with respect to an asymmetric loss function. Formally, PQGMs measure accuracy as
equation[equation omitted — 130 chars of source]
where $X_{-a}=(1,X_{V\backslash\{a\}}')'$, and the asymmetric loss function $\rho_{\tau}(t)=(\ensuremath{\tau-1\{t\leqslant0\}})t$ is the check function used in KoenkerBassett1978.
Importantly, PQGMs are concerned with the best linear predictor under the asymmetric loss function $\rho_\tau$ which is a specification that is widely used in practice. This is a fundamental distinction with respect to CIQGMs discussed in Section (ref) where the specification of the conditional quantile was approximately a linear function of transformations of $X_{V\backslash \{a\}}$.\footnote{In Section (ref) the vector $Z^a$ in equation ((ref)) collects the functions of the vector $X_{V\backslash \{a\}}$.} Indeed, we note that under suitable conditions the linear predictor that solves the minimization problem in ((ref)) approximates the conditional quantile regression as shown in BCF2011. (In fact, the conditional quantile function would be linear if $X_V$ was jointly Gaussian distributed.) However, PQGMs do not assume that the conditional quantile function of $X_a$ is well approximated by a linear function and instead it focuses on the best linear predictor.
We define that $X_b$ is predictively uninformative for $X_a$ given $X_{V\backslash\{a,b\}}$ if
$$ \mathcal{L}_a(\tau \mid V\backslash\{a\}) = \mathcal{L}_a(\tau \mid V\backslash\{a,b\}) \ \ \ \mbox{for all} \ \ \tau \in (0,1),$$
i.e., considering a linear function of $X_b$ will not improve our performance of predicting $X_a$ with respect to the asymmetric loss function $\rho_\tau$.
Again we can visualize the predictive relationship through a graph process indexed by $\tau \in (0,1)$. That is, for each $\tau \in (0,1)$ we have a directed graph $G^P(\tau)=(V,E^P(\tau))$, where an edge $(a,b)\in G^P(\tau)$ only if $X_b$ is predictively informative for $X_a$ given $X_{V\backslash\{a,b\}}$ at the quantile $\tau$. Finally, it is also convenient to define the PQGM associated with a subset $\mathcal{T}\subset(0,1)$ as $G^P(\mathcal{T})=(V,E^P(\mathcal{T}))$ where $$E^P(\mathcal{T})=\cup_{\tau \in \mathcal{T}}E^P(\tau).$$
$\mathcal{W}$-Conditional Quantile Graphical Models
In what follows, we discuss an extension of the QGMs discussed in Sections (ref) and (ref) to allows for conditioning on a (possible infinity) family of events $\varpi \in \mathcal{W}$.\footnote{With a slight abuse of notation, we let $\varpi$ to denote the event and also the index of such event. For example, we write ${\mathrm{P}}(\varpi)$ as a shorthand for ${\mathrm{P}}( W \in \Omega_\varpi)$.} Such extension is motivated by several applications in which the interdependence between the random variables in $X_V$ maybe substantially impacted by additional observable events (e.g. downside movements of the market). This general framework allows different forms of conditioning. The main implication of this extension is that QGMs are now graph processes indexed by $\tau \in \mathcal{T} \subset (0,1)$ and $\varpi \in \mathcal{W}$.
We define $X_a$ and $X_b$ are
$(\mathcal{T},\varpi)$-conditionally independent,
equation[equation omitted — 111 chars of source]
if for all $\tau \in \mathcal{T}$ we have
equation[equation omitted — 167 chars of source]
The conditional independence edge set associated with $(\tau,\varpi)$ is defined analogously as before. We denote them by $E^I(\tau,\varpi)$ and $E^I(\mathcal{T},\varpi) = \cup_{\tau \in \mathcal{T}} E^I(\tau,\varpi)$ for each $\varpi \in \mathcal{W}$.
The extension of PQGMs proceeds by defining the accuracy under the asymmetric loss function conditionally on $\varpi$. More precisely, we define
equation[equation omitted — 156 chars of source]
The prediction edge set associated with $(\tau,\varpi)$ is also defined analogously as before. We denote them by $E^P(\tau,\varpi)$ and $E^P(\mathcal{T},\varpi) = \cup_{\tau \in \mathcal{T}} E^P(\tau,\varpi)$, for each $\varpi \in \mathcal{W}$.
Estimators for High-Dimensional Quantile Graphical Models
In this section, we propose and discuss estimators for QGMs introduced in Section (ref). Throughout it is assumed that we observe a $d$-dimensional i.i.d. random vector $X_V$, namely $\{X_{iV} : i=1,\ldots,n\}$. Based on the data observed, unless additional assumptions are imposed we cannot estimate the quantities of interest for all $\tau \in (0,1)$. Instead, in what follows we will consider a (compact) set of quantile index $\mathcal{T}\subset (0,1)$. The estimators are intended to handle high dimensional models and a continuum of conditioning events in $\mathcal{W}$.
Estimators for CIQGMs
We discuss the specification and propose an estimator for CIQGMs. Although in general it is potentially hard to correctly specify coherent models, the following are simple examples.
example[Multivariate Gaussian Distribution]
Consider the Gaussian case, $X_V \sim N(\mu,\Sigma)$. It follows that for each $a\in V$, the conditional distribution $X_a\mid X_{V\backslash \{a\}}$ satisfies $$X_a\mid X_{V\backslash \{a\}} \sim N\left(\mu_a - \sum_{j\in V\backslash\{a\}} \frac{(\Sigma^{-1})_{aj}}{(\Sigma^{-1})_{aa}}(X_j-\mu_j), \frac{1}{(\Sigma^{-1})_{aa}}\right).$$
Therefore the conditional quantile function of $X_a$ is linear in $X_{V\backslash \{a\}}$ and is given by $$Q_{X_a}(\tau\vert X_{V\backslash \{a\}}) = \frac{\Phi^{-1}(\tau)}{(\Sigma^{-1})_{aa}^{1/2}}+\mu_a -\sum_{j\in V\backslash\{a\}} \frac{(\Sigma^{-1})_{aj}}{(\Sigma^{-1})_{aa}}(X_j-\mu_j).$$
example[Multivariate $t$-Distribution]
Consider the multivariate $t$ distribution case, $X_{V}\sim t_{p}(\mu,\Sigma,v)$,
with location $\mu$, scale matrix $\Sigma$, and degrees of freedom
$v$, as in ding2016conditional. It follows that for each
$a\in V$, the conditional distribution $X_{a}\vert X_{V\backslash\{a\}}$
satisfies
$$
X_{a}\mid X_{V\backslash\{a\}}\sim t_{p-1}\left(\mu_{a}-\sum_{j\in V\backslash\{a\}}\frac{(\Sigma^{-1})_{aj}}{(\Sigma^{-1})_{aa}}(X_{j}-\mu_{j}),\frac{v+d_{1}}{v+1}\frac{1}{(\Sigma^{-1})_{aa}},v+1\right).
$$
Therefore the conditional quantile function of $X_{a}$ is given by
$$Q_{X_a}(\tau\vert X_{V\backslash \{a\}})=\sqrt{\frac{v+d_{1}}{v+1}}\frac{F_{t_{v+1}}^{-1}(\tau)}{(\Sigma^{-1})_{aa}^{1/2}}+\mu_{a}-\sum_{j\in V\backslash\{a\}}\frac{(\Sigma^{-1})_{aj}}{(\Sigma^{-1})_{aa}}(X_{j}-\mu_{j}),$$
here $d_{1}=(X_{V\backslash\{a\}}-\mu_{V\backslash\{a\}})^{T}\Sigma_{aa}^{-1}(X_{V\backslash\{a\}}-\mu_{V\backslash\{a\}})$.
example[Multiplicative Error Model]
Consider $d=2$ so that $V=\{1,2\}$. Assume that $X_2$ and $\varepsilon$ are independent positive random variables. Assume further that they relate to $X_1$ as $$X_1 = \alpha + \varepsilon X_2.$$ In this case, we have that the conditional quantile functions are linear and given by $$Q_{X_1}(\tau\vert X_2) = \alpha + F^{-1}_\varepsilon(\tau) X_2 \ \ \ \mbox{and}\ \ \ Q_{X_2}(\tau\vert X_1) = (X_1 -\alpha)/F^{-1}_\varepsilon(1-\tau).$$
example[Additive Error Model]
Consider $d=2$ so that $V=\{1,2\}$. Let $X_2 \sim U(0,1)$ and $\varepsilon\sim U(0,1)$ be independent random variables. Also define the random variable $X_1$ as $$X_1 = \alpha + \beta X_2 + \varepsilon.$$ It follows that $Q_{X_1}(\tau\vert X_2) = \alpha + \beta X_2+\tau$. However, if $\beta = 0$, we have $ Q_{X_2}(\tau\vert X_1)=\tau$, and for $\beta > 0$, direct calculations yield that
$$ Q_{X_2}(\tau\vert X_1)= \left\{\begin{array}{l}
\frac{\tau}{\beta}(X_1-\alpha), \ \ \mbox{if } \ \ X_1 \leqslant \alpha + \beta\\
\tau + (1-\tau)(X_1-\alpha-\beta),\ \ \mbox{if } \ \ X_1 \geqslant \alpha + \beta\end{array}\right. $$ where we note that $X_1 \in [ \alpha, 1+\alpha+\beta]$.
example[Mixture of Gaussians]
Similar to the prior example, consider the case $X_V \mid \varpi \sim N(\mu_\varpi,\Sigma_\varpi)$ for each $\varpi \in \mathcal{W}$. It follows that for $a\in V$, the conditional distribution satisfies $$X_a\mid X_{V\backslash \{a\}}, \varpi \sim N\left(\mu_{\varpi a} - \sum_{j\in V\backslash\{a\}} \frac{(\Sigma^{-1})_{\varpi aj}}{(\Sigma^{-1})_{\varpi aa}}(X_j-\mu_{\varpi j}), \frac{1}{(\Sigma^{-1})_{\varpi aa}}\right).$$
Again the conditional quantile function of $X_a$ is linear in $X_{V\backslash \{a\}}$ and is given by $$Q_{X_a}(\tau\vert X_{V\backslash \{a\}},\varpi) = \frac{\Phi^{-1}(\tau)}{(\Sigma^{-1})_{\varpi aa}^{1/2}}+\mu_{\varpi a} -\sum_{j\in V\backslash\{a\}} \frac{(\Sigma^{-1})_{\varpi aj}}{(\Sigma^{-1})_{\varpi aa}}(X_j-\mu_{\varpi j}).$$
example[Monotone Transformations]
Consider the Gaussian case, for each $a\in V$, $X_a = h_a(Y_a)$ and $Y_V \sim N(\mu,\Sigma)$. It follows that for each $a\in V$, the conditional quantile function satisfies $$Q_{X_a}(\tau\vert X_{V\backslash \{a\}}) = h_a\left(\frac{\Phi^{-1}(\tau)}{(\Sigma^{-1})_{aa}^{1/2}}+\mu_a -\sum_{j\in V\backslash\{a\}} \frac{(\Sigma^{-1})_{aj}}{(\Sigma^{-1})_{aa}}(h_j^{-1}(X_j)-\mu_j)\right).$$
In particular, if $(h_a:a\in V)$ are monotone polynomials, the expression above is a sum of monomials with fractional and integer exponents.
Although a linear specification is correct for Examples (ref) and (ref), Example (ref) and (ref) illustrate that we need to consider a more general transformation of the covariates $X_V$ in the specification for each conditional quantile function. Nonetheless, specifications with additional non-linear terms can approximate non-drastic departures from normality.
We will consider a conditional quantile representation for each $a\in V$. It is based on transformations of the original covariates $X_{V\backslash\{a\}}$ that create a $p$-dimensional random vector $Z^a=Z^a(X_{V\backslash\{a\}})$ such that
equation[equation omitted — 179 chars of source]
where $r_{a\tau}$ denotes a small approximation error. For $b\in V\backslash\{a\}$ we let $I_a(b):= \{ j : Z^a_{j} \ \mbox{depends on} \ X_b\}$. That is, $I_a(b)$ contains the components of $Z^a$ that are functions of $X_b$. Under correct specification, if $X_a$ and $X_b$ are conditionally independent, we have $\beta_{a\tau j} = 0$ for all $j\in I_a(b)$, $\tau \in (0,1)$.
This allows us to connect the conditional independence quantile graph estimation problem with model selection with quantile regression. Indeed, the representation ((ref)) has been used in several quantile regression models, see Koenker2005. Under mild conditions this model allows us to identify the process $(\beta_{a\tau})_{\tau \in \mathcal{T}}$ as the solution of the following moment equation
equation[equation omitted — 114 chars of source]
In order to allow for a flexible specification, so that the approximation errors are negligible, it is attractive to consider a high-dimensional $Z^a$ where its dimension $p$ is possibly larger than the sample size $n$. In turn, having a large number of technical controls creates an estimation challenge if the number of coefficients $p$ is not negligible with respect to the sample size $n$. In such a high dimensional setting, a widely applicable condition that makes estimation possible is
approximate sparsity fan2011sparse,BCCH2012sparse,BCH2014inference. Formally we require
equation[equation omitted — 333 chars of source]
where the sparsity parameter $s$ of the model is allowed to grow (at a slower rate) as $n$ grows, and $f_{a\tau}=f_{X_a\mid X_{V\backslash\{a\}}}(Q_{X_a}(\tau\vert X_{V\backslash\{a\}})\vert X_{V\backslash\{a\}})$ denotes the conditional density function evaluated at the corresponding conditional quantile value. This sparsity also has implications on the maximum degree of the associated quantile graph.
Algorithm (ref) below contains our proposal to estimate $\beta_{a\tau}$, $a\in V$, $\tau \in \mathcal{T}$. It is based on three procedures in order to overcome high-dimensionality. In the first step, we apply a (post-)$\ell_1$-penalized quantile regression. The second step applies (post-)Lasso where the data is weighted by the conditional density function at the conditional quantile.\footnote{We note that an estimate for $f_{a\tau}$ is available from $\ell_1$-penalized quantile regression estimators for $\tau+h$ and $\tau-h$ where $h$ is a bandwidth parameter, see Koenker2005,BCK2013robustQR and Comment (ref).} Finally, the third step relies on constructing (orthogonal) score function that provides immunity to (unavoidable) model selection mistakes.
There are several parameters that need to be specified for Algorithm (ref). The penalty parameter $\lambda_{V\mathcal{T}}$ is chosen to be larger than the $\ell_\infty$-norm of the (rescaled) score at the true quantile function. The work in BC-SparseQR exploits the fact that this quantity is pivotal in their setting. Here, additional correlation structure would have an impact and the distribution is pivotal only for each $a\in V$. The penalty is based on the maximum of the quantiles of the following random variables (each with pivotal distribution), for $a\in V$
equation[equation omitted — 207 chars of source]
where $\{U_i: i=1,\ldots,n\}$ are i.i.d. uniform $(0,1)$ random variables, and $\widehat\sigma_{aj}^Z=\{{\mathbb{E}_n}[(Z_j^a)^2]\}^{1/2}$ for $j\in[p]$. The penalty parameter $\lambda_{V\mathcal{T}}$ is defined as
$$\lambda_{V\mathcal{T}} := \max_{a\in V} \Lambda_{a\mathcal{T}}(1-\xi/|V| \mid Z^a ), $$
that is, the maximum of the $1-\xi/|V|$ conditional quantile of $\Lambda_{a\mathcal{T}}$ given in ((ref)). Regarding the penalty term for the weighted Lasso in Step 2, we recommend a (theoretically valid) iterative choice. We refer to Appendix (ref) for the implementation details of the algorithm. We denote $\|\beta\|_{1,\widehat\sigma^Z}:=\sum_j \widehat\sigma_{aj}^Z|\beta_j|$ the standardized version of the $\ell_1$-norm.
algorithm[algorithm omitted — 1,159 chars of source]
Algorithm (ref) above has been studied in BCK2013robustQR where it is applied to a single triple $(a,\tau, j)$, and we have used the following parameter space for $\alpha$, $\mathcal{A}_{a\tau j} = \{ \alpha \in {\mathbb{R}} : |\alpha - \widetilde \beta_{a\tau j}| \leqslant 10 / \{\widehat \sigma_{aj}^Z \log n \} \}$. Under similar conditions, results that hold uniformly over $(a,\tau, j) \in V\times \mathcal{T}\times [p]$ are achievable (as shown in the next sections) building upon the tools developed in BC-SparseQR and chernozhukov2012gaussian. Algorithm (ref) is tailored to achieve good rates of convergence in the $\ell_\infty$-norm. In particular, under standard regularity conditions, with probability approaching to 1 we have
$$ \sup_{\tau \in \mathcal{T}} \| \beta_{a\tau} - \check\beta_{a\tau} \|_\infty \lesssim \sqrt{\frac{\log (p|V|n)}{n}}.$$
In order to create an estimate of $E^I(\tau)=\{ (a,b) \in V\times V: \max_{j\in I_a(b)}|\beta_{a\tau j}|>0\}$, we define $$\widehat E^I(\tau) = \left\{ (a,b)\in V\times V : \ \max_{j\in I_a(b)}\frac{|\check\beta_{a\tau j}|}{\mbox{se}(\check\beta_{a\tau j})} > \overline{\rm cv}\right\}$$ where $\mbox{se}(\check\beta_{a\tau j})=\{\tau(1-\tau){\mathbb{E}_n}[\widetilde v^2_{ia\tau j}]^{-1}\}^{1/2}$ with $\tilde{v}_{ia\tau j} = \widehat{f}_{ia\tau}\{ Z^a_{ij} - Z^a_{i,-j}\tilde{\gamma}^j_{a\tau}\}$, is an estimate of the standard deviation of the estimator, and the critical value $\overline{\rm cv}$ is set to account for the uniformity over $a\in V$, $\tau \in \mathcal{T}$, and $j\in[p]$. We discuss in the following sections a data driven procedure based on multiplier bootstrap that is theoretically valid in this high dimensional setting.
remark[Stepdown Procedure for $\overline{\rm cv}$] Setting a critical value $\overline{\rm cv}$ that accounts for the multiple hypotheses being tested plays an important role to estimate the graph $\widehat E^I(\tau)$. Further improvements can be obtained by considering the stepdown procedure of romano2005exact for multiple hypothesis testing that was studied for the high-dimensional case in chernozhukov2013gaussian. The procedure iteratively creates a suitable sequence of decreasing critical values. In each step only null hypotheses that were not rejected are considered to determine the critical value. Thus, as long as any hypothesis is rejected at a step, the critical value decreases and we continue to the next iteration. The procedure stops when no hypothesis in the current active set is rejected.
remark[Estimation of Conditional Density Function]
The algorithm above requires the conditional density function $f_{a\tau}$ which typically needs to be estimated in practice. It turns out that estimation of conditional quantiles yields a natural estimator for the conditional density function as
$$ f_{a\tau} = \frac{1}{\partial Q_{X_a}(\tau\vert Z^a)/\partial \tau}.$$
Therefore, based on $\ell_1$-penalized quantile regression estimates at the $\tau+h_n$ and $\tau-h_n$ quantile, where $h = h_n \to 0$ denotes a bandwidth parameter, we have\begin{equation} \widehat f_{a\tau} = \frac{2h}{\widehat Q_{X_a}(\tau+h\vert Z^a)-\widehat Q_{X_a}(\tau-h\vert Z^a)}\end{equation}
as an estimator of $f_{a\tau}$. Under smoothness conditions, it has a bias of order $h^2$. See BCK2013robustQR and the references therein for additional comments and estimators.
Estimators for PQGMs
In this section we propose an estimator for PQGMs in which case we are interested in the prediction of $X_a$, $a\in V$, using a linear combination of $X_{V\backslash\{a\}}$ under the asymmetric loss discussed in ((ref)). We will add an intercept as one of the variables for the sake of notation so that $X_{-a}=(1,X_{V\backslash\{a\}}')'$. Given the loss function $\rho_\tau$, the target $d$-dimensional vector of parameters $\beta_{a\tau}$ is defined as (part of) the solution of the following optimization problem
equation[equation omitted — 117 chars of source]
As we are interested in the case that $d$ is large, the use of high-dimensional tools to achieve consistent estimators is needed. The estimation procedure we proposed is based on $\ell_1$-penalized quantile regression but additional issues need to be considered to cope with the (non-vanishing) difference between the best linear predictor and the conditional quantile function. Again we consider models that satisfy an approximately sparse condition. Formally, we require the existence of sparse coefficients $\{\bar\beta_{a\tau}:a\in V, \tau \in \mathcal{T}\}$ such that
equation[equation omitted — 272 chars of source]
where (again) the sparsity parameter $s$ of the model is allowed to grow as $n$ grows. The high-dimensionality prevents us from using (standard) quantile regression methods and regularization methods are needed to achieve good prediction properties.
A key issue is to set the penalty parameter properly so that it bounds from above
equation[equation omitted — 165 chars of source]
However, it is important to note that we do not assume that the conditional quantile of $X_a$ is a linear function of $X_{-a}$. Under correct linear specification of the conditional quantile function, $\ell_1$-penalized quantile regression estimator has been studied in BC-SparseQR. The case that the conditional quantile function differs from a linear specification by vanishing approximation errors has been considered in kato2011 and BCK2013robustQR. The analysis proposed here aims to allow for non-vanishing misspecification of the quantile function relative to a linear specification while still guarantees good rates of convergence in the $\ell_2$-norm to the best linear specification. Thus the penalty parameter in the penalized quantile regression needs to account for such misspecification and is no longer pivotal as in BC-SparseQR.
In order to handle this issue we propose a two step estimation procedure. In the first step, the penalty parameter $\lambda_0$ is conservative and is set via bounds constructed based on symmetrization arguments, similar in spirit to vdGeer2008,Volume2013. This leads to $\lambda_0 = 2(1+1/16) \sqrt{2\log(8|V|^2/\xi)/n}$. Although this is conservative, under mild conditions this would lead to estimates that can be leverage to fine tune the penalty choice. The second step uses the preliminary estimator to bootstrap ((ref)) based on the tools in chernozhukov2013gaussian as follows. Specifically, for estimates $\widehat\varepsilon_{ia\tau}$ of the “noise" $\varepsilon_{ia\tau}=1\{X_{ia} \leqslant X_{i,-a}'\beta_{a\tau}\}-\tau$ for $i\in[n]$, for $a\in V$ define
equation[equation omitted — 251 chars of source]
where $(g_i)_{i=1}^n$ is a sequence of i.i.d. standard Gaussian random variables. The new penalty parameter $\bar\lambda_{V\mathcal{T}}$ is defined as
equation[equation omitted — 124 chars of source]
that is, the maximum of the $(1-\xi)$ conditional quantile of $\bar{\Lambda}_{a\mathcal{T}}$. The penalty choice above adapts to the unknown correlation structure across components and quantile indices. The following algorithm states the procedure where we denote weighted $\ell_1$-norms by $\|\beta\|_{1,\widehat\sigma^X}:=\sum_j\widehat{\sigma}_{aj}^X|\beta_j|$ with $\widehat{\sigma}_{aj}^X=\{{\mathbb{E}_n}[X_{j}^2]\}^{1/2}$ the standardized version of the $\ell_1$-norm and $\|\beta\|_{1,\widehat\varepsilon}:=\sum_j \widehat{\sigma}_{a\tau j}^{\varepsilon X}|\beta_j|$ with $ \widehat{\sigma}_{a\tau j}^{\varepsilon X} = \{{\mathbb{E}_n}[\widehat\varepsilon_{a\tau}^2X_{-a,j}^2]\}^{1/2}$ a norm based on the estimated residuals.
algorithm[algorithm omitted — 1,106 chars of source]
Under regularity conditions stated in Section (ref), with probability approaching 1, we have
$$ \max_{a\in V} \sup_{\tau \in \mathcal{T}} \| \beta_{a\tau} - \check\beta_{a\tau} \| \lesssim \sqrt{\frac{s\log (|V|n)}{n}}.$$
The estimate of the prediction quantile graph is given by the support of $(\check\beta_{a\tau})_{a\in V, \tau \in \mathcal{T}}$, namely $$\widehat E^P(\tau) = \left\{ (a,b)\in V\times V : \ |\widehat\beta_{a\tau b}| > \bar \lambda_{V\mathcal{T}} / \widehat{\sigma}_{a\tau b}^{\varepsilon X} \right \}.$$ That is, it is induced by covariates selected by the $\ell_1$-penalized estimator. Those thresholded estimators not only have the same rates of convergence as of the original penalized estimators but also possess additional sparsity guarantees.
Estimators for $\mathcal{W}$-Conditional Quantile Graphical Models
In order to handle the additional conditioning events $\Omega_\varpi$, $\varpi\in\mathcal{W}$, we propose to modify Algorithms (ref) and (ref) based on kernel smoothing. To that extent, we assume the observed data is of the form $\{(X_{iV}, W_i) : i=1,\ldots, n\}$, where $W_i$ might be defined through additional variables. Furthermore, we assume for each conditioning event $\varpi\in \mathcal{W}$ we have access to a kernel function $K_\varpi$ that is applied to $W$, to represent the relevant observations associated with $\varpi$ (recall that we denote ${\mathrm{P}}( W \in \Omega_\varpi)$ as ${\mathrm{P}}(\varpi)$). We assume that $K_\varpi(W) = 1\{W \in \Omega_\varpi\}$.
example[Stock Returns Under Market Downside Movements, continued] In Example (ref), we have $W$ as the market return and the conditioning event as $\Omega_\varpi=\{W\leqslant \varpi\}$ which is parameterized by $\varpi\in\mathcal{W}$, a closed interval in $ {\mathbb{R}}$. We might be interest in a fixed $\varpi$ or on a family of values $\varpi\in(-\bar \varpi, 0]$. The latter induces $\mathcal{W} = \{\Omega_\varpi=\{W\leqslant \varpi\} : \varpi\in(-\bar \varpi, 0] \}$. The kernel function is simply $K_{\varpi}(t) = 1\{t\leqslant \varpi\}$.
This framework encompasses the previous framework by having $K_\varpi(W) = 1$ for all $W$. Moreover, it allows for a richer class of estimands which require estimators whose properties should hold uniformly over $\varpi \in \mathcal{W}$ as well. Next we propose estimators for this setting, i.e. we generalize the previous methods to account for the additional conditioning on $\varpi\in \mathcal{W}$. In what follows, with a slight abuse of notation we use $\varpi$ to denote not only the index but also the event $\Omega_\varpi$. For further notational convenience, we denote $u=(a,\tau,\varpi)\in \mathcal{U} := V\times\mathcal{T}\times\mathcal{W}$ so that the set $\mathcal{U}$ collects all the three relevant indices. With $\widehat{\sigma}^{Z}_{a\varpi j} = \{{\mathbb{E}_n}[K_\varpi(W)(Z_j^a)^2]\}^{1/2}$, we define the following weighted $\ell_1$-norm $ \|\beta\|_{1,\varpi} = \sum_{j\in [p]} \widehat{\sigma}^{Z}_{a\varpi j} |\beta_j|.$ This norm is $\varpi$ dependent and provides the proper adjustments as we condition on different events associated with different $\varpi$'s.
We first consider estimators of CIQGMs conditional on the events in $\mathcal{W}$. In this setting, the model is correctly specified up to small approximation errors. The definition of the penalty parameter will be based on the random variable
$$ \Lambda_{a\mathcal{T}\mathcal{W}} = \sup_{\tau \in \mathcal{T}, \varpi \in \mathcal{W}} \max_{ j\in[p]} \left|\frac{{\mathbb{E}_n}[K_\varpi(W)(1\{U \leqslant \tau\} - \tau) Z^a_j]}{\sqrt{\tau(1-\tau)} \widehat{\sigma}^{Z}_{a\varpi j}}\right| $$
where $U_i$ are independent uniform $(0,1)$ random variables, and set the penalty $$\lambda_{V\mathcal{T}\mathcal{W}} = \max_{a\in V}\Lambda_{a\mathcal{T}\mathcal{W}}(1-\xi/\{|V|n^{1+2d_W}\}\vert Z^a, W),$$
that is, the maximum of the $(1-\xi/\{|V|n^{1+2d_W}\})$ conditional quantile of $\Lambda_{a\mathcal{T}\mathcal{W}} $. Algorithm (ref) provides the definition of the estimator. Here $\mathcal{A}_{uj} = \{ \alpha \in {\mathbb{R}} : |\alpha - \widetilde \beta_{uj}| \leqslant 10 / \{\widehat{\sigma}^{Z}_{a\varpi j} \log n \} \}$, and denote $\lambda_u := \lambda_{V\mathcal{T}\mathcal{W}}\sqrt{\tau(1-\tau)}$.
algorithm[algorithm omitted — 1,142 chars of source]
Next we consider estimators of PQGMs conditional on the events in $\mathcal{W}$. Similar to the previous case, for $a\in V$ define
equation[equation omitted — 296 chars of source]
where $(g_i)_{i=1}^n$ is a sequence of i.i.d. standard Gaussian random variables. The new penalty parameter $\bar\lambda_{V\mathcal{T}}$ is defined as
equation[equation omitted — 150 chars of source]
that is, the maximum of the $(1-\xi)$ conditional quantile of $\bar{\Lambda}_{a\mathcal{T}\mathcal{W}}$.
It will also be useful to define another weighted $\ell_1$-norm, $\|\beta\|_{1,\varpi\widehat\varepsilon}:= \sum_{j} \widehat{\sigma}_{a\tau \varpi j}^{\varepsilon X}|\beta_j|$ with $ \widehat{\sigma}_{a\tau \varpi j}^{\varepsilon X} = \{{\mathbb{E}_n}[K_\varpi(W)\widehat\varepsilon_{a\tau\varpi}^2X_{-a, j}^2]\}^{1/2}$. We also denote $\widehat{\sigma}^{X}_{a\varpi j}=\{{\mathbb{E}_n}[K_\varpi(W)X_{-a,j}^2]\}^{1/2}$. The penalty choice and weighted $\ell_1$-norm adapt to the unknown correlation structure across components and quantile indices. The following algorithm states the procedure, with $ \lambda_{0\mathcal{W}} = 2(1+1/16) \sqrt{2\log(8|V|^2\{ne/d_W\}^{2d_W}/\xi)/n}$.
algorithm[algorithm omitted — 1,215 chars of source]
remark[Computation of Penalty Parameter over $\mathcal{W}$] The penalty choices require one to maximize over $a\in V$, $\tau\in\mathcal{T}$ and $\varpi\in\mathcal{W}$. The set $V$ is discrete and does not pose a significant challenge. However both other sets are continuous and additional care is needed. In most applications we are concerned with the case that $\mathcal{W}$ is a low dimensional VC class of sets and it impacts the calculation only through indicator functions, which is precisely the case of $\mathcal{T}$. It follows that only a polynomial number (in $n$) of different values of $\tau$ and $\varpi$ would need to be considered. \footnote{
A class of sets is said to be a VC class, if the VC dimension is finite. In what follows we use that the VC dimension provides a way to control how much we can overfit the data and it will also lead to (theoretically valid) recommendations for the penalty parameters. For the formal definition of VC class, see vdV-W).}
Main Theoretical Results
This section is devoted to theoretical guarantees associated with the proposed estimators. We will establish rates of convergence results for the proposed estimators as well as the (uniform) validity of confidence regions. These results build upon and contribute to an increasing literature on the estimation of many processes of interest with (high-dimensional) nuisance parameters.
Throughout, we will provide results for the estimators of the $\mathcal{W}$-conditional quantile graphical models as those can be generalized the other models by setting $K_\varpi(W) = 1$. Although some of the tools are similar, CIQGMs and PQGMs require different estimators and are subject to different assumptions. Thus, substantial different analyses are required.
$\mathcal{W}$-Conditional CIQGM
For $u=(a,\tau,\varpi)\in \mathcal{U}$, define the $\tau$-conditional quantile function of $X_a$ given $X_{V\backslash\{a\}}$ and $\varpi$ as
equation[equation omitted — 105 chars of source]
where $Z^a$ is a $p$-dimensional vector of (known) transformations of $X_{V\backslash \{a\}}$, and $r_u$ is an approximation error. The event $\varpi \in \mathcal{W}$ will be used for further conditioning through the function $K_\varpi(W) = 1\{ W \in \varpi\}$.
We let $f_{X_a\mid X_{V\backslash \{a\}}, \varpi }(\cdot\vert X_{V\backslash \{a\}},\varpi)$ denote the conditional density function of $X_a$ given $X_{V\backslash \{a\}}$ and $\varpi\in\mathcal{W}$. We define $f_{u}:=f_{X_a\vert X_{V\backslash \{a\}}, \varpi }(Q_{X_{a}}(\tau\vert X_{V\backslash \{a\}}, \varpi )\vert X_{V\backslash \{a\}}, \varpi )$ as the value of the conditional density function evaluated at the $\tau$-conditional quantile. In our analysis we will consider for $u\in\mathcal{U}$
equation[equation omitted — 264 chars of source]
Moreover, for each $u\in\mathcal{U}$ and $j\in[p]$ we define
equation[equation omitted — 129 chars of source]
This provides a weighted projection to construct the residuals
$$ v_{uj} = f_u(Z^a_j - Z^a_{-j}\gamma_{u}^j) $$
that satisfy ${\mathrm{E}}[ f_{u} Z^a_{-j} v_{uj} \vert \varpi] =0$ for each $(u,j)\in \mathcal{U}\times[p]$.
The estimands of interest are $\beta_u \in {\mathbb{R}}^p $, $u \in \mathcal{U}$, and can be written as the solution of (a continuum of) moment equations. Letting $\beta_{uj}$ denote the $j$th component of $\beta_u$ so that $\beta_{uj} \in {\mathbb{R}}$ solves
$$ {\mathrm{E}}[\psi_{uj}(X,W,\beta,\eta_{uj}) ] = 0, $$
where the function $\psi_{uj}$ is given by
$$ \psi_{uj}(X,W,\beta,\eta_{uj}) = K_\varpi(W)( \tau - 1\{ X_a \leqslant Z^a_j \beta + Z^a_{-j}\eta_{uj}^{(1)} + \eta_{uj}^{(3)} \})f_u(Z_j^a-Z^a_{-j}\eta_{uj}^{(2)}),$$
and the true value of the nuisance parameter is given by $\eta_{uj}=(\eta_{uj}^{(1)},\eta_{uj}^{(2)},\eta_{uj}^{(3)})$ with $\eta_{uj}^{(1)}=\beta_{u,-j}$, $\eta_{uj}^{(2)}=\gamma^j_u$, and $\eta_{uj}^{(3)}=r_u$. In what follows $c, C$ denote some fixed constant, $\delta_n$ and $\Delta_n$ denote sequences go to zero with $\delta_n = n^{-\mu}$ for some sufficiently small $\mu$. Denote $\mu_\mathcal{W}=\inf_{\varpi\in\mathcal{W}}{\mathrm{P}}(\varpi)$.
{\bf Condition CI.} Let $u=(a, \tau, \varpi) \in \mathcal{U} := V \times \mathcal{T}\times \mathcal{W}$ and $(X_i,W_i)_{i=1}^n$ denote a sequence of independent and identically distributed random vectors generated accordingly to models ((ref)) and ((ref)):
(i) Suppose $\sup_{u\in\mathcal{U}, j\in[p]}\{\|\beta_u\|+\|\gamma^j_u\|\}\leqslant C$ and $\mathcal{T}$ is a fixed compact set:
(a) there exists $s=s_n$ such that $\sup_{u\in\mathcal{U}, j\in[p]}\{\|\beta_u\|_0 + \|\bar \gamma^j_u\|_0\}\leqslant s$, $\sup_{u\in\mathcal{U}, j\in[p]}\|\bar \gamma^j_u-\gamma^j_u\|+s^{-1/2}\|\bar \gamma^{j}_u-\gamma^{j}_u\|_1\leqslant C \{n^{-1}s\log(|V| p n)\}^{1/2}$, where $\bar \gamma^j_u$ is approximately sparse;
(b) the conditional distribution function of $X_a$ given $X_{V\backslash \{a\}}$ and $\varpi$ is absolutely continuous with continuously differentiable density $f_{X_a\vert X_{V\backslash \{a\}},\varpi}(t\vert X_{V\backslash \{a\}},\varpi)$ bounded by $\bar f$ and its derivative bounded by $\bar f'$ uniformly over $u\in\mathcal{U}$; (c) $|f_u-f_{u'}|\leqslant L_f\|u-u'\|$, $\|\beta_u-\beta_{u'}\|\leqslant L_\beta \|u-u'\|^\kappa$ with $\kappa \in [1/2,1]$, and ${\mathrm{E}}[ |K_\varpi(W)-K_{\varpi'}(W)|] \leqslant L_K\|\varpi-\varpi'\|$; (d) the VC dimension $d_W$ of the set $\mathcal{W}$ is fixed, $\{Q_{X_a}(\tau\vert X_{V\backslash \{a\}},\varpi) : (\tau,\varpi)\in \mathcal{T} \times \mathcal{W}\}$ is a VC-subgraph with VC-dimension $1+Cd_W$ for every $a\in V$;
(ii) The following moment conditions hold uniformly over $u\in\mathcal{U}$ and $j\in[p]$: ${\mathrm{E}}[|f_uv_{uj}Z^a_k|^2\vert \varpi]^{1/2}\leqslant C{\underline{f}}_u$, $\min_{a\in V}\inf_{\|\delta\|=1}{\mathrm{E}}[\{(X_a,Z^a)\delta\}^2 \vert \varpi]\geqslant c$, $ \max_{a\in V}\sup_{\|\delta\|=1}{\mathrm{E}}[ \{(X_a,Z^a)\delta\}^4 \vert \varpi ]\leqslant C$, ${\mathrm{E}}[f_u^2 (Z^a\delta)^{2}\vert \varpi] \leqslant C{\underline{f}}_u^2{\mathrm{E}}[(Z^a\delta)^{2}\vert \varpi ]$, $\max_{j,k}\frac{{\mathrm{E}}[|f_uv_{uj}Z^a_k|^3\vert \varpi]^{1/3}}{{\mathrm{E}}[|f_uv_{uj}Z^a_k|^2\vert \varpi]^{1/2}}\log^{1/2}(pn|V|) \leqslant \delta_n \{n {\mathrm{P}}(\varpi)\}^{1/6}$;
(iii) Furthermore, for some fixed $q\geqslant 4\vee (1+2d_W)$, $\sup_{u\in\,\mathcal{U},\|\delta\|=1} {\mathrm{E}}[|(X_a,Z^a)\delta|^2r_u^2 \vert\varpi]\leqslant C {\mathrm{E}}[r_u^2 \vert \varpi] \leqslant Cs/n$, $\max_{u\in\mathcal{U}, j\in[p]}|{\mathrm{E}}[f_ur_uv_{uj}\vert \varpi]|\leqslant \delta_n n^{-1/2}$, ${\mathrm{E}}[ \max_{i\leqslant n} \sup_{u\in \mathcal{U}}|K_\varpi(W)r_{iu}|^q ] \leqslant C$, and with probability $1-\Delta_n$, uniformly over $u\in \mathcal{U}, j\in[p]$: ${\mathbb{E}_n}[r_u^2v_{uj}^2\vert\varpi] + {\mathbb{E}_n}[r_u^2\vert\varpi]\lesssim n^{-1}s\log(p|V| n)$, ${\mathbb{E}_n}[K_\varpi(W)\{|r_u|+r_u^2\}(Z^a\delta)^2] \leqslant \delta_n {\mathbb{E}_n}[K_\varpi(W)f_u(Z^a\delta)^2]$;
(iv) For a fixed $q \geqslant 4 \lor (1+d_W)$, $\mathrm{diam}(\mathcal{W}) \leqslant n^{1/2q}$, ${\mathrm{E}}[ \max_{i\leqslant n}\|X_{iV}\|_\infty^q\vee \max_{a\in V}\|Z^a_i\|_\infty^q]^{1/q}/\mu_\mathcal{W}\leqslant M_n$, ${\mathrm{E}}[\max_{i\leqslant n} \sup_{u\in\mathcal{U}, j\in[p]} |v_{iuj}|^q ]^{1/q}\leqslant L_n $, $ (L_f+L_K)^2M_n^2\log^2(p|V|n)\leqslant \delta_n n\mu_\mathcal{W}^3{\underline{f}}_\mathcal{U}^6 $, $M_n^{4}\log(p|V|n) \log n\leqslant \delta_n^2 n \mu_\mathcal{W}^2{\underline{f}}_\mathcal{U}^2$, $s^2\log^2(p|V|n)\leqslant \delta_n^2 n {\underline{f}}_\mathcal{U}^4\mu_\mathcal{W}^6$,
$s^3\log^3(p|V|n)\leqslant \delta_n^4 n \underline{f}_{\mathcal{U}}^2\mu_\mathcal{W}^3$,
$L_n^2s\log^{3/2} (p|V|n) \leqslant \delta_n \underline{f}_{\mathcal{U}}(n\mu_\mathcal{W})^{1/2}$, $M_n s \sqrt{\log( p|V|n)} \leqslant \delta_n n^{1/2}\mu_\mathcal{W}{\underline{f}}_\mathcal{U}$.
Condition CI assumes various conditional moment conditions to allow for the estimation to be conditional on $\varpi \in \mathcal{W}$. Those are analogous to the (unconditional) conditions in the high-dimensional literature in quantile regression
models, BCK2013robustQR. In particular, condition CI(i) assumes smoothness of the density function, and of coefficients. Condition CI(ii) assumes conditions on the (conditional) population design matrices such as the ratio between eigenvalues. Condition CI(iii) pertains to the approximations errors and assumes mild moment conditions. Finally Condition CI(iv) provides sufficient conditions on the allowed growth of the model via $p$ and $|V|$ relative to the available sample size $n$. Note, Condition CI(iii) also assume $d_W$ is bounded by fixed $q$, and the proof can easily be extended to other cases.
Condition CI is a high level condition intended to allow approximate sparse models, approximation errors, tail events in $\mathcal{W}$, and to require only $q$ moments (going beyond sub-Gaussian variables). When applied to the special case of sub-Gaussian, exactly sparse, and singleton $\mathcal{W}$, it becomes a relatively standard assumption. For example, without approximation error, e.g. the multivariate Gaussian
case, Condition CI(iii) can be removed entirely. By allowing for a large number of variables (and transformations) the approximation errors can be controlled when the quantile functions belong to some smooth function class (e.g. Sobolev space). Although it is outside the scope of the current work, the ideas and results can be generalized
to dependent data, using results from chernozhukov2014clt.
Based on Condition CI, we derive our main results regarding the proposed estimator.
Moreover, we also establish new results for $\ell_1$-penalized quantile regression methods that hold uniformly over the indices $u \in {\mathcal{U}}$. The following theorems summarize these results.
theorem[{Uniform Rates of Convergence for $\mathcal{W}$-Conditional Penalized Quantile Regression}]
Under Condition CI, we have that with probability at least $1-o(1)$
$$ \|\widehat\beta_u - \beta_u\| \lesssim \sqrt{\frac{s(1+d_W)\log(p|V|n)}{n \underline{f}_u{\mathrm{P}}(\varpi)}}, \ \ \ \mbox{uniformly over $u=(a,\tau,\varpi) \in\mathcal{U}$}$$
Moreover, the thresholded estimator $\widehat\beta^{\bar\lambda}$, with $\bar\lambda=\sqrt{(1+d_W)\log(p|V|n)/n}$ and $\widehat\beta_{uj}^{\bar\lambda}=\widehat\beta_{uj}1\{|\widehat\beta_{uj}|>\\
\bar\lambda \widehat{\sigma}^{Z}_{a\varpi j}\}$, satisfies the same rate and $\|\widehat\beta^{\bar\lambda}\|_0\lesssim s$.
Theorem (ref) builds upon ideas in BC-SparseQR however the proof strategy is designed to derive rates that are adaptive to each $u\in \mathcal{U}$. Indeed the rates of convergence are $u$-dependent and they show a slower rate for rare events $\varpi \in \mathcal{W}$.
theorem[{Uniform Rates of Convergence for $\mathcal{W}$-Conditional Weighted Lasso}]
Under Condition CI, we have that with probability at least $1-o(1)$
$$ \|\widehat\gamma^j_u - \gamma^j_u\| \lesssim \frac{1}{\underline{f}_{u}} \sqrt{\frac{s(1+d_W)\log(p|V|n)}{n{\mathrm{P}}(\varpi)}} \ \ \ \mbox{and} \ \ \ \|\widehat\gamma_u^j\|_0 \lesssim s, \ \ \ \mbox{uniformly over $u=(a,\tau,\varpi) \in\mathcal{U}$, $j\in[p]$.}$$
The following result establishes a uniform Bahadur representation for the final estimators.
theorem[{Uniform Bahadur Representation for $\mathcal{W}$-Conditional CIQGM}] Under Condition CI, the estimator $(\check \beta_{uj})_{u \in \mathcal{U},j\in[p]}$ satisfies
$$ \sigma_{uj}^{-1}\sqrt{n}(\check\beta_{uj} - \beta_{uj}) = \mathbb{U}_n(u,j) + O_P(\delta_n) \text{ in } \ell^\infty(\mathcal{U}\times [p]),$$
where $\sigma^2_{uj}= \tau(1-\tau){\mathrm{E}} [K_\varpi(W)v_{uj}^2]^{-1}$ and
$$\mathbb{U}_n(u,j):=\frac{\{\tau(1-\tau){\mathrm{E}} [K_\varpi(W)v_{uj}^2]\}^{-1/2}}{\sqrt{n}} \sum_{i=1}^n ( \tau - 1\{U_i(a,\varpi)\leqslant \tau\})K_\varpi(W_i)v_{i,uj},$$ where $U_1(a,\varpi), \ldots, U_n(a,\varpi)$ are i.i.d. uniform $(0,1)$ random variables, independent of $v_{1,uj},\ldots, v_{n,uj}$.
Theorem (ref) plays a key role. However, it is important to note that the marginal distribution of $\mathbb{U}_n(u,j)$ is pivotal. Nonetheless, there is a non-trivial correlation structure between $U(a,\varpi)$ and $U(\tilde a,\tilde \varpi)$. In order to construct confidence regions with non-conservative guarantees, we rely on a multiplier bootstrap method. We will approximate the process $\mathcal{N}=(\mathcal{N}_{uj})_{u\in\mathcal{U}, j\in[p]}$ by the Gaussian multiplier bootstrap based on estimates $\widehat\psi_{uj}:= \{\tau(1-\tau){\mathbb{E}_n}[K_{\varpi}(W)\widehat{v}_{uj}^{2}]\}^{-1/2}(\tau-1\{ X_a \leqslant Z^a\widehat\beta_u\})K_\varpi(W_i)\widehat v_{uj}$ of $\bar\psi_{uj}(U,W)=\{\tau(1-\tau){\mathbb{E}_n}[K_{\varpi}(W)v_{uj}^{2}]\}^{-1/2}(\tau - 1\{ U(a,\varpi) \leqslant \tau\})K_\varpi(W)v_{uj}$, namely
$$ \widehat{\mathcal{G}} = (\widehat{\mathcal{G}}_{uj})_{u\in\mathcal{U},j\in[p]} = \left\{ \frac{1}{\sqrt{n}}\sum_{i=1}^ng_i\widehat\psi_{uj}(X_i,W_i)\right\}_{u\in\mathcal{U},j\in[p]}$$
where $(g_i)_{i=1}^n$ are independent standard normal random variables which are independent from the data $(W_i)_{i=1}^n$. Based on Theorem 5.2 of chernozhukov2013gaussian, the following result shows that the multiplier bootstrap provides a valid approximation to the large sample probability law of $\sqrt{n}(\check \beta_{uj}- \beta_{uj})_{u \in \mathcal{U},j\in[p]}$ which is suitable for the construction of uniform confidence bands over the set of indices associated with $I_a(b)$ for all $a,b \in V$. We let $\mathcal{P}_n$ denote the collection of distributions $P$ for the data such that Condition CI is satisfied for given $n$. This is the collection of all approximately sparse models where the above sparsity conditions, moment conditions, and growth conditions are satisfied.
corollary[{Gaussian Multiplier Bootstrap for $\mathcal{W}$-Conditional CIQGM}]
Under Condition CI with $\delta_n = o(\{ (1+d_W)\log(p|V|n)\}^{-1/2})$, and $(1+d_W)\log(p|V|n) = o(\{ (n/L_n^2)^{1/7}\wedge (n^{1-2/q}/L_n^2)^{1/3}\})$, we have that
$$ \sup_{P \in \mathcal{P}_n} \sup_{t,t' \in {\mathbb{R}}, u\in\mathcal{U}, b\in V} \left| {\mathrm{P}}_P\left( \max_{j\in I_a(b)}\frac{|\check\beta_{uj}-\beta_{uj}|}{n^{-1/2}\sigma_{uj}} \in [t,t'] \right) - {\mathrm{P}}_P\left( \max_{j\in I_a(b)}|\widehat{\mathcal{G}}_{uj}| \in [t,t']\mid (X_i,W_i)_{i=1}^n\right) \right| = o(1) $$
Corollary (ref) allows the construction of simultaneous confidence regions for the coefficients that are uniformly valid over the set of data generating processes induced by Condition CI. Based on the coefficients whose intervals do not overlap zero, we can construct a conditional independence graph process $\widehat E^{I}(\tau,\varpi), \tau \in \mathcal{T}, \varpi\in\mathcal{W}$ that contains the true conditional independence quantile graph with a specified probability.
$\mathcal{W}$-Conditional PQGM
In this section, we derive theoretical guarantees for the $\mathcal{W}$-conditional predictive quantile estimators uniformly over $u=(a,\tau,\varpi)\in \mathcal{U}$. For each $u\in \mathcal{U}$ the estimand of interest is $\beta_u \in {\mathbb{R}}^p $ that corresponds to the best linear predictor under asymmetric loss function, namely
equation[equation omitted — 118 chars of source]
where the event $\varpi \in \mathcal{W}$ is used for further conditioning. In the analysis below, the conditioning is implemented through the function $K_\varpi(W) = 1\{ W \in \varpi\}$.
In the analysis of this case, the main issue is to handle the inherent misspecification of the linear form $X_{-a}'\beta_u$ with respect to the true conditional quantile. The first consequence is to handle the identification condition. Given $X_{-a}$ and $\varpi\in\mathcal{W}$, we let $f_{u}:=f_{X_a\vert X_{-a}, \varpi } (X_{-a}'\beta_{u} \vert X_{-a}, \varpi )$ denote the value of the conditional density function evaluated at $X_{-a}'\beta_{u}$. In our analysis, we will consider
equation[equation omitted — 270 chars of source]
We remark that $\underline f_{u}$ defined in ((ref)) differs from ((ref)) which is the standard conditional density at the true quantile value. It turns out that Knight's identity can be used by exploiting the first order condition associated with the optimization problem ((ref)) which yields zero mean condition similar to the conditional quantile condition.
A second consequence of the misspecification is the lack of pivotality of the score. Such pivotal property was convenient in the previous section to define penalty parameters and to conduct inference. We will exploit bounds on the VC-dimension of the relevant classes of sets formally stated below.
{\bf Condition P.} Let $\mathcal{U} = V \times \mathcal{T}\times \mathcal{W}$ and $(X_i,W_i)_{i=1}^n$ denote a sequence of independent and identically distributed random vectors generated accordingly to models ((ref)):
(i) Suppose that $\sup_{u\in\mathcal{U}}\|\beta_u\|\leqslant C$ and $\mathcal{T}$ is a fixed compact set: (a) there exists $s=s_n$ and $\bar \beta_u$ such that $\sup_{u\in\mathcal{U}}\|\bar \beta_u\|_0 \leqslant s$, $\sup_{u\in\mathcal{U}}\|\bar \beta_u-\beta_u\|+s^{-1/2}\|\bar \beta_u-\beta_u\|_1 \leqslant \sqrt{s/n}$; (b) the conditional distribution function of $X_a$ given $X_{-a}$ and $\varpi$ is absolutely continuous with continuously differentiable density $f_{X_a\mid X_{ -a},\varpi}(t\mid X_{-a},\varpi)$ such that its values are bounded by $\bar f$ and its derivative is bounded by $\bar f'$ uniformly over $u\in\mathcal{U}$; (c) $|f_u-f_{u'}|\leqslant L_f\|u-u'\|$, $\|\beta_u-\beta_{u'}\|\leqslant L_\beta \|u-u'\|^\kappa$ with $\kappa \in [1/2,1]$, and ${\mathrm{E}}[ |K_\varpi(W)-K_{\varpi'}(W)|] \leqslant L_K\|\varpi-\varpi'\|$; (d) the VC dimension $d_W$ of the set $\mathcal{W}$ is fixed, $\{1\{ X_a \leqslant X_{-a}'\beta_u\} : (\tau,\varpi)\in \mathcal{T}\times \mathcal{W}\}$ is a VC-class with VC-dimension $1+d_W$ for every $a\in V$;
(ii) The following moment conditions hold uniformly over $u\in\mathcal{U}$: $\min_{a\in V}\inf_{\|\delta\|=1}{\mathrm{E}}[\{X_{-a}'\delta\}^2 \vert \varpi]\geqslant c$, $\max_{a\in V}\sup_{\|\delta\|=1}{\mathrm{E}}[ \{X_{-a}'\delta\}^4 \vert \varpi ]\leqslant C$;
((iii) With probability $1-\Delta_n$, uniformly over $u\in \mathcal{U}$ and $a\in V$: ${\mathbb{E}_n}[K_\varpi(W)\{|X_{-a}'(\bar\beta_u-\beta_u)|+|X_{-a}'(\bar\beta_u-\beta_u)|^2\}(Z^a\delta)^2] \leqslant \delta_n {\mathbb{E}_n}[K_\varpi(W)f_u(X_{-a}'\delta)^2]$;
(iv) For a fixed $q \geqslant 4 \lor (1+d_W)$, we have that: $\mathrm{diam}(\mathcal{W}) \leqslant n^{1/2q}$, $ {\mathrm{E}}[\max_{i\leqslant n}\|X_{iV}\|_\infty^q]^{1/q}/\mu_\mathcal{W} \leqslant M_n$, $M_n^2 \log^7(n|V|) \leqslant \delta_n n {\underline{f}}_\mathcal{U}^2\mu_\mathcal{W}^2$, $M_n^{4}\log(n|V|) \log n\leqslant \delta_n n \mu_\mathcal{W}$, $(L_f+L_K)^2M_n^2\log^2(|V|n)\leqslant \delta_n n \mu_\mathcal{W}^3{\underline{f}}_\mathcal{U}^6$, $M_n^2s\log^{3/2} (n|V|) \leqslant \delta_n \underline{f}_{\mathcal{U}}(n\mu_\mathcal{W})^{1/2}$, $M_n s \sqrt{\log(n |V|)} \leqslant \delta_n (n \mu_\mathcal{W})^{1/2}$, and $s^3\log^5(n|V|)\leqslant \delta_n n \underline{f}_{\mathcal{U}}^2\mu_\mathcal{W}^2$.
Condition P is a high-level condition. It allows to cover conditioning events $\varpi \in \mathcal{W}$ whose probability can decrease to zero (although slower than $n^{-1/4}$).
Next we derive our main results regarding the proposed estimator for the best linear predictor. These results are also new $\ell_1$-penalized quantile regression methods as it holds under possible misspecification of the conditional quantile function and hold uniformly over the indices $u\in \mathcal{U}$. The following theorem summarizes the result.
theorem[{Uniform Rates of Convergence for $\mathcal{W}$-Conditional Penalized Quantile Regression under Misspecification}]
Under Condition P, we have that with probability at least $1-o(1)$, uniformly over $u=(a,\tau,\varpi) \in\mathcal{U}$,
$$ \|\widehat\beta_u - \beta_u\| \lesssim \sqrt{\frac{s(1+d_W)\log(|V|n)}{n \underline{f}_u{\mathrm{P}}(\varpi)}}.$$
The data-driven choice of penalty parameter helps diminish the regularization bias and also allow to obtain sparse estimators with provably rates of convergence (through thresholding). Moreover, the $u$ specific penalty parameter combined with the new analysis yields an adaptive rate of convergence to each $u\in \mathcal{U}$ unlike previous works.
remark[Simultaneous Confidence Bands for Coefficients in PQGMs]
We note that in some applications we might be interested in constructing (simultaneous) confidence bands for the coefficients in PQGMs. In particular, this would include the cases practitioners are using a misspecified linear specification in a quantile regression model. Provided the conditional density function at $X_{-a}'\beta_u$ can be estimated, a version of Algorithm (ref) using the penalty parameters in Algorithm (ref) for the initial step can deliver such confidence regions via a multiplier bootstrap.
Application: International Financial Contagion and Systemic Risk
There is widespread disagreement about what finacial contagion entails, e.g. Forbes2002, diebold2015trans. The existing measures of contagion are mainly based on linear correlation and can only account for certain types of risk network structure or do not provide inference procedures for the large scale networks estimated. This paper defines contagion occurs whenever the quantile partial correlation from one country to another country is nonzero, i.e. the presence of edges in PQGM. This definition takes into account network spillover effects when identifying contagion and measuring systemic risk. Here, the weight of an edge indicates the strength of contagion effects. The estimated contagion network taking into account global interconnectedness is important for Eurozone financial regulators to identify globally systemically important EU countries, or for global financial portofolio diversification.
We revisit the analysis of of international
financial contagion, Claessens2001. We provide an alternative approach to the literature by visualizing tail interdependence via PQGM. As shown in Section (ref), our framework naturally extends the systemic risk measure CoVaR taking into account tail network spillover effects, hence after learning PQGM from data, we identify systemically important countries using our new systemic risk measures. We can also provide inference on networks estimated. To simplify the visualization, we provide graphical visualization for the confidence intervals of $\Delta CoVaR$s in Figure (ref) of Section (ref).
We focus on examining financial contagion through the volatility spillovers perspective, i.e. recovering volatility interconnectedness. Engle1993 reported that international stock markets are related through their volatilities
instead of returns. DieboldYilmaz2009 studied the return and volatility
spillovers of 19 countries and found differences in return and volatility
spillovers. \footnote{Modelling the time dependence in volatility is an important issue, although not the focus of this work, there is no uniform agreement on whether there is dependence once taking into account the heteroskedasticity of market returns or shifts in standard errors, see starica2005nonstationarities. }
We use average two-day equity index returns, \footnote{This is to control for the fact that markets in different countries are not open during the same hours. Results are robust to whether using two-day returns or using daily returns as in older versions of our work. Daily returns are also adjusted for weekends and holidays. } September 2009 to September 2013, from Morgan Stanley Capital International (MSCI).
The returns are all translated into dollar-equivalents as of September
6th 2013. \footnote{We calculate returns based on U.S. dollars since these were most frequently used in past work on contagion. } We use absolute returns as a proxy for volatility. \footnote{This has historically been used in the literature. While we do recognize that there are many different methods calculating volatility measures, volatility measuring itself is a large research study area and is outside the scope of the current work.} We have
a total of 45 countries in our sample, there are 21 developed markets
(Australia, Austria, Belgium, Canada, Denmark, France, Germany, Hong
Kong, Ireland, Italy, Japan, Netherlands, New Zealand, Norway, Portugal,
Singapore, Spain, Sweden, Switzerland, the United Kingdom, the United
States), 21 emerging markets (Brazil, Chile, Mexico, Greece, Israel,
China, Colombia, Czech Republic, Egypt, Hungary, India, Indonesia,
Korea, Malaysia, Peru, Philippines, Poland, Russia, Taiwan, Thailand,
Turkey), and 3 frontier markets (Argentina, Morocco, Jordan).
figure[figure omitted — 1,666 chars of source]
Contagion Networks.
Figure (ref) provides a full-sample analysis of global volatility network spillovers
at different tails. The networks are estimated via Algorithm (ref).
We denote 10% quantile as Low Tail, 50% quantile
as Median, 90% quantile as Up Tail. Results learnt from both PQGMs and GGM are presented. GGM or "Gaussian Graph" in Figure (ref) means the graph is estimated via graphical lasso (e.g., Friedman2008), and the final graph is chosen by Extended Bayesian Information Criterion (ebic), see Foygel2010.
Our purpose is to show the usefulness of PQGM in representing nonlinear
tail interdependence allowing for heteroscedasticity and to show that
PQGM can measure correlation asymmetry through looking at the tails
of the distribution (not specific to any model).
There are significant differences in the network structure in terms
of volatility spillovers when using PQGM and GGM. PQGM permits asymmetries in correlation dynamics, suited to investigate the presence of asymmetric responses. We find significant increase interdependence
at the up tail between the volatility series, that is we find downside correlations (high volatility) are much larger than
upside correlations (low volatility). This confirms findings in the finance
literature that financial markets become more interdependent during
high volatility periods.
We also find if two countries locate in the same geographic region,
with many similarities in terms of market structure and history, they
tend to be more closely connected (homophily effect as stated in network terminology),
while two economies locate in separate geographic regions are less
likely directly connected. In addition, we find among European Union member countries,
Germany appears to play a major role in the transmission of shocks
to others; while in Asia, Hong Kong, Thailand, and Singapore appear
to play major roles; and among all the north and south American countries,
Canada and US play major roles.
Systemic Risk.
With the estimated network, we can use different network statistics to measure the systemic risk contributions. Below we focus on the modified $\Delta CoVaR$ measure mentioned in Section (ref).
Figure (ref) provides German's $\Delta CoVaR$s with $\tau = 0.9$ and their $90\%$ uniform confidence intervals obtained via Corollary (ref). It reconfirms that France, Italy and UK contribute the most to German's $\Delta CoVaR$, means conditional on those countries being under distress relative to their median states, German would be affected the most. It is also interesting to find that other countries such as Netherlands can also have effects on German's $\Delta CoVaR$ although less statistically significant in terms of the magnitude of the effect.
figure[figure omitted — 175 chars of source]
figure[figure omitted — 167 chars of source]
In addition, we present net-$\Delta CoVaR$ discussed in Section (ref) with $\tau = 0.9$, i.e. the Up Tail, in Figure (ref) which shows that: globally, total
volatility spillovers from Germany and France to the
others are much larger than total volatility spillovers from the others
to them, and their net-$\Delta CoVaR$ are positive. Both Greece
and Spain have negative net-$\Delta CoVaR$.
appendix\section{Implementation Details of Algorithms}
This section provides details of the algorithms mentioned in Section (ref). Note, for the weighted-Lasso estimator, the choice of penalty level $\lambda := 1.1n^{-1/2}2\Phi^{-1}(1-\xi/N_n)$ and penalty loading $\widehat \Gamma_\tau= \text{diag}[ \widehat \Gamma_{\tau kk} , k \in [p]\backslash\{j\}]$
is a diagonal matrix defined by the following procedure: (1) Compute the post-Lasso estimator $\widetilde \gamma_{a\tau}^{j}$ based on $\lambda$ and initial values $ \widehat \Gamma_{\tau kk} = {\displaystyle \max_{i\leqslant n}} \|f_{ia\tau}Z^a_i\|_\infty \{{\mathbb{E}_n}[ |f_{a\tau}Z^a_k|^2]\}^{1/2}$. (2) Compute the residuals $\widehat v_{i a\tau j} = f_{ia\tau}(Z_{ij}^a - Z_{i,-j}^a\widetilde \gamma_{a\tau}^{j})$ and update the loadings \begin{equation} \widehat \Gamma_{\tau kk} = \sqrt{{\mathbb{E}_n}[f_{a\tau}^2 |Z_{k}^a \widehat{v}_{a\tau j}|^2]}, \ k \in [p]\backslash\{j\} \end{equation}
and use them to recompute the post-Lasso estimator $\widetilde \gamma_{a\tau}^{j}$. In the case of Algorithm (ref) we can take $N_n=|V|p^3n^3$, in the case of Algorithm (ref) we take $N_n=|V|p^2\{pn^3\}^{1+d_W}$. Denote $\widehat\sigma_{aj}^Z=\{{\mathbb{E}_n}[(Z_j^a)^2]\}^{1/2}$, $\widehat\sigma_{aj}^X = \{{\mathbb{E}_n}[X_{-a,j}^2]\}^{1/2}$, $\widehat\sigma_{a\varpi j}^Z= \{{\mathbb{E}_n}[K_\varpi(W)(Z^a_j)^2]\}^{1/2} $, and $\widehat\sigma_{a\varpi j}^X = \{{\mathbb{E}_n}[K_\varpi(W)X_{-a,j}^2]\}^{1/2}$.
\begin{table}[ht]
Detailed version of Algorithm (ref) (CIQGM)\\
For each $a \in V$, $\tau \in \mathcal{T}$, and $j\in [p]$, perform the following:
\begin{enumerate}
• Run Post-$\ell_1$-quantile regression of $X_a$ on $Z^a$; keep fitted value $ Z_{-j}^a\widetilde \beta_{a\tau,-j}$,
$$\begin{array}{l}
\widehat\beta_{a\tau}\in \arg\min_{\beta} {\mathbb{E}_n}[\rho_\tau(X_{a}-Z^a\beta)]+\lambda_{V\mathcal{T}} \sqrt{\tau(1-\tau)}\sum_{j=1}^p\widehat\sigma_{aj}^Z|\beta_j|\\
\widetilde\beta_{a\tau}\in \arg\min_{\beta} {\mathbb{E}_n}[\rho_\tau(X_{a}-Z^a\beta)] \ : \ \beta_j=0 \ \ \mbox{if} \ |\widehat\beta_{a\tau j}| \leqslant \lambda_{V\mathcal{T}}\sqrt{\tau(1-\tau)}/\widehat \sigma_{aj}^Z.\\
\end{array}$$
• Run Post-Lasso of $f_{a\tau} Z_{j}^a$ on $f_{a\tau} Z_{-j}^a$; keep the residual $\widetilde v_i:=f_{ia\tau}\{Z_{ij}^a-Z_{i,-j}^a\widetilde\gamma_{a\tau}^j\}$,
$$\begin{array}{l}
\widehat\gamma_{a\tau}^j\in \arg\min_\gamma{\mathbb{E}_n}[f_{a\tau}^2(Z_{j}^a-Z_{-j}^a\gamma)^2]+\lambda\|\widehat \Gamma_\tau\gamma\|_1\\
\widetilde\gamma_{a\tau}^j\in \arg\min_\gamma {\mathbb{E}_n}[f_{a\tau}^2(Z_{j}^a-Z_{-j}^a\gamma)^2] \ : \ {\rm support}(\gamma)\subseteq {\rm support}(\widehat\gamma_{a\tau}^j).\\
\end{array}$$
• Run Instrumental Quantile Regression of $X_{a} - Z_{-j}^a\widetilde \beta_{a\tau,-j}$ on $Z_{j}^a$
using $\widetilde v$ as the instrument for $Z_{j}^a$,
$$\begin{array}{l}
\displaystyle \check \beta_{a\tau,j}\in \arg\min_{\alpha \in \mathcal{A}_{a \tau j}} \frac{\{{\mathbb{E}_n}[(1\{X_{a} \leqslant Z_{j}^a\alpha+Z_{-j}^a\widetilde\beta_{a\tau,-j}\}-\tau)\widetilde v]\}^2}{{\mathbb{E}_n}[(1\{X_{a} \leqslant Z_{j}^a\alpha+Z_{-j}^a\widetilde\beta_{a\tau,-j}\}-\tau)^2\widetilde v^2]},
\end{array}$$
with $\mathcal{A}_{a\tau j} = \{ \alpha \in {\mathbb{R}} : |\alpha - \widetilde \beta_{a\tau j}| \leqslant 10/\{\widehat \sigma_{aj}^Z \log n\} \}$.
\end{enumerate}
\end{table}
\begin{table}[ht]
Detailed version of Algorithm (ref) (PQGM)\\
For each $a \in V$, and $\tau \in \mathcal{T}$, perform the following:
\begin{enumerate}
• Run Post-$\ell_1$-quantile regression of $X_a$ on $X_{-a}$,
$$\begin{array}{l}
\widehat\beta_{a\tau}\in \arg\min_{\beta} {\mathbb{E}_n}[\rho_\tau(X_{a}-X_{-a}'\beta)]+\lambda_0\sum_{j\in [d]}\widehat\sigma_{aj}^X|\beta_{j}| \\
\widetilde\beta_{a\tau}\in \arg\min_{\beta} {\mathbb{E}_n}[\rho_\tau(X_{a}-X_{-a}'\beta)] \ : \ \beta_j=0 \ \ \mbox{if} \ |\widehat\beta_{a\tau j}| \leqslant \lambda_{0}/\widehat \sigma_{aj}^X.\\
\end{array}
$$
• Set $\widehat \varepsilon_{ia\tau}= 1\{ X_{ia} \leqslant X_{i,-a}'\tilde{\beta}_{a\tau}\}-\tau$ for $i\in[n]$. Compute the penalty level $\bar{\lambda}_{V\mathcal{T}}$ via ((ref)).
• Run Post-$\ell_1$-quantile regression of $X_a$ on $X_{-a}$,
$$\begin{array}{l}
\widehat\beta_{a\tau} \in \arg\min_{\beta} {\mathbb{E}_n}[\rho_\tau(X_{a}-X_{-a}'\beta)]+\bar{\lambda}_{V\mathcal{T}}\sum_{j\in [d]}\{{\mathbb{E}_n}[\widehat\varepsilon_{a\tau}^2X_{-a,j}^2]\}^{1/2}|\beta_j| \\
\check\beta_{a\tau}\in \arg\min_{\beta} {\mathbb{E}_n}[\rho_\tau(X_{a}-X_{-a}'\beta)] \ : \ \beta_j=0 \ \ \mbox{if} \ |\widehat\beta_{a\tau j}| \leqslant \bar{\lambda}_{V\mathcal{T}}/\{{\mathbb{E}_n}[\widehat\varepsilon_{a\tau}^2X_{-a,j}^2]\}^{1/2}.\\
\end{array}$$
\end{enumerate}
\end{table}
\begin{table}[ht]
Detailed version of Algorithm (ref) ($\mathcal{W}$-Conditional CIQGM)\\
For each $u = (a,\tau,\varpi) \in \mathcal{U} = V\times \mathcal{T}\times \mathcal{W}$, and $j\in [p]$, perform the following:
\begin{enumerate}
• Run Post-$\ell_1$-quantile regression of $X_a$ on $Z^a$; keep fitted value $ Z^a_{-j}\widetilde \beta_{u,-j}$,
$$\begin{array}{l}
\widehat\beta_u\in \arg\min_{\beta} {\mathbb{E}_n}[K_\varpi(W)\rho_\tau(X_{a}-Z^a\beta)]+ \lambda_{u}\sum_{j=1}^{p}\widehat{\sigma}_{a\varpi j}^{Z}|\beta_{j}|\\
\widetilde\beta_u\in \arg\min_{\beta} {\mathbb{E}_n}[K_\varpi(W)\rho_\tau(X_{a}-Z^a\beta)] \ : \ \beta_j=0 \ \ \mbox{if} \ |\widehat\beta_{uj}| \leqslant \lambda_u/\widehat\sigma_{a\varpi j}^Z.\\
\end{array}$$
• Run Post-Lasso of $f_u Z_j^a$ on $f_u Z_{-j}^a$; keep the residual $\widetilde v:=f_u(Z^a_j-Z^a_{-j}\widetilde\gamma^j_u)$,
$$\begin{array}{l}
\widehat\gamma^j_u\in \arg\min_\theta {\mathbb{E}_n}[K_\varpi(W)f_u^2(Z^a_j-Z^a_{-j}\gamma)^2]+\lambda\|\widehat \Gamma_u\gamma\|_1\\
\widetilde\gamma^j_u\in \arg\min_\gamma {\mathbb{E}_n}[K_\varpi(W)f_u^2(Z^a_j-Z^a_{-j}\gamma)^2] \ : \ {\rm support}(\gamma)\subseteq {\rm support}(\widehat\gamma^j_u).\\
\end{array}$$
• Run Instrumental Quantile Regression of $X_{a} - Z^a_{-j}\widetilde \beta_{u,-j}$ on $Z^a_{j}$
using $\widetilde v$ as the instrument,
$$\begin{array}{l}
\displaystyle \check \beta_{uj}\in \arg\min_{\alpha \in \mathcal{A}_{uj}} \frac{\{{\mathbb{E}_n}[K_\varpi(W)(1\{X_{a} \leqslant Z^a_j\alpha+Z^a_{-j}\widetilde\beta_{u,-j}\}-\tau)\widetilde v]\}^2}{{\mathbb{E}_n}[K_\varpi(W)(1\{X_{a} \leqslant Z^a_j\alpha+Z^a_{-j}\widetilde\beta_{u,-j}\}-\tau)^2\widetilde v^2]}
\end{array}$$
where $\mathcal{A}_{uj}:= \{ \alpha \in {\mathbb{R}} : |\alpha-\widetilde\beta_{uj}|\leqslant 10/\{\widehat\sigma_{a\varpi j}^Z \log n\} \}$.
\end{enumerate}
\end{table}
\begin{table}[ht]
Detailed version of Algorithm (ref) ($\mathcal{W}$-Conditional PQGM)\\
For each $u = (a,\tau,\varpi) \in \mathcal{U} = V\times \mathcal{T}\times \mathcal{W}$ perform the following:
\begin{enumerate}
• Run Post-$\ell_1$-quantile regression of $X_a$ on $X_{-a}$,
$$\begin{array}{l}
\widehat\beta_u\in \arg\min_{\beta} {\mathbb{E}_n}[K_\varpi(W)\rho_\tau(X_{a}-X_{-a}'\beta)]+\lambda_{0\mathcal{W}}\sum_{j\in[d]}\widehat{\sigma}_{a\varpi j}^{X}|\beta_{j}|\\
\widetilde\beta_u\in \arg\min_{\beta} {\mathbb{E}_n}[K_\varpi(W)\rho_\tau(X_{a}-X_{-a}'\beta)] : \ \beta_j=0 \ \ \mbox{if} \ |\widehat\beta_{uj}| \leqslant \lambda_{0\mathcal{W}}/\widehat\sigma_{a\varpi j}^X. \\
\end{array}$$
• Set $\widehat\varepsilon_{iu}=1\{X_{ia} \leqslant X_{i,-a}'\widetilde\beta_{u}\}-\tau$ for $i\in[n]$, compute $\bar{\lambda}_{V\mathcal{T}\mathcal{W}}$ via ((ref)).\\
• Run Post-$\ell_1$-quantile regression of $X_a$ on $X_{-a}$,
$$\begin{array}{l}
\widehat\beta_u\in \arg\min_{\beta} {\mathbb{E}_n}[K_\varpi(W)\rho_\tau(X_{a}-X_{-a}'\beta)]+\bar{\lambda}_{V\mathcal{T}\mathcal{W}} \sum_{j\in[d]}\{{\mathbb{E}_n}[K_\varpi(W)\widehat\varepsilon_u^2X_{-a,j}^2]\}^{1/2}|\beta_j|.\\
\check\beta_{u}\in \arg\min_{\beta} {\mathbb{E}_n}[K_\varpi(W)\rho_\tau(X_{a}-X_{-a}'\beta)] \ : \ \beta_j=0 \ \ \mbox{if} \ |\widehat\beta_{uj}| \leqslant \bar{\lambda}_{V\mathcal{T}\mathcal{W}}/\{{\mathbb{E}_n}[K_\varpi(W)\widehat\varepsilon_u^2X_{-a,j}^2]\}^{1/2}.\\
\end{array}$$
\end{enumerate}
\end{table}
\section{Simulations of Quantile Graphical Models}
In this section, we perform numerical examples to illustrate the
performance of the estimators proposed for QGMs. We will consider several different
designs. In order to compare with other proposals we will consider both
Gaussian and non-Gaussian examples.
\subsection{Isotropic Non-Gaussian Example}
In general, the equivalence between a zero in the inverse covariance matrix and
a pair of conditional independent variables will break down for non-gaussian
distributions. The nonparanormal graphical models extends Gaussian Graphical Models
to Semiparametric Gaussian Copula models by transforming the variables
with smooth functions. We illustrate the applicability of CIQGM in representing
the conditional independence structure of a set of variables when the random variables
are not even jointly nonparanormal.
Consider i.i.d. copies of an $d$-dimensional random vector $\tilde{X}_V = (\tilde{X}_{1},\ldots, \tilde{X}_{d-1},\tilde{X}_{d})$
from the following multivariate normal distribution,
$\tilde{X}_V \sim N(0,I_{d\times d})$,
where $I_{d\times d}$ is the identity matrix. Further, we generate
\begin{equation}
X_d = -\ensuremath{\sqrt{\frac{2}{3\pi-2}}}+\ensuremath{\sqrt{\frac{\pi}{3\pi-2}}} \tilde{X}_{d-1}^{2}\vert \tilde{X}_{d}\vert.
\end{equation}
It follows that $\mathrm{E}[X_d]=\sqrt{\frac{\pi}{3\pi-2}}(\mathrm{E}[\vert \tilde{X}_{d}\vert]-\sqrt{2/\pi})=0$
and $\mathrm{Var}(X_d)=\frac{\pi}{3\pi-2}(\mathrm{E}[\tilde{X}_{d}^{2}\cdot \tilde{X}_{d-1}^{4}]-\frac{2}{\pi})=1$.
In addition, equation ((ref)) is a location-scale-shift
model in which the conditional median of the response is zero while
quantile functions other than the median are nonzero. We define
vector $X_V$ as
$$
X_V=(X_d, \tilde{X}_{1},...,\tilde{X}_{d-1})'.
$$
In this new set of variables, only $X_d$ and $\tilde{X}_{d-1}$ (i.e., node $1$ and $15$, when $d=15$) are not conditionally independent. Nonetheless, the
covariance matrix of $X_V$ is still $I_{d\times d}$.
Next we consider an example with $n=300$ and
$d=15$. We show graphs, in Figure (ref) and (ref), estimated by both CIQGM(s) and GGMs in this non-Gaussian setting.
\begin{figure}[ht]
\begin{minipage}{0.9\textwidth}
\begin{subfigure}[b]{0.3\textwidth}
\caption[]
{{ CIQGM(0.2)}}
\end{subfigure}
\begin{subfigure}[b]{0.3\textwidth}
\caption[]
{{ CIQGM(0.5)}}
\end{subfigure}
\begin{subfigure}[b]{0.3\textwidth}
\caption[]
{{ CIQGM(0.8)}}
\end{subfigure}
\vskip\baselineskip
\begin{subfigure}[b]{0.3\textwidth}
\caption[]
{{ CIQGM-Union}}
\end{subfigure}
\begin{subfigure}[b]{0.3\textwidth}
\caption[]
{{ Gaussian}}
\end{subfigure}
\begin{subfigure}[b]{0.3\textwidth}
\caption[]
{{ Nonparanormal}}
\end{subfigure}
\caption[ ]
{ QGM and GGM}
\end{minipage}
\end{figure}
In Figure (ref), Gaussian means the graph is estimated by using graphical lasso without any transformation
of $X_V$, and the final graph is chosen by Extended Bayesian
Information Criterion (ebic), see Foygel2010. Nonparanormal means the graph
is estimated using graphical lasso (likelihood based approach) with nonparanormal transformation
of $X_V$, see Liu2009, and again the final graph is chosen
by ebic. Both graphs are estimated using R-package \textbf{huge}.
In Figure (ref), as a robustness check, we also compare results produced by CIQGM with those produced by neighborhood
selection methods (pseudo-likelihood approach), e.g. TIGER of Liu2012b in R-package \textbf{flare}
the left graph is when choosing the turning parameter to be $\sqrt{\frac{\log d}{n}}$
while the right graph is when choosing the tuning parameter to be
$2\sqrt{\frac{\log d}{n}}$. Throughout, we use Tiger2
represent TIGER with penalty level $2\sqrt{\frac{\log d}{n}}$. As expected, GGM cannot detect the correct dependence structure when
the joint distribution is non-Gaussian while CIQGM can still represent
the right conditional independence structure.
\begin{figure}[ht]
\begin{minipage}{0.8\textwidth}
\begin{subfigure}[b]{0.4\textwidth}
\caption[]
{{ Tiger1}}
\end{subfigure}
\begin{subfigure}[b]{0.4\textwidth}
\caption[]
{{ Tiger2}}
\end{subfigure}
\caption[ ]
{ TIGER}
\end{minipage}
\end{figure}
\subsection{Gaussian Examples}
\subsubsection{Graph Recovery}
In this subsection, we start with comparing the numerical performance of QGM and other
methods, e.g. TIGER of Liu2012b and graphical lasso algorithm
(Glasso) of Friedman2008, in graph recovery using simulated datasets with different pairs of $(n,d)$. We start with one simulation for illustration purpose (the results are summarized in Figure (ref)), and then we show the performance of QGM through estimated degree distribution with 100 simulations (the results are summarized in Figure (ref)).
We mainly consider the Hub
graph, as mentioned in Liu2012b, which also corresponds to
the star network mentioned in Acemoglu2010,Acemoglu2013. In line with Liu2012b, we generate a $d$-dimensional sparse
graph $G^I=(V,E^I)$ represents the conditional independence structure
between the variables. In our simulations, we consider 12 settings
to compare these methods: (A) $n=200$, $d=10$; (B) $n=200$,
$d=20$; (C) $n=200$, $d=40$; (D) $n=400$, $d=10$; (E)
$n=400$, $d=20$; (F) $n=400$, $d=40$;
(G) $n=200$, $d=100$; (H) $n=200$, $d=200$; (I) $n=200$,
$d=400$; (J) $n=400$, $d=100$; (K) $n=400$, $d=200$; (L)
$n=400$, $d=400$. We adopt the following model for generating undirected
graphs and precision matrices.
\textbf{Hub graph.} The $d$ nodes are evenly partitioned into $d/20$
(or $d/10$ when $d<20$) disjoint groups with each group contains
$20$ (or $10$) nodes. Within each group, one node is selected as
the hub and we add edges between the hub and the other $19$ (or $9$)
nodes in that group. For example, the resulting graph has $190$ edges
when $d=200$ and $380$ edges when $d=400$. Once the graph is obtained,
we generate an adjacency matrix $E^I$ by setting the nonzero
off-diagonal elements to be 0.3 and the diagonal elements to be 0.
We calculate its smallest eigenvalue $\Lambda_{\min}(E^I)$.
The precision matrix is constructed as
\begin{equation}
\Theta=\mathbf{D}[E^I+(\vert\Lambda_{\min}(E^I)\vert+0.2)\cdot I_{d\times d}]\mathbf{D}
\end{equation}
where $\mathbf{D}\in\mathbb{R}^{d\times d}$ is a diagonal matrix
with $\mathbf{D}_{jj}=1$ for $j=1,...,d/2$ and $\mathbf{D}_{jj}=1.5$
for $j=d/2+1,...,d$. The covariance matrix $\Sigma:=\Theta^{-1}$is
then computed to generate the multivariate normal data: $X_{1},....,X_{d}\sim N(0,\Sigma)$.
Below we provide simulation results using different estimators: PQGM\footnote{Given the graphs are generated from multivariate Gaussian distribution we can use PQGM to simplify the computation.}, TIGER
and Glasso. We start with one simulation as an illustration:
\begin{figure}[ht]
\begin{minipage}{0.95\textwidth}
\begin{subfigure}[b]{0.3\textwidth}
\caption[]
{{ $n = 200, d = 10$}}
\end{subfigure}
\begin{subfigure}[b]{0.3\textwidth}
\caption[]
{{ $n = 200, d = 20$}}
\end{subfigure}
\begin{subfigure}[b]{0.3\textwidth}
\caption[]
{{ $n = 200, d = 40$}}
\end{subfigure}
\vskip\baselineskip
\begin{subfigure}[b]{0.3\textwidth}
\caption[]
{{ $n = 400, d = 10$}}
\end{subfigure}
\begin{subfigure}[b]{0.3\textwidth}
\caption[]
{{ $n = 400, d = 20$}}
\end{subfigure}
\begin{subfigure}[b]{0.3\textwidth}
\caption[]
{{ $n = 400, d = 40$}}
\end{subfigure}
\caption[ ]
{ One simulation}
\end{minipage}
\end{figure}
\begin{figure}[ht]\ContinuedFloat
\begin{minipage}{0.95\textwidth}
\begin{subfigure}[b]{0.3\textwidth}
\caption[]
{{ $n = 200, d = 100$}}
\end{subfigure}
\begin{subfigure}[b]{0.3\textwidth}
\caption[]
{{ $n = 200, d = 200$}}
\end{subfigure}
\begin{subfigure}[b]{0.3\textwidth}
\caption[]
{{ $n = 200, d = 400$}}
\end{subfigure}
\begin{subfigure}[b]{0.3\textwidth}
\caption[]
{{ $n = 400, d = 100$}}
\end{subfigure}
\begin{subfigure}[b]{0.3\textwidth}
\caption[]
{{ $n = 400, d = 200$}}
\end{subfigure}
\begin{subfigure}[b]{0.3\textwidth}
\caption[]
{{ $n = 400, d = 400$}}
\end{subfigure}
\caption[ ]
{ One simulation (Cont.)}
\end{minipage}
\end{figure}
Figure (ref) shows that: for the low dimensional cases, $d = {10, 20, 40}$, $n$ is large compared to $d$, CIQGM is comparable to TIGER and both are better than Glasso in terms of false positives; for the high dimensional cases, $d=\{100,200,300\}$, we can compare the performance of different graph estimators through looking at the “denseness” of the estimated graph (e.g., whether it is even or not), and again, both CIQGM and TIGER perform well in terms of graph recovery as compared to Glasso, and their performance are getting better when $n$ is increasing.
In what follows, Figure (ref) shows the degree distribution of true graph, the estimated ones, and the standard deviations of the degree difference (between the true graph and the estimated ones). It is based on simulations of Hub graph with $n=500$ and $d=40$. Simulated 100 times.
\begin{figure}[ht]
\begin{minipage}{0.9\textwidth}
\tabularnewline
\caption{Upper panel shows the degree distribution of the true graphs. Middle panel shows the degree distribution of the estimated graphs. Bottom panel shows the standard deviations of the degree difference (between the true graph and the estimated ones). Hub graph with $n=500$ and $d=40$. Simulated 100 times. }
\end{minipage}
\end{figure}
\subsubsection{Inference}
In this subsection Table (ref) shows the numerical performance of CIQGM, based on Algorithm (ref), on estimating Erd\H{o}s-R{\'e}nyi random graphs. More precisely, we construct approximate $90\%$ confidence intervals for $\beta_{ab}$ with $\tau = 0.5$, and we report the coverage probabilities. Note, in the jointly Gaussian distributed case, we have closed form solution of $\beta_{ab}$ as shown in Example (ref).
\textbf{Erd\H{o}s-R{\'e}nyi random graph.}
We add an edge between each pair of nodes with probability $\sqrt{\log d/n}/d$ independently.
Once the graph is obtained, we construct the adjacency matrix $E^I$ and generate the precision matrix $\Theta$ using ((ref)) but setting $\mathbf{D}_{jj}=1$ for $j=1,...,d/2$ and $\mathbf{D}_{jj}=1.5$ for $j=d/2+1,...,d$. We then invert $\Theta$ to get the covariance matrices $\Sigma:=\Theta^{-1}$ and generate the multivariate Gaussian data: $X_{1},....,X_{d}\sim N(0,\Sigma)$.
\begin{table}[ht]
\begin{minipage}{\textwidth}
\begin{threeparttable}
\begin{center}
\caption{ Erd\H{o}s-R{\'e}nyi Random Graph}
\begin{tabular}{ccccccc}
\hline
\hline
& $(a,b)$ & $n=200$ & & $n=500$ & & $n=1000$\tabularnewline
\hline
d =20 & (1,20) & 84.0 & & 87.5 & & 91.5\tabularnewline
& (10,11) & 84.0 & & 88.0 & & 92.5\tabularnewline
& (19, 20) & 86.5 & & 86.0 & & 90.0\tabularnewline
& ACP & 86.3 & & 89.4 & & 89.8\tabularnewline
& & & & & & \tabularnewline
d=50 & (1,50) & 86.5 & & 93.0 & & 90.5\tabularnewline
& (25,26) & 88.0 & & 87.0 & & 91.0\tabularnewline
& (49, 50) & 87.5 & & 90.5 & & 87.5\tabularnewline
& ACP & 86.7 & & 89.4 & & 90.1\tabularnewline
& & & & & & \tabularnewline
d=100 & (1,100) & 82.5 & & 81 & & 89.0\tabularnewline
& (50,51) & 85.0 & & 86 & & 92.0\tabularnewline
& (99, 100) & 78.5 & & 84 & & 87.0\tabularnewline
& ACP & 86.5 & & 86.8 & & 90\tabularnewline
\hline
\end{tabular}
\begin{tablenotes}
• $(a,b)$, coverage probability for $\beta_{ab}$; ACP, average coverage
probability for $\beta_{ab}$, with $a\in V$, $b\in V\backslash\{a\}$. Simulated 200 times.
\end{tablenotes}
\end{center}
\end{threeparttable}
\end{minipage}
\end{table}
\section{Proofs of Section 4}
\begin{proof}[Proof of Theorem (ref)]
By Lemma (ref), under Condition CI, for any $\theta$ such that $\|\theta\|_0\leqslant Cs\ell_n$, $\ell_n \to\infty$ slowly, we have that $$\|\sqrt{f_u}Z^a\theta\|_{n,\varpi}/\{{\mathrm{E}}[K_\varpi(W)f_u(Z^a\theta)^2]\}^{1/2} = 1+o_P(1).$$ Moreover, ${\mathrm{E}}[K_\varpi(W)f_u(Z^a\theta)^2]\geqslant {\underline{f}}_u {\mathrm{E}}[K_\varpi(W)(Z^a\theta)^2]$, ${\mathrm{E}}[K_\varpi(W)(Z^a\theta)^2]={\mathrm{E}}[(Z^a\theta)^2| \varpi]{\mathrm{P}}(\varpi)$, and ${\mathrm{E}}[(Z^a\theta)^2\mid \varpi] \geqslant c\|\theta\|^2$ by Condition CI. Lemma (ref) further implies that the ratio of the minimal and maximal eigenvalues of order $s\ell_n$ are bounded away from zero and from above uniformly over $\varpi\in \mathcal{W}$ and $a\in V$ with probability $1-o(1)$. Therefore, since $c\{{\mathrm{P}}(\varpi)\}^{1/2}\|\delta\|_1 \leqslant \|\delta\|_{1,\varpi} \leqslant C\{{\mathrm{P}}(\varpi)\}^{1/2}\|\delta\|_1$, we have $\kappa_{u,2\mathbf{c}} \geqslant c$ uniformly over $u\in\mathcal{U}$ with the same probability for $n$ large enough, see for instance BickelRitovTsybakov2009.
To establish rates of convergence of the estimator obtained in Step 1 we will apply Lemma (ref). Consider the events $\Omega_1, \Omega_2,$ and $\Omega_3$ as defined in ((ref)), ((ref)) and ((ref)). By the choice of $\lambda_u$ we have ${\mathrm{P}}(\Omega_1) \geqslant 1-o(1)$. By Condition CI with $\bar{R}_{u\xi} \leqslant Cs \log(p|V|n)/n$ and Lemma (ref) we have ${\mathrm{P}}(\Omega_2) \geqslant 1-o(1)$. Moreover, ${\mathrm{P}}(\Omega_3) \geqslant 1-o(1)$ by Lemma (ref) with $t_3 \leqslant C n^{-1/2} \sqrt{ (1+d_W)\log(p|V|nL_f) }$.
Using the same argument (with $Z^a$ replacing $X_{-a}$) as in ((ref)), ((ref)), and ((ref)), for
$$\delta\in A_u := \Delta_{\varpi,2\mathbf{c}}\cup\{v:\|v\|_{1,\varpi}\leqslant 2\mathbf{c}\bar R_{u\xi}/\lambda_u, \|\sqrt{f_u}Z^av\|_{n,\varpi} \geqslant C\sqrt{s(1+d_W)\log(p|V|n)/n}/\kappa_{u,2\mathbf{c}} \},$$
with the restricted set defined as $\Delta_{u, 2\tilde\mathbf{c}} = \{ \delta : \|\delta_{T^{c}_u}\|_1\leqslant 2\tilde\mathbf{c} \|\delta_{T_u}\|_1\}$ for $u\in \mathcal{U}$. we have $ \bar q_{A_u} \geqslant c ({\underline{f}}_\mathcal{U}^{3/2}/\bar f')\mu_\mathcal{W}^{1/2}/ \{\sqrt{s} \max_{a\in V, i\leqslant n}\|Z_i^a\|_\infty\}$ where $\max_{a\in V, i\leqslant n}\|Z_i^a\|_\infty \lesssim_P M_n$. Thus the conditions on $\bar q_{A_u}$ are satisfied since Condition CI assumes $M_n^2s^2\log(p|V|n) \leqslant n \mu_\mathcal{W}{\underline{f}}_\mathcal{U}^3$. The conditions on the approximation error are assumed in Condition CI.
Therefore, setting $\xi=1/\log n$, by Lemma (ref) we have uniformly over $u=(a,\tau,\varpi)\in \mathcal{U}$
\begin{equation}\begin{array}{rl}
\|\sqrt{f_u}Z^a(\widehat \beta_u -\beta_u)\|_{n,\varpi} & \lesssim \sqrt{(1+(t_3/\lambda_u)\bar R_{u\xi}} + (\lambda_u + t_3)\sqrt{s} \lesssim \sqrt{\frac{s(1+d_W)\log(p|V|n)}{n\tau(1-\tau)}} \\
\|\widehat \beta_u -\beta_u\|_{1,\varpi} & \lesssim s\sqrt{\frac{(1+d_W)\log(p|V|n)}{n}} \end{array}\end{equation}
here we used that $\lambda_u \leqslant C\sqrt{\frac{(1+d_W)\log(p|V|n)}{n}}$. Indeed by Lemma (ref) with $\tilde x_{ij} = K_\varpi(W_i)Z^a_{ij}$ and $\widehat\sigma_j \geqslant c{\mathrm{P}}(\varpi)^{1/2}$, we can bound $\Lambda_{a\tau \varpi}(1-\xi/\{|V|n^{1+2d_W}\}| X_{-a},W)$ under $M_n^2 \log(p|V|n/\{\tau(1-\tau)\}) = o(n \tau(1-\tau)\mu_\mathcal{W})$ for all $\tau\in\mathcal{T}$ and $\varpi \in \mathcal{W}$, and the bound on $\lambda_u$ follows from the union bound.
Let $\delta_u = \widehat \beta_u -\beta_u$. By triangle inequality it follows that
\begin{equation} \{{\mathrm{E}}[K_\varpi(W)f_u(Z^a\delta_u )^2]\}^{1/2} \leqslant \|\sqrt{f_u}Z^a\delta_u \|_{n,\varpi} + \|\delta_u\|_{1}\{ |({\mathbb{E}_n}-{\mathrm{E}})[K_\varpi(W)f_u(Z^a\delta_u)^2]/\|\delta_u\|_{1}^2]| \}^{1/2} \end{equation}
and the last term can be bounded by
$$\begin{array}{rl}{\displaystyle \sup_{\|\delta\|_{1} \leqslant 1}} |({\mathbb{E}_n}-{\mathrm{E}})[K_\varpi(W)f_u(Z^a\delta)^2/\|\delta\|_{1}^2]| & \leqslant \max_{k,j}|({\mathbb{E}_n}-{\mathrm{E}})[K_\varpi(W)f_uZ^a_{k}Z^a_{j}]| \\
&\lesssim \sqrt{\frac{(1+d_W)\log(p|V|n)}{n}}\end{array}$$
with probability $1-o(1)$ by Lemma (ref) under our conditions.
Combining the relations above with ((ref)), under $(1+d_W)s^2\log(p|V|n) = o(n)$ we have uniformly over $u\in\mathcal{U}$
$$\begin{array}{rl}
\|\delta_u\| & \lesssim \{{\mathrm{E}}[(Z^a\delta_u)^2\mid\varpi]\}^{1/2} \lesssim \{{\mathrm{P}}(\varpi)\}^{-1/2} \{{\mathrm{E}}[K_\varpi(W)(Z^a\delta_u)^2]\}^{1/2} \\
& \lesssim \{{\mathrm{P}}(\varpi){\underline{f}}_u\}^{-1/2}\{{\mathrm{E}}[K_\varpi(W)f_u(Z^a\delta_u)^2]\}^{1/2} \\
& \leqslant \{{\mathrm{P}}(\varpi){\underline{f}}_u\}^{-1/2}\|\sqrt{f_u}Z^a\delta_u\|_{n,\varpi} + \{{\mathrm{P}}(\varpi){\underline{f}}_u\}^{-1/2}\sqrt[4]{\frac{(1+d_W)\log(p|V|n)}{n}}\|\delta_u\|_{1}\\
& \leqslant C\sqrt{\frac{s(1+d_W)\log(p|V|n)}{n{\underline{f}}_u {\mathrm{P}}(\varpi) }}.\end{array}$$
given $\|\delta_u\|_{1}\leqslant \|\delta_u\|_{1,\varpi}/ {\mathrm{P}}(\varpi)^{1/2}$, ((ref)), and $s^2(1+d_W)\log(p|V|n) = o(n \mu_\mathcal{W}^4{\underline{f}}_\mathcal{U}^2)$ assumed in Condition CI.
Finally, let $\widehat\beta^{\bar\lambda}_u$ obtained by thresholding the estimator $\widehat\beta_u$ with $\bar\lambda:= \sqrt{(1+d_W)\log(p|V|n)/n}$ (note that each component is weighted by ${\mathbb{E}_n}[K_\varpi(W)(Z_j^a)^2]^{1/2}$). By Lemma (ref), we have with probability $1-o(1)$
$$
\begin{array}{rl}
\|Z^a(\widehat\beta^{\bar\lambda}_u - \beta_u)\|_{n,\varpi} \lesssim \sqrt{s(1+d_W)\log(p|V|n)/n} \\
\|\widehat\beta^{\bar\lambda}_u - \beta_u\|_{1,\varpi} \lesssim s\sqrt{(1+d_W)\log(p|V|n)/n}\\
|{\rm support}(\widehat\beta^{\bar\lambda}_u)| \lesssim s
\end{array}
$$ by the choice of $\bar\lambda$ and the rates in ((ref))
\end{proof}
\begin{proof}[Proof of Theorem (ref)]
We verify Assumption (ref) and Condition WL for the weighted Lasso model with index set $\mathcal{U} \times [p]$ where $Y_u = K_\varpi(W)Z^a_{j}$, $X_u=K_\varpi(W)Z^a_{-j}$, $\theta_u=\bar \gamma_u^j$, $a_u = (f_u,\bar r_{uj})$, $\bar r_{uj}=K_\varpi(W)Z^a_{-j}(\gamma_u^j-\bar\gamma_u^j)$, $S_{uj} = K_\varpi(W)f_u^2(Z^a_{j}-Z^a_{-j}\gamma_u^j)Z^a_{-j}=K_\varpi(W)f_uv_{uj}Z^a_{-j}$, and $w_u =K_\varpi(W)f_u^2$. We will take $N_n = |V|p^2\{pn^3\}^{1+d_W}$ in the definition of $\lambda$.
We first verify Condition WL. We have ${\mathrm{E}}[S_{ujk}^2] \leqslant \bar f^2{\mathrm{E}}[ |v_{uj}Z^a_{-j,k}|^2] \leqslant \bar f^2\{{\mathrm{E}}[ |v_{uj}|^4|Z^a_{-j,k}|^4]\}^{1/2} \leqslant C$ by the bounded fourth moment condition.
We have that $$\frac{{\mathrm{E}}[|S_{ujk}|^3]^{1/3}}{{\mathrm{E}}[|S_{ujk}|^2]^{1/2}} =\frac{{\mathrm{E}}[|S_{ujk}|^3\mid \varpi]^{1/3}}{{\mathrm{E}}[|S_{ujk}|^2 \mid \varpi]^{1/2}} \{ {\mathrm{P}}(\varpi)\}^{-1/6} = \frac{{\mathrm{E}}[|f_uv_{uj}Z_{-jk}^a|^3\mid \varpi]^{1/3}}{{\mathrm{E}}[|f_uv_{uj}Z^a_{-jk}|^2 \mid \varpi]^{1/2}} \{ {\mathrm{P}}(\varpi)\}^{-1/6} =:M_{uk} $$
By the choice of $N_n$ and $\Phi^{-1}(1-t)\leqslant C\sqrt{\log(1/t)}$, we have $M_{uk}\Phi^{-1}(1-\xi/\{2pN_n\}) \leqslant M_{uk}C(1+d_W)\log^{1/2}(pn |V|) \leqslant C\delta_n n^{1/6}$ where the last inequality holds by Condition CI so Condition WL(i) holds.
To verify Condition WL(ii) we will establish the validity of the choice of $N_n$. We will consider $u=(a,\tau,\varpi)\in \mathcal{U}$ and $u'=(a,\tau',\varpi')\in \mathcal{U}$. By Condition CI we have that \begin{equation}|f_u-f_{u'}| \leqslant L_f\|u-u'\| \ \ \mbox{and} \ \ {\mathrm{E}}[|K_\varpi(W)-K_{\varpi'}(W)|] \leqslant L_K\|\varpi-\varpi'\|.\end{equation}
Further, by Lemma (ref) we have
\begin{equation} \|\gamma_u^j-\gamma_{u'}^j\|\leqslant L_\gamma \{\|u-u'\|+\|\varpi-\varpi'\|^{1/2}\}.\end{equation}
By definition we have $$S_{ujk}-S_{u'jk}=\{K_\varpi(W)f_u^2-K_{\varpi'}(W)f_{u'}^2\}\{Z_j^a-Z^a_{-j}\gamma_u^j\}Z_k^a-K_{\varpi'}(W)f_{u'}^2\{Z^a_{-j}(\gamma_u^j-\gamma_{u'}^j)\}Z_k^a$$ and note that $f_u+f_{u'} \leqslant 2f_u+L\|u-u'\|$, $|Z_j-Z_{-j}^a\gamma_u^j|\cdot|Z_k^a| \leqslant |Z_j|^2+ 2|Z_k^a|^2 + |Z_{-j}^a\gamma_u^j|^2$. Moreover, $$\begin{array}{rl}
|K_\varpi(W)f_u^2-K_{\varpi'}(W)f_{u'}^2| & \leqslant K_\varpi(W)K_{\varpi'}(W)|f_u^2-f_{u'}^2| + (f_u+f_{u'})^2|K_\varpi(W)-K_{\varpi'}(W)|\\
& \leqslant 2 \bar f K_\varpi(W)K_{\varpi'}(W)|f_u-f_{u'}| + 4\bar f^2|K_\varpi(W)-K_{\varpi'}(W)| \\
\end{array}$$
Using these relations we have $|{\mathbb{E}_n}[ S_{ujk}-S_{u'jk} ]| \leqslant (I) + (II)$ where
$$\begin{array}{rl}
(I) & = {\mathbb{E}_n}[|K_\varpi(W)f_u^2-K_{\varpi'}(W)f_{u'}^2|\cdot |\{Z_j^a-Z^a_{-j}\gamma_u^j\}Z_k^a|] \\
& \leqslant \max_{i\leqslant n}\|Z_i^a\|_\infty^2(1+\|\gamma_u^j\|_1) {\mathbb{E}_n}[ 2\bar f K_\varpi(W)K_{\varpi'}(W)|f_u-f_{u'}| + 4\overline{f}^2|K_\varpi(W)-K_{\varpi'}(W)|]\\
& \leqslant (\bar f+\bar f^2) C\sqrt{s}\max_{i\leqslant n}\|Z_i^a\|_\infty^2\{ L_f \|u-u'\|+ {\mathbb{E}_n}[|K_\varpi(W)-K_{\varpi'}(W)|]\}\\
(II) & = {\mathbb{E}_n}[K_{\varpi'}(W)f_{u'}^2|Z^a_{-j}(\gamma_u^j-\gamma_{u'}^j)Z_k^a|] \\
& \leqslant \bar f^2 {\mathbb{E}_n}[ \|Z^a\|_\infty^2 ] \|\gamma_u^j-\gamma_{u'}^j\|_1 \\
& \leqslant \bar f^2 {\mathbb{E}_n}[ \|Z^a\|_\infty^2 ] \sqrt{p} L_\gamma\{ \|u-u'\|+\|\varpi-\varpi'\|^{1/2}\} \\
\end{array}$$
Moreover, we have that $\max_{i\leqslant n} \|Z_i^a\|_\infty^2 \lesssim_P M_n^2$. For $d_\mathcal{U}=\|\cdot\|$, an uniform $\epsilon$-cover of $\mathcal{U}$ satisfies $(6{\text{diam}}(\mathcal{U})/\epsilon)^{1+d_W} \geqslant N(\epsilon,\mathcal{U},\|\cdot\|)$. We will set $1/\epsilon = (1+\bar f^2)\{L_\gamma+L_f\}^2pnM_n^2\log^2(p|V|n)/\{\mu_\mathcal{W}{\underline{f}}_\mathcal{U}^2\} \leqslant pn^3$ so that with probability $1-o(1)$, for any pair $u,u'\in\mathcal{U}$, $\|u-u'\|\leqslant \epsilon$, we have
$$\begin{array}{rl}
|{\mathbb{E}_n}[ S_{ujk}-S_{u'jk} ]| & \lesssim (\bar f+\bar f^2) \sqrt{s}\max_{i\leqslant n}\|Z_i^a\|_\infty^2\{ L_f \epsilon+ {\mathbb{E}_n}[|K_\varpi(W)-K_{\varpi'}(W)|]\} \\
& +\bar f^2 {\mathbb{E}_n}[ \|Z_i^a\|_\infty^2 ] \sqrt{p} L_\gamma\{ \epsilon+ \epsilon^{1/2}\} \\
& \lesssim \delta_n n^{-1/2}\{\mu_\mathcal{W}\underline f_\mathcal{U}^2\}^{1/2} + \sqrt{s}M_n\log(n){\mathbb{E}_n}[|K_\varpi(W)-K_{\varpi'}(W)|]\\
\end{array}$$
by the choice of $\epsilon$.
To control the last term, note that $\mathcal{W}$ is a VC-class of events with VC dimension $d_W$. Thus by Lemma (ref), with probability $1-o(1)$
$$ \begin{array}{rl}
{\mathbb{E}_n}[|K_\varpi(W)-K_{\varpi'}(W)|] & \leqslant |({\mathbb{E}_n}-{\mathrm{E}})[|K_\varpi(W)-K_{\varpi'}(W)|]|+{\mathrm{E}}[|K_\varpi(W)-K_{\varpi'}(W)|] \\
&\displaystyle \leqslant \sup_{\varpi,\varpi'\in\mathcal{W},\|\varpi-\varpi'\|\leqslant \epsilon}|({\mathbb{E}_n}-{\mathrm{E}})[|K_\varpi(W)-K_{\varpi'}(W)|]|+ L_K\epsilon \\
& \lesssim \sqrt{\frac{d_W \log(n/\epsilon)}{n}}\epsilon^{1/2} + \frac{d_W\log(n/\epsilon)}{n} + L_K\epsilon
\end{array}
$$ which yields uniformly over $u\in \mathcal{U}$ and $j\in[p]$
\begin{equation} |{\mathbb{E}_n}[ S_{ujk}-S_{u'jk} ]| \lesssim \delta_n n^{-1/2} \mu_\mathcal{W}^{1/2}{\underline{f}}_\mathcal{U}\end{equation}
under $\sqrt{\epsilon d_W \log(n/\epsilon)}M_n\log n = o(\mu_\mathcal{W}^{1/2}{\underline{f}}_\mathcal{U})$ and $d_W\log(n/\epsilon)M_n\log n=o(n^{1/2}\mu_\mathcal{W}^{1/2}{\underline{f}}_\mathcal{U})$ assumed in Condition CI. In turn this implies
$$ \begin{array}{c}
\displaystyle \sup_{|u-u'|\leqslant \epsilon} \underset{j,k\in[p],j\neq k}{\max}\frac{|{\mathbb{E}_n}[S_{ujk}-S_{u'jk}]|}{{\mathrm{E}}[S_{ujk}^2]^{1/2}} \leqslant \delta_n n^{-1/2} \end{array}$$
since ${\mathrm{E}}[S_{ujk}^2] \geqslant c \mu_\mathcal{W}{\underline{f}}_\mathcal{U}^2$. Using the same choice of $\epsilon$, similar arguments also imply
\begin{equation} \begin{array}{c}
\ \sup_{|u-u'|\leqslant \epsilon}\underset{j,k\in[p],j\neq k}{\max}\frac{|{\mathrm{E}}[S_{ujk}^2-S_{u'jk}^2]|}{{\mathrm{E}}[S_{ujk}^2]} \leqslant \delta_n \end{array}\end{equation}
To establish the last requirement of Condition WL(ii), note that
\begin{equation} \begin{array}{rl}
\sup_{u\in \mathcal{U}} \max_{j,k\in [p], j\neq k} |({\mathbb{E}_n}-{\mathrm{E}})[S_{ujk}^2]| & \leqslant \sup_{u\in \mathcal{U}^\epsilon} \max_{j,k\in [p], j\neq k} |({\mathbb{E}_n}-{\mathrm{E}})[S_{ujk}^2]| + \Delta_n \\
\end{array} \end{equation}
where $ \Delta_n := \sup_{u,u'\in \mathcal{U}, \|u-u'\|\leqslant \epsilon} \max_{j,k \in [p], j\neq k} |({\mathbb{E}_n}-{\mathrm{E}})[S_{ujk}^2]-({\mathbb{E}_n}-{\mathrm{E}})[S_{u'jk}^2]|.$
To bound the first term, we will apply Corollary (ref) with $k=1$, $\widehat \mathcal{U} := \mathcal{U}^\epsilon\times[p]$ and the vector $\{(\bar X)_{uj}=S_{uj}, (u,j)\in \widehat \mathcal{U}\}$. In this case note that
$$K^2={\mathrm{E}}[ \max_{i\leqslant n} \sup_{u\in \mathcal{U}} \max_{j,k\in [p], j\neq k} S_{ujk}^2 ] \leqslant {\mathrm{E}}[ \max_{i\leqslant n} \sup_{u\in\mathcal{U},j\in[p]} |v_{iuj}|^2\|f_{iu}Z_i^a\|_\infty^2] \leqslant \bar f^2 M_n^2L_n^2.$$
Therefore, by Corollary (ref) and Markov inequality, we have with probability $1-o(1)$ that
$$ \begin{array}{rl}
\displaystyle \sup_{u\in \mathcal{U}} \max_{j,k\in [p], j\neq k} |({\mathbb{E}_n}-{\mathrm{E}})[S_{ujk}^2]|
& \leqslant C n^{-1/2}M_nL_n \log^{1/2}(p|V|n)\leqslant C\delta_n \mu_\mathcal{W}{\underline{f}}_\mathcal{U} + \Delta_n \end{array} $$
under $M_n^2L_n^2\log(p|V|n) \leqslant \delta_n n \mu_\mathcal{W}^2{\underline{f}}_\mathcal{U}^2$.
To control $\Delta_n$, note that
$$\begin{array}{rl}
|({\mathbb{E}_n}-{\mathrm{E}})[S_{ujk}^2]-({\mathbb{E}_n}-{\mathrm{E}})[S_{u'jk}^2]| & \leqslant |{\mathbb{E}_n}[S_{ujk}^2-S_{u'jk}^2]|+|{\mathrm{E}}[S_{ujk}^2-S_{u'jk}^2]|\\
& \leqslant {\mathbb{E}_n}[|S_{ujk}-S_{u'jk}|]\sup_{u\in\mathcal{U}}\max_{i\leqslant n}|2S_{iu'jk}|+|{\mathrm{E}}[S_{ujk}^2-S_{u'jk}^2]|\\
& \lesssim \delta_n n^{-1/2} \mu_\mathcal{W}^{1/2}{\underline{f}}_\mathcal{U} \bar f \sup_{u\in\mathcal{U}}\max_{i\leqslant n}|v_{iuj}|\|Z_i^a\|_\infty+\delta_n \mu_\mathcal{W}{\underline{f}}_\mathcal{U}^2\\
& \lesssim \delta_n n^{-1/2} \mu_\mathcal{W}^{1/2}{\underline{f}}_\mathcal{U} M_n L_n \log n+\delta_n \mu_\mathcal{W}{\underline{f}}_\mathcal{U}^2\\
\end{array}$$
with probability $1-o(1)$ where we used ((ref)) and ((ref)). Therefore $\Delta_n \lesssim \delta_n\mu_\mathcal{W}{\underline{f}}_\mathcal{U}^2$ with probability $1-o(1)$ as required.
To verify Assumption (ref)(a), note that $[\partial_\theta M_u(Y_u,X,\theta_u)-\partial_\theta M_u(Y_u,X,\theta_u,a_u)]'\delta= -f_u^2 K_\varpi(W)\bar r_{uj}Z_{-j}^a\delta$, so that by Cauchy-Schwartz, we have
$$\begin{array}{rl}
{\mathbb{E}_n}[\partial_\theta M_u(Y_u,X,\theta_u)-\partial_\theta M_u(Y_u,X,\theta_u,a_u)]'\delta & \leqslant \| f_u\bar r_{uj}\|_{n,\varpi}\| f_u Z_{-j}^a\delta\|_{n,\varpi} \leqslant C_{un} \| f_u Z_{-j}^a\delta\|_{n,\varpi} \end{array} $$
where we choose $C_{un}$ so that $\{ C_{un} \geqslant \max_{j\in[p]} \| f_u\bar r_{uj}\|_{n,\varpi}: u\in \mathcal{U}\}$ with probability $1-o(1)$. To bound $C_{un}$, by Lemma (ref), uniformly over $u\in\mathcal{U}, j\in[p]$ we have with probability $1-o(1)$ $$\|f_u\bar r_{uj}\|_{n,\varpi} = \|f_uZ_{-j}^a(\gamma_u^j-\bar\gamma_u^j)\|_{n,\varpi}\lesssim {\underline{f}}_u\{{\mathrm{P}}(\varpi)\}^{1/2}\{n^{-1} s\log(p|V|n)\}^{1/2}$$
so that setting $C_{un}={\underline{f}}_u\{{\mathrm{P}}(\varpi)\}^{1/2}\{n^{-1} s\log(p|V|n)\}^{1/2}$ suffices.
Next we show that Assumption (ref)(b) holds. First, by ((ref)) and the corresponding bounds, note the uniform convergence of the loadings
$$ \sup_{u\in \mathcal{U}, j, k\in[p], j\neq k} (|{\mathbb{E}_n}[ S_{ujk}^2] - {\mathrm{E}}[S_{ujk}^2] | + |({\mathbb{E}_n}-{\mathrm{E}})[ K_\varpi(W)f_u^2|Z_j^aZ_{-jk}^a|^2]|) \leqslant \delta_n \mu_\mathcal{W}{\underline{f}}_\mathcal{U}^2 $$ so that ${\mathbb{E}_n}[ S_{ujk}^2]/{\mathrm{E}}[S_{ujk}^2]= 1+o_P(1)$. It follows that $\tilde c$ is bounded above by a constant for $n$ large enough. Indeed, uniformly over $u\in\mathcal{U}$, $j\in[p]$, since $c{\underline{f}}_u\leqslant {\mathrm{E}}[|f_uv_{uj}Z^a_k|^2\mid \varpi]^{1/2}\leqslant C{\underline{f}}_u$, with probability $1-o(1)$ we have
$c {\underline{f}}_u{\mathrm{P}}(\varpi)^{1/2} \leqslant \widehat\Psi_{u0jj} \leqslant C {\underline{f}}_u{\mathrm{P}}(\varpi)^{1/2}$ so that $c/C \leqslant \|\widehat\Psi_{u0}\|_\infty\|\widehat\Psi_{u0}^{-1}\|_\infty \leqslant C/c$.
Assumption (ref)(c) follows directly from the choice of $M_u(Y_u,X_u,\theta)=K_\varpi(W)f_u^2(Z_j^a-Z^a_{-j}\theta)^2$ with $\bar q_{A_u} = \infty$.
The result for the rate of convergence then follows from Lemma (ref), namely
\begin{equation} \|f_u X_u'(\widehat\gamma_u^j-\gamma_u^j) \|_{n,\varpi} \lesssim \frac{\|\widehat \Psi_{u0}\|_\infty}{\bar\kappa_{u,2\mathbf{c}}}\sqrt{\frac{s\log(p|V|n)}{n}}+ C_{un} \lesssim \frac{{\underline{f}}_u{\mathrm{P}}(\varpi)^{1/2}}{\bar\kappa_{u,2\mathbf{c}}}\sqrt{\frac{s\log(p|V|n)}{n}} \end{equation}
By Lemma (ref) we have that for sparse vectors, $\|\theta\|_0\leqslant \ell_n s$ satisfies
$$ \|f_uZ^a_{-j}\theta\|_{n,\varpi}^2/{\mathrm{E}}[K_\varpi(W)f_u^2(Z^a_{-j}\theta)^2] = 1+o_P(1)$$
so that $\phi_{{\rm max}}(\ell_n s,uj) \leqslant C{\underline{f}}_u^2{\mathrm{P}}(\varpi)$ and $\widehat s_{uj} \leqslant \min_{m\in \mathcal{M}_u}\phi_{{\rm max}}(m,uj)L_u^2 \leqslant Cs$ provided $L_u^2 \lesssim s\{{\underline{f}}_u^2{\mathrm{P}}(\varpi)\}^{-1}$. Indeed, with probability $1-o(1)$, we have $\|\widehat\Psi_{u0}^{-1}\|_\infty \leqslant C{\underline{f}}_u^{-1}{\mathrm{P}}(\varpi)^{-1/2}$, so that $L_u\lesssim {\underline{f}}_u^{-1}{\mathrm{P}}(\varpi)^{-1/2} \frac{n}{\lambda} \{ C_{un} + L_{un} \}$. Moreover, we can take $C_{un} \lesssim {\underline{f}}_u\{{\mathrm{P}}(\varpi) n^{-1}s\log(p|V|n)\}^{1/2}$, and $L_{un} \lesssim \{n^{-1}s\log(p|V|n)\}^{1/2}$ in Assumption (ref) because
$$\begin{array}{rl}
& |\{{\mathbb{E}_n}[\partial_\gamma M_u(Y_u,X_u,\widehat\gamma_u^j)-\partial_\gamma M_u(Y_u,X_u,\gamma_u^j)]\}'\delta| \\
& = 2|{\mathbb{E}_n}[K_\varpi(W)f_u^2\{X_u'(\widehat\gamma_u^j-\gamma_u^j)\}X_u'\delta]|\\
&\leqslant 2\|f_u X_u'(\widehat\gamma_u^j-\gamma_u^j) \|_{n,\varpi} \|f_u X_u'\delta\|_{n,\varpi}=:L_{un}\|f_u X_u'\delta\|_{n,\varpi},\\
\end{array}$$
where the last inequality hold by ((ref)) since $\bar\kappa_{u,2\mathbf{c}} \geqslant c{\underline{f}}_u\{{\mathrm{P}}(\varpi)\}^{1/2}$. The bound on the restricted eigenvalue $\bar \kappa_{u,2\mathbf{c}}$ holds\footnote{Note that there are two restricted eigenvalues definitions, one used for the quantile regression ($\kappa_{u,2\mathbf{c}}$), and another used here for the weighted lasso ($\bar\kappa_{u,2\mathbf{c}}$). It is a consequence of the use of different norms.} by arguments similar to ((ref)) and using that $\|\delta\|_1\leqslant C\sqrt{s}\|\delta\|$ for any $\delta \in \Delta_{u, 2\mathbf{c}}$, and since for any $\|\delta\|=1$, we have
$$\begin{array}{rl}
c{\underline{f}}_u{\mathrm{P}}(\varpi) &\leqslant {\mathrm{E}}[K_\varpi(W)f_u(Z^a\delta)^2] \\
& \leqslant \{{\mathrm{E}}[K_\varpi(W)f_u^2(Z^a\delta)^2]\}^{1/2} \{{\mathrm{E}}[K_\varpi(W)(Z^a\delta)^2]\}^{1/2} \\
& \leqslant \{{\mathrm{E}}[K_\varpi(W)f_u^2(Z^a\delta)^2]\}^{1/2} C\{{\mathrm{P}}(\varpi)\}^{1/2} \end{array}$$
where the first inequality follows from the definition of ${\underline{f}}_u$, $\|\delta\|=1$, and Condition CI, so that we have $\{{\mathrm{E}}[K_\varpi(W)f_u^2(Z^a\delta)^2]\}^{1/2} \geqslant c'{\underline{f}}_u\{{\mathrm{P}}(\varpi)\}^{1/2}$.
Return to the rate of convergence we have by ((ref)) and $\bar\kappa_{u,2\mathbf{c}} \geqslant c{\underline{f}}_u\{{\mathrm{P}}(\varpi)\}^{1/2}$ that
\begin{equation} \|f_u X_u'(\widehat\gamma_u^j-\gamma_u^j) \|_{n,\varpi} \lesssim \frac{{\underline{f}}_u{\mathrm{P}}(\varpi)^{1/2}}{\bar\kappa_{u,2\mathbf{c}}}\sqrt{\frac{s\log(p|V|n)}{n}} \lesssim \sqrt{\frac{s\log(p|V|n)}{n}} \end{equation}
and the result follows by noting that $\|f_u X_u'(\widehat\gamma_u^j-\gamma_u^j) \|_{n,\varpi} \geqslant c {\underline{f}}_u{\mathrm{P}}(\varpi)^{1/2}\|\widehat\gamma_u^j-\gamma_u^j\|$ with probability $1-o(1)$ by arguments similar to ((ref)) under Condition CI.
The sparsity result follows from Lemma (ref). The result for Post Lasso follows from Lemma (ref) under the growth requirements in Condition CI.
\end{proof}
\begin{proof}[Proof of Theorem (ref)]
We will verify Assumptions (ref) and (ref), and the result follows from Theorem (ref).
The estimate of the nuisance parameter is constructed from the estimators in Steps 1 and 2 of the Algorithm.
For each $u=(a,\tau,\varpi)\in \mathcal{U}$ and $j\in [p]$, let $W_{uj}=(W,X_a,Z^a,v_{uj},r_u)$, where $v_{uj} = f_u(Z_j^a - Z_{-j}^a\gamma_u^j)$ and let $\theta_{uj} \in \Theta_{uj} = \{ \theta \in {\mathbb{R}} : |\theta-\beta_{uj}|\leqslant c/\log n \}$ (Assumption (ref)(i) holds). The score function is $$ \psi_{uj}(W_{uj},\theta,\eta_{uj}) = K_\varpi(W)\{ \tau - 1\{X_a \leqslant Z_j^a\theta+Z^a_{-j}\beta_{u,-j}+r_u\}\}f_u(Z_j^a - Z_{-j}^a\gamma_u^j)$$ where the nuisance parameter is $\eta_{uj}=(\beta_{u, -j},\gamma_u^j,r_u)$ and the last component is a function $r_{u}=r_{u}(X)$. Recall that $K_\varpi(W) \in \{0,1\}$ and let $
a_n = \max(n,p,|V|)$. Define the nuisance parameter set
${\mathcal{H}}_{uj} = \{ \eta =(\eta^{(1)},\eta^{(2)}, \eta^{(3)}): \|\eta - \eta_{uj}\|_e \leqslant \tau_n\}$ where
$\|\eta-\eta_{uj}\|_e = \|(\delta_\eta^{(1)},\delta_\eta^{(2)},\delta_\eta^{(3)})\|_e = \max \{ \|\delta_\eta^{(1)}\|, \|\delta_\eta^{(2)}\|, {\mathrm{E}}[|\delta_\eta^{(3)}|^2]^{1/2}\}$, and
$$ \tau_n := C\sup_{u\in\mathcal{U}} \frac{1}{1\wedge {\underline{f}}_\mathcal{U}}\sqrt{\frac{s \log a_n}{n \mu_\mathcal{W}}}$$
The differentiability of the mapping $(\theta,\eta) \in \Theta_{uj}\times {\mathcal{H}}_{uj}\mapsto {\mathrm{E}}\psi_{uj}(W_{uj},\theta,\eta)$ follows from the differentiability of the conditional probability distribution of $X_a$ given $X_{V\backslash \{a\}}$ and $\varpi$. Let $\eta=(\eta^{(1)},\eta^{(2)},\eta^{(3)})$, $\delta_\eta = (\delta_\eta^{(1)},\delta_\eta^{(2)},\delta_\eta^{(3)})$, and $\theta_{\bar r} = \theta + \bar{r}\delta_\theta$, $\eta_{\bar{r}} = \eta + \bar{r} \delta_\eta$.
To verify Assumption (ref)(v)(a) with $\alpha = 2$, for any $(\theta,\eta),(\bar \theta,\bar\eta)\in \Theta_{uj}\times {\mathcal{H}}_{uj}$ note that $f_{X_a\mid X_{-a},\varpi}$ is uniformly bounded from above by $\bar {f}$, therefore
$$\begin{array}{rl}
{\mathrm{E}}[ \{ \psi_{uj}(W_{uj},\theta,\eta)-\psi_{uj}(W_{uj},\bar \theta,\bar\eta)\}^2]^{1/2} \\
\leqslant \bar{f}{\mathrm{E}}[|Z^a_{-j}(\eta^{(2)}-\bar \eta^{(2)})|^2]^{1/2} + \bar{f}^2{\mathrm{E}}[ (Z^a_j-Z^a_{-j}\bar\eta^{(2)})^2\{ |\eta^{(3)}-\bar \eta^{(3)}|+ |Z^a_{-j}(\eta^{(1)}-\bar \eta^{(1)})|+ |Z_j^a(\theta-\bar\theta)|\}]^{1/2}\\
\leqslant C\|\eta^{(2)}-\bar \eta^{(2)}\| + \bar f{\mathrm{E}}[(Z^a_j-Z^a_{-j}\bar\eta^{(2)})^4]^{1/4}\{ {\mathrm{E}}[|\eta^{(3)}-\bar \eta^{(3)}|^2]^{1/4} + C\|\eta^{(1)}-\bar \eta^{(1)}\|+|\theta-\bar\theta| \}^{1/2}\\
\leqslant C'|\theta-\bar\theta|^{1/2}\vee \|\eta-\bar \eta\|_e^{1/2}
\end{array}$$
for some constance $C'<\infty$ since by Condition CI we have ${\mathrm{E}}[|Z^a\bar{\xi}|^2]^{1/2} \leqslant C\|\bar{\xi}\|$ for all vectors $\bar{\xi}$, and the conditions $\sup_{u\in\mathcal{U},j\in[p]}\|\gamma_u^j\|\leqslant C$, $\sup_{\theta \in \Theta_{uj}}|\theta|\leqslant C$, and $\sqrt{s \log(a_n)} \leqslant \delta_n \sqrt{n}$. This implies that $\|\eta^{(2)}-\bar\eta^{(2)}\|\leqslant \|\eta^{(2)}-\eta_{uj}^{(2)}\|+ \|\eta_{uj}^{(2)}-\bar\eta^{(2)}\| \leqslant 1$ so that $\|\eta^{(2)}-\bar\eta^{(2)}\|\leqslant \|\eta^{(2)}-\bar\eta^{(2)}\|^{1/2}$.
To verify Assumption (ref)(v)(b), let $t_{\bar r}= Z_j^a\theta_{\bar{r}}+Z^a_{-j}\eta_{\bar{r}}^{(1)}+\eta_{\bar r}^{(3)}$. We have
$$\begin{array}{rl}
\left. \partial_r{\mathrm{E}}(\psi_{uj}(W_{uj},\theta+r\delta_\theta,\eta+r\delta_\eta)) \right|_{r=\bar r} = \\ -{\mathrm{E}}[K_\varpi(W)f_{X_a\mid X_{-a},\varpi}(t_{\bar r})(Z_j^a - Z^a_{-j}\eta_{\bar{r}}^{(2)})\{Z_j^a\delta_\theta+Z_{-j}^a\delta_\eta^{(1)}+\delta_\eta^{(3)}\}] \\
-{\mathrm{E}}[K_\varpi(W)\{\tau-F_{X_a\mid X_{-a},\varpi}(t_{\bar r})\}Z_{-j}^a\delta_\eta^{(2)}] \end{array}$$
Applying Cauchy-Schwartz we have that
$$\begin{array}{lr}
\left|\left. \partial_r{\mathrm{E}}(\psi_{uj}(W_{uj},\theta+r\delta_\theta,\eta+r\delta_\eta)) \right|_{r=\bar r}\right|
\\
\leqslant \bar f {\mathrm{E}}[(Z_j^a - Z^a_{-j}\eta_{\bar{r}}^{(2)})^2]^{1/2}\{{\mathrm{E}}[ (Z_j^a)^2]^{1/2}|\delta_\theta|+{\mathrm{E}}[(Z_{-j}^a\delta_\eta^{(1)})^2]^{1/2}+{\mathrm{E}}[|\delta_\eta^{(3)}|^2]^{1/2}\} +\bar{f}{\mathrm{E}}[(Z_{-j}^a\delta_\eta^{(2)})^2]^{1/2}\\
\leqslant \bar B_{1n} (|\delta_\theta| \vee \|\eta-\eta_{uj}\|_e)\end{array}$$
where $\bar B_{1n} \leqslant C$ by the same arguments of bounded (second) moments of linear combinations.
Assumption (ref)(v)(c) follows similarly as
$$\begin{array}{rl}
\left. \partial_r^2{\mathrm{E}}(\psi_{uj}(W_{uj},\theta+r\delta_\theta,\eta+r\delta_\eta)) \right|_{r=\bar r} =\\
-{\mathrm{E}}[K_\varpi(W)f_{X_a\mid X_{-a},\varpi}'(t_{\bar r})(Z_j^a - Z^a_{-j}\eta_{\bar{r}}^{(2)})\{Z_j^a\delta_\theta+Z_{-j}^a\delta_\eta^{(1)}+\delta_\eta^{(3)}\}^2] \\
+2{\mathrm{E}}[K_\varpi(W)f_{X_a\mid X_{-a},\varpi}(t_{\bar r})(Z^a_{-j}\delta_\eta^{(2)})\{Z_j^a\delta_\theta+Z_{-j}^a\delta_\eta^{(1)}+\delta_\eta^{(3)}\}] \\
\end{array}$$
and under $|f'_{X_a\mid X_{-a},\varpi}|\leqslant \bar f'$, from Cauchy-Schwartz inequality we have
$$\begin{array}{ll}
\left|\left. \partial_r^2{\mathrm{E}}(\psi_{uj}(W_{uj},\theta+r\delta_\theta,\eta+r\delta_\eta)) \right|_{r=\bar r} \right| \\
\leqslant
|\bar f'_n{\mathrm{E}}[(Z_j^a - Z^a_{-j}\eta_{\bar{r}}^{(2)})^2]^{1/2}\{ {\mathrm{E}}[(Z_j^a)^4]|\delta_\theta|^2+{\mathrm{E}}[(Z_{-j}^a\delta_\eta^{(1)})^4]^{1/2}\}+C{\mathrm{E}}[\{\delta_\eta^{(3)}\}^2] \\
+2\bar f{\mathrm{E}}[(Z^a_{-j}\delta_\eta^{(2)})^2]^{1/2}\{{\mathrm{E}}[(Z_j^a)^2]^{1/2}|\delta_\theta|+{\mathrm{E}}[(Z_{-j}^a\delta_\eta^{(1)})^2]^{1/2}+{\mathrm{E}}[\{\delta_\eta^{(3)}\}^2]^{1/2}\} \\
\leqslant \bar B_{2n} (\delta_\theta^2\vee \|\eta-\eta_{uj}\|_e^2) \end{array}$$
where $\bar B_{2n} \leqslant C(1 + \bar f'_n)$ by the same arguments of bounded (fourth) moments as before and using that $|{\mathrm{E}}[(Z_j^a - Z^a_{-j}\eta_{\bar{r}}^{(2)})(\delta_\eta^{(3)})^2]|\leqslant \{ {\mathrm{E}}[(Z_j^a - Z^a_{-j}\eta_{\bar{r}}^{(2)})^2(\delta_\eta^{(3)})^2]\}^{1/2}{\mathrm{E}}[(\delta_\eta^{(3)})^2]^{1/2} \leqslant C{\mathrm{E}}[(\delta_\eta^{(3)})^2]$.
To verify the near orthogonality condition, note that for all $u\in\mathcal{U}$ and $j\in[p]$, since by definition $f_u = f_{X_a\mid X_{-a},\varpi}(Z^a\beta_u+r_u)$ we have
$$ |\mathrm{D}_{u,j,0}[\tilde \eta_{uj} - \eta_{uj}]| = |-{\mathrm{E}}[K_\varpi(W)f_u\{Z_{-j}^a(\tilde\eta^{(2)} - \eta_{uj}^{(2)})+ r_{u}\}v_{uj} ]| \leqslant \delta_n n^{-1/2} $$
by the relations ${\mathrm{E}}[K_\varpi(W)(\tau - F_{X_a\mid X_{-a},\varpi}(Z^a\beta_u+r_u))Z_{-j}^a]=0$ and ${\mathrm{E}}[K_\varpi(W)f_uZ_{-j}^av_{uj}]=0$ implied by the model, and $|{\mathrm{E}}[ K_\varpi(W) f_u r_{u}v_{uj} ]|\leqslant \delta_n n^{-1/2}$ by Condition CI. Thus, condition ((ref)) holds.
Furthermore, since $\Theta_{uj} \subset \theta_{uj} \pm C/\log n$, for $J_{uj}=\partial_\theta{\mathrm{E}}[\psi_{uj}(W_{uj},\theta_{uj},\eta_{uj})]={\mathrm{E}}[K_\varpi(W)f_uZ_j^av_{uj}]={\mathrm{E}}[K_\varpi(W)v_{uj}^2] ={\mathrm{E}}[v_{uj}^2| \varpi]{\mathrm{P}}(\varpi)$ as ${\mathrm{E}}[K_\varpi(W)f_uZ_{-j}^av_{uj}]=0$, we have that for all $\theta\in\Theta_{uj}$
$$ {\mathrm{E}}[\psi_{uj}(W_{uj},\theta,\eta_{uj})] = J_{uj}(\theta-\theta_{uj})+\frac{1}{2}\partial_\theta^2{\mathrm{E}}[\psi_{uj}(W_{uj},\bar\theta,\eta_{uj})](\theta-\theta_{uj})^2 $$ where $|\partial_\theta^2{\mathrm{E}}[\psi_{uj}(W_{uj},\bar\theta,\eta_{uj})]|\leqslant \bar f'{\mathrm{E}}[|Z_j^a|^2|v_{uj}|\mid \varpi] {\mathrm{P}}(\varpi) \leqslant \bar f'{\mathrm{E}}[|Z_j^a|^4| \varpi]^{1/2}{\mathrm{E}}[|v_{uj}|^2| \varpi]^{1/2} {\mathrm{P}}(\varpi) \leqslant C\bar f' {\mathrm{P}}(\varpi)$ so that for all $\theta\in\Theta_{uj}$ $$|{\mathrm{E}}[\psi_{uj}(W_{uj},\theta,\eta_{uj})]| \geqslant \{|{\mathrm{E}}[v_{uj}^2\mid \varpi]|- (C^2\bar f')/\log n\}{\mathrm{P}}(\varpi)|\theta-\theta_{uj}|$$
and we can take $j_n \geqslant c\inf_{\varpi\in\mathcal{W}}{\mathrm{P}}(\varpi) = c\mu_\mathcal{W}$.
Next we verify Assumption (ref) with ${\mathcal{H}}_{ujn} = \{ \eta =(\beta,\gamma,0) : \|\beta\|_0\leqslant Cs, \|\gamma\|_0\leqslant Cs, \|\beta-\beta_{u,-j}\| \leqslant C\tau_n, \|\gamma - \gamma_u^j\| \leqslant C\tau_n, \|\gamma-\gamma_u^j\|_1 \leqslant C\sqrt{s}\tau_n \}$.
We will show that $\widehat \eta_{uj} = (\widetilde\beta_{u,-j},\widetilde \gamma_u^j,0) \in {\mathcal{H}}_{ujn}$ with probability $1-o(1)$, uniformly over $u\in \mathcal{U}$ and $j\in [p]$.
Under Condition CI and the choice of penalty parameters, by Theorems (ref) and (ref), with probability $1-o(1)$, uniformly over $u\in\mathcal{U}$ we have
$$ \|\widetilde \beta_u - \beta_u\| \leqslant C\tau_n, \ \ \max_{j\in[p]}\sup_{u\in\mathcal{U}} \|\widetilde \gamma_{u}^j - \gamma_{u}^j\| \leqslant C\tau_n, \ \ \mbox{and} \ \ \max_{j\in[p]} \sup_{u\in\mathcal{U}} \|\widetilde \gamma_{u}^j\|_0 \leqslant \bar C s, $$
further by thresholding we can achieve $\sup_{u\in\mathcal{U}} \|\widetilde \beta_u\|_0 \leqslant \bar C s$ using Lemma (ref).
Next we establish the entropy bounds. For $\eta \in {\mathcal{H}}_{ujn}$ we have that $$\psi_{uj}(W_{uj},\theta,\eta)= K_\varpi(W)( \tau - 1\{ X_a \leqslant Z_j^a\theta+Z^a_{-j}\beta_{-j}\} )f_u\{Z_j^a - Z_{-j}^a\gamma\} $$
It follows that ${\mathcal{F}}_1\subset \mathcal{W}{\mathcal{G}}_1{\mathcal{G}}_2{\mathcal{G}}_3 \cup \bar{{\mathcal{F}}_0}$ where $\bar{{\mathcal{F}}_0}= \{ \psi_{uj}(W_{uj},\theta,\eta_{uj}) : u\in \mathcal{U}, j\in[p], \theta \in \Theta_{uj}\}$, ${\mathcal{G}}_1 =\{ \tau - 1\{ X_a\leqslant Z^a\beta\} : \|\beta\|_0 \leqslant Cs, \tau \in \mathcal{T}, a \in V\}$, ${\mathcal{G}}_2 =\{ Z^a \to Z^a(1,-\gamma), \|\gamma\|_0\leqslant Cs, \|\gamma\|\leqslant C, a\in V\}$, ${\mathcal{G}}_3 = \{ f_u : u \in \mathcal{U}\}$.
Under Condition CI, $\mathcal{W}$ is a VC class of sets with VC index $d_W$ (fixed). It follows that ${\mathcal{G}}_1$ and ${\mathcal{G}}_2$ are $p$ choose $O(s)$ VC-subgraph classes with VC indices at most
$O(s)$. Therefore, ${\rm ent}({\mathcal{G}}_1)\vee {\rm ent}({\mathcal{G}}_2)\vee {\rm ent}(\mathcal{W})\leqslant C s \log(a_n/\epsilon)+Cd_W\log(e/\epsilon)$ by Theorem 2.6.7 in vdV-W and by standard arguments. Also, since $f_u$ is Lipschitz in $u$ by Condition CI, we have ${\rm ent}({\mathcal{G}}_3)\leqslant (1+d_W)\log( a_nL_f/\epsilon)$. Moreover, an envelope $F_G$ for ${\mathcal{F}}_1$ satisfies $$\begin{array}{rl}
{\mathrm{E}}[F_G^q] & = {\mathrm{E}}[ \sup_{u\in\mathcal{U},j\in[p],\|\gamma-\gamma_u^j\|_1\leqslant C\sqrt{s}\tau_n} |v_{uj} - f_uZ^a_{-j}(\gamma-\gamma_u^j)|^q ]\\
& \leqslant 2^{q-1}{\mathrm{E}}[ \sup_{ u\in\mathcal{U}, j\in [p]} |v_{uj}|^q] + 2^{q-1}\bar f{\mathrm{E}}[\max_{a\in V}\|Z^a\|_\infty^q]\{C\sqrt{s}\tau_n\}^q \\
& \leqslant 2^{q-1}L_n^q + 2^{q-1}\bar f\{M_nC\sqrt{s}\tau_n\}^q \leqslant 2^{q}L_n^q \\
\end{array}$$ since $M_nC\sqrt{s}\tau_n \leqslant \delta_n L_n/\bar f$ and $\delta_n \leqslant 1$ for $n$ large.
Next we bound the entropy in $\bar{{\mathcal{F}}_0}$. Note that for any $\psi_{uj}(W_{uj},\theta,\eta_{uj})\in \bar{{\mathcal{F}}_0}$, there is some $\delta \in [-C,C]$ such that
$$\psi_{uj}(W_{uj},\theta,\eta_{uj})= K_\varpi(W)\{ \tau - 1\{ X_a \leqslant Z_j^a\delta+Q_{X_a}(\tau\mid X_{-a},\varpi)\} \}v_{uj}$$
and therefore $\bar{{\mathcal{F}}_0} \subset \mathcal{W}\{\mathcal{T}- \phi(\mathcal{V})\}\mathcal{L}$ where $\phi(t)=1\{t\leqslant 0\}$, $\mathcal{V}=\cup_{a\in V,j\in[p]}\mathcal{V}_{aj}$ with $$\mathcal{V}_{aj}:=\{ X_a - Z_j^a\delta - Q_{X_a}(\tau\mid X_{-a},\varpi): \tau\in \mathcal{T}, \varpi\in\mathcal{W}, |\delta| \leqslant C\},$$
and $\mathcal{L} = \cup_{a\in V, j\in[p]}(\mathcal{L}_{aj}+\{v_{\bar uj}\})$ where $\mathcal{L}_{aj} = \{ (X,W)\mapsto v_{uj} - v_{\bar uj} = f_uZ_{-j}^a(\gamma_u^j-\gamma_{\bar u}^j) : u\in\mathcal{U}\}$. Note that each $\mathcal{V}_{aj}$ is a VC subgraph class of functions with index $1+Cd_W$ as $\{Q_{X_a}(\tau| X_{-a},\varpi) : (\tau,\varpi)\in \mathcal{W}\times\mathcal{T}\}$ is a VC-subgraph with VC-dimension $Cd_W$ for every $a\in V$. Since $\phi$ is monotone, $\phi(\mathcal{V})$ is also the union of VC-dimension of order $1+Cd_W$.
Letting $F_1=1$ be an envelope for $\mathcal{W}$ and $\mathcal{T}-\phi(\mathcal{V})$. By Lemma (ref), it follows that $\|\gamma_u^j-\gamma_{u'}^j\|\leqslant L_\gamma \{\|u-u'\|+\|\varpi-\varpi'\|^{1/2}\}$ for some $L_\gamma$ satisfying $\log(L_\gamma) \leqslant C\log(p|V|n)$ under Condition CI. Therefore, $|v_{uj}-v_{\bar uj}| = |Z_{-j}^a(\gamma_u^j-\gamma_{\bar u}^j)|\leqslant \|Z^a\|_\infty \sqrt{p}\|\gamma_u^j-\gamma_{\bar u}^j\|$. For a choice of envelope $F_a = M_n^{-1}\|Z^a\|_\infty +2\sup_{u\in \mathcal{U}} |v_{uj}|$ which satisfies $\|F_a\|_{P,q} \lesssim L_n$, we have
$$\begin{array}{rl}
\log N(\epsilon \|F_a\|_{Q,2},\mathcal{L}_{aj},\|\cdot\|_{Q,2}) & \leqslant \log N(\frac{\epsilon}{M_n} \| \ \|Z^a\|_\infty\|_{Q,2},\mathcal{L}_{aj},\|\cdot\|_{Q,2}) \\
&\leqslant \log N(\epsilon /\{ M_n\sqrt{p}L_\gamma\},\mathcal{U},|\cdot|) \leqslant Cd_u\log(M_npL_\gamma/\epsilon)\end{array}$$
Since $\mathcal{L} = \cup_{a\in V,j\in[p]}(\mathcal{L}_{aj}+\{v_{\bar uj}\})$, taking $F_L = \max_{a\in V} F_a$, we have that
$$\begin{array}{rl}
\log N(\epsilon \|F_LF_1\|_{Q,2},\bar{\mathcal{F}}_0,\|\cdot\|_{Q,2}) & \leqslant \log N(\frac{\epsilon}{4} \|F_1\|_{Q,2},\mathcal{W},\|\cdot\|_{Q,2})+\log N(\frac{\epsilon}{4} \|F_1\|_{Q,2},\mathcal{T}-\phi(\mathcal{V}),\|\cdot\|_{Q,2})\\
& + \log \sum_{a\in V, j\in [p]}N(\frac{\epsilon}{2} \|F_a\|_{Q,2},\mathcal{L}_{aj},\|\cdot\|_{Q,2})\\
&\leqslant \log (p|V|) + 1+C'\{d_W+d_u\}\log(4eM_n|V|pL_\gamma/\epsilon)\\ \end{array}$$
where the last line follows from the previous bounds.
Next we verify the growth conditions in Assumption (ref) with the proposed ${\mathcal{F}}_1$ and $K_n \lesssim CL_n$. We take $s_{n(\mathcal{U},p)}= (1+ d_W)s$ and $a_n = \max\{ n,p,|V| \}$. Recall that $\bar B_{1n} \leqslant C$, $\bar B_{2n} \leqslant C$, $j_n\geqslant c \mu_\mathcal{W}$. Thus, we have $\sqrt{n}(\tau_n/j_n)^2 \lesssim \sqrt{n}\frac{s\log(p|V|n)}{n(1\wedge {\underline{f}}_\mathcal{U}^2)\mu_\mathcal{W}^3}\leqslant \delta_n$ under $s^2\log^2(p|V|n) \leqslant n(1\wedge {\underline{f}}_\mathcal{U}^4)\mu_\mathcal{W}^6$. Moreover, $(\tau_n/j_n)^{\alpha/2}\sqrt{s_{n(\mathcal{U},p)}\log(a_n)} \lesssim \sqrt[4]{\frac{(1+ d_W)^3s^3\log^3(p|V|n)}{n(1\wedge {\underline{f}}_\mathcal{U}^2)\mu_\mathcal{W}^3}}\lesssim \delta_n$ under $d_W$ fixed and $s^3\log^3(p|V|n) \leqslant \delta_n^4 n(1\wedge {\underline{f}}_\mathcal{U}^2)\mu_\mathcal{W}^3$ and $s_{n(\mathcal{U},p)}n^{-\frac{1}{2}}K_n\log(a_n)\log n \lesssim (1+ d_W)sn^{\frac{1}{q}-\frac{1}{2}}M_n\log(p|V|n)\log n \leqslant \delta_n$ under our conditions. Finally, the conditions of Corollary (ref) hold with $\rho_n = (1+d_W)$ since the score is the product of VC-subgraph classes of function with VC index bounded by $C(1+d_W)$.
\end{proof}
\begin{proof}[Proof of Theorem (ref)]
We will invoke Lemma (ref) with $\bar \beta_u$ as the estimand and $r_{iu}=X_{i,-a}'(\beta_u-\bar\beta_u)$, therefore ${\mathrm{E}}[ K_\varpi(W)(\tau - 1\{ X_a \leqslant X_{-a}'\bar\beta_u + r_{u} \})X_{-a}]=0$. To invoke the lemma we verify that the events $\Omega_1$, $\Omega_2$, $\Omega_3$ and $\Omega_4$ hold with probability $1-o(1)$
$$\begin{array}{rl}
\Omega_1 & :=\{ \lambda_u\geqslant c |S_{uj}|/\widehat\sigma_{a\varpi j}^{X}, \ \mbox{for all} \ \ u\in\mathcal{U}, j\in V\}, \\
\Omega_2 &:=\{\widehat R_u(\bar\beta_u) \leqslant \bar R_{u\xi} : u\in\mathcal{U}\}\\
\Omega_3 & :=\left\{ \sup_{u\in\mathcal{U}, 1/\sqrt{n} \leqslant \|\delta\|_{1,\varpi} \leqslant \sqrt{n}} |{\mathbb{E}_n}[g_u(\delta,X,W)-{\mathrm{E}}[g_u(\delta,X,W)\mid X_{-a},W]]|/\|\delta\|_{1,\varpi} \leqslant t_3 \right\}\\
\Omega_4 & :=\{ K_{u} \widehat\sigma_{a\varpi j}^{X} \geqslant |{\mathbb{E}_n}[h_{uj}(X_{-a},W)]|, \mbox{for all} \ u\in\mathcal{U}, j\in V\backslash\{a\} \}
\end{array}$$
where $g_u(\delta,X,W)= K_\varpi(W)\{\rho_\tau(X_a-X_{-a}'(\bar\beta_u+\delta))-\rho_\tau(X_a-X_{-a}'\bar\beta_u)\}$, $h_{uj}(X_{-a},W) ={\mathrm{E}}[ K_\varpi(W)\{\tau-F_{X_a\mid X_{-a},W }(X_{-a}'\bar\beta_u+r_u)\}X_{j}\mid X_{-a},W ]$.
By Lemma (ref) with $\xi=1/n$, by setting $\lambda_u=\lambda_0=c2(1+1/16) \sqrt{2\log(8|V|^2\{ne/d_W\}^{2d_W}n)/n}$, we have ${\mathrm{P}}(\Omega_1)=1-o(1)$. By Lemma (ref), setting $\bar R_{u\xi} = C s(1+d_W)\log(|V|n)/n$ we have ${\mathrm{P}}(\Omega_2)=1-o(1)$ for some $\xi = o(1)$.
By Lemma (ref) we have ${\mathrm{P}}(\Omega_3)=1-o(1)$ by setting $t_3 := C\sqrt{(1+d_W)\log\left(|V| n M_n/\xi \right)} $.
Finally, by Lemma (ref) with $K_u =C \sqrt{\frac{(1+d_W) \log(|V|n)}{n}}$ we have ${\mathrm{P}}(\Omega_4)=1-o(1)$.
Moreover, we have that $\|\bar \beta_u\|_{1,\varpi} \leqslant \sqrt{s}\|\bar \beta_u\|_{2,\varpi} \leqslant C\sqrt{s}= o(\sqrt{n})$ and $\frac{1}{\lambda_u(1-1/c)}\bar R_{u\xi} = o(\sqrt{n})$ for all $u\in\mathcal{U}$. Finally, we verify condition ((ref)) holds for all $$\delta\in A_u := \Delta_{u,2\mathbf{c}}\cup\{v:\|v\|_{1,\varpi}\leqslant 2\mathbf{c}\bar R_{u\xi}/\lambda_u, \|\sqrt{f_u}X_{-a}'v\|_{n,\varpi} \geqslant C\sqrt{s(1+d_W)\log(n|V|)/n}/\kappa_{u,2\mathbf{c}} \},$$ $\bar q_{A_u}/4 \geqslant (\sqrt{\bar f}+1) \|r_{u}\|_{n,\varpi} + \left[\lambda_u+t_3+K_u\right]\frac{3\mathbf{c}\sqrt{s}}{\kappa_{u,2\mathbf{c}}}$ and $\bar q_{A_u} \geqslant \{ 2\mathbf{c}\left(1+ \frac{t_3+K_u}{\lambda_u}\right)\bar R_{u\gamma} \}^{1/2}$.
Consider the matrices ${\mathbb{E}_n}[K_\varpi(W)f_uX_{-a}X_{-a}']$ and ${\mathrm{E}}[K_\varpi(W)f_uX_{-a}X_{-a}']$. By Lemma (ref), with probability $1-o(1)$, it follows that we can take $\eta = \eta_n = C M_n\sqrt{s(1+d_W)\log(|V|n)}\log(1+s)\{\log n\}/\sqrt{n}$ and $D_{kk} = 2\eta$ in Lemma (ref). (Note that we increase $\delta_n$ by a factor of $\sqrt{\log n}$.) Therefore, with at least the same probability we have (taking $s\geqslant 2$)
\begin{equation} \delta'{\mathbb{E}_n}[K_\varpi(W)f_uX_{-a}X_{-a}']\delta \geqslant \delta'{\mathrm{E}}[K_\varpi(W)f_uX_{-a}X_{-a}']\delta - 4\eta\|\delta\|_1^2/s\end{equation}
and by definition of ${\underline{f}}_u$ we have
$$ {\mathbb{E}_n}[K_\varpi(W)f_u|X_{-a}'\delta|^2] \geqslant {\underline{f}}_u {\mathrm{E}}[K_\varpi(W)|X_{-a}'\delta|^2] - 4\eta\|\delta\|_1^2/s \geqslant c{\underline{f}}_u {\mathrm{P}}(\varpi)\|\delta\|^2 - 4\eta\|\delta\|_1^2/s .$$
For $\delta \in \Delta_{u,2\mathbf{c}}$ we have $\|\delta\|_1 \leqslant C\|\delta\|_{1,\varpi}/\{{\mathrm{P}}(\varpi)\}^{1/2} \leqslant C' \|\delta_{T_u}\|_{1,\varpi}/\{{\mathrm{P}}(\varpi)\}^{1/2} \leqslant C' \sqrt{s} \|\delta_{T_u}\|_{2}$. Note that we can assume $\|\delta\| \geqslant c\sqrt{s(1+d_W)\log(n|V|)/n}$ otherwise we are done. So that for $\delta \in A_u \backslash \Delta_{u,2\mathbf{c}}$ we have that $\|\delta\|_1 / \|\delta\|_2 \leqslant C s\sqrt{\log(|V|n)/n} / \sqrt{s(1+d_W)\log(n|V|)/n} \leqslant C'\sqrt{s}$.
Similarly we have
\begin{equation} {\mathbb{E}_n}[K_\varpi(W)|X_{-a}'\delta|^2] \leqslant {\mathrm{E}}[K_\varpi(W)|X_{-a}'\delta|^2] + 4\eta\|\delta\|_1^2/s \leqslant C {\mathrm{P}}(\varpi)\|\delta\|^2 - 4\eta\|\delta\|_1^2/s.\end{equation}
Under the condition that $\eta = o({\underline{f}}_\mathcal{U} \mu_\mathcal{W})$, which holds by Condition P, for $n$ sufficiently large we have with probability $1-o(1)$ that
\begin{equation} \begin{array}{rl}
\bar q_{A_u} & \geqslant \frac{c}{\bar f'} \inf_{\delta\in A_u} \frac{{\mathbb{E}_n}[K_\varpi(W)f_u|X_{-a}'\delta|^2]^{3/2}}{{\mathbb{E}_n}[K_\varpi(W)|X_{-a}'\delta|^3]} \geqslant \frac{c}{\bar f'} \inf_{\delta\in A_u} \frac{{\mathbb{E}_n}[K_\varpi(W)f_u|X_{-a}'\delta|^2]^{3/2}}{{\mathbb{E}_n}[K_\varpi(W)|X_{-a}'\delta|^2]\max_{i\leqslant n}\|X_i\|_\infty\|\delta\|_1} \\
& \geqslant \frac{c}{\bar f'} \inf_{\delta\in A_u} \frac{ \{ c'{\underline{f}}_u {\mathrm{P}}(\varpi) \|\delta\|^2 \}^{3/2}}{C'{\mathrm{P}}(\varpi)\|\delta\|^2\max_{i\leqslant n}\|X_i\|_\infty\|\delta\|_1} \geqslant \frac{c}{\bar f'} \inf_{\delta\in A_u} \frac{ c'{\underline{f}}_u^{3/2} {\mathrm{P}}(\varpi)^{1/2} \|\delta\|}{C'\max_{i\leqslant n}\|X_i\|_\infty\|\delta\|_1} \\
& \geqslant C” \frac{{\underline{f}}_\mathcal{U}^{3/2}}{\bar f'} \frac{\mu_\mathcal{W}^{1/2}}{\sqrt{s}\max_{i\leqslant n}\|X_i\|_\infty}\\
\end{array}\end{equation}
where $\max_{i\leqslant n}\|X_i\|_\infty \leqslant \ell_n M_n$ with probability $1-o(1)$ for any $\ell_n \to \infty$.
Therefore, under the condition $M_n s\sqrt{\log(p|V|n)} = o( \sqrt{n \mu_\mathcal{W}})$ assumed in Condition P, the conditions on $\bar q_{A_u}$ are satisfied.
By Lemma (ref), we have uniformly over all $u=(a,\tau,\varpi)\in \mathcal{U} := V\times \mathcal{T}\times \mathcal{W}$
{ $$
\|\sqrt{f_u}X_{-a}'(\widehat\beta_u - \beta_u)\|_{n,\varpi} \leqslant C \sqrt{\frac{(1+d_W)\log(n|V|)}{n} }\frac{\sqrt{s}}{\kappa_{u,2\mathbf{c}}} \ \ \mbox{and} \ \ \|\widehat\beta_u - \beta_u\|_{1,\varpi} \leqslant C \sqrt{\frac{(1+d_W)\log(n|V|)}{n} }\frac{s}{\kappa_{u,2\mathbf{c}}} $$}
where $\kappa_{u,2\mathbf{c}}$ is bounded away from zero with probability $1-o(1)$ for $n$ sufficiently large.
Consider the thresholded estimators $\widehat\beta_u^{\bar \lambda}$ for $\bar \lambda =\{ (1+d_W)\log(n|V|)/n\}^{1/2}$. By Lemma (ref) we have $\|\widehat\beta_u^{\bar \lambda}\|_0 \leqslant C s$ and the same rates of convergence as $\widehat\beta_u$. Therefore, by refitting over the support of $\widehat\beta_u^{\bar \lambda}$ we have by Lemma (ref), the estimator $\widetilde \beta_u$ has the same rate of convergence where we used that $\widehat Q_u \leqslant \lambda_u\|\widehat\beta_u^{\bar \lambda}- \beta_u\|_{1,\varpi} \lesssim C s(1+d_W)\log(|V|n)/n$ (the other conditions of Lemma (ref) hold as for the conditions in Lemma (ref)).
Next we will invoke Lemma (ref) for the new penalty choice and penalty loadings. (We note that minor modifications cover the new penalty loadings.)
$$\begin{array}{rl}
\Omega_1 & :=\{ \lambda_u\geqslant c |S_{uj}|/\{{\mathbb{E}_n}[K_w(W)\varepsilon_u^2X_{-a,j}^2]\}^{1/2}, \ \mbox{for all} \ \ u\in\mathcal{U}, j\in V\}, \\
\Omega_2 &:=\{\widehat R_u(\bar\beta_u) \leqslant \bar R_{u\xi} : u\in\mathcal{U}\}\\
\Omega_3 & :=\left\{ \sup_{u\in\mathcal{U}, 1/\sqrt{n} \leqslant \|\delta\|_{1,\varpi} \leqslant \sqrt{n}} |{\mathbb{E}_n}[g_u(\delta,X,W)-{\mathrm{E}}[g_u(\delta,X,W)| X_{-a},W]]|/\{\theta_u\|\delta\|_{1,\varpi}\} \leqslant t_3 \right\}\\
\Omega_4 & :=\{ K_{u}\theta_u \widehat\sigma_{a\varpi j}^{X} \geqslant |{\mathbb{E}_n}[h_{uj}(X_{-a},W)]|, \mbox{for all} \ u\in\mathcal{U}, j\in V\backslash\{a\} \}\\
\Omega_5 & :=\{ \theta_u \geqslant \max_{j\in V}\widehat\sigma_{a\varpi j}^{X} / \{{\mathbb{E}_n}[K_\varpi(W)\varepsilon_u^2X_j^2]\}^{1/2} \}
\end{array}$$
where event $\Omega_5$ simply makes the relevant norms equivalent, $\|\cdot\|_{1,u}\leqslant \|\cdot\|_{1,\varpi}\leqslant \theta_u\|\cdot\|_{1,u}$. Note that we can always take $\theta_u \leqslant 1/\{\tau(1-\tau)\}\leqslant C$ since $\mathcal{T}$ is a fixed compact set.
Next we show that the bootstrap approximation of the score provides a valid choice of penalty parameter. Let $\widehat \varepsilon_u := 1\{ X_a \leqslant X_{-a}'\widetilde\beta_u\}-\tau$. For notational convenience for $u\in \mathcal{U}$, $j\in V\backslash\{a\}$ define $$\begin{array}{rl}
\displaystyle \bar\psi_{iuj}=\frac{K_\varpi(W_i)\varepsilon_{iu} X_{ij}}{{\mathrm{E}}[K_\varpi(W)\varepsilon_u^2X_j^2]^{1/2}}, \ \
\psi_{iuj}=\frac{K_\varpi(W_i)\varepsilon_{iu} X_{ij}}{{\mathbb{E}_n}[K_\varpi(W)\varepsilon_{u}^2X_j^2]^{1/2}}, \ \ \displaystyle \widehat \psi_{iuj} =\frac{K_\varpi(W_i)\widehat\varepsilon_{iu} X_{ij}}{{\mathbb{E}_n}[K_\varpi(W)\widehat\varepsilon_{u}^2X_j^2]^{1/2}}.
\end{array}$$
We will consider the following processes:
$$\begin{array}{rl}
\bar S_{uj} = \frac{1}{\sqrt{n}}\sum_{i=1}^n \bar \psi_{iuj}, \ \ S_{uj} = \frac{1}{\sqrt{n}}\sum_{i=1}^n \psi_{iuj}, \ \ \overline{\mathcal{G}}_{uj}= \frac{1}{\sqrt{n}}\sum_{i=1}^ng_i \bar \psi_{iuj}, \ \
\displaystyle \widehat{\mathcal{G}}_{uj} = \frac{1}{\sqrt{n}}\sum_{i=1}^ng_i \widehat \psi_{iuj}, \ \ \end{array}
$$ and $\mathcal{N}$ is a tight zero-mean Gaussian process with covariance operator given by ${\mathrm{E}}[\bar\psi_{uj}\bar\psi_{u'j'}]$. Their supremum are denoted by $\bar Z_{S}:=\sup_{u\in \mathcal{U}, j\in V\backslash\{a\}} | \bar S_{uj}|$, $Z_S:=\sup_{u\in \mathcal{U}, j\in V\backslash\{a\}} | S_{uj}|$, $\bar Z_G^*:=\sup_{u\in \mathcal{U}, j\in V\backslash\{a\}} | \overline{\mathcal{G}}_{uj}|$, $\widehat Z_G^*:=\sup_{u\in \mathcal{U}, j\in V\backslash\{a\}} | \widehat{\mathcal{G}}_{uj}|$, and $Z_N:=\sup_{u\in \mathcal{U}, j\in V\backslash\{a\}} | \mathcal{N}_{uj}|$.
The penalty choice should majorate $Z_S$ and we simulate via $\widehat Z_G^*$. We have that
$$ \begin{array}{rl}
| {\mathrm{P}}( Z_S \leqslant t ) - {\mathrm{P}}( \widehat Z_G^* \leqslant t )| & \leqslant | {\mathrm{P}}( Z_S \leqslant t ) - {\mathrm{P}}( \bar Z_S \leqslant t )|+| {\mathrm{P}}( \bar Z_S \leqslant t ) - {\mathrm{P}}( Z_N \leqslant t )|\\
& +| {\mathrm{P}}( Z_N \leqslant t ) - {\mathrm{P}}( \bar Z_G^* \leqslant t )|+| {\mathrm{P}}( \bar Z_G^* \leqslant t ) - {\mathrm{P}}( \widehat Z_G^* \leqslant t )|\\
\end{array}
$$
We proceed to bound each term. We have that
$$ \begin{array}{rl}
| Z_S - \bar Z_S| &\displaystyle \leqslant \bar Z_S \sup_{u\in\mathcal{U},j\in V\backslash\{a\}}\left| \frac{{\mathrm{E}}[K_\varpi(W)\varepsilon_{u}^2X_j^2]^{1/2}}{{\mathbb{E}_n}[K_\varpi(W)\varepsilon_{u}^2X_j^2]^{1/2}} - 1 \right| \\
& \displaystyle \leqslant \bar Z_S \sup_{u\in\mathcal{U},j\in V\backslash\{a\}}\left| \frac{({\mathbb{E}_n}-{\mathrm{E}})[K_\varpi(W)\varepsilon_{u}^2X_j^2]}{{\mathbb{E}_n}[K_\varpi(W)\varepsilon_{u}^2X_j^2]^{1/2}\{{\mathbb{E}_n}[K_\varpi(W)\varepsilon_{u}^2X_j^2]^{1/2}+{\mathrm{E}}[K_\varpi(W)\varepsilon_{u}^2X_j^2]^{1/2}\}} \right|\end{array} $$
Therefore, since $\{ 1\{ X_{a} \leqslant X_{-a}'\beta_u\} : u\in \mathcal{U}\}$ is a VC-subgraph of VC dimension $1+d_W$, and $\mathcal{W}$ is a VC class of sets of dimension $d_W$, we apply Lemma (ref) with envelope $F=\|X\|_\infty^2$ and $\sigma^2 \leqslant \max_{j\in V}{\mathrm{E}}[ X_j^4] \leqslant C$ to obtain with probability $1-o(1)$
$$ \sup_{u\in \mathcal{U}, j \in V\backslash\{a\} } |({\mathbb{E}_n}-{\mathrm{E}})[K_\varpi(W)\varepsilon_{u}^2X_j^2]| \lesssim \delta_{1n}':= \sqrt{\frac{(1+d_W)\log(|V|n)}{n}} + \frac{M_n^2(1+d_W)\log(|V|n)}{n} $$
where $\delta_{1n}=o(\mu_\mathcal{W}^2)$ under Condition P. Note that this implies that the denominator above is bounded away from zero by $c\mu_\mathcal{W}$. Therefore,
$$ | Z_S - \bar Z_S| \lesssim_P \delta_{1n} := \bar Z_S \delta_{1n}'/\mu_\mathcal{W}.$$
where $\bar Z_S \lesssim_P \{(1+d_W)\log(n|V|)\}^{1/2}$.
By Theorem 2.1 in chernozhukov2015noncenteredprocesses, since ${\mathrm{E}}[ \bar \psi_{uj}^4 ] \leqslant C$, there is a version of $Z_N$ such that
$$ | \bar Z_S - Z_N| \lesssim_P \delta_{2n}:=\left( \frac{M_n(1+d_W)\log(n|V|)}{n^{1/2}}+ \frac{M_n^{1/3}((1+d_W)\log(n|V|))^{2/3}}{n^{1/6}}\right) $$
and by Theorem 2.2 in chernozhukov2015noncenteredprocesses, there is also a version of
$$ | Z_N - \bar Z_G^*| \lesssim_P \left( \frac{M_n(1+d_W)\log(n|V|)}{n^{1/2}}+ \frac{M^{1/2}_n((1+d_W)\log(n|V|))^{3/4}}{n^{1/4}}\right) $$
Finally, we have that
$$ |\bar Z_G^* -\widehat Z_G^*|\leqslant \sup_{u\in\mathcal{U},j} \left| \frac{1}{\sqrt{n}}\sum_{i=1}^n g_i (\widehat \psi_{iuj} -\bar \psi_{iuj} ) \right| $$
where conditional on $(X_i,W_i), i=1,\ldots,n$, $\frac{1}{\sqrt{n}}\sum_{i=1}^n g_i (\widehat \psi_{iuj} -\bar \psi_{iuj} ) $ is a zero-mean Gaussian with variance ${\mathbb{E}_n}[(\widehat \psi_{iuj} -\bar \psi_{iuj} )^2] \leqslant \bar \delta_n^2$. Next we bound $\bar\delta_n$. We have
$$ \begin{array}{rl}
\bar \delta_n & \leqslant {\mathbb{E}_n}[(\widehat \psi_{uj} - \psi_{uj} )^2]^{1/2} + {\mathbb{E}_n}[( \psi_{uj} - \bar \psi_{uj} )^2]^{1/2} \leqslant {\mathbb{E}_n}[(\widehat \psi_{uj} - \psi_{uj} )^2]^{1/2} + \delta_{1n}/\mu_{\mathcal{W}},\\
\end{array}
$$
and
$$\begin{array}{rl}
{\mathbb{E}_n}[(\widehat \psi_{uj} - \psi_{uj} )^2]^{1/2} & \leqslant \frac{{\mathbb{E}_n}[(K_\varpi(W)X_{ij}|\widehat\varepsilon_u - \varepsilon_u|)^2]^{1/2}}{{\mathbb{E}_n}[K_\varpi(W)X_{ij}^2\widehat\varepsilon_u^2]^{1/2}} \\
& + \frac{{\mathbb{E}_n}[K_\varpi(W)X_{ij}^2\varepsilon_u^2]^{1/2}}{c{\mathrm{P}}(\varpi)} |{\mathbb{E}_n}[K_\varpi(W)X_{ij}^2\widehat\varepsilon_u^2]^{1/2}-{\mathbb{E}_n}[K_\varpi(W)X_{ij}^2\varepsilon_u^2]^{1/2}|\\
& \leqslant {\mathbb{E}_n}[K_\varpi(W)X_{ij}^2|\widehat\varepsilon_u - \varepsilon_u|^2]^{1/2} \left\{\frac{1}{{\mathbb{E}_n}[K_\varpi(W)X_{ij}^2\widehat\varepsilon_u^2]^{1/2}} + \frac{{\mathbb{E}_n}[K_\varpi(W)X_{ij}^2\varepsilon_u^2]^{1/2}}{c{\mathrm{P}}(\varpi)} \right\},\\
\end{array}$$
note that the term in the curly brackets is bounded by $C/{\mathrm{P}}(\varpi)^{1/2}$ with probability $1-o(1)$. To bound the other term note that $|\widehat\varepsilon_u - \varepsilon_u|^2 = | 1\{ X_a \leqslant X_{-a}'\widetilde\beta_u\} - 1\{ X_a \leqslant X_{-a}'\beta_u\}|$. Note that $\widehat \varepsilon_u = 1\{ X_a \leqslant X_{-a}'\widetilde\beta_u\} - \tau$ where $\|\widetilde\beta_u\|_0 \leqslant Cs$. Therefore, we have $\{ 1\{ X_a \leqslant X_{-a}'\widetilde\beta_u\} : u\in \mathcal{U}\} \subset \{ 1\{ X_a \leqslant X_{-a}'\beta\} : \|\beta\|_0\leqslant Cs\}$ which is the union of $\binom{|V|}{Cs}$ VC subgraph classes of functions with VC dimension $C's$. Moreover, we have
$$\begin{array}{rl}{\mathrm{E}}[K_\varpi(W)X_{ij}^2|\widehat\varepsilon_u - \varepsilon_u|^2] & = {\mathrm{E}}[K_\varpi(W)X_{ij}^2| 1\{ X_a \leqslant X_{-a}'\widetilde\beta_u\} - 1\{ X_a \leqslant X_{-a}'\beta_u\}|] \\
& \leqslant \bar f {\mathrm{E}}[K_\varpi(W)X_{ij}^2| X_{-a}'(\widetilde\beta_u -\beta_u)|] \\
& \leqslant \bar f {\mathrm{E}}[K_\varpi(W)X_{ij}^4]^{1/2}{\mathrm{E}}[K_\varpi(W)| X_{-a}'(\widetilde\beta_u-\beta_u)|^2]^{1/2} \\
& \leqslant C(\bar f/{\underline{f}}_\mathcal{U}^{1/2}) {\mathrm{P}}(\varpi)^{1/2} \sqrt{s(1+d_W)\log(n|V|)/n}\\
\end{array}$$
Therefore, by Lemma (ref), with probability $1-o(1)$ we have
$$ \left|({\mathbb{E}_n}-{\mathrm{E}})[K_\varpi(W)X_{ij}^2|\widehat\varepsilon_u - \varepsilon_u|^2]\right| \lesssim \sqrt{\frac{s(1+d_W)\log(n|V|)}{n}} C(\bar f/{\underline{f}}_\mathcal{U}^{1/2}) \sqrt{s(1+d_W)\log(n|V|)/n} $$
Under $\sqrt{s(1+d_W)\log(n|V|)/n} = o( {\underline{f}}_{\mathcal{U}}\mu_{\mathcal{W}})$ we have that with probability $1-o(1)$ that
$$ \bar\delta_n \leqslant C\{s(1+d_W)\log(n|V|)/n\}^{1/4}.$$
Therefore, using again the sparsity of $\widetilde \beta_u$ in the definition of $\widehat \psi_{iuj}$
$$\begin{array}{rl}
\displaystyle \sup_{u\in \mathcal{U}, j\in V\backslash\{a\}}\left| \frac{1}{\sqrt{n}}\sum_{i=1}^ng_i (\widehat \psi_{iuj} -\bar \psi_{iuj} ) \right| & \lesssim_P \bar \delta_n\sqrt{s(1+d_W)\log(|V|n)} \\
& \lesssim_P \delta_{3n}:=\{ s\log(|V|n)/n\}^{1/4}\sqrt{s(1+d_W)\log(|V|n)} \end{array}$$
The rest of the proof follows similarly to Corollary 2.2 in BCCW-ManyProcesses since under Condition P (and the bounds above) we have that $r_n:=\delta_{1n}+\delta_{2n}+\delta_{3n}=o( \{{\mathrm{E}}[Z_N]\}^{-1} ))$ where ${\mathrm{E}}[Z_N]\lesssim \{(1+d_W)\log(|V|n)\}^{1/2}$. Then we have $\sup_t | {\mathrm{P}}( Z_S \leqslant t ) - {\mathrm{P}}( \widehat Z_G^* \leqslant t )| = o_P(1)$ which in turn implies that
$$\begin{array}{rl}
{\mathrm{P}}( \Omega_1 ) & = {\mathrm{P}}( Z_S \leqslant \widehat c_G^*(\xi) ) \\
& \geqslant {\mathrm{P}}( \widehat Z_G^* \leqslant \widehat c_G^*(\xi) ) - | {\mathrm{P}}( Z_S \leqslant \widehat c_G^*(\xi) ) - {\mathrm{P}}( \widehat Z_G^* \leqslant \widehat c_G^*(\xi) )| \\
& \geqslant 1-\xi + o_P(1)\end{array}$$
Note that the occurrence of the events $\Omega_2$, $\Omega_3$ and $\Omega_4$ follows by similar arguments. The result follows by Lemma (ref), thresholding and applying Lemma (ref) and Lemma (ref) similarly to before.
\end{proof}
\section{Technical Lemmas for Conditional Independence Quantile Graphical Model}
Let $u=(a,\tau,\varpi)\in \mathcal{U}:= V\times \mathcal{T}\times \mathcal{W}$, and $T_u = {\rm support}(\beta_u)$ where $|T_u|\leqslant s$ for all $u\in\mathcal{U}$.
Define the pseudo-norms
$$ \|v\|_{n,\varpi}^2 := \frac{1}{n}\sum_{i=1}^n K_\varpi(W_i)(v_i)^2, \ \ \ \|\delta\|_{2,\varpi}:= \left\{\sum_{j=1}^p \{\widehat\sigma_{a\varpi j}^{Z}\}^2 |\delta_j|^2\right\}^{1/2}, \ \ \ \mbox{and} \ \ \ \|\delta\|_{1,\varpi}:= \sum_{j=1}^p \widehat\sigma_{a\varpi j}^Z|\delta_j|, $$ where $\widehat\sigma_{a\varpi j}^{Z} =\{{\mathbb{E}_n}[\{K_\varpi(W)Z_j^a\}^2]\}^{1/2}$. These pseudo-norms induce the following restricted eigenvalue as
$$ \kappa_{u,\mathbf{c}} = \min_{\|\delta_{T^c_u}\|_{1,\varpi} \leqslant \mathbf{c}\|\delta_{T_u}\|_{1,\varpi} }\frac{\|\sqrt{f_{u}}Z^a\delta\|_{n,\varpi}}{\|\delta\|_{1,\varpi}/\sqrt{s}}.$$
The restricted eigenvalue $\kappa_{u,\mathbf{c}}$ is an counterpart of the restricted eigenvalue proposed in BickelRitovTsybakov2009 for our setting. We note that $ \kappa_{u,\mathbf{c}}$ typically will vary with the events $\varpi\in\mathcal{W}$.
We will consider three key events in our analysis. Let
\begin{equation}\Omega_1:=\{ \lambda_u\geqslant c |S_{uj}|/\widehat\sigma_{a\varpi j}^Z, \ \mbox{for all} \ \ u\in\mathcal{U}, j\in [p]\}\end{equation} which occurs with probability at least $1-\xi$ by the choice of $\lambda_u$. For CIQGMs, we have $S_{uj}:={\mathbb{E}_n}[K_\varpi(W)(\tau-1\{X_a\leqslant Z^a\beta_u+r_{u}\}) Z^a_j]$, and $\lambda_u=\lambda_{V\mathcal{T}\mathcal{W}}\sqrt{\tau(1-\tau)}$. (In the case of PQGMs, we have $S_{uj}:= {\mathbb{E}_n}[K_\varpi(W)(\tau-1\{X_a\leqslant X_{-a}'\beta_u\}) X_{-a}]$, $\widehat\sigma_{a\varpi j}^X=\{{\mathbb{E}_n}[K_\varpi(W)X_{-a,j}^2]\}^{1/2}$ and $\lambda_u = \lambda_0$.)
To define the next event, for each $u\in \mathcal{U}$, consider the function defined as $$\widehat R_u(\beta_u) = {\mathbb{E}_n}[K_\varpi(W)\{\rho_u(X_a-Z^a\beta)-\rho_u(X_a-Z^a\beta_u-r_{u})-(\tau-1\{X_a\leqslant Z^a\beta_u+r_{u}\})(Z^a\beta-Z^a\beta_u-r_{u})\}]$$ in the case of CIQGMs. (In the case of PQGMs, we replace $Z^a$ with $X_{-a}$.) By convexity we have $\widehat R_u(\beta_u) \geqslant 0$. The event \begin{equation}\Omega_2 :=\{\widehat R_u(\beta_u) \leqslant \bar R_{u\xi} : u\in\mathcal{U}\}\end{equation} where $\bar R_{u\xi}$ are chosen so that $\Omega_2$ occurs with probability at least $1-\xi$. Note that by Lemma (ref), we have ${\mathbb{E}_n}[{\mathrm{E}}[\widehat R_u(\beta_u)\vert X_{-a},W]] \leqslant \bar f \|r_{u}\|_{n,\varpi}^2/2$ and with probability at least $1-\xi$, $\widehat R_u(\beta_u) \leqslant \bar R_{u\xi}:= 4\max\{\bar f \|r_{u}\|_{n,\varpi}^2, \ \|r_{u}\|_{n,\varpi}C\sqrt{\log(n^{1+d_W}p/\xi)/n}\} \leqslant C' s \log(n^{1+d_W}p/\xi) / n$.
Define $g_u(\delta,X,W)= K_\varpi(W)\{\rho_\tau(X_a-Z^a(\beta_u+\delta))-\rho_\tau(X_a-Z^a\beta_u)\}$ so that event $\Omega_3$ is defined as
\begin{equation} \Omega_3 := \left\{ \sup_{u\in\mathcal{U}, 1/\sqrt{n} \leqslant \|\delta\|_{1,\varpi} \leqslant \sqrt{n}} \frac{|{\mathbb{E}_n}[g_u(\delta,X,W)-{\mathrm{E}}[g_u(\delta,X,W)\vert X_{-a},W]]|}{\|\delta\|_{1,\varpi}} \leqslant t_3 \right\}\end{equation}
where $t_3$ is given in Lemma (ref) so that $\Omega_3$ holds with probability at least $1-\xi$.
\begin{lemma}
Suppose that $\Omega_1$, $\Omega_2$ and $\Omega_3$ holds. Further assume $2\frac{1+1/c}{1-1/c}\|\beta_u\|_{1,\varpi} + \frac{1}{\lambda_u(1-1/c)}\bar R_{u\xi} \leqslant \sqrt{n}$ for all $u\in\mathcal{U}$, and ((ref)) holds for all $\delta\in A_u := \Delta_{u,2\mathbf{c}}\cup\{v:\|v\|_{1,\varpi}\leqslant 2\mathbf{c}\bar R_{u\xi}/\lambda_u\}$, $\bar q_{A_u}/4 \geqslant (\sqrt{\bar f}+1) \|r_{u}\|_{n,\varpi} + \left[\lambda_u+t_3\right]\frac{3\mathbf{c}\sqrt{s}}{\kappa_{u,2\mathbf{c}}}$ and $\bar q_{A_u} \geqslant \{ 2\mathbf{c}\left(1+ \frac{t_3}{\lambda_u}\right)\bar R_{u\xi} \}^{1/2}$.
Then uniformly over all $u\in \mathcal{U}$ we have
{ $$\begin{array}{rl}
\|\sqrt{f_u}Z^a(\widehat\beta_u - \beta_u)\|_{n,\varpi} & \leqslant \sqrt{8\mathbf{c}\left(1+ \frac{t_3}{\lambda_u}\right)\bar R_{u\xi}} + (\bar f^{1/2}+1) \|r_{u}\|_{n,\varpi}+ \frac{3\mathbf{c}\lambda_u\sqrt{s}}{\kappa_{u,2\mathbf{c}}}+t_3\frac{(1+\mathbf{c})\sqrt{s}}{\kappa_{u,2\mathbf{c}}}\\
\|\widehat\beta_u - \beta_u\|_{1,\varpi} & \leqslant (1+2\mathbf{c})\sqrt{s}\|\sqrt{f_u}Z^a\delta_u\|_{n,\varpi}/\kappa_{u,2\mathbf{c}} + \frac{2\mathbf{c}}{\lambda_u}\bar R_{u\xi} \\
\end{array}$$}\end{lemma}
\begin{proof}[Proof of Lemma (ref)]
Let $u=(a,\tau,\varpi)\in \mathcal{U}$ and $\delta_u = \widehat\beta_u - \beta_u$.
By convexity and definition of $\widehat \beta_u$ we have
\begin{equation}\begin{array}{rl}
& \widehat R_u(\widehat\beta_u)-\widehat R_u(\beta_u)+S_u'\delta_u \\
& = {\mathbb{E}_n}[K_\varpi(W)\rho_u(X_a-Z^a\widehat\beta_u)] - {\mathbb{E}_n}[K_\varpi(W)\rho_u(X_a-Z^a\beta_u)] \\
& \leqslant \lambda_u\|\beta_u\|_{1,\varpi}-\lambda_u\|\widehat\beta_u\|_{1,\varpi}\end{array}\end{equation}
where $S_u$ is defined as in ((ref)) so that under $\Omega_1$ we have $\lambda_u\geqslant c|S_{uj}|/\widehat\sigma_{a\varpi j}^{Z}$.
Under $\Omega_1 \cap \Omega_2$, and since $\widehat R_u(\widehat{\beta}_u)\geqslant 0$, we have
\begin{equation}
\begin{array}{rl}
-\widehat R_u(\beta_u)-\frac{\lambda_u}{c}\|\delta_u\|_{1,\varpi} & \leqslant \widehat R_u(\beta_u+\delta_u)-\widehat R_u(\beta_u)+{\mathbb{E}_n}[K_\varpi(W)(\tau-1\{X_a\leqslant Z^a\beta_u+r_{u}\})Z^a\delta_u] \\
& = {\mathbb{E}_n}[K_\varpi(W)\rho_u(X_a-Z^a(\delta_u +\beta_u))] - {\mathbb{E}_n}[K_\varpi(W)\rho_u(X_a-Z^a\beta_u)] \\
& \leqslant \lambda_u\|\beta_u\|_{1,\varpi}-\lambda_u\|\delta_u +\beta_u\|_{1,\varpi}\end{array}\end{equation}
so that for $\mathbf{c} = (c+1)/(c-1)$
$$ \|\delta_{T^c_u}\|_{1,\varpi} \leqslant \mathbf{c} \|\delta_{T_u}\|_{1,\varpi} + \frac{c}{\lambda_u(c-1)}\widehat R_u(\beta_u).$$
To establish that $\delta_u \in A_u:=\Delta_{u,2\mathbf{c}}\cup \{ v : \|v\|_{1,\varpi} \leqslant 2\mathbf{c} \bar R_{u\xi}/\lambda_u\}$ we consider two cases. If $\|\delta_{T^c_u}\|_{1,\varpi} \geqslant 2\mathbf{c}\| \delta_{T_u}\|_{1,\varpi}$ we have
$$ \frac{1}{2}\|\delta_{T^c_u}\|_{1,\varpi} \leqslant \frac{c}{\lambda_u(c-1)}\widehat R_u(\beta_u)$$
and consequentially $$ \|\delta_u\|_{1,\varpi}\leqslant \{1+1/(2c)\}\|\delta_{T^c_u}\|_{1,\varpi} \leqslant \frac{2\mathbf{c}}{\lambda_u
}\widehat R_u(\beta_u).$$
Otherwise, we have $\|\delta_{T^c_u}\|_{1,\varpi} \leqslant 2\mathbf{c}\|\delta_{T_u}\|_{1,\varpi}$ which implies
$$\|\delta_u\|_{1,\varpi} \leqslant (1+2\mathbf{c})\|\delta_{T_u}\|_{1,\varpi} \leqslant (1+2\mathbf{c})\sqrt{s}\|\sqrt{f_u}Z^a\delta_u\|_{n,\varpi}/\kappa_{u,2\mathbf{c}}$$
by definition of $\kappa_{u,2\mathbf{c}}$. Thus we have $\delta_u \in A_u$ under $\Omega_1 \cap \Omega_2$.
Furthermore, ((ref)) also implies that
$$\begin{array}{rl}\|\delta_u +\beta_u\|_{1,\varpi} & \leqslant \|\beta_u\|_{1,\varpi}+\frac{1}{c}\|\delta_u\|_{1,\varpi}+\widehat R_u(\beta_u)/\lambda_u \\
&\leqslant (1+1/c)\|\beta_u\|_{1,\varpi}+(1/c)\|\delta_u +\beta_u\|_{1,\varpi}+\widehat R_u(\beta_u)/\lambda_u.\end{array}$$
which in turn establishes
$$\|\delta_u\|_{1,\varpi} \leqslant 2\frac{1+1/c}{1-1/c}\|\beta_u\|_{1,\varpi} + \frac{1}{\lambda_u(1-1/c)}\widehat R_u(\beta_u)\leqslant 2\frac{1+1/c}{1-1/c}\|\beta_u\|_{1,\varpi} + \frac{1}{\lambda_u(1-1/c)}\bar R_{u\xi} $$
where the last inequality holds under $\Omega_2$. Thus, $\|\delta_u\|_{1,\varpi}\leqslant \sqrt{n}$ under our condition.
In turn, $\delta_u$ is considered in the supremum that defines $\Omega_3$.
Under $\Omega_1 \cap \Omega_2 \cap \Omega_3$ we have
{ \begin{equation}\begin{array}{rl}
& {\mathbb{E}_n}[{\mathrm{E}}[K_\varpi(W)\{\rho_u(X_a-Z^a(\beta_u+ \delta_u))-\rho_u(X_a-Z^a\beta_u)\} \mid X_{-a}, W]\\
& \leqslant {\mathbb{E}_n}[K_\varpi(W)\{\rho_u(X_a-Z^a(\beta_u+ \delta_u))-\rho_u(X_a-Z^a\beta_u)\}] + t_3\|\delta_u\|_{1,\varpi}\\
& \leqslant \lambda_u\| \delta_u\|_{1,\varpi} + t_3\|\delta_u\|_{1,\varpi}\\
& \leqslant 2\mathbf{c}\left(1+ \frac{1}{\lambda_u}t_3\right)\bar R_{u\xi} + \|\sqrt{f_u}Z^a \delta_u\|_{n,\varpi}\left[\lambda_u+t_3\right] \frac{3\mathbf{c}\sqrt{s}}{\kappa_{u,2\mathbf{c}}}\\
\end{array}\end{equation}}
here we used the bound $\|\delta_u\|_{1,\varpi} \leqslant (1+2\mathbf{c})\sqrt{s}\|\sqrt{f_u}Z^a\delta_u\|_{n,\varpi}/\kappa_{u,2\mathbf{c}} + \frac{2\mathbf{c}}{\lambda_u}\bar R_{u\xi}$ under $\Omega_1\cap \Omega_2$.
Using Lemma (ref), since ((ref)) holds, we have for each $u\in \mathcal{U}$
$$\begin{array}{rl}
{\mathbb{E}_n}[{\mathrm{E}}[K_\varpi(W)\{\rho_u(X_a-Z^a(\beta_u+ \delta_u))-\rho_u(X_a-Z^a\beta_u)\} \mid X_{-a},W] \\
\geqslant - (\sqrt{\bar f}+1) \|r_{u}\|_{n,\varpi} \|\sqrt{f_u}Z^a \delta_u\|_{n,\varpi} - \sup_{u\in \mathcal{U}, j\in[p]} |{\mathbb{E}_n}[{\mathrm{E}}[ S_{uj} \vert X_{-a},W]/\widehat\sigma^{Z}_{a\varpi j} ]| \ \|\delta_u\|_{1,\varpi} \\
+ \frac{\|\sqrt{f_u}Z^a\delta_u\|_{n,\varpi}^2}{4} \wedge \{\bar q_{A_u}\|\sqrt{f_u}Z^a\delta_u\|_{n,\varpi}\}\end{array}$$
here we have ${\mathrm{E}}[ S_{iuj}\vert X_{i,-a},W_i]=0$ since $\tau = {\mathrm{P}}(X_a \leqslant Z^a\beta_u+r_u\vert X_{-a},W)$ by the definition of conditional quantile.
Note that for positive numbers $(t^2/4) \wedge q t \leqslant A + B t$ implies $t^2/4 \leqslant A+Bt$ provided $q/2 > B$ and $2q^2>A$. (Indeed, otherwise $(t^2/4) \geqslant q t$ so that $t\geqslant 4q$ which in turn implies that $2q^2 + qt/2 \leqslant (t^2/4) \wedge q t \leqslant A + Bt$.) Since $\bar q_{A_u}/4 \geqslant (\sqrt{\bar f}+1) \|r_{u}\|_{n,\varpi} + \left[ \{\lambda_u+t_3\}\frac{3\mathbf{c}\sqrt{s}}{\kappa_{u,2\mathbf{c}}}\right]$ and $\bar q_{A_u} \geqslant \{ 2\mathbf{c}\left(1+ \frac{t_3}{\lambda_u}\right)\bar R_{u\xi} \}^{1/2}$, the minimum on the right hand side is achieved by the quadratic part for all $u\in\mathcal{U}$. Therefore we have uniformly over $u\in\mathcal{U}$
{ $$ \frac{\|\sqrt{f_u}Z^a\delta_u\|_{n,\varpi}^2}{4} \leqslant 2\mathbf{c}\left(1+ \frac{t_3}{\lambda_u}\right)\bar R_{u\xi} + \|\sqrt{f_u}Z^a \delta_u\|_{n,\varpi}\left[(\sqrt{\bar f}+1) \|r_{u}\|_{n,\varpi}+ \{\lambda_u+t_3\}\frac{3\mathbf{c}\sqrt{s}}{\kappa_{u,2\mathbf{c}}}\right]
$$}
which implies that
{ $$\begin{array}{rl}
\|\sqrt{f_u}Z^a\delta_u\|_{n,\varpi} & \leqslant \sqrt{8\mathbf{c}\left(1+ \frac{t_3}{\lambda_u}\right)\bar R_{u\xi}} + \left[(\sqrt{\bar f}+1) \|r_{u}\|_{n,\varpi}+ \{\lambda_u+t_3\}\frac{3\mathbf{c}\sqrt{s}}{\kappa_{u,2\mathbf{c}}}\right].
\end{array}$$}
\end{proof}
\begin{lemma}[CIQGM, Event $\Omega_2$]
Under Condition CI we have ${\mathbb{E}_n}[{\mathrm{E}}[\widehat R_u(\beta_u)\vert X_{-a},\varpi]] \leqslant \bar f \|r_{u}\|_{n,\varpi}^2/2$, $\widehat R_u(\beta_u)\geqslant 0$ and
$${\mathrm{P}}\left( \sup_{u\in\mathcal{U}}\widehat R_u(\beta_u) \leqslant C\{1+ \bar f\} \{n^{-1}s(1+d_W)\log(p|V|n)\} \right) = 1-o(1).$$
\end{lemma}
\begin{proof}[Proof of Lemma (ref)]
We have that $\widehat R_u(\beta_u)\geqslant 0$ by convexity of $\rho_\tau$. Let $\varepsilon_{iu}=X_{ia} - Z_{i}^a \beta_u - r_{iu}$ where $\|\beta_u\|_0\leqslant s$ and $r_{iu}=Q_{X_a}(\tau\vert X_{-a},\varpi) - Z^a\beta_u$.
By Knight's identity ((ref)), $\widehat R_u(\beta_u) = -{\mathbb{E}_n}[ K_\varpi(W)r_{u} \int_0^1 1\{\varepsilon_{u} \leqslant -t r_{u}\} - 1\{\varepsilon_{u}\leqslant 0\} \ dt ]\geqslant 0$.
$$ \begin{array}{rl}
{\mathbb{E}_n}[{\mathrm{E}}[\widehat R_u(\beta_u)\vert X_{-a},\varpi] & = {\mathbb{E}_n}[ K_\varpi(W)r_{u} \int_0^1 F_{X_a\mid X_{-a},\varpi}(Z^a\beta_u+(1-t)r_u) - F_{X_a\mid X_{-a},\varpi}(Z^a\beta_u+r_u) \ dt]\\
& \leqslant {\mathbb{E}_n}[ K_\varpi(W)r_{u} \int_0^1 \bar f t r_{u} dt] \leqslant \bar f \|r_{u}\|_{n,\varpi}^2/2\leqslant C\bar f s/n.\end{array}$$
Since Condition CI assumes ${\mathrm{E}}[\|r_{u}\|_{n,\varpi}^2]\leqslant {\mathrm{P}}(\varpi)s/n$, by Markov's inequality we have ${\mathrm{P}}( \widehat R_u( \beta_u) \leqslant C\bar f s/n ) \geqslant 1/2$. Define $z_{iu} := -\int_0^1 1\{\varepsilon_{iu} \leqslant -t r_{iu}\} - 1\{\varepsilon_{iu}\leqslant 0\} \ dt$, so that $\widehat R_u(\beta_u) = {\mathbb{E}_n}[K_\varpi(W)r_{u}z_{u}]$ where $|z_{iu}|\leqslant 1$. By Lemma 2.3.7 in vdVaartWellner2007 (note that the Lemma does not require zero mean stochastic processes), for $t \geqslant 2 C\bar f s/n$ we have
$$ \frac{1}{2}{\mathrm{P}}\left( \sup_{u\in\mathcal{U}}|{\mathbb{E}_n}[K_\varpi(W)r_{u}z_{u}]| \geqslant t \right) \leqslant 2{\mathrm{P}}\left(\sup_{u\in\mathcal{U}}|{\mathbb{E}_n}[\varepsilon K_\varpi(W)r_{u}z_{u}]|>t/4\right) $$
where $\varepsilon_i, i=1,\ldots,n$ are Rademacher random variables independent of the data.
Next consider the class of functions $\mathcal{F}= \{ - K_\varpi(W)r_{u} (1\{\varepsilon_{iu} \leqslant -B_i r_{iu}\} - 1\{\varepsilon_{iu}\leqslant 0\} ) : u\in\mathcal{U}\}$ where $B_i\sim {\rm Uniform}(0,1)$ independent of $(X_i,W_i)_{i=1}^n$. It follows that $K_\varpi(W)r_{u}z_{u} = {\mathrm{E}}[ - K_\varpi(W)r_{u} (1\{\varepsilon_{iu} \leqslant -B_i r_{iu}\} - 1\{\varepsilon_{iu}\leqslant 0\} ) \vert X_i,W_i]$ where the expectation is taken over $B_i$ only. Thus we will bound the entropy of $\overline{\mathcal{F}}= \{ {\mathrm{E}}[ f\vert X,W]: f \in \mathcal{F} \}$ via Lemma (ref). Note that $\mathcal{R}:=\{ r_u = Q_{X_a}(\tau\vert X_{-a},\varpi) - Z^a\beta_u: u\in\mathcal{U}\}$ where $\mathcal{G}:=\{ Z^a\beta_u: u\in\mathcal{U}\}$ is contained in the union of at most $|V|\binom{p}{s}$ VC-classes of dimension $Cs$ and $\mathcal{H}:=\{Q_{X_a}(\tau\vert X_{-a},\varpi): u\in\mathcal{U}\}\}$ is the union of $|V|$ VC-class of functions of dimension $(1+d_W)$ by Condition CI. Finally note that $\mathcal{E}:=\{ \varepsilon_{iu}: u\in \mathcal{U}\} \subset \{ X_{ia}: a\in V\} - \mathcal{G} - \mathcal{R}$.
Therefore, we have
$$
\begin{array}{rl}
\sup_Q \log N(\epsilon \|\bar F\|_{Q,2}, \overline{\mathcal{F}}, \|\cdot\|_{Q,2}) & \leqslant \sup_Q \log N( (\epsilon/4)^2 \| F\|_{Q,2}, \mathcal{F}, \|\cdot\|_{Q,2})\\
& \leqslant \sup_Q \log N( \mbox{$\frac{1}{8}$}(\epsilon^2/16), \mathcal{W}, \|\cdot\|_{Q,2}) \\
&+ \sup_Q\log N( \mbox{$\frac{1}{8}$} (\epsilon^2/16) \| F\|_{Q,2}, \mathcal{R}, \|\cdot\|_{Q,2}) \\
&+ \sup_Q \log N( \mbox{$\frac{1}{8}$}(\epsilon^2/16), 1\{ \mathcal{E}+\{B\}\mathcal{R} \leqslant 0\} - 1\{ \mathcal{E} \leqslant 0\}, \|\cdot\|_{Q,2})\\
\end{array}
$$
We will apply Lemma (ref) with envelope $\bar F = \sup_{u\in\mathcal{U}} |K_\varpi(W)r_u|$, so that $ {\mathrm{E}}[\max_{i\leqslant n}\bar F_i^2]\leqslant C$, and $\sup_{u\in \mathcal{U}} {\mathrm{E}}[K_\varpi(W)r_{u}^2 ] \leqslant Cs/n =: \sigma^2$ by Condition CI. Thus, we have that with probability $1-o(1)$
$$ \sup_{u\in\mathcal{U}}|{\mathbb{E}_n}[\varepsilon K_\varpi(W)r_{u}z_{u}]| \lesssim \sqrt{\frac{s(1+d_W)\log(p|V|n)}{n}}\sqrt{\frac{s}{n}}+ \frac{s(1+d_W)\log(p|V|n)}{n}\lesssim \frac{s(1+d_W)\log(p|V|n)}{n}$$
under $M_n\sqrt{s^2/n} \leqslant C$.
\end{proof}
\begin{lemma}[CIQGM, Event $\Omega_3$]
For $u=(a,\tau,\varpi)\in \mathcal{U} := V\times \mathcal{T}\times \mathcal{W}$, define the function $g_u(\delta,X,W)= K_\varpi(W)\{\rho_\tau(X_a-Z^a(\beta_u+\delta))-\rho_\tau(X_a-Z^a\beta_u)\}$, and the event
$$ \Omega_3 := \left\{ \sup_{u\in\mathcal{U}, 1/\sqrt{n} \leqslant \|\delta\|_{1,\varpi} \leqslant \sqrt{n}} \frac{|{\mathbb{E}_n}[g_u(\delta,X,W)-{\mathrm{E}}[g_u(\delta,X,W)\mid X_{-a}, W]]|}{\|\delta\|_{1,\varpi}} < t_3 \right\}.$$
Then, under Condition CI we have $P(\Omega_3) \geqslant 1-\xi$ for any $t_3$ satisfying
$$ t_3\sqrt{n} \geqslant 12 + 16\sqrt{2\log(64|V|p^2 n^{3+2d_W}\log(n)L^{1+d_W/\kappa}_\beta M_n/\xi )} $$
\end{lemma}
\begin{proof}
We have that $\Omega_3^c := \{ \max_{a\in V} A_a \geqslant t_3\sqrt{n}\}$ for
$$A_a:= \sup_{(\tau,\varpi)\in\mathcal{T}\times\mathcal{W}, \underline{N} \leqslant \|\delta\|_{1,\varpi} \leqslant \bar{N}} \sqrt{n}\left|\frac{{\mathbb{E}_n}[g_u(\delta,X,W)-{\mathrm{E}}[g_u(\delta,X,W)\mid X_{-a},W]]}{\|\delta\|_{1,\varpi}}\right|. $$
Therefore, for $\underline{N} = 1/\sqrt{n}$ and $\bar{N} = \sqrt{n}$ we have by Lemma (ref) with $\rho = \kappa$, $L_{\eta} = L_{\beta}$, $\tilde{x} = Z^a$
$$\begin{array}{rl}
{\mathrm{P}}(\Omega_3^c) & = {\mathrm{P}}(\max_{a\in V} A_a \geqslant t_3\sqrt{n}) \\
& \leqslant |V|\max_{a\in V} {\mathrm{P}}(A_a \geqslant t_3\sqrt{n}) \\
& = |V|\max_{a\in V} {\mathrm{E}}_{X_{-a},W}\left\{ {\mathrm{P}}(A_a \geqslant t_3\sqrt{n} \mid X_{-a},W )\right\} \\
& \leqslant |V| \max_{a\in V} {\mathrm{E}}_{X_{-a},W}\left\{8p |\widehat{\mathcal{N}}|\cdot |\widehat{\mathcal{W}}|\cdot |\widehat{\mathcal{T}}| \exp(-(t_3\sqrt{n}/4-3)^2/32)\right\}\\
& \leqslant \exp(-(t_3\sqrt{n}/4-3)^2/32) |V|64p n^{1+d_W}\log(n)L_\beta {\mathrm{E}}_{X_{-a}}\left\{\frac{\max_{i\leqslant n}\|Z^a_i\|_\infty^{1+d_W/\kappa}}{\underline{N}^{1+d_W}}\right\}\\
& \leqslant \xi \end{array} $$
by the choice of $t_3$ and noting that $M_n^{(1+d_W/\kappa)/q} \geqslant {\mathrm{E}}_{X_{-a}}[\max_{i\leqslant n}\|Z^a_i\|_\infty^{1+d_W/\kappa}]$, $ 1+d_W/\kappa \leqslant q$ and $M_n \geqslant 1$.
\end{proof}
\begin{lemma}[CIQGM, Uniform Control of Approximation Error in Auxiliary Equation]
Under Condition CI, with probability $1-o(1)$ uniformly over $u\in\mathcal{U}$ and $j\in[p]$ we have
$$
{\mathbb{E}_n}[K_\varpi(W)f_u^2\{Z_{-j}^a(\gamma_u^j-\bar\gamma_u^j)\}^2] \lesssim {\underline{f}}_u^2{\mathrm{P}}(\varpi) \{n^{-1}s\log (p|V| n)\}^{1/2}.
$$
\end{lemma}
\begin{proof}
Define the class of functions ${\mathcal{G}} = \cup_{a\in V, j\in [p]}{\mathcal{G}}_{aj}$ with ${\mathcal{G}}_{aj}:=\{ Z_{-j}^a(\gamma_u^j-\bar\gamma_u^j) : \tau \in \mathcal{T}, \varpi\in \mathcal{W}\}$.
Under Condition CI we have $\sup_{u\in\mathcal{U}}\|\bar\gamma_u^j\|_0 \leqslant Cs$, $\sup_{u\in\mathcal{U}, j\in [p]}\|\bar \gamma_u^j-\gamma_u^j\| \vee \frac{\|\bar\gamma_u^j-\gamma_u^j\|_1}{\sqrt{s}}\leqslant \{n^{-1}s\log (p|V| n)\}^{1/2}.$
Without loss of generality we can set $\bar\gamma_{uk}^j = \gamma_{uk}^j$ for $k \in {\rm support}(\bar\gamma_u^j)$. Letting ${\mathcal{G}}_{aj,T}:=\{ Z_{-j}^a(\gamma_u^j-\gamma_{uT}^j) : \tau \in \mathcal{T}, \varpi\in \mathcal{W}\}$ for $T\subset \{1,\ldots,p\}$, it follows that $ {\mathcal{G}} \subset \cup_{a\in V, j\in [p]}\cup_{|T|\leqslant Cs}{\mathcal{G}}_{aj,T}$.
By Lemma (ref), we have $\|\gamma_u^j-\gamma_{u'}^j\|\leqslant L_\gamma (\|u-u'\|+\|u-u'\|^{1/2})$ for each $a\in V$, $j\in[p]$. (Note that although $\bar\gamma^j_u$ might not be Lipschitz in $u$, however, for each $T$, $\gamma_{uT}^j$ satisfies the same Lipschitz relation as $\gamma_u^j$, in fact $\|\bar\gamma_{uT}^j-\gamma_{u'T}^j\|\leqslant \|\gamma_u^j-\gamma_{u'}^j\|$ by construction.) Therefore, for each $T$ we have
$$\begin{array}{rl}
& \|\{Z_{-j}^a(\bar\gamma_{uT}^j-\gamma_u^j)\}^2-\{Z_{-j}^a(\bar\gamma_{u'T}^j-\gamma_{u'}^j)\}^2\|_{Q,2} \\
& \leqslant \|Z_{-j}^a(\bar\gamma_{uT}^j-\bar\gamma_{u'T}^j+\gamma_{u'}^j-\gamma_u^j)Z_{-j}^a(\bar\gamma_{uT}^j-\gamma_u^j+\bar\gamma_{u'T}^j-\gamma_{u'}^j)\|_{Q,2}\\
& \leqslant \|\|Z_{-j}^a\|_\infty^2\|_{Q,2} \|\bar\gamma_{uT}^j-\bar\gamma_{u'T}^j+\gamma_{u'}^j-\gamma_u^j\|_1\|\bar\gamma_{uT}^j-\gamma_u^j+\bar\gamma_{u'T}^j-\gamma_{u'}^j\|_1\\
& \leqslant 4\|\|Z_{-j}^a\|_\infty^2\|_{Q,2}\sup_{u\in\mathcal{U}}\|\bar\gamma_{uT}^j-\gamma_u^j\|_1 \sqrt{2p}\|\gamma_u^j-\gamma_{u'}^j\|\\
&\leqslant \|\|Z_{-j}^a\|_\infty^2\|_{Q,2}L_\gamma'(\|u-u'\|+\|u-u'\|^{1/2}).\end{array}$$
where $L_\gamma'=4\{n^{-1}s^2\log(p|V|n)\}^{1/2}\sqrt{2p}L_\gamma$. Thus, for the envelope $G = \max_{a\in V}\|Z^a\|_\infty^2\sup_{u\in\mathcal{U}} \|\bar\gamma_u^j-\gamma_u^j\|_1^2$ that
$$ \begin{array}{rl}\log N(\epsilon \|G\|_{Q,2}, {\mathcal{G}}, \|\cdot\|_{Q,2}) \leqslant Cs\log(|V|p)+\log N\left(\epsilon\frac{\sup_{u\in\mathcal{U}} \|\bar\gamma_u^j-\gamma_u^j\|_1^2}{L_\gamma'}, \mathcal{U}, d_\mathcal{U}\right) \leqslant Cs(1+d_W)^2 \log (L_\gamma'n/\epsilon). \end{array}$$
Next define the functions $\mathcal{W}_0=\{ K_\varpi(W)f_u^2 : u\in\mathcal{U} \}$, $\mathcal{W}_1=\{ {\mathrm{P}}( \varpi)^{-1} : \varpi \in \mathcal{W} \}$ and $\mathcal{W}_2=\{ K_\varpi(W) : \varpi \in \mathcal{W} \}$. We have that $\mathcal{W}_2$ is VC class with VC index $Cd_W$ and $\mathcal{W}_1$ is bounded by $\mu_\mathcal{W}^{-1}$ and covering number bounded by $(Cd_W/\{\mu_\mathcal{W}\epsilon\})^{1+d_W}$. Finally, since $|K_\varpi(W)f_u^2-K_{\varpi'}(W)f_{u'}^2|\leqslant K_\varpi(W)K_{\varpi'}(W)|f_u^2-f_{u'}^2| + \bar f^2 |K_\varpi(W)-K_{\varpi'}(W)|\leqslant 2\bar fL_f\|u-u'\|+\bar f^2 |K_\varpi(W)-K_{\varpi'}(W)|$, we have $N(\epsilon,\mathcal{U},|\cdot|) \leqslant (C(1+d_W)/\epsilon)^{1+d_W}$. Therefore, using standard bounds we have
$$ \log N(\epsilon \|\mu_\mathcal{W}^{-1}G\bar f\|_{Q,2}, \mathcal{W}_0\mathcal{W}_1\mathcal{W}_2{\mathcal{G}}, \|\cdot\|_{Q,2}) \lesssim s(1+d_W)^2 \log (L_\gamma' L_f n/\epsilon)$$
By Lemma (ref) we have that with probability $1-o(1)$ that
$$ \begin{array}{l} \sup_{u\in \mathcal{U}, j\in[p]} |({\mathbb{E}_n}-{\mathrm{E}})[f_u^2\{Z_{-j}^a(\gamma_u^j-\bar\gamma_u^j)\}^2/{\mathrm{P}}(\varpi)]| \\
\lesssim \sqrt{ \frac{s(1+d_W)^2 \log(p|V| n) \sup_{u\in\mathcal{U}}{\mathrm{E}}[K_\varpi(W)f_u^4\{Z_{-j}^a(\gamma_u^j-\bar\gamma_u^j)\}^4]/{\mathrm{P}}(\varpi)^2}{n}} + \frac{s(1+d_W)^2M_n^2\mu_\mathcal{W}^{-1}\sup_{u\in\mathcal{U}} \|\bar\gamma_u^j-\gamma_u^j\|_1^2 \log(p|V|n)}{n}\\
\lesssim \sqrt{\frac{s(1+d_W)^2\log(p|V|n)}{\mu_\mathcal{W}n}}\frac{s\log(p|V|n)}{n} + \frac{(1+d_W)^2M_n^2s^2\log(p|V|n)}{n\mu_\mathcal{W}}\frac{s\log(p|V|n)}{n}\\
\lesssim \frac{s\log(p|V|n)}{n}\mu_\mathcal{W}{\underline{f}}_\mathcal{U} \left\{ \sqrt{\frac{s(1+d_W)^2\log(p|V|n)}{\mu_\mathcal{W}^3{\underline{f}}_\mathcal{U}^2 n}} + \frac{(1+d_W)^2M_n^2s^2\log(p|V|n)}{n\mu_\mathcal{W}^2{\underline{f}}_\mathcal{U}} \right\}\\
\lesssim \frac{s\log(p|V|n)}{n}\mu_\mathcal{W}{\underline{f}}_\mathcal{U} \{ \delta_n^{1/2} + \delta_n^2 \}
\end{array}$$
here we used that ${\mathrm{E}}[f_u^4\{Z^a\delta\}^4\vert \varpi] \leqslant \bar f^4 {\mathrm{E}}[\{Z^a\delta\}^4\vert \varpi] \leqslant C\|\delta\|^4$, $\|\bar\gamma_u^j-\gamma_u^j\|+s^{-1/2}\|\bar\gamma_u^j-\gamma_u^j\|_1 \leqslant \{n^{-1}s\log(p|V|n)\}^{1/2}$, $ s(1+d_W)^2\log(p|V|n)\leqslant \delta_n n {\underline{f}}_\mathcal{U}^2\mu_\mathcal{W}^3$ and $(1+d_W)M_n s\log^{1/2}(p|V|n)\leqslant \delta_n n^{1/2}\mu_\mathcal{W}{\underline{f}}_\mathcal{U}$ by Condition CI. Furthermore, by Condition CI, the result follows from ${\mathrm{E}}[f_u^2\{Z_{-j}^a(\bar\gamma_u^j-\gamma_u^j)\}^2\vert \varpi] \leqslant C{\underline{f}}^2_u\|\bar\gamma_u^j-\gamma_u^j\|^2 \leqslant C {\underline{f}}^2_u n^{-1} s\log(p|V|n)$.
\end{proof}
\begin{lemma}
Under Condition CI, for $u=(a,\tau,\varpi) \in \mathcal{U}$ and $u'=(a,\tau',\varpi')\in \mathcal{U}$ we have that $$\|\gamma_u^j-\gamma_{u'}^j\| \leqslant \frac{C'}{\underline{f}_{u'}^2{\mathrm{P}}(\varpi')}\{ \bar f^2{\mathrm{E}}[\{K_{\varpi'}(W)-K_{\varpi}(W)\}^2]^{1/2}+ {\mathrm{E}}[K_\varpi(W)K_{\varpi'}(W)\{f_{u'}^2-f_u^2\}^2]^{1/2}\}. $$ In particular, we have $\|\gamma_u^j-\gamma_{u'}^j\| \leqslant L_\gamma \{ \|\varpi-\varpi'\|^{1/2}+\|u-u'\| \}$ for $L_\gamma = C\{L_f+L_K\}/\{{\underline{f}}_{\mathcal{U}}^2\mu_\mathcal{W}\}$ under ${\mathrm{E}}[|K_\varpi(W)-K_{\varpi'}(W)|] \leqslant L_K\|\varpi-\varpi'\|$, $K_\varpi(W)K_{\varpi'}(W)|f_{u'}-f_u|\leqslant L_f\|u'-u\|$, and $f_u \leqslant \bar f\leqslant C$.
\end{lemma}
\begin{proof} Let $u=(a,\tau,\varpi)$ and $u'=(a,\tau',\varpi')$. By Condition CI we have $$\begin{array}{rl}
\|\gamma_u^j-\gamma_{u'}^j\|^2 & \leqslant C {\mathrm{E}}[\{Z^a_{-j}(\gamma_u^j-\gamma_{u'}^j)\}^2\vert \varpi' ] \leqslant \{C/{\mathrm{P}}(\varpi')\} {\mathrm{E}}[K_{\varpi'}(W)\{Z^a_{-j}(\gamma_u^j-\gamma_{u'}^j)\}^2]\end{array}$$
To bound the last term of the right hand side above, by the definition of $\underline{f}_{u'}$ and Cauchy-Schwarz's inequality we have
$$\begin{array}{rl}
\underline{f}_{u'}{\mathrm{E}}[K_{\varpi'}(W)\{Z^a_{-j}(\gamma_u^j-\gamma_{u'}^j)\}^2] & \leqslant {\mathrm{E}}[K_{\varpi'}(W)f_{u'}\{Z^a_{-j}(\gamma_u^j-\gamma_{u'}^j)\}^2]\\
& \leqslant \{{\mathrm{E}}[K_{\varpi'}(W)f_{u'}^2\{Z^a_{-j}(\gamma_u^j-\gamma_{u'}^j)\}^2] \ {\mathrm{E}}[K_{\varpi'}(W)\{Z^a_{-j}(\gamma_u^j-\gamma_{u'}^j)\}^2]\}^{1/2}\\
\end{array}
$$
so that ${\mathrm{E}}[K_{\varpi'}(W)\{Z^a_{-j}(\gamma_u^j-\gamma_{u'}^j)\}^2]^{1/2}\leqslant \{{\mathrm{E}}[K_{\varpi'}(W)f_{u'}^2\{Z^a_{-j}(\gamma_u^j-\gamma_{u'}^j)\}^2]\}^{1/2}/\underline{f}_{u'}$. Therefore
\begin{equation}\|\gamma_u^j-\gamma_{u'}^j\|^2 \leqslant \{1/\underline{f}_{u'}\}^2\{C/{\mathrm{P}}(\varpi)\} {\mathrm{E}}[K_{\varpi'}(W)f_{u'}^2\{Z^a_{-j}(\gamma_u^j-\gamma_{u'}^j)\}^2].\end{equation}
We proceed to bound the last term. The optimality of $\gamma_u^j$ and $\gamma_{u'}^j$ yields
$$ {\mathrm{E}}[K_\varpi(W)f_u^2Z^a_{-j}(Z_j^a - Z^a_{-j}\gamma_u^j)]=0 \ \ \ \mbox{and} \ \ \ {\mathrm{E}}[K_{\varpi'}(W)f_{u'}^2Z^a_{-j}(Z_j^a - Z^a_{-j}\gamma_{u'}^j)]=0 $$
Therefore, we have \begin{equation}\begin{array}{rl}
{\mathrm{E}}[K_{\varpi'}(W)f_{u'}^2\{Z^a_{-j}(\gamma_u^j-\gamma_{u'}^j)\}Z^a_{-j}]& = -{\mathrm{E}}[K_{\varpi'}(W)f_{u'}^2\{Z^a_j-Z^a_{-j}\gamma_u^j\}Z^a_{-j}] \\
& = -{\mathrm{E}}[\{K_{\varpi'}(W)f_{u'}^2-K_{\varpi}(W)f_u^2\}\{Z^a_j-Z^a_{-j}\gamma_u^j\}Z^a_{-j}] \\
\end{array}
\end{equation}
Multiplying by $(\gamma_u^j-\gamma_{u'}^j)$ both sides of ((ref)), we have $$\begin{array}{rl}
&{\mathrm{E}}[K_{\varpi'}(W)f_{u'}^2\{Z^a_{-j}(\gamma_u^j-\gamma_{u'}^j)\}^2] \\
& \leqslant {\mathrm{E}}[\{K_{\varpi'}(W)f_{u'}^2-K_{\varpi}(W)f_u^2\}^2]^{1/2}\{{\mathrm{E}}[\{Z^a_j-Z^a_{-j}\gamma_u^j\}^2\{Z^a_{-j}(\gamma_u^j-\gamma_{u'}^j)\}^2]\}^{1/2}\\
& \leqslant {\mathrm{E}}[\{K_{\varpi'}(W)f_{u'}^2-K_{\varpi}(W)f_u^2\}^2]^{1/2}C \|\gamma_u^j-\gamma_{u'}^j\|\\
\end{array}$$
by the fourth moment assumption in Condition CI. By Condition CI, $f_u,f_{u'}\leqslant \bar f$, and it follows that
\begin{equation}|K_\varpi(W)f_u^2-K_{\varpi'}(W)f_{u'}^2| \leqslant K_\varpi(W)K_{\varpi'}(W)|f_u^2-f_{u'}^2|+\bar f^2 |K_{\varpi}(W)-K_{\varpi'}(W)|\end{equation}
From ((ref)) we obtain
$$\|\gamma_u^j-\gamma_{u'}^j\| \leqslant \frac{C'}{\underline{f}_{u'}^2{\mathrm{P}}(\varpi')}\{\bar f^2 {\mathrm{E}}[\{K_{\varpi'}(W)-K_{\varpi}(W)\}^2]^{1/2}+ {\mathrm{E}}[K_\varpi(W)K_{\varpi'}(W)\{f_{u'}^2-f_u^2\}^2]^{1/2}\}. $$
\end{proof}
\begin{lemma}
Let $\mathcal{U} = V \times \mathcal{T} \times \mathcal{W}$. Under Condition CI, for $m=1,2$, we have
$$ {\mathrm{E}}\left[ \sup_{u \in \mathcal{U}, \|\theta\|_0 \leqslant k, \|\theta\|=1 } | ({\mathbb{E}_n}-{\mathrm{E}})[K_\varpi(W)f_u^m(Z^a\theta)^2] | \right] \lesssim C \delta_n \sup_{u \in \mathcal{U}, \|\theta\|_0 \leqslant k, \|\theta\|=1 } \{ {\mathrm{E}}[K_\varpi(W)f_u^m(Z^a\theta)^2]\}^{1/2} $$
where $\delta_n = M_n\sqrt{k(1+d_W)C\log(p|V|n)}\log (1+k) \sqrt{\log n/n}$. Moreover, under Condition CI, $\delta_n = o(\mu_{\mathcal{W}})$.
\end{lemma}
\begin{proof}
By symmetrization we have
$${\mathrm{E}}\left[ \sup_{u \in \mathcal{U}, \|\theta\|_0 \leqslant k, \|\theta\|=1 } | ({\mathbb{E}_n}-{\mathrm{E}})[K_\varpi(W)f_u^m(Z^a\theta)^2] | \right] \leqslant 2 {\mathrm{E}}\left[\sup_{u \in \mathcal{U}, \|\theta\|_0 \leqslant k, \|\theta\|=1 } | {\mathbb{E}_n}[\varepsilon K_\varpi(W)f_u^m(Z^a\theta)^2] | \right] $$
where $\varepsilon_i$ are i.i.d. Rademacher random variables. We have that
$|K_\varpi(W) f_u^m-K_{\varpi'}(W)f_{u'}^m| \leqslant K_\varpi(W) K_{\varpi'}(W)|f_u - f_{u'} | (1 + 2\bar f)+\bar f^m|K_\varpi(W)- K_{\varpi'}(W)|$ for $m=1,2$ where $u$ and $u'$ have the same $a\in V$. However, conditional on $\{(W_i,X_i), i=1,\ldots,n\}$, $\{K_\varpi(W_i):i=1,\ldots,n, \varpi \in \mathcal{W}\}$ induces at most $n^{d_W}$ different sequences by Corollary 2.6.3 in vdV-W. This induces (at most) $n^{d_W}$ partitions of $\mathcal{W}$ such that $K_\varpi(W) = K_{\varpi'}(W)$ for any $\varpi,\varpi'$ in the same partition given the conditioning. Thus, for such suitable $\varpi'$ we have $|K_\varpi(W) f_u^m-K_{\varpi'}(W)f_{u'}^m| \leqslant K_\varpi(W) |f_{u} - f_{u'} |(1 + 2\bar f)$ for $m=1,2$. (Thus it suffices to create a net for each partition.) We can take a cover $\widehat{\mathcal{U}}$ of $V\times \mathcal{T}\times \mathcal{W}$ such that $\|u-u'\|\leqslant \{L_f(1+2\bar f)nk\max_{i\leqslant n}\|Z^a_i\|_\infty^2 \}^{-1}$ so that $|f_u-f_{u'}|(Z^a\theta)^2 \leqslant |f_u-f_{u'}|\|Z^a\|_\infty^2\|\theta\|_1^2 \leqslant|f_u-f_{u'}|\|Z^a\|_\infty^2k\|\theta\|^2$ which implies $$\begin{array}{rl}
\displaystyle \left|\sup_{u \in \mathcal{U}, \|\theta\|_0 \leqslant k, \|\theta\|=1 } | {\mathbb{E}_n}[\varepsilon K_\varpi(W)f_u^m(Z^a\theta)^2] | - \sup_{u \in \widehat{\mathcal{U}}, \|\theta\|_0 \leqslant k, \|\theta\|=1 } | {\mathbb{E}_n}[\varepsilon K_\varpi(W)f_u^m(Z^a\theta)^2] | \right| \leqslant n^{-1}\end{array}$$
Consequentially
$$ {\mathrm{E}}\left[\sup_{u \in \mathcal{U}, \|\theta\|_0 \leqslant k, \|\theta\|=1 } | {\mathbb{E}_n}[\varepsilon K_\varpi(W)f_u^m(Z^a\theta)^2] | \right] \leqslant {\mathrm{E}}\left[\sup_{u \in \widehat{\mathcal{U}}, \|\theta\|_0 \leqslant k, \|\theta\|=1 } | {\mathbb{E}_n}[\varepsilon K_\varpi(W)f_u^m(Z^a\theta)^2] | \right] + \frac{1}{n}$$
where $|\widehat{\mathcal{U}}| \leqslant |V| n^{d_W} \{L_f(1+2\bar f)n k\max_{i\leqslant n}\|Z^a_i\|_\infty^2\}^{(1+d_W)}$.
By Lemma (ref) with $K=K(W,X)= (1+\bar f^2)\sup_{a\in V}\max_{i\leqslant n}\|Z_i^a\|_\infty$ and
$$\begin{array}{rl}
\delta_n(W,X) & := \bar C K(W,X) \sqrt{k}\left(\sqrt{\log |\widehat\mathcal{U}|} + \sqrt{1+\log p} + \log k \sqrt{\log (p\vee n)} \sqrt{\log n} \right)/\sqrt{n} \\
& \lesssim K(W,X) \sqrt{k(1+d_W)C\log(p|V|nK(W,X))}\log (1+k) \sqrt{\log n}/\sqrt{n} \\
\end{array}$$
so that conditional on $(W,X)$ we have
$${\mathrm{E}}\left[\sup_{u \in \widehat{\mathcal{U}}, \|\theta\|_0 \leqslant k, \|\theta\|=1 } | {\mathbb{E}_n}[\varepsilon K_\varpi(W)f_u^m(Z^a\theta)^2] | \right]\lesssim \delta_n(W,X) \sup_{u \in \widehat{\mathcal{U}}, \|\theta\|_0 \leqslant k, \|\theta\|=1 }\sqrt{{\mathbb{E}_n}[K_\varpi(W)f_u^m(Z^a\theta)^2]} $$
Therefore,
$$\begin{array}{l}
\displaystyle {\mathrm{E}}\left[\sup_{u \in \mathcal{U}, \|\theta\|_0 \leqslant k, \|\theta\|=1 } | {\mathbb{E}_n}[\varepsilon K_\varpi(W)f_u^m(Z^a\theta)^2] | \right] \\
\displaystyle \leqslant {\mathrm{E}}_{W,X}\left[\delta_n(W,X) \sup_{u \in \widehat{\mathcal{U}}, \|\theta\|_0 \leqslant k, \|\theta\|=1 }\sqrt{{\mathbb{E}_n}[K_\varpi(W)f_u^m(Z^a\theta)^2]}\right] + \frac{1}{n} \\
\displaystyle \leqslant {\mathrm{E}}_{W,X}[\delta_n^2(W,X)] + {\mathrm{E}}_{W,X}[\delta_n^2(W,X)]^{1/2} \sup_{u \in \widehat{\mathcal{U}}, \|\theta\|_0 \leqslant k, \|\theta\|=1 }{\mathrm{E}}[K_\varpi(W)f_u^m(Z^a\theta)^2]^{1/2} + \frac{1}{n}
\end{array}
$$
Note that for a random variable $A\geqslant 1$, we have that ${\mathrm{E}}[A^2\sqrt{\log(CA)}]\leqslant {\mathrm{E}}[A^2]\sqrt{\log(C)} + {\mathrm{E}}[A^2\sqrt{\log(A)}] \leqslant {\mathrm{E}}[A^2]\sqrt{\log(C)} + {\mathrm{E}}[A^{2+1/4}]$.
Therefore, under Condition CI, since $q\geqslant 2+1/4$ in the definition of $M_n$, we have
$$ {\mathrm{E}}_{W,X}[\delta_n^2(W,X)]^{1/2} \lesssim M_n\sqrt{k(1+d_W)C\log(p|V|n)}\log (1+k) \sqrt{\log n}/\sqrt{n}.$$
The results follows by setting $\delta_n={\mathrm{E}}_{W,X}[\delta_n^2(W,X)]^{1/2}$.
\end{proof}
\section{Results for Prediction Quantile Graphical Models}
In the analysis of PQGM we also use the following event for some sequence $(K_u)_{u\in\mathcal{U}}$
\begin{equation} \begin{array}{c}\Omega_4 =\{ K_{u} \widehat{\sigma}^{X}_{a\varpi j} \geqslant {\mathbb{E}_n}[{\mathrm{E}}[ K_\varpi(W)(\tau-1\{X_a\leqslant X_{-a}'\beta_u+r_u\})X_{-a}\vert X_j,W ]], \ u\in\mathcal{U}, j\in V\backslash\{a\} \}.\end{array} \end{equation}
\begin{lemma}[Rate for PQGM]
Suppose that $\Omega_1$, $\Omega_2$, $\Omega_3$ and $\Omega_4$ hold. Further assume $2\frac{1+1/c}{1-1/c}\|\beta_u\|_{1,\varpi} + \frac{1}{\lambda_u(1-1/c)}\bar R_{u\xi} \leqslant \sqrt{n}$ for all $u\in\mathcal{U}$, and ((ref)) holds for all $\delta\in A_u := \Delta_{\varpi,2\mathbf{c}}\cup\{v:\|v\|_{1,\varpi}\leqslant 2\mathbf{c}\bar R_{u\xi}/\lambda_u\}$, $\bar q_{A_u}/4 \geqslant (\sqrt{\bar f}+1) \|r_{u}\|_{n,\varpi} + \left[\lambda_u+t_3+K_u\right]\frac{3\mathbf{c}\sqrt{s}}{\kappa_{u,2\mathbf{c}}}$ and $\bar q_{A_u} \geqslant \{ 2\mathbf{c}\left(1+ \frac{t_3+K_u}{\lambda_u}\right)\bar R_{u\xi} \}^{1/2}$.
Then uniformly over all $u=(a,\tau,\varpi)\in \mathcal{U} := V\times \mathcal{T}\times \mathcal{W}$, the $\|\cdot\|_{1,\varpi}$-penalized estimator $\widehat\beta_u$ satisfies
{ $$\begin{array}{rl}
\|\sqrt{f_u}X_{-a}'(\widehat\beta_u - \beta_u)\|_{n,\varpi} & \leqslant \sqrt{8\mathbf{c}\left(1+ \frac{t_3}{\lambda_u}\right)\bar R_{u\xi}} + (\bar f^{1/2}+1) \|r_{u}\|_{n,\varpi}+ [\lambda_u+t_3+K_u]\frac{3\mathbf{c}\sqrt{s}}{\kappa_{u,2\mathbf{c}}}\\
\|\widehat\beta_u - \beta_u\|_{1,\varpi} & \leqslant (1+2\mathbf{c})\sqrt{s}\|\sqrt{f_u}X_{-a}'\delta_u\|_{n,\varpi}/\kappa_{u,2\mathbf{c}} + \frac{2\mathbf{c}}{\lambda_u}\bar R_{u\xi} \\
\end{array}$$}\end{lemma}
\begin{proof}[Proof of Lemma (ref)]
The proof proceeds similarly to the proof of Lemma (ref) by defining
$$\begin{array}{rl}
\widehat R_u(\beta) & = {\mathbb{E}_n}[K_\varpi(W)\{\rho_u(X_a-X_{-a}'\beta)-\rho_u(X_a-X_{-a}'\beta_u-r_{u})\}]\\
& -{\mathbb{E}_n}[K_\varpi(W)\{(\tau-1\{X_a\leqslant X_{-a}'\beta_u+r_{u}\})(X_{-a}'\beta-X_{-a}'\beta_u-r_{u})\}].\end{array}$$
The same argument yields $\delta_u=\widehat\beta_u-\beta_u \in A_u := \Delta_{\varpi,2\mathbf{c}}\cup \{ v : \|v\|_{1,\varpi}\leqslant 2\mathbf{c} \bar R_{u\xi}/\lambda_u\}$ under $\Omega_1 \cap \Omega_2$. (Similarly we also have $\|\delta_u\|_{1,\varpi}\leqslant \sqrt{n}$.) Furthermore, under $\Omega_1\cap \Omega_2\cap\Omega_3$ we have that ((ref)) also holds which implies
$${\mathbb{E}_n}[{\mathrm{E}}[K_\varpi(W)\{\rho_u(X_a-X_{-a}'(\beta_u+ \delta_u))-\rho_u(X_a-X_{-a}'\beta_u)\} \vert X_{-a},W] \leqslant (\lambda_u+t_3)\|\delta_u\|_{1,\varpi}$$
Since the conditions of Lemma (ref) hold we have
$$\begin{array}{rl}
{\mathbb{E}_n}[{\mathrm{E}}[K_\varpi(W)\{\rho_\tau(X_a-X_{-a}'(\beta_u+ \delta_u))-\rho_\tau(X_a-X_{-a}'\beta_u)\} \vert X_{-a},W] \\
\geqslant - (\sqrt{\bar f}+1) \|r_{u}\|_{n,\varpi} \|\sqrt{f_u}X_{-a}' \delta\|_{n,\varpi} - K_u\|\delta_u\|_{1,\varpi} \\
+ \frac{\|\sqrt{f_u}X_{-a}'\delta_u\|_{n,\varpi}^2}{4} \wedge \bar q_{A_u}\|\sqrt{f_u}X_{-a}'\delta_u\|_{n,\varpi}
\end{array}$$
where $K_u$ is given in $\Omega_4$ which accounts for the misspecification the conditional quantile condition. Therefore, we have
$$\begin{array}{rl}
\frac{\|\sqrt{f_u}X_{-a}'\delta_u\|_{n,\varpi}^2}{4} \wedge \bar q_{A_u}\|\sqrt{f_u}X_{-a}'\delta_u\|_{n,\varpi} & \leqslant (\bar f^{1/2}+1) \|r_{u}\|_{n,\varpi} \|\sqrt{f_u}X_{-a}' \delta_u\|_{n,\varpi} + (\lambda_u + t_3 + K_u)\|\delta_u\|_{1,\varpi} \\
& \leqslant \{ (\bar f^{1/2}+1) \|r_{u}\|_{n,\varpi} +\frac{3\mathbf{c}\sqrt{s}}{\kappa_{u,2\mathbf{c}}}(\lambda_u + t_3 + K_u)\} \|\sqrt{f_u}X_{-a}' \delta_u\|_{n,\varpi}\\
& + (\lambda_u + t_3 + K_u)\frac{2\mathbf{c}}{\lambda_u}\bar R_{u\xi}
\end{array}$$
The result then follows with the same argument under the current assumptions that account for $K_u$.
\end{proof}
\begin{lemma}[PQGM, Event $\Omega_1$]
Under Condition P, we have
$$ {\mathrm{P}}\left( \sup_{u\in \mathcal{U},j\in [d]} \frac{|{\mathbb{E}_n}[K_\varpi(W)(\tau -1\{X_a\leqslant X_{-a}'\beta_u+r_u\})X_{-a, j}]|}{\widehat{\sigma}^{X}_{a\varpi j}} > t \right) \leqslant 8 |V|(\frac{ne}{d_W})^{2d_W} \exp\left( - \left\{\frac{t/(1+\bar{\delta}_n)}{2(1+1/16)}\right\}^2\right)$$
where $t\geqslant 4 \sup_{u\in \mathcal{U}} \{\mathrm{E}[K_\varpi(W)(\tau-1\{X_a\leqslant X_{-a}'\beta_u+r_u\})^2X_{-a, j}^2]\}^{1/2}$ and $\bar{\delta}_n = o(1)$. In particular, the RHS is less than $\xi$ if $t\geqslant 2(1+\bar{\delta}_n)(1+1/16) \sqrt{\log(8|V|\{ne/d_W\}^{2d_W}/\xi)}$.
\end{lemma}
\begin{proof}
Set $\sigma_{a\varpi j}^X := {\mathrm{E}}[K_\varpi(W)X_{-a,j}^2]^{1/2}$. We have that for any $\bar \delta_n \to 0$
\begin{equation} \begin{array}{rl}
{\mathrm{P}}( \lambda_0 \leqslant { \sup_{u\in \mathcal{U}, j\in [d]} }|{\mathbb{E}_n}[K_\varpi(W)(\tau -1\{X_a\leqslant X_{-a}'\beta_u\})X_{-a,j}]|/\widehat{\sigma}^{X}_{a\varpi j}) \\
\leqslant {\mathrm{P}}( \lambda_0 \leqslant (1+\bar \delta_n){ \sup_{u\in \mathcal{U}, j\in [d]} }|{\mathbb{E}_n}[K_\varpi(W)(\tau -1\{X_a\leqslant X_{-a}'\beta_u\})X_{-a,j}]|/ {\sigma}^{X}_{a\varpi j} )\\
+ {\mathrm{P}}( { \sup_{u \in \mathcal{U},j\in [d]} } {\sigma}^{X}_{a\varpi j}/\widehat{\sigma}^{X}_{a\varpi j} \geqslant (1+\bar \delta_n) ) \end{array} \end{equation}
To bound the last term in ((ref)), note that under Condition P, $ c\mu_\mathcal{W} \leqslant ({\sigma}^{X}_{a\varpi j})^2 \leqslant C$ and $\mathcal{W}$ is a VC class of set with VC dimension $d_W$. Therefore, by Lemma (ref) we have that with probability $1-o(1)$
\begin{equation} \sup_{u \in \mathcal{U}, j\in [d]} ({\mathbb{E}_n}-{\mathrm{E}})[K_\varpi(W)X_{-a,j}^2] \lesssim \sqrt{\frac{(1+d_W)\log(|V|M_n/\sigma_1)}{n}} + \frac{(1+d_W)M_n^2\log(|V|M_n/\sigma_1)}{n} \end{equation}
for ${\sigma}^2_{1} = \max_{u \in \mathcal{U}, j\in [d]}{\mathrm{E}}[K_\varpi(W)X_{-a,j}^2] \leqslant \max_{j\in V} {\mathrm{E}}[X_{-a,j}^2] \leqslant C$ and envelope $F=\|X\|_\infty^2$ so that $\|F\|_{P,2}\leqslant \|\max_{i\leqslant n}F_i\|_{P,2}\leqslant M_n^2$. Thus for $\bar \delta_n \to 0$, provided $(1+d_W)M_n^2\log(|V|n) = o(n^{1/2})$ and $ (1+d_W)\log(|V|n)=o(n\bar \delta_n^2 \mu_\mathcal{W}^2)$, so that the RHS of ((ref)) is $o(\bar\delta_n\mu_\mathcal{W})$, we have \begin{equation}{\mathrm{P}}\left( \frac{1}{1+\bar\delta_n} \leqslant \frac{\sigma_{a\varpi j}^X}{\widehat{\sigma}_{a\varpi j}^X } \leqslant (1+\bar\delta_n), \ \mbox{for all} \ u\in \mathcal{U}, j\in [d]\right) = 1-o(1). \end{equation}
Now we bound the first term of the RHS of ((ref)). and let $\sigma_2^2 = \sup_{u\in\mathcal{U},j\in [d]} {\rm Var}(K_\varpi(W)(\tau -1\{X_a\leqslant X_{-a}'\beta_u\})X_{-a,j}/\sigma_{a\varpi j}^X)\leqslant 1$. By symmetrization (adapting Lemma 2.3.7 in vdVaartWellner2007 to replace the “arbitrary" factor $2$ with $1+\bar \delta_n$), for $\delta := 1/(2(1+n\sigma_2^2/t^2))<1/2$ we have
$$ \begin{array}{rl}
(*):= {\mathrm{P}}( \sup_{u\in \mathcal{U}, j\in [d]} |\sum_{i=1}^n K_\varpi(W_i)(\tau -1\{X_{ia}\leqslant X_{i,-a}'\beta_u\})X_{i,-aj}/\sigma^{X}_{a\varpi j}| \geqslant t) \\ \leqslant 2 {\mathrm{P}}( \sup_{u\in \mathcal{U}, j\in [d]} |\sum_{i=1}^n\varepsilon_i K_\varpi(W_i)(\tau -1\{X_{ia}\leqslant X_{i,-a}'\beta_u\})X_{i,-aj}/\sigma^{X}_{a\varpi j}|\geqslant t\delta )\\
\leqslant 2 {\mathrm{P}}\left( {\displaystyle \sup_{u\in \mathcal{U}, j\in [d]}} \frac{|\sum_{i=1}^n\varepsilon_i K_\varpi(W_i)(\tau -1\{X_{ia}\leqslant X_{i,-a}'\beta_u\})X_{i,-aj}|}{\widehat{\sigma}^{X}_{a\varpi j} }\geqslant t\delta/(1+\bar\delta_n) \right) + o(1)
\end{array}$$
where $\varepsilon_i, i=1,\ldots,n$, are Rademacher random variables independent of the data, and the last inequality follows from ((ref)).
Therefore, by the union bound and symmetry, and iterated expectations we have
$$ (*)\leqslant 4 |V| \max_{j\in [d]} {\mathrm{E}}_{W,X}\left[ {\mathrm{P}}_{\varepsilon}\left( {\displaystyle \sup_{u\in \mathcal{U}}} \frac{\left|\sum_{i=1}^n\varepsilon_i K_\varpi(W_i)(\tau -1\{X_{ia}\leqslant X_{i,-a}'\beta_u\})X_{i,-aj}\right|}{\widehat{\sigma}^{X}_{a\varpi j} }\geqslant t\delta/(1+\bar{\delta}_n) \mid W,X \right) \right] $$
Next we use that $\{ \varpi \in \mathcal{W}\}$ is a VC class of sets with VC dimension bounded by $d_W$ and $\{ 1\{X_{a}\leqslant X_{-a}'\beta_u\} : (\tau,\varpi) \in \mathcal{T}\times \mathcal{W}\}$ is a VC class of sets with VC dimension bounded by $1+d_W$. By Corollary 2.6.3 in vdV-W, we have that conditionally on $(W_i,X_i)_{i=1}^n$, the set of (binary) sequences $\{ (K_\varpi(W_i))_{i=1,\ldots,n} : \varpi \in \mathcal{W}\}$ has at most $\sum_{j=0}^{d_W-1}\binom{n}{j}$ different values. Similarly, $\{ (1\{X_{ia}\leqslant X_{i,-a}'\beta_u\})_{i=1,\ldots,n} : u \in \mathcal{U}\}$ assumes at most $\sum_{j=0}^{d_W}\binom{n}{j}$ different values. Assuming that $n \geqslant d_W$, we have $\sum_{j=0}^{k}\binom{n}{j} \leqslant \{ne/k\}^{k}$ and
$$\begin{array}{l}
{\mathrm{P}}_\varepsilon\left( \sup_{u\in \mathcal{U}}\frac{\left|\sum_{i=1}^n\varepsilon_i K_\varpi(W_i)(\tau -1\{X_{ia}\leqslant X_{i,-a}'\beta_u\})X_{i,-aj}\right|}{\widehat{\sigma}^{X}_{a\varpi j} }\geqslant t\delta/(1+\bar{\delta}_n) \mid W,X \right) \\
\leqslant \{\frac{ne}{d_W-1}\}^{d_W-1}\{\frac{ne}{d_W}\}^{d_W}\sup_{u \in \mathcal{U}} {\mathrm{P}}_\varepsilon\left( \sup_{\tilde \tau \in \mathcal{T}} \frac{\left|\sum_{i=1}^n\varepsilon_i K_\varpi(W_i)(\tilde \tau -1\{X_{ia}\leqslant X_{i,-a}'\beta_u\})X_{i,-aj}\right|}{\widehat{\sigma}^{X}_{a\varpi j} }\geqslant t\delta/(1+\bar{\delta}_n) \mid W,X \right) \\
\leqslant \{ne/d_W\}^{2d_W}\sup_{u\in \mathcal{U}, \tilde \tau \in [\underline{\tau},\bar\tau]} {\mathrm{P}}_\varepsilon\left( \frac{\left|\sum_{i=1}^n\varepsilon_i K_\varpi(W_i)(\tilde \tau -1\{X_{ia}\leqslant X_{i,-a}'\beta_u\})X_{i,-aj}\right|}{\widehat{\sigma}^{X}_{a\varpi j} }\geqslant t\delta/(1+\bar{\delta}_n) \mid W,X \right) \\
\leqslant 2\{ne/d_W\}^{2d_W} \exp( - \{t\delta/[1+\bar{\delta}_n]\}^2)
\end{array}
$$
here we used that the expression is linear in $\tau$ and so it is maximized at the extremes. Combining the bounds in the last two displayed equations we have
$$ (*) \leqslant 8|V| \{ne/d_W\}^{2d_W} \exp( - \{t\delta/[1+\bar\delta_n]\}^2).$$
Therefore, setting $\lambda_0 = ct/n$ where $t\geqslant 4 \sqrt{n}\sigma_2$ and $t\geqslant 2(1+\bar\delta_n)(1+1/16) \sqrt{2\log(8p|V|\{ne/d_W\}^{2d_W}/\xi)}$. (Note that $t\geqslant 4 \sqrt{n}\sigma_2$ implies that $\delta \geqslant 1/\{2(1+1/16)\}$.)
\end{proof}
\begin{lemma}[PQGM, Event $\Omega_2$]
Under Condition P we have
$${\mathrm{P}}\left( \sup_{u\in\mathcal{U}}\widehat R_u(\bar \beta_u) \leqslant C\{1+ \bar f\} \{n^{-1}s(1+d_W)\log(|V|n)\} \right) = 1-o(1).$$
\end{lemma}
\begin{proof}[Proof of Lemma (ref)]
We have that $\widehat R_u(\bar \beta_u)\geqslant 0$ by convexity of $\rho_\tau$. Let $\varepsilon_{iu}=X_{ia} - X_{i,-a}\bar \beta_u - r_{iu}$ where $\|\bar\beta_u\|_0\leqslant s$ and $r_{iu}=X_{-a}'(\beta_u-\bar\beta_u)$. By Knight's identity ((ref)), $\widehat R_u(\bar\beta_u) = -{\mathbb{E}_n}[ K_\varpi(W)r_{u} \int_0^1 1\{\varepsilon_{u} \leqslant -t r_{u}\} - 1\{\varepsilon_{u}\leqslant 0\} \ dt ]\geqslant 0$.
$$ \begin{array}{rl}
{\mathbb{E}_n}[{\mathrm{E}}[\widehat R_u(\bar\beta_u)] & = {\mathbb{E}_n}[ K_\varpi(W)r_{u} \int_0^1 F_{X_a\mid X_{-a},\varpi}(X_{-a}'\bar\beta_u+(1-t)r_u) - F_{X_a\mid X_{-a},\varpi}(X_{-a}'\bar\beta_u+r_u) \ dt]\\
& \leqslant {\mathbb{E}_n}[K_\varpi(W)r_{u} \int_0^1 \bar f t r_{u} dt] \leqslant \bar f [\|r_{u}\|_{n,\varpi}^2]/2 \leqslant C\bar f s/n.\end{array}$$
Thus, by Markov's inequality we have $\inf_{u\in \mathcal{U}} \mathrm{P}( \widehat R_u(\bar \beta_u) \leqslant C\bar f s/n ) \geqslant 1/2$.
Define $z_{iu} := -\int_0^1 1\{\varepsilon_{iu} \leqslant -t r_{iu}\} - 1\{\varepsilon_{iu}\leqslant 0\} \ dt$, so that $\widehat R_u(\bar\beta_u) = {\mathbb{E}_n}[K_\varpi(W)r_{u}z_{u}]$ with $|z_{iu}|\leqslant 1$. By Lemma 2.3.7 in vdVaartWellner2007 (note that the Lemma does not require zero mean stochastic processes), for $t \geqslant 2 C\bar f s/n$ we have
$$ \frac{1}{2}\mathrm{P}\left( \sup_{u\in\mathcal{U}}|{\mathbb{E}_n}[K_\varpi(W)r_{u}z_{u}]| \geqslant t \right) \leqslant 2\mathrm{P}\left(\sup_{u\in\mathcal{U}}|{\mathbb{E}_n}[\varepsilon K_\varpi(W)r_{u}z_{u}]|>t/4\right) $$
where $\varepsilon_i, i=1,\ldots,n,$ are Rademacher random variables independent of the data.
Consider the class of functions $\mathcal{F}= \{ - K_\varpi(W)r_{u} (1\{\varepsilon_{iu} \leqslant -B_i r_{iu}\} - 1\{\varepsilon_{iu}\leqslant 0\} ) : u\in\mathcal{U}\}$ where $B_i\sim {\rm Uniform}(0,1)$ independent of $(X_i,W_i)_{i=1}^n$. It follows that $K_\varpi(W)r_{u}z_{u} = {\mathrm{E}}[ - K_\varpi(W)r_{u} (1\{\varepsilon_{iu} \leqslant -B_i r_{iu}\} - 1\{\varepsilon_{iu}\leqslant 0\} ) \vert X_i,W_i]$ where the expectation is taken over $B_i$ only. Thus we will bound the entropy of $\overline{\mathcal{F}}= \{ {\mathrm{E}}[ f\vert X,W] : f \in \mathcal{F} \}$ via Lemma (ref). Note that $\mathcal{R}:=\{ r_u = X_{-a}'\beta_u - X_{-a}'\bar\beta_u : u\in\mathcal{U}\}$ where $\mathcal{G}:=\{ X_{-a}'\bar\beta_u : u\in\mathcal{U}\}$ is contained in the union of at most $|V|\binom{p}{s}$ VC-classes of dimension $Cs$ and $\mathcal{H}:=\{X_{-a}'\beta_u : u\in\mathcal{U}\}\}$ is a VC-class of functions of dimension $(1+d_W)$ by Condition P. Finally note that $\mathcal{E}:=\{ \varepsilon_{iu} : u\in \mathcal{U}\} \subset \{ X_{ia} : a\in V\} - \mathcal{G} - \mathcal{R}$.
Therefore, we have
$$
\begin{array}{rl}
\sup_Q \log N(\epsilon \|\bar F\|_{Q,2}, \overline{\mathcal{F}}, \|\cdot\|_{Q,2}) & \leqslant \sup_Q \log N( (\epsilon/4)^2 \| F\|_{Q,2}, \mathcal{F}, \|\cdot\|_{Q,2})\\
& \leqslant \sup_Q \log N( \mbox{$\frac{1}{8}$}(\epsilon^2/16), \mathcal{W}, \|\cdot\|_{Q,2}) \\
&+ \sup_Q\log N( \mbox{$\frac{1}{8}$} (\epsilon^2/16) \| F\|_{Q,2}, \mathcal{R}, \|\cdot\|_{Q,2}) \\
&+ \sup_Q \log N( \mbox{$\frac{1}{8}$}(\epsilon^2/16), 1\{ \mathcal{E}+\{B\}\mathcal{R} \leqslant 0\} - 1\{ \mathcal{E} \leqslant 0\}, \|\cdot\|_{Q,2})\\
\end{array}
$$
By Lemma (ref) with envelope $\bar F = \|X\|_\infty\sup_{u\in\mathcal{U}}\|\beta_u-\bar\beta_u\|_1$, and $(\sigma^{r}_{\max})^2 = \sup_{u\in \mathcal{U}} {\mathrm{E}}[K_\varpi(W)r_{u}^2 ] \lesssim s/n$ by Condition P, we have that with probability $1-o(1)$
$$ \sup_{u\in\mathcal{U}}|{\mathbb{E}_n}[\varepsilon K_\varpi(W)r_{u}z_{u}]| \lesssim \sqrt{\frac{s(1+d_W)\log(|V|n)}{n}}\sqrt{\frac{s}{n}}+ \frac{M_n\sqrt{s^2/n}\log(|V|n)}{n}\lesssim \frac{s(1+d_W)\log(|V|n)}{n}$$
under $M_n\sqrt{s^2/n} \leqslant C$.
\end{proof}
\begin{lemma}[PQGM, Event $\Omega_3$]
Under Condition P, for $u=(a,\tau,\varpi)\in V\times \mathcal{T}\times \mathcal{W}$, define $g_u(\delta,X,W)= K_\varpi(W)\{\rho_\tau(X_a-X_{-a}'(\beta_u+\delta))-\rho_\tau(X_a-X_{-a}'\beta_u)\}$, and
$$ \Omega_3 := \left\{ \sup_{u\in\mathcal{U}, 1/\sqrt{n} \leqslant \|\delta\|_{1,\varpi} \leqslant \sqrt{n}} \frac{|{\mathbb{E}_n}[g_u(\delta,X,W)-{\mathrm{E}}[g_u(\delta,X,W)\vert X_{-a}, W]]|}{\|\delta\|_{1,\varpi}} < t_3 \right\}.$$
Then, under Condition CI we have $\mathrm{P}(\Omega_3) \geqslant 1-\xi$ for any
$$ t_3\sqrt{n} \geqslant 12 + 16\sqrt{2\log\left(64|V|^2 n^{1+d_W}\log(n)(L_{\beta}M_n\sqrt{n})^{1+d_W/\kappa} /\xi \right)} $$
\end{lemma}
\begin{proof}
We have that $\Omega_3^c := \{ \max_{a\in V} A_a \geqslant t_3\sqrt{n}\}$ for
$$A_a:= \sup_{(\tau,\varpi)\in\mathcal{T}\times\mathcal{W}, \underline{N} \leqslant \|\delta\|_{1,\varpi} \leqslant \bar{N}} \sqrt{n}\left|\frac{{\mathbb{E}_n}[g_u(\delta,X,W)-{\mathrm{E}}[g_u(\delta,X,W)\vert X_{-a},W]]}{\|\delta\|_{1,\varpi}}\right|. $$
We will apply Lemma (ref) with $\rho=\kappa$, $L_\eta=L_\beta$, $\tilde x = X_{-a}$ (so we take $p=|V|$), $\underline{N} = 1/\sqrt{n}$ and $\bar{N} = \sqrt{n}$. Therefore, we have by Lemma (ref) and the union bound
$$\begin{array}{rl}
\mathrm{P}(\Omega_3^c) & = \mathrm{P}(\max_{a\in V} A_a \geqslant t_3\sqrt{n}) \\
& \leqslant |V|\max_{a\in V} \mathrm{P}(A_a \geqslant t_3\sqrt{n}) \\
& = |V|\max_{a\in V} {\mathrm{E}}_{X_{-a},W}\left\{ \mathrm{P}(A_a \geqslant t_3\sqrt{n} \mid X_{-a},W )\right\} \\
& \leqslant |V| \max_{a\in V} {\mathrm{E}}_{X_{-a},W}\left\{8|V| \ |\widehat{\mathcal{N}}|\cdot |\widehat{\mathcal{W}}|\cdot |\widehat{\mathcal{T}}| \exp(-(t_3\sqrt{n}/4-3)^2/32)\right\}\\
& \leqslant C |V|^2 n^{1+d_W}\log(n)L_\beta^{1+d_W/\kappa} {\mathrm{E}}_{X_{-a}}\left\{\frac{\max_{i\leqslant n}\|X_{i,-a}\|_\infty^{1+d_W/\kappa}}{\underline{N}^{1+d_W/\kappa}}\right\}\exp(-(t_3\sqrt{n}/4-3)^2/32) \\
& \leqslant \xi \end{array} $$
where the last step follows by the choice of $t_3$.
\end{proof}
\begin{lemma}[PQGM, Event $\Omega_4$]
Under Condition P, and setting $K_u =C \sqrt{\frac{(1+d_W) \log(|V|n)}{n}}$, we have that $ {\mathrm{P}}( \Omega_4 ) = 1-o(1)$.
\end{lemma}
\begin{proof}
First note that by Lemma (ref) we have that with probability $1-o(1)$
$$ \sup_{\varpi \in \mathcal{W},j\in V} ({\mathbb{E}_n}-{\mathrm{E}})[K_\varpi(W)X_{-a,j}^2] \lesssim \sqrt{\frac{(1+d_W)\log(|V|M_n/\sigma)}{n}} + \frac{(1+d_W)M_n^2\log(|V|M_n/\sigma)}{n} $$
and $\sigma^{X}_{a\varpi j}\geqslant c {\mathrm{P}}(\varpi)$. Under $(1+d_W)\log(|V|M_n/\sigma)\leqslant \delta_n^2 \mu^2_{\mathcal{W}}$ and $(1+d_W)M_n^2\log(|V|M_n/\sigma)\leqslant \delta_n n \mu_{\mathcal{W}}$, we have that
$|({\mathbb{E}_n}-{\mathrm{E}})[K_\varpi(W)X_{-a,j}^2]| = o( \sigma^X_{a\varpi j})$ for all $u\in \mathcal{U}$. Therefore, we have
$$
\begin{array}{rl}
& {\mathrm{P}}( \sup_{u \in \mathcal{U}} |{\mathbb{E}_n}[h_{uj}(X_{-a},W)]|/\{K_u \widehat{\sigma}^{X}_{a\varpi j}\} > 1 ) \\
& \leqslant {\mathrm{P}}( \sup_{u \in \mathcal{U}} |{\mathbb{E}_n}[h_{uj}(X_{-a},W)]|/\{K_u \sigma^{X}_{a\varpi j}\} > 1+ O(\delta_n)) + o(1)\\
\end{array}
$$
Applying Lemma (ref) to $\mathcal{F}=\{ h_{uj}(X_{-a},W)/\sigma^{X}_{a\varpi j}: u \in \mathcal{U}\}$. For convenience define $\bar{\mathcal{H}}_j= \{ h_{uj}(X_{-a},W) : u \in \mathcal{U}\}$ and $\bar{\mathcal{K}}_j:=\{ {\mathrm{E}}[K_\varpi(W)X_j^2] : \varpi \in \mathcal{W}\}$. Note that $\bar{\mathcal{K}}_j$ has covering numbers bounded by the covering number of $\mathcal{K}_j:=\{ K_\varpi(W)X_j^2 : \varpi \in \mathcal{W}\}$ hence $\sup_Q \log N( \epsilon \|\bar K_j\|_{Q,2}, \bar{\mathcal{K}}_j, \|\cdot\|_{Q,2}) \leqslant \log \sup_{\tilde Q} N((\epsilon/4)^2\|F\|_{\tilde Q,2}, \mathcal{K}_j,\|\cdot\|_{\tilde Q,2})$ by Lemma (ref). Similarly, Lemma (ref) also allows us to bound covering numbers of $\bar{\mathcal{H}}_j$ via covering numbers of $\mathcal{H}_j= \{ K_\varpi(W)(\tau-1\{X_a\leqslant X_{-a}'\beta_u\})X_j : u \in \mathcal{U}\}$.
$$
\begin{array}{l}
\sup_Q \log N( \epsilon \|F\|_{Q,2}, \mathcal{F}, \|\cdot\|_{Q,2}) \leqslant p \max_{j\in[p]}\sup_Q \log N( \epsilon \|F_j\|_{Q,2}, \mathcal{F}_j, \|\cdot\|_{Q,2})\\
\leqslant p \max_{j\in[p]}\sup_Q \{ \log N( (1/2)\epsilon \|\bar H_j\|_{Q,2}, \bar{\mathcal{H}}_j, \|\cdot\|_{Q,2}) + \sup_Q \log N( (1/2)\epsilon c \mu_{\mathcal{W}}^{1/2}, 1/\bar{\mathcal{K}}_j^{1/2}, \|\cdot\|_{Q,2})\}\\
\leqslant p \max_{j\in[p]}\sup_Q \{ \log N( (1/2)\epsilon \|\bar H_j\|_{Q,2}, \bar{\mathcal{H}}_j, \|\cdot\|_{Q,2}) + \sup_Q \log N( (1/2)\epsilon c \mu_{\mathcal{W}}^{3/2}, \bar{\mathcal{K}}_j^{1/2}, \|\cdot\|_{Q,2})\}\\
\leqslant p \max_{j\in[p]}\sup_Q \{ \log N( (1/2)\epsilon \|\bar H_j\|_{Q,2}, \bar{\mathcal{H}}_j, \|\cdot\|_{Q,2}) + \sup_Q \log N( C(1/2)\epsilon c\mu_{\mathcal{W}}^{3/2}, \bar{\mathcal{K}}_j, \|\cdot\|_{Q,2})\}\\
\leqslant p \max_{j\in[p]}\sup_Q \{ \log N( (1/2)\epsilon \|\bar H_j\|_{Q,2}, \bar {\mathcal{H}}_j, \|\cdot\|_{Q,2}) + \sup_Q \log N( 1/(2C)\epsilon c\mu_{\mathcal{W}}^{3/2}, \bar{\mathcal{K}}_j, \|\cdot\|_{Q,2})\}\\
\leqslant p \max_{j\in[p]}\sup_{\tilde Q} \log N( (1/4)\epsilon^2 \|H_j\|_{Q,2}, \mathcal{H}_j, \|\cdot\|_{Q,2}) \\
+ p \max_{j\in[p]}\sup_{\tilde Q} \log N( (1/(4C^2)\epsilon^2 c^2\mu_{\mathcal{W}}^{3}, \mathcal{K}_j, \|\cdot\|_{\tilde Q,2}) \\ \end{array}
$$
where $F_j = c\|X\|_\infty/\mu_{\mathcal{W}}^{1/2}$, $H_j=\|X\|_\infty$. Since $\mathcal{K}_j$ is the product of a VC subgraph of dimension $d_W$ with a single function, and $\mathcal{H}_j$ is the product of two VC subgraph of dimension $1+d_W$ and a single function, by Lemma (ref) with $\sigma^2 = 1$, we have with probability $1-o(1)$
$$\sup_{u\in \mathcal{U}} \left| \frac{({\mathbb{E}_n}-{\mathrm{E}})[h_{uj}(X_{-a},W)]}{{\mathrm{E}}[K_\varpi(W)X_j^2]^{1/2}}\right| \leqslant C \sqrt{\frac{(1+d_W) \log(|V|n)}{n}} + C \frac{M_n (1+d_W)\log(|V|n)}{n\mu_{\mathcal{W}}^{1/2}}. $$
Thus, under $M_n(1+d_W)\log(|V|n) \leqslant n^{1/2}\mu_{\mathcal{W}}^{1/2}$ we have that we can take $K_u = C \sqrt{\frac{(1+d_W) \log(|V|n)}{n}}$.
\end{proof}
\section{Technical Results for High-Dimensional Quantile Regression}
In this section we provide technical results for high-dimensional quantile regression. It is based on a sample $(\tilde y_i, \tilde x_i, W_i)_{i=1}^n$, independent across $i$, $\rho_\tau(t) = (\tau-1\{t \leqslant 0\})t$, $\tau \in \mathcal{T} \subset (0,1)$ a compact interval, and a family of indicator functions $K_w(W)=1$ if $W \in \Omega_\varpi$, $K_w(W)=0$ otherwise, here $\Omega_\varpi\in\mathcal{W}$. For convenience we index the sets $\Omega_\varpi$ by $\varpi \in B_W \subset {\mathbb{R}}^{d_W}$ where we normalize the diameter of $B_W$ to be less or equal than 1/6. Let $f_{\tilde y|\tilde x,r_u,\varpi}(\cdot)$ denote the conditional density function, $f_{\tilde y|\tilde x,r_u,\varpi}(\cdot)\leqslant \bar f$, $|f_{\tilde y|\tilde x,r_u,\varpi}'(\cdot)|\leqslant \bar f'$ and $f_u:=f_{\tilde y|\tilde x,r_u,\varpi}(\tilde x'\eta_u)$. Moreover, we assume that
\begin{equation}
\|\eta_{u}-\eta_{\tilde u}\|_1 \leqslant L_\eta\{|\tau-\tilde \tau|+ \|\varpi-\tilde\varpi\|^\rho \}.\end{equation}
Although the results can be applied more generally, these results will be used for $(\eta_u,r_u), u=(\tilde y, \tau,\varpi) \in \mathcal{U} :=\{\tilde y\}\times \mathcal{T}\times \mathcal{W}$ satisfying
$${\mathrm{E}}[ K_\varpi(W)(\tau-1\{\tilde y \leqslant \tilde x'\eta_u + r_u \})\tilde x] = 0.$$
Note that this generality is flexible enough to allow us to cover the case that the $\tau$-conditional quantile function $Q_{\tilde y}(\tau\vert \tilde x, \varpi ) = \tilde x'\tilde \eta_u+\tilde r_{u}$ by setting $\eta_u = \tilde\eta_u$ and $r_u=\tilde r_u$ in which case ${\mathrm{E}}[(\tau-1\{\tilde y \leqslant \tilde x'\eta_u + r_u \})\vert \tilde x, \varpi] = 0$. It also covers the case that
$$ \tilde \eta_u \in \arg\min_\beta {\mathrm{E}}[ K_\varpi(W) \rho_\tau(\tilde y - \tilde x'\beta)] $$
so that ${\mathrm{E}}[ K_\varpi(W)(\tau-1\{\tilde y \leqslant \tilde x'\tilde \eta_u \})\tilde x]=0$ holds by the first order condition by setting $\eta_u=\tilde \eta_u$ and $r_u=0$. Moreover, it also covers the case that we work with a sparse approximation $\bar \eta_u$ of $\tilde \eta_u$ by setting $\eta_u=\bar\eta_u$ and $r_u=\tilde x'(\tilde\eta_u-\bar\eta_u)$.
\begin{lemma}[Identification Lemma]
For $u=(a,\tau,\varpi)\in\mathcal{U}$, and a subset $A_u\subset {\mathbb{R}}^p$ let
\begin{equation} \bar q_{A_u} = 1/(2\bar{f'}) \cdot \inf_{ \delta \in A_u} {\mathbb{E}_n}\[K_\varpi(W)f_u|\tilde x'\delta|^2\right]^{3/2}/{\mathbb{E}_n}\[K_\varpi(W)|\tilde x'\delta|^3\right]\end{equation}
and assume that for all $\delta \in A_u$
\begin{equation}{\mathbb{E}_n}\left[ K_\varpi(W)|r_{u}|\cdot |\tilde x'\delta|^2\right] + {\mathbb{E}_n}\left[ K_\varpi(W)r_{u}^2\cdot |\tilde x'\delta|^2\right] \leqslant 1/(4\bar f'){\mathbb{E}_n}[K_\varpi(W)f_u|\tilde x'\delta|^2].\end{equation}
Then we have
$$\begin{array}{l} {\mathbb{E}_n}[{\mathrm{E}}[K_\varpi(W)\rho_\tau(\tilde y-\tilde x'(\eta_u + \delta))\mid \tilde x,r_u,W ]] - {\mathbb{E}_n}[{\mathrm{E}}[K_\varpi(W)\rho_\tau(\tilde y-\tilde x'\eta_u)\mid \tilde x,r_u,W]] \\
\geqslant \frac{\|\sqrt{f_u}\tilde x'\delta\|_{n,\varpi}^2}{4} \wedge \left\{ \bar q_{A_u}\|\sqrt{f_u}\tilde x'\delta\|_{n,\varpi}\right\} - K_{n2}\|\sqrt{f_u}\tilde x'\delta\|_{n,\varpi}-K_{n1}\|\delta\|_{1,\varpi}.\end{array}$$
where $K_{n2}:=(\bar f^{1/2}+1)\|r_{u}\|_{n,\varpi}$ and $K_{n1} := \sup_{u\in\mathcal{U}, j\in[p]} \frac{|{\mathbb{E}_n}[{\mathrm{E}}[K_\varpi(W)(\tau-1\{\tilde y \leqslant \tilde x'\eta_u+r_u\})\tilde x_j \mid \tilde x, W]]|}{\{{\mathbb{E}_n}[K_\varpi(W)\tilde x_j^2]\}^{1/2}}$.\end{lemma}
\begin{proof}[Proof of Lemma (ref)] Let $T_u = {\rm support}(\eta_u)$, and $Q_u(\eta):={\mathbb{E}_n}{\mathrm{E}}[K_\varpi(W)\rho_\tau(\tilde y-\tilde x'\eta)\mid \tilde x,r_u,W]$. The proof proceeds in steps.
Step 1. (Minoration) Define the maximal radius over which the criterion function can be minorated by a quadratic function
$$ r_{A_u} = \sup_{r} \left\{\begin{array}{rl}
r \ : & Q_u(\eta_u+ \delta) - Q_u(\eta_u) + K_{n2}\|\sqrt{f_u}\tilde x'\delta\|_{n,\varpi}+ K_{n1}\|\delta\|_{1,\varpi} \geqslant \frac{1}{4} \|\sqrt{f_u}\tilde x' \delta\|^{2}_{n,\varpi}, \\
& \forall \delta\in A_u, \ \|\sqrt{f_u}\tilde x' \delta\|_{n,\varpi} \leqslant r \end{array}\right\}.$$
Step 2 below shows that $r_{A_u} \geqslant \bar q_{A_u}$. By construction of $r_{A_u}$ and the convexity of $Q_u(\cdot)$, $\|\cdot \|_{1,\varpi}$ and $\|\cdot \|_{n,\varpi}$,
$$ \begin{array}{lll}
&& Q_u(\eta_u + \delta) - Q_u(\eta_u) +K_{n2}\|\sqrt{f_u}\tilde x'\delta\|_{n,\varpi} + K_{n1}\|\delta\|_{1,\varpi} \geqslant \\
&& \geqslant \frac{\|\sqrt{f_u}\tilde x'\delta\|^2_{n,\varpi}}{4} \wedge \left\{ \frac{\|\sqrt{f_u}\tilde x'\delta\|_{n,\varpi}}{r_{A_u}} \cdot \begin{array}{rl}
\inf_{\tilde\delta} & Q_u(\eta_u+\tilde \delta) - Q_u(\eta_u) +K_{n2}\|\sqrt{f_u}\tilde x'\tilde\delta\|_{n,\varpi} + K_{n1}\|\tilde\delta\|_{1,\varpi}\\
& {\tilde \delta\in A_u, \|\sqrt{f_u}\tilde x' \tilde \delta\|_u \geqslant r_{A_u}} \end{array} \right\}\\
&& \geqslant \frac{\|\sqrt{f_u}\tilde x'\delta\|^2_{n,\varpi}}{4} \wedge \left\{ \frac{\|\sqrt{f_u}\tilde x'\delta\|_{n,\varpi}}{r_{A_u}} \frac{r_{A_u}^2}{4}\right\}
\\
&& \geqslant \frac{\|\sqrt{f_u}\tilde x'\delta\|_{n,\varpi}^2}{4} \wedge \left\{ \bar q_{A_u}\|\sqrt{f_u}\tilde x'\delta\|_{n,\varpi}\right\}.
\end{array}$$
Step 2. ($r_{A_u} \geqslant \bar q_{A_u}$) Let $F_{\tilde y \mid \tilde x,r_u,\varpi}$ denote the conditional distribution of
$\tilde y$ given $\tilde x,r_u,\varpi$. From Knight1998,
for any two scalars $w$ and $v$ the Knight's identity is
\begin{equation}
\rho_\tau(w-v) - \rho_\tau(w) = -v (\tau - 1\{w\leqslant 0\}) + \int_0^v(
1\{w\leqslant z\} - 1\{w\leqslant 0\})dz.
\end{equation}
Using ((ref)) with $w=\tilde y_i - \tilde x_i'\eta_u$ and $v =
\tilde x_i'\delta$ and taking expectations with respect to $\tilde y$, we have
$$\begin{array}{rl}
Q_u(\eta_u + \delta) - Q_u(\eta_u) = & -{\mathbb{E}_n}[{\mathrm{E}}\[K_\varpi(W)(\tau-1\{\tilde y\leqslant \tilde x'\eta_u\})\tilde x_i'\delta \mid \tilde x, r_u, W ]\right] \\
& + {\mathbb{E}_n}\left[ \int_0^{K_\varpi(W)\tilde x'\delta} F_{\tilde y|\tilde x,r_u,\varpi}(\tilde x'\eta_u + t) - F_{\tilde y|\tilde x,r_u,\varpi}(\tilde x'\eta_u) dt \right].\end{array}$$
Using the law of iterated expectations and mean value
expansion, the relation
$$\begin{array}{rl}
& |{\mathbb{E}_n}[{\mathrm{E}}\[K_\varpi(W)(\tau-1\{\tilde y\leqslant \tilde x'\eta_u\})\tilde x'\delta \mid \tilde x, r_u, W \right]] | \\
& = |{\mathbb{E}_n}\[K_\varpi(W)\{F_{\tilde y\mid \tilde x,r_u,\varpi}(\tilde x'\eta_u+r_u)-F_{\tilde y\mid \tilde x,r_u,\varpi}(\tilde x'\eta_u)\}\tilde x'\delta \right] \\
& + {\mathbb{E}_n}\[K_\varpi(W)\{\tau -F_{\tilde y\mid \tilde x,r_u,\varpi}(\tilde x'\eta_u+r_u)\}\tilde x'\delta \mid \tilde x, r_u, W \right] |\\
& \leqslant {\mathbb{E}_n}[K_\varpi(W)f_u|r_u| \ |\tilde x'\delta|]+\bar f'{\mathbb{E}_n}[K_\varpi(W)|r_u|^2 |\tilde x'\delta|] + K_{n1}\|\delta\|_{1,\varpi}\\
& \leqslant \|\sqrt{f_u}r_u\|_{n,\varpi}\|\sqrt{f_u}\tilde x'\delta\|_{n,\varpi} + \bar f'\|r_u\|_{n,\varpi}\|r_u \tilde x'\delta\|_{n,\varpi}+ K_{n1}\|\delta\|_{1,\varpi}\\
& \leqslant (\bar f^{1/2}+1)\|r_u\|_{n,\varpi}\|\sqrt{f_u}\tilde x'\delta\|_{n,\varpi}+ K_{n1}\|\delta\|_{1,\varpi} \end{array}$$ where we used our assumption on the approximation error and we have $K_{n2}=(\bar f^{1/2}+1)\|r_u\|_{n,\varpi}$. With that and similar arguments we obtain for $\tilde t_{\tilde x_i,t} \in [0,t]$
\begin{equation}
\begin{array}{rcl}
&& Q_u(\eta_u + \delta) - Q_u(\eta_u) + K_{n2}\|\sqrt{f_u}\tilde x'\delta\|_{n,\varpi}+K_{n1}\|\delta\|_{1,\varpi} \geqslant \\
&& Q_u(\eta_u + \delta) - Q_u(\eta_u) +{\mathbb{E}_n}[{\mathrm{E}}\[K_\varpi(W)(\tau-1\{\tilde y\leqslant \tilde x'\eta_u\})\tilde x'\delta \mid \tilde x, r_u, W ]\right] = \\
&& = {\mathbb{E}_n}\left[ \int_0^{K_\varpi(W)\tilde x'\delta} F_{\tilde y|\tilde x,r_u,\varpi}(\tilde x'\eta_u + t) - F_{\tilde y|\tilde x,r_u,\varpi}(\tilde x'\eta_u) dt \right] \\
&& = {\mathbb{E}_n}\left[ \int_0^{K_\varpi(W)\tilde x'\delta} tf_{\tilde y|\tilde x,r_u,\varpi}(\tilde x'\eta_u) + \frac{t^2}{2}f'_{\tilde y|\tilde x,r_u,\varpi}(\tilde x'\eta_u+\tilde t_{\tilde x,t}) dt \right] \\
&& \geqslant \frac{1}{2}\|\sqrt{f_u}\tilde x'\delta\|_{n,\varpi}^2 - \frac{1}{6}\bar f ' {\mathbb{E}_n}[K_\varpi(W)|\tilde x'\delta|^3] - {\mathbb{E}_n}\left[ \int_0^{K_\varpi(W)\tilde x_i'\delta} t[f_{\tilde y\mid \tilde x,r_u,\varpi}(\tilde x'\eta_u)-f_{\tilde y\mid \tilde x,r_u,\varpi}(\tilde x'\eta_u + r_u)]dt\right]\\
&&\geqslant \frac{1}{4}\|\sqrt{f_u}\tilde x'\delta\|_{n,\varpi}^2 + \frac{1}{4}\|\sqrt{f_u}\tilde x'\delta\|_{n,\varpi}^2 - \frac{1}{6} \bar f' {\mathbb{E}_n}[K_\varpi(W)|\tilde x'\delta|^3]-(\bar f'/2) {\mathbb{E}_n}\[K_\varpi(W) |\tilde r_u|\cdot |\tilde x'\delta|^2\right].\\
\end{array}
\end{equation}
Moreover, by assumption we have
\begin{equation}\begin{array}{rl}
{\mathbb{E}_n}\left[ K_\varpi(W) |r_{u}|\cdot |\tilde x'\delta|^2\right] \leqslant \frac{1}{4\bar f'}{\mathbb{E}_n}[K_\varpi(W)f_u|\tilde x'\delta|^2] \\
\end{array} \end{equation}
Note that for any $\delta$ such that $ \|\sqrt{f_u}\tilde x'\delta\|_{n,\varpi} \leqslant \bar q_{A_u}$ we have $$ \|\sqrt{f_u}\tilde x'\delta\|_{n,\varpi}\leqslant \bar q_{A_u} \leqslant
1/(2\bar{f'}) \cdot {\mathbb{E}_n}\[K_\varpi(W)f_u|\tilde x'\delta|^2\right]^{3/2}/{\mathbb{E}_n}\[K_\varpi(W)|\tilde x'\delta|^3\right].$$
It follows that $(1/6)\bar f'{\mathbb{E}_n}[K_\varpi(W)|\tilde x'\delta|^3] \leqslant
(1/8) {\mathbb{E}_n}[K_\varpi(W)f_u|\tilde x'\delta|^2]$. Combining this with ((ref)) we have
\begin{equation}\frac{1}{4}{\mathbb{E}_n}[K_\varpi(W)f_u|\tilde x'\delta|^2] - \frac{\bar f'}{6} {\mathbb{E}_n}[K_\varpi(W)|\tilde x'\delta|^3]-\frac{\bar f'}{2} {\mathbb{E}_n}\[K_\varpi(W) |r_u|\cdot |\tilde x'\delta|^2\right] \geqslant 0.\end{equation}
Combining ((ref)) and ((ref)) we have $r_{A_u} \geqslant \bar q_{A_u}$.
\end{proof}
\begin{lemma} Let $\mathcal{W}$ be a VC-class of sets with VC-index $d_W$. Conditional on $\{(W_i,\tilde x_i), i=1,\ldots,n\}$ we have
{\tiny \begin{eqnarray*}
P_{\tilde y}\left( \sup_{ {\tiny \begin{array}{c}\tau\in\mathcal{T},\varpi\in\mathcal{W}, \\ \underline{N}\leqslant\|\delta\|_{1,\varpi}\leqslant \bar N\end{array}}} \left| \mathbb{G}_n\left( K_\varpi(W)\frac{\rho_\tau(\tilde y-\tilde x'(\eta_u+\delta))-\rho_\tau(\tilde y-\tilde x'\eta_u)}{\|\delta\|_{1,\varpi}}\right) \right| \geqslant M \mid (W_i,\tilde x_i)_{i=1}^n \right) \leqslant S_n\exp(-(M/4-3)^2/32)
\end{eqnarray*}}
where $ S_n \leqslant 8p |\widehat{\mathcal{N}}|\cdot |\widehat{\mathcal{W}}|\cdot |\widehat{\mathcal{T}}|$, with
$$ |\widehat{\mathcal{N}}|\leqslant 1+\left\lfloor 3\sqrt{n}\log(\bar N/ \underline{N}) \right\rfloor, \ \ |\widehat{\mathcal{T}}|\leqslant 2\sqrt{n}\frac{\max_{i\leqslant n}\|\tilde x_i\|_\infty}{\underline{N}}L_\eta, \ \ |\widehat{\mathcal{W}}|\leqslant n^{d_W}+\left\{2\sqrt{n}\frac{\max_{i\leqslant n}\|\tilde x_i\|_\infty}{\underline{N}}L_\eta\right\}^{d_W/\rho}.$$
\end{lemma}
\begin{proof}[Proof of Lemma (ref)]
Let $g_{i\tau\varpi}(b) = K_\varpi(W_i)\{\rho_\tau(\tilde y_i-\tilde x_i'\eta_{\tau\varpi} +b)-\rho_\tau(\tilde y_i-\tilde x_i'\eta_{\tau\varpi})\}\leqslant K_\varpi(W_i)|b|$ since $K_\varpi(W_i)\in\{0,1\}$. Note that $|g_{i\tau\varpi}(b)- g_{i\tau\varpi}(a)|\leqslant K_\varpi(W_i)|b-a|$. To easy the notation we omit the conditioning on $(\tilde x_i,W_i)$ from the probabilities.
For any $\delta\in {\mathbb{R}}^p$, since $\rho_\tau$ is $1$-Lipschitz, we have
$$ \begin{array}{rl}
{\rm Var}\left( \mathbb{G}_n\left( \frac{g_{\tau\varpi}(\tilde x'\delta)}{\|\delta\|_{1,\varpi}}\right) \right) & \leqslant \frac{{\mathbb{E}_n}[\{g_{\tau\varpi}(\tilde x'\delta)\}^2] }{\|\tilde x'\delta\|_{n,\varpi}^2} \leqslant \frac{{\mathbb{E}_n}[|K_\varpi(W)\tilde x'\delta|^2] }{\|\tilde x'\delta\|_{n,\varpi}^2} = 1 \end{array}$$
since by definition $\|\delta\|_{1,\varpi}=\sum_j\|\delta_j\|_{1,\varpi}=\sum_j\|\tilde x'_j\delta_j\|_{n,\varpi}\geqslant \|\tilde x'\delta\|_{n,\varpi}$.
Since we are conditioning on $(W_i,\tilde x_i)_{i=1}^n$ the process is independent across $i$. Then, by Lemma 2.3.7 in vdV-W (Symmetrization for Probabilities) we have for any $M>1$
{ $$\begin{array}{rl}\displaystyle {\mathrm{P}}\left( \sup_{ \tau \in\mathcal{T}, \varpi\in\mathcal{W}, \underline{N}\leqslant \|\delta\|_{1,\varpi}\leqslant \bar{N}} \left| \mathbb{G}_n\left( \frac{g_{\tau\varpi}(\tilde x'\delta)}{\|\delta\|_{1,\varpi}}\right) \right| \geqslant M \right) \\
\displaystyle \leqslant \frac{2}{1-M^{-2}}P\left( \sup_{ \tau \in\mathcal{T}, \varpi\in\mathcal{W}, \underline{N}\leqslant \|\delta\|_{1,\varpi}\leqslant \bar{N}} \left| \mathbb{G}_n^o\left( \frac{g_{\tau\varpi}(\tilde x'\delta)}{\|\delta\|_{1,\varpi}}\right) \right| \geqslant M/4 \right)\end{array}$$} where $\mathbb{G}_n^o$ is the symmetrized process.
Consider $ \mathcal{F}_{t,\tau,\varpi} = \{ \delta : \|\delta\|_{1,\varpi} = t\}$. We will consider the families of $\mathcal{F}_{t,\tau,\varpi}$ for $t \in [\underline{N},\bar N]$, $\tau \in \mathcal{T}$ and $\varpi\in\mathcal{W}$.
We will construct a finite net $\widehat{\mathcal{T}}\times \widehat{\mathcal{W}}\times \widehat{\mathcal{N}}$ of $\mathcal{T}\times \mathcal{W}\times[\underline{N},\bar{N}]$ such that
$$\begin{array}{ll}
\displaystyle\sup_{\tau \in \mathcal{T},\varpi\in\mathcal{W},t\in[\underline{N},\bar{N}],\delta \in \mathcal{F}_{t,\tau,\varpi}} \left| \mathbb{G}_n^o\left( \frac{g_{\tau\varpi}(\tilde x'\delta)}{\|\delta\|_{1,\varpi}}\right) \right| &
\leqslant \displaystyle 3 + \sup_{\tau\in\widehat{\mathcal{T}}, \varpi\in\widehat{\mathcal{W}},t\in \widehat{\mathcal{N}}}\sup_{\delta \in \mathcal{F}_{t,\tau,\varpi}} \left| \mathbb{G}_n^o\left( \frac{g_{\tau\varpi}(\tilde x'\delta)}{t}\right) \right|
\displaystyle =: 3 + \mathcal{A}^o.\end{array}$$
By triangle inequality we have
\begin{equation} \begin{array}{rl}
\left| \mathbb{G}_n^o\left( \frac{g_{\tau\varpi}(\tilde x'\delta)}{t} - \frac{g_{\tilde \tau\tilde\varpi}(\tilde x'\tilde\delta)}{\tilde t} \right) \right|& \leqslant \left| \mathbb{G}_n^o\left( \frac{g_{\tau\varpi}(\tilde x'\delta)}{t} - \frac{g_{\tilde \tau\varpi}(\tilde x'\delta)}{t} \right) \right| +\left| \mathbb{G}_n^o\left( \frac{g_{\tilde\tau\varpi}(\tilde x'\delta)}{t} - \frac{g_{\tilde\tau\tilde\varpi}(\tilde x'\delta)}{t} \right) \right| \\
& +\left| \mathbb{G}_n^o\left( \frac{g_{\tilde\tau\tilde\varpi}(\tilde x'\delta)}{t} - \frac{g_{\tilde\tau\tilde\varpi}(\tilde x'\tilde\delta)}{\tilde t} \right) \right|\end{array} \end{equation}
The first term in ((ref)) is such that
\begin{equation}\begin{array}{rl}
\left| \mathbb{G}_n^o\left( \frac{g_{\tau\varpi}(\tilde x'\delta)}{t} - \frac{g_{\tilde\tau\varpi}(\tilde x'\delta)}{t} \right) \right| & \leqslant \frac{2\sqrt{n}}{t}{\mathbb{E}_n}[ K_\varpi(W) |\tilde x'(\eta_{\tau\varpi}-\eta_{\tilde\tau\varpi})|] \\ & \leqslant \frac{2\sqrt{n}}{\underline{N}}\max_{i\leqslant n}\|\tilde x_i\|_\infty{\mathbb{E}_n}[K_\varpi(W)]\|\eta_{\tau\varpi} - \eta_{\tilde\tau\varpi}\|_1 \\
& \leqslant \frac{2\sqrt{n}}{\underline{N}}\max_{i\leqslant n}\|\tilde x_i\|_\infty{\mathbb{E}_n}[K_\varpi(W)] L_\eta|\tau - \tau'|.
\end{array}\end{equation}
Define a net $\widehat{\mathcal{T}}=\{\tau_1,\ldots,\tau_T\}$ such that
$$ |\tau_{k+1} - \tau_k | \leqslant \left\{2\sqrt{n}\frac{\max_{i\leqslant n}\|\tilde x_i\|_\infty}{\underline{N}}L_\eta\right\}^{-1}. $$
To bound the second term in ((ref)), note that $\mathcal{W}$ is a VC-class. Therefore, by Corollary 2.6.3 in vdV-W we have that conditional on $(W_i)_{i=1}^n$, there are at most $n^{d_W}$ different sets $\varpi \in \mathcal{W}$ that induce a different sequence $\{K_\varpi(W_1),\ldots,K_\varpi(W_n)\}$. Thus we can choose a (data-dependent) cover $\widehat{\mathcal{W}}$ with at most $n^{d_W}$ values of $\varpi$. Further, similarly to ((ref)) we have $\|\eta_{\tilde\tau\varpi}-\eta_{\tilde\tau\tilde\varpi}\|_1 \leqslant L_\eta\|\varpi-\tilde \varpi\|^\rho$ and
\begin{equation}\begin{array}{rl}
\left| \mathbb{G}_n^o\left( \frac{g_{\tilde\tau\varpi}(\tilde x'\delta)}{t} - \frac{g_{\tilde\tau\tilde\varpi}(\tilde x'\delta)}{t} \right) \right| & \leqslant \frac{2\sqrt{n}}{t}{\mathbb{E}_n}[ K_\varpi(W) |\tilde x'(\eta_{\tilde\tau\varpi}-\eta_{\tilde\tau\tilde\varpi})|] + \frac{2\sqrt{n}}{t}{\mathbb{E}_n}[ |K_\varpi(W)-K_{\tilde \varpi}(W) | |\tilde x'\delta|] \\ & \leqslant \frac{2\sqrt{n}}{\underline{N}}\max_{i\leqslant n}\|\tilde x_i\|_\infty{\mathbb{E}_n}[K_\varpi(W)]\|\eta_{\tilde \tau\varpi} - \eta_{\tilde\tau\tilde \varpi}\|_1\\
& \leqslant \frac{2\sqrt{n}}{\underline{N}}\max_{i\leqslant n}\|\tilde x_i\|_\infty{\mathbb{E}_n}[K_\varpi(W)] L_\eta\|\varpi - \tilde\varpi\|^\rho.
\end{array}\end{equation}
We define a net $\widehat{\mathcal{W}}$ such that $|\widehat{\mathcal{W}}|\leqslant n^{d_W}+\left\{2\sqrt{n}\frac{\max_{i\leqslant n}\|\tilde x_i\|_\infty}{\underline{N}}L_\eta\right\}^{d_W/\rho}$
To bound the third term in ((ref)), note that for any $\delta \in \mathcal{F}_{t,\tau,\varpi}$, $t \leqslant \tilde t$, by considering $\tilde\delta :=\delta(\tilde t/t) \in \mathcal{F}_{\tilde t,\tau,\varpi}$ we have
$$\begin{array}{rl}
\left| \mathbb{G}_n^o\left( \frac{g_{\tau\varpi}(\tilde x'\delta)}{t} - \frac{g_{\tau\varpi}(\tilde x'\delta(\tilde t/t))}{\tilde t} \right) \right| & \leqslant \left| \mathbb{G}_n^o\left( \frac{g_{\tau\varpi}(\tilde x'\delta)}{t} - \frac{g_{\tau\varpi}(\tilde x'\delta(\tilde t/t))}{t} \right) \right| + \left| \mathbb{G}_n^o\left( \frac{g_{\tau\varpi}(\tilde x'\delta(\tilde t/t))}{t} - \frac{g_{\tau\varpi}(\tilde x'\delta(\tilde t/t))}{\tilde t} \right) \right|\\
& = \frac{1}{t}\left| \mathbb{G}_n^o\left( g_{\tau\varpi}(\tilde x'\delta) - g_{\tau\varpi}(\tilde x'\delta[\tilde t/t]) \right) \right| + \left| \mathbb{G}_n^o\left( g_{\tau\varpi}(\tilde x'\delta(\tilde t/t)) \right) \right|\cdot \left| \frac{1}{t} - \frac{1}{\tilde t}\right| \\
& \leqslant \sqrt{n}{\mathbb{E}_n}\left(\frac{|K_\varpi(W)\tilde x'\delta|}{t}\right) \frac{|t-\tilde t|}{t} + \sqrt{n}{\mathbb{E}_n}\left( |K_\varpi(W)\tilde x'\delta|\right) \frac{\tilde t}{t}\left| \frac{1}{t} - \frac{1}{\tilde t}\right|\\
& = 2 \sqrt{n}{\mathbb{E}_n}\left(\frac{|K_\varpi(W)\tilde x'\delta|}{t}\right) \left| \frac{t-\tilde t}{ t} \right| \leqslant 2\sqrt{n} \left| \frac{t-\tilde t}{t} \right|.
\end{array}
$$
We let $\widehat{\mathcal{N}}$ be a $\varepsilon$-net $\{\underline{N}=:t_1,t_2,\ldots,t_K:= \bar N\}$ of $[\underline{N},\bar N]$ such that $|t_k-t_{k+1}|/t_k \leqslant 1/(2\sqrt{n})$. Note that we can achieve that with $|\widehat{\mathcal{N}}|\leqslant 1 + \left\lfloor 3\sqrt{n}\log(\bar N/ \underline{N}) \right\rfloor$.
By Markov bound, we have
$$ \begin{array}{rl}
\mathrm{P}( \mathcal{A}^o \geqslant K ) & \leqslant \min_{\psi\geqslant 0} \exp(-\psi K){\mathrm{E}}[ \exp(\psi\mathcal{A}^o)]\\
& \leqslant 8p|\widehat{\mathcal{T}}|\cdot|\widehat{\mathcal{W}}|\cdot|\widehat{\mathcal{N}}|\min_{\psi\geqslant 0} \exp(-\psi K)\exp\left( 8\psi^2\right)\\
&\leqslant 8p|\widehat{\mathcal{T}}|\cdot|\widehat{\mathcal{W}}|\cdot|\widehat{\mathcal{N}}|\exp (-K^2/32 )
\end{array}
$$ here we set $\psi = K/16 $ and bound ${\mathrm{E}}[ \exp(\psi\mathcal{A}^o)]$ as follows
{ $$\begin{array}{rl}
\displaystyle {\mathrm{E}}\left[ \exp\left( \psi \mathcal{A}^o \right)\right]&
\displaystyle \leqslant_{(1)} 2|\widehat{\mathcal{T}}|\cdot|\widehat{\mathcal{W}}|\cdot|\widehat{\mathcal{N}}|\sup_{(\tau,\varpi,t)\in \widehat{\mathcal{T}}\times\widehat{\mathcal{W}}\times\widehat{\mathcal{N}}} {\mathrm{E}}\left[ \exp\left( \psi \sup_{\|\delta\|_{1,\varpi}=t} \mathbb{G}_n^o\left( \frac{g_{\tau\varpi}(\tilde x'\delta)}{t}\right) \right)\right] \\
& \displaystyle \leqslant_{(2)} 2|\widehat{\mathcal{T}}|\cdot|\widehat{\mathcal{W}}|\cdot|\widehat{\mathcal{N}}|\sup_{(\tau,\varpi,t)\in \widehat{\mathcal{T}}\times\widehat{\mathcal{W}}\times\widehat{\mathcal{N}}} {\mathrm{E}}\left[ \exp\left( 2\psi \sup_{\|\delta\|_{1,\varpi}=t} \mathbb{G}_n^o\left(\frac{K_\varpi(W)\tilde x'\delta}{t}\right) \right)\right] \\
& \displaystyle \leqslant_{(3)} 2|\widehat{\mathcal{T}}|\cdot|\widehat{\mathcal{W}}|\cdot|\widehat{\mathcal{N}}|\sup_{(\tau,\varpi,t)\in \widehat{\mathcal{T}}\times\widehat{\mathcal{W}}\times\widehat{\mathcal{N}}} {\mathrm{E}}\left[ \exp\left( 2\psi \left[\sup_{\|\delta\|_{1,\varpi}=t}\frac{\|\delta\|_{1,\varpi}}{t}\max_{j\leqslant p}\frac{|\mathbb{G}_n^o(K_\varpi(W)\tilde x_{j})|}{\{{\mathbb{E}_n}[K_\varpi(W)\tilde x_j^2]\}^{1/2}} \right]\right)\right] \\
& \displaystyle =_{(4)} 2|\widehat{\mathcal{T}}|\cdot|\widehat{\mathcal{W}}|\cdot|\widehat{\mathcal{N}}|\sup_{(\tau,\varpi,t)\in \widehat{\mathcal{T}}\times\widehat{\mathcal{W}}\times\widehat{\mathcal{N}}} {\mathrm{E}}\left[ \exp\left( 2\psi \left[\max_{j\leqslant p}\frac{|\mathbb{G}_n^o(K_\varpi(W)\tilde x_{j})|}{\{{\mathbb{E}_n}[K_\varpi(W)\tilde x_j^2]\}^{1/2}} \right]\right)\right] \\
& \displaystyle \leqslant_{(5)} 4p|\widehat{\mathcal{T}}|\cdot|\widehat{\mathcal{W}}|\cdot|\widehat{\mathcal{N}}|\max_{j\leqslant p}\sup_{\varpi\in\widehat{\mathcal{W}}}{\mathrm{E}}\left[ \exp\left( 4\psi \frac{\mathbb{G}_n^o(K_\varpi(W)\tilde x_{j})}{\{{\mathbb{E}_n}[K_\varpi(W)\tilde x_j^2]\}^{1/2}} \right)\right] \\
& \displaystyle \leqslant_{(6)} 8p|\widehat{\mathcal{T}}|\cdot|\widehat{\mathcal{W}}|\cdot|\widehat{\mathcal{N}}|\exp\left( 8\psi^2 \right) \\
\end{array} $$}
here (1) follows by $\exp(\max_{i\in I} |z_i|) \leqslant 2|I|\max_{i\in I} \exp(z_i)$, (2) by contraction principle (apply Theorem 4.12 LedouxTalagrandBook with $t_i=K_\varpi(W_i)\tilde x_i'\delta$, and $\phi_i(t_i)=\rho_\tau(K_\varpi(W_i)\tilde y_i-K_\varpi(W_i)\tilde x_i'\eta_\tau + t_i)-\rho_\tau(K_\varpi(W_i)\tilde y_i-K_\varpi(W_i)\tilde x_i'\eta_\tau)$ so that $|\phi_i(s)-\phi_i(t)|\leqslant |s-t|$ and $\phi_i(0)=0$,
(3) follows by $$|\mathbb{G}_n^o(K_\varpi(W)\tilde x'\delta)| \leqslant \|\delta\|_{1,\varpi}\max_{j\leqslant p}|\mathbb{G}_n^o(K_\varpi(W)\tilde x_j)/\{{\mathbb{E}_n}[K_\varpi(W)\tilde x_j^2]\}^{1/2}|,$$ (4) by the definition of suprema, (5) we again use $\exp(\max_{i\in I} |z_i|) \leqslant 2|I|\max_{i\in I} \exp(z_i)$, and (6) $\exp(z)+\exp(-z)\leqslant 2\exp(z^2/2)$.
\end{proof}
\begin{lemma}[Estimation Error of Refitted Quantile Regression] Consider an arbitrary vector $\widehat\eta_u$ and suppose $\|\eta_u\|_0\leqslant s$. Let $\|r_{iu} \leqslant \bar r_u\|_{n,\varpi}$, $|{\rm support}(\widehat\eta_u)| \leqslant \widehat s_u$ and ${\mathbb{E}_n}[K_\varpi(W)\{\rho_\tau(\tilde y_i - \tilde x_i'\widehat\eta_u)-\rho_\tau(\tilde y_i-\tilde x_i'\eta_u)\}] \leqslant \widehat Q_u$ for all $u\in\mathcal{U}$ hold. Furthermore, suppose that
{ $$\sup_{\footnotesize {\tiny u=(\tau,\varpi)\in\mathcal{U}}} \left| {\mathbb{E}_n}\left( K_\varpi(W)\frac{\rho_\tau(\tilde y-\tilde x'\widetilde\eta_u)-\rho_\tau(\tilde y-\tilde x'\eta_u)}{\|\widetilde\eta_u-\eta_u\|_{1,\varpi}} - {\mathrm{E}}\left[ K_\varpi(W)\frac{\rho_\tau(\tilde y-\tilde x'\widetilde\eta_u)-\rho_\tau(\tilde y-\tilde x'\eta_u)}{\|\widetilde\eta_u-\eta_u\|_{1,\varpi}} \mid W, \tilde x\right]\right) \right| \leqslant \frac{t_3}{\sqrt{n}}. $$}
Under these events, we have for $n$ large enough,
$$\|\sqrt{f_u}\tilde x_i'(\widetilde \eta_u - \eta_u)\|_{n,\varpi} \lesssim \widetilde N_u:= \sqrt{ \frac{ (\widehat s_u + s)}{\phi_{{\rm min}}(u,\widehat s_u+s)}}(K_{n1}+t_3/\sqrt{n}) +K_{n2} + \bar f \bar r_u + \widehat Q^{1/2}_u$$
where $\phi_{{\rm min}}(u,k)=\inf_{\|\delta\|_0=k} \|\sqrt{f_u}\tilde x'\delta\|_{n,\varpi}^2/\|\delta\|^2$, provided that
\begin{equation} \sup_{u\in \mathcal{U}, \|\bar\delta\|_0\leqslant \widehat s_u+s} \frac{\bar f'{\mathbb{E}_n}[K_\varpi(W)(|r_u|+|r_u|^2)|\tilde x'\bar\delta|^2]}{{\mathbb{E}_n}[K_\varpi(W)f_u|\tilde x'\bar \delta|^2]} +\widetilde N_u/\bar q_{A_u} \to 0.\end{equation}
where $A_u = \{ \delta \in {\mathbb{R}}^p : \|\delta\|_0 \leqslant \widehat s_u + s\}$.\end{lemma}
\begin{proof}[Proof of Lemma (ref)]
Let $ \ \widehat \delta_u = \widehat \eta_u - \eta_u$ which satisfies $\|\widehat \delta_u\|_0\leqslant \widehat s_u + s$. By optimality of $\widetilde
\eta_u$ in the refitted quantile regression we have
\begin{equation} \begin{array}{rl}
{\mathbb{E}_n}[K_\varpi(W)\rho_\tau( \tilde y_i-\tilde x_i'\widetilde \eta_u )] - {\mathbb{E}_n}[K_\varpi(W)\rho_\tau(\tilde y_i-\tilde x_i'\eta_u )] &\\
\leqslant {\mathbb{E}_n}[K_\varpi(W)\rho_\tau(\tilde y_i-\tilde x_i'\widehat \eta_u )] - {\mathbb{E}_n}[K_\varpi(W)\rho_\tau(\tilde y_i-\tilde x_i'\eta_u )] \leqslant \widehat Q_u\end{array}\end{equation}
where the second inequality holds by assumption.
Moreover, by assumption, uniformly over $u\in \mathcal{U}$, we have conditional on $(W_i,\tilde x_i, r_{iu})_{i=1}^n$ that
\begin{equation}
\left| \mathbb{G}_n\left( K_\varpi(W)\frac{\rho_\tau(\tilde y-\tilde x'(\eta_u+\widetilde\delta_u))-\rho_\tau(\tilde y-\tilde x'\eta_u)}{\|\widetilde\delta_u\|_{1,\varpi}}\right) \right| \leqslant t_3.\end{equation}
Thus combining relations
((ref)) and ((ref)), we have $$
{\mathbb{E}_n}[{\mathrm{E}}[K_\varpi(W)\{\rho_u(\tilde y-\tilde x'(\eta_u+\widetilde \delta_u) )-\rho_u(\tilde y-\tilde x'\eta_u )\}\vert \tilde x, \tilde r, \varpi]] \leqslant
\|\widetilde \delta_u\|_{1,\varpi}t_3/\sqrt{n} + \widehat Q_u .
$$
Invoking the sparse identifiability relation of Lemma (ref), since the required condition on the approximation errors $r_u$'s holds by assumption ((ref)), for $n$ large enough $$ \displaystyle
\frac{\|\sqrt{f_u}\tilde x'\widetilde\delta_u\|_{n,\varpi}^2}{4} \wedge \left\{ \bar q_{A_u}\|\sqrt{f_u}\tilde x'\widetilde\delta_u\|_{n,\varpi}\right\} \leqslant \displaystyle K_{n2}\|\sqrt{f_u}\tilde x'\widetilde\delta_u\|_{n,\varpi} + \|\widetilde \delta_u\|_{1,\varpi} (K_{n1}+t_3/\sqrt{n}) + \widehat Q_u,
$$ where $\bar q_{A_u}$ is defined with $A_u:=\{ \delta : \|\delta\|_0 \leqslant \widehat s_u+s\}$. Moreover, by the sparsity of $\tilde\delta_u$ we have $\|\widetilde \delta_u\|_{1,\varpi}\leqslant \sqrt{(\widehat s_u+s)/\phi_{{\rm min}}(u,\widehat s_u+s)}\|\sqrt{f_u}\tilde x'\widetilde\delta_u\|_{n,\varpi}$ so that we have for $t=\|\sqrt{f_u}\tilde x'\widetilde\delta_u\|_{n,\varpi}$,
$$ \begin{array}{rcl} \displaystyle
\frac{t^2}{4} \wedge \left\{ \bar q_{A_u}t\right\} & \leqslant & t(\displaystyle K_{n2} + \sqrt{(\widehat s_u+s)/\phi_{{\rm min}}(u,\widehat s_u+s)} \{K_{n1}+t_3/\sqrt{n}\}) + \widehat Q_u.
\end{array}$$
Note that for positive numbers $(t^2/4) \wedge (\bar q_{A_u} t) \leqslant A + B t$ implies $t^2/4 \leqslant A+Bt$ provided $\bar q_{A_u}/2 > B$ and $2\bar q_{A_u}^2>A$. (Indeed, otherwise $(t^2/4) \geqslant q t$ so that $t\geqslant 4q$, which in turn implies that $2\bar q_{A_u} ^2 + \bar q_{A_u} t/2 \leqslant (t^2/4) \wedge \bar q_{A_u} t \leqslant A + Bt$.) Note that $\bar q_{A_u}/2 > B$ and $2\bar q_{A_u}^2>A$ is implied by condition ((ref)) when we set $A = \widehat Q_u$ and $B = (\displaystyle K_{n2} + \sqrt{(\widehat s_u+s)/\phi_{{\rm min}}(u,\widehat s_u+s_u)} \{K_{1n}+t_3/\sqrt{n}\})$.
Thus the minimum is achieved in the quadratic part. Therefore, for $n$ sufficiently large, we have
$$\|\sqrt{f_u}\tilde x'\widetilde\delta_u\|_{n,\varpi} \leqslant \widehat Q_u^{1/2} + K_{n2} + (K_{n1}+t_3/\sqrt{n})\sqrt{(\widehat s_u+s)/\phi_{{\rm min}}(u,\widehat s_u+s_u)}.$$
\end{proof}
Under the condition $\max_{i\leqslant n} \|\tilde x_i\|_\infty^2\log(n\vee p) = o(n \min_{\tau\in\mathcal{T}}\tau(1-\tau))$, the next result provides new bounds for the data driven penalty choice parameter when the quantile indices in $\mathcal{T}$ can approach the extremes.
\begin{lemma}[Pivotal Penalty Parameter Bound]
Let $\underline{\tau}=\min_{\tau\in\mathcal{T}}\tau(1-\tau)$ and $K_n=\max_{i\leqslant n, j\in[p]}|\tilde x_{ij}/\widehat\sigma_j|$, $\widehat\sigma_j = {\mathbb{E}_n}[\tilde x_j^2]^{1/2}$. Under $K_n^2\log(p/\underline{\tau})=o(n\underline{\tau})$, for $n$ large enough we have that for some constant $\bar C$
$$\Lambda(1-\xi\vert \tilde x_1,\ldots,\tilde x_n) \leqslant \bar{C}\sqrt{\frac{\log(16p/(\underline{\tau}\xi))}{n}}$$ where $\Lambda(1-\xi\vert \tilde x_1,\ldots,\tilde x_n)$ is the $1-\xi$ quantile of $ \max_{j\in [p]}\sup_{\tau \in \mathcal{T}} \left|\frac{\sum_{i=1}^n\tilde x_{ij}(\tau - 1\{U_i \leqslant \tau\})}{\widehat\sigma_j\sqrt{\tau(1-\tau)}}\right|$ conditional on $\tilde x_1,\ldots,\tilde x_n$, and $U_i$ are independent uniform$(0,1)$ random variables.
\end{lemma}
\begin{proof}
Conditional on $\tilde x_1,\ldots, \tilde x_n$, letting $\widehat \sigma_j^2 = {\mathbb{E}_n}[x_{j}^2]$, we have that $$n\Lambda = \max_{j\in [p]} \sup_{\tau \in \mathcal{T}} \left|\frac{\sum_{i=1}^n\tilde x_{j}(\tau - 1\{U \leqslant \tau\})}{\widehat\sigma_j\sqrt{\tau(1-\tau)}}\right|.$$
Step 1. (Entropy Calculation) Let $\mathcal{F} = \{ \tilde x_{ij}(\tau - 1\{U_i \leqslant \tau\})/\widehat\sigma_j : \tau \in \mathcal{T}, j\in[p] \}$, $h_\tau = \sqrt{\tau(1-\tau)}$, and $\mathcal{G}=\{ f_\tau/ h_\tau : \tau \in \mathcal{T}\}$.
We have that
$$\begin{array}{rl}
d(f_\tau/h_\tau, f_{\bar \tau}/h_{\bar\tau})& \leqslant d(f_\tau, f_{\bar \tau})/h_\tau + d(f_{\bar \tau}/h_\tau, f_{\bar \tau}/h_{\bar\tau})\\
& \leqslant d(f_\tau, f_{\bar \tau})/h_\tau + d(0,f_{\bar \tau}/h_{\bar\tau})|h_\tau - h_{\bar \tau}|/h_\tau\end{array}$$
Therefore, since $\|F\|_Q\leqslant \|G\|_Q$ by $h_\tau \leqslant 1$, and $d(0,f_{\bar \tau}/h_{\bar\tau})\leqslant 1/h_{\bar\tau}$ we have
$$N(\epsilon\|G\|_Q,\mathcal{G},Q)\leqslant N(\epsilon\|F\|_Q/\{2 \min_{\tau\in \mathcal{T}} h_\tau\},\mathcal{F},Q)N(\epsilon/\{2\min_{\tau\in \mathcal{T}} h_\tau^2\}, \mathcal{T},|\cdot|).$$
Thus we have for some constants $K$ and $v$ that $$N(\epsilon\|G\|_Q,\mathcal{G},Q) \leqslant p( K / \{\epsilon \min_{\tau\in \mathcal{T}} h_\tau^2\} )^v.$$
Step 2.(Symmetrization) Since we have ${\mathrm{E}}[g^2] = 1$ for all $g\in\mathcal{G}$, by Lemma 2.3.7 in vdV-W we have
$$
{\mathrm{P}}( \Lambda \geqslant t\sqrt{n} ) \leqslant 4 {\mathrm{P}}( \max_{j\leqslant p} \sup_{\tau\in\mathcal{T}}\left|\mathbb{G}_n^o(g)\right| \geqslant t/4 )
$$ here $\mathbb{G}_n^o:\mathcal{G}\to \mathbb{R}$ is the symmetrized process generated by Rademacher variables. Conditional on $(x_1,u_1),\ldots,(x_n,u_n)$, we have that $\{\mathbb{G}_n^o(g):g\in\mathcal{G}\}$ is sub-Gaussian with respect to the $L_2(\mathbb{P}_n)$-norm by the Hoeffding inequality. Thus, by Lemma 16 in BC-SparseQR, for $\delta_n^2 = \sup_{g\in\mathcal{G}}{\mathbb{E}_n}[g^2]$ and $\bar\delta_n = \delta_n/\|G\|_{\mathbb{P}_n}$, we have
$$ {\mathrm{P}}( \sup_{g\in\mathcal{G}}|\mathbb{G}_n^o(g)|> C K\delta_n\sqrt{\log(pK/\underline{\tau})}\mid \{\tilde x_i,U_i\}_{i=1}^n) \leqslant \int_0^{\bar\delta_n/2}\epsilon^{-1} \{ p( K / \{\epsilon \min_{\tau\in \mathcal{T}} h_\tau^2\} )^v \}^{-C^2+1}d\epsilon$$
for some universal constant $K$.
In order to control $\delta_n$, note that
$ \delta_n^2 = \sup_{g\in\mathcal{G}}\frac{1}{\sqrt{n}}\mathbb{G}_n(g^2)+\mathrm{E}[g^2].$ In turn, since $\sup_{g\in\mathcal{G}}{\mathbb{E}_n}[g^4]\leqslant \delta_n^2\max_{i\leqslant n}G_i^2$, we have
$$ {\mathrm{P}}( \sup_{g\in\mathcal{G}}|\mathbb{G}_n^o(g^2)|> C\bar K \delta_n\max_{i\leqslant n}G_i\sqrt{\log(pK/\underline{\tau})}\mid \{\tilde x_i,U_i\}_{i=1}^n) \leqslant \int_0^{\bar\delta_n/2}\epsilon^{-1} \{p( K / \{\epsilon \underline{\tau}\} )^v \}^{-C^2+1}d\epsilon.$$
Thus with probability $1-\int_0^{1/2}\epsilon^{-1} \{ p( K/\epsilon \underline{\tau} )^v \}^{-C^2+1}d\epsilon$, since ${\mathrm{E}}[g^2]=1$ and $\max_{i\leqslant n}G_i\leqslant K_n/\sqrt{\underline{\tau}}$, we have $$\delta_n \leqslant 1 + \frac{C'K_n\sqrt{\log(pK/\underline{\tau})}}{\sqrt{n}\sqrt{\underline{\tau}}}.$$
Therefore, under $K_n\sqrt{\log(pK/\underline{\tau})}= o(\sqrt{n}\sqrt{\underline{\tau}})$, conditionally on $\{\tilde x_i\}_{i=1}^n$ and $n$ sufficiently large, with probability $1-2\int_0^{1/2}\epsilon^{-1} \{ p( K / \{\epsilon \underline{\tau}\} )^v \}^{-C^2+1}d\epsilon$ we have that
$$\sup_{g\in\mathcal{G}}|\mathbb{G}_n^o(g)| \leqslant 2C K\sqrt{\log(pK/\underline{\tau})}$$
The stated bound follows since for $C>2$ $$2\int_0^{1/2}\epsilon^{-1} \{ p( K / \{\epsilon \underline{\tau}\} )^v \}^{-C^2+1}d\epsilon \leqslant \{p/\underline{\tau}\}^{-C^2+1}2\int_0^{1/2}\epsilon^{-2+C^2}d\epsilon\leqslant \{p/\underline{\tau}\}^{-C^2+1}.$$
\end{proof}
\section{Inequalities}
\begin{lemma}[Transfer principle, oliveira2013lower]
Let $\widehat\Sigma$ and $\Sigma$ be $p\times p$ matrices with non-negative diagonal entries, and assume that for some $\eta \in (0,1)$ and $s \leqslant p$ we have
$$ \forall v\in {\mathbb{R}}^p, \|v\|_0 \leqslant s, v'\widehat\Sigma v \geqslant (1-\eta)v'\Sigma v $$
Let $D$ be a diagonal matrix such that $D_{kk} \geqslant \widehat \Sigma_{kk} - (1-\eta)\Sigma_{kk}$. Then for all $\delta \in {\mathbb{R}}^p$ we have
$$ \delta'\widehat\Sigma \delta \geqslant (1-\eta)\delta'\Sigma \delta - \|D^{1/2}\delta\|_1^2/(s-1).$$\end{lemma}
\begin{lemma}
Consider $\widehat\beta_u$ and $\beta_u$ with $\|\beta_u\|_0\leqslant s$. Denote by $\widehat \beta^{\lambda}_{u}$ the vector with $\widehat\beta^\lambda_{uj}=\widehat \beta_{uj}1\{\widehat\sigma^Z_{a\varpi j}|\widehat \beta_{uj}|\geqslant \lambda\}$ where $\widehat\sigma^Z_{a\varpi j}=\{{\mathbb{E}_n}[K_\varpi(W)(Z_j^a)^2]\}^{1/2}$. We have that
$$\begin{array}{rl}
\|\widehat \beta^\lambda_u - \beta_u\|_{1,\varpi} & \leqslant \|\widehat \beta_u - \beta_u \|_{1,\varpi}+\lambda s \\
|{\rm support}(\widehat\beta_u^\lambda)| & \leqslant s + \|\widehat\beta_u-\beta_u\|_{1,\varpi}/\lambda\\
\|Z^a(\widehat \beta_u^\lambda-\beta_u)\|_{n,\varpi} & \leqslant \|Z^a(\widehat \beta_u-\beta_u)\|_{n,\varpi} + \sqrt{\tilde{\phi}_{{\rm max}}(s,\varpi)} \{ 2\sqrt{s}\lambda + \|\widehat\beta_u-\beta_u\|_{1,\varpi}/\sqrt{s}\}\end{array}$$
here $\tilde{\phi}_{{\rm max}}(m,\varpi)= \sup_{1\leqslant \|\theta\|_0\leqslant m}\|\tilde Z^a\theta\|_{n,\varpi}/\|\theta\|$ and $\tilde Z^a_{ij} = Z^a_{ij}/\{{\mathbb{E}_n}[K_\varpi(W)(Z^a_j)^2]\}^{1/2}$. \end{lemma}
\begin{proof}
Let $T_u = {\rm support}(\beta_u)$. The first relation follows from the triangle inequality
$$\begin{array}{rl}
\|\widehat \beta^\lambda_u - \beta_u\|_{1,\varpi} &= \|(\widehat \beta^\lambda_u - \beta_u)_{T_u}\|_{1,\varpi} + \|(\widehat \beta^\lambda_u)_{T_u^c}\|_{1,\varpi} \\
&\leqslant \|(\widehat \beta^\lambda_u - \widehat\beta_u)_{T_u}\|_{1,\varpi} + \|(\widehat \beta_u - \beta_u)_{T_u}\|_{1,\varpi}+ \|(\widehat \beta^\lambda_u)_{T_u^c}\|_{1,\varpi} \\
& \leqslant \lambda s + \|(\widehat \beta_u - \beta_u)_{T_u}\|_{1,\varpi}+ \|(\widehat \beta_u)_{T_u^c}\|_{1,\varpi}\\
& = \lambda s + \|\widehat \beta_u - \beta_u\|_{1,\varpi} \end{array}$$
To show the second result note that $\|\widehat\beta_u-\beta_u\|_{1,\varpi} \geqslant \{|{\rm support}(\widehat\beta_u^\lambda)|-s\}\lambda$. Therefore, $$\begin{array}{rl}
|{\rm support}(\widehat\beta_u^\lambda)| \leqslant s + \|\widehat\beta_u-\beta_u\|_{1,\varpi}/\lambda
\end{array}$$
which yields the result.
To show the third bound, we start using the triangle inequality
$$ \|Z^a(\widehat\beta^\lambda_u-\beta_u)\|_{n,\varpi} \leqslant \|Z^a(\widehat\beta^\lambda_u-\widehat\beta_u)\|_{n,\varpi} + \|Z^a(\widehat\beta_u-\beta_u)\|_{n,\varpi}.$$
Without loss of generality, assume that the components are ordered so that $|(\widehat\beta^\lambda_u-\widehat\beta_u)_j|\widehat\sigma_{uj}$ is decreasing. Let $T_1$ be the set of $s$ indices corresponding to the largest values of $|(\widehat\beta^\lambda_u-\widehat\beta_u)_j|\widehat\sigma_{uj}$. Similarly define $T_k$ as the set of $s$ indices corresponding to the largest values of $|(\widehat\beta^\lambda_u-\widehat\beta_u)_j|\widehat\sigma_{uj}$ outside $\cup_{m=1}^{k-1}T_m$. Therefore, $\widehat\beta^\lambda_u-\widehat\beta_u=\sum_{k=1}^{\left\lceil p/s \right\rceil} (\widehat\beta^\lambda_u-\widehat\beta_u)_{T_k}$. Moreover,
given the monotonicity of the components, $\|(\widehat\beta^\lambda_u-\widehat\beta_u)_{T_k}\|_{2,\varpi} \leqslant \|(\widehat\beta^\lambda_u-\widehat\beta_u)_{T_{k-1}}\|_1/\sqrt{s}$. Then, we have
$$ \begin{array}{rl}
\|Z^a(\widehat\beta^\lambda_u-\widehat\beta_u)\|_{n,\varpi} & = \|Z^a\sum_{k=1}^{\left\lceil p/s \right\rceil} (\widehat\beta^\lambda_u-\widehat\beta_u)_{T_k}\|_{n,\varpi}\\
& \leqslant \|Z^a(\widehat\beta^\lambda_u-\widehat\beta_u)_{T_1}\|_{n,\varpi} + \sum_{k\geqslant 2}\|Z^a(\widehat\beta^\lambda_u-\widehat\beta_u)_{T_k}\|_{n,\varpi}\\
& \leqslant \sqrt{\tilde{\phi}_{{\rm max}}(s,\varpi)}\|(\widehat\beta^\lambda_u-\widehat\beta_u)_{T_1}\|_{2,\varpi}+\sqrt{\tilde{\phi}_{{\rm max}}(s,\varpi)}\sum_{k\geqslant 2}\|(\widehat\beta^\lambda_u-\widehat\beta_u)_{T_k}\|_{2,\varpi}\\
& \leqslant \sqrt{\tilde{\phi}_{{\rm max}}(s,\varpi)}\lambda\sqrt{s} + \sqrt{\tilde{\phi}_{{\rm max}}(s,\varpi)}\sum_{k\geqslant 1}\|(\widehat\beta^\lambda_u-\widehat\beta_u)_{T_k}\|_{1,\varpi}/\sqrt{s}\\
& = \sqrt{\tilde{\phi}_{{\rm max}}(s,\varpi)}\lambda\sqrt{s} + \sqrt{\tilde{\phi}_{{\rm max}}(s,\varpi)} \|\widehat\beta^\lambda_u-\widehat\beta_u\|_{1,\varpi}/\sqrt{s} \\
& \leqslant \sqrt{\tilde{\phi}_{{\rm max}}(s,\varpi)} \{ 2\lambda\sqrt{s} + \|\widehat\beta_u-\beta_u\|_{1,\varpi}/\sqrt{s}\} \\ \end{array}$$
here the last inequality follows from the first result and the triangle inequality.
\end{proof}
\begin{lemma}[Supremum of Sparse Vectors on Symmetrized Random Matrices]
Let $\widehat \mathcal{U}$ denote a finite set and $(X_{iu})_{u\in \widehat \mathcal{U}}$, $i=1,\ldots, n$, be fixed vectors such that $X_{iu} \in {\mathbb{R}}^p$ and $\max_{1\leqslant i\leqslant n}\max_{u\in \widehat\mathcal{U}}\|X_{iu}\|_\infty \leqslant K$. Furthermore define $$\delta_n:= \bar C K \sqrt{k}\left(\sqrt{\log |\widehat\mathcal{U}|} + \sqrt{1+\log p} + \log k \sqrt{\log (p\vee n)} \sqrt{\log n} \right)/\sqrt{n},$$ where $\bar C$ is a universal constant. Then,
$$ {\mathrm{E}} \left[ \sup_{\|\theta\|_0\leqslant k, \|\theta\| =1}\max_{u\in\widehat\mathcal{U}} \left| {\mathbb{E}_n} [ \varepsilon(\theta'X_{u})^2] \right| \right] \leqslant \delta_n \sup_{\|\theta\|_0\leqslant k, \|\theta\| =1, u\in\widehat \mathcal{U}} \sqrt{{\mathbb{E}_n}[(\theta'X_{u})^2]}. $$
\end{lemma}
\begin{proof}
See BCCW-ManyProcesses for the proof.
\end{proof}
\begin{corollary}[Supremum of Sparse Vectors on Many Random Matrices]
Let $\widehat \mathcal{U}$ denote a finite set and $(X_{iu})_{u\in \widehat \mathcal{U}}$, $i=1,\ldots, n$, be independent (across i) random vectors such that $X_{iu} \in {\mathbb{R}}^p$ and $$\sqrt{{\mathrm{E}}[ \max_{1\leqslant i\leqslant n}\max_{u\in \widehat\mathcal{U}}\|X_{iu}\|_\infty^2]} \leqslant K.$$ Furthermore define $$\delta_n:= \bar C K \sqrt{k}\left(\sqrt{\log |\widehat\mathcal{U}|} + \sqrt{1+\log p} + \log k \sqrt{\log (p\vee n)} \sqrt{\log n} \right)/\sqrt{n},$$ here $\bar C$ is a universal constant. Then,
$$ {\mathrm{E}}\left[ \sup_{\|\theta\|_0\leqslant k, \|\theta\| =1}\max_{u\in\widehat\mathcal{U}} \left| {\mathbb{E}_n}\left[ (\theta'X_{u})^2 - {\mathrm{E}}[(\theta'X_{u})^2] \right]\right|\right] \leqslant \delta_n^2 + \delta_n \sup_{\|\theta\|_0\leqslant k, \|\theta\| =1, u\in\widehat \mathcal{U}} \sqrt{{\mathbb{E}_n}[{\mathrm{E}}[(\theta'X_{u})^2]]}. $$
\end{corollary}
We will also use the following result of chernozhukov2012gaussian.
\begin{lemma}[Maximal Inequality]
Work with the setup above. Suppose that $F\geqslant \sup_{f \in \mathcal{F}}|f|$ is a measurable envelope for $\mathcal{F}$
with $\| F\|_{P,q} < \infty$ for some $q \geqslant 2$. Let $M = \max_{i\leqslant n} F(W_i)$ and $\sigma^{2} > 0$ be any positive constant such that $\sup_{f \in \mathcal{F}} \| f \|_{P,2}^{2} \leqslant \sigma^{2} \leqslant \| F \|_{P,2}^{2}$. Suppose that there exist constants $a \geqslant e$ and $v \geqslant 1$ such that
\begin{equation*}
\log \sup_{Q} N(\epsilon \| F \|_{Q,2}, \mathcal{F}, \| \cdot \|_{Q,2}) \leqslant v \log (a/\epsilon), \ 0 < \epsilon \leqslant 1.
\end{equation*}
Then
\begin{equation*}
{\mathrm{E}}_P [ \sup_{f\in \mathcal{F}} | \mathbb{G}_n(f)| ] \leqslant K \left( \sqrt{v\sigma^{2} \log \left ( \frac{a \| F \|_{P,2}}{\sigma} \right ) } + \frac{v\| M \|_{P, 2}}{\sqrt{n}} \log \left ( \frac{a \| F \|_{P,2}}{\sigma} \right ) \right),
\end{equation*}
here $K$ is an absolute constant. Moreover, for every $t \geqslant 1$, with probability $> 1-t^{-q/2}$,
\begin{multline*}
\sup_{f\in \mathcal{F}} | \mathbb{G}_n(f)| \leqslant (1+\alpha) {\mathrm{E}}_P [ \sup_{f\in \mathcal{F}} | \mathbb{G}_n(f)| ] + K(q) \Big [ (\sigma + n^{-1/2} \| M \|_{P,q}) \sqrt{t}
+ \alpha^{-1} n^{-1/2} \| M \|_{P,2}t \Big ], \
\end{multline*}
$\forall \alpha > 0$ where $K(q) > 0$ is a constant depends only on $q$. In particular, setting $a \geqslant n$ and $t = \log n$,
with probability $> 1- c(\log n)^{-1}$,
\begin{equation}
\sup_{f\in \mathcal{F}} | \mathbb{G}_n(f)| \leqslant K(q,c) \left ( \sigma \sqrt{v \log \left ( \frac{a \| F \|_{P,2}}{\sigma} \right ) } + \frac{v
\| M \|_{P,q} } {\sqrt{n}}\log \left ( \frac{a \| F \|_{P,2}}{\sigma} \right ) \right),
\end{equation}
here $ \| M \|_{P,q} \leqslant n^{1/q} \| F\|_{P,q}$ and $K(q,c) > 0$ is a constant depending only on $q$ and $c$.
\end{lemma}
\section{Confidence Regions for Function-Valued Parameters Based on Moment Conditions}
For completeness, in this section we collect an adaptation of the results of BCCW-ManyProcesses that are invoked in our proofs. The main difference is the weakening of the identification condition (which is allowed to decrease to zero, see the parameter $j_n$ in Condition (ref) below). We are interested in function-valued target parameters indexed by $u \in \mathcal{U} \subset \mathbb{R}^{d_u}$. The true value of the target parameter is denoted by $$\theta^0 = (\theta_{u j})_{u \in \mathcal{U}, j\in[\tilde p]}, \ \ \text{where} \ \ \theta_{u j} \in \Theta_{u j} \text{ for each }
u \in \mathcal{U} \ \ \mbox{and} \ \ j\in[\tilde p].$$
For each $u \in \mathcal{U}$ and $j\in [\tilde p]$, the parameter $\theta_{uj}$ is characterized as the solution to the following moment condition:
\begin{equation}
{\mathrm{E}}[ \psi_{u j}(W_{uj}, \theta_{uj}, \eta_{uj} )] = 0,
\end{equation}
where $W_{uj}$ is a random vector that takes values in a Borel set $\mathcal{W}_{uj} \subset \mathbb{R}^{d_w}$, $\eta^0=(\eta_{uj})_{u\in\mathcal{U},j\in[\tilde p]}$ is a nuisance parameter where $\eta_{uj} \in T_{uj}$ a convex set, and the moment function
\begin{equation}
\psi_{u j}: \mathcal{W}_{uj} \times \Theta_{uj} \times T_{ujn} \mapsto \mathbb{R}, \ \ (w, \theta, t) \mapsto \psi_{uj}(w, \theta, t)\end{equation}
is a Borel measurable map.
We assume that the (continuum) nuisance parameter $\eta^0$ can be
modelled and estimated by $\widehat\eta = (\widehat\eta_{uj})_{u\in\mathcal{U}, j\in[\tilde p]}$. We will discuss examples where the corresponding $\eta^0$ can be estimated using modern regularization and post-selection methods such as Lasso and Post-Lasso (although other procedures can be applied). The estimator $\check \theta_{u j}$ of $\theta_{u j}$ is constructed as any approximate $\epsilon_n$-solution in $\Theta_{uj}$ to a sample analog of the moment condition ((ref)), i.e.,
\begin{equation}
\max_{j\in[\tilde p]} \sup_{u \in \mathcal{U}} \left\{ |{\mathbb{E}_n}[ \psi_{uj}(W_{uj}, \check \theta_{uj}, \widehat \eta_{uj} ) ] | - \inf_{\theta_j \in \Theta_{uj}}|{\mathbb{E}_n}[ \psi_{u j}(W_{uj}, \theta, \widehat \eta_{uj} ) ] | \right\} \leqslant \epsilon_n = o_P(n^{-1/2}\delta_n).
\end{equation}
As discussed before, we rely on an orthogonality condition for regular estimation of $\theta_{uj}$, which we will state next.
\begin{definition}[\textbf{Near Orthogonality Condition}] For each $u \in \mathcal{U}$ and $j\in[\tilde p]$, we say that $\psi_{uj}$ obeys a general form of orthogonality with respect to ${\mathcal{H}}_{uj}$ uniformly in $u \in \mathcal{U}$, if the following conditions hold: the G{\^a}teaux derivative map
$$
\mathrm{D}_{u,j,\bar r}[\tilde \eta_{uj} - \eta_{uj}]:= \left. \partial_r {\mathrm{E}} \Bigg ( \psi_{u j} \Big\{ W_{uj}, \theta_{u j}, \eta_{uj}+ r \Big [\tilde \eta_{uj} - \eta_{uj}\Big] \Big\} \Bigg )\right|_{r=\bar r}
$$
exists for all $r \in [0,1)$, $\tilde \eta \in {\mathcal{H}}_{uj}$, $j\in\tilde p$, and $u \in \mathcal{U}$ and vanishes at $r=0$, namely,
\begin{equation}
|\mathrm{D}_{u,j,0}[\tilde \eta_{uj} - \eta_{uj}]|\leqslant \delta_n n^{-1/2} \ \ \text{ for all } \tilde \eta_{uj} \in {\mathcal{H}}_{uj}.
\end{equation}
\end{definition}
In what follows, we shall denote by $c_0$, $c$, and $C$ some positive constants.
\begin{assumption}[Moment Condition]
Consider a random element $W$, taking values in a measure space $(\mathcal{W}, \mathcal{A}_\mathcal{W})$, with law determined by a probability measure $P \in {\mathcal{P}}_n$.
The observed data $((W_{iu})_{u \in \mathcal{U}})_{i=1}^{n}$ consist of $n$ i.i.d. copies of a random element $(W_{u})_{u \in \mathcal{U}}$ which is generated as a suitably measurable transformation with respect to $W$ and $u$. Uniformly for all $n \geqslant n_0$ and $P \in \mathcal{P}_n$, the following conditions hold: (i) The true parameter value $\theta_{u j}$ obeys ((ref)) and is interior relative to $\Theta_{u j}$, namely there is a ball of radius $C n^{-1/2}{u_n} \log n $ centered at $\theta_{u j}$ contained in $\Theta_{u j}$ for all $u \in \mathcal{U}$, $j\in[\tilde p]$ with ${u_n}:= {\mathrm{E}}[\sup_{u \in \mathcal{U}, j\in [\tilde p]}|\sqrt{n}{\mathbb{E}_n}[\psi_{uj}(W_{uj}, \theta_{uj}, \eta_{uj} )]|]$; (ii) For each $u \in \mathcal{U}$ and $j \in [\tilde p]$, the map $ (\theta,\eta) \in \Theta_{uj} \times {\mathcal{H}}_{uj} \mapsto {\mathrm{E}}[\psi_{uj}(W_{uj}, \theta,\eta )]|$ is twice continuously differentiable; (iii) For all $u \in {\mathcal{U}}$ and $j \in [\tilde p]$, the moment function $\psi_{uj}$ obeys the orthogonality condition given in Definition (ref) for the set ${\mathcal{H}}_{uj} ={\mathcal{H}}_{ujn}$ specified in Assumption (ref); (iv) The following identifiability condition holds: $|{\mathrm{E}}[\psi_{u j}(W_{uj}, \theta, \eta_{uj})]| \geqslant \frac{1}{2}|J_{uj} (\theta- \theta_{u j})| \wedge c_0\ \text{ for all } \theta \in \Theta_{uj},$ with $J_{u j} := \left.\partial_\theta {\mathrm{E}}[ \psi_{u j} (W_{uj}, \theta, \eta_{uj})]\right|_{\theta=\theta_{uj}}$ satisfies $0<j_n<|J_{uj}| <C <\infty $ for all $u \in \mathcal{U}$ and $j\in [\tilde p]$; (v) The following smoothness conditions holds \begin{itemize}
• $\sup_{u \in \mathcal{U}, j \in [\tilde p],(\theta, \bar \theta) \in \Theta_{uj}^2, (\eta,\bar\eta) \in {\mathcal{H}}_{ujn}^2} \ \ \frac{{\mathrm{E}}[ \{ \psi_{uj}(W_{uj}, \theta,\eta) - \psi_{uj}(W_{uj}, \bar\theta,\bar\eta)\}^2]}{ \{ |\theta-\bar\theta|\vee \| \eta - \bar \eta\|_e\}^{\alpha}}\leqslant C$,
•
$ \sup_{u \in \mathcal{U}, (\theta, \eta) \in \Theta_{uj}\times {\mathcal{H}}_{ujn}, r\in[0,1)} \ \ \ \ \ | \partial_{r} {\mathrm{E}} \left [ \psi_{uj}(W_{uj}, \theta,\eta_{uj}+r\{\eta-\eta_{uj}\}) \right ]|/\|\eta-\eta_{uj}\|_e \leqslant \bar B_{1n}$,
• $\sup_{u \in \mathcal{U}, j \in [\tilde p], (\theta,\eta) \in \Theta_{uj}\times {\mathcal{H}}_{ujn}, r\in[0,1)} \frac{|\partial_{r}^2 {\mathrm{E}}[\psi_{uj}(W_{uj}, \theta_{uj}+r\{\theta-\theta_{uj}\},\eta_{uj}+r\{\eta - \eta_{uj}\})]|}{\{|\theta-\theta_{uj}|^2 \vee \|\eta-\eta_{uj}\|_{e}^2 \}} \leqslant \bar B_{2n}.$
\end{itemize}
\end{assumption}
Next we state assumptions on the nuisance functions. In what follows, let $\Delta_n \searrow 0$, $\delta_n \searrow 0$, and $\tau_n \searrow 0$ be sequences of constants approaching zero from above at a speed at most polynomial in $n$ (for example, $\delta_n \geqslant 1/n^c$ for some $c > 0$). \\
\begin{assumption}[Estimation of Nuisance Functions]
The following conditions hold for each $n \geqslant n_0$ and all $P \in \mathcal{P}_n$. The estimated functions $\widehat \eta_{uj} \in {\mathcal{H}}_{ujn}$ with probability at least $1- \Delta_n$,
${\mathcal{H}}_{ujn}$ is the set of measurable maps $\tilde \eta_{uj}$ such that
$$
\sup_{u\in\mathcal{U}}\max_{j\in[\tilde p]}\| \tilde \eta_{uj} - \eta_{uj}\|_{e} \leqslant \tau_n,
$$
here the $e$-norm is the same as in Assumption (ref), and whose complexity does not grow too quickly in the sense that
$\mathcal{F}_1 = \{ \psi_{uj}(W_{uj}, \theta, \eta): u \in \mathcal{U}, j \in [\tilde p], \theta \in \Theta_{uj}, \eta \in {\mathcal{H}}_{ujn} \cup \{\eta_{uj}\} \}$
is suitably measurable and its uniform covering entropy obeys:
$$
\sup_Q \log N(\epsilon \|F_1\|_{Q,2}, \mathcal{F}_1, \| \cdot \|_{Q,2}) \leqslant s_{n(\mathcal{U},\tilde p)} ( \log (a_n/\epsilon)) \vee 0,
$$
where $F_1(W)$ is an envelope for $\mathcal{F}_1$ which is measurable with respect to $W$ and satisfies $F_1(W)\geqslant \sup_{u\in \mathcal{U}, j\in[\tilde p], \theta\in \Theta_{uj}, \eta\in {\mathcal{H}}_{ujn}}|\psi_{uj}(W_{uj},\theta,\eta)|$ and $\|F_1\|_{P,q}\leqslant K_{n}$ for $q\geqslant 2$. The complexity characteristics $a_n \geqslant \max(n, K_n, \mathrm{e}) $ and $s_{n(\mathcal{U},\tilde p)} \geqslant 1$ obey the growth conditions:
$$
\begin{array}{rl}
n^{-1/2} \sqrt{ s_{n(\mathcal{U},\tilde p)} \log (a_n) } + n^{-1} s_{n(\mathcal{U},\tilde p)} n^{\frac{1}{q}} K_{n} \log (a_n) \leqslant \tau_n \\
\{(1\vee \bar B_{1n})(\tau_n/j_n)\}^{\alpha/2} \sqrt{ s_{n(\mathcal{U},\tilde p)} \log (a_n)} + s_{n(\mathcal{U},\tilde p)}n^{\frac{1}{q}-\frac{1}{2}}K_n\log (a_n) \log n \leqslant \delta_n,\\
\text{ and } \ \ \sqrt{n} \bar B_{2n} (1\vee \bar B_{1n}) (\tau_n/j_n)^2 \leqslant \delta_n
\end{array}$$
here $\bar B_{1n}$, $\bar B_{2n}$, $j_n$, $q$ and $\alpha$ are defined in Assumption (ref).
\end{assumption}
\begin{theorem}[{Uniform Bahadur representation for a Continuum of Target Parameters}] Under Assumptions (ref) and (ref), for an estimator $(\check \theta_{uj})_{u \in \mathcal{U},j\in[\tilde p]}$ that obeys equation ((ref)),
$$ \sqrt{n}\sigma_{uj}^{-1}(\check\theta_{uj} - \theta_{uj}) = \mathbb{G}_n \bar \psi_{uj} + O_P(\delta_n) \text{ in } \ell^\infty(\mathcal{U}\times[\tilde p]), \text{ uniformly in $P \in \mathcal{P}_n$},$$
here $\bar \psi_{uj}(W):= - \sigma_{uj}^{-1}J^{-1}_{uj} \psi_{uj}(W_{uj}, \theta_{uj}, \eta_{uj})$ and $\sigma_{uj}^2 = {\mathrm{E}}[J^{-2}_{uj} \psi_{uj}^2(W_{uj}, \theta_{uj}, \eta_{uj})]$.
\end{theorem}
The uniform Bahadur representation derived in Theorem (ref) is useful in the construction of simultaneous confidence bands for $(\theta_{uj})_{u\in\mathcal{U}, j\in[\tilde p]}$. This is achieved by new high-dimensional central limit theorems that have recently been developed in chernozhukov2013gaussian and chernozhukov2012gaussian. We will make use of the following regularity condition. In what follows $\bar{\delta}_n$ and $\Delta_n$ are fixed sequences going to zero, and we denote $\widehat \psi_{uj}(W):= - \widehat\sigma_{uj}^{-1}\widehat J_{uj}^{-1} \psi_{uj}(W_{uj}, \check \theta_{uj}, \widehat \eta_{uj})$ be the estimators of $\bar \psi_{uj}(W)$, with $\widehat J_{uj}$ and $\widehat\sigma_{uj}$ being suitable estimators of $J_{uj}$ and $\sigma_{uj}$.
In what follows, $\|\cdot\|_{\mathbb{P}_n,2}$ denotes the empirical $L_2(\mathbb{P}_n)$-norm with $\mathbb{P}_n$ as the empirical measure of the data.
\begin{assumption}[Score Regularity]
The following conditions hold for each $n \geqslant n_0$ and all $P \in \mathcal{P}_n$. (i) The class of function induced by the score $\mathcal{F}_0 = \{ \bar \psi_{uj}(W): u \in \mathcal{U}, j \in [\tilde p] \}$
is suitably measurable and its uniform covering entropy obeys:
$$
\sup_Q \log N(\epsilon \|F_0\|_{Q,2}, \mathcal{F}_0, \| \cdot \|_{Q,2}) \leqslant \varrho_{n} ( \log (A_n/\epsilon)) \vee 0,
$$
here $F_0(W)$ is an envelope for $\mathcal{F}_0$ which is measurable with respect to $W$ and satisfies $F_0(W)\geqslant \sup_{u\in \mathcal{U}, j\in[\tilde p]}|\bar \psi_{uj}(W)|$ and $\|F_0\|_{P,q}\leqslant L_{n}$ for $q\geqslant 4$. Furthermore, $c \leqslant \sup_{u\in \mathcal{U}, j\in [\tilde p]} {\mathrm{E}}[ |\bar \psi_{uj}(W)|^k] \leqslant CL_n^{k-2}$ for $k=2,3,4$. (ii) The set $\widehat{\mathcal{F}}_0 = \{ \bar \psi_{uj}(W)-\widehat\psi_{uj}(W): u \in \mathcal{U}, j \in [\tilde p] \}$ satisfies the conditions
$
\log N(\epsilon, \widehat{\mathcal{F}}_0, \|\cdot\|_{\mathbb{P}_n,2}) \leqslant \bar \varrho_{n} ( \log (\bar A_n/\epsilon)) \vee 0,
$ and $\sup_{u\in\mathcal{U},j\in[\tilde p]}{\mathbb{E}_n}[ \{ \bar \psi_{uj}(W) - \widehat \psi_{uj}(W)\}^2] \leqslant \bar\delta_n\{\rho_n\bar\rho_n\log(A_n\vee n)\log(\bar A_n\vee n)\}^{-1}$ with probability $1-\Delta_n$.
\end{assumption}
Assumption (ref) imposes condition on the class of functions induced by $\bar\psi_{uj}$ and on its estimators $\widehat\psi_{uj}$. Typically the bound $L_n$ on the moment of the envelope is smaller than $K_n$, and in many settings $\bar\rho_n=\rho_n \lesssim d_{\mathcal{U}}$ the dimension of $\mathcal{U}$.
Next let $\mathcal{N}$ denote a mean zero Gaussian process indexed by $\mathcal{U} \times [\tilde p]$ with covariance operator given by ${\mathrm{E}}[ \bar\psi_{uj}(W)\bar\psi_{u'j'}(W)]$ for $j,j' \in [\tilde p]$ and $u,u'\in\mathcal{U}$. Because of the high-dimensionality, indeed $\tilde p$ can be larger than the sample size $n$, the central limit theorem will be uniformly valid over “rectangles". This class of sets are rich enough to construct many confidence regions of interest in applications accounting for multiple testing. Let $\mathcal{R}$ denote the set of rectangles $R = \{ z \in {\mathbb{R}}^{\tilde p} : \max_{j\in A} z_j \leqslant t, \max_{j\in B} (-z_j) \leqslant t\}$ for all $A, B \subset [\tilde p]$ and $t\in {\mathbb{R}}$. The following result is a consequence of Theorem (ref) above and Corollary 2.2 of chernozhukov2012comparison.
\begin{corollary} Under Assumptions (ref), (ref) and Assumption (ref)(i), with $\delta_n = o(\{ \rho_n\log(A_n\vee n)\}^{-1/2})$, and $\rho_n\log(A_n\vee n) =o(\{ (n/L_n^2)^{1/7}\wedge (n^{1-2/q}/L_n^2)^{1/3}\})$, we have that
$$ \sup_{P \in \mathcal{P}_n} \sup_{R \in \mathcal{R}} \left| {\mathrm{P}}_P\left( \{\sup_{u\in\mathcal{U}}n^{1/2}\sigma_{uj}^{-1}(\check\theta_{uj}-\theta_{uj})\}_{j=1}^{\tilde p} \in R \right) - {\mathrm{P}}_P(\mathcal{N} \in R) \right| = o(1). $$
\end{corollary}
In order to derive a method to build confidence regions we approximate the process $\mathcal{N}$ by the Gaussian multiplier bootstrap based on estimates $\widehat\psi_{uj}$ of $\bar\psi_{uj}$, namely
$$ \widehat{\mathcal{G}} = (\widehat{\mathcal{G}}_{uj})_{u\in\mathcal{U}, j\in [\tilde p]} = \left\{ \frac{1}{\sqrt{n}}\sum_{i=1}^ng_i\widehat\psi_{uj}(W_i)\right\}_{u\in\mathcal{U}, j\in [\tilde p]}$$
here $(g_i)_{i=1}^n$ are independent standard normal random variables which are independent from the data $(W_i)_{i=1}^n$. Based on Theorem 5.2 of chernozhukov2013gaussian, the following result shows that the multiplier bootstrap provides a valid approximation to the large sample probability law of $\sqrt{n}(\check \theta_{uj}- \theta_{uj})_{u \in \mathcal{U},j\in[\tilde p]}$ over rectangles.
\begin{corollary}[\textbf{Uniform Validity of Gaussian Multiplier Bootstrap}]
Under Assumptions (ref), (ref) and Assumption (ref), with $\delta_n = o(\{ (1+d_{W})\rho_n\log(A_n\vee n)\}^{-1/2})$ and $\rho_n\log(A_n\vee n) = o(\{ (n/L_n^2)^{1/7}\wedge (n^{1-2/q}/L_n^2)^{1/3}\})$, we have that
$$ \sup_{P \in \mathcal{P}_n} \sup_{R \in \mathcal{R}} \left| {\mathrm{P}}_P\left( \{\sup_{u\in\mathcal{U}}n^{1/2}\sigma_{uj}^{-1}(\check\theta_{uj}-\theta_{uj})\}_{j=1}^{\tilde p} \in R \right) - {\mathrm{P}}_P(\widehat{\mathcal{G}} \in R\mid (W_i)_{i=1}^n) \right| = o(1) $$
\end{corollary}
\section{Continuum of $\ell_1$-Penalized M-Estimators}
For the reader's convenience, this section collects results on the estimation of a continuum of estimation of high-dimensional models via $\ell_1$-penalized estimators. We refer to BCCW-ManyProcesses for the proofs.
Consider a data generating process with a response variable $(Y_{u})_{u\in \mathcal{U}}$ and observable covariates $(X_u)_{u\in\mathcal{U}}$ satisfies for each $u\in \mathcal{U}$,
\begin{equation}\theta_u \in \arg\min_{\theta \in {\mathbb{R}}^p}{\mathrm{E}}[M_u(Y_{u},X_u,\theta,a_u)], \end{equation}
here $\theta_u$ is a $p$-dimensional vector, $a_u$ is a nuisance function that capture the misspecification of the model, $M_u$ is a pre-specified function, and the $p_u$-dimensional ($p_u\leqslant p$) covariate $X_u$ could have been constructed based on transformations of other variables. This implies that
$$\partial_\theta{\mathrm{E}}[M_u(Y_{u},X_u,\theta_u,a_u)]=0 \ \ \mbox{for all} \ u\in\mathcal{U}.$$ The solution $\theta_u$ is assumed to be sparse in the sense that for some process $(\theta_u)_{u\in\mathcal{U}}$ satisfies
$$\|\theta_u\|_0\leqslant s \ \mbox{for all} \ u\in\mathcal{U}.$$
Because of the nuisance function, such sparsity assumption is very mild and formulation ((ref)) encompasses several cases of interest including approximate sparse models. We focus on the estimation of $(\theta_u)_{u\in \mathcal{U}}$ and we assume that an estimate $\widehat a_u$ of the nuisance function $a_u$ is available and the criterion $M_u(Y_{u},X_u,\theta_u):=M_u(Y_{u},X_u,\theta_u,\widehat a_u)$ is used as a proxy for $M_u(Y_{u},X,\theta_u,a_u)$.
In the case of linear regression we have $M_u(y,x,\theta) = \frac{1}{2}(y-x'\theta)^2$. In the logistic regression case, we have
$M_u(y,x,\theta) = -\{1(y=1)\log {G}( x'\theta ) + 1(y=0)\log(1-{G}( x'\theta))\}$
with ${G}$ is the logistic link function ${G}(t)=\exp(t)/\{1+\exp(t)\}$. Additional examples include quantile regression models for $u\in (0,1)$.
\begin{example}[Quantile Regression Model]
Consider a data generating process $Y = F^{-1}_{Y\mid X}(U) = X'\theta_U + r_U(X)$, with $U\sim {\rm Unif}(0,1)$, and $X$ is a $p$-dimensional vector of covariates. The criterion $M_u(y,x,\theta) = ( u - 1\{ y \leqslant x'\theta\} )(y - x'\theta)$ with the (trivial) estimate $\widehat a_u = 0$ for the nuisance parameter $a_u = r_u$.
\end{example}
\begin{example}[Lasso with Estimated Weights]
We consider a linear model defined as
$ f_uY = f_uX'\theta_u + \bar r_u + \zeta_u, \ \ {\mathrm{E}}[f_uX\zeta_u] = 0$, here $X$ are $\bar p$-dimensional covariates, $\theta_u$ is a $s$-sparse vector, and $\bar r_u$ is an approximation error satisfies $\sup_{u\in\mathcal{U}}{\mathbb{E}_n}[\bar r_u^2] \lesssim_P s\log \bar p/n$. In this setting, $(Y,X)$ are observed and only an estimator $\widehat f_u$ of $f_u$ is available. This corresponds to nuisance parameter $a_u = (f_u,\bar r_u)$ and $\widehat a_u = (\widehat f_u,0)$ so that ${\mathbb{E}_n}[ M_u(Y,X,\theta,a_u)] = {\mathbb{E}_n}[ f_u^2(Y-X'\theta-\bar r_u)^2]$ and ${\mathbb{E}_n}[M_u(Y,X,\theta)]={\mathbb{E}_n}[\widehat f_u^2(Y-X'\theta)^2]$.
\end{example}
We assume that $n$ i.i.d. observations from dgps with ((ref)) holds, $\{( Y_{iu}, X_{iu})_{u\in\mathcal{U}}\}_{i=1}^n$, are observed to estimate $(\theta_u)_{u\in\mathcal{U}}$. For each $u\in\mathcal{U}$, a penalty level $\lambda$, and a diagonal matrix of penalty loadings $\widehat\Psi_u,$ we define the $\ell_1$-penalized $M_u$-estimator (Weighed-Lasso) as
\begin{equation} \widehat\theta_u \in \arg\min_{\theta} {\mathbb{E}_n}[M_u(Y_{u},X_u,\theta)] + \frac{\lambda}{n}\|\widehat\Psi_u\theta\|_1. \end{equation}
Furthermore, for each $u\in\mathcal{U}$, the post-penalized estimator (Post-Lasso) based on a set of covariates $\widetilde T_u$ is then defined as
\begin{equation}\widetilde\theta_u \in \arg\min_{\theta} {\mathbb{E}_n}[M_u(Y_{u},X_u,\theta)] \ \ : \ \ {\rm support}(\theta)\subseteq \widetilde T_u.\end{equation}
Potentially, the set $\widetilde T_u$ contains ${\rm support}(\widehat\theta_u)$ and possibly additional variables deemed as important (although in that case the total number of additional variables should also obey the same growth conditions that $s$ obeys). We will set $\widetilde T_u = {\rm support}(\widehat\theta_u)$ unless otherwise noted.
In order to handle the functional response data, the penalty level $\lambda$ and penalty loading $\widehat\Psi_u= {\rm diag}(\{\widehat l_{u k}, k=1,\ldots,p\})$ need to be set to control selection errors
uniformly over $u\in\mathcal{U}$. The choice of loading matrix is problem specific and we suggest to mimic the following “ideal" choice $\widehat\Psi_{u0} = {\rm diag}(\{ l_{u k}, k=1,\ldots,p\})$ with
\begin{equation} l_{uk} = \{{\mathbb{E}_n}\left[\{\partial_{\theta_k} M_u(Y_{u},X_u,\theta_u,a_u)\}^2\right]\}^{1/2} \end{equation} which is motivated by the use of self-normalized moderate deviation theory. In that case, it is suitable to set
$\lambda$ so that with high probability
\begin{equation} \frac{\lambda}{n} \geqslant c \sup_{u\in \mathcal{U}}\left\|\widehat \Psi^{-1}_{u0} {\mathbb{E}_n}\left[\partial_\theta M_u(Y_{u},X_u,\theta_u, a_u) \right] \right\|_\infty,
\end{equation} here $c>1$ is a fixed constant. Indeed, in the case that $\mathcal{U}$ is a singleton the choice above is similar to BickelRitovTsybakov2009, BC-PostLASSO, and BCW-SqLASSO. This approach was first employed for a continuum of indices $\mathcal{U}$ in the context of $\ell_1$-penalized quantile regression processes by BC-SparseQR.
To implement ((ref)), we propose setting the penalty level as
\begin{equation}\lambda = c \sqrt{n} \Phi^{-1}(1-\xi/\{2p N_n \}),
\end{equation}
here $N_n$ is a measure of the class of functions indexed by $\mathcal{U}$, $1-\xi$ (with $\xi = o(1)$) is a confidence level associated with the probability of event ((ref)), and $c>1$ is a slack constant. In many settings we can take $N_n = n^{d_{\mathcal{U}}}$. If the set $\mathcal{U}$ is a singleton, $N_n = 1$ suffices which corresponds to what is used in BCH2014inference.
\subsection{Generic Finite Sample Bounds}
In this subsection we derive finite sample bounds based on Assumption (ref) below. This assumption provides sufficient conditions that are implied by a variety of settings including generalized linear models.
\begin{assumption}[M-Estimation Conditions]
Let $\{(Y_{iu}, X_{iu}, u\in \mathcal{U}),i=1,\ldots, n\}$ be $n$ i.i.d. observations of the model ((ref)) and let $T_u = {\rm support}(\theta_u)$, here $\Vert T_u \Vert _0\leqslant s$, $u\in\mathcal{U}$. With probability $1-\Delta_n$ we have that for all $u\in \mathcal{U}$ there are weights $w_u=w_u(Y_u,X_u)$ and $C_{un}$ such that:
\begin{itemize}
• $|{\mathbb{E}_n}[\partial_{\theta} M_u(Y_u,X_u,\theta_u)-\partial_{\theta} M_u(Y_u,X_u,\theta_u,a_u)]'\delta|\leqslant C_{un}\|\sqrt{w_{u}} X_u'\delta\|_{\mathbb{P}_n,2} $;
• $\ell \widehat\Psi_{u0} \leqslant \widehat\Psi_u \leqslant L\widehat\Psi_{u0}$ for $\ell > 1/c$, and let $\tilde c = \frac{Lc+1}{\ell c-1}\sup_{u\in \mathcal{U}}\|\widehat \Psi_{u0}\|_\infty\|\widehat \Psi_{u0}^{-1}\|_\infty$;
• for all $\delta \in A_u$ there is $\bar q_{A_u}>0$ such that
$$\begin{array}{c}{\mathbb{E}_n}[M_u(Y_u,X_u,\theta_u + \delta)] - {\mathbb{E}_n}[M_u(Y_u,X_u,\theta_u)] -{\mathbb{E}_n}[\partial_{\theta} M_u(Y_u,X_u,\theta_u)]'\delta +2C_{un}\|\sqrt{w_{u}} X_u'\delta\|_{\mathbb{P}_n,2} \\ \geqslant \left\{\|\sqrt{w_{u}} X_u'\delta\|_{\mathbb{P}_n,2}^2\right\} \wedge \left\{ \bar q_{A_u}\|\sqrt{w_{u}} X_u'\delta\|_{\mathbb{P}_n,2}\right\}. \end{array}$$
\end{itemize}
\end{assumption}
In many applications we take the weights to be $w_u = w_u(X_u) = 1$ but we allow for more general weights. Assumption (ref)(a) bounds the impact of estimating the nuisance functions uniformly over $u\in\mathcal{U}$. In the setting with $s$-sparse estimands, we typically have $C_{un} \lesssim \{n^{-1}s \log (pn)\}^{1/2}$. The loadings $\widehat\Psi_u$ are assumed larger (but not too much larger) than the ideal choice $\widehat\Psi_{u0}$ defined in ((ref)). This is formalized in Assumption (ref)(b). Assumption (ref)(c) is an identification condition that will be imposed for specific choices of $A_u$ and $q_{A_u}$. It relates to conditions in the literature derived for the case of a singleton $\mathcal{U}$ and no nuisance functions, see the restricted strong convexity\footnote{Assumption (ref) (a) and (c) could have been stated with $\{C_{un}/\sqrt{s}\}\|\delta\|_1$ instead of $C_{un}\|\sqrt{w_{u}} X_u'\delta\|_{\mathbb{P}_n,2}$.} used in negahban2012unified and the non-linear impact coefficients used in BC-SparseQR and BCK2013robustQR.
The following results establish rates of convergence for the $\ell_1$-penalized solution with estimated nuisance functions ((ref)), sparsity bounds and rates of convergence for the post-selection refitted estimator ((ref)). They are based on restricted eigenvalue type conditions and sparse eigenvalue conditions. With the restricted eigenvalue is defined as $\bar\kappa_{u,2\tilde\mathbf{c}} = \inf_{\delta\in \Delta_{u, 2\tilde\mathbf{c}}} \|\sqrt{w_{u}} X_u'\delta\|_{\mathbb{P}_n,2}/\|\delta_{T_u}\|$ In the results for sparsity and post-selection refitted models, the minimum and maximum sparse eigenvalues,
$$ \phi_{{\rm min}}(m,u) = \min_{1\leqslant \|\delta\|_0\leqslant m}\frac{\|\sqrt{w_u}X_u'\delta\|_{\mathbb{P}_n,2}^2}{\|\delta\|^2} \ \ \mbox{and} \ \ \phi_{{\rm max}}(m,u) = \max_{1\leqslant \|\delta\|_0\leqslant m}\frac{\|\sqrt{w_{u}}X_u'\delta\|_{\mathbb{P}_n,2}^2}{\|\delta\|^2},$$
are also relevant quantities to characterize the behavior of the estimators.
\begin{lemma}
Suppose that Assumption (ref) holds with $\delta \in A_u = \{ \delta : \|\delta_{T^c_u}\|_1\leqslant 2\tilde\mathbf{c} \|\delta_{T_u}\|_1\} \cup \{ \delta : \|\delta\|_1 \leqslant \frac{6c\|\widehat\Psi_{u0}^{-1}\|_\infty}{\ell c - 1}\frac{n}{\lambda}C_{un}\|\sqrt{w_{u}} X_u'\delta\|_{\mathbb{P}_n,2}\}$ and $\bar q_{A_u}>3\left\{(L+\frac{1}{c})\|\widehat\Psi_{u0}\|_\infty\frac{\lambda\sqrt{s}}{n\bar\kappa_{u,2\tilde \mathbf{c}}}+ 9\tilde \mathbf{c} C_{un}\right\}.$ Suppose that $\lambda$ satisfies condition ((ref)) with probability $1-\Delta_n$. Then, with probability $1-2\Delta_n$ we have uniformly over $u\in \mathcal{U}$
$$\begin{array}{rl}
\|\sqrt{w_{u}} X_u'(\widehat \theta_u - \theta_u)\|_{\mathbb{P}_n,2} & \leqslant 3\left\{(L+\frac{1}{c})\|\widehat\Psi_{u0}\|_\infty\frac{\lambda\sqrt{s}}{n\bar\kappa_{u,2\tilde \mathbf{c}}}+ 9\tilde \mathbf{c} C_{un}\right\}\\
\|\widehat \theta_u - \theta_u\|_1 & \leqslant 3\left\{ \frac{(1+2\tilde\mathbf{c})\sqrt{s}}{\bar\kappa_{u,2\tilde \mathbf{c}}}+ \frac{6c\|\widehat\Psi_{u0}^{-1}\|_\infty}{\ell c - 1}\frac{n}{\lambda}C_{un}\right\} \left\{(L+\frac{1}{c})\|\widehat\Psi_{u0}\|_\infty\frac{\lambda\sqrt{s}}{n\bar\kappa_{u,2\tilde \mathbf{c}}}+ 9\tilde \mathbf{c} C_{un}\right\}
\end{array}
$$
\end{lemma}
\begin{lemma}[M-Estimation Sparsity]
In addition to conditions of Lemma (ref), assume that with probability $1-\Delta_n$ for all $u\in\mathcal{U}$ and $\delta \in {\mathbb{R}}^p$ we have
$$ |\{{\mathbb{E}_n}[\partial_\theta M_u(Y_u,X_u,\widehat\theta_u)-\partial_\theta M_u(Y_u,X_u,\theta_u)]\}'\delta| \leqslant L_{un}\|\sqrt{w_{u}}X_u'\delta\|_{\mathbb{P}_n,2}.$$
Let $\mathcal{M}_u=\{ m \in \mathbb{N} : m \geqslant 2 \phi_{{\rm max}}(m,u) L_u^2 \}$ with $L_u =\frac{c\|\widehat\Psi_{u0}^{-1}\|_\infty }{c\ell-1} \frac{n}{\lambda}\left\{ C_{un} + L_{un} \right\}$, then with probability $1-3\Delta_n$ we have that
$$ \widehat s_u \leqslant \min_{m\in \mathcal{M}_u} \phi_{{\rm max}}(m,u) L_u^2 \ \ \mbox{ for all } \ u\in \mathcal{U}.$$
\end{lemma}
\begin{lemma}
Let $\widetilde T_u, u\in \mathcal{U},$ be the support used for post penalized estimator ((ref)) and $\tilde s_u = \Vert \widetilde T_u\Vert_0 $ its cardinality. In addition to conditions of Lemma (ref), suppose that Assumption (ref)(c) holds also for $A_u=\{ \delta : \|\delta\|_0 \leqslant \tilde s_u + s\}$ with probability $1-\Delta_n$, $\bar q_{A_u}>2\left\{ \frac{\sqrt{\tilde s_u+s_u}\|{\mathbb{E}_n}[S_u]\|_\infty}{\sqrt{\phi_{{\rm min}}(\tilde s_u+s_u,u)}} + 3C_{un}\right\}$ and $\bar q_{A_u}>2\{{\mathbb{E}_n}[M_u(Y_u,X_u,\tilde \theta_u)] - {\mathbb{E}_n}[M_u(Y_u,X_u,\theta_u)]\}_+^{1/2}$. Then, we have uniformly over $u\in\mathcal{U}$
$$ \|\sqrt{w_{u}} X_u'(\tilde \theta_u-\theta_u)\|_{\mathbb{P}_n,2} \leqslant \{{\mathbb{E}_n}[M_u(Y_u,X_u,\tilde \theta_u)] - {\mathbb{E}_n}[M_u(Y_u,X_u,\theta_u)]\}_+^{1/2} + \frac{\sqrt{\tilde s_u+s_u}\|{\mathbb{E}_n}[S_u]\|_\infty}{\sqrt{\phi_{{\rm min}}(\tilde s_u+s_u,u)}} + 3C_{un}.$$
\end{lemma}
In Lemma (ref), if $\widetilde T_u={\rm support}(\widehat \theta_u)$, we have that
$${\mathbb{E}_n}[M_u(Y_u,X_u,\tilde \theta_u)] - {\mathbb{E}_n}[M_u(Y_u,X_u,\theta_u)] \leqslant {\mathbb{E}_n}[M_u(Y_u,X_u,\widehat \theta_u)] - {\mathbb{E}_n}[M_u(Y_u,X_u,\theta_u)] \leqslant \lambda C' \|\widehat\theta_u - \theta_u\|_1$$ and $\sup_{u\in\mathcal{U}}\|{\mathbb{E}_n}[S_u]\|_\infty \leqslant C' \lambda$ with high probability, $C' \leqslant L \sup_{u\in\mathcal{U}}\|\widehat \Psi_{u0}\|_\infty$.
These results generalize important results of the $\ell_1$-penalized estimators to the case of functional response data and estimated of nuisance functions.
A key assumption in Lemmas (ref)-(ref) is that the choice of $\lambda$ satisfies ((ref)). We next provide a set of simple generic conditions that will imply the validity of the proposed choice. These generic conditions can be verified in many applications of interest.
{\bf Condition WL.} {\it For each $u\in \mathcal{U}$, let $S_{u} = \partial_\theta M_u(Y_u,X_u,\theta_u,a_u)$, suppose that:\\
(i) ${\displaystyle\sup_{u\in\mathcal{U}} \max_{k\leqslant p} } \{{\mathrm{E}}[|S_{uk}|^3]\}^{1/3}/\{{\mathrm{E}}[|S_{uk}|^2]\}^{1/2}\Phi^{-1}(1-\xi/\{2p N_n\}) \leqslant \delta_n n^{1/6}$, for all $u\in \mathcal{U}$, $k\in[p]$;\\
(ii) $N_n \geqslant N(\epsilon,\mathcal{U},d_\mathcal{U})$, here $\epsilon$ is such that with probability $1-\Delta_n$:\\ ${\displaystyle \sup_{d_\mathcal{U}(u,u')\leqslant \epsilon} \max_{k\leqslant p}} \frac{\| {\mathbb{E}_n}[ S_{u}-S_{u'} ]\|_\infty}{{\mathrm{E}}[|S_{uk}|^2]^{1/2}}\leqslant \delta_n n^{-\frac{1}{2}}$, and
${\displaystyle \sup_{d_\mathcal{U}(u,u')\leqslant \epsilon} \max_{k\leqslant p}} \ \frac{|{\mathrm{E}}[S_{uk}^2-S_{u'k}^2]| +|({\mathbb{E}_n}-{\mathrm{E}})[S_{uk}^2]|}{{\mathrm{E}}[|S_{uk}|^2]}\leqslant \delta_n$.
}
The following technical lemma justifies the choice of penalty level $\lambda$. It is based on self-normalized moderate deviation theory.
\begin{lemma}[Choice of $\lambda$]
Suppose Condition WL holds, let $c'>c>1$ be constants, $\xi \in [1/n,1/\log n]$, and $\lambda = c'\sqrt{n}\Phi^{-1}(1-\xi/\{2pN_n\})$. Then for $n \geqslant n_0$ large enough depends only on Condition WL,
$${\mathrm{P}}\left ( \lambda/n \geqslant c \sup_{u\in\mathcal{U}}\|\widehat \Psi^{-1}_{u0}{\mathbb{E}_n}[ \partial_\theta M_u(Y_u,X_u,\theta_u,a_u) ]\|_\infty \right ) \geqslant 1-\xi - o(\xi)-\Delta_n.$$
\end{lemma}
We note that Condition WL(ii) contains high level conditions. See BCFH2013program for examples that satisfy these conditions. The following corollary summarizes these results for many applications of interest in well behaved designs.
\begin{corollary}[Rates under Simple Conditions]
Suppose that with probability $1-o(1)$ we have that $C_{un} \vee L_{un} \leqslant C\{n^{-1}s\log(pn)\}^{1/2}$, $(Lc+1)/(\ell c-1)\leqslant C$, $w_u = 1$, and Condition WL holds with $\log N_n \leqslant C\log(pn)$. Further suppose that with probability $1-o(1)$ the sparse minimal and maximal eigenvalues are well behaved, $c\leqslant \phi_{{\rm min}}(s \ell_n,u ) \leqslant \phi_{{\rm max}}(s \ell_n,u ) \leqslant C$ for some $\ell_n\to \infty$ uniformly over $u\in\mathcal{U}$. Then with probability $1-o(1)$ we have
$$ \sup_{u\in\mathcal{U}}\|X_u'(\widehat \theta_u - \theta_u)\|_{\mathbb{P}_n,2} \lesssim \sqrt{\frac{s\log(p n)}{n}}, \ \ \sup_{u\in\mathcal{U}}\|\widehat \theta_u - \theta_u\|_1 \lesssim \sqrt{\frac{s^2\log(p n)}{n}}, \ \ \mbox{and} \ \ \sup_{u\in\mathcal{U}}\|\widehat \theta_u \|_0 \lesssim s.$$
Moreover, if $\widetilde T_u={\rm support}(\widehat \theta_u)$, we have that
$$ \sup_{u\in\mathcal{U}}\|X_u'(\tilde \theta_u - \theta_u)\|_{\mathbb{P}_n,2} \lesssim \sqrt{\frac{s\log(p n)}{n}}$$
\end{corollary}
\section{Bounds on Covering entropy}
Let $ ( W_i)_{i=1}^n$ be a sequence of independent copies of a random element $ W$ taking values in a measurable space $({\mathcal{W}}, \mathcal{A}_{{\mathcal{W}}})$ according to a probability law $P$. Let $\mathcal{F}$ be a set of suitably measurable functions $f\colon {\mathcal{W}} \to \mathbb{R}$, equipped with a measurable envelope $F\colon \mathcal{W} \to \mathbb{R}$. The proofs for the following lemmas can be found in BCFH2013program.
\begin{lemma}[Algebra for Covering Entropies] \text Work with the setup above. \\
(1) Let $\mathcal{F}$ be a VC subgraph class with a finite VC index $k$ or any
other class whose entropy is bounded above by that of such a VC subgraph class, then
the uniform entropy numbers of ${\mathcal{F}}$ obey
\begin{equation*}
\sup_{Q} \log N(\epsilon \|F\|_{Q,2}, \mathcal{F}, \| \cdot \|_{Q,2}) \lesssim \{1+ k \log (1/\epsilon)\}\vee 0
\newline
\end{equation*}
(2) For any measurable classes of functions $\mathcal{F}$ and $\mathcal{F}^{\prime
} $ mapping $\mathcal{W}$ to $\Bbb{R}$,
\begin{align*}
&\log N(\epsilon \Vert F+F^{\prime }\Vert _{Q,2},\mathcal{F}+\mathcal{F}^{\prime
}, \| \cdot \|_{Q,2})\leqslant \log N\left(\mbox{$ \frac{\epsilon }{2}$}\Vert F\Vert _{Q,2},\mathcal{F}, \| \cdot \|_{Q,2}\right)
+ \log N\left( \mbox{$ \frac{\epsilon }{2}$}\Vert F^{\prime }\Vert _{Q,2},\mathcal{F}^{\prime
}, \| \cdot \|_{Q,2}\right), \\
&\log N(\epsilon \Vert F\cdot F^{\prime }\Vert _{Q,2},\mathcal{F}\cdot \mathcal{F}^{\prime
}, \| \cdot \|_{Q,2})\leqslant \log N\left( \mbox{$ \frac{\epsilon }{2}$}\Vert F\Vert _{Q,2},\mathcal{F}, \| \cdot \|_{Q,2}\right)
+ \log N\left( \mbox{$ \frac{\epsilon }{2}$}\Vert F^{\prime }\Vert _{Q,2},\mathcal{F}^{\prime
}, \| \cdot \|_{Q,2}\right), \\
& N(\epsilon \Vert F\vee F^{\prime }\Vert _{Q,2},\mathcal{F}\cup \mathcal{F}^{\prime
}, \| \cdot \|_{Q,2})\leqslant N\left(\epsilon\Vert F\Vert _{Q,2},\mathcal{F}, \| \cdot \|_{Q,2}\right)
+ N\left( \epsilon\Vert F^{\prime }\Vert _{Q,2},\mathcal{F}^{\prime
}, \| \cdot \|_{Q,2}\right).
\end{align*}
(3) For any measurable class of functions $
\mathcal{F}$ and a fixed function $f$ mapping $\mathcal{W}$ to $\Bbb{R}$,
\begin{equation*}
\log \sup_{Q} N(\epsilon \Vert |f|\cdot F\Vert _{Q,2},f\cdot\mathcal{F}, \| \cdot \|_{Q,2})\leqslant \log \sup_{Q} N\left(
\epsilon /2\Vert F\Vert _{Q,2},\mathcal{F}, \| \cdot \|_{Q,2}\right)
\end{equation*}
(4) Given measurable classes $\mathcal{F}_j$ and envelopes $F_j$, $j=1,\ldots,k$, mapping $\mathcal{W}$ to $\Bbb{R}$, a function $\phi\colon\Bbb{R}^k\to\Bbb{R}$ such that for $f_j,g_j\in\mathcal{F}_j$,
$ |\phi(f_1,\ldots,f_k) - \phi(g_1,\ldots,g_k) | \leqslant \sum_{j=1}^k L_j(x)|f_j(x)-g_j(x)|$, $L_j(x)\geqslant 0$, and fixed functions $\bar f_j \in \mathcal{F}_j$, the class of functions $\mathcal{L}=\{\phi(f_1,\ldots,f_k)-\phi(\bar f_1,\ldots,\bar f_k)\colon f_j \in\mathcal{F}_j, j=1,\ldots,k\}$ satisfies
\begin{equation*}
\log \sup_Q N\left(\epsilon\Big\|\sum_{j=1}^kL_jF_j\Big\|_{Q,2},\mathcal{L}, \| \cdot \|_{Q,2}\right)\leqslant \sum_{j=1}^k\log \sup_Q N\left(
\mbox{$\frac{\epsilon}{k}$}\|F_j\|_{Q,2},\mathcal{F}_j, \| \cdot \|_{Q,2}\right).
\end{equation*}
\end{lemma}
\begin{proof}
See Lemma L.1 in BCFH2013program.
\end{proof}
\begin{lemma}[Covering Entropy for Classes obtained as Conditional Expectations]
Let $\mathcal F$ denote a class of measurable functions $f\colon \mathcal{W}\times \mathcal{Y} \to \Bbb{R}$ with a measurable envelope $F$. For a given $f \in \mathcal{F}$, let $\bar f\colon \mathcal{W} \to \Bbb{R}$ be the function $\bar f (w) := \int f(w,y) d\mu_{w}(y)$ here $\mu_{w}$ is a regular conditional probability distribution over $y \in \mathcal{Y}$ conditional on $w\in\mathcal{W}$. Set $\bar{\mathcal{F}} = \{ \bar f \colon f \in \mathcal{F}\}$ and let $\bar F(w):=\int F(w,y) d\mu_w(y)$ be an envelope for $\bar{\mathcal{F}}$. Then, for $r, s \geqslant 1$,
$$
\log \sup_{Q} N(\epsilon \| \bar F\Vert _{Q,r}, \bar{\mathcal{F}}, \| \cdot \|_{Q,r}) \leqslant \log \sup_{\widetilde Q} N((\epsilon/4)^r \| F\Vert _{\widetilde Q,s}, \mathcal \mathcal{F} , \| \cdot \|_{\widetilde Q,s}),
$$ here $Q$ belongs to the set of finitely-discrete probability measures over $\mathcal{W}$ such that $0<\| \bar F\Vert _{Q,r}< \infty$, and $\widetilde Q$ belongs to the set of finitely-discrete probability measures over $\mathcal{W}\times \mathcal{Y}$ such that $0<\| F\Vert _{\widetilde Q,s}< \infty$. In particular, for every $\epsilon > 0$ and any $k\geqslant 1$,
$$
\log \sup_{Q} N(\epsilon, \bar{\mathcal{F}}, \| \cdot \|_{Q,k}) \leqslant \log \sup_{\widetilde Q} N(\epsilon/2, \mathcal \mathcal{F} , \| \cdot \|_{\widetilde Q,k} ).
$$
\end{lemma}
\begin{proof}
See Lemma L.2 in BCFH2013program.
\end{proof}