EconBase
← Back to paper

Generalized Linear Models with Structured Sparsity Estimators

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

72,784 characters · 10 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Generalized Linear Models With Structured Sparsity Estimators

\sloppy

abstractIn this paper, we introduce structured sparsity estimators in Generalized Linear Models. Structured sparsity estimators in the least squares loss are introduced by Stucky and van de Geer (2018) recently for fixed design and normal errors. We extend their results to debiased structured sparsity estimators with Generalized Linear Model based loss. Structured sparsity estimation means penalized loss functions with a possible sparsity structure used in the chosen norm. These include weighted group lasso, lasso and norms generated from convex cones. The significant difficulty is that it is not clear how to prove two oracle inequalities. The first one is for the initial penalized Generalized Linear Model estimator. Since it is not clear how a particular feasible-weighted nodewise regression may fit in an oracle inequality for penalized Generalized Linear Model, we need a second oracle inequality to get oracle bounds for the approximate inverse for the sample estimate of second-order partial derivative of Generalized Linear Model. Our contributions are fivefold: 1. We generalize the existing oracle inequality results in penalized Generalized Linear Models by proving the underlying conditions rather than assuming them. One of the key issues is the proof of a sample one-point margin condition and its use in an oracle inequality. 2. Our results cover even non sub-Gaussian errors and regressors. 3. We provide a feasible weighted nodewise regression proof which generalizes the results in the literature from a simple $l_1$ norm usage to norms generated from convex cones. 4. We realize that norms used in feasible nodewise regression proofs should be weaker or equal to the norms in penalized Generalized Linear Model loss. 5. We can debias the first step estimator via getting an approximate inverse of the singular-sample second order partial derivative of Generalized Linear Model loss. With this debiasing, we can get uniformly consistent estimators and asymptotically honest confidence intervals for parameters of interest. Our simulations also show good power and excellent size of the tests based on structured sparsity estimation.

Introduction

Generalized Linear Models (GLM) have been utilized in empirical work heavily both in econometrics and statistics. Recently, attention has been shifting to models when the number of parameters, $p$, exceeds the sample size, $n$. In a seminal paper, van de Geer et al. (2014) propose a debiased GLM with $l_1$ penalty. They were able to provide confidence intervals for parameters under high-level conditions. There were two significant issues that they solved in the literature with their article. First, they propose a formula for debiased GLM, and provide normal limits for the estimators of coefficients in the model. Then they also solved how to estimate for the inverse of the second-order partial derivative of GLM loss. This estimate is used in the formula for debiased GLM, and the standard plug-in estimators face the ill-posed inverse problem; hence they are not usable. They provide a non-standard solution based on a weighted version of nodewise regression, which was very difficult since the standard nodewise regression was not feasible. Recently Jankova and van de Geer (2016) generalize the debiased GLM with differentiable loss functions to possible non-differentiable GLM loss. Ning and Liu (2017) consider decorrelated M estimators with convex penalty, in a similar vein and use a different technique to debias than previous two papers cited above. They decorrelate one variable's effect on the other, and they use a Dantzig-based estimation of a specific moment function to get inference for target coefficients. Specifically, their inference centers on a low dimensional parameter, where the nuisance parameters are high dimensional. In their general theorem, there is the assumption of a consistency of moments with a specified rate. It is not also clear how this high dimensional moment estimation can have good power-size properties on inference.

Shi et al. (2019) introduce general inference in lasso type penalties in GLM. They analyze constrained partial regularization to get likelihood ratio type tests. They do not cover debiased GLM. Some simulation problems in coverage probabilities of certain parameters for debiased lasso for GLM is analyzed in Xia et al. (2020). They provide a solution when $p<n$ case.

One of the penalties that we analyze, as a sub-case of structured sparsity-based penalties, is the group lasso by Yuan and Lin (2006). Also, a weighted group lasso penalty for logistical regression is proposed by Meier et al. (2008). This last penalty is weighted by group size and combines $l_1$ and $l_2$ penalties. Lounici et al. (2011) provide oracle inequalities, for the least-squares loss an $l_1$ error bound for group lasso. The $l_1$ error bound increases with the true number of groups. Recently, Mitra and Zhang (2016) consider a debiased group lasso penalty in the least-squares loss. They use bounded regressors with subgaussian errors. They provide inference for structural parameters.

In this paper, we contribute to the literature that is cited above in several ways. One essential contribution is that GLM loss with structured sparsity estimators is amenable to inference in parameters of interest. We use a debiased GLM with penalties coming from structured sparsity-based norms. Our paper extends the least-squares loss with structured sparsity estimators as shown in Stucky and van de Geer (2018). Stucky and van de Geer (2018) use non-random covariates and normal errors, which is essential to their proof technique.

GLM case is not easy since fixed design with normal errors in the least-squares loss makes debiasing and inference much easier to construct and handle. Note that Stucky and van de Geer (2018) proof in the case of least squares loss with structured sparsity estimators do not carry over to GLM loss with structured sparsity estimators with random covariates, and non-normal also non-sub Gaussian errors. So inference on debiased GLM coefficients with structured sparsity-based norms is not trivial to handle. To overcome the difficulties, we start with extending the existing oracle inequality results for GLM loss in chapters 7 and 12 of van de Geer (2016). Theorems 7.2 and 12.2 in van de Geer (2016) exist either under strong conditions that have to be verified or a sub-case of GLM loss in a simplified design. We realize that sample versions of these strong conditions hold with probability approaching one with our proofs. However, these conditions are not easy to verify. The key to our proof is our introduction of a sample version of one-point margin condition (i.e. this is a condition that governs the loss function behavior in a neighborhood of true value of the parameters). We see that one-point margin condition introduces additional terms in an oracle inequality proof, so we change the existing oracle inequality proofs to consider these difficulties. Next, to get an approximate inverse of the sample second order partial derivative of GLM loss we introduce a feasible weighted nodewise regression with a convex cone based norm. In that sense, we extend the results on $l_1$ norm of van de Geer et al. (2014) to structured sparsity based norms. To get sharper bounds on our intermediate results, we also realize that the nodewise regression norm has to be weaker than or equal to norm of the penalized GLM loss. This sequencing of norms is a new finding for debiased estimators in high dimensions and can be helpful in other contexts.

As an output of our approach, we can test many restrictions and also have uniform-honest confidence intervals for our parameters. We also extend the previous literature to regressors with bounded moments and non-sub Gaussian errors. As a sub-case, we also consider a debiased weighted group lasso estimator in GLM loss.

There have been papers in debiasing lasso type estimators and providing confidence intervals in the recent literature. Starting with Belloni et al. (2014, 2016, 2017), Chernozhukov et al. (2018), Caner and Kock (2018, 2019), van de Geer et al. (2014) provide various ways debiasing in treatment effects, generalized linear models, least squares and GMM based models. For panel data related debiasing, we see papers by Kock (2016), Kock and Tang (2019). In terms of quantile regression, we see contributions by Chiang and Sasaki (2019).

Section 2 presents penalized general linear models with structured sparsity-based norm penalties. Section 3 provides a formula for how to debiasing in this new framework. Sections 4-5 offer a new oracle inequality for structured sparsity-based norm penalty and a feasible weighted nodewise regression technique. Section 6 provides a limit for increasing number of coefficients, and in Section 7, there is a sub-case of debiased weighted group lasso in GLM. Section 8 shows simulations that analyze test size, power, coverage in a limited exercise.

Penalized Generalized Linear Models

In this section, we introduce penalized Generalized Linear Models (GLM). Our penalty will extend $l_1$ penalty or elastic net penalty in GLM estimation. Our extension involves more general structured norms that will be tied to the sparsity properties of the parameter vector. These types of estimators are analyzed in van de Geer (2016) formally in the least-squares case and also in more detail for least squares case in Stucky and van de Geer (2018). The assumptions in these studies for the least-squares case assume fixed design, and normal errors. The case for GLM with structured sparsity-inducing norms has not been studied. The techniques in the least-squares case are not helpful in our case since we want to have random design with non-normal errors. GLM with structured sparsity-inducing norms in high dimensions will form the baseline estimator in our case, and we extend that to a debiased version where we can test restrictions and form confidence intervals. We follow van de Geer et al. (2014), where they study GLM with $l_1$ norm. Consider the regressors $X_i \in {\cal X} \in R^p$, and the outcome $y_i \in {\cal Y} \subseteq R$, for $i=1,\cdots, n$. The data is iid across $i=1,\cdots, n$. Regressor matrix $X$ is $n \times p$. The loss function is:

equation[equation omitted — 116 chars of source]

which is convex in $\beta$. The parameter space ${\cal B}$ is a convex subset of $R^p$. So the loss function can be represented either as in the left of ((ref)) or with the expression on the right of ((ref)). Define the first and second-order partial derivatives \[ \dot{\rho}_{\beta}:= \frac{\partial}{\partial \beta} \rho_{\beta} (y_i, X_i), \quad \ddot{\rho}_{\beta}:= \frac{\partial}{\partial \beta \partial \beta'} \rho_{\beta} (y_i, X_i),\] or in an alternative format

equation[equation omitted — 149 chars of source]

where $\dot{\rho}(.,.), \ddot{\rho}(.,.)$ are the partial derivative of function $\rho(.,.)$ with respect to second element.

Let

equation[equation omitted — 75 chars of source]

and

equation[equation omitted — 81 chars of source]

So $\beta_0$ is defined as the optimizer of the expected loss function. Define ${\cal B}_{local}$ as a convex subset of ${\cal B}$, which is in a local neighborhood of $\beta_0$. The local neighborhood will be defined in terms of the norm that we will use. This local set is needed since one of the main proofs in the appendix depends on the estimator to be in this local neighborhood (i.e. one point margin condition, Lemma (ref)).

Let $\Omega(.)$ be a norm on $R^p$. We specify its properties immediately below, but first define the $\Omega$ structured sparsity GLM estimator as \[ \hat{\beta}:= argmin_{\beta \in {\cal B} } \left[ \frac{1}{n} \sum_{i=1}^n \rho_{\beta} (y_i, X_i) + \lambda \Omega (\beta) \right],\] with $\lambda > 0$ as a tuning parameter. The norms that we analyze should have weak decomposability property. Weak decomposability will be a key requirement on $\Omega(.)$ and explained immediately below in Definition (ref). We use definition 6.1 of van de Geer (2016), and this is also defining an allowed set $S$. To that effect, divide the set $J=\{1,2,\cdots, p\}$ into $S$ and its mutually exclusive complement $S^c$. In other words $J = S \cup S^c$. Let $|S|$ represent the cardinality of the index set $S$. Now define another norm $\Omega^{S^c}(.)$ on $R^{p - |S|}$. Also define $\beta_S$ as a vector with entries equal to zero for elements with indices when $j \notin S$. Also define $\beta_{S^c}$ as the vector with all elements with indices, $j$ inside the set $S$, set to zero, all elements with indices belonging to $S^c$ are kept, $S^c:=\{ j \in \{1,2,\cdots, p\}: j \notin S\}$.

definitiona(Definition 6.1, van de Geer (2016). Fix some set $S$. We say that norm $\Omega$ is weakly decomposable for the set S if there exists a norm $\Omega^{S^c}$ on $R^{p- |S|}$ such that for all $\beta \in R^p$ \[ \Omega (\beta) \ge \Omega (\beta_S) + \Omega^{S^c} (\beta_{S^c}).\]
definitiona(Definition 6.1, van de Geer (2016)). We say that $S$ is an allowed set if $\Omega$ is weakly decomposable for the set S.

To give an example: for $l_1$ norm any subset $S$ of $J$ is an allowed set, and $\Omega^{S^c}(.)$ is again the $l_1$ norm. So $\|\beta\|_1= \| \beta_{S} \|_1 + \| \beta_{S^c} \|_1$. We will also give examples of weakly decomposable norms in this section. To give a broad example, all norms generated from convex cones are weakly decomposable; see section 6.9 of van de Geer (2016). Some of the specific examples of norms generated from convex cones are weighted group lasso norm, lasso, wedge norm, and concavity inducing norms. We give two examples of such norms.

Example 1. The first one is a weighted group lasso norm. The variables are grouped disjointly, and the penalty is designed accordingly. Let $\{G_j\}_{j=1}^m$ be a partition of $\{1, \cdots, p\}$ into disjoint $m$ groups. For a parameter vector $\beta \in R^p$, the weighted group lasso norm is:

\[ \| \beta \|_{wgl}:= \sum_{j=1}^m \sqrt{|G_j|} \| \beta_{G_j} \|_2.\] So $\beta$ vector is grouped into m disjoint groups, and size of the group $G_j$ is: $| G_j|$. Any union of groups can be an allowed set in weighted group norm.

Example 2. Another example is the wedge norm in section 6.9 of van de Geer (2016). Consider the convex cone, ${\cal A}:= \{ a_1 \ge a_2 \ge a_3 ..\cdots a_p > 0\}$, with $\| \beta \|_W:= \min_{a_j \in {\cal A}} \frac{1}{2} \sum_{j=1}^p \left( \frac{\beta_j^2}{a_j} + a_j \right).$ An allowed set is the first $s$ elements in $\beta$.

For any norm, $\Omega$, not necessarily weakly decomposable we know that by triangle inequality \[ \Omega (\beta) \le \Omega (\beta_S) + \Omega (\beta^{S^c}),\] so clearly for weakly decomposable $\Omega(.)$, we have $\Omega (\beta^{S^c}) \ge \Omega^{S^c} (\beta_{S^c})$. By Chapter 6 of van de Geer (2016) dual norm of $\Omega(.)$ is defined as \[ \Omega_* (w):= \max_{ \Omega (\beta) \le 1} |w'\beta|, \quad w \in R^p.\]

We need few more concepts regarding norms. This is taken from Section 6.4 of van de Geer (2016).

definitiona(Stronger norm). If $\Omega(.)$ and $\underline{\Omega}(.)$ are any two norms on $R^p$, and if we have \[ \Omega (\beta) \ge \underline{\Omega} (\beta), \quad \forall \beta \in R^p,\] we say that $\Omega$ is a stronger norm than $\underline{\Omega}$.

We also see that

equation[equation omitted — 153 chars of source]

where $\underline{\Omega}_*$ is the dual norm of $\underline{\Omega}$. Stronger norm definition is applicable to all norms regardless of their weak decomposability or not. As in section 6.4 of van de Geer (2016) we define the following lower bound norm for $\Omega(.)$. Formally define $\Omega^{S^c}(.)$, which is mentioned in Definition 1, as the largest norm among the norms $\underline{\Omega}^{S^c}(.)$ for which \[ \Omega(\beta) \ge \Omega(\beta_S) + \underline{\Omega}^{S^c} (\beta_{S^c}),\] hence define

equation[equation omitted — 124 chars of source]

We define $S_0$ as the indices of the active set. This is defined with respect to a particular norm $\Omega(.)$. $S_0$ should be an allowed set and carry all the indices with nonzero elements in the model. To clarify the last statement, to give an example, these elements can be indices of individual non-zero true coefficients in lasso via $l_1$ norm, $S_{0}:= \{ j: |\beta_{j0}| \neq 0 \}$, where $\beta_{j0}$ represents true value of $j$ th coefficient where $j=1,\cdots, p$, where $p$ is the total number of coefficients. For the weighted group lasso norm, these indices with nonzero elements are the indices of the active (non-zero) groups, so $S_0:=\{ j: \| \beta_{0,G_{j}} \|_2 \neq 0 \}$, where $\beta_{0,G_{j}}$ represents the true coefficients of $j$ th group, where $j=1,\cdots, m$, where $m$ is the total number of groups in the model. We define the sparsity as $s_0$, which is the cardinality of $S_0$, $s_0:= |S_0|$. Let $l_0$ ball ${\cal B}_{l_0} (s_0) := \{ \| \beta_0 \|_{l_0} \le s_0 \}$. Define the effective sparsity condition, or sometimes called $\Omega$ effective sparsity as follows.

definitionaEffective sparsity. (Definition 4.3 of van de Geer (2014)). Suppose $S$ is an allowed set. Let $L>0$ be some constant. The effective sparsity is \[ \Gamma^2 (L,S) := [min \{ E \| X \beta_S - X \beta_{S^c} \|_n^2 : \Omega (\beta_S) = 1, \Omega^{S^c} (\beta_{S^c}) \le L\}.]^{-1}.\] This is the inverse of the more familiar $\Omega$-eigenvalue condition.

This effective sparsity is defined as a population condition, compared to the sample version of van de Geer (2014), but Definition 7.5 of van de Geer (2016) has a general population version. The sample version of effective sparsity is defined in Appendix, and also a variant of this population effective sparsity is given in Appendix.

Debiased GLM Structured Sparsity Estimator

In this section we introduce a debiased version of GLM structured sparsity estimator. But first, we define by using differentiability of the objective function with ((ref)), \[ \hat{\Sigma}_{\hat{\beta}}:= \frac{1}{n} \sum_{i=1}^n \ddot{\rho}_{\hat{\beta}} (y_i, X_i)= \frac{1}{n} \sum_{i=1}^n X_{\hat{\beta},i} X_{\hat{\beta},i}',\] where $X_{\hat{\beta},i}:= X_i w_{\hat{\beta},i}$, in which

equation[equation omitted — 93 chars of source]

by equation ((ref)). Also see that $w_{\beta_0,i}$ is defined in the same way as in ((ref)), second order partial derivative depends on $\beta_0$ there. This sample moment estimator plays a crucial role in our derivations. In the case of least-squares loss, this corresponds to the empirical Gram matrix. Note that in our GLM loss case when $p>n$, $\hat{\Sigma}_{\hat{\beta}}$ is singular.

Furthermore define a $p \times p$ matrix $\hat{\Theta}$ which will be defined later as output of a nodewise regression. $\hat{\Theta}$ will be used as an approximate inverse for $\hat{\Sigma}_{\hat{\beta}}$. Section 5 considers the form and theory behind $\hat{\Theta}$. Our debiased estimator is:

eqnarray[eqnarray omitted — 262 chars of source]

where $\dot{\rho}_{\hat{\beta}} (y_i, X_i)$ is the partial derivative of our GLM loss function with respect to $\beta$ and evaluated at $\hat{\beta}$, and we use ((ref)) for the last equivalent definition.

We extend this debiased estimator to structured sparsity penalties. A slightly different formula is given in the previous literature, for the least-squares loss with structured sparsity penalty. The previous literature uses nuclear norm regularized multi-nodewise regression in Definitions 3-4 of Stucky and van de Geer (2018). We realized that if the design is fixed and with normal errors in the least-squares context, their nuclear-norm-based debiased estimator is easy to come up with limits. That structure is not amenable in GLM, with random design.

For testing in high dimensions, define a $p \times 1$ vector $\alpha$ such that $\| \alpha \|_2 =1$, and let ${\cal H}:=\{ j=1,\cdots, p: \alpha_j \neq 0\}$ with cardinality $|{\cal H}| = h$. $h$ will increase with sample size and we will precisely define this rate through our assumptions, and $h$ will be the number of restrictions that are tested and $h < p$. Clearly

equation[equation omitted — 75 chars of source]

since $\| \alpha \|_2 =1$, and ${\cal H}$ definition with using the norm inequality that puts an upper bound on $l_1$ norm in terms of $l_2$ norm. In the remaining sections we consider the following as the numerator of our test statistic:

equation[equation omitted — 217 chars of source]

The denominator of our test statistic will be

equation[equation omitted — 184 chars of source]

Assumptions

We provide the main assumptions used throughout the paper, and needed for oracle inequality in Theorem (ref). Another set of assumptions will be provided in their sections related to nodewise regression and limit theorem.

assumptionAThe data $X_i,y_i$ are iid across $i=1,\cdots,n$. Furthermore $\max_{1 \le j \le p} E | X_{1j}|^{r_x} \le C < \infty$, where $r_x \ge 4$ and $C>0$ is a positive constant. Also the effective sparsity is bounded away from infinity: $0< c \le \Gamma^2(2, S_0) \le C < \infty$ with $c>0$ which is a positive constant.
assumptionA(i). Define \[M_1:= \max_{1 \le i \le n} \max_{1 \le j \le p} \max_{1 \le l \le p} | X_{ij} X_{il} - E X_{ij} X_{il}|,\] and \[ M_2:= \max_{1 \le i \le n} \max_{1 \le j \le p} | X_{ij}|,\] Then we assume \[ \max \left( \frac{\sqrt{E M_1^2} \sqrt{lnp}}{\sqrt{n}}, \frac{\sqrt{E M_2^2} \sqrt{lnp}}{\sqrt{n}}\right) = O(1).\] (ii). \[ s_0 \sqrt{lnp/n} \to 0.\]
assumptionAThere exists a positive constant $C_{\rho}$, which depends on the shape of the second order partial derivative $\ddot{\rho}(.)$, and $\kappa$ all positive constants such that \[\ddot{\rho} (y_i, X_i' \beta)\ge 1 /C_{\rho}^2,\] for all $ \sup_{\beta_0 \in {\cal B}_{l_0} (s_0)} | X_i' (\beta - \beta_0)| \le \kappa.$
assumptionAThe derivatives $\dot{\rho} (y,a) = \frac{\partial \rho(y,a)}{ \partial a }$, $\ddot{\rho}(y,a) = \frac{\partial \rho(y,a)}{ \partial a^2 }$ exist for all $y,a$ and for some $\delta$ neighborhood of $X_i' \beta_0$, $\delta>0$ (i). \[ \sup_{\beta_0 \in {\cal B}_{l_0} (s_0)} \max_{a_0 \in \{X_i' \beta_0 \}} \sup_{y_i} | \dot{\rho} (y_i,a_0)| = O(1).\] (ii). \[ \sup_{\beta_0 \in {\cal B}_{l_0} (s_0)} \max_{a_0 \in \{X_i' \beta_0 \}} \sup_{ | a - a_0 | \le \delta} \sup_{y_i} | \ddot{\rho}(y_i,a) | = O(1).\] (iii). Also $\ddot{\rho}(y,a)$ is Lipschitz \[ \sup_{\beta_0 \in {\cal B}_{l_0} (s_0)} \max_{ a_0 \in \{ x_i' \beta_0\} } \sup_{ | a - a_0| \cup | \hat{a} - a_0| \le \delta} \sup_{y_i} \frac{ | \ddot{\rho} (y_i, a) - \ddot{\rho} (y_i, \hat{a}) | }{|a- \hat{a} | } \le 1.\]

We discuss the assumptions here. Assumption (ref) is standard, and effective sparsity is used to control certain matrix's singularity. Our proofs and remarks after Theorem (ref) will show how it is related to compatibility condition, which is more familiar in high dimensional statistics. Assumption (ref) is needed for concentration inequalities that we use, and these inequalities are from Chernozhukov et al. (2017). Assumption (ref) provides a lower bound on the second-order partial derivative of the loss function and is needed to control the one-point margin condition, which will be explained in the Appendix. This condition is also used in Chapter 12 of van de Geer (2016). It is possible to relax this condition, but this will lengthen the proofs immensely, so we avoided that. Assumption (ref) puts structure on GLM loss, and this is used as Assumption C.1 in van de Geer et al. (2014). We strengthened this to uniform over ball ${\cal B}_{l_0} (s_0)$. We also think that it is possible to get rid of bounded first-order partial derivative at $\beta_0$, and bounded second-order partial derivative bound in a uniform neighborhood of $\beta_0$ by using Assumption (ref) type of moments of these derivatives.

$\underline{\Omega}$ bound

One of the crucial elements in the paper is the $\underline{\Omega}$ bound for our estimator. This bound will be used in de-biasing, and the former literature takes this type of result given or allows fixed-random regressors in restrictive data setups. Theorems 9.19 and Corollary 9.20 of Wainwright (2019) provide oracle bounds under very restrictive conditions on the data and eigenvalue type conditions. Another paper related to our paper is the logistic result in $l_1$ norm in Theorem 12.2 of van de Geer (2016), which provides sample eigenvalue conditions with restrictive conditions on the data set. Compared to these results before us, we provide a different proof based on primitive assumptions in a general norm-setting. Even though the estimator is obtained by using $\Omega$ bound, and we are interested in $\underline{\Omega}$ bound, which is a weaker norm than $\Omega$ ($\underline{\Omega} \le \Omega$). Also, some of the key difficulties in obtaining such a result are that an empirical process result has to be established, sample one-point margin condition has to be proved, and since this is new, it has to be shown that sample one point margin condition does not impediment the oracle inequality proof. Details are in the Appendix.

We define a positive sequence $t_1$, which is defined in ((ref)), and $t_1= O(\sqrt{lnp/n})$. As in p.107 of van de Geer (2016) we take $\beta_{S_0}$ as "relevant coefficients" in $\beta_0$, and treat $\beta_{S_0^c}$ as "irrelevant smallish-like" part of $\beta_0$. This can be thought of nonzero coefficients as $\beta_{S_0}$, and local-to-zero and zero coefficients in $\beta_{S_0^c}$. Specifically, we formalize a condition in Theorem (ref)(ii) below for $\beta_{S_0^c}$ in terms of the weakly-decomposable norm that we use. A form of weak-sparsity will be imposed for asymptotic results. Define $l_0$ ball ${\cal B}_{l_0} (s_0):= \{ \| \beta_0 \|_{l_0} \le s_0 \}$.

theorema(i). Under Assumptions (ref)-(ref), with sufficiently large $n$ \[ \underline{\Omega} (\hat{\beta} - \beta_0) \le (18 \lambda) C_{\rho}^2 \Gamma^2 (2,S_0) + 32 \Omega(\beta_{S_0^c}),\] with probability at least $1- \frac{3}{p^{2c}} - \frac{1}{p^c} - \frac{7 C }{4 (lnp)^2}= 1- o(1)$. (ii). Also our Remark 2 below will show that, with assuming $\sup_{\beta_0 \in {\cal B }_{l_0} (s_0)} \Omega (\beta_{S_0^c}) \to 0$, then \[ \sup_{\beta_0 \in {\cal B}_{l_0} (s_0)} \underline{\Omega} ( \hat{\beta} - \beta_0) = O_p ( s_0 \sqrt{\frac{lnp}{n}}) = o_p (1).\]

Remarks. 1. First, we want to rewrite the upper bound in terms of the population version of the compatibility constant, where the literature is familiar with. Define the compatibility constant as in Definition 6.2 of van de Geer (2016) as \[ \phi^2 (L,S) := min\{ |S| E [\| X \beta_S - X \beta_{S_c} \|_n^2]: \Omega(\beta_S) =1, \Omega^{S^c} (\beta_{S^c}) \le L\},\] where $S, S^c$ are any allowed set and its complement respectively. Next by Definition (ref) and the above expression

equation[equation omitted — 68 chars of source]

Also, the empirical version of the equality in ((ref)) is on p.81 of van de Geer (2016), just before section 6.6 there. Using ((ref)) we can write the upper bound in terms of sparsity of the coefficients explicitly. In that respect by at $S= S_0$, with $L=2$ \[ \Gamma^2 (2, S_0) = \frac{|S_0|}{\phi^2 (2, S_0)}.\]

2. First, impose the weak-sparsity assumption, uniformly over ${\cal B}_{l_0} (s_0)$, we impose $\Omega (\beta_{S_0^c}) \to 0$. Next to get an asymptotic sense from our bound, $\lambda_e$ is a positive sequence defined in ((ref)), we can have $\lambda = 16 \lambda_e = O (\sqrt{lnp/n})$ as shown in Lemma (ref) in Appendix, and in Lemma (ref) of Appendix we also have $t_1= O (\sqrt{lnp/n})$, then the upper bound

equation[equation omitted — 109 chars of source]

when we replace $\Gamma^2 (2, S_0) \le C < \infty$, with $\phi^2 (2, S_0) \ge c > 0$ in Assumption (ref) and since $\lambda |S_0| = \lambda s_0 = o(1)$ by Assumption (ref).

3. Even though the estimator optimizes over $\Omega$ norm, the bound is in weaker $\underline{\Omega}$ norm, which is needed for the debiased estimator.

These results imply

equation[equation omitted — 152 chars of source]

4. One issue is the cost of the generality of the results. An alternative technique could have used a different approach and it may have been possible to get a better rate than in ((ref)). Our proof technique, on the other hand, is very general and uses the ranking of norms in Lemma (ref).

Nodewise Regression in Structured Sparsity Estimators

We start with definitions of several matrices used in nodewise regression with norm $\Omega(.)$. So we generalize the results in van de Geer et al. (2014) from $l_1$ norm to a more general norm structure designated by $\Omega(.)$. Next, we show that how nodewise regression be carried, and last we show that nodewise regression provides an approximate inverse of singular sample moment matrix, $\hat{\Sigma}_{\hat{\beta}}$ which is defined in section (ref) in GLM structure.

We extend the definitions in section (ref). Define $X_{\hat{\beta}}:= W_{\hat{\beta}}X$, where $W_{\hat{\beta}}:= diag (w_{\hat{\beta},1},\cdots , w_{\hat{\beta},n})'$ which is a $n \times n$ diagonal matrix. Note that $w_{\hat{\beta},i}:= \sqrt{\ddot{\rho} (y_i, X_i' \hat{\beta})}$ for $i=1,\cdots, n$. See that $j$ th column of $X_{\hat{\beta}}$ is denoted as $X_{\hat{\beta},j}: n \times 1$, and $X_{\hat{\beta}, -j}:n \times p-1$ is defined as all columns of $X_{\hat{\beta}}$ except $j$ th one. We define $ W_{\beta_0}:= diag (w_{\beta_0,1},\cdots, w_{\beta_0,i}, \cdots, w_{\beta_0,n})$ which is $n \times n$ diagonal matrix, with $w_{\beta_0,i}:= \sqrt{\ddot{\rho} (y_i, X_i' \beta_0)}$, for $i=1,\cdots,n$. Define $n \times p$ matrix: $X_{\beta_0} := W_{\beta_0} X$, and $j$ th column of that matrix as $X_{\beta_0,j}$, and all the columns except j th one as: $X_{\beta_0,-j}$, Define $\Theta:= \Sigma_{\beta_0}^{-1}$, where $\Sigma_{\beta_0}:= E X_{\beta_0, i} X_{\beta_0,i}'$, where $X_{\beta_0,i}'$ is the $i$ th row of $n \times p$ matrix $X_{\beta_0}$. So $X_{\beta_0,i}':= X_i' w_{\beta_0,i}$. $X_{\beta_0, i}$ is the column version of the row $X_{\beta_0, i}'$. Define $\gamma_{\beta_0,j}$ as $\gamma_j$ that minimizes $E [ X_{\beta_0,j} - X_{\beta_0, -j} \gamma_j]^2$.

We can write the following from p.3 of the supplement of van de Geer et al. (2014)

equation[equation omitted — 99 chars of source]

where

equation[equation omitted — 68 chars of source]

By the analysis in p.157 of Caner and Kock (2018) and ((ref)) we get the relation between $\Theta$ and regression coefficient $\gamma_{\beta_0,j}$, and the scalar $\tau_j^2$. Note that $\tau_{j}^2:= \frac{1}{\Theta_{j,j}}$, where $\Theta_{j,j}$ is the $j$ th main diagonal element of $\Theta$. With that analysis we get $\Theta= C_j/ \tau_{j}^2$, where $C_j$ is a $ p \times 1$ vector, with 1 in $j$ th cell, and the rest of $C_j$ is defined as $-\gamma_{\beta_0,j}$, $j=1,\cdots, p$, so at $j=1$ for example, $C_1:= (1, - \gamma_{\beta_0,1}')'$.

Hence, as shown in the proof of Theorem 3.2 in van de Geer et al. (2014), premultiply ((ref)) by $W_{\hat{\beta}} W_{\beta_0}^{-1}$ to have

equation[equation omitted — 137 chars of source]

We have the definition, for $j=1,\cdots,p$

equation[equation omitted — 190 chars of source]

where we will impose $\lambda_{nw}=\lambda_j$ for each $j=1,\cdots,p$, and $\lambda_{nw}$ is a positive sequence and its rate will be determined in the proofs. Define the nodewise regression estimates in the same form as in van de Geer et al. (2014). Define $\hat{\Theta}_j:= \hat{C}_j/\hat{\tau}_j^2$, with $\hat{C}_j$ defined as a vector with $1$ in $j$ th cell, and all other $p-1$ cells are $-\hat{\gamma}_{\hat{\beta},j}$ vector from the nodewise regression, and $\hat{\tau}_j^2:=\frac{X_{\hat{\beta},j}' (X_{\hat{\beta},j} - X_{\hat{\beta},-j} \hat{\gamma}_{\hat{\beta},j})}{n}$.

One word of caution is that recently van de Geer (2016) and Stucky and van de Geer (2018) use nuclear norm loss with a sum of norms over the restrictions tested instead of nodewise regression. This type of analysis works well due to the fixed design nature of regressors and normal errors via the least-squares loss in the main structural parameter estimation. The technique did not carry out to our more general random design with non-normal errors and generalized linear model.

One of the key issues is the penalty in the nodewise regression. We propose $\underline{\Omega} (.)$ norm instead of $\Omega (.)$ norm. The main reason is that dual norms have the following inequality $\underline{\Omega}_{*} \ge \Omega_{*}$ when $\underline{\Omega} \le \Omega$ by definitions of these norms. The proofs use dual norm inequality, and due to Theorem (ref) result, they use $\underline{\Omega}$ norm bounds. If we had operated with $\Omega$ based bounds in nodewise regression we had to convert them still to $\underline{\Omega}$ results which can be done via large upper bounds as shown in Lemma 3 of Stucky and van de Geer (2018) since this results in a larger bound, which the bounds depend on sparsity, so usage of $\Omega$ is not advised. In that sense, our proposal is new and will result in better- smaller bounds and faster convergence rates of the error to zero in certain proofs regarding central limit theorem type result. Specifically, this can be seen in Step 2 of proof of Theorem (ref), the equation before ((ref)). In summary, we provide a new approach to debiasing. If the main penalty in the loss function of the interest is $\Omega$ as in section 2, then the nodewise regression has to run with a weaker norm: $\underline{\Omega}: \Omega \ge \underline{\Omega}$. A similar approach is suggested by Stucky and van de Geer (2018) by using gauge functions as norms (weakest possible decomposable norm) in forming precision matrix estimate, with fixed design and the least squares loss. Their setup is different and does not overlap with us, since we only analyze norms generated from cones in our main equation, and use their properties to our advantage in the proofs such as Lemma (ref).

Now we form an inequality that will help us in the proofs of debiased GLM with structured sparsity. Start with $\hat{\tau}_j^2$ definition and divide each side by $\hat{\tau}_j^2$, and also using $X_{\hat{\beta},j}, X_{\hat{\beta},-j}$ definitions in $X_{\hat{\beta}}$

equation[equation omitted — 271 chars of source]

where we use the definition $\hat{\Theta}_j:= \hat{C}_j/\hat{\tau}_j^2 $, with $(X_{\hat{\beta},j} - X_{\hat{\beta},-j} \hat{\gamma}_{\hat{\beta},j}) = X_{\hat{\beta}} \hat{C}_j$ by using $X_{\hat{\beta}}, \hat{C}_j$ definitions. Denoting $\hat{z}$ as the sub-differential, and getting the KKT conditions from ((ref)) \[ \hat{z}_j = \frac{X_{\hat{\beta},-j}' (X_{\hat{\beta},j}- X_{\hat{\beta},-j} \hat{\gamma}_{\hat{\beta},j})}{ n \lambda_{nw}}.\] and $\underline{\Omega}_* (\hat{z}_j) \le 1$ for $j=1,\cdots,p$, $\underline{\Omega}_*(.)$ is the dual norm for $\underline{\Omega}(.)$. From KKT conditions \[ \frac{ \underline{\Omega}_* \left( X_{\hat{\beta},-j}' (X_{\hat{\beta},j}- X_{\hat{\beta},-j} \hat{\gamma}_{\hat{\beta},j})\right)}{n } \le \lambda_{nw},\] which implies by dividing each side by $\hat{\tau}_j^2$, and using $(X_{\hat{\beta},j}- X_{\hat{\beta},-j} \hat{\gamma}_{\hat{\beta},j}) = X_{\hat{\beta}} \hat{C}_j$, and by $\hat{\Theta}_j$ above

equation[equation omitted — 156 chars of source]

Combine ((ref))((ref)), and $\hat{\Sigma}_{\hat{\beta}}:= X_{\hat{\beta}}' X_{\hat{\beta}}/n$, for $j=1,\cdots, p$

equation[equation omitted — 138 chars of source]

So we show that in dual norm $\hat{\Sigma}_{\hat{\beta}}$ has an approximate inverse $\hat{\Theta}$. Equation ((ref)) is a new result and can be used in other contexts. Next we put forward our assumptions for this nodewise regression result. Before the next assumption we define the following terms. Let the number of restrictions in $\beta_0$ vector as $h$, and with indices of ${\cal H}:= \{ j: H_0: \beta_j = \beta_{0j} \}$. We can let $h$ grow with $n$, but $h < p$. Let $X_{\beta_0,-j,i,k}$ represent $X_{\beta_0,-j}$ ($n \times (p-1$)) matrix $(i,k)$ th element, and $\eta_{\beta_0, j,i}$ is the ($n \times 1$) $\eta_{\beta_0,j}$ vector's $i$ th element

\[ M_3:= \sup_{\beta_0 \in {\cal B}_{l_0} (s_0)} \max_{1 \le i \le n} \max_{j \in {\cal H}} \max_{1 \le k \le p-1} | X_{\beta_0,-j, i,k} \eta_{\beta_0,j,i} |.\]

\[ M_4:= \sup_{\beta_0 \in {\cal B}_{l_0} (s_0)} \max_{1 \le i \le n} \max_{j \in {\cal H}} | \eta_{\beta_0, j,i}^2 - E \eta_{\beta_0, j,i}^2|.\]

Since $h < p$, so $ln ph < ln p^2 = 2 ln p$.

We index $\gamma_{\beta_0,j}$ into $\gamma_{\beta_0, S_j}$, where $S_j$ represents all the indices with nonzero components in $\gamma_{\beta_0,j}$, and $\gamma_{\beta_0, S_j^c}$ where $S_j^c$ represents the indices with all the local to zero, and zero coefficients. We also provide a compatibility condition, for $j=1,\cdots, p$,

equation[equation omitted — 237 chars of source]
assumptionA(i). $inf_{\beta_0 \in {\cal B}_{l_0} (s_0)} Eigmin (\Sigma_{\beta_0}) \ge c > 0$. Also $\sup_{\beta_0 \in {\cal B}_{l_0} (s_0)}\max_{1 \le i \le n} \max_{1 \le j \le p} E | \eta_{\beta_0,i,j}|^r \le C < \infty$, for $r > 8$. (ii). \[ \frac{\sqrt{E M_3^2} \sqrt{lnp}}{\sqrt{n}} = O (1).\] \[ \frac{\sqrt{E M_4^2} \sqrt{ln h}}{\sqrt{n}} = O (1).\]

Define $\bar{s}:= \max_{1 \le j \le p} |S_j|$. Define two positive sequences, $H_n:= O (h^{2/r} n^{2/r})$, and $K_n := O (p^{2/r_x} n^{2/r_x})$. Define a known sequence $g_n $ which depends on the norm that is analyzed and the sample size, and $g_n$ is a nondecreasing function in $n$. To give an example, if $l_1$ norm is used for $\underline{\Omega}(.)$ then $g_n=\bar{s}^{1/2}$ by (B.55) of Caner and Kock (2018). Formally, $g_n$ is defined in Assumption (ref)(iii).

assumptionA(i). \[ g_n\bar{s}^{1/2} \sqrt{\frac{lnp}{n}} (\max(\bar{s}, H_n^2 s_0^2) = o(1).\] (ii).\[ K_n \bar{s} s_0 \sqrt{\frac{lnp}{n}} = o(1).\] (iii). \[ \sup_{\beta_0 \in {\cal B}_{l_0} (s_0)} \max_{1 \le j \le p} \underline{\Omega} (\gamma_{\beta_0,j}) = O ( g_n) .\]

Assumptions (ref), (ref) only relate to nodewise regression. Assumption (ref) uses an eigenvalue condition and implies compatibility condition ((ref)) via Lemma 4.1 of van de Geer (2014), and it is different in form and elements from the effective sparsity condition in Definition 4. This difference stems from the nature of nodewise regression, which is described above and our oracle inequality proof in Lemma 1 below. Our Assumption (ref) is a strengthened version of eigenvalue assumption for nodewise regression in van de Geer et al. (2014) due to uniformity in ${\cal B}_{l_0} (s_0)$ in our case. Cross product of moments rate assumption can be relaxed at the expense of lengthening the proofs via marginal moment conditions. Assumption (ref) is a sparsity type assumption that also replaces Assumption (ref)(ii).

We take a specific example to show that Assumption (ref)(i)-(ii) is holding, without some of the constants to simplify the issue. Let $\bar{s}= ln n, s_0 = ln n, p=2n, H_n = n^{2/r}, K_n = (2 ln n) ^{2/r_x} n^{4/r_x}$. Then with $l_1$ norm,$g_n = \sqrt{\bar{s}}$, and \[\bar{s} \sqrt{ln p/n} \max (\bar{s}, H_n^2 s_0^2)= O (\frac{(ln n)^{3/2}}{n^{1/2}} max(ln n, n^{4/r} (ln n)^2)) =O \left( (ln n)^{7/2} n^{4/r - 1/2}\right) = o(1) \] with $r > 8$, and $K_n \bar{s} s_0 \sqrt{lnp/n} = O ( (ln n)^{\frac{5}{2} + \frac{2}{r_x}} n^{4/r_x - 1/2}) = o(1)$ with $r_x > 8$. Assumption (ref)(iii), and $g_n$ can be shown in other contexts than $l_1$ norm, for example in weighted group lasso norm, $g_n = m \sqrt{| G|}$, where $m$ is the number of groups, and $|G|$ is the largest group size.

The following lemma is an essential result and shows the estimation of the rows of the precision matrix with nodewise regression, and the estimators are consistent. Define

equation[equation omitted — 103 chars of source]
lemmaUnder Assumptions (ref) with $r_x > 8$, (ref)(i),(ref)-(ref), with the following weak sparsity condition ((ref)), uniformly over ${\cal B}_{l_0} (s_0)$ \begin{equation} \max_{j \in {\cal H}} \Omega^{S^c} (\gamma_{\beta_0, S_j^c})=o(d_n)=o(1), \end{equation} then \[ \max_{j \in {\cal H}} \underline{\Omega} (\hat{\Theta}_j - \Theta_j) = O_p \left( g_n \sqrt{\bar{s}} \sqrt{\frac{lnp}{n}} \max(\bar{s}, H_n^2 s_0^2)\right) =o_p (1).\] This result is also valid uniformly over $l_0$ ball ${\cal B}_{l_0} (s_0)$.

Remarks. 1. This is a new lemma in the literature and establishes general norm bounds on nodewise regression estimates for GLM based estimators. In this sense, this provides a general result for estimating the inverse of the second-order partial derivative of GLM objective function. The usage of nodewise regression is necessitated by the singularity of the sample second-order partial derivative of GLM objective function. The closest to this result is in Theorem 3.2 of van de Geer et al. (2014) with $l_1$ norm bounds in GLM. Our proof also extends Theorem 6.1 of van de Geer (2016) proof for linear loss, with high-level conditions to generalized linear models with primitive conditions, and for a weaker norm, $\underline{\Omega}$.

2. The limit on van de Geer et al. (2014) depends on strong assumptions such as uniformly bounded regressors, uniformly bounded product of nodewise regression coefficient with regressor, and the knowledge of oracle inequalities in GLM in prediction norm as well as $l_1$ norm. Our result generalizes their results to regressors with moment bounds, and also there is no need for the product of regressors and the nodewise coefficient to be uniformly bounded. Also we obtain oracle inequalities in our Theorem (ref).

3. The cost to a more general proof will be slightly different rates compared to $l_1$ norm. We have a different proof technique than van de Geer et al. (2014), and benefiting from a maximal inequality that is due to Chernozhukov et al. (2017). In case of $l_1$ norm Theorem 3.2 in van de Geer et al. (2014) under the strong assumptions provide a rate of $\max(K \sqrt{\bar{s} lnp/n}, K^4 s_0 \sqrt{lnp/n})$, where they need a $\lambda = O ( K \sqrt{lnp/n})$, where $K$ is the rate for the uniformly bounded regressors in their case, i.e. $\max_{i,j} | X_{i,j}| = O( K)$, which is their Assumption D.1. In our Lemma (ref) above, our rate is $\max( \bar{s}^2 \sqrt{lnp/n}, \bar{s} s_0^2 H_n^2 \sqrt{lnp/n})$, since in $l_1$ case $g(\bar{s}) = \bar{s}^{1/2}$ as can be shown via analysis in p.159 of Caner and Kock (2018). In $l_1$ case it is not clear which proof technique will provide a sharper rate, since assumptions are different, and our proof is geared toward a general norm result, the bounds/proofs are different, hence not resulting in the same rate for $l_1$ in both cases.

4. Also, an interesting point is that whether a different weaker norm can also be useful in this lemma. In other words if we have $\bar{\Omega}$ such that $l_1 \le \bar{\Omega} \le \underline{\Omega}$. Proof of this lemma clarifies that such a $\bar{\Omega}$ proof will go through. Essentially, a very good choice can be $l_1$ norm which provides a sharper bound, unless there is a specially structured sparsity for nodewise regression. This norm choice also can be seen by Assumption (ref)(iii).

Limit

In this section we provide a limit result. But before that, for variance-covariance estimation we need the following Assumption which is a stricter version of Assumption (ref)(i). Define \[ M_5:= \max_{1 \le i \le n} \max_{1 \le k \le p } \max_{1 \le l \le p} | X_{ik}^2 X_{il}^2 - E X_{ik}^2 X_{il}^2|.\]

assumptionA(i). \[ \frac{\sqrt{E M_5^2} \sqrt{lnp}}{\sqrt{n}} = O(1).\] (ii). Set $r_x>8$ in Assumption (ref), and let $a \wedge b = min (a,b)$ \[ \frac{(h\bar{s})^{(r_x/4)+1} \wedge (h \bar{s})^{r_x/4} p }{n^{(r_x/4)-1}} = o(1).\] (iii). \[ h g_n \bar{s} \frac{lnp}{\sqrt{n}} \max(\bar{s}, H_n^2 s_0^2) = o(1).\] (iv). \[ (h \bar{s})^{1/2} K_n s_0^2 \frac{lnp}{n^{1/2}} = o(1).\]
assumptionA\[\inf_{\beta_0 \in {\cal B}_{l_0} (s_0)} Eigmin (E X_i X_i' \dot{\rho} (y_i, X_i' \beta_0)^2 ) \ge c > 0,\] \[ \sup_{\beta_0 \in {\cal B}_{l_0} (s_0)} Eigmax ( \Sigma_{\beta_0}) \le C < \infty,\] where $c, C$ are positive constants.

Assumptions (ref), (ref) are needed for central limit theorem result. Specifically, we use $r_x>8$. Assumption (ref) is a standard assumption on population moments by taking into account GLM nature of our problem. Assumption (ref) can be weakened easily by use of uniformly bounded weights and moment conditions on regressors. To see that Assumption (ref) is feasible, we set up the following example. Let $h= ln (n), \bar{s}=ln(n), s_0 = ln (n), p=2n, r_x=9$, then Assumption (ref)(ii) holds since $\frac{(ln n)^{13/2}}{n^{5 /4}} \to 0$. For Assumption (ref)(iii), with $l_1$ norm for $g_n= \sqrt{\bar{s}}$, with $\bar{s} = ln (n), s_0 = ln (n)$ and $r=9$, $h = ln (n)$, $p =2n$, we have $H_n = O ([ln (n)]^{2/9} n^{2/9})$, so $\max( \bar{s}, H_n^2 s_0^2) = O \left( [ln (n)]^{4/9} n^{4/9} [ln (n)]^2\right)$, then \[ h \bar{s}^{3/2} \frac{lnp}{\sqrt{n}} H_n^2 s_0^2 = O \left( \frac{[ln (n)]^{(7/2)+ (22/9)} n^{4/9}}{n^{1/2}}\right) = o(1).\] To show Assumption (ref)(iv), with the setup in (iii), $r_x = 9$, $K_n = O ( n^{4/9})$, then $[ln (n)]^4 n^{4/9}/n^{1/2} \to 0$ provides (iv).

We provide our main result, which is a central limit theorem for debiased GLM structured sparsity estimators. As far as we know, this is a new result in the literature where we have general weakly decomposable norms. We want to test the null of $\beta_j= \beta_{j0}$ for $j \in {\cal H}$.

theoremaUnder Assumptions (ref)-(ref)(i), (ref)-(ref), with $\sup_{\beta_0 \in {\cal B}_{l_0} (s_0)} \max_{j \in {\cal H}} \Omega^{S_c} (\gamma_{\beta_0, S_j^c})=o(d_n)=o(1)$ (i). Uniformly over $l_0$ ball ${\cal B}_{l_0} (s_0)$ \[\frac{n^{1/2} \alpha' (\hat{b} - \beta_{0})}{\hat{V}_{\alpha}} \stackrel{d}{\to} N(0,1),\] where $\hat{V}_{\alpha}^2:= \alpha' \hat{\Theta} [\frac{1}{n} \sum_{i=1}^n X_i X_i' \dot{\rho} (y_i, X_i' \hat{\beta})^2] \hat{\Theta}' \alpha.$ (ii). \[ \sup_{\beta_0 \in {\cal B}_{l_0} (s_0)} | \hat{V}_{\alpha}^2 - V_{\alpha}^2| = o_p (1),\] with $V_{\alpha}^2:= \alpha' \Theta [ E X_i X_i' \dot{\rho} (y_i, X_i' \beta_0)^2] \Theta \alpha$.

Remarks. 1. This theorem extends Theorem 3.3 in van de Geer et al.et al. (2014) from $l_1$ norm to $\Omega (.)$ norm under weaker conditions. The main issues are: a). We need to show a new oracle inequality for $\underline{\Omega}$ based norm as in Theorem (ref). b). Then we need a feasible nodewise regression via the proof in Lemma (ref). The reason that Lemma (ref) is needed rather than simple usage of proof of Theorem (ref) is that weights and feasible nodewise regression introduce a technical issue so that we need to extend least squares loss proof in Theorem 6.1 of van de Geer (2016) to GLM in our Lemma (ref).

2. Nonlinear restrictions may be another topic, but we think it will take more space for this paper, so we did not cover it.

3. A good choice to get better approximation rates can be the usage of $l_1$ norm in nodewise regression, regardless of the penalty norm for the initial estimator $\hat{\beta}$ in section 2, this point can be seen in the proofs by ((ref)) ((ref)) and the step 2 of proof of Theorem (ref) by Assumption (ref)(iii) with Lemma A1(ii), in terms of vector norms: $l_1 \le \underline{\Omega} \le \Omega$.

We provide a theorem that provides uniform confidence intervals for our parameters. The proof uses Theorem (ref) and follows the proof of Theorem 3 in Caner and Kock (2018). So no proof will be given. Let $\Phi (t)$ be the cdf of a standard normal distribution, and $z_{1 - \delta/2}$ is the $1 - \delta/2$ percentile of the standard normal distribution, and let $diam ([a,b])=b-a$ be the length of the interval $[a,b]$ in the real line, for all $j=1,\cdots, p$ we have the following Theorem.

theoremaUnder Assumption (ref) with $r_x>8$, Assumptions (ref)(i), (ref)-(ref), with $\sup_{\beta_0 \in {\cal B}_{l_0} (s_0)} \max_{j \in {\cal H}} \Omega (\gamma_{\beta_0, j})=o(d_n)=o(1)$ (i). \[ \sup_{t \in R} \sup_{\beta_0 \in {\cal B}_{l_0} (s_0)} \left| P \left( \frac{n^{1/2} \alpha' (\hat{b}- \beta_{0})}{\sqrt{ \alpha' \hat{\Theta} [\frac{1}{n} \sum_{i=1}^n X_i X_i' \dot{\rho} (y_i, X_i' \hat{\beta})^2] \hat{\Theta}' \alpha}} \le t \right) - \Phi (t) \right| \to 0.\] (ii). For each $j=1,\cdots, p$ \[ \lim_{ n \to \infty} \inf_{\beta_0 \in {\cal B}_{l_0} (s_0)} P \left( \beta_{j0} \in [ \hat{b}_j - z_{1 - \delta/2} \frac{\hat{\sigma}_j}{n^{1/2}}, \hat{b}_j + z_{1 - \delta/2} \frac{\hat{\sigma}_j}{n^{1/2}}] \right) = 1 - \delta.\] (iii). \[ \sup_{\beta_0 \in {\cal B}_{l_0} (s_0) }diam \left( [ \hat{b}_j - z_{1 - \delta/2} \frac{\hat{\sigma}_j}{n^{1/2}}, \hat{b}_j + z_{1 - \delta/2} \frac{\hat{\sigma}_j}{n^{1/2}}]\right) = O_p (\frac{1}{n^{1/2}}).\]

Example: Logistic Loss with Debiased Weighted Group Lasso

In this part of the paper, we follow our theorems with an example. This sub-case of our main theorems will be an analysis of the logistical loss function with a weighted group lasso norm. Then this estimator will be debiased, and we form confidence intervals for the coefficients that are of interest. Yuan and Lin (2006) introduces group lasso to capture the relation between the outcome variable and group of variables, rather than the individual variables. Oracle inequalities are proved by Lounici et al. (2011), and the debiased weighted group lasso estimator is analyzed by Mitra and Zhang (2016). These three papers handled the linear loss function with group structure. Meier et al. (2008) analyze weighted group lasso for logistic regression and provide maximal inequalities. So we extend this literature by providing a debiased weighted group lasso in a logistical loss context. First, we start with the penalty function and then show the more familiar logistical loss. Our penalty is the weighted group lasso norm in p.89 of van de Geer (2016). We follow the description and the properties of this norm from p.89-90 in van de Geer (2016). Let $\beta_{G_j}$ represent the $\beta$ vector that correspond to group $G_j$ entries, and there are $m$ groups in total. We assume groups are disjoint $G_j \cap G_k = \emptyset$. So $ \cup_{j=1}^m G_j = \{1, \cdots, p\}$. Group size of group $G_j$ is the cardinality of the group $|G_j|$. Let us denote the maximum group size by $g:= \max_{1 \le j \le m} |G_j |$. The weighted group lasso norm is defined as

equation[equation omitted — 94 chars of source]

where each group is weighted by the square root of its cardinality, in this way large groups are penalized proportionately to their size. Penalization occurs for each group, hence a group with its all members are included or excluded from the regression. The dual norm for weighted group lasso norm is, for a general vector $\omega$

equation[equation omitted — 114 chars of source]

where $\omega_{G_j}$ represent all elements correspond to $G_j$. Also any group is an allowed set, as well as their unions, since the weighted group lasso is weakly decomposable. Furthermore this norm is also decomposable. This means \[ \sum_{j=1}^m \sqrt{|G_j |} \| \beta_{G_j }\|_2 = \sum_{j \in S} \sqrt{|G_j|} \| \beta_{G_j} \|_2 + \sum_{j \in S^c} \sqrt{|G_j|} \|\beta_{G_j} \|_2,\] and $S$ can be a subset of all groups, say $S:= \{ 2, 4, 5\}$, and the remainder is $S_c:= \{ 1, 3, 6\}$, if $m=6$. So groups \{2,4,5\} and \{1,3,6\} can be decomposed into two separate sets $S$, $S^c$. Denote the active set with $S_0$, where this carries the indices of the active (relevant-nonzero groups), whereas $s_0$ is the cardinality, which is the number of relevant-active groups. We setup a logistic loss function with group structure as in Meier et al. (2008). Let $y_i$ be iid across i and take values of 0 or 1 , and $X_i$ is also iid across i and is a $p \times 1 $ vector. Denote $X_{i, G_j}$ as the predictors in $G_j$ th group at $i$ th observation, and $\beta_{G_j}$ represent the parameter vector corresponding to th $G_j$ th group, here $p_j:=| G_j|$ are the number of elements in $j$ th group, i.e. number of parameters in $\beta_{G_j}$.\[ \rho_{\beta} (y_i, X_i):= \rho (y_i, X_i' \beta) = - y_i \sum_{j=1}^m X_{i,G_j}' \beta_{G_j} + ln (1 + exp (\sum_{j=1}^m X_{i,G_j}' \beta_{G_j} )),\] where $X_{i,G_j}: p_j \times 1$, and $\sum_{j=1}^m p_j = p$. Let $\beta:=( \beta_{G_1}, \cdots, \beta_{G_j} , \cdots, \beta_{G_m})'$ which is a $p \times 1$ vector with $\beta_{G_j}: p_j \times 1$, for $j=1,\cdots, m$. The logistic loss with weighted group norm estimator is:

equation[equation omitted — 272 chars of source]

We want to carefully analyze whether Assumptions (ref)-(ref) are verified. First we see that Assumptions (ref)-(ref) are still needed, and $s_0$ is the number of relevant-active groups, which is the cardinality of $S_0:\{ j: \| \beta_{G_{j0}}\|_2 \neq 0 \}$. Assumption (ref) is also needed with $C_{\rho} \ge 1$ since \[ \ddot{\rho} (y_i, X_i' \beta ) = \frac{exp (X_i' \beta)}{[1 + exp (X_i'\beta)]^2},\] where $X_i' \beta = \sum_{j=1}^m X_{i, G_j}' \beta_{G_j}$. Assumption (ref) holds in this case of logistic loss. We start with Assumption (ref)(i). See that \[ \dot{\rho} (y_i, X_i' \beta_0) = - y_i + \frac{exp(X_i' \beta_0)}{1 + exp (X_i'\beta_0)},\] and clearly since $y_i$ is binary, with zero or one value, $| \dot{\rho} (y_i, X_i' \beta_0)| \le 2$ by triangle inequality. So (i) is satisfied. Next, for Assumption (ref)(ii), we have, $a:= X_i' \beta= \sum_{j=1}^m X_{i, G_{j}}' \beta_{G_j}$

equation[equation omitted — 85 chars of source]

hence $|\ddot{\rho} (y_i, a)| \le 1$, so (ii) is verified. Then Assumption (ref)(iii) is clearly satisfied since we have third degree differentiable in $a= X_i' \beta$ for $\rho(.)$ in logistical loss. Third order partial derivative is bounded, and with mean value theorem we get Lipschitz continuity in $a$.

corollaryUnder Assumptions (ref)-(ref) \[ \| \hat{\beta}_{LL} - \beta_0 \|_{wgl}:= \sum_{j=1}^m \sqrt{|G_j|} \| \hat{\beta}_{G_j} - \beta_{0, G_j} \|_2 = O_p ( s_0 \frac{\sqrt{lnp}}{\sqrt{n}}),\] where $s_0$ is the number of active-relevant (nonzero in $l_2$ norm) groups, which is the cardinality of $S_0=\{ j: \| \beta_{0,G_j} \|_2 \neq 0 \}$. The result is uniform over $l_0$ ball ${\cal B}_{l_0} (s_0) = \{ | \{ j: \| \beta_{0, G_j}\|_2 \neq 0 | \le s_0 \}$, where $|.|$ represents the cardinality of an index set.

Remark. Note that we have $\sqrt{lnp}$, where $p$ is the dimension of regressors. This rate may be the cost that we incur with our proof. So our general proof technique may have a cost in the rate, albeit a mild one. We now compare our results with the ones we can find in the literature. Note that we cite the two examples using the least-squares loss unlike our GLM loss. Our example and two comparison examples use the group lasso norm. For group norm, under uniformly bounded empirical Gram matrix, with non-normal errors, Theorem 8.1 of Lounici et al. (2011) has \[ s_0 \frac{ \sqrt{ (ln m)^{3/2 + \delta}}}{\sqrt{n}},\] where $\delta >0$ is a positive constant. This result is (8.3) of Lounici et al. (2011) with one task (T=1) there which provides the group structure equivalent to us. So main difference between the rates is comparing $lnp$ with $ (ln m)^{3/2 + \delta}$, so with large number of groups $m$, our and their result will be similar or we do better, otherwise when $m$ is small, the estimation error may be smaller than our result. Mitra and Zhang (2016) on the other hand find the rate \[ \| \hat{\beta} - \beta_0 \|_{wgl} = O_p ( \frac{l + s_0 ln m}{n}),\] where $l$ is the cardinality of the largest group. This last result is derived under uniformly bounded regressor assumption. So if $l$ is close to $n$ our rate seems better, otherwise, their rate is very good.

All the other assumptions are not tied to penalties. Now we define the debiased logistic estimator with a weighted group lasso penalty. Formally

equation[equation omitted — 216 chars of source]

As mentioned above to get $\hat{\Theta}$ we follow ((ref)) and the paragraph below that, we can use $l_1$ norm for nodewise regression. The term, which is summed in the last parenthesis, is the partial derivative of logistic loss with respect to $\beta$, and this point is made in ((ref)). Set sparsity in the precision matrix for all $j \in S_j^c: \gamma_{\beta_0, S_j^c} = 0$ to simplify the expressions in the corollary below, although weak sparsity is allowed as shown in Theorems above. We provide the limit for a debiased logistic estimator with weighted group norm.As far as we know, this is a new result in the literature.

corollaryUnder Assumptions (ref)-(ref)(i), (ref), (ref)--(ref), uniformly over $l_0$ ball ${\cal B}_{l_0} (s_0)$, \[\frac{n^{1/2} \alpha' (\hat{b}_{LL} - \beta_{0})}{\hat{V}_{\alpha}} \stackrel{d}{\to} N(0,1),\] where $\hat{V}_{\alpha}^2:= \alpha' \hat{\Theta} \left[\frac{1}{n} \sum_{i=1}^n X_i X_i' \left( -y_i+ \frac{exp (X_i' \hat{\beta}_{LL})}{(1 + exp (X_i' \hat{\beta}_{LL}))} \right)^2 \right] \hat{\Theta}' \alpha.$

Monte Carlo

In this section, we consider the performance of the debiased weighted group lasso with logistical loss that is described in the previous section. We consider two main setups. They will differ in terms of number of groups. Setup 1 will have 5 groups, and Setup 2 will have 10 groups. In each setup, we want to see the size and power of the test, and coverage of zero and nonzero coefficients. Since the computations are time-consuming, we use 100 iterations for each exercise.

Setup1: There are 5 groups and one intercept, which is not included in the groups (the intercept is not penalized). Let $p+1$ represent the total number of parameters fitted. Also, just to give an example, for $l$ th group, the row vector can be represented as: $\beta_{0,g_l}' = 0_{p/5}'$, which is all zeros, with dimension of the $l$ th group as $p/5$. Given that our parameter set is: \[ \beta_0:=(\beta_{0,1}=0, \beta_{0,g_1}' = 0_2', \beta_{0,g_2}'= 0_{p/5+10}', \beta_{0,g_3}'=1_{p/5}', \beta_{0, g_4}'= 0_{p/5}', \beta_{0,g_5}'=0_{2*p/5-12}')'.\]

For each $i_l=1,\cdots, n_l$, where $n_l$ is the observations in $l$ th group. Across $i_l$, the data is iid, with $X_{i_l}$ (which is $p$ is multivariate normal and $X_{i_l} \sim N(0, \Sigma)$ with $k,j$ th element in $\Sigma$) \[ \Sigma_{k,j} = \rho^{|k-j|},\] with $\rho=0.5, 0.75$. Inside the groups, the regressors are correlated, but outside there is independence. Also each group has the same multivariate normal distribution, but as described independent from other groups.

In setup 2, we deviate from setup 1. Here we just cover the differences between two setups. In this design, we have 10 groups to measure the effect of number of groups in our analysis.

eqnarray*[eqnarray* omitted — 330 chars of source]

Tuning parameter choice is essential, and for the weighted group lasso estimation, we use the procedure outlined by Meier et al. (2008). First, we set up a grid of $\lambda$ choices, and let $\lambda_{max}$ represent the $\lambda$ when used in weighted group lasso (logistical loss), will provide all zero parameter estimates. The "grplasso-R" program by Meier (2020) computes both $\lambda_{max}$ and the weighted group lasso in the logistical loss. Our grid of $\lambda$ choices are \[ \Lambda:= \{ \lambda_{max}*0.3, \lambda_{max}*(0.3)^2, \cdots, \lambda_{max}*(0.3)^{25} \}.\] These are $25$ possibilities in total. This type of grid is very similar to the one used in p.66 of Meier et al. (2008), and the idea is taken from that paper. For the weighted group lasso estimator with logistical loss, the tuning parameter choice is given by p.66 of Meier et al. (2008). This is similar to a two-fold cross-validation exercise. We give a broad outline of the procedure to choose $\lambda$ in Section 7 above, and form our estimator in Monte Carlo.

1. From the first half of the data $(i=1,\cdots, n/2))$ pick coefficient estimates by applying the weighted group lasso in ((ref)) for each $\lambda \in \Lambda$ above.

2. Use these estimates in the second half of the sample $i=n/2+1,\cdots, n$ in the unpenalized logistical loss, (i.e. the term ((ref)) without the weighted group lasso penalty).

3. Pick the $\lambda$ that provides the minimum in step 2 above. Denote this $\lambda$ as $\lambda_o$.

4. Run ((ref)) with full sample with $\lambda_o$ and get coefficient estimates $\hat{\beta}_{ LL}$.

5. We describe now the nodewise regression to get $\hat{\Theta}$. For this purpose we form weighted regressors $X_{\hat{\beta}_{LL}, j}:= W_{\hat{\beta}_{LL}} X_j$, and $X_{\hat{\beta}_{LL},-j}:= W_{\hat{\beta}_{LL}} X_{-j}$ with $W_{\hat{\beta}_{LL}}:= diag (\hat{w}_{\hat{\beta}_{LL},1}, \cdots, \hat{w}_{\hat{\beta}_{LL},i}, \cdots, \hat{w}_{\hat{\beta}_{LL},n})$ with $\hat{w}_{{\beta}_{LL},i}:=\sqrt{\ddot{\rho} (y_i, X_i' \hat{\beta}_{LL})}$ where we use ((ref)).

6. Then we run ((ref)) with $l_1$ penalty, and to choose tuning parameters, we use five-fold cross-validation.

7. We can then form $\hat{\Theta}$ as described in ((ref)) and below that equation.

8. Use $\hat{\beta}_{LL}, \hat{\Theta}$ in the formula for the debiased weighted group lasso estimator in ((ref)) and get $\hat{b}_{LL}$.

9. After getting $\hat{b}_{LL}$ the test formation and coverage can be seen in Sections 6-7.

We consider four different targets. First, we report the size of the test at $5\%$ with $h=2$ restrictions, and we test $H_0: \beta_2 =0, \beta_3 =0$, $\beta_{1}$ is the intercept, so we test basically whether group 1 is significant or not. For the power exercise, we test $H_0: \beta_2 = 0.5, \beta_3 =0.5$, which is a mild deviation from the true parameters. We also check the coverage of nonzero and nonzero parameter by checking $\beta_{0, g3,3} =1$ which is the third group's third coefficient for the nonzero parameter, and $\beta_{0,2}=0$ for the zero parameter, which is the first coefficient of group 2.

All else is the same for setup 2, except for the coverage of nonzero parameter exercise, which we check $\beta_{0, j}=1$, $j=p/10+16$. Tables 1-2 report the results. All cells in tables report percentages.

In both setups, we cover five combinations of sample size with number of parameters: $(n=100, p=100,) (n=150,p=100), (n=150,p=200), (n=300, p=200), (n=300,p=400)$. This type of setup is chosen since it can analyze three issues: 1. at fixed $n$, what will be the role of increase in $p$ on our metrics?, 2. at fixed $p$ what will be the role of increase in $n$ on our metrics? 3. when we increase $p,n$ simultaneously what will be the effect on our metrics?

table[table omitted — 750 chars of source]
table[table omitted — 751 chars of source]

Tables show that our test has a very good size across two different group sizes with different correlation structures for regressors. To give an example with five groups in Table 1, with $p=400, n=300$, the size of the tests are 3% and 0% at 5% levels at $\rho=0.5, \rho=0.75$ respectively. We see more varying power results. The test has good power with $\rho=0.50$ structure in both Tables 1-2. The power is between 76-100%. However, with $\rho=0.75$, a larger correlation among regressors, we see that the power declines. At $p=400, n=300$, in Table 2, with 10 groups and $\rho=0.75$, the power is 70%; however with $\rho=0.5$ in Table 2, the power is at 100%. At 95% ideal coverage level, we see that in Table 1, for zero parameter, the coverage is very good at $p=200, n=150$, and they are at 97%, and 100% level with $\rho=0.5, \rho=0.75$ respectively. For the nonzero parameter, the coverage is good at $p=100$, but deteriorates at $p=400$.

To answer the questions about increasing sample size-parameter dimension, Tables 1-2 show that when we keep $n=300$ and increase the number of parameters $p$ from 200 to 400, the size improves at $\rho=0.5$ from 6% to 2-3%, the power is stable at 99-100%, coverage of zero parameter is stable at 98-100%, however, the coverage of nonzero parameter deteriorates from 76-89% to 41-45%. The same type of results is more stable at $\rho=0.75$. Then we also see, if we fix $p=200$ and increase $n$ from 150 to 300 in Tables 1-2, at $\rho=0.50$, the size is stable at 4-6%, the power improves from 91-99% to 99-100%, the coverage of zero parameter is stable at 97-100%, and the coverage of the nonzero parameter improves from 62-67% to 76-89%. In the case of $\rho=0.75$, with the same question of increasing $n$ with fixed $p$, the power improves, and the other metrics are stable. The last question is, what may happen when we jointly increase $n,p$ from $n=150,p=200$ to $n=300, p=400$? The answer is very similar to the first question, for example, at Tables 1-2, with $\rho=0.50$, we see size decline from 4-5% to 2-3%, and the power improves from 91-99% to 99-100%, the coverage of zero parameter increases from 97% to 99-100%, however, the coverage of nonzero parameter declines from 62-67% to 41-45%.

Conclusion

In this paper, we propose structured penalty functions in generalized linear models. Using a feasible-weighted nodewise regression with the same or weaker penalty norm than the original problem, we estimate the inverse of the second order partial derivative of the loss function. Using this approximate inverse of the second-order partial-derivative we get a debiased GLM -structured sparsity estimator. We build uniformly valid confidence intervals around the parameters using the debiased estimate. A sub-case of debiased logistical loss with weighted group lasso penalty is analyzed. For future work, M-estimation in sparse structured framework can be considered.

\setcounter{equation}{0}\setcounter{lemma}{0} \baselineskip=15pt