Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
190,958 characters · 21 sections · 185 citation commands
High-Dimensional Econometrics and Regularized GMM
In this chapter, we review some of the main ideas and concepts from the literature on estimation and inference in high dimensions. High-dimensional models naturally arise in many contexts. First, empirical researchers may want to build more flexible models in an effort to approximate real phenomena better. Second, they may want to use more “flexible" controls to make conditional exogeneity more plausible in an effort to (more plausibly) identify causal/structural effects. Third, researchers may want to analyze policy effects on very high-dimensional outcomes and/or across many groups. Fourth, researchers may wish to leverage high-dimensional exclusion restrictions (“many instruments") in an effort to pin down structural parameters better. These and other contexts motivate the set of methods and results we overview in this chapter. In addition to providing an overview of useful tools, we develop some new results in order to make existing results more useful for applications in econometrics. We note that, since the literature on high-dimensional estimation and inference is large, we have opted to review only {\em some} of the main results from this literature. In this regard, our exposition complements other reviews, e.g. BC11b, FLQ11, CS17, CHS. For a textbook-level treatment, we refer an interested reader, for example, to BvdG11, G15, HTW15, and G16.
High-dimensionality typically refers to a setting where the number of parameters in a model is non-negligible compared to the sample size available. The presence of a large number of parameters often necessitates us to design estimation and inference methods that are different from those used in classical, low-dimensional, settings. High-dimensional models have always been of interest in econometrics and have recently been gaining in popularity. The recent interest in these models is due to both the availability of rich, modern data sets and to advances in the analysis of high-dimensional settings, such as the emergence of high-dimensional central limit theorems and regularization and post-regularization methods for estimation and inference.
We split the chapter into two parts. In the first part, we consider inference using the {\em Many Approximate Means} (MAM) framework. In particular, we assume that we have a potentially high-dimensional vector of parameters $$ \theta_0:=(\theta_{0 1},\dots,\theta_{0 p})^\prime\in\mathbb{R}^p $$ and its estimator $$ \hat{\theta}:=(\hat{\theta}_1,\dots,\hat{\theta}_p)^\prime\in\mathbb{R}^p, $$ having an approximately linear form,
where $Z_1,\dots,Z_n$ are independent zero-mean random vectors in $\mathbb{R}^p$, sometimes referred to as “influence functions”, and $r_n\in\mathbb{R}^p$ is a vector of linearization errors that are asymptotically negligible; see the next section for the formal requirement. The vectors $Z_1,\dots,Z_n$ are either directly observable or can be consistently estimated. Here, we allow for the case $p\gg n$.
This framework is rather general and covers, in particular, the case of testing multiple means with $r_{n}=0$. More generally, this framework covers multiple linear and non-linear $M$-estimators and also accommodates many de-biased estimators; see, e.g., HeShao00 for explicit conditions giving rise to linearization ((ref)) in low-dimensional settings and BCK:biometrika for conditions in the high-dimensional settings with $p \gg n$.
While conceptually easy-to-understand, the MAM framework allows us to present fundamental concepts in high-dimensional settings:
The first concept here includes simultaneous confidence interval construction for all (or some) components of the vector $\theta_0 = (\theta_{0 1},\dots,\theta_{0 p})'$. As we explain in the next section, constructing {\em simultaneous} confidence intervals is especially important in the high-dimensional settings and we explain how to construct such intervals. This concept also includes multiple testing with family-wise error rate (FWER) control, where we simultaneously test hypotheses about different components of the vector $\theta_0$ and we want to make sure that the probability of at least one false null rejection does not exceed the pre-specified level $\alpha$. The second concept includes multiple testing with false discovery rate (FDR) control, where we simultaneously test hypotheses about different components of $\theta_0$ and we want to make sure that the fraction of falsely rejected null hypotheses among all rejected null hypotheses does not exceed the pre-specified level $\alpha$, at least in expectation. FDR control is more liberal than FWER control, so procedures with FDR control typically have larger power than those with FWER control. This higher power may be particularly important, for example, in genoeconomics, where procedures with FWER control often fail to find any association between the outcome variables and genes; see Example (ref) below for the details. The third concept includes estimation of linear functionals of the vector $\theta_0$. We will show that estimating such functionals sometimes requires forms of regularization, and we will explain the details of $\ell_1$-regularization. This discussion will prepare us for the more ambitious problems arising in the second part of the chapter.
To perform the tasks described above, we will use some fundamental tools:
One of the main goals of the first part of this chapter will therefore be to provide statements and discussion of these key tools in a simple but interesting framework. Outside of being useful in the MAM framework, these tools play an important role in the theory of high-dimensional estimation and inference more generally.
We next review some simple motivating examples that fall into the MAM framework.
In the second part of the chapter, we study estimation and inference in the high-dimensional GMM setting, where both the number of moment equations and the dimensionality of the parameter of interest may be large. Specifically, we consider a random vector $X\in\mathbb R^{d_x}$, a vector of parameters $\theta\in\mathbb R^p$, and a vector-valued score function $g(X, \theta)$ mapping $\mathbb{R}^{d_x}\times\mathbb R^p$ into $\mathbb R^m$, for some $m\geq p$. For the moment function
we assume that the true parameter value $\theta_0$ satisfies
We are then interested in estimating $\theta_0$ and carrying out inference on $\theta_0$ using a random sample $X_1,\dots, X_n$ from the distribution of $X$. We allow for the case $m\gg n$ and $p\gg n$.
We develop a Regularized GMM estimator (RGMM) of $\theta_0$ and study its properties under various structural assumptions, such as sparsity or approximate sparsity of $\theta_0$. This novel estimator extends the Dantzig selector of Cand\`{e}s and Tao in CT07 that was developed specifically for estimating linear mean regression models.
To gain intuition behind the RGMM estimator, we also consider a general minimum distance estimation problem, where the parameter $\theta_0$ is known to satisfy (ref) but the function $g(\theta)$ does not necessarily take the form (ref). Assuming that an estimator of $g(\theta)$ is available, we formulate a Regularized Minimum Distance (RMD) estimator of $\theta_0$ and develop its properties under easily-interpretable high-level conditions. Specializing these conditions for the GMM setting then allows us to derive properties of the RGMM estimator under relatively low-level conditions.
Like other estimators developed for high-dimensional models, such as Lasso, the RGMM estimator is suitable for coping high-dimensionality of the problem but has a complicated asymptotic distribution, making inference based on this estimator problematic. We therefore also develop a Double/Debiased RGMM estimator (DRGMM) that is asymptotically linear, and thus fits into the MAM framework, reemphasizing the role of the MAM framework, and reducing the problem of inference on $\theta_0$ to our analysis in the first part of the paper. Importantly, for our DRGMM estimator, we also consider a version with the optimal weighting matrix.
{\bf Notation.} In what follows, all models and probability measures $P$ can be indexed by the sample size $n$, so that models and their dimensions can change with $n$, allowing dimensionality to increase with $n$. We use “wp $\to 1$" to abbreviate the phrase “with probability that converges to 1", and we use arrows $\to_{{\mathrm{P}}}$ and $\leadsto_{{\mathrm{P}}}$ to denote convergence in probability and in distribution, respectively. The symbol $\sim$ means “distributed as". The notation $a \lesssim b$ means $a = O(b)$ and $a \lesssim_{{\mathrm{P}}} b$ means $a = O_{{\mathrm{P}}}(b)$. We also use the notation $a \vee b = \max \{ a, b \}$ and $a \wedge b = \min \{ a , b \}$. For any $a\in\mathbb R$, $\lfloor a \rfloor$ denotes the largest integer that is smaller than or equal to $a$, and $\lceil a\rceil$ denotes the smallest integer that is larger than or equal to $a$. For a positive integer $m$, $[m] = \{1,\dots, m\}$.
Next, for any vector $x = (x_1,\dots,x_p)'$, we denote the $\ell_{1}$ and $\ell_{2}$ norms of $x$ by $\|x\|_1 = \sum_{j=1}^p|x_j|$ and $\| x \|_{2} = (\sum_{j=1}^p x_j^2)^{1/2}$, respectively. The $\ell_{0}$-“norm" of $x$, $\|x\|_0$, denotes the number of non-zero components of the vector $x$. Moreover, for any vector $x = (x_1,\dots,x_p)'$ in $\mathbb R^p$ and any set of indices $T\subset\{1,\dots,p\}$, we use $x_T = (x_{T 1},\dots,x_{T p})'$ to denote the vector in $\mathbb R^p$ such that $x_{T j} = x_j$ for $j\in T$ and $x_{T j} = 0$ for $j\in T^c$, where $T^c = \{ 1,\dots, p \} \setminus T$. For any matrix $A$ of $p$ columns, we use $\|A\|$ to denote the operator norm of $A$: $\| A \| = \sup_{x \in \mathbb{R}^{p}, \| x \|_{2}=1} \| Ax \|_{2}$.
The transpose of a column vector $x$ is denoted by $x'$. For a differentiable map $\mathbb{R}^{d} \ni x \mapsto f(x) \in \mathbb{R}^{k}$, we use $\partial_{x'} f$ to denote the $k\times d$ Jacobian matrix $\partial f / \partial x' = (\partial f_{i}/\partial x_{j})_{1 \le i \le k,1 \le j \le d}$, and we correspondingly use the expression $\partial_{x'} f(x_0)$ to denote $\partial_{x'} f (x) \mid_{x = x_0}$, etc. When we have an event $A$ whose occurrence depends on two independent random vectors, $X$ and $Y$, we use ${\mathrm{P}}_X(A)$ to denote the probability of $A$ with respect to the distribution of $X$, holding $Y$ fixed. For given $Z_{1},\dots,Z_{n}$, we use the notation $Z_{1}^{n} = (Z_{1},\dots,Z_{n})$. We use $\Phi$ and $\phi$ to denote the cdf and pdf of the standard normal distribution.
Finally, we use standard empirical process theory notation. In particular, ${\mathbb{E}_n}[\cdot]$ abbreviates the average $n^{-1}\sum_{i=1}^n[\cdot]$ over index $i=1,\dots,n$, e.g. ${\mathbb{E}_n}[f(z_i)]$ denotes $n^{-1}\sum_{i=1}^n f(z_i)$. Also, if $Z$ is a random vector with law $P$ and support $\mathcal Z$, $(Z_i)_{i\in[n]}$ is a random sample from the distribution of $Z$, and $\mathcal F$ is a class of functions $f\colon\mathcal Z\to\mathbb R$, then $\mathbb{G}_n f := \mathbb{G}_n f (Z) := n^{-1/2} \sum_{i=1}^n (f(Z_i) - {\mathrm{E}}[f(Z)])$ for all $f\in\mathcal F$.
Suppose that we have a parameter $\theta_0 = (\theta_{0 1},\dots,\theta_{0 p})' \in\mathbb R^p$ and an estimator $\hat{\theta}=(\hat{\theta}_1,\dots,\hat{\theta}_p)^\prime\in\mathbb{R}^p$ of this parameter that has an approximately linear form:
where $Z_1,\dots,Z_n$ are independent zero-mean random vectors in $\mathbb{R}^p$, sometimes called the “influence functions," and $r_n=(r_{n1},\dots,r_{np})^\prime\in\mathbb{R}^p$ is a vector of linearization errors that are small in the sense that
with a more precise requirement provided in Condition A. The vectors $Z_1,\dots,Z_n$ may not be directly observable, and we assume some estimators $\hat{Z}_1,\dots,\hat{Z}_n$ of these vectors are available in this case. In this section, we are interested in carrying out different types of inference on $\theta_0$. We are primarily interested in the case where $p$ is larger or much larger than $n$, but the results below apply when $p$ is smaller than $n$ as well. Throughout the chapter, we refer to this setting as the {\em Many Approximate Means} (MAM) framework.
In this section, we review results from the literature on the high-dimensional Central Limit Theorem (CLT), high-dimensional bootstrap theorems, moderate deviations for self-normalized sums, simultaneous confidence intervals, multiple testing with the Family-Wise Error Rate (FWER) control, and multiple testing with the False Discovery Rate (FDR) control. All results to be reviewed below exist in the literature for the case of many {\em exact} means, where the approximation errors are not present, $r_n = 0$, and the vectors $Z_i$ are observed. We extend these results to allow for many {\em approximate} means and also for unobservable but estimable vectors $Z_i$, i.e. we extend the results to cover the MAM framework. This extension is important because many estimators we work with are {\em asymptotically} linear but do not have to be linear in finite samples.
At the end of this section, we also consider the problem of estimating linear functionals of $\theta_0$, which motivates such concepts as sparsity and $\ell_1$-regularization and prepares us for the discussion in the second part of the chapter.
To perform inference on $\theta_0$, we first need to develop a distributional approximation for $$ \sqrt n(\hat\theta - \theta_0). $$ When $p$ is fixed and $n$ gets large, $\sqrt n(\hat\theta - \theta_0)$ converges in distribution to a zero-mean Gaussian random vector by a classical CLT but here we are interested in the case with $p\to\infty$ or even $p/n\to\infty$ as $n\to\infty$ making classical CLTs inapplicable. We therefore rely on the high-dimensional CLT developed in CCK13,CCK15,CCK17. To state the result, and also to extend it to allow for the MAM framework, we will use the following regularity conditions. Let $(B_n)_{n\geq 1}$, $(\delta_n)_{n\geq 1}$, and $(\beta_n)_{n\geq 1}$ be given sequences of constants satisfying $B_n\geq 1$, $\delta_n \searrow 0$, and $\beta_n \searrow 0$. Here, $B_n$ is allowed to grow to infinity as $n$ gets large.
Condition M. (i) $n^{-1}\sum_{i=1}^n {\mathrm{E}}[Z_{i j}^2] \geq 1$ for all $j \in [p]$ and (ii) $n^{-1}\sum_{i=1}^n{\mathrm{E}}[|Z_{i j}|^{2 + k}] \leq B^k_n$ for all $j\in[p]$ and $k = 1,2$.
Since $r_n$ is asymptotically negligible, in the sense that (ref) holds, it follows from (ref) that, for all $j\in[p]$, the asymptotic variance of $\sqrt n(\hat \theta_j - \theta_{0 j})$ is equal to $n^{-1}\sum_{i=1}^n {\mathrm{E}}[Z_{i j}^2] \geq 1$. Thus, the first part of Condition M requires that this variance is bounded away from zero. Such a condition precludes existence of super-efficient estimators and is typically imposed even in classical settings, where $p$ is small relative to $n$. The second part of Condition M imposes the mild requirement that $n^{-1}\sum_{i=1}^n {\mathrm{E}}[|Z_{i j}|^3]$ and $n^{-1}\sum_{i=1}^n {\mathrm{E}}[|Z_{i j}|^4]$ do not increase too quickly with $n$.
Condition E. Either of the following moment bounds holds:
}
The first part of Condition E.1 requires that $Z_{i j}$'s have light tails. In particular, under Condition E.1, the tails have to be sub-exponential: $$ {\mathrm{P}}(|Z_{i j}| > x) = {\mathrm{P}}(\exp(|Z_{i j}|/ B_n) > \exp(x/B_n)) \leq 2\exp(-x/B_n),\quad \text{ for all }x>0, $$ by Markov's inequality. In fact, Lemma 2.2.1 in VW96 shows that if $$ {\mathrm{P}}(|Z_{i j}| > x) \leq 2\exp(-x/C_n),\quad\text{ for all }x>0, $$ for some $C_n>0$, then ${\mathrm{E}}[\exp(|Z_{i j}|/B_n)] \leq 2$ holds with $B_n = 3C_n$. Thus, the first part of Condition E.1 is equivalent to all $Z_{i j}$'s having sub-exponential tails. The first part of Condition E.2, on the other hand, allows for heavy-tails but imposes some moment conditions on $\max_{j\in[p]}|Z_{i j}|$. Conditions E.1 and E.2 are therefore non-nested. The second parts of Conditions E.1 and E.2 impose restrictions on how fast $B_n$ and $p$ can grow. Note that we never impose E.1 and E.2 simultaneously.
Condition A. (i) The linearization errors obey ${\mathrm{P}}( \max_{j\in[p]}|r_{n j}| > \delta_n / \log^{1/2}(p n)) \leq \beta_{n}$, and (ii) the estimates of the influence functions obey ${\mathrm{P}}( \max_{j\in[p]}{\mathbb{E}_n}[(\hat{Z}_{ij}-Z_{ij})^2] > \delta^2_{n}/\log^2(p n)) \leq \beta_n$.
The first part of Condition A requires that the approximation errors in the vector $r_n$ are asymptotically negligible, and clarifies (ref). Note that if (ref) holds, then it is rather standard to show that there exist {\em some} sequences of positive constants $(\delta_n)_{n\geq 1}$ and $(\beta_n)_{n\geq 1}$ satisfying $\delta_n\searrow 0$ and $\beta_n\searrow 0$ such that the first part of Condition A holds. The second part of Condition A requires the estimators $\hat Z_{i j}$ of $Z_{i j}$ to be sufficiently precise. Again, if $\hat Z_{i j}$'s satisfy
then there exist {\em some} $(\delta_n)_{n\geq 1}$ and $(\beta_n)_{n\geq 1}$ satisfying $\delta_n\searrow 0$ and $\beta_n\searrow0$ such that the second part of Condition A holds.
In order to state a key CLT result, let $\mathcal A$ be the class of all (closed) rectangles in $\mathbb R^p$, i.e. sets $A$ of the form $$ A = \Big\{w = (w_1,\dots,w_p)'\in\mathbb R^p\colon w_{l j} \leq w_j \leq w_{r j}\text{ for all }j \in [p] \Big\}, $$ where $w_l = (w_{ l 1},\dots,w_{l p})'$ and $w_r = (w_{r 1},\dots,w_{r p})'$ are two vectors such that $w_{l j} \leq w_{r j}$ for all $j\in[p]$. (Here, both $w_{l j}$ and $w_{r j}$ can take values of $-\infty$ or $+\infty$.) Denote $V := n^{-1}\sum_{i=1}^n{\mathrm{E}}[Z_i Z_i']$, and let $N(0,V)$ be a zero-mean Gaussian random vector in $\mathbb R^p$ with covariance matrix $V$. The following theorem establishes the Gaussian approximation for the distribution of $\sqrt n(\hat\theta - \theta_0)$, which extends Proposition 2.1 in CCK17 to allow for many {\em approximate} means.
It is useful to note that Theorem (ref) allows $p$ to be larger or much larger than $n$. For example, the theorem implies that if $Z_i$'s are i.i.d zero-mean random vectors with each component bounded in absolute value by some constant $C$ (independent of $n$) almost surely and the variance of each component bounded from below by one, then
as long as $\log^7 p = o(n)$ and (ref) and (ref) hold. Thus, Theorem (ref) shows that Gaussian approximation over rectangles is possible even if $p$ is {\em exponentially} large in $n$.
Note, however, that the Gaussian approximation here is stated only for the probability of $\sqrt n(\hat \theta - \theta_0)$ hitting rectangles $A\in\mathcal A$. The same Gaussian approximation may not hold if we look at more general classes of sets, e.g. all (Borel measurable) convex sets. In fact, it is known that if we replace the class of all rectangles $\mathcal A$ in (ref) by the class of all Borel measurable convex sets, then we must assume that $p = o(n^{1/3})$, meaning $p\ll n$, in order to satisfy (ref); see discussion on p. 2310 of CCK17. On the other hand, as we will see below, the class of all rectangles is large enough to make Theorem (ref) useful in many applications. See also Remark (ref) below on how we can extend the class of rectangles and still allow for $p\gg n$.
The Gaussian approximation result of Theorem (ref) is useful as applications below indicate, but does not immediately give a practical distributional approximation since the covariance matrix $V$ is typically unknown. We therefore also consider bootstrap approximations. In particular, we consider the Gaussian (or multiplier) and empirical (or nonparametric) types of bootstrap. For the Gaussian bootstrap, let $e = (e_1,\dots,e_n)'$ be a vector consisting of i.i.d. $N(0,1)$ random variables independent of the data yielding the estimator $\hat\theta$. A Gaussian bootstrap draw of the estimator $\hat\theta$ is then defined as
Alternatively, letting $e = (e_{1},\dots,e_{n})'$ be a vector following the multinomial distribution with parameters $n$ and success probabilities $1/n,\dots,1/n$ independent of the data, we can define an empirical bootstrap draw of the estimator $\hat\theta$ as
Equivalently, the empirical bootstrap draw of $\hat{\theta}$ can be constructed as
where $\hat{Z}_{1}^{*},\dots,\hat{Z}_{n}^{*}$ are an i.i.d. sample from the empirical distribution of $\hat{Z}_{1},\dots,\hat{Z}_{n}$, and $\overline{\hat{Z}} = n^{-1} \sum_{i=1}^{n} \hat{Z}_{i}$. Indeed, the latter expression ((ref)) reduces to the former expression ((ref)) by setting each $e_{i}$ as the number of times that $\hat{Z}_{i}$ is “redrawn” in the bootstrap sample, and the vector $e=(e_1,\dots,e_n)'$ then follows the multinomial distribution with parameters $n$ and success probabilities $1/n,\dots,1/n$ independent of the data.
The following theorems show that the distribution of the bootstrap draw with respect to $e$ approximates the Gaussian distribution given in Theorem (ref).
Figure (ref) illustrates Theorems (ref), (ref), and (ref) for rectangles $A$ of a particular type: $$ A = \Big\{w = (w_1,\dots,w_p)'\in\mathbb R^p\colon -x \leq w_j\leq x\text{ for all }j\in[p]\Big\},\quad x\geq 0. $$ Specifically, Figure (ref) plots $$ {\mathrm{P}}(\sqrt n\| \hat\theta - \theta_0 \|_{\infty} \leq x)\text{ against }{\mathrm{P}}(\| N(0,V) \|_{\infty}\leq x) $$ and $$ {\mathrm{P}}(\sqrt n\| \hat\theta - \theta_0 \|_{\infty} \leq x)\text{ against }{\mathrm{P}}_e(\sqrt n\|\hat\theta^* - \hat\theta\|_{\infty} \leq x) $$ as $x$ varies from $0$ to $\infty$ for different values of $n$ and $p$ and a distribution of $Z_{i}$'s motivated by the problem of selecting the regularization parameter of the RMD estimator in Section (ref), where $\hat\theta^*$ is either the Gaussian or the empirical bootstrap draw. The figure indicates that both Gaussian and bootstrap approximations in Theorems (ref), (ref), and (ref) are rather precise.
Another useful result for inference in high-dimensional settings is a moderate deviation theorem for self-normalized sums, which we present below. This result typically leads to conservative inference but requires very weak moment conditions. In particular, it does not require Condition E.
Since for any $x>0$, we have $1 - \Phi(x) \leq \exp(-x^2/2)$ by Proposition 2.5 in D14, it follows that $\Phi^{-1}(1 - 1/p n) \leq \sqrt{2\log(p n)}$, and so setting $\alpha = 1/n$ in (ref) gives the following corollary of Theorem (ref):
If $Z_{i j}$'s are all bounded, or at least sub-Gaussian, it is straightforward to show by combining the union bound and exponential inequalities, such as those of Hoeffding or Bernstein, that
uniformly over $j\in[p]$. Comparing (ref) and (ref) now reveals an interesting feature of Theorem (ref) and Corollary (ref): replacing the {\em true value} $V_{j j}$ of the asymptotic variance of $\sqrt n(\hat\theta_j - \theta_{0 j})$ by an {\em estimator} ${\mathbb{E}_n}[\hat Z_{i j}^2]$ allows us to obtain the same bound, $O_P(\sqrt{\log(p n)})$, for the normalized version of $\sqrt n(\hat\theta_j - \theta_{0 j})$ without imposing strong moment conditions on the data, such as boundedness, since Corollary (ref) only assumes four finite moments of the $Z_{i j}$'s (via Condition M). Results of this form were used previously by BCCH12 in the theory of high-dimensional estimation via Lasso to allow for noise with heavy tails. Also, CCK13b used such results to develop computationally efficient tests of many moment inequalities for heavy-tailed data.
When only one $\theta_{0 j}$ is of interest, it follows from standard arguments that $$ \frac{\sqrt n(\hat\theta_j - \theta_{0 j})}{({\mathbb{E}_n}[\hat Z_{i j}^2])^{1/2}}\leadsto_{{\mathrm{P}}} N(0,1) $$ as long as $r_{n j} = o_P(1)$ and ${\mathbb{E}_n}[(\hat Z_{i j} - Z_{i j})^2] = o_P(1)$. We can thus, e.g., construct a two-sided confidence interval for $\theta_{0 j}$ with asymptotic coverage $1 - \alpha$ for some $\alpha\in(0,1)$ as $$ CS_j(1-\alpha) = \left[\hat\theta_j - ({\mathbb{E}_n}[\hat Z_{i j}^2])^{1/2}\frac{z_{\alpha/2}}{\sqrt n};\ \ \hat\theta_j + ({\mathbb{E}_n}[\hat Z_{i j}^2])^{1/2}\frac{z_{\alpha/2}}{\sqrt n} \right], $$ where $z_{\alpha/2}$ denotes the $(1-\alpha/2)$-quantile of $N(0,1)$, i.e., $1 - \Phi(z_{\alpha/2}) = \alpha/2$. However, the confidence intervals above are too optimistic when many components $\theta_{0 j}$ of the parameter vector $\theta_0$ are of interest, and it is likely that one or several $\theta_{0 j}$'s will fall out of their respective confidence intervals. Therefore, to obtain valid inferential statements, we need to carry out a multiplicity adjustment to explicitly take into account that many $\theta_{0 j}$ are of interest. In this subsection, we demonstrate how to perform this adjustment and construct {\em simultaneous} confidence intervals for multiple components of $\theta_0$ using Theorems (ref)--(ref).
The following quantity will play an important role in our analysis: $$ \lambda(1 - \alpha):=(1 - \alpha)\text{ quantile of }\|\sqrt n W(\hat\theta - \theta_0)\|_{\infty}, $$ where $W := \text{diag}(w_1,\dots,w_p)$ is a diagonal, potentially unknown, weighting matrix. For concreteness, for all $j \in [p]$, we often set $w_j = V_{j j}^{-1/2}$, which normalizes each $\sqrt n w_j(\hat\theta_j - \theta_{0 j})$ to have asymptotic variance one, or $w_j = 1$, which simplifies the analysis.
If we knew $\lambda(1 - \alpha)$ and $W$, we would be able to use confidence intervals $$ CS =\prod_{j\in[p]} CS_j, \quad CS_j := \left[\hat\theta_j - \frac{\lambda(1 - \alpha)}{ w_j\sqrt n}; \ \ \hat\theta_j + \frac{\lambda(1- \alpha)}{ w_j \sqrt n} \right], $$ since they clearly satisfy the desired coverage condition,
In practice, however, $\lambda(1 - \alpha)$ is typically unknown and has to be estimated from the data. To this end, let $\hat W:=\text{diag}(\hat w_1,\dots,\hat w_p)$ be an estimator of $W$ and let $$ \hat\lambda(1 - \alpha) := (1 - \alpha)\text{ quantile of }\| \sqrt n \hat W(\hat\theta^* - \hat\theta) \|_{\infty}\mid \hat W, (\hat Z_i)_{i=1}^n, $$ where $\hat\theta^*$ is obtained via the Gaussian bootstrap, (ref). The case where $\hat\theta^*$ is obtained via the empirical bootstrap, (ref), can be analyzed similarly. We then can define feasible confidence intervals as
Below, we will show that these feasible confidence intervals still satisfy (ref), under certain regularity conditions. To prove this claim, we impose the following condition:
Condition W. (i) For some $C_W\geq 1$, the diagonal elements of the matrix $W = \text{diag}(w_1,\dots,w_p)$ satisfy $C_W^{-1} V_{j j}^{-1/2} \leq w_j \leq C_W$ for all $j \in [p]$, and (ii) the estimator $\hat W = \text{diag}(\hat w_1,\dots,\hat w_p)$ of the matrix $W = \text{diag}(w_1,\dots,w_p)$ satisfies ${\mathrm{P}}( \max_{j\in[p]}|\hat w_j - w_j|^2 (1+ {\mathbb{E}_n}[\hat Z_{i j}^2]) > \delta^2_{n}/\log^2(p n)) \leq \beta_n$.
This condition holds trivially with $C_W = 1$ if we set $\hat w_j = w_j = 1$ for all $j\in[p]$ (recall that by Condition M, we have $V_{j j}\geq 1$ for all $j\in[p]$). Also, we show that Conditions M, E, and A imply Condition W, with possibly different $\delta_n$ and $\beta_n$, if we set $w_j = V_{j j}^{-1/2}$ and $\hat w_j = ({\mathbb{E}_n}[\hat Z_{i j}^2])^{-1/2}$ for all $j \in [p]$ as a part of the proof of Theorem (ref) below.
The key observation that allows us to show that the confidence intervals (ref) satisfy the desired coverage condition (ref) will be to show that the bootstrap quantile function $\hat\lambda(1 - \alpha)$, as well as the original quantile function $\lambda(1 - \alpha)$, can be approximated by the Gaussian quantile function, $$ \lambda^g(1 - \alpha) := (1 - \alpha)\text{ quantile of }\| W N(0,V) \|_{\infty}. $$ It is this place, where Theorems (ref)-(ref) play a key role. Formally, we have the following results.
In this subsection, we are interested in simultaneously testing hypotheses about different components of $\theta_0$. For concreteness, for each $j \in [p]$, we consider testing $$ H_j:\theta_{0 j}\leq\bar\theta_{0j}\text{ against }H_j^\prime:\theta_{0 j}> \bar\theta_{0j} $$ for some given value $\bar\theta_{0 j}$, where $H_j$ is the null and $H_j'$ is the alternative. The results below also apply for testing $H_j:\theta_{0 j} = \bar\theta_{0 j}$ against $H_j':\theta_{0 j}\neq\bar\theta_{0 j}$ with obvious modifications of test statistics and critical values.
Since we are interested in testing these hypotheses simultaneously for all $j \in [p]$, we seek a procedure that would reject {\em at least} one true null hypothesis with probability not larger than $\alpha + o(1)$, uniformly over a large class of data-generating processes and, in particular, uniformly over the set of true null hypotheses. In the literature, procedures with this property are said to have {\em strong control of the Family-Wise Error Rate} (FWER).
More formally, let $\mathcal P$ be a set of probability measures for the distribution of the data corresponding to different data generating processes, and let $P\in\mathcal P$ be the true probability measure. Each null hypothesis $H_{j}$ is equivalent to $P\in\mathcal P_{j}$ for some subset $\mathcal P_{j}$ of $\mathcal P$. Let $\mathcal{W}:=\{1,\dots,p\}$ and for $w\subset \mathcal{W}$ denote $\mathcal P^{w}:=(\cap_{j\in w}\mathcal P_{j})\cap(\cap_{j\notin w}\mathcal P_{j}^c)$ where $\mathcal P_{j}^c:=\mathcal P\backslash\mathcal P_{j}$. In words, $\mathcal P_j$ is the set of probability measures corresponding to the $j$th null hypothesis $H_j$ being true and $\mathcal P^w$ is the set of probability measures such that all null hypotheses $H_j$ with $j\in w$ are true and all null hypotheses $H_j$ with $j\notin w$ are false.
Corresponding to this notation, strong FWER control means
where ${\mathrm{P}}_{P}$ denotes the probability distribution generated by the probability measure of the data $P$. We seek a procedure that satisfies (ref).
We consider three different (but related) procedures: the Bonferroni, Bonferroni-Holm, and Romano-Wolf procedures. All three procedures will be based on the $t$-statistics,
For each $j \in [p]$, the {\em Bonferroni} procedure rejects $H_j$ if $t_j > \Phi^{-1}(1 - \alpha/p)$. It is clear why this procedure satisfies (ref): For any $w\subset \mathcal W$ and any $P\in\mathcal P^w$, $$ \max_{j\in w} t_j \leq \max_{j\in w}\frac{\sqrt{n}(\hat{\theta}_j-\theta_{0 j})}{({\mathbb{E}_n}[\hat Z_{i j}^2])^{1/2}} \leq \max_{j\in[p]} \frac{\sqrt{n}(\hat{\theta}_j-\theta_{0 j})}{({\mathbb{E}_n}[\hat Z_{i j}^2])^{1/2}}, $$ and under the conditions of Theorem (ref), $$ {\mathrm{P}}_P\left( \max_{j\in[p]}\frac{\sqrt{n}(\hat{\theta}_j-\theta_{0 j})}{({\mathbb{E}_n}[\hat Z_{i j}^2])^{1/2}} > \Phi^{-1}(1 - \alpha/p) \right) \leq \alpha + o(1), $$ where $o(1)$ does not depend on $(w,P)$. This result is formally stated in Theorem (ref).
The Bonferroni procedure is a one-step method, which determines the hypotheses $H_j$ to be rejected in just one step. This procedure can be improved by employing multi-step methods. In particular, so-called {\em stepdown} methods also have strong FWER control but may reject some $H_j$'s that are not rejected by the Bonferroni procedure, thus yielding important power improvements. We will consider the following form of the stepdown methods:
Here, we obtain the {\em Bonferroni-Holm} procedure, suggested in H79, by setting $$ (BH)\qquad c_{1-\alpha,w} = \Phi^{-1}(1 - \alpha/|w|) $$ for all $w\subset\mathcal W=\{1,\dots,p\}$, where $|w|$ denotes the number of elements in $w$, and we obtain the {\em Romano-Wolf} procedure, suggested in RW05, by setting $$ (RW)\qquad c_{1 - \alpha,w} = (1 - \alpha)\text{ quantile of }\max_{j\in w}\frac{\sqrt n(\hat\theta^* - \hat\theta)}{({\mathbb{E}_n}[\hat Z_{i j}^2])^{1/2}}\mid (\hat Z_i)_{i=1}^n, $$ where $\hat\theta^*$ is a bootstrap version of $\hat\theta$. In what follows, we maintain that $\hat\theta^*$ is obtained with the Gaussian bootstrap, (ref), though the case of the empirical bootstrap can also be considered. To show that the Bonferroni-Holm and Romano-Wolf procedures have the strong FWER control (ref), we will use the moderate deviation result in Theorem (ref) and the high-dimensional CLT and bootstrap results in Theorems (ref) and (ref), respectively.
In RW05, Romano and Wolf proved the following general result regarding the strong FWER control of the stepdown methods: If the critical values $c_{1-\alpha,w}$ satisfy
then the stepdown method described above satisfies ((ref)). Indeed, let $w$ be the set of true null hypotheses and suppose that the method rejects at least one of these hypotheses. Let $l$ be the step when the method rejects a true null hypothesis for the first time, and let $H_{j_0}$ be this hypothesis. Clearly, we have $w(l)\supset w$. It then follows from (ref) that
Combining these inequalities with ((ref)) yields ((ref)).
We now formally establish strong FWER control of the three procedures described above.
As in the previous subsection, here we are again interested in testing
simultaneously for all $j = 1,\dots,p$. In this subsection, however, we seek a procedure with a different type of control. In particular, we look for procedures controlling the {\em the False Discovery Rate} (FDR).
When $p$ is large or very large, the requirement of controlling the FWER, which we had in the previous subsection, is often considered too stringent. This is because large values of $p$ lead to large critical values $c_{1-\alpha,\mathcal W}$ used by the Bonferroni-Holm and Romano-Wolf procedures in the first step; and, when $p$ is very large, it often turns out that these procedures simply fail to reject any null hypothesis. In BH95, Benjamini and Hochberg offer an alternative, weaker, requirement: Instead of controlling the FWER, they suggest to control the False Discovery Rate (FDR), which is defined as the expected proportion of falsely rejected null hypotheses among all rejected null hypotheses.
To define the FDR formally, let $\mathcal H_0\subset\{1,\dots,p\}$ be the set of indices $j$ corresponding to the true null hypotheses $H_j$, and suppose that we have a procedure for testing (ref) simultaneously for all $j \in[p]$ that rejects null hypotheses $H_j$ with $j\in\mathcal H_R\subset\{1,\dots,p\}$. Thus, the total number of rejected null hypotheses is $|\mathcal H_R|$, and the total number of falsely rejected null hypotheses is $|\mathcal H_0\cap\mathcal H_R|$. We then can define the {\em False Discovery Proportion} (FDP) by $$ FDP:=\frac{|\mathcal H_0\cap \mathcal H_R|}{|\mathcal H_R| \vee 1}. $$ Here, we put $|\mathcal H_R\vee 1$ instead of $|\mathcal H_R|$ in the denominator to avoid division by zero since it may happen that all null hypotheses are accepted, so that $|\mathcal H_R| = 0$, and we want to define $FDP = 0$ when $|\mathcal H_R| = 0$. The FDR is then defined as the expected value of the FDP: $$ FDR = {\mathrm{E}}[FDP]. $$ We would like to have a procedure with the FDR being as small as possible and with large power in terms of rejecting false null hypotheses; so, for $\alpha\in(0,1)$, we seek a procedure controlling the FDR in the sense that $$ FDR \leq \alpha + o(1). $$ Like in the previous subsection, we can also say that we have a strong control of FDR over a set $\mathcal P$ of probability measures for the distribution of the data corresponding to different data generating processes if
where we use index $P$ to emphasize that the FDR depends on the distribution of the data represented by the probability measure $P$.
Next, we describe the {\em Benjamini-Hochberg} procedure, suggested in BH95, for testing (ref) simultaneously for all $j \in [p]$ with the FDR control. Recall the $t$-statistics $t_j$ defined in (ref), and let $t_{(1)} \geq \dots \geq t_{(p)}$ be the ordered sequence of $t_j$'s. Also, define $t_{(0)} := +\infty$. Then let $$ \hat k := \max\Big\{ j=0,1,\dots,p\colon 1 - \Phi(t_{(j)}) \leq \alpha j / p \Big\}. $$ The Benjamini-Hochberg procedure rejects all $H_{j}$ with $t_{j} \geq t_{(\hat k)}$. For simplicity, to prove that this procedure controls the FDR, we will assume in this subsection that the random vectors $Z_i$, $i\in[n]$, are i.i.d., all having the same distribution as that of some random vector $Z^0 = (Z_1^0,\dots,Z_p^0)$. In addition, we will impose the following conditions:
Condition C. For some $0<r<1$ and $0 < \rho < (1-r)/(1 + r)$, (i) the correlation between any two components of $Z^0$ is bounded in absolute value by $r$: $\max_{1\leq j < k \leq p}|{\mathrm{E}}[Z^0_j Z^0_k]| \leq r$, and (ii) for all $j \in [p]$, we have $|{\mathrm{E}}[Z^0_j Z^0_k]| \geq (\log p)^{-3}$ for at most $p^\rho$ different values of $k \in [p]$.
Condition F. (i) There exists $\mathcal H = \mathcal H_n \subset \{1,\dots,p\}$ such that $|\mathcal H| \geq \log\log p$ and $\sqrt n(\theta_{0 j} - \bar\theta_{0 j}) / V_{j j}^{1/2} \geq 2\sqrt{\log p} $ for all $j\in \mathcal H$; (ii) for some $\gamma\in(0,1)$, the number of false null hypotheses $H_j$, say $p_1$, satisfied $p_1\leq \gamma p$.
Conditions C and F are adapted from LS14. Condition C restricts dependence between $Z_{1}^{0},\dots,Z_{p}^{0}$ and essentially corresponds to Condition (C1) in LS14. Condition C implies that each $t$-statistic $t_j$ is highly correlated with at most $p^{\rho}$ other $t$-statistics and is only weakly correlated with all other $t$-statistics. Condition F is a technical condition that restricts the number of false null hypotheses to be not too small and not too large.
Theorem (ref) extends Theorem 4.1 in LS14 to allow for many {\em approximate} means. The proof of this theorem relies critically on the moderate deviation result in Theorem (ref). It is worth mentioning that LS14 finds a phase transition phenomenon that the Benjamini-Hochberg method can not control the FDR when $\log p \ge c_0 n^{1/3}$ for some constant $c_{0} > 0$; see Corollary 2.1 in LS14.
We next consider the problem of estimating linear functionals of the vector $\theta_0$, say $a'\theta_0$, where $a\in\mathbb R^p$ is a vector of loadings. These functionals are of interest in many settings. For instance, as we discussed in Example (ref) in the Introduction, the conditional average treatment effects may take the form of such functionals. We consider two cases separately: $\|a\|_1 = 1$ and $\|a\|_2 = 1$. As we will see, the analysis of the first case is straightforward and the results discussed so far immediately apply in this case. On the other hand, we will see that the second case is substantially more complicated, and treating it will require introducing some form of regularization. It is this second case which will help us to develop intuition for the results to be discussed in the second part of the chapter.
Inference using the MAM framework extends to the case where we are interested in linear functionals of $\theta_0$ of the form $$ a'\theta_0, \text{ with } \|a\|_1 = 1. $$ Indeed, by H\"{o}lder's inequality, $$ \sup_{\|a\|_1 = 1} |a'(\hat \theta - \theta_0)| = \| \hat \theta - \theta_0\|_\infty, $$ which can be small even if $p$ is much larger than $n$. For instance, in the “ideal noise model” where
with $I_p$ denoting the $p\times p$ identity matrix, we have that $$ \| \hat \theta - \theta_0\|_\infty \lesssim_P \sqrt{\log p/n}, $$ so the error $\|\hat\theta - \theta_0\|_{\infty}$ is small if $\log p$ is much smaller than $n$.
Moreover, we have for a large collection of functionals $a_k'\theta_0$, indexed by $a_k$ with $\|a_k\|_1 =1$ and $k \in [p]$ that under (ref), $$ \sqrt n a_k'(\hat \theta - \theta_0) = \frac{1}{\sqrt{n}} \sum_{i=1}^n a_k'Z_{i} + o_P(1/\sqrt{\log(pn)}) \quad \text{ uniformly in $k \in [p]$. } $$ It follows that $(a_k'\hat \theta_0)_{k=1}^p$ are approximate means themselves, so we can use the “many approximate means approach" for the inference on these functionals as well.
Here we again consider inference on linear functionals $a'\theta_0$. However, in contrast to the discussion above, we assume that $\|a \|_2 = 1$ instead of $\|a\|_1 = 1$. Using only the condition that $\|a\|_2 = 1$, the previous “reduction back to many means" is not possible. Indeed, $$ \sup_{\|a\|_2 = 1}| a'(\hat \theta-\theta_0)| = \| \hat \theta - \theta_0\|_2, $$ so estimating all such functionals uniformly well will require having $\| \hat \theta - \theta_0\|_2$ small, which is typically not possible. For example, even in the ideal noise model (ref), we have $$ \| \hat \theta - \theta_0\|_2^2 = \chi^2(p)/n = (p/n) (1 + O_P(1/\sqrt{p})), $$ which diverges to infinity if $p$ increases more rapidly than $n$.
We therefore need to replace the estimator $\hat\theta$ by another estimator, say $\tilde\theta$, which performs favorably even when $p$ is large relative to $n$ in the sense that $$ \| \tilde \theta - \theta_0 \|^2_{2} \text{ is small}. $$ In order to achieve this goal, we will need two key ingredients. First, we will need to assume that $\theta_0$ has some structure which can help to reduce the complexity of the estimation. Second, we will need to construct an estimator $\tilde\theta$ that is employing this structure to reduce the estimation error. In what follows, we shall rely on approximate sparsity as the structure on $\theta_0$ and make use of $\ell_1$-norm regularization to obtain the estimator $\tilde \theta$.
A simple structure that provides useful intuition is the {\em exact sparsity structure with separation from zero}. Under this structure, $\theta_0$ has $s$ large components of size bigger than the maximal estimation error $\rho = \| \hat \theta - \theta_0\|_\infty$ and all remaining components {\em exactly} equal to zero. If $\theta_0$ has this structure and the maximal estimation error $r$ is known, we can use a simple estimator $\tilde\theta = (\tilde\theta_1,\dots,\tilde\theta_p)'$ with $$ \tilde\theta_{j} =
$$ for all $j\in[p]$. This estimator might be referred to as a ``selection-based estimator.'' In practice, $\rho$ will typically be unknown, but its distribution can be estimated by the bootstrap, as discussed in Section \ref{MAM: CLT}. We can then replace $\rho$ in the definition of $\tilde \theta$ by an estimate of the $(1 - \alpha)$-quantile of its distribution for some $\alpha = \alpha_n\to 0$. Figure (ref) illustrates this selection-based estimator.
Assuming exact sparsity with separation structure provides a substantial dimension reduction that greatly eases the task of learning $\theta_0$ as long as $s$ is small. For example, in the ideal noise model, letting $T\subset\{1,\dots,p\}$ denote the set of indices of non-zero components of the vector $\theta_0$, the use of $\tilde\theta$ defined in the previous paragraph under this structure leads to the following bound:
which can be small provided that $s$ is much smaller than $n$. However, the exact sparsity with separation structure seems unrealistic as a model for real econometric applications. A structure in which all parameters are either exactly zero or magically align themselves to be larger in magnitude than $\rho$ seems extremely unintuitive and unlikely to correspond to most sensible economic models. We will not work with this structure further, though we will consider an exact sparsity structure with {\em no separation} since it helps to convey some of the main ideas of the theory of high-dimensional estimation.
More generally, we will consider an {\em approximately} sparse structure where the coefficients, sorted in non-increasing order in terms of absolute size, smoothly decline in magnitude towards zero. For this structure, we expect to achieve a rate of convergence similar to that in (ref). Intuitively, under an approximately sparse structure, the unregularized estimator $\hat \theta$ still informs us about the components $\theta_{0 j}$ that can not be distinguished from zero. We can thus use this information to set the regularized estimator $\tilde \theta_j$ for such components to zero. In such cases, we will be making an error by rounding those coefficients to zero, but the error will be negligible as long as the coefficients decrease to zero sufficiently quickly. We develop these results formally below.
For a given estimator $\hat\theta$, we consider the following $\ell_1$-regularization procedure:
where we minimize the $\ell_1$ norm of the coefficients subject to them deviating from the initial estimates, $\hat\theta$, in the $\ell_\infty$ norm by at most $\lambda$. In (ref), $\lambda$ is the regularization parameter which controls the shrinkage in the estimator $\tilde\theta$. At one extreme, setting $\lambda = 0$ results in no regularization and yields $\tilde \theta = \hat \theta$. At the other extreme, setting $\lambda =\infty$ produces maximal regularization and results in $\tilde \theta = 0$. In general, we aim to set $\lambda$ to be of the order of the estimation error $\| \hat \theta - \theta_0 \|_\infty$ as in (ref) below. Figure (ref) illustrates the regularized estimator $\tilde\theta$.
To establish properties of the estimator $\tilde\theta$ in (ref), observe that the optimization problem in (ref) separates into $p$ independent problems: For each $j \in [p]$,
It follows that the explicit solution $\tilde\theta$ of the optimization problem in (ref) is given by $$ \tilde \theta_j = (|\hat \theta_j | - \lambda )_+ \text{sign} (\hat \theta_j),\quad j\in[p], $$ where for any $a\in\mathbb R$, we use $(a)_{+}$ to denote $a1\{a>0\}$. This solution is known as the soft-thresholded estimator. In this simple setting, it also coincides with the well-known Lasso and Dantzig selector estimators.
To carry out (ref), we need to choose the regularization parameter $\lambda$. We will assume that $\lambda$ is chosen so that
As follows from the discussion above, we can approximate such a $\lambda$ either via self-normalized moderate deviations,
or via the bootstrap as outlined in Section (ref). In the ideal noise model ((ref)), we can also choose $\lambda$ as
We analyze the estimator $\tilde\theta$ in (ref) under three different conditions.
Condition ES. The parameter $\theta_0$ is exactly sparse: There exists $T \subset \{1,...,p\}$ with cardinality $s$ such that $\theta_{0j} \neq 0$ only for $j\in T$.
Condition AS. The parameter $\theta_0$ is approximately sparse: For some $A>0$ and $a>1/2$, the non-increasing rearrangement $(|\theta_0|^*_{j})_{j\in[p]}$ of absolute values of coefficients $(|\theta_{0j}|)_{j\in[p]}$ obeys $$ | \theta_0|^*_j \leq {A} j^{-{a}}, \quad j\in[p]. $$
Condition DM. The parameter $\theta_0$ has bounded $\ell_1$ norm: $\| \theta_0\|_1 \leq K$ for some $K>0$.
Figure (ref) illustrates these conditions. Speaking informally, ES can be thought of as a special case of AS, and AS can be thought of a special case of DM when $a>1$. More formally, Condition ES implies Condition AS with any $(a,A)$ such that $A\geq \|\theta_0\|_{\infty} s^a$; and, as long as $a>1$, Condition AS implies Condition DM with any $K\geq A a/(a - 1)$. While Conditions ES and AS require $\theta_0$ to be either sparse or approximately sparse, Condition DM allows $\theta_0$ to be “dense" -- to have many elements that are all of similar, small size -- such that neither sparsity nor approximate sparsity holds. When working with Condition AS, it will be convenient to denote $s = \lceil(A/\lambda)^{1/a}\rceil$, which can be thought of as the "effective" dimension of the approximately sparse $\theta_0$.
Condition AS on approximate sparsity can also be compared with conditions typically imposed in the literature on nonparametric series estimation, e.g. N97, where the {\em unordered} sequence of coefficients $(\theta_{0 j})_{j\in[p]}$ is often required to obey $|\theta_{0 j}| \leq A j^{-a}$, which is referred to as a smoothness condition. The approximate sparsity condition thus can be considered as a relaxation of the smoothness condition.
We now establish a bound on the estimation error $\tilde\theta - \theta_0$ in the $\ell_q$ norm, where $q\geq 1$.
Note that the condition $\lambda \lesssim \sqrt{ \log p/n}$ in Corollary (ref) can easily be satisfied in the ideal noise model and in many other models, as long as $\max_{j\in[p]}{\mathbb{E}_n}[\hat Z_{i j}^2] \lesssim 1$; see (ref) and (ref).
Corollary (ref) illustrates the power of regularization in high-dimensional settings when $\theta_0$ has some structure. In particular, Conditions ES and AS both imply that the estimation error of $\tilde\theta$ in the $\ell_2$ norm satisfies $$ \| \tilde \theta- \theta_0 \|_2 \lesssim_P \sqrt{s \log p/n}, $$ where $s$ is the effective dimension in the case of Condition AS as long as $\lambda$ is chosen appropriately. Hence, the ambient dimension $p$ affects the rate only through a log factor, and the effective dimension appears in a very natural form through the ratio $\sqrt{s/n}$. Thus, under either of these two conditions, we have consistency in the $\ell_2$ norm as long as $s \log p/n$ tends to zero. Under Condition DM, the estimation error satisfies $$ \|\tilde\theta - \theta_0\|_2 \lesssim_P \sqrt K(\log p/n)^{1/4}, $$ which tends to zero as long as $K^2\log p/n \to 0$.
{\bf Proof of Theorem (ref).} It follows from (ref) that with probability at least $1-\alpha$, we have $\| \hat \theta - \theta_0\|_{\infty} \leq \lambda$, in which case
by (ref). This gives the first asserted claim.
To prove the second claim, we assume that (ref) holds. Then, under Condition DM, $\| \theta_0\|_1 \leq K$, and so $$ \| \tilde \theta - \theta_0 \|_q^q \leq \|\tilde \theta - \theta_0\|_1 \|\tilde \theta - \theta_{0} \|^{q-1}_\infty \leq (2 K) (2 \lambda)^{q-1} \leq 2^q K \lambda^{q - 1} $$ by (ref), which gives (i).
Further, under Condition ES, (ref) implies that $|\tilde\theta_j| \leq |\theta_{0 j}| = 0$ for all $j\in T^c$, and so $$ \| \tilde \theta - \theta_0 \|_q^q = \| (\tilde \theta - \theta_0)_T\|_q^q \leq s \| \tilde \theta - \theta_0\|_\infty^q \leq s (2 \lambda)^q, $$ which gives (ii).
Finally, consider the case of AS and assume, without loss of generality, that components of $\theta_0$ are decreasing in absolute values, $|\theta_{ 0 1}| \geq \dots\geq |\theta_{0 p}|$, so that $|\theta_{0 j}| \leq A j^{-a}$ for all $j\in[p]$. Then, denoting $\bar T := \{ j \in [p]: A j^{-a} > \lambda \}$, it follows from the triangle inequality and (ref) that
Here, $ |\bar T| = s -1 $, and so
Also,
where the last inequality is established by replacing the sum by an integral, as formally shown in Lemma (ref) in Appendix (ref). Combining (ref), (ref), and (ref) gives (iii) and completes the proof of the theorem. {\tiny \ensuremath{\blacksquare} }
In this section, we consider the minimum distance estimation problem. We develop a Regularized Minimum Distance (RMD) estimator and study its properties.
Suppose that we have a target moment function $\theta \mapsto g(\theta)$, mapping $ \Theta \subset \Bbb{R}^p$ to $\Bbb{R}^m$, and its empirical version $\theta \mapsto \hat g(\theta)$, also mapping $ \Theta \subset \Bbb{R}^p$ to $\Bbb{R}^m$, where both $p$ and $m$ may be large. Assume that $\theta_0$ is the unique solution of the following equation: $$ g(\theta_0) = 0. $$ We are interested in estimating $\theta_0$ using the empirical version $\theta\mapsto\hat g(\theta)$ of the function $\theta\mapsto g(\theta)$.
We define the RMD estimator $\hat \theta$ as a solution of the optimization problem
where $\lambda$ is a regularization parameter. We will choose $\lambda$ so that a solution of (ref) exists with large probability; and in the event that $\|\hat g(\theta)\|_{\infty} > \lambda$ for all $\theta\in\Theta$ so that the optimization problem (ref) has no solution, we can set $\hat\theta$ to be equal to any particular element of $\Theta$. In the linear mean regression model, the RMD estimator reduces to the Dantzig Selector proposed by Cand\`{e}s and Tao in CT07.
Let $\alpha\in(0,1)$ be a constant, which should be thought of as some small number. We will assume that the regularization parameter $\lambda$ satisfies the following condition:
{\bf Condition L}. The regularization parameter $\lambda$ is such that
}
When $\hat g_j(\theta_0)\sim N(0,\sigma^2/n)$ for all $j\in[m]$, setting $\lambda = n^{-1/2}\sigma\Phi^{-1}(1-\alpha/(2m))$ is sufficient to satisfy Condition L. This Gaussianity of the moment conditions occurs, for example, in the high-dimensional linear regression model with homoscedastic Gaussian noise under appropriate normalization of the covariates; e.g. see BC11b. More generally, we can use the self-normalization method or the bootstrap to approximate $\lambda$ as discussed in Section 2, provided that a preliminary estimator of $\theta_0$ is available.
The key consequence of Condition L is that $\theta_0$ is feasible in the optimization problem ((ref)) with probability at least $1 - \alpha$, in which case a solution $\hat\theta$ of this optimization problem exists and, by optimality, satisfies $\|\hat \theta\|_1\leq \|\theta_0\|_1$. As we discuss below, this property is crucial to handle high-dimensional models.
Denote $$ \mathcal{R}(\theta_0) := \{\theta \in \Theta: \|\theta\|_1\leq \|\theta_0\|_1\}, $$ which we sometimes refer to as the restricted set. Also, let there be some sequences of positive constants $(\epsilon_n)_{n\geq 1}$ and $(\delta_n)_{n\geq 1}$ satisfying $\epsilon_n\searrow 0$ and $\delta_n\searrow 0$. To establish properties of the RMD estimator $\hat\theta$, we will use the following high-level conditions:
{\bf Condition EMC}. The empirical moment function concentrates around the target moment function: $$\sup_{\theta \in \mathcal{R}(\theta_0)}\| \hat g(\theta) - g(\theta)\|_\infty \leq \epsilon_n \text{ with probability at least } 1-\delta_n.$$
{\bf Condition MID}. The target moment function obeys the following identifiability condition: $$\{\| g(\theta) - g(\theta_0) \|_\infty \leq \epsilon, \theta \in \mathcal{R}(\theta_0)\} \text{ implies } \| \theta - \theta_0 \|_\ell \leq r( \epsilon; \theta_0, \ell), $$ for all $\epsilon>0$, where $\epsilon \mapsto r( \epsilon; \theta_0, \ell)$ is a weakly increasing rate function, mapping $[0,\infty)$ to $[0, \infty)$ and depending on the true value $\theta_0$ and the semi-norm of interest $\ell$.
Conditions EMC and MID encode the key blocks we need for the results. A large part of this section will be devoted to the verification of these conditions in particular examples. Condition EMC will be verified using empirical process methods, where contraction inequalities play a big role. Condition MID encodes both local and global identification of $\theta_0$. It states that if $g(\theta)$ is close to $g(\theta_0)$ in the $\ell_{\infty}$ norm and $\theta$ is weakly smaller than $\theta_0$ in terms of the $\ell_1$ norm, then $\theta$ is close to $\theta_0$ in the semi-norm of interest $\ell$. As we explain below, the validity of this condition depends on the interplay between the structure of $g$, the semi-norm $\ell$, and the structure of $\theta_0$. To appreciate the latter point, we note that if $\theta_0=0$, then Condition MID always holds with $r(\epsilon; 0, \ell) = 0$ for all semi-norms $\ell$. This means that $\theta_0 = 0$ is identified in the restricted set $\mathcal R(\theta_0)$ under no other assumptions on $g$. The structure of $g$ and $\theta_0$ will begin to play an important role when $\theta_0 \neq 0$, as we discuss below.
Using these conditions, we can immediately obtain the following elementary but important result on the properties of the RMD estimator:
{\bf Proof of Proposition (ref).} Consider the event that $\lambda\geq \| \hat g(\theta_0)\|_\infty$, $\hat\theta\in\mathcal R(\theta_0)$, and $\| \hat g(\hat \theta) - g(\hat \theta)\|_\infty \leq \epsilon_n$. By the union bound and Condition EMC, this event occurs with probability at least $1-\alpha -\delta_n$ since Condition L implies that with probability at least $1 - \alpha$, we have $\lambda \geq \|\hat g(\theta_0)\|_{\infty}$ and $\hat\theta\in\mathcal R(\theta_0)$. On this event, we have, by the definition of the RMD estimator, $\|\hat g(\hat\theta)\|_{\infty} \leq \lambda$, and so
by the triangle inequality. In turn, (ref) implies (ref) via Condition MID since $g(\theta_0) = 0$, which gives the asserted claim. {\tiny \ensuremath{\blacksquare} }
In what follows, let $$ G:= (\partial/\partial \theta') g(\theta)|_{\theta = \theta_0} $$ be the Jacobian matrix. This matrix plays an important role in encoding the local information about $\theta_0$. For convenience, for all $j\in[m]$, we will use $G_j(\theta)$ to denote the $j$th row of the matrix $G(\theta)$. Regarding the semi-norm $\ell$, we shall be focusing mainly on the $\ell_1$ norm $\|\cdot\|_1$, $\ell_2$ norm $\| \cdot \|_2$, and the Fisher norm $\| \cdot \|_{F}$, i.e. $\ell \in\{\ell_1,\ell_2,F\}$, where we define the Fisher norm by $$ \| \delta \|_F := |\delta' (G' G)^{1/2} \delta|^{1/2},\quad \delta\in\mathbb R^p. $$
To make Proposition (ref) operational, we need to verify Conditions EMC and MID and provide suitable choices of $\epsilon_n$, $\delta_n$, and the rate function $r$. We will do so separately in linear and non-linear models.
Here we consider the case where $\theta\mapsto g(\theta)$ is linear:
Two lead examples of this case are the linear mean regression model and the linear instrumental variables (IV) model:
We first discuss Condition MID. We consider three cases as in Section 2:
In the dense model, where $\|\theta_0\|_1 \leq K$, we can immediately deduce that whenever $G$ is symmetric and non-negative definite, Condition MID holds with $\ell = F$ and $$ r(\epsilon; \theta_0, F) \leq \sqrt{2K\epsilon}. $$ Indeed, to prove this inequality, observe that for any $\theta\in\mathcal R(\theta_0)$ satisfying $\|g(\theta) - g(\theta_0)\|_{\infty}\leq \epsilon$, we have
since $\theta\in\mathcal R(\theta_0)$ implies that $\|\theta - \theta_0\|_1 \leq \|\theta\|_1 + \|\theta_0\|_1 \leq 2\|\theta_0\|_1\leq 2K$ via the triangle inequality. This result will imply “slow" rates of convergence of the estimator $\hat\theta$, though these slow rates will be sufficient in some applications.
In the exactly sparse or approximately sparse models, we proceed as follows. Define a modulus of continuity $$ k(\theta_0, \ell) := \inf_{\theta \in \mathcal{R}(\theta_0): \| \theta - \theta_0\|_\ell >0 } \| G (\theta- \theta_0) \|_\infty/ \| \theta- \theta_0 \|_{\ell}, $$ which we can also call an identifiability factor. Whenever $k(\theta_0,\ell)\neq0$, we immediately obtain that Condition MID holds with $$ r(\epsilon; \theta_0, \ell) \leq k(\theta_0, \ell)^{-1} \epsilon. $$ There exist methods in the literature to show that $k(\theta_0,\ell)\neq0$ and to bound $k(\theta_0,\ell)$ from below in the exactly sparse model. When we consider the approximately sparse model, we will first sparsify $\theta_0$ to $\theta_0(\epsilon) = (\theta_{0 1}(\epsilon),\dots,\theta_{0 p}(\epsilon))'$ with $$ \theta_{0 j}(\epsilon) :=
$$ where $s := \lceil (A/\epsilon)^{1/a}\rceil$ and $\Delta := \sum_{j=1}^p |\theta_{0 j}|1\{A j^{-a} \leq \epsilon\}$. We will then provide a bound on $k^{-1}(\theta_0,\ell)\epsilon$ in terms of $k^{-1}(\theta_0(\epsilon), \ell)\epsilon$ and the approximation error $\| \theta_0 - \theta_0(\epsilon)\|_\ell$. Note that we assume that the components of $\theta_0$ are decreasing in absolute values, $|\theta_{ 0 1}| \geq \dots\geq |\theta_{0 p}|$ in this construction, which is without loss of generality because we do not use this information in the estimation. We also inflate components of $\theta_{0 j}(\epsilon)$ with $A j^{-a} > \epsilon$ by using $\theta_{0 j} + sign(\theta_{0 j})\Delta/(s-1)$ instead of $\theta_{0 j}$ to make sure that $\theta\in\mathcal R(\theta_0)$ implies $\theta\in\mathcal R(\theta_0(\epsilon))$, which will be important in the verification of Condition MID. Finally, we assume that $A > \epsilon$ to make sure that $s - 1 \geq 1$.
To present one possible lower bound for $k(\theta_0,\ell)$ in the exactly sparse model, we introduce some further notation. For $J\subset\{1,\dots,m\}$ and $H\subset\{1,\dots,p\}$, let $G_{J,H} = (G_{j,h})_{j\in J,h\in H}$ be the submatrix of $G$ consisting of all rows $j\in J$ and all columns $h\in H$ of $G$. Also, for an integer $l\geq s$, define the $l$-sparse smallest and $l$-sparse largest singular values of $G$ by
respectively, where $\sigma_{\min}(G_{J,H})$ and $\sigma_{\max}(G_{J,H})$ are the smallest and the largest singular values of the matrix $G_{J,H}$. We then have the following lower bound on $k(\theta_0,\ell)$, which is an immediate consequence of Theorem 1 in BCHN17:
In Lemma (ref), the quantity $\mu_n$ appears as a key factor determining the modulus of continuity $k(\theta_0,\ell)$. When $\mu_n$ is bounded away from zero, we say that we have a strongly identified model. Otherwise, i.e. when $\mu_n>0$ is drifting towards zero, we say that we have a non-strongly identified model. The quantity $\mu_n$ in turn can be easily bounded in linear regression models:
{\bf Example (ref)} (Linear Regression Models, Continued). In the linear mean regression model (ref), assume that all eigenvalues of $G = - {\mathrm{E}} W W'$ are bounded in absolute values from above and away from zero uniformly over $n$. Then for any integer $l\geq s$, we have $$ \sigma_{\min}(l) = \min_{|H|\leq l}\max_{|J|\leq l} \sigma_{\min}(G_{J,H}) \geq \min_{|H|\leq l}\sigma_{\min}(G_{H,H}) \geq \sigma_{\min}(G) $$ and, similarly, $$ \sigma_{\max}(l) = \max_{|H|\leq l}\max_{|J|\leq l} \sigma_{\max}(G_{J,H}) \leq \sigma_{\max}(G) $$ by the standard properties of eigenvalues of symmetric matrices. Hence, we can choose $\mu_n$ in Lemma (ref) to be bounded away from zero, which means that the linear mean regression model is strongly identified with $k(\theta_0,\ell_q)$ satisfying $k(\theta_0,\ell_q)^{-1} \leq C s^{1/q}$ for some constant $C>0$.
In the linear IV regression model (ref), assume first that there exist constants $\bar\sigma \geq\underline\sigma > 0$ such that for all $s/\log n\leq l \leq s\log n$ and any combination of $l$ covariates $W_{H_1},\dots,W_{H_l}$ from the vector $W = (W_1,\dots,W_p)'$, there exists a combination of $l$ instruments $Z_{J_1},\dots,Z_{J_l}$ from the vector $Z = (Z_1,\dots,Z_m)'$ such that the matrix ${\mathrm{E}}[(Z_{J_1},\dots,Z_{J_l})'(W_{H_1},\dots,W_{H_l})]$ has singular values bounded from above by $\bar\sigma$ and from below by $\underline\sigma$, which means that $Z_{J_1},\dots,Z_{J_l}$ are strong instruments for the covariates $W_{H_1},\dots,W_{H_l}$. Then again we can choose $\mu_n$ in Lemma (ref) to be bounded away from zero, which again implies that we have a strongly identified model with $k(\theta_0,\ell_q)$ satisfying $k(\theta_0,\ell_q)^{-1} \leq C s^{1/q}$. On the other hand, if for some covariate $W_j$, we only have weak instruments, $\mu_n$ will drift towards zero, and we obtain a non-strongly identified model.{\tiny \ensuremath{\blacksquare} }
In the analysis below, we use the bound (ref) as a starting point. In particular, letting there be a sequence of constants $(L_n)_{n\geq 1}$ satisfying $L_n \geq 1$ for all $n\geq 1$, we impose the following assumption:
{\bf Condition LID}. For $q\in\{1,2\}$, either of the following conditions hold: (a) $\theta_0$ obeys Condition ES and the bound $k (\theta_0, \ell_q) \geq s^{-1/q} \mu_n$ holds, or (b) $\theta_0$ obeys Condition AS, the bound $k (\theta_0(\epsilon), \ell_q) \geq s^{-1/q} \mu_n$ holds for $s = \lceil (A/\epsilon)^{1/a} \rceil $, and $\|G_j\|_1 \leq L_n$ for each $j\in[m]$, where $G_j$ denotes the $j$th row of $G$.
We use the notation LID as a shorthand for LID$(\theta_0,G)$ since the condition is indexed by both the parameter of interest $\theta_0$ and the matrix $G$. We will make the dependence explicit when we invoke the condition to other parameters.
We can now verify Condition MID:
{\bf Proof of Lemma (ref).} The first claim is immediate from the definition of the identifiability factor $k(\theta_0,\ell_q)$. To show the second claim, assume that Condition AS holds, fix $q\in\{1,2\}$, and take any $\theta\in\mathcal R(\theta_0)$ such that $\|g(\theta) - g(\theta_0)\|_{\infty} = \|G(\theta - \theta_0)\|_{\infty} \leq \epsilon$. By the triangle inequality,
Below, we bound the two terms on the right-hand side of this inequality.
To bound $\|\theta_0(\epsilon) - \theta_0\|_q$, we have
by Condition AS and Lemma (ref) in Appendix (ref). Thus,
To bound $\| \theta - \theta_0(\epsilon) \|_q$, we have
where the last inequality follows from (ref) and Condition AS. Therefore, since $\theta\in\mathcal R(\theta_0)$ implies $\theta\in\mathcal R(\theta_0(\epsilon))$, we have by Condition LID that $$ \| \theta - \theta_0(\epsilon) \|_q \leq C_{a,q} L_n\epsilon s^{1/q} \mu_n^{-1}. $$ Combining the bounds on $\|\theta_0(\epsilon) - \theta_0\|_q$ and $\| \theta - \theta_0(\epsilon) \|_q$ above with (ref) gives the second asserted claim. {\tiny \ensuremath{\blacksquare} }
We next verify Condition EMC. Let there be a sequence of constants $(\ell_n)_{n\geq 1}$ satisfying $\ell_n\geq 1$ for all $n \geq 1$. Consider the following condition:
{\bf Condition ELM}. {\em The empirical moment function is linear, $$\hat g(\theta) = \hat G \theta + \hat g(0),$$ and with probability at least $1-\delta_n$, we have $$ \max_{j\in[m]} \|\hat G_{j} - G_{j} \|_{\infty} \vee \| \hat g(0) - g(0) \|_{\infty} \leq \ell_{n}/\sqrt n. $$ }
{\bf Example (ref)} (Regularized GMM, Continued). When specialized to GMM problems, Condition ELM means that the score function is linear, $$ g(X, \theta) = G(X) \theta + g(X,0), $$ where $G(X):=(\partial/\partial\theta')g(X,\theta)|_{\theta = \theta_0}$, and with probability at least $1-\delta_n$, we have $$ \max_{j\in[m]} \|\Bbb{G}_n G_{j}(X)\|_{\infty} \vee \| \Bbb{G}_n g(X, 0)\|_{\infty} \leq \ell_{n}. $$ This form will be useful below to verify Condition ELM in the linear regression models. {\tiny \ensuremath{\blacksquare} }
Condition ELM is plausible. It is implied by many sufficient conditions based on self-normalized moderate deviations and high-dimensional central limit theorems, as reviewed in Section 2. In particular, it is possible to choose
in many cases, as illustrated in the case of linear regression models below. We highlight the slow growth of $\ell_n$ with respect to the number of parameters $p$ and the number moment functions $m$, which is critical to allow the analysis to handle high-dimensional models.
{\bf Proof of Lemma (ref).} Conditions DM and ELM imply that with probability at least $1 - \delta_n$,
which gives the asserted claim. {\tiny \ensuremath{\blacksquare} }
{\bf Example (ref)} (Linear Regression Models, Continued). Here, we verify Condition ELM for linear regression models under primitive conditions. Since the mean regression model is a special case of the IV regression model, we only consider the latter. In the linear IV model, we have $g(0)={\mathrm{E}}[Y Z]$ and $G=-{\mathrm{E}}[Z W']$. Suppose that $Z \in \mathbb{R}^m$ and $W\in \mathbb{R}^p$ are such that the following moment condition holds for some $\sigma > 0$: $$ \max_{j\in[m],k\in[p]} \Big({\mathrm{E}}[Y^4]+{\mathrm{E}}[Z_j^4]+{\mathrm{E}}[W_k^4]\Big) \leq \sigma^2. $$ Also, suppose that $$ {\mathrm{E}}\left[\|(Y,Z',W')'\|_\infty^{8}\right] \leq M_n^{4} $$ for some $M_n$, possibly growing to infinity, but such that $n^{-1/2} M_n^2\log(p\vee m \vee n)\leq \sigma^2$. Let $(Y_i,Z_i,W_i)_{i=1}^n$ be a random sample from the distribution of $(Y,Z,W)$. Then by H\"{o}lder's inequality, $$ \sum_{i=1}^n {\mathrm{E}}[(Y_i Z_{i j})^2] \leq \sum_{i=1}^n ({\mathrm{E}}[Y_i^4]{\mathrm{E}}[Z_{i j}^4])^{1/2} \leq n\sigma^2 $$ for all $j\in[m]$. Hence, by Lemma (ref),
for some universal constant $A>0$. Therefore, applying Lemma (ref) with $s = 4$, $t = \sigma\sqrt{3n\log n}$, and $\sigma^2$ replaced by $n\sigma^2$ shows that there exist universal constants $c,C > 0$ such that with probability at least $1 - c/\log^2 n$, $$ \|\mathbb G_n g(X,0)\|_{\infty} \leq C\sigma\sqrt{\log(p\vee n)}, $$ By the same argument, again with probability at least $1 - c/\log^2n$, we also have $$ \max_{j\in[m]}\|\mathbb G_n G_j(X)\|_{\infty} \leq C\sigma\sqrt{\log(p\vee m\vee n)}. $$ Condition ELM thus holds with $\delta_n = 2c/\log^2 n$ and $$ \ell_n = C\sigma\sqrt{\log(p\vee m\vee n)}, $$ which is in accord with (ref). {\tiny \ensuremath{\blacksquare} }
Summarizing the results in Proposition (ref), Lemma (ref), and Lemma (ref), we obtain the following theorem:
Theorem (ref) implies $\ell_q$-rates of convergence of the RMD estimator and characterizes how the sparsity of $\theta_0$ impacts these rates. The impact of the overall number of coefficients is bounded by the factor $\ell_n$ which typically grows logarithmically with the number of coefficients $p$ and the number of moment conditions $m$. This is effectively the impact of not knowing the support of $\theta_0$. This implies that we can achieve consistent estimators even if $(p\vee m)\gg n$ since $\ell_n$ can grow much slower than root-$n$.
We would like to find some useful conditions for nonlinear models, where the rates of convergence of the RMD estimator will be similar to what we have in the linear case.
{\bf Condition NLID}. Assume that Condition LID holds and that the target moment function $g$ satisfies a restricted non-linearity condition around $\theta = \theta_0$; namely $$ \{\| g(\theta) - g(\theta_0) \|_\infty \leq \epsilon, \theta \in \mathcal{R}(\theta_0)\} \text{ implies } \| G (\theta- \theta_0) \|_\infty/2 \leq \epsilon $$ for all $\epsilon \leq \epsilon^*$, where $\epsilon^*$ is a tolerance parameter, measuring the degree of the linearity of the problem, with $\epsilon^* = \infty$ in the linear case.
This condition is the gradient version of restricted convexity for $\ell_1$-penalized M-estimators in BC11 and NRWY14.
This lemma can be proven using the same argument as that leading to Lemma (ref), so we omit the proof.
Next, let there be sequences of positive constants $(B_{1 n})_{n\geq 1}$ and $(B_{2 n})_{n\geq 1}$. We consider an example of a sufficient condition that allows us to bound the empirical error in estimating $\theta_0$.
{\bf Condition ENM}. (i) The target and empirical moment functions have the form $g(\theta) = {\mathrm{E}}[g(X,\theta)]$ and $\hat g(\theta) = {\mathbb{E}_n}[g(X,\theta)]$, respectively, where $g(X,\theta) = (g_1(X,\theta),\dots,g_m(X,\theta))'$ is a vector of score functions, corresponding to the RGMM example. (ii) The score functions have the index form:
where $\tilde g_j$ is a measurable map from $\Bbb{R}^{d_x} \times \Bbb{R}$ to $\Bbb{R}$ for all $j\in[m]$, $Z_u$ is a measurable map from $\mathbb R^{d_x}$ to $\mathbb R^{p_u}$ for all $u\in[\bar u]$, $u(j)\in[\bar u]$ for all $j\in[m]$, $\vartheta_u$ is $p_u$-dimensional subvector of $\theta$ for all $u\in[\bar u]$, and $p = p_1 + \dots + p_{\bar u}$. (iii) The score functions are Lipschitz in the second argument, namely
with probability one, where $L_j$ is a measurable map from $\mathbb R^{d_x}$ to $\mathbb R_{+}$ for all $j\in [m]$. (iv) Finally, we have
for all $\theta\in\mathcal R(\theta_0)$ and $j\in[m]$ and with probability at least $1 - \delta_n/6$,
}
In many applications, the number of indices $Z_u(X)'\vartheta_u$ is small. As examples, the common class of single-index models clearly use just one index, $Z_u(X)'\vartheta_ u = Z(X)' \theta$; and we would have two indexes - the supply index $Z_1(X)'\vartheta_1$ and the demand index $Z_2(X)'\vartheta_2$ - if we estimate linear supply and demand equations. The Lipschitz condition allows us to use Ledoux-Talagrand type contraction inequalities for bounding the error. Condition ENM, though plausible in a number of applications, is strong. Its chief appeal is in immediately providing useful bounds on the empirical error. One could also obtain bounds on the empirical error through the use of maximal inequalities that carefully exploit the geometry and entropy properties of the set of functions $\{g(X,\theta)\colon \theta \in \mathcal{R}(\theta_0)\}$.
Next, we demonstrate how Condition ENM can be verified in particular examples. Specifically, we consider logistic regression and nonlinear IV regression models.
{\bf Example (ref)} (Nonlinear IV Regression Model, Continued). Consider the model $$ {\mathrm{E}}[f(Y,W'\theta_0)\mid Z] = 0, $$ where $Y\in\mathbb R$ is an outcome variable, $W\in\mathbb R^p$ is a vector of endogenous covariates, $Z\in\mathbb R^m$ is a vector of instruments, $f\colon \mathbb R^2\to\mathbb R$ is some known function, and $\theta_0\in\mathbb R^p$ is a vector of parameters of interest. A vector of score functions associated with this model is $$ g(X,\theta) = Z f(Y,W'\theta),\quad X = (Y,W',Z')', $$ and so $g_j(X,\theta) = \tilde g_j(X,\theta)$, where $$ \tilde g_j(X, t) = Z_j f(Y, t),\quad t\in\mathbb R, \ j\in[m]. $$ Suppose that the function $f\colon\mathbb R^2\to\mathbb R$ is Lipschitz in its second argument: $$ |f(Y,t) - f(Y,\tilde t)| \leq \gamma(Y)|t - \tilde t|,\quad \text{for all }t,\tilde t\in\mathbb R, $$ with probability one. Suppose also that for some $\sigma > 0$,
Then (ref) holds with $L_j(X) = Z_j\gamma(Y)$ for all $j\in[m]$. Also, for any $j\in[m]$ and $\theta\in\mathcal R(\theta_0)$,
and so (ref) holds for all
Further, like in Example (ref), by (ref), it follows from Lemmas (ref) and (ref) that $$ \max_{j\in[m], \ k\in[p]}{\mathbb{E}_n}[\gamma^2(Y)Z_j^2W_k^2] \leq C_1\sigma^4 $$ with probability at least $1 - c_1/\log^2(n)$, and by (ref), it follows from Lemmas (ref) and (ref) that $$ \|\mathbb G_n(g(X,\theta_0))\|_{\infty} \leq C_2 \sigma \sqrt{\log(m n)} $$ with probability at least $1 - c_2/\log^2(n)$, where $c_1$, $C_1$, $c_2$ and $C_2$ are universal constants. Conclude, by the union bound, that (ref) holds with probability at least $1 - \delta_n/6$ as long as we set $$ B_{2 n}^2 \geq C_1\sigma^4, \ \delta_n = 6(c_1 + c_2)/\log^2 n, \ \ell_n = C_2\sigma\sqrt{\log(m n)}. $$ We have thus verified all assumptions of Condition ENM. Note also that the conditions we give here are sufficient but sometimes are not necessary. For example, if we assume that the function $f\colon \mathbb R^2\to\mathbb R$ is bounded in absolute value by a constant $C$, then (ref) holds for all $$ B_n^2 \geq 4C\max_{j\in[m]}{\mathrm{E}}[Z_j^2]. $$ Depending on the setting, this bound can be better than (ref). {\tiny \ensuremath{\blacksquare} }
The following result is an immediate corollary of Proposition (ref) and Lemmas (ref) and (ref).
Theorem (ref) shows that, under sparsity conditions, the RGMM estimator can be consistent for $\theta_0$ in the $\ell_q$-norm in nonlinear models. Importantly, the dependence of the convergence rates on the total number of parameters and moment conditions is controlled by $\ell_n$ and $\tilde\ell_n$, which typically grow logarithmically with $p$ and $m$ as shown in Examples (ref) and (ref). Thus consistency is possible even for high-dimensional models when the number of parameters exceeds the sample size. These results also highlight the different rates of convergence for different norms of interest. In particular, the RGMM estimator has good rates of convergence in the $\ell_1$ and $\ell_2$ norms. However, the rate of convergence of the RGMM estimator in the max-norm is not optimal in many cases of interest; and additional tools, and estimators, are needed to obtain good estimators for that case.
Section (ref) focused on rates of convergence of regularized minimum distance estimators. We now turn to providing inferential statements for parameters of interest where we will leverage the obtained rates of convergence of the RGMM estimator. Our development will emphasize inference via the Neyman orthogonality principle. Specifically, we aim to construct moment equations $M(\alpha; \eta) = 0$ for the target parameters $\alpha \in \Bbb{R}^{p_1}$ given the nuisance parameter $\eta \in \Bbb{R}^p$ such that the true value $\alpha_0$ of the parameter $\alpha$ obeys
where $\eta_0$ is the true value of the nuisance parameter and such that the equations are first-order insensitive to local perturbations of the nuisance parameter $\eta$ around the true value:
We refer to the latter property as the Neyman orthogonality condition.
Given $M$, estimation of or inference about $\alpha_0$ will then be based on some estimator $\hat M$ of $M$, where we plug-in a potentially biased, regularized estimator $\hat \eta$ in place of the unknown $\eta_0$. The role of the Neyman orthogonality condition, ((ref)), is precisely to mitigate the impact of the use of such biased estimators (or other non-regular estimators) of $\eta_0$ on the estimation of $\alpha_0$. In many settings, basing estimation and inference for $\alpha_0$ on estimating equations with the Neyman orthogonality property will allow us
Here we take the target parameter and the nuisance parameter to be the same, namely $$ \alpha = \theta, \quad \eta = \theta. $$ The estimator for the nuisance parameter will be $\hat \eta = \hat \theta$, the RGMM estimator from the previous subsections.
We can construct the Neyman orthogonal equations $M(\alpha; \eta)$ for the pair $(\alpha, \eta)$ as follows. First we define an optimal moment selection matrix: $$ \gamma_0 = G ' \Omega^{-1}, $$where $\Omega = {\mathrm{E}} g(X,\theta_0)g(X,\theta_0)'$ and $G=\left.(\partial/\partial \theta')g(\theta)\right|_{\theta=\theta_0}$. This $p \times m$ moment selection matrix can be used to collapse $m$-dimensional moment equations $g(\theta_0) = 0$ to $p$-dimensional moment equations $\gamma_0 g(\theta_0) = 0$. We will focus on the optimal moment selection matrix, although, in principle, sub-optimal moment selection matrices could be used in practice as well. For example, estimation of the $m$ by $m$ matrix $\Omega$ and its inverse might be a limiting factor in some settings, and one might consider using ${\rm diag}(\Omega)$ instead as its inverse is trivial to compute.
Given the moment selection matrix, we define
We then have that at the true parameter values $$ M(\theta_0; \theta_0) = 0, $$ and the Neyman orthogonality condition holds: $$ \partial_{\eta'} M(\theta_0; \eta) \Big |_{\eta = \theta_0} = - \gamma_0 G + \gamma_0 G = 0. $$
Heuristically, if we somehow knew $\gamma_0$, $G$, and $\eta= \theta_0$, we could define an "oracle" linear estimator of $\alpha_0$, $\bar \theta$, as the root of $$ \bar M(\bar\theta; \theta_0) = \gamma_0 G (\bar \theta - \theta_0) + \gamma_0 \hat g( \theta_0) = 0; $$ that is $$ \sqrt{n}(\bar \theta - \theta_0) = - (\gamma_0 G)^{-1} \gamma_0 \sqrt{n}\hat g( \theta_0). $$ This estimator is linear, so it obeys $$ \sqrt{n}(\bar \theta - \theta_0) \approx_d N(0, V), \quad V = (G'\Omega^{-1} G)^{-1}, $$ over the sets in $\mathcal{A}$, the class of all rectangles in $\mathbb{R}^p$, under the CLT conditions of Section 2. The variance matrix $V$ here is the optimal variance matrix for GMM.
The above construction is infeasible because of the many unknowns, including the true value of the target parameter, appearing in it. To make construction feasible, we can plug-in estimators corresponding to the unknowns: instead:
Specific choices of estimators $\hat \mu$ and $\hat \gamma$ will be discussed later. Given estimators of all unknown quantities, we define a “two-step" estimator of the target parameter by solving the estimated Neyman-orthogonal equation: $$ \hat M(\theta; \hat \theta) = \hat \mu^{-1} (\theta - \hat \theta) + \hat \gamma \hat g(\hat \theta) = 0. $$ The resulting solution to this equation, $\check \theta$, provides an estimator of the target parameter which we refer to as the double/debiased regularized GMM (DRGMM) estimator:
By exploiting the Neyman orthogonality property and further assumptions on the problem, we can show that this estimator approximates the infeasible “oracle" estimator defined above in the sense that
Hence, the DRGMM estimator is also approximately linear and is therefore an approximate mean, so we have
over the class of all rectangles in $\mathbb{R}^p$ under the conditions of Theorem 2.1 in Section 2. Indeed, given this construction, we are back to the MAM framework. We can thus use the inferential tools from Section 2 for immediate construction of simultaneous confidence bands and hypothesis testing with control of FWER or FDR. Inference done in this way will be optimal in the sense that the variance matrix $V$ can not be generally improved by using any other moment selection matrix $\bar \gamma$ in place of $\gamma_0$. Optimality may also be attained in other semi-parametric senses, which we do not discuss.
Here we take the target parameter, $\alpha$, and the nuisance parameter, $\eta$, to be the different: $$ \alpha = \alpha, \quad \eta = \theta. $$ The true value of the parameter is given by $(\alpha_0', \eta_0')'$ and solves $$ g(\alpha_0, \theta_0) = 0. $$ Here, we are thinking of a situation where $\eta_0$ is strongly identified and can be well-estimated by RGMM while $\alpha_0$ is only weakly or partially identified. We would thus like to use a robust testing approach to test values of $\alpha_0$ and then invert to construct a confidence set for $\alpha_0$.
The estimator for the nuisance parameter will be $\hat \eta = \hat \theta$, the RGMM estimator from Section (ref). We can construct the Neyman orthogonal equations $M(\alpha; \eta)$ for the pair $(\alpha, \eta)$ as follows. First, we define a $p' \times m$ moment selection matrix for $\alpha_0$, $$ \xi_0, \text{ e.g. } \xi_0 = I \text{ or } \xi_0 = G_\alpha \Omega^{-1}, \quad G_\alpha = \partial_{\alpha'} g(\alpha_0, \theta_0), $$ where $p' \geq \dim(\alpha)$. The latter matrix will be optimal when $\alpha_0$ is strongly identified.
Given the moment selection matrix, we define
as in CHS. Using this estimating equation, we have that, at the true values of the parameters, $$ M(\alpha_0; \theta_0) = 0 $$ using $g(\alpha_0, \theta_0) = 0$ and that the Neyman orthogonality condition holds: $$ \partial_{\eta'} M(\theta_0; \theta) \Big |_{\theta= \theta_0} = (\xi_0 - \xi_0 G ( G' \Omega^{-1} G )^{-1} G \Omega^{-1}) G = 0. $$
Heuristically, if we knew $\theta_0$ and all the extra parameters used in forming $M$, we could use the oracle Neyman-orthogonal score for testing $\alpha_0$: $$ \sqrt{ n} \bar M(\alpha_0; \theta_0) = (\xi_0 - \xi_0 G \mu_0 \gamma_0) \sqrt{n}\hat g(\alpha_0, \theta_0). $$ This quantity is clearly linear in $\sqrt{n}\hat g(\alpha_0, \theta_0)$, so it obeys $$ \sqrt{ n} \bar M(\alpha_0; \theta_0) \approx_d N(0, V_M), \quad V_M = (\xi_0 - \xi_0 G \mu_0 \gamma_0) \Omega (\xi_0 - \xi_0 G \mu_0 \gamma_0)', $$ over the sets in $\mathcal{A}$, the class of all rectangles in $\mathbb{R}^p$, under the CLT conditions of Section 2.
The above construction is again clearly infeasible. As outlined in Section (ref), we can make the construction feasible by plugging in estimators for the various missing unknowns. Specifically, we will
We discuss estimation of $\hat \mu$, $\hat \gamma$, and $\hat \xi$ further in Section (ref).
Given plug-in estimates of the unknown quantities, we define the Neyman-orthogonal score function for testing $\alpha_0$: $$ \sqrt{n} \hat M(\alpha_0; \hat \theta) = \sqrt{n} (\hat \xi_0 - \hat \xi_0 \hat G \hat \mu_0 \hat \gamma_0) \hat g(\alpha_0, \hat \theta). $$ By exploiting Neyman orthogonality property and further assumptions on the problem, we can show that this feasible score approximates the infeasible “oracle" score,
Hence, the feasible score is approximately linear and is therefore an approximate mean. We then have
over the class of all rectangles in $\mathbb{R}^p$ under the conditions of Theorem 2.1 in Section 2. We are thus back within the setting outlined in Section 2 and may use the inferential tools from Section 2 for construction of simultaneous confidence bands and hypothesis testing with control of FWER or FDR. Given that we can provide valid inferential statements based on ((ref)) for any $\alpha_0$, we can invert to obtain confidence regions.
In order to analyze the estimator ((ref)), we can write, using elementary expansions and some algebra,
where
In ((ref)), $\tilde G = \{- \partial_{\theta'} \hat g_k(\hat \theta^*_k)\}_{k=1}^m$ denotes a $m\times p$ matrix with rows $ - \partial_{\theta'} \hat g_k(\hat \theta^*_k)$, $k \in [m]$, where each row is evaluated at a point $\hat \theta^*_k$ on the line between $\hat \theta$ and $\theta_0$.
Note that because of Neyman orthogonality property we expect the term $r_1$ to be small, in fact if we knew $(\mu_0, \gamma_0, G)$ the term would vanish. In linear moment models, $\hat G = \tilde G$, so that the second term vanishes, $r_2 = 0$. The last term can also vanish under mild conditions. We analyze the structure of these remainder terms in the lemma given below.
To fix ideas, we record a trivial proposition.
A crucial step to using Proposition (ref) is to ensure the approximation error $r$ is small. The following lemma is useful for thinking about estimators of $\gamma_0$ and $\mu_0$ which are suitably well-behaved to obtain small approximation errors. Of course, the lemma only suggests one possible direction, and there are other strategies for estimating $\gamma_0$ and $\mu_0$ to explore. In what follows, we use $v_j$ to denote the $j$th row of some matrix $v$.
A plausible approach is then to construct estimators $\hat \gamma$, $\hat \mu$, and $\hat \theta$ such that the upper bounds $\bar r_1$, $\bar r_2$, and $\bar r_3$ given in Lemma (ref) approach zero sufficiently fast. A natural choice of $\hat\theta$ is given by the RGMM estimator discussed in Section (ref). We can also obtain estimators $\hat G$ and $\hat \Omega$ by plug-in expressions. We now turn to estimating $\gamma_0$ and $\mu_0$.
First, we consider one potential estimator for $\gamma_0$. We do not attempt to use a standard plug-in estimate since $\hat\Omega$ will not be full rank in high-dimensional settings, so its inverse will be ill-posed even if $\Omega^{-1}$ is well-behaved. Instead, we define the estimator $\hat \gamma$ as the solution to the program:
where $\lambda^\gamma$ is a vector of regularization parameters. By allowing $\lambda^\gamma>0$, the constraint in ((ref)) requires solving the inverse problem only approximately, which is needed to handle the rank deficiency of $\hat \Omega$.
We proceed similarly in our proposed estimator for $\mu_0$. We define the estimator $\hat \mu$ as the solution to the following program:
where $e_j$ is a coordinate vector with 1 in the $j$-th position and 0 elsewhere and $\lambda^\mu$ is a vector of regularization parameters. Again, the use of positive regularization parameters $\lambda^\mu>0$ allows us to work with approximate solutions which are needed to cope with the high-dimensionality.
We now summarize an algorithm for constructing the estimator $\check\theta$.
{\bf Algorithm for DRGMM.}\\ {\it Step 1. Compute the RGMM estimator $\hat \theta$.\\ Step 2. Use the plug-in rules $\hat G = \partial_{\theta'} \hat g(\hat \theta)$ and $\hat \Omega = {\mathbb{E}_n} g(X,\hat \theta) g(X,\hat \theta)'$.\\ Step 3. Obtain the estimator $\hat\gamma$ as defined in ((ref)).\\ Step 4. Obtain the estimator $\hat\mu$ as defined in ((ref)).\\ Step 5. Update the initial RGMM estimator $\check\theta = \hat\theta - \hat\mu\hat\gamma \hat g(\hat\theta)$. }
We note that the regularized problems ((ref)) and ((ref)) can be cast as linear programming problems and can be solved separately by row $j\in [p]$. Both features are convenient from a computational perspective.
Next we proceed to analyze the properties of the estimators. The following lemma provides high-level conditions on the estimators of $G$ and $\Omega$ and on the penalty choices to derive the needed $\ell_1$-rates of convergence for the rows of $\hat \gamma$ and $\hat \mu$.
Lemma (ref) builds upon the theory of RGMM with linear score functions established in Section (ref). As expected, condition LID is assumed to hold for the different Jacobian matrices. Lemma (ref) also highlights sufficient conditions on how the penalty parameters should be chosen. Moreover, it assumes that the estimators $\hat \Omega$ and $\hat G$ have good rates of convergence in the $\ell_\infty$-norm.
The following lemmas complement the result of Lemma (ref) by providing conditions and explicit bounds on the rates of convergence for the estimators of $\Omega$ and $G$ in the linear and non-linear case.
The moment assumptions in Lemma (ref) are quite standard and allow for $m\gg n$. Requirement (iii) relies on the rate of convergence of $\hat \theta$ which impacts the estimation of $\Omega$ only. We note that there are examples in which we can bypass this term such as the homoskedastic linear instrumental variable case discussed in Theorem (ref).
Next we provide conditions to derive bounds on the estimation error of $G$ and $\Omega$ for the non-linear case. The conditions will assume a Lipschitz condition on the score and its derivative as stated in ((ref)) and ((ref)) of condition ENM.
The moment conditions are quite standard. The bounds depend on the $\ell_1$ and $\ell_2$-rates of convergence of the initial RGMM estimator $\hat \theta$.
The following theorems provide results that builds upon the RGMM estimator discussed in Section (ref) and builds upon the previous lemmas to deliver the approximate linear expansion ((ref)).
We begin with a result for the homoskedastic linear instrumental variables model. Let $Y = W'\theta_0 + \epsilon$ with ${\mathrm{E}} \epsilon Z = 0$, ${\mathrm{E}} \epsilon^2 ZZ' = \sigma^2 {\mathrm{E}} ZZ'$. Then, using the moment function $g(\theta) = {\mathrm{E}} [ (Y- W'\theta) Z]$, we have $G = -{\mathrm{E}} Z W'$, $g(0) = {\mathrm{E}} Y Z$, and $\Omega = \sigma^2{\mathrm{E}} ZZ'$. In this homoskedastic setting, we will compute $\hat \gamma$ in ((ref)) with $\hat\Omega ={\mathbb{E}_n} ZZ'$.
Theorem (ref) derives an approximate linear representation for the estimator that immediately allows us to construct simultaneous confidence regions for all parameters under the conditions of Section 2. The result exploits the homoskedasticity and bypasses the need to estimate $\sigma^2$. In turn this allows the representation to hold under the mild sparsity requirement of $n^{-1}s^2\log^2(pmn) \leq u_n^2$.
Under more stringent requirements, the next result considers the non-linear case where the functions $g_k$ satisfies a Lipschitz condition; see condition ENM in Section (ref).
Theorem (ref) provides one approach to constructing estimators with suitable linearization based on the RGMM estimator and the estimators ((ref)) and ((ref)) for the nuisance parameters. Under suitable choice of penalty parameters, the requirement $n^{-1}s^3\log^2(pmn) \leq u_n$ where $u_n = o(\log^{-1/2}(p))$ suffices to ensure that approximation errors do not distort the asymptotic coverage of (rectangular) confidence regions. We note that the derivation of practical choices of penalty parameters has drawn considerable attention in the literature. Although Theorem (ref) allows us to postulate a choice $\bar \lambda$ that allows us to cover a class of $s$-sparse models, it would be of interest to obtain adaptive rules that are theoretically valid and practical. See Remark (ref) below for some initial discussion.
The literature on CLT with increasing dimensions is quite broad, and we refer the reader to Appendix I in CCK13 for an extensive review on this topic prior to the publication of CCK13. In the following discussion, we mainly focus on the development after CCK13. The high-dimensional CLT and bootstrap results in this chapter, namely Theorems (ref)--(ref), build upon Proposition 2.1, Corollary 4.2, and Proposition 4.3, respectively, in CCK17, which, in turn, improves on the results of their earlier paper CCK13. Both papers use some important technical tools, such as anti-concentration inequalities and Gaussian comparison theorems, obtained in CCK15. W14 provides a helpful exposition of these results targeting mathematically oriented graduate students. CCK13 also provides several useful applications of these results, including the choice of the regularization parameter for the Dantzig selector, specification testing with a parametric model under the null and a nonparametric one under the alternative, and multiple testing with FWER control. Another useful application is testing many moment inequalities, where the number of moment inequalities is larger than the sample size, which is studied in CCK13b in detail.
There are several extensions of high-dimensional CLT and bootstrap results of CCK13,CCK17. DZ17 show that Condition E, which is imposed in Theorems (ref)--(ref), can be slightly improved if we are only concerned with inference based on the empirical bootstrap or if we use the multiplier bootstrap with Gaussian weights $e_i$ replaced by weights satisfying
In particular, they show that the term $\log^7(p n)/n$ in Condition E can be replaced by the term $\log^5(p n)/n$. To obtain their results, they circumvent the Gaussian approximation and work directly with the bootstrap approximation. ZW17,CCK13b,ZC17b develop time series extensions of the high-dimensional CLT. Ch17,ChKa17 develop extensions of the high-dimensional CLT and bootstrap theorems to $U$-statistics and randomized incomplete $U$-statistics, respectively (Ch17 focuses on the second order case). Ko17 studies Gaussian approximation to a high-dimensional vector of smooth Wiener functionals by combining the techniques developed in CCK13,CCK15,CCK17 and Malliavin calculus.
An important feature of Theorems (ref)--(ref) is that they provide distributional approximation results for the class of {\em rectangles}. There are also many related results in the literature if we are interested in other classes of sets. For example, B03,Be05 show that a result like (ref) with $\mathcal A$ being the class of all {\em convex sets} is possible under certain moment conditions if $p = o(n^{2/7})$. More formally, Be05 proves the following: Let $Z_{1},\dots,Z_{n}$ be zero-mean independent random vectors in $\mathbb{R}^{p}$ and suppose that the covariance matrix of $S_{n}^{Z} = n^{-1/2} \sum_{i=1}^{n}Z_{i}$, $V= n^{-1}\sum_{i=1}^{n} {\mathrm{E}}[Z_{i}Z_{i}']$, is invertible; then
where $K$ is a universal constant. In the simplest case where $V=I$ and $\| Z_{i} \|_{2} \le C \sqrt{p}$ for all $i\in[n]$ and some constant (independent of $n$), the right-hand side on ((ref)) is $O(p^{7/4}n^{-1/2})$, which is $o(1)$ if $p=o(n^{2/7})$. This result is straightforward to use for inference since the covariance matrix $V$ can be accurately estimated as long as $p\ll n$. Recently, Zh17 improves on the Bentskus condition in the case where $Z_{1},\dots,Z_{n}$ are i.i.d. with identity covariance matrix, and such that $\| Z_{i} \|_{2} \le C \sqrt{p}$ for all $i\in[n]$ and some constant $C$; under these conditions Zh17 shows that the left-hand side of ((ref)) is approaching zero provided that $p=o(n^{2/5})$ up to log factors. Z16 shows that the multiplier bootstrap inference over all {\em centered balls} with weights $e_i$ satisfying (ref) is possible if $p = o(n^{1/2})$. See also M93 for some early results regarding the multiplier bootstrap with weights $e_i$ satisfying (ref).
Theorem (ref) on moderate deviations of self-normalized sums extends the results of JSW03 to show that the moderate deviation inequality holds, up to some corrections, even if the original random variables in the denominator of the self-normalized sum are replaced by suitable estimators. This is particularly helpful when we allow for many approximate means as opposed to many exact means; see BCCH12 for an application. Textbook-level treatment of the theory of self-normalized sums can be found in PLS09.
Theorem (ref) on simultaneous confidence intervals is a rather simple application of Theorems (ref)--(ref). Similar results formulated in terms of particular applications can be found for example in BCK:biometrika, BCCW17, and BCHN17. Simultaneous confidence intervals also constitute an important research topic in the literature on nonparametric estimation and inference. Useful references on this literature are provided in CCK14b.
Theorems (ref) and (ref) on multiple testing with FWER control build upon RW05 with the key difference that we allow for $p\to\infty$, and in particular $p/n\to\infty$, as $n\to\infty$. A closely related analog of these theorems can be found in CCK13 but our conditions here are somewhat weaker than those in CCK13. W00 explains importance of FWER control in multiple testing. A clean textbook-level treatment of multiple testing with FWER control can be found in LR05. A useful discussion of multiple testing in experimental economics can be found in LSX15. RS10 draws a connection between multiple testing with the FWER control and constructing sets covering the identified sets with a prescribed probability in partially identified models and provide a stepdown procedure for constructing such sets. Note also that our formulation of the Bonferroni-Holm procedure is slightly different but equivalent to the commonly used formulation, e.g. in LR05. We have changed the formulation to facilitate the comparison between the Bonferroni-Holm and Romano-Wolf procedures.
Theorem (ref) on multiple testing with FDR control generalizes the results of LS14 to allow for many approximate means. In turn, LS14 generalizes the original results of BH95 on the Benjamini-Hochberg procedure to allow for the unknown distribution of the data and also to allow for some dependence between the $t$-statistics. Moreover, LS14 uses moderate deviation for self-normalized sums theory to allow for testing in ultra-high dimensions. Theorem (ref) is also closely related to the results in LL14, who considers multiple testing with FDR control for a specific setting: variable selection in a high-dimensional regression model. A useful discussion of multiple testing with FDR control and other types of control can be found in RSW08. A textbook-level treatment is provided in G15.
Our Condition C for the analysis of the Benjamini-Hochberg procedure requires that each $t$-statistic is correlated with a relatively small set of other $t$-statistics. The procedure, however, remains valid under the so-called positive regression dependency condition; see BY01 for details. There also has been a lot of research about related procedures; see e.g. RSW08b. A radically different procedure, which also allows for FDR control but which we did not consider in this chapter, is the knockoff filter of BC15. This alternative procedure is designed specifically for the variable selection problem in the linear mean regression model and requires the number of covariates to be smaller than the sample size but does not restrict dependence between covariates in any way. See also CFJL17 where the knockoff filter is modified to allow for the high-dimensional regression model, with the number of covariates exceeding the sample size, in exchange for some other conditions.
We present our high-dimensional estimation results within the context of $\ell_1$-regularized minimum distance estimation focusing on the case where the parameter vector exhibits an approximately sparse structure. The estimation results clearly build upon fundamental work for $\ell_1$-penalized regression of FF93 and T96. This initial work has been expanded in many directions; see, for example, the textbook treatment of BvdG11 as well as CT07, BRT09, BC11, BCW11, BCCH12, BCH14, and BCFVH17 for results most closely related to the approach taken in this review.
More generally, providing methods for estimating high-dimensional models has been an active area of research for quite some time, and there is a large collection of methods available within the literature. ESL and CASI provide useful textbook introductions to a wide array of methods that are useful in high-dimensional contexts. Developing new techniques for estimation in high-dimensional settings is also still an active area of research, so the list of methods available to researchers continues to expand. Further exploring the use of these procedures in economic applications and the impact of their use on inference about structural parameters seems like a useful avenue to pursue.
Methods for obtaining valid inferential statements following regularization in high-dimensional settings has been an active area of research in the recent statistics and econometrics literature. Early work on inference in high-dimensional settings focused on the exact sparsity structure with separation (from zero) discussed in Section (ref); see, e.g., FanLi2001 for an early paper or FanLv2010 for a more recent review. A consequence of sparsity with strong separation from zero is that model selection does not impact the asymptotic distribution of the parameters estimated in the selected model, under regularity conditions. This property allows one to do inference using standard approximate distributions for the parameters of the selected model ignoring that model selection was done. While convenient, inferential results obtained relying on this structure may perform very poorly in more realistic approximately sparse structures as was noted in a series of papers; see, for example, LP08a and LP08b.
The more recent work on inference about model parameters following the use of regularization, including the procedure outlined in this chapter, has focused on providing procedures that remain valid without maintaining exact sparsity with separation. As noted in Section (ref), a key element in obtaining valid inferential statements is the use of estimating equations that satisfy the Neyman orthogonality condition. This idea dates at least to N59 who used the idea of projecting the score that identifies the parameter of interest onto the ortho-complement of the tangent space for nuisance parameters in the construction of the $C(\alpha)$, or orthogonal score, statistic. This idea also plays a key role in semiparametric and targeted learning theory; see, for example, Andrews94, Newey94, vdv98, SRR99, and vdLR11.
Within the high-dimensional context, much of the work on inference focuses on inference for prespecified low-dimensional parameters in the presence of high-dimensional nuisance parameters when $\ell_1$ regularization or variable selection methods are used to estimate the nuisance parameters. BCH10b considers inference about parameters on a low-dimensional set of endogenous variables following selection of instruments from a high-dimensional set using lasso in a homoscedastic, Gaussian IV model. Their approach relies on the fact that the moment condition underlying IV estimation is Neyman orthogonal. These ideas were further developed in the context of providing uniformly valid inference about the parameters on endogenous variables in the IV context with many instruments to allow non-Gaussian heteroscedastic disturbances in BCCH12. BCH14, which to our knowledge provides the first formal statement of the Neyman orthogonality condition in the high-dimensional setting, covers inference on the parametric components of the partially linear model and average treatment effects. See also BCH10a, F15, K18, BCHK16, and BCFVH17, among others, for further applications and generalizations explicitly making use of Neyman orthogonal estimating equations. As noted above, Neyman orthogonal estimating equations are closely related to Neyman's $C(\alpha)$-statistic. The use of $C(\alpha)$ statistics for testing and estimation with high-dimensional approximately sparse models was first explored in the context of quantile regression in BCK:biometrika and in the context of high-dimensional generalized linear models by BCW16. Other uses of $C(\alpha)$-statistics or close variants include those in VSW:score, NL:sparc, YNL:score, and NL17. Finally, a different strand of the literature has focused on ex-post “de-biasing” of estimators to enable valid inference as opposed to directly basing estimation and inference on orthogonal estimating equations. While seemingly distinct, the de-biasing approach is the same as approximately solving orthogonal estimating equations; see, for example, discussion in CHS. Important seminal contributions following the de-biasing approach are ZZ14, vdGBRD14, and JM14.
Rather than focus on a pre-specified low-dimensional parameter, one may also do inference for high-dimensional parameters. WR09 and MMB09 use sample splitting to provide procedures for multiple inference in high-dimensional settings that can control FWER and FDR under strong conditions. NvdG13 also consider the construction of confidence sets for the entire parameter vector in a sparse high-dimensional regression using sample splitting ideas. vdGBRD14 suggest using Bonferroni-Holm in conjunction with the de-sparsified lasso using a limiting distribution derived under homoscedastic Gaussian errors. As discussed in this chapter, the high-dimensional CLT and bootstrap results of CCK13 are broadly applicable for inference about high-dimensional parameters. BCK:biometrika provides an early use of these results for construction of a simultaneous confidence rectangle for many target parameters within a rich class of models estimated using orthogonal estimating equations; see also CHS:hdm which implements inference for high-dimensional treatment or structural effects following BCK:biometrika. DBZ16 and ZC17 also consider bootstrap inference for many parameters estimated with debiased estimators in high-dimensional models building on CCK13. More recently, BCCW17 extends BCK:biometrika to provide valid inference for many functional parameters, and BCK:eiv consider inference for many parameters in a high-dimensional linear model with errors in variables. Finally, CG16, ZB:lin, ZB:proj, and HKM consider different approaches which allow testing hypotheses about functionals that may involve the entire high-dimensional parameter vector within different high-dimensional contexts.
There is also a rapidly growing body of research focused on learning economically interesting parameters that uses different high-dimensional methods and/or aims to provide reliable inferential statement under relatively weaker conditions. Data-adaptive estimation of nuisance functions, with an emphasis on using high-dimensional methods, is advocated in the targeted learning literature under the nomenclature “super learner” though many formal results in this literature are obtained in low-dimensional settings; see, e.g., vdl:super, vdLR11, and ZvdL11. AI16 is an important example that uses tree-based methods and sample splitting for estimating and performing inference about heterogeneous treatment effects. WA:rf considers estimation and inference for heterogeneous treatment effects using a variant of random forests with formal results established in a low-dimensional context. In a high-dimensional setting, CCDDHNR18 use Neyman orthogonal estimating equations and sample splitting to provide a generic procedure for inference about low-dimensional parameters under weak conditions that allow for the use of wide variety of high-dimensional, machine learning methods. Sample splitting and orthogonal estimating equations are also employed in CGST18, which considers inference for high-dimensional conditional heterogeneous treatment effects. These ideas are also extended in CDDF18 which provides inference for a variety of useful functionals of heterogeneous treatment effects, such as the best linear predictor of the conditional average treatment effect function, estimated via generic high-dimensional methods using sample splitting while accounting for uncertainty introduced from the sample splits under very mild conditions. BCHN17 provides multipurpose inference methods for high-dimensional causal effects which cover endogenous treatments. AIW:bal use a reweighting after regression adjustment via lasso which allows valid inference for the average treatment effect to be performed under very weak conditions on the propensity score as long as treatment and control conditional mean functions are linear. CNR17 considers inference for linear functionals of conditional expectations with estimated Riesz representers under very weak conditions. Note that among applications of CNR17 is estimation of average treatment effects, and that the conditions of CNR17 also impose weak assumptions on the propensity score. The advantage of CNR17's approach over that in AIW:bal is that the former explicitly allows the tradeoff in the rate of estimating the inverse propensity score with the rate of estimating the regression function, allowing misspecification of both functions. The regularity conditions are also substantively weaker, for example allowing regressions functions be completely non-sparse when the inverse propensity score is well-approximated.
Most of the work in the recent literature on high-dimensional estimation and inference relies on approximate sparsity to provide dimension reduction and the corresponding use of sparsity-based estimators. Dense models are appealing in many settings and may be usefully employed in more moderate-dimensional settings. The many weak-instrument regime, popular in econometrics since at least bekker, provides one such example. Approaches which provide valid inference for structural parameters within this setting can all be viewed as making use of regularization to avoid dramatic overfitting in the relationship between endogenous variables and the many available instruments. See, for example, CI04, Okui11, Carr12, and HK14 for approaches that explicitly use regularized first-stage estimation within a dense model framework. CJN16 and CJN17 consider inference for a low-dimensional set of coefficients in a linear model with number of variables proportional to but smaller than the sample size in a framework allowing for the coefficients on the nuisance variables to be dense. Within the same framework, CJM17 extend this work to address inference about parameters estimated using two-step procedures where the first step is a linear regression with many variables.
We conclude these bibliographic notes by noting that the references to the high-dimensional literature provided above are necessarily selective. The literature on high-dimensional estimation and inference is large and rapidly expanding, and it is impractical to give more than a cursory overview highlighting a few examples. The goal of this review is to provide readers with a few key papers in several areas and a taste of existing results.