EconBase
← Back to paper

Culling the herd of moments with penalized empirical likelihood

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

134,829 characters · 10 sections · 105 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Culling the Herd of Moments with Penalized Empirical Likelihood

\bibpunct{(}{)}{,}{a}{;}

abstractModels defined by moment conditions are at the center of structural econometric estimation, but economic theory is mostly agnostic about moment selection. While a large pool of valid moments can potentially improve estimation efficiency, in the meantime a few invalid ones may undermine consistency. This paper investigates the empirical likelihood estimation of these moment-defined models in high-dimensional settings. We propose a penalized empirical likelihood (PEL) estimation and establish its oracle property with consistent detection of invalid moments. The PEL estimator is asymptotically normally distributed, and a projected PEL procedure further eliminates its asymptotic bias and provides more accurate normal approximation to the finite sample behavior. Simulation exercises demonstrate excellent numerical performance of these methods in estimation and inference.

{\it Keywords}: Empirical likelihood; Estimating equations; High-dimensional statistical methods; Misspecification; Moment selection; Penalized likelihood

\thispagestyle{empty}

\setcounter{page}{1}

spacing{1.5}

Introduction

Economists' perennial pursuit of structural mechanisms leads to models defined by moments. These models can be written in a semiparametric form $\mathbb{E}\{{\mathbf g} ({\mathbf X}_i;\boldsymbol{\theta}_0)\}={\mathbf 0},$ where ${\mathbf g} = (g_j)_{j \in \{1,\ldots, r\} }$ is a vector of $r$ estimating functions, $\boldsymbol{\theta}_0$ is a vector of unknown parameters, and ${\mathbf X}_i$ is observed data. To estimate these models, the most popular method is generalized method of moments (GMM) hansen1982large. Empirical likelihood (EL) QinLawless1994 is a competitive alternative to GMM, thanks to its nice statistical properties. Both GMM and EL are essential building blocks of modern econometrics anatolyev2011methods.

Ideally, economists count on economic theory to guide the choice of variables and moments. However, the truth is that most economic theories are parsimonious abstractions and rarely pinpoint these choices in data-rich environments. The indeterminacy of moment selection brings about three related issues. The first is weak moments stock2012survey, which threatens identification of the true value $\boldsymbol{\theta}_0$ when multiple $\boldsymbol{\theta}$'s satisfy $\mathbb{E}\{{\mathbf g} ({\mathbf X}_i;\boldsymbol{\theta})\} \approx {\mathbf 0}$. Practitioners respond to the concerns of weak moments by adding more moment conditions in the hope to strengthen identification, causing the second issue of many moments roodman2009note. The hazard of many moments is the possible inclusion of invalid moments, meaning $\mathbb{E}\{g_j ({\mathbf X}_i;\boldsymbol{\theta}_0)\}\neq 0$ for some $j\in \{1,\ldots,r\}$ murray2006avoiding, which is the third issue. These cited survey papers highlight the unease incurred by the three challenges and econometricians' efforts in coping with them.

When an underlying economic theory is ambivalent, empirical results based on it can be controversial and susceptible to cherry-picking. In such a circumstance, of great importance are data-driven methods to guide and discipline moment selection. In the low-dimensional settings, andrews2001consistent propose the GMM information criteria, and hong2003generalized follow with the counterpart for EL. Information criteria are evaluated exhaustively at all combinations of moments, and the computation becomes infeasible when there are many potential moments. To overcome this challenge, liao2013adaptive ushers the adaptive Lasso shrinkage into GMM to select among a finite number of moments, and ChengLiao2015 further extend it to deal with a diverging number of moments and accommodate invalid ones.

Unprecedented progress in computation and information technology fuel an arms race between the sheer size of the data and the scale of empirical models. In the era of big data, on the one hand the cost of data collection and processing is tremendously lowered and rich datasets open new perspectives to inspect a myriad of problems; on the other hand economists attempt to build general models to capture various sources of heterogeneity in observational data. Empirical applications abound with models of many potential moments. For instance, eaton2011 create 1360 moments to estimate a structural trade model, altonji2013modeling match 2429 moments implied by a model of earning dynamics, and early works of linear instrumental variables (IV) models produce thousands of instruments by interacting variables angrist1992effect. Although the sample sizes in these examples are non-trivial, the proliferation of moments calls for a moment selection procedure capable of handling high-dimensional moments at a magnitude unrestricted by the sample size. In particular, invalid ones that jeopardize consistency must be identified and “culled” from the herd of moments.

The quadratic form of the usual GMM criterion function is incompatible with high-dimensional moments shi2016econometric, shi2016estimation, and thus belloni2018high regularize GMM with the sup-norm. In this paper, we contribute the most general and versatile procedure for high-dimensional nonlinear settings, to the best of our knowledge. We first develop a penalized empirical likelihood (PEL) solution ChangTangWu2018 to deal with valid and invalid moments simultaneously. We neutralize the invalid moments by an auxiliary parameter, following liao2013adaptive and ChengLiao2015. We establish the rate of convergence and asymptotic normality of the PEL estimator, and show the efficiency gain from incorporating extra valid moments. Under suitable conditions, it transpires that PEL enjoys the oracle property of consistent moment selection and parameter selection.

The asymptotic normal distribution of the PEL estimator involves a bias term caused by the high-dimensional moments. To spare the estimation of the bias term, we can take further actions to project out the influence of the high-dimensional nuisance parameter in the PEL estimator, which is called projected PEL (PPEL) Chang2020. The asymptotic normality of the PPEL estimator is free of bias, which facilitates statistical inference of the structural parameter as well as the validity of moments. Invoking statistical learning to assist our decisions, our method fits well in the recent trend of machine learning for the automatic selection of moments and variables.

Although this paper follows ChangTangWu2018 for estimation and Chang2020 for inference procedures, the key insight lies in the observations that the high-dimensional auxiliary parameter, which signifies the magnitude of misspecification, can be incorporated in the EL method as an additional high-dimensional parameter. In this paper, “high-dimensional” means that the numbers of parameters and/or moments are larger than the sample size, which goes beyond the scope of liao2013adaptive and ChengLiao2015. On the other hand, in order to adapt ChangTangWu2018 and Chang2020 to accommodate misspecified moments, we must deal with the distinctive roles of the main parameter of interest and the auxiliary parameter in identification and the technical challenges induced by them. Differences from ChangTangWu2018 are highlighted in Section (ref) about the rates of converges of the two components of the parameters, and those from Chang2020 are elaborated in Section (ref) about the ways of confidence region construction.

Literature review. Our paper stands on strands of literature, which are too vast to survey exhaustively. The accumulation of moments started from the linear IV model angrist1990lifetime, angrist1991does. The linear IV model motivates theoretical research on issues of many IV bekker1994alternative, weak IV (stock2005testing, andrews2012estimation), many weak IV chao2011asymptotic, hansen2014instrumental, invalid moments and many invalid IV kolesar2015identification,windmeijer2019use, to name a few. In high-dimensional contexts, belloni2012sparse use Lasso method for IV selection in the first stage, and belloni2014inference deal with post-selection inference. Utilizing the linear structure, gold2020inference and caner2018high provide inferential procedures for low-dimensional parameters in models with high-dimensional endogenous variables and high-dimensional IVs. Our method includes the linear IV model as a special case. In particular, the case of high-dimensional structural parameters is elaborated in Section (ref).

The proliferation of moments spreads from linear IV models to nonlinear models. For example, in empirical industrial organization researchers bring in moments from various resources, some of which are guided by economic theory, to mitigate the concerns of weak identification and improve estimation efficiency ackerberg2007econometric. In empirical macroeconomics, identification failure and moment misspecification are common issues mavroeidis2005identification. Under the GMM framework, inference under weak moments (stock2000gmm, kleibergen2005testing, andrews2020optimal), estimation under many weak moments han2006gmm, and robust procedures for invalid moments (ditraglia2016using, caner2018adaptive) have been developed.

EL's attractive theoretical properties are studied extensively (kitamura2001asymptotic, otsu2010bahadur, matsushita2013second, ChangChenChen2015). otsu2006generalized and newey2009generalized deal with its inference under weak IV, and caner2015hybrid select instruments in linear IV models. Penalization schemes on EL have been introduced by otsu2007penalized, tang2010penalized and ChangTangWu2018.

Organization. The rest of the paper is organized as follows. Section (ref) introduces the model and the analytic framework. We first derive the asymptotic properties of our estimation in the low-dimensional case in Section 3, extend them to the high-dimensional structural parameter in Section 4, and we further refine PEL with projection to eliminate its bias in Section 5. The theoretical results are supported in Section 6 by Monte Carlo simulations. The influential study of the determinants of economic outcomes after colonialism is revisited in Section 7. Section 8 concludes the paper. Due to the limitations of space, the proofs and the technical details of secondary importance are relegated into the supplementary materials.

Notations. We conclude Introduction with notations used throughout the paper. “Low-dimensional” is referred to the cases that the number of parameters or moments is much smaller than the sample size $n$, whereas “high-dimensional” goes the opposite. For two sequences of positive numbers $\{a_n\}$ and $\{b_n\}$, we write $a_n\lesssim b_n$ or $b_n\gtrsim a_n$ if there exists a positive constant $c$ such that $\limsup_{n\rightarrow\infty}a_n/b_n\leq c$, and write $a_n\ll b_n$ or $b_n\gg a_n$ if $\limsup_{n\rightarrow\infty}a_n/b_n=0$.

Denote by $1(\cdot)$ the indicator function. For a positive integer $q$, we write $[q]=\{1,\ldots,q\}$. For a $q\times q$ symmetric matrix ${\mathbf M}$, denote by $\lambda_{\min}({\mathbf M})$ and $\lambda_{\max}({\mathbf M})$ the smallest and largest eigenvalues of ${\mathbf M}$, respectively. For a $q_1\times q_2$ matrix ${\mathbf B}=(b_{i,j})_{q_1\times q_2}$, let ${\mathbf B}^{ \mathrm{\scriptscriptstyle T} }$ be its transpose, ${\mathbf B}^{\otimes2}={\mathbf B}{\mathbf B}^{ \mathrm{\scriptscriptstyle T} }$, ${\mathbf B}^{\circ\kappa}=(|b_{i,j}|^\kappa)_{q_1\times q_2}$ for any $\kappa>0$, $|{\mathbf B}|_\infty=\max_{i\in[q_1],j\in[q_2]}|b_{i,j}|$ be the sup-norm, and $\|{\mathbf B}\|_2=\lambda_{\max}^{1/2}({\mathbf B}^{\otimes2})$ be the spectral norm. Specifically, if $q_2=1$, we use $|{\mathbf B}|_\infty = \max_{i\in[q_1]}|b_{i,1}|$, $|{\mathbf B}|_1=\sum_{i=1}^{q_1}|b_{i,1}|$ and $|{\mathbf B}|_2=(\sum_{i=1}^{q_1}b_{i,1}^2)^{1/2}$ to denote the $L_{\infty}$-norm, $L_1$-norm and $L_2$-norm of the $q_1$-dimensional vector ${\mathbf B}$, respectively. For two square matrices ${\mathbf M}_1$ and ${\mathbf M}_2$, we say ${\mathbf M}_1 \le {\mathbf M}_2$ if $({\mathbf M}_2 - {\mathbf M}_1)$ is a positive semi-definite matrix.

The population mean is denoted by $\mathbb{E}(\cdot)$, and the sample mean by $\mathbb{E}_n(\cdot)= n^{-1}\sum_{i=1}^n \{\cdot\}$. For a given index set $\mathcal{L}$, let $| \mathcal{L} | $ be its cardinality. For a generic multivariate function ${\mathbf h}(\cdot;\cdot)$, we denote by ${\mathbf h}_{{ \mathcal{\scriptscriptstyle L} }}(\cdot;\cdot)$ the subvector of ${\mathbf h}(\cdot;\cdot)$ collecting the components indexed by $\mathcal{L}$. Analogously, we write ${\mathbf a}_{{ \mathcal{\scriptscriptstyle L} }}$ as the corresponding subvector of ${\mathbf a}$. For simplicity and when no confusion arises, we use the generic notation ${\mathbf h}_i(\boldsymbol{\theta})$ as the equivalence to ${\mathbf h}({\mathbf X}_i;\boldsymbol{\theta})$, and $\nabla_{\boldsymbol{\theta}} {\mathbf h}_i(\boldsymbol{\theta}) $ for the first-order partial derivative of ${\mathbf h}_i(\boldsymbol{\theta})$ with respect to $\boldsymbol{\theta}$. Denote by $h_{i,k}(\boldsymbol{\theta})$ the $k$-th component of ${\mathbf h}_i(\boldsymbol{\theta})$, and by $\nabla^2_{\boldsymbol{\theta}}h_{i,k}(\boldsymbol{\theta}) $ the second derivative of $h_{i,k}(\boldsymbol{\theta})$ with respect to $\boldsymbol{\theta}$. Let $\bar {\mathbf h}(\boldsymbol{\theta})=\mathbb{E}_n\{{\mathbf h}_i (\boldsymbol{\theta})\}$, and write its $k$-th component as $\bar h_k(\boldsymbol{\theta})= \mathbb{E}_n\{h_{i,k}(\boldsymbol{\theta})\}$. Analogously, let ${\mathbf h}_{i,{ \mathcal{\scriptscriptstyle L} }}(\boldsymbol{\theta})={\mathbf h}_{{ \mathcal{\scriptscriptstyle L} }}({\mathbf X}_i;\boldsymbol{\theta})$ and $\bar{{\mathbf h}}_{{ \mathcal{\scriptscriptstyle L} }}(\boldsymbol{\theta})= \mathbb{E}_n \{{\mathbf h}_{i,{ \mathcal{\scriptscriptstyle L} }}(\boldsymbol{\theta})\}$.

Empirical likelihood with a herd of moments

In this section we introduce the model and the EL estimation. Let ${\mathbf X}_1,\ldots,{\mathbf X}_n$ be $d$-dimensional independent and identically distributed generic observations, and $\boldsymbol{\theta}=(\theta_1,\ldots,\theta_p)^{ \mathrm{\scriptscriptstyle T} }$ be a $p$-dimensional parameter taking values in $\boldsymbol{\Theta} \subset \mathbb{R}^p$. For a set of $r_1$ estimating functions $ {\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}(\cdot;\cdot)=\{ g_{j}^{ \mathcal{\scriptscriptstyle (I)} } (\cdot; \cdot) \}_{j \in \mathcal{I}} $, the information of the model parameter $\boldsymbol{\theta}$ is collected by the unbiased moment condition

equation[equation omitted — 247 chars of source]

at the unknown true parameter $\boldsymbol{\theta}_0\in\boldsymbol{\Theta}$, where $r_1\geq p$ is necessary for identifying $\boldsymbol{\theta}_0$. The superscript $(\mathcal{I})$ labels this initial set, to be distinguished from the other set $(\mathcal{D})$ in Section (ref).

EL estimation with valid moments

Motivated from empirical applications in asset pricing, the two-step GMM hansen1982large was the default estimating method for the moment-defined model ((ref)). Intensive theoretical studies and numerical evidence in 1980's and 90's revealed some undesirable finite-sample properties of the two-step GMM altonji1996small. EL and the continuously updating GMM (CUE) hansen1996finite emerged as competitive solutions, and they were later unified as members of generalized empirical likelihood newey2003higher.

This paper focuses on EL with estimating equations:

equation*[equation* omitted — 237 chars of source]

proposed by QinLawless1994 based on the seminal idea of EL Owen1988,Owen1990. Maximizing $L(\boldsymbol{\theta})$ with respect to $\boldsymbol{\theta}$ delivers the EL estimator $ \hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }}=\arg\max_{\boldsymbol{\theta}\in\boldsymbol{\Theta}}L(\boldsymbol{\theta})$, which can be carried out equivalently by solving the corresponding dual problem

equation[equation omitted — 454 chars of source]

where $\hat{\Lambda}_{n}^{{ \mathcal{\scriptscriptstyle (I)} }}(\boldsymbol{\theta})=\{\boldsymbol{\lambda}\in\mathbb{R}^{r_1}:\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}_i(\boldsymbol{\theta})\in\mathcal {V}~\textrm{for any}~i\in[n]\}$ and $\mathcal{V}$ is an open interval containing zero.

To fix ideas, we first study the asymptotic property of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }}$ under regularity conditions. When the sample size $n$ grows, we adopt the asymptotic framework of Hjort2009 and ChangChenChen2015 to take the observations $\{{\mathbf g}_i^{{ \mathcal{\scriptscriptstyle (I)} }}(\boldsymbol{\theta})\}_{i=1}^n$ as a multi-index array, where $r_1$, $p$ and $d$ may depend on $n$. Proposition A.1 in the supplementary materials shows that the standard asymptotic normality for $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }}$ holds under mild regularity conditions.

Oracle EL estimation

In applied econometrics, there has been a tendency of assembling many IVs or creating many moments to identify the parameter of interest. In these applications, researchers often have some ideas about the relative importance of moments; in the meantime, when researchers work with more and more moments, some invalid ones may creep in.

exampleeaton2011 deem as the key moments 128($=2^7$) combinations of the largest 7 trade partners of France, which are more important than the other 1232 moments. angrist1990lifetime treats the date of birth (DOB) as the key IV and it is complemented by the IVs generated by interactions; recently kolesar2015identification raise the potential invalidity among these interaction terms.\footnote{Identification of the simple linear IV model requires the IVs satisfying the orthogonality condition and the relevance condition. However, the more relevant an IV to the endogenous variables, the more likely it is that the so-called “IV” is correlated with the structural error, thereby violating orthogonality. There is a thin line between a valid IV and an invalid one. }

To put into an analytic framework a herd of extra moments with unknown validity ex ante, suppose that $r_2$ estimating functions ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}=\{g_{j}^{ \mathcal{\scriptscriptstyle (D)} }\}_{j\in \mathcal{D}}$ are partitioned into two groups: \[ \mathcal{A}=\{j\in \mathcal{D}:\mathbb{E}\{g^{ \mathcal{\scriptscriptstyle (D)} }_{i,j}(\boldsymbol{\theta}_0)\}=0\}~~~\textrm{and}~~~ \mathcal{A}^{{\mathrm{c}}}=\{j\in \mathcal{D}:\mathbb{E}\{g^{ \mathcal{\scriptscriptstyle (D)} }_{i,j}(\boldsymbol{\theta}_0)\}\neq 0\}\,. \] Here the estimating functions in the set $\mathcal{A}$ are correctly specified, which can help improve the efficiency in estimating $\boldsymbol{\theta}_0$. On the contrary, those in $\mathcal{A}^{{\mathrm{c}}}$ are misspecified, and they can only undermine the identification of $\boldsymbol{\theta}_0$.

If there is an “oracle” that reveals which components of ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}$ belong to $\mathcal{A}$, i.e., the valid estimating functions, we can collect all valid estimating functions indexed by $\mathcal{H}:=\mathcal{I}\cup\mathcal{A}$ to estimate $\boldsymbol{\theta}_0$. Denote ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (H)} }} =\{ {\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} },{ \mathrm{\scriptscriptstyle T} }},{\mathbf g}^{{ \mathcal{\scriptscriptstyle (A)} },{ \mathrm{\scriptscriptstyle T} }}\}^{ \mathrm{\scriptscriptstyle T} }$ and $h=|\mathcal{H}|$ as the total number of valid moments. The associated EL estimation for $\boldsymbol{\theta}_0$ based on the estimating functions ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (H)} }}$ is given by

equation*[equation* omitted — 448 chars of source]

where $\hat{\Lambda}_{n}^{{ \mathcal{\scriptscriptstyle (H)} }}(\boldsymbol{\theta})=\{\boldsymbol{\lambda}\in\mathbb{R}^h:\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (H)} }}_i(\boldsymbol{\theta})\in\mathcal {V}~\textrm{for any}~i\in[n]\}$. By the same arguments as those in Proposition A.1 in the supplementary materials, the following Proposition (ref) verifies the asymptotic normality for $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (H)} }}$.

proposition[Oracle property] Assume that ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (H)} }} $ satisfies the conditions {\rm(A.1)--(A.5)} in the supplementary materials associated with ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (H)} }}$. If $h^{3}n^{-1+2/\gamma}=o(1)$ and $h^{3}p^2n^{-1}=o(1)$, then $ \sqrt{n}\boldsymbol{\alpha}^{ \mathrm{\scriptscriptstyle T} }\{{\mathbf J}^{{ \mathcal{\scriptscriptstyle (H)} }}\}^{1/2}\{\hat{\boldsymbol{\theta}}_{{{ \mathrm{\scriptscriptstyle EL} }}}^{{ \mathcal{\scriptscriptstyle (H)} }}-\boldsymbol{\theta}_0\}\xrightarrow{d}\mathcal{N}(0,1)$ as $n\rightarrow\infty$ for any $\boldsymbol{\alpha}\in\mathbb{R}^p$ with $|\boldsymbol{\alpha}|_2=1$, where $ {\mathbf J}^{{ \mathcal{\scriptscriptstyle (H)} }}=([\mathbb{E}\{\nabla_{\boldsymbol{\theta}}{\mathbf g}^{{ \mathcal{\scriptscriptstyle (H)} }}_i(\boldsymbol{\theta}_0)\}]^{ \mathrm{\scriptscriptstyle T} }\{{\mathbf V}^{{ \mathcal{\scriptscriptstyle (H)} }}(\boldsymbol{\theta}_0)\}^{-1/2})^{\otimes2} $ with ${\mathbf V}^{{ \mathcal{\scriptscriptstyle (H)} }}(\boldsymbol{\theta}_0)=\mathbb{E}\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (H)} }}_i(\boldsymbol{\theta}_0)^{\otimes2}\}$.

Obviously $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (H)} }}$ is more efficient than $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }}$ thanks for the additional moment restrictions in $\mathcal{A}$, which disclose more information about the parameter and thus tie down the estimation variability. This is parallel to hall2007information in GMM for finite numbers of parameters and moments. In reality, however, we are oblivious to $\mathcal{A}$, and therefore we cannot blindly summon all $r_2$ estimating functions in ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}$ to estimate $\boldsymbol{\theta}_0$. Following liao2013adaptive, we introduce an auxiliary parameter $\boldsymbol{\xi}=\mathbb{E}\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}_i(\boldsymbol{\theta})\}$ to neutralize the misspecified moments. Denote the augmented parameter $\boldsymbol{\psi}=(\boldsymbol{\theta}^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$, and the augmented parameter space $\boldsymbol{\Psi} = \boldsymbol{\Theta}\times\mathbf{\Upsilon}$. Let $r=r_1+r_2$ and stack these $r$ estimating functions as ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}({\mathbf X};\boldsymbol{\psi})=\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}({\mathbf X};\boldsymbol{\theta})^{ \mathrm{\scriptscriptstyle T} },{\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}({\mathbf X};\boldsymbol{\theta})^{ \mathrm{\scriptscriptstyle T} }-\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} }\}^{ \mathrm{\scriptscriptstyle T} }$. Then $\boldsymbol{\psi}_0=(\boldsymbol{\theta}_0^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\xi}_0^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ with $\boldsymbol{\xi}_0=\mathbb{E}\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}_i(\boldsymbol{\theta}_0)\}$ can be identified by

equation[equation omitted — 354 chars of source]

Can we directly include the auxiliary parameter $\boldsymbol{\xi}$ into the EL estimation? The answer is negative. If the auxiliary parameter is not regularized, the EL estimators with and without $\boldsymbol{\xi}$ are the same, up to numerical errors.

propositionIf $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }} $ is the unique solution of {\rm(ref)}, then $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (T)} }} = \hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }} $, where $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (T)} }}$ is the associated subvector of $\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (T)} }} =\arg\min_{\boldsymbol{\psi}\in \boldsymbol{\Psi} }\max_{\boldsymbol{\lambda}\in\hat{\Lambda}_{n}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi}) }\sum_{i=1}^n\log\{1+\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\boldsymbol{\psi})\}$ for the estimation of $\boldsymbol{\theta}_0$ with $\hat{\Lambda}_{n}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})=\{\boldsymbol{\lambda}\in\mathbb{R}^r:\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\boldsymbol{\psi})\in\mathcal {V}~\textrm{for any}~i\in[n]\}$.

Proposition (ref) states the equivalence between $\hat\boldsymbol{\theta}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathrm{\scriptscriptstyle EL} }}$ and $\hat\boldsymbol{\theta}^{{ \mathcal{\scriptscriptstyle (I)} }}_{{ \mathrm{\scriptscriptstyle EL} }}$. In the next section, we will show that desirable efficiency as in the oracle estimator $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (H)} }}$ can be achieved if we extend liao2013adaptive's idea of shrinking the auxiliary parameter\footnote{ In GMM estimation under a fixed $r$, liao2013adaptive shrinks $\boldsymbol{\xi}$ toward zero using the adaptive Lasso zou2006adaptive. } by further penalizing the Lagrange multipliers associated with the high-dimensional moments.

High-dimensional moments with low-dimensional $\boldsymbol{\theta}$

We consider in this section the model (ref) with a low-dimensional parameter $\boldsymbol{\theta}$, low-dimensional estimating functions ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}$ and high-dimensional estimating functions ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}$. In this setting, $p$ and $r_1$ are either fixed or diverge at some slow polynomial rates of $n$, whereas $r_2$ can grow much larger than $n$. The components in the parameter $\boldsymbol{\theta}$ are indexed by $\mathcal{P}$. All moments together are indexed by $\mathcal{T} = \mathcal{I} \cup \mathcal{D}$, in which the researcher knows that those in $\mathcal{I}$ are correctly specified in the sense that $\mathbb{E}\{ {\mathbf g}_i^{{ \mathcal{\scriptscriptstyle (I)} }}(\boldsymbol{\theta}_0)\} = {\mathbf 0}$, whereas she is uncertain whether those in $\mathcal{D} = \mathcal{A}\cup \mathcal{A}^{{\mathrm{c}}} $ satisfy $\mathbb{E}\{ {\mathbf g}_i^{{ \mathcal{\scriptscriptstyle (D)} }}(\boldsymbol{\theta}_0)\} = {\mathbf 0} $ or not. As there is a one-to-one relationship between $\boldsymbol{\xi}$ and the uncertain moments, we use $\mathcal{D}$ to index the components in $\boldsymbol{\xi}$ as well.

We propose the following PEL to simultaneously estimate the unknown parameter $\boldsymbol{\theta}_0$ and determine the validity of the estimating functions in ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}$:

equation[equation omitted — 688 chars of source]

where $\hat{\Lambda}_{n}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})=\{\boldsymbol{\lambda}\in\mathbb{R}^{r}:\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\boldsymbol{\psi})\in\mathcal {V}~\textrm{for any}~i\in[n]\}$, and $P_{1,\pi}(\cdot)$ and $P_{2,\nu}(\cdot)$ are two penalty functions with tuning parameters $\pi$ and $\nu$, respectively. With the penalty function $P_{2,\nu}(\cdot)$ and appropriately selected tuning parameter $\nu$, the estimator $(\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} },\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ is associated with a sparse Lagrange multiplier $\boldsymbol{\lambda}$. Since the sparse $\boldsymbol{\lambda}$ invokes a subset of the estimating functions ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}(\cdot;\cdot)$, it digests the high-dimensional moments as long as the number of nonzero components in $\boldsymbol{\lambda}$ is small. On the other hand, the penalty $P_{1,\pi}(\cdot)$ is applied to identify which components in ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}(\cdot;\cdot)$ are correctly specified and can estimate $(\xi_{0,k})_{k\in\mathcal{A}}$ exactly as $0$ with high probability.

These penalties originally appeared in ChangTangWu2018 though, there are several important differences. First, given the low-dimensional $\boldsymbol{\theta}$, it is unnecessary to assume sparsity on $\boldsymbol{\theta}_0$ for identification. As a result, our procedure here in (ref) only penalizes $\boldsymbol{\xi}$ and $\boldsymbol{\lambda}_{\mathcal{D}}$, not the entire parameter $\boldsymbol{\psi}=(\boldsymbol{\theta}^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ and the associated Lagrange multiplier $\boldsymbol{\lambda}$. Were $\boldsymbol{\lambda}_{\mathcal{I}}$ penalized, we would rule out some components of ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}(\cdot;\cdot)$ and thus may suffer loss of estimation efficiency for $\boldsymbol{\theta}_0$. Second, the identification condition invoked in this paper deviates from that in ChangTangWu2018. Our Condition (ref) for identification is imposed on $\boldsymbol{\theta}_0$, rather than the whole parameter $\boldsymbol{\psi}_0$. In contrast, ChangTangWu2018 specify identification condition for the parameter entity. While estimators in ChangTangWu2018 share the same rate of convergence, we must deal with the disparate rates of convergence of the main component $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ and the auxiliary $\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$, which significantly complicates the theoretical analysis of (ref), as detailed in Proposition (ref) below.

For any penalty function $P_\tau(\cdot)$ with a tuning parameter $\tau$, let $\rho(t;\tau)=\tau^{-1}P_\tau(t)$ for any $t\in[0,\infty)$ and $\tau\in(0,\infty)$. We assume that the two penalty functions $P_{1,\pi}(\cdot)$ and $P_{2,\nu}(\cdot)$ involved in (ref) belong to the following class:

equation[equation omitted — 336 chars of source]

This class $\mathscr{P}$, considered in LvFan2009, is broad and general. The commonly used $L_1$ penalty, SCAD penalty FanLi2001 and MCP penalty Zhang2010 are all included in $\mathscr{P}$. When $P_\tau(\cdot)\in\mathscr{P}$, we write the associated $\rho'(0^+;\tau)$ as $\rho'(0^+)$ for simplification.

To establish the limiting distribution of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ in (ref), we assume the following regularity conditions.

asThere exists a universal constant $K_1>0$ such that $$\inf_{\boldsymbol{\theta}\in\boldsymbol{\Theta}:\,|\boldsymbol{\theta}-\boldsymbol{\theta}_{0}|_\infty>\varepsilon}|\mathbb{E}\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}_i(\boldsymbol{\theta})\}|_\infty\ge K_1 \varepsilon$$ for any $\varepsilon>0$.
asThere exist universal constants $K_2>0$ and $\gamma>4$ such that $$ \max_{j\in\mathcal{I}}\mathbb{E}\bigg\{\sup_{\boldsymbol{\theta}\in\boldsymbol{\Theta}}|g^{{ \mathcal{\scriptscriptstyle (I)} }}_{i,j}(\boldsymbol{\theta})|^\gamma\bigg\}+\max_{j\in\mathcal{D}}\mathbb{E}\bigg\{\sup_{\boldsymbol{\theta}\in\boldsymbol{\Theta}}|g^{{ \mathcal{\scriptscriptstyle (D)} }}_{i,j}(\boldsymbol{\theta})|^\gamma\bigg\}\le K_2\,.$$
asFor each $j\in \mathcal{T}$, the function $g^{{ \mathcal{\scriptscriptstyle (T)} }}_{j}({\mathbf X};\boldsymbol{\psi})$ is twice continuously differentiable with respect to $\boldsymbol{\psi} \in\boldsymbol{\Psi}$ for any ${\mathbf X}$. For $\gamma$ specified in Condition (ref), $$ \sup_{\boldsymbol{\psi}\in\boldsymbol{\Psi}} \bigg( | \mathbb{E}_n[\{\nabla_{\boldsymbol{\psi}} {\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_{i}(\boldsymbol{\psi})\}^{\circ2}] |_{\infty} + \max_{j\in \mathcal{T}} | \mathbb{E}_n[\{\nabla^2_{\boldsymbol{\psi}}g^{{ \mathcal{\scriptscriptstyle (T)} }}_{i,j}(\boldsymbol{\psi}) \}^{\circ2}] |_{\infty} + | \mathbb{E}_n[\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_{i}(\boldsymbol{\psi})\}^{\circ\gamma}]|_{\infty}\bigg) =O_{ \mathrm{p} }(1)\,.$$ There is some universal constant $K_3>0$ such that $\sup_{\boldsymbol{\theta}\in\boldsymbol{\Theta}} |\mathbb{E} \{ \nabla_{\boldsymbol{\theta}} {\mathbf g}_{i,{ \mathcal{\scriptscriptstyle A} }^{{\mathrm{c}}}}^{ \mathcal{\scriptscriptstyle (D)} }(\boldsymbol{\theta}) \} |_\infty \le K_3$.

For any index set $\mathcal{F}\subset\mathcal{T}$ and $\boldsymbol{\psi}\in\boldsymbol{\Psi}$, define $ {\mathbf V}_{{ \mathcal{\scriptscriptstyle F} }}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})=\mathbb{E}\{{\mathbf g}_{i,{ \mathcal{\scriptscriptstyle F} }}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})^{\otimes2}\}$. When $\mathcal{F}=\mathcal{T}$, we write ${\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})={\mathbf V}_{{ \mathcal{\scriptscriptstyle T} }}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})$ for conciseness.

asThere exists a universal constant $K_4>1$ such that $K_4^{-1}<\lambda_{\rm min}\{{\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi}_0)\}\le\lambda_{\rm max}\{{\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi}_0)\}<K_4$.

Conditions (ref)--(ref) are standard regularity assumptions in the literature. Condition (ref) is an identification assumption of the estimating equations in the known set $\mathcal{I}$. Write $\boldsymbol{\xi}_0= (\xi_{0,k})_{k\in\mathcal{D}}$, and we continue by defining

align[align omitted — 141 chars of source]

in this section, where $\aleph_n=(n^{-1}\log r)^{1/2}$. Suppose:

equation[equation omitted — 252 chars of source]

to control the bias induced by $P_{1,\pi}(\cdot)$ on $\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$. With the assumption $\phi_n=o(\min_{k\in\mathcal{A}^{{\mathrm{c}}}}|\xi_{0,k}|)$ that the nonzero components of $\boldsymbol{\xi}_0$ do not diminish to zero too fast, (ref) can be replaced by

align[align omitted — 125 chars of source]

for some constant $c\in(0,1)$. If we select $P_{1,\pi}(\cdot)$ as an asymptotically unbiased penalty such as SCAD or MCP, we have $\chi_n=0$ in (ref) when

align[align omitted — 102 chars of source]

To simplify the presentation, in this section we assume that (ref) holds and $\chi_n=0$ in (ref).\footnote{ If (ref) is violated, the asymptotic normality of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ in Theorem (ref) below will still hold under (ref) along with more complicated notations to spell out the restrictions. }

To allocate the parameter of interest $\boldsymbol{\theta}$ and those invalid moments in $\mathcal{A}^{{\mathrm{c}}}$, we define an index set $\mathcal{S}=\mathcal{P}\cup \mathcal{A}^{{\mathrm{c}}}$ and its cardinality $s=|\mathcal{S}|$. For any $\boldsymbol{\psi} \in\boldsymbol{\Psi}$, the set $\mathcal{S}$ picks out $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }}=(\boldsymbol{\theta}^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\xi}_{{ \mathcal{\scriptscriptstyle A} }^{{\mathrm{c}}}}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$. Since $\mathcal{A} = \mathcal{S}^{{\mathrm{c}}} = (\mathcal{P}\cup\mathcal{D}) \backslash \mathcal{S}$, the corresponding auxiliary coefficients can be written as $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}=\boldsymbol{\xi}_{{ \mathcal{\scriptscriptstyle A} }}$ and furthermore $\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}={\mathbf 0}$ as $\mathcal{A}$ is the set of valid moments. Given some constant $C_*\in(0,1)$, define $\mathcal{M}^*_{\boldsymbol{\psi}}= \mathcal{I} \cup \mathcal{D}^*_{\boldsymbol{\psi}}$ with $\mathcal{D}^*_{\boldsymbol{\psi}}=\{ j\in \mathcal{D}: |\bar{g}^{{ \mathcal{\scriptscriptstyle (T)} }}_{j}(\boldsymbol{\psi})|\ge C_*\nu\rho'_2(0^+)\}$. We assume the existence of a sequence $\ell_n\to\infty$ such that $$ \mathbb{P}\bigg(\max_{\boldsymbol{\psi}\in\boldsymbol{\Psi}:\,|\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }}-\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle S} }}|_\infty\le c_n,\,|\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}|_1\leq \aleph_n }|\mathcal{M}^*_{\boldsymbol{\psi}}| \le \ell_n\bigg) \to 1 $$ with some $c_n\gg\phi_n$.

Let $b_{1,n}=\max\{a_n,r_{1}\aleph_n^2\}$ and $b_{2,n}=\max\{b_{1,n},\nu^2 \}$. Then $ \phi_n=\max\{pb_{1,n}^{1/2}, b_{2,n}^{1/2} \}$. Define $\boldsymbol{\Psi}_*=\{\boldsymbol{\psi}=(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }}^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }:|\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }}-\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle S} }}|_\infty\le\varepsilon,|\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}|_1\le \aleph_n \}$ for some fixed $\varepsilon>0$. Consider

align[align omitted — 476 chars of source]

Proposition (ref) shows that such a $\hat{\boldsymbol{\psi}}$ is a sparse local minimizer for (ref).

propositionLet $P_{1,\pi}(\cdot),P_{2,\nu}(\cdot)\in\mathscr{P}$ for $\mathscr{P}$ defined in {\rm((ref))}, and $P_{2,\nu}(\cdot)$ be convex with bounded second derivative around $0$. For $\hat\boldsymbol{\psi}$ defined as {\rm(ref)}, assume there exists a constant $\tilde{c}\in(C_*,1)$ such that $\mathbb{P}[\cup_{j\in\mathcal{T}}\{|\bar{g}_j^{ \mathcal{\scriptscriptstyle (T)} }(\hat\boldsymbol{\psi})|\in[\tilde{c}\nu\rho'_2(0^+),\nu\rho'_2(0^+))\}]\rightarrow0$. Under Conditions {\rm (ref)--(ref)} and (ref), if $\log r=o(n^{1/3})$, $s^2\ell_n\phi_n^{2}=o(1)$, $b_{2,n}=o(n^{-2/\gamma})$, $\ell_n\aleph_n=o(\nu)$ and $\ell_n^{1/2}\aleph_n=o(\pi)$, then with probability approaching one this $\hat\boldsymbol{\psi}=(\hat\boldsymbol{\theta}^{ \mathrm{\scriptscriptstyle T} },\hat\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} }_{{ \mathcal{\scriptscriptstyle A} }^{{\mathrm{c}}}},\hat\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} }_{ \mathcal{\scriptscriptstyle A} })^{ \mathrm{\scriptscriptstyle T} }$ provides a sparse local minimizer for the nonconvex optimization (ref) such that {\rm(i)} $|\hat{\boldsymbol{\theta}}-\boldsymbol{\theta}_0|_\infty=O_{ \mathrm{p} }(b_{1,n}^{1/2})$, {\rm(ii)} $|\hat{\boldsymbol{\xi}}_{{ \mathcal{\scriptscriptstyle A} }^{{\mathrm{c}}}}-\boldsymbol{\xi}_{0,{ \mathcal{\scriptscriptstyle A} }^{{\mathrm{c}}}}|_\infty=O_{ \mathrm{p} }(\phi_n) $, and {\rm(iii)} $\mathbb{P}(\hat{\boldsymbol{\xi}}_{{ \mathcal{\scriptscriptstyle A} }}={\mathbf 0})\to 1$ as $n\to\infty$.

In the rest of this section, we focus on the sparse local minimizer $\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }}=(\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} },\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ specified in (ref). We proceed with additional regularity conditions.

asThere exists a universal constant $K_5>1$ such that $K_5^{-1}<\lambda_{\rm min}\{{\mathbf Q}_{{ \mathcal{\scriptscriptstyle I} }\cup{ \mathcal{\scriptscriptstyle B} }_1,{ \mathcal{\scriptscriptstyle B} }_2}^{{ \mathcal{\scriptscriptstyle (T)} }}\}\le\lambda_{\rm max}\{{\mathbf Q}_{{ \mathcal{\scriptscriptstyle I} }\cup{ \mathcal{\scriptscriptstyle B} }_1,{ \mathcal{\scriptscriptstyle B} }_2}^{{ \mathcal{\scriptscriptstyle (T)} }}\}<K_5$ for any $\mathcal{B}_2\subset \mathcal{B}_1 \subset \mathcal{D}$ with $p+|\mathcal{B}_2| \le r_1+|\mathcal{B}_1| \le \ell_n$, where ${\mathbf Q}_{{ \mathcal{\scriptscriptstyle I} }\cup{ \mathcal{\scriptscriptstyle B} }_1,{ \mathcal{\scriptscriptstyle B} }_2}^{{ \mathcal{\scriptscriptstyle (T)} }}=([\mathbb{E}\{\nabla_{\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle P}}\cup{ \mathcal{\scriptscriptstyle B} }_2}}{\mathbf g}_{i,{ \mathcal{\scriptscriptstyle I} }\cup{ \mathcal{\scriptscriptstyle B} }_1}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi}_0)\}]^{ \mathrm{\scriptscriptstyle T} })^{\otimes2}$.

Condition (ref) is the sparse Riesz condition (zhang2008sparsity, chen2008extended) in our setting to deal with the high-dimensional $\boldsymbol{\psi}$ when $p+r_2>n$. Let $\hat\boldsymbol{\lambda}(\hat\boldsymbol{\psi}_{ \mathrm{\scriptscriptstyle PEL} })$ be the $r\times 1$ vector of Lagrange multiplier defined at $\hat\boldsymbol{\psi}_{ \mathrm{\scriptscriptstyle PEL} }$:

align*[align* omitted — 515 chars of source]

Write $\hat{\boldsymbol{\lambda}}:=\hat\boldsymbol{\lambda}(\hat\boldsymbol{\psi}_{ \mathrm{\scriptscriptstyle PEL} })=(\hat{\lambda}_1,\ldots,\hat{\lambda}_r)^{ \mathrm{\scriptscriptstyle T} }$ and $\rho_2(t;\nu)=\nu^{-1}P_{2,\nu}(t)$. Let $\mathcal{R}_n= \mathcal{I} \cup \{ j\in \mathcal{D}: \hat{\lambda}_j \neq 0 \} $ be the set of the estimated binding moments. Since only those $\lambda_j$'s associated with $\mathcal{D}$ are penalized, it follows that \[ \frac{1}{n}\sum_{i=1}^n\frac{g_{i,j}^{{ \mathcal{\scriptscriptstyle (T)} }}(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})}{1+\hat{\boldsymbol{\lambda}}_{{ \mathcal{\scriptscriptstyle R} }_n}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}_{i,{ \mathcal{\scriptscriptstyle R} }_n}^{{ \mathcal{\scriptscriptstyle (T)} }}(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})}=\left\{

aligned0\,, &if j\in \mathcal{I} \,,\\ \nu\rho_2'(|\hat{\lambda}_j|;\nu)\rm sgn(\hat\lambda_j) \,, &if j\in\mathcal{D} with \hat{\lambda}_j\neq 0\,.

\right. \] For any unbinding moment $j\in\mathcal{R}_n^{{\mathrm{c}}}$, define

equation[equation omitted — 445 chars of source]

If $\hat{\lambda}_j=0$ for some $j\in\mathcal{D}$, the subdifferential at $\hat{\lambda}_j$ has to include the zero element bertsekas1997nonlinear. That is, $\hat{\eta}_j\in[-\nu\rho_2'(0^+),\nu\rho_2'(0^+)]$ for any $j\in\mathcal{R}_n^{{\mathrm{c}}}$. In our theoretical analysis, we impose the next condition.

as$ \mathbb{P}[\cup_{j\in\mathcal{R}_n^{{\mathrm{c}}}}\{|\hat{\eta}_j|=\nu\rho_2'(0^+)\}]\rightarrow0 $ as $n\rightarrow\infty$.
remarkCondition (ref) requires that $\hat{\eta}_j$ $(j\in\mathcal{R}_n^{{\mathrm{c}}})$ does not lie on the boundary with probability approaching one, which is realistic in practice. If the distribution function of $\hat{\eta}_j$ is continuous at $\pm\nu\rho_2'(0^+)$, we then have $\mathbb{P}\{|\hat{\eta}_j|=\nu\rho_2'(0^+)\}=0$. Condition (ref) makes sure that $\hat{\boldsymbol{\lambda}}(\boldsymbol{\psi})$ is continuously differentiable at $\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ with probability approaching one. See Lemma A.5 in the supplementary materials.

Define $\mathcal{A}_*=\{ j \in \mathcal{A}:\hat{\lambda}_j\neq 0\}$ and $\mathcal{A}_{*,{\mathrm{c}}}=\{ j\in \mathcal{A}^{{\mathrm{c}}}:\hat{\lambda}_j\neq 0\}$. Let $\mathcal{I}^*=\mathcal{I} \cup\mathcal{A}_*$, and then $\mathcal{R}_n$ can be decomposed into two disjoint parts $\mathcal{I}^*$ and $\mathcal{A}_{*,{\mathrm{c}}}$. Furthermore, we define $\mathcal{S}_*= \mathcal{P} \cup \mathcal{A}_{*,{\mathrm{c}}}$. The fact $\mathcal{S}_*\subset\mathcal{S}$ implies that $|\mathcal{S}_*|=p+|\mathcal{A}_{*,{\mathrm{c}}}|\le |\mathcal{S}|=s$. For any $\boldsymbol{\psi}\in\boldsymbol{\Psi}$, we have $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }_*}=(\boldsymbol{\theta}^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} }_{{{ \mathcal{\scriptscriptstyle A} }}_{*,{\mathrm{c}}}})^{ \mathrm{\scriptscriptstyle T} }$. Define

align*[align* omitted — 441 chars of source]

with ${\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}(\boldsymbol{\theta}_0)=\mathbb{E}\{{\mathbf g}^{ \mathcal{\scriptscriptstyle (T)} }_{i,{ \mathcal{\scriptscriptstyle I} }^*}(\boldsymbol{\theta}_0)^{\otimes2}\}$, and $\hat\boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n}=\{{\mathbf J}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle R} }_n}\}^{-1}[\mathbb{E}\{\nabla_{\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }_*}}{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_{i,{ \mathcal{\scriptscriptstyle R} }_n}(\boldsymbol{\psi}_0)\}]^{ \mathrm{\scriptscriptstyle T} }\{{\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle R} }_n}(\boldsymbol{\psi}_0)\}^{-1}\hat\boldsymbol{\eta}_{{ \mathcal{\scriptscriptstyle R} }_n}$, where ${\mathbf J}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle R} }_n}=([\mathbb{E}\{\nabla_{\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }_*}}{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_{i,{ \mathcal{\scriptscriptstyle R} }_n}(\boldsymbol{\psi}_0)\}]^{ \mathrm{\scriptscriptstyle T} }\{{\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle R} }_n}(\boldsymbol{\psi}_0)\}^{-1/2})^{\otimes2}$, $\hat\boldsymbol{\eta}=(\hat{\eta}_j)_{j\in \mathcal{T}}$ with $\hat\eta_j=0$ for $j\in \mathcal{I}$, $\hat\eta_j=\nu\rho'_2(|\hat\lambda_j|;\nu){\rm sgn}(\hat\lambda_j)$ for $j\in \mathcal{D}$ with $\hat\lambda_j\neq0$, and $\hat\eta_j$ specified in (ref) for $j\in\mathcal{R}_n^{{\mathrm{c}}}$. The limiting distribution of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ is stated in Theorem (ref).

theoremAssume the conditions of Proposition {\rm(ref)} hold. Under Conditions {\rm (ref)} and {\rm(ref)}, if $\ell_n^{3/2}\log r=o(n^{1/2-1/\gamma})$ and $\ell_nn^{1/2}s^{3/2} \phi_n \nu=o(1)$, then $ n^{1/2}\boldsymbol{\alpha}^{ \mathrm{\scriptscriptstyle T} }\{{\mathbf J}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}\}^{1/2}\{\hat\boldsymbol{\theta}_{{{ \mathrm{\scriptscriptstyle PEL} }}}-\boldsymbol{\theta}_{0} -\hat\boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n,(1)}\}\xrightarrow{d}\mathcal{N}(0,1)$ as $n\to\infty$ for any $\boldsymbol{\alpha}\in\mathbb{R}^p$ with $|\boldsymbol{\alpha}|_2=1$, where $\hat\boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n,(1)}$ is the first $p$ components of $\hat\boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n}$ and $(a_n,\phi_n)$ are given in (ref).
remarkNotice that $a_n\lesssim s\pi$. To satisfy the restrictions in Theorem (ref), it is sufficient for $(n,r,p,\ell_n,s)$ to have: $r = o\{\exp(n^{1/3})\}$, $\ell_n=o\{n^{(\gamma-2)/(3\gamma)}(\log r)^{-2/3}\}$, $\ell_n^{5/2}s^{3/2}(\log r)\max\{p,\ell_n^{1/2}\}=o(n^{1/2})$, the tuning parameters $\nu$ and $\pi$ satisfying $ \ell_n^{1/2}\aleph_n \ll\pi\ll \min\{s^{-1}n^{-2/\gamma},\ell_n^{-4}s^{-4}p^{-2}(\log r)^{-1}\}$, $ \ell_n\aleph_n \ll \nu \ll \min\{(\ell_n^3s^3p^2\log r)^{-1/2},n^{-1/\gamma},(\ell_n^2ns^3)^{-1/4} \}$ and $\nu^2\pi\ll(\ell_n^2ns^4p^2)^{-1}$.

Thanks to the extra valid moments in ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}$, the asymptotic covariance of $\hat\boldsymbol{\theta}_{{ \mathrm{\scriptscriptstyle PEL} }}$ is smaller than that of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{ \mathcal{\scriptscriptstyle (I)} }$ which utilizes the information in ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}$ only. On the other hand, given that $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }}$ provides asymptotic normality under the set of correctly specified moments $\mathcal{I}$, one may question whether it is worthwhile to consider all the available moments among which some may risk misspecification. This is a trade-off between robustness and efficiency, a recurrent theme in modern econometrics. Although it inevitably depends on the data generating process (DGP), in big data environments the efficiency gain may be substantial. Benefits are demonstrated in the simulation exercises and the empirical application via very simple econometric models.\footnote{ See the RMSEs in Table (ref), the length of confidence intervals in Figure (ref), and the standard deviations in Table (ref). }

remarkThe setting with the sets $\mathcal{I}$ and $\mathcal{D}= \mathcal{A}\cup \mathcal{A}^{{\mathrm{c}}}$ is the same as liao2013adaptive. When estimating ((ref)) with low-dimensional moments based on GMM, ChengLiao2015 require the number of estimating functions $r=o(n^{1/3})$, and the signal of invalid estimating functions $\min_{k\in\mathcal{A}^{{\mathrm{c}}}}|\mathbb{E}\{g_{i,k}^{{ \mathcal{\scriptscriptstyle (D)} }}(\boldsymbol{\theta}_0)\}|\gg (rn^{-1})^{1/2}$. Our conditions on the relative size of the dimensions are more general. A sufficient condition for (ref) is $\min_{k\in\mathcal{A}^{{\mathrm{c}}}}|\mathbb{E}\{g_{i,k}^{{ \mathcal{\scriptscriptstyle (D)} }}(\boldsymbol{\theta}_0)\}|\gg \max\{\nu,p(s\pi)^{1/2},pr_1^{1/2}\aleph_n\}$, in view of $a_n\lesssim s\pi$. Under the restrictions discussed in Remark (ref), if we select $\nu$ and $\pi$ sufficiently close to $\ell_n\aleph_n$ and $ \ell_n^{1/2}\aleph_n$, respectively, then (ref) holds provided that $\min_{k\in\mathcal{A}^{{\mathrm{c}}}}|\mathbb{E}\{g_{i,k}^{{ \mathcal{\scriptscriptstyle (D)} }}(\boldsymbol{\theta}_0)\}|\gg ps^{1/2}\ell_n^{1/4}\aleph_n^{1/2}$. This is similar to the “beta-min condition” which is necessary for consistent variable selection by shrinkage methods buhlmann2013statistical.
remarkUnder low-dimensional moments the asymptotic normality involves no bias; see liao2013adaptive and ChengLiao2015. In contrast, Theorem (ref) here makes clear that the high-dimensional moments incur an additional asymptotic bias $\hat \boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n,(1)}$. To conduct hypothesis testing or construct confidence region about $\boldsymbol{\theta}$, in principle we can estimate and correct the bias $\hat{\boldsymbol{\zeta}}_{{ \mathcal{\scriptscriptstyle R} }_n,(1)}$. Such a direct bias correction approach, nevertheless, is undesirable in practice and in theory. The bias term involves multiplication and inverse of large matrices, which are difficult to estimate with precision in finite samples. Illustrated in our simulation experiments in Section (ref), we are unsatisfied with the asymptotic normality approximation after estimating and correcting the bias. Thus, we view Theorem (ref) as a characterization of the asymptotic behavior of $\hat\boldsymbol{\theta}_{{{ \mathrm{\scriptscriptstyle PEL} }}}$, but we do not encourage carrying out statistical inference based on it. Instead, Section (ref) recommends PPEL for inference.

While the augmented parameter $\boldsymbol{\xi}$ essentially validates all moments in $\mathcal{D}$, the efficiency gain comes from the penalty on $\boldsymbol{\xi}$ that shrinks some $\hat{\xi}_{k}$, $k \in \mathcal{A}$, to zero, thereby confirming the validity of these associated moments. To discuss moment selection, we write $\hat\boldsymbol{\xi}_{{ \mathrm{\scriptscriptstyle PEL} }}=(\hat{\xi}_k)_{k\in \mathcal{D}}$. If all valid estimating functions in $\mathcal{A}$ are selected by the optimization (ref), i.e., $\hat{\xi}_k = 0$ for all $k\in \mathcal{A}$, then the asymptotic covariance of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ is $\{{\mathbf J}^{{ \mathcal{\scriptscriptstyle (H)} }}\}^{-1}$.

Remind that $\{{\mathbf J}^{{ \mathcal{\scriptscriptstyle (H)} }}\}^{-1}$ in Proposition (ref) is the semiparametric efficiency bound for the estimation of $\boldsymbol{\theta}_0$ under the oracle, and Proposition (ref) shows that $\mathbb{P}(\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle A} }}={\mathbf 0})\rightarrow1$ as $n\rightarrow\infty$, which provides a natural moment selection criterion

align[align omitted — 77 chars of source]

to identify the valid estimating functions in $\mathcal{A}$. Notice that Proposition (ref) also indicates that $|\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle A} }^{{\mathrm{c}}}}-\boldsymbol{\xi}_{0,{ \mathcal{\scriptscriptstyle A} }^{{\mathrm{c}}}}|_\infty=O_{ \mathrm{p} }(\phi_n)$. Together with (ref), we have $\mathbb{P}[\cup_{k\in\mathcal{A}^{{\mathrm{c}}}}\{\hat{\xi}_k=0\}]\rightarrow0$ as $n\rightarrow\infty$. Based on these arguments, Theorem (ref) supports our proposed moments selection criterion.

theoremUnder the conditions of Proposition {\rm(ref)}, it holds that $\mathbb{P}(\mathcal{\hat A}=\mathcal{A})\to1$ as $n\rightarrow\infty$.

Up to now, we have established the asymptotic properties of the PEL estimator for the moment-defined model ((ref)) with high-dimensional moments and low-dimensional $\boldsymbol{\theta}$. The next section further extends the estimation to a high-dimensional structural parameter.

High-dimensional moments with high-dimensional $\boldsymbol{\theta}$

A high-dimensional parameter $\boldsymbol{\theta}$ is present when many control variables are included in the structural model. For example, the leading case of the linear IV model takes only one endogenous variable and it is accompanied by a few corresponding IVs. Nevertheless, such a simple setting may still involve many exogenous control variables in the main equation, and these control variables are natural instruments for themselves. An empirical example can be found in blundell1993we. fan2014endogeneity propose the focused GMM for such a linear IV model, which is a special case of the moment-defined model.

In this section, $p$ and $r_2$ are both allowed to be much larger than $n$, whereas $r_1$ is fixed or diverges at some slow polynomial rate of $n$. We assume $\boldsymbol{\theta}_0$ sparse in the sense that most of its components are zero. Write $\boldsymbol{\theta}_0=(\theta_{0,l})_{l \in \mathcal{P}} $ and define the active set $\mathcal{P}_{\sharp}=\{l \in \mathcal{P}:\theta_{0,l}\neq0\}$ with cardinality $p_{\sharp}=|\mathcal{P}_{\sharp}|$. To obtain a sparse estimate of $\boldsymbol{\theta}_0$, we add to (ref) a penalty on $\boldsymbol{\theta}$ to produce the following optimization problem:

equation[equation omitted — 794 chars of source]

With slight abuse of notation, we keep using $(\hat{\boldsymbol{\theta}}_{{{ \mathrm{\scriptscriptstyle PEL} }}}^{ \mathrm{\scriptscriptstyle T} },\hat{\boldsymbol{\xi}}_{{{ \mathrm{\scriptscriptstyle PEL} }}}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ to denote the solution to (ref).

As $p$ can be bigger than $r_1$ in high dimension, here we update Condition (ref) for the identification of $\boldsymbol{\theta}_0$. In view of ChangTangWu2018, we impose Condition (ref) below for the identification of the nonzero components of $\boldsymbol{\theta}_0$.

\addtocounter{as}{-6}

as(i) There exists a universal constant $K'_1>0$ such that $$\inf_{\boldsymbol{\theta}=(\boldsymbol{\theta}_{{ \mathcal{\scriptscriptstyle P}}_\sharp}^{{ \mathrm{\scriptscriptstyle T} }},\boldsymbol{\theta}_{{ \mathcal{\scriptscriptstyle P}}_\sharp^{{\mathrm{c}}}}^{{ \mathrm{\scriptscriptstyle T} }})^{{ \mathrm{\scriptscriptstyle T} }}\in\boldsymbol{\Theta}:\,|\boldsymbol{\theta}_{{ \mathcal{\scriptscriptstyle P}}_\sharp}-\boldsymbol{\theta}_{0,{ \mathcal{\scriptscriptstyle P}}_\sharp}|_\infty>\varepsilon, \,\boldsymbol{\theta}_{{ \mathcal{\scriptscriptstyle P}}_\sharp^{{\mathrm{c}}}}={\mathbf 0}}|\mathbb{E}\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}_i(\boldsymbol{\theta})\}|_\infty\ge K'_1 \varepsilon$$ for any $\varepsilon>0$. (ii) There exists a universal constant $K_3'>0$ such that $\sup_{\boldsymbol{\theta}\in\boldsymbol{\Theta}}|\mathbb{E}\{\nabla_{\boldsymbol{\theta}_{{ \mathcal{\scriptscriptstyle P}}_\sharp^{{\mathrm{c}}}}}{\mathbf g}_i^{{ \mathcal{\scriptscriptstyle (I)} }}(\boldsymbol{\theta})\}|_\infty\leq K_3'$.

As we have shown in Section (ref), the index set $\mathcal{S}$ and the quantities $a_n$ and $\phi_n$ play key roles in the theoretical analysis of the low-dimensional $\boldsymbol{\theta}$ in (ref). Under the current high-dimensional setting, we define

align[align omitted — 237 chars of source]

This $\mathcal{S}$ indexes all nonzero components in the true augmented parameter $\boldsymbol{\psi}_0$. Compared to its counterpart $a_n$ in Section (ref), the newly defined $a_n$ here is amended with an extra term $\sum_{l\in\mathcal{P}}P_{1,\pi}(|\theta_{0,l}|)$ due to the penalty imposed on $\boldsymbol{\theta}$. In the current high-dimensional setting, we also redefine

align[align omitted — 98 chars of source]

with $a_n$ in (ref). To control the bias introduced by $P_{1,\pi}(\cdot)$ on $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ and $\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$, similar to (ref), we assume that among the elements in $\boldsymbol{\psi}_0=(\psi_{0,1},\ldots,\psi_{0,p+r_2})^{ \mathrm{\scriptscriptstyle T} }$, the minimal active signal is

align[align omitted — 86 chars of source]

with $\phi_n$ in (ref) and $\max_{{k\in \mathcal{S}}}\sup_{c|\psi_{0,k}|<t<c^{-1}|\psi_{0,k}|}P'_{1,\pi}(t)=0$ for some constant $c\in(0,1)$.\footnote{See the arguments below (ref) for the validity of these assumptions.} Given the newly defined $\mathcal{S}$ and $\phi_n$, we can further update $\ell_n$, $\mathcal{R}_n$, $\mathcal{A}_*$, $\mathcal{A}_{*,{\mathrm{c}}}$ and $\mathcal{I}^*$ in the same manner as their counterparts in Section (ref). Moreover, we also define ${\mathbf J}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle R} }_n}$ and $\hat\boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n}$ as specified in Section (ref) with $\mathcal{S}_*=\mathcal{P}_{\sharp}\cup \mathcal{A}_{*,{\mathrm{c}}}$. For any $\boldsymbol{\psi}\in\boldsymbol{\Psi}$, we have $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }_*}=(\boldsymbol{\theta}_{{ \mathcal{\scriptscriptstyle P}}_{\sharp}}^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} }_{{{ \mathcal{\scriptscriptstyle A} }}_{*,{\mathrm{c}}}})^{ \mathrm{\scriptscriptstyle T} }$.

Proposition A.2 in the supplementary materials shows that there exists a sparse local minimizer $(\hat\boldsymbol{\theta}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} },\hat\boldsymbol{\xi}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }\in\boldsymbol{\Psi}$ for the nonconvex optimization (ref) such that $\mathbb{P}(\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle P}}^{{\mathrm{c}}}_{\sharp}}={\mathbf 0})\to1$ as $n\to\infty$, which means all zero components of $\boldsymbol{\theta}_0$ can be estimated exactly as zero with high probability. The limiting distribution of such $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle P}}_{\sharp}}$ and the moment selection outcomes are stated in Theorem (ref) below.

theoremLet $P_{1,\pi}(\cdot), P_{2,\nu}(\cdot)\in\mathscr{P}$ for $\mathscr{P}$ defined in (ref), and $P_{2,\nu}(\cdot)$ be convex with bounded second derivative around $0$. For the sparse local minimizer $\hat\boldsymbol{\psi}_{{{ \mathrm{\scriptscriptstyle PEL} }}}=(\hat{\boldsymbol{\theta}}_{{{ \mathrm{\scriptscriptstyle PEL} }}}^{ \mathrm{\scriptscriptstyle T} },\hat{\boldsymbol{\xi}}_{{{ \mathrm{\scriptscriptstyle PEL} }}}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ for {\rm(ref)} specified in Proposition {\rm A.2} in the supplementary materials, assume there exists a constant $\tilde{c}\in(C_*,1)$ such that $\mathbb{P}[\cup_{j\in\mathcal{T}}\{|\bar{g}_j^{ \mathcal{\scriptscriptstyle (T)} }(\hat\boldsymbol{\psi}_{{{ \mathrm{\scriptscriptstyle PEL} }}})|\in[\tilde{c}\nu\rho'_2(0^+),\nu\rho'_2(0^+))\}]\rightarrow0$. Suppose Conditions {\rm(ref)}, {\rm (ref)--(ref)} and (ref) hold. Furthermore, assume Condition {\rm(ref)} holds with the newly defined $\ell_n$, replacing $\mathcal{P}$ by $\mathcal{P}_{\sharp}$ and replacing $p$ by $p_{\sharp}$, and Condition {\rm(ref)} holds with the newly defined $\mathcal{R}_n$. If $\log r=o(n^{1/3})$, $\max\{a_n,\nu^2\}=o(n^{-2/\gamma})$, $\ell_n^{3/2}\log r=o(n^{1/2-1/\gamma})$, $\ell_nn^{1/2}s^{3/2} \phi_n \nu=o(1)$ and $\ell_n\aleph_n=o(\min\{\nu,\pi\})$, then we have $$ n^{1/2}\boldsymbol{\alpha}^{ \mathrm{\scriptscriptstyle T} }\{{\mathbf W}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}\}^{1/2}\{\hat\boldsymbol{\theta}_{{{ \mathrm{\scriptscriptstyle PEL} }},{ \mathcal{\scriptscriptstyle P}}_{\sharp}}-\boldsymbol{\theta}_{0,{ \mathcal{\scriptscriptstyle P}}_{\sharp}}-\hat\boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n,(1)}\}\xrightarrow{d}\mathcal{N}(0,1) $$ for any $\boldsymbol{\alpha}\in\mathbb{R}^{p_\sharp}$ with $|\boldsymbol{\alpha}|_2=1$ as $n\to\infty$, where $\hat\boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n,(1)}$ is the first $p_{\sharp}$ components of $\hat\boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n}$, $(a_n,\phi_n)$ are given by (ref) and (ref), respectively, and $ {\mathbf W}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}=([\mathbb{E}\{\nabla_{\boldsymbol{\theta}_{{ \mathcal{\scriptscriptstyle P}}_{\sharp}}}{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_{i,{ \mathcal{\scriptscriptstyle I} }^*}(\boldsymbol{\theta}_0)\}]^{ \mathrm{\scriptscriptstyle T} }\{{\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}(\boldsymbol{\theta}_0)\}^{-1/2})^{\otimes2}$ with ${\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}(\boldsymbol{\theta}_0)=\mathbb{E}\{{\mathbf g}^{ \mathcal{\scriptscriptstyle (T)} }_{i,{ \mathcal{\scriptscriptstyle I} }^*}(\boldsymbol{\theta}_0)^{\otimes2}\}$. Moreover, under the conditions of Proposition {\rm A.2} in the supplementary materials, it holds that $\mathbb{P}(\mathcal{\hat A}=\mathcal{A})\to1$ as $n\rightarrow\infty$, where $\hat{\mathcal{A}}$ is specified in (ref).
remarkInstead of assuming $ \ell_n^{1/2} \aleph_n =o(\pi)$ as in Theorem (ref), here Theorem (ref) strengthens $\ell_n \aleph_n=o(\pi)$. As shown in Section A.5 of the supplementary materials, this stronger condition guarantees that $\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}$, the zero components of $\boldsymbol{\psi}_0$, can be shrunk to zero with probability approaching one in the high-dimensional $\boldsymbol{\theta}$ setting.

Compared to (ref), a penalty on $\boldsymbol{\theta}$ is added to (ref). To appreciate the benefit from this additional penalty on the sparse parameter, without loss of generality, we write $\boldsymbol{\theta}_{0}=(\boldsymbol{\theta}_{0,{ \mathcal{\scriptscriptstyle P}}_{\sharp}}^{ \mathrm{\scriptscriptstyle T} },{\mathbf 0}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ and the inverse of ${\mathbf J}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}$ in Theorem (ref) in the following compatible partitioned matrix

align*[align* omitted — 600 chars of source]

where $[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{11}$ is a $p_{\sharp}\times p_{\sharp}$ matrix. The asymptotic covariance of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle P}}_{\sharp}}$ by (ref) is $[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{11}$, and that in (ref) is $\{{\mathbf W}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}\}^{-1}$ according to Theorem (ref). Notice that $\{{\mathbf W}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}\}^{-1}=[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{11} -[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{12}[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{22}^{-1}[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{21} \le [\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{11}$ by the inverse of a partitioned matrix. Implicitly restricting the parameter space, the penalty on $\boldsymbol{\theta}$ gains efficiency for the estimation of $\boldsymbol{\theta}_{0,{ \mathcal{\scriptscriptstyle P}}_{\sharp}}$, the nonzero components of $\boldsymbol{\theta}_0$.

remarkTheorem (ref) spells out the limiting distribution and efficiency gain for the estimator of the nonzero components in $\boldsymbol{\theta}_0$. If one is interested in some coefficient in $\mathcal{P}_{\sharp}^{{\mathrm{c}}}$, we have consistency in that $\mathbb{P}(\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle P}}^{{\mathrm{c}}}_{\sharp}}={\mathbf 0})\to1$ as $n\to\infty$, but the limiting distribution is irregular due to the shrinkage.\footnote{This is a generic property shared by procedures of the oracle properties, for example SCAD FanLi2001 and the adaptive Lasso zou2006adaptive. } Furthermore, parallel to Theorem (ref) the asymptotic bias is again present. Similar to Remark (ref), we do not suggest conducting statistical inference predicated on this characterization of the asymptotic behavior of normality.

Statistical inference is important in applied econometrics when researchers intend to assess whether the estimated result supports or rejects a hypothesized value of the parameter. The next section proposes an inferential procedure based on a projection of estimating functions, which is free of asymptotic biases.

Confidence regions for a subset of $\boldsymbol{\psi}$

Suppose we are interested in the inference for a subset of parameters $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$ for some generic small subset $\mathcal{M} \subset (\mathcal{P} \cup \mathcal{D}) $ with $|\mathcal{M}|=m$. Here we allow $m$ to be fixed or diverge slowly with the sample size $n$. What is novel here is that $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$ is allowed to contain part of the auxiliary parameter $\boldsymbol{\xi}$, for which our method will provide a formal statistical inference for the validity of a subset of the high-dimensional moment restrictions. In contrast, liao2013adaptive offers asymptotic normality for the structural parameter $\boldsymbol{\theta}$ under low-dimensional moments but does not characterize the asymptotic distribution of the auxiliary parameter $\boldsymbol{\xi}$.

A key technical issue is how to deal with the high-dimensional nuisance parameter $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{\mathrm{c}}}$, where $\mathcal{M}^{{\mathrm{c}}}= (\mathcal{P} \cup \mathcal{D}) \backslash \mathcal{M}$. Our approach follows Chang2020 to project out the influence of the nuisance parameter. Such idea was also used in ning2017general and neykov2018unified. Given an initial estimate $\boldsymbol{\psi}^*$ for $\boldsymbol{\psi}_0$, we first determine a linear transformation matrix ${\mathbf A}_n = ({\mathbf a}_k^n)_{k\in\mathcal{M}}^{{ \mathrm{\scriptscriptstyle T} }} \in \mathbb{R}^{m\times r}$ with each row defined as

align[align omitted — 323 chars of source]

for $k \in \mathcal{M}$, where $\varsigma\rightarrow0$ as $n\rightarrow\infty$ is a tuning parameter, and $\{\boldsymbol{\gamma}_k\}_{k\in \mathcal{M}}$ is a basis of the linear space $\{ {\mathbf b}=(b_j)_{j\in \mathcal{P}\cup\mathcal{D} }: {\mathbf b}_{\mathcal{M}^{{\mathrm{c}}}} = \mathbf{0} \}$. In practice, we can specify the $(p+r_2)$-dimensional vector $\boldsymbol{\gamma}_k = ( 1(j=k))_{j \in \mathcal{P} \cup \mathcal{D} }$ with all zero elements except one unit entry. As we will discuss in Remark (ref), the solution from (ref) or (ref) can serve as the initial estimator $\boldsymbol{\psi}^*$.

Based on ${\mathbf A}_n$ as in ((ref)), we then obtain the new $m$-dimensional estimating functions ${\mathbf f}^{{\mathbf A}_n}(\cdot;\cdot)$ by projecting ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}(\cdot;\cdot)$ on ${\mathbf A}_n$:

align*[align* omitted — 175 chars of source]

Write $\boldsymbol{\gamma}_k=(\gamma_{k,j})_{j\in \mathcal{P}\cup\mathcal{D}}$ and define an $m\times(p+r_2)$ matrix $\boldsymbol{\Gamma}=(\gamma_{k,j})_{k\in { \mathcal{\scriptscriptstyle M} }, \, j \in \mathcal{P} \cup \mathcal{D}}$. The definition of ${\mathbf A}_n$ implies that $|\nabla_{\boldsymbol{\psi}}\bar{{\mathbf f}}^{{\mathbf A}_n}(\boldsymbol{\psi}^*)-\boldsymbol{\Gamma}|_\infty\leq \varsigma$. Since all the components in the $j$-th column of $\boldsymbol{\Gamma}$ are zero for $j\in\mathcal{M}^{{\mathrm{c}}}$, the newly defined estimating functions ${\mathbf f}^{{\mathbf A}_n}$ is uninformative of the nuisance parameter $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}$. Informative is ${\mathbf f}^{{\mathbf A}_n}$ of the parameter $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$ of interest, as for $j\in\mathcal{M}$ some components in the $j$-th column of $\boldsymbol{\Gamma}$ must be nonzero.

Given the projected estimating functions ${\mathbf f}^{{\mathbf A}_n}$, it is possible to directly borrow from Chang2020 to construct the confidence region of $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$ by the EL ratio \[ w_n(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }})=2\max_{\boldsymbol{\lambda}\in\tilde{\Lambda}_n(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }})}\sum_{i=1}^n\log\{1+\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf f}_i^{{\mathbf A}_n}(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }},\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^*)\} \] with respect to $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$, where $\tilde{\Lambda}_n(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }})=\{\boldsymbol{\lambda} \in\mathbb{R}^m : \boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} } {\mathbf f}^{{\mathbf A}_n}_i(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }},\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^*) \in \mathcal{V}~\textrm{for any}~i\in[n]\}$. Since $w_n(\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle M} }}) \xrightarrow{d}\chi_m^2$ as $n\rightarrow\infty$ for a fixed $m$, the set $\{\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}\in\mathbb{R}^m:w_n(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }})\leq \chi_{m,1-\alpha}^2\}$ provides a $100(1-\alpha)\%$ confidence region for $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$, where $\chi_{m,1-\alpha}^2$ is the $(1-\alpha)$-quantile of $\chi_m^2$ distribution. Nonetheless, the finite-sample performance of such an asymptotically valid confidence region depends crucially on the convexity of $w_n(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }})$ near $\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle M} }}$. If convexity fails in a finite sample, this EL-ratio-based confidence region will be voluminous in magnitude due to numerical instability. Verifying the convexity condition is onerous, in particular when $m$ is large, for ${\mathbf f}_i^{{\mathbf A}_n}(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }},\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^*)$ may be a nonlinear function of $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$.

To secure a stable confidence region for $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$, in this paper we deviate from Chang2020 and instead recommend a re-estimation procedure for a confidence region based on the asymptotic normality of the estimator. Let

align*[align* omitted — 541 chars of source]

where $\boldsymbol{\Psi}^*_{ \mathcal{\scriptscriptstyle M} } = \{\boldsymbol{\psi}_{ \mathcal{\scriptscriptstyle M} }: |\boldsymbol{\psi}_{ \mathcal{\scriptscriptstyle M} } - \boldsymbol{\psi}^*_{ \mathcal{\scriptscriptstyle M} } |_1 \le O_{ \mathrm{p} }(\varpi_{1,n}) \}$ for some $\varpi_{1,n}\rightarrow0$ such that $|\boldsymbol{\psi}^*_{ \mathcal{\scriptscriptstyle M} }-\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle M} }}|_1 = O_{ \mathrm{p} }(\varpi_{1,n})$. Since $\boldsymbol{\psi}_{ \mathcal{\scriptscriptstyle M} }^*$ is an initial consistent estimator of $\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle M} }}$, it is reasonable to search for $\tilde{\boldsymbol{\psi}}_{{ \mathcal{\scriptscriptstyle M} }}$ in a small neighborhood of $\boldsymbol{\psi}^*_{ \mathcal{\scriptscriptstyle M} }$. We will specify $\varpi_{1,n}$ for some specific choices of $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}^*$ in Remark (ref).

To derive the limiting distribution of $\tilde\boldsymbol{\psi}_{ \mathcal{\scriptscriptstyle M} }$, we assume the following condition.

\addtocounter{as}{5}

asFor each $k\in \mathcal{M}$, there is a non-random ${\mathbf a}_k$ satisfying $[\mathbb{E}\{\nabla_{\boldsymbol{\psi}}{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\boldsymbol{\psi}_0)\}]^{ \mathrm{\scriptscriptstyle T} } {\mathbf a}_k = \boldsymbol{\gamma}_k$, $| {\mathbf a}_k |_1\le K_6$ for some universal constant $K_6 > 0$, and $ \max_{k\in { \mathcal{\scriptscriptstyle M} } }|{\mathbf a}_k^n - {\mathbf a}_k|_1 = O_{ \mathrm{p} }(\omega_n)$ for some $\omega_n \to 0$. Let ${\mathbf A}=({\mathbf a}_k)_{k\in\mathcal{M}}^{{ \mathrm{\scriptscriptstyle T} }} \in \mathbb{R}^{m\times r}$. The eigenvalues of ${\mathbf A}^{\otimes2}$ are uniformly bounded away from zero and infinity.
remarkLet $\boldsymbol{\Xi}=\mathbb{E}\{\nabla_{\boldsymbol{\psi}}{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\boldsymbol{\psi}_0)\}$ and $\widehat{\boldsymbol{\Xi}}=\nabla_{\boldsymbol{\psi}}\bar{{\mathbf g}}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi}^*)$. It follows from the existence of ${\mathbf a}_k$ that $\boldsymbol{\gamma}_k=\widehat{\boldsymbol{\Xi}}^{ \mathrm{\scriptscriptstyle T} }{\mathbf a}_k+(\boldsymbol{\Xi}-\widehat{\boldsymbol{\Xi}})^{ \mathrm{\scriptscriptstyle T} }{\mathbf a}_k=\widehat{\boldsymbol{\Xi}}^{ \mathrm{\scriptscriptstyle T} }{\mathbf a}_k+\boldsymbol{\varepsilon}_k$, where $\boldsymbol{\varepsilon}_k = (\boldsymbol{\Xi}-\widehat{\boldsymbol{\Xi}})^{ \mathrm{\scriptscriptstyle T} }{\mathbf a}_k$. Some mild conditions ensure $|\widehat{\boldsymbol{\Xi}}-\boldsymbol{\Xi}|_\infty=o_{ \mathrm{p} }(1)$. This, together with the assumption $|{\mathbf a}_k|_1\leq K_6$, implies that $\boldsymbol{\varepsilon}_k$ is stochastically small uniformly over all the components such that $|\boldsymbol{\varepsilon}_k|_\infty=o_{ \mathrm{p} }(1)$, which can be viewed as an attempt to recover a nonrandom ${\mathbf a}_k$ with no noise asymptotically (CandesTao2007, Bickeletal2009). It follows that $|{\mathbf a}_k^n-{\mathbf a}_k|_1=o_{ \mathrm{p} }(1)$ if $\widehat{\boldsymbol{\Xi}}$ satisfies the routine conditions for sparse signal recovering. Furthermore, the constant $K_6$ in Condition (ref) may be replaced by some diverging $\varphi_n$ and our main results remain valid.

Theorem (ref) gives the limiting distribution of $\tilde{\boldsymbol{\psi}}_{{ \mathcal{\scriptscriptstyle M} }}$ with a generic initial estimator $\boldsymbol{\psi}^*$.

theoremLet $|\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}^* - \boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle M} }}|_1 = O_{ \mathrm{p} }(\varpi_{1,n})$, $|\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^* - \boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}|_1 = O_{ \mathrm{p} }(\varpi_{2,n})$ for some $\varpi_{1,n}\to0$ and $\varpi_{2,n}\to0$. Under Conditions {\rm(ref)}--{\rm(ref)} and {\rm(ref)}, if $m=o(n^{\min\{1/5,(\gamma-2)/(3\gamma)\}})$, $m\omega_n^2(m^2+\log r)=o(1)$, $m\varpi_{1,n}=o(1)$ and $n m \varpi_{2,n}^2 (\varsigma^2 + \varpi_{1,n}^2 + \varpi_{2,n}^2) = o(1)$, then $ n^{1/2}\boldsymbol{\alpha}^{ \mathrm{\scriptscriptstyle T} }(\hat{\mathbf J}^*)^{1/2}(\tilde{\boldsymbol{\psi}}_{{ \mathcal{\scriptscriptstyle M} }} - \boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle M} }})\xrightarrow{d} \mathcal{N}(0,1)$ as $n\to\infty$ for any $\boldsymbol{\alpha} \in \mathbb{R}^m$ with $| \boldsymbol{\alpha} |_2 = 1$, where $ \hat{\mathbf J}^* = [ \{ \nabla_{\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}} \bar{{\mathbf f}}^{{\mathbf A}_n}(\tilde\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }},\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^*) \}^{ \mathrm{\scriptscriptstyle T} } \{ \widehat{{\mathbf V}}_{{\mathbf f}^{{\mathbf A}_n}}(\tilde\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }},\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^*) \}^{-1/2} ]^{\otimes2} $ with $\widehat{{\mathbf V}}_{{\mathbf f}^{{\mathbf A}_n}}(\tilde\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }},\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^*) = \mathbb{E}_n\{{\mathbf f}^{{\mathbf A}_n}_i(\tilde\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }},\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^*)^{\otimes2}\}$.

The above theorem is stated for $\tilde{\boldsymbol{\psi}}_{{ \mathcal{\scriptscriptstyle M} }} $. It includes the estimation of $\boldsymbol{\theta}_{{ \mathcal{\scriptscriptstyle M} }}$ as a special case if one's research interest falls on the main parameter for economic interpretation, and makes it possible to infer the validity of a subset of moments in view of selecting $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}=\boldsymbol{\xi}_{{ \mathcal{\scriptscriptstyle M} }}$.

remarkWe verify that the PEL estimator is qualified to serve as $\boldsymbol{\psi}^*$ in Theorem (ref). Define $\bar{s}=|\mathcal{S}\cap\mathcal{M}|$, and thus $|\mathcal{S}\cap\mathcal{M}^{{\mathrm{c}}}|=s-\bar{s}$. If $\boldsymbol{\theta}_0$ is low-dimensional and we choose $\boldsymbol{\psi}^*=\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ in (ref), we have $\varpi_{1,n}=\bar{s}\phi_n$ and $\varpi_{2,n}=(s-\bar{s})\phi_n$ due to $|\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle S} }}-\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle S} }}|_\infty=O_{ \mathrm{p} }(\phi_n)$ and $\mathbb{P}(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}={\mathbf 0})\rightarrow1$. Then $ms\phi_n=o(1)$ and $nms^2\phi_n^{2}(\varsigma^2+s^2\phi_n^{2})=o(1)$ are sufficient for the restrictions imposed on $\varpi_{1,n}$ and $\varpi_{2,n}$. Given $s$ and $\phi_n$ in Theorem (ref), if $m$, $\omega_n$ and $\varsigma$ satisfy $m=o(n^{\min\{1/5,(\gamma-2)/(3\gamma)\}})$, $m\omega_n^2(m^2+\log r)=o(1)$, $ms\phi_n=o(1)$ and $nms^2\phi_n^{2}(\varsigma^2+s^2\phi_n^{2})=o(1)$, the asymptotic normality of Theorem (ref) holds under this choice of the initial value $\boldsymbol{\psi}^* = \hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$. Analogously, if $\boldsymbol{\psi}^*=\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ in (ref) when $\boldsymbol{\theta}_0$ is of high dimension, Theorem (ref) also holds provided that $m$, $\omega_n$ and $\varsigma$ satisfy the same restrictions with the newly defined $s$ and $\phi_n$ in Section (ref).

So far, we have established statistical inference results based on the PPEL estimator $\tilde{\boldsymbol{\psi}}_{{ \mathcal{\scriptscriptstyle M} }}$ when we are interested in a small subset $\mathcal{M}$ of $\boldsymbol{\psi}$. In the next section, we check the finite sample performance via simulations.

Numerical studies

One of the most important models in econometrics is the linear IV regression. We design a linear IV model here to mimic the empirical application in Section (ref), whereas simulation results of a nonlinear panel regression with time-varying individual heterogeneity are presented in the supplementary materials.

Suppose that the researcher has at hand a dataset of $n$ independent observations of a vector $\left(y_{i},x_{i},{\mathbf z}_{i},w_{1i},{\mathbf w}_{2i},{\mathbf w}_{3i}\right)$, and is interested in estimating the main structural equation

equation[equation omitted — 150 chars of source]

where $x_{i}$ is a scalar endogenous variable, and ${\mathbf z}_{i}$ is a $d_{z}\times1$ vector of exogenous variables (including the intercept). Such an equation with a scalar endogenous variable is the leading case of IV regressions andrews2019weak. Due to space limitations, we report the results when we specify ${\mathbf z}_i=(1,z_{1,i},z_{2,i})^{ \mathrm{\scriptscriptstyle T} }\in\mathbb{R}^3$ with $( z_{1,i}, z_{2,i} )^{ \mathrm{\scriptscriptstyle T} } \sim \mathcal{N}({\mathbf 0},\rm{{\mathbf I}}_2)$, and the coefficients $ (\beta_x, \boldsymbol{\beta}^{ \mathrm{\scriptscriptstyle T} }_{\mathbf z})^{ \mathrm{\scriptscriptstyle T} }=(0.5,0.5,0.5,0.5)^{ \mathrm{\scriptscriptstyle T} }$ in a plausible setting. The numerical performance are robust across our experiments when the parameters are varied.

The corresponding reduced-form equation is $ x_{i}=\gamma_{w1}w_{1i}+\boldsymbol{\gamma}^{ \mathrm{\scriptscriptstyle T} }_{{\mathbf w}2}{\mathbf w}_{2i}+\boldsymbol{\gamma}_{{\mathbf z}}^{ \mathrm{\scriptscriptstyle T} }{\mathbf z}_{i}+u_{i}$, where $w_{1i}\in\mathbb{R}$ and ${\mathbf w}_{2i}\in\mathbb{R}^{d_{{\mathbf w}2}}$ are excluded instruments. We specify $ (w_{1i}, {\mathbf w}_{2i}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} } \sim \mathcal{N}({\mathbf 0},{\rm\bf I}_{d_{{\mathbf w}2}+1})$, $ \gamma_{w1}=0.8$, $\boldsymbol{\gamma}_{{\mathbf w}2}=(\gamma_{{\mathbf w}2,j})_{j\in[d_{{\mathbf w}2}]}$ with $\gamma_{{\mathbf w}2,j}=0.4-0.3(j-1)/(d_{{\mathbf w}2}-1)$, and $\boldsymbol{\gamma}_{\mathbf z}=(0.8,0.8,0.8)^{ \mathrm{\scriptscriptstyle T} }$. The two error terms $(\epsilon_i,u_i)^{{ \mathrm{\scriptscriptstyle T} }}$ are generated from the two-dimensional normal distribution with mean $0$, covariance $1$ and correlation $0.5$ which are independent of $\left({\mathbf z}_{i},w_{1i},{\mathbf w}_{2i}\right)$. An additional vector ${\mathbf w}_{3i}=(w_{3i,j})_{j\in[s]}\in\mathbb{R}^{s}$, which serves as the invalid IV, follows $ w_{3i,j}=\delta_j\epsilon_{i}+v_{i}$, where $v_{i}\sim \mathcal{N}\left(0,1\right)$ is independent of $\left(\epsilon_{i},u_{i}\right)$, and $\delta_j \neq 0$ controls the strength of correlation. Write $\boldsymbol{\delta}=(\delta_j)_{j\in[s]}$. To emulate the scenarios of weak, moderate and strong correlations between $\epsilon_i$ and ${\mathbf w}_{3i}$, we set $\delta_j=0.3+0.2(j-1)/(s-1)$, $0.5+0.2(j-1)/(s-1)$ and $0.7+0.2(j-1)/(s-1)$, respectively. It is expected that the smaller is the coefficient, the less severe is misspecification, so estimation is more prone to selection mistakes in finite samples.

While the researcher is confident about the validity of $w_{1i}$ in that $ {\mathbf g}_i^{{ \mathcal{\scriptscriptstyle (I)} }} (\boldsymbol{\theta})= (w_{1i}, {\mathbf z}_{i}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} } \times \left(y_{i}-\beta_{x}x_{i}-\boldsymbol{\beta}_{{\mathbf z}}^{ \mathrm{\scriptscriptstyle T} }{\mathbf z}_{i}\right), $ where ${\mathbf z}_i$ is self-instrumented, she is uncertain about the validity of $({\mathbf w}_{2i}, {\mathbf w}_{3i})$ and therefore must detect the invalid IVs. The two classes of undetermined (to the researcher) IVs consist of $ {\mathbf g}_i^{{ \mathcal{\scriptscriptstyle (D)} }} (\boldsymbol{\theta})= ({\mathbf w}_{2i}^{ \mathrm{\scriptscriptstyle T} }, {\mathbf w}_{3i}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} } \times \left(y_{i}-\beta_{x}x_{i}-\boldsymbol{\beta}_{{\mathbf z}}^{ \mathrm{\scriptscriptstyle T} }{\mathbf z}_{i}\right). $ According to our DGP design, ${\mathbf w}_{2i}$ is the valid IV vector so the moments involving ${\mathbf w}_{2i}$ equal to zero under the true parameter; those moments involving ${\mathbf w}_{3i}$ are invalid.

Let $d_w =d_{{\mathbf w}2}+s+1$ be the number of all the IVs, among which $s$ IVs are invalid. In the low-dimensional setting, we fix $s=6$ and consider $(n,d_w)=(100,50)$ and $(200,100)$. In the high-dimensional setting, we vary the number of invalid IVs to be $s\in\{6,8,13\}$ for $(n,d_w)=(100,120)$, and $s\in\{6,12,17\}$ for $(n,d_w)=(200,240)$, where $s$ is specified by rounding, in addition, $2\log n$ and $3 n^{1/3}$.

In terms of the numerical implementation, we compute the PEL estimates by the modified two-layer coordinate descent algorithm ChangTangWu2018. The SCAD penalty is used for both $P_{1,\pi}(\cdot)$ and $P_{2,\nu}(\cdot)$ in (ref) for all the numerical experiments in this paper with local quadratic approximation FanLi2001, and the tuning parameters $\nu$ and $\pi$ are chosen by the Bayesian information criterion (BIC). Specifically, we use the following BIC type function:

align[align omitted — 204 chars of source]

where $\mathrm{df}(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})$ denotes the number of nonzero elements in $\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$, and the log likelihood term is the EL ratio $ \ell(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})=2\sum_{i=1}^n\log\{1+\hat\boldsymbol{\lambda}(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})\}$.\footnote{leng2012penalized employ this BIC to choose the tuning parameter for the high-dimensional parameter $\boldsymbol{\psi}$, and ChangTangWu2018 apply it to select multiple tuning parameters. We follow their practice as (ref) remains valid in our setting where two tuning parameters are included. }

table[table omitted — 5,008 chars of source]

We first report results of moment selection by our selection criterion in (ref). Let “FP” (false positive) denote the frequency that the valid moments being not selected, and “FN” (false negative) denote the frequency that the invalid moments being selected. In Table (ref), PEL and DB-PEL denote the moment selection criterion based on the PEL estimator and its de-biased version, respectively. As expected, the strength of correlations between $\epsilon_i$ and ${\mathbf w}_{3i}$ does not affect FP, whereas FN quickly vanishes as the correlations get stronger. In all cases, larger sample sizes help reduce the chance of moment selection error.

table[table omitted — 12,010 chars of source]

The parameter of interest in the linear IV model is $\beta_{x}$ in (ref) as it characterizes the “causal effect” of the endogenous variable, which bears economic interpretation. The root-mean-square error (RMSE), bias (BIAS) and standard deviation (STD) are calculated for PEL, DB-PEL and the classical two-stage least squares (2SLS) for (ref), and an oracle estimator is added for comparison.\footnote{ The oracle EL in Section (ref) is for low-dimensional parameters and moments. To handle the large pool of orthogonal IVs, the oracle estimator here is a 2SLS estimator taking advantage of a few most relevant IVs---those with the top $0.1n$ big coefficients $\gamma_{\mathbf{w}2,j}$ in the reduced-form equation. } The results are summarized in Table (ref). The performance of PEL and the oracle is comparable, and the gaps are narrowed when the sample size is increased from $n=100$ to $n=200$, suggesting the capacity for PEL to mimic the oracle by absorbing the signal from the valid moments and in the meantime keeping the invalid ones at bay. The RMSE of 2SLS is significantly larger than those of PEL and DB-PEL.

The left panel of Figure (ref) plots the empirical cumulative distribution functions (ECDF) of the DB-PEL estimates. The dotted curve corresponds to the case of $(n,d_w,s)=(100, 120, 8)$, the dashed curve to that of $(n,d_w,s)=(200,240,12)$, and the solid curve is the cumulative distribution function (CDF) of $\mathcal{N}(0,1)$ for comparison. A better normal approximation can be obtained by the PPEL estimate, as shown in the right panel of Figure (ref), with its tuning parameter $\varsigma=0.08(n^{-1}\log p)^{1/2}$. We also present the confidence intervals (CI) according to PPEL, DB-PEL and 2SLS, as reported in Table (ref), for the 90%, 95% and 99% levels, where the CIs for DB-PEL and PPEL are predicated on the asymptotic normality from Theorems (ref) and (ref), respectively, and the CIs for 2SLS are based on its asymptotic normality as in standard textbooks. PPEL's coverage probability to the nominal counterpart is the best among the three estimators, and is much better than PEL. 2SLS's coverage probability seems too high in the high-dimensional setting, though it is acceptable in the low-dimensional case. Figure (ref) shows that the width of 2SLS's CI is much wider than that of PPEL, which is caused by efficiency loss from abandoning the potentially valid estimating equations.

figure[figure omitted — 538 chars of source]
table[table omitted — 9,093 chars of source]

Empirical application

Colonialism was widespread over the globe prior to the Second World War. While European institutions which protected private properties and checked government powers were successfully replicated in a few colonies, many others suffered expropriation of natural and human resources. In an influential study, Acemoglu2001 (AJR, henceforth) systematically explored the relationship between Europeans' mortality rates and the modes of institutions.

We revisit AJR's open-access dataset. AJR's main structural equation is $ y_i = \gamma_y + \beta_x x_i+ \beta_z z_i + \epsilon_i$, where the dependent variable $y_i$ is {\tt logarithm GDP per capita in 1995}, and the key explanatory variable of interest $x_i$ is {\tt average protection against expropriation risk}, a continuous variable of a scale between 0 and 10. The {\tt latitude} of a colony, denoted as $z_i$, is an additional control variable. For substantial measurement errors in the institutional index $x_i$ can bias the OLS estimator, the credibility of the empirical evidence counts on IVs.

AJR compiled from historical documents a seminal variable {\tt logarithm of European settler mortality}, and argued forcefully that it qualifies as a valid IV for the endogenous variable $x_i$. Their empirical evidence was based on the 2SLS regression in cross-sectional data of countries. The baseline result in AJR's Column (2) of Table 4 (p.1386) corresponds to the three estimating functions $ {\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}_i(\boldsymbol{\theta})= (w_{1i}, z_i, 1)^{{ \mathrm{\scriptscriptstyle T} }} \times (y_i -\gamma_y - \beta_x x_i - \beta_z z_i)$ by the notations in our simulation, and the 2SLS reports $\hat{\beta}_x = 1.00$ with STD 0.22.

There are another 11 potential health variables and institutional variables, which serve as potential IVs. AJR were uncertain about their validity, and they experimented in their Table 7 (p.1392) and Table 8 (p.1394) the empirical results under various IV configurations. These variables, denoted as ${\mathbf w}_2$, are (i) {\tt Malaria in 1994}, (ii) {\tt Yellow fever}, (iii) {\tt Life expectancy}, (iv) {\tt Infant mortality}, (v) {\tt Mean temperature}, (vi) {\tt Distance from coast}, (vii) {\tt European settlements in 1990}, (viii) {\tt Democracy (1st year of independence)}, (ix) {\tt Constraint on executive (1st year of independence)}, (x) {\tt Democracy in 1900}, and (xi) {\tt Constraint on executive in 1900}. (i)--(iv) are health variables, (v)--(vi) are geographic variables, and (vii)--(xi) are institutional variables. All these extra IVs are associated with the estimating functions $ {\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}_i (\boldsymbol{\theta})={\mathbf w}_{2i} \times (y_i-\gamma_y - \beta_x x_i - \beta_z z_i) \, . $ Given the moderate sample size of 56 countries, a total of 12 IVs is non-trivial.

Our PPEL estimates the main equation as \[ \widehat{\mbox{log GDP per capita}} = \underset{(1.194)}{2.048} + \underset{(0.126)}{0.945} \times \mathrm{institution} - \underset{(0.977)}{0.785} \times \mbox {lattitude}\,. \] For the estimation and inference of the key parameter of interest $\beta_x$, we compare the results of PEL, DB-PEL, PPEL and the conventional 2SLS. Its point estimation (PE), STD, and CIs are presented in Table (ref). Enjoying the efficiency gain from the extra IVs, the STD of PPEL is 0.126 when we select the tuning parameter $\varsigma=\zeta_c(n^{-1}\log p)^{1/2}$ with $\zeta_c=0.08$ (the boldface row), as suggested by our simulations. This STD achieves a 37% reduction relative to that of the plain 2SLS with the sole valid IV. Moreover, when we vary the constant $\zeta_c=0.08$ in $\varsigma$ as 0.04, 0.06, 0.12 and 0.16, PPEL is rather robust over the wide range of tuning parameters.

table[table omitted — 1,155 chars of source]

Our method is particularly important in unifying the IV selection in AJR's Tables 7 and 8 into a single set of automatically selected IVs. Among the 11 potential IVs, our moment selection criterion (ref) invalidates all institutional variables. Five IVs survive the testing: the climate variable {\tt mean temperature}, the geographic variable {\tt distance from coast}, and three health variables {\tt dummy of yellow fever}, {\tt infant mortality} and {\tt life expectancy}. All the other health and institutional variables are assessed as endogenous and are unsuitable for IVs in this study. The justification for {\tt mean temperature} and {\tt distance from coast} is straightforward because humans were unable to interfere with these natural conditions in the era of colonialism. Furthermore, AJR argued for the validity of {\tt dummy of yellow fever}, though due to concerns of lack of variation they did not employ it as the main IV (AJR's p.1393, Paragraph 2). Our variable selection result provides supportive evidence to AJR's heuristics.

Conclusion

This paper considers a general setting of an economic structural model with many potential moments, some of which may be invalid. These invalid moments must be disciplined in order to estimate the structural parameter consistently. We propose a PEL approach to estimate the parameter of interest while coping with the invalid moments. We show that the PEL estimator is normally distributed asymptotically, and invalid moments can be consistently detected thanks to the oracle property. To overcome the difficulty of estimating the bias in the limiting distribution of the PEL estimator, we further devise the PPEL approach for statistical inference of a low-dimensional object of interest, which is useful for hypothesis testing and confidence region construction. Simulation exercises are carried out to demonstrate excellent finite sample performance of our methods. We revisit an empirical application concerning economic development and shed new insight about its candidate instruments.