EconBase
← Back to paper

Culling the herd of moments with penalized empirical likelihood

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

134,829 characters

Culling the Herd of Moments with Penalized Empirical Likelihood




\bibliographystyle{agsm}
\bibpunct{(}{)}{,}{a}{}{;}

\maketitle
\begin{abstract}
Models defined by moment conditions are at the center of structural econometric estimation, but economic theory is mostly agnostic about moment selection. While a large pool of valid moments can potentially improve estimation efficiency, in the meantime a few invalid ones may undermine consistency. This paper investigates the empirical likelihood estimation of these moment-defined models in high-dimensional settings. We propose a penalized empirical likelihood (PEL) estimation and establish its oracle property with consistent detection of invalid moments. The PEL estimator is asymptotically normally distributed, and a projected PEL procedure further eliminates its asymptotic bias and provides more accurate normal approximation to the finite sample behavior. Simulation exercises demonstrate excellent numerical performance of these methods in estimation and inference.


\end{abstract}

\noindent
 {\it Keywords}: Empirical likelihood; Estimating equations; High-dimensional statistical methods; Misspecification; Moment selection; Penalized likelihood

\vfill

\thispagestyle{empty}




\newpage

\setcounter{page}{1}

\begin{spacing}{1.5}

\section{Introduction} \label{s1}


Economists' perennial pursuit of structural mechanisms leads to
models defined by moments.
These models can be written in a semiparametric form
$\mathbb{E}\{{\mathbf g} ({\mathbf X}_i;\boldsymbol{\theta}_0)\}={\mathbf 0},$ where
${\mathbf g} = (g_j)_{j \in \{1,\ldots, r\} }$ is a vector of $r$ estimating functions,
$\boldsymbol{\theta}_0$ is a vector of unknown parameters,  and ${\mathbf X}_i$ is observed data.
To estimate these models, the most popular method is
\emph{generalized method of moments}
(GMM) \citep{hansen1982large}.
\emph{Empirical likelihood} (EL) \citep{QinLawless1994} is a competitive alternative to GMM, thanks to its nice statistical properties.
Both GMM and EL are essential building blocks of modern econometrics
\citep{anatolyev2011methods}.



Ideally, economists count on economic theory to guide the choice of variables and moments.
However, the truth is that most economic theories are parsimonious abstractions and rarely pinpoint these choices in data-rich environments.
The indeterminacy of moment selection brings about three related issues.
The first is \emph{weak moments} \citep{stock2012survey}, which
threatens identification of the true value $\boldsymbol{\theta}_0$
when multiple $\boldsymbol{\theta}$'s satisfy $\mathbb{E}\{{\mathbf g} ({\mathbf X}_i;\boldsymbol{\theta})\} \approx {\mathbf 0}$.
Practitioners respond to the concerns of weak moments by adding more moment conditions in the hope
to strengthen identification, causing the second issue of
\emph{many moments} \citep{roodman2009note}.
The hazard of many moments is the possible inclusion of \emph{invalid moments}, meaning
 $\mathbb{E}\{g_j ({\mathbf X}_i;\boldsymbol{\theta}_0)\}\neq 0$ for some $j\in \{1,\ldots,r\}$ \citep{murray2006avoiding},
 which is the third issue.
These cited survey papers highlight the unease incurred by the three challenges
and econometricians' efforts in coping with them.



When an underlying economic theory is ambivalent, empirical results based on it can be controversial
and susceptible to cherry-picking.
In such a circumstance,
of great importance are data-driven methods to guide and discipline moment selection.
In the low-dimensional settings,
\citet{andrews2001consistent} propose the GMM information criteria, and
\citet{hong2003generalized} follow with the counterpart for EL.
Information criteria are evaluated exhaustively at all
combinations of moments, and the computation becomes infeasible when
there are many potential moments.
To overcome this challenge,
\citet{liao2013adaptive} ushers
the adaptive Lasso shrinkage into GMM to select among
a finite number of moments, and \citet{ChengLiao2015}
further extend it to deal with a diverging number of moments and accommodate invalid ones.



Unprecedented progress in computation and information technology
fuel an arms race between the sheer size of the data and the scale
of empirical models.
In the era of big data,
on the one hand the cost of data collection and processing is tremendously lowered and
rich datasets open new perspectives to inspect a myriad of problems; on the other hand
economists attempt to build general models to capture various sources of heterogeneity
in observational data.
Empirical applications abound with models of many potential moments.
For instance, \citet{eaton2011} create 1360 moments to estimate a structural trade model,
\citet{altonji2013modeling} match 2429 moments implied by a model of earning dynamics,
and early works of linear instrumental variables (IV) models produce thousands of instruments by interacting variables  \citep{angrist1992effect}.
Although the sample sizes in these examples are non-trivial,
the proliferation of moments calls for a moment selection procedure capable
of handling high-dimensional moments
at a magnitude unrestricted by the sample size.
In particular, invalid ones that jeopardize consistency must be identified
and ``culled'' from the herd of moments.

The quadratic form of the usual GMM criterion function is incompatible with
high-dimensional moments \citep{shi2016econometric, shi2016estimation},
and thus \citet{belloni2018high}  regularize GMM with the sup-norm.
In this paper, we contribute
the most general and versatile procedure for high-dimensional nonlinear settings,
to the best of our knowledge.
We first develop a \emph{penalized empirical likelihood} (PEL) solution \citep{ChangTangWu2018} to deal with valid and invalid moments simultaneously.
We neutralize the invalid moments
by an auxiliary parameter, following \citet{liao2013adaptive} and \citet{ChengLiao2015}.
We establish the rate of convergence and asymptotic normality of the PEL estimator, and show
the efficiency gain from incorporating extra valid moments.
Under suitable conditions, it transpires that PEL enjoys the oracle property
of consistent moment selection and parameter selection.


The asymptotic normal distribution of the PEL estimator involves a bias term caused by the high-dimensional moments.
To spare the estimation of the bias term,
we can take further actions to project out the influence of the high-dimensional nuisance parameter in the PEL estimator,
which is called \emph{projected PEL} (PPEL) \citep{Chang2020}.
The asymptotic normality of the PPEL estimator is free of bias,
which facilitates statistical inference
of the structural parameter as well as the validity of moments.
Invoking statistical learning to assist our decisions, our method fits well in the recent trend of machine learning for the automatic selection of moments and variables.



Although this paper follows \citet{ChangTangWu2018} for estimation and \citet{Chang2020} for inference procedures,
the key insight lies in the observations that the high-dimensional auxiliary parameter, which signifies
the magnitude of misspecification, can be incorporated in the
EL method as an additional high-dimensional parameter.
In this paper, ``high-dimensional'' means that the numbers of parameters and/or moments are larger than the sample size, which goes beyond the scope of \citet{liao2013adaptive} and \citet{ChengLiao2015}.
On the other hand, in order to adapt \citet{ChangTangWu2018} and \citet{Chang2020} to accommodate
misspecified moments, we must deal with the distinctive roles of the main parameter of interest and
the auxiliary parameter in identification and the technical challenges induced by them.
Differences from \citet{ChangTangWu2018} are highlighted in Section \ref{se:pel}
about the rates of converges of the two components of the parameters, and those from \citet{Chang2020}
are elaborated in Section \ref{sec:infer} about the ways of confidence region construction.







\textbf{Literature review}.
Our paper stands on strands of literature,
which are too vast to survey exhaustively.
The accumulation of moments started from the linear IV model
\citep{angrist1990lifetime, angrist1991does}. The linear IV
model motivates theoretical research on issues of many IV
\citep{bekker1994alternative}, weak IV
(\citealp{stock2005testing}, \citealp{andrews2012estimation}),
many weak IV \citep{chao2011asymptotic, hansen2014instrumental},
invalid moments and many invalid IV
\citep{kolesar2015identification,windmeijer2019use}, to name a few.
In high-dimensional contexts, \citet{belloni2012sparse} use Lasso method for IV selection
in the first stage, and \citet{belloni2014inference} deal with post-selection inference.
Utilizing the linear structure, \citet{gold2020inference} and \citet{caner2018high}
provide inferential procedures for low-dimensional parameters
in models with high-dimensional endogenous variables and high-dimensional IVs.
Our method includes the linear IV model as a special case. In particular, the case of high-dimensional
structural parameters is elaborated in Section \ref{se:hpel}.


The proliferation of moments spreads from linear IV models to nonlinear
models. For example, in empirical industrial organization
researchers bring in moments from various resources,
some of which are guided by economic theory, to
mitigate the concerns of weak identification and improve estimation efficiency
\citep{ackerberg2007econometric}.
In empirical macroeconomics, identification failure and moment misspecification are
common issues \citep{mavroeidis2005identification}.
Under the GMM framework,
inference under weak moments (\citealp{stock2000gmm}, \citealp{kleibergen2005testing}, \citealp{andrews2020optimal}),
estimation under many weak moments \citep{han2006gmm}, and robust procedures for
invalid moments (\citealp{ditraglia2016using}, \citealp{caner2018adaptive}) have been developed.


EL's attractive theoretical properties are studied extensively
(\citealp{kitamura2001asymptotic}, \citealp{otsu2010bahadur}, \citealp{matsushita2013second}, \citealp{ChangChenChen2015}).
\citet{otsu2006generalized} and \citet{newey2009generalized} deal with its inference under weak IV,
and \citet{caner2015hybrid} select instruments in linear IV models.
Penalization schemes on EL have been introduced by \citet{otsu2007penalized}, \citet{tang2010penalized} and \citet{ChangTangWu2018}.




\textbf{Organization.} The rest of the paper is organized as follows. Section \ref{s2} introduces
the model and the analytic framework.
We first derive the asymptotic properties of our estimation in the low-dimensional case in Section 3,
extend them to the high-dimensional structural parameter in Section 4,
and we further refine PEL with projection to eliminate its bias in Section 5.
The theoretical results are supported in Section 6 by Monte Carlo simulations.
The influential study of the determinants of economic outcomes after colonialism is revisited
in Section 7. Section 8 concludes the paper.
Due to the limitations of space, the proofs and the technical details of secondary importance are relegated into the supplementary materials.

\textbf{Notations.} We conclude Introduction with notations used throughout the paper.
``Low-dimensional'' is referred to the cases that the number of parameters or moments is much smaller than the sample size $n$, whereas ``high-dimensional'' goes the opposite.
For two sequences of positive numbers $\{a_n\}$ and $\{b_n\}$, we write $a_n\lesssim b_n$ or $b_n\gtrsim a_n$
if there exists a positive constant $c$ such that $\limsup_{n\rightarrow\infty}a_n/b_n\leq c$,
and write $a_n\ll b_n$ or $b_n\gg a_n$ if $\limsup_{n\rightarrow\infty}a_n/b_n=0$.


Denote by $1(\cdot)$ the indicator function. For a positive integer $q$, we write $[q]=\{1,\ldots,q\}$. For a $q\times q$ symmetric matrix ${\mathbf M}$, denote by $\lambda_{\min}({\mathbf M})$ and $\lambda_{\max}({\mathbf M})$ the smallest and largest eigenvalues of ${\mathbf M}$, respectively.
For a $q_1\times q_2$ matrix ${\mathbf B}=(b_{i,j})_{q_1\times q_2}$, let ${\mathbf B}^{ \mathrm{\scriptscriptstyle T} }$ be its transpose, ${\mathbf B}^{\otimes2}={\mathbf B}{\mathbf B}^{ \mathrm{\scriptscriptstyle T} }$, ${\mathbf B}^{\circ\kappa}=(|b_{i,j}|^\kappa)_{q_1\times q_2}$ for any $\kappa>0$, $|{\mathbf B}|_\infty=\max_{i\in[q_1],j\in[q_2]}|b_{i,j}|$ be the sup-norm,
and $\|{\mathbf B}\|_2=\lambda_{\max}^{1/2}({\mathbf B}^{\otimes2})$ be the spectral norm.
Specifically, if $q_2=1$, we use $|{\mathbf B}|_\infty = \max_{i\in[q_1]}|b_{i,1}|$, $|{\mathbf B}|_1=\sum_{i=1}^{q_1}|b_{i,1}|$ and $|{\mathbf B}|_2=(\sum_{i=1}^{q_1}b_{i,1}^2)^{1/2}$ to denote the $L_{\infty}$-norm, $L_1$-norm and $L_2$-norm of the $q_1$-dimensional vector ${\mathbf B}$, respectively.
For two square matrices ${\mathbf M}_1$ and ${\mathbf M}_2$, we say ${\mathbf M}_1 \le {\mathbf M}_2$ if $({\mathbf M}_2 - {\mathbf M}_1)$ is a positive semi-definite matrix.

The population mean is denoted by $\mathbb{E}(\cdot)$, and the sample mean by $\mathbb{E}_n(\cdot)= n^{-1}\sum_{i=1}^n \{\cdot\}$.
For a given index set $\mathcal{L}$, let $| \mathcal{L} | $ be its cardinality.
For a generic multivariate function ${\mathbf h}(\cdot;\cdot)$, we denote by ${\mathbf h}_{{ \mathcal{\scriptscriptstyle L} }}(\cdot;\cdot)$ the subvector of ${\mathbf h}(\cdot;\cdot)$ collecting the components indexed by $\mathcal{L}$.
Analogously, we write ${\mathbf a}_{{ \mathcal{\scriptscriptstyle L} }}$ as the corresponding subvector of ${\mathbf a}$.
For simplicity and when no confusion arises, we use the generic notation ${\mathbf h}_i(\boldsymbol{\theta})$ as the equivalence to ${\mathbf h}({\mathbf X}_i;\boldsymbol{\theta})$,
and $\nabla_{\boldsymbol{\theta}} {\mathbf h}_i(\boldsymbol{\theta}) $ for the first-order partial derivative of ${\mathbf h}_i(\boldsymbol{\theta})$ with respect to $\boldsymbol{\theta}$.
Denote by $h_{i,k}(\boldsymbol{\theta})$ the $k$-th component of ${\mathbf h}_i(\boldsymbol{\theta})$,
and by $\nabla^2_{\boldsymbol{\theta}}h_{i,k}(\boldsymbol{\theta}) $ the second derivative of $h_{i,k}(\boldsymbol{\theta})$ with respect to $\boldsymbol{\theta}$.
Let $\bar {\mathbf h}(\boldsymbol{\theta})=\mathbb{E}_n\{{\mathbf h}_i (\boldsymbol{\theta})\}$,
and write its $k$-th component as $\bar h_k(\boldsymbol{\theta})= \mathbb{E}_n\{h_{i,k}(\boldsymbol{\theta})\}$.
Analogously, let ${\mathbf h}_{i,{ \mathcal{\scriptscriptstyle L} }}(\boldsymbol{\theta})={\mathbf h}_{{ \mathcal{\scriptscriptstyle L} }}({\mathbf X}_i;\boldsymbol{\theta})$
and $\bar{{\mathbf h}}_{{ \mathcal{\scriptscriptstyle L} }}(\boldsymbol{\theta})= \mathbb{E}_n \{{\mathbf h}_{i,{ \mathcal{\scriptscriptstyle L} }}(\boldsymbol{\theta})\}$.



\section{Empirical likelihood with a herd of moments}
\label{s2}

In this section we introduce the model and the EL estimation. Let ${\mathbf X}_1,\ldots,{\mathbf X}_n$ be $d$-dimensional independent and identically distributed generic observations, and $\boldsymbol{\theta}=(\theta_1,\ldots,\theta_p)^{ \mathrm{\scriptscriptstyle T} }$ be a $p$-dimensional parameter taking values in $\boldsymbol{\Theta} \subset \mathbb{R}^p$. For a set of $r_1$ estimating functions
    $
    {\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}(\cdot;\cdot)=\{ g_{j}^{ \mathcal{\scriptscriptstyle (I)} } (\cdot; \cdot) \}_{j \in \mathcal{I}}
    $,
the information of the model parameter $\boldsymbol{\theta}$ is collected by the unbiased moment condition
    \begin{equation}\label{eq:esteq}
    {\mathbf 0} = \mathbb{E}\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}({\mathbf X}_i;\boldsymbol{\theta}_0)\} = \mathbb{E}\{{\mathbf g}_i^{{ \mathcal{\scriptscriptstyle (I)} }}(\boldsymbol{\theta}_0)\}
    \end{equation}
at the unknown true parameter $\boldsymbol{\theta}_0\in\boldsymbol{\Theta}$, where $r_1\geq p$ is necessary for identifying $\boldsymbol{\theta}_0$.
The superscript $(\mathcal{I})$ labels this \emph{initial} set, to be distinguished from the other
set $(\mathcal{D})$ in Section \ref{sec:add}.


\subsection{EL estimation with valid moments}

Motivated from empirical applications in asset pricing,
the two-step GMM \citep{hansen1982large} was the default estimating method for the moment-defined model (\ref{eq:esteq}).
Intensive theoretical studies and numerical evidence in 1980's and 90's revealed some undesirable finite-sample properties of
the two-step GMM \citep{altonji1996small}.
EL and the continuously updating GMM (CUE) \citep{hansen1996finite} emerged as competitive solutions,
and they were later unified as members of \emph{generalized empirical likelihood} \citep{newey2003higher}.


This paper focuses on EL with estimating equations:
    \begin{equation*}\label{eq:el}
    L(\boldsymbol{\theta})=\max\bigg\{\prod_{i=1}^n\pi_i:\pi_i>0\,,~\sum_{i=1}^n\pi_i=1\,,~\sum_{i=1}^n\pi_i{\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}_i(\boldsymbol{\theta})={\mathbf 0}\bigg\}\,,
    \end{equation*}
proposed by \cite{QinLawless1994}  based on the seminal idea of EL \citep{Owen1988,Owen1990}.
Maximizing $L(\boldsymbol{\theta})$ with respect to $\boldsymbol{\theta}$ delivers the EL estimator $ \hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }}=\arg\max_{\boldsymbol{\theta}\in\boldsymbol{\Theta}}L(\boldsymbol{\theta})$, which can be carried out equivalently by solving the corresponding dual problem
    \begin{equation}\label{eq:elestk}
    \hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }}=\arg\min_{\boldsymbol{\theta}\in\boldsymbol{\Theta}}\max_{\boldsymbol{\lambda}\in\hat{\Lambda}_{n}^{{ \mathcal{\scriptscriptstyle (I)} }}(\boldsymbol{\theta})}\sum_{i=1}^n\log\{1+\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}_i(\boldsymbol{\theta})\}\,,
    \end{equation}
where $\hat{\Lambda}_{n}^{{ \mathcal{\scriptscriptstyle (I)} }}(\boldsymbol{\theta})=\{\boldsymbol{\lambda}\in\mathbb{R}^{r_1}:\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}_i(\boldsymbol{\theta})\in\mathcal {V}~\textrm{for any}~i\in[n]\}$ and  $\mathcal{V}$ is an open interval containing zero.

To fix ideas, we first study the asymptotic property of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }}$ under
regularity conditions.
When the sample size $n$ grows,  we adopt the asymptotic framework of  \cite{Hjort2009} and \cite{ChangChenChen2015} to take
the observations $\{{\mathbf g}_i^{{ \mathcal{\scriptscriptstyle (I)} }}(\boldsymbol{\theta})\}_{i=1}^n$ as a multi-index array, where $r_1$, $p$ and $d$ may depend on $n$.
Proposition A.1 in the supplementary materials shows that the standard asymptotic normality for
$\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }}$ holds under
mild regularity conditions.






\subsection{Oracle EL estimation} \label{sec:add}

In applied econometrics, there has been a tendency of assembling many IVs or creating
many moments to identify the parameter of interest.
In these applications, researchers often have some
ideas about the relative importance of moments;
in the meantime, when researchers work with more and more moments,
some invalid ones may creep in.


\begin{example}
\citet{eaton2011} deem as the key moments 128($=2^7$) combinations of the largest 7 trade partners of France,
which are more important than the other 1232 moments.
\citet{angrist1990lifetime} treats the date of birth (DOB) as the key IV and it is complemented by the IVs generated by interactions;  recently \citet{kolesar2015identification} raise the potential invalidity
among these interaction terms.\footnote{Identification of the simple linear IV model requires the IVs satisfying the orthogonality condition and the relevance condition.
However, the more relevant an IV to the endogenous variables,
the more likely it is that the so-called ``IV'' is correlated with the structural
error, thereby violating orthogonality. There is a thin line between a valid IV and an invalid one.
}
\end{example}





To put into an analytic framework a herd of extra moments with unknown validity \emph{ex ante},
suppose that $r_2$ estimating functions
${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}=\{g_{j}^{ \mathcal{\scriptscriptstyle (D)} }\}_{j\in \mathcal{D}}$ are partitioned into two groups:
\[
\mathcal{A}=\{j\in \mathcal{D}:\mathbb{E}\{g^{ \mathcal{\scriptscriptstyle (D)} }_{i,j}(\boldsymbol{\theta}_0)\}=0\}~~~\textrm{and}~~~
\mathcal{A}^{{\mathrm{c}}}=\{j\in \mathcal{D}:\mathbb{E}\{g^{ \mathcal{\scriptscriptstyle (D)} }_{i,j}(\boldsymbol{\theta}_0)\}\neq 0\}\,.
\]
Here the estimating functions in the set $\mathcal{A}$ are correctly specified, which can help improve the efficiency in estimating $\boldsymbol{\theta}_0$.
On the contrary, those in $\mathcal{A}^{{\mathrm{c}}}$ are misspecified, and they can only undermine the identification of $\boldsymbol{\theta}_0$.






If there is an ``oracle'' that reveals which components of ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}$ belong to $\mathcal{A}$, i.e., the valid estimating functions, we can collect all valid estimating functions indexed by $\mathcal{H}:=\mathcal{I}\cup\mathcal{A}$ to estimate $\boldsymbol{\theta}_0$.
Denote ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (H)} }} =\{ {\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} },{ \mathrm{\scriptscriptstyle T} }},{\mathbf g}^{{ \mathcal{\scriptscriptstyle (A)} },{ \mathrm{\scriptscriptstyle T} }}\}^{ \mathrm{\scriptscriptstyle T} }$
 and $h=|\mathcal{H}|$ as the total number of valid moments. The associated EL estimation for $\boldsymbol{\theta}_0$ based on the estimating functions ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (H)} }}$ is given by
\begin{equation*}\label{eq:elest11}
\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (H)} }}=\arg\min_{\boldsymbol{\theta}\in\boldsymbol{\Theta}}\max_{\boldsymbol{\lambda}\in\hat{\Lambda}_{n}^{{ \mathcal{\scriptscriptstyle (H)} }}(\boldsymbol{\theta})}\sum_{i=1}^n\log\{1+\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (H)} }}_i(\boldsymbol{\theta})\}\,,
\end{equation*}
where $\hat{\Lambda}_{n}^{{ \mathcal{\scriptscriptstyle (H)} }}(\boldsymbol{\theta})=\{\boldsymbol{\lambda}\in\mathbb{R}^h:\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (H)} }}_i(\boldsymbol{\theta})\in\mathcal {V}~\textrm{for any}~i\in[n]\}$.
By the same arguments as those in Proposition A.1 in the supplementary materials,
the following Proposition \ref{cy1} verifies the asymptotic normality for $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (H)} }}$.



\begin{proposition}[Oracle property]\label{cy1}
Assume that ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (H)} }} $ satisfies the conditions {\rm(A.1)--(A.5)} in the supplementary materials associated with ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (H)} }}$. If $h^{3}n^{-1+2/\gamma}=o(1)$ and $h^{3}p^2n^{-1}=o(1)$, then
$
\sqrt{n}\boldsymbol{\alpha}^{ \mathrm{\scriptscriptstyle T} }\{{\mathbf J}^{{ \mathcal{\scriptscriptstyle (H)} }}\}^{1/2}\{\hat{\boldsymbol{\theta}}_{{{ \mathrm{\scriptscriptstyle EL} }}}^{{ \mathcal{\scriptscriptstyle (H)} }}-\boldsymbol{\theta}_0\}\xrightarrow{d}\mathcal{N}(0,1)$
as $n\rightarrow\infty$ for any $\boldsymbol{\alpha}\in\mathbb{R}^p$ with $|\boldsymbol{\alpha}|_2=1$,  where
$
{\mathbf J}^{{ \mathcal{\scriptscriptstyle (H)} }}=([\mathbb{E}\{\nabla_{\boldsymbol{\theta}}{\mathbf g}^{{ \mathcal{\scriptscriptstyle (H)} }}_i(\boldsymbol{\theta}_0)\}]^{ \mathrm{\scriptscriptstyle T} }\{{\mathbf V}^{{ \mathcal{\scriptscriptstyle (H)} }}(\boldsymbol{\theta}_0)\}^{-1/2})^{\otimes2}
$
with ${\mathbf V}^{{ \mathcal{\scriptscriptstyle (H)} }}(\boldsymbol{\theta}_0)=\mathbb{E}\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (H)} }}_i(\boldsymbol{\theta}_0)^{\otimes2}\}$.
\end{proposition}


Obviously $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (H)} }}$ is more efficient than $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }}$
thanks for the additional moment restrictions in $\mathcal{A}$, which
disclose more information about the parameter and thus tie down the estimation variability.
This is parallel to \citet{hall2007information} in GMM for finite numbers of parameters and moments.
In reality, however, we are oblivious to $\mathcal{A}$, and therefore
we cannot blindly summon all $r_2$ estimating functions in ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}$
to estimate $\boldsymbol{\theta}_0$.
Following \citet{liao2013adaptive}, we introduce an auxiliary parameter
$\boldsymbol{\xi}=\mathbb{E}\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}_i(\boldsymbol{\theta})\}$ to neutralize the misspecified moments.
Denote the augmented parameter $\boldsymbol{\psi}=(\boldsymbol{\theta}^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$,
and the augmented parameter space $\boldsymbol{\Psi} = \boldsymbol{\Theta}\times\mathbf{\Upsilon}$.
Let $r=r_1+r_2$ and stack these $r$ estimating functions as
${\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}({\mathbf X};\boldsymbol{\psi})=\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}({\mathbf X};\boldsymbol{\theta})^{ \mathrm{\scriptscriptstyle T} },{\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}({\mathbf X};\boldsymbol{\theta})^{ \mathrm{\scriptscriptstyle T} }-\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} }\}^{ \mathrm{\scriptscriptstyle T} }$.
Then $\boldsymbol{\psi}_0=(\boldsymbol{\theta}_0^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\xi}_0^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ with $\boldsymbol{\xi}_0=\mathbb{E}\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}_i(\boldsymbol{\theta}_0)\}$ can be identified by
\begin{equation}\label{eq:esteq2}
\mathbb{E}\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\boldsymbol{\psi}_0)\}
=
\mathbb{E}
\begin{Bmatrix}
{\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}_i(\boldsymbol{\theta}_0) \\
{\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}_i(\boldsymbol{\theta}_0)-\boldsymbol{\xi}_0
\end{Bmatrix}
={\mathbf 0}\, .
\end{equation}

Can we directly include the auxiliary parameter $\boldsymbol{\xi}$ into the EL estimation? The answer is negative. If the auxiliary parameter is not regularized,
the EL estimators with and without $\boldsymbol{\xi}$ are the same, up to numerical errors.
\begin{proposition}\label{prop11}
If $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }} $ is the unique solution of {\rm\eqref{eq:elestk}}, then $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (T)} }} = \hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }} $, where $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (T)} }}$ is the associated subvector of
$\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (T)} }} =\arg\min_{\boldsymbol{\psi}\in \boldsymbol{\Psi}
    }\max_{\boldsymbol{\lambda}\in\hat{\Lambda}_{n}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})
    }\sum_{i=1}^n\log\{1+\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\boldsymbol{\psi})\}$
for the estimation of $\boldsymbol{\theta}_0$ with $\hat{\Lambda}_{n}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})=\{\boldsymbol{\lambda}\in\mathbb{R}^r:\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\boldsymbol{\psi})\in\mathcal {V}~\textrm{for any}~i\in[n]\}$.
\end{proposition}


Proposition \ref{prop11} states the equivalence between $\hat\boldsymbol{\theta}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathrm{\scriptscriptstyle EL} }}$ and $\hat\boldsymbol{\theta}^{{ \mathcal{\scriptscriptstyle (I)} }}_{{ \mathrm{\scriptscriptstyle EL} }}$. In the next section, we will show that desirable efficiency as in
the oracle estimator $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (H)} }}$
can be achieved if we extend \citet{liao2013adaptive}'s idea of shrinking the auxiliary parameter\footnote{
In GMM estimation under a fixed $r$,  \citet{liao2013adaptive} shrinks $\boldsymbol{\xi}$ toward zero using the adaptive Lasso \citep{zou2006adaptive}.
}
by further penalizing the Lagrange multipliers associated with the high-dimensional moments.



\section{High-dimensional moments with low-dimensional $\boldsymbol{\theta}$}\label{se:pel}

We consider in this section the model \eqref{eq:esteq2} with a low-dimensional parameter $\boldsymbol{\theta}$,
low-dimensional estimating functions ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}$ and
high-dimensional estimating functions ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}$. In this setting, $p$ and $r_1$ are either fixed or diverge at some slow polynomial rates of $n$,  whereas $r_2$ can grow much larger than $n$.
The components in the parameter $\boldsymbol{\theta}$ are indexed by $\mathcal{P}$.
All moments together are indexed by $\mathcal{T} = \mathcal{I} \cup \mathcal{D}$, in which
the researcher knows that those in $\mathcal{I}$ are correctly specified
in the sense that $\mathbb{E}\{ {\mathbf g}_i^{{ \mathcal{\scriptscriptstyle (I)} }}(\boldsymbol{\theta}_0)\} = {\mathbf 0}$,
whereas she is uncertain whether those in $\mathcal{D} = \mathcal{A}\cup \mathcal{A}^{{\mathrm{c}}} $
satisfy $\mathbb{E}\{ {\mathbf g}_i^{{ \mathcal{\scriptscriptstyle (D)} }}(\boldsymbol{\theta}_0)\} = {\mathbf 0} $ or not.
As there is a one-to-one relationship between $\boldsymbol{\xi}$ and the uncertain moments, we use $\mathcal{D}$ to index the components in $\boldsymbol{\xi}$ as well.


We propose the following PEL
to simultaneously estimate the unknown parameter $\boldsymbol{\theta}_0$ and determine the validity of the estimating functions in
${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}$:
\begin{equation}\label{eq:est11}
  (\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} },\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }
 =
 \arg\min_{\boldsymbol{\psi}\in \boldsymbol{\Psi}}\max_{\boldsymbol{\lambda}\in\hat{\Lambda}_{n}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})}\bigg[\frac{1}{n}\sum_{i=1}^n\log\{1+\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\boldsymbol{\psi})\} -\sum_{j\in \mathcal{D} } P_{2,\nu}(|\lambda_j|)
 +\sum_{k\in \mathcal{D} }P_{1,\pi}(|\xi_k|)\bigg]\,,
\end{equation}
where $\hat{\Lambda}_{n}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})=\{\boldsymbol{\lambda}\in\mathbb{R}^{r}:\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\boldsymbol{\psi})\in\mathcal {V}~\textrm{for any}~i\in[n]\}$, and
$P_{1,\pi}(\cdot)$ and $P_{2,\nu}(\cdot)$ are two penalty functions with tuning parameters $\pi$ and $\nu$, respectively.
With the penalty function $P_{2,\nu}(\cdot)$ and appropriately selected tuning parameter $\nu$, the estimator $(\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} },\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ is associated with a sparse Lagrange multiplier $\boldsymbol{\lambda}$.
Since the sparse $\boldsymbol{\lambda}$ invokes a subset of the estimating functions ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}(\cdot;\cdot)$,
it digests the high-dimensional moments as long as the number of nonzero components in $\boldsymbol{\lambda}$ is small.
On the other hand, the penalty $P_{1,\pi}(\cdot)$ is applied to identify which components in ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}(\cdot;\cdot)$ are correctly specified and can estimate $(\xi_{0,k})_{k\in\mathcal{A}}$ exactly as $0$ with high probability.

These penalties originally appeared in \citet{ChangTangWu2018} though,
there are several important differences.
First, given the low-dimensional $\boldsymbol{\theta}$, it is unnecessary to assume sparsity on $\boldsymbol{\theta}_0$ for identification.
As a result, our procedure here in \eqref{eq:est11} only penalizes $\boldsymbol{\xi}$ and $\boldsymbol{\lambda}_{\mathcal{D}}$, not the entire parameter $\boldsymbol{\psi}=(\boldsymbol{\theta}^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ and the associated Lagrange multiplier $\boldsymbol{\lambda}$.
Were $\boldsymbol{\lambda}_{\mathcal{I}}$ penalized, we would rule out some components of ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}(\cdot;\cdot)$
and thus may suffer loss of estimation efficiency for $\boldsymbol{\theta}_0$.
Second, the identification condition invoked in this paper
deviates from that in \citet{ChangTangWu2018}. Our Condition \ref{A.1} for identification is imposed on $\boldsymbol{\theta}_0$,
rather than the whole parameter $\boldsymbol{\psi}_0$.
In contrast, \citet{ChangTangWu2018} specify identification condition for the parameter entity.
While estimators in \citet{ChangTangWu2018} share the same rate of convergence,
we must deal with the disparate rates of convergence of the main component $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ and the auxiliary $\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$,
which significantly complicates the theoretical analysis of \eqref{eq:est11},
as detailed in Proposition \ref{thm1} below.



For any penalty function $P_\tau(\cdot)$ with a tuning parameter $\tau$, let $\rho(t;\tau)=\tau^{-1}P_\tau(t)$ for any $t\in[0,\infty)$ and $\tau\in(0,\infty)$. We assume that the two penalty functions $P_{1,\pi}(\cdot)$ and $P_{2,\nu}(\cdot)$ involved in \eqref{eq:est11} belong to the following class:
\begin{equation}\label{eq:classp}
\begin{split}
\mathscr{P}=\{P_\tau(\cdot):&~\rho(t;\tau)~\mbox{is increasing in}~t\in[0,\infty)~\mbox{and has continuous derivative}\\
&~\rho'(t;\tau)~\mbox{for any}~t\in(0,\infty)~\mbox{with} ~\rho'(0^+;\tau)\in(0,\infty),~\mbox{where}\\
&~\rho'(0^+;\tau)~\mbox{is independent of}~\tau\}\,.
\end{split}
\end{equation}
This class $\mathscr{P}$, considered in \cite{LvFan2009}, is broad and general. The commonly used $L_1$ penalty, SCAD penalty \citep{FanLi2001} and MCP penalty \citep{Zhang2010} are all included in $\mathscr{P}$. When $P_\tau(\cdot)\in\mathscr{P}$, we write the associated $\rho'(0^+;\tau)$ as $\rho'(0^+)$ for simplification.



To establish the limiting distribution of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ in \eqref{eq:est11}, we assume the following regularity conditions.


\begin{as}\label{A.1}
There exists a universal constant $K_1>0$ such that
$$\inf_{\boldsymbol{\theta}\in\boldsymbol{\Theta}:\,|\boldsymbol{\theta}-\boldsymbol{\theta}_{0}|_\infty>\varepsilon}|\mathbb{E}\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}_i(\boldsymbol{\theta})\}|_\infty\ge K_1 \varepsilon$$ for any $\varepsilon>0$.
\end{as}



\begin{as}\label{A.2}
There exist universal constants $K_2>0$ and $\gamma>4$ such that
$$
\max_{j\in\mathcal{I}}\mathbb{E}\bigg\{\sup_{\boldsymbol{\theta}\in\boldsymbol{\Theta}}|g^{{ \mathcal{\scriptscriptstyle (I)} }}_{i,j}(\boldsymbol{\theta})|^\gamma\bigg\}+\max_{j\in\mathcal{D}}\mathbb{E}\bigg\{\sup_{\boldsymbol{\theta}\in\boldsymbol{\Theta}}|g^{{ \mathcal{\scriptscriptstyle (D)} }}_{i,j}(\boldsymbol{\theta})|^\gamma\bigg\}\le K_2\,.$$
\end{as}


\begin{as}\label{A.4}
For each $j\in \mathcal{T}$, the function $g^{{ \mathcal{\scriptscriptstyle (T)} }}_{j}({\mathbf X};\boldsymbol{\psi})$ is twice continuously differentiable with respect to $\boldsymbol{\psi} \in\boldsymbol{\Psi}$ for any ${\mathbf X}$. For $\gamma$ specified in Condition \ref{A.2}, $$
\sup_{\boldsymbol{\psi}\in\boldsymbol{\Psi}}
\bigg(
    |   \mathbb{E}_n[\{\nabla_{\boldsymbol{\psi}}  {\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_{i}(\boldsymbol{\psi})\}^{\circ2}]
    |_{\infty} + \max_{j\in \mathcal{T}}
 | \mathbb{E}_n[\{\nabla^2_{\boldsymbol{\psi}}g^{{ \mathcal{\scriptscriptstyle (T)} }}_{i,j}(\boldsymbol{\psi}) \}^{\circ2}]
 |_{\infty} +
| \mathbb{E}_n[\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_{i}(\boldsymbol{\psi})\}^{\circ\gamma}]|_{\infty}\bigg)
 =O_{ \mathrm{p} }(1)\,.$$
There is some universal constant $K_3>0$ such that
$\sup_{\boldsymbol{\theta}\in\boldsymbol{\Theta}}
|\mathbb{E} \{ \nabla_{\boldsymbol{\theta}}   {\mathbf g}_{i,{ \mathcal{\scriptscriptstyle A} }^{{\mathrm{c}}}}^{ \mathcal{\scriptscriptstyle (D)} }(\boldsymbol{\theta}) \} |_\infty \le K_3$.
\end{as}


For any index set $\mathcal{F}\subset\mathcal{T}$ and $\boldsymbol{\psi}\in\boldsymbol{\Psi}$, define $
{\mathbf V}_{{ \mathcal{\scriptscriptstyle F} }}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})=\mathbb{E}\{{\mathbf g}_{i,{ \mathcal{\scriptscriptstyle F} }}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})^{\otimes2}\}$. When $\mathcal{F}=\mathcal{T}$, we write ${\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})={\mathbf V}_{{ \mathcal{\scriptscriptstyle T} }}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})$ for conciseness.


\begin{as}\label{A.3}
There exists a universal constant $K_4>1$ such that $K_4^{-1}<\lambda_{\rm min}\{{\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi}_0)\}\le\lambda_{\rm max}\{{\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi}_0)\}<K_4$.
\end{as}

Conditions \ref{A.1}--\ref{A.3} are standard regularity assumptions in the literature.
Condition \ref{A.1} is an identification assumption of the estimating equations in the known set $\mathcal{I}$.
Write $\boldsymbol{\xi}_0= (\xi_{0,k})_{k\in\mathcal{D}}$, and we continue by defining
\begin{align}\label{eq:an}
a_n=\sum_{k\in\mathcal{D}}P_{1,\pi}(|\xi_{0,k}|)~~\textrm{and}~~
\phi_n=\max\{pa_n^{1/2},pr_1^{1/2}\aleph_n,\nu\}\,
\end{align}
in this section, where $\aleph_n=(n^{-1}\log r)^{1/2}$.
Suppose:
\begin{equation}\label{eq:xi1}
\mbox{\begin{minipage}[c]{.55\linewidth}
There exist $\chi_n\to0$ and $c_n\to0$ with $\phi_nc_n^{-1}\to0$ such that $
\max_{k\in\mathcal{A}^{{\mathrm{c}}}}\sup_{0<t<|\xi_{0,k}|+c_n}P'_{1,\pi}(t)=O(\chi_n)$
\end{minipage} }
\end{equation}
to control the bias induced by $P_{1,\pi}(\cdot)$ on $\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$.
With the assumption $\phi_n=o(\min_{k\in\mathcal{A}^{{\mathrm{c}}}}|\xi_{0,k}|)$ that
the nonzero components of $\boldsymbol{\xi}_0$ do not diminish to zero too fast, \eqref{eq:xi1} can be replaced by
\begin{align}\label{eq:xi}
\max_{k\in\mathcal{A}^{{\mathrm{c}}}}\sup_{c|\xi_{0,k}|<t<c^{-1}|\xi_{0,k}|}P'_{1,\pi}(t)=O(\chi_n)
\end{align}
for some constant $c\in(0,1)$. If we select $P_{1,\pi}(\cdot)$ as an asymptotically unbiased penalty such as SCAD or MCP, we have $\chi_n=0$ in \eqref{eq:xi} when
\begin{align}\label{eq:signal}
\min_{k\in\mathcal{A}^{{\mathrm{c}}}}|\xi_{0,k}|\gg\max\{\phi_n,\pi\}\,.
\end{align}
To simplify the presentation,  in this section we assume that \eqref{eq:signal} holds and
 $\chi_n=0$ in \eqref{eq:xi}.\footnote{
If \eqref{eq:signal} is violated,  the asymptotic normality of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ in Theorem \ref{thm2} below will still hold under \eqref{eq:xi1} along with
more complicated notations to spell out the restrictions.
}

To allocate the parameter of interest $\boldsymbol{\theta}$ and those invalid moments in $\mathcal{A}^{{\mathrm{c}}}$,
we define an index set
$\mathcal{S}=\mathcal{P}\cup \mathcal{A}^{{\mathrm{c}}}$ and its cardinality $s=|\mathcal{S}|$.
For any $\boldsymbol{\psi} \in\boldsymbol{\Psi}$, the set $\mathcal{S}$ picks out $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }}=(\boldsymbol{\theta}^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\xi}_{{ \mathcal{\scriptscriptstyle A} }^{{\mathrm{c}}}}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$.
Since  $\mathcal{A} = \mathcal{S}^{{\mathrm{c}}} = (\mathcal{P}\cup\mathcal{D}) \backslash \mathcal{S}$,
the corresponding auxiliary coefficients can be written as
 $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}=\boldsymbol{\xi}_{{ \mathcal{\scriptscriptstyle A} }}$ and furthermore
$\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}={\mathbf 0}$  as $\mathcal{A}$ is the set of valid moments. Given some constant $C_*\in(0,1)$,
define $\mathcal{M}^*_{\boldsymbol{\psi}}= \mathcal{I} \cup  \mathcal{D}^*_{\boldsymbol{\psi}}$
with $\mathcal{D}^*_{\boldsymbol{\psi}}=\{ j\in \mathcal{D}: |\bar{g}^{{ \mathcal{\scriptscriptstyle (T)} }}_{j}(\boldsymbol{\psi})|\ge C_*\nu\rho'_2(0^+)\}$.
We assume the existence of a sequence $\ell_n\to\infty$ such that
$$
\mathbb{P}\bigg(\max_{\boldsymbol{\psi}\in\boldsymbol{\Psi}:\,|\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }}-\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle S} }}|_\infty\le c_n,\,|\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}|_1\leq \aleph_n }|\mathcal{M}^*_{\boldsymbol{\psi}}| \le \ell_n\bigg) \to 1
$$
with some $c_n\gg\phi_n$.


Let $b_{1,n}=\max\{a_n,r_{1}\aleph_n^2\}$ and $b_{2,n}=\max\{b_{1,n},\nu^2 \}$. Then
$
\phi_n=\max\{pb_{1,n}^{1/2}, b_{2,n}^{1/2} \}$.
Define $\boldsymbol{\Psi}_*=\{\boldsymbol{\psi}=(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }}^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }:|\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }}-\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle S} }}|_\infty\le\varepsilon,|\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}|_1\le \aleph_n \}$
for some fixed $\varepsilon>0$. Consider
\begin{align}\label{eq:locest}
\hat{\boldsymbol{\psi}}=\arg\min_{\boldsymbol{\psi}\in\boldsymbol{\Psi}_*}\max_{\boldsymbol{\lambda}\in\hat{\Lambda}_{n}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})}\bigg[\frac{1}{n}\sum_{i=1}^n\log\{1+\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\boldsymbol{\psi})\} -\sum_{j\in \mathcal{D} } P_{2,\nu}(|\lambda_j|)
 +\sum_{k\in \mathcal{D} }P_{1,\pi}(|\xi_k|)\bigg]\,.
\end{align}
Proposition \ref{thm1} shows that such a $\hat{\boldsymbol{\psi}}$ is a sparse local minimizer for \eqref{eq:est11}.

\begin{proposition} \label{thm1}
Let $P_{1,\pi}(\cdot),P_{2,\nu}(\cdot)\in\mathscr{P}$ for $\mathscr{P}$ defined in {\rm(\ref{eq:classp})}, and $P_{2,\nu}(\cdot)$ be convex with bounded second derivative around $0$.
For $\hat\boldsymbol{\psi}$ defined as {\rm\eqref{eq:locest}}, assume there exists a constant $\tilde{c}\in(C_*,1)$ such that $\mathbb{P}[\cup_{j\in\mathcal{T}}\{|\bar{g}_j^{ \mathcal{\scriptscriptstyle (T)} }(\hat\boldsymbol{\psi})|\in[\tilde{c}\nu\rho'_2(0^+),\nu\rho'_2(0^+))\}]\rightarrow0$.  Under Conditions {\rm \ref{A.1}--\ref{A.3}} and \eqref{eq:signal}, if $\log r=o(n^{1/3})$, $s^2\ell_n\phi_n^{2}=o(1)$, $b_{2,n}=o(n^{-2/\gamma})$, $\ell_n\aleph_n=o(\nu)$ and $\ell_n^{1/2}\aleph_n=o(\pi)$,
 then with probability approaching one this $\hat\boldsymbol{\psi}=(\hat\boldsymbol{\theta}^{ \mathrm{\scriptscriptstyle T} },\hat\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} }_{{ \mathcal{\scriptscriptstyle A} }^{{\mathrm{c}}}},\hat\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} }_{ \mathcal{\scriptscriptstyle A} })^{ \mathrm{\scriptscriptstyle T} }$  provides a sparse local minimizer for the nonconvex optimization \eqref{eq:est11} such that {\rm(i)} $|\hat{\boldsymbol{\theta}}-\boldsymbol{\theta}_0|_\infty=O_{ \mathrm{p} }(b_{1,n}^{1/2})$, {\rm(ii)} $|\hat{\boldsymbol{\xi}}_{{ \mathcal{\scriptscriptstyle A} }^{{\mathrm{c}}}}-\boldsymbol{\xi}_{0,{ \mathcal{\scriptscriptstyle A} }^{{\mathrm{c}}}}|_\infty=O_{ \mathrm{p} }(\phi_n) $, and {\rm(iii)} $\mathbb{P}(\hat{\boldsymbol{\xi}}_{{ \mathcal{\scriptscriptstyle A} }}={\mathbf 0})\to 1$ as $n\to\infty$.
\end{proposition}

 In the rest of this section, we focus on the sparse local minimizer $\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }}=(\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} },\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ specified in \eqref{eq:locest}. We proceed with additional regularity conditions.

\begin{as}\label{A.7}
There exists a universal constant $K_5>1$ such that $K_5^{-1}<\lambda_{\rm min}\{{\mathbf Q}_{{ \mathcal{\scriptscriptstyle I} }\cup{ \mathcal{\scriptscriptstyle B} }_1,{ \mathcal{\scriptscriptstyle B} }_2}^{{ \mathcal{\scriptscriptstyle (T)} }}\}\le\lambda_{\rm max}\{{\mathbf Q}_{{ \mathcal{\scriptscriptstyle I} }\cup{ \mathcal{\scriptscriptstyle B} }_1,{ \mathcal{\scriptscriptstyle B} }_2}^{{ \mathcal{\scriptscriptstyle (T)} }}\}<K_5$ for any $\mathcal{B}_2\subset \mathcal{B}_1 \subset \mathcal{D}$ with $p+|\mathcal{B}_2| \le r_1+|\mathcal{B}_1| \le \ell_n$,
where ${\mathbf Q}_{{ \mathcal{\scriptscriptstyle I} }\cup{ \mathcal{\scriptscriptstyle B} }_1,{ \mathcal{\scriptscriptstyle B} }_2}^{{ \mathcal{\scriptscriptstyle (T)} }}=([\mathbb{E}\{\nabla_{\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle P}}\cup{ \mathcal{\scriptscriptstyle B} }_2}}{\mathbf g}_{i,{ \mathcal{\scriptscriptstyle I} }\cup{ \mathcal{\scriptscriptstyle B} }_1}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi}_0)\}]^{ \mathrm{\scriptscriptstyle T} })^{\otimes2}$.
\end{as}
Condition \ref{A.7} is the sparse Riesz condition (\citealp{zhang2008sparsity}, \citealp{chen2008extended}) in our setting
to deal with the high-dimensional $\boldsymbol{\psi}$ when $p+r_2>n$. Let $\hat\boldsymbol{\lambda}(\hat\boldsymbol{\psi}_{ \mathrm{\scriptscriptstyle PEL} })$ be the $r\times 1$ vector of
Lagrange multiplier defined at $\hat\boldsymbol{\psi}_{ \mathrm{\scriptscriptstyle PEL} }$:
\begin{align*}
\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\psi}}_{ \mathrm{\scriptscriptstyle PEL} })=\arg\max_{\boldsymbol{\lambda}\in\hat{\Lambda}_{n}^{{ \mathcal{\scriptscriptstyle (T)} }}(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})}\bigg[\frac{1}{n}\sum_{i=1}^n\log\{1+\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\hat{\boldsymbol{\psi}}_{ \mathrm{\scriptscriptstyle PEL} })\}-\sum_{j\in\mathcal{D}}P_{2,\nu}(|\lambda_j|)\bigg]\,.
\end{align*}
Write $\hat{\boldsymbol{\lambda}}:=\hat\boldsymbol{\lambda}(\hat\boldsymbol{\psi}_{ \mathrm{\scriptscriptstyle PEL} })=(\hat{\lambda}_1,\ldots,\hat{\lambda}_r)^{ \mathrm{\scriptscriptstyle T} }$ and $\rho_2(t;\nu)=\nu^{-1}P_{2,\nu}(t)$. Let
$\mathcal{R}_n= \mathcal{I} \cup \{ j\in \mathcal{D}: \hat{\lambda}_j \neq 0 \} $
be the set of the estimated binding moments.
Since only those $\lambda_j$'s associated with $\mathcal{D}$ are penalized, it follows that
\[
\frac{1}{n}\sum_{i=1}^n\frac{g_{i,j}^{{ \mathcal{\scriptscriptstyle (T)} }}(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})}{1+\hat{\boldsymbol{\lambda}}_{{ \mathcal{\scriptscriptstyle R} }_n}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}_{i,{ \mathcal{\scriptscriptstyle R} }_n}^{{ \mathcal{\scriptscriptstyle (T)} }}(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})}=\left\{\begin{aligned}
0\,,~~~~~~~~~~~~~~&\textrm{if}~j\in \mathcal{I} \,,\\
\nu\rho_2'(|\hat{\lambda}_j|;\nu)\mbox{\rm sgn}(\hat\lambda_j) \,,~~~&\textrm{if}~ j\in\mathcal{D} ~\textrm{with}~\hat{\lambda}_j\neq 0\,.
\end{aligned} \right.
\]
For any unbinding moment $j\in\mathcal{R}_n^{{\mathrm{c}}}$,
define
\begin{equation}\label{eq:bod}
\hat{\eta}_j := \frac{1}{n}\sum_{i=1}^n\frac{g_{i,j}^{{ \mathcal{\scriptscriptstyle (T)} }}(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})}{1+\hat{\boldsymbol{\lambda}}_{{ \mathcal{\scriptscriptstyle R} }_n}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}_{i,{ \mathcal{\scriptscriptstyle R} }_n}^{{ \mathcal{\scriptscriptstyle (T)} }}(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})}\,.
\end{equation}
If $\hat{\lambda}_j=0$ for some $j\in\mathcal{D}$, the subdifferential at $\hat{\lambda}_j$ has to include the zero element \citep{bertsekas1997nonlinear}. That is, $\hat{\eta}_j\in[-\nu\rho_2'(0^+),\nu\rho_2'(0^+)]$ for any $j\in\mathcal{R}_n^{{\mathrm{c}}}$. In our theoretical analysis, we impose the next condition.
\begin{as}\label{AA1}
$
\mathbb{P}[\cup_{j\in\mathcal{R}_n^{{\mathrm{c}}}}\{|\hat{\eta}_j|=\nu\rho_2'(0^+)\}]\rightarrow0
$
as $n\rightarrow\infty$.
\end{as}

\begin{remark}
Condition \ref{AA1} requires that $\hat{\eta}_j$ $(j\in\mathcal{R}_n^{{\mathrm{c}}})$ does not lie on the boundary with probability approaching one, which is realistic in practice. If the distribution function of $\hat{\eta}_j$ is continuous at $\pm\nu\rho_2'(0^+)$,  we then have $\mathbb{P}\{|\hat{\eta}_j|=\nu\rho_2'(0^+)\}=0$. Condition \ref{AA1} makes sure that $\hat{\boldsymbol{\lambda}}(\boldsymbol{\psi})$ is continuously differentiable at $\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ with probability approaching one. See Lemma A.5 in the supplementary materials.
\end{remark}

Define $\mathcal{A}_*=\{ j \in \mathcal{A}:\hat{\lambda}_j\neq 0\}$ and
$\mathcal{A}_{*,{\mathrm{c}}}=\{ j\in \mathcal{A}^{{\mathrm{c}}}:\hat{\lambda}_j\neq 0\}$.
Let $\mathcal{I}^*=\mathcal{I} \cup\mathcal{A}_*$,
and then  $\mathcal{R}_n$ can be decomposed into two disjoint parts $\mathcal{I}^*$
and $\mathcal{A}_{*,{\mathrm{c}}}$.
Furthermore, we define $\mathcal{S}_*= \mathcal{P} \cup  \mathcal{A}_{*,{\mathrm{c}}}$.
The fact $\mathcal{S}_*\subset\mathcal{S}$ implies that $|\mathcal{S}_*|=p+|\mathcal{A}_{*,{\mathrm{c}}}|\le |\mathcal{S}|=s$.
For any $\boldsymbol{\psi}\in\boldsymbol{\Psi}$, we have $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }_*}=(\boldsymbol{\theta}^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} }_{{{ \mathcal{\scriptscriptstyle A} }}_{*,{\mathrm{c}}}})^{ \mathrm{\scriptscriptstyle T} }$. Define
\begin{align*}
{\mathbf J}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}&=\big([\mathbb{E}\{\nabla_{\boldsymbol{\theta}}{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_{i,{ \mathcal{\scriptscriptstyle I} }^*}(\boldsymbol{\theta}_0)\}]^{ \mathrm{\scriptscriptstyle T} }\{{\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}(\boldsymbol{\theta}_0)\}^{-1/2}\big)^{\otimes2}
\end{align*}
with ${\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}(\boldsymbol{\theta}_0)=\mathbb{E}\{{\mathbf g}^{ \mathcal{\scriptscriptstyle (T)} }_{i,{ \mathcal{\scriptscriptstyle I} }^*}(\boldsymbol{\theta}_0)^{\otimes2}\}$,
and
$\hat\boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n}=\{{\mathbf J}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle R} }_n}\}^{-1}[\mathbb{E}\{\nabla_{\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }_*}}{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_{i,{ \mathcal{\scriptscriptstyle R} }_n}(\boldsymbol{\psi}_0)\}]^{ \mathrm{\scriptscriptstyle T} }\{{\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle R} }_n}(\boldsymbol{\psi}_0)\}^{-1}\hat\boldsymbol{\eta}_{{ \mathcal{\scriptscriptstyle R} }_n}$,
where ${\mathbf J}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle R} }_n}=([\mathbb{E}\{\nabla_{\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }_*}}{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_{i,{ \mathcal{\scriptscriptstyle R} }_n}(\boldsymbol{\psi}_0)\}]^{ \mathrm{\scriptscriptstyle T} }\{{\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle R} }_n}(\boldsymbol{\psi}_0)\}^{-1/2})^{\otimes2}$,
$\hat\boldsymbol{\eta}=(\hat{\eta}_j)_{j\in \mathcal{T}}$ with $\hat\eta_j=0$
for $j\in \mathcal{I}$,
$\hat\eta_j=\nu\rho'_2(|\hat\lambda_j|;\nu){\rm sgn}(\hat\lambda_j)$ for $j\in \mathcal{D}$ with $\hat\lambda_j\neq0$, and $\hat\eta_j$ specified in \eqref{eq:bod} for $j\in\mathcal{R}_n^{{\mathrm{c}}}$.
The limiting distribution of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ is stated in Theorem \ref{thm2}.


\begin{theorem}\label{thm2}
Assume the conditions of Proposition {\rm\ref{thm1}} hold. Under Conditions {\rm \ref{A.7}} and {\rm\ref{AA1}}, if
$\ell_n^{3/2}\log r=o(n^{1/2-1/\gamma})$ and
$\ell_nn^{1/2}s^{3/2} \phi_n \nu=o(1)$,
then
$
n^{1/2}\boldsymbol{\alpha}^{ \mathrm{\scriptscriptstyle T} }\{{\mathbf J}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}\}^{1/2}\{\hat\boldsymbol{\theta}_{{{ \mathrm{\scriptscriptstyle PEL} }}}-\boldsymbol{\theta}_{0}
-\hat\boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n,(1)}\}\xrightarrow{d}\mathcal{N}(0,1)$
as $n\to\infty$ for any $\boldsymbol{\alpha}\in\mathbb{R}^p$ with $|\boldsymbol{\alpha}|_2=1$,  where $\hat\boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n,(1)}$ is the first $p$ components of $\hat\boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n}$ and $(a_n,\phi_n)$ are given in \eqref{eq:an}.
\end{theorem}

\begin{remark}\label{rek:3.4}
Notice that $a_n\lesssim s\pi$. To satisfy the restrictions in Theorem \ref{thm2},
it is sufficient for $(n,r,p,\ell_n,s)$ to have:  $r = o\{\exp(n^{1/3})\}$,
$\ell_n=o\{n^{(\gamma-2)/(3\gamma)}(\log r)^{-2/3}\}$,
$\ell_n^{5/2}s^{3/2}(\log r)\max\{p,\ell_n^{1/2}\}=o(n^{1/2})$,
the tuning parameters $\nu$ and $\pi$ satisfying
$
\ell_n^{1/2}\aleph_n \ll\pi\ll \min\{s^{-1}n^{-2/\gamma},\ell_n^{-4}s^{-4}p^{-2}(\log r)^{-1}\}$, $
\ell_n\aleph_n  \ll \nu \ll \min\{(\ell_n^3s^3p^2\log r)^{-1/2},n^{-1/\gamma},(\ell_n^2ns^3)^{-1/4} \}$
and $\nu^2\pi\ll(\ell_n^2ns^4p^2)^{-1}$.
\end{remark}





Thanks to the extra valid moments in ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}$, the asymptotic covariance of $\hat\boldsymbol{\theta}_{{ \mathrm{\scriptscriptstyle PEL} }}$ is smaller than that of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{ \mathcal{\scriptscriptstyle (I)} }$ which utilizes the information in ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}$ only.
On the other hand,
given that $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }}$ provides asymptotic normality under
the set of correctly specified moments $\mathcal{I}$,
one may question whether it is worthwhile to consider all the available moments among which some may risk misspecification.
This is a trade-off between robustness and efficiency,
a recurrent theme in modern econometrics.
Although it inevitably depends on the data generating process (DGP),
in big data environments the efficiency gain may be substantial.
Benefits are demonstrated in the simulation exercises and the empirical application via very simple econometric models.\footnote{
See the RMSEs in Table \ref{tab5}, the length of confidence intervals in Figure \ref{fig2}, and the
standard deviations in Table \ref{tab:r1}.
}




\begin{remark}
The setting with the sets $\mathcal{I}$ and $\mathcal{D}= \mathcal{A}\cup \mathcal{A}^{{\mathrm{c}}}$ is the same as \citet{liao2013adaptive}.
When estimating (\ref{eq:esteq2}) with low-dimensional moments based on GMM, \cite{ChengLiao2015} require the number of estimating functions $r=o(n^{1/3})$,
and the signal of invalid estimating functions
$\min_{k\in\mathcal{A}^{{\mathrm{c}}}}|\mathbb{E}\{g_{i,k}^{{ \mathcal{\scriptscriptstyle (D)} }}(\boldsymbol{\theta}_0)\}|\gg (rn^{-1})^{1/2}$.
Our conditions on the relative size of the dimensions are more general. A sufficient condition for \eqref{eq:signal} is $\min_{k\in\mathcal{A}^{{\mathrm{c}}}}|\mathbb{E}\{g_{i,k}^{{ \mathcal{\scriptscriptstyle (D)} }}(\boldsymbol{\theta}_0)\}|\gg \max\{\nu,p(s\pi)^{1/2},pr_1^{1/2}\aleph_n\}$,
in view of $a_n\lesssim s\pi$.
Under the restrictions discussed in Remark \ref{rek:3.4},
if we select $\nu$ and $\pi$ sufficiently close to $\ell_n\aleph_n$ and $ \ell_n^{1/2}\aleph_n$,
respectively, then \eqref{eq:signal} holds provided that $\min_{k\in\mathcal{A}^{{\mathrm{c}}}}|\mathbb{E}\{g_{i,k}^{{ \mathcal{\scriptscriptstyle (D)} }}(\boldsymbol{\theta}_0)\}|\gg ps^{1/2}\ell_n^{1/4}\aleph_n^{1/2}$. This is similar to the ``beta-min condition'' which is necessary for consistent variable selection
by shrinkage methods \citep{buhlmann2013statistical}.
\end{remark}







\begin{remark}\label{rmk_thm2}
Under low-dimensional moments the asymptotic normality involves no bias;
see \citet[Theorem 3.5]{liao2013adaptive} and \citet[Theorem 3.3]{ChengLiao2015}.
In contrast, Theorem \ref{thm2} here makes clear that
the high-dimensional moments incur an
additional asymptotic bias $\hat \boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n,(1)}$.
To conduct hypothesis testing or construct confidence region
about  $\boldsymbol{\theta}$,
in principle we can estimate and correct the bias $\hat{\boldsymbol{\zeta}}_{{ \mathcal{\scriptscriptstyle R} }_n,(1)}$.
Such a direct bias correction approach, nevertheless, is undesirable in practice and in theory.
The bias term involves multiplication and inverse of large matrices,
which are difficult to estimate with precision in finite samples.
Illustrated in our simulation experiments in Section \ref{se:numst}, we are unsatisfied with the asymptotic normality approximation after estimating and correcting the bias.
Thus, we view Theorem \ref{thm2} as a characterization of the asymptotic behavior of
$\hat\boldsymbol{\theta}_{{{ \mathrm{\scriptscriptstyle PEL} }}}$, but we do not encourage carrying out statistical inference based on it.
Instead, Section \ref{sec:infer} recommends PPEL for inference.
\end{remark}





While the augmented parameter $\boldsymbol{\xi}$ essentially validates all moments in $\mathcal{D}$,
the efficiency gain comes
from the penalty on $\boldsymbol{\xi}$ that shrinks some $\hat{\xi}_{k}$, $k \in \mathcal{A}$, to zero,  thereby confirming the validity of these associated moments.
To discuss moment selection, we write $\hat\boldsymbol{\xi}_{{ \mathrm{\scriptscriptstyle PEL} }}=(\hat{\xi}_k)_{k\in \mathcal{D}}$.
If all valid estimating functions in $\mathcal{A}$ are selected by the optimization \eqref{eq:est11}, i.e., $\hat{\xi}_k = 0$ for all $k\in \mathcal{A}$,
then the asymptotic covariance of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ is $\{{\mathbf J}^{{ \mathcal{\scriptscriptstyle (H)} }}\}^{-1}$.

Remind that $\{{\mathbf J}^{{ \mathcal{\scriptscriptstyle (H)} }}\}^{-1}$ in Proposition \ref{cy1} is the semiparametric efficiency bound for the estimation of $\boldsymbol{\theta}_0$ under the oracle, and
Proposition \ref{thm1} shows that $\mathbb{P}(\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle A} }}={\mathbf 0})\rightarrow1$ as $n\rightarrow\infty$, which provides a natural moment selection criterion
\begin{align}\label{eq:A}
\mathcal{\hat A}=\{k \in \mathcal{D}:\hat\xi_k=0\}\,
\end{align}
to identify the valid estimating functions in $\mathcal{A}$.
Notice that Proposition \ref{thm1} also indicates that $|\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle A} }^{{\mathrm{c}}}}-\boldsymbol{\xi}_{0,{ \mathcal{\scriptscriptstyle A} }^{{\mathrm{c}}}}|_\infty=O_{ \mathrm{p} }(\phi_n)$. Together with \eqref{eq:signal}, we have $\mathbb{P}[\cup_{k\in\mathcal{A}^{{\mathrm{c}}}}\{\hat{\xi}_k=0\}]\rightarrow0$ as $n\rightarrow\infty$.
Based on these arguments, Theorem \ref{thmoracle} supports our proposed moments selection criterion.

\begin{theorem}\label{thmoracle}
Under the conditions of Proposition {\rm\ref{thm1}}, it holds that $\mathbb{P}(\mathcal{\hat A}=\mathcal{A})\to1$ as $n\rightarrow\infty$.
\end{theorem}

Up to now, we have established the asymptotic properties of the PEL estimator for the moment-defined model (\ref{eq:esteq2}) with high-dimensional moments and low-dimensional $\boldsymbol{\theta}$.
The next section further extends the estimation to a high-dimensional structural parameter.








\section{High-dimensional moments with high-dimensional $\boldsymbol{\theta}$}\label{se:hpel}

A high-dimensional parameter $\boldsymbol{\theta}$ is present when many control variables are included in the structural model.
For example, the leading case of the linear IV model takes only one endogenous variable and it is accompanied by a few corresponding IVs. Nevertheless, such a simple setting may still involve many exogenous control variables in the main equation,
and these control variables are natural instruments for themselves.
An empirical example can be found in \citet{blundell1993we}.
\citet{fan2014endogeneity} propose the \emph{focused GMM} for such a linear IV model, which is a special case of the moment-defined model.

In this section,  $p$ and $r_2$ are both allowed to be much larger than $n$,
whereas $r_1$ is fixed or diverges at some slow polynomial rate of $n$.
We assume $\boldsymbol{\theta}_0$ sparse in the sense that most of its components are zero.
Write $\boldsymbol{\theta}_0=(\theta_{0,l})_{l \in \mathcal{P}} $
and define the active set $\mathcal{P}_{\sharp}=\{l \in \mathcal{P}:\theta_{0,l}\neq0\}$ with cardinality $p_{\sharp}=|\mathcal{P}_{\sharp}|$.
To obtain a sparse estimate of $\boldsymbol{\theta}_0$, we add to \eqref{eq:est11} a penalty on $\boldsymbol{\theta}$ to produce the following optimization problem:
\begin{equation}\label{eq:est22}
\begin{split}
    (\hat{\boldsymbol{\theta}}_{{{ \mathrm{\scriptscriptstyle PEL} }}}^{ \mathrm{\scriptscriptstyle T} },\hat{\boldsymbol{\xi}}_{{{ \mathrm{\scriptscriptstyle PEL} }}}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} } = \arg\min_{\boldsymbol{\psi}\in\boldsymbol{\Psi}}\max_{\boldsymbol{\lambda}\in\hat{\Lambda}^{{ \mathcal{\scriptscriptstyle (T)} }}_{n}(\boldsymbol{\psi})}\bigg[\frac{1}{n}\sum_{i=1}^n\log\{1+\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\boldsymbol{\psi})\}
    & -\sum_{j \in \mathcal{D} }  P_{2,\nu}(|\lambda_j|)\\
    +~  \sum_{k \in \mathcal{D} } P_{1,\pi}(|\xi_k|)
    & +\sum_{l \in \mathcal{P} } P_{1,\pi}(|\theta_l|)\bigg]\,.
\end{split}
    \end{equation}
With slight abuse of notation, we keep using $(\hat{\boldsymbol{\theta}}_{{{ \mathrm{\scriptscriptstyle PEL} }}}^{ \mathrm{\scriptscriptstyle T} },\hat{\boldsymbol{\xi}}_{{{ \mathrm{\scriptscriptstyle PEL} }}}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ to denote the solution to \eqref{eq:est22}.

As $p$ can be bigger than $r_1$ in high dimension,
here we update Condition \ref{A.1} for the identification of $\boldsymbol{\theta}_0$.
In view of \citet{ChangTangWu2018}, we impose Condition \ref{A.1h} below for the identification of the nonzero components of $\boldsymbol{\theta}_0$.



\addtocounter{as}{-6}
\begin{as}\label{A.1h}
(i) There exists a universal constant $K'_1>0$ such that
$$\inf_{\boldsymbol{\theta}=(\boldsymbol{\theta}_{{ \mathcal{\scriptscriptstyle P}}_\sharp}^{{ \mathrm{\scriptscriptstyle T} }},\boldsymbol{\theta}_{{ \mathcal{\scriptscriptstyle P}}_\sharp^{{\mathrm{c}}}}^{{ \mathrm{\scriptscriptstyle T} }})^{{ \mathrm{\scriptscriptstyle T} }}\in\boldsymbol{\Theta}:\,|\boldsymbol{\theta}_{{ \mathcal{\scriptscriptstyle P}}_\sharp}-\boldsymbol{\theta}_{0,{ \mathcal{\scriptscriptstyle P}}_\sharp}|_\infty>\varepsilon, \,\boldsymbol{\theta}_{{ \mathcal{\scriptscriptstyle P}}_\sharp^{{\mathrm{c}}}}={\mathbf 0}}|\mathbb{E}\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}_i(\boldsymbol{\theta})\}|_\infty\ge K'_1 \varepsilon$$ for any $\varepsilon>0$. (ii) There exists a universal constant $K_3'>0$ such that $\sup_{\boldsymbol{\theta}\in\boldsymbol{\Theta}}|\mathbb{E}\{\nabla_{\boldsymbol{\theta}_{{ \mathcal{\scriptscriptstyle P}}_\sharp^{{\mathrm{c}}}}}{\mathbf g}_i^{{ \mathcal{\scriptscriptstyle (I)} }}(\boldsymbol{\theta})\}|_\infty\leq K_3'$.
\end{as}


As we have shown in Section \ref{se:pel}, the index set $\mathcal{S}$ and the quantities $a_n$ and $\phi_n$ play key roles in the theoretical analysis
of the low-dimensional $\boldsymbol{\theta}$ in \eqref{eq:est11}. Under the current high-dimensional setting, we define
\begin{align}
a_n=\sum_{k\in\mathcal{D}}P_{1,\pi}(|\xi_{0,k}|)+\sum_{l\in\mathcal{P}}P_{1,\pi}(|\theta_{0,l}|)~~\textrm{and}~~\mathcal{S}=\mathcal{P}_{\sharp}\cup\mathcal{A}^{{\mathrm{c}}}~~\textrm{with}~~s=|\mathcal{S}|\,.\label{eq:annew}\end{align}
This $\mathcal{S}$ indexes all nonzero components in the true augmented parameter $\boldsymbol{\psi}_0$.
Compared to its counterpart $a_n$ in Section \ref{se:pel},
the newly defined $a_n$ here is amended with an extra term $\sum_{l\in\mathcal{P}}P_{1,\pi}(|\theta_{0,l}|)$ due to the penalty imposed on $\boldsymbol{\theta}$.
In the current high-dimensional setting, we also redefine
\begin{align}\label{eq:phinnew}
\phi_n=\max\{p_{\sharp} a_n^{1/2},p_{\sharp}r_1^{1/2}\aleph_n,\nu\}
\end{align}
with $a_n$ in \eqref{eq:annew}.
To control the bias introduced by $P_{1,\pi}(\cdot)$ on $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ and $\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$, similar to \eqref{eq:signal}, we assume that among the elements in $\boldsymbol{\psi}_0=(\psi_{0,1},\ldots,\psi_{0,p+r_2})^{ \mathrm{\scriptscriptstyle T} }$, the minimal active signal is
\begin{align}\label{eq:signal:2}
\min_{k\in\mathcal{S}}|\psi_{0,k}|\gg\max\{\phi_n,\pi\}\end{align}
with $\phi_n$ in \eqref{eq:phinnew} and $\max_{{k\in \mathcal{S}}}\sup_{c|\psi_{0,k}|<t<c^{-1}|\psi_{0,k}|}P'_{1,\pi}(t)=0$
for some constant $c\in(0,1)$.\footnote{See the arguments below \eqref{eq:xi} for the validity of these assumptions.}
Given the newly defined $\mathcal{S}$ and $\phi_n$, we can further update $\ell_n$,
 $\mathcal{R}_n$, $\mathcal{A}_*$, $\mathcal{A}_{*,{\mathrm{c}}}$ and $\mathcal{I}^*$ in the same manner as
their counterparts in Section \ref{se:pel}.
Moreover, we also define ${\mathbf J}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle R} }_n}$
and $\hat\boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n}$ as specified in Section \ref{se:pel}
with $\mathcal{S}_*=\mathcal{P}_{\sharp}\cup \mathcal{A}_{*,{\mathrm{c}}}$. For any $\boldsymbol{\psi}\in\boldsymbol{\Psi}$, we have $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }_*}=(\boldsymbol{\theta}_{{ \mathcal{\scriptscriptstyle P}}_{\sharp}}^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} }_{{{ \mathcal{\scriptscriptstyle A} }}_{*,{\mathrm{c}}}})^{ \mathrm{\scriptscriptstyle T} }$.

Proposition A.2 in the supplementary materials shows that there exists a sparse local minimizer $(\hat\boldsymbol{\theta}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} },\hat\boldsymbol{\xi}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }\in\boldsymbol{\Psi}$ for the nonconvex optimization \eqref{eq:est22} such that $\mathbb{P}(\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle P}}^{{\mathrm{c}}}_{\sharp}}={\mathbf 0})\to1$ as $n\to\infty$, which means all zero components of $\boldsymbol{\theta}_0$ can be estimated  exactly as zero with high probability. The limiting distribution of such $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle P}}_{\sharp}}$ and the moment selection outcomes are stated in Theorem \ref{thm3} below.



\begin{theorem}\label{thm3}
Let $P_{1,\pi}(\cdot), P_{2,\nu}(\cdot)\in\mathscr{P}$ for $\mathscr{P}$ defined in \eqref{eq:classp},
and $P_{2,\nu}(\cdot)$ be convex with bounded second derivative around $0$. For the sparse local minimizer $\hat\boldsymbol{\psi}_{{{ \mathrm{\scriptscriptstyle PEL} }}}=(\hat{\boldsymbol{\theta}}_{{{ \mathrm{\scriptscriptstyle PEL} }}}^{ \mathrm{\scriptscriptstyle T} },\hat{\boldsymbol{\xi}}_{{{ \mathrm{\scriptscriptstyle PEL} }}}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ for {\rm\eqref{eq:est22}} specified in Proposition {\rm A.2} in the supplementary materials, assume there exists a constant $\tilde{c}\in(C_*,1)$ such that $\mathbb{P}[\cup_{j\in\mathcal{T}}\{|\bar{g}_j^{ \mathcal{\scriptscriptstyle (T)} }(\hat\boldsymbol{\psi}_{{{ \mathrm{\scriptscriptstyle PEL} }}})|\in[\tilde{c}\nu\rho'_2(0^+),\nu\rho'_2(0^+))\}]\rightarrow0$.
Suppose Conditions {\rm\ref{A.1h}}, {\rm \ref{A.2}--\ref{A.3}} and \eqref{eq:signal:2}  hold.
Furthermore, assume Condition {\rm\ref{A.7}} holds with the newly defined $\ell_n$,
replacing $\mathcal{P}$ by
$\mathcal{P}_{\sharp}$ and replacing $p$ by $p_{\sharp}$, and Condition {\rm\ref{AA1}} holds with the newly defined $\mathcal{R}_n$.
If $\log r=o(n^{1/3})$,  $\max\{a_n,\nu^2\}=o(n^{-2/\gamma})$, $\ell_n^{3/2}\log r=o(n^{1/2-1/\gamma})$, $\ell_nn^{1/2}s^{3/2} \phi_n \nu=o(1)$ and $\ell_n\aleph_n=o(\min\{\nu,\pi\})$, then we have
$$
    n^{1/2}\boldsymbol{\alpha}^{ \mathrm{\scriptscriptstyle T} }\{{\mathbf W}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}\}^{1/2}\{\hat\boldsymbol{\theta}_{{{ \mathrm{\scriptscriptstyle PEL} }},{ \mathcal{\scriptscriptstyle P}}_{\sharp}}-\boldsymbol{\theta}_{0,{ \mathcal{\scriptscriptstyle P}}_{\sharp}}-\hat\boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n,(1)}\}\xrightarrow{d}\mathcal{N}(0,1)
$$
for any $\boldsymbol{\alpha}\in\mathbb{R}^{p_\sharp}$ with $|\boldsymbol{\alpha}|_2=1$ as $n\to\infty$, where $\hat\boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n,(1)}$ is the first $p_{\sharp}$ components of $\hat\boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n}$,  $(a_n,\phi_n)$ are given by \eqref{eq:annew} and \eqref{eq:phinnew}, respectively, and $
    {\mathbf W}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}=([\mathbb{E}\{\nabla_{\boldsymbol{\theta}_{{ \mathcal{\scriptscriptstyle P}}_{\sharp}}}{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_{i,{ \mathcal{\scriptscriptstyle I} }^*}(\boldsymbol{\theta}_0)\}]^{ \mathrm{\scriptscriptstyle T} }\{{\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}(\boldsymbol{\theta}_0)\}^{-1/2})^{\otimes2}$ with ${\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}(\boldsymbol{\theta}_0)=\mathbb{E}\{{\mathbf g}^{ \mathcal{\scriptscriptstyle (T)} }_{i,{ \mathcal{\scriptscriptstyle I} }^*}(\boldsymbol{\theta}_0)^{\otimes2}\}$.
Moreover, under the conditions of Proposition {\rm A.2} in the supplementary materials, it holds that $\mathbb{P}(\mathcal{\hat A}=\mathcal{A})\to1$ as $n\rightarrow\infty$, where $\hat{\mathcal{A}}$ is specified in \eqref{eq:A}.
\end{theorem}


\begin{remark}
Instead of assuming $ \ell_n^{1/2} \aleph_n =o(\pi)$ as in Theorem \ref{thm2}, here Theorem \ref{thm3} strengthens $\ell_n  \aleph_n=o(\pi)$.
As shown in Section A.5 of the supplementary materials, this stronger condition guarantees that $\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}$, the zero components of $\boldsymbol{\psi}_0$, can be shrunk to zero with probability approaching one in the high-dimensional $\boldsymbol{\theta}$ setting.


\end{remark}



Compared to \eqref{eq:est11}, a penalty on $\boldsymbol{\theta}$ is added to \eqref{eq:est22}.
To appreciate the benefit from this additional penalty on the sparse parameter,
without loss of generality, we write $\boldsymbol{\theta}_{0}=(\boldsymbol{\theta}_{0,{ \mathcal{\scriptscriptstyle P}}_{\sharp}}^{ \mathrm{\scriptscriptstyle T} },{\mathbf 0}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ and
the inverse of ${\mathbf J}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}$ in Theorem \ref{thm2} in the following compatible partitioned matrix
\begin{align*}
\{{\mathbf J}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}\}^{-1}=\left(\begin{array}{cc}
~ [\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{11}    &  [\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{12} \\  ~[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{21}   &  [\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{22}
\end{array}\right)\,,
\end{align*}
where $[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{11}$ is a $p_{\sharp}\times p_{\sharp}$ matrix.
The asymptotic covariance of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle P}}_{\sharp}}$ by \eqref{eq:est11} is $[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{11}$,
and that in \eqref{eq:est22} is $\{{\mathbf W}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}\}^{-1}$ according to Theorem \ref{thm3}.
Notice that
$\{{\mathbf W}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}\}^{-1}=[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{11} -[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{12}[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{22}^{-1}[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{21} \le [\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{11}$
by the inverse of a partitioned matrix.
Implicitly restricting the parameter space, the penalty on  $\boldsymbol{\theta}$ gains efficiency for the estimation of $\boldsymbol{\theta}_{0,{ \mathcal{\scriptscriptstyle P}}_{\sharp}}$, the nonzero components of $\boldsymbol{\theta}_0$.


\begin{remark}

Theorem \ref{thm3} spells out the limiting distribution and efficiency gain for the estimator of the nonzero components in $\boldsymbol{\theta}_0$.
If one is interested in some coefficient in $\mathcal{P}_{\sharp}^{{\mathrm{c}}}$,
we have consistency in that
$\mathbb{P}(\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle P}}^{{\mathrm{c}}}_{\sharp}}={\mathbf 0})\to1$ as $n\to\infty$, but the limiting distribution
is irregular due to the shrinkage.\footnote{This is a generic property shared by procedures of the oracle properties, for example
SCAD \citep{FanLi2001} and the adaptive Lasso \citep{zou2006adaptive}.
}
Furthermore, parallel to Theorem \ref{thm2} the asymptotic bias is again present.
Similar to Remark \ref{rmk_thm2}, we do not suggest
conducting statistical inference predicated on this characterization of the asymptotic behavior of normality.
\end{remark}



Statistical inference is important in applied econometrics when researchers intend to assess
whether the estimated result supports or rejects a hypothesized value of the parameter.
The next section proposes an inferential procedure based on a projection of estimating functions, which is free of asymptotic biases.

\section{Confidence regions for a subset of $\boldsymbol{\psi}$}
\label{sec:infer}

Suppose we are interested in the inference for a subset of parameters $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$ for some generic small subset
$\mathcal{M} \subset (\mathcal{P} \cup \mathcal{D}) $ with $|\mathcal{M}|=m$.
Here we allow $m$ to be fixed or diverge slowly with the sample size $n$.
What is novel here is that $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$ is allowed to contain part of the auxiliary parameter $\boldsymbol{\xi}$,
for which our method will provide a formal statistical inference for the validity of a subset of the
high-dimensional moment restrictions.
In contrast, \citet{liao2013adaptive} offers asymptotic normality for the structural parameter $\boldsymbol{\theta}$ under low-dimensional moments but does not characterize the asymptotic distribution of the auxiliary parameter $\boldsymbol{\xi}$.

A key technical issue is how to deal with the high-dimensional nuisance parameter $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{\mathrm{c}}}$,
where $\mathcal{M}^{{\mathrm{c}}}= (\mathcal{P} \cup \mathcal{D}) \backslash \mathcal{M}$.
Our approach follows \citet{Chang2020} to project out the influence of the
nuisance parameter. Such idea was also used in \citet{ning2017general} and \citet{neykov2018unified}.
Given an initial estimate $\boldsymbol{\psi}^*$ for $\boldsymbol{\psi}_0$, we first determine a linear transformation matrix ${\mathbf A}_n = ({\mathbf a}_k^n)_{k\in\mathcal{M}}^{{ \mathrm{\scriptscriptstyle T} }} \in \mathbb{R}^{m\times r}$ with each row defined as
\begin{align}\label{eq:optim}
{\mathbf a}_k^n = \arg\min_{{\mathbf u}\in\mathbb{R}^r}|{\mathbf u}|_1\quad \mathrm{s.t.} \quad |\{\nabla_{\boldsymbol{\psi}}\bar{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi}^*)\}^{ \mathrm{\scriptscriptstyle T} } {\mathbf u} - \boldsymbol{\gamma}_k|_\infty \le \varsigma
\end{align}
for $k \in \mathcal{M}$,
where $\varsigma\rightarrow0$ as $n\rightarrow\infty$ is a tuning parameter,
and $\{\boldsymbol{\gamma}_k\}_{k\in \mathcal{M}}$ is a basis of the linear space
$\{
{\mathbf b}=(b_j)_{j\in \mathcal{P}\cup\mathcal{D}  }:
{\mathbf b}_{\mathcal{M}^{{\mathrm{c}}}} = \mathbf{0}
\}$.
In practice, we can specify the $(p+r_2)$-dimensional vector
$\boldsymbol{\gamma}_k = ( 1(j=k))_{j \in \mathcal{P} \cup \mathcal{D} }$
 with all zero elements except one unit entry.
As we will discuss in Remark \ref{initial_psi},
the solution from \eqref{eq:est11} or \eqref{eq:est22} can serve as the initial estimator $\boldsymbol{\psi}^*$.

Based on ${\mathbf A}_n$ as in (\ref{eq:optim}), we then obtain the new $m$-dimensional estimating functions
 ${\mathbf f}^{{\mathbf A}_n}(\cdot;\cdot)$ by projecting ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}(\cdot;\cdot)$ on ${\mathbf A}_n$:
\begin{align*}
{\mathbf f}^{{\mathbf A}_n}({\mathbf X};\boldsymbol{\psi})  = {\mathbf A}_n {\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}({\mathbf X};\boldsymbol{\psi}) \,.
\end{align*}
Write $\boldsymbol{\gamma}_k=(\gamma_{k,j})_{j\in \mathcal{P}\cup\mathcal{D}}$ and
define an $m\times(p+r_2)$ matrix $\boldsymbol{\Gamma}=(\gamma_{k,j})_{k\in { \mathcal{\scriptscriptstyle M} },
\, j \in \mathcal{P} \cup \mathcal{D}}$.
The definition of ${\mathbf A}_n$ implies that $|\nabla_{\boldsymbol{\psi}}\bar{{\mathbf f}}^{{\mathbf A}_n}(\boldsymbol{\psi}^*)-\boldsymbol{\Gamma}|_\infty\leq \varsigma$.
Since all the components in the $j$-th column of $\boldsymbol{\Gamma}$ are zero
for $j\in\mathcal{M}^{{\mathrm{c}}}$, the newly defined estimating functions ${\mathbf f}^{{\mathbf A}_n}$ is uninformative of the nuisance parameter
 $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}$.  Informative is  ${\mathbf f}^{{\mathbf A}_n}$ of the parameter $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$ of interest,
as for $j\in\mathcal{M}$ some components in the $j$-th column of $\boldsymbol{\Gamma}$ must be nonzero.

Given the projected estimating functions ${\mathbf f}^{{\mathbf A}_n}$,
it is possible to directly borrow from \citet{Chang2020} to construct the confidence region of $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$ by the EL ratio
\[
w_n(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }})=2\max_{\boldsymbol{\lambda}\in\tilde{\Lambda}_n(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }})}\sum_{i=1}^n\log\{1+\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf f}_i^{{\mathbf A}_n}(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }},\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^*)\}
\]
with respect to $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$, where $\tilde{\Lambda}_n(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }})=\{\boldsymbol{\lambda}  \in\mathbb{R}^m : \boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} } {\mathbf f}^{{\mathbf A}_n}_i(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }},\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^*) \in \mathcal{V}~\textrm{for any}~i\in[n]\}$.
Since $w_n(\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle M} }}) \xrightarrow{d}\chi_m^2$ as $n\rightarrow\infty$ for a fixed $m$,
the set $\{\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}\in\mathbb{R}^m:w_n(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }})\leq \chi_{m,1-\alpha}^2\}$ provides a $100(1-\alpha)\%$ confidence region  for $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$, where $\chi_{m,1-\alpha}^2$ is the $(1-\alpha)$-quantile of $\chi_m^2$ distribution.
Nonetheless, the finite-sample performance  of such an asymptotically valid confidence region depends crucially on the convexity of $w_n(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }})$ near $\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle M} }}$.
If convexity fails in a finite sample, this EL-ratio-based confidence region will be
voluminous in magnitude due to numerical instability.
Verifying the convexity condition is onerous, in particular when $m$ is large, for ${\mathbf f}_i^{{\mathbf A}_n}(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }},\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^*)$ may be a nonlinear function of $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$.

To secure a stable confidence region for $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$, in this paper we deviate from \citet{Chang2020} and instead
recommend a re-estimation procedure for a confidence region based on the asymptotic normality of the estimator.
Let
\begin{align*}
\tilde{\boldsymbol{\psi}}_{{ \mathcal{\scriptscriptstyle M} }} = \arg\min_{\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}\in\boldsymbol{\Psi}^*_{{ \mathcal{\scriptscriptstyle M} }}}\max_{\boldsymbol{\lambda}\in\tilde\Lambda_n(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }})}\sum_{i=1}^n\log\{1+\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} } {\mathbf f}^{{\mathbf A}_n}_i(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }},\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^*)\}\,,
\end{align*}
where $\boldsymbol{\Psi}^*_{ \mathcal{\scriptscriptstyle M} } = \{\boldsymbol{\psi}_{ \mathcal{\scriptscriptstyle M} }: |\boldsymbol{\psi}_{ \mathcal{\scriptscriptstyle M} } - \boldsymbol{\psi}^*_{ \mathcal{\scriptscriptstyle M} } |_1 \le O_{ \mathrm{p} }(\varpi_{1,n}) \}$ for some $\varpi_{1,n}\rightarrow0$ such that $|\boldsymbol{\psi}^*_{ \mathcal{\scriptscriptstyle M} }-\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle M} }}|_1 = O_{ \mathrm{p} }(\varpi_{1,n})$.
Since $\boldsymbol{\psi}_{ \mathcal{\scriptscriptstyle M} }^*$ is an initial consistent estimator of $\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle M} }}$,
it is reasonable to search for $\tilde{\boldsymbol{\psi}}_{{ \mathcal{\scriptscriptstyle M} }}$
in a small neighborhood of $\boldsymbol{\psi}^*_{ \mathcal{\scriptscriptstyle M} }$.
We will specify $\varpi_{1,n}$ for some specific choices of $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}^*$ in Remark \ref{initial_psi}.

To derive the limiting distribution of
$\tilde\boldsymbol{\psi}_{ \mathcal{\scriptscriptstyle M} }$, we assume the following condition.

\addtocounter{as}{5}
\begin{as}\label{A.10}
For each $k\in \mathcal{M}$, there is a non-random ${\mathbf a}_k$ satisfying $[\mathbb{E}\{\nabla_{\boldsymbol{\psi}}{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\boldsymbol{\psi}_0)\}]^{ \mathrm{\scriptscriptstyle T} } {\mathbf a}_k = \boldsymbol{\gamma}_k$, $| {\mathbf a}_k |_1\le K_6$  for some universal constant $K_6 > 0$, and $ \max_{k\in { \mathcal{\scriptscriptstyle M} } }|{\mathbf a}_k^n - {\mathbf a}_k|_1 = O_{ \mathrm{p} }(\omega_n)$ for some $\omega_n \to 0$. Let ${\mathbf A}=({\mathbf a}_k)_{k\in\mathcal{M}}^{{ \mathrm{\scriptscriptstyle T} }} \in \mathbb{R}^{m\times r}$. The eigenvalues of ${\mathbf A}^{\otimes2}$ are uniformly bounded away from zero and infinity.
\end{as}

\begin{remark}
Let $\boldsymbol{\Xi}=\mathbb{E}\{\nabla_{\boldsymbol{\psi}}{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\boldsymbol{\psi}_0)\}$ and $\widehat{\boldsymbol{\Xi}}=\nabla_{\boldsymbol{\psi}}\bar{{\mathbf g}}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi}^*)$. It follows from the existence of ${\mathbf a}_k$ that
$\boldsymbol{\gamma}_k=\widehat{\boldsymbol{\Xi}}^{ \mathrm{\scriptscriptstyle T} }{\mathbf a}_k+(\boldsymbol{\Xi}-\widehat{\boldsymbol{\Xi}})^{ \mathrm{\scriptscriptstyle T} }{\mathbf a}_k=\widehat{\boldsymbol{\Xi}}^{ \mathrm{\scriptscriptstyle T} }{\mathbf a}_k+\boldsymbol{\varepsilon}_k$,
where $\boldsymbol{\varepsilon}_k = (\boldsymbol{\Xi}-\widehat{\boldsymbol{\Xi}})^{ \mathrm{\scriptscriptstyle T} }{\mathbf a}_k$.
Some mild conditions ensure $|\widehat{\boldsymbol{\Xi}}-\boldsymbol{\Xi}|_\infty=o_{ \mathrm{p} }(1)$. This, together with the assumption $|{\mathbf a}_k|_1\leq K_6$, implies that $\boldsymbol{\varepsilon}_k$ is stochastically small uniformly over all the components such that $|\boldsymbol{\varepsilon}_k|_\infty=o_{ \mathrm{p} }(1)$,
which  can be viewed as an attempt to recover a nonrandom ${\mathbf a}_k$ with no noise asymptotically (\citealp{CandesTao2007}, \citealp{Bickeletal2009}).
It follows that $|{\mathbf a}_k^n-{\mathbf a}_k|_1=o_{ \mathrm{p} }(1)$
if $\widehat{\boldsymbol{\Xi}}$ satisfies the routine conditions for sparse signal recovering.
Furthermore, the constant $K_6$ in Condition \ref{A.10} may be replaced by some diverging $\varphi_n$ and our main results remain valid.
\end{remark}

Theorem \ref{thm:clt} gives the limiting distribution of $\tilde{\boldsymbol{\psi}}_{{ \mathcal{\scriptscriptstyle M} }}$
with a generic initial estimator $\boldsymbol{\psi}^*$.


\begin{theorem}\label{thm:clt}
Let $|\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}^* - \boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle M} }}|_1 = O_{ \mathrm{p} }(\varpi_{1,n})$, $|\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^* - \boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}|_1 = O_{ \mathrm{p} }(\varpi_{2,n})$ for some $\varpi_{1,n}\to0$ and $\varpi_{2,n}\to0$. Under Conditions {\rm\ref{A.2}}--{\rm\ref{A.3}} and {\rm\ref{A.10}}, if $m=o(n^{\min\{1/5,(\gamma-2)/(3\gamma)\}})$,
$m\omega_n^2(m^2+\log r)=o(1)$,
$m\varpi_{1,n}=o(1)$ and $n m \varpi_{2,n}^2 (\varsigma^2 + \varpi_{1,n}^2 + \varpi_{2,n}^2) = o(1)$, then
$
n^{1/2}\boldsymbol{\alpha}^{ \mathrm{\scriptscriptstyle T} }(\hat{\mathbf J}^*)^{1/2}(\tilde{\boldsymbol{\psi}}_{{ \mathcal{\scriptscriptstyle M} }} - \boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle M} }})\xrightarrow{d} \mathcal{N}(0,1)$ as $n\to\infty$
for any $\boldsymbol{\alpha} \in \mathbb{R}^m$ with $| \boldsymbol{\alpha} |_2 = 1$,  where
$
\hat{\mathbf J}^* = [ \{ \nabla_{\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}} \bar{{\mathbf f}}^{{\mathbf A}_n}(\tilde\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }},\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^*) \}^{ \mathrm{\scriptscriptstyle T} } \{ \widehat{{\mathbf V}}_{{\mathbf f}^{{\mathbf A}_n}}(\tilde\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }},\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^*) \}^{-1/2} ]^{\otimes2}
$
with $\widehat{{\mathbf V}}_{{\mathbf f}^{{\mathbf A}_n}}(\tilde\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }},\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^*)
= \mathbb{E}_n\{{\mathbf f}^{{\mathbf A}_n}_i(\tilde\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }},\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^*)^{\otimes2}\}$.
\end{theorem}


The above theorem is stated for $\tilde{\boldsymbol{\psi}}_{{ \mathcal{\scriptscriptstyle M} }} $. It includes the estimation of $\boldsymbol{\theta}_{{ \mathcal{\scriptscriptstyle M} }}$ as a special case if one's research interest falls on the main parameter for economic interpretation, and makes it possible to infer the validity of a subset of moments in view of selecting $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}=\boldsymbol{\xi}_{{ \mathcal{\scriptscriptstyle M} }}$.


\begin{remark}\label{initial_psi}
We verify that the PEL estimator is qualified to serve as $\boldsymbol{\psi}^*$ in Theorem \ref{thm:clt}.
Define $\bar{s}=|\mathcal{S}\cap\mathcal{M}|$, and thus $|\mathcal{S}\cap\mathcal{M}^{{\mathrm{c}}}|=s-\bar{s}$.
If $\boldsymbol{\theta}_0$ is low-dimensional and we choose $\boldsymbol{\psi}^*=\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ in \eqref{eq:est11},
we have $\varpi_{1,n}=\bar{s}\phi_n$ and $\varpi_{2,n}=(s-\bar{s})\phi_n$
due to $|\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle S} }}-\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle S} }}|_\infty=O_{ \mathrm{p} }(\phi_n)$ and $\mathbb{P}(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}={\mathbf 0})\rightarrow1$.
Then $ms\phi_n=o(1)$ and $nms^2\phi_n^{2}(\varsigma^2+s^2\phi_n^{2})=o(1)$ are sufficient for the restrictions imposed on $\varpi_{1,n}$ and $\varpi_{2,n}$.
Given $s$ and $\phi_n$ in Theorem \ref{thm2}, if $m$, $\omega_n$ and $\varsigma$ satisfy  $m=o(n^{\min\{1/5,(\gamma-2)/(3\gamma)\}})$, $m\omega_n^2(m^2+\log r)=o(1)$, $ms\phi_n=o(1)$ and  $nms^2\phi_n^{2}(\varsigma^2+s^2\phi_n^{2})=o(1)$, the asymptotic normality of Theorem \ref{thm:clt} holds
under this choice of the initial value $\boldsymbol{\psi}^* = \hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$.
Analogously, if  $\boldsymbol{\psi}^*=\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ in \eqref{eq:est22} when $\boldsymbol{\theta}_0$ is of high dimension, Theorem \ref{thm:clt} also holds provided that $m$, $\omega_n$ and $\varsigma$ satisfy the same restrictions with the newly defined $s$ and $\phi_n$  in Section \ref{se:hpel}.
\end{remark}



So far, we have established statistical inference results based on the PPEL estimator $\tilde{\boldsymbol{\psi}}_{{ \mathcal{\scriptscriptstyle M} }}$ when we are interested in a small subset
$\mathcal{M}$ of $\boldsymbol{\psi}$.
In the next section, we check the finite sample performance via simulations.


\section{Numerical studies}\label{se:numst}


One of the most important models in econometrics is the linear IV regression.
We design a linear IV model here to mimic the empirical application in Section \ref{emp_app}, whereas
simulation results of a nonlinear  panel regression with time-varying individual heterogeneity are presented in the supplementary materials.

Suppose that the researcher has at hand a dataset
of $n$ independent observations of a vector $\left(y_{i},x_{i},{\mathbf z}_{i},w_{1i},{\mathbf w}_{2i},{\mathbf w}_{3i}\right)$, and is interested in estimating the main structural equation
\begin{equation}
y_{i}=\beta_{x}x_{i}+\boldsymbol{\beta}_{{\mathbf z}}^{ \mathrm{\scriptscriptstyle T} }{\mathbf z}_{i}+\epsilon_{i}\,,
\label{eq:main}
\end{equation}
where $x_{i}$ is a scalar endogenous variable, and
${\mathbf z}_{i}$ is a $d_{z}\times1$ vector of exogenous variables (including the intercept).
Such an equation with a scalar endogenous variable is the leading case of IV regressions \citep{andrews2019weak}.
Due to space limitations, we report the results when we specify
${\mathbf z}_i=(1,z_{1,i},z_{2,i})^{ \mathrm{\scriptscriptstyle T} }\in\mathbb{R}^3$ with $( z_{1,i}, z_{2,i} )^{ \mathrm{\scriptscriptstyle T} } \sim \mathcal{N}({\mathbf 0},\rm{{\mathbf I}}_2)$, and
the coefficients $ (\beta_x, \boldsymbol{\beta}^{ \mathrm{\scriptscriptstyle T} }_{\mathbf z})^{ \mathrm{\scriptscriptstyle T} }=(0.5,0.5,0.5,0.5)^{ \mathrm{\scriptscriptstyle T} }$ in a plausible setting.
The numerical performance are robust across our experiments when the parameters are varied.

The corresponding reduced-form equation is
$
x_{i}=\gamma_{w1}w_{1i}+\boldsymbol{\gamma}^{ \mathrm{\scriptscriptstyle T} }_{{\mathbf w}2}{\mathbf w}_{2i}+\boldsymbol{\gamma}_{{\mathbf z}}^{ \mathrm{\scriptscriptstyle T} }{\mathbf z}_{i}+u_{i}$,
where $w_{1i}\in\mathbb{R}$ and ${\mathbf w}_{2i}\in\mathbb{R}^{d_{{\mathbf w}2}}$ are excluded instruments.
We specify
$ (w_{1i}, {\mathbf w}_{2i}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} } \sim \mathcal{N}({\mathbf 0},{\rm\bf I}_{d_{{\mathbf w}2}+1})$,
$ \gamma_{w1}=0.8$, $\boldsymbol{\gamma}_{{\mathbf w}2}=(\gamma_{{\mathbf w}2,j})_{j\in[d_{{\mathbf w}2}]}$ with $\gamma_{{\mathbf w}2,j}=0.4-0.3(j-1)/(d_{{\mathbf w}2}-1)$,
and $\boldsymbol{\gamma}_{\mathbf z}=(0.8,0.8,0.8)^{ \mathrm{\scriptscriptstyle T} }$.
The two error terms $(\epsilon_i,u_i)^{{ \mathrm{\scriptscriptstyle T} }}$ are generated from the two-dimensional normal distribution with mean $0$, covariance $1$ and correlation $0.5$ which are independent of $\left({\mathbf z}_{i},w_{1i},{\mathbf w}_{2i}\right)$.
An additional vector ${\mathbf w}_{3i}=(w_{3i,j})_{j\in[s]}\in\mathbb{R}^{s}$, which serves as the invalid IV, follows
$
w_{3i,j}=\delta_j\epsilon_{i}+v_{i}$,
where $v_{i}\sim \mathcal{N}\left(0,1\right)$ is
independent of $\left(\epsilon_{i},u_{i}\right)$, and $\delta_j \neq 0$ controls the strength of correlation.
Write $\boldsymbol{\delta}=(\delta_j)_{j\in[s]}$. To emulate the scenarios of \emph{weak}, \emph{moderate} and \emph{strong} correlations between $\epsilon_i$ and ${\mathbf w}_{3i}$, we set $\delta_j=0.3+0.2(j-1)/(s-1)$, $0.5+0.2(j-1)/(s-1)$ and $0.7+0.2(j-1)/(s-1)$, respectively.
It is expected that the smaller is the coefficient, the less severe is misspecification, so estimation is more prone to selection mistakes
in finite samples.

While the researcher is confident about the validity of
$w_{1i}$ in that
$
{\mathbf g}_i^{{ \mathcal{\scriptscriptstyle (I)} }} (\boldsymbol{\theta})=
(w_{1i},  {\mathbf z}_{i}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }
\times \left(y_{i}-\beta_{x}x_{i}-\boldsymbol{\beta}_{{\mathbf z}}^{ \mathrm{\scriptscriptstyle T} }{\mathbf z}_{i}\right),
$
where ${\mathbf z}_i$ is self-instrumented,
she is uncertain about the validity of $({\mathbf w}_{2i}, {\mathbf w}_{3i})$ and
therefore must detect the invalid IVs.
The two classes of undetermined (to the researcher) IVs consist of
$
{\mathbf g}_i^{{ \mathcal{\scriptscriptstyle (D)} }} (\boldsymbol{\theta})=
({\mathbf w}_{2i}^{ \mathrm{\scriptscriptstyle T} }, {\mathbf w}_{3i}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }
\times \left(y_{i}-\beta_{x}x_{i}-\boldsymbol{\beta}_{{\mathbf z}}^{ \mathrm{\scriptscriptstyle T} }{\mathbf z}_{i}\right).
$
According to our DGP design, ${\mathbf w}_{2i}$ is the valid IV vector so the moments
involving  ${\mathbf w}_{2i}$ equal to zero under the true parameter;
those moments involving ${\mathbf w}_{3i}$ are invalid.


Let $d_w =d_{{\mathbf w}2}+s+1$ be the number of all the IVs, among which
$s$ IVs are invalid. In the low-dimensional setting, we fix $s=6$ and
consider $(n,d_w)=(100,50)$ and $(200,100)$.
In the high-dimensional setting, we vary the number of invalid IVs to be
$s\in\{6,8,13\}$ for $(n,d_w)=(100,120)$, and $s\in\{6,12,17\}$ for $(n,d_w)=(200,240)$, where $s$ is specified by rounding, in addition, $2\log n$ and $3 n^{1/3}$.

In terms of the numerical implementation, we compute the PEL estimates by  the
\emph{modified two-layer coordinate descent algorithm} \citep{ChangTangWu2018}.
The SCAD penalty is used for both $P_{1,\pi}(\cdot)$ and $P_{2,\nu}(\cdot)$ in \eqref{eq:est11} for all the numerical experiments in this paper with local quadratic approximation \citep{FanLi2001}, and the tuning parameters $\nu$ and $\pi$ are chosen by the Bayesian information criterion (BIC).
Specifically, we use the following BIC type function:
\begin{align}\label{eq:BIC}
\mathrm{BIC} = \ell(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }}) + (\log n) \cdot \mathrm{df}(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})\,,
\end{align}
where $\mathrm{df}(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})$ denotes the number of nonzero elements in $\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$, and the log likelihood term is the EL ratio
$
\ell(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})=2\sum_{i=1}^n\log\{1+\hat\boldsymbol{\lambda}(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})\}$.\footnote{\cite{leng2012penalized} employ this BIC to choose the tuning parameter for the high-dimensional parameter $\boldsymbol{\psi}$, and \cite{ChangTangWu2018} apply it to select multiple tuning parameters. We follow their practice as \eqref{eq:BIC} remains valid in our setting where two tuning parameters are included. }






\begin{table}[ht]
 \small
\centering \caption{ PEL's performance in moment selection} \label{tab1}
\begin{spacing}{1.4}
\begin{tabular}{llllllllll}  \hline
               & Correlation: & \multicolumn{2}{c}{weak}                        &                      & \multicolumn{2}{c}{moderate}                    &                      & \multicolumn{2}{c}{strong}                      \\
$(n, d_w, s)$     & Method                   & \multicolumn{1}{c}{FP} & \multicolumn{1}{c}{FN} & \multicolumn{1}{c}{} & \multicolumn{1}{c}{FP} & \multicolumn{1}{c}{FN} & \multicolumn{1}{c}{} & \multicolumn{1}{c}{FP} & \multicolumn{1}{c}{FN} \\ \hline
\multicolumn{10}{c}{Panel A: low-dimensional setting}                                                                                                                                                                                         \\
(100, 50, 6)   & PEL                      & 0.1995                 & 0.0267                 &                      & 0.1933                 & 0.0003                 &                      & 0.2359                 & 0.0000                 \\
               & DB-PEL                   & 0.1983                 & 0.0273                 &                      & 0.1923                 & 0.0003                 &                      & 0.2343                 & 0.0000                 \\
(200, 100, 6)  & PEL                      & 0.0942                 & 0.0033                 &                      & 0.1012                 & 0.0000                 &                      & 0.0971                 & 0.0000                 \\
         & DB-PEL                   & 0.0930                 & 0.0033                 &                      & 0.1003                 & 0.0000                 &                      & 0.0962                 & 0.0000                 \\ \hline
\multicolumn{10}{c}{Panel B: high-dimensional setting}                                                                                                                                                                                        \\
(100, 120, 6)  & PEL                      & 0.1067                 & 0.0550                 &                      & 0.0850                 & 0.0047                 &                      & 0.1141                 & 0.0013                 \\
               & DB-PEL                   & 0.1061                 & 0.0580                 &                      & 0.0845                 & 0.0050                 &                      & 0.1131                 & 0.0013                 \\
(200, 240, 6)  & PEL                      & 0.0478                 & 0.0073                 &                      & 0.0476                 & 0.0007                 &                      & 0.0444                 & 0.0007                 \\
         & DB-PEL                   & 0.0468                 & 0.0093                 &                      & 0.0464                 & 0.0007                 &                      & 0.0435                 & 0.0007                 \\ \hline
(100, 120, 8)  & PEL                      & 0.1009                 & 0.0675                 &                      & 0.0872                 & 0.0063                 &                      & 0.1087                 & 0.0010                 \\
               & DB-PEL                   & 0.1004                 & 0.0693                 &                      & 0.0863                 & 0.0065                 &                      & 0.1074                 & 0.0015                 \\
(200, 240, 12) & PEL                      & 0.0498                 & 0.0072                 &                      & 0.0497                 & 0.0003                 &                      & 0.0494                 & 0.0003                 \\
               & DB-PEL                   & 0.0486                 & 0.0087                 &                      & 0.0488                 & 0.0003                 &                      & 0.0485                 & 0.0003                 \\ \hline
(100, 120, 13) & PEL                      & 0.0709                 & 0.0892                 &                      & 0.0558                 & 0.0168                 &                      & 0.1288                 & 0.0042                 \\
               & DB-PEL                   & 0.0706                 & 0.0905                 &                      & 0.0554                 & 0.0175                 &                      & 0.1268                 & 0.0043                 \\
(200, 240, 17) & PEL                      & 0.0402                 & 0.0086                 &                      & 0.0469                 & 0.0002                 &                      & 0.0456                 & 0.0000                 \\         & DB-PEL                   & 0.0394                 & 0.0116                 &                      & 0.0458                 & 0.0002                 &                      & 0.0447                 & 0.0000                \\ \hline
\end{tabular}
\end{spacing}
\end{table}


We first report results of moment selection by our selection criterion in \eqref{eq:A}. Let ``FP'' (false positive) denote the frequency that the valid moments being not selected, and ``FN'' (false negative) denote the frequency that the invalid moments being selected.
In Table \ref{tab1}, PEL and DB-PEL denote the moment selection criterion based on the PEL estimator and its de-biased version, respectively.
As expected, the strength of correlations between $\epsilon_i$ and ${\mathbf w}_{3i}$ does not affect FP,
whereas FN quickly vanishes as the correlations get stronger.
In all cases, larger sample sizes help reduce the chance of moment selection error.



\begin{table}[htbp]
\setlength\tabcolsep{3pt}
\small
\centering
\caption{Point estimations for $\beta_x$}\label{tab5}
\begin{spacing}{1.2}
\begin{tabular}{llrrrrrrrrrrr} \hline
               &Correlation: & \multicolumn{3}{c}{weak}                                                      &                      & \multicolumn{3}{c}{moderate}                                                  &                      & \multicolumn{3}{c}{strong}                                                    \\
$(n, d_w, s)$     & Method                   & \multicolumn{1}{c}{RMSE} & \multicolumn{1}{c}{BIAS} & \multicolumn{1}{c}{STD} & \multicolumn{1}{c}{} & \multicolumn{1}{c}{RMSE} & \multicolumn{1}{c}{BIAS} & \multicolumn{1}{c}{STD} & \multicolumn{1}{c}{} & \multicolumn{1}{c}{RMSE} & \multicolumn{1}{c}{BIAS} & \multicolumn{1}{c}{STD} \\ \hline
\multicolumn{13}{c}{Panel A: low-dimensional setting}                                                                                                                                                                                                                                                                                   \\
(100, 50, 6)              & Oracle                   & 0.074                    & 0.020                    & 0.071                   &                      & 0.074                    & 0.020                    & 0.071                   &                      & 0.074                    & 0.020                    & 0.071                   \\
               & PEL                      & 0.088                    & 0.030                    & 0.082                   &                      & 0.092                    & 0.032                    & 0.086                   &                      & 0.081                    & 0.019                    & 0.079                   \\
               & DB-PEL                   & 0.080                    & 0.032                    & 0.073                   &                      & 0.082                    & 0.035                    & 0.075                   &                      & 0.077                    & 0.026                    & 0.073                   \\
               & 2SLS                     & 0.145                    & -0.004                   & 0.145                   &                      & 0.145                    & -0.004                   & 0.145                   &                      & 0.145                    & -0.004                   & 0.145                   \\
(200, 100, 6)  & Oracle                   & 0.039                    & 0.010                    & 0.038                   &                      & 0.039                    & 0.010                    & 0.038                   &                      & 0.039                    & 0.010                    & 0.038                   \\ & PEL                      & 0.044                    & 0.013                    & 0.042                   &                      & 0.047                    & 0.013                    & 0.045                   &                      & 0.045                    & 0.013                    & 0.044                   \\
               & DB-PEL                   & 0.042                    & 0.021                    & 0.036                   &                      & 0.042                    & 0.020                    & 0.037                   &                      & 0.042                    & 0.020                    & 0.037                   \\
         & 2SLS                     & 0.101                    & -0.003                   & 0.101                   &                      & 0.101                    & -0.003                   & 0.101                   &                      & 0.101                    & -0.003                   & 0.101                   \\ \hline
\multicolumn{13}{c}{Panel B: high-dimensional setting}                                                                                                                                                                                                                                                                                  \\
(100, 120, 6)  & Oracle                   & 0.067                    & 0.018                    & 0.064                   &                      & 0.067                    & 0.018                    & 0.064                   &                      & 0.067                    & 0.018                    & 0.064                   \\& PEL                      & 0.090                    & 0.031                    & 0.084                   &                      & 0.098                    & 0.040                    & 0.090                   &                      & 0.092                    & 0.033                    & 0.087                   \\
               & DB-PEL                   & 0.083                    & 0.029                    & 0.078                   &                      & 0.082                    & 0.031                    & 0.076                   &                      & 0.077                    & 0.026                    & 0.073                   \\
               & 2SLS                     & 0.181                    & -0.016                   & 0.180                   &                      & 0.181                    & -0.016                   & 0.180                   &                      & 0.181                    & -0.016                   & 0.180                   \\
(200, 240, 6)  & Oracle                   & 0.035                    & 0.012                    & 0.033                   &                      & 0.035                    & 0.012                    & 0.033                   &                      & 0.035                    & 0.012                    & 0.033                   \\& PEL                      & 0.045                    & 0.014                    & 0.043                   &                      & 0.044                    & 0.014                    & 0.042                   &                      & 0.037                    & 0.014                    & 0.034                   \\
               & DB-PEL                   & 0.043                    & 0.017                    & 0.040                   &                      & 0.042                    & 0.017                    & 0.039                   &                      & 0.035                    & 0.016                    & 0.031                   \\
        & 2SLS                     & 0.128                    & 0.002                    & 0.128                   &                      & 0.128                    & 0.002                    & 0.128                   &                      & 0.128                    & 0.002                    & 0.128                   \\ \hline
(100, 120, 8)  & Oracle                   & 0.063                    & 0.015                    & 0.061                   &                      & 0.063                    & 0.015                    & 0.061                   &                      & 0.063                    & 0.015                    & 0.061                   \\& PEL                      & 0.110                    & 0.036                    & 0.104                   &                      & 0.116                    & 0.045                    & 0.107                   &                      & 0.112                    & 0.039                    & 0.105                   \\
               & DB-PEL                   & 0.100                    & 0.033                    & 0.094                   &                      & 0.098                    & 0.038                    & 0.090                   &                      & 0.092                    & 0.032                    & 0.087                   \\
               & 2SLS                     & 0.211                    & -0.011                   & 0.211                   &                      & 0.211                    & -0.011                   & 0.211                   &                      & 0.211                    & -0.011                   & 0.211                   \\
(200, 240, 12) & Oracle                   & 0.034                    & 0.010                    & 0.033                   &                      & 0.034                    & 0.010                    & 0.033                   &                      & 0.034                    & 0.010                    & 0.033                   \\& PEL                      & 0.038                    & 0.013                    & 0.036                   &                      & 0.036                    & 0.013                    & 0.033                   &                      & 0.036                    & 0.013                    & 0.033                   \\
               & DB-PEL                   & 0.039                    & 0.017                    & 0.035                   &                      & 0.038                    & 0.016                    & 0.034                   &                      & 0.038                    & 0.016                    & 0.034                   \\
          & 2SLS                     & 0.126                    & -0.005                   & 0.126                   &                      & 0.126                    & -0.005                   & 0.126                   &                      & 0.126                    & -0.005                   & 0.126                   \\ \hline
(100, 120, 13) & Oracle                   & 0.065                    & 0.020                    & 0.062                   &                      & 0.065                    & 0.020                    & 0.062                   &                      & 0.065                    & 0.020                    & 0.062                   \\& PEL                      & 0.117                    & 0.051                    & 0.105                   &                      & 0.139                    & 0.069                    & 0.121                   &                      & 0.110                    & 0.036                    & 0.104                   \\
               & DB-PEL                   & 0.103                    & 0.045                    & 0.092                   &                      & 0.113                    & 0.054                    & 0.099                   &                      & 0.093                    & 0.033                    & 0.087                   \\
               & 2SLS                     & 0.200                    & -0.008                   & 0.200                   &                      & 0.200                    & -0.008                   & 0.200                   &                      & 0.200                    & -0.008                   & 0.200                   \\
(200, 240, 17) & Oracle                   & 0.036                    & 0.012                    & 0.034                   &                      & 0.036                    & 0.012                    & 0.034                   &                      & 0.036                    & 0.012                    & 0.034                  \\ & PEL                      & 0.047                    & 0.013                    & 0.045                   &                      & 0.052                    & 0.011                    & 0.051                   &                      & 0.041                    & 0.011                    & 0.039                   \\
               & DB-PEL                   & 0.044                    & 0.016                    & 0.041                   &                      & 0.053                    & 0.015                    & 0.051                   &                      & 0.039                    & 0.015                    & 0.036                   \\
        & 2SLS                     & 0.133                    & -0.014                   & 0.132                   &                      & 0.133                    & -0.014                   & 0.132                   &                      & 0.133                    & -0.014                   & 0.132         \\ \hline
\end{tabular}

\end{spacing}
\end{table}


The parameter of interest in the linear IV model is $\beta_{x}$ in \eqref{eq:main} as it characterizes the ``causal effect'' of the endogenous variable, which bears economic interpretation. The root-mean-square error (RMSE), bias (BIAS) and standard deviation (STD) are calculated for
PEL, DB-PEL and the classical two-stage least squares (2SLS)  for \eqref{eq:main},
and an oracle estimator is added for comparison.\footnote{
The oracle EL in Section \ref{sec:add} is for low-dimensional parameters and moments.
To handle the large pool of orthogonal IVs,
the oracle estimator here is a 2SLS estimator taking advantage of a few most relevant IVs---those with the top $0.1n$ big coefficients $\gamma_{\mathbf{w}2,j}$ in the reduced-form equation.
}
The results are summarized in Table \ref{tab5}.
The performance of PEL and the oracle is comparable,
and the gaps are narrowed when the sample size is increased from $n=100$ to $n=200$, suggesting the capacity for PEL to mimic the oracle by absorbing the signal from the valid moments and in the
meantime keeping the invalid ones at bay.
The RMSE of 2SLS is significantly larger than those of PEL and DB-PEL.


The left panel of Figure \ref{fig1} plots the empirical cumulative distribution functions (ECDF) of the DB-PEL estimates. The dotted curve corresponds to the case of $(n,d_w,s)=(100, 120, 8)$, the dashed curve to
that of $(n,d_w,s)=(200,240,12)$, and the solid curve is the cumulative distribution function (CDF) of $\mathcal{N}(0,1)$
for comparison.
A better normal approximation can be obtained by the PPEL estimate, as shown in the right panel of
Figure \ref{fig1}, with its tuning parameter $\varsigma=0.08(n^{-1}\log p)^{1/2}$.
We also present the confidence intervals (CI) according to PPEL, DB-PEL and 2SLS, as reported in Table \ref{tab9},
for the 90\%, 95\% and 99\% levels, where the CIs for DB-PEL and PPEL are predicated on the asymptotic normality from Theorems \ref{thm2} and \ref{thm:clt}, respectively, and the CIs for 2SLS are based on its asymptotic normality as in standard textbooks.
PPEL's coverage probability to the nominal counterpart
is the best among the three estimators, and is much better than PEL.  2SLS's
coverage probability seems too high in the high-dimensional setting, though it is acceptable in the low-dimensional case.
Figure \ref{fig2} shows that the width of 2SLS's CI is much wider than that of PPEL, which is caused by efficiency loss
from abandoning the potentially valid estimating equations.

\begin{figure}[htbp]
\small
  \centering
  \includegraphics[width=0.7\textwidth]{Rplot_INF.png}
  \caption{ ECDF of DB-PEL (left) and PPEL (right) of $\beta_x$ with moderate correlation between $\epsilon_i$ and ${\mathbf w}_{3i}$}
  \label{fig1}

  \includegraphics[width=0.7\textwidth]{Rplot_CI.png}
  \caption{Width of $90\%$ (left), $95\%$ (middle), and $99\%$ (right) CI for $\beta_x$ estimated by PPEL, DB-PEL and 2SLS under moderate correlation between $\epsilon_i$ and ${\mathbf w}_{3i}$ with $(n,d_w,s)=(100,120,8)$}
  \label{fig2}
\end{figure}





\begin{table}[htbp]
\setlength\tabcolsep{3pt}
\small
\centering
\caption{Coverage probabilities for the CIs of $\beta_x$}
\label{tab9}
\begin{spacing}{1.4}

\begin{tabular}{lllllllllllll} \hline
               & Correlation: & \multicolumn{3}{c}{weak}                                                 &                      & \multicolumn{3}{c}{moderate}                                             &                      & \multicolumn{3}{c}{strong}                                               \\
$(n, d_w, s)$     & Method                   & \multicolumn{1}{c}{90} & \multicolumn{1}{c}{95} & \multicolumn{1}{c}{99} & \multicolumn{1}{c}{} & \multicolumn{1}{c}{90} & \multicolumn{1}{c}{95} & \multicolumn{1}{c}{99} & \multicolumn{1}{c}{} & \multicolumn{1}{c}{90} & \multicolumn{1}{c}{95} & \multicolumn{1}{c}{99} \\ \hline
\multicolumn{13}{c}{Panel A: low-dimensional setting}                                                                                                                                                                                                                                                                    \\
(100, 50, 6)   & PPEL                     & 0.894                  & 0.940                  & 0.988                  &                      & 0.890                  & 0.940                  & 0.986                  &                      & 0.888                  & 0.942                  & 0.986                  \\
               & DB-PEL                      & 0.736                  & 0.830                  & 0.926                  &                      & 0.726                  & 0.794                  & 0.916                  &                      & 0.744                  & 0.818                  & 0.902                  \\
               & 2SLS                     & 0.912                  & 0.964                  & 0.994                  &                      & 0.912                  & 0.964                  & 0.994                  &                      & 0.912                  & 0.964                  & 0.994                  \\
(200, 100, 6)  & PPEL                     & 0.902                  & 0.950                  & 0.994                  &                      & 0.902                  & 0.948                  & 0.994                  &                      & 0.902                  & 0.946                  & 0.994                  \\
               & DB-PEL                      & 0.762                  & 0.838                  & 0.934                  &                      & 0.734                  & 0.842                  & 0.940                  &                      & 0.752                  & 0.852                  & 0.928                  \\
               & 2SLS                     & 0.920                  & 0.956                  & 1.000                  &                      & 0.920                  & 0.956                  & 1.000                  &                      & 0.920                  & 0.956                  & 1.000                  \\ \hline
\multicolumn{13}{c}{Panel B: high-dimensional setting}                                                                                                                                                                                                                                                                   \\
(100, 120, 6)  & PPEL                     & 0.908                  & 0.956                  & 0.996                  &                      & 0.892                  & 0.948                  & 0.994                  &                      & 0.888                  & 0.940                  & 0.992                  \\
               & DB-PEL                      & 0.746                  & 0.832                  & 0.922                  &                      & 0.784                  & 0.850                  & 0.926                  &                      & 0.768                  & 0.836                  & 0.936                  \\
               & 2SLS                     & 0.956                  & 0.982                  & 1.000                  &                      & 0.956                  & 0.982                  & 1.000                  &                      & 0.956                  & 0.982                  & 1.000                  \\
(200, 240, 6)  & PPEL                     & 0.914                  & 0.964                  & 0.988                  &                      & 0.916                  & 0.960                  & 0.988                  &                      & 0.906                  & 0.954                  & 0.988                  \\
               & DB-PEL                      & 0.754                  & 0.846                  & 0.966                  &                      & 0.778                  & 0.854                  & 0.958                  &                      & 0.776                  & 0.866                  & 0.966                  \\
         & 2SLS                     & 0.966                  & 0.992                  & 1.000                  &                      & 0.966                  & 0.992                  & 1.000                  &                      & 0.966                  & 0.992                  & 1.000                  \\ \hline
(100, 120, 8)  & PPEL                     & 0.888                  & 0.942                  & 0.992                  &                      & 0.870                  & 0.930                  & 0.986                  &                      & 0.852                  & 0.924                  & 0.980                  \\
               & DB-PEL                      & 0.778                  & 0.836                  & 0.904                  &                      & 0.738                  & 0.818                  & 0.920                  &                      & 0.730                  & 0.806                  & 0.910                  \\
               & 2SLS                     & 0.960                  & 0.988                  & 1.000                  &                      & 0.960                  & 0.988                  & 1.000                  &                      & 0.960                  & 0.988                  & 1.000                  \\
(200, 240, 12) & PPEL                     & 0.924                  & 0.962                  & 0.990                  &                      & 0.914                  & 0.958                  & 0.986                  &                      & 0.892                  & 0.950                  & 0.982                  \\
               & DB-PEL                      & 0.794                  & 0.872                  & 0.954                  &                      & 0.788                  & 0.864                  & 0.946                  &                      & 0.764                  & 0.860                  & 0.938                  \\
         & 2SLS                     & 0.972                  & 0.998                  & 1.000                  &                      & 0.972                  & 0.998                  & 1.000                  &                      & 0.972                  & 0.998                  & 1.000                  \\ \hline
(100, 120, 13) & PPEL                     & 0.884                  & 0.940                  & 0.992                  &                      & 0.852                  & 0.914                  & 0.976                  &                      & 0.862                  & 0.912                  & 0.974                  \\
               & DB-PEL                      & 0.788                  & 0.862                  & 0.938                  &                      & 0.744                  & 0.798                  & 0.904                  &                      & 0.688                  & 0.768                  & 0.882                  \\
               & 2SLS                     & 0.942                  & 0.986                  & 0.996                  &                      & 0.942                  & 0.986                  & 0.996                  &                      & 0.942                  & 0.986                  & 0.996                  \\
(200, 240, 17) & PPEL                     & 0.916                  & 0.956                  & 0.984                  &                      & 0.914                  & 0.958                  & 0.982                  &                      & 0.886                  & 0.938                  & 0.982                  \\
               & DB-PEL                      & 0.792                  & 0.884                  & 0.968                  &                      & 0.780                  & 0.866                  & 0.954                  &                      & 0.798                  & 0.864                  & 0.952                  \\
         & 2SLS                     & 0.974                  & 0.988                  & 1.000                  &                      & 0.974                  & 0.988                  & 1.000                  &                      & 0.974                  & 0.988                  & 1.000         \\ \hline
\end{tabular}

\end{spacing}
\end{table}


\section{Empirical application}\label{emp_app}

Colonialism was widespread over the globe prior to the Second World War.
While European institutions which protected private properties and checked
government powers
were successfully replicated in a few colonies,
many others suffered expropriation of natural and human resources.
In an influential study, \cite{Acemoglu2001} (AJR, henceforth) systematically explored the relationship
between Europeans' mortality rates and the modes of institutions.

We revisit AJR's open-access dataset.
AJR's main structural equation is
$
y_i = \gamma_y + \beta_x x_i+ \beta_z z_i + \epsilon_i$,
where the dependent variable $y_i$ is {\tt logarithm GDP per capita in 1995},
and the key explanatory variable of interest $x_i$ is {\tt average protection against expropriation risk}, a continuous variable of a scale between 0 and 10. The {\tt latitude} of a colony, denoted as $z_i$, is an additional control variable.
For substantial measurement errors in the institutional index $x_i$ can bias the OLS estimator,
the credibility of the empirical evidence counts on IVs.

AJR compiled from historical documents a seminal variable {\tt logarithm of European settler mortality},
and argued forcefully that it qualifies as a valid IV for the endogenous variable $x_i$.
Their empirical evidence was based on the 2SLS regression in cross-sectional data of countries.
The baseline result in AJR's Column (2) of Table 4 (p.1386) corresponds to the three estimating functions
$
{\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}_i(\boldsymbol{\theta})= (w_{1i}, z_i, 1)^{{ \mathrm{\scriptscriptstyle T} }} \times (y_i -\gamma_y - \beta_x x_i - \beta_z z_i)$ by the notations in our simulation,
and the 2SLS reports $\hat{\beta}_x = 1.00$ with STD 0.22.

There are another 11 potential health variables and institutional variables,
which serve as potential IVs. AJR were uncertain about their validity,
and they experimented in their Table 7 (p.1392) and Table 8 (p.1394) the empirical results under various IV configurations.
These variables, denoted as ${\mathbf w}_2$, are
(i) {\tt Malaria in 1994},
(ii) {\tt Yellow fever},
(iii) {\tt Life expectancy},
(iv) {\tt Infant mortality},
(v) {\tt Mean temperature},
(vi) {\tt Distance from coast},
(vii) {\tt European settlements in 1990},
(viii) {\tt Democracy (1st year of independence)},
(ix) {\tt Constraint on executive (1st year of independence)},
(x) {\tt Democracy in 1900}, and
(xi) {\tt Constraint on executive in 1900}.
(i)--(iv) are health variables, (v)--(vi) are geographic variables, and (vii)--(xi) are institutional variables.
All these extra IVs are associated with the estimating functions
$
{\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}_i (\boldsymbol{\theta})={\mathbf w}_{2i} \times (y_i-\gamma_y - \beta_x x_i - \beta_z z_i) \, .
$
Given the moderate sample size of 56 countries,
a total of 12 IVs is non-trivial.





Our PPEL estimates the main equation as
\[
\widehat{\mbox{log GDP per capita}} = \underset{(1.194)}{2.048} + \underset{(0.126)}{0.945} \times \mathrm{institution}
 - \underset{(0.977)}{0.785} \times  \mbox {lattitude}\,.
\]
For the estimation and inference of
the key parameter of interest $\beta_x$, we compare the results of  PEL, DB-PEL, PPEL and the conventional 2SLS.
Its point estimation (PE), STD, and CIs are presented in Table \ref{tab:r1}.
Enjoying the efficiency gain from the extra IVs, the STD of  PPEL is 0.126
when we select the tuning parameter $\varsigma=\zeta_c(n^{-1}\log p)^{1/2}$ with $\zeta_c=0.08$ (the boldface row), as suggested by our simulations.
This STD achieves a 37\% reduction relative to that of the plain 2SLS with the sole valid IV.
Moreover,
when we vary the constant $\zeta_c=0.08$ in $\varsigma$ as 0.04, 0.06, 0.12 and 0.16, PPEL is rather robust over the wide range of tuning parameters.




\begin{table}[htbp]
\centering
\small

\caption{ Estimation of the effect of institution ($\beta_x$)  }\label{tab:r1}
\begin{spacing}{1.4}


\begin{tabular}{lc|ccc}\hline
                      & $\zeta_c$       & PE    & STD   & 95\% CI       \\ \hline
PEL                   & NA            & 0.937 & 0.078 & (0.786, 1.090) \\
DB-PEL                & NA            & 0.938 & 0.078 & (0.786, 1.090) \\
2SLS                  & NA            & 0.945 & 0.200 & (0.553, 1.338) \\\hline
\multirow{5}{*}{PPEL} & 0.04          & 0.942 & 0.159 & (0.631, 1.254) \\
                      & 0.06          & 0.941 & 0.136 & (0.675, 1.207) \\
                      & \textbf{0.08} & \textbf{0.945} & \textbf{0.126} & \textbf{(0.698, 1.193)} \\
                      & 0.12          & 0.964 & 0.150 & (0.669, 1.259) \\
                      & 0.16          & 0.967 & 0.152 & (0.669, 1.266) \\ \hline
\end{tabular}


\begin{flushleft}
\footnotesize{Note: Our sample size is 56, after removing countries with missing variables from the original sample of 64 countries,
so the 2SLS point estimate is 0.945, slightly different from AJR's 1.00.}
\end{flushleft}

\end{spacing}
\end{table}


Our method is particularly important in unifying the IV selection in AJR's Tables 7 and 8 into a single set of
automatically selected IVs.
Among the 11 potential IVs, our moment selection criterion \eqref{eq:A} invalidates
all institutional variables. Five IVs survive the testing: the climate variable {\tt mean temperature}, the geographic variable {\tt distance from coast}, and
three health variables {\tt dummy of yellow fever}, {\tt infant mortality} and {\tt life expectancy}.
All the other health and institutional variables are assessed as endogenous and are unsuitable for IVs in this study.
The justification for {\tt mean temperature} and {\tt distance from coast} is straightforward because humans were unable to interfere with these natural conditions in the era of colonialism.
Furthermore, AJR argued for the validity of {\tt dummy of yellow fever}, though due to concerns of lack of variation they did not employ it as the main IV (AJR's p.1393, Paragraph 2). Our variable selection result provides supportive evidence
to AJR's heuristics.

\section{Conclusion}
\label{s6}
This paper considers a general setting of an economic structural model
with many potential moments,
some of which may be invalid. These invalid moments must be disciplined in order to estimate the
structural parameter consistently.
We propose a PEL approach to estimate the parameter of interest while coping with the invalid moments.
We show that the PEL estimator is normally distributed asymptotically, and invalid moments can be consistently
detected thanks to the oracle property.
To overcome the difficulty of estimating the bias in the limiting distribution of the PEL estimator, we further devise
the PPEL approach for statistical inference of a low-dimensional object of interest,
which is useful for hypothesis testing and confidence region construction.
Simulation exercises are carried out to demonstrate excellent finite sample performance of our methods.
We revisit an empirical application concerning economic development and shed new insight about
its candidate instruments.