Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
134,829 characters · 10 sections · 105 citation commands
Culling the Herd of Moments with Penalized Empirical Likelihood
\bibpunct{(}{)}{,}{a}{;}
{\it Keywords}: Empirical likelihood; Estimating equations; High-dimensional statistical methods; Misspecification; Moment selection; Penalized likelihood
\thispagestyle{empty}
\setcounter{page}{1}
Economists' perennial pursuit of structural mechanisms leads to models defined by moments. These models can be written in a semiparametric form $\mathbb{E}\{{\mathbf g} ({\mathbf X}_i;\boldsymbol{\theta}_0)\}={\mathbf 0},$ where ${\mathbf g} = (g_j)_{j \in \{1,\ldots, r\} }$ is a vector of $r$ estimating functions, $\boldsymbol{\theta}_0$ is a vector of unknown parameters, and ${\mathbf X}_i$ is observed data. To estimate these models, the most popular method is generalized method of moments (GMM) hansen1982large. Empirical likelihood (EL) QinLawless1994 is a competitive alternative to GMM, thanks to its nice statistical properties. Both GMM and EL are essential building blocks of modern econometrics anatolyev2011methods.
Ideally, economists count on economic theory to guide the choice of variables and moments. However, the truth is that most economic theories are parsimonious abstractions and rarely pinpoint these choices in data-rich environments. The indeterminacy of moment selection brings about three related issues. The first is weak moments stock2012survey, which threatens identification of the true value $\boldsymbol{\theta}_0$ when multiple $\boldsymbol{\theta}$'s satisfy $\mathbb{E}\{{\mathbf g} ({\mathbf X}_i;\boldsymbol{\theta})\} \approx {\mathbf 0}$. Practitioners respond to the concerns of weak moments by adding more moment conditions in the hope to strengthen identification, causing the second issue of many moments roodman2009note. The hazard of many moments is the possible inclusion of invalid moments, meaning $\mathbb{E}\{g_j ({\mathbf X}_i;\boldsymbol{\theta}_0)\}\neq 0$ for some $j\in \{1,\ldots,r\}$ murray2006avoiding, which is the third issue. These cited survey papers highlight the unease incurred by the three challenges and econometricians' efforts in coping with them.
When an underlying economic theory is ambivalent, empirical results based on it can be controversial and susceptible to cherry-picking. In such a circumstance, of great importance are data-driven methods to guide and discipline moment selection. In the low-dimensional settings, andrews2001consistent propose the GMM information criteria, and hong2003generalized follow with the counterpart for EL. Information criteria are evaluated exhaustively at all combinations of moments, and the computation becomes infeasible when there are many potential moments. To overcome this challenge, liao2013adaptive ushers the adaptive Lasso shrinkage into GMM to select among a finite number of moments, and ChengLiao2015 further extend it to deal with a diverging number of moments and accommodate invalid ones.
Unprecedented progress in computation and information technology fuel an arms race between the sheer size of the data and the scale of empirical models. In the era of big data, on the one hand the cost of data collection and processing is tremendously lowered and rich datasets open new perspectives to inspect a myriad of problems; on the other hand economists attempt to build general models to capture various sources of heterogeneity in observational data. Empirical applications abound with models of many potential moments. For instance, eaton2011 create 1360 moments to estimate a structural trade model, altonji2013modeling match 2429 moments implied by a model of earning dynamics, and early works of linear instrumental variables (IV) models produce thousands of instruments by interacting variables angrist1992effect. Although the sample sizes in these examples are non-trivial, the proliferation of moments calls for a moment selection procedure capable of handling high-dimensional moments at a magnitude unrestricted by the sample size. In particular, invalid ones that jeopardize consistency must be identified and “culled” from the herd of moments.
The quadratic form of the usual GMM criterion function is incompatible with high-dimensional moments shi2016econometric, shi2016estimation, and thus belloni2018high regularize GMM with the sup-norm. In this paper, we contribute the most general and versatile procedure for high-dimensional nonlinear settings, to the best of our knowledge. We first develop a penalized empirical likelihood (PEL) solution ChangTangWu2018 to deal with valid and invalid moments simultaneously. We neutralize the invalid moments by an auxiliary parameter, following liao2013adaptive and ChengLiao2015. We establish the rate of convergence and asymptotic normality of the PEL estimator, and show the efficiency gain from incorporating extra valid moments. Under suitable conditions, it transpires that PEL enjoys the oracle property of consistent moment selection and parameter selection.
The asymptotic normal distribution of the PEL estimator involves a bias term caused by the high-dimensional moments. To spare the estimation of the bias term, we can take further actions to project out the influence of the high-dimensional nuisance parameter in the PEL estimator, which is called projected PEL (PPEL) Chang2020. The asymptotic normality of the PPEL estimator is free of bias, which facilitates statistical inference of the structural parameter as well as the validity of moments. Invoking statistical learning to assist our decisions, our method fits well in the recent trend of machine learning for the automatic selection of moments and variables.
Although this paper follows ChangTangWu2018 for estimation and Chang2020 for inference procedures, the key insight lies in the observations that the high-dimensional auxiliary parameter, which signifies the magnitude of misspecification, can be incorporated in the EL method as an additional high-dimensional parameter. In this paper, “high-dimensional” means that the numbers of parameters and/or moments are larger than the sample size, which goes beyond the scope of liao2013adaptive and ChengLiao2015. On the other hand, in order to adapt ChangTangWu2018 and Chang2020 to accommodate misspecified moments, we must deal with the distinctive roles of the main parameter of interest and the auxiliary parameter in identification and the technical challenges induced by them. Differences from ChangTangWu2018 are highlighted in Section (ref) about the rates of converges of the two components of the parameters, and those from Chang2020 are elaborated in Section (ref) about the ways of confidence region construction.
Literature review. Our paper stands on strands of literature, which are too vast to survey exhaustively. The accumulation of moments started from the linear IV model angrist1990lifetime, angrist1991does. The linear IV model motivates theoretical research on issues of many IV bekker1994alternative, weak IV (stock2005testing, andrews2012estimation), many weak IV chao2011asymptotic, hansen2014instrumental, invalid moments and many invalid IV kolesar2015identification,windmeijer2019use, to name a few. In high-dimensional contexts, belloni2012sparse use Lasso method for IV selection in the first stage, and belloni2014inference deal with post-selection inference. Utilizing the linear structure, gold2020inference and caner2018high provide inferential procedures for low-dimensional parameters in models with high-dimensional endogenous variables and high-dimensional IVs. Our method includes the linear IV model as a special case. In particular, the case of high-dimensional structural parameters is elaborated in Section (ref).
The proliferation of moments spreads from linear IV models to nonlinear models. For example, in empirical industrial organization researchers bring in moments from various resources, some of which are guided by economic theory, to mitigate the concerns of weak identification and improve estimation efficiency ackerberg2007econometric. In empirical macroeconomics, identification failure and moment misspecification are common issues mavroeidis2005identification. Under the GMM framework, inference under weak moments (stock2000gmm, kleibergen2005testing, andrews2020optimal), estimation under many weak moments han2006gmm, and robust procedures for invalid moments (ditraglia2016using, caner2018adaptive) have been developed.
EL's attractive theoretical properties are studied extensively (kitamura2001asymptotic, otsu2010bahadur, matsushita2013second, ChangChenChen2015). otsu2006generalized and newey2009generalized deal with its inference under weak IV, and caner2015hybrid select instruments in linear IV models. Penalization schemes on EL have been introduced by otsu2007penalized, tang2010penalized and ChangTangWu2018.
Organization. The rest of the paper is organized as follows. Section (ref) introduces the model and the analytic framework. We first derive the asymptotic properties of our estimation in the low-dimensional case in Section 3, extend them to the high-dimensional structural parameter in Section 4, and we further refine PEL with projection to eliminate its bias in Section 5. The theoretical results are supported in Section 6 by Monte Carlo simulations. The influential study of the determinants of economic outcomes after colonialism is revisited in Section 7. Section 8 concludes the paper. Due to the limitations of space, the proofs and the technical details of secondary importance are relegated into the supplementary materials.
Notations. We conclude Introduction with notations used throughout the paper. “Low-dimensional” is referred to the cases that the number of parameters or moments is much smaller than the sample size $n$, whereas “high-dimensional” goes the opposite. For two sequences of positive numbers $\{a_n\}$ and $\{b_n\}$, we write $a_n\lesssim b_n$ or $b_n\gtrsim a_n$ if there exists a positive constant $c$ such that $\limsup_{n\rightarrow\infty}a_n/b_n\leq c$, and write $a_n\ll b_n$ or $b_n\gg a_n$ if $\limsup_{n\rightarrow\infty}a_n/b_n=0$.
Denote by $1(\cdot)$ the indicator function. For a positive integer $q$, we write $[q]=\{1,\ldots,q\}$. For a $q\times q$ symmetric matrix ${\mathbf M}$, denote by $\lambda_{\min}({\mathbf M})$ and $\lambda_{\max}({\mathbf M})$ the smallest and largest eigenvalues of ${\mathbf M}$, respectively. For a $q_1\times q_2$ matrix ${\mathbf B}=(b_{i,j})_{q_1\times q_2}$, let ${\mathbf B}^{ \mathrm{\scriptscriptstyle T} }$ be its transpose, ${\mathbf B}^{\otimes2}={\mathbf B}{\mathbf B}^{ \mathrm{\scriptscriptstyle T} }$, ${\mathbf B}^{\circ\kappa}=(|b_{i,j}|^\kappa)_{q_1\times q_2}$ for any $\kappa>0$, $|{\mathbf B}|_\infty=\max_{i\in[q_1],j\in[q_2]}|b_{i,j}|$ be the sup-norm, and $\|{\mathbf B}\|_2=\lambda_{\max}^{1/2}({\mathbf B}^{\otimes2})$ be the spectral norm. Specifically, if $q_2=1$, we use $|{\mathbf B}|_\infty = \max_{i\in[q_1]}|b_{i,1}|$, $|{\mathbf B}|_1=\sum_{i=1}^{q_1}|b_{i,1}|$ and $|{\mathbf B}|_2=(\sum_{i=1}^{q_1}b_{i,1}^2)^{1/2}$ to denote the $L_{\infty}$-norm, $L_1$-norm and $L_2$-norm of the $q_1$-dimensional vector ${\mathbf B}$, respectively. For two square matrices ${\mathbf M}_1$ and ${\mathbf M}_2$, we say ${\mathbf M}_1 \le {\mathbf M}_2$ if $({\mathbf M}_2 - {\mathbf M}_1)$ is a positive semi-definite matrix.
The population mean is denoted by $\mathbb{E}(\cdot)$, and the sample mean by $\mathbb{E}_n(\cdot)= n^{-1}\sum_{i=1}^n \{\cdot\}$. For a given index set $\mathcal{L}$, let $| \mathcal{L} | $ be its cardinality. For a generic multivariate function ${\mathbf h}(\cdot;\cdot)$, we denote by ${\mathbf h}_{{ \mathcal{\scriptscriptstyle L} }}(\cdot;\cdot)$ the subvector of ${\mathbf h}(\cdot;\cdot)$ collecting the components indexed by $\mathcal{L}$. Analogously, we write ${\mathbf a}_{{ \mathcal{\scriptscriptstyle L} }}$ as the corresponding subvector of ${\mathbf a}$. For simplicity and when no confusion arises, we use the generic notation ${\mathbf h}_i(\boldsymbol{\theta})$ as the equivalence to ${\mathbf h}({\mathbf X}_i;\boldsymbol{\theta})$, and $\nabla_{\boldsymbol{\theta}} {\mathbf h}_i(\boldsymbol{\theta}) $ for the first-order partial derivative of ${\mathbf h}_i(\boldsymbol{\theta})$ with respect to $\boldsymbol{\theta}$. Denote by $h_{i,k}(\boldsymbol{\theta})$ the $k$-th component of ${\mathbf h}_i(\boldsymbol{\theta})$, and by $\nabla^2_{\boldsymbol{\theta}}h_{i,k}(\boldsymbol{\theta}) $ the second derivative of $h_{i,k}(\boldsymbol{\theta})$ with respect to $\boldsymbol{\theta}$. Let $\bar {\mathbf h}(\boldsymbol{\theta})=\mathbb{E}_n\{{\mathbf h}_i (\boldsymbol{\theta})\}$, and write its $k$-th component as $\bar h_k(\boldsymbol{\theta})= \mathbb{E}_n\{h_{i,k}(\boldsymbol{\theta})\}$. Analogously, let ${\mathbf h}_{i,{ \mathcal{\scriptscriptstyle L} }}(\boldsymbol{\theta})={\mathbf h}_{{ \mathcal{\scriptscriptstyle L} }}({\mathbf X}_i;\boldsymbol{\theta})$ and $\bar{{\mathbf h}}_{{ \mathcal{\scriptscriptstyle L} }}(\boldsymbol{\theta})= \mathbb{E}_n \{{\mathbf h}_{i,{ \mathcal{\scriptscriptstyle L} }}(\boldsymbol{\theta})\}$.
In this section we introduce the model and the EL estimation. Let ${\mathbf X}_1,\ldots,{\mathbf X}_n$ be $d$-dimensional independent and identically distributed generic observations, and $\boldsymbol{\theta}=(\theta_1,\ldots,\theta_p)^{ \mathrm{\scriptscriptstyle T} }$ be a $p$-dimensional parameter taking values in $\boldsymbol{\Theta} \subset \mathbb{R}^p$. For a set of $r_1$ estimating functions $ {\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}(\cdot;\cdot)=\{ g_{j}^{ \mathcal{\scriptscriptstyle (I)} } (\cdot; \cdot) \}_{j \in \mathcal{I}} $, the information of the model parameter $\boldsymbol{\theta}$ is collected by the unbiased moment condition
at the unknown true parameter $\boldsymbol{\theta}_0\in\boldsymbol{\Theta}$, where $r_1\geq p$ is necessary for identifying $\boldsymbol{\theta}_0$. The superscript $(\mathcal{I})$ labels this initial set, to be distinguished from the other set $(\mathcal{D})$ in Section (ref).
Motivated from empirical applications in asset pricing, the two-step GMM hansen1982large was the default estimating method for the moment-defined model ((ref)). Intensive theoretical studies and numerical evidence in 1980's and 90's revealed some undesirable finite-sample properties of the two-step GMM altonji1996small. EL and the continuously updating GMM (CUE) hansen1996finite emerged as competitive solutions, and they were later unified as members of generalized empirical likelihood newey2003higher.
This paper focuses on EL with estimating equations:
proposed by QinLawless1994 based on the seminal idea of EL Owen1988,Owen1990. Maximizing $L(\boldsymbol{\theta})$ with respect to $\boldsymbol{\theta}$ delivers the EL estimator $ \hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }}=\arg\max_{\boldsymbol{\theta}\in\boldsymbol{\Theta}}L(\boldsymbol{\theta})$, which can be carried out equivalently by solving the corresponding dual problem
where $\hat{\Lambda}_{n}^{{ \mathcal{\scriptscriptstyle (I)} }}(\boldsymbol{\theta})=\{\boldsymbol{\lambda}\in\mathbb{R}^{r_1}:\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}_i(\boldsymbol{\theta})\in\mathcal {V}~\textrm{for any}~i\in[n]\}$ and $\mathcal{V}$ is an open interval containing zero.
To fix ideas, we first study the asymptotic property of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }}$ under regularity conditions. When the sample size $n$ grows, we adopt the asymptotic framework of Hjort2009 and ChangChenChen2015 to take the observations $\{{\mathbf g}_i^{{ \mathcal{\scriptscriptstyle (I)} }}(\boldsymbol{\theta})\}_{i=1}^n$ as a multi-index array, where $r_1$, $p$ and $d$ may depend on $n$. Proposition A.1 in the supplementary materials shows that the standard asymptotic normality for $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }}$ holds under mild regularity conditions.
In applied econometrics, there has been a tendency of assembling many IVs or creating many moments to identify the parameter of interest. In these applications, researchers often have some ideas about the relative importance of moments; in the meantime, when researchers work with more and more moments, some invalid ones may creep in.
To put into an analytic framework a herd of extra moments with unknown validity ex ante, suppose that $r_2$ estimating functions ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}=\{g_{j}^{ \mathcal{\scriptscriptstyle (D)} }\}_{j\in \mathcal{D}}$ are partitioned into two groups: \[ \mathcal{A}=\{j\in \mathcal{D}:\mathbb{E}\{g^{ \mathcal{\scriptscriptstyle (D)} }_{i,j}(\boldsymbol{\theta}_0)\}=0\}~~~\textrm{and}~~~ \mathcal{A}^{{\mathrm{c}}}=\{j\in \mathcal{D}:\mathbb{E}\{g^{ \mathcal{\scriptscriptstyle (D)} }_{i,j}(\boldsymbol{\theta}_0)\}\neq 0\}\,. \] Here the estimating functions in the set $\mathcal{A}$ are correctly specified, which can help improve the efficiency in estimating $\boldsymbol{\theta}_0$. On the contrary, those in $\mathcal{A}^{{\mathrm{c}}}$ are misspecified, and they can only undermine the identification of $\boldsymbol{\theta}_0$.
If there is an “oracle” that reveals which components of ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}$ belong to $\mathcal{A}$, i.e., the valid estimating functions, we can collect all valid estimating functions indexed by $\mathcal{H}:=\mathcal{I}\cup\mathcal{A}$ to estimate $\boldsymbol{\theta}_0$. Denote ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (H)} }} =\{ {\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} },{ \mathrm{\scriptscriptstyle T} }},{\mathbf g}^{{ \mathcal{\scriptscriptstyle (A)} },{ \mathrm{\scriptscriptstyle T} }}\}^{ \mathrm{\scriptscriptstyle T} }$ and $h=|\mathcal{H}|$ as the total number of valid moments. The associated EL estimation for $\boldsymbol{\theta}_0$ based on the estimating functions ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (H)} }}$ is given by
where $\hat{\Lambda}_{n}^{{ \mathcal{\scriptscriptstyle (H)} }}(\boldsymbol{\theta})=\{\boldsymbol{\lambda}\in\mathbb{R}^h:\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (H)} }}_i(\boldsymbol{\theta})\in\mathcal {V}~\textrm{for any}~i\in[n]\}$. By the same arguments as those in Proposition A.1 in the supplementary materials, the following Proposition (ref) verifies the asymptotic normality for $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (H)} }}$.
Obviously $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (H)} }}$ is more efficient than $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }}$ thanks for the additional moment restrictions in $\mathcal{A}$, which disclose more information about the parameter and thus tie down the estimation variability. This is parallel to hall2007information in GMM for finite numbers of parameters and moments. In reality, however, we are oblivious to $\mathcal{A}$, and therefore we cannot blindly summon all $r_2$ estimating functions in ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}$ to estimate $\boldsymbol{\theta}_0$. Following liao2013adaptive, we introduce an auxiliary parameter $\boldsymbol{\xi}=\mathbb{E}\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}_i(\boldsymbol{\theta})\}$ to neutralize the misspecified moments. Denote the augmented parameter $\boldsymbol{\psi}=(\boldsymbol{\theta}^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$, and the augmented parameter space $\boldsymbol{\Psi} = \boldsymbol{\Theta}\times\mathbf{\Upsilon}$. Let $r=r_1+r_2$ and stack these $r$ estimating functions as ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}({\mathbf X};\boldsymbol{\psi})=\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}({\mathbf X};\boldsymbol{\theta})^{ \mathrm{\scriptscriptstyle T} },{\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}({\mathbf X};\boldsymbol{\theta})^{ \mathrm{\scriptscriptstyle T} }-\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} }\}^{ \mathrm{\scriptscriptstyle T} }$. Then $\boldsymbol{\psi}_0=(\boldsymbol{\theta}_0^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\xi}_0^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ with $\boldsymbol{\xi}_0=\mathbb{E}\{{\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}_i(\boldsymbol{\theta}_0)\}$ can be identified by
Can we directly include the auxiliary parameter $\boldsymbol{\xi}$ into the EL estimation? The answer is negative. If the auxiliary parameter is not regularized, the EL estimators with and without $\boldsymbol{\xi}$ are the same, up to numerical errors.
Proposition (ref) states the equivalence between $\hat\boldsymbol{\theta}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathrm{\scriptscriptstyle EL} }}$ and $\hat\boldsymbol{\theta}^{{ \mathcal{\scriptscriptstyle (I)} }}_{{ \mathrm{\scriptscriptstyle EL} }}$. In the next section, we will show that desirable efficiency as in the oracle estimator $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (H)} }}$ can be achieved if we extend liao2013adaptive's idea of shrinking the auxiliary parameter\footnote{ In GMM estimation under a fixed $r$, liao2013adaptive shrinks $\boldsymbol{\xi}$ toward zero using the adaptive Lasso zou2006adaptive. } by further penalizing the Lagrange multipliers associated with the high-dimensional moments.
We consider in this section the model (ref) with a low-dimensional parameter $\boldsymbol{\theta}$, low-dimensional estimating functions ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}$ and high-dimensional estimating functions ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}$. In this setting, $p$ and $r_1$ are either fixed or diverge at some slow polynomial rates of $n$, whereas $r_2$ can grow much larger than $n$. The components in the parameter $\boldsymbol{\theta}$ are indexed by $\mathcal{P}$. All moments together are indexed by $\mathcal{T} = \mathcal{I} \cup \mathcal{D}$, in which the researcher knows that those in $\mathcal{I}$ are correctly specified in the sense that $\mathbb{E}\{ {\mathbf g}_i^{{ \mathcal{\scriptscriptstyle (I)} }}(\boldsymbol{\theta}_0)\} = {\mathbf 0}$, whereas she is uncertain whether those in $\mathcal{D} = \mathcal{A}\cup \mathcal{A}^{{\mathrm{c}}} $ satisfy $\mathbb{E}\{ {\mathbf g}_i^{{ \mathcal{\scriptscriptstyle (D)} }}(\boldsymbol{\theta}_0)\} = {\mathbf 0} $ or not. As there is a one-to-one relationship between $\boldsymbol{\xi}$ and the uncertain moments, we use $\mathcal{D}$ to index the components in $\boldsymbol{\xi}$ as well.
We propose the following PEL to simultaneously estimate the unknown parameter $\boldsymbol{\theta}_0$ and determine the validity of the estimating functions in ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}$:
where $\hat{\Lambda}_{n}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})=\{\boldsymbol{\lambda}\in\mathbb{R}^{r}:\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\boldsymbol{\psi})\in\mathcal {V}~\textrm{for any}~i\in[n]\}$, and $P_{1,\pi}(\cdot)$ and $P_{2,\nu}(\cdot)$ are two penalty functions with tuning parameters $\pi$ and $\nu$, respectively. With the penalty function $P_{2,\nu}(\cdot)$ and appropriately selected tuning parameter $\nu$, the estimator $(\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} },\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ is associated with a sparse Lagrange multiplier $\boldsymbol{\lambda}$. Since the sparse $\boldsymbol{\lambda}$ invokes a subset of the estimating functions ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}(\cdot;\cdot)$, it digests the high-dimensional moments as long as the number of nonzero components in $\boldsymbol{\lambda}$ is small. On the other hand, the penalty $P_{1,\pi}(\cdot)$ is applied to identify which components in ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}(\cdot;\cdot)$ are correctly specified and can estimate $(\xi_{0,k})_{k\in\mathcal{A}}$ exactly as $0$ with high probability.
These penalties originally appeared in ChangTangWu2018 though, there are several important differences. First, given the low-dimensional $\boldsymbol{\theta}$, it is unnecessary to assume sparsity on $\boldsymbol{\theta}_0$ for identification. As a result, our procedure here in (ref) only penalizes $\boldsymbol{\xi}$ and $\boldsymbol{\lambda}_{\mathcal{D}}$, not the entire parameter $\boldsymbol{\psi}=(\boldsymbol{\theta}^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ and the associated Lagrange multiplier $\boldsymbol{\lambda}$. Were $\boldsymbol{\lambda}_{\mathcal{I}}$ penalized, we would rule out some components of ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}(\cdot;\cdot)$ and thus may suffer loss of estimation efficiency for $\boldsymbol{\theta}_0$. Second, the identification condition invoked in this paper deviates from that in ChangTangWu2018. Our Condition (ref) for identification is imposed on $\boldsymbol{\theta}_0$, rather than the whole parameter $\boldsymbol{\psi}_0$. In contrast, ChangTangWu2018 specify identification condition for the parameter entity. While estimators in ChangTangWu2018 share the same rate of convergence, we must deal with the disparate rates of convergence of the main component $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ and the auxiliary $\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$, which significantly complicates the theoretical analysis of (ref), as detailed in Proposition (ref) below.
For any penalty function $P_\tau(\cdot)$ with a tuning parameter $\tau$, let $\rho(t;\tau)=\tau^{-1}P_\tau(t)$ for any $t\in[0,\infty)$ and $\tau\in(0,\infty)$. We assume that the two penalty functions $P_{1,\pi}(\cdot)$ and $P_{2,\nu}(\cdot)$ involved in (ref) belong to the following class:
This class $\mathscr{P}$, considered in LvFan2009, is broad and general. The commonly used $L_1$ penalty, SCAD penalty FanLi2001 and MCP penalty Zhang2010 are all included in $\mathscr{P}$. When $P_\tau(\cdot)\in\mathscr{P}$, we write the associated $\rho'(0^+;\tau)$ as $\rho'(0^+)$ for simplification.
To establish the limiting distribution of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ in (ref), we assume the following regularity conditions.
For any index set $\mathcal{F}\subset\mathcal{T}$ and $\boldsymbol{\psi}\in\boldsymbol{\Psi}$, define $ {\mathbf V}_{{ \mathcal{\scriptscriptstyle F} }}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})=\mathbb{E}\{{\mathbf g}_{i,{ \mathcal{\scriptscriptstyle F} }}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})^{\otimes2}\}$. When $\mathcal{F}=\mathcal{T}$, we write ${\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})={\mathbf V}_{{ \mathcal{\scriptscriptstyle T} }}^{{ \mathcal{\scriptscriptstyle (T)} }}(\boldsymbol{\psi})$ for conciseness.
Conditions (ref)--(ref) are standard regularity assumptions in the literature. Condition (ref) is an identification assumption of the estimating equations in the known set $\mathcal{I}$. Write $\boldsymbol{\xi}_0= (\xi_{0,k})_{k\in\mathcal{D}}$, and we continue by defining
in this section, where $\aleph_n=(n^{-1}\log r)^{1/2}$. Suppose:
to control the bias induced by $P_{1,\pi}(\cdot)$ on $\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$. With the assumption $\phi_n=o(\min_{k\in\mathcal{A}^{{\mathrm{c}}}}|\xi_{0,k}|)$ that the nonzero components of $\boldsymbol{\xi}_0$ do not diminish to zero too fast, (ref) can be replaced by
for some constant $c\in(0,1)$. If we select $P_{1,\pi}(\cdot)$ as an asymptotically unbiased penalty such as SCAD or MCP, we have $\chi_n=0$ in (ref) when
To simplify the presentation, in this section we assume that (ref) holds and $\chi_n=0$ in (ref).\footnote{ If (ref) is violated, the asymptotic normality of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ in Theorem (ref) below will still hold under (ref) along with more complicated notations to spell out the restrictions. }
To allocate the parameter of interest $\boldsymbol{\theta}$ and those invalid moments in $\mathcal{A}^{{\mathrm{c}}}$, we define an index set $\mathcal{S}=\mathcal{P}\cup \mathcal{A}^{{\mathrm{c}}}$ and its cardinality $s=|\mathcal{S}|$. For any $\boldsymbol{\psi} \in\boldsymbol{\Psi}$, the set $\mathcal{S}$ picks out $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }}=(\boldsymbol{\theta}^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\xi}_{{ \mathcal{\scriptscriptstyle A} }^{{\mathrm{c}}}}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$. Since $\mathcal{A} = \mathcal{S}^{{\mathrm{c}}} = (\mathcal{P}\cup\mathcal{D}) \backslash \mathcal{S}$, the corresponding auxiliary coefficients can be written as $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}=\boldsymbol{\xi}_{{ \mathcal{\scriptscriptstyle A} }}$ and furthermore $\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}={\mathbf 0}$ as $\mathcal{A}$ is the set of valid moments. Given some constant $C_*\in(0,1)$, define $\mathcal{M}^*_{\boldsymbol{\psi}}= \mathcal{I} \cup \mathcal{D}^*_{\boldsymbol{\psi}}$ with $\mathcal{D}^*_{\boldsymbol{\psi}}=\{ j\in \mathcal{D}: |\bar{g}^{{ \mathcal{\scriptscriptstyle (T)} }}_{j}(\boldsymbol{\psi})|\ge C_*\nu\rho'_2(0^+)\}$. We assume the existence of a sequence $\ell_n\to\infty$ such that $$ \mathbb{P}\bigg(\max_{\boldsymbol{\psi}\in\boldsymbol{\Psi}:\,|\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }}-\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle S} }}|_\infty\le c_n,\,|\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}|_1\leq \aleph_n }|\mathcal{M}^*_{\boldsymbol{\psi}}| \le \ell_n\bigg) \to 1 $$ with some $c_n\gg\phi_n$.
Let $b_{1,n}=\max\{a_n,r_{1}\aleph_n^2\}$ and $b_{2,n}=\max\{b_{1,n},\nu^2 \}$. Then $ \phi_n=\max\{pb_{1,n}^{1/2}, b_{2,n}^{1/2} \}$. Define $\boldsymbol{\Psi}_*=\{\boldsymbol{\psi}=(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }}^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }:|\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }}-\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle S} }}|_\infty\le\varepsilon,|\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }^{{\mathrm{c}}}}|_1\le \aleph_n \}$ for some fixed $\varepsilon>0$. Consider
Proposition (ref) shows that such a $\hat{\boldsymbol{\psi}}$ is a sparse local minimizer for (ref).
In the rest of this section, we focus on the sparse local minimizer $\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }}=(\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} },\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ specified in (ref). We proceed with additional regularity conditions.
Condition (ref) is the sparse Riesz condition (zhang2008sparsity, chen2008extended) in our setting to deal with the high-dimensional $\boldsymbol{\psi}$ when $p+r_2>n$. Let $\hat\boldsymbol{\lambda}(\hat\boldsymbol{\psi}_{ \mathrm{\scriptscriptstyle PEL} })$ be the $r\times 1$ vector of Lagrange multiplier defined at $\hat\boldsymbol{\psi}_{ \mathrm{\scriptscriptstyle PEL} }$:
Write $\hat{\boldsymbol{\lambda}}:=\hat\boldsymbol{\lambda}(\hat\boldsymbol{\psi}_{ \mathrm{\scriptscriptstyle PEL} })=(\hat{\lambda}_1,\ldots,\hat{\lambda}_r)^{ \mathrm{\scriptscriptstyle T} }$ and $\rho_2(t;\nu)=\nu^{-1}P_{2,\nu}(t)$. Let $\mathcal{R}_n= \mathcal{I} \cup \{ j\in \mathcal{D}: \hat{\lambda}_j \neq 0 \} $ be the set of the estimated binding moments. Since only those $\lambda_j$'s associated with $\mathcal{D}$ are penalized, it follows that \[ \frac{1}{n}\sum_{i=1}^n\frac{g_{i,j}^{{ \mathcal{\scriptscriptstyle (T)} }}(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})}{1+\hat{\boldsymbol{\lambda}}_{{ \mathcal{\scriptscriptstyle R} }_n}^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}_{i,{ \mathcal{\scriptscriptstyle R} }_n}^{{ \mathcal{\scriptscriptstyle (T)} }}(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})}=\left\{
\right. \] For any unbinding moment $j\in\mathcal{R}_n^{{\mathrm{c}}}$, define
If $\hat{\lambda}_j=0$ for some $j\in\mathcal{D}$, the subdifferential at $\hat{\lambda}_j$ has to include the zero element bertsekas1997nonlinear. That is, $\hat{\eta}_j\in[-\nu\rho_2'(0^+),\nu\rho_2'(0^+)]$ for any $j\in\mathcal{R}_n^{{\mathrm{c}}}$. In our theoretical analysis, we impose the next condition.
Define $\mathcal{A}_*=\{ j \in \mathcal{A}:\hat{\lambda}_j\neq 0\}$ and $\mathcal{A}_{*,{\mathrm{c}}}=\{ j\in \mathcal{A}^{{\mathrm{c}}}:\hat{\lambda}_j\neq 0\}$. Let $\mathcal{I}^*=\mathcal{I} \cup\mathcal{A}_*$, and then $\mathcal{R}_n$ can be decomposed into two disjoint parts $\mathcal{I}^*$ and $\mathcal{A}_{*,{\mathrm{c}}}$. Furthermore, we define $\mathcal{S}_*= \mathcal{P} \cup \mathcal{A}_{*,{\mathrm{c}}}$. The fact $\mathcal{S}_*\subset\mathcal{S}$ implies that $|\mathcal{S}_*|=p+|\mathcal{A}_{*,{\mathrm{c}}}|\le |\mathcal{S}|=s$. For any $\boldsymbol{\psi}\in\boldsymbol{\Psi}$, we have $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }_*}=(\boldsymbol{\theta}^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} }_{{{ \mathcal{\scriptscriptstyle A} }}_{*,{\mathrm{c}}}})^{ \mathrm{\scriptscriptstyle T} }$. Define
with ${\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}(\boldsymbol{\theta}_0)=\mathbb{E}\{{\mathbf g}^{ \mathcal{\scriptscriptstyle (T)} }_{i,{ \mathcal{\scriptscriptstyle I} }^*}(\boldsymbol{\theta}_0)^{\otimes2}\}$, and $\hat\boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n}=\{{\mathbf J}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle R} }_n}\}^{-1}[\mathbb{E}\{\nabla_{\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }_*}}{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_{i,{ \mathcal{\scriptscriptstyle R} }_n}(\boldsymbol{\psi}_0)\}]^{ \mathrm{\scriptscriptstyle T} }\{{\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle R} }_n}(\boldsymbol{\psi}_0)\}^{-1}\hat\boldsymbol{\eta}_{{ \mathcal{\scriptscriptstyle R} }_n}$, where ${\mathbf J}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle R} }_n}=([\mathbb{E}\{\nabla_{\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }_*}}{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_{i,{ \mathcal{\scriptscriptstyle R} }_n}(\boldsymbol{\psi}_0)\}]^{ \mathrm{\scriptscriptstyle T} }\{{\mathbf V}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle R} }_n}(\boldsymbol{\psi}_0)\}^{-1/2})^{\otimes2}$, $\hat\boldsymbol{\eta}=(\hat{\eta}_j)_{j\in \mathcal{T}}$ with $\hat\eta_j=0$ for $j\in \mathcal{I}$, $\hat\eta_j=\nu\rho'_2(|\hat\lambda_j|;\nu){\rm sgn}(\hat\lambda_j)$ for $j\in \mathcal{D}$ with $\hat\lambda_j\neq0$, and $\hat\eta_j$ specified in (ref) for $j\in\mathcal{R}_n^{{\mathrm{c}}}$. The limiting distribution of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ is stated in Theorem (ref).
Thanks to the extra valid moments in ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}$, the asymptotic covariance of $\hat\boldsymbol{\theta}_{{ \mathrm{\scriptscriptstyle PEL} }}$ is smaller than that of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{ \mathcal{\scriptscriptstyle (I)} }$ which utilizes the information in ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}$ only. On the other hand, given that $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle EL} }}^{{ \mathcal{\scriptscriptstyle (I)} }}$ provides asymptotic normality under the set of correctly specified moments $\mathcal{I}$, one may question whether it is worthwhile to consider all the available moments among which some may risk misspecification. This is a trade-off between robustness and efficiency, a recurrent theme in modern econometrics. Although it inevitably depends on the data generating process (DGP), in big data environments the efficiency gain may be substantial. Benefits are demonstrated in the simulation exercises and the empirical application via very simple econometric models.\footnote{ See the RMSEs in Table (ref), the length of confidence intervals in Figure (ref), and the standard deviations in Table (ref). }
While the augmented parameter $\boldsymbol{\xi}$ essentially validates all moments in $\mathcal{D}$, the efficiency gain comes from the penalty on $\boldsymbol{\xi}$ that shrinks some $\hat{\xi}_{k}$, $k \in \mathcal{A}$, to zero, thereby confirming the validity of these associated moments. To discuss moment selection, we write $\hat\boldsymbol{\xi}_{{ \mathrm{\scriptscriptstyle PEL} }}=(\hat{\xi}_k)_{k\in \mathcal{D}}$. If all valid estimating functions in $\mathcal{A}$ are selected by the optimization (ref), i.e., $\hat{\xi}_k = 0$ for all $k\in \mathcal{A}$, then the asymptotic covariance of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ is $\{{\mathbf J}^{{ \mathcal{\scriptscriptstyle (H)} }}\}^{-1}$.
Remind that $\{{\mathbf J}^{{ \mathcal{\scriptscriptstyle (H)} }}\}^{-1}$ in Proposition (ref) is the semiparametric efficiency bound for the estimation of $\boldsymbol{\theta}_0$ under the oracle, and Proposition (ref) shows that $\mathbb{P}(\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle A} }}={\mathbf 0})\rightarrow1$ as $n\rightarrow\infty$, which provides a natural moment selection criterion
to identify the valid estimating functions in $\mathcal{A}$. Notice that Proposition (ref) also indicates that $|\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle A} }^{{\mathrm{c}}}}-\boldsymbol{\xi}_{0,{ \mathcal{\scriptscriptstyle A} }^{{\mathrm{c}}}}|_\infty=O_{ \mathrm{p} }(\phi_n)$. Together with (ref), we have $\mathbb{P}[\cup_{k\in\mathcal{A}^{{\mathrm{c}}}}\{\hat{\xi}_k=0\}]\rightarrow0$ as $n\rightarrow\infty$. Based on these arguments, Theorem (ref) supports our proposed moments selection criterion.
Up to now, we have established the asymptotic properties of the PEL estimator for the moment-defined model ((ref)) with high-dimensional moments and low-dimensional $\boldsymbol{\theta}$. The next section further extends the estimation to a high-dimensional structural parameter.
A high-dimensional parameter $\boldsymbol{\theta}$ is present when many control variables are included in the structural model. For example, the leading case of the linear IV model takes only one endogenous variable and it is accompanied by a few corresponding IVs. Nevertheless, such a simple setting may still involve many exogenous control variables in the main equation, and these control variables are natural instruments for themselves. An empirical example can be found in blundell1993we. fan2014endogeneity propose the focused GMM for such a linear IV model, which is a special case of the moment-defined model.
In this section, $p$ and $r_2$ are both allowed to be much larger than $n$, whereas $r_1$ is fixed or diverges at some slow polynomial rate of $n$. We assume $\boldsymbol{\theta}_0$ sparse in the sense that most of its components are zero. Write $\boldsymbol{\theta}_0=(\theta_{0,l})_{l \in \mathcal{P}} $ and define the active set $\mathcal{P}_{\sharp}=\{l \in \mathcal{P}:\theta_{0,l}\neq0\}$ with cardinality $p_{\sharp}=|\mathcal{P}_{\sharp}|$. To obtain a sparse estimate of $\boldsymbol{\theta}_0$, we add to (ref) a penalty on $\boldsymbol{\theta}$ to produce the following optimization problem:
With slight abuse of notation, we keep using $(\hat{\boldsymbol{\theta}}_{{{ \mathrm{\scriptscriptstyle PEL} }}}^{ \mathrm{\scriptscriptstyle T} },\hat{\boldsymbol{\xi}}_{{{ \mathrm{\scriptscriptstyle PEL} }}}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ to denote the solution to (ref).
As $p$ can be bigger than $r_1$ in high dimension, here we update Condition (ref) for the identification of $\boldsymbol{\theta}_0$. In view of ChangTangWu2018, we impose Condition (ref) below for the identification of the nonzero components of $\boldsymbol{\theta}_0$.
\addtocounter{as}{-6}
As we have shown in Section (ref), the index set $\mathcal{S}$ and the quantities $a_n$ and $\phi_n$ play key roles in the theoretical analysis of the low-dimensional $\boldsymbol{\theta}$ in (ref). Under the current high-dimensional setting, we define
This $\mathcal{S}$ indexes all nonzero components in the true augmented parameter $\boldsymbol{\psi}_0$. Compared to its counterpart $a_n$ in Section (ref), the newly defined $a_n$ here is amended with an extra term $\sum_{l\in\mathcal{P}}P_{1,\pi}(|\theta_{0,l}|)$ due to the penalty imposed on $\boldsymbol{\theta}$. In the current high-dimensional setting, we also redefine
with $a_n$ in (ref). To control the bias introduced by $P_{1,\pi}(\cdot)$ on $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} }}$ and $\hat{\boldsymbol{\xi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$, similar to (ref), we assume that among the elements in $\boldsymbol{\psi}_0=(\psi_{0,1},\ldots,\psi_{0,p+r_2})^{ \mathrm{\scriptscriptstyle T} }$, the minimal active signal is
with $\phi_n$ in (ref) and $\max_{{k\in \mathcal{S}}}\sup_{c|\psi_{0,k}|<t<c^{-1}|\psi_{0,k}|}P'_{1,\pi}(t)=0$ for some constant $c\in(0,1)$.\footnote{See the arguments below (ref) for the validity of these assumptions.} Given the newly defined $\mathcal{S}$ and $\phi_n$, we can further update $\ell_n$, $\mathcal{R}_n$, $\mathcal{A}_*$, $\mathcal{A}_{*,{\mathrm{c}}}$ and $\mathcal{I}^*$ in the same manner as their counterparts in Section (ref). Moreover, we also define ${\mathbf J}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle R} }_n}$ and $\hat\boldsymbol{\zeta}_{{ \mathcal{\scriptscriptstyle R} }_n}$ as specified in Section (ref) with $\mathcal{S}_*=\mathcal{P}_{\sharp}\cup \mathcal{A}_{*,{\mathrm{c}}}$. For any $\boldsymbol{\psi}\in\boldsymbol{\Psi}$, we have $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle S} }_*}=(\boldsymbol{\theta}_{{ \mathcal{\scriptscriptstyle P}}_{\sharp}}^{ \mathrm{\scriptscriptstyle T} },\boldsymbol{\xi}^{ \mathrm{\scriptscriptstyle T} }_{{{ \mathcal{\scriptscriptstyle A} }}_{*,{\mathrm{c}}}})^{ \mathrm{\scriptscriptstyle T} }$.
Proposition A.2 in the supplementary materials shows that there exists a sparse local minimizer $(\hat\boldsymbol{\theta}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} },\hat\boldsymbol{\xi}_{{ \mathrm{\scriptscriptstyle PEL} }}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }\in\boldsymbol{\Psi}$ for the nonconvex optimization (ref) such that $\mathbb{P}(\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle P}}^{{\mathrm{c}}}_{\sharp}}={\mathbf 0})\to1$ as $n\to\infty$, which means all zero components of $\boldsymbol{\theta}_0$ can be estimated exactly as zero with high probability. The limiting distribution of such $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle P}}_{\sharp}}$ and the moment selection outcomes are stated in Theorem (ref) below.
Compared to (ref), a penalty on $\boldsymbol{\theta}$ is added to (ref). To appreciate the benefit from this additional penalty on the sparse parameter, without loss of generality, we write $\boldsymbol{\theta}_{0}=(\boldsymbol{\theta}_{0,{ \mathcal{\scriptscriptstyle P}}_{\sharp}}^{ \mathrm{\scriptscriptstyle T} },{\mathbf 0}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} }$ and the inverse of ${\mathbf J}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}$ in Theorem (ref) in the following compatible partitioned matrix
where $[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{11}$ is a $p_{\sharp}\times p_{\sharp}$ matrix. The asymptotic covariance of $\hat{\boldsymbol{\theta}}_{{ \mathrm{\scriptscriptstyle PEL} },{ \mathcal{\scriptscriptstyle P}}_{\sharp}}$ by (ref) is $[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{11}$, and that in (ref) is $\{{\mathbf W}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}\}^{-1}$ according to Theorem (ref). Notice that $\{{\mathbf W}^{{ \mathcal{\scriptscriptstyle (T)} }}_{{ \mathcal{\scriptscriptstyle I} }^*}\}^{-1}=[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{11} -[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{12}[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{22}^{-1}[\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{21} \le [\{{\mathbf J}_{{ \mathcal{\scriptscriptstyle I} }^*}^{{ \mathcal{\scriptscriptstyle (T)} }}\}^{-1}]_{11}$ by the inverse of a partitioned matrix. Implicitly restricting the parameter space, the penalty on $\boldsymbol{\theta}$ gains efficiency for the estimation of $\boldsymbol{\theta}_{0,{ \mathcal{\scriptscriptstyle P}}_{\sharp}}$, the nonzero components of $\boldsymbol{\theta}_0$.
Statistical inference is important in applied econometrics when researchers intend to assess whether the estimated result supports or rejects a hypothesized value of the parameter. The next section proposes an inferential procedure based on a projection of estimating functions, which is free of asymptotic biases.
Suppose we are interested in the inference for a subset of parameters $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$ for some generic small subset $\mathcal{M} \subset (\mathcal{P} \cup \mathcal{D}) $ with $|\mathcal{M}|=m$. Here we allow $m$ to be fixed or diverge slowly with the sample size $n$. What is novel here is that $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$ is allowed to contain part of the auxiliary parameter $\boldsymbol{\xi}$, for which our method will provide a formal statistical inference for the validity of a subset of the high-dimensional moment restrictions. In contrast, liao2013adaptive offers asymptotic normality for the structural parameter $\boldsymbol{\theta}$ under low-dimensional moments but does not characterize the asymptotic distribution of the auxiliary parameter $\boldsymbol{\xi}$.
A key technical issue is how to deal with the high-dimensional nuisance parameter $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{\mathrm{c}}}$, where $\mathcal{M}^{{\mathrm{c}}}= (\mathcal{P} \cup \mathcal{D}) \backslash \mathcal{M}$. Our approach follows Chang2020 to project out the influence of the nuisance parameter. Such idea was also used in ning2017general and neykov2018unified. Given an initial estimate $\boldsymbol{\psi}^*$ for $\boldsymbol{\psi}_0$, we first determine a linear transformation matrix ${\mathbf A}_n = ({\mathbf a}_k^n)_{k\in\mathcal{M}}^{{ \mathrm{\scriptscriptstyle T} }} \in \mathbb{R}^{m\times r}$ with each row defined as
for $k \in \mathcal{M}$, where $\varsigma\rightarrow0$ as $n\rightarrow\infty$ is a tuning parameter, and $\{\boldsymbol{\gamma}_k\}_{k\in \mathcal{M}}$ is a basis of the linear space $\{ {\mathbf b}=(b_j)_{j\in \mathcal{P}\cup\mathcal{D} }: {\mathbf b}_{\mathcal{M}^{{\mathrm{c}}}} = \mathbf{0} \}$. In practice, we can specify the $(p+r_2)$-dimensional vector $\boldsymbol{\gamma}_k = ( 1(j=k))_{j \in \mathcal{P} \cup \mathcal{D} }$ with all zero elements except one unit entry. As we will discuss in Remark (ref), the solution from (ref) or (ref) can serve as the initial estimator $\boldsymbol{\psi}^*$.
Based on ${\mathbf A}_n$ as in ((ref)), we then obtain the new $m$-dimensional estimating functions ${\mathbf f}^{{\mathbf A}_n}(\cdot;\cdot)$ by projecting ${\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}(\cdot;\cdot)$ on ${\mathbf A}_n$:
Write $\boldsymbol{\gamma}_k=(\gamma_{k,j})_{j\in \mathcal{P}\cup\mathcal{D}}$ and define an $m\times(p+r_2)$ matrix $\boldsymbol{\Gamma}=(\gamma_{k,j})_{k\in { \mathcal{\scriptscriptstyle M} }, \, j \in \mathcal{P} \cup \mathcal{D}}$. The definition of ${\mathbf A}_n$ implies that $|\nabla_{\boldsymbol{\psi}}\bar{{\mathbf f}}^{{\mathbf A}_n}(\boldsymbol{\psi}^*)-\boldsymbol{\Gamma}|_\infty\leq \varsigma$. Since all the components in the $j$-th column of $\boldsymbol{\Gamma}$ are zero for $j\in\mathcal{M}^{{\mathrm{c}}}$, the newly defined estimating functions ${\mathbf f}^{{\mathbf A}_n}$ is uninformative of the nuisance parameter $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}$. Informative is ${\mathbf f}^{{\mathbf A}_n}$ of the parameter $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$ of interest, as for $j\in\mathcal{M}$ some components in the $j$-th column of $\boldsymbol{\Gamma}$ must be nonzero.
Given the projected estimating functions ${\mathbf f}^{{\mathbf A}_n}$, it is possible to directly borrow from Chang2020 to construct the confidence region of $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$ by the EL ratio \[ w_n(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }})=2\max_{\boldsymbol{\lambda}\in\tilde{\Lambda}_n(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }})}\sum_{i=1}^n\log\{1+\boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} }{\mathbf f}_i^{{\mathbf A}_n}(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }},\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^*)\} \] with respect to $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$, where $\tilde{\Lambda}_n(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }})=\{\boldsymbol{\lambda} \in\mathbb{R}^m : \boldsymbol{\lambda}^{ \mathrm{\scriptscriptstyle T} } {\mathbf f}^{{\mathbf A}_n}_i(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }},\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^*) \in \mathcal{V}~\textrm{for any}~i\in[n]\}$. Since $w_n(\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle M} }}) \xrightarrow{d}\chi_m^2$ as $n\rightarrow\infty$ for a fixed $m$, the set $\{\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}\in\mathbb{R}^m:w_n(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }})\leq \chi_{m,1-\alpha}^2\}$ provides a $100(1-\alpha)\%$ confidence region for $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$, where $\chi_{m,1-\alpha}^2$ is the $(1-\alpha)$-quantile of $\chi_m^2$ distribution. Nonetheless, the finite-sample performance of such an asymptotically valid confidence region depends crucially on the convexity of $w_n(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }})$ near $\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle M} }}$. If convexity fails in a finite sample, this EL-ratio-based confidence region will be voluminous in magnitude due to numerical instability. Verifying the convexity condition is onerous, in particular when $m$ is large, for ${\mathbf f}_i^{{\mathbf A}_n}(\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }},\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }^{{\mathrm{c}}}}^*)$ may be a nonlinear function of $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$.
To secure a stable confidence region for $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}$, in this paper we deviate from Chang2020 and instead recommend a re-estimation procedure for a confidence region based on the asymptotic normality of the estimator. Let
where $\boldsymbol{\Psi}^*_{ \mathcal{\scriptscriptstyle M} } = \{\boldsymbol{\psi}_{ \mathcal{\scriptscriptstyle M} }: |\boldsymbol{\psi}_{ \mathcal{\scriptscriptstyle M} } - \boldsymbol{\psi}^*_{ \mathcal{\scriptscriptstyle M} } |_1 \le O_{ \mathrm{p} }(\varpi_{1,n}) \}$ for some $\varpi_{1,n}\rightarrow0$ such that $|\boldsymbol{\psi}^*_{ \mathcal{\scriptscriptstyle M} }-\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle M} }}|_1 = O_{ \mathrm{p} }(\varpi_{1,n})$. Since $\boldsymbol{\psi}_{ \mathcal{\scriptscriptstyle M} }^*$ is an initial consistent estimator of $\boldsymbol{\psi}_{0,{ \mathcal{\scriptscriptstyle M} }}$, it is reasonable to search for $\tilde{\boldsymbol{\psi}}_{{ \mathcal{\scriptscriptstyle M} }}$ in a small neighborhood of $\boldsymbol{\psi}^*_{ \mathcal{\scriptscriptstyle M} }$. We will specify $\varpi_{1,n}$ for some specific choices of $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}^*$ in Remark (ref).
To derive the limiting distribution of $\tilde\boldsymbol{\psi}_{ \mathcal{\scriptscriptstyle M} }$, we assume the following condition.
\addtocounter{as}{5}
Theorem (ref) gives the limiting distribution of $\tilde{\boldsymbol{\psi}}_{{ \mathcal{\scriptscriptstyle M} }}$ with a generic initial estimator $\boldsymbol{\psi}^*$.
The above theorem is stated for $\tilde{\boldsymbol{\psi}}_{{ \mathcal{\scriptscriptstyle M} }} $. It includes the estimation of $\boldsymbol{\theta}_{{ \mathcal{\scriptscriptstyle M} }}$ as a special case if one's research interest falls on the main parameter for economic interpretation, and makes it possible to infer the validity of a subset of moments in view of selecting $\boldsymbol{\psi}_{{ \mathcal{\scriptscriptstyle M} }}=\boldsymbol{\xi}_{{ \mathcal{\scriptscriptstyle M} }}$.
So far, we have established statistical inference results based on the PPEL estimator $\tilde{\boldsymbol{\psi}}_{{ \mathcal{\scriptscriptstyle M} }}$ when we are interested in a small subset $\mathcal{M}$ of $\boldsymbol{\psi}$. In the next section, we check the finite sample performance via simulations.
One of the most important models in econometrics is the linear IV regression. We design a linear IV model here to mimic the empirical application in Section (ref), whereas simulation results of a nonlinear panel regression with time-varying individual heterogeneity are presented in the supplementary materials.
Suppose that the researcher has at hand a dataset of $n$ independent observations of a vector $\left(y_{i},x_{i},{\mathbf z}_{i},w_{1i},{\mathbf w}_{2i},{\mathbf w}_{3i}\right)$, and is interested in estimating the main structural equation
where $x_{i}$ is a scalar endogenous variable, and ${\mathbf z}_{i}$ is a $d_{z}\times1$ vector of exogenous variables (including the intercept). Such an equation with a scalar endogenous variable is the leading case of IV regressions andrews2019weak. Due to space limitations, we report the results when we specify ${\mathbf z}_i=(1,z_{1,i},z_{2,i})^{ \mathrm{\scriptscriptstyle T} }\in\mathbb{R}^3$ with $( z_{1,i}, z_{2,i} )^{ \mathrm{\scriptscriptstyle T} } \sim \mathcal{N}({\mathbf 0},\rm{{\mathbf I}}_2)$, and the coefficients $ (\beta_x, \boldsymbol{\beta}^{ \mathrm{\scriptscriptstyle T} }_{\mathbf z})^{ \mathrm{\scriptscriptstyle T} }=(0.5,0.5,0.5,0.5)^{ \mathrm{\scriptscriptstyle T} }$ in a plausible setting. The numerical performance are robust across our experiments when the parameters are varied.
The corresponding reduced-form equation is $ x_{i}=\gamma_{w1}w_{1i}+\boldsymbol{\gamma}^{ \mathrm{\scriptscriptstyle T} }_{{\mathbf w}2}{\mathbf w}_{2i}+\boldsymbol{\gamma}_{{\mathbf z}}^{ \mathrm{\scriptscriptstyle T} }{\mathbf z}_{i}+u_{i}$, where $w_{1i}\in\mathbb{R}$ and ${\mathbf w}_{2i}\in\mathbb{R}^{d_{{\mathbf w}2}}$ are excluded instruments. We specify $ (w_{1i}, {\mathbf w}_{2i}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} } \sim \mathcal{N}({\mathbf 0},{\rm\bf I}_{d_{{\mathbf w}2}+1})$, $ \gamma_{w1}=0.8$, $\boldsymbol{\gamma}_{{\mathbf w}2}=(\gamma_{{\mathbf w}2,j})_{j\in[d_{{\mathbf w}2}]}$ with $\gamma_{{\mathbf w}2,j}=0.4-0.3(j-1)/(d_{{\mathbf w}2}-1)$, and $\boldsymbol{\gamma}_{\mathbf z}=(0.8,0.8,0.8)^{ \mathrm{\scriptscriptstyle T} }$. The two error terms $(\epsilon_i,u_i)^{{ \mathrm{\scriptscriptstyle T} }}$ are generated from the two-dimensional normal distribution with mean $0$, covariance $1$ and correlation $0.5$ which are independent of $\left({\mathbf z}_{i},w_{1i},{\mathbf w}_{2i}\right)$. An additional vector ${\mathbf w}_{3i}=(w_{3i,j})_{j\in[s]}\in\mathbb{R}^{s}$, which serves as the invalid IV, follows $ w_{3i,j}=\delta_j\epsilon_{i}+v_{i}$, where $v_{i}\sim \mathcal{N}\left(0,1\right)$ is independent of $\left(\epsilon_{i},u_{i}\right)$, and $\delta_j \neq 0$ controls the strength of correlation. Write $\boldsymbol{\delta}=(\delta_j)_{j\in[s]}$. To emulate the scenarios of weak, moderate and strong correlations between $\epsilon_i$ and ${\mathbf w}_{3i}$, we set $\delta_j=0.3+0.2(j-1)/(s-1)$, $0.5+0.2(j-1)/(s-1)$ and $0.7+0.2(j-1)/(s-1)$, respectively. It is expected that the smaller is the coefficient, the less severe is misspecification, so estimation is more prone to selection mistakes in finite samples.
While the researcher is confident about the validity of $w_{1i}$ in that $ {\mathbf g}_i^{{ \mathcal{\scriptscriptstyle (I)} }} (\boldsymbol{\theta})= (w_{1i}, {\mathbf z}_{i}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} } \times \left(y_{i}-\beta_{x}x_{i}-\boldsymbol{\beta}_{{\mathbf z}}^{ \mathrm{\scriptscriptstyle T} }{\mathbf z}_{i}\right), $ where ${\mathbf z}_i$ is self-instrumented, she is uncertain about the validity of $({\mathbf w}_{2i}, {\mathbf w}_{3i})$ and therefore must detect the invalid IVs. The two classes of undetermined (to the researcher) IVs consist of $ {\mathbf g}_i^{{ \mathcal{\scriptscriptstyle (D)} }} (\boldsymbol{\theta})= ({\mathbf w}_{2i}^{ \mathrm{\scriptscriptstyle T} }, {\mathbf w}_{3i}^{ \mathrm{\scriptscriptstyle T} })^{ \mathrm{\scriptscriptstyle T} } \times \left(y_{i}-\beta_{x}x_{i}-\boldsymbol{\beta}_{{\mathbf z}}^{ \mathrm{\scriptscriptstyle T} }{\mathbf z}_{i}\right). $ According to our DGP design, ${\mathbf w}_{2i}$ is the valid IV vector so the moments involving ${\mathbf w}_{2i}$ equal to zero under the true parameter; those moments involving ${\mathbf w}_{3i}$ are invalid.
Let $d_w =d_{{\mathbf w}2}+s+1$ be the number of all the IVs, among which $s$ IVs are invalid. In the low-dimensional setting, we fix $s=6$ and consider $(n,d_w)=(100,50)$ and $(200,100)$. In the high-dimensional setting, we vary the number of invalid IVs to be $s\in\{6,8,13\}$ for $(n,d_w)=(100,120)$, and $s\in\{6,12,17\}$ for $(n,d_w)=(200,240)$, where $s$ is specified by rounding, in addition, $2\log n$ and $3 n^{1/3}$.
In terms of the numerical implementation, we compute the PEL estimates by the modified two-layer coordinate descent algorithm ChangTangWu2018. The SCAD penalty is used for both $P_{1,\pi}(\cdot)$ and $P_{2,\nu}(\cdot)$ in (ref) for all the numerical experiments in this paper with local quadratic approximation FanLi2001, and the tuning parameters $\nu$ and $\pi$ are chosen by the Bayesian information criterion (BIC). Specifically, we use the following BIC type function:
where $\mathrm{df}(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})$ denotes the number of nonzero elements in $\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }}$, and the log likelihood term is the EL ratio $ \ell(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})=2\sum_{i=1}^n\log\{1+\hat\boldsymbol{\lambda}(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})^{ \mathrm{\scriptscriptstyle T} }{\mathbf g}^{{ \mathcal{\scriptscriptstyle (T)} }}_i(\hat{\boldsymbol{\psi}}_{{ \mathrm{\scriptscriptstyle PEL} }})\}$.\footnote{leng2012penalized employ this BIC to choose the tuning parameter for the high-dimensional parameter $\boldsymbol{\psi}$, and ChangTangWu2018 apply it to select multiple tuning parameters. We follow their practice as (ref) remains valid in our setting where two tuning parameters are included. }
We first report results of moment selection by our selection criterion in (ref). Let “FP” (false positive) denote the frequency that the valid moments being not selected, and “FN” (false negative) denote the frequency that the invalid moments being selected. In Table (ref), PEL and DB-PEL denote the moment selection criterion based on the PEL estimator and its de-biased version, respectively. As expected, the strength of correlations between $\epsilon_i$ and ${\mathbf w}_{3i}$ does not affect FP, whereas FN quickly vanishes as the correlations get stronger. In all cases, larger sample sizes help reduce the chance of moment selection error.
The parameter of interest in the linear IV model is $\beta_{x}$ in (ref) as it characterizes the “causal effect” of the endogenous variable, which bears economic interpretation. The root-mean-square error (RMSE), bias (BIAS) and standard deviation (STD) are calculated for PEL, DB-PEL and the classical two-stage least squares (2SLS) for (ref), and an oracle estimator is added for comparison.\footnote{ The oracle EL in Section (ref) is for low-dimensional parameters and moments. To handle the large pool of orthogonal IVs, the oracle estimator here is a 2SLS estimator taking advantage of a few most relevant IVs---those with the top $0.1n$ big coefficients $\gamma_{\mathbf{w}2,j}$ in the reduced-form equation. } The results are summarized in Table (ref). The performance of PEL and the oracle is comparable, and the gaps are narrowed when the sample size is increased from $n=100$ to $n=200$, suggesting the capacity for PEL to mimic the oracle by absorbing the signal from the valid moments and in the meantime keeping the invalid ones at bay. The RMSE of 2SLS is significantly larger than those of PEL and DB-PEL.
The left panel of Figure (ref) plots the empirical cumulative distribution functions (ECDF) of the DB-PEL estimates. The dotted curve corresponds to the case of $(n,d_w,s)=(100, 120, 8)$, the dashed curve to that of $(n,d_w,s)=(200,240,12)$, and the solid curve is the cumulative distribution function (CDF) of $\mathcal{N}(0,1)$ for comparison. A better normal approximation can be obtained by the PPEL estimate, as shown in the right panel of Figure (ref), with its tuning parameter $\varsigma=0.08(n^{-1}\log p)^{1/2}$. We also present the confidence intervals (CI) according to PPEL, DB-PEL and 2SLS, as reported in Table (ref), for the 90%, 95% and 99% levels, where the CIs for DB-PEL and PPEL are predicated on the asymptotic normality from Theorems (ref) and (ref), respectively, and the CIs for 2SLS are based on its asymptotic normality as in standard textbooks. PPEL's coverage probability to the nominal counterpart is the best among the three estimators, and is much better than PEL. 2SLS's coverage probability seems too high in the high-dimensional setting, though it is acceptable in the low-dimensional case. Figure (ref) shows that the width of 2SLS's CI is much wider than that of PPEL, which is caused by efficiency loss from abandoning the potentially valid estimating equations.
Colonialism was widespread over the globe prior to the Second World War. While European institutions which protected private properties and checked government powers were successfully replicated in a few colonies, many others suffered expropriation of natural and human resources. In an influential study, Acemoglu2001 (AJR, henceforth) systematically explored the relationship between Europeans' mortality rates and the modes of institutions.
We revisit AJR's open-access dataset. AJR's main structural equation is $ y_i = \gamma_y + \beta_x x_i+ \beta_z z_i + \epsilon_i$, where the dependent variable $y_i$ is {\tt logarithm GDP per capita in 1995}, and the key explanatory variable of interest $x_i$ is {\tt average protection against expropriation risk}, a continuous variable of a scale between 0 and 10. The {\tt latitude} of a colony, denoted as $z_i$, is an additional control variable. For substantial measurement errors in the institutional index $x_i$ can bias the OLS estimator, the credibility of the empirical evidence counts on IVs.
AJR compiled from historical documents a seminal variable {\tt logarithm of European settler mortality}, and argued forcefully that it qualifies as a valid IV for the endogenous variable $x_i$. Their empirical evidence was based on the 2SLS regression in cross-sectional data of countries. The baseline result in AJR's Column (2) of Table 4 (p.1386) corresponds to the three estimating functions $ {\mathbf g}^{{ \mathcal{\scriptscriptstyle (I)} }}_i(\boldsymbol{\theta})= (w_{1i}, z_i, 1)^{{ \mathrm{\scriptscriptstyle T} }} \times (y_i -\gamma_y - \beta_x x_i - \beta_z z_i)$ by the notations in our simulation, and the 2SLS reports $\hat{\beta}_x = 1.00$ with STD 0.22.
There are another 11 potential health variables and institutional variables, which serve as potential IVs. AJR were uncertain about their validity, and they experimented in their Table 7 (p.1392) and Table 8 (p.1394) the empirical results under various IV configurations. These variables, denoted as ${\mathbf w}_2$, are (i) {\tt Malaria in 1994}, (ii) {\tt Yellow fever}, (iii) {\tt Life expectancy}, (iv) {\tt Infant mortality}, (v) {\tt Mean temperature}, (vi) {\tt Distance from coast}, (vii) {\tt European settlements in 1990}, (viii) {\tt Democracy (1st year of independence)}, (ix) {\tt Constraint on executive (1st year of independence)}, (x) {\tt Democracy in 1900}, and (xi) {\tt Constraint on executive in 1900}. (i)--(iv) are health variables, (v)--(vi) are geographic variables, and (vii)--(xi) are institutional variables. All these extra IVs are associated with the estimating functions $ {\mathbf g}^{{ \mathcal{\scriptscriptstyle (D)} }}_i (\boldsymbol{\theta})={\mathbf w}_{2i} \times (y_i-\gamma_y - \beta_x x_i - \beta_z z_i) \, . $ Given the moderate sample size of 56 countries, a total of 12 IVs is non-trivial.
Our PPEL estimates the main equation as \[ \widehat{\mbox{log GDP per capita}} = \underset{(1.194)}{2.048} + \underset{(0.126)}{0.945} \times \mathrm{institution} - \underset{(0.977)}{0.785} \times \mbox {lattitude}\,. \] For the estimation and inference of the key parameter of interest $\beta_x$, we compare the results of PEL, DB-PEL, PPEL and the conventional 2SLS. Its point estimation (PE), STD, and CIs are presented in Table (ref). Enjoying the efficiency gain from the extra IVs, the STD of PPEL is 0.126 when we select the tuning parameter $\varsigma=\zeta_c(n^{-1}\log p)^{1/2}$ with $\zeta_c=0.08$ (the boldface row), as suggested by our simulations. This STD achieves a 37% reduction relative to that of the plain 2SLS with the sole valid IV. Moreover, when we vary the constant $\zeta_c=0.08$ in $\varsigma$ as 0.04, 0.06, 0.12 and 0.16, PPEL is rather robust over the wide range of tuning parameters.
Our method is particularly important in unifying the IV selection in AJR's Tables 7 and 8 into a single set of automatically selected IVs. Among the 11 potential IVs, our moment selection criterion (ref) invalidates all institutional variables. Five IVs survive the testing: the climate variable {\tt mean temperature}, the geographic variable {\tt distance from coast}, and three health variables {\tt dummy of yellow fever}, {\tt infant mortality} and {\tt life expectancy}. All the other health and institutional variables are assessed as endogenous and are unsuitable for IVs in this study. The justification for {\tt mean temperature} and {\tt distance from coast} is straightforward because humans were unable to interfere with these natural conditions in the era of colonialism. Furthermore, AJR argued for the validity of {\tt dummy of yellow fever}, though due to concerns of lack of variation they did not employ it as the main IV (AJR's p.1393, Paragraph 2). Our variable selection result provides supportive evidence to AJR's heuristics.
This paper considers a general setting of an economic structural model with many potential moments, some of which may be invalid. These invalid moments must be disciplined in order to estimate the structural parameter consistently. We propose a PEL approach to estimate the parameter of interest while coping with the invalid moments. We show that the PEL estimator is normally distributed asymptotically, and invalid moments can be consistently detected thanks to the oracle property. To overcome the difficulty of estimating the bias in the limiting distribution of the PEL estimator, we further devise the PPEL approach for statistical inference of a low-dimensional object of interest, which is useful for hypothesis testing and confidence region construction. Simulation exercises are carried out to demonstrate excellent finite sample performance of our methods. We revisit an empirical application concerning economic development and shed new insight about its candidate instruments.