The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
107,940 characters
\begin{center}
\bigskip
{\huge Valid Post-Selection and Post-Regularization Inference: An Elementary, General Approach}\\
\end{center}
\title[Post-Selection and Post-Regularization Inference]{}
\author{Victor Chernozhukov, Christian Hansen, and Martin Spindler}\thanks{Chernozhukov: Massachussets Institute of Technology, 50 Memorial Drive, E52-361B, Cambridge, MA 02142, [email removed].
Hansen: University of Chicago Booth School of Business, 5807 S. Woodlawn Ave., Chicago, IL 60637, [email removed]. Spindler: Munich Center for the Economics of Aging, Amalienstr. 33, 80799 Munich, Germany, [email removed]. }
\date{December, 2014. We thank Denis Chetverikov, Mert Demirer, Anna Mikusheva, seminar participants and the discussant Susan Athey at
the AEA Session on Machine Learning in Economics and Econometrics, CEME conference on Non-Standard Problems in Econometrics,
Berlin Statistics Seminar, and the students from MIT's 14.387 Applied Econometrics Class for useful comments.}
\maketitle
\begin{footnotesize}
\textbf{Abstract.} We present an expository, general analysis of valid post-selection or post-regularization inference about a low-dimensional target parameter in the presence of a very high-dimensional nuisance parameter that is estimated using selection or regularization methods. Our analysis provides a set of high-level conditions under which inference for the low-dimensional parameter based on testing or point estimation methods will be regular despite selection or regularization biases occurring in the estimation of the high-dimensional nuisance parameter. The results may be applied to establish uniform validity of post-selection or post-regularization inference procedures for low-dimensional target parameters over large classes of models. The high-level conditions allow one to clearly see the types of structure needed to achieve valid post-regularization inference and encompass many existing results. A key element of the structure we employ and discuss in detail is the use of so-called orthogonal or ``immunized'' estimating equations that are locally insensitive to small mistakes in estimation of the high-dimensional nuisance parameter. As an illustration, we use the high-level conditions to provide readily verifiable sufficient conditions for a class of affine-quadratic models that include the usual linear model and linear instrumental variables model as special cases. As a further application and illustration, we use these results to provide an analysis of post-selection inference in a linear instrumental variables model with many regressors and many instruments. We conclude with a review of other developments in post-selection inference and note that many of the developments can be viewed as special cases of the general encompassing framework of orthogonal estimating equations provided in this paper.
\textbf{Key words:} Neyman, orthogonalization, $C(\alpha)$ statistics, optimal instrument, optimal score, optimal moment, post-selection and post-regularization inference, efficiency, optimality
\end{footnotesize}
\section{Introduction}
Analysis of high-dimensional models, models in which the number of parameters to be estimated is large relative to the sample size, is becoming increasingly important. Such models arise naturally in readily available high-dimensional data which have many measured characteristics available per individual observation as in, for example, large survey data sets, scanner data, and text data. Such models also arise naturally even in data with a small number of measured characteristics in situations where the exact functional form with which the observed variables enter the model is unknown. Examples of this scenario include semiparametric models with nonparametric nuisance functions. More generally, models with many parameters relative to the sample size often arise when attempting to model complex phenomena.
The key concept underlying the analysis of high-dimensional models is that regularization, such as model selection or shrinkage of model parameters, is necessary if one is to draw meaningful conclusions from the data. For example, the need for regularization is obvious in a linear regression model with the number of right-hand-side variables greater than the sample size, but arises far more generally in any setting in which the number of parameters is not small relative to the sample size. Given the importance of the use of regularization in analyzing high-dimensional models, it is then important to explicitly account for the impact of this regularization on the behavior of estimators if one wishes to accurately characterize their finite-sample behavior. The use of such regularization techniques may easily invalidate conventional approaches to inference about model parameters and other interesting target parameters. A major goal of this paper is to present a general, formal framework that provides guidance about setting up estimating equations and making appropriate use of regularization devices so that inference about parameters of interest will remain valid in the presence of data-dependent model selection or other approaches to regularization.
It is important to note that understanding estimators' behavior in high-dimensional settings is also useful in conventional low-dimensional settings. As noted above, dealing formally with high-dimensional models requires that one explicitly accounts for model selection or other forms of regularization. Providing results that explicitly account for this regularization then allows us to accommodate and coherently account for the fact that low-dimensional models estimated in practice are often the result of specification searches. As in the high-dimensional setting, failure to account for this variable selection will invalidate the usual inference procedures, whereas the approach that we outline will remain valid and can easily be applied in conventional low-dimensional settings.
The chief goal of this overview paper is to offer a general framework that encompasses many existing results regarding inference on model parameters in high-dimensional models. The encompassing framework we present and the key theoretical results are new, although they are clearly heavily influenced and foreshadowed by previous, more specialized results. As an application of the framework, we also present new results on inference in a reasonably broad class of models, termed affine-quadratic models, that includes the usual linear model and linear instrumental variables (IV) model and then apply these results to provide new ones regarding post-regularization inference on the parameters on endogenous variables in a linear instrumental variables model with very many instruments and controls (and also allowing for some misspecification). We also provide a discussion of previous research that aims to highlight that many existing results fall within the general framework.
Formally, we present a series of results for obtaining valid inferential statements about a low-dimensional parameter of interest, $\alpha$, in the presence of a high-dimensional nuisance parameter $\eta$. The general approach we offer relies on two fundamental elements. First, it is important that estimating equations used to draw inferences about $\alpha$ satisfy a key orthogonality or immunization condition.\footnote{We refer to the condition as an orthogonality or immunization condition as orthogonality is a much used term and our usage differs from some other usage in defining orthogonality conditions used in econometrics.} For example, when estimation and inference for $\alpha$ are based on the empirical analog of a theoretical system of equations
$$\mathrm{M}(\alpha, \eta)=0,$$
we show that setting up the equations in a manner such that the orthogonality or immunization condition
$$\partial_\eta \mathrm{M}(\alpha, \eta) = 0$$
holds is an important element in providing an inferential procedure for $\alpha$ that remains valid when $\eta$ is estimated using regularization. We note that this condition can generally be established. For example, we can apply Neyman's classic orthogonalized score in likelihood settings; see, e.g. \cite{Neyman59} and \cite{Neyman1979}. We also describe an extension of this classic approach to the GMM setting. In general, applying this orthogonalization will introduce additional nuisance parameters that will be treated as part of $\eta$.
The second key element of our approach is the use of high-quality, structured estimators of $\eta$. Crucially, additional structure on $\eta$ is needed for informative inference to proceed, and it is thus important to use estimation strategies that leverage and perform well under the desired structure. An example of a structure that has been usefully employed in the recent literature is approximate sparsity, e.g. \cite{BCCH12}. Within this framework, $\eta$ is well approximated by a sparse vector which suggests the use of a sparse estimator such as the Lasso (\cite{FF:1993} and \cite{T1996}).
The Lasso estimator solves the general problem
$$
\widehat\eta_L = \arg\min_{\eta } \ \ell(\mathrm{data}, \eta) + \lambda\sum_{j=1}^{p} |\psi_j\eta_j|,
$$
where $\ell(\mathrm{data}, \eta)$ is some general loss function that depends on the data and the parameter $\eta$, $\lambda$ is a penalty level, and $\psi_j$'s are penalty loadings. The leading example is the usual linear model in which $\ell(\mathrm{data},\eta) = \sum_{i=1}^{n} (y_i - x_i'\eta)^2$ is the usual least-squares loss, with $y_i$ denoting the outcome of interest for observation $i$ and $x_i$ denoting predictor variables, and we provide further discussion of this example in the appendix. Other examples of $\ell(\mathrm{data}, \eta)$ include suitable loss functions corresponding to well-known M-estimators, the negative of the log-likelihood, and GMM criterion functions. This estimator and related methods such as those in \cite{CandesTao2007}, \cite{MY2007}, \cite{BickelRitovTsybakov2009}, \cite{BC-PostLASSO}, and \cite{BCW-SqLASSO} are computationally efficient and have been shown to have good estimation properties even when perfect variable selection is not feasible under approximate sparsity. These good estimation properties then translate into providing ``good enough'' estimates of $\eta$ to result in valid inference about $\alpha$ when coupled with orthogonal estimating equations as discussed above.
Finally, it is important to note that the general results we present do not require or leverage approximate sparsity or sparsity-based estimation strategies. We provide this discussion here simply as an example and because the structure offers one concrete setting in which the general results we establish may be applied.
In the remainder of this paper, we present the main results. In Sections 2 and 3, we provide our general set of results that may be used to establish uniform validity of inference about low-dimensional parameters of interest in the presence of high-dimensional nuisance parameters. We provide the framework in Section 2, and then discuss how to achieve the key orthogonality condition in Section 3. In Sections 4 and 5, we provide details about establishing the necessary results for the estimation quality of $\eta$ within the approximately sparse framework. The analysis in Section 4 pertains to a reasonably general class of affine-quadratic models, and the analysis of Section 5 specializes this result to the case of estimating the parameters on a vector of endogenous variables in a linear instrumental variables model with very many potential control variables and very many potential instruments. The analysis in Section 5 thus extends results from \cite{BCCH12} and \cite{BelloniChernozhukovHansen2011}. We also provide a brief simulation example and an empirical example that looks at logit demand estimation within the linear many instrument and many control setting in Section 5. We conclude with a literature review in Section 6.
\textbf{Notation.} We use ``wp $\to 1$" to abbreviate the phrase ``with probability
that converges to 1", and we use the arrows $\to_{{\mathrm{P}}_n}$ and $\leadsto_{{\mathrm{P}}_n}$ to denote convergence in
probability and in distribution under the sequence of probability measures $\{{\mathrm{P}}_n\}$. The symbol $\sim$
means ``distributed as". The notation $a \lesssim b$ means that $a = O(b)$ and $a \lesssim_{{\mathrm{P}}_n} b$ means that $a = O_{{\mathrm{P}}_n}(b)$.
The $\ell_{2}$ and $\ell_{1}$ norms are denoted by
$\|\cdot\|$ and $\| \cdot \|_{1}$, respectively; and the $\ell_{0}$-``norm", $\|\cdot\|_0$, denotes the number of non-zero components of a vector. When applied to a matrix, $\|\cdot\|$ denotes the operator norm.
We use the notation $a \vee b = \max( a, b)$ and $a \wedge b = \min(a , b)$. Here and below, ${\mathbb{E}_n}[\cdot]$ abbreviates the average $n^{-1}\sum_{i=1}^n[\cdot]$ over index $i$. That is, ${\mathbb{E}_n}[f(w_i)]$ denotes $n^{-1}\sum_{i=1}^n[f(w_i)]$. In what follows, we use the $m$-sparse norm of a matrix $Q$ defined as
$$
\|Q\|_{\mathsf{sp}(m)} = \sup\{ | b'Qb|/\|b\|^2 : \|b\|_0 \leq m, \|b\| \neq 0\}.
$$
We also consider the pointwise norm of a square matrix matrix $Q$ at a point $x \neq 0$:
$$
\|Q\|_{\mathsf{pw}(x)} = |x'Qx|/\|x\|^2.
$$
For a differentiable map $x \mapsto f(x)$, mapping $\mathbb{R}^d$ to $\mathbb{R}^k$, we use $\partial_{x'} f$ to abbreviate the partial derivatives $(\partial/\partial x') f$, and we correspondingly use the expression
$\partial_{x'} f(x_0)$ to mean $\partial_{x'} f (x) \mid_{x = x_0}$, etc. We use $x'$ to denote the transpose of a column vector $x$.
\section{A Testing and Estimation Approach to Valid Post-Selection and Post-Regularization Inference}
\subsection{The Setting} We assume that estimation is based on the first $n$ elements $(w_{i,n})_{i=1}^n$ of the \textit{stationary} data-stream $(w_{i,n})_{i=1}^\infty$ which lives on the probability space $(\Omega, \mathcal{A}, {\mathrm{P}}_n)$. The data points
$w_{i,n}$ take values in a measurable space $\mathcal{W}$ for each $i$ and $n$. Here, ${\mathrm{P}}_n$, the probability law or data-generating process, can change with $n$. We allow the law to change with $n$ to claim robustness or uniform validity of results with respect to perturbations of such laws. Thus the data, all parameters, estimators, and other quantities are indexed by $n$, but we typically suppress this dependence to simplify notation.
The target parameter value $\alpha=\alpha_0$ is assumed to solve the system of theoretical equations
$$
{\mathrm{M}}(\alpha, \eta_0) = 0, $$
where $\mathrm{M}= (\mathrm{M}_l)_{l=1}^k$ is a measurable map from $\mathcal{A}\times \mathcal{H}$
to $\mathbb{R}^{k}$ and $\mathcal{A}\times \mathcal{H}$ are some convex subsets of $\mathbb{R}^d \times \mathbb{R}^p$. Here the dimension $d$ of the target parameter $\alpha \in \mathcal{A}$ and the number of equations $k$ are assumed to be fixed and
the dimension $p=p_n$ of the nuisance parameter $\eta \in \mathcal{H}$ is allowed to be very high, potentially much larger than $n$.
To handle the high-dimensional nuisance parameter $\eta$, we employ structured assumptions and selection or
regularization methods appropriate for the structure to estimate $\eta_0$.
Given an appropriate estimator $\hat \eta$, we can
construct an estimator $\hat \alpha$ as an approximate solution
to the estimating equation:
$$
\| \hat{\mathrm{M}}(\hat \alpha, \hat \eta)\| \leq \inf_{\alpha \in \mathcal{A}} \| \hat{\mathrm{M}}(\alpha, \hat \eta)\| + o(n^{-1/2})
$$
where $\hat{\mathrm{M}} = (\hat{\mathrm{M}}_l)_{l=1}^k$ is the empirical analog of theoretical equations $\mathrm{M}$, which is a measurable map from
$\mathcal{W}^n\times \mathcal{A}\times \mathcal{H}$ to $\mathbb{R}^k$.
We can also use $\hat{\mathrm{M}}(\alpha, \hat \eta)$ to test hypotheses about $\alpha_0$ and then
invert the tests to construct confidence sets.
It is not required in the formulation above, but a typical case is when $\hat \mathrm{M}$ and $\mathrm{M}$
are formed as theoretical and empirical moment functions:
$$
\quad {\mathrm{M}}(\alpha, \eta) := {\mathrm{E}}[\psi(w_i, \alpha, \eta)],
\quad \hat{\mathrm{M}}(\alpha, \eta) := {\mathbb{E}_n}[\psi(w_i, \alpha, \eta)],
$$
where $\psi = (\psi_l)_{l=1}^k$ is a measurable map from $\mathcal{W}\times \mathcal{A}\times \mathcal{H}$
to $\mathbb{R}^{k}$. Of course, there are many problems that do not fall in the moment condition framework.
\subsection{Valid Inference via Testing}
A simple introduction to the inferential problem is via the testing problem in which
we would like to test some hypothesis about the true parameter value $\alpha_0$.
By inverting the test,
we create a confidence set for $\alpha_0$. The key condition
for the validity of this confidence region is adaptivity, which can be ensured
by using orthogonal estimating equations and using structured assumptions on the high-dimensional
nuisance parameter.\footnote{We refer to \cite{Bickel1982} for a definition of and introduction to adaptivity.}
The key condition enabling us to perform valid inference on $\alpha_0$
is the \textit{adaptivity} condition:
\begin{equation}\label{eq:adaptivity}
\sqrt{n}(\hat{\mathrm{M}}(\alpha_0, \hat \eta) - \hat{\mathrm{M}}(\alpha_0, \eta_0)) \to_{{\mathrm{P}}_n} 0.
\end{equation}
This condition states that using $\sqrt{n}\hat{\mathrm{M}}(\alpha_0, \hat \eta)$ is as good as
using $\sqrt{n}\hat{\mathrm{M}}(\alpha_0, \eta_0)$, at least to the first order. This condition may hold
despite using estimators $\hat \eta$ that are not asymptotically
linear and are non-regular. Verification of adaptivity may involve substantial work as illustrated below.
A key requirement that often arises is the \textit{orthogonality} or \textit{immunization} condition:
\begin{equation}\label{eq:ortho}
\partial_{\eta'} {\mathrm{M}}(\alpha_0, \eta_0) = 0.
\end{equation}
This condition states that the equations are locally insensitive to small perturbations
of the nuisance parameter around the true parameter values. In several important models, this condition
is equivalent to the double-robustness condition
(\cite{robins:dr}). Additional assumptions regarding the \textit{quality of
estimation} of $\eta_0$ are also needed and are highlighted below.
The adaptivity condition immediately allows
us to use the statistic $\sqrt{n}\hat{\mathrm{M}}(\alpha_0, \hat \eta)$ to perform inference.
Indeed, suppose we have that
\begin{equation}\label{eq:normality}
\Omega^{-1/2}(\alpha_0)\sqrt{n}\hat{\mathrm{M}}(\alpha_0, \eta_0) \leadsto_{{\mathrm{P}}_n} \mathcal{N}(0,I_k)
\end{equation}
for some positive definite $\Omega(\alpha) = \operatorname{Var}( \sqrt{n}\hat{\mathrm{M}}(\alpha, \eta_0))$.
This condition can be verified using central limit theorems for triangular arrays. Such
theorems are available for independently and identically distributed (i.i.d.) as well as dependent and clustered data. Suppose further that there exists $\hat \Omega(\alpha)$ such that
\begin{equation}\label{eq:variance}
\hat \Omega^{-1/2}(\alpha_0) \Omega^{1/2}(\alpha_0) \to_{{\mathrm{P}}_n} I_k.
\end{equation}
It is then immediate that the following score statistic, evaluated at $\alpha = \alpha_0$, is asymptotically normal,
\begin{equation}\label{eq:normality2}
S(\alpha) := \hat \Omega^{-1/2}_n(\alpha)\sqrt{n}\hat{\mathrm{M}}(\alpha, \hat \eta) \leadsto_{{\mathrm{P}}_n} \mathcal{N}(0,I_k),
\end{equation}
and that the quadratic form of this score statistic is asymptotically $\chi^2$
with $k$ degrees of freedom:
\begin{equation}\label{eq:chi-square}
C(\alpha_0) = \| S(\alpha_0)\|^2 \leadsto_{{\mathrm{P}}_n} \chi^2(k).
\end{equation}
The statistic given in (\ref{eq:chi-square}) simply corresponds to a quadratic form in appropriately normalized statistics that have the desired immunization or orthogonality condition. We refer to this statistic as a ``generalized $C(\alpha)$-statistic'' in honor of Neyman's fundamental contributions, e.g. \cite{Neyman59} and \cite{Neyman1979}, because, in likelihood settings, the statistic (\ref{eq:chi-square}) reduces to Neyman's $C(\alpha)$-statistic and the generalized score $S(\alpha_0)$ given in (\ref{eq:normality2}) reduces to Neyman's orthogonalized score. We demonstrate these relationships in the special case of likelihood models in Section 3.1 and provide a generalization to GMM models in Section 3.2. Both of these examples serve to illustrate the construction of appropriate statistics in different settings, but we note that the framework applies far more generally.
The following elementary result is an immediate consequence of the preceding discussion. \\
\begin{proposition}[Valid Inference After Selection or Regularizaton] \textit{Consider a sequence $\{\mathbf{P}_n\}$ of sets
of probability laws such that for each sequence $\{{\mathrm{P}}_n\} \in \{\mathbf{P}_n\}$ the
adaptivity condition (\ref{eq:adaptivity}), the normality condition (\ref{eq:normality}), and the variance consistency condition (\ref{eq:variance}) hold.
Then $\mathsf{CR}_{1-a} = \{ \alpha \in \mathcal{A}: C(\alpha) \leq c(1-a) \}$, where
$c(1-a)$ is the $1-a$-quantile of a $\chi^2(k)$, is a uniformly valid confidence
interval for $\alpha_0$ in the sense that}
$$
\lim_{n \to \infty} \sup_{{\mathrm{P}} \in \mathbf{P}_n} |{\mathrm{P}} ( \alpha_0 \in \mathsf{CR}_{1-a} ) - (1-a) | = 0.
$$
\end{proposition}
We remark here that in order to make the uniformity claim interesting we should insist that
the sets of probability laws $\mathbf{P}_n$ are non-decreasing in $n$, i.e.
$\mathbf{P}_{\bar n} \subseteq \mathbf{P}_{n}$ whenever $\bar n \leq n$.
Proof. For any sequence of positive constants $\epsilon_n$ approaching $0$, let
${\mathrm{P}}_n \in \mathbf{P}_n$ be any sequence such that
$$
|{\mathrm{P}}_n ( \alpha_0 \in \mathsf{CR}_{1-a} ) - (1-a) | +\epsilon_n \geq \sup_{{\mathrm{P}} \in \mathbf{P}_n} |{\mathrm{P}} ( \alpha_0 \in \mathsf{CR}_{1-a} ) - (1-a) |.
$$
By conditions (\ref{eq:normality}) and (\ref{eq:variance}) we have that
$$
{\mathrm{P}}_n ( \alpha_0 \in \mathsf{CR}_{1-a} ) ={\mathrm{P}}_n ( C(\alpha_0) \leq c(1-a) ) \to \mathbb{P} (\chi^2(k) \leq c(1-a)) = 1-a,
$$
which implies the conclusion from the preceding display. \hfill{\tiny \ensuremath{\blacksquare} }
\subsection{Valid Inference via Adaptive Estimation} Suppose that $\mathrm{M}(\alpha_0, \eta_0) =0$ holds
for $\alpha_0 \in \mathcal{A}$. We consider an estimator $\hat \alpha \in \mathcal{A}$
that is an approximate minimizer of the map $\alpha \mapsto \| \hat{\mathrm{M}}(\alpha, \hat \eta)\|$ in the sense that
\begin{equation}\label{eq:minimization}
\| \hat \mathrm{M}(\hat \alpha, \hat \eta)\| \leq \inf_{\alpha \in \mathcal{A}} \| \hat \mathrm{M}(\alpha, \hat \eta)\| + o(n^{-1/2}).
\end{equation}
In order to analyze this estimator, we
assume that the derivatives
$\Gamma_1 : = \partial_{\alpha'} \mathrm{M}(\alpha_0, \eta_0)$ and $\partial_{\eta'} \mathrm{M}(\alpha, \eta_0)$
exist. We assume that $\alpha_0$ is interior relative
to the parameter space $\mathcal{A}$; namely, for some $\ell_n \to \infty$ such that $\ell_n/\sqrt{n} \to 0$,
\begin{equation}\label{eq:interior}
\{ \alpha \in \mathbb{R}^d: \| \alpha - \alpha_0\| \leq \ell_n/\sqrt{n} \} \subset \mathcal{A}.
\end{equation}
We also assume that the following local-global identifiability condition holds: For some constant $c>0$,
\begin{equation}
2 \| \mathrm{M}(\alpha, \eta_0)\| \geq \| \Gamma_1 (\alpha - \alpha_0) \| \wedge c \ \ \forall \alpha \in \mathcal{A}, \quad \mathrm{mineig}{(\Gamma_1'\Gamma_1)} \geq c.\label{eq:id}
\end{equation}
Furthermore, for $\Omega = \operatorname{Var} ( \sqrt{n} \hat \mathrm{M}(\alpha_0, \eta_0) )$, we suppose that the
central limit theorem,
\begin{equation}\label{eq:normal score}
\Omega^{-1/2} \sqrt{n} \hat \mathrm{M}(\alpha_0, \eta_0) \leadsto_{{\mathrm{P}}_n} \mathcal{N}(0, I),
\end{equation}
and the stability condition,
\begin{equation}\label{eq:stability}
\| \Gamma_1' \Gamma_1 \| + \|\Omega\| + \|\Omega^{-1}\| \lesssim 1,
\end{equation}
hold.
Assume that for some sequence of positive numbers $\{r_n\}$ such that $r_n \to 0$ and $r_n n^{1/2} \to \infty$, the following stochastic equicontinuity and continuity conditions hold:
\begin{eqnarray}\label{eq:se1}
&& \sup_{ \alpha \in \mathcal{A} } \frac{\| \hat \mathrm{M}(\alpha, \hat \eta) - \mathrm{M}(\alpha, \hat \eta) \| + \| \mathrm{M}(\alpha, \hat \eta) - \mathrm{M}(\alpha, \eta_0) \|}{r_n + \| \hat \mathrm{M} (\alpha, \hat \eta)\| + \|\mathrm{M}(\alpha, \eta_0)\| } \to_{{\mathrm{P}}_n} 0,\\
&& \sup_{ \| \alpha - \alpha_0\| \leq r_n} \frac{\| \hat \mathrm{M}(\alpha, \hat \eta) - \mathrm{M}(\alpha, \hat \eta) - \hat \mathrm{M}(\alpha_0, \eta_0) \|}{n^{-1/2} + \| \hat \mathrm{M} (\alpha, \hat \eta)\| + \|\mathrm{M}(\alpha, \eta_0)\| } \to_{{\mathrm{P}}_n} 0.\label{eq:se2}
\end{eqnarray}
Suppose that uniformly for all $\alpha \neq \alpha_0$ such that $\| \alpha - \alpha_0\| \leq r_n \to 0$, the following
conditions on the smoothness of $\mathrm{M}$ and the quality of the estimator $\hat \eta$ hold, as $n \to \infty$:
\begin{align}\label{eq:derivatives}
\begin{array}{lll}
&& \| \mathrm{M}(\alpha, \eta_0) - \mathrm{M}(\alpha_0, \eta_0)- \Gamma_1 [\alpha - \alpha_0]\| \| \alpha - \alpha_0\|^{-1} \to 0,\\
&& \sqrt{n} \| \mathrm{M}(\alpha, \hat \eta) - \mathrm{M}(\alpha, \eta_0)- \partial_{\eta'} \mathrm{M} (\alpha, \eta_0) [\hat \eta - \eta_0]\| \to_{{\mathrm{P}}_n} 0, \\
&& \| \{ \partial_{\eta'} \mathrm{M} (\alpha, \eta_0) - \partial_{\eta'} \mathrm{M} (\alpha_0, \eta_0)\} [\hat \eta - \eta_0] \| \|\alpha - \alpha_0\|^{-1}
\to_{{\mathrm{P}}_n} 0. \end{array} \end{align}
Finally, as before, we assume that the orthogonality condition
\begin{equation}\label{eq:ortho-e}
\partial_{\eta'} \mathrm{M} (\alpha_0, \eta_0) = 0
\end{equation}
holds.
The above conditions extend the analysis of \cite{pakes:pollard} and \cite{chen:linton:k}, which
in turn extended Huber's (\citeyear{huber}) classical results on Z-estimators. These conditions allow
for both smooth and non-smooth systems of estimating equations. The identifiability condition imposed above
is mild and holds for broad classes of identifiable models. The equicontinuity and smoothness
conditions imposed above require mild smoothness on the function $\mathrm{M}$
and also require that $\hat \eta$ is a good-quality estimator of $\eta_0$. In particular,
these conditions will often require that $\hat \eta$ converges to $\eta_0$ at a faster rate than $n^{-1/4}$
as demonstrated, for example, in the next section. However, the rate condition alone is not sufficient for adaptivity. We also need
the orthogonality condition (\ref{eq:ortho-e}). In addition, it is required that $\hat \eta \in \mathcal{H}_n$,
where $\mathcal{H}_n$ is a set whose complexity does not grow too quickly with the sample
size, to verify the stochastic equicontinuity condition; see, e.g., \cite{BCFH:Policy} and \cite{BCK-LAD}.
In the next section, we use the sparsity of $\hat \eta$ to control this complexity. Note that
conditions (\ref{eq:se1})-(\ref{eq:se2})
can be simplified by leaving only $r_n$ and $n^{-1/2}$ in the denominator, though
this simplification would then require imposing compactness on $\mathcal{A}$ even in linear problems.
\begin{proposition}[Valid Inference via Adaptive Estimation after Selection or Regularization]\label{prop:estimation}
Consider a sequence $\{\mathbf{P}_n\}$ of sets
of probability laws such that for each sequence $\{{\mathrm{P}}_n\} \in \{\mathbf{P}_n\}$
conditions (\ref{eq:minimization})-(\ref{eq:ortho-e}) hold. Then
$$
\sqrt{n}(\hat \alpha -\alpha_0) + [\Gamma_1'\Gamma_1]^{-1} \Gamma_1' \sqrt{n} \hat \mathrm{M}( \alpha_0, \eta_0) \to_{{\mathrm{P}}_n}0.
$$
In addition, for $V_n := (\Gamma'_1 \Gamma_1)^{-1} \Gamma_1' \Omega \Gamma_1(\Gamma'_1 \Gamma_1)^{-1},$
we have that
$$
\lim_{n \to \infty} \sup_{{\mathrm{P}} \in \mathbf{P}_n} \sup_{R \in \mathcal{R}} |{\mathrm{P}} ( V_n^{-1/2} (\hat \alpha - \alpha_0) \in R) - \mathbb{P}( \mathcal{N}(0,I) \in R) | = 0,
$$
where $\mathcal{R}$ is a collection of all convex sets. Moreover, the result continues to apply if $V_n$ is replaced by a consistent estimator $\hat V_n$ such that
$\hat V_n - V_n \to_{{\mathrm{P}}_n} 0$ under each sequence $\{{\mathrm{P}}_n\}$. Thus, $ \mathsf{CR}^l_{1-a} = [
l'\hat \alpha \pm c(1-a/2) (l'\hat V_nl/n)^{1/2}]$ where $c(1-a/2)$ is the $(1-a/2)$-quantile
of $\mathcal{N}(0,1)$ is a uniformly valid confidence set for $l'\alpha_0$:
$$
\lim_{n \to \infty} \sup_{{\mathrm{P}} \in \mathbf{P}_n} |{\mathrm{P}} ( l'\alpha_0 \in \mathsf{CR}^l_{1-a} ) - (1-a) | = 0.
$$
\end{proposition}
Note that the above formulation implicitly accommodates weighting options. Suppose
$\mathrm{M}^o$ and $\hat \mathrm{M}^o$ are the original theoretical and empirical systems of equations,
and let $\Gamma_1^o = \partial_{\alpha'} \mathrm{M}^o (\alpha_0, \eta_0)$ be the original Jacobian.
We could consider $k \times k$ positive-definite weight matrices $\mathrm{A}$ and $\hat{\mathrm{A}}$ such that
\begin{equation}\label{eq:A}
\|\mathrm{A}^2\| + \| (\mathrm{A}^2)^{-1} \| \lesssim 1, \quad \|\hat{\mathrm{A}}^2 - \mathrm{A}^2 \| \to_{{\mathrm{P}}_n} 0.
\end{equation}
For example, we may wish to use the optimal weighting matrix $
\mathrm{A}^2 = \operatorname{Var}(\sqrt{n} \hat{\mathrm{M}}^o(\alpha_0, \eta_0))^{-1}$ which can be estimated by $\hat \mathrm{A}^2$ obtained
using a preliminary estimator $\hat \alpha^o$ resulting from solving the problem with some non-optimal weighting
matrix such as $I$. We can then simply redefine the system of equations and the Jacobian
according to
\begin{equation}\label{eq:redefine}
\mathrm{M}(\alpha, \eta) = \mathrm{A} \mathrm{M}^o(\alpha, \eta), \quad \hat{\mathrm{M}}(\alpha, \eta) = \hat {\mathrm{A}} \hat{\mathrm{M}}^o(\alpha, \eta), \quad \Gamma_1 = \mathrm{A} \Gamma_1^o.
\end{equation}
\bigskip
\begin{proposition}[Adaptive Estimation via Weighted Equations]\label{prop:weighted}
Consider a sequence $\{\mathbf{P}_n\}$ of sets
of probability laws such that for each sequence $\{{\mathrm{P}}_n\} \in \{\mathbf{P}_n\}$
the
conditions of Proposition \ref{prop:estimation} hold for the original pair of
systems of equations $(\mathrm{M}^o, \hat{\mathrm{M}}^o)$ and that (\ref{eq:A}) holds. Then these
conditions also hold for the new pair $(\mathrm{M}, \hat{\mathrm{M}})$ in (\ref{eq:redefine}), so that all the conclusions
of Proposition \ref{prop:estimation} apply to the resulting approximate argmin estimator $\hat \alpha$. In particular, if
we use $\mathrm{A}^2 = \operatorname{Var}(\sqrt{n} \hat{\mathrm{M}}^o(\alpha_0, \eta_0))^{-1}$ and $\hat \mathrm{A}^2 - \mathrm{A}^2 \to_{{\mathrm{P}}_n} 0$, then
the large sample variance $V_n$ simplifies to
$
V_n = (\Gamma_1'\Gamma_1)^{-1}.
$
\end{proposition}
\subsection{Inference via Adaptive ``One-Step" Estimation}
We next consider a ``one-step" estimator. To define the estimator, we start with an initial
estimator $\tilde \alpha$ that satisfies, for $r_n = o(n^{-1/4})$,
\begin{equation}\label{eq:crude}
{\mathrm{P}}_n(\| \tilde \alpha - \alpha_0 \| \leq r_n) \to 1.
\end{equation}
The one-step estimator $\check\alpha$ then solves a linearized version of (\ref{eq:minimization}):
\begin{equation}\label{eq:one-step}
\check \alpha = \tilde \alpha - [\hat \Gamma_1'\hat \Gamma_1]^{-1} \hat \Gamma_1' \hat \mathrm{M}(\tilde \alpha, \hat \eta)
\end{equation}
where $\hat \Gamma_1$ is an estimator of $\Gamma_1$ such that
\begin{equation}\label{eq:consistent gamma}
{\mathrm{P}}_n(\|\hat \Gamma_1 - \Gamma_1\| \leq r_n) \to 1.
\end{equation}
Since the one-step estimator is considerably more crude than the argmin estimator, we need
to impose additional smoothness conditions. Specifically, we suppose that uniformly for all $\alpha \neq \alpha_0$ such that $\| \alpha - \alpha_0\| \leq r_n \to 0$, the following strengthened conditions on stochastic equicontinuity,
smoothness of $\mathrm{M}$ and the quality of the estimator $\hat \eta$ hold, as $n \to \infty$:
\begin{align}\label{eq:derivatives2}
\begin{array}{lll}
&& n^{1/2}\| \hat \mathrm{M}(\alpha, \hat \eta) - \mathrm{M}(\alpha, \hat \eta) - \hat \mathrm{M}(\alpha_0, \eta_0) \| \to_{{\mathrm{P}}_n} 0,\\
&& \| \mathrm{M}(\alpha, \eta_0) - \mathrm{M}(\alpha_0, \eta_0)- \Gamma_1 [\alpha - \alpha_0]\| \| \alpha - \alpha_0\|^{-2} \lesssim 1,\\
&& \sqrt{n} \| \mathrm{M}(\alpha, \hat \eta) - \mathrm{M}(\alpha, \eta_0)- \partial_{\eta'} \mathrm{M} (\alpha, \eta_0) [\hat \eta - \eta_0]\| \to_{{\mathrm{P}}_n} 0, \\
&& \sqrt{n} \| \{ \partial_{\eta'} \mathrm{M} (\alpha, \eta_0) - \partial_{\eta'} \mathrm{M} (\alpha_0, \eta_0)\} [\hat \eta - \eta_0] \|
\to_{{\mathrm{P}}_n} 0. \end{array} \end{align}
\bigskip
\begin{proposition}[Valid Inference via Adaptive One-Step Estimators]\label{prop:one-step}
Consider a sequence $\{\mathbf{P}_n\}$ of sets
of probability laws such that for each sequence $\{{\mathrm{P}}_n\} \in \{\mathbf{P}_n\}$ the
conditions of Proposition \ref{prop:estimation}
as well as (\ref{eq:crude}), (\ref{eq:consistent gamma}), and (\ref{eq:derivatives2}) hold. Then
the one-step estimator $\check \alpha$ defined by (\ref{eq:one-step}) is first order
equivalent to the argmin estimator $\hat \alpha$:
$$
\sqrt{n}(\check \alpha - \hat \alpha) \to_{{\mathrm{P}}_n} 0.
$$
Consequently, all conclusions of Proposition \ref{prop:estimation} apply to
$\check \alpha$ in place of $\hat \alpha$.
\end{proposition}
The one-step estimator requires stronger regularity conditions than the argmin estimator.
Moreover, there is finite-sample evidence (e.g. \cite{BCY-honest}) that in
practical problems the argmin estimator often works much better, since the one-step estimator
typically suffers from higher-order biases. This problem could be alleviated somewhat
by iterating on the one-step estimator, treating the previous iteration as the ``crude" start $\tilde \alpha$
for the next iteration.
\section{Achieving Orthogonality Using Neyman's Orthogonalization}
Here we describe orthogonalization ideas that go back at least to \cite{Neyman59}; see also \cite{Neyman1979}.
Neyman's idea was to project the score
that identifies the parameter of interest onto the ortho-complement of the tangent
space for the nuisance parameter. This projection underlies semi-parametric
efficiency theory, which is concerned particularly with the case in which $\eta$
is infinite-dimensional, cf. \cite{vdV}. Here we consider finite-dimensional $\eta$ of high dimension;
for discussion of infinite-dimensional $\eta$ in an approximately sparse setting, see \cite{BCFH:Policy} and \cite{BCK-LAD}.
\subsection{The Classical Likelihood Case} In likelihood settings, the construction of orthogonal equations was proposed by \cite{Neyman59} who used them in construction of his celebrated $C(\alpha)$-statistic. The $C(\alpha)$-statistic, or the orthogonal score statistic, was first explicitly utilized for testing (and also for setting up estimation) in high-dimensional sparse models in \cite{BCK-LAD} and \cite{BCK-QR},
in the context of quantile regression, and \cite{BCY-honest} in the context of logistic regression and other generalized linear models. More recent uses of $C(\alpha)$-statistics (or close variants) include those by \cite{witten:score}, \cite{hanliu1}, and \cite{hanliu2}.
Suppose that the (possibly conditional, possibly quasi) log-likelihood function associated with observation
$w_i$ is $\ell(w_i, \alpha, \beta)$, where $\alpha \in \mathcal{A} \subset \mathbb{R}^{d}$
is the target parameter and $\beta \in \mathcal{B} \subset \mathbb{R}^{p_0}$ is the nuisance parameter.
Under regularity conditions, the true parameter values $\gamma_0=(\alpha_0', \beta_0)'$ obey
\begin{equation}\label{eq: foc lik}
{\mathrm{E}} [\partial_\alpha \ell (w_i, \alpha_0, \beta_0)]=0, \quad {\mathrm{E}}[\partial_{\beta} \ell (w_i, \alpha_0, \beta_0)] =0.
\end{equation}
Now consider the moment function
\begin{equation}
{\mathrm{M}}(\alpha, \eta) = {\mathrm{E}}[\psi(w_i, \alpha, \eta)] , \ \ \psi(w_i, \alpha, \eta) = \partial_\alpha \ell (w_i, \alpha, \beta) - \mu \partial_{\beta} \ell (w_i, \alpha, \beta).
\end{equation}
Here the nuisance parameter is $$ \eta= (\beta', \textrm{vec}(\mu)')' \in \mathcal{B} \times \mathcal{D} \subset \mathbb{R}^{p}, \quad p=p_0 + dp_0,$$
where $\mu$ is the $d \times p_0$ \textit{orthogonalization} parameter matrix whose
true value $\mu_0$ solves the equation:
\begin{equation}\label{eq: ortho gono}
J_{\alpha\beta} - \mu J_{\beta \beta} =0 \ (\text{ i.e., } \mu_0 = J_{\alpha\beta}J^{-1}_{\beta \beta}),
\end{equation}
where, for $\gamma := (\alpha', \beta')'$ and $\gamma_0 := (\alpha_0', \beta_0')'$,
\begin{eqnarray*}
J := - \partial_{\gamma'} {\mathrm{E}} [ \partial_{\gamma} \ell(w_i, \gamma)\ ] \vert_{\gamma = \gamma_0} &=: & \left(
\begin{array}{cc}
J_{\alpha \alpha} & J_{\alpha \beta} \\
J_{\beta \alpha} & J_{\beta \beta} \\
\end{array}
\right).
\end{eqnarray*}
Note that $\mu_0$ not only creates the necessary orthogonality but also creates
\begin{itemize}
\item the \textit{optimal score} (in statistical language)
\item or, equivalently, the \textit{optimal instrument/moment} (in econometric language)\footnote{The connection
between optimal instruments/moments and likelihood/score has been elucidated by the fundamental work of \cite{chamberlain}.}
\end{itemize}
for inference about $\alpha_0$.
Provided $\mu_0$ is well-defined, we have by (\ref{eq: foc lik}) that
$${\mathrm{M}}(\alpha_0, \eta_0) = 0.$$ Moreover, the function $\mathrm{M}$ has the desired orthogonality property:\begin{equation}\label{eq: ortho lik}
\partial_{\eta'} {\mathrm{M}}(\alpha_0, \eta_0) = \Big [J_{\alpha \beta} - \mu_0 J_{\beta \beta}; \ F {\mathrm{E}}[\partial_\beta \ell (w_i, \alpha_0, \beta_0)] \Big ] =0,
\end{equation}
where $F$ is a tensor operator, such that $F x = \partial \mu x/ \partial \mathrm{vec(\mu)}' \mid_{\mu = \mu_0}$
is a $d \times (d p_0)$ matrix for any vector $x$ in $\mathbb{R}^{p_0}$. Note that the orthogonality property holds for Neyman's construction even if the likelihood is misspecified. That is, $\ell(w_i, \gamma_0)$
may be a quasi-likelihood, and the data need not be i.i.d. and may, for example,
exhibit complex dependence over $i$.
An alternative way to define $\mu_0$ arises by considering
that, under the correct specification and sufficient regularity, the information matrix equality holds and yields
\begin{eqnarray*}
J = J^0 & := & {\mathrm{E}} [ \partial_\gamma \ell(w_i, \gamma) \partial_\gamma \ell(w_i, \gamma) '] \vert_{\gamma = \gamma_0} \\
& = & \left . \left( \begin{array}{cc}
{\mathrm{E}}[ \partial_\alpha \ell(w_i, \gamma) \partial_{\alpha} \ell(w_i, \gamma)'] & {\mathrm{E}}[ \partial_\alpha \ell(w_i, \gamma) \partial_\beta \ell(w_i, \gamma)'] \\
{\mathrm{E}}[ \partial_\beta \ell(w_i, \gamma) \partial_\alpha \ell(w_i, \gamma)']& {\mathrm{E}}[ \partial_\beta \ell(w_i, \gamma) \partial_\beta \ell(w_i, \gamma)']\\
\end{array}
\right) \right |_{\gamma = \gamma_0}, \\
&=: & \left(
\begin{array}{cc}
J^0_{\alpha \alpha} & J^0_{ \alpha \beta} \\
J^{0}_{\beta \alpha} & J^0_{\beta \beta} \\
\end{array}
\right).
\end{eqnarray*}
Hence define $\mu^*_0 = J^0_{\alpha\beta} J^{0-1}_{\beta \beta}$ as the population \textit{projection coefficient} of the score for the main parameter $ \partial_\alpha \ell(w_i, \gamma_0)$ on the score for the nuisance parameter $\partial_\beta \ell(w_i, \gamma_0)$:
\begin{equation}\label{projection}
\partial_\alpha \ell(w_i, \gamma_0) = \mu^*_0 \partial_\beta \ell(w_i, \gamma_0) + \varrho, \ \ {\mathrm{E}}[ \varrho \partial_\beta \ell(w_i, \gamma_0)']=0.
\end{equation}
We can see this construction as the non-linear version of Frisch-Waugh's ``partialling out" from the linear regression model.
It is important to note that under misspecification the information matrix equality generally does not hold, and this projection approach does not provide valid orthogonalization.
\begin{lemma}[Neyman's orthogonalization for (quasi-) likelihood scores]
Suppose that for each $\gamma = (\alpha, \beta) \in \mathcal{A}\times\mathcal{B}$, the derivative $\partial_\gamma \ell(w_i, \gamma)$ exists and is continuous at $\gamma$ with probability one, and obeys the dominance condition ${\mathrm{E}} \sup_{\gamma \in \mathcal{A}\times\mathcal{B}}\| \partial_\gamma \ell(w_i, \gamma) \|^2< \infty$. Suppose that
condition (\ref{eq: foc lik}) holds for some (quasi-) true value $(\alpha_0, \beta_0)$. Then, (i) if $J$ exists and is finite and $J_{\beta \beta}$ is invertible, then the orthogonality condition (\ref{eq: ortho lik}) holds; (ii) if
the information matrix equality holds, namely $J= J^0$, then the orthogonality condition (\ref{eq: ortho lik}) holds for the projection parameter $\mu^*_0$ in place of the orthogonalization parameter matrix $\mu_0$.
\end{lemma}
The claim follows immediately from the computations above.
\bigskip
With the formulations given above Neyman's $C(\alpha)$-statistic takes the form
$$
C(\alpha) = \| S(\alpha)\|_2^2, \quad S (\alpha) = \hat \Omega^{-1/2} (\alpha, \hat \eta) \sqrt{n} \hat{\mathrm{M}}(\alpha, \hat \eta),
$$
where $\hat{\mathrm{M}}(\alpha, \hat \eta) = {\mathbb{E}_n} [\psi(w_i, \alpha, \hat \eta)]$ as before,
$ \Omega (\alpha, \eta_0) = \mathrm{Var} ( \sqrt{n} \hat{\mathrm{M}}(\alpha, \eta_0))$,
and $\hat \Omega (\alpha, \hat \eta)$ and $\hat \eta$ are suitable estimators
based on sparsity or other structured assumptions. The estimator is then
$$
\hat \alpha = \arg\inf_{\alpha \in \mathcal{A}} C(\alpha) = \arg\inf_{\alpha \in \mathcal{A}} \| \sqrt{n}\hat{\mathrm{M}}(\alpha, \hat \eta) \|,
$$
provided that $\hat \Omega (\alpha, \hat \eta)$ is positive definite for each $\alpha \in \mathcal{A}$. If the conditions
of Section 2 hold, we have that
\begin{equation}\label{mle:result}
C(\alpha) \leadsto \chi^2(d), \quad V_n^{-1/2}\sqrt{n}(\hat \alpha - \alpha_0) \leadsto \mathcal{N}(0,I),
\end{equation}
where $V_n = \Gamma_1^{-1} \Omega (\alpha_0, \eta_0)\Gamma_1^{-1}$ and $\Gamma_1 =
J_{\alpha \alpha} - \mu_0 J'_{\alpha \beta}$. Under the correct specification and i.i.d. sampling,
the variance matrix $V_n$ further reduces to the optimal
variance $$\Gamma_1^{-1}= (J_{\alpha \alpha} - J_{\alpha \beta} J^{-1}_{\beta \beta }J'_{\alpha \beta})^{-1},$$ of the
first $d$ components of the maximum likelihood estimator in a Gaussian shift experiment with observation $Z \sim \mathcal{N}(h, J^{-1}_0)$. Likewise, the result (\ref{mle:result}) also holds for
the one-step estimator $\check \alpha$ of Section 2 in place of $\hat \alpha$ as long as the conditions in Section 2 hold.
Provided that sparsity or its generalizations are plausible assumptions to make regarding
$\eta_0$, the formulations above naturally lend themselves to sparse estimation. For example, \cite{BCY-honest} used
penalized and post-penalized maximum likelihood to estimate $\beta_0$, and used the information matrix
equality to estimate the orthogonalization parameter matrix $\mu^*_0$ by using Lasso or Post-Lasso estimation
of the projection equation (\ref{projection}). It is also possible to estimate $\mu_0$ directly by
finding approximate sparse solutions to the empirical analog of the system of equations
$
J_{\alpha \beta} - \mu J_{\beta \beta}= 0$ using $\ell_1$-penalized estimation, as, e.g., in \cite{vdGBRD:AsymptoticConfidenceSets}, or post-$\ell_1$-penalized estimation.
\subsection{Achieving Orthogonality in GMM Problems }
Here we consider $\gamma_0 = (\alpha_0', \beta_0')'$ that solve
the system of equations: $$
{\mathrm{E}} [m( w_i, \alpha_0, \beta_0)] = 0,
$$
where $m: \mathcal{W} \times \mathcal{A} \times \mathcal{B} \mapsto \mathbb{R}^k$,
$\mathcal{A}\times\mathcal{B}$ is a convex subset of $\mathbb{R}^{d} \times \mathbb{R}^{p_0}$,
and $k \geq d+ p_0$ is the number of moments. The orthogonal moment equation is
\begin{equation}\label{eq: GMM1}
{\mathrm{M}}(\alpha, \eta) = {\mathrm{E}}[\psi(w_i, \alpha, \eta)] , \ \ \psi(w_i, \alpha, \eta) =
\mu m(w_i, \alpha, \beta).
\end{equation}
The nuisance parameter is $$ \eta= (\beta', \textrm{vec}(\mu)')' \in \mathcal{B} \times \mathcal{D} \subset \mathbb{R}^{p}, \quad p=p_0 + dk,$$
where $\mu$ is the $d \times k$ orthogonalization parameter matrix. The ``true value" of $\mu$ is $$
\mu_0=(G_{\alpha}' \Omega_m^{-1} - G_{\alpha}' \Omega_m^{-1} G_\beta (G_\beta' \Omega_m^{-1} G_\beta)^{-1} G_\beta '\Omega^{-1}_m ),
$$
where, for $\gamma = (\alpha', \beta')'$ and $\gamma_0 = (\alpha_0', \beta_0')'$,
$$
G_{\gamma} = \partial_{\gamma'} {\mathrm{E}}[m(w_i, \alpha, \beta)] \Big |_{\gamma= \gamma_0}= \Big [\partial_{\alpha'} {\mathrm{E}}[m(w_i, \alpha, \beta)], \partial_{\beta'} {\mathrm{E}}[m(w_i, \alpha, \beta)] \Big ] \Big |_{\gamma= \gamma_0}= : \Big [G_{\alpha}, G_{\beta} \Big],
$$
and
$$
\Omega_m = \operatorname{Var}( \sqrt{n} {\mathbb{E}_n}[m(w_i, \alpha_0, \beta_0) ] ).
$$
As before, we can interpret $\mu_0$ as an operator creating orthogonality while building
\begin{itemize}
\item the \textit{optimal instrument/moment} (in econometric language),
\item or, equivalently,
the \textit{optimal score} function (in statistical language).\footnote{Cf.~previous footnote.}
\end{itemize}
The resulting moment function has the required orthogonality property; namely, the first derivative with respect to the nuisance
parameter when evaluated at the true parameter values is zero:
\begin{equation}\label{GMM:ortho}
\partial_{\eta'}{\mathrm{M}}(\alpha_0, \eta)\vert_{\eta = \eta_0} = [\mu_0 G_\beta, F {\mathrm{E}}[m(w_i, \alpha_0, \beta_0)]] = 0,
\end{equation}
where $F$ is a tensor operator, such that $F x = \partial \mu x/ \partial \mathrm{vec(\mu)}' \mid_{\mu = \mu_0}$
is a $d \times (dk)$ matrix for any vector $x$ in $\mathbb{R}^{k}$.
Estimation and inference on $\alpha_0$ can be based on the empirical analog of (\ref{eq: GMM1}):
$$
\hat{\mathrm{M}}(\alpha, \hat \eta) = {\mathbb{E}_n}[ \psi (w_i, \alpha, \hat \eta) ],
$$
where $\hat \eta$ is a post-selection or other regularized estimator of $\eta_0$.
Note that the previous framework of (quasi)-likelihood is incorporated as a special case
with $$m(w_i, \alpha, \beta) = [ \partial_{\alpha} \ell(w_i, \alpha)', \partial_{\beta} \ell(w_i, \beta)']'.$$
With the formulations above, Neyman's $C(\alpha)$-statistic takes the form:
$$
C(\alpha) = \| S(\alpha)\|_2^2, \quad S (\alpha) = \hat \Omega^{-1/2} (\alpha, \hat \eta) \sqrt{n} \hat{\mathrm{M}}(\alpha, \hat \eta),
$$
where $\hat{\mathrm{M}}(\alpha, \hat \eta) = {\mathbb{E}_n} [\psi(w_i, \alpha, \hat \eta)]$ as before,
$ \Omega (\alpha, \eta_0) = \mathrm{Var} ( \sqrt{n} \hat{\mathrm{M}}(\alpha, \eta_0))$,
and $\hat \Omega (\alpha, \hat \eta)$ and $\hat \eta$ are suitable estimators
based on structured assumptions. The estimator is then
$$
\hat \alpha = \arg\inf_{\alpha \in \mathcal{A}} C(\alpha) = \arg\inf_{\alpha \in \mathcal{A}} \| \sqrt{n}\hat{\mathrm{M}}(\alpha, \hat \eta) \|,
$$
provided that $\hat \Omega (\alpha, \hat \eta)$ is positive definite for each $\alpha \in \mathcal{A}$. If the high-level conditions
of Section 2 hold, we have that
\begin{equation}\label{GMM:result}
C(\alpha) \leadsto_{{\mathrm{P}}_n} \chi^2(d), \quad V_n^{-1/2}\sqrt{n}(\hat \alpha - \alpha) \leadsto_{{\mathrm{P}}_n} \mathcal{N}(0,I),
\end{equation}
where $V_n =(\Gamma_1')^{-1} \Omega (\alpha_0, \eta_0)(\Gamma_1)^{-1}$ coincides with the optimal variance
for GMM; here $\Gamma_1 = \mu_0 G_{\alpha}$. Likewise, the same result (\ref{GMM:result}) holds for
the one-step estimator $\check \alpha$ of Section 2 in place of $\hat \alpha$ as long as the conditions in Section 2 hold.
In particular, the variance $V_n$
corresponds to the variance of the first $d$ components
of the maximum likelihood estimator in the normal shift experiment with the observation
$Z \sim \mathcal{N}(h, (G_\gamma'\Omega_m^{-1}G_\gamma)^{-1})$.
The above
is a generic outline of the properties that are expected for inference using orthogonalized GMM equations under structured assumptions. The problem of inference in GMM under sparsity is a very delicate matter due to the complex form of the orthogonalization parameters.
One approach to the problem is developed in \cite{GMM:sparse}.
\section{Achieving Adaptivity In Affine-Quadratic Models via Approximate Sparsity}
Here we take orthogonality as given
and explain how we can use approximate sparsity to achieve the adaptivity property (\ref{eq:adaptivity}).
\subsection{The Affine-Quadratic Model} We analyze the case in which $\hat \mathrm{M}$ and $\mathrm{M}$ are \textit{affine} in $\alpha$ and \textit{affine-quadratic} in $\eta$. Specifically, we suppose that for
all $\alpha$
$$
\hat \mathrm{M} (\alpha, \eta) = \hat \Gamma_1(\eta)\alpha + \hat \Gamma_2(\eta), \quad
\mathrm{M} (\alpha, \eta) = \Gamma_1(\eta)\alpha + \Gamma_2(\eta),
$$
where the orthogonality condition holds,
\begin{align*}
\partial_{\eta'} \mathrm{M} (\alpha_0, \eta_0) = 0,
\end{align*}
and $\eta \mapsto \hat \Gamma_j (\eta)$ and $\eta \mapsto \Gamma_j (\eta)$ are affine-quadratic in $\eta$ for $j=1$ and $j=2$. That is, we will have that all second-order derivatives of $\hat \Gamma_j (\eta)$ and $\Gamma_j (\eta)$ for $j = 1$ and $j=2$ are constant over the convex parameter
space $\mathcal{H}$ for $\eta$.
This setting is both useful, including most widely used linear models as a special case, and pedagogical, permitting simple illustration of the key issues that arise in treating the general problem. The derivations
given below easily generalize to more complicated models, but we defer the details to the interested reader.
The estimator in this case is
\begin{equation}\label{eq:linear estimator}
\hat \alpha = \arg\min_{\alpha \in \mathbb{R}^d} \|\hat \mathrm{M} (\alpha, \hat \eta) \|^2= - [\hat \Gamma_1(\hat \eta)'\hat \Gamma_1(\hat \eta)]^{-1} \hat \Gamma_1(\hat \eta)' \hat \Gamma_2(\hat \eta),
\end{equation}
provided the inverse is well-defined. It follows that
\begin{equation}\label{eq:linear estimator represent}
\sqrt{n} (\hat \alpha - \alpha_0) = - [\hat \Gamma_1(\hat \eta)'\hat \Gamma_1(\hat \eta)]^{-1} \hat \Gamma_1(\hat \eta)' \sqrt{n} \hat \mathrm{M}(\alpha_0, \hat \eta).
\end{equation}
This estimator is adaptive if, for $\Gamma_1:= \Gamma_1(\eta_0)$,
$$
\sqrt{n} (\hat \alpha - \alpha_0) + [\Gamma_1' \Gamma_1]^{-1} \Gamma_1' \sqrt{n}\hat \mathrm{M}(\alpha_0, \eta_0) \to_{{\mathrm{P}}_n} 0,
$$
which occurs under the conditions in (\ref{eq:normal score}) and (\ref{eq:stability}) if
\begin{equation}\label{eq:adaptivity estimation linear}
\sqrt{n}(\hat \mathrm{M}(\alpha_0, \hat \eta) - \hat \mathrm{M}(\alpha_0, \eta_0)) \to_{{\mathrm{P}}_n} 0, \quad \hat \Gamma_1(\hat \eta) - \Gamma_1(\eta_0)
\to_{{\mathrm{P}}_n} 0.
\end{equation}
Therefore, the problem of the adaptivity of the estimator is directly connected to the problem of the adaptivity of testing hypotheses
about $\alpha_0$.
\bigskip
\begin{lemma}[Adaptive Testing and Estimation in Affine-Quadratic Models] \label{prop:affine}
Consider a sequence $\{\mathbf{P}_n\}$ of sets of probability laws such that for each sequence $\{{\mathrm{P}}_n\} \in \{\mathbf{P}_n\}$, conditions stated in the first paragraph of Section 4.1, condition (\ref{eq:adaptivity estimation linear}), the asymptotic normality condition
(\ref{eq:normal score}), the stability condition (\ref{eq:stability}), and condition (\ref{eq:variance}) hold. Then
all the conditions of Propositions 1 and 2 hold. Moreover, the conclusions of Proposition 1 hold, and the conclusions of Proposition \ref{prop:estimation} hold for the estimator $\hat \alpha$ in (\ref{eq:linear estimator}).
\end{lemma}
\subsection{Adaptivity for Testing via Approximate Sparsity}
Assuming the
orthogonality condition holds, we follow \cite{BCCH12} in using approximate
sparsity to achieve the adaptivity property (\ref{eq:adaptivity}) for the testing problem in the
affine-quadratic models.
We can expand each element $\hat{\mathrm{M}}_j$ of $\hat{\mathrm{M}}= (\hat \mathrm{M}_j )_{j=1}^k$
as follows:
\begin{equation}\label{eq:decompose}
\sqrt{n}(\hat{\mathrm{M}}_j(\alpha_0, \hat \eta) - \hat{\mathrm{M}}_j(\alpha_0, \eta_0)) = T_{1,j} + T_{2,j} +T_{3,j},
\end{equation}
where
\begin{equation}\label{eq:terms}
\begin{array}{lll}
& & T_{1,j}:= \sqrt{n}\partial_{\eta} {\mathrm{M}}_j(\alpha_0, \eta_0) '(\hat \eta - \eta_0), \\
&& T_{2,j}:= \sqrt{n} (\partial_\eta \hat{\mathrm{M}}_j(\alpha_0, \eta_0) - \partial_\eta {\mathrm{M}}_j(\alpha_0, \eta_0) ) '(\hat \eta - \eta_0), \\
& & T_{3,j}:= \sqrt{n} 2^{-1} (\hat \eta - \eta_0)' \partial_\eta \partial_{\eta'}\hat{\mathrm{M}}_j(\alpha_0) (\hat \eta - \eta_0).
\end{array}
\end{equation}
The term $T_{1,j}$ vanishes precisely because of orthogonality, i.e. $$T_{1,j}=0.$$ However, terms $T_{2,j}$ and $T_{3,j}$
need not vanish. In order to show that they are asymptotically negligible, we need to impose further structure on the problem.
\textbf{Structure 1: Exact Sparsity.} We first consider the case of using an exact sparsity structure where $\|\eta_0\|_0 \leq s$ and $s=s_n \geq 1$ can depend on $n$. We then use
estimators $\hat \eta$ that exploit the sparsity structure.
Suppose that the following bounds hold with probability $1 - o(1)$ under ${\mathrm{P}}_n$:
\begin{equation}\label{eq:sparse}
\begin{array}{c}
\|\hat \eta\|_0 \lesssim s, \quad \|\eta_0 \|_0 \leq s, \\
\| \hat \eta- \eta_0\|_2 \lesssim \sqrt{(s/n) \log (pn)}, \quad \|\hat \eta - \eta_0\|_1
\lesssim \sqrt{(s^2/n) \log (pn)}.
\end{array}
\end{equation}
These conditions are typical performance bounds which hold for many sparsity-based
estimators such as Lasso, post-Lasso, and their extensions.
We suppose further that the moderate deviation bound
\begin{equation}\label{eq:SNMD}
\bar T_{2,j}=\| \sqrt{n} (\partial_{\eta'} \hat{\mathrm{M}}_j(\alpha_0, \eta_0) - \partial_{\eta'} {\mathrm{M}}_j(\alpha_0, \eta_0) )\|_{\infty} \lesssim_{{\mathrm{P}}_n} \sqrt{\log (p n)}
\end{equation}
holds and that the sparse norm of the second-derivative matrix is bounded:
\begin{equation}\label{eq:BDD}
\bar T_{3,j}= \| \partial_\eta \partial_{\eta'} \hat{\mathrm{M}}_j ( \alpha_0) \|_{\mathsf{sp}(\ell_n s)} \lesssim_{{\mathrm{P}}_n} 1
\end{equation}
where $\ell_n \to \infty$ but $\ell_n =o( \log n)$.
Following \cite{BCCH12}, we can verify condition (\ref{eq:SNMD}) using the moderate deviation theory for self-normalized sums (e.g., \cite{jing:etal}), which allows us to avoid making highly restrictive subgaussian or gaussian tail assumptions. Likewise, following \cite{BCCH12}, we can verify the second condition using laws of large numbers for large
matrices acting on sparse vectors as in \cite{rudelson:vershynin} and \cite{RudelsonZhou2011}; see Lemma 7. Indeed, condition (\ref{eq:BDD}) holds if
$$
\begin{array}{ll}
\| \partial_\eta \partial_{\eta'} \hat{\mathrm{M}}_j ( \alpha_0) -\partial_\eta \partial_{\eta'} {\mathrm{M}}_j(\alpha_0) \|_{\mathsf{sp}(\ell_n s)} \to_{{\mathrm{P}}_n} 0, \
\|\partial_\eta \partial_{\eta'} {\mathrm{M}}_j ( \alpha_0) \|_{\mathsf{sp}(\ell_n s)} \lesssim 1.
\end{array}
$$
The above analysis immediately implies the following elementary result.
\bigskip
\begin{lemma}[Elementary Adaptivity for Testing via Sparsity]\label{lemma:ad testing via sparsity} Let $\{ {\mathrm{P}}_n\}$ be a sequence of probability laws. Assume (i) $\eta \mapsto \hat \mathrm{M}(\alpha_0, \eta)$ and $\eta \mapsto \mathrm{M}(\alpha_0, \eta)$
are affine-quadratic in $\eta$ and the orthogonality condition holds, (ii) that the conditions on sparsity and the quality of estimation (\ref{eq:sparse}) hold, and the sparsity index obeys
\begin{equation}\label{strong}
s^2 \log (pn)^2/n \to 0,
\end{equation}
(iii) that the moderate deviation bound (\ref{eq:SNMD}) holds, and (iv) the sparse norm of the second
derivatives matrix is bounded as in (\ref{eq:BDD}). Then the adaptivity condition (\ref{eq:adaptivity}) holds for the sequence $\{ {\mathrm{P}}_n\}$.
\end{lemma}
We note that (\ref{strong}) requires that the true value of the nuisance parameter is sufficiently sparse,
which we can relax in some special cases to the requirement $s \log (pn)^{c}/n \to 0$, for some constant $c$, by using
sample-splitting techniques; see \cite{BCCH12}. However, this requirement seems unavoidable in general.
Proof. We note above that $T_{1,j} = 0$ by orthogonality. Under (\ref{eq:sparse})-(\ref{eq:SNMD}) if $s^2 \log (pn)^2/n \to 0$, then $T_{2,j}$ vanishes in probability, as by H\"{o}lder's inequality,
$$
T_{2,j} \leq \bar T_{2,j} \|\hat \eta - \eta_0\|_1 \lesssim_{{\mathrm{P}}_n} \sqrt{s^2 \log (pn)^2/n} \to_{{\mathrm{P}}_n} 0.
$$
Also, if $s^2 \log (pn)^2/n \to 0$, then $T_{3,j}$ vanishes in probability, since by H\"{o}lder's inequality
and for sufficiently large $n$,
$$
T_{3,j} \leq \bar T_{3,j} \|\hat \eta - \eta_0\|^2 \lesssim_{{\mathrm{P}}_n} \sqrt{n} s \log (pn)/n \to_{{\mathrm{P}}_n} 0.
$$
The conclusion follows from (\ref{eq:decompose}). \hfill{\tiny \ensuremath{\blacksquare} }
\textbf{Structure 2. Approximate Sparsity.} Following \cite{BCCH12},
we next consider an approximate sparsity structure. Approximate sparsity imposes that,
given a constant $c>0$, we can decompose $\eta_0$ into a
sparse component $\eta^m_r$ and a ``small" non-sparse component $\eta^r$:
\begin{align}\label{eq:app sparsity}
\begin{array}{c}
\eta_0 = \eta^m_{0} + \eta^r_{0}, \ \mathrm{support}(\eta^m_{0}) \cap \mathrm{support}(\eta^r_{0}) = \emptyset, \\
\| \eta_{0}^m\|_0 \leq s, \ \|\eta_{0}^r\|_2 \leq c \sqrt{s/n}, \ \|\eta^r_{0}\|_1 \leq c\sqrt{s^2/n}.
\end{array}
\end{align}
This condition allows for much more realistic and
richer models than can be accommodated under exact sparsity. For example,
$\eta_0$ needs not have \textit{any} zero components at all under approximate sparsity.
In Section 5, we provide an example in which (\ref{eq:app sparsity}) arises from a
more primitive condition that the absolute values $\{|\eta_{0j}|, j =1,...,p\}$, sorted in decreasing order,
decay at a polynomial speed with respect to $j$.
Suppose that we have an estimator $\hat \eta$ such that
with probability $1 - o(1)$ under ${\mathrm{P}}_n$ the following bounds hold:
\begin{equation}\label{eq:sparse2}
\begin{array}{c}
\|\hat \eta\|_0 \lesssim s,
\| \hat \eta- \eta^m_0\|_2 \lesssim \sqrt{(s/n) \log (pn)}, \quad \|\hat \eta - \eta^m_0\|_1
\lesssim \sqrt{(s^2/n) \log (pn)}.
\end{array}
\end{equation}
This condition is again a standard performance bound expected to hold for sparsity-based
estimators under approximate sparsity conditions; see \cite{BCCH12}. Note that by the approximate sparsity condition, we also have that, with probability $1- o(1)$ under ${\mathrm{P}}_n$,
\begin{equation}\label{eq:sparse overall performance}
\| \hat \eta- \eta_0\|_2 \lesssim \sqrt{(s/n) \log (pn)}, \quad \|\hat \eta - \eta_0\|_1
\lesssim \sqrt{(s^2/n) \log (pn)}.
\end{equation}
We can employ the same moderate deviation and bounded sparse norm
conditions as in the previous subsection. In addition, we require the pointwise norm of the second-derivatives matrix to be bounded. Specifically,
for any deterministic vector $a \neq 0$, we require
\begin{equation}\label{eq:pointwise BDD}
\ \|\partial_\eta \partial_{\eta'}\hat{\mathrm{M}}_j(\alpha_0)\|_{{\sf pw}(a)} \lesssim_{{\mathrm{P}}_n} 1.
\end{equation}
This condition can be easily verified using ordinary laws of large numbers.
\bigskip
\begin{lemma}[Elementary Adaptivity for Testing via Approximate Sparsity]\label{lemma:adapt approx sparsity} Let $\{ {\mathrm{P}}_n\}$ be a sequence of probability laws. Assume (i) $\eta \mapsto \hat \mathrm{M}(\alpha_0, \eta)$ and $\eta \mapsto \mathrm{M}(\alpha_0, \eta)$
are affine-quadratic in $\eta$ and the orthogonality condition holds, (ii) that the conditions on approximate sparsity (\ref{eq:app sparsity}) and the quality of estimation (\ref{eq:sparse2}) hold, and the sparsity index obeys
$$
s^2 \log (pn)^2/n \to 0,
$$
(iii) that the moderate deviation bound (\ref{eq:SNMD}) holds, (iv) the sparse norm of the second
derivatives matrix is bounded as in (\ref{eq:BDD}), and (v) the pointwise norm of the second derivative
matrix is bounded as in (\ref{eq:pointwise BDD}). Then the adaptivity condition (\ref{eq:adaptivity}) holds:
$$\sqrt{n}(\hat{\mathrm{M}}(\alpha_0, \hat \eta) - \hat{\mathrm{M}}(\alpha_0, \eta_0)) \to_{{\mathrm{P}}_n} 0.$$
\end{lemma}
\subsection{Adaptivity for Estimation via Approximate Sparsity} We work with the approximate sparsity setup and the affine-quadratic model introduced in the previous subsections.
In addition to the previous assumptions, we impose the following conditions on the components
$\partial_\eta {\Gamma}_{1,ml}$ of $\partial_\eta {\Gamma}_{1}$, where $m=1,...,k$ and $l=1,...,d$,. First, we need the following
deviation and boundedness condition: For each $m$ and $l$,
\begin{equation}\label{eq:SNMD estimation}
\|\partial_\eta \hat{\Gamma}_{1,ml}(\eta_0) - \partial_\eta {\Gamma}_{1,ml}(\eta_0)\|_{\infty}
\lesssim_{{\mathrm{P}}_n} 1, \quad \|\partial_{\eta} {\Gamma}_{1, ml}(\eta_0) \|_{\infty} \lesssim 1.
\end{equation}
Second, we require the sparse and pointwise norms of the following second-derivative
matrices be stochastically bounded: For each $m$ and $l$,
\begin{equation}\label{eq:Gamma BDD}
\|\partial_\eta \partial_{\eta'}\hat{\Gamma}_{1, ml}\|_{\mathsf{sp}(\ell_n s)} + \|\partial_\eta \partial_{\eta'}\hat{\Gamma}_{1, ml}\|_{\mathsf{pw}(a)} \lesssim_{{\mathrm{P}}_n} 1,
\end{equation}
where $a \neq 0$ is any deterministic vector. Both of these conditions are mild. They can be verified using self-normalized
moderate deviation theorems and by using laws of large numbers for matrices
as discussed in the previous subsection.
\begin{lemma}[Elementary Adaptivity for Estimation via Approximate Sparsity]\label{lemma:est adapt approx sparsity}
Consider a sequence $\{{\mathrm{P}}_n\}$ for which the conditions of the previous lemma hold. In addition assume that the deviation
bound (\ref{eq:SNMD estimation}) holds and the sparse norm and pointwise norms of the second
derivatives matrices are stochastically bounded as in (\ref{eq:Gamma BDD}).
Then the adaptivity condition (\ref{eq:adaptivity estimation linear}) holds for the testing and estimation problem
in the affine-quadratic model.\end{lemma}
\section{Analysis of the IV Model with Very Many Control and Instrumental Variables}
Note that in the following
we write $w \perp v$ to denote $\operatorname{Cov}(w,v) = 0$.
Consider the linear instrumental variable model with response variable:
\begin{align}\label{ystructure}
\begin{array}{lll}
y_i = d_i'\alpha_0 + x_i' \beta_0 + \varepsilon_i, & {\mathrm{E}}[\varepsilon_i]= 0, & \varepsilon_i \perp (z_i, x_i),
\end{array}
\end{align}
where $y_i$ is the response variable, $d_i = (d_{ik})_{k=1}^{p^d}$ is a $p^d$-vector of endogenous variables,
such that
\begin{align}
\label{dstructure}
\begin{array}{lll}
d_{i1} \ = x_{i}'\gamma_{01} \ + z_{i}'\delta_{01} \ + u_{i1}, & \ {\mathrm{E}}[u_{i1}]= 0, & \ u_{i1} \perp (z_i, x_i), \\
\vdots & \vdots & \vdots \\
d_{ip^d} = x_{i}'\gamma_{0p^d} + z_{i}'\delta_{0p^d} + u_{ip^d}, & {\mathrm{E}}[u_{ip^d}]= 0, & u_{ip^d} \perp (z_i, x_i). \\
\end{array} \end{align}
Here $x_i=(x_{ij})_{j=1}^{p^x}$ is a $p^x$-vector of exogenous control variables, including a constant, and $z_i
=(z_i)_{i=1}^{p^z}$ is a $p^z$-vector of instrumental variables.
We will have $n$ \textit{i.i.d.} draws of $w_i = (y_i, d'_i, x'_i, z'_i)'$
obeying this system of equations. We also assume that $\operatorname{Var}(w_i)$ is finite throughout so that the model is well defined.
The parameter value $\alpha_0$ is our target. We allow $p^x= p_n^x \gg n$ and $p^z = p_n^z \gg n$,
but we maintain that $p^d$ is fixed in our analysis. This model includes the case of many instruments and small number of controls considered by \cite{BCCH12} as a special case, and the analysis readily accommodates the case of many controls and no instruments -- i.e. the linear
regression model -- considered by \cite{BCH2011:InferenceGauss,BelloniChernozhukovHansen2011} and \cite{ZhangZhang:CI}. For the latter, we simply set $p_n^z = 0$ and impose the additional condition $\varepsilon_i \perp u_i$
for $u_i = (u_{ij})_{j=1}^{p_d}$, which together with $\varepsilon_i \perp x_i$ implies that $\varepsilon_i \perp d_i$. We also note that the condition $\varepsilon_i \perp x_i, z_i$ is weaker than the condition ${\mathrm{E}}[\varepsilon_i|x_i,z_i]=0$, which allows for some misspecification of the model.
We may have that $z_i$ and $x_i$ are correlated so that $z_i$ are valid instruments only after controlling for $x_i$; specifically, we let $z_i = \Pi x_i + \zeta_i,$ for $\Pi$ a $p_n^z \times p_n^x$ matrix and $\zeta_i$ a $p_n^z$-vector of unobservables with $x_i \perp \zeta_{i}$. Substituting this expression for $z_i$ as a function of $x_i$ into (\ref{ystructure}) gives a system for $y_i$ and $d_i$ that depends only on $x_i$:
\begin{align}
\begin{array}{lll}
y_i \ = x_{i}'\theta_0 \ + \ \rho^y_i, & {\mathrm{E}}[\rho^y_i]\ = 0, & \rho^y_i \perp x_i, \\
\ & \ & \\
d_{i1} \ \ = x_{i}'\vartheta_{01} + \rho^d_{i1}, & {\mathrm{E}}[\rho^d_{i1}]= 0, & \rho^d_{i1} \perp x_i, \\
\vdots & \vdots & \vdots \\
d_{ip^{d}} = x_{i}'\vartheta_{0p^d} + \rho^d_{ip^d}, & {\mathrm{E}}[\rho^d_{ip^d}]= 0, & \rho^d_{ip^d} \perp x_i. \\
\end{array}
\end{align}
Because the dimension $p=p_n$ of $$\eta_0 = ({\theta}'_0, ({\vartheta}_{0k}', \gamma_{0k}', \delta_{0k}')_{k=1}^{p^d} )'$$ may be larger than $n$, informative estimation and inference about $\alpha_0$ is impossible without imposing restrictions on $\eta_0$.
To state our assumptions, we fix a collection of positive constants $(\mathsf{a}, \mathsf{A}, \mathsf{c}, \mathsf{C})$, where $\mathsf{a}>1$, and a sequence of constants $\delta_n \searrow 0$ and $\ell_n \nearrow \infty$. These constants will not vary with ${\mathrm{P}}$, but rather we will work with collections of ${\mathrm{P}}$ defined by these constants.
\textsc{Condition AS.1} \textit{ We assume that
$\eta_0$ is approximately sparse, namely that the decreasing rearrangement $(|\eta_0|^*_{j})_{j=1}^p $
of absolute values of coefficients $(|\eta_{0j}|)_{j=1}^p$ obeys}
\begin{align}
| \eta_0|^*_j \leq \mathsf{A} j^{-\mathsf{a}}, \ \ \mathsf{a} > 1, \ \ j = 1,...,p.
\end{align}
Given this assumption we can decompose $\eta_0$ into a
sparse component $\eta^m_0$ and small non-sparse component $ \eta^r_{0}$:
\begin{align}\label{decomposition}
\begin{array}{c}
\eta_0 = \eta^m_{0} + \eta^r_{0}, \ \mathrm{support}(\eta^m_{0}) \cap \mathrm{support}(\eta^r_{0}) = \emptyset, \\
\| \eta^m_{0}\|_0 \leq s, \ \|\eta_{0}^r\|_2 \leq c \sqrt{s/n}, \ \|\eta^r_{0}\|_1 \leq c \sqrt{s^2/n},\\
s = c n^{\frac{1}{2\mathsf{a}}},
\end{array}
\end{align}
where the constant $c$ depends only on $(\mathsf{a}, \mathsf{A})$.
\textsc{Condition AS.2} \textit{We assume that }
\begin{equation}
s^2 \log(pn)^2/n \leq o(1).
\end{equation}
We shall perform inference on $\alpha_0$ using the empirical analog of theoretical
equations:
\begin{align}\label{moment}
{\mathrm{M}}(\alpha_0, \eta_0) = 0, \quad {\mathrm{M}}(\alpha, \eta) := \textrm{E} \left[ \psi(w_i, \alpha, \eta)\right],
\end{align}
where $\psi= (\psi_k)_{k=1}^{p^d}$ is defined by
$$\psi_{k}(w_i, \alpha, \eta) := \left (y_i - x_{i}'\theta- \sum_{\bar k=1}^{p^d}(d_{i \bar k}-x_{i}' \vartheta_{\bar k}) \alpha_{\bar k}\right)
(x_{i}' \gamma_k+ z_{i}' \delta_k - x_{i}'\vartheta_k).$$
We can verify that the following orthogonality condition holds:
\begin{align}
\label{orthogonality} \partial_{\eta'} {\mathrm{M}}(\alpha_0, \eta) \Big \vert_{\eta= \eta_0} =0.
\end{align}
This means that missing the true value $\eta_0$ by a small amount does not invalidate the moment condition.
Therefore, the moment condition will be relatively insensitive to
non-regular estimation of $\eta_0$.
We denote the empirical analog of (\ref{moment}) as
\begin{align}
\label{emp moment}
\hat{\mathrm{M}}(\alpha, \hat \eta) = 0, \quad \hat{\mathrm{M}}(\alpha, \eta) := {\mathbb{E}_n} \left[ \psi_i(\alpha, \eta)\right].
\end{align}
Inference based on this condition can be shown to be immunized against small selection mistakes by virtue of orthogonality.
The above formulation is a special case of the linear-affine model. Indeed, here we have
$$
\mathrm{M}(\alpha, \eta) = \Gamma_1(\eta) \alpha + \Gamma_2(\eta), \quad \hat{\mathrm{M}}(\alpha, \eta) = \hat \Gamma_1(\eta) \alpha + \hat \Gamma_2(\eta),
$$ $$
\Gamma_1(\eta) = {\mathrm{E}}[ \psi^a(w_i, \eta )], \quad \hat \Gamma_1(\eta) = {\mathbb{E}_n}[\psi^a(w_i, \eta ) ],
$$ $$
\Gamma_2(\eta) = {\mathrm{E}}[ \psi^b(w_i, \eta )], \quad \hat \Gamma_2(\eta) = {\mathbb{E}_n}[\psi^b(w_i, \eta ) ],
$$
where $$\psi^a_{k, \bar k}(w_i, \eta) = - (d_{i\bar k}-x_{i}' \vartheta_{\bar k})(x_{i}' \gamma_k+ z_{i}' \delta_k - x_{i}'\vartheta_k) ,$$$$\psi^b_{k}(w_i, \eta) = (y_i - x_{i}'\theta)(x_{i}' \gamma_k+ z_{i}' \delta_k - x_{i}'\vartheta_k).$$
Consequently we can use the results of the previous section. In order to do so we need to
provide a suitable estimator for $\eta_0$. Here we use the Lasso and Post-Lasso estimators,
as defined in \cite{BCCH12}, to deal with non-normal errors and heteroscedasticity.
\begin{algorithm}[Estimation of $\eta_0$] (1) For each $k$, do Lasso or Post-Lasso Regression of $d_{ik}$ on $x_i, z_i$ to obtain $\hat \gamma_k$ and $\hat \delta_k$. (2) Do Lasso or Post-Lasso Regression of $y_i$ on $x_i$ to get $\hat \theta$. (3) Do Lasso or Post-Lasso Regression of $\hat d_{ik} = x_i'\hat\gamma_k + z_i'\hat\delta_k$ on $x_i$ to get $\hat \vartheta_k$. The estimator of $\eta_0$ is given by $\hat \eta = ({\hat \theta}', ({\hat \vartheta}_{k}', \hat \gamma_{0k}', \hat \delta_{k}')_{k=1}^{p^d} )'.$
\end{algorithm}
We then use
$$
\hat \Omega(\alpha, \hat \eta) = {\mathbb{E}_n} [\psi(w_i, \alpha, \hat \eta) \psi(w_i, \alpha, \hat \eta)'].
$$
to estimate the variance matrix $\Omega(\alpha, \eta_0)= {\mathbb{E}_n} [\psi(w_i, \alpha, \eta_0) \psi(w_i, \alpha, \eta_0)'].$
We formulate the orthogonal score statistic and the $C(\alpha)$-statistic,
\begin{equation}\label{eq:normality3}
S(\alpha) := \hat \Omega^{-1/2}_n(\alpha, \hat \eta)\sqrt{n}\hat{\mathrm{M}}(\alpha, \hat \eta), \quad
C(\alpha) = \| S(\alpha)\|^2,
\end{equation}
as well as our estimator $\hat \alpha$:
$$
\hat \alpha = \arg\min_{\alpha \in \mathcal{A}} \| \sqrt{n}\hat{\mathrm{M}}(\alpha, \hat \eta) \|^2.
$$
Note also that $\hat \alpha = \arg\min_{\alpha \in \mathcal{A}} C(\alpha)$ under mild conditions, since we work
with ``exactly identified" systems of equations. We also need to specify a variance estimator
$\hat V_n$ for the large sample variance $V_n$ of $ \hat \alpha$. We set $\hat V_n = (\hat \Gamma_1(\hat \eta)')^{-1}
\hat \Omega(\hat \alpha, \hat \eta) (\hat \Gamma_1(\hat \eta))^{-1}$.
To estimate the nuisance parameter we impose the following condition.
Let $f_{i} := (f_{ij})_{j=1}^{p_{f}} := (x_i', z_i')'$; $h_{i} : = (h_{il})_{l=1}^{p_h} := (y_i, d'_i, \bar d'_i)'$
where $\bar d_i = (\bar d_{ik})_{k=1}^{p^d}$ and $\bar d_{ik} := x_i'\gamma_{0k} + z_i'\delta_{0k}$;
$v_{i} = (v_{il})_{l=1}^{p_h} := (\varepsilon_i, \rho^y_i, {\rho_i^{d}}', {\varrho_i}')'$
where $\varrho_i = (\varrho_{ik})_{k=1}^{p^d}$ and $\varrho_{ik} := d_{ik} - \bar d_{ik}$.
Let $\tilde h_{i} := h_i - {\mathrm{E}}[h_i]$.
\textsc{Condition RF.} \textit{ (i) The eigenvalues of ${\mathrm{E}}[f_i f_i']$ are bounded from
above by $\mathsf{C}$ and from below by $\mathsf{c}$. For all $j$ and $l$,
(ii) ${\mathrm{E}}[h^2_{il}] +{\mathrm{E}}[ |f^2_{ij} \tilde h^2_{il}|] + 1/{\mathrm{E}}[f_{ij}^2 v_{il}^2] \leq \mathsf{C}$ and
${\mathrm{E}}[ |f^2_{ij} v^2_{il}|] \leq {\mathrm{E}}[ |f^2_{ij} \tilde h^2_{il}|]$, (iii) $
{\mathrm{E}}[|f_{ij}^3 v_{il}^3|]^2 \log^3 (p n)/n \leq \delta_n$, and (iv) $s \log (pn)/n \leq \delta_n$.
With probability no less than $1- \delta_n$, we have that (v) $ \max_{i \leq n, j} f_{ij}^2 [s^2 \log (pn)]/n \leq \delta_n
$ and $ \max_{l, j } |({\mathbb{E}_n} - {\mathrm{E}})[f_{ij}^2 v^2_{il}] | + | ({\mathbb{E}_n} - {\mathrm{E}})[f^2_{ij}\tilde h^2_{il}]| \leq \delta_n$
and (vi) $\|{\mathbb{E}_n}[f_i f_i']- {\mathrm{E}}[f_i f_i']\|_{\mathsf{sp}(\ell_n s)} \leq \delta_n$.}
The conditions are motivated by those given in \cite{BCCH12}.
The current conditions are made slightly stronger to account
for the fact that we use zero covariance conditions in formulating the moments. Some conditions
could be easily relaxed at a cost of more complicated exposition.
To estimate the variance matrix and establish asymptotic normality, we also need
the following condition. Let $\mathsf{q}>4$ be a fixed constant.
\textsc{Condition SM.} \textit{ For each $l$ and $k$, (i) $\displaystyle {\mathrm{E}} [|h_{il}|^\mathsf{q}] + {\mathrm{E}}[|v_{il}|^\mathsf{q}] \leq \mathsf{C}$, (ii) $ \mathsf{c} \leq {\mathrm{E}}[\varepsilon_i^2\mid x_i, z_i]\leq \mathsf{C}$, $\mathsf{c}< {\mathrm{E}}[{\varrho_{ik}^2}\mid x_i,z_i]\leq \mathsf{C}$ a.s., (iii) $\sup_{\alpha \in \mathcal{A}} \|\alpha\|_2 \leq \mathsf{C}$.}
Under the conditions set forth above, we have the following result on validity of post-selection
and post-regularization
inference using the $C(\alpha)$-statistic and estimators derived from it.
\begin{proposition}[Valid Inference in Large Linear Models using $C(\alpha)$-statistics]
\label{prop: IV Calpha} Let $\mathbf{P}_n$ be the collection of all ${\mathrm{P}}$ such that Conditions AS.1-2, SM, and RF hold for the given $n$. Then uniformly in ${\mathrm{P}} \in \mathbf{P}_n$,
$S(\alpha_0) \leadsto \mathcal{N}(0, I)$, and $C(\alpha_0) \leadsto \chi^2(p^d)$. As a consequence,
the confidence set $\mathsf{CR}_{1-a} = \{ \alpha \in \mathcal{A}: C(\alpha) \leq c(1-a) \}$, where
$c(1-a)$ is the $1-a$-quantile of a $\chi^2(p^d)$ is uniformly valid for $\alpha_0$, in the sense that
$$
\lim_{n \to \infty} \sup_{{\mathrm{P}} \in \mathbf{P}_n} |{\mathrm{P}} ( \alpha_0 \in \mathsf{CR}_{1-a} ) - (1-a) | = 0.
$$
Furthermore, for $V_n = (\Gamma_1')^{-1} \Omega(\alpha_0,\eta_0) (\Gamma_1)^{-1},$
we have that
$$
\lim_{n \to \infty} \sup_{{\mathrm{P}} \in \mathbf{P}_n} \sup_{R \in \mathcal{R}} |{\mathrm{P}} ( V_n^{-1/2} (\hat \alpha - \alpha_0) \in R) - \mathbb{P}( \mathcal{N}(0,I) \in R) | = 0,
$$
where $\mathcal{R}$ is the collection of all convex sets. Moreover, the result continues to apply if $V_n$ is replaced by $\hat V_n$. Thus, $ \mathsf{CR}^l_{1-a} = [
l'\hat \alpha \pm c(1-a/2) (l'\hat V_nl/n)^{1/2}]$, where $c(1-a/2)$ is the $(1-a/2)$-quantile
of a $\mathcal{N}(0,1)$, provides a uniformly valid confidence set for $l'\alpha_0$:
$$
\lim_{n \to \infty} \sup_{{\mathrm{P}} \in \mathbf{P}_n} |{\mathrm{P}} ( l'\alpha_0 \in \mathsf{CR}^l_{1-a} ) - (1-a) | = 0.
$$
\end{proposition}
\subsection{Simulation Illustration}
In this section, we provide results from a small Monte Carlo simulation to illustrate the performance of the estimator resulting from the application of Algorithm 1 in a small sample setting. As comparison, we report results from two commonly used
``unprincipled" alternatives for which uniformly valid inference over the class of approximately sparse models does not hold. Simulation parameters were chosen so that approximate sparsity holds but exact sparsity is violated in such a way that we expected the unprincipled procedures to perform poorly.
For our simulation, we generate data as $n$ iid draws from the model
\begin{align*}
\left. \begin{array}{ll}
y_i &= \alpha d_i + x_i' \beta + 2\varepsilon_i \\
d_i &= x_{i}'\gamma + z_{i}'\delta + u_i \\
z_i &= \Pi x_i + .125\zeta_i \\
\end{array} \right | \quad \left( \begin{array}{c} \varepsilon_i \\ u_i \\ \zeta_i \\ x_i \end{array} \right )
\sim \mathcal{N}\left(0 , \left( \begin{array}{cccc} 1 & .6 & 0 & 0 \\ .6 & 1 & 0 & 0 \\ 0 & 0 & I_{p_n^z} & 0 \\ 0 & 0 & 0 & \Sigma \end{array} \right) \right),
\end{align*}
where $\Sigma$ is a $p_n^x \times p_n^x$ matrix with $\Sigma_{kj} = (0.5)^{|j-k|}$ and $I_{p_n^z}$ is a $p_n^z \times p_n^z$ identity matrix. We set the number of potential controls variables ($p_n^x$) to 200, the number of instruments ($p_n^z$) to 150, and the number of observations ($n$) to 200. For model coefficients, we set $\alpha = 0$, $\beta = \gamma$ as $p_n^x-$vectors with entries $\beta_j = \gamma_j = 1/(9\nu)$, $\nu = {4/9 + \sum_{j = 5}^{p_n^x} 1/j^2}$ for $j \le 4$ and $\beta_j = \gamma_j = 1/(j^2\nu)$ for $j > 4$, $\delta$ as a $p_n^z-$vector with entries $\delta_j = \frac{3}{j^2}$, and $\Pi = [I_{p_n^z} \ , \ 0_{p_n^z \times (p_n^x-p_n^z)}].$ We report results based on 1000 simulation replications.
We provide results for four different estimators - an infeasible Oracle estimator that knows the nuisance parameters $\eta$ (Oracle), two naive estimators, and the proposed ``Double-Selection'' estimator. The results for the proposed ``Double-Selection'' procedure are obtained following Algorithm 1 using Post-Lasso at every step. To obtain the Oracle results, we run standard IV regression of $y_i - \textrm{E}[y_i|x_i]$ on $d_i-\textrm{E}[d_i|x_i]$ using the single instrument $\zeta_i' \delta$. The expected values are obtained from the model above and $\zeta_i' \delta$ provides the information in the instruments that is unrelated to the controls.
The two naive alternatives offer unprincipled, although potentially intuitive alternatives. The first naive estimator follows Algorithm 1 but replaces Lasso/Post-Lasso with stepwise regression with a p-value for entry of .05 and a p-value for removal of .10 (Stepwise). The second naive estimator (Non-orthogonal) corresponds to using a moment condition that does not satisfy the orthogonality condition described previously but will produce valid inference when perfect model selection in the regression of $d$ on $x$ and $z$ is possible or perfect model selection in the regression of $y$ on $x$ is possible and an instrument is selected in the $d$ on $x$ and $z$ regression.\footnote{Specifically, for the second naive alternative (Non-orthogonal), we first do Lasso regression of $d$ on $x$ and $z$ to obtain Lasso estimates of the coefficients $\gamma$ and $\delta$. Denote these estimates as $\hat\gamma_L$ and $\hat\delta_L$, and denote the indices of the coefficients estimated to be non-zero as $\hat{I}^d_x = \{j : \hat\gamma_{Lj} \ne 0\}$ and $\hat{I}^d_z = \{j : \hat\delta_{Lj} \ne 0\}$. We then run Lasso regression of $y$ on $x$ to learn the identities of controls that predict the outcome. We denote the Lasso estimates as $\hat\theta_L$ and keep track of the indices of the coefficients estimated to be non-zero as $\hat{I}^y_x = \{j : \hat\theta_{Lj} \ne 0\}$. We then take the union of the controls selected in either step $\hat{I}_x = \hat{I}^y_x \cup \hat{I}^d_x$. The estimator of $\alpha$ is then obtained as the usual 2SLS estimator of $y_i$ on $d_i$ using all selected elements from $x_i$, $x_{ij}$ such that $j \in \hat{I}_x$, as controls and the selected elements from $z_i$, $z_{ij}$ such that $j \in \hat{I}^d_z$, as instruments. }
All of the Lasso and Post-Lasso estimates are obtained using the data-dependent penalty level from \cite{BC-PostLASSO}. This penalty level depends on a standard deviation that is estimated adapting the iterative algorithm described in \cite{BCCH12} Appendix A using Post-Lasso at each iteration. For inference in all cases, we use standard t-tests based on conventional homoscedastic IV standard errors obtained from the final IV step performed in each strategy.
We display the simulation results in Figure \ref{SimulationDistributions}, and we report the median bias (Bias), median absolute deviation (MAD), and size of 5\% level tests (Size) for each procedure in Table \ref{SimulationTable}. For each estimator, we plot the simulation estimate of the sampling distribution of the estimator centered around the true parameter and scaled by the estimated standard error. With this standardization, usual asymptotic approximations would suggest that these curves should line up with a $\mathcal{N}(0,1)$ density function which is displayed as the bold solid line in the figure. We can see that the Oracle estimator and the Double-Selection estimator are centered correctly and line up reasonably well with the $\mathcal{N}(0,1)$, although both estimators exhibit some mild skewness. It is interesting that the sampling distributions of the Oracle and Double-Selection estimators are very similar as predicted by the theory. In contrast, both of the naive estimators are centered far from zero, and it is clear that the asymptotic approximation provides a very poor guide to the finite sample distribution of these estimators in the design considered.
\begin{figure}\label{SimulationDistributions}
\includegraphics[width=\columnwidth]{AnnualReviewSimHistograms}
\caption{The figure presents the histogram of the estimator from each method centered around the true parameters and scaled by the estimated standard error from the simulation experiment. The red curve is the pdf of a standard normal which will correspond to the sampling distribution of the estimator under the asymptotic approximation. Each panel is labeled with the corresponding estimator from the simulation.}
\end{figure}
\begin{table}
\begin{center}
\caption{Summary of Simulation Results for the Estimation of $\alpha$}\label{SimulationTable}
\begin{tabular}{l c c c}
\hline \hline
Method & Bias & MAD & Size \\
\hline
Oracle & 0.015 & 0.247 & 0.043 \\
Stepwise & 0.282 & 0.368 & 0.261 \\
Non-orthogonal & 0.084 & 0.112 & 0.189 \\
Double-Selection & 0.069 & 0.243 & 0.053 \\
\hline
\hline
\end{tabular}
\end{center}
\begin{flushleft}
\footnotesize{This table summarizes the simulation results from a linear IV model with many instruments and controls. Estimators include an infeasible oracle as a benchmark (Oracle), two naive alternatives (Stepwise and Non-orthogonal) described in the text, and our proposed feasible valid procedure (Double-Selection). Median bias (Bias), median absolute deviation (MAD), and size for 5\% level tests (Size) are reported.}
\end{flushleft}
\end{table}
The poor inferential performance of the two naive estimators is driven by different phenomena. The unprincipled use of stepwise regression fails to control spurious inclusion of irrelevant variables which leads to inclusion of many essentially irrelevant variables, resulting in many-instrument-type problems (e.g. \cite{NeweyEtAl-JIVE}). In addition, the spuriously included variables are those most highly correlated to the noise within sample which adds an additional type of ``endogeneity bias''. The failure of the ``Non-orthogonal'' method is driven by the fact that perfect model selection is not possible within the present design: Here we have
model selection mistakes in which the control variables that are correlated to the instruments but only moderately correlated to the outcome and endogenous variable are missed. Such exclusions result in standard omitted variables bias in the estimator for the parameter of interest and substantial size distortions. The additional step in the Double-Selection procedure can be viewed as a way to guard against such mistakes. Overall, the results illustrate the uniformity claims made in the preceding section. The feasible Double-Selection procedure following from Algorithm 1 performs similarly to the semi-parametrically efficient infeasible Oracle. We obtain good inferential properties with the asymptotic approximation providing a fairly good guide to the behavior of the estimator despite working in a setting in which perfect model selection is impossible. Although simply illustrative of the theory, the results are reassuring and in line with extensive simulations in the linear model with many controls provided in \cite{BelloniChernozhukovHansen2011}, in the instrumental variables model with many instruments and a small number of controls provided in \cite{BCCH12}, and in linear panel data models provided in \cite{BCHK:FE}.
\subsection{Empirical Illustration: Logit Demand Estimation}
As further illustration of the approach, we provide a brief empirical example in which we estimate the coefficients in a simple logit model of demand for automobiles using market share data. Our example is based on the data and most basic strategy from \cite{BLP}. Specifically, we estimate the parameters from the model
\begin{align*}
\log(s_{it}) - \log(s_{0t}) &= \alpha_0 p_{it} + x_{it}'\beta_0 + \varepsilon_{it}, \\
p_{it} &= z_{it}'\delta_0 + x_{it}'\gamma_0 + u_{it},
\end{align*}
where $s_{it}$ is the market share of product $i$ in market $t$ with product zero denoting the outside option, $p_{it}$ is price and is treated as endogenous, $x_{it}$ are observed included product characteristics, and $z_{it}$ are instruments. One could also adapt the proposed variable selection procedures to extensions of this model such as the nested logit model or models allowing for random coefficients; see, e.g., \cite{BLPLASSO} for an example with a random coefficient.
In our example, we use the same set of product characteristics ($x$-variables) as used in obtaining the basic results in \cite{BLP}. Specifically, we use five variables in $x_{it}$: a constant, an air conditioning dummy, horsepower divided by weight, miles per dollar, and vehicle size. We refer to these five variables as the baseline set of controls.
We also adopt the argument from \cite{BLP} to form our potential instruments. \cite{BLP} argue that that characteristics of other products will satisfy an exclusion restriction, $\textrm{E}[\varepsilon_{it}|x_{j\tau}] = 0$ for any $\tau$ and $j \ne i$, and thus that any function of characteristics of other products may be used as instrument for price. This condition leaves a very high-dimensional set of potential instruments as any combination of functions of $\{x_{j\tau}\}_{j \ne i, \tau \ge 1}$ may be used to instrument for $p_{it}$. To reduce the dimensionality, \cite{BLP} use intuition and an exchangeability argument to motivate consideration of a small number of these potential instruments formed by taking sums of product characteristics formed by summing over products excluding product $i$. Specifically, we form baseline instruments by taking
$$
z_{k,it} = \left(\sum_{r \ne i, r \in \mathcal{I}_f} x_{k,rt} , \sum_{r \ne i, r \notin \mathcal{I}_f} x_{k,rt} \right)
$$
where $x_{k,it}$ is the $k^{th}$ element of vector $x_{it}$ and $\mathcal{I}_f$ denotes the set of products produced by firm $f$.
This choice yields a vector $z_{it}$ consisting of 10 instruments. We refer to this set of instruments as the baseline instruments.
Although the choice of the baseline instruments and controls is motivated by good intuition and economic theory, we note that theory does not clearly state which product characteristics or instruments should be used in the model. Theory also fails to indicate the functional form with which any such variables should enter the model. The high-dimensional methods outlined in this paper offer one strategy to help address these concerns that complements the economic intuition motivating the baseline controls and instruments. As an illustration, we consider an expanded set of controls and instruments. We augment the set of potential controls with all first order interactions of the baseline variables, quadratics and cubics in all continuous baseline variables, and a time trend that yields a total of 24 $x$-variables. We refer to these as the augmented controls. We then take sums of these characteristics as potential instruments following the original strategy which yields 48 potential instruments.
We present estimation results in Table \ref{EmpiricalTable}. We report results obtained by applying the method outlined in Algorithm 1 using just the baseline set of five product characterstics and 10 instruments in the row labeled ``Baseline 2SLS with Selection'' and results obtained by applying the method to the augmented set of 24 controls and 48 instruments in the row labeled ``Augmented 2SLS with Selection.'' In each case, we apply the method outlined in Algorithm 1 using post-Lasso in each step and forcing the intercept to be included in all models. We employ the heteroscedasticity robust version of Post-Lasso of \cite{BCCH12} following the implementation algorithm provided in Appendix A of \cite{BCCH12}. For comparison, we also report OLS and 2SLS estimates using only the baseline variables in ``Baseline OLS'' and ``Baseline 2SLS,'' respectively; and we report OLS and 2SLS estimates using the augmented variable set in ``Augmented OLS'' and ``Augmented 2SLS,'' respectively. All standard errors are conventional heteroscedasticity robust standard errors.
\begin{table}
\begin{center}
\caption{Estimates of Price Coefficient}\label{EmpiricalTable}
\begin{tabular}{l c c c}
\hline\hline
& Price Coefficient & Standard Error & Number Inelastic \\
\cline{2-4}
\text{} & \multicolumn{3}{c}{\textit{Estimates Without Selection}} \\
Baseline OLS & -0.089 & 0.004 & 1502 \\
Baseline 2SLS & -0.142 & 0.012 & 670 \\
Augmented OLS & -0.099 & 0.005 & 1405 \\
Augmented 2SLS & -0.127 & 0.014 & 874 \\
\text{} & \multicolumn{3}{c}{\textit{2SLS Estimates With ``Double Selection"}} \\
Baseline 2SLS Selection & -0.185 & 0.014 & 139 \\
Augmented 2SLS Selection & -0.221 & 0.015 & 12 \\
\hline
\hline
\end{tabular}
\end{center}
\begin{flushleft}
\footnotesize{This table reports estimates of the coefficient on price (``Price Coefficient'') along with the estimated standard error (``Standard Error'') obtained using different sets of controls and instruments. The rows ``Baseline OLS'' and ``Baseline 2SLS'' respectively provide OLS and 2SLS results using the baseline set of variables (5 controls and 10 instruments) described in the text. The rows ``Augmented OLS,'' ``Augemented 2SLS ''are defined similarly but use the augmented set of variables described in the text (24 controls and 48 instruments). The rows ``Baseline 2SLS with Selection'' and ``Augmented 2SLS with Selection''
applies the ``double selection" approach developed in this paper to select a set of controls and instruments and perform valid post-selection inference about the estimated price coefficient where selection occurs considering only the baseline variables. For each procedure, we also report the point estimate of the number of products for which demand is estimated to be inelastic in the column ``Number Inelastic.''}
\end{flushleft}
\end{table}
Considering first estimates of the price coefficient, we see that the estimated price coefficient increases in magnitude as we move from OLS to 2SLS and then to the selection based results. After selection using only the original variables, we estimate the price coefficient to be -.185 with an estimated standard error of .014 compared to an OLS estimate of -.089 with estimated standard error of .004 and 2SLS estimate of -.142 with estimated standard error of .012. In this case, all five controls are selected in the log-share on controls regression, all five controls but only four instruments are selected in the price on controls and instruments regression, and four of the controls are selected for the price on controls relationship. The difference between the baseline results is thus largely driven by the difference in instrument sets. The change in the estimated coefficient is consistent with the wisdom from the many-instrument literature that inclusion of irrelevant instruments biases 2SLS toward OLS.
With the larger set of variables, our post-model-selection estimator of the price coefficient is -.221 with an estimated standard error of .015 compared to the OLS estimate of -.099 with an estimated standard error of .005 and 2SLS estimate of -.127 with an estimated standard error of .014. Here, we see some evidence that the original set of controls may have been overly parsimonious as we select some terms that were not included in the baseline variable set. We also see a closer agreement between the OLS estimate and 2SLS estimate without selection which is likely driven by the larger number of instruments considered and the usual bias towards OLS seen in 2SLS with many weak or irrelevant instruments. In the log-share on controls regression, we have eight control variables selected; and we have seven controls and only four instruments selected in the price on controls and instrument regression. We also have 13 variables selected for the price on controls relationship. The selection of these additional variables suggests that there is important nonlinearity missed by the baseline set of variables.
The most interesting feature of the results is that estimates of own-price elasticities become more plausible as we move from the baseline results to the results based on variable selection with a large number of controls. Recall that facing inelastic demand is inconsistent with profit maximizing price choice within the present context, so theory would predict that demand should be elastic for all products. However, the baseline point estimates imply inelastic demand for 670 products. When we use the larger set of instruments without selection, the number of products for which we estimate inelastic demand increases to 874 with the increase generated by the 2SLS coefficient estimate moving back towards the OLS estimate.
The use of the variable selection results provides results closer to the theoretical prediction. The point estimates based on selection from only the baseline variables imply inelastic demand for 139 products, and we estimate inelastic demand for only 12 products using the results based on selection from the larger set of variables. Thus,
the new methods provide the most reasonable estimates of own-price elasticities.
We conclude by noting that the simple specification above suffers from the usual drawbacks of the logit demand model. However, the example illustrates how the application of the methods we outlined may be used in the estimation of structural parameters in economics and adds to the plausibility of the resulting estimates. In this example, we see that we obtain more sensible estimates of key parameters with at most a modest cost in increased estimation uncertainty after applying the methods in this paper while considering a flexible set of variables.
\section{Overview of Related Literature}
Inference following model selection or regularization more generally has been an active area of research in econometrics and statistics for the last several years. In this section, we provide a brief overview of this literature highlighting some key developments. This review is necessarily selective due to the large number of papers available and the rapid pace at which new papers are appearing. We choose to focus on papers that deal specifically with high-dimensional nuisance parameter settings, and note that the ideas in these papers apply in low dimensional settings as well.
Early work on inference in high-dimensional settings focused on inference based on the so-called perfect recovery; see, e.g., \cite{FanLi2001} for an early paper, \cite{FanLv2010:Review} for a more recent review, and \cite{bulman:book} for a textbook treatment. A consequence of this property is that model selection does not impact the asymptotic distribution of the parameters estimated in the selected model. This feature allows one to do inference using standard approximate distributions for the parameters of the selected model ignoring that model selection was done. While convenient and fruitful in many applications (e.g. signal processing), such results effectively rely on strong conditions that imply that one will be able to perfectly select the correct model. For example, such results in linear models require the so called ``beta-min condition" (\cite{bulman:book}) that all but a small number of coefficients are exactly zero and the remaining non-zero coefficients are bounded away from zero, effectively ruling out variables that have small, non-zero coefficients. Such conditions seem implausible in many applications, especially in econometrics, and relying on such conditions produces asymptotic approximations that may provide very poor approximations to finite-sample distributions of estimators as they are not uniformly valid over sequences of models that include even minor deviations from conditions implying perfect model selection. The concern about the lack of uniform validity of inference based on oracle properties was raised in a series of papers, including \cite{leeb:potscher:review} and \cite{leeb:potscher:hodges} among many others,
and the more recent work on post-model-selection inference has been focused on offering procedures that provide uniformly valid inference over interesting (large) classes of models that include cases where perfect model selection will not be possible.
To our knowledge, the first work to formally and expressly address the problem of obtaining uniformly valid inference following model selection is \cite{BellChernHans:Gauss} which considered inference about parameters on a low-dimensional set of endogenous variables following selection of instruments from among a high-dimensional set of potential instruments in a homoscedastic, Gaussian instrumental variables (IV) model. The approach does not rely on implausible ``beta-min" conditions which imply perfect model selection but instead relies on the fact that the moment condition underlying IV estimation satisfies the \textit{orthogonality condition} (\ref{eq:ortho}) and the use of high-quality variable selection methods. These ideas were further developed in the context of providing uniformly valid inference about the parameters on endogenous variables in the IV context with many instruments to allow non-Gaussian heteroscedastic disturbances in \cite{BCCH12}. These principles have also been applied in \cite{BCH2011:InferenceGauss}, who developed approaches for regression and IV models with Gaussian errors; \cite{BelloniChernozhukovHansen2011} (ArXiv 2011), which covers estimation of the parametric components of the partially linear model, estimation of average treatment effects, and provides a formal statement of the orthogonality condition (\ref{eq:ortho}); \cite{Farrell:JMP} which covers average treatment effects with discrete, multi-valued treatments; \cite{Kozbur:JMP} which covers additive nonparametric models; and \cite{BCHK:FE} which extends the IV and partially linear model results to allow for fixed effects panel data and clustered dependence structures. The most recent, general approach is provided in \cite{BCFH:Policy} where inference about parameters defined by a continuum of orthogonalized estimating equations with infinite-dimensional nusiance parameters is analyzed and positive results on inference are developed. The framework in \cite{BCFH:Policy} is general enough to cover the aforementioned papers and many other parametric and semi-parametric models considered in economics.
As noted above, providing uniformly valid inference following model selection is closely related to use of Neyman's $C(\alpha)$-statistic. Valid confidence regions can be obtained by inverting tests based on these statistics, and minimizers of $C(\alpha)$-statistics may be used as point estimators. The use of $C(\alpha)$ statistics for testing and estimation in high-dimensional approximately sparse models was first explored in the context of high-dimensional quantile regression in \cite{BCK-LAD} (Oberwolfach, 2012) and \cite{BCK-QR} and in the context of high-dimensional logistic regression and other high-dimensional generalized linear models by \cite{BCY-honest}. More recent uses of $C(\alpha)$-statistics (or close variants, under different names) include those in \cite{witten:score}, \cite{hanliu1}, and \cite{hanliu2} among others.
There have also been parallel developments based upon ex-post ``de-biasing'' of estimators. This approach is mathematically equivalent to doing classical ``one-step" corrections in the general framework of Section 2.
Indeed, while at first glance this ``de-biasing" approach may appear distinct from that taken in the papers listed above in this section, it is the same as approximately solving -- by doing one Gauss-Newton step -- orthogonal estimating equations satisfying (\ref{eq:ortho}).
The general results of Section 2 suggest that these approaches -- the exact solving and ``one-step" solving -- are generally first-order asymptotically equivalent, though higher-order differences may persist. To the best of our knowledge, the ``one-step" correction approach was first
employed in high-dimensional sparse models by \cite{ZhangZhang:CI} (ArXiv 2011) which covers the homoscedastic linear model (as well as in several follow-up works by the authors). This approach has been further used in \cite{vdGBRD:AsymptoticConfidenceSets} (ArXiv 2013) which covers homoscedastic linear models and some generalized linear models, and \cite{JM:ConfidenceIntervals} (ArXiv 2013) which offers a related, though somewhat different approach.
Note that \cite{BCK-LAD} and \cite{BCY-honest} also offer results on ``one-step" corrections as part of their analysis
of estimation and inference based upon the orthogonal estimating equations. We would not expect that the use of orthogonal estimating equations or the use of ``one-step" corrections to dominate each other in all cases, though computational evidence in \cite{BCY-honest} suggests that the use of exact solutions to orthogonal estimating equations may be preferable to approximate solutions obtained from ``one-step" corrections in the contexts considered in that paper.
Another branch of the recent literature takes a complementary, but logically distinct, approach that aims at doing valid inference for the parameters of a ``pseudo-true'' model that results from the use of a model selection procedure, see \cite{BerkEtAl13}.
Specifically, this approach conditions on a model selected by a data-dependent rule and then attempts to do inference -- conditional on the selection event -- for the parameters of the selected model, which may deviate from the ``true'' model that generated the data. Related developments within this approach appear in \cite{AdaptiveGraphLasso}, \cite{LeeTaylorScreening}, \cite{LeeEtAlLasso}, \cite{LockhartEtAlLasso}, \cite{LoftusStepwise}, \cite{TaylorEtAlAdaptive}, and \cite{FST14}.
It seems intellectually very interesting to combine the developments of the present paper (and other preceding papers cited above) with developments in this literature.
The previously mentioned work focuses on doing inference for low dimensional parameters in the presence of high dimensional nuisance parameters. There have also been developments on performing inference for high dimensional parameters.
\cite{Chern:SG} proposed inverting a Lasso performance bound in order to construct a simultaneous, Scheff\'{e}-style
confidence band on all parameters. An interesting feature of this approach is that it uses
weaker design conditions than many other approaches but requires the data analyst to supply explicit bounds on restricted eigenvalues. \cite{GautierTsybakovHDIV} (ArXiv 2011)
and \cite{CCK:AOS13} employ similar ideas while also working with various generalizations of restricted eigenvalues.
\cite{vdg:nickl} construct confidence ellipsoids for the entire parameter vector using sample splitting ideas.
Somewhat related to this literature are the results of \cite{BCK-LAD} who use the orthogonal
estimating equations framework with infinite-dimensional nuisance parameters and construct a simultaneous confidence rectangle
for many target parameters where the number of target parameters could be much larger than the sample size.
They relied upon the high-dimensional central limit theorems and bootstrap results established in \cite{CCK:AOS13}.
Most of the aforementioned results rely on (approximate) sparsity and related sparsity-based estimators. Some examples of the use of alternative regularization schemes are available in the many instrument literature in econometrics. For example, \cite{chamberlain:imbens:reiv} use a shrinkage estimator resulting from use of a Gaussian random coefficients structure over first-stage coefficients, and \cite{okui:manyiv} uses ridge regression for estimating the first-stage regression in a framework where the instruments may be ordered in terms of relevance. \cite{carrasco:regularizedIV} employs a different strategy based on directly regularizing the inverse that appears in the definition of the 2SLS estimator allowing for a number of moment conditions that are larger than the sample size; see also \cite{carrasco:LIML}. The theoretical development in \cite{carrasco:regularizedIV} relies on restrictions on the covariance structure of the instruments rather than on the coefficients of the instruments. \cite{RJIVE} considers a combination of ridge-regularization and the jackknife to provide a procedure that is valid allowing for the number of instruments to be greater than the sample size under weak restrictions on the covariance structure of the instruments and the first-stage coefficients. In all cases, the orthogonality condition holds allowing root-$n$ consistent and asymptotically normal estimation of the main parameter $\alpha$.
Many other interesting procedures beyond those mentioned in this review have been developed for estimating high-dimensional models; see, e.g. \cite{elements:book} for a textbook review. Developing new techniques for estimation in high-dimensional settings is also still an active area of research, so the list of methods available to researchers continues to expand. The use of these procedures and the impact of their use on inference about low-dimensional parameters of interest is an interesting research direction to explore. It seems likely that many of these procedures will provide sufficiently high-quality estimates that they may be used for estimating the high-dimensional nuisance parameters $\eta$ in the present setting.