The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
72,963 characters
Choosing the Dictionary and Penalty for IV-LASSO
\global\let\@title
\global\let\@author
\global\let\@date
\global\let\and
\@setupmaketitle
\begin{abstract}
Estimating the first stage of an instrumental variables (IV) model with the least
absolute shrinkage and selection operator (LASSO)
requires choosing a dictionary of technical instruments and a penalty level.
First-order asymptotic theory offers no guidance on these choices, as any consistent implementation yields a structural parameter estimator with the same limiting distribution. In finite samples, however, these choices can have a substantial impact on the resulting structural parameter estimate. Working in a model with a single endogenous regressor and homoskedastic Gaussian errors, we use first- and second-order Stein identities to derive the approximate mean squared error (AMSE) of the instrumental-variables LASSO (IV-LASSO) estimator, which can be consistently estimated and used to rank a prespecified list of dictionary--penalty candidates. The AMSE reveals a bias--variance trade-off: more complex first-stage fits better approximate the conditional mean of the endogenous variable but are also more correlated with the structural errors, with complexity measured by the degrees of freedom of the LASSO fit. The weight on this bias rises with the endogeneity of the regressor, a quantity that neither plug-in nor cross-validation penalty rules take into account.
Despite the AMSE being derived in a Gaussian model, penalty selection by minimizing the feasible AMSE criterion delivers up to a one-third lower mean squared error compared to cross-validation and plug-in penalty rules in Gaussian and non-Gaussian simulation designs calibrated to the data of \citet{gilchrist-glassberg-2016}.
\end{abstract}
\begin{center}
\textsc{Keywords:} Instrumental Variables, LASSO, Model Selection,
Second-Order Asymptotics\\[1mm]
\textsc{JEL Codes:} C26, C36, C52
\end{center}
\section{Introduction}
\label{sec:introduction}
When a large set of technical instruments is available, LASSO makes estimation of an
IV first stage feasible but leaves applied researchers to choose how heavily to penalize the fit and which dictionary of technical instruments to use. Under standard sparse-first-stage conditions, \citet{BCCH-2012} show that all sufficiently accurate LASSO fits yield IV estimators with the same limiting distribution. Standard asymptotic theory
therefore does not rank admissible implementations, even though their
finite-sample performance can differ. \Cref{fig:bcch-paths-intro} illustrates
this issue in the eminent-domain application of \citet{BCCH-2012}. Across
three nested first-stage dictionaries, the IV estimate changes materially
along the penalty path. Moreover, different methods of selecting the first-stage penalty parameter can change whether the 95\% confidence interval for the resulting structural parameter estimate contains zero.
For example, for the Case-Shiller home price index outcome, using cross-validation \citep{ChetverikovLiaoChernozhukov-2021} yields a significant result on a restricted instrument dictionary while using Stein's Unbiased Risk Estimate (SURE, \citet{TibshiraniTaylor-2012}) does not.
Given this variation, how should an applied researcher choose among possible first-stage implementations?
\begin{figure}[!htbp]
\centering
\includegraphics[width=\textwidth]{figures/bcch_full_path_intro}
\caption{IV Estimates and Conventional Significance across First-Stage
Implementations}
\label{fig:bcch-paths-intro}
\par\medskip\parbox{\linewidth}{\footnotesize\emph{Notes:}\ Lines trace IV-LASSO estimates over LASSO penalties within three nested dictionaries. Original is the published dictionary; Restricted removes cross-products, and Expanded adds all nondegenerate second-order
products. Filled markers denote pointwise 95\% heteroskedasticity robust confidence intervals that exclude zero. The solid gray line and dashed bounds show the published IV-post-LASSO estimate and confidence interval. The panels report the two outcomes for which that interval excludes zero (183 Case--Shiller and 110 non-metro circuit-year observations).}
\end{figure}
Existing first-stage tuning rules are not designed to maximize the precision of resulting structural parameter estimates. The calibrated penalty rule of \citet{BCCH-2012} and the bootstrap-after-cross-validation method of \citet{ChetverikovSorensen-2025} are designed to control first-stage estimation error, while
SURE and ordinary cross-validation target first-stage prediction
\citep{TibshiraniTaylor-2012,ChetverikovLiaoChernozhukov-2021}. These
objectives do not account for how first-stage estimation noise enters the
structural estimator. A LASSO fit with a lighter penalty or richer dictionary may better approximate the true first-stage, but a more complex fit will also be more correlated with the endogenous variable. When the first-stage and structural
errors are correlated, this in turn introduces bias into the structural estimate. The optimal first-stage implementation therefore depends not only on goodness of first-stage fit and instrument strength, but also on the level of endogeneity.
This paper develops a principled method for choosing the first-stage implementation by developing second-order asymptotic theory for the full-sample IV-LASSO estimator. We focus on the leading case of a single endogenous regressor and work in a homoskedastic Gaussian model. Following the higher-order tradition of \citet{Nagar-1959} and \citet{DonaldNewey-2001}, we prespecify a finite list of first-stage LASSO dictionary--penalty candidates and rank them by the approximate mean squared error (AMSE) of their resulting IV-LASSO estimators. The candidate specific part of the AMSE can be decomposed into a first-stage approximation cost and a complexity cost, with complexity measured by a collinearity-robust analogue of the number of instruments selected by the LASSO. The size of this complexity cost rises with the squared covariance
between the first-stage and structural errors, reflecting the discussion above. Our characterization yields a
feasible criterion that can be constructed from each candidate's fitted first-stage. Minimizing this criterion selects an estimator whose AMSE is as small as that of the best candidate in the list, up to vanishing relative error.
Deriving the criterion is nonstandard. Unlike ordinary least-squares, the LASSO coefficient estimate does not yield a closed-form expression and the LASSO fit is not a fully differentiable function of the data. These features prevent standard tools from being applied to analyze higher-order terms in the asymptotic expansion. We overcome these obstacles by leveraging a geometric characterization of the LASSO-fit from \citet{TibshiraniTaylor-2012} along with newly developed second-order Stein identities from \citet{BellecZhang-2021} to examine otherwise problematic second-order terms. In addition to the distributional restrictions stated above, the result requires standard sparsity conditions to ensure convergence of the first-stage estimates as well as a partial beta-min condition which allows for consistent estimation of the complexity of the LASSO fits. This partial beta-min condition requires that a growing core of first-stage coefficients must be sufficiently strong and is notably weaker than previous beta-min conditions in the literature on model selection via the LASSO \citep{ZhaoYu-2006,Wainwright-2009,BelloniChernozhukov-2013}.
We evaluate the performance of the proposed criterion through a simulation study calibrated to the data of \citet{gilchrist-glassberg-2016}, who use weather conditions to instrument for opening-weekend movie ticket sales. We hold the original instruments fixed and vary the strength of identification and degree of endogeneity.
At low levels of endogeneity, cross-validation and the proposed criterion perform comparably, in line with the theory developed in this paper. As the level of endogeneity rises, however, cross-validation fails to account for the bias passed on to the second-stage estimator. In contrast, our proposed criterion adjusts for the endogeneity and selects first-stages with lower levels of complexity, resulting in structural parameter estimates with up to a third lower mean-squared error than their cross-validation counterparts in designs with the highest levels of endogeneity. Interestingly, the relative performance of the criterion is essentially the same in both Gaussian and non-Gaussian designs. This suggests that, although our theory is developed in a Gaussian model, the AMSE of the IV-LASSO estimator in a general model may be well approximated by its AMSE in the Gaussian model.
We contribute to a growing literature on the use of high-dimensional or ``machine-learning'' methods in econometric analysis \citep{BelloniChernozhukovHansen-2014,CCDDHNR-2018,WagerAthey-2018,FarrellLiangMisra-2021} as well as to an established literature analyzing higher-order properties of econometric estimators \citep{Sargan1976,Rothenberg1984,NeweySmith2004,HahnHausmanKuersteiner2004}. A common feature of double/debiased machine learning estimators is that all suitably convergent nuisance parameter estimation procedures yield the same asymptotic distribution for the resulting target parameter estimator. Existing asymptotic theory therefore provides little guidance on how to choose among implementations satisfying certain high-level conditions. To our knowledge, we provide the first second-order mean squared error approximation for an econometric estimator with a high-dimensional LASSO first stage.\footnote{\citet{Velez-2026} develops second-order approximations for double/debiased machine learning estimators to study the choice of the number of cross-fitting folds. His higher-order analysis requires the first-stage estimators to admit a stochastic linear expansion. Our analysis instead directly accounts for the shrinkage and variable selection induced by the LASSO.} We show that the higher-order terms distinguishing alternative first-stage implementations can be consistently estimated and use them to select the first-stage dictionary and penalty. Our results thus extend the literature on selecting and regularizing instruments according to the approximate mean squared error of the structural parameter estimator \citep{DonaldNewey-2001,KuersteinerOkui-2010,Okui-2011,Carrasco-2012} to high-dimensional LASSO first stages.
The rest of this paper proceeds as follows. \Cref{sec:framework} sets
up the model, the candidate first stages, and the effective dimension.
\Cref{sec:results} states the conditions and the two main results.
\Cref{sec:empirical-application} reports the calibrated simulation
study. All proofs are deferred to the appendix.
\emph{Notation.} For \(m\in\mathbb N\), let \([m]:=\{1,\ldots,m\}\). For
an array \(\{a_i\}_{i\in[n]}\), define
\begin{equation*}
\mathbb{E}_n[a_i]:=\frac1n\sum_{i\in[n]}a_i,
\qquad
\overline{\mathbb{E}}[a_i]:=\mathbb{E}_n[\mathbb{E}[a_i]].
\end{equation*}
\section{Model and Setup}
\label{sec:framework}
To simplify exposition, we present the analysis assuming that any exogenous regressors have already been partialled out of both the outcome and the endogenous regressor.\footnote{An appendix providing a formal treatment of the model with controls is available on request.} Our analysis assumes there is only a single endogenous variable and does not cover the general case where \(x_i\) is vector-valued. Along with the first-stage, the IV model can be written as a system of simultaneous equations:
\begin{equation}
\label{eq:model}
\begin{split}
y_i &= \beta x_i + \varepsilon_i \\
x_i &= \Pi_i + v_i
\end{split}
\end{equation}
We assume the researcher observes \(n\) independent observations of the outcome \(y_i \in \SR\), the endogenous variable \(x_i \in \SR\), and instruments \(z_i \in \SR^{p}\) but neither the structural error \(\varepsilon_i \in \SR\) nor the first-stage errors \(v_i \in \SR\). The structural error \(\varepsilon_i \in \SR\) is assumed to be conditional-mean independent of the instruments, \(\mathbb{E}[\varepsilon_i|z_i] = 0\) and we let \(\Pi_i = \mathbb{E}[x_i | z_i]\) denote the conditional expectation of the endogenous variable given the instruments so that \(\mathbb{E}[v_i | z_i] = 0\). We collect the random variables in the vectors
\begin{align*}
x &\coloneqq (x_1,\dots,x_n)' \in \SR^n &
\Pi &\coloneqq (\Pi_1,\dots,\Pi_n)' \in \SR^n \\
\varepsilon &\coloneqq (\varepsilon_1,\dots,\varepsilon_n)' \in \SR^n &
v &\coloneqq (v_1,\dots,v_n)'\in \SR^n
\end{align*}
Throughout the theoretical analysis, we condition on the realized instruments \(\{z_i\}_{i=1}^n\) and treat them as fixed. All expectations and probability statements below are therefore understood as conditional on this realization, i.e, \(\Pi_i = \mathbb{E}[x_i]\) and \(\mathbb{E}[\varepsilon_i] = \mathbb{E}[v_i] = 0\). We will assume strong identification and conditionally homoskedastic Gaussian errors. Formally,
\begin{assumption}[Conditional Gaussian Sampling]
\label{assm:gaussian} Assume
\begin{enumerate}[(i)]
\item The errors \(\{(\varepsilon_i, v_i)\}_{i=1}^n\) are independent of the instruments and independently and identically distributed according to
\begin{equation}
\label{eq:gaussian-sampling}
\begin{pmatrix}\varepsilon_i\\v_i\end{pmatrix}
\sim N(0,\Omega), \qquad
\Omega :=\begin{pmatrix} \sigma_\varepsilon^2&\sigma_{\varepsilon v}\\ \sigma_{\varepsilon v}&\sigma_v^2
\end{pmatrix},
\end{equation}
where \(\Omega\) is a fixed nonsingular matrix.
\item For constants \(0<\underline H<\overline H<\infty\),
\begin{equation}
\label{eq:strength-endogeneity}
\underline H\le H:=\mathbb{E}_n[\Pi_i^2]\le\overline H, \qquad \sigma_{\varepsilon v}\ne0.
\end{equation}
\end{enumerate}
\end{assumption}
\Cref{assm:gaussian}(i) encodes the assumption of conditionally homoskedastic Gaussian errors while the first part of \Cref{assm:gaussian}(ii) encodes the strong identification assumption. The second part of \Cref{assm:gaussian}(ii) is a mild technical condition that assumes that there is an endogeneity problem to begin with, i.e \(\beta\) cannot be recovered from a simple regression of the outcome on the endogenous variable.
Under conditional homoskedasticity, the conditional mean \(\Pi_i\) is the optimal instrument for estimating \(\beta\): among IV estimators based on a nonrandom instrument, using \(\Pi_i\) itself minimizes the variance of the leading term.\footnote{For nonrandom \(g\in\SR^n\) with \(\mathbb{E}_n[g_i^2]\) bounded and \(\mathbb{E}_n[g_i\Pi_i]\) bounded away from zero, the IV estimator \(\widehat\beta_g:=\mathbb{E}_n[g_iy_i]/\mathbb{E}_n[g_ix_i]\) satisfies \(\sqrt n(\widehat\beta_g-\beta) =\sqrt n\,\mathbb{E}_n[g_i\varepsilon_i]/\mathbb{E}_n[g_i\Pi_i]+o_p(1)\), and the leading term is distributed exactly as \(N(0,\sigma_\varepsilon^2\mathbb{E}_n[g_i^2]/\mathbb{E}_n[g_i\Pi_i]^2)\). By Cauchy--Schwarz this variance is at least \(\sigma_\varepsilon^2/H\), with equality if and only if \(g\) is proportional to \(\Pi\). When the instruments are instead modeled as random draws, \(\sigma_\varepsilon^2/\mathbb{E}[\Pi_i^2]\) is the semiparametric efficiency bound for \(\beta\) under conditional homoskedasticity \citep{Chamberlain-1987,Newey-1990}; \(\sigma_\varepsilon^2/H\)
is its fixed-design counterpart.} Since \(\Pi_i\) is unknown, feasible estimation replaces it with a first-stage estimate \(\widehat\Pi_i\). First-order asymptotics provide no guidance on this choice. Any estimator \(\widehat\beta\) based on a sufficiently well-estimated first-stage (\Cref{sec:results} states sufficient conditions) admits the same expansion,
\begin{equation}
\label{eq:first-order}
\sqrt n(\widehat\beta-\beta)
=\frac{\sqrt n\,\mathbb{E}_n[\Pi_i\varepsilon_i]}{H}+o_p(1),
\end{equation}
whose leading term does not involve the first-stage estimate and, conditional on the
instruments, is distributed exactly as \(N(0,\sigma_\varepsilon^2/H)\) at every sample size. Comparing first-stage estimators thus requires analyzing terms hidden in the remainder. We provide a feasible criterion for this comparison when the first stage is estimated by the LASSO \citep{Tibshirani-1996}, as in \citet{BCCH-2012}, a leading approach when the instrument set is high-dimensional or when the researcher wants to use a high-dimensional basis of technical instruments.
As will be described below, LASSO estimation of the first stage is characterized by a choice of dictionary, i.e., the choice of instrument basis, and a penalty parameter. The criterion developed in \Cref{sec:results} ranks these candidate dictionary--penalty pairs by analyzing second-order terms in a mean-squared-error expansion of the resulting second-stage estimator. Our analysis is closely related to that of \citet{DonaldNewey-2001}, who provide a similar criterion when the first stage is estimated by ordinary least squares (OLS) on a chosen instrument basis.
\subsection{Candidate First Stages and the IV Estimator}
We assume that the researcher is interested in selecting a first-stage specification \(c\) from a fixed number \(J\) of candidates. Each candidate \(c\in[J]\) consists of a dictionary of technical instruments and a penalty level \(\lambda_c>0\). For each observation, candidate \(c\)'s dictionary maps the observed instruments \(z_i\) into a vector of technical instruments \(z_{c,i}:=(z_{c,i1},\dots,z_{c,ip_c})'\in\SR^{p_c}\), whose entries are fixed functions of \(z_i\), such as powers, interactions, or indicators. The dictionary dimension \(p_c\) may exceed the sample size \(n\). Although \(\lambda_c\) is nonrandom, both the dictionary and the penalty level may depend on \(n\). We collect the resulting technical instruments in the matrix
\begin{equation*}
Z_c:=(z_{c,1},\dots,z_{c,n})'\in\SR^{n\times p_c},
\qquad
Z_{c,B}:=(Z_c)_{\cdot,B}
\quad\text{for }B\subseteq[p_c],
\end{equation*}
Thus, the \(i\)th row of \(Z_c\) is \(z_{c,i}'\), while
\(Z_{c,B}\in\SR^{n\times |B|}\) contains the columns indexed by \(B\).
For candidate \(c\) and an arbitrary response vector
\(w\coloneqq(w_1,\dots,w_n)\in\SR^n\), define the LASSO coefficient
estimate and fitted vector \citep{Tibshirani-1996} by
\begin{equation}
\widehat\pi_c(w)
\in\operatorname*{arg\,min}_{\pi\in\mathbb R^{p_c}}
\left\{
\frac12\mathbb{E}_n[(w_i-z_{c,i}'\pi)^2]
+\lambda_c\lVert\pi\rVert_1
\right\},
\qquad
\mu_c(w):=Z_c\widehat\pi_c(w).
\label{eq:lasso}
\end{equation}
When the columns of \(Z_c\) are linearly dependent, the coefficient
minimizer need not be unique, but the fitted vector \(\mu_c(w)\) remains
unique \citep{TibshiraniTaylor-2012,Tibshirani-2013}. Applying this
procedure to the endogenous variable \(x\), define
\(\widehat\pi_c:=\widehat\pi_c(x)\),
\(\widehat\Pi_c:=\mu_c(x)\), and
\(\widehat\Pi_{c,i}:=z_{c,i}'\widehat\pi_c\). The corresponding
second-stage IV estimator is
\begin{equation}
\label{eq:iv-estimator}
\widehat\beta_c
\coloneqq
\begin{cases}
\displaystyle
\frac{\mathbb{E}_n[\widehat\Pi_{c,i}y_i]}
{\mathbb{E}_n[\widehat\Pi_{c,i}x_i]},
&\mathbb{E}_n[\widehat\Pi_{c,i}^2]>0,\\[2mm]
0,
&\mathbb{E}_n[\widehat\Pi_{c,i}^2]=0.
\end{cases}
\end{equation}
A natural criterion for comparing the candidate estimators
\(\widehat\beta_c\) is their scaled mean squared error (MSE), \(\mathbb{E}[n(\widehat\beta_c-\beta)^2].\)
The exact MSE of an IV estimator, however, need not exist in finite samples.
To see this, consider the infeasible estimator that uses the optimal instrument
\(\Pi_i\) in place of \(\widehat\Pi_{c,i}\) in \eqref{eq:iv-estimator}. This
estimator satisfies
\begin{equation}
\label{eq:infeasible-decomp}
\widehat\beta-\beta
=\frac{\mathbb{E}_n[\Pi_i\varepsilon_i]}{\mathbb{E}_n[\Pi_ix_i]}.
\end{equation}
Under \Cref{assm:gaussian}, the numerator and denominator on the right-hand
side are jointly Gaussian. Because \(\Omega\) is nonsingular, the numerator
is not an exact linear function of the denominator, and their ratio does not have a
finite second moment. Thus, even the infeasible estimator based on the optimal
instrument has no exact MSE. More generally, when the exact MSE exists, it
may be neither tractable nor consistently estimable.
Given this restriction, we instead extend the higher-order approach of
\citet{DonaldNewey-2001} to the present LASSO setting. Under suitable
conditions, we show that the scaled squared error of each estimator \(\widehat\beta_c\)
admits the decomposition
\begin{equation}
\label{eq:decomp}
n (\widehat\beta_c - \beta)^2 = C_0 + Q_c + r_c.
\end{equation}
Here, \(C_0\) is common to all candidates and captures the component of
squared error corresponding to the leading term in \eqref{eq:first-order}.
By contrast, \(Q_c\) captures the second-order contribution of candidate
\(c\)'s first-stage procedure, while \(r_c\) is asymptotically negligible in
the sense that
\(
\max_{c\in[J]}|r_c|/\mathbb{E}[Q_c]\to_p 0.
\)
We define \(\mathbb{E}[C_0+Q_c]\) as the approximate mean squared error (AMSE) of
\(\sqrt n(\widehat\beta_c-\beta)\). For the infeasible estimator based on the
optimal instrument, the expectation of the candidate-specific component is
zero. Hence, \(\bar Q_c:=\mathbb{E}[Q_c]\) measures the excess AMSE of candidate
\(c\) relative to that benchmark. Because \(C_0\) is common across candidates, ranking candidates by AMSE is
equivalent to ranking them by \(\bar Q_c\). In \Cref{sec:results}, we derive a
closed-form asymptotic approximation to \(\bar Q_c\), construct a consistent
feasible counterpart, and use it to select among the first-order-equivalent
candidate estimators.
\subsection{The Effective Dimension}
\label{sec:effective-dimension}
A key element of the criterion in \Cref{sec:results} is a measure of the complexity, or ``degrees of freedom,'' of each first-stage estimate. When the first stage is OLS on a linearly independent set of instruments, complexity is simply the number of instruments used. For the LASSO, however, the number of selected instruments is not always well defined. When dictionary columns are collinear, the coefficient solution need not be unique, and different solutions may have different supports, even though the fitted values and the residual are unique.
Let \(u_c(w):=w-\mu_c(w)\) denote the LASSO residual and define the equicorrelation set by
\begin{equation}
E_c(w)
:=\left\{
j\in[p_c]:
\left|\mathbb{E}_n[z_{c,ij}u_{c,i}(w)]\right|=\lambda_c
\right\},
\label{eq:equicorrelation}
\end{equation}
This is the set of dictionary columns for which the LASSO Karush--Kuhn--Tucker (KKT) inequality binds. Because the residual is unique, \(E_c(w)\) does not depend on which coefficient solution is used, and it contains the support of every solution.
We define the effective dimension of the first-stage fit as
\begin{equation}
d_c:=\operatorname{rank}(Z_{c,E_c(x)}).
\label{eq:realized-rank}
\end{equation}
The quantity \(d_c\) is observable and invariant to the choice of coefficient solution. When the LASSO solution is unique, \(d_c\) agrees with the number of selected instruments almost surely under our sampling assumption.\footnote{A sufficient condition for uniqueness is that the dictionary columns are in general position, which holds with probability one when the dictionary matrix has a joint continuous distribution \citep{Tibshirani-2013}.} As shown by \citet{TibshiraniTaylor-2012}, the degrees of freedom of the LASSO fit are given by \(\mathbb{E}[d_c]\).
\section{Main Results}
\label{sec:results}
After stating the conditions required for our analysis, this section formally establishes the approximate mean-squared error decomposition in \eqref{eq:decomp} and provides a closed-form asymptotic approximation to the expectation of its candidate-specific term, which serves as an infeasible selection criterion (\Cref{thm:main}). We then show that this infeasible criterion, and thus the expectation of the candidate-specific term itself, can be consistently estimated and use the resulting feasible criterion to select the optimal first-stage implementation (\Cref{thm:feasible}). Throughout, a candidate \(c\in[J]\) is a dictionary--penalty pair as described in \Cref{sec:framework}, and the number of candidates \(J\) is treated as fixed.
\subsection{Conditions on the Candidate First Stages}
\label{sec:primitive}
Along with the Gaussian sampling and strong identification conditions in \Cref{assm:gaussian}, our analysis proceeds under two additional sets of high-level assumptions. The first ensures that each candidate LASSO estimation procedure described in \eqref{eq:lasso} consistently estimates the first stage and thus that each candidate IV estimator in \eqref{eq:iv-estimator} is first-order equivalent. This set of high-level assumptions is standard in the literature on \(\ell_1\)-penalized estimation \citep{BickelRitovTsybakov-2009,VandeGeerBuhlmann-2009,BCCH-2012}. The second is a partial beta-min condition used to establish concentration of the LASSO effective dimension. It is similar in spirit to the growing-basis condition in \citet{DonaldNewey-2001} but does not require exact support recovery by the LASSO.
To present the assumptions, we begin by introducing the notion of an approximately sparse representation. We say that the first stage \(\Pi_i\) admits an approximately sparse representation in the candidate dictionary \(Z_c\) if there is a nonrandom coefficient vector \(\pi_c^0\in\SR^{p_c}\) such that
\begin{equation}
\Pi_i=z_{c,i}'\pi_c^0+\xi_{c,i},
\qquad
T_c:=\operatorname{supp}(\pi_c^0),
\qquad
s_c:=|T_c|,
\label{eq:sparse-approximation}
\end{equation}
where \(\operatorname{supp}(b):=\{j:b_j\ne0\}\) and \(s_c\ge1\). The rate restrictions on \(s_c\) and the approximation error \(\xi_{c,i}\) are stated below. As in \citet{BCCH-2012}, approximate sparsity can be interpreted either as an assumption that only a few of the instruments in the dictionary are important for explaining variation in the endogenous variable, or as an assumption that the conditional mean \(\Pi_i\) can be accurately approximated using only a small number of dictionary terms.
For a vector \(\delta\) and an index set \(T\), write \(\delta_T:=(\delta_j)_{j\in T}\)
and \(\lVert\delta\rVert_0:=|\operatorname{supp}(\delta)|\). Fix a constant
\(c_\lambda>1\), common to all candidates, and set
\[
L_\lambda:=\frac{c_\lambda+1}{c_\lambda-1}.
\]
Following
\citet{BCCH-2012}, define the restricted eigenvalue \citep{BickelRitovTsybakov-2009,VandeGeerBuhlmann-2009}
\begin{equation}
\kappa_c
:=\min_{\substack{\delta\ne0\\
\lVert\delta_{T_c^c}\rVert_1
\le L_\lambda\lVert\delta_{T_c}\rVert_1}}
\frac{\sqrt{s_c\mathbb{E}_n[(z_{c,i}'\delta)^2]}}
{\lVert\delta_{T_c}\rVert_1},
\label{eq:restricted-eigenvalue}
\end{equation}
and the upper and lower \(m\)-sparse eigenvalues
\begin{equation}
\phi_{\max,c}(m)
:=\max_{\substack{\delta\ne0\\
\lVert\delta\rVert_0\le m}}
\frac{\mathbb{E}_n[(z_{c,i}'\delta)^2]}{\lVert\delta\rVert_2^2},
\label{eq:sparse-eigenvalue}
\end{equation}
\begin{equation}
\phi_{\min,c}(m)
:=\min_{\substack{\delta\ne0\\
1\le\lVert\delta\rVert_0\le m}}
\frac{\mathbb{E}_n[(z_{c,i}'\delta)^2]}{\lVert\delta\rVert_2^2}.
\label{eq:lower-sparse-eigenvalue}
\end{equation}
The restricted eigenvalue bounds the design from below only over a cone of
vectors whose weight is concentrated on the support of the coefficient vector \(\pi_c^0\). Bounds on sparse eigenvalues require that small collections of dictionary columns be well conditioned. Conditions stated in terms of these quantities are weaker than restricting the full dictionary as larger collections of dictionary columns may still be collinear. In particular, bounds on restricted and sparse eigenvalues do not rule out cases where \(p_c \gg n\).
\Needspace{10\baselineskip}
\begin{assumption}[First-Stage Convergence]
\label{assm:primitive}
There are constants \(0\le C_a<\infty\), \(\alpha>0\),
\(0<\kappa_0\le1\),
\(0<\underline\phi\le1\le\overline\phi<\infty\), and
\(M>4L_\lambda^2\overline\phi/\kappa_0^2\), common to all candidates,
such that the following hold for each \(c \in [J]\):
\begin{enumerate}[(i)]
\item The dictionary columns are normalized such that \(\mathbb{E}_n[z_{c,ij}^2]=1\) for each \(j\in[p_c]\).
\item The approximation error satisfies
\[
\mathbb{E}_n[\xi_{c,i}^2]^{1/2}
\le C_a\sqrt{\frac{s_c}{n}}.
\]
\item The penalty satisfies
\[
\lambda_c
\ge c_\lambda\sqrt{2}\sigma_v
\sqrt{
\frac{\log(2p_c)+\alpha\log\log(p_c\vee n)}{n}
}.
\]
\item The design satisfies
\begin{equation}
\label{eq:design-conditions}
\kappa_c\ge\kappa_0,
\qquad
\phi_{\max,c}(\lceil Ms_c\rceil\wedge p_c) \le\overline\phi,
\qquad
\phi_{\min,c}(\lceil(M+1)s_c\rceil\wedge p_c) \ge\underline\phi.
\end{equation}
\item The sparsity index, \(s_c\), and penalty \(\lambda_c\) satisfy
\begin{equation}
\frac{s_c}{\sqrt n}\to 0,
\qquad
\sqrt{s_c}\lambda_c\to 0.
\label{eq:growth-rates}
\end{equation}
\end{enumerate}
\end{assumption}
Versions of each condition in \Cref{assm:primitive} appear in
\citet{BCCH-2012}, who also use them to establish convergence of first-stage
LASSO estimates. \Cref{assm:primitive}(i) is a normalization and can always be
satisfied by rescaling the columns of the dictionary.
\Cref{assm:primitive}(ii) requires that the approximately sparse
representation in \eqref{eq:sparse-approximation} be accurate in the sense
that the approximation error is \(O(\sqrt{s_c/n})\) in prediction norm.
The lower bound in \Cref{assm:primitive}(iii) ensures that the penalty dominates the empirical score with high probability, yielding the usual LASSO prediction-error bounds. Conditions of this form motivate the penalty choices in \citet{BickelRitovTsybakov-2009}, \citet{VandeGeerBuhlmann-2009}, \citet{BCCH-2012}, and \citet{ChetverikovSorensen-2025}. \Cref{assm:primitive}(iv) imposes the
restricted and sparse eigenvalue conditions discussed above. Finally,
part~(v) imposes two rate restrictions. Together with part~(ii),
\(\sqrt{s_c}\lambda_c\to0\) yields first-stage prediction consistency, while
\(s_c/\sqrt n\to0\) controls terms that arise in the higher-order expansion.
Combined with \Cref{assm:gaussian}, these conditions ensure that \eqref{eq:first-order} holds for every candidate, so the resulting IV estimators are first-order equivalent. Beyond guaranteeing suitable first-stage convergence of the candidate first-stage estimators, these conditions do little to restrict the candidate list. Penalties may vary between the bounds in parts~(iii) and~(v), and dictionaries may differ in both content and dimension. The criterion developed in \Cref{sec:leading-result} ranks these first-order-equivalent candidates.
The feasible version of the criterion, developed in \Cref{sec:feasible}, uses the observed effective dimension \(d_c\) in place of its unknown moments. The following condition justifies this substitution.
\Needspace{12\baselineskip}
\begin{assumption}[Partial Beta-Min]
\label{assm:strong-core}
For each candidate, let
\[
\underline b_{n,c}
:=\frac{4}{\sqrt{\underline\phi}}
\left\{
\frac{(1+c_\lambda^{-1})\lambda_c\sqrt{s_c}}{\kappa_0}
+C_a\sqrt{\frac{s_c}{n}}
\right\},
\]
where \(\kappa_0\) and \(\underline\phi\) are as in
Assumption~\ref{assm:primitive}(iv). Define the set of strong coordinates by
\(A_c:=\{j\in T_c:|\pi_{c,j}^0|\ge\underline b_{n,c}\}\) with
\(k_c:=|A_c|\). There is a fixed constant \(C_s\ge1\) such that,
uniformly over \(c\in[J]\),
\begin{equation}
s_c\le C_sk_c,
\qquad
k_c\to\infty,
\qquad
\frac{\log(p_c/k_c)}{k_c}
\to 0.
\label{eq:strong-core-growth}
\end{equation}
\end{assumption}
Assumption~\ref{assm:strong-core} imposes a beta-min condition on a subset of the non-zero coefficients. In the literature on consistent model selection via the LASSO, beta-min conditions are imposed on the entire coefficient support \(T_c\) and, together with conditions on the design, imply recovery of the true support \citep{ZhaoYu-2006,Wainwright-2009}. The condition here plays a different role. The strong coordinates are selected with probability approaching one. Together with the concentration bound for \(d_c\), this yields ratio consistency of the observed effective dimension for its expectation; see \Cref{prop:ratio}. It does not impose a beta-min lower bound on the coefficients in \(T_c\setminus A_c\), nor does it require these coefficients to be selected with probability approaching one. The requirement that \(k_c\to\infty\) parallels the condition in \citet{DonaldNewey-2001} that the number of instruments used to estimate the first stage grows with the sample size, and plays a similar role in the technical analysis. Because \(\underline b_{n,c}\to0\), Assumption~\ref{assm:strong-core} allows the coefficients in the strong core to shrink with \(n\), even though the size of that core must diverge.
\subsection{The AMSE Criterion}
\label{sec:leading-result}
This subsection formally establishes the decomposition in \eqref{eq:decomp} and presents the infeasible criterion, a closed-form asymptotic approximation to the expectation of the candidate-specific term. Exact expressions for the terms in \eqref{eq:decomp} are given in the appendix but are not needed for the characterization below.
Two population quantities enter the infeasible criterion. The first is a measure of first-stage strength,
\begin{equation}
h_c:=\overline{\mathbb{E}}[\widehat\Pi_{c,i}\Pi_i],
\label{eq:h-def}
\end{equation}
the population cross-moment between the fitted values and the optimal instrument. This is the population counterpart of the denominator of the IV estimator in \eqref{eq:iv-estimator}. Under the conditions of \Cref{sec:primitive}, \(h_c\) converges to the second moment of the optimal instrument, \(H\), and so is bounded away from zero in large samples.
The second quantity measures how well the fitted instrument approximates the optimal instrument. Because the IV estimator is unchanged when its instrument is multiplied by a nonzero constant, a fitted instrument proportional to the optimal instrument should register no error. We therefore measure the first-stage approximation error against the best rescaling of the optimal instrument,
\begin{equation}
\mathcal A_c
:=\min_{a\in\mathbb R}
\overline{\mathbb{E}}\!\left[
(\widehat\Pi_{c,i}-a\Pi_i)^2
\right]
=\overline{\mathbb{E}}[\widehat\Pi_{c,i}^2]-\frac{h_c^2}{H}.
\label{eq:amse-approx}
\end{equation}
The minimized value is the mean squared residual from a population regression of the fitted instrument on the optimal instrument. It reflects both the error from approximating the optimal instrument with the dictionary and the shrinkage and noise introduced by LASSO estimation. The ratio \(\mathcal A_c/h_c^2\) appearing in the infeasible criterion below is invariant to rescaling of the fitted instrument.
The closed-form approximation to the expectation of the candidate-specific term combines these two quantities with certain error moments and moments of the effective dimension of the LASSO fit defined in \Cref{sec:effective-dimension},
\begin{equation}
S_c
:=\frac{\sigma_\varepsilon^2\mathcal A_c}{h_c^2}
+\frac{\sigma_{\varepsilon v}^2\,\mathbb{E}[d_c^2+d_c]}
{nh_c^2}.
\label{eq:closed-form-criterion}
\end{equation}
We refer to \(S_c\) as the infeasible criterion. Each of its components is a population quantity, so the infeasible criterion cannot be computed directly from the data. \Cref{sec:feasible} constructs a consistent feasible counterpart used to select a candidate.
\begin{theorem}[Infeasible AMSE Criterion]
\label{thm:main}
Suppose Assumptions~\ref{assm:gaussian}, \ref{assm:primitive}, and
\ref{assm:strong-core} hold, with \(J<\infty\) fixed. Then, for all
sufficiently large \(n\), there exist random variables \(C_0\),
\(\{Q_c\}_{c\in[J]}\), and \(\{r_c\}_{c \in [J]}\) such that
\begin{enumerate}[(i)]
\item For every candidate,
\begin{equation}
n(\widehat\beta_c-\beta)^2
=C_0+Q_c+r_c,
\label{eq:squared-decomposition}
\end{equation}
where \(C_0\) does not depend on the candidate,
\(\mathbb{E}[C_0]=\sigma_\varepsilon^2/H+O(n^{-1})\), and the remainder
satisfies \(\max_{c\in[J]}|r_c|/\mathbb{E}[Q_c]\longrightarrow_p0\).
\item Uniformly over candidates, \(\mathbb{E}[Q_c]>0\)
and
\begin{equation}
\mathbb{E}[Q_c]=S_c\left(1+o(1)\right).
\label{eq:leading-characterization}
\end{equation}
\end{enumerate}
\end{theorem}
\Cref{thm:main}(i) establishes the validity of the decomposition in \eqref{eq:decomp}, while \Cref{thm:main}(ii) shows that \(S_c\) approximates the candidate-specific AMSE component \(\bar Q_c=\mathbb{E}[Q_c]\) with vanishing relative error. Together, these results justify ranking candidates by \(S_c\): minimizing \(S_c\) is asymptotically equivalent to minimizing \(\bar Q_c\), and hence AMSE, over the candidate list.
The infeasible criterion consists of two terms. The first is proportional to \(\mathcal A_c\), which measures how well the candidate LASSO approximates the optimal instrument.
This term vanishes only for a candidate estimator that is always proportional to the true first-stage \(\Pi\). The second term in \(S_c\) is proportional to \(\mathbb{E}[d_c^2 + d_c]\), a measure of the complexity of the LASSO fitted model.
A richer dictionary or a lighter penalty may reduce \(\mathcal A_c\) while increasing the effective dimension. To examine this tradeoff, write \(\rho=\operatorname{Corr}(\varepsilon_i,v_i)\), so that \(\sigma_{\varepsilon v}=\rho\sigma_\varepsilon\sigma_v\), and rewrite \eqref{eq:closed-form-criterion} as
\begin{equation}
S_c
=\frac{\sigma_\varepsilon^2}{h_c^2}
\left\{
\mathcal A_c
+\rho^2\sigma_v^2\,
\frac{\mathbb{E}[d_c^2+d_c]}{n}
\right\}.
\label{eq:criterion-exchange-rate}
\end{equation}
In this form the second term is weighted by \(\rho^2\sigma_v^2\), which is proportional to the squared covariance between the structural and first-stage errors. When the regressor is close to exogenous, the bias from even a complex fitted model is small and candidates are ranked mainly by the approximation term.\footnote{Exact exogeneity is ruled out, since Assumption~\ref{assm:gaussian}(ii) requires \(\sigma_{\varepsilon v}\ne0\). When \(\sigma_{\varepsilon v}=0\) the second term vanishes altogether and the remaining terms require a different expansion.} When the regressor is highly endogenous, the same complexity produces a larger bias and the criterion favors heavier regularization. This is the LASSO analogue of the many-instrument bias examined in \citet{Bekker-1994} and \citet{DonaldNewey-2001}.
Cross-validation on the first stage can be thought of as selecting the candidate that minimizes \(\mathcal A_c\) while ignoring the many-instrument bias.\footnote{Cross-validation targets a first stage estimator \(\widehat\Pi\) that minimizes \(\bar\mathbb{E}[(\widehat\Pi_i - \Pi_i)^2]\), which is an upper bound on \(\calA_c\).}
When endogeneity is weak, the complexity term receives little weight, so the cross-validation criterion and the AMSE criterion may rank candidates similarly. As endogeneity strengthens, the AMSE criterion places greater weight on first-stage complexity, whereas cross-validation does not. In the simulations of \Cref{sec:application-simulation}, the effective dimension selected by the proposed rule falls sharply as the error correlation rises, while the cross-validation choice is unchanged because its objective uses only the first stage. The plug-in penalty of \citet{BCCH-2012} is chosen to dominate the first-stage noise with high probability so that \Cref{assm:primitive}(iii) is satisfied, but its construction does not take into account the measure of fit \(\calA_c\) nor the level of endogeneity.
\subsection{Feasible Candidate Selection}
\label{sec:feasible}
Although \(S_c\) is infeasible, we construct an observable criterion that estimates it up to a candidate-independent centering term and therefore preserves its ranking asymptotically. The only additional inputs not obtained from the candidate fits are the structural error moments \(\sigma_\varepsilon^2\) and \(\sigma_{\varepsilon v}\). We estimate these moments using residuals from a pilot candidate \(\check c\in[J]\).
For each candidate, the observed IV denominator
\begin{equation}
\label{eq:sample-denominator}
\widehat h_c:=\mathbb{E}_n[\widehat\Pi_{c,i}x_i]
\end{equation}
will serve as an estimator for first-stage cross-moment \(h_c\). The KKT conditions for the LASSO optimization problem in \eqref{eq:lasso} yield \(\widehat h_c =\mathbb{E}_n[\widehat\Pi_{c,i}^2]+\lambda_c\lVert\widehat\pi_c\rVert_1\ge0\), so \(\widehat h_c=0\) if and only if \(\widehat\pi_c = 0\). The structural error moments are estimated with residuals from the pilot candidate \(\check c\). Setting \(\check\beta:=\widehat\beta_{\check c}\), let
\begin{equation}
\label{eq:nuisance-estimators}
\widehat\sigma_\varepsilon^2
:=\mathbb{E}_n[(y_i-\check\beta x_i)^2]
\qquad \text{and} \qquad
\widehat\sigma_{\varepsilon v}
:=\mathbb{E}_n[(y_i-\check\beta x_i)x_i].
\end{equation}
We put together these estimates to construct the feasible criterion
\begin{equation}
\label{eq:feasible-score}
\widehat S_c
:=\frac{
\widehat\sigma_\varepsilon^2
\mathbb{E}_n[\widehat\Pi_{c,i}^2]
+n^{-1}
\widehat\sigma_{\varepsilon v}^{\,2}
(d_c^2+d_c)
}{
\widehat h_c^2
},
\end{equation}
with the convention \(\widehat S_c:=+\infty\) when \(\widehat h_c=0\). The selected
candidate minimizes the feasible criterion,
\begin{equation}
\widehat c
:=
\min\left(
\operatorname*{arg\,min}_{c\in[J]}
\widehat S_c
\right),
\label{eq:selected-candidate}
\end{equation}
with ties broken by the smallest index.
The feasible criterion \(\widehat S_c\) replaces the moment \(\mathbb{E}[d_c^2+d_c]\) in the infeasible \(S_c\) with the realized value \(d_c^2+d_c\). This replacement is justified by
the following proposition.
\begin{prop}[Ratio Consistency of the Effective Dimension]
\label{prop:ratio}
Suppose Assumptions~\ref{assm:gaussian}, \ref{assm:primitive}, and
\ref{assm:strong-core} hold, with \(J<\infty\) fixed. Uniformly over
candidates,
\begin{align}
\frac{d_c}{\mathbb{E}[d_c]}
&\longrightarrow_p1,
\label{eq:rank-ratio-main}\\
\frac{d_c^2+d_c}{\mathbb{E}[d_c^2+d_c]}
&\longrightarrow_p1.
\label{eq:rank-factor-ratio-main}
\end{align}
\end{prop}
The result in \Cref{prop:ratio} relies on the condition from \Cref{assm:strong-core} that the number of nonzero coefficients in \(\pi_c^0\) diverges. Under this condition, second-order Stein identities from \citet{BellecZhang-2021} show that the variance of the effective dimension grows more slowly than its squared mean, yielding \eqref{eq:rank-ratio-main} and \eqref{eq:rank-factor-ratio-main}. \Cref{prop:ratio} does not require the
LASSO estimator to perfectly recover the support of the underlying
coefficient vector \(\pi_c^0\) in \eqref{eq:sparse-approximation}.
The numerator of \(\widehat S_c\) contains the fitted second moment
\(\mathbb{E}_n[\widehat\Pi_{c,i}^2]\) rather than an estimate of the approximation
error \(\mathcal A_c\). By the decomposition in \eqref{eq:amse-approx}
applied to sample moments, this substitution adds the constant
\(\widehat\sigma_\varepsilon^2/H\) relative to a criterion that directly uses an empirical analogue of \(\mathcal A_c\). This added constant does not depend on the candidate and thus does not affect the minimizer of the feasible criterion.\footnote{The added constant estimates
\(\sigma_\varepsilon^2/H\), which \Cref{thm:main}(i) shows is the leading
expectation of the common term in \eqref{eq:decomp}. The feasible
criterion therefore estimates the AMSE of each candidate rather than the
candidate-specific term alone.} The feasible criterion also uses the denominator \(\widehat h_c^2\) instead of the true \(h_c^2\). To first-order, the error from using this estimate of the denominator is \(-2\widehat\sigma_\varepsilon^2\mathbb{E}_n[\Pi_iv_i]/H^2\), which also does not depend on the candidate. Let \(B_n\) collect these two candidate-independent terms,
\begin{equation}
\label{eq:common-term}
B_n \coloneqq\widehat\sigma_\varepsilon^2
\left(\frac1H-\frac{2\mathbb{E}_n[\Pi_iv_i]}{H^2}\right).
\end{equation}
\Cref{thm:feasible} below shows that, uniformly over candidates,
\begin{equation}
\widehat S_c
=B_n+\bar Q_c\left(1+o_p(1)\right).
\end{equation}
Minimizing \(\widehat S_c\) is then asymptotically equivalent to minimizing
\(\bar Q_c\). Note that \(B_n\) itself need not be observed and may change with the sample size.
\begin{theorem}[Feasible AMSE Selection]
\label{thm:feasible}
Suppose Assumptions~\ref{assm:gaussian}, \ref{assm:primitive}, and
\ref{assm:strong-core} hold, with \(J<\infty\) fixed.
\begin{enumerate}[(i)]
\item Up to the common centering \(B_n\), the feasible criterion
consistently estimates the infeasible criterion,
\begin{equation}
\max_{c\in[J]}
\frac{|\widehat S_c-B_n-S_c|}{S_c}
\longrightarrow_p0,
\label{eq:feasible-consistency}
\end{equation}
and thus the expectation of the candidate-specific term,
\begin{equation}
\max_{c\in[J]}
\frac{|\widehat S_c-B_n-\bar Q_c|}{\bar Q_c}
\longrightarrow_p0.
\label{eq:feasible-expansion}
\end{equation}
\item The selected candidate \(\widehat c\) satisfies
\begin{equation}
\frac{\bar Q_{\widehat c}}
{\min_{c\in[J]}\bar Q_c}
\longrightarrow_p1.
\label{eq:oracle-equivalence}
\end{equation}
\end{enumerate}
\end{theorem}
By \Cref{thm:feasible}(ii), the selected candidate attains the smallest candidate-specific AMSE component in the candidate list up to vanishing relative error.
The conditional Gaussianity and homoskedasticity restrictions in \Cref{assm:gaussian} are used to derive the AMSE ranking. If the errors are non-Gaussian or heteroskedastic, \Cref{thm:feasible} no longer guarantees that the selected candidate is AMSE-optimal. If, under suitable alternative conditions, all candidates continue to admit the common first-order expansion in \eqref{eq:first-order} uniformly over the fixed candidate list, then the selected estimator retains the corresponding asymptotic distribution and first-order inference remains valid.
\section{Simulation Study}
\label{sec:empirical-application}
In this section, we examine the finite-sample performance of the proposed
criterion and compare to first-stage selection via cross-validation and the plug-in penalty of \citet{BCCH-2012}. The simulations are calibrated to the data of \citet{gilchrist-glassberg-2016}, who study the effect of social spillovers on movie consumption. The authors originally instrument for ticket sales on a given opening weekend day with 48 linearly independent instruments that measure national weather conditions. We hold these initial instrumnts as fixed and consider two main error regimes. We first consider Gaussian
errors, which matches the setting of \Cref{sec:results}, and then repeat the exercise with heavier-tailed Laplace errors to asses the sensitivity of our proposed criterion based selection to non-Gaussian errors. \Cref{sec:sim-diagnostics} provides implementation details and complete
results.
\subsection{Calibrated Gaussian Designs}
\label{sec:application-simulation}
We hold the observed weather instruments fixed and generate the outcome
and endogenous variable from the IV model described in \eqref{eq:model} with \(\beta=1\). The true first stage \(\Pi\) is proportional to the fitted values from a least-squares regression of opening-weekend sales on the original 48 linearly independent instruments. We rescale these fitted values to vary identification strength, targeting expected first-stage \(F\) statistics of \(3.8\), \(6.6\), \(12\), or \(25\) at the original sample size of \(n=1{,}671\). We also evaluate the performance of our proposed criterion on a smaller, fixed subsample of \(n=800\). Holding the instruments and first-stage direction fixed lets us examine how each rule's performance changes with identification strength and endogeneity in an empirically plausible setting.
The errors are generated homoskedastic Gaussian, with
\(\varepsilon_i=\rho v_i+\sqrt{1-\rho^2}\,e_i\) for \(v_i, e_i \overset{\text{iid}}{\sim} N(0,1)\). The parameter \(\rho\) controls the degree of endogeneity. We consider a value calibrated to the data, \(\rho\approx-0.29\), along with alternate values \(\rho=-0.40,-0.50,\ldots,-0.90\) to examine how the performance of various criterion changes with the level of endogeneity. Each of the 56 designs is simulated 1,000 times. We report the average of \(n(\widehat\beta-\beta)^2\), which we refer to as either the risk or the mean squared error.
Each candidate first stage implementation uses one of three nested dictionaries: either 34 temperature instruments (Temperature), the original 48 weather instruments (Original), or 524 instruments including temperature--weather interactions (Expanded). For each dictionary, we consider 13 multiples of the plug-in penalty of \citet{BCCH-2012} (``BCCH''), giving 39 candidates in total. This plug-in penalty scale uses the known first-stage error variance in place of an estimate. The proposed criterion selects among these candidate dictionary-penalty pairs by minimizing \(\widehat S_c\) in \eqref{eq:feasible-score}, with estimated error moments coming from a pilot estimator of \(\beta\). Cross-validation chooses from the same candidates using ten-fold first-stage prediction error, keeping observations from the same opening weekend in the same fold. BCCH uses its plug-in penalty on the Original dictionary. When it selects no instruments, we use the pilot estimate instead.
\begin{figure}[!htp]
\centering
\includegraphics[width=\textwidth]{figures/gs_calibrated_summary}
\caption{Relative risk and effective dimension under Gaussian errors.}
\label{fig:gs-calibrated-summary}
\par\medskip\parbox{\linewidth}{\footnotesize\emph{Notes:}\ Results use \(n=1{,}671\) and 1,000 replications. Panels A
and C report risk relative to the lowest risk among the three rules;
Panels B and D report mean effective dimension. Panels A--B use
\(F=6.6\) and Panels C--D use \(F=12\). The horizontal axis is the
absolute error correlation \(|\rho|\).}
\end{figure}
\Cref{fig:gs-calibrated-summary} shows how the rules respond to
endogeneity at the two intermediate levels of identification strength, \(F = 6.6\) and \(F = 12\). Cross-validation performs well when the regressor is only
mildly endogenous, but its relative risk increases as endogeneity rises. In contrast, the BCCH rules seems to impose a high degree of regularization and, as such, performs poorly when endogeneity is low but similarly to the criterion based rule at the highest levels of endogeneity. Both rules do not account for the level of endogeneity and so their mean effective dimension remains the same across all values of \(|\rho|\).
By comparasion, the proposed criterion adjusts the first-stage complexity based on the level of endogeneity, consistent with the increasing complexity cost seen in \eqref{eq:criterion-exchange-rate}. At both \(F = 6.6\) and \(F = 12\), the mean criterion effective dimension falls by nearly four-fold as \(|\rho|\) ranges from \(0.29\) to \(0.90\). As a result of this adjustment, the criterion is able to achieve nearly optimal performance across the entire range of regimes considered in \Cref{fig:gs-calibrated-summary}. The improvements in relative risk from using this criterion can be sizeable. At low levels of endogeneity, the criterion's risk is about 40\% lower than that of the BCCH plug-in rule while at the highest level of endogeneity, the criterion's risk is nearly 35\% lower than that of cross-validation.
\begin{table}[!htp]
\centering
\begin{threeparttable}
\caption{Relative risk, Gaussian errors.}
\label{tab:gs-compact}
\setlength{\tabcolsep}{9pt}
\begin{tabular}{ll S[table-format=2.3] S[table-format=2.3] S[table-format=2.3] S[table-format=2.3] S[table-format=2.3]}
\toprule
Design & \shortstack[l]{Summary\\Measure} & \multicolumn{1}{c}{\shortstack{Feasible\\Criterion}} & \multicolumn{1}{c}{\shortstack{Known\\Moments}} & \multicolumn{1}{c}{\shortstack{Infeasible\\Criterion}} & \multicolumn{1}{c}{\shortstack{Cross-\\validation}} & \multicolumn{1}{c}{BCCH} \\
\midrule
\multicolumn{7}{l}{\textit{Panel A: All $F$ and $\rho$, by sample size}} \\
\addlinespace
$n=800$ & Average & 1.045 & {--} & {--} & 1.135 & 1.189 \\
& Maximum & 1.266 & {--} & {--} & 1.665 & 1.833 \\
\addlinespace
$n=1{,}671$ & Average & 1.026 & {--} & {--} & 1.106 & 1.194 \\
& Maximum & 1.128 & {--} & {--} & 1.551 & 2.045 \\
\midrule
\multicolumn{7}{l}{\textit{Panel B: $n=1{,}671$, by first-stage strength}} \\
\addlinespace
$F=3.8$ & Average & 1.040 & {--} & {--} & 1.174 & 1.297 \\
& Maximum & 1.128 & {--} & {--} & 1.551 & 2.045 \\
\addlinespace
$F=6.6$ & Average & 1.017 & {--} & {--} & 1.171 & 1.191 \\
& Maximum & 1.050 & {--} & {--} & 1.501 & 1.684 \\
\addlinespace
$F=12$ & Average & 1.016 & {--} & {--} & 1.055 & 1.142 \\
& Maximum & 1.055 & {--} & {--} & 1.214 & 1.445 \\
\addlinespace
$F=25$ & Average & 1.030 & {--} & {--} & 1.023 & 1.146 \\
& Maximum & 1.072 & {--} & {--} & 1.117 & 1.379 \\
\midrule
\multicolumn{7}{l}{\textit{Panel C: All cells, relative to the oracle}} \\
\addlinespace
Full grid & Average & 1.164 & 1.116 & 1.000 & 1.268 & 1.326 \\
& Maximum & 1.440 & 1.289 & 1.000 & 1.811 & 2.010 \\
\bottomrule
\end{tabular}
\begin{tabnotes}
Risk is the mean of \(n(\widehat\beta-\beta)^2\) over 1,000 replications.
Panels A and B present risk relative to the best of the three feasible rules in each design. Panel A averages or maximizes over all 28 designs at each \(n\). Panel B averages over the seven values of \(\rho\) at each \(F\) at \(n = 1671\). Panel C reports risk relative to the candidate minimizing the infeasible \(S_c\). Known moments evaluates the feasible criterion at population error moments. Details of these infeasible benchmarks are given in \Cref{sec:sim-diagnostics}.
\end{tabnotes}
\end{threeparttable}
\end{table}
\Cref{tab:gs-compact} summarizes performance across all designs. The ``Known Moments'' column uses the proposed criterion with the error moments set to their population values, \(\sigma_\varepsilon^2=1\) and \(\sigma_{\varepsilon v}=\rho\). The ``Infeasible Criterion'' column selects the candidate minimizing the infeasible criterion \(S_c\), with its population moments computed from 1,000 independent simulation replications. Compared to the Cross-Validation and BCCH selection procedures, the feasible criterion has the smallest average and maximum relative risk at both sample sizes, over essentially all identification strength regimes, and over the entire grid of considered DGPs. At the original sample size, its largest relative risk is about 13\% above that of the best rule, compared with 55\% for cross-validation and 105\% for BCCH. As identification strengthens, the criterion and cross-validation perform more similarly. At \(n=1{,}671\) and \(F=25\), cross-validation has slightly lower average relative risk. This is consistent with first-stage estimation error having less effect on the structural estimate as identification strengthens. Intuitively, the behavior of the structural parameter estimate is driven more by the first-stage ``signal'' than first-stage estimation noise under stronger indentification.
Panel C also shows the effect of estimating the error moments used in
the criterion. Using their population values reduces average risk
relative to the infeasible criterion from 1.164 to 1.116, removing roughly 30\% of
the average excess risk and suggesting that the cost of estimating the error-variances is somewhat substantial at the \citet{gilchrist-glassberg-2016} sample size. Even with known error moments, the feasible
criterion still uses estimated first-stage quantities to rank
candidates, and its risk remains above that of the infeasible criterion in this
comparison. The
\subsection{Robustness to Non-Gaussian Errors}
\label{sec:application-laplace}
Since the theory in \Cref{sec:results} is derived under Gaussian errors, a reasonable question might be whether the proposed criterion is still useful in models with non-Gaussian errors. In other words, is the characterization of IV-LASSO AMSE in the Gaussian model informative about the AMSE of IV-LASSO in more general models. To assess this, we repeat the simulations of \Cref{sec:application-simulation} with variance-one Laplace errors. We retain the same first stages and candidate list and rerun the simulation experiment by drawing first-stage and structural errors from the heavier-tailed Laplace distribution.
\begin{figure}[!htp]
\centering
\includegraphics[width=\textwidth]{figures/gs_laplace_summary}
\caption{Relative risk and effective dimension under Laplace errors.}
\label{fig:gs-laplace-summary}
\par\medskip\parbox{\linewidth}{\footnotesize\emph{Notes:}\ Results use \(n=1{,}671\) and 1,000 replications. Panels A
and C report risk relative to the lowest risk among the three rules;
Panels B and D report mean effective dimension. Panels A--B use
\(F=6.6\) and Panels C--D use \(F=12\). The horizontal axis is the
absolute error correlation \(|\rho|\).}
\end{figure}
\Cref{fig:gs-laplace-summary} shows a similar response to endogeneity
as under Gaussian errors. At \(n=1{,}671\) and \(F=6.6\),
cross-validation a slightly lower risk than the proposed criterion at the three smallest values of \(|\rho|\), while the proposed criterion has the lowest risk from \(|\rho|=0.60\) through \(0.90\). At the highest level of endogeneity, the criterion's risk is about 28\% below that of cross-validation. The degrees of freedom of the criterion selected first-stage falls from 35.6 to 11.7, closely matching the decline under Gaussian errors. The similarity in selected dimensions suggests that use of the non-Gaussian errors has little effect on how heavily first-stage complexity is punished.
\begin{table}[!htp]
\centering
\begin{threeparttable}
\caption{Relative risk, Laplace errors.}
\label{tab:gs-laplace-compact}
\setlength{\tabcolsep}{9pt}
\begin{tabular}{ll S[table-format=2.3] S[table-format=2.3] S[table-format=2.3] S[table-format=2.3] S[table-format=2.3]}
\toprule
Design & \shortstack[l]{Summary\\Measure} & \multicolumn{1}{c}{\shortstack{Feasible\\Criterion}} & \multicolumn{1}{c}{\shortstack{Known\\Moments}} & \multicolumn{1}{c}{\shortstack{Infeasible\\Criterion}} & \multicolumn{1}{c}{\shortstack{Cross-\\validation}} & \multicolumn{1}{c}{BCCH} \\
\midrule
\multicolumn{7}{l}{\textit{Panel A: All $F$ and $\rho$, by sample size}} \\
\addlinespace
$n=800$ & Average & 1.044 & {--} & {--} & 1.130 & 1.205 \\
& Maximum & 1.261 & {--} & {--} & 1.631 & 2.010 \\
\addlinespace
$n=1{,}671$ & Average & 1.029 & {--} & {--} & 1.111 & 1.179 \\
& Maximum & 1.134 & {--} & {--} & 1.627 & 1.694 \\
\midrule
\multicolumn{7}{l}{\textit{Panel B: $n=1{,}671$, by first-stage strength}} \\
\addlinespace
$F=3.8$ & Average & 1.061 & {--} & {--} & 1.224 & 1.160 \\
& Maximum & 1.134 & {--} & {--} & 1.627 & 1.683 \\
\addlinespace
$F=6.6$ & Average & 1.020 & {--} & {--} & 1.132 & 1.226 \\
& Maximum & 1.058 & {--} & {--} & 1.385 & 1.694 \\
\addlinespace
$F=12$ & Average & 1.006 & {--} & {--} & 1.078 & 1.140 \\
& Maximum & 1.030 & {--} & {--} & 1.245 & 1.423 \\
\addlinespace
$F=25$ & Average & 1.028 & {--} & {--} & 1.011 & 1.191 \\
& Maximum & 1.049 & {--} & {--} & 1.079 & 1.435 \\
\midrule
\multicolumn{7}{l}{\textit{Panel C: All cells, relative to the oracle}} \\
\addlinespace
Full grid & Average & 1.162 & 1.115 & 1.000 & 1.263 & 1.325 \\
& Maximum & 1.432 & 1.288 & 1.000 & 1.784 & 1.860 \\
\bottomrule
\end{tabular}
\begin{tabnotes}
Results are computed as in \Cref{tab:gs-compact}, using 1,000 replications under Laplace errors. Risk is the mean of \(n(\widehat\beta-\beta)^2\) over 1,000 replications. Panels A and B present risk relative to the best of the three feasible rules in each design. Panel A averages or maximizes over all 28 designs at each \(n\). Panel B averages over the seven values of \(\rho\) at each \(F\) at \(n = 1671\). Panel C reports risk relative to the candidate minimizing the infeasible \(S_c\). Known moments evaluates the feasible criterion at population error moments. Details of these infeasible benchmarks are given in \Cref{sec:sim-diagnostics}.
\end{tabnotes}
\end{threeparttable}
\end{table}
This similarity between results in the Gaussian and non-Gaussian regimes is echoed in \Cref{tab:gs-laplace-compact} which shows that the feasible criterion again has the smallest average and maximum relative risk at both sample sizes. At \(n=1{,}671\), its risk is on average 2.9\% above the lowest risk in each design, compared with 2.6\% under Gaussian errors. The relative performance of the competing rules does change in some designs. At \(F=3.8\) and the original sample size, BCCH has lower average relative risk than cross-validation under Laplace errors, while the reverse holds under Gaussian errors. The proposed criterion has the lowest average relative risk in both cases. At the strongest level of identification, cross-validation again performs slightly better on average.
The comparison with the infeasible criterion is also similar across error
distributions. The criterion's average risk relative to the infeasible criterion is
1.162 under Laplace errors, compared with 1.164 under Gaussian errors.
As in the Gaussian design, using the population error moments in place of their estimates removes roughly 30\% of average
excess risk . Together, these results
suggest that the AMSE of the IV-LASSO estimator in a more general model
may be well approximated by its AMSE in the Gaussian model, in which
case the ranking of candidates derived in \Cref{sec:results} remains
informative.
\clearpage
\begin{thebibliography}{}
\bibitem[\protect\citeauthoryear{Bekker}{Bekker}{1994}]{Bekker-1994}
Bekker, P.~A. (1994).
\newblock Alternative approximations to the distributions of instrumental
variable estimators.
\newblock {\em Econometrica\/}~{\em 62\/}(3), 657--681.
\bibitem[\protect\citeauthoryear{Bellec and Zhang}{Bellec and
Zhang}{2021}]{BellecZhang-2021}
Bellec, P.~C. and C.-H. Zhang (2021).
\newblock Second-order {Stein}: {SURE} for {SURE} and other applications in
high-dimensional inference.
\newblock {\em The Annals of Statistics\/}~{\em 49\/}(4), 1864--1903.
\bibitem[\protect\citeauthoryear{Belloni, Chen, Chernozhukov, and
Hansen}{Belloni et~al.}{2012}]{BCCH-2012}
Belloni, A., D.~Chen, V.~Chernozhukov, and C.~Hansen (2012).
\newblock Sparse models and methods for optimal instruments with an application
to eminent domain.
\newblock {\em Econometrica\/}~{\em 80\/}(6), 2369--2429.
\bibitem[\protect\citeauthoryear{Belloni and Chernozhukov}{Belloni and
Chernozhukov}{2013}]{BelloniChernozhukov-2013}
Belloni, A. and V.~Chernozhukov (2013).
\newblock Least squares after model selection in high-dimensional sparse
models.
\newblock {\em Bernoulli\/}~{\em 19\/}(2), 521--547.
\bibitem[\protect\citeauthoryear{Belloni, Chernozhukov, and Hansen}{Belloni
et~al.}{2014}]{BelloniChernozhukovHansen-2014}
Belloni, A., V.~Chernozhukov, and C.~Hansen (2014).
\newblock Inference on treatment effects after selection among high-dimensional
controls.
\newblock {\em The Review of Economic Studies\/}~{\em 81\/}(2), 608--650.
\bibitem[\protect\citeauthoryear{Bickel, Ritov, and Tsybakov}{Bickel
et~al.}{2009}]{BickelRitovTsybakov-2009}
Bickel, P.~J., Y.~Ritov, and A.~B. Tsybakov (2009).
\newblock Simultaneous analysis of {Lasso} and {Dantzig} selector.
\newblock {\em The Annals of Statistics\/}~{\em 37\/}(4), 1705--1732.
\bibitem[\protect\citeauthoryear{Bogachev}{Bogachev}{1998}]{Bogachev-1998}
Bogachev, V.~I. (1998).
\newblock {\em Gaussian Measures}, Volume~62 of {\em Mathematical Surveys and
Monographs}.
\newblock Providence, RI: American Mathematical Society.
\bibitem[\protect\citeauthoryear{Carrasco}{Carrasco}{2012}]{Carrasco-2012}
Carrasco, M. (2012).
\newblock A regularization approach to the many instruments problem.
\newblock {\em Journal of Econometrics\/}~{\em 170\/}(2), 383--398.
\bibitem[\protect\citeauthoryear{Chamberlain}{Chamberlain}{1987}]{Chamberlain-1987}
Chamberlain, G. (1987).
\newblock Asymptotic efficiency in estimation with conditional moment
restrictions.
\newblock {\em Journal of Econometrics\/}~{\em 34\/}(3), 305--334.
\bibitem[\protect\citeauthoryear{Chernozhukov, Chetverikov, Demirer, Duflo,
Hansen, Newey, and Robins}{Chernozhukov et~al.}{2018}]{CCDDHNR-2018}
Chernozhukov, V., D.~Chetverikov, M.~Demirer, E.~Duflo, C.~Hansen, W.~Newey,
and J.~Robins (2018).
\newblock Double/debiased machine learning for treatment and structural
parameters.
\newblock {\em The Econometrics Journal\/}~{\em 21\/}(1), C1--C68.
\bibitem[\protect\citeauthoryear{Chetverikov, Liao, and
Chernozhukov}{Chetverikov et~al.}{2021}]{ChetverikovLiaoChernozhukov-2021}
Chetverikov, D., Z.~Liao, and V.~Chernozhukov (2021).
\newblock On cross-validated {Lasso} in high dimensions.
\newblock {\em The Annals of Statistics\/}~{\em 49\/}(3), 1300--1317.
\bibitem[\protect\citeauthoryear{Chetverikov and S{\o}rensen}{Chetverikov and
S{\o}rensen}{2025}]{ChetverikovSorensen-2025}
Chetverikov, D. and J.~R.-V. S{\o}rensen (2025).
\newblock Selecting penalty parameters of high-dimensional {M}-estimators using
bootstrapping after cross validation.
\newblock {\em Journal of Political Economy\/}~{\em 133\/}(10), 3208--3248.
\bibitem[\protect\citeauthoryear{Donald and Newey}{Donald and
Newey}{2001}]{DonaldNewey-2001}
Donald, S.~G. and W.~K. Newey (2001).
\newblock Choosing the number of instruments.
\newblock {\em Econometrica\/}~{\em 69\/}(5), 1161--1191.
\bibitem[\protect\citeauthoryear{Farrell, Liang, and Misra}{Farrell
et~al.}{2021}]{FarrellLiangMisra-2021}
Farrell, M.~H., T.~Liang, and S.~Misra (2021).
\newblock Deep neural networks for estimation and inference.
\newblock {\em Econometrica\/}~{\em 89\/}(1), 181--213.
\bibitem[\protect\citeauthoryear{Gilchrist and Sands}{Gilchrist and
Sands}{2016}]{gilchrist-glassberg-2016}
Gilchrist, D.~S. and E.~G. Sands (2016).
\newblock Something to talk about: Social spillovers in movie consumption.
\newblock {\em Journal of Political Economy\/}~{\em 124\/}(5), 1339--1382.
\bibitem[\protect\citeauthoryear{Hahn, Hausman, and Kuersteiner}{Hahn
et~al.}{2004}]{HahnHausmanKuersteiner2004}
Hahn, J., J.~Hausman, and G.~Kuersteiner (2004).
\newblock Estimation with weak instruments: Accuracy of higher-order bias and
{MSE} approximations.
\newblock {\em The Econometrics Journal\/}~{\em 7\/}(1), 272--306.
\bibitem[\protect\citeauthoryear{Heinonen}{Heinonen}{2005}]{Heinonen-2005}
Heinonen, J. (2005).
\newblock Lectures on {Lipschitz} analysis.
\newblock Report 100, University of Jyv{\"a}skyl{\"a}, Jyv{\"a}skyl{\"a},
Finland.
\bibitem[\protect\citeauthoryear{Isserlis}{Isserlis}{1918}]{Isserlis-1918}
Isserlis, L. (1918).
\newblock On a formula for the product-moment coefficient of any order of a
normal frequency distribution in any number of variables.
\newblock {\em Biometrika\/}~{\em 12\/}(1--2), 134--139.
\bibitem[\protect\citeauthoryear{Kuersteiner and Okui}{Kuersteiner and
Okui}{2010}]{KuersteinerOkui-2010}
Kuersteiner, G. and R.~Okui (2010).
\newblock Constructing optimal instruments by first-stage prediction averaging.
\newblock {\em Econometrica\/}~{\em 78\/}(2), 697--718.
\bibitem[\protect\citeauthoryear{Laurent and Massart}{Laurent and
Massart}{2000}]{LaurentMassart-2000}
Laurent, B. and P.~Massart (2000).
\newblock Adaptive estimation of a quadratic functional by model selection.
\newblock {\em The Annals of Statistics\/}~{\em 28\/}(5), 1302--1338.
\bibitem[\protect\citeauthoryear{Nagar}{Nagar}{1959}]{Nagar-1959}
Nagar, A.~L. (1959).
\newblock The bias and moment matrix of the general {$k$}-class estimators of
the parameters in simultaneous equations.
\newblock {\em Econometrica\/}~{\em 27\/}(4), 575--595.
\bibitem[\protect\citeauthoryear{Newey}{Newey}{1990}]{Newey-1990}
Newey, W.~K. (1990).
\newblock Efficient instrumental variables estimation of nonlinear models.
\newblock {\em Econometrica\/}~{\em 58\/}(4), 809--837.
\bibitem[\protect\citeauthoryear{Newey and Smith}{Newey and
Smith}{2004}]{NeweySmith2004}
Newey, W.~K. and R.~J. Smith (2004).
\newblock Higher order properties of {GMM} and generalized empirical likelihood
estimators.
\newblock {\em Econometrica\/}~{\em 72\/}(1), 219--255.
\bibitem[\protect\citeauthoryear{Okui}{Okui}{2011}]{Okui-2011}
Okui, R. (2011).
\newblock Instrumental variable estimation in the presence of many moment
conditions.
\newblock {\em Journal of Econometrics\/}~{\em 165\/}(1), 70--86.
\bibitem[\protect\citeauthoryear{Rothenberg}{Rothenberg}{1984}]{Rothenberg1984}
Rothenberg, T.~J. (1984).
\newblock Approximating the distributions of econometric estimators and test
statistics.
\newblock In Z.~Griliches and M.~D. Intriligator (Eds.), {\em Handbook of
Econometrics}, Volume~2, Chapter~15, pp.\ 881--935. Elsevier.
\bibitem[\protect\citeauthoryear{Sargan}{Sargan}{1976}]{Sargan1976}
Sargan, J.~D. (1976).
\newblock Econometric estimators and the {Edgeworth} approximation.
\newblock {\em Econometrica\/}~{\em 44\/}(3), 421--448.
\bibitem[\protect\citeauthoryear{Stein}{Stein}{1981}]{Stein-1981}
Stein, C.~M. (1981).
\newblock Estimation of the mean of a multivariate normal distribution.
\newblock {\em The Annals of Statistics\/}~{\em 9\/}(6), 1135--1151.
\bibitem[\protect\citeauthoryear{Tibshirani}{Tibshirani}{1996}]{Tibshirani-1996}
Tibshirani, R. (1996).
\newblock Regression shrinkage and selection via the {Lasso}.
\newblock {\em Journal of the Royal Statistical Society: Series B\/}~{\em
58\/}(1), 267--288.
\bibitem[\protect\citeauthoryear{Tibshirani}{Tibshirani}{2013}]{Tibshirani-2013}
Tibshirani, R.~J. (2013).
\newblock The {Lasso} problem and uniqueness.
\newblock {\em Electronic Journal of Statistics\/}~{\em 7}, 1456--1490.
\bibitem[\protect\citeauthoryear{Tibshirani and Taylor}{Tibshirani and
Taylor}{2012}]{TibshiraniTaylor-2012}
Tibshirani, R.~J. and J.~Taylor (2012).
\newblock Degrees of freedom in {Lasso} problems.
\newblock {\em The Annals of Statistics\/}~{\em 40\/}(2), 1198--1232.
\bibitem[\protect\citeauthoryear{van~de Geer and B{\"u}hlmann}{van~de Geer and
B{\"u}hlmann}{2009}]{VandeGeerBuhlmann-2009}
van~de Geer, S.~A. and P.~B{\"u}hlmann (2009).
\newblock On the conditions used to prove oracle results for the {Lasso}.
\newblock {\em Electronic Journal of Statistics\/}~{\em 3}, 1360--1392.
\bibitem[\protect\citeauthoryear{Velez}{Velez}{2026}]{Velez-2026}
Velez, A. (2026).
\newblock Debiased machine learning with many cross-fitting folds.
\newblock Working paper, Cornell University. August 2, 2026.
\bibitem[\protect\citeauthoryear{Wager and Athey}{Wager and
Athey}{2018}]{WagerAthey-2018}
Wager, S. and S.~Athey (2018).
\newblock Estimation and inference of heterogeneous treatment effects using
random forests.
\newblock {\em Journal of the American Statistical Association\/}~{\em
113\/}(523), 1228--1242.
\bibitem[\protect\citeauthoryear{Wainwright}{Wainwright}{2009}]{Wainwright-2009}
Wainwright, M.~J. (2009).
\newblock Sharp thresholds for high-dimensional and noisy sparsity recovery
using {$\ell_1$}-constrained quadratic programming ({Lasso}).
\newblock {\em IEEE Transactions on Information Theory\/}~{\em 55\/}(5),
2183--2202.
\bibitem[\protect\citeauthoryear{Zhao and Yu}{Zhao and Yu}{2006}]{ZhaoYu-2006}
Zhao, P. and B.~Yu (2006).
\newblock On model selection consistency of {Lasso}.
\newblock {\em Journal of Machine Learning Research\/}~{\em 7}, 2541--2563.
\end{thebibliography}
\clearpage