EconBase
← Back to paper

Binary response model with many weak instruments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

102,372 characters · 18 sections · 170 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Binary response model with many weak instruments

abstractThis paper considers an endogenous binary response model with many weak instruments. \textcolor{black}{We employ a control function approach and a regularization scheme to obtain better estimation results for the endogenous binary response model in the presence of many weak instruments.} Two consistent and asymptotically normally distributed estimators are provided, each of which is called a regularized conditional maximum likelihood estimator (RCMLE) and a regularized nonlinear least squares estimator (RNLSE). Monte Carlo simulations show that the proposed estimators outperform the existing ones when there are many weak instruments. We use the proposed estimation method to examine the effect of family income on college completion. JEL codes: C31, C35 Keywords: Regularization, Control function, Function-valued instrumental variables, Probit

Introduction

The conventional two-stage least squares (TSLS) estimation method has been widely used in empirical studies in economics and other fields of social science, even when the dependent variable is binary (see, e.g., \citealp*{Miguel2004}; Norton2008; \citealp*{Nunn2014}; Bastian2018). In spite of it not being recommended from a theoretical perspective, the empirical popularity of TSLS in binary response models is attributed not only to its ease of interpretation, but also to the fact that some essential statistical properties of estimators for {\it{nonlinear}} binary response models have only been studied in limited settings. For example, in contrast to the growing body of work on instrumental variable estimators for linear models with many instruments, little attention has been given to the statistical properties of estimators for binary response models in a similar context. Moreover, only a little is known when there are weak or nearly weak instruments, see e.g., Magnusson2010 and Dufour2018.

\textcolor{black}{This paper aims to develop a valid estimation and inference method for the endogenous binary response model when the dimension of the instrumental variable (or sometimes called the number of instruments) is very large.} In principle, using a large number of instruments is helpful for improving estimation efficiency. However, in spite of this advantage, this is not always recommended in finite samples. \textcolor{black}{This is because the use of many instruments can result in a less stable inverse of the covariance (see Examples (ref) and (ref)) and/or introduce bias (see Bekker1994). In such a case, the usual asymptotic approximation becomes less reliable. Therefore, in practice, a different approach is necessary to effectively utilize the information from a large number of instruments.} For example, one may suggest using a few instruments selected by a certain criterion, see, e.g., donald2001choosing, \textcolor{black}{Belloni}, and Caner2014. However, in the current study, we allow each element of a large number of instruments to be \textsl{nearly weak} (see Newey2009a). This consideration makes the consistency results of variable selection procedures in the aforementioned papers no longer valid even in linear models. Therefore, to handle the scenario with \textsl{“{many weak instruments}”}, we adopt a regularization scheme similar to those considered by Hausman2011, Carrasco2012, Hansen2014, Carrasco2015, and Han2019. Our regularization strategy is preferred to variable selection procedures for several reasons, which will be detailed in Section (ref).

Alongside the regularization scheme, we consider estimators similar to the two-stage conditional maximum likelihood estimator (2SCMLE) proposed by rivers1988limited, which is one of the most popular estimators for endogenous binary response models (see, e.g., Wooldridge2010). Specifically, as detailed in Section (ref), we use the first-stage residual, obtained using a regularization method, as an additional regressor in the second stage. Then, we define two estimators, each of which is called the regularized conditional maximum likelihood estimator (RCMLE) and the regularized nonlinear least squares estimator (RNLSE). Their asymptotic properties are studied under a parametric assumption similar to that in rivers1988limited. \textcolor{black}{While this parametric approach may not be satisfactory from a theoretical standpoint, a large number of citations to rivers1988limited suggest a clear preference for such an approach in the literature, at least when compared to semi- and non-parametric approaches. In addition, the parametric approach simplifies the asymptotic analysis of our estimators. This is partly because we do not need to estimate an infinite-dimensional parameter, such as the link function, under the parametric assumption. In this paper, we generalize the existing approach to allow for a possibly infinite-dimensional instrument, which is associated with an infinite-dimensional parameter in the first stage. Thus, a parametric assumption simplifies our analysis by reducing the number of infinite-dimensional parameters.}

\textcolor{black}{We consider nearly weak instruments in the framework of Newey2009a.} These instruments are commonly observed in empirical studies involving a binary dependent variable. For example, Miguel2004, Norton2008, Nunn2014, Frijters2009, and Bastian2018 report small values of the first-stage F test statistic that is known to be closely related to the weakness of instruments in the linear model (\citealp*{Staiger1997}; Stocka). Although this statistic is not a perfect measure of instrument strength in binary response models (e.g., \citealp*{frazier2020weak}), such small F test statistics suggest that their instrumental variables might be weak. However, in spite of its practical importance, only a few studies concern the issue of (nearly) weak instruments in endogenous binary response models. frazier2020weak study a way to test the weakness of instruments using a distorted J test under a parametric assumption. Andrews2013 investigate asymptotic properties of the generalized method of moments (GMM) estimator under weak identification and apply their estimation approach to an endogenous probit model. However, even in these articles, the case with many weak instruments has not been studied.

This paper is technically different from previous studies on many weak instruments. Specifically, we allow the instrumental variable to be function-valued; examples of function-valued random variables include the age-specific fertility rate (Florence2015), the skill-specific share of immigrants (seong2021), and the continuum of moments of an exogenous variable (Carrasco2012). Functional instrumental variables have received recent attention in the econometrics literature, but have mostly been used in linear models (e.g., Carrasco2012; Florence2015; Carrasco2015a, Carrasco2015; Benatia2017; seong2021). However, in contrast to the cases considered therein, our moment condition is nonlinear and not additively separable from the error term. Moreover, to address the endogeneity, we employ the control function approach. The limiting distributions of our estimators turn out to depend on the limiting distribution of the first-stage estimator that is given by an operator acting on a possibly infinite-dimensional space in our setting (see Section (ref)). These differences complicate our asymptotic analysis and make it technically distinguished from those of estimators for linear models with many weak instruments (e.g., Hansen2014; Carrasco2012; Carrasco2015; Benatia2017) and GMM estimators associated with an additively separable second stage (e.g., Hausman2011; Andrews2013; Han2019).

This paper is relevant to practitioners who need to control for endogeneity, using many weak instruments when the outcome is binary. We revisit the work by Bastian2018 and use our estimators to study the effect of family income in childhood and adolescent years on college completion in young adulthood. Following the authors, we first use three measures of Earned Income Tax Credit (EITC) exposure to address the endogeneity of family income, and then employ interactions of the EITC measures with 47 state-of-residence dummies as additional instruments. \textcolor{black}{Similar to the bias of TSLS estimates toward the OLS observed in the literature on many weak instruments in linear models, the 2SCMLE proposed by rivers1988limited tends to be biased toward the naive probit estimator that does not account for the endogeneity underlying the model. In contrast, our estimates appear to be robust to the inclusion of additional instruments.}

The paper is organized as follows. In Section (ref), we discuss the model of interest. In Section (ref), we propose our estimators and present their asymptotic properties. Section (ref) summarizes Monte Carlo simulation results. The empirical application is given in Section (ref). Section (ref) concludes. All proofs of the results given in this paper are provided in the Online Supplement. We include an additional simulation result and empirical application in the same supplement.

The Model

We consider the following model with independent and identically distributed (iid) data.

equation[equation omitted — 143 chars of source]

where $1\{ A \}$ is the indicator function that takes 1 if $ A$ is true and 0 otherwise. Thus, the outcome variable $y_i$ is binary. The $d_e$-dimensional explanatory variable $Y_{2i}$ consists of endogenous and exogenous variables; if a row of $Y_{2i}$ is exogenous, the corresponding row of $V_i$ is equal to zero. \textcolor{black}{The instrument $Z_i$ ($=Z(x_i)$) is a function of $x_i \in \mathbb{R}^{d_x}$ satisfying $\mathbb E[u_i | x_i] = 0$ and $\mathbb E \left[V_i | x_i \right] = 0$.}

The above model is similar to the standard endogenous binary response model considered by e.g., rivers1988limited, Rothe2009, and Blundell_Powell2004, except that it allows for a general class of random variables, such as a function-valued random variable, to be considered as the instrument. \textcolor{black}{The technical details of $Z_i$ for the general case will be discussed in Section (ref).} Instead, in this section, we provide examples of various types of instruments that can be considered in our setting. First, if $x_i =(x_{1i}, \ldots, x_{d_x i})'$, then $Z_i $ may be given by $ x_i$, a standard $d_x$-dimensional vector of different instruments (see, e.g., Examples (ref) and (ref)). If $x_i$ is a scalar, we may let $Z_i$ be $(1, x_i ,x_i^2,\ldots, x_i^K)'$. If so, our first-stage estimator can be understood as a non-parametric estimator of $\mathbb E [Y_{2i} | x_i ] $ in some cases, see Carrasco2012 and Antoine2014. \textcolor{black}{Lastly, in Section (ref), we allow for a function-valued random variable to be considered as the instrument $Z_i$, as discussed in recent articles including Carrasco2012, Florence2015 and Benatia2017.} Examples of function-valued random variables include (but are not limited to) the density of $x_i$, the continuum of moments of $x_i$ (Carrasco2012), the rainfall growth curve (Example (ref)), and the skill-specific share of immigrants (seong2021).

\textcolor{black}{The potential weakness of instruments is characterized by the first-stage parameter $\Pi_n$, whose magnitude may depend on $n$. This will be further detailed in Assumption (ref). A similar setting in linear models can be found in Chao2005, Hansen2014, and Carrasco2015, among many others. The parameter $\Pi_n$ will be represented by a matrix or a linear operator with rank $d_e$, depending on the instrumental variable used in the analysis. That is, regardless of the dimension of $Z_i$, $\Pi_n Z_i$ is a $d_e$-dimensional vector that combines information from the instrumental variable of a large, and possibly infinite, dimension. This will be further detailed in Section (ref).}

Let \textcolor{black}{$\eta_i =u_i - \mathbb E[u_i |x_i, V_i] = u_i+ V_i ' \psi_0$}. Then, the outcome equation can be written as follows:

equation[equation omitted — 99 chars of source]

\textcolor{black}{If $\eta_i |(x_i , V_i )\sim \mathcal N(0,\sigma_\eta ^2)$\footnote{\textcolor{black}{The distributional assumption is satisfied if $(u_i, V_i')'|x_i \sim \mathcal N(0, \mathbf{\Sigma} )$, as in rivers1988limited.}} and the variable $V_i$ is available, \hyperref[{model:3eq}]{\tagform@{\ref*{model:3eq}}} can be estimated by maximum likelihood estimation (MLE).} In this regard, $V_i$ is often called the control function in the literature. However, in practice, $V_i$ is not observable and has to be replaced by its estimate. rivers1988limited suggested replacing it with the residual from first-stage linear regression, denoted $\widehat{V}_i$. Then, the parameters in \hyperref[{model:3eq}]{\tagform@{\ref*{model:3eq}}} are estimable by maximizing the conditional likelihood of $y_i | (x_i, \widehat V_i)$ under a parametric assumption. This approach has become one of the most popular estimation methods for endogenous binary response models and was extended by Rothe2009 to the semi-parametric case.

Although the aforementioned estimators generally have good asymptotic properties, it is not likely to be the case in our setting. This is because, to ensure such a good property, the difference between $\widehat{V}_i$ and $ V_i$ should be asymptotically negligible. \textcolor{black}{This requirement is not likely to be satisfied, especially when $Z_i$ is of a large dimension or has a nearly singular covariance matrix,} see, e.g., Examples (ref) and (ref). In these cases, the first-stage estimator may be asymptotically biased, resulting in $\widehat{V}_i$ being asymptotically different from $V_i$. Another concern on \citepos{rivers1988limited} approach is that the asymptotic distribution of their estimator depends on the asymptotic distribution of the first-stage estimator. \textcolor{black}{In Section (ref), the first-stage parameter and its estimator are given by linear operators from a possibly infinite-dimensional Hilbert space to $\mathbb R^{d_e}$, whereas in rivers1988limited, they are given by matrices. This difference makes it difficult to apply their asymptotic results in our setting. }

Below, we provide examples where \citepos{rivers1988limited} approach may not be appropriate due to the nearly singular covariance of $Z_i$, even with a large sample size. \textcolor{black}{In the examples, the instrument $Z_i$ is given by $(W_i' , \widetilde{Z}_i ' )'$, where $W_i$ and $\widetilde{Z}_i$ denote the included and excluded exogenous variables, respectively. $Y_{2i}$ in \hyperref[{model:1eq}]{\tagform@{\ref*{model:1eq}}} is understood as a vector that includes $W_i$ and endogenous variables (whose notations are specific to each context).} We use the condition number (the ratio of the largest to smallest eigenvalues) of a covariance matrix as a measure of how close the matrix is to being singular.

example[]\normalfont Bastian2018 studied the impact of family income in childhood and adolescent years on academic achievement in early adulthood using the linear probability model. However, from a theoretical perspective, their result may be improved by using a binary response model; this is because the linear model often produces (i) the predicted probability outside the interval $[0,1]$, (ii) constant marginal effects of explanatory variables on the probability of outcome occurring, regardless of the values of the explanatory variables, and (iii) less efficient estimation results due to heteroscedasticity. Moreover, when endogeneity is present, theoretical results developed for the linear model do not always remain valid in the binary response model (li2019binary; frazier2020weak). Hence, as an alternative, one may consider the following model.\begin{equation*} y_i = {1} \{ \beta_0 + I_{i}'\beta_1 + W_{1i} ' \beta_2 + W_{2i,t} '\beta_3 \geq u_i \} ,\quad I_{i} = \pi_{0} + \widetilde Z_i ' \pi_1 + W_{1i} ' \pi_{2} + W_{2i,t} ' \pi_{3} + V_{i} , \end{equation*} where $y_i$ is 1 if individual $i$ is a college graduate and 0 otherwise. $I_i = ( I_{i,\text{(0-5)}}, I_{i,\text{(6-12)}},$ $I_{i,\text{(13-18)}})'$, where $I_{i,(a)}$ is the family income for individual $i$ at age interval $a$. $W_{1i}$ and $W_{2i}$ are vectors of exogenous variables detailed in Section (ref). As instruments, the authors suggested using measures of EITC exposure for individual $i$ in three age intervals: 0-5, 6-12, and 13-18. That is, $\widetilde Z_i = (\text{EITC}_{i,(\text{0-5})},\text{EITC}_{i,(\text{6-12})}\text{EITC}_{i,(\text{13-18})})'$. In Section (ref), we use additional instruments generated by interacting the EITC measures with state-of-residence dummies. This approach is expected to improve estimation efficiency by reducing variances of estimators. It also enables us to account for heterogeneous effects of EITC exposure on family income that may be influenced by the state of residence, thereby mitigating potential bias attributed to, e.g., state-specific costs of living. The maximum number of instruments is 144, and the sample size is 2,654. The largest and smallest eigenvalues of the sample covariance are approximately 1.94 and $9.35 \cdot 10^{-18}$, indicating a very large condition number. \textcolor{black}{This large condition number suggests that the inverse of the covariance is less stable, and the existing asymptotic approximation using this inverse may not perform well. Here, regularization will improve estimation results by stabilizing the inverse of the covariance.}
example[]\normalfont Angrist1991 has been revisited in numerous articles, particularly in the many-instruments literature (e.g., donald2001choosing; Carrasco2012; Hansen2014). Suppose we are interested in the effect of educational achievement on employment probability, rather than on earnings as studied by Angrist1991. One may consider the model below, where $y_i = 1$ if individual $i$ is employed for a full year and 0 otherwise.\begin{equation*} y_i = 1\{ \beta_0 + Educ_i \beta_1 + W_i ' \beta_2 \geq u_i \},\quad\quad Educ_i = \widetilde Z_{i} ' \Pi_1 + W_i ' \Pi_2 + V_i. \end{equation*} The variable $\text{Educ}_i$ is the years of schooling of individual $i$. $W_i$ consists of a constant, 9 year-of-birth dummies (YoB), and 50 state-of-birth dummies (SoB). As in Angrist1991, let $\widetilde Z_i$ be the 180-dimensional vector of 3 quarter-of-birth dummies and their interactions with YoB and SoB. The condition number of the covariance of $ (\widetilde{Z}_i ' , W_i ') ' $ is approximately 4,941, even after standardization. This suggests near singularity, even with a large sample size of 329,509.

Regularization with many weak instruments

We use regularization to resolve the aforementioned issues related to many, and nearly weak, instruments. There are various regularization methods available in the literature. A popular approach is reducing the number of instruments by (i) choosing the optimal number of instruments that minimizes a model selection criterion (e.g., donald2001choosing) or (ii) using a variable selection procedure such as Lasso (e.g., Belloni) or Elastic-net (e.g., Caner2014). This approach works particularly well if the signal from instruments is (i) concentrated on a few of them (called the sparsity condition) and (ii) strong enough to consistently select a set of instruments with non-zero first-stage coefficients. However, despite its popularity, this is not always the best approach to use many instruments. For example, the sparsity condition may not hold when the variables of interest are highly correlated with each other. In such a case, there could be a model selection mistake of choosing irrelevant instruments. \textcolor{black}{The selection mistake is not desirable especially in finite samples, because the resulting estimators could be less efficient. For additional concerns related to the selection mistake, see Hansen2014.} Hence, it would be advisable to avoid using the variable selection procedure if such a selection mistake is likely to occur. In this regard, the variable selection procedure may not be appropriate in our setting, because (i) we consider many (nearly) weak instruments and (ii) in such a scenario, variable selection procedures tend to select no instrument or almost randomly choose a set of instruments (see, e.g., Belloni). Furthermore, as in Examples (ref) and (ref), if instruments are given by a set of dummies, it is not recommended to choose a few of them; this is because (i) a selected set of instruments will depend on how the variables are encoded and (ii) the categorical variables are not likely to be represented by only a few of them.

In this paper, we regularize the covariance of instruments. Our regularization scheme, formally introduced in Section (ref), is preferred to the above-mentioned variable selection procedures for several reasons. Firstly, our approach does not need a sparsity condition, and secondly, is free from the aforementioned concern on selection mistakes. Moreover, in practice, valid instruments are likely to be (nearly) weak. In this case, it has been shown that the asymptotic properties of the estimators obtained by regularizing the covariance of instruments remain valid (e.g., Hausman2011; Hansen2014; Carrasco2015a), while the variable selection procedures no longer produce consistent estimation results. In contrast to the popularity of regularizing the covariance in linear models (Carrasco2012; Hansen2014; Carrasco2015), its application to binary response models has not been fully explored.

Estimator and Asymptotic Properties

We first consider the simple case $Z_i \in \mathbb R^{d_z}$ for $d_e \leq d_z <\infty$. In this case, regularization is useful if the covariance of $Z_i$ is nearly singular, as in Examples (ref) and (ref). In Section (ref), we further extend the discussion to accommodate function-valued $Z_i$. In that case, not only is regularization required to address the singularity of the covariance, but a novel asymptotic analysis is also needed to establish the asymptotic properties of our estimators. In the following sections, $Z_i$ is assumed to be centered without loss of generality.

Simple case: a finite-dimensional instrumental variable

Estimator

Let $\mathcal K$ and $\mathcal K_n$ denote the covariance of $Z_i$ and its sample counterpart where $\mathcal K = \mathbb E[Z_iZ_i']$ and $\mathcal K_n =n^{-1}\sum_{i=1} ^n Z_i Z_i' $. In Section (ref), we will adapt these definitions for the general case. \textcolor{black}{Given the model \hyperref[{model:1eq}]{\tagform@{\ref*{model:1eq}}} and the exogeneity of $Z_i$ with respect to $V_i$ ($\mathbb E[V_i Z_i']=0$), we have the following first-stage moment condition:}

equation[equation omitted — 165 chars of source]

In many econometric studies, it has been assumed that the covariance $\mathcal K$ is invertible to solve \hyperref[{mod2:eq1:finite}]{\tagform@{\ref*{mod2:eq1:finite}}}. However, this assumption is not likely to be satisfied when the dimension of $Z_i$ is large, as in Examples (ref) and (ref). Hence, instead of the inverse of $\mathcal K$, we consider its regularized inverse. Here, we focus on Tikhonov regularization (in Section (ref), we extend it to various regularization methods satisfying certain conditions). Specifically, for a regularization parameter $\alpha >0$ and the identity matrix of dimension $d_z$, denoted $\mathcal I_{d_z}$, $\mathcal K_\alpha ^{-1}$ denotes the regularized inverse of $\mathcal K$ such that

equation*[equation* omitted — 107 chars of source]

The regularized solution to \hyperref[{mod2:eq1:finite}]{\tagform@{\ref*{mod2:eq1:finite}}} (denoted $\Pi_\alpha$) and its sample counterpart (denoted $\widehat \Pi_{\alpha}$) are given by

equation[equation omitted — 364 chars of source]

Using $\widehat{\Pi}_\alpha$ in \hyperref[{eq:first-est:finite}]{\tagform@{\ref*{eq:first-est:finite}}}, we compute the predicted value of $Y_{2i} $, denoted $\widehat{\gamma}_{i,\alpha}$, and the first-stage residual $\widehat{V}_{i,\alpha}$, i.e., $\widehat{\gamma}_{i,\alpha} = \widehat \Pi_{\alpha} Z_i $ and $\widehat{V}_{i,\alpha} = Y_{2i} -\widehat \Pi_{\alpha} Z_i$. We use them as regressors in the second stage for mathematical convenience.

Suppose $(u_i -\mathbb E[u_i|x_i, V_i])| (x_i, V_i) = u_i +V_i'\psi_0 | (x_i, V_i) \sim \mathcal {N} (0,\sigma_\eta ^2)$. Here, what benefits us is not the conditional normality of $u_i$, but rather the fact that the endogeneity of $Y_{2i}$ affects the second stage only through the conditional expectation of $u_i$. The normality assumption is employed only to facilitate our discussion, and may be replaced by another distributional assumption. We then normalize the conditional variance $\sigma_\eta ^2 $ to $1$ to achieve the identification of the second-stage parameters, see Manski1988 for details. Lastly, let $ g_{i} = ( \gamma_{i} ' , V_{i} ' )' = ( (\Pi_n Z_i)' , (Y_{2i} - \Pi_n Z_i)')'$, $\widehat{g}_{i,\alpha} = (\widehat{\gamma}_{i,\alpha}', \widehat V_{i,\alpha}')'$ and $\theta = ( \beta ' , \beta '+ \psi ' )'$. In contrast to $\widehat{g}_{i,\alpha}$, the vector $g_i$ is not defined with the regularized inverse. The RCMLE (denoted $\widehat{\theta}_M$) and the RNLSE (denoted $\widehat{\theta}_N$) are the estimators of the true coefficients $\theta_{0M}$ and $\theta_{0N}$, respectively. $\widehat \theta_{ j }$ and $ \theta_{ 0j }$ are defined as follows: for $j \in \{ M,N\}$,

align[align omitted — 538 chars of source]

and the function $m_{j} (\theta , g_i ) $ is $ y_i\log\Phi (g_i ' \theta ) + (1-y_i )\log (1-\Phi (g_i '\theta ) )$ (resp.\ $ -\frac{1}{2} (y_i - \Phi (g_{i} ' \theta ) )^2$) if $j=M$ (resp.\ $j=N$). The above expectation is taken with respect to the distribution of $g_i$ and $y_i$ for sample size $n$, but we suppress the index $n$ unless it is necessary. Under the identification conditions in Section (ref), we have $\theta_{0M} = \theta_{0N} = \theta_0$.

remark\normalfont $\mathbb E [y_i | x_i ,V_{i,\alpha} ] $ is in general different from $\mathbb E [y_i | x_i ,V_{i} ] $. To see this in detail, let $\gamma_{i,\alpha} = \Pi_\alpha Z_i$, and $V_{i,\alpha} = Y_{2i} - \gamma_{i,\alpha}$. Then, under the aforementioned assumptions, we have \begin{equation} \mathbb E [y_i | x_i ,V_{i,\alpha} ] = \Phi ( \gamma_{i,\alpha} ' \beta_0 + V_{i,\alpha} ' (\beta_0 + \psi_0)- \psi _0' \Pi_{n}(\mathcal I_{d_z} - \mathcal K_\alpha^{-1} \mathcal K) Z_i ). \end{equation} If $\alpha>0$ is fixed, the conditional expectation in \hyperref[{model:6eq}]{\tagform@{\ref*{model:6eq}}} depends not only on the error $V_{i,\alpha}$ but also on a bias term associated with $\Pi_{n}(\mathcal I_{d_z} - \mathcal K_\alpha^{-1} \mathcal K) $. Thus, to mitigate its potential effect on the asymptotic properties of our estimators, $\alpha$ is assumed to shrink to zero as $n \hspace{-0.2em}\to\hspace{-0.2em} \infty$ throughout the paper.
remark\normalfont In this paper, we apply the regularization approach to the conditional maximum likelihood and nonlinear least squares (NLS) estimation methods. These methods are chosen because of their ease of implementation and popularity in empirical research. While other estimation methods, such as the instrumental variable probit and the limited information maximum likelihood in rivers1988limited, are also applicable, we do not pursue them further in the current paper.

Asymptotic Properties

This section presents the asymptotic properties of our estimators. Hereafter, we use $\lambda_{\min}(A)$ to denote the minimum eigenvalue of a matrix $A$.

assumptionFor $\Pi_{n}$ satisfying \hyperref[{mod2:eq1:finite}]{\tagform@{\ref*{mod2:eq1:finite}}}, there exists a bounded matrix $\Pi_0 : \mathbb R^{d_z} \hspace{-0.2em}\to\hspace{-0.2em} \mathbb R^{d_e}$ such that $\Pi_{n} = {\Lambda}_n \Pi_0 /\sqrt{n}$ where ${\Lambda}_n = \widetilde {\Lambda} \text{diag}(\mu_{1n} , \ldots, \mu_{d_{e}n})$. The matrix $\widetilde {\Lambda}: \mathbb R^{d_e} \to \mathbb R^{d_e} $ is bounded, and $\lambda_{\min}(\widetilde {\Lambda} \widetilde {\Lambda} ' )$ is bounded away from zero. For each $j$, either $\mu_{jn} = \sqrt{n}$ (strong) or $\mu_{jn} / \sqrt{n} \hspace{-0.2em}\to\hspace{-0.2em}0$ (weak)\textcolor{black}{, and $\mu_{m,n} = \underset{1 \leq j \leq d_{e}}{\min} \mu_{jn} \hspace{-0.2em}\to\hspace{-0.2em} +\infty$ as $n \hspace{-0.2em}\to\hspace{-0.2em}+ \infty$}.

Hereafter, we let $f_i = \Pi_0 Z_i$. \textcolor{black}{The above assumption is similar to Assumption 1 in Newey2009a, under which many weak moment asymptotics are studied for GMM models with linear and nonlinear moment conditions. Considering that the GMM class is general enough to encompass the MLE and the NLS (NEWEY1994), this framework would be suitable for studying consequences of many weak instruments in our setting.} A similar assumption can be found in the literature on many weak instruments in linear models, including Chao2005, Hansen2014, and Carrasco2015, to mention only a few.

Assumption (ref) essentially rules out weak instruments in the sense of Staiger1997. Specifically, under the assumption, each row of $f_i$ could be either (nearly) weak or strong, depending on the value of $\mu_{jn}$. If $\mu_{jn} = \sqrt{n}$ for all $j$, then every element of $f_i$ is strong, and regularization will play a role of reducing the many instruments bias discussed in Bekker1994. The above assumption is also related to the nearly weak moment condition in AR2009.

assumption\begin{enumerate*}[(i)] • $\{ u_i, V_{i}', x_i' \}_{i=1} ^n$ is iid and \textcolor{black}{$(u_i, V_i')'|x_i $ follows the joint normal distribution with mean zero and finite positive definite covariance $\mathbf \Sigma_{uV}$, and satisfies that $\mathbb E[u_i|x_i, V_i] = -\psi_0'V_i$ and $\mathop{\hbox{\rm Var}}(u_i |x_i, V_i) =1$;} • \textcolor{black}{There exists $c>0$ such that $\mathbb E [\Vert Z_i \Vert^2]\leq c $ and $\lambda_{\min}(\mathbb E [f_i f_i ' ] ) \geq 1/c$.} \end{enumerate*}
assumptionThe parameter space of $\theta$, denoted $ \Theta$, is compact. $\theta_{0}$ is the unique point satisfying Assumption (ref).(ref) and the model \hyperref[{model:1eq}]{\tagform@{\ref*{model:1eq}}}, and is in the interior of $\Theta$.

\textcolor{black}{Assumption (ref) is similar to Assumption 1 in rivers1988limited. As they mentioned, this assumption is stronger than what we need to control for the endogeneity of $Y_{2i}$, and the assumption can be weaken by assuming that $(u_i-\mathbb E[u_i|x_i, V_i])| (x_i, V_i) =(u_i+\psi_0 ' V_i)| (x_i, V_i) \sim\mathcal N ( 0, 1 )$;} that is, the endogeneity of $Y_{2i}$ arises through $\mathbb E[u_{i}| x_i,V_i]$, which enables us to address the endogeneity via the inclusion of $V_i$. We normalize the conditional variance of $u_i| (x_i,V_i )$ to 1 for identification (see Manski1988). A different normalization could have been chosen, but we select the current normalization because it makes the distribution of $u_i |(x_i,V_i )$ free from nuisance parameters.

\textcolor{black}{Assumptions (ref) and (ref) ensure the identification of $\theta_0$ (NEWEY1994), and similar conditions can be found in frazier2020weak which concerns the endogenous binary response model with weak instruments.} Under the conditions, $\theta_{0M} = \theta_{0N} = \theta_0$.

We then discuss a condition on $\Pi_0$; in the condition below, $\operatorname{ran} \mathcal K$ denotes the range of $\mathcal K$.

assumptionLet $\pi_{\ell,0}$ be the $\ell$th row of $\Pi_0$. Then, for $\ell = 1,\ldots, d_e $, $\pi_{\ell,0}\in \operatorname{ran} \mathcal K$.

As detailed later in Section (ref), a primitive condition for Assumption (ref) is that $(\pi_{\ell,0} ' \varphi_j )^{2}$ decreases at a rate slower than the decaying rate of the square of eigenvalues of $\mathcal K$, where $\{\varphi_j\}_{j \geq 1}$ denotes the eigenvectors of $\mathcal K$. This assumption is crucial for identifying $\Pi_0$ without assuming the invertibility of $ \mathcal K$ (see Carrasco2007; Benatia2017 for details) and for controlling the bias discussed in Remark (ref). \textcolor{black}{Moreover, point identification of $\Pi_0$ simplifies our asymptotic analysis by providing an explicit form of the regularization bias and the asymptotic distribution of the first-stage estimator, which are directly related to the asymptotic distributions of our estimators.}

The theorem below presents the consistency of $\widehat{\theta}_j$ under the aforementioned assumptions.

theorem\normalfont Suppose that $\alpha \hspace{-0.2em}\to\hspace{-0.2em} 0$ and $\mu_{m,n} ^2 \alpha^{1/2}\hspace{-0.2em}\to\hspace{-0.2em} \infty$ as $n \hspace{-0.2em}\to\hspace{-0.2em} \infty$, and Assumptions (ref) to (ref) are satisfied. Then, $ \widehat \theta_{ j} - \theta_{0} \overset{p}\rightarrow 0 $ as $n \hspace{-0.2em}\to\hspace{-0.2em} \infty$ for $j \in\{ M,N \}$.

The conditions on $\alpha$ and $\mu_{m,n}$ in Theorem (ref) are related to the weakness of $f_i$. If $\alpha = O(n^{-a})$ and $\mu_{m,n} = O(n^{\ell})$ for some positive $a$ and $\ell\leq 1/2$, then the condition on $\mu_{m,n}^2\alpha^{1/2} $ suggests $0<a< 2$. This is a necessary condition for $\alpha$ to ensure the consistency of $\widehat{\theta}_j$. A similar condition can be found in Carrasco2015 and Hansen2014, where in the latter, the inverse of the number of instruments plays a role similar to $\alpha$ in the current study. The convergence rate of $\alpha$ allows us to control for the bias associated with the nearly weak signal of $f_i$ and the regularized inverse simultaneously.

Next, we discuss the asymptotic distribution of $\widehat \theta_j$. We follow rivers1988limited and address the endogeneity by using the estimated control function. As a result, the limiting distribution of $\widehat\theta_j$ turns out to depend on that of the first-stage estimator. However, in our setting, obtaining the limiting distribution of the first-stage estimator is not straightforward because of the use of a regularized inverse and the bias discussed in Remark (ref). Although a similar scenario has been considered for linear models (Carrasco2012; Hansen2014; Carrasco2015a,Carrasco2015), the asymptotic approaches therein cannot be applied in our context because of some distinct features of our models, such as nonlinearity and the use of the control function approach. To resolve these issues and to encompass the general case in Section (ref), we adapt the asymptotic approach in Chen2003. To this end, we employ the following assumption, in which a.s.n.\ denotes almost surely for $n$ large enough.

assumption\begin{enumerate*}[(i)] • For all $\theta \in \Theta$, $0< \inf_i\Phi (g_i ' \theta )\leq \sup_i\Phi (g_i ' \theta )<1$ a.s.n.; • There exists $c>0$ such that $ \sup_i \Vert f_i\Vert \leq c $ a.s.n.\ and $ \mathbb E[\Vert Z_i\Vert ^4 ] \leq c$. \end{enumerate*}

Assumption (ref) is necessary to bound a particular quantity in our proof, but can be relaxed by using a trimming term as in Rothe2009. Thus, the condition may not be restrictive in practice.

{We then present the asymptotic normality result. Let $\mathcal S_n^{-1}$ denote $\text{diag}(\sqrt{n} {\Lambda}_n ^{-1} , \mathcal I_{d_e}) $ and $g_{0i}= \mathcal S_n ^{-1} g_i = (f_i ' , V_i ' )'$. Moreover, let $\sigma_{\psi_0} ^2=\psi_0 ' \mathbb E[V_iV_i'|x_i]\psi_0 = \psi_0 ' \mathbf \Sigma_V\psi_0$, $m_{1M} ^2 (\theta, g_i) = - \dot{m}_{2M} (\theta, g_i)$, $ \dot{m}_{2M} (\theta, g_i) = - \left((y_i - \Phi (g_i'\theta))\phi (g_i ' \theta )/\left( \Phi (g_i ' \theta )(1-\Phi (g_i ' \theta )) \right)\right)^2 $, $ m_{1N} ^2 (\theta , g_i) =$ $ \left((y_i - \Phi( g_i ' \theta ))\phi\left(g_i ' \theta\right)\right)^2 $, and $ \dot{m}_{2N} (\theta, g_i) =\phi ^2 (g_i ' \theta )$. The matrix $\mathcal C_j $ is the solution to $\mathcal K^{1/2}\mathcal C_j' = \mathbb E[ \dot{m}_{2j}(\theta_0,g_i)Z_i g_{0i}' ]$, and $e_{\ell}$ denotes the unit vector whose $\ell$th element is equal to one. Lastly, $\Gamma _{1,0j}=\mathbb E[ \dot{m}_{2j}(\theta_0,g_i)g_i g_{i}' ]$ and $\mathcal J_j =\mathcal J_{1,j} + \mathcal J_{2,j}$ where $\mathcal J_{1,j} = \mathbb E [ {m}_{1j} ^2 ( \theta_{0} , g_i ) g_{0i} g_{0i} ' ] $ and $\mathcal J_{2,j}$ is a $2d_e \times 2d_e$ matrix whose $(\ell,k)$th element is given by $\sigma_{\psi_0} ^2 e_\ell ' \mathcal C_j \mathcal C_j' e_k$.}

theorem\normalfont Suppose that $\alpha \to 0$, $\mu_{m,n} ^2 \alpha\to\infty$ and $\sqrt{n} \alpha \to0$ as $n \hspace{-0.2em}\to\hspace{-0.2em}\infty$, and Assumptions (ref) to (ref) hold. Then, $\mathcal W_{0j} ^{-1/2} \sqrt{n} \mathcal S_n ^\prime \left(\widehat \theta_j - \theta_{0} \right) \overset{d}\rightarrow \mathcal N \left( 0,\mathcal I_{2d_e} \right) $ for $j \in \{M,N\}$, where $ \mathcal W_{0j} = (\mathcal S_n ^{-1}\Gamma_{1,0j}\mathcal S_{n}^{\prime -1} ) ^{-1} \mathcal J_j (\mathcal S_n ^{-1}\Gamma_{1,0j} \mathcal S_n ^{\prime -1} ) ^{-1}$.

The conditions in Theorem (ref) are closely related to $\mu_{m,n}$ in Assumption (ref). Specifically, if $\mu_{m,n} = O(n ^{\ell})$, the theorem remains valid when $1/4 <\ell \leq1/2$. Similar results can be found in AR2009 and HR2020; these articles show that GMM estimators are asymptotically normally distributed if their moment conditions shrink to zero at such a rate. \textcolor{black}{In linear models, $\mu_{m,n}$ is closely related to the convergence rate of the concentration parameter, and conditions on the number of instruments or the regularization parameter, in relation to $\mu_{m,n}$, have been employed in the literature (see Carrasco2015 for a comprehensive review). To the best of our understanding, the conditions are not directly testable, because $\mu_{m,n}$ is neither observable nor estimable. Consequently, the condition on $\mu_{m,n} ^2 \alpha$ should be interpreted as a theoretical requirement for $\alpha$ to shrink slowly depending on the weakness of instruments.} The second condition $\sqrt{n}\alpha\to 0$ suggests a certain convergence rate of $\alpha$. This condition is required to mitigate a potential effect of the asymptotic bias of $\widehat{\Pi}_\alpha $ on the limiting distribution of $\widehat{\theta}_j$.

remark\normalfont \textcolor{black}{In Theorem (ref), $\Lambda_n$ in $\mathcal S_n$ can be understood as the weight determining the convergence rate of $\widetilde{\Lambda} \widehat{\beta}_j$, as in Chao_et_al. This weighting matrix is unfortunately neither estimable nor observable, because $\Lambda_n$ is not available in practice. Furthermore, the above asymptotic normality result tells us that if there exists a weak instrument (i.e., $\mu_{m,n} /\sqrt{n} = o(1)$), the asymptotic covariance in rivers1988limited ($\mathcal S_n^{\prime -1} \mathcal W_{0j} \mathcal S_n^{-1}$ in our notation) is divergent. However, the matrix $\mathcal S_n$ does not play any role in the Wald-based testing procedure, thereby allowing us to conduct inference on $\theta_0$. For example, let $H_0: \theta_0 = \theta$ be the null hypothesis of interest, and for each $j \in \{M,N\}$, suppose there exists $\widetilde{\mathcal W}_j$ satisfying $ \mathcal S_n' \widetilde{\mathcal W}_{j} \mathcal S_n = \mathcal W_{0j} + o_p(1)$. That is, $\widetilde{\mathcal W}_j$ properly normalized by $\mathcal S_n$ is a consistent estimator of $\mathcal W_{0j}$. (An example of $\widetilde{\mathcal W}_j$ can be found in Section (ref).) Then, as in ANTOINE2012350 and frazier2020weak, the weight is canceled out in the Wald testing procedure; specifically, $ n(\widehat{\theta}_j - \theta)' \widetilde{\mathcal W}_j ^{-1} (\widehat{\theta}_j - \theta) = n (\widehat{\theta}_j - \theta)' \mathcal S_n (\mathcal S_n ' \widetilde{\mathcal W}_{j}\mathcal S_n)^{-1} \mathcal S_n ' (\widehat{\theta}_j - \theta)\to_d\chi^2 (2d_e) $ under $H_0$ where $\chi^2 (2d_e)$ is the chi-square distribution with $2d_e$ degrees of freedom. }
remark\normalfont If there exist non-zero $r_n $ and $\varsigma_\ast \in \mathbb R^{2d_e}\backslash \{0\}$ such that $ r_n n^{-1/2}\mathcal S_n ^{-1} \varsigma \to \varsigma_\ast $ for some $\varsigma \in \mathbb R ^{2d_e}\backslash \{0\}$, then $ (\varsigma'\widetilde{\mathcal W}_{j}\varsigma) ^{-1/2}\sqrt{n} \varsigma'(\widehat{\theta}_j - \theta_0) \overset{d}\rightarrow N(0,1) $\textcolor{black}{; as in Remark (ref), the scaling terms are canceled out and thus $r_n $ and $\mathcal S_n$ have no effect on obtaining the limiting distribution of the test statistic, see Newey2009a.} As an application, we may study the endogeneity of $Y_{2i}$ by testing $H_0: \psi_0 = 0$. The endogeneity could be tested by setting $\varsigma = (-\iota_{d_e}', \iota_{d_e}')'$ for the $d_e$-dimensional vector of ones $\iota_{d_e}$, if $\mu_{m,n} \Lambda_n ^{-1} \to \Lambda^{-1}$ and $\Lambda ^{-1}\iota_{d_e} \neq 0 $.

General case: a functional instrumental variable

Details on the instrumental variable

Our definition of $Z_i$ in the general case is similar to that in Carrasco2012 and Carrasco2015a,Carrasco2015. \textcolor{black}{Specifically, we let $\mathcal H$ denote the Hilbert space of square integrable functions defined on a compact interval $\mathfrak C \in\mathbb R$ with inner product $\langle h_1, h_2\rangle _{\mathcal H} = \int_{\mathfrak C} h_1 (t)h_2 (t) dt$ for $h_1,h_2 \in \mathcal H$ and its induced norm $\left\Vert h_1 \right\Vert_{\mathcal H} = \langle h_1, h_1\rangle_{\mathcal H}^{1/2}$. We let $Z_i= \{Z (x_i, t), t\in \mathfrak C\}$ be a random variable that takes values in $\mathcal H$. That is, if $(\Omega,\mathbb F, \mathbb P)$ denotes the underlying probability space, $Z_i $ is a measurable function from $ \Omega$ to $\mathcal H$ equipped with the Borel $\sigma$-field (see HK2012 for details on random elements in $\mathcal H$). The index $t $ is the deterministic argument of the function-valued instrumental variable and is suppressed for notational simplicity.\footnote{\textcolor{black}{More generally, we may consider the Hilbert space of square integrable functions with respect to $\tau$ where $\tau $ is some positive, continuous and uniformly bounded measure on $\mathfrak C$, as in Carrasco2012 and Carrasco2015a, Carrasco2015. The results given in this paper can be applied without modification to this case, because any separable Hilbert space is isomorphic to $\mathcal H$ (Conway1994). }}}

Example (ref) and Appendix (ref) of the Online Supplement provide examples of function-valued $Z_i$.

example\normalfont Miguel2004 studied the effect of economic growth on the occurrence of civil war in sub-Saharan Africa. The authors suggested addressing the potential endogeneity of economic growth using annual rainfall growth because (i) rainfall variations are exogenous and (ii) a large volume of the economy in this area relies on agriculture. However, agricultural production may be better predicted if the instrument is not simply given by the annual {\it{average}}, but by a {\it{function}} of rainfall variations over the year, considering the importance of daily rainfall variations in agricultural productivity. To detail this idea, let $x_{is} (t) = 1-\dot x_{is} (t)/ \dot x_{is-1} (t)$ where $\dot x_{is}(t)$ is the daily rainfall in country $i$ on the $t$th day of year $s$. For each $i$ and $s$, we can obtain the rainfall growth curve, denoted $Z_{is}$, by smoothing $\{ x_{is} (t) \}_{t=1} ^{365}$.\footnote{\textcolor{black}{For example, one may consider $Z_{is}$ as a continuous function such that $Z_{is}(t/365) = x_{is}(t )$ for $t \in[0,365]$. Because $x_{is}(t)$ is observable only on integer points of $t$, the entire curve on year $s$ needs to be obtained by smoothing the realizations $\{ x_{is} (t) \}_{t=1} ^{365}$. A similar example can be found in HK2012.}} Then, one may consider the following model.\begin{equation*} conflict_{is} = 1\{ \beta_0 + growth_{is}\beta + W_{is}'\gamma \geq u_{is} \},\quadgrowth_{is} = \Pi_{n} Z_{is} + W_{is} ' \pi_2 + V_{is}. \end{equation*} The variable $\text{conflict}_{is} $ is equal to 1 if country $i$ experiences a civil conflict in year $s$, and 0 otherwise. $\text{growth}_{is}$ is economic growth in country $i$ in year $s$. $W_{is}$ is a set of exogenous variables. In this case, the first-stage parameter $\Pi_{n}$ will be represented by a linear map from $\mathcal H$ to $\mathbb R$.

The covariance of $Z_i$, denoted $\mathcal K$, and its sample counterpart, denoted $\mathcal K_n$, are given by \[ \mathcal K = \mathbb E [ Z_i \otimes Z_i ] \quad\text{and}\quad \mathcal K_n =n^{-1} \sum_{i=1} ^n Z_i \otimes Z_i, \] where $\otimes$ signifies the tensor product in $\mathcal H$ satisfying that $h_1 \otimes h_2 (\cdot) = \langle h_1, \cdot \rangle_{\mathcal H} h_2$ for any $ h_1,h_2 \in \mathcal H$. The tensor product is a natural generalization of the outer product in Euclidean spaces, and thus, $\mathcal K$ (resp.\ $\mathcal K_n$) can also be understood as a generalization of the covariance (resp.\ sample covariance) matrix in Section (ref). \textcolor{black}{The integral representation of $\mathcal K$ would be useful for understanding the covariance operator and the tensor product. For any $h \in \mathcal H$, $\mathcal Kh$ is an element in $\mathcal H$ and satisfies that

equation*[equation* omitted — 247 chars of source]

where $\mathtt{k}(t, \tilde t) =\mathbb E[ Z_i(t)Z_i (\tilde t)]$ is called the covariance kernel.} We note that $\mathcal K$ (resp.\ $\mathcal K_n$) allows the spectral decomposition with respect to its eigenvalues and eigenfunctions, denoted $\{ \kappa_j , \varphi_j \}_{j \geq 1 }$ (resp.\ $\{ \widehat \kappa_j , \widehat \varphi_j \}_{j \geq 1 }$). Assume that the eigenvalues are in descending order. Given the consistency of $\mathcal K_n$, $\widehat\kappa_j $ and $\widehat{\varphi}_j $ are consistent estimators of $\kappa_j$ and $\varphi_j$ (Bosq2000).

Estimator

We consider the following moment condition, which is similar to \hyperref[{mod2:eq1:finite}]{\tagform@{\ref*{mod2:eq1:finite}}}.

equation[equation omitted — 166 chars of source]

\textcolor{black}{That is, for any $h\in \mathcal H$, $ \mathbb E[\langle Z_i , h\rangle_{\mathcal H} Y_{2i}] = \Pi_n \mathbb E[\langle Z_i, h \rangle_{\mathcal H} Z_i] = \Pi_n \mathcal K h$.} As discussed by Carrasco2007, the unique solution to \hyperref[{mod2:eq1}]{\tagform@{\ref*{mod2:eq1}}} exists only when $ \mathcal K $ is bijective and its inverse is continuous. These conditions are not likely to be satisfied when $Z_i$ is an $\mathcal H$-valued random variable. Thus, we use a regularized inverse of $\mathcal K$, denoted $\mathcal K_\alpha ^{-1}$, to obtain a solution to \hyperref[{mod2:eq1}]{\tagform@{\ref*{mod2:eq1}}}. The regularized inverse considered here satisfies $\underset{\alpha \to 0}{\lim} \mathcal K_{\alpha} ^{-1} \mathcal K h = h $ for any $h \in \mathcal H $ (see, e.g., Kress1999) and has the following representation:

equation[equation omitted — 138 chars of source]

The conditions on $q(\kappa_j,\alpha)$ will be detailed in Assumption (ref). This representation encompasses various types of regularization schemes, such as (i) Tikhonov ($q(\kappa_j , \alpha) = \kappa_j ^2 /(\kappa_j ^2 +\alpha)$), (ii) ridge ($ q(\kappa_j, \alpha) = \kappa_j /(\kappa_j +\alpha)$),\footnote{\textcolor{black}{Tikhonov regularization is considered as an extension of ridge regularization to infinite dimensions. Specifically, as detailed in Carrasco2007, Tikhonov regularization is related to the method of moments estimation, whereas ridge regularization estimates parameters that minimize the residual sum of squares. That is, the two strategies are associated with different estimation methods, although they are asymptotically equivalent as $\alpha \to 0$. Hence, different conditions are required to use them in the analysis.}} and (iii) spectral cut-off ($q(\kappa_j , \alpha) = 1\{ \kappa_j ^2 \geq \alpha \}$). The sample counterpart of $\mathcal K_\alpha ^{-1}$, denoted $ \mathcal K_{n\alpha} ^{-1}$, is similarly defined by replacing $\{ \kappa_j , \varphi_j \}_{j \geq 1}$ in \hyperref[{kinv}]{\tagform@{\ref*{kinv}}} with $\{ \widehat\kappa_j , \widehat\varphi_j \}_{j \geq 1}$, i.e., $ \mathcal K_{n\alpha} ^{-1} = \sum_{j=1} ^n \widehat \kappa_j ^{-1} q(\widehat \kappa_j , \alpha) \widehat \varphi_j \otimes \widehat \varphi_j$. Then, for $\alpha>0$, the solution to \hyperref[{mod2:eq1}]{\tagform@{\ref*{mod2:eq1}}} (denoted $\Pi_\alpha$) and its sample counterpart (denoted $\widehat \Pi_{\alpha}$) are given as follows.

equation[equation omitted — 293 chars of source]

Although $\Pi_\alpha $ and $\widehat{\Pi}_\alpha$ in \hyperref[{eq:first-est}]{\tagform@{\ref*{eq:first-est}}} act on a Hilbert space of infinite dimension, the predicted value of $Y_{2i} $ and the first-stage residual, which are computed using $\widehat\Pi_\alpha$ in \hyperref[{eq:first-est}]{\tagform@{\ref*{eq:first-est}}}, are still $d_e$-dimensional. Therefore, we can obtain our estimators by solving \hyperref[{model:5eq}]{\tagform@{\ref*{model:5eq}}} as described in Section (ref).

Asymptotic Properties

assumption\begin{enumerate*}[(i)] • \textcolor{black}{Assumptions (ref), (ref) and (ref) hold for a bounded linear operator $\Pi_0 : \mathcal H \hspace{-0.2em}\to\hspace{-0.2em} \mathbb R^{d_e}$ and with $ \Vert Z_i \Vert $ being replaced by $\Vert Z_i \Vert_{\mathcal H}$.} • For $\{\pi_{\ell , 0} \hspace{-0.2em}=\hspace{-0.2em} \Pi_{0} ^\ast e_{\ell}\}_{\ell=1} ^{d_e}$, there is $\rho \geq 1$ such that $ \sum_{j=1} ^\infty \langle \varphi_j , \pi_{\ell,0} \rangle_{\mathcal H} ^2 /\kappa_{j} ^{2\rho } < \infty$, where $\Pi_0 ^\ast$ is the adjoint of $\Pi_0$. \end{enumerate*}
assumption$\mathcal K_{\alpha} ^{-1}$ allows the representation in \hyperref[{kinv}]{\tagform@{\ref*{kinv}}}. For $\kappa \geq 0$ and $\alpha >0$, the function $q (\kappa , \alpha ) $ in \hyperref[{kinv}]{\tagform@{\ref*{kinv}}} satisfies $ q (\kappa,\alpha)\in [0,1]$, $ \alpha q (\kappa , \alpha ) \leq c \kappa $ and $\sup_{\kappa}\kappa^{2\tilde\rho} (q(\kappa , \alpha)-1) \leq$ $ c\alpha^{\min \{\tilde \rho,1\} } $ for some $\tilde \rho >0$ and $c>0$. Moreover, $\underset{\alpha \to 0}{\lim} q(\kappa , \alpha) = 1$ if $\kappa > 0$.

Assumption (ref).(ref) is equivalent to the former conditions, except that the first-stage parameter and the norm of $Z_i$ are adapted to the general case. \textcolor{black}{The vector $f_i$ can be viewed as the (infeasible) optimal instrument that summarizes the information from infinite-dimensional instrument $Z_i$.} \textcolor{black}{The $d_e \times d_e$ matrix $\Lambda_n$ determines the identification strength of $Z_i$ in the direction to each $\pi_{\ell,0}$. As an example, suppose that

equation*[equation* omitted — 376 chars of source]

Let $[ a]_{i}$ denote the $i$th row of vector $a$. Then, the first stage can be written by

equation*[equation* omitted — 269 chars of source]

}\textcolor{black}{Thus, the reduced-form coefficient of the component of $Z_i$ with respect to $\pi_{1,0}$, $\langle Z_i, \pi_{1,0}\rangle_{\mathcal H}$, is strongly identified. In contrast, $[\beta]_2$ is strongly identified if $\mu_{2,n}=\sqrt{n}$, while weakly identified otherwise. Thus, $\Lambda_n$ allows for reduced-form coefficients to be associated with different identification powers. For further discussion when $\mathcal H = \mathbb R^{d_z}$, refer to Chao_et_al.}

Assumption (ref).(ref) is needed to identify $\Pi_0$ without the injectivity of $ \mathcal K$; see Carrasco2007 and Benatia2017. As briefly mentioned in Section (ref), this assumption restricts the relative decaying rate of $\mathcal K$'s eigenvalues with respect to its Fourier coefficients $\{\langle\pi_{\ell,0},\varphi_j\rangle_{\mathcal H}\}_{j \geq 1} $. It is particularly essential for proving that $n^{-1}\sum_{i=1} ^n \Vert \widehat g_{i,\alpha} - g_i \Vert ^2 =o_p(1)$, which plays an important role in obtaining our main results. Assumption (ref) describes the conditions on $\mathcal K_{\alpha} ^{-1}$. Many popular regularization methods, such as Tikhonov and spectral cut-off, satisfy the conditions (see Carrasco2007). Therefore, Assumption (ref) is not restrictive in practice.

The following theorem states the consistency of $\widehat{\theta}_j$ in the general case.

theorem\normalfont Suppose that $\alpha \hspace{-0.2em}\to\hspace{-0.2em} 0$ and $\mu_{m,n} ^2 \alpha\hspace{-0.2em}\to\hspace{-0.2em} \infty$ as $n \hspace{-0.2em}\to\hspace{-0.2em} \infty$, and Assumptions (ref) to (ref) are satisfied. Then, $ \widehat \theta_{ j} - \theta_{0} \overset{p}\rightarrow 0 $ as $n \hspace{-0.2em}\to\hspace{-0.2em} \infty$ for $j \in\{ M,N \}$.

\textcolor{black}{The condition on $\mu_{m,n}$ and $\alpha $ in Theorem (ref) requires a slower convergence rate of $\alpha$ compared to that in Theorem (ref). This stronger condition is required mainly due to the additional conditions on the regularized inverse in Assumption (ref).} In fact, some regularization methods, such as Tikhonov, spectral cut-off, and Landweber Fridman, satisfy $\alpha^{1/2} q(\kappa,\alpha)\leq c \kappa$ for some $c >0$, and for them, the condition on $\mu_{m,n}^2 \alpha$ in Theorem (ref) can be replaced by its weaker counterpart in Theorem (ref).

Next, we discuss the asymptotic distributions of the proposed estimators. As discussed in Section (ref), the asymptotic distributions cannot be straightforwardly obtained using the approaches in rivers1988limited, Carrasco2012, and Carrasco2015a,Carrasco2015; this is not only because of some distinctive properties of our model which are described in the previous section, but also because of the fact that $\widehat{\gamma}_{i,\alpha}$ and $\widehat{V}_{i,\alpha}$ now depend on an infinite-dimensional parameter estimator, $\widehat \Pi_\alpha$. This necessitates an approach to asymptotic analysis that differs from those in previous articles. Our setting is similar to the one studied by Chen2003, in which the criterion function contains an infinite-dimensional parameter estimate. Chen2003 discuss the primitive conditions for the asymptotic normality of an estimator obtained in such a setting. Here, we make use of their approach and study the asymptotic distributions of our estimators.

To this end, we employ additional conditions. We first introduce a condition on the integrabiliy of the covering number of $\mathcal H$ with respect to $ \Vert\cdot\Vert_{\mathcal H} $ (denoted $\mathfrak N (\epsilon,\mathcal H, \Vert \cdot\Vert_{\mathcal H} ) $), i.e., $\mathfrak N (\epsilon,\mathcal H, \Vert \cdot \Vert_{\mathcal H} ) = \inf \{ n \in \mathbb N : \mathcal H \subset \bigcup_{k=1} ^n \mathcal B_k (\epsilon ) \}$ where $\mathcal B_k (\epsilon ) = \{ h \in \mathcal H : \Vert h - h_k \Vert_{\mathcal H} < \epsilon , h_k \in \mathcal H \}$.

assumption$ \mathcal H$ satisfies $\int_0 ^\infty \sqrt{\log \mathfrak N (\epsilon , \mathcal H , \| \cdot \|_{\mathcal H} )} d\epsilon <\infty$.

In Chen2003, the covering number of an infinite-dimensional parameter space should satisfy the integrability condition similar to Assumption (ref). In our setting, the parameter space is given by the space of bounded linear operators from $\mathcal H$ to $\mathbb R^{d_e}$. Thus, it is not straightforward to verify \citepos{Chen2003} condition. In contrast, Assumption (ref) pertains only to the domain of our infinite-dimensional parameter, making it easier to verify. In the current paper, the restriction on the domain of $\Pi_0$ is sufficient, because its rank is finite (see Lemma (ref) in the Online Supplement). This assumption is likely to hold in many practical choices of $\mathcal H$; for example, it holds if $Z_i$ is finite-dimensional (as discussed in Section (ref)), or \textcolor{black}{is smooth enough and takes values in the space of functions with many finite derivatives}, see, e.g., Vaart1996.

assumption\begin{enumerate*}[(i)] • Assumption (ref) holds; • For $\kappa \geq 0$ and $\alpha >0$, $q(\kappa, \alpha)$ in \hyperref[{kinv}]{\tagform@{\ref*{kinv}}} satisfies $ \alpha^{1/2} q (\kappa , \alpha ) \leq c\kappa $ for some constant $c>0$; • Assumption (ref).(ref) holds with $\rho \geq 3/2$. \end{enumerate*}

Assumption (ref) is needed to address the bias discussed in Remark (ref). Assumption (ref).(ref) implies that not all regularization methods are applicable in our setting. For example, ridge regularization may fail to satisfy the condition since its conservative upper bound of $ q(\kappa,\alpha)$ is $\kappa O( \alpha ^{-1})$. On the other hand, other aforementioned methods, such as Tikhonov and spectral cut-off, satisfy the condition. Hence, in spite of its popularity, ridge regularization is less preferred in this paper.

Assumption (ref).(ref) can be relaxed if $\operatorname{ran} \mathcal V_j ^\ast$, the range of the adjoint operator of $ \mathcal V_j $, belongs to $ \operatorname{ran} \mathcal K$, the range of $\mathcal K$, where $\mathcal V_j $ is defined by $ \mathbb E [ \dot m_{2j} \left( \theta_0, g_i\right) Z_i \otimes g_{0i} ]$. Appendix (ref) in the Online Supplement shows that the closure of $\operatorname{ran} \mathcal V_j ^\ast$, denoted $ \operatorname{cl} (\operatorname{ran} \mathcal V_{j}^\ast)$, is always a subspace of the closure of $\operatorname{ran} \mathcal K$, denoted $\operatorname{cl} (\operatorname{ran} \mathcal K)$. If $Z_i$ is finite-dimensional or can be represented by a finite number of basis functions, $\operatorname{cl}(\operatorname{ran} \mathcal V_{j}^\ast) = \operatorname{ran} \mathcal V_j ^\ast $ (resp.\ $ \operatorname{cl} (\operatorname{ran} \mathcal K)=\operatorname{ran} \mathcal K$). In this case, Assumption (ref).(ref) will be satisfied. Hence, the assumption is not restrictive as it appears.

In this section, $\mathcal C_j$ is a linear operator from $\mathcal H$ to $\mathbb R^{d_e}$ satisfying $ \mathcal K^{1/2}\mathcal C_j ^\ast = \mathcal V_{j} ^\ast $, and the $(\ell, k)$th element of $\mathcal J_{2,j}$ is given by $\sigma_{\psi_0} ^2 \langle \mathcal C_j ^{\ast} e_{\ell}, \mathcal C_{j}^{\ast}e_k\rangle_{\mathcal H}$. The operator $\mathcal C_j$ is unique and bounded, even when $\mathcal H$ is a Hilbert space of infinite dimension (baker1973joint).

theorem\normalfont Suppose that $\alpha \hspace{-0.2em}\to\hspace{-0.2em} 0$, $\mu_{m,n} ^2 \alpha\hspace{-0.2em}\to\hspace{-0.2em}\infty$ and $\sqrt{n} \alpha\hspace{-0.2em}\to\hspace{-0.2em}0$ as $n \hspace{-0.2em}\to\hspace{-0.2em}\infty$, and Assumptions (ref) to (ref) are satisfied. Then, $ \mathcal W_{0j} ^{-1/2} \sqrt{n} \mathcal S_n ^\prime \left(\widehat \theta_j - \theta_{0} \right) \overset{d}\rightarrow \mathcal N \left( 0, \mathcal I_{2d_e} \right)$ for $j \in \{M,N\}$, where $ \mathcal W_{0j} = (\mathcal S_n ^{-1}\Gamma_{1,0j}\mathcal S_{n}^{\prime -1} ) ^{-1} \mathcal J_j (\mathcal S_n ^{-1}\Gamma_{1,0j} \mathcal S_n ^{\prime -1} ) ^{-1}$.

The asymptotic covariance $\mathcal W_{0j}$ is nonsingular irrespective of the dimension of $Z_i$. Thus, we can conduct inference on $\theta_0$ and $\varsigma'\theta_0$ as discussed in Remarks (ref) and (ref).

Estimation of Asymptotic Variances and Average Structural Functions

\textcolor{black}{ Remark (ref) tells us that the Wald testing procedure enables us to conduct inference on $\theta_0$ without prior knowledge of $\mathcal S_n$, as long as $\widetilde{W}_{j}$, a consistent estimator of $\mathcal{W}_{0j}$ after proper normalization, is available. $\widetilde{W}_j$ is an essential input for the testing procedure. In this section, we discuss its example.}

\textcolor{black}{Let $\widehat \Gamma_{1,j} = n^{-1}\sum_{i=1} ^n \dot m_{2j} ( \widehat \theta_j,\widehat g_{i,\alpha} ) \widehat g_{i,\alpha} \widehat g_{i,\alpha} '$, $\widehat {\mathcal J}_{1,j} = n^{-1}\sum_{i=1} ^n m_{1j} ^2 (\widehat \theta_j, \widehat g_{i,\alpha} ) \widehat g_{i,\alpha} \widehat g_{i,\alpha} '$, and $\widehat {\mathcal J}_{2,j} = \widehat {\sigma}_{j}^2 \widehat {\mathcal V}_{jn} \mathcal K_{n\alpha} ^{-1}\mathcal K_n \mathcal K_{n\alpha} ^{-1} \widehat {\mathcal V}_{jn} ^\ast$ where $ \widehat {\sigma}_{j}^2 = n^{-1}\sum_{i=1} ^n (\widehat \psi_j ' \widehat V_{i,\alpha} )^2 $ and $ \widehat {\mathcal V}_{jn} =n^{-1} \sum_{i=1} ^n \dot m_{2j} ( \widehat \theta_j , \widehat g_{i,\alpha } ) Z_i \otimes \widehat g_{i,\alpha} $. The theorem below summarizes Lemmas (ref)--(ref) in the Online Supplement. }

theorem\normalfont Let $\widehat {\mathcal W}_{j } = \widehat \Gamma_{1,j} ^{-1} (\widehat {\mathcal J}_{1,j} + \widehat {\mathcal J}_{2,j}) \widehat \Gamma_{1,j} ^{-1} $ and suppose that the conditions in Theorem (ref) hold. Then, for each $j$, $ \Vert \mathcal S_n ^\prime\widehat {\mathcal W}_{j}\mathcal S_n- \mathcal W_{0j} \Vert_{\text{HS}} = o_p(1) $, \textcolor{black}{where $\Vert A \Vert_{\text{HS}}= (\operatorname{tr}(A'A))^{1/2}$ for a matrix $A$.}
remark\normalfont We note that interest often lies in how the outcome probability changes according to a change in $Y_2$. In practice, this is often measured by two estimates, the average structural function (ASF) and the average partial effect (APE), each of which is given by \[{\text{ASF}} (y_{2}) := \frac{1}{n} \sum_{i=1} ^n\Phi ( y_{2} ' \widehat \beta + \widehat V_{i,\alpha} ' \widehat \psi )\quad\text{and}\quad\text{APE} (y_2) := \frac{1}{n} \sum_{i=1} ^n \phi (y_{2} ' \widehat \beta + \widehat V_{i,\alpha} ' \widehat \psi ) \widehat{\beta} ,\] see Wooldridge2010, Blundell_Powell2004 and Rothe2009. The ASF (resp.\ APE) is a consistent estimator of $\mathbb E_{V}[\mathbb P(y=1 | Y_2 = y_2, V)]$ (resp.\ $\mathbb E_{V}[ (\partial\mathbb P(y = 1 | Y_2, V)/\partial Y_2 )_{Y_2 = y_2} ]$).

Mean Squared Error and the Choice of $\alpha$

A practical challenge in implementing our estimation procedure is the selection of the regularization parameter. A theoretically grounded approach would be to choose $\alpha$ in such a way as to minimize the conditional mean squared error (MSE), defined by $\mathbb E[ \Vert \widehat{\theta}_j -\theta_0 \Vert^2 |x ]$, as in donald2001choosing and Carrasco2012. However, our analysis of the MSE is more complicated than theirs because of the nonlinearity and the use of the estimated control function whose asymptotic properties rely on a possibly infinite-dimensional random element. In particular, we need more conditions to control for the linearization error. Thus, to reduce the complexity, we consider the conditional MSE of an alternative estimator $\overline\theta$, defined as the solution to a linear problem. In this section, the subscript $j$ indicating estimators is suppressed to simplify the explanation. \textcolor{black}{In the simple case considered in Section (ref), $\overline\theta$ is given by the solution to the following:

equation*[equation* omitted — 348 chars of source]

where $\pi$, $\widehat{\pi}_\alpha$, and $\pi_n$ denote the vectorizations of $\Pi'$, $\widehat\Pi_{\alpha}'$ and $\Pi_{n}'$, respectively.} \textcolor{black}{We defer its formal definition for the general case to \hyperref[{thetaoverline: def}]{\tagform@{\ref*{thetaoverline: def}}} in Appendix (ref) of the Online Supplement, as it involves additional notations.} In our setting, the asymptotic distribution of $ \overline\theta$ is equivalent to that of $ \widehat{\theta} $. Hence, studying $\overline \theta$ would be enough for the purpose of discussing a way to choose $\alpha$.

The proposition below summarizes an upper bound of $\overline\theta$'s conditional MSE; we can provide more details if $\mu_{jn}$'s are known. While the following result may look conservative, it is still useful in practice due to two reasons: (i) the unknown nature of $\{\mu_{jn}\}_{j=1} ^{2d_e}$ and (ii) the fact that the rate below represents the outcome when there is at least one strong instrument.

proposition\normalfont \textcolor{black}{Suppose that the conditions in Theorem (ref) hold.} Then, $\mathbb E[ \Vert \mathcal S_n ' (\overline\theta - \theta_0)\Vert ^2 |x]= O_p(n^{-1}\alpha^{-1/2} + \alpha^2).$

The MSE of $\overline \theta$ turns out to be minimized where the MSE of $\widehat{\Pi}_\alpha$ is minimized. Due to this, the above convergence rate is similar to that in the functional linear model studied by Benatia2017, who suggested choosing $\alpha$ to minimize the conditional MSE. In our setting, the optimal $\alpha$ satisfies $\alpha \sim n^{-2/5}$. However, this rate may not be preferred, because it does not ensure the asymptotic normality of our estimators. Our asymptotic normality results require a fast convergence of $\widehat{\Pi}_n$'s bias, even at the cost of a larger variance; if we choose $\alpha$ as in Benatia2017, the limiting distribution of $\widehat{\theta}_j$ will depend on the regularization bias of the first-stage estimator, which may not be desirable in practice. Hence, a slower convergence of $\overline \theta$ (and that of $\widehat \theta_j$), resulting from not choosing the optimal rate of $\alpha$, should be understood as the cost of implementing inference with a possibly infinite-dimensional instrumental variable. In spite of its limitation, the above criterion still provides us with a practical guideline to select the regularization parameter. We defer its detailed discussion to Appendix (ref) in the Online Supplement.

Simulation Study

In this section, we study the finite sample performance of our estimators via Monte Carlo simulations when $Z_i \in \mathbb R^{K}$. The simulation results for the general case is separately discussed in Appendix (ref) of the Online Supplement.

Experiment 1: Gaussian Instrument

We consider the following data generating process (DGP).

equation[equation omitted — 426 chars of source]

where $\beta_1 = 1$, $\beta_2 = -1$, $\rho = 0.6$, and $\sigma_1 ^2 = 1 / (1-\rho ^2 )$. Thus, $\eta_i = u_i +\rho \sigma_2 ^{-1}\sigma_1 v_i \sim_\text{iid}\mathcal N (0,1 )$. The variable $z_{1i}$ is the first element of $Z_{i}\in \mathbb R^K$, and $Z_i \sim_{\text{iid}} \mathcal N (0, \Sigma_Z)$ where $\Sigma_Z = [ \sigma_z ^2 \rho_z ^{ |i-j |} ]$, $\sigma_z ^2 = 0.5$ and $\rho_z = 0.7$. We let $\sigma_2^2 = 1-\pi ' \Sigma_Z \pi$ so that the unconditional variance of $y_{2i}$ becomes 1.

We set $\pi = c_\ast (\iota_{\lfloor s K \rfloor} ', 0_{K-\lfloor s K \rfloor})'$, where $\lfloor \cdot \rfloor$ is the floor function. The parameter $s$ determines the sparsity of $\pi$. We consider sparse ($s=0.2$) and dense ($s=0.8$) cases. We set $K$ to $ 50$. The values of $c_\ast $ are chosen to have specific values of the concentration parameter, $\mu^2 = n \pi ' \Sigma_Z \pi/ (1-\pi ' \Sigma_Z \pi )$, as in Belloni. We consider two values of $\mu^2$: 30 and 60. Under the design, the infeasible first-stage F test statistic, given by $\mu^2 / \lfloor s K \rfloor$, ranges from 0.375 to 6. \textcolor{black}{The concentration parameter may not be a perfect measure of instrument strength for binary response models. However, considering its relevance in indicating the weakness of instruments in linear models, standard approaches to estimating endogenous binary response models, such as that of rivers1988limited, may not produce reliable estimation results when $\mu^2$ is small. A similar scenario with small $\mu^2$ has been considered by Magnusson2010 and Dufour2018 who study weak instruments in endogenous binary response models. Later in the section, we experiment with different values of $\mu^2$, and the results confirm the distortion of an existing estimator, similar to the TSLS, when $\mu^2$ is small. Hence, this simulation setting would be enough to study finite sample performance of our estimators when the existing estimation approach does not work well.}

We consider two RCMLEs computed with different regularization methods: the TRCMLE (Tikhonov) and SCRCMLE (spectral cut-off).\footnote{The results from the RNLSE are similar to those from the RCMLE. Hence, they are omitted.} Their regularization parameters are chosen as described in Appendix (ref) in the Online Supplement with $\alpha$ being selected from 25 equally spaced points between $\mathtt{c}_a n^{-0.6}0.001$ and $\mathtt{c}_a n^{-0.6}$.\footnote{Specifically, we compute \hyperref[{eq: mse: tmp}]{\tagform@{\ref*{eq: mse: tmp}}} with Generalized cross-validation in Carrasco2012 for the RNLSE.} The constant $\mathtt{c}_a$ is set to $\overline{\mathtt{c}}_a \max\{ 0.1, 1/\delta \}$ where $\delta $ is the first-stage F test. The constant $\overline{\mathtt{c}}_a$ is given by $\operatorname{tr}(\Sigma_Z ' \Sigma_Z)^{1/2} $ for the TRCMLE and its square for the SCRCMLE. This is designed to have a larger regularization parameter when a smaller concentration parameter is given. We compare the performance of our estimators with four alternatives: \citepos{rivers1988limited} 2SCMLE, the Inf.2SCMLE (to be detailed shortly), the naive probit estimator, and the Tikhonov-regularized TSLS estimator (TTSLS). Similar to the TSLS, the 2SCMLE is expected to be biased as $K$ increases. To separate any effect of this bias from that induced by weak instruments, we compute the 2SCMLE using instruments with non-zero first-stage coefficients only. This modified estimator is called the Inf.2SCMLE. If there is no distortion from weak instruments, the infeasible 2SCMLE is likely to perform well when $\mu^2 = 30$ and $s=0.3$. In contrast, if the estimator is not affected by the many instruments bias, it will perform well when $\mu^2 = 60$ and $s=0.8$. The TTSLS is similar to the estimators proposed by Carrasco2012 and Carrasco2015a, Carrasco2015, except for the use of the first-stage residual as an additional regressor in the second stage.\footnote{The variance is computed with the heteroscedasticity-robust covariance estimator $\text{HC}_1$ in mackinnon1985. We use the regularization parameter chosen for the TRCMLE.} This estimator is considered because of the popularity of linear probability models in empirical studies. As its true parameter is different from that of all the other estimators, we evaluate its performance at $\mathbb E [\phi (g_i ' \theta ) ] \beta_1$ so that a reasonable comparison can be made (see cameron2005microeconometrics). For each estimator, we compute the median bias (Med.Bias) and median absolute deviation (MAD). We also test $H_0: \beta_1 = 1$ at 5% significance level as explained in Remark (ref). The rejection rate is reported in columns labeled “RP” in Table (ref).

table[table omitted — 2,319 chars of source]

In Table (ref), our estimators exhibit better performance compared to the 2SCMLE and Inf.2SCMLE. Specifically, the 2SCMLE is severely biased and does not have reasonable rejection rates. In contrast, our estimators have smaller median bias and correct size in most cases considered in the table. Another interesting observation is that the SCRCMLE outperforms the TRCMLE in the dense design, although the difference diminishes with a larger sample size. The better performance of the SCRCMLE in the small sample may be related to the relative stability of the regularized inverse when the signal is dense. Similar results can be found in seong2021.

The Inf.2SCMLE does not perform well when the signal is either dense ($s=0.8$) or weak ($\mu^2= 30$) in Table (ref). The distortion in the dense design may be related to the many instruments bias; the estimator uses 40 instruments when the signal is dense. Meanwhile, the distortion when $\mu^2= 30$ casts doubt on the reliability of \citepos{rivers1988limited} 2SCMLE when the concentration parameter is small. This is consistent with the ASF estimates reported in Figure (ref). In the figure, the ASFs of the 2SCMLE shift toward those of the probit estimator as $\mu^2$ decreases.

In Table (ref), the TTSLS and the TRCMLE produce similar estimation results. However, the TTSLS always exhibits rejection rates above the nominal level. Furthermore, as reported in Figure (ref), the TTSLS does not provide good ASF estimates, especially at the boundaries of the support of $y_2$, whereas the TRCMLE yields estimates close to the true values even at the boundaries.

Next, we examine the performance of the estimators under the same DGP but with different values of $K$, $\mu^2$ and $s$. In the first experiment, we change $K$ from 10 to 100 while keeping $s$ and $\mu^2$ fixed at $ 0.5$ and $ 50$, respectively. Secondly, we set $K=50$ and $\mu^2 = 50$, and change $s$ from 0.1 to 1. Lastly, we fix $K= 50$ and $s = 0.5$, and change $\mu^2$ from 25 to 250. The sample size is 200. We focus on three estimators: the TRCMLE, 2SCMLE, and probit estimator.

The estimation results are reported in Figure (ref). The TRCMLE appears to have the smallest median bias and correct size in all the considered cases. On the other hand, as $K$ increases, the 2SCMLE and the probit estimator tend to perform similarly. This may not be surprising, since in the linear model, it is known that the TSLS shifts toward the OLS estimator as $K$ increases. \textcolor{black}{A similar convergence is observed as $\mu^2$ decreases, which suggests that \citepos{rivers1988limited} estimator may not be reliable when the concentration parameter is small.}

Factor Model

Here, we consider the case where only $\widetilde Z_i \in \mathbb R^{\tilde{K}}$ is observed rather than the true instrument $Z_i \in\mathbb R^K$. Let $M$ be a $ \widetilde K \times K$ matrix consisting of elements that are randomly drawn from $\text{Unif}[-1,1]$. This matrix is fixed across simulations, and let $\widetilde Z_i =M Z_i + \widetilde V_i$ where $\widetilde V_i \sim_{\text{iid}}\mathcal N (0, \widetilde\sigma ^2 \mathcal I_{\widetilde K} )$ and $\widetilde\sigma = 0.3$. The measurement error $\widetilde V_i$ is independent of $u_i$ and $V_i$. The parameters are set as follows: $K$=5, $\widetilde{K}=100$, $s=1$, $\rho_z = 0$, and $\sigma_z = 1$. We consider two different values of $\mu^2$: 30 and 60.

table[table omitted — 1,487 chars of source]

The Inf.2SCMLE is computed with $Z_i$, while the others are computed using $\widetilde Z_i$. The simulation results are reported in Table (ref) (the estimated ASFs are similar to those in Figure (ref) and reported in the Online Supplement). The results closely align with those in Section (ref); our estimators exhibit considerably smaller median bias and reasonable rejection rates compared to the others. The SCRCMLE tends to have a larger bias but a smaller MAD with a larger sample size, suggesting a potential bias-variance tradeoff. The Inf.2SCMLE, computed with $Z_i$, performs considerably well, whereas the 2SCMLE exhibits a substantially large median bias and poor size control.

Empirical Example: Bastian2018

Recent studies on the EITC have shown that it has a considerable impact on its recipients (see meyer2001welfare; Eissa2004) and their children (see Dahl2012). In particular, Bastian2018 provided empirical evidence that increasing EITC exposure positively impacts family earnings, consequently leading to better long-term academic achievement for children. We revisit their work and complement it as follows: First, we reexamine the weakness of their instruments. The first-stage F statistics reported in Bastian2018 range between 3.2 and 5.1, suggesting that their instruments are potentially weak. However, this measure may not be perfect considering the nonlinear nature of the binary response model. Thus, we use the test proposed by frazier2020weak to study their weakness. Second, we utilize additional instruments to increase estimation efficiency and account for heterogeneous effects of EITC exposure on family income depending on the state of residence. Lastly, we employ a parametric binary response model instead of the linear model; as detailed in Example (ref), this approach is expected to produce better estimation results from a theoretical perspective.

We use \citepos{Bastian2018} data from the 1968-2013 waves of the Panel Study of Income Dynamics. Our model is similar to theirs and is given as follows:

equation[equation omitted — 225 chars of source]

where $I_i = ( I_{i,\text{(0-5)}} , I_{i,\text{(6-12)}} , I_{i,\text{(13-18)}})'$ and each $I_{i,(a)}$ is the family income of individual $i$ at age interval $a$. The outcome variable $y_i$ is 1 if individual $i$ is a college graduate and 0 otherwise.\footnote{College graduation is assessed when individuals reach the age of 26.} $W_{1i}$ is an 11-dimensional vector of personal characteristics: age, age square, the number of siblings at age 18, and indicators for black, Hispanic, female, ever-married parents, and whether the individual's mother and father completed high school or are at least some college-educated. $W_{2i}$, which is measured at age 18, is state-by-year economic indicators: per capita GDP, the unemployment rate, the top marginal income tax rate, the minimum wage, maximum welfare benefits, spending on higher education, and tax revenue. All the continuous variables are standardized before analysis.

We consider two models with different instruments. In Model 1, we follow Bastian2018 and let $\widetilde Z_i$ be $(\text{EITC}_{i,(\text{0-5})}, \text{EITC}_{i,(\text{6-12})}, \text{EITC}_{i,(\text{13-18})})'$, where $\text{EITC}_{i,(a)}$ is the standardized measure of EITC exposure of individual $i$ at age interval $a$.\footnote{The measures of EITC exposure and family income are in thousands of 2013 dollars and discounted at a 3% annual rate from age 18. A detailed description can be found in Bastian2018.} As EITC exposure is measured in childhood and adolescent years, this could affect children's long-term educational attainment only through $I_i$; see Bastian2018 for more discussion on its validity. In Model 2, we use 144 instruments, consisting of the measures of EITC exposure and their interactions with 47 state-of-residence dummies (observed in the sample). These additional instruments are expected to help us address the potential bias associated with the heterogeneous effects of EITC transfers on family income caused by, such as, the state-specific cost of living. The sample size is 2,654.

We first apply \citepos{frazier2020weak} distorted J test to study if our instruments are weak in the sense of Staiger1997. This is needed to ensure the conditions in Theorem (ref). However, their test is not designed for the case where (i) the (sample) covariance of instruments is nearly singular and (ii) multiple endogenous variables are present. Hence, we apply the test only to Model 1 separately for each age interval $a$, after replacing $I_i$ in \hyperref[{model:eitc}]{\tagform@{\ref*{model:eitc}}} with $I_{i,(a)}$. The computed distorted J tests are 37.15, 21.81, and 42.76 for the age intervals 0-5, 6-12, and 13-18, respectively. The critical value at 5% significance level is around $ 9.48 $.\footnote{To conduct the test, $\delta_n$, $a_i$ and $b_i$ in frazier2020weak are set to $\widehat{\psi}/\log(\log(n))$, $( I_{i,(a)}, W_{1i}', W_{2i} ', \widetilde Z_i' , 0_{k}')'$, and $(0_{k}',1, W_{1i}', W_{2i}', \widetilde Z_i')'$ where $k = \dim(W_{1i})+\dim(W_{2i})+\dim(\widetilde Z_{i})+1=23$.} Therefore, the tests suggest that our model is not weakly identified in the sense of Staiger1997.

We consider the TRCMLE and the SCRCMLE. Their regularization parameters are chosen as in Section (ref). Specifically, $\alpha$ is chosen from 25 equally spaced points between $ \overline{\texttt{c}}_a n^{-0.6} 10^{-4}$ and $\overline{\texttt{c}}_a n^{-0.6} 10^{-1}$. We also consider \citepos{rivers1988limited} 2SCMLE and the probit estimator. In this section, we further consider another estimator similar to the post-Lasso estimator proposed by Belloni. This estimator is included due to its popularity in empirical research. To compute it, we first select instruments with non-zero first-stage coefficients using the Lasso procedure. Then, we refit the first stage with the selected instruments and compute the first-stage fitted values and residuals. Lastly, a linear model is estimated using the fitted values and residuals as regressors. The results are reported in the column labeled “Lasso” in Table (ref).

table[table omitted — 2,199 chars of source]

In Table (ref), our and \citepos{rivers1988limited} approaches yield similar estimates in Model 1. However, substantial discrepancies arise in Model 2. While our estimators exhibit similar coefficient estimates in both models, \citepos{rivers1988limited} estimates tend to converge toward the probit estimates in Model 2. This is consistent with the findings in Section (ref). Moreover, all the instrumental variable estimators report smaller standard errors in Model 2, suggesting that the large set of instruments in Model 2 contains additional information useful for predicting family income.

Similar observations can be found in Figure (ref), which reports the estimated probability of college completion depending on family income for each age interval. For clarity, let $\mu_{{a}}$ (resp.\ $ \sigma_{{a}}$) denote the sample mean (resp.\ standard deviation) of $I_{i(a)}$. Figure (ref) shows how the probability changes as $I_{i(a)}$ varies from $\mu_{{a}}-\sigma_{a}$ to $\mu_{{a}}+\sigma_{{a}}$, while other variables remain fixed at their medians. There is no significant change in the probabilities computed with the TRCMLE in both models. In contrast, the probability computed with the 2SCMLE substantially shifts toward that of the probit estimator in Model 2. Similar results are observed in the APE estimates in the Online Supplement.

Next, we focus on the Lasso estimates. As mentioned in Section (ref), parameters in linear probability models are different from those in nonlinear binary response models. Hence, to make a reasonable comparison of the coefficient estimates, we compute the APEs at the sample mean of $Y_{2i}$ and report them in Table (ref) in Appendix (ref) of the Online Supplement. However, even in the table, Lasso estimates tend to be biased toward probit estimates in Model 2. This bias may be attributed to factors discussed in Section (ref) and to a slower decreasing rate of the eigenvalues of $\mathcal K_n$.

Table (ref) reports the exogeneity testing results conducted as detailed in Remark (ref). The $p$-values are computed from the $\chi^2(3)$ distribution. The results indicate the presence of endogeneity at the 1% significance level regardless of the number of instruments and estimation methods. This explains the difference in coefficient estimates between the probit estimator and the others.

The estimation results in Table (ref) are similar to those in Bastian2018; increasing family income between ages 13 and 18 positively influences early adulthood college completion. While the authors found a similar effect, their estimates, reported in column 3 of their Table 6, are not significant. This discrepancy may arise from different models and estimation methods, such as the absence of the large number of fixed effects in our approach. Table (ref) also reports a positive (negative) impact of an increase in family income between birth and age 5 (ages 6 and 12) on college completion. Bastian2018 reported similar mixed coefficient signs and noted that these may be attributed to cross-age correlations in EITC exposure.

Conclusion

This paper proposes an estimation procedure to resolve the issue of many (nearly) weak instruments in endogenous binary response models. The proposed estimators are easy to compute and have good asymptotic properties. Monte Carlo studies suggest that our estimators outperform the existing estimators when there are many (nearly) weak instruments. We apply the estimators to study the effect of family income in childhood and adolescent years on college completion.

\onehalfspacing

figure[figure omitted — 387 chars of source]
figure[figure omitted — 1,029 chars of source]
figure[figure omitted — 478 chars of source]
center[center omitted — 103 chars of source]