Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
111,517 characters · 12 sections · 140 citation commands
A Bootstrap Specification Test for Semiparametric Models with Generated Regressors
\thispagestyle{empty}
Keywords: Hypothesis Testing, Bootstrap, Generated Regressors, Semiparametric Model, Control Function, Bias Correction. \\ JEL Classification: C01, C12, C14
\setcounter{page}{2}
Checking the correct specification of a model is empirically relevant, as a misspecified model can yield biased and inconsistent estimates and provide a misleading counterfactual analysis. In this paper we contribute to the literature by providing a specification test for semiparametric models with nonparametrically generated regressors. Such regressors are not observed by the researcher but are nonparametrically identified and estimable. Examples of semiparametric models with generated variables are common in empirical frameworks. They include endogenous models with control function rivers_limited_1988, blundell_endogeneity_2004, newey_nonparametric_1999, semiparametric sample selection models, extension of tobit models escanciano_identification_2016, or semiparametric empirical games with incomplete information aradillas-lopez_pairwise-difference_2012, lewbel_identification_2015. \\ Let $Y\in \mathbb R $, $Z\in \mathbb R ^p$, and $X\in \mathbb R^{p_X} $. Our goal is to test the null hypothesis
against its logical complement $\mathcal H _1=\mathcal H _0 ^c$, where $\nu:\mathbb{R}^{p_X+1}\mapsto\mathbb{R}^{p_\nu}$ and $q:B\times\mathbb{R}^{p_X+1}\mapsto\mathbb{R}^d$ are known vector-valued functions, the $\sigma$-field of $q(\beta,X,H(Z))$ is contained in the $\sigma$-field of $\nu(X,H(Z))$ for all $\beta\in B$, $B$ is a real set, and
is a nonparametric function, with $D\in \mathbb{R}$. The null hypothesis is featured by the presence of the generated variable $H(Z)$. This means that $H(Z)$ is latent, but it is identified and estimable in the first-step nonparametric regression ((ref)). The conditional moment restriction in ((ref)) arises from the above mentioned empirical models, see Section (ref) for details.
Our main contributions are twofold. First, we develop a test for the null hypothesis in ((ref)) featured by the presence of generated variables. Due to the presence of generated variables, the null hypothesis in ((ref)) cannot be tested by using existing methods. Second, we construct and show the validity of a novel wild bootstrap procedure to compute the critical values. Our wild bootstrap procedure (i) uses the information under the null hypothesis and (ii) differently from existing procedures for inference in the presence of generated variables does not require estimating a nonparametric derivative.
Our first main contribution is thus to propose a test for the conditional moment restriction in ((ref)). Such a moment restriction is semiparametric, in the sense that the conditional expectations of $Y$ on both sides of the equation are not restricted to have a specific functional form. If this moment condition did not contain the generated variable $H(Z)$, a specification test could be based on procedures already available in the literature for models where all regressors are observed, see e.g. delgado_significance_2001, xia_goodness--fit_2004. However, since the generated variable $H(Z)$ is not observed, these tests cannot be implemented in the setting considered here.\\ To construct our test statistic we proceed in three steps: in a first step we estimate the generated variable $H(Z)$, in a second step we estimate the parameter $\beta_0$, and in a third step we estimate the right hand side of Equation ((ref)). This allows us to compute the residuals of the semiparametric model. Our test statistic is then based on an { empirical process} involving the estimated residuals, see Section (ref) for details. A specific feature of this setting is that replacing the generated regressor $H(Z)$ with its estimate in the first step introduces an estimation error that impacts on the asymptotic behavior of our test statistic. Hence, it needs to be taken into account. A similar result is obtained in hahn_asymptotic_2013, mammen_semiparametric_2016, and hahn_nonparametric_2018 in an estimation context. Our context however is different from theirs because we have to deal with an empirical process. Moreover, compared to existing specification tests based on empirical processes delgado_significance_2001, xia_goodness--fit_2004, when establishing the asymptotic behavior of our test statistic we are faced with the technical challenge of dealing with an estimation error coming from the generated variables that is not present in existing specification tests.\\ A distinguishing feature of our test is the presence of bias corrections for the nonparametric estimators of the first and third step. Bias corrections have been proposed for the purpose of nonparametric estimation in, e.g., buhlmann_boosting_2003, di_marzio_boosting_2008, and park_l2_2009. We show that in our setting these bias corrections guarantee a { small bias property} of our test statistic. This means that the bias of our test statistic converges to zero faster than the bias of the nonparametric estimators on which it is based, see also newey_twicing_2004. To the best of our knowledge, { this is a novel approach in semiparametric models with generated variables}. The small bias property of our statistic is both theoretically appealing and practically relevant. It is theoretically appealing, as it avoids undersmoothing. This means that the bandwidths “optimal" for estimation can be employed in our specification test. The small bias property is also practically relevant: in our simulation experiment we show that, thanks to the bias corrections and the small bias property, our test is more stable to the selection of smoothing parameters than a test that does not use bias corrections and does not have a small bias property. Alternative approaches developed in chernozhukov_locally_2016 or escanciano_asymptotic_2018-2 could be used in our context to obtain a small bias property of the test statistic. However, as we explain in Section (ref), implementing these approaches in our context would be difficult: { due to the presence of generated variables}, they would require estimating a nonparametric derivative. Since nonparametric derivatives have slow convergence rates, such an estimation would complicate the practical implementation of the specification test. Differently, our bias corrections do not require estimating nonparametric derivatives and are easy to implement in the presence of generated variables. In other words, a further contribution of this paper is to show how to obtain a small bias property in a semiparametric context with generated variables, without estimating nonparametric derivatives.
Our second main contribution is to construct a novel wild bootstrap procedure to compute valid critical values for our test. We show that asymptotically our statistic converges to an intricate distribution depending on unknown features of the data generating process. So, the asymptotic distribution cannot be directly employed to obtain the critical values. This problem also arises in the semiparametric specification tests of, e.g., delgado_significance_2001 and xia_goodness--fit_2004 who develop bootstrap procedures to obtain the critical values. However, their bootstrap procedure are not valid in our context due to the presence of generated regressors. Since their bootstrap procedures are developed for settings { without} generated variables, they cannot replicate the estimation error arising from the generated regressors. Differently, in our setting the generated variables introduce an estimation error that impacts on the asymptotic behavior of the test statistic, and hence it needs to be taken into account. \\ Other procedures are developed in the literature to adjust semiparametric methods for the presence of generated variables, see wooldridge_econometric_2010, hahn_asymptotic_2013, hahn_nonparametric_2018. Such methods consist in estimating the adjustment terms due the generated variables in the influence function representation of the statistic, and then correct the standard errors of the test. Using this approach in our testing problem would be difficult for two reasons. First, it requires the statistic to be asymptotically pivotal, while this property does not hold in our setting. Second, in our framework estimating the adjustment terms coming from the generated variables requires estimating a nonparametric derivative. As highlighted earlier, this is not appealing in practice.\\ Thus, in this paper we construct a novel wild bootstrap procedure for obtaining valid critical values. In particular, our contribution is to develop a wild bootstrap method that does not require estimating a nonparametric derivative and can replicate the estimation error coming from the generated variables. To the best of our knowledge this is a novel contribution in the literature on generated regressors. The bootstrap we develop here imposes the null hypothesis when resampling the observations. This feature is attractive, as it means that we are using all the information available when constructing the critical values for testing. We develop the wild bootstrap test so that also the bootstrapped statistic has a small bias property, as well as the sample statistic.
Our last contribution is to develop an empirical method for the bandwidth selection in our testing problem. The method we propose chooses the bandwidth minimizing the “distance" between the estimated semiparametric model and the null hypothesis. In our simulation experiment, we show that this method provides a reliable inference in moderate samples.\\
Related literature. Beyond the studies already cited, there is a large literature analyzing the problem of estimation with generated variables. Early work includes pagan_econometric_1984, ahn_distribution_1993, ahn_semiparametric_1997. More recent results are presented in li_semiparametric_2002, chen_estimation_2003, rothe_semiparametric_2009, sperlich_note_2009, mammen_nonparametric_2012, gutknecht_testing_2016, mammen_semiparametric_2016, bravo_two-step_2020, hahn_nonparametric_2018, vanhems_estimation_2019. The impact of generated regressors on the asymptotic distribution of a finite dimensional estimator is analyzed in newey_asymptotic_1994 and hahn_asymptotic_2013. escanciano_uniform_2014 obtain an expansion of the residuals from a regression involving variables estimated in a preliminary step. These papers focus on estimation and do not address specification testing in the presence of generated regressors. \\ From a statistical point of view, our work is related to escanciano_uniform_2014 with two important distinctions. First, in this paper we develop a wild bootstrap test, while escanciano_uniform_2014 are not concerned with constructing a bootstrap procedure. Second, the residuals at the basis of our statistic involve bias correction terms which are not present in their context. This allows us to avoid undersmoothing and to obtain a small bias property of our statistic.\\ Finally, our work is related to the literature on specification testing in semiparametric models, see bierens_consistent_1982, bierens_consistent_1990, fan_consistent_1996, bierens_asymptotic_1997, \,stute_nonparametric_1997\,, delgado_consistent_2006, einmahl_specification_2008, lavergne_breaking_2008\,,\, delgado_distribution-free_2008\,,\, escanciano_testing_2010\,,\, neumeyer_estimating_2010, lavergne_significance_2015. We use an approach similar to bierens_asymptotic_1997, stinchcombe_consistent_1998, and delgado_significance_2001. The distinctive feature of our work is the presence of generated variables that introduce extra terms in the asymptotic expansion of our statistic.\\
Organization of the paper. Section (ref) starts by constructing the test statistic. Then, it sets up the assumptions and obtains the asymptotic behavior of our statistic. Section (ref) details the construction of our wild bootstrap procedure and shows its validity for computing the critical values. The main applications of our test are in Section (ref). Section (ref) first describes the practical implementation of our test and our method for the bandwidth selection. Then, it provides evidence about the finite sample behavior of our test and finally shows its application to a real data set. All proofs are relegated to Appendix (ref) and (ref). The supplementary material contains auxiliary results.
To use a more compact notation, we define
Let us assume that $W(\beta)$ is continuously distributed for any $\beta \in B$, and let $f_{W(\beta)}$ be the density function of $W(\beta)$ with respect to the Lebesgue measure. Since the $\sigma$-field of $W(\beta_0)$ is a subset of the $\sigma$-field of $\nu(X,H(Z))$, the null hypothesis in ((ref)) is equivalent to
for some $\beta_0\in B$. We have introduced the density function $f_{W(\beta_0)}$ to avoid a random denominator in the estimation of $G_{W(\beta_0)}$, see Section (ref) for details. Equation ((ref)) is a conditional moment restriction. To test such an equation we consider an { equivalent continuum} of { unconditional} moments. Let us assume that $\nu(X,H(Z))$ is a bounded random variable. Then, the null hypothesis in ((ref)) is equivalent to
where $\mathcal S$ is a set containing a neighborhood of the origin, $\varphi_s(X,H(Z)):=\varphi(s^T \nu(X,H(Z)))$, and $\varphi$ is a suitably chosen weighting function. In Assumption (ref) we report the precise conditions that $\varphi$ must satisfy to guarantee the equivalence between ((ref)) and ((ref)). For example, bierens_consistent_1982 shows such an equivalence with $\varphi(u)=\exp(\,\sqrt{-1}\, u\,)$, bierens_consistent_1990 uses the real exponential $\varphi(u)=\exp(u)$, while stinchcombe_consistent_1998 use more general weighting functions. Now, the formulation of the null hypothesis into a continuum of unconditional moments is quite convenient, as it allows us to construct a test without estimating $\mathbb E \{Y|\nu(X,H(Z))\}$. Differently, constructing a test based on the conditional moments in ((ref)) or ((ref)) would require estimating $\mathbb E \{Y|\nu(X,H(Z))\}$. From a practical point of view this would complicate the implementation of the test, as we should select additional smoothing parameters, say bandwidths. Moreover, in the presence of an important dimension of $\nu(X,H(Z))$ such a test would suffer from a curse of dimensionality. \\ Let us assume to observe an iid sample $\{Y_i,X_i,Z_i,D_i\}_{i=1}^n$ from $(Y,X,Z,D)$, and let us denote with $\mathbb P_n$ the empirical mean operator, i.e. $\mathbb P_n g(Y,X,Z,D)=n^{-1}\sum_{i=1}^n g(Y_i,X_i,Z_i,D_i)$ for any function $g$. Also, let $\mu$ be a finite measure supported on $\mathcal S $. Since the null hypothesis in ((ref)) is equivalent to
if we knew $G_{W(\beta_0)}$ and $f_{W(\beta_0)}$ we could test ((ref)) by using the Cramer-Von Mises functional
However, such an “oracle" statistic cannot be used in practice, as $G_{W(\beta_0)}$ and $f_{W(\beta_0)}$ are unknown and must be estimated. Now, if we could observe the generated regressor $H(Z)$ a test for ((ref)) could be constructed in two steps. In a first step, we could estimate $\beta_0$ by say $\widehat \beta$. In a second step, given $W(\widehat \beta)=q(\widehat \beta,X,H (Z))$, we could regress nonparametrically $Y$ on $W(\widehat \beta)$ to get an estimate of $G_{W(\beta_0)}$, and we could construct a nonparametric density estimator of $f_{W(\beta_0)}$ based on $W(\widehat \beta)$. Then, a statistic as in ((ref)) based on such estimators could be used for testing ((ref)). A test of this type is proposed in xia_goodness--fit_2004 and delgado_significance_2001 who develop bootstrap tests for equations similar to ((ref)) { when all regressors are observed}. However, since the generated variable $H(Z)$ is not observed, such tests cannot be implemented in our context.\\ Due to the presence of the generated variable $H(Z)$, we compute a { feasible} test statistic by a { three} step estimation. In a first step, on the basis of ((ref)), we regress nonparametrically $D$ on $Z$ so as to get an estimate of $H$, say $\widetilde H$. In a second step, we obtain an estimator of $\beta_0$, say $\widehat \beta$. Finally, in a third step, given an estimate of the generated regressor $\widetilde W(\widehat \beta)=q(\widehat \beta, X, \widetilde H (Z))$, we regress nonparametrically $Y$ on $\widetilde W(\widehat \beta)$ to get an estimate of $G_{W(\beta_0)}$, and we construct a nonparametric density estimator of $f_{W(\beta_0)}$ based on $\widetilde W (\widehat \beta)$. Plugging these estimators into ((ref)) gives a feasible test statistic. In the next section we detail such a three-step estimation procedure.
First step estimation. Let $K_H$ be a kernel on $\mathbb{R}^p$ and $h_H$ be a bandwidth. We define
The estimation of the generated variable $H(Z)$ could be based on $\widehat{H}(z):=\widehat{T}^{D}_Z(z)/\widehat{f}_Z(z)$. However, a test based on the first-step estimator $\widehat H$ would necessitate some undersmoothing, see Section (ref) for details. This would complicate the implementation of the test, as it would make tricky the choice of the first step bandwidth. To avoid these problems and allow for a { small-bias property} of our test statistic, we use an $L_2$ boosted or, equivalently, a bias corrected estimator in the first step. These types of estimators have been introduced in buhlmann_boosting_2003 for spline smoothing and di_marzio_boosting_2008 and park_l2_2009 for kernel smoothing. xia_goodness--fit_2004 and lapenta_encompassing_2022 used similar bias corrections for specification tests { when all regressors are observed}. The procedure works as follows. Given the initial estimator $\widehat H$, we compute the residuals $D_i-\widehat H (Z_i)$ for $i=1,\ldots,n$. Then, we construct a kernel estimator $\widehat{T}^{D-\widehat H}_Z(z)/\widehat f_Z (z)$, where $\widehat{T}^{D-\widehat H}_Z$ is defined similarly as in ((ref)), with the difference that the projected variable is $D-\widehat{H}(Z)$ instead of $D$. Here $-\widehat{T}^{D-\widehat H}_Z(z)/\widehat f_Z (z)$ is an estimate of the bias of $\widehat H(z)$, see xia_goodness--fit_2004. Finally, we consider the updated or bias-corrected estimator $\widehat H + \widehat{T}^{D-\widehat H}_Z/\widehat f_Z $. Such a procedure can be iterated a finite number of times, as proposed in buhlmann_boosting_2003, di_marzio_boosting_2008, and park_l2_2009. Here, we restrict ourselves to a one-step boosting: this is enough for our purpose to avoid undersmoothing and guarantee a small-bias property of the test statistic. \\ To precisely define the estimator, let use denote with $\mathbb I \{A\}$ the indicator function of the event $A$, and let us assume that $(X,Z)$ admits a joint density $f$ with respect to the Lebesgue measure. We define
$(\tau_n)_n$ is a sequence of positive numbers converging to zero, whose features are specified in Section (ref). Indeed, $K_0$ and $h_0$ denote a kernel and a bandwidth, while $\widehat f$ is the kernel estimator of the joint density of $(X,Z)$. We introduce the trimming $\widehat t$ in $\widehat{T}^{D - \widehat H}_Z$ to control for the random denominator of $\widehat H$. Then, we estimate $H$ by
As noticed in xia_goodness--fit_2004, $-\widehat T ^{D-\widehat H}_Z/\widehat f_Z$ can be interpreted as an estimator of the bias of $\widehat H$. So, $\widetilde H$ can be thought of as a bias-corrected estimator. di_marzio_boosting_2008 show that when bias corrections or, equivalently, boosting iterations are applied to kernel estimators, the bias decreases exponentially fast while the variance increases exponentially slow. Thus, when applied for estimation purposes these estimators are practically useful, as a kernel bias corrected estimator displays a lower mean-squared error than a kernel estimator that is not bias corrected.\footnote{Precise conditions can be found in di_marzio_boosting_2008 and park_l2_2009.} In our case, we do { not} apply these bias corrections for estimation purposes. As we argue in Section (ref), our motivation for using these bias corrections is to obtain a { small bias property} of the test statistic and to avoid undersmoothing.\\
Second-step estimation. In the second step, we estimate $\beta_0$ on the basis of the estimator of the generated variable $H(Z)$. Let $\widetilde W_i (\beta):=q(\beta,X_i, \widetilde{H}(Z_i))$ for $i=1,\ldots,n$, and let us recall that $d$ is the dimension of $\widetilde W(\beta)$. For a kernel $K$ and a bandwidth $h$, we define
Once again, the trimming $\widehat t$ is used in the above expressions to control for a random denominator in $\widetilde W(\beta)$, i.e. in $\widetilde H$. Then, an estimator of $\mathbb E \{Y|W(\beta)=w\}$ is
We estimate $\beta_0$ by a Semiparametric Least-Squares (SLS) principle:
We remark that in the above objective function only the first step estimator $\widetilde H$ is bias corrected, while $\widehat G _{\widetilde W (\beta)}$ does not contain bias corrections. As we show in Lemma (ref), this is enough to obtain a small-bias property of our test statistic. \\ A SLS estimator is also used in the specification test proposed by xia_goodness--fit_2004, where all variables are observed. Differently from xia_goodness--fit_2004, our context is featured by the generated variable $H(Z)$, so the SLS estimator used in our test is based on the first-step nonparametric estimator $\widetilde H$. This implies that the estimation error from $\widetilde H$ will appear in the influence-function representation of $\widehat \beta$, see Lemma (ref). A SLS estimator for index models with generated variables was also proposed in escanciano_identification_2016. However, in our case $\widehat \beta$ is based on the bias-corrected estimator $\widetilde H$: this allows us to avoid undersmoothing in the first step estimation and to obtain a small-bias property of the test statistic. \\
Third-step estimation. For $\widehat \varepsilon _i := Y_i - \widehat{G}_{\widetilde W (\widehat \beta)}(\widetilde W _i (\widehat \beta))$ we let
The role of the trimming $\widehat t$ in the above display is to control for a random denominator in $\widehat G _{\widetilde W (\widehat \beta)}$ and $\widetilde W(\widehat \beta)$. Similarly as in the first step estimation, $-\widehat{T}^{\widehat \varepsilon}_{\widetilde W (\widehat \beta)}/\widehat f _{\widetilde W (\widehat \beta)}$ is an estimate of the bias of $\widehat G _{\widetilde W (\widehat \beta)}$. So, we define the bias-corrected estimator of $G_{W(\beta_0)}$ as
Denoting the bias-corrected residuals as $\widetilde \varepsilon _i :=Y_i-\widetilde{G}_{\widetilde W (\widehat \beta)}(\widetilde W_i (\widehat \beta))$, our feasible statistic is given by
Let us comment on the above expression. The weighting by the estimated density $\widehat f_{\widetilde W(\widehat \beta)}$ allows us to get rid of the random denominator in $\widetilde G _{\widetilde W (\widehat \beta)}$. This has a double role. First, it “stabilizes" the behavior of $S_n$ by avoiding a denominator that can be close to zero if observations on the “tails" are included in the computation of $S_n$. Second, avoiding a random denominator in $\widetilde G _{\widetilde W (\widehat \beta)}$ simplifies our proofs for obtaining the asymptotic behavior of $S_n$. However, { due to the presence of generated variables}, we still have to control for a random denominator in $\widetilde W (\widehat \beta)$, i.e. in $\widetilde H (Z)$. This is the reason why we include the trimming $\widehat t$ in the computation of $S_n$.\\ The asymptotic behavior of $S_n$ is determined by the empirical process $\sqrt{n} \, \mathbb P _n\, \widetilde \varepsilon$ $\,\widehat f_{\widetilde W (\widehat \beta)}(\widetilde W (\widehat \beta))$ $\widehat t\,\, \varphi_s(X,\widetilde H (Z))$. As mentioned earlier, a Cramer-Von Mises statistic similar to $S_n$ has already been used for testing equations similar to ((ref)) { without generated variables}, see xia_goodness--fit_2004 and delgado_significance_2001. Differently, our context is featured by the presence of the generated regressors $H(Z)$ that are replaced by the first-step estimates $\widetilde H (Z)$. Such a replacement introduces an estimation error that needs to be taken into account. So, when obtaining the influence function representation of the empirical process at the basis of $S_n$, we are faced with the challenge of controlling for an estimation error that is not present in xia_goodness--fit_2004 and delgado_significance_2001. The estimation error from $\widetilde H (Z)$ is not negligible, as it will appear in the influence function representation of the empirical process at the basis of $S_n$, see the next section for details.
Let us start by introducing some definitions. We define the differential operator \[ \partial^{l}g(w) = \frac{\partial^{|l|}}{\partial^{l_{1}}w_{1} \ldots \partial^{l_{d}}w_{d}} g(w)\, , \qquad l=(l_{1},..,l_{d})'\, , \qquad |l| = l_{1}+..+l_{d} \, . \]
We next introduce the assumptions and then we comment on them.
Let us define the following convergence rates
The rate $d_0$ denotes the uniform convergence rate of the kernel density estimator $\widehat{f}$. The rate $d_H$ is the uniform convergence rate of the first step kernel density estimator $\widehat{f}_Z$. Finally, $d_G$ is the uniform convergence rate of the third step kernel density estimator if $H(Z)$ was observed.
Since the framework is featured by generated regressors, we need some conditions on the rates at which the densities of the variables go to zero on the tails. So, let $p_n^{W(\beta_0)}:=P(f_{W(\beta_0)}(W(\beta_0))\leq 3 \tau_n /2)$ , $p_n^Z:=P(f_Z(Z)\leq 3 \tau_n /2)$, $p_n^{X,Z}:=P(f(X,Z)\leq 3 \tau_n /2)$, and $p_n:=\max(p_n^{W(\beta_0)}, p_n^Z, p_n^{X,Z})$ . Also, let us define the sets
Assumption (ref)(ii) is an index restriction satisfied in the applications of our test described in Section (ref). In general, such an index restriction is implied by distributional assumptions on unobservables that can be justified by economic arguments, see Section (ref) for details.
Assumptions (ref)(iii)(iv) set the conditions that the weighting function $\varphi_s$ must satisfy to guarantee the equivalence between the conditional moment restriction in ((ref)) and the continuum of unconditional moments in ((ref)). By building on bierens_consistent_1982, bierens_econometric_2017 shows that Assumptions (ref)(iii)(iv) and the boundedness of $\nu(X,H(Z))$ (implied by Assumption (ref)(i) and (ref)(iii)(v)) are sufficient for such an equivalence to hold. Several choices of $\varphi_s$ have already been discussed in Section (ref). More choices can be found in stinchcombe_consistent_1998. Assumption (ref) imposes smoothing conditions that are common in the literature on semi and nonparametric testing. \\ Assumption (ref) sets the main conditions on the bandwidths for the first step estimation of $H$ and the third step estimation of $G_{W(\beta_0)}$. These conditions have a twofold role. First, together with Assumptions (ref) and (ref) they imply that the first step estimator of $H$ and the third step estimator of $G_{W(\beta_0)}$ are contained in a smooth class of functions with a bounded entropy. This is needed to obtain the asymptotic stochastic equicontinuity of the empirical process at the basis of the statistic $S_n$. Second, the bandwidth conditions in Assumption (ref) guarantee that the first step estimator $\widetilde H$, the third step estimator $\widehat G _{\widetilde W (\widehat \beta)}$, and its first-order derivative $\partial \widehat G _{\widetilde W (\widehat \beta)}$ have an $n^{-1/4}$ convergence rate towards their targets. While the $n^{-1/4}$ consistency of the nonparametric estimators is common in the literature on semiparametric testing, the $n^{-1/4}$ consistency of the first-order derivative of $\widehat G _{\widetilde W (\widehat \beta)}$ is specific to our framework with generated regressors. Heuristically, to handle the estimation error from the generated regressor $\widetilde H(Z)$, we combine the $n^{-1/4}$ consistency of the first step estimator $\widetilde H$ with the $n^{-1/4}$ consistency of the third step derivative $\partial \widehat G _{\widetilde W (\widehat \beta)}$ to get a first-order approximation of the type
The above expansion allows us to disentangle the estimation error coming from the third step and the estimation error due to the generated variable $\widetilde H(Z)$.
An important aspect is that Assumption (ref) avoids undersmoothing and allows for a { small bias property} of the empirical process at the basis of $S_n$. To explain these features, let us abstract from the appearance of the trimming rate $\tau_n$. Then, Assumption (ref) requires that $n h^{4(r-1)}=o(1)$ and $n h_H^{4 r_H}=o(1)$. These conditions have two consequences. First, we avoid undersmoothing in the sense that the test can be implemented with bandwidths that are “optimal" in terms of estimation.\footnote{ Here by “optimal" bandwidths we refer to the bandwidths minimizing the Mean Squared Errors of the nonparametric estimators. So, for the first step estimation the “optimal" bandwidth is proportional to $n^{-1/(2r_H + p)}$, while for the third step estimation the “optimal" bandwidth is proportional to $n^{-1/(2 r + d)}$. See li_nonparametric_2006. } Second, the empirical process at the basis of the statistic has a { small bias property}. This means that such an empirical process remains $\sqrt{n}$ consistent although the biases of the nonparametric estimators on which it is based do not converge to zero at a $n^{-1/2}$ rate. See newey_twicing_2004. This feature is possible thanks to the bias corrections (or equivalently the $L_2$ boosting corrections) used in the first step and the third step estimation. Without such bias corrections, the conditions $n h^{4(r-1)}=o(1)$ and $n h_H^{4 r_H}=o(1)$ should be “augmented" by $n h^{2\,r}=o(1)$ and $n h_H^{2 r_H}=o(1)$ to guarantee that the bias of the empirical process disappears fast enough, see e.g. delgado_significance_2001 and escanciano_uniform_2014. With such rates we would have to use undersmoothed bandwidths, and the empirical process at the basis of $S_n$ would no longer have a small-bias property. \\ The bias corrections used here are similar to those used in xia_goodness--fit_2004 and lapenta_encompassing_2022 who develop tests where all regressors are observed. Here, we show that these bias corrections (or equivalently the $L_2$ boosting corrections) guarantee a small bias property in a { semiparametric context} featured by the presence of { nonparametrically generated regressors}.\footnote{Due to the presence of nonparametrically generated regressors, in our framework the bias corrections of the first step estimator enter the bias corrections of the third step estimator. Thus, the way we handle the bias corrections to obtain the expansion of Proposition (ref) is fundamentally different from xia_goodness--fit_2004 and lapenta_encompassing_2022 who consider tests where { all regressors are observed}.}\\ The small bias property of the empirical process does not only have a theoretical appeal, but it is also attractive from a practical point of view: as we show in Section (ref), this property implies that when the bandwidth is perturbed our test is more stable in terms of size control than a test without a small bias property. See also newey_twicing_2004. In other words, from a practical standpoint, thanks to the small bias property we avoid an excessive sensitivity of the test with respect to the smoothing parameter. \\ Assumption (ref)(i)-(iv) gathers the conditions guaranteeing that the trimming disappears fast enough to avoid a bias coming from trimming. Similar conditions can be found in escanciano_uniform_2014. Assumption (ref)(v) implies a convexity feature needed to limit the entropy of the class of functions that asymptotically contains the nonparametric estimators. \\ Finally Assumption (ref) imposes the existence of a unique pseudo-true value of the finite dimensional parameter $\beta_0$. Such a condition is common to any specification test for nonlinear models where the finite dimensional parameter is estimated by optimizing a nonlinear objective function, see bierens_consistent_1982, lavergne_smooth_2013, escanciano_asymptotic_2018-2. Assumption (ref) plays a double role: under $\mathcal H _0$ it ensures the identification of $\beta_0$, while under $\mathcal H _1$ it ensures that the estimator $\widehat \beta$ has a well defined limit in probability. Now, under $\mathcal H_0$ Assumption (ref) can be shown by using normalization conditions on $B$ and support conditions on $X,Z$. A common normalization condition is that the first component of $\beta_0$, or equivalently the first coordinate of $B$, is set to one, see blundell_endogeneity_2004, rothe_semiparametric_2009, escanciano_identification_2016. Differently, under $\mathcal H_1$ what is needed for our test to have power is a well defined limit in probability of $\widehat \beta$. We could directly impose these conditions, but we choose to keep Assumption (ref) for presentation purposes.\\
To present the asymptotic behavior of the statistic, we need to define several objects. In the definitions below we drop the arguments of the functions when such arguments are clear from the context.\footnote{In ((ref)) for notational simplicity we use: $\partial G_{W(\beta)}(W(\beta)):=\partial_w G_{W(\beta)}(w)|_{w=W(\beta)}$, $\nabla_{\beta^T} G_{W(\beta)}(W(\beta))=\partial_{\beta^T} G_{W(\beta)}(W(\beta))+\partial^T G_{W(\beta)}(W(\beta))\partial_{\beta^T} W(\beta) $, and $\partial_H q(\beta,X,H(Z)):=\partial_u q(\beta,X,u)|_{u=H(Z)}$. }
The following proposition shows the influence function representation (IFR) of the empirical processes at the basis of the statistic $S_n$.
The above proposition is proved in Appendix (ref). A similar IFR is obtained in escanciano_uniform_2014 but under different conditions on the bandwidths and using different estimators. In particular, compared to escanciano_uniform_2014 our estimators contain bias corrections and our empirical process has a small bias property. Accordingly, the way we handle the bias terms appearing in the expansion of our empirical process is fundamentally different from the methods used in escanciano_uniform_2014.
The IFR obtained in Proposition (ref) allows us to get more insights on the connection between the small bias property of our empirical process and the presence of generated variables. The second term in the IFR of the empirical process is due to the generated variable: the fact that we replace the unobserved regressor $H(Z)$ with the generated regressor $\widetilde H (Z)$ introduces additional terms in the first-order asymptotics of our empirical process. Thus, the initial empirical process is not { locally robust} with respect to the first step estimation, in the sense that the Hadamard derivative of the moment condition in ((ref)) with respect to $H$ is non null. In other words, such a moment condition is not Neyman-orthogonal with respect to the first step estimation, see chernozhukov_locally_2016 and escanciano_asymptotic_2018-2. Now, as an alternative to the bias ($L_2$ boosting) corrections, we could get a small bias property of our empirical process by using a “locally robust" approach, as proposed in chernozhukov_locally_2016 and escanciano_asymptotic_2018-2. This would require estimating the second term appearing in the IFR of Proposition (ref) (due to the generated variable) and subtracting it to the empirical process at the basis of $S_n$. The new empirical process would be “locally robust", in the sense that it would be the empirical counterpart of a moment condition whose Hadamard derivative with respect to $H$ is zero. Then, as a test statistic we could consider the Cramer-Von Mises functional of such a { locally robust empirical process}. This new empirical process would have a small bias property as well as ours, see chernozhukov_locally_2016 for more insights. However, the implementation of such a locally robust approach in our semiparametric framework with generated variables would be difficult. Constructing a locally robust empirical process would require estimating the term $a_s$. This would require (i) to select additional smoothing parameters and (ii) to estimate the { nonparametric} derivative $\partial G _{W(\beta_0)}$. Since in practice it is not convenient to estimate nonparametric derivatives, such an estimation could complicate the practical implementation of our test. Differently, our $L_2$ boosting procedure does not require estimating nonparametric derivatives. Thus, we are proposing an attractive and convenient method to obtain the small bias property in a semiparametric context with generated variables.\\
From Proposition (ref) we can get the asymptotic distribution of $S_n$ under $\mathcal H _0$. Since $\varphi_s$ is Lipschitz in $s$, the IFR appearing in Proposition (ref) is Donsker. Hence, the empirical process at the basis of the statistic will converge weakly to a tight zero-mean Gaussian process $\mathbb G _s$ valued in $\ell^\infty (\mathcal S)$, with $\ell^{\infty}(\mathcal S)$ denoting the space of functionals on $\mathcal S$ endowed with the uniform norm. The Gaussian process $\mathbb{G}_s$ is characterized by the collection of covariances
Hence, by the continuity of the Cramer-Von Mises functional we get that under $\mathcal H _0$
From the previous two displays, the asymptotic distribution of $S_n$ depends on the unknown data generating process and we cannot use it for testing.\\ This problem is often encountered in semiparametric testing. As highlighted in the previous section, delgado_significance_2001 and xia_goodness--fit_2004 test an equation similar to ((ref)) but without generated variables. Their test statistics also converge to a functional of a zero-mean Gaussian process, and to obtain critical values they propose a bootstrap procedure. Differently from theirs, our context is featured by the presence of generated regressors. As shown in Proposition (ref) the generated variables introduce an estimation error that impacts on the asymptotic behavior of our test statistic. Thus, the bootstrap procedures proposed in delgado_significance_2001 and xia_goodness--fit_2004 cannot be applied in our testing problem, as they can't reproduce the estimation error due to the generated regressors. So, in the next section we motivate and develop a novel bootstrap procedure to get the critical values.
Before developing our bootstrap test, we will discuss some alternative methods that could be used to construct the critical values. This allows us to better motivate our bootstrap procedure. There are general methods available in the literature to adjust semiparametric tests for the presence of generated covariates, see pagan_econometric_1984, wooldridge_econometric_2010, hahn_asymptotic_2013, hahn_nonparametric_2018. They essentially work in two steps: (i) estimate the adjustment terms in the IFR of the statistic that are due to the generated variables; (ii) by using these estimated adjustment terms, correct the “standard errors" of the test statistic and construct the critical values for the test. We could adapt such a procedure to our framework, but this would be difficult for two reasons. First, this general procedure holds for a “studentized" statistic that is asymptotically pivotal, while in our empirical process framework constructing an asymptotically pivotal statistic is tricky, see song_testing_2010. Second, given the IFR in Proposition (ref), implementing such a procedure in our context requires estimating the weighting function $a_s$ defined in ((ref)). As noticed in the previous section, this requires selecting additional smoothing parameters and estimating a nonparametric derivative. So, such a procedure would be practically difficult to implement in our context for the reasons discussed earlier.\\ A second alternative would be to construct a weighted bootstrap procedure. See, e.g., huang_flexible_2016. By letting $\{\zeta_i\,:\,i=1,\ldots,n\}$ be a sequence of iid weights with mean zero, unit variance, and with a known distribution, the weighted bootstrap would be based on
where the “hatted" elements stand for estimated objects. The previous process is just a re-weighted version of the IFR in Proposition (ref). Conditionally on the sample data, the above process is expected to mimic the null behavior of the empirical process at the basis of our statistic. So, the idea would be to use it to construct valid critical values. However, such a procedure is not attractive in our case: due to the presence of generated variables, $a_s$ involves a nonparametric derivative that needs to be estimated, see Equation ((ref)).\\
In view of the previous remarks, our goal in this section is to develop a bootstrap method that (i) does not require estimating a nonparametric derivative, (ii) that can reproduce/mimic the estimation error due to the generated regressors, (iii) that resamples the observations by imposing the null hypothesis in the resampling scheme, and (iv) that is based on an empirical process with a small bias property. \\ To this end, let $\widehat W_i (\beta):=q(\beta,X_i,\widehat H (Z_i))$, where $\widehat H$ has been introduced in Section (ref). We define
where $\widehat T ^Y _{\widehat W (\beta)}$ and $\widehat f _{\widehat W (\beta)}$ are defined similarly as in ((ref)) with $\widehat W (\beta)$ replacing $\widetilde W (\beta)$.\\ Let $\{\xi_i\,:\,i=1,\ldots,n\}$ be iid copies of a random variable $\xi$ having a known distribution with $\mathbb E \xi=0$ and $\mathbb E \xi^2=1$, and let $\{\xi_i\,:\,i=1,\ldots,n\}$ be independent from the sample data $\{Y_i,X_i,Z_i,D_i\}_{i=1}^n$. The bootstrap Data Generating Process (DGP) is
for $i=1,\ldots,n$. In the “bootstrap world" only the weights $\{\xi_i\,:\,i=1,\ldots,n\}$ are random and the sample data is fixed. The bootstrap sample is $\{Y_i^*,D_i^{*},X_i,Z_i\}_{i=1}^n$. Below we describe how to construct the bootstrap version of the statistic by using such a sample. For the moment, let us highlight some important features of the bootstrap DGP in ((ref)). First, the bootstrap DGP is made by two equations. If there were no generated variables, the first equation alone would be enough to construct valid critical values. However, due to the presence of generated variables, a bootstrap DGP based only on the first equation would provide a misspecified inference, as it could not mimic the estimation error coming from the generated regressors. This is the reason for introducing the second equation in ((ref)): its role is to reproduce the estimation error from the generated variable $\widetilde H (Z)$ and to mimic the behavior of the second term in the IFR of Proposition (ref).\\ Second, in some of the applications described in Section (ref), the variable $D$ is a component of $X$. Thus, $D$ enters the true variable $W(\beta_0)=q(\beta_0,X,H(Z))$ in the population DGP. However, in the bootstrap DGP the variable $\widehat W (\widehat \beta)=q(\widehat \beta,X,\widehat H (Z))$ will not contain the bootstrap version of $D$, i.e. $D^*$, but only $D$. As we show in Proposition (ref), this is enough for providing a valid bootstrap inference. \\ Third, the observations $Y^*_i$ and $D^{*}_i$ cannot be resampled independently, but they must be generated by the { same} bootstrap weight $\xi_i$. This is an important feature, as it guarantees that the covariance of the bootstrap errors $[Y^*_i-\widehat G _{\widehat W (\widehat \beta)}(\widehat W _i (\widehat \beta))$ $ , D^*_i - \widehat H (Z_i)]$ has the same structure as the covariance of the population errors $[\,Y_i-G_{W(\beta_0)}(W_i(\beta_0))\,,$ $D_i-H(Z_i)\,]$. Since the former errors determine the behavior of the bootstrap version of $S_n$ and the latter errors determine the null behavior of $S_n$, such a feature is fundamental to guarantee the validity of the bootstrap inference. See Proposition (ref) ahead.\\ Fourth, the bootstrap DGP in ((ref)) does not involve the estimators based on bias corrections, i.e. $\widetilde H (Z)$ and $\widetilde G _{\widetilde W (\widehat \beta)}$, but only the non-corrected estimators $\widehat H(Z)$ and $\widehat G _{\widehat W (\widehat \beta)}$. As we show in Proposition (ref), this guarantees that also the bootstrapped empirical process will enjoy a small bias property. Differently, if we included the bias corrected estimators in ((ref)) the bootstrapped empirical process would not enjoy the small bias property.\\ Finally, we notice that the bootstrap DGP in ((ref)) guarantees the validity of the null hypothesis in the “bootstrap world". In particular, let us denote with $\mathbb E ^*$ the expectation in the bootstrap world, where only the weights $\{\xi_i\,:\,i=1,\ldots,n\}$ are random while the sample data is fixed. Then, ((ref)) implies that $\mathbb E ^* \{Y^*|\nu(X,\widehat H (Z))\}$ $=\mathbb E ^* \{Y^*|q(\widehat \beta, X,\widehat H (Z))\}$. This equality represents bootstrap version of $\mathcal H _0$. The fact that we impose the null hypothesis in the bootstrap world is an attractive feature, as it means that we are using all the “available information" when resampling our observations. \\ In what follows we show how to construct the bootstrap version of $S_n$. Similarly to Section (ref), this is done by a three steps procedure.\\
\noindentFirst step estimation for the bootstrap. Let
We use the trimming $\widehat t$ in the above expressions to control for a random denominator in $D^{*}$, see ((ref)). Then, the bootstrap counterpart of $\widetilde H$ is
\noindentSecond step estimation for the bootstrap. Let $\widetilde W ^* _i (\beta):=q(\beta,X_i,\widetilde H ^* (Z_i))$. Then, we define
and let $\widehat f_{\widetilde W ^* (\beta)}$ be defined similarly as in Equation ((ref)), obviously with $\widetilde W ^* (\beta)$ replacing $\widetilde W (\beta)$. Again, we use the trimming $\widehat t$ to control for the random denominator in $Y^*$ and $\widetilde W ^*(\beta)$. Then, the bootstrap counterpart of $\widehat G _{\widetilde W (\beta)}$ is
So, we let the bootstrap counterpart of $\widehat \beta$ be
Third step estimation for the bootstrap. Let us define $\widehat \varepsilon ^* _i := Y_i^* - \widehat G ^* _{\widetilde W ^* (\widehat \beta ^*)}(\widetilde W ^* _i (\widehat \beta ^*))$ for $i=1,\ldots,n$. Let $\widehat T ^{\widehat \varepsilon^*}_{\widetilde W ^* (\beta)}$ be defined similarly as in ((ref)), obviously with $\widehat \varepsilon^*$ replacing $Y^*$. Then, the bootstrap counterpart of $\widetilde G _{\widetilde W (\widehat \beta)}$ is
The bootstrap counterpart of the bias corrected residual will be $\widetilde \varepsilon ^*_i:=Y_i^*-\widetilde G ^*_{\widetilde W ^* (\widehat \beta ^*)}(\widetilde W ^*_i(\widehat \beta ^*))$. So, the bootstrap counterpart of the statistic is
Let us denote with $\Pr^*$ the probability where only the bootstrap weights $\{\xi_i\,:\,i=1,\ldots,n\}$ are random while the sample data is fixed. A test at the $\alpha$ nominal level will be based on the quantile
In practice, one can compute $\widehat c_{1-\alpha}$ by a Monte Carlo procedure that goes as follows:
If the statistic $S_n$ is larger than $\widehat c_{1-\alpha}$, then the null $\mathcal H _0$ will be rejected at the $\alpha$ nominal level. We recommend choosing $B$ such that $(1-\alpha)B$ is an integer, see davison1997bootstrap. \\
The proposition below shows the validity of the bootstrap test just proposed.
The above result is proved in Appendix (ref). Let us comment on the bootstrap expansion obtained in $(i)$. The IFR of the bootstrapped empirical process is a reweighted version of the IFR obtained in Proposition (ref). The weights are represented by the bootstrap weights $\{\xi_i\,:\,i=1,\ldots,n\}$. Notice that the covariance function of the bootstrapped IFR is the same as the covariance function in ((ref)). This guarantees that under $\mathcal H _0$ the bootstrapped IFR will have the same behavior as the population IFR of Proposition (ref). Thus, the bootstrapped statistic $S_n^*$ will mimic the null behavior of $S_n$ and the bootstrap inference will be valid. For such a feature to hold, it is fundamental that the bootstrap weights in the two equations in ((ref)) are the same. \\ Finally, we notice that Proposition (ref) is obtained under the bandwidth conditions of Assumption (ref). Thus, also the bootstrapped empirical process at the basis of $S_n^*$ has a small bias property, in the sense that it remains $\sqrt{n}$ consistent even though the bias of the nonparametric estimators on which it is based does not convergence to zero at a $n^{-1/2}$ rate.
In this section we present the main semiparametric models our test can be applied to.\\
Binary choice models with control functions. Control function methods are popular in econometrics for estimating binary choice models with endogenous regressors. They have been first introduced by rivers_limited_1988. They consider
where $Y$ denotes the binary choice of an agent, $g_1(X)=\theta_0^T X$ for a parameter $\theta_0$, $X=(D,Z_1^T)^T$, and the unobserved error $u$ is independent from $Z_1$ but correlated with $D$. To handle such a correlation, rivers_limited_1988 assume that there exists an instrument $Z_2$ such that for $Z=(Z_1^T,Z_2^T)^T$
The residual $V$ is called { control function}, as it allows controlling for the endogeneity of $D$, and $g_2(V,\epsilon)=-\gamma_0 V + \epsilon$. In addition to this structure, rivers_limited_1988 assume that $(\epsilon,V)$ are jointly normal and $\mathbb E \{D|Z\}$ is a linear function of $Z$. See also Section 4 in wooldridge_control_2015. Instead, one could consider a model given by ((ref)) and ((ref)), without imposing the normality of $(\epsilon,V)$ and the linearity of $\mathbb E \{D|Z\}$. This could be seen as a semiparametric version of rivers_limited_1988's setting. Equations ((ref)) and ((ref)) imply that $\mathbb E \{Y|X,Z\}=\mathbb E \{Y|X,V\}$ and that
The equation $\mathbb E \{Y|X,Z\}=\mathbb E \{Y|X,V\}$ is an exclusion restriction that is a direct consequence of ((ref)) and the control function assumption in ((ref)). It is obtained without imposing the parametric restrictions on $g_1$ and $g_2$, and it can be justified by economic arguments. Differently, the equation in the above display is implied by the parametric restrictions on $g_1$ and $g_2$ introduced to limit the curse of dimensionality. Our test can be applied to check the correct specification of such an equation. In particular, the equation in the previous display is a specific version of Equation ((ref)), where $\nu(X,H(Z))=(X,D-H(Z))$, $q((\theta_0,\gamma_0),X,H(Z))=(\theta_0^T X + \gamma_0(D-H(Z)))$, and the generated variable is $H(Z)=\mathbb E \{D|Z\}$. \\
In the same spirit, blundell_endogeneity_2004 consider a model given by Equation ((ref)), $X=(D,Z_1^T)^T$, $Z=(Z_1^T,Z_2^T)^T$, and the following exclusion restrictions
where $\sim$ stands for equality in distribution, and $V$ is as in ((ref)). Moreover, they assume that $g_1(X)=\theta_0^T X$ for a parameter $\theta_0$. The above display, ((ref)), and the parametric structure of $g_1$ imply that $\mathbb E \{Y|X,Z\}=\mathbb E \{Y|X,V\}$ and that
The equation $\mathbb E \{Y|X,Z\}=\mathbb E \{Y|X,V\}$ is a direct consequence of ((ref)) and the conditional independence restrictions in ((ref)). It does not need the single index structure of $g_1$. Hence, it can be justified by economic arguments. Differently, the parametric restriction on $g_1$ implies Equation ((ref)). Such a restriction is introduced for convenience (i.e. to limit the curse of dimensionality) and cannot be justified by economic arguments. Our test can be used to check the validity of such a restriction. In particular, this restriction is a particular case of ((ref)), with $\nu(X,H(Z))=(X,D-H(Z))$, $q(\theta_0,X,H(Z))=(\theta_0^T X,D-H(Z))$, and $H(Z)=\mathbb E \{D|Z\}$ as the generated variable . \\
Separable single index models with endogenous regressors. Let $Y$ be a continuous variable, $X=(D,Z_1^T)^T$, and $Z=(Z_1^T,Z_2^T)^T$. newey_nonparametric_1999 consider the model
and $V=D-\mathbb E \{D|Z\}$. Here $\epsilon$ is an unobserved error term, $G_0$ is a nonparametric function, $\beta_0$ is a vector of parameters of interest, and $D$ is an endogenous variable possibly correlated with $\epsilon$. The second part of the previous display is a mean-independence condition allowing to control for the endogeneity of $D$. Identification of $G_0$ and $\beta_0$ is discussed in newey_nonparametric_1999. Similarly to the previous subsection, this model implies that $\mathbb E \{Y|X,Z\}=\mathbb E \{Y|X,V\}$ and that $\mathbb E \{Y|X,V\}=\mathbb E \{Y|X^T\beta_0,V\}$. The equation $\mathbb E \{Y|X,Z\}=\mathbb E \{Y|X,V\}$ does not depend on the index restriction on $G_0$ and is a direct consequence of the mean independence condition in ((ref)). Hence, it can be justified by economic arguments. Differently, the equation $\mathbb E \{Y|X,V\}=\mathbb E \{Y|X^T\beta_0,V\}$ is a parametric restriction that can be checked by using our test, with $H(Z)=\mathbb E \{D|Z\}$ as the generated variable, $\nu(X,H(Z))=(X,D-H(Z))$, and $q(\beta_0,X,H(Z))=(\beta_0^T X,D-H(Z))$.\\
Models with sample selection. escanciano_identification_2016 consider a semiparametric model with sample selection and possibly truncation on the response variable. Let $\widetilde Y \in \mathbb R$ be a scalar random variable denoting an agent's decision, $(X,Z)$ be a vector of covariates, $D\in \{0,1\}$ be a selection variable, and $(\epsilon,u)$ be unobserved error terms. The model is
where $\psi$ is a known function, $g(X)=\beta_0^T X$, but $H$ is nonparametric. Without loss of generality we can assume $u\sim U[0,1]$, so $H$ is identified as $H(Z)=\mathbb E \{D|Z\}$ and denotes the propensity score. Because of selection, only $Y=\widetilde Y D$ is observed. Let us denote with $F$ the joint distribution of $(\epsilon,u)$. With $F$ nonparametric, this setup is a semiparametric generalization of Heckman's sample selection model. Putting $\psi(g(X) ,\epsilon)=\mathbb I \{g(X) \geq \epsilon\}$ gives a binary choice model with sample selection. If instead $\psi(g(X) ,\epsilon)=\max \{0,g(X) + \epsilon\}$ we get a truncated regression model with sample selection. This model is also known as double hurdle model, see escanciano_identification_2016 and cragg_statistical_1971. By the independence between $(\epsilon,u)$ and $(X,Z)$ we have $\mathbb E \{Y|X,Z\}=\int \psi(g(X), \epsilon)\, $ $\mathbb I \{H(Z)\geq u\} d\,F(u,\epsilon)$. Hence, the model implies $\mathbb E \{Y|X,Z\}= \mathbb E \{Y|X,H(Z)\}$ and
The equation $\mathbb E \{Y|X,Z\}= \mathbb E \{Y|X,H(Z)\}$ is obtained from ((ref)) without assuming the linear structure of $g$. Differently, the equation in the above display comes from the parametric restriction on $g$. Such an equation can be checked by using our test. In terms of the notation of Section (ref), our test can be applied to this framework by putting $\nu(X,H(Z))=(X,H(Z))$, $q(\beta_0,X,H(Z))=(\beta_0^T X, H(Z))$, and setting the generated variable as $H(Z)=\mathbb E \{D|Z\}$.\\
Semiparametric games with incomplete information. aradillas-lopez_pairwise-difference_2012 and lewbel_identification_2015 study identification and estimation in binary games with incomplete information and linear payoffs. For simplicity, let us consider two players indexed by $j\in \{1,2\}$. Each must take a binary decision, say $a_j\in \{0,1\}$. Let $X_j$ be the vector of observed covariates entering agent $j$'s payoff function and let $Z=(X_1^T,X_2^T)^T$. Each player has a private information $u_j$ that neither the other player nor the researcher can observe. It is assumed that $(u_1,u_2)\perp Z$ and $u_1\perp u_2$. The payoff function of player $j$ is
with $g_j(X_j, a_{-j})=\gamma_j^T X_j + \alpha_j a_{-j} $. We let $H_j(Z)=\mathbb E \{a_j|Z\}$. Under the hypothesis that a unique Bayesian-Nash equilibrium is played, the model implies that $\mathbb E \{a_j |X_1,X_2\}=\mathbb E \{a_j|X_j,$ $H_{-j}(Z)\}$ and that
for $j=1,2$. The equation $\mathbb E \{a_j |X_1,X_2\}=\mathbb E \{a_j|X_j,H_{-j}(Z)\}$ follows directly from the assumption that a unique Bayesian-Nash equilibrium is played and does not need the linearity of the profit function $g_j$. This can be justified by economic arguments. Differently, the equation in the previous display follows from a parametric restriction on $g_j$. Our test can be used to check such a condition. In terms of the notation of Section (ref), our test can be applied to this model with $\nu(X,H(Z))=(X_j,H_{-j}(Z))$ $q((\gamma_j,\alpha_j), X_j,H_{-j}(Z))=\gamma_j^T X _j + \alpha_j \, H_{-j}(Z)$, $D=a_{-j}$, and $H_{-j}(Z)=\mathbb{E}\{a_{-j}|Z\}$ as a generated variable.
In the first part of this section we show how to implement our test in practice. In the second part, we study its behavior in small samples. Finally, in the third part we apply our test to a real data example.
To show the practical implementation of our test, we provide the computational details for the three steps procedure described in Section (ref). We suggest to prior trim the 1% more extreme the observations on $X,Z$.\footnote{ Specifically, we trim those observations falling beyond the 99% quantile of the empirical distribution of $(|X|,|Z|)$. } Let us now describe the first step. The bandwidth for $\widetilde H $ is set according to the Silverman's rule of thumb, so $h_H=\widehat {\sigma}_Z\, n^{1/(2+2\,r_H)}$ with $\widehat \sigma _Z$ denoting the estimated standard deviations of the components of $Z$. The kernel $K_H$ is set to a second order Gaussian kernel. Then, we compute $\widetilde H $ as in ((ref)). \\ For the second step of Section (ref), we need to compute the estimator $\widehat \beta$ by minimizing the objective function in ((ref)). Notice that the estimator $\widehat G_{\widetilde W (\beta)}$ in ((ref)) depends on both $\beta$ and the bandwidth $h$. We follow the standard approach in semiparametric estimation, see delecroix_semiparametric_2006, rothe_semiparametric_2009, escanciano_identification_2016, maistre_nonparametric_2018, and in practice we choose $\widehat \beta$ by solving the joint minimization problem
where $\widehat G_{\widetilde W (\beta)}^{-i}$ denotes the leave-$i$-out version of $\widehat G_{\widetilde W (\beta)}$ defined as\footnote{For notational convenience we are dropping the dependence of $\widehat G _{\widetilde W (\beta)}^{-i}$ from $h$. }
The $\widehat \beta$ from ((ref)) is the estimator of $\beta_0$. \\ Let us now describe the computational details of the third step of Section (ref). To compute the statistic in ((ref)) we need to select (i) a weighting function $\varphi_s$, a measure $\mu$ , and a set $\cal S$, and (ii) a bandwidth $h$ and a kernel $K$ for $\widetilde G_{\widetilde W (\widehat \beta)}$ and $\widehat f _{\widetilde W (\widehat \beta)}$. We select $\varphi_s$, $\mu$, and $\cal S$ to make the computation of the integral in ((ref)) fast and simple, see below for details. So, we set $\varphi_s(\nu(X,H(Z)))=\exp(\,\sqrt{-1}\,s^T\, \nu(X,H(Z))\,)$, $\mu$ to the standard multivariate Gaussian density, and $\mathcal {S}=\mathbb{R}^{\dim\nu(X,H(Z))}$. We set the kernel $K$ to the second order Gaussian kernel. Finally, to choose the bandwidth $h$ we minimize the distance between the semiparametric model and the null hypothesis, so as to compute $S_n$ in an “optimistic" way. To this end, let us define the leave-$i$-out version of $\widetilde G_{\widetilde W (\widehat \beta)}$ as follows
where $\widehat G ^{-i}_{\widetilde W (\widehat \beta)}$ is defined in ((ref)), $\widehat f^{-i}_{\widetilde W (\widehat \beta)}$ is the leave-$i$-out version of $\widehat f_{\widetilde W (\widehat \beta)}$ (defined similarly as in ((ref))), and
Then, the bandwidth $h$ for testing is selected as follows
The intuition is simple: we are choosing the bandwidth for testing that makes the model “close" to $\mathcal H _0$, where the distance is measured by a leave-one-out version of the Cramer-Von Mises metric. The procedure can be seen as an adaptation of the Cross-validation principle to our testing problem: while the Cross Validation chooses the bandwidth minimizing the distance between the leave-one-out residuals and zero, in our testing problem we choose the bandwidth to minimize the distance between the semiparametric model and the null hypothesis. This distance is measured by a leave-one-out version of the Cramer-Von Mises functional. As we explain below, this way of choosing the bandwidth allows us to compare an “optimistic" statistic computed with the sample data to an “optimistic" statistic computed with the bootstrap data, see below for details.\footnote{As an alternative, we could choose the bandwidth to maximize the value of the statistic in the hope to increase the power. This approach is suggested in escanciano_uniform_2014 for an asymptotically pivotal test statistic. However, such a procedure would appear to be tricky in our case, since (i) our statistic is not asymptotically pivotal and (ii) we also have to choose the bandwidth in the bootstrap sample.}\\ Thanks to the choice of $\varphi_s$, $\mu$, and $\mathcal{S}$, the integral in the optimization problem ((ref)) has a simple closed form expression given by
where $\exp(-\|\cdot\|^2/2)$ is the characteristic function of the standard multivariate normal. Since the above double sum depends non linearly on $h$, the optimization in ((ref)) must be carried out numerically.\\ Once $\widehat h$ has been obtained from the minimization in ((ref)), we can use it to compute the third step estimators $\widetilde{G}_{\widetilde W (\widehat \beta)}$ and $\widehat f _{\widetilde{W}(\widehat \beta)}$ as in Section (ref). Thus, given the residuals $\widetilde \varepsilon_i =Y_i - \widetilde G_{\widetilde W (\widehat \beta)}(\widetilde W_i (\widehat \beta))$ with $i=1,\ldots,n$, we get the test statistic:
Let us now detail the computations for the bootstrapped statistic $S_n^*$ described in Section (ref). First, to generate bootstrap observations from ((ref)) we need to construct the estimates $\widehat H$ and $\widehat G_{\widehat W (\widehat \beta)}$, and we need to define a distribution for the bootstrap weights $\{\xi_i\,:\,i=1,\ldots,n\,\}$. The estimators $\widehat H$ and $\widehat G_{\widehat W (\widehat \beta)}$ entering the bootstrap DGP in ((ref)) are computed by using second order Gaussian kernels and the bandwidths $(h_H,\widehat h)$ obtained from the sample data. We suggest to generate the bootstrap weights from a two-points Rademacher distribution: $\Pr(\xi=1)=\Pr(\xi=-1)=1/2$. The choice of this distribution for the bootstrap weights is motivated by its good performance in other contexts, see davidson_wild_2008, djogbenou_asymptotic_2019.\footnote{ I thank an anonymous referee for having suggested the use of these weights. } Then, given the bootstrap sample $\{Y_i^*,D_i^{*},X_i,Z_i\, :\, i=1,\ldots,n\}$ obtained from ((ref)), we use { the same procedure as in the sample} to compute the bootstrap version of the test statistic, $S_n^*$. Notice that for each bootstrap sample/iteration we will have to minimize the bootstrap version of ((ref)) to get $\widehat \beta^*$ and the bootstrap version of ((ref)) to get $\widehat h ^*$. Finally, the critical values are obtained by the Monte Carlo procedure described in Section (ref). \\ At this point, it is worth to spend some words on the bandwidth selection procedure we have proposed. As previously highlighted, $\widehat h$ is selected to minimize the distance between the model and $\mathcal H _0$. This gives an “optimistic" value for the sample statistic $S_n$. However, also in the “bootstrap world" the bandwidth $h$ is chosen to minimize the distance between $\mathcal H _0$ and the semiparametric model. Thus, also in the bootstrap world the statistic $S_n^*$ is computed with an optimistic view. Now, the optimistic view of $S_n$ will reflect the reality only when the null is true. Differently, the optimistic view of $S_n^*$ will { always} reflect the reality in the bootstrap world: since we are imposing the null $\mathcal H _0$ in the resampling scheme, the null is always satisfied in the bootstrap world. This implies that when $\mathcal H _0$ holds, the behavior of $S_n$ will reflect the behavior of $S_n^*$, allowing to control the size of the test. However, when the null hypothesis $\mathcal{H}_0$ does not hold, $S_n$ will be too large as compared to the distribution of $S_n^*$, and the test will reject $\mathcal H _0$.
In this section we study the small sample performances of our test in a Monte Carlo experiment. We specify a Data Generating Process (DGP) in line with the binary choice model with a control function described in Section (ref), and in particular in Equations ((ref)) and ((ref)). So,
$(Z^{in}, Z^{ex} , V,\epsilon)$ are mutually independent. The error terms $(u,V,\epsilon)$ are not observed. $u$ is correlated with the endogenous regressor $D$ through $V$, so $V$ plays the role of a control function. The regressors $Z^{in}$ and $Z^{ex}$ are exogenous: the former is included in the equation for $Y$, while the latter is an excluded variable playing the role of an instruments. The control variable $V$ is unobserved but can be estimated by estimating $H$. The functional form of $H$ is unknown to the researcher, so $H$ must be estimated nonparametrically. \\ We set $V\sim \mathcal{N}(0,1)$, $\alpha=(1,1,1)/\sqrt{2}$, $Z^{in}\sim \mathcal{N}(0,1)$, and we draw $Z^{ex}$ from an exponential distribution truncated from above at 3 and standardized to have zero mean and unit variance. The parameter $p$ controls the departure from the null hypothesis, see below for details. \\ To check the robustness of our test with respect to different DGPs, we consider three different specifications for the distribution of $\epsilon$:
DGP 1 delivers a rescaled probit model with a distribution of $\epsilon$ that is unimodal and symmetric about zero. DGP 2 gives a unimodal distribution of $\epsilon$ with positive asymmetry and left skewness. Finally, DGP 3 generates $\epsilon$ according to a mixture between two Gaussians, delivering a distribution of $\epsilon$ that is bimodal and left-skewed. The three distributions for $\epsilon$ and ((ref)) are built so as to guarantee realistic features of the simulated data. Across these DGPs it holds that $\operatorname{Var} (u)\approx 8$ , Corr$(u,V)\approx 0.35$ , Corr$(u,D)\approx 0.25$ , and $\operatorname{Var} (D + Z^{in})\approx 4.5$. \\ We assume that the researcher specifies the model as
$V=D-\mathbb{E}\{D|Z\}$, and $\epsilon$ independent from all the other variables. The distribution of $(\epsilon,V)$ is nonparametric. Similarly to Section (ref), to check the correct specification of this model the researcher needs to test
where $X=(D,Z^{in})$. Given the DGP in ((ref)), the null hypothesis holds if $p=0$. As long as $p\neq 0 $ we are under the alternative $\mathcal{H}_1$. The magnitude of $p$ measures the departure from the null hypothesis. \\
In the reminder of this section our goal is threefold. First, we want to analyze the practical advantages of the bias corrections in our statistic guaranteeing the small-bias properties of our test. Second, we want to check the capacity of our test to correctly control the size under $\mathcal H _0$. Third, we want to analyze its power properties. \\ Let us start from the first task. To check the advantages of using bias corrections and having a small bias property of our statistic, we will compare the performances of two tests under different bandwidth choices. The first test is the one proposed in this paper that uses the { bias corrected} estimators $\widetilde H$ and $\widetilde G _{\widetilde W (\widehat \beta)}$. We will call it { BC Test}. The second test employs the procedure described in this paper, with the only exception that it does { not} use the bias corrected estimators, but it employs the { uncorrected} estimators $\widehat H$ and $\widehat G _{\widehat W (\widehat \beta)}$. We will call this test { UN Test}. We compare the ability of these two tests to control the size under different bandwidth choices. So, in the third step estimation we will use the bandwidth rule $h=\,C\, \widehat \sigma(D + \widehat \theta Z^{in} + \widehat \gamma (D-\widetilde H (Z))) n^{-1/6}$, where $\widehat \sigma$ is the estimated standard deviation of its argument and $C$ is a constant. This allows us to study the stability of each test under different bandwidth choices by letting the constant $C$ vary.\footnote{We stress that for this comparison the bandwidth $h$ in the third step is not computed as in ((ref)), but according to the rule $h=C\, \widehat \sigma(D + \widehat \theta Z^{in} + \widehat \gamma (D-\widetilde H (Z)))\, n^{-1/6}$.} We have ran 10000 simulations under the null $\mathcal H _0$, i.e. by setting $p=0$, for $n=400$. Since the procedure is intense from a computational point of view, see Section (ref), we used the Warp-Speed method proposed by davidson_improving_2007 and studied in giacomini_warp-speed_2013. This consists in drawing one bootstrap sample for each Monte Carlo simulation, and then use the entire set of bootstrapped statistics to compute the bootstrap p-values associated to each original statistic. We let the constant $C$ vary in the range of values $\{0.5,1,1.5,2,2.5\}$. For each of these values and for each simulated sample, the { BC Test} and the { UN Test} give two respective p-values and hence two respective decisions about the rejection of the null $\mathcal H _0$. The results are reported in Figures (ref) and (ref) for DGP 1. For a reason of space, we omit the results for DGP 2 and DGP 3 since they are qualitatively similar to DGP 1. For each value of $C$, Figures (ref) and (ref) contain the Error in Rejection Probability (ERP) of each test: this is the difference between the empirical rejection frequency and the nominal size of the test under the null $\mathcal H _0$. An ideal test would display a null ERP for each nominal size and for each value of $C$. This comparison allows us to evaluate (i) the ability of each test in controlling the size under different bandwidth choices, and (ii) the importance of employing bias corrections for estimating the null distribution of the statistic by wild bootstrapping. In other words, this gives us a measure of the stability/robustness of each bootstrap test with respect to the bandwidth. For $C=0.5$ up to $C=1$ both the { BC Test} and the { UN Test} display a good empirical size control for any nominal size. For $C=1.5$ and $C=2$ the size control starts deteriorating for the { UN Test} but remains quite stable for the { BC Test} at any nominal size. This means that the bootstrap does { not} manage to provide a good estimation of the null distribution of the statistic when bias corrections are { not} employed. This feature becomes even more clear for $C=2.5$: in this case the { UN Test} has a large ERP, while the { BC Test} shows a good size control for nominal levels up to 10 percent. These results show that employing bias corrections enables the bootstrap test to be stable with respect to the bandwidth choice. Also, these bias corrections allow for a satisfactory approximation of the null distribution of the statistic by our wild bootstrap procedure. Thus, such corrections do not only have the attractive theoretical feature of guaranteeing a small bias property of the test statistic, but they also have remarkable practical advantages.\\ Let us now turn to our second and third task. Here, we analyze the performance of our test in terms of size control and power properties when the third step bandwidth $h$ is selected as in ((ref)). We label such a test as { BC Test ($\widehat h$)}. We compare the performance of this test with the { BC Test} described previously in this subsection and applied with $C=1$. This bandwidth rule corresponds to the Silverman's Rule of Thumb. We label this test as { BC Test($C=1$)}. We ran 10000 Monte Carlo replications for $n=400$ and $n=800$. Also in this case the procedure is intense from a computational point of view: we need to run a nonlinear optimization for obtaining $\widehat \beta$ and a nonlinear optimization for obtaining $\widehat h$, for both the sample and the bootstrap data. So, to speed up computations we use the warp speed method previously described. We compute the empirical rejection frequencies for the { BC Test ($\widehat h$)} and the { BC Test($C=1$)} under $\mathcal H _0$, i.e. by setting $p=0$. This allows us to evaluate the performance of the test in terms of size control for each sample size. To evaluate the performance of the test in terms of power, we use $p=3$ and $p=4$. Indeed, the larger $p$ the more we depart from the null hypothesis. The results are reported in Tables (ref), (ref), and (ref) for the nominal sizes of 5% and 10%. We notice that the { BC Test ($\widehat h$)} shows a good size control under $\mathcal H _0$ across the different DGPs considered, also for the limited sample size of $n=400$. The test also displays a satisfactory power across the different DGPs. For a given sample size, the power of the test increases as we depart from the null hypothesis, i.e. when $p$ increases. Similarly, for a given level of $p$, the power of the test increases as the sample size grows. This simulation experiment shows that the bootstrap { BC Test ($\widehat h$)} based on the empirical bandwidth selection in ((ref)) manages to control well the size under $\mathcal H _0$ and has a satisfactory power.
We illustrate the practical implementation of our procedure in a real-data example. We test the specification of a semiparametric model describing women's labor force participation, see wooldridge_control_2015. Such a model studies the impact of non-labor income on women's labor force participation decisions. The data we use are taken from the 1991 wave of the Current Population Survey.\footnote{I am very grateful to Jeffrey Wooldridge for having shared his data set.} The sample consists of 2,947 women who do not have kids under the age of 6 and have a positive non labor income. We set up the model and the notation similarly as in Section (ref). The dependent variable $Y$ is an indicator that equals one if the woman participates to the labor force and 0 otherwise. $D$ represents the household's “other sources of income". The vector of controls $Z^{in}$ includes the logarithm of the woman's experience and a dummy indicating college or above education for the woman. We treat $D$ (other sources of income) as endogenous. Following wooldridge_control_2015, we set $Z^{ex}$ to the husband's level of education. This is a dummy indicating college or above education, and it is used as an instrument to control for the endogeneity of $D$. We normalize to unity the coefficient attached to $D$. \\ We apply our test to check the correct specification of a semiparametric index model similar to the one described in the previous section. Thus, we test the hypothesis in ((ref)). To implement the test in this real data example, we apply the steps described in Section (ref). We consider the { BC Test}($\widehat h$) where the bandwidth is computed as in Equation ((ref)). Notice that the nonlinear optimization for the bandwidth has to be done with the sample data and for each bootstrap replication. The number of bootstrap iterations is set to $B=999$. In Table (ref), we report the value of the test statistic and the 99% bootstrap quantile. This is the quantile of the bootstrap distribution of $S_n$. Since the value of the statistic is well above the 99% bootstrapped quantile, the correct specification of the model must be rejected at the 1% nominal level.
This is a revised version of a chapter of my PhD thesis. I thank my supervisor Pascal Lavergne and the components of my PhD committee Jean-Pierre Florens, Ingrid Van Keilegom, and Juan Carlos Escanciano for their comments and suggestions. I also thank Xavier d'Haultfoeuille and Jad Beyhum for their comments. Finally, I thank three anonymous referees for their comments. This research has received financial support from the European Research Council under the European Community's Seventh Program FP7/2007-2013 grant agreement N. 295298.