Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
22,436 characters · 3 sections · 29 citation commands
Bootstrap Inference on Partially Linear Binary Choice Model
In this paper, we propose a bootstrap inference procedure for the partially linear binary choice model studied by krief2014integrated. As introduced in krief2014integrated, this model is useful for estimating structural equations where nonlinearity is suspected, possibly arising from diminishing marginal returns, different life cycle regimes, or hectic physical phenomena. This model may also encompass the endogenous binary response model with a median control function restriction blundell2004endogeneity and the binary response model with a partially nonadditive error. krief2014integrated proposes a two-stage smoothed maximum score (SMS) estimator for the coefficient vector in the model and provides the asymptotic distribution for the estimator.
For inference of the model, horowitz2002bootstrap and cao2021smoothed show that the hypothesis testing results in finite samples based on the SMS estimator using analytic asymptotic approximation may not be reliable if the sample size is not sufficiently large. We follow the idea of horowitz2002bootstrap and cao2021smoothed and propose a bootstrap inference approach for the partially linear binary choice model. Our simulation results show that the proposed method can significantly address the over rejection issue caused by using the analytic asymptotic approximation, while maintaining a relatively high level of power.
We consider the partially linear binary choice model in krief2014integrated:
where $Y$ is a binary dependent variable, $X$ is a $d\times 1$ vector of independent variables, $\beta_X$ is the corresponding vector of unknown coefficients, $V$ is a continuously distributed univariate, $\phi $ is an unknown real-valued function, and $\epsilon$ is an unobservable error term.
For the purpose of identification, we follow krief2014integrated and assume that the first component of $X$, denoted by $X_{1}$, conditional on both the remaining components of $X$, denoted by $\widetilde{X}$, and $V$ admit a distribution function absolutely continuous with respect to the Lebesgue measure almost surely. Following horowitz1992smoothed, krief2014integrated, and chen2015binary, we assume $\beta_{1}=1$ for scale normalization, where $\beta_{1}$ is the coefficient of $X_{1}$. Let $\beta$ denote the coefficients for $\Tilde{X}$.
Under the assumption that the conditional median of $\epsilon$ given $X$ and $V$ is zero, i.e., $\operatorname{Med}(\epsilon|X,V)=0$, krief2014integrated proposed the integrated kernel-weighted smoothed maximum score (IKWSMS) estimator for $\beta$ in a two-stage procedure. Let $ G(\cdot) $ be a continuous function that corresponds to the integral of a $r$-th order kernel function for some integer $r\geq 4$.\footnote{See, for example, the definition of a $r$-th order kernel function in li2006nonparametric.} We assume that $\sup_{t\in\mathbb{R}} |G(t)|<M $ for some $ M<\infty $, $ \lim_{t\rightarrow-\infty}G(t)=0 $, and $ \lim_{t\rightarrow\infty}G(t)=1 $. Let $K(\cdot)$ be another kernel function. Suppose we observe an i.i.d.\ sample $\{(Y_i,X_i,V_i)\}_{i=1}^n$ from the distribution of $(Y,X,V)$.
In the first stage, we fix a $v$ in the support $\mathcal{V}$ of $V$ and estimate $\beta$ for $v$ by
where $W_i^T=(1,\widetilde{X}_i^T)$, $e_{d}=[O,I_{d-1}]$ is a $(d-1)\times d$ matrix with the first column being a zero vector, $I_{d-1}$ is the $(d-1)\times (d-1)$ identity matrix, $\Theta\subset \mathbb{R}^d$ is some compact set, and $h$ and $h_{v}$ are the smoothing parameters that converge to $0$ as $n\rightarrow\infty$. In this way, we obtain the estimator $\beta_{n}(v)$ for every $v\in\mathcal{V}$. In the second stage, the IKWSMS estimator, denoted by $\widehat{\beta}$, is defined as
where $\tau(\cdot)$ is a known weighting function on the real line with compact support that satisfies
For each $v$, define $L(v)=X^T\beta_X+\phi(v)$. Let $F_{\epsilon|W,L(v),V}$, $f_{V|W,L(v)}$, and $f_{L(v)|W}$ be the conditional cumulative distribution function (CDF) of $\epsilon$ given $(W,L(v),V)$, the conditional probability density funciton (PDF) of $V$ given $(W,L(v))$, and the conditional PDF of $L(v)$ given $W$, respectively. Then we define the random elements
and $T_{W}^{(1)}(l,v)=\partial T_{W}(l,v)/\partial l$ whenever these quantities exist. Also, we define $Z=X^T\beta_X+\phi(V)$. Then we let $F_{\epsilon|Z,W,V}$ and $f_{Z|W,V}$ be the conditional CDF of $\epsilon$ given $(Z,W,V)$ and the condition PDF of $Z$ given $(W,V)$, respectively. We define the random elements
and
for $j=1,\ldots,r$, provided these derivatives exist. Furthermore, for every $v$ such that all the relevant elements exist, we define
and
Under certain conditions, krief2014integrated showed that
where $\lambda=\lim_{n\rightarrow\infty}\sqrt{nh}h^{r}<\infty$.
It has been shown in horowitz2002bootstrap and cao2021smoothed that hypothesis testing of the SMS estimator based on asymptotic critical values may not have good finite sample performances. This is because in nonparametric and semiparametric estimations, large samples may be required to achieve reasonable agreement between the finite sample properties and the results of asymptotic theory. We next suggest using a bootstrap method for inference that could provide refinements to the empirical size of the test in finite samples.
For inference, we first consider testing the hypothesis
where $\beta_{j}$ denotes the $j$th component of $\beta$ and $\beta_{0j}$ is some prespecified value. To avoid computing the complicated bias term when conducting the test, we follow horowitz2002bootstrap and use the undersmoothed bandwidth $h$ such that $h$ converges to zero sufficiently fast so that $\lambda=0$. Thus, the test statistic for the $H_0$ in (ref) can be constructed as
where $\widehat{\beta}_j$ denotes the $j$th component of $\widehat{\beta}$ and $\widehat{\Omega}_{j}$ is the $(j,j)$ component of $\widehat{\Omega}$, a consistent estimator of $\Omega$. Based on Lemma 2 and the proof of Theorem 1 in krief2014integrated, such an estimator may be constructed as
where
and $H(V_{j})$ is the Hessian matrix of the objective function of $\theta$ in ((ref)) evaluated at $\widehat{\theta}(V_j)$. Algorithm (ref) illustrates the bootstrap procedure for the test.
We are also interested in a more general hypothesis
where $\mathcal{F}$ can be a vector valued function with continuous first derivatives. One special case is that $\mathcal{F}(b)=Rb$ with $b\in\mathbb{R}^{d-1}$ for some suitable matrix $R$ which incorporates the case where $H_{0}: \beta_{j}=\beta_{0j}$. Let $\mathcal{F}'(b)=\partial\mathcal{F}(b)/\partial b^T$ for all suitable $b\in\mathbb{R}^{d-1}$. Then we define the test statistic for the $H_0$ in (ref) as
Algorithm (ref) illustrates the bootstrap procedure for the test.
The asymptotic properties of the bootstrap tests in Algorithms (ref) and (ref) may be proved analogously to Theorems 4.3 in horowitz2002bootstrap. To avoid theoretical complications, we omit the proofs in the paper and show the finite sample properties of the tests in Monte Carlo simulations.
In this section, we present the Monte Carlo investigation of the finite sample performance of the proposed method. We design simulations based on those in horowitz2002bootstrap and krief2014integrated. Specifically, we consider estimating $\beta$ in the following model:
where $\beta=1$, $\phi(v)=\cos(2\pi v)$ for every $v$, $V=\Phi(\xi)$ with $\Phi(\cdot)$ being the cumulative distribution function of the standard normal random variable, and $(X_{1},X_{2},\xi)$ is a standard normal triplet with correlation coefficient $0.2$. We consider five distributions for the error $\epsilon$ as follows:
In all designs, $\epsilon$ is normalized with variance $1$. Every Monte Carlo experiment consists of $500$ replications. The warp-speed method of giacomini2013warp is employed to expedite the simulations. Similar to krief2014integrated, the smoothing function for the indicator function is
The derivative of $G$ is a kernel function with order $r=4$ muller1984smooth. The first-stage kernel function $K$ is
which is of order $8$ pagan1999nonparametric. The bandwidth selection method follows krief2014integrated (Silverman-like rule of thumb in silverman1986density). Specifically, for every $v$, we let $h=1.22 R_{L}n^{-1/(2r+1)}$ and $h_{v}=0.8R_{V}n^{-3/25}$, where $R_{L}$ and $R_{V}$ denote the interquartile ranges between the first and third quartiles of $\{X_{i1}+(1,X_{i2})\widehat{\theta}(v)\}_{i=1}^{n}$ and $\{V_{i}\}_{i=1}^{n}$, respectively, and $\widehat\theta(v)$ is obtained with $(h_0,h_{v0})=(n^{-1/(2r+1)},n^{-3/25})$ by
The weighting function $\tau(v)=1.25\cdot\mathds{1}(0.1\leq v \leq 0.9)$.
First, we run simulations to show the size control of the test. Both the asymptotic and bootstrap critical values are used to compute the empirical size of the $t$ test for $H_{0}: \beta=1$. The sample size $n$ is set to $1000$. The results of the experiments are shown in Table (ref). In contrast to the results obtained from using the asymptotic critical values, the empirical sizes obtained from using the bootstrap critical values are well controlled. Furthermore, the empirical sizes obtained from using the bootstrap critical values show little sensitivity for undersmoothed bandwidths $0.75h$ and $0.5h$. They are all well controlled for undersmoothed bandwidths.
We also investigate the empirical power of the test. The sample sizes we consider are $n\in\{250,500,1000\}$. Table (ref) reports the rejection rates of the $t$ test for the false hypothesis $H_{0}: \beta=0$. The empirical powers obtained from using the asymptotic critical values exceed those obtained from using the bootstrap critical values. This is natural because of the upward size distortions arising from using the asymptotic critical values. But we note that the differences are small. Furthermore, the empirical power of the proposed bootstrap test increases as the sample size increases, which demonstrates the good finite sample power property of the proposed test.