EconBase
← Back to paper

Bootstrap Inference on Partially Linear Binary Choice Model

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

22,436 characters · 3 sections · 29 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Bootstrap Inference on Partially Linear Binary Choice Model

abstractThe partially linear binary choice model can be used for estimating structural equations where nonlinearity may appear due to diminishing marginal returns, different life cycle regimes, or hectic physical phenomena. The inference procedure for this model based on the analytic asymptotic approximation could be unreliable in finite samples if the sample size is not sufficiently large. This paper proposes a bootstrap inference approach for the model. Monte Carlo simulations show that the proposed inference method performs well in finite samples compared to the procedure based on the asymptotic approximation. Keywords: Partially linear binary choice model, bootstrap inference, good finite sample properties

Introduction

In this paper, we propose a bootstrap inference procedure for the partially linear binary choice model studied by krief2014integrated. As introduced in krief2014integrated, this model is useful for estimating structural equations where nonlinearity is suspected, possibly arising from diminishing marginal returns, different life cycle regimes, or hectic physical phenomena. This model may also encompass the endogenous binary response model with a median control function restriction blundell2004endogeneity and the binary response model with a partially nonadditive error. krief2014integrated proposes a two-stage smoothed maximum score (SMS) estimator for the coefficient vector in the model and provides the asymptotic distribution for the estimator.

For inference of the model, horowitz2002bootstrap and cao2021smoothed show that the hypothesis testing results in finite samples based on the SMS estimator using analytic asymptotic approximation may not be reliable if the sample size is not sufficiently large. We follow the idea of horowitz2002bootstrap and cao2021smoothed and propose a bootstrap inference approach for the partially linear binary choice model. Our simulation results show that the proposed method can significantly address the over rejection issue caused by using the analytic asymptotic approximation, while maintaining a relatively high level of power.

Setup and Inference Procedure

We consider the partially linear binary choice model in krief2014integrated:

equation[equation omitted — 82 chars of source]

where $Y$ is a binary dependent variable, $X$ is a $d\times 1$ vector of independent variables, $\beta_X$ is the corresponding vector of unknown coefficients, $V$ is a continuously distributed univariate, $\phi $ is an unknown real-valued function, and $\epsilon$ is an unobservable error term.

For the purpose of identification, we follow krief2014integrated and assume that the first component of $X$, denoted by $X_{1}$, conditional on both the remaining components of $X$, denoted by $\widetilde{X}$, and $V$ admit a distribution function absolutely continuous with respect to the Lebesgue measure almost surely. Following horowitz1992smoothed, krief2014integrated, and chen2015binary, we assume $\beta_{1}=1$ for scale normalization, where $\beta_{1}$ is the coefficient of $X_{1}$. Let $\beta$ denote the coefficients for $\Tilde{X}$.

Under the assumption that the conditional median of $\epsilon$ given $X$ and $V$ is zero, i.e., $\operatorname{Med}(\epsilon|X,V)=0$, krief2014integrated proposed the integrated kernel-weighted smoothed maximum score (IKWSMS) estimator for $\beta$ in a two-stage procedure. Let $ G(\cdot) $ be a continuous function that corresponds to the integral of a $r$-th order kernel function for some integer $r\geq 4$.\footnote{See, for example, the definition of a $r$-th order kernel function in li2006nonparametric.} We assume that $\sup_{t\in\mathbb{R}} |G(t)|<M $ for some $ M<\infty $, $ \lim_{t\rightarrow-\infty}G(t)=0 $, and $ \lim_{t\rightarrow\infty}G(t)=1 $. Let $K(\cdot)$ be another kernel function. Suppose we observe an i.i.d.\ sample $\{(Y_i,X_i,V_i)\}_{i=1}^n$ from the distribution of $(Y,X,V)$.

In the first stage, we fix a $v$ in the support $\mathcal{V}$ of $V$ and estimate $\beta$ for $v$ by

equation[equation omitted — 221 chars of source]

where $W_i^T=(1,\widetilde{X}_i^T)$, $e_{d}=[O,I_{d-1}]$ is a $(d-1)\times d$ matrix with the first column being a zero vector, $I_{d-1}$ is the $(d-1)\times (d-1)$ identity matrix, $\Theta\subset \mathbb{R}^d$ is some compact set, and $h$ and $h_{v}$ are the smoothing parameters that converge to $0$ as $n\rightarrow\infty$. In this way, we obtain the estimator $\beta_{n}(v)$ for every $v\in\mathcal{V}$. In the second stage, the IKWSMS estimator, denoted by $\widehat{\beta}$, is defined as

equation[equation omitted — 98 chars of source]

where $\tau(\cdot)$ is a known weighting function on the real line with compact support that satisfies

equation*[equation* omitted — 50 chars of source]

For each $v$, define $L(v)=X^T\beta_X+\phi(v)$. Let $F_{\epsilon|W,L(v),V}$, $f_{V|W,L(v)}$, and $f_{L(v)|W}$ be the conditional cumulative distribution function (CDF) of $\epsilon$ given $(W,L(v),V)$, the conditional probability density funciton (PDF) of $V$ given $(W,L(v))$, and the conditional PDF of $L(v)$ given $W$, respectively. Then we define the random elements

equation*[equation* omitted — 102 chars of source]

and $T_{W}^{(1)}(l,v)=\partial T_{W}(l,v)/\partial l$ whenever these quantities exist. Also, we define $Z=X^T\beta_X+\phi(V)$. Then we let $F_{\epsilon|Z,W,V}$ and $f_{Z|W,V}$ be the conditional CDF of $\epsilon$ given $(Z,W,V)$ and the condition PDF of $Z$ given $(W,V)$, respectively. We define the random elements

equation*[equation* omitted — 126 chars of source]

and

equation*[equation* omitted — 94 chars of source]

for $j=1,\ldots,r$, provided these derivatives exist. Furthermore, for every $v$ such that all the relevant elements exist, we define

equation*[equation* omitted — 87 chars of source]
align*[align* omitted — 188 chars of source]

and

equation*[equation* omitted — 122 chars of source]

Under certain conditions, krief2014integrated showed that

equation[equation omitted — 105 chars of source]

where $\lambda=\lim_{n\rightarrow\infty}\sqrt{nh}h^{r}<\infty$.

It has been shown in horowitz2002bootstrap and cao2021smoothed that hypothesis testing of the SMS estimator based on asymptotic critical values may not have good finite sample performances. This is because in nonparametric and semiparametric estimations, large samples may be required to achieve reasonable agreement between the finite sample properties and the results of asymptotic theory. We next suggest using a bootstrap method for inference that could provide refinements to the empirical size of the test in finite samples.

For inference, we first consider testing the hypothesis

align[align omitted — 116 chars of source]

where $\beta_{j}$ denotes the $j$th component of $\beta$ and $\beta_{0j}$ is some prespecified value. To avoid computing the complicated bias term when conducting the test, we follow horowitz2002bootstrap and use the undersmoothed bandwidth $h$ such that $h$ converges to zero sufficiently fast so that $\lambda=0$. Thus, the test statistic for the $H_0$ in (ref) can be constructed as

equation[equation omitted — 132 chars of source]

where $\widehat{\beta}_j$ denotes the $j$th component of $\widehat{\beta}$ and $\widehat{\Omega}_{j}$ is the $(j,j)$ component of $\widehat{\Omega}$, a consistent estimator of $\Omega$. Based on Lemma 2 and the proof of Theorem 1 in krief2014integrated, such an estimator may be constructed as

equation*[equation* omitted — 222 chars of source]

where

equation[equation omitted — 234 chars of source]
equation*[equation* omitted — 69 chars of source]

and $H(V_{j})$ is the Hessian matrix of the objective function of $\theta$ in ((ref)) evaluated at $\widehat{\theta}(V_j)$. Algorithm (ref) illustrates the bootstrap procedure for the test.

algorithm[algorithm omitted — 1,219 chars of source]

We are also interested in a more general hypothesis

align[align omitted — 113 chars of source]

where $\mathcal{F}$ can be a vector valued function with continuous first derivatives. One special case is that $\mathcal{F}(b)=Rb$ with $b\in\mathbb{R}^{d-1}$ for some suitable matrix $R$ which incorporates the case where $H_{0}: \beta_{j}=\beta_{0j}$. Let $\mathcal{F}'(b)=\partial\mathcal{F}(b)/\partial b^T$ for all suitable $b\in\mathbb{R}^{d-1}$. Then we define the test statistic for the $H_0$ in (ref) as

align[align omitted — 203 chars of source]

Algorithm (ref) illustrates the bootstrap procedure for the test.

algorithm[algorithm omitted — 1,369 chars of source]

The asymptotic properties of the bootstrap tests in Algorithms (ref) and (ref) may be proved analogously to Theorems 4.3 in horowitz2002bootstrap. To avoid theoretical complications, we omit the proofs in the paper and show the finite sample properties of the tests in Monte Carlo simulations.

Monte Carlo Simulations

In this section, we present the Monte Carlo investigation of the finite sample performance of the proposed method. We design simulations based on those in horowitz2002bootstrap and krief2014integrated. Specifically, we consider estimating $\beta$ in the following model:

equation*[equation* omitted — 72 chars of source]

where $\beta=1$, $\phi(v)=\cos(2\pi v)$ for every $v$, $V=\Phi(\xi)$ with $\Phi(\cdot)$ being the cumulative distribution function of the standard normal random variable, and $(X_{1},X_{2},\xi)$ is a standard normal triplet with correlation coefficient $0.2$. We consider five distributions for the error $\epsilon$ as follows:

enumerate• UN: $\epsilon\sim \mathrm{Unif}(-\sqrt{3},\sqrt{3})$ • NR: $\epsilon\sim \mathcal{N}(0,1)$ • T3: $\epsilon\sim t_3$ (Student's $ t $ distribution with 3 degrees of freedom) • LG: $\epsilon\sim \mathrm{Logistic}(0,\pi/3)$ (logistic distribution) • HE: $\epsilon=(1+X_{1}^{2}+X_{2}^{2}+V^{2})e$, where $e\sim \mathrm{Logistic}(0,\pi/3)$ and $e\perp (X_1,X_2,V)$

In all designs, $\epsilon$ is normalized with variance $1$. Every Monte Carlo experiment consists of $500$ replications. The warp-speed method of giacomini2013warp is employed to expedite the simulations. Similar to krief2014integrated, the smoothing function for the indicator function is

equation*[equation* omitted — 156 chars of source]

The derivative of $G$ is a kernel function with order $r=4$ muller1984smooth. The first-stage kernel function $K$ is

equation*[equation* omitted — 82 chars of source]

which is of order $8$ pagan1999nonparametric. The bandwidth selection method follows krief2014integrated (Silverman-like rule of thumb in silverman1986density). Specifically, for every $v$, we let $h=1.22 R_{L}n^{-1/(2r+1)}$ and $h_{v}=0.8R_{V}n^{-3/25}$, where $R_{L}$ and $R_{V}$ denote the interquartile ranges between the first and third quartiles of $\{X_{i1}+(1,X_{i2})\widehat{\theta}(v)\}_{i=1}^{n}$ and $\{V_{i}\}_{i=1}^{n}$, respectively, and $\widehat\theta(v)$ is obtained with $(h_0,h_{v0})=(n^{-1/(2r+1)},n^{-3/25})$ by

equation[equation omitted — 230 chars of source]

The weighting function $\tau(v)=1.25\cdot\mathds{1}(0.1\leq v \leq 0.9)$.

First, we run simulations to show the size control of the test. Both the asymptotic and bootstrap critical values are used to compute the empirical size of the $t$ test for $H_{0}: \beta=1$. The sample size $n$ is set to $1000$. The results of the experiments are shown in Table (ref). In contrast to the results obtained from using the asymptotic critical values, the empirical sizes obtained from using the bootstrap critical values are well controlled. Furthermore, the empirical sizes obtained from using the bootstrap critical values show little sensitivity for undersmoothed bandwidths $0.75h$ and $0.5h$. They are all well controlled for undersmoothed bandwidths.

table[table omitted — 2,212 chars of source]

We also investigate the empirical power of the test. The sample sizes we consider are $n\in\{250,500,1000\}$. Table (ref) reports the rejection rates of the $t$ test for the false hypothesis $H_{0}: \beta=0$. The empirical powers obtained from using the asymptotic critical values exceed those obtained from using the bootstrap critical values. This is natural because of the upward size distortions arising from using the asymptotic critical values. But we note that the differences are small. Furthermore, the empirical power of the proposed bootstrap test increases as the sample size increases, which demonstrates the good finite sample power property of the proposed test.

table[table omitted — 1,956 chars of source]