Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
96,827 characters · 13 sections · 89 citation commands
Bootstraps for Dynamic Panel Threshold Models
\onehalfspacing
\sloppy
Threshold regression models are widely used in empirical research, and their usefulness has grown substantially with extenstions to the panel data settings. Estimation and inference methods for the threshold model in non-dynamic panels were developed by hansen_threshold_1999 and wang2015fixed. Dynamic panel threshold models were considered by seo_dynamic_2016, which proposes the generalized method of moments (GMM) estimation by generalizing the arellano_tests_1991 dynamic panel estimator. More recently, a latent group structure in the parameters of the panel threshold model was investigated by miao2020panel.
Applications of the panel threshold models cover numerous topics in economics. The effect of debt on economic growth is a well-known example that has been analyzed using panel threshold models, e.g., adam_fiscal_2005, cecchetti_real_2011, and chudik_is_2017. Another example is the threshold effect of inflation on economic growth such as the works by khan_threshold_2001, rousseau_inflation_2002, bick_2010, and kremer_inflation_2013. The benefit of foreign direct investment to productivity growth that depends on the regime determined by absorptive capacity is studied by girma_absorptive_2005 using firm-level panel data.
In empirical applications of threshold regression models, inference is usually performed after imposing an assumption about whether the model is continuous or not. Continuous threshold models that have kinks at the tipping points have received active research attention, e.g., Hansen_2017, seo_estimation_2019 and yang_panel_2020. In the literature, kink threshold models are analyzed for estimators that impose the continuity restriction as in chan_tsay_1998, Hansen_2017, and zhang_2017. On the other hand, unrestricted estimators are commonly used for discontinuous threshold models as in hansen_sample_2000. However, hidalgo_robust_2019 showed that the unrestricted least squares estimator possesses a different asymptotic property in the absence of discontinuity. Specifically, while the unrestricted model is not misspecified under continuity, failing to impose the restriction results in incorrect inference without proper care.
In the empirical literature, there has been mixed use of kink/discontinuous threshold models without much consideration of a possible specification error. Among the empirical examples referred to previously, khan_threshold_2001 use a continuous threshold model and impose continuity on their estimation procedure. They claim that the continuous model is desirable to prevent small changes in inflation rate from yielding different impacts around the threshold level. On the other hand, bick_2010 claims that the discontinuous threshold model is more appropriate for the same research question since overlooking a regime-dependent intercept can result in omitted variable bias. However, both of them do not provide econometric evidence that supports their choice of models.
For the dynamic panel threshold model, asymptotic normality of the GMM estimator is derived by seo_dynamic_2016 under the fixed $T$ scheme. However, the asymptotic normality is valid only for the discontinuous models since it requires a full rank condition on the Jacobian of the population moment, which is violated in continuous models. Although the continuity-restricted estimator described in seo_estimation_2019 is asymptotically normal, it may be problematic since empirical researchers often do not agree about whether their threshold models should have a kink or a jump at the threshold as in khan_threshold_2001 and bick_2010. Therefore, we focus on the unrestricted GMM estimator and bootstrap inference methods which do not require any pretest on continuity or prior knowledge about continuity of true models.
We first show that when the true model is continuous, the asymptotic normality of the unrestricted GMM estimator breaks down and the convergence rate of the threshold estimator becomes $n^{1/4}$-rate, which is slower than the standard $\sqrt{n}$-rate. Moreover, the standard nonparametric bootstrap is inconsistent in this case because the Jacobian from the bootstrap distribution does not degenerate fast enough due to the slow convergence rate of the threshold estimator.
We propose two different bootstrap methods to obtain confidence intervals for the parameters that are consistent regardless of whether the true model is continuous or not. One is for the threshold location, and the other is for the coefficients. The two bootstrap methods achieve consistency irrespective of the continuity of the model by adaptively setting the recentering parameter at the bootstrap for GMM introduced by hall_bootstrap_1996. This means that our bootstrap moment function achieves zero not at the sample estimator but at the parameter values that we propose. In the bootstrap for the threshold location, we employ a grid bootstrap to fix the recentering parameter. The grid bootstrap was originally proposed by hansen_grid_1999 for inference on an autoregressive parameter and applies test inversion. In case of the bootstrap for the coefficients, the recentering parameter is set to adjust the unrestricted estimator by a data driven criterion on the model's continuity. We also introduce a bootstrap test of model continuity.
Furthermore, we establish the uniform validity of the grid bootstrap for the unknown continuity (or discontinuity) of the threshold model. The importance of uniform validity is well recognized in the literature, notably in the works of mikusheva_2007, andrewsguggenberger2009, and romanoshaikh2012, among others, who have studied the uniformity of resampling procedures. In particular, mikusheva_2007 showed the uniform validity of the grid bootstrap for linear autoregressive models. Our work extends the advantage of the grid bootstrap to a broader class of nonstandard inference problems characterized by Jacobian degeneracy.
A set of Monte Carlo simulations demonstrate that the grid bootstrap performs favorably for inference on the threshold location, not only when the model is continuous but also when it includes a jump for various jump sizes. However, inference on the coefficients turns out to be more challenging. Our residual bootstrap confidence intervals for the coefficients, based on the lower and upper quantiles of bootstrap distributions, tend to exhibit undercoverage, even though they generally provide higher coverage rates than the standard nonparametric bootstrap.
We apply our inference methods to the dynamic firm investment model, whose static version was studied by fazzari_financing_1988 and hansen_threshold_1999 among others. It takes financial constraints into account via the threshold effect to determine a firm's investment decision.
In the literature, dovonon_testing_2013 and DovononHall2018 also deal with the degeneracy of the Jacobian in the context of the common conditional heteroskedasticity testing problem. In addition, a bootstrap based test for the common conditional heteroskedasticity feature was proposed by dovonon_bootstrapping_2017. However, their works do not deal with a discontinuous criterion function. Moreover, dovonon_testing_2013 and dovonon_bootstrapping_2017 study testing null hypothesis that always induces the degeneracy of the first-order derivative, while DovononHall2018 study the asymptotic distribution of an estimator when the degeneracy holds within the model. Therefore, they do not have to address the uncertainty associated with the potential degeneracy of the Jacobian.
Meanwhile, there is also a substantial body of literature on singularity-robust inference such as andrews_estimation_2012, andrews_gmm_2014 and han_estimation_2019, among many others. They are motivated by weak or non-identification problems, where models are not point identified. In contrast, we focus on an inference problem that does not involve identification failure even though the Jacobian of the moment function can become singular. andrews_identification-_2019 study more general singular cases than non-identification, but their approach requires differentiability of sample moments for subvector inference. Since our model exhibits discontinuity, the method of andrews_identification-_2019 is not applicable.
This paper is organized as follows. (ref) explains the dynamic panel threshold model. (ref) presents the asymptotic distribution theories of the estimators and test statistics related to the threshold location and continuity. (ref) proposes the bootstrap methods. (ref) reports Monte Carlo simulation results. (ref) contains an empirical application. (ref) concludes. The mathematical proofs and technical details are left to the Appendix.
We consider the dynamic panel threshold model,
where $1\leq i \leq n$, $1\leq t \leq T$, and $x_{it}\in\mathbb{R}^{p}$ is a regressor vector that includes $y_{i,t-1}$ and $q_{it}$. The threshold variable $q_{it}\in\mathbb{R}$ is allowed to be endogenous and is the last element of $x_{it}$.\footnote{Our analysis still holds if researchers have two sets of regressors $x_{1it}$ and $x_{2it}$ such that $y_{it}=x_{1it}'\beta+(1,x_{2it}')\delta1\{q_{it}>\gamma\}+\eta_{i}+\epsilon_{it}$ where $q_{it}$ is an element of $x_{2it}$. However, this paper sticks to the current form to keep the exposition simple.} We partition $x_{it} $ and such that $x_{it}=(\xi_{it}',q_{it})'\in\mathbb{R}^{p}$.
When $x_{it}$ consists of the lagged dependent variables, the model becomes the well-known self-exciting threshold autoregressive (TAR) model popularized by chan1985use. The static version where the lagged dependent variables are excluded from $x_{it}$ was considered by hansen_threshold_1999, while the current dynamic model was studied by seo_dynamic_2016.
The parameter $\gamma \in \Gamma $ denotes the threshold location, where $\Gamma$ is a compact set in $\mathbb{R}$, and $\alpha=(\beta',\delta')'\in A \subset \mathbb{R}^{2p+1}$ denotes the collection of coefficients. Let $\theta=(\alpha',\gamma)=(\beta',\delta',\gamma)'\in \Theta=A\bigtimes\Gamma $ denote the vector of all the parameters. The fixed effect $\eta_{i}$ is constant across time for each individual in the panel data. It is not identified but is eliminated after first-differencing for the GMM estimation. The idiosyncratic error $\epsilon_{it}$ is independent across individuals but can be dependent across time.
For the estimation, we use the GMM after the first-difference transformation
where
Let $z_{it}$ denote a set of instrumental variables at time $t$ such that $E[z_{it}\Delta\epsilon_{it}]$ becomes a zero vector, which may include lagged dependent variables $y_{it-2},...,y_{i1}$ and certain lagged variables of covariates $x_{it}$ and/or $q_{it}$, depending on the assumptions regarding exogeneity of those variables.
Then, we can define a vector of moment functions for the GMM estimation,
where $k\geq dim(\theta)=2p+2$ and $t_{0}\geq 2$ is the earliest period that the regressor and instrument can be defined. For example, $k=(T-1)(T-2)/2$ when $z_{it}=(y_{it-2},...,y_{i1})'$ and $t_{0}=3$. Denote the population moment by $g_{0}(\theta)=E[g_{i}(\theta)]$ and the sample moment by $$\bar{g}_{n}(\theta)=\frac{1}{n}\sum_{i=1}^{n}g_{i}(\theta).$$ We write $g_{i}$ instead of $g_{i}(\theta_{0})$ for simplicity of notations.
We consider the two-stage GMM estimation of the dynamic panel threshold model. In the first stage, we get an initial estimate by $\hat{\theta}_{(1)}=\arg\min_{\theta\in\Theta}\bar{g}_{n}(\theta)'\bar{g}_{n}(\theta)$ to compute a weight matrix
and obtain the second stage estimator
where $\hat{Q}_{n}(\theta)=\bar{g}_{n}(\theta)'W_{n}\bar{g}_{n}(\theta)$. seo_dynamic_2016 proposed averaging of a class of GMM estimators that are constructed from randomized first stage estimators. We do not pursue the averaging since our primary goal is the bootstrap inference.
In practice, the grid search algorithm is employed to compute the estimates. Note that when $\gamma$ is given, $\hat{\alpha}(\gamma)=\arg\min_{\alpha\in A}\hat{Q}_{n}(\alpha,\gamma)$ can be easily computed because the problem becomes the estimation of a linear dynamic panel model. Then, $\hat{\gamma}$ minimizes the profiled criterion $\tilde{Q}_{n}(\gamma)=\hat{Q}_{n}(\hat{\alpha}(\gamma),\gamma)$ over the grid of $\Gamma$.
Let $\theta_{0}=(\alpha_{0}',\gamma_{0})'=(\beta_{0}',\delta_{0}',\gamma_{0})'$ denote the true parameter value that lies in the interior of $\Theta$. For the point identification of $\theta_0$, $g_{0}(\theta)=0_{k}$ should hold if and only if $\theta=\theta_{0}$, where $0_{k}=(0,...,0)'\in\mathbb{R}^{k}$. Let
and $M_{i}(\gamma)=\left[
\right]$. Define $M_{0}(\gamma)=E[M_{i}(\gamma)]$, $M_{10}=E[M_{1i}]$, $M_{20}(\gamma)=E[M_{2i}(\gamma)]$, $\bar{M}_{n}(\gamma)=n^{-1}\sum_{i=1}^{n}M_{i}(\gamma)$, $\bar{M}_{1n}=n^{-1}\sum_{i=1}^{n}M_{1i}$, and $\bar{M}_{2n}(\gamma)=n^{-1}\sum_{i=1}^{n}M_{2i}(\gamma)$. We write $M_{0}$, $M_{20}$ and $\bar{M}_{n}$ instead of $M_{0}(\gamma_{0})$, $M_{20}(\gamma_{0})$ and $\bar{M}_{n}(\gamma_{0})$, respectively, for simplicity of notation. The identification condition is stated in (ref) that follows.
(ref) (i) is the identification condition for the coefficients once the true threshold location is identified. This means that instruments should be relevant to the first-differenced regressors appearing in $\eqref{eq:fd}$ when $\gamma=\gamma_{0}$.
(ref) (ii) is for the identification of the threshold location, which excludes the possibility of $\delta_{0}=0_{p+1}$. In the standard GMM problem, it is usually assumed that the Jacobian of $g_{0}(\theta)$ at $\theta_{0}$ is of full column rank for both the point identification and the asymptotic normality of the GMM estimator. The condition (ii) does not require the full rank condition on the Jacobian, which is related to the presence of a jump in the threshold model, and thus it generalizes the identification conditions in seo_dynamic_2016. When the model is continuous and has a kink at the threshold location, the last column of the Jacobian matrix, which is the first-order derivative with respect to $\gamma$ at the true parameter, becomes a zero vector. The exact formula for the Jacobian is given later in this section. This degeneracy does not violate the condition (ii), but it fails the asymptotic normality of the standard GMM estimator, which relies on the linearization of $g_{0}(\theta)$ near $\theta_{0}$ as in newey_chapter_1994.
To define the continuity, recall that $q_{it}$ is the last element of $x_{it}$ such that $x_{it}=(\xi_{it}',q_{it})'\in\mathbb{R}^{p}$. Accordingly, partition $\delta = (\delta_{1},\delta_{2}',\delta_{3})'$, where $\delta_{2}\in\mathbb{R}^{p-1}$ and $\delta_{1},\delta_{3}\in\mathbb{R}$, and $\delta_{0}=(\delta_{10},\delta_{20}',\delta_{30})'$. Hence, $\delta_{3}$ is the change in the coefficient of the threshold variable when the threshold variable surpasses the tipping point. Likewise, $\delta_{2}$ and $\delta_{1}$ are the changes in the coefficients for the other regressors, $\xi_{it}$, and the intercept, respectively. The continuity of the dynamic panel threshold model is formally given in (ref).
Note that this definition of continuity requires that $\delta_3 \neq 0$; otherwise, $\delta=0_{p+1}$.
The rank of the first-order derivative matrix, say $D_1$, of $g_{0}(\theta)$ at $\theta=\theta_{0}$ is crucial to the standard asymptotic normality of the GMM estimator. Let $G$ denote the first-order derivative of $g_{0}(\theta)$ with respect to $\gamma$ at $\theta=\theta_{0}$. Then,
where the conditional expectation $E_{t}[\cdot|q] = E[\cdot|q_{it}=q]$ and the density function $f_t(\cdot)$ of $q_{it}$ are assumed to exist. The derivation of $G$ is provided in the proof of (ref). Note that the first-order derivative of $g_{0}(\theta)$ with respect to $\alpha$ at $\theta=\theta_{0}$ is $M_{0}$. The linear independence of $G$ from the other columns in $D_1$ is required for the standard linear approximation $$g_{0}(\theta)\approx D_1 (\theta -\theta_0 ) = M_{0}(\alpha-\alpha_{0})+G(\gamma-\gamma_{0}).$$ Recall that the vector $G$ can be written as the product of the matrix $G_0$ and the vector $\delta_0$, (ref), and the first and last columns of $G_0$ are linearly dependent due to conditioning on $q_{it}=\gamma_0$ and $q_{it-1}=\gamma_{0}$. Then, the standard rank condition on the first derivative matrix $D_1$ can follow from a more primitive rank condition on $\left[
\right]$, which requires the linear independence of all columns in $M_0$ and all but the last column of $G_0$. Even if the primitive condition is met, however, the continuity restriction makes $G=0_{k}$ since $E_{s}[z_{it}(1,x_{is}')\delta_{0}|\gamma_{0}]=(\delta_{10}+\delta_{30}\gamma_{0})E_{s}[z_{it}|\gamma_{0}]=0$ for $s=t-1,t$, which leads to degeneracy of $D_{1}$.
When the rank condition fails due to the continuity, the expansion becomes $$g_{0}(\theta)\approx M_{0}(\alpha-\alpha_{0})+H(\gamma-\gamma_{0})^{2},$$ where
The detailed derivation is given in the proof of (ref). It is worth noting that $H$ is identical to the first column of $G_0$ up to a constant multiple. Then, the rank condition on $\left[
\right]$ is implied by the rank condition on $\left[
\right]$. Thus, the rank condition on $\left[
\right]$ can be viewed as a sufficient condition for both Assumptions \ref{ass:loc} and \ref{ass:locDis} in the next section, apart from the continuity restriction on $\theta$. Next section formalizes this discussion and presents the asymptotic distribution of the GMM estimator $\hat{\theta}$ under the continuity.
This section considers the asymptotic analysis when $T$ is fixed, the data are independent and identically distributed across $i$, and $n\rightarrow\infty$. Specifically, the data for each individual $i$ is determined by the realization of $\{(z_{it},x_{it},\epsilon_{it})_{t=1}^{T},y_{i0},\eta_{i}\}$, where $y_{i0}$ denotes an initial value. We make the following assumptions.
Assumptions (ref) and (ref) are similar to Assumptions 1 and 2 in seo_dynamic_2016 except for the differentiability conditions in (ref) which allow the second-order derivative of the population moment to be defined. Since the regressors include lagged dependent variables, (ref) requires the individual fixed effects and initial values to have finite fourth moments, too. The assumption also includes the conditions in (ref). (ref) is a rank condition for a nondegenerate asymptotic distribution when the underlying model is continuous. This condition may be viewed as less restrictive than the standard rank assumption as discussed in the previous section where $G$ and $H$ are defined. For easy reference, we restate the standard full rank assumption for the asymptotic normality of the GMM estimator for the discontinuous threshold regression below.
In a simple model, where $y_{it}=x_{it}'\beta+(\delta_{1}+\delta_{3}q_{it})1\{q_{it}>\gamma\} +\eta_{i}+\epsilon_{it}$, both Assumptions (ref) and (ref) require $\left[
\right]$ to have full rank, where $G_{01}$ is the first column of $G_{0}$ in \eqref{eq:firstderivative}, because $G=(\delta_{10}+\delta_{30}\gamma_{0})G_{01}$ while $H={\delta_{30}}G_{01}/2$.
(ref) below establishes the asymptotic distribution of the GMM estimator when the dynamic panel threshold model is continuous.
We observe that the convergence rate of $\hat{\gamma}$ is $n^{1/4}$, which is slower than the standard $\sqrt{n}$-rate. Meanwhile, seo_dynamic_2016 show the $\sqrt{n}$-convergence rate for $\hat{\gamma}$ when the model is discontinuous. Intuitively, it would be more difficult to detect the precise threshold location when there is a kink than when there is a jump at the tipping point. More technically, when the threshold model is discontinuous and the Jacobian is not singular, the limit of the GMM objective function admits a quadratic approximation with respect to $\gamma$ at the true value, while the limit admits a quartic approximation for the continuous model. Hence, the limit objective function becomes flatter in $\gamma$ at the true value resulting in the slower convergence rate. On the other hand, hidalgo_robust_2019 showed that the least squares criterion converges to a limit which is quadratic near the true $\gamma$ if the model is continuous and has a kink otherwise.
Moreover, we can observe that the asymptotic distribution of $\hat{\alpha}$ is also shifting to a non-normal distribution. Hence, standard inference methods based on the asymptotic normality become invalid for the continuous dynamic panel threshold model.
The asymptotic distribution of the GMM estimator is identical to the distribution reported in Theorem 1 (b) in DovononHall2018, which studies a smooth GMM problem with the degeneracy of the Jacobian. (ref) shows that even though the criterion of our threshold model is discontinuous with respect to the parameter $\gamma$, the same asymptotic distribution as that of DovononHall2018 appears. Meanwhile, dovonon_bootstrapping_2017 show that the standard nonparametric bootstrap becomes invalid when the Jacobian degenerates. To address this issue, we propose different bootstrap methods in (ref) for inference of the parameters.
The censored normal distribution also appears in andrews2002generalized which studies the estimation of a parameter on a boundary. Heuristically, because our analysis depends on the second-order derivative of $\gamma$ for the local polynomial expansion of $g_{0}(\theta)$ near $\theta_{0}$, only the asymptotic distribution of $(\hat{\gamma}-\gamma_{0})^{2}$ can be derived. Since $(\hat{\gamma}-\gamma_{0})^{2}$ should be nonnegative, the asymptotic censored normal distribution appears as in andrews2002generalized.
The asymptotic distribution in (ref) can be used for parameter inference when the true model is continuous, but the estimator is obtained without imposing the continuity restriction. As discussed in seo_dynamic_2016, $M_{0}$ and $\Omega$ can be consistently estimated, while $H$ can be nonparametrically estimated similarly to $G$. Then, it is straightforward to simulate the limit distribution of (ref) by generating random numbers for $U$ and $V$. However, there are several drawbacks to that approach, and hence we do not recommend it. First, empirical researchers might construct confidence intervals based on (ref) when they cannot reject the continuity. However, Leeb_Potscher_2005 show that confidence intervals after model selection are subject to size-distortion. Second, even if the true model is known to be continuous, the continuity-restricted estimator explained in seo_estimation_2019 is more efficient and asymptotically normal. Therefore, using the continuity-restricted estimator for estimation and inference is preferable. Finally, the nonparametric estimation of $H$ requires a tuning parameter and has a slower convergence rate.
seo_dynamic_2016 derived the asymptotic distribution of the GMM estimator and proposed an inference method when the underlying model is discontinuous. When the true model is discontinuous and Assumptions (ref), (ref), and (ref) hold,
$\Omega$ can be estimated by $\hat{\Omega}=\frac{1}{n}\sum_{i=1}^{n}[g_{i}(\hat{\theta})g_{i}(\hat{\theta})']-\bar{g}_{n}(\hat{\theta})\bar{g}_{n}(\hat{\theta})'$. Note that $D_{1}=\left[
\right]$, and $M_{0}$ can be estimated by $\bar{M}_{n}(\hat{\gamma})$, while the estimation of $G$ involves nonparametric estimation of the conditional means and densities. See section 4 of \cite{seo_dynamic_2016} for more details. Note that $(\hat{D}_{1}\hat{\Omega}^{-1}\hat{D}_{1})^{-1}$ diverges when the model is continuous since the last column of $\hat{D}_{1}$ converges to a zero vector when it is consistent. This paper does not study the behavior of the asymptotic confidence intervals when the true model is continuous.
Since the asymptotic distribution of the threshold estimator is not standard, we consider the GMM distance test introduced by newey_hypothesis_1987 for a hypothesis on the location of the threshold. Let the test statistic for the threshold location at $\gamma$ be \[\mathcal{D}_{n}(\gamma)=n(\min_{\alpha\in A}\hat{Q}_{n}(\alpha,\gamma)-\hat{Q}_{n}(\hat{\theta})),\] and let $\chi^{2}_{1}$ denote the chi-square distribution with 1 degree of freedom.
(ref) (i) presents the asymptotic distribution of the distance statistic under the continuity. Due to censoring, the asymptotic distribution becomes a mixture of the $\chi^{2}_{1}$ distribution with weight 1/2 and zero with weight 1/2.
Meanwhile, the chi-square limit in (ref) (ii) extends newey_hypothesis_1987 for a discontinuous moment function. seo_dynamic_2016 did not study the distance statistic.
(ref) (iii) shows that the GMM distance test for the threshold location is consistent. It also serves as the consistency of a bootstrap test together with (ref) since the bootstrap statistic is stochastically bounded whether or not the threshold location is true.
Since the limit distribution depends on the continuity of the model, we introduce a bootstrap in (ref), which is valid regardless of the model continuity. Furthermore, (ref) establishes the uniform validity of the bootstrap inference for the threshold location under some simplifying assumptions.
We propose a test for the continuity of the threshold model, similar to the approach used by Gonzalo_Wolf_TAR or Hidalgo_et_at_2022 in the threshold regression literature. While empirical researchers may employ the test to select a model, we utilize the test to modify the standard nonparametric bootstrap to make the bootstrap valid irrespective of the model continuity. Details of the use of the continuity test statistic in the bootstrap method are explained in (ref).
The continuity hypothesis is a joint hypothesis. We employ the GMM distance test. Let $\tilde{\theta}=\arg\min_{\theta\in\Theta_{c}} \hat{Q}_{n}(\theta)$ be the continuity-restricted estimator. The GMM distance test statistic is \[\mathcal{T}_{n}=n(\hat{Q}_{n}(\tilde{\theta})-\hat{Q}_{n}(\hat{\theta})).\]
While the limit distribution in (ref) (i) is non-standard, it can be simulated to obtain critical values for the test using consistent plug-in sample analogue estimators, e.g., $\hat{\Omega}=\frac{1}{n}\sum_{i=1}^{n}[g_{i}(\hat{\theta})g_{i}(\hat{\theta})']-\bar{g}_{n}(\hat{\theta})\bar{g}_{n}(\hat{\theta})'$, $\hat{M}_{1}=\bar{M}_{1n}$, $\hat{M}_{2}=\bar{M}_{2n}(\hat{\gamma})$, etc. Another way to obtain the critical values is via a bootstrap method, which is introduced in (ref).
(ref) (ii) shows that the continuity test is consistent. It also implies the consistency of the bootstrap test together with (ref), which shows that the bootstrap test statistic is stochastically bounded even when the true model is not continuous. The divergence rate of $\mathcal{T}_{n}$, which is faster than $n^{m}$ for any $0\leq m<1$, is exploited to modify the standard nonparametric bootstrap for the coefficients as detailed in (ref).
As usual, the superscript “*” denotes the bootstrap quantities or the convergence of bootstrap statistics under the bootstrap probability law conditional on the original sample. For example, $E^{*}$ denotes the expectation with respect to the bootstrap probability law conditional on the data. “$\xrightarrow{d^{*}}$, in $P$” denotes the distributional convergence of bootstrap statistics under the bootstrap probability law with probability approaching one. We write “$\nu_{n}^{*}=O_{p}^{*}(1)$, in $P$” if a sequence $\nu_{n}^{*}$ is stochastically bounded under the bootstrap probability law with probability approaching one. More details are written in (ref). Let $\widehat{F}^{*-1}_{n}(\varphi;S^{*})$ denote the empirical $\varphi$ quantile of a bootstrap statistic $S^{*}$.
This section introduces three different bootstrap schemes. The first bootstrap is for constructing bootstrap confidence interval(CI)s for the threshold, while the second bootstrap is for constructing bootstrap CIs for the coefficients. Both methods aim to provide valid inferences, regardless of whether the model is continuous or not. The third bootstrap is for testing continuity of the threshold model. The three bootstrap methods can be represented by means of (ref) with suitable choices of $\theta_0^{*}=(\beta_{0}^{*\prime},\delta_{0}^{*\prime},\gamma_{0}^{*})'$.
In step 1, we resample the regressors, the instruments, and the residuals jointly to maintain the dependence among them, unlike in the usual residual bootstrap. See e.g., giannerini2024validity for the description of the standard residual bootstrap, which resamples the residuals only, and the wild bootstrap for the testing of linearity in the threshold regression. There could be other ways of resampling not mentioned here and we do not attempt to decide which is the best here.
The parameter $\theta_{0}^{*}$ is used in step 2 of (ref) to generate the dependent variables in the bootstrap samples. In step 4, recentering of the bootstrap sample moment is done by subtracting $\bar{g}_{n}(\hat{\theta})=(\frac{1}{n}\sum_{i=1}^{n}z_{it_{0}}'\widehat{\Delta\epsilon}_{it_{0}},...,\frac{1}{n}\sum_{i=1}^{n}z_{iT}'\widehat{\Delta\epsilon}_{iT})'$. Note that the expectation of $\bar{g}_{n}^{*}(\theta)$ by the bootstrap probability law conditional on the data becomes zero when $\theta=\theta_{0}^{*}$ due to the recentering, which can be easily checked from the following equations: $g_{it}^{*}(\theta_{0}^{*})=z_{it}^{*}(\Delta y_{it}^{*}-\Delta x_{it}^{*}\beta_{0}^{*}-1_{it}(\gamma_{0}^{*})'X_{it}^{*}\delta_{0}^{*})=z_{it}^{*}\widehat{\Delta\epsilon}_{it}^{*}$ and $E^{*}[g_{it}^{*}(\theta_{0}^{*})] = n^{-1}\sum_{i=1}^n z_{it} \widehat{\Delta\epsilon}_{it}$ for $t=t_{0},...,T$.
A different choice of $\theta_{0}^{*}$ leads to a different bootstrap. For example, if $\theta_{0}^{*}=\hat{\theta}$, then the bootstrap becomes the standard nonparametric bootstrap in hall_bootstrap_1996 because $\Delta y_{it}^{*}=\Delta y_{i^{*}t}$ holds true for $i=1,...,n$ and $t=t_{0},...,T$ in step 2. Note that, for $\theta_{0}^{*}$ not equal to $\hat{\theta}$, step 2 of (ref) generates $\Delta y_{it}^{*}$'s that are generally different from $\Delta y_{i^{*}t}$'s. The following subsections detail three different choices of $\theta_{0}^{*}$ for three different inference problems.
To construct CIs for the threshold location, we propose to employ the grid bootstrap method introduced by hansen_grid_1999 for autoregressive models. Let $\Gamma_{n}=\{\gamma_{\ell}\in\Gamma:\ell=1,...,L\}$ be a grid of the candidate thresholds. The grid bootstrap constructs the confidence set by inverting the bootstrap threshold location tests over $\Gamma_{n}$. Specifically, a sequence of hypothesis tests for the hypothesized threshold locations in $\Gamma_{n}$ are performed by the bootstrap that imposes the null to generate bootstrap samples.
The null imposed bootstrap at a point $\gamma_{\ell}\in \Gamma_{n}$ can be implemented by setting $\theta_{0}^{*}=(\hat{\alpha}(\gamma_{\ell})',\gamma_{\ell})'$ in (ref), and the bootstrap test statistic is \[\mathcal{D}_{n}^{*}(\gamma_{\ell})=n(\min_{\alpha\in A}\hat{Q}_{n}^{*}(\alpha,\gamma_{\ell})-\min_{\theta\in\Theta}\hat{Q}_{n}^{*}(\theta)).\] The null hypothesis $\mathcal{H}_{0}:\gamma=\gamma_{\ell}$ is rejected at size $\tau$ if $\mathcal{D}_{n}(\gamma_{\ell})> \widehat{F}^{*-1}_{n}(1-\tau;\mathcal{D}_{n}^{*}(\gamma_{\ell}))$. Consequently, after running the null imposed bootstrap for each point in $\Gamma_{n}$, we can construct the $100(1-\tau)$% confidence set of $\gamma$ by
Note that the confidence set is not necessarily a connected set, even though researchers can convexify the set to get a connected CI. The CI does not become an empty set because $\mathcal{D}_{n}(\hat{\gamma})=0$ while $\mathcal{D}_{n}^{*}(\hat{\gamma})\geq 0$. The consistency of the grid bootstrap method is implied by (ref) that follows.
(ref) (i) and (ii) show that the limit distribution of the bootstrap test statistic, conditional on the data, is identical to that of the sample test statistic regardless of the continuity of the true model. Therefore, the CI for the threshold location by the grid bootstrap, (ref), achieves an exact coverage rate for both continuous and discontinuous models asymptotically. Specifically, $\lim_{n\rightarrow\infty}P(\gamma_{0}\in CI_{n,1-\tau}^{grid})=1-\tau$ for both cases (i) and (ii). (ref) (iii) says that the bootstrap test statistic is still stochastically bounded, conditionally on the data, under the alternative. As (ref) (iii) shows that the sample test statistic is stochastically unbounded under the alternative, the grid bootstrap CI has power against fixed alternatives.
We extend (ref) to the uniform validity of the grid bootstrap, which is important for good finite sample performance when the model is nearly continuous. We establish the uniform validity for the following simplified specification for analytical tractability:
where $\theta=(\beta',\delta',\gamma)'$ and $\delta=(\delta_{1},\delta_{3})'$ in this subsection.
This section briefly states the uniformity result of the grid bootstrap and gives a heuristic justification. Our derivation follows andrews_cheng_guggenberger_2020. It is highly complicated and involves more technical conditions, which are stated in (ref).
Specifically, we establish in (ref) that
where $P_{\phi}$ is the probability law when the model is specified by $\phi=(\theta,F)$ and $F$ is the distribution of $\{\eta_{i},y_{i0},(z_{it},x_{it},\epsilon_{it})_{t=1}^{T}\}$. The collection of probabilistic models $\Phi_{0}$ includes both continuous and discontinuous threshold models. More detailed discussions of technical assumptions about $\Phi_{0}$ are given in (ref).
For the uniformity analysis, we need to consider drifting sequences of true parameters $\phi_{0n}=(\theta_{0n},F_{0n})$ such that $\theta_{0n}\rightarrow\theta_{0,\infty}$ and $F_{0n}\rightarrow F_{0,\infty}$. Here, the distance between $F_{0n}$ and $F_{0,\infty}$ is induced by a specific choice of norm that is explained in (ref). To show the uniform validity of the grid bootstrap CI, we need to verify that the limit distribution of $\mathcal{D}_{n}^{*}(\gamma_{0n})$ conditional on the data is identical to the limit distribution of $\mathcal{D}_{n}(\gamma_{0n})$ under all the above drifting sequences of models. Our analysis finds that the limit distribution of the threshold location test statistic under the true null, i.e., the limit distribution of $\mathcal{D}_{n}(\gamma_{0n})$, is determined by $\zeta=\lim_{n\rightarrow\infty} n^{1/4}(\delta_{10n}+\delta_{30n}\gamma_{0n})$; see (ref) for details. When $\zeta=0$, the limit distribution of $\mathcal{D}_{n}(\gamma_{0n})$ is as described in (ref) (i). In contrast, when $|\zeta|=\infty$, the limit distribution is the $\chi^{2}_{1}$-distribution as in (ref) (ii). When $\zeta$ is finite and nonzero, then $\mathcal{D}_{n}(\gamma_{0n})$ has a nonstandard limit distribution that depends on $\zeta$.
Therefore, if $\theta_{0n}^{*}$ comprises a sequence of true parameters for a bootstrap scheme, then $n^{1/4}(\delta_{10n}^{*}+\delta_{30n}^{*}\gamma_{0n}^{*})$ should consistently estimate $\zeta$ for the bootstrap statistics to exhibit the same asymptotic behavior as the sample statistics.
Note that under the grid bootstrap scheme, the bootstrap test statistic $\mathcal{D}_{n}^{*}(\gamma_{0n})$ is drawn from the bootstrap that imposes the null threshold location $\gamma_{0n}$. The true parameter of the bootstrap data generating process (dgp) is $\theta_{0n}^{*}=(\hat{\alpha}_{n}(\gamma_{0n})',\gamma_{0n})'$. The restricted estimator satisfies $\|\hat{\alpha}(\gamma_{0n})-\alpha_{0n}\|=O_{p}(n^{-1/2})$, as the problem becomes estimating a standard linear dynamic panel model, and hence $n^{1/4}(\hat{\delta}_{1n}(\gamma_{0n})+\hat{\delta}_{3n}(\gamma_{0n})\gamma_{0n})=\zeta+o_{p}(1)$. Therefore, $\mathcal{D}_{n}^{*}(\gamma_{0n})$ conditionally converges to the limit distribution of $\mathcal{D}_{n}(\gamma_{0n})$, which leads to the uniform validity of the grid bootstrap confidence intervals. In contrast, $\hat{\theta}$ does not satisfy this property for some $\zeta$ and the bootstrap building on $\hat{\theta}$ is not uniformly valid.
The bootstrap CIs for the coefficients can be obtained by applying (ref) with $\theta_{0}^{*}$ set as
where $\tilde{\theta}=\arg\min_{\theta\in\Theta_{c}}\hat{Q}_{n}(\theta)$ is the continuity-restricted estimator. $\hat{C}$ is some estimated quantile, such as the $50$th percentile, of the limit distribution of the continuity test statistic $\mathcal{T}_{n}$ when the model is continuous. $\hat{C}$ can be obtained either by methods in (ref) or (ref). As long as $\hat{C}=O_{p}(1)$, the asymptotic validity of the residual bootstrap holds. Since $w_{n}=O_{p}(n^{-1/4})$ if the true model is continuous, and $w_{n}=1+o_p(1)$ if the model is discontinuous, the true parameter value for the bootstrap adapts to the model continuity.
After collecting the bootstrap estimators \[\hat{\theta}^{*}=(\hat{\alpha}^{*\prime},\hat{\gamma}^{*})'=\arg\min_{\theta\in\Theta}\hat{Q}_{n}^{*}(\theta),\] we can construct the CIs for the coefficients using the percentiles of either $|\hat{\alpha}_j^{*}-\alpha_{j0}^{*}|$ or $(\hat{\alpha}_j^{*}-\alpha_{j0}^{*})$. Here, $\hat{\alpha}^{*}_{j}$ and $\alpha^{*}_{j0}$ are the $j$th elements of $\hat{\alpha}^{*}$ and $\alpha^{*}_{0}$, respectively. The $100(1-\tau)$% CI for the $j$th element of the coefficients, $\alpha_{j}$, can be constructed by
or
which leads to a symmetric CI.
According to (ref) that follows, both CIs are asymptotically (pointwise) valid, and they should provide similar coverage rates close to the nominal rate for any fixed data generating process in large sample. However, our Monte Carlo experiments in (ref) show big differences in coverage rates between the two confidence intervals. Specifically, (ref) shows severe undercoverage while (ref) seems to provide much higher coverage rates. This phenomenon also appears for the nonparametric bootstrap. We provide further numerical investigation of the phenomenon in (ref), which suggests challenges for reliable bootstrap inference for the coefficients.
The asymptotic distributions of the bootstrap estimators in (ref), conditional on the data, match those of the sample estimators for both continuous and discontinuous cases. Therefore, the residual bootstrap CI becomes asymptotically valid in a pointwise sense, regardless of whether the model is continuous or discontinuous. {We acknowledge that (ref) does not guarantee the uniform validity of the bootstrap CI. The difficulty in establishing the uniform validity lies in analyzing asymptotic behaviors of $\mathcal{T}_{n}$ and $w_{n}$ for drifting sequences of the true models. $\mathcal{T}_{n}$ already exhibits an irregular limit distribution even in the pointwise setup, as shown in (ref) (i). This paper does not provide a theoretical analysis of whether the uniformity of the residual bootstrap can be achieved. Instead, we conduct Monte Carlo experiments for nearly continuous cases in (ref) and leaves theoretical work on the uniformity of the bootstrap method to future research.}
The key motivation for setting $\theta_{0}^{*}$, the true parameter of the bootstrap dgp, by (ref) is to make $\delta^{*}_{10}+\delta^{*}_{30}\gamma^{*}_{0}$ degenerate fast enough when the underlying model is continuous. The $n^{1/4}$ convergence rate of the unrestricted estimator $\hat{\gamma}$ to $\gamma_0$ is not sufficiently fast. To see this, let the first-derivative of the population moment with respect to $\gamma$ at $\theta$ be
for which we recall that $x_{it}=(\xi_{it}',q_{it})'$ and that $G(\theta_{0})=0_{k}$ under continuity. For the validity of a bootstrap method under continuity, the degeneracy of the Jacobian should be mimicked by the bootstrap dgp. In our residual bootstrap method, the Jacobian is $G(\theta_{0}^{*})=O_p(n^{-1/2})$. However, it is $G(\hat{\theta})=O_p(n^{-1/4})$ for the standard nonparametric bootstrap. This fails the standard nonparametric bootstrap. More formal treatment of the invalidity of the standard nonparametric bootstrap is given in (ref).
It is not difficult to check $G(\hat{\theta})=O_{p}(n^{-1/4})$ but not $o_{p}(n^{-1/4})$ under continuity, which is directly implied by $n^{1/4}(\hat{\delta}_{1}+\hat{\delta}_{3}\hat{\gamma})=O_{p}(1)$ but not $o_{p}(1)$ due to (ref). Meanwhile, in our residual bootstrap method, $\delta_{10}^{*}+\delta_{30}^{*}\gamma_{0}^{*}= w_{n}(\hat{\delta}_{1}+\hat{\delta}_{3}\hat{\gamma})+o_{p}(n^{-1/2})=O_{p}(n^{-1/2})$ and $\delta_{20}^{*}=w_{n}\hat{\delta}_{2}=O_{p}(n^{-3/4})$, which leads to $G(\theta_{0}^{*})=O_{p}(n^{-1/2})$. The exact formula for $\delta_{10}^{*}+\delta_{30}^{*}\gamma_{0}^{*}$ is provided in the comment after (ref).
According to the proof of (ref) in (ref), $(\delta_{10}^{*}+\delta_{30}^{*}\gamma_{0}^{*})=O_{p}(n^{-1/2})$ is sufficient for the first-order asymptotic validity when the true model is continuous. This requirement is explicitly stated in the conditions of (ref). While our choice of $n^{1/4}$ decay rate for $w_{n}$ guarantees this condition, it remains an open question whether there exists a rate of decay for $w_{n}$ that ensures uniform validity.
The idea of shrinking the first-order derivative in our bootstrap is closely related to other bootstrap methods developed for the case when asymptotic distributions of estimators are irregular. For example, chatterjee2011bootstrapping propose a bootstrap method for the lasso estimator, and cavaliere2022bootstrap study bootstrap inference on the boundary of a parameter space. Both papers set up the model where the problem appears if the true parameter value is zero, and they obtain true parameters of bootstrap dgps by thresholding unrestricted estimators, i.e., $\theta_{j0}^{*}=\hat{\theta}_{j}1\{|\hat{\theta}_{j}|>c_{n}\}$, where $c_{n}$ converges to zero in a proper rate.
The critical value for the continuity test introduced in (ref) can also be obtained by bootstrapping. Recall that $\tilde{\theta}=\arg\min_{\theta\in\Theta_{c}} \hat{Q}_{n}(\theta)$ is the continuity-restricted estimator. By setting $\theta_{0}^{*}=\tilde{\theta}$ in (ref), and collecting the bootstrap test statistic \[\mathcal{T}_{n}^{*}=n\left(\min_{\theta\in\Theta_{c}}\hat{Q}_{n}^{*}(\theta)-\min_{\theta\in\Theta}\hat{Q}_{n}^{*}(\theta)\right),\] we can get the critical value using the empirical quantile of $\mathcal{T}_{n}^{*}$. To run the bootstrap continuity test at size $\tau$, reject the continuity if $\mathcal{T}_{n}>\widehat{F}^{*-1}_{n}(1-\tau;\mathcal{T}_{n}^{*})$, where $\widehat{F}^{*-1}_{n}(1-\tau;\mathcal{T}_{n}^{*})$ is the empirical $(1-\tau)$ quantile of $\mathcal{T}_{n}^{*}$. The consistency of the bootstrap is implied by (ref) that follows.
(ref) (i) shows that the limit distribution of $\mathcal{T}_{n}^{*}$, conditional on the data, is identical to that of $\mathcal{T}_{n}$ under the null hypothesis. Moreover, (ref) (ii) says that $\mathcal{T}_{n}^{*}$ is still stochastically bounded, conditionally on the data, when the true model is discontinuous. As $\mathcal{T}_{n}$ is shown to be stochastically unbounded under the alternative, according to (ref) (ii), the bootstrap continuity test has power against fixed alternatives.
This section presents Monte Carlo simulations to investigate finite sample performances of our bootstrap methods. The data are generated by
with $\beta_{2}=0.6$, $\beta_{3}=1$, $\delta_{2}=0$, $\delta_{3}=2$, $\gamma=0.25$, $\sigma=0.5$, $\rho=0.7$, and $\rho_{eu}=0.5$. Note that (ref) implies that the threshold variable is weakly exogenous. That is, $E[e_{it}|q_{is}]=0$ for $s\leq t$ while $E[e_{it}|q_{is}]\neq 0$ for $s\geq t+1$. Additional results when the threshold variable is weakly endogenous are also presented in (ref). This section focuses on comparing bootstrap methods, while results based on the asymptotic method by seo_dynamic_2016 are reported in (ref).
To investigate how coverage rates of CIs change depending on continuity, we try different values of $\delta_{1}\in\{-0.5, -0.4, -0.3, 0, 0.5\}$, which implies different degrees of (dis)continuity $\delta_{1}+\delta_{3}\gamma\in\{0,0.1,0.2,0.5,1\}$. If $\delta_{1}=-0.5$, then $\delta_{1}+\delta_{3}\gamma=0$ and the model is continuous. Otherwise, the model is discontinuous. As near continuous designs, we try $\delta_{1}+\delta_{3}\gamma=0.1, 0.2$ and check for any poor CI performance. We generate samples of size $n\in\{400, 800, 1600\}$ and $T=6$. The number of repetitions for the Monte Carlo simulations is 2000. We use $z_{it}=(y_{it-2},...,y_{i1},q_{it-1},...,q_{i1})'$ for $t=t_{0},\dots, T$ as instruments. Since $t_{0}=3$, the total number of the instruments becomes 24. The number of bootstrap repetitions is set at 500 for each bootstrap method.
We begin with examining the finite sample coverage probabilities of bootstrap CIs for the threshold location. Specifically, the grid bootstrap CI (Grid-B) is compared with both percentile nonparametric bootstrap CI (NP-B) and symmetric percentile nonparametric bootstrap CI (NP-B(S)) that are defined as follows:
(ref) reports the coverage rates of 95% CIs for the threshold location. First, it shows that the bootstrap CI by NP-B is subject to severe undercoverage in all cases. This is the case even when $\delta_{1}+\delta_{3}\gamma=1$, despite the theoretical validity of NP-B when the model is discontinuous. Meanwhile, NP-B(S) exhibits extreme over-coverage in all cases. The large discrepancy between NP-B and NP-B(S) suggests that the distribution of the nonparametric bootstrap statistic $\hat{\gamma}^{*}-\hat{\gamma}$ poorly approximates that of $\hat{\gamma}-\gamma_{0}$, undermining its reliability for inference.
In contrast, (ref) shows that Grid-B provides more reasonable coverage rates. It seems that a larger jump yields coverage rates closer to the nominal level as a bigger jump is easier to detect. As expected from the uniform validity of Grid-B against near continuity, coverage rates remain valid for all the parameter values, if somewhat over-coveraged near continuity or under smaller sample sizes.
Compared to Grid-B, NP-B(S) exhibits higher coverage probabilities that are one or almost one for all cases. It indicates that NP-B(S) CIs are overly wide and non-informative. To investigate this further, we examine some power properties as reported in (ref). It shows that the NP-B(S) based test for the threshold location is trivial for many parametrizations, specifically when the design is continuous or near-continuous. In contrast, the Grid-B test is more powerful, oftentime twice more powerful than the NP-B(S) test. We report test power instead of CI lengths because of the computational burden associated with Grid-B, which constructs CIs by test inversion.
Next, we examine the coverage probabilities of the regression coefficients using different bootstrap CIs. We first report results for percentile bootstrap CIs that use the lower and upper quantiles of the bootstrap distributions. (ref) reports the coverage rates of the percentile CIs using the residual bootstrap (R-B) defined as (ref), and the standard nonparametric bootstrap (NP-B) defined as
$\hat{C}$ in (ref) is set as the 50th percentile of the bootstrap distribution of the test statistic $\mathcal{T}_{n}$ under the null hypothesis that the model is continuous, using the bootstrap method explained in (ref) with 500 repetitions.
As in the threshold inference case, the percentile CIs for the coefficients constructed using NP-B exhibit undercoverage across all specifications and sample sizes. Even when $\delta_{1} + \delta_{3}\gamma = 1$, where the model is discontinuous and NP-B is theoretically valid, the undercoverage remains severe. Although R-B yields higher coverage rates than NP-B, they still fall short of the nominal 95% level. As reported in (ref), R-B results in wider average CI lengths compared to NP-B, partly accounting for its improved coverage. (ref) presents the results with a much larger sample size, $n=10000$, and $\delta_{1}+\delta_{3}\gamma\in\{0,1\}$. When $n=10000$, the coverage rates of R-B approach the nominal level, although undercoverage persists for some coefficients.
Finally, we report the coverage rates of symmetric percentile CIs for the coefficients that are constructed using the nonparametric bootstrap (NP-B(S)) defined as
and the residual bootstrap (R-B(S)) defined as (ref). Tables (ref) and (ref) show the coverage rates and the ratios of the average lengths of CIs by the two bootstrap methods.
When the symmetric percentile CIs are used for the coefficients, (ref) shows that the coverage rates increase, as also observed in (ref). However, R-B(S) yields lower coverage rates than NP-B(S) and even produces undercoverage for the dgp reported in (ref). Nevertheless, R-B(S) tends to return wider CIs than NP-B(S) according to (ref).
NP-B(S) may appear to be the most suitable method for inference on the coefficients, given its higher coverage rates and shorter average CI lengths. However, a more detailed numerical analysis in (ref) reveals an undesirable property of the nonparametric bootstrap: the conditional distribution of the bootstrap statistic $\sqrt{n}(\hat{\alpha}^{*}-\hat{\alpha})$ is not centered at zero. This misalignment also results in the unexpected relationship between the coverage and the average length of CIs in Tables (ref) and (ref), which is further illustrated in (ref). These findings highlight the difficulty of reliable inference for the coefficients $\beta$ and $\delta$. A more comprehensive theoretical and methodological investigation is needed to address these challenges in future research.
Our empirical example examines a firm's investment decision model that incorporates financial constraints, as in hansen_threshold_1999 and seo_dynamic_2016. In a perfect financial market, firms can borrow as much money as they need to finance their investment projects, regardless of their financial conditions. Therefore, the financial conditions of firms are irrelevant to their investment decisions. However, in an imperfect financial market, some firms may be restricted in their access to external financing. These firms are said to be financially constrained. Financially constrained firms are more sensitive to the availability of internal financing, as they cannot rely on external financing to fund their investment projects.
fazzari_financing_1988 argue that firms' investments are positively related to their cash flow if they are financially constrained, where those firms are identified by low dividend payments. hansen_threshold_1999 applies the threshold panel regression more systematically to show that a more positive relationship between investment and cash flow is present for firms with higher leverage.
Since there are multiple candidate measures of the financial constraint for the threshold variable, we compare the following three dynamic panel threshold models:
where $\xi_{it-1}=(I_{it-1},CF_{it},PPE_{it-1},ROA_{it-1})'$. Here, $I_{it}$ is investment, $CF_{it}$ is cash flow, $PPE_{it}$ is property, plant and equipment, and $ROA_{it}$ is return on assets. $I_{it}$, $CF_{it}$ and $PPE_{it}$ are normalized by total assets. We have two candidate threshold variables, $LEV_{it}$ and $TQ_{it}$, which are leverage and Tobin's Q, respectively. Choice of the regressors and threshold variables is based on previous works like hansen_threshold_1999 and lang1996leverage. Note that the regression model (ref) is nested within (ref) and it is closer to a continuous threshold model.
Unlike the previous works, we do not need to assume either continuity or discontinuity for valid inferences since the bootstrap methods in this paper are adaptive to each case. With an assumption that the regressors are predetermined, we use the variables dated one period before as instruments. Hence, the instruments include $I_{t-2}$, $CF_{t-1}$, $PPE_{t-2}$, $ROA_{t-2}$ added by $LEV_{t-2}$ or $TQ_{t-2}$ for each period.
We construct a balanced panel of 1459 U.S. firms, excluding finance and utility firms, from 2010 to 2019 available in Compustat. To deal with extreme values, we drop firms if any of their non-threshold variables' values fall within the top or bottom 0.5% tails. Moreover, we exclude firms whose Tobin's Q is larger than 5 for more than 5 years when the threshold variable is Tobin's Q, leaving 1222 firms in the sample. Meanwhile, STREBULAEV20131 claims that firms with large CEO ownership or CEO-friendly boards show persistent zero-leverage behavior. To prevent our threshold regression from capturing corporate governance characteristics rather than financial constraints, we exclude firms whose leverage is zero for more than half of the time periods when leverage is the threshold variable, leaving 1056 firms in the sample.
(ref) reports the estimates and 95% CIs for (ref) and (ref), and (ref) for (ref). (ref) visualizes how the grid bootstrap CIs are obtained. The CIs for the coefficients are constructed by using the percentiles obtained from the residual bootstrap, defined as (ref)\footnote{The symmetric percentile CIs via residual bootstrap that use the 0.95 quantiles of $|\hat{\alpha}_{j}^{*}-\alpha_{j0}^{*}|$'s return similar results, unlike in Monte Carlo results from (ref). We report them in (ref).}. $\hat{C}$ for the precentile bootstrap is set at the 50th percentile of the bootstrap statistic for the continuity test, explained in (ref). For the threshold locations, the CIs are obtained by the grid bootstrap with convexification. For the grid bootstrap, we make 500 bootstrap draws for each grid point. The grids of the threshold locations have 81 points from the 10th percentile to the 90th percentile of the threshold variables, and there are equal number of observations between two consecutive points. (ref) and (ref) also report the bootstrap p-values for the continuity and linearity tests by the bootstrap methods explained in (ref) and (ref), respectively. The null hypothesis of the linearity test is $\mathcal{H}_{0}:\delta=(0,...,0)'$, which implies no threshold effects.
We find supporting evidence for the presence of the threshold effect when the threshold variable is Tobin's Q, but the statistical evidence is not strong for the leverage threshold model. (ref) and (ref) report the bootstrap p-values at .135, .011, and .011, for specifications (ref) - (ref), respectively. The statistical evidence to reject the continuity is not trivial for all specifications and gets stronger when it is the restricted model using Tobin's Q. The estimated bootstrap p-values are .028 and .004 for the unrestricted and the restricted using Tobin's Q. Furthermore, the confidence interval for the threshold location is narrower for the restricted model (ref) than for the unrestricted model (ref).
A notable finding concerning the coefficients estimates is that the relationship between cash flow and investment is positive and has larger magnitude for the low Tobin's Q firms and the high leverage firms compared to their other respective regimes, although they are not statistically significant at 5% level. Even though the sign and magnitude of the estimates align with the observations by lang1996leverage and hansen_threshold_1999 that a firm is subject to financial constraints when its Tobin's Q is low or leverage is high, there is uncertainty in the interpretation of our results due to the lack of statistical significance.
Next, the autoregressive coefficient of the lagged investment is significant at 5% level in the low leverage regime and is larger than in the high leverage regime. This lends supporting evidence for the presence of asymmetric dynamics in investment, akin to the dynamics of leverage analyzed by dang_asymmetric_2012. In the meantime, we note that the autoregressive coefficients for the low and high leverage regimes in Column (a) are 0.778 and -0.154, respectively, which appear more extreme than findings of the literature where the estimates are between 0.1 and 0.5, e.g., blundell_investment_1992. The autoregressive coefficients in the Column (b) are more in line with these estimates. Since the changes of the estimated coefficients in Column (b) are moderate, we also estimate the restricted model (ref).
Turning to (ref), we observe that the differences between the coefficients of the two regimes become significant at 5% level, and the CI for the threshold location becomes narrower while the estimate of the threshold location remains close to the estimate under the unrestricted model. The autoregressive coefficient of the lagged investment and the sensitivity of investment to both cash flow and return on assets are all positive and significant. The effect of Tobin's Q is both positive and significant for both high and low Tobin's Q regimes, but it almost disappears once it surpasses the threshold location. This suggests that low Tobin's Q is related to low investment but higher Tobin's Q does not cause higher investment once it reaches some level.
This paper studies the asymptotic properties of the GMM estimator in dynamic panel threshold models, showing that the limiting distribution depends critically on whether the true model exhibits a kink or a jump at the threshold. We demonstrate that the standard nonparametric bootstrap is inconsistent when the true model has a kink. To address this, we propose alternative bootstrap procedures for constructing confidence intervals for the threshold location and the model coefficients, which are shown to be consistent regardless of the model's continuity. In particular, we establish that the grid bootstrap for the threshold parameter is uniformly valid. Monte Carlo simulations confirm that the grid bootstrap outperforms the standard bootstrap in finite samples.
Several directions remain for future research. Our simulation results reveal highly asymmetric bootstrap distributions for the coefficient estimates, which distort finite sample inference. This highlights the need for a more thorough theoretical understanding of the bootstrap's behavior. In particular, whether uniform validity of the bootstrap for coefficients is achievable remains an important open question. Extensions of our bootstrap algorithms to incorporate latent group structures, interactive fixed effects, or threshold indices, as studied in miao2020panel, miao2020panel_interactive, and seo2007smoothed, lee2021factor, respectively, would also be valuable.