Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
70,410 characters · 7 sections · 62 citation commands
Testing for a Threshold in Models with Endogenous Regressors
In the aftermath of the 2008 financial crisis, there has been a surge in the macroeconomic literature investigating whether the response of many key macroeconomic variables to monetary and fiscal policies depends on the state of the economy - see, among others, auerbachgorodnichenko2013, owyangetal2013, Caggiano:2015, Cug:2015, rameyzubairy2018, Alloza:2022 and Jo:2022 for fiscal policy examples, and Santoro:2014, Barnichon:2018, Jorda:2020, Alpanda:2021, Bruns:2021 and Klepacz for monetary policy examples. These papers model state dependence in various ways, including via threshold models, in which case the state dependence is typically driven by a particular variable such as the unemployment rate, interest rates, or credit conditions.
Threshold models were also widely used in economics to model unemployment, growth, bank profits, asset prices, exchange rates, and interest rates; see hansen2011 for a survey of economic applications. While threshold models with exogenous regressors have been widely studied and their asymptotic properties are well known\footnote{See inter alia tong1990, hansen1996,hansen1999,hansen2000 and gonwolf2005 for inference, gonpit2002 for multiple threshold regression and model selection, canerhansen2001 and gonpit2006 for threshold regression with unit roots, seolin2007 for smoothed estimators of threshold models, leeetal2011 for testing for thresholds, and hansen2016 for threshold regressions with a kink.}, the literature on threshold models with endogenous regressors remains relatively scarce.\footnote{For some contributions with endogenous regressors, see inter alia: for time-series, canerhansen2004, who consider exogenous threshold variables and kourtellosetal2013 who consider endogenous threshold variables; for cross-sections and (short) panels, seoshin2016 (and references therein), yu2018 and mcadam2019, who consider endogenous threshold variables.} Nevertheless, in many applications, the regressors are endogenous and the existence of a threshold has important policy implications. For example, among the empirical papers cited above, owyangetal2013, Cug:2015, rameyzubairy2018 and Jo:2022 use a threshold model with endogenous regressors, where the state dependence of the macroeconomic response is driven by a threshold variable being above or below a certain a-priori fixed value. rameyzubairy2018 (RZ henceforth) used a threshold model with endogenous regressors to investigate whether the government spending multiplier is larger in recessions, where recessions were defined by the unemployment rate being below or above a threshold parameter. This has important policy implications, because if the government spending multiplier is larger (above one) in recessions, it implies that governments should spend more in recessions to boost the economy.
In their analysis, RZ fix this threshold parameter at an unemployment rate of $6.5\%$.\footnote{This is based on the Federal Reserve's use of this threshold in a policy announcement. They later do robustness checks with a larger threshold, and modelled time-varying thresholds.} As the threshold parameter is typically unknown, we revisit their question and test for an unknown threshold, using - to our knowledge - the only parametric test available for linear time series models with endogenous regressors that directly applies to the RZ model. This test was proposed in canerhansen2004 (CH henceforth). CH first compute a Wald test statistic for all candidate threshold values between the $\epsilon$ and $(1-\epsilon)$ quantiles of the threshold variable, then take the maximum over this sequence to obtain a test for the null hypothesis of no threshold against the alternative hypothesis of an unknown threshold in (otherwise) linear models with endogenous regressors and exogenous threshold variables.
Our simulations show that this test has serious size distortions, with rejection frequencies up to three times the nominal size in small samples (see Tables (ref) and (ref)), accompanied by a reversal to severe under-rejections for larger samples of $1000$ observations. Tables (ref) and (ref) show that these size distortions are already present in just-identified models with strong instruments and homoskedastic data. We identify two problems with the CH test that lead to these size distortions, and proceed to correct them.
The first problem is illustrated in Figure (ref), where we test for an unknown threshold in the RZ model, and plot the sequence of the CH test statistics over the candidate threshold values, along with the same sequence for three tests we propose.\footnote{Section (ref) explains how these tests are calculated. Section (ref) describes the threshold estimator, the model and the data. } The plot shows erratic behavior of the CH test sequence, switching frequently between low and high values, especially around the sample edges, but starting already at the 25% and the 75% sample quantiles of the threshold variable. Therefore, the CH test, the maximum of the plotted sequence of tests, can change by a large amount when slightly changing the trimming.\footnote{Note that this is not due to the actual threshold estimate being between cut-off points: if there was a threshold, its consistent estimate, based on 2SLS or in CH, with 25% cut-off, is $8.33$; however, in our application in Section (ref), and in line with RZ, we do not find evidence of such a threshold.} This is problematic for its application in practice, as in general, it may lead to both over- or under-rejection of the null hypothesis, especially since this non-monotonic behavior is not well replicated by the bootstrap critical values even for samples of $1000$ observations, as shown in our simulations. \FloatBarrier
\FloatBarrier We identify the source of this problem to lie in the computation of the variance estimator in the middle of each Wald test for a candidate threshold. The residuals in the variance estimator are obtained with sub-sample parameter estimators, using observations only below or above each candidate threshold value. When the threshold value is close to the sample edges, these residuals can be very inaccurate approximations of the true underlying errors, because of the slow convergence rate of the sub-sample estimators employed to obtain them. We correct this by obtaining the residuals with full-sample estimators instead. Figure (ref) shows that all three test statistics we propose no longer display this non-monotonic behavior, whether computed with generalized method of moment estimators (GMM) estimators as in CH, or with two-stage least squares estimators (2SLS).
A second, yet related issue arises in the construction of the critical values of the CH test. The critical values of unknown threshold tests typically depend on the data, and therefore need to be simulated or bootstrapped. CH propose to bootstrap the critical values via a wild fixed regressor bootstrap and prove the bootstrap validity of their test. However, just like the variance estimator, the bootstrap residuals (and therefore the bootstrap samples) are computed with estimators under the alternative of each candidate threshold value.\footnote{Note that bootstrapping under the alternative is not necessary even when the variance estimator is computed with residuals under each alternative hypothesis of a candidate threshold value.}
While bootstrapping under the alternative does not affect the asymptotic validity of the CH test, it is problematic for two reasons. First, it is computationally much more intensive than computing the bootstrap samples just once, under the null hypothesis, using full-sample estimators. This is because for each bootstrapped test, one needs to compute many bootstrap samples corresponding to each candidate threshold value. Second, just like their sample equivalents, the bootstrap residuals will be inaccurate at the sample edges due to slow convergence of the sub-sample estimators used to employ them. When taking the maximum over the sequence of bootstrapped tests, then doing so for many bootstrap samples, the bootstrapped critical values can become highly unreliable for the original test statistic, even for sample sizes up to $1000$ observations. Tables (ref) and (ref) in the simulation section show severe under-rejection of the null hypothesis for sample sizes of $1000$ observations. They also show that bootstrapping under the null hypothesis fixes this issue, leading to correctly sized tests, but only if the variance correction discussed earlier is also employed.\footnote{Note that all test statistics for an unknown threshold we consider are non-pivotal, so one cannot expect any bootstrap to provide asymptotic refinements. While for the (trimmed) edges of the sample, the residuals computed with sub-sample estimators and their bootstrapped version are clearly inaccurately estimating the true underlying errors, because of slow convergence of the sub-sample estimator employed to construct them, this is not the case for the middle of the sample. Because both our tests and the CH test take the maximum over all candidate threshold values, around the (trimmed) sample edges or not, it is not possible to derive uniform asymptotic refinements of our tests over the CH test; these refinements will only hold for candidate threshold values around the sample edges. We would like to thank a referee for raising this issue.}
In this paper, we propose three test statistics for testing the null hypothesis of an unknown threshold in threshold models with endogenous regressors and exogenous regressors, and because both their computation and the bootstrap is different than CH, we derive for all three tests their asymptotic distribution and bootstrap validity. The first test we propose is similar to the CH test and uses sub-sample GMM estimators, but, unlike CH, employs a different variance estimator and a null bootstrap. The other two tests are a likelihood ratio (LR) test and a Wald test, both based on 2SLS estimators. The 2SLS estimators are not conventional and therefore not a special case of the sub-sample GMM estimators used in the CH test, because they use additional information about the first stage being either linear or having itself a threshold, while the GMM estimators do not use this information by construction. Therefore, the resulting test statistics can be equally reliable to the test based on GMM estimators, as shown in our simulations. Because the 2SLS with a first stage threshold require consistent estimators of the first stage threshold parameter, as a by-product of our analysis, we also prove the consistency of ordinary least-squares threshold estimators with a fixed threshold, a result we could not find in the extant literature, only for very specific regression models.\footnote{ See Theorem (ref) in the Online Supplement.}
Our paper is closely related to several papers in the change-point literature. bch2019 study the same 2SLS-based test statistics as this paper but for change-points. They also prove bootstrap validity of their tests, however we employ different proof techniques in this paper because the threshold variable is typically correlated with regressors, while the change-points are not, and the asymptotic distributions will also be different. mm2014 use information about change-points in the first stage to improve the power of tests for moment conditions, while we use similar information to improve the size of our tests. antoineboldea2015b and antoineboldea2018 also use a full sample first stage or change points in the first stage for more efficient estimation, while we focus on testing.\\ It should be noted that we allow for endogenous regressors, but not for endogenous threshold variables. For the latter, see inter alia kourtellosetal2013, yu2018, mcadam2019 and liao2019. To account for regressor endogeneity, we use instruments for constructing parametric test statistics for thresholds. As a result, our tests have nontrivial local power for $O(T^{-1/2})$ threshold shifts, where $T$ is the sample size. This is in contrast to yu2018, who do not use instruments, but rather local shifts around the threshold to construct a nonparametric threshold test. As a result, their test covers more general functional forms, at the cost of losing power in $O(T^{-1/2})$ neighborhoods. Additionally, the later paper focuses on cross-sectional models, while our tests are applicable to both cross-sectional models and time series models.\\ In the empirical application, using the same data and model specification as in RZ, we revisit the question of whether the government spending multipliers are larger in recessions. As in RZ, we cannot rule out that the cumulative government spending multipliers are the same in recessions and expansions. However, we estimate the threshold unemployment rate to be $8.3\%$, rather than $6.5\%$ as imposed in RZ. This new threshold causes the military spending instrument constructed in RZ to become weaker for deep recessions, suggesting that this instrument is probably most informative at moderate unemployment rates somewhere between $6.5\%$ and $8.3\%$. \\ The paper is organized as follows. Section 2 describes the model, the CH test and our test statistics, the proposed bootstrap, as well as the assumptions and all the bootstrap validity results. Section 3 contains simulations and Section 4 contains the empirical application. Section 5 concludes. The Online Supplement, at the end of this document, contains all the proofs.
Our framework is a linear model with a possible threshold at $\gamma^0$:
where $y_t$ is the scalar dependent variable, $x_t$ is a $p_1\times1$ vector of endogenous variables, $z_{1t}$ a $p_2\times1$ vector of exogenous variables including the intercept and possibly lags of $y_t$, $q_t$ is the scalar exogenous threshold variable, $\mathbf{1}[\cdot]$ is the indicator function, $w_t=(x_t^\top,z_{1t}^\top)^\top$ and $\theta_i^0=(\theta_{ix}^{0\top},\theta_{iz}^{0\top})^\top$. Let $\gamma^0\in \Gamma$, a strict subset of the support of $q_t$, and let $p=p_1+p_2$. The threshold variable is assumed exogenous and it can be a function of the exogenous regressors. As in CH, the first stage can be a linear model:
or a threshold model:
where $\rho^0 \in \Gamma$ is a threshold not necessarily coinciding with $\gamma^0$, and $z_t$ are $q \times 1$ strong and valid instruments, including $z_{1t}$, with $q-p_2 \geq p_1$. We assume that $\mathbb E[(\epsilon_t, u_t^\top)|\mathfrak F_{t}]=0$, where $\mathfrak F_t=\sigma\{z_{t-s},v_{t-s-1},q_{t-s}|s\geq0\}$, so that equation (ref) can be estimated by either 2SLS or by GMM. \\ We are interested in testing for an unknown threshold, i.e. the null hypothesis $\mathbb H_0: \theta_1^0=\theta_2^0 = \theta^0$. CH proposed a test based on GMM estimators of $\theta_i^0, (i=1,2)$ for each $\gamma \in \Gamma$. Because $z_t$ and $q_t$ are exogenous, the moment conditions
hold for all $\gamma \in \Gamma$. Based on these moment conditions, they construct the two-step GMM estimators:
with $\sum_{1\gamma}(\cdot) = \sum_{t=1}^T (\cdot) \indgam$, $\sum_{2\gamma}(\cdot) = \sum_{t=1}^T (\cdot)\indinvgam$, $\hat N_{i\gamma} = T^{-1} \sum_{i\gamma} w_t z_t^\top$ and with $\hat H_{\epsilon,i\gamma} = T^{-1} \sum_{i\gamma} \hat \epsilon_{t,\gamma,(1)}^2 z_t z_t^\top$. Here, $\hat \epsilon_{t,\gamma,(1)} = y_t - w_t^\top \hat \theta_{1\gamma,(1)}\indgam - w_t^\top \hat \theta_{2\gamma,(1)}\indinvgam$ are the first step GMM residuals for each $\gamma$, and $\hat \theta_{i\gamma,(1)}$ are consistent first-step versions of $\hat \theta_{i \gamma,(2)}$, for example by replacing $\hat H_{\epsilon,i\gamma}$ with $\hat M_{i\gamma} = T^{-1} \sum_{i\gamma} z_t z_t^\top$. These estimators can be used to construct a Wald test for each $\gamma$, and taking the supremum of this sequence of Wald tests over $\gamma \in \Gamma$ yields the test in CH:
where $\hat V_{\gamma,(1)} = \sum_{i=1}^2 \Big(\hat N_{i\gamma} \hat H_{\epsilon,i\gamma}^{-1} \hat N_{i\gamma}^\top\Big)^{-1}$.\\ As shown in CH, the asymptotic distribution of the test statistic (ref) is non-pivotal and therefore needs to be simulated/bootstrapped for a given application. CH propose to generate new pseudo-dependent variables $y^b_{t,\gamma}=\hat\epsilon_{t,\gamma,(2)}\eta_t$, where $\eta_t\stackrel{iid}\sim\mathcal N(0,1)$ and $\hat\epsilon_{t,\gamma,(2)}$ denote the second step GMM residuals for each value of $\gamma$, and recalculate the test statistic (ref) for each $\gamma$ using $y^b_{t,\gamma}$ instead of $y_t$, and then for many bootstrap samples.\footnote{Note that the pseudo-dependent variables are generated without adding back the estimated mean to the bootstrap residuals. This is inconsequential to the analysis because the test statistic is based on mean differences across regimes of low or high $q_t$, and these are zero under the null of no threshold.} Even though CH prove validity of their bootstrap procedure in large samples, Tables (ref) and (ref) show that their bootstrap does not replicate well the empirical distribution of the test statistic in finite samples, being severely over-/undersized for small/large samples.\\ Tables (ref) and (ref) in the simulation section show that these size distortions are due to two interacting phenomena: the type of bootstrap employed, and the way the heteroskedasticity-robust variance estimators are computed. We therefore employ two corrections. First, we adjust the bootstrap such that the pseudo-dependent variable $y_t^b$ is constructed using full-sample residuals. That is, we replace $y_t^b=\hat\epsilon_{t,\gamma,(2)}\eta_t$ by $y_t^b=\hat\epsilon_{t,(2)}\eta_t$ where $\hat\epsilon_{t,(2)}=y_t-w_t^\top\hat\theta_{(2)}$ and $\hat\theta_{(2)}$ is the second step GMM estimate under $\mathbb H_0$. This gets rid of the undersizing of the CH test statistic documented in the simulation section: the residuals become more accurate around the sample edges as they are not constructed with sub-sample estimators. However, the simulations now indicate that the test is oversized (see Tables (ref) and (ref) , column “Mix").
Therefore, we employ a second correction, where the heteroskedasticity-robust variance estimators are also computed with full-sample parameter estimates. More exactly, rather than using $\hat\epsilon_{t,\gamma,(1)}$ and $\hat\epsilon_{t,\gamma,(2)}$ in the expression for $\hat H_{\epsilon,i\gamma}$, we use $\hat\epsilon_{t,(1)}=y_t-w_t^\top\hat\theta_{(1)}$ instead, where $\hat\theta_{(1)}$ is the first step full-sample GMM estimate (so we redefine $\hat H_{\epsilon,i\gamma}=\sum_{i\gamma} \hat\epsilon_{t,(1)}^2z_tz_t^\top$). As Tables 1 and 2 show (column “BR", “bootstrap/rectification”), this yields correctly sized sample test statistics in all samples considered.
Note that both effects that we correct for are due to unstable estimates of the residuals at the sample edges below/above the 15%/85%-quantiles of the empirical distribution of $q_t$. The test employing these two corrections is denoted by $WG_{T,BR}=\supl_{\gamma\in\Gamma}WG_{T,BR}(\gamma)$.\\
We also consider two 2SLS-based test statistics, because the GMM estimators involved in the computation of the tests above do not use information about the linearity or lack of linearity of the first stage. Therefore, they are not more efficient than the 2SLS estimators that use this information (see antoineboldea2015b for a formal proof of this statement for change-point models), so there is no reason to expect that 2SLS-based tests will be inferior to the GMM-based tests.
The likelihood-ratio type and a Wald-type test statistic for $\mathbb H_0:\,\theta_1^0=\theta_2^0$ based on 2SLS estimators are:
where $SSR_0=\sum_{t=1}^T(y_t-\hat w_t^\top\hat\theta)^2$, with $\hat\theta=(\sum_{t=1}^T\hat w_t\hat w_t^\top)^{-1}(\sum_{t=1}^T\hat w_ty_t)$ the full-sample 2SLS estimator, $SSR_1(\gamma)=\sum_{i=1}^2\sum_{i\gamma}(y_t-\hat w_t^\top\hat\theta_{i\gamma})^2$, with $\hat\theta_{i\gamma}=(\sum_{i\gamma}\hat w_t\hat w_t^\top)^{-1}(\sum_{i\gamma}\hat w_ty_t)$ the split-sample 2SLS estimators. Here, $\hat w_t=(\hat x_t^\top,z_{1t}^\top)^\top$ stacks the predicted endogenous variables $\hat x_t$ and the exogenous variables $z_{1t}$. The predicted endogenous variables are obtained either via estimating the linear first stage equation (ref):
or via estimating the threshold first-stage equation (ref):
Lastly, $\hat V_\gamma\inp V_\gamma=\lim\Var[T^{1/2}(\hat\theta_{1\gamma}-\hat\theta_{2\gamma})]$.\footnote{ The explicit expressions for $\hat V_\gamma$ and $V_\gamma$ are given in Online Supplement Section (ref), and Definition (ref) for a linear first stage, and in Online Supplement Section (ref), Definition (ref) for a threshold first stage, together with the expressions in the asymptotic distributions of the 2SLS test-statistics.} Unlike the sup Wald test in halletal2012, which is the change-point counterpart of the test here, our test -- through the way $\hat V_\gamma$ is defined -- takes into account that the 2SLS estimators $\hat\theta_{1\gamma}$ and $\hat\theta_{2\gamma}$ are correlated through either a full-sample first-stage or through misalignment of $\rho^0$ and $\gamma$. Moreover, as in the case of CH's GMM-test, the 2SLS test-statistics are non-pivotal and, therefore, need to be simulated/bootstrapped. The next subsection describes the bootstrap we propose and contains results for asymptotic validity of this bootstrap for all three tests proposed.
The bootstrap employed for both CH GMM test and our GMM test is a wild bootstrap with fixed regressors because it does not bootstrap the regressors $w_t, x_t$ and the instruments $z_{1,t}$. We already alluded to the proposed change in the bootstrap procedure for the CH test in the previous section. These changes are summarized in the Algorithms (ref) and (ref) below. The difference between the CH test and our test are highlighted in lines 3--6 of the below algorithms. Since CH construct their pseudo-dependent variable for each $\gamma$ separately, the for-loop over $\gamma$ starts already in line 3 of Algorithm (ref), as opposed to line 5 in Algorithm (ref) when the same pseudo-dependent variable is used for all values of $\gamma$. Line 6 in both algorithms indicates the difference in constructing the heteroskedasticity-robust variance estimators $\hat H_{\epsilon,i\gamma}^b$ and in $\hat H_{\epsilon,i\gamma}$.\\
\hfil
Algorithm 3 below describes the wild fixed regressor bootstrap for our proposed 2SLS test-statistics, where regressors $w_t$ and instruments $z_{1,t}$ are kept fixed in the bootstrap. The first stage linearity or lack thereof is taken into account in computing $\hat x_t$ - equation (ref) or (ref) respectively. For these tests, we need to know whether the first stage is linear or not; however, this is not necessarily a drawback in empirical work, because such knowledge is required for estimating the threshold parameter $\gamma^0$ consistently (see CH).\\
We now derive the asymptotic properties of our tests\footnote{We focus on the GMM-based test; the asymptotic distributions of the 2SLS tests are given in the Online Supplement, Sections (ref) and (ref).} and show their bootstrap validity. First define $g_t = x_t - u_t$, $h_t = y_t-\epsilon_t$, $M_{1\gamma}=E[z_tz_t^\top\indgam]$, $M=\plim_{\gamma\to\infty}M_{1\gamma}=E[z_tz_t^\top ]$, $M_{2\gamma}=M-M_{1\gamma},$ and $v_t=(\epsilon_t,u_t^\top)^\top$. Let $\|\cdot \|$ be the Euclidean norm. The following assumptions are similar to CH.
Most of these assumptions are also used in CH. Assumption (ref) $(a)$ is typically needed for nonlinear models. Assumption (ref) $(b)$ is also needed, as the only uniform law of large numbers and functional central limit theorem for partial sums in $\indgam$ that we are aware of derives from hansen1996 and require strict stationarity (see Lemma (ref)-(ref) in the Online Supplement). Assumption (ref) $(c)$ is a typical moment condition. Assumption (ref) $(d)$ is slightly different than CH: they also impose that $M_{1\gamma}$ is p.d. for all $\gamma$, but we require that the increments in $M_{1\gamma}$ are p.d. in the limit with eigenvalues bounded away from zero. The latter is technical in nature and required to obtain quantities bounded in probability in order to provide a self-contained proof of super-consistency of $\hat \rho$ in a threshold first stage model. Assumptions (ref) $(e)$ is standard in the threshold literature, and (ref) $(g)$ is an identification condition for a possible threshold in the first stage. Assumption (ref) $(f)$ is needed for uniqueness of the asymptotic distributions of the test statistics proposed.\\ With this assumption, we first show that employing the new heteroskedasticity-robust estimators do not alter the distribution of the CH test.
To show the validity of the null bootstrap, we require the following additional assumption:
Assumption (ref) $(a)$ is common for the wild bootstrap bch2019, and typical choices for $\eta_t$ are the normal distribution, the Rademacher distribution, and the asymmetric two-point distribution in mammen1993. CH propose using the normal distribution, but we use both the normal distribution and the mammen1993 distribution, as the latter yields better results for the GMM-Wald test, see Tables (ref)--(ref). Assumption (ref) $(b)$ is only needed for $\hat H_{\epsilon,i\gamma}^b$ to weakly converge to $H_{\epsilon,i\gamma}^b$ in probability under the bootstrap measure. Theorem (ref) proves the asymptotic validity of the null bootstrap for $WG_{T,BR}$.
For the 2SLS-based test-statistics, the asymptotic distributions are cumbersome and not of main interest. Therefore, we relegate these results to the Online Supplement, Sections (ref) and (ref). However, in order to derive these asymptotic distributions, we also provide in the Online Supplement, Theorem 5, a self-contained proof of super-consistency of the first stage (ordinary least-squares) threshold parameter estimate $\hat\rho$. This was also shown in chan1993, but for a threshold autoregressive model where $z_t$ and $q_t$ are lags of $x_t$. This proof may be of interest in its own right, as it extends proof techniques from change point analysis to threshold models.
The asymptotic distributions of the 2SLS based tests are also non-pivotal, and we conclude this section by stating the asymptotic validity of the bootstrap for these tests.
Consider the following data generating process (DGP) for $t=1,\ldots, T$:
where $z_t\stackrel{iid}{\sim}\mathcal N(1,1)$, $q_t=z_t+1$, and $z_t,\,x_t$, and $q_t$ are scalars. We set $\delta_\Pi=0$ for a linear first stage (LFS) and $\delta_\Pi\in\{-0.5,0.5,1\}$ for a threshold first stage (TFS) with $\rho^0=1.75$. Under the null hypothesis, $\delta_x=0$, and under the alternative hypothesis, $\delta_x =0.25$ with $\gamma^0=2.25$.\footnote{Note that because of just-identification, there is no difference between the first and the second-step GMM estimators, therefore $\hat \theta_{(2)} =\hat \theta_{(1)} = (\sum_{t=1}^T w_t z_t^\top )^{-1} ( T^{-1} \suml_{t=1}^T z_t y_t )$, and $\hat \theta_{i\gamma,(2)} = \hat \theta_{i\gamma,(1)} = (\hat N_{i\gamma}^\top)^{-1} ( T^{-1} \suml_{i\gamma} z_t y_t )$.} To generate $\epsilon_t$, we define $e_t\stackrel{iid}\sim\cN(0,1)$ and consider the following three cases.\\ In case (a), the errors are homoskedastic, i.e. $\epsilon_t=e_t$, and the econometrician knows this. Therefore, we use the i.i.d. bootstrap instead of the wild bootstrap, and make two adjustments to the computation of the test statistics. First, $v_t^{b}\stackrel{iid}\sim\mathcal N(0,\hat{\Sigma}_v)$ with $\hat\Sigma_v=T^{-1}\sum_{t=1}^T\hat v_t\hat v_t^\top$ for 2SLS. For GMM, $\epsilon_t^b \stackrel{iid}\sim\mathcal N(0,\hat\sigma^2_\epsilon)$ given the data, with $\hat\sigma^2_\epsilon=T^{-1}\sum_{t=1}^T\hat\epsilon_{t,(1)}^2$, and $\hat \epsilon_{t,(1)} = y_t- w_t^\top \hat \theta$, with $\hat \theta$ the full-sample 2SLS estimator. Second, all heteroskedasticity-robust estimators are replaced by their homoskedastic analogs. For example, $E[z_tz_t^\top\epsilon_t^2\indgam]$ is no longer estimated by $T^{-1}\sum_{1\gamma} z_tz_t^\top\hat\epsilon_t^2\indgam$, but by $\hat\sigma_\epsilon^2(T^{-1}\sum_{1\gamma}z_tz_t^\top\indgam)$. The same applies to the case of CH's original bootstrap, except that the residuals are computed for each value of $\gamma$ rather than under $\mathbb H_0$.\\ In case (b), the errors are still homoskedastic, i.e. $\epsilon_t=e_t$, but this is not known to the econometrician. Therefore, the heteroskedasticity-robust variance estimators described in Sections (ref) and (ref) are employed. \\ In case (c), the errors are conditional heteroskedastic, i.e. $\epsilon_t=e_t\cdot z_t/\sqrt{2}$ with $Var(e_t)=Var(u_t)=1$ and $Cov(u_t,e_t)=0.5$, and heteroskedasticity-robust variance estimators are employed.\\ In Tables (ref) and (ref) cases (b)-(c), the bootstrap is performed using $\eta_t\stackrel{iid}{\sim}\mathcal N(0,1)$. For all other results, $\eta_t \stackrel{iid}{\sim}(0,1)$ with draws from the asymmetric two-point distribution proposed by mammen1993. In all cases besides (a), we use the wild bootstrap as described in Section (ref).\\ There are $500$ bootstrap samples. For each simulation, we compute the 95% quantile of the bootstrap distribution of the test statistic, and if the test in the original sample is above this quantile, we reject, else we do not reject. $\gamma$ is varied between all sample realizations of $q_t$ from its $15\%$ quantile to its $85\%$ quantile. We report the rejection frequency of each test statistic in $1000$ simulations under the null and at 5% nominal size (Tables (ref)--(ref)), and under the alternative we plot the size-adjusted power, where the size-adjustment is made relative to the null DGPs described above (Figure (ref)).\\ \\ Tables (ref) and (ref) show that the bootstrap procedure originally proposed by CH has heavy size distortions in both directions. In particular, the test moves from being heavily oversized in small samples to severely undersized in large samples (columns “CH”). This originates from imprecise residual estimates and imprecise $\hat H_{\epsilon,i\gamma}$ for small and moderate sample sizes pertinent to applications. In particular, when $\gamma$ is close to the 15% or 85% quantiles of $q_t$, there is not enough data to obtain precise residuals and precisely estimate $H_{\epsilon,i\gamma}$. Moreover, changing the CH bootstrap to a null bootstrap results in oversized tests for all considered sample sizes (columns “Mix”). This problem is rectified by modifying $\hat H_{\epsilon,i\gamma}$ in the original test statistic, as evident from columns “BR” in Tables (ref) and (ref), where the empirical sizes are much closer to the nominal size. Tables (ref) and (ref) show that the 2SLS tests are in almost all cases close to nominal sizes, even in small samples, and that there is no clear ranking among the three proposed tests.
\FloatBarrier
\FloatBarrier
We also assess the power of all tests. For a large threshold $\delta_x=1$, all tests have power virtually equal to one even for sample sizes of $T=250$ and therefore we do not report these results. Figure (ref) shows the power properties for a small threshold of $\delta_x=0.25$. In small samples, the Wald tests dominate the LR test for all cases (a)-(c). Note that this is not necessarily for classical reasons of correcting for heteroskedasticity, as all tests are non-pivotal and bootstrapped. The power differences among all three tests vanish as the sample size grows. Therefore, we argue that all the tests proposed provide reliable alternatives in moderate samples pertinent to macroeconomic applications.
\FloatBarrier
\FloatBarrier
In this section, we revisit the question whether government spending is more effective in recessions, and address it as in RZ, using exactly the same data and model specifications, except that we test and estimate an unknown threshold rather than imposing it. For simplicity, we first focus on the instantaneous government spending multiplier $\theta_{g,i}(i=1,2)$, estimated similarly to RZ from:
where $y_t$ is real GDP divided by trend GDP, $g_t$ is real government spending divided by trend GDP -- which is endogenous and instrumented by military spending news $m_t$ -- and the threshold variable is $q_{t}$, the first lag of the unemployment rate. The exogenous regressors $z_{1t}$ are also included in $z_t$ and contain an intercept and four lags of $g_t,\,y_t,\,m_t$. Thus, $z_t=[z_{1t}^\top, m_t^\top]^\top$.\\ The data is from the RZ replication package.\footnote{\url{http://econweb.ucsd.edu/ vramey/research/Ramey_Zubairy_replication_codes.zip}} For details on the data construction, instrument validity, or interpretation of $\theta_{g,i}(i=1,2)$ as cumulative spending multipliers, we refer the interested reader to RZ. \\ Letting $\theta_i^0=(\theta_{g,i}, \theta_{z,i}^\top)^\top$ and $w_t = (g_t, z_{1,t}^\top)^\top$, the RZ estimators of $\theta_i^0$ are exactly the just-identified GMM (or instrumental variables, IV henceforth) estimators $\hat \theta_{i\gamma,(1)}$ defined in Section (ref), but evaluated in RZ at $\gamma=6.5$ (and ignoring the first stage which is irrelevant for conventional IV estimators).\footnote{All numbers referring to unemployment rates, such as $6.5$, should be interpreted as percentages: $6.5\%$.} The threshold $\gamma= 6.5$ is chosen by RZ as in owyangetal2013, based on the US Federal Reserve use of this threshold in its policy announcement; RZ also do a robustness check with a threshold of $8.0$. Since it is unclear why $6.5$ or $8.0$ would be the threshold that defines recessions versus expansions, we do not assume that the threshold $\gamma^0$ is known or even that there is a threshold $\gamma^0$; we instead test for the presence of $\gamma^0$ first. \\ The 2SLS tests require first estimating $\rho^0$ in equation (ref). Table (ref) reports the multivariate threshold estimates $\hat \rho$ described in Section (ref), along with the decisions of a LFS or a TFS based on the BIC3 criterion proposed in gonpit2002 and on the ordinary least-squares (OLS) versions of $LR_T^b$ and $W_T^b$ tests described in Section (ref), which were proposed in hansen1996. The estimate of $\rho^0$ change with the cut-off considered, but there is considerable evidence of a threshold in the first stage. The maximizer of the OLS version of $LR_T(\gamma)$ is exactly $\hat \rho$, a consistent estimator of $\rho^0$ as shown in Theorem (ref). Therefore, we use a TFS with $\hat \rho$ in Table (ref). \\ \FloatBarrier
Given $\hat \rho$ obtained for each cut-off, we test for an unknown threshold in equation (ref). Table (ref) shows that the LR test rejects the null. The 2SLS Wald test and our modified GMM Wald-test do not reject (but their values are relatively close to the critical values at certain cut-offs). From Figure (ref), it is evident that the sequence of all our test statistics are relatively flat for all values of $\gamma$. The CH test also never rejects the null, but its sequence is not flat: its value is relatively large at 10% trimming, and its critical values are very large at all trimming levels. This is in line with our simulations, which indicated that the tests are undersized at 500 observations, the number of observations in our sample. Its erratic behavior near the sample edges was further illustrated in Figure 1.\\ Because Equations (ref)--(ref) control for several lags -- in line with the RZ specification -- we choose the 25% cut-off results with $\hat \rho=4.0636$ and $\hat \gamma=8.3363$, where the latter is the 2SLS threshold estimate proposed in CH (or, equivalently, the implicit maximizer of the $LR_T(\gamma)$ quantity in this paper).\footnote{The confidence sets for both these thresholds obtained by inverting the likelihood ratio tests in hansen2000 and CH, or by simulating the asymptotic distribution in CH, are very tight when using the default nonparametric kernel. However, since both estimators are close to the 25% cut-off, and increase ($\hat \gamma$) or decrease ($\hat \rho$) when decreasing the cut-offs used, we can only interpret these estimators as close to the lower bounds of the true threshold values that are identified in the sample.} \\ We could conclude based on Figure (ref) and Table (ref) that there is little evidence that the instantaneous multipliers are different in recessions and expansions. In what follows, we also show that there is little evidence that the multipliers at other horizons than zero are different. To that end, as in RZ, we compute the cumulative government spending multipliers $\theta_{g,i}^h (i=1,2)$ at horizon $h=1,\ldots,H$ from the IV regression:
where $\suml_{h=0}^H g_{t+h}$ is instrumented by $m_t$.\footnote{It is unclear how to use the TFS specification (ref) to obtain cumulative government spending multipliers at $h>0$, because of the misalignment between the first and the second stage threshold, and we leave this to future research.}\\ Tables (ref)-(ref) show the RZ multipliers (using $6.5$ and $8$ - robustness check in RZ - as thresholds), and our multipliers for fifteen quarters ahead, calculated exactly as in RZ but with $\hat \gamma=8.3383$. We also report classical heteroskedasticity and autocorrelation (HAC) robust standard errors, weak instrument HAC robust confidence sets, and classical and weak-instrument HAC-robust tests for the difference in multipliers at the imposed thresholds. These tables show that in all cases, there is no evidence that government spending multipliers are different in recessions, once the possibility of weak instruments is taken into account at all horizons. \FloatBarrier
\FloatBarrier
\FloatBarrier
We therefore assess the possibility of weak instruments at various horizons in Figures (ref) and (ref), plotting the effective F-statistic for the null hypothesis of weak instruments in each regime across horizons. These figures show evidence of weak instruments in both regimes at short horizons, for all thresholds. This also holds for the effective F-statistics minus their critical value for our TFS specification with $\hat \rho=4.0636$: they are equal to approximately $-19$ for $q_t \leq 4.0636$ (101 observations), and $-17.5$ for $q_t> 4.0636$ (399 observations), so well below zero. Therefore, the weak instrument robust p-values should be used, even for Table (ref), at shorter horizons. Hence, once weak instruments are accounted for, there is no evidence that government spending multipliers are different in recessions, both in our paper and in RZ.\\ What we do learn from the analysis is that military spending news becomes a weaker instrument for longer horizons when the threshold increases from $6.5$ to $8.0$ or to $8.3363$, and therefore that the instrument relevance is not robust to the threshold used. This is also indicated in Figure (ref), which shows that, except for the World War II period, the news variable does not exhibit much variation when the unemployment rate is above $8.3363$. This suggests that the RZ military news instrument is more informative for intermediate values of unemployment, so for "normal" recessions rather than "deep" recessions.
In this paper we proposed two adjustments to the GMM Wald test of canerhansen2004, and two new 2SLS test statistics for threshold detection in linear models with endogenous regressors and exogenous thresholds. We derived the asymptotic validity of their null bootstrap equivalents, and showed through simulations and an application that these tests have better finite sample properties than the test proposed in canerhansen2004. \\ rothfelderboldea2016 show in their Theorem 1 that under conditional homoskedasticity and one endogenous regressor, the 2SLS estimators with a linear first stage or a threshold first stage can be more efficient than the GMM estimators that ignore this information. It would be interesting to assess when this efficiency carries over to more general settings, and whether there exists an optimal GMM estimator that uses similar information from the first stage as the 2SLS estimators.