EconBase
← Back to paper

Uniform Inference in High-Dimensional Threshold Regression Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

116,119 characters · 21 sections · 99 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Uniform Inference in High-Dimensional Threshold Regression Models

abstractWe develop a uniform inference theory for high-dimensional slope parameters in threshold regression models, allowing for either cross-sectional or time series data. We first establish oracle inequalities for prediction errors, and $\ell_1$ estimation errors for the Lasso estimator of the slope parameters and the threshold parameter, accommodating heteroskedastic non-subgaussian error terms and non-subgaussian covariates. Next, we derive the asymptotic distribution of tests involving an increasing number of slope parameters by debiasing (or desparsifying) the Lasso estimator in cases with no threshold effect and with a fixed threshold effect. We show that the asymptotic distributions in both cases are the same, allowing us to perform uniform inference without specifying whether the model is a linear or threshold regression. Additionally, we extend the theory to accommodate time series data under the near-epoch dependence assumption. Finally, we identify statistically significant factors influencing cross-country economic growth and quantify the effects of military news shocks on US government spending and GDP, while also estimating a data-driven threshold point in both applications. \\ \\ {\bf Keywords:} Sample splitting, Model selection, High-dimensional inference, Oracle inequalities. \\ {\bf JEL Codes:} C12, C13, C24.

\onehalfspacing

Introduction

Consider the following threshold regression model

equation[equation omitted — 138 chars of source]

where $X_{i}$ is a $p\times 1$ covariate vector, and $Q_{i}$ is the threshold variable determining regime switching; for example, rich countries may follow a different economic growth pattern from poor countries. $\tau_0$ is the unknown threshold parameter, and $U_{i}$ is the error term. In this paper, we focus on uniform inference for high-dimensional regression parameters $(\beta _{0},\delta _{0}),$ allowing for $p>n.$ The threshold autoregression (TAR) model, with the lag of the series as the threshold variable, was formally introduced by tong1980threshold to analyze cyclical time series data. It is a class of non-linear time series models and is parsimonious for nonparametric model estimation. \footnote{For a survey paper, see hansen2011threshold.} \footnote{chan1993consistency and chan1998limiting study the limiting properties of the least square estimators in the threshold autoregression model.} potter1995nonlinear applies it to study the properties of US GNP and finds that the response of output to shocks is asymmetric throughout different stages of the business cycle.

Subsequently, threshold regression is utilized by hansen2000 to identify multiple regimes based on a particular predetermined variable, allowing for either time series or cross-sectional data. Since then, there has been growing interest in reanalyzing previous empirical applications using threshold models, particularly when multiple equilibria may exist. For example, lee2016 consider cross-country economic growth behaviors initially analyzed by durlauf1995multiple; yu2021threshold and lee2023threshold examine race-based tipping behavior in residential segregation discussed in card2008tipping; and canerdebt, aj2013, and debt2017 investigate the effect of government debt on economic growth originally studied by reinhart2010growth. In this paper, we confirm the existence of multiple steady states in cross-country economic growth by showing that some threshold-effect coefficients are significantly different from zero. In addition, we apply the high-dimensional local projection threshold model to find a “data-driven” threshold point that defines the state of the economy and reestimate the impulse response to a military spending news shock in government spending and GDP.

The main contribution of this paper is to develop a uniform inference procedure for an increasing number of slope parameters in high-dimensional threshold regression models, a class of parsimonious nonlinear regression models. To the best of our knowledge, we are the first to establish that the debiased Lasso estimator achieves uniform convergence over a large class of parameters without the need to pre-specify the existence of a threshold effect, even when the number of covariates grows much faster than the sample size. In contrast, the existing literature has primarily focused on applying debiased methods within high-dimensional linear regression frameworks. Meanwhile, we demonstrate that the asymptotic distributions of the tests are identical in cases with no threshold effect and with a fixed threshold effect. Specifically, we derive oracle inequalities for prediction errors and $\ell_1$ estimation errors for the Lasso estimator of the slope parameters and the threshold parameter under more general conditions, allowing for heteroskedastic non-subgaussian error terms and non-subgaussian covariates when studying cross-sectional data. Moreover, we further extend the framework to high-dimensional time series threshold regression models and establish uniform inference theory under the near-epoch dependence assumption.

This work focuses on a high-dimensional framework. Firstly, variable selection is necessary to identify threshold effects. A linear model that incorporates a broader set of regressors can outperform a statistical model that emphasizes threshold effects with a specific set of covariates, as highlighted by lee2016. Meanwhile, economic theory often provides guidance on a set of variables that are likely to be relevant,, but does not specify precisely which variables are truly important or the functional form through which they should enter the model. This ambiguity leaves researchers with the challenge of selecting an appropriate set of control variables from a potentially large pool, which may include not only the raw regressors available in the data but also their interactions and various nonlinear transformations. We thus study high-dimensional settings to keep the model free from variable selection. High dimensionality may also result from addressing the issue of confoundedness (doi:10.1080/07350015.2016.1204919) and avoiding the non-invertibility of a structural moving average model (doi:10.1198/016214502388618960). Specifically, when we study local projection threshold regression, which is a special case of time series threshold regression, the number of covariates naturally becomes large due to the inclusion of multiple lags to control for autocorrelation.\footnote{See hdlp for further discussion on impulse response analysis with a large number of variables.} Additionally, when we apply high-dimensional threshold models to empirical applications, due to sample splitting, the total number of parameters may be larger than the sample size in the regime with the fewest observations, particularly when multiple threshold points exist, leading to poor estimation and out-of-sample prediction in finite samples. However, traditional estimation and inferential methods, such as OLS and MLE, are no longer valid even in high-dimensional linear regression models. Many methods are available for high-dimensional estimation and variable selection, for example, Lasso in tibshirani1996regression. We apply Lasso to estimate the high-dimensional threshold regression (allowing for $p>n$), as in lee2016 and canerkock2017.

In this paper, we first study threshold regression with cross-sectional data, allowing for heteroskedastic non-subgaussian error terms and non-subgaussian covariates. We use the concentration inequality \footnote{The concentration inequality originates from chernozhukov2014gaussian and chernozhukov2015comparison, as formulated in Lemma 2 of chiang_rodrigue_sasaki_2023} for the partial sum of random variables that we propose in Lemma (ref) to derive oracle inequalities for both the prediction error and $\ell_1$ estimation error for coefficients, which are qualitatively same as those in lee2016.

Next, we construct the desparsified Lasso estimator by using nodewise regression to estimate the empirical precision matrix. In our proof, we maintain the dependence assumption between covariates and the threshold variable, and we apply the inverse of a 2 $\times$ 2 block matrix to construct the precision matrix when the threshold effect may exist.\footnote{The independence assumption would significantly simplify the proof, but it is uncommon in empirical applications.} Based on the inference theory of canerkock2018 and our oracle inequalities, we establish the asymptotic distribution of tests involving an increasing number of slope parameters in the cases with no threshold effect and with a fixed threshold effect. We show that the asymptotic distributions in both cases are identical. We also provide a uniformly consistent covariance matrix estimator in both cases. \footnote{There is a slight difference between the limits of their asymptotic variances since, in the case of a fixed effect, there is a true value for the threshold parameter.} We further construct asymptotically valid confidence intervals for the interest of the slope parameter, which are uniformly valid and contract at the optimal rate. Moreover, we develop the uniform inference theory for the debiased Lasso estimator to the setting of the high-dimensional time series threshold regression model under the near-epoch-dependence assumption, with local projection threshold regression as a special case.

{\bf{Relation to literature.}} The existing literature on high-dimensional threshold regression has focused on deriving oracle inequalities for the prediction errors and estimation errors for the Lasso estimator of the slope parameters and the threshold parameter in the case of fixed design with gaussian errors (lee2016), and on model selection consistency in the case of random design with sub-gaussian covariates and errors (canerkock2017). However, high-dimensional inference is another important topic in statistics and econometrics; for example, the estimation of impulse response functions is an essential part of econometric inference in time series models. Thus, in this paper, we perform uniform inference for high-dimensional threshold regression parameters, allowing for either cross-sectional or time series data, by applying the de-sparsified method of geer2014 to complement the existing literature. This method desparsifies the estimator by constructing a reasonable approximate inverse of the singular empirical Gram matrix, thereby removing the bias from the estimation of the shrinkage method. Our asymptotic result is uniformly valid over the class of sparse models, where $s_0$ represents the sparsity level, which can grow with $n.$

A growing body of literature applies the desparsified method of geer2014 to perform uniform inference in high-dimensional regression models, motivated by the insight of Leeb_Pötscher_2005 that failing to account for the model selection step can lead to invalid statistical inference. Gold2018 desparsified the Lasso estimator based on a two-stage least squares estimation, allowing both numbers of instruments and of regressors to exceed the sample size. semenova2023inference desparsified the orthogonal Lasso estimator in their third stage when heterogeneous treatment effects are present. Additionally, 10.1093/jjfinec/nbac023, ADAMEK20231114, and hdlp constructed the desparsified Lasso estimator in high-dimensional time series models. The desparsified method has also been applied in high-dimensional panel data models to perform uniform inference, as shown in works by KOCK2016, kock_tang_2019, and chiang_rodrigue_sasaki_2023. However, all of these studies test hypotheses for a bounded number of parameters. canerkock2018 considered hypotheses involving an increasing number of parameters in linear regression models. We contribute to this strand of literature by developing a uniform inference framework in threshold regression models, a class of nonlinear regressions, accommodating an increasing number of slope parameters.

{\bf Organization:} The rest of the paper is organized as follows. Section (ref) recalls the Lasso estimator of lee2016 and establishes oracle inequalities under weaker conditions on covariates and error terms. We construct the debiased Lasso estimator and derive the uniformly asymptotic distribution of hypothesis tests in Section (ref). Section (ref) develops the uniform inference theory for high-dimensional time series threshold regression models. In Section (ref), we investigate finite sample properties of our debiased Lasso estimator, followed by two empirical applications. Section (ref) concludes. All proofs are deferred to the Appendix.

Notation

Denote the $\ell _{q}$ norm of a vector $a$ by $\left| a\right| _{q}$ and the empirical norm of $a \in \mathbb{R}^n$ by $||a||_n:= \left(n^{-1}\sum_{i=1}^n a_i^2\right)^{1/2}.$ For any $m\times n$ matrix $A,$ the induced $l_1$-norm and $l_\infty$-norm of $A$ are defined as $\Vert A\Vert_{l_1}:=\max_{1\leq j\leq n}\sum_{i=1}^m|A_{ij}|$ and $\Vert A\Vert_{l_\infty}:=\max_{1\leq i\leq m}\sum_{j=1}^n|A_{ij}|,$ respectively. Additionally, define $\Vert A\Vert{}_\infty:=\max_{1\leq i\leq m, 1\leq j\leq n}|A_{ij}|.$

For $a\in\mathbb{R}^n,$ denote the cardinality of $J(a)$ by $|J(a)|,$ where $J(a)=\{j=1,...,n: a_j \neq 0\}.$ Let $a_M$ denote the vector in $\mathbb{R}^n$ that has the same coordinates as $a$ on $M$ and zero coordinates on $M^c.$ Let the superscript $^{\left( j\right) }$ denote the $j$th element of a vector or the $j$th column of a matrix.

Finally, define $f_{(\alpha ,\tau )}(x,q):=x^{\prime }\beta +x^{\prime }\delta \bm{1}\{q<\tau \},$ $f_{0}(x,q):=x^{\prime }\beta _{0}+x^{\prime }\delta _{0}\bm{1}\{q<\tau _{0}\},$ and $\widehat{f}(x,q):=x^{\prime }\widehat{\beta } +x^{\prime }\widehat{\delta }\bm{1}\{q<\widehat{\tau }\}$. The prediction norm is defined as $\left\Vert \widehat{f}-f_{0}\right\Vert _{n} :=\sqrt{\frac{1}{n} \sum_{i=1}^n \left( \widehat{f}(X_i,Q_i) - f_{0}(X_i,Q_i) \right)^2 }.$

The literature refers to the method as either the “debiased” or the “desparsified” Lasso estimator; for clarity, we consistently use the term “debiased Lasso estimator” throughout the remainder of this paper.

The Lasso Estimator and Oracle Inequalities

Lasso Estimation

The threshold regression model (ref) can be rewritten as

align[align omitted — 231 chars of source]

$Q_{i} $ is the threshold variable that splits the sample into two regimes and $\delta_0$ represents the threshold effect between two regimes. The model (ref) thus captures a regime switch based on the observable variable $Q_i.$ The parameter $\tau_0$ is the unknown threshold parameter, which lies within a compact parameter space $T=[t_0,t_1].$ There is no threshold effect when $\delta_0 = 0,$ and the model reduces to a linear regression.

Denoting a $(2p\times 1)$ vector by $\bm{X}_{i}(\tau )=(X_{i}^{\prime }, X_{i}^{\prime }\bm{1}\{Q_i <\tau \})^{\prime }$ and an $(n\times 2p)$ matrix by $\bm{{X}}(\tau),$ where the $i$-th row is $\bm{X}_{i}(\tau )^{\prime }.$ Let ${X}$ and ${X}(\tau)$ denote the first and last $p$ columns of $\bm{{X}}(\tau),$ respectively. Thus, we can rewrite (ref) as

equation[equation omitted — 106 chars of source]

where $\alpha_{0}=(\beta _{0}^{\prime },\delta _{0}^{\prime })^{\prime }.$ In this paper, our interest lies in performing uniform inference for the high-dimensional slope parameter $\alpha_0,$ allowing for $p>n.$

Let $\bm{Y}:= (Y_{1},\ldots ,Y_{n})^{\prime }.$ The residual sum of squares is

equation[equation omitted — 256 chars of source]

The Lasso estimator for threshold regression can thus be defined as the one-step minimizer:

equation[equation omitted — 287 chars of source]

where $\mathcal{B}$ is the parameter space for $\alpha_0,$ and $\lambda$ is a tuning parameter. The $(2p \times 2p)$ diagonal weighting matrix is denoted as follows:

equation[equation omitted — 134 chars of source]

Furthermore, we can rewrite the penalty term as

equation*[equation* omitted — 428 chars of source]

Meanwhile, the one-step minimizer $(\widehat{\alpha }, \widehat{\tau })$ in (ref) can be regarded as a two-step minimizer:

(i) For each $\tau \in \mathbb{T}$, $\widehat{\alpha }(\tau )$ is defined as

align[align omitted — 209 chars of source]

(ii) Define $\widehat{\tau }$ as the estimator of $\tau _{0}$ such that:

equation[equation omitted — 221 chars of source]

Note that these estimators are weighted Lasso estimators that use a data-dependent $\ell_1$ penalty to balance covariates. chiang_rodrigue_sasaki_2023 summarize various ways to impose weights depending on different situations in Remark B.1. Additionally, in practice, $\widehat \tau$ is selected from the potential values of threshold variable $Q$ over $\{Q_1,\dots, Q_n\}.$ \footnote{If $n$ is very large, $\mathbb{T}$ can be approximated by a grid of $N$ evaluation points; see p.4 in hansen2000.} This selection results in an interval, and the maximum of the interval is chosen as the estimator $\widehat{\tau}.$

Oracle Inequalities

After recalling the Lasso estimator, we proceed to establish the oracle inequalities for the estimators in (ref). First, we make the following assumptions, some of which are modified from lee2016.

assmLet $\left\{ X_i, Q_i, U_i\right\}_{i=1}^n$ denote a sequence of independently distributed random variables. (i) For the parameter space $\mathcal {B}$ for $\alpha_0$, any $\alpha := (\alpha_1, \dots, \alpha_{2p}) \in \mathcal {B} \subset \mathbb{R}^{2p},$ including $\alpha_0$, satisfies $|\alpha|_{\infty}\le C_1,$ for some constant $C_1 >0$. Furthermore, $\frac{s_0^2|\delta_0|_1^2\log{p}}{n}=o_p(1).$ (ii) The threshold variable $Q_i,$ $i=1,...,n,$ is continuously distributed on $[0,1]$ with intensity function $f_Q(\tau).$ The parameter $\tau_0$ lies in $\mathbb{T}=[t_0, t_1],$ where $0<t_0<t_1<1.$ (iii) The covariates $X_i,$ $i=1,...,n,$ satisfy $\max_{1\le j \le p}E\left[\left(X_i^{(j)}\right)^4\right]\le C_2^4$ and $\min_{1\le j\le p}\\ E\left[\left(X_i^{(j)}\left(t_0\right)\right)^2\right]\ge C_3^2,$ for some constants $ C_2$ and $C_3$. Additionally, $ E\left[X_i^{(j)} X_i^{(l)} \vert Q_i=\tau\right]$ is continuous and bounded when $\tau$ is in a neighborhood of $\tau_0,$ for all $1\le j,l \le p$. (iv) The error terms $U_i,$ $i=1,...,n,$ satisfy $ E(U_i|X_i,Q_i) = 0$ and $\max_{1\leq i \leq n}E(U_i^4)\le C_4<\infty,$ for a positive constant $C_4$. Additionally, $ E\left[X_i^{(j)} X_i^{(l)} U_i^{2} \vert Q_i=\tau\right]$ is continuous and bounded when $\tau$ is in a neighborhood of $\tau_0,$ for all $1\le j,l \le p$. (v) $\frac{\sqrt{EM_{UX}^2}\sqrt{\log{p}}}{\sqrt{n}}=o_p(1),$ where $M_{UX}= \max_{1\le i\le n} \max_{1\le j \le p} \left|U_{i}X_{i}^{(j)}\right|$. (vi) $\frac{\sqrt{EM_{XX}^2}\sqrt{\log{p}}}{\sqrt{n}}=o_p(1),$ where $M_{XX}= \max_{1\le i\le n} \max_{1\le j,l \le p} \left| X_i^{(j)} X_i^{(l)}\right|$.

The first part of Assumption (ref) (i) restricts the magnitude of slope parameters. The second part further implies that $s_0$ and $|\delta_0|_1$ may grow with $n$, which ensures that Assumption 6 in Appendix E of lee2016 holds for sufficiently large $n$. Assumption (ref) (ii) ensures that there are no ties among the $Q_i$s. We can empirically transform the distribution of the threshold variables to a uniform distribution. Suppose that the threshold variable $\{\tilde{Q}\}$ has a continuous distribution for which the cumulative distribution function is $F_{\tilde{Q}}$. The probability integral transform implies that the random variable $Q$ has a standard uniform distribution where $Q$ is defined as $Q=F_{\tilde{Q}}(\tilde{Q}).$ To transform the marginals, we compute $Q_i=\widehat{F}_{\tilde{Q}}(\tilde{ Q}_i)=\frac{\text{rank of $\tilde{Q}_i$ among $\left\{\tilde{ Q}_i\right\}_{i=1}^n$ }}{n},$ where $\widehat{F}_{\tilde{Q}}$ denotes the empirical distribution functions of the data $\left\{\tilde{ Q}_i\right\}_{i=1}^n$. In particular, as a result of a continuous distribution, there is no tie among $\left\{\tilde{ Q}_i\right\}_{i=1}^n$. \footnote{We maintain the dependence between $Q_i$ and $X_i$ in the proof, and we will show in Section (ref) that the performance of our estimator does not depend on whether $Q_i$ is among the components of $X_i.$} Applying the Cauchy-Schwarz inequality under Assumptions (ref) (iii) and (iv) yields $\max_{1\le j,l \le p} E\left[X_i^{(j)} X_i^{(l)}\right]\le C_2^2$ uniformly in $i$; $\max_{1\le j \le p} \mathrm{Var} \left(U_iX_{i}^{(j)}\right)$, $\max_{1\le j,l\le p} \mathrm{Var} \left(X_i^{(j)} X_i^{(l)}\right),$ $\max_{1\le j,l\le p} \mathrm{Var}\left(X_i^{(j)} X_i^{(l)}\bm{1}\{Q_i <\tau_0 \}\right),$ and \\ $\max_{1\le j \le p} \mathrm{Var} \left(X_{i}^{(j)}(t_0)\right)^2$ are bounded uniformly in $i$.

remAssumption (ref) imposes weaker conditions on the covariates and error terms compared to the fixed covariates and gaussian errors in lee2016 and the sub-gaussian covariates and errors in canerkock2017, as it allows for heteroskedastic non-subgaussian error terms and non-subgaussian covariates.

Define

equation[equation omitted — 82 chars of source]

as the tuning parameter in ((ref)) for a constant $A \geq 0$ and a fixed constant $\mu\in(0,1).$

lemSuppose that Assumption (ref) holds. Let $(\widehat{\alpha},\widehat{\tau})$ be the Lasso estimator defined by (ref). Then, with probability at least $1-C(logn)^{-1},$ we have \begin{align} \left\Vert \widehat{f}-f_{0}\right\Vert _{n} & \leq \sqrt{ (6+2\mu_2)C_1 \sqrt{C_2^2+\mu_1\lambda}}\sqrt{s_0 \lambda}. \end{align}

Lemma (ref) provides a non-asymptotic upper bound on the prediction error, regardless of whether the specification is a linear or threshold regression, as in Theorem 1 of lee2016. The prediction error is consistent as $n\to\infty$, $p\to\infty,$ and $s_0\lambda\to 0.$ The lemma plays an important role in deriving oracle inequalities in Theorem (ref) for linear models and Theorem (ref) for threshold models.

Next, we impose the standard assumptions in high-dimensional regression models. To this end, we define the population covariance matrix $\bm{\Sigma}(\tau)=E\left[1/n \sum_{i=1}^n\bm{X}_i(\tau)\bm{X}_i(\tau)'\right]$, $\bm {M}={E}[1/n \sum_{i=1}^n{X_i}{X_i}']$, $\bm {M}(\tau)=E\left[1/n \sum_{i=1}^n {X_i}(\tau){X_i}(\tau) '\right],$ and $\bm {N}(\tau)=\bm {M}-\bm {M}(\tau)$. $\bm{\Sigma}(\tau)$ can be represented as a $2\times2$ matrix due to the properties of the indicator function, i.e.,

equation*[equation* omitted — 153 chars of source]

Meanwhile, we define the following population uniform adaptive restricted eigenvalue and impose certain assumptions,

\scalebox{0.85}{\parbox{0.1\linewidth}{

equation*[equation* omitted — 360 chars of source]

}}

assm(i) $\bm{M}(\tau)$ and $\bm{N}(\tau)$ are non-singular.\\ (ii) (Uniform Adaptive Restricted Eigenvalue Condition) For a positive number $c_0,$ and some set $\mathbb{S} \subset \mathbb{R}$, the following condition holds \begin{equation} \kappa(s_0, c_0, \mathbb{S},\bm{\Sigma})>0. \end{equation}

Assumption (ref) (i) is a standard assumption for model estimation. Assumption (ref) (ii) is a uniform adaptive restricted eigenvalue condition, which is a common and high-level condition in the literature of high-dimensional econometrics and statistics. This condition can be relaxed if $\bm{\Sigma}(\tau)$ has full rank. Moreover, $\bm{\Sigma}(\tau)$ is invertible by applying Theorem 2.1 (ii) in lu2002inverses under Assumption (ref) (i). We then can do the gaussian elimination to obtain

equation[equation omitted — 245 chars of source]

Thus, Assumption (ref) (ii) holds under the non-singularity conditions for $\bm{M}(\tau)$ and $\bm{N}(\tau).$ Lemma (ref) shows that $1/n \bm {X }(\tau)\bm{X }(\tau)'$ uniformly converges to $\bm{\Sigma}(\tau);$ therefore, the empirical adaptive restricted condition holds as stated in Lemma (ref).

Given that $\tau_0$ is unknown, we impose that the restricted eigenvalue condition holds uniformly over $\tau.$ Here, we analyze two separate cases. When $\delta_0=0,$ Assumption (ref) is required to hold uniformly with $\mathbb{S} =\mathbb{T},$ the entire parameter space for $\tau_0,$ since $\tau_0$ is not identified. When $\delta_0\ne0,$ this condition holds uniformly in a neighborhood of $\tau_0$ for the identification of $\tau_0$. The Uniform Adaptive Restricted Eigenvalue (UARE) Condition is applied to tighten the bound in Lemma (ref) for establishing the oracle inequalities for the prediction error as well as the $\ell _{1}$ estimation error for the parameters. Although we consider two cases separately, similar to lee2016, we can make predictions and estimate $\alpha_0$ without pretesting the existence of the threshold effect.

Case I. No Threshold Effect.

In the case where $\delta_0=0$, the true model simplifies to a linear model $Y_{i}= X_{i}^{\prime }\beta _{0}+U_{i}.$ The model ((ref)) is thus much more over-parameterized, but we can still estimate the slope parameter vector $\alpha_0$ precisely, as shown in Theorem (ref).

thmSupposed that $\delta_0=0$ and that Assumptions (ref)-(ref) hold with $\kappa = \\ \kappa \left(s_0, \frac{1+\mu}{1-\mu },\mathbb{T},\bm{\Sigma}\right).$ Let $(\widehat{\alpha},\widehat{\tau})$ be the Lasso estimator from (ref) with $\lambda$ satisfying (ref). Then, as $n\rightarrow \infty,$ with probability at least $1-C(logn)^{-1},$ we have \begin{equation*} \begin{aligned} &\left\Vert \widehat{f}-f_{0}\right\Vert _{n} \leq \frac{2\sqrt{2}}{\kappa}\left( \sqrt{C_2^2+\mu_1\lambda}\right)\sqrt{s_0}\lambda, \\ &\left| \widehat{\alpha }-\alpha _{0}\right| _{1} \le\frac{4\sqrt{2}}{\left( 1-\mu \right)\kappa ^{2}}\frac{C_2^2+\mu_1\lambda}{\sqrt{C_3^2-\mu_1\lambda}}{s_0\lambda }. \end{aligned} \end{equation*} Furthermore, these bounds are valid uniformly over the $l_0$-ball $$\mathcal{A}^{({1})}_{\ell_0}(s_0)=\left\{\alpha_0\in \mathbb{R}^{2p} \mid|\beta_0|_{\infty}\le C_1,\left|\beta_0\right|_0 \le s_0 ,\delta_0=0\right\}.$$

The bound on the prediction norm in Theorem (ref) is much tighter than that in Lemma (ref). Compared with the oracle inequalities established in the high-dimensional linear model literature (e.g. Bickel2009, geer2014), Theorem (ref) provides results of the same magnitude, indicating that our estimation procedure remains valid for linear models, despite being more overparameterized due to the inclusion of the additional parameters $\delta$ and $\tau$.

Case II. Fixed Threshold Effect.

In the case where $\delta_0\ne0,$ we assume that the true model has a well-identified and fixed threshold effect.

assm[Identifiability under Sparsity and Discontinuity of Regression] For any $\eta$ and $\tau $ such that $ \eta<\left\vert \tau -\tau _{0}\right\vert $ and $\alpha \in \left\{ \mathcal{\alpha }:\left|\alpha \right|_0 \le s_0 \right\} $, there exists a constant $C_4>0$ such that, with probability approaching one, \begin{equation*} \left\Vert f_{\left( \alpha ,\tau \right) }-f_0 \right\Vert _{n}^{2} > C_4\eta. \end{equation*}

Assumption (ref) states identifiability of $\tau_0.$ Its validity was studied in Appendix B.1 (pages A7–A8) of lee2016 when Assumption (ref) holds. \footnote{We omit the restriction $\eta\ge\min_{i}\left\vert Q_{i}-\tau_0\right\vert$ since $\eta > \min_{i}\left\vert Q_{i}-\tau_0\right\vert$ holds in the random design with a continuous threshold variable $Q$.} When $\tau_0$ is known, the UARE condition is only required to hold uniformly in a neighborhood of $\tau_0.$ We derive an upper bound for $|\widehat{\tau}-\tau_0|$ in Lemma (ref) and thus define $$\eta^\ast = \frac{ 2 (3+\mu_2)C_{1}}{C_4} \sqrt{C_2^2+\mu_1\lambda} s_0\lambda, \quad \mathbb{S} = \left\{ \left\vert \tau -\tau _{0}\right\vert \leq \eta ^{\ast }\right\}.$$

assm[Smoothness of Design] For any $\eta >0,$ there exists a constant $C_5<\infty $ such that with probability to one, \begin{align} \sup_{1 \le j,l\le p}\sup_{\left\vert \tau -\tau _{0}\right\vert <\eta } \frac{1}{n}\sum_{i=1}^{n}\left\vert X_{i}^{\left( j\right) } X_{i}^{\left( l\right) }\right\vert \left\vert \bm{1}\{Q_i <\tau_0 \} -\bm{1}\{Q_i <\tau\}\right\vert\leq C_5\eta, \end{align}
assm[Well-defined second moments] For any $\eta$ such that \\ $1/n \leq \eta \leq \sqrt{ (6+2\mu_2)C_1 \sqrt{C_2^2+\mu_1\lambda}}\sqrt{s_0 \lambda},$ $h_n^2({\eta})$ is bounded with probability approaching one, where \begin{equation} h_n^2({\eta}) = \frac{1}{2n\eta}\sum_{i= \min\{1,[n\left( \tau _{0}-\eta \right)]\}}^{\max\{[n\left( \tau _{0}+\eta] \right),n\}} \left(X_i'\delta_0\right)^2, \end{equation} and $[\cdot]$ denotes an integer part of any real number.

Assumptions (ref) and (ref) are similar to those in lee2016 and canerkock2017. Lemma (ref) shows that Assumption (ref) holds automatically under Assumptions (ref) and (ref).

thmSuppose that $\delta_0\ne0$ and that Assumptions (ref) and (ref) hold with $\kappa =\kappa (s_0, \frac{2+\mu}{1-\mu },\mathbb{S},\bm{\Sigma} ).$ Furthermore, suppose that Assumptions (ref), (ref) and (ref) hold. Let $(\widehat{\alpha },\widehat{\tau })$ be the Lasso estimator from (ref) with $\lambda$ satisfying (ref). Then, $n\rightarrow \infty,$ with probability at least $1-C(logn)^{-1},$ we have \begin{equation*} \begin{aligned} &\left\Vert \widehat{f}-f_{0}\right\Vert _{n} \leq 6\frac{\sqrt{C_2^2+\mu\lambda}}{\kappa} \sqrt{s_0}\lambda, \\ &\left| \widehat{\alpha }-\alpha _{0}\right| _{1} \leq \frac{36(C_2^2+\mu\lambda)}{\kappa^2(1-\mu)\sqrt{C_3^2-\mu\lambda}} s_0\lambda , \\ &\left\vert \widehat{\tau}-\tau _{0}\right\vert \leq \left( \frac{3\left( 1+\mu \right) \sqrt{\left(C_2^2+\mu\lambda\right)}}{(1-\mu )\sqrt{\left(C_3^2-\mu\lambda\right)}}+1\right) \frac{12(C_2^2+\mu\lambda)}{\kappa^2C_4} s_0\lambda^2. \end{aligned} \end{equation*}Furthermore, these bounds are valid uniformly over the $l_0$-ball $$\mathcal{A}^{({2})}_{\ell_0}(s_0)=\left\{\alpha_0\in \mathbb{R}^{2p} \mid|\alpha_0|_{\infty}\le C_1,\left|\alpha_0 \right|_0 \le s_0 ,\delta_0\ne0 \right\}.$$

We establish inequalities of the same order (up to a constant) in Theorem (ref) as those in Theorem (ref). These results hold uniformly over the parameter space $\mathcal{B}_{\ell_0}(s_0)$, defined as $$\mathcal{B}_{\ell_0}(s_0)= \mathcal{A}^{({1})}_{\ell_0}(s_0)\cup\mathcal{A}^{({2})}_{\ell_0}(s_0)=\left\{\alpha_0\in \mathbb{R}^{2p} \mid|\alpha_0|_{\infty}\le C,\left|\alpha_0\right|_0\le s_0\right\}.$$ Regarding the super-consistency of $\widehat{\tau}$, lee2016 notes that the least squares objective function is locally linear rather than locally quadratic, in a neighborhood of $\tau_0.$

remWe derive the asymptotic independence between $\widehat{\tau}$ and $\widehat{\alpha}$ based on the separability of the objective function in Section (ref). The estimation error bound for $\tau_0$ can be further tightened using the two-step estimation procedure proposed in Lee03072018, which also derives the asymptotic distribution of $\widehat{\tau},$ which is beyond the focus of the work.

The main contribution of this section is that we use concentration inequality to establish oracle inequalities with non-subgaussian random regressors and heteroskedastic non-subgaussian error terms. These oracle inequalities are the basis for developing the uniform inference theory.

The Debiased Lasso Estimator and Uniform Inference

The Debiased Lasso Estimator

To perform uniform inference for the slope parameters, we first construct the debiased Lasso estimator proposed by geer2014 in our high-dimensional threshold regression model as follows: \footnote{This estimator is obtained by inverting the Karush-Kuhn-Tucker (KKT) conditions.}

equation[equation omitted — 223 chars of source]

where $\widehat{\alpha}(\widehat{\tau})$ is obtained from (ref), and $\widehat{\bm{\Theta}}(\widehat{\tau})$ is an approximate inverse of the empirical Gram matrix $\widehat{\bm{\Sigma}}(\widehat{\tau}) = \bm {X}(\widehat{\tau})\bm {X}(\widehat{\tau})' /n$ because $\widehat{\bm{\Sigma}}(\widehat{\tau})$ is singular in our high-dimensional model.

We apply nodewise regression following Meinshausen2006 to construct an approximate inverse matrix. Since the inverse is built using (ref), nodewise regression is applied once or twice, depending on the invertibility of the empirical analogs of $\bm{M}(\tau)$ and $\bm{N}(\tau)$. For example, when $\tau n$ or $(1 - \tau) n$ is smaller than $p$, or when strong multicollinearity exists, $\widehat{\bm{M}}(\tau)$ and $\widehat{\bm{N}}(\tau)$ may become singular, necessitating multiple applications of nodewise regression.

The following discussion will focus on the debiased Lasso estimator under two cases.

Case I. No Threshold Effect

In the case where $\delta_0=0$, the true model simplifies to a linear model $Y = X\beta_0 + U.$ Substituting this into $\eqref{despCLASSO}$ yields

equation[equation omitted — 797 chars of source]

where $\Delta(\tau) = \sqrt{n} \left( \widehat{\bm{\Theta}}(\tau) \widehat{\bm{\Sigma}}(\tau) - I_{2p} \right) \left(\widehat{\alpha}(\widehat \tau) - \alpha_0\right).$ The second equality holds due to $\delta_0 = 0.$

We define a $(2p\times1)$ vector $g$ with $|g|_2 = 1$ and let $H =\{j=1,\dots,2p\mid g_j\ne0\}$ with cardinality $|H| = h<p.$ $H$ is the index set of the coefficients involved in the hypothesis to be tested. We allow for $h\rightarrow \infty$ but require $h/n\rightarrow 0,$ as $n \rightarrow \infty.$ Our test $g'\alpha=g'\alpha_0$ could thus involve an increasing number of parameters. By the Cauchy–Schwarz inequality, we have $|g|_1\le\sqrt{h}$. In particular, $g=e_j$ represents the case where we test only a single coefficient, where $e_j$ is the $2p \times 1$ unit vector with the $j$-th element being one.

Our focus is on

equation[equation omitted — 180 chars of source]

and we will derive its asymptotic distribution by applying a central limit theorem to $g'\widehat{\bm{\Theta}}(\widehat{\tau})\bm {X}(\widehat{\tau})'U/n^{1/2}$ and by showing that $g'\Delta(\widehat{\tau})$ is asymptotically negligible.

Case II. Fixed Threshold Effect

This subsection explores the case where the threshold effect is well-identified and fixed. Following a procedure similar to Section (ref), we substitute $Y = \bm {X}( \tau_0) \alpha_0 + U$ into $\eqref{despCLASSO},$ yielding

equation[equation omitted — 424 chars of source]

In this case, our focus is on

equation[equation omitted — 360 chars of source]

and we will derive its asymptotic distribution by applying a central limit theorem to $g'\widehat{\bm{\Theta}}(\widehat{\tau})\bm {X}(\widehat{\tau})'U/n^{1/2}$ and by showing that $g'\widehat{\bm{\Theta}}\left(\widehat{\tau})(\bm {X}(\widehat{\tau})'\bm {X}(\tau_0)-\bm {X}(\widehat{\tau})'\bm {X}(\widehat{\tau})\right)\alpha_0 /n^{1/2}$ and $g'\Delta(\widehat{\tau})$ are asymptotically negligible.

Constructing the Approximate Inverse \texorpdfstring{$\widehat{\bm{\Theta}}(\tau)$}{widehat{bm{Theta}}(tau)}

In this subsection, we formalize the process of constructing the approximate inverse matrix $\widehat{\bm{\Theta}}(\tau)$ of the singular empirical Gram matrix. The approach closely follows that of geer2014, with the additional requirement of verifying that the specified conditions are met.

We seek a well-behaved $\widehat{\bm{\Theta}}({\tau})$ and examine the asymptotic properties of $\widehat{\bm{\Theta}}(\tau)$ uniformly across $\tau\in \mathbb{T}$. Recalling (ref), we have $$\bm{\Theta}(\tau) =\bm{\Sigma}(\tau)^{-1}= {

bmatrix[bmatrix omitted — 151 chars of source]

}.$$ Define $\bm {A}(\tau)=\bm {M}(\tau)^{-1}$ and $\bm {B}(\tau)=\bm {N}(\tau)^{-1}.$ We construct the approximate inverse $\widehat{\bm {A}}(\tau)$ of $\widehat{\bm {M}}(\tau)$ and $\widehat{\bm {B}}(\tau)$ of $\widehat{\bm {N}}(\tau),$ where $\widehat{\bm{M}}(\tau) = \frac{1}{n}\sum_{i=1}^{n}{X_i}{X_i}'\bm{1}\{Q_i <\tau\}$ and $\widehat{\bm{N}}(\tau) = \frac{1}{n}\sum_{i=1}^{n}{X_i}{X_i}'\bm{1}\{Q_i \ge\tau \},$ to build the approximate inverse matrix $\widehat{\bm{\Theta}}({\tau}).$

Denote the $(p\times 1)$ vector by $\widetilde{X}^{(j)}_{i}(\tau )=X_{i}^{(j)}\bm{1}\{Q_i \ge\tau \}$ and the $(n\times p)$ matrix by $ \widetilde{X}(\tau).$ Let ${X}^{(-j)}(\tau )$ and $ \widetilde{X}^{(-j)}(\tau)$ denote the submatrices of $ X(\tau)$ and $ \widetilde{X}(\tau),$ respectively, without the $j$-th column. We study the following nodewise regression models with covariates orthogonal to the error terms in $L_2$ for all $ j = 1,\dots, p$, we consider $$X^{(j)}(\tau)=X^{(-j)}(\tau)'{\gamma}_{0,j}(\tau)+\upsilon^{(j)},$$ $$\widetilde{X}^{(j)}(\tau)=\widetilde{X}^{(-j)}(\tau)'\widetilde{\gamma}_{0,j}(\tau)+\widetilde{\upsilon}^{(j)}. \footnote{See Appendix B of \cite{canerkock2018} for the details about the covariance matrix's representation of the regression coefficients.}$$ $\upsilon^{(j)}$ and $\widetilde{\upsilon}^{(j)}$ may be functions of $\tau$ even though they are independence of $Q$.

We then impose the following assumption to control the tail distribution of $\left|\upsilon^{(j)} _{i} X_{i}^{(l)}\right|$ and $\left|\widetilde{\upsilon}^{(j)} _{i}X_{i}^{(l)}\right|,$ allowing us to apply the oracle inequalities in Section (ref) to the nodewise regressions.

assm(i) $\max_{1 \le j \le p } | \gamma_{0,j} |_{\infty} \le C$ and $\max_{1 \le j \le p } | \widetilde\gamma_{0,j} |_{\infty} \le C'$, for some positive constants $C$ and $C'.$\\ (ii) For $i=1,...,n,$ and $j=1,...,p,$ $ E\left[\upsilon^{(j)}_{i}\middle|X_i,Q_i\right] = 0$ and $ E\left[\widetilde{\upsilon}^{(j)}_{i}\middle|X_i,Q_i\right]= 0;$ $E\left[ \left(\upsilon^{(j)}_{i}\right)^{2} \right]$ and $E\left[ \left(\widetilde{\upsilon}^{(j)}_{i}\right)^{2}\right]$ are uniformly bounded in $j = 1,\dots, p.$\\ (iii) $\frac{\sqrt{EM_{\upsilon X}^2}\sqrt{\log{p}}}{\sqrt{n}}<\infty$ and $\frac{\sqrt{EM_{\widetilde \upsilon X}^2}\sqrt{\log{p}}}{\sqrt{n}}<\infty$, where $M_{\upsilon X}= \max_{1\le i\le n} \max_{1\le l \le p} \left|\upsilon ^{(j)} _{i}X_{i}^{(l)} \right|$ and $M_{\widetilde \upsilon X}= \max_{1\le i\le n} \max_{1\le l \le p} \left|\widetilde \upsilon ^{(j)} _{i}X_{i}^{(l)} \right|.$

Now, we begin constructing $\widehat{\bm {A}}(\tau).$ Given any $\tau \in \mathbb{T} $, for each $j=1,...,p,$ the Lasso estimator for the nodewise regression is given by

equation[equation omitted — 217 chars of source]

where $\bm{\widehat{\Gamma}_j}(\tau ):=\text{diag}\left\{ \left\Vert X^{(l)}(\tau )\right\Vert _{n}, l=1,...,p,l\ne j\right\},$ with components of $\widehat{\gamma}_j(\tau)=\{\widehat{\gamma}_j^{(k)}(\tau);\ k=1,...,p,\ k\neq j\}$. We choose $\lambda_{node,j} = \lambda_{node},$ required for the validity of Lemma (ref). Define

equation[equation omitted — 396 chars of source]

and $\widehat{\bm{Z}}(\tau)^2 = diag \left( \widehat{z}_1(\tau)^2, \dots, \widehat{z}_p(\tau)^2\right)$, where

equation[equation omitted — 206 chars of source]

We thus construct

equation[equation omitted — 92 chars of source]

Next, we show that $\widehat{\bm {A}}(\tau)$ is an approximate inverse matrix of $\widehat{\bm {M}}(\tau)$. Let $\widehat{A}_j(\tau)$ denote the $j$-th row of $\widehat{\bm {A}}(\tau)$. Thus, $\widehat{A}_j(\tau)=\widehat{C_j}(\tau)/\widehat{z}_j(\tau)^2$. Denoting by $\widetilde{e}_j$ the $j$-th unit vector, the KKT conditions imply that

equation[equation omitted — 195 chars of source]

Similarly, we use the same process to construct $\widehat{\bm {B}}(\tau).$ Given any $\tau \in \mathbb{T},$ for each $j=1,...,p,$ define

align*[align* omitted — 470 chars of source]

with components of $\widehat{\widetilde{\gamma}}_{j}(\tau) =\{\widehat{\widetilde{\gamma}}_j^{(k)}(\tau): k=1,...,p, k \neq j\}$. Meanwhile, define

equation[equation omitted — 466 chars of source]

and $\widehat{\widetilde{\bm{Z}}}(\tau)^2 = \text{diag}\left(\widehat{\widetilde{z}}_1(\tau)^2, \dots, \widehat{\widetilde{z}}_p(\tau)^2\right)$ with

equation*[equation* omitted — 280 chars of source]

We then construct

equation[equation omitted — 132 chars of source]

Therefore, we obtain

equation[equation omitted — 251 chars of source]

and provide the asymptotic properties of $\widehat{\bm{\Theta}}(\tau)$ in the following. Define $\Bar{s}= \sup_{\tau\in\mathbb{T}}\max_ {j\in H} s_j(\tau)$, $s_j(\tau)=\vert S_j(\tau)\vert$, and $ S_j(\tau) =\{ i =1,...,2p: \Theta_{j,i}(\tau)\ne 0\},$

lemSuppose that Assumptions (ref)-(ref) hold and set $\lambda_{node}=\frac{C}{\mu}\sqrt{\frac{\log{p}}{n}} $. Then, \begin{align} \max_{j\in H}\sup_{\tau\in \mathbb{T}} \left|\widehat{\Theta}_{j}(\tau) -\Theta_{j}(\tau) \right|_1&=O_p \left(\Bar{s} \sqrt{\frac{\log{p}}{n}}\right)\\ \max_{j\in H}\sup_{\tau\in \mathbb{T}} \left|\widehat{\Theta}_{j}(\tau) -\Theta_{j}(\tau) \right|_2&=O_p \left(\sqrt{\frac{\Bar{s} \log{p}}{n}}\right)\\ \max_{j\in H}\sup_{\tau\in \mathbb{T}} \left|\widehat{\Theta}_{j}(\tau) \right|_{1}&=O_p \left(\sqrt{\Bar{s}}\right)\\ \max_{j\in H}\sup_{\tau\in \mathbb{T}} \left| \widehat{\Theta}_{j}(\tau)'\widehat{\bm{\Sigma}}(\tau) - {e}_j' \right|_{\infty}&=O_p \left(\sqrt{\frac{\log{p}}{n}}\right) \end{align}

We derive the approximation error that arises from the inversion of the covariance matrix. For instance, ((ref)) establishes an upper bound on the maximal absolute entry of the $j$-th row of $\widehat{\bm{\Theta}}(\tau)'\widehat{\bm{\Sigma}}(\tau)-\bm{I}_{2p}$. These bounds hold uniformly over the entire $\tau$ parameter space and are valid regardless of whether the underlying model is a linear or threshold regression. Moreover, they provide sufficient conditions to ensure that $g'\Delta(\tau)$ in ((ref)) and ((ref)) is asymptotically negligible.

Uniform Inference for the Debiased Lasso Estimator

This section derives the asymptotic distribution of tests in cases with no threshold effect and a fixed threshold effect, showing that their distributions are the same. Furthermore, we construct uniform confidence intervals for the parameters of interest that contract at the optimal rate. To this end, we first impose some assumptions to establish the validity of the asymptotically gaussian inference.

assm(i) $\max_{1\le j \le p}E\left[\left(X_i^{(j)}\right)^{12}\right]$ and $E\left[U_i^{8}\right]$ are bounded uniformly in i. $\frac{\sqrt{EM_{X^6}^2}\sqrt{\log{p}}}{\sqrt{n}}=o_p(1)$, $\frac{\sqrt{EM_{X^2U^2}^2}\sqrt{\log{p}}}{\sqrt{n}}=o_p(1)$ and $\frac{\sqrt{EM_{X^4U^2}^2}\sqrt{\log{p}}}{\sqrt{n}}=o_p(1),$ \\ where $M_{X^6}= \max_{1\le i\le n} \max_{1\le k,l,j\le p} \left|\left({X}^{(k)}_i{X}^{(l)}_i{X}^{(j)}_i\right)^2\right|$, $M_{X^2U^2}=\\ \max_{1\le i\le n} \max_{1\le j,l \le p}\left|X_i^{(j)} X_i^{(l)}U_{i}^2\right|,$ and $M_{X^4U^2}=\max_{1\le i\le n} \max_{1\le j,l \le p} \left|\left( X_i^{(j)} X_i^{(l)}U_{i}\right)^2\right|.$ (ii)$$(h)^{\frac{3}{2}}s_0^2\Bar{s}^2\frac{\log{p}}{\sqrt{n}}=o_{p}(1); \quad \frac{(h\Bar{s})^{3}}{n}=o_p(1). $$ (iii) $ \kappa(s_0, c_0, \mathbb{T} ,\bm{{\Sigma}}_{xu})$ and $ {\kappa}(s_0, c_0, \mathbb{T} ,\bm{\Sigma}) $ are bounded away from zero.\\ $\phi_{\max}(\bm {\Sigma}_{xu}(\tau)) $ and $\phi_{\max}(\bm{\Sigma}(\tau)) $ are bounded from above, for $\tau \in \mathbb{T}.$

Assumption (ref) provides sufficient conditions for applying the central limit theorem to establish the asymptotic distribution of the statistics. Assumption (ref) (i) controls the tail behavior of the covariates and the error terms. By Assumption (ref) (i), $\max_{1\le j,l \le p}\mathrm{Var} \left(X_i^{(j)} X_i^{(l)}U_{i}^2\right)$, $\max_{1\le k,l,j\le p}\mathrm{Var}\left({X}^{(k)}_i{X}^{(l)}_i{X}^{(j)}_i\right)^2,$ and $\max_{1\le j ,l\le p} \mathrm{Var} \left( X_i^{(j)} X_i^{(l)}U_{i}\right)^2$ are bounded from above uniformly in $i$. Assumption (ref) (ii) restricts the number of covariates ($p$), the number of parameters included in conducting joint inference ($h$), the sparsity of the population covariance matrix ($\Bar{s}$), and the sparsity of the slope parameters ($s_0$). Notably, the second part of Assumption (ref) (ii) is to verify the Lyapunov condition. Assumption (ref) (iii) restricts the eigenvalues of $\bm{\Sigma}_{xu}(\tau) $ and $\bm{\Sigma}(\tau)$, where $\bm{\Sigma}_{xu}(\tau) = E\left[1/n\sum_{i=1}^n\bm{X}_i(\tau)\bm{X}_i'(\tau){U}_i^2\right].$

thmSuppose that Assumptions (ref) to (ref) hold. Then, as $n\rightarrow \infty,$ we have \begin{equation} \frac{\sqrt{n}g'(\widehat{a}\left(\widehat{\tau}) -\alpha_{0 }\right)}{\sqrt{g'\widehat{\bm{\Theta}}(\widehat{\tau})\widehat{\bm{\Sigma}}_{xu} (\widehat{\tau})\widehat{\bm{\Theta}}(\widehat{\tau})'g}}\stackrel{d}{\to}N(0,1), \end{equation} uniformly in $\alpha_0\in\mathcal{B}_{\ell_0}(s_0).$ Furthermore, \begin{equation} \sup_{\alpha_0\in\mathcal{A}^{({1})}_{\ell_0}(s_0)}\left|g'\widehat{\bm{\Theta}}(\widehat{\tau})\widehat{\bm{\Sigma}}_{xu} (\widehat{\tau})\widehat{\bm{\Theta}}(\widehat{\tau})'g-g'\bm{\Theta}(\widehat{\tau})\bm{\Sigma}_{xu} (\widehat{\tau})\bm{\Theta}(\widehat{\tau})'g\right| = o_p(1) \end{equation} \begin{equation} \sup_{\alpha_0\in\mathcal{A}^{({2})}_{\ell_0}(s_0)}\left|g'\widehat{\bm{\Theta}}(\widehat{\tau})\widehat{\bm{\Sigma}}_{xu} (\widehat{\tau})\widehat{\bm{\Theta}}(\widehat{\tau})'g-g'\bm{\Theta}(\tau_0)\bm{\Sigma}_{xu} (\tau_0)\bm{\Theta}(\tau_0)'g\right| = o_p(1), \end{equation} where $\widehat{\bm{\Sigma}}_{xu}(\widehat{\tau})=\frac{1}{n}\sum_{i=1}^n \bm{X}_i(\widehat{\tau})\bm{X}_i(\widehat{\tau})'\left(\widehat{U}_i(\widehat{\tau})\right)^2.$

Thus, we establish the asymptotic distribution of tests involving an increasing number of slope parameters in the cases with no threshold effect and a fixed threshold effect, showing that their asymptotic distributions are identical. Additionally, we provide a uniformly consistent covariance matrix estimator for both cases. However, there is a slight difference between the limits of their asymptotic variances, since there is a true value for the threshold parameter in the latter case. Because we lack prior knowledge of the existence of a threshold effect, we simultaneously impose the assumptions of Theorems (ref) and (ref) to establish Theorem (ref). Here, the number of parameters involved in hypotheses is allowed to grow to infinity at a rate restricted by Assumption (ref)(ii).

In the case of a fixed number of parameters being tested, by (ref), we have

equation[equation omitted — 291 chars of source]

for a fixed cardinality $h.$ Thus, a $\chi^2$ test can be applied to test a hypothesis involving $h$ parameters simultaneously.

Furthermore, the asymptotic result can be applied to test for a threshold effect based on the active parameters. Define the active set as $\widehat J=\{j:\widehat{\delta}_j \neq 0\}$ and $\widehat s = |\widehat J|,$ where $\widehat s$ can grow with $n$ and is of the same magnitude of $s_0.$ Recall that $H$ is the index set of $g;$ let $H = \widehat J,$ and assign equal weights to the non-zero elements of $g.$ The null hypothesis is $H_0: \delta_{\widehat{J},0}=0.$ Subsequently, a t-test can be performed on $g'\widehat{a}(\widehat{\tau})$ to test for the existence of a threshold effect.

Next, we establish confidence intervals for the parameters of interest. Let $\varPhi(t)$ denote the cumulative distribution function (CDF) of the standard normal distribution, and let $z_{1-\alpha/2}$ be the $1-\alpha/2$ percentile of the standard normal distribution. Define $\widehat{\sigma}_j(\widehat{\tau})=\sqrt{e_j'\widehat{\bm{\Theta}}(\widehat{\tau})\widehat{\bm{\Sigma}}_{xu}(\widehat{\tau}) \widehat{\bm{\Theta}}(\widehat{\tau})'e_j}$ for all $ j\in\{1,..., 2p \}.$ Let $\text{diam}([a, b])$ denote the length of the interval $[a, b]\subset\mathbb{R}.$

thmSuppose that Assumptions (ref), (ref), (ref), (ref), (ref) and (ref) hold. Then, as $n \rightarrow \infty,$ we have \begin{equation} \sup_{t\in\mathbb{R}}\sup_{\alpha_0\in\mathcal{B}_{\ell_0}(s_0)}\left|\mathbb{P}\left\{\frac{\sqrt{n}g'(\widehat{a}(\widehat{\tau}) -\alpha_{0 })}{\sqrt{g'\widehat{\bm{\Theta}}( \widehat{\tau})\widehat{\bm{\Sigma}}_{xu}(\widehat{\tau}) \widehat{\bm{\Theta}}(\widehat{\tau})'g}}\le t\right\}-\varPhi(t)\right| {\to}0. \end{equation} Furthermore, for all $ j\in\{1,\dots,2p \},$ \begin{equation} \lim_{n\to\infty} \inf_{\alpha_0\in\mathcal{B}_{\ell_0}(s_0)} \mathbb{P}\left\{\alpha_{0}^{(j)}\in \left[\widehat{a}^{(j)}(\widehat{\tau})-z_{1-\frac{\alpha}{2}}\frac{\widehat{\sigma}_j( \widehat{\tau})}{\sqrt{n}},\widehat{a}^{(j)}( \widehat{\tau})+z_{1-\frac{\alpha}{2}}\frac{\widehat{\sigma}_j(\widehat{\tau})}{\sqrt{n}}\right]\right\}=1-\alpha, \end{equation} and \begin{equation} \sup_{\alpha_0\in\mathcal{B}_{\ell_0}(s_0)}diam\left(\left[\widehat{a}^{(j)}(\widehat{\tau})-z_{1-\frac{\alpha}{2}}\frac{\widehat{\sigma}_j(\widehat{\tau})}{\sqrt{n}},\widehat{a}^{(j)}(\widehat{\tau})+z_{1-\frac{\alpha}{2}}\frac{\widehat{\sigma}_j( \widehat{\tau})}{\sqrt{n}}\right]\right)=O_p\left(\frac{1}{\sqrt{n}}\right). \end{equation}

Therefore, we show that the convergence of a linear combination of the parameters of $\widehat{a}(\widehat{\tau})$ to the standard normal distribution is uniformly valid over the $\ell_0$-ball with a radius of at most $s_0.$ Researchers can perform uniform inference for high-dimensional slope parameters without specifying whether the specification is a linear or threshold regression. In addition, these confidence intervals are asymptotically honest and contract at the optimal rate.

remOur results can be readily extended to panel data models with fixed effects under strict exogeneity, as this framework accommodates heteroskedastic error terms and provides uniformly consistent estimators of the asymptotic covariance matrix.

Time Series Threshold Model

We establish the uniform inference theory for the debiased Lasso estimator in the high-dimensional time series threshold regression model, extending the model of ADAMEK20231114 by allowing for the existence of a threshold effect, with local projection threshold regression as a special case. We now assume that $(Y_i, X_i, Q_i)$ is a sequence of dependent data while still focusing on the model ((ref)) as follows:

equation[equation omitted — 108 chars of source]

Oracle Inequalities and Uniform Inference

We continue studying the Lasso estimator from equation ((ref)) without adding weights, as the data will be standardized. Some assumptions are from ADAMEK20231114, but we list them here for completeness and briefly discuss them. In the notation, $C$ represents an arbitrary positive finite constant, and its value may vary from line to line.

assmLet $\left\{ X_i, Q_i, U_i\right\}_{i=1}^n$ denote a sequence of random variables. Define $\boldsymbol{W}_i = (X_i^\prime, U_i)^\prime$ and suppose that there exist some constants $\bar m>m>2$, and $d\geq \max\{1,(\bar m/m-1)/(\bar m-2)\}$ such that (i) $\mathrm{E}\left[\bm {W}_i\right]=0,$ $\mathrm{E}\left[X_i U_i\right]=0,$ $\mathrm{E}\left[Q_iU_i\right]=0,$ and $\max_{1\leq j\leq p+1,\ 1\leq i\leq N}E\left|W_{i}^{(j)}\right|^{2\bar m} \leq C,$ for a positive constant $C.$ (ii) Let $\boldsymbol{s}_{N,i}$ denote a $k(N)$-dimensional triangular array that is $\alpha$-mixing of size $-d/(1/m-1/\bar{m})$ with $\sigma\text{-field}$ $\mathcal{F}^{\boldsymbol{s}}_i:=\sigma\left\lbrace\boldsymbol{s}_{n,i},\boldsymbol{s}_{n,i-1},\dots\right\rbrace$ such that $\boldsymbol{W}_i$ is $\mathcal{F}^{\boldsymbol{s}}_i$-measurable. The process $\left\lbrace W_{i}^{(j)}\right\rbrace$ is $L_{2m}$-near-epoch-dependent (NED) of size $-d$ on $\boldsymbol{s}_{N,i }$ with positive bounded NED constants \footnote{A sequence $\boldsymbol{W}_{N}$ is of size $-d$ if $\bm{W}_N = O(N^{-d-\varepsilon})$ for some $\varepsilon>0.$}, uniformly over $j=1,\ldots,n + 1.$ (iii) The threshold variable $Q_i,$ $i=1,...,n,$ is continuously distributed. The parameter $\tau_0$ lies in $\mathbb{T}=[t_0, t_1],$ where $0<t_0<t_1<1.$ (iv) $ E\left[X_i^{(j)} X_i^{(l)} \vert Q_i=\tau\right]$ and $E\left[X_i^{(j)} X_i^{(l)} U_i^{2} \vert Q_i=\tau\right]$ are continuous and bounded when $\tau$ is in a neighborhood of $\tau_0,$ for all $1\le j,l \le p$.

This assumption is the time series analogue of Assumption (ref). NED is a general form of dependence, including examples such as mixing, martingale, and mixingale dependence. It can be approximated by a mixing process.

We assume that $\alpha_0$ is weakly sparse, which is a more general condition than the exact sparsity assumption for cross-sectional data. Additionally, establishing oracle inequalities under this condition for independent data is not difficult.

assmFor some $0\leq r<1$ and sparsity level $s_r$, define the $2p$-dimensional sparse compact parameter space \begin{equation*} \mathcal{B}_{2p}(r, s_r) :=\left\{\alpha\in \mathbb{R}^{2p}: |\alpha|_r^r \leq s_r, \; |\alpha|_{\infty} \leq C, \, \exists C < \infty \right\}, \end{equation*} and assume that $\alpha_0\in\mathcal{B}_{2p}(r, s_r)$.

In addition, we define $\mathcal{A}^{({1})}_{2p}(r,s_r)=\left\{\alpha\in \mathbb{R}^{2p}: |\alpha|_r^r \leq s_r, \; |\alpha|_{\infty} \leq C, \delta_0=0\right\}$ and $\mathcal{A}^{({2})}_{2p}(r,s_r)\\ =\left\{\alpha\in \mathbb{R}^{2p}: |\alpha|_r^r \leq s_r, \; |\alpha|_{\infty} \leq C, \delta_0\ne0\right\}.$ Thus, $\mathcal{B}_{2p}(r, s_r)=\mathcal{A}^{({1})}_{2p}(r,s_r) \cup \mathcal{A}^{({2})}_{2p}(r,s_r).$

We impose the following standard compatibility conditions for high-dimensional regression models, as in ADAMEK20231114, which are stronger than the uniform adaptive restricted eigenvalue conditions in Assumption (ref) (ii), while they are still considered regular conditions for establishing the consistency of the Lasso estimator.

assmRecall $\boldsymbol\Sigma(\tau):=1/n\sum\limits_{i=1}^n\mathrm{E}\left[\boldsymbol{X}_i(\tau)\boldsymbol{X}_i(\tau)'\right].$ For a general index set $S$ with cardinality $|S|$, define the compatibility constant \begin{equation*} \phi_{\boldsymbol{\Sigma}(\tau)}^2(S):=\min_{\tau \in \mathbb{S}}\min_{\left\{ \gamma\neq 0:| \gamma_{S^c}|_1\leq 3| \gamma_{S}|_1\right\}}\left\{\frac{| S|\gamma'{\boldsymbol{\Sigma}(\tau)} \gamma}{| \gamma_{S}|^2_1}\right\}. \end{equation*} Assume that $\phi_{{\boldsymbol{\Sigma}}(\tau)}^2(S_\lambda)\geq 1/C$, which implies that \begin{equation*} |\gamma_{S_\lambda}|^2_1\leq\frac{| S_\lambda|\gamma'{\boldsymbol{\Sigma}(\tau)} \gamma}{\phi_{{\boldsymbol{\Sigma}}(\tau)}^2(S_\lambda)} \leq C\vert S_\lambda\vert \gamma'{\boldsymbol{\Sigma}(\tau)} \gamma, \end{equation*} for all $\gamma$ satisfying $|\gamma_{S^c_\lambda}|_1\leq 3| \gamma_{S_\lambda}|_1\neq 0.$

The compatibility conditions for $\bm{M}(\tau)$ and $\bm{M}$ can be defined accordingly.

thmSuppose that Assumptions (ref), (ref), (ref), (ref), (ref) and the conditions of Lemma (ref) hold. Then, as $n\rightarrow \infty,$ with probability at least $1-C\log \log n^{-1},$ we have \begin{equation*} \begin{aligned} &\left|\left|\widehat f -f_0 \right|\right|_{n}^2 \leq C\lambda^{2-r}s_r, \\ & |\widehat \alpha - \alpha_0|_1 \leq C\lambda^{1-r}s_r, \end{aligned} \end{equation*} when the fixed threshold effect exists, \begin{equation*} |\widehat\tau - \tau_0|_1 \leq C\lambda^{2-r}s_r. \end{equation*}

Theorem (ref) provides oracle inequalities for dependent data, in comparison to Theorems (ref) and (ref). We do not separate the results into two cases, as the oracle inequalities are qualitatively the same whether or not a fixed threshold effect exists. Additionally, we provide a non-asymptotic bound for $\widehat\tau$ when the threshold effect does exist.

We utilize the same nodewise regression approach as in Section (ref) to construct the approximate inverses of the empirical Gram matrices $\widehat{\bm M}(\tau)$ and $\widehat{\bm N}(\tau),$ denoted by $\widehat {\bm A}(\tau)$ and $\widehat {\bm B}(\tau),$ for obtaining the debiased Lasso estimator. Thus, for all $j = 1,\dots, p,$ $$X^{(j)}(\tau)=X^{(-j)}(\tau)'{\gamma}_{0,j}(\tau)+\upsilon^{(j)},$$ $$\widetilde{X}^{(j)}(\tau)=\widetilde{X}^{(-j)}(\tau)'\widetilde{\gamma}_{0,j}(\tau)+\widetilde{\upsilon}^{(j)}.$$

We provide the following assumptions for applying Theorem (ref) to the nodewise Lasso regressions.

assm(i) For some $0\leq r<1$ and sparsity levels $s_{r}^{(j)},$ $\widetilde{s}_{r}^{(j)},$ let $\gamma_{0,j}\in\mathcal{B}_{p-1}\left(r,s_{r}^{(j)}\right)$ and $\widetilde{\gamma}_{0,j}\in\mathcal{B}_{p-1}\left(r,\widetilde{s}_{r}^{(j)}\right),$ $\forall j=1,...,p.$\\ (ii) For $i=1,...,n,$ and $j=1,...,p,$ $E\left[\upsilon_{i}^{(j)}\right]=0;$ $E\left[\widetilde{\upsilon}_{i}^{(j)}\right]=0,$ $ E\left[\upsilon^{(j)}_{i}\middle|X_i,Q_i\right] = 0$ and $ E\left[\widetilde{\upsilon}^{(j)}_{i}\middle|X_i,Q_i\right]= 0;$ $\max_{1\leq j\leq p,\ 1\leq i\leq n} E \left[\left|\upsilon_{i}^{(j)}\right|^{2\bar m} \right] \leq C,$ $\max_{1\leq j\leq p,\ 1\leq i\leq n} E \left[\left|\widetilde{\upsilon}_{i}^{(j)}\right|^{2\bar m} \right]\\ \leq C'$.\\ (iii) $\bm{M}(\tau)$ and $\bm{N}(\tau)$ are non-singular.

Assumption (ref) is analogous to Assumptions 4 and 5 in ADAMEK20231114. Assumption (ref) (iii) implies that $\bm{\Sigma}(\tau)$ is invertible, as discussed below Assumption (ref).

Therefore, we have

equation*[equation* omitted — 665 chars of source]

as in (ref).

Recall the debiased Lasso estimators from Section (ref): in the case of no threshold effect, $$\sqrt{n}g'(\widehat{a}(\widehat{\tau}) -\alpha_{0})=g'\widehat{\bm{\Theta}}(\widehat{\tau})\bm {X}(\widehat{\tau})'U/n^{1/2} -g'\Delta(\widehat{\tau}),$$ and in the case of fixed threshold effect,

equation*[equation* omitted — 349 chars of source]

Define $\bm{Z}(\tau_0)^2 = diag \left( z_1(\tau_0)^2, \dots, z_p(\tau_0)^2\right)$ and $\widetilde{\bm{Z}}(\tau_0)^2 = diag \left( \widetilde{z}_1(\tau_0)^2, \dots, \widetilde{z}_p(\tau_0)^2\right),$ where $z_j(\tau_0)^2:=1/n\sum_{i=1}^nE\left[\left(\upsilon_{i}^{(j)}(\tau_0)\right)^2\right]$ and $\widetilde{z}_j(\tau_0)^2:=1/n\sum_{i=1}^nE\left[\left(\widetilde{\upsilon}_{i}^{(j)}(\tau_0)\right)^2\right]$ from population nodewise regression. Furthermore, we define the long-run covariance matrices ${\boldsymbol{\Omega}}_{p,n} = E\left[1/n\left(\sum_{i=1}^n \boldsymbol{w}_i\right)\left(\sum_{i=1}^n \boldsymbol{w}'_i\right)\right],$ ${\widetilde{\boldsymbol{\Omega}}}_{p,n} = E\left[1/n\left(\sum_{i=1}^n \widetilde{\boldsymbol{w}}_i\right)\left(\sum_{i=1}^n \widetilde{\boldsymbol{w}}'_i\right)\right],$ and $\overline{\boldsymbol{\Omega}}_{p,n} = E\left[1/n\left(\sum_{i=1}^n \boldsymbol{w}_i\right)\left(\sum_{i=1}^n \widetilde{\boldsymbol{w}}'_i\right)\right],$ where $\boldsymbol{w}_i=\left(v_{i}^{(1)}u_i,\dots,v_{i}^{(p)}u_i\right)'$ and $\widetilde{\boldsymbol{w}}_i=\left(\widetilde{v}_{i}^{(1)}u_i,\dots,\widetilde{v}_{i}^{(p)}u_i\right)'.$ $\boldsymbol{\Omega}_{p,n},$ ${\widetilde{\boldsymbol{\Omega}}}_{p,n},$ and $\overline{\boldsymbol{\Omega}}_{p,n}$ can be rewritten as $\boldsymbol{\Omega}_{p,n}=\boldsymbol{\Xi}(0)+\\ \sum_{l=1}^{n-1}\left(\boldsymbol{\Xi}(l)+\boldsymbol{\Xi}^\prime(l)\right),$ $\widetilde{\boldsymbol{\Omega}}_{p,n}=\widetilde{\boldsymbol{\Xi}}(0)+\sum_{l=1}^{n-1}\left(\widetilde{\boldsymbol{\Xi}}(l)+\widetilde{\boldsymbol{\Xi}}^\prime(l)\right)$ and $\overline{\boldsymbol{\Omega}}_{p,n}=\overline{\boldsymbol{\Xi}}(0)+\\\sum_{l=1}^{n-1}\left(\overline{\boldsymbol{\Xi}}(l)+ \overline{\boldsymbol{\Xi}}^\prime(l)\right)$ where $\boldsymbol{\Xi}(l) = \frac{1}{n}\sum_{i=l+1}^nE\left[ \boldsymbol{w}_i \boldsymbol{w}_{i-l}^\prime\right],$ $\widetilde{\boldsymbol{\Xi}}(l) = \frac{1}{n}\sum_{i=l+1}^nE\left[ \widetilde{\boldsymbol{w}}_i \widetilde{\boldsymbol{w}}_{i-l}^\prime\right],$ and $\overline{\boldsymbol{\Xi}}(l) = \frac{1}{n}\sum_{i=l+1}^nE\left[{\boldsymbol{w}}_i \widetilde{\boldsymbol{w}}_{i-l}^\prime\right]$, as in ADAMEK20231114.

Let $\underline \lambda = \min_{j=1,...,2p} \lambda_j$ and $\bar \lambda = \max_{j=1,...,2p} \lambda_j,$ satisfy (ref) and let $\bar s_r = \\ \max\left\{\max_{j=1,\dots,p}s_r^{(j)},\max_{j=1,\dots,p}\widetilde{s}_r^{(j)}\right\}.$ Define $\lambda_{\min}= \min\{\lambda, \underset{\bar{}}{\lambda}\},$ $\lambda_{\max}= \max\{\lambda, \bar{\lambda}\},$ $s_{r,\max} = \max\{s_r, \bar{s}_r\}.$ Similar to Section (ref), we define $g$ as a $(2p\times1)$ vector with $|g|_2 = 1$ and let $H =\{j=1,\dots,2p\mid g_j\ne0\}$ with cardinality $|H| = h<C.$ $H$ contains the indices of the coefficients involved in the hypothesis to be tested. The following theorem establishes the asymptotic normality of the debiased Lasso estimator.

thmSuppose that Assumptions (ref), (ref) and (ref) to (ref) hold, that $s_{r,max}^{3/2}log p/\sqrt{n}\rightarrow 0,$ and that the smallest eigenvalues of $\boldsymbol{\Omega}_{p,n},$ ${\widetilde{\boldsymbol{\Omega}}}_{p,n}$ and $\overline{\boldsymbol{\Omega}}_{p,n}$ are bounded away from 0. Furthermore, assume that $\lambda_{\max}^2\leq(\log\log n) \lambda_{\min}^{r}\left[\sqrt{n}s_{r,\max}\right]^{-1}$, and \begin{equation*} \begin{split} 0<r<1:&\quad\lambda_{\min}\geq {(\log \log n)}\left[s_{r,\max}\left(\frac{p^{\left(\frac{2}{d}+\frac{2}{m-1}\right)}}{\sqrt{n}}\right)^{\frac{1}{\left(\frac{1}{d}+\frac{m}{m-1}\right)}}\right]^{\frac{1}{r}}\\ r=0:&\quad s_{0,\max}\leq {(\log \log n)^{-1}}\left[\frac{\sqrt{n}}{p^{\left(\frac{2}{d}+\frac{2}{m-1}\right)}}\right]^{\frac{1}{\left(\frac{1}{d}+\frac{m}{m-1}\right)}},\quad \lambda_{\min}\geq {(\log \log n)}\frac{p^{1/m}}{\sqrt{n}}. \end{split} \end{equation*} Then, as $n \rightarrow \infty,$ we have \begin{equation*} \frac{\sqrt{n}g'(\widehat{a}(\widehat\tau)-\alpha_0)}{\sqrt{g'\boldsymbol{\Psi}\left(\widehat{\tau}\right)g}}\overset{d}{\to}N\left(0,1\right), \end{equation*} uniformly in $\alpha_0\in\mathcal{B}_{2p}(r, s_r)$, where $\boldsymbol{\Psi}\left(\tau\right):=$ \scalebox{0.70}{\parbox{0.1\linewidth}{\begin{equation*} \begin{aligned} \begin{bmatrix}\widetilde{\bm{Z}}(\tau)^{-2}\widetilde{\boldsymbol{\Omega}}_{p,n}\widetilde{\bm{Z}}(\tau)^{-2} & \widetilde{\bm{Z}}(\tau)^{-2}\overline{\boldsymbol{\Omega}}_{p,n}\bm{Z}(\tau)^{-2}-\widetilde{\bm{Z}}(\tau)^{-2}\widetilde{\boldsymbol{\Omega}}_{p,n}\widetilde{\bm{Z}}(\tau)^{-2}\\ \widetilde{\bm{Z}}(\tau)^{-2}\overline{\boldsymbol{\Omega}}_{p,n}\bm{Z}(\tau)^{-2}-\widetilde{\bm{Z}}(\tau)^{-2}\widetilde{\boldsymbol{\Omega}}_{p,n}\widetilde{\bm{Z}}(\tau)^{-2}& \bm{Z}(\tau)^{-2}\boldsymbol{\Omega}_{p,n}\bm{Z}(\tau)^{-2}+\widetilde{\bm{Z}}(\tau)^{-2}\widetilde{\boldsymbol{\Omega}}_{p,n}\widetilde{\bm{Z}}(\tau)^{-2}-2\widetilde{\bm{Z}}(\tau)^{-2}\overline{\boldsymbol{\Omega}}_{p,n}\bm{Z}(\tau)^{-2}\end{bmatrix}. \end{aligned} \end{equation*}}}
remWe establish the uniform asymptotic normality of the debiased Lasso estimator based on a finite number of parameters, a common consideration in time series inference. We can include an increasing number of tested parameters at the cost of assuming an $\alpha$-mixing process instead of the NED framework. \footnote{See Section 4.3 in ADAMEK20231114 for further discussion.}

To estimate the asymptotic variance, we consider the long-run variance kernel estimator, as in ADAMEK20231114, $\widehat{\boldsymbol{\Omega}}=\widehat{\boldsymbol{\Xi}}(0)+\sum\limits_{l=1}^{\widehat{k}_n-1} K\left(\frac{l}{\widehat{k}_n}\right) \left(\widehat{\boldsymbol{\Xi}}(l)+\widehat{\boldsymbol{\Xi}}^\prime(l)\right),$ $\widehat{\widetilde{\boldsymbol{\Omega}}}=\widehat{\widetilde{\boldsymbol{\Xi}}}(0)+\sum\limits_{l=1}^{\widetilde{k}_n-1} K\left(\frac{l}{\widetilde{k}_n}\right) \left(\widehat{\widetilde{\boldsymbol{\Xi}}}(l)+\widehat{\widetilde{\boldsymbol{\Xi}}}^\prime(l)\right),$ and $\widehat{\overline{\boldsymbol{\Omega}}}=\widehat{\overline{\boldsymbol{\Xi}}}(0)+\sum\limits_{l=1}^{\overline{k}_n-1} K\left(\frac{l}{\overline{k}_n}\right) \left(\widehat{\overline{\boldsymbol{\Xi}}}(l)+\widehat{\overline{\boldsymbol{\Xi}}}^\prime(l)\right),$ where $\widehat{\boldsymbol{\Xi}}(l)=\frac{1}{n-l}\sum\limits_{i=l+1}^{n}\widehat{\boldsymbol{w}}_i \widehat{\boldsymbol{w}}_{i-l}^\prime$ with $\widehat{w}_{i}^{(j)}=\widehat{v}_{i}^{(j)}\widehat{u}_i,$ $\widehat{\widetilde{\boldsymbol{\Xi}}}(l)=\frac{1}{n-l}\sum\limits_{i=l+1}^{n}\widehat{\widetilde{\boldsymbol{w}}}_i \widehat{\widetilde{\boldsymbol{w}}}_{i-l}^\prime$ with $\widehat{\widetilde{w}}_{i}^{(j)}=\widehat{\widetilde{v}}_{i}^{(j)}\widehat{u}_i,$ and $\widehat{\overline{\boldsymbol{\Xi}}}(l)=\frac{1}{n-l}\sum\limits_{i=l+1}^{n}\widehat{\boldsymbol{w}}_i \widehat{\widetilde{\boldsymbol{w}}}_{i-l}^\prime,$ the kernel $K(\cdot)$ can be taken as the Bartlett kernel $K(l/k_n) = \left(1-\frac{l}{k_n}\right)$ (NeweyWest87) and the bandwidths $k_n,$ $\widetilde{k}_n$ and $\overline{k}_n$ should increase with the sample size at an appropriate rate. Define $k_n = \max\left\{\widehat{k}_n, \widetilde{k}_n, \overline{k}_n\right\}.$

thmTake $\widehat{\boldsymbol{\Omega}},$ $\widehat{\widetilde{\boldsymbol{\Omega}}}$ and $\widehat{\overline{\boldsymbol{\Omega}}}$ with $k_n\to\infty$ as $n\to\infty$, such that $k_nh^2(\sqrt{n}h^2)^{-\frac{1}{1/d+m/(m-2)}} \\ \rightarrow0$. Suppose that \begin{equation*}\begin{split} &\lambda_{\max}^{2-r}\leq (\log \log n)^{-1} \min\left\lbrace\left[\sqrt{k_n}\sqrt{n}s_{r,\max}\right]^{-1}\right.,\left[k_nh^{1/m}n^{1/m}s_{r,\max}\right]^{-1},\\ &\qquad\qquad\qquad\qquad\quad\quad\left.\left[k_n^2h^{3/m}n^{(3-m)/m}s_{r,\max}\right]^{-1},\left[k_n^{2/3}h^{1/(3m)}n^{(m+1)/3m}s_{r,\max}\right]^{-1}\right\rbrace,\\ &\lambda_{\max}^2\leq (\log \log n)^{-1} \lambda_{\min}^{r}\left[\sqrt{n}h^{2/m}s_{r,\max}\right]^{-1}, and \\ \end{split}\end{equation*} \begin{equation*} \begin{split} 0<r<1:&\quad\lambda_{\min}\geq (\log \log n)\left[s_{r,\max}\left(\frac{(hp)^{\left(\frac{2}{d}+\frac{2}{m-1}\right)}}{\sqrt{n}}\right)^{\frac{1}{\left(\frac{1}{d}+\frac{m}{m-1}\right)}}\right]^{\frac{1}{r}},\\ r=0:&\quad s_{0,\max}\leq (\log \log n)^{-1}\left[\frac{\sqrt{n}}{(hp)^{\left(\frac{2}{d}+\frac{2}{m-1}\right)}}\right]^{\frac{1}{\left(\frac{1}{d}+\frac{m}{m-1}\right)}},\quad \lambda_{\min}\geq (\log \log n)\frac{(hp)^{1/m}}{\sqrt{n}}, \end{split} \end{equation*} and that Assumptions (ref), (ref) and (ref) to (ref) hold, then, we have \begin{equation} \sup_{\alpha^0\in\mathcal{A}^{({1})}_{2p}(r,s_r)}\left|g'\widehat{\boldsymbol{\Psi}}\left(\widehat{\tau}\right)g - g'\boldsymbol{\Psi}\left(\widehat{\tau}\right)g\right|_1=o_p(1), \end{equation} \begin{equation} \sup_{\alpha^0\in\mathcal{A}^{({2})}_{2p}(r,s_r)}\left| g'\widehat{\boldsymbol{\Psi}}\left(\widehat{\tau}\right)g - g'\boldsymbol{\Psi}\left(\tau_0\right)g\right|_1=o_p(1), \end{equation} where $\widehat{\boldsymbol{\Psi}}\left(\widehat{\tau}\right)=$ \scalebox{0.70}{\parbox{0.1\linewidth}{\begin{equation*} \begin{aligned}\begin{bmatrix}\widehat{\widetilde{\bm{Z}}}(\widehat{\tau})^{-2}\widehat{\widetilde{\boldsymbol{\Omega}}}\widehat{\widetilde{\bm{Z}}}(\widehat{\tau})^{-2} & \widehat{\widetilde{\bm{Z}}}(\tau)^{-2}\widehat{\overline{\boldsymbol{\Omega}}}\widehat{\bm{Z}}(\tau)^{-2}-\widehat{\widetilde{\bm{Z}}}(\widehat{\tau})^{-2}\widehat{\widetilde{\boldsymbol{\Omega}}}\widehat{\widetilde{\bm{Z}}}(\widehat{\tau})^{-2}\\ \widehat{\widetilde{\bm{Z}}}(\tau)^{-2}\widehat{\overline{\boldsymbol{\Omega}}}\widehat{\bm{Z}}(\tau)^{-2}-\widehat{\widetilde{\bm{Z}}}(\widehat{\tau})^{-2}\widehat{\widetilde{\boldsymbol{\Omega}}}\widehat{\widetilde{\bm{Z}}}(\widehat{\tau})^{-2}& \widehat{\bm{Z}}(\widehat{\tau})^{-2}\widehat{\boldsymbol{\Omega}}\widehat{\bm{Z}}(\widehat{\tau})^{-2}+\widehat{\widetilde{\bm{Z}}}(\widehat{\tau})^{-2}\widehat{\widetilde{\boldsymbol{\Omega}}}\widehat{\widetilde{\bm{Z}}}(\widehat{\tau})^{-2}-2\widehat{\widetilde{\bm{Z}}}(\tau)^{-2}\widehat{\overline{\boldsymbol{\Omega}}}\widehat{\bm{Z}}(\tau)^{-2}\end{bmatrix}. \end{aligned} \end{equation*}}}

We provide a uniformly consistent covariance matrix estimator in the cases with a fixed threshold effect and without the threshold effect in Theorem (ref). Similar to (ref) and (ref), there is a slight difference between the limits of the two asymptotic variances, since there is a true value for the threshold parameter in the case of a fixed threshold effect.

rem10.1093/jjfinec/nbac023 provides faster convergence rates of the heteroskedasticity and autocorrelation consistent estimator.
thmSuppose that Assumptions (ref), (ref) and (ref) to (ref) hold, that$s_{r,max}^{3/2}log p/\sqrt{n}\rightarrow 0,$ that the smallest eigenvalues of $\boldsymbol{\Omega}_{p,n},$ $\overline{\boldsymbol{\Omega}}_{p,n}$ and that $\widetilde{\boldsymbol{\Omega}}_{p,n}$ are bounded away from 0, and $k_nn^{-\frac{1}{2/d+2m/(m-2)}}\rightarrow0$ for some $k_n\to\infty$. Further, assume that $\lambda\sim\lambda_{\max}\sim\lambda_{\min}$, and that \begin{equation*} \begin{split} 0<r<1:&\quad (\log \log n)^{-1} s_{r,\max}^{1/r}\left[\frac{p^{\left(\frac{2}{d}+\frac{2}{m-1}\right)}}{\sqrt{n}}\right]^{\frac{1}{r\left(\frac{1}{d}+\frac{m}{m-1}\right)}}\leq\lambda\leq\ \log \log n \left[k_n^2\sqrt{n}s_{r,\max}\right]^{-1/(2-r)},\\ r=0:&\quad\ (\log \log n)^{-1} \frac{p^{1/m}}{\sqrt{n}}\leq\lambda\leq \log \log n \left[k_n^2\sqrt{n}s_{0,\max}\right]^{-1/2}. \end{split} \end{equation*} Assume that $k_n^rs_{r,\max}p^{\left(2-r\right)\left(\frac{d+m-1}{dm+m-1}\right)}n^{\frac{1}{4}\left(r-\frac{d(m-1)(2-r)}{dm+m-1}\right)}\rightarrow0$, and that $k_n^2s_{0,\max}\frac{p^{2/m}}{\sqrt{n}}\rightarrow0$ if $r=0$. Then, we have \begin{equation*} \begin{aligned} &\sup_{t\in\mathbb{R}} \sup_{\alpha_0 \in \mathcal{B}_{2p}(r,s_r)} \left|P\left(\frac{\sqrt{n}g'(\widehat{a}(\widehat\tau)-\alpha_0)}{\sqrt{g'\widehat{\boldsymbol{\Psi}}\left(\widehat{\tau}\right)g}}\leq t\right)-\Phi(t)\right| \rightarrow 0. \end{aligned} \end{equation*}

Similar to the result in Theorem (ref), we show that the convergence of a linear combination of the parameters of the debiased estimator $\widehat{a}(\widehat{\tau})$ to the standard normal distribution is uniformly valid over the $\ell_r$-ball. This allows researchers to perform uniform inference without specifying whether the specification is a linear or threshold regression.

Local Projection Inference

In this section, we develop the uniform inference theory for the debiased impulse response parameters in the high-dimensional local projection threshold (HDLPT) model. We focus on the following local projection threshold regression:

\scalebox{0.78}{\parbox{0.1\linewidth}{

align[align omitted — 727 chars of source]

}}

where $h = 0, 1, \ldots, h_{\max},$ \footnote{We assume that the unknown threshold parameter ($\tau_0$) is the same when estimating the impulse response function across different horizons and that the number of horizons is finite.} $\beta_h =(\beta_{h,0}, \phi_{h}, \rho_{h}, \boldsymbol{\eta}_h', \boldsymbol{\Delta}_{h,k}')'$ represents the projection parameters when the threshold variable is above the threshold point $\tau_0$, while $\beta_{h,0} + \delta_{h,0}, \phi_{h}+\delta_{h,x,0}, \rho_{h}+\delta_{h,y,0}, \boldsymbol{\eta}_h+\delta_{h,\boldsymbol{\eta},0}, \boldsymbol{\Delta}_{h,k}+\delta_{h,k,\boldsymbol{\Delta},0}$ are the projection parameters when the threshold variable is below the threshold point. $U_{h,i}$ is the projection error and $\bm{z}_i=\left(\bm{w}_{s,i}^\prime, Y_i, x_i, \bm{w}_{f,i}^\prime\right)^\prime$ includes the response $Y_i$, the shock variable $x_i,$ and the vectors of control variables consisting of “slow" variables $\bm{w}_{s,i}\in\mathbb{R}^{n_s},$ and the “fast" variables $\bm{w}_{f,i}\in\mathbb{R}^{n_f}$ for identification purposes. We are interested in $\phi_{h}$ and $\phi_{h}+\delta_{h,x,0}$, either of which represents the response at horizon $h$ of $y_i$ after an impulse in $x_i.$ When $\delta_0 = (\delta_{h,x,0},\delta_{h,y,0},\delta_{h,\boldsymbol{\eta},0}^\prime,\delta_{h,k,\boldsymbol{\Delta},0}^\prime)^\prime=0,$ it reduces to the local projection regression, similar to equation (1) in hdlp. We focus on a small number of parameters, allowing us to rewrite equation ((ref)) as

equation[equation omitted — 289 chars of source]

and furthermore,

equation*[equation* omitted — 133 chars of source]

We now have two groups of parameters. The parameters of interest are $\alpha_{\mathcal{H},0}=(\beta_{H,0},\delta_{H,0}),$ which belong to the first group while the second group includes the parameters for control variables, where $\mathcal{H}$ is the index set representing the $2H$ variables of interest. Without loss of generality, we order the variables in $\bm{X}(\tau) = (\bm{X}_\mathcal{H}(\tau),\bm{X}_\mathcal{-H}(\tau)).$ We then apply the penalization method, as in hdlp, penalizing only the parameters for control variables $\alpha_{-\mathcal{H},0}.$ Thus, shrinkage bias does not affect the unpenalized parameters of interest. We suppress the dependence on $h$ since each horizon of the LPs is estimated separately. The Lasso estimator is given as follows:

equation[equation omitted — 364 chars of source]

where $\bm{D}$ is an $2p\times 2p$ diagonal matrix with $\bm{D}_{i,i}=0$ for $i\in \mathcal{H}$ and $\bm{D}_{i,i}=1$. \footnote{As in Section 2.1 of hdlp, the Lasso estimator can be derived by

equation*[equation* omitted — 571 chars of source]

where $\bm{M}_{(\bm{X}_{\mathcal{H}}(\tau))}:=I-\bm{X}_\mathcal{H}(\tau)(\bm{X}_\mathcal{H}(\tau)^\prime\bm{X}_\mathcal{H}(\tau))^{-1}\bm{X}_\mathcal{H}(\tau)^\prime,$ and $\widehat{\bm{\Sigma}}_\mathcal{H}:=\bm{X}_\mathcal{H}(\tau)^\prime\bm{X}_\mathcal{H}(\tau)/n$.} The oracle inequalities are qualitatively the same as those in Theorem (ref).

Next, we consider the debiased Lasso estimator

equation[equation omitted — 239 chars of source]

where $\widehat{\bm{\Theta}}(\widehat{\tau})$ is an $2H\times 2p$ submatrix of an approximate inverse of $\widehat{\bm{\Sigma}}(\widehat{\tau}):=\bm{X}(\widehat{\tau})^\prime\bm{X}(\widehat{\tau})/n$. We still use nodewise regression and follow the same process as in Section (ref) to obtain $\widehat{\bm{\Theta}}(\widehat\tau).$ We construct

equation*[equation* omitted — 549 chars of source]

and $\widehat{\boldsymbol{Z}}_{H}(\tau)^2:=\text{diag}\left(\widehat{z}_1(\tau)^2,\dots,\widehat{z}_H(\tau)^{2}\right)$, where $\widehat{z}_j(\tau)^2:= ||X^{(j)}(\tau)-X^{(-j)} (\tau)\widehat{\gamma}_j(\tau)||_n^2 + 2\lambda_j |\widehat{\gamma}_j(\tau)|_1,$ we thus obtain $\widehat {A}_{H}(\tau) = \widehat{\boldsymbol{Z}}_{H}(\tau)^{-2}\widehat{\boldsymbol{C}}_{H}(\tau).$ Similarly, we have $\widehat {B}_{H}(\tau) = \widehat{\widetilde{\boldsymbol{Z}}}_{H}(\tau)^{-2}\widehat{\widetilde{\boldsymbol{C}}}_{H}(\tau).$ Define $\boldsymbol{\Omega}_{H,p,n},$ $\overline{\boldsymbol{\Omega}}_{H,p,n},$ and $\widetilde{\boldsymbol{\Omega}}_{H,p,n}$ as the top-left $H\times H$ submatrices of $\boldsymbol{\Omega}_{p,n},$ $\overline{\boldsymbol{\Omega}}_{p,n}$ and $\widetilde{\boldsymbol{\Omega}}_{p,n},$ respectively.

thmSuppose that Assumptions (ref), (ref) and (ref) to (ref) hold, that $H \leq C,$ that $s_{r,max}^{3/2}log p/\sqrt{n}\rightarrow 0,$ that the smallest eigenvalues of $\boldsymbol{\Omega}_{H,p,n},$ $\overline{\boldsymbol{\Omega}}_{H,p,n},$ and that $\widetilde{\boldsymbol{\Omega}}_{H,p,n}$ are bounded away from 0, and $k_nn^{-\frac{1}{2/d+2m/(m-2)}}\rightarrow0$ for some $k_n\to\infty$. Further, assume that $\lambda\sim\lambda_{\max}\sim\lambda_{\min}$, and that \begin{equation*} \begin{split} 0<r<1:&\quad (\log \log n)^{-1} s_{r,\max}^{1/r}\left[\frac{p^{\left(\frac{2}{d}+\frac{2}{m-1}\right)}}{\sqrt{n}}\right]^{\frac{1}{r\left(\frac{1}{d}+\frac{m}{m-1}\right)}}\leq\lambda\leq\ \log \log n \left[k_n^2\sqrt{n}s_{r,\max}\right]^{-1/(2-r)},\\ r=0:&\quad\ (\log \log n)^{-1} \frac{p^{1/m}}{\sqrt{n}}\leq\lambda\leq \log \log n \left[k_n^2\sqrt{n}s_{0,\max}\right]^{-1/2}. \end{split} \end{equation*} Assume that $k_n^rs_{r,\max}p^{\left(2-r\right)\left(\frac{d+m-1}{dm+m-1}\right)}n^{\frac{1}{4}\left(r-\frac{d(m-1)(2-r)}{dm+m-1}\right)}\rightarrow0$, and that $k_n^2s_{0,\max}\frac{p^{2/m}}{\sqrt{n}}\rightarrow0$ if $r=0$. Then, for $g \in \mathbb{R}^{\mathcal{H}},$ we have \begin{equation*} \begin{aligned} \sup_{t\in\mathbb{R}} \sup_{\alpha_0 \in \boldsymbol{B}_{2p}(r,s_r)} \left|P\left(\frac{\sqrt{n}g'(\widehat{a}_{\mathcal{H}}(\widehat\tau)-\alpha_{\mathcal{H},0})}{\sqrt{g'\widehat{\boldsymbol{\Psi}}_{\mathcal{H}}(\widehat{\tau})g}}\leq t\right)-\Phi(t)\right|=o_p(1), \end{aligned} \end{equation*} where $\widehat{\boldsymbol{\Psi}}_\mathcal{H}\left(\widehat{\tau}\right)=$ \scalebox{0.65}{\parbox{0.1\linewidth}{\begin{equation*} \begin{aligned}\begin{bmatrix}\widehat{\widetilde{\bm{Z}}}_\mathcal{H}(\widehat{\tau})^{-2}\widehat{\widetilde{\boldsymbol{\Omega}}}_\mathcal{H}\widehat{\widetilde{\bm{Z}}}_\mathcal{H}(\widehat{\tau})^{-2} & \widehat{\widetilde{\bm{Z}}}_\mathcal{H}(\tau)^{-2}\widehat{\overline{\boldsymbol{\Omega}}}_\mathcal{H}\widehat{\bm{Z}}_\mathcal{H}(\tau)^{-2}-\widehat{\widetilde{\bm{Z}}}_\mathcal{H}(\widehat{\tau})^{-2}\widehat{\widetilde{\boldsymbol{\Omega}}}_\mathcal{H}\widehat{\widetilde{\bm{Z}}}_\mathcal{H}(\widehat{\tau})^{-2}\\ \widehat{\widetilde{\bm{Z}}}_\mathcal{H}(\tau)^{-2}\widehat{\overline{\boldsymbol{\Omega}}}_\mathcal{H}\widehat{\bm{Z}}_\mathcal{H}(\tau)^{-2}-\widehat{\widetilde{\bm{Z}}}_\mathcal{H}(\widehat{\tau})^{-2}\widehat{\widetilde{\boldsymbol{\Omega}}}_\mathcal{H}\widehat{\widetilde{\bm{Z}}}_\mathcal{H}(\widehat{\tau})^{-2}& \widehat{\bm{Z}}_\mathcal{H}(\widehat{\tau})^{-2}\widehat{\boldsymbol{\Omega}}_\mathcal{H}\widehat{\bm{Z}}_\mathcal{H}(\widehat{\tau})^{-2}+\widehat{\widetilde{\bm{Z}}}_\mathcal{H}(\widehat{\tau})^{-2}\widehat{\widetilde{\boldsymbol{\Omega}}}_\mathcal{H}\widehat{\widetilde{\bm{Z}}}_\mathcal{H}(\widehat{\tau})^{-2}-2\widehat{\widetilde{\bm{Z}}}_\mathcal{H}(\tau)^{-2}\widehat{\overline{\boldsymbol{\Omega}}}_\mathcal{H}\widehat{\bm{Z}}_\mathcal{H}(\tau)^{-2}\end{bmatrix}. \end{aligned} \end{equation*}}}

We also use an autocorrelation robust Newey-West long-run covariance estimator, as in Section (ref).

The uniformly consistent covariance results for the cases with no threshold effect and a fixed threshold effect are similar to those in Theorem (ref). The proof of Theorem (ref) follows from the proofs of Theorem (ref), and Theorem 1 in hdlp. Therefore, we omit the detailed proof.

Monte Carlo Simulation and Applications

We study the finite sample properties of the proposed debiased Lasso estimator for high-dimensional threshold regression through Monte Carlo experiments. We compare our debiased threshold Lasso (DTLasso) estimators for threshold regression with those from linear regression in cases with no threshold effect and a fixed threshold effect. Additionally, we apply our method to two empirical applications, one related to the multiple steady states of economic growth by durlauf1995multiple, and the other concerning the effect of a military spending news shock on government spending and GDP by ramey2018government.

Monte Carlo Simulation

We first describe the data-generating process. We consider the threshold regression model ((ref)), where the rows of the design matrix are i.i.d. realizations of $N(0,\bm{\Sigma}),$ with $\Sigma_{j,k} = 0.5^{|j-k|},$ a Toeplitz structure, and the error terms are $t$ distributed with 10 degrees of freedom. When the threshold variable $Q_i$ is independent of $X_i$, we take $Q_i \sim \text{uniform}(0,1)$. We also consider the case where the threshold variable correlates with the covariates. We take $\tau_0=0.5$ unless otherwise specified. We use the grid search method to find $\tau$ from 0.15 to 0.85 by steps of 0.01. Without loss of generality, we assume that $\beta_{0}$ is a $p\times 1$ vector with the first $s_0$ elements being $b$, the remaining $p-s_0$ elements being zeros and that $\delta_0$ is a $p\times 1$ vector with the first $s_0$ elements being 0, the next $s_0$ elements being $b1$, and the remaining elements being zeros.

To choose the tuning parameters $\lambda$, we use the generalized information criterion (GIC) proposed by GIC. We utilize GIC and ten-fold cross-validation to select the tuning parameters $\lambda$. However, according to our simulation results, cross-validation does not significantly enhance the quality of the results, while the processing time is considerably longer. Hence, we select both $\lambda$ and $\lambda_{node}$ based on GIC.

In Figures (ref) and (ref), we plot the constructed 95% confidence intervals for the realizations $(n,2p,s_0,b, b1, \rho_{Q,X^{(2)}})=(400,600,15,2,1,0.5)$ and $(n,2p,s_0,b, b1, \rho_{Q,X^{(2)}})=(400,600,15,2,0,0.5)$ for both threshold and linear regression models. DTLasso estimator for threshold regression performs much better than the debiased estimator for linear regression when there is a fixed threshold effect. Even when the threshold effect does not exist, our estimator for the threshold regression still performs comparably to that for the linear regression.

figure[figure omitted — 646 chars of source]
figure[figure omitted — 649 chars of source]

Additionally, we consider 100 independent realizations for each parameter $\alpha_{0,i}$ for each model specification. We focus only on the parameters $\beta_{0,i}$ as we study two separate cases: one with a fixed threshold effect and the other without a threshold effect ($\delta_{0,i}=0$). We compute the average length of the corresponding confidence interval $\rm{Avglength} \left(J_i(\beta)\right)$, and the average

equation[equation omitted — 83 chars of source]

We also compute the average length of intervals for the active and inactive parameters,

equation[equation omitted — 163 chars of source]

and the average coverage for individual parameters,

equation[equation omitted — 347 chars of source]

where $\widehat {\mathbb{P}}$ denotes the empirical probability computed based on $100$ realizations for each configuration. The results are reported in Table (ref). The debiased estimator performs well in terms of the presence of a fixed threshold effect, the correlation between the threshold variable and covariates, and varying magnitudes of effects, different levels of sparsity, the location of the threshold point, as well as different sample sizes and numbers of covariates.

table*[table* omitted — 1,282 chars of source]

Additionally, we consider a test for the family of hypotheses $\left\{H_{0}^{(j)}: \, \alpha_{0}^{(j)} = 0\right\}$ for $j=1,\dots,2p.$ We report the familywise error rate (FWER) based on the Bonferroni-Holm procedure and the empirical power,

equation*[equation* omitted — 101 chars of source]

We compare our results from threshold models with those from linear regression in Table (ref). The FWER based on the threshold model is close to the preassigned significance level of 0.05 and is robust to the magnitude of the threshold effect and the level of sparsity. DTLasso estimator for the threshold model has more power than that for the linear model, even when the threshold effect is small.

table*[table* omitted — 613 chars of source]

Economic Growth Rate

durlauf1995multiple provide a theoretical background for the existence of multiple steady states in economic growth models. They also consider a broad set of control variables to check the robustness, but lee2016 argue that this approach still restricts variable selection. Therefore, they apply the Lasso method to simultaneously select covariates and choose between linear and threshold models with high-dimensional data. To further identify the relevant covariates, we continue applying a threshold regression model to study countries' economic growth and analyze the significance of covariates by the debiased Lasso estimator. Our setup follows Equation (3.1) in lee2016,

equation[equation omitted — 228 chars of source]

where $\mathit{gr}_{i}$ is the annualized GDP growth rate for each country $i$ during the period 1960-1985, $\mathit{lgdp60}_{i}$ represents the log GDP in 1960, and $ X_{i}$ is a vector of additional covariates, including education, demographic characteristics, market openness, politics, and interaction terms. Table (ref) from lee2016 provides a detailed introduction to the covariates. We use either the initial GDP or the adult literacy rate in 1960 as the threshold variable $Q_{i}$ following durlauf1995multiple and lee2016. The grid interval ranges from the 10th to the 90th percentiles of the threshold variable. We utilize covariates from lee2016 and the dataset originating from barro1994data and durlauf1995multiple. With initial GDP as the threshold variable, we have 80 countries with 46 covariates (including a constant term). With literacy rate as the threshold variable, we have 70 countries with 47 covariates. The traditional OLS and MLE are no longer valid because of $2p > n.$

table*[table* omitted — 3,488 chars of source]

Table (ref) shows the results with initial GDP as the threshold variable. We provide additional findings based on the adult literacy rate as the threshold variable are presented in Table (ref) in Appendix (ref). Firstly, we confirm the presence of a threshold effect, as some values of $\delta_2$ are significantly different from 0. This finding provides evidence supporting the existence of multiple steady states in growth models and implies that the rate of growth convergence may differ across regimes defined by varying levels of initial GDP. Secondly, the significance of the covariates varies across regimes. For example, a higher proportion of secondary school completion accelerates a developing country's economic growth, whereas participating in an external war exerts a significantly negative effect on growth performance. Finally, we compare our findings with those reported in lee2016 and observe that our model identifies a larger set of significant variables, such as the average years of primary schooling among the female population in 1960 and the percentage of primary schooling completed in the female population in 1960. Furthermore, for covariates with significant debiased estimates, we report their standard errors to support valid post-selection inference.

Government Spending and GDP

ramey2018government provide theoretical and empirical background on the effect of a military news shock on government spending and GDP and use state-dependent LP to estimate impulse responses. Later, hdlp reestimate the impulse responses through a state-dependenct HDLP specification while including more lags for a robustness check. However, the threshold point defining the state of the economy in both literatures is based on a predetermined standard. For example, defining the state as slack when the unemployment rate exceeds 6.5 percent (the US Federal Reserve's standard). A state is defined as a zero lower bound (ZLB) when the 3-month Treasury bill rates are below 0.5 percent. We thus apply a high-dimensional local projection threshold model to find a theoretically more accurate threshold point to define the state and estimate the impulse responses accordingly. The model is as follows,

equation[equation omitted — 283 chars of source]

where $Y_i$ includes real per capita GDP and government spending, $x_i$ is the military spending news shock, $\boldsymbol{z}_i$ includes lags of the news, GDP, government spending, and tax. We use a quarterly dataset from 1889Q1 to 2015Q4.\footnote{The dataset is available at \url{https://econweb.ucsd.edu/ vramey/research.html\#govt}.} Section II.B in ramey2018government provides a more detailed description of the data and variables. The time series length is 161, spanning a long U.S. history. It includes many prolonged periods of slack state and extended periods of near-zero bound state, allowing us to estimate the impulse responses over reasonable horizons with non-changing states. We use one lag of the 3-month Treasury bill rate or one lag of the unemployment rate as the threshold variable. Section IV. A. and Section V.A in ramey2018government present the narrative reasons for the choice of the variable to define the state of the economy.

figure*[figure* omitted — 320 chars of source]

When one lag of the 3-month Treasury bill rate is the threshold variable, we take $K = 40$, and define the range of the threshold parameter as the interval spanning from the 10th to the 90th sample quantile for the threshold variable. We then define the state as ZLB when the rate is below 1.02$\%$ (corresponding to the threshold derived when $h=0$) and the normal state when it exceeds 1.02$\%$. Figure (ref) shows the impulse responses to a military spending news shock in government spending and GDP. The results indicate that both government spending and GDP exhibit significantly stronger responses during ZLB compared to normal states. Moreover, at its peak, the response of government spending to a military spending news shock exceeds the corresponding response of GDP. Compared to Figure 11 in ramey2018government, the peaks of both government spending and GDP occur two quarters earlier, and the magnitudes of these peaks are also larger. However, in the normal state, our responses are much more subtle.

When one lag of the unemployment rate is the threshold variable, we take $K = 40,$ and define the range of the threshold parameter as the interval spanning from the 10th to the 90th sample quantile for the threshold variable. We define the state as the low unemployment state when the unemployment rate is below 4.58$\%$ (corresponding to the threshold derived when $h=0$) and the high unemployment state (slack state) when the unemployment rate exceeds 4.58$\%$. Figure (ref) shows the impulse responses to a military spending news shock in government spending and GDP. The results indicate that both government spending and GDP exhibit significantly larger responses during high-unemployment states compared to low-unemployment states. Moreover, the response of government spending to a military spending news shock reaches its peak earlier, is more persistent, and exhibits a slightly larger peak magnitude than the corresponding response of GDP. Compared to Figure 5 in ramey2018government and Figure 4 in hdlp, in the low unemployment state, the patterns are very similar for both government spending and GDP. In the slack state, for government spending, the peak of the responses occurs earlier but lasts longer, and the peak magnitude is smaller. For GDP, compared to Figure 5 in ramey2018government, the peak occurs 5 periods earlier, and the peak magnitude is smaller. While comparing to Figure 4 in hdlp, the peak of the responses occurs 2 periods earlier, and the peak magnitude is smaller in the slack state. In addition, the impulse responses are slight at horizon 0 in the linear, high, and low unemployment states.

Our results are more robust as we include more lags and are less sensitive to the number of lags due to the Lasso estimator, which addresses variable selection. In contrast, ramey2018government only include four lags and do not perform robustness checks for the number of lags. In contrast, ramey2018government include only four lags and do not perform robustness checks on the number of lags, and their specification does not involve any variable selection procedure.

figure*[figure* omitted — 313 chars of source]

Conclusion

This paper proposes a debiased Lasso estimator for high-dimensional slope parameters in threshold regression models, allowing for either cross-sectional or time series data. We derive the asymptotic distribution of tests involving an increasing number of slope parameters and construct uniformly valid confidence bands. We show that the asymptotic distributions are the same in the cases with no threshold effect and a fixed threshold effect. Our study allows for less restrictive assumptions than existing research in high-dimensional threshold models, accommodating heteroskedastic non-subgaussian error terms and non-subgaussian covariates. Future research directions could include considering multiple threshold points or multiple threshold variables in the current framework and developing a uniform inference theory in panel data models with threshold effects.