Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
67,907 characters · 16 sections · 64 citation commands
Inference on many jumps in nonparametric panel regression models
\def\spacingset#1{ {#1}} \spacingset{1}
\if11 \fi
\if01 {
} \fi
{\it Keywords:} high-dimensional time series; temporal and cross-sectional dependence; changepoint analysis; Gaussian approximation; simultaneous tests; threshold regression.
\spacingset{1.9}
Recently, there has been a notable surge in the study and application of changepoint analysis in high-dimensional time series data (see, e.g., cho2015multiple, cao2018multi, chen2022inference, and li2024, along with the references cited therein). Detecting changepoints in regression models is important and has wide applications in the field of biology pastor1998use, agricultural sciences freeman1998credit, econometrics li2024, etc. The purpose of these studies is to estimate not only where the threshold effect exists but also the magnitude of the effects.
{ There is extensive literature on parametric high-dimensional changepoint regression analysis. li2016panel, manner2019testing wang2021statistically, rinaldo2021localizing, and cho2024detection investigate changepoint analysis of coefficients in high-dimensional linear regression models using penalization methods. lee2016lasso examine a high-dimensional regression model with a changepoint due to a covariate threshold in a cross-sectional study. These studies majorly focus on parametric models with high-dimensional covariates. Similar to lee2016lasso, this paper also considers changepoint regression due to a covariate threshold. However, our focus is on nonparametric regression with a high-dimensional changepoint effect driven by heterogeneity in the panel data structure. Specifically, we {investigate the heterogeneous nonparametric panel data with jumps in the conditional mean functions: }
where {$X_{jt}$ is a one-dimensional threshold variable,} $Y_{jt}\in \mathbb{R}^1$ and {$U_{jt}\in \mathbb{R}^d$} are the outcome and covariate variables that can be observed or unobserved for individual $j\in \lbrack N]:=\{1,\ldots ,N\}$ at time point $t\in \lbrack T],$} $h_{j}(\cdot,\cdot )$, $\tau_j(\cdot,\cdot)$ and $\sigma _{j}(\cdot,\cdot)$ are continuous functions, $c_{0j}$ is a threshold value that is either known or unknown and $\epsilon_{jt}$ is the error term. Our main focus is to test whether changepoint effects related to $\tau _{j}(X_{jt},U_{jt})$ exist and to determine if the jump sizes are identical uniformly across individuals $j$ under general spatial temporal dependence assumptions.
As a motivating example for our model, we explore an application involving stock lagged returns and volatility using threshold regression. Threshold regression, dating back to chan1985multiple and more recently studied in works such as duffy2023stationarity, is mainly formulated within a linear parametric setting. Our approach naturally extends this framework to the non-parametric setting. Specifically, in our first application, revising the news impact curve of engle1993measuring, $Y_{jt}$ is the volatility of firm $j$'s stock at time $t$ and $X_{jt}$ is the lagged stock return. A trader would like to discover whether there are asymmetric effects of $X_{jt}$ on $Y_{jt}$, as well as whether and where a discontinuous relationship occurs, for trading and risk management purposes. If there are multiple stocks jump, are they all identical uniformly? In Section (ref) in our empirical section, we identify significant threshold effects, demonstrating that our methodology is crucial for predicting stock volatility and effectively managing risk. In another application, we study the effect of party incumbency as determined by the 50% threshold in the two-party vote share in the US House Elections. Our findings confirm the presence of a significant positive incumbency effect as studied by lee2008randomized, and we detect significant heterogeneity in the threshold effect across the US states.
The methodological contributions are twofold. First, we develop algorithms to test for overall threshold effects, examine the potential heterogeneity effects and derive the asymptotic properties of the resulting test statistics. Due to the use of local smoothing, neither the weak dependence along the time dimension nor the weak or strong dependence along the individual/group dimension enters the asymptotic distribution of the final statistic. Interestingly, we find strong cross-sectional dependence is allowed in our error terms. As a practical recommendation, we find that under fairly weak assumptions, if there are clustering dependency patterns in $e_{jt}$, the correlation in the estimated statistics can be safely ignored, and only adjustments for heteroskedasticity are needed. Second, the standard central limit theorem is challenging to apply to high-dimensional time series when $N$ significantly exceeds the effective sample size. Previously, chernozhukov2015comparison, chernozhukov_central_2017, zhang2017gaussian, chernozhukov2019inference, and others addressed this issue by introducing high-dimensional Gaussian approximation (GA) methods. However, the existing literature cannot address our problem. The key difficulty is to incorporate possibly non-identically distributed time series under general spatial and temporal dependency structures and more importantly latent variables. In this paper, we propose a new high dimensional GA for weighted statistics that address all these issues. These GA results are essential for determining the critical values of our supremum test.
Our paper is related to three strands of literature in statistics. First, we add to the extensive body of work on threshold regression for high dimensional data. For example, miao2020panel_a and miao2020panel consider the estimation of a panel threshold regression model with interactive fixed effects (IFEs) and latent group structures, respectively; massacci2022high study inference for high-dimensional threshold regressions with common stochastic trends. The aforementioned literature does not account for potential heterogeneous threshold effects. Ignoring these effects may bias estimators or lead to efficiency loss when estimating homogeneous threshold effects in finite samples. Second, our paper contributes to the literature on testing homogeneity in a panel data setting. For example, phillips2003dynamic, pesaran2008testing and su2013testing propose tests for slope coefficient heterogeneity in a linear regression model with or without cross-sectional dependence. In the presence of slope heterogeneity, su2016identifying and su2018identifying model the slope heterogeneity via latent group structures. barassi2023threshold study a heterogeneous panel threshold regression model with IFEs. Above literature all focus on parametric setting. Our work complements these research streams by adopting nonparametric approaches, which address size and power distortions caused by model misspecifications. Third, our nonparametric approach is related with literature in trends estimation. Inference on smooth trend functions in nonparametric models for time series data is a well-established field; see zhang2012inference, karmakar2022simultaneous, gao2024time, chen2018testing, among others. Distinctly, while the above literature examines nonlinear dynamic models, they does not address discontinuity in trends, particularly heterogeneous jumps. The nonparametric literature on jumps in trends mostly focuses on independent and identically distributed (i.i.d.) settings, without accounting for heterogeneity or high dimensionality; see, for example, qiu1998discontinuous, muller1999discontinuous, and spokoiny1998estimation.
Another distinct feature is that we explicitly handle the case with missing covariates or models with misspecified variables. Theoretically, validating such an inference procedure requires working with a reduced-form model that has specific variance-covariance structures. This aspect of theoretical validity has not been explored in the time-varying trend literature. Establishing the validity of uniform inference is a major theoretical contribution that sets our work apart in the study of varying trends.
The rest of the paper is organized as follows. Section (ref) provides the model setup and the practical steps of the testing procedure. Section (ref) presents the assumptions and delivers the main theorems. In Section (ref) we provide simulation results.\footnote{The code for our method is available in the R package hdthreshold (\url{https://cran.r-project.org/web/packages/hdthreshold/index.html}), with an illustration accessible at \url{https://rpubs.com/Lk1110/1259698}.} Real data applications are discussed in Section (ref), and the paper concludes in Section (ref). The Supplementary Material includes additional theorems, proofs of all results, extra simulation outcomes, and tables for the empirical applications.
Notation. For a vector $v=(v_{1},\ldots,v_{d})\in \mathbb{R}^{d}$ and $q>0$, we denote $|v|_{q}=( \sum_{i=1}^{d}|v_{i}|^{q})^{1/q}$ and $|v|_{\infty }=\max_{1\leq i\leq d}|v_{i}|$. For a matrix $A=(a_{i,j})_{1\leq i\leq m,1\leq j\leq n}$, we define the max norm $|A|_{\text{max}}=\max_{i,j}|a_{i,j}|$. For $s>0$ and a random vector $X$, we say $X\in \mathcal{L}^{s}$ if $\lVert X\rVert _{s}=[ \mathbb{E}(|X|^{s})]^{1/s}<\infty $. For two positive number sequences $ (a_{n})$ and $(b_{n})$, we say $a_{n}=O(b_{n})$ or $a_{n}\lesssim b_{n}$ (resp. $a_{n}\asymp b_{n}$) if there exists $C>0$ such that $a_{n}/b_{n}\leq C$ (resp. $1/C\leq a_{n}/b_{n}\leq C$) for all large $n$, and say $ a_{n}=o(b_{n})$ if $a_{n}/b_{n}\rightarrow 0$ as $n\rightarrow \infty $. We set $(X_{n})$ and $(Y_{n})$ to be two sequences of random variables. Write $ X_{n}=O_{\mathbb{P}}(Y_{n})$ if for any $\epsilon >0$, there exists $C>0$ such that $\mathbb{P}(X_{n}/Y_{n}\leq C)>1-\epsilon $ for all large $n$, and say $X_{n}=o_{\mathbb{P}}(Y_{n})$ if $X_{n}/Y_{n}\rightarrow 0$ in probability as $n\rightarrow \infty $.
{ In this section, we present the model, the hypotheses and estimators, and the test procedure. Subsections (ref) and (ref) are concerned with the known threshold case. In Subsection (ref) we focus on the case in which the threshold locations $c_{0j}$ are unknown. The algorithms for our testing procedures are listed in Subsection (ref). }
{ In this subsection, we formulate our model. Recall the full model in ((ref)) where $U_{jt} \in \mathbb{R}^{d}$ is a latent random vector that can be arbitrarily correlated with $X_{jt},$ $\epsilon _{jt}$ is a mean zero term that is independent of $\left( X_{jt},U_{jt}\right) $.
Without loss of generality set $c_{0j} = 0$ for the known case. The full model in (ref) can be rewritten into the following reduced form model:
where $\tilde{\tau}_{j}(X_{jt})=\mathbb{E}[\tau _{j}(X_{jt},U_{jt})|X_{jt}]$, $\tilde{h}_{j}(X_{jt})=\mathbb{E[}h_{j}(X_{jt},U_{jt})|X_{jt}]$, $\varepsilon _{j}(X_{jt},U_{jt})=h_{j}(X_{jt},U_{jt})-\tilde{h} _{j}(X_{jt})+[\tau _{j}(X_{jt},U_{jt})-\tilde{\tau}_{j}(X_{jt})]\mathbf{1} _{\{X_{jt}\geq c_{0j}\}}$ and the reduced form error
The noise term $e_{jt}$ has a complex structure, influenced by both the latent variable and by the heterogeneous volatility of the observations. This intricate structure, combined with the temporal and cross-sectional dependencies of the noise, makes the inference process particularly challenging.
A few remarks are in order. First, we have $\mathbb{E}(e_{jt}| X_{jt})=0$ and the conditional expectation of the observed outcome given $X_{jt}$ is $ \mathbb{E}(Y_{jt}|X_{jt})=\tilde{h}_{j}(X_{jt})+\tilde{\tau}_{j}(X_{jt}) \mathbf{1}_{\{X_{jt}\geq c_{0j}\}}.$ Second, in Section (ref) we show the possibilities of extending our model to add more covariates and incoporating fixed effects. Third, ($X_{jt},U_{jt}$) can be lagged variables. For example, in our application on modeling stock volatility (Section (ref)), $X_{jt} = Y_{j(t-1)}$, functions $h_{j}(\cdot ,\cdot ) $, {$\tau _{j}(\cdot ,\cdot )$} and $\sigma _{j}(\cdot ,\cdot )$ represent baseline effect, the jump effect and the volatility function, respectively. The $U_{jt}$ are unobserved risk factors or variables we exclude in the nonparametric regression. Modeling $U_{jt}$ is important in this circumstance as the volatility of a single stock is most likely affected by other risk factors such as news events or unobserved market factors.
This paper aims to explore methods for conducting simultaneous inference on threshold effect estimators in a large $N$ and large $T$ setup. In particular, we allow $N\gg T.$ Our focus is on the threshold effect, starting with the case where the threshold value $c_{0j}$ is known and then extending the analysis to the unknown case in Subsections (ref) and (ref).
{ Assuming that $X_{jt}$ is a continuous random variable, the jump effects can be identified as a \textquotedblleft gap" in the conditional expectations and their derivatives at the thresholds. Let $\mu _{j}(x):=\mathbb{E}(Y_{jt}|X_{jt}=x)$ and $\mu _{j}(c_{0j}-)=\lim_{x\uparrow c_{0j}}\mu _{j}(x)$ and $\mu _{j}(c_{0j}+)=\lim_{x\downarrow c_{0j}}\mu _{j}(x).$ Consider $\partial _{+}\mu _{j}(\cdot )$ (resp. $\partial _{-}\mu _{j}(\cdot )$) as the right (resp. left) derivative of function $\mu _{j}(\cdot ).$ The threshold effect for the individual $j$ is defined as
} Due to continuity of function $h_j(\cdot,\cdot)$ and $\tau_j(\cdot,\cdot)$, we have $\gamma _{j}=\tilde{\tau}_{j}(c_{0j}),\gamma _{j}^{[1]}=\tilde{\tau}'_{j}(c_{0j}).$ ($\tilde{\tau}'_{j}(.)$ denotes the derivative of $\tilde{\tau}_{j}(.)$.)
In this section, we introduce the hypotheses and estimators. We would like to conduct simultaneous inferences on $\gamma _{j}$ and $\gamma_j^{[1]}$ for $ j\in \left[ N\right] $. The importance of studying individual/group-specific threshold effects is discussed in zimmert2019nonparametric.
{ First, we are interested in testing the existence of the threshold effects ($\gamma_j$ or $\gamma_j^{[1]}$). The null and alternative hypotheses are defined as follows:
The motivation for considering such a test is crucial. In Figure (ref), we present a simple comparison with the pooled threshold test in the literature, for example in fong2017chngpt. It is clear that the conventional pooled tests exhibit very low power due to signal cancellation, with performance only comparable in cases of strictly positive signals.
}
When $N$ is large, assuming homogeneous threshold effects $\gamma_j$ across all $j$ is restrictive, and inferences based on this assumption can be misleading if the threshold effects are heterogeneous. Therefore, it is important to test for heterogeneous threshold effects. In this case, the null and alternative hypotheses are:
{We can similarly define $H_0^{(1)[1]}$ and $H_0^{(2)[1]}$ for the existence and homogeneity tests of the derivative, respectively. }Rejecting of $H_{0}^{\left( 2\right) }$ suggests the presence of heterogeneous threshold effects. Moreover, providing simultaneous confidence intervals for the threshold effects is also tempting. If we would like to study the asymptotics of $\hat{ \gamma}=(\hat{\gamma}_{1},\ldots,\hat{\gamma}_{N})^{\top }$, the standard central limit theorem may fail due to the high dimensionality $N$. One of the key theoretical contributions is to provide a general framework for making uniform inferences on heterogeneous threshold effects.
{
To estimate $\gamma _{j} $, we need to estimate both $\mu _{j}(c_{0j}+)$ and $\mu _{j}(c_{0j}-).$ To this aim, it is natural to adopt local linear estimators, {see for example fan2018local}. Denote the estimators:
where }$K_{jt}:=K((X_{jt}-c_{0j})/b_{j}),${ \ $b_{j}$ is a bandwidth parameter, and $K(\cdot )$ is a kernel function. For $l=0,1,2$, define the weights as
where
and $w_{j,b}^{-}(u,v)$ (resp. $S_{jl,b}^{-}(u)$, $w_{j,b}^{-[1]}(u,v)$) is the same as $ w_{j,b}^{+}(u,v)$ (resp. $S_{jl,b}^{+}(u)$, $w_{j,b}^{+[1]}(u,v)$) with $+$ and $\geq $ therein replaced by $-$ and $<,$ respectively. {\color{red}
}
We can thus evaluate the threshold effect near the cut-off point and define our local linear estimator of the average threshold effect for individual/group $j$ as
where the weights take the following form $ w_{jt,b}^{+}=w_{j,b}^{+}(c_{0j},X_{jt})$ and $ w_{jt,b}^{-}=w_{j,b}^{-}(c_{0j},X_{jt})$. } { Thus we have,
We shall focus on $\hat{\gamma}_{j}$ from now on, and the algorithms and theorems concerning the statistical properties of $\hat{\gamma}_{j}^{[1]}$ are developed in Section (ref) in the Appendix.
} Under $H_{0}^{\left( 1\right) }:$ $\gamma _{j}=0$ for all $j\in \left[ N\right] $, we are exclusively testing for the overall significance of threshold effects. Since the functions $\tilde{h}_{j}(\cdot)$ and $\tilde{\tau}_{j}(\cdot)$ are continuous, by (ref), the difference between our estimator and the parameter can be approximately expressed in terms of a weighted average of the error sequence,
where $e_{jt}$ is the composite error in ((ref) ). Under $H_{0}^{\left( 1\right) }$, one would expect $\left\vert \hat{ \gamma}_{j}\right\vert $ to be small for all $j\in \left[ N\right] $. {This motivates us to consider the test statistic
where
}is the conditional variance {of }$(Tb_{j})^{1/2}\hat{\gamma}_{j}$ for standardization purposes. For now, we assume $v_{j}^{2}$ is given, in general, it is unknown and has to be replaced by its estimate. We discuss a consistent estimator of $v_{j}^{2}$ in Subsection (ref). Under the assumptions in {Section (ref)}, it can be shown that all asymptotic results based on $v_{j}^{2}$ remain valid when it is replaced by the consistent estimate $\hat{v}_{j}^{2}$.
Similarly, under $H_{0}^{\left( 2\right) }$, one would expect $\left\vert \hat{\gamma}_{j}-\overline{\hat{\gamma}}\right\vert $ to be small for all $j\in \left[ N\right] ,$ where $\overline{\hat{\gamma}}=\frac{1 }{N}\sum_{j=1}^{N}\hat{\gamma}_{j}$. This motivates us to consider the test statistic
where $\tilde{v}_{j}^{2}=(1-1/N)^{2}v_{j}^{2}+\sum_{i\neq j}^{N}v_{i}^{2}/N^{2}$. For practical implementation, we introduce a feasible version of $\mathcal{Q}$ in Algorithm (ref) in Section (ref) of the Supplementary Material.
{ It shall be noted that in case the breakpoints $c_{0j}$'s are not known, we shall develop an algorithm to estimate them. We therefore generalize our methodology to estimate unknown breakpoints that vary over }$ {\normalsize j}$. To this end, we search over a grid $[c_1,c_{2},\ldots ,c_K]$ for each individual $j\in \lbrack N]$. Let
Further, let $v_{j}^{2}(c_{i})=(Tb_{j})\sum_{t=1}^{T}w_{jt,b}^{2}(c_{i})\mathrm{Var} \big(e_{jt}|X_{jt}\big)$ and define
We propose to estimate $c_{0j}$ by $\hat{c}_{j}:=\arg \max_{1\leq i\leq K}(Tb_{j})^{1/2}|\hat{\gamma}_{j}(c_{i})/\hat{v}_{j}(c_{i})|\ $for $j\in \left[ N\right] ,${ \ where }$\hat{v}_{j}(c_{i})${ \ is a consistent estimator of }$v_{j}(c_{i})$ which will be introduced in Algorithm (ref) in Section (ref) of the Supplementary Material.
In this subsection, we present the steps of the test procedures. To proceed, we begin by discussing the estimation of the error variance $\sigma _{e,j}^{2}$ of $e_{jt}$ when $c_{0j}$'s are known.
The estimator $\hat{\sigma}_{e,j}^2$ is used to form $\hat{v}_{j}^{2}$ in the Algorithms (ref) and (ref) in the case of known $c_{0j}$, which are applied to conduct simultaneous tests for $H_{0}^{\left( 1\right) }$ and $ H_{0}^{\left( 2\right) }$. These algorithms are supported by the theoretical GA results in Section (ref) and Section (ref) in Supplementary Material. The following Algorithm (ref) is for testing $H_0^{(1)}$.
Furthermore, Algorithm (ref) in Section (ref) in the Supplementary Material is used for testing the homogeneity hypothesis $H_0^{(2)}$. For the case where the threshold is unknown, the corresponding algorithm is detailed in Algorithm (ref) in the same Section. Note that $\mathcal{\hat{I}}$ in Algorithm (ref) (resp. $\mathcal{\hat{Q}}$ in Algorithm (ref), $\mathcal{\hat{I}}^{C}$ in Algorithm {(ref)}) is a feasible version of $\mathcal{I}$ (resp. $\mathcal{Q}$, $\mathcal{{I}}^{C}.$). In Section (ref), we will study the asymptotic properties of $\mathcal{\hat{I}}$, $\mathcal{\hat{Q}}$ and $ \mathcal{\hat{I}}^{C}.$
Given the well-documented cross-sectional and serial dependence in high dimensional data, it is crucial to develop tests that account for both types of dependence. { It is surprising that, according to the above algorithm, $q_{\alpha}$ can be obtained by simulating only i.i.d. Gaussian random variables, without considering the complicated dependency structure of $e_{jt}$. Recall from ((ref)), we see that the reduced-form error $e_{jt}$ inherits nontrivial dependency structure in $ \epsilon _{jt}$. {However, this dependence does not affect the calculation of critical values, even in the case of fairly strong cross-sectional dependence (e.g., a factor structure). This is mainly due to the localized nature of our kernel-type estimators $\hat{\gamma }_{j}$'s in our test statistic. In fact, we can show that under some mild conditions, as long as the joint density of $X_{it}$ and $X_{jt}$ is not degenerate for any $\left( i,j\right) $ pair (see Assumption (ref) in Section (ref)), the covariance terms between $\hat{\gamma}_{i}$ and $\hat{\gamma}_{j}$ will be of lower order compared to the variance terms, $O(1/T)$. For a more detailed discussion on the key intuition on the impact of the dependence structure in a simplified setting, we refer to Section (ref) of the Supplementary Material.} For a discussion on the issue of bandwidth selection, we refer to Remark (ref) in Section (ref) in the Supplementary Material.}
In this section, we present our main results, with a particular focus on Theorem (ref), which provide Gaussian approximation results that determine the critical values for testing the existence of threshold effects. Theorem (ref) extends the result to unknown threshold case and Theorem (ref) shows the consistency of the estimated threshold locations. Theorem (ref) shows GA for testing homogeneity. Theorems (ref) and (ref) derive the consistency of variance to ensure the validity of our test statistics. Theorem (ref) analyzes the power of the tests for both the threshold effects and homogeneity. Results regarding the derivative cases are presented in Subsection (ref).
In this subsection, we state the assumptions for the asymptotic analyses in Section (ref) and Section (ref) in the Supplementary Material. First of all, we will pose a general spatial and temporal dependence assumption on the data generating processes. As for the covariate variable $X_{jt}$ and the confounding variable $U_{jt} \in \mathbb{R}^d$, we assume that they take the form
where $\tilde{H}_{j}=(\tilde{H}_{j1},\ldots, \tilde{H}_{jd})^\top$, $\{\tilde{H}_{ji}\}_{j\in \lbrack N], i\in \lbrack d]}$ and $\left\{ H_{j}\right\} _{j\in \lbrack N]}$ are functions so that $X_{jt}$ and $U_{jt}$ are well defined. For $t\in \mathbb{Z}$, $e_{t}\in\mathbb{R}$ denote i.i.d. innovations, and $\tilde{e} _{t}\in\mathbb{R}^{\tilde d}$ are i.i.d. random vectors with some constant $\tilde d>0$, independent of $\left\{ e_{t}\right\} _{t\in \mathbb{Z}}$. Denote
For the innovation part $\epsilon _{t}=(\epsilon_{t1},\ldots,\epsilon_{tN})^\top$ as in ((ref)), the dependence among $\epsilon _{t}:=(\epsilon _{1t},\ldots ,\epsilon _{Nt})^{\top }$ is allowed to be strong (e.g., with a factor structure) along the cross-sectional dimension and it has to be weak along the time dimension, as we only exploit the sample size in the time series dimension for the estimation of heterogeneous threshold effects. Specifically, we assume that $\epsilon _{t}$ follows an MA($\infty $) process as follows:
where $\eta _{t}=(\eta _{1t},\ldots ,\eta _{\tilde{N}t})^{\top }\in \mathbb{R }^{\tilde{N}},$ $t\in \mathbb{Z},$ are i.i.d. random vectors with zero mean and identity covariance matrix $I_{\tilde{N}}$, and $A_{k}\in \mathbb{R} ^{N\times \tilde{N}}$ are real-valued matrices with $\tilde{N}\leq cN$ for some constant $c>0$. In particular, we allow $\tilde{N}=N$ so that the $A_{k}$'s become square matrices. We assume that $\left\{ e_{t},\tilde{e} _{t}\right\} _{t\in \mathbb{Z}}$ are independent of $\left\{ \eta _{t}\right\} _{t\in \mathbb{Z}}.$
Next, we impose some conditions on the elements of $\eta _{t}.$
{ The following assumption imposes some conditions on the temporal dependence for the processes $(X_{jt})_{t\in \mathbb{Z}}$ and $ (\epsilon _{jt})_{t\in \mathbb{Z}}.$ } {Denote $\mathcal{C}_j$ as a set of search values for the time series $j$ containing at least the threshold values $c_{0j}$. If the true threshold is given, then $\mathcal{C}_j = \{c_{0j}\}$ is a single point in subsections (ref) and (ref). We will specify $\mathcal{C}_j$ later in the theorems for the unknown case.} Let $g_{j}(\cdot )$ be the density function of $X_{jt}.$ Recall $\mathcal F_s$ in (ref), for $t\geq s+1,$ let $g_{j,t}(x|\mathcal{F}_{s}):=d\mathbb{P}(X_{jt}\leq x| \mathcal{F}_{s})/dx.$
Assumption (ref) $(i)$-$(ii)$ essentially assumes weak temporal dependence for the processes $(X_{jt})_{t\in \mathbb{Z}}$ and $(\epsilon _{jt})_{t\in \mathbb{Z}}$ for any $j\in \left[ N\right]$. We allow strong cross-sectional dependence over the $j$ dimension. We shall also note that $\beta $ indicates the strength of the dependency structure of the error processes.
Furthermore, let $g_{j_{1},j_{2}}(\cdot ,\cdot |\mathcal{F}_{t-1})$ be the conditional joint density of $X_{j_{1}t}$ and $X_{j_{2}t}$ given $\mathcal{F}_{t-1}.$ Finally, the following assumption imposes a condition on $ g_{j_{1},j_{2}}(\cdot ,\cdot |\mathcal{F}_{t-1}).$
Let
Assumptions (ref)-(ref) in Section (ref) in the Supplementary Material are concerning standard assumptions on boundedness, smoothness and kernel functions involved.
{If the true threshold is given, then $\mathcal{C}_j = \{c_{0j}\}$ is a single point in Subsections (ref) and (ref).} {We first consider the GA result for $\mathcal{I}$ defined in ( (ref)). Define $d_{j}=(Tb_j)^{1/2}\gamma _{j}/v_{j}$ and $\underline{d} =(d_{1},d_{2},\ldots ,d_{N})^{\top }$. Further define $Z$ as a centered Gaussian random vector with identity covariance matrix. Under the null }$H_{0}^{\left( 1\right) },$ the bias term $\underline{d}$ becomes a zero vector. The following theorem states that the infeasible statistic $\mathcal{I}$ can be approximated well by the maximum of Gaussian random variables. Denote
Define }$\mathcal{R}_{NT}=\Delta +(\underline{b}T^{1-2/q})^{-1/3}\mathrm{log} (NT)+\mathrm{log}(NT)N^{1/q}T^{-\beta }$ under Assumption 3.1$(i)$, and $ \mathcal{R}_{NT}=\Delta +\mathrm{log}(NT)T^{-\beta }$ under Assumption 3.1$(ii)$.
Now we list the related rate requirement on $T$, $\bar b$ and $N$ under different moment conditions. Under Assumption (ref)$(i)$, if $T\bar{b}^{5}\mathrm{log}(N)\rightarrow 0,$ $\underline{b} ^{-1}T^{2/q-1}\mathrm{log}^{3}(NT)\rightarrow 0$ and $\mathrm{log} (NT)N^{1/q}T^{-\beta }\rightarrow 0,$ then $\Delta \rightarrow 0$ and $\mathcal{R}_{NT}\to 0$. Under Assumption (ref)$(ii)$, if $T\bar{b}^{5}\mathrm{log}(N)\rightarrow 0,$ $(T\underline{b})^{-1} \mathrm{log}^{7}(NT)\rightarrow 0,$ $\bar b\mathrm {log}^3(N)\rightarrow0$ and $\mathrm{log}(NT)T^{-\beta }\rightarrow 0,$ then $\Delta \rightarrow 0$ and $\mathcal{R}_{NT}\to 0$.
{ Theorem (ref) lays down the foundation for testing the null hypothesis $H_{0}^{\left( 1\right) }.$ The proof of Theorem (ref) is technically challenging and involved for two reasons. First, we allow for both temporal and cross-sectional dependence among $\left\{ \left( X_{jt},\epsilon _{jt}\right) \right\} .$ Second, the latent variable \(U_{jt}\) makes the error term \(e_{jt}\) in ((ref)) highly complex. Such complicated structure makes it impossible to apply any existing GA results or their proof strategies directly. In Section (ref) of the online supplement, we outline the main idea used in the proof of the above theorem, highlighting that the key step is proving the GA result in Lemma (ref). We note that the term $\mathrm{log}^{7/6}(NT)(\underline{b} T)^{-1/6} + \mathrm{log}^{1/2}(N)T^{1/2}\bar{b}^{5/2}$ arises from extending the original GA results in chernozhukov_central_2017. The term $\mathrm{log}(NT)N^{1/q}T^{-\beta }$ is due to the temporal dependence and the moment condition of the {innovation term $\eta _{jt}$}. The term $\bar{b}^{1/3} \mathrm{log}^{2/3}(N)$ results from the covariance approximation and Gaussian comparison. } {Here we adopt maximum based test statistic, extending to other type of statistic can also be considered, see for example the statistics in fan2015power and giessing2020bootstrapping.}
Next, we consider the GA result for {$\mathcal{Q}$} defined in ((ref)). Note that
where $\tilde{d}_{j}=(Tb_{j})^{1/2}(\gamma _{j}-\frac{1}{N}\sum_{i=1}^{N}\gamma _{i})/ \tilde{v}_{j}.$ The approximation is valid due to the smoothness of $\tilde h_j(\cdot)$ and $\tilde \tau_j(\cdot)$, the remainder is addressed in Subsection (ref). Let $\tilde{\underline{d}}=(\tilde{d}_{1},\tilde{d} _{2},\ldots ,\tilde{d}_{N})^{\top }.$ We deliver the GA results for $\mathcal{Q}$ in Supplementary Material (ref) and the proofs in Section (ref). Furthermore, the theoretical results regarding the derivative are given in Section (ref) and the proofs in Section (ref). In addition, in Section (ref), we establish the validity of the test statistics when variance estimators are utilized.
The most interesting finding is that in the GA result (ref), neither temporal nor cross-sectional dependencies appear in $|Z+\underline{d}|_\infty$, despite allowing strong cross-sectional dependence for the underlying data generating process $X_{jt}, U_{jt}, \epsilon_{jt}$ within $\mathcal{I}$. The theoretical intuition behind this result is that local smoothing effectively whitens the noise. In practice, this observation suggests that, when the Assumption (ref) hold for cross-sectional dependence, the correlation in the estimated statistics can be safely ignored, requiring only adjustments for heteroskedasticity.
In this section, we present theoretical results in the case of unknown thresholds. For simplicity of notation, we define $\mathcal{C}_j$ as the unified interval $\mathcal{C}_j = [c_{\min}, c_{\max}]$, where $c_{\min}$ and $c_{\max}$ are constants in $\mathbb{R}$. The theorem still holds even when $\mathcal C_j$ differs for different $j$. Recall that $[c_1,\ldots, c_K]$ is the grid we search over. Let $c_1=c_{\min}$ and $c_{K}=c_{\max}.$ As our statistics aggregate over different threshold grids, the problem becomes more demanding. Let $\Delta_{c,\min }=\min_{1\leq i\leq {K-1}}|c_{i+1}-c_{i}|$ and $\Delta_{c,\max}=\max_{1\leq i\leq {K-1}}|c_{i+1}-c_{i}|$.
Recall $w_{jt,b}(c_{i})$ in (ref) and define ${\gamma }_{j}(c_{i})=\sum_{t=1}^{T}w_{jt,b}(c_{i}) \gamma _{j}$. Let $$d_{jK}=(\gamma _{j}(c_{1})/v_{j}(c_{1}),\ldots ,\gamma _{j}(c_{K})/v_{j}(c_{K}))^{\top }$$ and $\underline{d} ^{C}=(d_{1K}^{\top },d_{2K}^{\top },\ldots ,d_{NK}^{\top })^{\top }.$ Define $Z^{C}=\left( {{Z_{1}^{C\top },...,Z_{N}^{C\top }}}\right) ^{\top }$ as an $NK$-vector of mean zero Gaussian distributed random vector with covariance matrix $\Sigma^C$ specified below. For each $j\in \lbrack N]$ and $1\leq i_{1},i_{2}\leq K$, the covariance between the $i_{1}$- and $i_{2}$-th elements of $Z_{j}^C$ {{is given by{
}}}${\normalsize {{Z_{j_1}^C}}}$ and $Z_{j_2}^C$ are independent for $ j_1\neq j_2$, so that the matrix $\Sigma^C$ is block diagonal. Note that $ \underline{d}^{C}$ denotes a bias term which is a zero vector under the null hypothesis. In addition, when $2\max_{1\leq j\leq N}b_{j}<\Delta_{c,\min },$ one can ensure the matrix to be diagonal. The following theorem establishes the asymptotic property of our test statistics with unknown thresholds.
The proof of the above theorem is a generalization of Theorem (ref) with a different variance-covariance structure. It should be noted that the conclusion of the above theorem follows the same pattern as Theorem (ref), except that $N$ is replaced with $NK$. The following theorem provides the theoretical support for the consistency of the threshold estimator. Note that we only focus on the location $j$ where $\gamma_j$ does not equal to $0$, and we require those breaks to be significant.
If the true break $c_{0j}$ does not fall in the grid, then the difference is controlled by $\max\{\Delta_{c,\max}, \mathrm {log}(N)/(T\gamma_j^2)\}.$ Under the assumption of a minimum break signal, we can achieve consistency in detecting the break. The rate is expected to depend on both the signal strength $\gamma_j$ and the available sample size $T$. Note that the precision of threshold estimation is not affected by the bandwidth $b_j$. This rate is in line with the high-dimensional changepoint literature, see for example Theorem 2 in li2024.
{ To further evaluate the feasibility of the test statistic $\mathcal{\hat{I}}^{C}$, we will also examine the consistency of the variance estimator $\hat{\sigma}_{e,j}^2(c_i)$ defined in Section (ref) of the Supplementary Material. Estimating the variance in the presence of an unknown threshold requires additional caution, particularly when accounting for jumps, see Section (ref). We will consider a truncation operation to eliminate the contamination from jumps. The consistency of the estimator will be established in a similar manner. For completeness, we provide Theorems (ref) and (ref) in Section (ref) of the Supplementary Material as counterparts to Theorems (ref) and (ref), specifically addressing the case with an unknown threshold. }
In this section, we analyze the finite sample performance of our GA results. In particular, we study the size and power of our uniform testing procedures in different settings. The data generating processes(DGPs) 1-7 are presented in Section (ref) in Supplementary Materials.
Throughout the study, we use a local linear estimator with a uniform kernel and select the bandwidth according to the procedure described in calonico2014robust. We consider three significance levels, namely, $\alpha =0.1,$ $0.05,$ and $0.01$. The results are based on $1000$ Monte Carlo iterations. All tables and figures are presented in Section (ref) in Supplementary Materials.
{ The results of simultaneous jump effects testing for DGP1--DGP3 are displayed in Table (ref). First, the empirical size closely matches the nominal size across all $N$ values and for large $T$ in all DGPs. For DGP3, we observe slightly oversize when $N=10\ $ which may be caused by the presence of heteroskedasticity in comparison with results with DGP2. Overall, these simulation results confirm our theoretical findings in Theorem (ref). In particular, it shows that the test procedure, which is based on critical values neglecting the dependence in the covariance structure, has fairly accurate size control even in the case of strong cross-sectional dependence in the innovations. Second, the local power properties of the test are reasonably well for all DGPs under investigation, and the power increases as either $N$ or $T$ increases. This result confirms that our testing procedure is well-suited to detect sparse alternatives.
In addition, we achieve comparable simulation performance in terms of size and power results for the test of homogeneity of the threshold effects in Table (ref). The results confirm our Gaussian approximation result in Theorem (ref). For robustness, we also consider three additional settings with even stronger cross-sectional dependence for DGP4--6 in Table (ref). The results show that size and power are indeed robust to strong cross-sectional dependence. Furthermore, Table (ref) demonstrates that detecting the existence of jump effects simultaneously across cross-sectional dimensions in the case of unknown thresholds has size and power properties as good as in the known threshold case, confirming our theoretical results from Theorem (ref). We additionally show the good finite sample performance for the estimation accuracy of the threshold location estimator in Table (ref), providing Monte Carlo evidence supporting the consistency result of the threshold in Theorem (ref).}
For comparison, one direct method is the sup-Wald or sup-likelihood ratio test from the pooled threshold regression literature, neither of which accounts for nonlinearity or heterogeneity (see, for example, andrews1993tests, hansen2000sample and fong2017chngpt). Notably, recent studies, such as barassi2023threshold, propose linear models with heterogeneous threshold effects in panel settings. However, they do not provide uniform testing methods. To facilitate comparison, we firstly restrict ourself to linear threshold model DGP7. And under the alternative we sample a fraction of coefficients $\gamma_j$ from a standard normal distribution. As shown in Table (ref), our method has similar performance as the sup-Wald type statistic under the null hypothesis, while it has much better power under the alternative. See also Figure (ref) in Section (ref). The reasons for this are twofold. First, our uniform testing procedure is robust towards settings with sparse signal, in which a pooled test suffers from dilution of signals. And second, our uniform testing procedure does not have the problem of signal cancellation in case the signals for different $j$'s have opposite signs and cancel each other out. As a second comparison, we consider the nonparametric pooled test statistic under DGP3. Again, we draw the non-zero threshold coefficients from a standard normal. Such a setup without accounting for parameter heterogeneity is covered such as in spokoiny1998estimation, yang2014jump and chiou2018nonparametric. As shown in Table (ref), similar to the linear case, our method works with better power. The above superior performance is mainly due to the sparsity of the signal under the alternative hypothesis. In addition, our method has better power even for dense alternative comparing to the pooled type test statistic when there might exist signal cancellation as illustrated in Figure (ref).
We also consider the test of existence of a threshold effect in the derivative under DGP1--3, to achieve similar size and power performance, oversmoothing is needed, as shown in Table (ref).
In this section, we apply our testing procedure to two empirical applications. First, we study possible jumps in the news impact curve for constituents of the S&P 500. As a second application, we study the impact of the incumbency effect in U.S. House elections.
In this application, we revisit the news impact curve introduced by engle1993measuring. In particular, we want to identify possible jumps in the relationship between news events in the form of past returns on the volatility of stocks. The application serves as an illustration of our uniform testing procedure when the threshold $c_{0j}$ is unknown. Originally proposed in the context of autoregressive volatility models (e.g., ARCH), engle1993measuring and linton2005estimating consider partially nonparametric specifications for the news impact curve which are able to capture possible asymmetries. A limitation of these specifications is the restriction to continuous functions. However, zakoian1994threshold and fornari1997sign advocate for (parametric) GARCH specifications that allow for discontinuities in the way past returns impact volatility. We therefore want to investigate this issue by testing for the presence of jumps.
In our model setup, the dependent variable $Y_{jt}$ is the Garman-Klass volatility of stock $j$ at time $t$, and $X_{jt}$ is the lagged stock return. We want to test the null hypothesis $H_0^{(1)}:\gamma_i=0$ for all $i\in[N]$ and all $c_{0j}$, versus the alternative $H_a^{(1)}:\gamma_i\neq0$ for some $i$ and some $c_{0j}$. Our testing procedure is robust to the presence of unobserved risk factors, $U_{jt}$, which are more than likely to occur in our application. Our data includes the daily stock returns and volatility data of the S&P 500 constituents from January 2010 to May 2024, which we obtained from Kaggle \footnote{https://www.kaggle.com/datasets/andrewmvd/sp-500-stocks.}. Therefore we have a fairly high-dimensional setting with $N=500$ and $T_j$ is the number of days of stock $j$ being a member of the stock index. See Figure (ref) for a visualization of the time series of returns and volatility for the stock of AT&T. As candidates for the threshold location, we consider $c_{i}\in\{-0.01,-0.005,0,0.005,0.01\}$.
The main findings are summarized in Table (ref). We find significant results for $11$ out of the $500$ stocks we consider. For stocks with significant threshold effects, the locations of the estimated thresholds vary, and none of the significant threshold effects occur precisely at zero. This contradicts the rationale of the sign-switching ARCH model of fornari1997sign, which models a jump effect based on the sign of the lagged returns, and thus expects a significant threshold effect at zero. For $c_{i}=-0.01$, we get significant negative effects at the $1\%$ level for five stocks, including Allstate Corporation and American Express. On the other hand, we obtain significant positive effects for the $c_{i}=0.01$ threshold for four stocks, e.g. AT&T and McKesson. These results reaffirm the findings of chen2011news, who find that both bad and good news have an increasing effect on volatility. Additionally, we get another negatively significant coefficient at the $-0.005$ threshold location and another positively significant coefficient at $c_{i}=0.005$. Overall, we get evidence for the presence of jump effects in the news impact curve. Figure (ref) shows a visualization of the significant positive threshold effect for AT&T at $c_i=0.01$, and as a comparison we show the fit for Atamai for which we did not find a significant effect at any threshold location.
Furthermore, we are interested in studying the predictive advantage of our method compared to a baseline model without jumps. The baseline model can be written as $Y_{jt}=\tilde{h}_j(X_{jt})+e_{jt}$, where $\tilde h_j(\cdot)$ and $e_{jt}$ are defined in (ref). For this purpose, we split the data into two parts, with training data from 2010--2019 and test data from 2020--2024. We restrict our analysis to the stocks identified in Table (ref) with significant threshold effects at $1\%$ level. The mean squared error of prediction is lower for all but one of the 11 stocks we consider. The same is true if we restrict our analysis to these observations close to the respective threshold, $c_{0j}$. We conclude that including threshold effects can slightly improve the predictive accuracy with a median gain of $1.6\%$ in predictive power.
In the second application, we are interested in estimating the advantage a candidate has in the U.S. House elections when the seat is currently occupied by the party of the candidate \footnote{We use the data provided by the electiondata. The data include the vote shares of all candidates in the elections for the US House of Representatives from 1976--2020. Due to redistricting at the start of each decade, we have to exclude years ending with a `0' or `2'.}. This is called the party incumbency effect. We follow the research design of lee2008randomized. See caughey2011elections for an analysis of more recent elections. The dependent variable $Y_{jt}$ is the democratic two-party vote share in year $ j $ in district $t$. The covariate $X_{jt}$ is the difference in the two-party vote share in the previous election. Since the winner of the election is determined by the 0 threshold, our model setup is appropriate. We want to test the one-sided null hypothesis $H_{0}^{(1)}:\gamma _{i}<0$ for all $ i\in \lbrack N]$, versus the alternative hypothesis $H_{a}^{(1)} $: $\gamma _{i}\geq 0$ for some $i$. Unlike lee2008randomized who pools the data together to estimate a single incumbency effect, we conduct a separate analysis for different election years and also a separate analysis for different states by interchanging the roles of $j$ and $t$ and allowing for possible heterogeneity of incumbency effects along either the $j$ or $t$ dimension. In addition, it is easy to see our theories extend to the unbalanced data in which case we can use $T_{j}$ to denote the number of observations associated with individual/group $j$.
First, we are interested in the election year-specific effects. The election year-specific results are reported in Table (ref) in Appendix (ref). Our analysis covers 13 election years, so $N=13$ and $T_{j}$ is equal to the number of districts included in the sample of the respective election year/group. For each $j,$ the number of effective observations is given by the number of non-zero weights in the local linear estimation, which depends on the bandwidth parameters $b_{j}$'s which are chosen as in the simulation section. Table (ref) reports the results for the effect estimate ($\hat{\gamma}_{j}$)$,$ the standard error ($(T_{j}b_{j})^{-1/2}\hat{v}_{j}$)$,$ the individual test statistic ($(T_{j}b_{j})^{1/2} \hat{\gamma}_{j}/\hat{v}_{j}$)$,$ the number of observations $T_{j}$ (Obs) and the number of effective observations (Eobs) as well. We find five years with significant effects at the $1\%$ level and one year which is significant at the $5\%$ level by using the simulated critical value for $\mathcal{\hat{I}}$. Since we are interested in the maximum of the test statistics in our uniform testing procedure, the null hypothesis of no effect ($H_{0}^{\left( 1\right) }$) can be rejected at the $1\%$ level. We can observe that the estimated incumbency effect is stronger in the period from 1978 to 1998, while in the most recent elections the effect is insignificant. However, the sign of the estimated coefficients is positive for all election years. We also run the test for the homogeneity of incumbency effects ($H_{0}^{\left( 2\right) }$) over the 13 election years, see Table (ref). We rely on the median ($\hat{\gamma}_{0.5}$) as a more robust estimator of the average effect across election years. The test statistic is given by $\hat{\mathcal{ Q}}=3.011$, which implies that we can reject the null hypothesis of homogeneous effects at a 5% level.
Now, we are interested in the state-specific effects. We restrict our attention to states with a sufficient number of elections held in the sample. We choose a threshold of 50 elections. This gives us $N=29$ states in our analysis, and $T_{j}$ now denotes the number of elections included for the $j$th state. We argue that this setting constitutes a fairly high-dimensional setting since the effective number of observations is smaller than $N$ for a fairly large proportion of the states. The results are reported in Table (ref). Consistent with the previous results, almost all of the estimated coefficients are positive. By using the critical values for the simultaneous tests, we find significant effects for six states, for Missouri at $1\%$ level, and for California, Iowa, Virginia, Alabama and South Carolina at $5\% $. Again, we can reject the null hypothesis of no jump effect for our uniform testing procedure at a $1\%$ level. The significant states tend to belong to those with the largest number of observations. However, for the other two largest states, New York and Texas, the estimated effect is rather small and far from being significant. The incumbency effect seems to be stronger in some states compared to others. See Figure (ref) in the Supplement. for a visualization of the estimated threshold effects for the states of California and Pennsylvania. We also run the test for the homogeneity of incumbency effects ($H_{0}^{\left( 2\right) }$) over the 29 states. The results are displayed in Table (ref) and the test statistic is given by $\hat{\mathcal{Q}}=3.011$, which implies that we can reject the null hypothesis of homogeneous effects at the $5\%$ level. We find heterogeneity in the effects across election periods and states.
This paper focuses on the estimation and inference of heterogeneous threshold effects in nonparametric high-dimensional data, accounting for both cross-sectional and time dependencies in the covariates and error terms. We propose a test to assess the significance of the threshold effect and another to evaluate whether the threshold effects are homogeneous across individuals or groups. Additionally, we introduce a consistent method for estimating unknown threshold locations. Our tests are based on high-dimensional Gaussian approximation results. Despite the complexity of the underlying dependency structures, we show that the variance-covariance structure of the threshold effect estimators has a simple analytical expression under general dependence conditions. Unlike changepoint analysis in time series, our approach considers a setup involving latent variables, leading to a significantly different development of the GA results. Simulations demonstrate significant power enhancement compared to the existing literature under various settings.
{12pt} \spacingset{1.5}