EconBase
← Back to paper

Testing the Number of Regimes in Markov Regime Switching Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

139,856 characters · 21 sections · 71 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Testing the Number of Regimes in Markov Regime Switching Models Hiroyuki Kasahara Vancouver School of Economics University of British Columbia [email removed] Katsumi Shimotsu Faculty of Economics University of Tokyo [email removed]

{ \singlespacing }

abstractMarkov regime switching models have been used in numerous empirical studies in economics and finance. However, the asymptotic distribution of the likelihood ratio test statistic for testing the number of regimes in Markov regime switching models has been an unresolved problem. This paper derives the asymptotic distribution of the likelihood ratio test statistic for testing the null hypothesis of $M_0$ regimes against the alternative hypothesis of $M_0 + 1$ regimes for any $M_0\geq 1$ both under the null hypothesis and under local alternatives. We show that the contiguous alternatives converge to the null hypothesis at a rate of $n^{-1/8}$ in regime switching models with normal density. The asymptotic validity of the parametric bootstrap is also established.

Key words: Differentiable in quadratic mean expansion; likelihood ratio test; Markov regime switching model; parametric bootstrap.

Introduction

The Markov regime switching model has been a popular framework for empirical work in economics and finance. Following the seminal contribution by hamilton89em, it has been used in numerous empirical studies to model, for example, the business cycle hamilton05stlouis, morleypiger12restat, stock market volatility hamiltonsusmel94joe, international equity markets angbekaert02rfs, okimoto08jfqa, monetary policy schorfheide05red, simszha06aer, bianchi13restud, and economic growth kahnrich07jme. Comprehensive theoretical accounts and surveys of applications are provided by hamilton08palgrave, hamilton16hdbk and angtimmermann12annual.

The number of regimes is an important parameter in applications of Markov regime switching models. Despite its importance, however, testing for the number of regimes in Markov regime switching models has been an unsolved problem because the standard asymptotic analysis of the likelihood ratio test statistic (LRTS) breaks down because of problems such as unidentifiable parameters, the true parameter being on the boundary of the parameter space, and the degeneracy of the Fisher information matrix. Testing the number of regimes for Markov regime switching models with normal density, which are popular in empirical applications, poses a further difficulty because normal density has the undesirable mathematical property that the second-order derivative with respect to the mean parameter is linearly dependent on the first derivative with respect to the variance parameter, leading to further singularity.

This paper proposes the likelihood ratio test of the null hypothesis of $M_0$ regimes against the alternative hypothesis of $M_0 + 1$ regimes for any $M_0\geq 1$ and derives its asymptotic distribution. To the best of our knowledge, the asymptotic distribution of the LRTS has not been derived for testing the null hypothesis of $M_0$ regimes with $M_0 \geq 2$. To test the null hypothesis of no regime switching, namely $M_0=1$, hansen92jae derives an upper bound of the asymptotic distribution of the LRTS, and garcia98ier also studies this problem. carrasco14em propose an information matrix-type test for parameter constancy in general dynamic models including regime switching models. chowhite07em derive the asymptotic distribution of the quasi-LRTS for testing the single regime against two regimes by rewriting the model as a two-component mixture model, thereby ignoring the temporal dependence of the regimes.\footnote{cartersteigerwald12em show that ignoring temporal dependence may render the quasi-maximum likelihood estimator inconsistent.} quzhuo17wp extend the analysis of chowhite07em and derive the asymptotic distribution of the LRTS that properly takes into account the temporal dependence of the regimes under some restrictions on the transition probabilities of latent regimes. marmer08empirical and dufour17emreviews develop tests for the null hypothesis of no regime switching by using different approaches from the LRTS. The studies discussed above focus on testing the single regime against two regimes. To the best of our knowledge, however, the asymptotic distribution of the LRTS for testing the null hypothesis of $M_0$ regimes with $M_0 \geq 2$ remains unknown.

Several papers in the literature consider tests when some parameters are not identified under the null hypothesis. These include davies77bm, davies87bm, andrewsploberger94em, andrewsploberger95as, hansen96em, andrews01em, and liushao03as, among others. Estimation and testing with a degenerate Fisher information matrix are investigated in an iid setting by chesher84em, leechesher86joe, rotnitzky00bernoulli, and gukoenkervolgushev17wp, among others. chen14joe examine uniform inference on the mixing probability in mixture models.

To facilitate the analysis herein, we develop a version of Le Cam's differentiable in quadratic mean (DQM) expansion that expands the likelihood ratio under the loss of identifiability, while adopting the reparameterization and higher-order expansion of kasaharashimotsu15jasa. In an iid setting, liushao03as develop a DQM expansion under the loss of identifiability in terms of the generalized score function. We extend liushao03as to accommodate dependent and heterogeneous data as well as modify them to fit our context of parametric regime switching models. Using a DQM-type expansion has an advantage over the “classical” approach based on the Taylor expansion up to the Hessian term because deriving a higher-order expansion becomes tedious as the order of expansion increases in a Markov regime switching model. Furthermore, regime switching models with normal components are not covered by liushao03as because their Theorem 4.1 assumes that the generalized score function is obtained by expanding the likelihood ratio twice, whereas our Section (ref) shows that the score function is a function of the fourth derivative of the likelihood ratio in the normal case.

Our approach follows dmr04as [DMR hereafter], who derive the asymptotic distribution of the maximum likelihood estimator (MLE) of regime switching models. We express the higher-order derivatives of the period density ratios in terms of the conditional expectation of the derivatives of the period complete-data log-density, i.e., the log-density when the state variable is observable, by applying the missing information principle woodbury71biometrics,louis82jrssb and extending the analysis of DMR. We then show that these derivatives of the period density ratios can be approximated by a stationary, ergodic, and square integrable martingale difference sequence by conditioning on the infinite past, and this approximation is shown to satisfy the regularity conditions for our DQM expansion.

We first derive the asymptotic null distribution of the LRTS for testing $H_0: M=1$ against $H_A: M=2$. When the regime-specific density is not normal, the log-likelihood function is locally approximated by a quadratic function of the second-order polynomials of the reparameterized parameters. When the density is normal, the degree of deficiency of the Fisher information matrix and required order of expansion depend on the value of the unidentified parameter; in particular, when the latent regime variables are serially uncorrelated, the model reduces to a finite mixture normal model in which the fourth-order DQM expansion is necessary to derive a quadratic approximation of the log-likelihood function. We expand the log-likelihood with respect to the judiciously chosen polynomials of the reparameterized parameters---which involves the fourth-order polynomials---to obtain a uniform approximation of the log-likelihood function in quadratic form and derive the asymptotic null distribution of the LRTS by maximizing the quadratic form under a set of cone constraints building on the results of andrews99em,andrews01em.

To derive the asymptotic null distribution of the LRTS for testing $H_0: M=M_0$ against $H_A: M=M_0+1$ for $M_0\geq 2$, we partition a set of parameters that describes the true null model in the alternative model into $M_0$ subsets, each of which corresponds to a specific way of generating the null model. We show that the asymptotic distribution of the LRTS is characterized by the maximum of the $M_0$ random variables, each of which represents the LRTS for testing each of the $M_0$ subsets.

We also derive the asymptotic distribution of the LRTS under local alternatives. carrasco14em show that the contiguous local alternatives of their tests are of order $n^{-1/4}$, where $n$ is the sample size. In a related problem of testing the number of components in finite mixture normal regression models, kasaharashimotsu15jasa show that the contiguous local alternatives are of order $n^{-1/8}$ chenli09as, chenlifu12jasa, honguyen16as. We show that the value of the unidentified parameter affects the convergence rate of the contiguous local alternatives. When the regime-specific density is normal, some contiguous local alternatives are of the order $n^{-1/8}$, and the LRT is shown to have non-trivial power against them. The tests of carrasco14em do not have power against such alternatives, whereas the test of quzhuo17wp rules out such alternatives because of their restriction on the parameter space.

The asymptotic validity of the parametric bootstrap is also established both under the null hypothesis and under local alternatives. The simulations show that our bootstrap LRT has good finite sample properties. Our results also imply that the bootstrap LRT is valid for testing the number of hidden states in hidden Markov models because this paper's model includes the hidden Markov model as a special case. Although several papers have analyzed the asymptotic property of the MLE of the hidden Markov model,\footnote{See, for example, leroux92spa, francq98stat, krishnamurthy98jtsa, brr98as, jp99as, legrandmevel00math, and doucmatias01bernouiil.} the asymptotic distribution of the LRTS for testing the number of hidden states has been an open question. \footnote{gassiat00esaim show that the LRTS for testing $H_0: M=1$ against $H_A: M=2$ diverges when state-specific densities have known and distinct parameter values. dannemann08cjstat analyze the modified quasi-LRTS for testing the null of two states against three.}

The remainder of this paper is organized as follows. After introducing the notation and assumptions in Section 2, we discuss the degeneracy of the Fisher information matrix and loss of identifiability in regime switching models in Section 3. Section 4 establishes the DQM-type expansion. Section 5 presents the uniform convergence of the derivatives of the density ratios. Sections 6 and 7 derive the asymptotic null distribution of the LRTS. Section 8 derives the asymptotic distribution under local alternatives. Section 9 establishes the consistency of the parametric bootstrap. Section 10 reports the results from the simulations and an empirical application, using U.S. GDP per capita quarterly growth rate data. Section 11 collects the proofs and auxiliary results.

Notation and assumptions

Let $:=$ denote “equals by definition.” Let $\Rightarrow$ denote the weak convergence of a sequence of stochastic processes indexed by $\pi$ for some space $\Pi$. For a matrix $B$, let $\lambda_{\min}(B) $ and $\lambda_{\max}(B)$ be the smallest and largest eigenvalues of $B$, respectively. For a $k$-dimensional vector $x = (x_1,\ldots,x_k)'$ and a matrix $B$, define $|x| := \sqrt{x'x}$ and $|B| := \sqrt{\lambda_{\max}(B'B)}$. For a $k\times 1$ vector $a=(a_1,\ldots,a_k)'$ and a function $f(a)$, let $\nabla_{a} f(a):=(\partial f(a)/\partial a_{1},\ldots,\partial f(a)/\partial a_{k})'$, and let $\nabla_{a}^j f(a)$ denote a collection of derivatives of the form $(\partial^j/\partial a_{i_1}\partial a_{i_2} \ldots \partial a_{i_j})f(a)$. Let $\mathbb{I}\{A\}$ denote an indicator function that takes the value 1 when $A$ is true and 0 otherwise. $\mathcal{C}$ denotes a generic non-negative finite constant whose value may change from one expression to another. Let $a \vee b :=\max\{a,b\}$ and $a \wedge b :=\min\{a,b\}$. Let $\lfloor x \rfloor$ denote the largest integer less than or equal to $x$, and define $(x)_+ := \max\{x,0\}$. Given a sequence $\{f_k\}_{k=1}^n$, let $\nu_n(f_k) := n^{-1/2} \sum_{k=1}^n [f_k - \mathbb{E}_{\vartheta^*}(f_k)]$. For a sequence $X_{n\varepsilon}$ indexed by $n=1,2,\ldots$ and $\varepsilon$, we write $X_{n\varepsilon} = O_{p\varepsilon}(a_n)$ if, for any $\delta>0$, there exist $\varepsilon>0$ and $M, n_0 <\infty$ such that $\mathbb{P}(|X_{n\varepsilon}/a_n| \leq M) \geq 1- \delta$ for all $n > n_0$, and we write $X_{n\varepsilon} = o_{p\varepsilon}(a_n)$ if, for any $\delta_1,\delta_2>0$, there exist $\varepsilon>0$ and $n_0$ such that $\mathbb{P}(|X_{n\varepsilon}/a_n| \leq \delta_1) \geq 1- \delta_2$ for all $n > n_0$. Loosely speaking, $X_{n\varepsilon} = O_{p\varepsilon}(a_n)$ and $X_{n\varepsilon} = o_{p\varepsilon}(a_n)$ mean that $X_{n\varepsilon} = O_{p}(a_n)$ and $X_{n\varepsilon} = o_{p}(a_n)$ when $\varepsilon$ is sufficiently small, respectively. All limits are taken as $n \to \infty$ unless stated otherwise. The proofs of all the propositions and lemmas are presented in the appendix.

Consider the Markov regime switching process defined by a discrete-time stochastic process $\{(X_k,Y_k,W_k)\}$, where $(X_k,Y_k,W_k)$ takes values in a set $\mathcal{X}_M\times \mathcal{Y}\times \mathcal{W}$ with $\mathcal{Y}\subset \mathbb{R}^{q_y}$ and $\mathcal{W}\subset \mathbb{R}^{q_w}$, and let $\mathcal{B}(\mathcal{X}_M\times\mathcal{Y}\times\mathcal{W})$ denote the associated Borel $\sigma$-field. For a stochastic process $\{Z_k\}$ and $a<b$, define ${\bf Z}_{a}^b:=(Z_a,Z_{a+1},\ldots,Z_b)$. Denote $\overline {\bf Y}_{k-1}:=(Y_{k-1},\ldots,Y_{k-s})$ for a fixed integer $s$ and $\overline {\bf Y}^b_{a}:=(\overline {\bf Y}_{a},\overline {\bf Y}_{a+1},\ldots,\overline {\bf Y}_{b})$. Here, $Y_k$ is an observable variable, $X_k$ is an unobservable state variable, $\overline {\bf Y}_{k-1}$ is the lagged $Y_k$'s used as a covariate, and $W_k$ is a weakly exogenous covariate. DMR's model does not include $W_k$.

assumption(a) $\{X_k\}_{k=0}^\infty$ is a first-order Markov chain with the state space $\mathcal{X}_M:=\{1,2,\ldots,M\}$. (b) For each $k\geq 1$, $X_k$ is independent of $({\bf X}_{0}^{k-2},\overline {\bf Y}_{0}^{k-1},{\bf W}_{0}^\infty)$ given $X_{k-1}$. (c) For each $k \geq 1$, $Y_k$ is conditionally independent of $( {\bf Y}_{-s+1}^{k-s-1}, {\bf X}_{0}^{k-1}, {\bf W}_{0}^{k-1},{\bf W}_{k+1}^{\infty})$ given $(\overline {\bf Y}_{k-1},X_k,W_k)$. (d) $ {\bf W}_{1}^{\infty}$ is conditionally independent of $(\overline {\bf Y}_{0},X_{0})$ given $W_0$.\footnote{Assumptions (ref)(a)--(d) imply that $W_k$ is conditionally independent of $({\bf X}_{0}^{k-1},\overline{\bf Y}_{0}^{k-1})$ given ${\bf W}_{0}^{k-1}$.} (e) $\{(X_k,Y_k,W_k)\}_{k=0}^{\infty}$ is a strictly stationary ergodic process.

When $W_k$ is absent, DMR provide a sufficient condition for the ergodicity of $(X_k,Y_k)$ in their Assumption (A2). We assume the ergodicity of $(X_k,Y_k,W_k)$ for brevity.

The unobservable Markov chain $\{X_k\}$ is called the regime. The integer $M$ represents the number of regimes specified in the model. The parameter $\vartheta_M=(\vartheta_{M,y}',\vartheta_{M,x}')'$ belongs to $\Theta_M= \Theta_{M,y}\times \Theta_{M,x}$, a compact subset of $\mathbb{R}^{q_M}$. $\vartheta_{M,x}$ contains the parameter of the transition probability of $X_k$, which we denote by $q_{\vartheta_{M,x}}(x_{k-1},x_k):= \mathbb{P}(X_{k}=x_k|X_{k-1}=x_{k-1})$. Let $p_{ij}:= q_{\vartheta_{M,x}}(i,j)$ for $i=1,\ldots,M$ and $j=1,\ldots,M-1$, and $q_{\vartheta_{M,x}}(i,M)$ is determined by $q_{\vartheta_{M,x}}(i,M) = 1-\sum_{j = 1}^{M - 1} p_{ij}$. $\vartheta_{M,y}=(\theta_1',\ldots,\theta_M',\gamma')'$ contains the parameter of the conditional density of $Y_k$ given $(\overline {\bf Y}_{k-1},X_k,W_k)$, which is given by $g_{\vartheta_{M,y}}(y_k|\overline{\bf y}_{k-1},x_k,w_k) := \sum_{j \in \mathcal{X}_M} \mathbb{I}\{x_k=j\}f(y_k|\overline{\bf{y}}_{k-1},w_k;\gamma,\theta_j)$. Here, $\gamma$ is the structural parameter that does not vary across regimes, $\theta_j$ is the regime-specific parameter that varies across regimes, and $f(y_k|\overline{\bf{y}}_{k-1},w_k;\gamma,\theta_j)$ is the conditional density of $y_k$ given $(\overline{\bf{y}}_{k-1},w_k)$ when $x_k=j$. Let

align*[align* omitted — 293 chars of source]

We assume $\Theta_{M,y}=\Theta_{\theta} \times \cdots \times \Theta_\theta\times \Theta_\gamma$, and the true parameter value is denoted by $\vartheta_M^*$.

We make the following assumptions that correspond to (A1)--(A3) in DMR.

assumption(a) $0<\sigma_{-}:=\inf_{\vartheta_{M,x}\in\Theta_{M,x}} \min_{x,x'\in \mathcal{X}_M} q_{\vartheta_{M,x}}(x,x')$ and \\$\sigma_+:=\sup_{\vartheta_{M,x}\in\Theta_{M,x}}\max_{x,x'\in\mathcal{X}_M}q_{\vartheta_{M,x}}(x,x')<1$ for each $M$. (b) For all $y'\in \mathcal{Y}$, $\overline y\in \mathcal{Y}^s$, and $w\in \mathcal{W}$, $0<\inf_{\vartheta_y\in\Theta_{M,y}} \sum_{x\in \mathcal{X}_M} g_{\vartheta_{M,y}}(y'|\overline y,x,w)$ and $\sup_{\vartheta_{M,y}\in\Theta_{M,y}} \sum_{x\in \mathcal{X}_M} g_{\vartheta_{M,y}}(y'|\overline y,x,w) <\infty$. (c) $b_+:=\sup_{\vartheta_{M,y}\in\Theta_y} \sup_{\overline {\bf y}_0,y_1,w,x} g_{\vartheta_{M,y}}(y_1|\overline {\bf y}_0,x,w)<\infty$ and $\mathbb{E}_{\vartheta^*}(|\log b_{-}(\overline{\bf{Y}}_0^1,W_1)|)<\infty$, where $b_{-}(\overline{\bf y}_0^1,w_1):=\inf_{\vartheta_{M,y}\in\Theta_{M,y}} \sum_{x\in\mathcal{X}_M} g_{\vartheta_{M,y}}(y_1|\overline{\bf y}_0,x,w_1)$.

As discussed on p. 2260 of DMR, Assumption (ref)(a) implies that the Markov chain $\{X_k\}$ has a unique invariant distribution and is uniformly ergodic for all $\theta_{M,x} \in \Theta_{M,x}$.\footnote{Assumptions (ref)(c) and (ref)(a) are also employed in DMR. As discussed in kasaharashimotsu17hamilton, these assumptions together rule out models in which the conditional density $Y_k$ depends on both current and lagged regimes. kasaharashimotsu17hamilton show the asymptotic normality of the MLE while relaxing Assumption (ref)(a) to allow for $\inf_{\vartheta_{M,x}\in\Theta_{M,x}} \min_{x,x'\in \mathcal{X}_M} q_{\vartheta_{M,x}}(x,x')=0$. It is possible to derive the asymptotic distribution of the LRT under similar assumptions to kasaharashimotsu17hamilton, albeit with a tedious derivation.} For notational brevity, we drop the subscript $M$ from $\mathcal{X}_M$, $\vartheta_{M}$, $\Theta_{M}$, etc., unless it is important to clarify the specific value of $M$. Assumptions (ref)(b) and (c) imply that $\{Z_k\}_{k=0}^{\infty}:=\{(X_k,\overline{\bf Y}_{k})\}_{k=0}^{\infty}$ is a Markov chain on $\mathcal{Z}:=\mathcal{X}\times\mathcal{Y}^s$ given $\{W_k\}_{k=0}^{\infty}$, and $Z_k$ is conditionally independent of $({\bf Z}_0^{k-2},{\bf W}_{0}^{k-1},{\bf W}_{k+1}^{\infty})$ given $(Z_{k-1},W_k)$. Consequently, Lemma 1, Corollary 1, and Lemma 9 of DMR go through even in the presence of $\{W_k\}_{k=0}^{\infty}$. Because $\{(Z_k,W_k)\}_{k=0}^\infty$ is stationary, we can and will extend $\{(Z_k,W_k)\}_{k=0}^\infty$ to a stationary process $\{(Z_k,W_k)\}_{k=-\infty}^\infty$ with doubly infinite time. We denote the probability measure and associated expectation of $\{(Z_k,W_k)\}_{k=\infty}^\infty$ under stationarity by $\mathbb{P}_\vartheta$ and $\mathbb{E}_\vartheta$, respectively.\footnote{DMR use $\overline{\mathbb{P}}_\vartheta$ and $\overline{\mathbb{E}}_\vartheta$ to denote probability and expectation under stationarity on $\{Z_k\}_{k=\infty}^\infty$, because their Section 7 deals with the case when $Z_0$ is drawn from an arbitrary distribution. Because we assume $\{(Z_k,W_k)\}_{k=\infty}^\infty$ is stationary, we use notations such as $\mathbb{P}_\vartheta$ and $\mathbb{E}_\vartheta$ without an overline for simplicity.}

Under Assumptions (ref)(a)--(d), the density of ${\bf Y}_{1}^n$ given $X_0=x_0$, $\overline{\bf Y}_{0}$ and ${\bf W}_{0}^n$ is given by

equation[equation omitted — 223 chars of source]

Define the conditional log-likelihood function and stationary log-likelihood function as

align*[align* omitted — 368 chars of source]

where we use the fact that $p_\vartheta (Y_k| \overline{\bf Y}_{0}^{k-1},{\bf W}_{0}^{n},x_0) = p_\vartheta (Y_k| \overline{\bf Y}_{0}^{k-1},{\bf W}_{0}^{k},x_0)$ and $p_\vartheta (Y_k| \overline{\bf Y}_{0}^{k-1},{\bf W}_{0}^{n}) = p_\vartheta (Y_k| \overline{\bf Y}_{0}^{k-1},{\bf W}_{0}^{k})$, which follows from Assumption (ref). Note that

align*[align* omitted — 422 chars of source]

and $\mathbb{P}_\vartheta(x_{k-1}|\overline{\bf Y}_{0}^{k-1},{\bf W}_{0}^{k-1}) = \sum_{x_0 \in \mathcal{X}}\mathbb{P}_\vartheta(x_{k-1}|\overline{\bf Y}_{0}^{k-1},{\bf W}_{0}^{k-1},x_0)\mathbb{P}_\vartheta(x_0|\overline{\bf Y}_{0}^{k-1},{\bf W}_{0}^{k-1})$. Let $\rho:=1-\sigma_-/\sigma_+\in [0,1)$. Lemma (ref)(a) in the appendix shows that, for all probability measures $\mu_1$ and $\mu_2$ on $\mathcal{B}(\mathcal{X})$ and all $(\overline{\bf y}_{0}^{k-1},{\bf w}_{0}^{k-1})$,

equation[equation omitted — 309 chars of source]

Consequently, $ p_\vartheta (Y_k| \overline{\bf Y}_{0}^{k-1},{\bf W}_{0}^{k-1},x_0) - p_\vartheta (Y_k| \overline{\bf Y}_{0}^{k-1},{\bf W}_{0}^{k-1})$ goes to zero at an exponential rate as $k\rightarrow \infty$. Therefore, as shown in the following proposition, the difference between $\ell_n(\theta,x_0)$ and $\ell_n(\theta)$ is bounded by a deterministic constant, and the maximum of $\ell_n(\vartheta,x_0)$ and the maximum of $\ell_n(\vartheta)$ are asymptotically equivalent.

propositionUnder Assumptions (ref) and (ref), for all $x_0\in \mathcal{X}$, \[ \sup_{\vartheta\in \Theta} |\ell_n(\vartheta,x_0)- \ell_n(\vartheta)| \leq 1/(1-\rho)^2 \quad \mathbb{P}_{\theta^*}\text{-a.s}. \]

As discussed on p. 2263 of DMR, the stationary density $p_\vartheta (Y_k| \overline{\bf Y}_{0}^{k-1},{\bf W}_{0}^{k})$ is not available in closed form for some models with autoregression. For this reason, we consider the log-likelihood function when the initial distribution of $X_0$ follows some distribution \\$\xi_M \in \Xi_M:=\{\xi(x_0)_{x_0 \in \mathcal{X}_M} :\xi(x_0)\geq 0$ and $\sum_{x_0 \in \mathcal{X}_M}\xi(x_0)=1\}$.

Define the MLE, $\hat\vartheta_{M,\xi_M}$, by the maximizer of the log-likelihood:

equation[equation omitted — 174 chars of source]

where $p_{\vartheta_{M}}({\bf Y}_1^n| \overline{\bf Y}_{0},{\bf W}_{0}^n,x_0)$ is given in ((ref)). We define the number of regimes by the smallest number $M$ such that the data density admits the representation ((ref)). Our objective is to test $H_0: M=M_0$ against $H_A: M=M_0+1$.

Degeneracy of the Fisher information matrix and non-identifiability under the null hypothesis

Consider testing $H_0:M=1$ against $H_A:M=2$ in a two-regime model. The null hypothesis can be written as $H_0:\theta_1^*=\theta_2^*$.\footnote{The null hypothesis of $H_0: M=1$ also holds when $(p_{11},p_{22}) = (1,0)$. We impose Assumption (ref)(a) to bound $p_{jj}$ away from 0 and 1 because the log-likelihood function may become unbounded as $p_{11}$ or $p_{22}$ tends to 1 in view of gassiat00esaim.} When $\theta_1=\theta_2$, the parameter $\vartheta_{2,x}$ is not identified because $Y_k$ has the same distribution across regimes. Furthermore, Section (ref) shows that when $\theta_1=\theta_2$, the scores with respect to $\theta_1$ and $\theta_2$ are linearly dependent, so that the Fisher information matrix is degenerate.

The log-likelihood function of Markov regime switching models with normal density has further degeneracy. In a two-regime model where $Y_k$ in the $j$-th regime follows $N(\mu_j,\sigma_j^2)$, the model reduces to a heteroscedastic normal mixture model when the $X_k$'s are serially uncorrelated, i.e., $p_{11}=1-p_{22}$. kasaharashimotsu15jasa show that in a heteroscedastic normal mixture model, the first and second derivatives of the log-likelihood function are linearly dependent and the score function is a function of the fourth-order derivative. Consequently, one needs to expand the log-likelihood function four times to derive the score function.

Quadratic expansion under the loss of identifiability

When testing the number of regimes by the LRT, a part of $\vartheta$ is not identified under the null hypothesis. Let $\pi$ denote the part of $\vartheta$ that is not identified under the null, and split $\vartheta$ as $\vartheta = (\psi',\pi')'$. For example, in testing $H_0: M=1$ against $H_A: M=2$, we have $\psi=\vartheta_{2,y}$ and $\pi=\vartheta_{2,x}$. We denote the conditional log-likelihood function as $\ell_n(\psi,\pi, x_0) := \ell_n(\vartheta, x_0)$ and use $p_\vartheta$ and $p_{\psi\pi}$ interchangeably.

Denote the true parameter value of $\psi$ by $\psi^*$, and denote the set of $(\psi,\pi)$ corresponding to the null hypothesis by $\Gamma^*= \{(\psi,\pi)\in\Theta: \psi=\psi^*\}$. Let $t_\vartheta$ be a continuous function of $\vartheta$ such that $t_{\vartheta}=0$ if and only if $\psi=\psi^*$. For $\varepsilon>0$, define the neighborhood of $\Gamma^*$ by \[ \mathcal{N}_{\varepsilon} := \{ \vartheta \in \Theta: |t_\vartheta|< \varepsilon\}. \] When the MLE is consistent, the asymptotic distribution of the LRTS is determined by the local properties of the likelihood functions in $\mathcal{N}_{\varepsilon}$.

We establish a general quadratic expansion of the log-likelihood function $\ell_n(\psi,\pi,\xi):= \ell_n(\vartheta,\xi)$ defined in ((ref)) around $\ell_n(\psi^*,\pi,\xi)$ that expresses $\ell_n(\psi,\pi,\xi)-\ell_n(\psi^*,\pi,\xi)$ as a quadratic function of $t_\vartheta$. Once we derive a quadratic expansion, the asymptotic distribution of the LRTS can be characterized by taking its supremum with respect to $t_\vartheta$ under an appropriate constraint and using the results of andrews99em,andrews01em.

Denote the conditional density ratio by

equation[equation omitted — 205 chars of source]

so that $\ell_n(\psi,\pi, x_0) - \ell_n(\psi^*,\pi, x_0) = \sum_{k=1}^n\log l_{\vartheta k x_0}$. We assume that $l_{\vartheta k x_0}$ can be expanded around $l_{\vartheta^* k x_0}=1$ as follows. With a slight abuse of the notation, let $P_n(f_k) := n^{-1} \sum_{k=1}^n f_k$ and recall $\nu_n(f_k) := n^{-1/2} \sum_{k=1}^n [f_k - \mathbb{E}_{\vartheta^*}(f_k)]$.

assumptionFor all $k=1,\ldots,n$, $l_{\vartheta k x_0} -1$ admits an expansion \begin{equation} l_{\vartheta k x_0} -1 = t_\vartheta' s_{\pi k} + r_{\vartheta k} + u_{\vartheta k x_0}, \end{equation} where $t_\vartheta$ satisfies $\psi \to \psi^*$ if $t_\vartheta \to 0$ and $(s_{\pi k}, r_{\vartheta k}, u_{\vartheta k x_0})$ satisfy, for some $C \in (0,\infty)$, $\delta >0$, $\varepsilon >0$, and $\rho \in (0,1)$, (a) $\mathbb{E}_{\vartheta^*}\sup_{\pi \in \Theta_\pi} \left|s_{\pi k}\right|^{2+\delta} < C$, (b) $\sup_{\pi \in \Theta_\pi}| P_n(s_{\pi k}s_{\pi k}') - \mathcal{I}_\pi| = o_p(1)$ with $0<\inf_{\pi\in \Theta_{\pi} }\lambda_{\min}(\mathcal{I}_\pi) \leq \sup_{\pi\in \Theta_{\pi} }\lambda_{\max}(\mathcal{I}_\pi)<\infty$, (c) $\mathbb{E}_{\vartheta^*}[\sup_{\vartheta \in \mathcal{N}_\varepsilon} |r_{\vartheta k}/(|t_\vartheta||\psi-\psi^*|)|^2] < \infty$, (d) $\sup_{\vartheta \in \mathcal{N}_\varepsilon} [ \nu_n(r_{\vartheta k})/(|t_\vartheta||\psi-\psi^*|)] = O_p(1)$, (e) $\sup_{x_0 \in \mathcal{X}}\sup_{\vartheta \in \mathcal{N}_\varepsilon} P_n (|u_{\vartheta k x_0}|/|\psi-\psi^*|)^j = O_p(n^{-1})$ for $j=1,2,3$, (f) $\sup_{x_0 \in \mathcal{X}}\sup_{\vartheta \in \mathcal{N}_\varepsilon} P_n (|s_{\pi k}||u_{\vartheta k x_0}|/|\psi-\psi^*|) = O_p(n^{-1})$, (g) $\sup_{\vartheta \in \mathcal{N}_{\varepsilon}} |\nu_n(s_{\pi k})| =O_p(1)$.

We first establish an expansion of $\ell_n(\psi,\pi, x_0)$ in the neighborhood of $\mathcal{N}_{c/\sqrt{n}}$ for any $c>0$.

propositionSuppose that Assumptions (ref) (a)--(f) hold. Then, for any $c>0$, \begin{equation*} \sup_{x_0 \in \mathcal{X}} \sup_{\vartheta \in \mathcal{N}_{c/\sqrt{n}}} \left| \ell_n(\psi,\pi, x_0) - \ell_n(\psi^*,\pi, x_0) - \sqrt{n}t_\vartheta' \nu_n (s_{\pi k}) + n t_\vartheta' \mathcal{I}_\pi t_\vartheta/2 \right| = o_p(1). \end{equation*}

The following proposition expands $\ell_n(\psi,\pi, x_0)$ in $A_{n\varepsilon}(x_0,\eta) := \{\vartheta \in \mathcal{N}_{\varepsilon}: \ell_n(\psi,\pi, x_0)-\ell_n(\psi^*,\pi, x_0) \geq -\eta \}$ for some $\eta \in [0,\infty)$. This proposition is useful for deriving the asymptotic distribution of the LRTS because a consistent MLE is in $ A_{n\varepsilon}(x_0,\eta)$ by definition. Let $A_{n\varepsilon c}(x_0,\eta) := A_{n\varepsilon }(x_0,\eta) \cup \mathcal{N}_{c/\sqrt{n}}$.

propositionSuppose that Assumption (ref) holds. Then, for any $\eta>0$, (a) $\sup_{x_0 \in \mathcal{X}} \sup_{\vartheta \in A_{n\varepsilon}(x_0,\eta)} |t_\vartheta| = O_{p \varepsilon }(n^{-1/2})$, and (b) for any $c>0$, \begin{equation*} \sup_{x_0 \in \mathcal{X}} \sup_{\vartheta \in A_{n\varepsilon c}(x_0,\eta) }\left|\ell_n(\psi,\pi, x_0)-\ell_n(\psi^*,\pi, x_0) - \sqrt{n} t_\vartheta' \nu_n (s_{\pi k}) + n t_\vartheta' \mathcal{I}_\pi t_\vartheta/2 \right| = o_{p \varepsilon }(1). \end{equation*}

The following corollary of Propositions (ref) and (ref) shows that $\ell_n(\vartheta,\xi)$ defined in ((ref)) admits a similar expansion to $\ell_n(\vartheta,x_0)$ for all $\xi$. Consequently, the asymptotic distribution of the LRTS does not depend on $\xi$, and $\ell_n(\vartheta,\xi)$ may be maximized in $\vartheta$ while fixing $\xi$ or jointly in $\vartheta$ and $\xi$. Let $A_{n\varepsilon } (\xi,\eta) := \{\vartheta \in \mathcal{N}_{\varepsilon }: \ell_n(\psi,\pi,\xi) - \ell_n(\psi^*,\pi,\xi) \geq -\eta \}$ and $A_{n\varepsilon c} (\xi,\eta):= A_{n\varepsilon} (\xi,\eta)\cup \mathcal{N}_{c/\sqrt{n}}$, which includes a consistent MLE with any $\xi$.

corollary(a) Under the assumptions of Proposition (ref), we have \\ $\sup_{\xi \in \Xi}\sup_{\vartheta \in \mathcal{N}_{c/\sqrt{n}}} \left| \ell_n(\psi,\pi,\xi) - \ell_n(\psi^*,\pi,\xi) - \sqrt{n}t_\vartheta' \nu_n (s_{\pi k}) + n t_\vartheta' \mathcal{I}_\pi t_\vartheta/2 \right| = o_p(1)$ for any $c>0$. (b) Under the assumptions of Proposition (ref), for any $\eta>0$ and $c>0$, $\sup_{\xi\in \Xi}\sup_{\vartheta \in A_{n\varepsilon} (\xi,\eta) } |t_\vartheta| = O_{p \varepsilon }(n^{-1/2})$ and $\sup_{\xi\in \Xi}\sup_{\vartheta \in A_{n\varepsilon c} (\xi,\eta) } | \ell_n(\psi,\pi,\xi) - \ell_n(\psi^*,\pi,\xi) - \sqrt{n}t_\vartheta' \nu_n (s_{\pi k}) + n t_\vartheta' \mathcal{I}_\pi t_\vartheta/2 | = o_{p \varepsilon }(1)$.

Uniform convergence of the derivatives of the log-density and density ratios

In this section, we establish approximations that enable us to apply the results in Section (ref) to the log-likelihood function of regime switching models. Because of the presence of singularity, the expansion ((ref)) of the density ratio $l_{\vartheta k x_0}$ involves the higher-order derivatives of the density ratios $\nabla_{\psi}^j l_{\vartheta k x_0}$ with $j\geq 2$. Note that $\nabla_\psi^j l_{\vartheta k x_0}$ can be expressed in terms of the derivatives of log-densities, $\nabla_\psi^j \log p_{\psi\pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1},{\bf W}_{0}^k,x_0)$. We show that these derivatives are approximated by their stationary counterpart that condition on the infinite past $(\overline{\bf{Y}}_{-\infty}^{k-1},{\bf W}_{-\infty}^{k})$ in place of $(\overline{\bf{Y}}_{0}^{k-1},{\bf W}_{0}^{k})$. Consequently, the sequence $\{\nabla_\psi^j l_{\vartheta k x_0}\}_{k=0}^{\infty}$ is approximated by a stationary martingale difference sequence.

For $1 \leq k \leq n$ and $m \geq 0$, let

equation[equation omitted — 308 chars of source]

denote the stationary density of ${\bf Y}_{-m +1 }^k$ associated with $\vartheta$ conditional on $\{\overline{\bf{Y}}_{-m},{\bf W}_{-m}^{k}\}$, where $X_{-m}$ is drawn from its true conditional stationary distribution $\mathbb{P}_{\vartheta^*}(x_{-m}|\overline{\bf{Y}}_{-m}^{k-1},{\bf W}_{-m}^{k})$. Let $\overline p_{\vartheta}(Y_k|\overline{\bf{Y}}_{-m}^{k-1},{\bf W}_{-m}^{k}) : = \overline p_{\vartheta}({\bf Y}_{-m+1}^k|\overline{\bf{Y}}_{-m},{\bf W}_{-m}^{k})/\overline p_{\vartheta}({\bf Y}_{-m+1}^{k-1}|\overline{\bf{Y}}_{-m},{\bf W}_{-m}^{k-1})$ denote the associated conditional density of $Y_k$ given $(\overline{\bf{Y}}_{-m}^{k-1},{\bf W}_{-m}^{k})$.\footnote{Note that DMR use the same notation $\overline p_{\vartheta}(\cdot|\overline{\bf{Y}}_{-m}^{k-1})$ for a different purpose. On p.\ 2263 and in some other (but not all) places, DMR use $\overline p_{\vartheta}(Y_k|\overline{\bf{Y}}_{0}^{k-1})$ to denote an (ordinary) stationary conditional distribution of $Y_k$. }

Define the density ratio as $l_{k,m,x}(\vartheta) := p_{\vartheta}(Y_k|\overline{\bf{Y}}_{-m}^{k-1},{\bf W}_{-m}^{k},X_{-m}=x)/p_{\vartheta^*}(Y_k|\overline{\bf{Y}}_{-m}^{k-1},{\bf W}_{-m}^{k},X_{-m}=x)$. For $j=1,2,\ldots,6$, $1 \leq k \leq n$, $m \geq 0$, and $x \in \mathcal{X}$, define the derivatives of the log-densities and density ratios by, with suppressing the subscript $\vartheta$ from $\nabla_\vartheta^j$ for brevity,

align*[align* omitted — 716 chars of source]

The following assumption corresponds to (A6)--(A8) in DMR and is tailored to our setting where some elements of $\vartheta_x^*$ are not identified and $\mathcal{X}$ is finite. Note that Assumptions (A6) and (A7) in DMR pertaining to $q_{\vartheta_x}(x,x')$ hold in our case because the $p_{ij}$'s are bounded away from 0 and 1. Let $G_{\vartheta k} := \sum_{x_k\in\mathcal{X}}g_{\vartheta_y}(Y_k|\overline{\bf{Y}}_{k-1}, x_k, W_k)$. $G_{\vartheta k}$ satisfies Assumption (ref)(b) in general when $\mathcal{N}^*$ is sufficiently small.

assumptionThere exists a positive real $\delta$ such that on $\mathcal{N}^*:= \{\vartheta\in\Theta: |\vartheta_y-\vartheta_y^*| < \delta\}$ the following conditions hold: (a) For all $(\overline{\bf{y}},y',x,w)\in \mathcal{Y}^s \times \mathcal{Y} \times \mathcal{X} \times \mathcal{W}$, $g_{\vartheta_y}(y'|\overline{\bf{y}},x,w)$ is six times continuously differentiable on $\mathcal{N}^*$. (b) $\mathbb{E}_{\vartheta^*}[\sup_{\vartheta\in \mathcal{N}^*} \sup_{x\in \mathcal{X}} | \nabla^j\log g_{\vartheta_y}(Y_1|\overline{\bf{Y}}_0,x,W) |^{2q_j}]< \infty$ for $j=1,2,\ldots,6$ and $\mathbb{E}_{\vartheta^*}\sup_{\vartheta \in \mathcal{N}^*} |G_{\vartheta k}/G_{\vartheta^* k}|^{q_g} < \infty$ with $q_1=6q_0, q_2=5q_0, \ldots, q_6=q_0$, where $q_0=(1+\varepsilon) q_\vartheta$ and $q_g=(1+\varepsilon)q_\vartheta/\varepsilon$ for some $\varepsilon>0$ and $q_\vartheta >\max\{3,\text{dim}(\vartheta)\}$. (c) For almost all $(\overline{\bf{y}},y',w) \in \mathcal{Y}^s \times \mathcal{Y}\times \mathcal{W}$, $\sup_{\vartheta\in \mathcal{N}^*} g_{\vartheta_y}(y'|\overline{\bf{y}},x,w) < \infty$ and, for almost all $(\overline{\bf{y}},x,w)\in \mathcal{Y}^s\times \mathcal{X} \times \mathcal{W}$, for $j=1,2,\ldots,6$, there exist functions $f^j_{\overline{\bf{y}},w,x}: \mathcal{Y} \rightarrow \mathbb{R}^+$ in $L^1$ such that $|\nabla^j g_{\vartheta_y}(y'|\overline{\bf{y}},x,w) |\leq f^j_{\overline{\bf{y}}, x, w}(y')$ for all $\vartheta\in \mathcal{N}^*$.

The following proposition shows that $\{\nabla^j \ell_{k,m,x}(\vartheta)\}_{m\geq 0}$ and $\{\nabla^j \overline \ell_{k,m}(\vartheta)\}_{m\geq 0}$ are $L^{r_j}(\mathbb{P}_{\vartheta^*})$-Cauchy sequences that converge to $\nabla^j \ell_{k,\infty}(\vartheta)$ $\mathbb{P}_{\vartheta^*}$-a.s.\ and in $L^{r_j}(\mathbb{P}_{\vartheta^*})$ uniformly in $\vartheta\in \mathcal{N}^*$ and $x \in \mathcal{X}$.

propositionUnder Assumptions (ref), (ref), and (ref), for $j=1,\ldots,6$, there exist random variables $(K_j, \{M_{j,k} \}_{k=1}^n )\in L^{r_j}(\mathbb{P}_{\vartheta^*})$ and $\rho_* \in (0,1)$ such that, for all $1 \leq k \leq n$ and $m' \geq m \geq 0$, \begin{align*} (a) &\quad \sup_{x\in\mathcal{X}}\sup_{\vartheta\in \mathcal{N}^*}|\nabla^j\ell_{k,m,x}(\vartheta) - \nabla^j\overline\ell_{k,m}(\vartheta) | \leq K_j (k+m)^7\rho_*^{k+m-1}\quad $\mathbb{P}_{\vartheta^*}$-a.s.,\\ (b) &\quad \sup_{x\in\mathcal{X}}\sup_{\vartheta\in \mathcal{N}^*}|\nabla^j\ell_{k,m,x}(\vartheta) - \nabla^j \ell_{k,m',x}(\vartheta) | \leq K_j(k+m)^7\rho_*^{k+m-1}\quad $\mathbb{P}_{\vartheta^*}$-a.s.,\\ (c) & \quad \sup_{m \geq 0}\sup_{x\in\mathcal{X}}\sup_{\vartheta\in \mathcal{N}^*}|\nabla^j\ell_{k,m,x}(\vartheta)| + \sup_{m \geq 0}\sup_{\vartheta\in \mathcal{N}^*}|\nabla^j\overline\ell_{k,m}(\vartheta)|\leq M_{j, k} \quad $\mathbb{P}_{\vartheta^*}$-a.s., \end{align*} where $r_1 = 6q_0$, $r_2 = 3q_0$, $r_3 = 2q_0$, $r_4 =3q_0/2$, $r_5 =6q_0/5$, and $r_6 =q_0$. (d) Uniformly in $\vartheta\in \mathcal{N}^*$ and $x \in \mathcal{X}$, $\nabla^j\ell_{k,m,x}(\vartheta)$ and $\nabla^j\overline\ell_{k,m}(\vartheta)$ converge $\mathbb{P}_{\vartheta^*}$-a.s.\ and in $L^{r_j}(\mathbb{P}_{\vartheta^*})$ to $\nabla^j \ell_{k,\infty}(\vartheta) \in L^{r_j}(\mathbb{P}_{\vartheta^*})$ as $m\rightarrow \infty$.

As shown by the following proposition, we may prove the uniform convergence of the derivatives of the density ratios by expressing them as polynomials of the derivatives of the log-density and applying Proposition (ref) and H\"older's inequality.

propositionUnder Assumptions (ref), (ref), and (ref), for $j=1,\ldots,6$, there exist random variables $\{K_{j,k}\}_{k=1}^n \in L^{q_\vartheta}(\mathbb{P}_{\vartheta^*})$ and $\rho_* \in (0,1)$ such that, for all $1 \leq k \leq n$ and $m' \geq m \geq 0$, \begin{align*} (a) &\quad \sup_{x\in\mathcal{X}}\sup_{\vartheta\in \mathcal{N}^*}|\nabla^j l_{k,m,x}(\vartheta) - \nabla^j \overline l_{k,m}(\vartheta) | \leq K_{j,k} (k+m)^{7}\rho_*^{k+m-1} \quad \mathbb{P}_{\vartheta^*}-a.s.,\\ (b) &\quad \sup_{x\in\mathcal{X}}\sup_{\vartheta\in \mathcal{N}^*}|\nabla^j l_{k,m,x}(\vartheta) - \nabla^j l_{k,m',x}(\vartheta) | \leq K_{j,k} (k+m)^{7}\rho_*^{k+m-1} \quad $\mathbb{P}_{\vartheta^*}$-a.s., \\ (c) &\quad \sup_{m\geq 0}\sup_{x\in\mathcal{X}}\sup_{\vartheta\in \mathcal{N}^*}|\nabla^j l_{k,m,x}(\vartheta)| + \sup_{m\geq 0}\sup_{\vartheta\in \mathcal{N}^*}|\nabla^j\overline{l}_{k,m}(\vartheta)| \leq K_{j,k} \quad \mathbb{P}_{\vartheta^*}-a.s. \end{align*} (d) Uniformly in $\vartheta\in \mathcal{N}^*$ and $x \in \mathcal{X}$, $\nabla^j l_{k,m,x}(\vartheta)$ and $\nabla^j\overline l_{k,m}(\vartheta)$ converge $\mathbb{P}_{\vartheta^*}$-a.s.\ and in $L^{q_\vartheta}(\mathbb{P}_{\vartheta^*})$ to $\nabla^j l_{k,\infty}(\vartheta) \in L^{q_\vartheta}(\mathbb{P}_{\vartheta^*})$ as $m\rightarrow \infty$. (e) $\sup_{\vartheta\in \mathcal{N}^*}|\nabla^j \overline l_{k,0}(\vartheta) - \nabla^j l_{k,\infty}(\vartheta) | \leq K_{j,k} k^{7}\rho_*^{k-1}$ $\mathbb{P}_{\vartheta^*}$-a.s.

When we apply the results in Section (ref) to regime switching models, $l_{k,0,x}(\vartheta)$ corresponds to $l_{\vartheta k x_0}$ on the left-hand side of ((ref)), and $s_{\pi k}$ in ((ref)) is a function of the $\nabla^j \overline l_{k, 0}(\vartheta)$'s. Proposition (ref) and the dominated convergence theorem for conditional expectations durrett10book imply that $\mathbb{E}_{\vartheta^*}[\nabla^j l_{k,\infty}(\vartheta)|\overline{\bf{Y}}_{-\infty}^{k-1}]=0$ for all $\vartheta\in \mathcal{N}^*$. Therefore, $\{\nabla^j l_{k,\infty}(\vartheta)\}_{k=-\infty}^{\infty}$ is a stationary, ergodic, and square integrable martingale difference sequence, and $\{\nabla^j l_{k,\infty}(\vartheta)\}_{j=1}^5$ satisfies Assumption (ref)(a)(b)(g).

Testing homogeneity

Before developing the LRT of $M_0$ components, we analyze a simpler case of testing the null hypothesis $H_0: M=1$ against $H_A:M=2$ when the data are from $H_0$. We assume that the parameter space for $\vartheta_{2,x}=(p_{11},p_{22})'$ takes the form $[\epsilon, 1-\epsilon]^2$ for a small $\epsilon \in (0,1/2)$. Denote the true parameter in the one-regime model by $\vartheta_1^*:= ((\theta^*)',(\gamma^*)')'$. The two-regime model gives rise to the true density $p_{\vartheta_1^*}({\bf Y}_1^n| \overline{{\bf Y}}_{0}, x_0)$ if the parameter $\vartheta_2=(\theta_1,\theta_2,\gamma,p_{11},p_{22})'$ lies in a subset of the parameter space \[ \Gamma^* := \left\{ (\theta_1,\theta_2,\gamma,p_{11},p_{22})\in \Theta_{ 2}: \theta_1=\theta_2=\theta^*\ \text{and}\ \gamma=\gamma^* \right\}. \] Note that $(p_{11},p_{22})$ is not identified under $H_0$.

Let $\ell_n(\vartheta_2,\xi_2) := \log \left( \sum_{x_0=1}^2 p_{\vartheta_2}({\bf Y}_1^n| \overline{\bf Y}_{0},{\bf W}_{0}^n,x_0) \xi_2(x_0) \right)$ denote the two-regime log-likelihood for a given initial distribution $\xi_2(x_0) \in \Xi_2$, and let $\hat\vartheta_2:= \arg\max_{\vartheta_2 \in \Theta_{2}} \ell_n(\vartheta_2,\xi_2)$ denote the MLE of $\vartheta_2$ given $\xi_2$, where $\xi_2$ is suppressed from $\hat \vartheta_2$ because $\xi_2$ does not matter asymptotically. Let $\hat \vartheta_1$ denote the one-regime MLE that maximizes the one-regime log-likelihood function $\ell_{0,n}(\vartheta_1) := \sum_{k=1}^n \log f (Y_k|\overline{\bf Y}_{k-1},W_k;\gamma,\theta)$ under the constraint $\vartheta_1 = (\theta',\gamma')' \in \Theta_1$.

We introduce the following assumption for the consistency of $\hat \vartheta_1$ and $\hat \vartheta_2$. Assumption (ref)(b) corresponds to Assumption (A4) of DMR. Assumption (ref)(c) is a standard identification condition for the one-regime model. Assumption (ref)(d) implies that the Kullback--Leibler divergence between $ p_{\vartheta_{1}^*}(Y_1 |\overline{\bf{Y}}_{-m}^0,{\bf W}_{-m}^0)$ and $ p_{\vartheta_{2}}(Y_1 |\overline{\bf{Y}}_{-m}^0,{\bf W}_{-m}^0)$ is 0 if and only if $\vartheta_2 \in \Gamma^*$.

assumption(a) $\Theta_1$ and $\Theta_2$ are compact, and $\vartheta_1^*$ is in the interior of $\Theta_1$. (b) For all $(x,x') \in \mathcal{X}$ and all $(\overline{{\bf y}},y',w)\in \mathcal{Y}^s\times \mathcal{Y}\times \mathcal{W}$, $f(y'|\overline{\bf{y}}_0,w;\gamma,\theta)$ is continuous in $(\gamma,\theta)$. (c) If $\vartheta_{1}\neq \vartheta_1^*$, then $\mathbb{P}_{\vartheta^*_{1}}\left(f(Y_1|\overline{\bf{Y}}_0,W_1;\gamma,\theta) \neq f(Y_1|\overline{\bf{Y}}_0,W_1;\gamma^*,\theta^*) \right)>0$. (d) $\mathbb{E}_{\vartheta^*_{1}}[ \log p_{\vartheta_{2}}(Y_1 |\overline{\bf{Y}}_{-m}^0,{\bf W}_{-m}^1)] = \mathbb{E}_{\vartheta^*_{1}}[\log p_{\vartheta_{1}^*}(Y_1 |\overline{\bf{Y}}_{-m}^0,{\bf W}_{-m}^1) ]$ for all $m \geq 0$ if and only if $\vartheta_{2}\in \Gamma^*$.

The following proposition shows the consistency of the MLEs of $\vartheta_1^*$ and $\vartheta_{2,y}^*$.

propositionSuppose that Assumptions (ref), (ref), and (ref) hold. Then, under the null hypothesis of $M=1$, $\hat\vartheta_1 \overset{p}{\rightarrow} \vartheta_1^*$ and $\inf_{\vartheta_2\in\Gamma^*}|\hat\vartheta_2-\vartheta_2|\overset{p}{\rightarrow} 0$.

Let $LR_n:=2[\ell_n(\hat\vartheta_2,\xi_2) - \ell_{0,n}(\hat\vartheta_1)]$ denote the LRTS for testing $H_0:M_0=1$ against $H_A:M_0=2$. Following the notation of Section (ref), we split $\vartheta_2$ into $\vartheta_2=(\psi,\pi)$, where $\pi$ is the part of $\vartheta$ not identified under the null hypothesis; the elements of $\psi$ are delineated later. In the current setting, $\pi$ corresponds to $\vartheta_{2,x}=(p_{11},p_{22})'$. Define $\varrho := \text{corr}_{\vartheta_{2,x}} (X_k,X_{k+1}) = p_{11}+p_{22}-1$ and $\alpha := \mathbb{P}_{\vartheta_{2,x}}(X_k=1)= (1-p_{22})/(2-p_{11}-p_{22})$. The parameter spaces for $\varrho$ and $\alpha$ under restriction $p_{11},p_{22} \in [\epsilon, 1-\epsilon]$ are given by $\Theta_{\varrho}:= [-1+2\epsilon,1-2\epsilon]$ and $\Theta_{\alpha}:= [\epsilon, 1-\epsilon]$, respectively. Because the mapping from $(p_{11},p_{22})$ to $(\varrho,\alpha)$ is one-to-one, we reparameterize $\pi$ as $\pi := (\varrho,\alpha)' \in \Theta_{\pi }:=\Theta_{\varrho }\times \Theta_{\alpha }$, and let $p_{\psi\pi}(\cdot|\cdot) := p_{\vartheta_2} (\cdot|\cdot)$. Henceforth, we suppress ${\bf W}_{0}^n$ for notational brevity and write, for example, $p_{\psi\pi}({\bf Y}_1^n| \overline{{\bf Y}}_{0}, {\bf W}_{0}^n,x_0)$ as $p_{\psi\pi}({\bf Y}_1^n|\overline{{\bf Y}}_{0},x_0)$ and $p_{\psi\pi} (y_k,x_k| \overline{\bf{y}}_{k-1},x_{k-1},w_k)$ as $p_{\psi\pi} (y_k,x_k| \overline{\bf{y}}_{k-1},x_{k-1})$ when doing so does not cause confusion.

We derive the asymptotic distribution of the LRTS by applying Corollary (ref) to $\ell_n(\psi,\pi,\xi_2)$ and representing $s_{\pi k}$ and $t_\vartheta$ in ((ref)) in terms of $\vartheta$, $f(Y_k|\overline{\bf{Y}}_{k-1};\gamma,\theta)$, and its derivatives; $s_{\pi k}$ involves higher-order derivatives, and $t_\vartheta$ consists of the functions of the polynomials of (reparameterized) $\vartheta$. Section (ref) analyzes the case when the regime-specific distribution of $y_k$ is not normal with unknown variance. Section (ref) analyzes the case when the regime-specific distribution $y_k$ is normal with regime-specific and unknown variance. Section (ref) handles the normal distribution where the variance is unknown and common across regimes. Note that because $\overline{\bf Y}_{-\infty}^\infty$ and ${\bf X}_{-\infty}^\infty$ are independent when $\psi = \psi^*$, we have

equation[equation omitted — 163 chars of source]

Define $q_k := \mathbb{I}\{X_k=1\}$, so that $\alpha = \mathbb{E}_{\psi^*\pi}[q_k]$.

Non-normal distribution

When we apply Corollary (ref) to regime switching models, $s_{\pi k}$ is a function of \\$\nabla^j\overline p_{\psi^*\pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})/\overline p_{\psi^*\pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})$'s with $ \overline p_{\psi\pi} (Y^k_{1}| \overline{{\bf Y}}_{0})$ defined in (\ref{p_bar}). In order to express $\nabla^j\overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1}) / \overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})$ in terms of $\nabla^j f(y|x;\gamma,\theta)$ via the Louis information principle (Lemma \ref{louis} in the appendix), we first derive the derivatives of the \emph{complete data} conditional density $p_{\vartheta_2} (y_k,x_k| \overline{\bf{y}}_{k-1},x_{k-1}) = g_{\vartheta_{2,y}}(y_k|\overline{\bf{y}}_{k-1}, x_k) q_{\vartheta_{2,x}}(x_{k-1},x_k)=\sum_{j =1}^2 \mathbb{I}\{x_k=j\}f(y_k|\overline{\bf{y}}_{k-1};\gamma,\theta_j)q_{\vartheta_{2,x}}(x_{k-1},x_k)$.

Consider the following reparameterization. Let

equation[equation omitted — 387 chars of source]

Let $\eta := (\gamma',\nu')'$ and $\psi_{\alpha}:=(\eta',\lambda')' \in \Theta_\eta \times \Theta_\lambda$. Under the null hypothesis of one regime, the true value of $\psi_{\alpha}$ is given by $\psi_{\alpha}^*:=(\gamma^*,\theta^*,0)'$. Henceforth, we suppress the subscript $\alpha$ from $\psi_{\alpha}$. Using this definition of $\psi$, let $\vartheta_2 := (\psi',\pi')' \in \Theta_\psi \times \Theta_\pi$. By using reparameterization ((ref)) and noting that $q_k = \mathbb{I}\{x_k=1\}$, we have $p_{\psi\pi} (y_k,x_k| \overline{\bf{y}}_{k-1},x_{k-1}) = g_{\psi}(y_k|\overline{\bf{y}}_{k-1}, x_k) q_{\pi}(x_{k-1},x_k)$ and

equation[equation omitted — 140 chars of source]

Henceforth, let $f^*_{k}$, $\nabla f^*_{k}$, $g^*_{k}$, and $\nabla g^*_{k}$ denote $f(Y_k|\overline{\bf{Y}}_{k-1};\gamma^*,\theta^*)$, $\nabla f(Y_k|\overline{\bf{Y}}_{k-1};\gamma^*,\theta^*)$, $g_{\psi^*}(Y_k|\overline{\bf{Y}}_{k-1},X_k) $, and $\nabla g_{\psi^*}(Y_k|\overline{\bf{Y}}_{k-1},X_k)$, respectively, and similarly for $\log f^*_{k}$ ,$\nabla \log f^*_{k}$, $\log g^*_{k}$, and $\nabla \log g^*_{k}$. Expanding $g_{\psi}(Y_k|\overline{\bf{Y}}_{k-1},X_k)$ twice with respect to $\psi=(\gamma',\nu',\lambda')'$ and evaluating at $\psi^*$ gives

equation[equation omitted — 352 chars of source]

Recall $\varrho := \text{corr}_{\vartheta_2^*}(q_k,q_{k+1})$. Observe that $q_k$ satisfies

equation[equation omitted — 364 chars of source]

where the first three results follow from the property of a Bernoulli random variable, and the last result follows from hamilton94book. Then, it follows from ((ref)) and ((ref)) that

equation[equation omitted — 294 chars of source]

From Lemma (ref), $\log p_{\psi\pi}(y_k, x_k|\overline{{\bf y}}_{k-1},x_{k-1}) = \log g_{\psi}(y_k|\overline{\bf{y}}_{k-1},x_k) + \log q_\pi(x_{k-1},x_k)$, and the definition of $\overline p_{\psi \pi} ( Y_1^k| \overline{{\bf Y}}_{0} )$ in ((ref)), we obtain \[

aligned&\frac{\nabla_\psi \overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})}{\overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})} =\nabla_\psi \log \overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1}) = \sum_{t=1}^k \mathbb{E}_{\vartheta^*} \left[\nabla_\psi \log g^*_{t} \middle|\overline{{\bf Y}}_{0}^k \right] - \sum_{t=1}^{k-1} \mathbb{E}_{\vartheta^*} \left[\nabla_\psi \log g^*_{t} \middle|\overline{{\bf Y}}_{0}^{k-1} \right].

\] Applying ((ref)), ((ref)), and $g_k^* =f_k^*$ to the right-hand side gives

equation[equation omitted — 395 chars of source]

Similarly, it follows from Lemma (ref), ((ref)), ((ref)), ((ref)), and $g_k^* =f_k^*$ that

align[align omitted — 1,449 chars of source]

Because the first-order derivative with respect to $\lambda$ is identically equal to zero in ((ref)), the unique elements of $\nabla_{\eta} \overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1}) /\overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1}) $ and $\nabla_{\lambda\lambda'} \overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})/\overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})$ constitute the generalized score $s_{\pi k}$ in Corollary (ref). This score is approximated by a stationary martingale difference sequence, where the approximation error satisfies Assumption (ref).

We collect some notations. Recall $\psi = (\eta',\lambda')'$ and $\eta = (\gamma',\nu')'$. For a $q \times 1$ vector $\lambda$ and a $q \times q$ matrix $s$, define the $q_\lambda \times 1$ vectors $v(\lambda)$ and $V(s)$ as

equation[equation omitted — 336 chars of source]

Noting that $\alpha(1-\alpha)>0$ for $\alpha\in\Theta_{\alpha}$, define, with $t_{\lambda}(\lambda,\pi):=\alpha(1-\alpha)v(\lambda)$,

equation[equation omitted — 505 chars of source]

and $s_{\lambda \varrho k} := V(s_{\lambda\lambda \varrho k})$ with

equation[equation omitted — 500 chars of source]

Here, $s_{\varrho k}$ in ((ref)) depends on $\varrho$ but not on $\alpha$ and corresponds to $s_{\pi k}$ in Corollary (ref). The following proposition shows that the log-likelihood function is approximated by a quadratic function of $\sqrt{n}t(\psi,\pi)$. Let $\mathcal{N}_{\varepsilon } := \{ \vartheta_2 \in \Theta_{2 }: |t(\psi,\pi)|< \varepsilon \}$. Let $A_{n \varepsilon }(\xi) := \{\vartheta \in \mathcal{N}_{\varepsilon}: \ell_n(\psi,\pi,\xi) - \ell_n(\psi^*,\pi,\xi) \geq 0\} $ and $A_{n\varepsilon c}(\xi) := A_{n\varepsilon }(\xi)\cup \mathcal{N}_{c/\sqrt{n}}$, where we suppress the subscript $2$ from $\xi_2$. We use this definition of $A_{n\varepsilon c}(\xi)$ through Sections (ref)--(ref). As shown in Sections (ref) and (ref), Assumption (ref) does not hold for regime switching models with a normal distribution.

assumption$0< \inf_{\varrho\in \Theta_{\varrho}} \lambda_{\min}(\mathcal{I}_{\varrho}) \leq \sup_{\varrho\in\Theta_{\varrho}} \lambda_{\max}(\mathcal{I}_\varrho) < \infty$ for $\mathcal{I}_{\varrho}= \lim_{k\rightarrow \infty} \mathbb{E}_{\vartheta^*}(s_{\varrho k}s_{\varrho k}')$, where $s_{\varrho k}$ is given in ((ref)).
propositionSuppose Assumptions (ref), (ref), (ref), (ref), and (ref) hold. Then, under the null hypothesis of $M=1$, (a) $\sup_{\xi}\sup_{\vartheta \in A_{n \varepsilon }(\xi)} |t(\psi,\pi)| = O_{p \varepsilon }(n^{-1/2})$; and (b) for any $c>0$, \begin{equation} \sup_{\xi \in \Xi}\sup_{\vartheta \in A_{n \varepsilon c}(\xi) } \left| \ell_n(\psi,\pi,\xi) - \ell_n(\psi^*,\pi,\xi) - \sqrt{n}t(\psi,\pi)' \nu_n (s_{\varrho k}) + n t(\psi,\pi)' \mathcal{I}_\varrho t(\psi,\pi)/2 \right| = o_{p \varepsilon} (1). \end{equation}

We proceed to derive the asymptotic distribution of the LRTS. With $s_{\varrho k}$ defined in ((ref)), define

equation[equation omitted — 862 chars of source]

where $G_{\lambda. \eta \varrho}$ is a $q_\lambda$-vector mean zero Gaussian process indexed by $\varrho$ with $\text{cov}(G_{\lambda. \eta\varrho_1},G_{\lambda. \eta\varrho_2}) = \mathcal{I}_{\lambda.\eta \varrho_1 \varrho_2}$. Define the set of admissible values of $\sqrt{n}\alpha(1-\alpha)v(\lambda)$ when $n\rightarrow \infty$ by $v(\mathbb{R}^q):= \{ x \in \mathbb{R}^{q_\lambda}: x = v(\lambda) \text{ for some } \lambda \in \mathbb{R}^q\}$. Define $\tilde t_{\lambda \varrho}$ by

equation[equation omitted — 302 chars of source]

The following proposition establishes the asymptotic null distribution of the LRTS.

propositionSuppose Assumptions (ref), (ref), (ref), (ref), and (ref) hold. Then, under the null hypothesis of $M=1$, $LR_n \overset{d}{\rightarrow} \sup_{ \varrho \in\Theta_{\varrho}} \left(\tilde t_{\lambda \varrho}' \mathcal{I}_{\lambda.\eta \varrho} \tilde t_{\lambda \varrho} \right)$.

In Proposition (ref), the LRTS and its asymptotic distribution depend on the choice of $\epsilon$ because $\Theta_\varrho=[-1+2\epsilon,1-2\epsilon]$. It is possible to develop a version of the EM test chenli09as, chenlifu12jasa, kasaharashimotsu15jasa in this context that does not impose an explicit restriction on the parameter space for $p_{11}$ and $p_{22}$; however, we leave such an extension to future research.

remarkWhen applied to the Markov regime switching model, the tests of carrasco14em use the residuals from projecting $\nabla_{\theta \theta'} f_k / f_k + 2 \sum_{t=1}^{k-1} \varrho^{k-t} (\nabla_{\theta} f_{t} / f_{t}) (\nabla_{\theta'} f_{k} / f_{k})$ on $\nabla_{\theta} f_k / f_k$, where both are evaluated at the one-regime MLE. Therefore, in the non-normal case, the LRT and tests of carrasco14em are based on the same score function.

Heteroscedastic normal distribution

Suppose that $Y_k\in \mathbb{R}$ in the $j$-th regime follows a normal distribution with regime-specific intercept $\mu_j$ and variance $\sigma_j^2$. We split $\theta_j$ into $\theta_j = (\zeta_j,\sigma_j^2)'= (\mu_j,\beta_j',\sigma_j^2)'$, and write the density of the $j$-th regime as

equation[equation omitted — 261 chars of source]

for some function $\varpi$. In many applications, $\varpi$ is a linear function of $\gamma$ and $\beta_j$, e.g., $\varpi(\overline{\bf{y}}_{k-1},w_k;\gamma,\beta_j)= (\overline{\bf{y}}_{k-1})'\beta_j + w_k'\gamma$. Consider the following reparameterization introduced in kasaharashimotsu15jasa ($\theta$ in Kasahara and Shimotsu corresponds to $\zeta$ here):

equation[equation omitted — 364 chars of source]

where $\nu_{\zeta}=(\nu_\mu,\nu_{\beta}')'$, $\lambda_{\zeta}=(\lambda_\mu,\lambda_{\beta}')'$, $C_1 := -(1/3)(1 + \alpha)$, and $C_2 := (1/3)(2 - \alpha)$, so that $C_1=C_2-1$. Collect the reparameterized parameters, except for $\alpha$, into one vector $\psi_{\alpha}$. As in Section (ref), we suppress the subscript $\alpha$ from $\psi_{\alpha}$. Let the reparameterized density be

equation[equation omitted — 245 chars of source]

Let $\psi := (\eta',\lambda')' \in \Theta_\psi = \Theta_\eta \times \Theta_\lambda$, where $\eta := (\gamma',\nu_\zeta',\nu_\sigma)'$ and $\lambda := (\lambda_\zeta',\lambda_\sigma)'$. Because the likelihood function of a normal mixture model is unbounded when $\sigma_j \rightarrow 0$ hartigan85book, we impose $\sigma_j \geq \epsilon_\sigma$ for a small $\epsilon_\sigma >0$ in $\Theta_\psi$. We proceed to derive the derivatives of $g_{\psi}(Y_k|\overline{\bf{Y}}_{k-1},X_k)$ evaluated at $\psi^*$. $\nabla_\psi g_k^*$, $\nabla_{\lambda\eta'}g_k^*$, and $\nabla_{\lambda\lambda'}g_k^*$ are the same as those given in ((ref)) except for $\nabla_{\lambda_\mu^2}g_k^*$ and that those with respect to $\lambda_\sigma^j$ are multiplied by $2^j$. The higher-order derivatives of $g_{\psi}(Y_k|\overline{\bf{Y}}_{k-1},X_k)$ with respect to $\lambda_\mu$ are derived by following kasaharashimotsu15jasa. From Lemma (ref) and the fact that normal density $f(\mu,\sigma^2)$ satisfies

equation[equation omitted — 351 chars of source]

we have

equation[equation omitted — 111 chars of source]

where

align*[align* omitted — 223 chars of source]

It follows from $\mathbb{E}_{\vartheta^*}[ q_k|\overline{{\bf Y}}_{-\infty}^n]=\alpha$, ((ref)), and elementary calculation that

equation[equation omitted — 648 chars of source]

with $b(\alpha): = -(2/3) (\alpha^2 - \alpha + 1) < 0$. Hence, $\mathbb{E}_{\vartheta^*}[ \nabla_{\lambda_\sigma^2} g_k^*|\overline{{\bf Y}}_{-\infty}^k]$ and $\mathbb{E}_{\vartheta^*}[ \nabla_{\lambda_\mu^4} g_k^*|\overline{{\bf Y}}_{-\infty}^k]$ are linearly dependent.

We proceed to derive a representation of $\nabla^j \overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})/\overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})$ in terms of $\nabla^j f_k^*$. Repeating the calculation leading to ((ref))--((ref)) and using ((ref)) gives the following. First, ((ref)) and ((ref)) still hold; second, the elements of $\nabla_{\lambda\lambda'} \overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})/\overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})$ except for the $(1,1)$-th element are given by ((ref)) after adjusting that the derivative with respect to $\lambda_{\sigma}$ must be multiplied by 2 (e.g., $\mathbb{E}_{\vartheta^*}[\nabla_{\lambda_\sigma} g_k^* |\overline{{\bf Y}}_{-\infty}^n] = 2 \nabla_{\sigma^2}f_k^*$ and $\mathbb{E}_{\vartheta^*}[\nabla_{\lambda_\sigma\lambda_\mu} g_k^* |\overline{{\bf Y}}_{-\infty}^n] = 2 \nabla_{\sigma^2\mu}f_k^*$); third,

equation[equation omitted — 328 chars of source]

When $\varrho \neq 0$, $\nabla_{\lambda_\mu^2} \overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})/\overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})$ is a non-degenerate random variable as in the non-normal case. When $\varrho=0$, however, $\nabla_{\lambda_\mu^2} \overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})/\overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})$ becomes identically equal to 0, and indeed the first non-zero derivative with respect to $\lambda_\mu$ is the fourth derivative.

Because of this degeneracy, we derive the asymptotic distribution of the LRTS by expanding $\ell_n(\psi,\pi,\xi)-\ell_n(\psi^*,\pi,\xi)$ four times. It is not correct, however, to simply approximate $\ell_n(\psi,\pi,\xi)-\ell_n(\psi^*,\pi,\xi)$ by a quadratic function of $\lambda_\mu^2$ (and other terms) when $\varrho \neq 0$ and a quadratic function of $\lambda_\mu^4$ when $\varrho=0$. This results in discontinuity at $\varrho=0$ and fails to provide a valid uniform approximation. We establish a uniform approximation by expanding $\ell_n(\psi,\pi,\xi)$ four times but expressing $\ell_n(\psi,\pi,\xi)$ in terms of $\varrho\lambda_\mu^2$, $\lambda_\mu^4$, and other terms.

For $m \geq 0$, define $\zeta_{k,m}(\varrho):= \sum_{t=-m+1}^{k-1} \varrho^{k-t-1} 2 \nabla_{\mu} f_{t}^*\nabla_{\mu} f_{k}^*/ f_{t}^* f_{k}^*$. Then, we can write ((ref)) as

equation[equation omitted — 352 chars of source]

Note that $\zeta_{k,m}(\varrho)$ satisfies $\mathbb{E}_{\vartheta^*}[\zeta_{k,m}(\varrho)| \overline{{\bf Y}}_{-m}^{k-1}]=0$ and is non-degenerate even when $\varrho=0$. Define $v(\lambda_\beta)$ as $v(\lambda)$ in ((ref)) but replacing $\lambda$ with $\lambda_\beta$. Collect the relevant parameters as

equation[equation omitted — 118 chars of source]

where

equation[equation omitted — 288 chars of source]

with $b(\alpha)= -(2/3)(\alpha^2-\alpha+1)<0$. Recall $\theta_j = (\zeta_j',\sigma_j^2)'= (\mu_j,\beta_j',\sigma_j^2)'$. Similarly to ((ref)), define the elements of the generalized score by

equation[equation omitted — 629 chars of source]

Define the generalized score as

equation[equation omitted — 549 chars of source]

The following proposition establishes a uniform approximation of the log-likelihood ratio.

assumption(a) $0< \inf_{\varrho\in\Theta_{\varrho}} \lambda_{\min}(\mathcal{I}_\varrho) \leq \sup_{\varrho\in\Theta_{\varrho}}\lambda_{\max}(\mathcal{I}_\varrho) < \infty$ for $\mathcal{I}_{\varrho}= \lim_{k\rightarrow \infty } \mathbb{E}_{\vartheta^*}(s_{\varrho k}s_{\varrho k}')$, where $s_{\varrho k}$ is given in ((ref)). (b) $\sigma_1^*,\sigma_2^* >\epsilon_\sigma$.
propositionSuppose Assumptions (ref), (ref), (ref), (ref), and (ref) hold and the density of the $j$-th regime is given by ((ref)). Then, under the null hypothesis of $M=1$, (a) $\sup_{\vartheta \in A_{n \varepsilon}(\xi)} |t(\psi,\pi)| = O_{p \varepsilon }(n^{-1/2})$; and (b) for any $c>0$, \begin{equation} \sup_{\xi \in \Xi} \sup_{\vartheta \in A_{n\varepsilon c}(\xi) } \left| \ell_n(\psi,\pi,\xi) - \ell_n(\psi^*,\pi,\xi) - \sqrt{n}t(\psi,\pi)' \nu_n (s_{\varrho k}) + n t(\psi,\pi)' \mathcal{I}_\varrho t(\psi,\pi)/2 \right| = o_{p \varepsilon }(1). \end{equation}

The asymptotic null distribution of the LRTS is characterized by the supremum of $ 2 t_{\lambda}' G_{\lambda.\eta\varrho}- t_{\lambda}' \mathcal{I}_{\lambda.\eta \varrho} t_{\lambda}$, where $G_{\lambda. \eta \varrho}$ and $\mathcal{I}_{\lambda.\eta \varrho}$ are defined analogously to those in ((ref)) but with $s_{\varrho k}$ defined in ((ref)), and the supremum is taken with respect to $t_{\lambda}$ and $\varrho\in\Theta_{\varrho}$ under the constraint implied by the limit of the set of possible values of $\sqrt{n}t_{\lambda}(\lambda,\pi)$ as $n\rightarrow\infty$. This constraint is given by the union of $\Lambda_{\lambda}^1$ and $\Lambda_{\lambda \varrho}^2$, where $q_\beta := \dim(\beta)$, $q_{\lambda}:= 3 +2q_\beta+q_\beta(q_\beta+1)/2$, and

equation[equation omitted — 936 chars of source]

Note that $\Lambda_{\lambda\varrho}^2$ depends on $\varrho$, whereas $\Lambda_{\lambda}^1$ does not depend on $\varrho$. Heuristically, $\Lambda_{\lambda}^1$ and $\Lambda_{\lambda\varrho}^2$ correspond to the limits of the set of possible values of $\sqrt{n}t_{\lambda}(\lambda,\pi)$ when $\liminf_{n\to\infty}n^{1/8}|\lambda_{\mu}| > 0$ and $\lambda_{\mu}=o(n^{-1/8})$, respectively. When $\liminf_{n\to\infty}n^{1/8}|\lambda_{\mu}| > 0$, we have $(\hat\lambda_\sigma, \hat\lambda_\beta) =O_p(n^{-3/8})$ because $t_{\lambda}(\hat\lambda,\pi)=O_p(n^{-1/2})$. Further, the set of possible values of $\sqrt{n} \varrho\lambda_{\mu}^2$ converges to $\mathbb{R}$ because $\varrho$ can be arbitrarily small. Consequently, the limit of $\sqrt{n}t_{\lambda}(\lambda,\pi)$ is characterized by $\Lambda_{\lambda}^1$.

Define $Z_{\lambda \varrho}$ and $\mathcal{I}_{\lambda.\eta \varrho}$ as in ((ref)) but with $s_{\pi k}$ defined in ((ref)). Let $Z_{\lambda 0}$ and $\mathcal{I}_{\lambda.\eta 0}$ denote $Z_{\lambda \varrho}$ and $\mathcal{I}_{\lambda.\eta \varrho}$ evaluated at $\varrho=0$. Define $\tilde{t}_{\lambda}^1$ and $\tilde{t}_{\lambda\varrho}^2$ by

equation[equation omitted — 587 chars of source]

The following proposition establishes the asymptotic null distribution of the LRTS.

propositionSuppose that the assumptions in Proposition (ref) hold. Then, under the null hypothesis of $M=1$, $LR_n \overset{d}{\rightarrow} \max\{ \mathbb{I}\{\varrho=0\} (\tilde t_{\lambda }^1)' \mathcal{I}_{\lambda.\eta 0} \tilde t_{\lambda }^1, \sup_{\varrho\in\Theta_{\varrho}} (\tilde t_{\lambda \varrho}^2)' \mathcal{I}_{\lambda.\eta \varrho} \tilde t_{\lambda \varrho}^2 \}$.
remarkquzhuo17wp derive the asymptotic distribution of the LRTS under the restriction that $\varrho\geq \epsilon>0$.
remarkIt is possible to extend our analysis to the exponential-LR type tests studied by andrewsploberger94em and carrasco14em.

Homoscedastic normal distribution

Suppose that $Y_k\in \mathbb{R}$ in the $j$-th regime follows a normal distribution with the regime-specific intercept $\mu_j$ but with common variance $\sigma^2$. We split $\gamma$ and $\theta_j$ into $\gamma=(\tilde \gamma',\sigma^2)'$ and $\theta_j=(\mu_j,\beta_j')'$, and write the density of the $j$-th regime as

equation[equation omitted — 273 chars of source]

for some function $\varpi$. Consider the following reparameterization:

equation[equation omitted — 267 chars of source]

where $\nu_{\theta}=(\nu_\mu,\nu_\beta')'$ and $\lambda=(\lambda_\mu,\lambda_\beta')'$. Collect the reparameterized parameters, except for $\alpha$, into one vector $\psi_{\alpha}$. Suppressing $\alpha$ from $\psi_{\alpha}$, let the reparameterized density be

equation[equation omitted — 212 chars of source]

Let $\eta = (\tilde\gamma',\nu_\theta',\nu_\sigma)'$; then, the first and second derivatives of $g_{\psi}(y_k|\overline{\bf{y}}_{k-1},x_k)$ with respect to $\eta$ and $\lambda$ are the same as those given in ((ref)) except for $\nabla_{\lambda_\mu^2}g_{\psi}(y_k|\overline{\bf{y}}_{k-1},x_k)$. We derive the higher-order derivatives of $g_{\psi}(y_k|\overline{\bf{y}}_{k-1},x_k)$ with respect to $\lambda_\mu$. From Lemmas (ref) and ((ref)), we obtain

equation[equation omitted — 262 chars of source]

where $d_{0k} :=1$, $d_{1k} := q_k - \alpha$, $d_{2k} := (q_k-\alpha)^2-\alpha(1-\alpha)$, $d_{3k} := (q_k-\alpha)^3 - 3(q_k-\alpha)\alpha(1-\alpha)$, and $d_{4k} := (q_k - \alpha)^4 -6 (q_k - \alpha)^2 \alpha(1-\alpha) + 3\alpha^2(1 - \alpha)^2$. It follows from $\mathbb{E}_{\vartheta^*}[ q_k|\overline{{\bf Y}}_{-\infty}^n]=\alpha$, ((ref)), and elementary calculation that

equation[equation omitted — 440 chars of source]

Repeating the calculation leading to ((ref))--((ref)) and using ((ref)) gives the following. First, ((ref)) and ((ref)) still hold; second, the elements of $\nabla_{\lambda\lambda'} \overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})/\overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})$ are given by ((ref)) except for the $(1,1)$-th element; third, $\nabla_{\lambda_\mu^2} \overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})/\overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})$ is given by ((ref)). Further, Lemma (ref) in the appendix shows that when $\varrho=0$, $\nabla_{\lambda_\mu^3} \overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})/\overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1}) = \alpha(1-\alpha)(1-2\alpha) \nabla_{\mu^3}f_k^*/f_k^*$ and $\nabla_{\lambda_\mu^4} \overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})/\overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1}) = \alpha(1-\alpha)(1-6\alpha+6\alpha^2) \nabla_{\mu^4}f_k^*/f_k^*$. Because $\nabla_{\lambda_\mu^3} \overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})/\overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})=0$ when $\alpha=1/2$ and $\varrho=0$, we expand $\ell_n(\psi,\pi,\xi)$ four times and express it in terms of $\varrho\lambda_\mu^2$, $(1-2\alpha)\lambda_\mu^3$, $\lambda_\mu^4$, and other terms to establish a uniform approximation.

Collect the relevant parameters as

equation[equation omitted — 361 chars of source]

Define the generalized score as

equation[equation omitted — 483 chars of source]

where $\zeta_{k,m}(\varrho)$ is defined as in ((ref)), $s_{\lambda_{\mu}^i k}:=\nabla_{\mu^i}f_k^*/f_k^*$ for $i=3,4$, and $s_{\lambda_{\beta \mu} \varrho k}$ and $s_{\lambda_{\beta\beta} \varrho k}$ are defined as in ((ref)) but using the density ((ref)) in place of ((ref)). Define, with $q_\beta := \dim(\beta)$ and $q_{\lambda}:=3 + q_\beta+q_\beta(q_\beta+1)/2$,

equation[equation omitted — 750 chars of source]

The following two propositions correspond to Propositions (ref) and (ref), establishing a uniform approximation of the log-likelihood ratio and asymptotic distribution of the LRTS.

assumption$0< \inf_{\varrho \in \Theta_{\varrho}} \lambda_{\min}(\mathcal{I}_\varrho) \leq \sup_{\varrho \in \Theta_{\varrho}}\lambda_{\max}(\mathcal{I}_\varrho) < \infty$ for $\mathcal{I}_{\varrho}= \lim_{k\rightarrow \infty} \mathbb{E}_{\vartheta^*}(s_{\varrho k}s_{\varrho k}')$, where $s_{\varrho k}$ is given in ((ref)).
propositionSuppose Assumptions (ref), (ref), (ref), (ref), and (ref) hold and the density of the $j$-th regime is given by ((ref)). Then, statements (a) and (b) of Proposition (ref) hold.
propositionSuppose that the assumptions in Proposition (ref) hold. Then, under the null hypothesis of $M=1$, $LR_n \overset{d}{\rightarrow} \max\{ \mathbb{I}\{\varrho=0\} (\tilde t_{\lambda }^1)' \mathcal{I}_{\lambda.\eta 0} \tilde t_{\lambda }^1, \sup_{\varrho\in\Theta_\varrho} (\tilde t_{\lambda \varrho}^2)' \mathcal{I}_{\lambda.\eta \varrho} \tilde t_{\lambda \varrho}^2 \}$, where $\tilde{t}_{\lambda}^1$ and $\tilde{t}_{\lambda\varrho}^2$ are defined as in ((ref)) but in terms of $(Z_{\lambda \varrho}, \mathcal{I}_{\lambda.\eta \varrho}, Z_{\lambda 0},\mathcal{I}_{\lambda.\eta 0})$ constructed with $s_{\varrho k}$ defined in ((ref)) and $\Lambda_{\lambda }^1$ and $\Lambda_{\lambda \varrho}^2$ defined in ((ref)).

Testing $H_0:M=M_0$ against $H_A:M=M_0+1$ for $M_0 \geq 2$

In this section, we derive the asymptotic distribution of the LRTS for testing the null hypothesis of $M_0$ regimes against the alternative of $M_0+1$ regimes for general $M_0 \geq 2$. We suppress the covariate ${\bf W}_{a}^b$ unless confusion might arise.

Let $\vartheta_{M_0}^*=((\vartheta_{M_0,x}^*)',(\vartheta_{M_0,y}^*)')'$ denote the parameter of the $M_0$-regime model, where $\vartheta_{M_0,x}^*$ contains $p_{ij}^*= q_{\vartheta^*_{M_0,x}}(i,j)>0$ for $i= 1,\ldots,M_0$ and $j=1,\ldots,M_0-1$, and $\vartheta_{M_0,y}^* = ((\theta_1^*)',\ldots,(\theta_{M_0}^*)',(\gamma^*)')'$. We assume $\max_{i}\sum_{j=1}^{M_0-1}p_{ij}^*<1$ and $\theta_{1}^*<\ldots< \theta_{M_0}^*$ for identification. The true $M_0$-regime conditional density of ${\bf Y}_1^n$ given $\overline{\bf Y}_{0}$ and $x_0$ is

equation[equation omitted — 215 chars of source]

where $p_{\vartheta_{M_0}^*} (y_k,x_k| \overline{\bf{y}}_{k-1},x_{k-1}) = g_{\vartheta_{M_0,y}^*}(y_k|\overline{\bf{y}}_{k-1}, x_k) q_{\vartheta_{M_0,x}^*}(x_{k-1},x_k)$ with $g_{\vartheta_{M_0,y}^*}(y_k|\overline{\bf{y}}_{k-1},x_k) = \sum_{j=1,\ldots,M_0} \mathbb{I}\{x_k=j\}f(y_k|\overline{\bf{y}}_{k-1};\gamma,\theta_j^*)$.

Let the conditional density of ${\bf Y}_{1}^n$ of an $(M_0+1)$-regime model be

equation[equation omitted — 220 chars of source]

where $p_{\vartheta_{M_0+1}}(y_k,x_k|\overline{\bf y}_{k-1},x_{k-1})$ is defined similarly to $p_{\vartheta_{M_0}^*}(y_k,x_k|\overline{\bf y}_{k-1},x_{k-1})$ with $\vartheta_{M_0+1,x}:= \{p_{ij}\}_{i=1,\ldots,M_0+1,j=1,\ldots,M_0}$ and $\vartheta_{M_0+1,y} := (\theta_1',\ldots,\theta_{M_0+1}',\gamma')'$. We assume that $\min_{i,j}p_{ij} \geq \epsilon$ for some $\epsilon\in (0,1/2)$.

Write the null hypothesis as $H_0 = \cup_{m=1}^{M_0} H_{0m}$ with \[ H_{0m} : \theta_1 < \cdots < \theta_{m} = \theta_{m + 1} < \cdots < \theta_{M_0 + 1}. \] Define the set of values of $\vartheta_{M_0+1}$ that yields the true density ((ref)) under $\mathbb{P}_{\vartheta^*_{M_0}}$ as $\Upsilon^*:=\{\vartheta_{M_0 + 1} \in \Theta_{M_0+1,\epsilon}: p_{\vartheta_{M_0+1}}({\bf Y}_1^n| \overline{\bf Y}_{0}, x_0) = p_{\vartheta_{M_0}^*}({\bf Y}_1^n| \overline{\bf Y}_{0}, x_0)\ \mathbb{P}_{\vartheta^*_{M_0}}\text{-a.s.}\}$. Under $H_{0m}$, the $(M_0 + 1)$-regime model ((ref)) generates the true $M_0$-regime density ((ref)) if $\theta_m = \theta_{m + 1} = \theta_{m}^{*}$ and the transition matrix of $X_k$ reduces to that of the true $M_0$-regime model.

We reparameterize the transition probability of $X_k$ by writing $\vartheta_{M_0+1,x}$ as $\vartheta_{M_0+1,x} = (\vartheta_{xm}',\pi_{xm}')'$, where $\vartheta_{xm}$ is point identified under $H_{0m}$, while $\pi_{xm}$ is not point identified under $H_{0m}$. The transition probability of $X_k$ under $\vartheta_{M_0+1,x}$ equals the transition probability of $X_k$ under $\vartheta_{M_0,x}^*$ if and only if $\vartheta_{xm} = \vartheta_{xm}^*$. The detailed derivation including the definition of $\vartheta_{xm}^*$ is provided in Section (ref) in the appendix. Define the subset of $\Upsilon^*$ that corresponds to $H_{0m}$ as

align*[align* omitted — 324 chars of source]

then, $\Upsilon^*= \Upsilon_{1}^* \cup \cdots \cup \Upsilon_{M_0}^*$ holds.

For $M=M_0,M_0+1$, let $\ell_n(\vartheta_M,\xi_M) := \log \left( \sum_{x_0=1}^M p_{\vartheta_M}({\bf Y}_1^n| \overline{\bf Y}_{0},x_0) \xi_M(x_0) \right)$ denote the $M$-regime log-likelihood for a given initial distribution $\xi_M(x_0) \in \Xi_M$. We treat $\xi_M(x_0)$ as fixed. Let $\hat\vartheta_{M_0}:= \arg\max_{\vartheta_{M_0} \in \Theta_{{M_0}}} \ell_n(\vartheta_{M_0},\xi_{M_0})$ and $\hat\vartheta_{M_0+1}:= \arg\max_{\vartheta_{M_0+1} \in \Theta_{M_0+1}} \ell_n(\vartheta_{M_0+1},\xi_{M_0+1})$. The following proposition shows that the MLE is consistent in the sense that the distance between $\hat{\vartheta}_{M_0+1}$ and $\Upsilon^*$ tends to 0 in probability. The proof of Proposition (ref) is essentially the same as the proof of Proposition (ref) and hence is omitted.

assumption(a) $\Theta_{M_0}$ and $\Theta_{M_0+1}$ are compact, and $\vartheta_{M_0}^*$ is in the interior of $\Theta_{M_0}$. (b) For all $(x,x') \in \mathcal{X}$ and all $(\overline{{\bf y}},y',w)\in \mathcal{Y}^s\times \mathcal{Y}\times \mathcal{W}$, $f(y'|\overline{\bf{y}}_0,w;\gamma,\theta)$ is continuous in $(\gamma,\theta)$. (c) $\mathbb{E}_{\vartheta^*_{M_0}}[ \log ( p_{\vartheta_{M_0}}(Y_1 |\overline{\bf{Y}}_{-m}^0,{\bf W}_{-m}^1)] = \mathbb{E}_{\vartheta^*_{M_0}}[ \log p_{\vartheta_{M_0}^*}(Y_1 |\overline{\bf{Y}}_{-m}^0,{\bf W}_{-m}^1) ]$ for all $m \geq 0$ if and only if $\vartheta_{M_0} = \vartheta_{M_0}^*$. (d) $\mathbb{E}_{\vartheta^*_{M_0}}[ \log ( p_{\vartheta_{M_0+1}}(Y_1 |\overline{\bf{Y}}_{-m}^0,{\bf W}_{-m}^0) ] = \mathbb{E}_{\vartheta^*_{M_0}} [\log p_{\vartheta_{M_0}^*}(Y_1 |\overline{\bf{Y}}_{-m}^0,{\bf W}_{-m}^1) ]$ for all $m \geq 0$ if and only if $\vartheta_{M_0+1}\in \Upsilon^*$.
propositionSuppose Assumptions (ref), (ref), and (ref) hold. Then, under the null hypothesis of $M=M_0$, $\hat\vartheta_{M_0} \overset{p}{\rightarrow} \vartheta_{M_0}^*$ and $\inf_{\vartheta_{M_0+1} \in \Upsilon^*} |\hat{\vartheta}_{M_0+1}-\vartheta_{M_0+1}|\overset{p}{\rightarrow} 0$.

Let $LR_{M_0,n}:=2[\ell_n(\hat\vartheta_{M_0+1},\xi_{M_0+1})-\ell_n(\hat\vartheta_{M_0},\xi_{M_0})]$ denote the LRTS for testing $H_0:M=M_0$ against $H_A:M=M_0+1$. We proceed to derive the asymptotic distribution of the LRTS by analyzing the behavior of the LRTS when $\vartheta_{M_0+1} \in \Upsilon^*_m$ for each $m$. Define $J_m :=\{m,m+1\}$. Observe that if ${\bf X}_1^k \in J_m^k$, then ${\bf X}_1^k$ follows a two-state Markov chain on $J_m$ whose transition probability is characterized by $\alpha_m := \mathbb{P}_{\vartheta_{M_0+1}}(X_k=m|X_k \in J_m)$ and $\varrho_m := \text{corr}_{\vartheta_{M_0+1}}(X_{k-1},X_{k}|(X_{k-1},X_{k}) \in J_m^2)$. Collect reparameterized $\pi_{xm}$ into $\pi_{xm} := (\varrho_m,\alpha_m, \phi_{m}')'$, where $\phi_m$ does not affect the transition probability of ${\bf X}_1^k$ when ${\bf X}_1^k \in J_m^{k}$. See Section (ref) in the appendix for the detailed derivation.

Define $q_{kj} := \mathbb{I}\{X_k=j\}$; then, we can write $\alpha_m$ and $\varrho_m$ as $\alpha_m = \mathbb{E}_{\vartheta_{M_0+1}}(q_{km}|X_k \in J_m)$ and $\varrho_m =\text{corr}_{\vartheta_{M_0+1}}(q_{k-1,m},q_{km}|(X_{k-1},X_{k}) \in J_m^2)$. Because $\overline{{\bf Y}}_{-\infty}^\infty$ provides no information for distinguishing between $X_k=m$ and $X_k={m+1}$ if $\theta_m = \theta_{m+1}$, we can write $\alpha_m$ and $\varrho_m$ as

equation[equation omitted — 280 chars of source]

Non-normal distribution

For non-normal component distributions, consider the following reparameterization similar to ((ref)):

equation*[equation* omitted — 166 chars of source]

Collect the reparameterized identified parameters into one vector $\psi_m: = (\eta_m',\lambda_m')'$, where $\eta_m=(\gamma',\{\theta_j'\}_{j=1}^{m-1}, \nu_m',\{\theta_j'\}_{j=m+2}^{M_0+1},\vartheta_{xm}')'$, so that the reparameterized $(M_0+1)$-regime log-likelihood function is $\ell_n(\psi_m,\pi_{xm},\xi_{M_0+1})$. Let $\psi_m^{*}=(\eta_m^*,\lambda_m^*)=((\vartheta_{M_0}^*)',0')'$ denote the value of $\psi_{m}$ under $H_{0m}$. Define the reparameterized conditional density of $y_k$ as

align*[align* omitted — 249 chars of source]

where $\overline J_m := \{1,\ldots,M_0+1\}\setminus J_m$. Let $f_{mk}^{*}$ denote $f(Y_k|\overline{\bf{Y}}_{k-1};\gamma^*,\theta_m^*)$. It follows from ((ref)) and the law of iterated expectations that

equation[equation omitted — 1,321 chars of source]

where the second equality holds because $g_{\psi_m^*}(Y_k|\overline{\bf{Y}}_{k-1}, X_k)=f_{mk}^*$ if $X_k \in J_m$, and the last equality holds because, conditional on $\{{\bf X}_{t_1}^{t_2} \in J_m^{t_2-t_1+1},\overline{\bf Y}_{-\infty}^n\} $, ${\bf X}_{t_1}^{t_2}$ is a two-state stationary Markov process with parameter $(\alpha_m,\varrho_m)$.

Let $g_{0k}^*$, $q_{0k}^*$, and $\overline p_{0k}^*$ denote $g_{\vartheta_{M_0,y}^*}(Y_k,X_k| \overline{{\bf Y}}_{k-1},X_{k-1})$, $q_{\vartheta_{M_0,x}^*}(X_{k-1},X_k)$, and $\overline p_{\vartheta_{M_0}^*}(Y_k| \overline{{\bf Y}}_{0}^{k-1})$. Let $\nabla g_{0k}^*$ denote the derivative of $g_{\vartheta_{M_0,y}}(Y_k,X_k| \overline{{\bf Y}}_{k-1},X_{k-1})$ evaluated at $\vartheta_{M_0,y}^*$, and define $\nabla q_{0k}^*$ and $\nabla \overline p_{0k}^*$ similarly. Repeating a derivation similar to ((ref))--((ref)) but using ((ref)) in place of ((ref)), we obtain

equation[equation omitted — 658 chars of source]
align[align omitted — 1,037 chars of source]

Define $\tilde \varrho := (\varrho_1,\ldots,\varrho_{M_0})'$, define $t_{\lambda}(\lambda_m,\pi_m)$ as $t_{\lambda}(\lambda,\pi)$ in ((ref)) by replacing $(\lambda,\pi)$ with $(\lambda_m,\pi_m)$, and let

equation[equation omitted — 592 chars of source]

and $s_{\lambda \varrho_m k}^m := V(s_{\lambda\lambda \varrho_m k}^m)$, where $s_{\lambda\lambda \varrho_m k}^m $ is defined similarly to ((ref)) as

equation[equation omitted — 552 chars of source]

Similarly to ((ref)), define

equation[equation omitted — 1,185 chars of source]

where ${G}_{\lambda.\eta \tilde \varrho}=((G_{\lambda.\eta \varrho_1}^1)',\ldots,(G_{\lambda.\eta\varrho_{M_0}}^{M_0})')'$ is an $M_0 q_\lambda$-vector mean zero Gaussian process with \\$cov({G}_{\lambda.\eta \tilde\varrho_1},{G}_{\lambda.\eta \tilde \varrho_2}) = \tilde{\mathcal{I}}_{\lambda.\eta \tilde\varrho_1\tilde\varrho_2}$. Note that $G_{\lambda.\eta \tilde \varrho}$ corresponds to the residuals from projecting $\tilde s_{\lambda \tilde \varrho k}$ on $\tilde s_{\eta k}$. Define $\tilde t_{\lambda \varrho_m}^m$ by

equation*[equation* omitted — 322 chars of source]

The following proposition gives the asymptotic null distribution of the LRTS. Under the stated assumptions, the log-likelihood function permits a quadratic approximation in the neighborhood of $\Upsilon_{m}^*$ similar to the one in Proposition (ref). Define $A_{n\varepsilon c}^m(\xi) := \{\vartheta_{M_0+1} \in \Theta_{M_0+1} : \{ \ell_n(\psi_m,\pi_m,\xi) - \ell_n(\psi_m^*,\pi_m,\xi) \geq 0 \} \land |t(\psi_m,\pi_m)|< \varepsilon \} \cup \mathcal{N}_{c/\sqrt{n}}$. Under $H_0: M=M_0$, for any $c>0$, for $m=1,\ldots,M_0$, and uniformly in $\xi \in \Xi$ and $\vartheta_{M_0+1} \in A_{n\varepsilon c}^m(\xi)$, \[ \ell_n(\psi_m,\pi_m,\xi) - \ell_n(\psi_m^*,\pi_m,\xi) \ - \sqrt{n}t(\psi_m,\pi_m)' \nu_n (s_{\varrho_m k}) + n t(\psi_m,\pi_m)' \mathcal{I}_{\varrho_m} t(\psi_m,\pi_m)/2 = o_{p \varepsilon} (1), \] where $s_{ \varrho_m k} := (\tilde{s}_{\eta k}',({s}_{\lambda \varrho_m k}^m)')'$ and $\mathcal{I}_{\varrho_m} = \lim_{k\rightarrow\infty} \mathbb{E}_{\vartheta^*_{M_0}}(s_{ \varrho_m k}s_{ \varrho_m k}')$. Consequently, the LRTS is asymptotically distributed as the maximum of the $M_0$ random variables, each of which represents the asymptotic distribution of the LRTS that tests $H_{0m}$. Denote the parameter space for $\varrho_m$ by $\Theta_{\varrho_m}$, and let $\tilde\Theta_{\varrho}:=\Theta_{\varrho_1}\times\ldots\times\Theta_{\varrho_{M_0}}$.

assumption$0< \inf_{\tilde \varrho \in \tilde\Theta_{\varrho}} \lambda_{\min}(\tilde{\mathcal{I}}_{\tilde\varrho}) \leq \sup_{\tilde \varrho \in \tilde\Theta_{\varrho}} \lambda_{\max}(\tilde{\mathcal{I}}_{\tilde\varrho}) < \infty$ for $\tilde{\mathcal{I}}_{\tilde\varrho}:= \lim_{k\rightarrow\infty}\mathbb{E}_{\vartheta^*_{M_0}}(\tilde{s}_{\tilde\varrho k}\tilde{s}_{\tilde\varrho k}')$, where $\tilde{s}_{\tilde\varrho k}$ is given in ((ref)).
propositionSuppose Assumptions (ref), (ref), (ref), (ref), and (ref) hold. Then, under $H_0: M=M_0$, $LR_{M_0,n}\overset{d}{\rightarrow} \max_{m=1,\ldots,M_0}\left\{ \sup_{\varrho_m \in \Theta_{\varrho}^m} \left( (\tilde t_{\lambda \varrho_m}^m)' \tilde{\mathcal{I}}_{\lambda.\eta \varrho_m}^m \tilde t_{\lambda \varrho_m}^m \right) \right\} $.

Heteroscedastic normal distribution

As in Section (ref), we assume that $Y_k \in \mathbb{R}$ in the $j$-th regime follows a normal distribution with the regime-specific intercept and variance of which density is given by ((ref)). Consider the following reparameterization similar to ((ref)):

equation*[equation* omitted — 397 chars of source]

where $\nu_{\zeta m}=(\nu_\mu,\nu_{\beta}')'$, $\lambda_{\zeta m}=(\lambda_{\mu m},\lambda_{\beta m}')'$, $C_1 := -(1/3)(1 + \alpha_m)$, and $C_2 := (1/3)(2 - \alpha_m)$. As in Section (ref), we collect the reparameterized identified parameters into $\psi_m: = (\eta_m',\lambda_m')'$, where $\eta_m=(\gamma',\{\theta_j'\}_{j=1}^{m-1}, \nu_{\zeta m}', \nu_{\sigma m},\{\theta_j'\}_{j=m+2}^{M_0+1},\vartheta_{xm}')'$ and $\lambda_m:=(\lambda_{\zeta m}',\lambda_{\sigma m})'$. Similar to ((ref)), define the reparameterized conditional density of $y_k$ as

equation*[equation* omitted — 406 chars of source]

Let $g_{mk}^*$, $f_{mk}^*$, $\nabla g_{mk}^*$, and $\nabla f_{mk}^*$ denote $g_{\psi_m^{*}}(Y_k|\overline{\bf{Y}}_{k-1}, X_k)$, $f(Y_k|\overline{\bf Y}_{k-1};\gamma^*,\theta_m^*)$, $\nabla g_{\psi_m^{*}}(Y_k|\overline{\bf{Y}}_{k-1}, X_k)$, and $\nabla f(Y_k|\overline{\bf Y}_{k-1};\gamma^*,\theta_m^*)$. From ((ref)) and a derivation similar to ((ref)), we obtain the following result that corresponds to ((ref)) in testing homogeneity:

equation[equation omitted — 660 chars of source]

Repeating the calculation leading to ((ref))--((ref)) and using ((ref)) gives the following. First, ((ref)) and ((ref)) still hold; second, the elements of $\nabla_{\lambda_m\lambda_m'} \overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})/\overline p_{\psi^* \pi} (Y_k| \overline{{\bf Y}}_{0}^{k-1})$ except for the $(1,1)$-th element are given by ((ref)) while adjusting the derivative with respect to $\lambda_{\sigma m}$ by multiplying by 2; third,

equation*[equation* omitted — 403 chars of source]

For $m \geq 0$, define $\zeta_{k,m}^m(\varrho_m):= \sum_{t=-m+1}^{k-1} \varrho_m^{k-t-1} 2 (\nabla_{\mu} f_{mt}^* \nabla_{\mu} f_{mk}^* / f_{mt}^* f_{mk}^*) \mathbb{P}_{\vartheta^*_{M_0}}((X_t,X_k) \in J_m^2|\overline{{\bf Y}}_{0}^k)$. Similarly to ((ref)), define the elements of the generalized score as

equation[equation omitted — 866 chars of source]

Similarly to ((ref)), define $\tilde{s}_{\tilde \varrho k}$ as in ((ref)) by redefining $s_{\lambda \varrho_m k}^m$ in ((ref)) as

equation[equation omitted — 349 chars of source]

Define ${\mathcal{I}}_{\lambda.\eta \varrho_m}^m$ and $Z^m_{\lambda \varrho_m}$ as in ((ref)) with $s_{\lambda \varrho_m k}^m$ defined in ((ref)). Let $Z_{\lambda 0}^m$ and $\mathcal{I}_{\lambda.\eta 0}^m$ denote $Z_{\lambda \varrho_m}^m$ and $\mathcal{I}_{\lambda.\eta \varrho_m}^m$ evaluated at $\varrho_m=0$. Define $\Lambda_\lambda^1$ as in ((ref)), and define $\Lambda_{\lambda \varrho_m}^2$ as in ((ref)) by replacing $\varrho$ with $\varrho_m$. Similar to ((ref)), define $\tilde{t}_{\lambda}^{m1}$ and $\tilde{t}_{\lambda \varrho_m}^{m2}$ by $r_{\lambda}(\tilde{t}_{\lambda}^{m1}) = \inf_{t_{\lambda} \in \Lambda_{\lambda}^1}r_{\lambda}^m(t_{\lambda})$ and $r_{\lambda\varrho_m}(\tilde{t}_{\lambda \varrho_m}^{m2}) = \inf_{t_{\lambda} \in \Lambda_{\lambda\varrho_m}^2}r_{\lambda\varrho_m}^m(t_{\lambda})$, where $r_{\lambda}^m(t_{\lambda}) := (t_{\lambda} -Z_{\lambda 0}^m)' \mathcal{I}_{\lambda.\eta 0}^m (t_{\lambda} -Z_{\lambda 0}^m)$ and $r_{\lambda\varrho_m}^m(t_{\lambda}) := (t_{\lambda} -Z_{\lambda \varrho_m}^m)' \mathcal{I}_{\lambda.\eta \varrho_m}^m (t_{\lambda} -Z_{\lambda \varrho_m}^m)$.

The following proposition establishes the asymptotic null distribution of the LRTS. As in the non-normal case, the LRTS is asymptotically distributed as the maximum of the $M_0$ random variables.

assumptionAssumption (ref) holds when $\tilde{s}_{\tilde \varrho_m k}$ is given in ((ref)).
propositionSuppose Assumptions (ref), (ref), (ref), (ref), and (ref) hold and the component density of the $j$-th regime is given by ((ref)). Then, under $H_0:m=M_0$, $LR_{M_0,n} \overset{d}{\rightarrow} \max_{m=1,\ldots,M_0}\{ \max\{ \mathbb{I}\{\varrho_m=0\} (\tilde t_{\lambda }^{m1})' \mathcal{I}_{\lambda.\eta 0}^m \tilde t_{\lambda }^{m1}, \sup_{\varrho_m \in \Theta_{\varrho}^m} (\tilde t_{\lambda \varrho_m}^{m2})' \mathcal{I}_{\lambda.\eta \varrho_m}^m \tilde t_{\lambda \varrho_m}^{m2} \}\}$.

Homoscedastic normal distribution

As in Section (ref), we assume that $Y_k\in \mathbb{R}$ in the $j$-th regime follows a normal distribution with the regime-specific intercept and common variance whose density is given by ((ref)).

The asymptotic distribution of the LRTS is derived by using a reparameterization

equation*[equation* omitted — 276 chars of source]

similar to ((ref)) and following the derivation in Sections (ref) and (ref). For brevity, we omit the details of the derivation. Define $s_{\lambda\lambda \varrho_m k}^m$ as in ((ref)), and denote each element of $s_{\lambda\lambda \varrho_m k}^m$ as

equation*[equation* omitted — 207 chars of source]

Similarly to ((ref)), define $\tilde{s}_{\tilde \varrho k}$ as in ((ref)) by redefining $s_{\lambda \varrho_m k}^m$ in ((ref)) as

equation[equation omitted — 272 chars of source]

where $s_{\lambda_{\mu}^i k}^m:= \mathbb{P}_{\vartheta^*_{M_0} (X_k \in J_m|\overline{{\bf Y}}_{0}^k)} \nabla_{\mu^i}f(Y_k|\overline{\bf Y}_{k-1};\gamma^*,\theta_m^*)/f(Y_k|\overline{\bf Y}_{k-1};\gamma^*,\theta_m^*)$ for $i=3,4$.

The following proposition establishes the asymptotic null distribution of the LRTS.

assumptionAssumption (ref) holds when $\tilde{s}_{\tilde \varrho_m k}$ is given in ((ref)).
propositionSuppose Assumptions (ref), (ref), (ref), (ref), and (ref) hold and the component density of the $j$-th regime is given by ((ref)). Then, under $H_0:m=M_0$, $ LR_{M_0,n} \overset{d}{\rightarrow} \max_{m=1,\ldots,M_0}\{ \max\{ \mathbb{I}\{\varrho_m=0\} (\tilde t_{\lambda }^{m1})' \mathcal{I}_{\lambda.\eta 0}^m \tilde t_{\lambda }^{m1}, \sup_{\varrho_m \in \Theta_{\varrho_m\epsilon}} (\tilde t_{\lambda \varrho_m}^{m2})' \mathcal{I}_{\lambda.\eta \varrho_m}^m \tilde t_{\lambda \varrho_m}^{m2} \}\}$, where $\tilde{t}_{\lambda}^{m1}$ and $\tilde{t}_{\lambda \varrho_m}^{m2}$ are defined as in Proposition (ref) but in terms of $(Z_{\lambda \varrho_m}^m, \mathcal{I}_{\lambda.\eta \varrho_m}^m, Z_{\lambda 0}^m, \mathcal{I}_{\lambda.\eta 0}^m)$ constructed with $s_{\lambda \varrho_m k}^m$ given in ((ref)) and $\Lambda_\lambda^1$ and $\Lambda_{\lambda \varrho_m}^2$ defined as in ((ref)) but replacing $\varrho$ with $\varrho_m$.

Asymptotic distribution under local alternatives

In this section, we derive the asymptotic distribution of our LRTS under local alternatives. While we focus on the case of testing $H_0: M=1$ against $H_A: M=2$, it is straightforward to extend the analysis to the case of testing $H_0: M=M_0$ against $H_A: M=M_0+1$ for $M_0\geq 2$.

Given $\pi\in \Theta_{\pi}$, we define a local parameter $h:=\sqrt{n}t(\psi,\pi)$, so that \[ h =

pmatrix[pmatrix omitted — 33 chars of source]

=

pmatrix[pmatrix omitted — 73 chars of source]

, \] where $t_{\lambda}(\lambda,\pi)$ differs across the different models and is given by ((ref)), ((ref)), and ((ref)). Given $h=(h_\eta',h_{\lambda}')'$ and $\pi\in\Theta_{\pi}$, we consider the sequence of contiguous local alternatives $\vartheta_n = (\psi_n',\pi_n')' = (\eta_n',\lambda_n',\pi_n')'\in\Theta_{\eta}\times\Theta_\lambda\times\Theta_{\pi}$ such that

equation[equation omitted — 174 chars of source]

Let $\mathbb{P}_{\vartheta,x_0}^n$ be the probability measure on $\{Y_k\}_{k=1}^n$ under $\vartheta$ conditional on the value of $\overline{\bf Y}_0$, $X_0$, and ${\bf W}_1^n$. Then, the log-likelihood ratio is given by \[ \log \frac{d\mathbb{P}_{\vartheta_n,x_0}^n}{d \mathbb{P}_{\vartheta^*,x_0}^n} = \ell_n(\psi_n,\pi_n,x_0)-\ell_n(\psi^*,\pi,x_0)= \log \left( \frac{\sum_{{\bf x}_1^{n}} \prod_{k=1}^n f_k(\eta_n, \lambda_n) q_{\pi_n}(x_{k-1},x_k) }{ \prod_{k=1}^n f_k(\eta^*,0) } \right), \] where $f_k(\eta,\lambda)$ is defined by the right-hand side of ((ref)), ((ref)), and ((ref)) for the models of the non-normal distribution, heteroscedastic normal distribution, and homoscedastic normal distribution, respectively. The following result follows from Le Cam's first and third lemmas and facilitates the derivation of the asymptotic distribution of the LRTS under $\mathbb{P}_{\vartheta_n,x_0}^n$.

propositionSuppose that the assumptions of Propositions (ref), (ref), and (ref) hold for the models of the non-normal, heteroscedastic normal, and homoscedastic normal distributions, respectively. Then, uniformly in $x_0 \in \mathcal{X}$, (a) $\mathbb{P}_{\vartheta_n,x_0}^n$ is mutually contiguous with respect to $\mathbb{P}_{\vartheta^*,x_0}^n$, and (b) under $\mathbb{P}_{\vartheta_n,x_0}^n$, we have $\log (d\mathbb{P}_{\vartheta_n,x_0}^n / d \mathbb{P}_{\vartheta^*,x_0}^n) = h' \nu_n(s_{\varrho_n k}) - \frac{1}{2}h' \mathcal{I}_{\varrho} h + o_{p}(1)$ with $\nu_n(s_{\varrho_n k}) \overset{d}{\rightarrow} N(\mathcal{I}_{ \varrho}h, \mathcal{I}_{ \varrho})$.

Non-normal distribution

For the non-normal distribution, the sequence of contiguous local alternatives is given by $\lambda_n = \bar \lambda/n^{1/4}$ because then $h_\lambda = \sqrt{n}\alpha(1-\alpha)v(\lambda_n)=\alpha(1-\alpha)v(\bar\lambda)$ holds. The following proposition derives the asymptotic distribution of the LRTS for the non-normal distribution under $H_{1 n}: (\pi_n,\eta_n,\lambda_n) = (\bar\pi,\eta^*,\bar \lambda/n^{1/4})$.

propositionSuppose that the assumptions of Proposition (ref) hold. For $\bar \pi\in\Theta_{\pi}$ and $\bar \lambda\neq 0$, define $h_{\lambda}:= \bar\alpha(1-\bar\alpha) v(\bar \lambda)$. Then, under $H_{1 n}: (\pi_n,\eta_n,\lambda_n) = (\bar\pi,\eta^*,\bar \lambda/n^{1/4})$, we have $LR_n \overset{d}{\rightarrow} \sup_{\varrho \in \Theta_{\varrho}} (\tilde t_{\lambda\varrho h})' \mathcal{I}_{\lambda.\eta\varrho} \tilde t_{\lambda\varrho h}$, where $\tilde t_{\lambda \varrho h}$ is defined as in ((ref)) but replacing $Z_{\lambda\varrho}$ in ((ref)) with $(\mathcal{I}_{\lambda.\eta\varrho})^{-1} G_{\lambda.\eta\varrho}+ h_{\lambda}$.

Heteroscedastic normal distribution

For the model with the heteroscedastic normal distribution, the sequences of contiguous local alternatives characterized by ((ref)) include the local alternatives of order $n^{-1/8}$.

propositionSuppose that the assumptions of Proposition (ref) hold for model ((ref)). For $\bar \varrho \in (-1,1)$, $\bar \alpha \in (0,1)$, and $\bar \lambda := (\bar \lambda_\mu,\bar \lambda_\sigma,\bar \lambda_\beta')'\neq (0,0,0)'$, let \begin{align*} & H_{1 n}^a: (\varrho_n,\alpha_n,\eta_n,\lambda_{\mu n}, \lambda_{\sigma n},\lambda_{\beta n}) = (\bar \varrho/n^{1/4}, \bar\alpha,\eta^*, \bar \lambda_\mu/ n^{1/8}, \bar \lambda_{\sigma}/n^{3/8},\bar \lambda_{\beta}/n^{3/8}), \\ & H_{1 n}^b: (\varrho_n,\alpha_n,\eta_n,\lambda_{\mu n}, \lambda_{\sigma n},\lambda_{\beta n}) = (\bar \varrho, \bar\alpha,\eta^*, \bar \lambda_\mu/ n^{1/4}, \bar \lambda_{\sigma}/n^{1/4},\bar \lambda_{\beta}/n^{1/4}), \end{align*} and define \begin{align*} h_{\lambda}^a: &= \bar\alpha(1-\bar\alpha) \times (\bar \varrho \bar \lambda_\mu^2, \bar \lambda_\mu \bar \lambda_\sigma, b(\bar\alpha) \bar \lambda_\mu^4/12, \bar \lambda_{\beta}' \bar \lambda_\mu, 0, 0)', \\ h_{\lambda}^b: &= \bar\alpha(1-\bar\alpha) \times (\bar \varrho\bar \lambda_\mu^2, \bar \lambda_\mu \bar \lambda_\sigma, \bar \lambda_\sigma^2, \bar \lambda_{\beta}' \bar \lambda_\mu, \bar \lambda_\beta' \bar \lambda_{\sigma}, v( \bar \lambda_\beta)')'. \end{align*} Then, for $j \in \{a,b\}$, under $H_{1 n}^j$, we have $LR_n \overset{d}{\rightarrow} \max\{ \mathbb{I}\{\varrho=0\} (\tilde t_{\lambda h}^{1j})' \mathcal{I}_{\lambda.\eta 0} \tilde t_{\lambda h}^{1j}, \sup_{\varrho \in \Theta_{\varrho}} (\tilde t_{\lambda \varrho h}^{2j})' \mathcal{I}_{\lambda.\eta \varrho} \tilde t_{\lambda \varrho h}^{2j} \}$, where $\tilde t_{\lambda h}^{1j}$ and $\tilde t_{\lambda \varrho h}^{2j}$ are defined as in ((ref)) but replacing $Z_{\lambda\varrho}$ with $(\mathcal{I}_{\lambda.\eta\varrho})^{-1} G_{\lambda.\eta\varrho}+ h_{\lambda}^j$.

In the local alternative $H_{1 n}^a$, $\varrho_n$ converges to $0$, and $\lambda_{\mu n}$ converges to 0 at a slower rate than $n^{-1/4}$. Our test has non-trivial power against these local alternatives in the neighborhood of $\varrho=0$. By contrast, the test of carrasco14em does not have power against the local alternatives in the neighborhood of $\varrho=0$, as discussed in Section 5 of carrasco14em. The test proposed by quzhuo17wp assumes that $\varrho$ is bounded away from zero and hence their test rules out $H_{1n}^a$.

Homoscedastic normal distribution

The local alternatives for the model with the homoscedastic distribution also include those of order $n^{-1/8}$ in the neighborhood of $\varrho=0$.

propositionSuppose that the assumptions of Proposition (ref) hold for model ((ref)). For $\bar \varrho \in (-1,1)$, $\bar \alpha \in (0,1)$, $\Delta_\alpha \neq 0$, and $\bar \lambda := (\bar \lambda_\mu,\bar \lambda_\beta')'\neq (0,0)'$, let \begin{align*} & H_{1 n}^a: (\varrho_n,\alpha_n,\eta_n,\lambda_{\mu n}, \lambda_{\beta n}) = (\bar \varrho/n^{1/4}, 1/2+\Delta_{\alpha}/n^{1/8}, \eta^*, \bar \lambda_\mu/ n^{1/8}, \bar \lambda_{\beta}/n^{3/8}), \\ & H_{1 n}^b: (\varrho_n,\alpha_n,\eta_n,\lambda_{\mu n}, \lambda_{\beta n}) = (\bar \varrho, \bar\alpha,\eta^*, \bar \lambda_\mu/ n^{1/4}, \bar \lambda_{\beta}/n^{1/4}), \end{align*} and define $h_{\lambda}^a := (1/4) \times (\bar\varrho \bar\lambda_{\mu}^2,\Delta_{\alpha} \bar\lambda_{\mu}^3,-\bar\lambda_\mu^4/2, \bar\lambda_\beta'\bar\lambda_\mu, 0)'$ and $h_{\lambda}^b := \bar\alpha(1-\bar\alpha) \times (\bar \varrho\bar \lambda_\mu^2, 0, 0, \bar \lambda_{\beta}' \bar \lambda_\mu, v( \bar \lambda_\beta)')'$. For $j=\{a,b\}$, define $\tilde t_{\lambda h}^{1j}$ and $\tilde t_{\lambda \varrho h}^{2j}$ as in ((ref)) but replacing $Z_{\lambda\varrho}$ with $(\mathcal{I}_{\lambda.\eta\varrho})^{-1} G_{\lambda.\eta\varrho}+ h_{\lambda}^j$, where $\mathcal{I}_{\lambda.\eta\varrho}$ and $G_{\lambda.\eta\varrho}$ are constructed with $s_{\varrho k}$ defined in ((ref)), and $\Lambda_{\lambda }^1$ and $\Lambda_{\lambda \varrho}^2$ are defined in ((ref)). Then, under $H_{1 n}^j$, we have $LR_n \overset{d}{\rightarrow} \max\{ \mathbb{I}\{\varrho=0\} (\tilde t_{\lambda h}^{1j})' \mathcal{I}_{\lambda.\eta 0} \tilde t_{\lambda h}^{1j}, \sup_{\varrho \in \Theta_{\varrho}} (\tilde t_{\lambda \varrho h}^{2j})' \mathcal{I}_{\lambda.\eta \varrho} \tilde t_{\lambda \varrho h}^{2j} \}$.

Parametric bootstrap

We consider the following parametric bootstrap to obtain the bootstrap critical value $c_{\alpha,B}$ and bootstrap $p$-value of our LRTS for testing $H_0: M=M_0$ against $H_A: M=M_0+1$.

enumerate• Using the observed data, compute $\hat \vartheta_{M_0}$, $\hat\vartheta_{M_0+1}$, and $LR_{M_0,n}$. • Given $\hat \vartheta_{M_0}$ and $\xi_{M_0}$, generate $B$ independent samples $\{Y_1^b,\ldots,Y_n^b\}_{b=1}^B$ under $H_0$ with $\vartheta_{M_0}=\hat\vartheta_{M_0}$ conditional on the observed value of $\overline{\bf Y}_{0}$ and ${\bf W}_{1}^n$. • For each simulated sample $\{Y_k^b\}_{k=1}^n$ with $(\overline{\bf Y}_0,{\bf W}_1^n)$, estimate $\hat \vartheta_{M_0}^b$ and $\hat\vartheta_{M_0+1}^b$ as in Step 1, and let $LR_{M_0,n}^b := 2[\ell_n(\hat\vartheta_{M_0+1}^b ,\xi_{M_0+1})-\ell_n(\hat\vartheta_{M_0}^b ,\xi_{M_0})]$ for $b=1,\ldots,B$. • Let $c_{\alpha,B}$ be the $(1-\alpha)$ quantile of $\{LR_{M_0,n}^b\}_{b=1}^B$, and define the bootstrap $p$-value as $B^{-1}\sum_{b=1}^B \mathbb{I}\{ LR_{M_0,n}^b > LR_{M_0, n}\}$.

The following proposition shows the consistency of the bootstrap critical values $c_{\alpha,B}$ for testing $H_0: M_0=1$. We omit the result for testing $H_0: M_0\geq 2$; it is straightforward to extend the analysis to the case for $M_0\geq 2$ with more tedious notations.

propositionSuppose that the assumptions of Propositions (ref), (ref), and (ref) hold for the models of the non-normal, heteroscedastic normal, and homoscedastic normal distributions, respectively. Then, the bootstrap critical values $c_{\alpha,B}$ converge to the asymptotic critical values in probability as $n$ and $B$ go to infinity under $H_0$ and under the local alternatives described in Propositions (ref), (ref), and (ref).

Simulations and empirical application

Simulations

We consider the following two models:

align[align omitted — 272 chars of source]

where $X_k\in\{1,\ldots,M\}$ with $p_{ij} = p(X_k=i|X_{k-1}=j)$. Model 1 in ((ref)) is similar to the model used in chowhite07em. This model has a switching intercept but the variance parameter $\sigma^2$ does not switch across regimes. In Model 2 in ((ref)), both the intercept and the variance parameters switch across regimes.

We compare the size and power property of our bootstrap LRT and that of the QLR test of chowhite07em with $\Theta_\mu = [-2,2]$ and the supTS test of carrasco14em with $\rho \in [-0.9. 0.9]$. The critical values are computed by bootstrap with $B=199$ bootstrap samples. Note that this comparison favors the LRT over the supTS test because the supTS test is designed to detect general parameter variation including Markov chain.

In Model 2, we set the lower bound of $\sigma$ as $\epsilon_\sigma = 0.01\hat \sigma$, where $\hat \sigma$ is the estimate of $\sigma$ from one-regime model. We also find that adding a penalty term to the log-likelihood function improves the finite sample property of the LRT. The penalty term prevents $\sigma_j$ from taking an extremely small value and takes the form $-a_n \sum_{j = 1}^{M} \{\hat\sigma^2/\sigma_j^2+\log(\sigma_j^2/\hat\sigma^2) - 1\}$. We set $a_n = 20n^{-1/2}$ and compute the test statistic using the log-likelihood function without the penalty term. Because the penalty term and its derivatives are $o_p(1)$ from the compactness of $\Theta_M$, adding this penalty term does not affect the consistency of the MLE or the asymptotic distribution of the LRTS.

We first examine the rejection frequency of $H_0:M=1$ against $H_A:M=2$ when the data are generated by $H_0: M=1$ with $(\beta,\mu,\sigma)=(0.5,0,1)$. The first panel in Table (ref) reports the rejection frequency of the bootstrap tests at the nominal 10%, 5%, and 1% levels over $3000$ replications with $n=200$ and $500$. Overall, all the tests have good sizes.

Table (ref) reports the power of the three tests for testing the null hypothesis of $M=1$ at the nominal level of 5%. We generate $3000$ data sets for $n= 500$ under the alternative hypothesis of $M=2$ by setting $\mu_1=0.2, 0.6$, and $1.0$ and $\mu_2=-\mu_1$, while $(p_{11},p_{22})=(0.25,0.25)$, $(0.50,0.50)$, $(0.70,0.70)$, and $(0.90,0.90)$. We set $\sigma=1$ for Model 1 and $(\sigma_1,\sigma_2)=(1.1,0.9)$ for Model 2. For Model 1, the power of all the tests increases as $\mu_1$ increases except for the supTS test with $(p_{11},p_{22})=(0.9,0.9)$. As $(p_{11},p_{22})$ moves away from $(0.5,0.5)$, the power of the LRT increases, whereas the QLRT has decreasing power. The LRT performs better than the supTS and QLR tests except for the case with $(p_{11},p_{22})=(0.25,0.25)$, where the supTS test performs very well, and the case with $(p_{11},p_{22})=(0.5,0.5)$, where the QLRT outperforms the LRT, because the true model is a finite mixture in this case. The last three columns of Table (ref) report the power of the LRT and supTS test to detect alternative models with switching variances (i.e., Model 2 with $M=2$). We did not examine the power of the QLRT because this test assumes non-switching variance. The LRT has stronger power than the supTS test in most cases.

The second panel in Table (ref) reports the rejection frequency of the LRT for testing $H_0:M=2$ against $H_A:M=3$ when the data are generated under the null hypothesis, showing its good size property, while neither the QLRT nor the supTS test is applicable for testing $H_0:M=2$ against $H_A:M=3$.

Table (ref) reports the power of our LRT for testing the null hypothesis of $M=2$ at the nominal level of 5%. We generate $3000$ data sets for $n= 500$ under the alternative hypothesis of $M=3$ across different values of $(\mu_1,\mu_2,\mu_3)$ and $(p_{11},p_{22},p_{33})$ with $p_{ij} = (1-p_{ii})/2$ for $j\neq i$, where we set $(\beta,\sigma)=(0.5,1.0)$ for Model 1 and $(\beta,\sigma_1,\sigma_2) = (0.5,0.9,1.2)$ for Model 2. Similar to the case of $H_0:M=1$, the power of the LRT for testing $H_0: M=2$ against $H_A: M=3$ increases when the alternative is further away from $H_0$ or when latent regimes become more persistent.

Empirical example

Using U.S. GDP per capita quarterly growth rate data from 1960Q1 to 2014Q4, we estimate the regime switching models with common variance (i.e., Model 1 in ((ref))) and with switching variances (i.e., Model 2 in ((ref))) for $M=1$, $2$, $3$, and $4$ and sequentially test the null hypothesis of $M=M_0$ against the alternative hypothesis $M=M_0+1$ for $M_0=1$, $2$, $3$, and $4$.\footnote{For both models, we restrict the parameter values for the transition probabilities by setting $\epsilon=0.05$ to prevent the issue of unbounded likelihood.} We also report the Akaike information criteria (AIC) and Bayesian information criteria (BIC) for reference, although, to our best knowledge, the consistency of the AIC and BIC for selecting the number of regimes has not been established in the literature.

Table (ref) reports the selected number of regimes by the AIC, BIC, and LRT. For Model ((ref)) with common variance, our LRT selects $M=4$, while the AIC and BIC select $M=3$ and $M=1$, respectively. For Model ((ref)) with switching variance, both the LRT and the AIC select $M=3$, while the BIC selects $M=2$.

Panel A of Table (ref) and Figure (ref) report the parameter estimates and posterior probabilities of being in each regime for the model with common variance for $M=2$, $3$, and $4$. Across the different specifications in $M$, the estimated values of $\mu_1$, $\mu_2, \ldots, \mu_M$ are well separated in the common variance model, indicating that each regime represents a booming or recession period with different degrees. In Figure (ref), when the number of regimes is specified as $M=2$, the posterior probability of the “recession” regime (Regime 1) against that of the “booming” regime (Regime 2) sharply rises during the collapse of Lehman Brothers in 2008 and then declines after 2009. When the number of regimes is specified as $M=3$, in addition to the “recession” and “booming” regimes corresponding to Regimes 1 and 2, respectively, the regime with a rapid change in the growth rates from low to high is captured by Regime 3; for the model with $M=3$ in Figure (ref), the posterior probability of Regime 3 rises in late 2009 when the U.S. economy started to recover from the Lehman collapse. When the number of regimes is specified as $M=4$, Regime 1 now captures a rapid change in the growth rates from high to low, where the posterior probability of Regime 1 becomes high when the growth rate of the U.S. economy rapidly declined in the middle of the Lehman collapse. The LRT selects the model with four regimes, which capture the rapid changes in the growth rates of U.S. GDP per capita during the Lehman collapse period.

Panel B of Table (ref) and Figure (ref) report the parameter estimates and posterior probabilities of being in each regime for the model with switching variance, respectively. When the number of regimes is specified as $M=2$, the variance parameter estimates are very different between the two regimes, while the intercept estimates are similar, indicating that Regime 1 is the “low volatility” regime, while Regime 2 is the “high volatility” regime. When the number of regimes is specified as $M=3$, different regimes capture different states of the U.S. economy in terms of both growth rates and volatilities.\footnote{We may test the null hypothesis of $\sigma_1=\sigma_2=\sigma_3$ in the model with switching variance given $M=3$ by the standard LRT with the critical value obtained from the chi-square distribution with two degrees of freedom. With $LRTS = 2\times (-297.01+307.39)=20.76$, the null hypothesis of $\sigma_1=\sigma_2=\sigma_3$ is rejected at the 1% significance level, suggesting that the model with switching variance is more appropriate than the model with common variance.} Regime 1 is characterized by the negative value of the intercept with high volatility, capturing a recession period. Regime 2 is characterized by the positive value of the intercept with low volatility, capturing a booming/stable economy. Regime 3 is characterized by a high value of the intercept and high variance, capturing both a rapid recovery in the growth rates and high volatility in the aftermath of the Lehman collapse in 2009. The LRT selects the model with three regimes when the model is specified with switching variance.