Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
79,444 characters · 17 sections · 98 citation commands
Asymptotic Properties of the Maximum Likelihood Estimator in Regime-Switching Models with Time-Varying Transition Probabilities
Keywords: {\em Regime-switching model, Time-varying transition probability, Asymptotic property, Maximum likelihood estimator.}
\setcounter{equation}{0} \setcounter{theorem}{0} \setcounter{assumption}{0} \setcounter{proposition}{0} \setcounter{corollary}{0} \setcounter{lemma}{0} \setcounter{example}{0} \setcounter{remark}{0}
Regime-switching models have been applied extensively since hamilton1989new to study how time-series patterns change across different underlying economic states, such as boom and recession, high-volatility and low-volatility financial market environments, and active and passive monetary and fiscal policies. This class of models features a bivariate process $(S_t,Y_t)$, where $(S_t)$ is an unobservable Markov chain determining the regime in each period, and $(Y_t)$ is an observable process whose conditional distribution is governed by the underlying regime $(S_t)$.
In basic Markov-switching models, like hamilton1989new, the unobservable regime process is assumed to follow a homogeneous Markov chain. It implies the transition probability and expected duration of each regime are constant through all periods, regardless of the level of the observable and how long the regime has lasted. The setup is restrictive in application and conflicts with empirical findings, such as those of burns1969progress and diebold1992have. The basic Markov-switching model is then extended in different ways to include the information of observations into the regime transition probabilities. When the observations take different values over time, the transition probabilities change accordingly, or simply, the model possesses time-varying transition probabilities (TVTP). diebold1994regime modelled the transition probability as a function of predetermined variables $(X_t)$, which can be chosen as economic covariates helpful to predict the regime change. The model is then applied widely; see, e.g., filardo1994business, bekaert1995time, gray1996modeling, filardo1998business, ang2008term. chib2004non, kim2008estimation, and chang2017new proposed another way to model the regime process. They introduced an autoregressive latent process, and the regime at each period is determined according to whether the latent process takes a value above or below a threshold. The information of the observable process $(Y_t)$ can be incorporated into the regime transition probabilities by allowing the innovations to $(Y_t)$ to be correlated with the innovations to the latent process. The information of predetermined variables can also be incorporated by embedding them to the latent process. This approach is then applied and extended; see, e.g., cheng2018new, song2018volatility, and chang2021origins.
The literature on the asymptotic properties of the maximum likelihood estimator (MLE) of TVTP regime-switching models is sparse, which might be owing to its theoretical challenges. The general difficulty in the proof of asymptotic theories with regime-switching models is that the predictive densities of the observable given past realizations do not form a stationary sequence, and thus, the ergodic theorem does not directly apply. The strategy, originated from baum1966statistical, is to approximate the log-likelihood function by the partial sum of a stationary ergodic sequence. The cornerstone of the approximation is the almost surely geometrically decaying bound of the mixing rate of the conditional chain $(S|Y,X)$ in probability and in $L^p$. douc2004asymptotic and kasahara2019asymptotic showed the approximation and the asymptotic properties of the MLE only in basic Markov-switching models.\footnote{baum1966statistical, leroux1992maximum, bickel1996inference, bickel1998asymptotic, jensen1999asymptotic, le2000exponential and douc2001asymptotics contribute to the asymptotic theories with the basic Markov-switching models less general than douc2004asymptotic and kasahara2019asymptotic. Their models, usually referred to as hidden Markov models, do not allow autoregression and ($Y_t$) are conditionally independent given the current regime.} kasahara2019asymptotic considered a more general model than douc2004asymptotic in that they allowed the observations to depend on lagged regimes and some elements of the regime transition probabilities to be zero. kasahara2019asymptotic showed only the approximation in probability and resorted to the dominated convergence theorem to show the approximation in $L^p$ by imposing some high-level assumptions on the moments of the period score and Hessian. The time-varying regime transition probability makes it more complex to show the bound, because it enters the mixing rate, causes the rate to approach unity as the transition probability approaches zero, and makes the rate non-decaying.
The theoretical works on TVTP regime-switching models include ailliot2015consistency and pouzo2018maximum. ailliot2015consistency showed consistency of the MLE. pouzo2018maximum showed consistency and asymptotic normality of the MLE under possible misspecification. Some of their assumptions are either too restrictive to hold in some widely-applied models or difficult to be verified in practice. First, they required the one-step transition probability of regimes to be positive. This assumption is key to showing the geometrically decaying bound of the mixing rate, because otherwise the bound would be unity and non-decaying. This assumption, however, excludes the models where the conditional distribution of $Y_t$ depends on lagged regimes. This assumption also precludes some elements of the transition probabilities being zero. Second, ailliot2015consistency particularly assumed the transition probabilities to be bounded away from zero. This assumption does not hold in the models such as diebold1994regime, where transition probability is a logistic function of $X_t$ and approaches zero as $X_t$ goes to infinity. Instead, pouzo2018maximum assumed the time-varying transition probabilities not to be too close to zero. Their assumption, though less restrictive than the one in ailliot2015consistency, involved sums of infinite terms, which might be difficult to verify in practice. Third, pouzo2018maximum made assumptions on the support of the conditional distribution of $Y_t$, which rules out the widely applied models with normal distributions.
This study shows consistency and asymptotic normality of the MLE in a wide range of TVTP regime-switching models. The main theoretical contribution of this study is that we show the mixing rate is decaying geometrically with a large probability and in $L^p$ by assuming there is a small probability that the observations take extreme values. Compared to kasahara2019asymptotic, our results in the approximation in $L^p$ do not rely on any high-level assumptions. We show the approximation in $L^p$ with the same assumptions that are used to show the approximation in probability. Compared to ailliot2015consistency and pouzo2018maximum, our model imposes weak conditions on the regime transition probabilities, which is verified to hold in models of diebold1994regime and chang2017new, and imposes weak conditions on the conditional density of $Y_t$, which allow for the density of a normal distribution. Then for some widely-applied models, such as filardo1994business, which extended the mean-switching AR(4) model in hamilton1989new to allow for time-varying transition probabilities, our theory applies while the theories of ailliot2015consistency and pouzo2018maximum do not.
We use simulation to examine how well the asymptotic distribution approximates the distribution of the MLE in finite samples. From the theoretical results, we can construct consistent estimates of the asymptotic variance of the MLE either from the negative of the inverse of the Hessian matrix or from the outer product of the score. We compare their performance and find that it is preferable to make inference based on the outer product of the score instead of the Hessian matrix.
The rest of this paper is organized as follows. Section (ref) lists the main assumptions and examples of TVTP regime-switching models. Section (ref) shows the geometrically decaying bounds of the mixing rate of the conditional chain in probability and in $L^p$. Sections (ref) and (ref) show consistency and asymptotic normality of the MLE, respectively. Section (ref) verifies the assumptions hold in a regime-switching autoregressive process with logistic or probit transition probabilities in diebold1994regime or transition probabilities in chang2017new. Section (ref) reports the simulation results. Section (ref) presents an empirical illustration and examines the contribution of including leading economic indicators into transition probabilities to describing U.S. industrial production. Section (ref) concludes.
\setcounter{equation}{0} \setcounter{theorem}{0} \setcounter{assumption}{0} \setcounter{proposition}{0} \setcounter{corollary}{0} \setcounter{lemma}{0} \setcounter{example}{0} \setcounter{remark}{0}
Notation: Let $\triangleq$ denote “equals by definition". Let “i.o." stand for “infinitely often". For a matrix or vector $M$, $\|M\|\triangleq\sum|M_{ij}|$. For two probability measures $\mu_1$ and $\mu_2$, define the total variation distance between $\mu_1$ and $\mu_2$ as $\|\mu_1-\mu_2\|_{TV}\triangleq\sup_A|\mu_1(A)-\mu_2(A)|$. Let $\mathbbm{1}\{A\}$ denote an indicator function that takes the value of one when $A$ is true and zero otherwise. Let $a\wedge b\triangleq\min\{a,b\}$. Let $\lfloor x\rfloor$ denote the largest integer less than or equal to $x$. For any $\{x_i\}$, define $\sum_{i=a}^b x_i\triangleq0$ and $\prod_{i=a}^b x_i\triangleq1$ when $b<a$. For a square matrix $A$, its spectral radius is denoted as $\rho(A)$. $\Phi(\cdot)$ and $\varphi(\cdot)$ are the distribution function and the density function of $\mathbb{N}(0,1)$, respectively.
We are interested in regime-switching models where the transition probabilities are allowed to include information from observations. In a regime-switching model, the conditional distribution of the observable process $(Y_t)$ is governed by an unobservable regime process $(S_t)$:
where $\overline{\mathbf{Y}}_t\triangleq(Y_t,\ldots,Y_{t-p+1})'$, $X_t$ is a predetermined variable (vector), $U_t$ is an independent and identically distributed (i.i.d.) sequence of random variables, and $f_\theta$ is a family of functions indexed by $\theta$. We allow $S_t$ to be a vector
so that the conditional density of $Y_t$ can depend on both the current and lagged regimes. One example is the autoregressive model with switching in mean and variance:
where $\gamma(z)=1-\gamma_1z-\ldots-\gamma_pz^p$. The seminal mean-switching AR(4) model of hamilton1989new is a special case of ((ref)) with $p=4,d=5$ and no switching in the variance.
We model the regime transition probability as
where $q_\theta$ is a family of probabilities indexed by $\theta$. We allow the regime transition probability ((ref)) to include information from observations, $\overline{\mathbf{Y}}_{t-1}$ and $X_t$. In the literature, diebold1994regime and chang2017new are two examples of such transition probabilities.
The above two models specify the transition of $\tilde{S}_t$ in ((ref)) and ((ref)). To obtain ((ref)) when $d\geq 2$, we write
For short notation, we define
for $n\geq m$. We similarly define $\tilde{\mathbf{S}}_m^n$, $\mathbf{S}_m^n$, and $\mathbf{X}_m^n$.
We assume that $\{S_t\}_{t=0}^{\infty}$ takes a value in a discrete set $\mathbb{S}$ and use $\mathcal{P}(\mathbb{S})$ to denote the power set of $\mathbb{S}$. For each $t\geq 1$ and given $(\overline{\mathbf{Y}}_{t-1},S_{t-1},X_t)$, $S_t$ is conditionally independent of $(\mathbf{Y}_{-p+1}^{t-p-1},\allowbreak \mathbf{S}_{0}^{t-2},\allowbreak \mathbf{X}_1^{t-1},\mathbf{X}_{t+1}^\infty)$. The transition probability is $q_\theta(s_t|S_{t-1},\overline{\mathbf{Y}}_{t-1},X_t)$. $\{Y_t\}_{t=-p+1}^{\infty}$ takes a value in a set $\mathbb{Y}$, which is separable and metrizable by a complete metric. Let $\overline{\mathbb{Y}}\triangleq \mathbb{Y}^p$. For each $t\geq1$ and given $(\overline{\mathbf{Y}}_{t-1},S_{t},X_t)$, $Y_t$ is independent of $(\mathbf{Y}_{-p+1}^{t-p-1},\mathbf{S}_{0}^{t-1},\mathbf{X}_1^{t-1},\allowbreak \mathbf{X}_{t+1}^\infty)$. The conditional law has a density $g_\theta(y_t|\overline{\mathbf{Y}}_{t-1},S_t,X_t)$ with respect to some fixed $\sigma-$finite measure $\nu$ on the Borel $\sigma-$field $\mathcal{B}(\mathbb{Y})$. $\{X_t\}_{t=1}^\infty$ takes a value in a set $\mathbb{X}$. Under the setup, conditional on $\mathbf{X}_{1}^{\infty}$, $(S_t,\overline{\mathbf{Y}}_t)$ follows a Markov chain with transition density
Moreover, we assume conditionally on $X_t$, $\{X_k\}_{k\geq t+1}$ is independent of $\{Y_k\}_{k\leq t}$ and $\{S_k\}_{k\leq t}$, and conditionally on $X_{t}$, $\{X_k\}_{k\leq t-1}$ is independent of $\{Y_k\}_{k\geq t}$ and $\{S_k\}_{k\geq t}$. Then $p_\theta(\mathbf{Y}_1^t|\overline{\mathbf{Y}}_0,S_0=s_0,\mathbf{X}_1^n) =p_\theta(\mathbf{Y}_1^t|\overline{\mathbf{Y}}_0,S_0=s_0,\mathbf{X}_1^t), p_\theta(\mathbf{Y}_1^t|\overline{\mathbf{Y}}_0,\mathbf{X}_1^n)=p_\theta(\mathbf{Y}_1^t|\overline{\mathbf{Y}}_0,\mathbf{X}_1^t). $ The following are the basic assumptions.
\setcounter{equation}{0} \setcounter{theorem}{0} \setcounter{assumption}{0} \setcounter{proposition}{0} \setcounter{corollary}{0} \setcounter{lemma}{0} \setcounter{example}{0} \setcounter{remark}{0}
This study works with the conditional likelihood function given initial observations $\overline{\mathbf{Y}}_0=(Y_0,\dots,Y_{-p+1})$, (unobservable) initial regime $S_0$, and predetermined variables $\mathbf{X}_1^t$, owing to the difficulties obtaining the closed-form expression of the unconditional stationary likelihood function. We can write the conditional log-likelihood function as
with the predictive density
for $t\geq1$. When the number of observations is $n+p$, we condition on the first $p$ observations, and arbitrarily choose the initial regime $s_{0}$. The aim of this study is to show consistency and asymptotic normality of the MLE $\hat{\theta}_{n,s_0}=\operatorname*{arg\,max}_\theta\ell_n(\theta,s_0)$ with any choice of $s_0$, even when it is not the true underlying initial regime.
Consistency and asymptotic normality follow if we can show the following two results: (a) the normalized log-likelihood $n^{-1}\ell_n(\theta,s_0)$ converges to a deterministic function $\ell(\theta)$ uniformly with respect to $\theta$, and $\theta^*$ is a well-separated point of the maximum of $\ell(\theta)$; and (b) a central limit theorem for the Fisher score function and a locally uniform law of large numbers for the observed Fisher information.
The updating distribution $\mathbb{P}_\theta(S_{t-1}=s_{t-1}|S_0=s_0,\overline{\mathbf{Y}}_0^{t-1},\mathbf{X}_1^{t-1})$ in ((ref)) is not stationary ergodic and thus, the predictive density is not stationary ergodic. Therefore, the ergodic theorem cannot be applied directly to conclude a law of large numbers or central limit theorems. The strategy, originated from baum1966statistical, is to approximate the predictive density with a stationary ergodic sequence. This section gives the theoretical foundations of the approximation.
We choose the stationary ergodic sequence to be
((ref)) differs from ((ref)) in that the updating distribution now depends on the whole history of $(Y_t,X_t)$ from past infinity and does not depend on the initial regime. We can extend $(Y_t,X_t)$ to doubly infinite time $\{Y_t,X_t\}_{t=-\infty}^{+\infty}$ because of Assumption (ref). Stationary ergodicity of ((ref)) follows from Theorem 7.1.3 in durrett2013probability. The approximation builds on the almost surely geometrically decaying bound of the mixing rate of the conditional chain $(S|Y,X)$. The bound guarantees that the influence of the observations and initial regime far in the past quickly vanishes and thus, the difference between the exact predictive and approximated predictive log densities becomes asymptotically negligible. Lemma (ref) establishes an upper bound for the mixing rate.
$\sigma_-(\overline{\mathbf{Y}}_{k-r}^{k-1},\mathbf{X}_{k-r+1}^{k})$ might approach zero when the observations take extreme values, as in the examples of diebold1994regime and chang2017new. Then in such extreme cases, the upper bound in ((ref)) is unity and does not decay over time. In Lemma (ref), We show that under Assumption (ref), which restricts the probability of such extreme cases, the upper bound in ((ref)) is decaying geometrically with a large probability and in $L^p$.
\setcounter{equation}{0} \setcounter{theorem}{0} \setcounter{assumption}{0} \setcounter{proposition}{0} \setcounter{corollary}{0} \setcounter{lemma}{0} \setcounter{example}{0} \setcounter{remark}{0}
This sections proves consistency of the MLE $\hat{\theta}_{n,s_0}=\operatorname*{arg\,max}_\theta \ell_n(\theta,s_0)$. We first show that the normalized log-likelihood $n^{-1}\ell_n(\theta,s_0)$ converges to a deterministic function $\ell(\theta)$ uniformly with respect to $\theta$. The first step is to approximate $n^{-1}\ell_n(\theta,s_0)$ with the sample mean of a $\mathbb{P}_{\theta^*}$-stationary ergodic sequence of random variables. The construction of the stationary ergodic sequence is as follows. Define
so that $n^{-1}\ell_n(\theta,s_0)=\frac{1}{n}\sum_{t=1}^n\Delta_{t,0,s_0}(\theta)$. We show that $\{\Delta_{t,m,s}(\theta)\}_{m\geq0}$ is a uniformly Cauchy sequence with respect to $\theta\in\Theta$ $\mathbb{P}_{\theta^*}$-a.s. in ((ref)) of Lemma (ref).
((ref)) is derived from the results in Section 3. To see this, we take $\mu_1(s)=\delta_{s_i}(s)$ and $\mu_2(s)=\mathbb{P}_{\theta}(S_{-m}=s|S_{-m'}=s_j,\overline{\mathbf{Y}}_{-m'}^n,\mathbf{X}_{-m'}^n)$ in Lemma (ref). Then ((ref)) gives an upper bound for the distance between $\mathbb{P}_{\theta}(S_k\in\cdot|S_{-m}=s_i,\overline{\mathbf{Y}}_{-m}^n,\mathbf{X}_{-m}^n)$ and $\mathbb{P}_{\theta}(S_k\in\cdot|S_{-m'}=s_j,\overline{\mathbf{Y}}_{-m'}^n,\mathbf{X}_{-m'}^n)$. ((ref)) follows by combining Lemmas (ref) and (ref). We give the details in the online appendix.
((ref)) indicates $\{\Delta_{t,m,s}(\theta)\}_{m\geq0}$ converges uniformly $\mathbb{P}_{\theta^*}$-a.s. and the limit does not depend on $s$. Define $\Delta_{t,\infty}(\theta)\triangleq\lim_{m\rightarrow\infty}\Delta_{t,m,s}(\theta)$, which is well-defined in $L^1(\mathbb{P}_{\theta^*})$. From ((ref)), $\Delta_{t,\infty}(\theta)=\lim_{m\rightarrow\infty}\Delta_{t,m}(\theta)$. From Assumption (ref), $\{\Delta_{t,\infty}\}_{t\geq0}$ is stationary ergodic, and the ergodic theorem applies:
Let $m=0$ and $m'\rightarrow\infty$ in ((ref)), which yields
((ref)) shows the difference between $n^{-1}\ell_n(\theta,s_0)$ and $\frac{1}{n}\sum_{t=1}^n \Delta_{t,\infty}(\theta)$ is asymptotically negligible. Then $n^{-1}\ell_n(\theta,s_0)$ also converges to $\ell(\theta)$ $\mathbb{P}_{\theta^*}$-a.s. From Assumption (ref), we can show the continuity of $\ell(\theta)$. Then $n^{-1}\ell_n(\theta,s_0)$ converges to $\ell(\theta)$ uniformly with respect to $\theta$, which is shown in Proposition (ref).
Consistency follows once we show that $\theta^*$ is a well-separated point of the maximum of $\ell(\theta)$. Theorem (ref) summarizes the finding and the details of proof are in the online appendix.
\setcounter{equation}{0} \setcounter{theorem}{0} \setcounter{assumption}{0} \setcounter{proposition}{0} \setcounter{corollary}{0} \setcounter{lemma}{0} \setcounter{example}{0} \setcounter{remark}{0}
This section establishes asymptotic normality of the MLE. We need additional regularity assumptions. Assume a positive real $\delta$ exists such that on $G\triangleq \{\theta\in\Theta: |\theta-\theta^*|<\delta\}$, the following conditions hold:
This subsection shows asymptotic normality of the score function. First, we show that the score function can be approximated by a sequence of integrable martingale increments. We use the Fisher identity (cappe2005inference) to write the score function as the expectation of the complete score conditional on the observed data:
where $\phi_{\theta,t}$ is the short-hand notation for
We write the period score function as
so that $\nabla_\theta \ell_n(\theta,s_0)=\sum_{t=1}^n\dot{\Delta}_{t,0,s_0}(\theta)$. We can similarly define
The stationary conditional score is constructed by conditioning on the whole history of $(Y_t,X_t)$ starting from past infinity: $\dot{\Delta}_{t,\infty}(\theta)\triangleq\lim_{m\rightarrow\infty} \dot{\Delta}_{t,m}(\theta)$. Define the filtration $\mathcal{F}_t=\sigma\big((\overline{\mathbf{Y}}_k,X_{k+1}),-\infty<k\leq t\big)$ for $t\in\mathbb{Z}$. ((ref))-((ref)) show that $\{\dot{\Delta}_{t,\infty}(\theta^*)\}_{t=-\infty}^{\infty}$ is an $(\mathcal{F},\mathbb{P}_{\theta^*})$-adapted stationary ergodic and square integrable martingale increment sequence. The central limit theorem for the sums of such a sequence shows that
where $I(\theta^*)\triangleq \mathbb{E}_{\theta^*}[\dot{\Delta}_{0,\infty}(\theta^*)\dot{\Delta}_{0,\infty}(\theta^*)']$ is the asymptotic Fisher information matrix.
Lemma (ref) shows that $\dot{\Delta}_{t,0,s_0}(\theta^*)$ can be approximated by $\dot{\Delta}_{t,\infty}(\theta^*)$ in $L^2(\mathbb{P}_{\theta^*})$. The approximation depends on the distance between regime processes that start from different time points, whose upper bound decays geometrically in $L^2(\mathbb{P}_{\theta^*})$ according to ((ref)).
$n^{-\frac{1}{2}}\sum_{t=1}^n\dot{\Delta}_{t,0,s_0}(\theta^*)$ has the same limiting distribution as $n^{-\frac{1}{2}}\sum_{t=1}^n\dot{\Delta}_{t,\infty}(\theta^*)$ by Lemma (ref).
This subsection presents the law of large numbers for the observed Fisher information. We use the Louis missing information principle (louis1982finding) to express the observed Fisher information in terms of the Hessian of the complete log-likelihood function:
where $\dot{\phi}_{\theta,t}$ is the short-hand notation for
We approximate the expressions in ((ref)) with stationary ergodic sequences $\mathbb{P}_{\theta^*}$-a.s. and in $L^1(\mathbb{P}_{\theta^*})$. Similar to Lemma (ref), we use ((ref)) to show the approximation in $L^1(\mathbb{P}_{\theta^*})$, while kasahara2019asymptotic imposed high-level assumptions on the moments of derivatives of predictive densities. The construction of the stationary ergodic sequences is similar to the one in Subsection (ref), and thus we leave the details in the online appendix and directly give the theoretical finding.
Theorems (ref) and (ref) together yield the following theorem of asymptotic normality.
From Theorems (ref), (ref), and (ref), the negative of the inverse of the Hessian matrix at the MLE value is a consistent estimate of the asymptotic variance. Next proposition gives another consistent estimate constructed from the outer product of the score.
A third consistent estimate, provided in cappe2005inference and also motivated from Proposition (ref), can be constructed from the outer product of the demeaned score
We compare their performance in finite samples through simulation in Section (ref).
\setcounter{equation}{0} \setcounter{theorem}{0} \setcounter{assumption}{0} \setcounter{proposition}{0} \setcounter{corollary}{0} \setcounter{lemma}{0} \setcounter{example}{0} \setcounter{remark}{0}
This section verifies the assumptions hold in an autoregressive process with regime-dependent means, variances, and autoregressive coefficients
and transition probabilities of Examples (ref) and (ref).
francq2001stationarity gave conditions for Markov-switching autoregressive models to be stationary ergodic. We extend their result to TVTP regime-switching models.
((ref)) covers a wide range of regime-switching autoregressive models. For instance, in ((ref)), if $p=k$ and $d=k+1$, we can rewrite ((ref)) as $\overline{\mathbf{Y}}_t=A(\tilde{S}_t)\overline{\mathbf{Y}}_{t-1}+B(S_t,X_t)$ with
and $B(S_t,X_t)=(\mu(\tilde{S}_t)-\sum_{j=1}^k\gamma_j(\tilde{S}_t)\mu(\tilde{S}_{t-j})+\sigma(\tilde{S}_t)U_t,0,\ldots,0)'.$
If the autoregressive coefficient does not depend on $\tilde{S}_t$, i.e., $A(\tilde{S}_t)=A$, then (c) is the usual stationarity condition of autoregressive processes. For instance, for the seminal hamilton1989new model, we only need to check whether $\rho(A)<1$. In the following analysis of this subsection, we consider the case where the autoregressive coefficient does depend on $\tilde{S}_t$.
In Example (ref), (a) in Theorem (ref) holds if, for instance, $X_t$ follows $X_t=\eta X_{t-1}+\xi_t$, $\xi_t\sim $ i.i.d. $\mathbb{N}(0,\Sigma)$ with $\|\eta\|<1$, and the transition probabilities are strictly positive and continuous with respect to $X_t$. For (c') in Theorem (ref), the expectation of the random matrix might not be easily computed analytically, even for the widely applied logistic and probit functions. Instead, we can verify (c') numerically for given parameter values.
For Example (ref), note that the transition probability is essentially a function of $(\tilde{S}_t,\tilde{S}_{t-1},U_{t-1})$. Then we express the transition probability as $\tilde{q}_\theta(\tilde{S}_t|\tilde{S}_{t-1},U_{t-1})$, so that we can apply Theorem (ref). From ((ref))--((ref)), $(\tilde{S}_t,U_{t-1})$ is strictly stationary ergodic. Next, we derive an expression of (c') in Theorem (ref). Using $\tilde{q}_{ij}\triangleq\mathbb{E}_{\theta^*}[\tilde{q}_\theta(\tilde{S}_t=j|\tilde{S}_{t-1}=i,U_{t-1})]=(1-j)w+j(1-w)$, where $i,j=0,1$ and
(c') is equivalent to $\|M\|<1$ with
We give a sufficient condition for Assumption (ref) that is easier to verify in practice.
First consider Example (ref) with $\tilde{q}_\theta(\tilde{S}_t=1|\tilde{S}_{t-1}=1,X_t)=\frac{\exp(X'_t\beta_1)}{1+\exp(X'_t\beta_1)}$, $\tilde{q}_\theta(\tilde{S}_t=1|\tilde{S}_{t-1}=2,X_t)=\frac{\exp(X'_t\beta_2)}{1+\exp(X'_t\beta_2)}$, $\mathbb{E}_{\theta^*}[\|X_t\|^2]<\infty$, and $U_t\sim \text{i.i.d. }\mathbb{N}(0,1)$ in ((ref)). In the definition of $\overline{\mathbf{Y}}_t$, $p=k$. In the definition of $S_t$, $d=k+1$. We choose $r=d$ and $\delta=1$. Then
It follows that
Then it suffices to show that
((ref)) follows from the normal distribution of $U_t$ and $\mathbb{E}_{\theta^*}[U_t^4]<\infty$. The following lemma shows that ((ref)) is satisfied.
Next consider Example (ref) with $\tilde{q}_\theta(\tilde{S}_t=1|\tilde{S}_{t-1}=1,X_t)=\Phi(X'_t\beta_1)$, $\tilde{q}_\theta(\tilde{S}_t=1|\tilde{S}_{t-1}=2,X_t)=\Phi(X'_t\beta_2)$, $\mathbb{E}_{\theta^*}[\|X_t\|^4]<\infty$, and $U_t\sim \text{i.i.d. }\mathbb{N}(0,1)$ in ((ref)). In this example, in the definition of $S_t$, $d=k+1$. We choose $r=d$ and $\delta=1$. Similar to the case of logistic functions, it suffices to show ((ref)) and ((ref)). ((ref)) follows from the normal distribution of $U_t$ and $\mathbb{E}_{\theta^*}[U_t^4]<\infty$. The following lemma shows that ((ref)) is satisfied.
Last, consider Example (ref). Because $(Y_t,\tilde{S}_t)$ is a Markov process of order $k+1$, in the definition of $S_t$, $d=k+1$, and in the definition of $\overline{\mathbf{Y}}_t$, $p=k+1$. We choose $r=d$ and $\delta=1$. Then
From $ \mathbb{E}_{\theta^*}[|\log\sigma_-(\overline{\mathbf{Y}}_0^{d-1})|^2]\leq\mathbb{E}_{\theta^*}[|\sum_{\ell=1}^d\inf_{\theta}\min_{\tilde{\mathbf{s}}_{\ell-d}^\ell}\log\tilde{q}_\theta(\tilde{s}_\ell|\tilde{s}_{\ell-1},\ldots,\tilde{s}_{\ell-d},\overline{\mathbf{Y}}_{\ell-1})|^2], $ it suffices to show $\mathbb{E}_{\theta^*}[|\inf_{\theta}\min_{\tilde{\mathbf{s}}_{\ell-d}^\ell}\log\tilde{q}_\theta(\tilde{s}_\ell|\tilde{s}_{\ell-1},\ldots,\tilde{s}_{\ell-d},\overline{\mathbf{Y}}_{\ell-1})|^2]<\infty$, which is shown in the following lemma.
Write the conditional density as
It is a class of finite mixtures of the $n$-fold product densities $\prod_{t=1}^n g_\theta(Y_t|\overline{\mathbf{Y}}_{t-1},s_t,X_t)$ with mixing distributions $\prod_{t=2}^n q_\theta(s_t|s_{t-1},\overline{\mathbf{Y}}_{t-1},X_t)\mathbb{P}_\theta(S_1=s_1|\overline{\mathbf{Y}}_0,\mathbf{X}_{-p+1}^1)$. Because the class of finite mixtures of the normal family $g_\theta(\cdot)$ is identifiable (cappe2005inference), Assumption (ref) holds for ((ref)) if the following two conditions hold. First, $\mu(\tilde{s}_t)+\sum_{j=1}^{k}\gamma_j(\tilde{s}_{t})(Y_{t-j}-\mu(\tilde{s}_{t-j}))\neq \mu(\tilde{s}_t')+\sum_{j=1}^{k}\gamma_j(\tilde{s}_{t}')(Y_{t-j}-\mu(\tilde{s}_{t-j}'))$ or $\sigma(\tilde{s}_t)\neq \sigma(\tilde{s}_t')$, for $\tilde{\mathbf{s}}_{t-k}^t\neq \tilde{\mathbf{s}}^{\prime t}_{t-k}$. That is, the means and variances in different regimes are distinct up to renumbering of regimes. Second, $q_\theta(\cdot)=q_{\theta^*}(\cdot)$ if and only if $\theta=\theta^*$.
\setcounter{equation}{0} \setcounter{theorem}{0} \setcounter{assumption}{0} \setcounter{proposition}{0} \setcounter{corollary}{0} \setcounter{lemma}{0} \setcounter{example}{0} \setcounter{remark}{0}
We use simulation to examine how well the asymptotic distribution approximates the distribution of the MLE in finite samples. We consider the hamilton1989new regime-switching model with transition probabilities $\tilde{q}_\theta(\tilde{S}_t=1|\tilde{S}_{t-1}=1,X_t)=\frac{\exp(\beta_{10}+\beta_{11}{X}_t)}{1+\exp(\beta_{10}+\beta_{11}{X}_t)}$ and $\tilde{q}_\theta(\tilde{S}_t=1|\tilde{S}_{t-1}=2,X_t)=\frac{\exp(\beta_{20}+\beta_{21}{X}_t)}{1+\exp(\beta_{20}+\beta_{21}{X}_t)}$. The data generating process of $X_t$ is $X_t=aX_{t-1}+\xi_t$ with $a=0.4$ and $\xi_t\sim$ i.i.d.$\mathbb{N}(0,1)$. The true parameter is taken from the estimation result in Section (ref), $(\mu(1)^*,\mu(2)^*,\gamma_1^*,\gamma_2^*,\gamma_3^*,\gamma_4^*,\sigma^*,\beta_{10}^*,\allowbreak \beta_{11}^*,\beta_{20}^*,\beta_{21}^*)=(-2.33, 0.16, 0.08, 0.17, 0.15,0.005,0.50,\allowbreak -1.70,\allowbreak -1.61,-5.66,-4.85)$. As commented in Remark (ref), this model is not covered by the theory of ailliot2015consistency and pouzo2018maximum but by our theory, and we have shown the assumptions hold in Section (ref).
The sample size ranges from 200 to 1000. For each sample size, we generate 1000 data sets. For each data, we use the Broyden--Fletcher--Goldfarb--Shanno algorithm to find the MLE. We construct the confidence interval from the Hessian matrix in ((ref)), the outer product of the score in ((ref)), and the outer product of the demeaned score in ((ref)).\footnote{The Hessian matrix is computed with the method in cappe2005inference and is generalized to allow for TVTP specifications. The score is computed following hamilton1996specification.}
Table (ref) reports the frequency at which the 95% confidence interval contains the true parameter. It is worth mentioning that in some replications, the diagonal elements of the inverse of ((ref)) are negative, and we fail to construct confidence intervals. We count it as the intervals not including the true parameters.\footnote{The numbers of replications with negative diagonal elements in the inverse of ((ref)) are respectively 7, 7, and 0 when $n= 200$, 600, and 1000.} The coverage frequency of the interval constructed from ((ref)) is the same as the one constructed from ((ref)), so we put their results together in Panel B of Table (ref) to save space.\footnote{The reason why the converge frequencies based on ((ref)) and ((ref)) are the same might be that the score mean $\frac{1}{n}\sum_{t=1}^n\dot{\Delta}_{t,0,s}(\hat{\theta}_{n,s})$ is of small scale. The average of the absolute value of the score mean across 1000 simulation experiments is no more than $1\times 10^{-7}$ when $n=200, 600,$ and 1000.} In Panel A, the coverage frequency of $\sigma$ at $n=1000$, though improves slightly from $n=200$, is still much lower than 95%. In Panel B, the coverage frequencies of all parameters improve as $n$ increases from 200 to 1000 and are generally close to 95%. In this simulation experiment, the confidence interval constructed from the outer product of the (demeaned) score outperforms the one constructed from the Hessian matrix. Considering also the huge efforts in the computation of the Hessian matrix, it might be preferable to make inference based on the outer product of the score instead of the Hessian matrix.
\setcounter{equation}{0} \setcounter{theorem}{0} \setcounter{assumption}{0} \setcounter{proposition}{0} \setcounter{corollary}{0} \setcounter{lemma}{0} \setcounter{example}{0} \setcounter{remark}{0}
This section examines the contribution of including leading economic indicators into regime transition probabilities to describing the expansion and contraction of U.S. industrial production. In the study of business cycles with TVTP regime-switching models, the predetermined variables are usually those considered to be useful as business-cycle predictors (filardo1994business), and leading economic indicators are possible candidates. The Organisation for Economic Co-operation and Development (OECD) database provides three leading indicators: the composite leading indicator (CLI), the business confidence index (BCI), and the consumer confidence index (CCI). As far as we know, there have not been studies on the performance of the three indicators in identifying the turning points of business cycles, and thus we fill the gap.
We model the monthly U.S. industrial production log growth rate as a mean-switching AR(4) process with two regimes as in hamilton1989new, and transition probabilities are $\tilde{q}_\theta(\tilde{S}_t=1|\tilde{S}_{t-1}=1,X_t)=\frac{\exp(\beta_{10}+\beta_{11}{X}_t)}{1+\exp(\beta_{10}+\beta_{11}{X}_t)}$ and $\tilde{q}_\theta(\tilde{S}_t=1|\tilde{S}_{t-1}=2,X_t)=\frac{\exp(\beta_{20}+\beta_{21}{X}_t)}{1+\exp(\beta_{20}+\beta_{21}{X}_t)}.$ $X_t$ is taken to be the lagged demeaned log growth rate of montly CLI, BCI, or CCI. The data of U.S. industrial production, CLI, BCI, and CCI is downloaded from the OECD website. The sample period is Jan 1984--Dec 2019. We choose the sample period because there seem structural breaks in the variance of industrial production respectively in 1984 and 2020. We proceed to compare the performance of the three indicators from three perspectives: hypothesis tests of including the indicators, the response of transition probabilities to changes in the indicators, and the inferred probability of recessions.
Table (ref) reports the estimation result. Based on the simulation results in Section (ref), we compute the standard errors according to ((ref)). Because regime-switching models are identified up to renumbering of regimes, we restrict $\mu(1)<\mu(2)$. Then regime 1 and 2 can be interpreted as recession and boom, respectively. The last row reports the $p$-value of the likelihood ratio test of not including predetermined variables. At the 1% level of significance, we can reject the null hypothesis in the case of CLI but fail to reject the null in the case of BCI and CCI.
Figure (ref) plots the inferred transition probabilities. It is worth noting that the transition probability from boom to recession responds strongly to changes in CLI. In fact, the transition probability to recession from boom is close to one during the early 1990s recession and the Great Recession.
Figure (ref) plots the smoothed probability of the recession regime and the NBER recession periods. All of the models produce high probabilities of recession during the Great Recession. Only the case of CLI produces high probabilities of recession in the early 1990s recession. All except the case of CLI gives false signals of recession in early 2005. All of the models, however, fail to identify the early 2000s recession.
The above comparison shows that CLI might be preferable from the perspective of hypothesis testing and its success at identifying the early 1990s recession. We also consider the model that includes all three leading indicators with $\tilde{q}_\theta(\tilde{S}_t=1|\tilde{S}_{t-1}=1,X_t)=\frac{\exp(\beta_{10}+\beta_{11}CLI_{t-1}+\beta_{12}BCI_{t-1}+\beta_{13}CCI_{t-1})}{1+\exp(\beta_{10}+\beta_{11}CLI_{t-1}+\beta_{12}BCI_{t-1}+\beta_{13}CCI_{t-1})}$, and $\tilde{q}_\theta(\tilde{S}_t=1|\tilde{S}_{t-1}=2,X_t)=\frac{\exp(\beta_{20}+\beta_{21}CLI_{t-1}+\beta_{22}BCI_{t-1}+\beta_{23}CCI_{t-1})}{1+\exp(\beta_{20}+\beta_{21}CLI_{t-1}+\beta_{22}BCI_{t-1}+\beta_{23}CCI_{t-1})}.$ The $p$-value of the likelihood ratio test of $\beta_{12}=\beta_{13}=\beta_{22}=\beta_{23}=0$ is 0.0459. From Figure (ref), the smoothed probability of recession is nearly 0.5 during the 1990s recession. Thus, the case of all three leading economic indicators hardly beats the case of CLI only.
\setcounter{equation}{0} \setcounter{theorem}{0} This study shows consistency and asymptotic normality of the MLE in TVTP regime-switching models. The proof follows from the geometrically decaying bound of the mixing rate of $(S|Y,X)$ with a large probability and in $L^p$, which is shown based on the assumption that there is a small probability for the observations to take extreme values. We verify the assumptions to hold in TVTP regime-switching autoregressive models with transition probabilities in diebold1994regime and chang2017new. We examine how well the asymptotic distribution approximates the distribution of the MLE in finite samples. We compare the estimates of the asymptotic variance constructed from the negative of the inverse Hessian matrix and the inverse of the outer product of the score. The simulation results favour the latter. We apply the TVTP regime-switching model to U.S. industrial production data and investigate the performance of three leading economic indicators in identifying business-cycle turning points.
We are grateful to Yoosoon Chang, Yoshihiko Nishiyama, and Joon Y. Park for their guidance in writing the paper. All errors are our own.