The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
111,351 characters
Asymptotic Properties of the Maximum Likelihood Estimator for Markov-switching Observation-driven Models
\maketitle
\begin{abstract}
A Markov-switching observation-driven model is a stochastic process $((S_t,Y_t))_{t \in \mathbb{Z}}$ where $(S_t)_{t \in \mathbb{Z}}$ is an unobserved Markov chain on a finite set and $(Y_t)_{t \in \mathbb{Z}}$ is an observed stochastic process such that the conditional distribution of $Y_t$ given $(Y_\tau)_{\tau \leq t-1}$ and $(S_\tau)_{\tau \leq t}$ depends on $(Y_\tau)_{\tau \leq t-1}$ and $S_t$. In this paper, we prove consistency and asymptotic normality of the maximum likelihood estimator for such model. As a special case, we also give conditions under which the maximum likelihood estimator for the widely applied Markov-switching generalised autoregressive conditional heteroscedasticity model introduced by \cite*{HaasMittnikPaolella2004b} is consistent and asymptotically normal.
\vspace{1.5ex} \\ \noindent
\textbf{Keywords:} State Space Models, Hidden Markov Models, Maximum Likelihood Estimation, Consistency, Asymptotic Normality, Markov-switching Generalised Autoregressive Conditional Heteroscedasticity Models.
\vspace{1.5ex} \\ \noindent
\textbf{JEL Classifications:} C12, C13, C22, C32, C58.
\end{abstract}
\section{Introduction} \label{Introduction}
State space models and their extensions are ubiquitous in economics and finance. A state space model is a stochastic process $((X_t,Y_t))_{t \in \mathbb{Z}}$ where $(X_t)_{t \in \mathbb{Z}}$ is an unobserved Markov process taking values in $\textup{X}$ and $(Y_t)_{t \in \mathbb{Z}}$ is an observed stochastic process taking values in $\textup{Y}$ such that the conditional distribution of $Y_t$ given $\mathbf{Y}_{-\infty}^{t-1}$ where $\mathbf{Y}_{i}^{j} := (Y_{i},...,Y_{j})$ and $\mathbf{X}_{-\infty}^{t}$ where $\mathbf{X}_{i}^{j} := (X_i,...,X_j)$ depends only on $X_t$. It is called a hidden Markov model when $X_t = S_t$ and $(S_t)_{t \in \mathbb{Z}}$ is a Markov chain on a finite set.
The arguably most well-known extension of the state space model is the autoregressive state space model of order $p \in \mathbb{N}$ in which the conditional distribution of $Y_t$ given $\mathbf{Y}_{-\infty}^{t-1}$ and $\mathbf{X}_{-\infty}^{t}$ depends on both $\mathbf{Y}_{t-p}^{t-1}$ and $X_t$. An example of an autoregressive state space model is the seminal Markov-switching autoregressive model introduced by \cite{Hamilton1989} to model economic growth where $\textup{X}$ is finite. Another example is the Markov-switching autoregressive conditional heteroscedasticity (ARCH) model introduced independently by \cite{Cai1994} and \cite{HamiltonSusmel1994} to model financial returns where $\textup{X}$ is also finite. See, for instance, \cite{Hamilton2010} and \cite{AngTimmermann2012} for more examples of autoregressive state space models in economics and finance, respectively.
Another extension of the state space model that has gained popularity recently is the observation-driven state space model in which the conditional distribution of $Y_t$ given $\mathbf{Y}_{-\infty}^{t-1}$ and $\mathbf{X}_{-\infty}^{t}$ now depends on both $\mathbf{Y}_{-\infty}^{t-1}$ and $X_t$. An example of an observation-driven state space model is the Markov-switching generalised ARCH (GARCH) model introduced by \cite{HaasMittnikPaolella2004b} where $\textup{X}$ is finite.\footnote{Note that there exist two types of Markov-switching GARCH models namely the one considered by \cite{FrancqRoussignolZakoian2001} in which the conditional distribution of $Y_t$ given $\mathbf{Y}_{-\infty}^{t-1}$ and $\mathbf{X}_{-\infty}^{t}$ depends on both $\mathbf{Y}_{-\infty}^{t-1}$ and $\mathbf{X}_{-\infty}^{t}$ and the one by \cite{HaasMittnikPaolella2004b} just discussed. See Section \ref{sec:terminology} for details.} Yet another example from finance is the score-driven state space model introduced by \cite{MonachePetrellaVenditti2021} where $\textup{X}$ is not finite. For examples of observation-driven state space models in economics, see, for instance, \cite{MonachePetrellaVenditti2016} and \cite{AngeliniGorgi2018} where $\textup{X}$ is not finite. For more examples of observation-driven state space models in finance, see, for instance, \cite{HaasMittnikPaolella2004a}, \cite{Ardia2009}, \cite{BrodaHaasKrausePaolellaSteude2013}, \cite{ArdiaBluteauBoudtCatania2018}, \cite{HaasLiu2018}, \cite{BernardiCatania2019}, and \cite{Walden2019} where $\textup{X}$ is finite or \cite{BuccheriBormettiCorsiLillo2021} and \cite{BuccheriCorsi2021} where $\textup{X}$ is not finite.
Statistical inference for state space models and their extensions – including estimation, which is typically done by maximum likelihood estimation – is therefore of significant practical importance. The asymptotic properties of the maximum likelihood estimator (MLE) for observation-driven state space models have, however, attracted almost no attention in the literature despite the increasing popularity of such models. In this paper, we prove both consistency (Theorem \ref{TheoremConsistency}) and asymptotic normality (Theorem \ref{TheoremAsymptoticNormality}) of the MLE for an observation-driven state space model where $\textup{X}$ is finite, called a Markov-switching observation-driven model. To the best of our knowledge, these results are the first of their kind in the literature. As a special case, we also give conditions under which the MLE for the widely applied Markov-switching GARCH model by \cite{HaasMittnikPaolella2004b} is both consistent (Theorem \ref{theo:c}) and asymptotically normal (Theorem \ref{theo:an}). Again, this is new to the literature and extends \cite{KandjiMisko2024}, who gave conditions under which it is only consistent.
In contrast, the asymptotic properties of the MLE for autoregressive state space models, that is, the models in which the conditional distribution of $Y_t$ given $\mathbf{Y}_{-\infty}^{t-1}$ and $\mathbf{X}_{-\infty}^{t}$ depends \textit{only} on $\mathbf{Y}_{t-p}^{t-1}$ and $X_t$, have attracted much attention in the literature. For autoregressive state space models where $\textup{X}$ is finite, consistency of the MLE was proved by \cite{FrancqRoussignol1998}, \cite{KrishnamurthyRyden1998}, and \cite{FrancqRoussignolZakoian2001}. This was later generalised by \cite{DoucMoulinesRyden2004}, who proved consistency and asymptotic normality of the MLE for autoregressive state space models where $\textup{X}$ is compact and not necessarily finite, a seminal result which can be used to give conditions under which both the MLE for the Markov-switching autoregressive model and the Markov-switching ARCH model is consistent and asymptotically normal. More recently, \cite{KasaharaShimotsu2019} relax some of the assumptions in \cite{DoucMoulinesRyden2004}.\footnote{Consistency and asymptotic normality of the MLE for hidden Markov models was proved by \cite{Leroux1992} and \cite{BickelRitovRyden1998}, respectively. Local consistency and asymptotic normality of the MLE for state space models where $\textup{X}$ is compact was proved by \cite{JensenPetersen1999}, and global consistency of the MLE for general state space models was proved by \cite{DoucMoulinesOlssonVanHandel2011}. See \cite{DoucMoulinesOlssonVanHandel2011} for more references on the asymptotic properties of the MLE for state space models.} In comparison to \cite{DoucMoulinesRyden2004} and \cite{KasaharaShimotsu2019}, we prove consistency and asymptotic normality of the MLE for models in which the conditional distribution of $Y_t$ given $\mathbf{Y}_{-\infty}^{t-1}$ and $\mathbf{X}_{-\infty}^{t}$ depends on \textit{both} $\mathbf{Y}_{-\infty}^{t-1}$ and $X_t$, but with $\textup{X}$ finite. This result can, in contrast to the ones in \cite{DoucMoulinesRyden2004} and \cite{KasaharaShimotsu2019}, be used to give conditions under which also the MLE for the widely applied Markov-switching GARCH model by \cite{HaasMittnikPaolella2004b} is consistent and asymptotically normal.
In the proofs of consistency and asymptotic normality of the MLE for the model, the fact that the time-varying parameters and the filter forget their initialisations asymptotically (Lemmas \ref{LemmaInvertibilityX} and \ref{LemmaInvertibilityFilter}, respectively) is crucial, similarly to \cite{DoucMoulinesRyden2004} (Corollary 1) and \cite{KasaharaShimotsu2019} (Lemma 1). However, differently from \cite{DoucMoulinesRyden2004} and \cite{KasaharaShimotsu2019}, who use theory for Markov chains, we prove these results using theory for stochastic difference equations since the latter gives lower-level conditions under which the MLE for the model is consistent and asymptotically normal, which are easier to verify. Theory for stochastic difference equations is also usually used for standard observation-driven models, see, for instance, \cite{BerkesHorvathKokoszka2003}, \cite{FrancqZakoian2004}, \cite{StraumannMikosch2006}, \cite{BlasquesGorgiKoopmanWintenberger2018}, and \cite{BlasquesVanBrummelenKoopmanLucas2022}.
The rest of the paper is organised as follows. Section \ref{Model} introduces the Markov-switching observation-driven model, and Section \ref{Examples1} gives some examples of Markov-switching observation-driven models. In Section \ref{SE}, the probabilistic properties of the model is studied. The asymptotic properties of the MLE for the model is then studied in Section \ref{CAN}. Section \ref{Examples2} studies both the asymptotic and finite-sample properties of the MLE for the Markov-switching GARCH model by \cite{HaasMittnikPaolella2004b}, the latter in a Monte Carlo simulation study. Section \ref{Conclusion} concludes. All proofs except the ones in the main text are collected in the appendix.
\section{The Markov-switching Observation-driven Model} \label{Model}
\subsection{Model}
A Markov-switching observation-driven model is a stochastic process $((S_t,Y_t))_{t \in \mathbb{Z}}$ where $(S_t)_{t \in \mathbb{Z}}$ is an unobserved Markov chain taking values in $\{1,...,J\}$ with transition probabilities
\begin{equation*}
p_{ij} := \mathbb{P}(S_{t+1} = j \mid S_{t} = i), \quad i,j \in \{1,...,J\},
\end{equation*}
and $(Y_t)_{t \in \mathbb{Z}}$ is an observed stochastic process taking values in $\mathcal{Y} \subseteq \mathbb{R}$ such that the conditional distribution of $Y_t$ given $\mathbf{Y}_{-\infty}^{t-1}$ and $\mathbf{S}_{-\infty}^{t}$ depends only on $\mathbf{Y}_{-\infty}^{t-1}$ and $S_t$ as follows
\begin{equation*}
Y_t \mid (\mathbf{Y}_{-\infty}^{t-1},S_t) \sim \mathcal{D}_{S_t} (X_{S_{t},t},\boldsymbol{\upsilon}_{S_t}),
\end{equation*}
where $X_{j,t}, j \in \{1,...,J\}$ is a time-varying parameter taking values in a complete set $\mathcal{X}_j \subseteq \mathbb{R}$ given by
\begin{equation*}
X_{j,t+1} = \phi_{j} (Y_t,X_{j,t};\boldsymbol{\upsilon}_{j}),
\end{equation*}
and $\boldsymbol{\upsilon}_{j}, j \in \{1,...,J\}$ is a vector of constant parameters taking values in a set $\boldsymbol{\Upsilon}_{j} \subseteq \mathbb{R}^{d_{j}}$. If $(S_t)_{t \in \mathbb{Z}}$ is an independent and identically distributed (i.i.d.) chain, that is, if
\begin{equation*}
p_{1j} = \cdots = p_{Jj}
\end{equation*}
for all $j \in \{1,...,J\}$, then the Markov-switching observation-driven model is called a mixture observation-driven model.
In the Markov-switching observation-driven model, filtering, prediction, and smoothing of the unobserved Markov chain $(S_t)_{t \in \mathbb{Z}}$, that is, computation of the conditional distribution
\begin{equation*}
\pi_{j,t \mid s} := \mathbb{P}(S_{t} = j \mid \mathbf{Y}_{-\infty}^{s}), \quad j \in \{1,...,J\},
\end{equation*}
which is called the filtering distribution when $t = s$, the predictive distribution when $t > s$, and the smoothing distribution when $t < s$, is done as in the Markov-switching autoregressive model and the hidden Markov model. The one-step-ahead prediction is given by
\begin{equation*}
\pi_{j,t+1 \mid t} = \sum_{i=1}^{J} p_{ij} \pi_{i,t \mid t},
\end{equation*}
and the filter is given by
\begin{equation*}
\pi_{j,t \mid t} = \frac{\pi_{j,t \mid t-1} f_{j} (Y_t;X_{j,t},\boldsymbol{\upsilon}_{j})}{f(Y_t)},
\end{equation*}
where $f(y),y \in \mathcal{Y}$ is the conditional probability density function (pdf) of $Y_t$ given $\mathbf{Y}_{-\infty}^{t-1}$ given by
\begin{equation*}
f(y) = \sum_{k=1}^{J} \pi_{k,t \mid t-1} f_{k} (y;X_{k,t},\boldsymbol{\upsilon}_{k}),
\end{equation*}
and $f_{j} (y;X_{j,t},\boldsymbol{\upsilon}_{j}),y \in \mathcal{Y}$ is the conditional pdf of $Y_t$ given $\mathbf{Y}_{-\infty}^{t-1}$ and $S_t = j$, see \cite{Hamilton1994} for details.\footnote{More generally, the $h$-step-ahead prediction is given by $\pi_{j,t+h \mid t} = \sum_{i=1}^{J} p_{ij}^{(h)} \pi_{i,t \mid t}$ where $p_{ij}^{(h)} := \mathbb{P}(S_{t+h} = j \mid S_{t} = i)$.} Let $\boldsymbol{\pi}_{t \mid s} := (\pi_{1,t \mid s},...,\pi_{J,t \mid s})^{\prime}$. Then,
\begin{equation*}
\boldsymbol{\pi}_{t+1 \mid t} = \mathbf{P}^{\prime} \boldsymbol{\pi}_{t \mid t},
\end{equation*}
where $\mathbf{P}$ is the transition probability matrix given by
\begin{equation*}
\mathbf{P}
:=
\begin{bmatrix}
p_{11} & \cdots & p_{1J} \\
\vdots & \ddots & \vdots \\
p_{J1} & \cdots & p_{JJ} \\
\end{bmatrix},
\end{equation*}
and
\begin{equation*}
\boldsymbol{\pi}_{t \mid t} = \mathbf{F}_t (\boldsymbol{\pi}_{t \mid t-1}) \boldsymbol{\pi}_{t \mid t-1},
\end{equation*}
where $\mathbf{F}_t(\boldsymbol{\pi}_{t \mid t-1})$ is a diagonal matrix with generic element
\begin{equation*}
[\mathbf{F}_t (\boldsymbol{\pi}_{t \mid t-1})]_{ii} := \frac{f_{i} (Y_t;X_{i,t},\boldsymbol{\upsilon}_{i})}{\sum_{k=1}^{J} \pi_{k,t \mid t-1} f_{k} (Y_t;X_{k,t},\boldsymbol{\upsilon}_{k})}, \quad i \in \{1,...,J\}.
\end{equation*}
Moreover, the smoother is given by
\begin{equation*}
\pi_{j,t \mid T} = \pi_{j,t \mid t} \sum_{i=1}^{J} p_{ji} \frac{\pi_{i,t+1 \mid T}}{\pi_{i,t+1 \mid t}}, \quad t < T,
\end{equation*}
see \cite{Hamilton1994} for details once again. Prediction of the observed stochastic process $(Y_t)_{t \in \mathbb{Z}}$ is also done as in the Markov-switching autoregressive model and the hidden Markov model.
Finally, note that the Markov-switching observation-driven model reduces to the Markov-switching autoregressive model of order $1$ if $\phi_{j} (Y_t,X_{j,t};\boldsymbol{\upsilon}_{j}) = \phi_{j} (Y_t;\boldsymbol{\upsilon}_{j})$ for all $j \in \{1,...,J\}$ and to the hidden Markov model if $X_{j,t} \equiv X_{j}$ for all $j \in \{1,...,J\}$.\footnote{All results continue to hold in the case where $X_{j,t+1} = \phi_{j} (Y_{t},...,Y_{t-p+1},X_{j,t};\boldsymbol{\upsilon}_{j})$ – in which case the Markov-switching observation-driven model reduces to the Markov-switching autoregressive model of order $p \in \mathbb{N}$ if $\phi_{j} (Y_{t},...,Y_{t-p+1},X_{j,t};\boldsymbol{\upsilon}_{j}) = \phi_{j} (Y_{t},...,Y_{t-p+1};\boldsymbol{\upsilon}_{j})$ for all $j \in \{1,...,J\}$ – but, to ease the exposition, we consider the case where $X_{j,t+1} = \phi_{j} (Y_{t},X_{j,t};\boldsymbol{\upsilon}_{j})$.}
\subsection{Terminology} \label{sec:terminology}
A word on terminology. Let $(S_t)_{t \in \mathbb{Z}}$ be an unobserved Markov chain. In the previous section, we defined a Markov-switching observation-driven model as a stochastic process $((S_t,Y_t))_{t \in \mathbb{Z}}$ where $(Y_t)_{t \in \mathbb{Z}}$ is an observed stochastic process such that the conditional distribution of $Y_t$ given $\mathbf{Y}_{-\infty}^{t-1}$ and $\mathbf{S}_{-\infty}^{t}$ depends only on $\mathbf{Y}_{-\infty}^{t-1}$ and $S_t$ as follows
\begin{equation}
Y_t \mid (\mathbf{Y}_{-\infty}^{t-1},S_t) \sim \mathcal{D}_{S_t} (X_{S_{t},t},\boldsymbol{\upsilon}_{S_t}), \label{eq:MSODM1}
\end{equation}
with
\begin{equation}
X_{S_t,t} = \phi_{S_t} (Y_{t-1},X_{S_t,t-1};\boldsymbol{\upsilon}_{S_t}); \label{eq:MSODM2}
\end{equation}
its dependence structure is illustrated in Figure \ref{fig:1}. As mentioned in the introduction, the Markov-switching GARCH model by \cite{HaasMittnikPaolella2004b} is an example of such a model, see the next section for details. In such a Markov-switching observation-driven model, filtering, prediction, and smoothing of the unobserved Markov chain $(S_t)_{t \in \mathbb{Z}}$ and thus prediction of the observed stochastic process $(Y_t)_{t \in \mathbb{Z}}$ can be done via the formulae in the previous section.
\begin{figure}[htbp]
\centering
\begin{tikzpicture}[minimum size = 1.25cm]
\node[left] at (-3,-1.5) {$\dots$};
\node[draw, circle] (S0) at (0,-1.5) {$S_{t-1}$};
\node[draw, circle] (S1) at (3,-1.5) {$S_t$};
\node[right] at (6,-1.5) {$\dots$};
\node[left] at (-3,1.5) {$\dots$};
\node[draw, circle] (Y0) at (0,1.5) {$Y_{t-1}$};
\node[draw, circle] (Y1) at (3,1.5) {$Y_t$};
\node[right] at (6,1.5) {$\dots$};
\draw[->] (-3,-1.5) -- (S0);
\draw[->] (S0) -- (S1);
\draw[->] (S1) -- (6,-1.5);
\draw[->] (-3,1.5) -- (Y0);
\draw[->] (-3,1.5) to [out=30, in=150] (Y1);
\draw[->] (-3,1.5) to[out=30, in=150] (6,1.5);
\draw[->] (Y0) -- (Y1);
\draw[->] (Y0) to[out=30, in=150] (6,1.5);
\draw[->] (Y1) -- (6,1.5);
\draw[->] (S0) -- (Y0);
\draw[->] (S1) -- (Y1);
\end{tikzpicture}
\captionsetup{font=footnotesize}
\caption{The dependence structure of the Markov-switching observation-driven model in Equations \eqref{eq:MSODM1} and \eqref{eq:MSODM2}.}
\label{fig:1}
\end{figure}
One could, however, also have defined a Markov-switching observation-driven model as a stochastic process $((S_t,Y_t))_{t \in \mathbb{Z}}$ where $(Y_t)_{t \in \mathbb{Z}}$ is an observed stochastic process such that the conditional distribution of $Y_t$ given $\mathbf{Y}_{-\infty}^{t-1}$ and $\mathbf{S}_{-\infty}^{t}$ depends on both $\mathbf{Y}_{-\infty}^{t-1}$ and $\mathbf{S}_{-\infty}^{t}$ as follows
\begin{equation}
Y_t \mid (\mathbf{Y}_{-\infty}^{t-1},\mathbf{S}_{-\infty}^{t}) \sim \mathcal{D}_{S_t} (X_{S_{t},t},\boldsymbol{\upsilon}_{S_t}), \label{eq:MSODM3}
\end{equation}
with
\begin{equation}
X_{S_t,t} = \phi_{S_t} (Y_{t-1},X_{S_{t-1},t-1};\boldsymbol{\upsilon}_{S_t}). \label{eq:MSODM4}
\end{equation}
The Markov-switching GARCH model considered by \cite{FrancqRoussignolZakoian2001} is an example of such a model. The dependence structure of such a Markov-switching observation-driven model is illustrated in Figure \ref{fig:2}. In such a model, filtering, prediction, and smoothing of the unobserved Markov chain $(S_t)_{t \in \mathbb{Z}}$ can, however, not be done due to the so-called path-dependence problem, a problem which makes such a model complicated, if not close to impossible, to handle. We stress that we consider the former, and \textit{not} the latter, Markov-switching observation-driven model.
\begin{figure}[htbp]
\centering
\begin{tikzpicture}[minimum size = 1.25cm]
\node[left] at (-3,-1.5) {$\dots$};
\node[draw, circle] (S0) at (0,-1.5) {$S_{t-1}$};
\node[draw, circle] (S1) at (3,-1.5) {$S_t$};
\node[right] at (6,-1.5) {$\dots$};
\node[left] at (-3,1.5) {$\dots$};
\node[draw, circle] (Y0) at (0,1.5) {$Y_{t-1}$};
\node[draw, circle] (Y1) at (3,1.5) {$Y_t$};
\node[right] at (6,1.5) {$\dots$};
\draw[->] (-3,-1.5) -- (S0);
\draw[->] (S0) -- (S1);
\draw[->] (S1) -- (6,-1.5);
\draw[->] (-3,1.5) -- (Y0);
\draw[->] (-3,1.5) to [out=30, in=150] (Y1);
\draw[->] (-3,1.5) to[out=30, in=150] (6,1.5);
\draw[->] (Y0) -- (Y1);
\draw[->] (Y0) to[out=30, in=150] (6,1.5);
\draw[->] (Y1) -- (6,1.5);
\draw[->] (-3,-1.5) -- (Y0);
\draw[->] (-3,-1.5) -- (Y1);
\draw[->] (-3,-1.5) -- (6,1.5);
\draw[->] (S0) -- (Y0);
\draw[->] (S0) -- (Y1);
\draw[->] (S0) -- (6,1.5);
\draw[->] (S1) -- (Y1);
\draw[->] (S1) -- (6,1.5);
\end{tikzpicture}
\captionsetup{font=footnotesize}
\caption{The dependence structure of the Markov-switching observation-driven model in Equations \eqref{eq:MSODM3} and \eqref{eq:MSODM4}.}
\label{fig:2}
\end{figure}
\section{Examples of Markov-switching Observation-driven Models} \label{Examples1}
In this section, we give some examples of Markov-switching observation-driven models.
\begin{example} \label{Example1}
An example of a Markov-switching observation-driven model is
\begin{equation}
Y_t = X_{S_t,t} + \sigma_{S_t} \varepsilon_t, \label{ARINF}
\end{equation}
where $(\varepsilon_t)_{t \in \mathbb{Z}}$ is a sequence of independent standard normal distributed random variables independent of $(S_t)_{t \in \mathbb{Z}}$ and, for each $j \in \{1,...,J\}$,
\begin{equation*}
X_{j,t+1} = \omega_{j} + \alpha_{j} Y_{t} + \beta_{j} X_{j,t},
\end{equation*}
where $\omega_{j} \in \mathbb{R}$, $\alpha_{j} \in \mathbb{R}$, $\beta_{j} \in \mathbb{R}$, and $\sigma_j^2 > 0$. Here, $\mathcal{Y} = \mathbb{R}$, $\mathcal{D}_{S_t}$ is the normal distribution with mean $X_{S_t,t}$ and variance $\sigma_{S_t}^{2}$, $\mathcal{X}_j = \mathbb{R}$, $\phi_{j} (y,x_j;\boldsymbol{\upsilon}_j) = [\boldsymbol{\upsilon}_j]_1 + [\boldsymbol{\upsilon}_j]_2 y + [\boldsymbol{\upsilon}_j]_3 x_j$, $\boldsymbol{\upsilon}_{j} = (\omega_{j},\alpha_{j},\beta_{j},\sigma_{j}^{2})^{\prime}$, and $\boldsymbol{\Upsilon}_{j} = \mathbb{R} \times \mathbb{R} \times \mathbb{R} \times (0,\infty)$.
A related model is the Markov-switching autoregressive model of order $p \in \mathbb{N}$ by \citet{Hamilton1989}. This model is given by
\begin{equation*}
Y_t = a_{S_t} + \sum_{i=1}^{p} b_{S_t}^{(i)} Y_{t-i} + \sigma_{S_t} \varepsilon_t,
\end{equation*}
where $(\varepsilon_t)_{t \in \mathbb{Z}}$ is a sequence of independent standard normal distributed random variables independent of $(S_t)_{t \in \mathbb{Z}}$ as above.
It can be shown that if $(Y_t)_{t \in \mathbb{Z}}$ is stationary and ergodic with $\mathbb{E} [\log^{+}|Y_t|] < \infty$ where $\log^{+} x := \max(\log x,0)$ for all $x > 0$ and $|\beta_j| < 1$ for all $j \in \{1,...,J\}$, then
\begin{equation*}
X_{j,t} = \frac{\omega_j}{1-\beta_j} + \sum_{i=1}^{\infty} \alpha_j \beta_{j}^{i-1} Y_{t-i}
\end{equation*}
for all $j \in \{1,...,J\}$. The Markov-switching observation-driven model in Equation \eqref{ARINF} can thus be thought of as a Markov-switching autoregressive model of order infinity given by
\begin{equation*}
Y_t = a_{S_t} + \sum_{i=1}^{\infty} b_{S_t}^{(i)} Y_{t-i} + \sigma_{S_t} \varepsilon_t,
\end{equation*}
where
\begin{equation*}
a_{S_t} := \frac{\omega_{S_t}}{1-\beta_{S_t}} \quad \text{and} \quad b_{S_t}^{(i)} := \alpha_{S_t} \beta_{S_t}^{i-1}.
\end{equation*}
\end{example}
\begin{example}
The Markov-switching GARCH model by \citet{HaasMittnikPaolella2004b} is given by
\begin{equation*}
Y_t = \sqrt{X_{S_t,t}} \varepsilon_t,
\end{equation*}
where $(\varepsilon_t)_{t \in \mathbb{Z}}$ is a sequence of independent standard normal distributed random variables independent of $(S_t)_{t \in \mathbb{Z}}$ and, for each $j \in \{1,...,J\}$,
\begin{equation*}
X_{j,t+1} = \omega_{j} + \alpha_{j} Y_{t}^2 + \beta_{j} X_{j,t},
\end{equation*}
where $\omega_{j} > 0$, $\alpha_{j} \geq 0$, and $\beta_{j} \geq 0$. This is also an example of a Markov-switching observation-driven model where $\mathcal{Y} = \mathbb{R}$, $\mathcal{D}_{S_t}$ is the normal distribution with mean zero and variance $X_{S_t,t}$, $\mathcal{X}_j = \left[ 0, \infty \right)$, $\phi_{j} (y,x_j;\boldsymbol{\upsilon}_j) = [\boldsymbol{\upsilon}_j]_1 + [\boldsymbol{\upsilon}_j]_2 y^2 + [\boldsymbol{\upsilon}_j]_3 x_j$, $\boldsymbol{\upsilon}_{j} = (\omega_{j},\alpha_{j},\beta_{j})^{\prime}$, and $\boldsymbol{\Upsilon}_{j} = [0,\infty) \times [0,\infty) \times [0,\infty)$. Note that it reduces to the mixture GARCH model by \citet{HaasMittnikPaolella2004a} if $(S_t)_{t \in \mathbb{Z}}$ is an i.i.d. chain.
The Markov-switching GARCH model is not the only Markov-switching conditional heteroscedasticity model that the Markov-switching observation-driven model nests. Consider, for instance, the Markov-switching asymmetric power GARCH model given by
\begin{equation*}
Y_t = X_{S_t,t}^{\frac{1}{\delta}} \varepsilon_t, \quad \delta > 0,
\end{equation*}
where $(\varepsilon_t)_{t \in \mathbb{Z}}$ is as above and, for each $j \in \{1,...,J\}$,
\begin{equation*}
X_{j,t+1} = \omega_{j} + \alpha_{j} \left( \left| Y_{t} \right| - \xi_j Y_{t} \right)^{\delta} + \beta_{j} X_{j,t},
\end{equation*}
where $\omega_{j} > 0$, $\alpha_{j} \geq 0$, $|\xi_j| \leq 1$, and $\beta_{j} \geq 0$, which reduces to the Markov-switching GJR-GARCH model if $\delta = 2$, the Markov-switching threshold GARCH model if $\delta = 1$, and the standard Markov-switching GARCH model above if $\delta = 2$ and $\xi_j = 0$ for all $j \in \{1,...,J\}$. This is also an example of a Markov-switching observation-driven model where $\mathcal{Y} = \mathbb{R}$, $\mathcal{D}_{S_t}$ is the normal distribution with mean zero and variance $X_{S_t,t}^{\frac{2}{\delta}}$, $\mathcal{X}_j = \left[ 0, \infty \right)$, $\phi_{j} (y,x_j;\boldsymbol{\upsilon}_j) = [\boldsymbol{\upsilon}_j]_1 + [\boldsymbol{\upsilon}_j]_2 (|y| - [\boldsymbol{\upsilon}_j]_3 y)^{\delta} + [\boldsymbol{\upsilon}_j]_4 x_j$, $\boldsymbol{\upsilon}_{j} = (\omega_{j},\alpha_{j},\xi_j,\beta_{j})^{\prime}$, and $\boldsymbol{\Upsilon}_{j} = [0,\infty) \times [0,\infty) \times [-1,1] \times [0,\infty)$.
\end{example}
\begin{example}
Let $y \mapsto F(y;x,\tilde{\boldsymbol{\upsilon}})$ be a cumulative distribution function (cdf) with support $\mathcal{N} \subseteq [0,\infty)$, for instance, $\mathcal{N} = \{0,1\}$, $\mathcal{N} = \mathbb{N}$, or $\mathcal{N} = [0,1]$, indexed by the mean $x$ and a vector of parameters $\tilde{\boldsymbol{\upsilon}}$, and assume that
\begin{equation*}
x \leq x^{*} \quad \Rightarrow \quad F^{-} (u;x,\tilde{\boldsymbol{\upsilon}}) \leq F^{-} (u;x^{*},\tilde{\boldsymbol{\upsilon}})
\end{equation*}
for all $u \in (0,1)$ where $F^{-} (u;x,\tilde{\boldsymbol{\upsilon}}) := \inf \{ y \in \mathcal{N} : F(y;x,\tilde{\boldsymbol{\upsilon}}) \geq u\}$.
The (present-regime dependent) Markov-switching positive linear conditional mean model by \citet{AknoucheFrancq2022} is given by
\begin{equation*}
Y_t = F^{-}_{S_t} (U_t;X_{S_t,t},\tilde{\boldsymbol{\upsilon}}_{S_t}),
\end{equation*}
where $(U_t)_{t \in \mathbb{Z}}$ is a sequence of independent uniform distributed random variables on $[0,1]$ independent of $(S_t)_{t \in \mathbb{Z}}$ and, for each $j \in \{1,...,J\}$,
\begin{equation*}
X_{j,t+1} = \omega_{j} + \alpha_{j} Y_{t} + \beta_{j} X_{j,t},
\end{equation*}
where $\omega_{j} > 0$, $\alpha_{j} \geq 0$, $\beta_{j} \geq 0$, and $\tilde{\boldsymbol{\upsilon}}_{j} \in \tilde{\boldsymbol{\Upsilon}}_{j}$. This is another example of a Markov-switching observation-driven model where $\mathcal{Y} = \mathcal{N}$, $F_{S_t}$ is the cdf of $\mathcal{D}_{S_t}$, $\mathcal{X}_j = \left[ 0, \infty \right)$, $\phi_{j} (y,x_j;\boldsymbol{\upsilon}_j) = [\boldsymbol{\upsilon}_j]_1 + [\boldsymbol{\upsilon}_j]_2 y + [\boldsymbol{\upsilon}_j]_3 x_j$, $\boldsymbol{\upsilon}_{j} = (\omega_{j},\alpha_{j},\beta_{j},\tilde{\boldsymbol{\upsilon}}_{j})^{\prime}$, and $\boldsymbol{\Upsilon}_{j} = [0,\infty) \times [0,\infty) \times [0,\infty) \times \tilde{\boldsymbol{\Upsilon}}_{j}$.
\end{example}
\begin{example}
A random variable $Y$ with support $[-\pi,\pi]$ is said to be von Mises distributed with location $x \in [-\pi,\pi]$ and concentration $\tilde{\upsilon} > 0$ if the pdf of $Y$ is given by
\begin{equation*}
f(y;x,\tilde{\upsilon}) = \frac{1}{2 \pi I_0(\tilde{\upsilon})} \exp(\tilde{\upsilon} \cos(y - x)), \quad y \in [-\pi,\pi],
\end{equation*}
where $I_0(\cdot)$ is the modified Bessel function of the first kind of order 0.
Yet another example of a Markov-switching observation-driven model is the Markov-switching score-driven model
\begin{equation*}
Y_t = \left( X_{S_t,t} + \varepsilon_{S_t,t} \right) \textup{mod} \, (2 \pi) - \pi,
\end{equation*}
where, for each $j \in \{1,...,J\}$, $(\varepsilon_{j,t})_{t \in \mathbb{Z}}$ is a sequence of independent von Mises distributed random variables with location $0$ and concentration $\tilde{\upsilon}_j > 0$, which is independent of $(\varepsilon_{i,t})_{t \in \mathbb{Z}}$ for all $i \in \{1,...,J\}$ such that $i \neq j$ and $(S_t)_{t \in \mathbb{Z}}$, and
\begin{equation*}
X_{j,t+1} = \omega_j + \alpha_j \tilde{\upsilon}_j \sin \left(Y_t - X_{j,t}\right) + \beta_j X_{j,t},
\end{equation*}
where $\omega_j \in \mathbb{R}$, $\alpha_j \in \mathbb{R}$, and $\beta_j \in \mathbb{R}$. The model is similar to the regime-switching score-driven model by \cite{HarveyPalumbo2023}, which is a generalisation of the score-driven model by \cite{HarveyHurnPalumboThiele2024}. Indeed, the Markov-switching score-driven model is a Markov-switching observation-driven model where $\mathcal{Y} = [-\pi,\pi]$, $\mathcal{D}_{S_t}$ is the von Mises distribution with location $X_{S_t,t}$ and concentration $\tilde{\upsilon}_{S_t}$, $\mathcal{X}_j = \mathbb{R}$, $\phi_{j} (y,x_j;\boldsymbol{\upsilon}_j) = [\boldsymbol{\upsilon}_j]_1 + [\boldsymbol{\upsilon}_j]_2 [\boldsymbol{\upsilon}_j]_4 \sin (y-x_j) + [\boldsymbol{\upsilon}_j]_3 x_j$, $\boldsymbol{\upsilon}_{j} = (\omega_{j},\alpha_{j},\beta_{j},\tilde{\upsilon}_{j})^{\prime}$, and $\boldsymbol{\Upsilon}_{j} = \mathbb{R} \times \mathbb{R} \times \mathbb{R} \times (0,\infty)$.
\end{example}
\section{Probabilistic Properties of the Model} \label{SE}
First, we study the probabilistic properties of the Markov-switching observation-driven model. We restrict our attention to the Markov-switching observation-driven models that can be written as
\begin{equation}
Y_t = \mathbf{1}_{S_t} \mathbf{g} (\boldsymbol{\varepsilon}_t;\mathbf{X}_t,\boldsymbol{\upsilon}). \label{Y}
\end{equation}
Here, $\mathbf{1}_{S_t} := (1_{\{S_t = 1\}},...,1_{\{S_t = J\}})$ where $(S_t)_{t \in \mathbb{Z}}$ is a stationary, irreducible, and aperiodic (thus ergodic) Markov chain taking values in $\{1,...,J\}$ with transition probabilities $p_{ij}$. Moreover, $\mathbf{g} (\boldsymbol{\varepsilon}_t;\mathbf{X}_t,\boldsymbol{\upsilon}) := (g_1(\varepsilon_{1,t};X_{1,t},\boldsymbol{\upsilon}_{1}),...,g_J(\varepsilon_{J,t};X_{J,t},\boldsymbol{\upsilon}_{J}))^{\prime}$ where $\boldsymbol{\varepsilon}_t := (\varepsilon_{1,t},...,\varepsilon_{J,t})^{\prime}$, for all $j \in \{1,...,J\}$, $(\varepsilon_{j,t})_{t \in \mathbb{Z}}$ is a sequence of i.i.d. random variables taking values in $\mathcal{E}_j$ with distribution $\mathcal{D}_{j}^{\varepsilon}(\boldsymbol{\upsilon}_{j})$, which is independent of $(\varepsilon_{i,t})_{t \in \mathbb{Z}}$ for all $i \in \{1,...,J\}$ such that $i \neq j$, $\mathbf{X}_t := (X_{1,t},...,X_{J,t})^{\prime}$ is given by
\begin{equation*}
\mathbf{X}_{t+1} = \boldsymbol{\phi}(S_t,\boldsymbol{\varepsilon}_t,\mathbf{X}_{t};\boldsymbol{\upsilon})
\end{equation*}
with
\begin{equation*}
[\boldsymbol{\phi}(S_t,\boldsymbol{\varepsilon}_t,\mathbf{X}_{t};\boldsymbol{\upsilon})]_j := \phi_{j} (\mathbf{1}_{S_t} \mathbf{g} (\boldsymbol{\varepsilon}_t;\mathbf{X}_t,\boldsymbol{\upsilon}),X_{j,t};\boldsymbol{\upsilon}_{j}), \quad j \in \{1,...,J\},
\end{equation*}
and $\boldsymbol{\upsilon} := (\boldsymbol{\upsilon}_{1},...,\boldsymbol{\upsilon}_{J})^{\prime}$. Finally, $(S_t)_{t \in \mathbb{Z}}$ and $(\varepsilon_{j,t})_{t \in \mathbb{Z}}$ are independent for all $j \in \{1,...,J\}$.\footnote{All examples in Section \ref{Examples1} can be written like this because if $\varepsilon_{i,t} \overset{d}{=} \varepsilon_{j,t}$ for all $i,j \in \{1,...,J\}$, then let $\varepsilon_{j,t} = \varepsilon_t$ for all $j \in \{1,...,J\}$ where $(\varepsilon_{t})_{t \in \mathbb{Z}}$ is a sequence of i.i.d. random variables taking values in $\mathcal{E}$ with distribution $\mathcal{D}^{\varepsilon}(\boldsymbol{\upsilon})$.}
\subsection{Stationarity and Ergodicity} \label{StationarityAndErgodicity}
Theorem \ref{TheoremStationarityAndErgodicity}, which follows from an application of Theorem 3.1 in \citet{Bougerol1993}, gives conditions under which the model is stationary and ergodic. Let
\begin{equation*}
\Lambda (\boldsymbol{\phi}_t) := \sup_{\underset{\mathbf{x} \neq \mathbf{y}}{\mathbf{x},\mathbf{y} \in \mathcal{X}_{1} \times \cdots \times \mathcal{X}_{J}}} \frac{|| \boldsymbol{\phi}_t (\mathbf{x}) - \boldsymbol{\phi}_t (\mathbf{y}) ||_{2}}{|| \mathbf{x} - \mathbf{y} ||_{2}},
\end{equation*}
where $\boldsymbol{\phi}_t (\mathbf{x}) := \boldsymbol{\phi}(S_t,\boldsymbol{\varepsilon}_t,\mathbf{x};\boldsymbol{\upsilon})$; here and in the following, $|| \mathbf{x} ||_{p} := (\sum_{i=1}^{n} | x_i |^p)^{1/p}, \mathbf{x} \in \mathbb{R}^{n}$.
\begin{theorem} \label{TheoremStationarityAndErgodicity}
Assume that
\begin{enumerate}[(i)]
\item there exists an $\mathbf{x} \in \mathcal{X}_{1} \times \cdots \times \mathcal{X}_{J}$ such that $\mathbb{E} [\log^{+} || \boldsymbol{\phi}_t (\mathbf{x}) - \mathbf{x} ||_{2}] < \infty$,
\item $\mathbb{E} [\log^{+} \Lambda (\boldsymbol{\phi}_t) ] < \infty$, and
\item there exists an $r \in \mathbb{N}$ such that
\begin{equation*}
- \infty \leq \mathbb{E} \left[ \log \Lambda \left( \boldsymbol{\phi}_t^{(r)} \right) \right] < 0,
\end{equation*}
where $\boldsymbol{\phi}_t^{(r)} (\mathbf{x}) := \boldsymbol{\phi}_t \circ \cdots \circ \boldsymbol{\phi}_{t-r+1} (\mathbf{x})$.
\end{enumerate}
Then, $(Y_t)_{t \in \mathbb{Z}}$ is stationary and ergodic.
\end{theorem}
\noindent
Conditions (i) and (ii) are standard regularity conditions, and Condition (iii) is a standard contraction condition – not in the strict sense, but in expectation.
\section{Asymptotic Properties of the Maximum Likelihood Estimator} \label{CAN}
We now study the asymptotic properties of the MLE for the Markov-switching observation-driven model discussed in the previous section.
Assume that a sample $(Y_t)_{t=1}^{T}$ from the Markov-switching observation-driven model $(Y_t)_{t \in \mathbb{Z}}$ given by Equation \eqref{Y} with $\boldsymbol{\theta} = \boldsymbol{\theta}_0$ is observed. Here,
\begin{equation*}
\boldsymbol{\theta} := (p_{ij}, i = 1,...,J, j = 1,...,J-1, \boldsymbol{\upsilon}_j, j = 1,...,J)^{\prime}
\end{equation*}
is the parameter vector since $\sum_{j=1}^{J} p_{ij} = 1$ for all $i \in \{1,...,J\}$ and
\begin{equation*}
\boldsymbol{\Theta} \subset \left\lbrace \boldsymbol{\theta} \in \mathbb{R}^d : p_{ij} > 0, i = 1,...,J, j = 1,...,J-1, \sum_{j=1}^{J-1} p_{ij} < 1, i = 1,...,J, \boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_{j}, j = 1,...,J \right\rbrace
\end{equation*}
with $d := J(J-1) + \sum_{j=1}^{J} d_j$ is the parameter space. The MLE $\hat{\boldsymbol{\theta}}_T$ of $\boldsymbol{\theta}_0$ is given by
\begin{equation*}
\hat{\boldsymbol{\theta}}_T = \underset{\boldsymbol{\theta} \in \boldsymbol{\Theta}}{\arg \max} \, \hat{L}_T (\boldsymbol{\theta}).
\end{equation*}
Here, $\hat{L}_T (\boldsymbol{\theta})$ is the \textit{initialised} and thus non-stationary log-likelihood function given by
\begin{equation*}
\hat{L}_T (\boldsymbol{\theta}) = \frac{1}{T} \sum_{t=1}^{T} \log \hat{f} (Y_t;\boldsymbol{\theta})
\end{equation*}
with
\begin{equation*}
\hat{f} (Y_t;\boldsymbol{\theta}) = \sum_{j=1}^{J} \hat{\pi}_{j,t \mid t-1} (\boldsymbol{\theta}) f_{j} (Y_t;\hat{X}_{j,t} (\boldsymbol{\upsilon}_j),\boldsymbol{\upsilon}_j),
\end{equation*}
where, for each $j \in \{1,...,J\}$, $(\hat{X}_{j,t} (\boldsymbol{\upsilon}_j))_{t \in \mathbb{N}}$ is given by
\begin{equation*}
\hat{X}_{j,t+1} (\boldsymbol{\upsilon}_j) = \phi_j (Y_t,\hat{X}_{j,t} (\boldsymbol{\upsilon}_j);\boldsymbol{\upsilon}_{j})
\end{equation*}
for some initialisation $\hat{X}_{j,1} (\boldsymbol{\upsilon}_j) \in \mathcal{X}_{j}$ and $(\hat{\boldsymbol{\pi}}_{t \mid t-1} (\boldsymbol{\theta}))_{t \in \mathbb{N}}$ is given by
\begin{equation*}
\hat{\boldsymbol{\pi}}_{t+1 \mid t} (\boldsymbol{\theta}) = \mathbf{P}^{\prime} \hat{\boldsymbol{\pi}}_{t \mid t} (\boldsymbol{\theta})
\end{equation*}
with
\begin{equation*}
\hat{\boldsymbol{\pi}}_{t \mid t} (\boldsymbol{\theta}) = \hat{\mathbf{F}}_t (\hat{\boldsymbol{\pi}}_{t \mid t-1} (\boldsymbol{\theta});\boldsymbol{\theta}) \hat{\boldsymbol{\pi}}_{t \mid t-1} (\boldsymbol{\theta})
\end{equation*}
for some initialisation $\hat{\boldsymbol{\pi}}_{0 \mid 0} (\boldsymbol{\theta}) \in \mathcal{S}$ with $\mathcal{S} := \{\mathbf{x} \in \mathbb{R}^{J} : x_j \geq 0, j=1,...,J, \sum_{j=1}^{J} x_j = 1 \}$ where
\begin{equation*}
[\hat{\mathbf{F}}_t (\mathbf{s})]_{ii} := \frac{f_{i} (Y_t;\hat{X}_{i,t} (\boldsymbol{\upsilon}_i),\boldsymbol{\upsilon}_{i})}{\sum_{k=1}^{J} s_{k} f_{k} (Y_t;\hat{X}_{k,t} (\boldsymbol{\upsilon}_k),\boldsymbol{\upsilon}_{k})}, \quad i \in \{1,...,J\}.
\end{equation*}
\subsection{Consistency} \label{Consistency}
Consistency follows from classical arguments if the \textit{initialised} and thus non-stationary log-likelihood function $\hat{L}_T (\boldsymbol{\theta})$ converges uniformly almost surely (a.s.) to a function, say, $L(\boldsymbol{\theta})$, which is uniquely maximised at $\boldsymbol{\theta}_0$. We show that $\hat{L}_T (\boldsymbol{\theta})$ converges uniformly a.s. to $L(\boldsymbol{\theta})$ in two steps, the first step being the main difficulty and hence focus. First, we show that the difference between $\hat{L}_T (\boldsymbol{\theta})$ and the \textit{non-initialised}, stationary, and ergodic log-likelihood function
\begin{equation*}
L_T (\boldsymbol{\theta}) = \frac{1}{T} \sum_{t=1}^{T} \log f (Y_t;\boldsymbol{\theta})
\end{equation*}
with
\begin{equation*}
f (Y_t;\boldsymbol{\theta}) = \sum_{j=1}^{J} \pi_{j,t \mid t-1} (\boldsymbol{\theta}) f_{j} (Y_t;X_{j,t} (\boldsymbol{\upsilon}_j),\boldsymbol{\upsilon}_j),
\end{equation*}
where, for each $j \in \{1,...,J\}$, $(X_{j,t} (\boldsymbol{\upsilon}_j))_{t \in \mathbb{Z}}$ is given in Lemma \ref{LemmaInvertibilityX} and $(\boldsymbol{\pi}_{t \mid t-1} (\boldsymbol{\theta}))_{t \in \mathbb{Z}}$ is given in Corollary \ref{CorollaryInvertibilityPrediction} converges uniformly a.s. to zero by showing that the time-varying parameters and the predictor forget their initialisations asymptotically, that is, that, for each $j \in \{1,...,J\}$, the difference between $(\hat{X}_{j,t} (\boldsymbol{\upsilon}_j))_{t \in \mathbb{N}}$ and $(X_{j,t} (\boldsymbol{\upsilon}_j))_{t \in \mathbb{Z}}$ and the one between $(\hat{\boldsymbol{\pi}}_{t \mid t-1} (\boldsymbol{\theta}))_{t \in \mathbb{N}}$ and $(\boldsymbol{\pi}_{t \mid t-1} (\boldsymbol{\theta}))_{t \in \mathbb{Z}}$ converge uniformly exponentially fast a.s. (e.a.s.) to zero.\footnote{A sequence of random matrices $(\mathbf{Z}_t)_{t \in \mathbb{Z}}$ is said to converge to zero e.a.s. if there exists a $\gamma > 1$ such that $\gamma^t || \mathbf{Z}_t ||_{p,p} \overset{a.s.}{\rightarrow} 0 \quad \text{as} \quad t \rightarrow \infty$ where $|| \mathbf{X} ||_{p,p} := (\sum_{i=1}^{n} \sum_{j=1}^{m} | x_{ij} |^p)^{1/p}, \mathbf{X} \in \mathbb{R}^{n \times m}$.} Then, we show that $L_T (\boldsymbol{\theta})$ converges uniformly a.s. to
\begin{equation*}
L(\boldsymbol{\theta}) = \mathbb{E} [\log f (Y_t;\boldsymbol{\theta})]
\end{equation*}
by using standard arguments. We also show that $L(\boldsymbol{\theta})$ is uniquely maximised at $\boldsymbol{\theta}_0$ by using standard arguments.
We assume the following.
\begin{assumption} \label{AssumptionY}
The conditions in Theorem \ref{TheoremStationarityAndErgodicity} hold for $\boldsymbol{\theta} = \boldsymbol{\theta}_0$.
\end{assumption}
\begin{assumption} \label{AssumptionTheta1}
$\boldsymbol{\Theta}$ is compact.
\end{assumption}
\begin{assumption} \label{AssumptionF1}
For each $j \in \{1,...,J\}$,
\begin{enumerate} [(i)]
\item $(x_j,\boldsymbol{\upsilon}_j) \mapsto f_{j} (y;x_j,\boldsymbol{\upsilon}_j)$ is continuous for all $y \in \mathcal{Y}$ and
\item $x_j \mapsto f_{j} (y;x_j,\boldsymbol{\upsilon}_j)$ is differentiable for all $y \in \mathcal{Y}$ and $\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_{j}$.
\end{enumerate}
\end{assumption}
\begin{assumption} \label{AssumptionPhi1}
For each $j \in \{1,...,J\}$,
\begin{enumerate} [(i)]
\item $(x_j,\boldsymbol{\upsilon}_j) \mapsto \phi_{j} (y,x_j;\boldsymbol{\upsilon}_j)$ is continuous for all $y \in \mathcal{Y}$ and
\item $x_j \mapsto \phi_{j} (y,x_j;\boldsymbol{\upsilon}_j)$ is differentiable for all $y \in \mathcal{Y}$ and $\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_{j}$.
\end{enumerate}
\end{assumption}
\noindent
Assumption \ref{AssumptionY} implies that $(Y_t)_{t \in \mathbb{Z}}$ is stationary and ergodic. Assumption \ref{AssumptionTheta1} is a standard regularity condition. So are Assumptions \ref{AssumptionF1} and \ref{AssumptionPhi1}; Condition (i) in Assumption \ref{AssumptionF1} is similar to Assumption 4 in \cite{DoucMoulinesRyden2004} and Assumption 6(a) in \cite{KasaharaShimotsu2019}, Condition (ii) in Assumption \ref{AssumptionF1} is an extra regularity condition, which is used to prove that the filter/predictor forgets its initialisation asymptotically, and Assumption \ref{AssumptionPhi1} is similar to the assumptions in \cite{BlasquesGorgiKoopmanWintenberger2018}.
The following lemma, which follows from an application of Theorem 3.1 in \citet{Bougerol1993}, gives conditions under which, for each $j \in \{1,...,J\}$, the difference between $(\hat{X}_{j,t} (\boldsymbol{\upsilon}_j))_{t \in \mathbb{N}}$ and $(X_{j,t} (\boldsymbol{\upsilon}_j))_{t \in \mathbb{Z}}$ converges uniformly e.a.s. to zero. For each $j \in \{1,...,J\}$, let
\begin{equation*}
\Lambda_{j,t} (\boldsymbol{\upsilon}_j) := \sup_{x_j \in \mathcal{X}_{j}} \left| \nabla_{x_j} \phi_{j} (Y_t,x_j;\boldsymbol{\upsilon}_j) \right|.
\end{equation*}
\begin{lemma} \label{LemmaInvertibilityX}
Assume that Assumptions \ref{AssumptionY}, \ref{AssumptionTheta1}, and \ref{AssumptionPhi1} hold. Moreover, assume that, for each $j \in \{1,...,J\}$,
\begin{enumerate} [(i)]
\item there exists an $x_j \in \mathcal{X}_{j}$ such that $\mathbb{E} [\log^{+} \sup_{\boldsymbol{\upsilon}_{j} \in \boldsymbol{\Upsilon}_{j}} | \phi_{j} (Y_t,x_j;\boldsymbol{\upsilon}_{j}) - x_j | ] < \infty$,
\item $\mathbb{E} [\log^{+} \sup_{\boldsymbol{\upsilon}_{j} \in \boldsymbol{\Upsilon}_{j}} \Lambda_{j,t} (\boldsymbol{\upsilon}_{j})] < \infty$, and
\item $-\infty \leq \mathbb{E} [\log \sup_{\boldsymbol{\upsilon}_{j} \in \boldsymbol{\Upsilon}_{j}} \Lambda_{j,t} (\boldsymbol{\upsilon}_{j})] < 0$.
\end{enumerate}
Then, for each $j \in \{1,...,J\}$, $(X_{j,t} (\boldsymbol{\upsilon}_j))_{t \in \mathbb{Z}}$ given by
\begin{equation*}
X_{j,t+1} (\boldsymbol{\upsilon}_j) = \phi_j (Y_t,X_{j,t} (\boldsymbol{\upsilon}_j);\boldsymbol{\upsilon}_{j})
\end{equation*}
is stationary and ergodic for all $\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_{j}$ and
\begin{equation*}
\sup_{\boldsymbol{\upsilon}_{j} \in \boldsymbol{\Upsilon}_{j}} \left| \hat{X}_{j,t} (\boldsymbol{\upsilon}_j) - X_{j,t} (\boldsymbol{\upsilon}_j) \right| \overset{e.a.s.}{\rightarrow} 0 \quad \text{as} \quad t \rightarrow \infty
\end{equation*}
for any initialisation $\hat{X}_{j,1} (\boldsymbol{\upsilon}_j) \in \mathcal{X}_{j}$.
\end{lemma}
\noindent
As above, Conditions (i) and (ii) are standard regularity conditions, and Condition (iii) is a standard contraction condition. They are similar to the conditions in \cite{BlasquesGorgiKoopmanWintenberger2018}.
Theorem 3.1 in \citet{Bougerol1993} can, however, not be used for $(\hat{\boldsymbol{\pi}}_{t \mid t-1} (\boldsymbol{\theta}))_{t \in \mathbb{N}}$ because $(\hat{\boldsymbol{\pi}}_{t \mid t-1} (\boldsymbol{\theta}))_{t \in \mathbb{N}}$ depends on $(\hat{X}_{1,t} (\boldsymbol{\upsilon}_1))_{t \in \mathbb{N}},...,(\hat{X}_{J,t} (\boldsymbol{\upsilon}_J))_{t \in \mathbb{N}}$ which are non-stationary. The following lemma, which gives conditions under which the difference between $(\hat{\boldsymbol{\pi}}_{t \mid t} (\boldsymbol{\theta}))_{t \in \mathbb{N}_0}$ and $(\boldsymbol{\pi}_{t \mid t} (\boldsymbol{\theta}))_{t \in \mathbb{Z}}$ converges uniformly e.a.s. to zero, follows instead from an application of Propositions 1 and 3 in \citet{Krabbe2025}, which, when combined, is a generalisation of Theorem 2.10 in \cite{StraumannMikosch2006}.\footnote{Propositions 1 and 3 in \citet{Krabbe2025} generalise the result in Theorem 2.10 in \cite{StraumannMikosch2006}, which is about stochastic difference equations taking values in a real or complex separable Banach space, to stochastic difference equations taking values in a complete subspace of a real or complex separable Banach space.}
\begin{lemma} \label{LemmaInvertibilityFilter}
Assume that Assumptions \ref{AssumptionY}-\ref{AssumptionPhi1} and the conditions in Lemma \ref{LemmaInvertibilityX} hold. Moreover, assume that for each $j \in \{1,...,J\}$, there exists an $m_j > 0$ such that
\begin{equation*}
\mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_j} \sup_{x_j \in \mathcal{X}_{j}} \left| \nabla_{x_j} \log f_{j} (Y_t;x_j,\boldsymbol{\upsilon}_j) \right|^{m_j} \right] < \infty.
\end{equation*}
Then, $(\boldsymbol{\pi}_{t \mid t} (\boldsymbol{\theta}))_{t \in \mathbb{Z}}$ given by
\begin{equation*}
\boldsymbol{\pi}_{t \mid t} (\boldsymbol{\theta}) = \mathbf{F}_t (\mathbf{P}^{\prime} \boldsymbol{\pi}_{t-1 \mid t-1} (\boldsymbol{\theta});\boldsymbol{\theta}) \mathbf{P}^{\prime} \boldsymbol{\pi}_{t-1 \mid t-1} (\boldsymbol{\theta})
\end{equation*}
where
\begin{equation*}
[\mathbf{F}_t (\mathbf{s})]_{ii} := \frac{f_{i} (Y_t;X_{i,t} (\boldsymbol{\upsilon}_i),\boldsymbol{\upsilon}_{i})}{\sum_{k=1}^{J} s_{k} f_{k} (Y_t;X_{k,t} (\boldsymbol{\upsilon}_k),\boldsymbol{\upsilon}_{k})}, \quad i \in \{1,...,J\}
\end{equation*}
is stationary and ergodic for all $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ and
\begin{equation*}
\sup_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} \left| \left| \hat{\boldsymbol{\pi}}_{t \mid t} (\boldsymbol{\theta}) - \boldsymbol{\pi}_{t \mid t} (\boldsymbol{\theta}) \right| \right|_{2} \overset{e.a.s.}{\rightarrow} 0 \quad \text{as} \quad t \rightarrow \infty
\end{equation*}
for any initialisation $\hat{\boldsymbol{\pi}}_{0 \mid 0} (\boldsymbol{\theta}) \in \mathcal{S}$.
\end{lemma}
\noindent
The next corollary is a direct consequence hereof.
\begin{corollary} \label{CorollaryInvertibilityPrediction}
Under the assumptions in Lemma \ref{LemmaInvertibilityFilter}, $(\boldsymbol{\pi}_{t \mid t-1} (\boldsymbol{\theta}))_{t \in \mathbb{Z}}$ given by
\begin{equation*}
\boldsymbol{\pi}_{t+1 \mid t} (\boldsymbol{\theta}) = \mathbf{P}^{\prime} \boldsymbol{\pi}_{t \mid t} (\boldsymbol{\theta})
\end{equation*}
is stationary and ergodic for all $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ and
\begin{equation*}
\sup_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} \left| \left| \hat{\boldsymbol{\pi}}_{t \mid t-1} (\boldsymbol{\theta}) - \boldsymbol{\pi}_{t \mid t-1} (\boldsymbol{\theta}) \right| \right|_{2} \overset{e.a.s.}{\rightarrow} 0 \quad \text{as} \quad t \rightarrow \infty
\end{equation*}
for any initialisation $\hat{\boldsymbol{\pi}}_{0 \mid 0} (\boldsymbol{\theta}) \in \mathcal{S}$.
\end{corollary}
\begin{remark} \label{Remark}
Note that, although $\boldsymbol{\pi}_{t \mid t} (\boldsymbol{\theta}) \in \mathcal{S}$ for all $t \in \mathbb{Z}$, $\boldsymbol{\pi}_{t \mid t-1} (\boldsymbol{\theta}) \in \mathcal{S}_{\boldsymbol{\theta}}$ for all $t \in \mathbb{Z}$ where $\mathcal{S}_{\boldsymbol{\theta}} := \{\mathbf{x} \in \mathbb{R}^{J} : x_j \geq \min_{i \in \{1,...,J\}} p_{ij}, j=1,...,J, \sum_{j=1}^{J} x_j = 1 \}$.
\end{remark}
\noindent
The result in Lemma \ref{LemmaInvertibilityFilter}/Corollary \ref{CorollaryInvertibilityPrediction} is not surprising. The Markov chain itself forgets its initialisation asymptotically (in case it is initialised) as $p_{ij} > 0$ for all $i,j \in \{1,...,J\}$, so it is not surprising that the filter/predictor also forgets its initialisation asymptotically provided that the time-varying parameters do the same.
Moreover, we assume the following.
\begin{assumption} \label{AssumptionF2}
For each $j \in \{1,...,J\}$,
\begin{equation*}
\mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_j} | \log f_{j} (Y_t;X_{j,t}(\boldsymbol{\upsilon}_j),\boldsymbol{\upsilon}_j)| \right] < \infty.
\end{equation*}
\end{assumption}
\noindent
Assumption \ref{AssumptionF2} is a standard moment condition and is similar to Assumption 3 in \cite{DoucMoulinesRyden2004} and Assumption 4 in \cite{KasaharaShimotsu2019}.
Finally, we assume the following as in \citet{FrancqRoussignol1998} and \cite{FrancqRoussignolZakoian2001} where $f^{(m)} (\mathbf{y} ; \boldsymbol{\theta}), \mathbf{y} \in \mathcal{Y}^{m}$ denotes the conditional pdf of $\mathbf{Y}_{t-m+1}^{t}$ given $\mathbf{Y}_{-\infty}^{t-m}$.
\begin{assumption} \label{AssumptionIdentification}
There exists an $m \in \mathbb{N}$ such that
\begin{align*}
f^{(m)} (\mathbf{Y}_{t-m+1}^{t} ; \boldsymbol{\theta}) &= f^{(m)} (\mathbf{Y}_{t-m+1}^{t} ; \boldsymbol{\theta}_0) \quad a.s.
\shortintertext{implies that}
\boldsymbol{\theta} &= \boldsymbol{\theta}_0.
\end{align*}
\end{assumption}
\noindent
Assumption \ref{AssumptionIdentification} is a standard identification condition. Note that slightly different identification conditions are used in \cite{DoucMoulinesRyden2004} and \cite{KasaharaShimotsu2019}.
Theorem \ref{TheoremConsistency} gives conditions under which the MLE is consistent.
\begin{theorem} \label{TheoremConsistency}
Assume that Assumptions \ref{AssumptionY}-\ref{AssumptionIdentification} and the conditions in Lemmas \ref{LemmaInvertibilityX} and \ref{LemmaInvertibilityFilter} hold. Then,
\begin{equation*}
\hat{\boldsymbol{\theta}}_T \overset{a.s.}{\rightarrow} \boldsymbol{\theta}_0 \quad \text{as} \quad T \rightarrow \infty.
\end{equation*}
\end{theorem}
\subsection{Asymptotic Normality} \label{AsymptoticNormality}
Asymptotic normality follows from classical arguments as consistency if the \textit{initialised} and thus non-stationary score vector $\sqrt{T} \nabla_{\boldsymbol{\theta}} \hat{L}_T (\boldsymbol{\theta}_0)$ and observed Fisher information matrix $(-\nabla_{\boldsymbol{\theta} \boldsymbol{\theta}} \hat{L}_T (\boldsymbol{\theta}))$ converge in distribution to $\mathcal{N} (\mathbf{0},\mathbf{I}(\boldsymbol{\theta}_0))$ and uniformly a.s. to $\mathbf{I}(\boldsymbol{\theta})$, respectively, where
\begin{equation*}
\mathbf{I} (\boldsymbol{\theta}) = - \mathbb{E} [ \nabla_{\boldsymbol{\theta} \boldsymbol{\theta}} \log f (Y_t;\boldsymbol{\theta}) ]
\end{equation*}
is the Fisher information matrix. As above, we show both in two steps, the first steps being the main difficulties and hence focus. To show the former, we first show that the difference between $\sqrt{T} \nabla_{\boldsymbol{\theta}} \hat{L}_T (\boldsymbol{\theta})$ and the \textit{non-initialised}, stationary, and ergodic score vector $\sqrt{T} \nabla_{\boldsymbol{\theta}} L_T (\boldsymbol{\theta})$ converges uniformly a.s. to zero by showing that the first-order derivatives of the time-varying parameters and the predictor forget their initialisations asymptotically and then that $\sqrt{T} \nabla_{\boldsymbol{\theta}} L_T (\boldsymbol{\theta}_0)$ converges in distribution to $\mathcal{N} (\mathbf{0},\mathbf{I}(\boldsymbol{\theta}_0))$ by using standard arguments. To show the latter, we first show that the difference between $(-\nabla_{\boldsymbol{\theta} \boldsymbol{\theta}} \hat{L}_T (\boldsymbol{\theta}))$ and the \textit{non-initialised}, stationary, and ergodic observed Fisher information matrix $(-\nabla_{\boldsymbol{\theta} \boldsymbol{\theta}} L_T (\boldsymbol{\theta}))$ converges uniformly a.s. to zero by showing that also the second-order derivatives of the time-varying parameters and the predictor forget their initialisations asymptotically and then that $(-\nabla_{\boldsymbol{\theta} \boldsymbol{\theta}} L_T (\boldsymbol{\theta}))$ converges uniformly a.s. to $\mathbf{I}(\boldsymbol{\theta})$ by using standard arguments.
In the following, let, for a function $\boldsymbol{\upsilon} \mapsto f(X(\boldsymbol{\upsilon}),\boldsymbol{\upsilon}) : \boldsymbol{\Upsilon} \rightarrow \mathbb{R}$ where $\boldsymbol{\upsilon} \mapsto X(\boldsymbol{\upsilon}) : \boldsymbol{\Upsilon} \rightarrow \mathbb{R}$ is another function, $\bar{\nabla}_{x} f(X(\boldsymbol{\upsilon}),\boldsymbol{\upsilon})$ and $\bar{\nabla}_{\boldsymbol{\upsilon}} f(X(\boldsymbol{\upsilon}),\boldsymbol{\upsilon})$ be given by
\begin{equation*}
\bar{\nabla}_{x} f(X(\boldsymbol{\upsilon}),\boldsymbol{\upsilon}) := \left. \nabla_{\bar{x}} f(\bar{x},\bar{\boldsymbol{\upsilon}}) \right|_{\bar{x} = X(\boldsymbol{\upsilon}),\bar{\boldsymbol{\upsilon}} = \boldsymbol{\upsilon}} \quad \text{and} \quad \bar{\nabla}_{\boldsymbol{\upsilon}} f(X(\boldsymbol{\upsilon}),\boldsymbol{\upsilon}) := \left. \nabla_{ \bar{\boldsymbol{\upsilon}}} f(\bar{x},\bar{\boldsymbol{\upsilon}}) \right|_{\bar{x} = X(\boldsymbol{\upsilon}),\bar{\boldsymbol{\upsilon}} = \boldsymbol{\upsilon}},
\end{equation*}
$\bar{\nabla}_{\boldsymbol{\upsilon} x} f(X(\boldsymbol{\upsilon}),\boldsymbol{\upsilon})$ be given by
\begin{equation*}
\bar{\nabla}_{\boldsymbol{\upsilon} x} f(X(\boldsymbol{\upsilon}),\boldsymbol{\upsilon}) := \left. \nabla_{\bar{\boldsymbol{\upsilon}} \bar{x}} f(\bar{x},\bar{\boldsymbol{\upsilon}}) \right|_{\bar{x} = X(\boldsymbol{\upsilon}),\bar{\boldsymbol{\upsilon}} = \boldsymbol{\upsilon}},
\end{equation*}
and $\bar{\nabla}_{x x} f(X(\boldsymbol{\upsilon}),\boldsymbol{\upsilon})$ and $\bar{\nabla}_{\boldsymbol{\upsilon} \boldsymbol{\upsilon}} f(X(\boldsymbol{\upsilon}),\boldsymbol{\upsilon})$ be given by
\begin{equation*}
\bar{\nabla}_{x x} f(X(\boldsymbol{\upsilon}),\boldsymbol{\upsilon}) := \left. \nabla_{\bar{x} \bar{x}} f(\bar{x},\bar{\boldsymbol{\upsilon}}) \right|_{\bar{x} = X(\boldsymbol{\upsilon}),\bar{\boldsymbol{\upsilon}} = \boldsymbol{\upsilon}} \quad \text{and} \quad
\bar{\nabla}_{\boldsymbol{\upsilon} \boldsymbol{\upsilon}} f(X(\boldsymbol{\upsilon}),\boldsymbol{\upsilon}) := \left. \nabla_{\bar{\boldsymbol{\upsilon}} \bar{\boldsymbol{\upsilon}}} f(\bar{x},\bar{\boldsymbol{\upsilon}}) \right|_{\bar{x} = X(\boldsymbol{\upsilon}),\bar{\boldsymbol{\upsilon}} = \boldsymbol{\upsilon}}.
\end{equation*}
In addition to Assumptions \ref{AssumptionY}-\ref{AssumptionIdentification}, we assume the following.
\begin{assumption} \label{AssumptionTheta2}
$\boldsymbol{\theta}_0 \in \textup{int}(\boldsymbol{\Theta})$.
\end{assumption}
\begin{assumption} \label{AssumptionF3}
For each $j \in \{1,...,J\}$,
\begin{enumerate} [(i)]
\item $(x_j,\boldsymbol{\upsilon}_j) \mapsto f_{j} (y;x_j,\boldsymbol{\upsilon}_j)$ is twice continuously differentiable for all $y \in \mathcal{Y}$ and
\item $x_j \mapsto \nabla_{(x_j,\boldsymbol{\upsilon}_j)(x_j,\boldsymbol{\upsilon}_j)} f_{j} (y;x_j,\boldsymbol{\upsilon}_j)$ is differentiable for all $y \in \mathcal{Y}$ and $\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_{j}$.
\end{enumerate}
\end{assumption}
\begin{assumption} \label{AssumptionPhi2}
For each $j \in \{1,...,J\}$,
\begin{enumerate} [(i)]
\item $(x_j,\boldsymbol{\upsilon}_j) \mapsto \phi_{j} (y,x_j;\boldsymbol{\upsilon}_j)$ is twice continuously differentiable for all $y \in \mathcal{Y}$.
\end{enumerate}
\end{assumption}
\noindent
Assumption \ref{AssumptionTheta2} is a standard regularity condition, which implies that there exists an $\varepsilon > 0$ such that $\boldsymbol{\theta}_0 \in \textup{int}(\bar{\boldsymbol{\Theta}})$ where
\begin{equation*}
\bar{\boldsymbol{\Theta}} := \left\lbrace \boldsymbol{\theta} \in \boldsymbol{\Theta} : \left| \left| \boldsymbol{\theta} - \boldsymbol{\theta}_0 \right| \right|_2 \leq \varepsilon \right\rbrace \subset \textup{int}(\boldsymbol{\Theta}).
\end{equation*}
Assumptions \ref{AssumptionF3} and \ref{AssumptionPhi2} are also standard regularity conditions; Condition (i) in Assumption \ref{AssumptionF3} is similar to Assumption 6 in \cite{DoucMoulinesRyden2004} and Assumption 7(a) in \cite{KasaharaShimotsu2019}, and Condition (ii) in Assumption \ref{AssumptionF3} is an extra regularity condition, which is used to prove that the second-order derivative of the predictor forgets its initialisation asymptotically.
The next two lemmata, which follow from applications of Theorem 2.10 in \citet{StraumannMikosch2006}, give conditions under which, for each $j \in \{1,...,J\}$, the difference between $( \nabla_{\boldsymbol{\upsilon}_j} \hat{X}_{j,t} (\boldsymbol{\upsilon}_j) )_{t \in \mathbb{N}}$ and $( \nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) )_{t \in \mathbb{Z}}$ and the one between $( \nabla_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} \hat{X}_{j,t} (\boldsymbol{\upsilon}_j) )_{t \in \mathbb{N}}$ and $( \nabla_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) )_{t \in \mathbb{Z}}$ converge uniformly e.a.s. to zero.
\begin{lemma} \label{LemmaInvertibilityDX}
Assume that Assumptions \ref{AssumptionY}, \ref{AssumptionTheta1}, \ref{AssumptionPhi1}, \ref{AssumptionPhi2}, and the conditions in Lemma \ref{LemmaInvertibilityX} hold. Moreover, assume that, for each $j \in \{1,...,J\}$,
\begin{enumerate} [(i)]
\item $\mathbb{E} [\log^{+} \sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} || \bar{\nabla}_{\boldsymbol{\upsilon}_j} \phi_{j} (Y_t,X_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) ||_{2} ] < \infty$.
\end{enumerate}
Then, for each $j \in \{1,...,J\}$, $( \nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) )_{t \in \mathbb{Z}}$ is stationary and ergodic for all $\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_j$. Finally, assume that, for each $j \in \{1,...,J\}$,
\begin{enumerate} [(i)]
\item $\mathbb{E} [ \log^{+} \sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} || \nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) ||_{2} ] < \infty$,
\item $\sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} || \bar{\nabla}_{\boldsymbol{\upsilon}_j} \phi_{j} (Y_t,\hat{X}_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) - \bar{\nabla}_{\boldsymbol{\upsilon}_j} \phi_{j} (Y_t,X_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) ||_{2} \overset{e.a.s.}{\rightarrow} 0$ as $t \rightarrow \infty$, and
\item $\sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} | \bar{\nabla}_{x_j} \phi_{j} (Y_t,\hat{X}_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) - \bar{\nabla}_{x_j} \phi_{j} (Y_t,X_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) | \overset{e.a.s.}{\rightarrow} 0$ as $t \rightarrow \infty$.
\end{enumerate}
Then, for each $j \in \{1,...,J\}$,
\begin{equation*}
\sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} || \nabla_{\boldsymbol{\upsilon}_j} \hat{X}_{j,t} (\boldsymbol{\upsilon}_j) - \nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) ||_{2} \overset{e.a.s.}{\rightarrow} 0 \quad \text{as} \quad t \rightarrow \infty
\end{equation*}
for any initialisation $\nabla_{\boldsymbol{\upsilon}_j} \hat{X}_{j,1} (\boldsymbol{\upsilon}_j) \in \mathbb{R}^{d_j}$.
\end{lemma}
\begin{lemma} \label{LemmaInvertibilityDDX}
Assume that Assumptions \ref{AssumptionY}, \ref{AssumptionTheta1}, \ref{AssumptionPhi1}, \ref{AssumptionPhi2}, and the conditions in Lemmas \ref{LemmaInvertibilityX} and \ref{LemmaInvertibilityDX} hold. Moreover, assume that, for each $j \in \{1,...,J\}$,
\begin{enumerate} [(i)]
\item $\mathbb{E} [\log^{+} \sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} | \bar{\nabla}_{x_j x_j} \phi_{j} (Y_t,X_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) | ] < \infty$,
\item $\mathbb{E} [\log^{+} \sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} || \bar{\nabla}_{\boldsymbol{\upsilon}_j x_j} \phi_{j} (Y_t,X_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) ||_{2} ] < \infty$, and
\item $\mathbb{E} [\log^{+} \sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} || \bar{\nabla}_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} \phi_{j} (Y_t,X_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) ||_{2,2} ] < \infty$.
\end{enumerate}
Then, for each $j \in \{1,...,J\}$, $( \nabla_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) )_{t \in \mathbb{Z}}$ is stationary and ergodic for all $\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_j$. Finally, assume that, for each $j \in \{1,...,J\}$,
\begin{enumerate} [(i)]
\item $\mathbb{E} [ \log^{+} \sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} || \nabla_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) ||_{2,2} ] < \infty$,
\item $\sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} | \bar{\nabla}_{x_j x_j} \phi_{j} (Y_t,\hat{X}_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) - \bar{\nabla}_{x_j x_j} \phi_{j} (Y_t,X_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) | \overset{e.a.s.}{\rightarrow} 0$ as $t \rightarrow \infty$,
\item $\sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} || \bar{\nabla}_{\boldsymbol{\upsilon}_j x_j} \phi_{j} (Y_t,\hat{X}_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) - \bar{\nabla}_{\boldsymbol{\upsilon}_j x_j} \phi_{j} (Y_t,X_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) ||_{2} \overset{e.a.s.}{\rightarrow} 0$ as $t \rightarrow \infty$, and
\item $\sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} || \bar{\nabla}_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} \phi_{j} (Y_t,\hat{X}_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) - \bar{\nabla}_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} \phi_{j} (Y_t,X_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) ||_{2,2} \overset{e.a.s.}{\rightarrow} 0$ as $t \rightarrow \infty$.
\end{enumerate}
Then, for each $j \in \{1,...,J\}$,
\begin{equation*}
\sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} || \nabla_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} \hat{X}_{j,t} (\boldsymbol{\upsilon}_j) - \nabla_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) ||_{2,2} \overset{e.a.s.}{\rightarrow} 0 \quad \text{as} \quad t \rightarrow \infty
\end{equation*}
for any initialisation $\nabla_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} \hat{X}_{j,1} (\boldsymbol{\upsilon}_j) \in \mathbb{R}^{d_j \times d_j}$.
\end{lemma}
We also assume the following.
\begin{assumption} \label{AssumptionF4}
For each $j \in \{1,...,J\}$,
\begin{enumerate} [(i)]
\item $\mathbb{E} [ \sup\nolimits_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} | \bar{\nabla}_{x_j} \log f_{j} (Y_t;X_{j,t} (\boldsymbol{\upsilon}_j),\boldsymbol{\upsilon}_j) |^2 || \nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) ||_{2}^2 ] < \infty$,
\item $\mathbb{E} [ \sup\nolimits_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} || \bar{\nabla}_{\boldsymbol{\upsilon}_j} \log f_{j} (Y_t;X_{j,t} (\boldsymbol{\upsilon}_j),\boldsymbol{\upsilon}_j) ||_{2}^{2} ] < \infty$,
\item $\mathbb{E} [ \sup\nolimits_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} | \bar{\nabla}_{x_j x_j} \log f_{j} (Y_t;{X}_{j,t} (\boldsymbol{\upsilon}_j),\boldsymbol{\upsilon}_j) | || \nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) ||_{2}^2 ] < \infty$,
\item $\mathbb{E} [ \sup\nolimits_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} | \bar{\nabla}_{x_j} \log f_{j} (Y_t;{X}_{j,t} (\boldsymbol{\upsilon}_j),\boldsymbol{\upsilon}_j) | || \nabla_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) ||_{2,2} ] < \infty$,
\item $\mathbb{E} [ \sup\nolimits_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} || \bar{\nabla}_{\boldsymbol{\upsilon}_j x_j} \log f_{j} (Y_t;{X}_{j,t} (\boldsymbol{\upsilon}_j),\boldsymbol{\upsilon}_j) ||_{2} || \nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) ||_{2} ] < \infty$, and
\item $\mathbb{E} [ \sup\nolimits_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} || \bar{\nabla}_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} \log f_{j} (Y_t;{X}_{j,t} (\boldsymbol{\upsilon}_j),\boldsymbol{\upsilon}_j) ||_{2,2} ] < \infty$.
\end{enumerate}
\end{assumption}
\noindent
Assumption \ref{AssumptionF4} is standard moment conditions, which imply that, for each $j \in \{1,...,J\}$,
\begin{equation*}
\mathbb{E} [ \sup\nolimits_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_j} || \nabla_{\boldsymbol{\upsilon}_j} \log f_{j} (Y_t;X_{j,t}(\boldsymbol{\upsilon}_j),\boldsymbol{\upsilon}_j) ||_{2}^{2} ] < \infty
\end{equation*}
and
\begin{equation*}
\mathbb{E} [ \sup\nolimits_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_j} || \nabla_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} \log f_{j} (Y_t;X_{j,t}(\boldsymbol{\upsilon}_j),\boldsymbol{\upsilon}_j) ||_{2,2} ] < \infty,
\end{equation*}
and is thus similar to Assumption 7(b) in \cite{DoucMoulinesRyden2004} and Assumption 7(c) in \cite{KasaharaShimotsu2019}.
In the following two lemmata, let, for a function $\mathbf{x} \mapsto \mathbf{f}(\mathbf{x}) : \mathbb{R}^n \rightarrow \mathbb{R}^k$, $\nabla_{\mathbf{x}} \mathbf{f}(\mathbf{x})$ be given by
\begin{equation*}
[\nabla_{\mathbf{x}} \mathbf{f}(\mathbf{x})]_{(j-1)n+i} = \nabla_{x_i} f_j (\mathbf{x}), \quad (i,j) \in \{1,...,n\} \times \{1,...,k\}
\end{equation*}
and $\nabla_{\mathbf{x} \mathbf{x}} \mathbf{f}(\mathbf{x})$ be given by
\begin{equation*}
[\nabla_{\mathbf{x} \mathbf{x}} \mathbf{f}(\mathbf{x})]_{(j-1)n^2+(i-1)n+l} = \nabla_{x_i x_l} f_j (\mathbf{x}), \quad (i,j,l) \in \{1,...,n\} \times \{1,...,k\} \times \{1,...,n\}.
\end{equation*}
With this notation, the next two lemmata, which also follow from applications of Theorem 2.10 in \citet{StraumannMikosch2006}, give conditions under which the difference between $( \nabla_{\boldsymbol{\theta}} \hat{\boldsymbol{\pi}}_{t \mid t-1} (\boldsymbol{\theta}) )_{t \in \mathbb{N}}$ and $( \nabla_{\boldsymbol{\theta}} \boldsymbol{\pi}_{t \mid t-1} (\boldsymbol{\theta}) )_{t \in \mathbb{Z}}$ and the one between $( \nabla_{\boldsymbol{\theta} \boldsymbol{\theta}} \hat{\boldsymbol{\pi}}_{t \mid t-1} (\boldsymbol{\theta}) )_{t \in \mathbb{N}}$ and $( \nabla_{\boldsymbol{\theta} \boldsymbol{\theta}} \boldsymbol{\pi}_{t \mid t-1} (\boldsymbol{\theta}) )_{t \in \mathbb{Z}}$ converge uniformly e.a.s. to zero.
\begin{lemma} \label{LemmaInvertibilityDPrediction}
Assume that Assumptions \ref{AssumptionY}-\ref{AssumptionPhi1}, \ref{AssumptionF3}-\ref{AssumptionF4}, and the conditions in Lemmas \ref{LemmaInvertibilityX}-\ref{LemmaInvertibilityDX} hold. Moreover, assume that for each $j \in \{1,...,J\}$, there exists an $m_j > 0$ such that
\begin{enumerate} [(i)]
\item $\mathbb{E} [ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_j} \sup_{x_j \in \mathcal{X}_{j}} | | \bar{\nabla}_{\boldsymbol{\upsilon}_j} \log f_{j} (Y_t;x_j,\boldsymbol{\upsilon}_j) | |_{2}^{m_j} ] < \infty$,
\item $\mathbb{E} [ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_j} \sup_{x_j \in \mathcal{X}_{j}} | \bar{\nabla}_{x_j x_j} \log f_{j} (Y_t;x_j,\boldsymbol{\upsilon}_j) |^{m_j} ] < \infty$, and
\item $\mathbb{E} [ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_j} \sup_{x_j \in \mathcal{X}_{j}} | | \bar{\nabla}_{\boldsymbol{\upsilon}_j x_j} \log f_{j} (Y_t;x_j,\boldsymbol{\upsilon}_j) | |_{2}^{m_j} ] < \infty$.
\end{enumerate}
Then, $( \nabla_{\boldsymbol{\theta}} \boldsymbol{\pi}_{t \mid t-1} (\boldsymbol{\theta}) )_{t \in \mathbb{Z}}$ is stationary and ergodic for all $\boldsymbol{\theta} \in \bar{\boldsymbol{\Theta}}$ with $\mathbb{E} [ \sup_{\boldsymbol{\theta} \in \bar{\boldsymbol{\Theta}}} || \nabla_{\boldsymbol{\theta}} \boldsymbol{\pi}_{t \mid t-1} (\boldsymbol{\theta}) ||_{2}^{2} ] < \infty$ and
\begin{equation*}
\sup_{\boldsymbol{\theta} \in \bar{\boldsymbol{\Theta}}} | | \nabla_{\boldsymbol{\theta}} \hat{\boldsymbol{\pi}}_{t \mid t-1} (\boldsymbol{\theta}) - \nabla_{\boldsymbol{\theta}} \boldsymbol{\pi}_{t \mid t-1} (\boldsymbol{\theta}) | |_{2} \overset{e.a.s.}{\rightarrow} 0 \quad \text{as} \quad t \rightarrow \infty
\end{equation*}
for any initialisation $\nabla_{\boldsymbol{\theta}} \hat{\boldsymbol{\pi}}_{0 \mid 0} (\boldsymbol{\theta}) \in \mathbb{R}^{(J-1)d}$.
\end{lemma}
\begin{lemma} \label{LemmaInvertibilityDDPrediction}
Assume that Assumptions \ref{AssumptionY}-\ref{AssumptionPhi1}, \ref{AssumptionF3}-\ref{AssumptionF4}, and the conditions in Lemmas \ref{LemmaInvertibilityX}-\ref{LemmaInvertibilityDPrediction} hold. Moreover, assume that for each $j \in \{1,...,J\}$, there exists an $m_j > 0$ such that
\begin{enumerate} [(i)]
\item $\mathbb{E} [ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_j} \sup_{x_j \in \mathcal{X}_{j}} | | \bar{\nabla}_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} \log f_{j} (Y_t;x_j,\boldsymbol{\upsilon}_j) | |_{2,2}^{m_j} ] < \infty$,
\item $\mathbb{E} [ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_j} \sup_{x_j \in \mathcal{X}_{j}} | \bar{\nabla}_{x_j x_j x_j} \log f_{j} (Y_t;x_j,\boldsymbol{\upsilon}_j) |^{m_j} ] < \infty$,
\item $\mathbb{E} [ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_j} \sup_{x_j \in \mathcal{X}_{j}} | | \bar{\nabla}_{\boldsymbol{\upsilon}_j x_j x_j} \log f_{j} (Y_t;x_j,\boldsymbol{\upsilon}_j) | |_{2}^{m_j} ] < \infty$, and
\item $\mathbb{E} [ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_j} \sup_{x_j \in \mathcal{X}_{j}} | | \bar{\nabla}_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j x_j} \log f_{j} (Y_t;x_j,\boldsymbol{\upsilon}_j) | |_{2,2}^{m_j} ] < \infty$.
\end{enumerate}
Then, $( \nabla_{\boldsymbol{\theta} \boldsymbol{\theta}} \boldsymbol{\pi}_{t \mid t-1} (\boldsymbol{\theta}) )_{t \in \mathbb{Z}}$ is stationary and ergodic for all $\boldsymbol{\theta} \in \bar{\boldsymbol{\Theta}}$ with $\mathbb{E} [ \sup_{\boldsymbol{\theta} \in \bar{\boldsymbol{\Theta}}} || \nabla_{\boldsymbol{\theta} \boldsymbol{\theta}} \boldsymbol{\pi}_{t \mid t-1} (\boldsymbol{\theta}) ||_{2} ] < \infty$ and
\begin{equation*}
\sup_{\boldsymbol{\theta} \in \bar{\boldsymbol{\Theta}}} | | \nabla_{\boldsymbol{\theta} \boldsymbol{\theta}} \hat{\boldsymbol{\pi}}_{t \mid t-1} (\boldsymbol{\theta}) - \nabla_{\boldsymbol{\theta} \boldsymbol{\theta}} \boldsymbol{\pi}_{t \mid t-1} (\boldsymbol{\theta}) | |_{2} \overset{e.a.s.}{\rightarrow} 0 \quad \text{as} \quad t \rightarrow \infty
\end{equation*}
for any initialisation $\nabla_{\boldsymbol{\theta} \boldsymbol{\theta}} \hat{\boldsymbol{\pi}}_{0 \mid 0} (\boldsymbol{\theta}) \in \mathbb{R}^{(J-1)d^2}$.
\end{lemma}
Theorem \ref{TheoremAsymptoticNormality} gives conditions under which the MLE is also asymptotically normal.
\begin{theorem} \label{TheoremAsymptoticNormality}
Assume that Assumptions \ref{AssumptionY}-\ref{AssumptionF4} and the conditions in Lemmas \ref{LemmaInvertibilityX}-\ref{LemmaInvertibilityDDPrediction} hold and that $\mathbf{I} (\boldsymbol{\theta}_0)$ is invertible. Then,
\begin{equation*}
\sqrt{T} (\hat{\boldsymbol{\theta}}_T - \boldsymbol{\theta}_0) \overset{d}{\rightarrow} \mathcal{N} (\mathbf{0},\mathbf{I}(\boldsymbol{\theta}_0)^{-1}) \quad \text{as} \quad T \rightarrow \infty.
\end{equation*}
\end{theorem}
\noindent
\cite{DoucMoulinesRyden2004} also assume that the Fisher information matrix $\mathbf{I} (\boldsymbol{\theta}_0)$ is invertible; an assumption which seems to be missing in \cite{KasaharaShimotsu2019}.
Finally, Proposition \ref{prop:varcovarmatrix} shows that, under the assumptions in Theorem \ref{TheoremAsymptoticNormality}, the observed Fisher information matrix $(- \nabla_{\boldsymbol{\theta} \boldsymbol{\theta}} \hat{L}_T (\hat{\boldsymbol{\theta}}_T))$ is a consistent estimator of the Fisher information matrix $\mathbf{I} (\boldsymbol{\theta}_0)$.
\begin{proposition} \label{prop:varcovarmatrix}
Under the assumptions in Theorem \ref{TheoremAsymptoticNormality},
\begin{equation*}
\nabla_{\boldsymbol{\theta} \boldsymbol{\theta}} \hat{L}_T (\hat{\boldsymbol{\theta}}_T) \overset{a.s.}{\rightarrow} (-\mathbf{I}(\boldsymbol{\theta}_0)) \quad \text{as} \quad T \rightarrow \infty.
\end{equation*}
\end{proposition}
\section{The Markov-switching GARCH Model} \label{Examples2}
In this section, we study both the asymptotic and finite-sample properties of the MLE for the Markov-switching GARCH model by \cite{HaasMittnikPaolella2004b}, the latter in a Monte Carlo simulation study.\footnote{Recall that the Markov-switching GARCH model reduces to the mixture GARCH model by \citet{HaasMittnikPaolella2004a} if $(S_t)_{t \in \mathbb{Z}}$ is an i.i.d. chain.}
\subsection{Asymptotic Properties}
Recall that the Markov-switching GARCH model by \citet{HaasMittnikPaolella2004b} is given by
\begin{equation*}
Y_t = \sqrt{X_{S_t,t}} \varepsilon_t,
\end{equation*}
where $(S_t)_{t \in \mathbb{Z}}$ is a stationary, irreducible, and aperiodic (thus ergodic) Markov chain with transition probabilities $p_{ij,0} \in (0,1), i,j \in \{1,...,J\}$, for each $j \in \{1,...,J\}$,
\begin{equation*}
X_{j,t+1} = \omega_{j,0} + \alpha_{j,0} Y_{t}^2 + \beta_{j,0} X_{j,t},
\end{equation*}
where $\omega_{j,0} > 0$, $\alpha_{j,0} \geq 0$, and $\beta_{j,0} \geq 0$, $(\varepsilon_t)_{t \in \mathbb{Z}}$ is a sequence of independent normal distributed random variables with zero mean and unit variance, and $(S_t)_{t \in \mathbb{Z}}$ and $(\varepsilon_t)_{t \in \mathbb{Z}}$ are independent.
\cite{Liu2006} gave conditions under which the model is stationary and ergodic. In the following, $\mathbf{M}_0$ is a $J^2 \times J^2$ matrix given by
\begin{equation*}
\left[ \mathbf{M}_0 \right]_{ij} := p_{ji,0} \left( \boldsymbol{\alpha}_0 \mathbf{e}_i^{\prime} + \boldsymbol{\beta}_0 \right), \quad i,j \in \{1,...,J\},
\end{equation*}
where $\boldsymbol{\alpha}_0 := (\alpha_{1,0},...,\alpha_{J,0})^{\prime}$, $\boldsymbol{\beta}_0 := \textup{diag}(\beta_{1,0},...,\beta_{J,0})$, and $\mathbf{e}_i$ is the $i$'th unit vector in $\mathbb{R}^J$.
\begin{theorem} [\cite{Liu2006}] \label{theo:se}
$(Y_t)_{t \in \mathbb{Z}}$ is stationary and ergodic with $\mathbb{E} [Y_t^2] < \infty$ if and only if $\rho(\mathbf{M}_0) < 1$ where $\rho(\mathbf{M}_0)$ is the spectral radius of $\mathbf{M}_0$.
\end{theorem}
\noindent
Note that $\rho(\mathbf{M}_0) < 1$ does not imply that $\alpha_{j,0} + \beta_{j,0} < 1$ for all $j \in \{1,...,J\}$, see \cite{Liu2006}.
We now give conditions under which the MLE is consistent and asymptotically normal. The MLE is consistent under the following assumptions.
\begin{assumption} \label{ass:one}
$\boldsymbol{\theta}_0 \in \boldsymbol{\Theta}$.
\end{assumption}
\begin{assumption} \label{ass:se}
$\rho(\mathbf{M}_0) < 1$.
\end{assumption}
\begin{assumption} \label{ass:compact}
$\boldsymbol{\Theta}$ is compact.
\end{assumption}
\begin{assumption} \label{ass:inv}
For all $\boldsymbol{\theta} \in \boldsymbol{\Theta}$, $\beta_j < 1$ for all $j \in \{1,...,J\}$.
\end{assumption}
\begin{assumption} \label{ass:identification}
For all $\boldsymbol{\theta} \in \boldsymbol{\Theta}$, there exists an $m \in \mathbb{N}$ such that $f^{(m)} (\mathbf{Y}_{t-m+1}^{t} ; \boldsymbol{\theta}) = f^{(m)} (\mathbf{Y}_{t-m+1}^{t} ; \boldsymbol{\theta}_0)$ a.s. implies that $\boldsymbol{\theta} = \boldsymbol{\theta}_0$.
\end{assumption}
\begin{theorem} \label{theo:c}
If Assumptions \ref{ass:one}-\ref{ass:identification} hold, then
\begin{equation*}
\hat{\boldsymbol{\theta}}_T \overset{a.s.}{\rightarrow} \boldsymbol{\theta}_0 \quad \text{as} \quad T \rightarrow \infty.
\end{equation*}
\end{theorem}
\begin{proof}
In the following, $\underline{\boldsymbol{\theta}} = \inf_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} \boldsymbol{\theta}$ and $\overline{\boldsymbol{\theta}} = \sup_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} \boldsymbol{\theta}$. Note that $(Y_t)_{t \in \mathbb{Z}}$ is stationary and ergodic by Theorem \ref{theo:se}. We therefore only need to verify Assumptions \ref{AssumptionTheta1}-\ref{AssumptionIdentification} and the conditions in Lemmas \ref{LemmaInvertibilityX} and \ref{LemmaInvertibilityFilter}.
Let $j \in \{1,...,J\}$ be given. First, Assumption \ref{AssumptionTheta1} is true by assumption and Assumptions \ref{AssumptionF1} and \ref{AssumptionPhi1} are trivially satisfied.
We now verify the conditions in Lemmas \ref{LemmaInvertibilityX} and \ref{LemmaInvertibilityFilter}. First, by Lemma 2.2 in \cite{StraumannMikosch2006},
\begin{equation*}
\mathbb{E} \left[ \log^{+} \sup_{\boldsymbol{\upsilon}_{j} \in \boldsymbol{\Upsilon}_{j}} \left| \phi_{j} (Y_t,x_j;\boldsymbol{\upsilon}_{j}) - x_j \right| \right] = \mathbb{E} \left[ \log^{+} \sup_{\boldsymbol{\upsilon}_{j} \in \boldsymbol{\Upsilon}_{j}} \left| \omega_{j} + \alpha_j Y_t^2 + \beta_j x_j - x_j \right| \right] \leq C_j + 2 \mathbb{E} \left[ \log^{+} \left| Y_t \right| \right]
\end{equation*}
for all $x_j \in \mathcal{X}_{j}$ where $C_j = 6 \log 2 + \log^{+} \overline{\omega}_{j} + \log^{+} \overline{\alpha}_{j} + \log^{+} \overline{\beta}_{j} + 2 \log^{+} x_j < \infty$, so Condition (i) in Lemma \ref{LemmaInvertibilityX} is satisfied since $\mathbb{E} [Y_t^2] < \infty$ implies that $\mathbb{E} [\log^{+} |Y_t|] < \infty$ by Lemma 2.2 in \cite{StraumannMikosch2006}. Moreover,
\begin{equation*}
\mathbb{E} \left[ \log \sup_{\boldsymbol{\upsilon}_{j} \in \boldsymbol{\Upsilon}_{j}} \Lambda_{j,t} (\boldsymbol{\upsilon}_{j}) \right] = \mathbb{E} \left[ \log \sup_{\boldsymbol{\upsilon}_{j} \in \boldsymbol{\Upsilon}_{j}} \beta_j \right] = \log \overline{\beta}_j,
\end{equation*}
so Conditions (ii) and (iii) in Lemma \ref{LemmaInvertibilityX} are also satisfied. Finally,
\begin{equation*}
\mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_j} \sup_{x_j \in \mathcal{X}_{j}} \left| \nabla_{x_j} \log f_{j} (Y_t;x_j,\boldsymbol{\upsilon}_j) \right| \right] = \mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_j} \sup_{x_j \in \mathcal{X}_{j}} \left| - \frac{1}{2} \frac{1}{x_j} + \frac{1}{2} \frac{Y_t^2}{x_j^2} \right| \right] \leq \frac{1}{2\underline{x}_j} + \frac{1}{2\underline{x}_j^2} \mathbb{E} \left[ Y_t^2 \right],
\end{equation*}
where $\underline{x}_j = \frac{\underline{\omega}_j}{1-\underline{\beta}_j} > 0$, so the condition in Lemma \ref{LemmaInvertibilityFilter} is also satisfied since $\mathbb{E} [Y_t^2] < \infty$.
Moreover, note that
\begin{align*}
\mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_j} \left| \log f_{j} (Y_t;X_{j,t}(\boldsymbol{\upsilon}_j),\boldsymbol{\upsilon}_j) \right| \right] &= \mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_j} \left| -\frac{1}{2} \log 2\pi - \frac{1}{2} \log X_{j,t}(\boldsymbol{\upsilon}_j) - \frac{1}{2} \frac{Y_t^2}{X_{j,t}(\boldsymbol{\upsilon}_j)} \right| \right] \\
&\leq \frac{1}{2} \log 2\pi + \frac{1}{2} \mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_j} \left| \log X_{j,t}(\boldsymbol{\upsilon}_j) \right| \right]+ \frac{1}{2\underline{x}_j} \mathbb{E} \left[ Y_t^2 \right],
\end{align*}
so Assumption \ref{AssumptionF2} is also satisfied since $\mathbb{E} [ \sup_{\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_j} | \log X_{j,t}(\boldsymbol{\upsilon}_j) | ] < \infty$ and $\mathbb{E} [ Y_t^2 ] < \infty$. Indeed, $\mathbb{E} [ Y_t^2 ] < \infty$ implies that $\mathbb{E} [ \sup_{\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_j} | \log X_{j,t}(\boldsymbol{\upsilon}_j) | ] < \infty$, which we now show. Note that $\log x = \log^{+} x - \log^{-} x$ for all $x > 0$ where $\log^{+} x = \max(\log x,0)$ and $\log^{-} x = -\min(\log x,0)$. Thus,
\begin{equation*}
\mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_j} \left| \log X_{j,t}(\boldsymbol{\upsilon}_j) \right| \right] \leq \mathbb{E} \left[ \log^{+} \sup_{\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_j} X_{j,t}(\boldsymbol{\upsilon}_j) \right] + \log^{-} \underline{x}_j.
\end{equation*}
Now, by Lemma \ref{LemmaInvertibilityX},
\begin{equation}
X_{j,t}(\boldsymbol{\upsilon}_j) = \frac{\omega_j}{1-\beta_j} + \alpha_j \sum_{i=0}^{\infty} \beta_j^i Y_{t-1-i}^2 \quad a.s. \label{eq:FZ}
\end{equation}
Thus,
\begin{equation*}
\mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_j} X_{j,t}(\boldsymbol{\upsilon}_j) \right] \leq \frac{\overline{\omega}_j}{1-\overline{\beta}_j} + \frac{\overline{\alpha}_j}{1-\overline{\beta}_j} \mathbb{E} \left[ Y_{t-1}^2 \right].
\end{equation*}
Hence, $\mathbb{E} [ Y_t^2 ] < \infty$ implies that $\mathbb{E} [ \sup_{\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_j} | \log X_{j,t}(\boldsymbol{\upsilon}_j) | ] < \infty$ since $\mathbb{E} [ \sup_{\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_j} X_{j,t}(\boldsymbol{\upsilon}_j) ] < \infty$ implies that $\mathbb{E} [ \log^{+} \sup_{\boldsymbol{\upsilon}_j \in \boldsymbol{\Upsilon}_j} X_{j,t}(\boldsymbol{\upsilon}_j) ] < \infty$ by Lemma 2.2 in \cite{StraumannMikosch2006}. Finally, Assumption \ref{AssumptionIdentification} is true by assumption.
\end{proof}
\noindent
Assumptions \ref{ass:one}-\ref{ass:identification} are identical to the assumptions in \cite{KandjiMisko2024}, except that \cite{KandjiMisko2024} assume that Assumption \ref{ass:identification} holds with $m = 1$.
Although the above result is identical to the one in \cite{KandjiMisko2024}, the proof of it is different. To prove consistency of the MLE, \cite{KandjiMisko2024} use one representation of the likelihood function, which is similar to the one used in \cite{FrancqRoussignol1998} and \cite{FrancqRoussignolZakoian2001} (see also Equation (12.25) in \cite{FrancqZakoian2019}), whereas we use another representation of the likelihood function, which is similar to the one used in \cite{DoucMoulinesRyden2004} and \cite{KasaharaShimotsu2019} (see also Equation (12.30) in \cite{FrancqZakoian2019}). The advantage of the latter is that it can also be used to prove asymptotic normality of the MLE, which we now do.
To give conditions under which the MLE is also asymptotically normal, we need the following result where, for all $s \in \mathbb{N}$, $\boldsymbol{\Sigma}_0^{(\otimes s)}$ is a $J^{s+1} \times J^{s+1}$ matrix given by
\begin{equation*}
\left[ \boldsymbol{\Sigma}_0^{(\otimes s)} \right]_{ij} := p_{ji,0} \mathbb{E} \left[ \mathbf{A}_{it,0}^{(\otimes s)} \right], \quad i,j \in \{1,...,J\}
\end{equation*}
with
\begin{equation*}
\mathbf{A}_{it,0} := \varepsilon_t^2 \boldsymbol{\alpha}_0 \mathbf{e}_i^{\prime} + \boldsymbol{\beta}_0,
\end{equation*}
where $\otimes$ denotes the Kronecker product.
\begin{theorem} [\cite{Liu2006}] \label{theo:m}
If $\rho(\mathbf{M}_0) < 1$ and $\rho(\boldsymbol{\Sigma}_0^{(\otimes s)}) < 1$, then $\mathbb{E} [Y_t^{2s}] < \infty$.
\end{theorem}
In addition to the above assumptions, the MLE is asymptotically normal under the following assumptions.
\begin{assumption} \label{ass:int}
$\boldsymbol{\theta}_0 \in \textup{int}(\boldsymbol{\Theta})$.
\end{assumption}
\begin{assumption} \label{ass:m}
$\rho(\boldsymbol{\Sigma}_0^{(\otimes 3)}) < 1$.
\end{assumption}
\begin{theorem} \label{theo:an}
If Assumptions \ref{ass:one}-\ref{ass:m} hold and $\mathbf{I}(\boldsymbol{\theta}_0)$ is invertible, then
\begin{equation*}
\sqrt{T} (\hat{\boldsymbol{\theta}}_T - \boldsymbol{\theta}_0) \overset{d}{\rightarrow} \mathcal{N} (\mathbf{0},\mathbf{I}(\boldsymbol{\theta}_0)^{-1}) \quad \text{as} \quad T \rightarrow \infty.
\end{equation*}
\end{theorem}
\begin{proof}
In the following, $\underline{\boldsymbol{\theta}} = \inf_{\boldsymbol{\theta} \in \bar{\boldsymbol{\Theta}}} \boldsymbol{\theta}$ and $\overline{\boldsymbol{\theta}} = \sup_{\boldsymbol{\theta} \in \bar{\boldsymbol{\Theta}}} \boldsymbol{\theta}$. We need to verify Assumptions \ref{AssumptionTheta2}-\ref{AssumptionF4} and the conditions in Lemmas \ref{LemmaInvertibilityDX}-\ref{LemmaInvertibilityDDPrediction}.
Let, as in the proof of Theorem \ref{theo:c}, $j \in \{1,...,J\}$ be given. First, Assumption \ref{AssumptionTheta2} is true by assumption and Assumptions \ref{AssumptionF3} and \ref{AssumptionPhi2} are trivially satisfied.
We now verify the conditions in Lemma \ref{LemmaInvertibilityDX}. First,
\begin{align*}
\mathbb{E} \left[ \log^{+} \sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \left| \bar{\nabla}_{\boldsymbol{\upsilon}_j} \phi_{j} (Y_t,X_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) \right| \right|_{2} \right] &\leq \mathbb{E} \left[ \log^{+} \sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} \left( 1 + Y_t^2 + X_{j,t} (\boldsymbol{\upsilon}_{j}) \right) \right] \\
&\leq 4 \log 2 + 2 \mathbb{E} \left[ \log^{+} \left| Y_t \right| \right] + \mathbb{E} \left[ \log^{+} \sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} X_{j,t} (\boldsymbol{\upsilon}_{j}) \right],
\end{align*}
so Condition (i) is satisfied since $\mathbb{E} [ Y_t^6 ] < \infty$ implies that $\mathbb{E} [ \log^{+} | Y_t | ] < \infty$ and $\mathbb{E} [ \log^{+} \sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} X_{j,t} (\boldsymbol{\upsilon}_{j}) ] < \infty$, see the proof of Theorem \ref{theo:c}. We now show that $\mathbb{E} [ \sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} | | \nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) | |_{2}^{2} ] < \infty$, which implies that $\mathbb{E} [ \log^{+} \sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} | | \nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) | |_{2} ] < \infty$. Note that
\begin{equation*}
\nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) = \sum_{i=0}^{\infty} \beta_j^i \begin{bmatrix} 1 \\ Y_{t-1-i}^2 \\ X_{j,t-1-i} (\boldsymbol{\upsilon}_j) \end{bmatrix} \quad a.s.,
\end{equation*}
so, by Minkowskis inequality,
\begin{equation*}
\mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \left| \nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) \right| \right|_{2}^{2} \right]^{\frac{1}{2}} \leq \frac{1}{1-\overline{\beta}_j} + \frac{1}{1-\overline{\beta}_j} \mathbb{E} \left[ Y_{t-1}^4 \right]^{\frac{1}{2}} + \frac{1}{1-\overline{\beta}_j} \mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} X_{j,t-1}^{2} (\boldsymbol{\upsilon}_j) \right]^{\frac{1}{2}}.
\end{equation*}
Hence, $\mathbb{E} [ \sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} || \nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) ||_{2}^{2} ] < \infty$ since $\mathbb{E} [ Y_t^6] < \infty$ and $\mathbb{E}[ \sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} X_{j,t}^{2} (\boldsymbol{\upsilon}_j)] < \infty$, see the proof of Theorem \ref{theo:c}. Finally,
\begin{equation*}
\sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \left| \bar{\nabla}_{\boldsymbol{\upsilon}_j} \phi_{j} (Y_t,\hat{X}_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) - \bar{\nabla}_{\boldsymbol{\upsilon}_j} \phi_{j} (Y_t,X_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) \right| \right|_{2} = \sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \hat{X}_{j,t} (\boldsymbol{\upsilon}_{j}) - X_{j,t} (\boldsymbol{\upsilon}_{j}) \right|
\end{equation*}
and
\begin{equation*}
\sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \bar{\nabla}_{x_j} \phi_{j} (Y_t,\hat{X}_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) - \bar{\nabla}_{x_j} \phi_{j} (Y_t,X_{j,t} (\boldsymbol{\upsilon}_{j});\boldsymbol{\upsilon}_{j}) \right| = 0,
\end{equation*}
so Conditions (ii) and (iii) are also satisfied since $\sup_{\boldsymbol{\upsilon}_{j} \in \bar{\boldsymbol{\Upsilon}}_{j}} | \hat{X}_{j,t} (\boldsymbol{\upsilon}_{j}) - X_{j,t} (\boldsymbol{\upsilon}_{j}) | \overset{e.a.s.}{\rightarrow} 0$ as $t \rightarrow \infty$. The conditions in Lemma \ref{LemmaInvertibilityDDX} can be verified similarly.
Moving on to Assumption \ref{AssumptionF4}, by Hölders inequality, for all $m \in (1,\frac{6}{4}]$,
\begin{align*}
&\mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \bar{\nabla}_{x_j} \log f_{j} (Y_t;X_{j,t} (\boldsymbol{\upsilon}_j),\boldsymbol{\upsilon}_j) \right|^2 \left| \left| \nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) \right| \right|_{2}^2 \right] \\
&= \mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| - \frac{1}{2} + \frac{1}{2} \frac{Y_t^2}{X_{j,t}(\boldsymbol{\upsilon}_j)} \right|^2 \left| \left| \frac{\nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j)}{X_{j,t}(\boldsymbol{\upsilon}_j)} \right| \right|_{2}^2 \right] \\
&\leq \mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| - \frac{1}{2} + \frac{1}{2} \frac{Y_t^2}{X_{j,t}(\boldsymbol{\upsilon}_j)} \right|^{2m} \right]^{\frac{1}{m}} \mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \left| \frac{\nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j)}{X_{j,t}(\boldsymbol{\upsilon}_j)} \right| \right|_{2}^{2m^c} \right]^{\frac{1}{m^c}} \\
&\leq \left( 1 + \frac{1}{\underline{x}_j^{2m}} \mathbb{E} \left[ \left|Y_t\right|^{4m} \right] \right)^{\frac{1}{m}} \mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \left| \frac{\nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j)}{X_{j,t}(\boldsymbol{\upsilon}_j)} \right| \right|_{2}^{2m^c} \right]^{\frac{1}{m^c}}
\end{align*}
with $\underline{x}_j = \frac{\underline{\omega}_j}{1-\underline{\beta}_j}$ and $m^c = \frac{m}{m-1}$ where the last inequality follows from the inequality $|x+y|^p \leq 2^p |x|^p + 2^p |y|^p$ for all $x,y \in (-\infty,\infty)$ and $p \in (0,\infty)$. Note that, by Equation \eqref{eq:FZ}, for all $n \in (0,\frac{1}{m^c})$,
\begin{align*}
\frac{\nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j)}{X_{j,t} (\boldsymbol{\upsilon}_j)}
&=
\begin{bmatrix}
\frac{1}{1-\beta_j} \frac{1}{X_{j,t} (\boldsymbol{\upsilon}_j)} \\
\left( \sum_{i=0}^{\infty} \beta_j^{i} Y_{t-1-i}^2 \right) \frac{1}{X_{j,t} (\boldsymbol{\upsilon}_j)} \\
\left( \frac{\omega_j}{(1-\beta_j)^2} + \alpha_j \sum_{i=0}^{\infty} i \beta_j^{i-1} Y_{t-1-i}^2 \right) \frac{1}{X_{j,t} (\boldsymbol{\upsilon}_j)}
\end{bmatrix} \\
&\leq
\begin{bmatrix}
\frac{1}{\omega_j} \\
\frac{1}{\alpha_j} \\
\frac{1}{1-\beta_j} + \frac{1}{\beta_j} \sum_{i=0}^{\infty} i \frac{\alpha_j \beta_j^{i} Y_{t-1-i}^2}{\frac{\omega_j}{1-\beta_j} + \alpha_j \beta_j^{i} Y_{t-1-i}^2}
\end{bmatrix} \\
&\leq
\begin{bmatrix}
\frac{1}{\omega_j} \\
\frac{1}{\alpha_j} \\
\frac{1}{1-\beta_j} + \frac{1}{\omega_j^n} \alpha_j^n \frac{(1-\beta_j)^n}{\beta_j} \sum_{i=0}^{\infty} i \beta_j^{ni} \left|Y_{t-1-i}\right|^{2n}
\end{bmatrix} \quad a.s.,
\end{align*}
where the inequalities are elementwise and the last one follows from the inequality $\frac{x}{1+x} \leq x^p$ for all $x \in [0,\infty)$ and $p \in (0,1)$, so, by Minkowskis inequality,
\begin{equation*}
\mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \left| \frac{\nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j)}{X_{j,t}(\boldsymbol{\upsilon}_j)} \right| \right|_{2}^{2m^c} \right]^{\frac{1}{2m^c}} \leq A_j + B_j \mathbb{E} \left[ \left|Y_{t-1} \right|^{4nm^c} \right]^{\frac{1}{2m^c}},
\end{equation*}
where $A_j = \frac{1}{\underline{\omega}_j} + \frac{1}{\underline{\alpha}_j} + \frac{1}{1-\overline{\beta}_j}$ and $B_j = \frac{1}{\underline{\omega}_j^n} \overline{\alpha}_j^n \frac{(1-\underline{\beta}_j)^n \overline{\beta}_j^n}{\underline{\beta}_j (1-\overline{\beta}_j^n)^2}$. Hence, $\mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \left| \frac{\nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j)}{X_{j,t}(\boldsymbol{\upsilon}_j)} \right| \right|_{2}^{2m^c} \right] < \infty$ since $\mathbb{E} [ Y_t^6] < \infty$. The first condition in Assumption \ref{AssumptionF4} is thus satisfied since $\mathbb{E} [ Y_t^6] < \infty$ and $\mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \left| \frac{\nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j)}{X_{j,t}(\boldsymbol{\upsilon}_j)} \right| \right|_{2}^{2m^c} \right] < \infty$. Moreover,
\begin{align*}
&\mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \bar{\nabla}_{x_j x_j} \log f_{j} (Y_t;{X}_{j,t} (\boldsymbol{\upsilon}_j),\boldsymbol{\upsilon}_j) \right| \left| \left| \nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) \right| \right|_{2}^2 \right] \\
&= \mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \frac{1}{2} - \frac{Y_t^2}{{X}_{j,t} (\boldsymbol{\upsilon}_j)} \right| \left| \left| \frac{\nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j)}{{X}_{j,t} (\boldsymbol{\upsilon}_j)} \right| \right|_{2}^2 \right] \\
&\leq \mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \frac{1}{2} - \frac{Y_t^2}{{X}_{j,t} (\boldsymbol{\upsilon}_j)} \right|^{m} \right]^{\frac{1}{m}} \mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \left| \frac{\nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j)}{{X}_{j,t} (\boldsymbol{\upsilon}_j)} \right| \right|_{2}^{2m^c} \right]^{\frac{1}{m^c}} \\
&\leq \left( 1 + \frac{2^m}{\underline{x}_j^{m}} \mathbb{E} \left[ \left|Y_t\right|^{2m} \right] \right)^{\frac{1}{m}} \mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \left| \frac{\nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j)}{{X}_{j,t} (\boldsymbol{\upsilon}_j)} \right| \right|_{2}^{2m^c} \right]^{\frac{1}{m^c}},
\end{align*}
so the second condition in Assumption \ref{AssumptionF4} is also satisfied. Finally, by the Cauchy–Schwarz inequality,
\begin{align*}
&\mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \bar{\nabla}_{x_j} \log f_{j} (Y_t;{X}_{j,t} (\boldsymbol{\upsilon}_j),\boldsymbol{\upsilon}_j) \right| \left| \left| \nabla_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) \right| \right|_{2,2} \right] \\
&= \mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| - \frac{1}{2} \frac{1}{X_{j,t}(\boldsymbol{\upsilon}_j)} + \frac{1}{2} \frac{Y_t^2}{X_{j,t}^2(\boldsymbol{\upsilon}_j)} \right| \left| \left| \nabla_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) \right| \right|_{2,2} \right] \\
& \leq \mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| - \frac{1}{2} \frac{1}{X_{j,t}(\boldsymbol{\upsilon}_j)} + \frac{1}{2} \frac{Y_t^2}{X_{j,t}^2(\boldsymbol{\upsilon}_j)} \right|^2 \right]^{\frac{1}{2}} \mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \left| \nabla_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) \right| \right|_{2,2}^{2} \right]^{\frac{1}{2}} \\
& \leq \left( \frac{1}{\underline{x}_j^2} + \frac{1}{\underline{x}_j^4} \mathbb{E} \left[ Y_t^4 \right] \right)^{\frac{1}{2}} \mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \left| \nabla_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) \right| \right|_{2,2}^{2} \right]^{\frac{1}{2}},
\end{align*}
where the last inequality follows from the inequality $|x+y|^p \leq 2^p |x|^p + 2^p |y|^p$ for all $x,y \in (-\infty,\infty)$ and $p \in (0,\infty)$ as above. Note that
\begin{equation*}
\nabla_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) = \sum_{i=0}^{\infty} \beta_j^i \left( \nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) \begin{bmatrix} 0 & 0 & 1 \end{bmatrix} + \begin{bmatrix} 0 \\ 0 \\ 1 \end{bmatrix} \nabla_{\boldsymbol{\upsilon}_j^{\prime}} X_{j,t} (\boldsymbol{\upsilon}_j) \right) \quad a.s.,
\end{equation*}
so, by Minkowskis inequality,
\begin{equation*}
\mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \left| \nabla_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) \right| \right|_{2,2}^{2} \right]^{\frac{1}{2}} \leq \frac{2}{1-\overline{\beta}_j} \mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} \left| \left| \nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) \right| \right|_{2}^{2} \right]^{\frac{1}{2}}.
\end{equation*}
Hence, $\mathbb{E} [ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} || \nabla_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) ||_{2,2}^{2} ] < \infty$ since $\mathbb{E} [ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} || \nabla_{\boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) ||_{2}^{2} ] < \infty$. The third condition in Assumption \ref{AssumptionF4} is thus also satisfied since $\mathbb{E} [ Y_t^6 ] < \infty$ and $\mathbb{E} [ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_{j}} || \nabla_{\boldsymbol{\upsilon}_j \boldsymbol{\upsilon}_j} X_{j,t} (\boldsymbol{\upsilon}_j) ||_{2,2}^{2} ] < \infty$.
Finally, we verify the condition in Lemma \ref{LemmaInvertibilityDPrediction}. Note that
\begin{align*}
\mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_j} \sup_{x_j \in \mathcal{X}_{j}} \left| \bar{\nabla}_{x_j x_j} \log f_{j} (Y_t;x_j,\boldsymbol{\upsilon}_j) \right| \right] &= \mathbb{E} \left[ \sup_{\boldsymbol{\upsilon}_j \in \bar{\boldsymbol{\Upsilon}}_j} \sup_{x_j \in \mathcal{X}_{j}} \left| \frac{1}{2} \frac{1}{x_j^2} - \frac{Y_t^2}{x_j^3} \right| \right] \\
&\leq \frac{1}{2\underline{x}_j^2} + \frac{1}{\underline{x}_j^3} \mathbb{E} \left[ Y_t^2 \right],
\end{align*}
so that condition is also satisfied since $\mathbb{E} [ Y_t^6 ] < \infty$. The condition in Lemma \ref{LemmaInvertibilityDDPrediction} can be verified similarly.
\end{proof}
\begin{remark}
An inspection of the proof of Theorem \ref{theo:an} shows that we only need to assume that there exists an $\varepsilon > 0$ such that $\mathbb{E} [|Y_t|^{4+\varepsilon}] < \infty$. However, in the literature, there exist to the best of our knowledge no conditions under which this is directly the case. We thus assume that $\rho(\boldsymbol{\Sigma}_0^{(\otimes 3)}) < 1$, which implies that $\mathbb{E} [Y_t^{6}] < \infty$, a stronger assumption than the assumption that there exists an $\varepsilon > 0$ such that $\mathbb{E} [|Y_t|^{4+\varepsilon}] < \infty$.
\end{remark}
\subsection{Finite-sample Properties}
To study the finite-sample properties of the MLE for the Markov-switching GARCH model by \cite{HaasMittnikPaolella2004b}, we perform a Monte Carlo simulation study.
Tables \ref{tab:MC1} and \ref{tab:MC2} report the estimated means and standard deviations (in parentheses) of the estimated parameters of a range of two-state Markov-switching GARCH models together with the true parameters. The benchmark model is a two-state Markov-switching GARCH model with true parameters $\omega_{1,0} = 0.025$, $\alpha_{1,0} = 0.05$, $\beta_{1,0} = 0.90$, $\omega_{2,0} = 0.25$, $\alpha_{2,0} = 0.30$, $\beta_{2,0} = 0.60$, $p_{11,0} = 0.95$, and $p_{22,0} = 0.90$, which are similar to the estimated parameters often found when a two-state Markov-switching GARCH model is estimated on data. Indeed, when a two-state Markov-switching GARCH model is estimated on data, one state is often a persistent low-volatility state in which $\alpha$ is relatively low and $\beta$ is relatively high. The other is often a persistent high-volatility state in which $\alpha$ is relatively high and $\beta$ is relatively low, which, according to \cite{HaasMittnikPaolella2004b}, may indicate a tendency to overreact to news possibly due to a prevailing panic-like mood. The means and standard deviations of the estimated parameters are estimated from $2500$ replications where a replication consists of first simulating $T$ observations from the model and then estimating the model from the $T$ observations; the R package MSGARCH developed by \cite{ArdiaBluteauBoudtCataniaTrottier2019} is used for both purposes. The finite-sample properties of the MLE for the benchmark model are good. Indeed, the estimated means of the estimated parameters converge to the true parameters and the estimated standard deviations of the estimated parameters converge to zero.
Table \ref{tab:MC1} investigates both what happens when the two states of the benchmark model are more similar and what happens when the two states of the benchmark model are more different. In both cases, the finite-sample properties of the MLE are good; best, however, in the case where the two states are more different since, in this case, it is easier to determine whether an observation is from one state or the other making it easier to estimate the parameters in the two states.
\begin{table}[t]
\centering
\tiny
\begin{tabular}{ccccccccc}
\toprule
& $\omega_{1,0} = 0.025$ & $\alpha_{1,0} = 0.05$ & $\beta_{1,0} = 0.90$ & $\omega_{2,0} = 0.50$ & $\alpha_{2,0} = 0.40$ & $\beta_{2,0} = 0.50$ & $p_{11,0} = 0.95$ & $p_{22,0} = 0.90$ \\
\cmidrule(lr){1-9}
$T = 1250$ & $\underset{(0.108)}{0.059}$ & $\underset{(0.056)}{0.065}$ & $\underset{(0.174)}{0.844}$ & $\underset{(0.621)}{0.716}$ & $\underset{(0.148)}{0.402}$ & $\underset{(0.178)}{0.443}$ & $\underset{(0.071)}{0.945}$ & $\underset{(0.100)}{0.896}$ \vspace{0.1cm} \\
$T = 2500$ & $\underset{(0.042)}{0.031}$ & $\underset{(0.027)}{0.053}$ & $\underset{(0.078)}{0.889}$ & $\underset{(0.268)}{0.564}$ & $\underset{(0.092)}{0.403}$ & $\underset{(0.110)}{0.478}$ & $\underset{(0.046)}{0.944}$ & $\underset{(0.067)}{0.894}$ \vspace{0.1cm} \\
$T = 5000$ & $\underset{(0.009)}{0.026}$ & $\underset{(0.013)}{0.050}$ & $\underset{(0.021)}{0.900}$ & $\underset{(0.145)}{0.518}$ & $\underset{(0.061)}{0.400}$ & $\underset{(0.071)}{0.494}$ & $\underset{(0.022)}{0.947}$ & $\underset{(0.040)}{0.896}$ \vspace{0.1cm} \\
\cmidrule(lr){1-9}
& $\omega_{1,0} = 0.025$ & $\alpha_{1,0} = 0.05$ & $\beta_{1,0} = 0.90$ & $\omega_{2,0} = 0.25$ & $\alpha_{2,0} = 0.30$ & $\beta_{2,0} = 0.60$ & $p_{11,0} = 0.95$ & $p_{22,0} = 0.90$ \\
\cmidrule(lr){1-9}
$T = 1250$ & $\underset{(0.112)}{0.077}$ & $\underset{(0.060)}{0.070}$ & $\underset{(0.222)}{0.796}$ & $\underset{(0.553)}{0.500}$ & $\underset{(0.155)}{0.302}$ & $\underset{(0.217)}{0.509}$ & $\underset{(0.105)}{0.948}$ & $\underset{(0.119)}{0.908}$ \vspace{0.1cm} \\
$T = 2500$ & $\underset{(0.069)}{0.046}$ & $\underset{(0.041)}{0.059}$ & $\underset{(0.143)}{0.858}$ & $\underset{(0.302)}{0.350}$ & $\underset{(0.110)}{0.309}$ & $\underset{(0.152)}{0.555}$ & $\underset{(0.078)}{0.944}$ & $\underset{(0.100)}{0.896}$ \vspace{0.1cm} \\
$T = 5000$ & $\underset{(0.025)}{0.029}$ & $\underset{(0.022)}{0.052}$ & $\underset{(0.056)}{0.892}$ & $\underset{(0.160)}{0.289}$ & $\underset{(0.070)}{0.305}$ & $\underset{(0.096)}{0.580}$ & $\underset{(0.051)}{0.944}$ & $\underset{(0.070)}{0.893}$ \vspace{0.1cm} \\
\cmidrule(lr){1-9}
& $\omega_{1,0} = 0.025$ & $\alpha_{1,0} = 0.05$ & $\beta_{1,0} = 0.90$ & $\omega_{2,0} = 0.125$ & $\alpha_{2,0} = 0.20$ & $\beta_{2,0} = 0.70$ & $p_{11,0} = 0.95$ & $p_{22,0} = 0.90$ \\
\cmidrule(lr){1-9}
$T = 1250$ & $\underset{(0.112)}{0.093}$ & $\underset{(0.068)}{0.068}$ & $\underset{(0.260)}{0.741}$ & $\underset{(0.410)}{0.335}$ & $\underset{(0.146)}{0.191}$ & $\underset{(0.263)}{0.587}$ & $\underset{(0.121)}{0.954}$ & $\underset{(0.132)}{0.925}$ \vspace{0.1cm} \\
$T = 2500$ & $\underset{(0.084)}{0.066}$ & $\underset{(0.049)}{0.065}$ & $\underset{(0.205)}{0.801}$ & $\underset{(0.371)}{0.282}$ & $\underset{(0.134)}{0.204}$ & $\underset{(0.224)}{0.613}$ & $\underset{(0.114)}{0.951}$ & $\underset{(0.124)}{0.918}$ \vspace{0.1cm} \\
$T = 5000$ & $\underset{(0.048)}{0.043}$ & $\underset{(0.034)}{0.058}$ & $\underset{(0.123)}{0.859}$ & $\underset{(0.243)}{0.217}$ & $\underset{(0.094)}{0.207}$ & $\underset{(0.168)}{0.643}$ & $\underset{(0.081)}{0.953}$ & $\underset{(0.112)}{0.909}$ \vspace{0.1cm} \\
\cmidrule(lr){1-9}
\bottomrule
\end{tabular}
\captionsetup{font=footnotesize}
\caption{The estimated means and standard deviations (in parentheses) of the estimated parameters of three two-state Markov-switching GARCH models together with the true parameters.}
\label{tab:MC1}
\end{table}
Table \ref{tab:MC2} investigates what happens when the second state of the benchmark model is less persistent. Although the finite-sample properties of the MLEs in the first state are good, the finite-sample properties of the MLEs in the second state are, somewhat surprisingly, not entirely satisfactory. There is, however, a natural explanation for this. Because the second state is less persistent, it is less likely to observe a relatively long sequence of consecutive observations from the second state making it more difficult to estimate the parameters in the second state; something which is important to keep in mind when a two-state Markov-switching GARCH model is applied to data.
\begin{table}[t]
\centering
\tiny
\begin{tabular}{ccccccccc}
\toprule
& $\omega_{1,0} = 0.025$ & $\alpha_{1,0} = 0.05$ & $\beta_{1,0} = 0.90$ & $\omega_{2,0} = 0.25$ & $\alpha_{2,0} = 0.30$ & $\beta_{2,0} = 0.60$ & $p_{11,0} = 0.95$ & $p_{22,0} = 0.90$ \\
\cmidrule(lr){1-9}
$T = 1250$ & $\underset{(0.112)}{0.077}$ & $\underset{(0.060)}{0.070}$ & $\underset{(0.222)}{0.796}$ & $\underset{(0.553)}{0.500}$ & $\underset{(0.155)}{0.302}$ & $\underset{(0.217)}{0.509}$ & $\underset{(0.105)}{0.948}$ & $\underset{(0.119)}{0.908}$ \vspace{0.1cm} \\
$T = 2500$ & $\underset{(0.069)}{0.046}$ & $\underset{(0.041)}{0.059}$ & $\underset{(0.143)}{0.858}$ & $\underset{(0.302)}{0.350}$ & $\underset{(0.110)}{0.309}$ & $\underset{(0.152)}{0.555}$ & $\underset{(0.078)}{0.944}$ & $\underset{(0.100)}{0.896}$ \vspace{0.1cm} \\
$T = 5000$ & $\underset{(0.025)}{0.029}$ & $\underset{(0.022)}{0.052}$ & $\underset{(0.056)}{0.892}$ & $\underset{(0.160)}{0.289}$ & $\underset{(0.070)}{0.305}$ & $\underset{(0.096)}{0.580}$ & $\underset{(0.051)}{0.944}$ & $\underset{(0.070)}{0.893}$ \vspace{0.1cm} \\
\cmidrule(lr){1-9}
& $\omega_{1,0} = 0.025$ & $\alpha_{1,0} = 0.05$ & $\beta_{1,0} = 0.90$ & $\omega_{2,0} = 0.25$ & $\alpha_{2,0} = 0.30$ & $\beta_{2,0} = 0.60$ & $p_{11,0} = 0.95$ & $p_{22,0} = 0.70$ \\
\cmidrule(lr){1-9}
$T = 1250$ & $\underset{(0.114)}{0.087}$ & $\underset{(0.054)}{0.054}$ & $\underset{(0.266)}{0.757}$ & $\underset{(0.485)}{0.480}$ & $\underset{(0.204)}{0.228}$ & $\underset{(0.295)}{0.472}$ & $\underset{(0.101)}{0.956}$ & $\underset{(0.167)}{0.881}$ \vspace{0.1cm} \\
$T = 2500$ & $\underset{(0.075)}{0.053}$ & $\underset{(0.032)}{0.051}$ & $\underset{(0.179)}{0.838}$ & $\underset{(0.418)}{0.457}$ & $\underset{(0.175)}{0.259}$ & $\underset{(0.258)}{0.476}$ & $\underset{(0.090)}{0.953}$ & $\underset{(0.179)}{0.838}$ \vspace{0.1cm} \\
$T = 5000$ & $\underset{(0.050)}{0.037}$ & $\underset{(0.018)}{0.050}$ & $\underset{(0.116)}{0.875}$ & $\underset{(0.335)}{0.398}$ & $\underset{(0.138)}{0.284}$ & $\underset{(0.204)}{0.500}$ & $\underset{(0.072)}{0.949}$ & $\underset{(0.173)}{0.786}$ \vspace{0.1cm} \\
\cmidrule(lr){1-9}
& $\omega_{1,0} = 0.025$ & $\alpha_{1,0} = 0.05$ & $\beta_{1,0} = 0.90$ & $\omega_{2,0} = 0.25$ & $\alpha_{2,0} = 0.30$ & $\beta_{2,0} = 0.60$ & $p_{11,0} = 0.95$ & $p_{22,0} = 0.50$ \\
\cmidrule(lr){1-9}
$T = 1250$ & $\underset{(0.113)}{0.092}$ & $\underset{(0.058)}{0.051}$ & $\underset{(0.286)}{0.732}$ & $\underset{(0.416)}{0.405}$ & $\underset{(0.197)}{0.160}$ & $\underset{(0.332)}{0.516}$ & $\underset{(0.116)}{0.955}$ & $\underset{(0.186)}{0.890}$ \vspace{0.1cm} \\
$T = 2500$ & $\underset{(0.084)}{0.063}$ & $\underset{(0.039)}{0.050}$ & $\underset{(0.220)}{0.807}$ & $\underset{(0.424)}{0.433}$ & $\underset{(0.199)}{0.187}$ & $\underset{(0.314)}{0.499}$ & $\underset{(0.082)}{0.965}$ & $\underset{(0.202)}{0.858}$ \vspace{0.1cm} \\
$T = 5000$ & $\underset{(0.057)}{0.042}$ & $\underset{(0.022)}{0.048}$ & $\underset{(0.146)}{0.862}$ & $\underset{(0.384)}{0.440}$ & $\underset{(0.188)}{0.227}$ & $\underset{(0.285)}{0.475}$ & $\underset{(0.073)}{0.961}$ & $\underset{(0.232)}{0.787}$ \vspace{0.1cm} \\
\cmidrule(lr){1-9}
\bottomrule
\end{tabular}
\captionsetup{font=footnotesize}
\caption{The estimated means and standard deviations (in parentheses) of the estimated parameters of three two-state Markov-switching GARCH models together with the true parameters.}
\label{tab:MC2}
\end{table}
\section{Conclusion} \label{Conclusion}
State space models and their extensions are ubiquitous in economics and finance, so statistical inference for them – including estimation, which is typically done by maximum likelihood estimation – is of significant practical importance.
In this paper, we proved both consistency and asymptotic normality of the MLE for a Markov-switching observation-driven model, that is, an observation-driven state space model where $\textup{X}$ is finite. To the best of our knowledge, these results are the first of their kind in the literature. As a special case, we also gave conditions under which the MLE for the widely applied Markov-switching GARCH model by \cite{HaasMittnikPaolella2004b} is both consistent and asymptotically normal. Again, this is new to the literature and extends \cite{KandjiMisko2024}, who gave conditions under which it is only consistent.
Several extensions of this paper are possible. One is to generalise the results in this paper to a higher-order Markov-switching observation-driven model, that is, a model where $X_{j,t+1} = \phi_{j} (Y_{t},...,Y_{t-q+1},X_{j,t},...,X_{j,t-p+1};\boldsymbol{\upsilon}_{j})$. Another is to generalise them to a multivariate Markov-switching observation-driven model. Both follow by using similar arguments to the ones in this paper. A final, and very interesting, extension of this paper is to generalise the results to an observation-driven state space model where $\textup{X}$ is not necessarily finite to cover more of the examples in the introduction. We conjecture that the results in this paper can be extended to an observation-driven state space model where $\textup{X}$ is compact relatively easily by combining the arguments in this paper and the ones in \cite{DoucMoulinesRyden2004} and \cite{KasaharaShimotsu2019}. We, however, leave this for future research.
\bibliography{references}