EconBase
← Back to paper

Maximum Likelihood Estimation in Markov Regime-Switching Models with Covariate-Dependent Transition Probabilities

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

80,458 characters

Maximum Likelihood Estimation in Markov Regime-Switching Models with Covariate-Dependent Transition Probabilities



\title{Maximum Likelihood Estimation in Markov Regime-Switching Models with
Covariate-Dependent Transition Probabilities\thanks{We would like to thank
Noureddine El Karoui, Adityanand Guntuboyina, Michael Jansson, Ulrich
M\"{u}ller, and three anonymous referees for useful comments and suggestions.
All errors are our own.}}
\author{Demian Pouzo,\thanks{ e-mail: [email removed] (corresponding author).
Address: UC Berkeley, Dept. of Economics; 530 Evans Hall, Berkeley CA 94619,
USA.} Zacharias Psaradakis\thanks{ e-mail: [email removed]. Address:
Birkbeck, University of London, Dept. of Economics, Mathematics \& Statistics;
Malet Street, London WC1E 7HX, UK} and Martin Sola\thanks{ e-mail:
[email removed]. Address: Universidad Torcuato Di Tella, Dept. of Economics;
Figueroa Alcorta 7350, C1428 Buenos Aires, Argentina.}}
\maketitle

\begin{abstract}
\noindent This paper considers maximum likelihood (ML) estimation in a large
class of models with hidden Markov regimes. We investigate consistency of the
ML estimator and local asymptotic normality for the models under general
conditions which allow for autoregressive dynamics in the observable process,
Markov regime sequences with covariate-dependent transition matrices, and
possible model misspecification. A Monte Carlo study examines the
finite-sample properties of the ML estimator in correctly specified and
misspecified models. An empirical application is also discussed.\bigskip

\noindent\noindent\textit{Key words and phrases}: Autoregressive model;
consistency; covariate-dependent transition probabilities; hidden Markov
model; Markov-switching model; maximum likelihood; local asymptotic normality;
misspecified models.\bigskip\newpage

\end{abstract}

\section{Introduction}

Stochastic models with parameters that are subject to changes driven by an
unobservable Markov chain (the regime or state sequence) have attracted
considerable attention in many different areas, the influential work by
\cite{hamilton89} being a prominent example from the econometrics literature.
An important subclass of such models, the so-called hidden Markov models, in
which observations are conditionally independent given the regime sequence,
are also widely used in a variety of disciplines. A common assumption in these
models is that the unobservable Markov chain is temporally homogeneous.

In this paper, we focus on a larger class of models in which the hidden regime
process and the observation process (conditional on the regimes) are both
temporally inhomogeneous Markov chains. This is a useful generalization of
models with a time-invariant transition mechanism that has found numerous
applications, especially in economics and finance.\footnote{Examples include,
among many others, applications to the analysis of business-cycle fluctuations
(e.g., \cite{filardo94}, \cite{FilardoGordon98}, \cite{ravn99},
\cite{Simpson01}, \cite{rivas15}), interest rates and yields (e.g.,
\cite{Gray96}, \cite{BekaertH01}, \cite{ang02}, \cite{angbek02j},
\cite{pses21}), consumption growth (\cite{Whitelaw00}), currency crises (e.g.,
\cite{peria02}, \cite{mourat08}), and cryptocurrency returns (\cite{tan21}).}
Statistical inference in this class of models is predominantly likelihood
based, even though very little is known about the asymptotic properties of the
relevant inferential procedures. In a typical application, inference is
conducted on the implicit assumption that the maximum likelihood (ML)
estimator of unknown parameters has its familiar properties of consistency,
asymptotic normality, and asymptotic efficiency, and associated confidence
sets and hypotheses tests are constructed in the usual manner. It hardly needs
noting that, unless these properties of likelihood-based inferential
procedures known from regular parametric estimation problems hold for Markov
regime-switching models with covariate-dependent transition probabilities,
inferences drawn from them cannot be justified in any meaningful way and
should be interpreted very cautiously. Arguing, for instance, as is common in
applied work, that an economic variable is a useful leading indicator for
business-cycle phases because it appears to have a `statistically significant'
coefficient in the transition functions of a Markov-switching model for output
growth is problematic when little is known about the properties of the
relevant estimators and related tests.

The main contribution of this paper is to provide consistency and asymptotic
normality results for a large class of models that are relevant in
applications. Our approach allows for autoregressive dynamics in the
observable process, covariate-dependence in the transition functions of the
hidden regime process, and potential model misspecification. To the best of
our knowledge, the only asymptotic results available on ML estimation in
Markov regime-switching models with covariate-dependent transition
probabilities are those of \cite{aill13}, who investigate consistency of the
ML estimator in a correctly specified model (i.e., a model which contains the
data-generating process). Our results include both consistency of the ML
estimator and local asymptotic normality (LAN) for the model, from which
asymptotic normality of the ML estimator can be inferred. Unlike
\cite{aill13}, who allow for a general hidden state space, we require the
latter to be finite, but do not restrict the model to be correctly specified.
In doing so, we also extend some results of \cite{white82} for independent,
identically distributed (i.i.d.) data to the case of dependent observations
and for classes of parametric distributions associated with dynamic models
with hidden Markov regimes. As such stochastic specifications are typically
highly parametric, it is important to understand the properties of
likelihood-based inferential procedures in situations where the true
probability structure of the data does not necessarily lie within the
parametric family of distributions specified by the model. We show that the ML
estimator in our setting converges to the true parameter value if the model is
correctly specified and to a pseudo-true parameter set if the model is
misspecified.\footnote{We refer to estimators for the various models discussed
throughout as ML estimators even though they may be obtained from a
pseudo-likelihood based on a misspecified model.} We also show that the sample
log-likelihood satisfies the LAN property, establish an asymptotic linear
representation for the ML estimator, obtain the asymptotic distribution of the
estimator, and present results relating to consistent estimation of its
asymptotic covariance matrix. These are the most general results available for
Markov regime-switching models with autoregressive dynamics and
covariate-dependent transition probabilities.

In related earlier work, \cite{Mevel04} examine consistency and asymptotic
normality of the ML estimator in misspecified hidden Markov models with a
finite state space, while \cite{douc12} consider consistency under general
state spaces. \cite{bickel96}, \cite{bickel98}, \cite{Jensen99},
\cite{douc01}, and \cite{douc11} investigate consistency and/or asymptotic
normality in correctly specified hidden Markov models with regime sequences
defined on either a finite or a general state space. \cite{Francq98} and
\cite{Krishna98} consider consistency in correctly specified autoregressive
models with Markov regimes defined on a finite state space. \cite{douc04} and
\cite{Kasahara2019} investigate consistency and asymptotic normality in a
similar autoregressive setup, but allow the regime sequence to take values in
a space that is not necessarily countable. In all of these papers, the hidden
regime sequence is assumed to be a temporally homogeneous Markov chain.

In the sequel, we follow \cite{bickel98} and \cite{douc04} fairly closely in
terms of the technical tools and the arguments used to establish our results,
but our setup is more general in certain respects. Like \cite{bickel98}, we
consider models with a finite hidden state space, but allow for autoregressive
dynamics in the observation sequence, covariate-dependence in the transition
probabilities of the regime sequence, and potential model misspecification. In
\cite{douc04}, the hidden Markov chain is allowed to take values in a compact
topological space, but is restricted to be temporally homogeneous and the
model is assumed to be correctly specified. The cornerstone of the methods
used in these papers for establishing the asymptotic properties of the ML
estimator are mixing-type results for the unobservable regime sequence
conditional on the observation sequence (see also \cite{bickel96}). This is
also true for our approach, although the aforementioned results cannot be
invoked directly because they are established under the assumption of temporal
homogeneity of the hidden Markov chain. We, therefore, extend these results to
allow for more general Markov regime sequences; in particular, we establish
mixing-type results for the unobservable regime sequence given the observed
data, allowing for a particular form of covariate-dependence in the transition
kernels. This last result is, to our knowledge, novel and may be of interest
in its own right.

The remainder of the paper is organized as follows. Section~\ref{sec:models}
defines the class of models under consideration, gives sufficient conditions
for stationarity and ergodicity of the observation process, and describes the
estimation problem. Section~\ref{sec:consistent} investigates consistency of
the ML estimator in a general setting. Section~\ref{sec:anormal} contains
results on the LAN property of the model and the asymptotic normality of the
ML estimator. Section~\ref{sec:simulation} presents simulation results on the
finite-sample properties of estimators based on well-specified and
misspecified likelihoods. Section~\ref{sec:illustration} presents an
illustration using real-world data. Proofs of the main results are gathered in
an Appendix.


The following notational conventions are used throughout the paper. For an
infinite sequence $(V_{j})_{j}$, $V_{a}^{b}=(V_{a},\ldots,V_{b})$ for any
$a\leq b$; $\mathcal{P}(\mathbb{V})$ denotes the set of Borel probability
measures on a Polish space $\mathbb{V}$; for a probability measure $P$,
$E_{P}(\cdot)$ denotes expectation with respect to $P$, $o_{P}(\cdot)$ and
$O_{P}(\cdot)$ indicate order in probability under $P$, $\Rightarrow_{P}$
signifies weak convergence under $P$, and $L^{r}(P)$, $1\leq r<\infty$,
denotes the class of measurable functions integrable to order $r$ with respect
to $P$; $\nabla_{\vartheta}$ and $\nabla_{\vartheta}^{2}$ are the gradient and
Hessian operators, respectively, with respect to $\vartheta$; $\left\Vert
\cdot\right\Vert $ denotes the Euclidean norm of a vector or matrix;
$1\{\cdot\}$ denotes the indicator function; $\mathbb{N}$ denotes the set of
positive integers. Unless stated otherwise, limits are taken as the sample
size, $T$, diverges to infinity.

\section{Model and Estimation}

\label{sec:models}

\subsection{Statistical Model}

\label{sec:model}

Let $(X_{t},S_{t})_{t=0}^{\infty}$ be a discrete-time stochastic process such
that, for each $t\in\mathbb{N}$, $S_{t}\in\mathbb{S}\equiv\{s_{1}
,\ldots,s_{|\mathbb{S}|}\}\subset\mathbb{R}$ is the unobservable state and
$X_{t}\in\mathbb{X}\subseteq\mathbb{R}^{h}$, for some $h\in\mathbb{N}$, is the
observable state. Moreover, for each $t\in\mathbb{N}$,~the conditional
distribution of $X_{t}$, given $X_{0}^{t-1}$ and $S_{0}^{t}$, depends only on
$X_{t-1}$ and $S_{t}$, and the conditional distribution of $S_{t}$, given
$X_{0}^{t-1}$ and $S_{0}^{t-1}$, depends only on $X_{t-1}$ and $S_{t-1}$, so
that
\begin{align*}
&  X_{t}\mid(X_{0}^{t-1},S_{0}^{t})\sim P_{\ast}(X_{t-1},S_{t},\cdot),\\
&  S_{t}\mid(X_{0}^{t-1},S_{0}^{t-1})\sim Q_{\ast}(X_{t-1},S_{t-1},\cdot),
\end{align*}
with $(x,s)\mapsto P_{\ast}(x,s,\cdot)\in\mathcal{P}(\mathbb{X})$ and
$(x,s)\mapsto Q_{\ast}(x,s,\cdot)\in\mathcal{P}(\mathbb{S})$ denoting the true
transition probabilities. It is further assumed that, for each $(x,s)\in
\mathbb{X}\times\mathbb{S}$, $P_{\ast}(x,s,\cdot)$ admits a density $p_{\ast
}(x,s,\cdot)$ with respect to some $\sigma$-finite measure on $\mathbb{X}$.
Our framework imposes no additional restrictions on this measure; for
instance, it can be the Lebesgue measure (i.e., allow for continuous $X_{t}$)
or a counting measure (i.e., allow for discrete $X_{t}$).

The researcher's model is given by a family of transition probabilities
$(x,s)\mapsto P_{\theta}(x,s,\cdot)\in\mathcal{P}(\mathbb{X})$ and
$(x,s)\mapsto Q_{\theta}(x,s,\cdot)\in\mathcal{P}(\mathbb{S})$ indexed by an
(unknown) parameter $\theta\in\Theta\subseteq\mathbb{R}^{q}$, for some
$q\in\mathbb{N}$, such that, for each $\theta\in\Theta$,
\[
X_{t}\mid(X_{0}^{t-1},S_{0}^{t})\sim P_{\theta}(X_{t-1},S_{t},\cdot),
\]
\[
S_{t}\mid(X_{0}^{t-1},S_{0}^{t-1})\sim Q_{\theta}(X_{t-1},S_{t-1},\cdot),
\]
and, for each $(x,s)\in\mathbb{X}\times\mathbb{S}$, $P_{\theta}(x,s,\cdot)$
admits a density $p_{\theta}(x,s,\cdot)$ with respect to the same measure used
to define $p_{\ast}(x,s,\cdot)$.

A few remarks about this setup are worth making. First, and perhaps most
importantly, the unobservable states (regimes) are a Markov chain whose
transition kernel can depend on the lagged value of the observable state. This
is the main departure from prior literature which, with the exception of
\cite{aill13}, has focused on the case where the conditional distribution of
$S_{t}$, given $X_{0}^{t-1}$ and $S_{0}^{t-1}$, depends only on $S_{t-1}$.
Second, the model $\{(P_{\theta},Q_{\theta})\colon\theta\in\Theta\}$ is
allowed to be misspecified in the sense that $(P_{\ast},Q_{\ast}
)\notin\{(P_{\theta},Q_{\theta})\colon\theta\in\Theta\}$. This setup
encompasses a rich family of models that arise in econometric and statistical
applications, some examples of which are given below. Note that, although the
family $\{Q_{\theta}\colon\theta\in\Theta\}$ may or may not contain $Q_{\ast}
$, it is defined on the same finite state space $\mathbb{S}$ as $Q_{\ast}$, so
misspecification of the number of unobservable regimes is ruled
out.\footnote{Such misspecification would introduce additional complexities
(e.g., locally non-quadratic likelihood surfaces, non-identifiable parameters)
which are beyond the scope of this paper.} Finally, even though the
conditional distribution of $S_{t}$, given $X_{0}^{t-1}$ and $S_{0}^{t-1}$, is
assumed to depend only on $X_{t-1}$ and $S_{t-1}$, it is straightforward to
extend all our results to situations where this distribution is dependent on
additional higher-order lags of $X_{t}$ and/or $S_{t}$.

\begin{example}
[Hidden Markov Model with Covariate-Dependent Transition Probabilities]
\label{exa:HMM} Let $x=(y,z)\in\mathbb{X}=\mathbb{R}^{2}$ and $\mathbb{S}
=\{0,1\}$. Let $P_{\ast}$ be determined by the equations
\begin{align*}
Y_{t}  &  =\mu^{\ast}(S_{t})+\sigma^{\ast}(S_{t})U_{1,t},\\
Z_{t}  &  =\mu_{2}^{\ast}+\psi^{\ast}Z_{t-1}+\sigma_{2}^{\ast}U_{2,t},
\end{align*}
where $(U_{1,t},U_{2,t})_{t}$ are i.i.d. (independent of $(S_{t})_{t}$) with
zero mean and covariance matrix indexed by a parameter $\rho^{\ast}$, e.g.,
$\left[
\begin{array}
[c]{cc}
1 & \rho^{\ast}\\
\rho^{\ast} & 1
\end{array}
\right]  $. The transition probabilities of $(S_{t})_{t}$ are allowed to
depend on $Z_{t-1}$; for instance, $(s,z)\mapsto Q_{\ast}(z,s,s)\equiv
\Pr(S_{t}=s\mid Z_{t-1}=z,S_{t-1}=s)=[1+\exp(-\alpha_{s}^{\ast}-\beta
_{s}^{\ast}z)]^{-1}$ for $s\in\mathbb{S}$. The homogeneous specification with
$\beta_{0}^{\ast}=\beta_{1}^{\ast}=0$ has been used to model regime shifts in
a variety of economic and financial time series, including output growth
(\cite{AlbertChib93}, \cite{rivas15}), foreign exchange rates
(\cite{engelhamilton90}, \cite{bollen0}), and equity returns (\cite{ryden98},
\cite{ang02}). The homogeneity restriction is relaxed in \cite{dieb94},
\cite{Engel96}, and \cite{ang02}, among others, to allow the transition
probabilities to depend on $Z_{t-1}$. The results obtained here establish the
asymptotic properties of the ML estimator in this class of models. $\triangle$
\end{example}

\begin{example}
[Markov-Switching Autoregressive Model with Covariate-Dependent Transition
Probabilities]\label{exa:MSR} A useful generalization of the previous example
is one where the outcome equation is extended to
\[
Y_{t}=\mu^{\ast}(S_{t})+\phi^{\ast}Y_{t-1}+\sigma^{\ast}(S_{t})U_{1,t}.
\]
Variations of the model with $\beta_{0}^{\ast}=\beta_{1}^{\ast}=0$ have found
widespread application in economics (e.g., \cite{Hansen92},
\cite{McCulloghTsay94}, \cite{Murcia95}, \cite{angbekwei08}) and beyond.
Generalizations of the model without the restriction of time-invariant
transition probabilities are also very popular and have been used, for
example, in the modeling of output growth (\cite{rivas15}), interest rates
(\cite{angbek02j}), consumption growth (\cite{Whitelaw00}), and bond spreads
(\cite{pses21}). The results obtained here establish the asymptotic properties
of the ML estimator in this class of models. $\triangle$
\end{example}

\begin{example}
[Mixture Autoregressive Model]\label{exa:MAM}Let $x\in\mathbb{X}=\mathbb{R}$,
$\mathbb{S}=\{0,1\}$, and for each $t\in\mathbb{N}$, $\Pr(S_{t}=0\mid
X_{t-1})=G_{\ast}(X_{t-1})$ for some $x\mapsto G_{\ast}(x)\in\lbrack0,1]$ and
$X_{t}\sim P_{\ast}(X_{t-1},S_{t},\cdot)$. This specification implies a
conditional density for $X_{t}$, given $X_{t-1}$, that is a mixture of the
type
\[
x\mapsto p_{\ast}(x\mid x_{t-1})=G_{\ast}(x_{t-1})p_{\ast}(x_{t-1}
,0,x)+(1-G_{\ast}(x_{t-1}))p_{\ast}(x_{t-1},1,x).
\]
The covariate-dependence of the transition functions is reflected by the fact
that $G_{\ast}$ depends on $X_{t-1}$. Models which belong to the general class
of mixture autoregressive models (e.g., \cite{dss07}, \cite{Tadjuidje09},
\cite{dsps11}, \cite{saik15}) are covered by this framework. $\triangle$
\end{example}

The examples above, as well as those that follow, illustrate that in many
areas of application the stochastic process $(X_{t},S_{t})_{t=0}^{\infty}$ is
typically highly complex and it is natural/desirable to allow for feedback
from past realizations of the observable process $(X_{t})_{t=0}^{\infty}$ to
the law of the unobservable regime sequence $(S_{t})_{t=0}^{\infty}$; a
tractable way for modeling such feedback is to allow the transition kernel
$Q_{\theta}$ to depend on $X_{t-1}$. This feature adds an additional level of
complexity and with it sources of potential misspecification. For instance, a
common assumption in most applications is that the transition probabilities of
hidden regimes are time-invariant, an assumption which may result in
misspecification of the transitions functions. In applications that allow for
covariate-dependent transition functions, a somewhat more subtle and often
overlooked source of misspecification, associated with endogeneity of the
transition-driving covariates, may come into play. Specifically, in the
context of models such as those in Examples~\ref{exa:HMM} and \ref{exa:MSR},
contemporaneous correlation between $U_{1,t}$ and $U_{2,t}$ is typically
ignored and inference is based on the likelihood implied by the outcome
equation alone; we consider this case in more detail in Sections
\ref{sec:simulation2} and \ref{sec:illustration}.





Before discussing estimation of $\theta$, we give a result regarding the
mixing and ergodicity properties of $(X_{t})_{t=0}^{\infty}$. To do so, let
$\bar{P}_{\ast}^{\kappa}$ denote the true distribution over $(X_{t}
)_{t=0}^{\infty}$ when the distribution of $(X_{0},S_{0})$ is $\kappa$. Under
the following assumptions, Lemma~\ref{lem:sta-ergo} below ensures that there
exists a Borel probability measure on $\mathbb{X}\times\mathbb{S}$, denoted
henceforth by $\nu$, for which $(X_{t})_{t=0}^{\infty}$ is stationary and ergodic.

\begin{assumption}
\label{ass:BDD_Q} There exists a continuous function $\underline{q}
:\mathbb{X}\rightarrow\mathbb{R}_{+}\setminus\{0\}$ such that, for all
$Q\in\{Q_{\theta}\colon\theta\in\Theta\}\cup Q_{\ast}$, $Q(x,s,s^{\prime}
)\geq\underline{q}(x)$ for all $(s^{\prime},s,x)\in\mathbb{S}^{2}
\times\mathbb{X}$.
\end{assumption}

\begin{assumption}
\label{ass:fx-ergo} There exist constants $\lambda^{\prime}\in(0,1)$,
$\gamma\in(0,1)$, $b^{\prime}>0$ and $R>2b^{\prime}/(1-\gamma)$, a lower
semi-continuous function $\mathcal{U}:\mathbb{X}\rightarrow\lbrack1,\infty)$,
and a measure $\varpi\in\mathcal{P}(\mathbb{X})$ such that, for all
$s\in\mathbb{S}$: (i)$~\int_{\mathbb{X}}\mathcal{U}(x^{\prime})P_{\ast
}(x,s,dx^{\prime})\leq\gamma\mathcal{U}(x)+b^{\prime}1\{x\in A\}$, with
$A\equiv\{x\in\mathbb{X}\colon\mathcal{U}(x)\leq R\}$; (ii)$~A$ is bounded and
$\varpi(A)>0$; (iii)$~\inf_{x\in A}P_{\ast}(x,s,C)\geq\lambda^{\prime}
\varpi(C)$ for any Borel set $C\subseteq\mathbb{X}$.

\end{assumption}

The following lemma establishes stationarity, ergodicity, and $\beta$-mixing
of $(X_{t})_{t=0}^{\infty}$.

\begin{lemma}
\label{lem:sta-ergo} Suppose Assumptions \ref{ass:BDD_Q} and \ref{ass:fx-ergo}
hold. Then, there exists a $\nu\in\mathcal{P}(\mathbb{X}\times\mathbb{S})$
such that, under $\bar{P}_{\ast}^{\nu}$, $(X_{t})_{t=0}^{\infty}$ is
stationary, ergodic, and $\beta$-mixing with mixing coefficients $\beta
_{n}=O(\gamma^{n})$, $n\in\mathbb{N}$.
\end{lemma}

\begin{proof}
See Supplemental Material \ref{SM:Ergo}.
\end{proof}


The result follows in a standard manner by using Assumptions \ref{ass:BDD_Q}
and \ref{ass:fx-ergo} to establish that the implied transition kernel of the
joint process $(X_{t},S_{t})_{t=0}^{\infty}$ has a unique invariant
distribution and also that it is Harris recurrent and aperiodic. This fact, in
turn, is used to show that $(X_{t})_{t=0}^{\infty}$ is stationary, ergodic,
and $\beta$-mixing at a geometric rate.

\medspace


\begin{remark}
[Discussion of Assumptions \ref{ass:BDD_Q} and \ref{ass:fx-ergo}
]\label{rem:BDD-Q} Assumption \ref{ass:BDD_Q} is an extension of a common
assumption in the literature (cf. \cite{douc04}, \cite{aill13}) to the case
where the transition kernel of $(S_{t})_{t=0}^{\infty}$ depends on $X_{t-1}$.
Allowing the lower bound $\underline{q}$ to depend on $x$ is especially
relevant when the support of $X_{t}$ is unbounded because, while
$\underline{q}(x)>0$, it is allowed to converge to zero as $\left\Vert
x\right\Vert \rightarrow\infty$. Although this assumption is not innocuous, we
view it as mild because it accommodates the typical specifications used in the
literature, where $Q$ is parameterized by a standard Gaussian or logistic
cumulative distribution function and a single index $x^{\intercal}\beta$, with
$\beta$ restricted to take values in a bounded subset of a finite-dimensional
Euclidean space.

Assumption \ref{ass:fx-ergo}(iii) is an analogous condition for the transition
kernel $P_{\ast}$. By inspection of the proof of Lemma~\ref{lem:sta-ergo}, it
is easy to see that it suffices to obtain a minorization condition for the
\textquotedblleft joint\textquotedblright\ kernel, i.e., $\inf_{x\in A}
P_{\ast}(x,s^{\prime},C)Q(x,s,s^{\prime})\geq\lambda\widetilde{\varpi
}(C,s^{\prime})$ for any Borel set $C\subseteq\mathbb{X}$ and for some
$\widetilde{\varpi}\in\mathcal{P}(\mathbb{X}\times\mathbb{S})$ and $\lambda
\in(0,1)$. Thus, Assumptions \ref{ass:BDD_Q}(i) and \ref{ass:fx-ergo}(i) could
be relaxed; e.g., the former could be relaxed to $Q(x,s,s^{\prime}
)\geq\underline{q}(x)\varrho(s^{\prime})$, where $\varrho\in\mathcal{P}
(\mathbb{S})$, or the latter could be relaxed to $\inf_{x\in A}P_{\ast
}(x,s^{\prime},C)\geq\lambda^{\prime}\widetilde{\varpi}(C,s^{\prime})$, where
$\widetilde{\varpi}\in\mathcal{P}(\mathbb{X}\times\mathbb{S})$.

Assumption \ref{ass:fx-ergo}(i),(ii) is a so-called Foster--Lyapunov drift
condition; see \cite{meyn2012markov} and references therein for a discussion
of the assumption. $\triangle$
\end{remark}

\medspace


In view of Lemma~\ref{lem:sta-ergo}, under $\nu$, the process $(X_{t}
)_{t=0}^{\infty}$ can be extended to a two-sided sequence $(X_{t})_{t=-\infty
}^{\infty}$. With a slight abuse of notation, we still use $\bar{P}_{\ast
}^{\nu}$ to denote the true probability distribution over $(X_{t}
,S_{t})_{t=-\infty}^{\infty}$; $\bar{P}_{\theta}^{\nu}$ is defined analogously
for the model $(Q_{\theta},p_{\theta},\nu)$.\footnote{Throughout the text, we
use $\bar{P}_{\theta}^{\nu}$ to denote any marginal or conditional
probabilities associated with $\bar{P}_{\theta}^{\nu}$; the same holds for
$\bar{P}_{\ast}^{\nu}$.}

\subsection{Parameter Estimation}

\label{sec:estimation}

For any $T\in\mathbb{N}$, let $\ell_{T}^{\nu}:\mathbb{X}^{T+1}\times
\Theta\rightarrow\mathbb{R}$ be the sample criterion function given by
\begin{equation}
\ell_{T}^{\nu}(X_{0}^{T},\theta)=T^{-1}\sum_{t=1}^{T}\log p_{t}^{\nu}
(X_{t}\mid X_{0}^{t-1},\theta), \label{logl}
\end{equation}
where $p_{t}^{\nu}(X_{t}\mid X_{0}^{t-1},\theta)$ denotes the conditional
density of $X_{t}$ given $X_{0}^{t-1}$ for any $\theta\in\Theta$; the latter
is defined recursively as follows: for any $t\geq1,$
\[
p_{t}^{\nu}(X_{t}\mid X_{0}^{t-1},\theta)=\sum_{s^{\prime}\in\mathbb{S}}
\sum_{s\in\mathbb{S}}p_{\theta}(X_{t-1},s^{\prime},X_{t})Q_{\theta}
(X_{t-1},s,s^{\prime})\delta_{t}^{\theta,\nu}(s),
\]
and $s\mapsto\delta_{t}^{\theta,\nu}(s)\equiv\bar{P}_{\theta}^{\nu}
(S_{t-1}=s\mid X_{0}^{t-1})$. For each $t\geq2$ and any $s\in\mathbb{S}$,
$s\mapsto\delta_{t}^{\theta,\nu}(s)$ satisfies the recursion
\[
\delta_{t}^{\theta,\nu}(s)=\sum_{\tilde{s}\in\mathbb{S}}\frac{Q_{\theta
}(X_{t-1},\tilde{s},s)p_{\theta}(X_{t-2},\tilde{s},X_{t-1})\delta
_{t-1}^{\theta,\nu}(\tilde{s})}{\sum_{s^{\prime}\in\mathbb{S}}p_{\theta
}(X_{t-2},s^{\prime},X_{t-1})\delta_{t-1}^{\theta,\nu}(s^{\prime})},
\]
with $s\mapsto\delta_{1}^{\theta,\nu}(s)=\sum_{\tilde{s}\in\mathbb{S}
}Q_{\theta}(X_{0},\tilde{s},s)\nu(\tilde{s}|X_{0})$, where $\nu(\cdot|\cdot)$
is the conditional density corresponding to $\nu$.

For a given initial distribution $\kappa\in\mathcal{P}(\mathbb{X}
\times\mathbb{S})$ over $(X_{0},S_{0})$, we define our estimator as
$\hat{\theta}_{\kappa,T}$, where
\begin{equation}
\ell_{T}^{\kappa}(X_{0}^{T},\hat{\theta}_{\kappa,T})\geq\sup_{\theta\in\Theta
}\ell_{T}^{\kappa}(X_{0}^{T},\theta)-\eta_{T}, \label{mle}
\end{equation}
for some $\eta_{T}\geq0$ and $\eta_{T}=o(1)$.



\section{Consistency}

\label{sec:consistent}

The main result of this section establishes convergence of the estimator
$\hat{\theta}_{\nu,T}$ to the set of points in $\Theta$ that are closest to
the true model under the Kullback--Leibler information criterion.

Let $H^{\ast}:\Theta\rightarrow\mathbb{R}_{+}\cup\{\infty\}$ be the
Kullback--Leibler information criterion $\theta\mapsto H^{\ast}(\theta)$,
which is given by
\[
H^{\ast}(\theta)=E_{\bar{P}_{\ast}^{\nu}}\left[  \log\frac{p_{\ast}^{\nu
}(X_{0}\mid X_{-\infty}^{-1})}{p^{\nu}(X_{0}\mid X_{-\infty}^{-1},\theta
)}\right]  ,
\]
where, for any $\theta\in\Theta$, $p^{\nu}(X_{0}\mid X_{-\infty}^{-1},\theta)$
denotes the conditional density of $X_{0}$ given $X_{-\infty}^{-1}$ induced by
$(P_{\theta},Q_{\theta},\nu)$, and $p_{\ast}^{\nu}(X_{0}\mid X_{-\infty}
^{-1})$ is its counterpart induced by the true transition kernels $(P_{\ast
},Q_{\ast},\nu)$; we refer the reader to the Supplemental Material
\ref{SM:PDFs} for details about the construction of these objects and their properties.

Our results allow for misspecified models, and thus $p_{\ast}^{\nu}
\notin\{p^{\nu}(\cdot\mid\cdot,\theta):\theta\in\Theta\}$. Hence, as in
\cite{white82}, the relevant limiting set for our estimator is
\[
\Theta_{\ast}=\underset{\theta\in\Theta}{\arg\min}H^{\ast}(\theta),
\]
which is the \emph{pseudo-true parameter (set)} that minimizes the
Kullback--Leibler information criterion. Under Assumption \ref{ass:exist}
below, $\Theta_{\ast}$ is non-empty and compact by the Weierstrass Theorem.

\begin{assumption}
\label{ass:exist} (i) $\Theta$ is compact; (ii) $H^{\ast}$ exists and is lower semi-continuous.
\end{assumption}

The following additional assumptions are used to establish the main result. To
state these, for any $\delta>0$ and any $\dot{\theta}\in\Theta$, let
$B(\delta,\dot{\theta})\equiv\{\theta\in\Theta\colon||\dot{\theta}
-\theta||<\delta\}$.

\begin{assumption}
\label{ass:bdd} (i) For any $\epsilon>0$, there exists some $\delta>0$ such
that
\[
\max_{\dot{\theta}\in\Theta}E_{\bar{P}_{\ast}^{\nu}}\left[  \sup_{\theta\in
B(\delta,\dot{\theta})}\frac{p^{\nu}(X_{0}\mid X_{-\infty}^{-1},\theta
)}{p^{\nu}(X_{0}\mid X_{-\infty}^{-1},\dot{\theta})}\right]  \leq1+\epsilon;
\]
(ii) there exists a function $(x,x^{\prime})\mapsto C(x,x^{\prime}
)\in\mathbb{R}_{+}$ such that $\sup_{\theta\in\Theta}\frac{\max_{s\in
\mathbb{S}}p_{\theta}(X,s,X^{\prime})}{\min_{s\in\mathbb{S}}p_{\theta
}(X,s,X^{\prime})}\leq C(X,X^{\prime})$ and $\frac{\max_{s\in\mathbb{S}
}p_{\ast}(X,s,X^{\prime})}{\min_{s\in\mathbb{S}}p_{\ast}(X,s,X^{\prime})}\leq
C(X,X^{\prime})$ a.s.-$\bar{P}_{\ast}^{\nu}$.
\end{assumption}

\begin{assumption}
\label{ass:qbar-sum} $T^{-1}\sum_{t=1}^{T}\max\{1,C(X_{t-1},X_{t}
)\}\prod_{i=0}^{t-1}(1-\underline{q}(X_{i}))=o_{\bar{P}_{\nu}^{\ast}}(1).$
\end{assumption}

\begin{remark}
[Discussion of Assumptions \ref{ass:exist}, \ref{ass:bdd} and
\ref{ass:qbar-sum}]Assumption~\ref{ass:exist}(i) is standard.
Assumption~\ref{ass:exist}(ii) is high-level but can be obtained from
lower-level conditions (e.g., by following the reasoning in Proposition~11 of
\cite{douc12}).

Assumption \ref{ass:bdd}(i) is a high-level condition used for establishing
uniform law of large numbers results (see Lemma~\ref{lem:lik-conv} in Section
\ref{app:consistent}). Assumption \ref{ass:bdd}(ii) is akin to Assumption~A4
in \cite{bickel98}; it essentially restricts the support of $p_{\theta}$ and
$p_{\ast}$ for different values of the state variable.\footnote{The second
part of Assumption \ref{ass:bdd}(ii) is used to show that $p_{\ast}^{\nu
}(\cdot|X_{-\infty}^{-1})$ integrates to one.}

Assumption \ref{ass:qbar-sum} essentially requires that $\underline{q}$ is not
\textquotedblleft too close\textquotedblright\ to zero on average and that
$C(X_{t-1},X_{t})$ is finite a.s.-$\bar{P}_{\nu}^{\ast}$. For instance, if
$\underline{q}(x)\geq c$ for some $c>0$, then Assumption \ref{ass:qbar-sum} is
automatically satisfied provided $E_{\bar{P}_{\nu}^{\ast}}[C(X_{0}
,X_{1})]<\infty$. Moreover, by exploiting the fact that $(X_{t})_{t}$ is
$\beta$-mixing (see Lemma \ref{lem:sta-ergo}), Lemma \ref{lem:suff-A5andA8} in
the Supplemental Material \ref{SM:SuffAssumptions} provides sufficient
conditions of the form $E_{\bar{P}_{\nu}^{\ast}}\left[  \underline{q}
(X_{1})\right]  >0$ and $E_{\bar{P}_{\nu}^{\ast}}\left[  C(X_{1},X_{0}
)^{l}\right]  <\infty$ for some $l>1$.\footnote{We thank a referee and the
editor for suggestions on how to weaken Assumption \ref{ass:bdd}, which lead
to these sufficient conditions for Assumption \ref{ass:qbar-sum}.}
$\triangle$
\end{remark}

We now establish consistency of the estimator defined by (\ref{mle}).

\begin{theorem}
\label{thm:consistent} Suppose Assumptions \ref{ass:BDD_Q}--\ref{ass:bdd}
hold. Then, $d_{\Theta}(\hat{\theta}_{\nu,T},\Theta_{\ast})=o_{\bar{P}_{\ast
}^{\nu}}(1).$\footnote{For any set $A\subseteq\Theta$, $d_{\Theta}
(\theta,A)\equiv\inf_{\dot{\theta}\in A}||\theta-\dot{\theta}||$.}
\end{theorem}

\begin{proof}
See Appendix \ref{app:consistent}.
\end{proof}


This result is analogous to Theorem~2 in \cite{douc12} but for a somewhat
different setup; specifically, we allow for autoregressive dynamics and
covariate-dependent transition probabilities, but restrict $\mathbb{S}$ to be
finite.\footnote{Finiteness of $\mathbb{S}$ simplifies the proofs as we do not
have to be concerned with uniformity issues of certain quantities as functions
of the hidden state. By combining techniques available in the literature
(e.g., \cite{douc12}) -- which essentially amount to imposing requirements
like compactness of $\mathbb{S}$ and continuity -- with ours, we conjecture
that our results can be extended to the case where $\mathbb{S}$ is a compact
set.}

Clearly, if the model is correctly specified and point identified, i.e., there
exists a $\theta\in\Theta$ such that $(P_{\ast},Q_{\ast})=(P_{\theta
},Q_{\theta})$, then $\Theta_{\ast}=\{\theta\}$ and our estimator converges in
$\bar{P}_{\ast}^{\nu}$-probability to this point. If, however, the model is
misspecified, our estimator converges to
the set of parameters that is closest to the true set, when closeness is
measured by means of the Kullback--Leibler information criterion (cf.
\cite{white82}, \cite{douc12}).

To prove Theorem \ref{thm:consistent}, we first show that $T^{-1}\sum
_{t=1}^{T}\log p_{t}^{\nu}(X_{t}\mid X_{0}^{t-1},\theta)$ is well-approximated
by $T^{-1}\sum_{t=1}^{T}\log p^{\nu}(X_{t}\mid X_{-\infty}^{t-1},\theta)$ (see
Lemma~\ref{lem:lik-approx} in Appendix~\ref{app:consistent}). Second, relying
on ergodicity (Lemma~\ref{lem:sta-ergo}) and Assumption \ref{ass:bdd}, we
establish a \emph{uniform} law of large numbers for the latter quantity (see
Lemma~\ref{lem:lik-conv} in Appendix~\ref{app:consistent}). The proof of
consistency then follows the standard Wald approach.

The approximation result in the first step relies on \textquotedblleft
mixing\textquotedblright\ results for the process $(S_{t})_{t=-\infty}
^{\infty}$, given $(X_{t})_{t=-\infty}^{\infty}$. The following theorem, which
might be of independent interest, establishes such a \textquotedblleft
mixing\textquotedblright\ result in our setting.

\begin{theorem}
\label{thm:Q-ergo} Take any $(j,m)\in\mathbb{N}^{2}$. Suppose that, for any
$\theta\in\Theta$, there exist mappings $x\mapsto\varrho(x,\cdot
)\in\mathcal{P}(\mathbb{S})$ and $\underline{q}:\mathbb{X}\rightarrow
\mathbb{R}_{+}$ such that, for all $(s,s^{\prime})\in\mathbb{S}^{2}$,
\begin{equation}
Q_{\theta}(X,s,s^{\prime})\geq\underline{q}(X)\varrho(X,s^{\prime}
)\quad~a.s.\text{\textrm{-}}\bar{P}_{\ast}^{\nu}. \label{eqn:Q-ergo}
\end{equation}
Then,\footnote{For any $P,Q\in$ $\mathcal{P}(\mathbb{S})$, $||P-Q||_{1}
\equiv\sum_{s\in\mathbb{S}}|P(s)-Q(s)|$.}
\[
\max_{(b,c)\in\mathbb{S}^{2}}\left\Vert \bar{P}_{\theta}^{\nu}(S_{j+1}
=\cdot|S_{-m}=b,X_{-m}^{j})-\bar{P}_{\theta}^{\nu}(S_{j+1}=\cdot
|S_{-m}=c,X_{-m}^{j})\right\Vert _{1}\leq\prod_{n=-m}^{j}(1-\underline{q}
(X_{n})).
\]

\end{theorem}

\begin{proof}
See Appendix \ref{app:mixing}.
\end{proof}


This result is analogous to results in \cite{douc04} (e.g., their Lemma 1 and
Corollary 1) but for a more general transition function $Q_{\theta}$ (albeit
under the requirement that $\mathbb{S}$ is finite). Specifically, we allow
$Q_{\theta}$ to depend on $X$, and do not restrict the lower bound in
(\ref{eqn:Q-ergo}) to be uniform in $s^{\prime}$; these features are, to our
knowledge, novel.

The proof of Theorem~\ref{thm:Q-ergo} relies on bounding the Dobrushin
coefficient of the transition kernel $\bar{P}_{\theta}^{\nu}(S_{l+1}
=\cdot|S_{l}=\cdot,X_{-m}^{j})$ by $1-\underline{q}(X_{l})$, for each
$l\in\{-m,...,j\}$. The case where $\varrho(x,\cdot)$, in condition
(\ref{eqn:Q-ergo}), is uniformly bounded from below has been studied in the
literature, and such a bound can be obtained by elementary calculations. For
the general case, however, the standard technique cannot be used, and we
develop a novel, to our knowledge, approach based on coupling techniques; see
Appendix~\ref{app:mixing} for details.



\begin{remark}
Remark \ref{rem:BDD-Q} and the fact that Theorem \ref{thm:Q-ergo} is
established under condition (\ref{eqn:Q-ergo}), imply that, in regards to
consistency, Assumption \ref{ass:BDD_Q} could be replaced by the weaker
condition (\ref{eqn:Q-ergo}). This remark, however, does not extend to the LAN
results of Section~\ref{sec:anormal}, since we do not know whether Assumption
\ref{ass:BDD_Q} could be weakened to condition (\ref{eqn:Q-ergo}) in this
case. $\triangle$
\end{remark}



We discuss next a canonical example that encompasses many commonly used
specifications, such as those in Examples \ref{exa:HMM} and \ref{exa:MSR}. The
purpose of this example is to verify our assumptions under primitive and
low-level conditions.

\begin{example}
[Canonical Example]\label{exa:Canon} We verify the regularity conditions for a
model with $\mathbb{S}=\{0,1\}$ and
\begin{align*}
&  X_{t}=\mu(S_{t})+\Phi^{\intercal}X_{t-1}+\Sigma^{1/2}(S_{t})\varepsilon
_{t},\\
&  S_{t}\sim Q_{\bar{\vartheta}}(X_{t-1},S_{t-1},\cdot),
\end{align*}
where $(\varepsilon_{t})_{t}\sim\mathrm{i.i.d.}~\mathcal{N}(0,I)$, $I$ being
the identity matrix, and $x\mapsto Q_{\bar{\vartheta}}(x,s,s)\equiv\Pr
(S_{t}=s\mid X_{t-1}=x,S_{t-1}=s)=\Psi(x^{\intercal}\bar{\vartheta}_{s})$ for
$s\in\mathbb{S}$, where $\Psi$ is a full-support, continuous cumulative
distribution function (e.g., logistic or normal). It is assumed that the
parameter set $\Theta$ is compact and that any $\theta\equiv((\mu
(s),\Sigma(s))_{s\in\mathbb{S}},\Phi,\bar{\vartheta})$ in it is such that
$\Sigma(\cdot)$ and $\Phi\Phi^{\intercal}$ have eigenvalues uniformly bounded
away from zero and infinity. For simplicity, we assume that the true model is
indexed by $((\mu_{\ast}(s),\Sigma_{\ast}(s))_{s\in\mathbb{S}},\Phi_{\ast
},\bar{\vartheta}_{\ast})$, for which analogous conditions hold.



Let $q(x)\equiv\inf_{b}\Psi(x^{\intercal}b)$ and note that $q(x)>0$ for each
finite $x$ as $\Psi$ has full support.
Thus, Assumption \ref{ass:BDD_Q} holds and $E_{\bar{P}_{\nu}^{\ast}}[q(X)]>0$.
The conditions of Assumption \ref{ass:fx-ergo} follow by the results in
\cite{douc2004practical}; see Lemma \ref{lem:fx-ergo-holds} in the
Supplemental Material \ref{SM:examples} for the formal argument. In Lemma
\ref{lem:exa-CBound} in the Supplemental Material \ref{SM:examples}, it is
shown that
\[
C^{-1}\underline{p}(x,y)\leq f_{\mathcal{N}}(\{y-\Phi^{\intercal}
x-\mu(s)\}\Sigma^{-1/2}(s))\leq C\overline{p}(x,y),
\]
for some $C\geq1$, where $f_{\mathcal{N}}$ is the $\mathcal{N}(0,I)$
probability density function,
\begin{align*}
\underline{p}(x,y)\equiv &  \exp\{-0.5(a_{1}\left\Vert y\right\Vert ^{2}
+a_{2}\left\Vert x\right\Vert ^{2}+2a_{3}\left\Vert x\right\Vert \left\Vert
y\right\Vert )-a_{4}\left\Vert y\right\Vert -a_{5}\left\Vert x\right\Vert
\},\\
\overline{p}(x,y)\equiv &  \exp\{-0.5(b_{1}\left\Vert y\right\Vert ^{2}
+b_{2}\left\Vert x\right\Vert ^{2}-2b_{3}\left\Vert x\right\Vert \left\Vert
y\right\Vert )+b_{4}\left\Vert y\right\Vert +b_{5}\left\Vert x\right\Vert \},
\end{align*}
and $a_{i},b_{i}$ $(i=1,\ldots,5)$ are positive constants that are functions
of uniform bounds of $\Sigma(\cdot)$, $\mu(\cdot)$ and $\Phi$ over $\Theta$;
they are defined in Lemmas \ref{lem:exa-CBound} and \ref{lem:exa-CBound2} in
the Supplemental Material \ref{SM:SuffAssumptions}. Therefore, by
Lemma~\ref{lem:prop-ergoPDF} in the Supplemental Material \ref{SM:PDFs}, it
follows that Assumptions \ref{ass:exist} and \ref{ass:bdd} hold with
$(x,x^{\prime})\mapsto C(x,x^{\prime})\equiv\overline{p}(x,x^{\prime
})/\underline{p}(x,x^{\prime})$, provided that $E_{\bar{P}_{\nu}^{\ast}}
[\exp\{-0.5l((b_{1}-a_{1})\left\Vert Y\right\Vert ^{2}+(b_{2}-a_{2})\left\Vert
X\right\Vert ^{2}-2(b_{3}+a_{3})\left\Vert X\right\Vert \left\Vert
Y\right\Vert )+l(b_{4}+a_{4})\left\Vert y\right\Vert +l(b_{5}+a_{5})\left\Vert
X\right\Vert \}]<\infty$ for some $l\geq1$. Lemma \ref{lem:exa-CBound2} in the
Supplemental Material \ref{SM:examples} shows that the latter condition holds
under some restrictions on the eigenvalues of $\Sigma(\cdot)$, $\Sigma_{\ast
}(\cdot)$, $\Phi\Phi^{\intercal}$, and $\Phi_{\ast}\Phi_{\ast}^{\intercal}$
(see the aforementioned lemma and its subsequent remark for the exact
formulation and more details). Finally, under these conditions, Lemma
\ref{lem:suff-A5andA8} in the Supplemental Material \ref{SM:SuffAssumptions}
shows that Assumption \ref{ass:qbar-sum} holds. $\triangle$
\end{example}

Lastly, we note that further characterization and interpretation of the
pseudo-true parameter set $\Theta_{\ast}$ is not straightforward in our
general setup, and we believe that progress in this direction should be made
on a case-by-case basis. While pursuing this program is outside the scope of
the present paper, in Section \ref{sec:simulation} we present numerical
results that attempt to shed some light on the characterization of this set.
In addition, Example~\ref{exa:HMM-miss} in the Supplemental Material
\ref{SM:examplemixture} provides analytical results on the characterization of
$\Theta_{\ast}$ for a subclass of the models in Example \ref{exa:MSR}, namely
hidden Markov models with covariate-dependent transition probabilities, when
the (misspecified) model is a simple mixture.









\section{Asymptotic Distribution Theory}

\label{sec:anormal}

In this section, we establish a LAN property (\cite[Ch. II]{IbragHas81},
\cite{leCam2012}) for our model and an asymptotic linear representation for
our estimator, from which asymptotic normality of the estimator is inferred.
For this, we require the following assumptions.



\begin{assumption}
\label{ass:Theta-int} (i) $\Theta_{\ast}=\{\theta_{\ast}\}\subset
\mathrm{int}(\Theta)$; (ii) $\theta\mapsto p_{\theta}(X,s,X^{\prime})$ and
$\theta\mapsto Q_{\theta}(X,s,s^{\prime})$ are twice continuously
differentiable $a.s.$-$\bar{P}_{\ast}^{\nu}$ for all $(s,s^{\prime}
)\in\mathbb{S}^{2}$.
\end{assumption}

\begin{assumption}
\label{ass:deriva-bdd} For some $\delta>0$ and $a\geq1$, and for all
$(s^{\prime},s)\in\mathbb{S}^{2}$: (i)
\[
E_{\bar{P}_{\ast}^{\nu}}\left[  \sup_{\theta\in B(\delta,\theta_{\ast}
)}\left\Vert \nabla_{\theta}\log p_{\theta}(X,s,X^{\prime})\right\Vert
^{2a}\right]  <\infty\text{ and}~E_{\bar{P}_{\ast}^{\nu}}\left[  \sup
_{\theta\in B(\delta,\theta_{\ast})}\left\Vert \nabla_{\theta}\log Q_{\theta
}(X,s,s^{\prime})\right\Vert ^{2a}\right]  <\infty;
\]
(ii)
\[
E_{\bar{P}_{\ast}^{\nu}}\left[  \sup_{\theta\in B(\delta,\theta_{\ast}
)}\left\Vert \nabla_{\theta}^{2}\log p_{\theta}(X,s,X^{\prime})\right\Vert
^{2a}\right]  <\infty\text{ and}~E_{\bar{P}_{\ast}^{\nu}}\left[  \sup
_{\theta\in B(\delta,\theta_{\ast})}\left\Vert \nabla_{\theta}^{2}\log
Q_{\theta}(X,s,s^{\prime})\right\Vert ^{2a}\right]  <\infty.
\]

\end{assumption}

\begin{assumption}
\label{ass:q-sum-p} $\sum_{j=0}^{\infty}\left(  E_{\bar{P}_{\ast}^{\nu}
}\left[  \prod_{i=0}^{j}(1-\underline{q}(X_{i}))^{\frac{2a}{1-a}}\right]
\right)  ^{p\left(  \frac{1-a}{2a}\right)  }<\infty$ for some $p\in(0,2/3)$
(and the same $a\geq1$ that appears in Assumption \ref{ass:deriva-bdd}).
\end{assumption}

\begin{remark}
[Discussion of Assumptions \ref{ass:Theta-int}, \ref{ass:deriva-bdd} and
\ref{ass:q-sum-p}]Part (i) of Assumption~\ref{ass:Theta-int} is standard in
the literature. The restriction that $\Theta_{\ast}$ is a singleton could be
relaxed using the ideas of \cite{liu2003asymptotics} for non-identified ML
estimators. This extension, albeit interesting, would present nuances that are
beyond the scope of the present paper. Part~(ii) of
Assumption~\ref{ass:Theta-int} is also standard, and so is
Assumption~\ref{ass:deriva-bdd} (see \cite{bickel98} for a discussion).
Finally, Assumption~\ref{ass:q-sum-p} is a strengthening of Assumption
\ref{ass:qbar-sum}, and is required in order to establish the existence of a
random sequence $(\Delta_{t}(\theta_{\ast}))_{t}$ which approximates the
\textquotedblleft score\textquotedblright\ function well (in the sense of
Lemma~\ref{lem:score_approx} in Appendix \ref{app:LAR}). As was the case with
Assumption \ref{ass:qbar-sum}, Lemma \ref{lem:suff-A5andA8} in the
Supplemental Material \ref{SM:SuffAssumptions} provides sufficient conditions
of the form $E_{\bar{P}_{\nu}^{\ast}}\left[  (1-\underline{q}(X_{1}
))^{2a/(1-a)}\right]  <1$. $\triangle$
\end{remark}

The next theorem establishes a LAN-type property for the log-likelihood
criterion function defined in (\ref{logl}).

\begin{theorem}
\label{thm:LAN} Suppose Assumptions \ref{ass:BDD_Q}, \ref{ass:fx-ergo},
\ref{ass:Theta-int}, \ref{ass:deriva-bdd} and \ref{ass:q-sum-p} hold. Then,
there exists a stationary and ergodic process $(\Delta_{t}(\theta_{\ast}
))_{t}$ in $L^{2}(\bar{P}_{\ast}^{\nu})$, a sequence of negative definite
matrices $(\xi_{t}(\theta_{\ast}))_{t}$, and a compact set $K\subseteq\Theta$
that includes $0$, such that, for any $v\in K$,
\begin{align*}
\ell_{T}^{\nu}(X_{0}^{T},\theta_{\ast}+v)-\ell_{T}^{\nu}(X_{0}^{T}
,\theta_{\ast})=  &  v^{\intercal}\left(  T^{-1}\sum_{t=0}^{T}\Delta
_{t}(\theta_{\ast})+o_{\bar{P}_{\ast}^{\nu}}(T^{-1/2})\right) \\
&  +(1/2)v^{\intercal}\left(  T^{-1}\sum_{t=0}^{T}\xi_{t}(\theta_{\ast
})+o_{\bar{P}_{\ast}^{\nu}}(1)\right)  v+R_{T}(v),
\end{align*}
where $v\mapsto R_{T}(v)\in\mathbb{R}$ is such that $\lim_{\delta\rightarrow
0}\bar{P}_{\ast}^{\nu}\left(  \sup_{v\in B(\delta,0)}||v||^{-2}R_{T}
(v)\geq\epsilon\right)  =0$ for any $\epsilon>0$ and any $T\in\mathbb{N}$.
\end{theorem}

\begin{proof}
See Appendix \ref{app:LAR}.
\end{proof}


Theorem~\ref{thm:LAN} extends the results in \cite{bickel96} and
\cite{bickel98} (see their remark on p.~1620) to a more general setup which
allows for covariate-dependent transition probabilities, autoregressive
dynamics, and misspecified models. The proof develops along the same lines as
theirs. The main difference relates to the way in which one establishes that
the \textquotedblleft score\textquotedblright\ $\nabla_{\theta}\ell_{T}^{\nu
}(\cdot,\theta_{\ast})$ and the Hessian $\nabla_{\theta}^{2}\ell_{T}^{\nu
}(\cdot,\theta_{\ast})$ can be approximated by $T^{-1}\sum_{t=0}^{T}\Delta
_{t}(\theta_{\ast})$ and $T^{-1}\sum_{t=0}^{T}\zeta_{t}(\theta_{\ast})$,
respectively (see Lemmas \ref{lem:score_approxV2} and \ref{lem:Hess-approx} in
Appendix \ref{app:LAR}). As mentioned earlier, these approximations rely on
\textquotedblleft mixing\textquotedblright\ properties of the \emph{temporally
inhomogeneous} hidden Markov chain; see Lemma~\ref{lem:Ker-rev} in the
Supplemental Material \ref{SM:LAR}.

Theorem \ref{thm:LAN} may be used to establish the following asymptotic linear
representation for our estimator in terms of $(\Delta_{t}(\theta_{\ast
}))_{t=0}^{\infty}$ and $(\xi_{t}(\theta_{\ast}))_{t=0}^{\infty}$.

\begin{theorem}
\label{thm:anormal} Suppose Assumptions \ref{ass:BDD_Q}--\ref{ass:q-sum-p}
hold and $\eta_{T}=o(T^{-1})$. Then,
\[
\frac{\sqrt{T}(\hat{\theta}_{\nu,T}-\theta_{\ast})}{\sqrt{\mathrm{tr}
\{\Sigma_{T}(\theta_{\ast})\}}}=-\{E_{\bar{P}_{\ast}^{\nu}}[\xi_{1}
(\theta_{\ast})]+o_{\bar{P}_{\ast}^{\nu}}(1)\}^{-1}T^{-1/2}\sum_{t=0}^{T}
\frac{\Delta_{t}(\theta_{\ast})}{\sqrt{\mathrm{tr}\{\Sigma_{T}(\theta_{\ast
})\}}}+o_{\bar{P}_{\ast}^{\nu}}\left(  1\right)  ,
\]
where $\Sigma_{T}(\theta_{\ast})\equiv T^{-1}E_{\bar{P}_{\ast}^{\nu}}
[\{\sum_{t=0}^{T}\Delta_{t}(\theta_{\ast})\}\{\sum_{t=0}^{T}\Delta_{t}
(\theta_{\ast})^{\intercal}\}]$.
\end{theorem}

\begin{proof}
See Appendix \ref{app:LAR}.
\end{proof}


Theorem \ref{thm:anormal} readily implies that, if $T^{-1/2}\Sigma_{T}
(\theta_{\ast})^{-1/2}\sum_{t=0}^{T}\Delta_{t}(\theta_{\ast})\Rightarrow
_{\bar{P}_{\ast}^{\nu}}\mathcal{N}(0,I)$, then
\begin{equation}
\sqrt{T}\Sigma_{T}(\theta_{\ast})^{-1/2}E_{\bar{P}_{\ast}^{\nu}}[\xi
_{1}(\theta_{\ast})](\hat{\theta}_{\nu,T}-\theta_{\ast})\Rightarrow_{\bar
{P}_{\ast}^{\nu}}\mathcal{N}(0,I). \label{eqn:anormtheta}
\end{equation}
(If $\Sigma_{T}(\theta_{\ast})\rightarrow\Sigma(\theta_{\ast})$, $\Sigma
_{T}(\theta_{\ast})$ may be replaced by $\Sigma(\theta_{\ast})$ in these
statements). This result is akin to results in \cite{white82} and shares the
same features, i.e., the asymptotic covariance matrix $\Omega_{T}(\theta
_{\ast})\equiv(E_{\bar{P}_{\ast}^{\nu}}[\xi_{1}(\theta_{\ast})])^{-1}
\Sigma_{T}(\theta_{\ast})(E_{\bar{P}_{\ast}^{\nu}}[\xi_{1}(\theta_{\ast
})])^{-1}$ has a \textquotedblleft sandwich\textquotedblright\ form and the
familiar Fisher information equality does not necessarily hold (see also
\cite[Ch. 6]{white94}).

The next theorem complements Theorem \ref{thm:anormal} by presenting results
on consistent estimation of the asymptotic covariance matrix of $\hat{\theta
}_{\nu,T}$, vis., $\Omega_{T}(\theta_{\ast})$, in both correctly specified and
potentially misspecified models. In the latter case, the proposed estimator is
of the \textquotedblleft heteroskedasticity and autocorrelation
consistent\textquotedblright\ type. As in many other results for such
estimators (e.g., \cite{nw87}, \cite[Ch. 8]{white94}), consistency requires
some conditions over the score covariance process $(\Delta_{t+\tau}
(\cdot)\Delta_{t}(\cdot)^{\intercal})_{t=0}^{T}$, $\tau\in\mathbb{N}$, namely
continuity and law of large numbers results (see Lemma
\ref{lem:Charac.ScoreProcess} in the Supplemental Material \ref{sm:StdErrors}
). In what follows, the modulus of continuity is defined as
\[
\ddot{\varpi}(\delta^{\prime})\equiv\max_{t}\left\Vert \sup_{||\theta
-\theta_{\ast}||<\delta^{\prime}}\left\Vert \Delta_{t}(\theta)\Delta
_{0}(\theta)^{\intercal}-\Delta_{t}(\theta_{\ast})\Delta_{0}(\theta_{\ast
})^{\intercal}\right\Vert \right\Vert _{L^{1}(\bar{P}_{\ast}^{\nu})},
\]
for any $\delta^{\prime}\in(0,\delta]$, where $\delta$ is as in Assumption
\ref{ass:deriva-bdd}.\footnote{Lemma \ref{lem:Charac.ScoreProcess} in the
Supplementary Material \ref{sm:StdErrors} shows that the modulus of continuity
converges to zero as $\delta$ vanishes.}

\begin{theorem}
\label{thm:StdErrors} Suppose the assumptions of Theorem \ref{thm:anormal}
hold.


\begin{enumerate}
\item[(a)] If $\theta_{\ast}$ is such that, for any $t\geq0$ and $T\geq1$,
$p_{t}^{\nu}(\cdot\mid X_{t-T}^{t-1};\theta^{\ast})=p_{\ast}^{\nu}(\cdot\mid
X_{t-T}^{t-1})$, then
\[
||\Omega_{T}(\theta_{\ast})-\{-H_{T}^{-1}(\hat{\theta}_{\nu,T})\}||=o_{\bar
{P}_{\ast}^{\nu}}(1),
\]
where
\[
H_{T}(\theta)\equiv T^{-1}\sum_{t=1}^{T}\nabla_{\theta}^{2}\log p_{t}^{\nu
}(X_{t}|X_{0}^{t-1},\theta).
\]


\item[(b)] If, for any $l\geq0$,
$|| E_{\bar{P}_{\ast}^{\nu}}[ \Delta_{l}(\theta_{\ast}) \Delta_{0}
(\theta_{\ast})^{\intercal} ] || \leq\bar{\upsilon}(l)$
for some integrable function $l\mapsto\bar{\upsilon}(l)\in\mathbb{R}_{+}$,
then
\[
||\Omega_{T}(\theta_{\ast})-H_{T}^{-1}(\hat{\theta}_{\nu,T})J_{T}(\hat{\theta
}_{\nu,T})H_{T}^{-1}(\hat{\theta}_{\nu,T})||=o_{\bar{P}_{\ast}^{\nu}}(1),
\]

\end{enumerate}

where
\begin{align*}
J_{T}(\theta)  &  \equiv T^{-1}\sum_{t=1}^{T}\nabla_{\theta}\log p_{t}^{\nu
}(X_{t}|X_{0}^{t-1},\theta)\nabla_{\theta}\log p_{t}^{\nu}(X_{t}|X_{0}
^{t-1},\theta)^{\intercal}\\
&  +\sum_{\tau=1}^{L_{T}}\frac{\omega(\tau,L_{T})}{T-\tau}\sum_{t=1}^{T-\tau
}\nabla_{\theta}\log p_{t+\tau}^{\nu}(X_{t+\tau}|X_{0}^{t+\tau-1}
,\theta)\nabla_{\theta}\log p_{t}^{\nu}(X_{t}|X_{0}^{t-1},\theta)^{\intercal
}\\
&  +\sum_{\tau=1}^{L_{T}}\frac{\omega(\tau,L_{T})}{T-\tau}\sum_{t=1}^{T-\tau
}\nabla_{\theta}\log p_{t}^{\nu}(X_{t}|X_{0}^{t-1},\theta)\nabla_{\theta}\log
p_{t+\tau}^{\nu}(X_{t+\tau}|X_{0}^{t+\tau-1},\theta)^{\intercal},
\end{align*}
$\omega(\cdot,\cdot)$ are bounded real weights with $\lim_{T\rightarrow\infty
}\omega(\tau,L_{T})=1$ for all $\tau\geq1$, and $(L_{T})_{T=1}^{\infty
}\subseteq\mathbb{N}$ is such that $L_{T}(\ddot{\varpi}(T^{-1/2}\log\log
T)\log\log T+r_{T}+T^{-1/2})=o(1)$, with $(r_{T})_{T}$ being a positive
sequence converging to zero.
\end{theorem}

\begin{proof}
	See Supplemental Material \ref{sm:StdErrors}.
\end{proof}


Part (a) of Theorem~\ref{thm:StdErrors} deals with correctly specified models
and makes use of the Fisher information equality. Part (b) provides a
consistent covariance estimator in the general case of potentially
misspecified models. The conditions for the weights $\omega(\tau,L_{T})$ and
the tuning parameters $(L_{T})_{T}$ are standard for estimators of this
type.\footnote{The additional condition that $\omega(\tau,L_{T})=\sum
_{j=1+\tau}^{L_{T}}c(j,L_{T})c(j-\tau,L_{T})$ for some constants
$c(1,L_{T}),\ldots,c(L_{T},L_{T})$ guarantees that $J_{T}(\hat{\theta}_{\nu
,T})$ is positive semidefinite. Weights obtained from the commonly used
Bartlett, Parzen and quadratic-spectral kernels satisfy this condition.} The
terms $\ddot{\varpi}(T^{-1/2}\log\log T)\log\log T$ and $r_{T}$ in the growth
condition for $L_{T}$ are analogous to those appearing in \cite{nw87} and
arise as \textquotedblleft costs\textquotedblright\ of working with
$\nabla_{\theta}\log p_{t}^{\nu}(X_{t}|X_{0}^{t-1},\hat{\theta}_{\nu,T})$ as
opposed to $\nabla_{\theta}\log p_{t}^{\nu}(X_{t}|X_{0}^{t-1},\theta_{\ast})$,
and of working with sample averages as opposed to population means,
respectively.\footnote{The $\log\log T$ factor is used in order to avoid
working with constants.}
On the other hand, the $T^{-1/2}$ term does not appear in \cite{nw87} and
arises because we need to approximate $\Delta_{t}(\theta_{\ast})$, that
depends on $X_{-\infty}^{t}$, with the score, that depends on $X_{0}^{t}$.

Theorems~\ref{thm:anormal} and \ref{thm:StdErrors}, together with an
asymptotic normality result such as (\ref{eqn:anormtheta}), provide the means
for constructing asymptotically correct confidence sets and hypotheses tests
for $\theta_{\ast}$. In correctly specified models, $(\Delta_{t}(\theta_{\ast
}))_{t=0}^{\infty}$ is a martingale difference sequence, and thus result
(\ref{eqn:anormtheta}) can be obtained by invoking a martingale central limit
theorem. In potentially misspecified models, $(\Delta_{t}(\theta_{\ast
}))_{t=0}^{\infty}$ will not, in general, be a martingale difference sequence,
so one should use a different approach; in some situations, a central limit
theorem for $\beta$-mixing sequences can be used instead. The following
corollary formalizes this discussion.

\begin{corollary}
\label{cor:anormal} Suppose the assumptions of Theorem \ref{thm:anormal}
hold.


\begin{enumerate}
\item[(a)] If $\theta_{\ast}$ is such that, for any $t\geq0$ and $T\geq1$,
$p_{t}^{\nu}(\cdot\mid X_{t-T}^{t-1};\theta^{\ast})=p_{\ast}^{\nu}(\cdot\mid
X_{t-T}^{t-1})$, then
\[
\sqrt{T}\{-H_{T}(\hat{\theta}_{\nu,T})\}^{1/2}(\hat{\theta}_{\nu,T}
-\theta_{\ast})\Rightarrow_{\bar{P}_{\ast}^{\nu}}\mathcal{N}(0,I).
\]


\item[(b)] If there exists $\bar{L}>0$ such that, for any $k,T>\bar{L}$,
$p_{k}^{\nu}(X_{k}\mid X_{k-T}^{k-1};\theta^{\ast})=p_{k}^{\nu}(X_{k}\mid
X_{k-\bar{L}}^{k-1};\theta^{\ast})$, $\lim\inf_{T\rightarrow\infty}e_{\min
}(\Sigma_{T}(\theta_{\ast}))>0$,\footnote{For any real symmetric matrix $A$,
$e_{\min}(A)$ denotes its minimum eigenvalue.} and $E_{\bar{P}_{\ast}^{\nu}
}[\left\Vert \Delta_{1}(\theta_{\ast})\right\Vert ^{4+4\delta}]<\infty$ for
some $\delta>0$, then
\[
\sqrt{T}\hat{\Omega}_{T}(\hat{\theta}_{\nu,T})^{-1/2}(\hat{\theta}_{\nu
,T}-\theta_{\ast})\Rightarrow_{\bar{P}_{\ast}^{\nu}}\mathcal{N}(0,I),
\]
where $\hat{\Omega}_{T}(\hat{\theta}_{\nu,T})\equiv H_{T}^{-1}(\hat{\theta
}_{\nu,T})J_{T}(\hat{\theta}_{\nu,T})H_{T}^{-1}(\hat{\theta}_{\nu,T})$.
\end{enumerate}
\end{corollary}

\begin{proof}
	See Appendix \ref{app:LAR}.
\end{proof}


Regarding the asymptotic normality result, both parts of the corollary rely on
the structure of the process defined by the random variables $\nabla_{\theta
}\log p_{k}^{\nu}(X_{k}\mid X_{-\infty}^{k-1};\theta^{\ast})$; this is a
martingale difference sequence in part (a) and a geometrically $\beta$-mixing
sequence in part (b). The latter result is a consequence of the $\beta$-mixing
structure of $(X_{t})_{t=-\infty}^{\infty}$ and of the fact that the
conditional density at $\theta_{\ast}$ depends on a finite number of lags.
Examples of models for which such properties hold true are mixture models (cf.
Example~\ref{exa:HMM-miss} in the Supplemental Material
\ref{SM:examplemixture}) and mixture autoregressive models (cf.
Example~\ref{exa:MAM}).\footnote{At this level of generality, we cannot
establish an asymptotic normality result in the general case where the model
is misspecified and $\nabla_{\theta}\log p_{k}^{\nu}(X_{k}\mid X_{-\infty
}^{k-1};\theta^{\ast})$ depends on the entire $X_{-\infty}^{k-1}$. The reason
is that, although the process $X_{-\infty}^{\infty}$ is $\beta$-mixing, there
is no guarantee that $\nabla_{\theta}\log p_{k}^{\nu}(X_{k}\mid X_{-\infty
}^{k-1};\theta^{\ast})$ inherits these mixing properties, or that it is a
martingale difference sequence.}

Regarding the covariance estimators used in parts (a) and (b) of the
corollary, these rely on the corresponding parts of Theorem
\ref{thm:StdErrors}. The result in part (b) is established by exploiting the
$\beta$-mixing structure of the score process $(\Delta_{k}(\theta_{\ast
}))_{k=-\infty}^{\infty}$ and the finiteness of its $4+4\delta$ moments (under
$\bar{P}_{\ast}^{\nu}$) for some $\delta>0$. Lemma \ref{lem:suff.HAC.Rate} in
Appendix \ref{app:LAR} shows that $r_{T}=L_{T}T^{-1/2}$, and, since
$\theta\mapsto\log p_{\theta}$ and $\theta\mapsto\log Q_{\theta}$ are smooth
(see Assumptions \ref{ass:Theta-int} and \ref{ass:deriva-bdd}), it follows
that $\delta^{\prime}\mapsto\ddot{\varpi}(\delta^{\prime})=C\delta^{\prime}$
for some finite constant $C$. Hence, the growth condition on $(L_{T})_{T}$
translates into $L_{T}T^{-1/4}\log\log T=o(1)$, which is analogous to that in
\cite{nw87} (apart from the $\log\log T$ factor, which was introduced in
Theorem \ref{thm:StdErrors} for convenience).

The next example verifies the assumptions in the context of the models
considered in Example \ref{exa:Canon}.

\begin{example}
\label{exa:Canon.Inference copy(1)} In view of the results in Example
\ref{exa:Canon}, we only need to verify Assumptions \ref{ass:Theta-int}
--\ref{ass:q-sum-p}. Part (i) of Assumption \ref{ass:Theta-int} is standard
and is directly imposed, while part (ii) follows from the setup of the
example. Assumption \ref{ass:deriva-bdd} follows by the continuity of the
derivatives. Finally, Lemma \ref{lem:suff-A5andA8} in the Supplemental
Material \ref{SM:consistent} implies that Assumption \ref{ass:q-sum-p} holds.

Thus, Theorem \ref{thm:anormal} holds for the class of models considered in
Example \ref{exa:Canon}. In particular, in the correctly specified case,
Corollary \ref{cor:anormal}(a) guarantees asymptotic normality of the
studentized ML estimator of $\theta_{\ast}$, thereby providing the basis for
inference. These results are, to our knowledge, new in the context of
Markov-switching autoregressive models with covariate-dependent transition
probabilities. $\triangle$
\end{example}



\section{Monte Carlo Simulations}

\label{sec:simulation}

The objective in this section is twofold. First, to assess the quality of
approximations provided by our asymptotic results by examining the
finite-sample properties of the ML estimator and related statistics in a
correctly specified Markov-switching autoregressive model with
covariate-dependent transition probabilities. Second, to explore the effects
of a type of empirically relevant misspecification which involves the use of
an incomplete approximation to the likelihood function that ignores potential
contemporaneous correlation between the observation variable $(Y_{t})$ and the
variable $(Z_{t})$ upon the lagged value of which the transition probabilities depend.

Monte Carlo experiments are based on artificial data $(X_{t}=(Y_{t}
,Z_{t}))_{t}$ generated according to the equations
\begin{align}
Y_{t}  &  =\mu_{0}(1-S_{t})+\mu_{1}S_{t}+\phi Y_{t-1}+[\sigma_{0}
(1-S_{t})+\sigma_{1}S_{t}]U_{1,t},\label{mcy}\\
Z_{t}  &  =\mu_{2}+\psi Z_{t-1}+\sigma_{2}U_{2,t}, \label{mcz}
\end{align}
for $t\in\mathbb{N}$, with $X_{0}=(0.5,0.4)$, $\mu_{0}=-\mu_{1}=1$, $\phi
=0.9$, $\sigma_{0}=\sigma_{1}=1$, $\mu_{2}=0.2$, $\psi=0.8$, $\sigma_{2}=0.6$,
and $(U_{1,t},U_{2,t})_{t}\thicksim$ $\mathrm{i.i.d.}~\mathcal{N}\left(
\left[
\begin{array}
[c]{c}
0\\
0
\end{array}
\right]  ,\left[
\begin{array}
[c]{cc}
1 & \rho\\
\rho & 1
\end{array}
\right]  \right)  $ with $\rho\in\{0,0.8\}$. The regimes $(S_{t})_{t}$ are a
Markov chain on $\{0,1\}$, independent of $(U_{1,t},U_{2,t})_{t}$, with
transition probabilities
\begin{equation}
Q_{\theta}(z,s,s)\equiv\Pr(S_{t}=s\mid Z_{t-1}=z,S_{t-1}=s)=[1+\exp
(-\alpha_{s}-\beta_{s}z)]^{-1}, \label{tvtpm}
\end{equation}
for $s\in\{0,1\}$, where $\alpha_{0}=\alpha_{1}=2$ and $\beta_{0}=-\beta
_{1}=-0.5$. The model defined by (\ref{mcy})--(\ref{tvtpm}) is a prototypical
Markov-switching autoregressive model with covariate-dependent transition
probabilities (cf. Example \ref{exa:MSR}). In each of 1000 independent Monte
Carlo replications, $100+T$ data points for $(X_{t})_{t}$ are generated, with
$T\in\{200,800,1600,3200\}$, and the last $T$ points are used to compute
estimates of the parameters of interest. In order to conserve space, only a
selection of the results are reported (the full set of results is available
upon request).

\subsection{Correct Specification}

\label{sec:simulation1}

In the first set of experiments, we consider estimation of the parameters of
the model in (\ref{mcy})--(\ref{tvtpm}) using the likelihood function based on
the conditional distribution of $X_{t}$ given $X_{0}^{t-1}$. Table
\ref{tab:MC-1n} reports the deviation of the mean of the finite-sample
distributions of the ML estimators of the elements of $\vartheta=(\mu_{0}
,\mu_{1},\phi,\sigma_{0},\sigma_{1},\alpha_{0},\beta_{0},\alpha_{1},\beta
_{1})$ from the corresponding true parameter values (bias) when $\rho=0.8$. We
also report the ratio of the sampling standard deviation of the estimators to
the estimated standard errors (averaged across replications); the latter are
computed using the Hessian estimator (cf. Theorem~\ref{thm:StdErrors}
(a)).\footnote{Results for $\rho=0$ are not reported since they are very
similar to those for $\rho=0.8$.} Although the estimators of $\beta_{0}$ and
$\beta_{1}$ are somewhat biased when $T=200$, bias is insignificant in the
rest of the cases. Estimated standard errors are somewhat downwards biased in
most cases but, unless $T$ is small, the bias is not generally substantial and
decreases as $T$ increases.


{\footnotesize \begin{table}[ptb]
\caption{Bias and Standard Deviation of ML Estimators $(\rho=0.8)$}
\label{tab:MC-1n}
\begin{center}
{\footnotesize
\begin{tabular}
[c]{cccccccccc}\hline\hline
$T$ & $\mu_{0}$ & $\mu_{1}$ & $\alpha_{1}$ & $\beta_{1}$ & $\alpha_{0}$ &
$\beta_{0}$ & $\sigma_{0}$ & $\sigma_{1}$ & $\phi$\\\hline
\multicolumn{1}{l}{} & \multicolumn{9}{l}{Bias}\\
\multicolumn{1}{l}{200} & -0.007 & 0.023 & 0.004 & 0.126 & 0.043 & 0.209 &
-0.011 & -0.005 & -0.005\\
\multicolumn{1}{l}{800} & -0.002 & 0.004 & -0.002 & 0.023 & 0.002 & 0.026 &
-0.001 & 0.000 & -0.001\\
\multicolumn{1}{l}{1600} & -0.001 & 0.002 & -0.002 & 0.019 & 0.007 & 0.019 &
0.001 & 0.000 & -0.001\\
\multicolumn{1}{l}{3200} & 0.001 & 0.002 & 0.002 & -0.005 & 0.002 & 0.011 &
0.001 & 0.000 & 0.000\\
\multicolumn{1}{l}{} & \multicolumn{9}{l}{Ratio of sampling standard deviation
to estimated standard error}\\
\multicolumn{1}{l}{200} & 1.085 & 1.016 & 1.084 & 1.078 & 1.132 & 1.178 &
1.066 & 1.043 & 1.072\\
\multicolumn{1}{l}{800} & 0.983 & 0.991 & 0.980 & 1.024 & 1.044 & 1.018 &
1.012 & 1.010 & 1.045\\
\multicolumn{1}{l}{1600} & 1.021 & 1.002 & 0.964 & 0.955 & 1.033 & 1.042 &
1.000 & 0.999 & 0.997\\
\multicolumn{1}{l}{3200} & 1.026 & 0.979 & 1.010 & 1.017 & 0.990 & 0.955 &
1.012 & 1.020 & 1.020\\\hline\hline
\end{tabular}
}
\end{center}
\end{table}}

We also examine conventional hypothesis tests for $\vartheta$. Table
\ref{tab:MC-2n} reports the rejection frequencies of: (i)~a $t$-type test of
$\mathcal{H}_{0}:\vartheta_{j}=\vartheta_{j}^{\ast}$ versus $\mathcal{H}
_{1}:\vartheta_{j}\neq\vartheta_{j}^{\ast}$, where $\vartheta_{j}$ is the
$j$-th element of $\vartheta$ and $\vartheta_{j}^{\ast}$ is its true value;
(ii)$~$a $t$-type test of $\mathcal{H}_{0}:\vartheta_{j}=0$ versus
$\mathcal{H}_{1}:\vartheta_{j}\neq0$. These rejection frequencies are referred
to as \textquotedblleft size\textquotedblright\ and \textquotedblleft
power\textquotedblright, respectively, and are computed using the 0.975
standard-normal quantile as critical value.\footnote{Results should be
interpreted with caution in the case of $\mathcal{H}_{0}:\sigma_{i}=0$,
$i\in\{0,1\}$, because the null value of $\sigma_{i}$ is on the boundary of
the maintained hypothesis. Our asymptotic theory does not allow for parameters
that may lie on the boundary of the parameter space.} Tests tend to have
Type~I error probabilities which are generally close to the nominal 0.05
level, especially for $T>200$. Tests are also powerful enough to reject the
hypothesis of a zero parameter value, except in the case of $\beta_{0}$ and
$\beta_{1}$ with $T=200$. The distributions of studentized statistics
associated with the elements of the ML estimator of $\vartheta$ (ratio of
estimation error to corresponding estimated standard error) generally tend to
have mean and variance (not shown) that do not differ substantially from zero
and one, respectively, and Gaussianity is never rejected for $T>200$
.\footnote{Statistics associated with $\phi$, $\sigma_{0}$ and $\sigma_{1}$
appear to fare somewhat worse than others when $T=200$, a finding similar to
that reported in \cite{ps98} for models with a time-invariant transition
mechanism.}

{\footnotesize \begin{table}[ptb]
\caption{Size and Power of $t$-Type Tests $(\rho=0.8)$}
\label{tab:MC-2n}
\begin{center}
{\footnotesize
\begin{tabular}
[c]{cccccccccc}\hline\hline
$T$ & $\mu_{0}$ & $\mu_{1}$ & $\alpha_{1}$ & $\beta_{1}$ & $\alpha_{0}$ &
$\beta_{0}$ & $\sigma_{0}$ & $\sigma_{1}$ & $\phi$\\\hline
\multicolumn{1}{l}{} & \multicolumn{9}{l}{Size}\\
\multicolumn{1}{l}{200} & 0.063 & 0.076 & 0.069 & 0.069 & 0.081 & 0.043 &
0.084 & 0.088 & 0.108\\
\multicolumn{1}{l}{800} & 0.053 & 0.063 & 0.061 & 0.051 & 0.074 & 0.049 &
0.065 & 0.065 & 0.073\\
\multicolumn{1}{l}{1600} & 0.065 & 0.061 & 0.047 & 0.043 & 0.060 & 0.048 &
0.052 & 0.047 & 0.068\\
\multicolumn{1}{l}{3200} & 0.049 & 0.058 & 0.053 & 0.058 & 0.054 & 0.042 &
0.051 & 0.061 & 0.053\\
\multicolumn{1}{l}{} & \multicolumn{9}{l}{Power}\\
\multicolumn{1}{l}{200} & 0.999 & 1.000 & 0.990 & 0.261 & 0.934 & 0.207 &
1.000 & 1.000 & 1.000\\
\multicolumn{1}{l}{800} & 1.000 & 1.000 & 1.000 & 0.812 & 0.999 & 0.732 &
1.000 & 1.000 & 1.000\\
\multicolumn{1}{l}{1600} & 1.000 & 1.000 & 1.000 & 0.988 & 1.000 & 0.966 &
1.000 & 1.000 & 1.000\\
\multicolumn{1}{l}{3200} & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 &
1.000 & 1.000 & 1.000\\\hline\hline
\end{tabular}
}
\end{center}
\end{table}}

\subsection{Misspecification}

\label{sec:simulation2}

In the second set of experiments, we consider estimation of the parameter
$\vartheta=(\mu_{0},\mu_{1},\phi,\sigma_{0},\sigma_{1},\alpha_{0},\beta
_{0},\alpha_{1},\beta_{1})$ using the partial likelihood function based on the
conditional distribution of $Y_{t}$ given $(Y_{t-1},S_{t})$. Inference in
models like (\ref{mcy})--(\ref{tvtpm}) is predominantly based on such a
partial likelihood that ignores the equation for $Z_{t}$ (see, e.g.,
\cite{dieb94}, \cite{filardo94}). Formally, the misspecified model may be
viewed as defined by equations (\ref{mcy})--(\ref{tvtpm}), with the additional
assumption that $\rho=0$. Under this assumption, estimation of $\vartheta$ is
based on (\ref{mcy}) alone, the (potentially incorrect) rational behind this
approach being that, since the conditional distribution of $Y_{t}$ given
$(Y_{t-1},S_{t})$ and the transition probabilities of $(S_{t})_{t}$ depend
only on $Z_{t-1}$, (\ref{mcy}) may be analyzed, without loss of relevant
information, independently of $(Z_{t})_{t}$. Even though such an approach may
be appealing because of its relative simplicity, it is far from obvious that
it provides a valid way for conducting inference on $\vartheta$; it is
unclear, for example, what the limit point of the ML estimator based on the
partial likelihood might be when $\rho\neq0$. In earlier sections, we
considered a theoretical framework that acknowledges this source of
misspecification (among others) and provided tools for asymptotically valid
inference. We now quantify the implications of this misspecification in finite
samples. For brevity, we refer to the maximizer of the partial likelihood
function associated with the conditional distribution of $Y_{t}$ given
$(Y_{t-1},S_{t})$ as the `partial ML' estimator to distinguish it from the
`joint ML' estimator based on the joint model for the conditional distribution
of $X_{t}$ given $X_{0}^{t-1}$ (cf. Section~\ref{sec:simulation1}).



Table \ref{tab:MC-3n} shows the estimated bias of the partial ML estimators of
the elements of $\vartheta$ and the ratio of the sampling standard deviation
of the estimators to the estimated standard errors (averaged across
replications) when $\rho=0.8$. To reflect what is common practice in applied
research, standard errors are computed using the Hessian estimator (which
relies on the assumption of a correctly specified likelihood) instead of a
\textquotedblleft sandwich\textquotedblright\ estimator (which allows for
mispecification). It is immediately apparent that the partial ML estimator of
most of the parameters is considerably more biased than the joint ML
estimator. The differences between the two estimators are more pronounced for
parameters associated with the transition probabilities ($\alpha_{0}$,
$\beta_{0}$, $\alpha_{1}$, $\beta_{1}$), the partial ML estimators of which
are significantly biased even for $T=3200$. This suggests that the bias of the
partial ML estimator when $\rho\neq0$ is not associated only with small
samples, a finding that is consistent with our asymptotic results. Regarding
the accuracy of estimated standard errors, the latter are downwards biased in
most cases, the bias being somewhat larger than it is for joint ML estimators.
However, unless $T$ is small, this bias is not generally substantial and
declines as $T$ increases, despite the fact that standard errors are obtained
from the Hessian.\footnote{Results for $\rho=0$ (not shown) are not
substantially different from those obtained from the joint ML\ procedure. This
is not surprising since the joint and partial ML estimators are both
consistent for the true parameter value when $\rho=0$.}
{\footnotesize \begin{table}[ptb]
\caption{Bias and Standard Deviation of Partial ML Estimators $(\rho=0.8)$}
\label{tab:MC-3n}
\begin{center}
{\footnotesize
\begin{tabular}
[c]{cccccccccc}\hline\hline
$T$ & $\mu_{0}$ & $\mu_{1}$ & $\alpha_{1}$ & $\beta_{1}$ & $\alpha_{0}$ &
$\beta_{0}$ & $\sigma_{0}$ & $\sigma_{1}$ & $\phi$\\\hline
\multicolumn{1}{l}{} & \multicolumn{9}{l}{Bias}\\
\multicolumn{1}{l}{200} & 0.019 & 0.046 & -0.172 & 0.898 & 0.397 & 1.363 &
-0.056 & -0.039 & -0.001\\
\multicolumn{1}{l}{800} & 0.017 & 0.017 & -0.156 & 0.485 & 0.216 & 0.838 &
-0.025 & -0.024 & 0.002\\
\multicolumn{1}{l}{1600} & 0.008 & 0.011 & -0.150 & 0.446 & 0.211 & 0.824 &
-0.018 & -0.021 & 0.003\\
\multicolumn{1}{l}{3200} & 0.012 & 0.008 & -0.149 & 0.444 & 0.196 & 0.773 &
-0.020 & -0.019 & 0.003\\
\multicolumn{1}{l}{} & \multicolumn{9}{l}{Ratio of sampling standard deviation
to estimated standard error}\\
\multicolumn{1}{l}{200} & 1.244 & 1.135 & 1.395 & 1.585 & 1.349 & 1.406 &
1.161 & 1.062 & 1.250\\
\multicolumn{1}{l}{800} & 1.027 & 1.041 & 1.153 & 1.170 & 1.067 & 1.119 &
0.988 & 1.020 & 1.080\\
\multicolumn{1}{l}{1600} & 0.991 & 1.014 & 1.070 & 1.075 & 1.072 & 1.161 &
1.006 & 0.959 & 1.054\\
\multicolumn{1}{l}{3200} & 1.047 & 1.027 & 1.050 & 1.063 & 1.034 & 1.088 &
1.021 & 1.004 & 1.056\\\hline\hline
\end{tabular}
}
\end{center}
\end{table}}

We note that hypothesis tests based on studentized statistics analogous to
those considered in Section~\ref{sec:simulation1} (not shown) are unreliable
when partial ML estimates are used. This is especially true in the case of
parameters associated with the transition probabilities, the corresponding
tests being either excessively conservative or excessively liberal. Although
tests of this type are extensively used in applied work, they should be
interpreted with caution since the statistics on which they are based have an
asymptotically normal null distribution only when $\rho=0$. We also note that
using the \textquotedblleft sandwich\textquotedblright\ estimator of
Theorem~\ref{thm:StdErrors}(b) instead of the estimator based on the observed
information matrix is not without difficulty when $\rho\neq0$. As
\cite{Freedman06} points out, the use of such an estimator for inference is
unlikely to produce results that are any less misleading under
misspecification since the problem of bias/inconsistency of the ML estimator
for the true parameter value remains. It is indeed clear from the results in
Table~\ref{tab:MC-3n} that the bias of the partial ML estimator presents a
much more serious problem in our setting than the inaccuracy of conventionally
computed standard errors.



\section{Empirical Illustration}

\label{sec:illustration}

In this section, we present an empirical illustration based on a
regime-switching model of a type that is commonly used in
economics.\footnote{An additional empirical example, which examines the
predictive ability of an index of leading indicators for regime changes in
U.S. output growth, is discussed in the Supplemental
Material~\ref{SM:empiricalf}.} Specifically, we investigate the potential
contribution of the interest rate spread and the growth in tax revenues in
predicting regime changes in U.S. real output growth. The model is a variant
of the specification used in the simulations and is given by
\begin{align}
Y_{t}  &  =\mu_{0}(1-S_{t})+\mu_{1}S_{t}+\sum_{i=1}^{h_{1}}\phi_{i}
Y_{t-i}+\sigma_{1}U_{1,t},\label{eqy}\\
Z_{t}  &  =\mu_{2}+\sum_{i=1}^{h_{2}}\psi_{i}Z_{t-i}+\sigma_{2}U_{2,t},
\label{eqz}
\end{align}
for some $h_{1},h_{2}\in\mathbb{N}$, with the hidden, two-state Markov chain
$(S_{t})_{t}$ being governed by the transition probabilities
\begin{equation}
Q_{\theta}(z,s,s)\equiv\Pr(S_{t}=s\mid Z_{t-1}=z,S_{t-1}=s)=[1+\exp
(-\alpha_{s}-\beta_{s}z)]^{-1}, \label{eqtp}
\end{equation}
with $s\in\{0,1\}$, and $(U_{1,t},U_{2,t})_{t}$ postulated to be i.i.d.
$\mathcal{N}\left(  \left[
\begin{array}
[c]{c}
0\\
0
\end{array}
\right]  ,\left[
\begin{array}
[c]{cc}
1 & \rho\\
\rho & 1
\end{array}
\right]  \right)  $ and independent of $(S_{t})_{t}$. In (\ref{eqy}
)--(\ref{eqtp}), $Y_{t}$ stands for the growth rate of real gross domestic
product and $Z_{t}$ is either the spread between the 10-year Treasury note
rate and the 3-month Treasury bill rate or the growth rate of real government
receipts of direct and indirect taxes.\footnote{The model could be generalized
to allow for Markov changes in all the parameters. However, since $Z_{t}$ is
thought of here as a potential leading indicator for business-cycle phases, it
does not seem sensible to allow the parameters in both (\ref{eqy}) and
(\ref{eqz}) to be subject to changes driven by $(S_{t})_{t}$. Modeling regime
changes in $(Y_{t})_{t}$ and $(Z_{t})_{t}$ as being driven by two independent
Markov processes is more attractive, but we choose to abstract from this as it
is not directly related to the main problem under study.} The data are
quarterly and span the period 1954:3--2009:2.\footnote{Interest rate data are
taken from the FRED$^{\textregistered}$ database; output and tax data are
taken from \cite{auerbach12}. The likelihood ratio test of \cite{Hansen92}
rejects the hypothesis that $\mu_{0}=\mu_{1}$ in (\ref{eqy}).}

Since the aim is not only to assess the predictive ability of the interest
rate spread and tax revenues for regime changes in output growth but also to
examine whether treating these variables as exogenous yields results which are
different from those obtained from a joint model, we compute two sets of
estimates: partial ML estimates based on (\ref{eqy}) alone and joint ML
estimates based on the system (\ref{eqy})--(\ref{eqz}). We note that in
econometric models of the business cycle such as (\ref{eqy})--(\ref{eqtp}), it
is common to rely on partial ML estimation (see, e.g., \cite{filardo94}).
Parameter estimates are reported in Tables \ref{tab:app-1} and \ref{tab:app-2}
, with estimated standard errors given in parentheses; the latter are obtained
from the \textquotedblleft sandwich\textquotedblright\ estimator of
Theorem~\ref{thm:StdErrors}(b).\footnote{The weights $\omega(\cdot,L_{T})$ are
obtained from the Parzen kernel and $L_{T}$ is determined using the automatic
plug-in method of \cite{andrews91}.
\par
{}} On the basis of $t$-type tests based on joint ML estimates, at least one
of the parameters $(\beta_{0},\beta_{1})$ is significantly different from
zero, indicating that the spread and tax revenues contain significant
information about the probability of switching between the two regimes.

{\footnotesize \ \begin{table}[ptb]
\caption{ML Estimates (Output Growth, Interest Rate Spread)}
\label{tab:app-1}
{\footnotesize  \ \  }
\par
\begin{center}
{\footnotesize \ \
\begin{tabular}
[c]{lllllllll}\hline\hline
\multicolumn{2}{c}{Partial ML} &  &  &  & \multicolumn{4}{c}{Joint ML}\\\hline
$\mu_{0}$ & 0.0091 (0.0012) &  &  &  & $\mu_{0}$ & 0.0092 (0.0012) &  & \\
$\mu_{1}$ & 0.0001 (0.0013) &  &  &  & $\mu_{1}$ & 0.0004 (0.0013) & $\mu_{2}$
& 0.2956 (0.0984)\\
$\phi_{1}$ & 0.1543 (0.0761) &  &  &  & $\phi_{1}$ & 0.1449 (0.0763) &
$\psi_{1}$ & 0.8098 (0.0012)\\
$\phi_{2}$ & 0.0620 (0.0855) &  &  &  & $\phi_{2}$ & 0.0595 (0.0852) &
$\psi_{2}$ & -0.1065 (0.1251)\\
$\phi_{3}$ & -0.0468 (0.0867) &  &  &  & $\phi_{3}$ & -0.0435 (0.0759) &
$\psi_{3}$ & 0.3368 (0.1747)\\
$\phi_{4}$ & -0.0136 (0.0929) &  &  &  & $\phi_{4}$ & -0.0183 (0.0914) &
$\psi_{4}$ & -0.2495 (0.0859)\\
$\alpha_{0}$ & -1.3367 (3.8866) &  &  &  & $\alpha_{0}$ & -1.5210 (2.6935) &
$\sigma_{2}$ & 0.7053 (0.1091)\\
$\beta_{0}$ & 8.8363 (9.4778) &  &  &  & $\beta_{0}$ & 9.1169 (6.4769) &
$\rho$ & -0.0849 (0.1198)\\
$\alpha_{1}$ & 3.1927 (0.7646) &  &  &  & $\alpha_{1}$ & 3.1218 (0.7688) &  &
\\
$\beta_{1}$ & -1.0779 (0.4304) &  &  &  & $\beta_{1}$ & -1.0338 (0.4293) &  &
\\
$\sigma_{1}$ & 0.0077 (0.0006) &  &  &  & $\sigma_{1}$ & 0.0078 (0.0006) &  &
\\\hline\hline
\end{tabular}
}
\end{center}
\end{table}}

Regarding the implications of treating $Z_{t}$ as exogenous, the differences
between partial and joint ML estimates are substantial in the model with tax
revenues (especially for autoregressive coefficients and the parameters
associated with the transition probabilities) but much less so in the model
with the interest rate spread. This is not entirely unexpected in view of the
fact that the estimated value of the conditional correlation $\rho$ is
relatively large (0.6034) in the former model but much smaller ($-$0.0849),
and insignificantly different from zero, in the latter. Such findings are in
line with the analytical and simulation results presented in previous
sections. The relatively large estimate of $\rho$ in Table~\ref{tab:app-2}
also suggests that inference based on the partial ML estimator is potentially
misleading because of the likely bias of the estimator.

{\footnotesize \ \begin{table}[pth]
\caption{ML Estimates (Output Growth, Growth in Taxes)}
\label{tab:app-2}
{\footnotesize  \  }
\par
\begin{center}
{\footnotesize \ \
\begin{tabular}
[c]{lllllllll}\hline\hline
\multicolumn{2}{c}{Partial ML} &  &  &  & \multicolumn{4}{c}{Joint ML}\\\hline
$\mu_{0}$ & 0.0071 (0.0014) &  &  &  & $\mu_{0}$ & 0.0081 (0.0014) &  & \\
$\mu_{1}$ & -0.0119 (0.0034) &  &  &  & $\mu_{1}$ & -0.0100 (0.0075) &
$\mu_{2}$ & 0.5455 (0.2633)\\
$\phi_{1}$ & 0.2076 (0.0679) &  &  &  & $\phi_{1}$ & 0.0987 (0.0665) &
$\psi_{1}$ & 0.1723 (0.1220)\\
$\phi_{2}$ & 0.0709 (0.0953) &  &  &  & $\phi_{2}$ & 0.0754 (0.0833) &
$\sigma_{2}$ & 3.0458 (0.3082)\\
$\phi_{3}$ & -0.0530 (0.0754) &  &  &  & $\phi_{3}$ & -0.1168 (0.0610) &
$\rho$ & 0.6034 (0.0673)\\
$\phi_{4}$ & -0.0291 (0.0885) &  &  &  & $\phi_{4}$ & -0.0345 (0.0721) &  & \\
$\alpha_{0}$ & 3.4588 (0.6379) &  &  &  & $\alpha_{0}$ & 3.8835 (1.3467) &  &
\\
$\beta_{0}$ & 0.2754 (0.1237) &  &  &  & $\beta_{0}$ & 0.4061 (0.1379) &  & \\
$\alpha_{1}$ & 0.3852 (0.8704) &  &  &  & $\alpha_{1}$ & -2.5349 (4.7906) &  &
\\
$\beta_{1}$ & 0.2558 (0.1053) &  &  &  & $\beta_{1}$ & 0.0579 (0.1713) &  & \\
$\sigma_{1}$ & 0.0076 (0.0006) &  &  &  & $\sigma_{1}$ & 0.0084 (0.0009) &  &
\\\hline\hline
\end{tabular}
}
\end{center}
\end{table}}





\newpage