EconBase
← Back to paper

Maximum Likelihood Estimation in Markov Regime-Switching Models with Covariate-Dependent Transition Probabilities

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

80,468 characters · 10 sections · 89 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Maximum Likelihood Estimation in Markov Regime-Switching Models with Covariate-Dependent Transition Probabilities

abstractThis paper considers maximum likelihood (ML) estimation in a large class of models with hidden Markov regimes. We investigate consistency of the ML estimator and local asymptotic normality for the models under general conditions which allow for autoregressive dynamics in the observable process, Markov regime sequences with covariate-dependent transition matrices, and possible model misspecification. A Monte Carlo study examines the finite-sample properties of the ML estimator in correctly specified and misspecified models. An empirical application is also discussed. \noindentKey words and phrases: Autoregressive model; consistency; covariate-dependent transition probabilities; hidden Markov model; Markov-switching model; maximum likelihood; local asymptotic normality; misspecified models.

Introduction

Stochastic models with parameters that are subject to changes driven by an unobservable Markov chain (the regime or state sequence) have attracted considerable attention in many different areas, the influential work by hamilton89 being a prominent example from the econometrics literature. An important subclass of such models, the so-called hidden Markov models, in which observations are conditionally independent given the regime sequence, are also widely used in a variety of disciplines. A common assumption in these models is that the unobservable Markov chain is temporally homogeneous.

In this paper, we focus on a larger class of models in which the hidden regime process and the observation process (conditional on the regimes) are both temporally inhomogeneous Markov chains. This is a useful generalization of models with a time-invariant transition mechanism that has found numerous applications, especially in economics and finance.\footnote{Examples include, among many others, applications to the analysis of business-cycle fluctuations (e.g., filardo94, FilardoGordon98, ravn99, Simpson01, rivas15), interest rates and yields (e.g., Gray96, BekaertH01, ang02, angbek02j, pses21), consumption growth (Whitelaw00), currency crises (e.g., peria02, mourat08), and cryptocurrency returns (tan21).} Statistical inference in this class of models is predominantly likelihood based, even though very little is known about the asymptotic properties of the relevant inferential procedures. In a typical application, inference is conducted on the implicit assumption that the maximum likelihood (ML) estimator of unknown parameters has its familiar properties of consistency, asymptotic normality, and asymptotic efficiency, and associated confidence sets and hypotheses tests are constructed in the usual manner. It hardly needs noting that, unless these properties of likelihood-based inferential procedures known from regular parametric estimation problems hold for Markov regime-switching models with covariate-dependent transition probabilities, inferences drawn from them cannot be justified in any meaningful way and should be interpreted very cautiously. Arguing, for instance, as is common in applied work, that an economic variable is a useful leading indicator for business-cycle phases because it appears to have a `statistically significant' coefficient in the transition functions of a Markov-switching model for output growth is problematic when little is known about the properties of the relevant estimators and related tests.

The main contribution of this paper is to provide consistency and asymptotic normality results for a large class of models that are relevant in applications. Our approach allows for autoregressive dynamics in the observable process, covariate-dependence in the transition functions of the hidden regime process, and potential model misspecification. To the best of our knowledge, the only asymptotic results available on ML estimation in Markov regime-switching models with covariate-dependent transition probabilities are those of aill13, who investigate consistency of the ML estimator in a correctly specified model (i.e., a model which contains the data-generating process). Our results include both consistency of the ML estimator and local asymptotic normality (LAN) for the model, from which asymptotic normality of the ML estimator can be inferred. Unlike aill13, who allow for a general hidden state space, we require the latter to be finite, but do not restrict the model to be correctly specified. In doing so, we also extend some results of white82 for independent, identically distributed (i.i.d.) data to the case of dependent observations and for classes of parametric distributions associated with dynamic models with hidden Markov regimes. As such stochastic specifications are typically highly parametric, it is important to understand the properties of likelihood-based inferential procedures in situations where the true probability structure of the data does not necessarily lie within the parametric family of distributions specified by the model. We show that the ML estimator in our setting converges to the true parameter value if the model is correctly specified and to a pseudo-true parameter set if the model is misspecified.\footnote{We refer to estimators for the various models discussed throughout as ML estimators even though they may be obtained from a pseudo-likelihood based on a misspecified model.} We also show that the sample log-likelihood satisfies the LAN property, establish an asymptotic linear representation for the ML estimator, obtain the asymptotic distribution of the estimator, and present results relating to consistent estimation of its asymptotic covariance matrix. These are the most general results available for Markov regime-switching models with autoregressive dynamics and covariate-dependent transition probabilities.

In related earlier work, Mevel04 examine consistency and asymptotic normality of the ML estimator in misspecified hidden Markov models with a finite state space, while douc12 consider consistency under general state spaces. bickel96, bickel98, Jensen99, douc01, and douc11 investigate consistency and/or asymptotic normality in correctly specified hidden Markov models with regime sequences defined on either a finite or a general state space. Francq98 and Krishna98 consider consistency in correctly specified autoregressive models with Markov regimes defined on a finite state space. douc04 and Kasahara2019 investigate consistency and asymptotic normality in a similar autoregressive setup, but allow the regime sequence to take values in a space that is not necessarily countable. In all of these papers, the hidden regime sequence is assumed to be a temporally homogeneous Markov chain.

In the sequel, we follow bickel98 and douc04 fairly closely in terms of the technical tools and the arguments used to establish our results, but our setup is more general in certain respects. Like bickel98, we consider models with a finite hidden state space, but allow for autoregressive dynamics in the observation sequence, covariate-dependence in the transition probabilities of the regime sequence, and potential model misspecification. In douc04, the hidden Markov chain is allowed to take values in a compact topological space, but is restricted to be temporally homogeneous and the model is assumed to be correctly specified. The cornerstone of the methods used in these papers for establishing the asymptotic properties of the ML estimator are mixing-type results for the unobservable regime sequence conditional on the observation sequence (see also bickel96). This is also true for our approach, although the aforementioned results cannot be invoked directly because they are established under the assumption of temporal homogeneity of the hidden Markov chain. We, therefore, extend these results to allow for more general Markov regime sequences; in particular, we establish mixing-type results for the unobservable regime sequence given the observed data, allowing for a particular form of covariate-dependence in the transition kernels. This last result is, to our knowledge, novel and may be of interest in its own right.

The remainder of the paper is organized as follows. Section (ref) defines the class of models under consideration, gives sufficient conditions for stationarity and ergodicity of the observation process, and describes the estimation problem. Section (ref) investigates consistency of the ML estimator in a general setting. Section (ref) contains results on the LAN property of the model and the asymptotic normality of the ML estimator. Section (ref) presents simulation results on the finite-sample properties of estimators based on well-specified and misspecified likelihoods. Section (ref) presents an illustration using real-world data. Proofs of the main results are gathered in an Appendix.

The following notational conventions are used throughout the paper. For an infinite sequence $(V_{j})_{j}$, $V_{a}^{b}=(V_{a},\ldots,V_{b})$ for any $a\leq b$; $\mathcal{P}(\mathbb{V})$ denotes the set of Borel probability measures on a Polish space $\mathbb{V}$; for a probability measure $P$, $E_{P}(\cdot)$ denotes expectation with respect to $P$, $o_{P}(\cdot)$ and $O_{P}(\cdot)$ indicate order in probability under $P$, $\Rightarrow_{P}$ signifies weak convergence under $P$, and $L^{r}(P)$, $1\leq r<\infty$, denotes the class of measurable functions integrable to order $r$ with respect to $P$; $\nabla_{\vartheta}$ and $\nabla_{\vartheta}^{2}$ are the gradient and Hessian operators, respectively, with respect to $\vartheta$; $\left\Vert \cdot\right\Vert $ denotes the Euclidean norm of a vector or matrix; $1\{\cdot\}$ denotes the indicator function; $\mathbb{N}$ denotes the set of positive integers. Unless stated otherwise, limits are taken as the sample size, $T$, diverges to infinity.

Model and Estimation

Statistical Model

Let $(X_{t},S_{t})_{t=0}^{\infty}$ be a discrete-time stochastic process such that, for each $t\in\mathbb{N}$, $S_{t}\in\mathbb{S}\equiv\{s_{1} ,\ldots,s_{|\mathbb{S}|}\}\subset\mathbb{R}$ is the unobservable state and $X_{t}\in\mathbb{X}\subseteq\mathbb{R}^{h}$, for some $h\in\mathbb{N}$, is the observable state. Moreover, for each $t\in\mathbb{N}$, the conditional distribution of $X_{t}$, given $X_{0}^{t-1}$ and $S_{0}^{t}$, depends only on $X_{t-1}$ and $S_{t}$, and the conditional distribution of $S_{t}$, given $X_{0}^{t-1}$ and $S_{0}^{t-1}$, depends only on $X_{t-1}$ and $S_{t-1}$, so that

align*[align* omitted — 161 chars of source]

with $(x,s)\mapsto P_{\ast}(x,s,\cdot)\in\mathcal{P}(\mathbb{X})$ and $(x,s)\mapsto Q_{\ast}(x,s,\cdot)\in\mathcal{P}(\mathbb{S})$ denoting the true transition probabilities. It is further assumed that, for each $(x,s)\in \mathbb{X}\times\mathbb{S}$, $P_{\ast}(x,s,\cdot)$ admits a density $p_{\ast }(x,s,\cdot)$ with respect to some $\sigma$-finite measure on $\mathbb{X}$. Our framework imposes no additional restrictions on this measure; for instance, it can be the Lebesgue measure (i.e., allow for continuous $X_{t}$) or a counting measure (i.e., allow for discrete $X_{t}$).

The researcher's model is given by a family of transition probabilities $(x,s)\mapsto P_{\theta}(x,s,\cdot)\in\mathcal{P}(\mathbb{X})$ and $(x,s)\mapsto Q_{\theta}(x,s,\cdot)\in\mathcal{P}(\mathbb{S})$ indexed by an (unknown) parameter $\theta\in\Theta\subseteq\mathbb{R}^{q}$, for some $q\in\mathbb{N}$, such that, for each $\theta\in\Theta$, \[ X_{t}\mid(X_{0}^{t-1},S_{0}^{t})\sim P_{\theta}(X_{t-1},S_{t},\cdot), \] \[ S_{t}\mid(X_{0}^{t-1},S_{0}^{t-1})\sim Q_{\theta}(X_{t-1},S_{t-1},\cdot), \] and, for each $(x,s)\in\mathbb{X}\times\mathbb{S}$, $P_{\theta}(x,s,\cdot)$ admits a density $p_{\theta}(x,s,\cdot)$ with respect to the same measure used to define $p_{\ast}(x,s,\cdot)$.

A few remarks about this setup are worth making. First, and perhaps most importantly, the unobservable states (regimes) are a Markov chain whose transition kernel can depend on the lagged value of the observable state. This is the main departure from prior literature which, with the exception of aill13, has focused on the case where the conditional distribution of $S_{t}$, given $X_{0}^{t-1}$ and $S_{0}^{t-1}$, depends only on $S_{t-1}$. Second, the model $\{(P_{\theta},Q_{\theta})\colon\theta\in\Theta\}$ is allowed to be misspecified in the sense that $(P_{\ast},Q_{\ast} )\notin\{(P_{\theta},Q_{\theta})\colon\theta\in\Theta\}$. This setup encompasses a rich family of models that arise in econometric and statistical applications, some examples of which are given below. Note that, although the family $\{Q_{\theta}\colon\theta\in\Theta\}$ may or may not contain $Q_{\ast} $, it is defined on the same finite state space $\mathbb{S}$ as $Q_{\ast}$, so misspecification of the number of unobservable regimes is ruled out.\footnote{Such misspecification would introduce additional complexities (e.g., locally non-quadratic likelihood surfaces, non-identifiable parameters) which are beyond the scope of this paper.} Finally, even though the conditional distribution of $S_{t}$, given $X_{0}^{t-1}$ and $S_{0}^{t-1}$, is assumed to depend only on $X_{t-1}$ and $S_{t-1}$, it is straightforward to extend all our results to situations where this distribution is dependent on additional higher-order lags of $X_{t}$ and/or $S_{t}$.

example[Hidden Markov Model with Covariate-Dependent Transition Probabilities] Let $x=(y,z)\in\mathbb{X}=\mathbb{R}^{2}$ and $\mathbb{S} =\{0,1\}$. Let $P_{\ast}$ be determined by the equations \begin{align*} Y_{t} & =\mu^{\ast}(S_{t})+\sigma^{\ast}(S_{t})U_{1,t},\\ Z_{t} & =\mu_{2}^{\ast}+\psi^{\ast}Z_{t-1}+\sigma_{2}^{\ast}U_{2,t}, \end{align*} where $(U_{1,t},U_{2,t})_{t}$ are i.i.d. (independent of $(S_{t})_{t}$) with zero mean and covariance matrix indexed by a parameter $\rho^{\ast}$, e.g., $\left[ \begin{array} [c]{cc} 1 & \rho^{\ast}\\ \rho^{\ast} & 1 \end{array} \right] $. The transition probabilities of $(S_{t})_{t}$ are allowed to depend on $Z_{t-1}$; for instance, $(s,z)\mapsto Q_{\ast}(z,s,s)\equiv \Pr(S_{t}=s\mid Z_{t-1}=z,S_{t-1}=s)=[1+\exp(-\alpha_{s}^{\ast}-\beta _{s}^{\ast}z)]^{-1}$ for $s\in\mathbb{S}$. The homogeneous specification with $\beta_{0}^{\ast}=\beta_{1}^{\ast}=0$ has been used to model regime shifts in a variety of economic and financial time series, including output growth (AlbertChib93, rivas15), foreign exchange rates (engelhamilton90, bollen0), and equity returns (ryden98, ang02). The homogeneity restriction is relaxed in dieb94, Engel96, and ang02, among others, to allow the transition probabilities to depend on $Z_{t-1}$. The results obtained here establish the asymptotic properties of the ML estimator in this class of models. $\triangle$
example[Markov-Switching Autoregressive Model with Covariate-Dependent Transition Probabilities] A useful generalization of the previous example is one where the outcome equation is extended to \[ Y_{t}=\mu^{\ast}(S_{t})+\phi^{\ast}Y_{t-1}+\sigma^{\ast}(S_{t})U_{1,t}. \] Variations of the model with $\beta_{0}^{\ast}=\beta_{1}^{\ast}=0$ have found widespread application in economics (e.g., Hansen92, McCulloghTsay94, Murcia95, angbekwei08) and beyond. Generalizations of the model without the restriction of time-invariant transition probabilities are also very popular and have been used, for example, in the modeling of output growth (rivas15), interest rates (angbek02j), consumption growth (Whitelaw00), and bond spreads (pses21). The results obtained here establish the asymptotic properties of the ML estimator in this class of models. $\triangle$
example[Mixture Autoregressive Model]Let $x\in\mathbb{X}=\mathbb{R}$, $\mathbb{S}=\{0,1\}$, and for each $t\in\mathbb{N}$, $\Pr(S_{t}=0\mid X_{t-1})=G_{\ast}(X_{t-1})$ for some $x\mapsto G_{\ast}(x)\in\lbrack0,1]$ and $X_{t}\sim P_{\ast}(X_{t-1},S_{t},\cdot)$. This specification implies a conditional density for $X_{t}$, given $X_{t-1}$, that is a mixture of the type \[ x\mapsto p_{\ast}(x\mid x_{t-1})=G_{\ast}(x_{t-1})p_{\ast}(x_{t-1} ,0,x)+(1-G_{\ast}(x_{t-1}))p_{\ast}(x_{t-1},1,x). \] The covariate-dependence of the transition functions is reflected by the fact that $G_{\ast}$ depends on $X_{t-1}$. Models which belong to the general class of mixture autoregressive models (e.g., dss07, Tadjuidje09, dsps11, saik15) are covered by this framework. $\triangle$

The examples above, as well as those that follow, illustrate that in many areas of application the stochastic process $(X_{t},S_{t})_{t=0}^{\infty}$ is typically highly complex and it is natural/desirable to allow for feedback from past realizations of the observable process $(X_{t})_{t=0}^{\infty}$ to the law of the unobservable regime sequence $(S_{t})_{t=0}^{\infty}$; a tractable way for modeling such feedback is to allow the transition kernel $Q_{\theta}$ to depend on $X_{t-1}$. This feature adds an additional level of complexity and with it sources of potential misspecification. For instance, a common assumption in most applications is that the transition probabilities of hidden regimes are time-invariant, an assumption which may result in misspecification of the transitions functions. In applications that allow for covariate-dependent transition functions, a somewhat more subtle and often overlooked source of misspecification, associated with endogeneity of the transition-driving covariates, may come into play. Specifically, in the context of models such as those in Examples (ref) and (ref), contemporaneous correlation between $U_{1,t}$ and $U_{2,t}$ is typically ignored and inference is based on the likelihood implied by the outcome equation alone; we consider this case in more detail in Sections (ref) and (ref).

Before discussing estimation of $\theta$, we give a result regarding the mixing and ergodicity properties of $(X_{t})_{t=0}^{\infty}$. To do so, let $\bar{P}_{\ast}^{\kappa}$ denote the true distribution over $(X_{t} )_{t=0}^{\infty}$ when the distribution of $(X_{0},S_{0})$ is $\kappa$. Under the following assumptions, Lemma (ref) below ensures that there exists a Borel probability measure on $\mathbb{X}\times\mathbb{S}$, denoted henceforth by $\nu$, for which $(X_{t})_{t=0}^{\infty}$ is stationary and ergodic.

assumptionThere exists a continuous function $\underline{q} :\mathbb{X}\rightarrow\mathbb{R}_{+}\setminus\{0\}$ such that, for all $Q\in\{Q_{\theta}\colon\theta\in\Theta\}\cup Q_{\ast}$, $Q(x,s,s^{\prime} )\geq\underline{q}(x)$ for all $(s^{\prime},s,x)\in\mathbb{S}^{2} \times\mathbb{X}$.
assumptionThere exist constants $\lambda^{\prime}\in(0,1)$, $\gamma\in(0,1)$, $b^{\prime}>0$ and $R>2b^{\prime}/(1-\gamma)$, a lower semi-continuous function $\mathcal{U}:\mathbb{X}\rightarrow\lbrack1,\infty)$, and a measure $\varpi\in\mathcal{P}(\mathbb{X})$ such that, for all $s\in\mathbb{S}$: (i)$~\int_{\mathbb{X}}\mathcal{U}(x^{\prime})P_{\ast }(x,s,dx^{\prime})\leq\gamma\mathcal{U}(x)+b^{\prime}1\{x\in A\}$, with $A\equiv\{x\in\mathbb{X}\colon\mathcal{U}(x)\leq R\}$; (ii)$~A$ is bounded and $\varpi(A)>0$; (iii)$~\inf_{x\in A}P_{\ast}(x,s,C)\geq\lambda^{\prime} \varpi(C)$ for any Borel set $C\subseteq\mathbb{X}$.

The following lemma establishes stationarity, ergodicity, and $\beta$-mixing of $(X_{t})_{t=0}^{\infty}$.

lemmaSuppose Assumptions (ref) and (ref) hold. Then, there exists a $\nu\in\mathcal{P}(\mathbb{X}\times\mathbb{S})$ such that, under $\bar{P}_{\ast}^{\nu}$, $(X_{t})_{t=0}^{\infty}$ is stationary, ergodic, and $\beta$-mixing with mixing coefficients $\beta _{n}=O(\gamma^{n})$, $n\in\mathbb{N}$.
proofSee Supplemental Material (ref).

The result follows in a standard manner by using Assumptions (ref) and (ref) to establish that the implied transition kernel of the joint process $(X_{t},S_{t})_{t=0}^{\infty}$ has a unique invariant distribution and also that it is Harris recurrent and aperiodic. This fact, in turn, is used to show that $(X_{t})_{t=0}^{\infty}$ is stationary, ergodic, and $\beta$-mixing at a geometric rate.

\medspace

remark[Discussion of Assumptions (ref) and (ref) ] Assumption (ref) is an extension of a common assumption in the literature (cf. douc04, aill13) to the case where the transition kernel of $(S_{t})_{t=0}^{\infty}$ depends on $X_{t-1}$. Allowing the lower bound $\underline{q}$ to depend on $x$ is especially relevant when the support of $X_{t}$ is unbounded because, while $\underline{q}(x)>0$, it is allowed to converge to zero as $\left\Vert x\right\Vert \rightarrow\infty$. Although this assumption is not innocuous, we view it as mild because it accommodates the typical specifications used in the literature, where $Q$ is parameterized by a standard Gaussian or logistic cumulative distribution function and a single index $x^{\intercal}\beta$, with $\beta$ restricted to take values in a bounded subset of a finite-dimensional Euclidean space. Assumption (ref)(iii) is an analogous condition for the transition kernel $P_{\ast}$. By inspection of the proof of Lemma (ref), it is easy to see that it suffices to obtain a minorization condition for the \textquotedblleft joint\textquotedblright\ kernel, i.e., $\inf_{x\in A} P_{\ast}(x,s^{\prime},C)Q(x,s,s^{\prime})\geq\lambda\widetilde{\varpi }(C,s^{\prime})$ for any Borel set $C\subseteq\mathbb{X}$ and for some $\widetilde{\varpi}\in\mathcal{P}(\mathbb{X}\times\mathbb{S})$ and $\lambda \in(0,1)$. Thus, Assumptions (ref)(i) and (ref)(i) could be relaxed; e.g., the former could be relaxed to $Q(x,s,s^{\prime} )\geq\underline{q}(x)\varrho(s^{\prime})$, where $\varrho\in\mathcal{P} (\mathbb{S})$, or the latter could be relaxed to $\inf_{x\in A}P_{\ast }(x,s^{\prime},C)\geq\lambda^{\prime}\widetilde{\varpi}(C,s^{\prime})$, where $\widetilde{\varpi}\in\mathcal{P}(\mathbb{X}\times\mathbb{S})$. Assumption (ref)(i),(ii) is a so-called Foster--Lyapunov drift condition; see meyn2012markov and references therein for a discussion of the assumption. $\triangle$

\medspace

In view of Lemma (ref), under $\nu$, the process $(X_{t} )_{t=0}^{\infty}$ can be extended to a two-sided sequence $(X_{t})_{t=-\infty }^{\infty}$. With a slight abuse of notation, we still use $\bar{P}_{\ast }^{\nu}$ to denote the true probability distribution over $(X_{t} ,S_{t})_{t=-\infty}^{\infty}$; $\bar{P}_{\theta}^{\nu}$ is defined analogously for the model $(Q_{\theta},p_{\theta},\nu)$.\footnote{Throughout the text, we use $\bar{P}_{\theta}^{\nu}$ to denote any marginal or conditional probabilities associated with $\bar{P}_{\theta}^{\nu}$; the same holds for $\bar{P}_{\ast}^{\nu}$.}

Parameter Estimation

For any $T\in\mathbb{N}$, let $\ell_{T}^{\nu}:\mathbb{X}^{T+1}\times \Theta\rightarrow\mathbb{R}$ be the sample criterion function given by

equation[equation omitted — 130 chars of source]

where $p_{t}^{\nu}(X_{t}\mid X_{0}^{t-1},\theta)$ denotes the conditional density of $X_{t}$ given $X_{0}^{t-1}$ for any $\theta\in\Theta$; the latter is defined recursively as follows: for any $t\geq1,$ \[ p_{t}^{\nu}(X_{t}\mid X_{0}^{t-1},\theta)=\sum_{s^{\prime}\in\mathbb{S}} \sum_{s\in\mathbb{S}}p_{\theta}(X_{t-1},s^{\prime},X_{t})Q_{\theta} (X_{t-1},s,s^{\prime})\delta_{t}^{\theta,\nu}(s), \] and $s\mapsto\delta_{t}^{\theta,\nu}(s)\equiv\bar{P}_{\theta}^{\nu} (S_{t-1}=s\mid X_{0}^{t-1})$. For each $t\geq2$ and any $s\in\mathbb{S}$, $s\mapsto\delta_{t}^{\theta,\nu}(s)$ satisfies the recursion \[ \delta_{t}^{\theta,\nu}(s)=\sum_{\tilde{s}\in\mathbb{S}}\frac{Q_{\theta }(X_{t-1},\tilde{s},s)p_{\theta}(X_{t-2},\tilde{s},X_{t-1})\delta _{t-1}^{\theta,\nu}(\tilde{s})}{\sum_{s^{\prime}\in\mathbb{S}}p_{\theta }(X_{t-2},s^{\prime},X_{t-1})\delta_{t-1}^{\theta,\nu}(s^{\prime})}, \] with $s\mapsto\delta_{1}^{\theta,\nu}(s)=\sum_{\tilde{s}\in\mathbb{S} }Q_{\theta}(X_{0},\tilde{s},s)\nu(\tilde{s}|X_{0})$, where $\nu(\cdot|\cdot)$ is the conditional density corresponding to $\nu$.

For a given initial distribution $\kappa\in\mathcal{P}(\mathbb{X} \times\mathbb{S})$ over $(X_{0},S_{0})$, we define our estimator as $\hat{\theta}_{\kappa,T}$, where

equation[equation omitted — 152 chars of source]

for some $\eta_{T}\geq0$ and $\eta_{T}=o(1)$.

Consistency

The main result of this section establishes convergence of the estimator $\hat{\theta}_{\nu,T}$ to the set of points in $\Theta$ that are closest to the true model under the Kullback--Leibler information criterion.

Let $H^{\ast}:\Theta\rightarrow\mathbb{R}_{+}\cup\{\infty\}$ be the Kullback--Leibler information criterion $\theta\mapsto H^{\ast}(\theta)$, which is given by \[ H^{\ast}(\theta)=E_{\bar{P}_{\ast}^{\nu}}\left[ \log\frac{p_{\ast}^{\nu }(X_{0}\mid X_{-\infty}^{-1})}{p^{\nu}(X_{0}\mid X_{-\infty}^{-1},\theta )}\right] , \] where, for any $\theta\in\Theta$, $p^{\nu}(X_{0}\mid X_{-\infty}^{-1},\theta)$ denotes the conditional density of $X_{0}$ given $X_{-\infty}^{-1}$ induced by $(P_{\theta},Q_{\theta},\nu)$, and $p_{\ast}^{\nu}(X_{0}\mid X_{-\infty} ^{-1})$ is its counterpart induced by the true transition kernels $(P_{\ast },Q_{\ast},\nu)$; we refer the reader to the Supplemental Material (ref) for details about the construction of these objects and their properties.

Our results allow for misspecified models, and thus $p_{\ast}^{\nu} \notin\{p^{\nu}(\cdot\mid\cdot,\theta):\theta\in\Theta\}$. Hence, as in white82, the relevant limiting set for our estimator is \[ \Theta_{\ast}=\underset{\theta\in\Theta}{\arg\min}H^{\ast}(\theta), \] which is the pseudo-true parameter (set) that minimizes the Kullback--Leibler information criterion. Under Assumption (ref) below, $\Theta_{\ast}$ is non-empty and compact by the Weierstrass Theorem.

assumption(i) $\Theta$ is compact; (ii) $H^{\ast}$ exists and is lower semi-continuous.

The following additional assumptions are used to establish the main result. To state these, for any $\delta>0$ and any $\dot{\theta}\in\Theta$, let $B(\delta,\dot{\theta})\equiv\{\theta\in\Theta\colon||\dot{\theta} -\theta||<\delta\}$.

assumption(i) For any $\epsilon>0$, there exists some $\delta>0$ such that \[ \max_{\dot{\theta}\in\Theta}E_{\bar{P}_{\ast}^{\nu}}\left[ \sup_{\theta\in B(\delta,\dot{\theta})}\frac{p^{\nu}(X_{0}\mid X_{-\infty}^{-1},\theta )}{p^{\nu}(X_{0}\mid X_{-\infty}^{-1},\dot{\theta})}\right] \leq1+\epsilon; \] (ii) there exists a function $(x,x^{\prime})\mapsto C(x,x^{\prime} )\in\mathbb{R}_{+}$ such that $\sup_{\theta\in\Theta}\frac{\max_{s\in \mathbb{S}}p_{\theta}(X,s,X^{\prime})}{\min_{s\in\mathbb{S}}p_{\theta }(X,s,X^{\prime})}\leq C(X,X^{\prime})$ and $\frac{\max_{s\in\mathbb{S} }p_{\ast}(X,s,X^{\prime})}{\min_{s\in\mathbb{S}}p_{\ast}(X,s,X^{\prime})}\leq C(X,X^{\prime})$ a.s.-$\bar{P}_{\ast}^{\nu}$.
assumption$T^{-1}\sum_{t=1}^{T}\max\{1,C(X_{t-1},X_{t} )\}\prod_{i=0}^{t-1}(1-\underline{q}(X_{i}))=o_{\bar{P}_{\nu}^{\ast}}(1).$
remark[Discussion of Assumptions (ref), (ref) and (ref)]Assumption (ref)(i) is standard. Assumption (ref)(ii) is high-level but can be obtained from lower-level conditions (e.g., by following the reasoning in Proposition 11 of douc12). Assumption (ref)(i) is a high-level condition used for establishing uniform law of large numbers results (see Lemma (ref) in Section (ref)). Assumption (ref)(ii) is akin to Assumption A4 in bickel98; it essentially restricts the support of $p_{\theta}$ and $p_{\ast}$ for different values of the state variable.\footnote{The second part of Assumption (ref)(ii) is used to show that $p_{\ast}^{\nu }(\cdot|X_{-\infty}^{-1})$ integrates to one.} Assumption (ref) essentially requires that $\underline{q}$ is not \textquotedblleft too close\textquotedblright\ to zero on average and that $C(X_{t-1},X_{t})$ is finite a.s.-$\bar{P}_{\nu}^{\ast}$. For instance, if $\underline{q}(x)\geq c$ for some $c>0$, then Assumption (ref) is automatically satisfied provided $E_{\bar{P}_{\nu}^{\ast}}[C(X_{0} ,X_{1})]<\infty$. Moreover, by exploiting the fact that $(X_{t})_{t}$ is $\beta$-mixing (see Lemma (ref)), Lemma (ref) in the Supplemental Material (ref) provides sufficient conditions of the form $E_{\bar{P}_{\nu}^{\ast}}\left[ \underline{q} (X_{1})\right] >0$ and $E_{\bar{P}_{\nu}^{\ast}}\left[ C(X_{1},X_{0} )^{l}\right] <\infty$ for some $l>1$.\footnote{We thank a referee and the editor for suggestions on how to weaken Assumption (ref), which lead to these sufficient conditions for Assumption (ref).} $\triangle$

We now establish consistency of the estimator defined by ((ref)).

theoremSuppose Assumptions (ref)--(ref) hold. Then, $d_{\Theta}(\hat{\theta}_{\nu,T},\Theta_{\ast})=o_{\bar{P}_{\ast }^{\nu}}(1).$\footnote{For any set $A\subseteq\Theta$, $d_{\Theta} (\theta,A)\equiv\inf_{\dot{\theta}\in A}||\theta-\dot{\theta}||$.}
proofSee Appendix (ref).

This result is analogous to Theorem 2 in douc12 but for a somewhat different setup; specifically, we allow for autoregressive dynamics and covariate-dependent transition probabilities, but restrict $\mathbb{S}$ to be finite.\footnote{Finiteness of $\mathbb{S}$ simplifies the proofs as we do not have to be concerned with uniformity issues of certain quantities as functions of the hidden state. By combining techniques available in the literature (e.g., douc12) -- which essentially amount to imposing requirements like compactness of $\mathbb{S}$ and continuity -- with ours, we conjecture that our results can be extended to the case where $\mathbb{S}$ is a compact set.}

Clearly, if the model is correctly specified and point identified, i.e., there exists a $\theta\in\Theta$ such that $(P_{\ast},Q_{\ast})=(P_{\theta },Q_{\theta})$, then $\Theta_{\ast}=\{\theta\}$ and our estimator converges in $\bar{P}_{\ast}^{\nu}$-probability to this point. If, however, the model is misspecified, our estimator converges to the set of parameters that is closest to the true set, when closeness is measured by means of the Kullback--Leibler information criterion (cf. white82, douc12).

To prove Theorem (ref), we first show that $T^{-1}\sum _{t=1}^{T}\log p_{t}^{\nu}(X_{t}\mid X_{0}^{t-1},\theta)$ is well-approximated by $T^{-1}\sum_{t=1}^{T}\log p^{\nu}(X_{t}\mid X_{-\infty}^{t-1},\theta)$ (see Lemma (ref) in Appendix (ref)). Second, relying on ergodicity (Lemma (ref)) and Assumption (ref), we establish a uniform law of large numbers for the latter quantity (see Lemma (ref) in Appendix (ref)). The proof of consistency then follows the standard Wald approach.

The approximation result in the first step relies on \textquotedblleft mixing\textquotedblright\ results for the process $(S_{t})_{t=-\infty} ^{\infty}$, given $(X_{t})_{t=-\infty}^{\infty}$. The following theorem, which might be of independent interest, establishes such a \textquotedblleft mixing\textquotedblright\ result in our setting.

theoremTake any $(j,m)\in\mathbb{N}^{2}$. Suppose that, for any $\theta\in\Theta$, there exist mappings $x\mapsto\varrho(x,\cdot )\in\mathcal{P}(\mathbb{S})$ and $\underline{q}:\mathbb{X}\rightarrow \mathbb{R}_{+}$ such that, for all $(s,s^{\prime})\in\mathbb{S}^{2}$, \begin{equation} Q_{\theta}(X,s,s^{\prime})\geqq(X)\varrho(X,s^{\prime} )\quad a.s.-\bar{P}_{\ast}^{\nu}. \end{equation} Then,\footnote{For any $P,Q\in$ $\mathcal{P}(\mathbb{S})$, $||P-Q||_{1} \equiv\sum_{s\in\mathbb{S}}|P(s)-Q(s)|$.} \[ \max_{(b,c)\in\mathbb{S}^{2}}\left\Vert \bar{P}_{\theta}^{\nu}(S_{j+1} =\cdot|S_{-m}=b,X_{-m}^{j})-\bar{P}_{\theta}^{\nu}(S_{j+1}=\cdot |S_{-m}=c,X_{-m}^{j})\right\Vert _{1}\leq\prod_{n=-m}^{j}(1-\underline{q} (X_{n})). \]
proofSee Appendix (ref).

This result is analogous to results in douc04 (e.g., their Lemma 1 and Corollary 1) but for a more general transition function $Q_{\theta}$ (albeit under the requirement that $\mathbb{S}$ is finite). Specifically, we allow $Q_{\theta}$ to depend on $X$, and do not restrict the lower bound in ((ref)) to be uniform in $s^{\prime}$; these features are, to our knowledge, novel.

The proof of Theorem (ref) relies on bounding the Dobrushin coefficient of the transition kernel $\bar{P}_{\theta}^{\nu}(S_{l+1} =\cdot|S_{l}=\cdot,X_{-m}^{j})$ by $1-\underline{q}(X_{l})$, for each $l\in\{-m,...,j\}$. The case where $\varrho(x,\cdot)$, in condition ((ref)), is uniformly bounded from below has been studied in the literature, and such a bound can be obtained by elementary calculations. For the general case, however, the standard technique cannot be used, and we develop a novel, to our knowledge, approach based on coupling techniques; see Appendix (ref) for details.

remarkRemark (ref) and the fact that Theorem (ref) is established under condition ((ref)), imply that, in regards to consistency, Assumption (ref) could be replaced by the weaker condition ((ref)). This remark, however, does not extend to the LAN results of Section (ref), since we do not know whether Assumption (ref) could be weakened to condition ((ref)) in this case. $\triangle$

We discuss next a canonical example that encompasses many commonly used specifications, such as those in Examples (ref) and (ref). The purpose of this example is to verify our assumptions under primitive and low-level conditions.

example[Canonical Example] We verify the regularity conditions for a model with $\mathbb{S}=\{0,1\}$ and \begin{align*} & X_{t}=\mu(S_{t})+\Phi^{\intercal}X_{t-1}+\Sigma^{1/2}(S_{t})\varepsilon _{t},\\ & S_{t}\sim Q_{\bar{\vartheta}}(X_{t-1},S_{t-1},\cdot), \end{align*} where $(\varepsilon_{t})_{t}\sim\mathrm{i.i.d.}~\mathcal{N}(0,I)$, $I$ being the identity matrix, and $x\mapsto Q_{\bar{\vartheta}}(x,s,s)\equiv\Pr (S_{t}=s\mid X_{t-1}=x,S_{t-1}=s)=\Psi(x^{\intercal}\bar{\vartheta}_{s})$ for $s\in\mathbb{S}$, where $\Psi$ is a full-support, continuous cumulative distribution function (e.g., logistic or normal). It is assumed that the parameter set $\Theta$ is compact and that any $\theta\equiv((\mu (s),\Sigma(s))_{s\in\mathbb{S}},\Phi,\bar{\vartheta})$ in it is such that $\Sigma(\cdot)$ and $\Phi\Phi^{\intercal}$ have eigenvalues uniformly bounded away from zero and infinity. For simplicity, we assume that the true model is indexed by $((\mu_{\ast}(s),\Sigma_{\ast}(s))_{s\in\mathbb{S}},\Phi_{\ast },\bar{\vartheta}_{\ast})$, for which analogous conditions hold. Let $q(x)\equiv\inf_{b}\Psi(x^{\intercal}b)$ and note that $q(x)>0$ for each finite $x$ as $\Psi$ has full support. Thus, Assumption (ref) holds and $E_{\bar{P}_{\nu}^{\ast}}[q(X)]>0$. The conditions of Assumption (ref) follow by the results in douc2004practical; see Lemma (ref) in the Supplemental Material (ref) for the formal argument. In Lemma (ref) in the Supplemental Material (ref), it is shown that \[ C^{-1}\underline{p}(x,y)\leq f_{\mathcal{N}}(\{y-\Phi^{\intercal} x-\mu(s)\}\Sigma^{-1/2}(s))\leq C\overline{p}(x,y), \] for some $C\geq1$, where $f_{\mathcal{N}}$ is the $\mathcal{N}(0,I)$ probability density function, \begin{align*} p(x,y)\equiv & \exp\{-0.5(a_{1}\left\Vert y\right\Vert ^{2} +a_{2}\left\Vert x\right\Vert ^{2}+2a_{3}\left\Vert x\right\Vert \left\Vert y\right\Vert )-a_{4}\left\Vert y\right\Vert -a_{5}\left\Vert x\right\Vert \},\\ \overline{p}(x,y)\equiv & \exp\{-0.5(b_{1}\left\Vert y\right\Vert ^{2} +b_{2}\left\Vert x\right\Vert ^{2}-2b_{3}\left\Vert x\right\Vert \left\Vert y\right\Vert )+b_{4}\left\Vert y\right\Vert +b_{5}\left\Vert x\right\Vert \}, \end{align*} and $a_{i},b_{i}$ $(i=1,\ldots,5)$ are positive constants that are functions of uniform bounds of $\Sigma(\cdot)$, $\mu(\cdot)$ and $\Phi$ over $\Theta$; they are defined in Lemmas (ref) and (ref) in the Supplemental Material (ref). Therefore, by Lemma (ref) in the Supplemental Material (ref), it follows that Assumptions (ref) and (ref) hold with $(x,x^{\prime})\mapsto C(x,x^{\prime})\equiv\overline{p}(x,x^{\prime })/\underline{p}(x,x^{\prime})$, provided that $E_{\bar{P}_{\nu}^{\ast}} [\exp\{-0.5l((b_{1}-a_{1})\left\Vert Y\right\Vert ^{2}+(b_{2}-a_{2})\left\Vert X\right\Vert ^{2}-2(b_{3}+a_{3})\left\Vert X\right\Vert \left\Vert Y\right\Vert )+l(b_{4}+a_{4})\left\Vert y\right\Vert +l(b_{5}+a_{5})\left\Vert X\right\Vert \}]<\infty$ for some $l\geq1$. Lemma (ref) in the Supplemental Material (ref) shows that the latter condition holds under some restrictions on the eigenvalues of $\Sigma(\cdot)$, $\Sigma_{\ast }(\cdot)$, $\Phi\Phi^{\intercal}$, and $\Phi_{\ast}\Phi_{\ast}^{\intercal}$ (see the aforementioned lemma and its subsequent remark for the exact formulation and more details). Finally, under these conditions, Lemma (ref) in the Supplemental Material (ref) shows that Assumption (ref) holds. $\triangle$

Lastly, we note that further characterization and interpretation of the pseudo-true parameter set $\Theta_{\ast}$ is not straightforward in our general setup, and we believe that progress in this direction should be made on a case-by-case basis. While pursuing this program is outside the scope of the present paper, in Section (ref) we present numerical results that attempt to shed some light on the characterization of this set. In addition, Example (ref) in the Supplemental Material (ref) provides analytical results on the characterization of $\Theta_{\ast}$ for a subclass of the models in Example (ref), namely hidden Markov models with covariate-dependent transition probabilities, when the (misspecified) model is a simple mixture.

Asymptotic Distribution Theory

In this section, we establish a LAN property (IbragHas81, leCam2012) for our model and an asymptotic linear representation for our estimator, from which asymptotic normality of the estimator is inferred. For this, we require the following assumptions.

assumption(i) $\Theta_{\ast}=\{\theta_{\ast}\}\subset \mathrm{int}(\Theta)$; (ii) $\theta\mapsto p_{\theta}(X,s,X^{\prime})$ and $\theta\mapsto Q_{\theta}(X,s,s^{\prime})$ are twice continuously differentiable $a.s.$-$\bar{P}_{\ast}^{\nu}$ for all $(s,s^{\prime} )\in\mathbb{S}^{2}$.
assumptionFor some $\delta>0$ and $a\geq1$, and for all $(s^{\prime},s)\in\mathbb{S}^{2}$: (i) \[ E_{\bar{P}_{\ast}^{\nu}}\left[ \sup_{\theta\in B(\delta,\theta_{\ast} )}\left\Vert \nabla_{\theta}\log p_{\theta}(X,s,X^{\prime})\right\Vert ^{2a}\right] <\infty\text{ and}~E_{\bar{P}_{\ast}^{\nu}}\left[ \sup _{\theta\in B(\delta,\theta_{\ast})}\left\Vert \nabla_{\theta}\log Q_{\theta }(X,s,s^{\prime})\right\Vert ^{2a}\right] <\infty; \] (ii) \[ E_{\bar{P}_{\ast}^{\nu}}\left[ \sup_{\theta\in B(\delta,\theta_{\ast} )}\left\Vert \nabla_{\theta}^{2}\log p_{\theta}(X,s,X^{\prime})\right\Vert ^{2a}\right] <\infty\text{ and}~E_{\bar{P}_{\ast}^{\nu}}\left[ \sup _{\theta\in B(\delta,\theta_{\ast})}\left\Vert \nabla_{\theta}^{2}\log Q_{\theta}(X,s,s^{\prime})\right\Vert ^{2a}\right] <\infty. \]
assumption$\sum_{j=0}^{\infty}\left( E_{\bar{P}_{\ast}^{\nu} }\left[ \prod_{i=0}^{j}(1-\underline{q}(X_{i}))^{\frac{2a}{1-a}}\right] \right) ^{p\left( \frac{1-a}{2a}\right) }<\infty$ for some $p\in(0,2/3)$ (and the same $a\geq1$ that appears in Assumption (ref)).
remark[Discussion of Assumptions (ref), (ref) and (ref)]Part (i) of Assumption (ref) is standard in the literature. The restriction that $\Theta_{\ast}$ is a singleton could be relaxed using the ideas of liu2003asymptotics for non-identified ML estimators. This extension, albeit interesting, would present nuances that are beyond the scope of the present paper. Part (ii) of Assumption (ref) is also standard, and so is Assumption (ref) (see bickel98 for a discussion). Finally, Assumption (ref) is a strengthening of Assumption (ref), and is required in order to establish the existence of a random sequence $(\Delta_{t}(\theta_{\ast}))_{t}$ which approximates the \textquotedblleft score\textquotedblright\ function well (in the sense of Lemma (ref) in Appendix (ref)). As was the case with Assumption (ref), Lemma (ref) in the Supplemental Material (ref) provides sufficient conditions of the form $E_{\bar{P}_{\nu}^{\ast}}\left[ (1-\underline{q}(X_{1} ))^{2a/(1-a)}\right] <1$. $\triangle$

The next theorem establishes a LAN-type property for the log-likelihood criterion function defined in ((ref)).

theoremSuppose Assumptions (ref), (ref), (ref), (ref) and (ref) hold. Then, there exists a stationary and ergodic process $(\Delta_{t}(\theta_{\ast} ))_{t}$ in $L^{2}(\bar{P}_{\ast}^{\nu})$, a sequence of negative definite matrices $(\xi_{t}(\theta_{\ast}))_{t}$, and a compact set $K\subseteq\Theta$ that includes $0$, such that, for any $v\in K$, \begin{align*} \ell_{T}^{\nu}(X_{0}^{T},\theta_{\ast}+v)-\ell_{T}^{\nu}(X_{0}^{T} ,\theta_{\ast})= & v^{\intercal}\left( T^{-1}\sum_{t=0}^{T}\Delta _{t}(\theta_{\ast})+o_{\bar{P}_{\ast}^{\nu}}(T^{-1/2})\right) \\ & +(1/2)v^{\intercal}\left( T^{-1}\sum_{t=0}^{T}\xi_{t}(\theta_{\ast })+o_{\bar{P}_{\ast}^{\nu}}(1)\right) v+R_{T}(v), \end{align*} where $v\mapsto R_{T}(v)\in\mathbb{R}$ is such that $\lim_{\delta\rightarrow 0}\bar{P}_{\ast}^{\nu}\left( \sup_{v\in B(\delta,0)}||v||^{-2}R_{T} (v)\geq\epsilon\right) =0$ for any $\epsilon>0$ and any $T\in\mathbb{N}$.
proofSee Appendix (ref).

Theorem (ref) extends the results in bickel96 and bickel98 (see their remark on p. 1620) to a more general setup which allows for covariate-dependent transition probabilities, autoregressive dynamics, and misspecified models. The proof develops along the same lines as theirs. The main difference relates to the way in which one establishes that the \textquotedblleft score\textquotedblright\ $\nabla_{\theta}\ell_{T}^{\nu }(\cdot,\theta_{\ast})$ and the Hessian $\nabla_{\theta}^{2}\ell_{T}^{\nu }(\cdot,\theta_{\ast})$ can be approximated by $T^{-1}\sum_{t=0}^{T}\Delta _{t}(\theta_{\ast})$ and $T^{-1}\sum_{t=0}^{T}\zeta_{t}(\theta_{\ast})$, respectively (see Lemmas (ref) and (ref) in Appendix (ref)). As mentioned earlier, these approximations rely on \textquotedblleft mixing\textquotedblright\ properties of the temporally inhomogeneous hidden Markov chain; see Lemma (ref) in the Supplemental Material (ref).

Theorem (ref) may be used to establish the following asymptotic linear representation for our estimator in terms of $(\Delta_{t}(\theta_{\ast }))_{t=0}^{\infty}$ and $(\xi_{t}(\theta_{\ast}))_{t=0}^{\infty}$.

theoremSuppose Assumptions (ref)--(ref) hold and $\eta_{T}=o(T^{-1})$. Then, \[ \frac{\sqrt{T}(\hat{\theta}_{\nu,T}-\theta_{\ast})}{\sqrt{\mathrm{tr} \{\Sigma_{T}(\theta_{\ast})\}}}=-\{E_{\bar{P}_{\ast}^{\nu}}[\xi_{1} (\theta_{\ast})]+o_{\bar{P}_{\ast}^{\nu}}(1)\}^{-1}T^{-1/2}\sum_{t=0}^{T} \frac{\Delta_{t}(\theta_{\ast})}{\sqrt{\mathrm{tr}\{\Sigma_{T}(\theta_{\ast })\}}}+o_{\bar{P}_{\ast}^{\nu}}\left( 1\right) , \] where $\Sigma_{T}(\theta_{\ast})\equiv T^{-1}E_{\bar{P}_{\ast}^{\nu}} [\{\sum_{t=0}^{T}\Delta_{t}(\theta_{\ast})\}\{\sum_{t=0}^{T}\Delta_{t} (\theta_{\ast})^{\intercal}\}]$.
proofSee Appendix (ref).

Theorem (ref) readily implies that, if $T^{-1/2}\Sigma_{T} (\theta_{\ast})^{-1/2}\sum_{t=0}^{T}\Delta_{t}(\theta_{\ast})\Rightarrow _{\bar{P}_{\ast}^{\nu}}\mathcal{N}(0,I)$, then

equation[equation omitted — 216 chars of source]

(If $\Sigma_{T}(\theta_{\ast})\rightarrow\Sigma(\theta_{\ast})$, $\Sigma _{T}(\theta_{\ast})$ may be replaced by $\Sigma(\theta_{\ast})$ in these statements). This result is akin to results in white82 and shares the same features, i.e., the asymptotic covariance matrix $\Omega_{T}(\theta _{\ast})\equiv(E_{\bar{P}_{\ast}^{\nu}}[\xi_{1}(\theta_{\ast})])^{-1} \Sigma_{T}(\theta_{\ast})(E_{\bar{P}_{\ast}^{\nu}}[\xi_{1}(\theta_{\ast })])^{-1}$ has a \textquotedblleft sandwich\textquotedblright\ form and the familiar Fisher information equality does not necessarily hold (see also white94).

The next theorem complements Theorem (ref) by presenting results on consistent estimation of the asymptotic covariance matrix of $\hat{\theta }_{\nu,T}$, vis., $\Omega_{T}(\theta_{\ast})$, in both correctly specified and potentially misspecified models. In the latter case, the proposed estimator is of the \textquotedblleft heteroskedasticity and autocorrelation consistent\textquotedblright\ type. As in many other results for such estimators (e.g., nw87, white94), consistency requires some conditions over the score covariance process $(\Delta_{t+\tau} (\cdot)\Delta_{t}(\cdot)^{\intercal})_{t=0}^{T}$, $\tau\in\mathbb{N}$, namely continuity and law of large numbers results (see Lemma (ref) in the Supplemental Material (ref) ). In what follows, the modulus of continuity is defined as \[ \ddot{\varpi}(\delta^{\prime})\equiv\max_{t}\left\Vert \sup_{||\theta -\theta_{\ast}||<\delta^{\prime}}\left\Vert \Delta_{t}(\theta)\Delta _{0}(\theta)^{\intercal}-\Delta_{t}(\theta_{\ast})\Delta_{0}(\theta_{\ast })^{\intercal}\right\Vert \right\Vert _{L^{1}(\bar{P}_{\ast}^{\nu})}, \] for any $\delta^{\prime}\in(0,\delta]$, where $\delta$ is as in Assumption (ref).\footnote{Lemma (ref) in the Supplementary Material (ref) shows that the modulus of continuity converges to zero as $\delta$ vanishes.}

theoremSuppose the assumptions of Theorem (ref) hold. \begin{enumerate} • If $\theta_{\ast}$ is such that, for any $t\geq0$ and $T\geq1$, $p_{t}^{\nu}(\cdot\mid X_{t-T}^{t-1};\theta^{\ast})=p_{\ast}^{\nu}(\cdot\mid X_{t-T}^{t-1})$, then \[ ||\Omega_{T}(\theta_{\ast})-\{-H_{T}^{-1}(\hat{\theta}_{\nu,T})\}||=o_{\bar {P}_{\ast}^{\nu}}(1), \] where \[ H_{T}(\theta)\equiv T^{-1}\sum_{t=1}^{T}\nabla_{\theta}^{2}\log p_{t}^{\nu }(X_{t}|X_{0}^{t-1},\theta). \] • If, for any $l\geq0$, $|| E_{\bar{P}_{\ast}^{\nu}}[ \Delta_{l}(\theta_{\ast}) \Delta_{0} (\theta_{\ast})^{\intercal} ] || \leq\bar{\upsilon}(l)$ for some integrable function $l\mapsto\bar{\upsilon}(l)\in\mathbb{R}_{+}$, then \[ ||\Omega_{T}(\theta_{\ast})-H_{T}^{-1}(\hat{\theta}_{\nu,T})J_{T}(\hat{\theta }_{\nu,T})H_{T}^{-1}(\hat{\theta}_{\nu,T})||=o_{\bar{P}_{\ast}^{\nu}}(1), \] \end{enumerate} where \begin{align*} J_{T}(\theta) & \equiv T^{-1}\sum_{t=1}^{T}\nabla_{\theta}\log p_{t}^{\nu }(X_{t}|X_{0}^{t-1},\theta)\nabla_{\theta}\log p_{t}^{\nu}(X_{t}|X_{0} ^{t-1},\theta)^{\intercal}\\ & +\sum_{\tau=1}^{L_{T}}\frac{\omega(\tau,L_{T})}{T-\tau}\sum_{t=1}^{T-\tau }\nabla_{\theta}\log p_{t+\tau}^{\nu}(X_{t+\tau}|X_{0}^{t+\tau-1} ,\theta)\nabla_{\theta}\log p_{t}^{\nu}(X_{t}|X_{0}^{t-1},\theta)^{\intercal }\\ & +\sum_{\tau=1}^{L_{T}}\frac{\omega(\tau,L_{T})}{T-\tau}\sum_{t=1}^{T-\tau }\nabla_{\theta}\log p_{t}^{\nu}(X_{t}|X_{0}^{t-1},\theta)\nabla_{\theta}\log p_{t+\tau}^{\nu}(X_{t+\tau}|X_{0}^{t+\tau-1},\theta)^{\intercal}, \end{align*} $\omega(\cdot,\cdot)$ are bounded real weights with $\lim_{T\rightarrow\infty }\omega(\tau,L_{T})=1$ for all $\tau\geq1$, and $(L_{T})_{T=1}^{\infty }\subseteq\mathbb{N}$ is such that $L_{T}(\ddot{\varpi}(T^{-1/2}\log\log T)\log\log T+r_{T}+T^{-1/2})=o(1)$, with $(r_{T})_{T}$ being a positive sequence converging to zero.
proofSee Supplemental Material (ref).

Part (a) of Theorem (ref) deals with correctly specified models and makes use of the Fisher information equality. Part (b) provides a consistent covariance estimator in the general case of potentially misspecified models. The conditions for the weights $\omega(\tau,L_{T})$ and the tuning parameters $(L_{T})_{T}$ are standard for estimators of this type.\footnote{The additional condition that $\omega(\tau,L_{T})=\sum _{j=1+\tau}^{L_{T}}c(j,L_{T})c(j-\tau,L_{T})$ for some constants $c(1,L_{T}),\ldots,c(L_{T},L_{T})$ guarantees that $J_{T}(\hat{\theta}_{\nu ,T})$ is positive semidefinite. Weights obtained from the commonly used Bartlett, Parzen and quadratic-spectral kernels satisfy this condition.} The terms $\ddot{\varpi}(T^{-1/2}\log\log T)\log\log T$ and $r_{T}$ in the growth condition for $L_{T}$ are analogous to those appearing in nw87 and arise as \textquotedblleft costs\textquotedblright\ of working with $\nabla_{\theta}\log p_{t}^{\nu}(X_{t}|X_{0}^{t-1},\hat{\theta}_{\nu,T})$ as opposed to $\nabla_{\theta}\log p_{t}^{\nu}(X_{t}|X_{0}^{t-1},\theta_{\ast})$, and of working with sample averages as opposed to population means, respectively.\footnote{The $\log\log T$ factor is used in order to avoid working with constants.} On the other hand, the $T^{-1/2}$ term does not appear in nw87 and arises because we need to approximate $\Delta_{t}(\theta_{\ast})$, that depends on $X_{-\infty}^{t}$, with the score, that depends on $X_{0}^{t}$.

Theorems (ref) and (ref), together with an asymptotic normality result such as ((ref)), provide the means for constructing asymptotically correct confidence sets and hypotheses tests for $\theta_{\ast}$. In correctly specified models, $(\Delta_{t}(\theta_{\ast }))_{t=0}^{\infty}$ is a martingale difference sequence, and thus result ((ref)) can be obtained by invoking a martingale central limit theorem. In potentially misspecified models, $(\Delta_{t}(\theta_{\ast }))_{t=0}^{\infty}$ will not, in general, be a martingale difference sequence, so one should use a different approach; in some situations, a central limit theorem for $\beta$-mixing sequences can be used instead. The following corollary formalizes this discussion.

corollarySuppose the assumptions of Theorem (ref) hold. \begin{enumerate} • If $\theta_{\ast}$ is such that, for any $t\geq0$ and $T\geq1$, $p_{t}^{\nu}(\cdot\mid X_{t-T}^{t-1};\theta^{\ast})=p_{\ast}^{\nu}(\cdot\mid X_{t-T}^{t-1})$, then \[ \sqrt{T}\{-H_{T}(\hat{\theta}_{\nu,T})\}^{1/2}(\hat{\theta}_{\nu,T} -\theta_{\ast})\Rightarrow_{\bar{P}_{\ast}^{\nu}}\mathcal{N}(0,I). \] • If there exists $\bar{L}>0$ such that, for any $k,T>\bar{L}$, $p_{k}^{\nu}(X_{k}\mid X_{k-T}^{k-1};\theta^{\ast})=p_{k}^{\nu}(X_{k}\mid X_{k-\bar{L}}^{k-1};\theta^{\ast})$, $\lim\inf_{T\rightarrow\infty}e_{\min }(\Sigma_{T}(\theta_{\ast}))>0$,\footnote{For any real symmetric matrix $A$, $e_{\min}(A)$ denotes its minimum eigenvalue.} and $E_{\bar{P}_{\ast}^{\nu} }[\left\Vert \Delta_{1}(\theta_{\ast})\right\Vert ^{4+4\delta}]<\infty$ for some $\delta>0$, then \[ \sqrt{T}\hat{\Omega}_{T}(\hat{\theta}_{\nu,T})^{-1/2}(\hat{\theta}_{\nu ,T}-\theta_{\ast})\Rightarrow_{\bar{P}_{\ast}^{\nu}}\mathcal{N}(0,I), \] where $\hat{\Omega}_{T}(\hat{\theta}_{\nu,T})\equiv H_{T}^{-1}(\hat{\theta }_{\nu,T})J_{T}(\hat{\theta}_{\nu,T})H_{T}^{-1}(\hat{\theta}_{\nu,T})$. \end{enumerate}
proofSee Appendix (ref).

Regarding the asymptotic normality result, both parts of the corollary rely on the structure of the process defined by the random variables $\nabla_{\theta }\log p_{k}^{\nu}(X_{k}\mid X_{-\infty}^{k-1};\theta^{\ast})$; this is a martingale difference sequence in part (a) and a geometrically $\beta$-mixing sequence in part (b). The latter result is a consequence of the $\beta$-mixing structure of $(X_{t})_{t=-\infty}^{\infty}$ and of the fact that the conditional density at $\theta_{\ast}$ depends on a finite number of lags. Examples of models for which such properties hold true are mixture models (cf. Example (ref) in the Supplemental Material (ref)) and mixture autoregressive models (cf. Example (ref)).\footnote{At this level of generality, we cannot establish an asymptotic normality result in the general case where the model is misspecified and $\nabla_{\theta}\log p_{k}^{\nu}(X_{k}\mid X_{-\infty }^{k-1};\theta^{\ast})$ depends on the entire $X_{-\infty}^{k-1}$. The reason is that, although the process $X_{-\infty}^{\infty}$ is $\beta$-mixing, there is no guarantee that $\nabla_{\theta}\log p_{k}^{\nu}(X_{k}\mid X_{-\infty }^{k-1};\theta^{\ast})$ inherits these mixing properties, or that it is a martingale difference sequence.}

Regarding the covariance estimators used in parts (a) and (b) of the corollary, these rely on the corresponding parts of Theorem (ref). The result in part (b) is established by exploiting the $\beta$-mixing structure of the score process $(\Delta_{k}(\theta_{\ast }))_{k=-\infty}^{\infty}$ and the finiteness of its $4+4\delta$ moments (under $\bar{P}_{\ast}^{\nu}$) for some $\delta>0$. Lemma (ref) in Appendix (ref) shows that $r_{T}=L_{T}T^{-1/2}$, and, since $\theta\mapsto\log p_{\theta}$ and $\theta\mapsto\log Q_{\theta}$ are smooth (see Assumptions (ref) and (ref)), it follows that $\delta^{\prime}\mapsto\ddot{\varpi}(\delta^{\prime})=C\delta^{\prime}$ for some finite constant $C$. Hence, the growth condition on $(L_{T})_{T}$ translates into $L_{T}T^{-1/4}\log\log T=o(1)$, which is analogous to that in nw87 (apart from the $\log\log T$ factor, which was introduced in Theorem (ref) for convenience).

The next example verifies the assumptions in the context of the models considered in Example (ref).

exampleIn view of the results in Example (ref), we only need to verify Assumptions (ref) --(ref). Part (i) of Assumption (ref) is standard and is directly imposed, while part (ii) follows from the setup of the example. Assumption (ref) follows by the continuity of the derivatives. Finally, Lemma (ref) in the Supplemental Material (ref) implies that Assumption (ref) holds. Thus, Theorem (ref) holds for the class of models considered in Example (ref). In particular, in the correctly specified case, Corollary (ref)(a) guarantees asymptotic normality of the studentized ML estimator of $\theta_{\ast}$, thereby providing the basis for inference. These results are, to our knowledge, new in the context of Markov-switching autoregressive models with covariate-dependent transition probabilities. $\triangle$

Monte Carlo Simulations

The objective in this section is twofold. First, to assess the quality of approximations provided by our asymptotic results by examining the finite-sample properties of the ML estimator and related statistics in a correctly specified Markov-switching autoregressive model with covariate-dependent transition probabilities. Second, to explore the effects of a type of empirically relevant misspecification which involves the use of an incomplete approximation to the likelihood function that ignores potential contemporaneous correlation between the observation variable $(Y_{t})$ and the variable $(Z_{t})$ upon the lagged value of which the transition probabilities depend.

Monte Carlo experiments are based on artificial data $(X_{t}=(Y_{t} ,Z_{t}))_{t}$ generated according to the equations

align[align omitted — 189 chars of source]

for $t\in\mathbb{N}$, with $X_{0}=(0.5,0.4)$, $\mu_{0}=-\mu_{1}=1$, $\phi =0.9$, $\sigma_{0}=\sigma_{1}=1$, $\mu_{2}=0.2$, $\psi=0.8$, $\sigma_{2}=0.6$, and $(U_{1,t},U_{2,t})_{t}\thicksim$ $\mathrm{i.i.d.} \mathcal{N}\left( \left[

array[array omitted — 25 chars of source]

\right] ,\left[

array[array omitted — 40 chars of source]

\right] \right) $ with $\rho\in\{0,0.8\}$. The regimes $(S_{t})_{t}$ are a Markov chain on $\{0,1\}$, independent of $(U_{1,t},U_{2,t})_{t}$, with transition probabilities

equation[equation omitted — 129 chars of source]

for $s\in\{0,1\}$, where $\alpha_{0}=\alpha_{1}=2$ and $\beta_{0}=-\beta _{1}=-0.5$. The model defined by ((ref))--((ref)) is a prototypical Markov-switching autoregressive model with covariate-dependent transition probabilities (cf. Example (ref)). In each of 1000 independent Monte Carlo replications, $100+T$ data points for $(X_{t})_{t}$ are generated, with $T\in\{200,800,1600,3200\}$, and the last $T$ points are used to compute estimates of the parameters of interest. In order to conserve space, only a selection of the results are reported (the full set of results is available upon request).

Correct Specification

In the first set of experiments, we consider estimation of the parameters of the model in ((ref))--((ref)) using the likelihood function based on the conditional distribution of $X_{t}$ given $X_{0}^{t-1}$. Table (ref) reports the deviation of the mean of the finite-sample distributions of the ML estimators of the elements of $\vartheta=(\mu_{0} ,\mu_{1},\phi,\sigma_{0},\sigma_{1},\alpha_{0},\beta_{0},\alpha_{1},\beta _{1})$ from the corresponding true parameter values (bias) when $\rho=0.8$. We also report the ratio of the sampling standard deviation of the estimators to the estimated standard errors (averaged across replications); the latter are computed using the Hessian estimator (cf. Theorem (ref) (a)).\footnote{Results for $\rho=0$ are not reported since they are very similar to those for $\rho=0.8$.} Although the estimators of $\beta_{0}$ and $\beta_{1}$ are somewhat biased when $T=200$, bias is insignificant in the rest of the cases. Estimated standard errors are somewhat downwards biased in most cases but, unless $T$ is small, the bias is not generally substantial and decreases as $T$ increases.

{

table[table omitted — 1,311 chars of source]

}

We also examine conventional hypothesis tests for $\vartheta$. Table (ref) reports the rejection frequencies of: (i) a $t$-type test of $\mathcal{H}_{0}:\vartheta_{j}=\vartheta_{j}^{\ast}$ versus $\mathcal{H} _{1}:\vartheta_{j}\neq\vartheta_{j}^{\ast}$, where $\vartheta_{j}$ is the $j$-th element of $\vartheta$ and $\vartheta_{j}^{\ast}$ is its true value; (ii)$~$a $t$-type test of $\mathcal{H}_{0}:\vartheta_{j}=0$ versus $\mathcal{H}_{1}:\vartheta_{j}\neq0$. These rejection frequencies are referred to as \textquotedblleft size\textquotedblright\ and \textquotedblleft power\textquotedblright, respectively, and are computed using the 0.975 standard-normal quantile as critical value.\footnote{Results should be interpreted with caution in the case of $\mathcal{H}_{0}:\sigma_{i}=0$, $i\in\{0,1\}$, because the null value of $\sigma_{i}$ is on the boundary of the maintained hypothesis. Our asymptotic theory does not allow for parameters that may lie on the boundary of the parameter space.} Tests tend to have Type I error probabilities which are generally close to the nominal 0.05 level, especially for $T>200$. Tests are also powerful enough to reject the hypothesis of a zero parameter value, except in the case of $\beta_{0}$ and $\beta_{1}$ with $T=200$. The distributions of studentized statistics associated with the elements of the ML estimator of $\vartheta$ (ratio of estimation error to corresponding estimated standard error) generally tend to have mean and variance (not shown) that do not differ substantially from zero and one, respectively, and Gaussianity is never rejected for $T>200$ .\footnote{Statistics associated with $\phi$, $\sigma_{0}$ and $\sigma_{1}$ appear to fare somewhat worse than others when $T=200$, a finding similar to that reported in ps98 for models with a time-invariant transition mechanism.}

{

table[table omitted — 1,228 chars of source]

}

Misspecification

In the second set of experiments, we consider estimation of the parameter $\vartheta=(\mu_{0},\mu_{1},\phi,\sigma_{0},\sigma_{1},\alpha_{0},\beta _{0},\alpha_{1},\beta_{1})$ using the partial likelihood function based on the conditional distribution of $Y_{t}$ given $(Y_{t-1},S_{t})$. Inference in models like ((ref))--((ref)) is predominantly based on such a partial likelihood that ignores the equation for $Z_{t}$ (see, e.g., dieb94, filardo94). Formally, the misspecified model may be viewed as defined by equations ((ref))--((ref)), with the additional assumption that $\rho=0$. Under this assumption, estimation of $\vartheta$ is based on ((ref)) alone, the (potentially incorrect) rational behind this approach being that, since the conditional distribution of $Y_{t}$ given $(Y_{t-1},S_{t})$ and the transition probabilities of $(S_{t})_{t}$ depend only on $Z_{t-1}$, ((ref)) may be analyzed, without loss of relevant information, independently of $(Z_{t})_{t}$. Even though such an approach may be appealing because of its relative simplicity, it is far from obvious that it provides a valid way for conducting inference on $\vartheta$; it is unclear, for example, what the limit point of the ML estimator based on the partial likelihood might be when $\rho\neq0$. In earlier sections, we considered a theoretical framework that acknowledges this source of misspecification (among others) and provided tools for asymptotically valid inference. We now quantify the implications of this misspecification in finite samples. For brevity, we refer to the maximizer of the partial likelihood function associated with the conditional distribution of $Y_{t}$ given $(Y_{t-1},S_{t})$ as the `partial ML' estimator to distinguish it from the `joint ML' estimator based on the joint model for the conditional distribution of $X_{t}$ given $X_{0}^{t-1}$ (cf. Section (ref)).

Table (ref) shows the estimated bias of the partial ML estimators of the elements of $\vartheta$ and the ratio of the sampling standard deviation of the estimators to the estimated standard errors (averaged across replications) when $\rho=0.8$. To reflect what is common practice in applied research, standard errors are computed using the Hessian estimator (which relies on the assumption of a correctly specified likelihood) instead of a \textquotedblleft sandwich\textquotedblright\ estimator (which allows for mispecification). It is immediately apparent that the partial ML estimator of most of the parameters is considerably more biased than the joint ML estimator. The differences between the two estimators are more pronounced for parameters associated with the transition probabilities ($\alpha_{0}$, $\beta_{0}$, $\alpha_{1}$, $\beta_{1}$), the partial ML estimators of which are significantly biased even for $T=3200$. This suggests that the bias of the partial ML estimator when $\rho\neq0$ is not associated only with small samples, a finding that is consistent with our asymptotic results. Regarding the accuracy of estimated standard errors, the latter are downwards biased in most cases, the bias being somewhat larger than it is for joint ML estimators. However, unless $T$ is small, this bias is not generally substantial and declines as $T$ increases, despite the fact that standard errors are obtained from the Hessian.\footnote{Results for $\rho=0$ (not shown) are not substantially different from those obtained from the joint ML\ procedure. This is not surprising since the joint and partial ML estimators are both consistent for the true parameter value when $\rho=0$.} {

table[table omitted — 1,320 chars of source]

}

We note that hypothesis tests based on studentized statistics analogous to those considered in Section (ref) (not shown) are unreliable when partial ML estimates are used. This is especially true in the case of parameters associated with the transition probabilities, the corresponding tests being either excessively conservative or excessively liberal. Although tests of this type are extensively used in applied work, they should be interpreted with caution since the statistics on which they are based have an asymptotically normal null distribution only when $\rho=0$. We also note that using the \textquotedblleft sandwich\textquotedblright\ estimator of Theorem (ref)(b) instead of the estimator based on the observed information matrix is not without difficulty when $\rho\neq0$. As Freedman06 points out, the use of such an estimator for inference is unlikely to produce results that are any less misleading under misspecification since the problem of bias/inconsistency of the ML estimator for the true parameter value remains. It is indeed clear from the results in Table (ref) that the bias of the partial ML estimator presents a much more serious problem in our setting than the inaccuracy of conventionally computed standard errors.

Empirical Illustration

In this section, we present an empirical illustration based on a regime-switching model of a type that is commonly used in economics.\footnote{An additional empirical example, which examines the predictive ability of an index of leading indicators for regime changes in U.S. output growth, is discussed in the Supplemental Material (ref).} Specifically, we investigate the potential contribution of the interest rate spread and the growth in tax revenues in predicting regime changes in U.S. real output growth. The model is a variant of the specification used in the simulations and is given by

align[align omitted — 204 chars of source]

for some $h_{1},h_{2}\in\mathbb{N}$, with the hidden, two-state Markov chain $(S_{t})_{t}$ being governed by the transition probabilities

equation[equation omitted — 128 chars of source]

with $s\in\{0,1\}$, and $(U_{1,t},U_{2,t})_{t}$ postulated to be i.i.d. $\mathcal{N}\left( \left[

array[array omitted — 25 chars of source]

\right] ,\left[

array[array omitted — 40 chars of source]

\right] \right) $ and independent of $(S_{t})_{t}$. In (\ref{eqy} )--(\ref{eqtp}), $Y_{t}$ stands for the growth rate of real gross domestic product and $Z_{t}$ is either the spread between the 10-year Treasury note rate and the 3-month Treasury bill rate or the growth rate of real government receipts of direct and indirect taxes.\footnote{The model could be generalized to allow for Markov changes in all the parameters. However, since $Z_{t}$ is thought of here as a potential leading indicator for business-cycle phases, it does not seem sensible to allow the parameters in both (\ref{eqy}) and (\ref{eqz}) to be subject to changes driven by $(S_{t})_{t}$. Modeling regime changes in $(Y_{t})_{t}$ and $(Z_{t})_{t}$ as being driven by two independent Markov processes is more attractive, but we choose to abstract from this as it is not directly related to the main problem under study.} The data are quarterly and span the period 1954:3--2009:2.\footnote{Interest rate data are taken from the FRED$^{\textregistered}$ database; output and tax data are taken from \cite{auerbach12}. The likelihood ratio test of \cite{Hansen92} rejects the hypothesis that $\mu_{0}=\mu_{1}$ in ((ref)).}

Since the aim is not only to assess the predictive ability of the interest rate spread and tax revenues for regime changes in output growth but also to examine whether treating these variables as exogenous yields results which are different from those obtained from a joint model, we compute two sets of estimates: partial ML estimates based on ((ref)) alone and joint ML estimates based on the system ((ref))--((ref)). We note that in econometric models of the business cycle such as ((ref))--((ref)), it is common to rely on partial ML estimation (see, e.g., filardo94). Parameter estimates are reported in Tables (ref) and (ref) , with estimated standard errors given in parentheses; the latter are obtained from the \textquotedblleft sandwich\textquotedblright\ estimator of Theorem (ref)(b).\footnote{The weights $\omega(\cdot,L_{T})$ are obtained from the Parzen kernel and $L_{T}$ is determined using the automatic plug-in method of andrews91. } On the basis of $t$-type tests based on joint ML estimates, at least one of the parameters $(\beta_{0},\beta_{1})$ is significantly different from zero, indicating that the spread and tax revenues contain significant information about the probability of switching between the two regimes.

{ \

table[table omitted — 1,363 chars of source]

}

Regarding the implications of treating $Z_{t}$ as exogenous, the differences between partial and joint ML estimates are substantial in the model with tax revenues (especially for autoregressive coefficients and the parameters associated with the transition probabilities) but much less so in the model with the interest rate spread. This is not entirely unexpected in view of the fact that the estimated value of the conditional correlation $\rho$ is relatively large (0.6034) in the former model but much smaller ($-$0.0849), and insignificantly different from zero, in the latter. Such findings are in line with the analytical and simulation results presented in previous sections. The relatively large estimate of $\rho$ in Table (ref) also suggests that inference based on the partial ML estimator is potentially misleading because of the likely bias of the estimator.

{ \

table[table omitted — 1,277 chars of source]

}