Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
44,048 characters · 8 sections · 45 citation commands
On the Robustness of Mixture Models in the Presence of Hidden Markov Regimes with Covariate-Dependent Transition Probabilities
\baselineskip0.3in
Consistency and asymptotic normality of least-squares estimators in regression models in the presence of potential model misspecification --- e.g., misspecification of the response function or misspecification of the dynamic structure of the errors --- are well-established facts (see, e.g., DomowitzWhite82). Such fundamental results, together with the related classical work of Huber67, underpin a large body of literature exploring the feasibility of drawing valid and meaningful inferences from parametric models that need not necessarily contain the true data-generating process (DGP). Numerous results of this kind have been established for a wide variety of models and estimators, both in static and dynamic settings, ranging from inference procedures based on estimating equations and moment conditions (e.g., Bates85) to quasi-maximum-likelihood (QML) procedures for conditional mean, conditional variance and conditional quantile models (e.g., White82,white94, Levine83, Gourieroux84, NeweySteigerwald97, KOMUNJER2005).
This paper adds to the literature by presenting another example of robustness with respect to misspecification. Specifically, we consider the case of Hidden Markov Models (HMMs), where observable variables exhibit conditional independence given an underlying unobservable regime sequence (and, possibly, exogenous covariate sequences), focusing on situations where the dependence structure of the regime sequence is misspecified. In our set-up, the DGP is taken to be a generalized HMM that may include covariates and has a finite number of Markov regimes, but the postulated probability model is a finite mixture model, that is, an HMM with independent, identically distributed (i.i.d.) regimes. By considering the pseudo-true parameter set for the QML estimator in the (misspecified) mixture model, it is shown that the parameters of the conditional distribution of the observable response variables are consistently estimable even if the dependence of the unobservable regime sequence is not taken into account. A condition on the tail behavior of the characteristic function of the (standardized) conditional distribution of the observable responses is also provided under which the pseudo-true parameter for the QML estimator is a singleton set. An important distinguishing feature of our analysis is that the true regime sequence is allowed to be a temporally inhomogeneous Markov chain whose transition probabilities are functions of observable variables.
This case holds practical significance given the widespread use of both HMMs and mixture models. HMMs with temporally inhomogeneous regime sequences have found applications in diverse areas such as biology (e.g., ghavidel2015nonhomogeneous), economics (e.g., dieb94, Engel96), earth sciences (e.g., Hughes99), and engineering (e.g., ramesh1992modeling). Temporally homogenous variants of HMMs and of Markov-switching regression models are also used extensively in economics and finance (e.g., engelhamilton90, ryden98, Jeanne2000, bollen0), as well as in biology, computing, engineering and statistics (see Ephraim2002 and references therein). Statistical inference in such models is typically likelihood-based and the properties of QML procedures are, naturally, of much interest. Nevertheless, HMMs are inherently intricate and computationally demanding due to the need to account for the underlying correlated regime sequence and for the dependence of the conditional distribution on the current hidden regime. By demonstrating that it is feasible to use a mixture model --- a simpler and computationally less demanding framework --- while still estimating consistently the parameters of the conditional distribution of the observations, this paper offers a more accessible avenue for practitioners to follow without sacrificing the accuracy of parameter estimates.
In related recent work, pouzo22MLE considered the asymptotic properties of the QML estimator in a rich class of models with Markov regimes under general conditions which allow for autoregressive dynamics in the observation sequence, covariate-dependence in the transition probabilities of the hidden regime sequence, and potential model misspecification. The QML estimator was shown to be consistent for the pseudo-true parameter (set) that minimizes the Kullback--Leibler information measure. Unsurprisingly, identifying the possible limit of the QML estimator when the true probability structure of the data does not necessarily lie within the parametric family of distributions specified by the model is not always a feasible task in such a general set-up. This paper provides an answer in the simpler case of switching-regression models, HMMs and related mixture models. Consistency results for misspecified pure HMMs (with no covariates in the outcome equation) can also be found in Mevel04 and douc12. Unlike our analysis, which allows the regime transition probabilities to be time-dependent and driven by observable variables, these papers restrict attention to the case of time-invariant transition mechanisms.
In the next section, we introduce the DGP and statistical model of interest, and consider QML estimation of the parameters of the outcome equation of a misspecified generalized HMM. Section (ref) discusses numerical results from a simulation study. Section (ref) summarizes and concludes.
Consider a discrete-time stochastic process $\{(X_{t},S_{t})\}_{t\geq0}$ such that $X_{t}=(Y_{t},Z_{t},W_{t})$ is an observable variable with values in $\mathbb{X}\subset\mathbb{R}^{3}$ and $S_{t}$ is a latent variable with values in $\mathbb{S}:=\{1,2,\ldots,d\}\subset\mathbb{N}$ for some $d\geq2$. The variable $S_{t}$ is viewed as the hidden regime (or state) associated with index $t$, which is \textquotedblleft observable\textquotedblright\ only indirectly through its effect on $X_{t}$. The following assumptions are made about the DGP:
Instead of the Markov-switching structure of the DGP, the researcher's postulated parametric model is a family of finite mixture models (without Markov dependence). Specifically, the model is specified by assuming that the regime variables $\{S_{t}\}_{t\geq1}$ are i.i.d. with distribution
In addition, the observable variables $\{Y_{t}\}_{t\geq1}$ are assumed to satisfy the equations
where $\mu$, $\gamma$ and $\sigma>0$ are known real functions on $\mathbb{S}$ and $\{\varepsilon_{t}\}_{t\geq1}$ are i.i.d. random variables, independent of $\{(S_{t},W_{t})\}_{t\geq1}$, such that $\varepsilon_{1}$ has the same density $f$ as $U_{1,1}$. The mixture model defined by ((ref)) and ((ref)) is parameterized by $\theta:=(\pi(s),\bar{\vartheta}_{s})_{s\in\mathbb{S}}$, with $\pi(s):=(\mu(s),\gamma(s),\sigma(s))$, which is assumed to take values in a compact set $\Theta\subset\mathbb{R}^{q}$, $q>1$. We denote by $P_{\pi }(\cdot|W_{t},S_{t})$ the conditional distribution of $Y_{t}$ given $(W_{t},S_{t})$ that is implied by ((ref)); the corresponding conditional density is denoted by $p_{\pi}(\cdot|W_{t},S_{t})$.
\paragraph{Key aspects of our set-up:}
First, the DGP has a (generalized) HMM structure in which $\{Y_{t}\}_{t\geq0}$ are independent, conditionally on the regime sequence $\{S_{t}\}_{t\geq0}$ and an exogenous covariate sequence $\{W_{t}\}_{t\geq0}$ (having the Markov property), so that the conditional distribution of $Y_{t}$ given the regime and covariate sequences depends only on $(S_{t},W_{t})$. The inclusion of the exogenous covariate $W_{t}$ in ((ref)) and ((ref)) allows the study of the causal effect of $W$ on $Y$ under different regimes; this causal effect is captured by $\gamma^{\ast}$ and is estimable via the mixture specification ((ref))--((ref)). Exogeneity of $W$ (assumption 4 above) is essential for the results discussed in Section (ref) to hold and, hence, for consistent estimation of the causal effect $\gamma^{\ast}$ under the (erroneous) assumption of independent regimes.\footnote{The standard HMM formulation is a special case in which $W_{t}$ is absent from the outcome equation ((ref)).} Second, the true hidden regimes $\{S_{t}\}_{t\geq0}$ are a temporally inhomogeneous Markov chain whose transition probabilities depend on the lagged value of the observable variable $Z_{t}$. The sequence $\{Z_{t}\}_{t\geq0}$ has the Markov property and is not required to be exogenous, in the sense that $Z_{t}$ may be contemporaneously correlated with $U_{1,t}$. Third, the statistical model is misspecified, in the sense that the DGP is not a member of the family $\{(P_{\pi},Q_{\bar{\vartheta}})\colon (\pi,\bar{\vartheta})\in\Theta\}$; this is because the dynamic structure of the regimes is misspecified. As already discussed in Section (ref), this relatively simple set-up is of much practical interest since HMMs with temporally inhomogeneous regime sequences have found many applications. Mixture models with i.i.d. regimes are also widely used in many different fields (see McLachlan2000 and Schnatter06), including economics and econometrics (see Compiani2016).
It is worth noting that, although we focus on scalar responses and covariates for the sake of simplicity, all our results can be extended straightforwardly to cases where $X_{t}\in\mathbb{X}\subset\mathbb{R}^{h}$ with $h>3$. For example, $W_{t}$ may be a vector of covariates, which may include lagged values of $W_{t}$ in cases where dynamic causal effects are of interest. Similarly, $Z_{t}$ may be a vector of information variables that affect the dynamic profile of the regime transition probabilities, whose generating mechanism has a finite-order autoregressive structure.
Given observations $(X_{1},\ldots,X_{T})$, $T\geq1$, the quasi-log-likelihood function for the parameter $\theta$ is
The QML estimator $\hat{\theta}_{T}$ of $\theta$ is defined as an approximate maximizer of $\ell_{T}(\theta)$ over $\Theta$, so that \[ \ell_{T}(\hat{\theta}_{T})\geq\sup_{\theta\in\Theta}\ell_{T}(\theta)-\eta _{T}, \] for some sequence $\{\eta_{T}\}_{T\geq1}\subset\mathbb{R}_{+}$ converging to zero.
It is not too onerous to verify that, under assumptions that are common in the literature (e.g., Gaussianity of $U_{1,1}$ and $Q_{\ast}(s|z,s^{\prime })=G(\alpha_{s,s^{\prime}}+\beta_{s,s^{\prime}}z)$ for some continuous distribution function $G$ on $\mathbb{R}$ whose support is all of $\mathbb{R} $), the conditions of pouzo22MLE required for convergence of the QML estimator of $\theta$ to a well-defined limit are satisfied. Specifically, let \[ \theta\mapsto H^{\ast}(\theta):=\mathbb{E}_{\bar{P}_{\ast}}\left[ \ln\left( \frac{p_{\ast}(Y_{1}|W_{1})}{p_{\theta}(Y_{1}|W_{1})}\right) \right] \] be the Kullback--Leibler information function, where $p_{\theta}(Y_{1} |W_{1}):=\sum_{s\in\mathbb{S}}\bar{\vartheta}_{s}p_{\pi}(Y_{1}|W_{1},s)$ denotes the conditional density of $Y_{1}$ given $W_{1}$ induced by $(P_{\pi },Q_{\bar{\vartheta}})$ for each $(\pi,\bar{\vartheta})\in\Theta$, $p_{\ast }(Y_{1}|W_{1})$ denotes the conditional density of $Y_{1}$ given $W_{1}$ induced by the (true) DGP, and the expectation $\mathbb{E}_{\bar{P}_{\ast} }(\cdot)$ is with respect to the distribution $\bar{P}_{\ast}$ of $\{(X_{t},S_{t})\}_{t\geq0}$ induced by the (true) DGP. Then, we have
in $\bar{P}_{\ast}$-probability, where
is the pseudo-true parameter (set) and $\left\Vert \cdot\right\Vert $ denotes the Euclidean norm on $\mathbb{R}^{q}$ (cf. Theorem 1 of pouzo22MLE).
A sharper result can be established by considering the pseudo-true parameter $\Theta_{\ast}$ under the specified DGP. Together with ((ref)) and ((ref)), the following theorem shows that, despite the erroneous treatment of hidden regimes as independent, QML based on the (misspecified) mixture model provides consistent estimators of the true parameters of the outcome equation.
Theorem (ref) establishes that the true parameters $\pi^{\ast} (s):=(\mu_{1}^{\ast}(s),\sigma_{1}^{\ast}(s),\gamma^{\ast}(s))$, $s\in\mathbb{S}$, associated with the observation equation ((ref)), together with $(\bar{\vartheta}_{s}^{\ast})_{s\in\mathbb{S}}$, minimize the Kullback--Leibler information $H^{\ast}(\theta)$. The corollary that follows shows that this minimizer is unique, as long as $\theta_{\ast}:=(\pi^{\ast }(s),\bar{\vartheta}_{s}^{\ast})_{s\in\mathbb{S}}$ is identified in $\Theta$. In the present context, $\theta_{\ast}$ is said to be identified in $\Theta$ if, for any $\theta\in\Theta$ such that $p_{\theta}(\cdot|\cdot)=p_{\theta _{\ast}}(\cdot|\cdot)$ $\bar{P}_{\ast}$-almost surely, $\theta=\theta_{\ast}$ up to permutations.\footnote{Formally, $(\pi(\mathfrak{p}[s]),\bar{\vartheta }_{\mathfrak{p}[s]})=(\pi^{\ast}(s),\bar{\vartheta}_{s}^{\ast})$ for all $s\in\mathbb{S}$ and any permutation $\mathfrak{p}:\mathbb{S}\rightarrow \mathbb{S}$. The qualifier `up to permutations' reflects the fact that the problem remains unchanged if the indices of the regimes are permuted.}
To provide some (non-technical) intuition behind Theorem (ref) and its corollary, recall that when the misspecified mixture model is fitted to data using maximum likelihood techniques, the objective function which is maximized is a quasi-likelihood. The resulting QML estimator is an estimator for the parameters which make the model's implied conditional distribution of the response as close as possible --- measured in Kullback--Leibler divergence --- to the true conditional distribution, even when the model is misspecified. Despite the fact that the researcher ignores dependence of the regimes, the conditional distribution of the response variable given the covariates remains a mixture of distributions under both the true and misspecified models. This structural similarity allows the mixture model to match the key features of the conditional distribution of interest correctly and, thus, the QML estimator consistently recovers the parameters of the outcome equation. This result relies heavily on two key conditions: stationarity of the regimes (Assumption 2) and exogeneity of the covariates (Assumption 4). The former stationarity condition is crucial because it ensures stability of the marginal distribution of the regimes, which the mixture model tries to fit. Without stationarity, the limiting object of the QML estimation procedure could vary over time, and consistency would break down. The latter exogeneity condition is fundamental because it guarantees that the covariates do not \textquotedblleft carry\textquotedblright\ information about future states or future shocks into the noise of the outcome equation. If exogeneity failed, the conditional distribution of the response would not be properly captured by the simple mixture model and the QML estimator would be biased and inconsistent for the parameters of interest.
We conclude by remarking that when the minimizer $\theta_{\ast}$ of the Kullback--Leibler information function is identified in $\Theta$ (and belongs to the interior of $\Theta$), asymptotic normality of $\sqrt{T}(\hat{\theta }_{T}-\theta_{\ast})$ may be deduced from the results of pouzo22MLE under suitable differentiability and moment conditions. These conditions are satisfied, for example, in the case where $f$ is Gaussian and $Q_{\ast }(s|z,s^{\prime})=G(\alpha_{s,s^{\prime}}+\beta_{s,s^{\prime}}z)$ for some continuous distribution function $G$ on $\mathbb{R}$ whose support is all of $\mathbb{R}$. In the next subsection, we discuss a sufficient condition for the high-level identifiability requirement of Corollary (ref) and show that this condition holds in some commonly used models.
Identifiability of mixture models has been studied extensively in the literature following the original contribution of Teicher1963, who established identifiability of finite mixtures of distributions such as the one-dimensional Gaussian and gamma. Yakowitz1968 gave a necessary and sufficient condition for identifiability, which holds, for example, in the case of finite mixtures of multivariate Gaussian distributions. This condition was exploited by Holzmann2004 and Holzmann2006 to provide sufficient low-level identifiability conditions based on the tail behavior of the characteristic function of the component distributions.\footnote{A review of related results for parametric and nonparameteric models that incorporate mixture distributions can be found in Compiani2016.}
Using the approach of Holzmann2004 and Holzmann2006, we now present a low-level condition, based only on features of $f$, which is sufficient for the identifiability of $\theta_{\ast}$ required in Corollary (ref).
As an example, consider what is, arguably, the most widely used class of mixture models, namely those in which $f$ is Gaussian. In this case, $\varphi(\tau)=e^{-\tau^{2}/2}$, $\tau\in\mathbb{R}$, and $\varphi(a_{1} \tau)/\varphi(a_{2}\tau)=e^{-(a_{1}^{2}-a_{2}^{2})\tau^{2}/2}$, $a_{1}>a_{2}$, so the condition of Lemma (ref) is satisfied. Thus, in the Gaussian case, the QML estimator of the parameters of the mixture model ((ref))--((ref)) converges, in $\bar{P}_{\ast}$-probability, to $\theta_{\ast}$. This result remains valid for non-Gaussian distributions, including distributions with heavy tails (and finite variance). For instance, the result holds if $f$ is the density of a (rescaled) Student-$t$ distribution with degrees of freedom $\upsilon>2$ (see Example 1 in Holzmann2006).
The consistency results in ((ref))--((ref)) and in Theorem (ref) are quite general, in the sense that they cover misspecified generalized HMMs with temporally inhomogeneous regime sequences and arbitrary observation conditional densities. They imply that dependence of the regimes in such HMMs may be safely ignored as long as the parameters of interest are those of the conditional density of the observations given the regimes and the covariates. It is important to note, however, that care should be taken in estimating the asymptotic covariance matrix of the QML estimator since the inverse of the observed information matrix is not necessarily a consistent estimator in a misspecified model. Consistent estimation in this case typically requires the use of an empirical sandwich estimator that does not rely on the information matrix equality (cf. Theorem 5 of pouzo22MLE).
Treating the regimes as an independent sequence simplifies likelihood-based inference compared to the case of correlated Markov regimes. In the latter case, an added difficulty, as demonstrated by pouzo22MLE, is that consistent QML estimation of the true parameter values in a model with Markov regimes having covariate-dependent transition functions typically requires joint analysis of equations such as ((ref)) and the generating mechanism of $\{Z_{t}\}$, even if the parameters of interest are only those associated with ((ref)). Furthermore, as pointed out by hamilton16, rich parameterizations of the transition mechanism of the regime sequence may not necessarily be desirable when working with relatively short time series because of legitimate concerns relating to potential over-fitting and inaccurate statistical inference. In such cases, parsimonious specifications which provide good approximations to key features of the data --- and, in our setting, consistent estimates of the parameters of interest --- can be attractive and useful.
Note that, for a class of regime-switching models in which the regime sequence $\{S_{t}\}$ is a temporally homogeneous, two-state Markov chain, an observation analogous to that implied by Theorem (ref) was made by chowhite07. They argued that the parameters of a model for the conditional distribution of the observable variable $X_{t}$, given $(X_{0}^{t-1},S_{0}^{t})$, can be consistently estimated by QML based on a misspecified version of the model with i.i.d. regimes --- and exploited this result to construct a quasi-likelihood-ratio test of the null hypothesis of a single regime against the alternative hypothesis of two regimes. However, Carter2012 demonstrated that consistency of the QML estimator for the true parameters in such a setting does not, in fact, hold if the model and the DGP contain an autoregressive component. This observation remains true in our more general set-up with temporally inhomogeneous hidden regime sequences. Specifically, a result analogous to that in Theorem (ref) does not hold when lagged values of $Y_{t}$ are present as covariates in the outcome equations ((ref)) and ((ref)) (e.g., as in Markov-switching autoregressive models). In this case, misspecification of the dependence structure of the regimes will affect estimation of all the parameters, not just those associated with the transition functions of the regime sequence.
As a numerical illustration of the results discussed in Section (ref), we report here findings from a small Monte Carlo simulation study in which the effect on QML estimators of ignoring Markov dependence of hidden regimes is assessed.
In the experiments, artificial data are generated according to the generalized HMM defined by ((ref)), with regimes $\{S_{t}\}$ which form a Markov chain on $\mathbb{S}=\{1,2\}$ such that \[ \Pr(S_{t}=s|S_{t-1}=s,Z_{t-1}=z)=[1+\exp(-\alpha_{s}^{\ast}-\beta_{s}^{\ast }z)]^{-1},\quad s\in\{1,2\},\quad z\in\mathbb{R}, \] and with $\{Z_{t}\}$ and $\{W_{t}\}$ satisfying the autoregressive equations \[ Z_{t}=\mu_{2}^{\ast}+\psi^{\ast}Z_{t-1}+\sigma_{2}^{\ast}U_{2,t}, \] \[ W_{t}=\mu_{3}^{\ast}+\delta^{\ast}W_{t-1}+\sigma_{3}^{\ast}U_{3,t}. \] The noise variables $\{(U_{1,t},U_{2,t},U_{3,t})\}$ are i.i.d, Gaussian, independent of $\{S_{t}\}$, with mean zero and covariance matrix \[ \left[
\right] . \] The parameter values are $\alpha_{1}^{\ast}=\alpha_{2}^{\ast}=2$, $\beta _{1}^{\ast}=-\beta_{2}^{\ast}=0.5$, $\mu_{1}^{\ast}(1)=-\mu_{1}^{\ast}(2)=1$, $\gamma^{\ast}(1)=0.5$, $\gamma^{\ast}(2)=1$, $\sigma_{1}^{\ast}(1)=\sigma _{1}^{\ast}(2)=1$, $\mu_{2}^{\ast}=\mu_{3}^{\ast}=0.2$, $\psi^{\ast} =\delta^{\ast}=0.8$, $\sigma_{2}^{\ast}=\sigma_{3}^{\ast}=1$, and $\rho^{\ast },\omega^{\ast}\in\{0,0.65\}$.
For each of 1000 samples of size $T\in\{200,800,1600,3200\}$ from this DGP, estimates of the parameters of the outcome equation are obtained by maximizing the quasi-log-likelihood function ((ref)) associated with the mixture model ((ref))--((ref)), with $\Pr(S_{t}=1)=\bar{\vartheta}$ and $\varepsilon_{t}\sim\mathcal{N}(0,1)$. Monte Carlo estimates of the bias of the QML estimators of $\mu(1)$, $\mu(2)$, $\gamma(1)$, $\gamma(2)$, $\sigma(1)$ and $\sigma(2)$ are reported in Table (ref). We also report the ratio of the sampling standard deviation of the estimators to estimated standard errors (averaged across replications for each design point). The latter are computed using a sandwich\ estimator based on the Hessian and the gradient of the quasi-log-likelihood function (cf. pouzo22MLE), with weights obtained from the Parzen kernel and a data-dependent bandwidth selected by the plug-in method of andrews91.
The results for $\omega^{\ast}=0$ shown in the top panel of Table (ref) reveal that, although the estimators of $\mu(1)$ and $\mu(2)$ are somewhat biased in the smallest of the sample sizes considered, finite-sample bias becomes insignificant in the rest of the cases (regardless of the value of the correlation parameter $\rho^{\ast}$), as is to be expected in light of the result in Theorem (ref). Furthermore, unless the sample size is small, estimated standard errors are very accurate as approximations to the standard deviation of the QML estimators.
The bottom panel of Table (ref) contains results for a DGP with $\omega^{\ast}=0.65$. A non-zero value for the correlation parameter $\omega^{\ast}$ violates the exogeneity assumption about $W_{t}$ that is maintained throughout Section (ref) (and it is not obvious what the limit point of the QML estimator based on ((ref)) might be in this case). The simulation results show that estimators of the parameters of the outcome equation are significantly biased, even for the largest sample size considered in the simulations. Biases in this case are clearly a consequence of the mixture model being misspecified beyond the assumption of i.i.d. regimes, the additional source of misspecification being the incorrect assumption of uncorrelatedness of the covariate $W_{t}$ and the noise variable $U_{1,t}$. The results relating to the accuracy of the estimated standard errors are not substantially different from those obtained with $\omega^{\ast}=0$.
As pointed out in Section (ref), another situation in which ignoring Markov dependence of the regimes is costly involves outcome equations that contain autoregressive dynamics. To demonstrate numerically the difficulties in such a case, 1000 artificial samples of various sizes are generated according to the Markov-switching autoregression
with $\phi^{\ast}=0.9$; the remaining parameter values and the generating mechanisms of $\{Z_{t}\}$, $\{S_{t}\}$ and $\{(U_{1,t},U_{2,t})\}$ are the same as in earlier simulation experiments. For each artificial sample, the parameters of the regime-switching autoregressive model
are estimated by maximizing the quasi-log-likelihood function associated with it under the assumption that the regime variables $\{S_{t}\}$ are i.i.d., with $\Pr(S_{t}=1)=\bar{\vartheta}$, and the noise variables $\{\varepsilon_{t}\}$ are i.i.d., independent of $\{S_{t}\}$, with $\varepsilon_{t}\sim \mathcal{N}(0,1)$.
The Monte Carlo results reported in Table (ref) reveal substantial finite-sample bias in the case of the QML estimators of the intercepts $\mu(1)$ and $\mu(2)$. The QML estimators of $\sigma(1)$, $\sigma(2)$ and $\phi$ generally exhibit little bias, which may be partly due to the fact that the simulation design is such that the values of $\phi^{\ast}$ and $\sigma _{1}^{\ast}$ are the same regardless of the realized regime. Unlike the HMM case considered before, estimated standard errors are not always accurate as approximations to the finite-sample standard deviation of the QML estimators in the autoregressive model, even for a parameter such as $\sigma(1)$, which is estimated with little bias. We note that qualitatively similar results are obtained when, in addition to $Y_{t-1}$, an exogenous covariate $W_{t}$, generated as in the previous experiments, is included in the right-hand sides of ((ref)) and ((ref)). {
}
In this paper, we have considered QML estimation of the parameters of a generalized HMM with exogenous covariates and a finite hidden state space. A distinguishing feature of our approach is that it allows the regime sequence to be a temporally inhomogeneous Markov chain with covariate-dependent transition probabilities. It has been shown that a mixture model with independent regimes is robust in the presence of correlated Markov regimes, in the sense that the parameters of the outcome equation can be estimated consistently by maximizing the quasi-likelihood function associated with the misspecified mixture model.
One possible application of our main result is to exploit it to construct tests for the number of regimes in HMMs with covariate-dependent transition probabilities, adopting a QML-based approach analogous to that of chowhite07. As is well known, such testing problems are non-standard and typically involve unidentifiable nuisance parameters, parameters that lie on the boundary of the parameter space, singularity of the information matrix, and non-quadratic approximations to the log-likelihood function.