EconBase
← Back to paper

Testing for observation-dependent regime switching in mixture autoregressive models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

95,811 characters · 19 sections · 106 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Testing for observation-dependent regime switching in mixture autoregressive models

abstractTesting for regime switching when the regime switching probabilities are specified either as constants (`mixture models') or are governed by a finite-state Markov chain (`Markov switching models') are long-standing problems that have also attracted recent interest. This paper considers testing for regime switching when the regime switching probabilities are time-varying and depend on observed data (`observation-dependent regime switching'). Specifically, we consider the likelihood ratio test for observation-dependent regime switching in mixture autoregressive models. The testing problem is highly nonstandard, involving unidentified nuisance parameters under the null, parameters on the boundary, singular information matrices, and higher-order approximations of the log-likelihood. We derive the asymptotic null distribution of the likelihood ratio test statistic in a general mixture autoregressive setting using high-level conditions that allow for various forms of dependence of the regime switching probabilities on past observations, and we illustrate the theory using two particular mixture autoregressive models. The likelihood ratio test has a nonstandard asymptotic distribution that can easily be simulated, and Monte Carlo studies show the test to have satisfactory finite sample size and power properties. JEL classification: C12, C22, C52. Keywords: Likelihood ratio test, singular information matrix, higher-order approximation of the log-likelihood, logistic mixture autoregressive model, Gaussian mixture autoregressive model.

Introduction

Different regime switching models are in widespread use in economics, finance, and other fields. When the regime switching probabilities are constants, these models are often referred to as `mixture models', and when these probabilities depend on past regimes and are governed by a finite-state Markov chain, the term (time homogeneous) `Markov switching models' is typically used. In this paper, we are interested in the case where the regime switching probabilities depend on observed data, a case we refer to as `observation-dependent regime switching'. Models of this kind can be viewed as special cases of time inhomogeneous Markov switching models (in which regime switching probabilities depend on both past regimes and observed data). Overviews of regime switching models can be found, for example, in fruhwirth2006finite and hamilton2016macroeconomic. Of critical interest in all these models is whether the use of several regimes is warranted or if a single-regime model would suffice. Testing for regime switching in all these models is plagued by several irregular features such as unidentified parameters and parameters on the boundary and is consequently notoriously difficult.

Tests for Markov switching have been considered by several authors in the econometrics literature. hansen1992likelihood and garcia1998asymptotic both considered sup-type likelihood ratio (LR) tests in Markov switching models but they did not present complete solutions. Hansen derived a bound for the distribution of the LR statistic, leading to a conservative procedure, while Garcia did not treat all the non-standard features of the problem in detail. cho2007testing analyzed the use of a LR statistic for a mixture model to test for Markov-switching type regime switching. They found their test based on a mixture model to have power against Markov switching alternatives even though it ignores the temporal dependence of the Markov chain. \citet*{carrasco2014optimal} took a different approach and proposed an information matrix type test that they showed to be asymptotically optimal against Markov switching alternatives. Very recently, both Qu2017Likelihood and Kasahara2017Testing have studied the LR statistic for regime switching in Markov switching models.

Regarding testing for mixture type regime switching, the existing literature is extensive, and several early references can be found, for instance, in mclachlan2000finite. Most papers in this literature consider the case of independent observations without regressors. Notable exceptions allowing for regressors (but not dependent data) and having set-ups closer to the present paper are zhu2004hypothesis,zhu2006asymptotics and kasahara2015testing who consider (among other things) LR tests for regime switching. Further comparison to these works will be provided in later sections.

In contrast to testing for Markov switching or mixture type regime switching, there exists almost no literature on testing for observation-dependent regime switching. The only two exceptions we are aware of are the unpublished PhD thesis of jeffries1998logistic and the recent paper of shen2015inference. Jeffries's thesis, which appears to have gone largely unnoticed, analyses the LR test in a specific (first-order) mixture autoregressive model; we will discuss his work further in later sections. shen2015inference consider the case of independent observations with regressors and observation-dependent regime switching, and propose an `expectation maximization test' for regime switching.

In this paper we consider testing for observation-dependent regime switching in a time series context. Specifically, we analyze the asymptotic distribution of the LR test statistic for testing a linear autoregressive model against a two-regime mixture autoregressive model with observation-dependent regime switching. Mixture autoregressive models have been discussed for instance in wong2000mixture,wong2001logistic, \citet*{dueker2007contemporaneous}, \citet*{dueker2011multivariate}, and \citet*{kalliovirta2015gaussian,kalliovirta2016gaussian}; further discussion of this previous work will be provided in Section 2. Motivation for allowing the regime switching probabilities to depend on observed data stems, for instance, from the desire to associate changes in regime to observable economic variables.\footnote{See, for instance, hamilton2016macroeconomic, whose Handbook of Macroeconomics chapter begins “Many economic time series exhibit dramatic breaks associated with events such as economic recessions, financial panics, and currency crises. Such changes in regime may arise from tipping points or other nonlinear dynamics and are core to some of the most important questions in macroeconomics.” }\footnote{More general models in which the regime switching probabilities are allowed to depend on both past regimes and observable variables have also been considered, see, e.g., \citet*{diebold1994regime}, filardo1994business, and \citet*{kim2008estimation}.} Following kasahara2015testing it would also be possible to consider the more general testing problem that in a model with more than two regimes the number of regimes can be reduced. However, as even the case of two regimes is quite complex in our set-up, we leave this extension to future research.

We consider mixture autoregressive (MAR) models in a rather general setting employing high-level conditions that allow for various forms of observation-dependent regime switching. As specific examples, we treat the so-called logistic MAR (LMAR) model of wong2001logistic and (a version of the) Gaussian MAR (GMAR) model of kalliovirta2015gaussian in detail. The technical challenges we face in analyzing the LR test statistic are similar to those when testing for Markov switching and mixture type regime switching. First, there are nuisance parameters that are unidentified under the null hypothesis. This is the classical davies1977hypothesis,davies1987hypothesis type problem. Second, under the null hypothesis, there are parameters on the boundary of the permissible parameter space. Such problems (also allowing for unidentified nuisance parameters under the null) are discussed in andrews1999estimation,andrews2001testing. Third, the Fisher information matrix is (potentially) singular, preventing the use of conventional second-order expansions of the log-likelihood to analyze the LR test statistic. Such problems are discussed by \citet*{rotnitzky2000likelihood}, and suitable reparameterizations and higher-order expansions are needed to analyze the LR statistic. A particular challenge in the present paper is to deal with these three problems simultaneously. Similar problems were faced by kasahara2015testing, and inspired by their work we consider a suitably reparameterized model, write a higher-order expansion of the log-likelihood function as a quadratic function of the new parameters, and then derive the asymptotic distribution of the LR test statistic by slightly extending and adapting the arguments of andrews1999estimation,andrews2001testing and zhu2006asymptotics (who partially generalize results of Andrews). Our two examples demonstrate that, compared to the mixture type regime switching considered by kasahara2015testing, observation-dependent regime switching can either simplify or complicate the analysis of the LR test statistic.

We contribute to the literature in several ways. (1) To the best of our knowledge, apart from the unpublished PhD thesis of jeffries1998logistic, we are the first to study testing for observation-dependent regime switching using the LR test statistic and among the rather few to allow for dependent observations. (2) We provide a general framework to cover various forms of observation-dependent regime switching, making our results potentially applicable to several models not explicitly discussed in the present paper. (3) From a methodological perspective, we slightly extend and adapt certain arguments of andrews1999estimation,andrews2001testing and zhu2006asymptotics, which may be of independent interest.

The rest of the paper is organized as follows. Section 2 reviews mixture autoregressive models. Section 3 analyzes the LR test statistic for testing a linear autoregressive model against a two-regime mixture autoregressive model. Simulation-based critical values and a Monte Carlo study are discussed in Section 4, and Section 5 concludes. Appendices A\textendash C contain technical details and proofs. Supplementary Appendices D\textendash E, available from the authors upon request, contain further technical details omitted from the paper.

Finally, a few notational conventions are given. All vectors will be treated as column vectors and, for the sake of uncluttered notation, we shall write $x=(x_{1},\ldots,x_{n})$ for the (column) vector $x$ where the components $x_{i}$ may be either scalars or vectors (or both). For any vector or matrix $x$, the Euclidean norm is denoted by $\left\Vert x\right\Vert $. We let $X_{T\alpha}=o_{p\alpha}(1)$ and $X_{T\alpha}=O{}_{p\alpha}(1)$ stand for $\sup_{\alpha\in A}\left\Vert X_{T\alpha}\right\Vert =o_{p}(1)$ and $\sup_{\alpha\in A}\left\Vert X_{T\alpha}\right\Vert =O{}_{p}(1)$, respectively, and $\lambda_{\min}(\cdot)$ and $\lambda_{\max}(\cdot)$ for the smallest and largest eigenvalue of the indicated matrix.

Mixture autoregressive models

General formulation

Let $y_{t}$ ($t=1,2,\ldots$) be a real-valued time series of interest, and let $\mathcal{F}_{t-1}=\sigma\left(y_{s},\textrm{ }s<t\right)$ denote the $\sigma$\textendash algebra generated by past $y_{t}$'s. We use $P_{t-1}\left(\cdot\right)$ to signify the conditional probability of the indicated event given $\mathcal{F}_{t-1}$. In the general two component mixture autoregressive model we consider the $y_{t}$'s are generated by

equation[equation omitted — 263 chars of source]

where the parameters $\tilde{\sigma}_{1}$ and $\tilde{\sigma}_{2}$ are positive, and conditions required for the autoregressive parameters $\tilde{\phi}_{i}$ and $\tilde{\varphi}_{i}$ ($i=1,\ldots,p$) will be discussed later. Furthermore, $\varepsilon_{t}$ and $s_{t}$ are (unobserved) stochastic processes which satisfy the following conditions: (a) $\varepsilon_{t}$ is a sequence of independent standard normal random variables such that $\varepsilon_{t}$ is independent of $\{y_{t-j},\ j>0\}$, (b) $s_{t}$ is a sequence of Bernoulli random variables such that, for each $t$, $P_{t-1}(s_{t}=1)=\alpha_{t}$ with $\alpha_{t}$ a function of $\boldsymbol{y}_{t-1}=(y_{t-1},\ldots,y_{t-p})$, and (c) conditional on $\mathcal{F}_{t-1}$, $s_{t}$ and $\varepsilon_{t}$ are independent.

The conditional probabilities $\alpha_{t}$ and $1-\alpha_{t}$ ($=P_{t-1}(s_{t}=0)$) are referred to as mixing weights. They can be thought of as (conditional) probabilities that determine which one of the two autoregressive components of the mixture generates the next observation $y_{t}$. In condition (b) it is assumed that the mixing weight $\alpha_{t}$ (and hence also the conditional distribution of $y_{t}$ given its past) only depends on $p$ lags of $y_{t}$; allowing for more than $p$ lags in the mixing weight would be possible at the cost of more complicated notation.

We assume that of the original parameters $\tilde{\phi}=(\tilde{\phi}_{0},\tilde{\phi}_{1},\ldots,\tilde{\phi}_{p},\tilde{\sigma}_{1}^{2})$ and $\tilde{\varphi}=(\tilde{\varphi}_{0},\tilde{\varphi}_{1},\ldots,\tilde{\varphi}_{p},\tilde{\sigma}_{2}^{2})$ in the two regimes, $q_{1}$ parameters are a priori assumed the same in both regimes and the remaining $q_{2}$ parameters are potentially different in the two regimes (with $q_{1}+q_{2}=p+2$). For instance, one may assume that $\tilde{\phi}_{0}$ and $\tilde{\varphi}_{0}$ are equal, or alternatively that $\tilde{\sigma}_{1}^{2}$ and $\tilde{\sigma}_{2}^{2}$ are equal. If such an assumption is plausible, taking it into account when devising a test for regime switching will be advantageous (it will lead to a test with better power). To this end, let $\beta$ be a $q_{1}\times1$ vector of common parameters, and let $\phi$ and $\varphi$ be $q_{2}\times1$ vectors of (potentially) different parameters. Then, for some known $(p+2)$\textendash dimensional permutation matrix $P$, $(\beta,\phi)=P\tilde{\phi}$ and $(\beta,\varphi)=P\tilde{\varphi}$. For simplicity, we assume that $\beta$ and $\phi$ are variation-free, requiring the autoregressive coefficients $\tilde{\phi}_{1},\ldots,\tilde{\phi}_{p}$ to be contained in either $\beta$ or $\phi$ (the same variation-freeness is assumed of $\beta$ and $\varphi$). If there are no common coefficients in the two regimes, the parameter $\beta$ can be dropped and $\phi=\tilde{\phi}$ and $\varphi=\tilde{\varphi}$.

As for the mixing weight $\alpha_{t}$, in addition to past $y_{t}$'s it depends on unknown parameters which may include components of the parameter vector $(\beta,\phi,\varphi)$ and an additional parameter $\alpha$ (scalar or vector). When this dependence needs to be emphasized we use the notation $\alpha_{t}(\alpha,\beta,\phi,\varphi)$.

Using equation ((ref)) and the conditions following it, the conditional density function of $y_{t}$ given its past, $f(\cdot\mid\mathcal{F}_{t-1})$, is obtained as

equation[equation omitted — 124 chars of source]

where the notation $f_{t}(\beta,\phi)$ signifies the density function of a (univariate) normal distribution with mean $\tilde{\phi}_{0}+\tilde{\phi}_{1}y_{t-1}+\cdots+\tilde{\phi}_{p}y_{t-p}$ and variance $\tilde{\sigma}_{1}^{2}$ evaluated at $y_{t}$, that is,

equation[equation omitted — 234 chars of source]

with $\mathscr{{\scriptstyle \mathfrak{N}}}(u)=(2\bar{\pi})^{-1/2}\exp(-u^{2}/2)$ the density function of a standard normal random variable and $\bar{\pi}=3.14\ldots$ the number pi. The notation $f_{t}(\beta,\varphi)$ is defined similarly by using the parameters $\tilde{\varphi}_{i}$ $\left(i=0,\ldots,p\right)$ and $\tilde{\sigma}_{2}^{2}$ instead of $\tilde{\phi}_{i}$ $\left(i=0,\ldots,p\right)$ and $\tilde{\sigma}_{1}^{2}$. Thus, the distribution of $y_{t}$ given its past is specified as a mixture of two normal densities with time varying mixing weights $\alpha_{t}$ and $1-\alpha_{t}$.

Different mixture autoregressive models are obtained by different specifications of the mixing weights (or in our case the single mixing weight $\alpha_{t}$). In some of the proposed models more than two mixture components are allowed but for reasons to be discussed below we will not consider these extensions. If the mixing weights are assumed constant over time the general mixture autoregressive model introduced above reduces to (a two component version) of the MAR\ model studied by wong2000mixture. In the LMAR\ model of wong2001logistic, a logistic transformation of the two mixing weights is assumed to be a linear function of past observed variables. Related two-regime mixture models with time-varying mixing weights have also been considered by gourieroux2006stochastic, dueker2007contemporaneous and \citet*{bec2008acr} whereas lanne2003modeling and kalliovirta2015gaussian have considered mixture autoregressive models in which multiple regimes are allowed.

A common problem with the application of mixture autoregressive models is determining the value of the (usually unknown) number of component models or regimes. As discussed in the Introduction, several irregular features make this problem difficult and these difficulties are encountered even when the observations are a random sample from a mixture of (two) normal distributions. To our knowledge the only solution presented for mixture autoregressive models is provided for a simple first order case with no intercept terms in the unpublished PhD thesis of jeffries1998logistic. As discussed in the recent papers by kasahara2012testing,kasahara2015testing and the references therein, some of the difficulties involved stem from properties of the normal distribution.

The difficulties referred to above also partly explain the complexity of our testing problem and why we only consider test procedures that can be used to test the null hypothesis that a two component mixture autoregressive model reduces to a conventional linear autoregressive model. Following the ideas in zhu2006asymptotics and kasahara2012testing,kasahara2015testing, we derive a LR test in the general set-up described above and apply it to two particular cases, the LMAR model of wong2001logistic and the GMAR model of kalliovirta2015gaussian. Next, we shall discuss these two models in more detail.

Two particular examples

\paragraph{LMAR Example.}

The LMAR model of wong2001logistic is defined by specifying the mixing weight $\alpha_{t}$ as \[ \alpha_{t}^{L}=\alpha_{t}^{L}(\alpha)=\frac{\exp(\alpha_{0}+\alpha_{1}y_{t-1}+\cdots+\alpha_{r}y_{t-m})}{1+\exp(\alpha_{0}+\alpha_{1}y_{t-1}+\cdots+\alpha_{r}y_{t-m})}, \] where the vector $\alpha=(\alpha_{0},\alpha_{1},\ldots,\alpha_{m})$ contains $m+1$ unknown parameters and the order $m$ ($1\leq m\leq p$) is assumed known.

\paragraph{GMAR Example.}

In the GMAR model of kalliovirta2015gaussian the mixing weight is defined as

equation[equation omitted — 288 chars of source]

where $\alpha\in(0,1)$ is an unknown parameter and $\mathsf{n}_{p}(\boldsymbol{y}_{t-1};\cdot)$ denotes the density function of a particular $p$\textendash dimensional ($p\geq1$) normal distribution defined as follows.

First, define the auxiliary Gaussian AR($p$) processes (cf. equation ((ref))) \[ \nu{}_{1,t}=\tilde{\phi}_{0}+\sum_{i=1}^{p}\tilde{\phi}_{i}\nu_{1,t-i}+\tilde{\sigma}_{1}\varepsilon_{t}\ \ \ \textrm{and}\ \ \ \nu_{2,t}=\tilde{\varphi}_{0}+\sum_{i=1}^{p}\tilde{\varphi}_{i}\nu_{2,t-i}+\tilde{\sigma}_{2}\varepsilon_{t}, \] where the autoregressive coefficients are assumed to satisfy

equation[equation omitted — 286 chars of source]

This condition implies that the processes $\nu{}_{1,t}$ and $\nu{}_{2,t}$ are stationary and that each of the two component models in ((ref)) satisfies the usual stationarity condition of the conventional linear AR($p$) model. Now set $\mathbf{\boldsymbol{\nu}}_{m,t}=\left(\nu_{m,t},\ldots,\nu_{m,t-p+1}\right)$ and $\mathbf{1}_{p}=\left(1,\ldots,1\right)$ ($p\times1$), and let $\mu_{m}\mathbf{1}_{p}$ and $\mathbf{\Gamma}_{m,p}$ signify the mean vector and covariance matrix of $\mathbf{\boldsymbol{\nu}}_{m,t}$ ($m=1,2$).\footnote{We have $\mu_{1}=\tilde{\phi}_{0}/\tilde{\phi}\left(1\right)$ and $\mu_{2}=\tilde{\varphi}_{0}/\tilde{\varphi}\left(1\right)$, whereas each of $\mathbf{\Gamma}_{m,p}$, $(m=1,2)$, has the familiar form of being a $p\times p$ symmetric Toeplitz matrix with $\gamma_{m,0}=Cov[\nu_{m,t},\nu_{m,t}]$ along the main diagonal, and $\gamma_{m,i}=Cov[\nu_{m,t},\nu{}_{m,t-i}]$, $i=1,\ldots,p-1$, on the diagonals above and below the main diagonal. Similarly to $\mu_{1}$ and $\mu_{2}$ the elements of the covariance matrices $\mathbf{\Gamma}_{1,p}$ and $\mathbf{\Gamma}_{2,p}$ are treated as functions of the parameters $\tilde{\phi}$ and $\tilde{\varphi}$, respectively (for details of this dependence, see lutkepohl2005new).} The random vector $\mathbf{\boldsymbol{\nu}}_{1,t}$ follows the $p$\textendash dimensional multivariate normal distribution with density

equation[equation omitted — 368 chars of source]

and the density of $\mathbf{\boldsymbol{\nu}}_{2,t}$, denoted by $\mathsf{n}_{p}(\mathbf{\boldsymbol{\nu}}_{2,t};\tilde{\varphi})$, is defined similarly. Equation ((ref)) and conditions ((ref))\textendash ((ref)) define the (two component) GMAR model (condition ((ref)) is part of the definition of the model because it is used to define the mixing weights).

Test procedure

We now consider a test procedure of the null hypothesis that a two component mixture autoregressive model reduces to a conventional linear autoregressive model.

The null hypothesis and the LR test statistic

We denote the conditional density function corresponding to the unrestricted model as (see ((ref))) \[ f_{2,t}(\alpha,\beta,\phi,\varphi):=f_{2}(y_{t}\mid\boldsymbol{y}_{t-1};\alpha,\beta,\phi,\varphi):=\alpha_{t}(\alpha,\beta,\phi,\varphi)f_{t}(\beta,\phi)+\left(1-\alpha_{t}(\alpha,\beta,\phi,\varphi)\right)f_{t}(\beta,\varphi), \] where we now make the dependence of the mixing weight on the parameters explicit. With this notation the log-likelihood function of the model based on a sample $(y_{-p+1},\ldots,y_{T})$ (and conditional on the initial values $(y_{-p+1},\ldots,y_{0})$) is $L_{T}(\alpha,\beta,\phi,\varphi)=\sum_{t=1}^{T}l_{t}(\alpha,\beta,\phi,\varphi)$ where \[ l_{t}(\alpha,\beta,\phi,\varphi)=\log[f_{2,t}(\alpha,\beta,\phi,\varphi)]=\log\left[\alpha_{t}(\alpha,\beta,\phi,\varphi)f_{t}(\beta,\phi)+\left(1-\alpha_{t}(\alpha,\beta,\phi,\varphi)\right)f_{t}(\beta,\varphi)\right]. \] The following assumption provides conditions on the data generation process, the parameter space of $\left(\alpha,\beta,\phi,\varphi\right)$, and the mixing weight $\alpha_{t}\left(\alpha,\beta,\phi,\varphi\right)$.

assumption\begin{description} • The $y_{t}$'s are generated by a stationary linear Gaussian AR($p$) model with (the true but unknown) parameter value $\tilde{\phi}^{\ast}$ an interior point of $\tilde{\Phi}$, a compact subset of $\{\tilde{\phi}=(\tilde{\phi}_{0},\tilde{\phi}_{1},\ldots,\tilde{\phi}_{p},\tilde{\sigma}^{2})\in\mathbb{R}^{p+2}:\tilde{\phi}_{0}\in\mathbb{R};\,1-\sum_{i=1}^{p}\tilde{\phi}_{i}z^{i}\neq0\text{ \ for }\left\vert z\right\vert \leq1;\,\tilde{\sigma}^{2}\in(0,\infty)\}$. • The parameter space of $\left(\alpha,\beta,\phi,\varphi\right)$ is $A\times B\times\Phi\times\Phi$, where $A$ is a compact subset of $\mathbb{R}^{a}$ and $B$ and $\Phi$ are those compact subsets of $\mathbb{R}^{q_{1}}$ and $\mathbb{R}^{q_{2}}$, respectively, that satisfy $(\beta,\phi)\in B\times\Phi$ if and only if $P^{-1}(\beta,\phi)\in\tilde{\Phi}$ (here $P$ is as in the third paragraph of Section 2.1). • For all $t$ and all $\left(\alpha,\beta,\phi,\varphi\right)\in A\times B\times\Phi\times\Phi$, the mixing weight $\alpha_{t}\left(\alpha,\beta,\phi,\varphi\right)$, is $\sigma(\boldsymbol{y}{}_{t-1})$\textendash measurable (with $\sigma(\boldsymbol{y}{}_{t-1})$ denoting the $\sigma$\textendash algebra generated by $\boldsymbol{y}{}_{t-1}$) and satisfies $\alpha_{t}\left(\alpha,\beta,\phi,\varphi\right)\in(0,1)$. \end{description}

As our interest is to study the asymptotic null distribution of the LR test statistic, Assumption 1(i) requires the data to be generated by a stationary linear Gaussian AR($p$) model. Assuming a compact parameter space in Assumptions 1(i) and (ii) is a standard requirement which facilitates proofs. Assumption 1(ii) accommodates to the main cases of interest, namely $\beta=\tilde{\phi}_{0}$, $\beta=(\tilde{\phi}_{1},\ldots,\tilde{\phi}_{p})$, $\beta=\tilde{\sigma}^{2}$, or any combination of these.\footnote{Note that assuming the autoregressive parameters $\phi$ and $\varphi$ to have a common parameter space is made for simplicity and could be relaxed; for an example where such a relaxation would be needed, see the ACR model of bec2008acr.}

Assumption 1(iii) implies that our two-component mixture autoregressive model reduces to a linear autoregression only when $\phi=\varphi$, regardless of the values of $\alpha\in A$ and $\beta\in B$. The null hypothesis to be tested is therefore $\phi=\varphi$ and the alternative is $\phi\neq\varphi$ or, more precisely, \[ H_{0}:(\phi,\varphi)\in(\Phi\times\Phi)^{0},\;\alpha\in A,\,\beta\in B\qquad\textrm{vs.}\qquad H_{1}:(\phi,\varphi)\in(\Phi\times\Phi)\setminus(\Phi\times\Phi)^{0},\;\alpha\in A,\,\beta\in B, \] where \[ (\Phi\times\Phi)^{0}=\{(\phi,\varphi)\in\Phi\times\Phi:\phi=\varphi\}. \] Note that under the null hypothesis the parameter $\alpha$ vanishes from the likelihood function and is therefore unidentified.

Let $f_{t}^{0}(\tilde{\phi}):=f^{0}(y_{t}\mid\boldsymbol{y}_{t-1};\tilde{\phi})$ and $L_{T}^{0}(\tilde{\phi})$ denote the conditional density and log-likelihood corresponding to the restricted model, that is, \[ f_{t}^{0}(\tilde{\phi})=f_{2,t}(\alpha,\beta,\phi,\phi)=f_{t}(\tilde{\phi})\,\,\,\,\,\text{and}\,\,\,\,\,L_{T}^{0}(\tilde{\phi})=\sum_{t=1}^{T}l_{t}^{0}(\tilde{\phi})\,\,\,\,\,\text{with}\,\,\,\,\,l_{t}^{0}(\tilde{\phi})=\log[f_{t}(\tilde{\phi})] \] (here the superscript $0$ refers to the model restricted by the null hypothesis). Note that these quantities are obtained from a linear Gaussian AR($p$) model. As $f_{2,t}(\alpha,\beta^{*},\phi^{\ast},\phi^{\ast})=f_{t}(\tilde{\phi}^{\ast})$ for any $\alpha\in A$, in the unrestricted model the parameter vector $(\alpha,\beta^{*},\phi^{\ast},\phi^{\ast})$ corresponds to the true model for any $\alpha\in A$.

As already indicated, Assumption 1(iii) implies that the restriction $\phi=\varphi$ is the only possibility to formulate the null hypothesis. However, this is not necessarily the case if (against Assumption 1(iii)) the mixing weight $\alpha_{t}\left(\alpha,\beta,\phi,\varphi\right)$ were allowed to take the boundary values zero and one. Of our two examples this would be possible for the GMAR model of kalliovirta2015gaussian but not for the LMAR model of wong2001logistic. In the GMAR model $\alpha_{t}\left(\alpha,\beta,\phi,\varphi\right)$ takes the boundary values zero and one when the parameter $\alpha$ takes these values (see Section 2.2). In both cases a linear autoregression results and either the parameter $\phi$ or $\varphi$ is unidentified (see ((ref))) (the MAR model of wong2000mixture provides a similar example). It would be possible to obtain tests for the GMAR model by using the null hypotheses which specifies $\alpha=0$ or $\alpha=1$. However, as in kasahara2012testing (see also kasahara2015testing), this approach would require rather restrictive assumptions and would also lead to very complicated derivations.\footnote{See, for instance, the remarks following Proposition 5 in kasahara2012testing or property (a) on p. 1633 of kasahara2015testing.} Therefore, we will not consider this option.

As the parameter $\alpha$ is unidentified under the null hypothesis, the appropriate likelihood ratio type test statistic is \[ LR_{T}=\sup_{\alpha\in A}LR_{T}\left(\alpha\right), \] where, for each fixed $\alpha\in A$, \[ LR_{T}\left(\alpha\right)=2[\sup_{(\beta,\phi,\varphi)\in B\times\Phi\times\Phi}L_{T}(\alpha,\beta,\phi,\varphi)-\sup_{\tilde{\phi}\in\tilde{\Phi}}L_{T}^{0}(\tilde{\phi})]. \] To obtain an operational test statistic let, for each fixed $\alpha\in A$, $(\hat{\beta}_{T\alpha},\hat{\phi}_{T\alpha},\hat{\varphi}_{T\alpha})$ denote an (approximate) unrestricted maximum likelihood (ML) estimator of the parameter vector $\left(\beta,\phi,\varphi\right)$. We make the following assumption.

assumptionThe unrestricted ML estimator satisfies the following conditions: \begin{description} • $L_{T}(\alpha,\hat{\beta}_{T\alpha},\hat{\phi}_{T\alpha},\hat{\varphi}_{T\alpha})=\sup\nolimits _{(\beta,\phi,\varphi)\in B\times\Phi\times\Phi}L_{T}(\alpha,\beta,\phi,\varphi)+o_{p\alpha}(1)$, • $(\hat{\beta}_{T\alpha},\hat{\phi}_{T\alpha},\hat{\varphi}_{T\alpha})=(\beta^{*},\phi^{\ast},\phi^{\ast})+o_{p\alpha}(1)$. \end{description}

Assumption 2(i) means that $(\hat{\beta}_{T\alpha},\hat{\phi}_{T\alpha},\hat{\varphi}_{T\alpha})$ is assumed to maximize the likelihood function only asymptotically. This assumption is technical and made for ease of exposition (see andrews1999estimation and zhu2006asymptotics for similar assumptions in related problems). Assumption 2(ii) is a high level condition on (uniform) consistency of the ML estimator and is analogous to Assumption 1 of andrews2001testing. It has to be verified on a case by case basis (this is exemplified below for the LMAR model and GMAR model).

As for the term $\sup_{\tilde{\phi}\in\tilde{\Phi}}L_{T}^{0}(\tilde{\phi})$ in the LR test statistic, note that $L_{T}^{0}(\tilde{\phi})$ is the (conditional) log-likelihood function of a linear Gaussian AR($p$) model. Let $\hat{\tilde{\phi}}_{T}$ denote an (approximate) maximum likelihood estimator of the parameters of a linear Gaussian AR($p$) model, that is, $\hat{\tilde{\phi}}_{T}$ satisfies\footnote{Note that the parameter space for $\tilde{\phi}$ is the compact set $\tilde{\Phi}$ and not the entire stationarity region of a (causal) linear AR($p$) model. Asymptotically, also the OLS estimator can be used. } \[ L_{T}^{0}(\hat{\tilde{\phi}}_{T})=\sup_{\tilde{\phi}\in\tilde{\Phi}}L_{T}^{0}(\tilde{\phi})+o_{p}(1)\,\,\,\textrm{and}\,\,\,\hat{\tilde{\phi}}_{T}=\tilde{\phi}^{\ast}+o_{p}(1). \] Noting that $L_{T}(\alpha,\beta^{*},\phi^{\ast},\phi^{\ast})=L_{T}^{0}(\tilde{\phi}^{\ast})$ for any $\alpha$ now allows us to write $LR_{T}\left(\alpha\right)$ as

equation[equation omitted — 272 chars of source]

The analysis of the second term on the right hand side is standard while dealing with the first term is more demanding requiring a substantial amount of preparation.

Examples (continued)

\paragraph{LMAR Example. }

In the LMAR example, we assume there are no common parameters in the two regimes so that the parameter $\beta$ is omitted, $\phi=\tilde{\phi}$, $\varphi=\tilde{\varphi}$, $q_{1}=0$, and $q_{2}=p+2$. To satisfy conditions (ii) and (iii) in Assumption 1, $A$ can be any compact subset of $\{(\alpha_{0},\alpha_{1},\ldots,\alpha_{m})\in\mathbb{R}^{m+1}:(\alpha_{1},\ldots,\alpha_{m})\neq(0,\ldots,0)\}$ where $1\leq m\leq p$. This ensures that the mixing weight $\alpha_{t}^{L}$ is not equal to a constant. For the verification of Assumption 2(ii), see Appendix B.

When $m=0$, the mixing weight $\alpha_{t}^{L}$ (and hence $1-\alpha_{t}^{L}$) is constant and the LMAR model reduces to the MAR model of wong2000mixture. In this special case our testing problem requires different and more complicated analyses than in the `real' LMAR case where $m\geq1$ and $(\alpha_{1},\ldots,\alpha_{r})\neq(0,\ldots,0)$ (we shall discuss this point more later). Therefore, the conditions $m\geq1$ and $(\alpha_{1},\ldots,\alpha_{m})\neq(0,\ldots,0)$ will be assumed in the sequel. A similar restriction is made by jeffries1998logistic in his (first-order) logistic mixture autoregressive model to facilitate the derivation of the LR test (see the hypotheses at the end of p. 95 and the following discussion, as well as the end of p. 110).

\paragraph{GMAR Example. }

The GMAR model exemplifies the setting with common coefficients by assuming that the intercept terms in the two regimes are the same (note that this still allows for different means in the two regimes). As will be discussed in more detail in Section 3.3.1, this assumption is partly due to the fact that otherwise the derivation of the LR test would become extremely complicated. Hence, in this example $\beta=\tilde{\phi}_{0}\,(=\tilde{\varphi}_{0})$, $\phi=(\tilde{\phi}_{1},\ldots,\tilde{\phi}_{p},\tilde{\sigma}_{1}^{2})$, $\varphi=(\tilde{\varphi}_{1},\ldots,\tilde{\varphi}_{p},\tilde{\sigma}_{2}^{2})$, $q_{1}=1$, and $q_{2}=p+1$. To satisfy Assumptions 1(ii) and (iii) the parameter space $A$ of $\alpha$ can be any compact and convex subset of $(0,1)$ (this also rules out the possibility that $\alpha=0$ or $\alpha=1$ discussed after Assumption 1). For the verification of Assumption 2(ii), see Appendix C.

It may be worth noting that there are cases where the mixing weight $\alpha_{t}^{G}$ is time invariant and equals $\alpha$. If this happens the GMAR model reduces to the MAR model of wong2000mixture.\footnote{An example is when $p=1$, $\tilde{\phi}_{0}=\tilde{\varphi}_{0}=0$, and $\tilde{\sigma}_{1}^{2}/(1-\tilde{\phi}_{1}^{2})=\tilde{\sigma}_{2}^{2}/(1-\tilde{\varphi}_{1}^{2})$ where the last equality can hold even if $(\tilde{\phi},\tilde{\sigma}_{1}^{2})$ is different from $(\tilde{\varphi},\tilde{\sigma}_{2}^{2})$.} However, unlike in the case of the LMAR model this fact does not complicate the derivation of our test. The reason seems to be that in the GMAR model the reduction occurs only for certain values of the parameters $\tilde{\phi}$ and $\tilde{\varphi}$ whereas in the LMAR model it occurs for all values of $\tilde{\phi}$ and $\tilde{\varphi}$.

Reparameterized model

In standard testing problems the derivation of the asymptotic distribution of the LR test would rely on a quadratic expansion of the log-likelihood function $L_{T}(\alpha,\beta,\phi,\varphi)=\sum_{t=1}^{T}l_{t}(\alpha,\beta,\phi,\varphi)$; when the parameter $\alpha$ is not identified under the null hypothesis, the relevant derivatives in this expansion would be with respect to $(\beta,\phi,\varphi)$ for fixed values of $\alpha\in A$. In problems with a singular information matrix it turns out to be convenient to follow rotnitzky2000likelihood and kasahara2012testing,kasahara2015testing and employ an appropriately reparameterized model.

The employed reparameterization is model specific and aims to have two conveniences. First, it transforms the null hypothesis $\phi=\varphi$ into a point null hypothesis where some components of the parameter vector are restricted to zero and the rest are left unrestricted. Second, and more importantly, it simplifies derivations in cases where the conventional quadratic expansion of the log-likelihood function breaks down because, under the null hypothesis, the scores of the parameters $(\beta,\phi,\varphi)$ are linearly dependent and, consequently, the (Fisher) information matrix is singular. As will be seen later, this is the case for the GMAR model of kalliovirta2015gaussian but not for the LMAR model of wong2001logistic.

General requirements for the reparameterization are described in the following assumption. Only the parameters restricted by the null hypothesis, $\phi$ and $\varphi$, are reparameterized. The examples in this and the following subsection illustrate how the reparameterization could be chosen.

assumption\begin{description} • For every $\alpha\in A$, the mapping $(\pi,\varpi)=\boldsymbol{\pi}_{\alpha}(\phi,\varphi)$ from $\Phi\times\Phi$ to $\Pi_{\alpha}$ is one-to-one with $\boldsymbol{\pi}_{\alpha}$ and $\boldsymbol{\pi}_{\alpha}^{-1}$ continuous. • For every $\alpha\in A$, $\boldsymbol{\pi}_{\alpha}((\Phi\times\Phi)^{0})=\Phi\times\{0\}$ and $\boldsymbol{\pi}_{\alpha}(\phi^{\ast},\phi^{\ast})=(\pi^{\ast},0):=(\phi^{\ast},0)$. • $(\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})=(\beta^{*},\pi^{\ast},0)+o_{p\alpha}(1)$, where $(\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})\coloneqq\boldsymbol{\pi}_{\alpha}(\hat{\phi}_{T\alpha},\hat{\varphi}_{T\alpha})$. \end{description}

We sometimes refer to the reparameterization described in Assumption (ref) as the `$\pi$-parameterization' and the original reparameterization as the `$\phi$-parameterization'. Note that the transformed parameters $\pi$ and $\varpi$ generally depend on $\alpha$ but, for brevity, we suppress this dependence from the notation. The parameter space of $(\pi,\varpi)$ also depends on $\alpha$ and is given, for any $\alpha\in A$, by \[ \Pi_{\alpha}=\{(\pi,\varpi)\in\mathbf{\mathbb{R}}^{2q_{2}}:(\pi,\varpi)=\boldsymbol{\pi}_{\alpha}(\phi,\varphi)\;\;\textrm{for some}\;\;(\phi,\varphi)\in\Phi\times\Phi\}. \] By Assumption (ref)(ii), the null hypothesis $\phi=\varphi$ can be equivalently written as $\varpi=0$ or, more precisely, as \[ H_{0}:\pi\in\Phi,\;\varpi=0,\;\alpha\in A,\,\beta\in B\qquad\textrm{vs.}\qquad H_{1}:(\pi,\varpi)\in\Pi_{\alpha}\setminus(\Phi\times\{0\}),\;\alpha\in A,\,\beta\in B. \] Note that under $H_{0}$, the parameters $\beta$ and $\pi$ are identified, but $\alpha$ is not. As for Assumption (ref)(iii), it is a high level condition similar to Assumption (ref)(ii) from which it can be derived with appropriate additional assumptions. A simple Lipschitz condition similar to andrews1992generic, given in Lemma (ref) in Appendix A, is one possibility.

To develop further notation, partition $\boldsymbol{\pi}_{\alpha}^{-1}(\pi,\varpi)$ into two $q_{2}$\textendash dimensional components as $\boldsymbol{\pi}_{\alpha}^{-1}(\pi,\varpi)=(\boldsymbol{\pi}_{\alpha,1}^{-1}(\pi,\varpi),\boldsymbol{\pi}_{\alpha,2}^{-1}(\pi,\varpi))$, and define \[ \hspace{-10pt}f_{2,t}^{\pi}(\alpha,\beta,\pi,\varpi):=f_{2}(y_{t}\mid\boldsymbol{y}_{t-1};\alpha,\beta,\phi,\varphi)=\alpha_{t}^{\pi}(\alpha,\beta,\pi,\varpi)f_{t}(\beta,\boldsymbol{\pi}_{\alpha,1}^{-1}(\pi,\varpi))+(1-\alpha_{t}^{\pi}(\alpha,\beta,\pi,\varpi))f_{t}(\beta,\boldsymbol{\pi}_{\alpha,2}^{-1}(\pi,\varpi)), \] where $\alpha_{t}^{\pi}(\alpha,\beta,\pi,\varpi)=\alpha_{t}(\alpha,\beta,\boldsymbol{\pi}_{\alpha,1}^{-1}(\pi,\varpi),\boldsymbol{\pi}_{\alpha,2}^{-1}(\pi,\varpi))$ and the function $f_{t}(\cdot)$ is as in ((ref)). The log-likelihood function of the reparameterized model can now be expressed as

equation[equation omitted — 123 chars of source]

where $l_{t}^{\pi}(\alpha,\beta,\pi,\varpi)=\log[f_{2,t}^{\pi}(\alpha,\beta,\pi,\varpi)]$, and in the $\pi$\textendash parameterization equation ((ref)) reads as

equation[equation omitted — 279 chars of source]

Examples (continued)

\paragraph{LMAR Example. }

The reparameterization we employ in the LMAR model is \[ (\pi,\varpi)=\boldsymbol{\pi}_{\alpha}(\phi,\varphi)=(\phi,\phi-\varphi)\;\;\textrm{so that}\;\;(\phi,\varphi)=\boldsymbol{\pi}_{\alpha}^{-1}(\pi,\varpi)=(\pi,\pi-\varpi). \] Note that in this case the reparameterization (via $\boldsymbol{\pi}_{\alpha}(\cdot,\cdot)$) does not depend on $\alpha$, and the same is true for the parameter space of $(\pi,\varpi)$. Verification of Assumption (ref) is straightforward using Lemma (ref) (for details, see Appendix B). In the LMAR case, the only benefit of the reparameterization is to transform the null hypothesis into a point null hypothesis.

\paragraph{GMAR Example. }

In the GMAR model our reparameterization is obtained by setting, for any fixed $\alpha\in A$, \[ (\pi,\varpi)=\boldsymbol{\pi}_{\alpha}(\phi,\varphi)=(\alpha\phi+(1-\alpha)\varphi,\phi-\varphi)\;\;\textrm{so that}\;\;(\phi,\varphi)=\boldsymbol{\pi}_{\alpha}^{-1}(\pi,\varpi)=(\pi+(1-\alpha)\varpi,\pi-\alpha\varpi). \] Verification of Assumption (ref) is again straightforward using Lemma (ref) (for details, see Appendix C). In the GMAR model, simplifying the null hypothesis is not the only benefit of the reparameterization, as will be discussed next.

As discussed before Assumption (ref), the relevant derivatives when expanding $L_{T}(\alpha,\beta,\phi,\varphi)$ are with respect to $(\beta,\phi,\varphi)$ and, in the GMAR case, these derivatives are linearly dependent under the null hypothesis. To see this and how the reparameterization affects this feature, note first that straightforward differentiation yields \[ \nabla_{(\beta,\phi,\varphi)}l_{t}(\alpha,\beta,\phi,\phi)=\left(\alpha\frac{\nabla_{\beta}f_{t}(\beta,\phi)}{f_{t}(\beta,\phi)},\alpha\frac{\nabla_{\phi}f_{t}(\beta,\phi)}{f_{t}(\beta,\phi)},(1-\alpha)\frac{\nabla_{\phi}f_{t}(\beta,\phi)}{f_{t}(\beta,\phi)}\right), \] where the null hypothesis $\phi=\varphi$ is imposed and $\nabla$ denotes differentiation with respect to the indicated parameters. As $(f_{t}(\beta,\phi))^{-1}\nabla_{(\beta,\phi)}f_{t}(\beta,\phi)$ is the score vector obtained from a linear Gaussian AR($p$) model, it is clear that the covariance matrix of the $(2p+3)$\textendash dimensional vector $\nabla_{(\beta,\phi,\varphi)}l_{t}(\alpha,\beta,\phi,\phi)$, and hence the (Fisher) information matrix, is singular with rank $p+2$. In contrast to the above, in the $\pi$\textendash parameterization the score vector is given by (see Supplementary Appendix C) \[ \nabla_{(\beta,\pi,\varpi)}l_{t}^{\pi}(\alpha,\beta,\pi,0)=\left(\frac{\nabla_{\beta}f_{t}(\beta,\pi)}{f_{t}(\beta,\pi)},\frac{\nabla_{\pi}f_{t}(\beta,\pi)}{f_{t}(\beta,\pi)},\mathbf{0}_{p+1}\right) \] when the null hypothesis $\varpi=0$ is imposed. Now the score of $\varpi$ is identically zero so that the reparameterization simplifies linear dependencies of the scores which turns out to be very useful in subsequent asymptotic analyses.

Quadratic expansion of the (reparameterized) log-likelihood function

As alluded to above, in standard testing problems the asymptotic analysis of a LR test statistic is based on a second order Taylor expansion of the (average) log-likelihood function around the true parameter value. An essential assumption here is positive definiteness of the (limiting) information matrix but, as illustrated in the previous section, this assumption does not necessarily hold in our testing problem due to linear dependencies among the derivatives of the log-likelihood function. As in rotnitzky2000likelihood, zhu2006asymptotics, and kasahara2012testing,kasahara2015testing, we therefore consider a quadratic expansion of the log-likelihood function that is not based on a second order Taylor expansion but (possibly) on a higher order Taylor expansion. The need for higher-order derivatives is illustrated by the GMAR example: as the score of $\varpi$ is identically zero, the second derivative (which turns out to be linearly independent of the score of $(\beta,\pi)$) now provides the first (nontrivial) local approximation for $\varpi$.

The following assumption ensures that the (reparameterized) log-likelihood function ((ref)) is (at least) twice continuously differentiable.

assumptionFor some integer $k\geq2$, and for every fixed $\alpha\in A$, the functions $\alpha_{t}(\alpha,\beta,\phi,\varphi)$ and $\boldsymbol{\pi}_{\alpha}^{-1}(\pi,\varpi)$ are $k$ times continuously differentiable (with respect to $(\beta,\phi,\varphi)$ and $(\pi,\varpi)$ in the interior of $B\times\Phi\times\Phi$ and $\Pi_{\alpha}$, respectively).

In our general framework the reparameterized log-likelihood function is assumed to have, for each $\alpha\in A$, a quadratic expansion in a transformed parameter vector $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)$ around $(\beta^{*},\pi^{\ast},0)$ given by

multline[multline omitted — 372 chars of source]

To illustrate this expansion, suppose the information matrix is positive definite so that the quantities on the right hand side are (typically) based on a second order Taylor expansion with $S_{T\alpha}$ and $\mathcal{I}_{\alpha}$ functions of $(\alpha,\beta^{*},\pi^{*},0)$. As already mentioned, this is the case for the LMAR model where (the following will be justified shortly) the parameter $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)$ is independent of $\alpha$ and given by $(\pi-\pi^{\ast},\varpi)$ and, for each $\alpha\in A$, $S_{T\alpha}$ is the score vector, $\mathcal{I}_{\alpha}$ is the (positive definite) Fisher information matrix, and $R_{T}(\alpha,\beta,\pi,\varpi)$ is a remainder term. As the notation indicates, these three terms depend on $\alpha$, and in general they may involve partial derivatives of the log-likelihood function of order higher than two (this is the case for the GMAR model, as will be demonstrated shortly). Then it may also get more complicated to find the reparameterization of the previous subsection and the transformed parameter vector $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)$, as the examples of Kasahara and Shimotsu (2015, 2017) and the discussion on the GMAR model below show; one possibility is to consider the iterative procedure discussed by rotnitzky2000likelihood (for a recent illuminating illustration of this approach, see hallin2014skew).

Our next assumption provides further details on expansion ((ref)). We use $\Rightarrow$ to signify weak convergence of a sequence of stochastic processes on a function space. In the assumption below, the weak convergence of interest is that of the process $S_{T\alpha}$ (indexed by $\alpha\in A$) to a process $S_{\alpha}$. The two function spaces relevant in this paper are $\mathcal{B}(A,\mathbb{R}^{k})$ and $\mathcal{C}(A,\mathbb{R}^{k})$, the former is the space of all $\mathbb{R}^{k}$\textendash valued bounded functions defined on (the compact set) $A$ equipped with the uniform metric ($d(x,y)=\sup_{a\in A}\left\Vert x(a)-y(a)\right\Vert $), and the latter is the same but with the continuity of the functions (with respect to $\alpha\in A$) also assumed.

assumptionFor each $\alpha\in A$, the log-likelihood function $L_{T}^{\pi}(\alpha,\beta,\pi,\varpi)$ has a quadratic expansion given in ((ref)), where \begin{description} • for each $\alpha\in A$, $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)$ is a mapping from $B\times\Pi_{\alpha}$ to\\ $\Theta_{\alpha}=\{\boldsymbol{\theta}\in\mathbb{R}^{r}:\boldsymbol{\theta}=\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)\text{ for some }(\beta,\pi,\varpi)\in B\times\Pi_{\alpha}\}$ such that (a) $\boldsymbol{\theta}(\alpha,\beta^{*},\pi^{*},0)=0$ and (b) for all $\epsilon>0$, $\inf_{\alpha\in A}\inf$$_{(\beta,\pi,\varpi)\in B\times\Pi_{\alpha}:\left\Vert (\beta,\pi,\varpi)-(\beta^{*},\pi^{\ast},0)\right\Vert \geq\epsilon}\left\Vert \boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)\right\Vert \geq\delta_{\epsilon}$ for some $\delta_{\epsilon}>0$. • $S_{T\alpha}=\sum_{t=1}^{T}s_{t\alpha}$ is a sequence of $\mathbb{R}^{r}$\textendash valued $\mathcal{F}_{T}$\textendash measurable stochastic processes indexed by $\alpha\in A$; $S_{T\alpha}$ does not depend on $(\beta,\pi,\varpi)$; $S_{T\alpha}$ has sample paths that are continuous as functions of $\alpha$; the process $T^{-1/2}S_{T\alpha}$ obeys $T^{-1/2}S_{T\bullet}\Rightarrow S_{\bullet}$ for some mean zero $\mathbb{R}^{r}$-valued Gaussian process $\{S_{\alpha}:\alpha\in A\}$ that satisfies $E[S_{\alpha}S_{\alpha}']=E[s_{t\alpha}s_{t\alpha}']=\mathcal{I}_{\alpha}$ for all $\alpha\in A$ and has continuous sample paths (as functions of $\alpha$) with probability 1. • $\mathcal{I}_{\alpha}$ is, for each $\alpha\in A$, a non-random symmetric $r\times r$ matrix (independent of $(\beta,\pi,\varpi)$); $\mathcal{I}_{\alpha}$ is continuous as a function of $\alpha$ and such that $0<\inf_{\alpha\in A}\lambda_{\min}(\mathcal{I}_{\alpha})$, $\sup{}_{\alpha\in A}\lambda_{\max}(\mathcal{I}_{\alpha})<\infty$. • $R_{T}(\alpha,\beta,\pi,\varpi)$ is a remainder term such that \[ \sup_{(\beta,\pi,\varpi)\in B\times\Pi_{\alpha}:\left\Vert (\beta,\pi,\varpi)-(\beta^{*},\pi^{\ast},0)\right\Vert \leq\gamma_{T}}\frac{\lvert R_{T}(\alpha,\beta,\pi,\varpi)\rvert}{(1+\lVert T^{1/2}\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)\rVert)^{2}}=o_{p\alpha}(1) \] for all sequences of (non-random) positive scalars $\{\gamma_{T},\ T\geq1\}$ for which $\gamma_{T}\rightarrow0$ as $T\rightarrow\infty$. \end{description}

Assumption (ref)(i) describes the transformed parameter $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)$, with part (b) being an identification condition. Assumption (ref)(ii) is the main ingredient needed to derive the limiting distribution of our LR test whereas (ref)(iv) ensures that the remainder term $R_{T}(\alpha,\beta,\pi,\varpi)$ has no effect on the final result. Assumption (ref)(iii) imposes rather standard conditions on the counterpart of the information matrix.

As in andrews1999estimation,andrews2001testing, zhu2006asymptotics, and kasahara2012testing,kasahara2015testing, for further developments it will be convenient to write the expansion ((ref)) in an alternative form as

multline[multline omitted — 395 chars of source]

where $Z_{T\alpha}=\mathcal{I}_{\alpha}^{-1}T^{-1/2}S_{T\alpha}$. Assumptions 5(ii) and (iii) imply the following facts (that will be justified in the proof of Lemma (ref) in Appendix A): $Z_{T\alpha}$ is $\mathcal{F}_{T}$\textendash measurable, independent of $(\beta,\pi,\varpi)$, continuous as a function of $\alpha$ with probability 1, and $Z_{T\bullet}\Rightarrow Z_{\bullet}$ where the mean zero $\mathbb{R}^{r}$\textendash valued Gaussian process $Z_{\alpha}=\mathcal{I}_{\alpha}^{-1}S_{\alpha}$ satisfies $E[Z_{\alpha}Z_{\alpha}']=\mathcal{I}_{\alpha}^{-1}$ for all $\alpha\in A$ and has continuous sample paths (as functions of $\alpha$) with probability 1.

Examples (continued)

\paragraph{LMAR Example.}

For the LMAR model, expansion ((ref)) (with the unnecessary $\beta$ being dropped everywhere) is obtained from a standard second-order Taylor expansion. Specifically, for an arbitrary fixed $\alpha\in A$, a standard second-order Taylor expansion of $L_{T}^{\pi}(\alpha,\pi,\varpi)=\sum_{t=1}^{T}l_{t}^{\pi}(\alpha,\pi,\varpi)$ around $\left(\pi^{\ast},0\right)$ with respect to the parameters $(\pi,\varpi)$ yields

align[align omitted — 357 chars of source]

where $(\dot{\pi},\dot{\varpi})$ denotes a point between $(\pi,\varpi)$ and $(\pi^{\ast},0)$, $\nabla_{(\pi,\varpi)}L_{T}^{\pi}(\alpha,\pi^{\ast},0)=\sum_{t=1}^{T}\nabla_{(\pi,\varpi)}l_{t}^{\pi}(\alpha,\pi^{\ast},0)$ and $\nabla_{(\pi,\varpi)(\pi,\varpi)'}^{2}L_{T}^{\pi}(\alpha,\pi,\varpi)=\sum_{t=1}^{T}\nabla_{(\pi,\varpi)(\pi,\varpi)'}^{2}l_{t}^{\pi}(\alpha,\pi,\varpi)$ (explicit expressions for the required derivatives are provided in Appendix B), and $\nabla$ and $\nabla^{2}$ denote first and second order differentiation with respect to the indicated parameters. Set $\boldsymbol{\theta}(\alpha,\pi,\varpi)=(\pi-\pi^{\ast},\varpi)=(\theta,\vartheta)$ and note that the parameter space $\Theta_{\alpha}=\Theta$ is independent of $\alpha$ and contains the origin, corresponding to the true model, as an interior point. Then define the vector $S_{T\alpha}$ and the matrix $\mathcal{I}_{\alpha}$ as\footnote{In what follows, $\nabla f_{t}(\cdot)$ denotes differentiation of $f_{t}(\cdot)$ in ((ref)) with respect to $\tilde{\phi}=(\tilde{\phi}_{0},\tilde{\phi}_{1},\ldots,\tilde{\phi}_{p},\tilde{\sigma}_{1}^{2})$.}

align[align omitted — 881 chars of source]

Adding and subtracting terms and reorganizing, expansion ((ref)) can be written as

align[align omitted — 354 chars of source]

with the remainder term

equation[equation omitted — 294 chars of source]

These equations yield the expansion ((ref)) in the case of the LMAR model. For details of verifying Assumptions (ref) and (ref), we refer to Appendix B.

As mentioned in the LMAR example of Section 3.1.1, the treatment of the special case where the mixing weight $\alpha_{t}^{L}$ is constant is more complicated than that of the `real' LMAR case. Indeed, replacing the mixing weight $\alpha_{t}^{L}$ by a constant in the preceding expression of the score vector $S_{T\alpha}$ immediately shows that the second-order Taylor expansion ((ref)) breaks down because, contrary to Assumption 5(iii), the components of $S_{T\alpha}$ are not linearly independent and, consequently, the Fisher information matrix $\mathcal{I}_{\alpha}$ is singular. Thus, a higher order Taylor expansion is needed to analyze the LR test statistic.

To give an idea of how one could proceed, we first note that the partial derivatives of the log-likelihood function behave in the same way as their counterparts in kasahara2015testing where mixtures of normal regression models (with constant mixing weights) are considered (see particularly the discussion following their Proposition 1). This is due to the fact that in the special case of constant mixing weights the LMAR model is obtained from the model considered in kasahara2015testing by replacing the exogenous regressors therein by lagged values of $y_{t}$. Thus, the arguments employed in that paper could be used to obtain the asymptotic distribution of the LR test statistic. Instead of a conventional second-order Taylor expansion this would require a more complicated reparameterization and an expansion based on partial derivatives of the log-likelihood function up to order eight. As most of the details appear very similar to those in kasahara2015testing we have preferred not to pursue this matter in this paper.

The preceding discussion means that, in the case of the LMAR model, time varying mixing weights are beneficial when the purpose is to derive a LR test for the adequacy of a single-regime model. A similar observation was made already by jeffries1998logistic. However, this does not happen in all mixture autoregressive models with time varying mixing weights, as the following discussion on the GMAR model demonstrates.

\paragraph{GMAR Example. }

As alluded to earlier, in the case of the GMAR model the expansion ((ref)) cannot be based on a second order Taylor expansion of the log-likelihood function. A higher order expansion is required, and similarly to kasahara2012testing the appropriate order turns out to be the fourth one with the elements of $\nabla_{\beta}l_{t}^{\pi}(\alpha,\beta^{*},\pi^{*},0)$ and $\nabla_{\pi}l_{t}^{\pi}(\alpha,\beta^{*},\pi^{*},0)$ and the distinctive elements of $\nabla_{\varpi\varpi^{\prime}}l_{t}^{\pi}(\alpha,\beta^{*},\pi^{*},0)$ (suitably normalized) used to define the vector $S_{T\alpha}$. In Appendix C we present, for an arbitrary fixed $\alpha\in A$, the explicit form of a standard fourth-order Taylor expansion of $L_{T}^{\pi}(\alpha,\beta,\pi,\varpi)=\sum_{t=1}^{T}l_{t}^{\pi}(\alpha,\beta,\pi,\varpi)$ around $(\beta^{*},\pi^{*},0)$ with respect to the parameters $(\beta,\pi,\varpi)$. Therein we also demonstrate that this fourth-order Taylor expansion can be written as a quadratic expansion of the form ((ref)) (or ((ref))) with the different quantities appearing therein defined as follows.

Define the vector $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)$ in ((ref)) as \[ \boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)=

bmatrix[bmatrix omitted — 82 chars of source]

=

bmatrix[bmatrix omitted — 77 chars of source]

, \] where $\theta(\alpha,\beta,\pi,\varpi)$ is $(q_{1}+q_{2})\times1$ and $\vartheta(\alpha,\beta,\pi,\varpi)$ is $q_{\vartheta}\times1$ with $q_{\vartheta}=q_{2}(q_{2}+1)/2$ (where $q_{1}=1$ and $q_{2}=p+1$) and where the vector $v(\varpi)$ contains the unique elements of $\varpi\varpi^{\prime}$, that is, \[ v(\varpi)=(\varpi_{1}^{2},\ldots,\varpi_{q_{2}}^{2},\varpi_{1}\varpi_{2},\ldots,\varpi_{1}\varpi_{q_{2}},\varpi_{2}\varpi_{3},\ldots,\varpi_{q_{2}-1}\varpi_{q_{2}}) \] (note that $v(\varpi)$ is just a re-ordering of $vech(\varpi\varpi^{\prime})$). The parameter space \[ \Theta_{\alpha}=\{\boldsymbol{\theta}=(\theta,\vartheta)\in\mathbb{R}^{q_{1}+q_{2}+q_{\vartheta}}:\theta=(\beta-\beta^{\ast},\pi-\pi^{\ast}),\vartheta=\alpha(1-\alpha)v(\varpi)\text{ for some }(\beta,\pi,\varpi)\in B\times\Pi_{\alpha}\} \] now depends on $\alpha$ and has the origin, corresponding to the true model, as a boundary point (due to the particular shape of the range of $\vartheta=\alpha(1-\alpha)v(\varpi)$); both of these features will complicate the subsequent analysis.

Next set

equation[equation omitted — 335 chars of source]

with the component vectors $\tilde{\nabla}_{\theta}l_{t}^{\pi\ast}$ and $\tilde{\nabla}_{\vartheta}l_{t}^{\pi\ast}$ given by

align[align omitted — 693 chars of source]

where $c_{ij}=1/2$ if $i=j$ and $c_{ij}=1$ if $i\neq j$. Explicit expressions for $\tilde{\nabla}_{\theta}l_{t}^{\pi\ast}$ and $\tilde{\nabla}_{\vartheta}l_{t}^{\pi\ast}$ can be found in Appendix C, and from them it can be seen that $S_{T}$ depends on $(\beta^{*},\pi^{\ast})$ only and not on $(\alpha,\beta,\pi,\varpi)$. The same is true for the matrix $\mathcal{I}\,(=\mathcal{I_{\alpha}})$ ($(q_{1}+q_{2}+q_{\vartheta})\times(q_{1}+q_{2}+q_{\vartheta})$) whose expression is also given in Appendix C. Finally, an explicit expression of the remainder term $R_{T}(\alpha,\beta,\pi,\varpi)$ is given in Appendix C. For the verification of Assumptions (ref) and (ref), see Appendix C.

In the GMAR example we have assumed that the intercept terms $\tilde{\phi}_{0}$ and $\tilde{\varphi}_{0}$ in the two regimes are the same. We are now in a position to describe the difficulties that allowing for $\tilde{\phi}_{0}\neq\tilde{\varphi}_{0}$ (and, hence, dropping $\beta$) would entail. In this case, the additional parameter $\tilde{\varphi}_{0}$ would correspond to $\varpi_{1}$, the first component of $\varpi$. As in Section 3.2.1, it would again be the case that $\nabla_{\varpi_{1}}l_{t}^{\pi}(\alpha,\pi^{*},0)=0$, leading us to consider second derivatives. But now, due to the properties of the Gaussian distribution, it would be the case that $\nabla_{\varpi_{1}\varpi_{1}}^{2}l_{t}^{\pi}(\alpha,\pi^{*},0)$ is linearly dependent with the components of $\nabla_{\pi}l_{t}^{\pi}(\alpha,\pi^{*},0)$, making it unsuitable to serve as a component of $S_{T}$. A reparameterization more complicated than that used in Section 3.2.1 would be needed, with the aim of obtaining $\nabla_{\varpi_{1}\varpi_{1}}^{2}l_{t}^{\pi}(\alpha,\pi^{*},0)=0$ and, instead of $\nabla_{\varpi_{1}\varpi_{1}}^{2}l_{t}^{\pi}(\alpha,\pi^{*},0)$, using $\nabla_{\varpi_{1}\varpi_{1}\varpi_{1}}^{3}l_{t}^{\pi}(\alpha,\pi^{*},0)$ or perhaps $\nabla_{\varpi_{1}\varpi_{1}\varpi_{1}\varpi_{1}}^{4}l_{t}^{\pi}(\alpha,\pi^{*},0)$ as the counterpart of the score of the parameter $\varpi_{1}$. It turns out that (restricting the discussion to the case $p=1$ only) the third derivative is suitable when $\alpha\neq1/2$ and $\tilde{\phi}_{1}\neq-1/2$, but that fourth (or higher) order derivatives are needed when $\alpha=1/2$ or $\tilde{\phi}_{1}=-1/2$. Similar difficulties (involving situations comparable to the cases $\alpha=1/2$ vs. $\alpha\neq1/2$, but apparently not ones involving also a counterpart of $\tilde{\phi}_{1}$) were faced by cho2007testing and Kasahara2017Testing, whose analyses suggest that expanding the log-likelihood at least to the eighth order is required. As the required analysis gets excessively complicated, we have chosen to leave it for future research.

Asymptotic analysis of the quadratic expansion

We continue by analyzing the expansion ((ref)) evaluated at $(\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})$. Previously, a similar analysis is provided by andrews2001testing but his approach is not directly applicable in our setting. The reason for this is that in the quadratic expansion in ((ref)) the dependence of the parameter $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)$ and its parameter space $\Theta_{\alpha}$ on the nuisance parameter $\alpha$ is not compatible with the formulation of andrews2001testing. The results of zhu2006asymptotics probably cover our case, but instead of trying to verify the assumptions employed by these authors we prove the needed results by adapting the arguments used in andrews1999estimation,andrews2001testing and zhu2006asymptotics to our setting. We proceed in several steps.

\paragraph{Asymptotic insignificance of the remainder term.}

We first establish that the remainder term $R_{T}(\alpha,\beta,\pi,\varpi)$, when evaluated at $(\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})$, has no effect on the asymptotic distribution of the quadratic expansion. A crucial ingredient in showing this is showing that the transformed parameter vector $\boldsymbol{\theta}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})$ is root-$T$ consistent in the sense that $\lVert T^{1/2}\boldsymbol{\theta}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})\rVert=O_{p\alpha}(1)$. This, together with part (iv) of Assumption (ref) allows us to obtain the result $R_{T}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})=o_{p\alpha}(1)$. We collect these results in the following lemma.

lemIf Assumptions 1\textendash 5 hold\,\footnote{Here and in what follows, a subset of the listed assumptions would sometimes suffice for the stated results.}, then (i) $\lVert T^{1/2}\boldsymbol{\theta}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})\rVert=O_{p\alpha}(1)$, \\ (ii) $R_{T}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})=o_{p\alpha}(1)$, and (iii) \begin{align} & L_{T}^{\pi}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})-L_{T}^{\pi}(\alpha,\beta^{*},\pi^{*},0)\nonumber \\ & =\frac{1}{2}Z_{T\alpha}^{\prime}\mathcal{I}_{\alpha}Z_{T\alpha}-\frac{1}{2}\bigl[T^{1/2}\boldsymbol{\theta}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})-Z_{T\alpha}\bigr]'\mathcal{I}_{\alpha}\bigl[T^{1/2}\boldsymbol{\theta}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})-Z_{T\alpha}\bigr]+o_{p\alpha}(1). \end{align}

Note that assertion (iii) of Lemma (ref) is analogous to andrews1999estimation.

\paragraph{Maximization of the likelihood vs. minimization of a related quadratic form.}

The first two terms on the right hand side of ((ref)) provide an approximation to the (reparameterized and centered) log-likelihood function $L_{T}^{\pi}(\alpha,\beta,\pi,\varpi)-L_{T}^{\pi}(\alpha,\beta^{*},\pi^{*},0)$ evaluated at the (approximate unrestricted reparameterized) ML estimator $(\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})$ in equation ((ref)). For later developments it would be convenient if the ML estimator on the right hand side of ((ref)) could be replaced by an (approximate) minimizer of the quadratic form $[T^{1/2}\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)-Z_{T\alpha}]'\mathcal{I}_{\alpha}[T^{1/2}\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)-Z_{T\alpha}]$. In order to justify this replacement we first note that, by definition, (cf. andrews1999estimation) \[ \inf_{(\beta,\pi,\varpi)\in B\times\Pi_{\alpha}}\bigl\{\bigl[T^{1/2}\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)-Z_{T\alpha}\bigr]'\mathcal{I}_{\alpha}\bigl[T^{1/2}\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)-Z_{T\alpha}\bigr]\bigr\}=\inf_{\boldsymbol{\lambda}\in\Theta_{\alpha,T}}\left\{ (\boldsymbol{\lambda}-Z_{T\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\lambda}-Z_{T\alpha})\right\} \] where, for each $T$, \[ \Theta_{\alpha,T}=\{\boldsymbol{\lambda}\in\mathbb{R}^{r}:\boldsymbol{\lambda}=T^{1/2}\boldsymbol{\theta}\text{ for some }\boldsymbol{\theta}\in\Theta_{\alpha}\} \] and $\Theta_{\alpha}$ is as defined in Assumption (ref)(i). Next, for each $\alpha\in A$, let $\boldsymbol{\hat{\lambda}}_{T\alpha q}=T^{1/2}\boldsymbol{\theta}(\alpha,\hat{\beta}_{T\alpha q},\hat{\pi}_{T\alpha q},\hat{\varpi}_{T\alpha q})$ (with the additional `$q$' in the subscripts referring to quadratic form) denote an approximate minimizer of $(\boldsymbol{\lambda}-Z_{T\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\lambda}-Z_{T\alpha})$ over $\Theta_{\alpha,T}$, that is, (cf. zhu2006asymptotics)

equation[equation omitted — 355 chars of source]

Now we can state the following lemma justifying the discussed replacement of $(\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})$ with $\boldsymbol{\hat{\lambda}}_{T\alpha q}$.

lemIf Assumptions 1\textendash 5 hold, then \[ L_{T}^{\pi}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})-L_{T}^{\pi}(\alpha,\beta^{*},\pi^{*},0)=\frac{1}{2}Z_{T\alpha}^{\prime}\mathcal{I}_{\alpha}Z_{T\alpha}-\frac{1}{2}(\boldsymbol{\hat{\lambda}}_{T\alpha q}-Z_{T\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\hat{\lambda}}_{T\alpha q}-Z_{T\alpha})+o_{p\alpha}(1). \]

\paragraph{Approximating the parameter space with a cone.}

In the previous subsection the quadratic form $(\boldsymbol{\lambda}-Z_{T\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\lambda}-Z_{T\alpha})$ was minimized over the set $\Theta_{\alpha,T}$ which can be complicated and hence difficult to use. Therefore we next show that the quadratic form $(\boldsymbol{\lambda}-Z_{T\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\lambda}-Z_{T\alpha})$ can instead be minimized over a simpler set, and to this end we first introduce some terminology.

We say that a collection of sets $\{\Gamma_{\alpha},\ \alpha\in A\}$ (where for each $\alpha\in A$, $\Gamma_{\alpha}\subset\mathbb{R}^{r}$) is `locally (at the origin) uniformly equal' to a set $\Lambda\subset\mathbb{R}^{r}$ if there exists a $\delta>0$ such that $\Gamma_{\alpha}\cap(-\delta,\delta)^{r}=\Lambda\cap(-\delta,\delta)^{r}$ for all $\alpha\in A$. Note that `$\{\Gamma_{\alpha},\ \alpha\in A\}$ is locally uniformly equal to $\Lambda$' implies that (i) `for all $\alpha\in A$, $\Gamma_{\alpha}$ is locally equal to $\Lambda$ in the sense of andrews1999estimation', but the reverse does not hold; and also that (ii) $\{\Gamma_{\alpha},\ \alpha\in A\}$ is uniformly approximated by the set $\Lambda$ in the sense of zhu2006asymptotics. Finally, we say that a set $\Lambda\subset\mathbb{R}^{r}$ is a `cone' if $\lambda\in\Lambda$ implies that $a\lambda\in\Lambda$ for all positive real scalars $a$.

Based on the preceding discussion we state the following assumption.

assumptionThe collection of sets $\{\Theta_{\alpha},\ \alpha\in A\}$ is locally uniformly equal to a cone $\Lambda\;(\subset\mathbb{R}^{r})$.

Note that by Assumption (ref)(i)(a), $0\in\Theta_{\alpha}$ for all $\alpha\in A$, so that the cone $\Lambda$ in Assumption (ref) necessarily contains $0\;(\in\mathbb{R}^{r})$. The cone $\Lambda$ also does not depend on $\alpha.$ Now we can establish the following result.

lemIf Assumptions 1\textendash 6 hold, then \[ \inf_{\boldsymbol{\lambda}\in\Theta_{\alpha,T}}\left\{ (\boldsymbol{\lambda}-Z_{T\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\lambda}-Z_{T\alpha})\right\} =\inf_{\boldsymbol{\lambda}\in\Lambda}\left\{ (\boldsymbol{\lambda}-Z_{T\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\lambda}-Z_{T\alpha})\right\} +o_{p\alpha}(1). \]

\paragraph{Describing the limiting random variable.}

From Lemmas (ref) and (ref) and the definition of $\boldsymbol{\hat{\lambda}}_{T\alpha q}$ we can now conclude that

equation[equation omitted — 393 chars of source]

The assumed weak convergence of $S_{T\alpha}$ (and hence that of $Z_{\alpha}=\mathcal{I}_{\alpha}^{-1}S_{\alpha}$) allows us to derive the weak limit of this random process described in the following lemma.

lemIf Assumptions 1\textendash 6 hold, then \[ 2[L_{T}^{\pi}(\bullet,\hat{\beta}_{T\bullet},\hat{\pi}_{T\bullet},\hat{\varpi}_{T\bullet})-L_{T}^{\pi}(\bullet,\beta^{*},\pi^{\ast},0)]\Rightarrow Z_{\bullet}^{\prime}\mathcal{I}_{\bullet}Z_{\bullet}-\inf_{\boldsymbol{\lambda}\in\Lambda}\left\{ (\boldsymbol{\lambda}-Z_{\bullet})^{\prime}\mathcal{I}_{\bullet}(\boldsymbol{\lambda}-Z_{\bullet})\right\} . \]

The limiting random process in Lemma (ref) can be written in a somewhat simpler form (cf. andrews1999estimation and andrews2001testing). The motivation for this comes from the fact that in our applications $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)$ can be decomposed into two parts as $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)=(\theta(\alpha,\beta,\pi,\varpi),\vartheta(\alpha,\beta,\pi,\varpi))$ with $\theta(\alpha,\beta,\pi,\varpi)\in\mathbb{R}^{q_{\theta}}$ and $\vartheta(\alpha,\beta,\pi,\varpi)\in\mathbb{R}^{q_{\vartheta}}$ (with $q_{\theta}=q_{1}+q_{2}$ and $r=q_{\theta}+q_{\vartheta}$) such that (i) the values of $\theta(\alpha,\beta,\pi,\varpi)$ are not restricted by the null hypothesis and do not lie on the boundary of the parameter space and (ii) the values of $\vartheta(\alpha,\beta,\pi,\varpi)$ are restricted by the null hypothesis and potentially lie on the boundary of the parameter space. Specifically, we assume the following.

assumptionThe cone $\Lambda$ of Assumption (ref) satisfies $\Lambda=\mathbb{R}^{q_{\theta}}\times\Lambda_{\vartheta}$ with $\Lambda_{\vartheta}$ a cone in $\mathbb{R}^{q_{\vartheta}}$.

Partition $S_{\alpha}$, $Z_{\alpha}$, $\boldsymbol{\lambda}$, and $\mathcal{I}_{\alpha}$ conformably with the partition $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)=(\theta(\alpha,\beta,\pi,\varpi),\vartheta(\alpha,\beta,\pi,\varpi))$ as \[ S_{\alpha}=\left[

array[array omitted — 54 chars of source]

\right],\,\,\,Z_{\alpha}=\left[

array[array omitted — 54 chars of source]

\right],\,\,\,\boldsymbol{\lambda}=

bmatrix[bmatrix omitted — 78 chars of source]

,\,\,\,\mathcal{I}_{\alpha}=\left[

array[array omitted — 166 chars of source]

\right] \] and let $(\mathcal{I}_{\alpha}^{-1})_{\vartheta\vartheta}$ denote the ($q_{\vartheta}\times q_{\vartheta}$) bottom right block of $\mathcal{I}_{\alpha}^{-1}$. Assumption (ref) together with properties of partitioned matrices yields the following result.

lemIf Assumptions 1\textendash 7 hold, then \begin{align*} & Z_{\alpha}'\mathcal{I}_{\alpha}Z_{\alpha}-\inf_{\boldsymbol{\lambda}\in\Lambda}\left\{ (\boldsymbol{\lambda}-Z_{\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\lambda}-Z_{\alpha})\right\} \\ & =Z_{\vartheta\alpha}'(\mathcal{I}_{\alpha}^{-1})_{\vartheta\vartheta}^{-1}Z_{\vartheta\alpha}-\inf_{\boldsymbol{\lambda}_{\vartheta}\in\Lambda_{\vartheta}}\left\{ (\boldsymbol{\lambda}_{\vartheta}-Z_{\vartheta\alpha})'(\mathcal{I}_{\alpha}^{-1})_{\vartheta\vartheta}^{-1}(\boldsymbol{\lambda}_{\vartheta}-Z_{\vartheta\alpha})\right\} +S_{\theta\alpha}'\mathcal{I}_{\theta\theta\alpha}^{-1}S_{\theta\alpha}. \end{align*}

Explicit expressions for $(\mathcal{I}_{\alpha}^{-1})_{\vartheta\vartheta}$ and $Z_{\vartheta\alpha}$ in terms of $S_{\alpha}$ and $\mathcal{I}_{\alpha}$ are given in the proof of this lemma in Supplementary Appendix D.

The LR test statistic

Derivation of the test statistic

The previous subsection described the asymptotic behavior of $2[L_{T}^{\pi}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})-L_{T}^{\pi}(\alpha,\beta^{*},\pi^{*},0)]$, the first term in the expression of $LR_{T}(\alpha)$ in ((ref)). Now consider the second term, namely $2[L_{T}^{0}(\hat{\tilde{\phi}}_{T})-L_{T}^{0}(\tilde{\phi}^{\ast})]$, corresponding to the model restricted by the null hypothesis. Recall that $L_{T}^{0}(\tilde{\phi})=\sum_{t=1}^{T}l_{t}^{0}(\tilde{\phi})$ with $l_{t}^{0}(\tilde{\phi})=\log[f_{t}(\tilde{\phi})]$ so that $\nabla_{\tilde{\phi}}l_{t}^{0}(\tilde{\phi}^{\ast})=(f_{t}(\tilde{\phi}^{*}))^{-1}\nabla f_{t}(\tilde{\phi}^{*})$ with $\tilde{\phi}^{*}$ an interior point of $\tilde{\Phi}$. Denote the score vector and limiting information matrix by \[ S_{T}^{0}=\sum_{t=1}^{T}\frac{\nabla f_{t}(\tilde{\phi}^{\ast})}{f_{t}(\tilde{\phi}^{\ast})},\,\,\,\,\,\mathcal{I}^{0}=E\biggl[\frac{\nabla f_{t}(\tilde{\phi}^{\ast})}{f_{t}(\tilde{\phi}^{\ast})}\frac{\nabla'f_{t}(\tilde{\phi}^{\ast})}{f_{t}(\tilde{\phi}^{\ast})}\biggr], \] respectively. For the following assumption, partition the process $S_{T\alpha}$ of Assumption (ref) as $S_{T\alpha}=(S_{T\theta\alpha},S_{T\vartheta\alpha})$ (with $S_{T\theta\alpha}$ $q_{\theta}$\textendash dimensional and $S_{T\vartheta\alpha}$ $q_{\vartheta}$\textendash dimensional). The following simplifying assumption, which holds in our examples (see the expressions of $S_{T\alpha}$ in ((ref)) and ((ref))), allows us to obtain a neat expression for the likelihood ratio test statistic in Theorem 1 below.

assumption$S_{T\theta\alpha}=S_{T}^{0}$.

Together with the earlier assumptions, Assumption (ref) implies that $T^{-1/2}S_{T\theta\alpha}=T^{-1/2}S_{T}^{0}\stackrel{d}{\to}S^{0}$, a $q_{\theta}$\textendash dimensional Gaussian random vector with mean zero and covariance matrix $\mathcal{I}_{\theta\theta\alpha}=E[S^{0}S^{0\prime}]=\mathcal{I}^{0}$. Standard likelihood theory now implies the following result.

lemIf Assumptions 1\textendash 8 hold, then $2[L_{T}^{0}(\hat{\tilde{\phi}}_{T})-L_{T}^{0}(\tilde{\phi}^{\ast})]\overset{d}{\rightarrow}S^{0\prime}(\mathcal{I}^{0})^{-1}S^{0}$, and the convergence is joint with that in Lemma (ref).

The preceding results, in particular Lemmas (ref), (ref), and (ref), now yield the distribution of the LR test statistic in the following theorem.

thmIf Assumptions 1\textendash 8 hold, then \begin{description} • $LR_{T}\left(\bullet\right)\Rightarrow Z_{\vartheta\bullet}'(\mathcal{I}_{\bullet}^{-1})_{\vartheta\vartheta}^{-1}Z_{\vartheta\bullet}-\inf_{\boldsymbol{\lambda}_{\vartheta}\in\Lambda_{\vartheta}}\{(\boldsymbol{\lambda}_{\vartheta}-Z_{\vartheta\bullet})'(\mathcal{I}_{\bullet}^{-1})_{\vartheta\vartheta}^{-1}(\boldsymbol{\lambda}_{\vartheta}-Z_{\vartheta\bullet})\}$, and • $LR_{T}=\sup_{\alpha\in A}LR_{T}\left(\alpha\right)\overset{d}{\rightarrow}\sup_{\alpha\in A}\bigl\{ Z_{\vartheta\alpha}'(\mathcal{I}_{\alpha}^{-1})_{\vartheta\vartheta}^{-1}Z_{\vartheta\alpha}-\inf_{\boldsymbol{\lambda}_{\vartheta}\in\Lambda_{\vartheta}}\{(\boldsymbol{\lambda}_{\vartheta}-Z_{\vartheta\alpha})'(\mathcal{I}_{\alpha}^{-1})_{\vartheta\vartheta}^{-1}(\boldsymbol{\lambda}_{\vartheta}-Z_{\vartheta\alpha})\}\bigr\}$. \end{description}

This completes the derivation of the LR test statistic. The asymptotic distribution is similar to that in andrews2001testing. As we next discuss, this distribution simplifies in both the LMAR and the GMAR examples.

Examples (continued)

\paragraph{LMAR Example. }

As was noted in Section 3.3.1, the LMAR case is rather standard in the sense that a conventional second-order Taylor expansion with a nonsingular information matrix and with no parameters on the boundary was sufficient to study the LR test. The only nonstandard feature in this case is the presence of unidentified parameters under the null hypothesis. Validity of Assumptions (ref)\textendash (ref) is easy to check (see Appendix B) with the cone $\Lambda$ of Assumption (ref) equal to $\mathbb{R}^{r}$. Thus the infimum in the distribution of the $LR_{T}$ statistic in Theorem 1(ii) equals zero and the result therein simplifies to\footnote{This result could also be obtained from andrews1995admissibility as their assumptions 1\textendash 5 appear to be satisfied in the LMAR case.} \[ LR_{T}=\sup_{\alpha\in A}LR_{T}\left(\alpha\right)\overset{d}{\rightarrow}\sup_{\alpha\in A}\left\{ Z_{\vartheta\alpha}'(\mathcal{I}_{\alpha}^{-1})_{\vartheta\vartheta}^{-1}Z_{\vartheta\alpha}\right\} . \] For every fixed $\alpha\in A$, the quantity $Z_{\vartheta\alpha}'(\mathcal{I}_{\alpha}^{-1})_{\vartheta\vartheta}^{-1}Z_{\vartheta\alpha}$ is a chi-squared random variable, so that the limiting distribution is a supremum of a chi-squared process similarly as in, for example, davies1987hypothesis, hansen1996inference, and andrews2001testing.

\paragraph{GMAR Example. }

In Section 3.3.1 it was seen that in the GMAR example $Z_{\alpha}$ and $\mathcal{I}_{\alpha}$ do not depend on $\alpha$. As the cone $\Lambda$ of Assumption (ref) does not depend on $\alpha$ either, the weak limit of $LR_{T}\left(\alpha\right)$ does not depend on $\alpha$. Therefore the result of Theorem 1 (validity of Assumptions (ref)\textendash (ref) is checked in Appendix C) simplifies to \[ LR_{T}=\sup_{\alpha\in A}LR_{T}\left(\alpha\right)\overset{d}{\rightarrow}Z_{\vartheta}'(\mathcal{I}^{-1})_{\vartheta\vartheta}^{-1}Z_{\vartheta}-\inf_{\boldsymbol{\lambda}_{\vartheta}\in\Lambda_{\vartheta}}\left\{ (\boldsymbol{\lambda}_{\vartheta}-Z_{\vartheta})'(\mathcal{I}^{-1})_{\vartheta\vartheta}^{-1}(\boldsymbol{\lambda}_{\vartheta}-Z_{\vartheta})\right\} , \] where the unnecessary $\alpha$ has been dropped from the notation. Here $Z_{\vartheta}$ follows an $q_{\vartheta}$-variate Gaussian distribution with covariance matrix $(\mathcal{I}^{-1})_{\vartheta\vartheta}$, and the limiting distribution, which is sometimes referred to as the chi-bar-squared distribution, is similar to the one in kasahara2012testing. Note that the cone $\Lambda_{\vartheta}=v(\mathbb{\mathbb{R}}^{q_{2}})$ (see Appendix C) is not convex (in contrast to (at least most of) the examples in andrews2001testing, but similarly to kasahara2012testing) and the dimension of this cone, $q_{\vartheta}=q_{2}(q_{2}+1)/2$, may not be small either ($q_{\vartheta}=3,6,10,\ldots$ for $q_{2}=2,3,4,\ldots$).

Simulation-based critical values and a Monte Carlo study

Simulating the asymptotic null distribution

Similarly to hansen1996inference and andrews2001testing, the asymptotic null distribution of the LR statistic in Theorem 1 is typically application-specific and cannot be tabulated. Following these papers, we use simulation methods to obtain critical values of the asymptotic null distribution. The following procedure is based on hansen1996inference and is analogous to the one used by zhu2004hypothesis in a related mixture setting.\footnote{An alternative to this procedure is to use bootstrap. However, the validity of bootstrap in the presence of parameters on the boundary and singular information matrices is not clear (see, e.g., andrews2000inconsistency). Another reason for preferring the proposed simulation method is that repeated estimation of the mixture model under the alternative may be computationally rather demanding.}

Let $A_{G}$ be some finite grid of $\alpha$ values in $A$. For each fixed $\alpha\in A_{G}$, let $\hat{s}_{t\alpha}$ signify an empirical counterpart of $s_{t\alpha}$ (see Assumption 5) where the unknown parameter $\tilde{\phi}^{*}$ (or $(\beta^{*},\pi^{*})$) is replaced by its consistent estimator under the null, $\hat{\tilde{\phi}}_{T}$. (The specific forms of $\hat{s}_{t\alpha}$ in the LMAR and GMAR examples are provided in Appendices B and C, respectively.) Set $\hat{\mathcal{I}}_{T\alpha}=T^{-1}\sum_{t=1}^{T}\hat{s}_{t\alpha}\hat{s}_{t\alpha}'$. Now, for each $j=1,\ldots,J$ (where $J$ denotes the number of repetitions), do the following.

description• Generate a sequence $\{v_{tj}\}_{t=1}^{T}$ of $T$ i.i.d. $N(0,1)$ random variables. • For each $\alpha\in A_{G}$, set $\hat{S}_{T\alpha}^{j}=\sum_{t=1}^{T}\hat{s}_{t\alpha}v_{tj}$, $\hat{Z}_{T\alpha}^{j}=\hat{\mathcal{I}}_{T\alpha}^{-1}T^{-1/2}\hat{S}_{T\alpha}^{j}$, and (using similar partitioning notation as before) \[ \widehat{LR}_{T}^{j}(\alpha)=\hat{Z}_{T\vartheta\alpha}^{j\prime}(\hat{\mathcal{I}}_{T\alpha}^{-1})_{\vartheta\vartheta}^{-1}\hat{Z}_{T\vartheta\alpha}^{j}-\inf_{\boldsymbol{\lambda}_{\vartheta}\in\Lambda_{\vartheta}}\bigl\{(\boldsymbol{\lambda}_{\vartheta}-\hat{Z}_{T\vartheta\alpha}^{j})'(\hat{\mathcal{I}}_{T\alpha}^{-1})_{\vartheta\vartheta}^{-1}(\boldsymbol{\lambda}_{\vartheta}-\hat{Z}_{T\vartheta\alpha}^{j})\bigr\}; \] here the minimization of the quadratic form over the cone $\Lambda_{\vartheta}$ has to be performed numerically. • Set $\widehat{LR}_{T,A_{G}}^{j}=\max_{\alpha\in A_{G}}\widehat{LR}_{T}^{j}(\alpha)$.

This yields a sample $\{\widehat{LR}_{T,A_{G}}^{1},\ldots,\widehat{LR}_{T,A_{G}}^{J}\}$ of $J$ realizations. An approximate $p$\textendash value corresponding to an observed LR test statistic $LR_{T}$ is computed as $J^{-1}\sum_{j=1}^{J}\mathbf{1}(\widehat{LR}_{T,A_{G}}^{j}>LR_{T})$ (here $\mathbf{1}(\cdot)$ denotes the indicator function). The precision of this approximation can be controlled by choosing $J$ large enough, see hansen1996inference (in the illustration below we use $J=1000$).

A small Monte Carlo study

We now study the finite sample properties of the LR test statistics and the simulation-based critical values. The results are presented in Table 1. We consider two LR test statistics, one based on an estimated LMAR model, and another based on an estimated GMAR model (as in our two examples in the preceding sections). In all simulations, we use an autoregressive order $p=1$, $J=1000$ repetitions (see the previous subsection), and three different sample sizes: $T=250$, $500$, and $1000$.

The top part of Table 1 presents results for size simulations. Data is generated from an AR(1) model (for a range of different parameter values shown in Table 1) and AR(1), LMAR(1), and GMAR(1) models are estimated (LMAR with $m=1$; GMAR with the restriction $\tilde{\phi}_{0}=\tilde{\varphi}_{0}$; in estimation of the mixture models we use a genetic algorithm as singularity of the information matrix may render gradient based methods unreliable). Two LR test statistics are calculated based on the estimated LMAR and GMAR models, respectively, and labelled `LMAR $LR_{T}$' and `GMAR $LR_{T}$'. Simulation-based $p$\textendash values are computed based on the asymptotic distributions in Section 3.5.2 and using the simulation procedure in Section 4.1. Using nominal levels $10\%$, $5\%$, and $1\%$, a reject/not-reject decision is recorded. This exercise is repeated $1000$ times, and the six rightmost columns in Table 1 present the empirical rejection frequencies (for the LMAR $LR_{T}$ and GMAR $LR_{T}$ tests and the three nominal levels used).

As can be seen from the results in Table 1 (top part), the LMAR $LR_{T}$ test's size is satisfactory overall, typically being somewhat oversized for sample sizes $T=250$ and $500$, and somewhat conservative for the largest sample size ($T=1000$). The parameter values used in simulation do not seem to have a large effect on the size. The GMAR $LR_{T}$ test, on the other hand, appears to be moderately oversized across all sample sizes and parameter values used.

The lower part of Table 1 presents results for power simulations. Data is generated either from a GMAR model or from an LMAR model (for a range of different parameter values shown in Table 1), and empirical rejection frequencies are calculated as above. Both the LMAR $LR_{T}$ test and the GMAR $LR_{T}$ test appear to have good overall power. As expected, when the two regimes differ more from each other, the tests have higher power, and the same happens when sample size is increased. Besides having good power against the `right' alternatives, the tests also turn out to have decent power against `wrong' alternatives: When data is generated from the GMAR (resp., LMAR) model, the LMAR $LR_{T}$ (resp., GMAR $LR_{T}$) test rejects reasonably often (the GMAR $LR_{T}$ test in particular seems capable of picking up LMAR type regime switching). Naturally, the power of the tests may be inflated due to the tests being oversized.

As a computational remark we note that the LMAR $LR_{T}$ and GMAR $LR_{T}$ tests and their $p$\textendash values are reasonably straightforward to compute in a matter of seconds using a standard, modern desktop computer (for one particular model, one particular sample size, and $J=1000$ repetitions). The GMAR $LR_{T}$ test is computationally more demanding than the LMAR $LR_{T}$ test as it involves the minimization of a quadratic form over a cone which is not needed in the LMAR case (see Sections 3.5.2 and 4.1); this is also one potential reason for the less precise size of the GMAR $LR_{T}$ test.

table[table omitted — 69 chars of source]

Conclusions

This paper has studied the asymptotic distribution of the LR test statistic for testing a linear autoregressive model against a two-regime mixture autoregressive model. A distinguishing feature of the paper is that the regime switching probabilities are observation-dependent. Technical challenges resulting from unidentified parameters under the null, parameters on the boundary, and singularity of the information matrix were dealt with by considering an appropriately reparameterized model and higher-order expansions of the log-likelihood function. The resulting asymptotic distribution of the LR test statistic is non-standard and application-specific. Critical values can be obtained by a straightforward simulation procedure, and a Monte Carlo study indicated the proposed tests to have satisfactory size and power properties.

The general theory of the paper was illustrated using two concrete examples, the LMAR model of wong2001logistic and (a version of the) GMAR model of kalliovirta2015gaussian. Considering other mixture AR models, as well as the general GMAR model, is left for future research. This paper was concerned with testing linearity against a two-regime model, and considering tests of $M\geq2$ regimes versus $M+1$ regimes, similarly as in kasahara2015testing in a related setting, forms another interesting research topic.