EconBase
← Back to paper

Testing for observation-dependent regime switching in mixture autoregressive models

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

95,791 characters

Testing for observation-dependent regime switching in mixture autoregressive models



\title{\vspace{20pt}
Testing for observation-dependent regime switching\\
in mixture autoregressive models\thanks{The authors thank the Academy of Finland for financial support. Contact
addresses: Mika Meitz, Department of Political and Economic Studies,
University of Helsinki, P. O. Box 17, FI\textendash 00014 University
of Helsinki, Finland; e-mail: [email removed]. Pentti Saikkonen,
Department of Mathematics and Statistics, University of Helsinki,
P. O. Box 68, FI\textendash 00014 University of Helsinki, Finland;
e-mail: [email removed].}\vspace{20pt}
}

\author{Mika Meitz\\\small{University of Helsinki} \and Pentti Saikkonen\\\small{University of Helsinki}\vspace{20pt}
}

\date{October 2017}
\maketitle
\begin{abstract}
\noindent Testing for regime switching when the regime switching
probabilities are specified either as constants (`mixture models')
or are governed by a finite-state Markov chain (`Markov switching
models') are long-standing problems that have also attracted recent
interest. This paper considers testing for regime switching when the
regime switching probabilities are time-varying and depend on observed
data (`observation-dependent regime switching'). Specifically, we
consider the likelihood ratio test for observation-dependent regime
switching in mixture autoregressive models. The testing problem is
highly nonstandard, involving unidentified nuisance parameters under
the null, parameters on the boundary, singular information matrices,
and higher-order approximations of the log-likelihood. We derive the
asymptotic null distribution of the likelihood ratio test statistic
in a general mixture autoregressive setting using high-level conditions
that allow for various forms of dependence of the regime switching
probabilities on past observations, and we illustrate the theory using
two particular mixture autoregressive models. The likelihood ratio
test has a nonstandard asymptotic distribution that can easily be
simulated, and Monte Carlo studies show the test to have satisfactory
finite sample size and power properties.

\bigskip{}
\bigskip{}
\bigskip{}

\noindent \textbf{JEL classification:} C12, C22, C52.

\bigskip{}

\noindent \textbf{Keywords:} Likelihood ratio test, singular information
matrix, higher-order approximation of the log-likelihood, logistic
mixture autoregressive model, Gaussian mixture autoregressive model.
\end{abstract}
\vfill{}
\pagebreak{}

\section{Introduction}

Different regime switching models are in widespread use in economics,
finance, and other fields. When the regime switching probabilities
are constants, these models are often referred to as `mixture models',
and when these probabilities depend on past regimes and are governed
by a finite-state Markov chain, the term (time homogeneous) `Markov
switching models' is typically used. In this paper, we are interested
in the case where the regime switching probabilities depend on observed
data, a case we refer to as `observation-dependent regime switching'.
Models of this kind can be viewed as special cases of time inhomogeneous
Markov switching models (in which regime switching probabilities depend
on both past regimes and observed data). Overviews of regime switching
models can be found, for example, in \citet{fruhwirth2006finite}
and \citet{hamilton2016macroeconomic}. Of critical interest in all
these models is whether the use of several regimes is warranted or
if a single-regime model would suffice. Testing for regime switching
in all these models is plagued by several irregular features such
as unidentified parameters and parameters on the boundary and is consequently
notoriously difficult.

Tests for Markov switching have been considered by several authors
in the econometrics literature. \citet{hansen1992likelihood} and
\citet{garcia1998asymptotic} both considered sup-type likelihood
ratio (LR) tests in Markov switching models but they did not present
complete solutions. Hansen derived a bound for the distribution of
the LR statistic, leading to a conservative procedure, while Garcia
did not treat all the non-standard features of the problem in detail.
\citet{cho2007testing} analyzed the use of a LR statistic for a mixture
model to test for Markov-switching type regime switching. They found
their test based on a mixture model to have power against Markov switching
alternatives even though it ignores the temporal dependence of the
Markov chain. \citet*{carrasco2014optimal} took a different approach
and proposed an information matrix type test that they showed to be
asymptotically optimal against Markov switching alternatives. Very
recently, both \citet{Qu2017Likelihood} and \citet{Kasahara2017Testing}
have studied the LR statistic for regime switching in Markov switching
models.

Regarding testing for mixture type regime switching, the existing
literature is extensive, and several early references can be found,
for instance, in \citet[Sec. 6.5.1]{mclachlan2000finite}. Most papers
in this literature consider the case of independent observations without
regressors. Notable exceptions allowing for regressors (but not dependent
data) and having set-ups closer to the present paper are \citet{zhu2004hypothesis,zhu2006asymptotics}
and \citet{kasahara2015testing} who consider (among other things)
LR tests for regime switching. Further comparison to these works will
be provided in later sections.

In contrast to testing for Markov switching or mixture type regime
switching, there exists almost no literature on testing for observation-dependent
regime switching. The only two exceptions we are aware of are the
unpublished PhD thesis of \citet{jeffries1998logistic} and the recent
paper of \citet{shen2015inference}. Jeffries's thesis, which appears
to have gone largely unnoticed, analyses the LR test in a specific
(first-order) mixture autoregressive model; we will discuss his work
further in later sections. \citet{shen2015inference} consider the
case of independent observations with regressors and observation-dependent
regime switching, and propose an `expectation maximization test' for
regime switching.

In this paper we consider testing for observation-dependent regime
switching in a time series context. Specifically, we analyze the asymptotic
distribution of the LR test statistic for testing a linear autoregressive
model against a two-regime mixture autoregressive model with observation-dependent
regime switching. Mixture autoregressive models have been discussed
for instance in \citet{wong2000mixture,wong2001logistic}, \citet*{dueker2007contemporaneous},
\citet*{dueker2011multivariate}, and \citet*{kalliovirta2015gaussian,kalliovirta2016gaussian};
further discussion of this previous work will be provided in Section
2. Motivation for allowing the regime switching probabilities to depend
on observed data stems, for instance, from the desire to associate
changes in regime to observable economic variables.\footnote{See, for instance, \citet{hamilton2016macroeconomic}, whose \emph{Handbook
of Macroeconomics} chapter begins ``Many economic time series exhibit
dramatic breaks associated with events such as economic recessions,
financial panics, and currency crises. Such changes in regime may
arise from tipping points or other nonlinear dynamics and are core
to some of the most important questions in macroeconomics.'' }\footnote{More general models in which the regime switching probabilities are
allowed to depend on both past regimes and observable variables have
also been considered, see, e.g., \citet*{diebold1994regime}, \citet{filardo1994business},
and \citet*{kim2008estimation}.} Following \citet{kasahara2015testing} it would also be possible
to consider the more general testing problem that in a model with
more than two regimes the number of regimes can be reduced. However,
as even the case of two regimes is quite complex in our set-up, we
leave this extension to future research.

We consider mixture autoregressive (MAR) models in a rather general
setting employing high-level conditions that allow for various forms
of observation-dependent regime switching. As specific examples, we
treat the so-called logistic MAR (LMAR) model of \citet{wong2001logistic}
and (a version of the) Gaussian MAR (GMAR) model of \citet{kalliovirta2015gaussian}
in detail. The technical challenges we face in analyzing the LR test
statistic are similar to those when testing for Markov switching and
mixture type regime switching. First, there are nuisance parameters
that are unidentified under the null hypothesis. This is the classical
\citet{davies1977hypothesis,davies1987hypothesis} type problem. Second,
under the null hypothesis, there are parameters on the boundary of
the permissible parameter space. Such problems (also allowing for
unidentified nuisance parameters under the null) are discussed in
\citet{andrews1999estimation,andrews2001testing}. Third, the Fisher
information matrix is (potentially) singular, preventing the use of
conventional second-order expansions of the log-likelihood to analyze
the LR test statistic. Such problems are discussed by \citet*{rotnitzky2000likelihood},
and suitable reparameterizations and higher-order expansions are needed
to analyze the LR statistic. A particular challenge in the present
paper is to deal with these three problems simultaneously. Similar
problems were faced by \citet{kasahara2015testing}, and inspired
by their work we consider a suitably reparameterized model, write
a higher-order expansion of the log-likelihood function as a quadratic
function of the new parameters, and then derive the asymptotic distribution
of the LR test statistic by slightly extending and adapting the arguments
of \citet{andrews1999estimation,andrews2001testing} and \citet{zhu2006asymptotics}
(who partially generalize results of Andrews). Our two examples demonstrate
that, compared to the mixture type regime switching considered by
\citet{kasahara2015testing}, observation-dependent regime switching
can either simplify or complicate the analysis of the LR test statistic.

We contribute to the literature in several ways. (1) To the best of
our knowledge, apart from the unpublished PhD thesis of \citet{jeffries1998logistic},
we are the first to study testing for observation-dependent regime
switching using the LR test statistic and among the rather few to
allow for dependent observations. (2) We provide a general framework
to cover various forms of observation-dependent regime switching,
making our results potentially applicable to several models not explicitly
discussed in the present paper. (3) From a methodological perspective,
we slightly extend and adapt certain arguments of \citet{andrews1999estimation,andrews2001testing}
and \citet{zhu2006asymptotics}, which may be of independent interest.

The rest of the paper is organized as follows. Section 2 reviews mixture
autoregressive models. Section 3 analyzes the LR test statistic for
testing a linear autoregressive model against a two-regime mixture
autoregressive model. Simulation-based critical values and a Monte
Carlo study are discussed in Section 4, and Section 5 concludes. Appendices
A\textendash C contain technical details and proofs. Supplementary
Appendices D\textendash E, available from the authors upon request,
contain further technical details omitted from the paper.

Finally, a few notational conventions are given. All vectors will
be treated as column vectors and, for the sake of uncluttered notation,
we shall write $x=(x_{1},\ldots,x_{n})$ for the (column) vector $x$
where the components $x_{i}$ may be either scalars or vectors (or
both). For any vector or matrix $x$, the Euclidean norm is denoted
by $\left\Vert x\right\Vert $. We let $X_{T\alpha}=o_{p\alpha}(1)$
and $X_{T\alpha}=O{}_{p\alpha}(1)$ stand for $\sup_{\alpha\in A}\left\Vert X_{T\alpha}\right\Vert =o_{p}(1)$
and $\sup_{\alpha\in A}\left\Vert X_{T\alpha}\right\Vert =O{}_{p}(1)$,
respectively, and $\lambda_{\min}(\cdot)$ and $\lambda_{\max}(\cdot)$
for the smallest and largest eigenvalue of the indicated matrix.

\section{Mixture autoregressive models}

\subsection{General formulation}

Let $y_{t}$ ($t=1,2,\ldots$) be a real-valued time series of interest,
and let $\mathcal{F}_{t-1}=\sigma\left(y_{s},\textrm{ }s<t\right)$
denote the $\sigma$\textendash algebra generated by past $y_{t}$'s.
We use $P_{t-1}\left(\cdot\right)$ to signify the conditional probability
of the indicated event given $\mathcal{F}_{t-1}$. In the general
two component mixture autoregressive model we consider the $y_{t}$'s
are generated by
\begin{equation}
y_{t}=s_{t}\biggl(\tilde{\phi}_{0}+\sum_{i=1}^{p}\tilde{\phi}_{i}y_{t-i}+\tilde{\sigma}_{1}\varepsilon_{t}\biggr)+(1-s_{t})\biggl(\tilde{\varphi}_{0}+\sum_{i=1}^{p}\tilde{\varphi}_{i}y_{t-i}+\tilde{\sigma}_{2}\varepsilon_{t}\biggr),\label{Model 1}
\end{equation}
where the parameters $\tilde{\sigma}_{1}$ and $\tilde{\sigma}_{2}$
are positive, and conditions required for the autoregressive parameters
$\tilde{\phi}_{i}$ and $\tilde{\varphi}_{i}$ ($i=1,\ldots,p$) will
be discussed later. Furthermore, $\varepsilon_{t}$ and $s_{t}$ are
(unobserved) stochastic processes which satisfy the following conditions:
(a) $\varepsilon_{t}$ is a sequence of independent standard normal
random variables such that $\varepsilon_{t}$ is independent of $\{y_{t-j},\ j>0\}$,
(b) $s_{t}$ is a sequence of Bernoulli random variables such that,
for each $t$, $P_{t-1}(s_{t}=1)=\alpha_{t}$ with $\alpha_{t}$ a
function of $\boldsymbol{y}_{t-1}=(y_{t-1},\ldots,y_{t-p})$, and
(c) conditional on $\mathcal{F}_{t-1}$, $s_{t}$ and $\varepsilon_{t}$
are independent.

The conditional probabilities $\alpha_{t}$ and $1-\alpha_{t}$ ($=P_{t-1}(s_{t}=0)$)
are referred to as mixing weights. They can be thought of as (conditional)
probabilities that determine which one of the two autoregressive components
of the mixture generates the next observation $y_{t}$. In condition
(b) it is assumed that the mixing weight $\alpha_{t}$ (and hence
also the conditional distribution of $y_{t}$ given its past) only
depends on $p$ lags of $y_{t}$; allowing for more than $p$ lags
in the mixing weight would be possible at the cost of more complicated
notation.

We assume that of the original parameters $\tilde{\phi}=(\tilde{\phi}_{0},\tilde{\phi}_{1},\ldots,\tilde{\phi}_{p},\tilde{\sigma}_{1}^{2})$
and $\tilde{\varphi}=(\tilde{\varphi}_{0},\tilde{\varphi}_{1},\ldots,\tilde{\varphi}_{p},\tilde{\sigma}_{2}^{2})$
in the two regimes, $q_{1}$ parameters are a priori assumed the same
in both regimes and the remaining $q_{2}$ parameters are potentially
different in the two regimes (with $q_{1}+q_{2}=p+2$). For instance,
one may assume that $\tilde{\phi}_{0}$ and $\tilde{\varphi}_{0}$
are equal, or alternatively that $\tilde{\sigma}_{1}^{2}$ and $\tilde{\sigma}_{2}^{2}$
are equal. If such an assumption is plausible, taking it into account
when devising a test for regime switching will be advantageous (it
will lead to a test with better power). To this end, let $\beta$
be a $q_{1}\times1$ vector of common parameters, and let $\phi$
and $\varphi$ be $q_{2}\times1$ vectors of (potentially) different
parameters. Then, for some known $(p+2)$\textendash dimensional permutation
matrix $P$, $(\beta,\phi)=P\tilde{\phi}$ and $(\beta,\varphi)=P\tilde{\varphi}$.
For simplicity, we assume that $\beta$ and $\phi$ are variation-free,
requiring the autoregressive coefficients $\tilde{\phi}_{1},\ldots,\tilde{\phi}_{p}$
to be contained in either $\beta$ or $\phi$ (the same variation-freeness
is assumed of $\beta$ and $\varphi$). If there are no common coefficients
in the two regimes, the parameter $\beta$ can be dropped and $\phi=\tilde{\phi}$
and $\varphi=\tilde{\varphi}$.

As for the mixing weight $\alpha_{t}$, in addition to past $y_{t}$'s
it depends on unknown parameters which may include components of the
parameter vector $(\beta,\phi,\varphi)$ and an additional parameter
$\alpha$ (scalar or vector). When this dependence needs to be emphasized
we use the notation $\alpha_{t}(\alpha,\beta,\phi,\varphi)$.

Using equation (\ref{Model 1}) and the conditions following it, the
conditional density function of $y_{t}$ given its past, $f(\cdot\mid\mathcal{F}_{t-1})$,
is obtained as
\begin{equation}
f(y_{t}\mid\mathcal{F}_{t-1})=\alpha_{t}f_{t}(\beta,\phi)+(1-\alpha_{t})f_{t}(\beta,\varphi),\label{Model 2}
\end{equation}
where the notation $f_{t}(\beta,\phi)$ signifies the density function
of a (univariate) normal distribution with mean $\tilde{\phi}_{0}+\tilde{\phi}_{1}y_{t-1}+\cdots+\tilde{\phi}_{p}y_{t-p}$
and variance $\tilde{\sigma}_{1}^{2}$ evaluated at $y_{t}$, that
is,

\begin{equation}
f_{t}(\beta,\phi)=\frac{1}{\tilde{\sigma}_{1}}{\scriptstyle \mathscr{\mathfrak{N}}}\biggl(\frac{y_{t}-(\tilde{\phi}_{0}+\tilde{\phi}_{1}y_{t-1}+\cdots+\tilde{\phi}_{p}y_{t-p})}{\tilde{\sigma}_{1}}\biggr),\label{eq:f_t}
\end{equation}
with $\mathscr{{\scriptstyle \mathfrak{N}}}(u)=(2\bar{\pi})^{-1/2}\exp(-u^{2}/2)$
the density function of a standard normal random variable and $\bar{\pi}=3.14\ldots$
the number pi. The notation $f_{t}(\beta,\varphi)$ is defined similarly
by using the parameters $\tilde{\varphi}_{i}$ $\left(i=0,\ldots,p\right)$
and $\tilde{\sigma}_{2}^{2}$ instead of $\tilde{\phi}_{i}$ $\left(i=0,\ldots,p\right)$
and $\tilde{\sigma}_{1}^{2}$. Thus, the distribution of $y_{t}$
given its past is specified as a mixture of two normal densities with
time varying mixing weights $\alpha_{t}$ and $1-\alpha_{t}$.

Different mixture autoregressive models are obtained by different
specifications of the mixing weights (or in our case the single mixing
weight $\alpha_{t}$). In some of the proposed models more than two
mixture components are allowed but for reasons to be discussed below
we will not consider these extensions. If the mixing weights are assumed
constant over time the general mixture autoregressive model introduced
above reduces to (a two component version) of the MAR\ model studied
by \citet{wong2000mixture}. In the LMAR\ model of \citet{wong2001logistic},
a logistic transformation of the two mixing weights is assumed to
be a linear function of past observed variables. Related two-regime
mixture models with time-varying mixing weights have also been considered
by \citet{gourieroux2006stochastic}, \citet{dueker2007contemporaneous}
and \citet*{bec2008acr} whereas \citet{lanne2003modeling} and \citet{kalliovirta2015gaussian}
have considered mixture autoregressive models in which multiple regimes
are allowed.

A common problem with the application of mixture autoregressive models
is determining the value of the (usually unknown) number of component
models or regimes. As discussed in the Introduction, several irregular
features make this problem difficult and these difficulties are encountered
even when the observations are a random sample from a mixture of (two)
normal distributions. To our knowledge the only solution presented
for mixture autoregressive models is provided for a simple first order
case with no intercept terms in the unpublished PhD thesis of \citet{jeffries1998logistic}.
As discussed in the recent papers by \citet{kasahara2012testing,kasahara2015testing}
and the references therein, some of the difficulties involved stem
from properties of the normal distribution.

The difficulties referred to above also partly explain the complexity
of our testing problem and why we only consider test procedures that
can be used to test the null hypothesis that a two component mixture
autoregressive model reduces to a conventional linear autoregressive
model. Following the ideas in \citet{zhu2006asymptotics} and \citet{kasahara2012testing,kasahara2015testing},
we derive a LR test in the general set-up described above and apply
it to two particular cases, the LMAR model of \citet{wong2001logistic}
and the GMAR model of \citet{kalliovirta2015gaussian}. Next, we shall
discuss these two models in more detail.

\subsection{Two particular examples }

\paragraph{LMAR Example.}

\noindent The LMAR model of \citet{wong2001logistic} is defined by
specifying the mixing weight $\alpha_{t}$ as
\[
\alpha_{t}^{L}=\alpha_{t}^{L}(\alpha)=\frac{\exp(\alpha_{0}+\alpha_{1}y_{t-1}+\cdots+\alpha_{r}y_{t-m})}{1+\exp(\alpha_{0}+\alpha_{1}y_{t-1}+\cdots+\alpha_{r}y_{t-m})},
\]
where the vector $\alpha=(\alpha_{0},\alpha_{1},\ldots,\alpha_{m})$
contains $m+1$ unknown parameters and the order $m$ ($1\leq m\leq p$)
is assumed known.

\paragraph{GMAR Example.}

\noindent In the GMAR model of \citet{kalliovirta2015gaussian} the
mixing weight is defined as
\begin{equation}
\alpha_{t}^{G}=\alpha_{t}^{G}(\alpha,\tilde{\phi},\tilde{\varphi})=\frac{\alpha\mathsf{n}_{p}(\boldsymbol{y}_{t-1};\tilde{\phi})}{\alpha\mathsf{n}_{p}(\boldsymbol{y}_{t-1};\tilde{\phi})+(1-\alpha)\mathsf{n}_{p}(\boldsymbol{y}_{t-1};\tilde{\varphi})},\label{Mixing Weights}
\end{equation}
where $\alpha\in(0,1)$ is an unknown parameter and $\mathsf{n}_{p}(\boldsymbol{y}_{t-1};\cdot)$
denotes the density function of a particular $p$\textendash dimensional
($p\geq1$) normal distribution defined as follows.

First, define the auxiliary Gaussian AR($p$) processes (cf. equation
(\ref{Model 1}))
\[
\nu{}_{1,t}=\tilde{\phi}_{0}+\sum_{i=1}^{p}\tilde{\phi}_{i}\nu_{1,t-i}+\tilde{\sigma}_{1}\varepsilon_{t}\ \ \ \textrm{and}\ \ \ \nu_{2,t}=\tilde{\varphi}_{0}+\sum_{i=1}^{p}\tilde{\varphi}_{i}\nu_{2,t-i}+\tilde{\sigma}_{2}\varepsilon_{t},
\]
where the autoregressive coefficients are assumed to satisfy
\begin{equation}
\tilde{\phi}(z):=1-\sum_{i=1}^{p}\tilde{\phi}_{i}z^{i}\neq0\,\,\text{for}\,\,\left\vert z\right\vert \leq1\,\,\,\,\textrm{and}\,\,\,\,\tilde{\varphi}(z):=1-\sum_{i=1}^{p}\tilde{\varphi}_{i}z^{i}\neq0\,\,\text{for}\,\,\left\vert z\right\vert \leq1.\label{Stat. condition}
\end{equation}
This condition implies that the processes $\nu{}_{1,t}$ and $\nu{}_{2,t}$
are stationary and that each of the two component models in (\ref{Model 1})
satisfies the usual stationarity condition of the conventional linear
AR($p$) model. Now set $\mathbf{\boldsymbol{\nu}}_{m,t}=\left(\nu_{m,t},\ldots,\nu_{m,t-p+1}\right)$
and $\mathbf{1}_{p}=\left(1,\ldots,1\right)$ ($p\times1$), and let
$\mu_{m}\mathbf{1}_{p}$ and $\mathbf{\Gamma}_{m,p}$ signify the
mean vector and covariance matrix of $\mathbf{\boldsymbol{\nu}}_{m,t}$
($m=1,2$).\footnote{We have $\mu_{1}=\tilde{\phi}_{0}/\tilde{\phi}\left(1\right)$ and
$\mu_{2}=\tilde{\varphi}_{0}/\tilde{\varphi}\left(1\right)$, whereas
each of $\mathbf{\Gamma}_{m,p}$, $(m=1,2)$, has the familiar form
of being a $p\times p$ symmetric Toeplitz matrix with $\gamma_{m,0}=Cov[\nu_{m,t},\nu_{m,t}]$
along the main diagonal, and $\gamma_{m,i}=Cov[\nu_{m,t},\nu{}_{m,t-i}]$,
$i=1,\ldots,p-1$, on the diagonals above and below the main diagonal.
Similarly to $\mu_{1}$ and $\mu_{2}$ the elements of the covariance
matrices $\mathbf{\Gamma}_{1,p}$ and $\mathbf{\Gamma}_{2,p}$ are
treated as functions of the parameters $\tilde{\phi}$ and $\tilde{\varphi}$,
respectively (for details of this dependence, see \citet[eqn. (2.1.39)]{lutkepohl2005new}).} The random vector $\mathbf{\boldsymbol{\nu}}_{1,t}$ follows the
$p$\textendash dimensional multivariate normal distribution with
density
\begin{equation}
\mathsf{n}_{p}(\mathbf{\boldsymbol{\nu}}_{1,t};\tilde{\phi})=\left(2\bar{\pi}\right)^{-p/2}\det(\mathbf{\Gamma}_{1,p}\mathbf{)}^{-1/2}\exp\left\{ -\tfrac{1}{2}\left(\mathbf{\boldsymbol{\nu}}_{1,t}-\mu_{1}\mathbf{1}_{p}\right)^{\prime}\mathbf{\Gamma}_{1,p}^{-1}\left(\mathbf{\boldsymbol{\nu}}_{1,t}-\mu_{1}\mathbf{1}_{p}\right)\right\} ,\label{Normal p}
\end{equation}
and the density of $\mathbf{\boldsymbol{\nu}}_{2,t}$, denoted by
$\mathsf{n}_{p}(\mathbf{\boldsymbol{\nu}}_{2,t};\tilde{\varphi})$,
is defined similarly. Equation (\ref{Model 1}) and conditions (\ref{Mixing Weights})\textendash (\ref{Normal p})
define the (two component) GMAR model (condition (\ref{Stat. condition})
is part of the definition of the model because it is used to define
the mixing weights).

\section{Test procedure}

We now consider a test procedure of the null hypothesis that a two
component mixture autoregressive model reduces to a conventional linear
autoregressive model.

\subsection{The null hypothesis and the LR test statistic}

We denote the conditional density function corresponding to the unrestricted
model as (see (\ref{Model 2}))
\[
f_{2,t}(\alpha,\beta,\phi,\varphi):=f_{2}(y_{t}\mid\boldsymbol{y}_{t-1};\alpha,\beta,\phi,\varphi):=\alpha_{t}(\alpha,\beta,\phi,\varphi)f_{t}(\beta,\phi)+\left(1-\alpha_{t}(\alpha,\beta,\phi,\varphi)\right)f_{t}(\beta,\varphi),
\]
where we now make the dependence of the mixing weight on the parameters
explicit. With this notation the log-likelihood function of the model
based on a sample $(y_{-p+1},\ldots,y_{T})$ (and conditional on the
initial values $(y_{-p+1},\ldots,y_{0})$) is $L_{T}(\alpha,\beta,\phi,\varphi)=\sum_{t=1}^{T}l_{t}(\alpha,\beta,\phi,\varphi)$
where
\[
l_{t}(\alpha,\beta,\phi,\varphi)=\log[f_{2,t}(\alpha,\beta,\phi,\varphi)]=\log\left[\alpha_{t}(\alpha,\beta,\phi,\varphi)f_{t}(\beta,\phi)+\left(1-\alpha_{t}(\alpha,\beta,\phi,\varphi)\right)f_{t}(\beta,\varphi)\right].
\]
The following assumption provides conditions on the data generation
process, the parameter space of $\left(\alpha,\beta,\phi,\varphi\right)$,
and the mixing weight $\alpha_{t}\left(\alpha,\beta,\phi,\varphi\right)$.
\begin{assumption}
\label{assu:DGPParSpaceMixWeight}~\vspace{-5pt}

\begin{description}
\item [{(i)}] The $y_{t}$'s are generated by a stationary linear Gaussian
AR($p$) model with (the true but unknown) parameter value $\tilde{\phi}^{\ast}$
an interior point of $\tilde{\Phi}$, a compact subset of $\{\tilde{\phi}=(\tilde{\phi}_{0},\tilde{\phi}_{1},\ldots,\tilde{\phi}_{p},\tilde{\sigma}^{2})\in\mathbb{R}^{p+2}:\tilde{\phi}_{0}\in\mathbb{R};\,1-\sum_{i=1}^{p}\tilde{\phi}_{i}z^{i}\neq0\text{ \ for }\left\vert z\right\vert \leq1;\,\tilde{\sigma}^{2}\in(0,\infty)\}$.
\item [{(ii)}] The parameter space of $\left(\alpha,\beta,\phi,\varphi\right)$
is $A\times B\times\Phi\times\Phi$, where $A$ is a compact subset
of $\mathbb{R}^{a}$ and $B$ and $\Phi$ are those compact subsets
of $\mathbb{R}^{q_{1}}$ and $\mathbb{R}^{q_{2}}$, respectively,
that satisfy $(\beta,\phi)\in B\times\Phi$ if and only if $P^{-1}(\beta,\phi)\in\tilde{\Phi}$
(here $P$ is as in the third paragraph of Section 2.1).
\item [{(iii)}] For all $t$ and all $\left(\alpha,\beta,\phi,\varphi\right)\in A\times B\times\Phi\times\Phi$,
the mixing weight $\alpha_{t}\left(\alpha,\beta,\phi,\varphi\right)$,
is $\sigma(\boldsymbol{y}{}_{t-1})$\textendash measurable (with $\sigma(\boldsymbol{y}{}_{t-1})$
denoting the $\sigma$\textendash algebra generated by $\boldsymbol{y}{}_{t-1}$)
and satisfies $\alpha_{t}\left(\alpha,\beta,\phi,\varphi\right)\in(0,1)$.
\end{description}
\end{assumption}
As our interest is to study the asymptotic null distribution of the
LR test statistic, Assumption 1(i) requires the data to be generated
by a stationary linear Gaussian AR($p$) model. Assuming a compact
parameter space in Assumptions 1(i) and (ii) is a standard requirement
which facilitates proofs. Assumption 1(ii) accommodates to the main
cases of interest, namely $\beta=\tilde{\phi}_{0}$, $\beta=(\tilde{\phi}_{1},\ldots,\tilde{\phi}_{p})$,
$\beta=\tilde{\sigma}^{2}$, or any combination of these.\footnote{Note that assuming the autoregressive parameters $\phi$ and $\varphi$
to have a common parameter space is made for simplicity and could
be relaxed; for an example where such a relaxation would be needed,
see the ACR model of \citet{bec2008acr}.}

Assumption 1(iii) implies that our two-component mixture autoregressive
model reduces to a linear autoregression only when $\phi=\varphi$,
regardless of the values of $\alpha\in A$ and $\beta\in B$. The
null hypothesis to be tested is therefore $\phi=\varphi$ and the
alternative is $\phi\neq\varphi$ or, more precisely,
\[
H_{0}:(\phi,\varphi)\in(\Phi\times\Phi)^{0},\;\alpha\in A,\,\beta\in B\qquad\textrm{vs.}\qquad H_{1}:(\phi,\varphi)\in(\Phi\times\Phi)\setminus(\Phi\times\Phi)^{0},\;\alpha\in A,\,\beta\in B,
\]
where
\[
(\Phi\times\Phi)^{0}=\{(\phi,\varphi)\in\Phi\times\Phi:\phi=\varphi\}.
\]
Note that under the null hypothesis the parameter $\alpha$ vanishes
from the likelihood function and is therefore unidentified.

Let $f_{t}^{0}(\tilde{\phi}):=f^{0}(y_{t}\mid\boldsymbol{y}_{t-1};\tilde{\phi})$
and $L_{T}^{0}(\tilde{\phi})$ denote the conditional density and
log-likelihood corresponding to the restricted model, that is,
\[
f_{t}^{0}(\tilde{\phi})=f_{2,t}(\alpha,\beta,\phi,\phi)=f_{t}(\tilde{\phi})\,\,\,\,\,\text{and}\,\,\,\,\,L_{T}^{0}(\tilde{\phi})=\sum_{t=1}^{T}l_{t}^{0}(\tilde{\phi})\,\,\,\,\,\text{with}\,\,\,\,\,l_{t}^{0}(\tilde{\phi})=\log[f_{t}(\tilde{\phi})]
\]
(here the superscript $0$ refers to the model restricted by the null
hypothesis). Note that these quantities are obtained from a linear
Gaussian AR($p$) model. As $f_{2,t}(\alpha,\beta^{*},\phi^{\ast},\phi^{\ast})=f_{t}(\tilde{\phi}^{\ast})$
for any $\alpha\in A$, in the unrestricted model the parameter vector
$(\alpha,\beta^{*},\phi^{\ast},\phi^{\ast})$ corresponds to the true
model for any $\alpha\in A$.

As already indicated, Assumption 1(iii) implies that the restriction
$\phi=\varphi$ is the only possibility to formulate the null hypothesis.
However, this is not necessarily the case if (against Assumption 1(iii))
the mixing weight $\alpha_{t}\left(\alpha,\beta,\phi,\varphi\right)$
were allowed to take the boundary values zero and one. Of our two
examples this would be possible for the GMAR model of \citet{kalliovirta2015gaussian}
but not for the LMAR model of \citet{wong2001logistic}. In the GMAR
model $\alpha_{t}\left(\alpha,\beta,\phi,\varphi\right)$ takes the
boundary values zero and one when the parameter $\alpha$ takes these
values (see Section 2.2). In both cases a linear autoregression results
and either the parameter $\phi$ or $\varphi$ is unidentified (see
(\ref{Model 2})) (the MAR model of \citet{wong2000mixture} provides
a similar example). It would be possible to obtain tests for the GMAR
model by using the null hypotheses which specifies $\alpha=0$ or
$\alpha=1$. However, as in \citet{kasahara2012testing} (see also
\citet{kasahara2015testing}), this approach would require rather
restrictive assumptions and would also lead to very complicated derivations.\footnote{See, for instance, the remarks following Proposition 5 in \citet{kasahara2012testing}
or property (a) on p. 1633 of \citet{kasahara2015testing}.} Therefore, we will not consider this option.

As the parameter $\alpha$ is unidentified under the null hypothesis,
the appropriate likelihood ratio type test statistic is
\[
LR_{T}=\sup_{\alpha\in A}LR_{T}\left(\alpha\right),
\]
where, for each fixed $\alpha\in A$,
\[
LR_{T}\left(\alpha\right)=2[\sup_{(\beta,\phi,\varphi)\in B\times\Phi\times\Phi}L_{T}(\alpha,\beta,\phi,\varphi)-\sup_{\tilde{\phi}\in\tilde{\Phi}}L_{T}^{0}(\tilde{\phi})].
\]
To obtain an operational test statistic let, for each fixed $\alpha\in A$,
$(\hat{\beta}_{T\alpha},\hat{\phi}_{T\alpha},\hat{\varphi}_{T\alpha})$
denote an (approximate) unrestricted maximum likelihood (ML) estimator
of the parameter vector $\left(\beta,\phi,\varphi\right)$. We make
the following assumption.
\begin{assumption}
\label{assu:unrestricted_est}The unrestricted ML estimator satisfies
the following conditions:~

\begin{description}
\item [{(i)}] $L_{T}(\alpha,\hat{\beta}_{T\alpha},\hat{\phi}_{T\alpha},\hat{\varphi}_{T\alpha})=\sup\nolimits _{(\beta,\phi,\varphi)\in B\times\Phi\times\Phi}L_{T}(\alpha,\beta,\phi,\varphi)+o_{p\alpha}(1)$,
\item [{(ii)}] $(\hat{\beta}_{T\alpha},\hat{\phi}_{T\alpha},\hat{\varphi}_{T\alpha})=(\beta^{*},\phi^{\ast},\phi^{\ast})+o_{p\alpha}(1)$.
\end{description}
\end{assumption}
Assumption 2(i) means that $(\hat{\beta}_{T\alpha},\hat{\phi}_{T\alpha},\hat{\varphi}_{T\alpha})$
is assumed to maximize the likelihood function only asymptotically.
This assumption is technical and made for ease of exposition (see
\citet{andrews1999estimation} and \citet{zhu2006asymptotics} for
similar assumptions in related problems). Assumption 2(ii) is a high
level condition on (uniform) consistency of the ML estimator and is
analogous to Assumption 1 of \citet{andrews2001testing}. It has to
be verified on a case by case basis (this is exemplified below for
the LMAR model and GMAR model).

As for the term $\sup_{\tilde{\phi}\in\tilde{\Phi}}L_{T}^{0}(\tilde{\phi})$
in the LR test statistic, note that $L_{T}^{0}(\tilde{\phi})$ is
the (conditional) log-likelihood function of a linear Gaussian AR($p$)
model. Let $\hat{\tilde{\phi}}_{T}$ denote an (approximate) maximum
likelihood estimator of the parameters of a linear Gaussian AR($p$)
model, that is, $\hat{\tilde{\phi}}_{T}$ satisfies\footnote{Note that the parameter space for $\tilde{\phi}$ is the compact set
$\tilde{\Phi}$ and not the entire stationarity region of a (causal)
linear AR($p$) model. Asymptotically, also the OLS estimator can
be used. }
\[
L_{T}^{0}(\hat{\tilde{\phi}}_{T})=\sup_{\tilde{\phi}\in\tilde{\Phi}}L_{T}^{0}(\tilde{\phi})+o_{p}(1)\,\,\,\textrm{and}\,\,\,\hat{\tilde{\phi}}_{T}=\tilde{\phi}^{\ast}+o_{p}(1).
\]
Noting that $L_{T}(\alpha,\beta^{*},\phi^{\ast},\phi^{\ast})=L_{T}^{0}(\tilde{\phi}^{\ast})$
for any $\alpha$ now allows us to write $LR_{T}\left(\alpha\right)$
as
\begin{equation}
LR_{T}\left(\alpha\right)=2[L_{T}(\alpha,\hat{\beta}_{T\alpha},\hat{\phi}_{T\alpha},\hat{\varphi}_{T\alpha})-L_{T}(\alpha,\beta^{*},\phi^{\ast},\phi^{\ast})]-2[L_{T}^{0}(\hat{\tilde{\phi}}_{T})-L_{T}^{0}(\tilde{\phi}^{\ast})]+o_{p\alpha}(1).\label{eq:LR_1}
\end{equation}
The analysis of the second term on the right hand side is standard
while dealing with the first term is more demanding requiring a substantial
amount of preparation.

\subsubsection{Examples (continued)}

\paragraph{LMAR Example. }

In the LMAR example, we assume there are no common parameters in the
two regimes so that the parameter $\beta$ is omitted, $\phi=\tilde{\phi}$,
$\varphi=\tilde{\varphi}$, $q_{1}=0$, and $q_{2}=p+2$. To satisfy
conditions (ii) and (iii) in Assumption 1, $A$ can be any compact
subset of $\{(\alpha_{0},\alpha_{1},\ldots,\alpha_{m})\in\mathbb{R}^{m+1}:(\alpha_{1},\ldots,\alpha_{m})\neq(0,\ldots,0)\}$
where $1\leq m\leq p$. This ensures that the mixing weight $\alpha_{t}^{L}$
is not equal to a constant. For the verification of Assumption 2(ii),
see Appendix B.

When $m=0$, the mixing weight $\alpha_{t}^{L}$ (and hence $1-\alpha_{t}^{L}$)
is constant and the LMAR model reduces to the MAR model of \citet{wong2000mixture}.
In this special case our testing problem requires different and more
complicated analyses than in the `real' LMAR case where $m\geq1$
and $(\alpha_{1},\ldots,\alpha_{r})\neq(0,\ldots,0)$ (we shall discuss
this point more later). Therefore, the conditions $m\geq1$ and $(\alpha_{1},\ldots,\alpha_{m})\neq(0,\ldots,0)$
will be assumed in the sequel. A similar restriction is made by \citet{jeffries1998logistic}
in his (first-order) logistic mixture autoregressive model to facilitate
the derivation of the LR test (see the hypotheses at the end of p.
95 and the following discussion, as well as the end of p. 110).

\paragraph{GMAR Example. }

\noindent The GMAR model exemplifies the setting with common coefficients
by assuming that the intercept terms in the two regimes are the same
(note that this still allows for different means in the two regimes).
As will be discussed in more detail in Section 3.3.1, this assumption
is partly due to the fact that otherwise the derivation of the LR
test would become extremely complicated. Hence, in this example $\beta=\tilde{\phi}_{0}\,(=\tilde{\varphi}_{0})$,
$\phi=(\tilde{\phi}_{1},\ldots,\tilde{\phi}_{p},\tilde{\sigma}_{1}^{2})$,
$\varphi=(\tilde{\varphi}_{1},\ldots,\tilde{\varphi}_{p},\tilde{\sigma}_{2}^{2})$,
$q_{1}=1$, and $q_{2}=p+1$. To satisfy Assumptions 1(ii) and (iii)
the parameter space $A$ of $\alpha$ can be any compact and convex
subset of $(0,1)$ (this also rules out the possibility that $\alpha=0$
or $\alpha=1$ discussed after Assumption 1). For the verification
of Assumption 2(ii), see Appendix C.

It may be worth noting that there are cases where the mixing weight
$\alpha_{t}^{G}$ is time invariant and equals $\alpha$. If this
happens the GMAR model reduces to the MAR model of \citet{wong2000mixture}.\footnote{An example is when $p=1$, $\tilde{\phi}_{0}=\tilde{\varphi}_{0}=0$,
and $\tilde{\sigma}_{1}^{2}/(1-\tilde{\phi}_{1}^{2})=\tilde{\sigma}_{2}^{2}/(1-\tilde{\varphi}_{1}^{2})$
where the last equality can hold even if $(\tilde{\phi},\tilde{\sigma}_{1}^{2})$
is different from $(\tilde{\varphi},\tilde{\sigma}_{2}^{2})$.} However, unlike in the case of the LMAR model this fact does not
complicate the derivation of our test. The reason seems to be that
in the GMAR model the reduction occurs only for certain values of
the parameters $\tilde{\phi}$ and $\tilde{\varphi}$ whereas in the
LMAR model it occurs for all values of $\tilde{\phi}$ and $\tilde{\varphi}$.

\subsection{Reparameterized model}

In standard testing problems the derivation of the asymptotic distribution
of the LR test would rely on a quadratic expansion of the log-likelihood
function $L_{T}(\alpha,\beta,\phi,\varphi)=\sum_{t=1}^{T}l_{t}(\alpha,\beta,\phi,\varphi)$;
when the parameter $\alpha$ is not identified under the null hypothesis,
the relevant derivatives in this expansion would be with respect to
$(\beta,\phi,\varphi)$ for fixed values of $\alpha\in A$. In problems
with a singular information matrix it turns out to be convenient to
follow \citet{rotnitzky2000likelihood} and \citet{kasahara2012testing,kasahara2015testing}
and employ an appropriately reparameterized model.

The employed reparameterization is model specific and aims to have
two conveniences. First, it transforms the null hypothesis $\phi=\varphi$
into a point null hypothesis where some components of the parameter
vector are restricted to zero and the rest are left unrestricted.
Second, and more importantly, it simplifies derivations in cases where
the conventional quadratic expansion of the log-likelihood function
breaks down because, under the null hypothesis, the scores of the
parameters $(\beta,\phi,\varphi)$ are linearly dependent and, consequently,
the (Fisher) information matrix is singular. As will be seen later,
this is the case for the GMAR model of \citet{kalliovirta2015gaussian}
but not for the LMAR model of \citet{wong2001logistic}.

General requirements for the reparameterization are described in the
following assumption. Only the parameters restricted by the null hypothesis,
$\phi$ and $\varphi$, are reparameterized. The examples in this
and the following subsection illustrate how the reparameterization
could be chosen.

\pagebreak{}
\begin{assumption}
\label{assu:pi-repam}~\vspace{-5pt}

\begin{description}
\item [{(i)}] For every $\alpha\in A$, the mapping $(\pi,\varpi)=\boldsymbol{\pi}_{\alpha}(\phi,\varphi)$
from $\Phi\times\Phi$ to $\Pi_{\alpha}$ is one-to-one with $\boldsymbol{\pi}_{\alpha}$
and $\boldsymbol{\pi}_{\alpha}^{-1}$ continuous.
\item [{(ii)}] For every $\alpha\in A$, $\boldsymbol{\pi}_{\alpha}((\Phi\times\Phi)^{0})=\Phi\times\{0\}$
and $\boldsymbol{\pi}_{\alpha}(\phi^{\ast},\phi^{\ast})=(\pi^{\ast},0):=(\phi^{\ast},0)$.
\item [{(iii)}] $(\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})=(\beta^{*},\pi^{\ast},0)+o_{p\alpha}(1)$,
where $(\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})\coloneqq\boldsymbol{\pi}_{\alpha}(\hat{\phi}_{T\alpha},\hat{\varphi}_{T\alpha})$.
\end{description}
\end{assumption}
We sometimes refer to the reparameterization described in Assumption
\ref{assu:pi-repam} as the `$\pi$-parameterization' and the original
reparameterization as the `$\phi$-parameterization'. Note that the
transformed parameters $\pi$ and $\varpi$ generally depend on $\alpha$
but, for brevity, we suppress this dependence from the notation. The
parameter space of $(\pi,\varpi)$ also depends on $\alpha$ and is
given, for any $\alpha\in A$, by
\[
\Pi_{\alpha}=\{(\pi,\varpi)\in\mathbf{\mathbb{R}}^{2q_{2}}:(\pi,\varpi)=\boldsymbol{\pi}_{\alpha}(\phi,\varphi)\;\;\textrm{for some}\;\;(\phi,\varphi)\in\Phi\times\Phi\}.
\]
By Assumption \ref{assu:pi-repam}(ii), the null hypothesis $\phi=\varphi$
can be equivalently written as $\varpi=0$ or, more precisely, as
\[
H_{0}:\pi\in\Phi,\;\varpi=0,\;\alpha\in A,\,\beta\in B\qquad\textrm{vs.}\qquad H_{1}:(\pi,\varpi)\in\Pi_{\alpha}\setminus(\Phi\times\{0\}),\;\alpha\in A,\,\beta\in B.
\]
Note that under $H_{0}$, the parameters $\beta$ and $\pi$ are identified,
but $\alpha$ is not. As for Assumption \ref{assu:pi-repam}(iii),
it is a high level condition similar to Assumption \ref{assu:unrestricted_est}(ii)
from which it can be derived with appropriate additional assumptions.
A simple Lipschitz condition similar to \citet[Assumption SE-1(b)]{andrews1992generic},
given in Lemma \ref{lem:LipCondForUnifCons} in Appendix A, is one
possibility.

To develop further notation, partition $\boldsymbol{\pi}_{\alpha}^{-1}(\pi,\varpi)$
into two $q_{2}$\textendash dimensional components as $\boldsymbol{\pi}_{\alpha}^{-1}(\pi,\varpi)=(\boldsymbol{\pi}_{\alpha,1}^{-1}(\pi,\varpi),\boldsymbol{\pi}_{\alpha,2}^{-1}(\pi,\varpi))$,
and define
\[
\hspace{-10pt}f_{2,t}^{\pi}(\alpha,\beta,\pi,\varpi):=f_{2}(y_{t}\mid\boldsymbol{y}_{t-1};\alpha,\beta,\phi,\varphi)=\alpha_{t}^{\pi}(\alpha,\beta,\pi,\varpi)f_{t}(\beta,\boldsymbol{\pi}_{\alpha,1}^{-1}(\pi,\varpi))+(1-\alpha_{t}^{\pi}(\alpha,\beta,\pi,\varpi))f_{t}(\beta,\boldsymbol{\pi}_{\alpha,2}^{-1}(\pi,\varpi)),
\]
where $\alpha_{t}^{\pi}(\alpha,\beta,\pi,\varpi)=\alpha_{t}(\alpha,\beta,\boldsymbol{\pi}_{\alpha,1}^{-1}(\pi,\varpi),\boldsymbol{\pi}_{\alpha,2}^{-1}(\pi,\varpi))$
and the function $f_{t}(\cdot)$ is as in (\ref{Model 2}). The log-likelihood
function of the reparameterized model can now be expressed as
\begin{equation}
L_{T}^{\pi}(\alpha,\beta,\pi,\varpi)=\sum_{t=1}^{T}l_{t}^{\pi}(\alpha,\beta,\pi,\varpi),\label{eq:LogLikPi}
\end{equation}
where $l_{t}^{\pi}(\alpha,\beta,\pi,\varpi)=\log[f_{2,t}^{\pi}(\alpha,\beta,\pi,\varpi)]$,
and in the $\pi$\textendash parameterization equation (\ref{eq:LR_1})
reads as
\begin{equation}
LR_{T}\left(\alpha\right)=2[L_{T}^{\pi}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})-L_{T}^{\pi}(\alpha,\beta^{*},\phi^{\ast},0)]-2[L_{T}^{0}(\hat{\tilde{\phi}}_{T})-L_{T}^{0}(\tilde{\phi}^{\ast})]+o_{p\alpha}(1).\label{eq:LR_pi-repam}
\end{equation}

\subsubsection{Examples (continued)}

\paragraph{LMAR Example. }

The reparameterization we employ in the LMAR model is
\[
(\pi,\varpi)=\boldsymbol{\pi}_{\alpha}(\phi,\varphi)=(\phi,\phi-\varphi)\;\;\textrm{so that}\;\;(\phi,\varphi)=\boldsymbol{\pi}_{\alpha}^{-1}(\pi,\varpi)=(\pi,\pi-\varpi).
\]
Note that in this case the reparameterization (via $\boldsymbol{\pi}_{\alpha}(\cdot,\cdot)$)
does not depend on $\alpha$, and the same is true for the parameter
space of $(\pi,\varpi)$. Verification of Assumption \ref{assu:pi-repam}
is straightforward using Lemma \ref{lem:LipCondForUnifCons} (for
details, see Appendix B). In the LMAR case, the only benefit of the
reparameterization is to transform the null hypothesis into a point
null hypothesis.

\paragraph{GMAR Example. }

\noindent In the GMAR model our reparameterization is obtained by
setting, for any fixed $\alpha\in A$,
\[
(\pi,\varpi)=\boldsymbol{\pi}_{\alpha}(\phi,\varphi)=(\alpha\phi+(1-\alpha)\varphi,\phi-\varphi)\;\;\textrm{so that}\;\;(\phi,\varphi)=\boldsymbol{\pi}_{\alpha}^{-1}(\pi,\varpi)=(\pi+(1-\alpha)\varpi,\pi-\alpha\varpi).
\]
Verification of Assumption \ref{assu:pi-repam} is again straightforward
using Lemma \ref{lem:LipCondForUnifCons} (for details, see Appendix
C). In the GMAR model, simplifying the null hypothesis is not the
only benefit of the reparameterization, as will be discussed next.

As discussed before Assumption \ref{assu:pi-repam}, the relevant
derivatives when expanding $L_{T}(\alpha,\beta,\phi,\varphi)$ are
with respect to $(\beta,\phi,\varphi)$ and, in the GMAR case, these
derivatives are linearly dependent under the null hypothesis. To see
this and how the reparameterization affects this feature, note first
that straightforward differentiation yields
\[
\nabla_{(\beta,\phi,\varphi)}l_{t}(\alpha,\beta,\phi,\phi)=\left(\alpha\frac{\nabla_{\beta}f_{t}(\beta,\phi)}{f_{t}(\beta,\phi)},\alpha\frac{\nabla_{\phi}f_{t}(\beta,\phi)}{f_{t}(\beta,\phi)},(1-\alpha)\frac{\nabla_{\phi}f_{t}(\beta,\phi)}{f_{t}(\beta,\phi)}\right),
\]
where the null hypothesis $\phi=\varphi$ is imposed and $\nabla$
denotes differentiation with respect to the indicated parameters.
As $(f_{t}(\beta,\phi))^{-1}\nabla_{(\beta,\phi)}f_{t}(\beta,\phi)$
is the score vector obtained from a linear Gaussian AR($p$) model,
it is clear that the covariance matrix of the $(2p+3)$\textendash dimensional
vector $\nabla_{(\beta,\phi,\varphi)}l_{t}(\alpha,\beta,\phi,\phi)$,
and hence the (Fisher) information matrix, is singular with rank $p+2$.
In contrast to the above, in the $\pi$\textendash parameterization
the score vector is given by (see Supplementary Appendix C)
\[
\nabla_{(\beta,\pi,\varpi)}l_{t}^{\pi}(\alpha,\beta,\pi,0)=\left(\frac{\nabla_{\beta}f_{t}(\beta,\pi)}{f_{t}(\beta,\pi)},\frac{\nabla_{\pi}f_{t}(\beta,\pi)}{f_{t}(\beta,\pi)},\mathbf{0}_{p+1}\right)
\]
when the null hypothesis $\varpi=0$ is imposed. Now the score of
$\varpi$ is identically zero so that the reparameterization simplifies
linear dependencies of the scores which turns out to be very useful
in subsequent asymptotic analyses.

\subsection{Quadratic expansion of the (reparameterized) log-likelihood function}

As alluded to above, in standard testing problems the asymptotic analysis
of a LR test statistic is based on a second order Taylor expansion
of the (average) log-likelihood function around the true parameter
value. An essential assumption here is positive definiteness of the
(limiting) information matrix but, as illustrated in the previous
section, this assumption does not necessarily hold in our testing
problem due to linear dependencies among the derivatives of the log-likelihood
function. As in \citet{rotnitzky2000likelihood}, \citet{zhu2006asymptotics},
and \citet{kasahara2012testing,kasahara2015testing}, we therefore
consider a quadratic expansion of the log-likelihood function that
is not based on a second order Taylor expansion but (possibly) on
a higher order Taylor expansion. The need for higher-order derivatives
is illustrated by the GMAR example: as the score of $\varpi$ is identically
zero, the second derivative (which turns out to be linearly independent
of the score of $(\beta,\pi)$) now provides the first (nontrivial)
local approximation for $\varpi$.

The following assumption ensures that the (reparameterized) log-likelihood
function (\ref{eq:LogLikPi}) is (at least) twice continuously differentiable.
\begin{assumption}
\label{assu:cont_differentiability}For some integer $k\geq2$, and
for every fixed $\alpha\in A$, the functions $\alpha_{t}(\alpha,\beta,\phi,\varphi)$
and $\boldsymbol{\pi}_{\alpha}^{-1}(\pi,\varpi)$ are $k$ times continuously
differentiable (with respect to $(\beta,\phi,\varphi)$ and $(\pi,\varpi)$
in the interior of $B\times\Phi\times\Phi$ and $\Pi_{\alpha}$, respectively).
\end{assumption}
In our general framework the reparameterized log-likelihood function
is assumed to have, for each $\alpha\in A$, a quadratic expansion
in a transformed parameter vector $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)$
around $(\beta^{*},\pi^{\ast},0)$ given by
\begin{multline}
L_{T}^{\pi}(\alpha,\beta,\pi,\varpi)-L_{T}^{\pi}(\alpha,\beta^{*},\pi^{*},0)\\
=(T^{-1/2}S_{T\alpha})'[T^{1/2}\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)]-\frac{1}{2}[T^{1/2}\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)]^{\prime}\mathcal{I}_{\alpha}[T^{1/2}\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)]+R_{T}(\alpha,\beta,\pi,\varpi).\label{eq:Quadr-Exp_1}
\end{multline}
To illustrate this expansion, suppose the information matrix is positive
definite so that the quantities on the right hand side are (typically)
based on a second order Taylor expansion with $S_{T\alpha}$ and $\mathcal{I}_{\alpha}$
functions of $(\alpha,\beta^{*},\pi^{*},0)$. As already mentioned,
this is the case for the LMAR model where (the following will be justified
shortly) the parameter $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)$
is independent of $\alpha$ and given by $(\pi-\pi^{\ast},\varpi)$
and, for each $\alpha\in A$, $S_{T\alpha}$ is the score vector,
$\mathcal{I}_{\alpha}$ is the (positive definite) Fisher information
matrix, and $R_{T}(\alpha,\beta,\pi,\varpi)$ is a remainder term.
As the notation indicates, these three terms depend on $\alpha$,
and in general they may involve partial derivatives of the log-likelihood
function of order higher than two (this is the case for the GMAR model,
as will be demonstrated shortly). Then it may also get more complicated
to find the reparameterization of the previous subsection and the
transformed parameter vector $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)$,
as the examples of Kasahara and Shimotsu (2015, 2017) and the discussion
on the GMAR model below show; one possibility is to consider the iterative
procedure discussed by \citet[Sections 4.4 and 4.5]{rotnitzky2000likelihood}
(for a recent illuminating illustration of this approach, see \citet{hallin2014skew}).

Our next assumption provides further details on expansion (\ref{eq:Quadr-Exp_1}).
We use $\Rightarrow$ to signify weak convergence of a sequence of
stochastic processes on a function space. In the assumption below,
the weak convergence of interest is that of the process $S_{T\alpha}$
(indexed by $\alpha\in A$) to a process $S_{\alpha}$. The two function
spaces relevant in this paper are $\mathcal{B}(A,\mathbb{R}^{k})$
and $\mathcal{C}(A,\mathbb{R}^{k})$, the former is the space of all
$\mathbb{R}^{k}$\textendash valued bounded functions defined on (the
compact set) $A$ equipped with the uniform metric ($d(x,y)=\sup_{a\in A}\left\Vert x(a)-y(a)\right\Vert $),
and the latter is the same but with the continuity of the functions
(with respect to $\alpha\in A$) also assumed.
\begin{assumption}
\label{assu:quadexp}For each $\alpha\in A$, the log-likelihood function
$L_{T}^{\pi}(\alpha,\beta,\pi,\varpi)$ has a quadratic expansion
given in (\ref{eq:Quadr-Exp_1}), where

\begin{description}
\item [{(i)}] for each $\alpha\in A$, $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)$
is a mapping from $B\times\Pi_{\alpha}$ to\\
$\Theta_{\alpha}=\{\boldsymbol{\theta}\in\mathbb{R}^{r}:\boldsymbol{\theta}=\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)\text{ for some }(\beta,\pi,\varpi)\in B\times\Pi_{\alpha}\}$
such that (a) $\boldsymbol{\theta}(\alpha,\beta^{*},\pi^{*},0)=0$
and (b) for all $\epsilon>0$, $\inf_{\alpha\in A}\inf$$_{(\beta,\pi,\varpi)\in B\times\Pi_{\alpha}:\left\Vert (\beta,\pi,\varpi)-(\beta^{*},\pi^{\ast},0)\right\Vert \geq\epsilon}\left\Vert \boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)\right\Vert \geq\delta_{\epsilon}$
for some $\delta_{\epsilon}>0$.
\item [{(ii)}] $S_{T\alpha}=\sum_{t=1}^{T}s_{t\alpha}$ is a sequence of
$\mathbb{R}^{r}$\textendash valued $\mathcal{F}_{T}$\textendash measurable
stochastic processes indexed by $\alpha\in A$; $S_{T\alpha}$ does
not depend on $(\beta,\pi,\varpi)$; $S_{T\alpha}$ has sample paths
that are continuous as functions of $\alpha$; the process $T^{-1/2}S_{T\alpha}$
obeys $T^{-1/2}S_{T\bullet}\Rightarrow S_{\bullet}$ for some mean
zero $\mathbb{R}^{r}$-valued Gaussian process $\{S_{\alpha}:\alpha\in A\}$
that satisfies $E[S_{\alpha}S_{\alpha}']=E[s_{t\alpha}s_{t\alpha}']=\mathcal{I}_{\alpha}$
for all $\alpha\in A$ and has continuous sample paths (as functions
of $\alpha$) with probability 1.
\item [{(iii)}] $\mathcal{I}_{\alpha}$ is, for each $\alpha\in A$, a
non-random symmetric $r\times r$ matrix (independent of $(\beta,\pi,\varpi)$);
$\mathcal{I}_{\alpha}$ is continuous as a function of $\alpha$ and
such that $0<\inf_{\alpha\in A}\lambda_{\min}(\mathcal{I}_{\alpha})$,
$\sup{}_{\alpha\in A}\lambda_{\max}(\mathcal{I}_{\alpha})<\infty$\emph{.}
\item [{(iv)}] $R_{T}(\alpha,\beta,\pi,\varpi)$ is a remainder term such
that
\[
\sup_{(\beta,\pi,\varpi)\in B\times\Pi_{\alpha}:\left\Vert (\beta,\pi,\varpi)-(\beta^{*},\pi^{\ast},0)\right\Vert \leq\gamma_{T}}\frac{\lvert R_{T}(\alpha,\beta,\pi,\varpi)\rvert}{(1+\lVert T^{1/2}\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)\rVert)^{2}}=o_{p\alpha}(1)
\]
for all sequences of (non-random) positive scalars $\{\gamma_{T},\ T\geq1\}$
for which $\gamma_{T}\rightarrow0$ as $T\rightarrow\infty$.
\end{description}
\end{assumption}
Assumption \ref{assu:quadexp}(i) describes the transformed parameter
$\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)$, with part (b) being
an identification condition. Assumption \ref{assu:quadexp}(ii) is
the main ingredient needed to derive the limiting distribution of
our LR test whereas \ref{assu:quadexp}(iv) ensures that the remainder
term $R_{T}(\alpha,\beta,\pi,\varpi)$ has no effect on the final
result. Assumption \ref{assu:quadexp}(iii) imposes rather standard
conditions on the counterpart of the information matrix.

As in \citet{andrews1999estimation,andrews2001testing}, \citet{zhu2006asymptotics},
and \citet{kasahara2012testing,kasahara2015testing}, for further
developments it will be convenient to write the expansion (\ref{eq:Quadr-Exp_1})
in an alternative form as
\begin{multline}
L_{T}^{\pi}(\alpha,\beta,\pi,\varpi)-L_{T}^{\pi}(\alpha,\beta^{*},\pi^{*},0)\\
=\frac{1}{2}Z_{T\alpha}^{\prime}\mathcal{I}_{\alpha}Z_{T\alpha}-\frac{1}{2}\bigl[T^{1/2}\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)-Z_{T\alpha}\bigr]'\mathcal{I}_{\alpha}\bigl[T^{1/2}\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)-Z_{T\alpha}\bigr]+R_{T}(\alpha,\beta,\pi,\varpi),\label{eq:Quadr-Exp_2}
\end{multline}
where $Z_{T\alpha}=\mathcal{I}_{\alpha}^{-1}T^{-1/2}S_{T\alpha}$.
Assumptions 5(ii) and (iii) imply the following facts (that will be
justified in the proof of Lemma \ref{lem:AsInsigOfRemainder} in Appendix
A): $Z_{T\alpha}$ is $\mathcal{F}_{T}$\textendash measurable, independent
of $(\beta,\pi,\varpi)$, continuous as a function of $\alpha$ with
probability 1, and $Z_{T\bullet}\Rightarrow Z_{\bullet}$ where the
mean zero $\mathbb{R}^{r}$\textendash valued Gaussian process $Z_{\alpha}=\mathcal{I}_{\alpha}^{-1}S_{\alpha}$
satisfies $E[Z_{\alpha}Z_{\alpha}']=\mathcal{I}_{\alpha}^{-1}$ for
all $\alpha\in A$ and has continuous sample paths (as functions of
$\alpha$) with probability 1.

\subsubsection{Examples (continued)}

\paragraph{LMAR Example.}

For the LMAR model, expansion (\ref{eq:Quadr-Exp_1}) (with the unnecessary
$\beta$ being dropped everywhere) is obtained from a standard second-order
Taylor expansion. Specifically, for an arbitrary fixed $\alpha\in A$,
a standard second-order Taylor expansion of $L_{T}^{\pi}(\alpha,\pi,\varpi)=\sum_{t=1}^{T}l_{t}^{\pi}(\alpha,\pi,\varpi)$
around $\left(\pi^{\ast},0\right)$ with respect to the parameters
$(\pi,\varpi)$ yields
\begin{align}
L_{T}^{\pi}(\alpha,\pi,\varpi)-L_{T}^{\pi}(\alpha,\pi^{\ast},0) & =(\pi-\pi^{\ast},\varpi)^{\prime}\nabla_{(\pi,\varpi)}L_{T}^{\pi}(\alpha,\pi^{\ast},0)\nonumber \\
 & \qquad+\frac{1}{2}(\pi-\pi^{\ast},\varpi)^{\prime}\nabla_{(\pi,\varpi)(\pi,\varpi)'}^{2}L_{T}^{\pi}(\alpha,\dot{\pi},\dot{\varpi})(\pi-\pi^{\ast},\varpi),\label{LMAR_TaylorExp}
\end{align}
where $(\dot{\pi},\dot{\varpi})$ denotes a point between $(\pi,\varpi)$
and $(\pi^{\ast},0)$, $\nabla_{(\pi,\varpi)}L_{T}^{\pi}(\alpha,\pi^{\ast},0)=\sum_{t=1}^{T}\nabla_{(\pi,\varpi)}l_{t}^{\pi}(\alpha,\pi^{\ast},0)$
and $\nabla_{(\pi,\varpi)(\pi,\varpi)'}^{2}L_{T}^{\pi}(\alpha,\pi,\varpi)=\sum_{t=1}^{T}\nabla_{(\pi,\varpi)(\pi,\varpi)'}^{2}l_{t}^{\pi}(\alpha,\pi,\varpi)$
(explicit expressions for the required derivatives are provided in
Appendix B), and $\nabla$ and $\nabla^{2}$ denote first and second
order differentiation with respect to the indicated parameters. Set
$\boldsymbol{\theta}(\alpha,\pi,\varpi)=(\pi-\pi^{\ast},\varpi)=(\theta,\vartheta)$
and note that the parameter space $\Theta_{\alpha}=\Theta$ is independent
of $\alpha$ and contains the origin, corresponding to the true model,
as an interior point. Then define the vector $S_{T\alpha}$ and the
matrix $\mathcal{I}_{\alpha}$ as\footnote{In what follows, $\nabla f_{t}(\cdot)$ denotes differentiation of
$f_{t}(\cdot)$ in (\ref{eq:f_t}) with respect to $\tilde{\phi}=(\tilde{\phi}_{0},\tilde{\phi}_{1},\ldots,\tilde{\phi}_{p},\tilde{\sigma}_{1}^{2})$.}
\begin{align}
S_{T\alpha} & =\nabla_{(\pi,\varpi)}L_{T}^{\pi}(\alpha,\pi^{\ast},0)=\sum_{t=1}^{T}\left(\frac{\nabla f_{t}(\pi^{\ast})}{f_{t}(\pi^{\ast})},-(1-\alpha_{1,t}^{L}(\alpha))\frac{\nabla f_{t}(\pi^{\ast})}{f_{t}(\pi^{\ast})}\right),\label{eq:LMAR_S_Ta}\\
\mathcal{I}_{\alpha} & =\begin{bmatrix}E\left[\frac{\nabla f_{t}(\pi^{\ast})}{f_{t}(\pi^{\ast})}\frac{\nabla'f_{t}(\pi^{\ast})}{f_{t}(\pi^{\ast})}\right] & -E\left[(1-\alpha_{1,t}^{L}(\alpha))\frac{\nabla f_{t}(\pi^{\ast})}{f_{t}(\pi^{\ast})}\frac{\nabla'f_{t}(\pi^{\ast})}{f_{t}(\pi^{\ast})}\right]\\
-E\left[(1-\alpha_{1,t}^{L}(\alpha))\frac{\nabla f_{t}(\pi^{\ast})}{f_{t}(\pi^{\ast})}\frac{\nabla'f_{t}(\pi^{\ast})}{f_{t}(\pi^{\ast})}\right] & E\left[(1-\alpha_{1,t}^{L}(\alpha))^{2}\frac{\nabla f_{t}(\pi^{\ast})}{f_{t}(\pi^{\ast})}\frac{\nabla'f_{t}(\pi^{\ast})}{f_{t}(\pi^{\ast})}\right]
\end{bmatrix}.\nonumber
\end{align}
Adding and subtracting terms and reorganizing, expansion (\ref{LMAR_TaylorExp})
can be written as
\begin{align}
L_{T}^{\pi}(\alpha,\pi,\varpi)-L_{T}^{\pi}(\alpha,\pi^{*},0) & =(T^{-1/2}S_{T\alpha})'[T^{1/2}\boldsymbol{\theta}(\alpha,\pi,\varpi)]\nonumber \\
 & \qquad-\frac{1}{2}[T^{1/2}\boldsymbol{\theta}(\alpha,\pi,\varpi)]^{\prime}\mathcal{I}_{\alpha}[T^{1/2}\boldsymbol{\theta}(\alpha,\pi,\varpi)]+R_{T}(\alpha,\pi,\varpi),\label{LMAR_QuadraticExp}
\end{align}
with the remainder term
\begin{equation}
R_{T}(\alpha,\pi,\varpi)=\frac{1}{2}[T^{1/2}\boldsymbol{\theta}(\alpha,\pi,\varpi)]^{\prime}[T^{-1}\nabla_{(\pi,\varpi)(\pi,\varpi)'}^{2}L_{T}^{\pi}(\alpha,\dot{\pi},\dot{\varpi})-(-\mathcal{I}_{\alpha})][T^{1/2}\boldsymbol{\theta}(\alpha,\pi,\varpi)].\label{LMAR_RemainderTerm}
\end{equation}
These equations yield the expansion (\ref{eq:Quadr-Exp_1}) in the
case of the LMAR model. For details of verifying Assumptions \ref{assu:cont_differentiability}
and \ref{assu:quadexp}, we refer to Appendix B.

As mentioned in the LMAR example of Section 3.1.1, the treatment of
the special case where the mixing weight $\alpha_{t}^{L}$ is constant
is more complicated than that of the `real' LMAR case. Indeed, replacing
the mixing weight $\alpha_{t}^{L}$ by a constant in the preceding
expression of the score vector $S_{T\alpha}$ immediately shows that
the second-order Taylor expansion (\ref{LMAR_QuadraticExp}) breaks
down because, contrary to Assumption 5(iii), the components of $S_{T\alpha}$
are not linearly independent and, consequently, the Fisher information
matrix $\mathcal{I}_{\alpha}$ is singular. Thus, a higher order Taylor
expansion is needed to analyze the LR test statistic.

To give an idea of how one could proceed, we first note that the partial
derivatives of the log-likelihood function behave in the same way
as their counterparts in \citet{kasahara2015testing} where mixtures
of normal regression models (with constant mixing weights) are considered
(see particularly the discussion following their Proposition 1). This
is due to the fact that in the special case of constant mixing weights
the LMAR model is obtained from the model considered in \citet{kasahara2015testing}
by replacing the exogenous regressors therein by lagged values of
$y_{t}$. Thus, the arguments employed in that paper could be used
to obtain the asymptotic distribution of the LR test statistic. Instead
of a conventional second-order Taylor expansion this would require
a more complicated reparameterization and an expansion based on partial
derivatives of the log-likelihood function up to order eight. As most
of the details appear very similar to those in \citet{kasahara2015testing}
we have preferred not to pursue this matter in this paper.

The preceding discussion means that, in the case of the LMAR model,
time varying mixing weights are beneficial when the purpose is to
derive a LR test for the adequacy of a single-regime model. A similar
observation was made already by \citet[p. 80]{jeffries1998logistic}.
However, this does not happen in all mixture autoregressive models
with time varying mixing weights, as the following discussion on the
GMAR model demonstrates.

\paragraph{GMAR Example. }

As alluded to earlier, in the case of the GMAR model the expansion
(\ref{eq:Quadr-Exp_1}) cannot be based on a second order Taylor expansion
of the log-likelihood function. A higher order expansion is required,
and similarly to \citet{kasahara2012testing} the appropriate order
turns out to be the fourth one with the elements of $\nabla_{\beta}l_{t}^{\pi}(\alpha,\beta^{*},\pi^{*},0)$
and $\nabla_{\pi}l_{t}^{\pi}(\alpha,\beta^{*},\pi^{*},0)$ and the
distinctive elements of $\nabla_{\varpi\varpi^{\prime}}l_{t}^{\pi}(\alpha,\beta^{*},\pi^{*},0)$
(suitably normalized) used to define the vector $S_{T\alpha}$. In
Appendix C we present, for an arbitrary fixed $\alpha\in A$, the
explicit form of a standard fourth-order Taylor expansion of $L_{T}^{\pi}(\alpha,\beta,\pi,\varpi)=\sum_{t=1}^{T}l_{t}^{\pi}(\alpha,\beta,\pi,\varpi)$
around $(\beta^{*},\pi^{*},0)$ with respect to the parameters $(\beta,\pi,\varpi)$.
Therein we also demonstrate that this fourth-order Taylor expansion
can be written as a quadratic expansion of the form (\ref{eq:Quadr-Exp_1})
(or (\ref{eq:Quadr-Exp_2})) with the different quantities appearing
therein defined as follows.

Define the vector $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)$
in (\ref{eq:Quadr-Exp_1}) as
\[
\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)=\begin{bmatrix}\theta(\alpha,\beta,\pi,\varpi)\\
\vartheta(\alpha,\beta,\pi,\varpi)
\end{bmatrix}=\begin{bmatrix}\beta-\beta^{\ast}\\
\pi-\pi^{\ast}\\
\alpha(1-\alpha)v(\varpi)
\end{bmatrix},
\]
where $\theta(\alpha,\beta,\pi,\varpi)$ is $(q_{1}+q_{2})\times1$
and $\vartheta(\alpha,\beta,\pi,\varpi)$ is $q_{\vartheta}\times1$
with $q_{\vartheta}=q_{2}(q_{2}+1)/2$ (where $q_{1}=1$ and $q_{2}=p+1$)
and where the vector $v(\varpi)$ contains the unique elements of
$\varpi\varpi^{\prime}$, that is,
\[
v(\varpi)=(\varpi_{1}^{2},\ldots,\varpi_{q_{2}}^{2},\varpi_{1}\varpi_{2},\ldots,\varpi_{1}\varpi_{q_{2}},\varpi_{2}\varpi_{3},\ldots,\varpi_{q_{2}-1}\varpi_{q_{2}})
\]
(note that $v(\varpi)$ is just a re-ordering of $vech(\varpi\varpi^{\prime})$).
The parameter space
\[
\Theta_{\alpha}=\{\boldsymbol{\theta}=(\theta,\vartheta)\in\mathbb{R}^{q_{1}+q_{2}+q_{\vartheta}}:\theta=(\beta-\beta^{\ast},\pi-\pi^{\ast}),\vartheta=\alpha(1-\alpha)v(\varpi)\text{ for some }(\beta,\pi,\varpi)\in B\times\Pi_{\alpha}\}
\]
now depends on $\alpha$ and has the origin, corresponding to the
true model, as a boundary point (due to the particular shape of the
range of $\vartheta=\alpha(1-\alpha)v(\varpi)$); both of these features
will complicate the subsequent analysis.

Next set
\begin{equation}
S_{T}\,(=S_{T\alpha})=\sum_{t=1}^{T}\tilde{\nabla}_{\boldsymbol{\theta}}l_{t}^{\pi\ast}\ \ \ \ \ \text{where }\,\,\,\,\,\tilde{\nabla}_{\boldsymbol{\theta}}l_{t}^{\pi\ast}=(\tilde{\nabla}_{\theta}l_{t}^{\pi\ast},\tilde{\nabla}_{\vartheta}l_{t}^{\pi\ast})\ \ \ \ \ ((q_{1}+q_{2}+q_{\vartheta})\times1)\label{eq:GMAR_S_T}
\end{equation}
with the component vectors $\tilde{\nabla}_{\theta}l_{t}^{\pi\ast}$
and $\tilde{\nabla}_{\vartheta}l_{t}^{\pi\ast}$ given by
\begin{align}
\tilde{\nabla}_{\theta}l_{t}^{\pi\ast} & =(\nabla_{\beta}l_{t}^{\pi}(\alpha,\beta^{*},\pi^{\ast},0),\nabla_{\pi}l_{t}^{\pi}(\alpha,\beta^{*},\pi^{\ast},0)),\nonumber \\
\tilde{\nabla}_{\vartheta}l_{t}^{\pi\ast} & =\bigl(c_{11}\nabla_{\varpi_{1}\varpi_{1}}^{2}l_{t}^{\pi}(\alpha,\beta^{*},\pi^{\ast},0),\ldots,c_{q_{2}q_{2}}\nabla_{\varpi_{q_{2}}\varpi_{q_{2}}}^{2}l_{t}^{\pi}(\alpha,\beta^{*},\pi^{\ast},0),\nonumber \\
 & \ \ \ \ \ c_{12}\nabla_{\varpi_{1}\varpi_{2}}^{2}l_{t}^{\pi}(\alpha,\beta^{*},\pi^{\ast},0),\ldots,c_{q_{2}-1,q_{2}}\nabla_{\varpi_{q_{2}-1}\varpi_{q_{2}}}^{2}l_{t}^{\pi}(\alpha,\beta^{*},\pi^{\ast},0)\bigr)/(\alpha(1-\alpha)),\label{eq:GMAR_S_T_component3}
\end{align}
where $c_{ij}=1/2$ if $i=j$ and $c_{ij}=1$ if $i\neq j$. Explicit
expressions for $\tilde{\nabla}_{\theta}l_{t}^{\pi\ast}$ and $\tilde{\nabla}_{\vartheta}l_{t}^{\pi\ast}$
can be found in Appendix C, and from them it can be seen that $S_{T}$
depends on $(\beta^{*},\pi^{\ast})$ only and not on $(\alpha,\beta,\pi,\varpi)$.
The same is true for the matrix $\mathcal{I}\,(=\mathcal{I_{\alpha}})$
($(q_{1}+q_{2}+q_{\vartheta})\times(q_{1}+q_{2}+q_{\vartheta})$)
whose expression is also given in Appendix C. Finally, an explicit
expression of the remainder term $R_{T}(\alpha,\beta,\pi,\varpi)$
is given in Appendix C. For the verification of Assumptions \ref{assu:cont_differentiability}
and \ref{assu:quadexp}, see Appendix C.

In the GMAR example we have assumed that the intercept terms $\tilde{\phi}_{0}$
and $\tilde{\varphi}_{0}$ in the two regimes are the same. We are
now in a position to describe the difficulties that allowing for $\tilde{\phi}_{0}\neq\tilde{\varphi}_{0}$
(and, hence, dropping $\beta$) would entail. In this case, the additional
parameter $\tilde{\varphi}_{0}$ would correspond to $\varpi_{1}$,
the first component of $\varpi$. As in Section 3.2.1, it would again
be the case that $\nabla_{\varpi_{1}}l_{t}^{\pi}(\alpha,\pi^{*},0)=0$,
leading us to consider second derivatives. But now, due to the properties
of the Gaussian distribution, it would be the case that $\nabla_{\varpi_{1}\varpi_{1}}^{2}l_{t}^{\pi}(\alpha,\pi^{*},0)$
is linearly dependent with the components of $\nabla_{\pi}l_{t}^{\pi}(\alpha,\pi^{*},0)$,
making it unsuitable to serve as a component of $S_{T}$. A reparameterization
more complicated than that used in Section 3.2.1 would be needed,
with the aim of obtaining $\nabla_{\varpi_{1}\varpi_{1}}^{2}l_{t}^{\pi}(\alpha,\pi^{*},0)=0$
and, instead of $\nabla_{\varpi_{1}\varpi_{1}}^{2}l_{t}^{\pi}(\alpha,\pi^{*},0)$,
using $\nabla_{\varpi_{1}\varpi_{1}\varpi_{1}}^{3}l_{t}^{\pi}(\alpha,\pi^{*},0)$
or perhaps $\nabla_{\varpi_{1}\varpi_{1}\varpi_{1}\varpi_{1}}^{4}l_{t}^{\pi}(\alpha,\pi^{*},0)$
as the counterpart of the score of the parameter $\varpi_{1}$. It
turns out that (restricting the discussion to the case $p=1$ only)
the third derivative is suitable when $\alpha\neq1/2$ and $\tilde{\phi}_{1}\neq-1/2$,
but that fourth (or higher) order derivatives are needed when $\alpha=1/2$
or $\tilde{\phi}_{1}=-1/2$. Similar difficulties (involving situations
comparable to the cases $\alpha=1/2$ vs. $\alpha\neq1/2$, but apparently
not ones involving also a counterpart of $\tilde{\phi}_{1}$) were
faced by \citet[Sec. 2.3.3]{cho2007testing} and \citet[Sec. 6.2]{Kasahara2017Testing},
whose analyses suggest that expanding the log-likelihood at least
to the eighth order is required. As the required analysis gets excessively
complicated, we have chosen to leave it for future research.

\subsection{Asymptotic analysis of the quadratic expansion}

We continue by analyzing the expansion (\ref{eq:Quadr-Exp_2}) evaluated
at $(\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})$.
Previously, a similar analysis is provided by \citet{andrews2001testing}
but his approach is not directly applicable in our setting. The reason
for this is that in the quadratic expansion in (\ref{eq:Quadr-Exp_2})
the dependence of the parameter $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)$
and its parameter space $\Theta_{\alpha}$ on the nuisance parameter
$\alpha$ is not compatible with the formulation of \citet[eqn (3.3)]{andrews2001testing}.
The results of \citet{zhu2006asymptotics} probably cover our case,
but instead of trying to verify the assumptions employed by these
authors we prove the needed results by adapting the arguments used
in \citet{andrews1999estimation,andrews2001testing} and \citet{zhu2006asymptotics}
to our setting. We proceed in several steps.

\paragraph{Asymptotic insignificance of the remainder term.}

We first establish that the remainder term $R_{T}(\alpha,\beta,\pi,\varpi)$,
when evaluated at $(\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})$,
has no effect on the asymptotic distribution of the quadratic expansion.
A crucial ingredient in showing this is showing that the transformed
parameter vector $\boldsymbol{\theta}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})$
is root-$T$ consistent in the sense that $\lVert T^{1/2}\boldsymbol{\theta}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})\rVert=O_{p\alpha}(1)$.
This, together with part (iv) of Assumption \ref{assu:quadexp} allows
us to obtain the result $R_{T}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})=o_{p\alpha}(1)$.
We collect these results in the following lemma.
\begin{lem}
\label{lem:AsInsigOfRemainder}If Assumptions 1\textendash 5 hold\,\footnote{Here and in what follows, a subset of the listed assumptions would
sometimes suffice for the stated results.}, then (i) $\lVert T^{1/2}\boldsymbol{\theta}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})\rVert=O_{p\alpha}(1)$,
\\
(ii) $R_{T}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})=o_{p\alpha}(1)$,
and (iii)
\begin{align}
 & L_{T}^{\pi}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})-L_{T}^{\pi}(\alpha,\beta^{*},\pi^{*},0)\nonumber \\
 & =\frac{1}{2}Z_{T\alpha}^{\prime}\mathcal{I}_{\alpha}Z_{T\alpha}-\frac{1}{2}\bigl[T^{1/2}\boldsymbol{\theta}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})-Z_{T\alpha}\bigr]'\mathcal{I}_{\alpha}\bigl[T^{1/2}\boldsymbol{\theta}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})-Z_{T\alpha}\bigr]+o_{p\alpha}(1).\label{eq:Lemma 2 kaava}
\end{align}
\end{lem}
Note that assertion (iii) of Lemma \ref{lem:AsInsigOfRemainder} is
analogous to \citet[Theorem 2b]{andrews1999estimation}.

\paragraph{Maximization of the likelihood vs. minimization of a related quadratic
form.}

The first two terms on the right hand side of (\ref{eq:Quadr-Exp_2})
provide an approximation to the (reparameterized and centered) log-likelihood
function $L_{T}^{\pi}(\alpha,\beta,\pi,\varpi)-L_{T}^{\pi}(\alpha,\beta^{*},\pi^{*},0)$
evaluated at the (approximate unrestricted reparameterized) ML estimator
$(\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})$
in equation (\ref{eq:Lemma 2 kaava}). For later developments it would
be convenient if the ML estimator on the right hand side of (\ref{eq:Lemma 2 kaava})
could be replaced by an (approximate) minimizer of the quadratic form
$[T^{1/2}\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)-Z_{T\alpha}]'\mathcal{I}_{\alpha}[T^{1/2}\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)-Z_{T\alpha}]$.
In order to justify this replacement we first note that, by definition,
(cf. \citet[eqn. (3.6)]{andrews1999estimation})
\[
\inf_{(\beta,\pi,\varpi)\in B\times\Pi_{\alpha}}\bigl\{\bigl[T^{1/2}\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)-Z_{T\alpha}\bigr]'\mathcal{I}_{\alpha}\bigl[T^{1/2}\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)-Z_{T\alpha}\bigr]\bigr\}=\inf_{\boldsymbol{\lambda}\in\Theta_{\alpha,T}}\left\{ (\boldsymbol{\lambda}-Z_{T\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\lambda}-Z_{T\alpha})\right\}
\]
where, for each $T$,
\[
\Theta_{\alpha,T}=\{\boldsymbol{\lambda}\in\mathbb{R}^{r}:\boldsymbol{\lambda}=T^{1/2}\boldsymbol{\theta}\text{ for some }\boldsymbol{\theta}\in\Theta_{\alpha}\}
\]
and $\Theta_{\alpha}$ is as defined in Assumption \ref{assu:quadexp}(i).
Next, for each $\alpha\in A$, let $\boldsymbol{\hat{\lambda}}_{T\alpha q}=T^{1/2}\boldsymbol{\theta}(\alpha,\hat{\beta}_{T\alpha q},\hat{\pi}_{T\alpha q},\hat{\varpi}_{T\alpha q})$
(with the additional `$q$' in the subscripts referring to quadratic
form) denote an approximate minimizer of $(\boldsymbol{\lambda}-Z_{T\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\lambda}-Z_{T\alpha})$
over $\Theta_{\alpha,T}$, that is, (cf. \citet[eqn. (12)]{zhu2006asymptotics})
\begin{equation}
(\boldsymbol{\hat{\lambda}}_{T\alpha q}-Z_{T\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\hat{\lambda}}_{T\alpha q}-Z_{T\alpha})=\inf_{\boldsymbol{\lambda}\in\Theta_{\alpha,T}}\left\{ (\boldsymbol{\lambda}-Z_{T\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\lambda}-Z_{T\alpha})\right\} +o_{p\alpha}(1).\label{eq:QuadFormMinimizer}
\end{equation}
Now we can state the following lemma justifying the discussed replacement
of $(\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})$
with $\boldsymbol{\hat{\lambda}}_{T\alpha q}$.
\begin{lem}
\label{lem:MaxLikVsMinQuadForm}If Assumptions 1\textendash 5 hold,
then
\[
L_{T}^{\pi}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})-L_{T}^{\pi}(\alpha,\beta^{*},\pi^{*},0)=\frac{1}{2}Z_{T\alpha}^{\prime}\mathcal{I}_{\alpha}Z_{T\alpha}-\frac{1}{2}(\boldsymbol{\hat{\lambda}}_{T\alpha q}-Z_{T\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\hat{\lambda}}_{T\alpha q}-Z_{T\alpha})+o_{p\alpha}(1).
\]
\end{lem}

\paragraph{Approximating the parameter space with a cone.}

In the previous subsection the quadratic form $(\boldsymbol{\lambda}-Z_{T\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\lambda}-Z_{T\alpha})$
was minimized over the set $\Theta_{\alpha,T}$ which can be complicated
and hence difficult to use. Therefore we next show that the quadratic
form $(\boldsymbol{\lambda}-Z_{T\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\lambda}-Z_{T\alpha})$
can instead be minimized over a simpler set, and to this end we first
introduce some terminology.

We say that a collection of sets $\{\Gamma_{\alpha},\ \alpha\in A\}$
(where for each $\alpha\in A$, $\Gamma_{\alpha}\subset\mathbb{R}^{r}$)
is `locally (at the origin) uniformly equal' to a set $\Lambda\subset\mathbb{R}^{r}$
if there exists a $\delta>0$ such that $\Gamma_{\alpha}\cap(-\delta,\delta)^{r}=\Lambda\cap(-\delta,\delta)^{r}$
for all $\alpha\in A$. Note that `$\{\Gamma_{\alpha},\ \alpha\in A\}$
is locally uniformly equal to $\Lambda$' implies that (i) `for all
$\alpha\in A$, $\Gamma_{\alpha}$ is locally equal to $\Lambda$
in the sense of \citet[p.�1359]{andrews1999estimation}', but the
reverse does not hold; and also that (ii) $\{\Gamma_{\alpha},\ \alpha\in A\}$
is uniformly approximated by the set $\Lambda$ in the sense of \citet[Defn. 3]{zhu2006asymptotics}.
Finally, we say that a set $\Lambda\subset\mathbb{R}^{r}$ is a `cone'
if $\lambda\in\Lambda$ implies that $a\lambda\in\Lambda$ for all
positive real scalars $a$.

Based on the preceding discussion we state the following assumption.
\begin{assumption}
\label{assu:cone}The collection of sets $\{\Theta_{\alpha},\ \alpha\in A\}$
is locally uniformly equal to a cone $\Lambda\;(\subset\mathbb{R}^{r})$.
\end{assumption}
Note that by Assumption \ref{assu:quadexp}(i)(a), $0\in\Theta_{\alpha}$
for all $\alpha\in A$, so that the cone $\Lambda$ in Assumption
\ref{assu:cone} necessarily contains $0\;(\in\mathbb{R}^{r})$. The
cone $\Lambda$ also does not depend on $\alpha.$ Now we can establish
the following result.
\begin{lem}
\label{lem:Cone}If Assumptions 1\textendash 6 hold, then
\[
\inf_{\boldsymbol{\lambda}\in\Theta_{\alpha,T}}\left\{ (\boldsymbol{\lambda}-Z_{T\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\lambda}-Z_{T\alpha})\right\} =\inf_{\boldsymbol{\lambda}\in\Lambda}\left\{ (\boldsymbol{\lambda}-Z_{T\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\lambda}-Z_{T\alpha})\right\} +o_{p\alpha}(1).
\]
\end{lem}

\paragraph{Describing the limiting random variable.}

From Lemmas \ref{lem:MaxLikVsMinQuadForm} and \ref{lem:Cone} and
the definition of $\boldsymbol{\hat{\lambda}}_{T\alpha q}$ we can
now conclude that
\begin{equation}
2[L_{T}^{\pi}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})-L_{T}^{\pi}(\alpha,\beta^{*},\pi^{*},0)]=Z_{T\alpha}^{\prime}\mathcal{I}_{\alpha}Z_{T\alpha}-\inf_{\boldsymbol{\lambda}\in\Lambda}\left\{ (\boldsymbol{\lambda}-Z_{T\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\lambda}-Z_{T\alpha})\right\} +o_{p\alpha}(1).\label{eq:BeforeLemmaWeakConv}
\end{equation}
The assumed weak convergence of $S_{T\alpha}$ (and hence that of
$Z_{\alpha}=\mathcal{I}_{\alpha}^{-1}S_{\alpha}$) allows us to derive
the weak limit of this random process described in the following lemma.
\begin{lem}
\label{lem:WeakConv}If Assumptions 1\textendash 6 hold, then
\[
2[L_{T}^{\pi}(\bullet,\hat{\beta}_{T\bullet},\hat{\pi}_{T\bullet},\hat{\varpi}_{T\bullet})-L_{T}^{\pi}(\bullet,\beta^{*},\pi^{\ast},0)]\Rightarrow Z_{\bullet}^{\prime}\mathcal{I}_{\bullet}Z_{\bullet}-\inf_{\boldsymbol{\lambda}\in\Lambda}\left\{ (\boldsymbol{\lambda}-Z_{\bullet})^{\prime}\mathcal{I}_{\bullet}(\boldsymbol{\lambda}-Z_{\bullet})\right\} .
\]
\end{lem}
The limiting random process in Lemma \ref{lem:WeakConv} can be written
in a somewhat simpler form (cf. \citet[Thm�4]{andrews1999estimation}
and \citet[Thm 2]{andrews2001testing}). The motivation for this comes
from the fact that in our applications $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)$
can be decomposed into two parts as $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)=(\theta(\alpha,\beta,\pi,\varpi),\vartheta(\alpha,\beta,\pi,\varpi))$
with $\theta(\alpha,\beta,\pi,\varpi)\in\mathbb{R}^{q_{\theta}}$
and $\vartheta(\alpha,\beta,\pi,\varpi)\in\mathbb{R}^{q_{\vartheta}}$
(with $q_{\theta}=q_{1}+q_{2}$ and $r=q_{\theta}+q_{\vartheta}$)
such that (i) the values of $\theta(\alpha,\beta,\pi,\varpi)$ are
not restricted by the null hypothesis and do not lie on the boundary
of the parameter space and (ii) the values of $\vartheta(\alpha,\beta,\pi,\varpi)$
are restricted by the null hypothesis and potentially lie on the boundary
of the parameter space. Specifically, we assume the following.
\begin{assumption}
\label{assu:cone2}The cone $\Lambda$ of Assumption \ref{assu:cone}
satisfies $\Lambda=\mathbb{R}^{q_{\theta}}\times\Lambda_{\vartheta}$
with $\Lambda_{\vartheta}$ a cone in $\mathbb{R}^{q_{\vartheta}}$.
\end{assumption}
Partition $S_{\alpha}$, $Z_{\alpha}$, $\boldsymbol{\lambda}$, and
$\mathcal{I}_{\alpha}$ conformably with the partition $\boldsymbol{\theta}(\alpha,\beta,\pi,\varpi)=(\theta(\alpha,\beta,\pi,\varpi),\vartheta(\alpha,\beta,\pi,\varpi))$
as
\[
S_{\alpha}=\left[\begin{array}{c}
S_{\theta\alpha}\\
S_{\vartheta\alpha}
\end{array}\right],\,\,\,Z_{\alpha}=\left[\begin{array}{c}
Z_{\theta\alpha}\\
Z_{\vartheta\alpha}
\end{array}\right],\,\,\,\boldsymbol{\lambda}=\begin{bmatrix}\boldsymbol{\lambda}_{\theta}\\
\boldsymbol{\lambda}_{\vartheta}
\end{bmatrix},\,\,\,\mathcal{I}_{\alpha}=\left[\begin{array}{cc}
\mathcal{I}_{\theta\theta\alpha} & \mathcal{I}_{\theta\vartheta\alpha}\\
\mathcal{I}_{\vartheta\theta\alpha} & \mathcal{I}_{\vartheta\vartheta\alpha}
\end{array}\right]
\]
and let $(\mathcal{I}_{\alpha}^{-1})_{\vartheta\vartheta}$ denote
the ($q_{\vartheta}\times q_{\vartheta}$) bottom right block of $\mathcal{I}_{\alpha}^{-1}$.
Assumption \ref{assu:cone2} together with properties of partitioned
matrices yields the following result.
\begin{lem}
\label{lem:WeakLimitSubvec}If Assumptions 1\textendash 7 hold, then
\begin{align*}
 & Z_{\alpha}'\mathcal{I}_{\alpha}Z_{\alpha}-\inf_{\boldsymbol{\lambda}\in\Lambda}\left\{ (\boldsymbol{\lambda}-Z_{\alpha})^{\prime}\mathcal{I}_{\alpha}(\boldsymbol{\lambda}-Z_{\alpha})\right\} \\
 & =Z_{\vartheta\alpha}'(\mathcal{I}_{\alpha}^{-1})_{\vartheta\vartheta}^{-1}Z_{\vartheta\alpha}-\inf_{\boldsymbol{\lambda}_{\vartheta}\in\Lambda_{\vartheta}}\left\{ (\boldsymbol{\lambda}_{\vartheta}-Z_{\vartheta\alpha})'(\mathcal{I}_{\alpha}^{-1})_{\vartheta\vartheta}^{-1}(\boldsymbol{\lambda}_{\vartheta}-Z_{\vartheta\alpha})\right\} +S_{\theta\alpha}'\mathcal{I}_{\theta\theta\alpha}^{-1}S_{\theta\alpha}.
\end{align*}
\end{lem}
Explicit expressions for $(\mathcal{I}_{\alpha}^{-1})_{\vartheta\vartheta}$
and $Z_{\vartheta\alpha}$ in terms of $S_{\alpha}$ and $\mathcal{I}_{\alpha}$
are given in the proof of this lemma in Supplementary Appendix D.

\subsection{The LR test statistic}

\subsubsection{Derivation of the test statistic}

The previous subsection described the asymptotic behavior of $2[L_{T}^{\pi}(\alpha,\hat{\beta}_{T\alpha},\hat{\pi}_{T\alpha},\hat{\varpi}_{T\alpha})-L_{T}^{\pi}(\alpha,\beta^{*},\pi^{*},0)]$,
the first term in the expression of $LR_{T}(\alpha)$ in (\ref{eq:LR_pi-repam}).
Now consider the second term, namely $2[L_{T}^{0}(\hat{\tilde{\phi}}_{T})-L_{T}^{0}(\tilde{\phi}^{\ast})]$,
corresponding to the model restricted by the null hypothesis. Recall
that $L_{T}^{0}(\tilde{\phi})=\sum_{t=1}^{T}l_{t}^{0}(\tilde{\phi})$
with $l_{t}^{0}(\tilde{\phi})=\log[f_{t}(\tilde{\phi})]$ so that
$\nabla_{\tilde{\phi}}l_{t}^{0}(\tilde{\phi}^{\ast})=(f_{t}(\tilde{\phi}^{*}))^{-1}\nabla f_{t}(\tilde{\phi}^{*})$
with $\tilde{\phi}^{*}$ an interior point of $\tilde{\Phi}$. Denote
the score vector and limiting information matrix by
\[
S_{T}^{0}=\sum_{t=1}^{T}\frac{\nabla f_{t}(\tilde{\phi}^{\ast})}{f_{t}(\tilde{\phi}^{\ast})},\,\,\,\,\,\mathcal{I}^{0}=E\biggl[\frac{\nabla f_{t}(\tilde{\phi}^{\ast})}{f_{t}(\tilde{\phi}^{\ast})}\frac{\nabla'f_{t}(\tilde{\phi}^{\ast})}{f_{t}(\tilde{\phi}^{\ast})}\biggr],
\]
respectively. For the following assumption, partition the process
$S_{T\alpha}$ of Assumption \ref{assu:quadexp} as $S_{T\alpha}=(S_{T\theta\alpha},S_{T\vartheta\alpha})$
(with $S_{T\theta\alpha}$ $q_{\theta}$\textendash dimensional and
$S_{T\vartheta\alpha}$ $q_{\vartheta}$\textendash dimensional).
The following simplifying assumption, which holds in our examples
(see the expressions of $S_{T\alpha}$ in (\ref{eq:LMAR_S_Ta}) and
(\ref{eq:GMAR_S_T})), allows us to obtain a neat expression for the
likelihood ratio test statistic in Theorem 1 below.
\begin{assumption}
\label{assu:ScoreEqualsLinearScore}$S_{T\theta\alpha}=S_{T}^{0}$.
\end{assumption}
Together with the earlier assumptions, Assumption \ref{assu:ScoreEqualsLinearScore}
implies that $T^{-1/2}S_{T\theta\alpha}=T^{-1/2}S_{T}^{0}\stackrel{d}{\to}S^{0}$,
a $q_{\theta}$\textendash dimensional Gaussian random vector with
mean zero and covariance matrix $\mathcal{I}_{\theta\theta\alpha}=E[S^{0}S^{0\prime}]=\mathcal{I}^{0}$.
Standard likelihood theory now implies the following result.
\begin{lem}
\label{lem:WeakConvLinear}If Assumptions 1\textendash 8 hold, then
$2[L_{T}^{0}(\hat{\tilde{\phi}}_{T})-L_{T}^{0}(\tilde{\phi}^{\ast})]\overset{d}{\rightarrow}S^{0\prime}(\mathcal{I}^{0})^{-1}S^{0}$,
and the convergence is joint with that in Lemma \ref{lem:WeakConv}.
\end{lem}
The preceding results, in particular Lemmas \ref{lem:WeakConv}, \ref{lem:WeakLimitSubvec},
and \ref{lem:WeakConvLinear}, now yield the distribution of the LR
test statistic in the following theorem.
\begin{thm}
If Assumptions 1\textendash 8 hold, then
\begin{description}
\item [{(i)}] $LR_{T}\left(\bullet\right)\Rightarrow Z_{\vartheta\bullet}'(\mathcal{I}_{\bullet}^{-1})_{\vartheta\vartheta}^{-1}Z_{\vartheta\bullet}-\inf_{\boldsymbol{\lambda}_{\vartheta}\in\Lambda_{\vartheta}}\{(\boldsymbol{\lambda}_{\vartheta}-Z_{\vartheta\bullet})'(\mathcal{I}_{\bullet}^{-1})_{\vartheta\vartheta}^{-1}(\boldsymbol{\lambda}_{\vartheta}-Z_{\vartheta\bullet})\}$,
and
\item [{(ii)}] $LR_{T}=\sup_{\alpha\in A}LR_{T}\left(\alpha\right)\overset{d}{\rightarrow}\sup_{\alpha\in A}\bigl\{ Z_{\vartheta\alpha}'(\mathcal{I}_{\alpha}^{-1})_{\vartheta\vartheta}^{-1}Z_{\vartheta\alpha}-\inf_{\boldsymbol{\lambda}_{\vartheta}\in\Lambda_{\vartheta}}\{(\boldsymbol{\lambda}_{\vartheta}-Z_{\vartheta\alpha})'(\mathcal{I}_{\alpha}^{-1})_{\vartheta\vartheta}^{-1}(\boldsymbol{\lambda}_{\vartheta}-Z_{\vartheta\alpha})\}\bigr\}$.
\end{description}
\end{thm}
This completes the derivation of the LR test statistic. The asymptotic
distribution is similar to that in \citet[Thm 4]{andrews2001testing}.
As we next discuss, this distribution simplifies in both the LMAR
and the GMAR examples.

\subsubsection{Examples (continued) }

\paragraph{LMAR Example. }

As was noted in Section 3.3.1, the LMAR case is rather standard in
the sense that a conventional second-order Taylor expansion with a
nonsingular information matrix and with no parameters on the boundary
was sufficient to study the LR test. The only nonstandard feature
in this case is the presence of unidentified parameters under the
null hypothesis. Validity of Assumptions \ref{assu:cone}\textendash \ref{assu:ScoreEqualsLinearScore}
is easy to check (see Appendix B) with the cone $\Lambda$ of Assumption
\ref{assu:cone} equal to $\mathbb{R}^{r}$. Thus the infimum in the
distribution of the $LR_{T}$ statistic in Theorem 1(ii) equals zero
and the result therein simplifies to\footnote{This result could also be obtained from \citet[Sec 2.2, 2.4]{andrews1995admissibility}
as their assumptions 1\textendash 5 appear to be satisfied in the
LMAR case.}
\[
LR_{T}=\sup_{\alpha\in A}LR_{T}\left(\alpha\right)\overset{d}{\rightarrow}\sup_{\alpha\in A}\left\{ Z_{\vartheta\alpha}'(\mathcal{I}_{\alpha}^{-1})_{\vartheta\vartheta}^{-1}Z_{\vartheta\alpha}\right\} .
\]
For every fixed $\alpha\in A$, the quantity $Z_{\vartheta\alpha}'(\mathcal{I}_{\alpha}^{-1})_{\vartheta\vartheta}^{-1}Z_{\vartheta\alpha}$
is a chi-squared random variable, so that the limiting distribution
is a supremum of a chi-squared process similarly as in, for example,
\citet{davies1987hypothesis}, \citet[Thm 1]{hansen1996inference},
and \citet[eqn. (5.7)]{andrews2001testing}.

\paragraph{GMAR Example. }

\noindent In Section 3.3.1 it was seen that in the GMAR example $Z_{\alpha}$
and $\mathcal{I}_{\alpha}$ do not depend on $\alpha$. As the cone
$\Lambda$ of Assumption \ref{assu:cone} does not depend on $\alpha$
either, the weak limit of $LR_{T}\left(\alpha\right)$ does not depend
on $\alpha$. Therefore the result of Theorem 1 (validity of Assumptions
\ref{assu:cone}\textendash \ref{assu:ScoreEqualsLinearScore} is
checked in Appendix C) simplifies to
\[
LR_{T}=\sup_{\alpha\in A}LR_{T}\left(\alpha\right)\overset{d}{\rightarrow}Z_{\vartheta}'(\mathcal{I}^{-1})_{\vartheta\vartheta}^{-1}Z_{\vartheta}-\inf_{\boldsymbol{\lambda}_{\vartheta}\in\Lambda_{\vartheta}}\left\{ (\boldsymbol{\lambda}_{\vartheta}-Z_{\vartheta})'(\mathcal{I}^{-1})_{\vartheta\vartheta}^{-1}(\boldsymbol{\lambda}_{\vartheta}-Z_{\vartheta})\right\} ,
\]
where the unnecessary $\alpha$ has been dropped from the notation.
Here $Z_{\vartheta}$ follows an $q_{\vartheta}$-variate Gaussian
distribution with covariance matrix $(\mathcal{I}^{-1})_{\vartheta\vartheta}$,
and the limiting distribution, which is sometimes referred to as the
chi-bar-squared distribution, is similar to the one in \citet[Proposition 3c,d]{kasahara2012testing}.
Note that the cone $\Lambda_{\vartheta}=v(\mathbb{\mathbb{R}}^{q_{2}})$
(see Appendix C) is not convex (in contrast to (at least most of)
the examples in \citet{andrews2001testing}, but similarly to \citet[Proposition 3c,d]{kasahara2012testing})
and the dimension of this cone, $q_{\vartheta}=q_{2}(q_{2}+1)/2$,
may not be small either ($q_{\vartheta}=3,6,10,\ldots$ for $q_{2}=2,3,4,\ldots$).

\section{Simulation-based critical values and a Monte Carlo study}

\subsection{Simulating the asymptotic null distribution}

Similarly to \citet{hansen1996inference} and \citet{andrews2001testing},
the asymptotic null distribution of the LR statistic in Theorem 1
is typically application-specific and cannot be tabulated. Following
these papers, we use simulation methods to obtain critical values
of the asymptotic null distribution. The following procedure is based
on \citet{hansen1996inference} and is analogous to the one used by
\citet[Sec 2.1]{zhu2004hypothesis} in a related mixture setting.\footnote{An alternative to this procedure is to use bootstrap. However, the
validity of bootstrap in the presence of parameters on the boundary
and singular information matrices is not clear (see, e.g., \citet{andrews2000inconsistency}).
Another reason for preferring the proposed simulation method is that
repeated estimation of the mixture model under the alternative may
be computationally rather demanding.}

Let $A_{G}$ be some finite grid of $\alpha$ values in $A$. For
each fixed $\alpha\in A_{G}$, let $\hat{s}_{t\alpha}$ signify an
empirical counterpart of $s_{t\alpha}$ (see Assumption 5) where the
unknown parameter $\tilde{\phi}^{*}$ (or $(\beta^{*},\pi^{*})$)
is replaced by its consistent estimator under the null, $\hat{\tilde{\phi}}_{T}$.
(The specific forms of $\hat{s}_{t\alpha}$ in the LMAR and GMAR examples
are provided in Appendices B and C, respectively.) Set $\hat{\mathcal{I}}_{T\alpha}=T^{-1}\sum_{t=1}^{T}\hat{s}_{t\alpha}\hat{s}_{t\alpha}'$.
Now, for each $j=1,\ldots,J$ (where $J$ denotes the number of repetitions),
do the following.
\begin{description}
\item [{(i)}] Generate a sequence $\{v_{tj}\}_{t=1}^{T}$ of $T$ i.i.d.
$N(0,1)$ random variables.
\item [{(ii)}] For each $\alpha\in A_{G}$, set $\hat{S}_{T\alpha}^{j}=\sum_{t=1}^{T}\hat{s}_{t\alpha}v_{tj}$,
$\hat{Z}_{T\alpha}^{j}=\hat{\mathcal{I}}_{T\alpha}^{-1}T^{-1/2}\hat{S}_{T\alpha}^{j}$,
and (using similar partitioning notation as before)
\[
\widehat{LR}_{T}^{j}(\alpha)=\hat{Z}_{T\vartheta\alpha}^{j\prime}(\hat{\mathcal{I}}_{T\alpha}^{-1})_{\vartheta\vartheta}^{-1}\hat{Z}_{T\vartheta\alpha}^{j}-\inf_{\boldsymbol{\lambda}_{\vartheta}\in\Lambda_{\vartheta}}\bigl\{(\boldsymbol{\lambda}_{\vartheta}-\hat{Z}_{T\vartheta\alpha}^{j})'(\hat{\mathcal{I}}_{T\alpha}^{-1})_{\vartheta\vartheta}^{-1}(\boldsymbol{\lambda}_{\vartheta}-\hat{Z}_{T\vartheta\alpha}^{j})\bigr\};
\]
here the minimization of the quadratic form over the cone $\Lambda_{\vartheta}$
has to be performed numerically.
\item [{(iii)}] Set $\widehat{LR}_{T,A_{G}}^{j}=\max_{\alpha\in A_{G}}\widehat{LR}_{T}^{j}(\alpha)$.
\end{description}
This yields a sample $\{\widehat{LR}_{T,A_{G}}^{1},\ldots,\widehat{LR}_{T,A_{G}}^{J}\}$
of $J$ realizations. An approximate $p$\textendash value corresponding
to an observed LR test statistic $LR_{T}$ is computed as $J^{-1}\sum_{j=1}^{J}\mathbf{1}(\widehat{LR}_{T,A_{G}}^{j}>LR_{T})$
(here $\mathbf{1}(\cdot)$ denotes the indicator function). The precision
of this approximation can be controlled by choosing $J$ large enough,
see \citet{hansen1996inference} (in the illustration below we use
$J=1000$).

\subsection{A small Monte Carlo study}

We now study the finite sample properties of the LR test statistics
and the simulation-based critical values. The results are presented
in Table 1. We consider two LR test statistics, one based on an estimated
LMAR model, and another based on an estimated GMAR model (as in our
two examples in the preceding sections). In all simulations, we use
an autoregressive order $p=1$, $J=1000$ repetitions (see the previous
subsection), and three different sample sizes: $T=250$, $500$, and
$1000$.

The top part of Table 1 presents results for size simulations. Data
is generated from an AR(1) model (for a range of different parameter
values shown in Table 1) and AR(1), LMAR(1), and GMAR(1) models are
estimated (LMAR with $m=1$; GMAR with the restriction $\tilde{\phi}_{0}=\tilde{\varphi}_{0}$;
in estimation of the mixture models we use a genetic algorithm as
singularity of the information matrix may render gradient based methods
unreliable). Two LR test statistics are calculated based on the estimated
LMAR and GMAR models, respectively, and labelled `LMAR $LR_{T}$'
and `GMAR $LR_{T}$'. Simulation-based $p$\textendash values are
computed based on the asymptotic distributions in Section 3.5.2 and
using the simulation procedure in Section 4.1. Using nominal levels
$10\%$, $5\%$, and $1\%$, a reject/not-reject decision is recorded.
This exercise is repeated $1000$ times, and the six rightmost columns
in Table 1 present the empirical rejection frequencies (for the LMAR
$LR_{T}$ and GMAR $LR_{T}$ tests and the three nominal levels used).

As can be seen from the results in Table 1 (top part), the LMAR $LR_{T}$
test's size is satisfactory overall, typically being somewhat oversized
for sample sizes $T=250$ and $500$, and somewhat conservative for
the largest sample size ($T=1000$). The parameter values used in
simulation do not seem to have a large effect on the size. The GMAR
$LR_{T}$ test, on the other hand, appears to be moderately oversized
across all sample sizes and parameter values used.

The lower part of Table 1 presents results for power simulations.
Data is generated either from a GMAR model or from an LMAR model (for
a range of different parameter values shown in Table 1), and empirical
rejection frequencies are calculated as above. Both the LMAR $LR_{T}$
test and the GMAR $LR_{T}$ test appear to have good overall power.
As expected, when the two regimes differ more from each other, the
tests have higher power, and the same happens when sample size is
increased. Besides having good power against the `right' alternatives,
the tests also turn out to have decent power against `wrong' alternatives:
When data is generated from the GMAR (resp., LMAR) model, the LMAR
$LR_{T}$ (resp., GMAR $LR_{T}$) test rejects reasonably often (the
GMAR $LR_{T}$ test in particular seems capable of picking up LMAR
type regime switching). Naturally, the power of the tests may be inflated
due to the tests being oversized.

As a computational remark we note that the LMAR $LR_{T}$ and GMAR
$LR_{T}$ tests and their $p$\textendash values are reasonably straightforward
to compute in a matter of seconds using a standard, modern desktop
computer (for one particular model, one particular sample size, and
$J=1000$ repetitions). The GMAR $LR_{T}$ test is computationally
more demanding than the LMAR $LR_{T}$ test as it involves the minimization
of a quadratic form over a cone which is not needed in the LMAR case
(see Sections 3.5.2 and 4.1); this is also one potential reason for
the less precise size of the GMAR $LR_{T}$ test.

\begin{table}[p]
\includegraphics[scale=0.88]{SimulationResultsPrelim}
\end{table}

\section{Conclusions}

This paper has studied the asymptotic distribution of the LR test
statistic for testing a linear autoregressive model against a two-regime
mixture autoregressive model. A distinguishing feature of the paper
is that the regime switching probabilities are observation-dependent.
Technical challenges resulting from unidentified parameters under
the null, parameters on the boundary, and singularity of the information
matrix were dealt with by considering an appropriately reparameterized
model and higher-order expansions of the log-likelihood function.
The resulting asymptotic distribution of the LR test statistic is
non-standard and application-specific. Critical values can be obtained
by a straightforward simulation procedure, and a Monte Carlo study
indicated the proposed tests to have satisfactory size and power properties.

The general theory of the paper was illustrated using two concrete
examples, the LMAR model of \citet{wong2001logistic} and (a version
of the) GMAR model of \citet{kalliovirta2015gaussian}. Considering
other mixture AR models, as well as the general GMAR model, is left
for future research. This paper was concerned with testing linearity
against a two-regime model, and considering tests of $M\geq2$ regimes
versus $M+1$ regimes, similarly as in \citet{kasahara2015testing}
in a related setting, forms another interesting research topic.

\pagebreak{}