Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Model Selection for Explosive Models
abstractThis paper examines the limit properties of information criteria (such as
AIC, BIC, HQIC) for distinguishing between the unit root model and the
various kinds of explosive models. The explosive models include the
local-to-unit-root model, the mildly explosive model and the regular
explosive model. Initial conditions with different order of magnitude are
considered. Both the OLS estimator and the indirect inference estimator are
studied. It is found that BIC\ and HQIC, but not AIC, consistently select
the unit root model when data come from the unit root model. When data come
from the local-to-unit-root model, both BIC and HQIC\ select the wrong model
with probability approaching 1 while AIC has a positive probability of
selecting the right model in the limit. When data come from the regular
explosive model or from the mildly explosive model in the form of $
1+n^{\Greekmath 010B }/n$ with $\Greekmath 010B \in (0,1)$, all three information criteria
consistently select the true model. Indirect inference estimation can
increase or decrease the probability for information criteria to select the
right model asymptotically relative to OLS, depending on the information
criteria and the true model. Simulation results confirm our asymptotic
results in finite sample.
Keywords: Model Selection; Information Criteria; Local-to-unit-root
Model; Mildly Explosive Model; Unit Root Model; Indirect Inference.
Introduction
Information criteria have found a wide range of practical applications in
empirical work. Examples include choosing explanatory variables in
regression models and selecting lag lengths in time series models.
Frequently used information criteria are AIC of Akaike (1969, 1973), BIC of
Schwarz (1978), HQIC of Hannan and Quinn (1979). A major nice feature in
these information criteria is that the penalty term is trivial to compute
and hence the implementation of them is straightforward and can be made
automatic.
With a growing interest in nonstationarity in time series analysis,
researchers have examined the properties of information criteria in the
context of nonstationary models with the unit root behavior. An important
form of nonstationarity in time series involves explosive roots. Recent
global financial crisis has motivated researchers to study explosive
behavior in economic and financial time series; see, for example, Phillips
and Yu (2011), Phillips, Wu and Yu (2011) and Phillips, Shi and Yu (2015a,
b).
In this paper, we study the limit properties of information criteria for
distinguishing between the unit root model and the explosive models. The
information criteria considered in this paper have a general form and
include AIC, BIC and HQIC as the special cases. The impact of the initial
condition on the limit properties is examined by allowing for an initial
condition of three different orders of magnitude. Moreover, both the OLS
estimator and the indirect inference estimator are studied when
investigating the limit properties of information criteria. The motivation
for the use of indirect inference estimator comes from the existence of
finite sample bias in the OLS estimator and the ability that the indirect
inference method can reduce the bias.
It is found that information criteria consistently choose the unit root
model against the explosive alternatives when data comes from the unit root
model. Second, we prove that the probability for information criteria to
correctly select the explosive model models against the unit root model
depends crucially on both the degree of explosiveness and the size of the
penalty term in information criteria. Finally and surprisingly, we show that
indirect inference estimation can increase or decrease the probability for
information criteria to select the right model asymptotically relative to
OLS, depending on the information criteria and the true model.
The rest of this paper is organized as follows. Section 2 introduces the
models and information criteria, and briefly reviews the literature. Section
(ref) gives the limit properties of information criteria for
distinguishing models with an explosive root from the unit root model when
the OLS\ estimator is used. Section 4 gives the limit properties of
information criteria when the indirect inference estimator\ is used. Section
(ref) provides Monte Carlo evidence to support the theoretical
results. Section (ref) concludes. All the detailed proofs are
provided in the appendix. To compress notation, we denote $
\int\nolimits_{0}^{1}BdB$ and $\int\nolimits_{0}^{1}B^{2}$ in short for $
\int\nolimits_{0}^{1}B(r)dB(r)$ and $\int\nolimits_{0}^{1}B(r)^{2}dr$
respectively throughout the paper, and $\Rightarrow $ denotes weak
convergence.
Models, Information Criteria and A Literature Review
The model considered in the present paper is of the form:
equation[equation omitted — 153 chars of source]
where $u_{t}\overset{iid}{\sim }(0,\Greekmath 011B ^{2})$ and the model is
initialized at $t=0$ with some $X_{0}$. The autoregressive (AR) coefficient $
\Greekmath 011A _{n}$ is the crucial parameter that determines the dynamic behavior of $
X_{t}$. When $\Greekmath 011A _{n}=\Greekmath 011A $ and $\left\vert \Greekmath 011A \right\vert <1$, $X_{t}$
is stationary. When $\Greekmath 011A _{n}=1$, $X_{t}$ has a unit root (UR hereafter).
When $\Greekmath 011A _{n}=1-c_{n}/n=1-c/n$ for $c>0$, $X_{t}$ is near-stationary and
has a root that is local-to-unity (LTUS hereafter) (Phillips, 1987b; Chan
and Wei, 1987). When $\Greekmath 011A _{n}=\Greekmath 011A $ and $\left\vert \Greekmath 011A \right\vert >1$,
$X_{t}$ has an explosive root (EX hereafter). When $\Greekmath 011A
_{n}=1+c_{n}/n=1+c/n $ for $c>0$, $X_{t}$ is near-explosive and also has a
root that is local-to-unity (LTUE hereafter). When $\Greekmath 011A _{n}=1-c_{n}/n$ for
$c_{n}\rightarrow \infty $ but $c_{n}/n\searrow 0$, the root represents
moderate deviations from unity and $X_{t}$ is near-stationary (Phillips and
Magdalinos, 2007). When $\Greekmath 011A _{n}=1+c_{n}/n$ for $c_{n}\rightarrow \infty $
but $c_{n}/n\searrow 0$, $X_{t}$ is mildly explosive (hereafter ME).
The asymptotic properties of the OLS\ estimator of the AR coefficient in the
stationary AR(1) model is well known. The rate of convergence is $\sqrt{n}$
and the limiting distribution is Gaussian. Phillips (1987a) provided the
limiting theory for the OLS\ estimator in the UR model and the rate of
convergence is $n$. Phillips (1987b) and Chan and Wei (1987) established the
asymptotic theory for the LTUS and LTUE models. The asymptotic theory is
similar to that in the UR model and the rate of convergence is also $n$. In
the cases of UR and LTU, $u_{t}$ can be weakly dependent stationary.
Anderson (1959) studied the limiting distribution of the OLS\ estimator in
the EX model under the condition that $u_{t}\overset{iid}{\sim }\mathcal{N}
(0,\Greekmath 011B ^{2})$ and $X_{0}=0$. The limiting distribution is Cauchy and the
rate of convergence is $\Greekmath 011A ^{n}$. However, no invariance principle
applies. Assuming $X_{0}=o_{p}(\sqrt{n/c_{n}})$, Phillips and Magdalinos
(2007) developed the asymptotic theory for the model with $\Greekmath 011A
_{n}=1-c_{n}/n$ for $c_{n}\rightarrow \infty $ but $c_{n}/n\searrow 0$ and
showed that the asymptotic distribution is invariant to the error
distribution. The rate of convergence is $n/\sqrt{c_{n}}$. If $
c_{n}=n^{\Greekmath 010B }$ with $\Greekmath 010B \in (0,1)$, this rate of convergence bridges
that of UR/LTU models and that of the stationary process. Phillips and
Magdalinos (2007) also developed the asymptotic theory for the ME model. The
rate of convergence is $n\Greekmath 011A _{n}^{n}/c_{n}$. The limiting distribution is
Cauchy which is the same as in the EX model. Interestingly, in the ME case,
the asymptotic theory is independent of the initial condition as long as $
X_{0}=o_{p}(\sqrt{n/c_{n}})$.
It is known that the OLS estimator of $\Greekmath 011A _{n}$ is biased downward when $
\Greekmath 011A _{n}=1$ or when $\Greekmath 011A _{n}$ is in the vicinity of unity. In this case,
the indirect inference estimation is effective in reducing the bias.
Phillips (2012) derives the asymptotic theory of the indirect inference
estimator when the model is UR or LTU and $u_{t}\overset{iid}{\sim }\mathcal{
N}(0,\Greekmath 011B ^{2})$. The rate of convergence remains unchanged while the
limiting distribution is different from that of the OLS estimator.
Information criteria for model selection have been proposed by Akaike (1969,
1973), Schwarz (1978), Hannan and Quinn (1979), among many others. The
general form of these criteria is
equation*[equation* omitted — 82 chars of source]
where $k$ is the number of parameters to be estimated, $\widehat{\Greekmath 011B }
_{k}^{2}$ is the estimated $\Greekmath 011B ^{2}$ when $k$ parameters are estimated.
In general, $IC_{k}$ trades off the term that measures the goodness-of-fit
(i.e. $\log \widehat{\Greekmath 011B }_{k}^{2}$) and the penalty term that measures
the complexity of the model (i.e. $kp_{n}/n$). Coefficient $p_{n}=2,\log
n,2\log \log n$ corresponds to AIC of Akaike (1973), BIC\ of Schwarz (1978)
and HQIC of Hannan and Quinn (1979). Other forms of $p_{n}$ are possible.
In the time series literature, information criteria have been widely used to
select the lag length both in the family of stationary models and in the
family of nonstationary models; see for example, Ng and Perron (1995) and
Ploberger and Phillips (2003). The information criteria can also be used to
evaluate whether $\Greekmath 011A _{n}=1$ (i.e. $k=0$) or $\Greekmath 011A _{n}\neq 1$ (i.e. $k=1$
) in Model ((ref)). For example, Phillips (2008) obtained limit
properties of $IC_{k}$ for distinguishing between the unit root model and
the stationary model. Phillips and Lee (2015) show that BIC can successfully
distinguish the UR model from the ME model. This is a surprising result as
it is well known that BIC cannot consistently distinguish between the UR\
model and the LTU model; see Ploberger and Phillips (2003).
In this paper we focus our attention to distinguishability between the unit
root model and the three explosive models (i.e., LTUE, ME and EX) after the
candidate models are estimated by OLS or by the indirect inference method.
As a result, we make contributions in two strands of literature, explosive
time series and indirect inference.
To visually understand the difference between the UR\ model, the LTU model
and the ME model, we simulate a sample path of different length ($
n=100,200,500,1000$) with $y_{0}=0$, based on the same realizations of the
error process, iid $\mathcal{N}(0,1)$, from the following four models, $\Greekmath 011A
_{n}=1$ (UR), $\Greekmath 011A _{n}=1+1/n$ (LTUE), $\Greekmath 011A _{n}=1+n^{0.1}/n$ (ME1), and $
\Greekmath 011A _{n}=1+n^{0.5}/n$ (ME2). Figures 1-3 give the time series plot of UR
against LTU, UR against ME1, UR against ME2, respectively. It can be seen
from Figure 1 that it is very difficult to distinguish between the UR
process and the LTU process, even when the sample size is as large as 1,000.
When the sample size increases, the gap between the UR\ process and the two
ME processes becomes larger and larger, as apparent in Figure 2 and more so
in Figure 3.
figure[figure omitted — 180 chars of source]
figure[figure omitted — 193 chars of source]
figure[figure omitted — 191 chars of source]
Limit Properties Based on the OLS Estimator
When the data generating process (DGP) is the UR model, since $\Greekmath 011A _{n}=1$,
we set the parameter count to $k=0$. For the LTU model, the ME model and the
EX model, we need to estimate the AR coefficient and hence set the parameter
count to $k=1$. Throughout the paper we denote $\widehat{\Greekmath 011A }$ the OLS
estimator of $\Greekmath 011A $. $\widehat{k}_{IC}=0$ or $1$ means the information
criterion of the UR\ model is smaller or larger than that of the competing
model when $\Greekmath 011A $ is estimated by OLS. We aim to find the limit of the
following probabilities:
align[align omitted — 333 chars of source]
As shown in Phillips and Magdalinos (2009), the unit root asymptotic
distribution is sensitive to initial conditions in the distant past. To
understand how the initial condition affects the property of $\widehat{k}
_{IC}$, we follow Phillips and Magdalinos (2009) by assuming alternative
initial conditions.
assumption[IN]
The initial condition has the form
\begin{equation}
X_{0}(n)=\sum_{j=0}^{\Greekmath 0114 _{n}}u_{-j},
\end{equation}
where $\Greekmath 0114 _{n}$ is a sequence of integers satisfying $\Greekmath 0114
_{n}\rightarrow \infty $ and
\begin{equation}
\frac{\Greekmath 0114 _{n}}{n}\rightarrow \Greekmath 011C \in \left[ 0,\infty \right] \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, as
}n\rightarrow \infty .
\end{equation}
The following cases are distinguished:
\begin{enumerate}
• If $\Greekmath 011C = 0$, $X_0(n) $ is said to be a recent past
initialization.
• If $\Greekmath 011C \in \left(0, \infty\right)$, $X_0(n) $ is said to be a
distant past initialization.
• If $\Greekmath 011C =\infty $, $X_{0}(n)$ is said to be an infinite past
initialization.
\end{enumerate}
theoremUnder Assumption (ref) (i) or (ii) or (iii), we have
\begin{enumerate}
• when $p_{n}\rightarrow \infty $ and $p_{n}/n\rightarrow 0$ as $
n\rightarrow \infty $,
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=0\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ IC_{0}-IC_{1}\leq 0\right\} =1,
\\
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=0\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ IC_{0}-IC_{1}>0\right\} =0.
\end{align*}
• when $p_{n}=2$, the asymptotic distribution under the AIC
criterion is
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{AIC}=0|k=0\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ AIC_{0}-AIC_{1}\leq 0\right\}
=P\left( \Greekmath 0118 ^{2}<2\right) , \\
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{AIC}=1|k=0\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ AIC_{0}-AIC_{1}>0\right\}
=1-P\left( \Greekmath 0118 ^{2}<2\right) .
\end{align*}
where
\begin{equation*}
\Greekmath 0118 ^{2}=
\begin{cases}
\dfrac{\left( \int_{0}^{1}BdB\right) ^{2}}{\int_{0}^{1}B^{2}}, & \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }
\Greekmath 011C =0 \\
\dfrac{\left( \int_{0}^{1}B_{\Greekmath 011C }dB\right) ^{2}}{\int_{0}^{1}B_{\Greekmath 011C }^{2}}
, & \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }\Greekmath 011C \in (0,\infty ) \\
B(1)^{2}, & \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }\Greekmath 011C =\infty
\end{cases}
,
\end{equation*}
with $B(s)$ being a Brownian motion, and
\begin{equation*}
B_{\Greekmath 011C }(s)=B(s)+\sqrt{\Greekmath 011C }B_{0}(1),
\end{equation*}
with $B_{0}(s)$ being an independent Brownian motion.
\end{enumerate}
remarkTheorem (ref) is the same as Theorem 1 in Phillips (2008)
for distinguishing between the UR model and the stationary model. The
condition that $p_{n}\rightarrow \infty $ and $p_{n}/n\rightarrow 0$ covers
BIC and HQIC and hence, both BIC and HQIC can consistently select the UR
model. The AIC criterion is inconsistent and its asymptotic distribution
depends on $\Greekmath 0118 ^{2}$, the squared unit root $t$-statistic for the OLS
estimator.
remarkThe validity of Theorem (ref) does not require the iid
assumption for the error term $u_{t}$. If we follow Phillips (2008) by
denoting $F(L)=\sum_{j=0}^{\infty }F_{j}L^{j}$, with $F_{0}=1$ and $F(1)\neq
0$, and letting $u_{s}$ have Wold representation
\begin{equation}
u_{s}=F(L)\Greekmath 0122 _{s}=\sum_{j=0}^{\infty }F_{j}\Greekmath 0122 _{s-j}\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{
, with }\sum_{j=0}^{\infty }j^{1/2}\left\vert F_{j}\right\vert <\infty ,
\end{equation}
where $\Greekmath 0122 _{t}\overset{iid}{\sim }\left( 0,\Greekmath 011B _{\Greekmath 0122
}^{2}\right) $, the results in Theorem (ref) continue to hold.
However, both $B_{0}$ and $\Greekmath 0118 ^{2}$ need to be modified to accommodate the
dependence in $u_{t}$ as in Phillips (2008).
theoremLet Assumption (ref) (i) or (ii) holds. Assume the true
DGP is the LTUE model.
\begin{enumerate}
• When $p_{n}\rightarrow \infty $ and $p_{n}/n\rightarrow 0$ as $
n\rightarrow \infty $,
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left(
IC_{1}-IC_{0}\right) >0\right\} =1, \\
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left(
IC_{1}-IC_{0}\right) \leq 0\right\} =0.
\end{align*}
• When $p_{n}=2$, the asymptotic distribution of the AIC criterion
is
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{AIC}=0|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ {n}\left( AIC_{1}-AIC_{0}\right)
>0\right\} =1-P\left( \Greekmath 0110 ^{2}>2\right) , \\
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{AIC}=1|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ {n}\left( AIC_{1}-AIC_{0}\right)
\leq 0\right\} =P\left( \Greekmath 0110 ^{2}>2\right) ,
\end{align*}
where
\begin{equation*}
\Greekmath 0110 ^{2}=\frac{\left( \int_{0}^{1}J_{c}dB\right) ^{2}}{
\int_{0}^{1}J_{c}^{2}}+2{c}\int_{0}^{1}J_{c}dB+c^{2}\int_{0}^{1}J_{c}^{2},
\end{equation*}
with
\begin{equation*}
J_{c}(r)=\int_{0}^{r}\exp \left\{ c(r-s)\right\} dB(s).
\end{equation*}
\end{enumerate}
remarkTheorem (ref) shows that all the information criteria
are inconsistent in distinguishing between the LTUE model and the UR models
when data comes from the LTUE model. AIC selects the wrong model with
probability going to $1-P\left( \Greekmath 0110 ^{2}>2\right) $, which depends on the
localization constant $c$. This problem worsens for BIC and HQIC as the
probability of selecting the wrong model goes to one. Note that BIC is well
known to be blind to local alternatives; see, for example, Ploberger and
Phillips (2003).
theoremLet Assumption (ref) (i) or (ii) holds. Assume the true
DGP is the ME model.
\begin{enumerate}
• When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A
_{n}^{2n}}=0,$
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left(
IC_{1}-IC_{0}\right) >0\right\} =0, \\
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left(
IC_{1}-IC_{0}\right) \leq 0\right\} =1.
\end{align*}
• When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A
_{n}^{2n}}=\Greekmath 0119 \in (0,+\infty ),$
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left(
IC_{1}-IC_{0}\right) >0\right\} =P\left( \Greekmath 011F ^{2}(1)<4\Greekmath 0119 \right) , \\
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left(
IC_{1}-IC_{0}\right) \leq 0\right\} =1-P\left( \Greekmath 011F ^{2}(1)<4\Greekmath 0119 \right) .
\end{align*}
• When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A
_{n}^{2n}}\rightarrow +\infty ,$
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left(
IC_{1}-IC_{0}\right) >0\right\} =1, \\
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left(
IC_{1}-IC_{0}\right) \leq 0\right\} =0.
\end{align*}
\end{enumerate}
remarkTheorem (ref) shows that the limit probability of
selecting the correct model by information criteria under the ME model
depends critically on two parameters, $c_{n}$, $p_{n}$. As expected, the
larger $c_{n}$, the further the model away from the UR\ model and the higher
probability for the information criteria to select the correct model.
Interestingly, the smaller $p_{n}$, the higher probability for the
information criteria to select the correct model. From Phillips and
Magdalinos (2009), we know $\Greekmath 011A _{n}^{-n}=o(c_{n}^{-1})$ and hence $\Greekmath 011A
_{n}^{n}/c_{n}\rightarrow +\infty $. In the special case where $
c_{n}=n^{\Greekmath 010B }$, for $\Greekmath 010B \in (0,1)$, $\lim\limits_{n\rightarrow
\infty }p_{n}/\Greekmath 011A _{n}^{2n}=0$ no matter whether $p_{n}=2$ or $\log n$ or $
2\log \log n$. In this case, all the well-known information criteria can
consistently select the true model.
theoremLet Assumption (ref) (i) holds. Assume the true DGP is
the EX model.
\begin{enumerate}
• When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A ^{2n}}
=0,$
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A ^{2n}}\left(
IC_{1}-IC_{0}\right) >0\right\} =0, \\
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A ^{2n}}\left(
IC_{1}-IC_{0}\right) \leq 0\right\} =1.
\end{align*}
• When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A ^{2n}}
=\Greekmath 0119 \in (0,+\infty ),$
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A ^{2n}}\left(
IC_{1}-IC_{0}\right) >0\right\} =P\left( \Greekmath 011F ^{2}(1)<(1+\Greekmath 011A )^{2}\Greekmath 0119
\right) , \\
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A ^{2n}}\left(
IC_{1}-IC_{0}\right) \leq 0\right\} =1-P\left( \Greekmath 011F ^{2}(1)<(1+\Greekmath 011A )^{2}\Greekmath 0119
\right) .
\end{align*}
• When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A ^{2n}}
\rightarrow +\infty ,$
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left(
IC_{1}-IC_{0}\right) >0\right\} =1, \\
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left(
IC_{1}-IC_{0}\right) \leq 0\right\} =0.
\end{align*}
\end{enumerate}
remarkTheorem (ref) shows that the limit probability of
selecting the correct model by information criteria under the EX model
depends also critically on two parameters, $\Greekmath 011A $, $p_{n}$. As expected,
the larger $\Greekmath 011A $, the higher probability for the information criteria to
select the correct model. Interestingly, the smaller $p_{n}$, the higher
probability for the information criteria to select the correct model. If $
p_{n}=2$ or $\log n$ or $2\log \log n$, $\lim\limits_{n\rightarrow \infty
}p_{n}/\Greekmath 011A ^{2n}=0$ and hence case (1) applies, suggesting that all the
well-known information criteria can consistently select the true model.
Results in Theorem (ref) can be extended to cover the LTUE model and
the ME model with weakly dependent errors. The following proposition
establishes the results for the ME model.
propositionLet\ Assumption (ref) (i) or (ii) and the
assumption specified in Equation ((ref)) hold. Assume the true DGP is
the ME model.
\begin{enumerate}
• When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A
_{n}^{2n}}=0,$
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left(
IC_{1}-IC_{0}\right) >0\right\} =0, \\
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left(
IC_{1}-IC_{0}\right) \leq 0\right\} =1.
\end{align*}
• When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A
_{n}^{2n}}=\Greekmath 0119 \in (0,+\infty ),$
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left(
IC_{1}-IC_{0}\right) >0\right\} =P\left( \Greekmath 011F ^{2}(1)<\frac{4\Greekmath 0119 }{\Greekmath 0121
^{2}}\right) , \\
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left(
IC_{1}-IC_{0}\right) \leq 0\right\} =1-P\left( \Greekmath 011F ^{2}(1)<\frac{4\Greekmath 0119 }{
\Greekmath 0121 ^{2}}\right) .
\end{align*}
where $\Greekmath 0121 ^{2}=\left( \sum\nolimits_{j=0}^{\infty }F_{j}\right) ^{2}.$
• When $\dfrac{p_{n}}{\Greekmath 011A _{n}^{2n}}\rightarrow +\infty ,$
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left(
IC_{1}-IC_{0}\right) >0\right\} =1, \\
\lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left(
IC_{1}-IC_{0}\right) \leq 0\right\} =0.
\end{align*}
\end{enumerate}
Limit Properties Based on the Indirect Inference Estimator
The OLS estimator of $\Greekmath 011A _{n}$ in Model ((ref)) is known to be biased
and the bias is acute when $\Greekmath 011A _{n}$ is close to unity. To reduce the
bias, the indirect inference method of Smith (1993) and Gour\'{e}rioux et al
(1993) can be used if Model ((ref)) is fully specified. Phillips (2012)
derives the asymptotic theory of the indirect inference estimator when the
model is UR or LTU and $u_{t}\overset{iid}{\sim }\mathcal{N}(0,\Greekmath 011B ^{2})$
. Throughout the paper we denote $\breve{\Greekmath 011A }$ the indirect inference
estimator of $\Greekmath 011A $. Let $h(c)=c+g(c)$ and $g(c)=g^{-}(c)1_{\{c\leq
0\}}+g^{+}(c)1_{\{c>0\}}$ with
alignat*{2}
g^{-}(c)& = & & -\dfrac{3}{4}\int_{0}^{\infty }e^{-\frac{v}{4}
}k^{-}(v;c)^{1/2}dv+\dfrac{1}{4}\int_{0}^{\infty }e^{-\frac{v}{4}
}k^{-}(v;c)^{3/2}dv \\
& & & -\dfrac{e^{2c}}{8}\int_{0}^{\infty }e^{-\frac{5v}{4}
}k^{-}(v;c)^{3/2}vdv, \\
g^{+}(c)& = & & \dfrac{3}{4}\int_{0}^{\infty }e^{\frac{w}{4}
}k^{+}(w;c)^{1/2}dw-\dfrac{1}{4}\int_{0}^{\infty }e^{\frac{w}{4}
}k^{+}(w;c)^{3/2}dw \\
& & & -\dfrac{e^{2c}}{8}\int_{0}^{\infty }e^{\frac{5w}{4}
}k^{+}(w;c)^{3/2}wdw, \\
k^{-}(v;c)& = & & \dfrac{2v-4c}{v+e^{2c}ve^{-v}-4c}, \\
k^{+}(w;c)& = & & \dfrac{2w+4c}{w+e^{2c}we^{w}+4c}.
Phillips (2012) shows that under the UR model,
equation*[equation* omitted — 218 chars of source]
and under the LTUE model,
equation*[equation* omitted — 250 chars of source]
Let $\breve{k}_{IC}=0$ or $1$ mean the information criterion of the UR\
model is smaller or larger than that of the competing model when the model
is estimated by the indirect inference method. We aim to find is the limit
of the following probabilities:
align[align omitted — 325 chars of source]
theoremUnder Assumption (ref)(i) or (ii) or (iii), we have
\begin{enumerate}
• when $p_{n}\rightarrow \infty $ and $p_{n}/n\rightarrow 0$ as $
n\rightarrow \infty $,
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=0|k=0\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ IC_{0}-IC_{1}\leq 0\right\} =1,
\\
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=1|k=0\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ IC_{0}-IC_{1}>0\right\} =0;
\end{align*}
• when $p_{n}=2$, the asymptotic distribution under the AIC
criterion is
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{AIC}=0|k=0\right\} &
=P\left( \Greekmath 0126 ^{2}<2\right) , \\
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{AIC}=1|k=0\right\} &
=1-P\left( \Greekmath 0126 ^{2}<2\right) ,
\end{align*}
where
\begin{equation*}
\Greekmath 0126 ^{2}=
\begin{cases}
\int_{0}^{1}B^{2}\cdot h^{-1}\left( \left( \dfrac{\int_{0}^{1}BdB}{
\int_{0}^{1}B^{2}}\right) ^{2}\right) -2\int_{0}^{1}BdB\cdot h^{-1}\left(
\dfrac{\int_{0}^{1}BdB}{\int_{0}^{1}B^{2}}\right) , & \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }\Greekmath 011C =0 \\
\int_{0}^{1}B_{\Greekmath 011C }^{2}\cdot h^{-1}\left( \left( \dfrac{
\int_{0}^{1}B_{\Greekmath 011C }dB}{\int_{0}^{1}B_{\Greekmath 011C }^{2}}\right) ^{2}\right)
-2\int_{0}^{1}B_{\Greekmath 011C }dB\cdot h^{-1}\left( \dfrac{\int_{0}^{1}B_{\Greekmath 011C }dB}{
\int_{0}^{1}B_{\Greekmath 011C }^{2}}\right) , & \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }\Greekmath 011C \in (0,\infty ) \\
h^{-1}\left( \mathcal{C}\right) ^{2}B_{0}^{2}(1)-2h^{-1}\left( \mathcal{C}
\right) B(1)B_{0}(1), & \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }\Greekmath 011C =\infty
\end{cases}
,
\end{equation*}
with $\mathcal{C}$ being a standard Cauchy variate.
\end{enumerate}
remarkAccording to Theorem (ref), as long as $p_{n}\rightarrow
\infty $ and $p_{n}/n\rightarrow 0$, information criteria based on the
indirect inference estimator is consistent in selecting the UR\ model.
Hence, BIC and HQIC based on the indirect inference estimator can
consistently select the UR model. Like the AIC criterion that is based on
the OLS estimator, the AIC criterion based on the indirect inference
estimator continues to be inconsistent. However, its asymptotic distribution
depends on $\Greekmath 0126 ^{2}$, the squared unit root $t$-statistic for the
indirect inference estimator.
remarkAs shown in Phillips (2012), the squared unit root $t$
-statistic for the indirect inference estimator has a smaller variance than
that of the squared unit root $t$-statistic for the OLS estimator.
Consequently, $P\left( \Greekmath 0126 ^{2}<2\right) >P\left( \Greekmath 0118 ^{2}<2\right) $,
suggesting that AIC based on the indirect inference estimator can select the
true model (i.e. the UR model) with a larger probability than that based on
the OLS estimator.
theoremLet Assumption (ref) (i) or (ii) holds. Assume the true
DGP is the LTUE model.
\begin{enumerate}
• When $p_{n}\rightarrow \infty $ and $p_{n}/n\rightarrow 0$ as $
n\rightarrow \infty $,
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=0|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left(
IC_{1}-IC_{0}\right) >0\right\} =1, \\
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=1|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left(
IC_{1}-IC_{0}\right) \leq 0\right\} =0.
\end{align*}
• When $p_{n}=2$, the asymptotic distribution under the AIC
criterion is
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{AIC}
}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ {n}\left(
AIC_{1}-AIC_{0}\right) >0\right\} =1-P\left( \Greekmath 0123 ^{2}>2\right) , \\
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{AIC}
}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ {n}\left(
AIC_{1}-AIC_{0}\right) \leq 0\right\} =P\left( \Greekmath 0123 ^{2}>2\right) ,
\end{align*}
where
\begin{equation*}
\Greekmath 0123 ^{2}\equiv 2h^{-1}\left( \dfrac{\int_{0}^{1}J_{c}dB}{
\int_{0}^{1}J_{c}^{2}}+c\right) \left(
\int_{0}^{1}J_{c}dB+c\int_{0}^{1}J_{c}^{2}\right) -h^{-1}\left( \dfrac{
\int_{0}^{1}J_{c}dB}{\int_{0}^{1}J_{c}^{2}}+c\right)
^{2}\int_{0}^{1}J_{c}^{2}.
\end{equation*}
\end{enumerate}
remarkTheorem (ref) shows that all the information criteria
continue to be inconsistent in distinguishing between the LTUE model and the
UR models when data come from the LTUE model even when the indirect
inference estimation is employed. AIC selects the wrong model with
probability going to $1-P\left( \Greekmath 0123 ^{2}>2\right) $. Since the
variance of $\Greekmath 0110 ^{2}$ is bigger than that of $\Greekmath 0123 ^{2}$, the tail
probability of $\Greekmath 0110 ^{2}$ is larger than that of $\Greekmath 0123 ^{2}$,
suggesting that AIC based on OLS selects the true model (i.e. LTUE model)
with a greater\ probability than AIC\ based on the indirect inference
estimator. This is a rather surprising result and suggests that the
superiority in estimation does not necessarily translate to the superiority
in model selection.
theoremLet Assumption (ref) (i) or (ii) holds. Assume the true
DGP is the ME model.
\begin{enumerate}
• When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A
_{n}^{2n}}=0,$
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=0|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left(
IC_{1}-IC_{0}\right) >0\right\} =0, \\
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=1|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left(
IC_{1}-IC_{0}\right) \leq 0\right\} =1.
\end{align*}
• When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A
_{n}^{2n}}=\Greekmath 0119 \in (0,+\infty ),$
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=0|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left(
IC_{1}-IC_{0}\right) >0\right\} =P\left( \Greekmath 011F ^{2}(1)<4\Greekmath 0119 \right) , \\
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=1|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left(
IC_{1}-IC_{0}\right) \leq 0\right\} =1-P\left( \Greekmath 011F ^{2}(1)<4\Greekmath 0119 \right) .
\end{align*}
• When $\dfrac{p_{n}}{\Greekmath 011A _{n}^{2n}}\rightarrow +\infty ,$
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=0|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left(
IC_{1}-IC_{0}\right) >0\right\} =1, \\
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=1|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left(
IC_{1}-IC_{0}\right) \leq 0\right\} =0.
\end{align*}
\end{enumerate}
remarkThe results in Theorem (ref) are the same as those in
Theorem (ref), suggesting all the well-known information criteria can
consistently select the true model (i.e. ME model) when $c_{n}=n^{\Greekmath 010B }$,
for $\Greekmath 010B \in (0,1)$.
theoremLet Assumption (ref) (i) holds. Assume the true DGP is
the EX model.
\begin{enumerate}
• When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A ^{2n}}
=0,$
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=0|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A ^{2n}}\left(
IC_{1}-IC_{0}\right) >0\right\} =0, \\
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=1|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A ^{2n}}\left(
IC_{1}-IC_{0}\right) \leq 0\right\} =1.
\end{align*}
• When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A ^{2n}}
=\Greekmath 0119 \in (0,+\infty ),$
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=0|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A ^{2n}}\left(
IC_{1}-IC_{0}\right) >0\right\} =P\left( \Greekmath 011F ^{2}(1)<(1+\Greekmath 011A )^{2}\Greekmath 0119
\right) , \\
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=1|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A ^{2n}}\left(
IC_{1}-IC_{0}\right) \leq 0\right\} =1-P\left( \Greekmath 011F ^{2}(1)<(1+\Greekmath 011A )^{2}\Greekmath 0119
\right) .
\end{align*}
• When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A ^{2n}}
\rightarrow +\infty ,$
\begin{align*}
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=0|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left(
IC_{1}-IC_{0}\right) >0\right\} =1, \\
\lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=1|k=1\right\} &
=\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left(
IC_{1}-IC_{0}\right) \leq 0\right\} =0.
\end{align*}
\end{enumerate}
remarkThe results in Theorem (ref) are the same as those in
Theorem (ref), suggesting that all the well-known information
criteria can consistently select the true model (i.e. EX model).
Monte Carlo Study
In this section, we examine the performance of alternative information
criteria, namely, AIC, BIC and HQIC, in finite sample via simulated data and
check the reliability of the asymptotic results developed in Section 3 and
Section 4. In the simulation study, we use both OLS and the indirect
inference method to estimate $\Greekmath 011A _{n}$ from sample paths that are
simulated from different DGPs. In total we design four experiments. In the
first experiment we simulate data from the UR model. In the second
experiment we simulate data from the LTUE model with $c=1$ (i.e. $\Greekmath 011A
_{n}=1+1/n)$. In the third experiment we simulate data from two ME models
with $c_{n}=n^{0.1}$, $n^{0.3}$, respectively. In the last experiment we
simulate data from the EX model with $\Greekmath 011A =1.01,1.05$, respectively. In all
experiments, we simulate 10,000 sample paths with initial value $X_{0}=0$
and four sample sizes are considered, $n=100,200,500,1000$. In each
experiment, we report the fraction of the number of times in which the
correct model is selected out of 10,000 replications.
Table (ref) reports the results when the true DGP is UR. Several
results can be found here. First, the probability for BIC\ and HQIC to
select the true model grows as $n$ grows. However, the probability for AIC
to select the true model does not seem to increase or decrease as $n$ grows.
This observation is consistent with the asymptotic results reported in
Theorem (ref). Second, the probability for BIC to select the true
model is larger than that in HQIC which is in turn larger than AIC in these
four sample sizes. So we can conclude that the probability grows as $p_{n}$
increases since $2<2\log \log n<\log n$ when $100\leq n\leq 1000$. Third,
the probability implied by AIC based on the indirect inference estimator is
larger than that based on OLS. This finding is consistent with Theorem (ref) and Remark (ref).
table[table omitted — 686 chars of source]
Table (ref) report the results when the true DGP is the LTUE model
with $c_{n}=1$. Also reported is the value of $p_{n}/\Greekmath 011A _{n}^{2n}$.
Several results can be found here. First, the probability for BIC\ and HQIC
to select the true model becomes smaller as $n$ grows. However, the
probability for AIC to select the true model does not seem to increase or
decrease as $n$ grows. This observation is consistent with the asymptotic
results in Theorem (ref). Second, the probability implied by AIC
based on the indirect inference estimator is smaller than that based on OLS.
This finding is consistent with in Theorem (ref) and Remark (ref). Finally, it seems that AIC performs better than BIC and HQIC in
all cases.
table[table omitted — 1,055 chars of source]
Table (ref) report the results when the true DGP is the ME model with
$c_{n}=n^{0.1},n^{0.3}$. Also reported is the value of $p_{n}/\Greekmath 011A _{n}^{2n}$
. Several results can be found here. First, the probability for all three
information criteria to select the true model grows as $n$ increases. This
observation is consistent with the asymptotic results reported in Theorem
(ref) and Remark (ref). Second, comparing the results for $
c_{n}=n^{0.1}$ and those for $c_{n}=n^{0.3}$, the probability for all three
information criteria to select the true model increases when $c_{n}$ is
bigger. Third, the probability based on the indirect inference estimator is
smaller than that based on OLS. Finally, it seems that AIC performs better
than BIC and HQIC in all cases.
table[table omitted — 2,045 chars of source]
Table (ref) report the results when the true DGP is the EX model with $
\Greekmath 011A =1.01,1.05$. Also reported is the value of $p_{n}/\Greekmath 011A ^{2n}$. Several
results can be found here. First, when $\Greekmath 011A =1.01$, which is larger than
the unity by 1%, the probability for information criteria to select the
correct model is small in all cases when the sample size is small. However,
it grows very quickly with the sample size. When $\Greekmath 011A =1.05$, the
probability for information criteria to select the correct model is almost 1
in all cases even when the sample size is small and increases with the
sample size. Finally, it seems that AIC performs better than BIC and HQIC in
all cases.
table[table omitted — 2,123 chars of source]
Conclusion
This paper studies the limit properties of information criteria for
distinguishing between unit root model and three types of explosive models.
Both the OLS estimator and the indirect inference estimator are employed to
estimate the AR coefficient in the candidate model. This paper contributes
to the literature in three aspects. First, our results extends results in
the literature to the explosive side of the unit root, and we find that
information criteria consistently choose the unit root model when the unit
root model is the true model. Second, we show that the limiting
probabilities for information criteria to select the explosive model depends
on both the distance of autoregressive coefficient from unity and the size
of penalty term in the information criteria. When the penalty term is not
too large and the root is not too close to unit root, all the information
criteria consistently select the true model. It is known that the indirect
inference method is effective in reducing the bias in OLS estimation in all
cases as well as reducing the variance in OLS estimation in the UR\ model
and in the LTU model. However, when information criteria are used in
connection with the indirect inference estimation, the limiting
probabilities for information criteria to select the correct model can go up
or down relative to that with the OLS estimation, depending on the true DGP.
When the true DGP is the UR\ model, the indirect inference estimation
increases the probability. When the true DGP is the LTUE\ model or the ME
model or the EX model, the indirect inference estimation decreases the
probability. This rather surprising result suggests that the superiority in
estimation does not necessarily translate to the superiority in model
selection.