EconBase
← Back to paper

Model Selection for Explosive Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

52,332 characters · 6 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Model Selection for Explosive Models

abstractThis paper examines the limit properties of information criteria (such as AIC, BIC, HQIC) for distinguishing between the unit root model and the various kinds of explosive models. The explosive models include the local-to-unit-root model, the mildly explosive model and the regular explosive model. Initial conditions with different order of magnitude are considered. Both the OLS estimator and the indirect inference estimator are studied. It is found that BIC\ and HQIC, but not AIC, consistently select the unit root model when data come from the unit root model. When data come from the local-to-unit-root model, both BIC and HQIC\ select the wrong model with probability approaching 1 while AIC has a positive probability of selecting the right model in the limit. When data come from the regular explosive model or from the mildly explosive model in the form of $ 1+n^{\Greekmath 010B }/n$ with $\Greekmath 010B \in (0,1)$, all three information criteria consistently select the true model. Indirect inference estimation can increase or decrease the probability for information criteria to select the right model asymptotically relative to OLS, depending on the information criteria and the true model. Simulation results confirm our asymptotic results in finite sample. Keywords: Model Selection; Information Criteria; Local-to-unit-root Model; Mildly Explosive Model; Unit Root Model; Indirect Inference.

Introduction

Information criteria have found a wide range of practical applications in empirical work. Examples include choosing explanatory variables in regression models and selecting lag lengths in time series models. Frequently used information criteria are AIC of Akaike (1969, 1973), BIC of Schwarz (1978), HQIC of Hannan and Quinn (1979). A major nice feature in these information criteria is that the penalty term is trivial to compute and hence the implementation of them is straightforward and can be made automatic.

With a growing interest in nonstationarity in time series analysis, researchers have examined the properties of information criteria in the context of nonstationary models with the unit root behavior. An important form of nonstationarity in time series involves explosive roots. Recent global financial crisis has motivated researchers to study explosive behavior in economic and financial time series; see, for example, Phillips and Yu (2011), Phillips, Wu and Yu (2011) and Phillips, Shi and Yu (2015a, b).

In this paper, we study the limit properties of information criteria for distinguishing between the unit root model and the explosive models. The information criteria considered in this paper have a general form and include AIC, BIC and HQIC as the special cases. The impact of the initial condition on the limit properties is examined by allowing for an initial condition of three different orders of magnitude. Moreover, both the OLS estimator and the indirect inference estimator are studied when investigating the limit properties of information criteria. The motivation for the use of indirect inference estimator comes from the existence of finite sample bias in the OLS estimator and the ability that the indirect inference method can reduce the bias.

It is found that information criteria consistently choose the unit root model against the explosive alternatives when data comes from the unit root model. Second, we prove that the probability for information criteria to correctly select the explosive model models against the unit root model depends crucially on both the degree of explosiveness and the size of the penalty term in information criteria. Finally and surprisingly, we show that indirect inference estimation can increase or decrease the probability for information criteria to select the right model asymptotically relative to OLS, depending on the information criteria and the true model.

The rest of this paper is organized as follows. Section 2 introduces the models and information criteria, and briefly reviews the literature. Section (ref) gives the limit properties of information criteria for distinguishing models with an explosive root from the unit root model when the OLS\ estimator is used. Section 4 gives the limit properties of information criteria when the indirect inference estimator\ is used. Section (ref) provides Monte Carlo evidence to support the theoretical results. Section (ref) concludes. All the detailed proofs are provided in the appendix. To compress notation, we denote $ \int\nolimits_{0}^{1}BdB$ and $\int\nolimits_{0}^{1}B^{2}$ in short for $ \int\nolimits_{0}^{1}B(r)dB(r)$ and $\int\nolimits_{0}^{1}B(r)^{2}dr$ respectively throughout the paper, and $\Rightarrow $ denotes weak convergence.

Models, Information Criteria and A Literature Review

The model considered in the present paper is of the form:

equation[equation omitted — 153 chars of source]

where $u_{t}\overset{iid}{\sim }(0,\Greekmath 011B ^{2})$ and the model is initialized at $t=0$ with some $X_{0}$. The autoregressive (AR) coefficient $ \Greekmath 011A _{n}$ is the crucial parameter that determines the dynamic behavior of $ X_{t}$. When $\Greekmath 011A _{n}=\Greekmath 011A $ and $\left\vert \Greekmath 011A \right\vert <1$, $X_{t}$ is stationary. When $\Greekmath 011A _{n}=1$, $X_{t}$ has a unit root (UR hereafter). When $\Greekmath 011A _{n}=1-c_{n}/n=1-c/n$ for $c>0$, $X_{t}$ is near-stationary and has a root that is local-to-unity (LTUS hereafter) (Phillips, 1987b; Chan and Wei, 1987). When $\Greekmath 011A _{n}=\Greekmath 011A $ and $\left\vert \Greekmath 011A \right\vert >1$, $X_{t}$ has an explosive root (EX hereafter). When $\Greekmath 011A _{n}=1+c_{n}/n=1+c/n $ for $c>0$, $X_{t}$ is near-explosive and also has a root that is local-to-unity (LTUE hereafter). When $\Greekmath 011A _{n}=1-c_{n}/n$ for $c_{n}\rightarrow \infty $ but $c_{n}/n\searrow 0$, the root represents moderate deviations from unity and $X_{t}$ is near-stationary (Phillips and Magdalinos, 2007). When $\Greekmath 011A _{n}=1+c_{n}/n$ for $c_{n}\rightarrow \infty $ but $c_{n}/n\searrow 0$, $X_{t}$ is mildly explosive (hereafter ME).

The asymptotic properties of the OLS\ estimator of the AR coefficient in the stationary AR(1) model is well known. The rate of convergence is $\sqrt{n}$ and the limiting distribution is Gaussian. Phillips (1987a) provided the limiting theory for the OLS\ estimator in the UR model and the rate of convergence is $n$. Phillips (1987b) and Chan and Wei (1987) established the asymptotic theory for the LTUS and LTUE models. The asymptotic theory is similar to that in the UR model and the rate of convergence is also $n$. In the cases of UR and LTU, $u_{t}$ can be weakly dependent stationary. Anderson (1959) studied the limiting distribution of the OLS\ estimator in the EX model under the condition that $u_{t}\overset{iid}{\sim }\mathcal{N} (0,\Greekmath 011B ^{2})$ and $X_{0}=0$. The limiting distribution is Cauchy and the rate of convergence is $\Greekmath 011A ^{n}$. However, no invariance principle applies. Assuming $X_{0}=o_{p}(\sqrt{n/c_{n}})$, Phillips and Magdalinos (2007) developed the asymptotic theory for the model with $\Greekmath 011A _{n}=1-c_{n}/n$ for $c_{n}\rightarrow \infty $ but $c_{n}/n\searrow 0$ and showed that the asymptotic distribution is invariant to the error distribution. The rate of convergence is $n/\sqrt{c_{n}}$. If $ c_{n}=n^{\Greekmath 010B }$ with $\Greekmath 010B \in (0,1)$, this rate of convergence bridges that of UR/LTU models and that of the stationary process. Phillips and Magdalinos (2007) also developed the asymptotic theory for the ME model. The rate of convergence is $n\Greekmath 011A _{n}^{n}/c_{n}$. The limiting distribution is Cauchy which is the same as in the EX model. Interestingly, in the ME case, the asymptotic theory is independent of the initial condition as long as $ X_{0}=o_{p}(\sqrt{n/c_{n}})$.

It is known that the OLS estimator of $\Greekmath 011A _{n}$ is biased downward when $ \Greekmath 011A _{n}=1$ or when $\Greekmath 011A _{n}$ is in the vicinity of unity. In this case, the indirect inference estimation is effective in reducing the bias. Phillips (2012) derives the asymptotic theory of the indirect inference estimator when the model is UR or LTU and $u_{t}\overset{iid}{\sim }\mathcal{ N}(0,\Greekmath 011B ^{2})$. The rate of convergence remains unchanged while the limiting distribution is different from that of the OLS estimator.

Information criteria for model selection have been proposed by Akaike (1969, 1973), Schwarz (1978), Hannan and Quinn (1979), among many others. The general form of these criteria is

equation*[equation* omitted — 82 chars of source]

where $k$ is the number of parameters to be estimated, $\widehat{\Greekmath 011B } _{k}^{2}$ is the estimated $\Greekmath 011B ^{2}$ when $k$ parameters are estimated. In general, $IC_{k}$ trades off the term that measures the goodness-of-fit (i.e. $\log \widehat{\Greekmath 011B }_{k}^{2}$) and the penalty term that measures the complexity of the model (i.e. $kp_{n}/n$). Coefficient $p_{n}=2,\log n,2\log \log n$ corresponds to AIC of Akaike (1973), BIC\ of Schwarz (1978) and HQIC of Hannan and Quinn (1979). Other forms of $p_{n}$ are possible.

In the time series literature, information criteria have been widely used to select the lag length both in the family of stationary models and in the family of nonstationary models; see for example, Ng and Perron (1995) and Ploberger and Phillips (2003). The information criteria can also be used to evaluate whether $\Greekmath 011A _{n}=1$ (i.e. $k=0$) or $\Greekmath 011A _{n}\neq 1$ (i.e. $k=1$ ) in Model ((ref)). For example, Phillips (2008) obtained limit properties of $IC_{k}$ for distinguishing between the unit root model and the stationary model. Phillips and Lee (2015) show that BIC can successfully distinguish the UR model from the ME model. This is a surprising result as it is well known that BIC cannot consistently distinguish between the UR\ model and the LTU model; see Ploberger and Phillips (2003).

In this paper we focus our attention to distinguishability between the unit root model and the three explosive models (i.e., LTUE, ME and EX) after the candidate models are estimated by OLS or by the indirect inference method. As a result, we make contributions in two strands of literature, explosive time series and indirect inference.

To visually understand the difference between the UR\ model, the LTU model and the ME model, we simulate a sample path of different length ($ n=100,200,500,1000$) with $y_{0}=0$, based on the same realizations of the error process, iid $\mathcal{N}(0,1)$, from the following four models, $\Greekmath 011A _{n}=1$ (UR), $\Greekmath 011A _{n}=1+1/n$ (LTUE), $\Greekmath 011A _{n}=1+n^{0.1}/n$ (ME1), and $ \Greekmath 011A _{n}=1+n^{0.5}/n$ (ME2). Figures 1-3 give the time series plot of UR against LTU, UR against ME1, UR against ME2, respectively. It can be seen from Figure 1 that it is very difficult to distinguish between the UR process and the LTU process, even when the sample size is as large as 1,000. When the sample size increases, the gap between the UR\ process and the two ME processes becomes larger and larger, as apparent in Figure 2 and more so in Figure 3.

figure[figure omitted — 180 chars of source]
figure[figure omitted — 193 chars of source]
figure[figure omitted — 191 chars of source]

Limit Properties Based on the OLS Estimator

When the data generating process (DGP) is the UR model, since $\Greekmath 011A _{n}=1$, we set the parameter count to $k=0$. For the LTU model, the ME model and the EX model, we need to estimate the AR coefficient and hence set the parameter count to $k=1$. Throughout the paper we denote $\widehat{\Greekmath 011A }$ the OLS estimator of $\Greekmath 011A $. $\widehat{k}_{IC}=0$ or $1$ means the information criterion of the UR\ model is smaller or larger than that of the competing model when $\Greekmath 011A $ is estimated by OLS. We aim to find the limit of the following probabilities:

align[align omitted — 333 chars of source]

As shown in Phillips and Magdalinos (2009), the unit root asymptotic distribution is sensitive to initial conditions in the distant past. To understand how the initial condition affects the property of $\widehat{k} _{IC}$, we follow Phillips and Magdalinos (2009) by assuming alternative initial conditions.

assumption[IN] The initial condition has the form \begin{equation} X_{0}(n)=\sum_{j=0}^{\Greekmath 0114 _{n}}u_{-j}, \end{equation} where $\Greekmath 0114 _{n}$ is a sequence of integers satisfying $\Greekmath 0114 _{n}\rightarrow \infty $ and \begin{equation} \frac{\Greekmath 0114 _{n}}{n}\rightarrow \Greekmath 011C \in \left[ 0,\infty \right] \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, as }n\rightarrow \infty . \end{equation} The following cases are distinguished: \begin{enumerate} • If $\Greekmath 011C = 0$, $X_0(n) $ is said to be a recent past initialization. • If $\Greekmath 011C \in \left(0, \infty\right)$, $X_0(n) $ is said to be a distant past initialization. • If $\Greekmath 011C =\infty $, $X_{0}(n)$ is said to be an infinite past initialization. \end{enumerate}
theoremUnder Assumption (ref) (i) or (ii) or (iii), we have \begin{enumerate} • when $p_{n}\rightarrow \infty $ and $p_{n}/n\rightarrow 0$ as $ n\rightarrow \infty $, \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=0\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ IC_{0}-IC_{1}\leq 0\right\} =1, \\ \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=0\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ IC_{0}-IC_{1}>0\right\} =0. \end{align*} • when $p_{n}=2$, the asymptotic distribution under the AIC criterion is \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{AIC}=0|k=0\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ AIC_{0}-AIC_{1}\leq 0\right\} =P\left( \Greekmath 0118 ^{2}<2\right) , \\ \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{AIC}=1|k=0\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ AIC_{0}-AIC_{1}>0\right\} =1-P\left( \Greekmath 0118 ^{2}<2\right) . \end{align*} where \begin{equation*} \Greekmath 0118 ^{2}= \begin{cases} \dfrac{\left( \int_{0}^{1}BdB\right) ^{2}}{\int_{0}^{1}B^{2}}, & \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if } \Greekmath 011C =0 \\ \dfrac{\left( \int_{0}^{1}B_{\Greekmath 011C }dB\right) ^{2}}{\int_{0}^{1}B_{\Greekmath 011C }^{2}} , & \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }\Greekmath 011C \in (0,\infty ) \\ B(1)^{2}, & \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }\Greekmath 011C =\infty \end{cases} , \end{equation*} with $B(s)$ being a Brownian motion, and \begin{equation*} B_{\Greekmath 011C }(s)=B(s)+\sqrt{\Greekmath 011C }B_{0}(1), \end{equation*} with $B_{0}(s)$ being an independent Brownian motion. \end{enumerate}
remarkTheorem (ref) is the same as Theorem 1 in Phillips (2008) for distinguishing between the UR model and the stationary model. The condition that $p_{n}\rightarrow \infty $ and $p_{n}/n\rightarrow 0$ covers BIC and HQIC and hence, both BIC and HQIC can consistently select the UR model. The AIC criterion is inconsistent and its asymptotic distribution depends on $\Greekmath 0118 ^{2}$, the squared unit root $t$-statistic for the OLS estimator.
remarkThe validity of Theorem (ref) does not require the iid assumption for the error term $u_{t}$. If we follow Phillips (2008) by denoting $F(L)=\sum_{j=0}^{\infty }F_{j}L^{j}$, with $F_{0}=1$ and $F(1)\neq 0$, and letting $u_{s}$ have Wold representation \begin{equation} u_{s}=F(L)\Greekmath 0122 _{s}=\sum_{j=0}^{\infty }F_{j}\Greekmath 0122 _{s-j}\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ , with }\sum_{j=0}^{\infty }j^{1/2}\left\vert F_{j}\right\vert <\infty , \end{equation} where $\Greekmath 0122 _{t}\overset{iid}{\sim }\left( 0,\Greekmath 011B _{\Greekmath 0122 }^{2}\right) $, the results in Theorem (ref) continue to hold. However, both $B_{0}$ and $\Greekmath 0118 ^{2}$ need to be modified to accommodate the dependence in $u_{t}$ as in Phillips (2008).
theoremLet Assumption (ref) (i) or (ii) holds. Assume the true DGP is the LTUE model. \begin{enumerate} • When $p_{n}\rightarrow \infty $ and $p_{n}/n\rightarrow 0$ as $ n\rightarrow \infty $, \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left( IC_{1}-IC_{0}\right) >0\right\} =1, \\ \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left( IC_{1}-IC_{0}\right) \leq 0\right\} =0. \end{align*} • When $p_{n}=2$, the asymptotic distribution of the AIC criterion is \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{AIC}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ {n}\left( AIC_{1}-AIC_{0}\right) >0\right\} =1-P\left( \Greekmath 0110 ^{2}>2\right) , \\ \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{AIC}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ {n}\left( AIC_{1}-AIC_{0}\right) \leq 0\right\} =P\left( \Greekmath 0110 ^{2}>2\right) , \end{align*} where \begin{equation*} \Greekmath 0110 ^{2}=\frac{\left( \int_{0}^{1}J_{c}dB\right) ^{2}}{ \int_{0}^{1}J_{c}^{2}}+2{c}\int_{0}^{1}J_{c}dB+c^{2}\int_{0}^{1}J_{c}^{2}, \end{equation*} with \begin{equation*} J_{c}(r)=\int_{0}^{r}\exp \left\{ c(r-s)\right\} dB(s). \end{equation*} \end{enumerate}
remarkTheorem (ref) shows that all the information criteria are inconsistent in distinguishing between the LTUE model and the UR models when data comes from the LTUE model. AIC selects the wrong model with probability going to $1-P\left( \Greekmath 0110 ^{2}>2\right) $, which depends on the localization constant $c$. This problem worsens for BIC and HQIC as the probability of selecting the wrong model goes to one. Note that BIC is well known to be blind to local alternatives; see, for example, Ploberger and Phillips (2003).
theoremLet Assumption (ref) (i) or (ii) holds. Assume the true DGP is the ME model. \begin{enumerate} • When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A _{n}^{2n}}=0,$ \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left( IC_{1}-IC_{0}\right) >0\right\} =0, \\ \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left( IC_{1}-IC_{0}\right) \leq 0\right\} =1. \end{align*} • When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A _{n}^{2n}}=\Greekmath 0119 \in (0,+\infty ),$ \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left( IC_{1}-IC_{0}\right) >0\right\} =P\left( \Greekmath 011F ^{2}(1)<4\Greekmath 0119 \right) , \\ \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left( IC_{1}-IC_{0}\right) \leq 0\right\} =1-P\left( \Greekmath 011F ^{2}(1)<4\Greekmath 0119 \right) . \end{align*} • When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A _{n}^{2n}}\rightarrow +\infty ,$ \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left( IC_{1}-IC_{0}\right) >0\right\} =1, \\ \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left( IC_{1}-IC_{0}\right) \leq 0\right\} =0. \end{align*} \end{enumerate}
remarkTheorem (ref) shows that the limit probability of selecting the correct model by information criteria under the ME model depends critically on two parameters, $c_{n}$, $p_{n}$. As expected, the larger $c_{n}$, the further the model away from the UR\ model and the higher probability for the information criteria to select the correct model. Interestingly, the smaller $p_{n}$, the higher probability for the information criteria to select the correct model. From Phillips and Magdalinos (2009), we know $\Greekmath 011A _{n}^{-n}=o(c_{n}^{-1})$ and hence $\Greekmath 011A _{n}^{n}/c_{n}\rightarrow +\infty $. In the special case where $ c_{n}=n^{\Greekmath 010B }$, for $\Greekmath 010B \in (0,1)$, $\lim\limits_{n\rightarrow \infty }p_{n}/\Greekmath 011A _{n}^{2n}=0$ no matter whether $p_{n}=2$ or $\log n$ or $ 2\log \log n$. In this case, all the well-known information criteria can consistently select the true model.
theoremLet Assumption (ref) (i) holds. Assume the true DGP is the EX model. \begin{enumerate} • When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A ^{2n}} =0,$ \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A ^{2n}}\left( IC_{1}-IC_{0}\right) >0\right\} =0, \\ \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A ^{2n}}\left( IC_{1}-IC_{0}\right) \leq 0\right\} =1. \end{align*} • When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A ^{2n}} =\Greekmath 0119 \in (0,+\infty ),$ \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A ^{2n}}\left( IC_{1}-IC_{0}\right) >0\right\} =P\left( \Greekmath 011F ^{2}(1)<(1+\Greekmath 011A )^{2}\Greekmath 0119 \right) , \\ \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A ^{2n}}\left( IC_{1}-IC_{0}\right) \leq 0\right\} =1-P\left( \Greekmath 011F ^{2}(1)<(1+\Greekmath 011A )^{2}\Greekmath 0119 \right) . \end{align*} • When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A ^{2n}} \rightarrow +\infty ,$ \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left( IC_{1}-IC_{0}\right) >0\right\} =1, \\ \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left( IC_{1}-IC_{0}\right) \leq 0\right\} =0. \end{align*} \end{enumerate}
remarkTheorem (ref) shows that the limit probability of selecting the correct model by information criteria under the EX model depends also critically on two parameters, $\Greekmath 011A $, $p_{n}$. As expected, the larger $\Greekmath 011A $, the higher probability for the information criteria to select the correct model. Interestingly, the smaller $p_{n}$, the higher probability for the information criteria to select the correct model. If $ p_{n}=2$ or $\log n$ or $2\log \log n$, $\lim\limits_{n\rightarrow \infty }p_{n}/\Greekmath 011A ^{2n}=0$ and hence case (1) applies, suggesting that all the well-known information criteria can consistently select the true model.

Results in Theorem (ref) can be extended to cover the LTUE model and the ME model with weakly dependent errors. The following proposition establishes the results for the ME model.

propositionLet\ Assumption (ref) (i) or (ii) and the assumption specified in Equation ((ref)) hold. Assume the true DGP is the ME model. \begin{enumerate} • When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A _{n}^{2n}}=0,$ \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left( IC_{1}-IC_{0}\right) >0\right\} =0, \\ \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left( IC_{1}-IC_{0}\right) \leq 0\right\} =1. \end{align*} • When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A _{n}^{2n}}=\Greekmath 0119 \in (0,+\infty ),$ \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left( IC_{1}-IC_{0}\right) >0\right\} =P\left( \Greekmath 011F ^{2}(1)<\frac{4\Greekmath 0119 }{\Greekmath 0121 ^{2}}\right) , \\ \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left( IC_{1}-IC_{0}\right) \leq 0\right\} =1-P\left( \Greekmath 011F ^{2}(1)<\frac{4\Greekmath 0119 }{ \Greekmath 0121 ^{2}}\right) . \end{align*} where $\Greekmath 0121 ^{2}=\left( \sum\nolimits_{j=0}^{\infty }F_{j}\right) ^{2}.$ • When $\dfrac{p_{n}}{\Greekmath 011A _{n}^{2n}}\rightarrow +\infty ,$ \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left( IC_{1}-IC_{0}\right) >0\right\} =1, \\ \lim\limits_{n\rightarrow \infty }P\left\{ \widehat{k}_{IC}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left( IC_{1}-IC_{0}\right) \leq 0\right\} =0. \end{align*} \end{enumerate}

Limit Properties Based on the Indirect Inference Estimator

The OLS estimator of $\Greekmath 011A _{n}$ in Model ((ref)) is known to be biased and the bias is acute when $\Greekmath 011A _{n}$ is close to unity. To reduce the bias, the indirect inference method of Smith (1993) and Gour\'{e}rioux et al (1993) can be used if Model ((ref)) is fully specified. Phillips (2012) derives the asymptotic theory of the indirect inference estimator when the model is UR or LTU and $u_{t}\overset{iid}{\sim }\mathcal{N}(0,\Greekmath 011B ^{2})$ . Throughout the paper we denote $\breve{\Greekmath 011A }$ the indirect inference estimator of $\Greekmath 011A $. Let $h(c)=c+g(c)$ and $g(c)=g^{-}(c)1_{\{c\leq 0\}}+g^{+}(c)1_{\{c>0\}}$ with

alignat*{2} g^{-}(c)& = & & -\dfrac{3}{4}\int_{0}^{\infty }e^{-\frac{v}{4} }k^{-}(v;c)^{1/2}dv+\dfrac{1}{4}\int_{0}^{\infty }e^{-\frac{v}{4} }k^{-}(v;c)^{3/2}dv \\ & & & -\dfrac{e^{2c}}{8}\int_{0}^{\infty }e^{-\frac{5v}{4} }k^{-}(v;c)^{3/2}vdv, \\ g^{+}(c)& = & & \dfrac{3}{4}\int_{0}^{\infty }e^{\frac{w}{4} }k^{+}(w;c)^{1/2}dw-\dfrac{1}{4}\int_{0}^{\infty }e^{\frac{w}{4} }k^{+}(w;c)^{3/2}dw \\ & & & -\dfrac{e^{2c}}{8}\int_{0}^{\infty }e^{\frac{5w}{4} }k^{+}(w;c)^{3/2}wdw, \\ k^{-}(v;c)& = & & \dfrac{2v-4c}{v+e^{2c}ve^{-v}-4c}, \\ k^{+}(w;c)& = & & \dfrac{2w+4c}{w+e^{2c}we^{w}+4c}.

Phillips (2012) shows that under the UR model,

equation*[equation* omitted — 218 chars of source]

and under the LTUE model,

equation*[equation* omitted — 250 chars of source]

Let $\breve{k}_{IC}=0$ or $1$ mean the information criterion of the UR\ model is smaller or larger than that of the competing model when the model is estimated by the indirect inference method. We aim to find is the limit of the following probabilities:

align[align omitted — 325 chars of source]
theoremUnder Assumption (ref)(i) or (ii) or (iii), we have \begin{enumerate} • when $p_{n}\rightarrow \infty $ and $p_{n}/n\rightarrow 0$ as $ n\rightarrow \infty $, \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=0|k=0\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ IC_{0}-IC_{1}\leq 0\right\} =1, \\ \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=1|k=0\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ IC_{0}-IC_{1}>0\right\} =0; \end{align*} • when $p_{n}=2$, the asymptotic distribution under the AIC criterion is \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{AIC}=0|k=0\right\} & =P\left( \Greekmath 0126 ^{2}<2\right) , \\ \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{AIC}=1|k=0\right\} & =1-P\left( \Greekmath 0126 ^{2}<2\right) , \end{align*} where \begin{equation*} \Greekmath 0126 ^{2}= \begin{cases} \int_{0}^{1}B^{2}\cdot h^{-1}\left( \left( \dfrac{\int_{0}^{1}BdB}{ \int_{0}^{1}B^{2}}\right) ^{2}\right) -2\int_{0}^{1}BdB\cdot h^{-1}\left( \dfrac{\int_{0}^{1}BdB}{\int_{0}^{1}B^{2}}\right) , & \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }\Greekmath 011C =0 \\ \int_{0}^{1}B_{\Greekmath 011C }^{2}\cdot h^{-1}\left( \left( \dfrac{ \int_{0}^{1}B_{\Greekmath 011C }dB}{\int_{0}^{1}B_{\Greekmath 011C }^{2}}\right) ^{2}\right) -2\int_{0}^{1}B_{\Greekmath 011C }dB\cdot h^{-1}\left( \dfrac{\int_{0}^{1}B_{\Greekmath 011C }dB}{ \int_{0}^{1}B_{\Greekmath 011C }^{2}}\right) , & \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }\Greekmath 011C \in (0,\infty ) \\ h^{-1}\left( \mathcal{C}\right) ^{2}B_{0}^{2}(1)-2h^{-1}\left( \mathcal{C} \right) B(1)B_{0}(1), & \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{if }\Greekmath 011C =\infty \end{cases} , \end{equation*} with $\mathcal{C}$ being a standard Cauchy variate. \end{enumerate}
remarkAccording to Theorem (ref), as long as $p_{n}\rightarrow \infty $ and $p_{n}/n\rightarrow 0$, information criteria based on the indirect inference estimator is consistent in selecting the UR\ model. Hence, BIC and HQIC based on the indirect inference estimator can consistently select the UR model. Like the AIC criterion that is based on the OLS estimator, the AIC criterion based on the indirect inference estimator continues to be inconsistent. However, its asymptotic distribution depends on $\Greekmath 0126 ^{2}$, the squared unit root $t$-statistic for the indirect inference estimator.
remarkAs shown in Phillips (2012), the squared unit root $t$ -statistic for the indirect inference estimator has a smaller variance than that of the squared unit root $t$-statistic for the OLS estimator. Consequently, $P\left( \Greekmath 0126 ^{2}<2\right) >P\left( \Greekmath 0118 ^{2}<2\right) $, suggesting that AIC based on the indirect inference estimator can select the true model (i.e. the UR model) with a larger probability than that based on the OLS estimator.
theoremLet Assumption (ref) (i) or (ii) holds. Assume the true DGP is the LTUE model. \begin{enumerate} • When $p_{n}\rightarrow \infty $ and $p_{n}/n\rightarrow 0$ as $ n\rightarrow \infty $, \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left( IC_{1}-IC_{0}\right) >0\right\} =1, \\ \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left( IC_{1}-IC_{0}\right) \leq 0\right\} =0. \end{align*} • When $p_{n}=2$, the asymptotic distribution under the AIC criterion is \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{AIC} }=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ {n}\left( AIC_{1}-AIC_{0}\right) >0\right\} =1-P\left( \Greekmath 0123 ^{2}>2\right) , \\ \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{AIC} }=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ {n}\left( AIC_{1}-AIC_{0}\right) \leq 0\right\} =P\left( \Greekmath 0123 ^{2}>2\right) , \end{align*} where \begin{equation*} \Greekmath 0123 ^{2}\equiv 2h^{-1}\left( \dfrac{\int_{0}^{1}J_{c}dB}{ \int_{0}^{1}J_{c}^{2}}+c\right) \left( \int_{0}^{1}J_{c}dB+c\int_{0}^{1}J_{c}^{2}\right) -h^{-1}\left( \dfrac{ \int_{0}^{1}J_{c}dB}{\int_{0}^{1}J_{c}^{2}}+c\right) ^{2}\int_{0}^{1}J_{c}^{2}. \end{equation*} \end{enumerate}
remarkTheorem (ref) shows that all the information criteria continue to be inconsistent in distinguishing between the LTUE model and the UR models when data come from the LTUE model even when the indirect inference estimation is employed. AIC selects the wrong model with probability going to $1-P\left( \Greekmath 0123 ^{2}>2\right) $. Since the variance of $\Greekmath 0110 ^{2}$ is bigger than that of $\Greekmath 0123 ^{2}$, the tail probability of $\Greekmath 0110 ^{2}$ is larger than that of $\Greekmath 0123 ^{2}$, suggesting that AIC based on OLS selects the true model (i.e. LTUE model) with a greater\ probability than AIC\ based on the indirect inference estimator. This is a rather surprising result and suggests that the superiority in estimation does not necessarily translate to the superiority in model selection.
theoremLet Assumption (ref) (i) or (ii) holds. Assume the true DGP is the ME model. \begin{enumerate} • When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A _{n}^{2n}}=0,$ \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left( IC_{1}-IC_{0}\right) >0\right\} =0, \\ \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left( IC_{1}-IC_{0}\right) \leq 0\right\} =1. \end{align*} • When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A _{n}^{2n}}=\Greekmath 0119 \in (0,+\infty ),$ \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left( IC_{1}-IC_{0}\right) >0\right\} =P\left( \Greekmath 011F ^{2}(1)<4\Greekmath 0119 \right) , \\ \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A _{n}^{2n}}\left( IC_{1}-IC_{0}\right) \leq 0\right\} =1-P\left( \Greekmath 011F ^{2}(1)<4\Greekmath 0119 \right) . \end{align*} • When $\dfrac{p_{n}}{\Greekmath 011A _{n}^{2n}}\rightarrow +\infty ,$ \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left( IC_{1}-IC_{0}\right) >0\right\} =1, \\ \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left( IC_{1}-IC_{0}\right) \leq 0\right\} =0. \end{align*} \end{enumerate}
remarkThe results in Theorem (ref) are the same as those in Theorem (ref), suggesting all the well-known information criteria can consistently select the true model (i.e. ME model) when $c_{n}=n^{\Greekmath 010B }$, for $\Greekmath 010B \in (0,1)$.
theoremLet Assumption (ref) (i) holds. Assume the true DGP is the EX model. \begin{enumerate} • When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A ^{2n}} =0,$ \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A ^{2n}}\left( IC_{1}-IC_{0}\right) >0\right\} =0, \\ \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A ^{2n}}\left( IC_{1}-IC_{0}\right) \leq 0\right\} =1. \end{align*} • When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A ^{2n}} =\Greekmath 0119 \in (0,+\infty ),$ \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A ^{2n}}\left( IC_{1}-IC_{0}\right) >0\right\} =P\left( \Greekmath 011F ^{2}(1)<(1+\Greekmath 011A )^{2}\Greekmath 0119 \right) , \\ \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{\Greekmath 011A ^{2n}}\left( IC_{1}-IC_{0}\right) \leq 0\right\} =1-P\left( \Greekmath 011F ^{2}(1)<(1+\Greekmath 011A )^{2}\Greekmath 0119 \right) . \end{align*} • When $\lim\limits_{n\rightarrow \infty }\dfrac{p_{n}}{\Greekmath 011A ^{2n}} \rightarrow +\infty ,$ \begin{align*} \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=0|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left( IC_{1}-IC_{0}\right) >0\right\} =1, \\ \lim\limits_{n\rightarrow \infty }P\left\{ \breve{k}_{IC}=1|k=1\right\} & =\lim\limits_{n\rightarrow \infty }P\left\{ \dfrac{n}{p_{n}}\left( IC_{1}-IC_{0}\right) \leq 0\right\} =0. \end{align*} \end{enumerate}
remarkThe results in Theorem (ref) are the same as those in Theorem (ref), suggesting that all the well-known information criteria can consistently select the true model (i.e. EX model).

Monte Carlo Study

In this section, we examine the performance of alternative information criteria, namely, AIC, BIC and HQIC, in finite sample via simulated data and check the reliability of the asymptotic results developed in Section 3 and Section 4. In the simulation study, we use both OLS and the indirect inference method to estimate $\Greekmath 011A _{n}$ from sample paths that are simulated from different DGPs. In total we design four experiments. In the first experiment we simulate data from the UR model. In the second experiment we simulate data from the LTUE model with $c=1$ (i.e. $\Greekmath 011A _{n}=1+1/n)$. In the third experiment we simulate data from two ME models with $c_{n}=n^{0.1}$, $n^{0.3}$, respectively. In the last experiment we simulate data from the EX model with $\Greekmath 011A =1.01,1.05$, respectively. In all experiments, we simulate 10,000 sample paths with initial value $X_{0}=0$ and four sample sizes are considered, $n=100,200,500,1000$. In each experiment, we report the fraction of the number of times in which the correct model is selected out of 10,000 replications.

Table (ref) reports the results when the true DGP is UR. Several results can be found here. First, the probability for BIC\ and HQIC to select the true model grows as $n$ grows. However, the probability for AIC to select the true model does not seem to increase or decrease as $n$ grows. This observation is consistent with the asymptotic results reported in Theorem (ref). Second, the probability for BIC to select the true model is larger than that in HQIC which is in turn larger than AIC in these four sample sizes. So we can conclude that the probability grows as $p_{n}$ increases since $2<2\log \log n<\log n$ when $100\leq n\leq 1000$. Third, the probability implied by AIC based on the indirect inference estimator is larger than that based on OLS. This finding is consistent with Theorem (ref) and Remark (ref).

table[table omitted — 686 chars of source]

Table (ref) report the results when the true DGP is the LTUE model with $c_{n}=1$. Also reported is the value of $p_{n}/\Greekmath 011A _{n}^{2n}$. Several results can be found here. First, the probability for BIC\ and HQIC to select the true model becomes smaller as $n$ grows. However, the probability for AIC to select the true model does not seem to increase or decrease as $n$ grows. This observation is consistent with the asymptotic results in Theorem (ref). Second, the probability implied by AIC based on the indirect inference estimator is smaller than that based on OLS. This finding is consistent with in Theorem (ref) and Remark (ref). Finally, it seems that AIC performs better than BIC and HQIC in all cases.

table[table omitted — 1,055 chars of source]

Table (ref) report the results when the true DGP is the ME model with $c_{n}=n^{0.1},n^{0.3}$. Also reported is the value of $p_{n}/\Greekmath 011A _{n}^{2n}$ . Several results can be found here. First, the probability for all three information criteria to select the true model grows as $n$ increases. This observation is consistent with the asymptotic results reported in Theorem (ref) and Remark (ref). Second, comparing the results for $ c_{n}=n^{0.1}$ and those for $c_{n}=n^{0.3}$, the probability for all three information criteria to select the true model increases when $c_{n}$ is bigger. Third, the probability based on the indirect inference estimator is smaller than that based on OLS. Finally, it seems that AIC performs better than BIC and HQIC in all cases.

table[table omitted — 2,045 chars of source]

Table (ref) report the results when the true DGP is the EX model with $ \Greekmath 011A =1.01,1.05$. Also reported is the value of $p_{n}/\Greekmath 011A ^{2n}$. Several results can be found here. First, when $\Greekmath 011A =1.01$, which is larger than the unity by 1%, the probability for information criteria to select the correct model is small in all cases when the sample size is small. However, it grows very quickly with the sample size. When $\Greekmath 011A =1.05$, the probability for information criteria to select the correct model is almost 1 in all cases even when the sample size is small and increases with the sample size. Finally, it seems that AIC performs better than BIC and HQIC in all cases.

table[table omitted — 2,123 chars of source]

Conclusion

This paper studies the limit properties of information criteria for distinguishing between unit root model and three types of explosive models. Both the OLS estimator and the indirect inference estimator are employed to estimate the AR coefficient in the candidate model. This paper contributes to the literature in three aspects. First, our results extends results in the literature to the explosive side of the unit root, and we find that information criteria consistently choose the unit root model when the unit root model is the true model. Second, we show that the limiting probabilities for information criteria to select the explosive model depends on both the distance of autoregressive coefficient from unity and the size of penalty term in the information criteria. When the penalty term is not too large and the root is not too close to unit root, all the information criteria consistently select the true model. It is known that the indirect inference method is effective in reducing the bias in OLS estimation in all cases as well as reducing the variance in OLS estimation in the UR\ model and in the LTU model. However, when information criteria are used in connection with the indirect inference estimation, the limiting probabilities for information criteria to select the correct model can go up or down relative to that with the OLS estimation, depending on the true DGP. When the true DGP is the UR\ model, the indirect inference estimation increases the probability. When the true DGP is the LTUE\ model or the ME model or the EX model, the indirect inference estimation decreases the probability. This rather surprising result suggests that the superiority in estimation does not necessarily translate to the superiority in model selection.