EconBase
← Back to paper

A Characterization of the $M$-tests Under Nearly Integrated Nearly White Noise

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

68,526 characters

A Characterization of the $M$-tests Under Nearly Integrated Nearly White Noise


\singlespacing

\title{A Characterization of the $M$-tests Under Nearly Integrated Nearly White Noise}
\author{Aidan Wardak\footnote{Email: [email removed].}  and Sayar Karmakar\footnote{ Corresponding author. Email:[email removed]. The authors declare no competing interests.}}
\date{September 24, 2026}

\maketitle

\begin{abstract}
\normalsize\singlespacing
\noindent
We derive the limiting distributions of the $M$-test family of unit root statistics in the nearly integrated nearly white noise (NINW) framework introduced by \citet{NP1994} in the case of an unknown linear time trend. In the case of known long run variance (LRV), the limiting distributions of the $M^{GLS}$ tests are contaminated by additional noise terms as a result of quasi differencing whereas these terms are less present in the $M^{OLS}$ limiting distributions, both of which display conservative properties under conventional critical values. Furthermore, we prove the Gaussian power envelope in the NINW model is asymptotically equivalent to the standard envelope of \citet{ERS1996}, and that the oracle $M$-tests have inefficient power relative to this benchmark. We then derive the limiting distributions of the feasible statistics and show that the autoregressive estimate of the LRV commonly used overestimates the LRV, creating altered limiting distributions. Finally, finite sample simulations illustrate that, of the procedures considered, no uniformly satisfactory solution exists for handling a series with a large negative moving average coefficient.
\end{abstract}

\noindent\textbf{Keywords: } $M$-tests; GLS detrending; long-run variance

\section{Introduction}
Testing for the presence of a unit root is common practice, but the reliability of many unit root tests drastically deteriorates under data-generating processes with a moving average coefficient near $-1$. The original Dickey-Fuller test \citep{DickeyFuller1979} was extended to hopefully accommodate an error process such as this by \citet{SaidDickey1984}. \citet{PP1988} developed a nonparametric correction to these tests based on estimation of the LRV (the spectral density at frequency zero). Yet, these tests display severe size distortions when the moving average coefficient is near $-1$, as documented in \citet{Schwert1989}, \citet{Agiakloglou1992EMPIRICALEO}, \citet{Leybourne1999}, and more. The $M$-test family of statistics originating with \citet{Stock1990} and further developed by \citet{NgPerron1996}, were specifically designed to improve robustness in precisely this type of data generating process (DGP). \citet{NP1998} subsequently analyzed the autoregressive LRV estimator used to implement these tests and found substantial size improvement over kernel-based methods like \citet{NeweyWest1987}.

\citet{NgPerron2001} combined the typical $M$-test construction with the GLS detrending developed by \citet{DuFourKing1991} and \citet{ERS1996} to further improve the power of the $M$-tests. Under the standard local-to-unity framework, the resulting $M^{GLS}$ tests display local asymptotic power functions close to that of the Gaussian power envelope while retaining more favorable size properties than conventional tests under a large negative moving average coefficient. Ng and Perron also emphasize that implementation of unit root tests depends greatly on the lag truncation parameter $k$ \citep{Agiakloglou1996,NgPerron1995,LOPEZ1997}. They find when the moving average coefficient is large and negative, a relatively large autoregression is necessary to control size, while conventional criteria like the AIC and BIC tend to select too few lags. The modified information criterion of \citet{NgPerron2001} was designed to address this problem. Furthermore, \citet{PQ2007} suggest using OLS detrended data to construct the modified information criterion to solve the problem of power reversal; that is, the power of GLS-based tests decrease when the autoregressive coefficient strays further from unity as documented in \citet{seo2005improving}.

The present paper revisits these results under the nearly integrated nearly white noise local asymptotic framework of \citet{NP1994} and \citet{NgPerron1996}. In this framework, the autoregressive coefficient approaches unity and the moving average coefficient approaches $-1$. As a result, the LRV collapses to 0. This degeneracy is the defining feature of this framework; quantities that are otherwise asymptotically negligible under the standard local-to-unity framework can become first order when measured relative to the vanishing LRV. \citet{NgPerron1996}, \citet{NP1998}, and \citet{NgPerron2001} analyze aspects of the $M$-tests and lag selection under this framework, but do not fully characterize the behavior of the $M$-tests with an unknown linear time trend, the associated Gaussian power bound, and the exact relative behavior of the feasible LRV estimator jointly.

The first contribution of this paper is to derive the known LRV (oracle) distributions of the OLS and GLS based $M$-tests within the NINW framework. For OLS detrending, the resulting limiting distributions are a direct extension of the no-deterministics limiting distributions already derived by \citet{NgPerron1996}. However, the $M^{GLS}$ tests have limiting distributions that gain extra nuisance terms that remain asymptotically non-negligible as a result of quasi-differencing. The GLS limiting distributions therefore cannot be obtained by merely replacing the undetrended Ornstein-Uhlenbeck process with its standard GLS-detrended counterpart. The resulting OLS and GLS statistics display $\delta$ dependent support restrictions; for sufficiently small $\delta$, standard critical values lie outside these support bounds and the asymptotic rejection probability is hence exactly zero. Thus, the NINW problem can produce severe under-rejection even when the LRV is known.

The second contribution of this paper is to characterize the Gaussian power envelope under the NINW DGP. \citet{ERS1996} derive the standard Gaussian envelope under conditions that include a positive spectral density at frequency zero which fails under our framework. We find that the power envelope remains unchanged from the standard case and is hence not dependent on $\delta$. For every fixed $\delta >0$, exact whitening removes the nearly noninvertible MA component and the nuisance projection terms generated by the transformed intercept and trend either cancel between the null and point alternative or are offset by their normalization. Hence, the poor asymptotic power of the $M$-tests relative to this benchmark is a direct result of the inefficiency of the statistics, not an intrinsically harder testing problem.

The third contribution concerns the feasible implementation of the tests in which the LRV must be estimated. \citet{NP1998} prove that $s^2_{AR} \stackrel{\mathbb{P}}{\to} 0$ under NINW. \citet{NgPerron2001} subsequently show that, under NINW, the conditions $k^2 s^{2}_{AR} = O_\mathbb{P}(1)$ and $Ts^2_{AR} = O_\mathbb{P}(1)$ can both hold only at the lag rate $k/\sqrt{T} \to \kappa \in (0,\infty)$. Neither result, however, determines whether $s^2_{AR}$ is ratio consistent with the LRV, which is necessary for oracle and feasible limiting distribution equivalence as the LRV collapses to zero. We show that at this lag rate, the ratio of $s^2_{AR}$ and the true LRV approaches $\coth^2(\kappa\delta/2)$ asymptotically, and hence, $s^2_{AR}$ is not ratio consistent. This sharpens the earlier NINW consistency results by identifying the multiplicative distortion that remains at the rate required to keep the feasible tests bounded.

The asymptotic analysis is complemented by finite sample Monte Carlo experiments designed to examine the interaction between detrending, lag selection, and the admissible upper bound on $k$. Increasing the conventional lag cap to allow orders proportional to $\sqrt{T}$ can materially reduce the most extreme NINW oversizing, but this improvement comes with substantial losses in size-adjusted power. Furthermore, GLS based construction of the modified information criterion results in favorable size properties relative to an OLS based construction, but then the tests are subject to the power reversal problem. Finally, the $M^{GLS}$  and $M^{OLS}$ tests display similar size-adjusted power. These results illustrate a persistent size-power tradeoff rather than any uniformly satisfactory implementation.

The remainder of the paper is organized as follows. Section \ref{sec:model} introduces the NINW model, GLS detrending, and the statistics. Section \ref{sec:oracle} derives the oracle limiting distributions of the $M$-tests under NINW. Section \ref{sec:env} characterizes the Gaussian power envelope under NINW and analyzes the asymptotic power of the oracle $M$-tests. Section \ref{sec:feasible_ld} derives the feasible limiting distributions of the $M$-tests under NINW. Section \ref{sec:finite_sample} reports finite sample results. Section \ref{sec:conclusion} concludes. Proofs are relegated to the appendices.



\section{Model, GLS detrending, and statistics}\label{sec:model}

\subsection{Nearly integrated nearly white noise}
\begin{revision}
The data-generating process used throughout the paper is
\begin{equation}\label{eq:model}
\begin{aligned}
 y_t&=\beta_0+\beta_1t+u_t,\\
 u_t&=\rho_T u_{t-1}+v_t,\qquad \rho_T=1+\frac{c}{T},\\
 v_t&=e_t+\theta_Te_{t-1},\qquad \theta_T=-1+\frac{\delta}{\sqrt T},
 \qquad t=1,\ldots,T.
\end{aligned}
\end{equation}
The stochastic conditions used to derive the asymptotic results are collected next.

\begin{assumption}\label{ass:innovations}
The initialization is $u_0=e_0=0$. The innovations $\{e_t\}_{t=1}^T$ are i.i.d. with
\[
 \mathbb{E} e_t=0,\qquad \mathbb{E} e_t^2=\sigma_e^2\in(0,\infty),\qquad \mathbb{E}|e_t|^4<\infty.
\]
The local parameters satisfy fixed $c\leq0$ and fixed $\delta>0$ as $T\to\infty$. Whenever GLS detrending is used, $\bar{c}<0$ is fixed. Note that Gaussianity is not imposed here.
\end{assumption}

Observe the autoregressive coefficient approaches one at rate $T^{-1}$, while the moving-average coefficient approaches $-1$ at rate $T^{-1/2}$ (equivalently, the corresponding MA root approaches $+1$). This is the classical NINW model of \citet{NP1994} and \citet{NgPerron1996}, augmented here by an intercept and linear trend. Its defining feature is
\begin{equation}
    \omega^2_T = \sigma^2_e(1+\theta_T)^2 = \frac{\sigma^2_e \delta^2}{T}, \quad\text{so} \quad T\omega^2_T = \sigma^2_e\delta^2. \label{eq:lrv}
\end{equation}
Hence the LRV of $v_t$ tends to zero at rate $T^{-1}$.
\end{revision}

\subsection{GLS detrending}
Let $z_t = (1,t)'$ denote the deterministic regressors. Fix $\bar{c}<0$ and define $\bar{\alpha} = 1+\bar{c}/T$. For any series $y_t$, define its quasi-differenced counterpart by
\begin{align}
    y_1^{\bar{\alpha}} = y_1, \qquad y_t^{\bar{\alpha}} = y_t - \bar{\alpha} y_{t-1}, \quad t \geq 2, \label{eq:quasi_w}\\
    z_1^{\bar{\alpha}} = z_1, \qquad z_t^{\bar{\alpha}} = z_t - \bar{\alpha} z_{t-1}, \quad t \geq 2. \label{eq:quasi_z}
\end{align}
The GLS estimator of the deterministic coefficients is then
\begin{equation}
    \hat{\psi} = \arg \min_\psi \sum_{t=1}^T (y^{\bar{\alpha}}_t - \psi'z_t^{\bar{\alpha}})^2, \label{eq:gls_minimizer}
\end{equation}
so that the GLS detrended series is $\tilde{y}_t = y_t - \hat{\psi}'z_t.$ \citet{ERS1996} recommended $\bar{c} = -13.5$ for an unknown linear time trend. Observe that the first observation is left undifferenced; this fact is important for the limiting distributions of the statistics.

\subsection{The statistics}\label{sec:stats}
The $M^{GLS}$ family of statistics is defined as
\begin{align}
    &M\!Z_{\alpha}^{GLS} = \frac{T^{-1}\tilde{y}^2_T - s^2_{AR}}{2T^{-2}\sum_{t=2}^T\tilde{y}^2_{t-1}}, \qquad M\!S\!B^{GLS} = \left(\frac{T^{-2}\sum_{t=2}^T\tilde{y}^2_{t-1}}{s^2_{AR}} \right)^{1/2}, \label{eq:mza_msb}\\
    &M\!Z_{t}^{GLS} = M\!Z_{\alpha}^{GLS} \cdot M\!S\!B^{GLS}, \qquad M\!P_{T}^{GLS} = \frac{\bar{c}^2 T^{-2} \sum_{t=2}^T\tilde{y}^2_{t-1} + (1-\bar{c})T^{-1}\tilde{y}^2_T}{s^2_{AR}}, \label{eq:mzt_mpt}
\end{align}
where $s^2_{AR}$ is an autoregressive estimate of the LRV to be handled in more detail later. For their OLS counterparts, replace $\tilde{y}$ in the level moments by OLS detrended $y$; the common estimator $s^2_{AR}$ is still constructed from GLS residuals. For the following section, we let $s^2_{AR}$ be replaced by the true LRV $\omega_T^2$ to understand the behavior of the statistics when they are gifted the true LRV.

\section{Oracle limiting distributions}\label{sec:oracle}

\subsection{Limit objects}
Before expressing the limiting distributions, it is important to define some terms that will appear in the distributions. Let $W$ denote standard Brownian motion on [0,1] and, for $c \in \mathbb{R}$, let
\begin{align}
    J_c(r) = \int_0^r e^{(r-s)c} dW(s) \label{eq:jcr}\\
    J_c^\tau(r) = J_c(r) - \hat{a} - \hat{b}r, \label{eq:jcr_tau}
\end{align}
where $\hat{b} = 12 \int_0^1(s-\frac{1}{2})J_c(s)ds$, and $\hat{a} = \int_0^1J_c(s)ds - \frac{1}{2}\hat{b}$.

Let $\varepsilon_1$ and $\varepsilon_{\infty}$ denote the limiting variables associated with the standardized first and last sample innovations $e_1/\sigma_e$ and $e_T/\sigma_e$, respectively, as constructed in Lemma \ref{A.0}. They are independent draws from the distribution of $e_t/\sigma_e$ and are independent of the Brownian motion generating $J_c$. Note that our $\varepsilon_\infty $ is equivalent to $e_\infty$ in \citet[Thm.~3.1]{NgPerron1996}.

The terms generated by $\bar{c}$ are
\begin{equation}
    \lambda = \frac{1-\bar{c}}{1-\bar{c} + \bar{c}^2/3}, \qquad q_{\bar{c}} = \lambda\varepsilon_{\infty} - \frac{3-\lambda}{2}\varepsilon_1. \label{eq:limit_constants}
\end{equation}
Finally, the typical GLS functional
\begin{equation}
    V_{c,\bar{c}}(r) = J_c(r) - r \left( \lambda J_c(1) + 3(1-\lambda) \int_0^1 sJ_c(s)ds \right). \label{gls_func}
\end{equation}

\subsection{Limiting distributions of sample moments}\label{sec:lim_of_sm}

We first derive the limiting distributions of the two sample moments fundamental to the $M$-test family. Note that $\Rightarrow$ denotes convergence in distribution throughout the paper.
\begin{proposition}[Sample moment limits under OLS]\label{prop:sm_ols} Let $\{y_t\}$ be defined by \eqref{eq:model} and Assumption~\ref{ass:innovations} hold. For OLS-detrended data,
\begin{align}
    &T^{-1}\sum_{t=2}^T \tilde{y}^2_{t-1} \Rightarrow \sigma^2_e\left(1+\delta^2 \int_0^1 J_c^{\tau}(r)^2 dr \right) \label{eq:sm1_ols} \\
    &\tilde{y}_T \Rightarrow \sigma_e(\varepsilon_{\infty} + \delta J_c^{\tau}(1)). \label{eq:sm2_ols}
\end{align}

\end{proposition}

\begin{proposition}[Sample moment limits under GLS]\label{prop:sm_gls} Let $\{y_t\}$ be defined by \eqref{eq:model} and Assumption~\ref{ass:innovations} hold, with fixed $\bar{c}<0$. For GLS-detrended data,

\begin{align}
    &T^{-1}\sum_{t=2}^T \tilde{y}^2_{t-1} \Rightarrow \sigma^2_e \left[1+\int_0^1(\delta V_{c,\bar{c}}(r) - \varepsilon_1 - q_{\bar{c}}r)^2dr \right] \label{eq:sm1_gls} \\
    &\tilde{y}_T \Rightarrow \sigma_e \Big[(1-\lambda)\varepsilon_{\infty} + \frac{1-\lambda}{2}\varepsilon_1 + \delta V_{c, \bar{c}}(1)  \Big]. \label{eq:sm2_gls}
\end{align}
\end{proposition}

\begin{remark}
As can be seen in the previous propositions, we can simply replace the undetrended process $J_c$, with the corresponding detrended process $J_c^{\tau}$ in the no-deterministics sample moment limits of \citet{NgPerron1996} to obtain the limits under an unknown linear time trend. This works as the OLS fitted line to the innovations is asymptotically negligible, so only the fitted line to the near-integrated component survives in addition to $\varepsilon_\infty$. On the other hand, under GLS detrending, the fitted line to the innovations survives asymptotically, introducing additional nuisance terms to the sample moment limits. Consequently, the GLS limits cannot be obtained from the OLS limits by simply replacing $J_c^{\tau}$ with its GLS detrended counterpart $V_{c,\bar{c}}.$
\end{remark}

\subsection{Limiting distributions of statistics}\label{sec:limit_distributions}

\begin{theorem}[Oracle $M^{OLS}$ limits under NINW]\label{thm:oracle1} Let $\{y_t\}$ be defined by \eqref{eq:model} and Assumption~\ref{ass:innovations} hold. For OLS-detrended data and known LRV,
\begin{align}
    M\!Z_{\alpha}^{OLS} &\Rightarrow \frac{(\varepsilon_{\infty}+\delta J_c^\tau(1))^2 - \delta^2}{2 \left(1+\delta^2 \int_0^1J_c^{\tau}(r)^2dr\right)}, \notag\\
    M\!S\!B^{OLS} &\Rightarrow \frac{1}{\delta}\left(1+\delta^2 \int_0^1 J_c^{\tau}(r)^2dr\right)^{1/2} ,\label{eq:mza_msb_limt_ols} \\
    M\!Z_{t}^{OLS} &\Rightarrow \frac{(\varepsilon_{\infty}+\delta J_c^{\tau}(1))^2 - \delta^2}{2\delta \left(1 + \delta^2 \int_0^1 J_c^{\tau}(r)^2dr\right)^{1/2}} \label{eq:mzt_limit_ols}
\end{align}
These limiting distributions are precisely the limiting distributions derived by \citet{NgPerron1996} in the no-deterministics case, except we replace $J_c(r)$ with $J_c^{\tau}(r)$ whenever the former appears.
\end{theorem}
\begin{revision}
\begin{proof}
Let $D_T^{OLS}=T^{-1}\sum_{t=2}^T\tilde{y}_{t-1}^2$. Since $T\omega_T^2=\sigma_e^2\delta^2$, the definitions in \eqref{eq:mza_msb} imply
\[
 M\!Z_{\alpha}^{OLS}=\frac{\tilde{y}_T^2-T\omega_T^2}{2D_T^{OLS}},\qquad
 M\!S\!B^{OLS}=\left(\frac{D_T^{OLS}}{T\omega_T^2}\right)^{1/2}.
\]
Apply Proposition~\ref{prop:sm_ols} and the continuous-mapping theorem; the limit for $M\!Z_{t}^{OLS}$ follows from $M\!Z_{t}^{OLS}=M\!Z_{\alpha}^{OLS}M\!S\!B^{OLS}$.
\end{proof}
\end{revision}

\begin{theorem}[Oracle $M^{GLS}$ limits under NINW]\label{thm:oracle2} Let $\{y_t\}$ be defined by \eqref{eq:model} and Assumption~\ref{ass:innovations} hold, with fixed $\bar{c}<0$. For GLS-detrended data and known LRV,
\begin{align}
    M\!Z_{\alpha}^{GLS} &\Rightarrow \frac{[(1-\lambda)\varepsilon_{\infty} + \frac{1-\lambda}{2}\varepsilon_1+ \delta V_{c,\bar{c}}(1)]^2 - \delta^2}{2\left(1+ \int_0^1[\delta V_{c,\bar{c}}(r) - \varepsilon_1 - q_{\bar{c}}r]^2dr\right)}, \label{eq:mza_limit_gls} \\
    M\!S\!B^{GLS} &\Rightarrow \frac{1}{\delta}\left(1 + \int_0^1 [\delta V_{c,\bar{c}}(r) - \varepsilon_1 - q_{\bar{c}}r]^2dr \right)^{1/2} \label{eq:msb_limit_gls} \\
    M\!Z_{t}^{GLS} &\Rightarrow \frac{[(1-\lambda)\varepsilon_{\infty} + \frac{1-\lambda}{2}\varepsilon_1+ \delta V_{c,\bar{c}}(1)]^2 - \delta^2}{2\delta \left( 1 + \int_0^1[\delta V_{c,\bar{c}}(r) - \varepsilon_1 - q_{\bar{c}}r]^2dr\right)^{1/2}}, \label{eq:mzt_limit_gls} \\
    M\!P_{T}^{GLS} &\Rightarrow \frac{1}{\delta^2} \Bigg\{\bar{c}^2 \left(1 + \int_0^1[\delta V_{c,\bar{c}}(r) - \varepsilon_1 - q_{\bar{c}}r]^2dr\right) \notag\\
    &\qquad{} + (1-\bar{c}) \left( (1-\lambda)\varepsilon_{\infty} + \frac{1-\lambda}{2}\varepsilon_1 + \delta V_{c,\bar{c}}(1) \right)^2 \Bigg\} \label{eq:mpt_limit_gls}
\end{align}

\end{theorem}
\begin{revision}
\begin{proof}
Set $D_T^{GLS}=T^{-1}\sum_{t=2}^T\tilde{y}_{t-1}^2$. Multiplying the numerator and denominator of each statistic in \eqref{eq:mza_msb}--\eqref{eq:mzt_mpt} by the appropriate power of $T$, and using $T\omega_T^2=\sigma_e^2\delta^2$, reduces every statistic to a continuous function of $(D_T^{GLS},\tilde{y}_T)$. Proposition~\ref{prop:sm_gls} and continuous mapping then give the four displayed limits.
\end{proof}
\end{revision}

\begin{remark}
    Clearly, the limiting distributions of the tests, particularly the GLS based tests, carry extra nuisance terms not present in the standard limiting distributions of the tests. Yet,
    observe that as $\delta \rightarrow \infty$, we recover the standard limiting distributions of \citet{NgPerron2001} and \citet{Stock1990}. To see this, define
    \[
    D = 1+\int_0^1\left[ \delta V_{c,\bar{c}}(r) - \varepsilon_1 - q_{\bar{c}}r \right]^2dr, \qquad E = (1-\lambda)\varepsilon_{\infty} + \frac{1-\lambda}{2}\varepsilon_1 + \delta V_{c,\bar{c}}(1).
    \]
    Then
    \[
    \frac{D}{\delta^2} \xrightarrow{\mathrm{a.s.}} \int_0^1 V_{c,\bar{c}}(r)^2dr, \qquad \frac{E}{\delta} \xrightarrow{\mathrm{a.s.}} V_{c,\bar{c}}(1).
    \]
    Consequently,
    \[
    \frac{E^2-\delta^2}{2D} \xrightarrow[\delta\to\infty]{\mathrm{a.s.}} \frac{V_{c,\bar{c}}(1)^2 - 1}{2\int_0^1 V_{c,\bar{c}}(r)^2dr},
    \]
    the standard limiting distribution of $M\!Z_{\alpha}^{GLS}$ in the case of an unknown linear time trend. Similar arguments yield the same results for the other $M$-tests.

    The opposite regime is nonuniform. As $\delta \rightarrow 0$, all dependence on the true local-to-unity parameter $c$ enters through the vanishing term $\delta V_{c, \bar{c}}$ or $\delta J_c^{\tau}$. The surviving leading terms are instead determined by the boundary terms $\varepsilon_1$ and/or $\varepsilon_{\infty}$ and the fixed GLS design parameter $\bar{c}$. Thus, the leading null and local alternative laws become indistinguishable as $\delta \rightarrow 0$. This is the partly the mechanism behind the deterioration in size-adjusted power examined later.
\end{remark}

\begin{remark}\label{rem:innovation_dep}
    Since the $M$-test limiting distributions contain the terms $\varepsilon_1$ and $\varepsilon_{\infty}$, they depend on the full distribution of the standardized innovations, $e_t/\sigma_e$, beyond their first two moments. Hence Gaussian NINW critical values need not be valid for other innovation laws. Note that $M\!S\!B^{OLS}$ is the exception to this.

\end{remark}

\begin{corollary}[Support bounds]\label{cor:support_bounds}
Under the conditions of Theorems~\ref{thm:oracle1}--\ref{thm:oracle2}, for every fixed $c\leq0$ the oracle NINW limits satisfy
\begin{equation}
    M\!Z_{\alpha} \geq -\frac{\delta^2}{2}, \qquad
    M\!S\!B \geq \frac{1}{\delta}, \qquad
    M\!Z_{t} \geq -\frac{\delta}{2}, \qquad
    M\!P_{T}^{GLS} \geq \frac{\bar{c}^2}{\delta^2} \quad\text{a.s.} \label{eq:oracle_bounds}
\end{equation}
Consequently, for a critical value $cv$ the asymptotic rejection probability is exactly zero whenever
\[
 \delta<\sqrt{2|cv|},\qquad
 \delta<\frac1{cv},\qquad
 \delta<2|cv|,\qquad
 \delta<\frac{|\bar{c}|}{\sqrt{cv}},
\]
respectively. Moreover, the finite sample rejection probability converges to $0$ as $T \to \infty.$
\end{corollary}

\begin{remark}
For the linear time trend case, \citet{NgPerron2001} report 5\% critical values of -17.3, 0.168, -2.91, and 5.48 for $M\!Z_{\alpha}^{GLS}, M\!S\!B^{GLS}, M\!Z_{t}^{GLS}$ and $M\!P_{T}^{GLS}$, respectively. Corollary~\ref{cor:support_bounds} therefore gives
\[
\delta^*_{M\!Z_{\alpha}} \approx 5.88, \quad \delta^*_{M\!S\!B} \approx 5.95, \quad \delta^*_{M\!Z_{t}} \approx  5.82, \quad \delta^*_{M\!P_{T}} \approx 5.77.
\]
 Hence, all four GLS-based 5\% size tests enter their zero asymptotic rejection probability region at approximately $\delta^* =$ 5.8--6.0. The magnitude of this region can be made more concrete by relating it to the negative moving average coefficient itself. Since
 \[
 \theta_T = -1 + \frac{\delta}{\sqrt{T}},
 \]
 at $T=100$, the thresholds above correspond to
 \[
 \theta_T  \approx -0.412, \quad -0.405, \quad -0.418, \quad -0.423
 \]
 for $M\!Z_{\alpha}^{GLS}, M\!S\!B^{GLS}, M\!Z_{t}^{GLS}$, and $M\!P_{T}^{GLS}$, respectively. Thus, for example, along a NINW sequence with $\delta<5.88$, we have that for
 \[
 \theta_T<-1+\frac{5.88}{\sqrt{T}},
 \]
 the size and local asymptotic power of $M\!Z_{\alpha}^{GLS}$ computed using the 5\% asymptotic critical values converge to zero. For the $M^{OLS}$ tests, $\delta^* \approx 6.40$; an even greater region of zero asymptotic rejection probability.
\end{remark}

\section{Asymptotic power of the tests}\label{sec:env}

\subsection{The asymptotic Gaussian power envelope}

The limiting distributions of Section~\ref{sec:oracle} do not by themselves
determine whether the behavior of the $M$-tests reflects a harder testing
problem or inefficiency of the statistics. To distinguish the two, we
characterize the maximal asymptotic power attainable in the NINW
experiment under an unknown linear time trend, following the construction of \citet{ERS1996}. Their envelope
is derived under the condition that $\{v_t\}$ has a strictly positive
spectral density at frequency zero (their Condition~A). By \eqref{eq:lrv}
the LRV $\omega_T^2=\sigma_e^2\delta^2/T$ vanishes in the
limit, so that condition fails and equality of the two power bounds cannot
be assumed.

Assume, just in this section, that in addition to Assumption \ref{ass:innovations}$,\{e_t\}$ is Gaussian and that
$\Sigma_T=\operatorname{Cov}(v_1,\dots,v_T)$, is available to the
researcher. Granting this knowledge can only raise attainable power, so the
resulting bound is valid for any feasible trend invariant procedure. Let
$y^{\bar{\alpha}},Z^{\bar{\alpha}}$ denote the data and trend regressors quasi-differenced
at $\bar{\alpha}$ as in \eqref{eq:quasi_w}--\eqref{eq:quasi_z}, and let
$y^{1},Z^{1}$ denote the same construction at $\bar{\alpha}=1$, that is, first
differences with the initial observation retained. Define the Gaussian
criteria
\begin{equation}
L_T(\bar{\alpha},\psi)=\bigl(y^{\bar{\alpha}}-Z^{\bar{\alpha}}\psi\bigr)'\,\Sigma_T^{-1}\,
\bigl(y^{\bar{\alpha}}-Z^{\bar{\alpha}}\psi\bigr),
\qquad
L_T(1,\psi)=\bigl(y^{1}-Z^{1}\psi\bigr)'\,\Sigma_T^{-1}\,
\bigl(y^{1}-Z^{1}\psi\bigr),
\label{eq:ersobjective}
\end{equation}
each of which is, up to an additive constant, minus two times the Gaussian
log-likelihood under the alternative $c=\bar{c}$ and under the null $c=0$
(where $\bar{\alpha}=1$), respectively. The most powerful test of $H_0\colon c=0$
against the point alternative $c=\bar{c}$ that is invariant to the trend
coefficients rejects for small values of the likelihood-ratio statistic
\begin{equation}
L^{*,\tau}_T(\bar{c})
=\min_{\psi}L_T(\bar{\alpha},\psi)-\min_{\psi}L_T(1,\psi),
\label{eq:Lstar}
\end{equation}
the difference in weighted sums of squared residuals from two constrained
GLS regressions, one imposing the alternative and one imposing the null
\citep[eqn.~(6)]{ERS1996}. Because the asymptotically sufficient statistic for
$c$ is two-dimensional, no uniformly most powerful invariant test exists:
\eqref{eq:Lstar} defines an infinite family of point-optimal tests indexed
by $\bar{c}$, none dominating the others at every alternative.

Under standard errors satisfying Condition~A ($\{v_t\}$ has a positive spectral density at frequency zero), the rejection region of the
test indexed by $\bar{c}$ admits a continuous-time representation. With
$V_{c,\bar{c}}$ as in \eqref{gls_func}, the asymptotic power function of the
size-$\alpha$ test indexed by $\bar{c}$, evaluated at the true local
parameter $c$, is
\begin{equation}
\pi^{\tau}(c,\bar{c};\alpha)
=\Pr\!\left[\;\bar{c}^{2}\!\int_0^1 V_{c,\bar{c}}(r)^2\,dr
+(1-\bar{c})\,V_{c,\bar{c}}(1)^2 \,<\, b^{\tau}_{\alpha}(\bar{c})\right],
\label{eq:pitau}
\end{equation}
where $b^{\tau}_{\alpha}(\bar{c})$ is defined by the same probability
statement at $c=0$ equaling $\alpha$ \citep[eqn.~(8)]{ERS1996}. Note $c$ indexes the true local
alternative at which power is evaluated, while $\bar{c}$ indexes which member
of the point-optimal family is used. Since the test indexed by $\bar{c}$ is
most powerful invariant against $c=\bar{c}$, no invariant size-$\alpha$ test
can exceed $\pi^{\tau}(\bar{c},\bar{c};\alpha)$ at that alternative. The
pointwise upper bound on asymptotic power is therefore the diagonal,
\begin{equation}
\Pi^{\tau}_{\mathrm{ERS}}(c;\alpha)\equiv\pi^{\tau}(c,c;\alpha),
\label{eq:ers_envelope}
\end{equation}
the ERS linear-trend Gaussian power envelope.

Let $\Pi^{\tau}_{\mathrm{NINW}}(c,\delta;\alpha)$ denote the corresponding
object under \eqref{eq:model} and Assumption~\ref{ass:innovations}: the maximal
asymptotic power attainable at the fixed alternative $(c,\delta)$ by
size-$\alpha$ tests invariant to the intercept and linear trend. The
results of Section~\ref{sec:oracle} might suggest that this envelope is
degraded relative to \eqref{eq:ers_envelope} since the statistics' limits depend
on $\delta$, are contaminated by the boundary variates, and lose their
dependence on $c$ as $\delta\to0$. The following theorem shows that it is
not.

\begin{theorem}[Gaussian power envelope under NINW]\label{thm:power_envelope}
Let $\Pi^{\tau}_{NINW}(c,\delta;\alpha)$ be the Gaussian power envelope under \eqref{eq:model}, Assumption~\ref{ass:innovations}, and the Gaussian oracle information set above. Then, for every fixed $c<0$, $\delta>0$, and $0<\alpha<1$, we have
\begin{equation}
    \Pi^{\tau}_{NINW}(c, \delta;\alpha) = \Pi^{\tau}_{ERS}(c;\alpha). \label{eq:envelope_equivalence}
\end{equation}
Hence the asymptotic Gaussian power envelope under an unknown linear time trend is identical to the standard ERS power envelope under an unknown linear time trend, and moreover, it does not depend on $\delta$.
\end{theorem}
\begin{remark}
This theorem has a consequential interpretation. Although $\delta$ enters the limiting distributions of the NINW $M$-tests, it does not enter the Gaussian power envelope. Thus the NINW process does not offer unit root tests intrinsically less information than in the standard case. Hence any loss in power at lower $\delta$'s can be interpreted as a fault of the oracle $M$-tests, not the DGP. Note that attainability of the same envelope by a fully feasible procedure is a separate question.
\end{remark}

\subsection{Asymptotic power of the \texorpdfstring{$M$}{M}-tests}
Now that we have characterized the Gaussian power envelope under NINW, we examine the asymptotic power of the $M$-tests. Figure~\ref{fig:asymptotic_power_mpt} graphs the asymptotic power function of the $MP_T^{GLS}$ test at various $\delta$'s and the Gaussian power envelope. Power is simulated from the asymptotic limiting functionals using 100,000 null and 100,000 alternative Monte Carlo replications and a $50,000$ point discretization grid. The $M\!P_{T}$ curves use $\bar{c}=-13.5$ and $\delta$ specific critical values; the Gaussian power envelope is simulated point-wise, $\bar{c}=c$, at each local alternative. Size is set to 0.05.

\FloatBarrier

\begin{figure}[t]
    \centering
    \includegraphics[width=\textwidth]{no_sb_paper/figures/fig_mp_power_ninw_paper.pdf}
    \caption{Asymptotic power of $M\!P_{T}$ at various deltas compared to the Gaussian power envelope. The vertical line is $ c =-13.5.$}
    \label{fig:asymptotic_power_mpt}
\end{figure}

Figure~\ref{fig:asymptotic_power_mpt} illustrates that the power loss becomes increasingly severe as $\delta$ decreases. For $\delta=20$, the power curve is close to the envelope, consistent with the recovery of the standard limiting distribution as $\delta$ becomes large. At moderate to low values of $\delta$, however, a substantial loss of power exists. The inefficiency relative to the envelope is especially clear at $\delta=1$. Power remains below 0.2 as $c=-30$, whereas the envelope is almost at one.

It is important to emphasize that these power losses occur when the LRV is treated as known. Therefore, these results precede any difficulties with LRV estimation under NINW. In fact, a feasible $M^{GLS}$ test that shares the same limiting distributions as the oracle case would only inherit the intrinsic power inefficiency problem. As a further robustness check, we reoptimized $\bar{c}$ at $\delta=1$ for the various tests and found that power was still well below the envelope, despite this completely infeasible advantage. Thus, we can conclude that the $M^{GLS}$ tests are intrinsically inefficient in terms of asymptotic power under NINW.

Note that similar results hold for the OLS variants of the $M$-tests, with one key exception, $M\!S\!B^{OLS}$, which actually has a power function invariant to $\delta$. This was observed by \citet{NgPerron1996} in the no-deterministics case. The fact that the test's asymptotic power is invariant to $\delta$ can be derived from its limiting distribution. Let $A^{\tau}_c = \int_0^1J_c^{\tau}(r)^2dr$, then the limiting distribution from \eqref{eq:mza_msb_limt_ols} can be rewritten as $M\!S\!B^{OLS} \Rightarrow (A^{\tau}_c + \delta^{-2})^{1/2}$. Then, the corresponding lower tail null critical value is $(a^{\tau}_{\alpha}+\delta^{-2})^{1/2}$, where $a^{\tau}_{\alpha}$ is the lower $\alpha$-quantile of $A^{\tau}_0.$ Hence, the size corrected rejection event is $A^{\tau}_c < a^{\tau}_{\alpha}$ because the common term $\delta^{-2}$ cancels. Thus, the asymptotic power of $M\!S\!B^{OLS}$ is independent of $\delta$ and coincides with its standard local-to-unity power function. This property is unique to the $M\!S\!B^{OLS}$ test.

\FloatBarrier

\section{Feasible limiting distributions}\label{sec:feasible_ld}
The entire analysis so far treated the LRV $\omega^2_T=\sigma^2_e\delta^2/T$ as known. Obviously, the practitioner does not implement these oracle tests, but rather a feasible test which must use an estimator for the LRV. Since $\omega^2_T \to 0$ as $T \to \infty$, the condition for feasible and oracle equivalence (of limiting distributions) is stronger than just absolute consistency.

To make this clear, let $\hat{\omega}^2_T$ be an estimator of $\omega^2_T$. Absolute consistency implies
\begin{equation}
    \hat{\omega}^2_T - \omega^2_T \stackrel{\mathbb{P}}{\to} 0. \label{eq:abs_con}
\end{equation}
When $\omega^2_T$ is a fixed positive constant, \eqref{eq:abs_con} implies ratio consistency, but under NINW it does not. For example, suppose $\hat{\omega}^2_T=2\omega^2_T$. Then \eqref{eq:abs_con} is satisfied even though it remains twice the true LRV at all $T$. The relevant requirement is therefore
\begin{equation}
    \frac{\hat{\omega}^2_T}{\omega^2_T} \stackrel{\mathbb{P}}{\to} 1, \quad\text{equivalently,} \quad T\hat{\omega}^2_T \stackrel{\mathbb{P}}{\to} \sigma^2_e \delta^2. \label{eq:rel_con}
\end{equation}

We refer to \eqref{eq:rel_con} as ratio consistency. If it holds, then $\omega^2_T$ can be replaced by $\hat{\omega}^2_T$, leaving the oracle limiting distributions unchanged.

\subsection{The autoregressive long-run variance estimator}
We consider the autoregressive estimator used by \citet{NgPerron1996,NP1998} and \citet{NgPerron2001}. Throughout this subsection, $\tilde{y}_t$ denotes the GLS detrended series. The resulting $s^2_{AR}$ is used for both the OLS and GLS versions of the $M$-statistics; the distinction between them concerns the level moments, not the construction of the LRV estimator. For a truncation order $k$, estimate
\begin{equation}
    \Delta\tilde{y}_t = \hat{b}_0\tilde{y}_{t-1} + \sum_{j=1}^k \hat{b}_j\Delta\tilde{y}_{t-j} + \hat{e}_{t,k} \label{eq:ADF}
\end{equation}
by least squares over $t=k+2,\ldots,T$, and define
\begin{equation}
    \hat{b}(1)= \sum_{j=1}^k \hat{b}_j, \quad \hat{\sigma}^2_{ek}= \frac{1}{T}\sum_{t=k+2}^T \hat{e}^2_{t,k}, \quad s^2_{AR} = \frac{\hat{\sigma}^2_{ek}}{(1-\hat{b}(1))^2}. \label{eq:s^2_ar}
\end{equation}
\begin{revision}
The existing literature establishes important but distinct parts of the feasible problem. In the no-deterministics case, \citet{NgPerron1996} state that the autoregressive estimator yields the same NINW limits as the oracle case under $k\to\infty$ and $k/T\to0$. \citet{NP1998} subsequently establish that $s^2_{AR} \stackrel{\mathbb{P}}{\to} 0$ when the spectral density at frequency zero collapses. That is, they establish absolute consistency for the limiting value zero, but again, that by itself does not determine the ratio in \eqref{eq:rel_con}. Finally, \citet{NgPerron2001} establish that, under NINW, $k^2s^2_{AR}=O_\mathbb{P}(1)$ and $Ts^2_{AR} = O_\mathbb{P}(1)$ can both hold only if
\begin{equation}
    \frac{k}{\sqrt{T}} \to \kappa, \quad 0<\kappa<\infty. \label{eq:rate_con}
\end{equation}
Thus, $k \propto \sqrt{T}$ is the rate required to keep the feasible $M$-tests bounded. It remains to determine whether this rate also delivers ratio consistency.
\end{revision}

\begin{proposition}[$s^2_{AR}$ under NINW]\label{prop:feasible_s2ar}
Let $\{y_t\}$ be defined by \eqref{eq:model}, Assumption~\ref{ass:innovations} hold, and construct $s^2_{AR}$ from GLS-detrended data as in \eqref{eq:ADF}--\eqref{eq:s^2_ar}. If the deterministic integer sequence $k=k_T$ satisfies \eqref{eq:rate_con}, then
\begin{equation}
    \frac{s^2_{AR}}{\omega^2_T} \stackrel{\mathbb{P}}{\to} \coth^2\left(\frac{\kappa\delta}{2}\right), \label{eq:s^2_rel_con}
\end{equation}
equivalently, $Ts^2_{AR} \stackrel{\mathbb{P}}{\to}\sigma^2_e \delta^2\coth^2(\kappa\delta/2).$ It follows for finite $\kappa$ and $\delta$ that $\coth^2(\kappa\delta/2)>1$ so the oracle and feasible $M$-test limiting distributions differ.
\end{proposition}
This proposition leads us directly to the following two theorems.

\begin{theorem}[Feasible $M^{OLS}$ limiting distributions under NINW]\label{thm:feasible_mols}
{\color{revisioncolor}Let $\{y_t\}$ be defined by\eqref{eq:model} and Assumption~\ref{ass:innovations} hold. Use OLS-detrended data for the level moments and construct $s^2_{AR}$ from GLS-detrended data as in \eqref{eq:ADF}--\eqref{eq:s^2_ar}. If the deterministic lag sequence satisfies $k/\sqrt T\to\kappa\in(0,\infty)$, then}
\begin{align}
    M\!Z_{\alpha}^{OLS} &\Rightarrow \frac{(\varepsilon_{\infty}+\delta J_c^\tau(1))^2 - \delta^2 \coth^2(\frac{\kappa\delta}{2})}{2 \left(1+\delta^2 \int_0^1J_c^{\tau}(r)^2dr\right)}, \notag\\
    M\!S\!B^{OLS} &\Rightarrow \frac{\tanh(\frac{\kappa\delta}{2})}{\delta}\left(1+\delta^2 \int_0^1 J_c^{\tau}(r)^2dr\right)^{1/2} ,\label{eq:mza_msb_limt_ols_feasible} \\
    M\!Z_{t}^{OLS} &\Rightarrow \frac{(\varepsilon_{\infty}+\delta J_c^{\tau}(1))^2 - \delta^2 \coth^2(\frac{\kappa\delta}{2})}{2\delta \coth(\frac{\kappa\delta}{2})\left(1 + \delta^2 \int_0^1 J_c^{\tau}(r)^2dr\right)^{1/2}} \label{eq:mzt_limit_ols_feasible}
\end{align}

\end{theorem}
\begin{revision}
\begin{proof}
By Proposition~\ref{prop:feasible_s2ar}, $Ts^2_{AR}\stackrel{\mathbb{P}}{\to}\sigma_e^2\delta^2\coth^2(\kappa\delta/2)$. Combine this with Proposition~\ref{prop:sm_ols} in the algebraic definitions of $M\!Z_{\alpha}^{OLS}$ and $M\!S\!B^{OLS}$, and apply Slutsky's theorem. The $M\!Z_{t}^{OLS}$ limit follows by multiplication.
\end{proof}
\end{revision}

\begin{theorem}[Feasible $M^{GLS}$ limiting distributions under NINW]\label{thm:feasible_mgls}
{\color{revisioncolor}Let $\{y_t\}$ be defined by\eqref{eq:model} and Assumption~\ref{ass:innovations} hold. Use GLS-detrended data for the level moments and construct $s^2_{AR}$ from \eqref{eq:ADF}--\eqref{eq:s^2_ar}. If the deterministic lag sequence satisfies $k/\sqrt T\to\kappa\in(0,\infty)$, then}

\begin{align}
    M\!Z_{\alpha}^{GLS} &\Rightarrow \frac{[(1-\lambda)\varepsilon_{\infty} + \frac{1-\lambda}{2}\varepsilon_1+ \delta V_{c,\bar{c}}(1)]^2 - \delta^2\coth^2(\frac{\kappa\delta}{2})}{2\left(1+ \int_0^1[\delta V_{c,\bar{c}}(r) - \varepsilon_1 - q_{\bar{c}}r]^2dr\right)}, \label{eq:mza_limit_gls_feasible} \\
    M\!S\!B^{GLS} &\Rightarrow \frac{\tanh(\frac{\kappa\delta}{2})}{\delta}\left(1 + \int_0^1 [\delta V_{c,\bar{c}}(r) - \varepsilon_1 - q_{\bar{c}}r]^2dr \right)^{1/2} \label{eq:msb_limit_gls_feasible} \\
    M\!Z_{t}^{GLS} &\Rightarrow \frac{[(1-\lambda)\varepsilon_{\infty} + \frac{1-\lambda}{2}\varepsilon_1+ \delta V_{c,\bar{c}}(1)]^2 - \delta^2\coth^2(\frac{\kappa\delta}{2})}{2\delta\coth(\frac{\kappa\delta}{2}) \left( 1 + \int_0^1[\delta V_{c,\bar{c}}(r) - \varepsilon_1 - q_{\bar{c}}r]^2dr\right)^{1/2}}, \label{eq:mzt_limit_gls_feasible} \\
    M\!P_{T}^{GLS} &\Rightarrow \frac{\tanh^2(\frac{\kappa\delta}{2})}{\delta^2} \Bigg\{\bar{c}^2 \left(1 + \int_0^1[\delta V_{c,\bar{c}}(r) - \varepsilon_1 - q_{\bar{c}}r]^2dr\right) \notag\\
    &\qquad{} + (1-\bar{c}) \left( (1-\lambda)\varepsilon_{\infty} + \frac{1-\lambda}{2}\varepsilon_1 + \delta V_{c,\bar{c}}(1) \right)^2 \Bigg\} \label{eq:mpt_limit_gls_feasible}
\end{align}

\end{theorem}
\begin{revision}
\begin{proof}
Combine Proposition~\ref{prop:feasible_s2ar} with the GLS sample-moment limits in Proposition~\ref{prop:sm_gls}. Slutsky's theorem gives the limits for $M\!Z_{\alpha}^{GLS}$ and $M\!S\!B^{GLS}$; multiplication gives $M\!Z_{t}^{GLS}$, while direct substitution into \eqref{eq:mzt_mpt} gives the $M\!P_{T}^{GLS}$ limit.
\end{proof}
\end{revision}

\begin{remark}\label{rem:feasible_interpretation}
At the proportional lag rate \(k/\sqrt T\to\kappa\), the effect of
estimating the LRV is summarized by $\coth^2\left(\frac{\kappa\delta}{2}\right).$
Terms involving the LRV itself acquire this inflation
factor, while terms involving its inverse square root acquire the factor
$\tanh\left(\frac{\kappa\delta}{2}\right).$

For finite \(\kappa\), these factors differ from one, so the feasible
limits do not coincide with the oracle limits. They approach the oracle
limits only as \(\kappa\delta\) becomes large. Hence, when the selected lag order is sufficiently large relative to the decay rate of the moving average component, the feasible tests begin to inherit the oracle behavior. In particular, for values of $\delta$ that lie in the zero size region identified in Corollary \ref{cor:support_bounds}, conventional critical values can lead to under-rejection. It is important to note, however, that the feasible zero asymptotic rejection probability region, when it exists, will always be contained within the corresponding oracle region as a result of LRV overestimation counteracting the oracle under-rejection.
\end{remark}

\begin{remark}
\label{rem:kappadelta}
The coefficients in the autoregressive representation of the
near-canceling moving average component decay approximately as
\[
\left(1-\frac{\delta}{\sqrt T}\right)^j
\approx
\exp\left(-\frac{\delta j}{\sqrt T}\right).
\]
When $k/\sqrt T\to\kappa$, the last included coefficient is therefore
of approximate order $e^{-\kappa\delta}$. The product
$\kappa\delta$ measures how far the lag truncation extends into this
decaying sequence, which explains why the additional distortion caused
by LRV estimation depends on $\kappa$ and $\delta$
through their product.
\end{remark}

\section{Finite sample simulations}\label{sec:finite_sample}
The preceding oracle and feasible results suggest two sources of nonstandard behavior under NINW: (1) when the LRV is known, the $M$-tests undersize under conventional critical values for sufficiently small $\delta$, and (2) at the proportional lag rate, $s^2_{AR}$ asymptotically overestimates the LRV, increasing rejection. Hence, the finite sample analysis explores how these two factors interact under data-dependent lag selection rules.

Proposition \ref{prop:feasible_s2ar} and Theorems \ref{thm:feasible_mols}--\ref{thm:feasible_mgls} concern deterministic lag sequences satisfying $k/\sqrt T\to\kappa\in(0,\infty)$. The simulations examine how the mechanisms identified by these results manifest under data-dependent lag selection by MAIC and MBIC. Establishing limiting distributions for these selected-lag procedures is beyond the scope of this paper.


\subsection{Lag selection rules}
In this paper, we consider only the modified criteria $MAIC^{GLS}, MBIC^{GLS}, MAIC^{OLS},$ and $MBIC^{OLS}$ which originate from \citet{NgPerron2001} and \citet{PQ2007}, respectively. We do not consider the standard criteria like AIC and BIC as can be seen in \citet{akaike_new_1974} and \citet{BIC1974}, respectively, as \citet{NgPerron2001} already show their unsuitability under DGPs featuring large negative moving average coefficients.

The modified criteria choose a value $k_{MIC} = \arg\min_{k\in\{0,\ldots,k_{max}\}} MIC(k)$ where
\begin{equation}
    MIC(k) = \ln(\hat{\sigma}^2_k) + \frac{C_T(\tau_T(k)+k)}{T-k_{max}-1} \label{eq:mic},
\end{equation}
and
\begin{equation}
    \hat{\sigma}^2_k = \frac{1}{T-k_{max}-1}\sum_{t=k_{max}+2}^T \hat{e}^2_{t,k}, \quad \tau_T(k) = \frac{1}{\hat{\sigma}^2_k}\hat{b}^2_0\sum_{t=k_{max}+2}^T \tilde{y}^2_{t-1}. \label{eq:mic_components}
\end{equation}
$k_{max}$ is usually set to $k_{max}=[ 12(T/100)^{1/4}]$ in accordance with \citet{Schwert1989} and for MAIC $C_T=2$ whereas for MBIC $C_T=\ln(T-k_{max}-1)$. For the $MIC^{GLS}$ criteria, \eqref{eq:mic} is fed GLS detrended data while, for the $MIC^{OLS}$ criteria, it is fed OLS detrended data.

\subsection{Size results}\label{sec:size}
We hold $\delta$ fixed across sample sizes and $e_0=u_0=0$ with $e_t\stackrel{\mathrm{i.i.d.}}{\sim}N(0,1)$.Unless otherwise noted, 30,000 Monte Carlo replications are used. $s^2_{AR}$ as defined in \eqref{eq:s^2_ar} is constructed from GLS detrended data for all procedures. To allow lag orders at the rate required by Proposition \ref{prop:feasible_s2ar}, we set $k_{max}=[ 2\sqrt{T} ] $, which yields $k_{max} = 20,31,44$, for $T=100,250, 500$, respectively\footnote{This does not necessarily imply that this extended cap is preferable to the conventional cap. But, under the conventional cap, oversizing is more extreme at low delta and exhibits similar undersizing elsewhere.}. This choice differs from that of \citet{NgPerron2001} and \citet{PQ2007} as we find, for large negative moving average coefficients, $k$ will often bind to $k_{max}$ when the traditional cap is used.

\begin{sidewaystable}[p]
\centering
\normalsize\singlespacing
\begin{threeparttable}
\caption{Empirical size of the $MZ_\alpha^{OLS}$ and $MZ_\alpha^{GLS}$ tests using conventional 5\% asymptotic critical values under extended lag cap}
\label{tab:mza-delta-asymptotic-size}
\setlength{\tabcolsep}{5.5pt}
\begin{tabular}{c S[table-format=1.1] S[table-format=-1.3] *{8}{S[table-format=1.3]}}
\toprule
& &
& \multicolumn{4}{c}{$MZ_\alpha^{OLS}$}
& \multicolumn{4}{c}{$MZ_\alpha^{GLS}$} \\
\cmidrule(lr){4-7}\cmidrule(lr){8-11}
{$T$} & {$\delta$} & {$\theta$}
& {$MAIC^{GLS}$}
& {$MAIC^{OLS}$}
& {$MBIC^{GLS}$}
& {$MBIC^{OLS}$}
& {$MAIC^{GLS}$}
& {$MAIC^{OLS}$}
& {$MBIC^{GLS}$}
& {$MBIC^{OLS}$}
\\
\midrule
\multirow{10}{*}{100}
& 0.5 & -0.950 & 0.183 & 0.485 & 0.198 & 0.503 & 0.194 & 0.499 & 0.210 & 0.516 \\
& 1.0 & -0.900 & 0.124 & 0.297 & 0.135 & 0.313 & 0.137 & 0.312 & 0.150 & 0.328 \\
& 1.5 & -0.850 & 0.075 & 0.166 & 0.083 & 0.177 & 0.086 & 0.179 & 0.097 & 0.192 \\
& 2.0 & -0.800 & 0.053 & 0.103 & 0.060 & 0.111 & 0.065 & 0.117 & 0.074 & 0.127 \\
& 2.5 & -0.750 & 0.037 & 0.066 & 0.042 & 0.073 & 0.048 & 0.078 & 0.055 & 0.085 \\
& 3.0 & -0.700 & 0.031 & 0.050 & 0.035 & 0.056 & 0.042 & 0.062 & 0.049 & 0.069 \\
& 3.5 & -0.650 & 0.026 & 0.042 & 0.030 & 0.046 & 0.038 & 0.054 & 0.044 & 0.059 \\
& 4.0 & -0.600 & 0.026 & 0.037 & 0.030 & 0.041 & 0.039 & 0.049 & 0.045 & 0.054 \\
& 4.5 & -0.550 & 0.026 & 0.036 & 0.030 & 0.040 & 0.040 & 0.050 & 0.046 & 0.054 \\
& 5.0 & -0.500 & 0.026 & 0.034 & 0.030 & 0.038 & 0.039 & 0.047 & 0.045 & 0.052 \\
\specialrule{1.05pt}{2pt}{2pt}
\multirow{10}{*}{250}
& 0.5 & -0.968 & 0.059 & 0.353 & 0.064 & 0.369 & 0.063 & 0.356 & 0.069 & 0.372 \\
& 1.0 & -0.937 & 0.027 & 0.131 & 0.030 & 0.140 & 0.030 & 0.135 & 0.034 & 0.144 \\
& 1.5 & -0.905 & 0.013 & 0.048 & 0.015 & 0.053 & 0.016 & 0.052 & 0.018 & 0.057 \\
& 2.0 & -0.874 & 0.009 & 0.024 & 0.010 & 0.027 & 0.011 & 0.027 & 0.014 & 0.030 \\
& 2.5 & -0.842 & 0.006 & 0.014 & 0.007 & 0.016 & 0.008 & 0.017 & 0.011 & 0.020 \\
& 3.0 & -0.810 & 0.006 & 0.012 & 0.008 & 0.014 & 0.010 & 0.016 & 0.013 & 0.019 \\
& 3.5 & -0.779 & 0.007 & 0.012 & 0.010 & 0.014 & 0.012 & 0.017 & 0.016 & 0.020 \\
& 4.0 & -0.747 & 0.007 & 0.011 & 0.010 & 0.014 & 0.013 & 0.018 & 0.018 & 0.022 \\
& 4.5 & -0.715 & 0.010 & 0.013 & 0.014 & 0.016 & 0.017 & 0.020 & 0.023 & 0.025 \\
& 5.0 & -0.684 & 0.012 & 0.015 & 0.017 & 0.019 & 0.019 & 0.022 & 0.025 & 0.027 \\
\specialrule{1.05pt}{2pt}{2pt}
\multirow{10}{*}{500}
& 0.5 & -0.978 & 0.022 & 0.215 & 0.024 & 0.227 & 0.024 & 0.217 & 0.026 & 0.229 \\
& 1.0 & -0.955 & 0.006 & 0.041 & 0.007 & 0.045 & 0.007 & 0.043 & 0.008 & 0.047 \\
& 1.5 & -0.933 & 0.002 & 0.010 & 0.002 & 0.011 & 0.003 & 0.012 & 0.003 & 0.013 \\
& 2.0 & -0.911 & 0.002 & 0.004 & 0.002 & 0.005 & 0.002 & 0.006 & 0.003 & 0.006 \\
& 2.5 & -0.888 & 0.001 & 0.002 & 0.001 & 0.003 & 0.002 & 0.004 & 0.002 & 0.004 \\
& 3.0 & -0.866 & 0.001 & 0.002 & 0.002 & 0.003 & 0.002 & 0.004 & 0.003 & 0.005 \\
& 3.5 & -0.843 & 0.002 & 0.003 & 0.002 & 0.003 & 0.003 & 0.004 & 0.004 & 0.005 \\
& 4.0 & -0.821 & 0.002 & 0.003 & 0.004 & 0.004 & 0.004 & 0.005 & 0.006 & 0.007 \\
& 4.5 & -0.799 & 0.003 & 0.004 & 0.006 & 0.005 & 0.006 & 0.007 & 0.010 & 0.010 \\
& 5.0 & -0.776 & 0.004 & 0.005 & 0.009 & 0.007 & 0.008 & 0.009 & 0.013 & 0.013 \\
\bottomrule
\end{tabular}
\begin{tablenotes}[flushleft]\footnotesize\singlespacing
\item \textit{Notes:} Critical values are -20.5 for $M\!Z_{\alpha}^{OLS}$ and -17.3 for $M\!Z_{\alpha}^{GLS}.$
\end{tablenotes}
\end{threeparttable}
\end{sidewaystable}

 Table \ref{tab:mza-delta-asymptotic-size} shows empirical size results for $M\!Z_{\alpha}$ for brevity as all feasible tests display qualitatively similar results. The main pattern as $T$ increases is a transition from oversizing to undersizing at each fixed $\delta$, or undersizing to more extreme undersizing. Hence, a certain cell reaching near nominal size is more coincidental to the exact parameter specification rather than a well-calibrated test. In addition, $M\!Z_{\alpha}^{OLS}$ is generally more undersized (or less oversized) than $M\!Z_{\alpha}^{GLS}$ across all procedures. Also, the $MIC^{OLS}$ criteria oversize relative to their GLS counterparts, which is especially strong at small $\delta$ as a result of selecting significantly smaller $k$. Table~\ref{tab:mza-maicgls-summary-sqrt2T} provides a more focused explanation of these results for the $MAIC^{GLS}$ rule applied to $M\!Z_{\alpha}^{GLS}$ (although similar results hold for all specifications).

\begin{table}[t]
\centering
\normalsize\singlespacing
\begin{threeparttable}
\caption{Size breakdown for $M\!Z_{\alpha}^{GLS}$ with $MAIC^{GLS}$ under extended lag cap}
\label{tab:mza-maicgls-summary-sqrt2T}
\setlength{\tabcolsep}{8pt}
\begin{tabular}{c S[table-format=-1.3] S[table-format=1.1] S[table-format=1.3] S[table-format=1.3] S[table-format=1.3] S[table-format=2.1] S[table-format=1.3] S[table-format=2.2] S[table-format=1.3]}
\toprule
& & & \multicolumn{3}{c}{$MZ_\alpha^{GLS}$ size} & & & & \\
\cmidrule(lr){4-6}
{$T$} & {$\theta$} & {$\delta$} & {feasible} & {$R^{pop}_k\omega^2$} & {oracle} & {$k$} & {$k/\sqrt{T}$} & {$ s^2_{AR}/\omega^2$} & {$ H$} \\
\midrule
\multirow{4}{*}{100} & -0.950 & 0.5 & 0.194 & 0.129 & 0.000 & 10.0 & 1.000 & 16.80 & 1.152 \\
& -0.850 & 1.5 & 0.086 & 0.043 & 0.000 & 8.0 & 0.800 & 2.98 & 1.072 \\
& -0.750 & 2.5 & 0.048 & 0.018 & 0.000 & 6.0 & 0.600 & 1.93 & 1.001 \\
& -0.500 & 5.0 & 0.039 & 0.026 & 0.001 & 3.0 & 0.300 & 1.28 & 0.928 \\
\midrule
\multirow{4}{*}{250} & -0.968 & 0.5 & 0.063 & 0.032 & 0.000 & 21.0 & 1.328 & 11.29 & 1.187 \\
& -0.905 & 1.5 & 0.016 & 0.004 & 0.000 & 15.0 & 0.949 & 2.41 & 1.035 \\
& -0.842 & 2.5 & 0.008 & 0.001 & 0.000 & 11.0 & 0.696 & 1.77 & 0.996 \\
& -0.684 & 5.0 & 0.019 & 0.004 & 0.000 & 6.0 & 0.379 & 1.34 & 0.961 \\
\midrule
\multirow{4}{*}{500} & -0.978 & 0.5 & 0.024 & 0.010 & 0.000 & 32.0 & 1.431 & 10.25 & 1.134 \\
& -0.933 & 1.5 & 0.003 & 0.000 & 0.000 & 23.0 & 1.029 & 2.36 & 1.017 \\
& -0.888 & 2.5 & 0.002 & 0.000 & 0.000 & 16.0 & 0.716 & 1.77 & 0.988 \\
& -0.776 & 5.0 & 0.008 & 0.001 & 0.000 & 9.0 & 0.402 & 1.37 & 0.970 \\
\bottomrule
\end{tabular}
\begin{tablenotes}[flushleft]\footnotesize\singlespacing
\item \textit{Notes}: The feasible, \(R_k^{pop}\omega_T^2\), and oracle columns report rejection frequencies using \(s_{AR}^2\), the population AR(\(k\)) LRV evaluated at each replication’s selected lag, and the true LRV, respectively, in the same statistic with critical value \(-17.3\). The middle column retains population truncation distortion while removing LRV coefficient-estimation error. The columns \(k\), \(k/\sqrt T\), \(s_{AR}^2/\omega_T^2\), and \(H\) report medians across replications.
\end{tablenotes}
\end{threeparttable}
\end{table}

 Two important definitions are
 \begin{equation}
     H = \frac{s^2_{AR}/\omega_T^2}{R^{pop}_k}, \qquad R^{pop}_k = \frac{s^{2,pop}_{AR}}{\omega_T^2} = \frac{(1+(-\theta_T)^{k+1})(1+(-\theta_T)^{k+2})}{(1-(-\theta_T)^{k+1})(1-(-\theta_T)^{k+2})}, \label{eq:H}
 \end{equation}
 so that $R^{pop}_k$ is the inflation of the LRV estimate as a result of truncating the infinite AR process at lag $k$ even if the population AR coefficients are known\footnote{The expression for $R^{pop}_k$ follows from solving the population Yule--Walker equations for the AR($k$) best linear approximation to the stationary MA(1) process and evaluating its implied LRV. Note that under the lag rate in \ref{eq:rate_con}, $R^{pop}_k \to \coth^2(\kappa\delta/2).$}. Hence, $H$ (calculated per replication) measures the remaining finite sample distortion in $s^2_{AR}/\omega_T^2$ after accounting for this population truncation effect. Hence, $H=1$ implies that the realized LRV inflation is exactly that implied by the truncation at the selected lag. The oracle rejection rate is essentially zero throughout the table, confirming the strong conservatism of the statistic under conventional critical values. Population truncation offsets part of this conservatism. For example, at $T=100, \delta=0.5$, replacing the true LRV by $R^{pop}_k \omega_T^2$ raises rejection from near zero to 0.129, compared with feasible size of 0.194. The remaining difference between the latter two can be explained by estimation error beyond the population truncation benchmark. Similarly, the near nominal feasible size at $T=100, \delta=2.5$ combines an oracle size of near zero and a truncation-only size of 0.018, and an additional finite sample estimation effect. Thus, its proximity to 0.05 does not indicate a well-calibrated test.

 Median $s^2_{AR}/\omega_T^2$ is consistently over one and is particularly large at smaller $\delta$. Median $H$ is close to one, indicating that population AR truncation explains much of the LRV overestimation. This alone does not imply that truncation alone explains feasible size since rejection depends on the upper tail of $s^2_{AR}$, whereas $H$ reports a median. Indeed, the consistently higher feasible than truncation-only sizes illustrate the residual estimation variability is consequential.

 As $T$ increases, feasible size declines toward the conservative oracle behavior. For small and moderate $\delta$, this is accompanied by a slowing decline in LRV overestimation (at $\delta=5$ the overestimation ratio modestly rises). Overall, the results support that feasible size is a result of an interaction between intrinsic oracle conservatism and truncation induced LRV overestimation.

    The previous simulations hold $\delta$ fixed as $T$ increases, differing from the conventional practice of holding $\theta$ fixed. We repeated the size simulations, holding $\theta$ fixed at -0.9 and -0.8, for which $\delta_T = 0.1\sqrt{T}$ and $\delta_T = 0.2\sqrt{T}$, respectively. In this case, as $T$ increases, so does $\delta$ at each fixed $\theta$, so the DGP moves further away from the small $\delta$ NINW region as $T$ increases. We find that size exhibits a non-monotone pattern in this case: at $T=100$, the tests are oversized, and at $T=250,500,1000$, size is near 0, while by $T=10000$, size begins to increase towards the nominal level. This is consistent with the idea that at small sample sizes LRV overestimation dominates the intrinsic conservatism displayed by the statistics. In addition, large sample simulations were run with $\delta=0.5$ and $\delta=5$ for up to $T=20000$ and we find that feasible size decreases towards zero for all lag selection rules (albeit at different rates for each rule), consistent with Remark \ref{rem:feasible_interpretation}.


\subsection{Size-adjusted power results}\label{sec:power}
To assess the power of the tests, we set $c=-13.5$ for the local alternative. We now consider more specifically the trade-offs of using the extended cap or the conventional cap for $k$.

\begin{sidewaystable}[tp]
\centering
\normalsize\singlespacing
\begin{threeparttable}
\caption{Size-adjusted power of the $MZ_\alpha^{OLS}$ and $MZ_\alpha^{GLS}$ tests under extended lag cap}
\label{tab:mza-delta-sa-power-extended}
\setlength{\tabcolsep}{5.5pt}
\begin{tabular}{c S[table-format=1.1] S[table-format=-1.3] *{8}{S[table-format=1.3]}}
\toprule
& &
& \multicolumn{4}{c}{$MZ_\alpha^{OLS}$}
& \multicolumn{4}{c}{$MZ_\alpha^{GLS}$} \\
\cmidrule(lr){4-7}\cmidrule(lr){8-11}
{$T$} & {$\delta$} & {$\theta$}
& {$MAIC^{GLS}$}
& {$MAIC^{OLS}$}
& {$MBIC^{GLS}$}
& {$MBIC^{OLS}$}
& {$MAIC^{GLS}$}
& {$MAIC^{OLS}$}
& {$MBIC^{GLS}$}
& {$MBIC^{OLS}$}
\\
\midrule
\multirow{10}{*}{100}
& 0.5 & -0.950 & 0.105 & 0.123 & 0.104 & 0.119 & 0.112 & 0.111 & 0.111 & 0.110 \\
& 1.0 & -0.900 & 0.151 & 0.185 & 0.147 & 0.182 & 0.153 & 0.202 & 0.152 & 0.202 \\
& 1.5 & -0.850 & 0.174 & 0.259 & 0.179 & 0.254 & 0.174 & 0.279 & 0.179 & 0.281 \\
& 2.0 & -0.800 & 0.165 & 0.273 & 0.171 & 0.277 & 0.165 & 0.273 & 0.171 & 0.275 \\
& 2.5 & -0.750 & 0.171 & 0.254 & 0.174 & 0.256 & 0.172 & 0.254 & 0.175 & 0.255 \\
& 3.0 & -0.700 & 0.177 & 0.236 & 0.182 & 0.237 & 0.178 & 0.237 & 0.184 & 0.238 \\
& 3.5 & -0.650 & 0.178 & 0.226 & 0.183 & 0.229 & 0.181 & 0.228 & 0.186 & 0.233 \\
& 4.0 & -0.600 & 0.199 & 0.228 & 0.202 & 0.231 & 0.200 & 0.235 & 0.207 & 0.237 \\
& 4.5 & -0.550 & 0.207 & 0.231 & 0.212 & 0.235 & 0.212 & 0.235 & 0.221 & 0.240 \\
& 5.0 & -0.500 & 0.222 & 0.247 & 0.235 & 0.252 & 0.227 & 0.253 & 0.240 & 0.260 \\
\specialrule{1.05pt}{2pt}{2pt}
\multirow{10}{*}{250}
& 0.5 & -0.968 & 0.112 & 0.156 & 0.114 & 0.151 & 0.112 & 0.178 & 0.114 & 0.175 \\
& 1.0 & -0.937 & 0.110 & 0.270 & 0.112 & 0.269 & 0.110 & 0.265 & 0.112 & 0.264 \\
& 1.5 & -0.905 & 0.103 & 0.253 & 0.103 & 0.253 & 0.104 & 0.252 & 0.106 & 0.251 \\
& 2.0 & -0.874 & 0.112 & 0.218 & 0.109 & 0.220 & 0.113 & 0.214 & 0.115 & 0.215 \\
& 2.5 & -0.842 & 0.107 & 0.185 & 0.106 & 0.186 & 0.112 & 0.183 & 0.113 & 0.184 \\
& 3.0 & -0.810 & 0.124 & 0.184 & 0.127 & 0.183 & 0.132 & 0.183 & 0.136 & 0.181 \\
& 3.5 & -0.779 & 0.141 & 0.187 & 0.140 & 0.185 & 0.144 & 0.185 & 0.150 & 0.183 \\
& 4.0 & -0.747 & 0.162 & 0.198 & 0.166 & 0.203 & 0.168 & 0.199 & 0.178 & 0.204 \\
& 4.5 & -0.715 & 0.182 & 0.215 & 0.187 & 0.221 & 0.186 & 0.215 & 0.197 & 0.221 \\
& 5.0 & -0.684 & 0.197 & 0.224 & 0.198 & 0.227 & 0.207 & 0.229 & 0.216 & 0.233 \\
\specialrule{1.05pt}{2pt}{2pt}
\multirow{10}{*}{500}
& 0.5 & -0.978 & 0.088 & 0.209 & 0.088 & 0.206 & 0.088 & 0.217 & 0.088 & 0.215 \\
& 1.0 & -0.955 & 0.086 & 0.265 & 0.083 & 0.266 & 0.091 & 0.261 & 0.086 & 0.262 \\
& 1.5 & -0.933 & 0.080 & 0.196 & 0.067 & 0.195 & 0.086 & 0.191 & 0.078 & 0.190 \\
& 2.0 & -0.911 & 0.078 & 0.156 & 0.066 & 0.152 & 0.087 & 0.152 & 0.079 & 0.146 \\
& 2.5 & -0.888 & 0.094 & 0.151 & 0.081 & 0.145 & 0.102 & 0.149 & 0.098 & 0.142 \\
& 3.0 & -0.866 & 0.109 & 0.154 & 0.093 & 0.150 & 0.116 & 0.150 & 0.112 & 0.145 \\
& 3.5 & -0.843 & 0.134 & 0.172 & 0.119 & 0.168 & 0.144 & 0.172 & 0.141 & 0.165 \\
& 4.0 & -0.821 & 0.151 & 0.176 & 0.138 & 0.180 & 0.159 & 0.177 & 0.160 & 0.177 \\
& 4.5 & -0.799 & 0.177 & 0.199 & 0.160 & 0.202 & 0.187 & 0.201 & 0.188 & 0.203 \\
& 5.0 & -0.776 & 0.201 & 0.221 & 0.191 & 0.221 & 0.212 & 0.225 & 0.214 & 0.224 \\
\bottomrule
\end{tabular}
\begin{tablenotes}[flushleft]\footnotesize\singlespacing
\item \textit{Notes:} Size-adjusted power against the local alternative $\rho=1+c/T$ with $c=-13.5$ ($\rho=0.865, 0.946, 0.973$ for $T=100,\,250,\,500$, respectively). Lags are selected from $k\in[0,k_{max}]$ with $k_{max}=20, 31, 44$ for $T=100,\,250,\,500$, respectively.
\end{tablenotes}
\end{threeparttable}
\end{sidewaystable}
\begin{sidewaystable}[tp]
\centering
\normalsize\singlespacing
\begin{threeparttable}
\caption{Size-adjusted power of the $MZ_\alpha^{OLS}$ and $MZ_\alpha^{GLS}$ tests under conventional lag cap}
\label{tab:mza-delta-sa-power-conventional}
\setlength{\tabcolsep}{5.5pt}
\begin{tabular}{c S[table-format=1.1] S[table-format=-1.3] *{8}{S[table-format=1.3]}}
\toprule
& &
& \multicolumn{4}{c}{$MZ_\alpha^{OLS}$}
& \multicolumn{4}{c}{$MZ_\alpha^{GLS}$} \\
\cmidrule(lr){4-7}\cmidrule(lr){8-11}
{$T$} & {$\delta$} & {$\theta$}
& {$MAIC^{GLS}$}
& {$MAIC^{OLS}$}
& {$MBIC^{GLS}$}
& {$MBIC^{OLS}$}
& {$MAIC^{GLS}$}
& {$MAIC^{OLS}$}
& {$MBIC^{GLS}$}
& {$MBIC^{OLS}$}
\\
\midrule
\multirow{10}{*}{100}
& 0.5 & -0.950 & 0.120 & 0.115 & 0.119 & 0.113 & 0.119 & 0.107 & 0.118 & 0.105 \\
& 1.0 & -0.900 & 0.167 & 0.179 & 0.167 & 0.176 & 0.167 & 0.198 & 0.167 & 0.194 \\
& 1.5 & -0.850 & 0.182 & 0.252 & 0.184 & 0.248 & 0.182 & 0.277 & 0.185 & 0.274 \\
& 2.0 & -0.800 & 0.175 & 0.280 & 0.176 & 0.280 & 0.175 & 0.278 & 0.175 & 0.280 \\
& 2.5 & -0.750 & 0.175 & 0.262 & 0.179 & 0.264 & 0.175 & 0.262 & 0.179 & 0.263 \\
& 3.0 & -0.700 & 0.179 & 0.240 & 0.181 & 0.239 & 0.180 & 0.240 & 0.183 & 0.240 \\
& 3.5 & -0.650 & 0.184 & 0.232 & 0.186 & 0.234 & 0.188 & 0.234 & 0.189 & 0.238 \\
& 4.0 & -0.600 & 0.200 & 0.234 & 0.200 & 0.235 & 0.200 & 0.239 & 0.203 & 0.241 \\
& 4.5 & -0.550 & 0.212 & 0.237 & 0.213 & 0.240 & 0.216 & 0.244 & 0.219 & 0.247 \\
& 5.0 & -0.500 & 0.228 & 0.254 & 0.232 & 0.256 & 0.228 & 0.258 & 0.236 & 0.262 \\
\specialrule{1.05pt}{2pt}{2pt}
\multirow{10}{*}{250}
& 0.5 & -0.968 & 0.131 & 0.138 & 0.130 & 0.136 & 0.130 & 0.161 & 0.130 & 0.160 \\
& 1.0 & -0.937 & 0.145 & 0.259 & 0.145 & 0.254 & 0.144 & 0.255 & 0.145 & 0.252 \\
& 1.5 & -0.905 & 0.144 & 0.268 & 0.142 & 0.271 & 0.144 & 0.267 & 0.142 & 0.268 \\
& 2.0 & -0.874 & 0.161 & 0.240 & 0.154 & 0.240 & 0.161 & 0.236 & 0.159 & 0.237 \\
& 2.5 & -0.842 & 0.151 & 0.208 & 0.145 & 0.211 & 0.156 & 0.208 & 0.150 & 0.206 \\
& 3.0 & -0.810 & 0.161 & 0.205 & 0.151 & 0.200 & 0.167 & 0.204 & 0.161 & 0.199 \\
& 3.5 & -0.779 & 0.165 & 0.200 & 0.156 & 0.198 & 0.169 & 0.200 & 0.166 & 0.194 \\
& 4.0 & -0.747 & 0.186 & 0.218 & 0.177 & 0.213 & 0.193 & 0.220 & 0.191 & 0.216 \\
& 4.5 & -0.715 & 0.200 & 0.227 & 0.198 & 0.232 & 0.203 & 0.226 & 0.207 & 0.229 \\
& 5.0 & -0.684 & 0.211 & 0.234 & 0.209 & 0.237 & 0.220 & 0.239 & 0.224 & 0.241 \\
\specialrule{1.05pt}{2pt}{2pt}
\multirow{10}{*}{500}
& 0.5 & -0.978 & 0.117 & 0.171 & 0.118 & 0.168 & 0.117 & 0.186 & 0.118 & 0.184 \\
& 1.0 & -0.955 & 0.144 & 0.264 & 0.143 & 0.265 & 0.143 & 0.260 & 0.143 & 0.262 \\
& 1.5 & -0.933 & 0.166 & 0.227 & 0.163 & 0.229 & 0.165 & 0.220 & 0.165 & 0.222 \\
& 2.0 & -0.911 & 0.181 & 0.212 & 0.160 & 0.210 & 0.181 & 0.209 & 0.180 & 0.208 \\
& 2.5 & -0.888 & 0.186 & 0.208 & 0.155 & 0.203 & 0.188 & 0.204 & 0.177 & 0.200 \\
& 3.0 & -0.866 & 0.185 & 0.206 & 0.145 & 0.196 & 0.186 & 0.201 & 0.173 & 0.191 \\
& 3.5 & -0.843 & 0.196 & 0.218 & 0.159 & 0.207 & 0.204 & 0.217 & 0.184 & 0.201 \\
& 4.0 & -0.821 & 0.198 & 0.218 & 0.166 & 0.206 & 0.202 & 0.215 & 0.189 & 0.205 \\
& 4.5 & -0.799 & 0.210 & 0.230 & 0.182 & 0.218 & 0.225 & 0.234 & 0.212 & 0.222 \\
& 5.0 & -0.776 & 0.231 & 0.244 & 0.204 & 0.234 & 0.238 & 0.246 & 0.227 & 0.238 \\
\bottomrule
\end{tabular}
\begin{tablenotes}[flushleft]\footnotesize\singlespacing
\item \textit{Notes:} Size-adjusted power against the local alternative $\rho=1+c/T$ with $c=-13.5$ ($\rho=0.865, 0.946, 0.973$ for $T=100,\,250,\,500$, respectively). Lags are selected from $k\in[0,k_{max}]$ with $k_{max}=12, 15, 17$ for $T=100,\,250,\,500$, respectively.
\end{tablenotes}
\end{threeparttable}
\end{sidewaystable}

Tables~\ref{tab:mza-delta-sa-power-extended} and \ref{tab:mza-delta-sa-power-conventional} report size-adjusted power under the extended cap and conventional cap for $M\!Z_{\alpha}$, respectively\footnote{Note that $M\!S\!B^{OLS}$ produced similar power results to that of $M\!Z_{\alpha}$ in spite of its favorable oracle power function, demonstrating that oracle properties need not pass on to the feasible tests.}. A clear result of the extended cap is that power frequently deteriorates as the sample size increases. For $\delta=1.5,\dots,4.5$, power at $T=500$ is below its $T=100$ value for every procedure, and all but two of these procedures have power decrease monotonically. Since each cell is size-adjusted by construction, this decline reflects weaker discriminatory ability between the null and local alternative, not increasing conservatism under conventional critical values. On the other hand, the conventional cap generally permits greater size-adjusted power but weakly and not uniformly. In addition, the systematic decline in power as the sample size increases is less systematic: many procedures lose power from $T=100$ to $T=250$, but recover by $T=500$, although continued declines and other patterns are also present.

For a fixed lag selection rule, $M\!Z_{\alpha}^{GLS}$ has a slight power advantage over $M\!Z_{\alpha}^{OLS}$, but this ranking is not uniformly true. By contrast, OLS lag selection rules generally produce greater size-adjusted power than their GLS counterparts under both caps. This advantage is particularly apparent under the extended cap, but remains pertinent under the conventional cap as well. Differences between MAIC and MBIC are usually secondary to those associated with the cap and GLS or OLS construction of the information criterion though not entirely negligible.

A related pattern appears in Table~VI.B of \citet{NgPerron2001}. For $M\!Z_{\alpha}^{GLS}$ and $MAIC^{GLS}$, at $\theta=-0.8$, the table shows size-adjusted power declining as $T$ increases. Their simulations were fixed-$\theta$ while ours are fixed-$\delta$, so they are not directly comparable, but their results show that decreasing size-adjusted power with sample size is not unique to the simulation design present here.

We also ran simulations which replaced the Gaussian innovation distribution by standardized $t_5$ and Rademacher innovations. Rademacher innovations resulted in a substantial decrease in size-adjusted power while $t_5$ innovations slightly increased size-adjusted power in general. These results support the importance of the innovation distribution dependence observed in Remark \ref{rem:innovation_dep}.

Another complicating issue not shown in these simulations, is the problem of power reversal. \citet{PQ2007} recommend the OLS construction of the modified information criterion instead of GLS as a remedy, and our results show it to generally be favorable in terms of size adjusted power. However, the OLS criteria drastically oversize at low $\delta$ and severely undersize elsewhere, producing a severe size-power tradeoff.

\FloatBarrier

\section{Conclusion}\label{sec:conclusion}
We illustrate the limitations of the $M$-tests under nearly integrated nearly white noise and an unknown linear time trend. The tests' limiting distributions are contaminated by noise terms, with the $M^{GLS}$ tests being more so, causing dependence on the innovation distribution. When the LRV is known, conventional critical values result in a region of zero asymptotic rejection probability for sufficiently small $\delta$ and several tests exhibit substantial loss of asymptotic power. Yet, the Gaussian power envelope itself remains identical to the standard envelope in \citet{ERS1996}. Hence, power losses reflect inefficiency of the tests in utilizing the information available rather than a deterioration of the information itself. The oracle $M\!S\!B^{OLS}$ test's power function provides an exception.

Feasible implementation introduces a separate distortion. For deterministic lag orders satisfying $k/\sqrt{T} \to \kappa$ $(0 < \kappa < \infty)$, the autoregressive LRV estimator converges, relative to the true LRV, to $\coth^2(\kappa\delta/2)$. Thus, oracle and feasible limiting distribution equivalence does not hold under the lag rate which keeps the $M$-tests bounded. The feasible tests face two forces which pull in different directions: LRV overestimation which encourages rejection and oracle undersizing. This implies, near nominal size, when obtained, is a result of these two forces balancing rather than a well calibrated test.

The finite sample simulations under data-dependent lag selection rules shows the practical consequences of this interaction. The choice of lag cap and detrending for the modified information criterion result in vast differences in size and size-adjusted power, and limitations exist for all that are not resolved as the sample size increases. Among the procedures examined, no single one is uniformly satisfactory. These results motivate further research into procedures designed explicitly to handle a vanishing LRV.


\newpage