The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
109,893 characters
Inference in a Stationary/Nonstationary Autoregressive Time-Varying-Parameter Model
\global\long
\global\long \makeatletter \renewenvironment{proof}[1][\proofname]{\par\pushQED{\ \qedsymbol}\normalfont\topsep6\p@\@plus6\p@\relax\trivlist\item[\hskip\labelsep\bfseries#1\@addpunct{.}]\ignorespaces}{\popQED\endtrivlist\@endpefalse}
\makeatother
\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\def\mc#1{\mathscr{#1}}
\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\def\abs#1{\left|#1\right|}
\global\long\def\norm#1{\left\Vert #1\right\Vert }
\global\long\def\rest#1{\left.#1\right|}
\global\long\def\bracket#1#2{\left\langle #1\middle\vert#2\right\rangle }
\global\long\def\sandvich#1#2#3{\left\langle #1\middle\vert#2\middle\vert#3\right\rangle }
\global\long\def\turd#1{\frac{#1}{3}}
\global\long\global\long\def\sand#1{\left\lceil #1\right\vert }
\global\long\def\wich#1{\left\vert #1\right\rfloor }
\global\long\def\sandwich#1#2#3{\left\lceil #1\middle\vert#2\middle\vert#3\right\rfloor }
\global\long\def\abs#1{\left|#1\right|}
\global\long\def\norm#1{\left\Vert #1\right\Vert }
\global\long\def\rest#1{\left.#1\right|}
\global\long\def\inprod#1{\left\langle #1\right\rangle }
\global\long\def\ol#1{\overline{#1}}
\global\long\def\ul#1{\underline{#1}}
\global\long\def\td#1{\tilde{#1}}
\global\long\def\bs#1{\boldsymbol{#1}}
\global\long\global\long\global\long\global\long\global\long\global\long\global\long\begin{titlepage}
\clearpage\maketitle \thispagestyle{empty}
\begin{abstract}
This paper considers nonparametric estimation and inference in first-order
autoregressive (AR(1)) models with deterministically time-varying
parameters. A key feature of the proposed approach is to allow for
time-varying stationarity in some time periods, time-varying nonstationarity
(i.e., unit root or local-to-unit root behavior) in other periods,
and smooth transitions between the two. The estimation of the AR parameter
at any time point is based on a local least squares regression method,
where the relevant initial condition is endogenous. We obtain limit
distributions for the AR parameter estimator and t-statistic at a
given point $\tau$ in time when the parameter exhibits unit root,
local-to-unity, or stationary/stationary-like behavior at time $\tau.$
These results are used to construct confidence intervals and median-unbiased
interval estimators for the AR parameter at any specified point in
time. The confidence intervals have correct asymptotic coverage probabilities
with the coverage holding uniformly over stationary and nonstationary
behavior of the observations.
\bigskip{}
\noindent\textbf{Keywords:} Autoregressive time-varying-parameter
model, endogenous initial condition, nonparametric estimation, confidence
interval.
\end{abstract}
\end{titlepage}
\section{\protect\label{sec:MT-Intro}Introduction}
Autoregressive models --- stationary or nonstationary --- are workhorse
models in econometric time series. In this paper, we consider a deterministically
time-varying parameter (TVP) autoregressive model that allows for
stationary and non-stationary behavior at different points in the
time period of interest. Thus, the level of persistence of the time
series can change over the time period. The motivation for considering
such a model is that the economy is in continual transition due to
technological, institutional, political, and demographic changes.
The model considered allows for the intercept and error variance of
the model also to be time-varying, not just the AR parameter.
The estimation method employed is local least squares, which depends
on a tuning parameter $h$ that determines the local neighborhood
that is considered. We construct a confidence interval (CI) for the
value of the AR parameter at a time point $\tau$ via the inversion
of tests, as is common in the literature for constant parameter AR
models, e.g., see \citet*{stock1991confidence,andrews1993exactly,hansen1999grid,mikusheva2007uniform},
and \citet{andrews2014conditional}. We show that the CI's have correct
uniform asymptotic coverage probability for a parameter space that
allows the time series to be stationary in parts of the time period
and nonstationary in other parts, using the approach in \citet*{andrews2020generic}.
We also construct asymptotically median-unbiased interval estimators
(MUE's) in an analogous fashion.
In the TVP case, the initial condition is endogenous due to the choice
of the local neighborhood and depends on the potentially different
behavior of the time series prior to the local neighborhood. For a
given time point $\tau$ of interest, we find that the asymptotic
distributions of the LS estimator and t-statistic depend on the endogenous
initial condition in the local-to-unity case. This is analogous to
the asymptotic effect of the initial condition--under certain assumptions
on the initial condition--in constant parameter AR models, see \citet*{elliott1999efficient,elliott2001confidence},
and \citet{muller2003tests}. On the other hand, the endogenous initial
condition does not affect the asymptotic distributions when the AR
coefficient at time $\tau$ is more distant from one than local-to-unity.
We note that whether a time series is local-to-unity at time $\tau,$
or not, depends on $\tau$ and the chosen bandwidth $h.$ It is important
to provide asymptotic results that hold uniformly over a parameter
space that does not depend on $h,$ which we do.
We introduce a method for determining the tuning parameter $h$ based
on a forecast-error criterion. We provide conditions under which this
data-dependent choice of $h$ is asymptotically equivalent to an infeasible
choice that minimizes the unobserved ``empirical loss.'' These results
are similar to results for i.i.d. models given in \citet{li1987asymptotic}
and \citet{andrews1991asymptotic}.
We provide Monte Carlo simulation results for the methods introduced
in the paper. We consider true autoregressive functions whose shapes
are sinusoidal, linear, partly flat/partly linear, flat, and kinked
linear. We consider cases where the functions are close, or equal,
to one in some regions, but different from one by varying amounts
in other regions. We find that the proposed CI has reasonably good
coverage probabilities and short average lengths for most of the data-generating
processes considered. For example, nominal 95\% CI's are found to
have finite-sample coverage probabilities ranging from 92.5\% to 96\%
in 88.8\% of the cases, across 205 cases. The lowest coverage probabilities,
in the range from 87.5\% to 90\%, occur only in 2.0\% of the cases.
The MUE's are found to have very small finite-sample median bias across
the different cases considered. The magnitude of the data-dependent
choice of $h$ varies widely depending upon the shape of the autoregressive
function, as desired.
We provide some empirical applications of the methods to monthly inflation
and real exchange rates in several countries using data from the IMF
International Financial Statistics database. We find that the inflation
series exhibit noticeable time variation of the AR parameter across
the time period considered. We discuss how this relates to the literature
on the persistence properties of inflation. On the other hand, we
find that the real exchange rate series have nearly constant AR values
across time that are equal to, or close to, one. Hence, the TVP methods
are capable of producing constant AR values when it is appropriate
to do so. We discuss how these results relate to the literature on
the persistence of exchange rates. In Section \ref{sec:addl-Empirical-Results}
of the Supplemental Material, we also report results for interest
rates for several countries and results for a number of US macroeconomic
time series using the Federal Reserve Economic Database (FRED).
The results of the paper apply to a TVP-AR(1) model. For some time
series, a TVP-AR(p) model with $p>1$ may be more appropriate than
a TVP-AR(1) model. The methods introduced in the paper can be extended
to a TVP-AR(p) model, see Section \ref{sec:Extension-to-TVP-AR(p)}
of the Supplemental Material for details.
Relative to the literature, the contribution of this paper is to develop
methods for a deterministically time-varying parameter (TVP) AR(1)
model that allows time-varying stationarity in some time periods and
time-varying nonstationarity in other periods. The resulting model
is much more flexible than a constant parameter model. No paper in
the existing literature does this. The methods we employ are quite
similar to those employed in constant parameter AR(1) models that
impose stationarity or nonstationarity across the whole time period,
such as those referenced above. In particular, we invert tests of
null hypotheses concerning the AR coefficient (at a particular point
in time) and utilize the nonstandard asymptotic distributions of the
t-statistics for such null hypotheses to obtain critical values. Our
results differ from those obtained for constant parameter AR(1) models
in that (i) we consider a local neighborhood of the time point of
interest $\tau,$ indexed by a bandwidth parameter $h,$ and within
that time period the AR coefficient changes with time, which causes
biases that have to be accounted for, (ii) the initial condition of
the local neighborhood depends on the choice of the local neighborhood
and on the past behavior of the TVP-AR(1) process, which may differ
from its current behavior, which needs to be accounted for, and (iii)
we use the data to select a suitable local neighborhood using a forecast-error
criterion function. None of these features arise in a constant parameter
model.
The literature contains numerous papers that consider time-varying
parameter AR models. Some of these papers consider deterministic TVP's,
as in this paper, but they do not allow for stationarity in some time
periods and nonstationarity in others. References for stationary (i.e.,
short-range dependent) TVP AR models include \citet*{Rao1970,Grenier1983,dahlhaus1996kullback,Dahlhaus1997},
\citet{Dahlhaus1998}, \citet*{dahlhaus1999nonlinear}, \citet*{Moulines2005AoS},
\citet*{PhillipsXu2008Adaptive}, \citet*{ding2017ejs}, \citet{vanDelft2018EJS},
and \citet*{karmakar2022simultaneous}. References for nonstationary
(i.e., long-range dependent) TVP AR models include \citet*{bykhovskaya2018boundary,bykhovskaya2020point},
which focus on tests of a unit root null hypothesis against functional
local-to-unity alternatives, as opposed to CI's for an AR parameter,
which is the focus of this paper. No papers in the literature consider
CI's for an AR parameter in nonstationary TVP AR models.
The literature also includes papers on random coefficient (RC) AR
models and functional coefficient (FC) AR models (in which the AR
coefficient depends on observable variables). References for stationary
RC AR models includes \citet{NichollsQuinn1980}, \citet{NichollsQuinn1981},
\citet*{doan1984forecasting}, and \citet{cogley2005drifts}. References
for papers on RC AR models with random coefficients that follow a
nonstationary process include \citet{follmer1993microeconomic}, \citet*{GIRAITISKapetaniosYates2014rckernel,GiraitisKapetaniosYates2018},
and \citet*{tao2019random}. A reference for stationary FC AR models
is \citet*{CaiFanYao2000}. References for nonstationary FC AR models
includes \citet{Juhl2005}, \citet{Lieberman2012}, and \citet*{lieberman2014norming,lieberman2017multivariate,lieberman2018iv}.
This paper is organized as follows. Section \ref{sec:MT-Model=000020Setup}
introduces the TVP-AR(1) model. Section \ref{sec:MT-CI=000020AR=000020Par}
introduces the CI and MUE for the AR parameter at time $\tau.$ Section
\ref{sec:MT-Choice_of_h} introduces the data-dependent method for
choosing the bandwidth parameter $h$ based on a forecast-error criterion.
Section \ref{sec:MT-Monte-Carlo-Simulations} presents the Monte Carlo
simulation results. Section \ref{sec:MT-Empirical-Applications} presents
the empirical results. Section \ref{sec:MT-Asymptotics} shows that
the CI has correct uniform asymptotic coverage probability, the MUE
is asymptotically median unbiased, and the data-dependent method for
choosing the bandwidth parameter $h$ has some some desirable asymptotic
properties. It also provides the asymptotic behavior of the local
least squares estimator and t-statistic under a variety of drifting
sequences of distributions, which are used in the proof of the uniform
asymptotic coverage probability results and the asymptotic median
unbiasedness results. The Supplemental Material includes the critical
values of the limiting distribution of our t-statistic in Section
\ref{sec:SM-Critical-Values-J_psi}, the proofs of the results of
the paper in Section \ref{sec:Theory}, additional simulation results
in Section \ref{sec:Additional-Simulation-Results}, a description
of how to extend the methods to TVP-AR(p) models for $p>1$ in Section
\ref{sec:Extension-to-TVP-AR(p)}, and information about the empirical
applications and additional empirical results in Section \ref{sec:addl-Empirical-Results}.
All limits in this paper are as $n\rightarrow\infty.$ For notational
simplicity, but with some abuse of notation, we let $\cdot/nh$ denote
$\cdot/(nh)$ throughout the paper.
\section{\protect\label{sec:MT-Model=000020Setup}Model}
\numberwithin{equation}{section}
The TVP-AR(1) model we consider is
\begin{align}
Y_{t} & =\mu_{t}+Y_{t}^{*}\ \text{and}\nonumber \\
Y_{t}^{*} & =\rho_{t}Y_{t-1}^{*}+\sigma_{t}U_{t},\ \text{for\ }t=1,...,n,\label{eq:MT-tvp-model}
\end{align}
where $\rho_{t}\in\left[-1+\varepsilon_{1},1\right]$ for some $0<\varepsilon_{1}<2$.
The autoregressive parameter $\rho_{t}$ is allowed to vary with time
$t.$ A key feature of the model is that it allows for stationary,
unit root, or local-to-unity behavior at different points in time.
The errors $\left\{ U_{t}:t=0,1,...,n\right\} $ are a stationary
martingale difference sequence under $F$ with $E_{F}\left(\rest{U_{t}}\mathscr{G}_{t-1}\right)=0\ \text{a.s.}$,
$E_{F}\left(\rest{U_{t}^{2}}\mathscr{G}_{t-1}\right)=1\ \text{a.s.}$,
and $E_{F}\left(\rest{U_{t}^{4}}\mathscr{G}_{t-1}\right)<M\ \text{a.s.}$
for some $M\in\left(0,\infty\right)$, where $\mathscr{G}_{t}$ is
some non-decreasing sequence of $\sigma$-fields for which $\sigma\left(U_{0},...,U_{t},Y_{0}^{*}\right)\subseteq\mathscr{G}_{t}$
for $t=1,...,n$.
We assume $\rho_{t},$ $\mu_{t},$ and $\sigma_{t}^{2}$ satisfy
\begin{equation}
\rho_{t}:=\rho\left(t/n\right),\ \mu_{t}:=\mu\left(t/n\right),\ \text{and }\sigma_{t}^{2}:=\sigma^{2}\left(t/n\right),\label{eq:MT-rhomusigma}
\end{equation}
respectively, where $\rho\left(\cdot\right)\ $is a twice continuously
differentiable function on $\left[0,1\right]$ and $\mu\left(\cdot\right)\ $and
$\sigma^{2}\left(\cdot\right)$ are Lipschitz functions on $\left[0,1\right]$.
Given \eqref{eq:MT-rhomusigma}, $Y_{t}$, $Y_{t}^{*}$, $\rho_{t}$,
$\mu_{t}$, and $\sigma_{t}$ depend implicitly on $n$.
Let $\tau\in\left(0,1\right)$. We consider estimation and inference
concerning
\begin{equation}
\rho(\tau),\label{eq:MT-paramofinterest}
\end{equation}
which is the value of the autoregressive function $\rho(\cdot)$ at
the $\tau$ fraction of the way through the sample.
For ease of reading, the definition of the parameter space of functions
$\rho\left(\cdot\right)$, $\mu\left(\cdot\right),\ $and $\sigma^{2}\left(\cdot\right)$
that is considered, which includes some structure on the $\rho\left(\cdot\right)$
and $\mu\left(\cdot\right)\ $functions, is given in Section \ref{sec:MT-Parameter=000020Space}
below.
\section{Confidence Interval for the Autoregressive \protect \protect \\
Parameter \texorpdfstring{$\boldsymbol{\rho(\tau)}$}{rho(tau)}
\protect\label{sec:MT-CI=000020AR=000020Par}}
\subsection{Local Least Squares Estimator of \texorpdfstring{$\boldsymbol{\rho(\tau)}$}{rho(tau)}}
We employ a LS estimator of $\rho(\tau)$ based on the time periods
$t=T_{1},...,T_{2}$, where
\begin{equation}
T_{1}=\lfloor n\tau\rfloor-\lfloor nh/2\rfloor\ \text{and\ }T_{2}=\lfloor n\tau\rfloor+\lfloor nh/2\rfloor,\label{eq:MT-timeperiod}
\end{equation}
for a bandwidth parameter $h$. The total number of periods in $\left[T_{1},T_{2}\right]$
is within one of $nh$. For this range of time periods, the initial
condition time period is $T_{0}:=T_{1}-1.$
For the asymptotic results given below the bandwidth $h$ satisfies
$h\rightarrow0$ and $nh\rightarrow\infty$ as $n\rightarrow\infty$. In Section \ref{sec:MT-Choice_of_h}
below, we introduce a data-dependent bandwidth that is smaller or
larger depending on how wiggly or flat the true $\rho\left(\cdot\right)\ $
and $\mu\left(\cdot\right)\ $ functions are.
Define
\begin{equation}
\ol Y_{nh}:=\dfrac{1}{nh}\sum_{t=T_{1}}^{T_{2}}Y_{t}\ \text{and}\ \ol Y_{nh,-1}:=\dfrac{1}{nh}\sum_{t=T_{1}}^{T_{2}}Y_{t-1}.\label{eq:MT-Y_dagger}
\end{equation}
To estimate $\rho\left(\tau\right)$, we regress $Y_{t}$ on a constant
and $Y_{t-1}.$ The resulting local LS estimator $\widehat{\rho}_{n\tau}$
is
\begin{equation}
\widehat{\rho}_{n\tau}=\dfrac{\sum_{t=T_{1}}^{T_{2}}\left(Y_{t-1}-\ol Y_{nh,-1}\right)\left(Y_{t}-\ol Y_{nh}\right)}{\sum_{t=T_{1}}^{T_{2}}\left(Y_{t-1}-\ol Y_{nh,-1}\right)^{2}}.\label{eq:MT-rho_hat}
\end{equation}
\subsection{Confidence Interval for \texorpdfstring{$\boldsymbol{\rho(\tau)}$}{rho(tau)}
\protect\label{subsec:MT-CI=000020AR=000020Par} }
The CI for $\rho\left(\tau\right)$ that we consider is obtained by
inverting tests of null hypotheses of the form $H_{0}:\rho\left(\tau\right)=\rho_{0}$
for different values $\rho_{0}\in\left[-1+\varepsilon_{1},1\right]$.
The estimator of the time-varying variance $\sigma^{2}\left(\cdot\right)$
at $t/n=\tau$ is defined to be
\begin{equation}
\widehat{\sigma}_{n\tau}^{2}\coloneqq\left(nh\right)^{-1}\sum_{t=T_{1}}^{T_{2}}\left[Y_{t}-\ol Y_{nh}-\widehat{\rho}_{n\tau}\left(Y_{t-1}-\ol Y_{nh,-1}\right)\right]^{2}.\label{eq:MT-sigma=000020ntau=000020def}
\end{equation}
For arbitrary $\rho_{0}\in(-1,1],$ the t-statistic that is used to
construct the CI for $\rho(\tau)$ is
\begin{equation}
T_{n}\left(\rho_{0}\right)\coloneqq\dfrac{\left(nh\right)^{1/2}\left(\widehat{\rho}_{n\tau}-\rho_{0}\right)}{\widehat{s}_{n\tau}},\ \text{where }\widehat{s}_{n\tau}^{2}\coloneqq\widehat{\sigma}_{n\tau}^{2}/\left(nh\right)^{-1}\sum_{t=T_{1}}^{T_{2}}\left(Y_{t-1}-\ol Y_{nh,-1}\right)^{2}.\label{eq:MT-t-stat=000020def}
\end{equation}
Let $B(\cdot)$ denote a standard Brownian motion on $[0,1].$ Let $Z_{1}$
be a standard normal random variable that is independent of $B\left(\cdot\right)$.
Define
\begin{align}
I_{\psi}\left(s\right):= & \int_{0}^{s}\exp\left\{ -\left(s-r\right)\psi\right\} dB\left(r\right),\nonumber \\
I_{\psi}^{*}\left(s\right)\coloneqq & \begin{cases}
I_{\psi}\left(s\right)+\dfrac{1}{\sqrt{2\psi}}\exp\left(-\psi s\right)Z_{1} & \text{for }\psi>0\\
B\left(s\right) & \text{for }\psi=0,\text{ and}
\end{cases}\nonumber \\
I_{D,\psi}^{*}\left(s\right):= & I_{\psi}^{*}\left(s\right)-\int_{0}^{1}I_{\psi}^{*}\left(r\right)dr.\label{eq:MT-I_D,tau(r)=000020def}
\end{align}
The stochastic process $I_{\psi}\left(s\right)$ is an Ornstein-Uhlenbeck
process on $\left[0,1\right]$ with parameter $\psi.$
Below we consider sequences of functions $\left\{ \rho_{n}\left(\cdot\right),\mu_{n}\left(\cdot\right),\sigma_{n}\left(\cdot\right)\right\} _{n\geq1}$
and null hypotheses $H_{0}:\rho_{n}\left(\tau\right)=\rho_{0,n},$
where the null hypotheses values $\rho_{0,n}$ depend on $n$ for
$n\geq1.$ For suitable sequences $\left\{ \rho_{0,n}\right\} _{n\geq1}$,
we show that under $H_{0},$
\begin{equation}
T_{n}\left(\rho_{0,n}\right)\rightarrow_{d} J_{\psi}\ \text{for}\ \psi\in\left[0,\infty\right],\label{eq:MT-Asy=000020distn=000020T=000020stat}
\end{equation}
where $T_{n}\left(\rho_{0,n}\right)$ is defined with $\rho_{0,n}$
in place of $\rho_{0}$ in \eqref{eq:MT-t-stat=000020def}, $\psi$
depends on the sequence $\left\{ \rho_{0,n}\right\} _{n\geq1},$ and
$J_{\psi}$ is defined as follows. For $\psi=\infty$, which corresponds
to a ``stationary'' sequence $\left\{ \rho_{0,n}\right\} _{n\geq1}$,
$J_{\psi}$ has a $N(0,1)$ distribution. For $\psi\in\left[0,\infty\right)$,
which corresponds to a local-to-unity or unit root sequence $\left\{ \rho_{0,n}\right\} _{n\geq1}$,
\begin{equation}
J_{\psi}:=\left(\int_{0}^{1}I_{D,\psi}^{*2}\left(s\right)ds\right)^{-1/2}\int_{0}^{1}I_{D,\psi}^{*}\left(s\right)dB\left(s\right).\label{eq:MT-Distn=000020J_psi}
\end{equation}
For $\alpha\in(0,1),$ let $c_{\psi}\left(\alpha\right)$ denote the
$\alpha$ quantile of the distribution of $J_{\psi}$. For given $\alpha$,
we compute $c_{\psi}\left(\alpha\right)$, the $\alpha$-quantile
of $J_{\psi}$ in \eqref{eq:MT-Distn=000020J_psi}, by simulating
the asymptotic distribution $J_{\psi}$. To do so, $B=300,000$ independent
constant coefficient AR(1) sequences are generated with innovations
$U_{t}\sim_{iid}N\left(0,1\right)$, stationary start-up, $n=25,000$,
and $\rho=\exp\left(-\psi/n\right)$. For each sequence, the test
statistic $T_{n}\left(\rho\right)$ defined in \eqref{eq:MT-t-stat=000020def}
is calculated. Then the simulated estimate of $c_{\psi}\left(\alpha\right)$
is the $\alpha$-quantile of the empirical distribution of the $B=300,000$
realizations of the test statistic $T_{n}\left(\rho\right)$.
The nominal $1-\alpha$ equal-tailed two-sided CI for $\rho(\tau)$
is
\begin{align}
CI_{n,\tau} & :=\left\{ \rho_{0}\in\left[-1+\varepsilon_{1},1\right]:c_{\psi_{nh,\rho_{0}}}\left(\alpha/2\right)\leq T_{n}(\rho_{0})\leq c_{\psi_{nh,\rho_{0}}}\left(1-\alpha/2\right)\right\} ,\ \text{where}\ \nonumber \\
\psi_{nh,\rho_{0}} & :=-nh\ln\left(\rho_{0}\right)\ \text{for }\rho_{0}>0\ \text{and }\psi_{nh,\rho_{0}}:=\infty\ \text{for }\rho_{0}\leq0.\label{eq:MT-CI=000020defn}
\end{align}
The CI $CI_{n,\tau}$ can be computed by taking a fine grid of values
$\rho_{0}\in\left[-1+\varepsilon_{1},1\right]$ and comparing $T_{n}\left(\rho_{0}\right)$
to $c_{\psi_{nh,\rho_{0}}}\left(\alpha/2\right)$ and $c_{\psi_{nh,\rho_{0}}}\left(1-\alpha/2\right).$
Table \ref{tab:SM-Critical-Values_psi} provides the critical values
$c_{\psi_{nh,\rho_{0}}}\left(\alpha/2\right)$ and $c_{\psi_{nh,\rho_{0}}}\left(1-\alpha/2\right)$
for $\alpha=.05$, and .1. Given these critical values, computation of
equal-tailed two-sided 90\% and 95\% CI's for $\rho(\tau)$ is fast.
The correct asymptotic size and asymptotic similarity of the CI $CI_{n,\tau}$
are established in Theorem \ref{thm:MT-Cor=000020Asy=000020Size}
in Section \ref{sec:MT-Asymptotics} below.
\subsection{Median-Unbiased Interval Estimator of \texorpdfstring{$\boldsymbol{\rho(\tau)}$}{rho(tau)}\protect\label{subsec:MT-Median-Unbiased-Interval-Est}}
By definition, an estimator $\widehat{\theta}_{n}$ of a parameter
$\theta$ is median unbiased if $P\left(\widehat{\theta}_{n}\geq\theta\right)\geq1/2$
and $P\left(\widehat{\theta}_{n}\leq\theta\right)\geq1/2.$ Here we
introduce a median-unbiased interval estimator of $\rho\left(\tau\right)$
that satisfies an analogous condition. Also, with probability close
to one, the estimator is a single point.\footnote{If the .5 quantile of the asymptotic null distribution of the t-statistic
was strictly decreasing in $\rho_{0},$ or equivalently, $c_{\psi}\left(.5\right)$
was strictly increasing in $\psi$, then the proposed interval estimator
would be a point estimator with probability one. Because this condition
fails to hold exactly, but almost holds, there is a very small probability
that the estimator is a short interval, rather than a point.} Let $CI_{n,\tau}^{up}\left(.5\right)$ and $CI_{n,\tau}^{low}\left(.5\right)$
denote level $.5$ one-sided upper-bound and lower-bound CIs for $\rho\left(\tau\right)$,
respectively. By definition,
\begin{align}
CI_{n,\tau}^{up}\left(.5\right) & :=\left\{ \rho_{0}\in\left[-1+\varepsilon_{1},1\right]:c_{\psi_{nh,\rho_{0}}}\left(.5\right)\leq T_{n}\left(\rho_{0}\right)\right\} \ \text{and}\nonumber \\
CI_{n,\tau}^{low}\left(.5\right) & :=\left\{ \rho_{0}\in\left[-1+\varepsilon_{1},1\right]:T_{n}\left(\rho_{0}\right)\leq c_{\psi_{nh,\rho_{0}}}\left(.5\right)\right\} \ \text{for}\ \psi_{nh,\rho_{0}}\ \text{as\ in}\ \text{(\ref{eq:MT-CI=000020defn})}.\label{eq:MT-1-sided=000020CIs}
\end{align}
The median-unbiased interval estimator $\widetilde{\rho}_{n\tau}$
of $\rho\left(\tau\right)$ is defined by
\begin{align}
\widetilde{\rho}_{n\tau} & =\left[\widetilde{\rho}_{n\tau,low},\widetilde{\rho}_{n\tau,up}\right],\ \text{where}\nonumber \\
\widetilde{\rho}{}_{n\tau,up} & =\max\left\{ \rho_{0}:\rho_{0}\in CI_{n,\tau}^{up}\left(.5\right)\right\} \ \text{and}\nonumber \\
\widetilde{\rho}{}_{n\tau,low} & =\min\left\{ \rho_{0}:\rho_{0}\in CI_{n,\tau}^{low}\left(.5\right)\right\} .\label{eq:MT-MUE_Def}
\end{align}
By construction, $\widetilde{\rho}_{n\tau,low}\leq\widetilde{\rho}_{n\tau,up}.$\footnote{This holds because $\widetilde{\rho}{}_{n\tau,up}\geq\sup\left\{ \rho_{0}\in\left[-1+\varepsilon_{1},1\right]:c_{\psi_{nh,\rho_{0}}}\left(.5\right)=T_{n}\left(\rho_{0}\right)\right\} $
and $\widetilde{\rho}{}_{n\tau,low}$ is less than or equal to the
infimum of the values in the same set. } In addition, $\widetilde{\rho}_{n\tau}$ is a singleton whenever
the set $\left\{ \rho_{0}\in\left[-1+\varepsilon_{1},1\right]:T_{n}\left(\rho_{0}\right)=c_{\psi_{nh,\rho_{0}}}\left(.5\right)\right\} $
contains a single point, in which case $\widetilde{\rho}_{n\tau}$
equals this point. Simulations show that $\widetilde{\rho}_{n\tau}$
is a singleton with probability close to one and a very short interval
when it is not a singleton. Table \ref{tab:SM-Critical-Values_psi}
provides the critical values $c_{\psi_{nh,\rho_{0}}}\left(.5\right)$
for a wide range of $\psi$. Given these critical values, computation
of $\widetilde{\rho}_{n\tau}$ is fast.
The estimator $\widetilde{\rho}_{n\tau}$ has the following median-unbiasedness
property
\begin{align}
\liminf\limits_{n\rightarrow\infty}P\left(\widetilde{\rho}_{n\tau,up}\geq\rho\left(\tau\right)\right) & \geq1/2\ \textup{and}\nonumber \\
\liminf\limits_{n\rightarrow\infty}P\left(\widetilde{\rho}_{n\tau,low}\leq\rho\left(\tau\right)\right) & \geq1/2.\label{eq:MT-1-sided=000020CIs-1-1}
\end{align}
Furthermore, as shown in Corollary \ref{cor:MT-Med-Unbiased} in Section
\ref{sec:MT-Asymptotics} below, these asymptotic properties hold
in a uniform sense over the parameter space. This property is important
for ensuring that the asymptotic properties of $\widetilde{\rho}_{n\tau}$
reflect its finite-sample properties.
\section{Data-dependent Bandwidth Parameter \texorpdfstring{$\boldsymbol{h}$}{h}\protect\label{sec:MT-Choice_of_h}}
The CI for $\rho\left(\tau\right)$ proposed in Section \ref{subsec:MT-CI=000020AR=000020Par}
requires a tuning parameter $h$ that determines the interval of observations
used to construct the CI. This section introduces a data-dependent
method of selecting $h$ that minimizes a forecast-error criterion.
We present its asymptotic properties in Section \ref{subsec:MT-Asymptotic-Results-for-hhat}
under somewhat high-level conditions.
The forecast-error criterion is the average over $t=1,...,n$ of the
squared errors from forecasting $Y_{t}$ using the predicted value
$\widehat{Y}_{t}$. Specifically, for $t>\lfloor nh\rfloor,$ let
$(\widehat{\mu}_{t-1}(h),\widehat{\rho}_{t-1}(h))$ denote the LS
estimator of $(\mu_{t},\rho_{t})$ from the regression of $Y_{s}$
on a constant and $Y_{s-1}$ using the $nh$ observations $s=t-\lfloor nh\rfloor,...,t-1.$
For $t\leq\lfloor nh\rfloor,$ since there are fewer than $\lfloor nh\rfloor$
observations, $(\mu_{t},\rho_{t})$ is estimated using the LS estimator
$(\widehat{\mu}_{t-1}(h),\widehat{\rho}_{t-1}(h))$ based on the $\lfloor nh\rfloor$
observations $s=1,...,t-1,t+1,...,\lfloor nh\rfloor+1.$ The observation
$Y_{t}$ is predicted by $\widehat{\mu}_{t-1}(h)+Y_{t-1}\widehat{\rho}_{t-1}(h)$
for $t=1,...,n$. The forecast-error criterion is
\begin{equation}
FE_{n}(h):=n^{-1}\sum_{t=1}^{n}(Y_{t}-\widehat{\mu}_{t-1}(h)-Y_{t-1}\widehat{\rho}_{t-1}(h))^{2}.\label{eq:MT-Forecast=000020error=000020defn}
\end{equation}
We assume that $h$ is chosen from a finite set $\mathcal{H}_{n},$
which depends on $n$ and whose cardinality may depend on $n.$ By
definition, the data-dependent choice of $h,$ $\widehat{h},$ minimizes
$FE_{n}(h)$ over $\mathcal{H}_{n}$: \footnote{If the argmin is not unique, for specificity $\widehat{h},$ is taken
to be the largest argmin. But, non-uniqueness occurs with probability
zero provided the innovations have a continuous component.}
\begin{equation}
\widehat{h}:=\arg\min_{\mathcal{H}_{n}}FE_{n}\left(h\right).\label{eq:MT-h^hat=000020defn}
\end{equation}
\begin{rem}
The criterion $FE_{n}(h)$ has the desirable property of being an
average of unbiased risk estimators for $t>\lfloor nh\rfloor.$ This
holds because
\begin{eqnarray}
& & \hspace{-0.08in}E(Y_{t}-\widehat{\mu}_{t-1}(h)-Y_{t-1}\widehat{\rho}_{t-1}(h))^{2}\nonumber \\
& = & \hspace{-0.08in}E(U_{t}-(\widehat{\mu}_{t-1}(h)-\mu_{t})+Y_{t-1}(\widehat{\rho}_{t-1}(h)-\rho_{t})))^{2}\nonumber \\
& = & \hspace{-0.08in}E U_{t}^{2}-2E U_{t}(\widehat{\mu}_{t-1}(h)-\mu_{t}+Y_{t-1}(\widehat{\rho}_{t-1}(h)-\rho_{t}))+E(\widehat{\mu}_{t-1}(h)-\mu_{t}+Y_{t-1}(\widehat{\rho}_{t-1}(h)-\rho_{t}))^{2}\nonumber \\
& = & \hspace{-0.08in}E U_{t}^{2}+E(\widehat{\mu}_{t-1}(h)-\mu_{t}+Y_{t-1}(\widehat{\rho}_{t-1}(h)-\rho_{t}))^{2},
\end{eqnarray}
where the third equality holds for $t>\lfloor nh\rfloor$ since $(\widehat{\mu}_{t-1}(h),\widehat{\rho}_{t-1}(h))$
is a function of the innovations indexed by $s\leq t-1$ and $E(U_{t}|U_{t-1},...,U_{1},Y_{0}^{\ast})=0.$
For the relatively small number of initial observations with $t\leq\lfloor nh\rfloor,$
the third equality does not hold because $\widehat{\rho}_{t-1}(h)$
depends on observations $s$ with $s\geq t.$
In terms of computation, the criterion $FE_{n}(h)$ has the advantage
that the same $\widehat{h}$ is employed for multiple $\lfloor nh\rfloor$
values of interest. It also avoids introducing additional tuning parameters,
such as the length of a forecasting period.
\end{rem}
\begin{rem}
In Section \ref{subsec:MT-Asymptotic-Results-for-hhat}, we show that
the data-dependent $\widehat{h}$ value achieves the asymptotically
optimal trade-off between bias and variance obtained by the infeasible
$\widehat{h}_{opt}$ which minimizes the ``empirical loss'' as a
function of $h$ defined in \eqref{eq:MT-Empirical=000020loss=000020defn}
below. In consequence, undersmoothing $\widehat{h}$ should yield
a value of $h$ for which the bias is dominated by the variance asymptotically,
the CI's defined above have correct asymptotic size, and the median-unbiased
interval estimator is asymptotically median unbiased. We employ a
relatively sophisticated undersmoothing method. The objective of undersmoothing
(i.e., making $\widehat{h}_{us}$ smaller than $\widehat{h}$) is
to reduce the bias. The cost of undersmoothing is an increase in variance.
We want to undersmooth more when the cost of doing so in terms of
increasing the variance is relatively small, and we want to undersmooth
less when the cost is larger. This means that we need to take account
of the shape of the $FE_{n}(h)$ function. When it is flatter at its
minimum, we want to undersmooth more. This leads to the following
definition:
\begin{equation}
\widehat{h}_{us0}=\min\left\{ h\mathcal{\in H}_{n}:FE_{n}(h)\leq Q_{c_{1}}(FE_{n})\right\} ,\label{eq:MT-hhat_us0-def}
\end{equation}
where $Q_{c_{1}}(FE_{n})$ is the $c_{1}$ quantile of $FE_{n}$
among $h\mathcal{\in H}_{n}$ and we take $c_{1}=.2$ in the simulations
and applications. This definition does not guarantee that $\widehat{h}_{us0}$
is of a smaller order than $\widehat{h}$, which is necessary to ensure
proper asymptotics. In consequence, we modify the definition as follows:
\begin{equation}
\widehat{h}_{us}=\min\left\{ \widehat{h}_{us0},\widehat{h}_{us1}\right\} ,\label{eq:MT-hhat_us-def}
\end{equation}
where $\widehat{h}_{us1}=c_{2}n^{-a}\widehat{h}$ for $c_{2},a>0.$
In the simulations and applications, we use $c_{2}=1.5$ and $a=1/10.$
\end{rem}
\section{Monte Carlo Simulations\protect\label{sec:MT-Monte-Carlo-Simulations}}
In this section, we analyze the finite-sample performance of the methods
introduced above using Monte Carlo simulations. First, we describe
the data generating processes (DGP's) considered and how the CI and
MUE are implemented. Then, we report the results. We find that the
proposed CI has reasonably good coverage probabilities and short average
lengths for most of the DGP's considered. We also find that the MUE
performs quite well both when the true DGP of $\rho_{t}$ is time-varying
and flat.
\subsection{Simulation Setup and Methodology}
We consider 21 DGP's for $\rho_{t}$. The graphs of the $\rho_{t}$
functions are given in Figures \ref{fig:MT-Sim_CP_AL_MAD_1}--\ref{fig:MT-Sim_CP_AL_MAD_3}
and Figures \ref{fig:Sim_CP_AL_MAD_SM1}--\ref{fig:Sim_CP_AL_MAD_SM4},
which appear in the Supplemental Material. There are five categories
of $\rho_{t}$ functions: sinusoidal, linear, flat linear, flat, and
kinked linear. For the sinusoidal functions, we consider sinusoidal
functions for $\rho_{t}$ with 6 different ranges and shapes. For
example, ``sin 1.00-0.90-1.00'' corresponds to the sinusoidal function
where $\rho_{t}$ first achieves its maximum of 1.00 at $t/n=.20$,
then drops to 0.90 at $t/n=.60$ and finally increases to 1.00 at
$t/n=1$, with a frequency of 2.5$\pi$. For linear functions, we
consider 4 different linear DGP's for $\rho_{t}$ corresponding to
different values of the slopes and intercepts. For flat linear functions,
the first half of $\rho_{t}$ is flat and the second half is linear,
and 4 different cases are considered. For example, ``flat-lin 0.90-0.99''
means
\[
\rho_{t}=\rho(t/n)=\begin{cases}
.90 & \text{for }t/n\in\left[0,1/2\right],\\
.18t/n+.81 & \text{for }t/n\in(1/2,1].
\end{cases}
\]
For the flat functions, we consider 3 values .75, .90, and .99. In
the Supplemental Material, we also report results for kinked linear
functions, where the first and second halves of $\rho_{t}$ are linear,
but with slopes of different signs. For example, ``kinked 1.00-0.80-1.00''
is a DGP for $\rho_{t}$ where the first half of $\rho_{t}$ decreases
linearly from 1 to .80 while the second half increases linearly from
.80 to 1. We consider 4 different cases in this category.
For each of the 21 DGP's for $\rho_{t}$, we consider both constant
and time-varying DGP's for $\mu_{t}$ and $\sigma_{t}$ in \eqref{eq:MT-tvp-model}
except for that of ``sin 1.00-0.90-1.00'' with constant $\mu_{t}$
and $\sigma_{t}$ to present the legend for all the figures. Define $\mu_{t}^{*}=(1-\rho_{t})\mu_{t}$,
which allows one to rewrite \eqref{eq:MT-tvp-model} as
\begin{equation}
Y_{t}=\mu_{t}^{*}+\rho_{t}Y_{t-1}+\sigma_{t}U_{t},\text{ for }t=1,2,...,n.\label{eq:MT-Sim_Yt_DGP}
\end{equation}
When $\mu_{t}$ and $\sigma_{t}$ are constant, we take $\mu_{t}=0$
and $\sigma_{t}=1$, which implies $\mu_{t}^{*}=0$. When $\mu_{t}$
and $\sigma_{t}$ are time-varying, we generate $\mu_{t}$ such that
$\mu_{t}^{*}$ increases linearly from -.1 to .1, and $\sigma_{t}$
increases linearly from .95 to 1.05 as $t/n$ goes from 0 to 1. In
consequence, there are a total of 41 ($=21\times2-1$) DGP's for $\left(\rho_{t},\mu_{t}^{*},\sigma_{t}\right)$.
We take $U_{t}$ to be i.i.d. $N\left(0,1\right)$, initialize $Y_{0}$
by drawing it from a $N\left(0,\ol{\sigma}_{n}^{2}\right)$ distribution,
where $\ol{\sigma}_{n}=1/\left(1-\ol{\rho}_{n}^{2}\right)$ and $\ol{\rho}_{n}=n^{-1}\sum_{t=1}^{n}\rho_{t}$,
and form the $Y_{t}$ sequence as in \eqref{eq:MT-Sim_Yt_DGP}.
We compute a nominal .95 two-sided CI and MUE for $\rho(\tau)$ for
each of 5 time points of interest, indexed by $\tau\in\left\{ .2,.4,.6,.8,1\right\} $.
The CI and MUE are implemented as follows. First, we compute $\widehat{h}$
based on the method described in Section \ref{sec:MT-Choice_of_h}
and take the undersmoothed $\widehat{h}_{us}$ to be as defined in
(\ref{eq:MT-hhat_us-def}) with $c_{1}=.2,$ $c_{2}=1.5,$ and $a=1/10.$
In this step, the set of $nh$ values based on $\mathcal{H}_{n}$
is $\left\{ 140,155,...,500,650,...,1500\right\} $. Next, for each
candidate value $\rho_{0}\in\left\{ -1,-.995,...,.945,.95,.951,.952,...,1\right\} $,
we calculate the t-statistics $T_{n}\left(\rho_{0}\right)$ defined
in \eqref{eq:MT-t-stat=000020def} from the regression of $Y_{t}$
on a constant and $Y_{t-1}$ using the $n\widehat{h}_{us}$ observations
centered around $n\tau$. When $n\tau$ is close to the boundary,
we use $n\widehat{h}_{us}/2$ observations to the side that has abundant
data, and as many observations as available to the other side until
the boundary is hit. In consequence, a different number of observations
for different $\tau$ may be used even though the $n\widehat{h}_{us}$
value is the same. Comparing the $T_{n}\left(\rho_{0}\right)$ values
with corresponding critical values $c_{\psi_{nh,\rho_{0}}}\left(\alpha/2\right)$
and $c_{\psi_{nh,\rho_{0}}}\left(1-\alpha/2\right)$ (for $\psi_{nh,\rho_{0}}$
defined in \eqref{eq:MT-CI=000020defn}) at each candidate $\rho_{0}$
gives the nominal $1-\alpha$ equal-tailed two-sided CI for $\rho(\tau)$
as defined in \eqref{eq:MT-CI=000020defn}. We compute coverage probabilities
and average lengths of the CI's across all simulations for each point
of interest for each DGP.
To calculate the MUE for $\rho(\tau)$, we follow the procedure described
in Section \ref{subsec:MT-Median-Unbiased-Interval-Est} and use $\widetilde{\rho}_{n\tau}$
as the MUE of $\rho\left(\tau\right)$ when $\widetilde{\rho}_{n\tau}$
is a singleton. When $\widetilde{\rho}{}_{n\tau,up}\neq\widetilde{\rho}{}_{n\tau,low}$,
we take $\widetilde{\rho}{}_{n\tau,up}$ as the MUE for $\rho\left(\tau\right)$
based on two considerations. First, using the upper bound $\widetilde{\rho}{}_{n\tau,up}$
has the theoretical property that it is not median-biased towards
zero for positive $\rho\left(\tau\right)$ values because, by Corollary
\ref{cor:MT-Med-Unbiased}, $\liminf_{n\rightarrow\infty}\inf_{\lambda\in\Lambda_{n}}P_{\lambda}\left(\widetilde{\rho}_{n\tau,up}\geq\rho\left(\tau\right)\right)\geq1/2$.
Second, in the simulations we find that $\widetilde{\rho}_{n\tau}$
is much more likely to be an interval when the true $\rho\left(\tau\right)$
is close to one. We report the absolute median biases of the MUE and
the range of the MUE values across the simulations for each time point
of interest for each DGP in the figures.
For each of the 41 DGP's for $Y_{t}$, we run $M=5,000$ simulations.
The length of the $Y_{t}$ sequence is $n=1,500$. Given the 41 DGP's
and 5 time points of interest for each one, a total of 205 cases are
considered.
\subsection{Discussion of Results}
First, Table \ref{tab:MT-Range_of_CP} summarizes the CP results by
reporting the number and percentage of times that the CI CP's lie
in different ranges across the 205 cases considered. The nominal CP
is .95. Table \ref{tab:MT-Range_of_CP} shows that only in a small
fraction of cases, 2.0\%, is the CP quite low, i.e., in $[.875,.90).$
In the majority of cases, 92.2\%, the CP is $.925$ or larger, which
shows that the proposed method has reasonably good CP regardless of
whether the underlying DGP is curvy or flat and whether the point
of interest is in the middle or at the boundary of the time period.
\begin{table}
\begin{centering}
\caption{\protect\label{tab:MT-Range_of_CP}Distribution of CP's Across 205
Cases}
\par\end{centering}
\centering{}
\begin{tabular}{cccccc}
\toprule
CP Range & {[}.875,.90) & {[}.90,.925) & {[}.925,.94) & {[}.94,.96) & {[}.96,.965{]}\tabularnewline
\midrule
Number & 4 & 12 & 34 & 148 & 7\tabularnewline
Percent & 2.0\% & 5.9\% & 16.6\% & 72.2\% & 3.4\%\tabularnewline
\bottomrule
\end{tabular}
\end{table}
Second, we consider summary statistics for the absolute values of
the median biases (AMB) of the MUE, i.e., $\abs{\text{median}\left(\widetilde{\rho}_{n\tau,up}-\rho\left(\tau\right)\right)}$,
across the 205 cases considered. The mean and median of the absolute
median biases across all cases are .004 and .003, respectively. The
range is $[.000,.023]$. Thus, the magnitudes of the absolute median
biases of the MUE are generally small.
Third, we consider the values of $n\widehat{h}_{us}$ selected by
the data-dependent method. The mean, median, and range are 212, 207,
and {[}56, 433{]}, respectively. The $n\widehat{h}_{us}$ values vary
depending on the shape of the $\rho$ function considered. The flatter
is the true $\rho$ function, the larger are the $n\widehat{h}_{us}$
values.
Now, we describe the simulation results for the different $\rho$
functions considered. Figure \ref{fig:MT-Sim_CP_AL_MAD_1} provides
simulation results for five sinusoidal $\rho$ functions with constant
$\mu$ and $\sigma^{2}$ functions. The amplitudes of the sinusoidal
functions increase as one moves down the rows. In the first column
of graphs, the $\rho$ function's maximum value occurs at $\tau=1,$
which corresponds to $t=n\tau=n.$ In the second column of graphs,
the $\rho$ function's minimum value occurs at $\tau=1.$ Each graph
consists of an ``upper'' and a ``lower'' graph. The ``upper''
graph shows the true $\rho$ function, the average CI lower and upper
bounds at five points of interest $\tau=.2,$ $.4,$ $.6,$ $.8,$
$1.0,$ where $t=n\tau,$ and twice the average lower and upper absolute
deviations of the MUE (i.e., $2E|\widehat{\rho}_{t}-\rho_{t}|\mathbf{\mathbbm1}(\widehat{\rho}_{t}<\rho_{t})$
and $2E|\widehat{\rho}_{t}-\rho_{t}|\mathbf{\mathbbm1}(\widehat{\rho}_{t}>\rho_{t}),$
respectively, whose sum is somewhat analogous to two standard deviations)
at each of these points of interest. The difference between the average
upper and lower CI bounds gives the average lengths (AL's) of the
CI's. The ``lower'' graphs report the CI coverage probabilities
(CP's) at the five points of interest. The CI's all have nominal CP
.95.
The sinusoidal graphs in Figure \ref{fig:MT-Sim_CP_AL_MAD_1} show
the following: (i) All but three of the CP's range from .925 to .965,
which are close to the nominal CP of .95. (ii) The AL's are shortest
for $\rho(\tau)$ values closest to one, and longest for $\rho(\tau)$
values farthest from one. This reflects the $nh$-consistency of the
LS estimator in the (temporally local) unit root and local-to-unity
scenarios, and the $(nh)^{1/2}$-consistency of the LS estimator in
the (temporally local) stationary scenarios. (iii) The AL's are larger
for the $\tau=1$ than the $\tau<1$ cases because fewer observations
are used to construct the CI in the former cases, which reflects the
need to reduce the boundary bias in order to obtain a good CP. Combining
this fact with point (ii), we find the largest AL's occur when $\tau=1$
{\parfillskip=0pt\par}
\begin{figure}[H]
\captionsetup[subfigure]{position=top,font=scriptsize,singlelinecheck=off,justification=raggedright}\vspace{-2.5em}
\begin{centering}
\subfloat[sin 1.00-0.90-1.00, constant $\mu$ and $\sigma$]{\noindent\includegraphics[scale=0.36,viewport=0bp 0bp 576bp 576bp]{N1500/sin-1. 00-0. 90-1. 00_varmusigma=N_qtl=0. 2_N=1500_M=5000}
}\quad{}\subfloat[sin 0.90-1.00-0.90, constant $\mu$ and $\sigma$]{\includegraphics[scale=0.36]{N1500/sin-0. 90-1. 00-0. 90_varmusigma=N_qtl=0. 2_N=1500_M=5000}
}
\par\end{centering}
\vspace{-0.65em}
\begin{centering}
\subfloat[sin 1.00-0.80-1.00, constant $\mu$ and $\sigma$]{\includegraphics[scale=0.36]{N1500/sin-1. 00-0. 80-1. 00_varmusigma=N_qtl=0. 2_N=1500_M=5000}
}\quad{}\subfloat[sin 0.80-1.00-0.80, constant $\mu$ and $\sigma$]{\includegraphics[scale=0.36]{N1500/sin-0. 80-1. 00-0. 80_varmusigma=N_qtl=0. 2_N=1500_M=5000}
}
\par\end{centering}
\vspace{-0.65em}
\begin{centering}
\subfloat[sin 1.00-0.60-1.00, constant $\mu$ and $\sigma$]{\includegraphics[scale=0.36]{N1500/sin-1. 00-0. 60-1. 00_varmusigma=N_qtl=0. 2_N=1500_M=5000}
}\quad{}\captionsetup[subfloat]{position=top,font=scriptsize,singlelinecheck=off,justification=raggedright,labelformat=empty}\subfloat[Legend for the Figures]{\includegraphics[scale=0.36]{N1500/flat-0. 99_us_1. 5_varmusigma=Nlegendonly}
}
\par\end{centering}
\vspace{-0.75em}
\caption{\protect\label{fig:MT-Sim_CP_AL_MAD_1}CP's and AL's of CI's for $\rho\left(\tau\right)$
and MAD's of the MUE of $\rho\left(\tau\right)$}
\end{figure}
{\noindent}and $\rho(\tau)$ is far from one (which occurs in the
two graphs in the second column). (iv) The magnitudes of the MUE average
absolute deviations are roughly proportional to the AL's of the CI's.
(v) The simulation results for the sinusoidal functions with time-varying
$\mu$ and $\sigma^{2}$ functions, which are given in Figure \ref{fig:Sim_CP_AL_MAD_SM1}
in the Supplemental Material, are quite similar to those in Figure
\ref{fig:MT-Sim_CP_AL_MAD_1}.
Figure \ref{fig:MT-Sim_CP_AL_MAD_2}(a)-(d) provides results for linear
$\rho$ functions with constant $\mu$ and $\sigma^{2}$ functions.
The results are similar to those for the sinusoidal functions, although
all of the CP's lie between .924 and .957. Points (ii)-(v) above also
apply to the linear functions in Figure \ref{fig:MT-Sim_CP_AL_MAD_2}.
Results for time-varying $\mu$ and $\sigma^{2}$ are provided in
Figure \ref{fig:Sim_CP_AL_MAD_SM2} in the Supplemental Material,
and are similar to those in Figure \ref{fig:MT-Sim_CP_AL_MAD_2}.
Figures \ref{fig:MT-Sim_CP_AL_MAD_2}(e)-(f) and \ref{fig:MT-Sim_CP_AL_MAD_3}(a)-(b)
report results for flat-linear $\rho$ functions with $\rho(t)$ varying
between $.90$ and $.99$ in Figure \ref{fig:MT-Sim_CP_AL_MAD_2}(e)-(f)
and between $.80$ and $.99$ in Figure \ref{fig:MT-Sim_CP_AL_MAD_3}(a)-(b).
Figures \ref{fig:MT-Sim_CP_AL_MAD_2}(e) and \ref{fig:MT-Sim_CP_AL_MAD_3}(a)
report results for constant $\mu$ and $\sigma^{2};$ while Figures
\ref{fig:MT-Sim_CP_AL_MAD_2}(f) and \ref{fig:MT-Sim_CP_AL_MAD_3}(b)
report results for time-varying $\mu$ and $\sigma^{2}.$ The CP results
in Figures \ref{fig:MT-Sim_CP_AL_MAD_2}(e)-(f) and \ref{fig:MT-Sim_CP_AL_MAD_3}(a)-(b)
lie between .932 and .956.
Figures \ref{fig:Sim_CP_AL_MAD_SM2}(e)-(f) and \ref{fig:Sim_CP_AL_MAD_SM3}(a)-(b)
provide analogous results to those just discussed for flat-linear
$\rho$ functions, but for the case where the flat part has value
$.99,$ rather than $.80$ or $.90,$ and the linear part has a negative
slope, rather than a positive slope. The CP results for these cases
are lower than those for the flat-linear $\rho$ functions in Figures
\ref{fig:MT-Sim_CP_AL_MAD_2} and \ref{fig:MT-Sim_CP_AL_MAD_3}. For
example, Figure \ref{fig:Sim_CP_AL_MAD_SM2}(e) shows a CP of .905
at $\tau=.60.$ For the other four cases, the CP's are in the range
of .917 to .957.
Figure \ref{fig:MT-Sim_CP_AL_MAD_3}(c)-(f) considers flat $\rho$
functions. In Figure \ref{fig:MT-Sim_CP_AL_MAD_3}(c)-(d), $\rho=.99$
and the $\mu$ and $\sigma^{2}$ functions are constant and time-varying,
respectively. Figure \ref{fig:MT-Sim_CP_AL_MAD_3}(e)-(f) is analogous,
but with $\rho=.90.$ Figure \ref{fig:Sim_CP_AL_MAD_SM3}(c)-(d) also
is analogous, but with $\rho=.75.$ The CP's in the flat $\rho$ function
cases are all close to .95 and lie between .932 and .960. This occurs
because there is no bias accruing due to a time-varying $\rho$ function.
However, in the cases of flat $\rho,$ $\mu,$ and $\sigma^{2}$ functions,
the CI's are not as short as the oracle CI's that rely on knowledge
that the $\rho,$ $\mu,$ and $\sigma^{2}$ functions are all flat
and take $nh=n.$ For the case of $\rho=.99$ and flat $\mu$ and
$\sigma^{2}$ functions, the AL's of the TVP CI and the oracle CI
are $.022$ and $.012,$ respectively, at $\tau=.2$. For $\rho=.90,$
they are $.085$ and $.039$ at the same $\tau$. For $\rho=.75$,
they are $.124$ and $.062$. There is a substantial increase in the
average lengths in the flat cases using the methods of this paper,
in order to ensure good CP's in the near flat cases.
Lastly, Figures \ref{fig:Sim_CP_AL_MAD_SM3}(e)-(f) and \ref{fig:Sim_CP_AL_MAD_SM4}(a)-(f)
in the Supplemental Material provide results for kinked $\rho$ functions
that either linearly increase until $\tau=.5$ and then linearly decrease,
or linearly decrease until $\tau=.5$ and then linearly increase.
Both constant and time-varying $\mu$ and $\sigma^{2}$ functions
are considered. In two cases, the CP's are .904 and .906. In all other
cases, {\parfillskip=0pt\par}
\begin{figure}[H]
\captionsetup[subfigure]{position=top,font=scriptsize,singlelinecheck=off,justification=raggedright}\vspace{-2.5em}
\begin{centering}
\subfloat[linear 0.90-1.00, constant $\mu$ and $\sigma$]{\includegraphics[scale=0.36]{N1500/linear-0. 90-1. 00_varmusigma=N_qtl=0. 2_N=1500_M=5000}
}\quad{}\subfloat[linear 1.00-0.90, constant $\mu$ and $\sigma$]{\includegraphics[scale=0.36]{N1500/linear-1. 00-0. 90_varmusigma=N_qtl=0. 2_N=1500_M=5000}
}
\par\end{centering}
\vspace{-0.65em}
\begin{centering}
\subfloat[linear 0.60-0.90, constant $\mu$ and $\sigma$]{\includegraphics[scale=0.36]{N1500/linear-0. 60-0. 90_varmusigma=N_qtl=0. 2_N=1500_M=5000}
}\quad{}\subfloat[linear 0.90-0.60, constant $\mu$ and $\sigma$]{\includegraphics[scale=0.36]{N1500/linear-0. 90-0. 60_varmusigma=N_qtl=0. 2_N=1500_M=5000}
}
\par\end{centering}
\vspace{-0.65em}
\begin{centering}
\subfloat[flat-lin 0.90-0.99, constant $\mu$ and $\sigma$]{\includegraphics[scale=0.36]{N1500/flat-lin-0. 90-0. 99_varmusigma=N_qtl=0. 2_N=1500_M=5000}
}\quad{}\subfloat[flat-lin 0.90-0.99, time-varying $\mu$ and $\sigma$]{\includegraphics[scale=0.36]{N1500/flat-lin-0. 90-0. 99_varmusigma=T_qtl=0. 2_N=1500_M=5000}
}
\par\end{centering}
\vspace{-0.75em}
\caption{\protect\label{fig:MT-Sim_CP_AL_MAD_2}CP's and AL's of CI's for $\rho\left(\tau\right)$
and MAD's of the MUE of $\rho\left(\tau\right)$}
\end{figure}
\begin{figure}[H]
\captionsetup[subfigure]{position=top,font=scriptsize,singlelinecheck=off,justification=raggedright}\vspace{-2.5em}
\begin{centering}
\subfloat[flat-lin 0.80-0.99, constant $\mu$ and $\sigma$]{\includegraphics[scale=0.36]{N1500/flat-lin-0. 80-0. 99_varmusigma=N_qtl=0. 2_N=1500_M=5000}
}\quad{}\subfloat[flat-lin 0.80-0.99, time-varying $\mu$ and $\sigma$]{\includegraphics[scale=0.36]{N1500/flat-lin-0. 80-0. 99_varmusigma=T_qtl=0. 2_N=1500_M=5000}
}
\par\end{centering}
\vspace{-0.65em}
\begin{centering}
\subfloat[flat 0.99, constant $\mu$ and $\sigma$]{\includegraphics[scale=0.36]{N1500/flat-0. 99_varmusigma=N_qtl=0. 2_N=1500_M=5000}
}\quad{}\subfloat[flat 0.99, time-varying $\mu$ and $\sigma$]{\includegraphics[scale=0.36]{N1500/flat-0. 99_varmusigma=T_qtl=0. 2_N=1500_M=5000}
}
\par\end{centering}
\vspace{-0.65em}
\begin{centering}
\subfloat[flat 0.90, constant $\mu$ and $\sigma$]{\includegraphics[scale=0.36]{N1500/flat-0. 90_varmusigma=N_qtl=0. 2_N=1500_M=5000}
}\quad{}\subfloat[flat 0.90, time-varying $\mu$ and $\sigma$]{\includegraphics[scale=0.36]{N1500/flat-0. 90_varmusigma=T_qtl=0. 2_N=1500_M=5000}
}
\par\end{centering}
\vspace{-0.75em}
\caption{\protect\label{fig:MT-Sim_CP_AL_MAD_3}CP's and AL's of CI's for $\rho\left(\tau\right)$
and MAD's of the MUE of $\rho\left(\tau\right)$}
\end{figure}
{\noindent}the CP's are in the range of .930 to .953. When the difference
between the minimum and maximum $\rho$ value considered is large,
viz., $.4,$ the AL's of the CI's are large. This occurs because the
data-dependent choice of $h$ is small in order to avoid bias.
Overall, the CP's of the CI's are quite good with only a few being
as low as .88 or .89 out of the 205 cases considered. The AL's vary
substantially across scenarios depending on how close $\rho(\tau)$
is to one and how much the $\rho$ function varies with time, as is
to be expected.
\section{Empirical Applications\protect\label{sec:MT-Empirical-Applications}}
This section presents applications of the proposed methods to time
series of inflation and exchange rates in several countries. The data
comes from the IMF (International Financial Statistics (IFS)) database.
In the Supplemental Material, results are provided for some additional
countries and for interest rates. The Supplemental Material also provides
applications to eight macroeconomic series for the US, using the Federal
Reserve Economic Data (FRED).
In terms of computation, we set $nh_{\min}=.2n$ and $nh_{\max}=2n$,
where $n$ is the sample size. For $nh$ values between $nh_{\min}$
and $nh_{\text{mid}}$, where $nh_{\text{mid}}\coloneqq.5n$, we use
a grid size of $.02n$, while for $nh$ values between $nh_{\text{mid}}$
and $nh_{\max}$, we use a grid size of $.05n$ because the range
of $nh$ values is wider than between $nh_{\min}$ and $nh_{\text{mid}}$.
\subsection{Inflation\protect\label{subsec:MT-EMP_Inflation}}
First, we apply our method to the monthly inflation data, defined
as the percentage change in CPI over the previous month, see Figure
\ref{fig:MT-EMP_AR1_1}. We consider three countries: the US, Canada,
and Germany. The data span is Feb 1955 to Oct 2022, which contains
$n=813$ observations for each country. We compute the $n\widehat{h}_{us}$
values based on (\ref{eq:MT-hhat_us-def}). Given these, we compute
the MUE's and 90\% CI's for $\rho\left(t\right)$ at each $t$ and
for each country. As a robustness check, we multiply the undersmoothed
$n\widehat{h}_{us}$ by 1.5 and present the results in the right column
of each figure. The final $n\widehat{h}_{us}$ values used for the
computations are listed in the titles of the figures. For all of the
inflation series, Ljung-Box tests based on six lags of the residuals
from the AR(1) model fail to reject the null hypothesis of no autocorrelation
at the 5\% significance level, see Table \ref{tab:ACF_Test_AR1}.
For comparative purposes, we also fit a constant autoregression coefficient
AR(1) model to the data and report the estimated $\widehat{\rho}$
and its 90\% CI in the same graph. To obtain the constant parameter
MUE's denoted by the flat red solid lines and constant parameter CI's
denoted by the flat red dotted lines, we fix $n\widehat{h}_{us}=2n$
and apply our method. Note that the
\begin{figure}[H]
\captionsetup[subfigure]{position=top,font=scriptsize,singlelinecheck=off,justification=raggedright}\vspace{-1.2cm}
\begin{centering}
\subfloat[US Inflation, $n\widehat{h}_{us}=$ 125]{\includegraphics[scale=0.36]{\string"tvpar_emp_figures/AR1/US_inflation,_nh=125quantile\string".pdf}
}\quad{}\subfloat[US Inflation, $1.5n\widehat{h}_{us}=$ 188]{\includegraphics[scale=0.36]{\string"tvpar_emp_figures/AR1/US_inflation,_nh=188quantile\string".pdf}
}
\par\end{centering}
\begin{centering}
\subfloat[Canada Inflation, $n\widehat{h}_{us}=$ 125]{\noindent\includegraphics[scale=0.36]{\string"tvpar_emp_figures/AR1/Canada_inflation,_nh=125quantile\string".pdf}
}\quad{}\subfloat[Canada Inflation, $1.5n\widehat{h}_{us}=$ 188]{\includegraphics[scale=0.36]{\string"tvpar_emp_figures/AR1/Canada_inflation,_nh=188quantile\string".pdf}
}
\par\end{centering}
\begin{centering}
\subfloat[Germany Inflation, $n\widehat{h}_{us}=$ 125]{\includegraphics[scale=0.36]{\string"tvpar_emp_figures/AR1/Germany_inflation,_nh=125quantile\string".pdf}
}\quad{}\subfloat[Germany Inflation, $1.5n\widehat{h}_{us}=$ 188]{\includegraphics[scale=0.36]{\string"tvpar_emp_figures/AR1/Germany_inflation,_nh=188quantile\string".pdf}
}
\par\end{centering}
\caption{Estimates and 90\% CI's for the AR(1) Coefficient in TVP-AR(1) Models:
US, Canada, and Germany Inflation. The solid black line is the MUE's
of $\rho\left(t\right)$ in a TVP-AR(1) model, with the 90\% CI's
denoted by the dotted black line. The solid red line is the MUE's
of the AR coefficient in a constant parameter AR(1) model, with the
90\% CI's denoted by the dotted red line.\protect\label{fig:MT-EMP_AR1_1}}
\end{figure}
{\noindent}method to derive the constant parameter estimates is equivalent
to Mikusheva's \citeyearpar{mikusheva2007uniform} modification of
Stock's \citeyearpar{stock1991confidence} method. \citet{mikusheva2007uniform}
proves that these confidence sets are uniformly valid asymptotically
for non-time-varying $\rho\in(0,1]$.
Inflation persistence in the US has been extensively studied with
mixed results regarding its extent and causes. Studies like those
by \citet{cogley2001evolving} have noted a decrease in inflation
persistence following the early 1980s, attributing this change to
shifts in monetary policy, especially the Federal Reserve's (Fed)
increased focus on inflation targeting. This view is supported by
research that points to the Volcker disinflation period as a pivotal
time when the Fed's credibility was enhanced, leading to a stabilization
of inflation expectations, see \citet{sims1992interpreting}. We also
find a sharp decline from .78 to .12 in inflation persistence in the
US measured by the MUE of the TVP-AR coefficient between 1983 and
1995 in Figure \ref{fig:MT-EMP_AR1_1}(a), which is consistent with
the above empirical results.
Entering the era of 2000s, inflation persistence remains a central
topic in macroeconomic research. One line of research (e.g., \citet*{eggertsson2003zero,chung2012have})
concerns zero lower bound (ZLB) effects when the Fed set the interest
rates close to zero. The theory predicts that at ZLB, monetary policy
could be less effective in controlling inflation, thereby potentially
increasing inflation persistence if not coupled with assertive non-traditional
interventions. To overcome the challenges of the ZLB effects on the
effectiveness of its monetary policies, the Fed initiated a new practice
called ``forward guidance,'' where the future course of monetary
policy is communicated to the public by the central bank. Numerous
studies (e.g., \citet*{CAMPBELL2012BIP,negro2023JPE}) have suggested
that forward guidance is effective in enabling the Fed to better control
inflation, which could lead to lower inflation persistence. Along
this line, \citet{bernanke2020new} argues that forward guidance combined
with quantitative easing (QE) granted the Fed significantly more space
to provide accommodation when its standard policy rate was near zero.
More recently, \citet*{cole2023living} argue that despite the significant
efforts to make their policy credible, the credibility of most central
banks including the Fed has been generally declining, making monetary
policies aimed at controlling inflation less effective. A concrete
manifestation of their claim would be an increase in the inflation
persistence over a longer horizon.
We find an increase in inflation persistence in the US from .3 to
.6 measured by the MUE of TVP-AR coefficient between 2000 and 2008
in Figure \ref{fig:MT-EMP_AR1_1}(a), which seems to be mostly driven
by the ZLB effects despite the introduction of forward guidance during
this period. Later on, when more specific monetary policies (e.g.,
QE) were introduced to stimulate the economy following the 2008 financial
crisis, inflation persistence drops from .6 to .4 between 2010 and
2013, in line with the result of \citet{bernanke2020new}. With that
said, we observe in Figure \ref{fig:MT-EMP_AR1_1}(a) an upward trend
in inflation persistence in the US over a longer horizon between 2000
and 2022, which could be caused by a decline in the credibility of
the Fed as claimed by \citet*{cole2023living}. Overall, we conclude
that while the Fed's measures have been effective in controlling inflation
expectations and persistence, more nuanced policies are needed to
handle complications arising from the ZLB effects, heterogeneous beliefs
(\citet*{andrade2019forward}), structural changes (e.g., higher volatility
and shifts in consumer behavior and supply chains) caused by the Covid-19
and other factors.
We also apply our method to study inflation persistence in Canada
and Germany and obtain the following results. First, we find significant
time variation in inflation persistence as measured by the MUE's of
$\rho\left(t\right)$ for both countries since 1980s. Given the magnitude
of the time-variation, our findings show that inflation persistence
is more likely to be driven by changes in the regime of monetary policy
and credibility of central banks, rather than by nominal or real frictions
in the economy. Second, there seems to be a universal upward trend
since 2000 in inflation persistence for all countries under consideration,
echoing the findings of \citet*{cole2023living}, who argue that there
is a decline in credibility in central bank policies. Third, the introduction
of the Euro in 1999 marked a significant shift in inflation persistence
in Germany, with the European Central Bank (ECB) taking over monetary
policy. Inflation rates remained relatively stable but were influenced
by broader Eurozone policies and economic conditions. The early 2000s
saw moderate inflation, which aligned closely with the ECB\textquoteright s
target. Our method captures this regime change by showing drastically
different patterns of inflation persistence before and after 1999,
supporting the conclusion above that inflation persistence is more
likely to be driven by changes in the regime of monetary policy and
the credibility of central banks.
\subsection{Real Exchange Rate\protect\label{subsec:MT-EMP_Real-Exchange-Rate}}
The second application concerns the real exchange rate. The results
are given in Figure \ref{fig:MT-EMP_AR1_2}. We use US dollars (USD)
as the benchmark currency, and calculate the bilateral real exchange
rate $re_{it}$ of country $i$ at time $t$ as $rex_{it}=nex_{it}\times\frac{CPI_{it}}{CPI_{0t}}$,
where $nex_{it}$ is the nominal exchange rate (USD per domestic currency)
at time $t$, $CPI_{it}$ is the price level of country $i$ at time
$t$, and $CPI_{0t}$ is the price level of the US at time $t$. Therefore,
an increase in $rex_{it}$ represents an appreciation of country $i$'s
currency against USD. We report results for the UK, Sweden, and Switzerland.
The data span for the monthly real exchange rate dataset is Jan 1957
to Aug 2022. Thus, $n=788$ for each country. Similar to the inflation
series, we compute the $n\widehat{h}_{us}$ values based on (\ref{eq:MT-hhat_us-def}).
Then, we compute the MUE's and 90\% CI's for the{\noindent}
\begin{figure}[H]
\captionsetup[subfigure]{position=top,font=scriptsize,singlelinecheck=off,justification=raggedright}\vspace{-1cm}
\begin{centering}
\subfloat[UK Real Exchange Rate, $n\widehat{h}_{us}=$ 823]{\noindent\includegraphics[scale=0.36]{\string"tvpar_emp_figures/AR1/UK_real_exchange_rate,_nh=823quantile\string".pdf}
}\quad{}\subfloat[UK Real Exchange Rate, $1.5n\widehat{h}_{us}=$ 1,234]{\includegraphics[scale=0.36]{\string"tvpar_emp_figures/AR1/UK_real_exchange_rate,_nh=1234quantile\string".pdf}
}
\par\end{centering}
\begin{centering}
\subfloat[Sweden Real Exchange Rate, $n\widehat{h}_{us}=$ 823]{\includegraphics[scale=0.36]{\string"tvpar_emp_figures/AR1/Sweden_real_exchange_rate,_nh=823quantile\string".pdf}
}\quad{}\subfloat[Sweden Real Exchange Rate, $1.5n\widehat{h}_{us}=$ 1,234]{\includegraphics[scale=0.36]{\string"tvpar_emp_figures/AR1/Sweden_real_exchange_rate,_nh=1234quantile\string".pdf}
}
\par\end{centering}
\begin{centering}
\subfloat[Switzerland Real Exchange Rate, $n\widehat{h}_{us}=$ 393]{\includegraphics[scale=0.36]{\string"tvpar_emp_figures/AR1/Switzerland_real_exchange_rate,_nh=393quantile\string".pdf}
}\quad{}\subfloat[Switzerland Real Exchange Rate, $1.5n\widehat{h}_{us}=$ 590]{\includegraphics[scale=0.36]{\string"tvpar_emp_figures/AR1/Switzerland_real_exchange_rate,_nh=590quantile\string".pdf}
}
\par\end{centering}
\caption{Estimates and 90\% CI's for the AR(1) Coefficient in TVP-AR(1) Models:
UK, Sweden, and Switzerland Real Exchange Rate\protect\label{fig:MT-EMP_AR1_2}}
\end{figure}
{\noindent}AR coefficient, $\rho\left(t\right)$, at each $t$ and
for each country. For robustness purposes, we multiply the undersmoothed
$n\widehat{h}_{us}$ by 1.5 and report the results in the right column
of each figure. For the real exchange rate series considered in this
section, Ljung-Box tests based on six lags of the residuals from the
AR(1) model fail to reject the null hypothesis of no autocorrelation
at the 5\% significance level, see Table \ref{tab:ACF_Test_AR1}.
It has been known in the literature that real exchange rates in developed
countries tend to be highly persistent, with deviations from the PPP
level taking a long time to disappear (\citet{rogoff1996purchasing,engel2014exchange}).
There are several explanations in the literature for the real exchange
rate persistence, including nominal price rigidities, interest rate
inertia in monetary policy (\citet{benigno2004real}), heterogeneous
dynamics in subcomponents (\citet*{imbs2005ppp}), Balassa--Samuelson
effects (\citet{balassa1964purchasing,samuelson1964theoretical}),
to name a few. Figure \ref{fig:MT-EMP_AR1_2} presents the results
on real exchange rate persistence measured by the MUE's of $\rho\left(t\right)$
for the UK, Sweden, and Switzerland. Across all three countries the
MUE's are close to 1, showing that our method yields near constant
graphs in scenarios where that seems to be suitable. The 90\% CI's
for $\rho\left(t\right)$ are also very short. \citet*{chari2002can}
report estimates from a constant parameter autoregressive process
to lie between .76 and .87 for US bilateral real exchange rates against
nine developed European countries using data between 1972 and 1994.
They developed a general equilibrium model where the firms can only
set price once per year, which implies a very high level of price
stickiness in the economy. Yet, their model cannot generate the level
of persistence observed in the data. Allowing for a time-varying parameter
autoregressive process and using data from a longer time horizon,
our MUE's of $\rho\left(t\right)$ for the three currencies are even
higher at close to one, supporting their claim that nominal price
rigidities are not enough to explain the high real exchange rate persistence.
Practitioners would need to seek other possible explanations such
as the Balassa--Samuelson effects or heterogeneous dynamics in subcomponents.
While high real exchange rate persistence seems prevalent, there is
heterogeneity in the pattern of persistence across the countries we
consider. In particular, the MUE's of $\rho\left(t\right)$ for Switzerland
demonstrate a downward trend since early 1980s in Figure \ref{fig:MT-EMP_AR1_2}(e).
The Swiss National Bank (SNB) has adopted several significant monetary
policy changes since the 1990s, including the shift to a three-fold
target (price stability, 3-year inflation forecast, and a range for
the 3M Libor, see \citet*{jordan2010ten}) in the late 1990s and the
introduction of negative interest rates in 2015 when it was forced
to abandon a policy of defending the Swiss franc with a peg to the
euro. These policies aim to stabilize price levels and influence interest
rates, which can affect exchange rate dynamics and potentially reduce
persistence by promoting quicker adjustments to shocks. In Figure
\ref{fig:MT-EMP_AR1_2}(e) we observe a sharp decline in the MUE of
AR(1) coefficient between 1995 and 2002 and between 2015 and 2019,
consistent with the timing of SNB's monetary policy changes. Meanwhile,
the European debt crisis and subsequent economic turmoil in the Eurozone
led to significant safe-haven flows into the Swiss franc, prompting
the SNB to implement a cap on the franc\textquoteright s value against
the euro in 2011. This cap was removed in 2015. Such a cap on the
franc's value would increase its real exchange rate persistence since
the price of the franc could have fluctuated above the cap for the
duration of the policy. We also find the real exchange rate persistence
to be slightly increasing between 2011 and 2015 in Figure \ref{fig:MT-EMP_AR1_2}(e).
As shown by these results, our method can capture the major events
and policy changes that affect real exchange rate persistence reasonably
accurately.
\section{Asymptotics\protect\label{sec:MT-Asymptotics}}
This section establishes the correct uniform asymptotic size and asymptotic
similarity of the confidence interval $CI_{n,\tau}$ for $\rho(\tau)$,
the median-unbiased property of $\widetilde{\rho}_{n\tau}$, and some
asymptotic properties of the data-dependent bandwidth parameter $\widehat{h}$.
To prove these results, this section provides asymptotic results for
the LS estimator $\widehat{\rho}_{n\tau}$ and t-statistic $T_{n}\left(\rho_{0,n}\right)$
under drifting sequences of parameter values. All proofs of the results
stated below are given in Section \ref{sec:Theory} of the Supplemental
Material.
\subsection{\protect\label{sec:MT-Parameter=000020Space}Parameter Space}
Let $I_{a,r}:=\left[a-r,a+r\right]$ for $a\in\mathbb{R}$ and $r>0$. Let
$\lfloor x\rfloor$ denote the integer part of $x$.
We impose the following structure on the $\rho$ function: for some
$\varepsilon_{2},\varepsilon_{3}>0$,
\begin{equation}
\rho\left(s\right)=1-\kappa\left(s\right)/b\label{eq:MT-structure-of-rho}
\end{equation}
for $s\in I_{\tau,\varepsilon_{2}}$ and some $b\in\left[\varepsilon_{3},\infty\right]$,
where $\kappa\left(\cdot\right)$ is a nonnegative twice continuously
differentiable function on $I_{\tau,\varepsilon_{2}}$.
The parameter space for $\left(\rho,\mu,\sigma^{2},\kappa,b,F\right)$
is given by
$\Lambda_{n}=\{\lambda=\left(\rho,\mu,\sigma^{2},\kappa,b,F\right)$:
(i) $\rho,\ \mu,$ and $\sigma^{2}$ are Lipschitz functions from
$\left[0,1\right]$ to $\left[-1+\varepsilon_{1},1\right],\ \left[C_{2,L},C_{2,U}\right]$,
and $\left[C_{3,L},C_{3,U}\right]$, respectively, with Lipschitz
constants bounded by $L_{1}$, $L_{2}$, and $L_{3}$, respectively;
(ii) $\rho\left(s\right)=1-\kappa\left(s\right)/b$ for $s\in I_{\tau,\varepsilon_{2}}$
and $b\in[\varepsilon_{3},\infty]$, where $\kappa\left(\cdot\right)$
is a twice continuously differentiable function from $I_{\tau,\varepsilon_{2}}$
to $\left[0,C_{4}\right]$ with Lipschitz constant bounded by $L_{4}$
and $\kappa(\tau)\geq\varepsilon_{4}$;
(iii) $\mu\left(s\right)=C_{\mu1}\exp\left\{ -\eta\left(s\right)/b\right\} +C_{\mu2}$
for $s\in I_{\tau,\varepsilon_{2}}$ and $b\in[\varepsilon_{3},\infty]$,
where $\eta\left(\cdot\right)$ is a nonnegative Lipschitz function
with Lipschitz constant bounded by $L_{5}$;
(iv) $\left\{ U_{t}:t=0,1,...,n\right\} $ is a stationary martingale
difference sequence under $F$ with $E_{F}\left(\rest{U_{t}}\mathscr{G}_{t-1}\right)=0\ \text{a.s.}$,
$E_{F}\left(\rest{U_{t}^{2}}\mathscr{G}_{t-1}\right)=1\ \text{a.s.}$,
and $E_{F}\left[\rest{U_{t}^{4}}\mathscr{G}_{t-1}\right]<M\ \text{a.s.},$
where $\mathscr{G}_{t}$ is some non-decreasing sequence of $\sigma$-fields
for which $\sigma\left(U_{0},...,U_{t}\right)\subseteq\mathscr{G}_{t}$
and $\sigma\left(Y_{0}^{*}\right)\in\mathscr{G}_{t}$ for $t=1,...,n$;
(v) $E_{F}\left(Y_{0}^{*}\right)^{2}\leq C_{5}n$;
for some constants $0<\varepsilon_{1}<2$, $\varepsilon_{2},\varepsilon_{3},\varepsilon_{4}>0$,
$-\infty<C_{j,L}\leq C_{j,U}<\infty$ for $j=2,3$, $C_{3,L}>0$,
$0\leq C_{4}<\infty$, $0\leq C_{5}<\infty$, $L_{j}<\infty\ \forall j\leq5$,
$C_{\mu1},C_{\mu2}\in\left(-\infty,\infty\right),$ and $M\in\left(0,\infty\right)\}$.
$\ $
Note that the dependence on $n$ of $\Lambda_{n}$ is only through
part (v), which concerns the initial condition $Y_{0}^{*}$.\footnote{The parameter space $\Lambda_{n}$ could be made independent of $n$
by specifying that $E\left(Y_{0}^{*}\right)^{2}\leq C_{5}$. But,
allowing the bound on $E\left(Y_{0}^{*}\right)^{2}$ to be $C_{5}n$,
allows the parameter space to expand with $n$ and is in accord with
assumptions in the literature for AR models with a possible unit root
or near unit root.} Also, $\mu\left(s\right)$ is defined in such a way that it is allowed
to vary smoothly on $\mathbb{R}$. It is used in the proof of Lemma \ref{lem:MT-Max=000020Intertemporal=000020Difference}(d),
which is further used to bound the absolute difference between $\mu\left(t\right)$
and $\mu\left(s\right)$ for $t,s\in\left[T_{1},T_{2}\right]$ in
obtaining the asymptotic distribution of the t-statistic $T_{n}\left(\rho_{0,n}\right)$.
$\ $
To establish the correct asymptotic size of the CI for $\rho(\tau)$,
we need to consider sequences of parameters $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\right\} _{n\geq1}$
in $\Lambda_{n}.$ The parameter space $\Lambda_{n}$ is defined to
allow the sequence $\left\{ \rho_{n}(\tau)\right\} _{n\geq1}$ to
equal one for all $n\geq1,$ converge to one at any rate, or be bounded
away from one. More specifically, by the definition of $\rho_{n}(\cdot)$,
\begin{equation}
\rho_{n}\left(\tau\right)=1-\kappa_{n}\left(\tau\right)/b_{n}\ \text{for some }b_{n}\in\left[0,\infty\right].\label{eq:MT-rho_n(tau)}
\end{equation}
If $b_{n}\rightarrow\infty$, then $\rho_{n}\left(\tau\right)$ is local
to one. If $\limsup_{n\rightarrow\infty}b_{n}<\infty$, then $\rho_{n}\left(\tau\right)$
is bounded away from one. Furthermore, $\Lambda_{n}$ is defined so
that, for points $s\in\left[0,1\right]$ other than $\tau,$ $\rho_{n}\left(s\right)$
can converge to one at any rate or be bounded away from one, and the
distance-from-unity behavior may differ between different $s$ outside
of $I_{\tau,\varepsilon_{2}}.$
\subsection{Correct Asymptotic Size of the Confidence Interval for \texorpdfstring{$\boldsymbol{\rho(\tau)}$}{rho(tau)}}
The bandwidth $h$ is assumed to satisfy the following assumptions.
\begin{assumption}[\textbf{Bandwidth $\boldsymbol{h}$}]
\label{assu:MT-h} $h\rightarrow0$ and $nh\rightarrow\infty$ as $n\rightarrow\infty$.
\end{assumption}
\begin{assumption}[\textbf{Order of $\boldsymbol{h}$}]
\label{assu:MT-stationary-nh^5=000020->=0000200} $nh^{5}\rightarrow0$.\footnote{Assumption \ref{assu:MT-stationary-nh^5=000020->=0000200} is used
in the proofs of Lemmas \ref{lem:MT-Order=000020of=000020Y_T0}(b),
\ref{lem:MT-Order=000020of=000020Y_T0}(c), \ref{lem:MT-asympdist_dist_stationary-denom},
and \ref{lem:MT-asympdist_dist_stationary-numerator=000020}(a) below.}
\end{assumption}
Let $P_{\lambda}\left(\cdot\right)$ denote probability under $\lambda\in\Lambda_{n}.$
The correct asymptotic size and asymptotic similarity of the CI $CI_{n,\tau}$
are established in the following theorem. The theorem's proof relies
on the ``drifting pointwise'' asymptotic results given in Sections
\ref{sec:MT-Preliminaries}-\ref{sec:MT-Stationary=000020Case} below
and the generic asymptotic size results in \citet*{andrews2020generic}.
\begin{thm}[\textbf{Correct Asymptotic Size of $\boldsymbol{CI_{\boldsymbol{n,}\tau}}$}]
\label{thm:MT-Cor=000020Asy=000020Size}Under Assumptions \ref{assu:MT-h}
and \ref{assu:MT-stationary-nh^5=000020->=0000200},
\[
\liminf\limits_{n\rightarrow\infty}\inf_{\lambda\in\Lambda_{n}}P_{\lambda}\left(\rho\left(\tau\right)\in CI_{n,\tau}\right)=\limsup\limits_{n\rightarrow\infty}\sup_{\lambda\in\Lambda_{n}}P_{\lambda}\left(\rho\left(\tau\right)\in CI_{n,\tau}\right)=1-\alpha.
\]
\end{thm}
The median-unbiased interval estimator $\widetilde{\rho}_{n\tau}$
has the following median-unbiasedness property. This result is a corollary
to one-sided versions of Theorem \ref{thm:MT-Cor=000020Asy=000020Size}.\footnote{Corollary \ref{cor:MT-Med-Unbiased} holds because $CI_{n,\tau}^{up}\left(.5\right)$
and $CI_{n,\tau}^{low}\left(.5\right)$ both have coverage probabilities
of $1/2$ or greater by the proof of Theorem \ref{thm:MT-Cor=000020Asy=000020Size}
applied to these one-sided CIs and $\left(\widetilde{\rho}_{n\tau}\geq\rho\left(\tau\right)\right)$$\supset$
$CI_{n,\tau}^{up}\left(.5\right)$ and $\left(\widetilde{\rho}_{n\tau}\leq\rho\left(\tau\right)\right)\supset CI_{n,\tau}^{low}\left(.5\right).$}
\begin{cor}[\textbf{Asymptotic Median-Unbiasedness of $\boldsymbol{\boldsymbol{\widetilde{\rho}}_{\boldsymbol{n\tau}}}$}]
\label{cor:MT-Med-Unbiased}Under Assumptions \ref{assu:MT-h} and
\ref{assu:MT-stationary-nh^5=000020->=0000200},
\begin{align*}
\liminf\limits_{n\rightarrow\infty}\inf_{\lambda\in\Lambda_{n}}P_{\lambda}\left(\widetilde{\rho}_{n\tau,up}\geq\rho\left(\tau\right)\right) & \geq1/2\ \textup{and}\\
\liminf\limits_{n\rightarrow\infty}\inf_{\lambda\in\Lambda_{n}}P_{\lambda}\left(\widetilde{\rho}_{n\tau,low}\leq\rho\left(\tau\right)\right) & \geq1/2.
\end{align*}
\end{cor}
\subsection{\protect\label{sec:MT-Preliminaries}Preliminaries}
Define
\begin{equation}
c_{t,j}:=\prod_{k=0}^{j-1}\rho_{t-k}\ \text{for}\ 1\leq j\leq t,\ 1\leq t\leq n,\ \ \text{and}\ c_{t,0}:=1.\label{eq:MT-def_c_t,j}
\end{equation}
We consider the interval $I_{n\tau,nh/2}$ and decompose $Y_{t}^{*}$
into the sum of two parts
\begin{equation}
Y_{t}^{*}=Y_{t}^{0}+c_{t,t-T_{0}}Y_{T_{0}}^{*},\ \text{for }t\in I_{n\tau,nh/2},\label{eq:MT-decompY*}
\end{equation}
where $Y_{t}^{0}$ is an AR(1) process with the same time-varying
parameters $\rho_{t}$ as $Y_{t}^{*}$, but with a zero initial condition
at $T_{0}:=T_{1}-1$, and $c_{t,t-T_{0}}$ is defined in \eqref{eq:MT-def_c_t,j}.
By recursive substitution in \eqref{eq:MT-tvp-model}, we have
\begin{equation}
Y_{t}^{0}=\sum_{j=0}^{t-T_{1}}c_{t,j}\sigma_{t-j}U_{t-j}\ \text{and }Y_{t}^{*}=\sum_{j=0}^{t-T_{1}}c_{t,j}\sigma_{t-j}U_{t-j}+c_{t,t-T_{0}}Y_{T_{0}}^{*}\ \text{for }t=T_{1},...,T_{2}.\label{eq:MT-Y_t^0=000020and=000020Y_t^star}
\end{equation}
Then, from \eqref{eq:MT-tvp-model} and \eqref{eq:MT-decompY*}, we
have
\begin{align}
Y_{t} & =\mu_{t}+\sum_{j=0}^{t-T_{1}}c_{t,j}\sigma_{t-j}U_{t-j}+c_{t,t-T_{0}}Y_{T_{0}}^{*}\ \text{for }t=T_{1},...,T_{2}.\label{eq:MT-YtandYt0}
\end{align}
We consider sequences $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$$.$
The estimand of interest is $\rho_{n}\left(\tau\right)$ for $n\geq1.$
Let
\begin{align}
\rho_{n,t}:=\rho_{n}\left(t/n\right),\ & \mu_{n,t}:=\mu_{n}\left(t/n\right),\ \sigma_{n,t}^{2}:=\sigma_{n}^{2}\left(t/n\right),\nonumber \\
\rho_{n\tau}:=\rho_{n}\left(\tau\right),\ & \mu_{n\tau}:=\mu_{n}\left(\tau\right),\ \text{and }\sigma_{n\tau}^{2}:=\sigma_{n}^{2}\left(\tau\right).\label{eq:MT-rho_n,t=000020defn=0000202nd=000020time}
\end{align}
By part (ii) of the definition of $\Lambda_{n},$ for $t\in I_{n\tau,nh/2}$,
\begin{equation}
\rho_{n,t}:=\rho_{n}\left(t/n\right)=1-\kappa_{n}\left(t/n\right)/b_{n}\label{eq:MT-rho_n}
\end{equation}
for $n$ sufficiently large that $h/2+1/n\leq\varepsilon_{2}.$ For
notational simplicity, we drop the $n$ subscripts on $\rho_{n,t},$
$\mu_{n,t},$ and $\sigma_{n,t}^{2}$ below.
We consider sequences of null hypotheses $H_{0}:\rho_{n\tau}=\rho_{0,n}$
for $n\geq1.$ As noted above, the CI $CI_{n,\tau}$ is defined by
the inversion of tests of such null hypotheses.
$\ $
For some results, we assume the functions $\kappa_{n}\left(\cdot\right)$,
$\mu_{n}\left(\cdot\right)$, and $\sigma_{n}^{2}\left(\cdot\right)$
converge.
\begin{assumption}[\textbf{Limits of $\boldsymbol{\kappa_{n}\left(\cdot\right)}$, $\boldsymbol{\mu_{n}\left(\cdot\right)}$,
and $\boldsymbol{\sigma_{n}^{2}\left(\cdot\right)}$}]
\textbf{} \label{assu:MT-kappa_n_sigma_n=000020mu_n=000020convergence}
$\kappa_{n}\left(\cdot\right)$, $\mu_{n}\left(\cdot\right),$ and $\sigma_{n}^{2}\left(\cdot\right)$
restricted to $I_{\tau,\varepsilon_{2}}$ converge uniformly to some
functions $\kappa_{0}\left(\cdot\right)$, $\mu_{0}\left(\cdot\right),$
and $\sigma_{0}^{2}\left(\cdot\right)$ on $I_{\tau,\varepsilon_{2}}$,
respectively.
\end{assumption}
$\ $
For some results, we assume that $\left\{ b_{n}/\left(nh^{1/2}\right)\right\} _{n\geq1}$
converges.
\begin{assumption}[\textbf{Convergence of $\boldsymbol{b_{n}/\left(nh^{1/2}\right)}$}]
\textbf{} \label{assu:MT-bn=000020div=000020nh=000020converges}
$b_{n}/\left(nh^{1/2}\right)\to w_{0}$ for some $w_{0}\in\left[0,\infty\right]$.
\end{assumption}
$\ $
Assumptions \ref{assu:MT-kappa_n_sigma_n=000020mu_n=000020convergence}
and \ref{assu:MT-bn=000020div=000020nh=000020converges} are innocuous
because, to establish the uniform inference results in Theorem \ref{thm:MT-Cor=000020Asy=000020Size},
we show that it suffices to consider sequences $\left\{ \lambda_{n}\in\Lambda_{n}\right\} _{n\geq1}$
that satisfy these assumptions. Note that Theorem \ref{thm:MT-Cor=000020Asy=000020Size}
and Corollary \ref{cor:MT-Med-Unbiased} do not impose Assumption
\ref{assu:MT-kappa_n_sigma_n=000020mu_n=000020convergence} or \ref{assu:MT-bn=000020div=000020nh=000020converges}.
The following lemma bounds the maximum intertemporal difference between
the TVP $\rho_{t}$ and $\rho_{n\tau}$ for $t\in\left[T_{1},T_{2}\right]$.
\begin{lem}[{{\textbf{Maximum Intertemporal Differences on $\boldsymbol{\left[T_{1},T_{2}\right]}$}}}]
\label{lem:MT-Max=000020Intertemporal=000020Difference} Under Assumptions
\ref{assu:MT-h} and \ref{assu:MT-kappa_n_sigma_n=000020mu_n=000020convergence},
for a sequence $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$,
\begin{enumerate}
\item $\max_{t\in\left[T_{1},T_{2}\right]}\abs{\rho_{t}-\rho_{n\tau}}=O\left(h/b_{n}\right)$,
\item $\max_{t\in\left[T_{1},T_{2}\right]}\abs{\sigma_{t}^{2}-\sigma_{n\tau}^{2}}=O\left(h\right)$,
\item $\max_{t\in\left[T_{1},T_{2}\right]}\max_{0\leq j\leq t-T_{1}}\abs{c_{t,j}-\rho_{n\tau}^{j}}=O\left(nh^{2}/b_{n}\right)$,
$\text{ and}$
\item $\max_{t\in\left[T_{1},T_{2}\right]}\abs{\mu_{t}-\mu_{n\tau}}=O\left(h/b_{n}\right)$.
\end{enumerate}
\end{lem}
The next lemma provides several bounds on the endogenous initial condition
$Y_{T_{0}}^{*}$.
\begin{lem}[\textbf{Order of $\boldsymbol{Y_{T_{0}}^{*}}$}]
\label{lem:MT-Order=000020of=000020Y_T0} For a sequence $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$,
we have
\begin{enumerate}
\item $Y_{T_{0}}^{*}=O_{p}\left(n^{1/2}\right)$ under Assumption \ref{assu:MT-h},
\item $Y_{T_{0}}^{*}=o_{p}\left(b_{n}/\left(nh\right){}^{1/2}\right)$ under
Assumptions \ref{assu:MT-h}, \ref{assu:MT-stationary-nh^5=000020->=0000200},
\ref{assu:MT-kappa_n_sigma_n=000020mu_n=000020convergence}, and \ref{assu:MT-bn=000020div=000020nh=000020converges}
and $nh/b_{n}\to r_{0}=0$, and
\item $Y_{T_{0}}^{*}=O_{p}\left(b_{n}^{1/2}\right)$ under Assumptions \ref{assu:MT-h}
and \ref{assu:MT-stationary-nh^5=000020->=0000200} and $nh/b_{n}\to r_{0}=\infty$.
\end{enumerate}
\end{lem}
\begin{rem}
Lemma \ref{lem:MT-Order=000020of=000020Y_T0}(a) shows that part (v)
of $\Lambda_{n}$ implies that the endogenous initial condition $Y_{T_{0}}^{*}$
is $O_{p}\left(n^{1/2}\right)$. Part (b) of the lemma is used in
the ``very local-to-unity'' case in which $nh/b_{n}\to r_{0}=0$.
It provides a better bound than part (a) when $b_{n}/\left(nh^{1/2}\right)\to w_{0}<\infty$
because, in that case, the bound $o_{p}\left(b_{n}/\left(nh\right){}^{1/2}\right)$
is $o_{p}\left(n^{1/2}\right)$. Part (c) of the lemma provides a
better bound than part (a) in the ``stationary'' case in which $nh/b_{n}\to r_{0}=\infty$.
\end{rem}
\subsection{\protect\label{sec:MT-Local-to-Unit-root} Asymptotics in the Local-to-Unity
Case}
The local-to-unity case is characterized by $nh/b_{n}\rightarrow r_{0}\in\left[0,\infty\right)$.
In this section, we determine the asymptotic distributions of the
local LS estimator $\widehat{\rho}_{n\tau}$ and corresponding t-statistic
$T_{n}\left(\rho_{0,n}\right)$ in the local-to-unity case.
Define
\begin{equation}
t\left(s\right):=t_{n,\tau}\left(s\right)=T_{1}+\lfloor nhs\rfloor=\lfloor n\tau\rfloor-\lfloor nh/2\rfloor+\lfloor nhs\rfloor\label{eq:MT-def=000020t(s)}
\end{equation}
as a function of $s\in\left[0,1\right]$ for any fixed $\tau\in\left[0,1\right]$.
First, we obtain the asymptotic distribution of the zero-initial condition
process $\left(nh\right)^{-1/2}Y_{n,t\left(s\right)}^{0}$ for $s\in\left[0,1\right]$.
Then, we obtain the asymptotic distribution of the normalized initial
condition $Y_{T_{0}}^{*}$ in the $r_{0}\in\left(0,\infty\right)$
case. Lastly, we obtain the asymptotic distributions of $\widehat{\rho}_{n\tau}$
and $T_{n}\left(\rho_{0,n}\right)$. The last step includes dealing
with the initial condition $Y_{T_{0}}^{*}$ in the $r_{0}=0$ case.
\begin{lem}[\textbf{Asymptotic Distribution of $\boldsymbol{Y_{n,t\left(s\right)}^{0}}$}]
\label{lem:MT-y0limit} Under Assumptions \ref{assu:MT-h} and \ref{assu:MT-kappa_n_sigma_n=000020mu_n=000020convergence},
for a sequence $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$
and $nh/b_{n}\rightarrow r_{0}\in\left[0,\infty\right)$ $\left(\text{local-to-unity case}\right)$,
we have
\[
\left(nh\right)^{-1/2}Y_{n,t\left(s\right)}^{0}/\sigma_{0}\left(\tau\right)\Rightarrow I_{\psi}\left(s\right)\ for\ \psi=r_{0}\kappa_{0}(\tau),
\]
where $t\left(s\right)$ is defined in \eqref{eq:MT-def=000020t(s)},
$I_{\psi}\left(s\right)$ is defined in \eqref{eq:MT-I_D,tau(r)=000020def},
and ``$\Rightarrow$'' denotes weak convergence with respect to
the Skorohod metric.
\end{lem}
The following lemma is used to determine the effect of the initial
condition on the LS estimator and t-statistic when $r_{0}\in(0,\infty).$
\begin{lem}[\textbf{Asymptotic Distribution of the Initial Condition $\boldsymbol{Y_{T_{0}}^{*}}$}]
\label{lem:MT-T_0=000020Initial=000020Condition=000020Asy=000020Distn=000020Lem}\textbf{
}Under Assumptions \ref{assu:MT-h} and \ref{assu:MT-kappa_n_sigma_n=000020mu_n=000020convergence}\emph{,
}for a sequence $\{\lambda_{n}=(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n})\in\Lambda_{n}\}_{n\geq1}$
and $nh/b_{n}\rightarrow r_{0}\in(0,\infty)$ \emph{(}local-to-unity
case\emph{), }we have
\[
(2\psi/nh)^{1/2}Y_{T_{0}}^{\ast}/\sigma_{0}\left(\tau\right)\rightarrow_{d}Z_{1}\sim N(0,1),
\]
where $\psi:=r_{0}\kappa_{0}(\tau),$ $Z_{1}$ is independent of $B(\cdot),$
$I_{\psi}(\cdot)$ defined in \eqref{eq:MT-I_D,tau(r)=000020def}\emph{,}
and the convergence holds jointly with the convergence in Lemma \ref{lem:MT-y0limit}\emph{.}
\end{lem}
The following lemma is useful in determining the asymptotic properties
of $\widehat{\rho}_{n\tau}$ in the local-to-unity case.
\begin{lem}[\textbf{Convergence of Components in the Local-to-Unity Case}]
\label{lem:MT-asympdist_components_local=000020to=000020unity=000020case}
Under the null hypothesis $H_{0}:\rho_{n\tau}=\rho_{0,n}$, Assumptions
\ref{assu:MT-h} and \ref{assu:MT-kappa_n_sigma_n=000020mu_n=000020convergence},
and $nh/b_{n}\rightarrow r_{0}\in\left(0,\infty\right)$ $\left(\text{local-to-unity case}\right)$,
for a sequence $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$
and $\psi=r_{0}\kappa_{0}\left(\tau\right)$, the following results
hold jointly
\begin{enumerate}
\item $\left(nh\right)^{-1/2}Y_{t\left(s\right)}/\sigma_{0}\left(\tau\right)\Rightarrow I_{\psi}^{*}\left(s\right),$
where $t\left(s\right)=T_{1}+\lfloor nhs\rfloor$ for $s\in\left[0,1\right],$
\item $\left(nh\right)^{-3/2}\sum_{t=T_{1}}^{T_{2}}Y_{t-1}/\sigma_{0}\left(\tau\right)\rightarrow_{d}\int_{0}^{1}I_{\psi}^{*}\left(s\right)ds,$
\item $\left(nh\right)^{-2}\sum_{t=T_{1}}^{T_{2}}Y_{t-1}^{2}/\sigma_{0}^{2}\left(\tau\right)\rightarrow_{d}\int_{0}^{1}I_{\psi}^{*2}\left(s\right)ds,$
\item $\left(nh\right)^{-1/2}\sum_{t=T_{1}}^{T_{2}}U_{t}\sigma_{t}/\sigma_{0}\left(\tau\right)\rightarrow_{d}\int_{0}^{1}dB\left(s\right),$
\item $\left(nh\right)^{-1}\sum_{t=T_{1}}^{T_{2}}Y_{t-1}U_{t}\sigma_{t}/\sigma_{0}^{2}\left(\tau\right)\rightarrow_{d}\int_{0}^{1}I_{\psi}^{*}\left(s\right)dB\left(s\right)$,
\item $\left(nh\right)^{-2}\sum_{t=T_{1}}^{T_{2}}\left(Y_{t-1}^{0}/\sigma_{0}\left(\tau\right)\right)^{2}\rightarrow_{d}\int_{0}^{1}I_{\psi}^{2}\left(s\right)ds$,
and
\item when $r_{0}=0$, parts \textup{(a)--(c)} and \textup{(e)} hold with
$Y_{t\left(s\right)}$ \textup{(}$=\mu_{t\left(s\right)}+Y_{t\left(s\right)}^{0}+c_{t\left(s\right),t\left(s\right)-T_{0}}Y_{T_{0}}^{*}$\textup{)}
and $Y_{t-1}$ replaced by $\mu_{t\left(s\right)}+Y_{t\left(s\right)}^{0}$
and $\mu_{t-1}+Y_{t-1}^{0}$, respectively.
\end{enumerate}
\end{lem}
After proper re-scaling, we have
\begin{equation}
nh\left(\widehat{\rho}_{n\tau}-\rho_{0,n}\right)=\dfrac{\left(nh\right)^{-1}\sum_{t=T_{1}}^{T_{2}}\left(Y_{t-1}-\ol Y_{nh,-1}\right)\left(Y_{t}-\rho_{0,n}Y_{t-1}\right)}{\left(nh\right)^{-2}\sum_{t=T_{1}}^{T_{2}}\left(Y_{t-1}-\ol Y_{nh,-1}\right)^{2}}.\label{eq:MT-normalized=000020rho}
\end{equation}
The next theorem gives the asymptotic distribution of $nh\left(\widehat{\rho}_{n\tau}-\rho_{0,n}\right)$
and the t-statistic $T_{n}\left(\rho_{0,n}\right)$ in the local-to-unity
case.
\begin{thm}[\textbf{Asymptotic Distribution of Normalized $\boldsymbol{\widehat{\rho}_{n\tau}}$
and $\boldsymbol{t}$-Statistic in the Local-to-Unity Case}]
\label{thm:MT-asym_rho_hat_local=000020to=000020unity=000020case}Under
the null hypothesis $H_{0}:\rho_{n\tau}=\rho_{0,n}$, Assumptions
\ref{assu:MT-h}, \ref{assu:MT-stationary-nh^5=000020->=0000200},
and \ref{assu:MT-kappa_n_sigma_n=000020mu_n=000020convergence}, $nh/b_{n}\rightarrow r_{0}\in\left[0,\infty\right)$
$\left(\text{local-to-unity case}\right)$, and Assumption \ref{assu:MT-bn=000020div=000020nh=000020converges}
if $r_{0}=0$, for a sequence $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$
and $\psi=r_{0}\kappa_{0}\left(\tau\right)$, we have
\[
nh\left(\widehat{\rho}_{n\tau}-\rho_{0,n}\right)\rightarrow_{d}\left(\int_{0}^{1}I_{D,\psi}^{*2}\left(s\right)ds\right)^{-1}\int_{0}^{1}I_{D,\psi}^{*}\left(s\right)dB\left(s\right)
\]
and
\[
T_{n}\left(\rho_{0,n}\right)\rightarrow_{d}\left(\int_{0}^{1}I_{D,\psi}^{*2}\left(s\right)ds\right)^{-1/2}\int_{0}^{1}I_{D,\psi}^{*}\left(s\right)dB\left(s\right).
\]
\end{thm}
\begin{rem}
\label{rem:MT-subseq=000020loc=000020unity=000020thm}For any subsequence
$\{p_{n}\}_{n\geq1}$ of $\{n\}_{n\geq1}$, Lemmas \ref{lem:MT-Max=000020Intertemporal=000020Difference}--\ref{lem:MT-asympdist_components_local=000020to=000020unity=000020case}
and Theorem \ref{thm:MT-asym_rho_hat_local=000020to=000020unity=000020case}
hold with $p_{n}$ in place of $n$ throughout and $h_{p_{n}}$ in
place of $h=h_{n}.$
\end{rem}
\subsection{\protect\label{sec:MT-Stationary=000020Case}Asymptotics in the Stationary
Case}
The ``stationary'' case is characterized by $b_{n}=o\left(nh\right)$,
or equivalently, $nh/b_{n}\rightarrow r_{0}=\infty.$ The stationary
case is defined such that the t-statistic has a standard normal distribution
under $H_{0}:\rho_{n\tau}=\rho_{0,n}$ in the stationary case. If
$b_{n}$ is bounded, then $\rho_{n\tau}\leq C_{\rho}$ for some constant
$C_{\rho}<1$ for all $n\geq1$. If $b_{n}$ diverges to infinity,
then $\rho_{n\tau}$ goes to one at a rate slower than $1/nh$. Thus,
the stationary case includes some scenarios where $\rho_{n\tau}$
goes to one.
Define
\begin{equation}
\ol{\rho}_{n}:=\max\left\{ \exp\left\{ -\kappa_{0}\left(\tau\right)/\left(2b_{n}\right)\right\} ,-1+\varepsilon_{1}\right\} ,\label{eq:MT-def=000020rho^bar=000020ol}
\end{equation}
where $-1+\varepsilon_{1}$ is a lower bound on $\rho_{t}$ by the
definition of $\Lambda_{n}.$ We can bound $|\rho_{t}|$ for $t\in\left[T_{1},T_{2}\right]$
using $\ol{\rho}_{n}$. First, suppose $\rho_{t}\leq0$, then $-1+\varepsilon_{1}\leq\rho_{t}\leq0$,
which implies that $|\rho_{t}|\leq$$\ol{\rho}_{n}$, as desired.
Given this, in the following calculations we suppose $\rho_{t}\geq0$
for all $t\in[0,1]$ without loss of generality. Then, for $n$ sufficiently
large, we have
\begin{align}
\max_{t\in\left[T_{1},T_{2}\right]}\abs{\rho_{t}} & \leq\max_{t\in\left[T_{1},T_{2}\right]}\abs{\rho_{t}-\rho_{n\tau}}+\abs{\rho_{n\tau}-\exp\left\{ -\kappa_{0}\left(\tau\right)/b_{n}\right\} }+\exp\left\{ -\kappa_{0}\left(\tau\right)/b_{n}\right\} ,\nonumber \\
& \leq O\left(h/b_{n}\right)+o\left(1\right)/b_{n}+\exp\left\{ -\kappa_{0}\left(\tau\right)/b_{n}\right\} \leq\ol{\rho}_{n},\label{eq:MT-bound=000020max=000020rho_t=000020by=000020rho=000020bar}
\end{align}
where $b_{n}\geq\varepsilon_{3}>0$, the second inequality uses Lemma
\ref{lem:MT-Max=000020Intertemporal=000020Difference}(a), \eqref{eq:MT-rho_n},
Assumption \ref{assu:MT-kappa_n_sigma_n=000020mu_n=000020convergence},
and a mean value expansion of the $\exp\left\{ \cdot\right\} $ function,
and the last inequality holds using $\kappa_{0}(\tau)>0$ and the
fact that when $b_{n}\rightarrow\infty$, $\overline{\rho}_{n}-\exp\{-\kappa_{0}(\tau)/b_{n}\}\geq K/b_{n}$
for some constant $K>0.$
Equation \eqref{eq:MT-bound=000020max=000020rho_t=000020by=000020rho=000020bar}
implies
\begin{equation}
\abs{c_{t,j}}=\abs{\prod_{k=0}^{j-1}\rho_{t-k}}\leq\overline{\rho}_{n}^{j}\ \text{for }j=1,...,t-T_{1}+1\ \text{and }t=T_{1},...,T_{2}.\label{eq:MT-c_tj=000020bounded-gto_1}
\end{equation}
Furthermore, using Lemma A.1 of \citet*{GIRAITISKapetaniosYates2014rckernel}
and the definition of the parameter space $\Lambda_{n}$, under $H_{0}:\rho_{n\tau}=\rho_{0,n}$,
we have
\begin{equation}
\abs{c_{t,j}-\rho_{0,n}^{j}}\leq j\ol{\rho}_{n}^{j-1}\max_{k\in\left[T_{1},T_{2}\right]}\abs{\rho_{k}-\rho_{0,n}}=j\ol{\rho}_{n}^{j-1}O\left(h/b_{n}\right)\ \text{for }j=0,...,t-T_{0}\ \text{and }t=T_{1},...,T_{2},\label{eq:MT-ctj=000020-=000020rho^j=000020bound-gto_1}
\end{equation}
where the equality holds by Lemma \ref{lem:MT-Max=000020Intertemporal=000020Difference}(a).
The following results are used in the analysis:
\begin{align}
\sum_{t=0}^{\infty}\ol{\rho}_{n}^{t} & =\left(1-\ol{\rho}_{n}\right)^{-1}=O\left(b_{n}\right),\label{eq:MT-ps=000020rho^t}\\
\sum_{t=0}^{\infty}\ol{\rho}_{n}^{2t} & =\left(1-\ol{\rho}_{n}^{2}\right)^{-1}=O\left(b_{n}\right),\label{eq:MT-ps=000020rho^2t}\\
\sum_{t=0}^{\infty}t\ol{\rho}_{n}^{t} & =\ol{\rho}_{n}\left(1-\ol{\rho}_{n}\right)^{-2}=O\left(b_{n}^{2}\right),\label{eq:MT-ps=000020t*rho^t}\\
\sum_{t=0}^{\infty}t^{2}\ol{\rho}_{n}^{2t} & =\ol{\rho}_{n}^{2}\left(1+\ol{\rho}_{n}^{2}\right)\left(1-\ol{\rho}_{n}^{2}\right)^{-3}=O\left(b_{n}^{3}\right),\text{ and}\label{eq:MT-ps=000020t^2*rho^2t}\\
\sum_{t>s=0}^{nh}\ol{\rho}_{n}^{t-s} & =\sum_{t=1}^{nh}\sum_{s=0}^{t-1}\ol{\rho}_{n}^{t-s}=O\left(nhb_{n}\right).\label{eq:MT-ps=000020rho^(t-s)}
\end{align}
When $\rho_{0,n}=1-\kappa_{n}\left(\tau\right)/b_{n}$ with $b_{n}=o\left(nh\right)$
, one can replace $\ol{\rho}_{n}$ with $\rho_{0,n}$ in \eqref{eq:MT-ps=000020rho^t}--\eqref{eq:MT-ps=000020rho^(t-s)}
and the results still hold.
$\ $
We now establish the $N\left(0,1\right)$ asymptotic distribution
of $\widehat{\rho}_{n\tau}$ in the stationary case with the following
normalization:
\begin{align}
& \ \left(1-\rho_{0,n}^{2}\right)^{-1/2}\left(nh\right)^{1/2}\left(\widehat{\rho}_{n\tau}-\rho_{0,n}\right)\nonumber \\
= & \ \dfrac{\left(1-\rho_{0,n}^{2}\right)^{1/2}\left(nh\right)^{-1/2}\sum_{t=T_{1}}^{T_{2}}\left(Y_{t-1}-\ol Y_{nh,-1}\right)\left(Y_{t}-\rho_{0,n}Y_{t-1}\right)}{\left(1-\rho_{0,n}^{2}\right)\left(nh\right)^{-1}\sum_{t=T_{1}}^{T_{2}}\left(Y_{t-1}-\ol Y_{nh,-1}\right)^{2}}.\label{eq:MT-stationary=000020-=000020rho=000020hat-1}
\end{align}
Next, we prove a lemma on the asymptotic properties of the zero-initial
condition process $Y_{t}^{0}$ in the stationary case.
\begin{lem}[\textbf{Asymptotics in the Stationary Case}]
\label{lem:MT-Asymptotic=000020Properties=000020of=000020the=000020Components=000020Stationary}
Under the null hypothesis $H_{0}:\rho_{n\tau}=\rho_{0,n}$ and Assumptions
\ref{assu:MT-h} and \ref{assu:MT-kappa_n_sigma_n=000020mu_n=000020convergence},
for a sequence $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$
where $nh/b_{n}\rightarrow r_{0}=\infty$ and $\kappa_{0}\left(\tau\right)>0$
$\left(\text{stationary case}\right)$, the following results hold
jointly
\begin{enumerate}
\item $\left(1-\rho_{0,n}^{2}\right)^{1/2}\left(nh\right)^{-1}\sum_{t=T_{1}}^{T_{2}}Y_{t-1}^{0}/\sigma_{0}\left(\tau\right)\rightarrow_{p}0,$
\item $\left(1-\rho_{0,n}^{2}\right)\left(nh\right)^{-1}\sum_{t=T_{1}}^{T_{2}}\left(Y_{t-1}^{0}/\sigma_{0}\left(\tau\right)\right)^{2}\rightarrow_{p}1,$
and
\item $\left(1-\rho_{0,n}^{2}\right)^{1/2}\left(nh\right)^{-1/2}\sum_{t=T_{1}}^{T_{2}}Y_{t-1}^{0}\sigma_{t}U_{t}/\sigma_{0}^{2}\left(\tau\right)\rightarrow_{d} N\left(0,1\right),$
\end{enumerate}
where $Y_{t}^{0}$ is defined in \eqref{eq:MT-Y_t^0=000020and=000020Y_t^star}.
\end{lem}
Lemma \ref{lem:MT-Asymptotic=000020Properties=000020of=000020the=000020Components=000020Stationary}
concerns the zero-initial condition process $Y_{t}^{0}$ with time-varying
autoregressive parameter $\rho_{t}$, which is the foundation of our
asymptotic analysis of $\widehat{\rho}_{n\tau}$. Essentially Lemma
\ref{lem:MT-Asymptotic=000020Properties=000020of=000020the=000020Components=000020Stationary}
says that when $nh/b_{n}\rightarrow r_{0}=\infty$, the asymptotic distributions
of normalized sums of $Y_{t}^{0}$ behave in the same way as in an
AR process with constant $\rho<1$. The proof involves applications
of appropriate weak laws of large numbers and central limit theorems
for martingale difference triangular arrays and approximations of
TVPs. One can extend the results to allow the $U_{t}$ process to
be conditionally heteroskedastic, but this requires a different definition
of $\widehat{\sigma}_{n\tau}^{2}$ and complicates the analysis, see
\citet{andrews2014conditional}. For simplicity, we do not consider
this extension here.
$\ $
Define
\begin{equation}
\ol{\mu}_{nh}=\left(nh\right)^{-1}\sum_{t=T_{1}}^{T_{2}}\mu_{t}\ \text{and }\ol{\mu}_{nh,-1}=\left(nh\right)^{-1}\sum_{t=T_{1}}^{T_{2}}\mu_{t-1}.\label{eq:MT-def=000020-=000020olmu_nh}
\end{equation}
\begin{lem}[\textbf{Asymptotics of the Denominator in the Stationary Case}]
\label{lem:MT-asympdist_dist_stationary-denom} Under the null hypothesis
$H_{0}:\rho_{n\tau}=\rho_{0,n}$ and Assumptions \ref{assu:MT-h},
\ref{assu:MT-stationary-nh^5=000020->=0000200}, and \ref{assu:MT-kappa_n_sigma_n=000020mu_n=000020convergence},
for a sequence $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$
where $nh/b_{n}\rightarrow r_{0}=\infty$ and $\kappa_{0}\left(\tau\right)>0$
$\left(\text{stationary case}\right)$, the following results hold
\begin{enumerate}
\item $\left(1-\rho_{0,n}^{2}\right)\left(nh\right)^{-1}\sum_{t=T_{1}}^{T_{2}}\left(Y_{t-1}-\ol{\mu}_{nh,-1}\right)^{2}/\sigma_{0}^{2}\left(\tau\right)\rightarrow_{p}1$,
and
\item $\left(1-\rho_{0,n}^{2}\right)\left(\ol Y_{nh,-1}-\ol{\mu}_{nh,-1}\right)^{2}/\sigma_{0}^{2}\left(\tau\right)\rightarrow_{p}0$.
\end{enumerate}
\end{lem}
Lemma \ref{lem:MT-asympdist_dist_stationary-denom} concerns the asymptotic
distribution of the denominator of the normalized $\widehat{\rho}_{n\tau}$
in \eqref{eq:MT-stationary=000020-=000020rho=000020hat-1}. To bound
the difference between $Y_{t}^{0}$ and $Y_{t}$ and control for TVPs
asymptotically, we use various inequalities and approximations in
the proof of Lemma \ref{lem:MT-asympdist_dist_stationary-denom}.
\begin{lem}[\textbf{Asymptotics of the Numerator in the Stationary Case}]
\label{lem:MT-asympdist_dist_stationary-numerator=000020} Under
the null hypothesis $H_{0}:\rho_{n\tau}=\rho_{0,n}$ and Assumptions
\ref{assu:MT-h}, \ref{assu:MT-stationary-nh^5=000020->=0000200},
and \ref{assu:MT-kappa_n_sigma_n=000020mu_n=000020convergence}, for
a sequence $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$
where $nh/b_{n}\rightarrow r_{0}=\infty$ and $\kappa_{0}\left(\tau\right)>0$
$\left(\text{stationary case}\right)$, the following results hold
\begin{enumerate}
\item $\left(1-\rho_{0,n}^{2}\right)^{1/2}\left(nh\right)^{-1/2}\sum_{t=T_{1}}^{T_{2}}\left(Y_{t-1}-\ol{\mu}_{nh,-1}\right)\left[\begin{array}{c}
Y_{t}-\ol{\mu}_{nh}-\rho_{0,n}\left(Y_{t-1}-\ol{\mu}_{nh,-1}\right)\end{array}\right]/\sigma_{0}^{2}\left(\tau\right)$
$\rightarrow_{d} N\left(0,1\right)$, and
\item $\left(1-\rho_{0,n}^{2}\right)^{1/2}\left(nh\right)^{-1/2}\sum_{t=T_{1}}^{T_{2}}\left(\ol Y_{nh,-1}-\ol{\mu}_{nh,-1}\right)\left[\begin{array}{c}
Y_{t}-\ol{\mu}_{nh}-\rho_{0,n}\left(Y_{t-1}-\ol{\mu}_{nh,-1}\right)\end{array}\right]/\sigma_{0}^{2}\left(\tau\right)$
$\rightarrow_{p}0$.
\end{enumerate}
\end{lem}
Lemma \ref{lem:MT-asympdist_dist_stationary-numerator=000020} concerns
the asymptotic distribution of the numerator of the normalized $\widehat{\rho}_{n\tau}$
in \eqref{eq:MT-stationary=000020-=000020rho=000020hat-1}. The proof
of this lemma uses the results in Lemmas \ref{lem:MT-Order=000020of=000020Y_T0}(c)
and \ref{lem:MT-Asymptotic=000020Properties=000020of=000020the=000020Components=000020Stationary}.
$\ $
The next theorem provides the limit distributions of $\left(1-\rho_{0,n}^{2}\right)^{-1/2}\left(nh\right)^{1/2}\left(\widehat{\rho}_{n\tau}-\rho_{0,n}\right)$
and $T_{n}\left(\rho_{0,n}\right)$ in the stationary $\rho_{n\tau}$
case.
\begin{thm}[\textbf{Asymptotic Distribution of Normalized $\boldsymbol{\widehat{\rho}_{n\tau}}$
and $\boldsymbol{t}$-Statistic in the Stationary Case}]
\label{thm:MT-asympdist_dist_rho^hat-stationary} Under the null
hypothesis $H_{0}:\rho_{n\tau}=\rho_{0,n}$ and Assumptions \ref{assu:MT-h},
\ref{assu:MT-stationary-nh^5=000020->=0000200}, and \ref{assu:MT-kappa_n_sigma_n=000020mu_n=000020convergence},
for a sequence $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$
where $nh/b_{n}\rightarrow r_{0}=\infty$ and $\kappa_{0}\left(\tau\right)>0$
$\left(\text{stationary case}\right)$, we have
\[
\left(1-\rho_{0,n}^{2}\right)^{-1/2}\left(nh\right)^{1/2}\left(\widehat{\rho}_{n\tau}-\rho_{0,n}\right)\rightarrow_{d} N\left(0,1\right)
\]
and
\[
T_{n}\left(\rho_{0,n}\right)\rightarrow_{d} N\left(0,1\right).
\]
\end{thm}
\begin{rem}
\label{rem:MT-subseq=000020staty=000020case}For any subsequence $\{p_{n}\}_{n\geq1}$
of $\{n\}_{n\geq1}$, Lemmas \ref{lem:MT-Asymptotic=000020Properties=000020of=000020the=000020Components=000020Stationary}--\ref{lem:MT-asympdist_dist_stationary-numerator=000020}
and Theorem \ref{thm:MT-asympdist_dist_rho^hat-stationary} hold with
$p_{n}$ in place of $n$ throughout and $h_{p_{n}}$ in place of
$h=h_{n}.$
\end{rem}
\subsection{Asymptotic Results for \texorpdfstring{$\boldsymbol{\widehat{h}}$}{h}
\protect\label{subsec:MT-Asymptotic-Results-for-hhat}}
Here we give conditions under which $\widehat{h}$ defined in \eqref{eq:MT-h^hat=000020defn}
is asymptotically equivalent to the value $\widehat{h}_{opt}$ that
minimizes the ``empirical loss,'' which is unobserved. See \citet{li1987asymptotic}
and \citet{andrews1991asymptotic} for analogous results in i.i.d.
models. The empirical loss, $L_{n}(h),$ is
\begin{equation}
L_{n}(h):=n^{-1}\sum_{t=1}^{n}(\widehat{\mu}_{t-1}(h)-\mu_{t}+Y_{t-1}(\widehat{\rho}_{t-1}(h)-\rho_{t}))^{2}.\label{eq:MT-Empirical=000020loss=000020defn}
\end{equation}
We give a simple high-level condition under which
\begin{equation}
\frac{L_{n}(\widehat{h})}{L_{n}(\widehat{h}_{opt})}=\frac{L_{n}(\widehat{h})}{\inf_{h\in\mathcal{H}_{n}}L_{n}(h)}\rightarrow_{p}1.\label{eq:MT-Optimality=000020criterion}
\end{equation}
The $FE_{n}(h)$ criterion can be written as follows:
\begin{align}
FE_{n}(h)=\ & L_{n}(h)-2C_{n}(h)+E_{n},\text{ where }E_{n}:=n^{-1}\sum_{t=1}^{n}\sigma_{t}^{2}U_{t}^{2}\text{ and}\nonumber \\
C_{n}(h):=\ & n^{-1}\sum_{t=1}^{n}\sigma_{t}U_{t}(\widehat{\mu}_{t-1}(h)-\mu_{t}+Y_{t-1}(\widehat{\rho}_{t-1}(h)-\rho_{t})).\label{eq:MT-Decomp=000020os=000020FE=000020criterion}
\end{align}
The empirical loss $L_{n}(h)$ and the cross-product term $C_{n}(h)$
depend on $h,$ but the average squared error $E_{n}$ does not. Since
$E_{n}$ does not depend on $h,$ $\widehat{h}$ minimizes $L_{n}(h)-2C_{n}(h)$
over $\mathcal{H}_{n}.$
Under the following condition, $\widehat{h}$ is asymptotically equivalent
to the infeasible value $\widehat{h}_{opt}$ that minimizes $L_{n}(h).$
\begin{assumption}
\label{assu:MT-Asymp_h1}$\sup_{h\in\mathcal{H}_{n}}\frac{|C_{n}(h)|}{L_{n}(h)}\rightarrow_{p}0.$
\end{assumption}
\begin{lem}
\label{lem:MT-First=000020Data_Dependent=000020Bandwidth=000020Lem}Under
\emph{Assumption }\ref{assu:MT-Asymp_h1}, $\frac{L_{n}(\widehat{h})}{L_{n}(\widehat{h}_{opt})}\rightarrow_{p}1.$
\end{lem}
Now, we give a set of sufficient conditions for Assumption \ref{assu:MT-Asymp_h1}.
Let $h_{\min}=\min\{h:h\in\mathcal{H}_{n}\}$ and $h_{\max}=\max\{h:h\in\mathcal{H}_{n}\}.$
Typically, $h_{\min}$ and $h_{\max}$ depend on $n$ and decrease
to $0$ as $n\rightarrow\infty.$ Let $\xi_{n}$ denote the cardinality
of $\mathcal{H}_{n}.$ We decompose $L_{n}(h)$ and $C_{n}(h)$ into
the main components $L_{2n}(h)$ and $C_{2n}(h),$ respectively, which
depend on $t=nh_{\max}+1,...,n$ and for which $(\widehat{\mu}_{t-1}(h),\widehat{\rho}_{t-1}(h))$
depend only on random variables in $\mathcal{G}_{t-1},$ and the ``small
$t$'' boundary components $L_{1n}(h)$ and $C_{1n}(h),$ respectively,
which depend on $t=1,...,nh_{\max}$ for which $(\widehat{\mu}_{t-1}(h),\widehat{\rho}_{t-1}(h))$
depend on some random variables that are not in $\mathcal{G}_{t-1}.$
Define
\begin{align}
L_{1n}(h):=\ & n^{-1}\sum_{t=1}^{nh_{\max}}(\widehat{\mu}_{t-1}(h)-\mu_{t}+Y_{t-1}(\widehat{\rho}_{t-1}(h)-\rho_{t}))^{2},\text{ }\nonumber \\
L_{2n}(h):=\ & n^{-1}\sum_{t=nh_{\max}+1}^{n}(\widehat{\mu}_{t-1}(h)-\mu_{t}+Y_{t-1}(\widehat{\rho}_{t-1}(h)-\rho_{t}))^{2},\nonumber \\
C_{1n}(h):=\ & n^{-1}\sum_{t=1}^{nh_{\max}}\sigma_{t}U_{t}(\widehat{\mu}_{t-1}(h)-\mu_{t}+Y_{t-1}(\widehat{\rho}_{t-1}(h)-\rho_{t})),\text{ and}\nonumber \\
C_{2n}(h):=\ & n^{-1}\sum_{t=nh_{\max}+1}^{n}\sigma_{t}U_{t}(\widehat{\mu}_{t-1}(h)-\mu_{t}+Y_{t-1}(\widehat{\rho}_{t-1}(h)-\rho_{t})).\label{eq:MT-Partition=000020of=000020L=000020and=000020C}
\end{align}
The risk as a function of $h$ is $R_{n}(h)=E L_{n}(h).$ Let $R_{2n}(h)=E L_{2n}(h).$
The following assumption is sufficient for Assumption \ref{assu:MT-Asymp_h1}.
\begin{assumption}[\textbf{Sufficient Conditions for Assumption \ref{assu:MT-Asymp_h1}}]
\label{assu:MT-Assu_h2}
\end{assumption}
\begin{enumerate}
\item The true sequence of distributions is from the parameter spaces $\{\Lambda_{n}:n\geq1\}.$
\item $\sup_{h\in\mathcal{H}_{n}}\frac{|C_{1n}(h)|}{L_{n}(h)}\rightarrow_{p}0.$
\item $\sup_{h\in\mathcal{H}_{n}}|\frac{L_{2n}(h)}{R_{2n}(h)}-1|\rightarrow_{p}0.$
\item $h_{\max}\leq1-\varepsilon$ for $n$ large for some $\varepsilon>0.$
\item $\frac{n\inf_{h\in\mathcal{H}_{n}}R_{2n}(h)}{\xi_{n}}\rightarrow\infty.$
\end{enumerate}
Assumption \ref{assu:MT-Assu_h2}(b) implies that the ``small $t$''
boundary component of $C_{n}(h),$ which only depends on $t\leq nh_{\max},$
is asymptotically dominated by $L_{n}(h),$ which is based on all
$t\leq n.$ Assumption \ref{assu:MT-Assu_h2}(c) requires that the
variability of $L_{2n}(h)$ is small relative to its mean. In particular,
Assumption \ref{assu:MT-Assu_h2}(c) holds if $StdDev(L_{2n}(h))/E L_{2n}(h)=o(1).$
Assumption \ref{assu:MT-Assu_h2}(d) implies that the elements of
$\mathcal{H}_{n}$ are bounded away from one, which implies that $nh=n$
is not a feasible choice.
To interpret Assumption \ref{assu:MT-Assu_h2}(e), we give an intuitive
discussion of the magnitudes of $\inf_{h\in\mathcal{H}_{n}}L_{2n}(h)$
and $\inf_{h\in\mathcal{H}_{n}}R_{2n}(h).$ First, consider the case
where $\{Y_{t}\}_{t\leq n}$ displays unit root or local-to-unity
behavior across the whole time series. Then, $\widehat{\mu}_{t-1}(h)-\mu_{t}\thickapprox(nh)^{-1/2}$
and $Y_{t-1}(\widehat{\rho}_{t-1}(h)-\rho_{t})\thickapprox n^{1/2}(nh)^{-1}=(nh^{2})^{-1/2},$
$L_{2n}(h)\thickapprox(nh^{2})^{-1},$ and under suitable conditions,
$R_{2n}(h)\thickapprox(nh^{2})^{-1}.$ Second, in the case where $\{Y_{t}\}_{t\leq n}$
displays stationary behavior across the whole time series, $\widehat{\mu}_{t-1}(h)-\mu_{t}\thickapprox(nh)^{-1/2},$
$Y_{t-1}(\widehat{\rho}_{t-1}(h)-\rho_{t})\thickapprox(nh)^{-1/2},$
$L_{2n}(h)\thickapprox(nh)^{-1},$ and under suitable conditions $R_{2n}(h)\thickapprox(nh)^{-1}.$
Third, in the case where $\{Y_{t}\}_{t\leq n}$ displays behavior
that varies between unit root and stationarity across the time series,
the order of magnitude of $R_{2n}(h)$ is between $(nh^{2})^{-1}$
and $(nh)^{-1},$ which is bounded below by $(nh_{\max})^{-1}.$ Hence,
Assumption \ref{assu:MT-Assu_h2}(e) requires $n\cdot(nh_{\max})^{-1}/\xi_{n}\rightarrow\infty$
or $\xi_{n}=o(h_{\max}^{-1}).$ That is, the number $\xi_{n}$ of
values $h$ in $\mathcal{H}_{n}$ needs to be of smaller order than
the reciprocal of the maximum value $h_{\max}$ in $\mathcal{H}_{n}.$
\begin{lem}
\label{lem:MT-Second=000020Data_Dependent=000020Bandwidth=000020Lem}\emph{Assumption
}\ref{assu:MT-Assu_h2}\emph{ }implies \emph{Assumption }\ref{assu:MT-Asymp_h1}\emph{.}
\end{lem}
\begin{rem}
\noindent The proof of Lemma \ref{lem:MT-Second=000020Data_Dependent=000020Bandwidth=000020Lem}
shows that $E C_{2n}(h)=0$ and $Var(C_{2n}(h))\leq C_{3U}\times(n-nh)^{-1}E L_{2n}(h)$
$\forall h\in\mathcal{H}_{n},n\geq1,$ where $C_{3U}<\infty$ is the
bound on the variance function $\sigma^{2}$ in the definition of
$\Lambda_{n}.$
\end{rem}
\begin{rem}
\noindent Let $R_{1n}(h)=E L_{1n}(h).$ Under Assumption \ref{assu:MT-Assu_h2}
and $\sup_{h\in\mathcal{H}_{n}}\frac{R_{1n}(h)+E|C_{1n}(h)|}{R_{n}(h)}\rightarrow0,$
we also have: $\frac{R_{n}(\widehat{h})}{R_{n}(\widehat{h}_{opt})}\rightarrow_{p}1.$
\end{rem}
\newpage{}