EconBase
← Back to paper

Inference in a Stationary/Nonstationary Autoregressive Time-Varying-Parameter Model

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

109,893 characters · 20 sections · 25 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference in a Stationary/Nonstationary Autoregressive Time-Varying-Parameter Model

\global\long \global\long \makeatletter \renewenvironment{proof}[1][\proofname]{\pushQED{\ \qedsymbol}\normalfont\topsep6\p@\@plus6\p@\relax\trivlist\ignorespaces}{\popQED\endtrivlist\@endpefalse} \makeatother

\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\def\mc#1{\mathscr{#1}} \global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\def\abs#1{\left|#1\right|} \global\long\def\norm#1{\left\Vert #1\right\Vert } \global\long\def\rest#1{\left.#1\right|} \global\long\def\bracket#1#2{\left\langle #1\middle\vert#2\right\rangle } \global\long\def\sandvich#1#2#3{\left\langle #1\middle\vert#2\middle\vert#3\right\rangle } \global\long\def\turd#1{\frac{#1}{3}} \global\long\global\long\def\sand#1{\left\lceil #1\right\vert } \global\long\def\wich#1{\left\vert #1\right\rfloor } \global\long\def\sandwich#1#2#3{\left\lceil #1\middle\vert#2\middle\vert#3\right\rfloor } \global\long\def\abs#1{\left|#1\right|} \global\long\def\norm#1{\left\Vert #1\right\Vert } \global\long\def\rest#1{\left.#1\right|} \global\long\def\inprod#1{\left\langle #1\right\rangle } \global\long\def\ol#1{\overline{#1}} \global\long\def\ul#1{#1} \global\long\def\td#1{\tilde{#1}} \global\long\def\bs#1{\boldsymbol{#1}} \global\long\global\long\global\long\global\long\global\long\global\long\global\long

titlepage\thispagestyle{empty} \begin{abstract} This paper considers nonparametric estimation and inference in first-order autoregressive (AR(1)) models with deterministically time-varying parameters. A key feature of the proposed approach is to allow for time-varying stationarity in some time periods, time-varying nonstationarity (i.e., unit root or local-to-unit root behavior) in other periods, and smooth transitions between the two. The estimation of the AR parameter at any time point is based on a local least squares regression method, where the relevant initial condition is endogenous. We obtain limit distributions for the AR parameter estimator and t-statistic at a given point $\tau$ in time when the parameter exhibits unit root, local-to-unity, or stationary/stationary-like behavior at time $\tau.$ These results are used to construct confidence intervals and median-unbiased interval estimators for the AR parameter at any specified point in time. The confidence intervals have correct asymptotic coverage probabilities with the coverage holding uniformly over stationary and nonstationary behavior of the observations. \noindentKeywords: Autoregressive time-varying-parameter model, endogenous initial condition, nonparametric estimation, confidence interval. \end{abstract}

\protectIntroduction

Autoregressive models --- stationary or nonstationary --- are workhorse models in econometric time series. In this paper, we consider a deterministically time-varying parameter (TVP) autoregressive model that allows for stationary and non-stationary behavior at different points in the time period of interest. Thus, the level of persistence of the time series can change over the time period. The motivation for considering such a model is that the economy is in continual transition due to technological, institutional, political, and demographic changes. The model considered allows for the intercept and error variance of the model also to be time-varying, not just the AR parameter.

The estimation method employed is local least squares, which depends on a tuning parameter $h$ that determines the local neighborhood that is considered. We construct a confidence interval (CI) for the value of the AR parameter at a time point $\tau$ via the inversion of tests, as is common in the literature for constant parameter AR models, e.g., see \citet*{stock1991confidence,andrews1993exactly,hansen1999grid,mikusheva2007uniform}, and andrews2014conditional. We show that the CI's have correct uniform asymptotic coverage probability for a parameter space that allows the time series to be stationary in parts of the time period and nonstationary in other parts, using the approach in \citet*{andrews2020generic}. We also construct asymptotically median-unbiased interval estimators (MUE's) in an analogous fashion.

In the TVP case, the initial condition is endogenous due to the choice of the local neighborhood and depends on the potentially different behavior of the time series prior to the local neighborhood. For a given time point $\tau$ of interest, we find that the asymptotic distributions of the LS estimator and t-statistic depend on the endogenous initial condition in the local-to-unity case. This is analogous to the asymptotic effect of the initial condition--under certain assumptions on the initial condition--in constant parameter AR models, see \citet*{elliott1999efficient,elliott2001confidence}, and muller2003tests. On the other hand, the endogenous initial condition does not affect the asymptotic distributions when the AR coefficient at time $\tau$ is more distant from one than local-to-unity.

We note that whether a time series is local-to-unity at time $\tau,$ or not, depends on $\tau$ and the chosen bandwidth $h.$ It is important to provide asymptotic results that hold uniformly over a parameter space that does not depend on $h,$ which we do.

We introduce a method for determining the tuning parameter $h$ based on a forecast-error criterion. We provide conditions under which this data-dependent choice of $h$ is asymptotically equivalent to an infeasible choice that minimizes the unobserved “empirical loss.” These results are similar to results for i.i.d. models given in li1987asymptotic and andrews1991asymptotic.

We provide Monte Carlo simulation results for the methods introduced in the paper. We consider true autoregressive functions whose shapes are sinusoidal, linear, partly flat/partly linear, flat, and kinked linear. We consider cases where the functions are close, or equal, to one in some regions, but different from one by varying amounts in other regions. We find that the proposed CI has reasonably good coverage probabilities and short average lengths for most of the data-generating processes considered. For example, nominal 95% CI's are found to have finite-sample coverage probabilities ranging from 92.5% to 96% in 88.8% of the cases, across 205 cases. The lowest coverage probabilities, in the range from 87.5% to 90%, occur only in 2.0% of the cases. The MUE's are found to have very small finite-sample median bias across the different cases considered. The magnitude of the data-dependent choice of $h$ varies widely depending upon the shape of the autoregressive function, as desired.

We provide some empirical applications of the methods to monthly inflation and real exchange rates in several countries using data from the IMF International Financial Statistics database. We find that the inflation series exhibit noticeable time variation of the AR parameter across the time period considered. We discuss how this relates to the literature on the persistence properties of inflation. On the other hand, we find that the real exchange rate series have nearly constant AR values across time that are equal to, or close to, one. Hence, the TVP methods are capable of producing constant AR values when it is appropriate to do so. We discuss how these results relate to the literature on the persistence of exchange rates. In Section (ref) of the Supplemental Material, we also report results for interest rates for several countries and results for a number of US macroeconomic time series using the Federal Reserve Economic Database (FRED).

The results of the paper apply to a TVP-AR(1) model. For some time series, a TVP-AR(p) model with $p>1$ may be more appropriate than a TVP-AR(1) model. The methods introduced in the paper can be extended to a TVP-AR(p) model, see Section (ref) of the Supplemental Material for details.

Relative to the literature, the contribution of this paper is to develop methods for a deterministically time-varying parameter (TVP) AR(1) model that allows time-varying stationarity in some time periods and time-varying nonstationarity in other periods. The resulting model is much more flexible than a constant parameter model. No paper in the existing literature does this. The methods we employ are quite similar to those employed in constant parameter AR(1) models that impose stationarity or nonstationarity across the whole time period, such as those referenced above. In particular, we invert tests of null hypotheses concerning the AR coefficient (at a particular point in time) and utilize the nonstandard asymptotic distributions of the t-statistics for such null hypotheses to obtain critical values. Our results differ from those obtained for constant parameter AR(1) models in that (i) we consider a local neighborhood of the time point of interest $\tau,$ indexed by a bandwidth parameter $h,$ and within that time period the AR coefficient changes with time, which causes biases that have to be accounted for, (ii) the initial condition of the local neighborhood depends on the choice of the local neighborhood and on the past behavior of the TVP-AR(1) process, which may differ from its current behavior, which needs to be accounted for, and (iii) we use the data to select a suitable local neighborhood using a forecast-error criterion function. None of these features arise in a constant parameter model.

The literature contains numerous papers that consider time-varying parameter AR models. Some of these papers consider deterministic TVP's, as in this paper, but they do not allow for stationarity in some time periods and nonstationarity in others. References for stationary (i.e., short-range dependent) TVP AR models include \citet*{Rao1970,Grenier1983,dahlhaus1996kullback,Dahlhaus1997}, Dahlhaus1998, \citet*{dahlhaus1999nonlinear}, \citet*{Moulines2005AoS}, \citet*{PhillipsXu2008Adaptive}, \citet*{ding2017ejs}, vanDelft2018EJS, and \citet*{karmakar2022simultaneous}. References for nonstationary (i.e., long-range dependent) TVP AR models include \citet*{bykhovskaya2018boundary,bykhovskaya2020point}, which focus on tests of a unit root null hypothesis against functional local-to-unity alternatives, as opposed to CI's for an AR parameter, which is the focus of this paper. No papers in the literature consider CI's for an AR parameter in nonstationary TVP AR models.

The literature also includes papers on random coefficient (RC) AR models and functional coefficient (FC) AR models (in which the AR coefficient depends on observable variables). References for stationary RC AR models includes NichollsQuinn1980, NichollsQuinn1981, \citet*{doan1984forecasting}, and cogley2005drifts. References for papers on RC AR models with random coefficients that follow a nonstationary process include follmer1993microeconomic, \citet*{GIRAITISKapetaniosYates2014rckernel,GiraitisKapetaniosYates2018}, and \citet*{tao2019random}. A reference for stationary FC AR models is \citet*{CaiFanYao2000}. References for nonstationary FC AR models includes Juhl2005, Lieberman2012, and \citet*{lieberman2014norming,lieberman2017multivariate,lieberman2018iv}.

This paper is organized as follows. Section (ref) introduces the TVP-AR(1) model. Section (ref) introduces the CI and MUE for the AR parameter at time $\tau.$ Section (ref) introduces the data-dependent method for choosing the bandwidth parameter $h$ based on a forecast-error criterion. Section (ref) presents the Monte Carlo simulation results. Section (ref) presents the empirical results. Section (ref) shows that the CI has correct uniform asymptotic coverage probability, the MUE is asymptotically median unbiased, and the data-dependent method for choosing the bandwidth parameter $h$ has some some desirable asymptotic properties. It also provides the asymptotic behavior of the local least squares estimator and t-statistic under a variety of drifting sequences of distributions, which are used in the proof of the uniform asymptotic coverage probability results and the asymptotic median unbiasedness results. The Supplemental Material includes the critical values of the limiting distribution of our t-statistic in Section (ref), the proofs of the results of the paper in Section (ref), additional simulation results in Section (ref), a description of how to extend the methods to TVP-AR(p) models for $p>1$ in Section (ref), and information about the empirical applications and additional empirical results in Section (ref).

All limits in this paper are as $n\rightarrow\infty.$ For notational simplicity, but with some abuse of notation, we let $\cdot/nh$ denote $\cdot/(nh)$ throughout the paper.

\protectModel

\numberwithin{equation}{section}

The TVP-AR(1) model we consider is

align[align omitted — 160 chars of source]

where $\rho_{t}\in\left[-1+\varepsilon_{1},1\right]$ for some $0<\varepsilon_{1}<2$. The autoregressive parameter $\rho_{t}$ is allowed to vary with time $t.$ A key feature of the model is that it allows for stationary, unit root, or local-to-unity behavior at different points in time. The errors $\left\{ U_{t}:t=0,1,...,n\right\} $ are a stationary martingale difference sequence under $F$ with $E_{F}\left(\rest{U_{t}}\mathscr{G}_{t-1}\right)=0\ \text{a.s.}$, $E_{F}\left(\rest{U_{t}^{2}}\mathscr{G}_{t-1}\right)=1\ \text{a.s.}$, and $E_{F}\left(\rest{U_{t}^{4}}\mathscr{G}_{t-1}\right)<M\ \text{a.s.}$ for some $M\in\left(0,\infty\right)$, where $\mathscr{G}_{t}$ is some non-decreasing sequence of $\sigma$-fields for which $\sigma\left(U_{0},...,U_{t},Y_{0}^{*}\right)\subseteq\mathscr{G}_{t}$ for $t=1,...,n$.

We assume $\rho_{t},$ $\mu_{t},$ and $\sigma_{t}^{2}$ satisfy

equation[equation omitted — 158 chars of source]

respectively, where $\rho\left(\cdot\right)\ $is a twice continuously differentiable function on $\left[0,1\right]$ and $\mu\left(\cdot\right)\ $and $\sigma^{2}\left(\cdot\right)$ are Lipschitz functions on $\left[0,1\right]$. Given (ref), $Y_{t}$, $Y_{t}^{*}$, $\rho_{t}$, $\mu_{t}$, and $\sigma_{t}$ depend implicitly on $n$.

Let $\tau\in\left(0,1\right)$. We consider estimation and inference concerning

equation[equation omitted — 56 chars of source]

which is the value of the autoregressive function $\rho(\cdot)$ at the $\tau$ fraction of the way through the sample.

For ease of reading, the definition of the parameter space of functions $\rho\left(\cdot\right)$, $\mu\left(\cdot\right),\ $and $\sigma^{2}\left(\cdot\right)$ that is considered, which includes some structure on the $\rho\left(\cdot\right)$ and $\mu\left(\cdot\right)\ $functions, is given in Section (ref) below.

Confidence Interval for the Autoregressive \\ Parameter \texorpdfstring{$\boldsymbol{\rho(\tau)}$}{rho(tau)}

Local Least Squares Estimator of \texorpdfstring{$\boldsymbol{\rho(\tau)}$}{rho(tau)}

We employ a LS estimator of $\rho(\tau)$ based on the time periods $t=T_{1},...,T_{2}$, where

equation[equation omitted — 147 chars of source]

for a bandwidth parameter $h$. The total number of periods in $\left[T_{1},T_{2}\right]$ is within one of $nh$. For this range of time periods, the initial condition time period is $T_{0}:=T_{1}-1.$

For the asymptotic results given below the bandwidth $h$ satisfies $h\rightarrow0$ and $nh\rightarrow\infty$ as $n\rightarrow\infty$. In Section (ref) below, we introduce a data-dependent bandwidth that is smaller or larger depending on how wiggly or flat the true $\rho\left(\cdot\right)\ $ and $\mu\left(\cdot\right)\ $ functions are.

Define

equation[equation omitted — 162 chars of source]

To estimate $\rho\left(\tau\right)$, we regress $Y_{t}$ on a constant and $Y_{t-1}.$ The resulting local LS estimator $\widehat{\rho}_{n\tau}$ is

equation[equation omitted — 216 chars of source]

Confidence Interval for \texorpdfstring{$\boldsymbol{\rho(\tau)}$}{rho(tau)}

The CI for $\rho\left(\tau\right)$ that we consider is obtained by inverting tests of null hypotheses of the form $H_{0}:\rho\left(\tau\right)=\rho_{0}$ for different values $\rho_{0}\in\left[-1+\varepsilon_{1},1\right]$.

The estimator of the time-varying variance $\sigma^{2}\left(\cdot\right)$ at $t/n=\tau$ is defined to be

equation[equation omitted — 226 chars of source]

For arbitrary $\rho_{0}\in(-1,1],$ the t-statistic that is used to construct the CI for $\rho(\tau)$ is

equation[equation omitted — 333 chars of source]

Let $B(\cdot)$ denote a standard Brownian motion on $[0,1].$ Let $Z_{1}$ be a standard normal random variable that is independent of $B\left(\cdot\right)$. Define

align[align omitted — 481 chars of source]

The stochastic process $I_{\psi}\left(s\right)$ is an Ornstein-Uhlenbeck process on $\left[0,1\right]$ with parameter $\psi.$

Below we consider sequences of functions $\left\{ \rho_{n}\left(\cdot\right),\mu_{n}\left(\cdot\right),\sigma_{n}\left(\cdot\right)\right\} _{n\geq1}$ and null hypotheses $H_{0}:\rho_{n}\left(\tau\right)=\rho_{0,n},$ where the null hypotheses values $\rho_{0,n}$ depend on $n$ for $n\geq1.$ For suitable sequences $\left\{ \rho_{0,n}\right\} _{n\geq1}$, we show that under $H_{0},$

equation[equation omitted — 159 chars of source]

where $T_{n}\left(\rho_{0,n}\right)$ is defined with $\rho_{0,n}$ in place of $\rho_{0}$ in (ref), $\psi$ depends on the sequence $\left\{ \rho_{0,n}\right\} _{n\geq1},$ and $J_{\psi}$ is defined as follows. For $\psi=\infty$, which corresponds to a “stationary” sequence $\left\{ \rho_{0,n}\right\} _{n\geq1}$, $J_{\psi}$ has a $N(0,1)$ distribution. For $\psi\in\left[0,\infty\right)$, which corresponds to a local-to-unity or unit root sequence $\left\{ \rho_{0,n}\right\} _{n\geq1}$,

equation[equation omitted — 177 chars of source]

For $\alpha\in(0,1),$ let $c_{\psi}\left(\alpha\right)$ denote the $\alpha$ quantile of the distribution of $J_{\psi}$. For given $\alpha$, we compute $c_{\psi}\left(\alpha\right)$, the $\alpha$-quantile of $J_{\psi}$ in (ref), by simulating the asymptotic distribution $J_{\psi}$. To do so, $B=300,000$ independent constant coefficient AR(1) sequences are generated with innovations $U_{t}\sim_{iid}N\left(0,1\right)$, stationary start-up, $n=25,000$, and $\rho=\exp\left(-\psi/n\right)$. For each sequence, the test statistic $T_{n}\left(\rho\right)$ defined in (ref) is calculated. Then the simulated estimate of $c_{\psi}\left(\alpha\right)$ is the $\alpha$-quantile of the empirical distribution of the $B=300,000$ realizations of the test statistic $T_{n}\left(\rho\right)$.

The nominal $1-\alpha$ equal-tailed two-sided CI for $\rho(\tau)$ is

align[align omitted — 400 chars of source]

The CI $CI_{n,\tau}$ can be computed by taking a fine grid of values $\rho_{0}\in\left[-1+\varepsilon_{1},1\right]$ and comparing $T_{n}\left(\rho_{0}\right)$ to $c_{\psi_{nh,\rho_{0}}}\left(\alpha/2\right)$ and $c_{\psi_{nh,\rho_{0}}}\left(1-\alpha/2\right).$ Table (ref) provides the critical values $c_{\psi_{nh,\rho_{0}}}\left(\alpha/2\right)$ and $c_{\psi_{nh,\rho_{0}}}\left(1-\alpha/2\right)$ for $\alpha=.05$, and .1. Given these critical values, computation of equal-tailed two-sided 90% and 95% CI's for $\rho(\tau)$ is fast.

The correct asymptotic size and asymptotic similarity of the CI $CI_{n,\tau}$ are established in Theorem (ref) in Section (ref) below.

Median-Unbiased Interval Estimator of \texorpdfstring{$\boldsymbol{\rho(\tau)}$}{rho(tau)}

By definition, an estimator $\widehat{\theta}_{n}$ of a parameter $\theta$ is median unbiased if $P\left(\widehat{\theta}_{n}\geq\theta\right)\geq1/2$ and $P\left(\widehat{\theta}_{n}\leq\theta\right)\geq1/2.$ Here we introduce a median-unbiased interval estimator of $\rho\left(\tau\right)$ that satisfies an analogous condition. Also, with probability close to one, the estimator is a single point.\footnote{If the .5 quantile of the asymptotic null distribution of the t-statistic was strictly decreasing in $\rho_{0},$ or equivalently, $c_{\psi}\left(.5\right)$ was strictly increasing in $\psi$, then the proposed interval estimator would be a point estimator with probability one. Because this condition fails to hold exactly, but almost holds, there is a very small probability that the estimator is a short interval, rather than a point.} Let $CI_{n,\tau}^{up}\left(.5\right)$ and $CI_{n,\tau}^{low}\left(.5\right)$ denote level $.5$ one-sided upper-bound and lower-bound CIs for $\rho\left(\tau\right)$, respectively. By definition,

align[align omitted — 486 chars of source]

The median-unbiased interval estimator $\widetilde{\rho}_{n\tau}$ of $\rho\left(\tau\right)$ is defined by

align[align omitted — 399 chars of source]

By construction, $\widetilde{\rho}_{n\tau,low}\leq\widetilde{\rho}_{n\tau,up}.$\footnote{This holds because $\widetilde{\rho}{}_{n\tau,up}\geq\sup\left\{ \rho_{0}\in\left[-1+\varepsilon_{1},1\right]:c_{\psi_{nh,\rho_{0}}}\left(.5\right)=T_{n}\left(\rho_{0}\right)\right\} $ and $\widetilde{\rho}{}_{n\tau,low}$ is less than or equal to the infimum of the values in the same set. } In addition, $\widetilde{\rho}_{n\tau}$ is a singleton whenever the set $\left\{ \rho_{0}\in\left[-1+\varepsilon_{1},1\right]:T_{n}\left(\rho_{0}\right)=c_{\psi_{nh,\rho_{0}}}\left(.5\right)\right\} $ contains a single point, in which case $\widetilde{\rho}_{n\tau}$ equals this point. Simulations show that $\widetilde{\rho}_{n\tau}$ is a singleton with probability close to one and a very short interval when it is not a singleton. Table (ref) provides the critical values $c_{\psi_{nh,\rho_{0}}}\left(.5\right)$ for a wide range of $\psi$. Given these critical values, computation of $\widetilde{\rho}_{n\tau}$ is fast.

The estimator $\widetilde{\rho}_{n\tau}$ has the following median-unbiasedness property

align[align omitted — 299 chars of source]

Furthermore, as shown in Corollary (ref) in Section (ref) below, these asymptotic properties hold in a uniform sense over the parameter space. This property is important for ensuring that the asymptotic properties of $\widetilde{\rho}_{n\tau}$ reflect its finite-sample properties.

Data-dependent Bandwidth Parameter \texorpdfstring{$\boldsymbol{h}$}{h}

The CI for $\rho\left(\tau\right)$ proposed in Section (ref) requires a tuning parameter $h$ that determines the interval of observations used to construct the CI. This section introduces a data-dependent method of selecting $h$ that minimizes a forecast-error criterion. We present its asymptotic properties in Section (ref) under somewhat high-level conditions.

The forecast-error criterion is the average over $t=1,...,n$ of the squared errors from forecasting $Y_{t}$ using the predicted value $\widehat{Y}_{t}$. Specifically, for $t>\lfloor nh\rfloor,$ let $(\widehat{\mu}_{t-1}(h),\widehat{\rho}_{t-1}(h))$ denote the LS estimator of $(\mu_{t},\rho_{t})$ from the regression of $Y_{s}$ on a constant and $Y_{s-1}$ using the $nh$ observations $s=t-\lfloor nh\rfloor,...,t-1.$ For $t\leq\lfloor nh\rfloor,$ since there are fewer than $\lfloor nh\rfloor$ observations, $(\mu_{t},\rho_{t})$ is estimated using the LS estimator $(\widehat{\mu}_{t-1}(h),\widehat{\rho}_{t-1}(h))$ based on the $\lfloor nh\rfloor$ observations $s=1,...,t-1,t+1,...,\lfloor nh\rfloor+1.$ The observation $Y_{t}$ is predicted by $\widehat{\mu}_{t-1}(h)+Y_{t-1}\widehat{\rho}_{t-1}(h)$ for $t=1,...,n$. The forecast-error criterion is

equation[equation omitted — 158 chars of source]

We assume that $h$ is chosen from a finite set $\mathcal{H}_{n},$ which depends on $n$ and whose cardinality may depend on $n.$ By definition, the data-dependent choice of $h,$ $\widehat{h},$ minimizes $FE_{n}(h)$ over $\mathcal{H}_{n}$: \footnote{If the argmin is not unique, for specificity $\widehat{h},$ is taken to be the largest argmin. But, non-uniqueness occurs with probability zero provided the innovations have a continuous component.}

equation[equation omitted — 106 chars of source]
remThe criterion $FE_{n}(h)$ has the desirable property of being an average of unbiased risk estimators for $t>\lfloor nh\rfloor.$ This holds because \begin{eqnarray} & & E(Y_{t}-\widehat{\mu}_{t-1}(h)-Y_{t-1}\widehat{\rho}_{t-1}(h))^{2}\nonumber \\ & = & E(U_{t}-(\widehat{\mu}_{t-1}(h)-\mu_{t})+Y_{t-1}(\widehat{\rho}_{t-1}(h)-\rho_{t})))^{2}\nonumber \\ & = & E U_{t}^{2}-2E U_{t}(\widehat{\mu}_{t-1}(h)-\mu_{t}+Y_{t-1}(\widehat{\rho}_{t-1}(h)-\rho_{t}))+E(\widehat{\mu}_{t-1}(h)-\mu_{t}+Y_{t-1}(\widehat{\rho}_{t-1}(h)-\rho_{t}))^{2}\nonumber \\ & = & E U_{t}^{2}+E(\widehat{\mu}_{t-1}(h)-\mu_{t}+Y_{t-1}(\widehat{\rho}_{t-1}(h)-\rho_{t}))^{2}, \end{eqnarray} where the third equality holds for $t>\lfloor nh\rfloor$ since $(\widehat{\mu}_{t-1}(h),\widehat{\rho}_{t-1}(h))$ is a function of the innovations indexed by $s\leq t-1$ and $E(U_{t}|U_{t-1},...,U_{1},Y_{0}^{\ast})=0.$ For the relatively small number of initial observations with $t\leq\lfloor nh\rfloor,$ the third equality does not hold because $\widehat{\rho}_{t-1}(h)$ depends on observations $s$ with $s\geq t.$ In terms of computation, the criterion $FE_{n}(h)$ has the advantage that the same $\widehat{h}$ is employed for multiple $\lfloor nh\rfloor$ values of interest. It also avoids introducing additional tuning parameters, such as the length of a forecasting period.
remIn Section (ref), we show that the data-dependent $\widehat{h}$ value achieves the asymptotically optimal trade-off between bias and variance obtained by the infeasible $\widehat{h}_{opt}$ which minimizes the “empirical loss” as a function of $h$ defined in (ref) below. In consequence, undersmoothing $\widehat{h}$ should yield a value of $h$ for which the bias is dominated by the variance asymptotically, the CI's defined above have correct asymptotic size, and the median-unbiased interval estimator is asymptotically median unbiased. We employ a relatively sophisticated undersmoothing method. The objective of undersmoothing (i.e., making $\widehat{h}_{us}$ smaller than $\widehat{h}$) is to reduce the bias. The cost of undersmoothing is an increase in variance. We want to undersmooth more when the cost of doing so in terms of increasing the variance is relatively small, and we want to undersmooth less when the cost is larger. This means that we need to take account of the shape of the $FE_{n}(h)$ function. When it is flatter at its minimum, we want to undersmooth more. This leads to the following definition: \begin{equation} \widehat{h}_{us0}=\min\left\{ h\mathcal{\in H}_{n}:FE_{n}(h)\leq Q_{c_{1}}(FE_{n})\right\} , \end{equation} where $Q_{c_{1}}(FE_{n})$ is the $c_{1}$ quantile of $FE_{n}$ among $h\mathcal{\in H}_{n}$ and we take $c_{1}=.2$ in the simulations and applications. This definition does not guarantee that $\widehat{h}_{us0}$ is of a smaller order than $\widehat{h}$, which is necessary to ensure proper asymptotics. In consequence, we modify the definition as follows: \begin{equation} \widehat{h}_{us}=\min\left\{ \widehat{h}_{us0},\widehat{h}_{us1}\right\} , \end{equation} where $\widehat{h}_{us1}=c_{2}n^{-a}\widehat{h}$ for $c_{2},a>0.$ In the simulations and applications, we use $c_{2}=1.5$ and $a=1/10.$

Monte Carlo Simulations

In this section, we analyze the finite-sample performance of the methods introduced above using Monte Carlo simulations. First, we describe the data generating processes (DGP's) considered and how the CI and MUE are implemented. Then, we report the results. We find that the proposed CI has reasonably good coverage probabilities and short average lengths for most of the DGP's considered. We also find that the MUE performs quite well both when the true DGP of $\rho_{t}$ is time-varying and flat.

Simulation Setup and Methodology

We consider 21 DGP's for $\rho_{t}$. The graphs of the $\rho_{t}$ functions are given in Figures (ref)--(ref) and Figures (ref)--(ref), which appear in the Supplemental Material. There are five categories of $\rho_{t}$ functions: sinusoidal, linear, flat linear, flat, and kinked linear. For the sinusoidal functions, we consider sinusoidal functions for $\rho_{t}$ with 6 different ranges and shapes. For example, “sin 1.00-0.90-1.00” corresponds to the sinusoidal function where $\rho_{t}$ first achieves its maximum of 1.00 at $t/n=.20$, then drops to 0.90 at $t/n=.60$ and finally increases to 1.00 at $t/n=1$, with a frequency of 2.5$\pi$. For linear functions, we consider 4 different linear DGP's for $\rho_{t}$ corresponding to different values of the slopes and intercepts. For flat linear functions, the first half of $\rho_{t}$ is flat and the second half is linear, and 4 different cases are considered. For example, “flat-lin 0.90-0.99” means \[ \rho_{t}=\rho(t/n)=

cases.90 & for t/n\in\left[0,1/2\right],\\ .18t/n+.81 & for t/n\in(1/2,1].

\] For the flat functions, we consider 3 values .75, .90, and .99. In the Supplemental Material, we also report results for kinked linear functions, where the first and second halves of $\rho_{t}$ are linear, but with slopes of different signs. For example, “kinked 1.00-0.80-1.00” is a DGP for $\rho_{t}$ where the first half of $\rho_{t}$ decreases linearly from 1 to .80 while the second half increases linearly from .80 to 1. We consider 4 different cases in this category.

For each of the 21 DGP's for $\rho_{t}$, we consider both constant and time-varying DGP's for $\mu_{t}$ and $\sigma_{t}$ in (ref) except for that of “sin 1.00-0.90-1.00” with constant $\mu_{t}$ and $\sigma_{t}$ to present the legend for all the figures. Define $\mu_{t}^{*}=(1-\rho_{t})\mu_{t}$, which allows one to rewrite (ref) as

equation[equation omitted — 114 chars of source]

When $\mu_{t}$ and $\sigma_{t}$ are constant, we take $\mu_{t}=0$ and $\sigma_{t}=1$, which implies $\mu_{t}^{*}=0$. When $\mu_{t}$ and $\sigma_{t}$ are time-varying, we generate $\mu_{t}$ such that $\mu_{t}^{*}$ increases linearly from -.1 to .1, and $\sigma_{t}$ increases linearly from .95 to 1.05 as $t/n$ goes from 0 to 1. In consequence, there are a total of 41 ($=21\times2-1$) DGP's for $\left(\rho_{t},\mu_{t}^{*},\sigma_{t}\right)$. We take $U_{t}$ to be i.i.d. $N\left(0,1\right)$, initialize $Y_{0}$ by drawing it from a $N\left(0,\ol{\sigma}_{n}^{2}\right)$ distribution, where $\ol{\sigma}_{n}=1/\left(1-\ol{\rho}_{n}^{2}\right)$ and $\ol{\rho}_{n}=n^{-1}\sum_{t=1}^{n}\rho_{t}$, and form the $Y_{t}$ sequence as in (ref).

We compute a nominal .95 two-sided CI and MUE for $\rho(\tau)$ for each of 5 time points of interest, indexed by $\tau\in\left\{ .2,.4,.6,.8,1\right\} $. The CI and MUE are implemented as follows. First, we compute $\widehat{h}$ based on the method described in Section (ref) and take the undersmoothed $\widehat{h}_{us}$ to be as defined in ((ref)) with $c_{1}=.2,$ $c_{2}=1.5,$ and $a=1/10.$ In this step, the set of $nh$ values based on $\mathcal{H}_{n}$ is $\left\{ 140,155,...,500,650,...,1500\right\} $. Next, for each candidate value $\rho_{0}\in\left\{ -1,-.995,...,.945,.95,.951,.952,...,1\right\} $, we calculate the t-statistics $T_{n}\left(\rho_{0}\right)$ defined in (ref) from the regression of $Y_{t}$ on a constant and $Y_{t-1}$ using the $n\widehat{h}_{us}$ observations centered around $n\tau$. When $n\tau$ is close to the boundary, we use $n\widehat{h}_{us}/2$ observations to the side that has abundant data, and as many observations as available to the other side until the boundary is hit. In consequence, a different number of observations for different $\tau$ may be used even though the $n\widehat{h}_{us}$ value is the same. Comparing the $T_{n}\left(\rho_{0}\right)$ values with corresponding critical values $c_{\psi_{nh,\rho_{0}}}\left(\alpha/2\right)$ and $c_{\psi_{nh,\rho_{0}}}\left(1-\alpha/2\right)$ (for $\psi_{nh,\rho_{0}}$ defined in (ref)) at each candidate $\rho_{0}$ gives the nominal $1-\alpha$ equal-tailed two-sided CI for $\rho(\tau)$ as defined in (ref). We compute coverage probabilities and average lengths of the CI's across all simulations for each point of interest for each DGP.

To calculate the MUE for $\rho(\tau)$, we follow the procedure described in Section (ref) and use $\widetilde{\rho}_{n\tau}$ as the MUE of $\rho\left(\tau\right)$ when $\widetilde{\rho}_{n\tau}$ is a singleton. When $\widetilde{\rho}{}_{n\tau,up}\neq\widetilde{\rho}{}_{n\tau,low}$, we take $\widetilde{\rho}{}_{n\tau,up}$ as the MUE for $\rho\left(\tau\right)$ based on two considerations. First, using the upper bound $\widetilde{\rho}{}_{n\tau,up}$ has the theoretical property that it is not median-biased towards zero for positive $\rho\left(\tau\right)$ values because, by Corollary (ref), $\liminf_{n\rightarrow\infty}\inf_{\lambda\in\Lambda_{n}}P_{\lambda}\left(\widetilde{\rho}_{n\tau,up}\geq\rho\left(\tau\right)\right)\geq1/2$. Second, in the simulations we find that $\widetilde{\rho}_{n\tau}$ is much more likely to be an interval when the true $\rho\left(\tau\right)$ is close to one. We report the absolute median biases of the MUE and the range of the MUE values across the simulations for each time point of interest for each DGP in the figures.

For each of the 41 DGP's for $Y_{t}$, we run $M=5,000$ simulations. The length of the $Y_{t}$ sequence is $n=1,500$. Given the 41 DGP's and 5 time points of interest for each one, a total of 205 cases are considered.

Discussion of Results

First, Table (ref) summarizes the CP results by reporting the number and percentage of times that the CI CP's lie in different ranges across the 205 cases considered. The nominal CP is .95. Table (ref) shows that only in a small fraction of cases, 2.0%, is the CP quite low, i.e., in $[.875,.90).$ In the majority of cases, 92.2%, the CP is $.925$ or larger, which shows that the proposed method has reasonably good CP regardless of whether the underlying DGP is curvy or flat and whether the point of interest is in the middle or at the boundary of the time period.

table[table omitted — 424 chars of source]

Second, we consider summary statistics for the absolute values of the median biases (AMB) of the MUE, i.e., $\abs{\text{median}\left(\widetilde{\rho}_{n\tau,up}-\rho\left(\tau\right)\right)}$, across the 205 cases considered. The mean and median of the absolute median biases across all cases are .004 and .003, respectively. The range is $[.000,.023]$. Thus, the magnitudes of the absolute median biases of the MUE are generally small.

Third, we consider the values of $n\widehat{h}_{us}$ selected by the data-dependent method. The mean, median, and range are 212, 207, and {[}56, 433{]}, respectively. The $n\widehat{h}_{us}$ values vary depending on the shape of the $\rho$ function considered. The flatter is the true $\rho$ function, the larger are the $n\widehat{h}_{us}$ values.

Now, we describe the simulation results for the different $\rho$ functions considered. Figure (ref) provides simulation results for five sinusoidal $\rho$ functions with constant $\mu$ and $\sigma^{2}$ functions. The amplitudes of the sinusoidal functions increase as one moves down the rows. In the first column of graphs, the $\rho$ function's maximum value occurs at $\tau=1,$ which corresponds to $t=n\tau=n.$ In the second column of graphs, the $\rho$ function's minimum value occurs at $\tau=1.$ Each graph consists of an “upper” and a “lower” graph. The “upper” graph shows the true $\rho$ function, the average CI lower and upper bounds at five points of interest $\tau=.2,$ $.4,$ $.6,$ $.8,$ $1.0,$ where $t=n\tau,$ and twice the average lower and upper absolute deviations of the MUE (i.e., $2E|\widehat{\rho}_{t}-\rho_{t}|\mathbf{\mathbbm1}(\widehat{\rho}_{t}<\rho_{t})$ and $2E|\widehat{\rho}_{t}-\rho_{t}|\mathbf{\mathbbm1}(\widehat{\rho}_{t}>\rho_{t}),$ respectively, whose sum is somewhat analogous to two standard deviations) at each of these points of interest. The difference between the average upper and lower CI bounds gives the average lengths (AL's) of the CI's. The “lower” graphs report the CI coverage probabilities (CP's) at the five points of interest. The CI's all have nominal CP .95.

The sinusoidal graphs in Figure (ref) show the following: (i) All but three of the CP's range from .925 to .965, which are close to the nominal CP of .95. (ii) The AL's are shortest for $\rho(\tau)$ values closest to one, and longest for $\rho(\tau)$ values farthest from one. This reflects the $nh$-consistency of the LS estimator in the (temporally local) unit root and local-to-unity scenarios, and the $(nh)^{1/2}$-consistency of the LS estimator in the (temporally local) stationary scenarios. (iii) The AL's are larger for the $\tau=1$ than the $\tau<1$ cases because fewer observations are used to construct the CI in the former cases, which reflects the need to reduce the boundary bias in order to obtain a good CP. Combining this fact with point (ii), we find the largest AL's occur when $\tau=1$ {\parfillskip=0pt}

figure[figure omitted — 1,515 chars of source]

and $\rho(\tau)$ is far from one (which occurs in the two graphs in the second column). (iv) The magnitudes of the MUE average absolute deviations are roughly proportional to the AL's of the CI's. (v) The simulation results for the sinusoidal functions with time-varying $\mu$ and $\sigma^{2}$ functions, which are given in Figure (ref) in the Supplemental Material, are quite similar to those in Figure (ref).

Figure (ref)(a)-(d) provides results for linear $\rho$ functions with constant $\mu$ and $\sigma^{2}$ functions. The results are similar to those for the sinusoidal functions, although all of the CP's lie between .924 and .957. Points (ii)-(v) above also apply to the linear functions in Figure (ref). Results for time-varying $\mu$ and $\sigma^{2}$ are provided in Figure (ref) in the Supplemental Material, and are similar to those in Figure (ref).

Figures (ref)(e)-(f) and (ref)(a)-(b) report results for flat-linear $\rho$ functions with $\rho(t)$ varying between $.90$ and $.99$ in Figure (ref)(e)-(f) and between $.80$ and $.99$ in Figure (ref)(a)-(b). Figures (ref)(e) and (ref)(a) report results for constant $\mu$ and $\sigma^{2};$ while Figures (ref)(f) and (ref)(b) report results for time-varying $\mu$ and $\sigma^{2}.$ The CP results in Figures (ref)(e)-(f) and (ref)(a)-(b) lie between .932 and .956.

Figures (ref)(e)-(f) and (ref)(a)-(b) provide analogous results to those just discussed for flat-linear $\rho$ functions, but for the case where the flat part has value $.99,$ rather than $.80$ or $.90,$ and the linear part has a negative slope, rather than a positive slope. The CP results for these cases are lower than those for the flat-linear $\rho$ functions in Figures (ref) and (ref). For example, Figure (ref)(e) shows a CP of .905 at $\tau=.60.$ For the other four cases, the CP's are in the range of .917 to .957.

Figure (ref)(c)-(f) considers flat $\rho$ functions. In Figure (ref)(c)-(d), $\rho=.99$ and the $\mu$ and $\sigma^{2}$ functions are constant and time-varying, respectively. Figure (ref)(e)-(f) is analogous, but with $\rho=.90.$ Figure (ref)(c)-(d) also is analogous, but with $\rho=.75.$ The CP's in the flat $\rho$ function cases are all close to .95 and lie between .932 and .960. This occurs because there is no bias accruing due to a time-varying $\rho$ function. However, in the cases of flat $\rho,$ $\mu,$ and $\sigma^{2}$ functions, the CI's are not as short as the oracle CI's that rely on knowledge that the $\rho,$ $\mu,$ and $\sigma^{2}$ functions are all flat and take $nh=n.$ For the case of $\rho=.99$ and flat $\mu$ and $\sigma^{2}$ functions, the AL's of the TVP CI and the oracle CI are $.022$ and $.012,$ respectively, at $\tau=.2$. For $\rho=.90,$ they are $.085$ and $.039$ at the same $\tau$. For $\rho=.75$, they are $.124$ and $.062$. There is a substantial increase in the average lengths in the flat cases using the methods of this paper, in order to ensure good CP's in the near flat cases.

Lastly, Figures (ref)(e)-(f) and (ref)(a)-(f) in the Supplemental Material provide results for kinked $\rho$ functions that either linearly increase until $\tau=.5$ and then linearly decrease, or linearly decrease until $\tau=.5$ and then linearly increase. Both constant and time-varying $\mu$ and $\sigma^{2}$ functions are considered. In two cases, the CP's are .904 and .906. In all other cases, {\parfillskip=0pt}

figure[figure omitted — 1,383 chars of source]
figure[figure omitted — 1,331 chars of source]

the CP's are in the range of .930 to .953. When the difference between the minimum and maximum $\rho$ value considered is large, viz., $.4,$ the AL's of the CI's are large. This occurs because the data-dependent choice of $h$ is small in order to avoid bias.

Overall, the CP's of the CI's are quite good with only a few being as low as .88 or .89 out of the 205 cases considered. The AL's vary substantially across scenarios depending on how close $\rho(\tau)$ is to one and how much the $\rho$ function varies with time, as is to be expected.

Empirical Applications

This section presents applications of the proposed methods to time series of inflation and exchange rates in several countries. The data comes from the IMF (International Financial Statistics (IFS)) database.

In the Supplemental Material, results are provided for some additional countries and for interest rates. The Supplemental Material also provides applications to eight macroeconomic series for the US, using the Federal Reserve Economic Data (FRED).

In terms of computation, we set $nh_{\min}=.2n$ and $nh_{\max}=2n$, where $n$ is the sample size. For $nh$ values between $nh_{\min}$ and $nh_{\text{mid}}$, where $nh_{\text{mid}}\coloneqq.5n$, we use a grid size of $.02n$, while for $nh$ values between $nh_{\text{mid}}$ and $nh_{\max}$, we use a grid size of $.05n$ because the range of $nh$ values is wider than between $nh_{\min}$ and $nh_{\text{mid}}$.

Inflation

First, we apply our method to the monthly inflation data, defined as the percentage change in CPI over the previous month, see Figure (ref). We consider three countries: the US, Canada, and Germany. The data span is Feb 1955 to Oct 2022, which contains $n=813$ observations for each country. We compute the $n\widehat{h}_{us}$ values based on ((ref)). Given these, we compute the MUE's and 90% CI's for $\rho\left(t\right)$ at each $t$ and for each country. As a robustness check, we multiply the undersmoothed $n\widehat{h}_{us}$ by 1.5 and present the results in the right column of each figure. The final $n\widehat{h}_{us}$ values used for the computations are listed in the titles of the figures. For all of the inflation series, Ljung-Box tests based on six lags of the residuals from the AR(1) model fail to reject the null hypothesis of no autocorrelation at the 5% significance level, see Table (ref).

For comparative purposes, we also fit a constant autoregression coefficient AR(1) model to the data and report the estimated $\widehat{\rho}$ and its 90% CI in the same graph. To obtain the constant parameter MUE's denoted by the flat red solid lines and constant parameter CI's denoted by the flat red dotted lines, we fix $n\widehat{h}_{us}=2n$ and apply our method. Note that the

figure[figure omitted — 1,666 chars of source]

method to derive the constant parameter estimates is equivalent to Mikusheva's mikusheva2007uniform modification of Stock's stock1991confidence method. mikusheva2007uniform proves that these confidence sets are uniformly valid asymptotically for non-time-varying $\rho\in(0,1]$.

Inflation persistence in the US has been extensively studied with mixed results regarding its extent and causes. Studies like those by cogley2001evolving have noted a decrease in inflation persistence following the early 1980s, attributing this change to shifts in monetary policy, especially the Federal Reserve's (Fed) increased focus on inflation targeting. This view is supported by research that points to the Volcker disinflation period as a pivotal time when the Fed's credibility was enhanced, leading to a stabilization of inflation expectations, see sims1992interpreting. We also find a sharp decline from .78 to .12 in inflation persistence in the US measured by the MUE of the TVP-AR coefficient between 1983 and 1995 in Figure (ref)(a), which is consistent with the above empirical results.

Entering the era of 2000s, inflation persistence remains a central topic in macroeconomic research. One line of research (e.g., \citet*{eggertsson2003zero,chung2012have}) concerns zero lower bound (ZLB) effects when the Fed set the interest rates close to zero. The theory predicts that at ZLB, monetary policy could be less effective in controlling inflation, thereby potentially increasing inflation persistence if not coupled with assertive non-traditional interventions. To overcome the challenges of the ZLB effects on the effectiveness of its monetary policies, the Fed initiated a new practice called “forward guidance,” where the future course of monetary policy is communicated to the public by the central bank. Numerous studies (e.g., \citet*{CAMPBELL2012BIP,negro2023JPE}) have suggested that forward guidance is effective in enabling the Fed to better control inflation, which could lead to lower inflation persistence. Along this line, bernanke2020new argues that forward guidance combined with quantitative easing (QE) granted the Fed significantly more space to provide accommodation when its standard policy rate was near zero. More recently, \citet*{cole2023living} argue that despite the significant efforts to make their policy credible, the credibility of most central banks including the Fed has been generally declining, making monetary policies aimed at controlling inflation less effective. A concrete manifestation of their claim would be an increase in the inflation persistence over a longer horizon.

We find an increase in inflation persistence in the US from .3 to .6 measured by the MUE of TVP-AR coefficient between 2000 and 2008 in Figure (ref)(a), which seems to be mostly driven by the ZLB effects despite the introduction of forward guidance during this period. Later on, when more specific monetary policies (e.g., QE) were introduced to stimulate the economy following the 2008 financial crisis, inflation persistence drops from .6 to .4 between 2010 and 2013, in line with the result of bernanke2020new. With that said, we observe in Figure (ref)(a) an upward trend in inflation persistence in the US over a longer horizon between 2000 and 2022, which could be caused by a decline in the credibility of the Fed as claimed by \citet*{cole2023living}. Overall, we conclude that while the Fed's measures have been effective in controlling inflation expectations and persistence, more nuanced policies are needed to handle complications arising from the ZLB effects, heterogeneous beliefs (\citet*{andrade2019forward}), structural changes (e.g., higher volatility and shifts in consumer behavior and supply chains) caused by the Covid-19 and other factors.

We also apply our method to study inflation persistence in Canada and Germany and obtain the following results. First, we find significant time variation in inflation persistence as measured by the MUE's of $\rho\left(t\right)$ for both countries since 1980s. Given the magnitude of the time-variation, our findings show that inflation persistence is more likely to be driven by changes in the regime of monetary policy and credibility of central banks, rather than by nominal or real frictions in the economy. Second, there seems to be a universal upward trend since 2000 in inflation persistence for all countries under consideration, echoing the findings of \citet*{cole2023living}, who argue that there is a decline in credibility in central bank policies. Third, the introduction of the Euro in 1999 marked a significant shift in inflation persistence in Germany, with the European Central Bank (ECB) taking over monetary policy. Inflation rates remained relatively stable but were influenced by broader Eurozone policies and economic conditions. The early 2000s saw moderate inflation, which aligned closely with the ECB\textquoteright s target. Our method captures this regime change by showing drastically different patterns of inflation persistence before and after 1999, supporting the conclusion above that inflation persistence is more likely to be driven by changes in the regime of monetary policy and the credibility of central banks.

Real Exchange Rate

The second application concerns the real exchange rate. The results are given in Figure (ref). We use US dollars (USD) as the benchmark currency, and calculate the bilateral real exchange rate $re_{it}$ of country $i$ at time $t$ as $rex_{it}=nex_{it}\times\frac{CPI_{it}}{CPI_{0t}}$, where $nex_{it}$ is the nominal exchange rate (USD per domestic currency) at time $t$, $CPI_{it}$ is the price level of country $i$ at time $t$, and $CPI_{0t}$ is the price level of the US at time $t$. Therefore, an increase in $rex_{it}$ represents an appreciation of country $i$'s currency against USD. We report results for the UK, Sweden, and Switzerland. The data span for the monthly real exchange rate dataset is Jan 1957 to Aug 2022. Thus, $n=788$ for each country. Similar to the inflation series, we compute the $n\widehat{h}_{us}$ values based on ((ref)). Then, we compute the MUE's and 90% CI's for the

figure[figure omitted — 1,531 chars of source]

AR coefficient, $\rho\left(t\right)$, at each $t$ and for each country. For robustness purposes, we multiply the undersmoothed $n\widehat{h}_{us}$ by 1.5 and report the results in the right column of each figure. For the real exchange rate series considered in this section, Ljung-Box tests based on six lags of the residuals from the AR(1) model fail to reject the null hypothesis of no autocorrelation at the 5% significance level, see Table (ref).

It has been known in the literature that real exchange rates in developed countries tend to be highly persistent, with deviations from the PPP level taking a long time to disappear (rogoff1996purchasing,engel2014exchange). There are several explanations in the literature for the real exchange rate persistence, including nominal price rigidities, interest rate inertia in monetary policy (benigno2004real), heterogeneous dynamics in subcomponents (\citet*{imbs2005ppp}), Balassa--Samuelson effects (balassa1964purchasing,samuelson1964theoretical), to name a few. Figure (ref) presents the results on real exchange rate persistence measured by the MUE's of $\rho\left(t\right)$ for the UK, Sweden, and Switzerland. Across all three countries the MUE's are close to 1, showing that our method yields near constant graphs in scenarios where that seems to be suitable. The 90% CI's for $\rho\left(t\right)$ are also very short. \citet*{chari2002can} report estimates from a constant parameter autoregressive process to lie between .76 and .87 for US bilateral real exchange rates against nine developed European countries using data between 1972 and 1994. They developed a general equilibrium model where the firms can only set price once per year, which implies a very high level of price stickiness in the economy. Yet, their model cannot generate the level of persistence observed in the data. Allowing for a time-varying parameter autoregressive process and using data from a longer time horizon, our MUE's of $\rho\left(t\right)$ for the three currencies are even higher at close to one, supporting their claim that nominal price rigidities are not enough to explain the high real exchange rate persistence. Practitioners would need to seek other possible explanations such as the Balassa--Samuelson effects or heterogeneous dynamics in subcomponents.

While high real exchange rate persistence seems prevalent, there is heterogeneity in the pattern of persistence across the countries we consider. In particular, the MUE's of $\rho\left(t\right)$ for Switzerland demonstrate a downward trend since early 1980s in Figure (ref)(e). The Swiss National Bank (SNB) has adopted several significant monetary policy changes since the 1990s, including the shift to a three-fold target (price stability, 3-year inflation forecast, and a range for the 3M Libor, see \citet*{jordan2010ten}) in the late 1990s and the introduction of negative interest rates in 2015 when it was forced to abandon a policy of defending the Swiss franc with a peg to the euro. These policies aim to stabilize price levels and influence interest rates, which can affect exchange rate dynamics and potentially reduce persistence by promoting quicker adjustments to shocks. In Figure (ref)(e) we observe a sharp decline in the MUE of AR(1) coefficient between 1995 and 2002 and between 2015 and 2019, consistent with the timing of SNB's monetary policy changes. Meanwhile, the European debt crisis and subsequent economic turmoil in the Eurozone led to significant safe-haven flows into the Swiss franc, prompting the SNB to implement a cap on the franc\textquoteright s value against the euro in 2011. This cap was removed in 2015. Such a cap on the franc's value would increase its real exchange rate persistence since the price of the franc could have fluctuated above the cap for the duration of the policy. We also find the real exchange rate persistence to be slightly increasing between 2011 and 2015 in Figure (ref)(e). As shown by these results, our method can capture the major events and policy changes that affect real exchange rate persistence reasonably accurately.

Asymptotics

This section establishes the correct uniform asymptotic size and asymptotic similarity of the confidence interval $CI_{n,\tau}$ for $\rho(\tau)$, the median-unbiased property of $\widetilde{\rho}_{n\tau}$, and some asymptotic properties of the data-dependent bandwidth parameter $\widehat{h}$. To prove these results, this section provides asymptotic results for the LS estimator $\widehat{\rho}_{n\tau}$ and t-statistic $T_{n}\left(\rho_{0,n}\right)$ under drifting sequences of parameter values. All proofs of the results stated below are given in Section (ref) of the Supplemental Material.

\protectParameter Space

Let $I_{a,r}:=\left[a-r,a+r\right]$ for $a\in\mathbb{R}$ and $r>0$. Let $\lfloor x\rfloor$ denote the integer part of $x$.

We impose the following structure on the $\rho$ function: for some $\varepsilon_{2},\varepsilon_{3}>0$,

equation[equation omitted — 89 chars of source]

for $s\in I_{\tau,\varepsilon_{2}}$ and some $b\in\left[\varepsilon_{3},\infty\right]$, where $\kappa\left(\cdot\right)$ is a nonnegative twice continuously differentiable function on $I_{\tau,\varepsilon_{2}}$.

The parameter space for $\left(\rho,\mu,\sigma^{2},\kappa,b,F\right)$ is given by

$\Lambda_{n}=\{\lambda=\left(\rho,\mu,\sigma^{2},\kappa,b,F\right)$:

(i) $\rho,\ \mu,$ and $\sigma^{2}$ are Lipschitz functions from $\left[0,1\right]$ to $\left[-1+\varepsilon_{1},1\right],\ \left[C_{2,L},C_{2,U}\right]$, and $\left[C_{3,L},C_{3,U}\right]$, respectively, with Lipschitz constants bounded by $L_{1}$, $L_{2}$, and $L_{3}$, respectively;

(ii) $\rho\left(s\right)=1-\kappa\left(s\right)/b$ for $s\in I_{\tau,\varepsilon_{2}}$ and $b\in[\varepsilon_{3},\infty]$, where $\kappa\left(\cdot\right)$ is a twice continuously differentiable function from $I_{\tau,\varepsilon_{2}}$ to $\left[0,C_{4}\right]$ with Lipschitz constant bounded by $L_{4}$ and $\kappa(\tau)\geq\varepsilon_{4}$;

(iii) $\mu\left(s\right)=C_{\mu1}\exp\left\{ -\eta\left(s\right)/b\right\} +C_{\mu2}$ for $s\in I_{\tau,\varepsilon_{2}}$ and $b\in[\varepsilon_{3},\infty]$, where $\eta\left(\cdot\right)$ is a nonnegative Lipschitz function with Lipschitz constant bounded by $L_{5}$;

(iv) $\left\{ U_{t}:t=0,1,...,n\right\} $ is a stationary martingale difference sequence under $F$ with $E_{F}\left(\rest{U_{t}}\mathscr{G}_{t-1}\right)=0\ \text{a.s.}$, $E_{F}\left(\rest{U_{t}^{2}}\mathscr{G}_{t-1}\right)=1\ \text{a.s.}$, and $E_{F}\left[\rest{U_{t}^{4}}\mathscr{G}_{t-1}\right]<M\ \text{a.s.},$ where $\mathscr{G}_{t}$ is some non-decreasing sequence of $\sigma$-fields for which $\sigma\left(U_{0},...,U_{t}\right)\subseteq\mathscr{G}_{t}$ and $\sigma\left(Y_{0}^{*}\right)\in\mathscr{G}_{t}$ for $t=1,...,n$;

(v) $E_{F}\left(Y_{0}^{*}\right)^{2}\leq C_{5}n$;

for some constants $0<\varepsilon_{1}<2$, $\varepsilon_{2},\varepsilon_{3},\varepsilon_{4}>0$, $-\infty<C_{j,L}\leq C_{j,U}<\infty$ for $j=2,3$, $C_{3,L}>0$, $0\leq C_{4}<\infty$, $0\leq C_{5}<\infty$, $L_{j}<\infty\ \forall j\leq5$, $C_{\mu1},C_{\mu2}\in\left(-\infty,\infty\right),$ and $M\in\left(0,\infty\right)\}$.

$\ $

Note that the dependence on $n$ of $\Lambda_{n}$ is only through part (v), which concerns the initial condition $Y_{0}^{*}$.\footnote{The parameter space $\Lambda_{n}$ could be made independent of $n$ by specifying that $E\left(Y_{0}^{*}\right)^{2}\leq C_{5}$. But, allowing the bound on $E\left(Y_{0}^{*}\right)^{2}$ to be $C_{5}n$, allows the parameter space to expand with $n$ and is in accord with assumptions in the literature for AR models with a possible unit root or near unit root.} Also, $\mu\left(s\right)$ is defined in such a way that it is allowed to vary smoothly on $\mathbb{R}$. It is used in the proof of Lemma (ref)(d), which is further used to bound the absolute difference between $\mu\left(t\right)$ and $\mu\left(s\right)$ for $t,s\in\left[T_{1},T_{2}\right]$ in obtaining the asymptotic distribution of the t-statistic $T_{n}\left(\rho_{0,n}\right)$.

$\ $

To establish the correct asymptotic size of the CI for $\rho(\tau)$, we need to consider sequences of parameters $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\right\} _{n\geq1}$ in $\Lambda_{n}.$ The parameter space $\Lambda_{n}$ is defined to allow the sequence $\left\{ \rho_{n}(\tau)\right\} _{n\geq1}$ to equal one for all $n\geq1,$ converge to one at any rate, or be bounded away from one. More specifically, by the definition of $\rho_{n}(\cdot)$,

equation[equation omitted — 149 chars of source]

If $b_{n}\rightarrow\infty$, then $\rho_{n}\left(\tau\right)$ is local to one. If $\limsup_{n\rightarrow\infty}b_{n}<\infty$, then $\rho_{n}\left(\tau\right)$ is bounded away from one. Furthermore, $\Lambda_{n}$ is defined so that, for points $s\in\left[0,1\right]$ other than $\tau,$ $\rho_{n}\left(s\right)$ can converge to one at any rate or be bounded away from one, and the distance-from-unity behavior may differ between different $s$ outside of $I_{\tau,\varepsilon_{2}}.$

Correct Asymptotic Size of the Confidence Interval for \texorpdfstring{$\boldsymbol{\rho(\tau)}$}{rho(tau)}

The bandwidth $h$ is assumed to satisfy the following assumptions.

assumption[Bandwidth $\boldsymbol{h}$] $h\rightarrow0$ and $nh\rightarrow\infty$ as $n\rightarrow\infty$.
assumption[Order of $\boldsymbol{h}$] $nh^{5}\rightarrow0$.\footnote{Assumption (ref) is used in the proofs of Lemmas (ref)(b), (ref)(c), (ref), and (ref)(a) below.}

Let $P_{\lambda}\left(\cdot\right)$ denote probability under $\lambda\in\Lambda_{n}.$ The correct asymptotic size and asymptotic similarity of the CI $CI_{n,\tau}$ are established in the following theorem. The theorem's proof relies on the “drifting pointwise” asymptotic results given in Sections (ref)-(ref) below and the generic asymptotic size results in \citet*{andrews2020generic}.

thm[Correct Asymptotic Size of $\boldsymbol{CI_{\boldsymbol{n,}\tau}}$] Under Assumptions (ref) and (ref), \[ \liminf\limits_{n\rightarrow\infty}\inf_{\lambda\in\Lambda_{n}}P_{\lambda}\left(\rho\left(\tau\right)\in CI_{n,\tau}\right)=\limsup\limits_{n\rightarrow\infty}\sup_{\lambda\in\Lambda_{n}}P_{\lambda}\left(\rho\left(\tau\right)\in CI_{n,\tau}\right)=1-\alpha. \]

The median-unbiased interval estimator $\widetilde{\rho}_{n\tau}$ has the following median-unbiasedness property. This result is a corollary to one-sided versions of Theorem (ref).\footnote{Corollary (ref) holds because $CI_{n,\tau}^{up}\left(.5\right)$ and $CI_{n,\tau}^{low}\left(.5\right)$ both have coverage probabilities of $1/2$ or greater by the proof of Theorem (ref) applied to these one-sided CIs and $\left(\widetilde{\rho}_{n\tau}\geq\rho\left(\tau\right)\right)$$\supset$ $CI_{n,\tau}^{up}\left(.5\right)$ and $\left(\widetilde{\rho}_{n\tau}\leq\rho\left(\tau\right)\right)\supset CI_{n,\tau}^{low}\left(.5\right).$}

cor[Asymptotic Median-Unbiasedness of $\boldsymbol{\boldsymbol{\widetilde{\rho}}_{\boldsymbol{n\tau}}}$] Under Assumptions (ref) and (ref), \begin{align*} \liminf\limits_{n\rightarrow\infty}\inf_{\lambda\in\Lambda_{n}}P_{\lambda}\left(\widetilde{\rho}_{n\tau,up}\geq\rho\left(\tau\right)\right) & \geq1/2\ and\\ \liminf\limits_{n\rightarrow\infty}\inf_{\lambda\in\Lambda_{n}}P_{\lambda}\left(\widetilde{\rho}_{n\tau,low}\leq\rho\left(\tau\right)\right) & \geq1/2. \end{align*}

\protectPreliminaries

Define

equation[equation omitted — 146 chars of source]

We consider the interval $I_{n\tau,nh/2}$ and decompose $Y_{t}^{*}$ into the sum of two parts

equation[equation omitted — 118 chars of source]

where $Y_{t}^{0}$ is an AR(1) process with the same time-varying parameters $\rho_{t}$ as $Y_{t}^{*}$, but with a zero initial condition at $T_{0}:=T_{1}-1$, and $c_{t,t-T_{0}}$ is defined in (ref). By recursive substitution in (ref), we have

equation[equation omitted — 243 chars of source]

Then, from (ref) and (ref), we have

align[align omitted — 156 chars of source]

We consider sequences $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$$.$ The estimand of interest is $\rho_{n}\left(\tau\right)$ for $n\geq1.$

Let

align[align omitted — 355 chars of source]

By part (ii) of the definition of $\Lambda_{n},$ for $t\in I_{n\tau,nh/2}$,

equation[equation omitted — 106 chars of source]

for $n$ sufficiently large that $h/2+1/n\leq\varepsilon_{2}.$ For notational simplicity, we drop the $n$ subscripts on $\rho_{n,t},$ $\mu_{n,t},$ and $\sigma_{n,t}^{2}$ below.

We consider sequences of null hypotheses $H_{0}:\rho_{n\tau}=\rho_{0,n}$ for $n\geq1.$ As noted above, the CI $CI_{n,\tau}$ is defined by the inversion of tests of such null hypotheses.

$\ $

For some results, we assume the functions $\kappa_{n}\left(\cdot\right)$, $\mu_{n}\left(\cdot\right)$, and $\sigma_{n}^{2}\left(\cdot\right)$ converge.

assumption[Limits of $\boldsymbol{\kappa_{n}\left(\cdot\right)}$, $\boldsymbol{\mu_{n}\left(\cdot\right)}$, and $\boldsymbol{\sigma_{n}^{2}\left(\cdot\right)}$] $\kappa_{n}\left(\cdot\right)$, $\mu_{n}\left(\cdot\right),$ and $\sigma_{n}^{2}\left(\cdot\right)$ restricted to $I_{\tau,\varepsilon_{2}}$ converge uniformly to some functions $\kappa_{0}\left(\cdot\right)$, $\mu_{0}\left(\cdot\right),$ and $\sigma_{0}^{2}\left(\cdot\right)$ on $I_{\tau,\varepsilon_{2}}$, respectively.

$\ $

For some results, we assume that $\left\{ b_{n}/\left(nh^{1/2}\right)\right\} _{n\geq1}$ converges.

assumption[Convergence of $\boldsymbol{b_{n}/\left(nh^{1/2}\right)}$] $b_{n}/\left(nh^{1/2}\right)\to w_{0}$ for some $w_{0}\in\left[0,\infty\right]$.

$\ $

Assumptions (ref) and (ref) are innocuous because, to establish the uniform inference results in Theorem (ref), we show that it suffices to consider sequences $\left\{ \lambda_{n}\in\Lambda_{n}\right\} _{n\geq1}$ that satisfy these assumptions. Note that Theorem (ref) and Corollary (ref) do not impose Assumption (ref) or (ref).

The following lemma bounds the maximum intertemporal difference between the TVP $\rho_{t}$ and $\rho_{n\tau}$ for $t\in\left[T_{1},T_{2}\right]$.

lem[{{Maximum Intertemporal Differences on $\boldsymbol{\left[T_{1},T_{2}\right]}$}}] Under Assumptions (ref) and (ref), for a sequence $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$, \begin{enumerate} • $\max_{t\in\left[T_{1},T_{2}\right]}\abs{\rho_{t}-\rho_{n\tau}}=O\left(h/b_{n}\right)$, • $\max_{t\in\left[T_{1},T_{2}\right]}\abs{\sigma_{t}^{2}-\sigma_{n\tau}^{2}}=O\left(h\right)$, • $\max_{t\in\left[T_{1},T_{2}\right]}\max_{0\leq j\leq t-T_{1}}\abs{c_{t,j}-\rho_{n\tau}^{j}}=O\left(nh^{2}/b_{n}\right)$, $\text{ and}$$\max_{t\in\left[T_{1},T_{2}\right]}\abs{\mu_{t}-\mu_{n\tau}}=O\left(h/b_{n}\right)$. \end{enumerate}

The next lemma provides several bounds on the endogenous initial condition $Y_{T_{0}}^{*}$.

lem[Order of $\boldsymbol{Y_{T_{0}}^{*}}$] For a sequence $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$, we have \begin{enumerate} • $Y_{T_{0}}^{*}=O_{p}\left(n^{1/2}\right)$ under Assumption (ref), • $Y_{T_{0}}^{*}=o_{p}\left(b_{n}/\left(nh\right){}^{1/2}\right)$ under Assumptions (ref), (ref), (ref), and (ref) and $nh/b_{n}\to r_{0}=0$, and • $Y_{T_{0}}^{*}=O_{p}\left(b_{n}^{1/2}\right)$ under Assumptions (ref) and (ref) and $nh/b_{n}\to r_{0}=\infty$. \end{enumerate}
remLemma (ref)(a) shows that part (v) of $\Lambda_{n}$ implies that the endogenous initial condition $Y_{T_{0}}^{*}$ is $O_{p}\left(n^{1/2}\right)$. Part (b) of the lemma is used in the “very local-to-unity” case in which $nh/b_{n}\to r_{0}=0$. It provides a better bound than part (a) when $b_{n}/\left(nh^{1/2}\right)\to w_{0}<\infty$ because, in that case, the bound $o_{p}\left(b_{n}/\left(nh\right){}^{1/2}\right)$ is $o_{p}\left(n^{1/2}\right)$. Part (c) of the lemma provides a better bound than part (a) in the “stationary” case in which $nh/b_{n}\to r_{0}=\infty$.

Asymptotics in the Local-to-Unity Case

The local-to-unity case is characterized by $nh/b_{n}\rightarrow r_{0}\in\left[0,\infty\right)$. In this section, we determine the asymptotic distributions of the local LS estimator $\widehat{\rho}_{n\tau}$ and corresponding t-statistic $T_{n}\left(\rho_{0,n}\right)$ in the local-to-unity case.

Define

equation[equation omitted — 170 chars of source]

as a function of $s\in\left[0,1\right]$ for any fixed $\tau\in\left[0,1\right]$. First, we obtain the asymptotic distribution of the zero-initial condition process $\left(nh\right)^{-1/2}Y_{n,t\left(s\right)}^{0}$ for $s\in\left[0,1\right]$. Then, we obtain the asymptotic distribution of the normalized initial condition $Y_{T_{0}}^{*}$ in the $r_{0}\in\left(0,\infty\right)$ case. Lastly, we obtain the asymptotic distributions of $\widehat{\rho}_{n\tau}$ and $T_{n}\left(\rho_{0,n}\right)$. The last step includes dealing with the initial condition $Y_{T_{0}}^{*}$ in the $r_{0}=0$ case.

lem[Asymptotic Distribution of $\boldsymbol{Y_{n,t\left(s\right)}^{0}}$] Under Assumptions (ref) and (ref), for a sequence $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$ and $nh/b_{n}\rightarrow r_{0}\in\left[0,\infty\right)$ $\left(\text{local-to-unity case}\right)$, we have \[ \left(nh\right)^{-1/2}Y_{n,t\left(s\right)}^{0}/\sigma_{0}\left(\tau\right)\Rightarrow I_{\psi}\left(s\right)\ for\ \psi=r_{0}\kappa_{0}(\tau), \] where $t\left(s\right)$ is defined in (ref), $I_{\psi}\left(s\right)$ is defined in (ref), and “$\Rightarrow$” denotes weak convergence with respect to the Skorohod metric.

The following lemma is used to determine the effect of the initial condition on the LS estimator and t-statistic when $r_{0}\in(0,\infty).$

lem[Asymptotic Distribution of the Initial Condition $\boldsymbol{Y_{T_{0}}^{*}}$] Under Assumptions (ref) and (ref), for a sequence $\{\lambda_{n}=(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n})\in\Lambda_{n}\}_{n\geq1}$ and $nh/b_{n}\rightarrow r_{0}\in(0,\infty)$ (local-to-unity case), we have \[ (2\psi/nh)^{1/2}Y_{T_{0}}^{\ast}/\sigma_{0}\left(\tau\right)\rightarrow_{d}Z_{1}\sim N(0,1), \] where $\psi:=r_{0}\kappa_{0}(\tau),$ $Z_{1}$ is independent of $B(\cdot),$ $I_{\psi}(\cdot)$ defined in (ref), and the convergence holds jointly with the convergence in Lemma (ref)\emph{.}

The following lemma is useful in determining the asymptotic properties of $\widehat{\rho}_{n\tau}$ in the local-to-unity case.

lem[Convergence of Components in the Local-to-Unity Case] Under the null hypothesis $H_{0}:\rho_{n\tau}=\rho_{0,n}$, Assumptions (ref) and (ref), and $nh/b_{n}\rightarrow r_{0}\in\left(0,\infty\right)$ $\left(\text{local-to-unity case}\right)$, for a sequence $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$ and $\psi=r_{0}\kappa_{0}\left(\tau\right)$, the following results hold jointly \begin{enumerate} • $\left(nh\right)^{-1/2}Y_{t\left(s\right)}/\sigma_{0}\left(\tau\right)\Rightarrow I_{\psi}^{*}\left(s\right),$ where $t\left(s\right)=T_{1}+\lfloor nhs\rfloor$ for $s\in\left[0,1\right],$$\left(nh\right)^{-3/2}\sum_{t=T_{1}}^{T_{2}}Y_{t-1}/\sigma_{0}\left(\tau\right)\rightarrow_{d}\int_{0}^{1}I_{\psi}^{*}\left(s\right)ds,$$\left(nh\right)^{-2}\sum_{t=T_{1}}^{T_{2}}Y_{t-1}^{2}/\sigma_{0}^{2}\left(\tau\right)\rightarrow_{d}\int_{0}^{1}I_{\psi}^{*2}\left(s\right)ds,$$\left(nh\right)^{-1/2}\sum_{t=T_{1}}^{T_{2}}U_{t}\sigma_{t}/\sigma_{0}\left(\tau\right)\rightarrow_{d}\int_{0}^{1}dB\left(s\right),$$\left(nh\right)^{-1}\sum_{t=T_{1}}^{T_{2}}Y_{t-1}U_{t}\sigma_{t}/\sigma_{0}^{2}\left(\tau\right)\rightarrow_{d}\int_{0}^{1}I_{\psi}^{*}\left(s\right)dB\left(s\right)$, • $\left(nh\right)^{-2}\sum_{t=T_{1}}^{T_{2}}\left(Y_{t-1}^{0}/\sigma_{0}\left(\tau\right)\right)^{2}\rightarrow_{d}\int_{0}^{1}I_{\psi}^{2}\left(s\right)ds$, and • when $r_{0}=0$, parts (a)--(c) and (e) hold with $Y_{t\left(s\right)}$ ($=\mu_{t\left(s\right)}+Y_{t\left(s\right)}^{0}+c_{t\left(s\right),t\left(s\right)-T_{0}}Y_{T_{0}}^{*}$) and $Y_{t-1}$ replaced by $\mu_{t\left(s\right)}+Y_{t\left(s\right)}^{0}$ and $\mu_{t-1}+Y_{t-1}^{0}$, respectively. \end{enumerate}

After proper re-scaling, we have

equation[equation omitted — 302 chars of source]

The next theorem gives the asymptotic distribution of $nh\left(\widehat{\rho}_{n\tau}-\rho_{0,n}\right)$ and the t-statistic $T_{n}\left(\rho_{0,n}\right)$ in the local-to-unity case.

thm[Asymptotic Distribution of Normalized $\boldsymbol{\widehat{\rho}_{n\tau}}$ and $\boldsymbol{t}$-Statistic in the Local-to-Unity Case] Under the null hypothesis $H_{0}:\rho_{n\tau}=\rho_{0,n}$, Assumptions (ref), (ref), and (ref), $nh/b_{n}\rightarrow r_{0}\in\left[0,\infty\right)$ $\left(\text{local-to-unity case}\right)$, and Assumption (ref) if $r_{0}=0$, for a sequence $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$ and $\psi=r_{0}\kappa_{0}\left(\tau\right)$, we have \[ nh\left(\widehat{\rho}_{n\tau}-\rho_{0,n}\right)\rightarrow_{d}\left(\int_{0}^{1}I_{D,\psi}^{*2}\left(s\right)ds\right)^{-1}\int_{0}^{1}I_{D,\psi}^{*}\left(s\right)dB\left(s\right) \] and \[ T_{n}\left(\rho_{0,n}\right)\rightarrow_{d}\left(\int_{0}^{1}I_{D,\psi}^{*2}\left(s\right)ds\right)^{-1/2}\int_{0}^{1}I_{D,\psi}^{*}\left(s\right)dB\left(s\right). \]
remFor any subsequence $\{p_{n}\}_{n\geq1}$ of $\{n\}_{n\geq1}$, Lemmas (ref)--(ref) and Theorem (ref) hold with $p_{n}$ in place of $n$ throughout and $h_{p_{n}}$ in place of $h=h_{n}.$

\protectAsymptotics in the Stationary Case

The “stationary” case is characterized by $b_{n}=o\left(nh\right)$, or equivalently, $nh/b_{n}\rightarrow r_{0}=\infty.$ The stationary case is defined such that the t-statistic has a standard normal distribution under $H_{0}:\rho_{n\tau}=\rho_{0,n}$ in the stationary case. If $b_{n}$ is bounded, then $\rho_{n\tau}\leq C_{\rho}$ for some constant $C_{\rho}<1$ for all $n\geq1$. If $b_{n}$ diverges to infinity, then $\rho_{n\tau}$ goes to one at a rate slower than $1/nh$. Thus, the stationary case includes some scenarios where $\rho_{n\tau}$ goes to one.

Define

equation[equation omitted — 181 chars of source]

where $-1+\varepsilon_{1}$ is a lower bound on $\rho_{t}$ by the definition of $\Lambda_{n}.$ We can bound $|\rho_{t}|$ for $t\in\left[T_{1},T_{2}\right]$ using $\ol{\rho}_{n}$. First, suppose $\rho_{t}\leq0$, then $-1+\varepsilon_{1}\leq\rho_{t}\leq0$, which implies that $|\rho_{t}|\leq$$\ol{\rho}_{n}$, as desired. Given this, in the following calculations we suppose $\rho_{t}\geq0$ for all $t\in[0,1]$ without loss of generality. Then, for $n$ sufficiently large, we have

align[align omitted — 471 chars of source]

where $b_{n}\geq\varepsilon_{3}>0$, the second inequality uses Lemma (ref)(a), (ref), Assumption (ref), and a mean value expansion of the $\exp\left\{ \cdot\right\} $ function, and the last inequality holds using $\kappa_{0}(\tau)>0$ and the fact that when $b_{n}\rightarrow\infty$, $\overline{\rho}_{n}-\exp\{-\kappa_{0}(\tau)/b_{n}\}\geq K/b_{n}$ for some constant $K>0.$

Equation (ref) implies

equation[equation omitted — 189 chars of source]

Furthermore, using Lemma A.1 of \citet*{GIRAITISKapetaniosYates2014rckernel} and the definition of the parameter space $\Lambda_{n}$, under $H_{0}:\rho_{n\tau}=\rho_{0,n}$, we have

equation[equation omitted — 285 chars of source]

where the equality holds by Lemma (ref)(a).

The following results are used in the analysis:

align[align omitted — 763 chars of source]

When $\rho_{0,n}=1-\kappa_{n}\left(\tau\right)/b_{n}$ with $b_{n}=o\left(nh\right)$ , one can replace $\ol{\rho}_{n}$ with $\rho_{0,n}$ in (ref)--(ref) and the results still hold.

$\ $

We now establish the $N\left(0,1\right)$ asymptotic distribution of $\widehat{\rho}_{n\tau}$ in the stationary case with the following normalization:

align[align omitted — 463 chars of source]

Next, we prove a lemma on the asymptotic properties of the zero-initial condition process $Y_{t}^{0}$ in the stationary case.

lem[Asymptotics in the Stationary Case] Under the null hypothesis $H_{0}:\rho_{n\tau}=\rho_{0,n}$ and Assumptions (ref) and (ref), for a sequence $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$ where $nh/b_{n}\rightarrow r_{0}=\infty$ and $\kappa_{0}\left(\tau\right)>0$ $\left(\text{stationary case}\right)$, the following results hold jointly \begin{enumerate} • $\left(1-\rho_{0,n}^{2}\right)^{1/2}\left(nh\right)^{-1}\sum_{t=T_{1}}^{T_{2}}Y_{t-1}^{0}/\sigma_{0}\left(\tau\right)\rightarrow_{p}0,$$\left(1-\rho_{0,n}^{2}\right)\left(nh\right)^{-1}\sum_{t=T_{1}}^{T_{2}}\left(Y_{t-1}^{0}/\sigma_{0}\left(\tau\right)\right)^{2}\rightarrow_{p}1,$ and • $\left(1-\rho_{0,n}^{2}\right)^{1/2}\left(nh\right)^{-1/2}\sum_{t=T_{1}}^{T_{2}}Y_{t-1}^{0}\sigma_{t}U_{t}/\sigma_{0}^{2}\left(\tau\right)\rightarrow_{d} N\left(0,1\right),$ \end{enumerate} where $Y_{t}^{0}$ is defined in (ref).

Lemma (ref) concerns the zero-initial condition process $Y_{t}^{0}$ with time-varying autoregressive parameter $\rho_{t}$, which is the foundation of our asymptotic analysis of $\widehat{\rho}_{n\tau}$. Essentially Lemma (ref) says that when $nh/b_{n}\rightarrow r_{0}=\infty$, the asymptotic distributions of normalized sums of $Y_{t}^{0}$ behave in the same way as in an AR process with constant $\rho<1$. The proof involves applications of appropriate weak laws of large numbers and central limit theorems for martingale difference triangular arrays and approximations of TVPs. One can extend the results to allow the $U_{t}$ process to be conditionally heteroskedastic, but this requires a different definition of $\widehat{\sigma}_{n\tau}^{2}$ and complicates the analysis, see andrews2014conditional. For simplicity, we do not consider this extension here.

$\ $

Define

equation[equation omitted — 200 chars of source]
lem[Asymptotics of the Denominator in the Stationary Case] Under the null hypothesis $H_{0}:\rho_{n\tau}=\rho_{0,n}$ and Assumptions (ref), (ref), and (ref), for a sequence $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$ where $nh/b_{n}\rightarrow r_{0}=\infty$ and $\kappa_{0}\left(\tau\right)>0$ $\left(\text{stationary case}\right)$, the following results hold \begin{enumerate} • $\left(1-\rho_{0,n}^{2}\right)\left(nh\right)^{-1}\sum_{t=T_{1}}^{T_{2}}\left(Y_{t-1}-\ol{\mu}_{nh,-1}\right)^{2}/\sigma_{0}^{2}\left(\tau\right)\rightarrow_{p}1$, and • $\left(1-\rho_{0,n}^{2}\right)\left(\ol Y_{nh,-1}-\ol{\mu}_{nh,-1}\right)^{2}/\sigma_{0}^{2}\left(\tau\right)\rightarrow_{p}0$. \end{enumerate}

Lemma (ref) concerns the asymptotic distribution of the denominator of the normalized $\widehat{\rho}_{n\tau}$ in (ref). To bound the difference between $Y_{t}^{0}$ and $Y_{t}$ and control for TVPs asymptotically, we use various inequalities and approximations in the proof of Lemma (ref).

lem[Asymptotics of the Numerator in the Stationary Case] Under the null hypothesis $H_{0}:\rho_{n\tau}=\rho_{0,n}$ and Assumptions (ref), (ref), and (ref), for a sequence $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$ where $nh/b_{n}\rightarrow r_{0}=\infty$ and $\kappa_{0}\left(\tau\right)>0$ $\left(\text{stationary case}\right)$, the following results hold \begin{enumerate} • $\left(1-\rho_{0,n}^{2}\right)^{1/2}\left(nh\right)^{-1/2}\sum_{t=T_{1}}^{T_{2}}\left(Y_{t-1}-\ol{\mu}_{nh,-1}\right)\left[\begin{array}{c} Y_{t}-\ol{\mu}_{nh}-\rho_{0,n}\left(Y_{t-1}-\ol{\mu}_{nh,-1}\right)\end{array}\right]/\sigma_{0}^{2}\left(\tau\right)$ $\rightarrow_{d} N\left(0,1\right)$, and • $\left(1-\rho_{0,n}^{2}\right)^{1/2}\left(nh\right)^{-1/2}\sum_{t=T_{1}}^{T_{2}}\left(\ol Y_{nh,-1}-\ol{\mu}_{nh,-1}\right)\left[\begin{array}{c} Y_{t}-\ol{\mu}_{nh}-\rho_{0,n}\left(Y_{t-1}-\ol{\mu}_{nh,-1}\right)\end{array}\right]/\sigma_{0}^{2}\left(\tau\right)$ $\rightarrow_{p}0$. \end{enumerate}

Lemma (ref) concerns the asymptotic distribution of the numerator of the normalized $\widehat{\rho}_{n\tau}$ in (ref). The proof of this lemma uses the results in Lemmas (ref)(c) and (ref).

$\ $

The next theorem provides the limit distributions of $\left(1-\rho_{0,n}^{2}\right)^{-1/2}\left(nh\right)^{1/2}\left(\widehat{\rho}_{n\tau}-\rho_{0,n}\right)$ and $T_{n}\left(\rho_{0,n}\right)$ in the stationary $\rho_{n\tau}$ case.

thm[Asymptotic Distribution of Normalized $\boldsymbol{\widehat{\rho}_{n\tau}}$ and $\boldsymbol{t}$-Statistic in the Stationary Case] Under the null hypothesis $H_{0}:\rho_{n\tau}=\rho_{0,n}$ and Assumptions (ref), (ref), and (ref), for a sequence $\left\{ \lambda_{n}=\left(\rho_{n},\mu_{n},\sigma_{n}^{2},\kappa_{n},b_{n},F_{n}\right)\in\Lambda_{n}\right\} _{n\geq1}$ where $nh/b_{n}\rightarrow r_{0}=\infty$ and $\kappa_{0}\left(\tau\right)>0$ $\left(\text{stationary case}\right)$, we have \[ \left(1-\rho_{0,n}^{2}\right)^{-1/2}\left(nh\right)^{1/2}\left(\widehat{\rho}_{n\tau}-\rho_{0,n}\right)\rightarrow_{d} N\left(0,1\right) \] and \[ T_{n}\left(\rho_{0,n}\right)\rightarrow_{d} N\left(0,1\right). \]
remFor any subsequence $\{p_{n}\}_{n\geq1}$ of $\{n\}_{n\geq1}$, Lemmas (ref)--(ref) and Theorem (ref) hold with $p_{n}$ in place of $n$ throughout and $h_{p_{n}}$ in place of $h=h_{n}.$

Asymptotic Results for \texorpdfstring{$\boldsymbol{\widehat{h}}$}{h}

Here we give conditions under which $\widehat{h}$ defined in (ref) is asymptotically equivalent to the value $\widehat{h}_{opt}$ that minimizes the “empirical loss,” which is unobserved. See li1987asymptotic and andrews1991asymptotic for analogous results in i.i.d. models. The empirical loss, $L_{n}(h),$ is

equation[equation omitted — 170 chars of source]

We give a simple high-level condition under which

equation[equation omitted — 186 chars of source]

The $FE_{n}(h)$ criterion can be written as follows:

align[align omitted — 320 chars of source]

The empirical loss $L_{n}(h)$ and the cross-product term $C_{n}(h)$ depend on $h,$ but the average squared error $E_{n}$ does not. Since $E_{n}$ does not depend on $h,$ $\widehat{h}$ minimizes $L_{n}(h)-2C_{n}(h)$ over $\mathcal{H}_{n}.$

Under the following condition, $\widehat{h}$ is asymptotically equivalent to the infeasible value $\widehat{h}_{opt}$ that minimizes $L_{n}(h).$

assumption$\sup_{h\in\mathcal{H}_{n}}\frac{|C_{n}(h)|}{L_{n}(h)}\rightarrow_{p}0.$
lemUnder Assumption (ref), $\frac{L_{n}(\widehat{h})}{L_{n}(\widehat{h}_{opt})}\rightarrow_{p}1.$

Now, we give a set of sufficient conditions for Assumption (ref). Let $h_{\min}=\min\{h:h\in\mathcal{H}_{n}\}$ and $h_{\max}=\max\{h:h\in\mathcal{H}_{n}\}.$ Typically, $h_{\min}$ and $h_{\max}$ depend on $n$ and decrease to $0$ as $n\rightarrow\infty.$ Let $\xi_{n}$ denote the cardinality of $\mathcal{H}_{n}.$ We decompose $L_{n}(h)$ and $C_{n}(h)$ into the main components $L_{2n}(h)$ and $C_{2n}(h),$ respectively, which depend on $t=nh_{\max}+1,...,n$ and for which $(\widehat{\mu}_{t-1}(h),\widehat{\rho}_{t-1}(h))$ depend only on random variables in $\mathcal{G}_{t-1},$ and the “small $t$” boundary components $L_{1n}(h)$ and $C_{1n}(h),$ respectively, which depend on $t=1,...,nh_{\max}$ for which $(\widehat{\mu}_{t-1}(h),\widehat{\rho}_{t-1}(h))$ depend on some random variables that are not in $\mathcal{G}_{t-1}.$ Define

align[align omitted — 647 chars of source]

The risk as a function of $h$ is $R_{n}(h)=E L_{n}(h).$ Let $R_{2n}(h)=E L_{2n}(h).$

The following assumption is sufficient for Assumption (ref).

assumption[Sufficient Conditions for Assumption (ref)]
enumerate• The true sequence of distributions is from the parameter spaces $\{\Lambda_{n}:n\geq1\}.$$\sup_{h\in\mathcal{H}_{n}}\frac{|C_{1n}(h)|}{L_{n}(h)}\rightarrow_{p}0.$$\sup_{h\in\mathcal{H}_{n}}|\frac{L_{2n}(h)}{R_{2n}(h)}-1|\rightarrow_{p}0.$$h_{\max}\leq1-\varepsilon$ for $n$ large for some $\varepsilon>0.$$\frac{n\inf_{h\in\mathcal{H}_{n}}R_{2n}(h)}{\xi_{n}}\rightarrow\infty.$

Assumption (ref)(b) implies that the “small $t$” boundary component of $C_{n}(h),$ which only depends on $t\leq nh_{\max},$ is asymptotically dominated by $L_{n}(h),$ which is based on all $t\leq n.$ Assumption (ref)(c) requires that the variability of $L_{2n}(h)$ is small relative to its mean. In particular, Assumption (ref)(c) holds if $StdDev(L_{2n}(h))/E L_{2n}(h)=o(1).$ Assumption (ref)(d) implies that the elements of $\mathcal{H}_{n}$ are bounded away from one, which implies that $nh=n$ is not a feasible choice.

To interpret Assumption (ref)(e), we give an intuitive discussion of the magnitudes of $\inf_{h\in\mathcal{H}_{n}}L_{2n}(h)$ and $\inf_{h\in\mathcal{H}_{n}}R_{2n}(h).$ First, consider the case where $\{Y_{t}\}_{t\leq n}$ displays unit root or local-to-unity behavior across the whole time series. Then, $\widehat{\mu}_{t-1}(h)-\mu_{t}\thickapprox(nh)^{-1/2}$ and $Y_{t-1}(\widehat{\rho}_{t-1}(h)-\rho_{t})\thickapprox n^{1/2}(nh)^{-1}=(nh^{2})^{-1/2},$ $L_{2n}(h)\thickapprox(nh^{2})^{-1},$ and under suitable conditions, $R_{2n}(h)\thickapprox(nh^{2})^{-1}.$ Second, in the case where $\{Y_{t}\}_{t\leq n}$ displays stationary behavior across the whole time series, $\widehat{\mu}_{t-1}(h)-\mu_{t}\thickapprox(nh)^{-1/2},$ $Y_{t-1}(\widehat{\rho}_{t-1}(h)-\rho_{t})\thickapprox(nh)^{-1/2},$ $L_{2n}(h)\thickapprox(nh)^{-1},$ and under suitable conditions $R_{2n}(h)\thickapprox(nh)^{-1}.$ Third, in the case where $\{Y_{t}\}_{t\leq n}$ displays behavior that varies between unit root and stationarity across the time series, the order of magnitude of $R_{2n}(h)$ is between $(nh^{2})^{-1}$ and $(nh)^{-1},$ which is bounded below by $(nh_{\max})^{-1}.$ Hence, Assumption (ref)(e) requires $n\cdot(nh_{\max})^{-1}/\xi_{n}\rightarrow\infty$ or $\xi_{n}=o(h_{\max}^{-1}).$ That is, the number $\xi_{n}$ of values $h$ in $\mathcal{H}_{n}$ needs to be of smaller order than the reciprocal of the maximum value $h_{\max}$ in $\mathcal{H}_{n}.$

lemAssumption (ref) implies Assumption (ref).
remThe proof of Lemma (ref) shows that $E C_{2n}(h)=0$ and $Var(C_{2n}(h))\leq C_{3U}\times(n-nh)^{-1}E L_{2n}(h)$ $\forall h\in\mathcal{H}_{n},n\geq1,$ where $C_{3U}<\infty$ is the bound on the variance function $\sigma^{2}$ in the definition of $\Lambda_{n}.$
remLet $R_{1n}(h)=E L_{1n}(h).$ Under Assumption (ref) and $\sup_{h\in\mathcal{H}_{n}}\frac{R_{1n}(h)+E|C_{1n}(h)|}{R_{n}(h)}\rightarrow0,$ we also have: $\frac{R_{n}(\widehat{h})}{R_{n}(\widehat{h}_{opt})}\rightarrow_{p}1.$