EconBase
← Back to paper

Estimating Time-Varying Parameters of Various Smoothness in Linear Models via Kernel Regression

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

96,364 characters · 22 sections · 66 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Estimating Time-Varying Parameters of Various Smoothness in Linear Models via Kernel Regression

\onehalfspacing

titlingpage\begin{abstract} We study kernel-based estimation of nonparametric time-varying parameters (TVPs) in linear models. Our contributions are threefold. First, we establish consistency and asymptotic normality of the kernel-based estimator for a broad class of TVPs including deterministic smooth functions, the rescaled random walk, structural breaks, the threshold model and their mixtures. Our analysis exploits the smoothness of the TVP. Second, we show that the bandwidth rate must be determined according to the smoothness of the TVP. For example, the conventional $T^{-1/5}$ rate is valid only for sufficiently smooth TVPs, and the bandwidth should be proportional to $T^{-1/2}$ for random-walk TVPs, where $T$ is the sample size. We show this highlighting the overlooked fact that the bandwidth determines a trade-off between the convergence rate and the size of the class of TVPs that can be estimated. Third, we propose a data-driven procedure for bandwidth selection that is adaptive to the latent smoothness of the TVP. Simulations and an application to the capital asset pricing model suggest that the proposed method offers a unified approach to estimating a wide class of TVP models. \end{abstract} Keywords: Bandwidth, kernel estimation, random walk, structural break, time-varying parameter JEL Codes: C14, C22

Introduction

Parameter instabilities are widely observed in econometric analysis. One of the most common specifications for parameter changes is the following linear model:

align[align omitted — 99 chars of source]

where $T$ is the sample size, $p\times 1$ vector $x_t$ is the regressor, $p\times 1$ triangular array $\beta_{T,t}$ is the time-varying coefficient, and $\varepsilon_t$ is the disturbance.

In the literature, time-varying parameters are often estimated via kernel regression where observations are weighted by some kernel function. Starting from robinsonNonparametricEstimationTimeVarying1989, a large literature develops kernel-based estimation and inference for time-varying coefficient models; e.g., caiTrendingTimevaryingCoefficient2007, chenTestingSmoothStructural2012, zhangInferenceTimeVaryingRegression2012, inoueRollingWindowSelection2017, and friedrichSieveBootstrapInference2024. We follow this strand of literature and study kernel-based estimation of $\beta_{T,t}$.

Our contributions are threefold. First, we consider a broader class of time-varying parameters than is typically assumed. The most common assumption adopted in existing works is that $\beta_{T,t}$ is so smooth that it is continuously differentiable caiTrendingTimevaryingCoefficient2007,zhangInferenceTimeVaryingRegression2012,inoueRollingWindowSelection2017. However, smooth functions are not the only model for parameter instability popular in economics and statistics. The random walk model, in which $\beta_{T,t}$ is modeled as $\beta_{T,t} = \sum_{i=1}^{t}u_i$ with $u_i$ a transitory process, is a popular alternative nyblomTestingConstancyParameters1989,stockMedianUnbiasedEstimation1998,cogleyDriftsVolatilitiesMonetary2005. Another example is (abrupt) structural breaks in $\beta_{T,t}$ andrewsTestsParameterInstability1993,baiEstimatingTestingLinear1998. These two modeling schemes have received less attention in the literature on kernel-based estimation, and it is largely unknown what the consequence is if one applies kernel regression to these models.\footnote{giraitisInferenceStochasticTimevarying2014 and giraitisTimevaryingInstrumentalVariable2021 are among few exceptions. They show that random-walk type parameters can be estimated via a kernel-based method.}$^{,}$\footnote{pesaranSelectionEstimationWindow2007, pesaranOptimalForecastsPresence2013 and hiranoAnalyzingCrossvalidationForecasting2022 apply kernel-based approaches for random walk and structural break type parameter instabilities, but their focus is on optimal forecasting of $y_t$, rather than estimation of $\beta_{T,t}$.} We develop kernel-based estimation theory that covers a wide class of time-varying parameters, including smooth functions, the rescaled random walk, and structural breaks. This class also includes the threshold regression model of hansenSampleSplittingThreshold2000, which has rarely been considered in the context of kernel regression.

Let us emphasize that the class of time-varying parameters considered in this article also includes the mixtures of the aforementioned models. The relationship between $y_t$ and $x_t$, for example, may evolve smoothly but exhibit discontinuities during global financial crises or pandemics. The literature has acknowledged the importance of taking into account several types of parameter instability. For instance, mullerEfficientEstimationParameter2010 consider inference in models with time-varying parameters approximated by Gaussian processes and continuous functions possibly with finitely many jumps. Although their framework allows for nonlinear models and thus is more general than ours in this respect, they focus on small parameter instabilities (relative to those considered in this work). Therefore, large instabilities are not allowed in their model. chenTestingSmoothStructural2012 develop tests for smooth parameter changes with finitely many breaks but do not develop estimation theory for time-varying parameters of this type. kristensenNonparametricDetectionEstimation2012 proposes a nonparametric estimation method for time-varying coefficients by developing a framework that he argues allows for smooth functions, structural breaks, and the rescaled random walk. However, his analysis is restricted to smooth functional parameters only and cannot be extended to the other specifications. giraitisTimevaryingInstrumentalVariable2021 allow smooth deterministic functions, the rescaled random walk, and their mixture in an IV setting, but exclude (large) discontinuous breaks. Unlike these earlier works, we employ a general framework that accommodates all the aforementioned models and their mixtures.

We develop this framework by considering the class of time-varying parameters characterized by smoothness parameter $\alpha>0$, which generalizes the H\"{o}lder-class functions studied in the literature on nonparametric estimation tsybakovIntroductionNonparametricEstimation2009. For example, continuously differentiable functions have smoothness of $\alpha=1$ as in the usual H\"{o}lder condition, and the random walk divided by $\sqrt{T}$ has $\alpha=1/2$. As with the H\"{o}lder condition, a smaller $\alpha$ means more roughness of the path of the time-varying parameter. This implies that the random walk divided by $\sqrt{T}$ is less smooth than continuously differentiable functions.

Our second contribution is to discuss the role of the bandwidth and its implications on bandwidth selection in the above general setting; beyond the usual bias-variance trade-off inherent in nonparametric estimation, we demonstrate that the bandwidth determines a trade-off between the rate of convergence and the size of the class of time-varying parameters that can be estimated. While this fact is consistent with and may be inferred from earlier results on nonparametric estimation of H\"{o}lder-class functions, its implications on bandwidth selection seem underappreciated in the literature on time-varying parameter models.

Specifically, we show that the rate of the bandwidth should be determined according to the smoothness of $\beta_{T,t}$. We illustrate this argument through two examples. First, we show that the conventional $T^{-1/5}$-rate bandwidth, which is specialized to continuously differentiable functions and often used in the literature zhouSimultaneousInferenceLinear2010, is invalid if $\beta_{T,t}$ is less smooth. For example, we show that the bandwidth should be proportional to $T^{-1/2}$ if $\beta_{T,t}$ is the random walk divided by $\sqrt{T}$. Second, we demonstrate that, if the time-varying parameter experiences both smooth and abrupt parameter changes, the abrupt breaks of certain magnitudes are absorbed in smooth parameter changes so that the kernel-based estimation delivers valid inference, while discontinuous changes of a larger magnitude cause bias. We show that the bandwidth determines the break magnitudes at which this bias arises.

Our third contribution is to propose a data-driven bandwidth selection procedure. Unlike existing approaches that focus on smooth time-varying parameters and $T^{-1/5}$-rate bandwidths, the proposed method allows researchers to select the bandwidth from a wide range of candidate values, adaptively to the latent smoothness of the time-varying parameter. We evaluate its finite-sample performance via Monte Carlo simulations and illustrate the method using the capital asset pricing model (CAPM). In this application, our selection algorithm does not support the conventional $T^{-1/5}$-rate bandwidth, casting doubt on the extent to which this routinely selected bandwidth and the commonly used assumption of (continuously) differentiable parameters are justified. Furthermore, the proposed procedure partially supports a $T^{-1/2}$-rate bandwidth, producing an estimated trajectory close to that obtained from a Bayesian random-walk estimation.

The remainder of this paper is organized as follows. Section (ref) defines smoothness of time-varying parameters. Section (ref) establishes asymptotic properties of the kernel-based estimator. Section (ref) discusses the consequence of an improper bandwidth choice and develops a bandwidth selection method adaptive to the smoothness of the time-varying parameter. Section (ref) conducts Monte Carlo experiments, and Section (ref) gives a real data analysis. Section (ref) concludes. Mathematical proofs of the main results are relegated to Appendix A.

Notation: For any matrix $A$, $\|A\|=\mathrm{tr}(A'A)^{1/2}$ denotes the Frobenius norm of $A$. For any positive number $b$, $\lfloor b \rfloor$ denotes the integer part of $b$. $\stackrel{p}{\to}$ and $\stackrel{d}{\to}$ signify convergence in probability and convergence in distribution as $T\to\infty$, respectively. $\Rightarrow$ signifies weak convergence of the associated probability measures.

Smoothness of Time-Varying Parameters

We consider estimating $\beta_{T,t}$ by using the local constant (Nadaraya-Watson) estimator:

align[align omitted — 154 chars of source]

where $K(\cdot)$ is a kernel function and $h$ is the bandwidth parameter satisfying $h\to0$ and $Th\to \infty$ as $T\to \infty$. Assumptions on the data generating process and kernel $K$ will be detailed in Section (ref)

In discussing the asymptotic properties of $\hat{\beta}_t$, the smoothness of the path of $\beta_{T,t}$ has a decisive effect. In the following definition, we quantify the smoothness of $\beta_{T,t}$ by a single parameter $\alpha$.

definitionTriangular array $\beta_{T,t}$ such that $\beta_{T,t} = O_p(1)$ as $T \to \infty$ for all $t$ is said to belong to the class type-a $\mathrm{TVP}(\alpha)$ or type-b $\mathrm{TVP}(\alpha)$, if the following condition (a) or (b) holds, respectively: \begin{itemize} • There exists some real $\alpha>0$ such that for any sequence $\{a_T\}$ of positive integers satisfying $a_T = o(T)$ and $a_T \to \infty$ as $T \to \infty$, and for any $t$, \begin{align} \max_{j:|t-j|\leq a_T} \|\beta_{T,t}-\beta_{T,j}\| = O_p\Bigl(\Bigl(\frac{a_T}{T}\Bigr)^{\alpha}\Bigr), \ as \ T \to \infty. \end{align} • There exists some real $\alpha>0$ such that for any sequence $\{a_T\}$ of positive integers satisfying $a_T = o(T)$ and $a_T \to \infty$ as $T \to \infty$, and for any $t$, \begin{align} \max_{j:|t-j|\leq a_T} \|\beta_{T,t}-\beta_{T,j}\| = O_p\Bigl(\frac{1}{T^{\alpha}}\Bigr), \ as \ T \to \infty. \end{align} \end{itemize} Furthermore, if $\beta_{T,t}$ belongs to $\mathrm{TVP}(\alpha)$ (type-a or type-b) but does not belong to $\mathrm{TVP}(\beta)$ for all $\beta>\alpha$, then it is said to belong to $\mathrm{TVP}(\alpha)$ on the boundary.

Definition (ref) essentially controls by $\alpha$ the smoothness of the path of $\beta_{T,t}$ on any interval of any length of a smaller order than $T$. In typical applications, $a_T$ will be set $a_T = \lfloor Th \rfloor$. Definition (ref)(a) allows the difference between the values of $\beta_{T,t}$ at distinct time points to grow as the time points gets further apart, while Definition (ref)(b) does not.\footnote{Therefore, $\beta_{T,t}$ belongs to type-a $\mathrm{TVP}(\alpha)$ if it belongs to type-b $\mathrm{TVP}(\alpha)$.}

Determining $\alpha$ such that a given $\beta_{T,t}$ belongs to $\mathrm{TVP}(\alpha)$ on the boundary enables us to derive the largest possible bandwidth under which $\hat{\beta}_t - \beta_{T,t}$ is asymptotically normally distributed; see Theorem (ref) below for this point.

Because $a_T/T<1$ and $\alpha>0$, a smaller $\alpha$ permits larger differences $\|\beta_{T,t}-\beta_{T,j}\|$, resulting in $\beta_{T,t}$ possibly having a rougher path. Note that triangular arrays unbounded in probability are excluded from Definition (ref). We emphasize that Definition (ref) does not impose any parametric assumption on $\beta_{T,t}$ (other than smoothness $\alpha$), and that $\beta_{T,t}$ may be deterministic or stochastic. In addition, $\beta_{T,t}$ is allowed to have arbitrary correlation with $x_t$ and $\varepsilon_{t}$. Definition (ref) is quite general and accommodates many important time-varying parameters, as shown below.

remgiraitisTimevaryingInstrumentalVariable2021 develop a kernel-based instrumental variable method to estimate time-varying parameters. The classes of time-varying parameters they consider are essentially type-a $\mathrm{TVP}(1)$ and $\mathrm{TVP}(1/2)$, albeit with slightly different definitions. They do not consider time-varying parameters belonging to type-a $\mathrm{TVP}(\alpha)$ with $\alpha \neq 1/2,1$ or type-b $\mathrm{TVP}(\alpha)$.
exm[Continuously differentiable functions] A popular model for time-varying parameters is deterministic smooth functions, accompanied by the formulation $\beta_{T,t} = \beta(t/T)$ for some continuously differentiable function $\beta(\cdot)$ on $[0,1]$ caiTrendingTimevaryingCoefficient2007,zhangInferenceTimeVaryingRegression2012,chenTestingSmoothStructural2012. Under this formulation, the fact that $\sup_{0\leq r\leq 1}\|\beta'(r)\| \leq C$ for some constant $C>0$ implies that for any $s,t=1,2,\ldots,T$, $\|\beta_{T,t}-\beta_{T,s}\| = \|\beta(t/T) - \beta(s/T)\| \leq C |t-s|/T$ by the mean value theorem. Therefore, we have $\max_{j:|t-j|\leq a_T}\|\beta_{T,t} - \beta_{T,j}\| \leq Ca_T/T = O(a_T/T)$ uniformly in $t$, which implies $\beta_{T,t}$ belongs to the type-a $\mathrm{TVP}(1)$ class. Furthermore, if there exist some interval $(a,b)$ and constant $c>0$ such that $\inf_{x\in(a,b)}\|\beta'(x)\|\geq c$,\footnote{This condition excludes constant functions.} then we have $\max_{j:|t-j|\leq a_T}\|\beta_{T,t}-\beta_{T,j}\| \geq ca_T/T$ for $t\in (a,b)$ and sufficiently large $T$ by the mean value theorem, and hence $\max_{j:|t-j|\leq a_T}\|\beta_{T,t}-\beta_{T,j}\|$ is not $O((a_T/T)^{\alpha})$ for $\alpha>1$ for such $t$ and $T$. This implies $\beta_{T,t}$ belongs to type-a $\mathrm{TVP}(1)$ on the boundary. More generally, $\beta_{T,t}$ belongs to the type-a $\mathrm{TVP}(\alpha)$ class if it is H\"{o}lder continuous with exponent $\alpha$.
exm[The random walk] Researchers often assume that the parameters of interest follow the random walk nyblomTestingConstancyParameters1989. We consider the random walk scaled by $\sqrt{T}$: $\beta_{T,0} = \mu$ and $\beta_{T,t} = \mu + (1/\sqrt{T})\sum_{i=1}^{t}u_i, \ t\geq1$, where $\mu$ is a constant and $\{u_i\}$ is an i.i.d. sequence with $E[u_i]=0$ and $ V[u_i] = \Sigma_u > 0$. The functional central limit theorem (FCLT) implies that \begin{align} \beta_{T,\lfloor T\cdot\rfloor} = \mu + \frac{1}{\sqrt{T}}\sum_{i=1}^{\lfloor T\cdot \rfloor}u_i \Rightarrow \mu + \Sigma_u^{1/2}B_1(\cdot), \end{align} in the Skorokhod space $D^p_{[0,1]}$, where $B_1$ is a $p$-dimensional vector standard Brownian motion. Then, the following result holds: for any $a_T<t<T-a_T+1$, \begin{align} \max_{j:|t-j|\leq a_T} \|\beta_{T,t} - \beta_{T,j}\| &=\max\left\{ \max_{t-a_T\leq j\leq t-1} \|\beta_{T,t} - \beta_{T,j}\|, \max_{t+1\leq j \leq t+a_T} \|\beta_{T,j} - \beta_{T,t}\|\right\} \\ &= \max\left\{\max_{t-a_T\leq j\leq t-1}\left\|\frac{1}{\sqrt{T}}\sum_{i=j+1}^{t} u_i\right\|, \max_{t+1\leq j \leq t+a_T} \left\|\frac{1}{\sqrt{T}}\sum_{i=t+1}^{j} u_i\right\|\right\} \\ &\stackrel{d}{=}\max\left\{\max_{1\leq j\leq a_T}\left\|\frac{1}{\sqrt{T}}\sum_{i=1}^{j} u_i \right\|, \max_{1\leq j\leq a_T}\left\|\frac{1}{\sqrt{T}}\sum_{i=1}^{j} u_i^* \right\| \right\} \\ &=\sqrt{\frac{a_T}{T}} \ \max\left\{\sup_{0\leq r \leq 1}\left\|\frac{1}{\sqrt{a_T}}\sum_{i=1}^{\lfloor a_Tr \rfloor} u_i \right\|, \sup_{0\leq r \leq 1}\left\|\frac{1}{\sqrt{a_T}}\sum_{i=1}^{\lfloor a_Tr \rfloor} u_i^* \right\| \right\}, \end{align} where $u_i^*$ is an i.i.d. copy of $u_i$, and the third equality in distribution follows from the i.i.d property of $\{u_t\}$. Because $\sup_{0\leq r \leq 1}\|(a_T)^{-1/2}\sum_{i=1}^{\lfloor a_Tr \rfloor} u_i \| = O_p(1)$ by the continuous mapping theorem (CMT), $\max_{j:|t-j|\leq a_T} \|\beta_{T,t} - \beta_{T,j}\| = O_p(\sqrt{a_T/T})$. Moreover, since $\sup_{0\leq r \leq 1}\|(a_T)^{-1/2}\sum_{i=1}^{\lfloor a_Tr \rfloor} u_i \|\stackrel{d}{\to} \sup_{0\leq r \leq 1}\|B_1(r)\|$, where $\sup_{0\leq r \leq 1}\|B_1(r)\|>0$ a.s., it follows that $\max_{j:|t-j|\leq a_T} \|\beta_{T,t} - \beta_{T,j}\|$ is not $O_p((a_T/T)^{\alpha})$ for any $\alpha>1/2$.\footnote{This can be verified by using the strong approximation.} The same conclusion holds for the other $t$. Hence, the random walk divided by $\sqrt{T}$ belongs to the type-a $\mathrm{TVP}(1/2)$ class on the boundary. More generally, the random walk divided by $T^{\alpha}$ belongs to the type-a $\mathrm{TVP}(\alpha)$ class for $\alpha\geq 1/2$, while the random walk divided by $T^{\alpha}$ with $\alpha<1/2$ is excluded from Definition (ref) because it is unbounded in probability. Because the random walk divided by $\sqrt{T}$ does not belong to $\mathrm{TVP}(1)$, it is less smooth than continuously differentiable functions on $[0,1]$. This is intuitively because the random walk divided by $\sqrt{T}$ weakly converges to Brownian motion, which is nowhere differentiable almost surely.
remmullerEfficientEstimationParameter2010 study an inferential problem concerning time-varying parameters approximated by Gaussian processes and piece-wise continuous functions scaled by a factor of $T^{-1/2}$. Leading examples are $T^{-1/2}\beta(t/T)$ with $\beta(\cdot)$ continuous on $[0,1]$ and $T^{-1/2}B_1(t/T)$, which is approximately equivalent (in distribution) to $T^{-1}\sum_{i=1}^{t}u_i=O_p(1/\sqrt{T})$. Therefore, non-vanishing smooth functions and random walks are not considered in their framework.
exm[Structural breaks] Structural breaks in parameters have attracted attention casiniStructuralBreaksTime2018. Suppose time-varying coefficient $\beta_{T,t}$ experiences one abrupt break during the sample period: \begin{align} \beta_{T,t} = \begin{cases} \beta_1 & for \ t=1,2,\ldots,T_B \\ \beta_2 & for \ t=T_B + 1, T_B + 2,\ldots,T \\ \end{cases}, \end{align} where $T_B = \lfloor \tau_B T\rfloor, \ \tau_B \in (0,1)$, and $\| \beta_1 - \beta_2\| = \delta/T^{\alpha}$ for some $\delta>0$ and $\alpha>0$. Under this formulation, the break is of shrinking magnitude, as considered in baiEstimationChangePoint1997. $\beta_{T,t}$ belongs to the type-b $\mathrm{TVP}(\alpha)$ class on the boundary. Specifically, we have, for any $t \in\{1,\ldots,T_B-a_T\}\cup \{T_B+a_T+1,\ldots,T\}$, $\max_{j:|t-j|\leq a_T} \|\beta_{T,t}-\beta_{T,j}\| = 0$, and for any $t\in\{T_B-a_T+1,\ldots,T_B+a_T\}$, $\max_{j:|t-j|\leq a_T} \|\beta_{T,t}-\beta_{T,j}\| = \delta T^{-\alpha}$. Note that the asymptotically non-negligible discontinuity given by $\alpha=0$ is excluded from Definition (ref).
exm[Threshold models] hansenSampleSplittingThreshold2000 considers the threshold regression model obtained by letting $\beta_{T,t} = \theta_1 + \delta_T1\{q_t > \eta\}$, where $q_t$ is the threshold variable that determines the regime at time $t$, depending on whether it exceeds threshold parameter $\eta$. $\delta_T$, which hansenSampleSplittingThreshold2000 refers to as the threshold effect, expresses the magnitude of discontinuous changes in $\beta_{T,t}$. hansenSampleSplittingThreshold2000 assumes $\delta_T = c/T^{\alpha}$.\footnote{hansenSampleSplittingThreshold2000 also imposes $0<\alpha<1/2$, but this restriction is not necessary in our framework.} $\beta_{T,t}$ clearly satisfies $\|\beta_{T,t} - \beta_{T,j}\| \leq \|\delta_T\| = O_p(1/T^{\alpha})$, for all $t$ and $j$, which implies $\beta_{T,t}$ belongs to type-b $\mathrm{TVP}(\alpha)$.
exm[Mixed model] Suppose that $\beta_{T,t}$ is expressed as $\beta_{T,t}=\beta_{1,T,t} + \beta_{2,T,t}$, where $\beta_{1,T,t}$ is continuously differentiable and $\beta_{2,T,t} = \mu + (1/\sqrt{T})\sum_{i=1}^tu_i$ with $u_i$ defined as in Example 2. Then, it is straightforward to show that $\beta_{T,t}$ belongs to type-a $\mathrm{TVP}(1/2)$. More generally, for any finite positive integer $S$, if $\beta_{T,t}$ is expressed as the sum of $S$ time-varying parameters each of which belongs to the type-a $\mathrm{TVP}(\alpha_s)$ class ($s=1,\ldots,S)$, then $\beta_{T,t}$ belongs to type-a $\mathrm{TVP}(\min\{\alpha_1,\ldots,\alpha_S\})$.

Mixture models where $\beta_{T,t}$ is the sum of both type-a and type-b time-varying parameters will be considered in Section (ref).

Asymptotics

Assumptions

We suppose kernel $K(\cdot)$ satisfies the following condition:

asm$K(x)\geq0, \ x\in \mathbb{R}$, is Lipschitz continuous and has compact support $[-1,1]$. • $\int_{-1}^{1}K(x)dx = 1$.

Commonly used kernels such as the uniform density on $[-1,1]$ and the Epanechnikov kernel satisfy Assumption (ref). Following the arguments of giraitisInferenceStochasticTimevarying2014,giraitisTimevaryingInstrumentalVariable2021, kernels with non-compact support such as the Gaussian kernel are permitted under some stronger condition. We focus on kernels with a compact support as specified in condition (a) to avoid unessential complications. Note that $2Th$ is the effective sample size of the kernel-based estimation, since $K((t-i)/Th) = 0$ for $i$ such that $|t-i|> Th$.

Next, we impose the following assumption on model (ref).\footnote{For the definition of near epoch dependence (NED), see, e.g., davidsonStochasticLimitTheory1994.}

asm$\{(x_t', \varepsilon_t)\}_t$ is $\mathrm{L}_2$-NED of size $-(r-1)/(r-2)$ on an $\alpha$-mixing sequence of size $-r/(r-2)$ for some $r>2$, with respect to some positive constants $d_t$ satisfying $\sup_td_t < \infty$. Moreover, $\sup_t E[\|x_t\|^{2r}] + \sup_t E[|\varepsilon_t|^{2r}] < \infty$. • $\{x_t\varepsilon_t\}_t$ has mean zero and is serially uncorrelated. • For each $t=\lfloor Tr \rfloor, \ r\in(0,1)$, and $h$ such that $h\to0$ and $Th\to\infty$ as $T\to \infty$, there exist nonrandom symmetric matrices $\Omega(r)>0$ and $\Sigma(r)>0$ such that $(1/Th)\sum_{i=1}^{T}K((t-i)/Th)E[x_ix_i'] \to \Omega(r)$ and $\mathrm{Var}\Bigl((1/\sqrt{Th})\sum_{i=1}^{T}K((t-i)/Th)x_i\varepsilon_i\Bigr) \to \Sigma(r)$.

Assumption (ref)(a) allows the regressor and disturbance to be weakly serially dependent. The NED assumption is more general than mixing conditions commonly assumed in the literature caiTrendingTimevaryingCoefficient2007,chenTestingSmoothStructural2012,giraitisTimevaryingInstrumentalVariable2021,friedrichSieveBootstrapInference2024. Also note that we do not impose strict or covariance stationarity unlike earlier works caiTrendingTimevaryingCoefficient2007,chenTestingSmoothStructural2012,friedrichSieveBootstrapInference2024, and thus our framework allows for heteroskedasticity in $\varepsilon_t$. Assumption (ref)(b) requires that the product of regressors and disturbance be serially uncorrelated, which is satisfied when, for example, $\varepsilon_t$ is a martingale difference sequence (m.d.s.) with respect to $\mathcal{F}_{T,t} \coloneqq \sigma(\{x_{t+1},x_t,\varepsilon_t,x_{t-1},\varepsilon_{t-1},\ldots\})$. The assumption of no serial correlation or m.d.s. is common in the literature chenTestingSmoothStructural2012,kristensenNonparametricDetectionEstimation2012,giraitisTimevaryingInstrumentalVariable2021. Assumption (ref)(c) holds under Assumptions (ref)(a)-(b) if $x_t$ and $x_t\varepsilon_t$ are covariance-stationary (see Corollary (ref)).

Asymptotic properties of $\hat{\beta}_t$

In the following theorem, we establish the consistency and asymptotic normality of the kernel-based estimator, $\hat{\beta}_t$.

thmSuppose Assumptions (ref) and (ref) hold. Then, for $t=\lfloor Tr \rfloor, \ r\in (0,1)$, we have \begin{align} \sqrt{Th}(\hat{\beta}_t - \beta_{T,t} - R_{T,t}) \stackrel{d}{\to} N(0,\Omega(r)^{-1}\Sigma(r)\Omega(r)^{-1}), \end{align} where \begin{align} R_{T,t} = \begin{cases} O_p(h^{\alpha}) & if $\beta_{T,t}$ satisfies Definition (ref)(a) \\ O_p(T^{-\alpha}) & if $\beta_{T,t}$ satisfies Definition (ref)(b) \end{cases}. \end{align} In particular, for $h =cT^{\gamma}, \ c>0, \ \gamma\in(-1,0)$, we have \mathtoolsset{showonlyrefs=false} \begin{align} \sqrt{cT^{1+\gamma}}(\hat{\beta}_t - \beta_{T,t}) \stackrel{d}{\to} N(0,\Omega(r)^{-1}\Sigma(r)\Omega(r)^{-1}), \end{align} \mathtoolsset{showonlyrefs=true} for $\gamma \in \Gamma(\alpha)$, where \begin{align} \Gamma(\alpha) = \begin{cases} (-1,-\frac{1}{2\alpha+1}) & if $\beta_{T,t}$ satisfies Definition (ref)(a) \\ (-1,2\alpha -1)\cap (-1,0) & if $\beta_{T,t}$ satisfies Definition (ref)(b) \end{cases}. \end{align} If $\beta_{T,t}$ belongs to $\mathrm{TVP}(\alpha)$ on the boundary, $\gamma$ close to the right endpoint of $\Gamma(\alpha)$ gives asymptotic normality and the fastest possible convergence rate.
remWe do not derive the asymptotic distribution of $\hat{\beta}_t$ at boundary points (near $t=0$ and $t=T$), but the derivation will proceed along the lines of caiTrendingTimevaryingCoefficient2007. As shown by caiTrendingTimevaryingCoefficient2007, the local constant estimator suffers from a larger bias at boundary points than the local linear estimator if $\beta_{T,t}$ is continuously differentiable. However, as discussed soon later (in Example (ref) below), the local linear estimator is available only when $\beta_{T,t}$ is (continuously) differentiable and is not applicable to nondifferentiable time-varying parameters such as the random walk. To accommodate both differentiable and nondifferentiable time-varying parameters, we focus on the local constant estimator.
corSuppose Assumptions (ref) and (ref)(a)-(b) hold. Suppose also that $\{x_t\}_t$ and $\{x_t\varepsilon_t\}_t$ are covariance-stationary. Then, (ref)-(ref) hold with $\Omega(r)$ and $\Sigma(r)$ replaced by $\Omega \coloneqq E[x_1x_1']$ and $\Sigma \coloneqq \int_{-1}^{1}K(x)^2dxE[\varepsilon_1^2x_1x_1']$, respectively.

In what follows, we will set $h=cT^{\gamma}$ and call $\gamma$ (as well as $h$) the bandwidth parameter.

If $\beta_{T,t}$ belongs to type-a $\mathrm{TVP}(\alpha)$ on the boundary, then it can be estimated by setting $\gamma \approx -1/(2\alpha+1)$, yielding a convergence rate $\approx T^{-\alpha/(2\alpha+1)}$. The same convergence rate has been established in the literature for the minimax risk of kernel-based estimators, assuming that the parameter of interest belongs to a H\"{o}lder class with exponent $\alpha$ tsybakovIntroductionNonparametricEstimation2009. Theorem (ref) shows that an analogous result holds under our definition of smoothness, Definition (ref)(a).

Smoothness parameter $\alpha$ affects $\Gamma(\alpha)$, the set of bandwidth parameter $\gamma$ that yields $\sqrt{Th}$-consistency and asymptotic normality. This makes the rate of convergence $T^{(1+\gamma)/2}$ dependent on $\alpha$. Letting $\alpha \to 0$, the kernel-based estimation can accommodate time-varying parameters of arbitrary smoothness, but this is accompanied by $\Gamma(\alpha) \to -1$, resulting in the rate of convergence $T^{(1+\gamma)/2} \to 1$. In contrast, if we let $\alpha\to\infty$, then $\Gamma(\alpha)$ tends to $(-1,0)$, and the choice $\gamma \approx 0$ yields a nearly $\sqrt{T}$-rate convergence, but only highly smooth parameters can be estimated. This observation shows that there is a trade-off between the rate of convergence and the size of the class of the time-varying parameters that can be estimated.

Because $\Gamma(\alpha)$ is the set of $\gamma$ that yields $\sqrt{Th}$-consistency and asymptotic normality under given $\alpha$, we can obtain the set of $\alpha$ that leads to $\sqrt{Th}$-consistency and asymptotic normality of $\hat{\beta}_{t}$ under given $\gamma$, by inverting the expression of $\Gamma(\alpha)$. Letting $A(\gamma)$ denote such a set, we can say that $\hat{\beta}_t$ calculated using given $\gamma$ is $\sqrt{Th}$-consistent and asymptotically normal for time-varying parameters with smoothness $\alpha\in A(\gamma)$, where

align[align omitted — 299 chars of source]

Letting $\gamma \to -1$, $A(\gamma)$ tends to $(0,\infty)$, which implies that time-varying parameters with any smoothness $\alpha>0$ can be estimated, but the rate of convergence becomes $T^{(1+\gamma)/2} \to 1$. On the other hand, if we let $\gamma\uparrow0$, then the rate of convergence is as fast as $\sqrt{T}$, but $A(\gamma) \to \infty$ (the smoothness of constant parameters) in the type-a case. Hence, the bandwidth determines the trade-off between efficiency and robustness.

\setcounter{exm}{0}

exm[Continued] Because continuously differentiable $\beta_{T,t}$ belongs to the type-a $\mathrm{TVP}(1)$ class, for any $\gamma \in \Gamma(1) = (-1,-1/3)$, we have $\sqrt{cT^{1+\gamma}}(\hat{\beta}_t - \beta_{T,t}) \stackrel{d}{\to} N(0,\Omega(r)^{-1}\Sigma(r)\Omega(r)^{-1})$. Setting $\gamma \approx -1/3$ gives the fastest rate of convergence of $T^{1/3}$. If $\beta_{T,t}$ is twice continuously differentiable, and the kernel is symmetric, then the set of the admissible bandwidths, $\Gamma(\alpha)$, widens to $(-1,-1/5)$, giving the faster rate of convergence of $T^{2/5}$ caiTrendingTimevaryingCoefficient2007. In general, we will be able to enlarge $\Gamma(\alpha)$ to $(-1,-1/(4\alpha+1))\cup (-1,-1/3)$ in the type-a case if the following additional condition (mimicking the Taylor expansion) holds: \begin{align} \max_{j:|t-j|\leq a_T} \left\|\beta_{T,t} - \beta_{T,j} -c_t\left(\frac{t}{T} - \frac{j}{T}\right) \right\| = O_p\Bigl(\Bigl(\frac{a_T}{T}\Bigr)^{2\alpha}\Bigr), \end{align} for some (possibly random) bounded vector $c_t$. Condition (ref), however, essentially requires differentiability of $\beta_{T,t}$ with respect to time, which is not satisfied by, e.g., the random walk divided by $\sqrt{T}$, so that the enlarged version of $\Gamma(\alpha)$ is only available to a limited class of time-varying parameters. For the same reason, the local linear estimator, which is based on the Taylor expansion of $\beta_{T,t}$, is not applicable to nondifferentiable time-varying parameters.
exm[Continued] If $\beta_{T,t}$ is the random walk divided by $\sqrt{T}$, then $\Gamma(1/2)=(-1,-1/2)$, and thus the fastest rate of convergence given by $\gamma \approx -1/2$ is $T^{1/4}$, slower than $T^{2/5}$ in the continuously differentiable case. The same set of admissible bandwidths is derived by giraitisInferenceStochasticTimevarying2014, who consider a random-walk type time-varying coefficient in the context of univariate AR(1) models. Furthermore, we show in Appendix B that the bandwidth minimizing the MSE of $\hat{\beta}_t$ is proportional to $T^{-1/2}$ when $\beta_{T,t}$ is the random walk divided by $\sqrt{T}$. We prove this result under more restrictive conditions than Definition (ref), and Assumptions (ref) and (ref). Therefore, the choice of $\gamma = -1/2$ may also be justified as the minimizer of the MSE of $\hat{\beta}_t$.
exm[Continued] Suppose $\beta_{T,t}$ is defined as in (ref). Because $\beta_{T,t}$ belongs to the type-b $\mathrm{TVP}(\alpha)$ class, arbitrary $\gamma$ in $(-1,0)$ yields the $\sqrt{Th}$-consistency and asymptotic normality of $\hat{\beta}_t$ as long as $\alpha\geq 1/2$. In particular, setting $\gamma \approx 0$ gives a near $\sqrt{T}$-consistency.\footnote{In fact, setting exactly $\gamma=0$ yields $\sqrt{T}$-consistency and asymptotic normality if $\alpha>1/2$. In this case, each $\beta_{T,t}, \ t=1,\ldots,T$ is estimated by using the full sample, but $\hat{\beta}_t$ and $\hat{\beta}_s$ ($t\neq s$) may take different values. This is because the weighting scheme (based on the kernel, $K(\cdot)$) is different for different time points. If $K(\cdot)$ is the uniform kernel, $\hat{\beta}_t$ equals the full-sample OLS estimator for all $t$.} For the case of $\alpha \in (0,1/2)$, however, a smaller $\alpha$ leads to a larger discontinuity in $\beta_{T,t}$ and thus a slower rate of convergence (through a narrower $\Gamma(\alpha)$). Therefore, if $\beta_{T,t}$ experiences large structural breaks given by $\alpha<1/2$, and if there is no other source of instability in the path of $\beta_{T,t}$, then a conventional structural-break approach that achieves $\sqrt{T}$-consistency baiEstimatingTestingLinear1998 will be more suitable.
exm[Continued] The argument given in Example 3 also applies to the threshold model: When $\delta_T=O_p(1/T^{\alpha})$ with $\alpha\geq 1/2$, the kernel-based method delivers a $\sqrt{Th}$-consistent, asymptotically normal estimation of $\beta_{T,t}$, whereas hansenSampleSplittingThreshold2000's (hansenSampleSplittingThreshold2000) method should be used when $\alpha<1/2$ and the threshold effect solely determines the parameter path.

Estimation of variance-covariance matrices

To conduct inference, one needs to consistently estimate the asymptotic variance of $\hat{\beta}_t$. Under Assumptions (ref) and (ref), $\Omega(r)$ can be consistently estimated by

align[align omitted — 123 chars of source]

see Lemma (ref) in Appendix A. A natural estimator of $\Sigma(r)$ is

align[align omitted — 146 chars of source]

where $\hat{\varepsilon}_i = y_i - x_i'\hat{\beta}_i$.\footnote{If $\{x_t\varepsilon_t\}$ is serially correlated, $\Sigma(r)$ is typically the long-run variance of $\{x_t\varepsilon_t\}$. In this case, an appropriate estimator of $\Sigma(r)$ would be a nonparametric kernel estimator such as the Newey-West one, as suggested by caiTrendingTimevaryingCoefficient2007. We do not explore in this direction to save space.} To prove the consistency of $\hat{\Sigma}(r)$, however, the current assumptions are not sufficient. This is because for each $t=\lfloor Tr \rfloor$, the estimation errors $\hat{\beta}_{t+j}-\beta_{T,t+j}$, where $j\in[-\lfloor Th \rfloor, \lfloor Th \rfloor]$, are required to be asymptotically negligible uniformly over $j\in[-\lfloor Th \rfloor, \lfloor Th \rfloor]$, on which $\hat{\Sigma}(r)$ is calculated. To ensure the uniform consistency of $\hat{\beta}_{t+j}$ over $j\in[-\lfloor Th \rfloor, \lfloor Th \rfloor]$, we need the following additional conditions.

asmAssumption (ref) holds with part (a) replaced by the following condition: \begin{itemize} • $\{(x_t', \varepsilon_t)\}_t$ is $\mathrm{L}_2$-NED of size $-2(r-1)/(r-2)$ on an $\alpha$-mixing sequence of size $-2r/(r-2)$ for some $r>2$, with respect to some positive constants $d_t$ satisfying $\sup_td_t < \infty$. Moreover, $\sup_t E[\|x_t\|^{2r}] + \sup_t E[|\varepsilon_t|^{2r}] < \infty$. \end{itemize}

Assumption (ref)(a') strengthens Assumption (ref)(a) by increasing the decaying rates of the mixing and NED coefficients, essentially weakening the serial dependence of $\{(x_t', \varepsilon_t)\}_t$.

asmThere exists some constant $\rho>0$ such that $\inf_{t\geq1}\lambda'E[x_tx_t']\lambda\geq \rho\left\|\lambda\right\|^2$ for any $\lambda\neq0$.

Assumption (ref) requires that there be enough variation in the data, as it implies that the minimum eigenvalue of $E[x_tx_t']$ is bounded away from zero uniformly in $t$.

asmFor each $t=\lfloor Tr \rfloor$, $r\in(0,1)$, it holds that \begin{align} \max_{-\lfloor Th \rfloor\leq j \leq \lfloor Th \rfloor}\left\|\frac{1}{Th}\sum_{i=1}^TK\left(\frac{t+j-i}{Th}\right)x_ix_i'\left(\beta_{T,i}-\beta_{T,t+j}\right)\right\| = o_p(1). \end{align}

Assumption (ref) is a high-level one that ensures the uniform consistency of $\hat{\beta}_{t+j}$ over $j\in[-\lfloor Th \rfloor, \lfloor Th \rfloor]$. In Appendix C, we show that Assumption (ref) is satisfied in the time-varying models and under the implied bandwidths discussed in Examples (ref)-(ref).

thmSuppose Assumptions (ref) and (ref)-(ref) hold. Then, for each $t=\lfloor Tr \rfloor, \ r\in (0,1)$, we have $\hat{\Sigma}(r) \stackrel{p}{\to} \Sigma(r)$.

On Bandwidth Selection: Implications and a Guide

In Theorem (ref), we showed the set of admissible bandwidth rates depends on the smoothness $\alpha$ of $\beta_{T,t}$. This implies that an improperly selected bandwidth rate (given by $\gamma \notin \Gamma(\alpha)$) leads to misleading inference. In this section, we illustrate this implication through some examples where the evolutionary mechanism of $\beta_{T,t}$ is misspecified. We also discuss how to choose the bandwidth in empirical studies.

When random-walk $\beta_{T,t}$ is assumed to be continuously differentiable

Suppose one assumes $\beta_{T,t}$ is a continuously differentiable function and sets $\gamma \approx -1/3$, but the fact is that $\beta_{T,t}$ follows the random walk divided by $\sqrt{T}$. Using the results given in Theorem (ref), it is readily shown that the kernel-based estimator satisfies $\sqrt{cT^{1+\gamma}}(\hat{\beta}_t - \beta_{T,t}) = S_{T,t} + O_p(T^{1/2+\gamma})$, where $S_{T,t} \stackrel{d}{\to}N(0,\Omega(r)^{-1}\Sigma(r)\Omega(r)^{-1})$. Since the bias term is of order $O_p(T^{1/2+\gamma})$ and $\gamma > -1/2$, the difference $\hat{\beta}_{t} - \beta_{T,t}$ is dominated by the bias term. Because the bias term is not normal in general, confidence intervals based on a normal approximation will perform poorly.

In the literature on smooth (differentiable) time-varying parameters, researchers often use a rule-of-thumb or plug-in bandwidth $h=\mathrm{constant}\times T^{-1/5}$, or pick the bandwidth minimizing the cross-validation criterion over $h\in [c_1T^{-1/5},c_2T^{-1/5}]$ for some $0<c_1<c_2$ zhouSimultaneousInferenceLinear2010,zhangInferenceTimeVaryingRegression2012,kristensenNonparametricDetectionEstimation2012,chengBayesianBandwidthEstimation2019,sunPenalizedTimevaryingModel2023. Although these selection rules lead to an efficient estimation of $\beta_{T,t}$ as long as it is correctly specified as a continuously differentiable function, they will yield a biased estimation if $\beta_{T,t}$ is a random walk, or more generally, if $\beta_{T,t}$ does not belong to type-a $\mathrm{TVP}(1)$.

The effect of neglected breaks

Suppose $\beta_{T,t}=\mu_{T,t} + (1/\sqrt{T})\sum_{i=1}^{t}u_i$, where $u_t$ is defined as in Example (ref), and $\mu_{T,t}$ satisfies

align[align omitted — 177 chars of source]

with $T_B = \lfloor \tau_BT \rfloor$ and $\mu_2-\mu_1 = \delta/T^{\alpha}$. Then, we can show

align[align omitted — 262 chars of source]

where $R_{T,t}$ is defined in (ref) and (ref). The asymptotic order of the bias term, $R_{T,t}$, is $O_p(T^{\gamma/2})$ for $t$ outside the $\lfloor Th \rfloor$-neighborhood of break point $T_B$. On the $\lfloor Th \rfloor$-neighborhood of $T_B$, it is $O_p(T^{\gamma/2})$ if $\alpha\geq-\gamma/2$, while it is $O_p(T^{-\alpha})$ if $0<\alpha<-\gamma/2$.

Suppose we estimate $\beta_{T,t}$ by $\hat{\beta}_t$ assuming $\mu_{T,t} = \mu$, that is, the parameter instability is purely due to the zero-mean random walk. In this case, the (misleading) optimal rate of convergence is achieved by the choice of $\gamma=-1/2$, yielding $R_{T,t}=O_p(T^{-1/4})$ for $t\in[1,T_B-\lfloor Th \rfloor]\cup[T_B+1+\lfloor Th \rfloor,T]$, and

align[align omitted — 143 chars of source]

for $t\in[T_B-\lfloor Th \rfloor+1, T_B+\lfloor Th \rfloor]$.

When $\alpha\geq1/4$, the asymptotic order of $R_{T,t}$ is $O_p(T^{-1/4})$ for all $t$, the same order as in the pure random walk case (see (ref)), so that the choice $\gamma =-1/2$ is valid and leads to the fastest rate of convergence. In contrast, if $0<\alpha<1/4$, $R_{T,t} = O_p(T^{-\alpha})$ for $t=T_B \pm \lfloor rTh \rfloor, \ r\in [0,1]$. Because the asymptotically normal component of the decomposition of $\hat{\beta}_t - \beta_{T,t}$ is $O_p(T^{-(1+\gamma)/2}) = O_p(T^{-1/4})$, the asymptotic behavior of $\hat{\beta}_t - \beta_{T,t}$ is dominated by the bias term. Therefore, the structural break induces a severe bias in the kernel-based estimation on the $\lfloor Th \rfloor$-neighborhood of the discontinuity point $t=T_B$.

The above result tells us that, if we set $\gamma=-1/2$, abrupt breaks of size $1/T^{\alpha}$ are absorbed in random walk parameter instabilities if $\alpha\geq1/4$, while the abrupt breaks “stick out" and cause bias when $\alpha<1/4$.

A more general result can be derived if we invoke $\Gamma(\alpha)$ and $A(\gamma)$ defined in (ref) and (ref), respectively. Suppose that $\beta_{T,t}$ can be expressed as $\beta_{T,t}^1 + \mu_{T,t}$, where $\beta_{T,t}^1$ belongs to type-a $\mathrm{TVP}(\alpha_1)$ with $\alpha_1 \in (0,\infty)$, and $\mu_{T,t}$ is defined as in (ref) but the magnitude of the break is $\mu_2-\mu_1 = \delta/T^{\alpha_2}$. Then, $\hat{\beta}_t$ is $\sqrt{Th}$-consistent and asymptotically normal under any $\gamma \in \Gamma(\alpha_1)=(-1,-1/(2\alpha_1+1))$ for $t\in[1,T_B-\lfloor Th \rfloor]\cup[T_B+1+\lfloor Th \rfloor,T]$. On $[T_B-\lfloor Th \rfloor+1, T_B+\lfloor Th \rfloor]$, the abrupt break is absorbed in $\beta_{T,t}^1$, and the same $\gamma$ leads to $\sqrt{Th}$-consistency and asymptotic normality if $\alpha_2\in A(\gamma) = ((1+\gamma)/2, \infty)$, while the bias term dominates the asymptotically normal term if $\alpha_2<(1+\gamma)/2$.

A guide for bandwidth selection

In the previous subsections, we have observed that an improperly selected $\gamma$ leads to misleading inference. Therefore, care must be taken in determining the bandwidth parameter. Because bandwidth parameter $h$ takes the form of $h=cT^{\gamma}$ with $c>0$ and $\gamma<0$, we first discuss how to determine $\gamma$ and then how to select $c$.

If one can identify the evolutionary mechanism of $\beta_{T,t}$ based on some prior information, they may select $\gamma$ appropriately, referring to the theoretical results derived in Section (ref). For instance, if the random walk coefficient model is plausible, $\gamma = -1/2$ is an appealing choice. If $\beta_{T,t}$ is known to be twice continuously differentiable, then various methods for bandwidth selection proposed in the literature can be used to determine $\gamma$ and $c$ jointly (e.g., zhangInferenceTimeVaryingRegression2012).

remWhatever $\gamma \in (-1,0)$ may be selected, abrupt breaks and threshold effects of size $1/T^{\alpha}$ lead to biased estimation around the discontinuity points if $\alpha<(1+\gamma)/2$; recall Section (ref). To avoid facing bias around the discontinuity points, one may be tempted to split the sample using some test for structural breaks baiEstimatingTestingLinear1998 or hansenSampleSplittingThreshold2000's (hansenSampleSplittingThreshold2000) approach, and then apply kernel regression within each subsample. However, our simulation shows that these sample-splitting approaches may lead to a misleading conclusion if latent discontinuous changes are mixed with smooth parameter changes. According to the simulation results, structural break tests can both underestimate and overestimate the number of discontinuous changes with a nonnegligible (or large in some cases) probability. Underestimating the number of discontinuous breaks implies that some latent abrupt breaks are overlooked, and an overestimation implies that spurious abrupt breaks are detected. Therefore, conventional structural break tests probably are not suitable for detecting abrupt breaks if they are mixed with smooth parameter changes. See Appendix D for details.

Determining $\gamma$ is more delicate when there is no prior information that helps identify the evolutionary mechanism of $\beta_{T,t}$. Here, we propose two data-driven procedures to select the value of $\gamma$, which require prespecified lower and upper bounds $\underline{\gamma}, \overline{\gamma}\in(-1,0)$ to construct the set of candidate $\gamma$ values, $[\underline{\gamma},\overline{\gamma}]$. The theory developed in this article will serve as a guiding principle in determining $\underline{\gamma}$ and $\overline{\gamma}$.

The first procedure we propose is a naive cross-validation-based method: For each $\gamma\in[\underline{\gamma},\overline{\gamma}]$ (on some grid), calculate $\hat{\beta}_{-t,m}(\gamma)$, where $\hat{\beta}_{-t,m}(\gamma)$ are the leave-$(2m+1)$-out local constant estimators calculated without the data on $s\in[t-m,t+m]$ and with $h=T^{\gamma}$ for some $m\in\mathbb{N}\cup\{0\}$, compute the cross-validation criterion, $\mathrm{CV}(\gamma)\coloneqq T^{-1}\sum_{t=1}^T(y_t-x_t'\hat{\beta}_{-t,m}(\gamma))^2$, and then pick the minimizer of $\mathrm{CV}(\gamma)$. One may use the generalized cross-validation (GCV) considered in zhouSimultaneousInferenceLinear2010,zhangInferenceTimeVaryingRegression2012.

The second procedure is based on fixed-design wild bootstrap goncalvesBootstrappingAutoregressionsConditional2004.

algo[Bootstrap-based] \begin{itemize} • For each $\gamma_1\in[\underline{\gamma},\overline{\gamma}]$, calculate the local constant estimators with $h=h_1\coloneqq T^{\gamma_1}$, denoted by $\hat{\beta}_t(\gamma_1)$, and obtain residuals $\hat{\varepsilon}_t(\gamma_1)=y_t-x_t'\hat{\beta}_t(\gamma_1)$. • For each $\gamma_1\in[\underline{\gamma},\overline{\gamma}]$, apply fixed-design wild bootstrap to resample $y_t$: $y_t^*(\gamma_1)=x_t'\hat{\beta}_t(\gamma_1)+\varepsilon_t^*(\gamma_1)$, where $\varepsilon_t^*(\gamma_1)\coloneqq \eta_t\hat{\varepsilon}_t(\gamma_1)$ and $\eta_t \sim \mathrm{i.i.d.} \ N(0,1)$ independent of the data. For each $\gamma_2\leq\gamma_1$, calculate the local constant estimators with $h=h_2=T^{\gamma_2}$ using $(y_t^*(\gamma_1), x_t)$: \begin{align} \hat{\beta}_t^*(\gamma_1,\gamma_2) \coloneqq \left(\sum_{i=t-\lfloor Th_2\rfloor}^{t+\lfloor Th_2\rfloor}K\left(\frac{t-i}{Th_2}\right)x_ix_i'\right)^{-1} \sum_{i=t-\lfloor Th_2\rfloor}^{t+\lfloor Th_2\rfloor}K\left(\frac{t-i}{Th_2}\right)x_iy_i^*(\gamma_1). \end{align} • For each pair $(\gamma_1,\gamma_2)$, construct the $100(1-q)\%$ confidence intervals for $\hat{\beta}_t(\gamma_1)$ based on $\hat{\beta}_t^*(\gamma_1,\gamma_2)$, its standard error, and the quantile of $N(0,1)$, and compute the empirical coverage rates (obtained from $B$ bootstrap intervals), denoted by $\mathrm{CR}(\gamma_1,\gamma_2)$. • The selected value is the largest $\gamma_1$ such that $\mathrm{CR}(\gamma_1,\gamma_2)\geq 1-\bar{q}$ for all $\gamma_2\leq\gamma_1$ and some tolerance level $\bar{q}$; that is, $\hat{\gamma} = \max\Upsilon$, where \begin{align} \Upsilon \coloneqq \{\gamma_1:\gamma_1\in[\gamma,\overline{\gamma}], \mathrm{CR}(\gamma_1,\gamma_2)\geq 1-\bar{q} \ for all \ \gamma_2\in[\gamma,\gamma_1]\}. \end{align} If $\Upsilon$ is empty, then $\hat{\gamma} = \underline{\gamma}$. \end{itemize}

The rationale behind Algorithm (ref) is as follows. If $\gamma_1$ is sufficiently small that $\hat{\beta}_t(\gamma_1)$ is $\sqrt{Th_1}$-consistent for $\beta_{T,t}$, then $\hat{\varepsilon}_t(\gamma_1)$ are good approximations of unobserved $\varepsilon_t$, and bootstrap sample $y_t^*(\gamma_1)$ generated from $\varepsilon_t^*(\gamma_1)$ is “well-behaved". Treating $\hat{\beta}_t(\gamma_1)$ as the pseudo-true parameters, $\hat{\beta}_t^*(\gamma_1,\gamma_2)$ ($\gamma_2\leq\gamma_1$) are $\sqrt{Th_2}$-consistent for $\hat{\beta}_t(\gamma_1)$ and asymptotically normal under the bootstrap probability measure, in probability. Then, the confidence interval for $\hat{\beta}_t(\gamma_1)$ based on $\hat{\beta}_t^*(\gamma_1,\gamma_2)$ and $N(0,1)$ should attain empirical coverage rates close to the nominal confidence level, with high probability. In Step 4, we pick the largest $\gamma_1$ such that the above argument applies, so that the fastest possible convergence rate can be achieved.

To theoretically justify the above reasoning, we impose the following regularity condition, strengthening Assumption (ref):

asmFor each $t=\lfloor Tr \rfloor$, $r\in(0,1)$, it holds that \begin{align} \max_{-\lfloor Th_2 \rfloor\leq j \leq \lfloor Th_2 \rfloor}\left\|\frac{1}{Th_1}\sum_{i=1}^TK\left(\frac{t+j-i}{Th_1}\right)x_ix_i'\left(\beta_{T,i}-\beta_{T,t+j}\right)\right\| = o_p(1/\sqrt{Th_2}). \end{align}

If $\beta_{T,t}$ satisfies Condition (ref) given in Appendix C, then Assumption (ref) holds when $2\alpha\gamma_1+\gamma_2<-1$. If $\beta_{T,t}$ is a rescaled random walk, under Condition (ref) given in Appendix C, Lemma 5(iii) of giraitisTimevaryingInstrumentalVariable2021 shows that Assumption (ref) holds when $\gamma_1+\gamma_2<-1$.

thmSuppose that Assumptions (ref), (ref), (ref), and (ref) hold, and that $\beta_{T,t}$ belongs to type-a $\mathrm{TVP}(\alpha)$. If $\gamma_1\in\Gamma(\alpha)=(-1,-(2\alpha+1)^{-1})$, we have, for each $t=\lfloor Tr \rfloor, \ r\in (0,1)$, and for $\gamma_2\in[\underline{\gamma},\gamma_1]$, \begin{align} \sup_{x\in\mathbb{R}^p}\left|P^*\left(\sqrt{Th_2}\left(\hat{\beta}_t^*(\gamma_1,\gamma_2)-\hat{\beta}_t(\gamma_1)-R_{T,t}^*\right)\leq x\right) - P\left(Z\leq x\right)\right| \stackrel{p}{\to}0, \end{align} where $P^*$ denotes the probability measure induced by the fixed-design wild bootstrap, $Z\sim N(0, \Omega(r)^{-1}\Sigma(r)\Omega(r)^{-1})$, and \begin{align} R_{T,t}^* = \begin{cases} O_{p^*} \left(1/\sqrt{Th_2}\right) & \mathrm{if} \ \gamma_2=\gamma_1 \\ o_{p^*} \left(1/\sqrt{Th_2}\right) & \mathrm{if} \ \gamma_2<\gamma_1 \end{cases}, \end{align} with arbitrarily high probability for sufficiently large $T$.
remThe statement of Theorem (ref) still holds if $\beta_{T,t}$ has type-b discontinuities of size $1/T^{\alpha_1}$ with $\alpha_1\in A(\gamma_1)=((1+\gamma_1)/2, \infty)$.
remAlgorithm (ref) may reject a valid choice $\gamma_1\in\Gamma(\alpha)$ and suggest a conservative $\hat{\gamma}$ (undersmoothing) if the (bootstrap) distribution of $\hat{\beta}_t^*(\gamma_1,\gamma_1)-\hat{\beta}_t(\gamma_1)$ is poorly approximated by the normal distribution. There are two cases where this normal approximation is poor. First, if $T$ and $\gamma_1$ are small, the effective sample size can be quite small.\footnote{For example, if $\gamma_1=-1/2$, the effective sample size is as small as $2\lfloor Th\rfloor=28$ when $T=200$.} Second, the bias term, $R_{T,t}^*$, in Theorem (ref) may be of the same order as the asymptotically normal part of $\hat{\beta}_t^*(\gamma_1,\gamma_1)-\hat{\beta}_t(\gamma_1)$ and distort the distribution of $\hat{\beta}_t^*(\gamma_1,\gamma_1)-\hat{\beta}_t(\gamma_1)$.\footnote{A solution in the second case would be to correct the bias term, $R_{T,t}^*$, but this requires an explicit formula for $R_{T,t}^*$, which seems not possible under the quite general smoothness condition, Definition (ref). For example, bias formulae are typically derived assuming $\beta_{T,t}$ is twice continuously differentiable caiTrendingTimevaryingCoefficient2007,zhouSimultaneousInferenceLinear2010. Bias correction in our general framework is left for future research.} While a conservative $\hat{\gamma}$ yields an asymptotically unbiased estimation, it causes an efficiency loss. In spite of such a limitation, Algorithm (ref) leads to a more efficient estimation than the most conservative choice $\gamma=\underline{\gamma}$, at least when $T$ is sufficiently large; see the simulation results in Section (ref).

Once $\gamma$ is determined, one can select $c$ by minimizing some criterion that is a function of $c$. For instance, the cross-validation criterion, $\mathrm{CV}(c) \coloneqq T^{-1}\sum_{t=1}^T(y_t - x_t'\hat{\beta}_{-t}(c,\hat{\gamma}))^2$, where $\hat{\beta}_{-t}(c,\hat{\gamma})$ are the leave-one-out kernel estimators calculated under $h=cT^{\hat{\gamma}}$, can be used. Some other criteria may be used to determine $c$ such as the AIC as suggested in caiTrendingTimevaryingCoefficient2007 or the GCV considered in zhouSimultaneousInferenceLinear2010 and zhangInferenceTimeVaryingRegression2012.

Monte Carlo Simulation

In this section, we conduct three Monte Carlo experiments to verify the implications provided in Section (ref). We use the following DGP: $y_t = \beta_{T,t}x_t + \varepsilon_{t}, \ t=1,\ldots,T$, where $x_t = 0.5x_{t-1} + \varepsilon_{x,t}$ with $\varepsilon_{x,t} \sim \mathrm{i.i.d.} \ N(0,1)$, and $\beta_{T,t}$ is defined differently in different experiments. For the specification of $\varepsilon_{t}$, we consider two cases: $\varepsilon_{t} = u_t$, where $u_t \sim \mathrm{i.i.d.} \ N(0,1)$ (i.i.d. case) and $\varepsilon_{t} = \sigma_tu_t$ with $\sigma_t^2 = 0.1 + 0.3\varepsilon_{y,t-1}^2 + 0.6\sigma_{t-1}^2$ (GARCH case). To obtain $\hat{\beta}_t$, we use the Epanechnikov kernel $K(x) = 0.75(1-x^2)I(|x|\leq 1)$.

Simulation for Section (ref)

The first experiment is related to Section (ref), and $\beta_{T,t}$ is generated as the rescaled random walk: $\beta_{T,t} = T^{-1/2}\sum_{i=1}^{t}v_i$. We consider two DGPs for driver process $v_t$: (i)$v_t\sim \mathrm{i.i.d.} \ N(0,1)$ and (ii) $v_t \sim \mathrm{i.i.d.} \ \mathrm{log \ normal}$ with parameters $\mu=0$ and $ \sigma = 1$.\footnote{Specifically, $X$ follows a log normal distribution if $X=\exp(Z)$, where $Z \sim N(\mu,\sigma^2)$.} Four sample sizes are used: $T \in \{100,200,400,800\}$. To evaluate the global performance of $\hat{\beta}_{t}$, we calculate $\mathrm{MSE} = T^{-1}\sum_{t=1}^{T}(\hat{\beta}_t - \beta_{T,t})^2$ (the reported MSE is the mean MSE over 2000 replications). To evaluate the normal approximation given in Corollary (ref), we construct the $95\%$ confidence interval for $\beta_{T,0.5T}$ (the middle point of the sample). The variance estimators are $\hat{\Omega} \coloneqq T^{-1}\sum_{i=1}^{T}x_i^2$ and $\hat{\Sigma} \coloneqq \int_{-1}^{1}K(x)^2dx\times T^{-1}\sum_{i=1}^{T}\hat{\varepsilon}_i^2x_i^2$. We experiment with bandwidth parameter $h=T^{\gamma}$ and $\gamma \in \{-0.2,-0.33,-0.5,-0.55,-0.6,-0.7\}$, and evaluate the performance for each pair $(\gamma,T)$. We also analyze the performance of the data-driven selection procedures for $\gamma$ suggested in Section (ref).\footnote{For the cross-validation-based method, we select $\hat{\gamma}$ from $[-0.5,-0.2]$ using leave-three-out estimators. For the bootstrap-based method, we select $\hat{\gamma}$ from $\{-0.5,-0.4,-0.33,-0.2\}$ and set the tolerance level $\bar{q}=0.1$.} The results are presented in Tables (ref) and (ref).

Because $\beta_{T,t}$ is the random walk divided by $\sqrt{T}$, our theoretical results predict that an appropriate bandwidth is $\gamma \approx -1/2$, while the kernel-based estimator leads to poor inference when $\gamma > -1/2$. Our simulation result corroborates this analysis. First, consider the case where $\varepsilon_{t} \sim \mathrm{i.i.d.} \ N(0,1)$ (Table (ref)). In case (i) (Gaussian random-walk $\beta_{T,t}$), when $\gamma = -0.2$, the coverage rate is far below the 95% confidence level. What is worse, it deviates from 0.95 as $T$ increases. Note that the MSE is relatively large. When $\gamma=-1/3$, the MSE takes the smallest value for all $T$ considered, but the coverage rate is still too small. This result warns researchers against using these bandwidths unless they are confident that $\beta_{T,t}$ can be well approximated by smooth functions with smoothness parameter $\alpha=1$. For $\gamma\leq -1/2$, the interval estimation performs well with coverage rate being 85-90% and getting better as $T$ increases. However, $\gamma = -0.7$ leads to undercoverage when $T$ is small and the largest MSE for all $T$. $\gamma = -0.6$ also gives large MSEs. The choices $\gamma \approx -1/2$ lead to good coverage and small MSE, so that these choices are recommended for random-walk type parameters, or more generally, for time-varying parameters belonging to $\mathrm{TVP}(1/2)$ on the boundary.

Next consider the performance of $\gamma=\hat{\gamma}$ selected by data-dependent procedures. For the cross-validation method, the mean MSE is close to the smallest MSE attained by the deterministic choice of $\gamma=-0.33$, particularly when $T$ is large. On the other hand, the coverage ratio is far below the nominal level and takes values between 73% and 78%, although the coverage gradually improves as the sample size increases. For the bootstrap-based selection, the mean MSE and the coverage ratio is almost identical to those attained by the deterministic choice of $\gamma=-0.5$; that is, the MSE is relatively large when $T$ is small, but improves quickly as $T$ increases, and the coverage ratio is reasonably good. Based on these observations, the cross-validation method seems useful when a small MSE is desired, while the bootstrap-based method is a reasonable choice when unbiased estimation is prioritized.

The result for case (ii) (non-Gaussian random-walk $\beta_{T,t}$) is similar, so the same comment applies.

Results for the case where $\varepsilon_{t}$ is GARCH (Table (ref)) are similar to those for the i.i.d case. Hence, we do not repeat the same analysis.

Simulation for Section (ref)

The second experiment is for verifying the implication provided in Section (ref). In this simulation, we analyze the effect of (neglected) structural breaks. For this purpose, we generate $\beta_{T,t}$ according to $\beta_{T,t} = \mu_{T,t} + T^{-1/2}\sum_{i=1}^{t}v_i$, where $v_i \sim \mathrm{i.i.d.} \ N(0,1)$ and $\mu_{T,t}$ is an intercept term experiencing a break at $t=0.5T$. Specifically, we let

align[align omitted — 160 chars of source]

where $\alpha \in \{0.1,0.2,0.3,0.4\}$. A smaller $\alpha$ yields a larger break.

We consider estimating $\beta_{T,t}$ with the choice $h = T^{-1/2}$, reflecting the ignorance of the break. According to our theoretical analysis, the kernel-based estimator has a severe bias around $t=0.5T$ when $\alpha<0.25$, while breaks given by $\alpha>0.25$ have no effect asymptotically. To confirm this implication, we calculate the MSE and coverage rate of $\hat{\beta}_{t}$ for $t=\tau T$ with $\tau = 0.4,0.45,0.5,0.55,0.6$. The MSE is calculated for each $\tau$ as the mean squared error over 2000 replications, that is, $\mathrm{MSE}(\tau) = 2000^{-1}\sum_{i=1}^{2000}(\hat{\beta}_{\tau T}^{(i)} - \beta_{T,\tau T}^{(i)})^2$, where superscript $i$ signifies $\hat{\beta}_{\tau T}^{(i)}$ and $\beta_{T,\tau T}^{(i)}$ are obtained in the $i$th replication. We consider four sample sizes; (i) $T = 100$, (ii) $T=200$, (iii) $T=400$, and (iv) $T=800$. We use $\hat{\Omega}_t = (Th)^{-1}\sum_{i=1}^{T}K((t-i)/Th)x_i^2$ and $\hat{\Sigma}_t = (Th)^{-1}\sum_{i=1}^{T}K((t-i)/Th)^2\hat{\varepsilon}_i^2x_i^2$ as the variance estimators to evaluate the normal approximation given in Theorem (ref). Results are reported in Tables (ref) and (ref).

First, let us see the case of $\varepsilon_{t}$ being i.i.d. and $T=100$ (Table (ref), the row labeled (i)). The MSEs and coverage rates for $\tau = 0.4$ and $0.6$ are stable across $\alpha$. This is because the break only affects estimation around the discontinuity point, $t=0.5T$. The break has a severe effect on $\hat{\beta}_{\tau T}$ with $\tau=0.45,0.5,0.55$, both in terms of MSE and coverage. The smaller $\alpha$ is (i.e. the larger the break is), the worse the performance gets. Moreover, this effect is more profound for $\tau$ closer to $0.5$. In terms of the coverage rate, smaller breaks given by $\alpha \geq 0.25$ have a nonnegligible effect. This indicates that, although breaks of these magnitudes asymptotically have no impact, they do have nontrivial effects in finite samples.

For case (ii) ($T=200$), MSEs for $\tau=0.5$ and $\alpha<0.25$ are still large. Note that MSEs for $\tau = 0.45,0.55$ are comparable with those for $\tau = 0.4,0.6$. This is because the abrupt break affects $\hat{\beta}_{t}$ on the $Th$-neighborhood of the break date. Because $t=0.45T$ and $t=0.55T$ are outside the $Th$-neighborhood of $0.5T$, the performance of $\hat{\beta}_{\tau T}$ improves as $T$ increases for $\tau = 0.45,0.55$. $\hat{\beta}_{0.5T}$ also suffers from poor coverage for all $\alpha$. For the cases with $T=400,800$ (cases (iii) and (iv)), a similar comment applies. In particular, the coverage rates for $\alpha < 0.25$ and $\tau=0.5$ deteriorate as $T$ increases.

Examining the case with $\varepsilon_{t}$ being GARCH (see Table (ref)), the same conclusion is drawn, so the detail is omitted.

Balance between robustness and efficiency

In this subsection, we investigate the finite-sample performance of the data-driven bandwidth selection procedures in an environment where $\beta_{T,t}$ evolves smoothly but experiences a jump at some point. Specifically, we specify $\beta_{T,t}$ as $\beta_{T,t} = \beta(t/T)$, where $\beta(x) = x + \mu_T(x)$ with $\mu_T(x) = 0$ for $x\leq 0.5$ and $\mu_T(x) = 1.5/T^{0.4}$ for $x>0.5$. $\beta_{T,t}$ evolves smoothly and deterministically over time but experiences a break at the middle point. Recalling the theoretical analysis in Section (ref), the break of size $T^{-0.4}$ can be accommodated as long as $\gamma < -0.2$. We analyze the finite-sample performance of $\hat{\beta}_{T,t}$ with $\gamma \in \{-0.2,-0.33,-0.5\}$ and $\gamma$ selected by the data-driven methods. We are interested in (i) whether the data-driven procedures can select $\gamma < -0.2$ (unbiasedness) and (ii) whether the selected $\gamma$ is close to $-0.2$ (efficiency). We study the mean MSE and the empirical coverage ratio at the break point, $t=0.5T$. The results are reported in Table (ref). Because the results for the cases with i.i.d. error and GARCH error are qualitatively similar, we comment on the i.i.d. case only.

We first consider the deterministic $\gamma$. Although $\gamma=-0.2$ yields the smallest MSE, this choice results in undercoverage at the break point, as expected. For both choices $\gamma=-0.33,-0.5$, the coverage rate improves as $T$ increases, but $\gamma=-0.33$ gives a much smaller MSE. Next consider the data-dependent procedures. For the cross-validation method, the mean MSEs are almost identical to those obtained under $\gamma=-0.33$, but the empirical coverage rate is well below the nominal rate. For the bootstrap method, the mean MSEs and coverage rates are almost identical to those obtained by $\gamma=-0.5$ for $T\in\{100,200,400\}$. When $T=800$, however, the bootstrap-based procedure improves the MSE by about 20% compared to $\gamma=-0.5$ while maintaining the same level of coverage rate.

Empirical Application

In this section, we apply kernel regression to estimate the time-varying CAPM.\footnote{The R code used for the empirical application is available on the author's website (\url{https://sites.google.com/view/mikihito-nishi/home}).} Parameter instabilities are widely observed in the CAPM literature ghyselsStableFactorStructures1998,lewellenConditionalCAPMDoes2006,famaValuePremiumCAPM2006,angCAPMLongRun2007,angTestingConditionalFactor2012,guoTimeVaryingBetaValue2017. We consider estimating the following factor model:

align[align omitted — 77 chars of source]

where $R_{j,t}$ denotes the excess return of portfolio $j$ at time $t$, and $R_{M,t}$ is the market excess return. The coefficients alpha and beta are allowed to be time-varying.

Background

In the CAPM literature, parameter instability is often modeled by letting parameters depend on observable instruments. But results drawn from this approach tend to be sensitive to the choice of instruments ghyselsStableFactorStructures1998. To overcome this problem, researchers have proposed time-varying parameter models that do not utilize exogenous information.

Some assume that parameters experience abrupt changes, and others model parameter instability via the (near) random walk or smooth functions of time. For example, famaValuePremiumCAPM2006 and lewellenConditionalCAPMDoes2006 split the sample assuming that parameter changes occur based on calendar time (e.g., monthly or yearly), and apply OLS within subsamples. However, estimates obtained in this fashion suffer from bias if the timing of structural breaks is misspecified. angCAPMLongRun2007 use a Bayesian approach assuming (near) random walk alpha and beta. liTestingConditionalFactor2011 and angTestingConditionalFactor2012 treat the parameters as deterministic continuously differentiable functions of time.

Given the fact that continuously differentiable functions and the random walk can be estimated under $\gamma=-1/5$ and $\gamma=-1/2$, respectively, we set the lower and upper bounds for $\gamma$ as $\underline{\gamma}=-0.5$ and $\overline{\gamma}=-0.2$.

Data

All data are extracted from Kenneth French's website (\url{https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html}). Following liCombinedApproachInference2015, we form three portfolios denoted by G, V, and G-V, respectively, from the 25 size-B/M portfolios. G is the average of the five portfolios in the lowest B/M quintile, V is the average of the five portfolios in the highest B/M quintile, and V-G is simply their difference. All the data are monthly, spanning 1952:1-2019:12 ($T=816$).

Results

We use the Epanechnikov kernel and $\hat{\Omega}_t = (Th)^{-1}\sum_{i=1}^{T}K((t-i)/Th)x_ix_i'$ and $\hat{\Sigma}_t = (Th)^{-1}\sum_{i=1}^{T}K((t-i)/Th)^2\hat{\varepsilon}_i^2x_ix_i'$ as the variance estimators, where $x_t = (1,R_{M,t})'$. To save space, we only discuss the result for portfolio V-G. The results for portfolios G and V are given in Appendix E.

Selection of the bandwidth

We determine two tuning parameters for the bandwidth, $h=cT^{\gamma}$, as explained in Section (ref). For Algorithm (ref), we construct 95% bootstrap confidence intervals and set the tolerance level to be $\bar{q}=0.1$, giving the threshold of 90% empirical coverage rate.

First, we consider selecting $\gamma$. Figure (ref) depicts the CV criterion computed using leave-($2m+1$)-out estimators for $m=0,1,2$. For $m=0,1$, the minimum is attained at $\gamma=-0.5$, whereas $\gamma=-0.32$ is the minimizer when $m=2$. Since there is little reason to prefer some specific value of $m$ to other values, we also use Algorithm (ref) to seek further evidence. Reported in Table (ref) are the mean empirical coverage rates taken over $t=1,\ldots,T$, $\overline{\mathrm{CR}}(\gamma_1,\gamma_2)\coloneqq T^{-1}\sum_{t=1}^T\mathrm{CR}_t(\gamma_1,\gamma_2)$. Each empirical coverage rate is calculated using $200$ bootstrap samples. For $\gamma_1=-0.33$ and $\gamma_1=-0.4$, the empirical coverage rates exceed the threshold of 0.9 for all $\gamma_2\leq\gamma_1$, and hence $\gamma_1=-0.33$ is supported by this procedure. Given these results, we set $\hat{\gamma}=-0.33$ since both CV- and bootstrap-based procedures support this choice. It is noteworthy that $\gamma=-1/5$, a prevalent choice in the literature, is rejected by our selection algorithm. This result highlights the importance of including other $\gamma$ values in the set of candidate bandwidths.

Given $\gamma=\hat{\gamma}$, we determine scaling constant $c$ via cross-validation. The selected value, $\hat{c}$, is the minimizer of the cross-validation criterion $\mathrm{CV}(c)$ over $c\in \{0.5,0.55,\ldots,1.5\}$.

Interval estimation

In Figure (ref), we plot the estimated time-varying alpha and its 95% confidence band.\footnote{This confidence band is obtained by sequentially calculating the pointwise 95% confidence intervals and is not a uniform 95% confidence band.} The estimated alpha fluctuates around the value zero throughout the sample period, and the confidence band includes zero at all time points. Figure (ref) depicts the estimated time-varying beta. It starts with a positive value that is significantly different from zero and then fluctuates around zero up to $t=300$. Then, it starts to decrease and stays below zero with the confidence band excluding zero. It starts to increase from $t=600$, and fluctuates around zero from $t=660$ toward the end of the sample.

Comparison with the Bayesian estimate

The CV-based selection procedure suggests that $\gamma=-1/2$ is partly supported by the data. Noting that this choice accommodates parameters following the (rescaled) random walk, and that random walk parameters are often estimated via Bayesian methods, it is interesting to compare the kernel-based estimates obtained from $h=\hat{c}T^{-1/2}$ with the estimates obtained from a Bayesian procedure in which parameters are assumed to be the random walk.

Let $\theta_{t} \coloneqq (\alpha_{t}, \beta_{t})'$. In the Bayesian method, we estimate the time-varying alpha and beta by using the Markov Chain Monte Carlo algorithm, assuming that $\theta_{t} = \theta_{t-1} + u_t$, where $u_t \sim N(0,D^2)$ with $D^2 = \mathrm{diag}(D_1^2,D_2^2)$.\footnote{For computation, we use the R package walker developed by helskeWalkerBayesianGeneralized2023.} As the prior distributions for parameters $\theta_{0}$, $D$ and $\mathrm{Var}(\varepsilon_t)=\sigma_{\varepsilon}^2$, we suppose $\theta_{0} \sim N(\mu\mathbf{1}_2,\sigma^2\mathbf{I}_2)$, $D_i \sim \mathrm{Gamma}(v_1,v_2), \ i=1,2$, and $\sigma_\varepsilon \sim \mathrm{Gamma}(\nu_1,\nu_2)$. We consider three configurations of hyperparameters. For each configuration, $(\mu,\sigma,v_2,\nu_1,\nu_2)$ are set to $(\mu,\sigma,v_2,\nu_1,\nu_2)= (0,32,10^{-4},2,10^{-4})$. The value of $v_1$ is varied, and we set $v_1 = 1, 2,$ and $4$.\footnote{We also changed the values for $(\mu,\sigma,v_2,\nu_1,\nu_2)$, but the estimates were insensitive to these parameters.}

In Figure (ref), we compare the Bayesian estimates with the kernel-based estimates with $h=\hat{c}T^{-1/2}$. For the estimated alpha (Figure (ref)), the trajectory obtained from the kernel method is more volatile (with a larger amplitude) than that obtained from the Bayesian algorithm, but the trajectories seem to share the same frequency. More striking is the similarity between the estimates of the time-varying beta. The estimated trajectories obtained from the two distinct methods are almost indistinguishable throughout the sample period, irrespective of the value of $v_1$.

We also compare the quantitative performances of the kernel and Bayesian estimators in terms of the in-sample fit (SSR). Standardizing the SSR obtained from the kernel-based method to be 1, the relative SSR's for the Bayesian estimators with $v_1=1,2,$ and $4$ are 1.019, 1.000, and 0.953, respectively. The kernel estimator yields a comparable in-sample fit relative to the Bayesian estimator.

Conclusion

We studied kernel-based estimation of time-varying parameters over a wide range of smoothness. We set up a general framework that quantifies the smoothness of the time-varying parameter by a single parameter $\alpha$, and established consistency and asymptotic normality of the kernel-based estimator under this framework. The results cover many important time-varying parameter models, including continuously differentiable functions, the rescaled random walk, abrupt structural breaks, the threshold regression model, and their mixtures.

Our analysis also highlights an often-overlooked role of the bandwidth and its implications on bandwidth selection. Beyond the bias-variance trade-off, when the parameter may be nondifferentiable, the bandwidth determines a trade-off between the rate of convergence and the size of the class of time-varying parameters that can be estimated. Theory and simulations show that the appropriate bandwidth rate depends on the smoothness of the time-varying parameter. In particular, a conventional $T^{-1/5}$-rate bandwidth is invalid in the case of nondifferentiable time-varying parameters such as the random walk. Another important implication from our result is that abrupt breaks of certain magnitudes cause bias in the kernel-based estimation.

Taking into account the diversity of existing time-varying parameter models, we proposed a data-dependent bandwidth selection procedure that adapts to unknown smoothness of the time-varying parameter. Monte Carlo experiments and an application to the time-varying CAPM suggest that the proposed method serves as a unified approach to estimating a variety of time-varying parameter models.

table[table omitted — 2,792 chars of source]
table[table omitted — 2,807 chars of source]
table[table omitted — 2,949 chars of source]
table[table omitted — 2,942 chars of source]
table[table omitted — 2,150 chars of source]
table[table omitted — 1,147 chars of source]
figure[figure omitted — 419 chars of source]
figure[figure omitted — 612 chars of source]
figure[figure omitted — 828 chars of source]

\titleformat*{\section}