Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
96,364 characters · 22 sections · 66 citation commands
Estimating Time-Varying Parameters of Various Smoothness in Linear Models via Kernel Regression
\onehalfspacing
Parameter instabilities are widely observed in econometric analysis. One of the most common specifications for parameter changes is the following linear model:
where $T$ is the sample size, $p\times 1$ vector $x_t$ is the regressor, $p\times 1$ triangular array $\beta_{T,t}$ is the time-varying coefficient, and $\varepsilon_t$ is the disturbance.
In the literature, time-varying parameters are often estimated via kernel regression where observations are weighted by some kernel function. Starting from robinsonNonparametricEstimationTimeVarying1989, a large literature develops kernel-based estimation and inference for time-varying coefficient models; e.g., caiTrendingTimevaryingCoefficient2007, chenTestingSmoothStructural2012, zhangInferenceTimeVaryingRegression2012, inoueRollingWindowSelection2017, and friedrichSieveBootstrapInference2024. We follow this strand of literature and study kernel-based estimation of $\beta_{T,t}$.
Our contributions are threefold. First, we consider a broader class of time-varying parameters than is typically assumed. The most common assumption adopted in existing works is that $\beta_{T,t}$ is so smooth that it is continuously differentiable caiTrendingTimevaryingCoefficient2007,zhangInferenceTimeVaryingRegression2012,inoueRollingWindowSelection2017. However, smooth functions are not the only model for parameter instability popular in economics and statistics. The random walk model, in which $\beta_{T,t}$ is modeled as $\beta_{T,t} = \sum_{i=1}^{t}u_i$ with $u_i$ a transitory process, is a popular alternative nyblomTestingConstancyParameters1989,stockMedianUnbiasedEstimation1998,cogleyDriftsVolatilitiesMonetary2005. Another example is (abrupt) structural breaks in $\beta_{T,t}$ andrewsTestsParameterInstability1993,baiEstimatingTestingLinear1998. These two modeling schemes have received less attention in the literature on kernel-based estimation, and it is largely unknown what the consequence is if one applies kernel regression to these models.\footnote{giraitisInferenceStochasticTimevarying2014 and giraitisTimevaryingInstrumentalVariable2021 are among few exceptions. They show that random-walk type parameters can be estimated via a kernel-based method.}$^{,}$\footnote{pesaranSelectionEstimationWindow2007, pesaranOptimalForecastsPresence2013 and hiranoAnalyzingCrossvalidationForecasting2022 apply kernel-based approaches for random walk and structural break type parameter instabilities, but their focus is on optimal forecasting of $y_t$, rather than estimation of $\beta_{T,t}$.} We develop kernel-based estimation theory that covers a wide class of time-varying parameters, including smooth functions, the rescaled random walk, and structural breaks. This class also includes the threshold regression model of hansenSampleSplittingThreshold2000, which has rarely been considered in the context of kernel regression.
Let us emphasize that the class of time-varying parameters considered in this article also includes the mixtures of the aforementioned models. The relationship between $y_t$ and $x_t$, for example, may evolve smoothly but exhibit discontinuities during global financial crises or pandemics. The literature has acknowledged the importance of taking into account several types of parameter instability. For instance, mullerEfficientEstimationParameter2010 consider inference in models with time-varying parameters approximated by Gaussian processes and continuous functions possibly with finitely many jumps. Although their framework allows for nonlinear models and thus is more general than ours in this respect, they focus on small parameter instabilities (relative to those considered in this work). Therefore, large instabilities are not allowed in their model. chenTestingSmoothStructural2012 develop tests for smooth parameter changes with finitely many breaks but do not develop estimation theory for time-varying parameters of this type. kristensenNonparametricDetectionEstimation2012 proposes a nonparametric estimation method for time-varying coefficients by developing a framework that he argues allows for smooth functions, structural breaks, and the rescaled random walk. However, his analysis is restricted to smooth functional parameters only and cannot be extended to the other specifications. giraitisTimevaryingInstrumentalVariable2021 allow smooth deterministic functions, the rescaled random walk, and their mixture in an IV setting, but exclude (large) discontinuous breaks. Unlike these earlier works, we employ a general framework that accommodates all the aforementioned models and their mixtures.
We develop this framework by considering the class of time-varying parameters characterized by smoothness parameter $\alpha>0$, which generalizes the H\"{o}lder-class functions studied in the literature on nonparametric estimation tsybakovIntroductionNonparametricEstimation2009. For example, continuously differentiable functions have smoothness of $\alpha=1$ as in the usual H\"{o}lder condition, and the random walk divided by $\sqrt{T}$ has $\alpha=1/2$. As with the H\"{o}lder condition, a smaller $\alpha$ means more roughness of the path of the time-varying parameter. This implies that the random walk divided by $\sqrt{T}$ is less smooth than continuously differentiable functions.
Our second contribution is to discuss the role of the bandwidth and its implications on bandwidth selection in the above general setting; beyond the usual bias-variance trade-off inherent in nonparametric estimation, we demonstrate that the bandwidth determines a trade-off between the rate of convergence and the size of the class of time-varying parameters that can be estimated. While this fact is consistent with and may be inferred from earlier results on nonparametric estimation of H\"{o}lder-class functions, its implications on bandwidth selection seem underappreciated in the literature on time-varying parameter models.
Specifically, we show that the rate of the bandwidth should be determined according to the smoothness of $\beta_{T,t}$. We illustrate this argument through two examples. First, we show that the conventional $T^{-1/5}$-rate bandwidth, which is specialized to continuously differentiable functions and often used in the literature zhouSimultaneousInferenceLinear2010, is invalid if $\beta_{T,t}$ is less smooth. For example, we show that the bandwidth should be proportional to $T^{-1/2}$ if $\beta_{T,t}$ is the random walk divided by $\sqrt{T}$. Second, we demonstrate that, if the time-varying parameter experiences both smooth and abrupt parameter changes, the abrupt breaks of certain magnitudes are absorbed in smooth parameter changes so that the kernel-based estimation delivers valid inference, while discontinuous changes of a larger magnitude cause bias. We show that the bandwidth determines the break magnitudes at which this bias arises.
Our third contribution is to propose a data-driven bandwidth selection procedure. Unlike existing approaches that focus on smooth time-varying parameters and $T^{-1/5}$-rate bandwidths, the proposed method allows researchers to select the bandwidth from a wide range of candidate values, adaptively to the latent smoothness of the time-varying parameter. We evaluate its finite-sample performance via Monte Carlo simulations and illustrate the method using the capital asset pricing model (CAPM). In this application, our selection algorithm does not support the conventional $T^{-1/5}$-rate bandwidth, casting doubt on the extent to which this routinely selected bandwidth and the commonly used assumption of (continuously) differentiable parameters are justified. Furthermore, the proposed procedure partially supports a $T^{-1/2}$-rate bandwidth, producing an estimated trajectory close to that obtained from a Bayesian random-walk estimation.
The remainder of this paper is organized as follows. Section (ref) defines smoothness of time-varying parameters. Section (ref) establishes asymptotic properties of the kernel-based estimator. Section (ref) discusses the consequence of an improper bandwidth choice and develops a bandwidth selection method adaptive to the smoothness of the time-varying parameter. Section (ref) conducts Monte Carlo experiments, and Section (ref) gives a real data analysis. Section (ref) concludes. Mathematical proofs of the main results are relegated to Appendix A.
Notation: For any matrix $A$, $\|A\|=\mathrm{tr}(A'A)^{1/2}$ denotes the Frobenius norm of $A$. For any positive number $b$, $\lfloor b \rfloor$ denotes the integer part of $b$. $\stackrel{p}{\to}$ and $\stackrel{d}{\to}$ signify convergence in probability and convergence in distribution as $T\to\infty$, respectively. $\Rightarrow$ signifies weak convergence of the associated probability measures.
We consider estimating $\beta_{T,t}$ by using the local constant (Nadaraya-Watson) estimator:
where $K(\cdot)$ is a kernel function and $h$ is the bandwidth parameter satisfying $h\to0$ and $Th\to \infty$ as $T\to \infty$. Assumptions on the data generating process and kernel $K$ will be detailed in Section (ref)
In discussing the asymptotic properties of $\hat{\beta}_t$, the smoothness of the path of $\beta_{T,t}$ has a decisive effect. In the following definition, we quantify the smoothness of $\beta_{T,t}$ by a single parameter $\alpha$.
Definition (ref) essentially controls by $\alpha$ the smoothness of the path of $\beta_{T,t}$ on any interval of any length of a smaller order than $T$. In typical applications, $a_T$ will be set $a_T = \lfloor Th \rfloor$. Definition (ref)(a) allows the difference between the values of $\beta_{T,t}$ at distinct time points to grow as the time points gets further apart, while Definition (ref)(b) does not.\footnote{Therefore, $\beta_{T,t}$ belongs to type-a $\mathrm{TVP}(\alpha)$ if it belongs to type-b $\mathrm{TVP}(\alpha)$.}
Determining $\alpha$ such that a given $\beta_{T,t}$ belongs to $\mathrm{TVP}(\alpha)$ on the boundary enables us to derive the largest possible bandwidth under which $\hat{\beta}_t - \beta_{T,t}$ is asymptotically normally distributed; see Theorem (ref) below for this point.
Because $a_T/T<1$ and $\alpha>0$, a smaller $\alpha$ permits larger differences $\|\beta_{T,t}-\beta_{T,j}\|$, resulting in $\beta_{T,t}$ possibly having a rougher path. Note that triangular arrays unbounded in probability are excluded from Definition (ref). We emphasize that Definition (ref) does not impose any parametric assumption on $\beta_{T,t}$ (other than smoothness $\alpha$), and that $\beta_{T,t}$ may be deterministic or stochastic. In addition, $\beta_{T,t}$ is allowed to have arbitrary correlation with $x_t$ and $\varepsilon_{t}$. Definition (ref) is quite general and accommodates many important time-varying parameters, as shown below.
Mixture models where $\beta_{T,t}$ is the sum of both type-a and type-b time-varying parameters will be considered in Section (ref).
We suppose kernel $K(\cdot)$ satisfies the following condition:
Commonly used kernels such as the uniform density on $[-1,1]$ and the Epanechnikov kernel satisfy Assumption (ref). Following the arguments of giraitisInferenceStochasticTimevarying2014,giraitisTimevaryingInstrumentalVariable2021, kernels with non-compact support such as the Gaussian kernel are permitted under some stronger condition. We focus on kernels with a compact support as specified in condition (a) to avoid unessential complications. Note that $2Th$ is the effective sample size of the kernel-based estimation, since $K((t-i)/Th) = 0$ for $i$ such that $|t-i|> Th$.
Next, we impose the following assumption on model (ref).\footnote{For the definition of near epoch dependence (NED), see, e.g., davidsonStochasticLimitTheory1994.}
Assumption (ref)(a) allows the regressor and disturbance to be weakly serially dependent. The NED assumption is more general than mixing conditions commonly assumed in the literature caiTrendingTimevaryingCoefficient2007,chenTestingSmoothStructural2012,giraitisTimevaryingInstrumentalVariable2021,friedrichSieveBootstrapInference2024. Also note that we do not impose strict or covariance stationarity unlike earlier works caiTrendingTimevaryingCoefficient2007,chenTestingSmoothStructural2012,friedrichSieveBootstrapInference2024, and thus our framework allows for heteroskedasticity in $\varepsilon_t$. Assumption (ref)(b) requires that the product of regressors and disturbance be serially uncorrelated, which is satisfied when, for example, $\varepsilon_t$ is a martingale difference sequence (m.d.s.) with respect to $\mathcal{F}_{T,t} \coloneqq \sigma(\{x_{t+1},x_t,\varepsilon_t,x_{t-1},\varepsilon_{t-1},\ldots\})$. The assumption of no serial correlation or m.d.s. is common in the literature chenTestingSmoothStructural2012,kristensenNonparametricDetectionEstimation2012,giraitisTimevaryingInstrumentalVariable2021. Assumption (ref)(c) holds under Assumptions (ref)(a)-(b) if $x_t$ and $x_t\varepsilon_t$ are covariance-stationary (see Corollary (ref)).
In the following theorem, we establish the consistency and asymptotic normality of the kernel-based estimator, $\hat{\beta}_t$.
In what follows, we will set $h=cT^{\gamma}$ and call $\gamma$ (as well as $h$) the bandwidth parameter.
If $\beta_{T,t}$ belongs to type-a $\mathrm{TVP}(\alpha)$ on the boundary, then it can be estimated by setting $\gamma \approx -1/(2\alpha+1)$, yielding a convergence rate $\approx T^{-\alpha/(2\alpha+1)}$. The same convergence rate has been established in the literature for the minimax risk of kernel-based estimators, assuming that the parameter of interest belongs to a H\"{o}lder class with exponent $\alpha$ tsybakovIntroductionNonparametricEstimation2009. Theorem (ref) shows that an analogous result holds under our definition of smoothness, Definition (ref)(a).
Smoothness parameter $\alpha$ affects $\Gamma(\alpha)$, the set of bandwidth parameter $\gamma$ that yields $\sqrt{Th}$-consistency and asymptotic normality. This makes the rate of convergence $T^{(1+\gamma)/2}$ dependent on $\alpha$. Letting $\alpha \to 0$, the kernel-based estimation can accommodate time-varying parameters of arbitrary smoothness, but this is accompanied by $\Gamma(\alpha) \to -1$, resulting in the rate of convergence $T^{(1+\gamma)/2} \to 1$. In contrast, if we let $\alpha\to\infty$, then $\Gamma(\alpha)$ tends to $(-1,0)$, and the choice $\gamma \approx 0$ yields a nearly $\sqrt{T}$-rate convergence, but only highly smooth parameters can be estimated. This observation shows that there is a trade-off between the rate of convergence and the size of the class of the time-varying parameters that can be estimated.
Because $\Gamma(\alpha)$ is the set of $\gamma$ that yields $\sqrt{Th}$-consistency and asymptotic normality under given $\alpha$, we can obtain the set of $\alpha$ that leads to $\sqrt{Th}$-consistency and asymptotic normality of $\hat{\beta}_{t}$ under given $\gamma$, by inverting the expression of $\Gamma(\alpha)$. Letting $A(\gamma)$ denote such a set, we can say that $\hat{\beta}_t$ calculated using given $\gamma$ is $\sqrt{Th}$-consistent and asymptotically normal for time-varying parameters with smoothness $\alpha\in A(\gamma)$, where
Letting $\gamma \to -1$, $A(\gamma)$ tends to $(0,\infty)$, which implies that time-varying parameters with any smoothness $\alpha>0$ can be estimated, but the rate of convergence becomes $T^{(1+\gamma)/2} \to 1$. On the other hand, if we let $\gamma\uparrow0$, then the rate of convergence is as fast as $\sqrt{T}$, but $A(\gamma) \to \infty$ (the smoothness of constant parameters) in the type-a case. Hence, the bandwidth determines the trade-off between efficiency and robustness.
\setcounter{exm}{0}
To conduct inference, one needs to consistently estimate the asymptotic variance of $\hat{\beta}_t$. Under Assumptions (ref) and (ref), $\Omega(r)$ can be consistently estimated by
see Lemma (ref) in Appendix A. A natural estimator of $\Sigma(r)$ is
where $\hat{\varepsilon}_i = y_i - x_i'\hat{\beta}_i$.\footnote{If $\{x_t\varepsilon_t\}$ is serially correlated, $\Sigma(r)$ is typically the long-run variance of $\{x_t\varepsilon_t\}$. In this case, an appropriate estimator of $\Sigma(r)$ would be a nonparametric kernel estimator such as the Newey-West one, as suggested by caiTrendingTimevaryingCoefficient2007. We do not explore in this direction to save space.} To prove the consistency of $\hat{\Sigma}(r)$, however, the current assumptions are not sufficient. This is because for each $t=\lfloor Tr \rfloor$, the estimation errors $\hat{\beta}_{t+j}-\beta_{T,t+j}$, where $j\in[-\lfloor Th \rfloor, \lfloor Th \rfloor]$, are required to be asymptotically negligible uniformly over $j\in[-\lfloor Th \rfloor, \lfloor Th \rfloor]$, on which $\hat{\Sigma}(r)$ is calculated. To ensure the uniform consistency of $\hat{\beta}_{t+j}$ over $j\in[-\lfloor Th \rfloor, \lfloor Th \rfloor]$, we need the following additional conditions.
Assumption (ref)(a') strengthens Assumption (ref)(a) by increasing the decaying rates of the mixing and NED coefficients, essentially weakening the serial dependence of $\{(x_t', \varepsilon_t)\}_t$.
Assumption (ref) requires that there be enough variation in the data, as it implies that the minimum eigenvalue of $E[x_tx_t']$ is bounded away from zero uniformly in $t$.
Assumption (ref) is a high-level one that ensures the uniform consistency of $\hat{\beta}_{t+j}$ over $j\in[-\lfloor Th \rfloor, \lfloor Th \rfloor]$. In Appendix C, we show that Assumption (ref) is satisfied in the time-varying models and under the implied bandwidths discussed in Examples (ref)-(ref).
In Theorem (ref), we showed the set of admissible bandwidth rates depends on the smoothness $\alpha$ of $\beta_{T,t}$. This implies that an improperly selected bandwidth rate (given by $\gamma \notin \Gamma(\alpha)$) leads to misleading inference. In this section, we illustrate this implication through some examples where the evolutionary mechanism of $\beta_{T,t}$ is misspecified. We also discuss how to choose the bandwidth in empirical studies.
Suppose one assumes $\beta_{T,t}$ is a continuously differentiable function and sets $\gamma \approx -1/3$, but the fact is that $\beta_{T,t}$ follows the random walk divided by $\sqrt{T}$. Using the results given in Theorem (ref), it is readily shown that the kernel-based estimator satisfies $\sqrt{cT^{1+\gamma}}(\hat{\beta}_t - \beta_{T,t}) = S_{T,t} + O_p(T^{1/2+\gamma})$, where $S_{T,t} \stackrel{d}{\to}N(0,\Omega(r)^{-1}\Sigma(r)\Omega(r)^{-1})$. Since the bias term is of order $O_p(T^{1/2+\gamma})$ and $\gamma > -1/2$, the difference $\hat{\beta}_{t} - \beta_{T,t}$ is dominated by the bias term. Because the bias term is not normal in general, confidence intervals based on a normal approximation will perform poorly.
In the literature on smooth (differentiable) time-varying parameters, researchers often use a rule-of-thumb or plug-in bandwidth $h=\mathrm{constant}\times T^{-1/5}$, or pick the bandwidth minimizing the cross-validation criterion over $h\in [c_1T^{-1/5},c_2T^{-1/5}]$ for some $0<c_1<c_2$ zhouSimultaneousInferenceLinear2010,zhangInferenceTimeVaryingRegression2012,kristensenNonparametricDetectionEstimation2012,chengBayesianBandwidthEstimation2019,sunPenalizedTimevaryingModel2023. Although these selection rules lead to an efficient estimation of $\beta_{T,t}$ as long as it is correctly specified as a continuously differentiable function, they will yield a biased estimation if $\beta_{T,t}$ is a random walk, or more generally, if $\beta_{T,t}$ does not belong to type-a $\mathrm{TVP}(1)$.
Suppose $\beta_{T,t}=\mu_{T,t} + (1/\sqrt{T})\sum_{i=1}^{t}u_i$, where $u_t$ is defined as in Example (ref), and $\mu_{T,t}$ satisfies
with $T_B = \lfloor \tau_BT \rfloor$ and $\mu_2-\mu_1 = \delta/T^{\alpha}$. Then, we can show
where $R_{T,t}$ is defined in (ref) and (ref). The asymptotic order of the bias term, $R_{T,t}$, is $O_p(T^{\gamma/2})$ for $t$ outside the $\lfloor Th \rfloor$-neighborhood of break point $T_B$. On the $\lfloor Th \rfloor$-neighborhood of $T_B$, it is $O_p(T^{\gamma/2})$ if $\alpha\geq-\gamma/2$, while it is $O_p(T^{-\alpha})$ if $0<\alpha<-\gamma/2$.
Suppose we estimate $\beta_{T,t}$ by $\hat{\beta}_t$ assuming $\mu_{T,t} = \mu$, that is, the parameter instability is purely due to the zero-mean random walk. In this case, the (misleading) optimal rate of convergence is achieved by the choice of $\gamma=-1/2$, yielding $R_{T,t}=O_p(T^{-1/4})$ for $t\in[1,T_B-\lfloor Th \rfloor]\cup[T_B+1+\lfloor Th \rfloor,T]$, and
for $t\in[T_B-\lfloor Th \rfloor+1, T_B+\lfloor Th \rfloor]$.
When $\alpha\geq1/4$, the asymptotic order of $R_{T,t}$ is $O_p(T^{-1/4})$ for all $t$, the same order as in the pure random walk case (see (ref)), so that the choice $\gamma =-1/2$ is valid and leads to the fastest rate of convergence. In contrast, if $0<\alpha<1/4$, $R_{T,t} = O_p(T^{-\alpha})$ for $t=T_B \pm \lfloor rTh \rfloor, \ r\in [0,1]$. Because the asymptotically normal component of the decomposition of $\hat{\beta}_t - \beta_{T,t}$ is $O_p(T^{-(1+\gamma)/2}) = O_p(T^{-1/4})$, the asymptotic behavior of $\hat{\beta}_t - \beta_{T,t}$ is dominated by the bias term. Therefore, the structural break induces a severe bias in the kernel-based estimation on the $\lfloor Th \rfloor$-neighborhood of the discontinuity point $t=T_B$.
The above result tells us that, if we set $\gamma=-1/2$, abrupt breaks of size $1/T^{\alpha}$ are absorbed in random walk parameter instabilities if $\alpha\geq1/4$, while the abrupt breaks “stick out" and cause bias when $\alpha<1/4$.
A more general result can be derived if we invoke $\Gamma(\alpha)$ and $A(\gamma)$ defined in (ref) and (ref), respectively. Suppose that $\beta_{T,t}$ can be expressed as $\beta_{T,t}^1 + \mu_{T,t}$, where $\beta_{T,t}^1$ belongs to type-a $\mathrm{TVP}(\alpha_1)$ with $\alpha_1 \in (0,\infty)$, and $\mu_{T,t}$ is defined as in (ref) but the magnitude of the break is $\mu_2-\mu_1 = \delta/T^{\alpha_2}$. Then, $\hat{\beta}_t$ is $\sqrt{Th}$-consistent and asymptotically normal under any $\gamma \in \Gamma(\alpha_1)=(-1,-1/(2\alpha_1+1))$ for $t\in[1,T_B-\lfloor Th \rfloor]\cup[T_B+1+\lfloor Th \rfloor,T]$. On $[T_B-\lfloor Th \rfloor+1, T_B+\lfloor Th \rfloor]$, the abrupt break is absorbed in $\beta_{T,t}^1$, and the same $\gamma$ leads to $\sqrt{Th}$-consistency and asymptotic normality if $\alpha_2\in A(\gamma) = ((1+\gamma)/2, \infty)$, while the bias term dominates the asymptotically normal term if $\alpha_2<(1+\gamma)/2$.
In the previous subsections, we have observed that an improperly selected $\gamma$ leads to misleading inference. Therefore, care must be taken in determining the bandwidth parameter. Because bandwidth parameter $h$ takes the form of $h=cT^{\gamma}$ with $c>0$ and $\gamma<0$, we first discuss how to determine $\gamma$ and then how to select $c$.
If one can identify the evolutionary mechanism of $\beta_{T,t}$ based on some prior information, they may select $\gamma$ appropriately, referring to the theoretical results derived in Section (ref). For instance, if the random walk coefficient model is plausible, $\gamma = -1/2$ is an appealing choice. If $\beta_{T,t}$ is known to be twice continuously differentiable, then various methods for bandwidth selection proposed in the literature can be used to determine $\gamma$ and $c$ jointly (e.g., zhangInferenceTimeVaryingRegression2012).
Determining $\gamma$ is more delicate when there is no prior information that helps identify the evolutionary mechanism of $\beta_{T,t}$. Here, we propose two data-driven procedures to select the value of $\gamma$, which require prespecified lower and upper bounds $\underline{\gamma}, \overline{\gamma}\in(-1,0)$ to construct the set of candidate $\gamma$ values, $[\underline{\gamma},\overline{\gamma}]$. The theory developed in this article will serve as a guiding principle in determining $\underline{\gamma}$ and $\overline{\gamma}$.
The first procedure we propose is a naive cross-validation-based method: For each $\gamma\in[\underline{\gamma},\overline{\gamma}]$ (on some grid), calculate $\hat{\beta}_{-t,m}(\gamma)$, where $\hat{\beta}_{-t,m}(\gamma)$ are the leave-$(2m+1)$-out local constant estimators calculated without the data on $s\in[t-m,t+m]$ and with $h=T^{\gamma}$ for some $m\in\mathbb{N}\cup\{0\}$, compute the cross-validation criterion, $\mathrm{CV}(\gamma)\coloneqq T^{-1}\sum_{t=1}^T(y_t-x_t'\hat{\beta}_{-t,m}(\gamma))^2$, and then pick the minimizer of $\mathrm{CV}(\gamma)$. One may use the generalized cross-validation (GCV) considered in zhouSimultaneousInferenceLinear2010,zhangInferenceTimeVaryingRegression2012.
The second procedure is based on fixed-design wild bootstrap goncalvesBootstrappingAutoregressionsConditional2004.
The rationale behind Algorithm (ref) is as follows. If $\gamma_1$ is sufficiently small that $\hat{\beta}_t(\gamma_1)$ is $\sqrt{Th_1}$-consistent for $\beta_{T,t}$, then $\hat{\varepsilon}_t(\gamma_1)$ are good approximations of unobserved $\varepsilon_t$, and bootstrap sample $y_t^*(\gamma_1)$ generated from $\varepsilon_t^*(\gamma_1)$ is “well-behaved". Treating $\hat{\beta}_t(\gamma_1)$ as the pseudo-true parameters, $\hat{\beta}_t^*(\gamma_1,\gamma_2)$ ($\gamma_2\leq\gamma_1$) are $\sqrt{Th_2}$-consistent for $\hat{\beta}_t(\gamma_1)$ and asymptotically normal under the bootstrap probability measure, in probability. Then, the confidence interval for $\hat{\beta}_t(\gamma_1)$ based on $\hat{\beta}_t^*(\gamma_1,\gamma_2)$ and $N(0,1)$ should attain empirical coverage rates close to the nominal confidence level, with high probability. In Step 4, we pick the largest $\gamma_1$ such that the above argument applies, so that the fastest possible convergence rate can be achieved.
To theoretically justify the above reasoning, we impose the following regularity condition, strengthening Assumption (ref):
If $\beta_{T,t}$ satisfies Condition (ref) given in Appendix C, then Assumption (ref) holds when $2\alpha\gamma_1+\gamma_2<-1$. If $\beta_{T,t}$ is a rescaled random walk, under Condition (ref) given in Appendix C, Lemma 5(iii) of giraitisTimevaryingInstrumentalVariable2021 shows that Assumption (ref) holds when $\gamma_1+\gamma_2<-1$.
Once $\gamma$ is determined, one can select $c$ by minimizing some criterion that is a function of $c$. For instance, the cross-validation criterion, $\mathrm{CV}(c) \coloneqq T^{-1}\sum_{t=1}^T(y_t - x_t'\hat{\beta}_{-t}(c,\hat{\gamma}))^2$, where $\hat{\beta}_{-t}(c,\hat{\gamma})$ are the leave-one-out kernel estimators calculated under $h=cT^{\hat{\gamma}}$, can be used. Some other criteria may be used to determine $c$ such as the AIC as suggested in caiTrendingTimevaryingCoefficient2007 or the GCV considered in zhouSimultaneousInferenceLinear2010 and zhangInferenceTimeVaryingRegression2012.
In this section, we conduct three Monte Carlo experiments to verify the implications provided in Section (ref). We use the following DGP: $y_t = \beta_{T,t}x_t + \varepsilon_{t}, \ t=1,\ldots,T$, where $x_t = 0.5x_{t-1} + \varepsilon_{x,t}$ with $\varepsilon_{x,t} \sim \mathrm{i.i.d.} \ N(0,1)$, and $\beta_{T,t}$ is defined differently in different experiments. For the specification of $\varepsilon_{t}$, we consider two cases: $\varepsilon_{t} = u_t$, where $u_t \sim \mathrm{i.i.d.} \ N(0,1)$ (i.i.d. case) and $\varepsilon_{t} = \sigma_tu_t$ with $\sigma_t^2 = 0.1 + 0.3\varepsilon_{y,t-1}^2 + 0.6\sigma_{t-1}^2$ (GARCH case). To obtain $\hat{\beta}_t$, we use the Epanechnikov kernel $K(x) = 0.75(1-x^2)I(|x|\leq 1)$.
The first experiment is related to Section (ref), and $\beta_{T,t}$ is generated as the rescaled random walk: $\beta_{T,t} = T^{-1/2}\sum_{i=1}^{t}v_i$. We consider two DGPs for driver process $v_t$: (i)$v_t\sim \mathrm{i.i.d.} \ N(0,1)$ and (ii) $v_t \sim \mathrm{i.i.d.} \ \mathrm{log \ normal}$ with parameters $\mu=0$ and $ \sigma = 1$.\footnote{Specifically, $X$ follows a log normal distribution if $X=\exp(Z)$, where $Z \sim N(\mu,\sigma^2)$.} Four sample sizes are used: $T \in \{100,200,400,800\}$. To evaluate the global performance of $\hat{\beta}_{t}$, we calculate $\mathrm{MSE} = T^{-1}\sum_{t=1}^{T}(\hat{\beta}_t - \beta_{T,t})^2$ (the reported MSE is the mean MSE over 2000 replications). To evaluate the normal approximation given in Corollary (ref), we construct the $95\%$ confidence interval for $\beta_{T,0.5T}$ (the middle point of the sample). The variance estimators are $\hat{\Omega} \coloneqq T^{-1}\sum_{i=1}^{T}x_i^2$ and $\hat{\Sigma} \coloneqq \int_{-1}^{1}K(x)^2dx\times T^{-1}\sum_{i=1}^{T}\hat{\varepsilon}_i^2x_i^2$. We experiment with bandwidth parameter $h=T^{\gamma}$ and $\gamma \in \{-0.2,-0.33,-0.5,-0.55,-0.6,-0.7\}$, and evaluate the performance for each pair $(\gamma,T)$. We also analyze the performance of the data-driven selection procedures for $\gamma$ suggested in Section (ref).\footnote{For the cross-validation-based method, we select $\hat{\gamma}$ from $[-0.5,-0.2]$ using leave-three-out estimators. For the bootstrap-based method, we select $\hat{\gamma}$ from $\{-0.5,-0.4,-0.33,-0.2\}$ and set the tolerance level $\bar{q}=0.1$.} The results are presented in Tables (ref) and (ref).
Because $\beta_{T,t}$ is the random walk divided by $\sqrt{T}$, our theoretical results predict that an appropriate bandwidth is $\gamma \approx -1/2$, while the kernel-based estimator leads to poor inference when $\gamma > -1/2$. Our simulation result corroborates this analysis. First, consider the case where $\varepsilon_{t} \sim \mathrm{i.i.d.} \ N(0,1)$ (Table (ref)). In case (i) (Gaussian random-walk $\beta_{T,t}$), when $\gamma = -0.2$, the coverage rate is far below the 95% confidence level. What is worse, it deviates from 0.95 as $T$ increases. Note that the MSE is relatively large. When $\gamma=-1/3$, the MSE takes the smallest value for all $T$ considered, but the coverage rate is still too small. This result warns researchers against using these bandwidths unless they are confident that $\beta_{T,t}$ can be well approximated by smooth functions with smoothness parameter $\alpha=1$. For $\gamma\leq -1/2$, the interval estimation performs well with coverage rate being 85-90% and getting better as $T$ increases. However, $\gamma = -0.7$ leads to undercoverage when $T$ is small and the largest MSE for all $T$. $\gamma = -0.6$ also gives large MSEs. The choices $\gamma \approx -1/2$ lead to good coverage and small MSE, so that these choices are recommended for random-walk type parameters, or more generally, for time-varying parameters belonging to $\mathrm{TVP}(1/2)$ on the boundary.
Next consider the performance of $\gamma=\hat{\gamma}$ selected by data-dependent procedures. For the cross-validation method, the mean MSE is close to the smallest MSE attained by the deterministic choice of $\gamma=-0.33$, particularly when $T$ is large. On the other hand, the coverage ratio is far below the nominal level and takes values between 73% and 78%, although the coverage gradually improves as the sample size increases. For the bootstrap-based selection, the mean MSE and the coverage ratio is almost identical to those attained by the deterministic choice of $\gamma=-0.5$; that is, the MSE is relatively large when $T$ is small, but improves quickly as $T$ increases, and the coverage ratio is reasonably good. Based on these observations, the cross-validation method seems useful when a small MSE is desired, while the bootstrap-based method is a reasonable choice when unbiased estimation is prioritized.
The result for case (ii) (non-Gaussian random-walk $\beta_{T,t}$) is similar, so the same comment applies.
Results for the case where $\varepsilon_{t}$ is GARCH (Table (ref)) are similar to those for the i.i.d case. Hence, we do not repeat the same analysis.
The second experiment is for verifying the implication provided in Section (ref). In this simulation, we analyze the effect of (neglected) structural breaks. For this purpose, we generate $\beta_{T,t}$ according to $\beta_{T,t} = \mu_{T,t} + T^{-1/2}\sum_{i=1}^{t}v_i$, where $v_i \sim \mathrm{i.i.d.} \ N(0,1)$ and $\mu_{T,t}$ is an intercept term experiencing a break at $t=0.5T$. Specifically, we let
where $\alpha \in \{0.1,0.2,0.3,0.4\}$. A smaller $\alpha$ yields a larger break.
We consider estimating $\beta_{T,t}$ with the choice $h = T^{-1/2}$, reflecting the ignorance of the break. According to our theoretical analysis, the kernel-based estimator has a severe bias around $t=0.5T$ when $\alpha<0.25$, while breaks given by $\alpha>0.25$ have no effect asymptotically. To confirm this implication, we calculate the MSE and coverage rate of $\hat{\beta}_{t}$ for $t=\tau T$ with $\tau = 0.4,0.45,0.5,0.55,0.6$. The MSE is calculated for each $\tau$ as the mean squared error over 2000 replications, that is, $\mathrm{MSE}(\tau) = 2000^{-1}\sum_{i=1}^{2000}(\hat{\beta}_{\tau T}^{(i)} - \beta_{T,\tau T}^{(i)})^2$, where superscript $i$ signifies $\hat{\beta}_{\tau T}^{(i)}$ and $\beta_{T,\tau T}^{(i)}$ are obtained in the $i$th replication. We consider four sample sizes; (i) $T = 100$, (ii) $T=200$, (iii) $T=400$, and (iv) $T=800$. We use $\hat{\Omega}_t = (Th)^{-1}\sum_{i=1}^{T}K((t-i)/Th)x_i^2$ and $\hat{\Sigma}_t = (Th)^{-1}\sum_{i=1}^{T}K((t-i)/Th)^2\hat{\varepsilon}_i^2x_i^2$ as the variance estimators to evaluate the normal approximation given in Theorem (ref). Results are reported in Tables (ref) and (ref).
First, let us see the case of $\varepsilon_{t}$ being i.i.d. and $T=100$ (Table (ref), the row labeled (i)). The MSEs and coverage rates for $\tau = 0.4$ and $0.6$ are stable across $\alpha$. This is because the break only affects estimation around the discontinuity point, $t=0.5T$. The break has a severe effect on $\hat{\beta}_{\tau T}$ with $\tau=0.45,0.5,0.55$, both in terms of MSE and coverage. The smaller $\alpha$ is (i.e. the larger the break is), the worse the performance gets. Moreover, this effect is more profound for $\tau$ closer to $0.5$. In terms of the coverage rate, smaller breaks given by $\alpha \geq 0.25$ have a nonnegligible effect. This indicates that, although breaks of these magnitudes asymptotically have no impact, they do have nontrivial effects in finite samples.
For case (ii) ($T=200$), MSEs for $\tau=0.5$ and $\alpha<0.25$ are still large. Note that MSEs for $\tau = 0.45,0.55$ are comparable with those for $\tau = 0.4,0.6$. This is because the abrupt break affects $\hat{\beta}_{t}$ on the $Th$-neighborhood of the break date. Because $t=0.45T$ and $t=0.55T$ are outside the $Th$-neighborhood of $0.5T$, the performance of $\hat{\beta}_{\tau T}$ improves as $T$ increases for $\tau = 0.45,0.55$. $\hat{\beta}_{0.5T}$ also suffers from poor coverage for all $\alpha$. For the cases with $T=400,800$ (cases (iii) and (iv)), a similar comment applies. In particular, the coverage rates for $\alpha < 0.25$ and $\tau=0.5$ deteriorate as $T$ increases.
Examining the case with $\varepsilon_{t}$ being GARCH (see Table (ref)), the same conclusion is drawn, so the detail is omitted.
In this subsection, we investigate the finite-sample performance of the data-driven bandwidth selection procedures in an environment where $\beta_{T,t}$ evolves smoothly but experiences a jump at some point. Specifically, we specify $\beta_{T,t}$ as $\beta_{T,t} = \beta(t/T)$, where $\beta(x) = x + \mu_T(x)$ with $\mu_T(x) = 0$ for $x\leq 0.5$ and $\mu_T(x) = 1.5/T^{0.4}$ for $x>0.5$. $\beta_{T,t}$ evolves smoothly and deterministically over time but experiences a break at the middle point. Recalling the theoretical analysis in Section (ref), the break of size $T^{-0.4}$ can be accommodated as long as $\gamma < -0.2$. We analyze the finite-sample performance of $\hat{\beta}_{T,t}$ with $\gamma \in \{-0.2,-0.33,-0.5\}$ and $\gamma$ selected by the data-driven methods. We are interested in (i) whether the data-driven procedures can select $\gamma < -0.2$ (unbiasedness) and (ii) whether the selected $\gamma$ is close to $-0.2$ (efficiency). We study the mean MSE and the empirical coverage ratio at the break point, $t=0.5T$. The results are reported in Table (ref). Because the results for the cases with i.i.d. error and GARCH error are qualitatively similar, we comment on the i.i.d. case only.
We first consider the deterministic $\gamma$. Although $\gamma=-0.2$ yields the smallest MSE, this choice results in undercoverage at the break point, as expected. For both choices $\gamma=-0.33,-0.5$, the coverage rate improves as $T$ increases, but $\gamma=-0.33$ gives a much smaller MSE. Next consider the data-dependent procedures. For the cross-validation method, the mean MSEs are almost identical to those obtained under $\gamma=-0.33$, but the empirical coverage rate is well below the nominal rate. For the bootstrap method, the mean MSEs and coverage rates are almost identical to those obtained by $\gamma=-0.5$ for $T\in\{100,200,400\}$. When $T=800$, however, the bootstrap-based procedure improves the MSE by about 20% compared to $\gamma=-0.5$ while maintaining the same level of coverage rate.
In this section, we apply kernel regression to estimate the time-varying CAPM.\footnote{The R code used for the empirical application is available on the author's website (\url{https://sites.google.com/view/mikihito-nishi/home}).} Parameter instabilities are widely observed in the CAPM literature ghyselsStableFactorStructures1998,lewellenConditionalCAPMDoes2006,famaValuePremiumCAPM2006,angCAPMLongRun2007,angTestingConditionalFactor2012,guoTimeVaryingBetaValue2017. We consider estimating the following factor model:
where $R_{j,t}$ denotes the excess return of portfolio $j$ at time $t$, and $R_{M,t}$ is the market excess return. The coefficients alpha and beta are allowed to be time-varying.
In the CAPM literature, parameter instability is often modeled by letting parameters depend on observable instruments. But results drawn from this approach tend to be sensitive to the choice of instruments ghyselsStableFactorStructures1998. To overcome this problem, researchers have proposed time-varying parameter models that do not utilize exogenous information.
Some assume that parameters experience abrupt changes, and others model parameter instability via the (near) random walk or smooth functions of time. For example, famaValuePremiumCAPM2006 and lewellenConditionalCAPMDoes2006 split the sample assuming that parameter changes occur based on calendar time (e.g., monthly or yearly), and apply OLS within subsamples. However, estimates obtained in this fashion suffer from bias if the timing of structural breaks is misspecified. angCAPMLongRun2007 use a Bayesian approach assuming (near) random walk alpha and beta. liTestingConditionalFactor2011 and angTestingConditionalFactor2012 treat the parameters as deterministic continuously differentiable functions of time.
Given the fact that continuously differentiable functions and the random walk can be estimated under $\gamma=-1/5$ and $\gamma=-1/2$, respectively, we set the lower and upper bounds for $\gamma$ as $\underline{\gamma}=-0.5$ and $\overline{\gamma}=-0.2$.
All data are extracted from Kenneth French's website (\url{https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html}). Following liCombinedApproachInference2015, we form three portfolios denoted by G, V, and G-V, respectively, from the 25 size-B/M portfolios. G is the average of the five portfolios in the lowest B/M quintile, V is the average of the five portfolios in the highest B/M quintile, and V-G is simply their difference. All the data are monthly, spanning 1952:1-2019:12 ($T=816$).
We use the Epanechnikov kernel and $\hat{\Omega}_t = (Th)^{-1}\sum_{i=1}^{T}K((t-i)/Th)x_ix_i'$ and $\hat{\Sigma}_t = (Th)^{-1}\sum_{i=1}^{T}K((t-i)/Th)^2\hat{\varepsilon}_i^2x_ix_i'$ as the variance estimators, where $x_t = (1,R_{M,t})'$. To save space, we only discuss the result for portfolio V-G. The results for portfolios G and V are given in Appendix E.
We determine two tuning parameters for the bandwidth, $h=cT^{\gamma}$, as explained in Section (ref). For Algorithm (ref), we construct 95% bootstrap confidence intervals and set the tolerance level to be $\bar{q}=0.1$, giving the threshold of 90% empirical coverage rate.
First, we consider selecting $\gamma$. Figure (ref) depicts the CV criterion computed using leave-($2m+1$)-out estimators for $m=0,1,2$. For $m=0,1$, the minimum is attained at $\gamma=-0.5$, whereas $\gamma=-0.32$ is the minimizer when $m=2$. Since there is little reason to prefer some specific value of $m$ to other values, we also use Algorithm (ref) to seek further evidence. Reported in Table (ref) are the mean empirical coverage rates taken over $t=1,\ldots,T$, $\overline{\mathrm{CR}}(\gamma_1,\gamma_2)\coloneqq T^{-1}\sum_{t=1}^T\mathrm{CR}_t(\gamma_1,\gamma_2)$. Each empirical coverage rate is calculated using $200$ bootstrap samples. For $\gamma_1=-0.33$ and $\gamma_1=-0.4$, the empirical coverage rates exceed the threshold of 0.9 for all $\gamma_2\leq\gamma_1$, and hence $\gamma_1=-0.33$ is supported by this procedure. Given these results, we set $\hat{\gamma}=-0.33$ since both CV- and bootstrap-based procedures support this choice. It is noteworthy that $\gamma=-1/5$, a prevalent choice in the literature, is rejected by our selection algorithm. This result highlights the importance of including other $\gamma$ values in the set of candidate bandwidths.
Given $\gamma=\hat{\gamma}$, we determine scaling constant $c$ via cross-validation. The selected value, $\hat{c}$, is the minimizer of the cross-validation criterion $\mathrm{CV}(c)$ over $c\in \{0.5,0.55,\ldots,1.5\}$.
In Figure (ref), we plot the estimated time-varying alpha and its 95% confidence band.\footnote{This confidence band is obtained by sequentially calculating the pointwise 95% confidence intervals and is not a uniform 95% confidence band.} The estimated alpha fluctuates around the value zero throughout the sample period, and the confidence band includes zero at all time points. Figure (ref) depicts the estimated time-varying beta. It starts with a positive value that is significantly different from zero and then fluctuates around zero up to $t=300$. Then, it starts to decrease and stays below zero with the confidence band excluding zero. It starts to increase from $t=600$, and fluctuates around zero from $t=660$ toward the end of the sample.
The CV-based selection procedure suggests that $\gamma=-1/2$ is partly supported by the data. Noting that this choice accommodates parameters following the (rescaled) random walk, and that random walk parameters are often estimated via Bayesian methods, it is interesting to compare the kernel-based estimates obtained from $h=\hat{c}T^{-1/2}$ with the estimates obtained from a Bayesian procedure in which parameters are assumed to be the random walk.
Let $\theta_{t} \coloneqq (\alpha_{t}, \beta_{t})'$. In the Bayesian method, we estimate the time-varying alpha and beta by using the Markov Chain Monte Carlo algorithm, assuming that $\theta_{t} = \theta_{t-1} + u_t$, where $u_t \sim N(0,D^2)$ with $D^2 = \mathrm{diag}(D_1^2,D_2^2)$.\footnote{For computation, we use the R package walker developed by helskeWalkerBayesianGeneralized2023.} As the prior distributions for parameters $\theta_{0}$, $D$ and $\mathrm{Var}(\varepsilon_t)=\sigma_{\varepsilon}^2$, we suppose $\theta_{0} \sim N(\mu\mathbf{1}_2,\sigma^2\mathbf{I}_2)$, $D_i \sim \mathrm{Gamma}(v_1,v_2), \ i=1,2$, and $\sigma_\varepsilon \sim \mathrm{Gamma}(\nu_1,\nu_2)$. We consider three configurations of hyperparameters. For each configuration, $(\mu,\sigma,v_2,\nu_1,\nu_2)$ are set to $(\mu,\sigma,v_2,\nu_1,\nu_2)= (0,32,10^{-4},2,10^{-4})$. The value of $v_1$ is varied, and we set $v_1 = 1, 2,$ and $4$.\footnote{We also changed the values for $(\mu,\sigma,v_2,\nu_1,\nu_2)$, but the estimates were insensitive to these parameters.}
In Figure (ref), we compare the Bayesian estimates with the kernel-based estimates with $h=\hat{c}T^{-1/2}$. For the estimated alpha (Figure (ref)), the trajectory obtained from the kernel method is more volatile (with a larger amplitude) than that obtained from the Bayesian algorithm, but the trajectories seem to share the same frequency. More striking is the similarity between the estimates of the time-varying beta. The estimated trajectories obtained from the two distinct methods are almost indistinguishable throughout the sample period, irrespective of the value of $v_1$.
We also compare the quantitative performances of the kernel and Bayesian estimators in terms of the in-sample fit (SSR). Standardizing the SSR obtained from the kernel-based method to be 1, the relative SSR's for the Bayesian estimators with $v_1=1,2,$ and $4$ are 1.019, 1.000, and 0.953, respectively. The kernel estimator yields a comparable in-sample fit relative to the Bayesian estimator.
We studied kernel-based estimation of time-varying parameters over a wide range of smoothness. We set up a general framework that quantifies the smoothness of the time-varying parameter by a single parameter $\alpha$, and established consistency and asymptotic normality of the kernel-based estimator under this framework. The results cover many important time-varying parameter models, including continuously differentiable functions, the rescaled random walk, abrupt structural breaks, the threshold regression model, and their mixtures.
Our analysis also highlights an often-overlooked role of the bandwidth and its implications on bandwidth selection. Beyond the bias-variance trade-off, when the parameter may be nondifferentiable, the bandwidth determines a trade-off between the rate of convergence and the size of the class of time-varying parameters that can be estimated. Theory and simulations show that the appropriate bandwidth rate depends on the smoothness of the time-varying parameter. In particular, a conventional $T^{-1/5}$-rate bandwidth is invalid in the case of nondifferentiable time-varying parameters such as the random walk. Another important implication from our result is that abrupt breaks of certain magnitudes cause bias in the kernel-based estimation.
Taking into account the diversity of existing time-varying parameter models, we proposed a data-dependent bandwidth selection procedure that adapts to unknown smoothness of the time-varying parameter. Monte Carlo experiments and an application to the time-varying CAPM suggest that the proposed method serves as a unified approach to estimating a variety of time-varying parameter models.
\titleformat*{\section}