Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
106,508 characters · 13 sections · 75 citation commands
Local polynomial estimation of time-varying parameters in nonlinear models
There is ample empirical evidence of time-varying parameters in many econometric models; see, e.g., Inoue2011, giacomini2016, Caldara2012, Ghysels1990 and Christoffersen2012. Most studies aiming at accommodating this feature assume a fully parametric model for this time variation; one example of this is structural break models. This has the advantage that the time-varying version of a given model stays parametric and can be estimated using existing methods. The disadvantage is that the researcher runs the risk of choosing a misspecified model for the time variation.
To reduce this risk, methods that treat the problem of time--varying parameters as a structured nonparametric one have been developed: They assume that the sequence of time-varying parameters arise as values of an underlying function which is then estimated nonparametrically; one popular class of estimators that falls in this category are local estimators which includes the so-called rolling-window estimator. However, the existing literature has mostly focused on local constant (Nadaraya-Watson) kernel estimators of the time-varying parameters; see, e.g., dahlhaus2017 and bardet2022.
We here propose to estimate the time--varying parameters using local polynomial estimators since these are known to have a number of attractive properties compared to the local constant one; see, e.g., fan1995JASA . Under very weak restrictions on the model being estimated and the time series data being used, we develop an asymptotic theory for the estimators. Specifically, we show that they are normally distributed in large samples and provide a complete charactersation of the leading variance and bias components.
The class of estimators includes as special cases the local constant estimator and the local linear estimator. We find that the local constant estimator requires stronger regularity conditions to be well-behaved in large samples and will generally suffer from additional biases in the interior of the domain compared to the local linear estimator. These additional biases involve the so-called derivative process of the stationary approximation to data which is not present in the biases of the local linear one. Moreover, the local linear estimator enjoys the well-known automatic boundary adjustment property: At the beginning and end of the sample it will perform better than the local constant one. This feature is important since often the main interest is on the values that the time-varying parameters take at the end of the sample.
The two most closely related papers to ours are dahlhaus2017 and bardet2022 who develop a general theory for local constant estimators of time--varying parameters in Markov models and infinite memory models, respectively. Our framework encompasses theirs as special cases and we consider a broader class of estimators than they do. Moreover, our proof techniques are different and, when specialising to local constant estimators, allow us to arrive at the same results as they do under weaker conditions: First, our theory imposes minimum requirements on the data--generating process with the main assumption being that it is local stationary. Second, it imposes weaker restrictions on the bandwidth used in the estimation; in particular, we can allow for the bandwidth being chosen using standard bandwidth selection rules, such as cross--validation, which is not the case in dahlhaus2017 and bardet2022. Second, we characterise the leading bias term of the estimators, which is in contrast to dahlhaus2017 and bardet2022. To demonstrate these attractive features of our general theory, we apply it to a class of Markov models with exogenous co--variates.
Another important feature of our theory is that it also applies to discrete-valued time series models, such as Poisson autoregressions. Such models are not covered by the theories of dahlhaus2017 and bardet2022 since these require the model of interest to be smooth; this condition is not satisfied when the time series is discrete--valued. In contrast, our theory for local linear estimators impose very weak smoothness conditions on the model due to our new proof techniques and so applies to discrete--valued time series models without any modifications. For the local constant estimator, we combine the ideas of Truquet2019, who analyse time--varying discrete-valued time series models, with our main result to obtain a complete analysis of this estimator.
We also contribute to the literature on asymptotic analysis of local polynomial estimators of varying-coefficient models by extending existing results fan1995JASA,loader2006 to cover situations where the objective functions are non-concave. This proves to be a non--trivial extension, but at the same time an important one since the log--likelihood functions of many non-linear models are non-concave.
As an empirical application, we revisit the empirical study of agosto2016 where a Poisson autoregressive model with additional co--variates was used to model and analyze US defaults. Using the proposed methodology, we find substantial time--variation in the model parameters that the original study was unable to capture. In particular, we find that the ability of macroeconomic and financial variables to predict defaults have varied substantially over time. A battery of informal tests of the time--varying model against the time--invariant version finds strong support for the former.
The remainder of the paper is organized as follows: Framework and estimators are introduced in Section (ref). Section (ref) presents the asymptotic theory of the estimators. Section (ref) provides examples of the theory when applied to particular models. We present the results of two simulation studies and the empirical application in Sections (ref) and (ref), respectively. All proofs have been relegated to the Appendix.
We are given $n$ observations, $Z_{n,t}\in \left( \mathcal{Z},\left\Vert \cdot \right\Vert \right) $, $t=1,\ldots ,n$, where $\left( \mathcal{Z} ,\left\Vert \cdot \right\Vert \right) $ is a Banach space, from a time-series model characterised by a finite--dimensional vector of unknown parameters $\theta \in \Theta \subset \mathbb{R}^{d_{\theta }}$ to be estimated. In most applications, $Z_{n,t}$ will also be finite--dimensional but our theory allows for, e.g., functional data as well. We take as given an objective function $\ell _{n,t}\left( \theta \right) =\ell \left( \mathcal{Z}_{n,t};\theta \right) \in \mathbb{R}$, where $\mathcal{Z} _{n,t}=\left( Z_{n,t},,....Z_{n,0},Z_{n,-1},....\right) $. Since we do not observe the process before $t=1$, we here initialise the process at deterministic values chosen by us, $Z_{n,-t}=z_{-t}$, $t\geq 1$. Under regularity conditions stated below, the effect of the initial values will vanish asymptotically.
The objective function is assumed to identify the data--generating parameter as its maximiser, $\theta =\arg \max_{\theta ^{\prime }\in \Theta }\mathbb{E} \left[ \ell _{n,t}\left( \theta ^{\prime }\right) \right] $, if $\theta $ indeed was time--invariant and $\ell _{n,t}\left( \theta ^{\prime }\right) $ was stationary and ergodic. In this case, the natural estimator is to replace population expectations by sample ones and estimate $\theta $ by $ \hat{\theta}=\arg \max_{\theta ^{\prime }\in \Theta }\frac{1}{n} \sum_{t=1}^{n}\ell _{n,t}\left( \theta ^{\prime }\right) $.
Suppose now that in fact $\theta $ is varying over time so that $\mathcal{Z} _{n,t}$ is generated by $\theta _{n,t}=\theta (t/n)$, $t=1,...,n$, where $ \theta :\left[ 0,1\right] \mapsto \Theta $ is an unknown function that characterizes the time--variation in the parameters.\footnote{ Data points now depends on sample size $n$ through $\theta \left( t/n\right) $, which is why we write $Z_{n,t}$ instead of of $Z_{t}$.} At the same time, the objective function is still assumed to identify the parameter in the sense that $\theta \left( t/n\right) =\arg \max_{\theta ^{\prime }\in \Theta }\mathbb{E}\left[ \ell _{n,t}\left( \theta ^{\prime }\right) \right] $. We then propose to estimate $\theta \left( u\right) $ at any given value $u\in \left[ 0,1\right] $ using local polynomial estimators: First, for $t/n$ in a neighbourhood of $u$, we approximate $\theta \left( t/n\right) $ by the following polynomial of order $m\geq 0$,
where $\beta =\left( \beta _{1}^{\prime },...,\beta _{m+1}^{\prime }\right) ^{\prime }\in \mathbb{R}^{\left( m+1\right) d_{\theta }}$ with $\beta _{i+1}=\theta ^{\left( i\right) }\left( u\right) =\partial ^{i}\theta \left( u\right) /\partial u^{i}\in \mathbb{R}^{d_{\theta }}$ and
Next, to control the approximation error, $\theta \left( t/n\right) -\theta _{u,\beta }^{\ast }\left( t/n\right) $, we introduce a kernel weighted version of the "global" objective function evaluated at the polynomial approximation,
where $K_{b}\left( \cdot \right) =K\left( \cdot /b\right) /b$, $K:\mathbb{R} \mapsto \mathbb{R}$ is a kernel function, and $b=b_{n}>0$ a bandwidth. The kernel weights ensure that when $t/n-u$ is "large", the corresponding observations are down weighted in the estimation, thereby controlling for the aforementioned approximation error. We then estimate the polynomial coefficients by
where
The estimated $\beta $ coefficients are used as estimates of $\theta \left( u\right) $ and its first $m$ derivatives, $\hat{\theta}^{\left( i\right) }\left( u\right) =\hat{\beta}_{i+1}\left( u\right) $, $i=0,...,m$. When $m=0$ , we recover the standard local-constant estimator. The above class of estimators is similar to the ones considered in fan1995JASA for so--called varying--coefficient models, except that we consider time series models with the "regressor" that we smooth over being normalized time, $t/n$ , and do not restrict $\theta \mapsto \ell _{n,t}\left( \theta \right) $ to be convex.
The choice of the order of the polynomial, $m$, should reflect the degree of smoothness that we are willing to assume $u\mapsto \theta \left( u\right) $ has. If $\theta \left( u\right) $ is $m$ times differentiable, then we should use this $m$ in the estimation for optimal control of the bias in the nonparametric estimation. On the other hand, increasing the order of the polynomial tend to increase the variability of the resulting estimator, since more local parameters are introduced in the estimation. For a further dicussion of this issue, we refer the reader to Section 3.3 of Fan2018.
Our framework includes Markov processes and stochastic processes with infinite memory doukhan2008,bardet2022 as special cases. In Section (ref), we apply our general theory to the following class of Markov models with exogenous co--variates,
where $G:\mathcal{Y}^{q}\times \mathcal{X}\times \mathcal{E}\times \Theta $ is a known function, $X_{n,t-1}$ is a vector of exogenous co--variates and $ \varepsilon _{t}$ is a sequence of errors. For a given specification of $G$ and the distribution of $\varepsilon _{t}$, we can then derive the corresponding log--likelihood for the model with time--invariant parameters, $\ell _{n,t}\left( \theta \right) =\ell \left( Z_{n,t},X_{n,t-1};\theta \right) $, where $Z_{n,t}=\left( Y_{n,t},X_{n,t}\right) $. Below, we provide three examples of models that our theory applies to:
We here provide an asymptotic theory for $\hat{\beta}$. One complication of this analysis is that the components of $\hat{\beta}$ converge with different rates. We follow the existing literature and handle this issue by introducing a re--scaled version of $\hat{\beta}$; see, e.g., han2014 for a similar approach: Define $\hat{\alpha}=U_{n}\hat{\beta}=(\hat{\theta} \left( u\right) ^{\prime },b\hat{\theta}^{\left( 1\right) }\left( u\right) ^{\prime },...,b^{m}\hat{\theta}^{\left( m\right) }\left( u\right) ^{\prime })^{\prime }$, where
is a weighting matrix containing their relative convergence rates, Given that $U_{n}$ is non-singular, the estimation problem ((ref) ) is equivalent to solving
where $D_{m,b}\left( u\right) =D_{m}\left( u/b\right) $ and
with $\mathcal{K}$ denoting the support of $K$. We will then analyze the properties of $\hat{\alpha}$.
Due to the time--varying parameters, $Z_{n,t}$ will generally be non--stationary. To develop an asymptotic theory that allows for this feature, we will rely on the concept of local stationarity as introduced by dahlhaus1997; see also dahlhaus2006 and dahlhaus2017 . We first generalize this concept to sequences of random functions:
If $W_{n,t}\left( \theta \right) =W_{n,t}$ does not depend on any parameters, we write LS$\left( p,q\right) $. The above condition states that $W_{n,t}\left( \theta \right) $ may be non--stationary, but it is locally in time well--approximated by a stationary version $W_{t}^{\ast }\left( \theta |u\right) $. Compared to existing definitions of local stationarity, we allow for an additional term $\rho ^{t}$ to appear in the approximation error. This is needed in order to allow for the initial value of $ W_{n,t}\left( \theta \right) $ to be chosen arbitrarily. In contrast, by not including $\rho ^{t}$ in their defintions, most of the existing literature implicitly assumes that $W_{n,t}\left( \theta \right) $ has been initialised at $W_{n,0}\left( \theta \right) =W_{0}^{\ast }\left( \theta |u\right) $. When used in the analysis of local estimators, this latter definition implicitly requires that the data-generating process changes as the researcher varies $u$ in the local log-likelihood which is a rather peculiar assumption. In contrast, the above definition allows for $W_{n,0}\left( \theta \right) $ to be initialized at a given fixed value -- as long as the impact of this dies out with rate $\rho $.
For an example of how the additional error term appears in autoregressive models, we refer the reader to the proof of Lemma (ref) in Section (ref) which allows for arbitrary initialisation of the data-generating process. The additional error term due to different initializations is here assumed to decay geometrically and so our definition rules out long-memory type processes. This is mostly for simplicity and we expect that most of our results can be generalized to allow for slower decay rates.
We will then require that $\ell _{n,t}\left( \theta \right) $ is ULS$\left( p,q,\Theta \right) $ with stationary approximation $\ell _{t}^{\ast }\left( \theta |u\right) =\ell \left( \mathcal{Z}_{t}^{\ast }\left( u\right) ,\theta \right) $ where $\mathcal{Z}_{t}^{\ast }\left( u\right) =\left( Z_{t}^{\ast }\left( u\right) ,Z_{t-1}^{\ast }\left( u\right) ,....\right) $ is the stationary solution to the model being estimated when $\theta _{n,t}=\theta \left( u\right) $ is constant. To illustrate, consider again ((ref)). Under regularity conditions on $G$ and $\varepsilon _{t}$ (see Section (ref) for details), the stationary solution will in this case take the form
where we impose the high--level condition that the exogenous co--variates are locally stationary. If the data-generating process is locally stationary, it follows under great generality that the likelihood and its derivatives are also locally stationary, c.f. Section (ref).
The next step in our proof is to establish a uniform law of large numbers (ULLN) for the stationary approximation of $Q_{n}\left( \alpha |u\right) $, $ Q_{n}^{\ast }\left( \alpha |u\right) =\frac{1}{n}\sum_{t=1}^{n}K_{b}\left( t/n-u\right) \ell _{t}^{\ast }\left( D_{b}\left( t/n-u\right) \alpha |u\right) $. A sufficient condition for a ULLN to hold is that $\theta \mapsto \ell _{n,t}^{\ast }\left( \theta |u\right) $ is $L_{p}$ -continuous:
Imposing $L_{p}$-continuity w.r.t. $\theta $ is weaker than almost surely continuity: If $\theta \mapsto W_{t}^{\ast }\left( \theta |u\right) $ is almost surely continuous with $\mathbb{E}\left[ \sup_{\theta \in \Theta }\left\Vert W_{t}^{\ast }\left( \theta |u\right) \right\Vert ^{p}\right] <\infty $ the process is also $L_{p}$-continuous since $DW_{t}(\delta )=\sup_{\Vert \theta -\theta ^{\prime }\Vert \leq \delta }\left\Vert W_{t}^{\ast }\left( \theta |u\right) -W_{t}^{\ast }\left( \theta ^{\prime }|u\right) \right\Vert ^{p}$, $\delta >0$, will then satisfy $\lim_{\delta \rightarrow 0}DW_{t}(\delta )=0$ almost surely and so, by dominated convergence, $\lim_{\delta \rightarrow 0}\mathbb{E}[DW_{t}(\delta )]=0$. It is easily verified that $L_{p}$-continuity w.r.t. $\theta $ implies stochastic equicontinuity of $Q_{n}^{\ast }\left( \alpha |u\right) $ and so a ULLN holds, c.f. Lemma (ref)(i) in Appendix (ref).
We are now ready to state the regularity conditions used to show consistency:
Assumption (ref)(i) imposes stronger than usual assumptions on $ K$ and excludes, among others, the Gaussian kernel and higher-order kernels. It includes, on the other hand, the Epanechnikov and the triangular kernel. The restriction that $K\left( \cdot \right) \geq 0$ is used to ensure identification of the parameters when $m>0$; without this, identification is not necessarily guaranteed; see below for further discussion. For the analysis of the local constant estimator ($m=0$), all subsequent results will go through with $K$ having full support and taking negative values.
The compact support assumption greatly simplifies our analysis of local polynomial estimation of non-concave models: In order to establish uniform convergence of the likelihood we require $\Theta $ to be compact as is standard in the literature. But under this restriction, it is easily checked that $D_{m,b}\left( v\right) \alpha \notin \Theta $ as $b\rightarrow 0$ for any given $\alpha =\left( \alpha _{1},...,\alpha _{m+1}\right) $ with $ \alpha _{i}\neq 0$ for some $i\geq 1$ and any $v\neq 0$. Thus, to allow for kernels with unbounded support, we would generally need the parameter space $ \mathcal{A}$, as defined in ((ref)), to collapse at $\left\{ \left( \alpha _{1},0,...,0\right) :\alpha _{1}\in \Theta \right\} $ as $ b\rightarrow 0$. Such shrinking behaviour in turn means that a formal Taylor expansion of $\ell _{n,t}\left( D_{m,b}\left( v\right) \alpha \right) $ w.r.t. $\alpha $ is difficult to obtain and so standard arguments to establish asymptotic normality of $\hat{\alpha}$ cannot be applied. On the other hand, by restricting the support $\mathcal{K}$ to be compact, $ K_{b}\left( v\right) \ell _{n,t}\left( D_{m,b}\left( v\right) \alpha \right) $ is well-defined for all $\alpha \in \mathcal{A}$ and $v\in \mathbb{R}$ (where we set $K_{b}\left( v\right) \ell _{n,t}\left( D_{m,b}\left( v\right) \alpha \right) =0$ for $v/b\notin \mathcal{K}$). Moreover, $\left( \alpha _{1},0,...,0\right) $ is an interior point of $\mathcal{A}$ and so in our analysis of $\hat{\alpha}$ we can employ standard arguments involving a Taylor expansion of the score function around this point.
Assumption (ref) is standard in the analysis of \textquotedblleft global\textquotedblright\ extremum estimators of stationary models on the form $\tilde{\theta}\left( u\right) =\arg \max_{\theta \in \Theta }\sum_{t=1}^{n}\ell _{t}^{\ast }\left( \theta |u\right) /n$. In particular, for a given time series model, we can import existing results for verification of Assumption (ref) (ii)-(iii); see Section (ref) for more details. Assumption (ref) in conjunction with $K\left( \cdot \right) \geq 0$ ensures that the local polynomial estimator identifies $\theta \left( u\right) $. If we allow for kernels that take negative values, we have to replace (ref)(iii) with the following more abstract identification condition: The function $Q^{\ast }\left( \alpha |u\right) =\int K\left( v\right) \mathbb{E}\left[ \ell _{t}^{\ast }\left( D_{m}\left( v\right) \alpha |u\right) \right] dv$ satisfies $Q^{\ast }\left( \alpha |u\right) <Q^{\ast }\left( \left( \theta \left( u\right) ,0,...,0\right) |u\right) $ for any $\alpha \neq \left( \theta \left( u\right) ,0,...,0\right) $. We have not been able to provide primitive conditions for this to hold when $K$ can take negative values and so instead impose the positivity constraint on $K$.
If the objective function $\theta \mapsto \ell _{n,t}\left( \theta \right) $ is concave and $\Theta $ is convex, we can replace Assumption (ref) with the following pointwise versions: For any $\theta \in \Theta $, $\ell _{n,t}\left( \theta \right) $ is locally stationary and $\mathbb{E}\left[ |\ell _{t}^{\ast }\left( \theta |u\right) |\right] <\infty $; see Theorem 2.7 in newey1994. Under the above assumptions, the following consistency result holds:
Note that the above theorem only shows consistency of $\hat{\theta}\left( u\right) $ and so at this stage we cannot make any statements regarding $ \hat{\theta}^{\left( i\right) }\left( u\right) $, $i=1,...,m$. This is similar to other results for nonlinear extremum estimators that converge with different rates; see, e.g., Theorem 9 in han2014 where a global consistency result is only provided for the component with the fastest rate.
However, under additional regularity conditions on the quasi-likelihood function, we can provide a more precise analysis of the estimators, including local consistency of $\hat{\theta}^{\left( k\right) }\left( u\right) $, $1\leq k\leq m$. With $s_{n,t}\left( \theta \right) =\partial \ell _{n,t}\left( \theta \right) /\left( \partial \theta \right) \in \mathbb{ R}^{d_{\theta }}$ and $h_{n,t}\left( \theta \right) =\partial ^{2}\ell _{n,t}\left( \theta \right) /(\partial \theta \partial \theta ^{^{\prime }})\in \mathbb{R}^{d_{\theta }\times d_{\theta }}$, $D_{n,t}\left( u\right) =D_{m,b}\left( t/n-u\right) $ and $K_{n,t}\left( u\right) =K_{b}(t/n-u)$, the score and hessian of $Q_{n}\left( \alpha |u\right) $ are given by
It is easily checked that $\alpha _{0}:=U_{n}\beta _{0}$, where $\beta _{0}=(\theta \left( u\right) ^{\prime },\theta ^{\left( 1\right) }\left( u\right) ^{\prime },...,\theta ^{\left( m\right) }\left( u\right) ^{\prime })^{\prime }$, belongs to the interior of $\mathcal{A}$ for all $n$ large enough due to Assumption (ref)(ii) below in conjunction with Assumption (ref). Due to the consistency result, so will $\hat{\alpha}$ w.p.a. 1. Thus, $\hat{\alpha}$ will satisfy the first-order condition of ((ref)) which combined with the mean-value theorem yields
where $\bar{\alpha}$ is situated on the line segment connecting $\hat{\alpha} $ and $\alpha _{0}$. We then decompose the score function into a bias and variance component, $S_{n}\left( \alpha _{0}|u\right) =B_{n}\left( u\right) +S_{n}\left( u\right) $, where
$s_{n,t}=s_{n,t}\left( \theta \left( t/n\right) \right) $, and $ b_{n,t}=s_{n,t}\left( \theta _{u}^{\ast }\left( t/n\right) \right) -s_{n,t}\left( \theta \left( t/n\right) \right) $ with $\theta _{u}^{\ast }\left( t/n\right) $ defined in eq. ((ref)). This decomposition is different from the one usually employed in the analysis of kernel estimators of time-varying coefficients where $s_{n,t}\left( \theta \left( t/n\right) \right) $ is replaced by the stationary version of the score function evaluated at $\theta \left( u\right) $, $s_{t}^{\ast }\left( \theta \left( u\right) |u\right) $; see, e.g., dahlhaus2017 and dahlhaus2006. This "usual" choice has as consequence that the corresponding bias term in generally involves the time derivative process of the score function and so the resulting analysis tends to impose stronger regularity conditions on the model being estimated. By instead centering the analysis around $s_{n,t}$, we can obtain the leading term of the bias $ B_{n}\left( u\right) $ through a standard Taylor expansion w.r.t. $\theta $,
Thus, our approach allows for a simpler derivation of the leading bias and variance terms under the following weak regularity conditions, where here and in the following $p$ and $q$ have to satisfy $p\geq 1$ and $q>0$, but can otherwise vary depending on the particular application.
Similar to Assumption (ref), Assumption (ref) contains standard regularity conditions used in the analysis of regular parametric estimators on the form $\tilde{\theta}\left( u\right) =\arg \max_{\theta \in \Theta }\sum_{t=1}^{n}\ell _{t}^{\ast }\left( \theta |u\right) $. At the same time, Assumption (ref)(iii) is non-standard compared to the existing literature (as discussed above) and, together with (ref)(i), allows us to apply a novel martingale central limit theorem for locally stationary sequences to $S_{n}\left( u\right) $,
see Lemma (ref)(iii) in Appendix (ref). This result can be seen as a generalisation of the standard CLT for stationary and ergodic MGD's that allows for locally stationary processes. The MGD assumption amounts to assuming that the time-varying model is correctly specified and has to be verified on a case-by-case basis.
Finally, Assumption (ref)(ii) together with the expansion in eq. ((ref)) is used to derive the limits of $B_{n}\left( u\right) $ and $H_{n}\left( \bar{\alpha}|u\right) $,
where $\mu _{i}=\int_{\mathbb{R}}K\left( v\right) v^{m+i}D_{m}\left( v\right) dv$ and $\mathbb{K}_{i}=\int_{\mathbb{R}}K^{i}\left( v\right) D_{m}\left( v\right) D_{m}\left( v\right) ^{\prime }dv$, $i\geq 1$. Combining ((ref)), ((ref)) and ((ref)), we obtain:
Similar to existing results for local polynomial estimators in a cross-sectional setting, the leading bias term in ((ref)) only depends on $\theta ^{\left( m+1\right) }\left( u\right) $ and so the estimators adapt to the curvature of $\theta \left( u\right) $. The asymptotic variance in Theorem (ref) can be estimated using plug-in methods: It follows from the proof of Theorem (ref) that
satisfies $\hat{W}\left( u\right) \rightarrow ^{p}\mathbb{K}_{2}\otimes \Omega \left( u\right) $ while $H_{n}\left( \hat{\alpha}|u\right) \rightarrow ^{p}\mathbb{K}_{1}\otimes H\left( u\right) $.
Compared to most existing asymptotic results for the local constant estimator, such as dahlhaus2017, above result with $m\geq 1$ impose much weaker restrictions on the bandwidth. In particular, standard bandwidth selection rules can be employed here but not under most of the existing theories since their conditions require undersmoothing (that is, $ b\rightarrow 0$ at a faster rate than the optimal one). This is due to the fact that these theories do not provide a complete characterisation of the leading bias term. The few papers that do characterise the leading bias term, such as dahlhaus2006, require the so--called time derivatives of the stationary score function to exist and be well-behaved since these enter their bias expressions. Our conditions and results, on the other hand, do not require these and are analogous to the ones found in the literature on local polynomial likelihood estimators; see, e.g., Theorem 1b of fan1995JASA.
Equation ((ref)) holds for any value of $m\geq 0$ and $ i=0,...,m$. However, if $K$ is symmetric, then $\kappa _{1,i}=0$ when $m-i$ is even. In particular, for the local constant estimator ($m=i=0$), Theorem (ref) only informs us that the bias component of $\hat{\theta} \left( u\right) $ is $o_{p}\left( b\right) $ which is not a sharp rate. To obtain the leading bias term in the cases where $m-i$ is even, a higher-order expansion of $b_{n,t}$ in eq. ((ref)) is necessary. This expansion requires additional assumptions involving aforementioned time derivatives and standard derivatives w.r.t. $\theta $ of $h_{t}^{\ast }\left( \theta \left( u\right) |u\right) $. To present these, we need the following additional concept:
Our definition of time differentiability is slightly weaker compared to the one found in dahlhaus2017 and other papers where differentiability w.r.t. $u$ has to hold almost surely. With this definition in hand, we are ready to introduce the following additional regularity conditions in order to derive the leading bias term when $m-i$ is even:
The time-derivative $\partial _{u}h_{t}^{\ast }\left( \theta |u\right) $ will generally involve time-derivatives of the underlying stationary approximation of data: If $h_{t}^{\ast }\left( \theta |u\right) =h\left( \mathcal{Z}_{t}^{\ast }\left( u\right) ;\theta \right) $ for some function $ h $ which is differentiable w.r.t. $\mathcal{Z}_{t}^{\ast }\left( u\right) $ , then it takes the form $\partial _{u}h_{t}^{\ast }\left( \theta |u\right) =\sum_{i=0}^{\infty }\partial h\left( z_{0},z_{1},z_{2},....;\theta \right) /\left( \partial z_{i}\right) |_{z=\mathcal{Z}_{t}^{\ast }\left( u\right) }\times \partial _{u}Z_{t-i}^{\ast }\left( u\right) $, where $\partial _{u}Z_{i,t}^{\ast }\left( u\right) $ is the time derivative of $Z_{t}^{\ast }\left( u\right) $. The short memory condition imposed in Assumption (ref) is used to control the variance component of the first-order bias term derived in Theorem (ref). In Section (ref), we use the concept of $\tau $--weak dependence doukhan2008 to verify Assumption (ref).
Under the above additional assumptions, we obtain the following higher-order expansion of the bias component:
As a special case, we obtain the following result for the local constant estimator:
To our knowledge this is the first complete characterization of the leading bias term of local constant estimators in general time-varying parameter models. The final characterisation of the bias, $B_{1}(u)+B_{2}\left( u\right) =-\frac{1}{2}\partial _{u}^{2}\mathbb{E}\left[ s_{t}^{\ast }\left( \theta \left( v\right) |u\right) \right] _{v=u}$, corresponds to the one obtained in dahlhaus2006 for the time--varying ARCH model. This characterisation, however, requires $s_{t}^{\ast }\left( \theta |u\right) $ to be twice differentiable w.r.t. $u$, while our characterisation only requires $h_{t}^{\ast }\left( \theta |u\right) $ to be once differentiable w.r.t $u$.
Comparing Theorems (ref) and (ref), we see that the local linear and local constant estimators share convergence rate and asymptotic variance, but that the latter suffers from additional biases. This is consistent with the theory found for local constant and local linear estimators in a cross-sectional settingl; ; see, e.g., fan1995JASA.
The above theory for the local constant estimator does not cover discrete-valued time series models. Specifically, Theorem (ref) requires $h_{t}^{\ast }\left( \theta |u\right) $ to be differentiable w.r.t. $u$, c.f. Assumption (ref). This property rarely holds when $h_{t}^{\ast }\left( \theta |u\right) $ is a function of discrete--valued random variables since these are generally not smooth functions of the underlying parameters of the model; see Truquet2019 for more details. Thus, Theorem (ref) does not apply to, for example, the Poisson autoregressive model.
But Theorem (ref) still applies. We therefore combine the ideas of Truquet2019,Truquet2020 with Theorem (ref) to obtain a theory for the local constant estimator that also covers models with discrete-valued outcomes. This is achived by replacing Assumptions (ref) and (ref) with the following ones:
If $v\mapsto h_{t}^{\ast }\left( \theta \left( u\right) |v\right) $ is $ L_{1} $-differentiable w.r.t. $v$ at $u$, then $\partial _{v}\mathbb{E}\left[ h_{t}^{\ast }\left( \theta \left( u\right) |v\right) \right] _{v=u}=\mathbb{E }\left[ \partial _{v}h_{t}^{\ast }\left( \theta \left( u\right) |v\right) \right] _{v=u}$. Thus, Assumption (ref) is weaker than Assumption (ref) and is satisfied as long as the cumulative distribution function of $h_{t}^{\ast }\left( \theta \left( u\right) |v\right) $ is differentiable w.r.t. $v$, c.f. Lemma (ref) below. This property holds for\ many discrete-valued models, including Poisson autoregressions and dynamic discrete choice models, c.f. Truquet2019,Truquet2020.
Assumption (ref), on the other hand, is stronger than Assumption (ref). However, similar to Assumption (ref), part (i) is satisfied if $h_{t}^{\ast }\left( \theta \left( u\right) |v\right) $ is $\tau $-weakly dependent for $v$ in a neighbourhood of $u$ since this in turn implies that the joint process $ \left( h_{t}^{\ast }\left( \theta \left( u\right) |v_{1}\right) ,h_{t}^{\ast }\left( \theta \left( u\right) |v_{2}\right) \right) $ is weakly dependent for $\left( v_{1},v_{2}\right) $ in a neighbourhood of $\left( u,u\right) $. Moreover, part (ii) will hold under the same conditions that ensure Assumption (ref) holds, namely that the joint distribution function of $\left( h_{0}^{\ast }\left( \theta \left( u\right) |v_{1}\right) ,h_{t}^{\ast }\left( \theta \left( u\right) |v_{2}\right) \right) $ is differentiable.
The following result shows that the results for the local constant estimator remains essentially the same under Assumptions (ref)-- (ref) in place of (ref)--(ref):
We have already seen that the local linear estimator has smaller biases than the local constant one in the interior of its domain, $u\in \left( 0,1\right) $. Another well-known advantage of the local linear estimators in a cross-sectional setting is that they exhibit automatic boundary carpentering. This property also holds in our setting where the boundaries are $u=0$ and $u=1$. Since the results for $u=1$ are similar, we here only analyze the properties of the local constant ($m=0$) and the local linear ($ m=1$) estimators at $u=cb$ for some constant $c>0$. Combining the intermediate bias--variance analysis carried out in the proofs of Theorem (ref) and (ref) with the arguments of fan1995JASA, we find that the two estimators remain asymptotically normally distributed but their asymptotic biases and variances take different forms:
We refer to fan1995JASA for the precise expressions of the constants $\kappa _{i,j}^{c}$, $i=1,2$ and $j=0,1,2$. At the boundary, both the biases and variances of the local constant and local linear estimators are now different. While the difference between two asymptotic variances is a constant scale, compare $a_{0}$ and $a_{1}$ above, the biases are now of a different order: The local linear estimator still enjoys a bias of order $ O\left( b^{2}\right) $ while the bias of the local constant one blows up and becomes of order $O\left( b\right) $. Thus, the local constant estimator will generally suffer from significantly larger biases at the boundary compared to the local linear one.
To demonstrate the usefulness of our general results, we here apply our theory to Markov models with exogenous co--variates and infinite memory autoregressive models, respectively. In the process, we provide more primitive conditions for the critical assumption of ULS.
By inspection of our Assumptions (ref)--(ref), we observe that, for a given parametric model and estimator, most of the assumptions are easily verifiable using standard arguments known from the literature on regular parametric estimators. The only ones that are non--standard are Assumptions (ref) and (ref), which require the reserarcher to show that $\ell _{n,t}\left( \theta \right) $ and its first two derivatives are ULS. Similarly, Assumption (ref) is also a ULS requirement, while Assumption (ref) and/or (ref) will hold if $h_{t}^{\ast }\left( \theta \left( u\right) |u\right) $ is weakly dependent.
The following lemma provides sufficient conditions for a given transformation of a time series to be ULS if the underlying time series is ULS. It also shows that if the stationary version is $\tau $--weakly dependent then so is the transformation.
In our applications below, we will then require $\ell _{n,t}\left( \theta \right) $ and its first two derivatives to satisfy these conditions so that Assumptions (ref), (ref), (ref) and (ref) are satisfied.
Next, we here provide more primitive conditions for Assumption (ref) that can be be applied to discrete--valued random variables:
What remains is to show that $u\mapsto p\left( w;\theta ,u\right) $ is differentiable. This is fairly straightforward for Markov models, where we can import results from, e.g., Truquet2020 and VazquezAbad1992 ; see proof of Corollary (ref) for an example of this. A sufficient condition for $\int \left\vert f\left( w\right) \partial _{u}p\left( w;\theta ,u\right) \right\vert d\mu \left( w\right) <\infty $ is $\int \left\vert f\left( w\right) \right\vert \frac{\left\vert \partial _{u}p\left( w;\theta ,u\right) \right\vert }{p\left( w;\theta ,u\right) } p\left( w;\theta ,u\right) d\mu \left( w\right) <\infty $. For example, if $ \left\vert \partial _{u}p\left( w;\theta ,u\right) \right\vert /p\left( w;\theta ,u\right) \leq C\left( 1+\left\Vert w\right\Vert ^{r}\right) $, $ r\geq 0$, then we need $E\left[ \left\Vert W_{t}^{\ast }\left( \theta |u\right) \right\Vert ^{r}\left\Vert f\left( W_{t}^{\ast }\left( \theta |u\right) ;u\right) \right\Vert \right] <\infty $.
We summarise our findings in the following corollary:
If the model and estimator of interest is regular in the sense that its stationary version satisfies above assumptions, then all that remains to be shown is that $Z_{n,t}$ is locally stationary with $Z_{t}^{\ast }\left( u\right) $ being weakly dependent. The next subsection provides primitive conditions for this to hold for Markov models with exogenous co--variates, while the second subsection discuss application to infinite memory models. In the third subsection, we revisit the examples of Section (ref) and provide primitive conditions under which our main results apply to these.
We first consider $q$-Markov models without covariates on the form
where $G:\mathcal{Y}^{q}\times \mathcal{E}\times \Theta \mapsto \mathcal{Y}$ is some known mapping, $\varepsilon _{t}\in \mathcal{E}\subseteq \mathbb{R} ^{d_{\varepsilon }}$ is a sequence of i.i.d. errors, and $\theta \left( \cdot \right) \in \Theta $. Importantly, the initial value $Y_{n,0}$ can be arbitrarily chosen which is in contrast to most of the existing literature. Under regularity conditions, its corresponding stationary approximation $ Y_{t}^{\ast }\left( u\right) $ will solve
We impose the following assumptions:
This assumption is similar to the one found in dahlhaus2017, but we here allow for $p\neq \tilde{p}.$ Assuming we can verify $\mathbb{E}\left[ \lVert Y_{t}^{\ast }\left( u\right) \rVert ^{\tilde{p}r}\right] <\infty $ this allows us to show higher-order local stationarity ($\tilde{p}>p$). Furthermore, part (iv) only requires suitable moments of $Y_{n,0}$ but otherwise the process can be initialized at any given value while most of the existing literature implicitly assumes $Y_{n,0}=Y_{0}^{\ast }\left( u\right) $.
Next, we extend the results to the following class of $q$-Markov models with exogenous co-variates:
We could allow for $X_{n,t}$ to exhibit richer dynamics, e.g., allow for $ X_{n,t}$ to depend on lags of $Y_{n,t}$. However, this would lead to more complicated assumptions and so we here maintain ((ref)) for simplicity. We impose the following assumptions on ((ref))--( (ref)):
Combining Corollary (ref) with Lemma (ref), we obtain a set of easily verifiable primitive conditions for local M-estimators of time-varying parameters in Markov models to satisfy Theorems (ref) and (ref); see Section (ref) for examples of the verification procedure.
Next, we consider the following class of infinite memory models,
The local constant estimator of $\theta \left( u\right) $ was analysed in bardet2022 using techniques different from ours.\ Specifically, they provide primitive conditions under which $Y_{n,t}$ satisfies
where $Y_{t}^{\ast }\left( u\right) $ is the stationary solution to $ Y_{t}^{\ast }\left( u\right) =G\left( Y_{t-1}^{\ast }\left( u\right) ,Y_{t-2}^{\ast }\left( u\right) ,....,\varepsilon _{t};\theta \left( u\right) \right) $. But this result unfortunately does not suffice for our theory since this requires $Y_{n,t}$ to satisfy our version of local stationarity as given in Definition (ref). At the same time, bardet2022 do provide primitive conditions for the assumptions on the stationary version of Corollary (ref) to hold. Thus, under the high--level assumption that $Y_{n,t}$ satisfies our version of local stationarity, we can combine the results of bardet2022 with our Corollary (ref) to obtain a theory for local polynomial estimators of infinite memory models. Under this high--level condition, we are also able to provide a precise characterisation of the leading bias term of their estimator, something which was missing in the analysis of bardet2022. We leave it to future research to establish conditions under which $Y_{n,t}$ is locally stationary a la Definition (ref).
We here apply Corollary (ref) to Examples (ref)--(ref). Note here that the first and third examples cannot be analysed using existing theories since these rely on differentiability of the model. Moreover, none of the existing theories allow for co--variates to be included in the model. As such, the results below are new to the literature.
Throughout this section the following assumption is maintained where $\theta \left( u\right) $ and $\Theta $ are specified in each of the following examples:
For the time--varying TAR-X model in Example (ref), we obtain the following result:
This appears to be the first result for the threshold AR--X\ model in the literature. Note here that we do not need to restrict $\Theta $ to be compact since here $\ell _{n,t}\left( \theta \right) $ is concave. The Eq. ( (ref)) ensures that the model indeed has a locally stationary solution and is $\tau $-weakly dependent. The additional restrictions imposed in the second part of the corollary are used to show that the time--derivative of $\left( Y_{t}^{\ast }\left( u\right) ,X_{t}^{\ast }\left( u\right) \right) $ exists so that Assumption (ref) holds.
Next, consider the local Gaussian QMLE of the tv--ARCH-X model given in Example (ref). Under eq. ((ref)) below, there exists a locally stationary solution to the model which takes the form
where $\tilde{X}_{t}^{\ast }(u)=\left( 1,Y_{t-1}^{\ast }(u),...,Y_{t-q}^{\ast }(u),X_{t-1}^{\ast }\left( u\right) ^{\prime }\right) ^{\prime }$.
Our conditions on the model are more or less identical to bardet2022 but allows for exogenous regressors to be included. Our conditions are substantially weaker compared to dahlhaus2006 and inoue2019 who impose much stronger moment conditions.
Finally, we apply our results to the local MLE of the tv--PARX model in Example (ref). Under ((ref)) below, the tv--PARX\ process is locally stationary with stationary solution $Y_{t}^{\ast }(u)| \mathcal{F}_{t-1}^{\ast }(u) \sim \mathrm{Poisson}\left( \lambda _{t}^{\ast }\left( u\right) \right) $, where $\lambda _{t}^{\ast }\left( u\right) =\theta \left( u\right) ^{\prime }\tilde{X}_{t}^{\ast }\left( u\right) $ and $\tilde{X}_{t}^{\ast }(u)=\left( 1,Y_{t-1}^{\ast }(u),...,Y_{t-q}^{\ast }(u),X_{t-1}^{\ast }\left( u\right) ^{\prime }\right) ^{\prime }$.
In this section, we examine the finite-sample performances of the local constant and local linear estimators. All reported results are based on 1000 simulated data sets. The over--all performance of the estimators is evaluated using the mean absolute deviation error (MADE), $MADE_{i}:=\frac{1 }{n}\sum_{t=1}^{n}\mathbb{E}\left[ |\hat{\theta}_{i}\left( t/n\right) -\theta _{i}\left( t/n\right) |\right] $, as well as their integrated bias, variance, and mean squared error. All results are based on the Epanechnikov kernel and with the bandwidth chosen using the cross-validation method proposed in richter.
We first consider the time-varying ARCH(1) in eq. ((ref)) where $ \varepsilon \sim i.i.d.N\left( 0,1\right) $, $\omega \left( u\right) =0.7-0.5\sin \left( 4\pi u\right) $ and $\alpha \left( u\right) =0.45+0.4\sin \left( 4\pi u\right) $. We estimate $\omega \left( u\right) $ and $\alpha \left( u\right) $ using both Gaussian log-likelihood and the WLS method of fryzlewicz2008. Table (ref) reports the performance of the estimators. For all sample sizes, the local MLE's perform better than the local WLS estimators. For sample sizes of $n=250$ and $n=500$ , the local constant MLE performs as well as the local linear MLE in terms of the global measures, but the latter performs best in terms of IMSE and MADE for $n=1000$. Thus, the over--all superiority of the local linear estimator indicated by the theory appears to be a large--sample property.
To compare the performance of the estimators near the end of the sample, we also evaluate the bias of the estimators for the first and last 2.5% of time periods corresponding to $u\in \left[ 0,0.025\right] \cup \left[ 0.975,1 \right] $. This is reported as IBias2BD in Table (ref). As predicted by the theory, we find that relative to the local constant versions the local linear WLS and ML estimators enjoy significantly smaller biases near the boundaries of $\left[ 0,1\right] $ for all sample sizes.
We here report simulation results for the local linear MLE of the following PARX(1) model with an additional exogenous regressor $X_{n,t}$,
where $\omega\left(u\right)=0.6-0.3u+0.3\sin\left(2\pi u\right)$, $ \alpha\left(u\right)=0.3+0.3u-0.3\sin\left(2\pi u\right)$ and $ \gamma\left(u\right)=1-0.5\cos\left(\pi u\right)$. The dynamics of the exogenous regressor was chosen as
where either
Table (ref) reports the over-all performance of the estimators in terms of integrated squared bias, variance, MSE and MADE. The table shows that, in all sample sizes, the local linear estimator behaves well for both DGP1 and DGP2. All bias, variance, and MADE decrease as the sample size increases. Finally, similar to the case of the tvARCH model, the local linear estimator performs well near the boundaries; we leave out these results since they are similar to the ones reported for the tv-ARCH model above.
We here revisit the empirical analysis of US corporate defaults carried out in agosto2016 with the aim of examining whether there is evidence of structural instability in the time series. The data set consists of monthly number of bankruptcies among Moody's rated industrial firms in the United States for the period 1982--2011 ($n=360$ observations), collected from Moody's Credit Risk Calculator (CRC). Figure (ref) shows the time series of default counts together with its sample autocorrelation function, which reveals high temporal dependence in default counts and existence of default clusters over time.
With $Y_{n,t}\in \left\{ 0,1,2,\ldots \right\} $, $t\geq 1$ denoting the number of defaults in a given month, we use a log--PARX model to gain better understanding of the dynamics of $Y_{n,t}$ as a function of its own past, $ Y_{n,t-m}$, $m\geq 1$, but also in terms of additional covariates $ X_{n,t}\in \mathbb{R}^{d_{x}}$, which include relevant macroeconomic and financial factors as considered in agosto2016. We model $Y_{n,t}$ as a conditional Poisson distribution $Y_{n,t}|\mathcal{F}_{n,t-1}\sim \mathrm{ Poisson}\left( \lambda _{n,t}\left( \theta \left( t/n\right) \right) \right) $, $t=1,2,\ldots ,n$, where the intensity $\lambda _{n,t}\left( \theta \left( t/n\right) \right) $ depends on past counts, co--variates $X_{n,t-1}$ and a vector of time-varying parameters $\theta \left( t/n\right) $. Our favoured specification of $\lambda _{n,t}\left( \theta \right) $ is the log--PARX model of fokianos2011, which is here augmented by the chosen set of exogenous variables, $X_{n,t-1}$,
so that $\theta =\left( \omega ,\alpha _{1},...,\alpha _{p},\gamma ^{\prime }\right) ^{\prime }$.
We here deviate from agosto2016 that specifies $\lambda _{n,t}$ to be a linear function of past counts and factors. The reasons for us favouring the above log--specification over the linear one are three--fold: First, ( (ref)) do not impose positivity constraints on the parameters which facilitates the numerical computation of the estimators; second, it allows us to include any predictors we wish without the need of first transforming them to ensure that each component of the resulting $X_{n,t-1}$ is positive; third, when we estimate both models using the US default data, we found that the log--specification delivers a better fit.
Similar to agosto2016, we use $\sum_{i=1}^{p}\alpha _{i}\left( t/n\right) $ as a measure of contagion in the financial markets: If $ \sum_{i=1}^{p}\alpha _{i}\left( t/n\right) $ is large then firms defaulting today will lead to a large increase in the risk of other firms defaulting next period everything else equal.
As exogenous covariates, we consider the same financial, credit market, and macroeconomic variables as in agosto2016: Realized volatility ($RV$) computed using daily squared return on the S&P 500 index, the Leading Index released by the Federal Reserve ($LI$), year-to-year change in Industrial Production Index ($IP$), one-year return on the S&P 500 index ($SPX$), the three-month Treasury bill rate ($TB3$), and BAA Moody's rated to 10--year Treasury spread ($SP$).\footnote{ These covariates are tested for the existence of default covariates in das2007, duffie2007, and lando2010.} agosto2016 decompose $LI$ and $IP$ into their positive and negative parts to deal with above--mentioned issue of $X_{n,t-1}$ having to be positive in the linear specification. In contrast, no such transformations are needed for our log-specification. Moreover, as well as $RV$, we also consider the logarithm of $RV$, $\log \left( RV\right) $, to evaluate if the latter is a better predictor; again, this would not be possible in the linear specification of agosto2016.
The lag length $p$ is chosen using BIC where, for a given model, the log--likelihood is evaluated at the estimated time-varying parameters. \footnote{ As pointed out by a referee, we are unable to provide a formal justification for using BIC to choose the lag-length in our time-varying parameter version of the PARX model. So the application of this model selection criterion to our setting is very ad hoc} According to this version of BIC, the preferred specification is $p=2$. Importantly, by allowing for the parameters to be time--varying, a much more parsimonious model is selected by BIC: If we do model selection where for each model we restrict the estimated parameters to be constant over time, the preferred model is $p=6$. Moreover, the persistence, or contagion, of the estimated time--invariant version, as measured by $\sum_{i=1}^{p}\hat{\alpha}_{i}$, is substantially higher than for the time-varying version, as measured by $\sum_{i=1}^{p}\hat{\alpha} _{i}\left( u\right) $, $u\in \left[ 0,1\right] $. This is consistent with the findings reported in Hillebrand2005 for GARCH\ models: Neglecting structural changes in parameters causes the estimates of these to be suffering from a strong upward bias which in turn leads BIC to selecting a bigger model.
As benchmark, we first estimate the time--invariant version of ((ref)) with $p=2$. Table (ref) shows the estimation results for four different specifications of the time--invariant LLPARX(2) model: Column (1) contains the results when only $\left( RV,LI\right) $ are included; column (2) when only $\left( \log RV,LI\right) $ are included; column (3) when all covariates are included except for $\log \left( RV\right) $; and column (4) when all covariates except $RV$ are included. As in agosto2016, once we control for the information contained in $RV$ and $LI$ , none of the other four covariates are found to be relevant in predicting future defaults. We also observe that the two specifications using $\log \left( RV\right) $ appear to perform better than the ones using $RV$. Based on these results, our favoured specification of the time-varying version is to use $LI$ and $\log \left( RV\right) $ as exogenous variables:
Figure (ref) shows the time--series of the local linear estimates of the time-varying parameters of ((ref)) together with the time--invariant estimates reported in column (2) of Table (ref). Pointwise confidence bands are computed based on the asymptotic distribution derived in Theorem (ref). The Leading Index is pointwise significant for most of the sample period which highlights the link between macroeconomic activity and corporate defaults also found in agosto2016. At the same time, this link exhibits substantial time--variation, in particular at the end of the sample during the Great Recession. The link between realized volatility and defaults of industrial firms is less significant but also appears to be changing over time. Similar to agosto2016, the realized volatility and the Leading Index are strong explanatory variables during the Great Recession (2007--2011). However, differently, both of them are still relevant in the late 1980s and early 1990s. Especially, the effect of $\log \left( RV\right) $ tends to be negative in this period which cannot be captured by the aforementioned linear version of the PARX model.
Finally, judging from $\hat{\alpha}_{1}\left( t/n\right) +\hat{\alpha} _{2}\left( t/n\right) $, there is very little contagion in the financial markets from 1990s and onwards, which is somewhat consistent with the findings in agosto2016. However, at the end of the sample, we actually find a negative effect of today's log--defaults on tomorrow's default risk. Also note that the time--varying estimates, $\hat{\alpha} _{1}\left( t/n\right) $ and $\hat{\alpha}_{2}\left( t/n\right) $, remain below the corresponding time--invariant ones, $\hat{\alpha}_{1}$ and $\hat{ \alpha}_{2}$ as marked by the horizontal red lines, throughout the sample period. Again, this seems to indicate that ignoring time--variation in the parameters of PARX models lead to over estimation of the level of persistence/contagion.
To assess in--sample fit and whether the reported time--variation in the parameters is statistically significant, we carry out an array of graphical and quantitative diagnostic tools for time series. First, we plot in the left panel of Figure (ref) the actual default counts together with the predicted defaults $\hat{Y}_{n,t}:=\hat{\lambda} _{n,t}=\lambda _{n,t}\left( \hat{\theta}\left( t/n\right) \right) $. As can be seen from this plot, the time--varying LLPARX model captures the default counts dynamics well. In the right panel of Figure (ref), the sample autocorrelation function of the standardized Pearson residuals $ \hat{e}_{n,t}=\hat{\lambda}_{n,t}^{-1/2}\left( Y_{n,t}-\hat{\lambda} _{n,t}\right) $ is plotted. Under correct specification, $e_{n,t}$ should be white noise -- the plotted sample autocorrelation function supports this.
We also evaluate the adequacy of fit using the probability integral transform (PIT). We follow Davis2016 and compute the PIT's by
where $\left\{ \nu _{t}\right\} $ is a sequence of i.i.d. random variables from a standard uniform distribution, and $F_{n,t}$ is the CDF of a Poisson($ \hat{\lambda}_{n,t})$ distribution. Under correct model specification, $\hat{ u}_{n,t}$ is a sequence of i.i.d. random variables from the standard uniform distribution. Figure (ref) depicts the histogram of PIT which show that the tvLLPARX(2) model provides a better in--sample fit than the corresponding time--invariant model.
Finally, in Table (ref), we report the log-likelihood, AIC and BIC values and the $p$-value from a Kolmogorov-Smirnov test of the PIT's being uniformly distributed of each of four specifications in columns (1)--(4) of Table (ref) together with the time-varying versions of (2) and (4), labelled tv (2) and tv(4), respectively. From these, we see that allowing for time--varying parameters increase the in--sample fit dramatically. While this is not a formal statistical test of time--variation, it provides strong informal evidence of such. As mentioned earlier, we also see that specification (2) is favoured over (1), (3) and (4) in the time-invariant case, and that (2) is favoured over (4) in the time--varying case.
We complete the analysis by conducting a final sensitivity analysis of ((ref)). This is done by including the remaining exogenous covariates in addition to realized volatility and Leading Index in the time--varying version of the model:
The estimation results are provided in Figure (ref). The estimates, except for the realized volatility, are consistent with our baseline findings. The Leading Index remains highly significant, which shows that macroeconomic factors are relevant in predicting future defaults. The link between short--term interest rates and defaults of industrial firms changes over time. During the late 1980s and the Great Recession (2007--2011), interest rates play a role in determining the interest expense of firms. In the 1990s, similar to the finding in duffie2007, the sign of the coefficient for the short--term rate is consistent with the fact that the US Federal Reserve often increases the short--term rates to control business expansions. Controlling for $LI$ and $TB3$, other covariates are estimated to be insignificant for most of the sample period except for the period of the Great Recession (2007--2011), in which financial, credit market, and macroeconomic variables are significant explanators of the default intensity. These are novel findings that the original analysis of agosto2016 did not reveal.