Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
71,804 characters · 12 sections · 59 citation commands
Simultaneous Inference of a Partially Linear Model in Time Series
{15pt}
\newtheorem{definition}{Definition} \newtheorem{assumption}{Assumption} \newtheorem{theorem}{Theorem} \newtheorem{proposition}{Proposition} \newtheorem{corollary}{Corollary} \newtheorem{lemma}{Lemma} \newtheorem{remark}{Remark} \newtheorem{example}{Example}
Partially linear models are of interest in many practical problems. For example, in econometrics, engle_semiparametric_1986 modeled the electricity sales as the combination of a smooth function of temperature and a linear function of price and income; in materials science, green_semi-parametric_1985 used a semi-parametric generalized linear model to analyze the bioassay data for the study of flame retardants; in biology, liang_empirical_2009 applied generalized partially linear models to investigate the relationship between viral load and CD4$^+$ cell counts to understand AIDS pathogenesis. See other applications in hardle_partially_2000.
In this paper, we consider a partially linear time series regression model
where $(Z_i,X_i,Y_i)$ are observed stationary processes with $Z_i\in\mathbb R^{l}$, $X_i\in\mathbb R^d$ and $Y_i\in\mathbb R$, for $l,d\ge1$. Here $\boldsymbol \beta\in\mathbb R^{l}$ is a fixed vector of unknown parameters and $\mu(\cdot)$ [resp. $\sigma^2(\cdot)$] is an unknown smooth regression function (resp. conditional variance or volatility function) from $\mathbb R^d$ to $\mathbb R$. In addition, $\epsilon_i\in\mathbb R$ is an unobserved random error with mean zero, independent of the covariates $X_i$ and $Z_i$. In this work, we'll primarily concentrate on the conditional volatility $\sigma(X_i)$ for clarity's sake although our findings can be extended to $\sigma(X_i,Z_i)$. Compared to completely parametric or nonparametric specifications, a partially linear model in ((ref)) enjoys a flexible semi-parametric structure. The parametric components can provide easier interpretations of each variable to better characterize the underlying data-generating mechanism, while the additional nonparametric part allows a data-driven approximation with no specific structures imposed on the true regression function, which can avoid inconsistent estimators and faulty inferences due to model mis-specification in purely parametric models.
In the past decades, much attention has been directed to estimating and testing partially linear models. See, for instance, engle_semiparametric_1986,rice_convergence_1986,Robinson,speckman_kernel_1988,schick_root-n_1996 on the $\sqrt{n}$-consistent estimators of $\boldsymbol \beta$; gao_convergence_1995,fan_profile_2005,xie_scad-penalized_2009 on the inferences of $\boldsymbol \beta$. It is crucial to include the parametric component in equation ((ref)) because the parametric component is derived out of relevant theories and it is practically useful to identify the parametric component. One notable example is the Phillips curve with a time-varying natural unemployment rate (KHK:2014), where the parameter $\boldsymbol \beta$ captures the effect of the unemployment rate on the price inflation. Moreover, $\boldsymbol \beta$ can be employed to measure the impact of various demographic variables, such as age, gender, family size, and the residency type, on the household gasoline consumption, as demonstrated in the U.S. case study by kim_simultaneous_2021. In fact, this paper significantly extends the scope of both KHK:2014 and kim_simultaneous_2021 by introducing a multivariate, time-dependent, and stochastic nonlinear component into ((ref)). Furthermore, the parameter $\boldsymbol \beta$ in equation ((ref)) can represent the factor that determines the efficiency of the foreign currency market, as discussed in Section (ref). Considering these diverse roles played by the parametric component, it is essential and useful to include the linear parametric part for the estimation and the statistical inference of model ((ref)), instead of relying solely on the nonparametric part. It is worth noting that the purely nonparametric model corresponds to a special case of model ((ref)) where $\boldsymbol \beta$ equals zero.
The estimation of the nonparametric part has also been intensively studied, including the methods based on kernel, local linear and spline smoothers (hamilton_local_1997,yu_penalized_2002,fan_kernel-based_2003,aneiros-perez_local_2008). Several attempts have been made to the inference of $\mu$ in partially linear models, such as the consistency and asymptotic normality for the estimator of $\mu$ by liang_asymptotic_1997, the point-wise confidence intervals of $\mu$ based on empirical likelihood by liang_empirical_2009, and simultaneous confidence bands of multivariate function $\mu$ by KHK:2014,kim_simultaneous_2021. However, all the aforementioned literature focused on the {\it independent} or non-stochastic observations. No previous work investigated the simultaneous inference of $\mu$ in a partially linear time series model with {\it dependence} as ((ref)) that is commonly encountered in real data (hardle_partially_2000). Further, most of the studies on partially linear models assume the error terms in model ((ref)) to be homoskedastic, where the error term $\sigma(X_i)\epsilon_i$ is simply reduced to $\epsilon_i$ and the conditional variance is constant over time. This can be a shortcoming since it would rule out most macroeconomic and financial time series, at least for the application to asset pricing (nelson_conditional_1991), where returns may be uncorrelated but feature stochastic volatility. The current paper aims to fill in these gaps by providing theory for the simultaneous inference of the mean trend $\mu$ in ((ref)) under a general dependency structure, which also allows for conditional heteroskedasticity.
Specifically, we allow both the errors $\epsilon_i$ and the covariates $X_i$ in ((ref)) to be dependent over $i$. Let $X_i=(X_{i1},X_{i2},\ldots,X_{id})^{\top}$ be a stationary process of the form
where $v_i$ are independent and identically distributed (i.i.d.) random vectors in $\mathbb R^{d'}$ for $d'\ge1$ and $H=(H_1,H_2,\ldots,H_d)^{\top}$ is a measurable function such that $X_i$ is a well-defined. The nonlinear Wold representation ((ref)) allows a very general class of stationary processes, including linear processes such as vector autoregressive models (VAR) and autoregressive moving average (ARMA) models and nonlinear transforms such as bilinear models, Volterra processes, Markov chain models, threshold/exponential autoregressive models (TAR/EAR) and (generalized) autoregressive conditionally heteroscedastic (ARCH/GARCH) type models, etc. Within this framework, $v_i$ can be viewed as independent inputs of a physical system, and all the dependencies among the outputs $X_i$ result from the underlying data-generating mechanism $H(\cdot)$.
For the identification of model ((ref)), we assume that the error $\epsilon_i$ is independent of both covariates $X_i$ and $Z_i$. In particular, we shall proceed with the conditional expectation $\mathbb E(Y_i\mid X_i,Z_i)=Z_i^{\top}\boldsymbol \beta + \mu(X_i)$. Following the fixed design case in hardle_partially_2000, we let $Z_i$ be a function of $X_i$ plus another noise term (cf. Assumption (ref)). This means $\mathbb E(Y_i\mid X_i,Z_i)$ cannot be reduced to $\mathbb E(Y_i\mid X_i)$ and indicates that it is nontrivial and also challenging to extend the inference of $\mu(\cdot)$ in a purely nonparamteric model to that under a partially linear setting. If $Z_i=0$, then model ((ref)) is simply a nonparametric regression process as a special case, that is,
When $X_i=Y_{i-1}$ and $\epsilon_i$ are i.i.d. random noises, model ((ref)) incorporates many interesting linear and nonlinear autoregressive processes (AR), such as AR processes if $\mu(x)=ax$ for some real parameter $a$, and autoregressive conditional heteroscedastic (ARCH) processes if $\mu(x)=0$ and $\sigma^2(x)=\alpha_0+\alpha_1x^2$ for some non-negative real parameters $\alpha_0,\,\alpha_1\in\mathbb R$.
Many contributions have been made to constructing the SCR of $\mu(\cdot)$ in completely nonparametric models. For example, concerning independent data, johnston_probabilities_1982 was among the first to investigate the inferences of univariate mean regression functions; hardle_asymptotic_1989 derived simultaneous confidence bands for one-dimensional kernel M-estimators; HS:2010,GH:2012 constructed uniform confidence bands for conditional quantile and expectile functions, respectively. With regard to dependent cases, see inference of trends in a fixed design with $X_i=i/n$ by WZ:2007; confidence bands for the mean function in functional time series with physical dependence by CS2015; nonlinear regression model with nonstationary regressors by Li2017 and a time-varying nonlinear regression model by ZW2015. In particular, ZW:2008,LW:2010 proposed inferences of the univariate mean and volatility functions in a similar time series regression model in ((ref)), where they assumed that the error terms $\epsilon_i$ are i.i.d.. Our work can be viewed as a generalization of their results by extending the dependence structure of $\epsilon_i$ and by including an additional parametric part to accommodate a broader class of data-generating mechanisms. That is, our work is distinct from ZW:2008,LW:2010 in that our model framework is semi-parametric with the multivariate covariate $X_i$ and time-dependent $\epsilon_i$, while the framework in ZW:2008,LW:2010 is purely nonparametric with an univariate $X_i$ and an i.i.d. noise $\epsilon_i$.
Contributions: Here we summarize our three main contributions to the literature: Firstly, we extend the dependence structure of the error terms to a more general case by allowing $\epsilon_i$ to be {\it dependent over} $i$, while also accounting for the dependence among the covariates $X_i$ and the conditional heteroscedasticity. Secondly, different from the relevant studies relying on the Gumbel convergence to achieve the asymptotics of the statistics (zhao_kernel_2006,LW:2010), we provide a new testing methodology based on the multiplier bootstrap enlightened by chernozhukov_central_2017. This allows one to {\it avoid the notoriously slow convergence issue} associated with the Gumbel distribution. Thirdly, there is no previous work performing simultaneous inference of the {\it multivariate} $\mu(\cdot)$ in ((ref)), allowing the multivariate covariate $X_i\in\mathbb{R}^d$ with $d\geq 2$ under some general {\it time dependence} setting. This paper could be a complement to the non-parametric model validation problem with a general dependence structure and conditional heteroscedasticity, applicable in various multivariate scenarios.
Notation: For a vector $v=(v_1,...,v_d)\in\mathbb R^d$ and $q>0$, we denote $|v|_q=(\sum_{i=1}^d|v_i|^q)^{1/q}$ and $|v|_{\infty}=\max_{1\le i\le d}|v_i|$. For $s>0$ and a random vector $X$, we say $X\in\mathcal L^s$ if $\lVert X\rVert_s=[\mathbb E(|X|_2^s)]^{1/s}<\infty$. For two positive number sequences $(a_n)$ and $(b_n)$, we say $a_n=O(b_n)$ or $a_n\lesssim b_n$ (resp. $a_n\asymp b_n$) if there exists $C>0$ such that $a_n/b_n\le C$ (resp. $1/C\le a_n/b_n\le C$) for all large $n$, and say $a_n=o(b_n)$ if $a_n/b_n\rightarrow0$ as $n\rightarrow\infty$. We set $(X_n)$ and $(Y_n)$ to be two sequences of random variables. Write $X_n=O_{\mathbb P}(Y_n)$ if for $\forall \epsilon>0$, there exists $C>0$ such that $\mathbb P(|X_n/Y_n|\le C)>1-\epsilon$ for all large $n$, and say $X_n=o_{\mathbb P}(Y_n)$ if $X_n/Y_n\rightarrow 0$ in probability as $n\rightarrow\infty$. We denote the centered random variable $X$ by $\mathbb E_0(X)$, that is, $\mathbb E_0(X)=X-\mathbb E(X)$.
Roadmap: The rest of the paper is structured as follows. Section (ref) introduces the overall methodology to perform simultaneous inference of $\mu(\cdot)$ in ((ref)) with $d\geq 1$. The asymptotic properties of the proposed statistics and estimators as well as the implementation are provided in Sections (ref) and (ref). Section (ref) is devoted to a simulation study to evaluate the performance of our methods and Section (ref) offers an empirical application, the forward premium anomaly, to demonstrate the validity of the proposed methodology in practice. Section (ref) concludes the paper and discusses potential extensions for future research. The technical proofs are deferred to the Supplementary Materials.
In this section, we first introduce the definition of the simultaneous confidence region (SCR). Then, we shall follow with the estimator of the nonparametric trend function $\mu(\cdot)$ in ((ref)). Further, we illustrate our new methodology on constructing the SCR of $\mu(\cdot)$ based on this estimated $\mu(\cdot)$. The theoretical intuition of the proposed simultaneous inference is also provided.
To conduct simultaneous inference of the trend $\mu(\cdot)$ in model ((ref)), we shall construct the nonparametric simultaneous confidence region (SCR) for $\mu(\cdot)$. In particular, we consider deriving asymptotic SCR for $\mu(\cdot)$ over the region $\mathcal T_d=[T_{11},T_{12}]\times [T_{21},T_{22}]\times \cdots \times [T_{d1},T_{d2}] \subset \mathbb R^d$ with confidence level $100(1-\alpha)\%$, $\alpha\in(0,1)$. To this end, we shall find two functions $l_n(\cdot)$ and $r_n(\cdot)$ based on the observations $(Z_i,X_i,Y_i)$, $1\le i\le n$, such that
Given the SCR for $\mu(\cdot)$, we can verify whether $\mu(\cdot)$ is of some certain parametric form by testing the null hypothesis
against the alternative $\mathcal H_{\mathcal A}:\, \mu(\cdot)\neq\mu_{\theta}(\cdot)$, where $\theta\in\Theta$ for some parametric space $\Theta$ and $\mu_{\theta}(\cdot)$ is a multivariate parametric function. Specifically, one can test ((ref)) by checking whether the condition $l_n(x)\le\mu_\theta(x)\le r_n(x)$ holds for all $x\in\mathcal T_d$. If this condition does not hold for some $x\in\mathcal T_d$, then we reject the null hypothesis at level $\alpha$. The SCR-based inference is more preferred to other standard inferential procedures utilizing mean-integrated-squared-error (MISE) type statistics, for example, since it is more effective in suggesting the right function form of $\mu(\cdot)$ in ((ref)). When the null hypothesis in ((ref)) is rejected via an MISE-type test statistic, it would be rather difficult to figure out the reason for rejection, which, however, can be easily dealt with under our approach by locating graphically where the SCR is violated by $\mu_{\theta}(\cdot)$ under the null hypothesis.
Next, we provide an estimator for $\mu(\cdot)$ given the observed sample $(Z_i,X_i,Y_i)$, $1\le i\le n$. Let $x=(x_1,\ldots,x_d)^{\top}\in\mathbb R^d$ and set $K(\cdot)\ge0$ to be some kernel function with support $[-1,1]^d$. We consider the following optimization problem:
where $K_h(\cdot)=K(\cdot/h)/h^d$, $h$ is a bandwidth parameter with $h\rightarrow0$ and $h^dn\rightarrow\infty$, and $\hat\boldsymbol \beta$ is a consistent estimator of the unknown parameters $\boldsymbol \beta$ in ((ref)). We shall defer the details of $\hat\boldsymbol \beta$ to Section (ref). Here in ((ref)), we adopt the local constant estimator for the simplicity of notation. One can achieve similar results by applying the local linear estimator introduced in fan_local_1996. By solving ((ref)), we can obtain the Nadaraya-Watson estimator for $\mu(\cdot)$ which has the expression
where the weight function $w_h(x,X_i)$ is defined as
Moreover, we denote the consistent estimator of the volatility function $\sigma(\cdot)$ in ((ref)) by $\hat\sigma(\cdot)$, and we shall provide the detailed definition and consistency results of $\hat\sigma(\cdot)$ in Section (ref). Given $\hat\boldsymbol \beta$ and $\hat\sigma(\cdot)$, we consider the statistic $\sup_{x\in\mathcal T_d}\big|\hat\mu^*(x)- \mu(x)\big|/\hat\sigma(x)$ to construct the SCR of $\mu(x)$. We shall note that, due to the smoothness of $\mu(\cdot)$ and the consistency of $\hat\boldsymbol \beta$, this statistic can be approximately written into the supremum of a sum of dependent random fields conditioned on the covariates $X_i$, that is
It is non-trivial to investigate the asymptotic properties of ((ref)) when $X_i$ and $\epsilon_i$ are dependent over $i$. zhao_kernel_2006,LW:2010 have dealt with the case where the covariates $X_i$ are dependent while the errors $\epsilon_i$ are independent by establishing the Gumbel convergence. To address the more general dependency structure in our study, we propose an extension of the high-dimensional Gaussian approximation theorem introduced by chernozhukov_central_2017 to dependent processes with continuous index sets. By this generalized high-dimensional Gaussian approximation, we shall expect the limit distribution of our proposed statistic to be approximated by the one of the maximum of a centered Gaussian random vector $\hat\mathcal Z=(\hat \mathcal Z_1,\ldots,\hat\mathcal Z_n)^{\top}\in\mathbb R^n$, that is
We defer the detailed definition of the covariance matrix for $\hat\mathcal Z$ to ((ref)). Intuitively, the result in ((ref)) would enable us to find the critical value of our proposed statistic, and consequently facilitates the construction of the simultaneous confidence region for $\mu(\cdot)$. Specifically, we can approximate $l_n(x)$ and $r_n(x)$ in ((ref)) by the estimators
respectively, where $\hat q_{\alpha}$ is the $(1-\alpha)$-th empirical quantile of $\max_{1\le j\le n}|\hat\mathcal Z_j|/\sqrt{h^dn}$ given the significance level $\alpha\in(0,1)$, and it can be evaluated by the multiplier bootstrap (chernozhukov_central_2017). Based on the SCR in ((ref)), one can test whether the trend function $\mu(\cdot)$ is of any particular parametric form $\mu_\theta(\cdot)$, such as quadratic or cubic patterns, by evaluating whether $l_n(x)\le\mu_\theta(x)\le r_n(x)$ is satisfied for all $x\in\mathcal T_d$. If the SCR fails to {\it entirely} contain this parametric form, then we reject the null hypothesis ((ref)) at level $\alpha$. We shall provide the detailed steps for implementing the SCR construction at the end of Section (ref) after we introduce our main theorems.
This section is devoted to our main results on the asymptotic properties for the statistic $\sup_{x\in\mathcal T_d}\big|\hat\mu^*(x)- \mu(x)\big|/\hat\sigma(x)$, which provides the theoretical foundation for the construction of SCR. In Section (ref), we shall first establish the asymptotic distribution of the proposed statistic under the oracle setting, that is, assuming that the unknown parameters $\boldsymbol \beta$ and $\sigma(\cdot)$ in the statistic are the true ones. The case with $\boldsymbol \beta$ and $\sigma(\cdot)$ replaced by their consistent estimators $\hat\boldsymbol \beta$ and $\hat\sigma(\cdot)$, respectively, are dealt with in Section (ref), where we provide the consistency results for $\hat\boldsymbol \beta$ and $\hat\sigma(\cdot)$ and show the similar asymptotic distribution of the statistic.
We shall start with some regularity conditions which will be useful to establish our main theorems. First, we impose the smoothness condition on the trend function $\mu(\cdot)$ and assume that the volatility function $\sigma(\cdot)$ also varies smoothly and is bounded on the support $\mathcal T_d$.
Next, we shall specify the dependence structure of the error process $\{\epsilon_i\}_{1\le i\le n}$ and the covariates $X_1,\ldots,X_n$ in model ((ref)). In particular, throughout this paper, we assume that $\{\epsilon_i\}_{1\le i\le n}$ is MA($\infty$), which can be formulated as follows:
where $\eta_i\in\mathbb R$ are i.i.d. random variables with mean zero and unit variance, independent of $(X_i,Z_i)$ in ((ref)). The coefficients $a_k$, $k\ge0$, take values in $\mathbb R$ such that $\epsilon_i$ is a proper random variable. We assume that the innovations $\eta_i$ have finite $q$-th moment, for some $q\ge 4$. We define the absolute sum of the coefficients as $S=\big|\sum_{k\ge0}a_k\big|$. Then, the long-run variance of $\{\epsilon_i\}_{1\le i\le n}$ is
where $\gamma(k)=\mathbb E(\epsilon_i\epsilon_{i+k})$ is the autocovariance of innovations with lag $k\in\mathbb Z$.
Assumptions (ref) and (ref) post conditions on the moment and dependency structure of errors $\{\epsilon_i\}_{1\le i\le n}$. Specifically, the moment condition in Assumptions (ref) depends on $q$ which characterizes the heavy-tailedness of the noise, and a larger $q$ means a thinner tail. Assumption (ref) requires that the dependency strength of $\{\epsilon_i\}_i$ decays at a polynomial rate, which also ensures that the long-run variance of $\{\epsilon_i\}_i$ is finite.
For the covariates $X_1,\ldots,X_n$, we denote the density function of $X_{i}=(X_{i1},\ldots, X_{id})^{\top}$ by $g(x_1,\ldots, x_d)$. For $s=(s_1,s_2,\ldots,s_d)^{\top} \in \mathbb{R}^d$, we define
where the filtration $\mathcal{F}_i = (\ldots ,v_{i-1},v_{i})$, $i\in \mathbb Z$. We let $v_i'$ be an i.i.d. copy of $v_i$ and $\mathcal{F}_{i,\{k\}}$ be $\mathcal{F}_{i}$ with $v_k$ therein replaced by $v_k'$. We define $X_{ij,\{i-k\}}$ as a coupled version of $X_{ij}$ with the form $$X_{ij,\{i-k\}}=H_j(\ldots ,v_{i-k-1},v'_{i-k},v_{i-k+1},\ldots ,v_i).$$ In addition, we denote $X_{i,\{i-k\}}=(X_{i1,\{i-k\}},\ldots ,X_{id,\{i-k\}})^\top$. We impose assumptions on the moments and dependency structures of covariates $X_1,\ldots,X_n$ as follows.
Roughly speaking, the functional dependence measure $\theta_{k,s}$ defined in Assumption (ref) quantifies the dependence of $X_i$ on $v_{i-k}$ by measuring the distance between $g(x\mid\mathcal F_{i-1})$ and its coupled version $g(x\mid\mathcal F_{i-1,\{i-k\}})$. The physical dependence measure accounts for any measurable function of cumulative i.i.d. noises, which makes it applicable to a wide range of linear and nonlinear time series processes. To elaborate such conditional dependency structure, we shall consider a simple moving average (MA) example as follows.
Assumption (ref) imposes conditions on the boundary and smoothness of the density function $g(x)$ on the support $\mathcal T_d$. In the literature, conditions similar to Assumption (ref) have been commonly used; see, for instance, aneiros-perez_local_2008,kim_simultaneous_2021.
This subsection is devoted to our main results, the asymptotic distribution of the proposed statistic derived by Gaussian approximation. To explicitly illustrate our theory on the simultaneous inference of trend $\mu(\cdot)$, we shall first assume that $\boldsymbol \beta$ and $\sigma(\cdot)$ in model ((ref)) are both known. In particular, we define
and consider the statistic $\sup_{x\in\mathcal T_d}\big|\hat\mu(x)- \mu(x)\big|/\sigma(x)$. By Slutsky's theorem, the asymptotic distribution of the statistic still holds when replacing the true $\boldsymbol \beta$ and $\sigma(\cdot)$ therein by their consistent estimators $\hat\boldsymbol \beta$ and $\hat\sigma(\cdot)$, respectively. Hence, in this subsection, we shall first proceed under the oracle setting, and we refer to the implementation with theoretical guarantees to Section (ref).
As mentioned in Section (ref), we aim to provide the limiting distribution for the statistic $\sup_{x\in\mathcal T_d}\big|\hat\mu(x)- \mu(x)\big|/\sigma(x)$ by applying the high-dimensional Gaussian approximation. Specifically, we shall generate a centered Gaussian random field $\{\mathcal Z_t\}_{t\in\mathcal T_d}$ and utilize the distribution of $\sup_{t\in\mathcal T_d}|\mathcal Z_t|$ for the approximation. We denote the covariance matrix of $\sqrt{h^dn}\big|\hat\mu(x)- \mu(x)\big|/\sigma(x)$ conditioned on the covariates $X_i$, $1\le i\le n$, by $Q=(Q_{t,s})_{t,s\in\mathcal T_d}$, where $Q_{t,s}$ takes the form
where $c_{t,s,i,k} = \sigma(x_t)^{-1}\sigma(x_s)^{-1}\sigma(X_i)\sigma(X_{i+k})$. Further, in practical usage, to evaluate the covariance matrix $Q$, we would need to estimate the autocovariance $\gamma(k)$, for each $k\ge0$, while this is unrealistic when $k$ goes to infinity. Therefore, to establish a consistent estimator for $Q$, we define a truncated version of $Q$ as $Q^{(L)}=(Q_{t,s}^{(L)})_{t,s\in\mathcal T_d}$, for some large positive integer $L$, where $Q_{t,s}^{(L)}$ is defined as
When applied to real data, the autocovariance $\gamma(k)$ can be estimated by the notable existing methods and we defer the details of this estimation to Section (ref). Now we let $\{\mathcal Z_t\}_{t\in\mathcal T_d}$ be a Gaussian random field with mean zero and conditional covariance matrix $Q^{(L)}$ given the covariates $X_1,X_2,\ldots,X_n$. The first main theorem is stated as follows.
As shown by Theorem (ref), our approach differs from the one based on the Gumbel convergence in extreme value theory adopted by zhao_kernel_2006,LW:2010. We extend the high-dimensional Gaussian approximation in chernozhukov_central_2017 to dependent processes and provide a Gaussian limit distribution for $\sup_{x\in\mathcal T_d}|\hat\mu(x)-\mu(x)|/\sigma(x)$. Our method does not need the density estimate of covariate $X_i$ to build the test statistic, and thus offers easier implementation (see details in Section (ref)) than the competing ones.
In the previous sections, we assumed that $\boldsymbol \beta$, $\sigma(\cdot)$ in model ((ref)) and the conditional long-run covariance matrix of $\sup_{t\in\mathcal T_d}\big|\hat\mu(x)- \mu(x)\big|/\sigma(x)$ are all known, which, however, is not realistic in practice. Hence, we shall introduce the estimators $\hat\boldsymbol \beta$, $\hat\sigma(\cdot)$ and $\hat Q^{(L)}$ in this subsection, followed by the consistency results and the limit distribution of the statistic built on these estimates. The detailed instructions on how to construct SCR in practice are also provided in Section (ref).
The literature on estimating the parametric components in model ((ref)) has a long history. Robinson was among the first contributing to this problem, where he derived a least square estimator of $\boldsymbol \beta$ based on a Nadaraya-Waston kernel estimator of $\mu$ and provided the $\sqrt{n}$-consistency result of the estimate. In this study, we shall consider the same estimator for $\boldsymbol \beta$, and we will show that the $\sqrt{n}$-consistent rate still holds under our setting. Define the Robinson's estimator $\hat\boldsymbol \beta$ as follows,
where $\tilde Y_i$ and $\tilde Z_i$ are the kernel estimators of $Y_i$ and $Z_i$, respectively, which are
To establish the consistency result of the Robinson's estimator $\hat\boldsymbol \beta$, we shall introduce some regularity conditions on the covariates $Z_i$ of the parametric part in ((ref)). Specifically, following hardle_asymptotic_1989, we consider $Z_i$ as random design points and related to $X_i$ in the following way.
Assumption (ref) can be considered as a generalization of the conditions with weaker requirements compared to the ones introduced in hardle_asymptotic_1989, gao_convergence_1995 and Sun, where they assumed $u_i$ to be i.i.d. random vectors, while in this paper, we allow $u_i$ to be weakly dependent over $i$. In fact, Assumption (ref)(iii) implies the short-range dependence (SRD) of $u_i$. The following proposition asserts that, under such dependence conditions, we can consistently estimate $\boldsymbol \beta$ in the time series regression model ((ref)) by the Robinson's estimator at the rate of $\sqrt{n}$.
Concerning the conditional volatility function $\sigma(X_i)$ in ((ref)), for each $1\le i\le n$, we define the local constant estimator
Note that the bandwidth parameter $h$ in ((ref)) can be selected differently from the one used in the estimator for $\mu(\cdot)$. Similar methods have been applied in the existing studies; see, for example, fan_local_1996,zhao_kernel_2006. In this paper, we adopt the same bandwidth parameters in both $\hat\mu^*(\cdot)$ and $\hat\sigma(\cdot)$ for brevity. The consistency result of $\hat\sigma(\cdot)$ is stated as follows.
Recall that in Theorem (ref), $\mathcal Z_t$, $t\in\mathcal T_d$, is a Gaussian random field with mean zero and conditional covariance matrix $Q^{(L)}$ as defined in ((ref)). In practice, one shall estimate $Q^{(L)}$ by $\hat Q^{(L)}=(\hat Q_{j,j'}^{(L)})_{1\le j,j'\le n}$, that is
where $\hat c_{j,j',i,k} = \hat\sigma^{-1}(X_j)\hat\sigma^{-1}(X_{j'})\hat\sigma(X_i)\hat\sigma(X_{i+k})$, and $\hat\gamma(k)$ is a consistent estimate of $\gamma(k)$ which takes the form
with $\hat\epsilon_i=(Y_i-Z_i^{\top}\hat\boldsymbol \beta-\hat\mu^*(X_i))/\hat\sigma(X_i)$. The precision of the estimated autocovariance $\hat\gamma(k)$ has been well investigated in the literature; see, for example, shumway. Also, note that we adopt the estimator $\hat Q_{j,j'}^{(L)}$ which includes the term $\hat c_{j,j',i,k}$ as the estimator of $c_{j,j',i,k}$ appeared in the true $Q_{t,s}^{(L)}$ defined in ((ref)). Since $\hat\sigma(\cdot)$ is a consistent estimator of $\sigma(\cdot)$ by Proposition (ref) and $\sigma(\cdot)$ is bounded on the support $\mathcal T_d$ according to Assumption (ref), it can be shown that $\hat Q_{j,j'}^{(L)}$ is a consistent estimator of $Q_{j,j'}^{(L)}$. We state this result and provide the exact consistency rate as follows.
In real-data applications, one can simply take $L=\sqrt{n}$ to obtain a consistent estimator $\hat Q^{(L)} =(\hat Q_{j_1,j_2}^{(L)})_{1\le j_1,j_2\le n}$ to estimate the covariance matrix of the statistic $\sqrt{h^dn}|\hat\mu^*(x)-\mu(x)|/\hat\sigma(x)$ conditioned on the covariates $X_i$, $1\le i\le n$. In next section, we shall provide the Gaussian approximation for the test statistic $\sup_{x\in\mathcal T_d}|\hat\mu^*(x)-\mu(x)|/\hat\sigma(x)$ with the estimated covariance matrix applied.
Provided the consistent estimated long-run covariance matrix $\hat Q^{(L)}$ in Proposition (ref), we now introduce a sequence of Gaussian random variables $\hat\mathcal Z_j$, $1\le j\le n$, with mean zero and the conditional covariance matrix $\hat Q^{(L)}$, i.e., $\hat Q_{j,j'}^{(L)} = \mathbb E(\hat\mathcal Z_j\hat\mathcal Z_{j'})$, for each $1\le j,j'\le n$. We shall show that the Gaussian approximation result in Theorem (ref) still holds when we replace $Q_{t,s}^{(L)}$ therein by its estimator $\hat Q_{j,j'}^{(L)}$. In addition, Propositions (ref) and (ref) ensure that even when $\boldsymbol \beta$ and $\sigma(x)$ in model ((ref)) are both unknown, one shall still achieve the similar asymptotic distribution in Theorem (ref) for the test statistic $\sup_{x\in\mathcal T_d}\big|\hat\mu^*(x)- \mu(x)\big|/\hat\sigma(x)$ by employing the proposed estimators $\hat\boldsymbol \beta$ and $\hat\sigma(x)$. As a result, combining Propositions (ref)--(ref) facilitates the construction of SCR for practical use. For the theoretical guarantee, we show the Gaussian approximation result with $\hat\boldsymbol \beta$, $\hat\sigma(x)$ and $\hat Q^{(L)}$ plugged in as follows.
Recall the upper and lower bounds of the simultaneous confidence region, i.e., $\hat l_n(x)$ and $\hat r_n(x)$ defined in ((ref)). For the convenience of implementation, we provide the detailed steps of our proposed method for the construction of SCR as follows. Suppose that we have observed a sample $(Z_i,X_i,Y_i)$, $1\le i\le n$. Given the significance level $\alpha\in(0,1)$, we aim to construct the $(1-\alpha)$-th percentile SCR for $\mu(\cdot)$ in model ((ref)).
We present the simulation study with four different time series models in this section to illustrate the performance of our proposed SCR. Consider the following data-generating process:
where $\beta=0.5$, $\mu(x)=0.3+0.4\,x$ and $\sigma(x)=\big(0.1+0.1\,x^2\big)^{1/2}$. We shall compute the {\it coverage probability} of the SCR for $\mu(\cdot)$ in ((ref)) to evaluate our method. We set the covariate terms $x_i$ and $z_i$ to be
where $\delta_i$ and $u_i$ are both i.i.d. random variables following the standard normal distribution, and $\delta_i$ are independent of $u_i$. Note that this setting is in accordance with Assumption (ref). For the specifications of $\epsilon_i$ in ((ref)), we consider four commonly used models: (i) i.i.d. standard normal, (ii) an autoregressive (AR) process, (iii) a moving-average (MA) process and (iv) an autoregressive moving-average (ARMA) process, which take the forms of:
where $\eta_i$ are i.i.d. random variables following the standard normal distribution. For a range of bandwidths (see Table (ref)), we follow Steps 1--6, the procedures described in the previous section, to construct the 95% SCR of $\mu(\cdot)$ in ((ref)). The process is repeated $M$ times for each chosen bandwidth parameter. Given these $M$ different SCRs, we count how many of those SCRs contain the true $\mu(\cdot)$ in them, which gives us the desired coverage probabilities. Here we let $M=1,000$ (i.e. number of iterations) and $n=200$ (i.e. sample size). The simulation results under (i)--(iv) are summarized by Table (ref).
From Table (ref), we see that the coverage probabilities under various bandwidths are reasonably close to $0.95$, which is the nominal coverage rate at a 5 percent level. For some bandwidths, we have the coverage probabilities that are almost identical to the nominal level $95\%$. However, as the bandwidth gets extreme, the coverage probability deviates from the nominal $0.95$, as we can observe from Table (ref).
One of the useful applications for ((ref)) in time series analysis is the {\it forward premium anomaly}. Consider the celebrated forward premium regression (Fm:1984):
where $s_{t}$ is log of monthly {\it spot} exchange rate at time $t$ and $f_{1,t}$ is log of monthly {\it forward} exchange rate with one-month maturity at time $t$, respectively. Here $u_t$ is a model error. A vast amount of literature illustrates that the constant term $\mu$ in ((ref)), the foreign exchange risk premium, is related to {\it fundamentals} (DH:1985,HS:1986,Hod:1989,Mark:1995,BK:2006,Alv:2009,LRV:2014,BK:2015,KKK:2022), such as money growth, output growth, interest rates, conditional variance of money growth and so on. Given this, one can possibly model ((ref)) in the following flexible manner:
where
Here, $m_t$ and $y_t$ are log of domestic money stock and log of domestic production, respectively, and $m^*_t$ and $y^*_t$ are their foreign counterparts. Note here that $\mu$ in ((ref)) is replaced by $\mu(x_{1,t},\,x_{2,t})$, and that the conditional heteroscedasticity is introduced in ((ref)). Interestingly, Assumption (ref) for ((ref)) implies, quite intuitively, that the forward premium $f_{1,t}-s_t$ does depend on the fundamentals $x_{1,t}$ and $x_{2,t}$.
Although the theory of Uncovered Interest Parity (UIP) in international finance implies that $\mu(\cdot)=0$, numerous empirical studies actually show $\mu(\cdot)\not=0$. We refer to FT:1990 and Eg:1996 for an excellent review of the literature on this. Recently, BDKK:2023 revisited the issue and tested the UIP hypothesis to show the effect of modeling lagged spot returns and lagged forward premiums in ((ref)). Interestingly, ((ref)) is a specification of our partially linear model framework in ((ref)). Hence, the methodology developed in this study can be readily applied to decide whether or not the UIP condition holds for ((ref)).
The exchange rate data used in this study are monthly spots and 30-day forward exchange rate data for Australian Dollar (AUD), British Pound (GBP) and U.S. Dollar (USD), where USD is the numeraire currency. That is, we employ the AUD/USD and GBP/USD currency pairs in this empirical study. For the money and the output variables, we employ monthly M1 and monthly industrial production index home (i.e. U.S.) and abroad (i.e. Australia and U.K.). Note that the usual GDP series cannot be used here because the sampling frequency is monthly in this application. End-of-month observations from December 1988 through December 2016, including the 2008 financial crisis, are used in this study.
The estimation and inference results, including the simultaneous confidence region, are reported by Figures (ref)--(ref). First of all, the semi-parametric Robinson estimate of the UIP coefficient $\beta$ in ((ref)) and its standard error are $\hat\beta_R=0.362$ and $s.e.(\hat\beta_R)=0.753$ for AUD/USD, and $\hat\beta_R=0.802$ and $s.e.(\hat\beta_R)=0.776$ for GBP/USD. In contrast, the traditional OLS estimates of $\beta$ in ((ref)) are $\hat\beta_{OLS}=-0.029$ and $s.e.(\hat\beta_{OLS})=0.718$ for AUD/USD, and $\hat\beta_{OLS}=0.689$ and $s.e.(\hat\beta_{OLS})=0.641$ for GBP/USD. Interestingly, our semi-parametric estimates of the UIP coefficient for both AUD/USD and GBP/USD are closer to one, the value implied by UIP, with lower $t$-statistics than their OLS counterparts.
Figure (ref) shows the estimates of $\hat{\mu}(\cdot,\cdot)$ and $\sigma(\cdot,\cdot)$ (i.e. the nonlinear surfaces) for AUD/USD, while Figure (ref) shows the mean and volatility estimates for GBP/USD. As we can see from Figures (ref) and (ref), the estimates of both the mean and the volatility functions for each currency pair are {\it highly nonlinear}, which implies that ((ref)) may not be an appropriate parametric form for the underlying processes.
Figure (ref) represents the two-dimensional relationship between $\mu(\cdot,\cdot)$ and the relative output growth for AUD/USD when the relative money growth is fixed at a certain percentile. Similarly, Figure (ref) represents the relationship between $\mu(\cdot,\cdot)$ and the relative money growth for AUD/USD when the relative output growth is fixed at a certain percentile. Each panel in Figures (ref) and (ref) also includes the 95% SCR of the corresponding function estimate. The horizontal line represents the null hypothesis of $\mu(\cdot,\cdot)=0$.
In order to not reject the null hypothesis of zero risk premium at a 95% confidence level, the horizontal line fixed at zero must be {\it entirely} contained by $all$ SCRs in Figures (ref) and (ref). Clearly, Panels (a)--(c) in Figure (ref) and Panels (a)--(c) in Figure (ref) fail to contain the horizontal line entirely. That is, all of the panels in Figures (ref) and (ref) fail to contain the horizontal line entirely. Hence, the null hypothesis gets rejected at a 5% level, which is the key implication of Figures (ref) and (ref). Even a more general null hypothesis of {\it constant} risk premium is also rejected clearly at a 5 percent level because the SCRs in Figure (ref) cannot entirely contain any horizontal line in them.
The main reason for the rejection is because of the tendency that the local-linear estimate of $\mu(\cdot,\cdot)$ changes $nonlinearly$ in both the output growth and the money growth. Similarly, Figures (ref) and (ref) show the results corresponding to the GBP/USD pair. Again, the null hypothesis gets rejected at a 5% level because the 95% SCRs in all panels of Figures (ref) and (ref), except for Panel (b) of Figure (ref), fail to entirely contain the horizontal line fixed at zero.
Interestingly, these findings highlight the relative advantage of our SCR-based inference because, with the usual $p$-value of a MISE-type test statistic (HM:1993), it is rather difficult to figure out why the null gets rejected. Under our approach, however, we can easily locate exactly where the SCR fails to contain the null, so that we can propose an appropriate function for the underlying process more easily.
Intuitively, the trend estimates in this study make sense because either a relative increase in the domestic money supply (i.e. an increase in $x_{1,t}$) or a relative decrease in the domestic production (i.e. a decrease in $x_{2,t}$) will cause the foreign currency to appreciate, yielding a decrease in the risk premium for holding the appreciating currency. Hence the risk premium, $\mu(\cdot,\cdot)$, should increase with a relative increase in the domestic production or a relative decrease in the domestic money supply. This intuition appears to be generally in line with the results in Figures (ref)--(ref).
This paper proposed a new methodology to conduct simultaneous inference of trend in a semi-parametric partially linear time series model and the number of covariate terms in the nonlinear part is allowed to be two or higher. We generalized the high-dimensional Gaussian approximation theory (chernozhukov_central_2017) to construct the simultaneous confidence region (SCR) for the multivariate unknown trend in time series. This work can be viewed as an extension of the two-dimensional uniform confidence band (johnston_probabilities_1982,KHK:2014) and a generalization to the {\it time series} setting compared to kim_simultaneous_2021. The relating asymptotic properties of the introduced methodology are investigated and the finite-sample properties are studied through a simulation experiment. The developed methodology is applied to the forward premium regression (Fm:1984). The empirical analysis based on simultaneous inference confirms that the zero-risk-premium hypothesis for the AUD/USD and GBP/USD currency pairs is rejected at a 5$\%$ level, mainly due to the underlying nonlinear nature, as shown by Figures (ref)--(ref).
The current study can be extended to the case where all the model parameters are functions of random processes. For instance, the UIP coefficient in ((ref)) can be modeled as a function of fundamentals as well. Similarly, both the pricing error and the beta coefficient from a standard factor pricing model in finance can be modeled as functions of relevant state variables. The methodology developed in this study can be easily extended to handle such a framework. Additionally, one can conduct simultaneous inference of the unknown function in ((ref)) when the covariate terms are non-stationary processes. This would significantly generalize the results in this study. Further insight can be gained by extending the current work in these and other directions, and the authors are currently working on these issues.
\printbibliography