Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
108,295 characters · 19 sections · 92 citation commands
Functional Linear Projection and Impulse Response Analysis
After the initial development by Oscar2005, the local projection approach has become one of the foremost applications of time series analysis in the fields of macroeconomics and public policy for studying dynamic causal effects, known as structural impulse responses. The work by PW2021 provides more rigorous grounding for the approach as an important complement to structural vector autoregressive (SVAR) models, by establishing the theoretical equivalence between estimation results of the two. Practitioners have benefited from this result, as the local projection approach is easier to implement and interpret, computationally less burdensome, and less sensitive to model misspecification (see Oscar2005 for a detailed discussion).
This paper contributes to recent developments in this area by providing statistical inference methodologies to study responses of target variables when shocks are characterized by function-valued random variables, such as the functional monetary policy shocks in IR2021. We consider a linear projection of a target variable onto the space spanned by the variables in the conditioning set, including function-valued shocks. Similar to Oscar2005, this setup simplifies impulse response analysis into the estimation of a regression model. Our first result is the formal theoretical evidence for interpreting the parameters of interest in our regression model as structural impulse response functions of SVAR models involving a functional variable, under a certain identification condition. Considering the scarcity of the literature, our paper will be a valuable first step toward formally understanding functional structural shocks and their impact. Furthermore, while the equivalence for this special case has been formally established by Oscar2005 for finite dimensional VAR models, its extension to the case with a functional covariate has not yet been fully explored.
Our benchmark model to be studied is closely related to the functional local projection approach proposed by IR2021, but a crucial distinction exists. Specifically, in contrast to them, we do not require a parametric assumption on the structure of function-valued shocks and their coefficients. In IR2021, functional shocks are represented by a few known functional factors. This assumption crucially reduces the dimensionality of the variables to be analyzed, thereby making existing estimation and inference methodologies valid despite the infinite dimensionality of functional data. However, such a parametric structure is not always available in practice and could result in model misspecification; this point will be further demonstrated in Sections (ref) and (ref) by using an example. Considering practical challenges of exact identification of those factors in infinite dimensional function spaces, our approach could be an appealing alternative for practitioners.
The current paper is closely related to recently growing studies on the VAR model involving a functional covariate, including IR2021, BHCJ2023, and Chang_Chen_Schorfheide_2024. However, our focus on identification and the relationship between SVAR models and linear projection in the presence of a functional covariate distinguishes the current paper from theirs. Moreover, the aforementioned studies typically approximate a functional covariate using a few factors or basis functions and then apply identification strategies developed in the finite dimensional SVAR literature (e.g.,\ kilian2017structural), assuming that the approximation error is either zero or negligible as the sample size increases. Practitioners might prefer these approaches due to their ease of implementation. However, this simplified method does not always ensure the identification of infinite dimensional structural parameters on the entire space in which the parameters take values, thereby restricting our ability to fully comprehend the features of structural parameters. This point will be detailed in Section (ref) by using an example.
We propose an estimator based on the operator Schur complement and establish its asymptotic properties, including (local) asymptotic normality. As our model belongs to the so-called scalar-on-function models studied in, e.g., Hall2007, SHIN2009, Florence2015, imaizumi2018, and Babii2022, our estimator can be understood as a complement to the existing ones developed therein. However, existing estimators (i) first reduce the dimension of functional predictors and then use the dimension-reduced ones in the analysis (e.g., AP2006,SHIN2009), (ii) are designed without scalar-valued covariates (e.g., Hall2007, Florence2015, imaizumi2018), or (iii) do not account for time series dependence, which is crucial in our context (except for Babii2022, among the aforementioned articles). The approaches in (i) may not be preferred because they do not take into account covariation across variables in the dimension reduction procedure, while the latter two may limit practical applicability. Our estimator complements them by using a regularization method that allows us to account for covariation between scalar- and function-valued variables. This point is particularly relevant if the function-valued variable is given by a regressor without a structural interpretation, as in Section (ref). Later in Appendix (ref), we further extend our method to accommodate an endogenous functional predictor.
As discussed in Hall2007, Florence2015, and references therein, estimating functional linear regression models intrinsically involves an inverse problem similar to those in nonparametric estimation. Although we focus on a model that is linear in functional variables, these function-valued random variables can be represented by an infinite number of basis functions. In this regard, our paper is also related to studies on nonlinear impulse responses in, e.g., Oscar2005, kilian2017structural, and KP2024.
Although our main focus is on estimating impulse response coefficients when shocks are characterized by functions, our methodology could also be an empirically appealing alternative to prediction methods, such as that in Barbaglia2023, in that it allows us to utilize richer information, such as distributional information or observations at different frequencies. In particular, if a functional variable is given by a predictor without a structural interpretation, our model reduces to a predictive regression model with a functional predictor. As studied in Babii2022 and seong2021functional, functional predictors could lead to better forecasting outcomes by exploiting information overlooked in conventional estimation approaches.
In our empirical application, we first employ the quantile curve of the economic sentiment measure proposed by Barbaglia2023 and study the impact of perturbations given to the sentiment distribution on Total Nonfarm Payrolls in the US, without a structural interpretation. The sentiment measure is observed at a daily frequency and describes the presence of negative or positive tones in relation to specific terms, such as \textsl{economy} or \textsl{inflation}. The proposed estimator shows that the response depends not only on the magnitude of the perturbation, but also on its shape, which has not been observed previously. Moreover, we find that location shifts in the sentiment distribution have significant effects on the target variable in the short term. For instance, if the overall sentiment distribution shifts to the left, the average sentiment on the economy becomes more pessimistic, predicting significantly negative economic growth in the near future.
We then revisit the work by IR2021 on the impact of monetary policy shocks, which are characterized by shifts in yield curves on monetary policy announcement dates. We find that during conventional periods, inflation responses are generally insignificant; however, their impact tends to vary depending on how the shock affects interest rates at short or long maturities. That is, the impact of functional monetary policy shocks depends on their shapes and directions.
The rest of the paper is organized as follows. Section (ref) motivates our benchmark model and provides examples of functional variables. In Section (ref), we consider the SVAR model involving a functional variable and discuss when the parameters in our model can be interpreted as functional structural impulse responses. Section (ref) proposes our estimator and establishes its asymptotic properties. We apply our estimators to the empirical data in Section (ref). Section (ref) summarizes simulation results. Section (ref) concludes. The appendix includes proofs and an extension of the main theoretical results to instrumental variable (IV) estimation. In the Supplementary Material, we provide mathematical preliminaries and extension to functional SVAR models.
We let $y_t$ denote a scalar-valued target (dependent) variable and $\mathbf{w}_t= ( {w}_{1,t},\ldots, {w}_{m,t})'$ be a vector of exogenous (scalar-valued) control variables possibly containing lagged $y_t$'s. The variable $X_t = \{X_t (r) : r \in [0,1]\} $ denotes a function-valued variable that takes values in some function space $\mathcal H$ to be detailed shortly.\footnote{For the ease of explanation, we suppose that the domain of $X_t$ is $[0,1]$, but its extension to any arbitrary compact interval $[a,b]$ is straightforward.} Examples of $X_t$ include, but are not limited to, sentiment quantile curves in Example (ref) and functional monetary policy shocks in IR2021. In this paper, we are particularly interested in the response of $y_{t+h}$ when an additional perturbation (or shock) $\zeta$ is introduced to $X_t$; that is, for any $x \in \mathcal H$, we are interested in
In previous studies on time series analysis, including Oscar2005 and PW2021, $X_t$, the variable exposed to a perturbation, is characterized by a scalar or a vector, and the linear projection of $y_t$ onto the space spanned by $\{X_t, \mathbf w_t\}$ reduces the impulse response analysis to inference of projection coefficients. However, its extension to the case where $X_t$ is given by a function is not straightforward. For example, in our setup, the perturbation $\zeta$ in \hyperref[{eq: irf: general}]{\tagform@{\ref*{eq: irf: general}}} is specified as a function on $[0,1]$. Thus, the response of $y_t$ in \hyperref[{eq: irf: general}]{\tagform@{\ref*{eq: irf: general}}} depends not only on the magnitude of $\zeta$, but also on its shape. This feature has not been observed in the conventional studies on SVAR models.
We follow Oscar2005 and consider the following linear projection of the scalar target variable $y_t$ onto the space spanned by the variables in the conditioning set, including the function-valued $X_t$:
The functional covariate $X_t$ and its coefficient $\beta_h$ are assumed to take values in $\mathcal H=L^2[0,1]$ (the Hilbert space of square-integrable functions), which allows us to conveniently write $\int_{0}^{1} \beta_h(r)X_t(r) dr$ as the inner product $\langle \beta_h, X_t \rangle$ defined on $\mathcal H$.
Under the linear projection in \hyperref[{eq: model: benchmark}]{\tagform@{\ref*{eq: model: benchmark}}}, we identify \hyperref[{eq: irf: general}]{\tagform@{\ref*{eq: irf: general}}} as follows:
If $X_t$ is given by a functional structural shock (such as the functional monetary policy shock in IR2021) or if the true data generating process (DGP) follows a special structure in Section (ref), the parameter $\beta_h$ reduces to the structural impulse response function (SIRF) of $y_t$ when $X_t$ experiences a function-valued shock $\zeta$. In this regard, we extend the seminal work of Oscar2005 to allow for a functional variable.
If $X_t$ is identified as a functional structural shock, as in IR2021, the parameter $\beta_h$ in \hyperref[{eq: model: benchmark}]{\tagform@{\ref*{eq: model: benchmark}}} can be certainly interpreted as the SIRF of $y_t$ to that shock. However, its general link to the SIRF similar to those in Oscar2005 or PW2021 has not yet formally established in the functional setup, although \hyperref[{eq: irf: general}]{\tagform@{\ref*{eq: irf: general}}} and \hyperref[{eq: model: benchmark}]{\tagform@{\ref*{eq: model: benchmark}}} partly provide an intuitive guidance on its interpretation as a direct estimator of the reduced-form impulse response.
In this section, we show that the parameter $\beta_h$ in \hyperref[{def: local: irf}]{\tagform@{\ref*{def: local: irf}}} provides a crucial and interesting interpretation as a SIRF when the true DGP follows an autoregressive structure. This suggests the importance of our benchmark model \hyperref[{eq: model: benchmark}]{\tagform@{\ref*{eq: model: benchmark}}}, particularly as an extension of Oscar2005.
For the ease of exposition, we exclude $\mathbf w_t$ in this section. The results given in this section can be easily extended to include other covariates as long as they are finite dimensional. Thus we do not lose any generality by this simplification.
As mathematical preliminaries, Section (ref) of the Supplementary Material provides a brief review of definitions and properties of \(\mathcal{H}\)-valued random variable \(X_t\), linear operators on \(\mathcal{H}\), and the product Hilbert space \(\mathbb{R} \times \mathcal{H}\), where the tuple \(\left[
\right]\) takes values.\footnote{We treat the tuple \( \left[
\right] \), consisting of \(\mathbb{R}\)-valued and \(\mathcal{H}\)-valued random variables, as a “column vector” in the multivariate setting. This notation introduces no confusion or mathematical imprecision in the subsequent discussion. Similar notation will be used for any other tuples appearing later.} As detailed in the appendix, a linear operator \(\mathcal{D}\) on \(\mathbb{R} \times \mathcal{H}\) can be represented as an operator matrix, say $\mathcal{D} = \left[
\right].$ This operator maps $ \left[
\right] \in \mathbb{R} \times \mathcal{H} $ to $\left[
\right] \in \mathbb{R} \times \mathcal{H}$, analogous to transforming a \((2 \times 1)\) vector using a \((2 \times 2)\) matrix. For this reason, with a slight abuse of notation, we henceforth treat the \((\mathbb{R} \times \mathcal{H})\)-valued random element \( \left[
\right]\) as if it were a \((2 \times 1)\) random vector, and any linear operator on \(\mathbb{R} \times \mathcal{H}\) as if it were a \((2 \times 2)\) matrix. This simplification is valid in the considered setup without any loss of technical rigor (see Section (ref) of the Supplementary Material).
In this section, we suppose that the tuple $\left[
\right] \in \mathbb R \times \mathcal H$ follows the model below:
The above model is identical to the standard bivariate SVAR model other than that $X_t$ is a $\mathcal H$-valued random variable, and thus the coefficients $\mathcal A$ and $\mathcal B$ are in fact operator matrices. Specifically, $\alpha_{12}$ and $ \beta_{12}$ are given by linear maps from $ \mathcal H$ to $ \mathbb R$, while $\alpha_{21}$ and $\beta_{21}$ are maps from $\mathbb R$ to $\mathcal H$. The $(1,1)$-th (resp.\ $(2,2)$-th) elements, $I_1$ and $a_{11}$ (resp.\ $I_2$ and $a_{22}$), are linear operators acting on $\mathbb{R}$ (resp.\ $\mathcal H)$, where $I_1$ and $I_2$ denote the identity maps in the relevant spaces. The tuple of structural shocks $\left[
\right]$ is a $(\mathbb R \times \mathcal H)$-valued random element. Its covariance operator (see \hyperref[{eqcovop}]{\tagform@{\ref*{eqcovop}}} in the Supplementary Material) is given as follows:
where $\otimes$ denotes the tensor product, which generalizes the outer product in the Euclidean space, and $\sigma_{11}$ (resp.\ $\mathbf{\Sigma}_{22}$) is a linear operator acting on $\mathbb{R}$ (resp.\ $\mathcal H$).\footnote{Since $\mathbf{\Sigma}$ is an operator matrix whose $(i,j)$-th entry is a map from $\mathcal{H}_i$ to $\mathcal{H}_j$, with $\mathcal{H}_1 = \mathbb R$ and $\mathcal{H}_2=\mathcal{H}$, $\sigma_{11}$ is given by a linear map on $\mathbb{R}$. However, any linear map on $\mathbb{R}$ is nothing but a scalar multiplication, given by $c I_1$ for some $c \in \mathbb{R}$. Thus, there is little risk of confusion even if we understand $\sigma_{11}$ as a real-valued constant. } See Section (ref) of the Supplementary Material for details.
We first establish the fully functional identification of structural parameters and highlight its differences from existing approaches. Then, we link the SIRF implied by the SVAR structure to the coefficients in our benchmark model.
It can be shown that the operator $\mathcal B$ in \hyperref[{eq: model: svar}]{\tagform@{\ref*{eq: model: svar}}} is invertible and $\Gamma:=\mathcal B^{-1}\mathcal A$ can be well defined under the condition in Proposition (ref) that will appear shortly. Thus \hyperref[{eq: model: svar}]{\tagform@{\ref*{eq: model: svar}}} can be written into the following reduced-form VAR (RFVAR) model:
$\Sigma_{\varepsilon}$ denotes the covariance of $\left[
\right]$. In this section, we discuss conditions under which the structural parameters in \hyperref[{eq: model: svar}]{\tagform@{\ref*{eq: model: svar}}} can be identified from the reduced-form parameters in \hyperref[{eqreduced1}]{\tagform@{\ref*{eqreduced1}}}. Our assumption for achieving such identification is based on the standard causal ordering of the (contemporaneous) variables in the system, as in Sims1972,Sims1980. Later, we will show that the coefficients of our benchmark model reduces to the SIRFs under this identification scheme.
Although the conditions (ref) and (ref) in Proposition (ref) are seemingly similar to each other, they suggest different levels of complexity in estimating the outcome equations. Specifically, if there is no contemporaneous impact of function-valued $X_t$ to scalar-valued $y_t$ (i.e., $\beta_{12} = 0$), the outcome equation of $y_t$ reduces to $y_t = \alpha_{11} y_{t-1} + \alpha_{12}X_{t-1} + u_{1t}$, and thus it includes only one infinite dimensional parameter $\alpha_{12}$. On the other hand, under the second identification condition (i.e., $\beta_{21} = 0$), the outcome equation reduces to $y_t = \alpha_{11} y_{t-1} -\beta_{12}X_t + \alpha_{12}X_{t-1} + u_{1t}$, in which the number of infinite dimensional parameters doubles. Estimating these parameters involves solving an inverse problem that necessitates the use of a regularization scheme. Considering this, the resulting estimator from the second identification scheme is likely to suffer from a larger regularization bias.
Another interesting implication of Proposition (ref) is that the standard rank-based identification strategy cannot be straightforwardly translated into our setup involving a functional covariate. Instead, depending on the direction of the contemporaneous impact, we may additionally need an injectivity condition to identify structural parameters from the reduced-form parameters. For example, under the condition (ref) in Proposition (ref), $\Sigma_{\varepsilon,22} = \Sigma_{22}$ and $\beta_{12}$ satisfies
see our proof of Proposition (ref) in Appendix (ref). In this case, the injectivity of $\Sigma_{22}$ is required to uniquely identify $\beta_{12}$ from \hyperref[{eqinject}]{\tagform@{\ref*{eqinject}}} (see carrasco2007linear,Seo2024).
The two observations mentioned above contrast with identification conditions and their implications in the existing literature, such as Sims1972, where a different ordering of variables primarily affects their economic interpretation and the direction of contemporaneous shocks, rather than the level of asymptotic bias or computational burden. Thus, in the SVAR model with functional covariates, identification conditions need to be cautiously chosen with taking into account the above.
To understand the relationship between the coefficient in our benchmark model \hyperref[{eq: model: benchmark}]{\tagform@{\ref*{eq: model: benchmark}}} and the SIRF implied by \hyperref[{eq: model: svar}]{\tagform@{\ref*{eq: model: svar}}}, it is convenient to consider the MA($\infty$) representation. Under the invertibility of $\mathcal B$, the SVAR model in \hyperref[{eq: model: svar}]{\tagform@{\ref*{eq: model: svar}}} allows the reduced-form representation in \hyperref[{eqreduced1}]{\tagform@{\ref*{eqreduced1}}} and furthermore, we have
where $\mathbf{u}_t = \left[
\right] \in\mathbb R \times \mathcal H$ and $\Gamma ^0 = I$ which is the identity operator on $\mathbb{R}\times \mathcal H$. \hyperref[{eq: irf: general}]{\textup{\tagform@{\ref*{eq: irf: general}}}} and \hyperref[{eqrfvar}]{\textup{\tagform@{\ref*{eqrfvar}}}} suggest that the SIRF of $ \left[
\right]$ at horizon~$h$, when the perturbation $\widetilde\zeta \equiv \left[
\right] \in \mathbb{R}\times \mathcal H$ is introduced to $\mathbf{u}_t$, is given as follows:
Both $\Gamma ^h$ and $\mathcal B^{-1}$ are operator matrices mapping from $\mathbb R \times \mathcal H$ to $\mathbb R \times \mathcal H$, and the same holds for $\Gamma ^h\mathcal B^{-1}$. Therefore, $\Gamma ^h \mathcal B^{-1}$ allows the following representation:
where $P_{\mathbb{R}}$ denotes the projection map given by $P_{\mathbb{R}}\left[
\right] = x_1$ and $P_{\mathbb{R}}^\ast$ is its adjoint satisfying $P_{\mathbb{R}}^\ast(x_1) = \left[
\right]$. Similarly, $P_{\mathcal{H}}$ (resp.\ $P_{\mathcal{H}}^\ast$) is a projection map satisfying $P_{\mathcal{H}} \left[
\right] = x_2$ (resp.\ $P_{\mathcal{H}}^\ast (x_2) =\left[
\right] $) for $x_2 \in \mathcal H$. Given these projection maps, each element of $\Gamma ^h \mathcal B^{-1}$ is characterized by different linear maps\footnote{$\operatorname{SIRF}_{11,h} : \mathbb R \to \mathbb R$ (i.e., a scalar multiplication), $\operatorname{SIRF}_{12,h}: \mathcal H \to \mathbb R$, $\operatorname{SIRF}_{21,h}: \mathbb R \to \mathcal H$ and $\operatorname{SIRF}_{22,h}:\mathcal H \to \mathcal H$.} and measures an effect of an additional structural shock on $y_{t+h}$ or $X_{t+h}$. For example, the response of $y_{t+h}$ to a function-valued shock $\zeta\in \mathcal H$ is given by a scalar such that
The response of $X_{t+h}$ to the unit shock on $u_{1,t}$ is characterized by a function such that
These are often of interest to practitioners. The SIRFs in \hyperref[{eqpartialirf1}]{\tagform@{\ref*{eqpartialirf1}}} and \hyperref[{eqpartialirf2}]{\tagform@{\ref*{eqpartialirf2}}} can be interpreted similarly to those of standard SVAR models. However, because $X_t$ and $U_{2,t}$ are given by $\mathcal H$-valued functions of infinite dimension, the implementation of $\zeta$ to the structural error can be represented with an infinite number of basis functions. That is, for $\{ \xi_j\}_{j \geq 1}$, a set of orthonormal basis functions that span $\mathcal H$, we have
In this regard, the SIRF in \hyperref[{eqpartialirf1}]{\tagform@{\ref*{eqpartialirf1}}} will be interpreted as the response of $y_t$ when all the basis functions move jointly to the direction of $\zeta$. A similar observation was previously made by IR2021 under a parametric assumption, and it remains valid even in our setting without that assumption.
Proposition (ref)(ref) provides a theoretical foundation for interpreting the map $\langle \beta_h, \cdot \rangle$ in the benchmark model \hyperref[{eq: model: benchmark}]{\tagform@{\ref*{eq: model: benchmark}}} as $\operatorname{SIRF}_{12,h}$. Specifically, if the identification condition $\beta_{12} = 0$ holds and the true DGP follows \hyperref[{eq: model: svar}]{\tagform@{\ref*{eq: model: svar}}}, then the coefficient map $\langle \beta_h, \cdot \rangle$ represents the SIRF at horizon $h$ when the functional structural error associated with $X_t$ experiences a shock $\zeta$. It may be possible to estimate the SIRFs directly from the RFVAR model by applying an appropriate regularization scheme required for an infinite dimensional setup. We provide a brief outline of this approach in Section (ref) of the Supplementary Material. However, as detailed in the appendix, this approach makes statistical inference on the SIRFs much more challenging, while our benchmark model simplifies the inference procedure by exclusively focusing on the outcome equation of interest.
Proposition (ref)(ref) implies that $\operatorname{SIRF}_{21,h}$ can be characterized by a certain coefficient map in a function-on-function regression model with scalar control variables. We study how statistical inference can be implemented for this quantity in Section (ref) of the Supplementary Material.
In empirical studies involving functional random variables, it is a common practice to first approximate the functional variable as a finite dimensional vector and then apply existing estimation or inference methods. However, as Nielsen2023 pointed out in the context of cointegration tests for functional time series, this finite dimensional approximation can lead to misspecification errors, causing misleading estimation and interpretation. To illustrate this, suppose
where $ \widetilde \Upsilon_t = \left[
\right] \in \mathbb{R}\times \mathcal H$. To study the above model, a popular approach is to transform $X_t$ in $\widetilde\Upsilon_t$ into a finite dimensional vector, say $\mathbf{x}_t$, in advance and then estimate a VAR model similar to \hyperref[{eq: model: svar3}]{\textup{\tagform@{\ref*{eq: model: svar3}}}} with $(y_t,\mathbf{x}_t')'$. The transformation from $X_t$ to $\mathbf{x}_t$ may be expressed as a projection of $X_t$ onto a finite dimensional space such that $P_0{\widetilde\Upsilon}_t = (y_t, \langle X_t, \xi_1 \rangle,\ldots, \langle X_t, \xi_K \rangle)'$ with some finite $K$. $\xi_j$ may be either a pre-specified parametric function (IR2021) or a principal component (BHCJ2023). However, as shown by Nielsen2023, this projection operation does not preserve the VAR structure. Specifically, from \hyperref[{eq: model: svar3}]{\tagform@{\ref*{eq: model: svar3}}}, we have
If $P_0 \neq I$, $P_0\widetilde\Upsilon_{t}$ does not follow the VAR(1) structure in \hyperref[{eq: model: svar3}]{\tagform@{\ref*{eq: model: svar3}}} since it is generally correlated with $\mathbf{v}_t$, which potentially invalidates many existing estimation and inference methodologies developed for VAR models; for instance, Nielsen2023 show that the cointegration rank test of Johansen1991,Johansen1995 is generally misleading in this setup. This finite dimensional approximation can be justified only when $P_0$ is close enough to $I$ (unless we consider the special case where $X_t$ can be fully expressed by a finite number of known basis functions, which is essentially equivalent to a finite dimensional setup). This implies that \( K \) must be sufficiently large. In functional linear models, it is well known that an increase in \( K \) can significantly raise the variance of coefficient estimators while reducing the regularization bias. Thus, choosing \( K \) in advance without accordance with the desired inferential methods may introduce an unbalanced bias-variance trade-off.
To facilitate our discussion, it is useful to consider an alternative representation of our benchmark model in \hyperref[{eq: model: benchmark}]{\tagform@{\ref*{eq: model: benchmark}}}. We let $\widetilde{\mathcal H}$ denote the product Hilbert space, whose inner product is given by the sum of the inner products in $\mathbb{R}^m$ and $\mathcal H$ (see Section (ref) of the Supplementary Material). With a slight abuse of notation, we let $\langle \cdot, \cdot \rangle$ denote the inner product defined on any of $\mathbb{R}^m$, $\mathcal H$, and $\widetilde{\mathcal H}$ (e.g., if $h_1, h_2 \in \mathbb{R}^m$ then $\langle h_1,h_2 \rangle=h_1'h_2$). Because the inner product is inherently defined for elements in the same space, there is little risk of confusion following this simplification. For any elements $h_j \in \mathcal H_j$ and $h_k\in \mathcal H_k$, where $\mathcal H_j$ and $\mathcal H_k$ can be any of $\mathbb{R}^m$, $\mathcal H$, and $\widetilde{\mathcal H}$, we let $\otimes$ denote the tensor product, defined by $h_j\otimes h_k (\cdot) = \langle h_j,\cdot \rangle h_k$ (see Section (ref) of the Supplementary Material). In particular, if $h_1 = \left[
\right] \in \widetilde{\mathcal H}$ and $h_{2} = \left[
\right] \in \widetilde{\mathcal H}$, then $h_1\otimes h_2$ can be understood as an operator matrix given by $\left[
\right]$. We define the following $\widetilde{\mathcal H}$-valued random element $\Upsilon_t$ and its coefficient $\theta_h$:
Under this representation, $ \langle\theta_h, \Upsilon_t\rangle = \langle \alpha_h,\mathbf{w}_t \rangle + \langle \beta_h, X_t \rangle$, and \hyperref[{eq: model: benchmark}]{\tagform@{\ref*{eq: model: benchmark}}} is simplified into the following regression model:
We hereafter assume that the variables $y_{t+h}$, $X_t$ and $\mathbf w_t$ have zero means. Extending to the case where these means are unknown is straightforward by considering their demeaned values, under the assumptions detailed shortly. Furthermore, we assume that $X_t$ and $\mathbf{w}_t$ are exogenous with respect to $u_{h,t}$. Then, by its definition in \hyperref[{eqdefY}]{\tagform@{\ref*{eqdefY}}}, $\Upsilon_t$ satisfies the exogeneity condition such that $C_{\Upsilon u} = \mathbb{E}[u_{h,t}\Upsilon_t] = 0$. This in turn implies the following moment condition:
Under the identification condition to be detailed shortly, $\theta_h$ can be estimated from the sample counterpart of \hyperref[{eqpopmoment}]{\tagform@{\ref*{eqpopmoment}}} given by
where $\widehat{C}_{\Upsilon y}$ and $\widehat{C}_{\Upsilon\Upsilon}$ are given by $ \widehat{C}_{\Upsilon y} = T^{-1}\sum_{t=1}^T y_{t+h} \Upsilon_t$ and $ \widehat{C}_{\Upsilon \Upsilon} = T^{-1} \sum_{t=1}^T \Upsilon_t \otimes \Upsilon_t$. However, in our setup, $\bar{\theta}_h$ is not generally obtainable from \hyperref[{eqsammoment}]{\tagform@{\ref*{eqsammoment}}}, since $ \widehat{C}_{\Upsilon\Upsilon}$, an operator acting on $\widetilde{\mathcal H}$, is not invertible. We circumvent this issue by considering an estimator constructed using a regularized inverse of $\widehat{C}_{\Upsilon\Upsilon}$ based on its operator Schur complement. Specifically, we note that whenever convenient, ${C}_{\Upsilon\Upsilon}$ and $\widehat{C}_{\Upsilon\Upsilon}$ can be understood as operator matrices on $\widetilde{\mathcal H}$ such that
where $\Gamma_{ij} = \mathbb{E}[g_{jt}\otimes g_{it}]$, $\widehat{\Gamma}_{ij} = T^{-1}\sum_{t=1}^T g_{jt}\otimes g_{it}$, $g_{1t} = \mathbf{w}_t$, and $g_{2t} = X_t$. That is, for any $\mathbf{a} \in \mathbb R^m$, $\Gamma_{11}\mathbf{a} = \mathbb E[ \langle \mathbf w_t, \mathbf a \rangle \mathbf{w}_t]$ and $\Gamma_{21}\mathbf{a} = \mathbb E[ \langle \mathbf w_t, \mathbf a \rangle X_t]$. Similarly, for any $ h\in \mathcal{H}$, $\Gamma_{12} h = \mathbb E[\langle X_t, h\rangle \mathbf w_t]$ and $\Gamma_{22} h = \mathbb E[\langle X_t, h \rangle X_t]$. Note that $\Gamma_{11}$ is the covariance of $ \mathbf{w}_t$ which is assumed to be invertible throughout the paper (see Assumption (ref)). We then define the following operator Schur complements $\mathrm{S}$ (of ${C}_{\Upsilon\Upsilon}$) and $\widehat{\mathrm{S}}$ (of $\widehat{C}_{\Upsilon\Upsilon}$) (see Bart2007):
{The operator Schur complement $\mathrm{S}$ can be understood as the covariance of the residuals obtained by regressing $X_t$ on $\mathbf{w}_t$, and $\widehat{\mathrm{S}}$ is its sample counterpart (see fukumizu2004dimensionality).} Lastly, to introduce our regularization scheme, we further represent $\mathrm{S}$ (resp.\ $\widehat{\mathrm{S}}$) with respect to its eigenvalues and eigenvectors $\{ \lambda_j,\nu_j \}_{j \geq 1}$ (resp.\ $\{ \widehat\lambda_j, \widehat \nu_j \}_{j \geq 1}$) as follows:
where $\lambda_1\geq\lambda_2\geq\ldots \geq 0$ and $\widehat{\lambda}_1\geq\widehat{\lambda}_2\geq\ldots\geq 0$. This representation is possible as $\mathrm{S}$ and $\widehat{\mathrm{S}}$ are self-adjoint, nonnegative, and compact (see Bosq2000, p.\ 34). The empirical eigenelements $\{\widehat{\lambda}_j , \widehat{v}_j\}$ can be obtained using the functional principal component analysis (FPCA). From \hyperref[{eqshur0}]{\tagform@{\ref*{eqshur0}}}, noninvertibility of $\widehat{\mathrm{S}}$ is evident as its partial inverse $\widehat{\mathrm{S}}_K^{-1} = \sum_{j=1}^K \widehat{\lambda}_j ^{-1}\widehat{v}_j\otimes \widehat{v}_j$ grows without bound in the operator norm as $K$ increases. This consequently leads to noninvertibility of $\widehat{C}_{\Upsilon\Upsilon}$ (see Bart2007, p.\ 29). Therefore, to construct a regularized inverse of $\widehat{C}_{\Upsilon\Upsilon}$, we first construct regularized inverses of the operator Schur complements $\mathrm{S}$ and $\widehat{\mathrm{S}}$ as follows:
for the regularization parameter $\tau$, decaying to zero as $T$ increases. Then, we define the following regularized inverse $\widehat{C}_{\Upsilon\Upsilon,\operatorname{\mathrm{K}_{\tau}}}^{-1}$ of $\widehat{C}_{\Upsilon\Upsilon}$:
By using the regularized inverse \hyperref[{eqreginv}]{\tagform@{\ref*{eqreginv}}} and the moment condition \hyperref[{eqsammoment}]{\tagform@{\ref*{eqsammoment}}}, we define our estimator $\widehat\theta_h$ as follows:
Let $\ker A$ denote the kernel of $A$. To uniquely identify $\theta_h$ from \hyperref[{eqpopmoment}]{\tagform@{\ref*{eqpopmoment}}}, we assume the following:
Note that $\Upsilon_t$ involves the $\mathcal H$-valued random variable $X_t$. In the literature on functional data analysis, it is commonly assumed that the covariance of such a functional random variable allows infinitely many nonzero eigenvalues. Therefore, as deduced from Mas2007 and the Riesz representation theorem (Conway1994, p.\ 13), the parameter of interest in \hyperref[{eqpopmoment}]{\tagform@{\ref*{eqpopmoment}}}, $\theta_h$, is not uniquely identified as an element of $\widetilde{\mathcal H}$ if $\ker C_{\Upsilon \Upsilon} \neq \{0\}$. The failure of identification occurs because, for any $\psi \in \ker C_{\Upsilon\Upsilon}$, $C_{\Upsilon\Upsilon} ( \theta_h+\psi ) = C_{\Upsilon\Upsilon}\theta_h$. Thus, if $\theta_h$ satisfies \hyperref[{eqpopmoment}]{\tagform@{\ref*{eqpopmoment}}}, $\theta_h+\psi$ also satisfies it. Assumption (ref) prevents this failure of identification.
Our next assumption is related to asymptotic properties of $\widehat\theta_h$. Below, $\widehat{C}_{\Upsilon u} = T^{-1}\sum_{i=1} ^T u_{h,t} \Upsilon_t$ and $\Vert \cdot \Vert _{\operatorname{op}}$ is the operator norm (see Section (ref) of the Supplementary Material for its formal definition).
The $L^4$-$m$-approximability in Assumption (ref)(ref), whose formal definition is provided in Section (ref) of the Supplementary Material, is employed to use existing limit theorems. The assumption is not only widely adopted in the literature on stationary functional time series but also inclusive of many practical and interesting examples, such as the SVAR model in Section (ref). Assumption (ref)(ref) contains high-level conditions on the limiting behavior of $\widehat{C}_{\Upsilon \Upsilon}$ and $\widehat{C}_{\Upsilon u}$, and is not restrictive given the stationarity of $\{\Upsilon_t\}$ and $\{ u_{h,t}\Upsilon_t\}$. Some primitive sufficient conditions can be found in Bosq2000.
We impose the following conditions to characterize the rate of convergence of $\widehat\theta_h$.
When $\mathrm{S}$ is given by the standard covariance of the functional explanatory variable, Assumption (ref)(ref) reduces to a standard assumption commonly used in the literature on functional linear models. Given that $\sum_{j=1}^\infty \lambda_j < \infty$ must hold (see Lemma (ref)), the first condition $\lambda_j^2 \leq \mathtt{C} j^{-\rho}$ for some $\rho>2$ is natural (obviously, $\rho\leq 2$ may not result in the summability of $\{\lambda_j\}_{j\geq 1}$). By the latter condition, we require that the eigenvalues of $\mathrm{S}$ are well separated. As documented in a similar context concerning functional linear models (see e.g., Hall2007,imaizumi2018,seong2021functional), this separation is crucial for achieving sufficient accuracy in the estimation of eigenelements. Assumption (ref)(ref) can be understood as a smoothness condition on $\beta_h$ with respect to the eigenvectors $\{v_j\}_{j \geq 1} $. In contrast to $\beta_h$, a similar restriction is not necessary for $\alpha_h$ associated with the vector-valued random variable $\mathbf{w}_t$. Considering that $\beta_{h}$ is an element in a Hilbert space satisfying $\sum_{j=1}^\infty \langle \beta_h, v_j \rangle^2 <\infty$, Assumption (ref)(ref) does not seem too restrictive.
The following theorem states consistency and the rate of convergence of $\widehat\theta_h$.
In Theorem (ref), the choice of the regularization parameter $\tau$ to ensure the consistency of the proposed estimator depends on $\rho$, allowing $\tau$ to decay at a faster rate as $\rho$ increases. Assuming that the same regularization parameter is employed, an increase in $\rho$ implies a faster convergence of $\widehat{\theta}_h$ to $\theta_h$. It is worth noting that the consistency is established if $\tau$ satisfies $T\tau^{3} \to \infty$ (such as e.g., $\tau = T^{-1/3 + \epsilon}$ or $T^{-1/3} \log^{\epsilon} T$ for small $\epsilon>0$) as long as $\rho >2$ as assumed in Assumption (ref); of course, $\tau$ decaying at a slower rate can also be considered in practice, without affecting consistency. Nevertheless, this naive selection of $\tau$ may be particularly advantageous for practitioners with little knowledge of these eigenvalues.
Given that $\langle \beta_h, \zeta\rangle$ can naturally be interpreted as the response to a perturbation applied to $X_t$, as in \hyperref[{def: local: irf}]{\tagform@{\ref*{def: local: irf}}}, and in some special cases, such as in Proposition (ref), it can further be interpreted as $\operatorname{SIRF}_{12,h}(\zeta)$, it is of our interest to conduct statistical inference on this quantity. More generally, we consider inference on \( \langle \theta_h, \zeta \rangle \) for any \( \zeta \in \widetilde{\mathcal H} \). To clarify the perturbation \( \zeta \) in the subsequent discussion and avoid potential confusion regarding the space in which \( \zeta \) takes values, we let
By setting \( \zeta_1 = 0 \), inference on \( \langle \theta_h, \zeta \rangle \) reduces to inference on \( \langle \beta_h, \zeta_2 \rangle \) for \( \zeta_2 \in \mathcal H \). Meanwhile, if $\zeta_2 = 0$, it reduces to inference on $\langle \alpha_h, \zeta_1\rangle \equiv \alpha_h ' \zeta_1$.
Let $\widehat{u}_{h,t}= y_{t+h}- \langle \Upsilon_t, \widehat{\theta}_h\rangle$ and $\mathrm{k}(\cdot)$ be a standard weight function to be specified shortly. Then, under Assumption (ref), we define the following long-run covariance operator and its sample counterpart (see berkes2013weak): below, ${\mathcal U}_{h,t} = {u}_{h,t}\Upsilon_{t}$ and $\widehat{\mathcal U}_{h,t} = \widehat{u}_{h,t}\Upsilon_{t}$.
We also define the quantity $c_{m,j}$ as
where $\{ \bar{\lambda}_j \}$ and $ \{ \bar{v}_j\}$ denote the eigenvalues and eigenvectors of $C_{\Upsilon\Upsilon}$, which are generally different from the eigenelements of $\mathrm{S}$. The quantity satisfies $\sum_{j=1}^{m} c_{m,j}^2(\zeta) = 1$ for every $m$ unless $\langle \zeta,\bar{v}_j \rangle \neq 0$ for at least one $j$.
To establish local asymptotic normality of $\widehat\theta_h$, we employ the following assumption:
Assumption (ref)(ref) is adopted from horvath2013estimation and imposed for convenience in our proof of consistency of $\widehat{\Lambda}_{\mathcal U}$. Assumption (ref)(ref) is imposed to ensure nondegenerate convergence rate in our asymptotic normality result; in the special case where $\Lambda_{\mathcal U} = \sigma_u^2 C_{\Upsilon\Upsilon}$ for some $\sigma_u^2>0$,\footnote{This can happen when $\{u_t\}$ is a homogeneous martingale difference with respect to the filtration $\mathfrak F_t=\sigma(\{u_{s}\}_{s\leq t-1},\{\Upsilon_{s}\}_{s\leq t})$ as in the case considered by seong2021functional} the former condition $\zeta \notin \ker C_{\Upsilon\Upsilon}$ implies the latter, and hence Assumption (ref)(ref) can be simplified. Assumptions (ref)(ref) and (ref)(ref) are technical requirements facilitating our asymptotic analysis, and are not quite restrictive. In particular, as $\Lambda_{\mathcal U}$ is a covariance operator (Bosq2000), the quantity in Assumption (ref)(ref) converges to a nonnegative constant for every $\zeta \in \widetilde{\mathcal H}$. Thus, what we require of $\zeta$ is solely to ensure the positivity of the limit. Assumption (ref)(ref) is similar to standard assumptions found in the literature (see Bosq2000).
We employ the following assumption: below, we let $X_{w,t} = X_t-\Gamma_{21}\Gamma_{11}^{-1}\mathbf{w}_t$, $r_t(j,\ell) = \langle X_t,v_j \rangle \langle X_{w,t},v_{\ell} \rangle - \mathbb E[\langle X_t,v_j \rangle \langle X_{w,t},v_{\ell} \rangle]$ for $j,\ell \geq 1$, and $\mathtt{C}$ be a generic positive constant.
A similar but slightly different assumption can be found in seong2021functional and references therein. In particular, Assumption (ref)(ref) pertains to the smoothness of $\zeta$; similar to Assumption (ref)(ref), it is natural to consider $\delta >1/2$ as $\zeta_2\in\mathcal H$ and $ {\Gamma}_{11}^{-1}{\Gamma}_{12}$ is a bounded linear functional.\footnote{There exists unique element $h \in \mathcal H$ such that ${\Gamma}_{11}^{-1}{\Gamma}_{12}v_j = \langle h, v_j \rangle$ by the Riesz representation theorem (see Conway1994, p.\ 13), and $\sum_{j=1}^\infty\|{\Gamma}_{11}^{-1}{\Gamma}_{12}v_j\|^2 < \infty$ as $\{v_j\}$ is an orthonormal basis in $\mathcal H$.}
Lastly, we let $\widehat{P}_{\operatorname{\mathrm{K}_{\tau}}}=\widehat{C}_{\Upsilon\Upsilon,\operatorname{\mathrm{K}_{\tau}}}^{-1} \widehat{C}_{\Upsilon\Upsilon} $ and $P_{\operatorname{\mathrm{K}_{\tau}}} = {C}_{\Upsilon\Upsilon,\operatorname{\mathrm{K}_{\tau}}}^{-1} {C}_{\Upsilon\Upsilon}$, where ${C}_{\Upsilon\Upsilon,\operatorname{\mathrm{K}_{\tau}}}^{-1}$ is defined by replacing $\widehat\Gamma_{ij}$ and $\widehat\mathrm{S}_{\operatorname{\mathrm{K}_{\tau}}}^{-1}$ in \hyperref[{eqreginv}]{\tagform@{\ref*{eqreginv}}} with $\Gamma_{ij}$ and $\mathrm{S}_{\operatorname{\mathrm{K}_{\tau}}}^{-1}$ respectively. Then, $\widehat{P}_{\operatorname{\mathrm{K}_{\tau}}} $ and $P_{\operatorname{\mathrm{K}_{\tau}}} $ allow the following representations:
where $I_1$ (resp.\ $I_2$) is the identity map on $\mathbb{R}^{m}$ (resp.\ $\mathcal H$) and
That is, $\Pi_{\operatorname{\mathrm{K}_{\tau}}}$ is the projection onto the span of $\{v_j\}_{j=1}^{\operatorname{\mathrm{K}_{\tau}}}$ and $\widehat{\Pi}_{\operatorname{\mathrm{K}_{\tau}}}$ is its sample counterpart. Using $\widehat{P}_{\operatorname{\mathrm{K}_{\tau}}}$ and ${P}_{\operatorname{\mathrm{K}_{\tau}}}$, we decompose $\langle \widehat{\theta}_h-\theta_h,\zeta\rangle$ into $ \widehat{\Theta}_1+\widehat{\Theta}_{2A}+\widehat{\Theta}_{2B}$ such that
The quantity $\psi_{\operatorname{\mathrm{K}_{\tau}}}(\zeta)$ in Theorem (ref) plays an important role as a normalizing factor in our asymptotic normality result. With an obvious adaptation of the discussion given by seong2021functional, we know that it is convergent only on a strict subspace of $\widetilde{\mathcal H}$, and ${\psi}_{\operatorname{\mathrm{K}_{\tau}}}(\zeta)$ is likely to diverge (at a rate slower than $T$; see Remark (ref)) for $\zeta$ arbitrarily chosen by practitioners.
Similar to Theorem (ref), the decay rate of the regularization parameter in Theorem (ref) depends on $\rho$, and a larger value of $\rho$ leads to a faster convergence. Note that $\tau$ satisfying $T\tau^{4} \to \infty$ (e.g., $\tau = T^{-1/4 + \epsilon}$ or $T^{-1/4} \log^{\epsilon} T$ for small $\epsilon>0$) meets the requirement for Theorem (ref)(ref), provided that $\rho >2$ as assumed in Assumption (ref). Such a choice of $\tau$ may be preferred for practitioners. Of course, as in Theorem (ref), we may also consider $\tau$ decaying to zero at a slower rate, but this is not recommended due to the condition $T^{1/2}\tau^{(\delta+\varsigma-1)/\rho} \to 0$ for the result given in Theorem (ref)(ref). This condition can be understood as requiring sufficiently large $\varsigma$ and $\delta$ for a given $\tau$; moreover, as the decay rate of $\tau$ decreases, stricter requirements are imposed on $\varsigma$ and $\delta$.
If $\zeta$ is restricted to have a nonzero element only in $\mathbb{R}^m$ (i.e., $\zeta_2 = 0 \in \mathcal{H}$) and $\psi_{\operatorname{\mathrm{K}_{\tau}}}(\zeta) \to_p C_{\zeta} <\infty$, it can be shown that $\langle \widehat{\theta}_h - \theta_h, \zeta \rangle$ converges in distribution to a normal random variable at the rate of $\sqrt{T}$. However, we here emphasize that the former condition, $\zeta_2 = 0$, does not guarantee the latter condition, $\psi_{\operatorname{\mathrm{K}_{\tau}}}(\zeta) \to_p C_{\zeta}$. Thus, it is common to observe a nonparametric rate of convergence even for the parameters associated with the finite dimensional random variables in this setup (see Remark (ref) for more details).
In this section, we study how economic variables respond to random perturbations introduced to economic sentiment distribution in the US. To measure sentiment, we adopt \citepos{Barbaglia2023} data, which quantify sentiment on specific economic subjects from sentences published in major US news articles. The sentiment measure is defined by each token (i.e., word) and its neighboring sentences, and includes both intensity and tone for each subject. We refer interested readers to Barbaglia2023 for a detailed definition of the sentiment measure.
Among others, we particularly focus on the daily sentiment measure on the \textsl{economy}, and use it to construct monthly quantile sentiment curves. The quantile curves are represented by 31 Fourier basis functions, which become our functional predictor, and are reported in Figure (ref). As mentioned earlier in Section (ref) and also by Barbaglia2023, the sentiment quantile and its associated distribution seem to be closely related to business cycle fluctuations: during recessive periods, the overall quantiles tend to shift downward and exhibit a steeper slope coefficient, indicating a larger dispersion in the sentiment distribution. This could be interpreted as a higher level of disagreement about economic sentiment during recessive periods.
The dependent variable $y_t$ is specified to monthly Total Nonfarm Payroll (PAYEMS) in the US. The data are provided by the Federal Reserve Bank. We follow Barbaglia2023 and let $\mathbf{w}_t$ in \hyperref[{eq: model: benchmark}]{\tagform@{\ref*{eq: model: benchmark}}} be the vector consisting of the Chicago Fed National Activity Index (CFNAI), the National Financial Conditions Index (NFCI), and \citepos*{ADS2009} measure of economic activity (ADS). Because NFCI and ADS are observed at a higher frequency, we average them over months to match the frequency of the target variable. The sample runs from January 1984 to December 2021, with a total size of 456 observations.
The regularization parameter of our estimator is chosen by roughly considering $\rho$ satisfying Assumption (ref). Specifically, if $\mathtt{C}=1$, Assumption (ref)(ref) tells us that $\rho \geq \rho^\ast \equiv -(\log (\lambda_j ^2 - \lambda_{j+1} ^2) /\log j)-1$. Thus, we set $\widetilde \rho$ to $\lceil 100 \rho ^\ast \rceil/100$, where $\lceil \cdot \rceil$ is the ceiling function. Then the regularization parameter is set to $0.01\Vert \widehat C_{\Upsilon\Upsilon}\Vert_{\text{HS}} T^{-\widetilde \rho/(\widetilde \rho+2)}$, where $\Vert \cdot \Vert_{\text{HS}}$ denotes the Hilbert-Schmidt norm. This approach may provide practical insights into the magnitude of $\rho$, thereby allowing us to choose $\tau$ without violating the assumption. We investigate the finite-sample performance of our estimator computed with this regularization parameter in Section (ref).
We first estimate the impulse responses of PAYEMS to shocks given to the sentiment distribution using our benchmark model in \hyperref[{eq: model: benchmark}]{\tagform@{\ref*{eq: model: benchmark}}} for $ h \in\{1,2,\ldots,12\}$, without a structural interpretation. We let $\zeta$ be the shocks considered in Figure (ref), with their norms normalized to one. The normalization is taken to ensure reasonable comparison between their effects. The estimation results are reported in Figure (ref), where solid lines indicate $\{ \langle \beta_h, \zeta \rangle \}_{h=1}^{12}$ in \hyperref[{eq: sirf}]{\tagform@{\ref*{eq: sirf}}} and the shaded area represents their 90% confidence intervals computed using the local asymptotic normality result in Theorem (ref). As mentioned earlier, the first two shocks are associated with positive and negative shifts in the sentiment distribution. Meanwhile, in the last two columns, we consider cases where sentiment about the economy exhibits greater dispersion, with and without a negative shift, respectively. The responses in Figure (ref) have different vertical scales, which may require caution in their interpretation.
Figure (ref) suggests that the effects of the distributional shocks tend to reach their peak after 6 to 8 months, and their magnitude tends to decrease as the forecasting horizon increases. In particular, when the average sentiment about the economy becomes more optimistic (resp.\ pessimistic), payroll growth appears to significantly increase (resp.\ decrease), especially during the first six months. The change in disagreement levels in sentiment appears to have a smaller effect on PAYEMS growth compared to locational shifts in the distribution.
Lastly, we study if our estimation approach produces any improvement compared to the existing estimators, by using the empirical median absolute prediction error (MAPE). The MAPEs are computed based on a rolling window with two different test sets, selected to include approximately 60% and 70% of the observations from the total sample. In addition to ours, denoted by SCInv, we consider two alternative estimators for comparison: PCA-FR and PCA-SVAR. The PCA-FR is obtained by applying \citepos{SHIN2009} estimation approach to our benchmark model. The PCA-SVAR is computed using the standard estimation methods developed for the recursive SVAR model in a finite dimensional setting, after pre-applying the FPCA-based dimension reduction to $X_t$.
Table (ref) summarizes MAPEs. To mitigate a potential effect of different choices of regularization parameters on estimation results, we fix $\tau$ so that all the estimators select the first two score functions, regardless of the forecasting horizon. This choice is based on empirical data, which suggests that approximately 99% of the variations are explained by the first two scores. Overall, it seems that our estimator outperforms the PCA-SVAR in terms of smaller MAPEs, and the superior performance is more significantly observed as the forecasting horizon increases. An interesting observation is that the PCA-FR approach also outperforms the PCA-SVAR, although both estimators are computed with the same empirical eigenelements, and thus the PCA-FR can be understood as a single-equation estimation of SVAR models. This may be attributed to the fact that, in the PCA-SVAR estimation, the sentiment quantile curve serves as both the target (i.e., dependent) and prediction (i.e., explanatory) variables, whereas it is only used as a predictor in the PCA-FR approach. Therefore, the PCA-SVAR estimator is likely to be subject to a larger bias associated with \textsl{double truncation}. Meanwhile, in this particular example, the two functional approaches, the SCInv and the PCA-FR, produce similar MAPEs regardless of forecasting horizons.
In this section, we employ \citepos{IR2021} functional monetary policy shocks and study the monetary policy impact on the inflation growth by using the proposed estimator. The data span the period from January 1995 to June 2016 and the functional shock is reported in Figure (ref). Under Assumption I in IR2021, the slope coefficient of the shock ($\beta_h$ in our notation) can be understood as the impulse response to monetary policy shocks, with the effect represented by a function. In this section, we overall follow IR2021. The control vector $\mathbf w_t$ is given by the set of the first two lagged dependent variables. $y_t$ is given by the inflation growth rate. The impulse response $\{ \langle \beta_h,\zeta\rangle \}_{h=1}^{20}$ is estimated with $\zeta$ corresponding to shocks observed on three specific dates of interest identified by the authors: 9/1998, 2/1999, and 1/2007. Note that the shocks and the changes in yield curves induced by them in the current study may differ slightly from those in IR2021 because we do not represent them into level and curvature components.
The estimated impulse responses and shifts in the yield curve caused by the shocks are reported in Figure (ref). The first observation is that the response seems to depend on the shape of the shock. For example, in September 1998 and January 2007, the yield curves experience decreases in overall maturities. Their effects reported in Figures (ref) and (ref) align with economic theory, which suggests that decreases in interest rates for all maturities lead to higher inflation growth. On the other hand, if the shock raises interest rates in general (e.g., February 1999), inflation tends to respond negatively, with the effect becoming more pronounced over a longer horizon. Although such a similar observation can be found in IR2021, it is worth noting that the results in Figure (ref) are not directly comparable with those reported in IR2021. This is because of the nonparametric nature of our estimation approach, which does not require a parametric restriction on either the parameter of interest or the functional variables. Therefore, our estimator should be understood as a complement to the estimator considered by the authors.
In this section, we study finite sample performance of estimators studied in Section (ref) using Monte Carlo experiments. Throughout the section, we consider the case with $T=250$ and $T=500$, and the total number of replications is set to 1,000.
In our simulation, the variable $X_t $ is designed to mimic the economic sentiment quantile functions in Section (ref). To this end, let $X_t = \sum_{j=1} ^{31} x_{j,t}\xi_j$ where $\xi_j$ denotes the $j$-th eigenvector of $X_t$'s variance and the $j$-th coordinate process $x_{j,t} $ is given by $ \langle X_t, \xi_j\rangle$. Assume that $x_{j,t}$ follows AR(1) such that
for $j=1, \ldots, 31$, where $e_{j,t}\sim_{iid} N(0,1)$. The parameter values $\alpha_j ^x$ and $\sigma_j$ are set to the estimates obtained from the economic sentiment quantile curves, with $c_{e,j}$ set to one for all $j$. Later in the section, we keep or exaggerate the variance of each coordinate process by setting $c_{e,j} = c_{e,-1} 1\{ j\neq 1\} + 1\{j=1\}$ and considering two different values of $c_{e,-1}$: (i) $c_{e,-1} = 1$ and (ii) $c_{e,-1} = 4$. In the latter case, the ordered eigenvalues (from largest to smallest) are scaled up, except for the first one, without altering their order. Hence, $\lambda_j$ exhibits a significantly slower diminishing tendency for $j$ that is not large, compared to the former case.
We assume $\{ y_t, X_t \}_{t=1} ^T$ follows the DGP such that
where $u_{t}\sim_{iid} N(0,1)$ and $\beta_h(\cdot) = \sum_{j=1} ^{J} \beta_{j,h} \xi_j(\cdot)$. In our empirical data, the first two coordinate processes ($x_{1,t}$ and $x_{2,t}$) explain more than 99% of the variations. Thus, for each $h$, the parameters $\{\alpha_h, \beta_{1,h}, \beta_{2,h}, \sigma_{u,h}\}$ are replaced by the estimates from the empirical data when $c_u=1$ and $J=2$. In the simulation, we increase $J$ to 31 and let the remaining coefficients $\{\beta_{j+2,h}\}_{j=1} ^{29}$ be given by $ \min \{ |\beta_{1,h}|, |\beta_{2,h}| \} \times 0.7^j$. This simulation design allows us to avoid the inverse problem in obtaining the coefficient estimates while keeping most variations in $X_t$, so that realizations from the DGP can mimic the empirical data. In the simulation, the constant $c_u$ is set to 0.5. We note that the coefficients in this section should not be interpreted as the coefficient estimates in Section (ref).
We consider three estimators: SCInv, SCInv$_1$ and PCA-FR. SCInv and SCInv$_1$ are differentiated in the choice of the regularization parameter. SCInv is computed as in Section (ref). The second estimator SCInv$_1$ and also \citepos{SHIN2009} PCA-FR are computed by using the AIC criterion detailed in SHIN2009.
The estimation results are summarized in Table (ref) for three forecasting horizons: $h=1,3,5$. We report the bias (Bias) and variance (Var) of each estimator relative to those computed using SCInv. Overall, the proposed estimator based on the naive choice of $\rho$ (SCInv) tends to produce the smallest bias and variance, particularly when $c_{e,-1}$ is large, which is related to the relative decay rate of $\lambda_j$. Meanwhile, the two estimators based on the same regularization parameter, SCInv$_{1}$ and PCA-FR, perform similarly.
We then study Theorem (ref) by the means of local asymptotic normality to study the impulse responses $\{\langle \beta_h, \zeta \rangle; h=1,3,5 \}$, where $\zeta$ is set to a constant function for simplicity. Table (ref) reports 95% coverage probabilities computed with SCInv for each forecasting horizon. Overall, the coverage probability is very close to the nominal level in both sample sizes. As discussed in Section (ref), to the best of the authors' knowledge, the assumptions on $\rho$, $\varsigma$ and $\delta$ in Theorem (ref) are not directly testable in practice, and thus the asymptotic bias terms $\widehat \Theta_{2A}$ and $\widehat \Theta_{2B}$ may not be asymptotically negligible. Nevertheless, the simulation results reported in Table (ref) suggest that those asymptotic bias would be small and thus do not distort testing results based on the asymptotic approximation in Theorem (ref).
Figure (ref) reports the pointwise estimates of the functional coefficient at $h=1$ and the impulse response estimates for $h=1,\ldots, 6$ when $c_{e,-1}=1$. The simulation results for $c_{e,-1}=4$ are similar; thus, we omit the figures to save space. The reported coefficient estimates are obtained by averaging the estimates across 1,000 simulations. Consistent with the observations in the tables, our estimator produces both functional coefficient estimates and the pointwise impulse response estimates close to the true values, demonstrating the practical applicability of our approach.
Lastly, we examine the performance of our estimator as an estimator of the SIRF suggested in Propositions (ref) and (ref), using a small-scale Monte Carlo simulation based on the following SVAR(1) process:
where $(u_{1,t}, u_{2,t}, u_{3,t} )' \sim_{iid} \mathcal N(0, \text{diag}(\sigma_1^2, \sigma_2^2, \sigma_3^2))$ and $x_{j,t}$ denotes the $j$-th coordinate process of economic sentiment quantiles. The functional predictor $X_t$ is assumed to be generated as follows.
where $ u_{j+1, t} \sim_{iid} N( {0}, \sigma_j ^2 ) $ across $j$ and $t$, and $\sigma_j ^2 = \sigma_3 0.8^j$ for $j = 3, \ldots, 31$, and $\xi_j$ are the eigenvectors of $X_t$'s covariance. The parameters and $\{\xi_j\}_{j=1} ^{31}$ are estimated from empirical data as in Section (ref). As before, this setup is designed to keep the structure of $X_t$ whose largest variations are determined by the first two coordinate processes. Then, in the simulation, we replace the variance structure of $\mathbf {u}_t= (u_{1,t},u_{2,t}, \ldots , u_{32,t})'$ with $c_{\ast}\text{diag}(c_1 \sigma_1^2, \sigma_2^2,\sigma_3^2, \ldots, \sigma_{32}^2 )$, where $c_{\ast}$ is the normalizing constant that allows us to keep the norm of the variance of $\mathbf {u}_t$ to be equal to one. The other constant $c_1$ determines the relative magnitude of the structural error associated with $y_t$, and thus it is inversely related to the signal-to-noise ratio. We consider three values of $c_1$: 1, 0.5, and 0.2. Along with our estimators, the SCInv and the SCInv$_{1}$, we consider the PCA-SVAR for comparison. As the PCA-SVAR produces the same estimation results with the PCA-FR in this setup, we omit the latter.
Table (ref) summarizes simulation results. Overall, the SCInv produces the smallest bias at the cost of a relatively large variance. As observed previously, the other two estimators, the SCInv$_1$ and the PCA-SVAR, report similar estimation results with each other, regardless of the value of $c_1$ and the sample size, as the difference between the two regularization schemes is not significant in this simulation design. Figure (ref) reports the true functional coefficient and its estimates (averaged across 1,000 replications) for each value of $c_1$ when $h=1$. In the figure, it is noticeable that the PCA-SVAR estimator tends to get closer to the true functional coefficient as $c_1$ decreases. On the other hand, our estimator, the SCInv (dotted), produces very robust estimation results close to the true regardless of the value of $c_1$.
In this paper, we study impulse response analysis with functional predictors and other scalar-valued covariates and propose new estimation and inference methodologies. We show that the proposed estimator allows for an interesting interpretation as a structural impulse response in some special cases. In our empirical application, we study how economic variables respond to a certain distributional shock on sentiment. The results are consistent with existing observations and show negative (resp.\ positive) responses of economic growth variables when sentiment distributions shift to the left (resp.\ right). Monte Carlo simulation results also confirm our theoretical findings.
\makeatletter \def\@seccntformat#1{ \csname the#1\endcsname.\quad } \makeatother