EconBase
← Back to paper

From rotational to scalar invariance: Enhancing identifiability in score-driven factor models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

65,777 characters · 13 sections · 38 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

From rotational to scalar invariance:\ Enhancing identifiability in score-driven factor models

{10pt} {10pt}

\newgeometry{left=2.5cm,right=2.5cm,top=1cm, bottom=2cm}

center[center omitted — 45 chars of source]
abstractWe show that, for a certain class of scaling matrices including the commonly used inverse square-root of the conditional Fisher Information, score-driven factor models are identifiable up to a multiplicative scalar constant under very mild restrictions.\ This result has no analogue in parameter-driven models, as it exploits the different structure of the score-driven factor dynamics.\ Consequently, score-driven models offer a clear advantage in terms of economic interpretability compared to parameter-driven factor models, which are identifiable only up to orthogonal transformations.\ Our restrictions are order-invariant and can be generalized to score-driven factor models with dynamic loadings and nonlinear factor models. We test extensively the identification strategy using simulated and real data.\ The empirical analysis on financial and macroeconomic data reveals a substantial increase of log-likelihood ratios and significantly improved out-of-sample forecast performance when switching from the classical restrictions adopted in the literature to our more flexible specifications. Keywords:\ Identification, Factor Models, Score-driven Models, Forecasting.\\ JEL codes:\ C51, C58.

\newgeometry{left=2.5cm,right=2.5cm,top=2cm, bottom=2cm}

Introduction

Given the increasing interconnectedness of the global economy, the dynamics of a panel of economic and financial time series can often be described by a set of few common factors.\ For this reason, the use of dynamic factor models has become increasingly popular in the econometric literature.\ In a dynamic factor model, the common factors may be either observable or latent.\ Assuming latent factors is regarded as less restrictive because it sidesteps the challenge of selecting the common factors from a vast array of potential explanatory variables.\ On the other hand, the use of unobservable factors generates identification issues, the solution of which poses significant interpretability problems.\ For instance, the estimated factors might be order-dependent or defined up to rotations, thereby complicating their economic or financial interpretation.

In this paper, we contribute to the literature on factor models identification. In particular, we focus on factor models where the factor dynamics are driven by the score of the conditional observation density.\ Using the language of Cox, these models are called observation-driven, meaning that the main source of variation of the factor dynamics are the past observations.\ In contrast, in parameter-driven models, the factor dynamics depend on their own source of uncertainty; see, e.g., geweke1977dynamic, chamberlain1983funds, StockWatson2011, doz2012quasi.\ Examples of score-driven factor models are given in creal2014observation, oh2018time, artemova, among others.\ Score-driven factor models are particular instances of the general class of score-driven models proposed by harvey2013dynamic and creal2013generalized.\ The main advantage of score-driven factor models with respect to parameter-driven specifications is that the exact likelihood can always be written in closed form.\ Moreover, the law of motion of the factors, being driven by the score of a non-normal conditional density, is naturally robust to the extreme observations which may occur during periods of financial and/or economic distress.

The main contribution of this paper is to show that score-driven factor models can be identified under significantly milder restrictions compared to parameter-driven factor models.\ This is due to the fact that, when subject to a linear non-singular transformation, the law of motion of the factors in a score-driven model is invariant only under two specific choices of the scaling matrix used to normalize the score.\ The first choice corresponds to the inverse of the conditional Fisher information matrix, while the second one corresponds to the identity matrix.\ In these two cases, multiplying the factors by a non-singular matrix leads to new factors following a score-driven recursion of the same form.\ Therefore, both the static and time-varying parameters are not identifiable, as in parameter-driven factor models.\ However, when the scaling matrix is set differently, for example as the commonly used square-root inverse of the conditional Fisher information, multiplying the factors by a non-singular matrix results in a new recursion that cannot be represented as the original score-driven recursion.\ In other words, under affine non-singular transformations, the original score-driven law of motion is not generally preserved.\ Thus, for this class of normalization matrices, score-driven factor models behave differently compared to parameter-driven factor models, and the lack of invariance can be exploited to identify the static parameters.

When the score is scaled through the class of normalization matrices for which the score-driven dynamics are not preserved, the factors and the loading matrix are identifiable up to a common multiplicative constant.\ This facilitates enormously their economic interpretability because a scalar transformation only scales the factors without altering their dynamics.\ Similarly, it leaves the relative magnitudes of the factor loadings unchanged, so that it is simple to quantify the impact each factor has on each variable.\ By comparison, parameter-driven factor models are identifiable only up to orthogonal transformations, and any identifying restriction amounts to selecting a specific rotation of the factors.\ Another implication of our results is that the estimated factors are order-invariant, meaning that they are independent from the order of the variables within the panel of time-series. This is an important advantage compared to other identification approaches in likelihood-based methods imposing a priori dependencies structures between the factors and the observable data relying on the ordering of the time-series.\

A second contribution of the paper is to show that the same identification scheme can be used in score-driven factor models with time-varying loadings and in nonlinear factor models.\ Empirical evidence of time-varying loadings has been found in macroeconomic data (mikkelsen2019consistent, xu2022testing, hillebrand2023exchange) and financial data ( adrian2009learning, kelly2019characteristics, giglio2022factor).\ Nonlinear factor models with score-driven dynamics are widely adopted to reduce dimensionality in multivariate models.\ One example is given by the dynamic correlation model proposed in creal2011dynamic, where the authors suggest to impose a factor structure for the time-varying parameters in order to reduce the size of the parameter space.

A relevant implication of the scalar invariance of score-driven factor models is that the static and dynamic parameters can be identified without imposing constraints on the factor loadings.\ The main identifying assumption needed is the diagonal structure of the matrix multiplying the scaled score, which is a standard restriction adopted in score-driven models.\ This leads to a higher flexibility in the specification of the model compared to parameter-driven factor models, where the factor loadings are generally constrained in order to guarantee identification.\ The higher flexibility is shown in our empirical application, where we find a significant increase in likelihood ratios when moving from the classical restrictions adopted in the literature to our specification.\ We also show that the factors extracted through our approach have a straightforward economic interpretation and produce superior out-of-sample forecasts of macroeconomic variables.

Another work addressing the problem of identifying the static and time-varying parameters in score-driven factor models is artemova.\ The main difference between our approach and theirs is that artemova normalizes the score through the inverse conditional Fisher information.\ This normalization preserves the structure of the score-driven law of motion when the time-varying parameters are multiplied by a non-singular matrix, leading to the same rotational indeterminacy found in parameter-driven factor models.\ Indeed, the identifying restrictions used in artemova coincide with the orthogonality conditions on the factor loading matrix suggested by bai2012statistical, which pick a specific rotation of the factors.\ In contrast, by normalizing the score through the inverse square-root of the conditional Fisher information, we exploit the different transformation properties of the score-driven law of motion in order to circumvent rotational indeterminacy and identify the model up to a common scalar constant.

The rest of the paper is organized as follows.\ In Section (ref), we discuss the main differences between parameter-driven and score-driven factor models when it comes to the identification of the static and dynamic parameters.\ Section (ref) illustrates the main result of the paper, i.e., it provides a set of restrictions enabling the identification of score-driven factor models up to a common scalar constant.\ In Section (ref), we show that the same restrictions are sufficient to identify the model in the presence of time-varying loadings and in nonlinear factor structures.\ Several Monte Carlo results supporting our conclusions are reported in Section (ref), while Section (ref) presents the empirical results.\ Finally, Section (ref) concludes.\ The proofs of the main results are collected in the Appendix section.

Identification in parameter-driven versus score-driven factor models

We start by illustrating the different behavior of parameter-driven and score-driven models when scaling the factors through a non-singular square matrix.\ Let us first consider the following parameter-driven factor model

align[align omitted — 179 chars of source]

where $\bm{y}_t$ is an $n\times 1$ vector, $\bm{\alpha}_t$ is a vector of $r<n$ unobservable common factors and $\bm{\Lambda}$ is an $n\times r$ matrix of factor loadings.\ The two white noises $\{\bm{\epsilon}_t\}_{t\in\mathbb{Z}}$, $\{\bm{\eta}_t\}_{t\in\mathbb{Z}}$ are independent, with covariance matrices $\bm{H}\in\mathbb{R}^{n\times n}$, $\bm{Q}\in\mathbb{R}^{r\times r}$, respectively.\ Let $\bm{T}\in\mathbb{R}^{r\times r}$ be a non-singular square matrix.\ The above model can be re-parameterized as follows

align[align omitted — 168 chars of source]

where $\bm{\Lambda}^{*}=\bm{\Lambda}\bm{T}$, $\bm{\alpha}_t^*=\bm{T}^{-1}\bm{\alpha}_t$, $\bm{\omega}^* = \bm{T}^{-1}\bm{\omega}$, $\bm{\Phi}^{*}=\bm{T}^{-1}\bm{\Phi}\bm{T}$, while $\bm{\eta}_t^*$ has covariance $\bm{Q}^*=\bm{T}^{-1}\bm{Q}\bm{T}^{-1\prime}$.\ Since the two parameterizations are observationally equivalent, the model is not identifiable without prior restrictions.\ Specifically, we need $r^2$ restrictions, equal to the number of degrees of freedom of $\bm{T}$, in order to fully identify the model parameters.\ A common identification strategy consists in restricting the covariance matrix of $\bm{\alpha}_t$, or that of $\bm{\eta}_t$, to be the identity matrix.\ This restriction leaves the factors identifiable up to an orthogonal transformation.\ We thus need to fix the remaining $r(r-1)/2$ free parameters, i.e., we need to specify the particular rotation being estimated.\ This can be done by imposing some structure on the factor loading matrix.\ Several types of restrictions are possible, some of which are order-invariant, while others depend on the ordering of the variables; see, e.g., harvey_1990 and bai2012statistical.\

Let us now consider the identification of the static parameters in score-driven factor models.\ Let $\bm{\mathcal{F}}_{t}=\sigma(\bm{y}_t,\bm{y}_{t-1},\dots,\bm{y}_1)$ denote the information filtration generated by the process $\{\bm{y}_t\}_{t\in\mathbb{N}}$.\ Let us consider a score-driven factor model of the following form

align[align omitted — 165 chars of source]

where $\bm{\epsilon}_t$ is a white noise with covariance $\bm{\Sigma}\in\mathbb{R}^{n\times n}$, $\bm{s}_t = \bm{S}_t \bm{\nabla}_t$, and

equation[equation omitted — 143 chars of source]

is the score of the predictive likelihood.\ The scaling matrix $\bm{S}_t$ is set as $\bm{S}_t =\left(\bm{\mathcal{I}}_{t|t-1}\right)^{-\beta}$, $\beta\in [0,1]$, where $\bm{\mathcal{I}}_{t|t-1}$ denotes the conditional Fisher information, defined as

equation[equation omitted — 104 chars of source]

Common choices of the exponent $\beta$ are $\beta=\left\{0,\frac{1}{2},1\right\}$.\ The time-varying parameter $\{\bm{f}_t\}_{t\in\mathbb{N}}$ is predictable with respect to the filtration $\bm{\mathcal{F}}_{t}$, implying that the predictive likelihood coincides with the conditional observation density $p(\bm{y}_t|\bm{f}_t)$ determined by the distribution of the idiosyncratic noise $\bm{\epsilon}_t$.\ The latter is left unspecified at the moment because the results we recover are independent from the choice of such a distribution.\ The score-driven recursion allows writing the likelihood in closed form via the prediction error decomposition, in a similar fashion to GARCH-type models, and to infer the static parameters through standard numerical optimization methods.\ Moreover, the update of the time-varying parameters based on the score provides robust estimates in the case of fat-tailed distributions; see also d2024dynamic, opschoor2024conditional and blasques2024maximum for some recent applications of score-driven models.

It is immediate to prove the following result.

propLet $\bm{T}\in\mathbb{R}^{r\times r}$ be non-singular and let us set $\overline{\bm{\Lambda}}=\bm{\Lambda}\bm{T}$, $\overline{\bm{f}}_t=\bm{T}^{-1}\bm{f}_t$. The score-driven factor model in Equations (ref), (ref) can be re-parameterized as follows: \begin{align} \bm{y}_t &= \overline{\bm{\Lambda}}\ \overline{\bm{f}}_t + \bm{\epsilon}_t\\ \overline{\bm{f}}_{t+1} &= \overline{\bm{c}} +\overline{\bm{A}}\ \left(\overline{\bm{\mathcal{I}}}_{t|t-1}\right)^{-\beta}(\bm{T}^{-1+\beta})^{\prime}\overline{\bm{\nabla}}_t + \overline{\bm{B}}\ \overline{\bm{f}}_t \end{align} where $\overline{\bm{c}}=\bm{T}^{-1}\bm{c}$, $\overline{\bm{A}}=\bm{T}^{-1}\bm{A}\bm{T}^{\beta}$, $\overline{\bm{B}}=\bm{T}^{-1}\bm{B}\bm{T}$, $\overline{\bm{\nabla}}_t = \left[\frac{\partial\log p(\bm{y}_t|\bm{\mathcal{F}}_{t-1},\overline{\bm{f}}_t)}{\partial\overline{\bm{f}}_t}\right]'$, $\overline{\bm{\mathcal{I}}}_{t|t-1}=\mathbb{E}[\overline{\bm{\nabla}}_t\overline{\bm{\nabla}}_t'|\bm{\mathcal{F}}_{t-1}]$.

This result shows that the dynamics of the transformed process $\{\overline{\bm{f}}_t\}_{t\in\mathbb{N}}$ cannot be written as a standard score-driven recursion of the form given in Equation (ref).\ This is because the two matrices $\left(\bm{\mathcal{I}}_{t|t-1}\right)^{-\beta}$ and $(\bm{T}^{-1+\beta})'$ do not generally commute, preventing the term $(\bm{T}^{-1+\beta})'$ in Equation (ref) from being absorbed into the matrix $\overline{\bm{A}}$ that multiplies the scaled score. Therefore, score-driven factor models behave differently compared to parameter-driven factor models, where, for any choice of $\bm{T}$, the transformed factors follow the same linear process describing the original factors.\

Among all the possible values of $\beta\in [0,1]$, there exist only two specific choices for which the term $(\bm{T}^{-1+\beta})'$ can be absorbed into the matrix $\overline{\bm{A}}$ for any non-singular matrix $\bm{T}\in\mathbb{R}^{n\times n}$.\ We discuss these two cases below.

itemize• Case $\beta=0$.\ We have $\bm{T}^{\beta}=\bm{I}_{r\times r}$ and $\left(\overline{\bm{\mathcal{I}}}_{t|t-1}\right)^{-\beta}=\bm{I}_{r\times r}$. Therefore, $\overline{\bm{A}}=\bm{T}^{-1}\bm{A}\bm{T}^{\beta}$ can be re-defined as $\overline{\overline{\bm{A}}}=\overline{\bm{A}}(\bm{T}^{-1})'= \bm{T}^{-1}\bm{A}(\bm{T}^{-1})'$, and Equation (ref) becomes \begin{equation} \overline{\bm{f}}_{t+1} = \overline{\bm{c}} + \overline{\overline{\bm{A}}}\ \overline{\bm{\nabla}}_t + \overline{\bm{B}}\ \overline{\bm{f}}_t, \end{equation} which has the same form of Equation (ref) for $\beta=0$. • Case $\beta=1$. We have $\overline{\bm{s}}_t=\left(\overline{\bm{\mathcal{I}}}_{t|t-1}\right)^{-1}\overline{\bm{\nabla}}_t$, and therefore Equation (ref) becomes \begin{equation} \overline{\bm{f}}_{t+1} = \overline{\bm{c}} +\overline{\bm{A}}\ \left(\overline{\bm{\mathcal{I}}}_{t|t-1}\right)^{-1}\overline{\bm{\nabla}}_t + \overline{\bm{B}}\ \overline{\bm{f}}_t, \end{equation} which has the same form of Equation (ref) for $\beta=1$.

In these two special cases, score-driven factor models are subject to the same identification issues of parameter-driven factor models because multiplying the factors by an arbitrary non-singular matrix $\bm{T}$ results in a recursion of similar form.\ For example, artemova considers the case $\beta=1$, and uses identifying restrictions similar to those applied in parameter-driven models to infer the static parameters.

Scalar identification

Let us now consider the case $\beta\in (0,1)$.\ For example, we may set $\beta=\frac{1}{2}$, meaning that the score is normalized by the inverse square-root of the conditional Fisher information.\ This normalization is considered in several score-driven specifications; see creal2014observation among others.\ Equation (ref) becomes

equation[equation omitted — 234 chars of source]

which has a different structure compared to the original factor dynamics in Equation (ref).\ This result can be used in order to identify the static parameters under much milder restrictions compared to those required in parameter-driven models.\ In particular, such restrictions must ensure that the two matrices $\left(\overline{\bm{\mathcal{I}}}_{t|t-1}\right)^{-\frac{1}{2}}$ and $(\bm{T}^{-\frac{1}{2}})'$ do not commute, so that $(\bm{T}^{-\frac{1}{2}})'$ cannot be absorbed in the matrix $\overline{\bm{A}}$.\ These restrictions are summarized in the following assumptions:

assThe matrix $\bm{A}$ is diagonal, with non-zero diagonal elements.
assThe conditional Fisher information matrix $\bm{\mathcal{I}}_{t|t-1}$ is not block diagonal.

The restriction in Assumption (ref) is very common in the score-driven literature.\ Assumption (ref) is generally verified when the factor loading matrix $\bm{\Lambda}$ and/or the covariance matrix $\bm{\Sigma}$ are unconstrained.\ To see this, observe that, if $\bm{\epsilon}_t$ has an elliptical distribution, e.g., it is a Student-$t$, the Fisher information matrix is proportional to $\bm{\Lambda}'\bm{\Sigma}^{-1}\bm{\Lambda}$; see Appendix (ref).\ Therefore, $\bm{\mathcal{I}}_{t|t-1}$ is generally not block diagonal when no prior restrictions on $\bm{\Lambda}$ and/or $\bm{\Sigma}$ are imposed.\ For example, $\bm{\mathcal{I}}_{t|t-1}$ is not block diagonal when $\bm{\Lambda}$ is unrestricted and $\bm{\Sigma}$ is diagonal.

thmLet $\beta\in (0,1)$.\ Under Assumptions (ref) and (ref), $\left(\overline{\bm{\mathcal{I}}}_{t|t-1}\right)^{-\beta}$ and $(\bm{T}^{-1+\beta})'$ commute if only if $\bm{T}$ is scalar, i.e., $\bm{T}=q\bm{I}_{r\times r}$, $q\in\mathbb{R}$.

The above result shows that, when the score is scaled through, e.g., the inverse square-root of the conditional Fisher information, the static parameters and the common factors are identifiable up to a common multiplicative scalar term under fairly general assumptions on the structure of the parameter space.\ Compared to parameter-driven models, where restricting the covariance of the factors to be the identity matrix leaves them defined up to a rotation, here restricting the matrix $\bm{A}$ to be diagonal identifies the factors and the static parameters up to a common multiplicative term, thus avoiding rotational indeterminacy.\ Clearly, this improves the interpretability of the model estimates because the effect of the common multiplicative factor is just to scale the factor dynamics.

To remove the indeterminacy associated with the scalar factor, it is enough to fix one parameter.\ For example, we may fix one of the entries of either $\bm{\Lambda}$ or $\bm{c}$.\ Specifically, we set

equation[equation omitted — 29 chars of source]

which implies $q=1$.\ This choice is more convenient than setting a constraint for the factor loadings because the latter would break the structure of $\bm{\Lambda}$ when the order of the time-series changes.\ On the contrary, setting $\bm{c}_1 = 1$ leads to order-invariant restrictions.\ This result holds for the wide class of elliptical distributions, as shown in the next proposition.

propLet us assume that $p(\bm{y}_t| \bm{\mathcal{F}}_{t-1},\bm{f}_t)$ is in the class of elliptical distributions, that is, $p(\bm{y}_t|\bm{\mathcal{F}}_{t-1},\bm{f}_t)=(\det{\bm{\Sigma}})^{-\frac{1}{2}}\psi ((\bm{y}_t-\bm{\Lambda}\bm{f}_t)'\bm{\Sigma}^{-1}(\bm{y}_t-\bm{\Lambda}\bm{f}_t))$, where $\psi$ is a scalar function.\ We further assume that $\bm{\Lambda}$ is unconstrained.\ Then the likelihood is invariant for transformations of the form $\bm{y}_t\to \bm{P}\bm{y}_t$, where $\bm{P}$ is a permutation matrix.\ Furthermore, the filtered dynamics of $\bm{f}_t$ are not affected by the permutation if the same starting value is used to initialize the filter of the permuted series.

Generalizations

Score-driven factor loadings

One of the advantages of the normalization based on the inverse square-root of the conditional Fisher information is that, under the same restrictions specified in Assumptions (ref), (ref), it is possible to resolve the rotational indeterminacy also in factor models with time-varying loadings.\ Let us consider the following dynamic factor model

equation[equation omitted — 93 chars of source]

where $\bm{\Lambda}_t$, $\bm{g}_t$ are both predictable with respect to $\bm{\mathcal{F}}_t$.\ Let $\bm{l}_t=\text{vec}(\bm{\Lambda}_t)$ be an $n r\times 1 $ vector obtained by stacking the elements of $\bm{\Lambda}_t$ in a column.\ The score-driven dynamics of $\bm{l}_t$ and $\bm{g}_t$ are given by

align[align omitted — 204 chars of source]

where

align[align omitted — 397 chars of source]

The two normalization matrices $\bm{S}_t^{(l)}$, $\bm{S}_t^{(g)}$ are set as follows

equation[equation omitted — 173 chars of source]

where $\alpha,\beta\in\ [0,1]$, and $\bm{\mathcal{I}}_{t|t-1}^{(l)}$ and $\bm{\mathcal{I}}_{t|t-1}^{(g)}$ are two conditional Fisher Information matrices

equation[equation omitted — 252 chars of source]

Proposition (ref) can be generalized as follows:

propLet $\bm{T}\in\mathbb{R}^{r\times r}$ be non-singular and let us set $\overline{\bm{\Lambda}}_t=\bm{\Lambda}_t\bm{T}$, $\overline{\bm{l}}_t=\emph{vec}(\bm{\overline{\Lambda}}_t)=(\bm{T}'\otimes \bm{I})\emph{vec}(\bm{\Lambda}_t)$, $\overline{\bm{g}}_t=\bm{T}^{-1}\bm{g}_t$. The score-driven factor model with time-varying loadings in Equations (ref), (ref), (ref) can be re-parameterized as follows: \begin{align} \bm{y}_t &= \overline{\bm{\Lambda}}_t\ \overline{\bm{g}}_t + \bm{\epsilon}_t\\ \overline{\bm{l}}_{t+1} &= \overline{\bm{c}}^{(l)}+ \overline{\bm{A}}^{(l)} \left(\overline{\bm{\mathcal{I}}}_{t|t-1}^{(l)}\right)^{-\alpha}(\bm{T}^{-\alpha+1}\otimes \bm{I})\overline{\bm{\nabla}}_t^{(l)} + \overline{\bm{B}}^{(l)} \overline{\bm{l}}_t \\ \overline{\bm{g}}_{t+1} &=\overline{\bm{c}}^{(g)}+\overline{\bm{A}}^{(g)} \left(\overline{\bm{\mathcal{I}}}_{t|t-1}^{(g)}\right)^{-\beta}(\bm{T}^{-1+\beta})^{\prime}\overline{\bm{\nabla}}_t^{(g)} +\overline{\bm{B}}^{(g)}\overline{\bm{g}}_t \end{align} where $\overline{\bm{c}}^{(l)}=(\bm{T}'\otimes \bm{I})\bm{c}^{(l)}$, $\overline{\bm{A}}^{(l)}=(\bm{T}'\otimes \bm{I})\bm{A}^{(l)}(\bm{T}'\otimes \bm{I})^{-\alpha}$, $\overline{\bm{B}}^{(l)}=(\bm{T}'\otimes \bm{I})\bm{B}^{(l)}(\bm{T}'\otimes \bm{I})^{-1}$, $\overline{\bm{c}}^{(g)}=\bm{T}^{-1}\bm{c}^{(g)}$, $\overline{\bm{A}}^{(g)}=\bm{T}^{-1}\bm{A}^{(g)}\bm{T}^{\beta}$, $\overline{\bm{B}}^{(g)}=\bm{T}^{-1}\bm{B}^{(g)}\bm{T}$, and \begin{align} \overline{\bm{\nabla}}_t^{(l)}&=\left[\frac{\partial\log\mathcal{L}(\bm{y}_t|\bm{\mathcal{F}}_{t-1},\overline{\bm{l}}_t,\overline{\bm{g}}_t)}{\partial\bm{\overline{l}}_t}\right]', \quad \overline{\bm{\mathcal{I}}}_{t|t-1}^{(l)}= \mathbb{E}[\overline{\bm{\nabla}}_t^{(l)}\overline{\bm{\nabla}}_t^{(l)\prime}|\bm{\mathcal{F}}_{t-1}]\\ \overline{\bm{\nabla}}_t^{(g)}&=\left[\frac{\partial\log\mathcal{L}(\bm{y}_t|\bm{\mathcal{F}}_{t-1},\overline{\bm{l}}_t,\overline{\bm{g}}_t)}{\partial\bm{\overline{g}}_t}\right]', \quad \overline{\bm{\mathcal{I}}}_{t|t-1}^{(g)}= \mathbb{E}[\overline{\bm{\nabla}}_t^{(g)}\overline{\bm{\nabla}}_t^{(g)\prime}|\bm{\mathcal{F}}_{t-1}] \end{align}

Also in this case, the dynamics of the transformed process $\overline{\bm{g}}_t$ cannot be represented as a score-driven recursion.\ In particular, for $\beta\in (0,1)$, it is still true that the law of motion of $\overline{\bm{g}}_t$ has a different structure relative to Equation (ref) due to the non-commutativity of the two matrices $(\bm{T}^{-1+\beta})'$ and $\left(\overline{\bm{\mathcal{I}}}_{t|t-1}^{(g)}\right)^{-\beta}$.\ Compared to the static case presented in the previous section, the main difference here is that the Fisher information $\overline{\bm{\mathcal{I}}}_{t|t-1}^{(g)}$ is generally time-varying.\ For instance, when the distribution of $\bm{\epsilon}_t$ is in the class of elliptical distributions, $\overline{\bm{\mathcal{I}}}_{t|t-1}^{(g)}$ is proportional to $\overline{\bm{\Lambda}}_t'\bm{\Sigma}^{-1}\overline{\bm{\Lambda}}_t$.\ The following two assumptions, which are the analogue of Assumptions (ref), (ref), are sufficient, also in this more general case, to prevent the two matrices $(\bm{T}^{-1+\beta})'$ and $\left(\overline{\bm{\mathcal{I}}}_{t|t-1}^{(g)}\right)^{-\beta}$ to commute.

assThe matrix $\bm{A}^{(g)}$ is diagonal, with non-zero diagonal elements.
assThe conditional Fisher information matrix $\bm{\mathcal{I}}_{t|t-1}^{(g)}$ is not block diagonal.
thmLet $\beta\in (0,1)$.\ Under Assumptions (ref) and (ref), $\left(\overline{\bm{\mathcal{I}}}_{t|t-1}\right)^{-\beta}$ and $(\bm{T}^{-1+\beta})'$ commute if only if $\bm{T}$ is scalar, i.e., $\bm{T}=q\bm{I}_{r\times r}$, $q\in\mathbb{R}$.

The main consequence of Theorem (ref) is that we can identify the model parameters up to a common multiplicative scalar term also in a dynamic setting with time-varying loadings.\ This is true because Assumptions (ref) and (ref) restrict the matrix $\bm{T}$ to be scalar, and thus also the parameters governing the dynamics of $\bm{\Lambda}_t$ can be identified up to a common scalar term.

We use the same identifying restriction presented in the static case, i.e., we fix the first element of $\bm{c}^{(l)}$

equation[equation omitted — 35 chars of source]

which implies $q=1$.\ It is worth to note that $q$ is uniquely determined independently from the normalization $\alpha\in [0,1]$ adopted to scale the score of the factor loading dynamics.\ In our simulation and empirical analysis, we set $\alpha=0$ since $\overline{\bm{\mathcal{I}}}_{t|t-1}^{(l)}$ is generally a singular matrix in models characterized by an elliptical distribution for the idiosyncratic noise $\bm{\epsilon}_t$.

Nonlinear factor models

Let $h:\mathbb{R}^r\to\mathbb{R}^d$ be a differentiable vector function. We consider the following nonlinear factor specification

align[align omitted — 253 chars of source]

where $p(\bm{y}_t|\bm{\mathcal{F}}_{t-1},\bm{h}(\bm{f}_t))$ denotes the conditional probability density of $\bm{y}_t$ given a time-varying vector $\bm{h}(\bm{f}_t)\in\mathbb{R}^d$, $\bm{\Lambda}\in\mathbb{R}^{d\times r}$, and $\bm{f_t}\in\mathbb{R}^r$ is a vector of common factors driving the dynamics of the time-varying parameters.\ Here, $\bm{s}_t$ is still the normalized score, i.e., $\bm{s}_t=\bm{S}_t\bm{\nabla}_t$, $\bm{S}_t=(\bm{\mathcal{I}}_{t|t-1})^{-\beta}$, where

equation[equation omitted — 231 chars of source]

It is straightforward to verify that multiplying $\bm{f}_t$ by a non-singular matrix results in the same transformation described in Proposition (ref).\ Therefore, we can adopt the restrictions described in Section (ref) to identify the static parameters.\ In general, the factor structure in models with nonlinear observation density has the role of reducing the number of parameters, which might otherwise increase very fast with the dimensionality.\ Some examples of nonlinear score-driven factor models are given in creal2011dynamic, who propose common factors driving the time-varying correlations in order to reduce the model dimensionality, and opschoor2021closed, who consider a copula model with dynamic common factors.\

Monte Carlo analysis

In this section, we perform an extensive Monte Carlo analysis to demonstrate that the model is identifiable under the scalar restriction discussed in Section (ref).\ To this end, we study the finite sample properties of the standard maximum-likelihood estimator in both the constant and time-varying loading cases. Throughout the study, we use the following Student-$t$ conditional density to derive the score-driven law of motion of the time-varying parameters:

equation[equation omitted — 345 chars of source]

where $\bm{\Lambda}$ is replaced by $\bm{\Lambda}_t$ in the time-varying loading case.\ Explicit expressions of the scores and conditional Fisher information for both $\bm{f}_t$ and $\bm{l}_t=\text{vec}(\bm{\Lambda}_t)$ are given in Appendix (ref).

Static loadings

We first consider a low-dimensional setting with $n=5$ time series and $r=2$ common factors.\ The parameters governing the score-driven factor dynamics are set as follows:

equation[equation omitted — 196 chars of source]

Moreover, we set $\nu=5$, $\bm{\Sigma}=\frac{1}{2}\bm{I}_{5\times 5}$, $\bm{\Lambda}_{ij}\sim\text{Uniform}(0,1)$, for $i=1,\dots,5$ and $j=1,2$.\ The first entry of $\bm{c}$ is kept fixed in order to resolve the scalar invariance and identify all the parameters.\ The covariance matrix $\bm{\Sigma}$ is scalar in this example, however no scalar restriction is imposed in the estimation, i.e., we estimate each diagonal entry of $\bm{\Sigma}$ independently from others.\ Note that, with this choice of the static parameters, our restrictions in Assumptions (ref), (ref) are satisfied.\ The initial values of the static parameters given as an input to the likelihood optimization are obtained by perturbing the true values with uniform random shifts ranging from 0 to 50%.

figure[figure omitted — 1,152 chars of source]

We generate $N=250$ replications of the model with different sample sizes $T=250,1000,4000$.\ The kernel density estimates of the maximum-likelihood estimates of $\nu$ and of each individual element of $\bm{c}$, $\bm{A}$, $\bm{B}$ are shown in Figure (ref), whereas those of $\bm{\Lambda}$ and $\bm{\Sigma}$ are reported in Appendix (ref) to save space.\ The results show that the maximum-likelihood estimator concentrates around the true values as the sample size increases, thus confirming that the model parameters are identifiable under the restriction.

We now consider a large dimensional setting with $n=100$ and $r=2$.\ The parameters $\bm{c}$, $\bm{A}$, $\bm{B}$, $\nu$ are chosen as in the previous study, whereas $\bm{\Sigma}$ and $\bm{\Lambda}$ are set as $\bm{\Sigma}=2\bm{I}_{100\times 100}$, and $\bm{\Lambda}_{ij}\sim\text{Uniform}(0,1)$, for $i=1,\dots,100$ and $j=1,2$.\ Note that the variance of the idiosyncratic noise is increased in order to achieve a signal-to-noise ratio comparable to that used in the previous setting.\ We generate $N=250$ replications of the model with $T=[250,1000,4000]$. Figure (ref) shows the kernel density estimates of the Frobenius distances between the true matrices and the maximum-likelihood estimates.\ For all model matrices, we note a significant drop in the Frobenius norms as the sample size increases, indicating that the estimates get closer to the true parameter values as the sample size increases.\

figure[figure omitted — 1,107 chars of source]

Time-varying loadings

As a second exercise, we study the finite-sample behavior of the score-driven factor model with time-varying coefficients in order to verify that the model parameters are correctly estimated under the same set of identifying restrictions.\ We consider a simulation setting with $n=5$ time-series variables and $r=2$ common factors.\ The parameters $\bm{c}^{(g)}$, $\bm{A}^{(g)}$, $\bm{B}^{(g)}$, $\bm{\Sigma}$, $\nu$ are set as in Section (ref).\ The parameters governing the loading dynamics are instead set as follows: $\bm{c}_i^{(l)}=0.1$, $\bm{A}_{ii}^{(l)}\sim \text{Uniform}(0,0.5)$, $\bm{B}_{ii}^{(l)}=0.9$, for $i=1,\dots,10$.\ In this setting, the two matrices $\bm{\Sigma}$, $\bm{\bm{B}}$ are scalar, however each diagonal element is estimated independently from others.\ In contrast, we estimate the same constant $\bm{c}_i^{(l)}$ for all the time-varying loadings.\ Other configurations are possible, for example we may estimate the same $\bm{A}_{ii}^{(l)}$ and $\bm{B}_{ii}^{(l)}$ for any loading and a different $\bm{c}_i^{(l)}$, or we may choose different $\bm{c}_i^{(l)}$, $\bm{A}_{ii}^{(l)}$ and $\bm{B}_{ii}^{(l)}$ for all loadings.\ The results we obtain in these different configurations do not different substantially from each others, and thus to save space we report here those related to the first configuration with different $\bm{A}_{ii}^{(l)}$ and $\bm{B}_{ii}^{(l)}$ and common $\bm{c}_i^{(l)}$.

The initial parameter values for the likelihood maximization are set as in Section (ref), i.e., by perturbing the true values with uniform random shifts ranging from 0 to 50%.\ The kernel density estimates of $\bm{c}^{(g)}$, $\bm{A}^{(g)}$, $\bm{B}^{(g)}$, $\nu$ are reported in Figure (ref), while the results for the remaining parameters are shown in Appendix (ref).\ We note that, similarly to the static case in Section (ref), all the static parameters are correctly estimated and the distribution of the maximum-likelihood estimator concentrates around the true values as the sample size increases.\ This confirms the results of Section (ref), showing that the model parameters are identifiable under the scalar restriction also when the factor loading matrix is time-varying.

figure[figure omitted — 1,197 chars of source]

Empirical Application

In the following empirical analyses, the main goal is to assess the potential advantages of the flexible dependence structure achieved through the introduced identification scheme to capture the dynamics of the financial and macroeconomic time series. To this end, we apply the proposed framework to two datasets: (i) a panel of 8 macro-financial time series previously applied in the literature to extract economic activity indicators (e.g., creal2014observation, artemova) and (ii) the daily returns series of 87 constituents of the S&P500.

In each application, we compare the benchmark score-driven factor model given by Equations (ref) and (ref) under the scale \(\beta=\frac{1}{2}\), fixed constant \(c_1\), and unrestricted loading matrix \(\bm{\Lambda}\), with a set of specifications with restricted loading matrix. In particular, we initially consider the common restrictions from the earlier literature to insure the identifiability of the model parameters; see, e.g., bai2012statistical and artemova2022score. In this regard, the lower-tringular structure is imposed on the first \(r\) rows of the loading matrix corresponding to the \(r\) factors, with the diagonal elements set equal to 1. To illustrate, let us consider \(N=12\) assets exposed to \(r=3\) latent factors. The resulting \(N \times r\) matrix of loadings, labeled as general lower-triangular (LT), is given by:

equation[equation omitted — 272 chars of source]

Then, we examine the loading structures in score-driven factor copula models of opschoor2021closed, ensuring the identification of the model parameters. They divide the assets into \(p\) groups according to their industry classification, assuming that all the assets within groups share the analogous dependence structure. As such, let us assume that \(N=12\) assets belong to \(p=3\) equally-sized groups. The first loading matrix is restricted to the lower-triangular form of group-specific loadings to \(r=p\) factors

equation[equation omitted — 240 chars of source]

where each \(\Tilde{\lambda}_{i,j}\), for the group \(i=1,2,3,\) and factor \(j=1,2,3,\) is a \(4 \times 1\) vector.

The second restriction implies the common exposure of all groups to the so-called common factor and individual loadings on the group-specific one for a total of \(p+1\) unique factor loadings. In the example above with \(p=3\) groups and \(r=p+1=4\) factors, the restrictions would be:

equation[equation omitted — 244 chars of source]

Ultimately, the third specification features all the groups exposed to \(r=2\) factors with the common and group-specific loadings on the first and the second factor, respectively. The resulting loading matrix is given by:

equation[equation omitted — 223 chars of source]

with \(i=1,2,3,\) and \(r=2\) factors.

Application to the macro-financial dataset

In the first part of the empirical analyses, we compare the score-driven factor model (Eq. (ref) and (ref)) with unrestricted loading matrix against the general lower-triangular (LT) restriction in Equation ((ref)) with respect to a set of \(N=8\) US macro and financial time series from January 1981 until August 2024 (\(T = 524\) months).\footnote{Given the dataset, the group-based restrictions in Equations (ref), (ref), and (ref), are excluded from this analysis.}

figure[figure omitted — 206 chars of source]

In particular, following creal2014observation, we consider the annual change in log industrial production (INDPRO), the annual change in the unemployment rate (UNRATE), the annualized S&P500 returns (S&PRet) and volatility (S&PVol), and the spread between the yield on Baa-rated bonds and the yield on 10-year Treasury bonds (BAA10YTB). In addition, the dataset contains the annual change in log retail sales (RETAIL), the annual change in the housing starts index (HOUSING), as well as the annual change in the survey-based consumer sentiment index (UMCSENT).\footnote{The data are retrieved from the FRED-MD macroeconomic database, while the S&PVol series is the annualized monthly realized volatility of the S&P500 Index constructed using the daily data from Yahoo Finance; see creal2014observation.} We plot the standardized series in Figure (ref).

To perform the estimation, we consider the Student-$t$ innovations, allowing for distinct constant terms and \(A\) parameters that govern the dynamics of factors, while the coefficients of the matrix \(B\) are the same. We initialize the parameters based on the PCA estimates for loadings and respective group-based averages for correspondingly restricted matrices. For the parameters governing the dynamics of factors, we set \(c\) and \(A\) equal to, respectively, the constant term and standard deviation estimates of an AR(1) model fitted to the principal components. Whereas, the initial value of the shared parameter \(B\) is set above 0.9 for all factors to reflect the persistence typically observed in the literature for such series.

table[table omitted — 1,207 chars of source]

The in-sample results (Table (ref)) evaluated in terms of the value of the maximized log-likelihood function (LLF), the Akaike information criterion (AIC), and the Bayesian information criterion (BIC), suggest that the 3F Full model provides the best fit to the data in terms of both AIC and BIC. Regardless of the number of factors, the models with unconstrained loadings, i.e., the so-called Full models, always significantly outperform the corresponding models with LT restrictions. As follows, all the restricted models are rejected against the unrestricted ones, with the $p$-value of each LR statistics below \(0.01\); see Table (ref).

Next, we assess the robustness of the estimated factor dynamics with respect to changes in the ordering of the observed variables. To this purpose, we examine the dynamics of the estimated factors of the best fitting model, 3F Full, and the corresponding restriction, 3F LT, under the shifted cross-sectional ordering of the data. Clearly, Figure (ref) confirms that the 3F Full fitted factors are order-invariant, whereas the law of motion of the 3F LT factors depends on the ordering (Figure (ref)). Moreover, the first fitted factor of the 3F Full model (Figure (ref)) captures the business cycle dynamics, with its troughs aligning with periods of economic crises, such as the dot-com bubble in 2002, the 2007-2008 financial crisis, and the 2020 COVID pandemic.

table[table omitted — 680 chars of source]

Finally, we perform a forecasting comparison of models based on their in-sample and out-of-sample MSE losses (Table (ref)). For the out-of-sample analysis, we use a rolling window approach, re-estimating the models over \(T_e\) = 312 months and generating one-month-ahead forecasts, for a total of 212 forecasts. The model rankings remain the same across both evaluations. Notably, the 4F Full model achieves the lowest average MSE loss, while each Full model outperforms its corresponding LT restriction.

table[table omitted — 1,260 chars of source]
figure[figure omitted — 253 chars of source]
figure[figure omitted — 245 chars of source]

Application to daily returns of S&P500 constituents

In the second part of the empirical application, we compare the score-driven factor model (Eq. (ref) and (ref)) with a full set of discussed specifications with restricted loading matrix, i.e., Equations (ref), (ref), (ref), and (ref).

Except for the LT restriction in Equation (ref), where the scale is identifiable, the parameter \(c_1\) is pre-set in the remaining restrictions in order to fix the scale of the model parameters. We estimate the models on the cross-section of \(N=87\) open-to-close log-returns of the S&P500 Index constituents over the period from January 2, 2001 until December 31, 2014 (\(T=3521\)) adopted by opschoor2021closed. Table (ref) in the Appendix provides a list of the ticker symbols and the corresponding industry classification for all the stocks.

The full set of candidate models includes the unrestricted score-driven factor models with 2, 5, and 11 factors (2F Full, 5F Full, and 11F Full, respectively), the 2 factor model with the common and group-specific loadings (2F GS; see (ref)), the 5 and 11 factor models with the general lower-triangular restriction (5F LT and 11F LT, respectively), coupled with the model with the group-specific lower-triangular loadings (10F GS-LT; see (ref)), and the specification with the common and group-specific factor loading matrix (11F GS; see (ref)). The choice of the number of factors is driven by \(p=10\) groups available in the dataset, together with the specifications of the loadings restrictions.

table[table omitted — 1,212 chars of source]

Table (ref) summarizes the full-sample estimation results.\footnote{The parameters are initialized in the same manner as for the macro-financial dataset.} We derive several conclusions. First, the 11F Full model outperforms all the other models in terms of both AIC and BIC criteria. Second, the models with restricted loadings are always inferior compared to the Full model with the same number of factors, with the best AIC and BIC among the former achieved under the general lower-triangular restriction (i.e., LT). Hence, reducing the model complexity by grouping the data based on the observed characteristics, such as the industry classification, significantly reduces the model fitting.

These results are supperted by the likelihood ratio (LR) tests shown in Table (ref).\ The table reports the LR statistics for testing the null hypothesis of a selected simpler version with restricted loadings against the unrestricted model that nests it. The $p$-value of each LR statistic is below 0.01, assuming a $\chi^2$ distribution with the reported number of degrees of freedom.

table[table omitted — 796 chars of source]

Given that the assumption of constant loadings may be considered restrictive, we perform the full in-sample analyses for the analogous set of models with time-varying loading dynamics in Equations (ref)-(ref). In particular, to avoid the parameter proliferation, we assume a scalar structure for \(\bm{A}^{(l)}\) and \(\bm{B}^{(l)}\), and target the unconditional mean \(\mathbb{E}[\bm{l}_t]\) by utilizing the corresponding static estimate of $\bm{\Lambda}$. As such, all the dynamic models have two additional parameters with respect to the static versions. The corresponding estimation results and LR tests are presented in Tables (ref) and (ref).

The comparison of the results in Tables (ref) and (ref) indicates that each dynamic model significantly outperforms the corresponding static version. Regarding the former, we confirm that the Full model with either 2, 5, or 11 factors is superior to the corresponding restricted versions in terms of both AIC and BIC. Furthermore, the dynamic 11F model with the general lower-triangular restriction, i.e., 11F LT, prevails over the models with group-based loading matrices, whereas the best-fitting model for each information criterion is the 11F Full model. Ultimately, the LR test results reported in Table (ref) show that all the restricted dynamic models are rejected against the corresponding unrestricted version at the 1% significance level.

table[table omitted — 1,275 chars of source]
table[table omitted — 863 chars of source]

In Figure (ref), we show an example of the fitted dynamic loadings. The loading paths exhibit notable shifts during the periods of financial and/or economic turbulence, e.g., the dot-com bubble, the Great Recession, etc., confirming that the assumption of the constant relationship between the returns of the selected S&P500 constituents and the underlying factors may be too restrictive.

figure[figure omitted — 224 chars of source]

We conclude that the flexibility of the score-driven framework under the introduced identification scheme significantly increases the model ability to capture the dependence dynamics of both financial and macroeconomic time series since the models with unrestricted loadings always outperform each of the restricted specifications in a highly statistically significant way.

Conclusions

This paper studies the identifiability of score-driven factor models, demonstrating that, under fairly general assumptions on the parameter space, the static and time-varying parameters can be determined up to multiplicative scalar constant. We further extend our results to include score-driven factor models with dynamic loadings and nonlinear factor models. This property distinguishes score-driven models from parameter-driven factor models, which are only identifiable up to rotations.\ The ability to identify score-driven models more precisely under minimal restrictions enhances their economic interpretability.

Our theoretical findings are validated through extensive simulations, confirming the practical applicability and robustness of the proposed methods. The results highlight the advantages of score-driven factor models in terms of both identifiability and ease of estimation, making them a valuable tool for economic and financial time series analysis.

In the empirical application, we verify that the score-driven framework under the introduced flexible identification scheme provides for significant in- and out-of-sample gains with respect to the typically applied restrictions in modelling both the financial and macroeconomic time series dynamics.