Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
47,406 characters · 8 sections · 54 citation commands
Retrieval from Mixed Sampling Frequency: Generic Identifiability in the Unit Root VAR
{\let\relax}
{\em Keywords:}{\em Keywords:} Mixed Frequency, REMIS, VAR, Cointegration, Vector Error Correction Model, Identifiability
{\em MSC:} 62M10, 62P20
Econometric analysis is often encountered with multivariate time series data sampled at mixed frequencies. Examples for treating this are Zadrozny1988, GhyselsEtAl2007[MIDAS-regression], anderson2012identifiability, Schorfheide2015Real-Time, ghysels2016macroeconomics, AndersonEtAl2016a and chambers2020frequency. Identifiability is a prerequisite for consistent estimation DeistlerSeifert1978, PoetscherPrucha1997 and often is needed for economic interpretation of effects related to particular model parameters. This article investigates {\em identifiability} of the model parameters in a Johansen1995 vector error correction model.\\ The general question is whether the internal characteristics, i.e. the model parameters $\theta$, can be retrieved from the external characteristics -- in our case observable second moments. Identifiability means that the mapping from the parameters to these second moments is injective. Often injectivity of this mapping can only be achieved for a certain subset of the parameterspace. Here, we prove that identifiability can be obtained for a generic subset of the parameterspace AndersonEtAl2016a.\\ As opposed to MIDAS-regression, where the observations at high frequency are considered as additional information, we consider mixed frequency as either a “missing-values” or a “dis-aggregation”-problem, by which we mean the following: We commence from an underlying high frequency system (e.g., a VECM) parameterised by $\theta$ for a multivariate process
with dimensions $n$, $n_f$ and $n_s$ for $y_t$, $y_t^f$ and $y_t^s$ respectively. Our aim is to identify and estimate the high frequency system from the observed (mixed frequency) data. The observational scheme is as follows: While the fast variables $y_t^f$ are observed at $t \in \mathbb{Z}$, for the slow variables $y_t^s$ we consider: \\ 1. Stock-Case: $y_t^s$ is observed only at $t \in N \mathbb Z$ for some sampling rate $N \geq 2$, hence we have a missing-value problem. \\ 2. Affine aggregation: we observe an affine transformation
where $c_i$ are known constant matrices for $i \geq 0$, $c_w$ is a known vector and $w_t$ is observed at $t \in N \mathbb Z$. A special case of affine aggregation are flow variables: For example suppose $y_t^s = GDP_t$, the monthly gross domestic product of a country. The quarterly GDP, $w_t$, is the sum of three monthly GDPs. We call $y_t^s$ latent whenever it is not directly observed. Hence, our aim is to retrieve the underlying high frequency parameters $\theta$ from data observed according to the observational schemes described above. \footnote{In this example we assume that the variable considered, $y_t$, is integrated of order one. If by contrast $(\log y_t)$ is integrated of order one, the affine approximation of Aadland2000 in combination with the methodology developed in this article can be applied.} \\ With the procedure described above, we are able to model all kinds of linear dynamic relationships between latent and observed variables, whereas the MIDAS ghysels2016macroeconomics approach only covers relationships between observed variables. After identifying the parameters one may interpolate missing values or dis-aggregate observations in a model based way by using the retrieved parameters of the underlying high frequency system.\\ Estimation of continuous time models from mixed frequency data are investigated in chambers2003asymptotic, chambers2016estimation, chambers2020frequency. In particular, chambers2003asymptotic, chambers2020frequency consider co-integrating regressions and show that the scaled estimators proposed, converge in distribution to functionals of Brownian motion and to stochastic integrals. Hence, the estimators are (weakly) consistent. Then, by GABRIELSEN1978 -- and for the case of strong consistency by DeistlerSeifert1978 -- the model parameters are identified.\\ For the stable vector auto-regessive model anderson2012identifiability and AndersonEtAl2016a either used the {\em blocking approach} Filler2010, ghysels2016macroeconomics or the {\em extended Yule-Walker equations} ChenZadrozny1998, AndersonEtAl2016a to show g-identifiability. For the same model class GersingAndDeistler2021 present an alternative proof for identifiability using the so-called canonical projection form. This idea is also applied in this paper. On the other hand, DeistlerEtAl2017 show that the parameters need not be identified in the auto-regressive-moving average (VARMA) case, if the order of the MA polynomial exceeds the order of the AR polynomial.\\ This article is organised as follows: Section (ref) starts with the vector error correction model developed in Johansen1995 as the underlying high frequency model. In Section (ref) we describe the observational schemes considered in detail. In particular, we introduce a stationary blocked process containing all observed variables. Section (ref) introduces conditions, which are later shown to be sufficient for identifiability. We prove that these conditions hold generically in the underlying high frequency parameterspace. Section (ref) extends the REMIS approach to the non-stationary case: Here, we use the result from chambers2020frequency that the cointegrating vectors can be identified from mixed frequency data. First, we derive a state-space representation of the blocked process that we call Canonical Projection Form (CPF). In this representation, the system matrices are simple transformations of the parameters of the underlying high-frequency model. After that we start from the unique factor of the spectrum of the blocked process ScherrerDeistler2018book to get an arbitrary minimal realisation for this factor and relate this to the CPF. From there we can retrieve the parameters of the underlying high frequency system using the structural properties of the CPF. Section (ref) adds deterministic terms. Finally, Section (ref) concludes.
In the first step, we introduce the class of underlying high frequency systems: We commence from a process which is integrated of order one and allows for cointegration. Suppose $(y_t)_{t \in \mathbb Z}$ is $n \times 1$ and a solution on $\mathbb Z$ of the vector error correction system:
where $(\nu_t)_{t \in \mathbb Z}$ is white noise and $\Pi$ is of rank $r > 0$ in the case of cointegrating relationships, but we also allow the case $r = 0$. Such solutions always exist and can be constructed as described in detail in bauerwagner2012. We obtain a unique factorisation of $\Pi = \alpha \beta'$ with $\alpha, \beta \in \mathbb{R}^{n \times r}$ applying the singular value decomposition to $\Pi$ in the following way:
where $Q$ is a non-singular matrix of elementary row operations that transforms $\tilde D V_1'$ into its reduced echelon form, such that $Q \tilde D V_1' =
$. We stack the parameters $\alpha, \beta, \Phi_1, ..., \Phi_{p-1}$ to a vector $\theta_{VECM} \in \mathbb{R}^{d}$, where $d = nr + (n - r)r +(p-1)n^2$.\\ We also have a VAR$(p)$ representation for $(y_t)$ of the form,
Throughout this article, we assume that $r$ and $p$ are known a priori. We obtain the representation in ((ref)) by the mapping $\psi$:
with $\theta_{AR} = \operatorname{vec}
$. On the other hand for a $\theta_{AR}$ which has a corresponding VECM representation, we compute $\theta_{VECM}$ as follows:
Now, define the polynomial matrix $a(z) = I_n - \mathcal{A}_1 z - \cdots - \mathcal{A}_p z^p$ where $z$ is a complex variable or the lag operator on $\mathbb Z$ depending on the context. For $\check{c} =
\in \mathbb{R}^{n \times r}$ and $\check{c}_\bot =
\in \mathbb{R}^{n \times (n-r)} $, $\beta_\bot := \big(I_n - \check{c} (\beta' \check{c})^{-1} \beta' \big) \check{c}_\bot$, and $\alpha_\bot$ defined analogously to $\beta_\bot$. We impose the following assumptions Johansen1995:
We define the parameterspace as follows:\footnote{We write $\mathbb R^d \Big|_{C1, C2}$ to denote the set of real vectors in $\mathbb R^d$ for which $C1$ and $C2$ hold.}
Note that under these assumptions $\psi$ is a homeomorphism. The set of $\operatorname{vech}{\Sigma_\nu}$ with $\Sigma_\nu \in \mathbb{R}^{n \times n}$, $\Sigma_\nu = \Sigma_\nu'$ and $\Sigma_\nu > 0$ (condition (C4) in Assumption (ref)) is denoted by $\Theta_2$. The overall parameterspace for the VAR$(p)$ representation is
We will also need the state-space representation of $(y_t)_{t \in \mathbb Z}$, which follows from ((ref)):
Note that ((ref)), ((ref)) is always controllable as $\Sigma_\nu$ and therefore $\Gamma (t) := \operatorname{\mathbb E} \big(X_{t+1} X_{t+1}'\big) $ are of full rank. The system ((ref)), ((ref)) is also observable whenever $\mathcal A_p$ is of full rank. This follows since $\mathcal A_p$ is nonsingular (and therefore $\mathcal A$ is non-singular) from the BPH-test (see Kailath1980 2.4.3). Hence under Assumption (ref) and if $\mathcal A_p$ is nonsingular the system ((ref)), ((ref)) is minimal. For details on controllability and observability see e.g. ScherrerDeistler2018book, chapter 7 or HannanAndDeistler2012, chapter 2.
A main challenge of the identifiability proof in the integrated case -- as opposed to the stationary case AndersonEtAl2016a -- is that the second moments of an integrated process {(that is, $\mathbb{E} y_s y_t$, $s,t \in \mathbb{Z}$)} are time dependent and cannot be estimated directly. Instead, for the sake of practical relevance of identifiability considerations, we identify from observable second moments of stationary transformations of the level process {(that is, $\left( y_t \right)_{t \in \mathbb{Z}}$).} \\ Suppose for the moment, that the matrix of cointegration vectors $\beta$ is known. Our proof commences from what we call the “blocked process”, where we distinguish between the Stock- and the Flow-case:\\ 1. Stock Variables: In this case for $t \in N \mathbb{Z}$, we get the co-stationary vector $\tilde{y}_t$ of “observed” random variables. We will use $\tilde{n} := r + n + (N-1) n_f$ for the dimension of $\tilde y_t$ henceforth. Let $u_t^\mathcal{S} :=\beta^{\prime} y_t$, $\Delta_N y_t := y_t -y_{t-N} = \sum_{j=0}^{N-1} \Delta y_{t-j}$, {and }\\
The blocked process $(\tilde y_t)$ is similar to the blocked process in AndersonEtAl2016a with the distinction that we added the variable $\beta'y_t = u_t^{\mathcal S}$ and take differences at lag $N$. Admittedly, the true $\beta$ is in fact not observed, however since $\beta$ can be estimated consistently, for the purpose of the analysis of identifiability we can assume $\beta'y_t$ to be observed.\\ 2. Flow Variables: In a similar way, we may consider the case where all slow variables are flow variables, in which case we are able to observe the temporal aggregate $w_t := \sum_{j=0}^{N-1} y_{t-j}^s$ at $t \in N \mathbb{Z}$. So
If all slow variables are flow variables, we can observe $\sum_{j=0}^{N-1} y_{t-j} = \left( w_t^{\prime} , \sum_{j=0}^{N-1} y_{t-j}^{f \prime} \right)^{\prime} $, $t \in N \mathbb{Z}$. Since $\beta^{\prime} y_t$ is stationary, we have that $\left( \beta^{\prime} y_t \right)_{t \in N \mathbb{Z}}$ and $u_t^\mathcal{F} := \beta^{\prime} \sum_{j=0}^{N-1} y_{t-j} \in \mathbb{R}^r$ are integrated of order zero. For the flow case we define the co-stationary vector process
We call the autocovariance function of the (stationary) blocked process
observed second moments, which can be consistently estimated from the data (if $\beta$ is known) under standard assumptions. \\ The motivation to consider this blocked process for identifiability is the following:\\ 1. We take differences at lag $N$ (as opposed to lag one) because these differences can be directly computed from the mixed frequency data and are stationary.\\ 2. Note that the set of observable autocovariances given mixed frequency data is
where the superscript “$\cdot$” is shorthand for $\mathcal{S}$ or $\mathcal{F}$. Note that these are exactly the second moments of the autocovariance function $\tilde \gamma$ of the blocked process defined in equations ((ref)) for the stock case. In an obvious way this is treated accordingly in the flow case ((ref)). So the blocked process “contains the whole second moment information available” from which we can identify. The same idea is also applied for the stationary case in AndersonEtAl2016a. \\ 3. Our interest in the particular blocked process ((ref)), ((ref)) having $u_t^\cdot$ in the first coordinates, originates in the fact that we can obtain a minimal representation for this process (see Section (ref)), where the parameters are fairly simple functions of the parameters of the underlying high frequency system. This will finally help us to retrieve the high frequency model parameters.\\ Next, we define the concept of generic identifiability. Here, identifiability is concerned with the problem whether the parameters of the underlying high frequency system ((ref)), ((ref)) or ((ref)) are uniquely determined from the observable second moments (defined below in this section). To be more precise, a subset $\Theta_I \subset \Theta$ is called identifiable, {if the mapping attaching the observable second moments to the parameters $\theta \in \Theta_I$ is injective. } In our setting identifiability for the whole set $\Theta$ cannot be obtained. To see this, we consider a simple example where $p=1$, $r=1$, and $n=2$, the first coordinate of $y_t$ is a fast variable, denoted $y_{t}^{f}$, while the second coordinate, $y_t^s$, is a slow stock variable. We assume that the cointegrating vector $\beta= \left(1,\beta_{s} \right)$ is known. Recall that the observed second moments are as described in equations ((ref)) and ((ref)). Let $\sigma_{ff}$, $\sigma_{fs}^{}= \sigma_{sf}$, and $\sigma_{ss}^{}$ denote the elements of the covariance matrix $\Sigma_\nu$. Appendix (ref) shows that there exist two parameter vectors $\theta^I := \left( \alpha_f^{I}, \alpha_s^{I}, 1, \beta_{s}^{}, \sigma_{ff}^{I}, \sigma_{fs}^{I}, \sigma_{ss}^{I} \right)^{\prime} \not= \theta^{II} := \left( \alpha_f^{II}, \alpha_s^{II}, 1, \beta_{s}^{}, \sigma_{ff}^{II}, \sigma_{fs}^{II}, \sigma_{ss}^{II} \right)^{\prime} $ such that all observable second moments are the same; hence in this case the mapping from the model parameters to observable second moments cannot be injective and the model parameters are not identified from observed second moments. In this example $\alpha_f^{I} = \alpha_f^{II}=0$. This implies that the fast coordinate follows a random walk and does not provide any information on the parameter $\alpha_s$, that is on how $\beta^{\prime}y_t$ affects $\Delta y_{t}^{s}$, $t \in 2 \mathbb{Z}$. However, in this paper we prove that identifiability holds for a so called generic subset of $\Theta$. Note that a set $\Theta_I \subset \Theta$ is called generic in $\Theta$, if it contains a subset that is open and dense in $\Theta$.\\ Let $\Theta_{I} := \left(G \cap \Theta_1 \right) \times \Theta_2$, where $G\subset \mathbb R^{n^2 p}$ is defined in Assumption (ref) below. In this paper we show firstly that $\Theta_I$ is generic in $\Theta$ (see Section (ref)) and secondly that the set of high frequency systems corresponding to $\Theta_I$ is identifiable from the observable second moments (see Section (ref)). Or formally, we show that
is injective on $\Theta_I \subset \Theta$. \\ Finally, in terms of identifiability, we may suppose without loss of generality that $\beta$ is known. For instance miller2016conditionally or chambers2020frequency propose estimators, accounting for stock and flow variables, respectively. The estimators of $\beta$ scaled by $T$ weakly converge to a random variable bounded in probability. Hence, e.g. by white2001asymptotic, the estimator is weakly consistent. By GABRIELSEN1978 the matrix of cointegrating vectors $\beta \in \mathbb{R}^{n \times r}$ is identified from mixed frequency observations given the assumptions imposed in chambers2020frequency or miller2016conditionally. These assumptions are only posed on the stochastic properties of the high frequency innovations $(\nu_t)_{t \in \mathbb Z}$ and therefore do not restrict our results on the genericity of the identifiability conditions from Section (ref). If strong consistency could be established for some estimator of $\beta$, the results of DeistlerSeifert1978 apply and $\beta$ is identified.
In this section we define the conditions that we need for identifiability and prove that these conditions result in a generic subset of the parameterspace. Define a set $G \subset \mathbb R^{n^2 p}$ by the following assumptions:
Assumption (I2) already follows from $\Sigma>0$. Recall that $\Theta_{I} = \left(G \cap \Theta_1 \right) \times \Theta_2$. These assumptions are similar to the stationary case considered in Felsenstein2014, AndersonEtAl2016a, AndersonEtAl2016b. There, the stability condition defines an open set $\Theta' \subset \mathbb R^{n^2 p }$. We also have a corresponding set $G'$ defining the identifiability conditions for the stationary case, which is generic in $\mathbb R^{n^2 p }$. Then, the intersection $\Theta' \cap G'$ is generic in $\Theta '$. However, in the integrated case, where unit roots occur, the situation is more intricate since neither $\Theta_1$ nor $G$ is open in $\mathbb R^{n^2 p}$. This follows from the fact that for a process with $n-r$ common trends, the $n-r$ eigenvalues of $\mathcal A$ in ((ref)) are equal to one [note that the eigenvalues of $\mathcal{A}$ are the reciprocals of the zeros of $a(z)$]. The following Theorem (ref) implies that the identifiability conditions are generically fulfilled in $\Theta$:
Since genericity is a topological property, it also holds for the homeomorphic parameterspace corresponding the vector error correction representation in ((ref)) defined by Assumption (ref).
In this section, we first define a canonical state-space representation for the blocked process running on $t \in N \mathbb Z$. We prove that this representation is minimal under our identifiability conditions. Then under an additional assumption on the lag order $p$, we show that the high frequency parameters are generically identifiabile. The proofs of minimality and identifiability make use of the canonical representation.\\ We follow HansenJohansen1999 and obtain from ((ref)) the following state-space system for $\beta^{\prime} y_t$ and first differences of $y_t$, that is $\Delta y_t = y_t - y_{t-1}$. Then,
By $m := r + n(p-1)$, we denote the dimension of $\underline{x}_t$. As we will see later, given that our identifiability conditions hold $m$ is also the McMillan degree of $(\tilde y_t)_{t \in N\mathbb Z}$. \\ According to the observational scheme, the slow variables $y_t^s$ are observed only every $N$-th period. We derive state-space representations for the processes ((ref)) and ((ref)) running on $t \in N \mathbb Z$:\\ 1. Case: Stock Variables: We define a new state vector $x_{t+1}$ in the following way, with the condition that $p \geq N + 2$: {
} By iterating the system ((ref)), ((ref)), we get the non-miniphase system (in the sense that the transfer-function is not causally invertible as the input dimension exceeds ($Nn$) the output dimension ($\tilde n$), noting that $\Sigma_\nu > 0$): {
} The matrices $B_{b, c} \in \mathbb{R}^{r+n(p-1) \times Nn} $ and $D_b \in \mathbb{R}^{r+n \times Nn}$ are obtained from $B$ and $A$.\\ 2. Case: Flow Variables: Next, we obtain the state vector $x_{t+1}$ for the flow case. Note that $y_{t-j} = y_t - \sum_{\ell=1}^{j} \Delta y_{t-\ell}$, such that $\sum_{j=0}^{N-1} y_{t-j} = \sum_{j=0}^{N-1} \left( y_t - \sum_{\ell=1}^{j} \Delta y_{t-\ell} \right) = N y_t - (N-1) \Delta y_{t-1} - \cdots - \Delta y_{t-N+1} $. Analogously to equation ((ref)), this yields for $p \geq 2N+1$ that {\scriptsize
} We use the same notation for $\tilde y_t$, $x_t$, $c$ for both cases. With this notation, we obtain the following state-space representation for blocked process in the flow case: {
} The matrix $D_{b,c} \in \mathbb{R}^{\tilde{n} \times Nn} $ follows from $D_b$, the matrix $c$ and the selection of the corresponding rows resulting in $\tilde{y}_t$. \\ 3. Case: Mixed Case: Consider the case where we have slow stock as well as slow flow variables: For example, if $\left( y_t \right)$ is a three-dimensional process, where $n_f=1$, $n_s=2$, $N=2$, $c_1=I_2$, and $c_2 =
$ in equation~(\ref{eq: def w_t}). Then $\beta^{\prime} \left(y_t^{f \prime},w_t^{\prime} \right)^{\prime} $ is (in general) not stationary. However, in special cases, such as separate cointegrating relationships among the slow flow variables only, or among the slow stock and fast variables only, etc. we can proceed similarly to the flow case. In the following we only consider the stock or the flow case. \\ The problem with the systems considered above is that the inputs $\nu^b_t$ are not the innovations of $\tilde{y}_t$. However, from the stable miniphase spectral factorisation, we only obtain transfer functions corresponding to systems in innovation form \citep[see, e.g.,][Chapter~7]{ScherrerDeistler2018book}. The following Theorem \ref{thm: th1aggregated} is the first step for obtaining a canonical state-space representation for the blocked process. A minimal state-space representation is called ``canonical'' if its parameters are uniquely determined from the transfer function. We introduce the following notation for specific subspaces of $L^2(\Omega, \mathcal A, P)$, the space of square integrable random variables on the underlying probability space $(\Omega, \mathcal A, P)$:
where $\overline{\operatorname{sp}}(\cdot)$ denotes the closed span and $\operatorname{proj}(v \mid U)$ the projection of $v$ on a closed subspace $U$ of $L^2$.
We call the representation in ((ref)), ((ref)) canonical projection form (CPF) of $\tilde{y}_t$. Note that the CPF provides an algorithm for computing the transfer function $\tilde k(\tilde z)$ of $(\tilde y_t)_{t \in N \mathbb Z}$ which corresponds to the Wold representation, where $\tilde z := z^N$.\\ Next we show that the system ((ref)) and ((ref)) is observable and controllable and therefore minimal HannanAndDeistler2012 for all $\theta \in \Theta_I$.
By Theorem (ref), we know that the McMillan degree of the transfer function of the blocked process $(\tilde y_t)_{t \in N \mathbb Z}$ corresponding to an underyling high-frequency VECM is $m = r + n (p - 1)$. This will be used in the proof of the subsequent Theorem (ref), where we can relate an arbitrary minimal realisation $(\bar A_{b,c}, \bar B_{b, c}, \bar C_{b, c})$ of the transfer function $\tilde k(\tilde z) = \left(\bar{C}_{b,c} \left( I_{m} \tilde{z}^{-1} - \bar{A}_{b,c} \right) \bar{B}_{b,c} + I_{\tilde{n}} \right) $ (where $\tilde z := z^N$) to the CPF $(A_{b, c}, \tilde B_{c}, C_{b, c})$. The minimal realisation $(\bar A_{b,c}, \bar B_{b, c}, \bar C_{b, c})$ can be either obtained by the spectral factorisation and e.g. the echelon realisation from the Hankel matrix of the transfer function HannanAndDeistler2012 or directly from the Hankel matrix of the observed second moments AndersonEtAl2016a. In the next step we relate the CPF to the underlying VECM/VAR -- exploiting the fact that the parameters $\theta$ of the underyling VECM reappear in the CPF.\\ Finally, we show that the parameters of the high frequency system are generically identifiable from the observed second moments, i.e. from $\tilde \gamma$.
Since by Theorem (ref), $\Theta_I$ is a generic subset of $\Theta$, we say that $\theta$ is generically identifiable from the observed autocovariance function $\tilde \gamma$. Theorems (ref) and (ref) imply that the representation ((ref)), ((ref)) is indeed canonical on $\pi(\Theta_I)$. Since the second moments of $(\tilde y_t)$ can be consistently estimated from the data under mild conditions, by the continuity of $\pi^{-1}$ it follows that we have a consistent estimator for $\theta$. The mapping $\pi^{-1}$ is also called realisation procedure, since we realise the system parameters from the external characteristics of the data, i.e. the second moments, the spectrum or the transfer function respectively.\\ Finally, we consider the question whether $\pi^{-1}(\pi (\Theta_I)) = \Theta_I$. This is important to ensure that outside that the identified parameter set $\Theta_I$ there are no elements, say $\theta_{\neg I}$, which result in the same observable second moments as some $\theta_{} \in \Theta_I$:
This section investigates the VECM
containing the deterministic terms $\mu_0$ and $\mu_1 t$. The five cases following from ((ref)), namely “$H_2(r)$, $H_1(r)$, $H_1^*(r)$, $H(r)$, and $H^*(r)$”, are obtained and defined in Johansen1995[page 81 and our Appendix (ref)]. Recall that the cointegrating vectors $\beta$ can be identified from mixed frequency data chambers2020frequency. For high frequency data $(\Delta y_t)_{t \in \mathbb Z}$ and $(\beta'y_t)_{t \in \mathbb Z}$ we can compute the expectations $\operatorname{\mathbb E} \Delta y_t$ and $\operatorname{\mathbb E} \beta^{\prime} y_t$, while for mixed frequency case we get $\operatorname{\mathbb E} \beta^{\prime} y_t$ and $\operatorname{\mathbb E} \Delta_N y_t = \mathbb{E} \left(y_t - y_{t-N} \right)$ for the stock case and $\operatorname{\mathbb E} \beta^{\prime} w_t = \operatorname{\mathbb E} \sum_{j=0}^{N-1} \beta^{\prime} y_{t-j}$ and $\operatorname{\mathbb E} \Delta_N^{\Sigma} y_t = \operatorname{\mathbb E} \sum_{j=0}^{N-1} \Delta_N y_{t-j}$ for the flow case, respectively. To identify the deterministic terms in ((ref)) we can proceed as follows:
This results in:
In this paper, we generalise the results on identifiability from mixed frequency data in AndersonEtAl2016a, AndersonEtAl2016b obtained for stationary VAR-systems to the case of unit-roots and cointegrating relationships. As is well known these systems have also a {\em vector error correction representation}. The corresponding parameterspaces are homeomorphic.\\ We commence from a solution of the (unstable) VAR system on the integers $\mathbb Z$ bauerwagner2012. Then we take differences at lag $N$ (which is the sampling rate of the slow/aggregated process) and stack these to what we call the “blocked process”. In addition, the blocked process also contains the stationary process $\beta'y_t$, where $\beta$ is the matrix of cointegrating relationships. This matrix is identified from mixed frequency data as already shown in chambers2020frequency. This blocked process is stationary and contains all relevant differences of the observations. \\ The contribution of this paper can be seen as an extension of the results in chambers2020frequency, by proving that also the remaining parameters of the vector error correction model (i.e. besides $\beta$) are (generically) identified from mixed frequency observations.\\ The identifiability proof consists of two steps: In the first step, we derive a state-space representation of the blocked process (“the canonical projection form”) which is minimal, in innovation form (both, for the stock and the flow case) and unique. In the second step, we derive an algorithm, that retrieves the parameters of the underlying high frequency system from the parameters of the canonical projection form.\\ We show that the conditions (Assumption (ref)) which are sufficient for identifiability are generic in the parameterspace. This is more intricate than in the stationary case, since the parameterspace is not an open subspace of the Euclidean space, due to the fact that we allow for unit roots. Since the VECM and the VAR parameterspaces are homeomorphic, the genericity result holds for both.\\ Finally, we show that all common cases of deterministic terms in the VECM can be reduced to the case of non-deterministic terms.