Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
42,431 characters · 5 sections · 16 citation commands
\centerline{Yun Liu$^1$ and Yeonwoo Rho$^{2}$\footnote[0]{ Address for correspondence: Yeonwoo Rho, Department of Mathematical Sciences, Michigan Technological University, Houghton, MI 49931, USA. (Email: [email removed])}} \centerline {\textsuperscript{1,2}\it Michigan Technological University} \centerline{\today}
In recent years, datasets that involve different sampling frequencies have drawn substantial attention in various fields. Several methods were introduced to handle mixed-frequency variables in a regression model. One conventional approach is time averaging of high-frequency variables, where high-frequency variables are aggregated using a predetermined fixed-weight function. Another is the autoregressive distributed lag (ADL) model, in which all high-frequency variables are used as regressors. The mixed data sampling (MIDAS) regression model 2004 was proposed to balance the complexity and the flexibility of these two approaches. In MIDAS models, the weight function is written as a nonlinear parametric function with a few parameters. The elements in the weight function do not move as freely as the ones in the ADL model due to the parametric restriction. They are still more flexible than those in time averaging since parameters in the weight function are determined by data. This idea of concise yet data-driven reduction of information embedded in high sampling frequency has driven a recent surge of interest in MIDAS models foroni2013survey.
However, MIDAS models involve nonlinear estimation. If the time averaging is good enough, there is no need to go through this nonlinear estimation. This motivates a specification test that helps decide between the time averaging and the MIDAS models. There have yet been only a handful of such tests. andreou presented a Durbin-Wu-Hausman (DWH) type test, designed to see whether there is an omitted variable bias caused by overlooking the MIDAS effect. Miller2016 presented two variable addition test (VAT) statistics. In particular, the second VAT statistic, called a modified VAT statistic, was designed for nonstationary high-frequency variables. Groenvik:Rho:2017 further extended Miller's first VAT statistic using a self-normalized approach.
In this paper, we shall further explore the DWH specification test introduced in andreou. In particular, the DWH test requires choosing appropriate instrumental variables, but there has not yet been a practical guidance so far. We shall propose a set of instrumental variables that is suitable for this test. Section (ref) presents details of such a choice, demonstrating its theoretical consistency when the frequency ratio is large enough. Section (ref) presents finite sample comparisons. All technical proofs and full simulation results can be found in the appendix.
The following notations are used consistently throughout the manuscript. Let $T$ be the sample size at low frequency, and $m$ be the frequency ratio between the two sampling frequencies. $\mathbf{j}_t$ is a $T\times1$ vector with the $t$-th element being 1 and the rest 0. $\mathbf{j}$ is a $T\times1$ vector of 1's. Symbols $\mathbf{y}=(y_{1},\cdots,y_{T})'$, $\mathbf{x}_{t,m}^{(m)}=\left(x_t,x_{t-1/m},\cdots,x_{t-(m-1)/m}\right)'$, and $\mathbf{z}_t=(z_{1,t},\cdots,z_{p,t})'$ are reserved for the low frequency variable, the high frequency variable, and $p$ instrumental variables, respectively. We use $\boldsymbol{\pi}=(\pi_1, \cdots, \pi_m)'$ to indicate an $m\times1$ weight vector to aggregate the high frequency variable such that $\pi_i\geq0$ and $\sum_{i=1}^m \pi_i=1$. The matrix $\mathbf{P}_{\mathbf{A}}=\mathbf{A}(\mathbf{A}'\mathbf{A})^{-1}\mathbf{A}'$ denotes the projection matrix onto the space spanned by the columns of $\mathbf{A}$, and $\mathbf{M}_{\mathbf{A}}=\mathbf{I}-\mathbf{P}_{\mathbf{A}}$. For convenience, we define the following matrices: $\mathbf{X}=[\mathbf{x}_{1,m}^{(m)},\cdots,\mathbf{x}_{T,m}^{(m)}]'$, $\mathbf{Z}=[\mathbf{z}_1,\cdots,\mathbf{z}_T]'$, and $\mathbf{X}^A=[\mathbf{j},\mathbf{X}\boldsymbol{\pi}_0]=[\mathbf{x}_1^A,\cdots,\mathbf{x}_T^A]'$, where $\mathbf{x}_t^A=(1,x_t^A)'$ is the $t$-th row of $\mathbf{X}^A$ and $\boldsymbol{\pi}_0$ is the predetermined weight vector.
Consider a dataset with different sampling frequencies. Let $\{y_t\}_{t=1}^T$ and $\{\mathbf{x}_{t,m}^{(m)}\}_{t=1}^T$ be the variables observed at lower and higher sampling frequencies, respectively. The MIDAS model is constructed, aiming to model low-frequency variable using high-frequency variable:
The error process $\{u_{t}\}$ is stationary and uncorrelated with $\{\mathbf{x}_{t,m}^{(m)}\}$. The vector $\boldsymbol{\pi}(\boldsymbol{\theta})=(\pi_1(\boldsymbol{\theta}),\ldots,\pi_m(\boldsymbol{\theta}))'$ consists of a function of a finite dimensional unknown parameter $\boldsymbol{\theta}$ such that $\pi_i(\boldsymbol{\theta})\geq0$ and $\sum_{i=1}^m\pi_i(\boldsymbol{\theta})=1$. This vector dictates how much weight would be assigned when aggregating the high-frequency variable, $\mathbf{x}_{t,m}^{(m)}$.
In a time averaging model, $\boldsymbol{\pi}=\boldsymbol{\pi}_0$ is a predetermined fixed-weight vector that does not depend on any unknown parameter $\boldsymbol{\theta}$. Without loss of generality, let the number of aggregated lags be the same as the frequency ratio $m$. Then the regression model ((ref)) becomes
We consider the test between time averaging ((ref)) and MIDAS aggregation ((ref)), i.e. $H_0:\boldsymbol\pi=\boldsymbol\pi_0$ versus $H_a:\boldsymbol\pi=\boldsymbol\pi(\theta)$. The two commonly used weights for time averaging are the flat aggregation $\boldsymbol{\pi}_0=(1/m,\ldots,1/m)'$ and the end-of-period sampling $\boldsymbol{\pi}_0=(1,0,\ldots,0)'$. In this article, a more general scenario of the end-of-period sampling is considered: a fixed number, $n$, of elements in $\boldsymbol{\pi}_0$ are assigned with positive values, where $n$ is independent of $m$. For brevity, we assign the first $n$ elements and leave the rest as zero, i.e. $\boldsymbol{\pi}_0=(\pi_{0,1},\ldots,\pi_{0,n},0,\ldots,0)'$ where $\pi_{0,i}>0$ for $i=1,\cdots,n$ and $\sum_{i=1}^{n}\pi_{0,i}=1$.
The least squares (LS) principle can be applied to estimate the parameters $\beta_0^A$ and $\beta_1^A$ in model ((ref)) when the null hypothesis is true. We call this estimator, $\widehat{\boldsymbol\beta}^A=(\widehat{\beta}_0^A,\widehat{\beta}_1^A)'=({\mathbf{X}^{A}}'\mathbf{X}^{A})^{-1}{\mathbf{X}^{A}}'\mathbf{y}$, the NULL-LS estimator. By comparing models ((ref)) and ((ref)), the error process ((ref)) can be rewritten as $u_{t}^A=u_t+\mathbf{j}'_t\mathbf{X}\left(\boldsymbol{\pi}(\boldsymbol{\theta})-\boldsymbol{\pi}_0\right)\beta_1$. Under the null, $u_{t}^A$ is uncorrelated with $x_t^A$ since $u_t^A=u_t$. However, under the alternative, $u_{t}^A$ is correlated with $x_t^A$ due to the omitted variable. Therefore, testing whether $\boldsymbol\pi=\boldsymbol\pi_0$ is equivalent to testing whether the NULL-LS estimator is consistent.
To test the consistency of the NULL-LS estimator using a DWH-type test, another estimator that is consistent under both the null and the alternative is required. This estimator may not be efficient under the null. See Lee:2010, for example. The two stage least squares (2SLS) estimator with proper instruments could be such an estimator. Assume that the instruments $\mathbf{z}_t$ are correlated with $x_t^A$, but uncorrelated with $u_{t}^A$. Consider a two stage regression model: the time-averaging model ((ref)) and an auxiliary regression of the flat aggregated term $x_t^A$ on the instrumental variable $\mathbf{z}_t$ given as
where $E\left(\varepsilon_t|x_t^A\right)=0$. The 2SLS estimator is $ \widehat{\boldsymbol\beta}=({\mathbf{X}^A}'\mathbf{P}_{\mathbf{Z}}\mathbf{X}^A)^{-1}({\mathbf{X}^A}'\mathbf{P}_{\mathbf{Z}}\mathbf{y})$. The bias of the 2SLS estimator $\widehat{\boldsymbol\beta}$ of $\boldsymbol\beta$ can be written as
where $\mathbf{u}^A=(u_1^A,\ldots,u_T^A)'$. The following Assumption (ref) is for the consistency of the NULL-LS under the null and for the consistency of the 2SLS estimator under both the null and the alternative.
Assumptions (ref)(ref) and (ref)(ref) ensure the consistency of the NULL-LS estimator. Assumption (ref)(ref) indicates that $\mathbf{X}^A$ has full column rank. Assumption (ref)(ref) implies the relation between the time-averaging term $\mathbf{X}^A$ and the error process $\mathbf{u}^A$, and their product should be asymptotically normal. Under the null, $\mathbf{X}^A$ and $\mathbf{u}^A$ should not be correlated, leading $E\left({\mathbf{x}_t^A}{{u}_t^A}\right)=0$. Under the alternative, $\mathbf{X}^A$ and $\mathbf{u}^A$ are allowed to be correlated, i.e., $E\left({\mathbf{x}_t^A}{{u}_t^A}\right)\neq\mathbf{0}$. The variance-covariance matrix $\boldsymbol{\Omega}$ in Assumption (ref)(ref) can be consistently estimated. This can be done, for example, using heteroskedasticity and autocorrelation consistent (HAC) estimators Newey:West:1987,Andrews:1991. Assumptions (ref)(ref)--(ref) hold under both hypotheses. These ensure the consistency of the 2SLS estimator. In particular, Assumption (ref)(ref) requires that $\mathbf{Z}$ and $\mathbf{u}^A$ should be uncorrelated. It is worth noting that the number of instrumental variables should be greater than or equal to the rank of $\mathbf{X}^A$. Refer to Ruud:2000 for more details and explanations.
Now we derive our test statistic. If Assumption (ref) holds, the asymptotic distributions of $\widehat{\boldsymbol{\beta}}^A$ under the null and $\widehat{\boldsymbol{\beta}}$ under both hypotheses can be written as followings:
where $\mathbf{V}^A=\mathbf{Q}_{XX}^{-1}\boldsymbol\Omega\mathbf{Q}_{XX}^{-1}$ and $\mathbf{V}=\left(\mathbf{Q}_{XZ}\mathbf{Q}_{ZZ}^{-1}\mathbf{Q}_{XZ}'\right)^{-1}\left(\mathbf{Q}_{XZ}\mathbf{Q}_{ZZ}^{-1}\boldsymbol\Sigma_{Zu}\mathbf{Q}_{ZZ}^{-1}\mathbf{Q}_{XZ}'\right)\left(\mathbf{Q}_{XZ}\mathbf{Q}_{ZZ}^{-1}\mathbf{Q}_{XZ}'\right)^{-1}$. Since both $\widehat{\boldsymbol\beta}^A$ and $\widehat{\boldsymbol\beta}$ are consistent under the null, the difference between the two estimators, $\widehat{\Delta}=\widehat{\boldsymbol\beta}-\widehat{\boldsymbol\beta}^A$ converges to zero in probability. The main idea of the DWH test is to test whether $\widehat{\Delta}$ is significantly different from $\mathbf{0}$. This is equivalent to test whether ${\mathbf{X}^A}'\mathbf{P}_{\mathbf{Z}}\mathbf{M}_{\mathbf{X}^A}\mathbf{y}$ is significantly different from $\mathbf{0}$, since $\widehat{\Delta}$ can be written as $\widehat{\Delta}=\widehat{\boldsymbol\beta}-\widehat{\boldsymbol\beta}^A =({\mathbf{X}^A}'\mathbf{P}_{\mathbf{Z}}\mathbf{X}^A)^{-1}({\mathbf{X}^A}'\mathbf{P}_{\mathbf{Z}}\mathbf{M}_{\mathbf{X}^A}\mathbf{y})$ and $({\mathbf{X}^A}'\mathbf{P}_{\mathbf{Z}}\mathbf{X}^A)^{-1}$ is positive definite.
We can easily see that
Thus, $(\mathbf{X}\boldsymbol{\pi}_0)'\mathbf{P}_{\mathbf{Z}}\mathbf{M}_{\mathbf{X}^A}\mathbf{y}$ should be approximately zero under the null. Let $\widehat{\boldsymbol{\varepsilon}}=\mathbf{M}_{\mathbf{Z}}\mathbf{X}\boldsymbol{\pi}_0$ and $\widehat{\mathbf{u}}^A=\mathbf{M}_{\mathbf{X}^A}\mathbf{y}$ indicate the fitted residuals from ((ref)). Consider a regression model $\widehat{\mathbf{u}}^A={\mathbf{X}^A}\boldsymbol\alpha+{\widehat{\boldsymbol\varepsilon}}\delta+\boldsymbol\upsilon$. Applying Frisch$-$Waugh$-$Lovell (FWL) theorem, the OLS estimator $\widehat{\delta}$ of $\delta$ is
Note that the latter part of $\widehat{\delta}$ can be derived as $(\mathbf{M}_{\mathbf{X}^A}\mathbf{M}_{\mathbf{Z}}\mathbf{X}\boldsymbol{\pi}_0)'\mathbf{M}_{\mathbf{X}^A}\widehat{\mathbf{u}}^A=(\mathbf{X}\boldsymbol{\pi}_0)'\mathbf{M}_{\mathbf{X}^A}{\mathbf{y}}-(\mathbf{X}\boldsymbol{\pi}_0)'\mathbf{P}_{\mathbf{Z}}\mathbf{M}_{\mathbf{X}^A}{\mathbf{y}}$. Since the third relation shown in ((ref)) indicates that $(\mathbf{X}\boldsymbol{\pi}_0)'\mathbf{M}_{\mathbf{X}^A}{\mathbf{y}}=0$, $\widehat{\delta}=0$ is equivalent to $(\mathbf{X}\boldsymbol{\pi}_0)'\mathbf{P}_{\mathbf{Z}}\mathbf{M}_{\mathbf{X}^A}\mathbf{y}=0$. Hence, testing whether $\widehat{\Delta}$ approaches to zero in probability can be viewed as testing if the coefficient $\widehat{\delta}$ is significantly different from zero. Consider the test statistic
where $\mathbf{b}'=-\left[(\mathbf{M}_{\mathbf{X}^A}\mathbf{M}_{\mathbf{Z}}\mathbf{X}\boldsymbol{\pi}_0)'(\mathbf{M}_{\mathbf{X}^A}\mathbf{M}_{\mathbf{Z}}\mathbf{X}\boldsymbol{\pi}_0)\right]^{-1}\left[(\mathbf{X}\boldsymbol{\pi}_0)'\mathbf{P}_{\mathbf{Z}}\mathbf{X}^A\right]$, and $\widehat{\mathbf{V}}$ and $\widehat{\mathbf{V}}^A$ are consistent estimators of $\mathbf{V}$ and $\mathbf{V}^A$, respectively.
The proof of Theorem (ref) is presented in Appendix (ref). It is worth noting that Assumption 1 holds only when the instruments $\mathbf{z}_t$ are chosen carefully. More specifically, $\mathbf{z}_t$ should be correlated with the time-averaging term, $x_t^A$, but uncorrelated with $u_{t}^A$. This is to ensure Assumptions (ref)(ref) and (ref)(ref). Otherwise, the consistency of the 2SLS estimator may not be guaranteed. However, in practice, it is difficult to find such instruments. andreou suggested using all or part of high-frequency variables as instruments. However, they did not provide any practical guidance that is theoretically supported. In fact, with their suggested choice of instruments, it is possible that the chosen instruments are correlated with the error process. In this case, the 2SLS estimators would not be consistent, which may lower the power. In what follows, we shall propose a set of instruments that is theoretically valid for the DWH-type specification test. To derive theoretical properties, we assume following conditions on the instruments and the data generating process.
Under Assumption (ref), the low-frequency response variable $\{y_t\}$ is viewed as an MIDAS aggregation of the underlying high-frequency true process $\{y_{t-k/m}\}$, where $y_{t-k/m}=\beta_0+x_{t-k/m}\beta_1+u_{t-k/m}$. Note that $\{y_{t-k/m}\}$ is not observed in practice.
If we choose too many high-frequency lags as instruments, it might lead to a problem of a large number of weak instruments. As a consequence, the 2SLS estimator may be biased towards the NULL-LS estimator. The bias tends to get worse when there are more excessive number of instruments compared to the number of endogenous regressors. A brief explanation is presented by Greene. Based on the number of the parameters in ((ref)) and the consideration on possibly weak instruments, we shall construct $p=2$ instrumental variables, $\mathbf{z}_t=(z_{1,t},z_{2,t})',\ t=1\cdots,T$, as linear combinations of the high-frequency regressor. Inspired by Miller2016, we propose to choose weights of the instruments $\mathbf{z}_t$ as the following two decreasing sequences:
It is worth noting that these weights are designed to decrease exponentially and linearly fast. This is to mimic the behaviors of the MIDAS weights with exponential Almon lag and beta polynomials. Then the two instrumental variables can be written in a vector form as $\mathbf{z}'_t={\mathbf{x}_{t,m}^{(m)}}'\Upsilon$, where $\Upsilon=[\Upsilon_{1},\hspace{0.07cm}\Upsilon_{2}]$. The following theorem demonstrates that the proposed instruments are approximately valid when the frequency ratio is large.
The proof of Theorem (ref) can be found in Appendix (ref). Under both the null and the alternative, it is easy to see that {${z}_{r,t}$ is correlated with ${x}_t^A$. The main result of Theorem (ref) is that ${z}_{r,t}$ and ${u}_t^A$ are asymptotically uncorrelated when the frequency ratio is large, with the rate $E(z_{r,t}u_t^A)=O(m^{-1})$. Hence, the 2SLS estimator using our choice of the instruments is consistent when the frequency ratio $m$ is large. On the other hand, when $m$ is small, $T^{-1}{\mathbf{Z}}'\mathbf{u}^A$ converges, in probability, to a nonzero constant. Thus, the DWH specification test with our choice of instruments would only work when $m$ is large enough. This explains the low power of our test in finite samples when $m$ is small in the next section.
In this section, we examine finite sample sizes and powers of our method and two other comparable methods in literature: the second test presented in andreou (AGK, hereafter) and the unmodified VAT test in Miller2016. We first briefly introduce algorithms of the methods in comparison.
Algorithm 1 [Our Method]
Algorithm 2 [Miller's Method]
To make the results comparable, we use a simulation setting similar to the one proposed by Miller2016. At high-frequency level, data are generated with $y_{t-k/m}=x_{t-k/m}\beta+u_{t-k/m}$ for $t=1,\ldots,T$, $k=0,\ldots,m-1$. The high-frequency processes $\{x_{t-k/m}\}$ and $\{u_{t-k/m}\}$ are generated as stationary AR(1) processes given by $u_{t-k/m}=cu_{t-(k+1)/m}+\eta_{t-k/m}$ and $x_{t-k/m}=dx_{t-(k+1)/m}+\eta^*_{t-k/m}$, where $\{\eta_{t-k/m}\}$ and $\{\eta^*_{t-k/m}\}$ are i.i.d. $N(0,1)$. Let $\beta=10$. Denote $\mathbf{y}_{t,m}^{(m)}=(y_t,y_{t-1/m}, \cdots, y_{t-(m-1)/m})'$ and $\mathbf{u}_{t,m}^{(m)}$ be the unobserved high-frequency response and the error process between time $t-1$ and $t$. Let $\boldsymbol{\pi}_0=\mathbf{j}/m$ and $\boldsymbol{\pi}(\theta)=(\pi_1(\theta),\cdots,\pi_m(\theta))$, where $\pi_j(\theta)$ is defined in Assumption (ref)(ref). The low-frequency processes are generated as $y_{t}={\mathbf{y}_{t,m}^{(m)}}'\boldsymbol\pi(\theta)$ and $u_{t}={\mathbf{u}_{t,m}^{(m)}}'\boldsymbol\pi(\theta)$. Here, $\theta=\theta_0=0$ indicates the flat aggregation, which corresponds to the null. If $\theta\neq0$, the weights are no longer flat. Let $\theta=\theta_0+k$ where $k\in\{0.1,0.2,\cdots,1.9,2.0\}$ represent MIDAS-type alternatives. The nominal level is $0.05$. $R=2000$ Monte Carlo replications are generated. The sample sizes is $T\in\{125,512\}$. The frequency ratio is $m\in\{4,150,365\}$.
Table 1 presents the empirical sizes and the powers of our method, the AGK method and Miller's method when $c\in\{0,0.8\}$ and $d=0$. The results of more comprehensive settings are presented in Appendix (ref), which are consistent to what we observe in Table 1. When $k=0$, sizes closest to 0.05 are presented in boldface. In all our simulation settings, all methods seem to have reasonable sizes. Our and the AGK method tend to have more cases in which sizes are closer to 0.05, while Miller's unmodified VAT tends to slightly over-reject.
When $k\neq0$, empirical rejection rates represent powers of the tests. Powers less than 0.9 are shown in boldface. When $m$ is small, our method is not as powerful as the AGK method or Miller's unmodified VAT. These two methods have much better performance under all alternatives. For $T=125$, when the high-frequency error is AR(1), our method is less powerful when the effect size is small ($k\leq0.6$), whether the high-frequency error is i.i.d. or not. When $T=512$, the power of our method is not very large when the effect size is not large enough. This observation is consistent with Theorem (ref). When $m$ is small, the 2SLS estimator would not be consistent using the chosen instruments. As a matter of fact, if $m=4$, the two weighted functions in constructing the instruments are almost identical. Therefore, when $m$ is small, $m=4$, the AGK method seems to be good enough by choosing the most recent two high-frequency variables (out of four). Miller's unmodified VAT is another attractive alternative when $m$ is small since it is as powerful as the AGK method.
However, when $m$ is large, the effect of a careful choice of instruments is more visible. When $m$ is $150$, the power of he AGK method never exceeds 0.90 for all alternatives. In the meantime, our method tends to have higher power under almost all alternatives. Miller's unmodified VAT tends to be just a little less powerful than our method for small effective sizes. Additionally, as the sample size increases ($T=512$), the AGK method becomes more powerful for large local alternatives, while all three methods reduce the power when the effective sizes are small. Except for a few small effect sizes with the AR(1) high-frequency error process, our method has the highest power for most cases. Similar conclusions can be drawn for $m=365$.
In this paper, we considered a DWH test to choose between the time-averaging models and MIDAS models. For the DWH test, the instruments need to be carefully chosen to avoid the problems involved with weak instruments and correlation with the error terms. However, there had not yet been a rigorous work regarding the proper choice of instruments. The main contribution of this paper is that a set of instruments has been proposed with a theoretical validation. In particular, the proposed instruments would only work when the frequency ratio is large enough. The Monte Carlo simulations reconfirm our theoretical findings. The DWH test with our proposed instruments is more powerful in finite samples compared to the one with a less careful choice of instruments. However, this is only the case when the frequency ratio is large enough. Therefore, our proposed specification test would be useful when handling two extremely different sampling frequencies such as monthly versus hourly observations. On the other hand, if the frequency ratio is very small, taking a few most recent high-frequency variables as the instruments or taking Miller's approach would be better.
The main purpose of this paper is to provide an insight on a proper choice of instruments. To keep the exposition concise, we limited the scope of the paper using somewhat strong assumptions. Now that we understand the behavior of the instruments better, an extension of this paper to accommodate more than one regressors and general data generating process is underway.
The authors are grateful to J. Isaac Miller and to the participants of the 48th annual meetings of Illinois Economics Association for constructive comments and suggestions. Superior, a high-performance computing infrastructure at Michigan Technological University, was used in obtaining the Monte Carlo simulation results. This work was partially supported by NSF grant CPS-1739422.