Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
53,580 characters · 13 sections · 32 citation commands
\justifying \sloppy \nohyphens
\stepcounter{section} In economics, the dependence of a variable $Y$ (the endogeneous variable) on another variable(s) $X$ (the exogeneous variable) is rarely instantaneous and very often, $Y$ responds to $X$ after a lapse of time. These lags can be due to a variety of factors, including the time it takes for individuals and businesses to adjust their behaviour in response to the change or the time it takes for the data to be collected and processed. Distributed Lag (DL) and Autoregressive Distributed Lag (ADL) models are commonly used for lagged relationships when endogenous and exogenous variables are observed at the same frequency. However, there are many practical situations where variables are observed at mixed frequencies, limiting the direct applicability of these models. For instance, daily job postings (high frequency) and monthly unemployment rate (low frequency) provide understanding of labour market dynamics. Relationship between inflation (monthly) and GDP (quarterly) provide an insight into economic condition of a country.
Approaches such as data averaging and State space model are designed to handle situations involving mixed data frequencies. But the first one involves “pre-filtering of data", due to which a lot of potentially useful information might be discarded. The second one consists of a system of two equations, a measurement equation which links observed series to a latent state process, and a state equation which describes the state process dynamics. The system of equations therefore can require a lot of parameters, for the measurement equation, the state dynamics and their error process. Therefore, state space model estimation can suffer from identification problems and high computational complexity. Hence, suggesting an alternative approach becomes important.
As an alternative, Mixed Data Sampling (MIDAS) regression model was introduced by article1. It provides flexibility in handling data sampled at different frequencies and for a straightforward forecast of a low-frequency variable based on lagged low-frequency and high-frequency data. This area attracted the attention of many researchers in early twenties. article2 enriched the MIDAS literature by introducing more general mixed-data structures, non-linearities, unequally spaced observations, and multiple equations.
MIDAS models were typically estimated via nonlinear least squares (NLS) (article3). Later, article15 estimated MIDAS regression models via Ordinary Least Square (OLS) with polynomial parameter profiling which appeared to be more appealing in terms of computation and forecasting. In estimation of macro indicators, generally advance estimates are released due to time consuming estimation process and subsequent revisions are released after a set pattern of lag. Resulting from measurement processes, these indicators are subject to error, which can bias results and reduce reliability. Therefore, corrective techniques have been proposed to improve robustness.
Although, in the presence of Measurement Error (ME) in the data, the linear model resembles a conventional regression model but the estimator so obtained, is inconsistent. In the existing literature, consistent estimators for regression coefficients are put forth, assuming the presence of some additional information (article42,article43, article22, inbook33 , article36,article46, and inbook30). article49 studied the robustness of the linear regression under ME models with replicated observations by assuming the ME to be normally distributed. Further, article45 provided a consistent estimator in the functional and structural ME model without assuming any distributional form of ME . Later, article50, article40, article39, article51,article52, article54 developed various statistical methods for analyzing data with measurement error. SINGH2012198, article53 proposed consistent estimators of regression coefficients, using linear restrictions in replicated measurement error model. As per the models containing lagged terms, article24 have addressed estimation of the parameters for linear autoregressive models in presence of additive and uncorrelated measurement errors.
To the best of our knowledge, the work on estimation of parameters of ADL-MIDAS (Autoregressive Distributed Lag - Mixed Data Sampling) model in the presence of ME has not been taken up so far. In this paper, we shall explore the effect of the presence of measurement error in both exogeneous and endogeneous variables on profile estimator of MIDAS regression model suggested by article15. Using profile log-likelihood approach along with corrected score methodology (article25), a consistent estimator for ADL-MIDAS is proposed.
The paper is structured into six sections. Section 1 is Introductory. Section 2 specifies the MIDAS model with its existing profile likelihood estimator. ADL-MIDAS ME model is introduced and the effect of ME on parameter estimation of the ADL-MIDAS model is shown in Section 3. Section 4 deals with corrected estimators of the parameters of ADL-MIDAS ME model. Section 5 delves into numerical efficiency, demonstrated through simulations. The conclusions are presented in Section 6.
\stepcounter{section} Let $\mathcal{Z}_{t}$ be an endogeneous variable sampled at some fixed, say annual, quarterly or monthly, sampling frequency and $\xi_{t} $ be an exogeneous variable sampled $m$ times faster than $\mathcal{Z}_{t}$ which means that if $\mathcal{Z}_{t}$ is sampled $T$ times then $\xi_{t} $ is sampled $mT$ times.Then, the class of ADL-MIDAS regression is obtained, when the endogeneous variable depends linearly on its own previous values in MIDAS model. By incorporating an autoregressive component of order $p$, the ADL-MIDAS model can be expressed as:
for $t = p,p+1,...,(T-1)$. Here,
An essential aspect of MIDAS model is the parsimonious parameterization of the lagged coefficients $c(k; \theta)$. If the parameters of the lagged polynomial are left unrestricted (known function $C(.)$ does not depend on $\theta$), then there would be many parameters to be estimated. As a way of addressing parameter proliferation in MIDAS regression, the coefficients of the MIDAS model are captured by a known function $ C(L^{1/m};\theta) $ of few parameters summarized in a vector $\theta$. Hence, to avoid parameter proliferation in case of long high-frequency lags, functional lag polynomials have been proposed. article2 provide an in-depth discussion regarding the specification of various polynomials, including the Beta polynomial, Almon lag polynomial, and Step functions.
In the matrix form, (ref) can be written as
where,
with $ \xi_{t}(\theta) = C(L^{1/m};\theta)\xi_{t}$ and $ \boldsymbol{\beta} = (a, \rho_{1},..., \rho_{p}, b)'. $
Assuming that $ \epsilon_{t} $ are iid $ N(0,\sigma^{2}_{\epsilon}) $, the log-likelihood can be written as
The estimator for parameters of (ref) are provided by article15 using profile likelihood approach by concentrating the likelihood with repsect to $\sigma_{\epsilon}^{2}$ first $\left( \text{providing } \hat{\sigma}_{\epsilon}^{2}=\boldsymbol{\mathcal{E}'\mathcal{E}}/T\right)$ and then obtaining $\boldsymbol{\hat{\beta}}_{\bar{\theta}}$ by fixing $\theta=\bar{\theta}$. Resultant estimator $\hat{\boldsymbol{\beta}}_{\bar{\theta}}$ has the same functional form as OLS estimator. Substituting back $\boldsymbol{\hat{\beta}}_{\bar{\theta}}$ into likelihood function (ref) and optimising for $\theta$, the resultant joint estimator for $(\boldsymbol{\beta},\theta,\sigma_{\epsilon}^{2})$ is provided as:
where, $\boldsymbol{\hat{\mathcal{E}}} = \left[\boldsymbol{\mathcal{Z}}-\boldsymbol{\Psi}(\hat{\theta})\hat{\boldsymbol{\beta}}(\hat{\theta})\right]$. Since profile estimator (ref) is obtained by maximising $\mathcal{L}$, we have $\dfrac{\partial \mathcal{L}}{\partial \hat{\boldsymbol{\gamma}}}=0$, where $\hat{\boldsymbol{\gamma}}=(\hat{\boldsymbol{\beta}}',\hat{\theta},\hat{\sigma}_{\epsilon}^{2})'$ is the profile estimator for the true value $\boldsymbol{\gamma}=(\boldsymbol{\beta}',\theta,\sigma_{\epsilon}^{2})'$. By Taylor expansion of gradient around the true value $\boldsymbol{\gamma}$, the following expression is obtained:
where, $\tilde{\boldsymbol{\gamma}}$ is on the line segment between $\boldsymbol{\gamma}$ and $\hat{\boldsymbol{\gamma}}$.
article15 proved the consistency of $\hat{\boldsymbol{\gamma}}$ by showing that $\dfrac{1}{T}\left(\dfrac{\partial \mathcal{L}}{\partial \boldsymbol{\gamma}}\right)\xrightarrow{}0$ and $\dfrac{1}{T}\left(\dfrac{\partial^{2} \mathcal{L}}{\partial \boldsymbol{\gamma} \partial \boldsymbol{\gamma}'}\right)$ tends to deterministic matrix as $T\xrightarrow{}\infty$.
\stepcounter{section}
In this section, the impact of measurement error on estimation of ADL-MIDAS estimator is discussed.
Under ADL-MIDAS model (ref), assume that the endogeneous and exogeneous variables $\mathcal{Z}_{t}$ and $\xi_{t}$ are unobservable and can be observed through $y_{t}$ and $x_{t}$ with additional measurement errors $u_{t}$ and $v_{t}$ respectively, such that
Substituting (ref) in (ref), we get ADL-MIDAS ME model, which can be written as:
In matrix form, (ref) can be written as
Where,
with $ v_{t}(\theta) = C(L^{1/m};\theta)v_{t}$ and $ x_{t}(\theta) = C(L^{1/m};\theta)x_{t}$. By substituting (ref) in (ref), we get
Where, $\boldsymbol{\mathcal{T}} = \boldsymbol{\mathcal{E}} + \boldsymbol{U} - \boldsymbol{V}(\theta)\boldsymbol{\beta}$. The model given by (ref) is ADL-MIDAS ME model.
Since, $\boldsymbol{Y}$ and $\boldsymbol{X}(\theta)$ are observed in place of $\boldsymbol{\mathcal{Z}}$ and $\boldsymbol{\Psi}(\theta)$, if presence of ME is ignored and existing estimator is used, then (ref) and (ref) can be written as
and
where, $\hat{\boldsymbol{\mathcal{T}}} = \left[\boldsymbol{Y}-\boldsymbol{X}(\hat{\theta}_{e})\hat{\boldsymbol{\beta}}_{e}(\hat{\theta}_{e})\right]$.
The estimator given in (ref) is consistent in the absence of measurement error. However, it is interesting to check whether the same estimator, that is, (ref) obtained using ME contaminated data still possesses the property of consistency or not.
Under ME contaminated data, the original likelihood function $\mathcal{L}$ converts to $\mathcal{L^{*}}$ given by (ref). Using $\mathcal{L^{*}}$ instead of $\mathcal{L}$ in (ref) and following the approach of article15, this will show whether there is any change in the asymptotic properties of the profile estimator after introduction of ME.
For simplicity segregating data from parameters, $\mathcal{L^{*}}$ given in (ref) can be written as:
Where, $ \boldsymbol{X}_{(T-p)\times (j^{max}+p+1)} $ = $
$ and $ \boldsymbol{\beta}_{M} = (a, \rho_{1},..., \rho_{p}, bc(0, \theta),..., bc(j^{max}-1,\theta))' $. The matrices $\boldsymbol{\Psi}$ and $\boldsymbol{V}$ can also be defined on similar lines as $\boldsymbol{X}$ with order $(T-p)\times(j^{max}+p+1)$.\\\\
Using $\mathcal{L^{*}}$ given by (ref) in (ref) and denoting the profile estimator obtained with measurement error as $\boldsymbol{\hat{\gamma}}_{e} = (\hat{\boldsymbol{\beta}}'_{e},\hat{\theta}_{e},\hat{\sigma}_{\epsilon(e)}^{2})'$, we obtain
where, $\boldsymbol{\gamma}= (\boldsymbol{\beta}',\theta,\sigma_{\epsilon}^{2})'$.
To study the asymptotic behaviour of $\boldsymbol{\gamma}$ in presence of measurement error, we assume the following:
Using the assumptions and results in the appendix along with (ref), it is observed that
This indicates that $ plim \hspace{1mm}\hat{\boldsymbol{\gamma}}_{e} \ne \boldsymbol{\gamma}$. Thus, the estimator (ref) looses the property of consistency when data is measured with errors. It is intersting to note that, $\boldsymbol{\Sigma}=0$ leads to original ADL-MIDAS model without measurement error. In this scenario, both elements on the right-hand side of equation (ref) will be equal to zero, thereby preserving the consistency of the estimator in the absence of measurement error.
In the subsequent section, a consistent estimator for ADL-MIDAS ME model (ref) is proposed.
\stepcounter{section}
Using (ref), the log-likelihood function (ref) in the presence of measurement error can be written as
Taking conditional expectation on both sides given $ \boldsymbol{\mathcal{Z}}$ and $\boldsymbol{\Psi}$ and using assumptions, we get
Using assumptions, it is evaluated that $ E\left[\boldsymbol{U'U}\right] = (T-P)\sigma_{u}^{2} $ and $ E\left[\boldsymbol{V}(\theta)'\boldsymbol{V}(\theta)\right] = (T-P) \boldsymbol{\Sigma}_{c} $ where, $ \boldsymbol{\Sigma}_{c} $ = $
,$ $ \boldsymbol{\Sigma_{v}} = \sigma_{v}^{2} \boldsymbol{I}_{j^{max} \times j^{max}}$, $\boldsymbol{C}$= $
'$ and $\boldsymbol{I} $ is an identity matrix. Thus,
The inconsistency of estimator (ref) arises because of the presence of the second term on right hand side of (ref) which arises due to measurement error. To address this, the log-likelihood function is modified using the corrected score methodology proposed by article25, and a new consitent estimator is obtained. Corrected log-likelihood is written as
It is interesting to note that the right hand side of (ref) contains measurement error variances $\sigma_{v}^{2}$ and $\sigma_{u}^{2}$ in addition to parameters of interest. For measurement error contaminated data, it is well documented in literature that the usual estimators are inconsistent and additional information is required for obtaining consistent estimators. Here the consistent estimator is proposed under the assumption of known $\sigma_{v}^{2}$ and $\sigma_{u}^{2}$. The variance of measurement errors may be available from past experience of investigator and/or similar studies conducted in the past etc.
In (ref), using parameter profiling method, that is, by concentrating the likelihood function first with respect to $\sigma_{\epsilon}^{2}$, we get
Substituting (ref) in (ref) and differentiating with respect to $\boldsymbol{\beta}$ for fixed $\theta=\bar{\theta}$, the resultant estimator is
In the above estimator, MIDAS polynomial parameter $'\theta'$ is profiled out and a closed form solution is obtained. Substituting estimator of $\boldsymbol{\beta}$ that is $\hat{\boldsymbol{\beta}}_{c}\bar{(\theta)}$ into (ref), the maximization of (ref) with respect to $\theta$ reduces to maximizing
On further simplification, we obtain
Hence, the optimized estimators of $\theta$, $\boldsymbol{\beta}$ and $\sigma_{\epsilon}^{2}$ are obtained as
In the preceeding discussion, maximization is subjected to the constraint that the regressors $\boldsymbol{X}(\theta)$ are chosen from a set of polynomials such as Beta or others, with weights that collectively sum up to one.
The large sample properties of estimators (ref) are discussed in the next subsection.
Denote by $\boldsymbol{\hat{\gamma}}_{c} = (\hat{\boldsymbol{\beta}}'_{c}, \hat{\theta}_{c}, \hat{\sigma}_{\epsilon c}^{2})'$, the corrected profile estimator of true parameter $\boldsymbol{\gamma}=(\boldsymbol{\beta}',\theta,\sigma_{\epsilon}^{2})'$. As estimator $\boldsymbol{\hat{\gamma}_{c}}$ is obtained by maximizing (ref), therefore $\frac{\partial \mathcal{L}_{c}^{*}}{\partial \hat{\boldsymbol{\gamma}_{c}}} =0$. Considering $\mathcal{L^{*}}$ given in (ref) and $\mathcal{L}_{c}^{*} $ given in (ref) (with expression $\boldsymbol{\beta'\Sigma_{c}\beta}=\boldsymbol{\beta}_{M}'\boldsymbol{\Sigma\beta}_{M}$). Following the steps of article15 and using Taylor expansion of gradient around the true value $\boldsymbol{\gamma}$, we get
where $\tilde{\boldsymbol{\gamma}}$ is on the line segment between $\boldsymbol{\gamma}$ and $\hat{\boldsymbol{\gamma}}_{c}$.
From appendix, it is clear that
In (ref), the first term converges to a deterministic matrix, whereas its second term converges to zero. Hence, the estimator (ref) so obtained using corrected log-likelihood function is consistent. The above results are summarised in the following Proposition,
The Monte-carlo simulation study has been carried out to demonstrate the effect of measurement error on the existing estimator $(\hat{\boldsymbol{\beta}}_{e}',\hat{\theta}_{e},\hat{\sigma}_{\epsilon(e)}^{2})$ and to evaluate the performance of corrected estimator $(\hat{\boldsymbol{\beta}}'_{c},\hat{\theta}_{c}, \hat{\sigma}_{\epsilon c}^{2})$.
In simulations, Beta polynomials introduced by article2, are employed to constrain coefficients within the MIDAS model. It relies on the Beta probability density function, which encompasses two parameters, specified as following:\\\\ $c(j;\theta_{1},\theta_{2})= \dfrac{f(\frac{j}{j^{max}},\theta_{1};\theta_{2})}{\sum_{j=0}^{j^{max}-1}f(\frac{j}{j^{max}},\theta_{1};\theta_{2})}$;\\\\ $f(x,\theta_{1},\theta_{2})=\dfrac{1}{B(\theta_{1},\theta_{2})}x^{\theta_{1}-1}(1-x)^{\theta_{2}-1}$, for any $\theta_{1}>0,\theta_{2}>0$\\\\ where $B(\theta_{1},\theta_{2})=\dfrac{\Gamma(\theta_{1})\Gamma(\theta_{2})}{\Gamma(\theta_{1}+\theta_{2})}$ and $\Gamma(\theta)=\int_{0}^{\infty}e^{-x}x^{\theta-1}dx$.
For ease of use, we confine $\theta$ to one-dimensional space, however, it is not necessary. Specific case of the MIDAS Beta polynomial by taking $\theta_{1}=1$ involves only one parameter. Estimating the single parameter $\theta_{2}$ with the restriction that it be larger than one, yields single-parameter downward sloping weights which are more flexible than exponential or geometric decay patterns (article15).
In this experiment, monthly (high frequency) and quarterly data (low frequency) are generated. The monthly data exhibits an AR(1) pattern, characterized by an autocorrelation of 0.8. Quarterly data is generated using MIDAS model given by (2.3) with $p=2$ and parameters fixed as $a=0$, $\rho_{1}=0.3,\rho_{2}=0.2$ and $b=1$. The measurement errors in monthly and quarterly data are generated using $N(0,\sigma_{v}^{2})$ and $N(0,\sigma_{u}^{2})$ and measurement error contaminated series are generated using (3.1). For simulation, various combinations of $T, j^{max}, \theta, \sigma^{2}_{v}, \sigma^{2}_{u}$ such as $T=(12, 24, 36, 48, 60, 72 ,96, 120)$, $j^{max}= (3, 6, 9, 12, 15, 18, 21, 24)$, $\theta= (2, 5, 10)$, $\sigma^{2}_{v}=(0.5,1,1.5)$ and $\sigma^{2}_{u}=(0.5,1,1.5)$ are considered.
The simulations are run for 1000 replications, where for each iteration $(\hat{\boldsymbol{\beta}}'_{e}, \hat{\theta}_{e}, \hat{\sigma}_{\epsilon(e)}^{2})$ and $(\hat{\boldsymbol{\beta}}'_{c},\hat{\theta}_{c}, \hat{\sigma}_{\epsilon c}^{2}) $ are estimated. Optimization of the MIDAS parameter $'\theta'$ is carried out using Golden section search methodology employing 50 iterations. One can go for more iterations but, optimum convergence is attained till 50 iterations. For each iteration, Square Error Matrix $(SEM)_{i}=(\hat{\boldsymbol{\beta}}_{i}-\boldsymbol{\beta})(\hat{\boldsymbol{\beta}}_{i}-{\boldsymbol{\beta}})'$ is computed.
In case of measurement error, consistent estimators may not have finite expectations (article31), therefore rather than relying on empirical expectations, empirical medians are used. Accordingly, the following measures are used for comparing estimators
For comparing the performance of $(\hat{\boldsymbol{\beta}}_{e}', \hat{\theta}_{e}, \hat{\sigma}_{\epsilon (e)}^{2})$ and $(\hat{\boldsymbol{\beta}}'_{c}, \hat{\theta}_{c}, \hat{\sigma}_{\epsilon c}^{2})$ as $T$ increases, the plots of NMedB, medB$(\theta)$ and medB$(\sigma_{\epsilon}^{2})$ are provided in Figures 1(a)-1(f).
Figure 1 shows that in case of the existing estimator, although NMedB and medB$(\theta)$ stablizes after decline, they do not tend towards zero as the sample size ($T$) increases. An important observation is that medB$(\sigma_{\epsilon}^{2})$ increases with an increase in sample size for the naive estimator and is near to zero for all sample sizes for the proposed estimator. The consequences are severe for estimator $\hat{\theta}$ when value of $\theta$ is large as can be observed from Figures 1(b) and 1(e). This confirms the theoretical finding that the naive estimator ($\hat{\boldsymbol{\beta}}'_{e},\hat{\theta}_{e}, \hat{\sigma}_{\epsilon(e)}^{2}$) are inconsistent in the presence of measurement contaminated data. On the other hand, the NMedB, medB$(\theta)$ and medB$(\sigma_{\epsilon}^{2})$for the proposed estimator ($\hat{\boldsymbol{\beta}}'_{c},\hat{\theta}_{c}, \hat{\sigma}_{\epsilon c}^{2}$) tend towards zero as sample size increases.
For observing the effect of $j^{max}$ on the proposed estimator $(\hat{\boldsymbol{\beta}'_{c}},\hat{\theta}_{c},\hat{\sigma}_{\epsilon c}^{2})$, Figures 2(a)-2(d) depict graphical presentation of NMedB, trMedSEM, medB$(\theta)$ and medB$(\sigma_{\epsilon}^{2})$ with respect to sample size for various values of $j^{max}$.
Figures 2(a) and 2(b) suggest that as sample size increases, the bias as well as variability of $\hat{\boldsymbol{\beta}}_{c}$ decline. However, for higher $j^{max}$, the values are higher, although the gap reduces with an increasing sample size. A similar trend is observed for medB$(\sigma_{\epsilon}^{2})$ up to a sample size of 24. But, for sample size greater than 24, the biases are almost equal and tend towards zero. Also, medB$(\theta)$ of $\hat{\theta}_{c}$ declines with an increase in sample size. However, the gap remains larger at higher sample sizes. This motivates us to further examine the effect of $j^{max}$ vis-a-vis $\theta$ and $T$ .
Examining the effect of various values of $\theta$ and $j^{max} $ on the characteristics of the proposed estimator, Figure 3(a) suggests that the NMedB increases with an increase in value of $j^{max} $. On the other hand, the higher the value of $\theta$, the lower will be NMedB of the proposed estimator, although the overall trend remains nearly same with respect to $j^{max} $. In context of the MIDAS regression model, $\theta$ plays a crucial role by assigning weights to the lagged values of exogeneous variables. It is important to note that using the weights based on Beta polynomials suggested by article2, the higher value of $\theta$ results in more weight being given to recent past values of $x_{t}$, while a lower value of $\theta$ also assigns substantial weights to distant past values of $x_{t}$ . It is interesting to note from Figure 3(c) that an increase in value of $j^{max}$ has serious consequences on the bias of $\hat{\theta}_{c}$, more particularly when the value of underlying parameter $\theta$ is high. It can also be noticed from Figure 3(d) that for lower value of $\theta$, the value of medB$(\sigma_{\epsilon}^{2})$ of the proposed estimator is small.
From Figure 4, it is evident that with an increase in the number of lags of $x_{t}$, NMedB, trMedSEM and medB$(\theta)$ increase. However, for the higher sample size, the bias and variability are comparatively lower. Measure of $\sigma_{\epsilon}^{2}$, medB$(\sigma_{\epsilon}^{2})$ shows different pattern. Figure 4(d) illustrates that as $j^{max}$ increases, medB$(\sigma_{\epsilon}^{2})$ decreases and then stabilizes. Conversely, with a larger sample size, the medB$(\sigma_{\epsilon}^{2})$ is also higher. From application point of view, although the selection of number of lags is based on some criteria like AIC, BIC etc., a parsimonious model is preferred. Employing large number of lags can lead to model complexity and computational burden. Viewing the need of parsimonious model with observations from Figures 3 and 4, the importance of selecting optimum number of lags of $x_{t}$ cannot be overlooked.
Evaluating the effect of magnitude of measurement error, the results with large measurement error variance are compared with low measurement error variance. Apart from this, the effects of measurement error in low frequency and high frequency variables are also examined separately. The effects of $\sigma_{u}^{2}$ and $\sigma_{v}^{2}$ on NMedB, trMedSEM, medB$(\theta)$, and medB$(\sigma_{\epsilon}^{2})$ are discussed, with more detailed results provided in Tables 1–4.
Comparing the values of NMedB, trMedSEM, medB$(\theta)$ and medB$(\sigma_{\epsilon}^{2})$ from Table 2 (higher $\sigma_{u}^{2}$) and Table 3 (higher $\sigma_{v}^{2}$) with corresponding values of Table 1 (low $\sigma_{u}^{2}$ and $\sigma_{v}^{2}$), it is interesting to note that measurement error in low frequency variable has higher amplification effect than that of high frequency variable on the properties of $\hat{\boldsymbol{\beta}_{c}}$. On the other hand, measurement error in high frequency variable have higher amplification effect on the properties of estimators $\hat{\theta}_{c}$ and $\hat{\sigma}_{\epsilon c}^{2}$.
This paper considers the MIDAS model which is a useful model for handling mixed frequency data. article15 proposed a profile likelihood estimator for the MIDAS model which simplifies and expedites computations. Our study shows that in presence of the measurement error contaminated data, the profile estimator is inconsistent. For such situations, we propose a new profile estimator using corrected score methodology, assuming the prior knowledge of measurement error variance. The consistency of the proposed estimator is established. Asymptotic properties of the proposed estimator have been explored. It is shown that the proposed estimator asymptotically follows normal distribution.
The small sample properties of the proposed estimator are also explored using simulations. It is observed that as sample size increases, the bias and variability of proposed estimator decline towards zero, whereas, for the estimator of article15, they stabilize above zero after decline. It is observed that for given sample size, the selection of lags of high frequency variable to be included in the model should be judicious as the higher number of lags inflate the bias and the variability of proposed estimator.
It is observed that the higher is the variance of measurement error, the higher will be the bias and variability in proposed estimator. The variance of measurement error in low frequency variable has higher amplification effect than measurement error variance of high frequency variable on estimator of regression coefficient ($\hat{\boldsymbol{\beta}}_{c}$). On other hand, the measurement error variance of high frequency variable has higher amplification effect on estimator of variance of equation error ($\hat{\sigma}_{\epsilon c}^{2}$) and hyperparameter ($\hat{\theta}_{c}$).
It is observed that, when distant past values of high frequency variables have substantial weights, the bias and variability of regression coefficient ($\hat{\boldsymbol{\beta}}_{c}$) are higher whereas $\hat{\theta}_{c}$ and $\hat{\sigma}_{\epsilon c}^{2}$ show lower bias. The situation is reversed when only recent past values of high frequency variables have substantial weights.
The first author acknowledges the University Grants Commission (UGC), Government of India, for providing financial assistance to conduct this research.