Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
69,885 characters · 18 sections · 97 citation commands
{.26in} \thispagestyle{empty}
This paper focuses on the Generalized Covariance Estimator proposed by gourieroux2022generalized. Extending the properties of this estimator to the misspecification cases can give us access to make inference in a large class of non-Gaussian non-linear time series models, such as causal-noncausal processes, Double Autoregressive (DAR) models, or mixed SVARs. Furthermore, we propose a test based on the GCov estimator, which does not rely on any distributional assumption for testing nested, overlapping, and non-nested hypotheses based on the properties of the estimator under misspecification. This test can contribute to model selection. Moreover, we extend the use of the GCov estimator by introducing a constrained GCov (CGCov) estimator. This estimator is useful for a broad range of models with constraints on the parameters, such as ARCH-GARCH and DAR models.
Misspecification is an inevitable issue in econometrics. The source of misspecification could come from parametric or non-parametric aspects of models and estimators. In the parametric part, we may encounter a misspecified model space; for instance, if your data has an ARMA(1,1) nature, but you fit ARCH-GARCH models. Another source of misspecification is order selections, where there is always a chance of overfitting or underfitting the true model. In parametric estimators, the parametric assumptions can also cause misspecification issues. For instance, consider a model with non-Gaussian errors in which the model parameters are estimated with Gaussian MLE.
Under misspecification, there are several challenges. The first challenge is making an inference. The asymptotic normality of the estimators and the variance are usually developed under correct specification (or we call it under the null); however, these results may not hold under misspecification. Recently, bonhomme2022minimizing suggested an approach for making inference in local misspecification. The second challenge is any hypothesis testing, such as a simple T-test, Wald test, likelihood ratio, non-nested tests, etc. All the asymptotic results of the well-known estimators are based on the correct specification, and it is possible that those could not be valid under misspecification. One may want to select a model among multiple model spaces; in that case, granger1995comments suggested that using information criteria like Akaike or BIC is more useful than testing different model spaces against each other. The other one may want to eliminate one model or only compare two; testing the model spaces is more effective. There is a vast literature on non-nested tests, including the Cox test [cox1961tests,cox1962further], J test [davidson1981several, davidson1983testing, davidson1984model], JA test [fisher1981alternative], encompassing test [gourieroux1983testing], and Vuong test [vuong1989likelihood, shi2015nondegenerate]. In this paper, we focus on the Wald-type and score-type tests proposed by gourieroux1982likelihood, which can be applied to both nested and non-nested cases. \footnote{For letliture review on non-nested tests see gourieroux1995testing and pesaran2001non.}
Alternatively, the problem of interest should not be limited to specifying the models; it could also involve specifying the distribution of the time series or the number of lags to consider. Specifically, to select the order of non-causality in the causal non-causal literature, the existing method based on the information criteria is misspecified [see gourieroux2018misspecification]. Therefore, depending on the problem of interest, the properties of an estimator under misspecification can be a tool to address such issues.
The parametric misspecification can be extended to models with constraints on the parameter space. The estimation of the parameters of interest on the boundaries needs more attention since we lose the asymptotic normality properties of the estimator, and this causes problems for inference or hypothesis testing. This is a well-developed problem in ARCH-GARCH models and estimators such as Maximum Likelihood or GMM [gourieroux1982likelihood, andrews1999estimation, andrews2001testing, francq2007quasi, francq2009testing, cavaliere2022bootstrap, cavaliere2024specification]. Here we develop the asymptotic properties of the GCov estimator when we are on the boundary of the parameter space, both under correct parametric estimation and misspecification. We then demonstrate that the GCov specification test provided by gourieroux2022generalized does not follow a chi-square distribution, and we need to use the bootstrap-based GCov test proposed by jasiak2023gcov. This development contributes to the estimation of constrained models without having a parametric assumption on the distribution of the error, such as DAR models.
The properties of the GCov estimator under misspecification and constraints extend the use of the GCov estimator and test statistics in nonlinear models such as causal-noncausal and DAR models. The causal-noncausal processes are useful to model time series with bubble patterns both in univariate [ giancaterini2025inference, de2025forecasting, hecq2020mixed, hecq2021forecasting, hecq2025non] and multivariate [cubadda2023detecting, cubadda2024optimization, davis2020noncausal, lanne2013noncausal, gourieroux2017noncausal, gourieroux2022generalized] frameworks. Based on these processes, we can detect the bubble periods [giancaterini2025bubble, blasques2025novel] and build portfolios that take advantage of bubble periods [hall2024modelling, giancaterini2025regularized].
This paper contributes to the estimation and specification tests of DAR models. Here we extend the traditional DAR(p) models proposed by ling2004estimation and ling2007double and use the augmented DAR(p,q) presented by jiang2020non. Based on the developments of the GCov estimator presented in this paper, we can extend the estimation of DAR models under correct specification and misspecification from QML [zhu2013quasi, li2023maximum] to a semiparametric approach and consequently provide robust specification test and model selection test in a more general DAR(p,q) framework and allowing the parameters be on the boundary on the constraint parameter set.
Outline: The rest of the paper is as follows: Section 2 briefly covers the GCov estimator and specification test. In Section 3, we develop the asymptotic properties of the GCov estimator under misspecification and discuss model selection tests. Section 4 introduces the constrained GCov estimator. Section 5 illustrates the finite sample properties of the proposed tests and estimators in the context of causal-noncausal and DAR models. Section 6 presents two real-world applications utilizing the consumer price index by final energy demand and the US 3-month Treasury bill. We conclude in Section 8.
Compared to the parametric approach, utilizing semi-parametric methods such as the Generalized Covariance estimator for estimating coefficients of noncausal processes has several benefits. gourieroux2017noncausal, gourieroux2022generalized propose a new semi-parametric method called the Generalized Covariance estimator (GCov), which is asymptotically consistent and normally distributed with known variance under the correct specification of the parametric model and non-parametric part of the estimator, considering (non)linear transformations of the residuals.
Let's consider the following stationary process within a semi-parametric model framework:
where, $\underline{Y_t} = (Y_t, Y_{t-1}, \ldots)$, and $u_t$ is an i.i.d. sequence. We assume that the function $g$ is known, while $\theta$ is an unknown parameter vector. The GCov estimator for estimating the vector $\theta$ is defined as follows:
where
Here, $\hat{\Gamma}^a(h; \theta)$ represents the covariance function between $a(g(\underline{Y_t}; \theta))$ and $a(g(\underline{Y_{t-h}}; \theta))$ and $a(.)$ includes transformations.
The GCov estimator is useful for estimating the parameters of nonlinear models in non-Gaussian frameworks. Recently, the GCov estimator has been used to estimate the parameters of the causal-noncausal models [gourieroux2022generalized, jasiak2023gcov]. To identify the univariate causal-noncausal process, consider the following process:
where the error term $\epsilon_t$ is non-Gaussian and i.i.d. sequence. The non-Gaussianity assumption is for the identification of the noncausal part from the causal part. The polynomial $\Phi(L)$ in the lag polynomial of order $r$ is backward-looking. However, in these models, we have the lead polynomial $\Psi(L^{-1})$ of order $s$, which is forward-looking and is the deviation from the traditional pure causal autoregressive. We can express the nature of this kind of model by focusing on the roots of causal and noncausal polynomials, which are outside and inside the unit circle, respectively.
Example 2.1: If a MAR(1,1) model is fitted to \(y_t\), defined as
\[(1 - \phi L)(1 - \psi L^{-1})y_t = \epsilon_t,\]
where the errors \(\epsilon_t\) are independent and identically distributed, satisfying \(E(|\epsilon_t|^{\delta}) < \infty\) for \(\delta > 0\), and the parameters \(\phi\) and \(\psi\) are two autoregressive coefficients that are strictly less than one. In this case, the parameter vector is defined as \(\theta = (\phi, \psi)'\).
This category of models can be extended to the causal-noncausal VAR models. Two sets of identifications exist for the mixed VAR process. The first one is proposed by lanne2013noncausal and follows the univariate representation
where $ \Phi(L)=I_n-\Phi_1 L- \Phi_2 L^2-...-\Phi_r L^r$ and $\Psi(L^{-1})=I_n - \Psi_1 L^{-1}- \Psi_2 L^{-2}-...-\Psi_sL^{-s}$. The condition here is $det\Phi(z)\neq 0$ for $|z|<1$ and $det\Psi(z)\neq 0$ for $|z|<1$. The second representation proposed by gourieroux2017noncausal and davis2020noncausal considers only the lag polynomial and allows the roots of the polynomial to be inside or outside of the unit circle. lanne2013noncausal give an example indicating these models are non-nested [see also giancaterini2023essays and gourieroux2024nonlinear].
gourieroux2022generalized propose a portmanteau test based on the GCov estimation, which has an asymptotic chi-square distribution. Consider the objective function we minimize in (ref) at the estimated parameter $\hat{\theta}$:
Then, for the null hypothesis of $$H_0 : \{\Gamma^a_0 (h,\hat{\theta}_T) =0, \; h=1,...,H\},$$ we have
which has a chi-square distribution with degrees of freedom equal to $H(KL)^2 -dim(\theta)$ where K is the number of linear and non-linear transformations, and L is the number of variables.
jasiak2023gcov extend the GCov test in several ways. First, they develop the asymptotic analysis of the GCov test for local alternatives and demonstrate that the test exhibits an asymptotically non-centered chi-square distribution if deviations from the null are local. Second, they propose a bootstrap GCov test that allows the use of estimators other than GCov to estimate the model's parameters.
The goal of this section is to develop the asymptotic properties of the semi-parametric GCov estimator under parametric model misspecification. Consequently, we present model selection tests based on the asymptotic properties of the GCov estimator under misspecification.
This study considers two specification families, denoted as $M_1$ and $M_2$. These two families could be used as model spaces or lag length. We address two key aspects: the relative positions of specification spaces and the position of truth in relation to those. First, we need clarification on the position of the specification spaces relative to each other. These positions can be broadly categorized into three main types. First, $M_1$ and $M_2$ are non-nested, which means we can not achieve any of them from the other space[Figure (ref)]. Second, one of them could be nested within the other. In this case, we refer to them as nested [Figure (ref)], and the last one occurs when there is an overlap among the spaces. Following liao2020nondegenerate, we call them overlapping non-nested [Figure (ref)].
Second, the position of the truth in comparison to the specification families is essential to understand whether we are under the correct specification or misspecification. While the actual truth remains unknown, models and tests often rely on assumptions about the truth's position within the specification spaces. Sometimes, we test two different model spaces when the truth lies outside of both, resulting in a misspecification [Figure (ref)]. However, some tests have been developed to tell us which model spaces are closer to the truth [vuong1989likelihoodand gourieroux1995testing]. In other scenarios, we have overlapping non-nested hypotheses, and the truth is in the overlapping part; then we have an identification issue since we can not identify the truth[Figure (ref)]. Alternatively, when dealing with nested hypotheses (let us say M2 is nested in M1), the truth is inside M2, so M1 is overfitting[Figure (ref)]. However, we assume that M1 and M2 are strictly non-nested, and the truth lies in one of them, without loss of generality, in M1 [Figure (ref)]. This assumption enables us to derive the asymptotic distributions of the test under the true null hypothesis.
Consider the following non-nested model spaces:
$$M_1 : g(Y_t; \theta) = u_t,$$ $$M_2 : h(Y_t; \beta) = v_t,$$
which $g$ and $h$ are known functions satisfying the assumption of the previous section and strictly non-nested. Without loss of generality, let us assume we are under the true null hypothesis (model spaces) of $M_1$. The parameters of interest are $\theta$ and $\beta$. Our time series satisfies all the assumptions of the GCov estimator, including Assumption 3.1.
Assumption 3.1:
-The process $Y_t$ is a strictly stationary sequence and the errors are i.i.d with true distribution of $f_0$ ($M_1$).
- The functions $g$ and $h$ are invertible respect to $Y_t$ and also differentiable.
Assumption 3.2: The distribution of $u_t$ and $v_t$ is identical, however, $v_t$ allows to have dependence structure. Transformed residuals under correct specification and misspecification have finite fourth moments.
Since the GCov estimator is semi-parametric, and we need to choose the transformations based on the characteristics of the errors, we consider the first part of Assumption 3.2, indicating identical distributions of $u_t$ and $v_t$ to facilitate the process of choosing transformations. However, this assumption can be relaxed by using GCov with many transformations as proposed in jasiak2023gcov.
Assumption 3.3: The pseudo-true value of the parameter, $b(\theta_0)$, and the finite sample pseudo-true value of the parameter,$b_T(\theta_0)$, exist, and both of them are unique and on the boundary of a compact set $\Theta$.
Assumption 3.4: The binding function $b(.)$ is one to one and the $\frac{\partial b}{\theta'}[\theta_0,f_0]$ is full-column rank.
Assumption 3.5: The matrices $$\sum_{h=1}^{H} \frac{\partial \mathrm{Tr}[R_a^2(h, \beta)] }{\partial \beta } [b(\theta_0)] \frac{\partial \mathrm{Tr}[R_a^2(h, \beta)] }{\partial \beta' } [b(\theta_0)],$$
and
$$\sum_{h=1}^{H} \frac{\partial^2 \mathrm{Tr}[R_a^2(h, \beta)] }{\partial \beta \partial \beta'}[b(\theta_0)]$$
are positive, semi-definite.
Assumption 3.3 provides the existence and uniqueness of the pseudo-true value of the parameter. Assumption 3.4 comes from the differentiability of the binding function. Assumption 3.5 ensures the well-behavior of variance.
Since we are under $M_1$, it means $g(.)$ is the true model, the true value of the parameter is $\theta_0$, and the estimate of the parameter $\hat{\theta}$ goes to $\theta_0$ asymptotically [gourieroux2022generalized]. However, when we fit the model in $M_2$, the value of the estimated parameter $\hat{\beta}$ is going to the pseudo-true value of the parameter, $b(\theta_0)$, asymptotically. We refer to the finite sample pseudo-true value of the parameter as $b_T(\theta_0)$. The values of the estimated parameters are obtained by the minimization of the GCov objective function based on different models and different parameters as follows:
and
Proposition 3.1: Under assumptions 3.1 to 3.5, the GCov estimator is consistent and has an asymptotically normal distribution around the pseudo-true value of the parameter:
where: $$\Omega^a_{22}(H,b(\theta_0))=J^a_{22}(H,b(\theta_0))^{-1}I^a_{22}(H,b(\theta_0))J^a_{22}(H,b(\theta_0))^{-1},$$
$$ J^{a}_{22} (H, b(\theta_0)) = \sum_{h=1}^{H} \frac{\partial^2 \mathrm{Tr}[R_a^2(h, \beta)] }{\partial \beta \partial \beta'}[b(\theta_0)],$$
$$ I^{a}_{22} (H,b(\theta_0))= \sum_{h=1}^{H} \frac{\partial \mathrm{Tr}[R_a^2(h, \beta)] }{\partial \beta } [b(\theta_0)] \frac{\partial \mathrm{Tr}[R_a^2(h, \beta)] }{\partial \beta' } [b(\theta_0)].$$
Proof: See Appendix A.
Remark 3.1: If we are in a correct specification, Proposition 3.1 is equivalent to the asymptotic properties of the GCov estimator developed in gourieroux2022generalized.
Comparing the asymptotic distribution of the GCov estimator under correct specification and misspecification, we can argue that both are asymptotically Normal; however, we lose the semi-parametric efficiency properties under misspecification. We have the following joint multivariate distribution in Corollary 3.1, based on Proposition 3.1 and following the approach in gourieroux1983testing for the PML estimator.
Corollary 3.1: If model $M_1$ is well specified and model $M_2$ is misspecified, then by Proposition 3.1 and asymptotic Normality of the GCov estimator under correct specification, the vector
\[ \sqrt{T}
, \]
has an asymptotically Normal distribution with mean zero and variance $$\Omega^a(H,\theta_0,b(\theta_0))=J^a (H,\theta_0,b(\theta_0))^{-1} I^a(H,\theta_0,b(\theta_0))J^a (H,\theta_0,b(\theta_0))^{-1},$$
where
\[ J^a (H,\theta_0,b(\theta_0))=
, \]
\[ I^a (H,\theta_0,b(\theta_0))=
, \]
$$ J^{a}_{11} (H, \theta_0) = \sum_{h=1}^{H} \frac{\partial^2 \mathrm{Tr}[R_a^2(h, \theta)] }{\partial \theta \partial \theta'}[\theta_0],$$
$$ I^{a}_{11} (H,\theta_0)= \sum_{h=1}^{H} \frac{\partial \mathrm{Tr}[R_a^2(h, \theta)] }{\partial \theta } [\theta_0] \frac{\partial \mathrm{Tr}[R_a^2(h, \theta)] }{\partial \theta '} [\theta_0],$$
$$ I^{a}_{12} (H,\theta_0,b(\theta_0))= \sum_{h=1}^{H} \frac{\partial \mathrm{Tr}[R_a^2(h, \theta)] }{\partial \theta } [\theta_0] \frac{\partial \mathrm{Tr}[R_a^2(h, \beta)] }{\partial \beta' } [b(\theta_0)] = I^{a}_{21} (H,\theta_0,b(\theta_0))' . $$
Propositions 3.1 and Corollary 3.1 give the asymptotic distribution of $\hat{\beta}_T$ around the pseudo-true value of the parameter, which is usually unknown to the researcher. Under a misspecified model, we have the parameter's pseudo-true value $b(\theta_0)$, the asymptotic pseudo-true value $b(\hat{\theta})$, and the finite sample's pseudo-true value $b_T(\hat{\theta})$.
Proposition 3.2: The GCov estimator has an asymptotically normal distribution around the asymptotic pseudo-true value $b(\hat{\theta}_T)$ and a finite sample pseudo-true value of parameter $b_T(\hat{\theta})$with variances
and
where
Proof: See Appendix A.
Even obtaining the closed form of the finite sample pseudo-true value of parameter $b_T(\hat{\theta})$ may not be feasible. Here, we present a simulation approach proposed by gourieroux1993indirect that can give an asymptotically consistent estimator of the simulated pseudo-true value.
In this approach, we have the following steps:
-We estimate the parameter under correct specification as $\hat{\theta}$ estimated parameter and $\hat{u_t}$ as fitted residuals.
-We resample the residuals to obtain $\hat{u^s_t}$ for $s=1,2,..., S$.
-We generate $y^s_t=g^{-1}(\hat{\theta},\hat{u^s_t})$.
- For each $y^s_t$ we fit misspecified $h(.)$ and estimate the parameter $\hat{\beta}^s$.
-Then we have
This simulation path is computationally time-consuming. gourieroux1993indirect suggested an alternative way that instead of estimating $\beta^s$ for $s=1,\dots,S$ we can simulation time series of dimension $TS$ and estimate $b_{TS}(\theta_0)$ which is equivalent to $b_{T,S}(\theta_0)$.
Remark 3.2: Based on the simulated finite sample pseudo-true value of parameter $b_{T, S}(\hat{\theta})$ and under pure time series model[gourieroux1995testing] we can argue that $\sqrt{T}(\hat{\beta}_T - b_{T, S}(\hat{\theta}) )$ and $\sqrt{T}(\hat{\beta}_T - b_{TS}(\hat{\theta}) )$ are asymptotically equivalent and normally distributed with mean zero and variance-covariance matrix equal to $$\Omega^a_{S}=\left( 1 + \frac{1}{S} \right) {J^a_{22}}^{-1} I^{a*}_{22} {J^a_{22}}^{-1}.$$
We can argue the model is correctly specified if we do not reject the null of i.i.d residuals based on the GCov specification test, and is misspecified if we reject the null hypothesis based on estimated models. This argument is sensitive to the number of lags included in the GCov objective function and also the transformations we consider. Consider $M_1$ as the correct specification and $M_2$ as the misspecified model space. With different transformations and also different numbers of transformations and lags, we expect $M_1$ to always be the correct specification based on Definition 1. However, there is a possibility of false acceptance of $M_2$ as the correct specification based on the non-informative transformations.
The GCov-based specification test has been proposed by gourieroux2022generalized, and jasiak2023gcov investigated the properties of this test specifically under local alternatives. Here, we can extend their work to a broader range of model selection tools, including testing one specification against others where the models may be nested, overlapping, or non-nested.
First, we focus on the Wald-type test method introduced by gourieroux1983testing and white1982maximum. This testing approach depends on the concept of the pseudo-true parameter value, initially developed by sawa1978information and white1982maximum. Subsequently, it has played a significant role in developing encompassing tests, as demonstrated by mizon1986encompassing and gourieroux1995testing. The regularity conditions in this subsection follow gallant1980statistical and burguete1980unification.
We introduce a Wald-type test based on the GCov estimator, which is useful for testing between two separate model spaces. We consider the hypotheses outlined in subsection 3.1. In this context, we use the GCov estimator because it offers several advantages in estimation under the non-Gaussian i.i.d. errors framework. The application of the GCov-based test finds particular utility in specifying the non-causality order of mixed Auto-regressive models.
Corollary 3.2: Based on Assumptions 3.1 to 3.5, and Propositions 3.1 and 3.2, we propose the following GCov-based Wald-type tests:
which all of them have asymptotically $\chi^2$ distribution with $d_{1}$ ,$d_{2}$ and $d_{3}$ as their degrees of freedom which are ranks of $\hat{\Omega}^a_{A}$ , $\hat{\Omega}^a_{F}$ and $\hat{\Omega}^a_{S}$, respectively. The consistency of the proposed tests holds if and only if $b(a(\beta_0)) \neq \beta_0$ [gourieroux1983testing,gourieroux1995testing].
The asymptotic distribution of $\xi_{T}^{W1} $, $\xi_{T}^{W2} $, and $\xi_{T}^{W3} $ is directly consequence of Proposition 3.2 and Corollary 3.1. If we are in the nested case, these statistics are reduced to the traditional Wald test but are now based on GCov.
Next, we want to propose a GCov-based score-type test. Following gourieroux1983testing approach for MLE we define
$$\hat{\lambda}_T^{(1)}= \sum_{h=1}^{H} \frac{\partial \mathrm{Tr}[R_a^2(h, \beta)] }{\partial \beta } [b(\hat{\theta})],$$
and
$$\hat{\lambda}_T^{(2)}= \sum_{h=1}^{H} \frac{\partial \mathrm{Tr}[R_a^2(h, \beta)] }{\partial \beta } [b_T(\hat{\theta})],$$
we want to construct the test that examines their departed from zero. Therefore, we have the following proposition regarding the asymptotic distributions of the score function.
Proposition 3.3: If the correct specification is in $M_1$ we have $$\sqrt{T}\hat{\lambda}_T^{(1)} \sim N(0,I^{a}_{22}- I^a_{21} {I^a_{11}}^{-1} I^a_{12}),$$
and
$$\sqrt{T}\hat{\lambda}_T^{(2)} \sim N(0,I^{a*}_{22}- I^a_{21} {I^a_{11}}^{-1} I^a_{12}).$$
Proof: See Appendix A.
Based on the asymptotic distribution of the GCov-based score functions defined previously, we can now construct an extension to the traditional score test, which is asymptotically equivalent to the extensions of the Wald test in the previous subsection.
Corollary 3.3: Based on Proposition 3.3 we have statistics
and
where both have asymptotically chi-square distribution with degrees of freedom equal to the rank of the variance of score functions.
In this section, we investigate the properties of the GCov estimator under inequality constraints. Some non-linear models in non-Gaussian time series frameworks have conditions that can count as constraints to the optimization problem. For example, the roots of the lag and lead polynomials of causal-noncausal models should satisfy the specific structure. Instead of the conditions of the models, sometimes the researcher is interested not only in the true specification but, in contrast, in the misspecified model that satisfies the specific conditions. An example is the investor who wants a portfolio based on noncausal components of VAR models, as suggested in hall2024modelling, but is opposed to short sales. Therefore, we should constrain the noncausal component to be positive in all elements. Another example is the DAR model.
Example 4.1: Consider DAR(1)
where $\eta_t$ is i.i.d, $\omega>0$, $\alpha\ge0$ and the necessary condition for stationary solution is $E\left(log|\phi+\eta_t\alpha|\right)<0$.
Consider the objective function of GCov estimator constrained by $r$ inequalities $q_r(\theta)\ge0$ :
If the true value of the parameter or the pseudo-true value of the parameter is inside the compact set that satisfies the constraint, then the distribution of the GCov is asymptotically Normal. However, if the true value or pseudo-true value of the parameter is not within the constraint set, the distribution will be the projection of the Normal distribution onto the set of parameters that satisfy the constraints. Here, we follow gourieroux1982likelihood, gourieroux1995statistics, and francq2007quasi approaches when the true value of the correctly specified model or the pseudo-true value of the parameter in the misspecified model is on the boundary of the parameter space based on constraints.
To solve the optimization problem provided in (ref), we can use Kuhn-Tucker multipliers instead of Lagrangian multipliers as suggested in gourieroux1995statistics. We rewrite the inequality-constrained optimization problem as a Hamiltonian function.
$$L_T^a(\theta,H)=\sum_{h=1}^{H} \mathrm{Tr}[R_a^2(h, \theta)]+ \sum_{r=1}^{R} \gamma_r q_r(\theta),$$
where $\gamma_r$ are Kuhn-Tucker multipliers. Consequently, we write the first-order conditions as
$$\frac{\partial L_T^a(\hat{\theta}^C_T,H) }{\partial \theta}=0 \Leftrightarrow \frac{\partial \sum_{h=1}^{H} \mathrm{Tr}[R_a^2(h, \hat{\theta}^C_T)] }{\partial \theta}+ \frac{\partial q'(\hat{\theta}^C_T) }{\partial \theta} \hat{\gamma}_T =0,$$
which gives the Kuhn-Tucker multipliers vector as
$$\hat{\gamma}_T= - \left(\frac{\partial q'(\hat{\theta}^C_T) }{\partial \theta}\right)^{-1} \frac{\partial \sum_{h=1}^{H} \mathrm{Tr}[R_a^2(h, \hat{\theta}^C_T)] }{\partial \theta}. $$
Here, we investigate the properties of CGCov under both correct and misspecified conditions when the true value of the parameter or pseudo-true value lies on the boundary of the parameter space. To get an asymptotic distribution of the CGCov estimator, we need to substitute assumption 3.3 with a stronger version of that, considering the existence of the finite sixth moment of the transformed residuals, and also an assumption to facilitate the cases where the parameter is precisely on the boundary.
Proposition 4.1: Under Assumptions provided in Appendix C, we have
i) $\hat{\theta}^C_T$ and $\hat{\beta}^C_T$ are consistent estimators of $\theta_0$ and $b(\theta_0)$, respectively.
ii)
where
when the true parameter is on the boundary. We have $\Lambda_i=\mathbb{R}$ if $\theta_{0i}$ is not on the boundary and $\Lambda_i$ equal to space satisfying $q_r$ constrain if $\theta_{0i}$ is on the boundary.
iii) The GCov specification test distribution is not asymptotically a chi-square distribution, and we have
$$L_T^a(\hat{\theta}_T^C,H)\to \lambda^{\Lambda'} J^a_{11} \lambda^{\Lambda}.$$
iv)
where
when the pseudo-true value of the parameter is on the boundary. We have $\Lambda_i=\mathbb{R}$ if $b(\theta_{0i})$ is not on the boundary and $\Lambda_i$ equal to space satisfying $q_r$ constrain if $b(\theta_{0i})$ is on the boundary.
v) The GCov specification test distribution is not asymptotically a chi-square distribution under misspecification and the pseudo-true value of the parameter on the boundary of the parameter space, and we have
$$L_T^a(\hat{\beta}_T^C,H)- L_T^a(b(\theta_0),H)\to \lambda^{\Lambda'} J^a_{22} \lambda^{\Lambda}.$$
Proof: See Appendix B.
Remark 4.1: If the constraints only are on the non-negativity of parameters like DAR models, then $\Lambda_i=[0,\infty)$ if $\theta_{0i}$ is on the boundary.
Remark 4.2: If the constraints are a function of more than one parameter, for instance, $q_1(\theta_1,\theta_2)=\theta_1+\theta_2\ge0$ then if the true parameters are on the boundary as $\theta_{01}=0.3$ and $\theta_{02}=-0.3$ then $\Lambda_1\times\Lambda_2$ is consist of all the $(\theta_1,\theta_2)$ satisfying $q_1$.
Remark 4.3: Since by Proposition 4.1, the CGCov does not have asymptotic normal distribution when the (pseudo)true value of the parameter is on the boundary, the GCov specification test proposed in gourieroux2022generalized based on CGCov is not asymptotically chi-square distributed anymore. Instead, we can use the bootstrap GCov test proposed by jasiak2023gcov, which only has the constancy assumption that we have with the CGCov estimator.
Based on Proposition 4.1, we cannot use the model selection test provided in Section 3 when we use CGCov and the (pseudo-)true value of the parameter is on the boundary. This problem arises specifically when there is an over-identified specification in the conditional volatility models, such as ARCH-GARCH models or DAR models. Then, the pseudo-true value of the parameter in the misspecified model will be zero, and it will be located on the boundary of the parameter space, based on the non-negativity assumption for parameters. To address this issue, we examine the asymptotic distribution of the test statistics proposed in Section 3, based on the CGCov estimator, under the condition that the (pseudo-)true value of the parameter is on the boundary.
Here, we focus on the constrained GCov estimator. Specifically, we investigate its use in the context of causal and noncausal models, but it is not limited to these. To estimate the parameters of MAR(r,s) in equation (ref) as $\Theta_T$ with respect to the assumption $|\lambda|<1$ and $|\gamma|<1$. Then we have
where the definition of $R_a^2(h, \Theta)$ provided in (ref).
The proposed constraint is on the roots of polynomials; however, we can transform constraint (23) to impose the new set of constraints on $\Theta$. We use the algorithm proposed by jury1964theory to convert the constraint on the roots to the constraints on the parameter. Consider the lag polynomial of order r:
$$\Phi(L)=1-\phi_1L-\phi_2L^2-...\phi_r L^r,$$
where the roots of this polynomial should be outside of the unit circle. This is equivalent to
$$L^r \Phi(L^{-1})=-\phi_r -\phi_{r-1}L-...-\phi_1L^{r-1}+L^r,$$
where the roots are inside the unit circle. For simplicity of notation, we rewrite it as $$F(z)=a_0+a_1z+...+a_rz^r,$$
where $z=L$, $a_r=1$, and $a_i=-\phi_{r-i}$ for $i=0,1,...,r-1.$ Then we construct matrix $X_{k} $ and $Y_{k}$ as follow
Then we can rewrite the determinants of $|X_k+Y_k|=A_k+B_k$ and $|X_k-Y_k|=A_k-B_k$ where $A_k$ and $B_k$ are stability constants[ jury1964theory]. The constraint that roots of $F(z)$ are inside the unit circle is equivalent to roots of $\Phi(L)$ be outside the unit circle for $r-odd$
$$F(1)>0, F(-1)<0$$ $$(-1)^{k(k+1)/2}(A_k \pm B_k)>0, k=2,4,6,\dots,r-1.$$
For $r-even$ are $$F(1)>0, F(-1)>0$$ $$(-1)^{k(k+1)/2}(A_k -B_k)>0,(-1)^{k(k+1)/2}(A_k +B_k)<0 ,k=1,2,5,\dots,r-1.$$ The same approach can be used for lead polynomials.
Example 4.2: Consider $MAR(3,3)$ as $$(1-\phi_1L-\phi_2L^2-\phi_3 L^3)(1-\psi_1L^{-1}-\psi_2 L^{-2} -\psi_3 L^{-3})y_t=\epsilon_t,$$
where the roots of a lag polynomial are outside, and the lead polynomial is inside the unit circle. The equivalent constraint on the parameters is:
Example 4.3: Consider the case that the researcher is only interested in fitting a pure noncausal mixed-VAR(1) process
where the roots of lagged polynomials are inside the unit circle. The representation of the root conditions on the parameters $\Phi$ is
$$ 1 <(\phi_{11}\phi_{22} -\phi_{12}\phi_{21})\quad \text{and} \quad|\phi_{11}+\phi_{22}|<1+(\phi_{11}\phi_{22} -\phi_{12}\phi_{21}),$$
or
$$ (\phi_{11}\phi_{22} -\phi_{12}\phi_{21})<0 \quad \text{and} \quad |\phi_{11}+\phi_{22}|<-1-(\phi_{11}\phi_{22} -\phi_{12}\phi_{21}).$$
This section investigates the use of the proposed estimator and tests in (non)causal and (non)invertible processes. These models satisfy the assumptions of the GCov estimator both under correct specification and misspecification.
The concept of misspecification has been introduced to these models by gourieroux2018misspecification by considering the misspecified order of lags and leads to the causal-noncausal process. Moreover, they also used PML and a misspecified distribution of the error term in the estimation process. An interesting aspect of MAR models is the achievability of the closed-form of binding functions under misspecification, as explored for special cases in gourieroux2018misspecification. In this chapter, we aim to extend their work to a broader context and relax the assumption of a known distribution by deriving a binding function based on the GCov as a semi-parametric estimator. It is worth noting that we cannot use the binding functions provided by gourieroux2018misspecification in this paper to apply the model selection test, as their binding functions are provided for the PML estimator. The fact that MAR(r,s) for $r>1$ and $s>1$ are non-nested and the existing model's selection criteria for MAR models are biased [gourieroux2018misspecification] gives a clear contribution of the GCov-based model's selection tests.
Remark 5.1: Consider the DGP of $MAR(r,s)$ and the misspecified model of $MAR(r-q,s+q)$ where we call $q$ as order of misspecification. The roots of the lag polynomial of correct specifications are $\lambda_1^{-1},...,\lambda_r^{-1}$ where $|\lambda|<1$ and the roots of a lead polynomial are $\gamma_1,...,\gamma_s$ where $|\gamma|<1$. Consider the q roots from the lag polynomial that will flip to the lead roots as the last q roots. The new set of roots under the misspecified model are $\lambda_1,...,\lambda_{r-q}$ and $\gamma_1,...,\gamma_{s+q}$ where $\gamma_{s+i}=\lambda_{r-i}$ for $i=1,...,q$. Then, the asymptotic binding functions of pseudo-true parameters are for $i=1,...,r-q$
$$b_{\phi_{i}}(\phi_{0,1},...,\phi_{0,r},\psi_{0,1},...,\psi_{0,s})= (-1)^{i+1} \sum_{j_1 =1}^{r-q} \sum_{j_1<j_2}^{r-q}...\sum_{j_{i-1}<j_{i}}^{r-q} \lambda_{j_1} \lambda_{j_2}... \lambda_{j_{i}},$$
and for $i=1,..,s+q$
$$b_{\psi_{i}}(\phi_{0,1},...,\phi_{0,r},\psi_{0,1},...,\psi_{0,s})= (-1)^{i+1} \sum_{j_1 =1}^{s+q} \sum_{j_1<j_2}^{s+q}...\sum_{j_{i-1}<j_{i}}^{s+q} \gamma_{j_1} \gamma_{j_2}... \gamma_{j_{i}}.$$
Consider that q out of r choices are possible for the flipping roots. Therefore, we have $\frac{r!}{(r-q)! q!}$ different possible sets of pseudo-true value of parameters with unconstrained GCov estimator.
Remark 5.1 is an extension of the pure causal representation of MAR(r,s) proposed by hecq2022spectral since the pure causal representation of MAR(r,s) could count as a misspecified model. Moreover, we advance the closed-form binding functions for any order of misspecification, which extends the work of gourieroux2018misspecification.
Example 5.1: To compare the pseudo-true value of the parameters based on GCov and ML estimators, we conduct same simulation as provided in gourieroux2018misspecification Figure 2 by generating a noncausal AR(1) with a Cauchy error distribution and with different autoregressive coefficients from 0.1 to 0.9. The number of observations is T=1000. Then we fit a causal AR(1) as a misspecified model and report the mean of the estimator for 1000 replications in Figure 4. Comparing Figure 4 with Figure 2 of gourieroux2018misspecification shows the advantage of using the GCov estimator in terms of not having discontinuity in the pseudo-true value of the parameter.
Remark 5.2: According to Remark 5.1, the pseudo-true values of the parameters under the misspecified parametric model are not in the interval that satisfies the assumptions of the model if the order of misspecification $q$ is non-zero. We can construct a Wald-type or Score-type test to choose between different causal and noncausal models. However, with a constrained GCov estimator, we have $b(a(\beta_0))\neq \beta_0$, which allows us to use the proposed tests. However, we do not have a close form of binding functions here. Therefore, we can not develop $\xi^{W1}_T$ and $\xi^{S1}_T$ test statistics.
Remark 5.3: consider model M1:MAR(r,s) wit true set of parameters $\alpha_0$ and model M2:MAR(r-q,s+q) with true set of parameter $\beta_0$. Then, if we are under the M1 specification, the binding function is $b(\alpha_0)$, and if we are under the M2 specification, we have $a(\beta_0)$ as binding functions. For any non-zero misspecification order $q$ we have $b(a(\beta_0))=\beta_0$ and $a(b(\alpha_0))=\alpha_0$ using the GCov estimator.
Remark 5.4: consider M1: $MAR(r,s)$ where $r+s=p$ and M2: $MAR(r',s')$ where $r'+s'=p'$ and $p<p'$. If we are under M1 specification, then $b(a(\beta_0))\neq \beta_0$.
Example 5.2: Consider Null hypothesis of M1:MAR(0,1) and the alternative hypothesis is the M2:MAR(0,2) model and we use constrained GCov estimator with $K=2$ and $H=3$. The DGP for empirical size is MAR(0,1) with $\psi=0.3$, and the DGP for empirical power is MAR(0,2) with $\psi_1=0.3$ and $\psi_2=0.6$. Both DGPs have t(4), t(5), and t(6) error distributions, and we run the simulation experiment for $T=100,200,500$ observations. Since we are in the nested test, we can use a Wald-type test with simplified $\Omega_A=J^a_{22}(b(\hat{\theta}))$. From Remark 4.1, we have
$$b_1 (\psi_{0,1})= \psi_{0,1},$$ and $$b_2 (\psi_{0,1})=0.$$
Therefore we can construct the $\hat{\xi}^{W1}_T$ as follow
$$\hat{\xi}^{W1}_T= T
' J^a_{22}
.$$
Since the rank of $J^a_{22}$ is equal to 2 we compare the $\hat{\xi}^{W1}_T$ with $\chi^2_{0.95}(2)$. We reject the null hypothesis if $\hat{\xi}^{W1}_T>\chi^2_{0.95}(2)$. Table 1 indicates the results of this test's empirical size and power.
In this subsection, we propose an alternative algorithm for choosing the correct specification in MAR models. The existing algorithm based on AIC criteria has a significant bias in certain situations, like the error being Cauchy distributed [gourieroux2018misspecification]. Here is the proposed algorithm:
1. Fit causal AR(p) for $p=1,2,3,\dots$, test the i.i.d residuals by GCov specification test, and choose the first p that gives you i.i.d residuals.
2. Fit all possible MAR(r,s) where $r+s=p$ with unconstrained GCov and choose the model that does not violate the roots assumption.
Example 5.3: Consider the MAR(1,1) with Cauchy and t(5) error distribution, $\phi=0.7$ and $\psi=0.2$ Similar to the example provided in hecq2022spectral. The number of observations is $T=500$, and we examine the identification algorithm based on the unconstrained GCov estimator in the second stage. First, we identify $p$, and then choose the causal order $r$ and the noncausal order $s$. The simulation results are based on 1000 replications, and the upper bound of p is five lags.
Table 2 shows the rate of choosing the correct specification, total lags-leads order $p=2$, and also the $MAR(r,s)$ possible specifications. For the DGP of $MAR(1,1) $, with Cauchy errors among those, we choose the correct $p$; with the probability of $0.985$, we choose the correct order of causal and noncausal. However, when we change the error distribution to $t(5)$ and get closer to the Normality, this rate decreases, and the possibility of choosing misspecified causal models increases, which aligns with the theory. The second row of Table 2 provides evidence of possible identification issues and the existence of partial identification in causal-noncausal processes when the error term is close to a normal distribution.
Remark 5.5: We can modify the second step of the model selection algorithm for causal-noncausal models by replacing the unconstrained GCov estimator with the constrained GCov. This way, the results are independent of initial values. Additionally, we can structure the Wald-type test to compare $MAR(2,0)$ and $MAR(1,1)$ based on the constrained GCov in cases where we choose both of them using the proposed algorithm.
In this subsection, we utilize CGCov to estimate agmented DAR(p,q) models presented by jiang2020non and present a model selection approach to select the optimal p and q. Consider the DAR(p,q) model as follows:
where $\eta_t$ is i.i.d, $w>0$ and $\alpha_i\geq0$ for $i=1,\dots,p$ and $y_t$ is strictly stationary. We can rewrite this process in the structure of Model M1 in subsection 3.1 as
$$M_1: g(Y_t;\theta)=u_t,$$
where
$$g(Y_t;\theta)= \frac{y_t-\phi_1 y_{t-1}-\dots-\phi_p y_{t-p}}{ \sqrt{\omega +\alpha_1 y_{t-1}^2+\dots+ \alpha_q y_{t-q}^2}}=\eta_t=u_t,$$
and
$$\theta=[\phi_1,\dots,\phi_p,w,\alpha_1,\dots,\alpha_p]'.$$
Since we have constraints, we need to use the CGCov estimator proposed in section 4. In the literature related to DAR models, it is usually assumed that $p=q$, and these models are referred to as DAR(p).
Here, we propose a new approach to select the order of DAR models based on a bootstrap-based GCov specification test similar to the algorithm we proposed earlier in subsection 4.1 for MAR models. The case of DAR models requires more careful attention, as we have constraints on parameters and must use the CGCov estimator for estimation. Fist Consider following algorithm to choose $max(p,q)$:
1. fit $DAR(i)$ for $i=1,2,3,\dots$, estimate the parameters by CGCov and then test the i.i.d $\hat{\eta}_t$ by GCov-based bootstrap test until you get i.i.d $\hat{\eta}_t$ and consider that lag as $p'=i$. Then, $max(p,q)=p'$.
2.0 If you are fitting $DAR(p)$ with $p=q$, then $p=p'$ and you can select the model. If you consider $DAR(p, q)$ models, follow these steps:
2.1 Fix $p=p'$ and fit $DAR(p',q)$ for $q=1,2,..,p'-1$. Then use the bootstrap-based GCov specification test. If $p>q$, then you can choose the first $q$ for which you do not reject the null hypothesis.
2.2 If $q>p$, then you will reject all the models in 2.1. Fix $q=p'$ and fit $DAR(p,p'$) for $q=1,2,...,p'-1$. Choose the first p-value in which you do not reject the null hypothesis.
2.3 If $p=q$, then you will reject all the models in 2.1 and 2.2. Therefore, your model is $DAR(p')$.
Example 5.6: Consider DAR(2,1)
$$y_t=\phi_1 y_{t-1}+\phi_2 y_{t-2} + \eta_t \sqrt{w+\alpha y_{t-1}^2},$$
where $\phi_1=0.4$, $\phi_2=0.2$, $w=1$, and $\alpha=0.4$. The $\eta_t$ is i.i.d with t(5) distribution, and we consider $T=1000$ observation. We use the DAR(p,q) models selection approach to choose $p$ and $q$. Table (ref) illustrates the estimated parameter properties for the misspecified models DAR(1) and DAR(2), as well as the correct specification DAR(2,1). Moreover, we report the probability of rejecting the model based on both the GCov test and the bootstrap-based GCov test. Finally, we use the proposed model selection algorithm for DAR models and provide the probabilities in the last column of Table (ref).
This subsection examines monthly data from the Producer Price Index by Commodity: Final Demand: Final Energy Demand (PPIDES) from November 2009 to January 2023. We detrend the series by regressing it on a constant and time. We employ a model selection algorithm proposed in Section 5.1.1 to find the total number of lags and leads. For the selection stage, we use both constrained and unconstrained GCov estimators to fit the causal and noncausal processes, aiming to highlight the importance of using constrained GCov and the potential for misspecification under unconstrained GCov.
First, by the NLSD test proposed by jasiak2023gcov, we show the existence of linear and nonlinear dependence in the time series, considering $K=2$ transformations of residuals and residuals square and $H=10$. The value of the test is 1245.1, and the chi-square 0.95 percent critical value is 55.76, which indicates the existence of dependence in the time series. Then, we fit $MAR(p,0)$ until we get i.i.d. residuals based on the GCov test. Table (ref) indicates that the total number of lags and leads equal to two yields i.i.d. residuals. However, this model violates the assumption of roots outside the unit circle once, providing evidence that $MAR(1,1)$ is the correct specification. For the second stage, we fit causal and noncausal processes using both constrained and unconstrained GCov estimators, with exactly the same initial optimization values. Table (ref) provides the results for both.
By comparing the UC and C panels of Table (ref), we can argue that the unconstrained GCov provides pseudo-true values of the parameter, which is equal to the inverse of the true coefficients in this case, and violates the model's assumptions. However, the constrained one directly estimates the true specification. Figure (ref) shows the fitted values of $MAR(1,1)$ and estimated residuals. Moreover, we report the ACF of the series, square series, residuals, and squared residuals in Figure (ref) to support the correct specification of $MAR(1,1)$. Finally, we use a Wald-type test to exclude under-fitting possibilities, which, in this case, is reduced to the simple T-test. Since the $\psi_2$ is insignificant, we conclude that $MAR(1,1)$ is the best model for the PPIDES series.
Here, we consider the US 3-month Treasury bill second market rate monthly data from January 1934 to April 2025, with a total sample of 1096 observations. Figure (ref) shows the series itself and the first difference of the series. We fit the DAR(p,q) model to the first difference of the series with the CGCov estimator and based on the model selection algorithm proposed previously. This application aligns with the work of jiang2020non, but instead of using weekly data for a specific window, we apply the DAR model to all available monthly data. We have K=4, including up to the fourth power of $\hat{\eta_t}$ and $H=10$. Based on the model selection approach, we landed on the DAR(1) model. Table (ref) provides the estimated parameter and also the GCov specification test. Since $\hat{w}$ is close to the boundary of the parameter space, the asymptotic distribution of the GCov specification test is not valid anymore. Therefore, we use the bootstrap CV. If we went with chi-square asymptotic distribution, we would reject the i.i.d $\hat{\eta_t}$; however, with bootstrap critical value, we do not reject the null hypothesis of i.i.d $\hat{\eta_t}$.
This paper investigates the properties of the GCov estimator under misspecification. We propose Wald-type and score-type tests based on the GCov estimator for model selection and provide their asymptotic distribution. Moreover, we develop an indirect GCov estimator and specification test for models that do not satisfy the GCov estimator assumptions. Specifically, we contribute to the literature on (non)causal processes with a broader range of estimation, model selection, and hypothesis testing tools. Finally, we propose a Constrained GCov estimator and develop its asymptotic distribution when the true value or pseudo-true value of the parameter is on the boundary of the parameter space.
This work can be extended in two ways. First, developing encompassing tests based on the GCov estimator that can choose between two misspecified models and recommend the one that is closer to the true specification. Second, we can extend the properties of the GCov to indirect inference for the estimation of noninvertible moving average models.