EconBase
← Back to paper

Estimation for conditional moment models based on martingale difference divergence

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

58,201 characters · 11 sections · 50 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Estimation for conditional moment models based on martingale difference divergence

\baselineskip=0.85 true cm

abstractWe provide a new estimation method for conditional moment models via the martingale difference divergence (MDD). Our MDD-based estimation method is formed in the framework of a continuum of unconditional moment restrictions. Unlike the existing estimation methods in this framework, the MDD-based estimation method adopts a non-integrable weighting function, which could grab more information from unconditional moment restrictions than the integrable weighting function to enhance the estimation efficiency. Due to the nature of shift-invariance in MDD, our MDD-based estimation method can not identify the intercept parameters. To overcome this identification issue, we further provide a two-step estimation procedure for the model with intercept parameters. Under regularity conditions, we establish the asymptotics of the proposed estimators, which are not only easy-to-implement with analytic asymptotic variances, but also applicable to time series data with an unspecified form of conditional heteroskedasticity. Finally, we illustrate the usefulness of the proposed estimators by simulations and two real examples.

{\it Keywords and phrases}: Conditional moment models; Martingale difference divergence; Time series model estimation.

{\it MOS subject classification: 62M10, 62M20}

Introduction

Conditional moment models play a pivotal role in various applications within economics and statistics. Typically, these models assume an unknown parameter $\theta_0$, which is subject to the following conditional moment restrictions C:1987:

equation[equation omitted — 86 chars of source]

where $\theta_0 \in \Theta \subset \mathbb R^d$, $Z_t$ is a $k$-dimensional random vector that may contain both endogenous and exogenous variables, $X_t$ is a $q$-dimensional random vector that contains the conditioning variables, and $h$ is a given smooth function mapping $\mathbb R^k\times\Theta$ into $\mathbb R^l$. To estimate $\theta_0$ in ((ref)), the conventional approach is to employ unconditional moment restrictions, such as the generalized method of moments (GMM) (Hansen:1982). To be specific, by transforming conditional moment restrictions in ((ref)) to unconditional ones, we can construct estimating moment functions $E(f(X_t)h(Z_t,\theta_0))=0$ for a $r\times l$ matrix of instrumental variables (IVs) $f(X_t)$. To attain the semiparametric efficiency bound, Newey:1990,Newey:1993 choose the optimal IVs by using the nearest neighbors or series expansions method. See the similar non-parametric implementation in DIN:2003 and HK:2011, and its further development on variational method of moments in BK:2003 for univariate $X_t$. Notwithstanding the efficiency of the aforementioned estimators, they have three drawbacks that possibly hinder their practical application. First, their asymptotic theory is only established for the independent and identically distributed (i.i.d.) sample $\{(X_t, Z_t)\}_{t=1}^{n}$. Second, their implementation requires the use of tuning parameters or a well-behaved preliminary estimator, or sometimes both. Third, despite the use of optimal IVs, they may fail to identify $\theta_0$ when $h(Z_t,\theta_0)$ is a nonlinear function.

To address the above drawbacks, DL:2004 propose an alternative method for estimating $\theta_0$ by considering a continuum of unconditional moment restrictions. These restrictions essentially correspond to an infinite number of IVs that can span a space of functions of $X_t$ (Bierens:1982,Bierens:1990). Specifically, DL:2004 consider the restrictions of

equation[equation omitted — 105 chars of source]

where $\mbox{I}(\cdot)$ is the indicator function, and $\mbox{I}(X_t\leq s)$ is an IV for each $s$. To account for the effect of all IVs in ((ref)), they estimate $\theta_0$ by minimizing the sample counterpart of

equation[equation omitted — 118 chars of source]

where $\|\cdot\|$ is the Frobenius norm, and the weighting function $w(s)$ is the density function of $X_t$. See also Escanciano:2006 for similar ideas in the i.i.d. setting. It is worth noting that the idea of DL:2004 has a linkage to that of CF:2000, which uses the exponential function $\exp(sX_t)$ as the IV for univariate $X_t$ and $s$ in a neighborhood of $0$. When $\{(X_t, Z_t)\}_{t=1}^{n}$ is an i.i.d. sample, the parameter estimation using the exponential function to form an infinite number of IVs is further investigated by LP:2013 for model ((ref)), Escanciano:2018 for linear models, and AS:2022 for partially linear models. When $\{(X_t, Z_t)\}_{t=1}^{n}$ is temporally dependent, an extension study of Escanciano:2018 can be found in CEG:2022 for linear models. Regardless whether indicator or exponential function is adopted to generate IVs, these estimation methods relying on a continuum of unconditional moment restrictions necessitate the condition that $w(s)$ is integrable. It is important to note that due to the integrability of $w(s)$, its value likely decays rapidly to zero as $|s|\to\infty$, especially when the dimension of $s$ is large (see Wang:2024). It turns out that the influence of IVs associated with larger values of $s$ is substantially diminished through the use of $w(s)$, bringing an adverse impact on the efficiency of model estimation.

Within the framework of a continuum of unconditional moment restrictions, our paper is motivated by the martingale difference divergence (MDD) introduced by SZ:2014. We propose to estimate $\theta_0$ using a non-integrable weighting function as:

equation[equation omitted — 110 chars of source]

where $c_q=\pi^{(1+q)/2}/\Gamma((1+q)/2)$ with $\Gamma(\cdot)$ being the gamma function. The proposed MDD-based estimator relies on the fact that under ((ref)),

equation[equation omitted — 104 chars of source]

From ((ref)), we consider a continuum of unconditional moment restrictions

equation[equation omitted — 181 chars of source]

where $\mathrm{i}$ is the imaginary unit, $\langle a, b\rangle$ is the inner product for $a, b\in \mathbb{R}^{q}$, and $\exp(\mathrm{i} \langle s, X_t\rangle)$ is an IV for each $s$. In view of ((ref)), the MDD-based estimator aims to minimize the sample counterpart of

equation[equation omitted — 195 chars of source]

where $w_*(s)$ is defined in ((ref)). Here, we denote $|x|_{l}=\sqrt{x^{\star}x}$ for a $l$-dimensional complex vector $x$, with $x^{\star}$ being the conjugate transpose of $x$. Notably, our MDD-based estimator essentially relies on the conditional moment restrictions in ((ref)) rather than those in ((ref)). Hence, unlike the aforementioned estimators, it has a different identification condition, which prevents identifying intercept parameters (if exist) for most of models. To deal with this identification issue on intercept parameters, we further propose a two-step estimation procedure for the model with intercept parameters. To be specific, we first estimate all of non-intercept parameters by the MDD-based estimation method at step one. Based on the MDD-based estimator from the step one, we then estimate all of intercept parameters using a straightforward moment estimation approach at step two.

Under regularity conditions that allow for time series sample $\{(X_t, Z_t)\}_{t=1}^{n}$, we provide the consistency and asymptotic normality of the proposed estimator. More importantly, we manage to derive its asymptotic variance in analytical form, facilitating the statistical inference procedure. To demonstrate the effectiveness of our proposed estimators, we conduct extensive simulation studies. The results illustrate a significant advantage in efficiency compared to the estimator proposed in DL:2004. This highlights the importance of the non-integrable weighting function $w_*(s)$ to construct parameter estimators for conditional moment time series models. Finally, we demonstrate the superiority of our proposed estimators over the estimator in DL:2004 by two real examples.

For univariate linear model with i.i.d. data sample, Tsyawo:2023 also applies $w_*(s)$ to construct an integrated conditional moment (ICM) estimator by minimizing the sample counterpart of

equation[equation omitted — 174 chars of source]

rooting in a continuum of unconditional moment restrictions $E\big\{h(Z_t,\theta_0)\exp(\mathrm{i} \langle s, X_t\rangle)\big\} =0$ for $s\in\mathbb{R}^{q}$. In view of ((ref))--((ref)), the MDD-based estimator is different from the ICM estimator by subtracting $E(h(Z_t,\theta_0))$ from $h(Z_t,\theta_0)$ to create unconditional moment restrictions. Due to the additional substraction, the sample counterpart of $\mbox{MDD}(\theta_0)$ in ((ref)) has an attractive integral form. This is key to deriving our asymptotics for time series data sample, as it enables the use of proof technique for the weak convergence in the functional space. However, such a crucial integral form does not exist for the sample counterpart of $\mbox{ICM}(\theta_0)$ in ((ref)), and Tsyawo:2023 proves the asymptotics of the ICM estimator by using a different proof technique that possibly only works for univariate linear model with i.i.d. data sample. Although the substraction makes the MDD-based estimator applicable for general time series models, it leads to an aforementioned identification issue on intercept parameters as a trade off. Owing to our well-shaped two-step estimation procedure, this identification issue is practically irrelevant in most of cases. Hence, our MDD-based estimation method could have a much larger application scope than the ICM estimation method in Tsyawo:2023 to deal with the multiple linear/nonlinear time series models, which are allowed to have an unspecified form of conditional heteroskedasticity.

The remaining paper proceeds as follows. Section (ref) introduces the MDD-based estimator and establishes its asymptotics. Section (ref) provides the two-step estimation procedure and its related asymptotics. Simulation results are reported in Section (ref). Two real examples are given in Section (ref). Concluding remarks are offered in Section (ref). Technical proofs are presented in the Appendix.

The MDD-based Estimator and its Asymptotics

The MDD-based Estimator

For any two vectors $V\in\mathbb{R}^{l}$ and $U\in\mathbb{R}^{q}$, the MDD of $V$ given $U$ (denoted by $\mbox{MDD}(V|U)$) in SZ:2014 is defined as follows:

flalign*\operatorname{MDD}(V|U)^{2}&\triangleq\int_{\mathbb{R}^{q}} \Big|g_{V, U}(s)-g_{V} g_{U}(s)\Big|_{l}^{2}w_*(s) ds \\ &=\int_{\mathbb{R}^{q}} \Big|E\big\{[V-E(V)]\exp(\mathrm{i} \langle s, U\rangle)\big\}\Big|_{l}^{2}w_*(s) ds,

where $w_*(s)$ is defined in ((ref)), $g_{V, U}(s)=E\big(V \exp(\textrm{i} \langle s, U\rangle)\big)$, $g_{V}={E}(V)$, and $g_{U}(s)=E\big(\exp(\textrm{i}\langle s, U\rangle)\big)$. SZ:2014 show that $\operatorname{MDD}(V|U)^{2}$ has an equivalent expression as

flalign\operatorname{MDD}(V|U)^{2}=-E\Big\{\big[V-E(V))^{\top}(V'-E(V')\big]\|U-U'\|\Big\},

where $(V', U')$ is an i.i.d. copy of $(V, U)$, and $A^{\top}$ is the transpose of a matrix $A$. Hence, $\mbox{MDD}(\theta_0)$ in ((ref)) can be re-written as

flalign*\begin{split} MDD(\theta_0)&=\operatorname{MDD}(h(Z_t,\theta_0)|X_t)^{2}\\ &=-E\Big\{\big[\big(h(Z_t,\theta_0)-E(h(Z_t,\theta_0))\big)^{\top}\big(h(Z_{t'},\theta_0)-E(h(Z_{t'},\theta_0))\big)\big]\|X_t-X_{t'}\|\Big\}, \end{split}

where $(Z_{t'}, X_{t'})$ is an i.i.d. copy of $(Z_t, X_t)$. According to the preceding formula, $\mbox{MDD}(\theta_0)$ has the following sample counterpart:

flalign\begin{split} MDD_n(\theta)&=-\frac{1}{n^2} \sum_{t=1}^{n}\sum_{t'=1}^{n} \Big\{\big[\big(h(Z_t,\theta)-\bar{h}(\theta)\big)^{\top}\big(h(Z_{t'},\theta)-\bar{h}(\theta)\big)\big]\|X_t-X_{t'}\|\Big\}, \end{split}

where $\bar{h}(\theta)=n^{-1}\sum_{t=1}^{n}h(Z_t,\theta)$. Then, our MDD-based estimator of $\theta_0$ (denoted by $\widehat{\theta}_n$) is defined as the minimizer of $\mbox{MDD}_n(\theta)$ for $\theta\in\Theta$, that is,

equation*[equation* omitted — 111 chars of source]

where $\Theta\subset \mathbb{R}^{d}$ is a parameter space.

Asymptotic Properties

Let $X_t$ and $Z_t$ in ((ref)) be two subvectors of $Y_t$. We make the following assumptions to derive the asymptotics of $\widehat{\theta}_n$.

assum$E\big[h(Z_t,\theta)-E(h(Z_t,\theta))|X_t\big]=0$ (a.s.) if and only if $\theta=\theta_0$.
assum$Y_t$ is strictly stationary and ergodic.
assum(i) $h(z,\theta)$ is twice continuously differentiable for each $(z,\theta) \in \mathbb{R}^{k}\times \Theta$; (ii) there exists a function $g(z)$ such that $\|\partial h(z,\theta)/\partial\theta\|\leq g(z)$ and $E\|g(Z_t)\|^4<\infty$; (iii) $E\|X_t\|^4 < \infty$ and $E\prod_{i=1}^q \|X_{it}\|^{2u_0} < \infty$ for some $u_0>1$, where $X_{it}$ is the $i$th entry of $X_t$.
assum$\Theta$ is compact and $\theta_0$ is an interior point of $\Theta$.
assum$h(Z_t,\theta_0)$ is a martingale difference sequence with respect to the filtration $\mathcal{F}_t=\sigma(Y_s,s\leq t)$.

Our assumptions above are similar to those in DL:2004, except that we need a different identification condition for $\theta_0$ in Assumption (ref). For the estimator in DL:2004 or other existing conditional moment estimators, the identification condition for $\theta_0$ is

flalignE(h(Z_t,\theta)|X_t)=0 (a.s.) if and only if \theta=\theta_0.

To illustrate the distinction between the identification conditions in Assumption (ref) and ((ref)), we consider a simple univariate linear model:

flalignZ_{1t}=\alpha_0+\beta_0 Z_{2t}+\varepsilon_t,

where $Z_t=(Z_{1t},Z_{2t})^{\top}$, $\theta_0=(\alpha_0,\beta_0)^{\top}$, and $E(\varepsilon_t|X_t)=0$. Model ((ref)) is corresponding to model ((ref)) with $h(Z_t,\theta_0)=Z_{1t}-\alpha_0-\beta_0 Z_{2t}$. Let $\theta=(\alpha, \beta)^{\top}$ and $h(Z_t,\theta)=Z_{1t}-\alpha-\beta Z_{2t}$. For this model, it is straightforward to see

flalign*E\big[h(Z_t,\theta)-E(h(Z_t,\theta))|X_t\big]&=(\beta_0-\beta)E[Z_{2t}-E(Z_{2t})|X_t],\\ E(h(Z_t,\theta)|X_t)&=(\alpha_0-\alpha)+(\beta_0-\beta)E(Z_{2t}|X_t).

Therefore, under some mild conditions on $X_t$ and $Z_{2t}$, the MDD-based estimator $\widehat{\theta}_n$ can only identify $\beta_0$, while the estimator in DL:2004 can identify both $\alpha_0$ and $\beta_0$. This simple example reveals that the identification condition in Assumption (ref) only allows us to deal with a model without intercept parameter. For the model with intercept parameter, a two-step estimation procedure can be given to estimate both non-intercept and intercept parameters (see more discussions for model ((ref)) below).

Define

flalignu(x)=E\Big\{\Big[\frac{\partial h(Z_t,\theta_0)}{\partial \theta}-E\Big(\frac{\partial h(Z_t,\theta_0)}{\partial \theta}\Big)\Big]\|X_t-x\|\Big\}.

Below, we show the consistency and asymptotic normality of $\widehat{\theta}_n$.

thmSuppose Assumptions (ref)--(ref) hold. Then, $\widehat{\theta}_n-\theta_0\overset{p}{\longrightarrow}0$ as $n\to\infty$.
thmSuppose Assumptions (ref)--(ref) hold. Then, \begin{equation*} \sqrt{n}(\widehat{\theta}_n-\theta_0)\overset{d}{\longrightarrow} N(0,\Omega^{-1}\Sigma\Omega^{-1}) \,\,\, as n\to\infty, \end{equation*} where \begin{flalign*} \Omega&=E\Big\{\Big[\frac{\partial h(Z_t,\theta_0)}{\partial \theta}-E\Big(\frac{\partial h(Z_t,\theta_0)}{\partial \theta}\Big)\Big]^\top u(X_t)\Big\},\\ \Sigma&=E\big\{[u(X_t)-E(u(X_t))]^\top h(Z_t,\theta_0)h(Z_t,\theta_0)^\top [u(X_t)-E(u(X_t))]\big\}. \end{flalign*}

To establish the asymptotic properties of $\widehat{\theta}_n$, our objective is to transform $\mbox{MDD}_n(\theta)$ into an integral form. For this purpose, we need an identity

flalign\|x\|=\int_{\mathbb{R}^{q}}\big[1-\cos(\langle s,x\rangle)\big]w_{*}(s)ds \,\, for all x\in \mathbb{R}^{q};

see SRB:2007. Using the identity in ((ref)) and the fact that

flalign\sum_{t=1}^{n}\sum_{t'=1}^{n} \Big\{\big[\big(h(Z_t,\theta)-\bar{h}(\theta)\big)^{\top}\big(h(Z_{t'},\theta)-\bar{h}(\theta)\big)\big]\Big\}=0,

we can show

flalign*\begin{split} MDD_n(\theta)&=\frac{1}{n^2} \sum_{t=1}^{n}\sum_{t'=1}^{n} \Big\{\big[\big(h(Z_t,\theta)-\bar{h}(\theta)\big)^{\top}\big(h(Z_{t'},\theta)-\bar{h}(\theta)\big)\big]\\ &\quad\quad\quad\quad \times \int_{\mathbb{R}^{q}}\cos(\langle s,X_{t'}-X_{t}\rangle)w_{*}(s)ds\Big\}\\ &=\frac{1}{n^2} \sum_{t=1}^{n}\sum_{t'=1}^{n} \Big\{\big[\big(h(Z_t,\theta)-\bar{h}(\theta)\big)^{\top}\big(h(Z_{t'},\theta)-\bar{h}(\theta)\big)\big]\\ &\quad\quad\quad\quad \times \int_{\mathbb{R}^{q}}\big[\cos(\langle s,X_{t'}-X_{t}\rangle)+i \sin(\langle s,X_{t'}-X_{t}\rangle)\big]w_{*}(s)ds\Big\}, \end{split}

where the last result holds because $\sin(\cdot)$ is an odd function. By observing that $\cos(\langle s,X_{t'}-X_{t}\rangle)+\textrm{i} \sin(\langle s,X_{t'}-X_{t}\rangle)=\exp(\textrm{i}\langle s,X_{t'}\rangle) \exp(-\textrm{i}\langle s,X_{t}\rangle)$, it entails

flalignMDD_n(\theta)&= \int_{\mathbb{R}^{q}} \Big|\mathcal{G}_n(s,\theta)\Big|_{l}^{2}w_*(s) ds =\int_{\mathbb{R}^{q}} \mathcal{G}_n(s,\theta)^{\star} \mathcal{G}_n(s,\theta)w_*(s) ds,

where

equation[equation omitted — 162 chars of source]

The integral form of $\mbox{MDD}_n(\theta)$ in ((ref)) enables us to investigate the asymptotics of $\widehat{\theta}_n$ by studying those of the process $\mathcal{G}_n(s,\theta)$. This feature makes our proof techniques applicable to time series samples, so it largely broadens the application scope of $\widehat{\theta}_n$ for studying the multiple linear/nonlinear time series models.

It is worth noting that our proof techniques can not apply to the ICM estimator in Tsyawo:2023. This is because the sample counterpart of $\mbox{ICM}(\theta_0)$ in ((ref)) is

flalign*ICM_n(\theta)&=-\frac{1}{n^2} \sum_{t=1}^{n}\sum_{t'=1}^{n} \Big\{\big[h(Z_t,\theta)^{\top}h(Z_{t'},\theta)\big]\|X_t-X_{t'}\|\Big\},

based on a similar argument as for ((ref)). However, it is not possible to transform $\mbox{ICM}_n(\theta)$ into a similar integral form as $\mbox{MDD}_n(\theta)$, since the result ((ref)) does not hold in the absence of $\bar{h}(\theta)$.

To see the linkage between our asymptotic normality result and that in Tsyawo:2023, we consider the following multiple linear model (without a constant vector):

flalignZ_{1t}=\Gamma_0 Z_{2t}+\varepsilon_t,

where $Z_{1t}\in \mathbb{R}^{l}$, $Z_{2t}\in\mathbb{R}^{(k-l)}$, $\Gamma_0\in\mathbb{R}^{l\times (k-l)}$, and $E(\varepsilon_t|X_t)=0$. Then, model ((ref)) is a special case of model ((ref)) with $Z_t=(Z_t^{\top},Z_{2t}^{\top})^{\top}$, $\theta_0=vec(\Gamma_0)$, and $h(Z_t,\theta_0)=Z_{1t}-\Gamma_0 Z_{2t}$. In this case, it is not hard to see that $\widehat{\theta}_n=vec(\widehat{\Gamma}_{n})$ has a closed-form solution: $$\widehat{\theta}_n=\Xi_{1n}^{-1}\Xi_{2n},$$ where $$\Xi_{1n}=\sum_{t=1}^n\sum_{t'=1}^n\Big[\mbox{I}_{l}\otimes Z_{2t}^{\top}-\frac{1}{n}\sum_{t=1}^n\big(\mbox{I}_{l}\otimes Z_{2t}^{\top}\big)\Big]^\top\Big[\mbox{I}_{l}\otimes Z_{2t'}^{\top}-\frac{1}{n}\sum_{t=1}^n\big(\mbox{I}_{l}\otimes Z_{2t}^{\top}\big)\Big]\left\|Z_{2t}-Z_{2t'}\right\|$$ and $$\Xi_{2n}=\sum_{t=1}^n\sum_{t'=1}^n\Big[\mbox{I}_{l}\otimes Z_{2t}^{\top}-\frac{1}{n}\sum_{t=1}^n\big(\mbox{I}_{l}\otimes Z_{2t}^{\top}\big)\Big]^\top \Big(Z_{1t'}-\frac{1}{n}\sum_{t=1}^n Z_{1t}\Big)\left\|Z_{2t}-Z_{2t'}\right\|. $$ In addition, for model ((ref)), we have $\partial h(Z_t,\theta_0)/\partial \theta=-\mbox{I}_{l}\otimes Z_{2t}^{\top}$, leading to

flalign*\Omega&=E\big\{\big[I_{l}\otimes Z_{2t}^{\top}-E\big(I_{l}\otimes Z_{2t}^{\top}\big)\big]^\top u(X_t)\big\},\\ \Sigma&=E\big\{[u(X_t)-E(u(X_t))]^\top \varepsilon_t\varepsilon_t^\top [u(X_t)-E(u(X_t))]\big\},

where $u(x)=E\big\{\big[\mbox{I}_{l}\otimes Z_{2t}^{\top}-E\big(\mbox{I}_{l}\otimes Z_{2t}^{\top}\big)\big]\|X_t-x\|\big\}$. In other words, if $E(u(X_t))=0$, $E(Z_{2t})=0$, and $l=1$, our asymptotic variance $\Omega^{-1}\Sigma\Omega^{-1}$ becomes that of Tsyawo:2023.

The Two-step Estimation Procedure

In this section, we provide a two-step estimation procedure for model ((ref)) with intercept parameters. Specifically, we consider a special model ((ref)) with

flalignh(Z_t,\theta_0)=\Big( \begin{array}{c} \theta_{10}\\ 0 \end{array}\Big) +m(Z_t,\theta_{20}),

where $\theta_0=(\theta_{10}^{\top},\theta_{20}^{\top})^{\top}$, $\theta_{10}$ is a $d_1$-dimensional vector of intercept parameters with $d_1\leq l$, $\theta_{20}$ is a $d_2$-dimensional vector of non-intercept parameters, and $m$ is a given function mapping $\mathbb R^k\times\Theta$ into $\mathbb R^l$. Clearly, model ((ref)) nests model ((ref)) with $d_1=l=1$, $d_2=1$, $\theta_{10}=-\alpha_0$, $\theta_{20}=\beta_0$, and $m(Z_t,\theta_{20})=Z_{1t}-\theta_{20}Z_{2t}$. Let $\theta=(\theta_{1}^{\top},\theta_{2}^{\top})^{\top}\in \Theta_1\times \Theta_2$, where $\Theta_i\subset \mathbb{R}^{d_i}$ for $i=1$ and 2. As discussed before, the MDD-based estimation method can not identify $\theta_{10}$, so it has to exclude the estimation of $\theta_{10}$ by simply replacing $h(Z_t,\theta)$ with $m(Z_t,\theta_{2})$. Since the value of $\mbox{MDD}_n(\theta)$ in ((ref)) is invariant to any shift on $h(Z_t,\theta)$, this exclusion implementation does not impact the MDD-based estimator of $\theta_{20}$ given by

flalign*\widehat{\theta}_{2n}=\mathop{\arg\min}\limits_{\theta_2 \in \Theta_2}\operatorname{MDD}_{2,n}(\theta_2),

where $\operatorname{MDD}_{2,n}(\theta_2)$ is defined in the same way as $\operatorname{MDD}_{n}(\theta)$, with $h(Z_t,\theta)$ replaced by $m(Z_t,\theta_2)$. After obtaining $\widehat{\theta}_{2n}$, we partition the functions $h$ and $m$ as follows:

flalignh(Z_t,\theta_0)= \Big( \begin{array}{c} h_1(Z_t,\theta_0)\\ h_2(Z_t,\theta_0) \end{array}\Big)\,\,\, and \,\,\, m(Z_t,\theta_{20})= \Big( \begin{array}{c} m_1(Z_t,\theta_{20})\\ m_2(Z_t,\theta_{20}) \end{array}\Big),

with $h_1\in \mathcal{R}^{d_1}$, $h_2\in \mathcal{R}^{l-d_1}$, $m_1\in \mathcal{R}^{d_1}$, and $m_2\in \mathcal{R}^{l-d_1}$. Then, we use the estimating moment functions $E(\theta_{10}+m_1(Z_t,\theta_{20}))=0$ to estimate $\theta_{10}$ by

flalign*\widehat{\theta}_{1n}=-\frac{1}{n}\sum_{t=1}^{n} m_1(Z_t,\widehat{\theta}_{2n}).

With a slight abuse of notation, we denote $\widehat{\theta}_n=(\widehat{\theta}_{1n}^{\top}, \widehat{\theta}_{2n}^{\top})^{\top}$ for model ((ref)). To present the asymptotic results of $\widehat{\theta}_{1n}$ and $\widehat{\theta}_{2n}$, we need the following notation:

flalign\begin{split} u_2(x)&=E\Big\{\Big[\frac{\partial m(Z_t,\theta_{20})}{\partial \theta_2}-E\Big(\frac{\partial m(Z_t,\theta_{20})}{\partial \theta_2}\Big)\Big]\|X_t-x\|\Big\},\\ \Omega_1&=E\big\{h(Z_t,\theta_0)h(Z_t,\theta_0)^\top [u_2(X_t)-E(u_2(X_t))]\big\},\\ \Sigma_1&=E\big\{h(Z_t,\theta_0)h(Z_t,\theta_0)^\top\big\},\\ \Omega_2&=E\Big\{\Big[\frac{\partial m(Z_t,\theta_{20})}{\partial \theta_2}-E\Big(\frac{\partial m(Z_t,\theta_{20})}{\partial \theta_2}\Big)\Big]^\top u_2(X_t)\Big\},\\ \Sigma_2&=E\big\{[u_2(X_t)-E(u_2(X_t))]^\top h(Z_t,\theta_0)h(Z_t,\theta_0)^\top [u_2(X_t)-E(u_2(X_t))]\big\},\\ J_t&= \begin{pmatrix} E\Big(\frac{\partial m_1(Z_t,\theta_{20})}{\partial\theta_2}\Big) \Omega_2^{-1} [u_2(X_t)-E(u_2(X_t))]^{\top}-\Upsilon\\ -\Omega_2^{-1} [u_2(X_t)-E(u_2(X_t))]^{\top} \end{pmatrix} \end{split}

with $\Upsilon=(\mbox{I}_{d_1},0)\in\mathcal{R}^{d_1\times l}$. Here, $\mbox{I}_{d_1}$ is the identity matrix of size $d_1$. Under an additional assumption for the identifiability of $\theta_{20}$, we are able to establish the joint asymptotic normality of $\widehat{\theta}_{1n}$ and $\widehat{\theta}_{2n}$.

assum$E\big[m(Z_t,\theta_2)-E(m(Z_t,\theta_2))|X_t\big]=0$ (a.s.) if and only if $\theta_2=\theta_{20}$.
thmSuppose Assumptions (ref)--(ref) and (ref) hold. Then, under model ((ref)), \begin{equation*} \sqrt{n}(\widehat{\theta}_n-\theta_0)\overset{d}{\longrightarrow} N(0,V) \,\,\, as n\to\infty, \end{equation*} where $V=E\big[J_t h(Z_t,\theta_0)h(Z_t,\theta_0)^\top J_t^{\top}\big]$. Particularly, \begin{equation*} \sqrt{n}(\widehat{\theta}_{1n}-\theta_{10})\overset{d}{\longrightarrow} N(0,V_1) and \sqrt{n}(\widehat{\theta}_{2n}-\theta_{20})\overset{d}{\longrightarrow} N(0,V_2)\,\,\, as n\to\infty, \end{equation*} where \begin{flalign*} V_1&=E\Big(\frac{\partial m_1(Z_t,\theta_{20})}{\partial\theta_2}\Big) \Omega_2^{-1}\Sigma_2\Omega_2^{-1}E\Big(\frac{\partial m_1(Z_t,\theta_{20})}{\partial\theta_2}\Big)^{\top}-\Upsilon\Omega_1\Omega_2^{-1}E\Big(\frac{\partial m_1(Z_t,\theta_{20})}{\partial\theta_2}\Big)^{\top}\\ &\quad-E\Big(\frac{\partial m_1(Z_t,\theta_{20})}{\partial\theta_2}\Big)\Omega_2^{-1}\Omega_1^{\top}\Upsilon^{\top} +\Upsilon\Sigma_1\Upsilon^{\top},\\ V_2&=\Omega_2^{-1}\Sigma_2\Omega_2^{-1}. \end{flalign*}

As the asymptotic variances in Theorems (ref) and (ref) have analytic expressions, they can be directly estimated by their sample counterparts. This leads to an easy-to-implement statistical inference for $\theta_0$, without the selection of any tuning parameters and preliminary estimator. Meanwhile, we should highlight that our technical assumptions for both theorems allow the data to have the conditional heteroskedasticity of unknown form, so it makes our estimation method have a large application scope to handle the multiple linear/nonlinear time series models.

Finally, we give some remarks on the estimation efficiency. First, although it is hard to make a formal efficiency comparison between $\widehat{\theta}_n$ and the estimator in DL:2004, the simulation studies in the next section demonstrate that $\widehat{\theta}_n$ is significantly more efficient than the estimator in DL:2004. Second, $\widehat{\theta}_n$ is not granted to be efficient. Following DL:2004, we can take $\widehat{\theta}_n$ as the preliminary estimator to obtain an efficient estimator $$\widetilde{\theta}_n=\widehat{\theta}_n-\Big[\frac{\partial^2 L_n(\widehat{\theta}_n)}{\partial\theta\partial\theta^{\top}}\Big]^{-1}\frac{\partial L_n(\widehat{\theta}_n)}{\partial \theta},$$ where $L_n(\theta)$ is the efficient objective function. However, the derivation of $L_n(\theta)$ usually relies on some nonparametric methods to obtain optimal IVs and so the selection of tuning parameters is unavoidable; see Newey:1993. Since our goal is to provide a user-friendly statistical inference for $\theta_0$, the study of $\widetilde{\theta}_n$ will not be given in this paper and is left for future research.

Simulations

In this section, we carry out simulation experiments to assess the finite-sample performance of our proposed MDD-based estimator $\widehat{\theta}_n$. First, we generate $1000$ Monte-Carlo repetitions of sample size $n=50$, $100$, and $200$ from the following twelve univariate data generating processes (DGPs):

itemize[label=] • DGP 1: $Z_{1t}=\theta_0Z_{2t}+\varepsilon_t$, where $\theta_0=1$, $Z_{2t}$ satisfies the order 1 autoregressive (AR(1)) model: $Z_{2t}=0.3Z_{2,t-1}+\eta_t$, and $(\varepsilon_t,\eta_t)^{\top}\overset{i.i.d.}{\sim} N(0,\mbox{I}_2)$; • DGP 2: $Z_{1t}=\theta_0Z_{2t}+\varepsilon_t$, where $\theta_0=1$, $Z_{2t}$ satisfies the AR(1) model: $Z_{2t}=0.3Z_{2,t-1}+\zeta_t$, $\varepsilon_t$ satisfies the order 1 autoregressive conditional heteroskedasticity (ARCH(1)) model: $\varepsilon_t=v_t^{1/2}\eta_t$ and $v_t=0.4+0.5\varepsilon_{t-1}^2$, and $(\zeta_t,\eta_t)^{\top}\overset{i.i.d.}{\sim} N(0,\mbox{I}_2)$; • DGP 3: $Z_{1t}=\sin(\theta_0Z_{2t})+\varepsilon_t$, where $\theta_0=1$, $\varepsilon_t\overset{i.i.d.}{\sim} N(0,1)$, $Z_{2t}\overset{i.i.d.}{\sim} U[-1,1]$, and $\varepsilon_t$ and $Z_{2t}$ are independent; • DGP 4: $Z_{1t}=\sin(\theta_0Z_{2t})+\varepsilon_t$, where $\theta_0=1$, $\varepsilon_t$ follows the ARCH(1) model: $\varepsilon_t=v_t^{1/2}\eta_t$ and $v_t=0.4+0.5\varepsilon_{t-1}^2$, $\eta_t\overset{i.i.d.}{\sim} N(0,1)$, $Z_{2t}\overset{i.i.d.}{\sim} U[-1,1]$, and $\eta_t$ and $Z_{2t}$ are independent; • DGP 5: $Z_{1t}=\text{sigmoid}(\theta_0Z_{2t})+\varepsilon_t$, where $\theta_0=1$, $\text{sigmoid}(x)=1/[1+\exp(-x)]$, and $\varepsilon_t$ and $Z_{2t}$ are generated as in DGP 3; • DGP 6: $Z_{1t}=\text{sigmoid}(\theta_0Z_{2t})+\varepsilon_t$, where $\theta_0=1$, and $\varepsilon_t$ and $Z_{2t}$ are generated as in DGP 4; • DGP 7: $Z_{1t}=\theta_0^2Z_{2t}+\theta_0Z_{2t}^2+\varepsilon_t$, where $\theta_0=5/4$ and $(\epsilon_t,Z_{2t})^{\top}\overset{i.i.d.}{\sim} N(0,\mbox{I}_2)$; • DGP 8: $Z_{1t}=\theta_0Z_{2t}+\varepsilon_t$, where $\theta_0=1$, $\varepsilon_t$ satisfies the AR(1) model: $\varepsilon_t=0.1\varepsilon_{t-1}+\eta_t$, and $(\eta_t,Z_{2t})^{\top}\overset{i.i.d.}{\sim} N(0,\mbox{I}_2)$; • DGP 9: $Z_{1t}=\theta_0Z_{1,t-1}+\varepsilon_t$, where $\theta_0=0.5$ and $\varepsilon_t \overset{i.i.d.}{\sim} t(7)$; • DGP 10: $Z_{1t}=\theta_0Z_{1,t-1}+\varepsilon_t$, where $\theta_0=0.5$ and $\varepsilon_t$ is generated as in DGP 2; • DGP 11: $Z_{1t}=\theta_{10}+\theta_{20}Z_{2t}+\varepsilon_t$, where $\theta_0=(\theta_{10},\theta_{20})^{\top}=(0.5, 1)^{\top}$, and $\varepsilon_t$ and $Z_{2t}$ are generated as in DGP 1; • DGP 12: $Z_{1t}=\theta_{10}+\theta_{20}Z_{2t}+\varepsilon_t$, where $\theta_0=(\theta_{10},\theta_{20})^{\top}=(0.5, 1)^{\top}$, $Z_{2t}=\zeta_t$, and $\varepsilon_t$ and $\zeta_t$ are generated as in DGP 2.

DGPs 1--10 consider ten models without intercept, while DPGs 11--12 regard two models with intercept. Among them, DGPs 1--8 and 11--12 are linear/nonlinear time series models or regressions with i.i.d. errors (DGPs 1, 3, 5, 7, and 11), conditionally heteroskedastic errors (DGPs 2, 4, 6, and 12), or correlated errors (DGP 8), and DGPs 9 and 10 are AR models with i.i.d. errors and conditionally heteroskedastic errors, respectively. We compute the MDD-based estimator $\widehat{\theta}_n$ of $\theta_0$ in DGPs 1--10, while we calculate $\widehat{\theta}_n=(\widehat{\theta}_{1n},\widehat{\theta}_{2n})^{\top}$ of $\theta_0$ in DGPs 11--12 by using the two-step estimation procedure. As a comparison, we also compute the estimator in DL:2004 for each DGP. For all considered estimators above, we use the conditioning variable $Z_{1,t-1}$ for DGPs 9--10 and $Z_{2t}$ for the other DGPs.

Based on the results from $1000$ repetitions, Table (ref) reports the bias, empirical standard deviation (ESD), and asymptotic standard deviation (ASD) of all considered estimators, where our proposed estimators and the estimator in DL:2004 are abbreviated as “MDD” and “DL”, respectively. From this table, we find that except for DGPs 5--6, both MDD and DL estimators have small biases, and the value of ASD is close to that of ESD for both estimators, in align with their related asymptotic normality results. For DGPs 5--6, both MDD and DL estimators have large biased when $n=50$ and $100$, and the bias of DL estimator remains large when $n=200$; in these two DGPs, the values of ASD and ESD for the MDD estimator get closed as $n$ increases to 200, whereas those for the DL estimator still have a significant difference for a large $n$ particularly when the errors are conditionally heteroskedastic. Note that our findings in DGPs 5--6 match the statement on page 163 of Tsay:2005 that the parameters in the sigmoid or smooth transition function are hard to estimate. Nevertheless, our simulations results in DGPs 5--6 show that the MDD estimator can outperform DL estimator significantly for the smooth transition model, especially when the sample size is small. Moreover, from Table (ref), we observe that the MDD estimator always has a much smaller value of ASD than the DL estimator, indicating a significant efficiency advantage of MDD estimator over DL estimator in all examined DGPs.

Next, we generate $1000$ Monte-Carlo repetitions of sample size $n=50$, $100$, and $200$ from the following four multivariate DGPs:

itemize[label=] • DGP 13: $Z_{1t}=A_0 Z_{2t}+\varepsilon_t$, where $A_0=\begin{pmatrix} \theta_{11,0}& \theta_{12,0}\\ \theta_{21,0}& \theta_{22,0}\end{pmatrix}$ with $\theta_0=(\theta_{11,0},\theta_{12,0},\theta_{21,0},\theta_{22,0})^{\top}$ $=(1,-1,1,2)^{\top}$, $Z_{2t}=\begin{pmatrix} 0.3 & 0\\ 0 & 0.2 \end{pmatrix}Z_{2,t-1}+\zeta_t$, and $(\zeta_t^{\top},\varepsilon_t^{\top})^{\top}\overset{i.i.d.}{\sim} N(0, \mbox{I}_{4})$; • DGP 14: $Z_{1t}=A_0 Z_{2t}+\varepsilon_t$, where $A_0$ and $Z_{2t}$ are defined as in DGP 13, $\varepsilon_t=(\varepsilon_{1,t},\varepsilon_{2,t})^{\top}$ follows the multivariate generalized ARCH (GARCH) model: $\varepsilon_t=V_t^{1/2}\eta_t$ and $V_t=\left(v_{i j,t}\right)_{i, j=1,2}$ is a $2\times 2$ symmetric matrix with $$ \left\{\begin{array}{l} v_{11,t}=0.1+0.8 v_{11,t-1}+0.1 \varepsilon_{1,t-1}^2, \\ v_{22,t}=0.1+0.8 v_{22,t-1}+0.1 \varepsilon_{2,t-1}^2, \\ v_{12,t}=0.7 \sqrt{v_{11,t} v_{22,t}}, \end{array}\right. $$ and $(\zeta_t^{\top},\eta_t^{\top})^{\top}\overset{i.i.d.}{\sim} N(0, \mbox{I}_{4})$; • DGP 15: $Z_{1t}=A_0 Z_{2t}+\varepsilon_t$, where $A_0$ and $Z_{2t}$ are defined as in DGP 13, $\varepsilon_t$ follows the vector AR(1) model: $\varepsilon_t=\begin{pmatrix} 0.2 & 0\\ 0 & 0.1 \end{pmatrix} \varepsilon_{t-1}+\eta_t$, and $(\zeta_t^{\top},\eta_{t}^{\top})^{\top}\overset{i.i.d.}{\sim} N(0, \mbox{I}_{4})$; • DGP 16: $Z_{1t}=A_0 Z_{1,t-1}+\varepsilon_t$, where $A_0$ is defined as in DGP 13 with $\theta_0=(0.6,-0.4,0.8,0.2)^{\top}$, and $\varepsilon_t\overset{i.i.d.}{\sim} N(0, \mbox{I}_{2})$.

DGPs 13, 14, and 15 are three time series models with i.i.d. errors, conditionally heteroskedastic errors, and correlated errors, respectively, and DGP 16 is a vector AR(1) model with i.i.d. errors. For all of these DGPs, we compute the MDD estimator $\widehat{\theta}_n$ with the conditioning variable $Z_{2t}$ for DGPs 13--15 and $Z_{1,t-1}$ for DGP 16. As a comparison, the DL estimator is also calculated with the same choices of conditioning variable as the MDD estimator.

Based on the results from $1000$ repetitions, Table (ref) reports the bias, ESD, and ASD of both MDD and DL estimators in DGPs 13--16. From this table, we can reach the similar conclusion as above that the MDD estimators are more efficient than the DL estimators in all examined DGPs.

table[table omitted — 5,791 chars of source]
table[table omitted — 6,046 chars of source]

Empirical Analysis

This section provides two real examples to demonstrate the importance of our proposed estimators. As before, the estimates from our two-step estimation method and the estimation method in DL:2004 are referred to as “MDD” and “DL” estimates, respectively.

A Univariate Example

In this subsection, we re-study the data set in LLZ:2016. This data set contains weekly closing prices of Hang Seng Index (HSI) from January 2000 to December 2007 with 418 observations in total. Define log-return (in percentage) $Y_t=100(\log P_t-\log P_{t-1})$, where $P_t$ is the closing price of HSI at time $t$. LLZ:2016 find that both conditional mean and variance of $y_t$ have the threshold effect, and they fit the conditional mean of $Y_t$ by an order 2 threshold autoregressive (TAR(2)) model with $Y_{t-1}$ being the threshold variable and $0$ being the threshold. Following their findings, we apply our two-step estimation method (with the conditioning variable $X_t=(Y_{t-1},...,Y_{t-4})^{\top}$) to obtain the fitted TAR(2) model below:

equation[equation omitted — 313 chars of source]

where the standard errors of MDD estimates are given in parentheses. From model ((ref)), we find that only the parameter for $Y_{t-2}$ in the regime with respect to $Y_{t-1}\leq 0$ is barely significantly different from zero at the 10% level. Note that the fitting result in LLZ:2016 gives a much stronger evidence for this significant parameter. However, their result highly relies on a specific conditionally heteroskedastic model of $Y_t$, whereas our result is robust allowing the conditional heteroskedasticity of $Y_t$ to have an unspecified form.

As a comparison, we also use the DL estimation method in DL:2004 to get the following fitted TAR(2) model:

equation[equation omitted — 312 chars of source]

where the same conditioning variable $X_t$ is adopted as for our two-step estimation method, and the standard errors of DL estimates are given in parentheses. Compared with the results in model ((ref)), the DL estimation is less efficient than the two-step estimation (as evidenced by the larger standard errors associated with the DL estimates in ((ref))), and it can not detect any significant parameter.

A Multivariate Example

In this subsection, we re-visit a benchmark data set in Tsay:2005, which consists of daily log returns of the SP500 index, the stock price of Cisco Systems, and the stock price of Intel Corporation from January 2, 1991 to December 31, 1999 with 2275 observations in total. We denote this 3-dimensional multivariate time series by $Y_t = (Y_{1t}, Y_{2t}, Y_{3t})^{\top}$. Tsay:2005 fits $Y_t$ via the following order 3 vector AR model:

equation[equation omitted — 94 chars of source]

where the above model is estimated by the least squares (LS) estimation method. However, since $\varepsilon_t$ has the conditional heteroskedasticity effect as shown in Tsay:2005, the standard errors of LS estimates are not reliable under the i.i.d. assumption of $\varepsilon_t$. To overcome this difficulty, we employ our two-step estimation method (with the conditioning variable $X_t=(Y_{t-1}^{\top},Y_{t-2}^{\top},Y_{t-3}^{\top})^{\top}$) to estimate model ((ref)), and present the corresponding MDD estimates in Table (ref). For comparison, the DL estimates (with the same conditioning variable $X_t$ as for the two-step estimation method) are also reported in this table. From Table (ref), we find that (i) the MDD estimates have much smaller values of standard errors than the DL estimates, lending a support that the MDD estimates are much more accurate than the DL estimates; and (ii) the MDD estimates are able to detect much more significant parameters in model ((ref)) than the DL estimates. Hence, the above findings illustrate the importance of non-integrable weighting function used by the two-step estimation method.

Overall, although both two-step and DL estimation methods allow for an unspecified form of conditional heteroskedasticity, our two real examples show that the former method could provide more accurate estimates than the latter one, and this advantage is more substantial when the data are multivariate.

table[table omitted — 2,188 chars of source]

Concluding Remarks

In this paper, we propose a new MDD-based estimation method for conditional moment models. This MDD-based estimation method roots in the idea of a continuum of unconditional moment restrictions, with a non-integrable weighting function to ensemble the information from different estimating moment functions. Under conditions that allow for time series data with conditional heteroskedasticity of unknown form, the proposed estimators by our MDD-based method are shown to be asymptotically normal with analytic asymptotic variances. Hence, all of our prosed estimators are easy for implementing statistical inference and have a large application scope for studying multiple linear/nonlinear time series models. Their importance is intensively demonstrated through simulations and two real examples. As a future work, it is interesting to extend our MDD-based estimation idea to the non-smooth function $h$ as in CP:2009,CP:2012 and GS:2012, and this may call for non-trivial technical treatments.

Acknowledgments

The codes and data used for this paper are accessible at “\url{https://github.com/hkusky/MDDEstimation.git}”.

\setcounter{equation}{0} \setcounter{section}{0}