EconBase
← Back to paper

Bilinear form test statistics for extremum estimation

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

24,350 characters · 7 sections · 24 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Bilinear form test statistics for extremum estimation

frontmatter\tnotetext[titlenote]{This research was supported by Fondecyt grant 11140433, Regione Autonoma della Sardegna Master and Back grant PRR-MAB-A2011-24192 (F. Crudu) and Fondecyt grant 1140580 (F. Osorio).} \ead{[email removed]} \ead{[email removed]} \cortext[cor1]{Correspondence to: Department of Economics and Statistics, University of Siena, Piazza San Francesco, 7/8 53100 Siena, Italy.} \address[DEPS]{Department of Economics and Statistics, University of Siena, Italy} \address[DMAT]{Departamento de Matem\'{a}tica, Universidad T\'{e}cnica Federico Santa Mar\'{i}a, Chile} \begin{abstract} This paper develops a set of test statistics based on bilinear forms in the context of the extremum estimation framework with particular interest in nonlinear hypothesis. We show that the proposed statistic converges to a conventional chi-square limit. A Monte Carlo experiment suggests that the test statistic works well in finite samples. \end{abstract} \begin{keyword} Extremum estimation \sep Gradient statistic \sep Bilinear form test \sep Nonlinear hypothesis. \JEL C12 \sep C14 \sep C69. \end{keyword}

Introduction

The purpose of this paper is to introduce a novel test statistic for extremum estimation (EE). In this very general setting Gourieroux:1995, Hayashi:2000, conventional test statistics are defined either in terms of differences (pseudo likelihood ratio or distance statistic) or in terms of quadratic forms (Wald, Lagrange multiplier also known as Rao's Rao:1948 score statistic). The test proposed in this paper is defined in terms of a bilinear form ($BF$). This approach is not entirely new as a bilinear form test for maximum likelihood was introduced by Terrell:2002 Lemonte:2016. Our test statistic has a conventional chi-square limit and, similarly to the Wald test, it is generally not invariant to the definition of the null hypothesis. It is, though, easy to see that in the context of linear models the $BF$ test is equal to the distance statistic, which is, on the other hand, invariant. Furthermore, when nonlinear models are involved our Monte Carlo simulations suggest that the discrepancy induced by equivalent definitions of the null hypothesis is relatively small when compared, e.g., to the Wald test. In the general case, the computational burden associated to the $BF$ statistic is comparable to that of the distance metric statistic, since both the estimator under the null and under the alternative must be calculated. To the best of our knowledge this is the first paper that deals with this problem in the context of EE.

The remainder of the paper unfolds as follows. Section (ref) contains the description of the test statistics for a generic, potentially nonlinear, null hypothesis and their asymptotic properties; the asymptotic results and the corresponding proofs are presented in a concise fashion and are mostly based on the results in Gourieroux:1995. In Section (ref) we study, via Monte Carlo experiments, the finite sample properties of the test in comparison with other more conventional EE test statistics. Section (ref) offers some conclusions while the appendices contain the proofs of the asymptotic results.

A bilinear form test statistic

Let us consider a scalar objective function $Q_n(\mbox{\boldmath $\beta$})$ that depends on a set of data ${\mbox{\boldmath $w$}_i}, i=1,\dots,n$ with ${\mbox{\boldmath $w$}_i}\in\mathbb{R}^k$ and $\mbox{\boldmath $\beta$}\in\mathcal{B} \subset\mathbb{R}^p$ where $\mathcal{B}$ is compact. The EE for our objective function can be defined as

equation[equation omitted — 159 chars of source]

Let us now suppose that we want to test the following null hypothesis

equation[equation omitted — 109 chars of source]

given that $\mbox{\boldmath $g$}:\mathbb{R}^p\to\mathbb{R}^q$ is a continuously differentiable function and $\mbox{\boldmath $G$}(\mbox{\boldmath $\beta$}) = \partial\mbox{\boldmath $g$}(\mbox{\boldmath $\beta$})/\partial\mbox{\boldmath $\beta$}^\top$ is a $q\times p$ matrix with $\operatorname{rk}(\mbox{\boldmath $G$}(\mbox{\boldmath $\beta$})) = q$. The resulting constrained estimator is defined as the solution of the Lagrangian problem

equation[equation omitted — 207 chars of source]

where $\mbox{\boldmath $\lambda$}$ denotes a vector of Lagrange multipliers. Hence,

equation[equation omitted — 210 chars of source]

The null hypothesis in Equation ((ref)) can be tested, for example, by means of the simple Wald ($W$) test, that only requires the unconstrained estimator or either the Lagrange multiplier ($LM$) test or the distance metric ($D$) statistic that both require the constrained estimator in Equation ((ref)). The $BF$ tests that we propose are generalizations of Terrell's gradient statistic Terrell:2002 to the EE context.\footnote{Sometimes the term gradient statistic is used to indicate the $LM$ test for GMM Ruud:2000. To avoid confusion we prefer the expression bilinear form test and the corresponding abbreviation $BF$.} Let us first define $\mbox{\boldmath $A$}_n(\mbox{\boldmath $\beta$}_0)\mathrel{\vcenter{\baselineskip0.5ex \lineskiplimit0pt \hbox{\scriptsize.}\hbox{\scriptsize.}}} = \partial^2 Q_n(\mbox{\boldmath $\beta$}_0)/\partial\mbox{\boldmath $\beta$}\partial\mbox{\boldmath $\beta$}^\top$ and assume that $\mbox{\boldmath $A$}_n(\mbox{\boldmath $\beta$}_0) \stackrel{\sf a.s.}{\to} \mbox{\boldmath $A$}$ uniformly. Let us also assume that \[ \sqrt{n}\,\frac{\partial Q_n(\mbox{\boldmath $\beta$}_0)}{\partial\mbox{\boldmath $\beta$}}\stackrel{\sf D}{\to} \mathsf{N}_p(\mbox{\boldmath $0$}, \mbox{\boldmath $B$}). \] Furthermore, let $\mbox{\boldmath $G$}\mathrel{\vcenter{\baselineskip0.5ex \lineskiplimit0pt \hbox{\scriptsize.}\hbox{\scriptsize.}}} =\mbox{\boldmath $G$}(\mbox{\boldmath $\beta$}_0)$, $\mbox{\boldmath $S$} = \mbox{\boldmath $G$}\{-\mbox{\boldmath $A$}\}^{-1} \mbox{\boldmath $G$}^\top$ and $\mbox{\boldmath $\Omega$} = \mbox{\boldmath $G$}\mbox{\boldmath $A$}^{-1}\mbox{\boldmath $B$}\mbox{\boldmath $A$}^{-1}\mbox{\boldmath $G$}^\top$. Then,

equation[equation omitted — 337 chars of source]

where $\widetilde{\mbox{\boldmath $\lambda$}}{}_n$ is the solution for $\mbox{\boldmath $\lambda$}$ in the Lagrangian problem defined by Equation ((ref)). The $BF$ statistic also has the following alternative formulations. Let $\mbox{\boldmath $G$}^+= \mbox{\boldmath $G$}^\top\{\mbox{\boldmath $G$}\mbox{\boldmath $G$}^\top\}^{-1}$ denote the Moore-Penrose inverse of $\mbox{\boldmath $G$}$ Magnus:2007. Then,

align[align omitted — 857 chars of source]

Let us define $\mbox{\boldmath $P$}_G \mathrel{\vcenter{\baselineskip0.5ex \lineskiplimit0pt \hbox{\scriptsize.}\hbox{\scriptsize.}}} = \mbox{\boldmath $G$}^+\mbox{\boldmath $G$}$ and assume that $\mbox{\boldmath $B$} = -\mbox{\boldmath $A$}$, which leads to $\mbox{\boldmath $S$} = \mbox{\boldmath $\Omega$}$. We then obtain the following specifications:

align[align omitted — 1,365 chars of source]

The assumption that $\mbox{\boldmath $B$} = -\mbox{\boldmath $A$}$ is not very restrictive as it may include as special cases maximum likelihood and GMM statistics Hayashi:2000. Next, we consider a quadratic objective function where this condition is satisfied.

rmkLet $Q_n(\mbox{\boldmath $\beta$}) = -{\textstyle\frac{1}{2}}\mbox{\boldmath $f$}_n^\top(\mbox{\boldmath $\beta$})\mbox{\boldmath $W$}^{-1}\mbox{\boldmath $f$}_n(\mbox{\boldmath $\beta$})$ where $\mbox{\boldmath $f$}_n(\mbox{\boldmath $\beta$})$ is a set of sample moment conditions and $\mbox{\boldmath $W$}$ is a conformable positive definite matrix, then \[ \frac{\partial Q_n(\mbox{\boldmath $\beta$})}{\partial\mbox{\boldmath $\beta$}} = -\mbox{\boldmath $F$}_n^\top(\mbox{\boldmath $\beta$})\mbox{\boldmath $W$}^{-1} \mbox{\boldmath $f$}_n(\mbox{\boldmath $\beta$}), \] with $\mbox{\boldmath $F$}_n(\mbox{\boldmath $\beta$}) = \partial\mbox{\boldmath $f$}_n(\mbox{\boldmath $\beta$})/\partial\mbox{\boldmath $\beta$}^\top$, and \[ \frac{\partial^2 Q_n(\mbox{\boldmath $\beta$})}{\partial\mbox{\boldmath $\beta$}\partial\mbox{\boldmath $\beta$}^\top} = -\Big[\frac{\partial\mbox{\boldmath $F$}_n^\top(\mbox{\boldmath $\beta$})}{\partial\mbox{\boldmath $\beta$}}\Big] \big[\mbox{\boldmath $W$}^{-1}\mbox{\boldmath $f$}_n(\mbox{\boldmath $\beta$})\big] - \mbox{\boldmath $F$}_n^\top(\mbox{\boldmath $\beta$})\mbox{\boldmath $W$}^{-1}\mbox{\boldmath $F$}_n(\mbox{\boldmath $\beta$}), \] where $[\cdot][\cdot]$ denotes array multiplication Wei:1998. If $\mbox{\boldmath $f$}_n(\mbox{\boldmath $\beta$}_0)$ converges to its expected value, i.e. zero, its derivatives converge almost surely to finite full rank matrices and $\sqrt{n}\mbox{\boldmath $f$}_n(\mbox{\boldmath $\beta$}_0) \stackrel{\sf D}{\to} \mathsf{N}(\mbox{\boldmath $0$},\mbox{\boldmath $W$})$, then we find $\mbox{\boldmath $B$} = \mbox{\boldmath $F$}^\top \mbox{\boldmath $W$}^{-1}\mbox{\boldmath $F$}$ and $\mbox{\boldmath $A$} = -\mbox{\boldmath $F$}^\top\mbox{\boldmath $W$}^{-1}\mbox{\boldmath $F$}$. Hence, $\mbox{\boldmath $B$}=-\mbox{\boldmath $A$}$ holds.

The following proposition shows that the $BF$ tests are asymptotically equivalent and have a conventional chi-square limit.

propositionUnder the assumptions of Property 24.16 and Property 24.10 in Gourieroux:1995, with $\mbox{\boldmath $g$}:\mathbb{R}^p\to\mathbb{R}^q$ being a continuously differentiable function and $\mbox{\boldmath $G$} (\mbox{\boldmath $\beta$}) = \partial\mbox{\boldmath $g$}(\mbox{\boldmath $\beta$})/\partial\mbox{\boldmath $\beta$}^\top$ a $q\times p$ matrix with $\operatorname{rk}(\mbox{\boldmath $G$}(\mbox{\boldmath $\beta$})) = q$, \[ BF_k\stackrel{\sf D}{\to}\chi^2_q, \qquad k=1,2,3. \] If, in addition, $\mbox{\boldmath $B$} = -\mbox{\boldmath $A$}$ holds, then \[ BF_k\stackrel{\sf D}{\to}\chi^2_q, \qquad k=4,5,6,7. \]
pfSee (ref).
rmkWhen $Q_n(\mbox{\boldmath $\beta$}) = \overline{\ell}_n(\mbox{\boldmath $\beta$})$ is the log-likelihood function we obtain that the $BF$ statistic is given by \begin{equation} BF = \boldmath $U$_n^\top(\widetilde{\boldmath $\beta$}_n)\boldmath $G$^+\boldmath $g$(\widehat{\boldmath $\beta$}_n), \end{equation} where $\mbox{\boldmath $U$}_n(\mbox{\boldmath $\beta$}) = \partial\overline{\ell}_n(\mbox{\boldmath $\beta$})/\partial\mbox{\boldmath $\beta$}$ denotes the score function. We must highlight that ((ref)) is an extension of the test proposed by Terrell:2002 to tackle nonlinear hypotheses.
rmkIt is interesting to see that in the case of the linear model, $D$ and $BF$ are equal. Let us consider, the example in Hansen:2006. The $BF$ statistic is \[ BF = (\mbox{\boldmath $y$} - \mbox{\boldmath $X$}\widetilde{\mbox{\boldmath $\beta$}}_n)^\top\mbox{\boldmath $X$}\mbox{\boldmath $B$}^{-1}\mbox{\boldmath $X$}^\top \mbox{\boldmath $X$}(\widehat{\mbox{\boldmath $\beta$}}_n - \widetilde{\mbox{\boldmath $\beta$}}_n). \] Since $\widehat{\mbox{\boldmath $\beta$}}_n = (\mbox{\boldmath $X$}^\top\mbox{\boldmath $X$})^{-1}\mbox{\boldmath $X$}^\top\mbox{\boldmath $y$}$ and $\mbox{\boldmath $X$}^\top (\mbox{\boldmath $y$} - \mbox{\boldmath $X$}\widehat{\mbox{\boldmath $\beta$}}_n) = \mbox{\boldmath $0$}$, it follows immediately that $BF = D$.

Next proposition establishes the asymptotic equivalence between $BF$ and $LM$ tests for nonlinear hypothesis Boos:1992.

propositionThe $BF$ test statistic in Equation ((ref)) and the Lagrange multiplier test statistic \[ LM \mathrel{\vcenter{\baselineskip0.5ex \lineskiplimit0pt \hbox{\scriptsize.}\hbox{\scriptsize.}}} = n\widetilde{\mbox{\boldmath $\lambda$}}{}_n^\top\mbox{\boldmath $S$}\mbox{\boldmath $\Omega$}^{-1}\mbox{\boldmath $S$}\widetilde{\mbox{\boldmath $\lambda$}}_n, \] are asymptotically equivalent under $H_0:\mbox{\boldmath $g$}(\mbox{\boldmath $\beta$}_0) = \mbox{\boldmath $0$}$. Their common asymptotic distribution is $\chi^2_q$.
pfSee (ref).
table*[table* omitted — 1,782 chars of source]

Monte Carlo simulations

To study the finite sample properties of the $BF$ statistic we consider two equivalent nonlinear null hypotheses, as in Gregory:1985 Hansen:2006, Lafontaine:1986. The $BF$ test, which is not invariant to the specification of the null, is compared against the $W$, $LM$ and $D$ statistics. While the first test is known to be not invariant, the last two tests are invariant and work well in finite samples (see, for instance, Dagenais:1991 and Hansen:2006). The performance of the tests is measured in terms of how close the empirical size is to the 5% nominal size and in terms of the discrepancy between the empirical sizes produced by competing equivalent hypotheses. Here, the distance metric statistic $D$ is defined as \[ D \mathrel{\vcenter{\baselineskip0.5ex \lineskiplimit0pt \hbox{\scriptsize.}\hbox{\scriptsize.}}} = n(Q_n(\widetilde{\mbox{\boldmath $\beta$}}_n) - Q_n(\widehat{\mbox{\boldmath $\beta$}}_n)), \] where $Q_n(\mbox{\boldmath $\beta$})$ is the objective function of the nonlinear least squares estimator. In our experiment the $BF$ statistic defined in Equation ((ref)) was used.

figure*[figure* omitted — 1,360 chars of source]

Setup

We consider the model specification \[ \mbox{\boldmath $y$} = \mbox{\boldmath $1$}_n\beta_1 + \mbox{\boldmath $x$}_2\beta_2 + \exp(\mbox{\boldmath $x$}_3\beta_3) + \mbox{\boldmath $\varepsilon$}, \] where $\mbox{\boldmath $1$}_n$ is a $n$-vector of ones, $\mbox{\boldmath $x$}_j\sim \mathsf{N}_n(\mbox{\boldmath $0$}, 0.16\,\mbox{\boldmath $I$})$, $j=2,3$ and $\mbox{\boldmath $\varepsilon$}\sim \mathsf{N}_n(\mbox{\boldmath $0$}, 0.16\,\mbox{\boldmath $I$})$. Moreover, we consider the following combinations of parameters \[ (\beta_1,\beta_2,\beta_3) \in \{(1,10,0.1),(1,5,0.2),(1,2,0.5),(1,1,1)\}, \] and sample sizes $n\in\{20,50,100,500\}$. We test two equivalent null hypotheses

equation[equation omitted — 69 chars of source]

and

equation[equation omitted — 62 chars of source]

The number of Monte Carlo replications is set to 5000. In addition, we compute the empirical power under the alternative hypotheses $H_1^A: \beta_2 - \delta/\beta_3 = 0$, and $H_1^B: \beta_2\beta_3 - \delta = 0$ for different values of $\delta$. The R code to perform the simulations described in this section and some additional results are available at github.\footnotemark[2]\footnotetext[2]{URL: \url{https://github.com/faosorios/BF_EE}}

Comments on the simulations

The results in Table (ref) suggest that the $BF$ test works well in finite samples even when the sample size is as small as $n=20$. In most of the considered cases the $BF$ test outperforms the distance statistic $D$ as well as the $LM$ test. It is worth noticing that, unlike $W$, the $BF$ test is not very sensitive to the specification of the null hypothesis. Empirical power of the $BF$ and $LM$ tests is displayed in Figure (ref). Using the $BF$ test may cause some loss of power. However, as expected, as the sample size increases the empirical power of the $LM$ and $BF$ tests becomes indistinguishable.

Concluding remarks

In this paper we introduced a set of bilinear form tests for EE that may be considered as a generalization of Terrell's gradient statistics Terrell:2002. The asymptotic distribution of the proposed tests is chi-square with degrees of freedom equal to the number of restrictions. A Monte Carlo experiment shows that the $BF$ test works well in finite samples and that it generally outperforms its competitors. Furthermore, while the $BF$ test is not generally invariant to the specification of the null, its finite sample performance seems to be only marginally affected by such a property. It is worth noticing that, despite the favorable finite sample properties, the $BF$ test requires the estimation of the parameters of interest both under the null and under the alternative. This feature may make it less attractive, from a computational point of view, when compared to some of its competitors, e.g. the $LM$ test, that only require the estimation of the parameters under the alternative. Nonetheless, this type of development offers yet another alternative for carrying out hypothesis tests in such general contexts as quadratic inference functions Qu:2000, generalized empirical likelihood Newey:2004, and maximum L$q$-likelihood estimation Ferrari:2010. We must emphasize that a detailed study of the properties of local power and invariance of the $BF$ test deserves further exploration along the lines of, for instance, Dagenais:1991 and Lemonte:2016.

Acknowledgements

The authors acknowledge the suggestions from an anonymous referee which helped to improve the manuscript.