The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
88,222 characters
A New Wald Test for Hypothesis Testing Based on MCMC outputs
\title{{\LARGE \textbf{A New Wald Test for Hypothesis Testing Based on MCMC outputs}}\thanks{
Li gratefully acknowledges the financial support of the Chinese Natural
Science Fund (No. 71271221)�� Yu would like to acknowledge the
financial support from Singapore Ministry of Education Academic Research
Fund Tier 2 under the grant number MOE2011-T2-2-096 and Tier 3 under the
grant number MOE2013-T3-1-017. Yong Li, Hanqing Advanced Institute of
Economics and Finance, Renmin University of China, Beijing, 100872, P.R.
China. Email: [email removed]. Xiao-Bin, Liu, School of Economics,
Singapore Management University, 90 Stamford Road, Singapore 178903. Jun Yu,
School of Economics and Lee Kong Chian School of Business, Singapore
Management University, 90 Stamford Road, Singapore 178903. Email:
[email removed]. URL: http://www.mysmu.edu/faculty/yujun/. }}
\author{Yong Li \\
\emph{Renmin University} \and Xiaobin Liu \\
\emph{Singapore Management University} \and Jun Yu \\
\emph{Singapore Management University} \and Tao Zeng\\
\emph{Wuhan University}}
\maketitle
\begin{abstract}
In this paper, a new and convenient $\chi^2$ wald test based on MCMC outputs is proposed for
hypothesis testing. The new statistic can be explained as MCMC version
of Wald test and has several important advantages that make it very
convenient in practical applications. First, it is well-defined under
improper prior distributions and avoids Jeffrey-Lindley's paradox. Second,
it's asymptotic distribution can be proved to follow the $\chi^2$
distribution so that the threshold values can be easily calibrated from this
distribution. Third, it's statistical error can be derived using the Markov
chain Monte Carlo (MCMC) approach. Fourth, most importantly, it is only
based on the posterior MCMC random samples drawn from the posterior
distribution. Hence, it is only the by-product of the posterior outputs and
very easy to compute. In addition, when the prior information is available,
the finite sample theory is derived for the proposed test statistic. At
last, the usefulness of the test is illustrated with several applications to
latent variable models widely used in economics and finance. \vskip0.4cm
\noindent \textit{JEL classification:} C11, C12 \noindent \newline
\textit{Keywords:} Bayesian $\chi^2$ test; Decision theory; Wald test;
Markov chain Monte Carlo; Latent variable models,
\end{abstract}
\vskip 2cm \baselineskip=17pt
\section{Introduction}
Latent variable models have been widely used in economics, finance, and many
other disciplines. Two typical models are the dynamic stochastic general
equilibrium models in macroeconomics and stochastic volatility models in
finance. The latent variable models are generally indexed by the latent
variable and the parameter. In many latent variable models, the latent
variable is generally high-dimensional so that the observed likelihood
function which is a marginal integral on the latent variable is often
intractable and becomes difficult to evaluate accurately. Consequently, the
statistical inference for latent variable models is nontrivial in practice.
In the recent years, Bayesian MCMC methods have been applied in more and
more applications in economics and finance due to that they make it possible
to fit increasingly complex models, especially latent variable models, see
Geweke, et al (2011) and reference therein.
In economic research, the point null hypothesis test is a fundamental topic
in statistical inference. Under the Bayesian paradigm, the Bayes factors
(BFs) are the corner-stone of Bayesian hypothesis testing (e.g.
Jeffreys,1961; Kass and Raftery 1995; Geweke, 2007). Unfortunately, the BFs
are not problem-free. First, the BFs are sensitive to the prior distribution
and subjects to the notorious Jeffreys-Lindley's paradox; see for example,
Kass and Raftery (1995), Poirier (1995), Robert (1993, 2001). Second, the
calculation of BFs generally involves the evaluation of marginal likelihood.
In many cases, the evaluation of marginal likelihood is often difficult.
Not surprisingly, some alternative strategies have been proposed to test a
point null hypothesis in the Bayesian literature. In recent years, on the
basis of the statistical decision theory, several interesting Bayesian
approaches to replace BFs have been developed for hypothesis testing. For
example, Bernardo and Rueda (2002, BR hereafter) demonstrated that BFs for
the Bayesian hypothesis testing can be regarded as a decision problem with a
simple zero-one discrete loss function. However, the zero-one discrete
function requires the use of non-regular (not absolutely continuous) prior
and this is why BF leads to Jeffreys-Lindley's paradox. BR further suggested
using a continuous loss function, based on the well-known continuous
Kullback-Leibler (KL) divergence function. As a result, it was shown in BR
that their Bayesian test statistic does not depend on any arbitrary constant
in the prior. However, BR's approach has some disadvantages. First, the
analytical expression of the KL loss function required by BR is not always
available, especially for latent variable models. Second, the test statistic
is not a pivotal quantity. Consequently, BR had to use subjective threshold
values to test the hypothesis.
To deal with the computational problem in BR in latent variable models, Li
and Yu (2012, LY hereafter) developed a new test statistic based on the $
\mathcal{Q}$ function in the Expectation-Maximization (EM) algorithm. LY
showed that the new statistic is well-defined under improper priors and easy
to compute for latent variable models. Following the idea of McCulloch
(1989), LY proposed to choose the threshold values based on the Bernoulli
distribution. However, like the test statistic proposed by BR, the test
statistic proposed by LY is not pivotal. Moreover, it is not clear if the
test statistic of LY can resolve Jeffreys-Lindley's paradox.
Based on the difference between the deviances, Li, Zeng and Yu (2014, LZY
hereafter) developed another Bayesian test statistic for hypothesis testing.
This test statistic is well-defined under improper priors, free of
Jeffreys-Lindley's paradox, and not difficult to compute. Moreover, its
asymptotic distribution can be derived and one may obtain the threshold
values from the asymptotic distribution. Unfortunately, in general the
asymptotic distribution depends on some unknown population parameters and
hence the test is not pivotal. With sharing the nice properties with Li,
Zeng and Yu (2014, LZY hereafter), Li, Liu and Yu (2015)(2015, LLY
hereafter) further proposed a pivotal Bayesian test statistic, based on a
quadratic loss function, to test a point null hypothesis within the
decision-theoretic framework. However, LLY required to evaluate the first
derivative of the observed log-likelihood. As to the latent variable models,
because the observed log-likelihood is often intractable, this still posed
some tedious computational efforts although there have been several
interesting methods for evaluating the first derivative, such as EM
algorithm, Kalman filter or Particle filter.
In the paper, we want to propose another novel, easy-to-implement Bayesian
statistic for hypothesis testing in the framework of latent variable models.
The new statistic can share the important advantages with LLY. First, it is
well-defined under improper prior distributions and avoids Jeffrey-Lindley's
paradox. Second, under some mild regularity conditions, the statistic is
asymptotically equivalent to the Wald test. Hence, from the large sample
theory, it's asymptotic distribution can be derived to follow the $\chi^2$
distribution so that the threshold values can be easily calibrated from this
distribution. Third, it's statistical error can be derived using the Markov
chain Monte Carlo (MCMC) approach. In addition, most importantly, compared
with the previous test statistics, it is extremely convenient for the latent
variable models. We don't need to evaluate the first-order derivative of the
observed log-likelihood function, which is time consuming and difficult for
the latent variable models. We just need the MCMC output of posterior
simulation. The only effort we should make is the inverse of the posterior
variance matrix of the interest parameter in hypothesis testing.
Fortunately, in most applications, the dimension of the interest parameter
is often not so high that our method can be easily applied. In addition,
when the prior information is available, we establish the finite sample
theoretical properties for the proposed test statistic.
The paper is organized as follows. Section 2 presents the Bayesian analysis
for latent variable models. Section 3 develops the new Bayesian test
statistic from the decisional viewpoint and establishes its finite and large
sample theoretical properties. Section 4 illustrates the new method by using
three real examples in economics and finance. Section 5 concludes the paper.
Appendix collects the proof of all the theoretical results.
\section{Bayesian analysis of latent variable models}
Without loss of generality, let $\mathbf{y}=(\mathbf{y}_{1},\mathbf{y}
_{2},\cdots ,\mathbf{y}_{n})^{T}$ denote observed variables and ${
\mbox{\boldmath${z}$}}=({\mbox{\boldmath${z}$}}_{1},{\mbox{\boldmath${z}$}}
_{2},\cdots ,{\mbox{\boldmath${z}$}}_{n})^{T},$ the latent variables. The
latent variable model is indexed by the parameter, ${\mbox{\boldmath${
\vartheta}$}}$. Let $p(\mathbf{y}|{\mbox{\boldmath${\vartheta}$}})$ be the
likelihood function of the observed data, and $p(\mathbf{y},{
\mbox{\boldmath${z}$}}|{\mbox{\boldmath${\vartheta}$}}),$ the complete
likelihood function. The relationship between these two likelihood functions
is:
\begin{equation}
p(\mathbf{y}|{\mbox{\boldmath${\vartheta}$}})=\int p(\mathbf{y},{
\mbox{\boldmath${z}$}}|{\mbox{\boldmath${\vartheta}$}})d{\mbox{
\boldmath${z}$}.} \label{eq01}
\end{equation}
In many latent variable modes, especially dynamic latent variable models,
the latent variable $\mathbf{z}$ is often dependent on the sample size.
Hence, the integral is high-dimensional and often does not have an
analytical expression so that it is generally very difficult to evaluate.
Consequently, the statistical inferences, such as estimation and hypothesis
testing, are difficult to implement if they are based on the popular maximum
likelihood approach.
In recent years, it has been documented that the latent variables models can
be simply and efficiently analyzed using MCMC techniques under the Bayesian
framework. For details about Bayesian analysis of latent variable models via
MCMC such as algorithms, examples and references, see Geweke, et al. (2011).
Let $p({\mbox{\boldmath${\vartheta}$}})$ be prior distribution of unknown
parameter ${\mbox{\boldmath${\vartheta}$}}$. Owing to the complexity induced
by latent variables, the observed likelihood $p(\mathbf{y}|{
\mbox{\boldmath${\vartheta}$}})$ is often intractable, hence it is almost
impossible to evaluate the expectation of the posterior density $p({
\mbox{\boldmath${\vartheta}$}}|\mathbf{y})$ directly. To alleviate this
difficulty, in the posterior analysis, the popular data-augmentation
strategy(Tanner and Wong, 1987) is applied to augment the observed variable $
\mathbf{y}$ with the latent variable $\mathbf{z}$. Then, the well-known
Gibbs sampler can be used to generate random samples from the joint
posterior distribution $p({\mbox{\boldmath${\vartheta}$}} ,\mathbf{z}|
\mathbf{y})$. More concretely, we start with an initial value $[{
\mbox{\boldmath${\vartheta}$}}^{(0)},,\mathbf{z}^{(0)}]$, and then simulates
one by one; at the $j$th iteration, with current values $[{
\mbox{\boldmath${\vartheta}$}}^{(j)},\mathbf{z}^{(j)}]:$
\begin{description}
\item ~~~~~~~~~~~~~~~(a) Generate ${\mbox{\boldmath${\vartheta}$}}^{(j+1)}$
from $p({\mbox{\boldmath${\vartheta}$}} |\mathbf{z}^{(j)},\mathbf{y})$;
\item ~~~~~~~~~~~~~~~(b) Generate $\mathbf{z}^{(j+1)}$ from $p(\mathbf{z}|{
\mbox{\boldmath${\vartheta}$}}^{(j+1)},\mathbf{z})$.
\end{description}
After the burning-in phase, that is, sufficiently many iterations of this
iteration procedure, the simulated random samples can be regarded as
efficient random observations from the joint posterior distribution $p({
\mbox{\boldmath${\vartheta}$}} ,\mathbf{z}|\mathbf{y})$.
The statistical inference can be established on the efficient random
observations drawn from the posterior distribution. Bayesian estimates of ${
\mbox{\boldmath${\vartheta}$}}$ and latent variables $\mathbf{z}$ as well as
their standard errors can be easily obtained via the corresponding sampling
mean and sample covariance matrix of the generated random observations.
Specifically, let $\{{\mbox{\boldmath${\vartheta}$}}^{(j)},\mathbf{z}
^{(j)},j=1,2,\cdots, J\}$ be effective random observations generated form
the joint posterior distribution $p({\mbox{\boldmath${\vartheta}$}},\mathbf{z
}|\mathbf{y})$. Then the joint Bayesian estimates of ${\mbox{\boldmath${
\vartheta}$}},\mathbf{z}$, as well as the estimates of their covariance
matrix can be obtained as follows:
\begin{eqnarray*}
& &{\mbox{\boldmath${{\widehat\vartheta}}$}}=\frac {1}{J}\sum_{j=1}^J{
\mbox{\boldmath${\theta}$}}^{(j)},~ \widehat{Var}({\mbox{\boldmath${
\vartheta}$}}|\mathbf{y})=\frac{1}{J}\sum_{j=1}^J ({\mbox{\boldmath${
\vartheta}$}}^{(j)}-{\mbox{\boldmath${{\widehat{\vartheta}}}$}})({
\mbox{\boldmath${\vartheta}$}}^{(j)}-{\mbox{\boldmath${\widehat{\vartheta}}$}
})^{\prime} \\
& &\mathbf{\widehat{z}}=\frac {1}{J}\sum_{j=1}^J\mathbf{z}^{(j)},~ \widehat{
Var}(\mathbf{z}|\mathbf{y})=\frac{1}{J}\sum_{j=1}^J (\mathbf{z}^{(j)}-
\mathbf{\widehat{z}})(\mathbf{z}^{(j)}-\mathbf{\widehat{z}})^{\prime}.
\end{eqnarray*}
\section{Bayesian Hypothesis Testing from the Decision Theory}
\subsection{Testing a point null hypothesis}
It is assumed that a probability model $M\equiv \{p(\mathbf{y}|{
\mbox{\boldmath${\theta}$}},{\mbox{\boldmath${\psi}$}})\}$ is used to fit
the data. We are concerned with a point null hypothesis testing problem
which may arise from the prediction of a particular theory. Let ${
\mbox{\boldmath${\theta}$}}\in \mathbf{\Theta }$ denote a vector of $p$
-dimensional parameters of interest and ${\mbox{\boldmath${\psi}$}}\in
\mathbf{\Psi }$ a vector of $q$-dimensional nuisance parameters. The problem
of testing a point null hypothesis is given by
\begin{equation}
\left\{
\begin{array}{cc}
H_{0}: & {\mbox{\boldmath${\theta}$}}={\mbox{\boldmath${\theta}$}}_{0} \\
H_{1}: & {\mbox{\boldmath${\theta}$}}\neq {\mbox{\boldmath${\theta}$}}_{0}
\end{array}
.\right. \label{eq03}
\end{equation}
The hypothesis testing may be formulated as a decision problem. It is
obvious that the decision space has two statistical decisions, to accept $
H_{0}$ (name it $d_{0}$) or to reject $H_{0}$ (name it $d_{1}$). Let $\{
\mathcal{L}[d_{i},({\mbox{\boldmath${\theta}$}},{\mbox{\boldmath${\psi}$}}
)],i=0,1\}$ be the loss function of statistical decision. Hence, a natural
statistical decision to reject $H_{0}$ can be made when the expected
posterior loss of accepting $H_{0}$ is sufficiently larger than the expected
posterior loss of rejecting $H_{0}$, i.e.,
\begin{equation*}
~~~\mathbf{T}(\mathbf{y},{\mbox{\boldmath${\theta}$}}_{0})=\int_{\Theta
}\int_{\Psi }\left\{ \mathcal{L}[d_{0},({\mbox{\boldmath${\theta}$}},{
\mbox{\boldmath${\psi}$}})]-\mathcal{L}[d_{1},({\mbox{\boldmath${\theta}$}},{
\mbox{\boldmath${\psi}$}})]\right\} p({\mbox{\boldmath${\theta}$}},{
\mbox{\boldmath${\psi}$}}|\mathbf{y})\mbox{d}{\mbox{\boldmath${\theta}$}
\mbox{d}\mbox{\boldmath${\psi}$}}>c\geq 0,
\end{equation*}
where $\mathbf{T}(\mathbf{y},{\mbox{\boldmath${\theta}$}}_{0})$ is a
Bayesian test statistic; $p({\mbox{\boldmath${\theta}$}},{
\mbox{\boldmath${\psi}$}}|\mathbf{y})$ the posterior distribution with some
given prior $p({\mbox{\boldmath${\theta}$}},{\mbox{\boldmath${\psi}$}})$; $c$
a threshold value. Let $\triangle \mathcal{L}[H_{0},({\mbox{\boldmath${
\theta}$}},{\mbox{\boldmath${\psi}$}})]=\mathcal{L}[d_{0},({
\mbox{\boldmath${\theta}$}},{\mbox{\boldmath${\psi}$}})]-\mathcal{L}[d_{1},({
\mbox{\boldmath${\theta}$}},{\mbox{\boldmath${\psi}$}})]$ be the net loss
difference function which can generally be used to measure the evidence
against $H_{0}$ as a function of $({\mbox{\boldmath${\theta}$}},{
\mbox{\boldmath${\psi}$}})$. Hence, the Bayesian test statistic can be
rewritten as
\begin{equation*}
~~~\mathbf{T}(\mathbf{y},{\mbox{\boldmath${\theta}$}}_{0})=E_{{
\mbox{\boldmath${\vartheta}$}}|\mathbf{y}}\left( \triangle \mathcal{L}
[H_{0},({\mbox{\boldmath${\theta}$}},{\mbox{\boldmath${\psi}$}})]\right).
\end{equation*}
\begin{remark}
When the equal prior $p\left( {\mbox{\boldmath${
\theta}$}} ={\mbox{\boldmath${\theta }$}}_{0}\right) =p\left( {
\mbox{\boldmath${\theta}$}} \neq {\mbox{\boldmath${\theta }$}}_{0}\right) =
\frac{1}{2}$ and the net loss function is taken as
\begin{equation*}
\Delta \mathcal{L}\left( H_{0},{\mbox{\boldmath${\theta}$}},{
\mbox{\boldmath${\psi }$}}\right) =
\begin{cases}
-1 & \text{if }{\mbox{\boldmath${\theta}$}} ={\mbox{\boldmath${\theta }$}}
_{0} \\
1, & \text{if }{\mbox{\boldmath${\theta}$}} \neq {\mbox{\boldmath${\theta }$}
}_{0}
\end{cases}
\end{equation*}
following BR (2002) and Li and Yu (2012), the Bayesian test statistic can be
given by
\begin{equation*}
T\left(\mathbf{y},{\mbox{\boldmath${\theta }$}}_{0}\right) =\int_{\Theta
}\int_{\Psi }\Delta \mathcal{L}\left( H_{0},{
\mbox{\boldmath${\theta ,\psi
}$}}\right) p\left( {\mbox{\boldmath${\theta ,\psi |\mathbf{y}}$}}\right) d{
\mbox{\boldmath${\theta d\psi}$}}>0
\end{equation*}
which is equivalent to the well known BFs (Kass and Raftery, 1995) as
\begin{equation*}
BF_{10}=\frac{p(\mathbf{y}|H_1)}{p(\mathbf{y}|H_0)}=\frac{\int p(\mathbf{y},
\mathbf{h},{\mbox{\boldmath${\vartheta}$}})\mbox{d}\mathbf{h}\mbox{d}{
\mbox{\boldmath${\vartheta}$}}}{\int p(\mathbf{y},\mathbf{h},{
\mbox{\boldmath${{\psi}}$}}|{\mbox{\boldmath${\theta}$}}_0)\mbox{d}\mathbf{h}
\mbox{d}{\mbox{\boldmath${{\psi}}$}}}>1
\end{equation*}
when rejecting the null hypothesis. In practice, the BFs are often served as
the gold statistics for hypothesis testing and the benchmark for the other
test statistics. However, the BFs have some theoretical and computational
difficulties. First, in the literature, it is well documented that it can
not be well defined when using improper priors and suffers from the
notorious Jeffreys-Lindley's paradox, see Poirier (1995), Robert (2001), Li
and Yu (2012), Li, Zeng and Yu (2014), etc. Second, the computation of $
BF_{10}$ requires to evaluate the marginal likelihood $p(\mathbf{y}|H_i),
i=0,1$. Clearly, for latent variable models, this often involves a
marginalization over the unknown latent variables $\mathbf{h}$ and the
parameter ${\mbox{\boldmath${\vartheta}$}}$. Furthermore, it is often a
high-dimensional integration and generally hard to do in practice although
there have been several interesting methods proposed in the literature for
computing BFs from the MCMC output; see, for example, Chib (1995), and Chib
and Jeliazkov (2001).
\end{remark}
\begin{remark}
Under decision theory framework, several papers have explored some effective
approaches to replace the BFs for point-null hypothesis testing. Poirier
(1997) developed a loss function approach for hypothesis testing for models
without latent variables. Bernardo and Rueda (2002) proposed an intrinsic
statistic for Bayesian hypothesis test based on the Kullback-Leibler (KL)
loss function. However, the analytical expression of the KL loss function
required by BR is not always available, especially for latent variable
models. Furthermore, the test statistic is not a pivotal quantity so that BR
had to use subjective threshold values for hypothesis testing. To deal with
latent variable models, Li and Yu (2012) proposed a Bayesian test statistic
based on the Q-function loss function within EM algorithm. LY showed that
the test statistic is well-defined under improper priors and easy to compute
for latent variable models. However, like the test statistic proposed by BR,
the test statistic proposed by LY is not pivotal. Moreover, it is not clear
if the test statistic of LY can resolve Jeffreys-Lindley's paradox. Li, Zeng
and Yu (2014) proposed another test statistic, which is a Bayesian version
of likelihood ratio test statistic. This test statistic is well-defined
under improper priors, free of Jeffreys-Lindley's paradox, and not difficult
to compute. Moreover, its asymptotic distribution can be derived and one may
obtain the threshold values from the asymptotic distribution. Unfortunately,
in general the asymptotic distribution depends on some unknown population
parameters and hence the test is not pivotal.
\end{remark}
\begin{remark}
In a recent paper, Li, Liu and Yu (2015) proposed a new Bayesian test
statistic with the following quadratic loss function
\begin{equation*}
\Delta l\left( H_{0},{\mbox{\boldmath${\theta ,\psi }$}}\right) =\left( {
\mbox{\boldmath${\theta }$}}-\bar{{\mbox{\boldmath${\theta
}$}}}\right) ^{\prime }C_{\theta \theta }\left( \bar{{\mbox{\boldmath${
\vartheta }$}}}_{0}\right) \left( {\mbox{\boldmath${\theta }$}}-\bar{{
\mbox{\boldmath${\theta }$}}}\right) ,
\end{equation*}
where $\bar{{\mbox{\boldmath${\vartheta }$}}}_{0}=\left( {
\mbox{\boldmath${\theta }$}}_{0},\bar{{\mbox{\boldmath${\psi }$}}}
_{0}\right) $ is the posterior mean under the null and $C_{\theta \theta
}\left( {\mbox{\boldmath${\vartheta }$}}\right) $ is the submatrix of $
C\left( {\mbox{\boldmath${\vartheta }$}}\right) =\left\{ \frac{\partial \log
p\left( \mathbf{y},{\mbox{\boldmath${\vartheta}$}} \right) }{\partial {
\mbox{\boldmath${\vartheta }$}}}\right\} \left\{ \frac{\partial \log p\left(
\mathbf{y},{\mbox{\boldmath${\vartheta}$}}\right) }{\partial {
\mbox{\boldmath${\vartheta }$}}}\right\} ^{\prime }$ with respect to
parameters ${\mbox{\boldmath${\theta }$}}$. With this loss function, they
showed that under some mild regularity conditions, the proposed Bayesian
test statistics followed a pivotal $\chi _{p}^{2}$ asymptotically, hence, it
is very easy to calibrate threshold values. Furthermore, this proposed test
statistic shared some nice properties with Li and Yu (2012), Li,Zeng and Yu
(2014), that is, this test statistic is well-defined under improper prior
and immune to Jefferys-Lindley's paradox. As to latent variable models,
obviously, the test statistic by Li, Liu and Yu (2015) needs to evaluate the
first-derivative of the observed likelihood function. As noted in section 2,
the observed likelihood function often generally doesn't have analytical
form so that it is not easy to do. Li, Liu and Yu (2015) showed that some
complex simulation algorithms such as EM algorithm, Kalman filter, Particle
filter have to be applied for evaluating the first derivative. \textbf{
Further, the standard error of the new statistic will be smaller than the
one in LLY.}
\end{remark}
\subsection{A new Bayesian $\protect\chi^2$ test from decision theory}
In this subsection, as to latent variable models, based on the decision
theory, we develop a new Bayesian $\chi^2$ test statistic for hypothesis
testing. The new test statistic can share the nice advantages with Li, Liu
and Yu (2015). For example, it can be well-defined under improper prior
distributions and avoids Jeffrey-Lindley's paradox. Furthermore, the
threshold values can be easily calibrated from the pivotal asymptotic
distribution and it's statistical error can be derived using MCMC approach.
Most importantly, the new test statistic can achieve other important
advantages over the existing approaches, such as, Li, et al (2015). Our new
contributions are twofold. As to latent variable models, it can be shown
that the new test statistic is only the by-product of the posterior outputs,
hence, very easy to compute. In addition, when the prior information is
available, we establish the finite sample theory.
As to any ${\mbox{\boldmath${\tilde\vartheta}$}}$ in support space of ${
\mbox{\boldmath${\vartheta}$}}$, let
\begin{equation*}
\mathbf{V}({\mbox{\boldmath${\tilde\vartheta}$}})=E\left[({
\mbox{\boldmath${\vartheta}$}}-{\mbox{\boldmath${\tilde\vartheta}$}})({
\mbox{\boldmath${\vartheta}$}}-{\mbox{\boldmath${\tilde\vartheta}$}}
)^{\prime}|\mathbf{y},H_1\right]=\int ({\mbox{\boldmath${\vartheta}$}}-{
\mbox{\boldmath${\tilde\vartheta}$}})({\mbox{\boldmath${\vartheta}$}}-{
\mbox{\boldmath${\tilde\vartheta}$}})^{\prime}p({\mbox{\boldmath${
\vartheta}$}}|\mathbf{y})\mbox{d}{\mbox{\boldmath${\vartheta}$}}
\end{equation*}
In this paper, under the statistical decision theory, we propose the
following net loss function for hypothesis testing
\begin{equation*}
\triangle \mathcal{L}[H_{0},({\mbox{\boldmath${\theta}$}},{
\mbox{\boldmath${\psi}$}})]=\left({\mbox{\boldmath${\theta}$}}-{
\mbox{\boldmath${\theta}$}}_{0}\right)^{\prime}\left[\mathbf{V}
_{\theta\theta}({\mbox{\boldmath${\bar\vartheta}$}})\right]^{-1}\left({
\mbox{\boldmath${\theta}$}}-{\mbox{\boldmath${\theta}$}}_{0}\right)
\end{equation*}
where $\mathbf{V}_{\theta\theta}({\mbox{\boldmath${\bar\vartheta}$}})$ is
the submatrix of $\mathbf{V}({\mbox{\boldmath${\bar\vartheta}$}})$
corresponding to ${\mbox{\boldmath${\theta}$}}$, $\left[\mathbf{V}
_{\theta\theta}({\mbox{\boldmath${\bar\vartheta}$}})\right]^{-1}$ is the
inverse matrix of $\mathbf{V}_{\theta\theta}({\mbox{\boldmath${\bar
\vartheta}$}})$ and ${\mbox{\boldmath${\bar\vartheta}$}}$ is the posterior
mean of ${\mbox{\boldmath${\vartheta}$}}$ under the alternative hypothesis $
H_1$. Then, we can define a Bayesian test statistic as follows:
\begin{equation}
\mathbf{T}(\mathbf{y},{\mbox{\boldmath${\theta}$}}_0)=\int \triangle
\mathcal{L}[H_{0},({\mbox{\boldmath${\theta}$}},{\mbox{\boldmath${\psi}$}}
)]p({\mbox{\boldmath${\vartheta}$}}|\mathbf{y})\mbox{d}{\mbox{\boldmath${
\vartheta}$}}=\int \left({\mbox{\boldmath${\theta}$}}-{\mbox{\boldmath${
\theta}$}}_{0}\right)^{\prime}\left[\mathbf{V}_{\theta\theta}({
\mbox{\boldmath${\bar\vartheta}$}})\right]^{-1}({\mbox{\boldmath${\bar
\vartheta}$}}) \left({\mbox{\boldmath${\theta}$}}-{\mbox{\boldmath${\theta}$}
}_{0}\right)p({\mbox{\boldmath${\vartheta}$}}|\mathbf{y})\mbox{d}{
\mbox{\boldmath${\vartheta}$}}
\end{equation}
\begin{remark}
When informative priors are not available, an objective prior or default
prior may be used. Often, $p ({\mbox{\boldmath${\theta}$}})$ is taken as
uninformative priors, such as Jeffreys or the reference prior (Jeffreys,
1961; Berger and Bernardo, 1992). These priors are generally improper, and
it follows that $p({\mbox{\boldmath${\vartheta}$}})=Af({\mbox{\boldmath${
\vartheta}$}})$ where $f({\mbox{\boldmath${\vartheta}$}})$ is a
nonintegrable function, and $A$ is an arbitrary positive constant. Since the
posterior distribution $p({\mbox{\boldmath${\vartheta}$}}|\mathbf{y})$ is
independent of an arbitrary constant in the prior distributions, and $
\mathbf{V}_{\theta\theta}({\mbox{\boldmath${\bar\vartheta}$}})$ is the
posterior covariance matrix of the interest parameter ${\mbox{\boldmath${
\theta}$}}$, hence, the statistic is independent of an arbitrary constant.
Consequently, our proposed test statistic $\mathbf{T}({\mbox{\boldmath${
\mathbf{y}}$}},{\mbox{\boldmath${\theta_0}$}})$ is independent on this
arbitrary positive constant and can be well-defined under improper priors.
\end{remark}
\begin{remark}
To see how the new statistic can avoid Jeffreys-Lindley's paradox, consider
the example discussed in Li, et al (2015). Let $y_1,y_2,\cdots,y_n\sim
N(\theta ,\sigma ^{2})$ with a known $\sigma ^{2}$ and we test the null
hypothesis $H_{0}:\theta =0$. Let the prior distribution of $\theta $ be $
N(\mu ,\tau ^{2})$. The prior distribution of $\theta $ can be set as $
N(\mu_0,\tau ^{2})$ with $\mu_0=0$. Suppose $\mathbf{y}=(y_{1},...,y_{n}),
\bar{y}=\frac{1}{n}\sum_{i=1}^{n}y_{i}$. We want to test the simple point
null hypothesis $H_{0}:\theta =0$. The posterior distribution of $\theta $
is $N(\mu(\mathbf{y}),\omega ^{2})$ with
\begin{eqnarray*}
\mu(\mathbf{y})=\frac{n\tau ^{2}\bar{y}}{\sigma ^{2}+n\tau ^{2}},\omega ^{2}=
\frac{\sigma ^{2}\tau ^{2}}{\sigma ^{2}+n\tau ^{2}},
\end{eqnarray*}
It can be shown that
\begin{eqnarray*}
& &2\log BF_{10}=\frac{n\tau^2}{n\tau^2+\sigma^2}\frac{n\bar{y}^2}{\sigma^2}
+\log\frac{\sigma^2} {n\tau ^{2}+\sigma^2} \\
& &\mathbf{T}(\mathbf{y}, \theta_{0})=\frac{n\tau^2}{n\tau^2+\sigma^2}\frac{n
\bar{y}^2}{\sigma^2}+1
\end{eqnarray*}
Clearly, when the prior information is very uninformative, as $
\tau^2\rightarrow +\infty$, we can get that $\log BF_{10}\rightarrow-\infty$
which means that the BFs always support the null hypothesis. This is
well-known as Jeffreys-Lindley's paradox in the Bayesian literature.
However, we can find that $\mathbf{T}(\mathbf{y},\theta_0)\rightarrow \frac{n
\bar{y}^2}{\sigma^2}+1$ as $\tau^2\rightarrow +\infty$. Hence, $\mathbf{T}(
\mathbf{y},\theta_0)$ is distributed as $\chi ^{2}(1)+1$ when $H_{0}$ is
true. Consequently, our proposed test statistic is immune to
Jeffreys-Lindley's paradox.
\end{remark}
\begin{remark}
The implementation of the Bayesian test statistic by Li,et al (2015)
requires the evaluation of the first derivative of the observed
log-likelihood function. As described in section 2, for latent variable
models, the observed likelihood function generally doesn't have analytical
form so that it is generally hard to get the fist derivative. Compared with
Li,et al (2015), the main advantage of the proposed test statistic in this
paper is that it is not highly computational intensive. From the equation
(3), we can easily observe that it doesn't require to evaluate the first
derivatives. From the computational perspective, our test statistic is only
involved of the posterior random samples and the inverse of the posterior
covariance matrix. In practice, through the latent variable $\mathbf{z}$ or
parameter ${\mbox{\boldmath${\vartheta}$}}$ may be high-dimensional, in
manly latent variable models, the interest parameter ${\mbox{\boldmath${
\theta}$}}$ is often low-dimensional. Hence, the proposed Bayesian test
statistic is only by-product of Bayesian posterior output, not requires
additional computational efforts. This is especially advantageous for latent
variable models.
\end{remark}
\subsection{Large sample theory for the Bayesian test statistic}
In this subsection, we establish the Bayesian large sample theory for the
proposed test statistic. Let $\left\{ z_{t}\right\} $ be a sequence of
random vectors defined on the probability space $(\Omega ,\mathcal{F},P)$
and $z^{t}$ be the collection of $\left( z_{1},z_{2},\ldots ,z_{t}\right) $.
Let $y_{t}$ denote an element of $z_{t}$ and write $z_{t}$ as $
(y_{t},w_{t}^{\prime })^{\prime }$, then we can write the conditional
likelihood function for $y_{t}$ as $f_{t}\left( y_{t}|x_{t},\vartheta
\right) $, where $x_{t}$ include some elements of $w_{t}$ and $z^{t-1}$.
Define $g_{t}\left( \vartheta \right) =g_{t}\left( z^{t},\vartheta \right)
=\log f_{t}\left( y_{t}|x_{t},\vartheta \right) $ to be the conditional
likelihood for $t$ observation and $\nabla ^{j}g_{t}\left( \vartheta \right)
$ as the $jth$ derivative of $g_{t}\left( \vartheta \right) $, we suppress
the subscript when $j=1$. The logarithm of posterior likelihood function is
\begin{equation*}
\mathcal{L}_{n}({\mbox{\boldmath${\vartheta}$}})=\log p({\mbox{\boldmath${
\vartheta}$}}|\mathbf{y}).
\end{equation*}
Furthermore, let $\dot{\mathcal{L}}_{n}({\mbox{\boldmath${\vartheta}$}}
)=\partial \log p({\mbox{\boldmath${\vartheta}$}}|\mathbf{y})/\partial {
\mbox{\boldmath${\vartheta}$}}$, $\ddot{\mathcal{L}}_{n}({
\mbox{\boldmath${\vartheta}$}})=\partial ^{2}\log p({\mbox{\boldmath${
\vartheta}$}}|\mathbf{y})/\partial {\mbox{\boldmath${\vartheta}$}}{\partial
\mbox{\boldmath${\vartheta}$}}^{\prime }$ and the negative Hessian matrix as
\begin{equation*}
\mathbf{I}({\mbox{\boldmath${\vartheta}$}})=-\frac{\partial ^{2}\log p(
\mathbf{y}|{\mbox{\boldmath${\vartheta}$}})}{\partial {\mbox{\boldmath${
\vartheta}$}}\partial {\mbox{\boldmath${\vartheta}$}}^{\prime }}.
\end{equation*}
Let the prior density to be $p({\mbox{\boldmath${\vartheta}$}})$, $\gamma({
\mbox{\boldmath${\vartheta}$}})=\log p({\mbox{\boldmath${\vartheta}$}})$ and
$\gamma^{{\mbox{\boldmath${\vartheta}$}}}({\mbox{\boldmath${\vartheta}$}}
)=\partial\log p({\mbox{\boldmath${\vartheta}$}})/\partial{
\mbox{\boldmath${\vartheta}$}}$. In order to derive the asymptotic
distribution of the proposed test statistic, following LZY (2014) and
LLY(2015), a set of regularity conditions are imposed in the following.
\begin{assumption}
There exists a finite sample size $n^{\ast }$, so that, for $n>n^{\ast }$,
there is a local maximum at ${\mbox{\boldmath${\widehat\vartheta_m}$}}$
(i.e., posterior mode) such that $\dot{\mathcal{L}}_{n}({\mbox{\boldmath${
\widehat\vartheta_m}$}})=0$ and $\ddot{\mathcal{L}}_{n}({\mbox{\boldmath${
\widehat\vartheta_m}$}})$ is negative definite.
\end{assumption}
\begin{assumption}
The largest eigenvalue $\lambda _{n}$ of $-\ddot{\mathcal{L}}_{n}^{-2}(
\widehat{\mbox{\boldmath${\vartheta}$}}_{m})$ goes to zero in probability as
$n\rightarrow \infty $.
\end{assumption}
\begin{assumption}
For any $\varepsilon >0$, there exists a positive number $\delta $, such that
\begin{equation}
\lim_{n\rightarrow \infty }P\left[ \sup_{{\mbox{\boldmath${\vartheta}$}} \in
B\left( {\mbox{\boldmath${\widehat\vartheta_m}$}},\text{ }\delta \right)
}\left\Vert \ddot{\mathcal{L}}^{-1}\left( {\mbox{\boldmath${\widehat
\vartheta_m}$}}\right) \left[ \ddot{\mathcal{L}}\left( {\mbox{\boldmath${
\widehat\vartheta}$}} \right) -\ddot{\mathcal{L}}\left( {\mbox{\boldmath${
\widehat\vartheta_m}$}}\right) \right] \right\Vert <\varepsilon \right] =1 .
\label{A3}
\end{equation}
\end{assumption}
\begin{assumption}
For any $\delta>0$,
\begin{equation*}
\int_{{\mbox{\boldmath${\Omega}$}}-B(\widehat{\mbox{\boldmath${\vartheta}$}}
_{m},\delta )}p({\mbox{\boldmath${\vartheta}$}}|\mathbf{y})d{{
\mbox{\boldmath${\vartheta}$}}}\rightarrow 0,
\end{equation*}
in probability as $n\rightarrow \infty $, where $\mathbf{\Omega }$ is the
support space of ${\mbox{\boldmath${\vartheta}$}}$.
\end{assumption}
\begin{assumption}
For any $\delta>0$,
\begin{equation*}
\int_{{\mbox{\boldmath${\Omega}$}}-B(\widehat{\mbox{\boldmath${\vartheta}$}}
_{m},\delta )}\left\Vert {\mbox{\boldmath${\vartheta}$}} \right\Vert ^{2}p({
\mbox{\boldmath${\vartheta}$}}|\mathbf{y})d{{\mbox{\boldmath${\vartheta}$}}}
= O_p(n^{-3}),
\end{equation*}
as $n\rightarrow \infty $, where $\mathbf{\Omega }$ is the support space of $
{\mbox{\boldmath${\vartheta}$}}$.
\end{assumption}
\begin{assumption}
Let $\mathcal{\vartheta }_{0}$ to be true value, $\mathcal{\vartheta }
_{0}\in int\left( \Theta \right) $ where $\Theta $ is a compact, separable
metric space.
\end{assumption}
\begin{assumption}
$\left\{w_{t},t=1,2,3,\ldots \right\} $ is an $\alpha$ mixing sequence that
satisfies, for $\mathcal{F}_{-\infty }^{t}=\sigma \left(
z_{t},z_{t-1},\ldots \right) $ and $\mathcal{F}_{t+m}^{\infty }=\sigma
\left( z_{t+m},z_{t+m+1},\ldots \right) $, the mixing coefficient $\alpha
\left( m\right) =O\left( m^{\frac{-r}{r-2}-\varepsilon }\right) $ for some $
\varepsilon >0$ and $r>2$.
\end{assumption}
\begin{assumption}
Let $N_{\delta }\left( \mathcal{\vartheta }_{\ast }\right) =\left\{ \mathcal{
\vartheta \in }\Theta :\left\Vert \mathcal{\vartheta -\vartheta }_{\ast
}\right\Vert \leq \delta \right\} $ for $\mathcal{\vartheta }_{\ast }\in
\Theta $, $\delta \geq 0$ and $0\leq j \leq s_{1}$, (i) $\sup_{\mathcal{
\vartheta \in }N_{\delta }\left( \mathcal{\vartheta }_{\ast }\right)
}\nabla^{j}g_{t}\left( \mathcal{\vartheta }\right) $ and $\inf_{\mathcal{
\vartheta \in }N_{\delta }\left( \mathcal{\vartheta }_{\ast }\right)
}\nabla^{j}g_{t}\left( \mathcal{\vartheta }\right) $ are measurable to $
\mathcal{F}_{-\infty }^{t}$ and strictly stationary; (ii) $E\left[ \sup_{
\mathcal{\vartheta \in }N_{\delta }\left( \mathcal{\vartheta }_{\ast
}\right) }\nabla^{j}g_{t}\left( \mathcal{\vartheta }\right) \right] <\infty $
and $E\left[ \inf_{\mathcal{\vartheta \in }N_{\delta }\left( \mathcal{
\vartheta }_{\ast }\right) }\nabla^{j}g_{t}\left( \mathcal{\vartheta }
\right) \right] >-\infty $; (iii) $\lim_{\delta \downarrow 0}E\left[ \sup_{
\mathcal{\vartheta \in }N_{\delta }\left( \mathcal{\vartheta }_{\ast
}\right) }\nabla^{j}g_{t}\left( \mathcal{\vartheta }\right) \right]
=\lim_{\delta \downarrow 0}E\left[ \inf_{\mathcal{\vartheta \in }N_{\delta
}\left( \mathcal{\vartheta }_{\ast }\right) }\nabla^{j}g_{t}\left( \mathcal{
\vartheta }\right) \right] =E\left[ \nabla^{j}g_{t}\left( \mathcal{\vartheta
}_{\ast }\right) \right] $.
\end{assumption}
\begin{assumption}
There exists a function $M_{t}(\omega _{t})$ such that for $0\leqslant j
\leqslant s_{2} $, all $\theta \in \mathcal{G}$ where $\mathcal{G}$ is an
open, convex set containing $\Theta $, $\bigtriangledown ^{j}g_{t}\left(
\vartheta \right) $ exists, $\sup_{t}\left\Vert M_{t}(\omega
_{t})\right\Vert ^{r+\delta }\leq M<\infty $ for some $\delta >0$.
\end{assumption}
\begin{assumption}
The prior density is continuous and $0<p(\vartheta)<\infty$ for all $
\vartheta\in\Theta$.
\end{assumption}
\begin{assumption}
For $0<j<s_{3}$, $E\left\Vert\bigtriangledown ^{j}\gamma\left( \vartheta
\right)\right\Vert=O(1) $ .
\end{assumption}
\begin{remark}
Regularity Assumptions 1-4 have been used to develop the Bayesian large
sample theory. This theory is proved by chen(1985), which states that the
posterior distribution is degenerate about the posterior mode and
asymptotically normal after suitable scaling, that is,
\begin{eqnarray*}
\left({\mbox{\boldmath${\vartheta}$}}-\widehat{\mbox{\boldmath${\vartheta}$}}
_{m}\right)|\mathbf{y}\overset{d}{\longrightarrow} N\left[0, -\ddot{\mathcal{
L}}_{n}^{-1}(\widehat{\mbox{\boldmath${\vartheta}$}}_{m})\right]
\end{eqnarray*}
More details, one can refer to Chen (1985). Bickel and Doksum (2006), and Le
Cam and Yang (2000), Ghosh (2003) presented another version of this theorem
on the basis of other similar regularity conditions. The main difference is
that the value at which the asymptotic posterior variance matrix is
evaluated. It is the posterior mode $\widehat{\mbox{\boldmath${\vartheta}$}}
_{m}$ in Chen (1985), the true value $\widehat{\mbox{\boldmath${\vartheta}$}}
_{0}$ in Bickel and Doksum (2006), and Le Cam and Yang (2000), Ghosh (2003)
and the MLE estimator $\widehat{{\mbox{\boldmath${\vartheta}$}}}$ in Kim
(1994) depending on different assumptions.
\end{remark}
\begin{remark}
Under Assumptions 1-4, conditional on the observed data $\mathbf{y}$, it can
be shown that
\begin{eqnarray*}
&&\bar{{\mbox{\boldmath${\vartheta}$}}}=E\left[ {\mbox{\boldmath${
\vartheta}$}}|\mathbf{y},H_1\right] =\widehat{\mbox{\boldmath${\vartheta}$}}
_{m}+o_p(n^{-1/2}), \\
&&\mathbf{V}\left(\widehat{\mbox{\boldmath${\vartheta}$}}_{m}\right) =E\left[
\left({\mbox{\boldmath${\vartheta}$}}-\widehat{\mbox{\boldmath${\vartheta}$}}
_{m}\right) \left({\mbox{\boldmath${\vartheta}$}}-\widehat{
\mbox{\boldmath${\vartheta}$}}_{m}\right) ^{^{\prime }}|\mathbf{y},H_1\right]
=-\ddot{\mathcal{L}}^{-1}_{n}\left(\widehat{{\mbox{\boldmath${\vartheta}$}}}
_{m}\right) +o_p(n^{-1}),
\end{eqnarray*}
where $\bar{{\mbox{\boldmath${\vartheta}$}}}$ is the posterior mean. These
conclusions have been given by Li,Zeng and Yu (2014).
\end{remark}
\begin{remark}
Following Rilstone, Srivatsava and Ullah(1996), Bester and Hansen (2006),
the assumptions 5-10 are used to justify the validity of high order Laplace
expansion. The assumption 5 is analogous to the analytical assumptions for
Laplace's method (kass et al., 1990), but we impose the higher order
constraints $O_p(n^{-3})$ other than $O_p(n^{-2})$, see also Miyata (2004,
2010). With these assumptions, we can get the standard form higher order
Laplace expansion of the order $O_p(n^{-2})$ in kass et al. (1990) to $
O_p(n^{-3})$, similar to the fully exponential form in Miyata (2004, 2010).
\end{remark}
Let $\widehat{{\mbox{\boldmath${\vartheta}$}}}$ to be the maximum likelihood
estimator of ${\mbox{\boldmath${\vartheta}$}}$ and $\widehat{{
\mbox{\boldmath${\theta}$}}}$ is the subvector of $\widehat{{
\mbox{\boldmath${\vartheta}$}}}$ corresponding to ${{\mbox{\boldmath${
\theta}$}}}$, under Assumptions 5-8 with $s_{1}=2$ and $s_{2}=2$, the Wald
statistic be
\begin{equation*}
\mathbf{Wald}=\left(\widehat{{\mbox{\boldmath${\theta}$}}}-{{
\mbox{\boldmath${\theta}$}}}_0\right)^{\prime}\left[-\ddot{\mathcal{L}}_{n,\theta\theta}^{-1}(\widehat{{\mbox{\boldmath${\vartheta}$}}})\right]
^{-1} \left(\widehat{{\mbox{\boldmath${\theta}$}}}-{{\mbox{\boldmath${
\theta}$}}}_0\right),
\end{equation*}
where $\ddot{\mathcal{L}}_{n,\theta\theta}^{-1}(\widehat{{\mbox{\boldmath${\vartheta}$}}})$ is the submatrix of $\ddot{\mathcal{L}}_{n}^{-1}(\widehat{{\mbox{\boldmath${\vartheta}$}}})$
corresponding to ${\mbox{\boldmath${\theta}$}}$.
\begin{theorem}
\label{thm1} Under Assumptions 6-11 with $s_{1}=s_{2}=s_{3}=3$, when the
likelihood dominates the prior such as $p({\mbox{\boldmath${\vartheta}$}}
)=O_p(1)$, under the null hypothesis, we can show that
\begin{eqnarray}
&&\mathbf{T}({\mbox{\boldmath${\mathbf{y}}$}},{\mbox{\boldmath${\theta}$}}
_{0})-p=\mathbf{Wald}+o_{p}(1)\overset{d}{\rightarrow}\chi^2(p)
\end{eqnarray}
\end{theorem}
\begin{remark}
In Theorem \ref{thm1}, we can see that under the null hypothesis, the
asymptotic distribution of $\mathbf{T}(\mathbf{y},{\mbox{\boldmath${\theta}$}
}_{0})$ always follows the $\chi ^{2}$ distribution, hence, is pivotal. As
to the proposed test statistic, we still need to specify some threshold
values, $c$ for implementing the test, that is,
\begin{equation*}
\text{Accept}~~H_{0}\text{ if }\mathbf{T}(\mathbf{y},{\mbox{\boldmath${
\theta}$}}_{0})\leq c;\text{ Reject}~~H_{0}\ \text{ if }\mathbf{T}(\mathbf{y}
,{\mbox{\boldmath${\theta}$}}_{0})>c.
\end{equation*}
Hence, this asymptotic $\chi^2$ distribution can be utilized conveniently to
calibrate threshold values.
\end{remark}
\begin{remark}
From this theorem, $\mathbf{T}(\mathbf{y},{\mbox{\boldmath${\theta}$}}_{0})$
may be regarded as the Bayesian version of the $\mathbf{Wald}$ statistic.
However, the $\mathbf{Wald}$ test statistic is a frequentist test which is
based on the maximum likelihood estimation of the model in the alternative
hypothesis whereas our test is a Bayesian test which is based on the
posterior quantities of the models under the alternative hypothesis.
\end{remark}
\begin{remark}
The implementation of the $\mathbf{Wald}$ test requires the ML estimation of
the model and evaluation of the second derivative of the observed likelihood
function under the alternative hypothesis. As described in section 2, for
latent variable models, the observed likelihood function generally doesn't
have analytical form so that it is generally hard to get the maximum
likelihood estimator and its corresponding second derivative. Hence, it is
difficult to apply the $\mathbf{Wald}$ test statistic for hypothesis
testing. However, our proposed Bayesian test statistic is only by-product of
posterior outputs. As long as the Bayesian MCMC methods are applicable, our
test can be implemented for latent variable models. In addition, from
equation (3), $\mathbf{T}({\mbox{\boldmath${\mathbf{y}}$}},{
\mbox{\boldmath${\theta_0}$}})$ can incorporate the prior information
through the posterior distribution directly, but $\mathbf{Wald}$ can not
incorporate the useful prior information.
\end{remark}
\begin{remark}
We use a simple example to illustrate the influence of the prior
distributions. Let $y_{1},...,y_{n}\sim N(\theta ,\sigma ^{2})$ with a known
variance $\sigma ^{2}=1$. The true value of $\theta $ is set at $\theta
_{0}=0.10$. The prior distribution of $\theta $ is set as $N(\mu _{0},\tau
^{2})$. The simple point null hypothesis is $H_{0}:\theta =0$. It can be
shown that
\begin{eqnarray*}
&&2\log BF_{10}=\frac{\sigma ^{2}\tau^2}{\sigma ^{2}+n\tau ^{2}}\left(\frac{n
\bar{y}}{\sigma^{2}}+\frac{\mu_0}{\tau^2}\right)^2+\log \frac{\sigma ^{2}}{
\sigma ^{2}+n\tau ^{2}}, \\
&&\mathbf{T}(\mathbf{y},\theta _{0})=\frac{\sigma ^{2}\tau^2}{\sigma
^{2}+n\tau ^{2}}\left(\frac{n\bar{y}}{\sigma^{2}}+\frac{\mu_0}{\tau^2}
\right)^2+1,\mathbf{Wald}=\frac{n\bar{y}^{2}}{\sigma ^{2}},
\end{eqnarray*}
where $\bar{y}=\frac{1}{n}\sum_{i=1}^{n}y_{i}$. When $n\longrightarrow
\infty $, $\mathbf{T}(\mathbf{y},\theta _{0})-1\longrightarrow $ $\mathbf{
Wald}$ and the asymptotic distribution for both $\mathbf{T}(\mathbf{y}
,\theta _{0})-1$ and $\mathbf{Wald}$ is $\chi ^{2}(1)$. Let us consider the
case that corresponds to an informative prior $N(0.10,10^{-3})$ and compare
it to the case that corresponds to a non-informative prior $N(0,10^{50})$.
Table 1 reports $2\log BF_{10}$, $\mathbf{T}(\mathbf{y},\theta _{0})$, and $
\mathbf{Wald}$ when $n=10,100,1000,10000$ under these two priors. It can be
seen that both the BF and the new test depend on the prior (although the BFs
tend to choose the wrong model under the vague prior even when the sample
size is very large) while the $\mathbf{Wald}$ test is independent of the
prior. When $n=10,100$, $\mathbf{T}(\mathbf{y},\theta _{0}) $ correctly
rejects the null hypothesis when the prior is informative but fails to
reject it when the prior is vague under the significant level 5\%. In this
case, the $\mathbf{Wald}$ test fails to reject the null hypothesis under
both priors.\footnote{
To implement the $\mathbf{Wald}$ test,we use the following Fisher's scale.
Let $\alpha $ be the critical level and $P=1-\alpha $. If $P$ is between
95\% and 97.5\%, the evidence for the alternative is \textquotedblleft
moderate\textquotedblright ; between 97.5\% and 99\%, \textquotedblleft
substantial\textquotedblright ; between 99\% and 99.5\%, \textquotedblleft
strong\textquotedblright ; between 99.5\% and 99.9\%, \textquotedblleft very
strong\textquotedblright ; larger than 99.9\%, \textquotedblleft
overwhelming\textquotedblright . To implement the BF we use Jeffreys' scale
instead. If $\log BF_{10}$ is less than 0, there is \textquotedblleft
negative\textquotedblright\ evidence for the alternative; between 0 and 1,
\textquotedblleft not worth more than a bare mention\textquotedblright ;
between 1 and 3, \textquotedblleft positive\textquotedblright ; between 3
and 5, \textquotedblleft strong\textquotedblright ; larger than 5,
\textquotedblleft very strong\textquotedblright .}
\begin{table}[tbph]
\caption{Comparison of $2\log BF_{10}$, $\mathbf{T}(\mathbf{y},\protect
\theta _{0})$, and $\mathbf{Wald}$}
\begin{center}
\vspace{1em}
\begin{tabular}{c|cccc|cccc}
\hline
Prior & \multicolumn{4}{|c|}{$N(0.10,10^{-3})$} & \multicolumn{4}{c}{$
N(0,10^{50})$} \\ \hline
$n$ & 10 & 100 & 1000 & 10000 & 10 & 100 & 1000 & 10000 \\ \hline
$2\log BF_{10}$ & 9.96 & 11.12 & 20.60 & 93.58 & -117.42 & -118.50 & -110.72
& -38.00 \\ \hline
$\mathbf{T}(\mathbf{y},\theta _{0})$ & 10.96 & 12.22 & 22.30 & 96.98 & 1.01
& 2.23 & 12.32 & 87.03 \\ \hline
$\mathbf{Wald}$ & 0.01 & 1.23 & 11.32 & 86.03 & 0.01 & 1.23 & 11.32 & 86.03
\\ \hline
\end{tabular}
\end{center}
\end{table}
\end{remark}
\begin{remark}
Under the null hypothesis, our statistic can be written as
\begin{equation*}
\mathbf{T}(\mathbf{y},\theta _{0})=1+n\frac{\overline{y}^{2}}{\sigma ^{2}}+2
\overline{y}\frac{\mu _{0}}{\tau ^{2}}-\overline{y}^{2}\frac{1}{\tau ^{2}}+
\frac{1}{n}\left( \frac{\mu _{0}}{\tau ^{2}}\right)
^{2}\sigma^{2}+O_{p}\left( n^{-3/2}\right)
\end{equation*}
where $2\overline{y}\frac{\mu _{0}}{\tau ^{2}}$ is the order of $n^{-1/2}$,
and $-\overline{y}^{2}\frac{1}{\tau ^{2}}+\frac{1}{n}\left( \frac{\mu _{0}}{
\tau ^{2}}\right) ^{2}\sigma^{2}$ has the order $n^{-1}$.
\end{remark}
Since $\mathbf{T}({\mbox{\boldmath${\mathbf{y}}$}},{
\mbox{\boldmath${
\theta_0}$}})$ is calculated by using the MCMC output, it is important to
assess the numerical standard error for measuring the magnitude of
simulation error.
\begin{corollary}
\label{corol1}
Given the posterior draws $\{{\mbox{\boldmath${\vartheta}$}}^{\left( j\right) },j=1,2,\cdots ,J\}$, the numerical standard error (NSE) of the statistic $\mathbf{T}(\mathbf{y},\theta _{0})$ is,
\begin{equation*}
NSE\left( \widehat{\mathbf{T}}\left( \mathbf{y},{\mbox{\boldmath${\theta}$}}
_{0}\right) \right) = \sqrt{\frac{\partial \widehat{\mathbf{T}}\left( \mathbf{y},{\
\mbox{\boldmath${\theta}$}}_{0}\right) }{\partial \widehat{\mathbf{h}}}
Var\left( \widehat{{\mbox{\boldmath${h}$}}}\right) \left( \frac{\partial
\widehat{\mathbf{T}}\left( \mathbf{y},{\mbox{\boldmath${\theta}$}}
_{0}\right) }{\partial \widehat{{\mbox{\boldmath${h}$}}}}\right) ^{\prime }},
\end{equation*}
where $\frac{\partial \widehat{\mathbf{T}}\left( \mathbf{y},{
\mbox{\boldmath${
\theta}$}}_{0}\right) }{\partial \widehat{{\mbox{\boldmath${h}$}}}} =
-vec({\mbox{\boldmath${A}$}}^\prime)^{\prime}\left(\widehat{{\mbox{\boldmath${H}$}}}
^{\prime-1}\otimes\widehat{{\mbox{\boldmath${H}$}}}^{-1}\right)\frac{
\partial \widehat{{\mbox{\boldmath${H}$}}}}{\partial \widehat{{\
\mbox{\boldmath${h}$}}}}$, $\widehat{{
\mbox{\boldmath${H}$}}}=\frac{1}{J}\sum_{j=1}^{J}\left( {\
\mbox{\boldmath${\theta}$}}^{\left( j\right) }-{
\mbox{\boldmath${\bar{
\theta}}$}}\right) \left( {\mbox{\boldmath${\theta}$}}^{\left(j\right) }-{\
\mbox{\boldmath${\bar{\theta}}$}}\right) ^{\prime }$, $\mathbf{A} = \left({\mbox{\boldmath${\bar\theta-\theta_{0}}$}}\right)\left({\mbox{\boldmath${\bar\theta-\theta_{0}}$}}
\right)^\prime$, $\widehat{\mathbf{h}}=vech\left(\widehat{{\mbox{\boldmath${H}$}}}\right)$, $\frac{ \partial \widehat{{\mbox{\boldmath${H}$}}}}{\partial \widehat{{\
\mbox{\boldmath${h}$}}}} = \left( \frac{\partial vec(\widehat{{
\mbox{\boldmath${H}$}}})}{ \partial \widehat{{\mbox{\boldmath${h}$}}}}
\right)$, $Var\left({\widehat{{\mbox{\boldmath${h}$}}}}\right)$ is the NSE of $\widehat{{\mbox{\boldmath${h}$}}}$.
\end{corollary}
The Corollary \ref{corol1} shows us how to compute the numerical standard error of the proposed statistic. For the NSE of $\widehat{{\mbox{\boldmath${h}$}}}$, $Var\left({\widehat{{\mbox{\boldmath${h}$}}}}\right)$, following Newey and West (1987), a consistent estimator can be given by
\begin{equation*}
Var(\widehat{\mathbf{h}})=\frac{1}{J}\left[ \Omega _{0}+\sum_{k=1}^{q}\left(
1-\frac{k}{q+1}\right) \left( \Omega _{k}+\Omega _{k}^{\prime }\right)
\right] ,
\end{equation*}
where
\begin{equation*}
\Omega _{k}=J^{-1}\sum_{j=k+1}^{J}\left( \mathbf{h}^{\left(j\right) }-
\widehat{\mathbf{h}}\right) \left( \mathbf{h}^{\left(j\right) }-\widehat{
\mathbf{h}}\right) ^{\prime }.
\end{equation*}
and the value of $q$ is always equal to 10.
\section{The Extension of the Test}
In this section, we extend the point-null hypothesis aforementioned
into the following problem,
\begin{equation*}
\begin{cases}
H_{0}: & R{\mbox{\boldmath${\vartheta}$}}_{0}={\mbox{\boldmath${r}$}}\\
H_{1}: & R{\mbox{\boldmath${\vartheta}$}}_{0}\neq{\mbox{\boldmath${r}$}}
\end{cases},
\end{equation*}
where $R$ is a $m\times\left(d+q\right)$ matrix, ${\mbox{\boldmath${r}$}}\in\mathbb{R}^{m}$.
This hypothesis problem is much more general than the previous one.
On the other hand, it can help us to study the relationship among
parameters. Further, for such problems, it is hard to use the Bayes factor. Hence, the extension here is meaningful.
For such problem, the frequentist Wald statistic is
\begin{equation*}
\text{\textbf{Wald}}=\left(R\widehat{{\mbox{\boldmath${\vartheta}$}}}-{\mbox{\boldmath${r}$}}\right)^{\prime}\left[R\left(-\ddot{\mathcal{L}}_{n}^{-1}\left(\widehat{{\mbox{\boldmath${\vartheta}$}}}\right)\right)R^{\prime}\right]^{-1}\left(R\widehat{{\mbox{\boldmath${\vartheta}$}}}-{\mbox{\boldmath${r}$}}\right),
\end{equation*}
where $\widehat{{\mbox{\boldmath${\vartheta}$}}}$ is the MLE estimator of ${\mbox{\boldmath${\vartheta}$}}$.
According to the decision theory, we define the net loss function
for such problem as
\begin{equation*}
\Delta\mathcal{L}\left(H_{0},{\mbox{\boldmath${\vartheta}$}}\right)=\left(R{\mbox{\boldmath${\vartheta}$}}-{\mbox{\boldmath${r}$}}\right)^{\prime}\left[RV\left({\mbox{\boldmath${\bar{\vartheta}}$}}\right)R^{\prime}\right]^{-1}\left(R\mathbf{{\mbox{\boldmath${\vartheta}$}}}-{\mbox{\boldmath${r}$}}\right),
\end{equation*}
where $\bar{{\mbox{\boldmath${\vartheta}$}}}$ is the posterior mean of ${\mbox{\boldmath${\vartheta}$}}$,
$V\left(\bar{{\mbox{\boldmath${\vartheta}$}}}\right)=E\left[\left.\left({\mbox{\boldmath${\vartheta}$}}-\bar{{\mbox{\boldmath${\vartheta}$}}}\right)\left({\mbox{\boldmath${\vartheta}$}}-\bar{{\mbox{\boldmath${\vartheta}$}}}\right)^{\prime}\right|{\mbox{\boldmath${y}$}},H_{1}\right]$.
Then the statistic is defined as
\begin{equation*}
{\mbox{\boldmath${T}$}}\left({\mbox{\boldmath${y}$}},{\mbox{\boldmath${r}$}}\right)=\int\Delta\mathcal{L}\left(H_{0},{\mbox{\boldmath${\vartheta}$}}\right)d{\mbox{\boldmath${\vartheta}$}}=\int\left(R{\mbox{\boldmath${\vartheta}$}}-{\mbox{\boldmath${r}$}}\right)^{\prime}\left[RV\left({\mbox{\boldmath${\bar{\vartheta}}$}}\right)R^{\prime}\right]^{-1}\left(R\mathbf{{\mbox{\boldmath${\vartheta}$}}}-{\mbox{\boldmath${r}$}}\right)d{\mbox{\boldmath${\vartheta}$}}.
\end{equation*}
\begin{theorem}
\label{thm2}Under Assumptions \textasciitilde{}\textasciitilde{}\textasciitilde{}\textasciitilde{},
when the likelihood information dominates the prior information, under
the null hypothesis,
\begin{equation*}
{\mbox{\boldmath${T}$}}\left({\mbox{\boldmath${y}$}},{\mbox{\boldmath${r}$}}\right)-m=\text{\textbf{Wald}}+o_{p}\text{\ensuremath{\left(1\right)} \ensuremath{\overset{d}{\rightarrow}}}\chi^{2}\left(m\right).
\end{equation*}
\end{theorem}
Similarly, for the statistic, the numerical standard error can be computed in the following corollary.
\begin{corollary}
\label{corol2}
Given the posterior draws $\{{\mbox{\boldmath${\vartheta}$}}^{\left( j\right) },j=1,2,\cdots ,J\}$, the numerical standard error (NSE) of the statistic $\mathbf{T}(\mathbf{y},{\mbox{\boldmath${r}$}})$ is,
\begin{equation*}
NSE\left( \widehat{\mathbf{T}}\left( \mathbf{y},{\mbox{\boldmath${r}$}}\right) \right) = \sqrt{\frac{\partial \widehat{\mathbf{T}}\left( \mathbf{y},
{\mbox{\boldmath${r}$}}\right) }{\partial \widehat{\mathbf{h}}}
Var\left( \widehat{{\mbox{\boldmath${h}$}}}\right) \left( \frac{\partial
\widehat{\mathbf{T}}\left( \mathbf{y},{\mbox{\boldmath${r}$}}\right) }{\partial \widehat{{\mbox{\boldmath${h}$}}}}\right) ^{\prime }},
\end{equation*}
where $\frac{\partial \widehat{\mathbf{T}}\left( \mathbf{y},
{\mbox{\boldmath${r}$}}\right) }{\partial \widehat{{\mbox{\boldmath${h}$}}}} =
-vec({\mbox{\boldmath${A}$}}^\prime)^{\prime}\left[\left(R\widehat{{\mbox{\boldmath${H}$}}}R^{\prime} \right)^{-1}\otimes\left(R\widehat{{\mbox{\boldmath${H}$}}}R^{\prime} \right)^{-1}\right]\left(R\otimes R\right)\frac{
\partial \widehat{{\mbox{\boldmath${H}$}}}}{\partial \widehat{{\
\mbox{\boldmath${h}$}}}}$, $\widehat{{
\mbox{\boldmath${H}$}}}=\frac{1}{J}\sum_{j=1}^{J}\left( {\mbox{\boldmath${\vartheta}$}}^{\left( j\right) }-\bar{{\mbox{\boldmath${\vartheta}$}}}\right) \left( {\mbox{\boldmath${\vartheta}$}}^{\left( j\right) }-\bar{{\mbox{\boldmath${\vartheta}$}}}\right) ^{\prime }$, $\mathbf{A} = \left(R{\mbox{\boldmath${\bar\vartheta} - {\mbox{\boldmath${r}$}}$}}\right)\left(R{\mbox{\boldmath${\bar\vartheta}$}} - {\mbox{\boldmath${r}$}}\right)^\prime$, $\widehat{\mathbf{h}}=vech\left(\widehat{{\mbox{\boldmath${H}$}}}\right)$, $\frac{ \partial \widehat{{\mbox{\boldmath${H}$}}}}{\partial \widehat{{\
\mbox{\boldmath${h}$}}}} = \left( \frac{\partial vec(\widehat{{
\mbox{\boldmath${H}$}}})}{ \partial \widehat{{\mbox{\boldmath${h}$}}}}
\right)$, $Var\left({\widehat{{\mbox{\boldmath${h}$}}}}\right)$ is the NSE of $\widehat{{\mbox{\boldmath${h}$}}}$.
\end{corollary}
For the NSE of $\widehat{{\mbox{\boldmath${h}$}}}$, $Var\left({\widehat{{\mbox{\boldmath${h}$}}}}\right)$, we can still follow the way proposed by Newey and West (1987) to evaluated.
\section{Simulation Studies}
In this section, we do two simulation studies to check the empirical size and power of the proposed test statistic. The first example is a simple simulation examination based on linear regression model where the our proposed test statistic has analytical expression. We compare the size and power of the new statistics with the Wald statistic. In the second example, we use the stochastic volatility model with leverage effect, where Wald statistic can not be used, to study the size and power of our statistic.
\subsection{The empirical power and size of ${\mbox{\boldmath${T}$}}\left({\mbox{\boldmath${y}$}},{{\mbox{\boldmath${\beta}$}}}_{0}\right)$ and Wald statistic for linear regression model}
In this subsection, we use the simple linear regression model to examine the empirical power and size of the proposed test statistic. The model we use is
\begin{equation*}
y_{i}={\mbox{\boldmath${x}$}}_{i}^{\prime}{\mbox{\boldmath${\beta}$}}+\epsilon_{i},\epsilon_{i}\sim N\left(0,\sigma^{2}\right),i=1,\dots,n.
\end{equation*}
with ${\mbox{\boldmath${x}$}}_{i1}=1$. Let ${\mbox{\boldmath${X}$}}=\left({\mbox{\boldmath${x}$}}_{1}^{\prime},\dots,{\mbox{\boldmath${x}$}}_{N}^{\prime}\right)^{\prime}$,
then we can rewrite the model in matrix form,
\begin{equation*}
{\mbox{\boldmath${y}$}}={\mbox{\boldmath${X}$}}{\mbox{\boldmath${\beta}$}}+{\mbox{\boldmath${\epsilon}$}},
\end{equation*}
where ${\mbox{\boldmath${y}$}}=\left(y_{1},\dots,y_{n}\right)^{\prime}$, ${\mbox{\boldmath${\epsilon}$}}=\left(\epsilon_{1},\dots,\epsilon_{n}\right)^{\prime}$.
We are interested in the subvector of ${\mbox{\boldmath${\beta}$}}$, $\breve{{\mbox{\boldmath${\beta}$}}}$,
then ${\mbox{\boldmath${\beta}$}}=\left(\breve{{\mbox{\boldmath${\beta}$}}}^{\prime},\tilde{{\mbox{\boldmath${\beta}$}}}^{\prime}\right)^{\prime}$.
Here we want to test $H_{0}:\breve{{\mbox{\boldmath${\beta}$}}}=\breve{{\mbox{\boldmath${\beta}$}}}_{0}$
against $H_{1}:\breve{{\mbox{\boldmath${\beta}$}}}\neq\breve{{\mbox{\boldmath${\beta}$}}}_{0}$ and $H_{0}:R{\mbox{\boldmath${\beta}$}}={{\mbox{\boldmath${r}$}}}$
against $H_{1}:R{\mbox{\boldmath${\beta}$}}\neq{{\mbox{\boldmath${r}$}}}$.
Assume that the prior distribution for ${\mbox{\boldmath${\beta}$}}$ and $\sigma^{2}$
are normal and inverse gamma, respectively,
\begin{equation*}
{\mbox{\boldmath${\beta}$}}|\sigma^{2}\sim N\left(\mu_{0},\sigma^{2}V_{0}\right),\sigma^{2}\sim IG\left(a,b\right),
\end{equation*}
where $\mu_{0}$, $V_{0}$ and $a$, $b$ are hyperparameters.
The proposed statistic ${\mbox{\boldmath${T}$}}\left({\mbox{\boldmath${y}$}},{{\mbox{\boldmath${\beta}$}}}_{0}\right)$ for the first problem is
\begin{equation*}
{\mbox{\boldmath${T}$}}\left({\mbox{\boldmath${y}$}},\breve{{\mbox{\boldmath${\beta}$}}}_{0}\right)=p+\frac{v-2}{2s}\left(\bar{\breve{{\mbox{\boldmath${\beta}$}}}}_{H_{1}}-\breve{{\mbox{\boldmath${\beta}$}}}_{0}\right)^{\prime}\breve{V}^{*-1}\left(\bar{\breve{{\mbox{\boldmath${\beta}$}}}}_{H_{1}}-\breve{{\mbox{\boldmath${\beta}$}}}_{0}\right),
\end{equation*}
where $v=2a+n$, $s=b+\frac{1}{2}\left(\mu_{0}^{\prime}V_{0}^{-1}\mu_{0}+{\mbox{\boldmath${y}$}}^{\prime}{\mbox{\boldmath${y}$}}-\mu^{*\prime}V^{*-1}\mu^{*}\right)$,
$V^{*}=\left(V_{0}^{-1}+{\mbox{\boldmath${X}$}}^{\prime}{\mbox{\boldmath${X}$}}\right)^{-1}$,$\mu^{*}=V^{*}\left(V_{0}^{-1}\tilde{\mu}+{\mbox{\boldmath${X}$}}^{\prime}{\mbox{\boldmath${y}$}}\right)$
and $\breve{V}^{*}$ the submatrix of $V^{*}$ corresponding to $\breve{{\mbox{\boldmath${\beta}$}}}$.
$p$ is the dimension of $\breve{{\mbox{\boldmath${\beta}$}}}$ and $\bar{\breve{{\mbox{\boldmath${\beta}$}}}}_{H_{1}}$
is the posterior mean of $\breve{{\mbox{\boldmath${\beta}$}}}$ under $H_{1}$. The details is given in the Appendix \ref{smlexpl1}.
For the second hypothesis problem, it can be readily derived that
the statistic is
\begin{equation*}
{\mbox{\boldmath${T}$}}\left({\mbox{\boldmath${y}$}},{\mbox{\boldmath${r}$}}\right)=m+\frac{v-2}{2s}\left(R\bar{{\mbox{\boldmath${\beta}$}}}_{H_{1}}-{\mbox{\boldmath${r}$}}\right)^{\prime}\left(RV^{*}R^{\prime}\right)^{-1}\left(R\bar{{\mbox{\boldmath${\beta}$}}}_{H_{1}}-{\mbox{\boldmath${r}$}}\right).
\end{equation*}
For simplicity, we consider the case in which ${\mbox{\boldmath${\beta}$}}=\left(\beta_{1},\beta_{2}, \beta_{3},\beta_{4}\right)$, ${\mbox{\boldmath${x}$}}_{i}=\left(x_{i1},x_{i2},x_{i3},x_{i4}\right)^{\prime}$, where $x_{i1}=1$, $x_{i1},x_{i2},x_{i3},x_{i4}\sim N\left(0,1\right)$. In order to compare the empirical power and size between the new statistics and Wald statistic, the parameter values we use to simulate data are designed as $\sigma^{2}=0.01, \beta_{1} = 0.3, \beta_{2} = 0.2, \beta_{3} = 0.1\gamma, \beta_{4} = 0.5\gamma$ for $\gamma = 0, 0.1, 0.3, 0.5$. The replication number is 1000 and we consider the circumstances where the sample sizes are $n=50,100,150$, respectively.
In each replication, given the sample size, after the data simulation, we consider the hypothesis problem that whether $\beta_{2}=0$,$\beta_{3}=0$ , $\beta_{2}=\beta_{3}=0$ and $\beta_{2} + \beta_{3}=0$. In order to estimate the parameters, the prior we use is
\begin{equation*}
\tilde{\mu}=\left(0,0,0,0\right)^{\prime},\tilde{V}=1000I_{4},
\end{equation*}
\begin{equation*}
a=0.0001,b=0.0001,
\end{equation*}
where $I_{4}$ is the $4\times4$ identity matrix. For each scenarios,
we draw 5000 samples from the posterior distribution and then use
the posterior samples to obtain the posterior mean.
Given the credit level $95\%$, the ratios of the replications that reject the null hypothesis are computed and listed in Table $\ref{smlexpl1table1}$ in different scenarios. From the table, on one hand, the empirical size for the new statistic is quite good and almost the same as Wald statistic. For all the hypothesis problems, the size is approaching $5\%$ as the sample size increase. On the other hand, the empirical power performs well similar to the Wald statitic. As the $\gamma$ becomes larger, which implies that the values of parameters are further away from zero, and the sample size increase, the empirical power of the new statistic goes to $100\%$. All in all, the empirical power and size of the new statistic are very good and almost the same as the those of Wald statistic.
\begin{table}[tbph]
\caption{The empirical sizes and powers for linear model in different scenarios}
\label{smlexpl1table1}
\begin{center}
\footnotesize
\begin{tabular}{c|c|cc|cc|cc|cc}
\hline
\multicolumn{1}{c}{} & & \multicolumn{2}{c|}{Empirical Size} & \multicolumn{6}{c}{Empirical Power}\\
\cline{3-10}
\multicolumn{1}{c}{} & & \multicolumn{2}{c|}{$\gamma=0$} & \multicolumn{2}{c|}{$\gamma=0.1$} & \multicolumn{2}{c|}{$\gamma=0.3$} & \multicolumn{2}{c}{$\gamma=0.5$}\\
\hline
& Null Hypothsis & $\boldsymbol{T}\left(\boldsymbol{y},\breve{\boldsymbol{\beta}}_{0}\right)$ & Wald & $\boldsymbol{T}\left(\boldsymbol{y},\breve{\boldsymbol{\beta}}_{0}\right)$ & Wald & $\boldsymbol{T}\left(\boldsymbol{y},\breve{\boldsymbol{\beta}}_{0}\right)$ & Wald & $\boldsymbol{T}\left(\boldsymbol{y},\breve{\boldsymbol{\beta}}_{0}\right)$ & Wald \\
\hline
\multirow{4}{*}{$n=50$} & $\beta_{3}=0$ & $4.50\%$ & $5.10\%$ & $10.40\%$ & $11.00\%$ & $55.80\%$ & $57.30\%$ & $92.00\%$ & $92.20\%$\\
& $\beta_{4}=0$ & $6.50\%$ & $7.10\%$ & $92.00\%$ & $92.5\%$ & $100\%$ & $100\%$ & $100\%$ & $100\%$\\
& $\beta_{3}=\beta_{4}=0$ & $6.60\%$ & $7.50\%$ & $88.80\%$ & $89.70\%$ & $100\%$ & $100\%$ & $100\%$ & $100\%$\\
& $\beta_{3}+\beta_{4}=0$ & $6.20\%$ & $6.70\%$ & $83.30\%$ & $84.00\%$ & $100\%$ & $100\%$ & $100\%$ & $100\%$\\
\hline
\multirow{4}{*}{$n=100$} & $\beta_{3}=0$ & $5.50\%$ & $5.80\%$ & $20.20\%$ & $20.40\%$ & $82.00
$ & $82.80\%$ & $99.90\%$ & $100\%$\\
& $\beta_{4}=0$ & $4.60\%$ & $5.00\%$ & $99.70\%$ & $99.70\%$ & $100\%$ & $100\%$ & $100\%$ & $100\%$\\
& $\beta_{3}=\beta_{4}=0$ & $5.70\%$ & $6.00\%$ & $99.50\%$ & $99.50\%$ & $100\%$ & $100\%$ & $100\%$ & $100\%$\\
& $\beta_{3}+\beta_{4}=0$ & $6.00\%$ & $6.20\%$ & $98.60\%$ & $98.60\%$ & $100\%$ & $100\%$ & $100\%$ & $100\%$\\
\hline
\multirow{4}{*}{$n=150$} & $\beta_{3}=0$ & $5.30\%$ & $5.40\%$ & $24.40
$ & $24.60\%$ & $95.90\%$ & $95.90\%$ & $100\%$ & $100\%$\\
& $\beta_{4}=0$ & $5.20\%$ & $5.30\%$ & $100\%$ & $100\%$ & $100\%$ & $100\%$ & $100\%$ & $100\%$\\
& $\beta_{3}=\beta_{4}=0$ & $5.40\%$ & $5.60\%$ & $100\%$ & $100\%$ & $100\%$ & $100\%$ & $100\%$ & $100\%$\\
& $\beta_{3}+\beta_{4}=0$ & $4.20\%$ & $4.20\%$ & $99.80\%$ & $99.80\%$ & $100\%$ & $100\%$ & $100\%$ & $100\%$\\
\hline
\end{tabular}
\end{center}
\end{table}
\subsection{The power and size of ${\mbox{\boldmath${T}$}}\left({\mbox{\boldmath${y}$}},{{\mbox{\boldmath${\beta}$}}}_{0}\right)$ for leverage stochastic volatility model}
In this subsection, we examine the empirical power and size of the new statistic in stochastic volatility model with leverage effect (LSV). It is a type of latent variable models, for which the usual frequentist hypothesis tests such as Wald test can not be applied. But as we emphasize above, our new statistic can be readily used for such models. The model we study is defined as follows,
\begin{equation*}
\begin{cases}
r_{t}=\exp\left(\frac{h_{t}}{2}\right)\epsilon_{t},\\
h_{t+1}=\mu+\phi\left(h_{t}-\mu\right)+\sigma\varepsilon_{t+1},
\end{cases}
\end{equation*}
with
\begin{equation*}
\left(\begin{array}{c}
\epsilon_{t}\\
\varepsilon_{t+1}
\end{array}\right)\sim N\left(\left(\begin{array}{c}
0\\
0
\end{array}\right),\left(\begin{array}{cc}
1 & \rho\\
\rho & 1
\end{array}\right)\right),
\end{equation*}
where $r_{t}$ is the data observed, $h_{t}$ the latent volatility at period $t$. $\rho$ is the leverage effect. $\mu$, $\phi$ and $\sigma$ are the parameters we need to estimate. In order to examine the empirical power and size of the new hypoethesis testing, we use several sets of parameter values to simulate the model. We consider $\mu = -10, \phi = 0.97, \sigma^2 = 0.025$ and besides, $\rho = 0,-0.1,-0.2,-0.4$, respectively. The number of replications is 500 with sample size $T=1000,1500,2000$, respectively.
Then given the sample size $T$, we would like to test whether $\rho=0$
or not. That is,
\begin{equation*}
H_{0}:\rho=0,\text{ }H_{1}:\rho\neq0.
\end{equation*}
The priors we use to estimate the model in each case are listed
in the following,
\begin{equation*}
\mu\sim N\left(0,100\right),\phi\sim Beta\left(1,1\right),\sigma^{-2}\sim\Gamma\left(0.001,0.001\right),\rho\sim U\left(-1,1\right).
\end{equation*}
We use the R2OpenBUGS package to estimate the parameters. We draw 30,000 samples and the first 10,000 is discarded. The remaining 20,000 samples are used to compute the posterior means and statistic. Given the credit level $95\%$, the ratios of the replications that reject the null hypothesis are computed and listed in Table $\ref{smlexpl2table2}$ given different sample size.
From the Table $\ref{smlexpl2table2}$, on one hand, we can find that the empirical power of the new statistic performs well increasingly as the sample size increases. On the other hand, the empirical size also approaching $5.4\%$ as the sample size increases. To conclude, even for latent variable models, in which case usual methods are unavailable, our new statistic also possesses satisfactory power and size properties.
\begin{table}[tbph]
\caption{The rejection ratios of the new statistic for LSV model given credit
level $95\%$}
\label{smlexpl2table2}
\begin{center}
\begin{tabular}{c|c|ccc}
\hline
& Empirical Size & \multicolumn{3}{c}{Empirical Power}\\
\hline
& $\rho=0$ & $\rho=-0.1$ & $\rho=-0.2$ & $\rho=-0.4$\\
\hline
$T=1000$ & $7.80\%$ & & & $80.60\%$\\
$T=1500$ & $7.00\%$ & & & $93.60\%$\\
$T=2000$ & $5.40\%$ & & & $98.20\%$\\
\hline
\end{tabular}
\end{center}
\end{table}
\section{Empirical Illustrations}
In this section, we illustrate the proposed test statistic using two popular examples in economics and finance. The first example is a Multi-level Probit model. In this example, the observed data likelihood is available in closed-form, facilitating the comparison of the BF, the statistic in LLY and our proposed test. The second example is a stochastic volatility model with leverage effect, which is a typical case of latent variable model, where the volatility is latent.
\subsection{Examining the marginal effects on Probit model}
Li (2006) proposed a Bayesian method to estimate a simultaneous equation model. In her model, the first part is ordered Probit model. The second part is a two-limit censored regression. She tried to examine the effect of high school education on income and unemployment period. Following her experiment, we use the same model and the same data set to implement our new test .
Let $z_{hi}$ denote the high school grade completed by individual $i$, and $
y_{hi}$ denote the latent outcome corresponding to $z_{hi}$, where $h$
labels the schooling outcome, $z_{hi}=1$ if individual $i$ dropped out of
high school after completing the ninth grade, $z_{hi}=2$ if he dropped out
after completing the tenth grade, $z_{hi}=3$ if he dropped out after
completing the eleventh grade, and $z_{hi}=4$ if he completed high school.
\begin{equation}
\begin{cases}
y_{hi}={\mbox{\boldmath${\beta}$}}_{h}^{\prime}{\mbox{\boldmath${x}$}}
_{hi}+\epsilon_{hi,} & \epsilon_{hi}\sim N\left(0,\sigma{}
_{h}^{2}\right),\gamma_{z_{hi}}<y_{hi}<\gamma_{z_{hi}+1} \\
\gamma_{1}=-\infty,\gamma_{2}=0, & \gamma_{2}<\gamma_{3}<\gamma_{4},
\gamma_{4}=1,\gamma_{5}=\infty
\end{cases}
,
\end{equation}
for $i=1,\dots,N,$ where ${\mbox{\boldmath${x}$}}_{hi}$ is a $k_{h}\times1$
vector of individual-level variables, including base year congnitive test
score, parental income, parental education, number of siblings, gender,
race, county level employment growth rate between 1980 and 1982, a
fourth-order polynomial in age and a fourth-order polynomial in the time
eligible to drop out. $\epsilon_{hi}$ is the individual-level random term, $
N\left(\mu,\sigma^{2}\right)$ the normal distribution with mean $\mu$ and
variance $\sigma^{2}$, $\sigma_{h}^{2}$ the variance of the
unobservables, $\left\{ \gamma_{j}\right\} _{j=1}^{5}$ are the cutoff
points, and $N$ is the total number of individuals.
Let $\omega_{ui}$ denote the proportion of time individual $i$ is
unemployed, and $y_{ui}$ the latent outcome corresponding to $\omega_{ui}$,
and $y_{ui}$ is limited as,
\begin{equation}
y_{ui}
\begin{cases}
\leq0 & \omega_{ui}=0 \\
=\omega_{ui} & 0<\omega_{ui}<1 \\
\geq 1 & \omega_{ui}=1
\end{cases}
,
\end{equation}
then the censored regression is,
\begin{equation}
y_{ui}={\mbox{\boldmath${\beta}$}}_{u}^{\prime}{\mbox{\boldmath${x}$}}_{ui}+{
\mbox{\boldmath${s}$}}_{i}^{\prime}{\mbox{\boldmath${\eta}$}}+
\epsilon_{ui},\epsilon_{ui}\sim N\left(0,\sigma{}_{u}^{2}\right)
\end{equation}
for $i=1,2,\dots,N$, where $x_{ui}$ is $k_{u}\times1$ vector of observed
variables, including base year cognitive test score, parental income,
parental education, number of siblings, gender, race, age and a dummy
variable indicating any post-secondary education. $\epsilon_{ui}$ is the
unobservable, and $\sigma_{u}^{2}$ is the variance.
In the model, ${\mbox{\boldmath${s}$}}_{i}$ is a $4\times1$ vector of dummy
variables indicating the high school grade completed by individual $i$. Let $
{\mbox{\boldmath${s}$}}_{i}=\left(s_{i,1},s_{i,2},s_{i,3},s_{i,4}\right)^{
\prime}$, then $s_{i,z_{hi}}=1$ and $s_{i,j}=0$, $j\neq z_{hi}$. ${
\mbox{\boldmath${\eta}$}}$ indicates the $4\times1$ vector of school
coefficients of $s_{i}$, which is different from the model in Li (2006).
The random terms are correlated,
\begin{equation*}
\left(
\begin{array}{c}
\epsilon_{hi} \\
\epsilon_{ui}
\end{array}
\right)\sim N\left(\left(
\begin{array}{c}
0 \\
0
\end{array}
\right),\left(
\begin{array}{cc}
\sigma_{h}^{2} & \sigma_{hu} \\
\sigma_{hu} & \sigma_{u}^{2}
\end{array}
\right)\right)=N\left(0_{2\times1},\Sigma\right).
\end{equation*}
In the paper, the author used Bayesian method to estimate the parameters.
The priors she used are listed in the following.
\begin{equation*}
{\mbox{\boldmath${\beta}$}}=\left({\mbox{\boldmath${\beta}$}}_{h}^{\prime},{
\mbox{\boldmath${\beta}$}}_{u}^{\prime}\right)^{\prime} \sim N\left({
\mbox{\boldmath${\beta}$}}_{0},V_{\beta}\right),\ \Sigma\sim
IW\left(\rho,\rho R\right),
\end{equation*}
\begin{equation*}
{\mbox{\boldmath${\eta}$}}\sim N\left({\mbox{\boldmath${\eta}$}}
_{0},V_{\eta}\right),\ \gamma_{3}\sim Beta\left(u_{1},u_{2}\right),
\end{equation*}
where ${\mbox{\boldmath${\beta}$}}_{0}=0_{k\times1}$, $k=k_{h}+k_{u}$, $
V_{\beta}=1000I_{k}$, $IW\left(\rho,\rho R\right)$ denotes the inverted
Wishart distribution with degrees of freedom parameter $\rho$ and scale
parameter $R$, $\rho=6$, $R=I_{2}$, ${\mbox{\boldmath${\eta}$}}_{0}=0_{4\times1}$,$V_{\eta} = I_{4}$ , $u_{1}=u_{2}=1$, $Beta\left(\alpha,\delta\right)$ denotes the Beta distribution.
The estimation is almost the same as the the Gibbs method proposed by
Li (2006). We run the MCMC for 20,000 times. After dropping the first 4000
samples and convergence checking, we treat the left 16,000 as the effective
draws. The posterior means and the posterior standard errors are reported in
Table $\ref{empiricalexpl1table1}$.
\begin{table}[tbp]
\caption{The Posterior Means and Standard Errors of Parameters(without the
dummy variables)}
\label{empiricalexpl1table1}
\begin{center}
\begin{tabular}{clrrl}
\hline
& \multirow{2}{*}{} & $E\left(\cdot|Data\right)$ & $SE\left(\cdot|Data\right)
$ & \\ \cline{3-4}
& & \multicolumn{2}{l}{High school completion $y_{h}$} & \\ \cline{2-4}
& Constant & 0.9474 & 0.2119 & \\
& Parental income & 0.0110 & 0.0262 & \\
& Base year cognitive test & 0.4413 & 0.0370 & \\
& Father\textquoteright s education & 0.0456 & 0.0131 & \\
& Mother\textquoteright s education & 0.0627 & 0.0159 & \\
& Number of siblings & -0.0370 & 0.0153 & \\
& Female & -0.0694 & 0.0534 & \\
& Minority & 0.3840 & 0.0664 & \\
& County employment growth & -0.0132 & 0.0047 & \\
& Age & -0.4150 & 0.0853 & \\
& $\mbox{Age}^{2}$ & -0.1887 & 0.0766 & \\
& $\mbox{Age}^{3}$ & -0.0333 & 0.0468 & \\
& $\mbox{Age}^{4}$ & 0.0311 & 0.0148 & \\
& Time eligible to drop out & 0.0932 & 0.0696 & \\
& $\mbox{Time}^{2}$ & 0.0905 & 0.0473 & \\
& $\mbox{Time}^{3}$ & -0.0090 & 0.0106 & \\
& $\mbox{Time}^{4}$ & -0.0094 & 0.0053 & \\
& & \multicolumn{2}{l}{Proportion of time unemployed $\omega_{u}$} & \\
\cline{2-4}
& Parental income & -0.0275 & 0.0056 & \\
& Base year cognitive test & -0.0392 & 0.0071 & \\
& Father\textquoteright s education & -0.0020 & 0.0025 & \\
& Mother\textquoteright s education & -0.0043 & 0.0030 & \\
& Number of siblings & 0.0049 & 0.0034 & \\
& Post-secondary education & -0.0113 & 0.0138 & \\
& Female & 0.0621 & 0.0112 & \\
& Minority & 0.0826 & 0.0131 & \\
& Age & -0.0058 & 0.0126 & \\
& Completing ninth grade(${\mbox{\boldmath${\eta}$}}_{1}$) & 0.1925 & 0.0705
& \\
& Completing tenth grade(${\mbox{\boldmath${\eta}$}}_{2}$) & 0.1211 & 0.0530
& \\
& Completing eleventh grade(${\mbox{\boldmath${\eta}$}}_{3}$) & 0.1187 &
0.0492 & \\
& Completing high school(${\mbox{\boldmath${\eta}$}}_{4}$) & 0.0083 & 0.0416
& \\
& & \multicolumn{2}{l}{Civariance matrix $\Sigma$} & \\ \cline{2-4}
& $\sigma_{h}^{2}$ & 0.9450 & 0.0914 & \\
& $\sigma_{u}^{2}$ & 0.1215 & 0.0039 & \\
& $\sigma_{hu}$ & -0.0099 & 0.0191 & \\
& & \multicolumn{2}{l}{Cutoff point} & \\ \cline{2-4}
& $\gamma_{3}$ & 0.6684 & 0.0220 & \\ \hline
\end{tabular}
\end{center}
\end{table}
In this example, we try to examine whether the marginal effects of father's
education $\left(\beta_{4}\right)$ and mother's education $\left(\beta_{5}
\right)$ on the completion of high school can be ignored or not. Since the ${
\mbox{\boldmath${T}$}}\left(Data,{\mbox{\boldmath${\theta}$}}_{0}\right)$
does not have analytical expression, according to the Appendix \ref{empexpl1}, we use the
MCMC output to approximate the statistic. Further, in order to compare the
statistic with the one in Li, Liu and Yu (2015) and the Bayes factor, we also
report the $\widehat{\log BF_{10}}$ and the $\widehat{{\mbox{\boldmath${T}$}}
}_{LLY}\left(Data,{\mbox{\boldmath${\theta}$}}_{0}\right)$ in the Table \ref{empiricalexpl1table2}.
In this case, the log-likelihood has closed-form expression. Hence, the corresponding numerical standard error for each statistic is also reported in the Table $\ref{empiricalexpl1table2}$.
\begin{table}[tbp]
\caption{The proposed statistic $\protect\widehat{{\mbox{\boldmath${T}$}}}
\left(Data,{\mbox{\boldmath${\theta}$}}_{0}\right)$,$\protect\widehat{{
\mbox{\boldmath${T}$}}}_{LLY}\left(Data,{\mbox{\boldmath${\theta}$}}
_{0}\right)$, $\protect\widehat{\log BF_{10}}$, their computing time(in
seconds), and the numerical standard errors.}
\label{empiricalexpl1table2}
\begin{center}
\begin{tabular}{cccc}
\hline
& \multicolumn{3}{c}{$\beta_{3}=\beta_{4}=0$} \\ \cline{2-4}
& Value & NSE & Time \\ \hline
$\widehat{{\mbox{\boldmath${T}$}}}\left(Data,{\mbox{\boldmath${\theta}$}}
_{0}\right)$ & 45.39 & 1.59 & 48311.31 \\
$\widehat{{\mbox{\boldmath${T}$}}}_{LLY}\left(Data,{\mbox{\boldmath${
\theta}$}}_{0}\right)$ & 2502.00 & 89.57 & 87385.55 \\
$\widehat{\log BF_{10}}$ & 5.2019 & 1.03 & 341175.45 \\ \hline
\end{tabular}
\end{center}
\end{table}
The result we obtained in the Table $\ref{empiricalexpl1table2}$ strongly prove the
advantages of the proposed statistic. The 99.99 percentile of $
\chi^{2}\left(2\right)$ is 18.42. Both the $\widehat{{\mbox{\boldmath${T}$}}}
\left(Data,{\mbox{\boldmath${\theta}$}}_{0}\right)$ and $\widehat{{
\mbox{\boldmath${T}$}}}_{LLY}\left(Data,{\mbox{\boldmath${\theta}$}}
_{0}\right)$ are much larger than 18.42, which indicates that the null
hypothesis is rejected under the $99.99\%$ probability level. Those results
are consistent with the value of $\widehat{\log BF_{10}}$, which strongly
supports the alternative hypothesis. Those three statistics all tell us that
the marginal effect of the parents' education on the high school completion
is not negligible. Further, they all have small numerical standard error
compared with the corresponding values.
What's more from the table we can learn is that the proposed statistic takes
much less time than the other two statistics. It takes around as half as the
time $\widehat{{\mbox{\boldmath${T}$}}}_{LLY}\left(Data,{\mbox{\boldmath${
\theta}$}}_{0}\right)$ used and one fifth of the time Bayes facor used.
\begin{remark}
From the example above, we can readily find that for the problem of high-dimensional parameters, the computation of the new statistic avoids the inversion of the large-scale information matrix, which is inevitable when we use Wald statistic. As we know, when the matrix is of large-scale, the information matrix may not be positive definite, therefore is singular. However, by using our new statistic, we only need compute the posterior covariance, which are the byproduct of the estimation procedure. Therefore, our statistic is superior to the Wald statistic for the hypothesis problems in a problem with many parameters.
\end{remark}
\subsection{Testing the leverage effect on the stochastic volatility models}
Stochastic volatility models are widely used in finance and economics. The
financial leverage effect is very important and documented in many
financial literature, see Black (1976). Following Yu (2005), the leverage
effects SV model is defined as follows:
\begin{equation*}
\begin{cases}
r_{t}=\exp\left(\frac{h_{t}}{2}\right)\epsilon_{t} \\
h_{t+1}=\mu+\phi \left(h_{t} - \mu\right)+\sigma\varepsilon_{t+1}
\end{cases}
\end{equation*}
with
\begin{equation*}
\left(
\begin{array}{c}
\epsilon_{t} \\
\varepsilon_{t+1}
\end{array}
\right)\overset{i.i.d.}{\sim}N\left(\left(
\begin{array}{c}
0 \\
0
\end{array}
\right),\left(
\begin{array}{cc}
1 & \rho \\
\rho & 1
\end{array}
\right)\right),
\end{equation*}
and $h_{0}=\mu$, where $r_{t}$ is the return at time $t$, $h_{t}$ the return
volatility at period t. In this model, $\rho$ is the parameter indicating the
leverage effect. When $\rho<0$, there is a negative relationship between the
expected future volatility and the current return (Yu, 2005). In particular,
volatility tends to rise in response to bad news but fall in response to
good news (Black, 1976). Hence, we construct the hypothesis, $H_{0}:\rho=0$,
to test whether the leverage effect exists or not.
In this example, we used two cases to illustrate how to use the proposed
statistic. And further, the statistic is also compared with the statistic
proposed by LLY, ${\mbox{\boldmath${T}$}}_{LLY}\left({\mbox{\boldmath${y}$}},
{\mbox{\boldmath${\theta}$}}_{0}\right)$ and Bayes factor. The derivation of
the computation is given in the Appendix \ref{empexpl1}.
In the first case, we use the data that consist of daily returns on
Pound/Dollar exchange rates from 01/10/81 to 28/06/85 with sample size 945. The series $r_{t}$ is
the daily mean-corrected returns. The R2OpenBUGS is used to estimate the
model with the following priors for each parameter:
\begin{equation*}
\mu\sim N\left(0,100\right),\phi\sim
Beta\left(1,1\right),\sigma^{-2}\sim\Gamma\left(0.001,0.001\right),\rho\sim
U\left(-1,1\right).
\end{equation*}
\begin{table}[tbph]
\caption{The posterior mean of parameter estimated in case 1}
\label{empiricalexpl2table1}
\begin{center}
\vspace{1em}
\begin{tabular}{ccccc}
\hline
& \multicolumn{2}{c}{$H_{1}$} & \multicolumn{2}{c}{$H_{0}$} \\ \hline
Parameter & Mean & SE & Mean & SE \\ \hline
$\mu$ & -0.5776 & 0.3487 & -0.6608 & 0.3164 \\
$\phi$ & 0.9849 & 0.0097 & 0.9793 & 0.0127 \\
$\rho$ & -0.0941 & 0.1507 & - & - \\
$\tau$ & 0.1553 & 0.0243 & 0.1618 & 0.0360 \\ \hline
\end{tabular}
\end{center}
\end{table}
We draw 50,000 from the posterior distribution and discard the first 20,000 as build-in period. Then we store every 5th value of the remaining samples as effective observations. The estimation
results are reported in Table $\ref{empiricalexpl2table1}$.
We aim to test whether there is leverage effect or not, hence the hypothesis problem is:
\begin{equation*}
H_{0}:\rho=0,\,\,H_{1}:\rho\neq0.
\end{equation*}
\begin{table}[tbph]
\caption{The statistic $\protect\widehat{{\mbox{\boldmath${T}$}}}_{LLY}\left(
{\ \mbox{\boldmath${y}$}},{\mbox{\boldmath${\theta}$}}_{0}\right)$, $\protect
\widehat{{\ \mbox{\boldmath${T}$}}}\left({\mbox{\boldmath${y}$}},{
\mbox{\boldmath${\theta}$}}_{0}\right)$, $\protect\widehat{\log BF_{10}}$,
their computing time (in seconds), and the numerical standard errors of the
first two statistics in case 1.}
\label{empiricalexpl2table2}
\begin{center}
\vspace{1em}
\begin{tabular}{cccc}
\hline
& $\widehat{{\mbox{\boldmath${T}$}}}\left({\mbox{\boldmath${y}$}},{
\mbox{\boldmath${\theta}$}}_{0}\right)$ & $\widehat{{\mbox{\boldmath${T}$}}}
_{LLY}\left({\mbox{\boldmath${y}$}},{\mbox{\boldmath${\theta}$}}_{0}\right)$
& $\widehat{\log BF_{10}}$ \\ \hline
Value & 1.3893 & 0.2883 & -10.1235 \\
NSE & 0.0255 & 0.2028 & - \\
Time used(s) & 1765.3591 & 2313.4812 & 5465.6422 \\ \hline
\end{tabular}
\end{center}
\end{table}
In Table $\ref{empiricalexpl2table2}$, we report the Bayes factor, the statistic in LLY, $
\widehat{{\mbox{\boldmath${T}$}}}_{LLY}\left({\mbox{\boldmath${y}$}},{
\mbox{\boldmath${\theta}$}}_{0}\right)$, and the proposed statistic, $
\widehat{{\mbox{\boldmath${T}$}}}\left({\mbox{\boldmath${y}$}},{
\mbox{\boldmath${\theta}$}}_{0}\right)$. The $\widehat{\log BF_{10}}$
strongly supports the null hypothesis, that is, there is not leverage
effect. Meanwhile, since ${\mbox{\boldmath${T}$}}_{LLY}\left({
\mbox{\boldmath${y}$}},{\mbox{\boldmath${\theta}$}}_{0}\right)$ follows a $
\chi^{2}\left(1\right)$ distribution, the value of this statistic with a
rather small NSE shows that it fails to reject the null hypothesis at the
95\% probability level. For the porposed statistic, ${\mbox{\boldmath${T}$}}
\left({\mbox{\boldmath${y}$}},{\mbox{\boldmath${\theta}$}}_{0}\right)-1
\overset{d}{\rightarrow}\chi^{2}\left(1\right)$, $\widehat{{
\mbox{\boldmath${T}$}}}\left({\mbox{\boldmath${y}$}},{\mbox{\boldmath${
\theta}$}}_{0}\right)-1$ is closed to $\widehat{{\mbox{\boldmath${T}$}}}
_{LLY}\left({\mbox{\boldmath${y}$}},{\mbox{\boldmath${\theta}$}}_{0}\right)$
with smaller NSE. Thus it can not reject the null hypothesis under 95\%
probability level. Thus, the outcomes of all three statistics are consistent.
In the second case, the data we use is 1,822 daily returns of the Standard \& Poor (S\&P) 500 index,
covering the period between January 3, 2005 and March 28, 2012. We use the same priors and similar method to estimate the model. And the parameter estimated are
listed in Table $\ref{empiricalexpl2table3}$, which are quite different from the first case.
\begin{table}[tbph]
\caption{The posterior mean of parameter estimated in case 2}
\label{empiricalexpl2table3}
\begin{center}
\vspace{1em}
\begin{tabular}{ccccc}
\hline
& \multicolumn{2}{c}{$H_{1}$} & \multicolumn{2}{c}{$H_{0}$} \\ \hline
Parameter & Mean & SE & Mean & SE \\ \hline
$\mu$ & -10.8800 & 0.1751 & -11.2200 & 0.3349 \\
$\phi$ & 0.9804 & 0.0039 & 0.9897 & 0.0042 \\
$\rho$ & -0.7151 & 0.0422 & - & - \\
$\tau$ & 0.2057 & 0.0178 & 0.1705 & 0.0169 \\ \hline
\end{tabular}
\end{center}
\end{table}
Again, the three statistics are reported in Table $\ref{empiricalexpl2table4}$. Contrary to
first case, all the statistics strongly support the null hypotheses, that is,
there is leverage effect in the data. For $\widehat{{\mbox{\boldmath${T}$}}}
\left({\mbox{\boldmath${y}$}},{\mbox{\boldmath${\theta}$}}_{0}\right)$ and $
\widehat{{\mbox{\boldmath${T}$}}}_{LLY}\left({\mbox{\boldmath${y}$}},{
\mbox{\boldmath${\theta}$}}_{0}\right)$, they all reject the null hypothesis
under the 99\% probability level. At the same time, the $\widehat{
\log BF_{10}}$ also strongly supports the alternative hypothesis. Therefore,
the results of all three statistics are consistent.
\begin{table}[tbph]
\caption{The statistic $\protect\widehat{{\mbox{\boldmath${T}$}}}_{LLY}\left(
{\ \mbox{\boldmath${y}$}},{\mbox{\boldmath${\theta}$}}_{0}\right)$, $\protect
\widehat{{\ \mbox{\boldmath${T}$}}}\left({\mbox{\boldmath${y}$}},{
\mbox{\boldmath${\theta}$}}_{0}\right)$, $\protect\widehat{\log BF_{10}}$,
their computing time (in seconds), and the numerical standard errors of the
first two statistics in case 2.}
\label{empiricalexpl2table3}
\begin{center}
\vspace{1em}
\begin{tabular}{cccc}
\hline
& $\widehat{{\mbox{\boldmath${T}$}}}\left({\mbox{\boldmath${y}$}},{
\mbox{\boldmath${\theta}$}}_{0}\right)$ & $\widehat{{\mbox{\boldmath${T}$}}}
_{LLY}\left({\mbox{\boldmath${y}$}},{\mbox{\boldmath${\theta}$}}_{0}\right)$
& $\widehat{\log BF_{10}}$ \\ \hline
Value & 287.7944 & 8.2419 & 51.9582 \\
NSE & 0.6915 & 0.6849 & - \\
Time used(s) & 5606.8825 & 6862.3672 & 13391.2791 \\ \hline
\end{tabular}
\end{center}
\end{table}
\section{Conclusion}
In this paper, a new $\chi^2$-type Bayesian test statistic is proposed to
test a point null hypothesis for latent variable models. The new statistic
can be explained as Bayesian version of Wald test. Compared with existing
literature, the proposed test statistic has achieved several important
advantages, hence, can appeal many practical applications. First, for latent
variable models, it is only by-product of posterior outputs, hence, is very
easy to compute, not require additional computational efforts after the
model is estimated using MCMC techniques. Second, it is well-defined under
improper prior distributions and avoids Jeffrey-Lindley's paradox. Third,
it's asymptotic distribution is pivotal so that the threshold values can be
easily obtained from the asymptotic chi-squared distribution. Through Monte
Carlo studies, it can be shown that the test power is almost equivalent to
Wald test, but better than the test statistic by Li, Liu and Yu (2014). If
the observed likelihood doesn't have analytical form, Wald test statistic is
very difficult to be applied, but, the proposed test statistic is also easy
to implement.
\textbf{The Bayes factor and the Bayesian test statistic with Li, Liu and Yu
(2015) need to be computed out and make an empirical comparison }
\vskip 0.5cm