Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
88,225 characters · 14 sections · 0 citation commands
A New Wald Test for Hypothesis Testing Based on MCMC outputs
\vskip 2cm \baselineskip=17pt
Latent variable models have been widely used in economics, finance, and many other disciplines. Two typical models are the dynamic stochastic general equilibrium models in macroeconomics and stochastic volatility models in finance. The latent variable models are generally indexed by the latent variable and the parameter. In many latent variable models, the latent variable is generally high-dimensional so that the observed likelihood function which is a marginal integral on the latent variable is often intractable and becomes difficult to evaluate accurately. Consequently, the statistical inference for latent variable models is nontrivial in practice. In the recent years, Bayesian MCMC methods have been applied in more and more applications in economics and finance due to that they make it possible to fit increasingly complex models, especially latent variable models, see Geweke, et al (2011) and reference therein.
In economic research, the point null hypothesis test is a fundamental topic in statistical inference. Under the Bayesian paradigm, the Bayes factors (BFs) are the corner-stone of Bayesian hypothesis testing (e.g. Jeffreys,1961; Kass and Raftery 1995; Geweke, 2007). Unfortunately, the BFs are not problem-free. First, the BFs are sensitive to the prior distribution and subjects to the notorious Jeffreys-Lindley's paradox; see for example, Kass and Raftery (1995), Poirier (1995), Robert (1993, 2001). Second, the calculation of BFs generally involves the evaluation of marginal likelihood. In many cases, the evaluation of marginal likelihood is often difficult.
Not surprisingly, some alternative strategies have been proposed to test a point null hypothesis in the Bayesian literature. In recent years, on the basis of the statistical decision theory, several interesting Bayesian approaches to replace BFs have been developed for hypothesis testing. For example, Bernardo and Rueda (2002, BR hereafter) demonstrated that BFs for the Bayesian hypothesis testing can be regarded as a decision problem with a simple zero-one discrete loss function. However, the zero-one discrete function requires the use of non-regular (not absolutely continuous) prior and this is why BF leads to Jeffreys-Lindley's paradox. BR further suggested using a continuous loss function, based on the well-known continuous Kullback-Leibler (KL) divergence function. As a result, it was shown in BR that their Bayesian test statistic does not depend on any arbitrary constant in the prior. However, BR's approach has some disadvantages. First, the analytical expression of the KL loss function required by BR is not always available, especially for latent variable models. Second, the test statistic is not a pivotal quantity. Consequently, BR had to use subjective threshold values to test the hypothesis.
To deal with the computational problem in BR in latent variable models, Li and Yu (2012, LY hereafter) developed a new test statistic based on the $ \mathcal{Q}$ function in the Expectation-Maximization (EM) algorithm. LY showed that the new statistic is well-defined under improper priors and easy to compute for latent variable models. Following the idea of McCulloch (1989), LY proposed to choose the threshold values based on the Bernoulli distribution. However, like the test statistic proposed by BR, the test statistic proposed by LY is not pivotal. Moreover, it is not clear if the test statistic of LY can resolve Jeffreys-Lindley's paradox.
Based on the difference between the deviances, Li, Zeng and Yu (2014, LZY hereafter) developed another Bayesian test statistic for hypothesis testing. This test statistic is well-defined under improper priors, free of Jeffreys-Lindley's paradox, and not difficult to compute. Moreover, its asymptotic distribution can be derived and one may obtain the threshold values from the asymptotic distribution. Unfortunately, in general the asymptotic distribution depends on some unknown population parameters and hence the test is not pivotal. With sharing the nice properties with Li, Zeng and Yu (2014, LZY hereafter), Li, Liu and Yu (2015)(2015, LLY hereafter) further proposed a pivotal Bayesian test statistic, based on a quadratic loss function, to test a point null hypothesis within the decision-theoretic framework. However, LLY required to evaluate the first derivative of the observed log-likelihood. As to the latent variable models, because the observed log-likelihood is often intractable, this still posed some tedious computational efforts although there have been several interesting methods for evaluating the first derivative, such as EM algorithm, Kalman filter or Particle filter.
In the paper, we want to propose another novel, easy-to-implement Bayesian statistic for hypothesis testing in the framework of latent variable models. The new statistic can share the important advantages with LLY. First, it is well-defined under improper prior distributions and avoids Jeffrey-Lindley's paradox. Second, under some mild regularity conditions, the statistic is asymptotically equivalent to the Wald test. Hence, from the large sample theory, it's asymptotic distribution can be derived to follow the $\chi^2$ distribution so that the threshold values can be easily calibrated from this distribution. Third, it's statistical error can be derived using the Markov chain Monte Carlo (MCMC) approach. In addition, most importantly, compared with the previous test statistics, it is extremely convenient for the latent variable models. We don't need to evaluate the first-order derivative of the observed log-likelihood function, which is time consuming and difficult for the latent variable models. We just need the MCMC output of posterior simulation. The only effort we should make is the inverse of the posterior variance matrix of the interest parameter in hypothesis testing. Fortunately, in most applications, the dimension of the interest parameter is often not so high that our method can be easily applied. In addition, when the prior information is available, we establish the finite sample theoretical properties for the proposed test statistic.
The paper is organized as follows. Section 2 presents the Bayesian analysis for latent variable models. Section 3 develops the new Bayesian test statistic from the decisional viewpoint and establishes its finite and large sample theoretical properties. Section 4 illustrates the new method by using three real examples in economics and finance. Section 5 concludes the paper. Appendix collects the proof of all the theoretical results.
Without loss of generality, let $\mathbf{y}=(\mathbf{y}_{1},\mathbf{y} _{2},\cdots ,\mathbf{y}_{n})^{T}$ denote observed variables and ${ \mbox{\boldmath${z}$}}=({\mbox{\boldmath${z}$}}_{1},{\mbox{\boldmath${z}$}} _{2},\cdots ,{\mbox{\boldmath${z}$}}_{n})^{T},$ the latent variables. The latent variable model is indexed by the parameter, ${\mbox{\boldmath${ \vartheta}$}}$. Let $p(\mathbf{y}|{\mbox{\boldmath${\vartheta}$}})$ be the likelihood function of the observed data, and $p(\mathbf{y},{ \mbox{\boldmath${z}$}}|{\mbox{\boldmath${\vartheta}$}}),$ the complete likelihood function. The relationship between these two likelihood functions is:
In many latent variable modes, especially dynamic latent variable models, the latent variable $\mathbf{z}$ is often dependent on the sample size. Hence, the integral is high-dimensional and often does not have an analytical expression so that it is generally very difficult to evaluate. Consequently, the statistical inferences, such as estimation and hypothesis testing, are difficult to implement if they are based on the popular maximum likelihood approach.
In recent years, it has been documented that the latent variables models can be simply and efficiently analyzed using MCMC techniques under the Bayesian framework. For details about Bayesian analysis of latent variable models via MCMC such as algorithms, examples and references, see Geweke, et al. (2011). Let $p({\mbox{\boldmath${\vartheta}$}})$ be prior distribution of unknown parameter ${\mbox{\boldmath${\vartheta}$}}$. Owing to the complexity induced by latent variables, the observed likelihood $p(\mathbf{y}|{ \mbox{\boldmath${\vartheta}$}})$ is often intractable, hence it is almost impossible to evaluate the expectation of the posterior density $p({ \mbox{\boldmath${\vartheta}$}}|\mathbf{y})$ directly. To alleviate this difficulty, in the posterior analysis, the popular data-augmentation strategy(Tanner and Wong, 1987) is applied to augment the observed variable $ \mathbf{y}$ with the latent variable $\mathbf{z}$. Then, the well-known Gibbs sampler can be used to generate random samples from the joint posterior distribution $p({\mbox{\boldmath${\vartheta}$}} ,\mathbf{z}| \mathbf{y})$. More concretely, we start with an initial value $[{ \mbox{\boldmath${\vartheta}$}}^{(0)},,\mathbf{z}^{(0)}]$, and then simulates one by one; at the $j$th iteration, with current values $[{ \mbox{\boldmath${\vartheta}$}}^{(j)},\mathbf{z}^{(j)}]:$
After the burning-in phase, that is, sufficiently many iterations of this iteration procedure, the simulated random samples can be regarded as efficient random observations from the joint posterior distribution $p({ \mbox{\boldmath${\vartheta}$}} ,\mathbf{z}|\mathbf{y})$.
The statistical inference can be established on the efficient random observations drawn from the posterior distribution. Bayesian estimates of ${ \mbox{\boldmath${\vartheta}$}}$ and latent variables $\mathbf{z}$ as well as their standard errors can be easily obtained via the corresponding sampling mean and sample covariance matrix of the generated random observations. Specifically, let $\{{\mbox{\boldmath${\vartheta}$}}^{(j)},\mathbf{z} ^{(j)},j=1,2,\cdots, J\}$ be effective random observations generated form the joint posterior distribution $p({\mbox{\boldmath${\vartheta}$}},\mathbf{z }|\mathbf{y})$. Then the joint Bayesian estimates of ${\mbox{\boldmath${ \vartheta}$}},\mathbf{z}$, as well as the estimates of their covariance matrix can be obtained as follows:
It is assumed that a probability model $M\equiv \{p(\mathbf{y}|{ \mbox{\boldmath${\theta}$}},{\mbox{\boldmath${\psi}$}})\}$ is used to fit the data. We are concerned with a point null hypothesis testing problem which may arise from the prediction of a particular theory. Let ${ \mbox{\boldmath${\theta}$}}\in \mathbf{\Theta }$ denote a vector of $p$ -dimensional parameters of interest and ${\mbox{\boldmath${\psi}$}}\in \mathbf{\Psi }$ a vector of $q$-dimensional nuisance parameters. The problem of testing a point null hypothesis is given by
The hypothesis testing may be formulated as a decision problem. It is obvious that the decision space has two statistical decisions, to accept $ H_{0}$ (name it $d_{0}$) or to reject $H_{0}$ (name it $d_{1}$). Let $\{ \mathcal{L}[d_{i},({\mbox{\boldmath${\theta}$}},{\mbox{\boldmath${\psi}$}} )],i=0,1\}$ be the loss function of statistical decision. Hence, a natural statistical decision to reject $H_{0}$ can be made when the expected posterior loss of accepting $H_{0}$ is sufficiently larger than the expected posterior loss of rejecting $H_{0}$, i.e.,
where $\mathbf{T}(\mathbf{y},{\mbox{\boldmath${\theta}$}}_{0})$ is a Bayesian test statistic; $p({\mbox{\boldmath${\theta}$}},{ \mbox{\boldmath${\psi}$}}|\mathbf{y})$ the posterior distribution with some given prior $p({\mbox{\boldmath${\theta}$}},{\mbox{\boldmath${\psi}$}})$; $c$ a threshold value. Let $\triangle \mathcal{L}[H_{0},({\mbox{\boldmath${ \theta}$}},{\mbox{\boldmath${\psi}$}})]=\mathcal{L}[d_{0},({ \mbox{\boldmath${\theta}$}},{\mbox{\boldmath${\psi}$}})]-\mathcal{L}[d_{1},({ \mbox{\boldmath${\theta}$}},{\mbox{\boldmath${\psi}$}})]$ be the net loss difference function which can generally be used to measure the evidence against $H_{0}$ as a function of $({\mbox{\boldmath${\theta}$}},{ \mbox{\boldmath${\psi}$}})$. Hence, the Bayesian test statistic can be rewritten as
In this subsection, as to latent variable models, based on the decision theory, we develop a new Bayesian $\chi^2$ test statistic for hypothesis testing. The new test statistic can share the nice advantages with Li, Liu and Yu (2015). For example, it can be well-defined under improper prior distributions and avoids Jeffrey-Lindley's paradox. Furthermore, the threshold values can be easily calibrated from the pivotal asymptotic distribution and it's statistical error can be derived using MCMC approach. Most importantly, the new test statistic can achieve other important advantages over the existing approaches, such as, Li, et al (2015). Our new contributions are twofold. As to latent variable models, it can be shown that the new test statistic is only the by-product of the posterior outputs, hence, very easy to compute. In addition, when the prior information is available, we establish the finite sample theory.
As to any ${\mbox{\boldmath${\tilde\vartheta}$}}$ in support space of ${ \mbox{\boldmath${\vartheta}$}}$, let
In this paper, under the statistical decision theory, we propose the following net loss function for hypothesis testing
where $\mathbf{V}_{\theta\theta}({\mbox{\boldmath${\bar\vartheta}$}})$ is the submatrix of $\mathbf{V}({\mbox{\boldmath${\bar\vartheta}$}})$ corresponding to ${\mbox{\boldmath${\theta}$}}$, $\left[\mathbf{V} _{\theta\theta}({\mbox{\boldmath${\bar\vartheta}$}})\right]^{-1}$ is the inverse matrix of $\mathbf{V}_{\theta\theta}({\mbox{\boldmath${\bar \vartheta}$}})$ and ${\mbox{\boldmath${\bar\vartheta}$}}$ is the posterior mean of ${\mbox{\boldmath${\vartheta}$}}$ under the alternative hypothesis $ H_1$. Then, we can define a Bayesian test statistic as follows:
In this subsection, we establish the Bayesian large sample theory for the proposed test statistic. Let $\left\{ z_{t}\right\} $ be a sequence of random vectors defined on the probability space $(\Omega ,\mathcal{F},P)$ and $z^{t}$ be the collection of $\left( z_{1},z_{2},\ldots ,z_{t}\right) $. Let $y_{t}$ denote an element of $z_{t}$ and write $z_{t}$ as $ (y_{t},w_{t}^{\prime })^{\prime }$, then we can write the conditional likelihood function for $y_{t}$ as $f_{t}\left( y_{t}|x_{t},\vartheta \right) $, where $x_{t}$ include some elements of $w_{t}$ and $z^{t-1}$. Define $g_{t}\left( \vartheta \right) =g_{t}\left( z^{t},\vartheta \right) =\log f_{t}\left( y_{t}|x_{t},\vartheta \right) $ to be the conditional likelihood for $t$ observation and $\nabla ^{j}g_{t}\left( \vartheta \right) $ as the $jth$ derivative of $g_{t}\left( \vartheta \right) $, we suppress the subscript when $j=1$. The logarithm of posterior likelihood function is
Furthermore, let $\dot{\mathcal{L}}_{n}({\mbox{\boldmath${\vartheta}$}} )=\partial \log p({\mbox{\boldmath${\vartheta}$}}|\mathbf{y})/\partial { \mbox{\boldmath${\vartheta}$}}$, $\ddot{\mathcal{L}}_{n}({ \mbox{\boldmath${\vartheta}$}})=\partial ^{2}\log p({\mbox{\boldmath${ \vartheta}$}}|\mathbf{y})/\partial {\mbox{\boldmath${\vartheta}$}}{\partial \mbox{\boldmath${\vartheta}$}}^{\prime }$ and the negative Hessian matrix as
Let the prior density to be $p({\mbox{\boldmath${\vartheta}$}})$, $\gamma({ \mbox{\boldmath${\vartheta}$}})=\log p({\mbox{\boldmath${\vartheta}$}})$ and $\gamma^{{\mbox{\boldmath${\vartheta}$}}}({\mbox{\boldmath${\vartheta}$}} )=\partial\log p({\mbox{\boldmath${\vartheta}$}})/\partial{ \mbox{\boldmath${\vartheta}$}}$. In order to derive the asymptotic distribution of the proposed test statistic, following LZY (2014) and LLY(2015), a set of regularity conditions are imposed in the following.
Let $\widehat{{\mbox{\boldmath${\vartheta}$}}}$ to be the maximum likelihood estimator of ${\mbox{\boldmath${\vartheta}$}}$ and $\widehat{{ \mbox{\boldmath${\theta}$}}}$ is the subvector of $\widehat{{ \mbox{\boldmath${\vartheta}$}}}$ corresponding to ${{\mbox{\boldmath${ \theta}$}}}$, under Assumptions 5-8 with $s_{1}=2$ and $s_{2}=2$, the Wald statistic be
where $\ddot{\mathcal{L}}_{n,\theta\theta}^{-1}(\widehat{{\mbox{\boldmath${\vartheta}$}}})$ is the submatrix of $\ddot{\mathcal{L}}_{n}^{-1}(\widehat{{\mbox{\boldmath${\vartheta}$}}})$ corresponding to ${\mbox{\boldmath${\theta}$}}$.
Since $\mathbf{T}({\mbox{\boldmath${\mathbf{y}}$}},{ \mbox{\boldmath${ \theta_0}$}})$ is calculated by using the MCMC output, it is important to assess the numerical standard error for measuring the magnitude of simulation error.
The Corollary (ref) shows us how to compute the numerical standard error of the proposed statistic. For the NSE of $\widehat{{\mbox{\boldmath${h}$}}}$, $Var\left({\widehat{{\mbox{\boldmath${h}$}}}}\right)$, following Newey and West (1987), a consistent estimator can be given by
where
and the value of $q$ is always equal to 10.
In this section, we extend the point-null hypothesis aforementioned into the following problem,
where $R$ is a $m\times\left(d+q\right)$ matrix, ${\mbox{\boldmath${r}$}}\in\mathbb{R}^{m}$. This hypothesis problem is much more general than the previous one. On the other hand, it can help us to study the relationship among parameters. Further, for such problems, it is hard to use the Bayes factor. Hence, the extension here is meaningful.
For such problem, the frequentist Wald statistic is
where $\widehat{{\mbox{\boldmath${\vartheta}$}}}$ is the MLE estimator of ${\mbox{\boldmath${\vartheta}$}}$.
According to the decision theory, we define the net loss function for such problem as
where $\bar{{\mbox{\boldmath${\vartheta}$}}}$ is the posterior mean of ${\mbox{\boldmath${\vartheta}$}}$, $V\left(\bar{{\mbox{\boldmath${\vartheta}$}}}\right)=E\left[\left.\left({\mbox{\boldmath${\vartheta}$}}-\bar{{\mbox{\boldmath${\vartheta}$}}}\right)\left({\mbox{\boldmath${\vartheta}$}}-\bar{{\mbox{\boldmath${\vartheta}$}}}\right)^{\prime}\right|{\mbox{\boldmath${y}$}},H_{1}\right]$. Then the statistic is defined as
Similarly, for the statistic, the numerical standard error can be computed in the following corollary.
For the NSE of $\widehat{{\mbox{\boldmath${h}$}}}$, $Var\left({\widehat{{\mbox{\boldmath${h}$}}}}\right)$, we can still follow the way proposed by Newey and West (1987) to evaluated.
In this section, we do two simulation studies to check the empirical size and power of the proposed test statistic. The first example is a simple simulation examination based on linear regression model where the our proposed test statistic has analytical expression. We compare the size and power of the new statistics with the Wald statistic. In the second example, we use the stochastic volatility model with leverage effect, where Wald statistic can not be used, to study the size and power of our statistic.
In this subsection, we use the simple linear regression model to examine the empirical power and size of the proposed test statistic. The model we use is
with ${\mbox{\boldmath${x}$}}_{i1}=1$. Let ${\mbox{\boldmath${X}$}}=\left({\mbox{\boldmath${x}$}}_{1}^{\prime},\dots,{\mbox{\boldmath${x}$}}_{N}^{\prime}\right)^{\prime}$, then we can rewrite the model in matrix form,
where ${\mbox{\boldmath${y}$}}=\left(y_{1},\dots,y_{n}\right)^{\prime}$, ${\mbox{\boldmath${\epsilon}$}}=\left(\epsilon_{1},\dots,\epsilon_{n}\right)^{\prime}$.
We are interested in the subvector of ${\mbox{\boldmath${\beta}$}}$, $\breve{{\mbox{\boldmath${\beta}$}}}$, then ${\mbox{\boldmath${\beta}$}}=\left(\breve{{\mbox{\boldmath${\beta}$}}}^{\prime},\tilde{{\mbox{\boldmath${\beta}$}}}^{\prime}\right)^{\prime}$. Here we want to test $H_{0}:\breve{{\mbox{\boldmath${\beta}$}}}=\breve{{\mbox{\boldmath${\beta}$}}}_{0}$ against $H_{1}:\breve{{\mbox{\boldmath${\beta}$}}}\neq\breve{{\mbox{\boldmath${\beta}$}}}_{0}$ and $H_{0}:R{\mbox{\boldmath${\beta}$}}={{\mbox{\boldmath${r}$}}}$ against $H_{1}:R{\mbox{\boldmath${\beta}$}}\neq{{\mbox{\boldmath${r}$}}}$. Assume that the prior distribution for ${\mbox{\boldmath${\beta}$}}$ and $\sigma^{2}$ are normal and inverse gamma, respectively,
where $\mu_{0}$, $V_{0}$ and $a$, $b$ are hyperparameters.
The proposed statistic ${\mbox{\boldmath${T}$}}\left({\mbox{\boldmath${y}$}},{{\mbox{\boldmath${\beta}$}}}_{0}\right)$ for the first problem is
where $v=2a+n$, $s=b+\frac{1}{2}\left(\mu_{0}^{\prime}V_{0}^{-1}\mu_{0}+{\mbox{\boldmath${y}$}}^{\prime}{\mbox{\boldmath${y}$}}-\mu^{*\prime}V^{*-1}\mu^{*}\right)$, $V^{*}=\left(V_{0}^{-1}+{\mbox{\boldmath${X}$}}^{\prime}{\mbox{\boldmath${X}$}}\right)^{-1}$,$\mu^{*}=V^{*}\left(V_{0}^{-1}\tilde{\mu}+{\mbox{\boldmath${X}$}}^{\prime}{\mbox{\boldmath${y}$}}\right)$ and $\breve{V}^{*}$ the submatrix of $V^{*}$ corresponding to $\breve{{\mbox{\boldmath${\beta}$}}}$. $p$ is the dimension of $\breve{{\mbox{\boldmath${\beta}$}}}$ and $\bar{\breve{{\mbox{\boldmath${\beta}$}}}}_{H_{1}}$ is the posterior mean of $\breve{{\mbox{\boldmath${\beta}$}}}$ under $H_{1}$. The details is given in the Appendix (ref).
For the second hypothesis problem, it can be readily derived that the statistic is
For simplicity, we consider the case in which ${\mbox{\boldmath${\beta}$}}=\left(\beta_{1},\beta_{2}, \beta_{3},\beta_{4}\right)$, ${\mbox{\boldmath${x}$}}_{i}=\left(x_{i1},x_{i2},x_{i3},x_{i4}\right)^{\prime}$, where $x_{i1}=1$, $x_{i1},x_{i2},x_{i3},x_{i4}\sim N\left(0,1\right)$. In order to compare the empirical power and size between the new statistics and Wald statistic, the parameter values we use to simulate data are designed as $\sigma^{2}=0.01, \beta_{1} = 0.3, \beta_{2} = 0.2, \beta_{3} = 0.1\gamma, \beta_{4} = 0.5\gamma$ for $\gamma = 0, 0.1, 0.3, 0.5$. The replication number is 1000 and we consider the circumstances where the sample sizes are $n=50,100,150$, respectively.
In each replication, given the sample size, after the data simulation, we consider the hypothesis problem that whether $\beta_{2}=0$,$\beta_{3}=0$ , $\beta_{2}=\beta_{3}=0$ and $\beta_{2} + \beta_{3}=0$. In order to estimate the parameters, the prior we use is
where $I_{4}$ is the $4\times4$ identity matrix. For each scenarios, we draw 5000 samples from the posterior distribution and then use the posterior samples to obtain the posterior mean.
Given the credit level $95\%$, the ratios of the replications that reject the null hypothesis are computed and listed in Table $\ref{smlexpl1table1}$ in different scenarios. From the table, on one hand, the empirical size for the new statistic is quite good and almost the same as Wald statistic. For all the hypothesis problems, the size is approaching $5\%$ as the sample size increase. On the other hand, the empirical power performs well similar to the Wald statitic. As the $\gamma$ becomes larger, which implies that the values of parameters are further away from zero, and the sample size increase, the empirical power of the new statistic goes to $100\%$. All in all, the empirical power and size of the new statistic are very good and almost the same as the those of Wald statistic.
In this subsection, we examine the empirical power and size of the new statistic in stochastic volatility model with leverage effect (LSV). It is a type of latent variable models, for which the usual frequentist hypothesis tests such as Wald test can not be applied. But as we emphasize above, our new statistic can be readily used for such models. The model we study is defined as follows,
with
where $r_{t}$ is the data observed, $h_{t}$ the latent volatility at period $t$. $\rho$ is the leverage effect. $\mu$, $\phi$ and $\sigma$ are the parameters we need to estimate. In order to examine the empirical power and size of the new hypoethesis testing, we use several sets of parameter values to simulate the model. We consider $\mu = -10, \phi = 0.97, \sigma^2 = 0.025$ and besides, $\rho = 0,-0.1,-0.2,-0.4$, respectively. The number of replications is 500 with sample size $T=1000,1500,2000$, respectively.
Then given the sample size $T$, we would like to test whether $\rho=0$ or not. That is,
The priors we use to estimate the model in each case are listed in the following,
We use the R2OpenBUGS package to estimate the parameters. We draw 30,000 samples and the first 10,000 is discarded. The remaining 20,000 samples are used to compute the posterior means and statistic. Given the credit level $95\%$, the ratios of the replications that reject the null hypothesis are computed and listed in Table $\ref{smlexpl2table2}$ given different sample size.
From the Table $\ref{smlexpl2table2}$, on one hand, we can find that the empirical power of the new statistic performs well increasingly as the sample size increases. On the other hand, the empirical size also approaching $5.4\%$ as the sample size increases. To conclude, even for latent variable models, in which case usual methods are unavailable, our new statistic also possesses satisfactory power and size properties.
In this section, we illustrate the proposed test statistic using two popular examples in economics and finance. The first example is a Multi-level Probit model. In this example, the observed data likelihood is available in closed-form, facilitating the comparison of the BF, the statistic in LLY and our proposed test. The second example is a stochastic volatility model with leverage effect, which is a typical case of latent variable model, where the volatility is latent.
Li (2006) proposed a Bayesian method to estimate a simultaneous equation model. In her model, the first part is ordered Probit model. The second part is a two-limit censored regression. She tried to examine the effect of high school education on income and unemployment period. Following her experiment, we use the same model and the same data set to implement our new test .
Let $z_{hi}$ denote the high school grade completed by individual $i$, and $ y_{hi}$ denote the latent outcome corresponding to $z_{hi}$, where $h$ labels the schooling outcome, $z_{hi}=1$ if individual $i$ dropped out of high school after completing the ninth grade, $z_{hi}=2$ if he dropped out after completing the tenth grade, $z_{hi}=3$ if he dropped out after completing the eleventh grade, and $z_{hi}=4$ if he completed high school.
for $i=1,\dots,N,$ where ${\mbox{\boldmath${x}$}}_{hi}$ is a $k_{h}\times1$ vector of individual-level variables, including base year congnitive test score, parental income, parental education, number of siblings, gender, race, county level employment growth rate between 1980 and 1982, a fourth-order polynomial in age and a fourth-order polynomial in the time eligible to drop out. $\epsilon_{hi}$ is the individual-level random term, $ N\left(\mu,\sigma^{2}\right)$ the normal distribution with mean $\mu$ and variance $\sigma^{2}$, $\sigma_{h}^{2}$ the variance of the unobservables, $\left\{ \gamma_{j}\right\} _{j=1}^{5}$ are the cutoff points, and $N$ is the total number of individuals.
Let $\omega_{ui}$ denote the proportion of time individual $i$ is unemployed, and $y_{ui}$ the latent outcome corresponding to $\omega_{ui}$, and $y_{ui}$ is limited as,
then the censored regression is,
for $i=1,2,\dots,N$, where $x_{ui}$ is $k_{u}\times1$ vector of observed variables, including base year cognitive test score, parental income, parental education, number of siblings, gender, race, age and a dummy variable indicating any post-secondary education. $\epsilon_{ui}$ is the unobservable, and $\sigma_{u}^{2}$ is the variance.
In the model, ${\mbox{\boldmath${s}$}}_{i}$ is a $4\times1$ vector of dummy variables indicating the high school grade completed by individual $i$. Let $ {\mbox{\boldmath${s}$}}_{i}=\left(s_{i,1},s_{i,2},s_{i,3},s_{i,4}\right)^{ \prime}$, then $s_{i,z_{hi}}=1$ and $s_{i,j}=0$, $j\neq z_{hi}$. ${ \mbox{\boldmath${\eta}$}}$ indicates the $4\times1$ vector of school coefficients of $s_{i}$, which is different from the model in Li (2006).
The random terms are correlated,
In the paper, the author used Bayesian method to estimate the parameters. The priors she used are listed in the following.
where ${\mbox{\boldmath${\beta}$}}_{0}=0_{k\times1}$, $k=k_{h}+k_{u}$, $ V_{\beta}=1000I_{k}$, $IW\left(\rho,\rho R\right)$ denotes the inverted Wishart distribution with degrees of freedom parameter $\rho$ and scale parameter $R$, $\rho=6$, $R=I_{2}$, ${\mbox{\boldmath${\eta}$}}_{0}=0_{4\times1}$,$V_{\eta} = I_{4}$ , $u_{1}=u_{2}=1$, $Beta\left(\alpha,\delta\right)$ denotes the Beta distribution.
The estimation is almost the same as the the Gibbs method proposed by Li (2006). We run the MCMC for 20,000 times. After dropping the first 4000 samples and convergence checking, we treat the left 16,000 as the effective draws. The posterior means and the posterior standard errors are reported in Table $\ref{empiricalexpl1table1}$.
In this example, we try to examine whether the marginal effects of father's education $\left(\beta_{4}\right)$ and mother's education $\left(\beta_{5} \right)$ on the completion of high school can be ignored or not. Since the ${ \mbox{\boldmath${T}$}}\left(Data,{\mbox{\boldmath${\theta}$}}_{0}\right)$ does not have analytical expression, according to the Appendix (ref), we use the MCMC output to approximate the statistic. Further, in order to compare the statistic with the one in Li, Liu and Yu (2015) and the Bayes factor, we also report the $\widehat{\log BF_{10}}$ and the $\widehat{{\mbox{\boldmath${T}$}} }_{LLY}\left(Data,{\mbox{\boldmath${\theta}$}}_{0}\right)$ in the Table (ref). In this case, the log-likelihood has closed-form expression. Hence, the corresponding numerical standard error for each statistic is also reported in the Table $\ref{empiricalexpl1table2}$.
The result we obtained in the Table $\ref{empiricalexpl1table2}$ strongly prove the advantages of the proposed statistic. The 99.99 percentile of $ \chi^{2}\left(2\right)$ is 18.42. Both the $\widehat{{\mbox{\boldmath${T}$}}} \left(Data,{\mbox{\boldmath${\theta}$}}_{0}\right)$ and $\widehat{{ \mbox{\boldmath${T}$}}}_{LLY}\left(Data,{\mbox{\boldmath${\theta}$}} _{0}\right)$ are much larger than 18.42, which indicates that the null hypothesis is rejected under the $99.99\%$ probability level. Those results are consistent with the value of $\widehat{\log BF_{10}}$, which strongly supports the alternative hypothesis. Those three statistics all tell us that the marginal effect of the parents' education on the high school completion is not negligible. Further, they all have small numerical standard error compared with the corresponding values.
What's more from the table we can learn is that the proposed statistic takes much less time than the other two statistics. It takes around as half as the time $\widehat{{\mbox{\boldmath${T}$}}}_{LLY}\left(Data,{\mbox{\boldmath${ \theta}$}}_{0}\right)$ used and one fifth of the time Bayes facor used.
Stochastic volatility models are widely used in finance and economics. The financial leverage effect is very important and documented in many financial literature, see Black (1976). Following Yu (2005), the leverage effects SV model is defined as follows:
with
and $h_{0}=\mu$, where $r_{t}$ is the return at time $t$, $h_{t}$ the return volatility at period t. In this model, $\rho$ is the parameter indicating the leverage effect. When $\rho<0$, there is a negative relationship between the expected future volatility and the current return (Yu, 2005). In particular, volatility tends to rise in response to bad news but fall in response to good news (Black, 1976). Hence, we construct the hypothesis, $H_{0}:\rho=0$, to test whether the leverage effect exists or not.
In this example, we used two cases to illustrate how to use the proposed statistic. And further, the statistic is also compared with the statistic proposed by LLY, ${\mbox{\boldmath${T}$}}_{LLY}\left({\mbox{\boldmath${y}$}}, {\mbox{\boldmath${\theta}$}}_{0}\right)$ and Bayes factor. The derivation of the computation is given in the Appendix (ref).
In the first case, we use the data that consist of daily returns on Pound/Dollar exchange rates from 01/10/81 to 28/06/85 with sample size 945. The series $r_{t}$ is the daily mean-corrected returns. The R2OpenBUGS is used to estimate the model with the following priors for each parameter:
We draw 50,000 from the posterior distribution and discard the first 20,000 as build-in period. Then we store every 5th value of the remaining samples as effective observations. The estimation results are reported in Table $\ref{empiricalexpl2table1}$.
We aim to test whether there is leverage effect or not, hence the hypothesis problem is:
In Table $\ref{empiricalexpl2table2}$, we report the Bayes factor, the statistic in LLY, $ \widehat{{\mbox{\boldmath${T}$}}}_{LLY}\left({\mbox{\boldmath${y}$}},{ \mbox{\boldmath${\theta}$}}_{0}\right)$, and the proposed statistic, $ \widehat{{\mbox{\boldmath${T}$}}}\left({\mbox{\boldmath${y}$}},{ \mbox{\boldmath${\theta}$}}_{0}\right)$. The $\widehat{\log BF_{10}}$ strongly supports the null hypothesis, that is, there is not leverage effect. Meanwhile, since ${\mbox{\boldmath${T}$}}_{LLY}\left({ \mbox{\boldmath${y}$}},{\mbox{\boldmath${\theta}$}}_{0}\right)$ follows a $ \chi^{2}\left(1\right)$ distribution, the value of this statistic with a rather small NSE shows that it fails to reject the null hypothesis at the 95% probability level. For the porposed statistic, ${\mbox{\boldmath${T}$}} \left({\mbox{\boldmath${y}$}},{\mbox{\boldmath${\theta}$}}_{0}\right)-1 \overset{d}{\rightarrow}\chi^{2}\left(1\right)$, $\widehat{{ \mbox{\boldmath${T}$}}}\left({\mbox{\boldmath${y}$}},{\mbox{\boldmath${ \theta}$}}_{0}\right)-1$ is closed to $\widehat{{\mbox{\boldmath${T}$}}} _{LLY}\left({\mbox{\boldmath${y}$}},{\mbox{\boldmath${\theta}$}}_{0}\right)$ with smaller NSE. Thus it can not reject the null hypothesis under 95% probability level. Thus, the outcomes of all three statistics are consistent.
In the second case, the data we use is 1,822 daily returns of the Standard & Poor (S&P) 500 index, covering the period between January 3, 2005 and March 28, 2012. We use the same priors and similar method to estimate the model. And the parameter estimated are listed in Table $\ref{empiricalexpl2table3}$, which are quite different from the first case.
Again, the three statistics are reported in Table $\ref{empiricalexpl2table4}$. Contrary to first case, all the statistics strongly support the null hypotheses, that is, there is leverage effect in the data. For $\widehat{{\mbox{\boldmath${T}$}}} \left({\mbox{\boldmath${y}$}},{\mbox{\boldmath${\theta}$}}_{0}\right)$ and $ \widehat{{\mbox{\boldmath${T}$}}}_{LLY}\left({\mbox{\boldmath${y}$}},{ \mbox{\boldmath${\theta}$}}_{0}\right)$, they all reject the null hypothesis under the 99% probability level. At the same time, the $\widehat{ \log BF_{10}}$ also strongly supports the alternative hypothesis. Therefore, the results of all three statistics are consistent.
In this paper, a new $\chi^2$-type Bayesian test statistic is proposed to test a point null hypothesis for latent variable models. The new statistic can be explained as Bayesian version of Wald test. Compared with existing literature, the proposed test statistic has achieved several important advantages, hence, can appeal many practical applications. First, for latent variable models, it is only by-product of posterior outputs, hence, is very easy to compute, not require additional computational efforts after the model is estimated using MCMC techniques. Second, it is well-defined under improper prior distributions and avoids Jeffrey-Lindley's paradox. Third, it's asymptotic distribution is pivotal so that the threshold values can be easily obtained from the asymptotic chi-squared distribution. Through Monte Carlo studies, it can be shown that the test power is almost equivalent to Wald test, but better than the test statistic by Li, Liu and Yu (2014). If the observed likelihood doesn't have analytical form, Wald test statistic is very difficult to be applied, but, the proposed test statistic is also easy to implement.
The Bayes factor and the Bayesian test statistic with Li, Liu and Yu (2015) need to be computed out and make an empirical comparison
\vskip 0.5cm