Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
92,268 characters · 13 sections · 40 citation commands
MinP Score Tests with an Inequality Constrained Parameter Space
Score tests were originally proposed by rao48. They are also known as Lagrange multiplier (LM) tests due to the work of ajssd1958 and silvey59. Score tests have the advantage of requiring only estimate of the model restricted by the null hypothesis, which often is much simpler than models defined under the alternative hypothesis. This advantage becomes more attractive when models defined under the alternative hypothesis are subject to inequality constraints that complicate model estimation and statistical inference. One-sided tests are more appropriate than two-sided tests when dealing with inequality constraint (see, e.g., ad01, frazak09, ketz18 and canipera18 and the early literature therein). One-sided score tests have been studied by many authors, including ghm82, wf89a,wf89b, leeking93, ss95, kingwu97, demsen98 and ad01. However, existing one-sided score tests are designed for jointly testing on the parameter of interest. For example, they can be used for testing the null hypothesis of no autoregressive conditional heteroskedasticity (ARCH) effects by assessing whether all the ARCH parameters are jointly zero.
This paper proposes what we term `MinP score' tests, where the test statistic is the minimum of a set of $p$-values. The minimand includes the $ p $-values for testing individual elements of parameter of interest using individual scores. Importantly, the set may be extended to include a $p$ -value of existing score tests. We shall show that, compared with existing tests, our tests have a competitive power. In addition, our test can have the further benefit of allowing, under some regularity conditions, for simultaneously testing individual components of the parameter of interest. For instance, when testing for ARCH the practitioner is able to detect which individual ARCH parameters are non-zero. In contrast, rejection of non-ARCH effects by existing tests does not reveal the evidence of which of the individual ARCH parameters are non-zero. This added benefit is appealing in particular as it allows for model identification within nested models through testing, without estimating candidate models except the model defined under the null hypothesis. To the best of our knowledge, no existing tests allow for model identification using only the restricted estimate under the null hypothesis.
The competitiveness of our proposed tests with regard to joint testing on the parameter of interest is supported by the admissibility property we establish. Our admissibility result is applicable to a sort of combined tests for one-sided testing that includes the extended MaxT (EMaxT) tests recently proposed by lu16. With regard to testing individual components of the parameter of interest, it is important to take into account multiplicity of Type I error roazwo10. We adopt the family wise error rate (FWER), i.e. the probability of wrongly rejecting at least one true individual hypotheses, for the Type I error control and show that the FWER control can be achieved under certain condition.
We illustrate our tests using three examples, all relevant in applications: linear regression models, ARCH models and random coefficient models. In the linear regression model example, we show that our tests can be used for selecting variables that only exert non-negative marginal effects on the dependent variable, as well as jointly testing if any such effects exist. In this application, the FWER is asymptotically controlled under conditions required for usual Type I error control in joint testing. Variable selection through testing in linear regression models has gained much interest in the recent literature (see, e.g., mckqia15 and related discussion papers). Our tests add contributions to the literature. In the context of ARCH (and random coefficient) models, we show our tests may be used for testing ARCH (and random coefficient) effects and selecting ARCH (and random coefficient) models if non-ARCH (non-random coefficient) effect is rejected. We provide sufficient conditions that lead to the asymptotic control of FWER in ARCH (and random coefficient) models. Simulation studies are carried out to compare finite sample performance of our tests with existing one-sided score tests in terms of joint testing, as well as the performance in model identification.
The remainder of the paper is organized as follows. In the next section we set up a general framework under maximisation of an objective function and introduce some assumptions. We provide a result of the limiting normal distribution for function of scores that serves as a basis for constructing our tests. Section (ref) presents our MinP score tests and focus on illustrating them for joint testing. Section (ref) studies multiple testing by our tests. Section (ref) provides three applications and conducts the simulation study. Concluding remarks are made in Section (ref). Technical proofs are provided in the appendix.
The following notation is used throughout the paper. Denote by $ \xrightarrow{p}$ convergence in probability and by $\xrightarrow{d}$ convergence in distribution. Let $\mathcal{C}_{k}=[0,\infty )^{k}$ and $ \mathcal{C}_{k}^{+}=(0,\infty )^{k}$. Denote by $I_{k}$ the $k$-dimension identity matrix and $0_{k}$ the $k\times 1$ vector of $0$ (we may simply use $I$ and $0$ without indicating the dimension if no confusion is deemed to arise). For a matrix $B$, $B_{ij}$ indicates its $(i,j)$-th element. Similarly, for a vector $s$, $s_{i}$ indicates its $i$-th element. A parameter symbol used in a subscript indicates the partitioned component(s) of a vector or a matrix corresponding to parameter involved. For example, $ s=(s_{\gamma }:s_{\psi })$ indicates the partition of $s$ according to the parameter $(\gamma :\psi )$. $\mathcal{J}_{\gamma \psi }$ is the left top corner block of $\mathcal{J}$ corresponding to $\gamma \ $and\ $\psi $. When an inequality operation is applied to vector(s) it means an elementwise relationship. Finally, $\mathbbm{1}(\cdot )$ denotes the usual indicator function and $\left\Vert \cdot \right\Vert $ denotes the Euclidean norm.
Let $Y_{T}$ be the data at sample size $T$ and $l_{T}(\theta )$ be the estimator objective function (e.g., the likelihood function) that depends on $Y_{T}$ and the $q$ dimensional parameter $\theta $. With $\theta $ partitioned as $\theta =(\gamma ^{\prime },\psi ^{\prime })^{\prime }$, interest is in inference on the $k\times 1$ vector $\gamma $, with $\psi $ being a vector of nuisance parameters. We assume that a subset of $\gamma $ may lie on the boundary of the parameter space, with all the remaining elements of $\gamma $ being interior points. Specifically, with the partition $\gamma =(\gamma _{1}^{\prime },\gamma _{2}^{\prime })^{\prime }$, we assume that the first subset $ \gamma _{1}\in \mathcal{C}_{k_{1}}$, $\gamma _{2}\in \mathcal{C} _{k-k_{1}}^{+}$ and $\psi \in \Psi \subset \mathbb{R}^{q-k}$ for some $1\leq k_{1}\leq k\leq q<\infty $. That is, $\gamma _{1}$ may be on the boundary of the cone $\mathcal{C}_{k_{1}}$ while $\gamma _{2}$ is assumed to be an interior point of the cone $\mathcal{C}_{k-k_{1}}$. The parameter space is denoted by $\Theta (k_{1})=\mathcal{C}_{k_{1}}\times \mathcal{C} _{k-k_{1}}^{+}\times \Psi $, and $\Theta (k)$ implies $\mathcal{C} _{k-k_{1}}^{+}=\emptyset $ (with $k_{1}=k$).
We consider two inference problems. The first, which we label `joint testing' or `global testing' interchangeably, concerns testing the null hypothesis $\mathsf{H}_{0}:\gamma =0$. This testing problem is complicated by the fact that under the null $\gamma $ is on the boundary of the parameter space. The second is model selection, i.e. identification of the non-zero elements of $\gamma $ once the global null hypothesis is rejected. This inference problem requires to deal with multiple testing on the elements of $\gamma $.
Let $\theta _{0}=(\gamma _{01}^{\prime },\gamma _{02}^{\prime },\psi _{0}^{\prime })^{\prime }$ denote the true value of $\theta \in \Theta (k_{1})$. Consider the parameter space restricted by
Let $\tilde{\theta}(k_{1})$ be the restricted parameter estimator, which satisfies
We define $\bar{\theta}_{0}=(0_{k_{1}}^{\prime },\bar{\gamma}_{02}^{\prime }, \bar{\psi}_{0}^{\prime })^{\prime }$ as $plim_{T\rightarrow \infty }\tilde{ \theta}(k_{1})\in \bar{\Theta}(k_{1})$ and denote this by the \textquotedblleft pseudo-true\textquotedblright\ value. When $\gamma _{01}=0$ , clearly the pseudo-true value $\bar{\theta}_{0}=\theta _{0}$, the true value. However, when $\gamma _{01}\neq 0$, these differ as will be exploited below.
Denote by $\mathcal{N}(\theta _{0},\epsilon )=\{\theta :\left\Vert \theta -\theta _{0}\right\Vert =\epsilon <aT^{-1/2}\}$ for some $a>0$, such that $ \epsilon =O(T^{-1/2})$. That is, $\mathcal{N}(\theta _{0},\epsilon )$ is an open sphere neighbourhood centered at $\theta _{0}$ with the radius $ \epsilon =O(T^{-1/2})$. Let $\Theta _{\epsilon }(\theta _{0},k_{1})=\mathcal{ N}(\theta _{0},\epsilon )\cap \Theta (k_{1})$ and $\bar{\Theta}_{\epsilon }(\theta _{0},k_{1})=\mathcal{N}(\theta _{0},\epsilon )\cap \bar{\Theta} (k_{1})$. Denote furthermore by $s_{T}(\theta )=\partial l_{T}(\theta )/\partial \theta $ the score function and by $\mathcal{J}_{T}(\theta )=-\partial ^{2}l_{T}(\theta )/\partial \theta \partial \theta ^{\prime }$ the negative Hessian matrix. We make the following assumption throughout.
We now introduce a result which will be used throughout. Let
and
Moreover, let $U_{T,\gamma _{1}}(\theta )$ be the sub-vector of $ U_{T}(\theta )$ and $\mathcal{G}_{T,\gamma _{1}\gamma _{1}}(\theta )$ be the sub-block matrix of $\mathcal{G}_{T}(\theta )$ corresponding to $\gamma _{1}$ . Because $\mathcal{\psi }$ is not subject to constraint in obtaining $ \tilde{\theta}(k_{1})$ it implies $s_{T,\psi }(\tilde{\theta}(k_{1}))=0$. Also, $s_{T,\gamma _{2}}(\tilde{\theta}(k_{1}))=0$ as $\gamma _{2}\in \mathcal{C}_{k-k_{1}}^{+}$ is assumed to be interior point of $\mathcal{C} _{k-k_{1}}$. Therefore,
where $\mathcal{J}_{T,\gamma _{1}\gamma _{1}}^{-1}(\theta )$ is the sub-block matrix of $\mathcal{J}_{T}^{-1}(\theta )$ corresponding to $\gamma _{1}$. The following lemma holds for $\theta _{0}$ in a neighbourhood of $ \bar{\theta}_{0}$.
Our score tests for both global and multiple testings are constructed based on $U_{T,\gamma }(\tilde{\theta}(k))$, where $\tilde{\theta}(k)$ is the constrained estimator in $\bar{\Theta}(k)=\{0\}^{k}\times \Psi $. We do not estimate $\tilde{\theta}(k_{1})$ for $k_{1}<k$. Nevertheless, Lemma (ref) provides the theoretical basis to study the FWER in multiple testing where the true null hypotheses may be only some of the elements of $\gamma $ being zero. To illustrate let us consider the ARCH(3) model:
where $u_{n}$ is assumed to be independently drawn from a random distribution with the mean $0$ and variance $1$ and $\theta =(\gamma ^{\prime },\omega )^{\prime }=(\alpha _{1},\alpha _{2},\alpha _{3},\omega )^{\prime }$. Our score tests are constructed based on $\tilde{\theta} (k)=(0,0,0,\tilde{\omega})^{\prime }$ with $\tilde{\gamma}=\tilde{\gamma} _{1}=(0,0,0)^{\prime }$, $\tilde{\gamma}_{2}=\emptyset $ and $\tilde{\psi}= \tilde{\omega}$. In the global testing $\mathsf{H}_{0}:\gamma =0$, Lemma (ref) provides the null distribution of $T^{1/2}U_{T,\gamma }(\tilde{\theta} (k))$ for asymptotically controlling the Type I error. In the multiple testing $\mathsf{H}_{0i}:\alpha _{i}=0$, $i\in \{1,2,3\}$, suppose the true null is $\alpha _{01}=0$, $\alpha _{02}=0$ and $\alpha _{03}>0$. Lemma (ref) provides the asymptotic distribution of $T^{1/2}U_{T,\gamma _{1}}( \tilde{\theta}(k_{1}))$, where $\gamma _{1}=(\alpha _{1},\alpha _{2})^{\prime }$ and $\gamma _{2}=\alpha _{3}$. Clearly, this distribution depends on not only $\alpha _{01}=0$ and $\alpha _{02}=0$, but $\alpha _{03}>0$ and $\omega _{0}$. The FWER control requires the probability of rejecting $\alpha _{1}=0$ and/or $\alpha _{2}=0$ is bounded by a designated level. Because the researcher does not know the true $\theta _{0}$, the FWER should be controlled at any possible $\theta _{0}$. We shall study our tests in ARCH models in more details in Section (ref).
In what follows, to simplify notation, $\tilde{\theta}$ should be understood as $\tilde{\theta}(k)$; if needed, $\tilde{\theta}(k_{1})$ is used to stress result holds for $k_{1}\leq k$.
The joint testing concerns all elements of $\gamma $ being $0$ or not. The global null and alternative hypotheses of interest are
Lemma (ref) above implies $U_{T,\gamma }(\tilde{\theta})=\gamma _{0}+o_{p}(1)$. As $\gamma _{0}\in \mathcal{C}_{k}\subset \Theta (k)$, it is desired for the test statistic constructed based on $U_{T,\gamma }(\tilde{ \theta})$ to reflect the cone restriction $\Theta (k)$. One may construct one-sided tests based on the test statistic ss05
where
It then follows from Lemma (ref) that under $\mathsf{H}_{0}$
where $Z_{\theta _{0}}$ is the $k$-dimension multivariate normal random variable with the mean $0$ and variance $\mathcal{G}_{\gamma \gamma }(\theta _{0})$.
Alternatively, one may construct joint $t$ type tests based on the test statistic
where
lu13. Under $\mathsf{H}_{0}$, $ \tilde{t}_{t}\xrightarrow{d}N(0,1)$. We suggest that $d_{T}$ be computed by Algorithm 1 in lu16 using the covariance $\mathcal{G}_{T,\gamma \gamma }(\tilde{\theta})$, so that the resulting $t$ test has the optimality of asymptotically maximizing the minimum power within the class of tests based on $\tilde{t}_{t}$.
We now discuss our proposed three MinP score tests. Let $U_{T,\gamma }=(U_{T,\gamma ,1},...,U_{T,\gamma ,k})^{\prime }$ and $\tilde{t}=(\tilde{t} _{1},...,\tilde{t}_{k})^{\prime }$, where $\tilde{t}_{i}=U_{T,\gamma ,i}( \tilde{\theta})/\sqrt{\mathcal{G}_{T,\gamma \gamma ,ii}(\tilde{\theta})}$ and $\mathcal{G}_{T,\gamma \gamma ,ii}(\tilde{\theta})$ is the $(i,i)$ th-element of $\mathcal{G}_{T,\gamma \gamma }(\tilde{\theta})$. Let $\tilde{p }_{i}=1-F_{T,i}(\tilde{t}_{i})$, $i\in \{1,..,k\}$, where $F_{T}(\cdot )$ is the cumulative distribution function (CDF) of $\tilde{t}$ under $\mathsf{H} _{0}$ and $F_{T,i}(\cdot )$ indicates the marginal distribution corresponding to $\tilde{t}_{i}$. Let $F_{T,c}(\cdot )$ and $F_{T,t}(\cdot )$ be the CDFs under $\mathsf{H}_{0}$ corresponding to $\tilde{t}_{c}$ and $ \tilde{t}_{t}$, respectively. Finally, let $\tilde{p}_{c}=1-F_{T,c}\{\tilde{t }_{c}\}$ and $\tilde{p}_{t}=1-F_{T,t}\{\tilde{t}_{t}\}$.
The first proposed MinP score test is based on the test statistic
The test statistic for the other two proposed MinP score tests is given by
where $\tilde{p}_{g}$ takes the value of either $\tilde{p}_{c}$ or $\tilde{p} _{t}$. We refer to the score test based on $\tilde{p}_{m1}$ as the `MinP-s' test, and the score tests based on $\tilde{p}_{m2}$ using $\tilde{p}_{c}$ and $\tilde{p}_{t}$ as `MinP-sc' and `MinP-st' tests, respectively.
Let $\tilde{c}_{mj}(\alpha )$ be the critical value at the $\alpha $ significance level such that for $\theta \in \bar{\Theta}(k)$
We then reject $\mathsf{H}_{0}$ if $\tilde{p}_{mj}\leq \tilde{c}_{mj}(\alpha )$; otherwise $\mathsf{H}_{0}$ is accepted. Algorithm (ref) below provides an algorithm for numerically computing $\tilde{p}_{c}$, $\tilde{p} _{t}$, $\tilde{p}_{i}$, $i=1,...,k$, and $\tilde{c}_{mj}(\alpha )$ via a valid bootstrap method. Examples of bootstrap implementations are given in Section 5 below.
We establish the asymptotically locally admissibility property of our tests. ad96 studied the admissibility of one-sided LR tests, Wald tests and LM tests by showing the optimal power of the tests for a given weighting function of the parameter space under the alternative hypothesis. A similar technique was adopted by andplo95, andmorsto06 and chehanjan09 for the study of admissibility of rotation invariant tests. marden82 (marden82, marden85) studied the admissibility of tests that combine tests of the same kind but applied to independent samples. His admissibility results are based on the principle of Bayes tests. Our proposed tests are a combination of different kinds of tests applied to the same sample. We derive our admissibility result by extending the work of birnbaum55 and stein56 (see, e.g., Theorem 6.7.1 of lehrom05) to the case of one-sided testing.
Let $\varphi (Y_{T})$ be a $\{0,1\}$-valued test for the global hypothesis $\mathsf{H}_{0}$. If $\lim_{T\rightarrow \infty }\sup_{\theta \in \bar{\Theta}_{\epsilon }(\bar{\theta}_{0},k)}E_{\theta }(\varphi (Y_{T}))\leq \alpha $, then the test $\varphi (Y_{T})$ is said to be a locally asymptotically level $\alpha $ test. Denote by $\varphi _{mj}(Y_{T})=\mathbbm{1}(\tilde{p}_{mj}\leq \tilde{c}_{mj}(\alpha ))$ the test function for the proposed MinP score tests. Let $\varphi _{mj}^{\prime }(Y_{T})$ be a locally asymptotically distinct test from $\varphi _{mj}(Y_{T})$ such that for each $\theta \in \Theta _{\epsilon }(\bar{\theta} _{0},k)$
Note that a distinct test $\varphi _{mj}^{\prime }(Y_{T})$ defined in (ref) implies that $\varphi _{mj}^{\prime }(Y_{T})$ cannot take the same values as $\varphi _{mj}(Y_{T})$ asymptotically, and the limiting acceptance region of the test $\varphi _{mj}(Y_{T})$ is different from that of $\varphi _{mj}^{\prime }(Y_{T})$.
Suppose the limiting acceptance region of $\varphi _{mj}(Y_{T})$ is a closed, convex and lower set $E$ such that if $u\leq v$ and $v\in E$, then $ u\in E$. Define a halfspace by $W_{e,\delta }=\{v:e^{\prime }\mathcal{G} _{\gamma \gamma }^{-1}(\theta _{0})v<\delta \}$, where $v\in \mathbb{R}^{k}$ , $e\in \mathcal{C}\backslash \{\gamma =0\}$ and $\delta \in \mathbb{R}$. Let $W=\{v:\bigcap_{e\in \mathcal{C}\backslash \{\gamma =0\}}W_{e,\delta }\}$ . Denote by $W^{\ast }(E)\subseteq W$ such that $E\subseteq W^{\ast }(E)\subseteq W^{\prime }(E)$ for all $W^{\prime }(E)$ that satisfies $ E\subseteq W^{\prime }(E)\subseteq W$. Therefore, $W^{\ast }(E)$ can be viewed as the intersection of the halfspaces $W_{e,\delta }\subseteq W$ for all $e\in \mathcal{C}\backslash \{\gamma =0\}$ that are tangent with $E$.
The limiting acceptance region of the MinP-s test is $E_{s}(c_{s})= \{v:v<c_{s}\}$, $c_{s}>0$. Consider the $k$ halfspaces $W_{e,\delta }$ with $ e$ and $v$ being the column vector of the identity matrix. Then $W^{\ast }(E_{s})=E_{s}$. For joint $t$ tests based on $\tilde{t}_{t}$ the limiting acceptance region is the halfspace defined by $E_{t}(c_{t})=\{v:d^{\prime } \mathcal{G}_{\gamma \gamma }^{-1}(\theta _{0})v<c_{t},d\in \mathcal{C} ^{+}\backslash \{\gamma =0\}$, $c_{t}>0$. Because that $d$ is a column vector with positive elements that $E_{t}$ cuts through the top corner of $ E_{s}$. Thus, for the MinP-st test the limiting acceptance region is $ E_{st}(c)=\{v:E_{s}(c)\cap E_{t}(c)\}$, $c>0$, and $W^{\ast }(E_{st})=E_{st}$ . For tests based on $\tilde{t}_{c}$ the limiting acceptance region is $ E_{c}(c_{c})=\{v:v^{\prime }\mathcal{G}_{\gamma \gamma }^{-1}(\theta _{0})v<c_{c}\}$, $c_{c}>0$. The limiting acceptance region of the MinP-sc test is then $E_{sc}=\{v:E_{s}(c_{s})\cap E_{c}(c_{c})\}$, where $c_{s}$ and $c_{c}$ are positive constants such that the overall Type I error control for testing $\mathsf{H}_{0}$ by MinP-sc tests is achieved. In this case it may be that $W^{\ast }(E_{sc})=E_{sc}$ or $W^{\ast }(E_{sc})=E_{s}$ (see Figure (ref) for illustration).
The MinP score tests described previously may be used for simultaneously testing on individual components of $\gamma =(\gamma _{1},...\gamma _{k})^{\prime }$. That is, consider testing multiple hypotheses
When a set of hypotheses is tested simultaneously it is important to control multiplicity of Type I errors as otherwise the probability of wrongly rejecting one hypothesis increases as the number of true hypotheses increases. The family-wise error rate (FWER) is a widely used Type I error probability control in multiple testing. Let
be the set containing the indices of true $\mathsf{H}_{0i}$. The FWER is the probability of rejecting at least one $\mathsf{H}_{0i}$, $i\in K_{0}$ under $ \bigcap_{i\in K_{0}}\mathsf{H}_{0i}$. We shall show that the FWER control is bounded by $\alpha $ for MinP score tests under some additional assumptions.
To illustrate, consider the following simple example. Suppose the sample $ \{y_{n},n=1,...,T\}$ is independently randomly drawn from a multivariate normal distribution with mean $\gamma $ and known variance $\Omega $. Then, the objective (log-likelihood)\ function is
As $\Omega $ is known, $s_{T}(\gamma )=\sum\nolimits_{n=1}^{T}\Omega ^{-1}(y_{n}-\gamma )$, $\mathcal{J}_{T}(\gamma )=T\Omega ^{-1}$ and $ U_{T}(\gamma )=T^{-1}\sum\nolimits_{n=1}^{T}(y_{n}-\gamma )$. Therefore, under $\mathsf{H}_{0}$, $U_{T}(\tilde{\gamma})=T^{-1}\sum \nolimits_{n=1}^{T}y_{n}$ and $T^{1/2}U_{T}(\tilde{\gamma})\sim N(\gamma ,\Omega )$. The MinP score test based on $\tilde{p}_{m1}$ is the same as MinP tests based on the sample average. In this case the joint distribution of $U_{T,i}(\tilde{\gamma})$, $i\in J\subset K$, for a subset $J$ under $ \bigcap_{i\in J}\mathsf{H}_{0i}$ is $N(0,\Omega _{J})$, where $\Omega _{J}$ contains the blocks of $\Omega $ corresponding to $J$. Obviously, the joint distribution of $U_{T,i}(\tilde{\gamma})$ is not affected by the truth or falsehood of the remaining $\mathsf{H}_{0i}$, $i\in K\setminus J$. This is an example of what is known as the subset pivotality condition (see, e.g., wesyou93). As a result, the stepdown procedure of MinP-s score tests based on $\tilde{p}_{m1}$ allows for the control of the FWER. However, this may not be the case in our set up as the covariance $\mathcal{G}_{\gamma \gamma }(\theta _{0})$ in (ref) may depend on $\gamma _{0}$. By setting $\gamma =0$ as the researcher would usually do in joint testing of the global null $\mathsf{H}_{0}$ is normally referred to as the weak control of FWER in the literature of multiple testing, which does not guarantee the control of the FWER in multiple testing (see, e.g., romwol05b).
Let $\tilde{c}_{J}(\alpha )$ be the critical value at the $\alpha $ significance level such that under $\bigcap_{i\in J}\mathsf{H}_{0i}$, $ J\subset K$,
where the notation $\tilde{p}_{i}(\tilde{\theta}(k))$ is used to stress $ \tilde{p}_{i}(\tilde{\theta}(k))$ is based on the estimate $\tilde{\theta} (k) $, not the estimate by restricting $\tilde{\gamma}_{i}=0$, $i\in J$. Note that if $J=K$ then $\tilde{c}_{J}(\alpha )$ is equal to $\tilde{c} _{m1}(\alpha )$ in the global testing.
MinP score tests for multiple testing are as follows. If the global hypothesis $\mathsf{H}_{0}$ is rejected then order $\tilde{p}_{i}$, $ i=1,...,k$, as follows:
and let $\mathsf{H}_{(01)}$, ..., $\mathsf{H}_{(0k)}$ be the corresponding hypotheses. Reject $\mathsf{H}_{(01)}$ if and only if $\tilde{p}_{(1)}\leq \tilde{c}_{mj}(\alpha )$. If $\mathsf{H}_{(01)}$ is rejected one would then continue to the stepdown procedure for the remaining $\mathsf{H}_{(02)}$, ..., $\mathsf{H}_{(0k)}$, as shown in Algorithm (ref) below.
Without loss of generality let the first $k_{0}$ elements of $\gamma _{0}$ containing the true $\mathsf{H}_{0i}:\gamma _{0i}=0$, i.e., $ K_{0}=\{1,...,k_{0}\}$. When $K_{0}=\emptyset $, the FWER is set to $0$ by default. When $K_{0}\neq \emptyset $, let $c_{K_{0}}(\alpha )$ be the critical value such that
We show in Theorem (ref) that, under the assumptions of Theorem (ref), if also Assumption (ref) holds, the FWER control for MinP score tests is bounded by $\alpha $.
Multiple tests of the hypotheses (ref) allow for model selection through testing. Let
where $\tilde{c}_{K_{1}}(\alpha )=\tilde{c}_{mj}(\alpha )$, be the set of the indices of $\mathsf{H}_{0i}$ that is rejected in the above stepdown procedure.
We implement our tests for applications in linear regression models, ARCH models and random coefficient models. Simulation studies are also conducted in order to examine the finite sample performance of our tests. The number of bootstrap repetitions $B$ in implementing Algorithm (ref) is set to $ 999$. The number of Monte Carlo replications in our simulation studies is set to $2000$. All the reported results are based on the nominal $0.05$ size. In all three models considered below $x_{n}=(x_{n}^{d},1)^{\prime }$, where $x_{n}^{d}$ is a scalar, $\beta =(1,1)^{\prime }$ and we observe the sample $\{y_{n},z_{n},x_{n}^{d},$ $n=1,...T\}$. Let $\tilde{\epsilon}_{n}$ be the regression residuals under $\mathsf{H}_{0}$. When implemented, the bootstrap is based on generating the bootstrap shocks $\epsilon _{n}^{\ast }$ by randomly drawn with replacement from $\tilde{\epsilon}_{n}$ and $ y_{n}^{\ast }$ from the regression model under $\mathsf{H}_{0}$ using $ \tilde{\theta}(k)$, $\epsilon _{n}^{\ast }$ and fixed covariates. Note that, in some applications, one may need to center the residuals in order to have bootstrap shocks with (conditionally on the original data) exact zero mean.
In all applications we report size and power with respect to global tests of $\mathsf{H}_{0}$. In terms of multiple testing of $\mathsf{H}_{0i}$, $i\in K$ , we report the FWER and the probability of rejecting individual $\mathsf{H} _{0i}$ by the stepdown multiple testing procedure described in Algorithm (ref).
Consider the model
where $\epsilon _{n}$ is assumed to be independently drawn from a random distribution with the mean $0$ and variance $\sigma ^{2}<\infty $, $z_{n}$ and $x_{n}$ are the $k\times 1$ and $(q-k)\times 1$ column vectors, respectively, and $\gamma \in \mathcal{C}\subset \mathbb{R}^{k}$ and $\beta \in \mathbb{R}^{q-k-1}$ with the last element being the intercept parameter. The objective function $l_{T}(\theta )=-\sum\nolimits_{n=1}^{T}(y_{n}-z_{n}^{\prime }\gamma -x_{n}^{\prime }\beta )^{2}$, where $\theta =(\gamma ^{\prime },\beta ^{\prime })^{\prime }$. Let $ \tilde{\epsilon}_{n}$ be the estimated residuals based on the constrained Ordinary Least Squares (OLS) estimate $\tilde{\theta}=(0^{\prime },\tilde{ \beta}^{\prime })^{\prime }$. To check whether Assumption (ref) holds in this example, we can make use of the following proposition.
In our simulation study we generate $ X^{d}=(z_{1},...,z_{T},x_{1}^{d},...,x_{T}^{d})^{\prime }$ as i.i.d. from $ N(0,\Sigma )$, where
$\rho $ is constant and $1_{k}$ is a $k\times 1$ vector of 1. Thus $ E(T^{-1}X^{d\prime }X^{d})^{-1}=\Sigma ^{-1}$, which has the structure of a correlation matrix with all the off-diagonal elements being $\rho $.
Tables (ref)--(ref) report the estimated probabilities of rejecting $\mathsf{H}_{0}$, FWER and $\mathsf{H}_{0i}$ with $\epsilon _{n}$ being i.i.d. (and independent of $X^{d}$) and following either the $N(0,1)$ distribution or the Student's $t$ distribution with 5 degrees of freedom; $k$ is set to $2$. Notice that for this data generating process, Assumptions (ref), (ref) and (ref) hold. Assumption (ref) also holds by Proposition (ref). Results show that MinP-sc and MinP-st score tests have competitive finite-sample performance in both size and power when compared with score tests based on $\tilde{t}_{c}$ and $\tilde{t}_{t}$ in terms of global testing. However, MinP-sc and MinP-st score tests have an added benefit of testing individual components of $\gamma $ with the FWER control. While MinP-s score tests may perform slightly better in multiple testing, they can have inferior power in global testing of $\mathsf{H}_{0}$, for example, in the case of $\rho =-0.45$.
Consider the regression model with ARCH disturbances
and $T\mathcal{G}_{T,\gamma \gamma }(\tilde{\theta})=I_{k}\mathbf{+}o_{p}(1)$ . Assumptions (ref), (ref) and (ref) are straightforward to verify. Assumption (ref) follows under a sufficient condition as stated in the following proposition.
We compare the performance of MinP score tests with that of score tests proposed in demsen98 and leeking93. The score test proposed in demsen98 is asymptotically equivalent to the score test based on $ \tilde{t}_{c}$ that has the limiting null distribution $\sum_{i=0}^{k}2^{-q} \binom{k}{i}\chi _{i}^{2}$. The score test proposed in leeking93 is asymptotically equivalent to the score test based on $\tilde{t}_{t}$. We generate $x_{n}=(x_{n}^{d},1)^{\prime }$, $x_{n}^{d}=0.8x_{n-1}^{d}+e_{n}$, $ e_{n}\sim i.i.d.N(0,4)$, and $u_{n}$ from either the standard normal distribution or the Student's $t$ with $5$ degrees of freedom. We add a burn-in period of 1000 observations for $\epsilon _{n}$, which are discarded in estimation. We adopt the constrained OLS estimate $\tilde{\theta} =(0^{\prime },\tilde{\beta}^{\prime },\tilde{\sigma}^{2})^{\prime }$.
Tables (ref) and (ref) report the estimated probabilities of rejecting $\mathsf{H}_{0}$, FWER and $\mathsf{H}_{0i}$. The results show that the global powers of MinP tests are competitive with those of demsen98's and leeking93's tests. Although the FWER control by our MinP tests may not always be guaranteed, the probability of identifying the false $\mathsf{H}_{0i}$ increases towards $1$ as sample size increases. {1.8pt}
Consider the model
where the random coefficient $\xi _{n}$ has the form
and $\xi $ and $\beta $ are a column vector of a fixed coefficient with the dimension of $k$ and $(q-k-1)$, respectively. The random variables $\epsilon _{n}\in \mathbb{R}$ and $\eta _{n}\in \mathbb{R}^{k}$ are unobserved independent errors across individual elements of $\eta _{n}$, between $ \epsilon _{n}$ and $\eta _{n}$, as well as across $n=1,...,T$ that satisfy $ E\epsilon _{n}=0$, $E\epsilon _{n}^{2}=\sigma _{\epsilon }^{2}$, $E(\eta _{n}|x_{n})=0$ and $E(\eta _{n}\eta _{n}^{\prime }|x_{n})=diag(\gamma )$, where $\gamma =(\sigma _{\eta ,1}^{2},...,\sigma _{\eta ,k}^{2})^{\prime }$.
The model may be rewritten as
where
Consider the quasi-log-likelihood function under normality as $l_{T}(\theta )\propto -\frac{1}{2}\sum_{n=1}^{T}\{\log \sigma _{\varpi _{n}}^{2}+\frac{ \varpi _{n}^{2}}{\sigma _{\varpi _{n}}^{2}}\}$, where $\sigma _{\varpi _{n}}^{2}=\gamma ^{\prime }z_{n}^{2}+\sigma _{\epsilon }^{2}$ and $ z_{n}^{2}=(z_{n,1}^{2},...,z_{n,k}^{2})^{\prime }$. Let $\lambda =(\gamma ^{\prime },\sigma _{\epsilon }^{2})^{\prime }$, $\psi =(\xi ^{\prime },\beta ^{\prime })^{\prime }$, $\theta =(\lambda ^{\prime },\psi ^{\prime })^{\prime }$ and $d_{n}=((z_{n}^{2})^{\prime },1)^{\prime }$. Since $ E(z_{n}^{\prime },x_{n}^{\prime })\varpi _{n}=0$ we have $T^{-1}\mathcal{J} _{T,\lambda \psi }(\theta )=\sum\nolimits_{n=1}^{T}(2\sigma _{\varpi _{n}}^{4})^{-1}d_{n}(z_{n}^{\prime },x_{n}^{\prime })\varpi _{n}=o_{p}(1)$ and $\mathcal{G}_{\lambda \lambda }(\theta )$ only involves $\mathcal{J} _{\lambda \lambda }(\theta )$ and $\mathcal{V}_{\lambda \lambda }(\theta )$. So let's consider $U_{T,\lambda }(\theta )=\mathcal{J}_{T,\lambda \lambda }^{-1}(\theta )s_{T,\lambda }(\theta )$ for constructing our tests with $ s_{T,\lambda }(\theta )=\sum_{n=1}^{T}\frac{\varpi _{n}^{2}-\sigma _{\varpi _{n}}^{2}}{2\sigma _{\varpi _{n}}^{4}}d_{n}$ and $\mathcal{J}_{T,\lambda \lambda }(\theta )=\sum\nolimits_{n=1}^{T}\frac{2\varpi _{n}^{2}-\sigma _{\varpi _{n}}^{2}}{2\sigma _{\varpi _{n}}^{6}}d_{n}d_{n}^{\prime }$. Hence, $U_{T,\gamma _{1}}(\tilde{\theta}(k_{1}))=\mathcal{J}_{T,\gamma _{1}\gamma _{1}}^{-1}(\tilde{\theta}(k_{1}))s_{T,\gamma _{1}}(\tilde{\theta}(k_{1}))$ and $\mathcal{G}_{\gamma _{1}\gamma _{1}}(\theta )$ is the upper left block corresponding to $\mathcal{G}_{\lambda \lambda }(\theta )$. Without loss of generality let $k_{0}=k_{1}$, so $\mathcal{G}_{\gamma _{1}\gamma _{1}}(\theta _{0})$ corresponds to the true $\gamma _{0i}=0$.
Under the global $\mathsf{H}_{0}$ we have $\tilde{\gamma}=0$, $\tilde{\sigma} _{\varpi _{n}}^{2}=\tilde{\sigma}_{\epsilon }^{2}$ and $E\varpi _{n}^{2}=\sigma _{\varpi _{n}}^{2}$, hence $T^{-1}\mathcal{J}_{T,\lambda \lambda }(\tilde{\theta})=(2\tilde{\sigma}_{\epsilon }^{4})^{-1}T^{-1}\sum\nolimits_{n=1}^{T}d_{n}d_{n}^{\prime }\mathbf{+} o_{p}(1)$ and $s_{T,\lambda }(\tilde{\theta})=(2\tilde{\sigma}_{\epsilon }^{4})^{-1}\sum\nolimits_{n=1}^{T}(\tilde{\varpi}_{n}^{2}-\tilde{\sigma} _{\epsilon }^{2})d_{n}$, where $\tilde{\varpi}_{n}$ is the residual computed based on $\tilde{\theta}$. For $i\neq j$, $i,j\in \{1,...,T\}$, it follows$E \tilde{\epsilon}_{n}^{2}=\omega _{0}$, $E\tilde{\epsilon}_{i}^{2}\tilde{ \epsilon}_{j}^{2}=E\tilde{\epsilon}_{i}^{2}E\tilde{\epsilon}_{j}^{2}$ and $E \tilde{\epsilon}_{i}^{4}=E\tilde{\epsilon}_{j}^{4}$. Let $\tilde{\varpi} _{n}^{2}=T^{-1}\sum\nolimits_{n=1}^{T}(y_{n}-z_{n}^{\prime }\tilde{\xi} +x_{n}^{\prime }\tilde{\beta})^{2}$ and $T^{-1}\sum \nolimits_{n=1}^{T}d_{n}d_{n}^{\prime }=\Sigma _{d}\mathbf{+}o_{p}(1)$. It can then be shown that $T\mathcal{G}_{T,\lambda \lambda }(\tilde{\theta})= \tilde{\Sigma}_{d}\mathbf{+}o_{p}(1)$, where $\tilde{\Sigma} _{d}=\sum\nolimits_{n=1}^{T}(y_{n}-z_{n}^{\prime }\tilde{\xi}+x_{n}^{\prime } \tilde{\beta})^{2}(\sum\nolimits_{n=1}^{T}d_{n}d_{n}^{\prime })^{-1}$. Assumption (ref) can be checked through the following proposition (where Assumptions (ref), (ref) and (ref) can be checked as in Andrews, 1999, Example 1).
Tables (ref)--(ref) report the estimated probabilities of rejecting $\mathsf{H}_{0}$, FWER and $\mathsf{H}_{0i}$ with $\epsilon _{n}\sim N(0,1)$ or $t_{5}$. We generate i.i.d $x_{n}^{d}$ from $N(0,4)$ and i.i.d. $z_{n}$ from the multivariate normal distribution with mean 0 and covariance matrix with diagonal entry $1$ and all off-diagonal elements being 0.2. The results show that MinP tests have a competitive finite sample performance compared with score tests based on $\tilde{t}_{c}$ in terms of global testing. Score tests based on $\tilde{t}_{t}$ can have considerable worse global power in some cases. With regard to multiple testing MinP score tests identify the false $\mathsf{H}_{0i}$ with probability increases towards $1$ as sample size increases although the FWER control may not always guaranteed. While MinP-s score tests may perform slightly better in multiple testing, such an advantage may be compromised by their global power in testing $\mathsf{H} _{0} $.
This paper proposes MinP score tests that allow for global and multiple testing of one-sided hypotheses. The simulation results suggest that the proposed tests are competitive with existing one-sided score tests with respect to global power. We further demonstrate that the proposed tests may be used for multiple testing and model identification among nested models. The feature of model identification is appealing when the model parameter is subject to inequality constraints. These tests allow for model identification without estimating candidate models except the one defined under the global null hypothesis, which is usually easy to obtain. We illustrated applications of the tests in linear regression, ARCH models, and random coefficient models. Admissibility and model identification properties of these tests provide a theoretical support of our testing approach to the practically important problem of one-side testing.