EconBase
← Back to paper

Extended MinP Tests for Global and Multiple testing

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

84,809 characters · 13 sections · 60 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Extending MinP Tests for Global and Multiple Testing

{3pt} {3pt}

frontmatter\address{UniSA Business School, University of South Australia} \begin{abstract} Empirical economic studies often involve multiple propositions or hypotheses, with researchers aiming to assess both the collective and individual evidence against these propositions or hypotheses. To rigorously assess this evidence, practitioners frequently employ tests with quadratic test statistics, such as $F$-tests and Wald tests, or tests based on minimum/maximum type test statistics. This paper introduces a combination test that merges these two classes of tests using the minimum $p$-value principle. The proposed test capitalizes on the global power advantages of both constituent tests while retaining the benefits of the stepdown procedure from minimum/maximum type tests. JEL classification: C12. \end{abstract} \begin{keyword} Familywise error rate\sep Evaluation of treatment effects\sep Global testing\sep Multiple testing\sep Step-down procedure \sep Tippett tests \end{keyword}

Introduction

Empirical economic studies often involve multiple propositions or hypotheses. In such studies, researchers are interested in evaluating the collective and individual evidence against these propositions or hypotheses. A prominent example in the recent literature is the evaluation of treatment effects that are measured by several variables. Two popular classes of statistical tests adopted by researchers are tests based on quadratic-form test statistics (QF tests), or tests based on the minimum of $p$-values, (or the maximum of the corresponding test statistics). QF tests evaluate overall treatment effects; they include the Wald, $F$- and Hotelling's $T^{2}$ tests (see chagne09, bendan16, angjorkue18, beaman21 and chosutzim22 for examples). The minimum $p$-values (MinP) tests evaluate the treatment effect relating to individual variable measurements, and look to control the multiplicity of Type I errors (see anderson08 , leesha14, gerhecetal14, lishxu16 and chosutzim22 for examples).

Generally speaking, QF tests have good global power in detecting overall treatment effects arising from the accumulation of many small individual treatment effects. However, a rejection by a QF test does not indicate which individual treatment effects are statistically significant, unless subsequent testing via the closed testing procedure finds further evidence (see, e.g., marpergar76, roazwo11, lu20cl and goehemsol19). In comparison, detection of any individual treatment effect in MinP tests implies an overall treatment effect. This tends to provide global power advantage over QF tests when one or a few individual treatment effects are present. MinP tests are popular in multiple testing; they can be viewed as a special case of closed tests, but have the computational advantage of sequential testing over general closed tests.

Naturally, the primary objective in the evaluation of treatment effects is the presence of the overall treatment effect. However, the researcher usually does not have a priori knowledge of the true distribution generating the data in an empirical investigation. Since QF and MinP tests have distinct global powers for detecting on the presence of the overall treatment effect, the researcher is likely to attempt both tests to examine empirical evidence. For example, the researcher may attempt the other test if one test does not provide the evidence he/she anticipates, or they may initially aim to identify as many possible variables that relate to the treatment effect as possible, but turns to QF tests after the initial attempt fails to identify any variables. In fact, we argue that the researcher should attempt both tests for a robustness check on how sensitive the empirical evidence is to the distribution generating the data.

If both the tests are carried out, but only the results of one test are reported, or a test is chosen based on sample information, then the error probability of claimed findings may not be controlled at the desired level. For example, if the researcher chooses to report the more significant evidence on the overall treatment effect from the two tests attempted, then the error probability of claiming positive findings when there are no treatment effects (Type I error) becomes inflated beyond the level the researcher intends to control. Over-reporting has drawn increasing concerns in the literature. Recent papers by christensenmiguel18 and young19 provide excellent accounts of the issue. In his influential paper, wh2000 draws attention to the dangerous practice of data snooping with primary concern on testing predictive superiority when different forecasting models are explored (see also hansenpr05, romwol05 and roshwo08).

This paper proposes an adjusted $p$-value method which offers a tool for robust and honest reporting. In fact, we show that the method can be viewed as a combined test constructed through the minimum $p$ value principle, which is a special case of the generalized mean of $p$-values for combining tests (cf. vovkwang20 and vovkwang22). We show that our combined tests are admissible if both constituent tests are admissible. Furthermore, we provide the necessary and sufficient condition under which the combined test improves the global power of the constituent tests. The combined test also retains the stepdown procedure of MinP tests when simultaneously testing individual treatment effects, and preserves the control of the familywise error rate (FWER) by the MinP test.

The $p$-values in the combined test can be easily computed based on the Bonferroni correction; more precise adjustments may be computed through resampling schemes (e.g., leesha14, churom16 and bugcansha19). Computational implementations are readily available in popular packages such as R and Stata (see brhowe16 and claromwol20), which can be easily extended to compute our adjusted $p$ -values through the inclusion of an additional $p$-value associated with a global test such as a QF test.

Related literature. The idea of combining different tests can be traced back to tippett31, pearson33 and fisher36. There has been renewed interest in combining tests in the recent literature. vovkwang20 and vovkwang22 studied a class of combination tests based on the generalized mean of $p$-values and suggested the closed testing procedure based on their combination tests for multiple testing. Our test is a Tippett type test constructed based on the minimum of the $p$-values of the constituent tests, and may be viewed as a special case of the generalized mean of $p$-values. lu16 proposed an extended MaxT (EMaxT) test by combining a sum test and a MaxT test for one-sided testing where the sum test is chosen to direct to the `middle' of the constrained parameter space under the global hypothesis. Our tests here can be viewed as a generalization of the EMaxT test to a more general setting, by combining p-values to allow for the adoption of a general global test, as well as retaining a stepdown procedure of multiple testing. Both the EMaxT test and the combined test proposed in this paper have a common feature of preserving the rejection region shapes of constituent tests, hence somewhat inheriting the respective strength of the constituent tests with regard to global power. In fact, the properties studied in this paper also apply to the EMaxT test. In high dimensional problems, combination tests of maximum type tests and QF tests have drawn interests in the recent literature due to their distinct global powers against sparse and dense alternatives. fanliayao15 proposed tests aiming at enhancing the global power of QF tests in high-dimensional settings. Their tests combine a QF test and some screened $t$-tests where the screened $t$-test components have an asymptotic size $0$. kockpre23power provide sufficient conditions for fanliayao15's test to achieve global power enhancement. hexuwupan21 and fejilixi20 derived approximated distributions of constituent tests in high-dimensional settings, and studied the properties of subsequent combination tests. However, these high-dimensional combination tests are not devised for multiple testing, but rather for global testing. There are also other combination tests proposed in the recent literature. andrewsI16 studied a class of combination tests for testing of weakly identified models by exploring global power advantages of the two distinct tests in combination. helmeicha19post constructed a combined test for multiple testing based on the marginal $p$-values conditional on a global test. Our paper contributes to this growing body of literature by proposing a combined test that allows for global and multiple testings simultaneously. In practice, the set of individual hypotheses rejected by our combined test is the same as that of the constituent MinP test once an individual hypothesis is rejected by our combined test, while revealing a potentially much stronger signal on global testing.

The remainder of the paper is organized as follows. We motivate and illustrate our combined test in the case of a multivariate normal location model in Section (ref). We then present the combined test in a more general setting in Section (ref) where some properties concerning the combined test are established. In Section (ref) simulation studies are provided to examine the performance of the combined test compared with other tests. Section (ref) provides a real data application for testing the effects that exercise has on seven biometric measures based on the data published in chagne09. Concluding remarks are made in Section (ref). Proofs are presented in the appendix.

Illustrations

To fix ideas, we begin with an example of testing the multivariate normal mean with a known covariance. Let $X=(X_{1},...,X_{k})^{\prime }\sim N(\mu ,\Sigma )$, where $\mu =(\mu _{1},...,\mu _{k})^{\prime }$ and $\Sigma $ has the structure of the equicorrelation matrix $\{\rho _{ij}\}$, $i,j\in K=\{1,\cdots ,k\}$, with $\rho _{ij}=\rho $, $-1<\rho <1$, when $i\neq j$ and $\rho _{ij}=1$, when $i=j$. The individual hypotheses are:

equation*[equation* omitted — 92 chars of source]

Let $\Phi (\cdot )$ be the cumulative distribution function (CDF) of the standard normal random variable and $F_{\chi _{k}^{2}}(\cdot )$ be the CDF of the central chi-square random variable with $k$ degrees of freedom, and $ F_{\chi _{k}^{2}}^{-1}(\cdot )$ be the inverse function of $F_{\chi _{k}^{2}} $. Denote the individual $p$-value by $\hat{p}_{i}(X_{i})=2\Phi (-\left\vert X_{i}\right\vert )$.

A simple example

Consider the two-sided testing with $k=2$. With the control of the FWER, the probability of rejecting at least one true $H_{i}$ (more detailed discussions on the FWER are provided in Section (ref)) , MinP tests would reject $H_{i}$ if

equation*[equation* omitted — 93 chars of source]

where $\alpha \in (0,1)$ is the significance level, and $c_{m}(\alpha )$ satisfies

equation*[equation* omitted — 93 chars of source]

Consider the global null hypothesis $H_{K}:\mu =0$ which is the intersection of all individual null hypotheses $\cap _{i=1}^{k}H_{i}$. Rejection of any $ H_{i}$ implies rejection of $H_{K}$. Therefore, MinP tests can be used to jointly test $H_{K}$ against $H_{K}^{\prime }:\mu \neq 0$ and reject $H_{K}$ if any $\hat{p}_{i}(X_{i})<c_{m}(\alpha )$ or $\min \{\hat{p} _{i}(X_{i}),i\in K\}<c_{m}(\alpha )$. If the researcher uses the Likelihood Ratio (LR) tests for testing $H_{K}$, the test statistic $X^{\prime }\Sigma ^{-1}X$ follows the null distribution $\chi _{k}^{2}$ and $H_{K}$ would be rejected if

equation*[equation* omitted — 86 chars of source]

Figure (ref) shows the comparison of the rejection regions of LR tests and MinP tests in the case of $k=2$ and $\rho =0$ with $\alpha =0.05$. The rejection region of MinP tests is

equation*[equation* omitted — 92 chars of source]

where $c_{m}(\alpha )=1-(1-\alpha )^{1/2}=0.0253$. $S_{m}(\alpha )$ represents the area outside the square box. The rejection region of LR tests is

equation*[equation* omitted — 62 chars of source]

which is the area outside the circle. For $X\in A(\alpha )=S_{g}(\alpha )\cap S_{m}^{c}(\alpha )$, where $S^{c}$ is the complement set of $S$, LR tests reject $H_{K}$, but MinP tests do not reject $H_{K}$. For $X\in B(\alpha )=S_{g}^{c}(\alpha )\cap S_{m}(\alpha )$, MinP tests reject $H_{K}$ , but LR tests do not reject $H_{K}$. MinP tests are more likely to reject $ H_{K}$ than LR tests when one of $\left\vert X_{i}\right\vert $, $i=1,2$, dominates the other. However, as the boundary of the rejection region of LR tests is defined by the circle $X_{1}^{2}+X_{2}^{2}=F_{\chi _{2}^{2}}^{-1}(0.95)=5.99$, LR tests are more likely to reject $H_{K}$ than MinP tests when none of $\left\vert X_{i}\right\vert $, $i=1,2$, dominates the other. If one chooses LR tests or MinP tests based on the sample information, it would inevitably lead to a data snooping problem. For example, if one chooses LR tests or MinP tests depending on the outcome of rejection, it would effectively lead to the enlarged rejection region as $ S_{g}(\alpha )\cup B(\alpha )$ or $S_{m}(\alpha )\cup A(\alpha )$. An enlarged rejection region implies an inflated size. For example, the inflated size is about $0.07$ at $\alpha =0.05$ when $\rho =0.9$ or $-0.9$.

With regards to multiple testing of $H_{i}$, $i=1,2$, MinP tests reveal the evidence on testing $H_{i}$ with the control of the FWER. A rejection of $ H_{K}$ by LR tests itself does not directly indicate which $H_{i}$ should be rejected unless $H_{i}$ is rejected as well. This is the so-called closure testing procedure in which the rejection of an $H_{i}$ requires rejections of all $H_{K_{i}}$, $\{i\}\subseteq K_{i}\subseteq K$, in a general case of $ k\geq 2$. The rejection region of closed tests in the case of $k=2$ is

equation*[equation* omitted — 96 chars of source]

where $S_{i}(\alpha )=\{X:\hat{p}_{i}(X_{i})\leq \alpha \}$. This rejection region is a strict subset of $S_{g}(\alpha )$; hence, the FWER of closed tests based on $S_{g}(\alpha )$ is strictly less than $\alpha $. As $k $ increases, the FWER control becomes more conservative and consequently, the capacity to detect false $H_{i}$ is reduced.

figure[figure omitted — 289 chars of source]

A combination of two tests

In the above $k=2$ case, it follows that

eqnarray*[eqnarray* omitted — 181 chars of source]

where $a\wedge b$ defines $\min (a,b)$, and we may use both the operations interchangeably throughout the paper. To maintain the relative strength of the global power of LR tests and MinP tests, one may preserve the shape of the combined rejection region by adjusting it through an $\alpha ^{\prime }<\alpha $ such that the combined rejection region has a probability of $ \alpha $ under $H_{K}$, namely,

equation[equation omitted — 186 chars of source]

This implies that for an observed $X=x$, the adjusted $p$-value $\hat{p} _{l}^{adj}(x)$, $l=g,m,i$, can be computed as

equation*[equation* omitted — 104 chars of source]

One may compute the null distribution

equation*[equation* omitted — 77 chars of source]

based on Monte Carlo simulation by randomly drawing $X$ from $N(\mu ,\Sigma ) $ if it is assumed known, or through the permutation or bootstrap resampling methods based on the sample. One may opt to ease the computational burden by computing the adjusted $p$-values based on, for example, the Bonferroni inequality. More detailed discussions are presented in Section (ref).

Let $\hat{p}_{c}^{adj}(X)=\hat{p}_{g}^{adj}(X)\wedge \hat{p}_{m}^{adj}(X)$. The rejection region based on the combined test is

equation*[equation* omitted — 71 chars of source]

which is equivalent to

equation*[equation* omitted — 97 chars of source]

where $\alpha ^{\prime }<\alpha $ satisfies ((ref)). The rejection rule of the combined test is to reject $H_{K}$ if $\hat{p}_{c}^{adj}(x)\leq \alpha $, not to reject it otherwise.

To compare the power performance of the combined test of $H_{K}$ with LR tests and MinP tests, we approximate power functions of tests considered in Section (ref) based on $1,000,000$ and $100,000$ independent random draws from $N(\mu ,\Sigma )$ for the cases of $k=2$ and $k\geqslant 2$, respectively. The exception is the global power function of LR tests which is computed as $\Pr \{\chi _{k}^{2}(r^{2})>F_{\chi _{k}^{2}}^{-1}(0.95)\}$, where $\chi _{k}^{2}(r^{2})$ is the chi-square random variable with $k$ degrees of freedom and the non-centrality parameter $r^{2}$. Figure (ref) presents the comparison of the global powers of testing $H_{K}$ for LR, MinP and the combined test with $\alpha =0.05$ in the bivariate case. We take $ \mu _{1}=r\cos \varphi $, $r=2$ and

equation*[equation* omitted — 243 chars of source]

so that $\mu ^{\prime }\Sigma ^{-1}\mu =r^{2}$. The comparison shows that LR tests have an overall global power advantage over MinP tests. However, MinP tests can outperform LR tests when either $\left\vert \mu _{1}\right\vert $ or $\left\vert \mu _{2}\right\vert $ dominates the other. The combined test somewhat inherits the respective strengths of LR and MinP tests with regard to global power. In relation to multiple testing, the combined test can outperform closed tests. However, MinP tests are more likely to reject $ H_{i} $, $i\in K$, than the combined test. This is because $\hat{p} _{g}(x)\wedge \hat{p}_{m}(x)\leq \hat{p}_{m}(x)$ for every $x\in X$ with the strict inequality holding for some $x\in X$. Figure (ref) presents a comparison of the average number of correctly rejected (ANCR) false $H_{i}$ for closed, MinP and the combined test.

To further illustrate that the combined test can share the power strength of MinP tests to some extent, Figures (ref) and (ref) present comparisons of the global power in testing $H_{K}$, as well as comparisons of the probability of rejecting $H_{1}$ in multiple testing as the number of hypotheses $k$ increases. The comparison in Figure (ref) is based on the case of $\Sigma =I$ and $\mu =(3,0,...,0)^{\prime }$, while that in Figure (ref) is based on the case of $\rho _{ij}=0.9$, $i\neq j$, $i,j\in K$, and $\mu =(3,...,3)^{\prime }$. The comparisons show that MinP tests have a clear power advantage over LR tests in both global and multiple testing. The advantage becomes increasingly apparent as $k$ increases. The combined test in such cases share some of the strength of MinP tests.

figure[figure omitted — 268 chars of source]
figure[figure omitted — 334 chars of source]
figure[figure omitted — 352 chars of source]
figure[figure omitted — 388 chars of source]

Stepdown procedure

As for MinP tests in multiple testing, an improved ability to reject more $H_{i}$, $i\in K$, is possible for the combined test through a stepdown testing procedure. For example, for the points

equation*[equation* omitted — 121 chars of source]

in the bivariate example, the combined test rejects $H_{1}$, but not $H_{2}$ . However, if we proceed in the same fashion as in the stepdown procedure of MinP tests, one would then reject $H_{2}$ in the second step because $\hat{p} _{2}(X_{2})<\alpha $. Although the combined test has a disadvantage compared with MinP tests in the first step of multiple testing, the combined test would have the same outcomes as MinP tests in the stepdown procedure if the null hypothesis with the smallest $\hat{p}_{i}(X_{i})$ is rejected by the combined test in the first step of the stepdown procedure.

The combined test

Setup

We study the combined test in a more general setting. Suppose that the sample $X^{(n)}$, where $n$ indicates sample size, is generated from the unknown distribution $P\in \mathbf{P}$, where $\mathbf{P}$ defines a set of probability distributions. Let $\hat{p}=\hat{p}(X^{(n)})$ be a $p$-value. Let $G^{(n)}(u,P)$, $u\in \lbrack 0,1]$, be a sequence of CDFs of $\hat{p}$ under $P\in \mathbf{P}$. Denote the test function by

equation*[equation* omitted — 92 chars of source]

where $\mathbf{1(\cdot )}$ is the usual indication function. Note that for ease of presentation, we consider nonrandomized tests in this paper.

Let $H_{i}$, $i\in K=\{1,...,k\}$, $k\geq 2$, be the individual null hypotheses and $H_{i}^{\prime }$ be the corresponding alternative hypotheses. Let the set of distribution under $H_{i}$ be $\mathbf{P} _{i}\subset \mathbf{P}$. Let $K_{i}\subseteq K$ be a sub-index set, and $ K_{\ast }\subseteq K$ be the set containing the indices of true $H_{i}$. (The subscript $i$ in $K_{i}$ will be useful when we discuss the stepdown procedure later on). Denote by $\mathbf{P}_{K}=\cap _{i\in K}\mathbf{P} _{i}\subset \mathbf{P}$ the set of null distributions corresponding to the global null hypothesis $H_{K}$ and by $\mathbf{P}_{K}^{\prime }=\mathbf{ P\smallsetminus P}_{K}$ the set of distributions corresponding to the alternative hypothesis $H_{K}^{\prime }$. Assume $\mathbf{P}_{K}\subset \mathbf{P}_{i}$, for all $i\in K$. That is, $\mathbf{P}_{K}$ is the strict subset of $\mathbf{P}_{i}$, for all $i\in K$.

A test $\phi ^{(n)}$ of $H_{K}$ is referred to as the asymptotic pointwise level-$\alpha $ test if

equation*[equation* omitted — 96 chars of source]

where $E_{P\in \mathbf{P}_{K}}(\cdot )$ is the expected value with respect to $P\in \mathbf{P}_{K}$. (In this paper, we restrict our attention to pointwise control.) In relation to multiple testing, the FWER is the probability of rejecting any $H_{i}$, $i\in K_{\ast }$, under the true $P\in \mathbf{P}_{K_{\ast }}$. That is,

equation*[equation* omitted — 105 chars of source]

The asymptotic pointwise FWER control at the level $\alpha $ based on the sample $X^{(n)}$ is achieved if

equation*[equation* omitted — 64 chars of source]

It is worth noting that any distribution restricted by a possible configuration of true and null hypotheses belong to $\mathbf{P}_{K_{\ast }}$ . The above definition of the FWER control is known as strong control of the FWER.

Since the true null set $K_{\ast }$ is typically unknown, nor is the true $ P\in \mathbf{P}_{K_{\ast }}$, assuming the true $P\in \mathbf{P}_{K}$ (which is referred to as the weak control in the literature) does not guarantee the control of the FWER (see, e.g., romwol05b). However, in many applications, the researcher may be able to assume the subset pivotality condition of wesyou93, which says that the true $P\in \mathbf{P} _{K_{\ast }}$ is not affected by whether $H_{i}$, $i\in K\mathbf{ \smallsetminus }K_{\ast }$, is true or not. Hence, the FWER can be controlled by assuming the true $P\in \mathbf{P}_{K}$. The subset pivotality condition is not a necessary condition for controlling FWER. romwol05b showed a weaker sufficient condition for controlling FWER that is also satisfied under the subset pivotality condition. We shall show later that the proposed combined test controls the FWER as long as the MinP test being combined controls the FWER.

Let the subscripts $c,g$, $m$, and $i$ indicate the combined, global, MinP and individual tests, respectively. The $p$-value of the combined test is

equation*[equation* omitted — 59 chars of source]

The adjusted $p$-value of the combined test for testing $H_{K}$ is defined as

equation*[equation* omitted — 89 chars of source]

or equivalently,

equation[equation omitted — 91 chars of source]

where

equation*[equation* omitted — 102 chars of source]

The limiting null distribution of $G_{c}^{(n)}(\cdot ,P\in \mathbf{P}_{K} \mathbf{)}$ is usually unknown for our combined tests, so we use $ G_{c}^{(n)}(\cdot )$ for computing $\hat{p}_{c}^{adj}$. The computation may be based on simulation methods by utilizing some limiting distribution, or based on resampling methods such as bootstrap, permutation, and subsampling.

The researcher may opt to ease the computational burden by computing conservative adjusted $p$-values based on the Bonferroni inequality as

eqnarray*[eqnarray* omitted — 141 chars of source]

or

eqnarray*[eqnarray* omitted — 160 chars of source]

then compute $\hat{p}_{c}^{adj}$ based on ((ref)). One may improve the conservativeness of the Bonferroni inequality for computing $\hat{p} _{m}^{adj}$ in testing $H_{K}$ by using inequalities such as the {\v{S}}id{ \'{a}}k inequality (sidak68) or the Simes inequality (simes86 and sarkar98). However, these inequalities are not generally applicable.

Global testing

The combined test rejects the global null hypothesis $H_{K}$ if $\hat{p} _{c}^{adj}\leq \alpha $; otherwise it accepts $H_{K}$. We first present the size property of the combined test. We refer to the procedure in which $\hat{ p}_{c}^{adj}$ is computed based on $G_{c}^{(n)}(\cdot )$ as $\phi _{c,1}^{(n)}$ and to the procedure based on the Bonferroni inequality as $ \phi _{c,2}^{(n)}$. The $\phi _{c,2}^{(n)}$ procedure may be conservative in the sense that the test size may be strictly less than $\alpha $, but it is computationally easy to implement. Note that to keep the presentation concise $\phi _{c}^{(n)}$ will be used to represent both $\phi _{c,1}^{(n)}$ and $\phi _{c,2}^{(n)}$ when no confusion is deemed to arise.

assumptionWith fixed $P\in \mathbf{P}_{K}$ and $l\in \{m,g,c\}$, \newline (i) $G_{l}^{(n)}(u,P\mathbf{)}\rightarrow G_{l}(u,P)$, \newline (ii) $G_{l}(u,P)$ is a continuous and strictly increasing function of $u\in \lbrack 0,1]$.\newline (iii) $G_{l}(u,P)\leq U(0,1)$, where $U(0,1)$ is the uniform distribution on $[0,1]$.
remarkWhen $\mathbf{P}_{K}$ is a singleton, that is, testing the simple null hypothesis, $G_{l}(u,P\in \mathbf{P}_{K})$ is typically the uniform distribution on $[0,1]$. In such a case, Assumption (ref)(iii) is met with equality. When $\mathbf{P}_{K}$ is composite, Assumption (ref)(iii) requires stochastic domination by the uniform random variable to ensure size control. For example, suppose $H_{i}:\mu _{i}\leq 0$, $i\in K$, in the multivariate normal example in Section (ref), the null distribution $ \mathbf{P}_{K}$ is composite. The least favourite null distribution at $\mu =0$ for testing $H_{K}$ is stochastically dominated by $U(0,1)$ representing the null distribution of the $p$-value with the true ${\mu }\in \{{\mu :}\mu _{i}\leq 0,i\in K\}$.

The following lemma concerns the size property of the combined test.

lemma(i) If Assumption (ref) holds for $l=c$, then $\limsup_{n\rightarrow \infty }E_{P\in \mathbf{P}_{K}}\phi _{c,1}^{(n)}\leq \alpha $. (ii) If Assumption (ref) holds for $l=m,g$, then $\limsup_{n\rightarrow \infty }E_{P\in \mathbf{P}_{K}}\phi _{c,2}^{(n)}\leq \alpha $, with the inequality holding strictly if \begin{equation} \liminf_{n\rightarrow \infty }E_{P\in \mathbf{P}_{K}}\phi _{l}^{(n)}(1-\phi _{l^{\prime }}^{(n)})>0 \end{equation} where $l,l^{\prime }\in \{g,m\}$ and $l\neq l^{\prime }$.
remarkAssumption (ref)(i) required for $l=c$ is to ensure that the estimation of $G_{c}^{(n)}(\hat{p}_{c},P\mathbf{)}$ for the test $\phi _{c,1}^{(n)}$ is consistent, while ((ref)) implies that $\phi _{l}^{(n)}$, $l\in \{g,m\} $ has limiting distinct rejection regions.
remarkWhen two tests are attempted, adjusting the $p$-value is important for overall control of test size, as illustrated in the simple example in Section (ref). In some cases, an inflated size can be much more severe. For example, if the LR test in the simple example in Section (ref) is replaced by the sum test of birovami09 that has the rejection region \begin{equation*} S_{g}(\alpha )=\{X:\left\vert X_{1}+X_{2}\right\vert >\sqrt{2(1+\rho )} c_{\Phi }(\alpha )\}, \end{equation*} where $c_{\Phi }(\alpha )$ is the $(1-\alpha /2)$th quantile of the standard normal distribution, the inflated size is about 0.21 when $\alpha =0.05$ and $\rho =0.9$.
remarkWhile the test $\phi _{c,1}^{(n)}$ may be able to achieve the size control asymptotically exactly at the level $\alpha $ in some cases, the conservativeness of the tests $\phi _{c,2}^{(n)}$ is reflected by their size control being asymptotically strictly less than $\alpha $.

We now turn to study some global properties. Consider a set of local alternatives

equation*[equation* omitted — 128 chars of source]

A test $\phi _{l}^{(n)}$ is asymptotically $d$-admissible if for any other test $\phi ^{(n)}$

equation[equation omitted — 209 chars of source]

and

equation[equation omitted — 178 chars of source]

jointly imply

equation*[equation* omitted — 112 chars of source]

for all $P\in \mathbf{P}$. The $d$-admissibility implies that $\phi _{l}^{(n)}$ can have better global power for some $P_{n}\in \mathbf{P} _{n,K}^{\prime } $ asymptotically compared with any other tests that do not have an asymptotically larger size.

theoremUnder Assumptions (ref) if both the two tests $\phi _{g}^{(n)}$ and $\phi _{m}^{(n)}$ are asymptotically $d$-admissible, then the combined tests $\phi _{c}^{(n)}$ is asymptotically $d$-admissible.

Although Theorem (ref) states that both the test $\phi _{c,1}^{(n)}$ and $\phi _{c,2}^{(n)}$\ are asymptotically $d$-admissible, the power of $ \phi _{c,2}^{(n)}$ may be asymptotically uniformly improved by $\phi _{c,1}^{(n)}$. This improvement typically occurs when $\phi _{c,2}^{(n)}$ has a test size strictly less than the nominal level $\alpha $. In comparison to $d$-admissibility, $\alpha $-admissibility is defined as follows: for any other level-$\alpha $ test $\phi ^{(n)}$, ((ref)) implies ((ref)) for all $P\in \{P_{n}:P_{n}\in \mathbf{P} _{n,K}^{\prime }\}$ (lehrom05). In other words, the $\alpha $ -admissibility of a test demands that there does not exist any other test that has better power for at least some $P\in \mathbf{P}_{K}$ and non-worse power for all other $P\in \mathbf{P}_{K}$.

theoremUnder Assumptions (ref) if both the two tests $\phi _{g}^{(n)}$ and $\phi _{m}^{(n)}$ are asymptotically $d$-admissible and $ \lim_{n\rightarrow \infty }E_{P\in \mathbf{P}_{K}}(\phi _{c}^{(n)})=\alpha $ , then the combined tests $\phi _{c}^{(n)}$ is asymptotically $\alpha $ -admissible.
remarkOne may construct $\phi _{m}^{(n)}$ based on the Bonferroni inequality and combine it through a Monte Carlo or resampling method such that the combined test achieves an asymptotically exact size $\alpha $. By Theorem (ref) the resultant combined test $\phi _{c,1}^{(n)}$ is asymptotically $\alpha $ -admissible.

The following theorem provides a necessary and sufficient condition under which the combined test enhances the power of a constituent test. Let $ \tilde{\phi}_{l}^{(n)}=\mathbf{1}(\hat{p}_{l}^{adj}\leq \alpha )$, $l\in \{g,m\}$, be the test function based on the adjusted $p$-value instead of the unadjusted $p$-value used in $\phi _{g}^{(n)}$ and $\phi _{m}^{(n)}$.

theorem\begin{equation*} \liminf_{n\rightarrow \infty }E_{P\in \mathbf{P}_{n,K}^{\prime }}(\phi _{c}^{(n)})>\limsup_{n\rightarrow \infty }E_{P\in \mathbf{P}_{n,K}^{\prime }}(\phi _{l}^{(n)}),\qquad l\in \{g,m\}, \end{equation*} holds if and only if \begin{equation} \liminf_{n\rightarrow \infty }E_{P\in \mathbf{P}_{n,K}^{\prime }}\psi ^{(n)}>0, \end{equation} where \begin{equation*} \psi ^{(n)}=(\tilde{\phi}_{l^{\prime }}^{(n)}-\tilde{\phi}_{l}^{(n)})\mathbf{ 1}(\tilde{\phi}_{l^{\prime }}^{(n)}>\tilde{\phi}_{l}^{(n)})-(\phi _{l}^{(n)}- \tilde{\phi}_{l}^{(n)}), \end{equation*} $l\neq l^{\prime }$, $l,l^{\prime }\in \{g,m\}$.

The condition ((ref)) indicates that the power gain to $\tilde{\phi} _{l^{\prime }}^{(n)}$ from $\tilde{\phi}_{l}^{(n)}$ is asymptotically sufficient to offset the power loss of $\phi _{l}^{(n)}$ caused by the adjustment of its $p$-value. This condition reveals the sources of the power gain and loss in relation to the global power improvement of the combined test. We illustrate the condition ((ref)) in the bivariate example in Section (ref). Let $X_{i}=\sqrt{n}\hat{\mu}_{i}$, $i=1,2$, and $ \sqrt{n}(\hat{\mu}-\mu )\sim N(0,\Sigma )$. Let $\tilde{S}_{g}(\alpha )=\{X: \hat{p}_{g}^{adj}(X)\leq \alpha ,\mu =0\}$ and $\tilde{S}_{m}(\alpha )=\{X: \hat{p}_{m}^{adj}(X_{i})\leq \alpha ,\mu =0\}$. Let

equation*[equation* omitted — 90 chars of source]
equation*[equation* omitted — 90 chars of source]

Without loss of generality we let $l=m$ and $l^{\prime }=g$. Then

equation*[equation* omitted — 161 chars of source]

while

equation*[equation* omitted — 127 chars of source]

If we decompose $S_{m}(\alpha )-\tilde{S}_{m}(\alpha )$ into the two exclusive regions

equation*[equation* omitted — 190 chars of source]

it follows

eqnarray*[eqnarray* omitted — 229 chars of source]

where $(\tilde{B}(\alpha )-\Delta _{1})$ and $\Delta _{2}$ are exclusive. The rejection region for $\phi _{c}^{(n)}$ is

equation*[equation* omitted — 145 chars of source]

For the combined test $\phi _{c}^{(n)}$ to exhibit a better global power in testing $H_{K}$ under $P\in \mathbf{P}_{n,K}^{\prime }$ which corresponds to the local alternatives $\sqrt{n}\mu $, the following condition must hold

equation*[equation* omitted — 173 chars of source]

which is equivalent to

equation[equation omitted — 167 chars of source]

When either $\mu _{i}$, $i=1,2$, dominates the other, the first term in the left hand side of ((ref)) is likely dominate the second term. Consequently, the combined test improves the global power of the MinP test.

Theorem (ref) demonstrates that the combined test can improve the global power of the constituent tests under the condition stated in (ref). For a $P\in \mathbf{P}_{n,K}^{\prime }$, this improvement is more probable for the constituent test with a lower global power compared to the other test. This result is formally articulated in the next theorem, which establishes that the combined test exhibits a more balanced global power, with its global power is bounded between that of two constituent tests.

theoremFor $P\in \mathbf{P}_{n,K}^{\prime }$ and $l\neq l^{\prime }$ , $l,l^{\prime }\in \{g,m\}$ if \begin{equation*} \liminf_{n\rightarrow \infty }E_{P\in \mathbf{P}_{n,K}^{\prime }}\psi ^{(n)}\geq 0,\quad and\quad \limsup_{n\rightarrow \infty }E_{P\in \mathbf{P}_{n,K}^{\prime }}\psi ^{(n)}\leq \liminf_{n\rightarrow \infty }E_{P\in \mathbf{P}_{n,K}^{\prime }}(\phi _{l^{\prime }}^{(n)}-\phi _{l}^{(n)}), \end{equation*} then, \begin{equation} \liminf_{n\rightarrow \infty }E_{P\in \mathbf{P}_{n,K}^{\prime }}(\phi _{c}^{(n)})\geq \limsup_{n\rightarrow \infty }E_{P\in \mathbf{P} _{n,K}^{\prime }}(\phi _{l}^{(n)}), \end{equation} \begin{equation} \limsup_{n\rightarrow \infty }E_{P\in \mathbf{P}_{n,K}^{\prime }}(\phi _{c}^{(n)})\leq \liminf_{n\rightarrow \infty }E_{P\in \mathbf{P} _{n,K}^{\prime }}(\phi _{l^{\prime }}^{(n)}). \end{equation}
comment\begin{remark} Without loss of generality, assuming $\liminf_{n\rightarrow \infty }E_{P\in \mathbf{P}_{n,K}^{\prime }}(\phi _{l^{\prime }}^{(n)})\geq \limsup_{n\rightarrow \infty }E_{P\in \mathbf{P}_{n,K}^{\prime }}(\phi _{l}^{(n)})$, that is, $\phi _{l^{\prime }}^{(n)}$ has an asymptotically better power than $\phi _{l}^{(n)}$, then from the proof of Theorem (ref), in order for the combined test to improve the global power of both constituent tests, it requires \begin{equation*} \liminf_{n\rightarrow \infty }E_{P\in \mathbf{P}_{n,K}^{\prime }}\{\psi ^{(n)}-(\phi _{l^{\prime }}^{(n)}-\phi _{l}^{(n)})\}>0, \end{equation*} That is, the power gain to $\tilde{\phi}_{l^{\prime }}^{(n)}$ from $\tilde{ \phi}_{l}^{(n)}$ less than the power loss of $\phi _{l}^{(n)}$ caused by the adjustment of its $p$-value is asymptotically larger than the difference of power of the two constituent tests. \end{remark}
remarkIn the simulation studies reported in Sections (ref) and (ref) we found that in most cases, the global power of the combined test is bounded between the global powers of two constituent tests.

Multiple testing

The combined test may be viewed as an extended MinP test by adding another $ p $-value for testing the global $H_{K}$ to the minimand set of MinP tests. If an individual null hypothesis $H_{i}$, $i\in K$, is rejected after it has been adjusted for the extended multiplicity, then the combined test proceeds to the usual stepdown procedure of MinP tests for multiple testing. It is well-known that the stepdown procedure has better power in rejecting false individual null hypotheses than the single-step procedure (c.f. romwol05b).

Let

equation*[equation* omitted — 73 chars of source]

denote the ordered $p$-values $\hat{p}_{i}$, $i\in K$, and let $H_{(1)}$, $ H_{(2)}$, ..., $H_{(k)}$, be the corresponding null hypotheses. If $\hat{p} _{m}^{adj}\leq \alpha $, which is equivalent to $\min (\hat{p} _{i}^{adj},i\in K)\leq \alpha $, where $\hat{p}_{i}^{adj}=G_{c}^{(n)}(\hat{p} _{i},P\in \mathbf{P}_{K}\mathbf{)}$, then reject $H_{(1)}$ and the combined test proceeds to the usual stepdown procedure of MinP tests of the remaining $H_{(2)}$, ..., $H_{(k)}$. Let $G_{m,K_{i}}^{(n)}(\cdot ,P\in \mathbf{P} _{K_{i}}\mathbf{)}$, $K_{i}\subset K$, be the CDF of $\min (\hat{p}_{i},i\in K_{i})$. Let the adjusted $p$-value associated with testing $H_{(i)}$ be

equation*[equation* omitted — 106 chars of source]

where $K_{i}=\{(i),...,(k)\}$. Then, usual MinP tests may be implemented for $(i)=(2),...,(k)$: reject $H_{(i)}$ if $\hat{p}_{m,(i)}^{adj}\leq \alpha $, stop otherwise. One may compute the adjusted $p$-values based on the Bonferroni inequality as

equation*[equation* omitted — 105 chars of source]

where $\left\vert \cdot \right\vert $ is the cardinality of $K_{i}$.

One would expect the combined test to inherit properties of multiple testing from MinP tests to some degree. Compared with the stepdown procedure of MinP tests, our combined test differs only in the first step of rejecting $ H_{(1)} $; the combined test rejects $H_{(1)}$ if $\hat{p} _{(1)}^{adj}=G_{c}^{(n)}(\hat{p}_{(1)},P\in \mathbf{P}_{K})\leq \alpha $ whereas MinP tests reject $H_{(1)}$ if $\hat{p}_{m,(1)}^{adj}=G_{m,K}^{(n)}( \hat{p}_{(1)},P\in \mathbf{P}_{K})\leq \alpha $. Since $G_{m,K}^{(n)}(u,P\in \mathbf{P}_{K})\leq G_{c}^{(n)}(u,P\in \mathbf{P}_{K})$ for every $u\in \lbrack 0,1]$, it follows $\hat{p}_{m,(1)}^{adj}\leq \hat{p}_{(1)}^{adj}$. Thus, $\hat{p}_{(1)}^{adj}\leq \alpha $ implies $\hat{p}_{m,(1)}^{adj}\leq \alpha $; a rejection of $H_{(1)}$ by the combined test implies the rejection by the MinP test. Once $H_{(1)}$ is rejected by the combined test both tests share the same stepdown procedure, hence share the same testing outcome in terms of testing the remaining $H_{i}$, $i\in K\smallsetminus \{(1)\}$.

The FWER control requires that the probability of rejecting at least one true $H_{i}$, $i\in K_{\ast }$, is bounded above by the designated level $ \alpha $. Because the true null set $K_{\ast }$ is typically unknown to the researcher, the control of the FWER is not guaranteed in the stepdown procedure. romwol05b showed that MinP tests control the FWER under a monotonicity condition, which is a weaker condition than the subset pivotality condition of wesyou93. The nest theorem shows that the combined test controls the FWER so long as the MinP test being combined controls the FWER.

theoremIf the limit superior of the FWER of the MinP test being combined is less than $\alpha $, then the limit superior of the FWER of the combined test is also less than $\alpha $ .
remarkAs $\hat{p}_{g}\wedge \hat{p}_{m}\leq \hat{p}_{m}$, it follows $\hat{p} _{m}^{adj}\geq \hat{p}_{m}$. Consequently, the ability to reject $H_{(1)}$ by the combined test may be compromised. Nevertheless, such a compromise may be rewarded with a marked improvement in the global power of testing $H_{K}$ by $\phi _{m}^{(n)}$.

A general combined procedure

To facilitate the application of our proposed combined test, this section summarizes the procedure as follows.

algorithm[algorithm omitted — 612 chars of source]

Monte Carlo studies

This section reports a simulation study on tests of the multivariate mean. The significance level is set to $0.05$. The number of replications is set to 2000. Let $X^{(n)}=\{X_{t}=(X_{t1},...,X_{tk})^{\prime },t=1,...,n\}$, where $X_{t}$ is an independent $k$-dimensional random vector from the multivariate normal distribution with the mean $\mu =(\mu _{1},...,\mu _{k})^{\prime }$ and covariance $\Sigma _{X}$. In our simulation study, we let $\Sigma _{X}$ have the correlation matrix structure $\{\rho _{ij}\}$, $ i,j\in K$ as in Section (ref). Let $\bar{X}=(\bar{X}_{1},...,\bar{X} _{k})^{\prime }$ be the studentized sample mean, and let $\hat{\Sigma}$ be the correlation matrix corresponding to $\hat{\Sigma}_{X}=(n-1)^{-1} \sum_{t=1}^{n}(X_{t}-\bar{X})(X_{t}-\bar{X})^{\prime }$. By the multivariate central limit theorem it follows that

equation*[equation* omitted — 121 chars of source]

where $\hat{\Sigma}^{1/2}$ is the matrix such that $\hat{\Sigma}^{1/2}\hat{ \Sigma}^{1/2}=\hat{\Sigma}$.

We conducted simulation studies on two-sided testing of $\mu $ to compare the performance of the $\phi _{c^{(1)}}^{(n)}$ test with that of other tests. Hotelling's $T^{2}$ tests are adopted for testing the global null hypothesis $H_{K}$, as well as for implementing the combined test and closed tests. For an individual test of $H_{i}:\mu _{i}=0$ against $H_{i}^{\prime }:\mu _{i}\neq 0$, $i\in K$, the test statistic is $T_{i}=\left\vert n^{ \frac{1}{2}}\bar{X}_{i}\right\vert $. We consider two approaches for computing $p$-values. One is based on the bootstrap method and the other is based on a limiting normal distribution.

In our Monte Carlo simulation study the design of the correlation matrix $ \Sigma _{X}=\left\{ \rho _{ij}\right\} $ has two structures. One takes the form of the equicorrelation matrix, and the other takes $\rho _{ij}=a^{\left\vert i-j\right\vert }$ with $a=0.5$ or $-0.5$. In the bootstrap approach $2000$ bootstrap samples are used. In the limiting distribution approach, we compute $\hat{p}_{i}=2\Phi (-\left\vert n^{\frac{1 }{2}}\bar{X}_{i}\right\vert )$ and $\hat{p}_{g}=1-G_{\chi _{k}^{2}}(n\bar{X} ^{\prime }\hat{\Sigma}^{-1}\bar{X})$, and use $10,000$ random draws to approximate the distribution $G_{c}^{(n)}(u,P)$ and $G_{m,K_{i}}^{(n)}(u,P) $ , $K_{i}\subseteq K$. Tables (ref)--(ref) report the results of the simulation study. The results for the global hypothesis testing are reported in terms of estimated sizes and powers. The results for the multiple hypothesis testing are reported in terms of the estimated FWER and the estimated ANCR false $H_{i}$. The notation $\mathbbm{1}_{k,m}$ represents the $k$-dimension column vector with the first $m$ elements being $1$ and the remaining elements being $0$. We observe that Hotelling's $T^{2}$ tests outperform MinP tests in global power in many instances, but not in all cases. The global power performance of MinP tests can increase by up to 20% in cases where the individual $X_{ti}$, $i\in K$, has an equal mean and they are positively correlated. This aligns with what we observed in Section (ref). With regard to the ANCR false $H_{i}$ in multiple testing, the closure procedure based on Hotelling's $T^{2} $ tests may outperform MinP tests in some cases but not in others. The combined test appears to have its global power bounded between the global powers of Hotelling's $T^{2}$ tests and MinP tests in most cases. In some cases the combined test considerably improves the global power of MinP tests with little compromise on the multiple testing power measured in terms of ANCR false $H_{i}$. The results based on the bootstrap approach reported in Tables (ref) and (ref) are very close to the results based on the limiting distribution approach reported in Tables (ref) and (ref) for the case $k=4$.

table*[table* omitted — 3,776 chars of source]
table*[table* omitted — 3,779 chars of source]
table*[table* omitted — 3,802 chars of source]
table*[table* omitted — 3,802 chars of source]
table*[table* omitted — 3,779 chars of source]
table*[table* omitted — 3,779 chars of source]

An empirical application

This section presents a real data application for evaluating the effectiveness of exercise using the data reported in chagne09. The data are available as supplementary materials on the Econometrica website and contain seven biometric measures of the participants. The measures are indicated in Table (ref). The participants are randomly divided into three groups: the control group (G1), the first treatment group who were paid \$25 to attend the gym once a week (G2), and the second treatment group who were paid an additional \$100 to attend the gym eight more times in the following four weeks (G3). G1 had 39 participants, G2 had 56 participants (after excluding one who had incomplete observations on some of variables) and G3 had 60 participants. See chagne09 for more details.

Denote by $X_{q,t_{q},i}$ the change from the initial measurement level to the final measurement level taken after 20 weeks for the $q$th group, the $ t_{q}$th participant and the $i$th measure with $q\in \{G1,G2,G3\}$, $ t_{q}\in \{1,...,n_{q}\}$, $n_{G1}=39$, $n_{G2}=56$, $n_{G3}=60$ and $i\in K=\{1,...,7\}$. Let $X_{q,t_{q}}=(X_{q,t_{q},1},...X_{q,t_{q},7})^{\prime }$ . Let $\bar{X}_{q}=n_{q}^{-1}\sum_{t_{q}=1}^{n_{q}}X_{q,t_{q}}$ be the sample mean for the $q$th group and $\hat{\Sigma}_{q}=(n_{q}-1)^{-1} \sum_{t_{q}=1}^{n_{q}}(X_{q,t_{q}}-\bar{X}_{q})(X_{q,t_{q}}-\bar{X} _{q})^{\prime }$ be the corresponding covariance matrix. Let $\mu _{q}=(\mu _{q,1},...,\mu _{q,k})^{\prime }$ be the population mean corresponding to $ \bar{X}_{q}$. churom16 studied permutation tests of multiple hypotheses

equation*[equation* omitted — 129 chars of source]

where $q_{1}$, $q_{2}\in \{G1,G2,G3\}$ and $q_{1}\neq q_{2}$, for each $i\in K$ across seven biometric measures. Their multiple tests were based on closed tests with the intersection null hypotheses $H_{J}$, $J\subseteq K$, being tested by either their modified Hotelling's $T^{2}$ test or their MaxT test, although they did not use the term `MaxT' as we do.

We implemented the combined test based on the modified Hotelling's $T^{2}$ test and the MaxT test proposed by churom16. The modified Hotelling's $T^{2}$ test had the test statistic

equation*[equation* omitted — 114 chars of source]

where $\bar{X}_{q_{1}q_{2}}=\bar{X}_{q_{1}}-\bar{X}_{q_{2}}$ and $\hat{\Sigma }_{q_{1}q_{2}}=\hat{\Sigma}_{q_{1}}+\frac{n_{q_{1}}}{n_{q_{2}}}\hat{\Sigma} _{q_{2}}$, for testing

equation*[equation* omitted — 121 chars of source]

The individual tests of $H_{i}$ used the test statistic

equation*[equation* omitted — 153 chars of source]

where $\hat{\Sigma}_{q_{1}q_{2}}^{(ii)}$ was the $(i,i)$th element of $\hat{ \Sigma}_{q_{1}q_{2}}$.

Following churom16, we generated $X^{d}=\bar{X}_{q_{1}q_{2}}$, $d\in D $, from the two-sample random permutations. For each $X^{d}\in D$, we bootstrapped $\hat{p}_{g}(X^{d})$ and $\hat{p}_{i}(X^{d})$ by following Algorithm 2.1 of churom16. The adjusted sample $p$-values ($\hat{p} _{g}^{adj}$ and $\hat{p}_{i}^{adj}$) were then computed as the proportions of the values in the permuted sample sequence $\{\hat{p}_{c}(X^{d}),d\in D\}$ that were less than or equal to the observed sample $\hat{p}_{g}$ and $\hat{p }_{i}$, respectively. The number of random permuted samples was set to $ 10,000$ (with $9999$ permuted samples generated plus the original sample). The number of bootstrap samples was set to $3000$.

We compared the results of the combined test with those of Chung and Romano's modified Hotelling's $T^{2}$ test and the MinP test, as well as the closed tests based on the modified Hotelling's $T^{2}$ test and the MinP test. Table (ref) presents the difference in the sample averages $\bar{X }_{q_{1}q_{2}}$ (column 2), the associated standard errors (s.e.) $ n_{q_{1}}^{-1/2}\sqrt{\hat{\Sigma}_{q_{1}q_{2}}^{(ii)}}$ (column 3) and $p$ -values. The $p$-values of the single-step MinP test are reported in column 4: they were computed as the proportion of $\{\min (\hat{p}_{i}(X^{d}),i\in K),d\in D\}$ that was less than or equal to the original sample $\hat{p}_{i}$ . The $p$-values of the modified Hotelling's $T^{2}$ test are also reported in column 4. The $p$-value of the closed tests for testing $H_{i}$ is reported as the largest $p$-value of those obtained in testing $H_{J}$ for all $J\subseteq K$ that involve $H_{i}$. Columns 5 and 6 report the $p$ -values of the closed tests based on the modified Hotelling's $T^{2}$ test and the MinP test, respectively. The adjusted $p$-values in the first step of the combined test are reported in column 7. The adjusted $p$-values of the combined test in the stepdown procedure are reported in column 8, where the adjustment was computed as the proportion of $\{\min (\hat{p} _{i}(X^{d}),i\in K_{i}),d\in D\}$ that was less than or equal to the smallest original sample $\hat{p}_{(i)}$ for each $(i)=(2),...,(k)$. We note that the $p$-values of the modified Hotelling's $T^{2}$ test and the MinP test reported here are slightly different from those reported in churom16. These differences may be attributed to the fact that churom16 appeared to have used 57 participants in the G2 group, while we used 56 participants. The difference may also be due to different random numbers in generating random permuted and bootstrap samples.

Although group comparisons reavealed different dominating effects, we observed that effects on some biometric measures dominated in all three group-wise comparisons. Therefore, it is not surprising that the MinP test tended to have a smaller $p$-value than the modified Hotelling's $T^{2}$ test in testing the global hypothesis $H_{K}$. However, we cannot determine if the stronger evidence presented by the MinP test was simply due to the sampling variation. The combined test, which combines the modified Hotelling's $T^{2}$ test and the MinP test, acts as a robust and honest check.

When comparing the control group with the first treatment group, the MinP test found insufficient evidence to suggest a difference between the two groups, specifically in body fat and pulse rate measures. In contrast, the modified Hotelling's $T^{2}$ test found insufficient evidence to suggest any difference between the two groups. The combined test revealed that the differences in body fat and pulse rate measures dominated those of the other measures, although they were not statistically significant. In comparing the control group with the second treatment group, the combined test, the MinP test and the modified Hotelling's $T^{2}$ test found significant evidence to suggest a difference between the two groups. Furthermore, all of the multiple testing procedures--namely the combined test, the MinP test and the closed procedures based on the modified Hotelling's $T^{2}$ test--indicated that the effect on body fat measure was significant. When comparing the first treatment group with the second treatment group, the MinP test and the modified Hotelling's $T^{2}$ test again yielded conflicting evidence in rejecting $H_{K}$; the modified Hotelling's $T^{2}$ test found no significant evidence, whereas the MinP test found significant evidence. However, unlike in the case of comparing the control group with the first treatment group, the combined test confirmed the findings of the MinP test, along with the closed procedure based on the MinP test, in both global and multiple hypothesis tests.

table*[table* omitted — 2,561 chars of source]

Conclusion

This paper proposes a combined test for global and multiple testing. The combined test exploits the global power advantages of the tests being combined while retaining the benefit of the stepdown procedure of MinP tests. It also offers a tool for the robust and honest evaluation of a global hypothesis when two tests with distinct power advantages are available for hypothesis testing. A simulation study is provided to illustrate the power performance of the combined test for global testing and multiple testing. An empirical application testing the effects of exercise is provided to illustrate the practical relevance of the proposed test.