Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
84,809 characters · 13 sections · 60 citation commands
Extending MinP Tests for Global and Multiple Testing
{3pt} {3pt}
Empirical economic studies often involve multiple propositions or hypotheses. In such studies, researchers are interested in evaluating the collective and individual evidence against these propositions or hypotheses. A prominent example in the recent literature is the evaluation of treatment effects that are measured by several variables. Two popular classes of statistical tests adopted by researchers are tests based on quadratic-form test statistics (QF tests), or tests based on the minimum of $p$-values, (or the maximum of the corresponding test statistics). QF tests evaluate overall treatment effects; they include the Wald, $F$- and Hotelling's $T^{2}$ tests (see chagne09, bendan16, angjorkue18, beaman21 and chosutzim22 for examples). The minimum $p$-values (MinP) tests evaluate the treatment effect relating to individual variable measurements, and look to control the multiplicity of Type I errors (see anderson08 , leesha14, gerhecetal14, lishxu16 and chosutzim22 for examples).
Generally speaking, QF tests have good global power in detecting overall treatment effects arising from the accumulation of many small individual treatment effects. However, a rejection by a QF test does not indicate which individual treatment effects are statistically significant, unless subsequent testing via the closed testing procedure finds further evidence (see, e.g., marpergar76, roazwo11, lu20cl and goehemsol19). In comparison, detection of any individual treatment effect in MinP tests implies an overall treatment effect. This tends to provide global power advantage over QF tests when one or a few individual treatment effects are present. MinP tests are popular in multiple testing; they can be viewed as a special case of closed tests, but have the computational advantage of sequential testing over general closed tests.
Naturally, the primary objective in the evaluation of treatment effects is the presence of the overall treatment effect. However, the researcher usually does not have a priori knowledge of the true distribution generating the data in an empirical investigation. Since QF and MinP tests have distinct global powers for detecting on the presence of the overall treatment effect, the researcher is likely to attempt both tests to examine empirical evidence. For example, the researcher may attempt the other test if one test does not provide the evidence he/she anticipates, or they may initially aim to identify as many possible variables that relate to the treatment effect as possible, but turns to QF tests after the initial attempt fails to identify any variables. In fact, we argue that the researcher should attempt both tests for a robustness check on how sensitive the empirical evidence is to the distribution generating the data.
If both the tests are carried out, but only the results of one test are reported, or a test is chosen based on sample information, then the error probability of claimed findings may not be controlled at the desired level. For example, if the researcher chooses to report the more significant evidence on the overall treatment effect from the two tests attempted, then the error probability of claiming positive findings when there are no treatment effects (Type I error) becomes inflated beyond the level the researcher intends to control. Over-reporting has drawn increasing concerns in the literature. Recent papers by christensenmiguel18 and young19 provide excellent accounts of the issue. In his influential paper, wh2000 draws attention to the dangerous practice of data snooping with primary concern on testing predictive superiority when different forecasting models are explored (see also hansenpr05, romwol05 and roshwo08).
This paper proposes an adjusted $p$-value method which offers a tool for robust and honest reporting. In fact, we show that the method can be viewed as a combined test constructed through the minimum $p$ value principle, which is a special case of the generalized mean of $p$-values for combining tests (cf. vovkwang20 and vovkwang22). We show that our combined tests are admissible if both constituent tests are admissible. Furthermore, we provide the necessary and sufficient condition under which the combined test improves the global power of the constituent tests. The combined test also retains the stepdown procedure of MinP tests when simultaneously testing individual treatment effects, and preserves the control of the familywise error rate (FWER) by the MinP test.
The $p$-values in the combined test can be easily computed based on the Bonferroni correction; more precise adjustments may be computed through resampling schemes (e.g., leesha14, churom16 and bugcansha19). Computational implementations are readily available in popular packages such as R and Stata (see brhowe16 and claromwol20), which can be easily extended to compute our adjusted $p$ -values through the inclusion of an additional $p$-value associated with a global test such as a QF test.
Related literature. The idea of combining different tests can be traced back to tippett31, pearson33 and fisher36. There has been renewed interest in combining tests in the recent literature. vovkwang20 and vovkwang22 studied a class of combination tests based on the generalized mean of $p$-values and suggested the closed testing procedure based on their combination tests for multiple testing. Our test is a Tippett type test constructed based on the minimum of the $p$-values of the constituent tests, and may be viewed as a special case of the generalized mean of $p$-values. lu16 proposed an extended MaxT (EMaxT) test by combining a sum test and a MaxT test for one-sided testing where the sum test is chosen to direct to the `middle' of the constrained parameter space under the global hypothesis. Our tests here can be viewed as a generalization of the EMaxT test to a more general setting, by combining p-values to allow for the adoption of a general global test, as well as retaining a stepdown procedure of multiple testing. Both the EMaxT test and the combined test proposed in this paper have a common feature of preserving the rejection region shapes of constituent tests, hence somewhat inheriting the respective strength of the constituent tests with regard to global power. In fact, the properties studied in this paper also apply to the EMaxT test. In high dimensional problems, combination tests of maximum type tests and QF tests have drawn interests in the recent literature due to their distinct global powers against sparse and dense alternatives. fanliayao15 proposed tests aiming at enhancing the global power of QF tests in high-dimensional settings. Their tests combine a QF test and some screened $t$-tests where the screened $t$-test components have an asymptotic size $0$. kockpre23power provide sufficient conditions for fanliayao15's test to achieve global power enhancement. hexuwupan21 and fejilixi20 derived approximated distributions of constituent tests in high-dimensional settings, and studied the properties of subsequent combination tests. However, these high-dimensional combination tests are not devised for multiple testing, but rather for global testing. There are also other combination tests proposed in the recent literature. andrewsI16 studied a class of combination tests for testing of weakly identified models by exploring global power advantages of the two distinct tests in combination. helmeicha19post constructed a combined test for multiple testing based on the marginal $p$-values conditional on a global test. Our paper contributes to this growing body of literature by proposing a combined test that allows for global and multiple testings simultaneously. In practice, the set of individual hypotheses rejected by our combined test is the same as that of the constituent MinP test once an individual hypothesis is rejected by our combined test, while revealing a potentially much stronger signal on global testing.
The remainder of the paper is organized as follows. We motivate and illustrate our combined test in the case of a multivariate normal location model in Section (ref). We then present the combined test in a more general setting in Section (ref) where some properties concerning the combined test are established. In Section (ref) simulation studies are provided to examine the performance of the combined test compared with other tests. Section (ref) provides a real data application for testing the effects that exercise has on seven biometric measures based on the data published in chagne09. Concluding remarks are made in Section (ref). Proofs are presented in the appendix.
To fix ideas, we begin with an example of testing the multivariate normal mean with a known covariance. Let $X=(X_{1},...,X_{k})^{\prime }\sim N(\mu ,\Sigma )$, where $\mu =(\mu _{1},...,\mu _{k})^{\prime }$ and $\Sigma $ has the structure of the equicorrelation matrix $\{\rho _{ij}\}$, $i,j\in K=\{1,\cdots ,k\}$, with $\rho _{ij}=\rho $, $-1<\rho <1$, when $i\neq j$ and $\rho _{ij}=1$, when $i=j$. The individual hypotheses are:
Let $\Phi (\cdot )$ be the cumulative distribution function (CDF) of the standard normal random variable and $F_{\chi _{k}^{2}}(\cdot )$ be the CDF of the central chi-square random variable with $k$ degrees of freedom, and $ F_{\chi _{k}^{2}}^{-1}(\cdot )$ be the inverse function of $F_{\chi _{k}^{2}} $. Denote the individual $p$-value by $\hat{p}_{i}(X_{i})=2\Phi (-\left\vert X_{i}\right\vert )$.
Consider the two-sided testing with $k=2$. With the control of the FWER, the probability of rejecting at least one true $H_{i}$ (more detailed discussions on the FWER are provided in Section (ref)) , MinP tests would reject $H_{i}$ if
where $\alpha \in (0,1)$ is the significance level, and $c_{m}(\alpha )$ satisfies
Consider the global null hypothesis $H_{K}:\mu =0$ which is the intersection of all individual null hypotheses $\cap _{i=1}^{k}H_{i}$. Rejection of any $ H_{i}$ implies rejection of $H_{K}$. Therefore, MinP tests can be used to jointly test $H_{K}$ against $H_{K}^{\prime }:\mu \neq 0$ and reject $H_{K}$ if any $\hat{p}_{i}(X_{i})<c_{m}(\alpha )$ or $\min \{\hat{p} _{i}(X_{i}),i\in K\}<c_{m}(\alpha )$. If the researcher uses the Likelihood Ratio (LR) tests for testing $H_{K}$, the test statistic $X^{\prime }\Sigma ^{-1}X$ follows the null distribution $\chi _{k}^{2}$ and $H_{K}$ would be rejected if
Figure (ref) shows the comparison of the rejection regions of LR tests and MinP tests in the case of $k=2$ and $\rho =0$ with $\alpha =0.05$. The rejection region of MinP tests is
where $c_{m}(\alpha )=1-(1-\alpha )^{1/2}=0.0253$. $S_{m}(\alpha )$ represents the area outside the square box. The rejection region of LR tests is
which is the area outside the circle. For $X\in A(\alpha )=S_{g}(\alpha )\cap S_{m}^{c}(\alpha )$, where $S^{c}$ is the complement set of $S$, LR tests reject $H_{K}$, but MinP tests do not reject $H_{K}$. For $X\in B(\alpha )=S_{g}^{c}(\alpha )\cap S_{m}(\alpha )$, MinP tests reject $H_{K}$ , but LR tests do not reject $H_{K}$. MinP tests are more likely to reject $ H_{K}$ than LR tests when one of $\left\vert X_{i}\right\vert $, $i=1,2$, dominates the other. However, as the boundary of the rejection region of LR tests is defined by the circle $X_{1}^{2}+X_{2}^{2}=F_{\chi _{2}^{2}}^{-1}(0.95)=5.99$, LR tests are more likely to reject $H_{K}$ than MinP tests when none of $\left\vert X_{i}\right\vert $, $i=1,2$, dominates the other. If one chooses LR tests or MinP tests based on the sample information, it would inevitably lead to a data snooping problem. For example, if one chooses LR tests or MinP tests depending on the outcome of rejection, it would effectively lead to the enlarged rejection region as $ S_{g}(\alpha )\cup B(\alpha )$ or $S_{m}(\alpha )\cup A(\alpha )$. An enlarged rejection region implies an inflated size. For example, the inflated size is about $0.07$ at $\alpha =0.05$ when $\rho =0.9$ or $-0.9$.
With regards to multiple testing of $H_{i}$, $i=1,2$, MinP tests reveal the evidence on testing $H_{i}$ with the control of the FWER. A rejection of $ H_{K}$ by LR tests itself does not directly indicate which $H_{i}$ should be rejected unless $H_{i}$ is rejected as well. This is the so-called closure testing procedure in which the rejection of an $H_{i}$ requires rejections of all $H_{K_{i}}$, $\{i\}\subseteq K_{i}\subseteq K$, in a general case of $ k\geq 2$. The rejection region of closed tests in the case of $k=2$ is
where $S_{i}(\alpha )=\{X:\hat{p}_{i}(X_{i})\leq \alpha \}$. This rejection region is a strict subset of $S_{g}(\alpha )$; hence, the FWER of closed tests based on $S_{g}(\alpha )$ is strictly less than $\alpha $. As $k $ increases, the FWER control becomes more conservative and consequently, the capacity to detect false $H_{i}$ is reduced.
In the above $k=2$ case, it follows that
where $a\wedge b$ defines $\min (a,b)$, and we may use both the operations interchangeably throughout the paper. To maintain the relative strength of the global power of LR tests and MinP tests, one may preserve the shape of the combined rejection region by adjusting it through an $\alpha ^{\prime }<\alpha $ such that the combined rejection region has a probability of $ \alpha $ under $H_{K}$, namely,
This implies that for an observed $X=x$, the adjusted $p$-value $\hat{p} _{l}^{adj}(x)$, $l=g,m,i$, can be computed as
One may compute the null distribution
based on Monte Carlo simulation by randomly drawing $X$ from $N(\mu ,\Sigma ) $ if it is assumed known, or through the permutation or bootstrap resampling methods based on the sample. One may opt to ease the computational burden by computing the adjusted $p$-values based on, for example, the Bonferroni inequality. More detailed discussions are presented in Section (ref).
Let $\hat{p}_{c}^{adj}(X)=\hat{p}_{g}^{adj}(X)\wedge \hat{p}_{m}^{adj}(X)$. The rejection region based on the combined test is
which is equivalent to
where $\alpha ^{\prime }<\alpha $ satisfies ((ref)). The rejection rule of the combined test is to reject $H_{K}$ if $\hat{p}_{c}^{adj}(x)\leq \alpha $, not to reject it otherwise.
To compare the power performance of the combined test of $H_{K}$ with LR tests and MinP tests, we approximate power functions of tests considered in Section (ref) based on $1,000,000$ and $100,000$ independent random draws from $N(\mu ,\Sigma )$ for the cases of $k=2$ and $k\geqslant 2$, respectively. The exception is the global power function of LR tests which is computed as $\Pr \{\chi _{k}^{2}(r^{2})>F_{\chi _{k}^{2}}^{-1}(0.95)\}$, where $\chi _{k}^{2}(r^{2})$ is the chi-square random variable with $k$ degrees of freedom and the non-centrality parameter $r^{2}$. Figure (ref) presents the comparison of the global powers of testing $H_{K}$ for LR, MinP and the combined test with $\alpha =0.05$ in the bivariate case. We take $ \mu _{1}=r\cos \varphi $, $r=2$ and
so that $\mu ^{\prime }\Sigma ^{-1}\mu =r^{2}$. The comparison shows that LR tests have an overall global power advantage over MinP tests. However, MinP tests can outperform LR tests when either $\left\vert \mu _{1}\right\vert $ or $\left\vert \mu _{2}\right\vert $ dominates the other. The combined test somewhat inherits the respective strengths of LR and MinP tests with regard to global power. In relation to multiple testing, the combined test can outperform closed tests. However, MinP tests are more likely to reject $ H_{i} $, $i\in K$, than the combined test. This is because $\hat{p} _{g}(x)\wedge \hat{p}_{m}(x)\leq \hat{p}_{m}(x)$ for every $x\in X$ with the strict inequality holding for some $x\in X$. Figure (ref) presents a comparison of the average number of correctly rejected (ANCR) false $H_{i}$ for closed, MinP and the combined test.
To further illustrate that the combined test can share the power strength of MinP tests to some extent, Figures (ref) and (ref) present comparisons of the global power in testing $H_{K}$, as well as comparisons of the probability of rejecting $H_{1}$ in multiple testing as the number of hypotheses $k$ increases. The comparison in Figure (ref) is based on the case of $\Sigma =I$ and $\mu =(3,0,...,0)^{\prime }$, while that in Figure (ref) is based on the case of $\rho _{ij}=0.9$, $i\neq j$, $i,j\in K$, and $\mu =(3,...,3)^{\prime }$. The comparisons show that MinP tests have a clear power advantage over LR tests in both global and multiple testing. The advantage becomes increasingly apparent as $k$ increases. The combined test in such cases share some of the strength of MinP tests.
As for MinP tests in multiple testing, an improved ability to reject more $H_{i}$, $i\in K$, is possible for the combined test through a stepdown testing procedure. For example, for the points
in the bivariate example, the combined test rejects $H_{1}$, but not $H_{2}$ . However, if we proceed in the same fashion as in the stepdown procedure of MinP tests, one would then reject $H_{2}$ in the second step because $\hat{p} _{2}(X_{2})<\alpha $. Although the combined test has a disadvantage compared with MinP tests in the first step of multiple testing, the combined test would have the same outcomes as MinP tests in the stepdown procedure if the null hypothesis with the smallest $\hat{p}_{i}(X_{i})$ is rejected by the combined test in the first step of the stepdown procedure.
We study the combined test in a more general setting. Suppose that the sample $X^{(n)}$, where $n$ indicates sample size, is generated from the unknown distribution $P\in \mathbf{P}$, where $\mathbf{P}$ defines a set of probability distributions. Let $\hat{p}=\hat{p}(X^{(n)})$ be a $p$-value. Let $G^{(n)}(u,P)$, $u\in \lbrack 0,1]$, be a sequence of CDFs of $\hat{p}$ under $P\in \mathbf{P}$. Denote the test function by
where $\mathbf{1(\cdot )}$ is the usual indication function. Note that for ease of presentation, we consider nonrandomized tests in this paper.
Let $H_{i}$, $i\in K=\{1,...,k\}$, $k\geq 2$, be the individual null hypotheses and $H_{i}^{\prime }$ be the corresponding alternative hypotheses. Let the set of distribution under $H_{i}$ be $\mathbf{P} _{i}\subset \mathbf{P}$. Let $K_{i}\subseteq K$ be a sub-index set, and $ K_{\ast }\subseteq K$ be the set containing the indices of true $H_{i}$. (The subscript $i$ in $K_{i}$ will be useful when we discuss the stepdown procedure later on). Denote by $\mathbf{P}_{K}=\cap _{i\in K}\mathbf{P} _{i}\subset \mathbf{P}$ the set of null distributions corresponding to the global null hypothesis $H_{K}$ and by $\mathbf{P}_{K}^{\prime }=\mathbf{ P\smallsetminus P}_{K}$ the set of distributions corresponding to the alternative hypothesis $H_{K}^{\prime }$. Assume $\mathbf{P}_{K}\subset \mathbf{P}_{i}$, for all $i\in K$. That is, $\mathbf{P}_{K}$ is the strict subset of $\mathbf{P}_{i}$, for all $i\in K$.
A test $\phi ^{(n)}$ of $H_{K}$ is referred to as the asymptotic pointwise level-$\alpha $ test if
where $E_{P\in \mathbf{P}_{K}}(\cdot )$ is the expected value with respect to $P\in \mathbf{P}_{K}$. (In this paper, we restrict our attention to pointwise control.) In relation to multiple testing, the FWER is the probability of rejecting any $H_{i}$, $i\in K_{\ast }$, under the true $P\in \mathbf{P}_{K_{\ast }}$. That is,
The asymptotic pointwise FWER control at the level $\alpha $ based on the sample $X^{(n)}$ is achieved if
It is worth noting that any distribution restricted by a possible configuration of true and null hypotheses belong to $\mathbf{P}_{K_{\ast }}$ . The above definition of the FWER control is known as strong control of the FWER.
Since the true null set $K_{\ast }$ is typically unknown, nor is the true $ P\in \mathbf{P}_{K_{\ast }}$, assuming the true $P\in \mathbf{P}_{K}$ (which is referred to as the weak control in the literature) does not guarantee the control of the FWER (see, e.g., romwol05b). However, in many applications, the researcher may be able to assume the subset pivotality condition of wesyou93, which says that the true $P\in \mathbf{P} _{K_{\ast }}$ is not affected by whether $H_{i}$, $i\in K\mathbf{ \smallsetminus }K_{\ast }$, is true or not. Hence, the FWER can be controlled by assuming the true $P\in \mathbf{P}_{K}$. The subset pivotality condition is not a necessary condition for controlling FWER. romwol05b showed a weaker sufficient condition for controlling FWER that is also satisfied under the subset pivotality condition. We shall show later that the proposed combined test controls the FWER as long as the MinP test being combined controls the FWER.
Let the subscripts $c,g$, $m$, and $i$ indicate the combined, global, MinP and individual tests, respectively. The $p$-value of the combined test is
The adjusted $p$-value of the combined test for testing $H_{K}$ is defined as
or equivalently,
where
The limiting null distribution of $G_{c}^{(n)}(\cdot ,P\in \mathbf{P}_{K} \mathbf{)}$ is usually unknown for our combined tests, so we use $ G_{c}^{(n)}(\cdot )$ for computing $\hat{p}_{c}^{adj}$. The computation may be based on simulation methods by utilizing some limiting distribution, or based on resampling methods such as bootstrap, permutation, and subsampling.
The researcher may opt to ease the computational burden by computing conservative adjusted $p$-values based on the Bonferroni inequality as
or
then compute $\hat{p}_{c}^{adj}$ based on ((ref)). One may improve the conservativeness of the Bonferroni inequality for computing $\hat{p} _{m}^{adj}$ in testing $H_{K}$ by using inequalities such as the {\v{S}}id{ \'{a}}k inequality (sidak68) or the Simes inequality (simes86 and sarkar98). However, these inequalities are not generally applicable.
The combined test rejects the global null hypothesis $H_{K}$ if $\hat{p} _{c}^{adj}\leq \alpha $; otherwise it accepts $H_{K}$. We first present the size property of the combined test. We refer to the procedure in which $\hat{ p}_{c}^{adj}$ is computed based on $G_{c}^{(n)}(\cdot )$ as $\phi _{c,1}^{(n)}$ and to the procedure based on the Bonferroni inequality as $ \phi _{c,2}^{(n)}$. The $\phi _{c,2}^{(n)}$ procedure may be conservative in the sense that the test size may be strictly less than $\alpha $, but it is computationally easy to implement. Note that to keep the presentation concise $\phi _{c}^{(n)}$ will be used to represent both $\phi _{c,1}^{(n)}$ and $\phi _{c,2}^{(n)}$ when no confusion is deemed to arise.
The following lemma concerns the size property of the combined test.
We now turn to study some global properties. Consider a set of local alternatives
A test $\phi _{l}^{(n)}$ is asymptotically $d$-admissible if for any other test $\phi ^{(n)}$
and
jointly imply
for all $P\in \mathbf{P}$. The $d$-admissibility implies that $\phi _{l}^{(n)}$ can have better global power for some $P_{n}\in \mathbf{P} _{n,K}^{\prime } $ asymptotically compared with any other tests that do not have an asymptotically larger size.
Although Theorem (ref) states that both the test $\phi _{c,1}^{(n)}$ and $\phi _{c,2}^{(n)}$\ are asymptotically $d$-admissible, the power of $ \phi _{c,2}^{(n)}$ may be asymptotically uniformly improved by $\phi _{c,1}^{(n)}$. This improvement typically occurs when $\phi _{c,2}^{(n)}$ has a test size strictly less than the nominal level $\alpha $. In comparison to $d$-admissibility, $\alpha $-admissibility is defined as follows: for any other level-$\alpha $ test $\phi ^{(n)}$, ((ref)) implies ((ref)) for all $P\in \{P_{n}:P_{n}\in \mathbf{P} _{n,K}^{\prime }\}$ (lehrom05). In other words, the $\alpha $ -admissibility of a test demands that there does not exist any other test that has better power for at least some $P\in \mathbf{P}_{K}$ and non-worse power for all other $P\in \mathbf{P}_{K}$.
The following theorem provides a necessary and sufficient condition under which the combined test enhances the power of a constituent test. Let $ \tilde{\phi}_{l}^{(n)}=\mathbf{1}(\hat{p}_{l}^{adj}\leq \alpha )$, $l\in \{g,m\}$, be the test function based on the adjusted $p$-value instead of the unadjusted $p$-value used in $\phi _{g}^{(n)}$ and $\phi _{m}^{(n)}$.
The condition ((ref)) indicates that the power gain to $\tilde{\phi} _{l^{\prime }}^{(n)}$ from $\tilde{\phi}_{l}^{(n)}$ is asymptotically sufficient to offset the power loss of $\phi _{l}^{(n)}$ caused by the adjustment of its $p$-value. This condition reveals the sources of the power gain and loss in relation to the global power improvement of the combined test. We illustrate the condition ((ref)) in the bivariate example in Section (ref). Let $X_{i}=\sqrt{n}\hat{\mu}_{i}$, $i=1,2$, and $ \sqrt{n}(\hat{\mu}-\mu )\sim N(0,\Sigma )$. Let $\tilde{S}_{g}(\alpha )=\{X: \hat{p}_{g}^{adj}(X)\leq \alpha ,\mu =0\}$ and $\tilde{S}_{m}(\alpha )=\{X: \hat{p}_{m}^{adj}(X_{i})\leq \alpha ,\mu =0\}$. Let
Without loss of generality we let $l=m$ and $l^{\prime }=g$. Then
while
If we decompose $S_{m}(\alpha )-\tilde{S}_{m}(\alpha )$ into the two exclusive regions
it follows
where $(\tilde{B}(\alpha )-\Delta _{1})$ and $\Delta _{2}$ are exclusive. The rejection region for $\phi _{c}^{(n)}$ is
For the combined test $\phi _{c}^{(n)}$ to exhibit a better global power in testing $H_{K}$ under $P\in \mathbf{P}_{n,K}^{\prime }$ which corresponds to the local alternatives $\sqrt{n}\mu $, the following condition must hold
which is equivalent to
When either $\mu _{i}$, $i=1,2$, dominates the other, the first term in the left hand side of ((ref)) is likely dominate the second term. Consequently, the combined test improves the global power of the MinP test.
Theorem (ref) demonstrates that the combined test can improve the global power of the constituent tests under the condition stated in (ref). For a $P\in \mathbf{P}_{n,K}^{\prime }$, this improvement is more probable for the constituent test with a lower global power compared to the other test. This result is formally articulated in the next theorem, which establishes that the combined test exhibits a more balanced global power, with its global power is bounded between that of two constituent tests.
The combined test may be viewed as an extended MinP test by adding another $ p $-value for testing the global $H_{K}$ to the minimand set of MinP tests. If an individual null hypothesis $H_{i}$, $i\in K$, is rejected after it has been adjusted for the extended multiplicity, then the combined test proceeds to the usual stepdown procedure of MinP tests for multiple testing. It is well-known that the stepdown procedure has better power in rejecting false individual null hypotheses than the single-step procedure (c.f. romwol05b).
Let
denote the ordered $p$-values $\hat{p}_{i}$, $i\in K$, and let $H_{(1)}$, $ H_{(2)}$, ..., $H_{(k)}$, be the corresponding null hypotheses. If $\hat{p} _{m}^{adj}\leq \alpha $, which is equivalent to $\min (\hat{p} _{i}^{adj},i\in K)\leq \alpha $, where $\hat{p}_{i}^{adj}=G_{c}^{(n)}(\hat{p} _{i},P\in \mathbf{P}_{K}\mathbf{)}$, then reject $H_{(1)}$ and the combined test proceeds to the usual stepdown procedure of MinP tests of the remaining $H_{(2)}$, ..., $H_{(k)}$. Let $G_{m,K_{i}}^{(n)}(\cdot ,P\in \mathbf{P} _{K_{i}}\mathbf{)}$, $K_{i}\subset K$, be the CDF of $\min (\hat{p}_{i},i\in K_{i})$. Let the adjusted $p$-value associated with testing $H_{(i)}$ be
where $K_{i}=\{(i),...,(k)\}$. Then, usual MinP tests may be implemented for $(i)=(2),...,(k)$: reject $H_{(i)}$ if $\hat{p}_{m,(i)}^{adj}\leq \alpha $, stop otherwise. One may compute the adjusted $p$-values based on the Bonferroni inequality as
where $\left\vert \cdot \right\vert $ is the cardinality of $K_{i}$.
One would expect the combined test to inherit properties of multiple testing from MinP tests to some degree. Compared with the stepdown procedure of MinP tests, our combined test differs only in the first step of rejecting $ H_{(1)} $; the combined test rejects $H_{(1)}$ if $\hat{p} _{(1)}^{adj}=G_{c}^{(n)}(\hat{p}_{(1)},P\in \mathbf{P}_{K})\leq \alpha $ whereas MinP tests reject $H_{(1)}$ if $\hat{p}_{m,(1)}^{adj}=G_{m,K}^{(n)}( \hat{p}_{(1)},P\in \mathbf{P}_{K})\leq \alpha $. Since $G_{m,K}^{(n)}(u,P\in \mathbf{P}_{K})\leq G_{c}^{(n)}(u,P\in \mathbf{P}_{K})$ for every $u\in \lbrack 0,1]$, it follows $\hat{p}_{m,(1)}^{adj}\leq \hat{p}_{(1)}^{adj}$. Thus, $\hat{p}_{(1)}^{adj}\leq \alpha $ implies $\hat{p}_{m,(1)}^{adj}\leq \alpha $; a rejection of $H_{(1)}$ by the combined test implies the rejection by the MinP test. Once $H_{(1)}$ is rejected by the combined test both tests share the same stepdown procedure, hence share the same testing outcome in terms of testing the remaining $H_{i}$, $i\in K\smallsetminus \{(1)\}$.
The FWER control requires that the probability of rejecting at least one true $H_{i}$, $i\in K_{\ast }$, is bounded above by the designated level $ \alpha $. Because the true null set $K_{\ast }$ is typically unknown to the researcher, the control of the FWER is not guaranteed in the stepdown procedure. romwol05b showed that MinP tests control the FWER under a monotonicity condition, which is a weaker condition than the subset pivotality condition of wesyou93. The nest theorem shows that the combined test controls the FWER so long as the MinP test being combined controls the FWER.
To facilitate the application of our proposed combined test, this section summarizes the procedure as follows.
This section reports a simulation study on tests of the multivariate mean. The significance level is set to $0.05$. The number of replications is set to 2000. Let $X^{(n)}=\{X_{t}=(X_{t1},...,X_{tk})^{\prime },t=1,...,n\}$, where $X_{t}$ is an independent $k$-dimensional random vector from the multivariate normal distribution with the mean $\mu =(\mu _{1},...,\mu _{k})^{\prime }$ and covariance $\Sigma _{X}$. In our simulation study, we let $\Sigma _{X}$ have the correlation matrix structure $\{\rho _{ij}\}$, $ i,j\in K$ as in Section (ref). Let $\bar{X}=(\bar{X}_{1},...,\bar{X} _{k})^{\prime }$ be the studentized sample mean, and let $\hat{\Sigma}$ be the correlation matrix corresponding to $\hat{\Sigma}_{X}=(n-1)^{-1} \sum_{t=1}^{n}(X_{t}-\bar{X})(X_{t}-\bar{X})^{\prime }$. By the multivariate central limit theorem it follows that
where $\hat{\Sigma}^{1/2}$ is the matrix such that $\hat{\Sigma}^{1/2}\hat{ \Sigma}^{1/2}=\hat{\Sigma}$.
We conducted simulation studies on two-sided testing of $\mu $ to compare the performance of the $\phi _{c^{(1)}}^{(n)}$ test with that of other tests. Hotelling's $T^{2}$ tests are adopted for testing the global null hypothesis $H_{K}$, as well as for implementing the combined test and closed tests. For an individual test of $H_{i}:\mu _{i}=0$ against $H_{i}^{\prime }:\mu _{i}\neq 0$, $i\in K$, the test statistic is $T_{i}=\left\vert n^{ \frac{1}{2}}\bar{X}_{i}\right\vert $. We consider two approaches for computing $p$-values. One is based on the bootstrap method and the other is based on a limiting normal distribution.
In our Monte Carlo simulation study the design of the correlation matrix $ \Sigma _{X}=\left\{ \rho _{ij}\right\} $ has two structures. One takes the form of the equicorrelation matrix, and the other takes $\rho _{ij}=a^{\left\vert i-j\right\vert }$ with $a=0.5$ or $-0.5$. In the bootstrap approach $2000$ bootstrap samples are used. In the limiting distribution approach, we compute $\hat{p}_{i}=2\Phi (-\left\vert n^{\frac{1 }{2}}\bar{X}_{i}\right\vert )$ and $\hat{p}_{g}=1-G_{\chi _{k}^{2}}(n\bar{X} ^{\prime }\hat{\Sigma}^{-1}\bar{X})$, and use $10,000$ random draws to approximate the distribution $G_{c}^{(n)}(u,P)$ and $G_{m,K_{i}}^{(n)}(u,P) $ , $K_{i}\subseteq K$. Tables (ref)--(ref) report the results of the simulation study. The results for the global hypothesis testing are reported in terms of estimated sizes and powers. The results for the multiple hypothesis testing are reported in terms of the estimated FWER and the estimated ANCR false $H_{i}$. The notation $\mathbbm{1}_{k,m}$ represents the $k$-dimension column vector with the first $m$ elements being $1$ and the remaining elements being $0$. We observe that Hotelling's $T^{2}$ tests outperform MinP tests in global power in many instances, but not in all cases. The global power performance of MinP tests can increase by up to 20% in cases where the individual $X_{ti}$, $i\in K$, has an equal mean and they are positively correlated. This aligns with what we observed in Section (ref). With regard to the ANCR false $H_{i}$ in multiple testing, the closure procedure based on Hotelling's $T^{2} $ tests may outperform MinP tests in some cases but not in others. The combined test appears to have its global power bounded between the global powers of Hotelling's $T^{2}$ tests and MinP tests in most cases. In some cases the combined test considerably improves the global power of MinP tests with little compromise on the multiple testing power measured in terms of ANCR false $H_{i}$. The results based on the bootstrap approach reported in Tables (ref) and (ref) are very close to the results based on the limiting distribution approach reported in Tables (ref) and (ref) for the case $k=4$.
This section presents a real data application for evaluating the effectiveness of exercise using the data reported in chagne09. The data are available as supplementary materials on the Econometrica website and contain seven biometric measures of the participants. The measures are indicated in Table (ref). The participants are randomly divided into three groups: the control group (G1), the first treatment group who were paid \$25 to attend the gym once a week (G2), and the second treatment group who were paid an additional \$100 to attend the gym eight more times in the following four weeks (G3). G1 had 39 participants, G2 had 56 participants (after excluding one who had incomplete observations on some of variables) and G3 had 60 participants. See chagne09 for more details.
Denote by $X_{q,t_{q},i}$ the change from the initial measurement level to the final measurement level taken after 20 weeks for the $q$th group, the $ t_{q}$th participant and the $i$th measure with $q\in \{G1,G2,G3\}$, $ t_{q}\in \{1,...,n_{q}\}$, $n_{G1}=39$, $n_{G2}=56$, $n_{G3}=60$ and $i\in K=\{1,...,7\}$. Let $X_{q,t_{q}}=(X_{q,t_{q},1},...X_{q,t_{q},7})^{\prime }$ . Let $\bar{X}_{q}=n_{q}^{-1}\sum_{t_{q}=1}^{n_{q}}X_{q,t_{q}}$ be the sample mean for the $q$th group and $\hat{\Sigma}_{q}=(n_{q}-1)^{-1} \sum_{t_{q}=1}^{n_{q}}(X_{q,t_{q}}-\bar{X}_{q})(X_{q,t_{q}}-\bar{X} _{q})^{\prime }$ be the corresponding covariance matrix. Let $\mu _{q}=(\mu _{q,1},...,\mu _{q,k})^{\prime }$ be the population mean corresponding to $ \bar{X}_{q}$. churom16 studied permutation tests of multiple hypotheses
where $q_{1}$, $q_{2}\in \{G1,G2,G3\}$ and $q_{1}\neq q_{2}$, for each $i\in K$ across seven biometric measures. Their multiple tests were based on closed tests with the intersection null hypotheses $H_{J}$, $J\subseteq K$, being tested by either their modified Hotelling's $T^{2}$ test or their MaxT test, although they did not use the term `MaxT' as we do.
We implemented the combined test based on the modified Hotelling's $T^{2}$ test and the MaxT test proposed by churom16. The modified Hotelling's $T^{2}$ test had the test statistic
where $\bar{X}_{q_{1}q_{2}}=\bar{X}_{q_{1}}-\bar{X}_{q_{2}}$ and $\hat{\Sigma }_{q_{1}q_{2}}=\hat{\Sigma}_{q_{1}}+\frac{n_{q_{1}}}{n_{q_{2}}}\hat{\Sigma} _{q_{2}}$, for testing
The individual tests of $H_{i}$ used the test statistic
where $\hat{\Sigma}_{q_{1}q_{2}}^{(ii)}$ was the $(i,i)$th element of $\hat{ \Sigma}_{q_{1}q_{2}}$.
Following churom16, we generated $X^{d}=\bar{X}_{q_{1}q_{2}}$, $d\in D $, from the two-sample random permutations. For each $X^{d}\in D$, we bootstrapped $\hat{p}_{g}(X^{d})$ and $\hat{p}_{i}(X^{d})$ by following Algorithm 2.1 of churom16. The adjusted sample $p$-values ($\hat{p} _{g}^{adj}$ and $\hat{p}_{i}^{adj}$) were then computed as the proportions of the values in the permuted sample sequence $\{\hat{p}_{c}(X^{d}),d\in D\}$ that were less than or equal to the observed sample $\hat{p}_{g}$ and $\hat{p }_{i}$, respectively. The number of random permuted samples was set to $ 10,000$ (with $9999$ permuted samples generated plus the original sample). The number of bootstrap samples was set to $3000$.
We compared the results of the combined test with those of Chung and Romano's modified Hotelling's $T^{2}$ test and the MinP test, as well as the closed tests based on the modified Hotelling's $T^{2}$ test and the MinP test. Table (ref) presents the difference in the sample averages $\bar{X }_{q_{1}q_{2}}$ (column 2), the associated standard errors (s.e.) $ n_{q_{1}}^{-1/2}\sqrt{\hat{\Sigma}_{q_{1}q_{2}}^{(ii)}}$ (column 3) and $p$ -values. The $p$-values of the single-step MinP test are reported in column 4: they were computed as the proportion of $\{\min (\hat{p}_{i}(X^{d}),i\in K),d\in D\}$ that was less than or equal to the original sample $\hat{p}_{i}$ . The $p$-values of the modified Hotelling's $T^{2}$ test are also reported in column 4. The $p$-value of the closed tests for testing $H_{i}$ is reported as the largest $p$-value of those obtained in testing $H_{J}$ for all $J\subseteq K$ that involve $H_{i}$. Columns 5 and 6 report the $p$ -values of the closed tests based on the modified Hotelling's $T^{2}$ test and the MinP test, respectively. The adjusted $p$-values in the first step of the combined test are reported in column 7. The adjusted $p$-values of the combined test in the stepdown procedure are reported in column 8, where the adjustment was computed as the proportion of $\{\min (\hat{p} _{i}(X^{d}),i\in K_{i}),d\in D\}$ that was less than or equal to the smallest original sample $\hat{p}_{(i)}$ for each $(i)=(2),...,(k)$. We note that the $p$-values of the modified Hotelling's $T^{2}$ test and the MinP test reported here are slightly different from those reported in churom16. These differences may be attributed to the fact that churom16 appeared to have used 57 participants in the G2 group, while we used 56 participants. The difference may also be due to different random numbers in generating random permuted and bootstrap samples.
Although group comparisons reavealed different dominating effects, we observed that effects on some biometric measures dominated in all three group-wise comparisons. Therefore, it is not surprising that the MinP test tended to have a smaller $p$-value than the modified Hotelling's $T^{2}$ test in testing the global hypothesis $H_{K}$. However, we cannot determine if the stronger evidence presented by the MinP test was simply due to the sampling variation. The combined test, which combines the modified Hotelling's $T^{2}$ test and the MinP test, acts as a robust and honest check.
When comparing the control group with the first treatment group, the MinP test found insufficient evidence to suggest a difference between the two groups, specifically in body fat and pulse rate measures. In contrast, the modified Hotelling's $T^{2}$ test found insufficient evidence to suggest any difference between the two groups. The combined test revealed that the differences in body fat and pulse rate measures dominated those of the other measures, although they were not statistically significant. In comparing the control group with the second treatment group, the combined test, the MinP test and the modified Hotelling's $T^{2}$ test found significant evidence to suggest a difference between the two groups. Furthermore, all of the multiple testing procedures--namely the combined test, the MinP test and the closed procedures based on the modified Hotelling's $T^{2}$ test--indicated that the effect on body fat measure was significant. When comparing the first treatment group with the second treatment group, the MinP test and the modified Hotelling's $T^{2}$ test again yielded conflicting evidence in rejecting $H_{K}$; the modified Hotelling's $T^{2}$ test found no significant evidence, whereas the MinP test found significant evidence. However, unlike in the case of comparing the control group with the first treatment group, the combined test confirmed the findings of the MinP test, along with the closed procedure based on the MinP test, in both global and multiple hypothesis tests.
This paper proposes a combined test for global and multiple testing. The combined test exploits the global power advantages of the tests being combined while retaining the benefit of the stepdown procedure of MinP tests. It also offers a tool for the robust and honest evaluation of a global hypothesis when two tests with distinct power advantages are available for hypothesis testing. A simulation study is provided to illustrate the power performance of the combined test for global testing and multiple testing. An empirical application testing the effects of exercise is provided to illustrate the practical relevance of the proposed test.