EconBase
← Back to paper

Sequential Cauchy Combination Test for Multiple Testing Problems with Financial Applications

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

73,956 characters · 15 sections · 73 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Sequential Cauchy Combination Test for Multiple Testing Problems with Financial Applications

\def\spacingset#1{ {#1}} \spacingset{1}

\if11 {

} \fi

\if01 {

center[center omitted — 121 chars of source]

} \fi

spacing{1.00} \begin{abstract} We introduce a simple tool to control for false discoveries and identify individual signals in scenarios involving many tests, dependent test statistics, and potentially sparse signals. The tool applies the Cauchy combination test recursively on a sequence of expanding subsets of $p$-values and is referred to as the sequential Cauchy combination test. While the original Cauchy combination test aims to make a global statement about a set of null hypotheses by summing transformed $p$-values, our sequential version determines which $p$-values trigger the rejection of the global null. The sequential test achieves strong familywise error rate control, exhibits less conservatism compared to existing controlling procedures when dealing with dependent test statistics, and provides a power boost. As illustrations, we revisit two well-known large-scale multiple testing problems in finance for which the test statistics have either serial dependence or cross-sectional dependence, namely monitoring drift bursts in asset prices and searching for assets with a nonzero alpha. In both applications, the sequential Cauchy combination test proves to be a preferable alternative. It overcomes many of the drawbacks inherent to inequality-based controlling procedures, extreme value approaches, resampling and screening methods, and it improves the power in simulations, leading to distinct empirical outcomes. \\ \noindentKeywords: Multiple hypothesis testing; Cauchy combination; High-dimensional; Sequential rejection; Sparse alternatives; Dependence; Drift Burst; Non-zero alpha\\ \noindentJEL classification: C12, C13, C58 \end{abstract}

\singlespacing

Introduction

There are many needle-in-a-haystack problems in empirical finance. There are tests to detect skilled funds barras2010false,giglio2021thousands, nonzero alpha stocks fan2015power, explanatory factors harvey2016and, profitable technical trading rules bajgrowicz2012technical,sullivan1999data, jumps and drift bursts in high-frequency asset prices lee2007jumps, leemykland2012jumps, bajgrowicz2015jumps,christensen2014fact,christensen2018drift, to name but a few examples. These statistical tests have in common that, in order to detect a signal, they require applying the same test repeatedly. This repetition arises either from the presence of a large cross-section of units (e.g., different funds, stocks, factors or trading rules) or because the test is required to be applied continuously over time (every minute or every day). It is commonly known that the simultaneous testing of multiple hypotheses is prone to a “false discovery problem". As more and more tests are performed, an increasing number of them will be significant purely due to chance.\footnote{A well-known example of disregarding the multiple testing problem is so-called “data snooping" or “p-hacking", which is a misuse of statistical testing: one exhaustively searches for signals without compensating for the number of inferences being made giglio2021thousands.} To counteract the inflation of false discoveries, researchers typically treat the hypotheses as a “family" and set a threshold which controls a combined measure of error across all tests, rather than the Type I error of an individual test. The procedures we focus on in this paper address this issue by imposing an upper bound, denoted as $\alpha$, on the probability of making at least one false discovery, thereby controlling the so-called “familywise error rate" shaffer1995multiple,goeman2014multiple.\footnote{Other procedures control, for example, the expected proportion of false discoveries or the so-called “the false discovery rate" benjamini1995,barras2010false,giglio2021thousands.}

The dependence between test statistics plays a crucial role in the effectiveness of multiple testing corrections. In scenarios where the test statistics are independent and adhere to a standard normal distribution under the null hypothesis (as observed in certain jump tests, for instance), there exist well-established statistical solutions.\footnote{Popular approaches to address the false discovery problem in jump detection include choosing a critical value corresponding to an extremely high quantile of the normal distribution andersen2007no or choosing another threshold based on the quantile of the asymptotic distribution of the maximums of the test statistics lee2007jumps,leemykland2012jumps.} However, in many applications in economics and finance, assuming independence among the test statistics is implausible, because many popular test statistics are constructed from overlapping rolling windows or are computed from stock returns that are likely to be driven by common factors. In scenarios with dependent test statistics, popular multiple testing corrections, such as those based the Gumbel distribution lee2007jumps or statistical inequalities such as the Bonferroni correction and its subsequent improvements holm1979simple,hommel1988stagewise,hochberg1988sharper protect against false discoveries, but are known to be overly conservative. The familywise error rate of these methods often turns out to be much smaller than the desired upper bound $\alpha$. Simulation-based methods have also been used to account for the observed correlation of the test statistics when setting a threshold white2000,romano2005exact,romano2005stepwise, but they are not ideal either. Aside from being computationally intensive, they impose a strong parametric assumption on the dependence structure christensen2018drift, which could be misspecified.

In this paper, we introduce a simple tool to control for false discoveries, while being agnostic about the dependence among the test statistics. The solution we propose only uses raw $p$-values and is built upon the Cauchy combination test of liu2020cauchy, which tackles the issue of dependence in the test statistics from another perspective. Their global Cauchy combination (GCC) test is grounded on a convenient theoretical property of Cauchy distributions, which states that linear combinations of these variates behave similarly to a standard Cauchy variate at extreme tails, regardless of the dependence structure. Drawing from this insight, liu2020cauchy propose a transformation of raw $p$-values, such that the transformed $p$-values follow a standard Cauchy distribution under the null hypothesis, and then construct a new test statistic as a linear combination of these transformed $p$-values, with its corresponding critical value derived from a Cauchy distribution. In doing so, they prove that the familywise error rate of the GCC test converges to the nominal size $\alpha$ as the significance level $\alpha$ tends to zero, when all hypotheses are true and test statistics have arbitrary dependency structures. This test is well-suited to deal with the challenges posed by correlation, high-dimensionality, and sparsity, but it is designed for inferences about a global hypothesis. It is not obvious how statements about individual hypotheses are to be made with this procedure.

We extend the pioneering work of liu2020cauchy by introducing a sequential version of the Cauchy combination test to pinpoint the individual hypotheses that trigger the rejection of the global null, enabling the identification of individual signals, such as, skilled funds, nonzero alpha stocks, explanatory factors, or timestamps of jumps and flash crashes. We apply the GCC test recursively on expanding subsets of $p$-values, starting from the largest and progressing to the smallest $p$-value. This process generates a sequence of Cauchy combination test statistics. The $p$-values associated with these test statistics are computed based on a standard Cauchy distribution. Individual violations are detected when the corresponding $p$-value is lower than a predefined threshold $\alpha$. We refer to this new testing procedure as the sequential Cauchy combination (SCC) test, which inherits all the convenient theoretical properties of the GCC test, including being agnostic about the dependence structure. We prove that the SCC test achieves strong familywise error rate control as the significance level $\alpha$ tends to zero, regardless of whether the number of individual hypotheses is fixed or infinite. Moreover, compared to the benchmark procedures, the familywise error rate of the SCC test is closer to the theoretical upper bound, which boosts the power and helps to better identify the individual signals.

To showcase the advantages of the sequential Cauchy combination test, we revisit two multiple testing problems in financial econometrics that exhibit non-trivial correlation structures in the test statistics, high dimensions, and sparse signals, which are common challenges in the field.

itemize• In the first example, we revisit the drift burst hypothesis of christensen2018drift, which aims to identify explosive trends in stock prices. The drift burst test statistic relies on ultrahigh-frequency data and is applied multiple times within a trading day. The test statistics are constructed from overlapping rolling windows and exhibit serial dependence. Drift bursts are rare events, and the strength of this signal varies over time. • In the second example, we test for multiple nonzero alphas within the fama2015five five-factor model framework. If the model fully explains asset returns, the estimated “alphas" should be statistically indistinguishable from zero. The presence of unknown common factors generates strong cross-sectional dependencies among the test statistics giglio2021thousands. Nonzero alphas are typically rare and weak fan2015power.

To assess the robustness of the SCC test against different forms of dependence, we conduct two sets of simulation studies. The first set involves directly generating test statistics with different correlation matrices. The second set involves generating data from a specific underlying process and computing a sequence of test statistics from the data, mimicking the situation in real-world empirical applications. Specifically, we simulate log prices of financial assets using a continuous-time drift bursting process in Example 1, and excess returns from a factor model in Example 2. The main findings from these simulations highlight that the SSC test outperforms other multiple testing corrections, including statistical inequality-based approaches, methods based on extreme value theory, resampling, and screening approaches. Despite its simplicity, the SCC test demonstrates superior properties in terms of minimizing conservativeness and maximizing successful detections.

The rest of the paper is organized as follows. Section (ref) introduces the general notation, definitions, and introduces both the global and sequential Cauchy combination tests. Section (ref) illustrates the finite sample performance of the sequential Cauchy combination test in a simulation experiment with different types of correlations, relative to other multiple testing corrections. Sections (ref) and (ref) revisit the two financial applications. Section (ref) concludes. Appendix (ref) contains the proofs. The Online Supplement provides detailed information on the benchmark procedures, along with additional descriptions and simulations of the drift burst test and the nonzero alpha test.

Multiple Hypothesis Testing with Correlated Test Statistics

In this section, we first introduce the notation and terminology used throughout the article regarding multiple hypothesis testing. We then review the global Cauchy combination test of liu2020cauchy, followed by our sequential version of the Cauchy combination test.

Setting

Let $H_{i}$ denote the $i^{\text{th}}$ null hypothesis of interest, with $i=1,...,d$. Here, $d$ denotes the total number of individual hypotheses, and $\mathcal{H}_{0}$ denotes the collection of null hypotheses of interest. To test the $d$ hypotheses, we can use the associated vector of test statistics $\bm{X}=(X_{1},X_{2},\ldots,X_{d})^{^{\prime }}$, one for each hypothesis being tested, or the corresponding raw $p$-values $p_{1},\ldots ,p_{d}$. The test statistics can be independent or dependent. For many popular tests, such as those described in Section (ref), the test statistics are constructed from rolling windows and exhibit strong serial correlation.

To ensure the validity of individual hypothesis testing, it is common practice to control the probability of falsely rejecting a single hypothesis that is true (known as a false positive or Type I error) at a pre-specified nominal $\alpha$-level. However, when dealing with a large number of hypotheses, the issue of multiplicity arises: if the Type I error of each individual test is controlled at the $\alpha$-level, the probability of having at least one false positive conclusion rises well above $\alpha$.

A classical global test circumvents the issue of multiplicity by replacing multiple tests with a single test. The corresponding global null hypothesis, denoted as $\mathcal{H}_0 = \bigcap_{i=1}^{d} H_{i}$, assumes that all elementary hypotheses are true, and the alternative hypothesis posits that at least one hypothesis is false. In the context of monitoring specific events like jumps or drift bursts within a fixed time period, such as a day, the global null hypothesis would reflect the absence of any such event occurring within that given timeframe. Although global tests serve their purpose by aggregating effects, they may not provide the means to differentiate among individual hypotheses. In the field of financial econometrics, we are often interested in precisely timestamping drift bursts or identifying skilled fund managers, which requires a more granular analysis beyond the scope of global tests.

Let $\mathcal{T}$ denote the set of true hypotheses, $\mathcal{F}$ denote the set of false hypotheses, and $\mathcal{R}$ denote the set of rejected hypotheses. The set of true and false hypotheses are unknown. A statistical test selects hypotheses to reject based on empirical data, and the corresponding set of discoveries in $\mathcal{R}$ should coincide with $\mathcal{F}$ as much as possible, while controlling the probability of making false discoveries. The objective of many multiple testing corrections is to control the familywise error rate (FWER), which constrains the probability of at least one false rejection within a family, denoted as $P[\mathcal{T} \cap \mathcal{R} \neq \varnothing]$. Ideally, the multiple testing correction should ensure that the FWER is not greater than the upper bound $\alpha$, while striving to keep it as close to $\alpha$ as possible. We concentrate on strong control of the FWER, allowing for the presence of some false hypotheses ($\mathcal{F} \neq \varnothing$), rather than weak FWER control, which assumes that all hypotheses of interest are true (i.e., $\mathcal{T} = \mathcal{H}_0$).

Global Cauchy combination test

The global Cauchy combination (GCC) test examines the global null hypothesis. The GCC test statistic is constructed from raw $p$-values of the individual test statistics $X_i$, which are uniformely distributed between $0$ and $1$ under the null hypothesis. The core idea of this test is first to transform these uniformly distributed $p$-values into standard Cauchy variates using the transformation formula $\tan \{(0.5-p_{i})\pi \}$, and then construct a new test statistic by taking the weighted sum of these transformed $p$-values. The new test statistic is denoted by $\tilde{T}$ and is defined as:

equation[equation omitted — 112 chars of source]

in which the $w_{i}$'s are non-negative weights that sum up to one. Throughout the paper, the weights $w_{i}$ are set to $1/d$, for $i=1,\ldots,d$, following liu2020cauchy.

When the raw $p$-values are independent or perfectly dependent, the new test statistic (ref) has a standard Cauchy distribution under the null, because the family of Cauchy densities is closed under convolutions. Even in cases of general dependence (whether weak, mild, or strong), the correlation structure has minimal impact on the tail behavior of the test statistics due to the heavy tails of the Cauchy distribution. Specifically, liu2020cauchy prove that:

equation[equation omitted — 125 chars of source]

in which $C$ is a standard Cauchy random variable, subject to certain regularity conditions on the test statistic.

The result expressed in (ref) suggests that, under the global null hypothesis, the tail of the Cauchy combination test statistic is approximately Cauchy distributed, under arbitrary dependence structures, so that a $p$-value of the Cauchy combination test, denoted $\widetilde{p}$, can be calculated from the standard Cauchy distribution. Suppose that we observe $\tilde{T}=t_{0}$, then:

equation[equation omitted — 92 chars of source]

Using the GCC $p$-values (ref), the tail result in (ref) can be equivalently stated as the actual size converging to the nominal size $\alpha$ as the significance level tends to zero:

equation[equation omitted — 122 chars of source]

The approximation is particularly accurate for small $\alpha$'s, which are of particular interest in large-scale testing problems such as Examples 1 and 2 in sections (ref) and (ref). Importantly, the GCC test achieves weak familywise error rate control regardless of the underlying correlation structure.

Figure (ref) illustrates that while the dependence among individual test statistics may affect the null distribution of the GCC test statistic (ref), its influence on the tail is minimal. To illustrate this point, we simulate a vector of test statistics $\bm{X}$ from a $d$-variate normal distribution with correlation matrix $\bm{\Sigma}$, i.e., $N_d(\bm{0}, \bm{\Sigma})$ with $\bm{\Sigma} = (\sigma_{ij})$ and $d = 300$. The diagonal elements $\sigma_{ii}=1$ for all $i=1,\ldots,d$ and the off-diagonal elements $\sigma_{ij} = \theta^{\abs{i-j}}$ for $i \neq j$, where $\theta$ takes the values of $0.2, 0.4, 0.8, 0.95$. We repeat the simulation $10^7$ times, and calculate two-sided $p$-values, the GCC test (ref) and the GCC $p$-value (ref) for each draw. The histogram of the $10^7$ GCC $p$-values is plotted in Figure (ref). When the autocorrelation is low (i.e., $\theta=0.2$), the distribution of $p$-values resembles a uniform distribution. As the autocorrelation increases, a pothole in the middle and a bump at the end of the histogram appear. However, regardless of magnitude of the autoregressive parameter, the percentage of $p$-values falling into the first bin remains consistently around $5$%, as guaranteed by the limit result described in (ref).

figure[figure omitted — 1,565 chars of source]

Sequential Cauchy Combination Test

The main contribution of this paper is the introduction of the sequential Cauchy combination test, which extends the GCC test of liu2020cauchy to make statements on the elementary hypotheses. To facilitate this, the raw $p$-values are sorted in ascending order, denoting them as $p_{(1)}\leq p_{(2)}\leq \ldots \leq p_{(d)}$, where $H_{(1)},H_{(2)},\ldots,H_{(d)}$ correspond to their respective null hypotheses. For the purpose of testing hypothesis $H_{\left( i\right) }$, we calculate a Cauchy combination test statistic, denoted as ${\normalsize \tilde{T}}_{\left( i\right) }$, using a subset of $p$-values running from $p_{(i)}$ to $p_{(d)}$ as:

equation[equation omitted — 171 chars of source]

where $w_j$ represents the weight assigned to each $p$-value in the subset. The corresponding $p$-value is computed as:

equation[equation omitted — 112 chars of source]

We reject the $i$th null hypothesis $H_{(i)}$ if $\widetilde{p}_{(i)}\leq\alpha$. Similar to the step-up procedure introduced by hommel1988stagewise, the SCC test leverages power across hypotheses: the test statistic $\tilde{T}_{(i)}$ is computed using the raw $p$-values associated with $\mathcal{H}_0^{(i)}=\bigcap_{j=i}^{d} H_{(j)}$.

A more prescriptive description of the SCC testing procedure is as follows:

tcolorbox\begin{minipage}{1\linewidth} SCC algorithm \begin{enumerate} • Calculate raw $p$-values $p_1, p_2,\ldots, p_d$ corresponding to the null hypotheses $H_{1}, H_{2},\ldots, H_{d} $. • Order the raw $p$-values in ascending order, $p_{(1)},p_{(2)},\ldots,p_{(d)}$, with their corresponding ordered null hypotheses $H_{(1)},H_{(2)},\ldots,H_{(d)}$. • Calculate the SCC test statistic $\tilde T_{(i)}$ and the transformed Cauchy $p$-values $\widetilde{p}_{(i)}$ from a subset of the ordered $p$-values $\left\{p_{(j)}\right\} _{j=i}^{d}$ using (ref) and (ref), respectively, for $i=1,\ldots,d$. • Construct the rejection set $\mathcal{R}=\left\{H_{\left(i\right)} : \widetilde{p}_{(i)}\leq \alpha\right\}$. \end{enumerate} \end{minipage}

Figure (ref) illustrates the mechanics of sequential Cauchy combination procedure using a simulated sequence of test statistics. The top row of the figure shows the raw and ordered $p$-values. Most of the observations correspond to the null hypothesis (represented by grey dots), while a few observations correspond to the alternative hypothesis (represented by black dots). The data-generating process is the same as the one used in Figure (ref), where $\theta=0.9$ and $d=100$. We add constant signals (nonzero mean) for five out of the 100 hypotheses, with a signal strength of $\pm2.806$. The sign of the signal aligns with the sign of the test statistic under the null, such that the signal always amplifies the magnitude of the test statistic. The bottom row of the figure plots the sequential Cauchy combination test statistics (ref) and their corresponding $p$-values (ref). In particular, the bottom right panel shows that the SCC $p$-values $\widetilde{p}_{(i)}$ decrease as $i$ decreases from $d$ to $1$. In this example, the SCC test rejects three out of the five alternative hypotheses and does not reject any under the null hypothesis. These rejections correspond to the 4$^\text{th}$, 29$^\text{th}$ and 46$^\text{th}$ hypotheses in the top left panel. It is worth noting that the smallest SCC $p$-value corresponds to the $p$-value of the GCC test in (ref), which performs the test on the largest set of hypotheses.

figure[figure omitted — 1,109 chars of source]

The sequential Cauchy combination testing procedure requires two assumptions. Let $\bm{X}=(X_{1},X_{2},\ldots,X_{d})^{^{\prime }}$ represent the vector of test statistics.

assumption(1) The original test statistics $(X_{i},X_{j})$, for any $1\leq i<j\leq d$, follow a bivariate normal distribution; (2) $E\left( \bm{X} \right) =0$.

The requirement of bivariate normality in Assumption (ref) is a condition weaker than joint normality, enabling the procedure to be applicable for high-dimensional settings. When the dimension $d$ increases at a certain rate with the sample size, the test statistics $\bm{X}$ may not jointly converge to a multivariate normal distribution due to a slower rate of convergence liu2020cauchy, and assuming joint normality becomes unrealistic in such settings.

liu2020cauchy show through simulations that the global Cauchy approximation remains accurate even when the normality assumption is violated, and follows a multivariate Student-$t$ distribution (with four degrees-of-freedom) instead. For a showcase example in finance where the test statistics are Student-$t$ distributed, we refer the reader to Section (ref).

assumptionLet $\mathbf{\Sigma }=corr\left( \bm{X}\right) $. (1) The largest eigenvalue of the correlation matrix $\lambda _{\max}\left( \mathbf{ \Sigma }\right) \leq C_0$ for some constant $C_0>0$; (2) $\max_{1\leq i<j\leq d}\left\{ \sigma _{i,j}^{2}\right\} \leq \sigma _{\max }^{2}<1$ for some constant $0<\sigma _{\max }^{2}<1$, where $\sigma _{i,j}$ is the $\left( i,j\right) $ element of $\mathbf{\Sigma }$.

Assumption (ref) on the correlation matrix becomes relevant when the number of hypotheses $d$ diverges to infinity. It imposes two conditions: boundedness of the largest eigenvalue of the correlation matrix and the absence of perfectly correlated test statistics. The conditions are frequently encountered in high-dimensional settings and general enough to encompass a wide range of tests.\footnote{However, this assumption excludes very strong dependence and latent factors shared by the test statistics in high-dimensional settings. Investigating the relaxation of this assumption is left for future research.}

theoremUnder Assumption (ref) for a fixed $d$ and Assumptions (ref) and (ref) for $d=o(h^{\eta})$ with $0<\eta<1/2$, as $\alpha\rightarrow 0$, the probability of the SCC testing procedure making at least one false rejection converges to $\alpha$, i.e., \begin{equation} \lim_{\alpha\rightarrow 0}\Pr\left\{\mathcal{R} \cap \mathcal{T}=\emptyset\right\} \rightarrow \alpha. \end{equation}

The proof of Theorem (ref) is provided in Appendix (ref). The theoretical result in (ref) for the SCC testing procedure stands in stark contrast to statistical inequality-based controlling procedures covered in the Online Supplement, which have the property: $$\Pr \left\{ \mathcal{R}\cap \mathcal{T\neq \emptyset }\right\} \leq \alpha.$$ See goeman2010sequential for a discussion of their theoretical properties. These inequality-based controlling procedures ensure that the likelihood of making at least one false discovery is bounded above by the pre-specified significance level, $\alpha$. Consequently, the SCC procedure exhibits less conservatism compared to inequality-based controlling procedures.

Simulations

In this section, we compare the performance of the SCC test against several popular multiple testing corrections in a simulation study, considering different forms of dependence. The benchmark procedures include four inequality-based approaches: the Bonferroni correction and its subsequent improvements proposed by holm1979simple, hommel1988stagewise and hochberg1988sharper, as well as the Gumbel approach. Detailed discussions of these benchmark procedures can be found in the Online Supplement.

Under the null hypothesis

We assess the statistical performance of the different multiple testing corrections under the null hypothesis. To measure the empirical familywise error rate, we conduct $S=10^4$ replications for each method $m$, and calculate $\widehat{FWER}_m$ as follows: \[ \widehat{FWER}_m=\frac{1}{S}\sum_{s=1}^{S} \bm{1}\left( \min_{i\in\left\{1,2,\cdots,d\right\}} \left\{p_{(m,i)}^{s}\right\}\leq\alpha\right), \] where $\bm{1}(.)$ is the indicator function, and $p_{(m,i)}^{(s)}$ represents the $p$-value of the $i$th hypothesis for method $m$ in the $s$th replication. When the test statistics exhibit strong dependence, we expect the $\widehat{FWER}$ of the SCC test to be closer to the nominal level $\alpha$ compared to the other procedures.

Under the null hypothesis, the test statistics $\bm{X}$ are generated from a $d$-variate normal distribution with zero mean and covariance matrix $\bm{\Sigma}$, i.e., $N_d(\bm{0}, \bm{\Sigma})$. We set the dimension $d$ to 100. The diagonal elements of the covariance matrix, $\sigma_{ii}$, are all equal to $1$, for $i=1,\ldots,d$. The off-diagonal element, $\sigma_{ij}$ with $i\neq j$, adhere to three specific specifications.

itemize• Model 1. Exponential decay: $\sigma_{ij} = \theta^{\abs{i-j}}$ with $\theta= 0.2, 0.4, 0.6,$ $0.8, 0.90, 0.95$. • Model 2. Polynomial decay: $\sigma_{ij} = \frac{1}{0.7 + \abs{i - j}^\theta}$ with $\theta = 1.0, 1.5, 2.0, 2.5$. • Model 3. Block-diagonal: $\bm{\Sigma} = \text{diag}\{A_1,\ldots, A_{d/10}\}$, for which each diagonal block $A_k$ is a $10 \times 10$ equi-correlation matrix with its off-diagonals $\sigma_{ij} = \theta$ and $\theta = 0.1, 0.3, 0.5, 0.7, 0.9$.

Models 1 and 2 also appear in liu2020cauchy. The exponentially decaying correlation structures in Model 1 are frequently observed in time series and financial econometrics. For instance, in Section (ref), the sequence of drift burst test statistics constructed from overlapping rolling windows, exhibits an autoregressive process and has an exponential decaying covariance structure, as shown by christensen2018drift. The block-diagonal structure in Model 3 is commonly used when testing high-dimensional factor pricing models fan2015power and emulates a cross-sectional dependence structure.

Table (ref) shows the superior performance of the SCC test in the presence of correlated test statistics under the null hypothesis. The empirical familywise error rate of the SCC test, reported in the last column, remains close to the nominal level $\alpha = 5\%$ across various correlation structures, demonstrating its robustness. These findings align with the theoretical discussions presented in Section (ref). In contrast, the inequality-based procedures and the Gumbel method exhibit greater conservatism, as evidenced by their substantially lower FWERs (although slight variations exist depending on the correlation pattern). The SCC test stands out as the only controlling procedure with an empirical FWER close to the nominal level (5%) across all three types of correlation structures in the test statistics.

table[table omitted — 2,467 chars of source]

Under the alternative hypothesis

Under the alternative hypothesis, the performance of the controlling procedures is assessed based on their global power (the percentage of replications that reject at least one hypothesis) and their successful detection rate (the percentage of overlapping hypotheses between the sets of false hypotheses and discoveries). Given the improved accuracy of the SCC procedure in controlling the FWER under the null hypothesis (as demonstrated in Table (ref)), we anticipate that the SCC test will exhibit higher power when applied under the alternative hypothesis.

The test statistic vector $\bm{X}$ is generated from a $d$-variate normal distribution with mean vector $\bm{\mu} = (\mu_i)$ and a correlation matrix $\bm{\Sigma} = (\sigma_{ij})$, i.e., $N_d (\bm{\mu}, \bm{\Sigma})$. We adopt the same correlation matrices $\bm{\Sigma}$ as discussed in Section (ref). The percentage of signals (i.e., non-zero $\mu_i$'s in the vector $\bm{\mu}$) is set to be 5% (specifically, out of the 100 hypotheses, 5 are under the alternative). All signals have the same strength $\abs{\mu_i} = \mu_0$ which is chosen to be relatively weak, i.e., $\pm2.1737$.\footnote{ The chosen signal strength ensures that the test power converges to unity as $d \rightarrow \infty$ in the case of sparse signals, following the result presented in Theorem 3 of liu2020cauchy. } The sign of the signal aligns with the sign of the test statistic under the null so that the signal always amplifies the magnitude of the test statistic.

table[table omitted — 2,832 chars of source]

Table (ref) shows the superior global power of the SCC test in the presence of correlated test statistics and sparse signals. The SCC test exhibits an approximate $10$% power enhancement compared to the runner-up method. When the significance level $\alpha$ is set to $5\%$, the power of the SCC test ranges between $69$% and $77$%. Although each statistical inequality-based method improves upon its predecessor in certain aspects, we do not observe a significant difference in the frequency of rejections among the four approaches. As anticipated, the Gumbel approach, which assumes independent test statistics, is the most conservative test.

Table (ref) tells a similar story with respect to the average numbers of successful detections. The SCC testing procedure successfully detects approximately 1.2 out of 5 hypotheses (or $24\%$) under the alternative, even with sparse and weak signals.\footnote{The average number of false detections is slightly higher (not reported) for the SCC test, amounting to falsely rejecting on average 0.1 (out of 95) true hypotheses.} Meanwhile, the average number of successful detection for the inequality-based procedures are around 0.95 (or $19\%$), whereas the value for the Gumbel procedure is about 0.81 (or $16.2\%$).

table[table omitted — 2,528 chars of source]

Example 1: Monitoring Drift Burst

The drift burst hypothesis, as proposed by christensen2018drift, postulates the existence of locally explosive trends in high-frequency asset prices, resembling phenomena like flash crashes or gradual jumps. An infamous example of a flash crash occurred on May 6, 2010, when the Dow Jones Industrial Average rapidly dropped nearly 1,000 points within minutes, only to recover most of the losses shortly thereafter.

The drift burst test serves to detect and timestamp these explosive trends. In our analysis, we compute the test statistic on a minute-by-minute basis. Since there are 6.5 trading hours per day, the test needs to be conducted 341 times per day, resulting in a multiple testing problem. The test statistics are computed using overlapping windows and are expected to exhibit high autocorrelation. Given that drift bursts are rare events, they can be considered as sparse signals, and the strength of these signals can vary. There are several approaches to deal with the false discovery problem in this context, including the benchmark procedures considered in Section (ref), the SCC, test and a resampling procedure suggested by christensen2018drift.

Section (ref) presents the drift burst hypothesis and the drift burst test as a means to detect these phenomena. In Section S2.3 of the Online Supplement, we conduct a simulation study. We apply the controlling procedures to detect drift bursts in real-world data, specifically the Nasdaq composite index and S&P 500 index ETFs, in Section (ref).

Drift Burst Hypothesis and Test

Under the null hypothesis of no drift burst, the frictionless log prices $P = (P_t)_{t \geq 0}$ follow an It\^o semi-martingale process defined on a filtered probability space $(\Omega, \mathcal{F}, (\mathcal{F})_{t \geq 0}, \ensuremath{\mathbb{P}})$:

eqnarray[eqnarray omitted — 68 chars of source]

where $\mu_t$ is the instantaneous drift, and the diffusive component consists of the spot volatility $\sigma_t$ and a standard Brownian motion $W_t$. The coefficients $\mu_t$ and $\sigma_t$ are locally bounded or “non-explosive". The volatility process is assumed to follow a heston1993closed-type dynamics:

equation*[equation* omitted — 106 chars of source]

where $B_t$ is a standard Brownian motion and $E\left( dW_{t},dB_{t}\right) =r dt$.

Under the alternative hypothesis, a drift-bursting term $\mu_t^\text{b}$ and a volatility-bursting component $\sigma_t^\text{b}$ are added to the standard It\^o semi-martingale process (ref), resulting in the following dynamics:

eqnarray[eqnarray omitted — 120 chars of source]

where $\abs{\mu_t^\text{b}} / \sigma_t^\text{b} \rightarrow \infty$ as $t \rightarrow \tau_{\text{b}}$, with $\tau_{\text{b}}$ denoting the drift burst time. An example of such an explosive process is given by:

eqnarray[eqnarray omitted — 330 chars of source]

where $2\Delta t$ represents the duration of the burst, $\alpha^\text{b}$ corresponds to the strength of the burst, $\beta^\text{b}$ is the strength of the volatility burst and $a$ and $b$ are constants. This data-generating process can capture various realistic patterns, including flash crashes and mildly explosive trends, and is used in the simulations in Section S2.3.

Observations are recorded at equidistant intervals $0 = t_0 < t_1 < \ldots <t_n = T$, where $T$ represents a fixed time period, such as one trading day consisting of 6.5 trading hours. The distance between two consecutive observations is $\Delta_n=t_{i+1}-t_{i}$. The observed log price, contaminated by noise, is denoted as $\widetilde{P}_{t_i} = P_{t_i} + \epsilon_{t_i}$, where the $\epsilon$ is an error term (noise) and independent from the latent log price $P$.

The noise-robust drift burst statistic christensen2018drift is defined as:

eqnarray[eqnarray omitted — 121 chars of source]

where $h$ is the bandwidth of the mean estimator, $\hat{\bar{\mu}}_{t_i}$ is a noise-robust estimator for the local drift, and $\hat{\bar{\sigma}}_{t_i}^{2}$ is a noise-robust estimator for the local variance. The test is applied on a coarse sampling grid, and the null hypothesis of the drift burst test asserts that `there is no drift bursting at time period $t_i$'. The drift burst test statistics are computed minute-by-minute using overlapping rolling windows. As a result, the dimension of the test statistic sequence $\bm{X}=\left\{X_i\right\}$ is large (with the parameters settings we end up with $d=341$ for the period of one day) and the tests are autocorrelated. For more information regarding the implementation of noise-robust estimators of the local drift and local variance, we refer the reader to the Online Supplement.

Under the null hypothesis, as the sampling frequency approaches infinity ($\Delta_n\rightarrow 0$), the test statistic (ref) converges to the standard normal distribution, i.e., $X_i \rightarrow^{d} N(0,1)$, This implies that the test statistic satisfies the necessary assumptions required by the Cauchy combination test when the sampling frequency is sufficiently high under the null. Under the alternative hypothesis, the test statistic diverges when the drift term explodes fast enough relative to the volatility, i.e., $\abs{X_i} \rightarrow \infty$, at the drift burst time.

To control the familywise error rate, christensen2018drift propose a resampling-based approach for generating critical values for the drift burst test. The resampling approach approximates the dependence structure of the test statistic sequence under the null hypothesis with an autoregressive order one (i.e., AR(1)) model and to obtain its distribution via simulation.

There are a few limitations associated with using simulated critical values in practice. Since each sequence of test statistics (corresponding to each day in an empirical analysis or each sample path in a Monte Carlo study) requires a unique critical value, the resampling procedure is computationally intensive.\footnote{To expedite the process, a table can be prepared in advance containing the quantile functions of the normalized maxima for various values of the autoregressive coefficient $\theta$ and dimensions $d$. However, an interpolation routine becomes necessary when the estimated first-order autocorrelation and dimension are not included in the table.} It also imposes a strong parametric assumption on the dependence structure, which could be misspecified. Additionally, resampling attempts to reproduce the dependence structure under the null hypothesis using possibly contaminated data christensen2018drift, which in turn can affect the estimation of the critical value. Therefore, caution should be exercised when interpreting the results, and the potential biases associated with the parameter estimation must be considered. See the Online Supplement for more detailed discussion of the resampling approach.

In the Online Supplement, we also evaluate the performance of the aforementioned controlling procedures in a simulation setting. Specifically, we evaluate the performance of the four inequality-based procedures, the Gumbel method, the resampling approach, and the SCC testing procedure. Instead of directly simulating the test statistics as done in Section (ref), we generate log prices from (ref) under the null hypothesis and (ref) under the alternative hypothesis, and then compute drift burst test statistics (ref) from the simulated prices. Two types of drifts bursts are considered: a V-shaped 20-minute flash crash and a three-day persistent expansion.

Our findings suggest that the SCC procedure is the most preferred option for monitoring drift bursts. The inequality-based methods and the Gumbel method tend to be conservative when applied to the autocorrelated drift burst test statistics. While the resampling procedure shows slightly better performance in terms of power and successful detection rate for flash crashes, the SCC procedure significantly outperforms the resampling method in identifying persistent expansions. Although there might be a small loss in power observed in some specific cases, it is a reasonable trade-off for having a robust approach to monitor drift bursts. Moreover, the SCC procedure procedure offers the advantage of being easier to implement and can accommodate arbitrary dependency structures without requiring simulations, making it a more practical choice overall.

Empirics

We apply the same controlling procedures to the drift burst test using data from the Nasdaq ETF (ticker: IXIC) and the S&P 500 ETF (ticker: SPY) covering the period from 1996 to 2020. The data was obtained from the Refinitiv Tick History Database at a one-second frequency, and we follow the data cleaning rules outlined in barndorff2009realized. We test for drift bursts on a minute-by-minute basis and control the familywise error rate at a level of $0.1$%.

The weekly prices of the Nasdaq and S&P 500 ETFs are illustrated in Figure (ref), with grey bars indicating weeks that exhibit drift bursts. Drift bursts are determined using the test proposed by christensen2018drift and the SCC controlling procedure. It is evident that the Nasdaq index experienced a higher number of drift bursts, particularly in the early 2000s following the collapse of the dot-com bubble. The S&P 500 index barely has any rejections.

figure[figure omitted — 802 chars of source]

Table (ref) reports the rejection frequencies of the drift burst test using the aforementioned controlling procedures. Compared to the benchmark procedures, the SCC test demonstrates a higher sensitivity in detecting drift burst days (i.e., at least one rejection within the day) and time intervals (i.e., the total number of rejections) in the Nasdaq index. In the case of the S&P 500 index ETF, the SCC testing procedure detects more drift burst days and intervals than the inequality-based and Gumbel methods, albeit slightly fewer than the resampling procedure. It is noteworthy that the bursting episodes in the S&P 500 are relatively short-lived, aligning with the Flash Crash data-generating process considered in our simulations. On the other hand, the Nasdaq index exhibits persistent drifts laurent2022unit, akin to the persistent expansion data-generating process.

table[table omitted — 1,246 chars of source]

Example 2: In Search of Nonzero Alpha Assets

The Capital Asset Pricing Model (CAPM) is a prominent risk model. However, in light of empirical evidence revealing systematic patterns in stock returns, often referred to as “anomalies", many additional risk factors have been introduced to explain average returns. These risk factors are said to represent some dimension of undiversifiable systematic risk that should be compensated with higher returns. If the factor model fully characterizes expected returns, the regression intercept (also known as the “alpha") should theoretically equal zero.

We search for nonzero alpha assets using the fama2015five five-factor model framework. The conventional approach is to run time series regressions on each individual asset, and subsequently perform individual tests on the estimated alphas. The number of assets that need to be tested simultaneously is large (e.g., all of the S&P 500 stocks) and the test statistics of the cross-sectional alphas are most likely correlated due to the presence of unknown common factors giglio2021thousands. There is a general consensus in the empirical finance literature that mispriced assets are rare fan2015power,giglio2021thousands. To tackle the challenge of multiple testing, several methods can be used. These include the benchmark procedures, the SCC test, and a screening procedure proposed by fan2015power. All of these procedures control the FWER and have the ability to identify individual violations.\footnote{An alternative objective is to control the false discovery rate, which is defined as the proportion of false discoveries (see e.g., barras2010false, barras2010false and giglio2021thousands, giglio2021thousands for two examples).}

Section (ref) provides an overview of the nonzero alpha hypothesis and introduces the corresponding test. Section S3.2 of the Online Supplement presents simulation results that compare the performance of the controlling procedures in the context of the nonzero alpha test. In Section (ref), we examine Fama-French portfolios formed on bivariate sorts and search for portfolios with a nonzero alpha.

Nonzero Alpha Hypothesis and Test

The multi-factor pricing model, motivated by the Arbitrage Pricing Theory ross1976arbitrage, postulates how financial returns are related to market risks. This model has enjoyed widespread application in asset pricing and portfolio management. Let $y_{it}$ be the excess return (i.e., real rate of return minus the risk-free rate) of the $i$th financial asset at day $t$ and consider the following linear regression model:

eqnarray[eqnarray omitted — 149 chars of source]

where $a_i$ is an intercept, often referred to as “alpha", $\bm{b}_i = (b_{i1}, \ldots, b_{iK})^{\prime }$ is a vector of factor sensitivities or loadings, also known as “betas", $\bm{f}_t = (f_{1t}, \ldots, f_{Kt})^{\prime }$ are observable factors, and $u_{it}$ is an idiosyncratic error which is uncorrelated with the factors. A well-known example of (ref) is the three-factor model of fama1992cross, which captures a substantial portion of the variation in the cross-section of average returns and absorbs a lot of the anomalies that have plagued the CAPM fama1996multifactor.

Our objective is to identify individual assets with a nonzero alpha. The null hypothesis of each asset is therefore $H_i: a_i=0$ (`there is no mispricing of asset $i$') and the alternative hypothesis is $a_i\neq 0$ (`asset $i$ is mispriced'), for $i=1,\ldots, d$. The most common way to test this null hypothesis is to use a simple $t$-statistic for $a_i$, i.e.,

eqnarray[eqnarray omitted — 89 chars of source]

where $\hat{a}_i$ is the estimated alpha and $\hat{ \sigma}_{\hat{a}_i}$ is the estimated standard error, for each asset $i = 1, ..., d$. Under the null hypothesis, the test statistic (ref) follows a Student-$t$ distribution, i.e., $X_i \sim t(\nu)$, with $\nu$ being the degrees of freedom. A viable detection strategy involves computing the test statistic (ref) for each asset in the cross-section and rejecting the null hypothesis when $\abs{X_i}$ exceeds a pre-specified quantile of the $t(\nu)$ distribution.

The multiplicity issue arises when dealing with a large number of assets. One important benchmark for controlling false discoveries in testing factor pricing models is the power enhancement test proposed by fan2015power. This global test employs a screening technique that incidentally identifies individual violations. We refer to this approach as the screening method and provide a comprehensive description of the procedure in the Online Supplement. We can also use other approaches such as the inequality-based, Gumbel method and SCC test. It is worth noting that the cross-sectional test statistics are likely to be cross-correlated giglio2021thousands, and hence the Gumbel method and inequality-based procedures are again expected to be conservative. While it is true that a Student-$t$ distribution does not exactly fulfill the assumptions of the Cauchy combination test, liu2020cauchy show through simulations that the Cauchy approximation remains accurate under such a departure from normality.

Once again, we assess the finite sample performance of the SCC testing procedure by comparing it with existing controlling procedures, which include the four inequality-based procedures, the Gumbel method and the screening approach, for the identification of nonzero alpha assets in both simulation settings and empirical applications. We exclude the resampling approach considered for the drift burst test, as it is not suitable for the current cross-sectional context. Complete details of the simulation designs and results can be found in the Online Supplement. Unlike the simulations in Section (ref), we simulate excess returns from the fama2015five five-factor model, with its parameters calibrated to the empirical data. We then compute the test statistics from the simulated data. Our findings show that the SCC test outperforms all other procedures in terms of controlling the FWER, global power, and successful detection rate. The Gumbel method and screening method tend to be most conservative in their outcomes.

Empirics: Kenneth French Portfolios

In the empirical analysis, we study the portfolios available in Kenneth French's Data Library.\footnote{\url{https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html}} Specifically, we focus on the $d = 100$ portfolios formed bivariately ($10\times10$) based on size and book-to-market (Size-BM), size and investment (Size-INV), and size and operating profitability (Size-OP). The left-hand side of equation (ref) comprises value-weighted portfolio excess returns and the right-hand side are the five factors in the fama2015five five-factor model (i.e., market, value, size, profitability and investment). The sample period spans from July 1963 to November 2022 and consists of $713$ monthly observations. To conduct the nonzero alpha test, we use a rolling window approach with a window size of $T = 240$ observations. We treat the $100$ portfolios as one family and control the familywise error rate at the $5\%$ level.

Table (ref) reports the average rejection frequencies of zero alpha null (across the rolling analysis) using various controlling procedures. As anticipated, there are few violations. For instance, in the case of the Size-BM portfolios, the SCC test shows an average rejection frequency of 2.60%. Overall, the SCC test consistently yields higher rejection frequencies compared to the inequality-based methods, the screening procedure, and the Gumbel method, in that order. These empirical results align with prior research on mispricing, confirming the rarity of nonzero alpha assets fama1996multifactor,fan2015power,giglio2021thousands. Nonetheless, the SCC testing procedure can detect more of these rare violations.

table[table omitted — 981 chars of source]

Evidently, the rejection numbers vary over time, as can be seen in Figure (ref). In the first decade of the sample period, there is almost zero rejection according to all procedures. The number of identified nonzero portfolios starts to increase in the late 1990s, reaching its peak during the dot-com bubble crash in the early 2000s, and declined afterwards. The SCC test identifies more violations than other procedures for about $40\%$ of the sample period. During the remaining periods, the rejection numbers of the SCC test are on par with the other procedures. While the results from the benchmark procedures are similar, the gap between the SCC test and other procedures can be very substantial. For instance, in March 2001, the SCC test detected 13 portfolios with nonzero alphas, while the benchmark procedures identify only 1 or 2 such portfolios. The findings above align with our theoretical expectations and are consistent with our previous simulations results, reinforing the notion that using SCC can yield significant improvements in testing outcomes.

figure[figure omitted — 391 chars of source]

Conclusions

We introduce a simple procedure to control for false discoveries and identify individual signals in scenarios involving many tests, dependent test statistics, and potentially sparse signals. The tool is agnostic to the underlying dependence structure and scalable to deal with high dimensions. Our approach is a sequential version of the global Cauchy combination test proposed by liu2020cauchy. By applying the global test recursively on a sequence of expanding subsets of ordered $p$-values, our sequential Cauchy combination test enables the identification of individual violations.

We show that the sequential Cauchy combination test achieves strong familywise error rate control and is less conservative compared to popular statistical inequality-based methods (such as the Bonferroni correction and subsequent improvements of holm1979simple, holm1979simple, hommel1988stagewise, hommel1988stagewise and hochberg1988sharper, hochberg1988sharper) and the Gumbel method.

The Cauchy transformation has proven its value in a genome-wide association study of Crohn's disease liu2020cauchy, but its applicability extends beyond genomics. We revisit two important needle-in-a-haystack problems in financial econometrics, where the test statistics have either serial or cross-sectional dependence: monitoring drift bursts and searching for nonzero alpha assets. The drift burst test of christensen2018drift detects the presence of explosive trends in asset prices using high-frequency intraday data. The test statistics are computed from overlapping windows, resulting in high autocorrelation. We also revisit the fama2015five multi-factor model to identify nonzero alpha financial assets. Detecting these rare nonzero alphas among a large group of financial assets is challenging, especially when the test statistics are likely to be cross-sectionally correlated. Without a proper controlling procedure, one might flag false discoveries or miss important signals. Our results indicate that the sequential Cauchy combination test is the a preferable method for both applications.

We emphasize that our sequential Cauchy combination test is not limited to financial econometrics. We anticipate its applicability to a wide range of hypothesis tests in fields such as economics, finance, medicine, marketing and climate studies, as it can handle various types of dependence effectively.