Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
258,560 characters · 17 sections · 55 citation commands
Inference on effect size after multiple hypothesis testing
\begingroup
\@thanks \endgroup \global\let\@thanks\@empty
Many empirical studies across fields --- from public health and education to psychology and economics --- estimate multiple effects of interventions or treatments, whether by examining different outcome measures, types of interventions, or population subgroups list2019multiple. Estimating multiple treatment effects allows researchers to validate their methodological approach, understand how effects vary across contexts, and identify the underlying mechanisms at work. These estimated effects are typically categorised as either statistically significant or insignificant using multiple hypothesis testing procedures. The significant effects are often interpreted as evidence for effective interventions, making them the primary focus of subsequent analysis. Researchers are particularly interested in their practical importance, as determined by the magnitude or size of these effects (for example, economists are concerned about an effect's “economic significance” and medical researchers about its “clinical significance”).\footnote{See, for example, the special issue in the Journal of Socio-Economics, Volume 44, Issue 5 for the discussions regarding the difference between statistical and economic significance. In particular, see ziliak2004size and the comment articles on it.} However, when attention is restricted to the significant effects, their estimated magnitudes are subject to selection bias, which invalidates conventional statistical inference methods.
This paper proposes new methods that provide valid statistical inference for significant effects by explicitly accounting for their selection based on statistical significance after multiple hypothesis testing. To determine significance, we consider a broad class of step-down and step-up multiple testing procedures, including classical methods such as those by holm1979simple, benjamini1995controlling,benjamini2001control and more powerful modern bootstrap-based methods such as romano2005stepwise,list2019multiple. The proposed methods can be applied if the underlying treatment effect estimators are asymptotically jointly Gaussian with a consistently estimable covariance matrix. Given these mild assumptions, our methods apply broadly, including to standard causal inference designs with observational or experimental data.
We use the term “selective reporting of significant results” to refer to the common practice of emphasising statistically significant estimates. In particular, summaries of research findings in the abstract or introduction of a paper typically highlight the significant results. By contrast, insignificant results are relegated to less prominent parts of the paper, such as passing mentions in the empirical section or in the appendix.\footnote{ For example, the introduction in karlan2007does --- on which one of our empirical illustrations builds --- summarises their subgroup analysis by reporting that revenue increased by 55% in “red” states, while there was “little effect” in “blue” states. Their introduction thus highlights the subgroup with a significant effect.} Further selection arises when the interpretation of a significant effect changes depending on the significance of other effects. \footnote{For instance, we may not interpret an effect as causal if we find the effect on a placebo outcome to be significant.} One might argue that the solution is to avoid selective summaries and interpretations altogether. Yet this does not resolve the problem; it merely shifts the burden of synthesising the results onto the reader, who then faces a selective inference problem of their own.
Our new point estimators and confidence intervals are grounded in the principle of conditional selective inference fithian_optimal_2017, lee2016exact. Specifically, we account for the fact that the same data are used for determining significance and estimating effect sizes by conditioning on a “selection event” that describes the “observed significance.” Thereby, we control statistical error rates “on average” across all possible versions of a study that analyse the same effects based on observed significance.
Our conditioning approach allows researchers to apply standard multiple testing protocols based on the full sample to determine significance. In contrast, recently proposed selective inference methods introduce randomisation into the selection stage to facilitate post-selection inference fithian_optimal_2017, tian_selective_2018, leiner2025data. In the context of selection by significance, this reduces the power of the significance test, akin to using a smaller sample.\footnote{In data carving fithian_optimal_2017, selection is based on a split sample. Data fission leiner2025data and randomised response selection tian_selective_2018 add noise at the selection stage.} Our approach is advantageous when establishing significance is the primary research goal and reporting effect sizes is of secondary concern. It also admits retrospective corrections of selection bias in existing studies that have already exhausted the full data for significance testing. Our methods can be applied “out of the box” to empirical analyses that already use multiple hypothesis testing. Randomisation-based methods require more careful planning, possibly requiring a pre-analysis plan that specifies the randomisation seed to safeguard against the randomness of selection results.
A key feature of our approach is that it scales well with the number of treatment effects, greatly expanding the range of possible applications. In one of our data examples, we estimate fund-specific “alphas” for outperforming funds identified through multiple hypothesis testing among more than three hundred mutual funds. Other methods of post-selection inference are computationally infeasible at this scale; for example, berk_valid_2013 state that their method can scale only to a dimension of “20, or slightly larger” (p. 16). Even alternative conditional inference methods, such as data carving fithian_optimal_2017 and selection using randomised responses tian_selective_2018, face numerical difficulties when the number of hypotheses at the selection stage is large, since they carry out a complex (numerical) integration over randomised selection events.
We achieve scalability through two key ingredients: additional conditioning on a nuisance parameter and an algorithm that, for $m$ treatment effects, computes the conditional support in $\mathcal{O}(m^3 \log m)$ time. Developing this algorithm is a central technical contribution of the paper. Unlike in other selective inference settings lee2016exact, tibshirani_exact_2016, our problem admits no simple closed-form expression for the conditional support. The underlying difficulty is the complex geometry of the selection event. Step-up and step-down testing proceed sequentially, and the significance thresholds change depending on which effects have been found significant so far.
We prove the uniform asymptotic validity of our procedures under the assumption that the estimated treatment effects uniformly converge to a joint Gaussian distribution as the sample size approaches infinity. The uniformity of our result is important because it allows for local asymptotics where some treatment effects are on the margin of significance even in large samples. This ensures that our asymptotic result provides a relevant approximation of finite samples where we see relatively small $t$-statistics.
In Monte Carlo simulations, we demonstrate the validity of our procedures in finite samples. Even in non-Gaussian settings, our methods capture selection bias and yield confidence intervals with adequate (conditional) coverage. By contrast, conventional confidence intervals that ignore the selection problem exhibit significant undercoverage. Bonferroni-corrected simultaneous confidence intervals undercover when significance is “marginal” and are overly conservative when effects are highly significant.
To demonstrate the behavior of our methods when applied to real data, we provide two empirical applications. We find that the direction and magnitude of the bias correction is not easily anticipated, which highlights the relevance of a formal statistical approach to address the selection problem. For marginally significant effects, our post-selection inference tends to correct the naive effect size estimates downward, potentially addressing concerns about inflated effect sizes in low-powered studies ioannidis2017power,kvarven2020comparing.
Our first application revisits the data from karlan2007does on the effects of a matching grant on charitable donations. We build on the earlier analysis by list2019multiple, who applied multiple testing procedures to identify significant treatments. Our methods provide inference on the effect sizes of these significant effects.
In our second application, we estimate mutual fund “alphas,” which measure performance beyond systematic risk factors. Using multiple testing, we identify funds with significantly positive alphas---commonly viewed as market outperformers---from a sample of more than 370 funds, highlighting that our method scales to settings with many estimated effects. We show that, even after accounting for selection based on significance, some outperformers still beat the market by an economically meaningful margin.
\paragraph*{Relation to the literature}
This paper is part of the selective inference literature. This rapidly growing literature is reviewed in zhang2022post. A selection of relevant studies includes berk_valid_2013, lockhart_significance_2014, fithian_optimal_2017, tibshirani_exact_2016, lee2016exact, bachoc_valid_2019, watanabe2021, jewell_testing_2022, terada2023, duy2023,sarfati2025post.
In this paper, we adopt the conditional inference approach, which controls statistical error under a conditional distribution. Alternative approaches target the unconditional distribution benjamini2005false,andrews_inference_2024,mccloskey2024hybrid, berk_valid_2013, zrnic2024locally.
The conditional inference approach is a natural fit for our problem since it allows us to control statistical error conditional on the observed significance. The observed significance provides the context in which the estimated effects are interpreted and determines the research questions that are emphasised when summarising the results. Inference based on the unconditional distribution averages statistical error over different contexts that focus on different research questions and hence attract different readers. fithian_optimal_2017 observe that “averaging our error rates across [\dots] two questions with two different interpretations, seems inappropriate” (p. 31).
Moreover, the conditional inference approach produces point estimates that capture selection bias. The unconditional approach proposed by benjamini2005false addresses only uncertainty, not bias. Therefore, their approach offers methods “for confidence intervals but none for point estimation” benjamini2010simultaneous.
Targeting the unconditional distribution can increase power goeman2024selection, zrnic2024locally, but dismisses the important role of the selection event in the interpretation of the estimated effects. Additionally, we empirically show that conditional inference can yield tighter confidence intervals than simultaneous inference, demonstrating that the potential power gains from unconditional inference are not uniform.
We use the “polyhedral” conditional inference method of lee2016exact as the basis for our procedures. Recently introduced randomisation-based methods fithian_optimal_2017, tian_selective_2018, leiner2025data provide alternative ways to carry out conditional selective inference and address some limitations of the polyhedral approach kivaranovic2021, kivaranovic2024. In addition, the data fusion approach of leiner2025data can even be applied in some non-Gaussian settings such as testing Poisson data. Randomisation methods rely on specific modifications at the selection stage, ruling out using out-of-the-box multiple testing procedures.
\paragraph*{Organisation of the paper} Section (ref) introduces our new procedures in a simplified setting. Section (ref) provides some intuition based on a univariate example. Section (ref) defines the multiple testing procedures covered by our methods. Section (ref) describes the computational algorithm that makes our methods scalable. Section (ref) discusses some properties of our procedures and extends our procedures to more general settings. Section (ref) provides an asymptotic justification of our procedures. Simulation results are included in Section (ref). Sections (ref) and (ref) provide our empirical applications. Section (ref) concludes. All proofs are in the online appendix.
Our goal is to make inferences on the parameters that we find significant among multiple, possibly many, estimated parameters. This section defines our formal setting and introduces our proposed procedures.
We begin by outlining the context of our study. To introduce the method, we assume for now that the parameter estimator is normally distributed with a known variance-covariance matrix. We identify significant estimates based on $t$-statistics.
Let $\theta = (\theta_1, \dotsc, \theta_m)'$ denote a vector of parameters. For concreteness, we consider the parameters in $\theta$ to be treatment effects corresponding to $m$ specifications, which may involve different treatments, outcome variables, and/or subpopulations.
We assume that there exists an estimator $\hat{\theta}$ for $\theta$ that is normally distributed with mean $\theta$ and known variance-covariance matrix $V$. We subsequently relax these assumptions and allow unknown $V$ (Section (ref)) and settings where Gaussianity holds only asymptotically (Section (ref)).
Significance is determined based on the vector of $t$-statistics for the estimated effects, given by $X = \operatorname{diag}^{-1/2} (\mathbf{v}) \hat{\theta}$, where $\mathbf{v}$ is the diagonal of $V$. Note that $X\sim N(\operatorname{diag}^{-1/2} (\mathbf{v})\theta, \Omega)$, where $\Omega = \operatorname{diag}^{-1/2} (\mathbf{v}) V \operatorname{diag}^{-1/2} (\mathbf{v})$ is the correlation matrix of $\hat \theta$.
We assume that the set of significant effects, denoted as $\hat{S} \subseteq \{1, \dotsc, m\}$, is determined by a step-down or step-up procedure that sequentially tests the significance of the estimated effects. A treatment effect $h$ is deemed significant and included in $\hat{S}$ if $\lvert X_h\rvert \geq \bar{x}_h$ (two-sided testing) or $X_h \geq \bar{x}_h$ (one-sided testing). The effect-specific critical values $\bar{x}_h$ depend on the data. We fully define the procedures covered by our methods in Section (ref).
Our approach follows the polyhedral method of lee2016exact and inverts a distribution function that conditions on significance and an additional nuisance parameter.
For a set $S$ of significant effects and a significant effect $s \in S$, we decompose $X$ as
where $Z^{(s)} = X - \Omega_{\bullet, s} X_s$ and $\Omega_{\bullet, s}$ is the $s$-th column of $\Omega$. By a property of the Gaussian distribution, $Z^{(s)}$ is independent of $X_s$, implying that the distribution of $X_s$ given $Z^{(s)}$ depends only on $\theta_s$ and not on the other elements of $\theta$.
A selection event can be represented as a subset of the support of the $t$-statistics $X$. Specifically, we derive the set $\mathcal{X} (S) \subset \mathbb{R}^m$ such that $\hat{S} = S$ if and only if $X \in \mathcal{X} (S)$. As we show below, $\mathcal{X}(S)$ has a complex geometry and can be costly to compute except when $m$ is very small. We circumvent direct evaluation of $\mathcal{X}(S)$ by conditioning on $Z^{(s)} = z$. In particular, we observe the significant effects $S$ if $X_s$ is contained in the (marginal) conditional support $\mathcal{X}_s (z, S)$, where
Below, we develop an efficient algorithm that computes $\mathcal{X}_s(z, S)$ in polynomial time.
Let $F_s(x_s \mid z, \theta_s, S)$ denote the distribution function of $X_s$ conditional on $Z^{(s)} = z$, and $\hat{S} = S$, given by $F_s(x_s \mid z, \theta_s, S) = F_s(x_s \mid z, \theta_s, V, \bar{x}, S)$, where
and $\Phi$ is the standard normal distribution function. This distribution is a truncated Gaussian distribution. Evaluating $F_s$ at the observed $x_s = X_s$, $z = Z^{(s)}$ and $S = \hat{S}$ gives the rosenblatt1952remarks transformation with the key property that
where $\mathcal{U}[0, 1]$ denotes the uniform distribution on the unit interval.\footnote{We formally derive this property in Lemma (ref) in Online Appendix (ref).}
For $p \in (0, 1)$, let $\tilde{\theta}^{(p)}_s$ be defined as the unique solution to
Uniqueness of the solution is established in Lemma (ref) in Online Appendix (ref). Our conditionally median-unbiased estimator is given by $\tilde{\theta}^{\text{ub}}_s = \tilde{\theta}^{(0.5)}_s$. Our conditional confidence interval with nominal coverage $1 - \alpha$ is given by $\left[ \tilde{\theta}^{(1 - \alpha/2)}_s, \tilde{\theta}^{(\alpha/2)}_s \right]$.
The conditionally median-unbiased estimator over- and underestimates the true effect $\theta_s$ with equal probability conditional on the selection event, as established in the following result.
The next result establishes that the conditional confidence interval covers the true effect $\theta_s$ with probability $1 - \alpha$ conditional on the selection event.
The conventional unconditional equal-tailed confidence interval for $\theta_s$, denoted $\text{CI}_\alpha(\theta_s)$, is given by $\hat{\theta}_s \pm \sqrt{\mathbf{v}_s} \Phi^{-1} \left(1 - \alpha/2 \right)$, where $\Phi^{-1}$ is the standard normal quantile function. Other selective confidence intervals modify the unconditional interval symmetrically by adjusting its width benjamini2005false,berk_valid_2013. In contrast, our adjustment is asymmetric, allowing the confidence interval to reflect selection bias (see, e.g., Table (ref) below).
We can use Bonferroni-correction to construct a joint conditional confidence set for all significant treatment effects. The previous theorem guarantees that
with probability $1 - \alpha$ conditional on $\hat{S} = S$, where $\bigtimes$ denotes the Cartesian product.
We now illustrate the selective inference problem and our solution in a simple setting with a single treatment effect ($m = 1$).
Suppose that the significance of $\hat{\theta}_1$ is established by a $t$-test, that is, $\hat{S} = \{1\}$ if $X_1 \in \mathcal{X}(\{1\})$, where $\mathcal{X}(\{1\}) = \{x \in \mathbb{R}: \abs{x} \geq \Phi^{-1} (1 - \beta/2)\}$ for two-sided testing, and $\mathcal{X}(\{1\}) = \{x \in \mathbb{R}: x \geq \Phi^{-1} (1 - \beta)\}$ for one-sided testing. Here, $\beta$ denotes the nominal size of the test.
Figure (ref) shows that the distribution of $X_1$ is a truncated Gaussian and typically asymmetric. This asymmetry implies that $\hat{\theta}_1$ is conditionally biased for the true effect size $\theta_1$.
In the two-sided case, the median-unbiased estimator $\tilde{\theta}_1^{\text{ub}}$ that solves $F_1(X_1 \mid \tilde{\theta}_1^{\text{ub}}, \{1\}) = 0.5$ will always correct the conventional estimator $\hat{\theta}_1$ toward zero. This aligns with the intuition that significant effects are positively selected (“winner's curse”), and that their magnitudes are overestimated.
To verify this analytically, let $X_1 > \bar{x}_1 = \Phi^{-1}(1 - \beta/2)$ and note that
Since $F_1(x \mid \theta_1, \{1\})$ is decreasing in $\theta_1$, $\tilde{\theta}_1^{\text{ub}}$ is smaller than $\hat{\theta}_1$.
This section defines the multiple-testing procedures used to determine the set $\hat{S}$ of jointly significant effects and characterises the set $\mathcal{X}(S)$ of $t$-statistics that lead to the selection of $\hat{S} = S$. For exposition, we focus on step-down procedures in the main text. Step-up procedures are covered in Online Appendix (ref).
The most commonly used step-down procedures control the family-wise error rate (FWER), that is, the probability of including at least one false positive in $\hat{S}$. For two-sided testing, FWER-control at level $\beta$ means that
The corresponding definition for one-sided testing is obtained by replacing $\{h: \theta_h = 0\}$ with $\{h: \theta_h \leq 0\}$ in the previous display.
A step-down rule proceeds by successively comparing the ordered test statistics against a sequence of thresholds. Let $(>, j)$ denote the (data-dependent) index of the test statistic with the $j$th-largest magnitude, that is, $\lvert X_{(>, 1)} \rvert \geq \lvert X_{(>, 2)} \rvert \geq \dotsm \geq \lvert X_{(>, m)} \rvert$. The thresholds are determined by a threshold function $\bar{x}$ that maps any $A \subseteq \{1, \dotsc, m\}$ to a positive number such that, for $A, B \subseteq \{1, \dotsc, m\}$ and $A \subseteq B$, $\bar{x} (A) \leq \bar{x} (B)$.
For some rules, the threshold function can be written as $\bar{x} (A) = \Phi^{-1} \left(1 - \beta_{\lvert A \rvert} / 2 \right)$. For example, under the Bonferroni procedure we have $\beta_j = \beta/m$, and under the Holm procedure $\beta_j = \beta/(m + 1 - j)$, where $\beta$ is the nominal FWER.\footnote{Additional examples are provided in Table (ref) of the Online Appendix.} For other rules, such as romano2005stepwise, $\bar{x}(A)$ depends on features of $A$ beyond its cardinality. This allows the procedure to exploit correlations among the test statistics, thereby increasing power.
A step-down procedure proceeds as follows:
{\singlespacing
} We denote the values of the test statistics $X$ for which Algorithm (ref) selects significant effects $S$ by $\mathcal{X}(S) \subseteq \mathbb{R}^m$.
For the case of two effects, Figure (ref) illustrates the region $\mathcal{X} (S)$ for $S = \{1\}$ (only first effect is significant), $S = \{2\}$ (only second effect is significant) and $S = \{1, 2\}$ (both effects are significant).
We introduce additional notation to prepare for a general characterisation of $\mathcal{X}(S)$. For $S \subseteq \{1, \dotsc, m\}$, let $\mathcal{E}(S)$ denote the class of bijections from $S$ onto $\{1, \dotsc, \lvert S \rvert\}$. For $\sigma \in \mathcal{E}(S)$, $\sigma(h)$ maps effect $h$ to its “rank” under the ordering $\sigma$. For example, the pre-image $\sigma^{-1} (\{1, 2\})$ gives the treatment effects with ranks $1$ and $2$. $S^{\mathsf{c}}$ is the complement of $S$ in $\{1, \dotsc, m\} $. The following lemma characterises $\mathcal{X}(S)$.
We use characterisation (ref) of $\mathcal{X}(S)$ in our theoretical analysis. Due to computational complexity, we do not directly evaluate $\mathcal{X}(S)$ when implementing our procedures. Our solution is discussed in the next section.
This section describes an efficient algorithm for computing the conditional support $\mathcal{X}_s(z, S)$ without evaluating $\mathcal{X}(S)$. Given that $\mathcal{X} (S)$ is a complex set, defined by iterating over $2^m$ permutations, and potentially non-convex set (see Figure (ref)), the existence of such an algorithm is not obvious.
Our algorithm has only polynomial complexity and evaluates $\mathcal{X}_s(z, S)$ in $\mathcal{O}(m^3 \log m)$ time, enabling applications with many effects, such as $m=371$ in our mutual fund study. In Online Appendix (ref), we provide an improved algorithm based on the bentley1979algorithms algorithm that reduces the computational cost to $\mathcal{O}(\tilde{n}_I \lvert S \rvert \log \lvert S \rvert)$, where $\lvert S \rvert$ is the number of significant effects and $\tilde{n}_I < S$ is a number that depends on the data.
For exposition, we present a simple version of the algorithm for two-sided step-down rules and the selection event $\hat{S} = S$. Improvements and generalisations of the algorithm are described in Online Appendix (ref). The algorithm approximates $\mathcal{X}_s(z, S)$ by its closure. The approximation error has Lebesgue measure zero and does not affect any of the statistics that we consider.
We first introduce some notation. Let $\mathbf{x}_z (x_s) = \Omega_{\bullet, s} x_s + z$ denote the vector-valued function of $x_s$ that gives the $t$-statistics if $X_s = x_s$ and $Z^{(s)} = z$. For $\sigma \in \mathcal{E}(S)$, let $\bar{x}_{\sigma, h} = \bar{x} \left(\sigma^{-1}(\{\sigma (h), \dotsc, \lvert S \rvert \}) \cup S^{\mathsf{c}} \right)$ for $h \in S$ and $\bar{x}_{\sigma, h} = \bar{x} \left( S^{\mathsf{c}}\right)$ for $h \notin S$.
Our algorithm is based on intervals $I$ over which the linear functions $\mathbf{x}_{z,h}(x_s)$, $h=1, \dotsc, m$, can be unambiguously ordered --- that is, they neither intersect nor change sign. These intervals are obtained by partitioning the real line at points where a function $\mathbf{x}_{z,h}(x_s)$ crosses the horizontal axis or where two functions $\mathbf{x}_{z,h}(x_s)$ and $\mathbf{x}_{z,h'}(x_s)$ intersect. In the online appendix, we describe refinements that avoid computing all such intersection points. The key observation underlying our algorithm is that the ordering of the $t$-statistics $\mathbf{x}_z(x_s)$ remains constant within each interval $I$. By Lemma (ref), whether $x_s \in \mathcal{X}_s(z, S)$ depends only on the permutation that reorders the components of $\mathbf{x}_z(x_s)$ by their absolute values. Consequently, the conditional support can be computed without enumerating all permutations in $\mathcal{E}(S)$. The number of permutations that must be considered is bounded by the number of intervals $I$, which is quadratic.
Our algorithm proceeds as follows:
{\singlespacing
}
The following result establishes the validity and efficiency of our algorithm.
An immediate corollary to Theorem (ref) is that the conditional support is a union of intervals. By merging overlapping intervals, this union can be written as a union of disjoint intervals $[u_k, \ell_k]$, $k = 1, \dotsc, n_I$ with $n_I \leq m(m+1)/2 + 1$. In practice, the number of disjoint intervals $n_I$ is typically very small. In Online Appendix (ref), we characterise scenarios in which $n_I = 1$ (i.e., a connected interval).
In this section, we examine when conditional inference coincides with conventional methods and how to adapt it to estimated variance–covariance matrices or threshold functions. We also discuss which selection events provide an appropriate summary of “observed significance.”
In certain situations, conditional and unconditional inference yield similar results. Very roughly speaking, this happens when the effects are “highly significant.” In a univariate setting, this means that the $t$-statistic of the effect of interest is large. In a multivariate setting, significance of an effect depends also on the $t$-statistics of other effects and a more careful examination of what “highly significant” means is needed.
The following theorem gives sufficient conditions for the convergence of the conditional confidence interval to the unconditional confidence interval. It holds for selection events $\hat{S} = S$ with a positive, but not necessarily high, probability in the limit.\footnote{In a different context concerning conditional inference, andrews_inference_2024 demonstrate that their conditional confidence intervals converge to unconditional confidence intervals when the selection event's probability approaches one (see Proposition 3 in their paper). Our framework allows for a similar conclusion to be drawn as a corollary to Theorem (ref).}
For a setting with (asymptotically) uncorrelated estimators, such as subgroup analysis, Theorem (ref) implies that conditional and unconditional confidence interval for a significant effect $s$ converge if we expect the $t$-statistic $X_s$ to be large.
In settings with correlated estimators, Theorem (ref) imposes more stringent conditions. Suppose that we conduct inference on $s$ conditional on $\hat{S} = S$ and that the effect $h \neq s$ is correlated with $s$. The theorem requires that the significance status of $h$ becomes asymptotically certain: if $h \in S$ then $h \in \hat{S}$ with probability approaching one and if $h \notin S$ then $h \notin \hat{S}$ with probability approaching one. For two-sided testing this rules out $h \notin S$ since the rejection probability of a two-sided test is bounded away from zero regardless of the true effect size.
A straightforward corollary to Theorem (ref) is that the conditional median-unbiased estimator converges to the unconditional point estimator under the same conditions.
In this section, we accommodate unknown variance-covariance matrices and data-dependent threshold functions. This requires two modifications. First, we replace the test statistics $X$ by studentised $t$-statistics based on a consistent estimator $\hat{V}$ of the covariance matrix. Second, we replace the fixed threshold function $\bar{x}$ by a data-dependent threshold function $\hat{\bar{x}}$. The latter modification is important because data-dependent threshold functions, such as those in romano2005stepwise,list2019multiple, yield more adaptive and potentially more powerful tests.
The studentised $t$-statistics are defined as $\hat{X} = \operatorname{diag}^{-1/2} (\hat{\mathbf{v}}) \hat{\theta}$, where $\hat{\mathbf{v}}$ is the diagonal of $\hat{V}$. Our inference is based on an estimate of the true distribution function $F_s(x_s \mid z, \theta_s, S)$ given by
where $\hat{Z}^{(s)} = \hat{X} - \hat{\Omega}_{\bullet, s} \hat{X}_s$ and $ \hat{\Omega} = \operatorname{diag}^{-1/2}(\hat{\mathbf{v}}) \hat{V} \operatorname{diag}^{-1/2}(\hat{\mathbf{v}}) $. We now define our conditional inference procedures as in Section (ref), with $\widehat{F}_s$ replacing $F_s$. Let $\hat{\theta}_s^{(p)}$ denote the unique solution to
The estimator $\hat{\theta}_s^{\text{ub}}$ and the confidence interval $\widehat{\text{CCI}}_{\alpha}(\theta_s \mid S)$ are now defined analogously to $\tilde{\theta}_s^{\text{ub}}$ and $\text{CCI}_{\alpha}(\theta_s \mid S)$ above, with $\tilde{\theta}_s^{(p)}$ replaced by $\hat{\theta}_s^{(p)}$. The theoretical analysis of the modified procedure is given in Section (ref).
What event describes “observed significance” in a multivariate setting is not obvious. We now motivate our focus on the event $\hat{S} = S$ and discuss alternative selection events that may be relevant in some applications.
By conditioning on $\hat{S} = S$, we fully take into account the outcome of the multiple testing procedure, including both which effects are significant and which are not.
This selection event is appropriate if the interpretation and framing of a significant effect $s \in S$ depends on whether other effects $h \neq s$ are found significant or insignificant, as is often the case in empirical studies.
For example, in Section (ref) below, we partially replicate the study of karlan2007does on the effects of a matching grant on charitable giving. An effect on the response rate captures whether the matching grant motivates more people to donate, whereas effects on the donation amount capture also the size of donations. If the response rate is insignificant, we may interpret the additional donations as coming from the existing donor pool. Other examples where insignificant effects shape the interpretation of empirical results include placebo tests (see our application in Online Appendix (ref)).
In some applications, conditioning on $\hat{S} \supseteq S $ is more suitable. This choice removes information about insignificance from the selection event. This is desirable, for example, in our mutual fund application. Funds with significantly positive alphas are funds that systematically outperform the market. The reporting on outperforming funds is unlikely to change when finding insignificant alphas for other funds. The event $\hat{S} \supseteq S$ accounts for the funds in $S$ to be classified as outperformers, but it is agnostic about the performance of the other funds. Even though $\hat{S} \supseteq S$ may appear to be an even more complex event than $\hat{S} = S$, our polynomial-time algorithm readily extends to this case (see Online Appendix (ref)).
One may also consider the selection event $s \in \hat{S}$ for a given significant effect $s \in S$. This is the most direct analogue of conditioning on $\hat{S} = \{1\}$ in the univariate case. However, this defines different selection events for different significant effects, making it difficult to conduct joint conditional inference on all significant effects. We mention this event for completeness but do not pursue it further in this paper.
We now provide asymptotic counterparts to Theorem (ref) and Theorem (ref) with estimated covariance matrices, data-dependent thresholds and asymptotically Gaussian estimators of the treatment effects. We allow all population quantities to depend on the probability measure and prove asymptotic results that are uniform over sequences of probability measures.
Let $\lVert \cdot \rVert$ denote the $L_2$-norm. For a sequence of classes of probability measures $\mathcal{P}_n$, assume:
Assumption (ref).(ref) weakens the assumption that the treatment effect estimators are exactly Gaussian to uniform asymptotic Gaussianity kasy2019uniformity. Assumptions (ref).(ref) and (ref).(ref) require the estimated variance-covariance matrix and the data-dependent threshold function to be uniformly consistent for their respective population counterparts.
The following theorem is an asymptotic counterpart to Theorem (ref).
This result establishes that $\hat{\theta}^{\text{ub}}_s$ is asymptotically median-unbiased for $\theta_s$ uniformly over $\mathcal{P}_n$ with the possible exception of sequences $P_n \in \mathcal{P}_n$ along which $P_n (\hat{S} = S)$ vanishes.
The uniformity of our result is important because it covers local-to-zero signal-to-noise ratios, where effective treatments are not always detected in the limit. In such cases, selectivity cannot be ignored, and conventional confidence intervals are invalid. Under conventional “pointwise” asymptotics, by contrast, effective treatments have diverging $t$-statistics and are therefore highly “significant” in the limit. This limit is a poor approximation of typical finite-sample settings, where $t$-statistics are often quite small even for effective treatments.
The statement of Theorem (ref) holds trivially when $P(\hat S=S)=0$. By focusing on cases where $ P(\hat S = S)$ is positive, we ensure that we do not condition on an event with zero probability in the asymptotic limit. Similar assumptions are made in other applications of conditional inference andrews_inference_2024.
The next result establishes that the conditional confidence interval $\widehat{\text{CCI}}_\alpha(\theta_s \mid S)$ has an asymptotic conditional coverage probability of $1-\alpha$.
This result immediately implies the unconditional validity of our conditional confidence interval. For example, for the joint confidence set (ref), it holds that
We use Monte Carlo simulations to assess the finite-sample performance of our conditional confidence interval and conditional median-unbiased estimator.
We consider two designs: one with normally distributed data (“Normal”) and one with chi-squared distributed data (“Chi-squared”). In both, we test $m = 5$ parameters with population values $\theta = (0.05, 0.03, 0.01, 0, 0)'$ and focus on inference for the first parameter when it is significant. Significance is determined using the Holm step-down procedure at a FWER of $\beta = 10\%$, based on $t$-tests computed from $n$ i.i.d. observations of an $m$-vector $Y_i$. The test statistic is $X_k = \sqrt{n} \bar{Y}_k / S_k$, where $\bar{Y}_k$ and $S_k^2$ denote the sample mean and variance. Components of $Y_i$ have pairwise correlation $\rho = 0.5$. The two designs differ only in the marginal distribution of $Y_i$.
In the “Normal” design, $Y_i$ follows a multivariate normal distribution with mean $\theta$ and a variance–covariance matrix with ones on the diagonal and $\rho$ off-diagonal. In the “Chi-squared” design, $Y_{ik} = (U_{ik}^2 - 1)/\sqrt{2} + \theta_k$, where $U_i$ is multivariate normal with mean zero and a variance–covariance matrix with ones on the diagonal and $\sqrt{\rho}$ off-diagonal, yielding correlation $\rho$ between $Y_{ik}$ and $Y_{ik'}$ for $k \neq k'$.
We consider $n \in \{100, 300, 500, 700, 900\}$. Each design uses 5{,}000 replications, increased to 20{,}000 for $n = 100$ to ensure enough cases in which the first parameter is selected. We compute the conditional, unconditional (“naive”), and Bonferroni-corrected confidence intervals, each with nominal coverage $(1 - \alpha) = 90\%$.
We simulate confidence interval coverage and estimator bias conditional on the first parameter being significant, that is, on the event $\{1\} \subseteq \hat{S}$. Table (ref) reports results for two-sided tests; additional results for one-sided tests appear in Online Appendix (ref). Selection probabilities---the chances that the first parameter is found significant---range from 2% to 21%.
We first examine coverage. The conditional intervals consistently achieve coverage close to the nominal 90% level. In the “Chi-squared” design, we observe light undercoverage at the smallest sample size ($n=100$), suggesting that, in this setting, somewhat larger samples are required for the asymptotic approximation to be fully accurate. For sample sizes of 300 or more, empirical coverage is satisfactory even with non-normal data. By contrast, naive confidence intervals undercover across all designs and sample sizes, particularly when selection is infrequent. Bonferroni intervals are robust to selection but reach nominal coverage only at large $n$ and tend to be overly conservative. In terms of precision, the conditional intervals are wider than the naive ones yet substantially shorter than the Bonferroni intervals, highlighting their efficiency gains.
Beyond providing valid confidence intervals, our approach also mitigates selection bias. The conditional bias of the conventional estimator ($0.084$ in the “Normal” design at $n=100$) is reduced to $-0.015$ by our conditional median-unbiased estimator. A possible explanation for the Bonferroni interval's conditional undercoverage is that it is centred on the conventional estimator, which exhibits substantially larger conditional bias (e.g., $0.084$ in the “Normal” design at $n=100$).
We illustrate our selective conditional inference methods by revisiting the field experiment conducted by karlan2007does, which studies the effectiveness of matching grants on charitable giving. The experiment is also reanalysed by list2019multiple using step-down multiple testing. We add to their analysis by inferring the effect sizes of the significant results. We focus on the binary treatment whether a potential donor was offered a matching grant and ignore heterogeneous characteristics of the matching grants.\footnote{See Section 4 of list2019multiple for details on the different grants.}
The details of the estimation procedures are given in Online Appendix (ref). To detect significant effects, we use three FWER-controlling procedures: the Bonferroni correction, the Holm procedure, and a bootstrap procedure based on romano2005stepwise. Our implementation of the RW2005 procedure is described in Online Appendix (ref). We set the nominal FWER to 10%. All reported confidence intervals use nominal coverage levels of 95%.
We study treatment effects corresponding to four different outcomes: the response rate, dollars given not including match, dollars given including match, and the change in the amount donated from the previous elicitation.\footnote{In Appendix (ref), we consider additional specifications that also comprise multiple subgroups and treatments.} For comparability, we measure effect size for all outcomes in terms of standard deviations.\footnote{For the response rate, this standard deviation is defined above and equal to 13%. For the other outcomes, we use $\$8.18$, the standard deviation of the amount donated in the control group.} Given that the signs of the effects are not obvious, we use two-sided testing.\footnote{For example, a matching grant may increase the amount the charity receives but decrease individual donations due to crowding out.}
The outcomes “Response rate” and “Dollars given incl match” are found to be significant by all three procedures. The RW2005 procedure detects an additional significant effect for “Dollars given not incl match.”
Our conditional inference results are summarised in Table (ref).
Under the selection event $\hat{S} = S$, our conditional inference post RW2005-selection carries out minor downward correction, consistent with positively selected significant effects. Inference after Holm and Bonferroni detection is more subtle, indicating a negative bias for the effect on “Dollars given incl match.” This outcome is positively correlated with the effect on “Dollars given not incl match,” which is insignificant. Test statistics and therefore significance depend on the true effect sizes and intrinsic randomness. Since the two outcomes are correlated, they share some of their randomness. To reconcile observing one effect as significant and the other as insignificant, the true effect for the significant effect should be larger than its unconditional estimate. The upward correction for the effect on “Dollars given incl match” is more pronounced for the Holm procedure than the Bonferroni procedure. This is because “Dollars given not incl match” is closer to the threshold of significance under the Holm procedure than under the Bonferroni procedure.\footnote{The Bonferroni $p$-value for “Dollars given not incl match” is 20%, the Holm $p$-value is just above 10%.}
The correlation structure among the outcomes is crucial for explaining the upward correction. For uncorrelated treatment effect estimators, the bias correction will always be downward, consistent with the positive “winner's curse” bias in the univariate case (see our results for multiple subgroups in Online Appendix (ref)).
In contrast to the selection event $\hat{S} = S$, the selection event $\hat{S} \supseteq S$ does not include information about insignificant effects. This information is crucial for the explanation of a possible upward correction. The results in Table (ref) demonstrate that all bias corrections are downward when we condition on $\hat{S} \supseteq S$. The downward corrections behave differently from the corresponding correction in a univariate setting where a more powerful test is associated with a smaller positive bias. In our multivariate setting, we observe a steeper downward correction for the Holm procedure than for the Bonferroni procedure, even though the former is more powerful. This result arises because the Holm procedure provides a threshold closer to the $t$-statistic for “Dollars given not including match” than the Bonferroni procedure.
Our second application uses our method to infer the performance of mutual funds identified as market outperformers through multiple hypothesis testing among hundreds of funds, illustrating the computational scalability of our approach.
The multiple testing challenge of detecting outperforming funds among numerous candidates has been examined by giglio2021thousands, and similar issues arise in other areas of finance harvey2016and, harvey2020evaluation, heath2023reusing.
We examine fixed-income mutual funds using the CRSP Survivor-Bias-Free U.S. Mutual Fund Database available through Wharton Research Data Services (WRDS). Our sample covers the period from January 2000 to April 2024. After excluding funds with missing data, we obtain monthly returns for 371 funds.
A mutual fund is said to outperform the market if it exhibits a significantly positive “alpha,” representing excess returns unexplained by exposure to systematic risk factors. Each fund's alpha is estimated using the Fama-French five-factor model fama2015five,
where $R_{i,t}$ denotes the fund's excess return over the risk-free rate, proxied by the beginning-of-month 30-day T-bill yield. The factors include the market excess return ($R_{m,t}$) and four zero-investment portfolios capturing size ($R_{smb,t}$), value ($R_{hml,t}$), profitability ($R_{rmw,t}$), and investment ($R_{cma,t}$) effects.
This model estimates 371 alphas and $371 \times 5$ factor loadings. To address the high dimensionality of the resulting covariance matrix, we employ the principal orthogonal complement thresholding (POET) method of fan2013large. The POET estimator is obtained by thresholding the remaining components of the sample covariance matrix (the principal orthogonal complement) after removing the first $K$ principal components. For our dataset, $K=1$ is selected using the criterion of bai2002determining, and we use the default soft-thresholding implementation in the R package POET Rpoet2016. This covariance matrix is used to compute $t$-statistics and subsequently for implementing our conditional inference procedures.
Following giglio2021thousands, we apply multiple hypothesis testing to identify funds with significantly positive alphas, interpreted as market outperformers. We employ the Holm (FWER control) and benjamini2001control (FDR control) procedures, setting the nominal FWER and FDR levels to 10%. Both procedures detect five funds with significantly positive alphas (see Table (ref)).
Having identified five significant funds, we now use our methods to quantify the extent to which they outperform the market. Because these few outperformers are found significant among a large pool of 371 funds, their estimated alphas are likely to be positively selected. Our procedures explicitly account for this issue and provide unbiased inference on the outperformers' alphas. Importantly, our algorithm can handle the complex selection event that arises from testing all 371 alphas jointly.
We condition on the event $\hat{S} \supseteq S$, that is, on finding the five outperformers but not on finding the remaining funds insignificant. This choice is natural because, as discussed in Section (ref), for which funds we find insignificant alphas does not affect our interpretation of the outperformers.
Table (ref) reports conditionally median-unbiased alphas and conditional confidence intervals for the outperforming funds, alongside conventional (unconditional) estimates and intervals for comparison. The nominal coverage level for all confidence intervals is set to $1 - \alpha = 0.9$. Conventional intervals are Bonferroni-corrected for the total number of funds ($m=371$), whereas conditional intervals are Bonferroni-corrected for the number of selected funds ($\lvert \hat{S} \rvert = 5$).
The median-unbiased estimates are only slightly smaller than their unconditional counterparts, suggesting that positive selection bias is of limited practical importance in this application. Conditional confidence intervals are generally narrower than their unconditional analogues, with length reductions of roughly 12-35% under the Holm-based procedure and 21-36% under the BY-based procedure, reflecting the reduced Bonferroni correction under conditional inference.
Testing alphas against zero identifies statistical rather than economic significance. We assess economic significance based on the lower bounds of the joint conditional confidence intervals, which exceed 0.135 for four of the five outperforming funds, indicating economically meaningful outperformance.
In this paper, we have proposed new methods for conducting inference on the magnitude of significant effects detected through multiple hypothesis testing. Our methods support a wide range of multiple testing procedures, scale to scenarios involving many estimated effects, and accommodate different selection events that operationalise the notion of “observed significance.”
While our primary focus is on inference after multiple hypothesis testing, our procedure can also be applied to conduct inference conditional on pre-test and placebo results. For instance, a pre-test is considered passed if the effects of one or more placebo outcomes are found to be insignificant. This type of test is commonly used, for example, when assessing parallel trends in a difference-in-differences design roth2022pretest. In the online appendix, we provide another example where property crime serves as a placebo outcome to test an empirical specification for estimating an effect on violent crime.
Our approach follows the multiple testing literature in assuming that the full set of tested hypotheses is known. Addressing situations where insignificant results are hidden requires a different framework andrews2019identification, berk_valid_2013. As shown by sarfati2025post, limiting attention to a pre-specified set of possible hypotheses is essential for valid conditional inference. This requirement implies that the framework cannot accommodate unrestricted data mining.
Finally, our numerical examples demonstrate the strong power properties of our procedure, and it would be worthwhile to explore these aspects further from a theoretical perspective. Our empirical illustrations in Sections (ref) and (ref) suggest that our procedure can produce informative confidence intervals. Additionally, the simulation results presented in Section (ref) confirm its strong power characteristics. It would be interesting to investigate whether the optimal selective inference discussed by fithian_optimal_2017 can be applied in this context.
We are grateful for comments from Erik Hjalmarsson, Mikael Lindahl and participants at CFE-CMStatistics 2022, the 4th Aarhus Workshop in Econometrics, IPDC2025 in Montpellier, and Econometrics Forum 2025 in Awara. Dzemski acknowledges financial support from Jan Wallanders och Tom Hedelius stiftelse samt Tore Browaldhs stiftelse under grant P24-0135. Okui acknowledges financial support from the Japan Society for the Promotion of Science under KAKENHI Grant Nos.23K25501. Wang acknowledges financial support from the Singapore Ministry of Education Tier 1 grants RG104/21, RG51/24, and NTU CoHASS Research Support Grant. The authors made use of the AI tools Grammarly and ChatGPT for language editing. All AI-generated suggestions were reviewed and verified by the authors, who take full responsibility for the content of the article. All errors remain our own.
\printbibliography
\singlespacing